Abstract Paper Portal of Conference on Neural Information Processing Systems (NeurIPS) 2026
Abstract: Diffusion models achieve state-of-the-art image synthesis, with their generative trajectories fundamentally exhibit a spectral bias, resolving low-frequency global structures early and high-frequency fine details later. Conventional stochastic differential equation (SDE) solvers fail to account for this dynamic, naively injecting uniform white noise throughout the entire process and misusing the finite energy budget. In this work, we establish a mathematical framework that reconsiders SDE inference as a targeted, frequency-decoupled energy transfer. Leveraging this framework, we introduce Colored Noise Sampling (CNS), a novel, training-free stochastic solver. Rather than injecting uniform white noise, CNS utilizes a dynamic, timestep- and frequency-dependent schedule that more-efficiently allocates injected energy toward structurally unresolved frequency bands. By actively exploiting the model's inherent spectral bias, CNS systematically steers the generated distribution toward the true data manifold. Extensive experiments demonstrate that CNS significantly outperforms standard ODE and SDE baselines as a strictly plug-and-play, inference-time sampler substitution across diverse architectures (SiT, JiT, FLUX). Compared to standard sampling on ImageNet-256, CNS achieves substantial unguided FID reductions, improving from 8.26 to 6.27 on SiT-XL/2, 32.39 to 26.69 on JiT-B/16, and 11.88 to 8.31 on JiT-H/16, while yielding consistent relative FID improvements with Classifier-Free Guidance.
Abstract: Generative modeling can be formulated as learning a mapping f such that its pushforward distribution matches the data distribution. The pushforward behavior can be carried out iteratively at inference time, for example, in diffusion/flow-based models. In this paper, we propose Drifting Models, a generative modeling framework that evolves the pushforward distribution during training and naturally admits one-step inference. We introduce a drifting field that governs the sample movement and achieves equilibrium when the distributions match. This leads to a training objective that allows the neural network optimizer to evolve the distribution. In experiments, our one-step generator achieves state-of-the-art results on ImageNet 256×256, with FID 1.54 in latent space and 1.61 in pixel space. We hope that our work opens up new opportunities for high-quality one-step generation.
Abstract: Recent AI regulations increasingly emphasize the need for mechanisms that preserve the utility of data for AI innovation while preventing misuse, particularly by enforcing purpose limitation in downstream AI applications. In practice, enforcing this principle remains challenging, as released data can be trivially fed into arbitrary models beyond its declared intent. Existing approaches attempt to mitigate this risk by either perturbing data or retraining models to limit unintended use. These strategies, however, offer no protection against inference by unknown or externally trained models, or fundamentally rely on control over the training or deployment. In this work, we introduce non-transferable examples (NTEs), recoded data that act as a task-level "ciphertext" decodable only by a designated model. Where adversarial examples exploit sensitive input directions, NTEs use the complementary insensitive subspace: a training-free, data-agnostic recoding within a model-specific low-sensitivity subspace preserves the authorized model's outputs while degrading unauthorized ones through subspace misalignment. We establish formal bounds certifying output fidelity for the authorized model and showing unauthorized degradation scales with measurable spectral misalignment between models. Empirically, NTEs preserve authorized-model performance across diverse vision backbones and vision-language models, while unauthorized models collapse even under adaptive reconstruction attacks. These results establish NTEs as a practical means to preserve intended data utility while preventing unauthorized exploitation. Our source code and visual demos are available at: https://github.com/model-specific/non-transferable-examples.
Authors:
Yifei Dong, Fengyi Wu, Yilong Dai, Lingdong Kong, Guangyu Chen, Qiyu Hu, Yetong Sha, Feng Liu, Siyu Huang, Qi Dai, Zhi-Qi ChengAbstract: Goal-conditioned visual navigation has been a long-standing testbed for embodied AI. We study a natural language-conditioned variant, language-conditioned visual navigation (LCVN), in which an embodied agent must follow a natural language instruction given only an initial egocentric observation. Without access to goal images, the agent must rely on language to shape its perception and continuous control. We introduce the LCVN Dataset, a benchmark of 39,016 trajectories and 117,048 human-verified instructions spanning diverse environments and instruction styles. Building on this benchmark, we study two complementary paradigms: (i) latent-imagination policy learning, in which a diffusion-based world model (LCVN-WM) imagines future observations and an actor-critic agent (LCVN-AC) learns its policy entirely within the imagined latent space; and (ii) unified autoregressive prediction, in which a single multimodal backbone (LCVN-Uni) jointly predicts actions and observations in one forward pass over a shared token sequence. Experiments show that two paradigms offer complementary strengths: latent imagination produces more temporally coherent rollouts, whereas unified prediction generalizes better to unseen environments. Targeted ablations further isolate the contributions of language guidance, conditioning signals, and instruction style, clarifying when language grounding versus dynamics modeling is the performance bottleneck. Together, these findings position LCVN as a concrete testbed for studying how language, imagination, and decision-making interact in embodied agents.
Abstract: We show that Fréchet Distance (FD), long considered impractical as a training objective, can in fact be effectively optimized in the representation space. Our idea is simple: decouple the population size for FD estimation (e.g., 50k) from the batch size for gradient computation (e.g., 1024). We term this approach FD-Loss. Optimizing FD-Loss reveals several surprising findings. First, post-training a base generator with FD-Loss in different representation spaces consistently improves visual quality. Under the Inception feature space, a one-step generator achieves 0.72 FID on ImageNet 256×256. Second, the same FD-Loss repurposes multi-step generators into strong one-step generators without teacher distillation, adversarial training or per-sample targets. Third, FID can misrank visual quality: modern representations can yield better samples despite worse Inception FID. This motivates FDr^k, a multi-representation metric. We hope this work will encourage further exploration of distributional distances in diverse representation spaces as both training objectives and evaluation metrics for generative models. Code and checkpoints will be open-sourced.
Abstract: Geometry is invariant to viewpoint, which makes any collection of images a redundant encoding of a single 3D state. Existing feed-forward reconstruction models fail to exploit this: per-view methods emit overlapping, unaligned pointmaps that grow linearly with input count, while global-latent methods commit to a fixed, low-resolution output. We introduce Surflo, which compresses a variable number of unposed RGB views into K latent tokens-one global state-and decodes oriented 3D surface points by independently transporting them from noise onto the surface via flow matching. This frees the output from any fixed grid or token budget: the same latent yields from a few thousand to a million points in a single forward pass. To suppress the local inconsistencies inherent to independent per-point decoding, an inference-time guidance term correlates nearby points by injecting a photometric gradient during ODE integration. Surflo matches or surpasses feed-forward baselines on surface metrics, run an order of magnitude faster than optimization-based methods that require hundreds of views, and is the only feed-forward approach to combine a global latent with arbitrary-resolution decoding.
Abstract: Reinforcement Learning (RL) post-training has become the standard for aligning generative models with human preferences, yet most methods rely on a single scalar reward. When multiple criteria matter, the prevailing practice of ``early scalarization'' collapses rewards into a fixed weighted sum. This commits the model to a single trade-off point at training time, providing no inference-time control over inherently conflicting goals -- such as prompt adherence versus source fidelity in image editing. We introduce ParetoSlider, a multi-objective RL (MORL) framework for Diffusion Models post-training that allows continuous reward control during inference. By training the model with continuously varying preference weights as a conditioning signal the model learns how its denoising trajectory should change as the requested reward priority changes. Consequently, we enable users to navigate optimal trade-offs at inference time without retraining or maintaining multiple checkpoints. We evaluate ParetoSlider across three state-of-the-art flow-matching backbones: SD3.5, FluxKontext, and LTX-2. Our single preference-conditioned model matches or exceeds the performance of baselines trained separately for fixed reward trade-offs, while uniquely providing fine-grained control over competing generative goals.
Abstract: Recent multimodal large language models (MLLMs) increasingly rely on visual chain-of-thought to perform region-grounded reasoning over images. However, existing approaches ground regions via either textified coordinates—causing modality mismatch and semantic fragmentation—or fixed-granularity patches that both limit precise region selection and often require non-trivial architectural changes. In this paper, we propose Numerical Visual Chain-of-Thought (NV-CoT), a framework that enables MLLMs to reason over images using continuous numerical coordinates. NV-CoT expands the MLLM action space from discrete vocabulary tokens to a continuous Euclidean space, allowing models to directly generate bounding-box coordinates as actions with only minimal architectural modification. The framework supports both supervised fine-tuning and reinforcement learning. In particular, we replace categorical token policies with a Gaussian (or Laplace) policy over coordinates and introduce stochasticity via reparameterized sampling, making NV-CoT fully compatible with GRPO-style policy optimization. Extensive experiments on three benchmarks against eight representative visual reasoning baselines demonstrate that NV-CoT significantly improves localization precision and final answer accuracy, while also accelerating training convergence, validating the effectiveness of continuous-action visual reasoning in MLLMs. The code is provided in the Supplementary Materials.
Abstract: Visual decoding from brain signals is a key challenge at the intersection of computer vision and neuroscience, requiring methods that bridge neural representations and computational models of vision. A field-wide goal is to achieve accurate, generalizable decoding from non-invasive, temporally resolved signals, including electroencephalography (EEG). A major obstacle towards this goal is the low signal-to-noise ratio of EEG and the substantial inter-subject variability, which render direct end-to-end EEG–image supervision weak and unstable. To address this, we introduce a tri-modal contrastive framework for EEG-based visual decoding that aligns EEG, visual, and textual representations within a unified latent space. Our approach follows a two-stage design. First, we pre-train an EEG encoder via masked reconstruction on unlabeled trials, learning spatio-temporal regularities that transfer robustly downstream. Second, we jointly align EEG, image, and LLM-generated textual descriptions through contrastive learning, where text supervision acts as a semantic regularizer that injects linguistic structure into the shared space without overwhelming the primary EEG–image signal. The encoder integrates subject-specific adaptation, graph-attention over channels, and temporal-spatial convolutional embeddings. On the Things-EEG2 200-way zero-shot benchmark, our framework achieves 54.1% Top-1 and 83.4% Top-5 accuracy, substantially exceeding the strongest prior baseline (32.4% / 64.0%), with paired Wilcoxon tests confirming significance (p < 0.01) over all in-subject baselines. We validate generalization on Things-MEG. Analysis reveals that compact embedding geometries (CN-CLIP) outperform much larger backbones, and that decoding aligns with established neurophysiology of visual processing. This work is a critical step towards robust, semantically-grounded visual decoding from non-invasive temporal neural signals.
Abstract: Natural-language instance navigation becomes challenging when the initial user request does not uniquely specify the target instance. A practical agent should reduce the user's burden by actively asking only the information needed to distinguish the intended target from similar distractors, rather than requiring a detailed description upfront. Existing approaches often fall short of this goal: they may stop at the first plausible candidate before sufficiently exploring alternatives, or, even after collecting multiple candidates, ask about the target's attributes derived from individual candidates rather than questions selected to distinguish candidates in the pool. As a result, despite the dialogue, the agent may still fail to distinguish the target from distractors, leading to premature decisions and lengthy user responses. We propose Proactive Instance Navigation with Comparative Judgment (ProCompNav), a two-stage framework that first constructs a candidate pool and then identifies the target through comparative judgment. At each round, ProCompNav extracts an attribute-value pair that splits the current pool, asks a binary yes/no question, and prunes all inconsistent candidates at once. This reframes disambiguation from open-ended target description to pool-level discriminative questioning, where each question is chosen to narrow the candidate set. On CoIN-Bench, ProCompNav improves Success Rate over interactive baselines with the same minimal input and non-interactive baselines with detailed descriptions, while substantially reducing Response Length. ProCompNav also achieves state-of-the-art Success Rate on TextNav, suggesting that comparative judgment is broadly useful for instance-level navigation among similar distractors.
PaperID: 11, Oral
Authors:
Sam McCallum, Zander Blasingame, Timothy Herschell, Niklas Rindtorff, Alexander Tong, James FosterAbstract: Flow and diffusion models generate high-quality samples in many modalities; however, many network evaluations are required during inference due to numerical integration of an underlying differential equation. Flow maps alleviate this problem by learning the solution map of the differential equation directly, enabling few-step sampling. Yet, current methods are restricted to approximating the solution map of ODEs. These methods can be used to learn the transition kernel of an SDE, thereby obtaining a solution map that recovers the marginal distributions of the process (weak convergence) rather than the solution path (strong convergence). We propose Strong Stochastic Flow Maps (SSFMs) as a novel framework for learning the strong solution map of additive-noise SDEs, directly generalizing deterministic flow maps to the stochastic setting. A polynomial approximation to Brownian motion is introduced and shown to converge pathwise. These results enable a simulation-free training objective for the solution map of diffusion models. We demonstrate that SSFMs outperform previous flow map methods on image generation and enable few-step sampling of molecular systems.
Abstract: Binary classification from positive-only samples is a variant of PAC learning in which the learner receives i.i.d. samples from the positive region of an unknown target concept, but is evaluated under the original distribution (which places mass on both positive and negative regions). This model dates back to (Natarajan, 1987, STOC), and the characterization of improper learning is well-known – it even appears in textbooks (Kearns and Vazirani, 1994, Exercise 3.7). The characterization of positive-only learning, however, has long remained open. In this work, we revisit and settle this question: a concept class is properly learnable from positive-only samples if and only if it has finite VC dimension and satisfies a new combinatorial condition, which we call uniform exterior separability. Together with several separation results, this characterization reveals a surprisingly rich landscape that differs sharply from standard PAC learning: proper and improper learning are separated, randomized and deterministic proper learning are separated, there are classes for which ERM is a learner, and finite VC dimension does not suffice even for non-uniform learning. Along the way, we introduce new combinatorial dimensions that we believe can be of broader interest in learning theory.
PaperID: 13, Oral
Abstract: Rotation averaging on the circle, equivalently robust angular synchronization, underlies a wide range of geometric estimation problems, for example, gravity-aligned Structure-from-Motion, multi-way point-cloud registration, and in-plane alignment in single-particle cryo-electron microscopy. We present the first algorithm that certifiably solves the robust maximum-consensus formulation of this problem to global optimality. Our approach discretises the circle into uniform bins, formulates a pairwise Markov random field and solves it via branch-and-bound with two key novelties: an ICM-guided label-pruning rule that collapses the branching factor to a handful of candidates, and a node-constrained relaxation bound orders of magnitude tighter than the standard per-edge bound. Together, they reduce the search from millions of nodes to a few hundred on typical instances. On all tested problems, the algorithm terminates with a proof of global optimality, typically in a few seconds. Across synthetic graphs and three real-world domains -- image-based SfM on IMC 2023/2024, multi-way LiDAR point-cloud registration on KITTI and NSS, and cryo-EM angular synchronization on EMPIAR-10166 -- the certified solution achieves sub-degree accuracy at outlier rates where other baselines fail. Our code will be made publicly available.
Abstract: We study Flow Matching in a semi-discrete setting where a Gaussian source is transported toward a discrete target supported on finitely many points. This semi-discrete regime is the theoretical setting behind the use of Flow Matching for generative modeling, where the target distribution is represented by a finite dataset. In this semi-discrete regime, the exact Flow Matching velocity field is available in closed form, which makes it possible to analyze the geometry induced by the terminal flow map independently of optimization and approximation effects. We investigate the terminal assignment regions, namely the preimages of the target atoms under the terminal flow. We show that these regions are open, simply connected and, under an additional assumption, homeomorphic to the unit ball. At the same time, a planar four-point example shows that these cells can differ sharply from Laguerre cells arising in semi-discrete optimal transport: they may be non-convex, have curved boundaries, and exhibit different boundedness and adjacency patterns. These results clarify the geometry intrinsically induced by the exact semi-discrete Flow Matching objective before neural approximation enters the picture.
PaperID: 15, Oral
Abstract: Current video-to-4D methods struggle with complex topology changes, transparent materials, thin structures, and inner surfaces. We present Trellis4D, a dynamic mesh generation framework by inheriting the expressive representation of Trellis2, adapting it from image-to-3D to video-conditioned 4D generation. Our design arises from two key questions: (a) how to enable Trellis2's frame-local attention to share information across frames while preserving its pretrained quality on rare cases such as transparent objects and inner surfaces, and (b) how to inject temporal information into a purely 3D positional encoding without breaking pretrained capabilities. We address (a) with a sliding-window cross-frame attention and anchor on the first frame. The first frame is generated by the base Trellis2 model and injected into our model, letting it inherit Trellis2's quality in rare cases through cross-frame attention. We address (b) with a 4D temporal encoding that repurposes redundant low-frequency spatial RoPE bands for time, extending the encoding from 3D with no additional parameters. Extensive experiments show the effectiveness of Trellis4D for high-quality dynamic mesh generation on ActionBench and our own challenging complex dynamics set.
Abstract: Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering the regime where compute is growing faster than high-quality data. We propose Compute-Data (CD) scaling laws, a unified framework that bridges compute-optimal scaling, in which data scales freely with compute, and data-optimal scaling, in which the corpus is fixed and compute can grow unbounded. CD scaling extends classic scaling by introducing a token effectiveness function that quantifies how much a derived token, produced for instance by multi-epoch repetition or paraphrasing, is worth relative to a fresh one, ranging from a perfect substitute to no value at all. Fitting \eta for two data expansion strategies (multi-epoch repetition, paraphrasing) across model sizes ranging from 14M to 600M parameters on the Dolma-3 corpus, we find that it is far from constant: it depends jointly on model size, the tokens-per-parameter ratio, and the amount of derived data, and saturates as the corpus is expanded. The functional form of \eta implies that substituting compute for data diminishes with both model size and data availability, and partitions training into three operational regimes: compute-bound, data-bound, and model-bound, showing that classic compute-optimal allocation is suboptimal across most of the practically relevant regime.
PaperID: 17, Oral
Authors: Luke J O'Connor
Abstract: At the heart of existing language model agents is a fixed orchestrator program which is responsible for the state transition between consecutive turns. This paper introduces \emphself-programmed execution (SPE), an agent architecture in which the model completion is itself the orchestrator program, and the harness evaluates this program but does not impose its own orchestration policy. I formalize this idea using agentic machines: an SPE state is one from which a model completion can load any state of an embedded copy of the machine, meaning that it is subject to no fixed turn-to-turn orchestration policy. Realizing SPE in practice is nontrivial because the same data is both model context and executable program. I therefore introduce \textscSpell, a Lisp-based language in which programs can edit and re-evaluate themselves, and effectful expressions like model invocations are structured such that re-evaluating an edited program does not replay its side effects. Experiments with existing models, not trained for SPE or \textscSpell, show that frontier models can operate in this regime and accomplish challenging agentic tasks. These results demonstrate how an LM can act as an agent without any fixed orchestration policy, and they raise the question of what self-orchestration strategies might be learned by a model trained for self-programmed execution.
Authors: Tim Joseph, Marcus Fechner, Philipp Stegmaier, Karam Daaboul, Marius Zöllner
Abstract: Intrinsic rewards for exploration in reinforcement learning condition on different contexts: lifelong rewards score each transition against accumulated experience but ignore within-rollout redundancy; episodic rewards penalize intra-trajectory repetition but discard lifetime progress. Hybrid methods combine both signals through heuristic weights or require Gaussian-process dynamics that do not scale beyond low-dimensional state spaces. Trajectory-level information gain decomposes into per-step terms that condition on the replay buffer and rollout prefix simultaneously, but remains intractable for deep models. We derive the Conditional Information Gain (CIG) reward as a tractable surrogate: a log-determinant objective over an ensemble disagreement kernel whose Cholesky factorization yields causal per-step rewards that retain both conditioning sets while scaling to high-dimensional state spaces. We instantiate CIG in a model-based setting, where rollouts are short and within-rollout corrections remain largely unexplored. Across twelve tasks spanning discrete (MiniGrid) and continuous control (OGBench), in both clean and stochastic-distractor settings, CIG outperforms or matches prior exploration methods while remaining robust to stochastic distractors.
PaperID: 19, Oral
Authors:
Vitus Benson, Fanny Yang, Markus ReichsteinAbstract: Earth observation foundation models produce global meter-resolution embeddings that underpin work across agriculture, ecology, hydrology, and beyond. Yet the embeddings themselves are opaque: each dimension entangles unrelated physical concepts, rendering their interpretability and robustness challenging for practitioners and scientists alike. Here, we propose sparse autoencoders (SAEs) as a principled way to decompose Earth embeddings into overcomplete dictionaries of sparse, approximately monosemantic features. We introduce an evaluation protocol suited to a domain where LLM-based auto-interpretability fails and spatial autocorrelation matters, scoring any SAE on reconstruction, sparsity, monosemantic recovery, and spatial coherence against a curated label registry of categorical and continuous reference products. On this protocol, three state-of-the-art SAE recipes (ReLU+\ell_1, BatchTopK, and Matryoshka BatchTopK) all recover monosemantic and spatially contiguous concepts from AlphaEarth and Tessera embeddings much better than raw embeddings, with Matryoshka the strongest of the three. Two applications follow: (i) unsupervised discovery of land-cover subconcepts (e.g., splitting water into river, lake, ocean), and (ii) improved area-of-applicability estimates for OOD detection. Together, these results are a step toward Earth observation models whose features can be inspected rather than only used, opening pathways for scientific discovery and improved trustworthiness in geospatial applications. Code and models will be released upon publication.
Abstract: Effective machine learning depends not only on how we model data, but also on what data we choose to collect. While large sequence models have revolutionized data modeling, the problem of automated data selection, or "intrinsic curiosity", remains a significant challenge. Classic approaches incentivize exploration by rewarding an agent based on its "learning progress", which measures how much a newly acquired observation improves a world model's predictive ability. However, evaluating these rewards traditionally requires expensive inner loops of gradient descent updates within each trajectory, rendering them computationally impractical at scale. In this work, we investigate whether the emergent in-context learning (ICL) capabilities of sequence models can eliminate this bottleneck by serving as immediate, update-free world models. Specifically, we evaluate whether an exploration policy can be trained to maximize learning progress, using solely the prediction errors and counterfactual context manipulations of an in-context learner. We first prove that in general Markov decision processes, this is in fact impossible in an unbiased way: the resulting intrinsic rewards either suffer from nuisance terms that bias their estimation of true learning progress, or they cannot be implemented using an in-context learner's prediction errors. Conversely, we prove a positive result for a broad subclass of non-temporal settings, encompassing active learning and Bayesian Experimental Design: here, ICL-derived rewards successfully bound and asymptotically converge to the true learning progress. We corroborate our theory with controlled experiments across continuous and symbolic environments, demonstrating that our ICL-driven framework successfully trains curious data-collection policies that explore optimally.
Abstract: Diffusion models generate samples by iteratively querying learned score estimates. A rapidly growing literature focuses on accelerating sampling by minimizing the number of score evaluations, yet the information-theoretic limits of such acceleration remain unclear. In this work, we establish the first score query lower bounds for diffusion sampling. We prove that for d-dimensional distributions, given access to score estimates with polynomial accuracy \varepsilon=d^-O(1) (in any L^p sense), any sampling algorithm requires \widetilde\Omega(\sqrtd) adaptive score queries. In particular, our proof shows that any sampler must search over \widetilde\Omega(\sqrtd) distinct noise levels, providing a formal explanation for why multiscale noise schedules are necessary in practice.
PaperID: 22, Oral
Abstract: Fine-tuned LLMs can learn complex and subtle behaviors. However, it can be difficult to ascertain what behaviors the model acquired during fine-tuning. We introduce LoRAcles: fine-tuned language models that take LoRA adapter weights as input, and answer natural language questions about them. We introduce a self-supervised pipeline for training LoRAcles: we first train LoRAs on small sets of pre-training documents, and then train LoRAcles on these LoRAs to answer questions about the documents. LoRAcles generalize far beyond their training data: they are state-of-the-art on AuditBench, a benchmark of models with sophisticated hidden behaviors; they can detect subtle changes in model organisms (fine-tunes deliberately constructed to encode a hidden behavior for auditing research), such as subliminal learning; and they are the first tool to achieve non-trivial performance at verbalizing semantic backdoor triggers. LoRAcles can also describe some learned behaviors from the weights of a full-parameter fine-tune. LoRAcles are a highly scalable method: we train LoRAcles for Qwen-3-14B and Llama-3.3-70B, and observe that performance scales smoothly with the size of our training dataset up to 100K LoRAs, though our method may need refinement in order to scale further. We note, though, that LoRAcles are prone to hallucinations and currently only elicit the most salient behavioral changes, limiting their utility beyond narrow fine-tunes.
Abstract: Tools from random matrix theory have become central to deep learning theory, using spectral information to provide mechanisms for modeling generalization, robustness, scaling, and failure modes. While often capable of modeling empirical behavior, practical computations are limited by matrix size, often imposing a restriction to models that are too small to be realistic. This motivates the inference of properties of larger models from the behavior of smaller ones. Free decompression (FD) is a recently proposed method for extrapolating spectral information across matrix sizes, but its utility is currently limited by strong assumptions that preclude its implementation on more realistic machine learning (ML) models. We use algebraic spectral curve theory to provide a general FD methodology for spectral densities whose Stieltjes transform satisfies an algebraic relation, a modeling assumption that is more likely to hold in practice. This recasts FD as an evolution along spectral curves which can be readily integrated. Our framework enables the expansion of spectral densities that have multiple or multi-modal bulks, that exist at multiple scales, and that contain atoms, all characteristic of real-world data and popular ML models. We demonstrate the efficacy of our framework on models of interest in modern ML, including Hessian and activation matrices associated with neural networks and large-scale diffusion models.
Abstract: Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce , a meticulously curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabulary taught above Grade 5. Training a 2B-parameter LLM from scratch on , a model with sufficient language competence for open-ended evaluation, yet with clear knowledge and capability boundaries mapped to interpretable curriculum guidelines. We release as a developmentally restricted sandbox to study how models acquire, represent, and use data under a well-defined training scope. We illustrate the sandbox's utility in a first suite of experiments on injecting new knowledge through post-training and in-context learning. These methods let better utilize existing knowledge, but do not raise out-of-scope capabilities. Our findings underscore the value of this controlled environment for future investigations.
PaperID: 25, Oral
Abstract: Multimodal learning beyond two modalities commonly leverages a specific modality (e.g., text) to bind other modalities. However, how to establish a more balanced representation space that approximates shared semantics while respecting the holistic geometry of n-modal data remains challenging. In this work, we present BaryBind, which aims to transport the specific modality towards the Wasserstein barycenter (WB) optimized across all modalities and introduces a volumetric alignment objective to establish a unified semantic space around the WB embedding. Specifically, we project specific modalities to the WB, which minimizes the average Wasserstein distances to multimodal distributions and serves as the anchor for subsequent alignment. We then construct a barycenter simplex, whose volume is taken as a similarity metric for global alignment centered at the WB. Extensive experiments show that BaryBind achieves competitive performance in text-video-audio retrieval, classification, videoQA, and cross-modal generation tasks, along with robustness under modality absence and scalability to more than three modalities.
Abstract: Learning in neural systems arises from synaptic changes that reshape the representations underlying behavior. While low-rank recurrent neural networks (RNNs) have emerged as a powerful framework for linking connectivity to function, a theoretical understanding of their learning process remains elusive. Here, we extend the low-rank framework from activity to learning by deriving gradient-descent dynamics directly in a reduced overlap space. We formulate a closed-form, low-dimensional system of ODEs that governs learning in this space, exact for linear RNNs and asymptotically exact for nonlinear RNNs in the large-N Gaussian limit. Central to our analysis is a distinction between two classes of overlaps: _loss-visible_ overlaps, which fully determine network activity, output, and loss, and _loss-invisible_ overlaps, which do not affect function but are required to describe learning. We illustrate the consequences of this decomposition through two phenomena. First, we show that learning can serve as a perturbation that exposes differences in connectivity between functionally equivalent networks. Second, we show that loss-invisible overlaps can act as memory variables that encode training history, and characterize the conditions under which this occurs. Finally, we present several testable predictions for biological learning experiments derived from our theory.
Abstract: The widespread adoption of AI assistants has prompted the development of privacy-aware platforms designed to extract insights from real-world usage. Their privacy protections primarily rely on layering multiple heuristic techniques, such as PII redaction, clustering, and aggregation. In this paper, we put their privacy claims to the test by presenting CLIOPATRA, the first attack against privacy-preserving LLM-based insights systems. Our attack involves an adversary that carefully designs and inserts malicious chats into the system to break multiple layers of protections and induce the leakage of sensitive information from a target user’s chat. We evaluate CLIOPATRA on one such platform, Anthropic’s Clio, and target synthetically generated medical chats to show that an adversary can successfully and confidently (with nearly 100% precision) extract the medical history contained in these chats in up to 65% of cases. We also show that CLIOPATRA can stealthily extract information by obfuscating the private information in the generated insights. Finally, we demonstrate that existing ad hoc mitigations, such as LLM-based privacy auditing, are unreliable and fail to detect major leaks. Taken together, our findings indicate that, even when layered, current heuristic protections are insufficient to adequately protect user data, and that prompt injection has been an understudied risk in LLM-based insight systems.
PaperID: 28, Oral
Abstract: Open-world semi-supervised learning aims to use limited labeled data from known classes to classify unlabeled samples that contain both known and novel classes. Most existing methods over-rely on known classes and tend to generalize incorrectly in complex open-world settings. This leads to degraded pseudo-label quality and the model is often overconfident in wrong predictions. To address this, this paper proposes a Mitigating Confidence Miscalibration Method in Open-World Semi-Supervised Learning (OpenMCM). By leveraging discriminative knowledge from known classes to guide learning in the novel class space, our method effectively mitigates overconfident misclassifications and improves discrimination accuracy. Furthermore, we propose an adaptive threshold calibration strategy to independently determine optimal decision boundary for each class. By integrating a high-reliable pseudo-label fusion mechanism, we enhance recognition stability for novel classes. Experiments on three benchmark datasets (i.e., CIFAR-10, CIFAR-100, and ImageNet-100) show that our framework achieves performance improvements over existing state-of-the-art methods.
Abstract: Joint prediction sets for multivariate time series should control a single event while adapting to cross-coordinate dependence. We study filtered conformal ellipsoids: a frozen state-space filter emits a one-step predictive mean and covariance, and split-conformal calibration is applied to the resulting Mahalanobis scores. The filter is used to choose the ellipsoid shape; conformal calibration chooses the scalar radius, so the construction benefits from a learned predictive covariance without relying on Gaussian tail probabilities for coverage. The main difficulty is that filtered scores are dependent and learned recurrent filters need not contract in their raw hidden state; we therefore analyse contraction in an observable predictive-law quotient that identifies hidden states producing the same future sequence of emitted Gaussian laws. Under a stable Bayes Gaussian-projection filter, covariance bounds, and a finite-horizon observability/Fisher condition, small excess Gaussian negative log-likelihood implies contraction of the learned emitted laws. Combined with a threshold-autocovariance envelope this yields a Chebyshev-type approximate coverage bound for filtered split-conformal prediction under dependence; a sharper Bernstein-type bound requires an additional geometric-mixing concentration assumption. Under Gaussian oracle realisability we also obtain a near-oracle log-volume comparison within the class of conditionally valid Gaussian ellipsoid rules. We instantiate the framework with a GCN-GRU filter with diagonal-plus-low-rank covariance. On moderate-size graph-native traffic benchmarks (METR-LA-20 and PEMS-BAY-50), the learned filter gives sharper at-target ellipsoids than static-covariance and non-filter baselines; at full-graph scale and on non-graph-native datasets, factor and copula baselines can be stronger.
Abstract: AI generated predictions increasingly inform decision making in critical tasks, and therefore must be trustworthy. One widely used measure of trustworthiness is calibration, which requires that the predictions match the true frequencies and can be treated like real probabilities of a given outcome. However, defining calibration is subtle, and designing good measures of calibration error has been an active topic of recent research. The first goal is to find calibration measures that are , meaning they can inform decision makers about their utility loss when predictions are treated as true probabilities, which is known as swap regret. The second goal is to find calibration measures that are , meaning that calibration error can be measured from a small sample of predictions and outcomes. Although these are very basic requirements, there is no existing calibration measure that fully satisfies both properties, and all existing measures relax actionability by bounding a weaker notion of swap regret, or relax testability by having suboptimal estimation error. We introduce a new calibration measure, Soft-Binned Calibration Decision Loss (SCDL), which we prove is fully actionable without weakening either requirement, and testable with nearly optimal error rate. In addition, SCDL satisfies other desired properties such as continuity and consistency. We also provide a set of experiments confirming that the theoretical advantages of SCDL compared to other measures lead to better performance in practice.
Abstract: We study the optimal scale at which real-valued function classes exhibit uniform convergence and learnability. Our main result establishes a scale-sensitive generalization of the fundamental theorem of PAC learning: for every bounded real-valued class \mathcalF and every \gamma>0, uniform convergence at scale \gamma, agnostic learnability at scale \gamma/2, and finiteness of the fat-shattering dimension at every scale \gamma'>\gamma are equivalent. This resolves a question by Anthony and Bartlett (Cambridge Univ. Press ’99) on the precise scales governing learnability, refuting a conjecture attributed there to Phil Long that a multiplicative 2-factor gap is unavoidable, and improves the upper bounds of Bartlett and Long (JCSS ’98), which incur such a loss. The key technical ingredient is a direct bound on empirical \ell_\infty covering numbers, avoiding the standard detour through packing numbers. As a consequence, we obtain sharp asymptotic metric-entropy bounds in terms of the fat-shattering scale \gamma: an O(\log^2 n) bound holds already at scale \gamma/2, while an O(\log n) bound holds at scale 2\gamma. We further show that the O(\log^2 n) bound is sometimes tight. These results resolve open questions by Alon et al. (JACM ’97) and Rudelson and Vershynin (Ann. of Math. ’06). As an application, we establish a sharp dichotomy for bounded integral probability metrics: every such IPM is either estimable or cannot be weakly evaluated within any multiplicative factor c<3, while 3-weak evaluability always holds, resolving an open question from Aiyer et al. (ICML '26). We conclude with several open questions on quantitative sample complexity and evaluability.
Authors:
Ivan Kobyzev, Abbas Ghaddar, Ali Nasiri-Sarvi, Lifeng Shang, Yufei CUIAbstract: State Space Models (SSMs) compress sequence history into a bounded recurrent state, making the resulting memory law a central architectural choice for long-context performance. Most modern SSMs rely on ODE-based dynamics that lead to exponential forgetting, limiting their ability to retain information over broad temporal ranges. We introduce FRAC, a selective SSM architecture derived from fractional dynamics that replaces this exponential decay with power-law long memory. To make fractional dynamics practical, FRAC approximates the heavy-tailed target kernel with a finite-state, log-spaced sum of exponential modes. This construction turns fractional memory into an efficient recurrent module with parallel training and prefill, while retaining bounded-state autoregressive decoding. Extensive experiments, including 1.3B-parameter language modeling, demonstrate that FRAC consistently improves long-context performance over state-of-the-art SSM baselines while staying competitive on short-context. These results show that fractional dynamics provide a practical and effective prior for long-context SSMs.
Abstract: Mixture-of-Experts (MoE) is now the dominant architecture for frontier language models, yet it requires all expert parameters to be loaded in memory, making it less preferable for memory-constrained deployment. Existing compression methods reduce the number of experts but the output remains an MoE model with the same fundamental limitation. We present the first systematic framework for converting a trained MoE into a standard fully dense architecture: experts are scored, selected, and grouped, then concatenated into a dense FFN and refined by knowledge distillation from the MoE teacher. We evaluate 7 scoring, 5 grouping, and 2 magnitude scaling methods across a range of selected expert counts on Qwen3-30B-A3B, yielding 350 configurations. We find that the choice of scoring method is the most impactful, with our novel diversity-aware scoring consistently outperforming prior methods on Qwen3-30B-A3B, DeepSeek-V2-Lite, and GPT-OSS-20B. Under a controlled comparison at matched parameter count, MoE-to-dense outperforms dense-to-dense pruning by +6.3 pp in average downstream accuracy after ~4B-token distillation at 1.6x faster training wall-clock speed.
Abstract: Parameter-efficient fine-tuning (PEFT) is a natural way to adapt pretrained vision-language-action (VLA) policies, but most adapter designs apply temporally static updates throughout a control rollout, overlooking the phase-dependent nature of continuous-action manipulation. Such policies traverse distinct regimes, including approach, contact transition, grasping, transport, and placement, each requiring different adaptation behaviors. We propose PhaseLoRA, a lightweight LoRA parameterization that conditions adaptation at each control step using two weakly supervised descriptors: fine-control tendency and event/boundary intensity. PhaseLoRA modulates the LoRA left factor in the action expert, allowing the effective low-rank update direction to vary over time while keeping the backbone largely frozen. On LIBERO, PhaseLoRA improves average success rate by 12.2 points over a matched-parameter high-rank LoRA baseline and outperforms stronger LoRA variants. Ablations and update-direction analyses show that the gain is not explained by parameter count, arbitrary temporal modulation, or scalar gating, but by control-regime-aligned changes in low-rank update directions. These results establish within-trajectory conditioning as an effective lightweight PEFT axis for continuous-action VLA policies.
PaperID: 35, Oral
Abstract: Concept-based explanation methods have emerged as a prominent framework for interpreting deep neural networks using human-understandable concepts. However, existing approaches are limited to assigning a scalar importance score to a concept, failing to capture the interaction between the concept and its visual context. We argue that overlooking this interaction leaves a predictor's reasoning only partially understood. To address this gap, we introduce IDEA, a novel global post-hoc explanation framework that explicitly models the interaction between a concept and its visual context. Through this modeling, IDEA decomposes concept-level predictive information into three atomic components: uniqueness (independent concept contribution), redundancy (information shared with the context), and synergy (information arising strictly from concept-context interaction). Consequently, IDEA extends interpretability beyond merely identifying influential concepts to diagnosing how a predictor reasons -- for instance, flagging potential context dependencies that scalar scores cannot reveal. IDEA requires no task-specific concept annotations and no access to classifier internals. Evaluations on controlled and real-world settings demonstrate that IDEA can surface reasoning patterns invisible to traditional scalar concept scores.
Abstract: Many emerging agentic paradigms require agents to collaborate with one another (or people) to achieve shared goals. However, existing approaches to learning policies for such collaborative problems produce brittle solutions that fail when paired with new partners. We attribute these failures to a combination of free-riding during training and a lack of strategic robustness. To address these problems, we study the concept of strategic risk aversion and interpret it as a principled inductive bias for generalizable cooperation with unseen partners. While strategically risk-averse players are robust to deviations in their partner's behavior by design, we show that, in collaborative games, they also (1) can have better equilibrium outcomes than those at classical game-theoretic concepts like Nash, and (2) exhibit less or no free-riding. Inspired by these insights, we develop a multi-agent reinforcement learning (MARL) algorithm that integrates strategic risk aversion into standard policy optimization methods. Our empirical results across collaborative benchmarks (including an LLM collaboration task) validate our theory and demonstrate that our approach consistently achieves reliable collaboration with heterogeneous and previously unseen partners across collaborative tasks.
Abstract: We introduce TiRex-2, a recurrent xLSTM-based time series foundation model that generalizes the univariate TiRex to multivariate forecasting with both past and future covariates. Real-world forecasting is inherently sequential: observations arrive continuously, variables evolve jointly, and a subset of covariates is known ahead of time. Existing Transformer-based time series foundation models capture cross-variate dependencies but incur quadratic complexity in context length and require full-history recomputation as new observations arrive. TiRex-2 addresses these limitations through a memory-centric recurrent design that operates at constant per-patch cost under streaming. The model combines a bidirectional time mixer with an asymmetric grouped-attention variate mixer, enabling the integration of future-known covariates while preserving strict causality over target variables. To our knowledge, this is the first time series foundation model that achieves this combination of properties. To support scalable multivariate pretraining, we propose a synthetic coupling pipeline that composes diverse multivariate samples on the fly from large univariate corpora. Empirically, TiRex-2 achieves state-of-the-art zero-shot performance on GIFT-Eval and fev-bench, remains stable when streamed to arbitrary context lengths, and maintains constant inference cost per patch. The model uses 38.4M active parameters in univariate mode, with an additional 44.1M parameters activated for multivariate forecasting.
PaperID: 38, Oral
Abstract: The use of algorithmic predictions in decision-making leads to a feedback loop where the models we deploy actively influence the data distributions we see, and later use to retrain on. This dynamic was formalized by Perdomo et al. 2020 in their work on performative prediction. Our main result is an unconditional reduction showing that any no-regret algorithm deployed in performative settings converges to a (mixed) performatively stable equilibrium: a solution in which models actively shape data distributions in ways that their own predictions look optimal in hindsight. Prior to our work, all positive results in this area imposed strong restrictions on how models influenced distributions. By using a martingale argument and allowing randomization, we avoid any assumption on how populations respond to predictions and sidestep recent hardness results showing that deterministic stable models are in general PPAD-hard to compute. Lastly, on a more conceptual note, our connection sheds light on why common algorithms, like gradient descent, are naturally stabilizing and prevent runaway feedback loops. We hope our work enables future technical transfer of ideas between online optimization and performativity.
Abstract: Recent text-only models demonstrate remarkable reasoning capabilities. Extending these to visual domains requires vision-language models to translate images into text descriptions. However, current models, trained to produce captions for human readers, often omit the precise details that reasoning systems require. This creates an interface mismatch: reasoners often fail not due to reasoning limitations but because they lack access to critical visual information. We propose Adaptive-Clarification Reinforcement Learning (AC-RL), which, through interaction, teaches vision models what information reasoners need. Our key insight is that clarification requests during training reveal information gaps; by penalizing success that requires clarification, we create pressure for comprehensive initial captions that enable the reasoner to solve the problem in a single pass. AC-RL improves average accuracy by 4.4 points over pretrained baselines across seven visual reasoning benchmarks, and analysis shows it would cut clarification requests by up to 44% if those were allowed. By treating clarification as a form of implicit supervision, AC-RL demonstrates that vision-language interfaces can be effectively learned through interaction alone, without requiring explicit annotations.
Abstract: Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate this process, as they can inspect datasets, generate code, and produce analyses end-to-end. However, we show that they frequently make subtle inferential errors that lead to incorrect conclusions despite correctly executed analyses. Existing benchmarks fail to capture this failure mode, as they rarely assess whether a reported p-value is statistically valid given the assumptions underlying the data. We address this gap by building P-Bench, a benchmark comprising 425 open-ended, realistic hypothesis-testing tasks spanning economics, biology, and medicine. Each task requires an agent to select a statistical method, compute a p-value, and draw a conclusion given only a scientific hypothesis and a dataset. We further introduce Fisher-R1, an open-weight LLM agent trained for rigorous hypothesis testing using synthetic tasks and reinforcement learning. On P-Bench, Fisher-R1-14B substantially improves over its backbone and outperforms strong proprietary and open-source baselines, including GPT-5.4 and DeepSeek-V4-Pro, achieving a 21% average relative improvement in single-trial success over DeepSeek-V4-Pro, with gains up to 26% on the most challenging tasks. Our results demonstrate that current LLM agents lack reliable statistical reasoning for hypothesis testing and that reinforcement learning on tasks with verified statistical reward substantially improves reliability.
Authors: Tristan Shah, Stas Tiomkin
Abstract: Intrinsic Motivation (IM) aims to train agents without external rewards, enabling useful behavior to emerge from the agent's interaction with its environment alone. However, the dominant IM approaches rely on information-theoretic quantities with designer-chosen variables, introducing bias and lacking a principled connection to dynamics or optimal control (OC). We introduce Controllable Information Production (CIP), a new foundation for IM explicitly grounded in dynamical systems and OC. CIP measures the rate at which an agent produces information, capturing controllable complexity without external knowledge or bias. CIP unifies IM and OC into a single framework, formalizing physical intelligence as the control of information production. It further reveals connections between the structure of the value function and Kolmogorov–Sinai entropy. CIP consistently outperforms prior IM methods on standard benchmarks in robot learning and solves tasks they fail on, including humanoid self-righting. These results support a general organizing principle: physical intelligence emerges from driving systems toward the edge of controllable chaos.
PaperID: 42, Oral
Authors:
Alessandro Burzio, Tobias Fischer, Sven Elflein, Qunjie Zhou, Riccardo de Lutio, Jiawei Ren, Jiahui Huang, Shengyu Huang, Marc Pollefeys, Laura Leal-Taixé, Zan Gojcic, Haithem TurkiAbstract: Recent feed-forward 3D reconstruction transformers have scaled to over a billion parameters, following the broader trend of increasing model capacity in computer vision. Yet emerging evidence suggests that contiguous transformer layers often behave like repeated applications of similar operations, and multi-view reconstruction transformers refine their predictions progressively across decoder depth. We posit that model depth partially buys iteration, paid for inefficiently in unique parameters, and instead make that iteration explicit in architecture. Our model, DéjàView, applies a single looped transformer block recurrently to per-view features for K refinement steps. Trained once, it exposes K as an inference-time compute knob, matching or outperforming substantially larger feed-forward baselines across five reconstruction benchmarks spanning indoor, outdoor, object-centric, and driving scenes, while using a fraction of their parameters and comparable or lower compute. Importantly, the same looped block formulation outperforms an otherwise identical variant with independent per-step parameters under matched training data and compute, suggesting that explicit iteration is not merely a compute-efficient substitute for capacity but a stronger inductive bias for multi-view 3D reconstruction.
Abstract: A common failure mode in long-horizon agentic test-time scaling is error propagation, where factual errors or invalid deductions introduced at intermediate steps persist in the agent's belief state and contaminate later reasoning. Existing test-time scaling methods provide limited control over this process, as they often rely on agents to detect their own mistakes, select among flawed trajectories, or refine solutions only after errors have already shaped the reasoning path. We propose ExComm, a communication protocol for exploration-stage agentic test-time scaling. ExComm is motivated by the empirical observation that the majority of intermediate errors in parallel agentic reasoning produce detectable cross-agent factual conflicts. Leveraging the iterative structure of agentic workflows, ExComm periodically audits agent belief states to detect such conflicts, resolves them through a dedicated tool-based verification loop, and returns concise, targeted feedback to the involved agents. Corrections are incorporated through soft belief updates, which append verified feedback rather than overwriting existing beliefs. Furthermore, to prevent collapsing trajectory diversity due to communication, ExComm further introduces a trajectory diversification module that redirects redundant trajectories toward orthogonal strategies. Experiments on AIME 2024, AIME 2025, and GAIA with Gemini-2.5-Flash-Lite and Qwen3.5-4B show that ExComm consistently outperforms strong test-time scaling baselines, achieving average performance gains of 5.7% and 5.0% over the best-performing baselines, respectively. Further analyses demonstrate improved error recovery, favorable scaling behavior, stronger diversity than adapted communication baselines, and the best performance-cost trade-off among the evaluated methods.
PaperID: 44, Oral
Authors:
Xiaoran Zhao, Weiming Liu, Lianyong Qi, Weiyi Zhong, Yuwen Liu, Xiaolong Xu, Haolong Xiang, Qiang Ni, Xuyun Zhang, Wanchun Dou, Xiaokang ZhouAbstract: Recommendation unlearning (RU) aims to remove the influence of specified user--item interactions from a trained recommender while preserving utility on the remaining data. Existing RU methods usually require access to the original training data, especially the remaining data, to recover recommendation utility after unlearning. However, access to such data is often restricted or unavailable in practice due to privacy constraints, storage costs, or third-party deployment, giving rise to source-free recommendation unlearning. To address this challenge, we propose an adaptive source-free recommendation unlearning ( first estimates the inaccessible gradient on the remaining data. It exploits the first-order optimality condition of the trained recommender and approximates the remaining gradient using the forgetting data and squared gradient accumulators from adaptive optimizers, e.g., Adam. then formulates unlearning as a gradient-constrained update, where the estimated remaining gradient is used as a constraint to limit utility degradation while maximizing the unlearning objective. adaptively adjusts the utility-preserving correction according to the current gradient geometry. We further theoretically prove that approaches an approximate Pareto stationary solution. Extensive experiments on three real-world datasets demonstrate that achieves strong performance in both unlearning and recommendation. The code is available at https://anonymous.4open.science/r/AdaSRU-94D9/.
PaperID: 45, Oral
Abstract: Recurrent neural networks are widely used in neuroscience and machine learning, but their weight matrices often appear dense and unstructured, defying clean interpretation. The Schur decomposition expresses any recurrent weight matrix in quasi-triangular form, allowing network dynamics to be interpreted as a latent circuit of "functionally feedforward" interactions between orthogonal modes. However, its application for interpreting recurrent weight matrices has been limited due to its well-known combinatorial non-uniqueness. We introduce SchurMO (Schur Manifold Optimization), a method for discovering structured Schur forms via Riemannian gradient descent on the manifold of orthogonal similarity transformations, enabling recovery of target motifs. Compared to a greedy baseline, SchurMO more reliably and more efficiently recovers ground-truth latent motifs (chain-like, banded, and modular structure) from noisy weight matrices. Applying SchurMO to nonlinear recurrent networks enables recovery of latent chain-like structure, providing a link between connectivity and the underlying dynamics of sequence production. Our results highlight SchurMO as a scalable, computationally efficient approach to discovering latent structure in the Schur forms of recurrent weights.
Abstract: Building capable medical vision-language models (VLMs) is challenging due to the scarcity of domain data and the difficulty of eliciting pre-trained medical visual knowledge. We present HuatuoGPT-5, a medical VLM built on a simple yet effective recipe: large-scale multimodal pre-training followed by reinforcement learning (RL). We first construct PubMedVision-Plus, the largest medical image-text dataset with 60M pairs from public PubMed literature, along with a clean 12M high-quality subset. After pre-training on this data, we find that traditional supervised fine-tuning (SFT) struggles to elicit the pre-trained knowledge. In contrast, by applying RL with verifiable and rubric-based instances synthesized directly from the pre-training data, we successfully incentivize the model to use this pre-trained medical visual knowledge. On PubMed-derived data, HuatuoGPT-5 outperforms similarly sized VLMs on multiple medical benchmarks and expert evaluations, significantly narrowing the gap between medical VLMs and leading proprietary VLMs. We hope these data, models, and findings can help accelerate open research on capable medical VLMs.
PaperID: 47, Oral
Authors: Israel R Soares, Pedro Benedetti, Thanda Shwe, Israel Mendonca, Masayoshi Aritsugi
Abstract: Empirical Risk Minimization treats training samples uniformly, which can degrade performance in noisy or heterogeneous settings. We introduce a reference-guided objective that uses per-sample error comparisons with a fixed reference model to reweight gradient contributions without relying on prediction imitation, while preserving gradient direction and ensuring bounded scaling. Theoretical analysis shows that the formulation induces adaptive weighting with controlled amplification. Experiments on image classification with synthetic label noise and on time-series forecasting demonstrate broad improvements across architectures, particularly in moderate-to-high noise and heterogeneous regimes. Results further indicate that the method is most effective when the error-ranking alignment between the baseline model and the reference model is low, suggesting that gains arise primarily from the diversity of the reference error ranking relative to the learner, rather than the reference's absolute performance.
PaperID: 48, Oral
Abstract: Uncertainty quantification for individual treatment effects (ITEs) is a daunting challenge in causal inference. Motivated by recent advances in conformal prediction, several works aim to construct distribution-free prediction sets for ITEs with desired coverage under standard assumptions such as strong ignorability and overlap. In this paper, we show that such goals are fundamentally unattainable in the presence of continuous covariates. Specifically, we establish finite-sample and asymptotic impossibility results demonstrating that any distribution-free prediction set achieving desired coverage for ITEs must be trivial, in the sense that it has infinite expected length. Our analysis relies on a connection between ITE inference and the hardness of conditional independence testing, and highlights the intrinsic limitations imposed by the missing data nature of causal inference. These results provide a new perspective on existing methods, clarifying that their apparent success necessarily relies on additional structural assumptions beyond standard causal assumptions.
Abstract: LLM-agent simulation offers a flexible computational tool for studying population response trajectories that depend on scenario events, memory, demographics, and evolving social context. However, full multi-round simulation scales linearly with both population size and horizon, requiring every agent to query the LLM at every round. We propose Adaptive Prototype Simulation (APS), a framework that treats scaling as a recurrent LLM-oracle allocation problem rather than a static sampling, propagation, or surrogate-modeling problem. APS preserves the specified LLM as the online transition oracle while querying adaptive core prototypes, selected singleton-tail agents, and shadow-audit agents. Prototype responses induce local response surfaces for nearby agents, reducing online LLM calls without replacing the underlying transition model. To control approximation bias, shadow-audit residual correction estimates propagation residuals for aggregate correction and future budget allocation, while tail-protected singleton routing directly queries selected isolated, heterogeneous, or high-curvature regions that are vulnerable to smoothing. We analyze APS as an estimator of the LLM-induced population transition and decompose its error into prototype-coverage error, shadow-audit residual-correction error, local-propagation bias, and temporal context mismatch. Under the reported protocols, APS gives lower reference-aligned distributional discrepancy than scale-oriented and same-budget baselines while reducing online LLM calls, with ablations and compact robustness checks diagnosing the main bias-control mechanisms. In a 10M-agent, multi-round public-opinion simulation, APS achieves a 381.1× reduction over full simulation, with reference-aligned final-round JSD 0.094 against the corresponding full-LLM reference.
PaperID: 50, Oral
Abstract: Large language model (LLM)-powered multi-agent systems have been shown to effectively solve real-world end-to-end problems through collaboration among diverse agents. However, as LLM capabilities advance and user demands continue to grow, the range of problems that agents can address expands dramatically, while the population of available agents and tools continues to increase, giving rise to open and evolving ecosystems of autonomous agents. In this paradigm, agents are likely to become increasingly specialized, each focusing on particular capabilities or problem types. In such a setting, a central challenge is how to efficiently assemble an appropriate team of agents to take on different responsibilities from a massive agent pool for a given task. In this work, we study this problem in a controlled setting with 300 task-oriented agents. SMART (Scalable Multi-Agent Role-conditioned Teaming) first plans a task-conditioned collaboration topology and candidate role set, then constructs a compact role-conditioned candidate pool, and finally uses a LLM-free Monte Carlo Tree Search (MCTS) to assign agents to roles online via metadata-based rollouts. This design enables high flexibility without incurring prohibitive inference costs. Across multiple experiments, our approach achieves superior performance on GSM8K (93.64%), MATH (61.73%), HumanEval (93.89%), and MBPP (87.68%) using gpt-4o-mini under a limited budget.
PaperID: 51, Oral
Abstract: World models are increasingly used as policy-in-the-loop imagination environments, where reliable rollouts require fine-grained controllability with respect to low-level robot actions. A key obstacle to scaling such models in robotics is that actions are not a universal language in pixel space: changes in visual environment, camera view, robot placement, or embodiment alter how the same numerical action manifests visually, leading to conflicting supervision under mixed training and brittle generalization at deployment. We introduce SyncWorld, an action-conditioned world model that serves as a zero-shot simulator across unseen environments without any additional training. SyncWorld leverages a visual calibration episode---paired frames and actions that showcase all the controllable degrees of freedom---to specify the setup-specific Action--Visual Mapping in context. Training with visual calibration contexts teaches the model to interpret actions through visual evidence and to leverage interaction history when explicit calibration is unavailable. Experiments show that SyncWorld can accurately simulate action outcomes in previously unseen settings, and that its capability of simulating rollouts enables test-time policy improvement without training.
PaperID: 52, Oral
Abstract: Minimum Description Length (MDL) formalizes the principle of Occam's razor by optimizing the total description length: L(\mathrmmodel)+L(\mathrmdata \ | \ \mathrmmodel). For sequential prediction, the MDL method repeatedly selects a model with a minimum objective score of the observed prefix for the next step prediction. Classical MDL prediction theory shows that exact optimization of the MDL objective indeed provides a strong compression guarantee that supports reliable prediction. However, practical machine learning usually can only find models by approximately optimizing the objective function. To bridge this gap, this paper addresses the following fundamental question: \emphUnder what forms of approximation and regularization does approximate MDL still guarantee reliable sequential prediction? This work offers a principled characterization. We prove that for any approximation with \emphadditive slack \(C\) of the more general form of the \emphbalanced MDL objective: \lambda\cdot L(\mathrmmodel)+L(\mathrmdata \ | \ \mathrmmodel), the cumulative expected squared prediction error is finite for all \lambda\ge1. The case \lambda>1 is proved by an affinity-telescoping argument, while the boundary case \lambda=1 is proved by a likelihood-ratio stopping argument based on exact static MDL bounds. Our results establish that classical MDL regularization remains robust to any fixed additive optimization error. Furthermore, we establish that our characterization of the approximate MDL framework is sharp: When 0<\lambda<1, overfits can happen to incur infinite cumulative expected error in the universal class of estimable measures, and hence a strong form of model-complexity regularization is necessary. In addition, model selection may fail in every regularized regime \lambda >0, under multiplicative approximation, and thus, additive approximation is both sufficient and essential.
PaperID: 53, Oral
Authors:
Longkang Li, Jin Tian, Kun ZhangAbstract: Protein-molecule virtual screening is increasingly cast as a problem of representation learning in a shared embedding space. Existing methods rely on dense holistic alignment, entangling invariant binding determinants with nuisance correlations and limiting transfer to new targets. It has been noted that binding in protein–molecule systems involves sparse cross-modality interactions: binding is governed by a small contact interface and a few decisive local interactions (e.g., hydrogen bonds, hydrophobic contacts, and salt bridges) rather than the global structures of the protein and molecule. We hypothesize that uncovering and leveraging sparse interaction patterns is critical for generalization beyond the training data, as these patterns are reusable and expected to improve performance across different scenarios. In this paper, we aim to identify and leverage sparse interaction patterns, and verify our hypothesis. Since the training data contain only observed binding pairs, we formalize this prior via a V-structure causal model under Heckman-style selection, and establish three theoretical results: (i) the latent concepts of interacting proteins and molecules are not identifiable without appropriate sparsity constraints;(ii) these concepts and their sparse interactions are component-wise identifiable under structural sparsity conditions; and (iii) a low-rank relaxation of these conditions yields subspace identifiability of the concepts and interactions. Inspired by these principles, we propose CausalBind with three implementation variants. Extensive experiments on DUD-E and LIT-PCBA benchmarks show that all variants consistently outperform strong baselines across all metrics, respectively.
PaperID: 54, Oral
Authors: Isabella Thiel, Juan M. Bello-Rivas, Yannis Kevrekidis, Themis Sapsis
Abstract: Extreme events in chaotic systems are difficult to learn from short trajectories because they are controlled by transient finite-time instability rather than by frequently observed bulk dynamics. We propose a mechanism-aware conditioning plug-in framework that turns a nudged coarse ensemble into a non-intrusive sensor of local instability geometry. In the small-noise regime, the ensemble covariance aggregates the same finite-time deformation kernels that govern local instability, providing a Jacobian-free proxy for the local amplification structure around a synchronized coarse trajectory. A small FiLM module injects statistics of this ensemble geometry into an otherwise unchanged backbone while leaving the coarse simulator unchanged. We demonstrate this interface in two distinct pipelines: a Transformer-style residual-attention corrector for a controlled low-dimensional chaotic system and a probabilistic recurrent STORN corrector for topographic two-layer quasi-geostrophic (QG) flow. In the low-dimensional benchmark, ensemble covariance directions co-activate with OTD modes and FiLM conditioning improves 99th-percentile exceedance-frequency errors over an identical no-context Transformer baseline. In QG, a fixed ensemble-conditioned FiLM-STORN model trained on only 50 time units substantially improves long-horizon rare-event statistics in the data-limited regime, including density-tail errors, exceedance frequencies, and spatial exceedance-area distributions relative to an unconditioned STORN trained on the same data; on averaged high-threshold exceedance diagnostics, it also outperforms the baseline STORN trained with 20 times more high-resolution data. These results show that local instability geometry is not merely interpretable post hoc, but an actionable conditioning signal for data-efficient rare-event emulation.
PaperID: 55, Oral
Abstract: Naturalistic language encoding models based on large language model features can predict cortical responses well, but their learned representations remain difficult to interpret. Three challenges are central: semantic selectivity is fundamentally non-identifiable in high-dimensional correlated feature spaces, cortical organization remains implicit in voxelwise weight maps, and cross-subject structure is usually addressed only after model fitting. We introduce a multisubject coupled sparse autoencoder that jointly reconstructs language-model features and predicts subject-specific brain responses, learning a sparse basis of semantic-cortical atoms that makes semantic structure, cortical organization, and cross-subject correspondence explicit within a single model. In naturalistic story-listening fMRI, the model retains 88% of matched dense-feature ridge brain prediction performance using 50 atoms and k=5 active atoms per TR. Relative to matched two-stage baselines, joint training yields the strongest overall tradeoff between prediction and interpretability. The learned atoms are semantically coherent and interpretable, recover more reproducible cross-subject brain maps than standard post hoc encoding-model analyses, align with large-scale cortical structure from Yeo~7 and Neurosynth, and transfer better than ridge to held-out subjects in the low-data regime. This framework provides a more interpretable basis for cognitive neuroscience analyses of semantic-to-brain encoding by linking semantic features to corresponding cortical response patterns across subjects.
Abstract: Evaluating Computer Use Agents (CUAs) on interactive environments is fraught with methodological pitfalls that the field has yet to systematically address. We show that a 1MB replay script that blindly executes a recorded action sequence without ever observing the screen outperforms frontier models on prominent static benchmarks, and prove that its expected success rate is exactly equal to the source agent's pass@k in deterministic environments. We trace this and other failures to two root causes: non-principled environment design (static, unsandboxed, or unreliably verified environments) and non-principled evaluation methodology (naïve aggregation and misuse of pass@k for stateful UI interactions). To address the first, we propose PRISM, five design principles for CUA environments (privileged verification, realistic environments, integrity-checked configurations, sandboxed execution, and multifactorial variability) and instantiate them in DigiWorld, a benchmark of 15 realistic sandboxed mobile applications able to evaluate agents in over 3.2 million verified unique configurations. To address the second, we develop an aggregation framework pairing Wilson score intervals with hierarchical bootstrap, producing confidence intervals that correctly account for the nested structure of CUA benchmarks, as we empirically demonstrate. All together, we show that principled environment design and rigorous evaluation methodology are not optional refinements but prerequisites for meaningful CUA research.
PaperID: 57, Oral
Abstract: Arena-style rankings of language models are a widely used evaluation framework in leaderboards like LMArena, research papers, and model evaluation in production. These rankings aim to realistically assess model performance using pairwise comparisons of anonymous models on user prompts, aggregating votes via the Bradley-Terry model to produce the final ranking. However, we find that these leaderboards are highly vulnerable to several manipulation strategies that unscrupulous providers could use to boost a model's rank: strategic voting (individual votes that boost a model's rank, including votes on irrelevant models), strategic prompting (choosing prompts that favor a model), and strategic nomination (boosting the rank of a model by inserting irrelevant models). We demonstrate these effects on public leaderboard data, analyze the circumstances that cause these vulnerabilities, and describe approaches to safeguard against each attack. Notably, these issues hold for any pairwise preference-based evaluation that uses the Bradley-Terry model. These manipulations boost the rank of nearly every tested model; each attack can boost a model by up to 15-21 places in a ranking and, combined, they can boost a model by up to 45 places.
PaperID: 58, Oral
Authors: Omer Wasim, Sami Davies, Shallu Tomer, Rathish Das
Abstract: In this paper, we introduce a semi-random model in online optimization, which interpolates naturally between the standard (adversarial) and random-order arrival model, which we call the online with a random sample (ORS) model: the full input may be chosen adversarially, but a random p-fraction of the input is selected and presented to the algorithm in random-order, before the remaining non-sampled elements are presented (in adversarial order). While similar in spirit to the AOS (adversarial order with a sample) model introduced by Kaplan, Naori and Raz, in our ORS model, the sampled p-fraction is revealed online instead of offline, and hence, competitiveness is measured with respect to the full input. Our central focus in the paper is applying the ORS model to the secretary problem and its natural (k,1) variant. For the secretary problem, we present an optimal pe^-p-competitive algorithm, while for the (k,1)-secretary problem, we obtain a competitive algorithm whose performance converges to optimal competitiveness as p\rightarrow 1. We also include a O(\log (1/p))-competitive algorithm for facility location.
PaperID: 59, Oral
Authors: Zach Cohen, Jan Drugowitsch
Abstract: Neural populations exhibit heterogeneous tuning: Neurons differ in their selectivity, gain, and in how their preferred inputs are distributed across stimulus or latent-variable space. While this heterogeneity has been studied extensively in the context of encoding precision, its implications for learning and generalization are less well-understood. Here, we develop a probabilistic theory that links a neural population's tuning statistics to its inductive bias: in large populations, the tuning statistics induce a kernel and, equivalently, a prior of a generative model over functions. This prior in turn predicts the population's ability to generalize across a distribution of related tasks. Our theory determines the population tuning statistics that minimize the generalization error over a distribution of functions. We further present a normative account of tuning adaptation inspired by our probabilistic approach, showing that tuning changes that maximize the marginal likelihood of observed data systematically reduce generalization error relative to non-adapting populations. Applied to the adaptation of hippocampal spatial representations in a reward learning task, we found that our theory could explain the over-representation of rewarding locations through experience in an environment with a localized reward. We further demonstrate how different components of the population's generative model differentially impact the speed and magnitude of tuning adaptation. Overall, our work proposes a novel connection between tuning variability and a population's ability to learn and to generalize across tasks and develops a Bayesian theory of tuning adaptation in neural populations during learning.
Abstract: Unsupervised pretraining has driven empirical advances in goal-conditioned reinforcement learning (GCRL), but its theoretical foundations remain poorly understood. In particular, an influential class of methods, mutual information skill learning (MISL), discovers behaviorally diverse skills that can later be used for downstream goal-reaching. However, it remains a theoretical mystery why skills learned through MISL should support goal-reaching. A subtle challenge is that both GCRL and MISL are umbrella terms: different GCRL tasks use distinct criteria for measuring goal-reaching performance, while different MISL methods optimize distinct notions of behavioral diversity. We address this challenge and unify GCRL and MISL as instances of control maximization. We identify three canonical GCRL formulations and prove that they are fundamentally inequivalent: they can induce incompatible optimal policies even in the same environment. Nevertheless, they all share a common interpretation: a well-performing goal-conditioned policy is one whose future trajectory is highly sensitive to the commanded goal, with the precise notion of sensitivity determined by the GCRL formulation. Noting that MISL objectives can be understood as measures of skill-sensitivity akin to goal-sensitivity, we show that MISL objectives are bounded by formulation-specific downstream goal-sensitivities. These bounds establish a precise correspondence between MISL methods and downstream GCRL tasks: for every GCRL formulation, there exists a matching MISL objective for which more diverse skills afford greater downstream goal sensitivity. Our results thus lay a theoretical foundation for RL pretraining and have important practical implications, such as suggesting which pretraining objectives to use when a user cares about a specific class of downstream tasks.
PaperID: 61, Oral
Abstract: Synthetic data can scale visual pretraining, but current selection criteria only partially explain which datasets train strong representations. We propose using the energy profile of a dataset---the distribution of log-probability levels it occupies under a normalised real-image density model---as an audit and curation principle for synthetic data. To operationalise this view, we introduce Energy-Level Distribution Distance (Eld), which compares real and synthetic energy distributions. Real images span sparse high-likelihood scenes and dense low-likelihood textures, so useful synthetic data should cover this energy range rather than collapse to likely but low-variation samples. Across diverse synthetic datasets, Eld exposes mismatches missed by FID and predicts pretraining performance beyond FID, LPIPS variation, and feature-space recall, including across CNN and ViT self-supervised settings. We then use Eld to resample and mix synthetic sources, matching the real-image energy histogram while preserving diversity and improving downstream transfer.
Abstract: We consider two-layer neural networks trained in the feature-learning regime using gradient descent, and relate the output of the finite-width network f_\hat\rho to its infinite-width counterpart f_\rho^MF, which evolves in the mean-field dynamics. While constant-time horizon bounds for \|f_\rho^MF -f_\hat\rho\| may be obtained via standard Grönwall estimates, the long-time behavior of the fluctuation is a more delicate matter. Uniform-in-time bounds often rely on (local) strong convexity in the landscape or Logarithmic Sobolev inequalities present in noisy gradient dynamics. % In this work, we study a noiseless setting and do not make assumptions on the geometry of the landscape near the optimum. Instead, we impose a weaker condition on the convergence rate of the mean-field deterministic Wasserstein-gradient-flow dynamics. In this work, we establish non-asymptotic weak propagation-of-chaos that holds uniformly in time, obtained by exploiting instead the convergence rate of the mean-field deterministic Wasserstein-gradient-flow dynamics. Specifically, denoting by L_t the mean-field loss at time t and m the number of neurons, under standard regularity assumptions and the condition \int_0^\infty L_t^1/2 dt =O_d(1), we obtain the uniform in time bound \|f_\rho^MF -f_\hat\rho\|^2 \lesssim \textpoly(d) m^-\min(1,c/6) whenever L_t \lesssim t^-c. Our result holds in a noiseless setting and does not make any assumptions on the geometry of the landscape near the optimum. A key implication of our result is that whenever the convergence rate of the mean-field, population-loss dynamics is faster than 1/t^2, we can attain a loss of \epsilon with only \textpoly(d/\epsilon) neurons, training samples, and GD steps.
Abstract: SO(3) equivariant graph neural networks have become the dominant paradigm for atomistic foundation models, achieving high accuracy and data efficiency by building rotational symmetry directly into the architecture. Yet the computational cost of their higher-order tensor operations creates a tough trade-off between model accuracy and inference efficiency. In this paper, we propose a structural pruning method for SO(3) equivariant atomistic foundation models to bridge this accuracy-efficiency gap. The pruning is applied along the channel and order dimensions, with each irreducible representation kept or removed as a complete block, thereby retaining SO(3) equivariance. Starting from a large checkpoint, the pruned model substantially reduces the inference cost while retaining higher accuracy than an independently trained small model. The pruned MACE-MP model outperforms the official from-scratch trained small model on 7 of 9 metrics on the Matbench Discovery leaderboard. In terms of efficiency, compressed MACE-MP and MACE-OFF models contain 1.5× to 4× fewer parameters and require 2.5× to 4× less pre-training compute than training a small model from scratch. For downstream applications, fine-tuning the pruned model reduces energy and force errors by 70.1% and 34.4% compared to training task-specific models from scratch across eight representative downstream datasets. We demonstrate that the method generalizes to other SO(3) equivariant architectures (SevenNet, eSCN) and can be combined with quantization and knowledge distillation for further gains.
Authors:
Boxin Wang, Chankyu Lee, Nayeon Lee, Sheng-Chieh Lin, Wenliang Dai, Yang Chen, Yangyi Chen, Zhuolin Yang, Zihan Liu, Mohammad Shoeybi, Bryan Catanzaro, Wei PingAbstract: Building general-purpose reasoning models with reinforcement learning (RL) entails substantial cross-domain heterogeneity, including large variation in inference-time response lengths and verification latency. Such variability complicates the RL infrastructure, slows training, and makes training curriculum (e.g., response length extension) and hyperparameter selection challenging. In this work, we propose cascaded domain-wise reinforcement learning (Cascade RL) to develop Nemotron-Cascade, capable of operating in both \emphinstruct and deep \emphthinking modes, without any performance gap relative to a thinking-only counterpart. Departing from conventional approaches that blend heterogeneous prompts from different domains, Cascade RL orchestrates sequential, domain-wise RL, reducing engineering complexity and delivering state-of-the-art performance across a wide range of benchmarks. Notably, RLHF for alignment, when used as a pre-step, boosts the model's reasoning ability far beyond mere preference optimization, and subsequent domain-wise RLVR stages rarely degrade the benchmark performance attained in earlier domains and may even improve it. Our 14B model, after RL, outperforms its SFT teacher, DeepSeek-R1-0528, on LiveCodeBench v5/v6/Pro and achieves silver-medal performance in the 2025 International Olympiad in Informatics (IOI). We transparently share our training and data recipes.
PaperID: 65, Oral
Authors: qiao huang, Ao Shen, Jie Du, Manning Wang
Abstract: Protein variant effect prediction is central to understanding protein function and guiding protein engineering, yet experimental labels remain costly. In many protein experiments, the variants that can be assayed are dictated by protocol constraints, while the variants that need accurate prediction may lie in a different or broader region of the landscape. This assay-prediction mismatch makes candidate-local criteria such as predicted function, uncertainty, or exploration value insufficient for measurement design. Given a protocol-defined accessible space and a limited budget, we ask which variants should be assayed so that the resulting data best improves landscape prediction. We propose TALK, which estimates each candidate’s contribution through posterior coupling between accessible and target variants and reduces redundant information within the batch. Across ProteinGym and GB1 settings spanning mutation-order transfer and different accessible-target relationships, TALK uses limited budgets more efficiently, yielding labels that better support subsequent prediction than baselines. These results point to a practical two-stage workflow: first, spend a limited assay budget on variants selected to be informative for the target landscape and train a predictive model from the measured labels; then use this model to score and prioritize high-potential variants.
PaperID: 66, Oral
Authors: Yao Rong, Vaibhav Unhelkar
Abstract: When auditing machine learning models, human experts sequentially inspect model predictions and explanations, gradually building a mental model of the model’s behavior to uncover errors. This manual process is critical for safe deployment in high-stakes domains but fundamentally limited in scalability. In this work, we propose a novel method to simulate this human auditing behavior for automating the auditing process. We formalize auditing as a sequential decision-making problem, modeled as a Markov decision process. Our state representation summarizes the auditor’s accumulated knowledge of model predictions and feature usage. Building on this formalization, we solve the MDP using reinforcement learning, yielding RLAuditor, which learns to select informative test samples to efficiently uncover model errors. We further connect our formalization to Theory of Mind, showing that our state mirrors how human auditors build mental models: its update rule is equivalent to Bayesian belief revision and it converges to the model’s behavioral pattern at an interpretable rate. We validate our approach on multiple ML models and datasets across different data modalities. In a human subject study, we demonstrate that RLAuditor helps human auditors produce more accurate audit reports in less time.
Abstract: Estimating the number of distinct elements in a data stream is well understood when repeated elements are identical. In modern settings, however, observations are high-dimensional and noisy, so repeated instances of the same object are only approximately similar -- for example, different images of the same individual may vary significantly at the pixel level. Classical sketches such as HyperLogLog rely on consistent hash values for identical elements and break down in this regime. Recent work on robust distinct counting in general metric spaces achieves \tilde\Theta(\sqrtn) memory, which is tight in the worst case. We show that substantially improved memory guarantees are possible under geometric structure common in learned representations. We introduce MaxSketch, a simple max-linear sketch built from random Gaussian projections, and prove that it succeeds in estimating the number of distinct latent objects. Concretely, we show that under this assumption m = \tildeO(\log n / \varepsilon^4) random projections (and hence \tildeO(\log n/\varepsilon^4) memory) suffice to recover the true distinct count within a (1+\varepsilon) factor. Experiments on image streams confirm that MaxSketch accurately estimates distinct counts and generalizes beyond the training regime. Our results bridge classical streaming algorithms and modern representation learning, showing how geometric structure can fundamentally reduce the complexity of distinct counting.
PaperID: 68, Oral
Authors: Joakim Andén, Justus Sagemüller
Abstract: Many projection-based imaging modalities reconstruct unknown objects from collections of noisy projections related by unknown in-plane rotations, making effective estimation dependent on rotationally consistent alignment and aggregation. In cryogenic electron microscopy (cryo-EM), the extremely low signal-to-noise ratio (SNR) makes such multi-image integration essential. We introduce the polar transformer, a permutation-invariant architecture for sets of images built around an \mathrmSO(2)^K-equivariant angular attention mechanism that jointly aligns and aggregates information across images. This architecture leverages a stable polar representation in which we define polar CNN feature extractors and evaluate attention weights over all relative angular shifts efficiently via FFT-based correlations, guaranteeing equivariance to independent in-plane rotations of each image in the set. We evaluate the polar transformer on simulated cryo-EM images using denoising as a quantitative proxy. Across different regimes, polar transformers outperform strong single-image and image-set baselines, achieving a reduction of up to 40% in relative MSE at an SNR of 0.02 and improving downstream ab initio reconstruction resolution by 20%.
PaperID: 69, Oral
Authors: Adi Mashiach
Abstract: Foundation models trained on large unlabeled corpora develop emergent capabilities — tasks that go beyond the training objective. We extend this paradigm to multimodal biology with Bonbon, a foundation model for molecular interactions trained on approximately 3.2 trillion tokens of paired protein-ligand sequences. From a single self-supervised objective, Bonbon exhibits four zero-shot emergent capabilities at hierarchical levels of resolution — functional (mechanism of action), residue (binding site localization, including orthosteric/allosteric discrimination), bond (covalent versus non-covalent engagement), and atom (active moiety identification at single-bond resolution). All four capabilities generalize to out-of-distribution proteins. Bonbon was trained on protein-ligand sequences alone, without any task-specific supervision and without structural coordinates. We conducted controlled scaling experiments across four model scales from 42M to 1.6B parameters; every capability evaluated shows a sharp jump at the largest scale on continuous metrics, even as pretraining loss decreases smoothly. This demonstrates that loss-based scaling laws alone are insufficient to predict when zero-shot emergent capabilities become available. The capabilities reported here emerge after training at a 1:2,000 parameter-to-token ratio to accommodate the sparsity of paired biological interaction data. The same frozen representations achieve state-of-the-art binding and affinity predictions at five orders of magnitude greater speed than structure-based methods. In prospective wet-lab validation, Bonbon achieves a 70% hit rate on de novo compounds, discovering novel chemotypes not previously documented as inhibitors for those targets. All of the tested compounds had less than 60% structural similarity to molecules in the training corpus.
Authors:
Shubham Agarwal, Alexander Krentsel, Shu Liu, Mert Cemri, Audrey Cheng, Rui Meng, Tomas Pfister, Chun-Liang Li, Sylvia Ratnasamy, Aditya Parameswaran, Matei A Zaharia, Ion Stoica, Mohsen LesaniAbstract: AI agents increasingly excel at generating, testing, and refining code. However, they fall short on tasks requiring formal guarantees of full coverage that testing alone cannot provide. Distributed systems are a prime example: properties such as consistency between reads and writes must hold under every possible interleaving of events. Mechanized formal verification can guarantee such correctness, but typically demands months to years of expert effort. As evidence, even SOTA coding agents (Claude Code with Opus 4.6, Codex with GPT-5.4) succeed on only 2/7 distributed key-value-store specifications. In this paper, we present the first effective approach to addressing this gap, Inductive Deductive Synthesis (IDS), which jointly and incrementally synthesizes implementation and proof, and learns from failed attempts to systematically try promising strategies. Built as an agentic LLM system, IDS achieves 7/7 in about 6.8 hours and 106 per spec on average, roughly 200× faster than expert effort and 17% cheaper than SOTA agents. IDS further incorporates performance feedback into the same loop, yielding implementations up to 3× faster than published verified systems.
Abstract: In the past year, custom and unreleased math reasoning models reached gold medal performance on the International Mathematical Olympiad (IMO). Similar performance was then reported using large-scale inference on publicly available models but at prohibitive costs (e.g., 3000 per problem). In this work, we present an inference pipeline that attains best-in-class performance on IMO-style math problems at an average inference cost orders of magnitude below competing methods while using only general-purpose off-the-shelf models. Our method relies on insights about grader failure in solver-grader pipelines, which we call the Cognitive Well (iterative refinement converging to a wrong solution that the solver as well as the pipeline's internal grader consider to be basically correct). Our pipeline addresses these failure modes through conjecture extraction, wherein candidate lemmas are isolated from generated solutions and independently verified alongside their negations in a fresh environment (context detachment). On IMO-ProofBench Advanced (PB-Adv), our pipeline achieves 67.1% performance using Gemini 3.0 Pro with an average cost per question of ~ \31. This surpasses the performance of the DeepThink IMO gold-winning model, and more than doubles the success rate of the next best publicly accessible pipeline, all at a fraction of the cost.
Abstract: Follow the regularized leader (FTRL) is a foundational algorithm in online learning whose regret properties have been extensively studied for decades. However, in stark contrast to mirror descent---and gradient descent in particular---the convergence of FTRL to first-order stationary points in constrained nonconvex optimization was hitherto unresolved. In this paper, we show that continuous-time FTRL can fail to converge even asymptotically in constrained optimization problems; that is, the Karush-Kuhn-Tucker (KKT) gap of FTRL can remain bounded away from zero indefinitely. This is especially surprising in light of the fact that continuous-time FTRL guarantees a monotonic improvement of the objective value. As a result, we establish that non-convergent behavior of no-regret dynamics---which besets general variational inequality problems---can occur even in potential systems. From a technical standpoint, our construction relies on an infinitely differentiable function where the gradient flow dynamics exhibit decelerating oscillations without ever converging pointwise, which we in turn embed into higher-dimensional FTRL dynamics that sustain a perpetually large KKT gap.
Authors: Krishna Khadka, Yu Lei, Raghu Kacker, D. R Kuhn
Abstract: Class Activation Mapping (CAM) and related saliency techniques are widely used to explain predictions of deep vision models, producing post-hoc heatmaps that highlight image regions a model relied on. However, existing CAM methods aggregate weighted contributions over the full set of internal units, including units whose contribution may be incidental. The resulting heatmaps are often diffuse and obscure which units are actually sufficient to preserve the prediction. We reframe visual explanation as the selection of a 1-minimal sufficient subset of internal representational units, feature maps in CNNs and patch tokens in ViTs, whose joint activation preserves the model's prediction. A subset is 1-minimal sufficient if removing any single unit changes the prediction. We propose DD-CAM, a gradient-free procedure adapted from delta debugging in software engineering that instantiates this formulation. DD-CAM exploits classifier-head structure with two regimes: iterated single-unit removal for non-interacting heads (GAP+FC), and recursive partition-and-reduce for interacting heads (multi-layer FC stacks and ViT self-attention). On 2,000 ImageNet validation images across eight CNN and ViT architectures, DD-CAM is the strongest method on 13 of 18 faithfulness metric-group cells against thirteen baselines. On the NIH ChestX-ray14 bounding-box set, DD-CAM improves IoU by 45% and F1 by 22% over the strongest baseline while producing single-region explanations. These results suggest that necessity-grounded internal-unit selection can produce more focused and faithful explanations than weighted aggregation over all units, with particular promise for safety-critical settings such as medical imaging.
PaperID: 74, Oral
Abstract: Accurate forecasting of Earth's dynamic systems is vital in geophysics, where neural operators offer a powerful data-driven paradigm. Standard operators often treat the Earth as a flat Euclidean plane, introducing geometric artifacts. In contrast, Spherical Neural Operators (SNOs) resolve this mismatch by operating within the spherical spectral space, naturally adapting to spherical geometries. However, for computational efficiency, existing SNOs typically rely on the rotation equivariance assumption. This assumption simplifies the transitions of physical states into isolated spectral modal evolutions, thereby over-idealizing the Earth as a homogeneous system. To solve this problem, we no longer rely on the rotation equivariance assumption that isolates spectral modal evolution, but directly derive an explicit modal interaction kernel and propose Diagonal Spherical Neural Operators named DIAGNO. DIAGNO explicitly parameterizes cross-modal interactions through spectral modal decomposition, thereby preserving the inherent heterogeneity of Earth dynamics. Furthermore, leveraging the unique diagonal structure of the spherical spectrum, we design inter- and intra-diagonal interaction mechanisms to natively capture essential zonal and meridional dynamics, respecting intrinsic geophysical regularities. Extensive experiments across simulated fluid dynamics and real-world Earth system reanalysis (including atmosphere and ocean) datasets demonstrate DIAGNO’s superior capability, consistently achieving state-of-the-art performance in multi-step forecasting. Beyond this, we validate DIAGNO’s practical analytical value in professional geophysical domains by assessing its fidelity in capturing climatological anomalies and maintaining physical consistency.
PaperID: 75, Oral
Authors:
Yixiang Fan, Zhiqiang Pu, Hao Ma, Dongmin Li, Weihao Zhang, Yinxing Dai, Shijie Zheng, Jingjing HuangAbstract: Mode connectivity studies reveal that high-performing neural networks are often connected by low-loss paths in parameter space. We investigate the analogue of this phenomenon in cooperative multi-agent reinforcement learning (MARL), where a solution is not a single model but a team of decentralized policies. We introduce coordination connectivity, a diagnostic evaluating whether linear interpolation between the actor parameters of two trained policy teams preserves high team return. The resulting coordination barrier measures whether two successful teams are separated by a performance cliff along the joint-policy path. We evaluate this diagnostic on the StarCraft Multi-Agent Challenge (SMAC) across five controlled training histories, isolating the effects of shared initialization, continued training, parameter decoupling, and HAPPO-style separate-actor specialization. Across three SMAC maps and four seeds, policy teams derived from a shared MAPPO checkpoint exhibit near-zero interpolation barriers, even after non-shared specialization. In contrast, interpolation paths involving independently trained HAPPO teams exhibit substantially larger barriers. Mixed-team assembly, single-agent hot-swap, clone-all, role-swap, and weight-space probes further indicate that shared-initialized agents specialize into distinct roles while remaining within a low-barrier coordination component. These results suggest that shared initialization implicitly structures the joint-policy space of cooperative MARL, providing a mode-connectivity-inspired lens for analyzing coordination, specialization, and same-source cross-version team compatibility.
PaperID: 76, Oral
Abstract: Anchor graphs are a standard route to scalable multi-view clustering, but most existing pipelines still depend on either a shared anchor set sampled from concatenated features or a separate anchor-alignment stage before fusion. Both choices are restrictive: the former hard-codes equal anchor budgets and importance at anchor-construction time, while the latter introduces a cumbersome correspondence problem and still struggles with unequal anchor numbers. We develop an alignment-free probabilistic formulation in which view-specific anchor graphs are noisy manifestations of a shared latent cluster identity rather than objects that require cross-view correspondence. Each view has its own anchor space, number of anchors and cluster prototypes; the only shared latent variable is the sample label. We instantiate this idea with a Gaussian latent-label model and a deterministic prototype M-step. The resulting variational EM algorithm admits closed-form updates with convergence guarantees, while experiments demonstrate competitive clustering performance and practical runtime on most datasets.
Abstract: We study online learning in the random-order model, where the multiset of loss functions is chosen adversarially but revealed in a uniformly random order. By extending the batch-to-online transformation of Dong and Yoshida (2023), we show that if an offline algorithm enjoys a (1+\varepsilon)-approximation guarantee, an average sensitivity bound controlled by a function \varphi(\varepsilon), and stability with respect to \varepsilon, then we can obtain a small-loss regret bound typically of order \tilde O(\varphi^\star(\mathrmOPT_T)), where \varphi^\star is the concave conjugate of \varphi, \mathrmOPT_T is the offline optimum over T rounds, and \tilde O hides polylogarithmic factors in T. Our result refines their original (1+\varepsilon)-approximate regret guarantee and applies to a broad class of problems, including online k-means clustering and online low-rank approximation. We further apply our approach to online submodular function minimization using (1\pm\varepsilon)-cut sparsifiers of submodular hypergraphs, obtaining a small-loss regret bound of \tilde O(n^3 + n^3/4\mathrmOPT_T^3/4), where n is the ground-set size; we also demonstrate its applicability to online \ell_1 regression. Our work sheds light on the power of sparsification and related algorithmic techniques in achieving small-loss regret bounds in the random-order model, without requiring structural assumptions on loss functions, such as linearity or smoothness.
PaperID: 78, Oral
Abstract: Tabular data is widely applied in critical domains, where tree-based models remain dominant, yet their performance is often constrained by manual heuristic splitting criteria and limited ensemble diversity. While hyperbolic geometry naturally fits the hierarchical topology of tabular data, existing hyperbolic algorithms are hindered by complex Riemannian optimization and the difficulty of defining subtree priors in non-Euclidean spaces. To address these challenges, we propose HOT, a hyperbolic random forest framework based on data decoupling. First, we leverage intrinsic geometric structures and LLM reasoning to guide subtree construction, replacing traditional manual heuristics with geometric-semantic priors. Second, we introduce an orthogonal hyperbolic subspace projection mechanism via tangent space isometry, effectively decoupling feature dependencies and maximizing ensemble diversity. Finally, we establish an end-to-end collaborative training paradigm where gradient residuals from a graph network directly guide the growth of new hyperbolic trees, bypassing expensive iterative optimization. Experiments on 13 benchmark datasets demonstrate that HOT significantly outperforms state-of-the-art baselines in both computational efficiency and predictive performance.
PaperID: 79, Oral
Authors: Meher Chaitanya Pindiprolu, My Le, Luana Ruiz
Abstract: We introduce Graph Cascades, a mesoscopic rewiring strategy for Graph Neural Networks (GNNs) and Graph Transformers (GTs) that captures intermediate-scale graph structure beyond purely local edges or fully global attention. Using contagion-based diffusion processes, Graph Cascades constructs, in \mathcalO(|V|+|E|) time, an auxiliary graph where node pairs supported by repeated multi-hop reinforcement are promoted to direct neighbors. We theoretically characterize when reinforcement-based rewiring helps: sufficient conditions under which a reinforcement-based edge-selection rule has higher label agreement than direct adjacency, an SBM witness in which two-hop reinforcement is perfectly homophilic, and formalizing mesoscopic connectivity via graph effective resistance. Empirically, across node-classification benchmarks, Graph Cascades improves multiple GNN and sparse-GT backbones, with the most reliable gains observed on heterophilic and moderate- to high-degree homophilic graphs. The theoretical conditions also identify regimes where mesoscopic rewiring is unlikely to be beneficial -- low-degree regular graphs and graphs with structural bottlenecks -- and these predictions match the observed failures. We additionally observe tight correlations between performance and structural properties in the rewired graphs.
Abstract: Whole-brain 4D fMRI generation is valuable for modeling functional brain dynamics, yet existing fMRI foundation models mainly target representation learning and downstream prediction rather than conditional predictive generation. We introduce BrainWorld, a structural-prior-conditioned generative model for whole-brain 4D fMRI dynamics. BrainWorld uses sMRI as subject-level anatomical context to guide future fMRI generation, integrating structural information into the denoising process rather than treating it as a parallel modality. Evaluated on 22 datasets spanning diverse cohorts and brain states, BrainWorld generates stable 4D fMRI trajectories up to 400 frames, improves downstream performance through generated-example augmentation, and learns transferable multimodal representations that outperform baselines. Together, these results establish BrainWorld as a condition-aware generative framework for long-horizon brain dynamics modeling and multimodal representation learning.
PaperID: 81, Oral
Authors:
Niclas A Göring, Shuofeng Zhang, Roi Holtzman, yoonsoo nam, Ard LouisAbstract: Every neuron in a trained network outputs a number that can be separated into its sign (positive or negative) and its magnitude (absolute size). We call the set of all neuron signs the sign code. This raises a key question: does classification depend more on the sign code or on the magnitudes? For overparameterized vision classifiers with sign-preserving activations, the answer is overwhelmingly the sign code. A k-nearest neighbor classifier operating only on this binary sign code achieves accuracy comparable to the full model across different architectures, activation functions, and datasets. Moreover, replacing all magnitudes in every layer with random noise while keeping the original signs barely changes residual network performance. In contrast, flipping just 20% of the signs reduces accuracy to chance. Models generalize well when same-class inputs share similar sign codes. When this overlap is disrupted by a regularizer, training accuracy remains high, but test performance collapses. A margin decomposition helps explain these results: the sign-only contribution to each logit grows as \sqrtd in network width, dominating the magnitude residual. The natural unit of analysis for these networks may not be through their continuous weight space but instead through their discrete sign codes.
Authors: Simon Kuang, Xinfan Lin
Abstract: We study the problem of propagating the mean and covariance of a general multivariate Gaussian distribution through a deep (residual) neural network using layer-by-layer moment matching. We close a longstanding gap by deriving exact moment matching for the probit, GeLU, ReLU (as a limit of GeLU), Heaviside (as a limit of probit), and sine activation functions; for both feedforward and generalized residual layers. On random networks, we find orders-of-magnitude improvements in the KL divergence error metric, up to a millionfold, over popular alternatives. On a variational Bayesian neural network, we show that our method attains hundredfold improvements in KL divergence from Monte Carlo ground truth over a state-of-the-art deterministic inference method. We also give an smooth-distance error bound showing that, under regularity assumptions, moment matching removes the leading low-variance errors and propagates higher-order local accuracy through the layers of a network.
PaperID: 83, Oral
Authors: Cameron Dufault, Scott Xu, Alan Moses
Abstract: Despite the increasing scale of genome language models (gLMs), their ability to decode the function of regulatory sequences remains unclear. gLM pretraining relies on sequence reconstruction, which may struggle due to the noisy, rapidly evolving nature of regulatory DNA. Self-supervised contrastive approaches provide a promising alternative. Inspired by language-image architectures like CLIP, we demonstrate contrastive promoter-protein pretraining (C3P). By learning to align promoters to their corresponding proteins, we leverage the rich representations of proteins learned by protein language models as supervisory signal for the learning of promoter representations. After training on 88 million bacterial promoter-protein pairs, we evaluate the predictive power of C3P-learned promoter representations for inference of curated regulatory annotations, finding multi-fold improvement over leading gLMs. We also introduce zero-shot co-regulated gene retrieval, the ability to find co-regulated genes in a genome using no experimental data. We find that compared to a randomly initialized baseline, C3P training consistently provides significant zero-shot performance gains, unlike gLMs. Scaling analysis reveals the potential for further improvement as well as the efficiency of C3P, which achieved strong performance at a fraction of the training cost of leading gLMs. In addition to demonstrating that C3P training is effective for learning representations of bacterial regulatory sequences, our strong zero-shot co-regulated gene retrieval performance suggests the possibility of decoding gene regulation for millions of bacteria from their genomes alone.
PaperID: 84, Oral
Abstract: Embodied agents make decisions by combining information from heterogeneous memory modules, including declarative memory for explicit knowledge, procedural memory for learned routines, and working memory for current observations. During interaction, the memory contents required for decision making change with the agent's state, goal, and task progress, while many contents remain shared across consecutive decisions. However, existing memory modules are typically accessed through separate interfaces, causing each updated memory context to be processed as a new input to the foundation model. This prevents selective key-value (KV) cache reuse for unchanged memory contents and forces full KV cache recomputation even when much of the memory context remains valid. We introduce HeMeR, an inference framework that performs KV reconciliation between dynamically updated memory contexts and reusable KV states. HeMeR tracks declarative, procedural, and working memory contents across action decisions, preserves KV states for unchanged contents, and reconciles only the states affected by memory-content updates. This enables efficient and stable inference as memory contents evolve during interaction. We evaluate HeMeR on ALFWorld, Habitat, and real-world robotic deployments requiring coordinated use of heterogeneous memory contents. On Real-world, HeMeR achieves 83.20% task success and reduces per-step inference time from 2.94 to 1.87 seconds compared with MemRL by preserving compatible KV states as task progress updates the memory context.
Abstract: Neural surrogate solvers of partial differential equations (PDEs) promise dramatic speedups over numerical methods, especially in scenarios requires many solves. However, current accuracy-based evaluations do not fully consider two central issues: (i) neural solvers incur substantial up-front costs for data generation, training, and tuning; and (ii) classical solvers can also generate low-fidelity solutions at a sufficiently low simulation cost. To explicitly account for these realities and fully incorporate end-to-end costs, we propose an evaluation framework centered on , a metric that counts the forward solves before a learned solver is cost-effective relative to an error-equivalent traditional solver. To evaluate this measure, we apply scaling laws to determine how much training budget to allocate to data generation and discuss how to achieve smooth error-matching in diverse settings. We evaluate the breakeven complexity of many modern neural PDE solvers on three PDEs on 2D periodic domains from APEBench and a novel benchmark of flows past multiple obstacles generated by the GPU-native PyFR code. Among other findings, our results suggest that neural PDE solvers will be on problems as problems get harder in terms of cost, dimension, rollout, physics regime (e.g. higher Re), etc.
Abstract: Current video foundation models, including the strongest self-supervised models such as V-JEPA2, fail to capture how humans organize social information in dynamic scenes. For example, across a range of diverse vision models tested, none were able to predict human similarity judgments to social video clips as well as a sentence embedding model of the caption text (MPNet). We show this gap in vision model performance can be closed by a compact behavioral supervisory signal. We introduce behavioral geometric supervision (BGS): a hybrid objective that constrains local and global pairwise embedding geometry to match the relational similarity structure across videos. We apply this method using a new human similarity dataset, containing 49,484 odd-one-out judgments from 250 naturalistic social video clips, and low-rank adaptation across four ViT backbones (V-JEPA 2/2.1, TimeSformer, VideoMAE, and CLIP). We find that one of the best fine-tuned models, V-JEPA 2.1, nearly triples in performance compared to the pre-trained baseline and reaches close to the noise ceiling, exceeding the strongest sentence-embedding baseline. In addition, finetuned models (i) capture unique variance in human judgments that caption-based language embeddings do not, (ii) develop interpretable social-affective attributes (valence, arousal, and dominance) despite never being trained on any of these attributes, (iii) zero-shot transfer to a separate dataset of out-of-distribution abstract social interactions, and (iv) shift spatial attention from scene context to socially informative regions (faces, gaze, and interacting bodies). A matched language-distillation control fails to reproduce these gains, ruling out caption transfer as the mechanism. Our results show how a modest amount of human behavioral data can steer video models toward human-like social visual understanding.
Authors:
Chiyu Zhang, Huiqin Yang, Bendong Jiang, Xiaolei Zhang, Yiran Zhao, Ruyi Chen, Lu Zhou, Xiaogang Xu, Jiafei Wu, Liming Fang, Zhe LiuAbstract: The rapid proliferation of LLM-based autonomous agents in real operating system environments introduces a qualitatively new category of safety risk beyond traditional content safety: \emphbehavior jailbreak, where an adversary induces an agent to execute dangerous OS-level operations with irreversible physical consequences. Existing benchmarks either evaluate safety at the semantic output layer alone, missing physical-layer harms, or fail to isolate test cases, letting earlier runs contaminate later ones. We present LITMUS (LLM-agents In-OS Testing for Measuring Unsafe Subversion), a benchmark that addresses both gaps through a semantic–physical dual verification mechanism and an OS-level state rollback design. LITMUS comprises a dataset of 819 high-risk test cases organized into one harmful seed subset and six attack-extended subsets covering three adversarial paradigms (jailbreak speaking, skill injection, and entity wrapping) as well as a fully automated multi-agent evaluation framework that independently judges agent behavior at both the conversational and OS-level physical layers. Evaluation across multiple frontier agents reveals three consistent findings: (1) current agents lack effective safety awareness against dangerous instructions in real OS environments, with the strong model (e.g. Claude Sonnet 4.6) still executing 40.64% of high-risk operations; (2) agents exhibit pervasive Execution Hallucination (EH), verbally refusing a request while the dangerous operation has already completed at the system level, a phenomenon invisible to every prior semantic-only evaluation framework; and (3) skill injection and entity wrapping attacks we designed achieve high success rates, exposing pronounced agent vulnerabilities to malicious skill interference and instruction obfuscation. LITMUS provides the first standardized platform for reproducible, physically grounded behavioral safety evaluation of LLM agents in real OS environments. The dataset and code are in the supplementary material.
Abstract: Similarity measures such as canonical correlation analysis (CCA), representational similarity analysis (RSA) and center kernel alignment (CKA), are widely used to compare the representational geometries used by different neural networks to solve the same task. Yet, because existing methods compare the , they may fail to capture subtle yet crucial distinctions between fundamentally different computations performed by biological and artificial neural networks. Here, we introduce metric similarity analysis (MSA), a novel method which leverages tools from Riemannian geometry to compare the intrinsic geometry of neural manifolds. Within our mathematical framework, we derive several properties of MSA, including coordinate, scale and rotation invariances. We show that MSA can be used to i) disentangle neural computations of deep networks in rich vs. lazy learning regimes, ii) compare the computation-through-dynamics of recurrent neural networks and state-space models during neuroscience working memory tasks, and iii) investigate the statistical manifolds of diffusion models. Overall, we introduce a mathematically grounded and broadly applicable framework to understand the mechanisms behind neural computations by contrasting their Riemannian geometries, with broad applicability to both neuroscience and machine learning.
PaperID: 89, Oral
Abstract: Large-scale text-to-image diffusion models can inadvertently generate harmful or sensitive content learned from uncurated training data, motivating the development of concept erasure methods that remove undesired concepts from pre-trained models. However, recent studies reveal that erased models remain vulnerable to adversarial prompt attacks that recover the supposedly removed concepts, highlighting the need for robust erasure techniques. In this work, we visualize the text embedding distributions of adversarial prompts and find that the difficulty of erasure is non-uniform. Existing methods effectively erase concepts in easier regions but leave localized residual holes in harder regions, which adversarial prompts consistently exploit. Motivated by this finding, we propose Entire-space Robust Erasure (EntiRE), an end-to-end framework that casts robust concept erasure as an out-of-distribution generalization problem. By adopting an invariant learning formulation regularized by total variation, EntiRE enforces more uniform erasure across the continuous prompt space, suppressing the localized residual regions. Extensive experiments on erasing nudity, artistic style, and object concepts demonstrate that EntiRE consistently outperforms state-of-the-art baselines in both robustness against adversarial prompt attacks and utility preservation.
PaperID: 90, Oral
Abstract: Diffusion language models (DLMs) generate text through iterative denoising, enabling parallel decoding, bidirectional conditioning, and editable intermediate states. Under harmful prompts, however, intermediate denoising states can contain risk-inducing tokens that are seemingly benign but increase the probability of an unsafe final response by shaping how the remaining masked positions are completed. Since committed tokens in DLMs can still be re-masked and rewritten, safety control can be applied during denoising by removing a small set of risk-inducing tokens before they drive subsequent generation toward unsafe responses. Existing DLM defenses mainly rely on model-level alignment, trajectory-level detection, or block-level repair, often requiring training cost, extra inference passes, or coarse regeneration. We propose CURE, a counterfactual value-driven test-time alignment method that keeps the base DLM fixed and selectively re-masks risk-inducing tokens. CURE trains a time-conditioned safety value model to estimate whether a partially denoised state will lead to an unsafe final response. During inference, CURE constructs counterfactual readout views that remove or isolate targeted tokens, estimates their contribution to future unsafe response, and re-masks only high-risk tokens for later rewriting. Across three DLMs and five jailbreak benchmarks, CURE reduces macro-average ASR from 47.20% to 1.36%, preserves utility on MMLU and GSM8K, and adds only 0.84 extra TFLOPs, achieving a strong safety-utility-efficiency trade-off. Code is available at \urlhttps://anonymous.4open.science/r/CURE.
PaperID: 91, Oral
Abstract: The Lipschitz constant of neural networks is widely used in robustness, control, generalization bounds, and training stability. As modern neural networks continue to scale, existing methods for Lipschitz estimation are either very time-consuming or overly loose. In this work, we focus on both fast and accurate Lipschitz estimation for large feedforward neural networks. We first reformulate the well-known semidefinite program (SDP) framework into an explicit recursive form, and introduce a key theoretical result, bound backpropagation: the SDP objective has upper and lower bounds that can be propagated backward along the recursive computation graph and depend only on partial variables. Leveraging this property, we decompose the original SDP into a sequence of small subproblems, each minimizing a combination of upper and lower bounds over partial variables. We prove that each subproblem is one-dimensional and convex, enabling fast solution via the secant method. Finally, we present the full FLIP algorithm and its complexity, yielding fast and tight Lipschitz estimation for large feedforward networks. Experiments on both randomly generated and trained networks show that FLIP improves speed and tightness by several orders of magnitude over prior methods, giving the tightest Lipschitz estimates for 100M-parameter networks in 3 seconds.
PaperID: 92, Oral
Authors:
Zhimin Mei, Gezheng Xu, Mohammad Hossein Moslemi, Ruiyi Fang, Nima Hosseini Dashtbayaz, Ziyan Wang, Jincheng Zhou, Haolin Jiang, Song Tang, Charles Ling, Boyu WangAbstract: Source-free domain adaptation (SFDA) adapts a pre-trained source model to an unlabeled target domain without access to source data. However, existing SFDA studies have predominantly focused on classification tasks, leaving time series forecasting (TSF), a fundamentally different problem with continuous outputs and complex temporal dependencies, largely unexplored. In this work, we identify unique characteristics and key challenges of SFDA in TSF: 1) forecasting labels correspond to future temporal segments of the same underlying signal, enabling us to split target input and construct a self-supervised learning task, thereby training an accurate short-term forecasting model to provide high-quality pseudo-labels; 2) however, such supervision remains inherently for long-horizon source model adaptation. To fill this gap, we propose the notion of Spectral Energy Allocation (SEA) pattern, defined over a short-long signal pair. The SEA pattern provides a structured correspondence that bridges signals across different temporal horizons, enabling reliable short-horizon pseudo-labels to guide long-horizon adaptation. Extensive experiments demonstrate that our method substantially outperforms baselines across different forecasting backbones.
PaperID: 93, Oral
Authors:
Deyuan (Mike) He, Ankush Desai, Sharad Malik, Aarti GuptaAbstract: LLM agents that execute multi-step tool calls must satisfy two objectives simultaneously: completing the user's task (utility) and conforming to domain policies and correctness constraints (safety). These objectives are in tension -- blocking an unsafe action prevents a violation but can leave the agent stranded with no principled way to recover. Existing approaches sacrifice one objective for the other. We introduce Soteria, a framework that reconciles safety and utility through verified hierarchical planning and runtime enforcement of specifications. Before execution, the agent generates a structured plan that is formally verified against the specifications prior to any tool invocation. During execution, the verified plan serves dual roles: it ensures that the agent takes specification-conformant trajectories and provides guidance when unsafe actions are blocked. Across multiple benchmarks covering a diverse range of tool-use tasks, Soteria achieves perfect specification conformance while improving utility by up to 5× over existing guardrail approaches, demonstrating that safety and utility need not be traded off.
Abstract: Modern distributed optimization methods mostly rely on traditional synchronous approaches, despite substantial recent progress in asynchronous optimization. We revisit Synchronous SGD and its robust variant, called m-Synchronous SGD, which can be interpreted as a method with a backup-worker strategy, and theoretically show that they are nearly optimal in many heterogeneous computation and communication scenarios, which is somewhat unexpected. We analyze the synchronous methods under random times and adversarial partial participation of workers, and prove that their time complexities are optimal in many practical regimes, up to logarithmic factors. While synchronous methods are not universal solutions and there exist tasks where asynchronous methods may be necessary, we show that they are sufficient for many modern heterogeneous computation and communication scenarios.
Authors:
Zhenyu Liang, Xiao Zhang, Tianchao Li, Jack C Cheng, Chi-Keung TangAbstract: Transmissive scenes are ubiquitous in daily life, yet reconstructing and rendering them remains highly challenging due to the inherent entanglement between near-field reflections from the surrounding environment on the transmissive surface, and the transmitted content of the scene behind it. This coupling gives rise to dual surface geometries and dual radiance components within each observation, posing ambiguities for standard methods. We present TransmissiveGS, a novel framework for disentangled reconstruction and rendering of transmissive scenes. Specifically, we model the scene with a dual-Gaussian representation and introduce a deferred shading function to jointly render the two Gaussian components. To separate reflection and transmission, we exploit the inherent multi-view inconsistency of reflections and leverage the residuals from reconstructing multi-view consistent content as cues for disentangled geometry and appearance modeling. We further propose a reflection light field that enables high-fidelity estimation of near-field reflections. During training, we introduce a high-frequency regularization to preserve fine details. We also contribute a new synthetic dataset for evaluating transmissive surface reconstruction. Experiments on both synthetic and real-world scenes demonstrate that TransmissiveGS consistently outperforms prior Gaussian Splatting-based methods in both reconstruction and rendering quality for transmissive scenes.
PaperID: 96, Oral
Authors:
Linru Zhang, Jun J Sim, Xiangning Wang, Jiahao Zhong, Kaiyu Zhou, Huanyi Ye, Yongsen Zheng, Lushan Song, Xiaojian Liang, Yingting Liu, Yujing Sun, HUAXIONG WANG, Pu Duan, Kwok-Yan LamAbstract: Fully Homomorphic Encryption (FHE) enables non-interactive, privacy-preserving transformer fine-tuning, protecting highly sensitive data as inputs to large language models. Prior work has predominantly pursued efficiency gains through architectural refinements, yet overall performance remains bounded by the substantial cost inherent in encrypted computation. To address these issues, we target the bottleneck at the ciphertext operation level. Concretely, we propose a novel packing strategy that supports batching over a flexible number of inputs, making it well-suited for fine-tuning workloads. In addition, we design operation-specific matrix evaluation algorithms that significantly reduce (i) the number of ciphertext rotations and relinearizations, and (ii) the number of bootstrapping invocations. Additionally, we exploit functional bootstrapping to fold the evaluation of element-wise activation functions directly into the bootstrapping step, thereby reducing overall computational complexity. As a result, our framework achieves a 30× reduction in ciphertext rotations and relinearizations across all matrix operations, as well as a 2.5× reduction in FHE bootstrapping invocations over the entire computation flow, compared to the latest state-of-the-art work. It is worth noting that our system completes fine-tuning of a 2-layer BERT-style model with a batch size of 16 in 650 seconds, achieving a 8× speedup over the state-of-the-art work.
Abstract: As Embodied Agents System (EAS) move into physical domains such as autonomous navigation and household assistance, reliable execution becomes critical for physical task completion and collaboration trust. Recent EAS reliability evaluation works typically focus on outcome-centric metrics as success rate or safety-only scores to assess an agent's performance, which collapse diverse execution trajectories into coarse outcomes. Therefore they ignore a critical property of EAS -- which we define as the Resilience -- that reflects how EASs recover, stabilize, and extend under perturbations and across iterative updates. The lack of resilience is particularly critical in open-world environments due to continuous unexpected disruptions, thus directly affecting the quality of EAS deployment. To address this problem, we gain insight from the resilience-engineering concepts to EAS groundings and propose a novel resilience evaluation framework that can be flexibly applied to any EAS. Specifically, we define the first comprehensive resilience metrics suite for EASs system that exposes Rebound, Stability, and Graceful Extensibility across embodied tasks execution, providing a practical grounding for EAS resilience analysis. We further implement the resilience evaluation layer that transforms execution evidence into assessments for comparison, diagnosis and optimization. Across 400 household tasks with 10 representative EAS baselines, our evaluation reveals process-level distinction hidden by outcome metrics, including recovery cost differences among successful episodes (\Delta C_\mathrmrec=25.2), increased instability under semantic perturbations, and task-family-specific degradation under stress. Metrics-guided optimizations reduce recovery cost by 42.94%, increase stability by 19.87%, and improve graceful extensibility completion by 10.08%, showing the diagnostic effect of resilience evaluation. Beyond these targeted gains, our results reveal trade-offs among resilience characteristics, suggesting that a resilient EAS construction should be evaluated and configured according to deployment-specific requirements. Our code and data are available at .
PaperID: 98, Oral
Authors: Yunchan Jeon, SOO GON KIM, Seungwook Kim, CheolWon Lee, Jongmin Lee
Abstract: Recovering 3D pose and shape of birds from single images is bottlenecked by the long tail of avian diversity: out of ~ 10,000 species, only a few dozen have 3D-labeled training data, and in-the-wild test sets routinely contain species never seen at training. Generalizing across clades is therefore a load-bearing problem. The current state of the art, AniMer+, mitigates it with a family-level contrastive loss that conveys only a binary same-family indicator per pair and discards the continuous structure of the species tree. We propose a taxonomic Ornstein--Uhlenbeck prior on the AVES shape code: a multivariate Gaussian whose covariance decays exponentially with Linnaean rank distance, with a learnable selection strength and noise scale. The prior generalizes family contrastive into a continuous, generative form, exposing the graded cross- and within-family similarity that a binary indicator cannot, and brings continuous comparative phylogenetics in as a biology-grounded inductive bias for the long tail. Our method consistently improves over AniMer+ on every reported avian benchmark cell, with the largest gains on the unseen-species CowBird benchmark. A permutation ablation isolates taxonomic structure from generic Gaussian smoothing, and a rank-granularity ablation identifies cross-order distinctions—the structure beyond a family-vs-not indicator—as the dominant source of the unseen-species gain.
PaperID: 99, Oral
Abstract: Current brain–machine interfaces (BMI) face fundamental limitations due to inherent latency from reliance on delayed motor cortical signals, and computational overhead, restricting their effectiveness in real-time applications such as rehabilitation therapy. Recent neuroscience indicates that prefrontal and sensory cortical activities precede motor execution, offering an opportunity for proactive intent prediction. However, challenges remain in acquiring multi-region neural data, efficiently decoding high-dimensional signals, and ensuring model interpretability. To address these, we developed a high-density electrocorticography (ECoG)-based paradigm based on marmosets and introduced an information-bottleneck-driven graph transformer (ECoG-IBGT), reframing neural decoding as graph classification. Our method achieves 99.29% accuracy up to 400 ms before action onset with inherent interpretability, laying the foundation for reliable, low-latency BMIs. The source code is available at https://anonymous.4open.science/r/ecog-ibgt-code-F2D2.
Abstract: 3D animal reconstruction in the wild remains challenging due to large species variation, frequent occlusions, and the prevalence of multi-animal scenes, while existing methods predominantly focus on single-animal settings. We present SAM 3D Animal, the first promptable framework for multi-animal 3D reconstruction from a single image. Built on the SMAL+ parametric animal model, our method jointly reconstructs multiple instances and supports flexible prompts in the form of keypoints and masks which enable more reliable disambiguation in crowded and occluded scenes. To train such a model, we further introduce Herd3D, a multi-animal 3D dataset containing over 5K images, designed to increase diversity in species, interactions, and occlusion patterns. Experiments on the Animal3D, APTv2, and Animal Kingdom datasets show that our framework achieves state-of-the-art results over both existing model-based and model-free methods, demonstrating a scalable and effective solution for prompt-driven animal 3D reconstruction in the wild.
PaperID: 101, Oral
Abstract: Multi-view video recordings are increasingly used to capture the 3D movements of animals in experimental settings, yet extracting rich 3D representations from these recordings remains challenging. Supervised pose estimation requires extensive manual annotation, while general-purpose 3D reconstruction models trained on generic scene datasets fail on the specialized imagery and sparse-view setting of laboratory experiments. We address these limitations with BEAST3D, a self-supervised pretraining framework that learns 3D visual representations from unlabeled, calibrated multi-view video. BEAST3D uses a vision transformer to predict 3D Gaussian splats that reconstruct held-out views through differentiable rendering, while simultaneously segmenting the animal from the background. BEAST3D reconstructs 3D structure with as few as four views by conditioning directly on known camera parameters---unlike general-purpose models, which must estimate camera geometry from dense overlapping viewpoints that are seldom available in lab settings. Through comprehensive evaluation across four species, we demonstrate that BEAST3D produces rich, viewpoint-invariant features that transfer effectively to three downstream tasks: novel view synthesis, which validates the quality of the learned 3D representations; multi-view pose estimation, which provides the sparse keypoint trajectories widely used in behavioral analysis; and neural encoding, which relates 3D behavioral features to simultaneously recorded neural activity. BEAST3D thus establishes a versatile framework for behavioral analysis that leverages 3D structure in modern multi-view laboratory recordings.
Abstract: Unified models (UMs) hold promise for their ability to understand and generate content across heterogeneous modalities. Compared to merely generating visual content, the use of UMs for interleaved cross-modal reasoning is more promising and valuable, e.g., for solving understanding problems that require dense visual thinking, improving visual generation through self-reflection, or modeling visual dynamics of the physical world guided by stepwise action interventions. However, existing UMs necessitate pixel decoding as a bridge due to their disjoint visual representations for understanding and generation, which is both ineffective and computationally cumbersome. In this paper, we introduce LatentUM, a novel unified model that represents all modalities within a shared semantic latent space, eliminating pixel-space mediation for model-internal reasoning over generated visual states. By making generated visual tokens directly interpretable to the model itself, LatentUM enables flexible interleaved cross-modal reasoning and generation. Empirically, LatentUM achieves state-of-the-art performance on the Visual Spatial Planning benchmark, pushes the limits of visual generation through self-reflection and demonstrates world modeling capability by predicting future visual states in the shared semantic latent space.
Abstract: Self-play red team is an established approach to improving AI safety in which different instances of the same model play attacker and defender roles in a zero-sum game, i.e., where the attacker tries to jailbreak the defender; if self-play converges to a Nash equilibrium, the model is guaranteed to respond safely within the settings of the game. Although the parameter sharing enforced by the use of the same model for the two roles improves stability and performance, it introduces fundamental theoretical and architectural limitations. We show that the set of Nash equilibria that can be reached corresponds to a broad class of behaviours that includes trivial always refuse strategies and oracle-like defenders, thus limiting practical applicability. We then show that when attacker and defender share and update the same base model, the dynamics collapse to self-consistency, so that attacks do not enforce adversarial pressure on the defender. In response, we propose Anchored Bipolicy Self-Play, which trains distinct role-specific LoRA adapters on top of a frozen base model, thereby maintaining stable optimisation while preserving adversarial pressure through explicit role separation. In relation to standard self-play, we show up to 100x greater parameter efficiency than fine-tuning and consistent improvements in safety compared to self-play fine-tuned models. We evaluate on Qwen2.5-3B, 7B,14B-IT models across widely used safety benchmarks, showing improved robustness without loss of reasoning ability. Cross-play experiments further show that our attacker and defender models are superior to self-play in terms of adversarial defence and safety.
PaperID: 104, Oral
Abstract: Post-training quantization (PTQ) is widely used to deploy large language models efficiently, but its effect on reasoning models is not well understood. Across math, coding, and science QA, we find that aggressive PTQ reduces accuracy while increasing chain-of-thought (CoT) length. Surprisingly, we show that in up to 52% of the quantized models' failures, models reach the right answer in intermediate reasoning steps but do not output it as a final answer. To understand why quantization leads to this increase in overthinking errors, we measure the token-level KL divergence between quantized and full-precision output distributions. Positions with high KL divergence correlate strongly with high next-token entropy, and at these positions quantized models disproportionately sample overthinking markers such as “wait", “but", and “alternatively". We show that simply introducing a training-free logit penalty on a curated set of overthinking markers can reduce CoT length by 12--23% while preserving or improving accuracy across 5 models (1.5B--32B parameters), 3 quantization methods, and 5 benchmarks, yielding the best Pareto frontier of accuracy against reasoning cost. Overthinking errors produced by quantized models are particularly reduced by up to 58%.
Authors:
Yuyang Ji, Yixuan Shen, Anil Jain, Xiaoming Liu, Feng LiuAbstract: Digital entities such as AI agents and humanoid robots increasingly operate alongside real humans, yet their identity infrastructure is based on credentials rather than embodied biometric identity. We introduce Biometric Identity Provisioning (BIP), a new problem and solution framework that addresses: given an enrollment gallery of real human identities, provision virtual identities that are non-colliding with every enrolled identity, maintain sufficient inter-class separability, and are realizable as high-fidelity face images. The key geometric insight is that real face identities occupy a low-dimensional subspace of the embedding hypersphere, leaving no residual subspace for virtual identities. Hence, virtual identities must instead be allocated as unclaimed gaps within the real face manifold itself. BIP is therefore a constrained packing problem: available gaps vastly exceed any foreseeable enrollment scale, and provisioned identities remain non-colliding even as new real identities are subsequently enrolled. Grounded in this geometry, our repulsion-based allocation is not bounded by any fixed provisioning count; we demonstrate 10M non-colliding virtual identity embeddings against a gallery of 360K real identities. Realizing these embeddings as face images requires a generator that operates outside the training distribution of real face images; we introduce GapGen, a gap-aware generator trained with a curriculum that progressively extends synthesis into non-colliding regions, validated at 1M photorealistic virtual face images. We further construct v-LFW, a virtual counterpart to LFW face dataset, with protocols for virtual face verification, cross-reality matching, real-vs-virtual detection, and unified recognition and detection.
PaperID: 106, Oral
Abstract: Understanding how neural population responses represent sensory information is a central problem in systems neuroscience. One approach is to define a representational geometry on stimulus space in which distances reflect how reliably stimuli can be distinguished from neural activity. However, different constructions of these distances can lead to qualitatively different conclusions about the neural code. Here, we show that a unique Riemannian representational geometry emerges from first principles governing how distances contract as stimulus resolution is lost through coarse-graining. This results in a multi-scale extension of the Fisher information metric, capturing encoding structure from fine stimulus details to coarse global distinctions. The resulting geometry is exactly related to the mutual information encoded by the population: well encoded stimulus directions -- those contributing more to mutual information -- are expanded, whereas poorly encoded directions are contracted. The metric tensor can be estimated using diffusion models, making the framework practical for large neural populations and high-dimensional stimuli. Applied to visual cortical responses to natural images, the eigenvectors of the metric tensor identify stimulus variations that contribute most to information transmission, yielding interpretable features that are robust to modelling choices. Together, these results provide a principled, information-theoretic framework for characterising neural population codes.
Abstract: Activation monitoring, which probes a model's internal states using lightweight classifiers, is an emerging tool for AI safety. However, its worst-case robustness under a misalignment threat model, where a model might learn to actively conceal its internal states, remains untested. We ask: could a model learn to evade previously unseen activation monitors? We demonstrate that fine-tuning can create : models capable of zero-shot evading activation monitors. Specifically, we fine-tune an LLM to evade monitors for benign concepts (e.g., HTML) when conditioned on a trigger phrase "You are being probed for concept". This learned mechanism generalizes zero-shot: substituting the concept with a safety-relevant term like 'deception' causes the model to successfully evade previously unseen safety monitors, even those trained post hoc on no-trigger activations from the frozen fine-tuned checkpoint. We validate this across diverse model families (Llama, Gemma, Qwen), finding that evasion is highly selective to the triggered concept and incurs minimal capability degradation. Mechanistic analyses suggest that the trigger induces an activation shift that moves representations away from probe-aligned directions. Our work provides a proof-of-concept for this failure mode of activation monitoring under misalignment threat models.
PaperID: 108, Oral
Abstract: AI-generated videos are becoming increasingly realistic, yet most detection methods focus on spatial artifacts within individual frames and require synthetic training data from known generators, causing them to fail on unseen generators. We identify two previously unexplored temporally distributed forensic microstructures that reveal whether a video was AI-generated. These traces arise from how a video's visual field and motion evolve over time, capturing subtle deviations that generators introduce when approximating real-world temporal dynamics. To expose them, we train two simple predictors on real video only, modeling visual evolution (VE-FSD) and motion evolution (ME-FSD), and analyze their prediction residuals. We fit 3D spatio-temporal autoregressive models to these residuals, producing compact descriptors we call Video Forensic Self-Descriptions (VFSD). VFSDs of real videos naturally cluster together and separate from AI-generated ones, enabling zero-shot detection and source attribution without synthetic training data. Experiments across multiple datasets and 35 generators show VFSD achieves state-of-the-art performance on both tasks.
PaperID: 109, Oral
Abstract: Recent mechanistic work has uncovered learned algorithms within neural networks, from modular arithmetic to search and planning in game-playing agents. But does algorithmic structure guarantee algorithmic behavior? We investigate this in Leela Chess Zero, the strongest neural chess engine, where prior work identified learned look-ahead. By extending the logit lens to its move-selecting policy network, we discover that correct puzzle solutions—including immediate checkmates—often appear in intermediate layers but are systematically overridden in the final output, a phenomenon we term . Replicating prior analyses on these positions, we find that look-ahead operates normally—future moves of the correct continuation are represented, causally important, and linearly decodable—ruling out a failure of the algorithm itself. Instead, late layers increasingly shift toward prioritizing safe play over aggression. To test whether this shift drives the override, we steer the model against these preferences and recover 61.7% of forgotten puzzles, providing causal evidence that safety priors override algorithmically computed solutions. These findings demonstrate that algorithmic structure does not guarantee algorithmic behavior: a model can internally solve a problem and still output the wrong answer.
PaperID: 110, Oral
Abstract: Despite the promise of Neural Combinatorial Optimization (NCO), recent literature has increasingly coupled with specific backbone architectures, limiting generic model design and fair cross-paradigm evaluation. We propose UECO, a simple yet effective ptimization with a structure-aware attention mixture scheme, where a lightweight module defined upon graph-based locality is interleaved with conventional (global) node-focused attention blocks to dynamically evolve edge representations and inject local topological information into attention scores. Instead of fusing static edge features in a single step, UECO thereby integrates local and global messages in a progressive and problem-agnostic manner. UECO is orthogonal to task-specific decoders, and works seamlessly with (AE) paradigms, across node- or edge-oriented problems, symmetric or asymmetric tasks, and sparse or dense graphs, for supervised, reinforcement, or generative learning. Extensive experiments show that, via plug-in substitution of encoders, UECO consistently improves solution and generalization quality over diverse backbones and task scales on TSP, ATSP, CVRP, ACVRP, MIS, MaxClique, and MaxCut.
PaperID: 111, Oral
Authors:
Xuekang Zhu, Kaiwen Feng, Ruifeng Wang, Xiwen Wang, Xiaochen Ma, Bo Du, Changjiang Jiang, Chenfan Qu, Songyu Ye, Jian liu, Ji-Zhe ZhouAbstract: Image Manipulation Localization (IML) is commonly formulated as a fully supervised learning task that estimates the optimal manipulation mask y for a given image x. In this work, we first reveal the latent nature of artifacts and thus reinterpret IML as a latent-variable problem, P(y| x)=\int P(y|z) P(z|x)dz, where z denotes the artifacts. Following this interpretation, we pinpoint the cause for the current IML models' insufficiency as their implicit artifacts modeling strategy, highlighting the necessity of modeling z in an explicit manner. Without direct labels, feature disentanglement is the most appropriate solution for this explicit modeling. Accordingly, we propose a two-stage learning paradigm with the Pairwise Artifacts Learning (PAL) and Standard Localization (SL) phases to estimate P(z|x) and P (y|z) via edit relations. To support our edit-relation-based learning, we further curate EditGroup-45K, a source-anchored dataset organized into edit groups for pair construction. Extensive experiments show that our PAL paradigm yields consistent improvements across diverse IML architectures, and empirical analyses further verify that PAL does capture artifacts explicitly through feature disentanglement. Code and dataset will be publicly available.
PaperID: 112, Oral
Abstract: The task of visual reasoning relative pose identification (VRRPI) has exposed critical weaknesses in the multi-view 3D spatial reasoning capabilities of modern vision-language models (VLMs). We present a systematic analysis of VLMs behavior in VRRPI, uncovering stable output biases where VLMs disproportionately predict specific motion patterns regardless of visual evidence. To mitigate these inherent biases, we propose self knowledge re-expression with iterative debiasing (SKR-ID), an annotation-free pipeline that adapts VLMs for VRRPI using only randomly extracted unannotated image pairs. SKR-ID transitions VLMs' output mechanism from generic next-token prediction to task-optimized classification heads. The pipeline proceeds in two stages: (1) a cold-start initialization using debiasing object-centric prompts to anchor motion inference in 2D visual evidence; and (2) a recursive logit-ranking mechanism that enforces statistical parity by using the median logit value as a dynamic threshold. Experimental evaluations across frontier VLMs on standard benchmarks demonstrate that SKR-ID significantly mitigates systematic biases and unearths latent 3D reasoning ability. Our method achieves substantial gains, improving open-sourced VLMs' Macro-F1 scores by at least 8.3% and at most 23.0%, proving that VLMs possess significant untapped multi-view 3D reasoning potential.
Abstract: Self-improving AI systems aim to reduce reliance on human engineering by learning to improve their own learning and problem-solving processes. Existing approaches to self-improvement rely on fixed, handcrafted meta-level mechanisms, fundamentally limiting how fast such systems can improve. The Darwin G\"odel Machine (DGM) achieves open-ended self-improvement in coding, but this relies on domain-specific alignment between task performance and self-modification skill. However, this alignment does not generally hold beyond coding domains. We introduce hyperagents, self-referential agents that integrate a task agent (which solves the target task) and a meta agent (which modifies itself and the task agent) into a single editable program. Crucially, the meta-level modification procedure is itself editable, enabling metacognitive self-modification, improving not only the task-solving behavior, but also the mechanism that generates future improvements. We instantiate this framework by extending DGM to create DGM-Hyperagents (DGM-H), eliminating the need for domain-specific alignment to potentially support self-accelerating progress on any computable task. Across diverse domains, the DGM-H improves performance over time and outperforms baselines without self-improvement or open-ended exploration, as well as prior self-improving systems. Furthermore, the DGM-H improves the process by which it generates new agents (e.g., persistent memory, performance tracking), and these meta-level improvements transfer across domains and accumulate across runs. All experiments were conducted with safety precautions (e.g., sandboxing, human oversight). We discuss what safety entails in this setting and the broader implications of self-improving systems. DGM-Hyperagents offer a glimpse of open-ended AI systems that do not merely search for better solutions, but continually improve their search for how to improve.
Abstract: Reinforcement learning is the standard way to post-train foundation models, but progress is bottlenecked by exploration. The usual approach perturbs sampled actions, which only moves the policy locally, yet hard tasks often require strategies very different from what the policy currently does. This is especially problematic when fine-tuning starts from a weak policy with low success rates and little useful behavior to refine. Many foundation-model policies, however, are conditioned on a natural-language prompt, and varying the prompt produces qualitatively different behaviors from the same policy. This suggests a different angle: perturb the prompt instead of the action. Our key insight is that a distribution over prompts induces a distribution over policies, with structure already shaped by the pretrained language prior. This turns exploration into a familiar problem: posterior sampling, where we maintain a belief over which prompts elicit success, sample from it to drive rollouts, and update it from what we observe. A vision-language model fits this role naturally, both proposing plausible prompts and reasoning about past rollouts to shift the distribution toward ones more likely to succeed. We call this Prompt-Driven Exploration (PDE). On hard tasks where action-space exploration fails to make progress, PDE bootstraps from nearly zero success to high success rates with far fewer environment interactions.
Abstract: We present the first method for intrinsic, continuous gradient-based optimization for models with p-adic parameters. To overcome the vanishing gradients caused by the discrete valuation of the p-adics, we lift parameters to the Berkovich affine line, a canonical analytification of the p-adic numbers that provides a path-connected space. This enables gradient descent that respects the non-Archimedean metric with a local transition cost that is linear. By extending losses to the Berkovich line, we get piecewise linear functions, corresponding to the cell decomposition of a tropical complex; gradients are efficiently computed via max-plus operations. We prove local convergence within tropical cells for linear models on the \mathbbQ_p subtree, and demonstrate the success of our method on synthetic tasks and on the p-adic Quillian semantic benchmark [Martins, 2025], where we achieve performance equal or better than \mathbbR-valued counterparts while training in comparable time.
Authors:
Jiachen Liu, Jiaxin Pei, Jintao Huang, Chenglei Si, Ao Qu, Robert Tang, Runyu Lu, Lichang Chen, Xiaoyan Bai, Haizhong Zheng, Shirui Chen, Zhiyang Chen, Haojie Ye, Yujuan Fu, Zexue He, Zijian Jin, Zhenyu Zhang, Shangquan Sun, Maestro Harmon, Dianzhuo Wang, Qian-Ze Zhu, Jiachen Sun, Mingyuan Wu, Baoyu Zhou, Chenyu You, Shijian Lu, Yiming Qiu, Fan Lai, Yuan Yuan, Yao Li, Junyuan Hong, Ruihao Zhu, Beidi Chen, Alex Pentland, Ang Chen, Mosharaf Chowdhury, Zechen ZhangAbstract: Scientific publication compresses a branching, iterative research process into a linear narrative, discarding the majority of what was discovered along the way. This compilation imposes two structural costs: a Storytelling Tax, where failed experiments, rejected hypotheses, and the branching exploration process are discarded to fit a linear narrative; and an Engineering Tax, where the gap between reviewer-sufficient prose and agent-sufficient specification leaves critical implementation details unwritten. Tolerable for human readers, these costs become critical when AI agents must understand, reproduce, and extend published work. We introduce the Agent-Native Research Artifact (ARA), a protocol that replaces the narrative paper with an agent-executable research package structured around four layers: scientific logic, executable code with full specifications, an exploration graph that preserves the failures compilation discards, and evidence grounding every claim in raw outputs. Three mechanisms support the ecosystem: a Live Research Manager that captures decisions and dead ends during ordinary development; an ARA Compiler that translates legacy PDFs and repos into ARAs; and an ARA-native review system that automates objective checks (analogous to a grammar checker for prose) so human reviewers can focus on significance, novelty, and taste. On PaperBench and RE-Bench, ARA raises question-answering accuracy from 72.4% to 93.7% and reproduction success from 57.4% to 64.4%. On RE-Bench's five open-ended extension tasks, preserved failure traces in ARA accelerate progress, but can also constrain a capable agent from stepping outside the prior-run box depending on the agent's capabilities.
Authors: Lars van der Laan, Mark van der Laan
Abstract: Prediction-powered inference uses black-box predictions to improve estimation from small labeled samples and large unlabeled samples, but raw prediction scores are often miscalibrated and can be inefficient regression adjustments. We study semisupervised mean estimation with a black-box score, a small labeled sample, and a large unlabeled sample. AIPW and PPI give valid inference for fixed or cross-fitted scores, but their efficiency depends on how well the score predicts the outcome. We propose \emphcalibrated prediction-powered inference: post-hoc calibrate the score on the labeled sample, then average the calibrated predictions over the pooled covariate sample. The estimator requires no retraining, takes a simple plug-in form, and has an exact AIPW representation. For linear calibration, we show first-order equivalence to PPI++. For isotonic calibration, we establish asymptotic normality, valid Wald inference, and ``calibeating" guarantees: isotonic post-processing improves the score as a predictor and as a first-order regression adjustment within the monotone class, and subsequent score-only post-processing yields no additional first-order gain. We also show that the original PPI estimator is a special case of AIPW and can be inefficient when the prediction score is already accurate. Simulations, benchmark reproductions, and an LLM-evaluation application show that calibrated estimators often improve on PPI and are competitive with AIPW and PPI++.
Abstract: Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs. We argue that this overlooks a key computation performed by particularly sparsity-sensitive neurons in the MLP up and gate projections: separating similar inputs into dissimilar outputs. This suggests that effective pruning should preserve not only activations, but also pairwise output differences. We introduce a family of difference-informed pruning methods built upon this principle. Wisp is a first-order, update-free method that scores weights using input-difference norms, and Wisp+ refines this score neuronwise using the input pairs each neuron separates most strongly. Finally, Whisper is a second-order method that uses a lightly regularized difference Hessian as its reconstruction objective. Across the Llama 2 and 3.1 families, our second-order variant consistently improves language modeling performance over strong reconstruction-based baselines, while our update-free variants improve over activation-aware update-free baselines, with gains stronger in more constrained settings. The improvements over Wanda and SparseGPT extend to structured sparsity, downstream evaluations, and other model families, suggesting that preserving input differentiation is a broadly useful signal for post-training LLM sparsification.
Authors:
Seungtae Nam, Jungwoo Kim, Gyeongjin Kang, Younggeun Lee, Seungkwon Yang, Eunbyung ParkAbstract: We present 3D consistency tokens, a compact set of 3D anchors that ground cross-view correspondence in explicit geometry. While feedforward Gaussian models amortize per-scene optimization into a single forward pass, their rendering quality still falls short of optimization-based pipelines. A primary cause lies in the relatively inaccurate correspondences they recover by comparing each image token to others across views based on appearance and semantic cues, which can be ambiguous in repetitive or semantically uniform regions. The proposed consistency tokens alleviate such ambiguity by routing image tokens that look alike but originate from different parts of the scene to distinct anchors, so that cross-view matches are established through shared 3D locations rather than appearance alone. We construct the consistency tokens from point clouds produced by either Structure-from-Motion or modern geometry foundation models, both of which can be noisy or incomplete. To address this, we adopt a bidirectional cross-attention architecture in which the two token sets co-refine one another, with image tokens gaining geometric grounding from the consistency tokens and the consistency tokens being corrected by the appearance cues of the images. On DL3DV-10K and in cross-dataset evaluations on Mip-NeRF 360 and Tanks & Temples, our approach surpasses prior feedforward models and attains quality competitive with optimization-based pipelines, while preserving single-pass efficiency.
Abstract: Multi-view transformers solve 3D vision tasks in a single feed-forward pass, yet their internal mechanisms remain poorly understood. We view their residual layers as progressively updating a latent state, and study how this state evolves during 3D reconstruction. To make the state geometrically interpretable, we train layer-wise probes that decode intermediate tokens into 3D pointmaps, yielding intermediate 3D reconstructions throughout the network. This view recasts feed-forward 3D reconstruction as progressive refinement rather than a single reconstruction step, with most layers improving reconstruction. Convergence speed varies across models and sample difficulty, with slower convergence on difficult samples. We further analyze these updates and find consistent, complementary roles for single-view and cross-view attention. Cross-view attention improves multi-view alignment, while single-view attention refines intra-view geometry and compensates for points visible in only one view. Finally, we identify cross-view attention heads that specialize in correspondence matching. We show that their correspondences are jointly refined with the reconstructed geometry, and that selected heads can be partially replaced by analytical correspondences. Our analysis provides a direct view of the reconstruction process inside multi-view transformers, and suggests a first step toward replacing emergent mechanisms with analytical operations. The code used for the analysis will be made available.
Abstract: The rapid progress of software engineering agents drives interest in applying them to open-ended problems such as GPU kernel optimization. However, we find that existing coding agents converge to local optima on real-world, expert-written kernels (e.g., FlashInfer). We introduce Evolving Agent Teams, a multi-agent framework that escapes such local optima. It features (1) a central supervisor that diagnoses each local optima and restructures the roles of a team of agents; and (2) a persistent file system that comprises a strategy tree, an evolution log, and per-team artifacts, grounding each restructuring in accumulated experience. Under a matched token budget, Evolving Agent Teams consistently outperforms strong coding-agent baselines. With an extended token budget, Evolving Agent Teams generates a CUDA kernel that reaches 1.2× mean speedup over the production FlashInfer MLA paged decode kernel and wins on 36 of 47 production workloads from FlashInfer-Bench (0.86--0.99× on the remaining 11).
Abstract: LLM agents are increasingly studied as systems that can self-evolve during deployment. A representative mechanism for realizing self-evolution is memory evolution, where agents maintain evolving memories derived from previous tasks and condition future inference on their memory. Although memory evolution is intended to improve agent performance, it also creates a feedback loop in which model-generated content can persist across tasks and influence future decisions. We study this loop as a delayed attack channel for memory-evolving agents. We introduce asynchronous agentic poisoning, a provider-side threat in which a malicious provider releases a model that appears safe under fresh-context evaluation but exhibits unsafe behavior after self-evolving through memory evolution. We realize this threat by fine-tuning the released model to silently embed a hidden marker into its otherwise normal responses, causing the marker to accumulate in the agent’s memory through memory evolution. When the accumulated markers are later retrieved into the input context, the model reacts to them as a backdoor trigger and shifts to attacker-specified behavior. This creates a self-triggered transition from safety-aligned behavior to attacker-specified behavior without requiring any post-deployment attacker interaction. Experimental results show that our method increases the harmful rate from below 5% to around 90% after memory evolution, while not substantially degrading the model’s self-evolution performance. These findings demonstrate that memory evolution creates a provider-side pathway for delayed compromise, allowing a model to appear safe at release time while becoming unsafe after deployment through the agent’s self-evolution process.
Authors:
Xiangdong Zhang, Xiaohan Qin, Tuo Dai, Xiaoming Shi, Huaijin Wu, Yebin Yang, Zhuo Xia, Shaofeng Zhang, Yu Wang, Yu Cheng, Junchi YanAbstract: Hyper-Connections (HC) extend the single residual stream of Transformers into N parallel streams, improving language model pre-training with modest additional FLOPs. Manifold-Constrained HC (mHC) stabilizes this formulation at scale by constraining residual mixing to doubly stochastic matrices. Large gains from N=1 to N=4 suggest residual-stream expansion as a promising scaling axis, yet existing HC-family methods typically stop at N=4. Our experiments on mHC reveal why: directly scaling N beyond 4 yields rapidly diminishing returns, as gains saturate while training compute grows sharply. We trace this to two bottlenecks: fixed-dimensional layer outputs cannot supply enough write-back information for more streams; meanwhile, generating the residual-mixing matrix makes the dominant cost scale cubically with N. To address both bottlenecks, we propose xHC (Expanded Hyper-Connections), the first HC-family method to achieve meaningful expansion beyond N=4. xHC enriches write-back signals with local contextual features along the token sequence and uses a sparse residual-stream architecture that updates only k out of N streams while preserving dense access to all streams. In language model pre-training, xHC with N=16 and k=4 substantially outperforms both mHC and the vanilla residual baseline. On a 69B MoE model, xHC improves the average downstream score by 5.7 points over mHC (vs.\ +1.3 for mHC over vanilla) with only 2.4% additional training FLOPs over the vanilla baseline. Scaling law experiments show that xHC matches the loss of a vanilla baseline trained with 1.50× the compute, and its continued improvement with larger N makes expansion rate a practical scaling axis for HC-family models.
Abstract: We study allowing large language models (LLMs) to process arbitrarily long prompts through the lens of inference-time scaling. We propose s), a general inference paradigm that treats long prompts as part of an external environment and allows the LLM to programmatically examine, decompose, and recursively call itself over snippets of the prompt. We find that RLMs can successfully process inputs more than an order of magnitude beyond model context window limits and, even for shorter prompts, dramatically outperform the quality of vanilla frontier LLMs and common long-context scaffolds (e.g., on GPT-5 by a median across the evaluated benchmarks of 26% against compaction, 130% against CodeAct with sub-calls, and 13% against Claude Code) across four diverse long-context tasks while having comparable cost. At a small scale, we post-train the first model around the RLM. Our model, , outperforms the underlying Qwen3-8B model by a median of 28% and even approaches the quality of vanilla GPT-5 on three long-context tasks.
Abstract: Flow-based generation in high-dimensional spaces is difficult because velocity prediction requires modeling high-dimensional noise, even when data has strong low-rank structure. We present Asymmetric Flow Modeling (AsymFlow), a rank-asymmetric velocity parameterization that restricts noise prediction to a low-rank subspace while keeping data prediction full-dimensional. From this asymmetric prediction, AsymFlow analytically recovers the full-dimensional velocity without changing the network architecture or training/sampling procedures. On ImageNet 256×256, AsymFlow achieves a leading 1.57 FID, outperforming prior DiT/JiT-like pixel diffusion models by a large margin. AsymFlow also provides the first-ever route for finetuning pretrained latent flow models into pixel-space models: a patch-level linear lift initializes a low-rank pixel model whose denoising trajectory preserves the latent model's high-level semantics and structure, so finetuning mainly improves low-level residuals rather than relearning pixel generation. We show that the pixel AsymFlow model finetuned from FLUX.2 klein 9B establishes a new state of the art for pixel-space text-to-image generation, beating its latent base on HPSv3, DPG-Bench, and GenEval while qualitatively showing substantially improved visual realism. Code and models will be released publicly.
Abstract: Diffusion and flow-based models have become the de facto approaches for generating continuous data, e.g., in domains such as images and videos. Their success has attracted growing interest in applying them to language modeling. Unlike their image-domain counterparts, today’s leading diffusion language models (DLMs) primarily operate over discrete tokens. In this paper, we show that continuous DLMs can be made effective with minimal adaptation to the discrete domain. We propose Embedded Language Flows (ELF), a class of diffusion models in continuous embedding space based on continuous-time Flow Matching. Unlike existing DLMs, ELF predominantly stays within the continuous embedding space until the final time step, where it maps to discrete tokens using a shared-weight network. This formulation makes it straightforward to adapt established techniques from image-domain diffusion models, e.g., classifier-free guidance (CFG). Experiments show that ELF substantially outperforms leading discrete and continuous DLMs, achieving better generation quality with fewer sampling steps. These results suggest that ELF offers a promising path toward effective continuous DLMs.
Authors:
Sonal Kumar, Prem Seetharaman, Ke Chen, Oriol Nieto, Jiaqi Su, Zhepei Wang, Rithesh Kumar, Dinesh Manocha, Nicholas J Bryan, Zeyu Jin, Justin SalamonAbstract: Large Audio Language Models struggle to disentangle overlapping events in complex acoustic scenes, yielding temporally inconsistent captions and frequent hallucinations. We introduce Timestamped Audio Captioner (TAC), a model that produces temporally grounded audio descriptions at varying degrees of detail and resolution. TAC is trained with a synthetic data pipeline that constructs challenging and dynamic mixtures from real-world audio sources, enabling robust learning under realistic polyphonic conditions. Across event detection, dense captioning, and speaker diarization, TAC outperforms most competing methods - including dedicated diarization systems with a low hallucination rate and accurate temporal grounding. We also introduce TAC-V, an audio-visual pipeline to generate semantically rich audio-visual descriptions. We then show that TAC serves as a "semantic bridge" for a text-only reasoner: a simple TAC→LLM and TAC-V→LLM cascade achieves state-of-the-art scores on benchmarks for both audio (MMAU-Pro, MMSU, MMAR) and audio-visual (DailyOmni, VideoHolmes) understanding and reasoning respectively. We encourage readers to see detailed qualitative results on our demo page: https://audioanon.github.io/tacmodel/.
Abstract: Recent video reasoning models increasingly produce spatio-temporal evidence chains that localize objects at specific timestamps. While these traces improve interpretability by grounding where and when evidence appears, they often leave the motion connecting observations, the how, implicit. This makes dynamic and trajectory-dependent claims difficult to supervise, verify, or penalize when unsupported by the video. We formalize this missing component as Spatial-Temporal-Trajectory (STT) reasoning and introduce Motion-o, a motion-centric extension to vision-language models (VLMs) that makes trajectories explicit and verifiable. Motion-o augments evidence chains with Motion Chain of Thought (MCoT), a structured pathway that represents object motion through a discrete \texttt\ tag summarizing direction, speed, and scale change. To supervise MCoT, we densify sparse spatio-temporal annotations into object tracks and derive motion descriptors from centroid displacement and box-area change. We then train with complementary rewards for trajectory consistency and visual grounding, including a perturbation-based signal that penalizes motion descriptions that remain unchanged when temporal evidence is removed. Across multiple video understanding benchmarks, Motion-o consistently improves trajectory-faithful reasoning without architectural modifications. These results suggest that an explicit motion interface can complement existing VLM pipelines by converting implicit dynamics into verifiable evidence.
Authors: Jiawen Dai, Yue Song
Abstract: Oscillations and synchronization are widely believed to play a fundamental role in representation and computation. However, existing machine learning approaches based on synchronization dynamics have largely been confined to specialized settings such as object discovery, with limited evidence of scalability to standard vision benchmarks or logic reasoning tasks. We propose the \emphWinfree Oscillatory Neural Network (\emphWONN), a dynamical neural architecture based on generalized Winfree dynamics. \emphWONN evolves representations on the torus (S^1)^d through structured oscillatory interactions, combining phase-based inductive biases with flexible and hierarchical interaction mechanisms instantiated as either fixed trigonometric mappings or learnable neural networks. We evaluate \emphWONN on image recognition and complex reasoning tasks, including CIFAR, ImageNet, Maze-hard, and Sudoku. Across these domains, \emphWONN achieves competitive or superior performance with strong parameter efficiency. In particular, \emphWONN is, to our knowledge, the first synchronization-based oscillatory architecture to scale competitively to ImageNet-1K. Furthermore, on Maze-hard, \emphWONN achieves 80.1% accuracy using only 1% of the parameters of prior state-of-the-art models. These results suggest that structured oscillatory dynamics provide a scalable and parameter-efficient alternative to conventional neural architectures.
Abstract: This paper proposes a novel crowd counting approach, the Gaussian Density Splatting Network (GDSNet). Unlike methods that rely on conventional, grid-based density maps and are sensitive to spatial resolution, GDSNet represents a crowd as a superposition of continuous 2D Gaussian primitives. Our approach is built upon two key contributions. First, we introduce a control-point-based fitting mechanism to structure the prediction of the Gaussian parameters. We design a method to allocate a set of control points that define local regions, from which features are pooled to regress each primitive's parameters. Second, we adapt a differentiable Gaussian Splatting framework to the counting task by parameterizing each primitive with geometric parameters and a scalar density mass. This formulation allows the network to be trained end-to-end via spatial matching of differentiably rendered density maps, naturally providing both local density supervision and global count optimization. Extensive evaluations on four standard benchmarks show GDSNet consistently outperforms the state of the art. Code will be released upon acceptance.
Abstract: When multi-agent systems (MAS) fail, identifying where the decisive error occurred is the first step for automated recovery to an earlier state. Error attribution remains a fundamental challenge due to the long interaction traces that large language model-based MAS generate. This paper presents a framework for error attribution based on conformal prediction (CP) which provides finite-sample, distribution-free coverage guarantees. We introduce new algorithms for filtration-based CP designed for sequential data such as agent trajectories. Unlike existing CP algorithms, our approach predicts sets that are contiguous sequences to enable efficient recovery and debugging. We verify our theoretical guarantees on a variety of agents and datasets, show that errors can be precisely isolated, then use prediction sets to rollback MAS to correct their own errors. Our overall approach is model-agnostic, and offers a principled uncertainty layer for MAS error attribution.
Abstract: Over the past few years, diffusion-based Schrödinger bridge models have been proposed to approximate optimal transport dynamics between two prescribed boundary distributions, with successful applications to generative modeling. More precisely, these methods aim to estimate a path measure whose initial and terminal marginals match the two boundary distributions, while minimizing the Kullback–Leibler divergence with respect to a reference Markov process. In this work, we consider the generalized Schrödinger bridge problem, in which the reference process is a twisted Brownian motion, that is, a Feynman–Kac transform of a Brownian motion induced by a time-dependent differentiable potential. Building on the Iterative Markovian Fitting (IMF) paradigm, and in particular on its special case Diffusion Schrödinger Bridge Matching (DSBM), which corresponds to the zero potential, we introduce Twisted Schrödinger Bridge Matching (TSBM), a diffusion-based method designed to handle both continuous- and discrete-time potentials. Unlike previous approaches, TSBM provides a rigorous extension of the IMF scheme to the generalized Schrödinger bridge problem. This derivation leads to a new bridge-matching loss that depends explicitly on the gradient of the potential and recovers the DSBM objective when the potential vanishes, yielding improved performance. We further introduce trajectory-based variance-reduction techniques that substantially stabilize optimization and may be useful beyond the present setting. Finally, we empirically demonstrate the benefits of TSBM on several experimental settings.
Abstract: Diffusion Language Models (DLMs) offer attractive advantages over Auto-Regressive (AR) models, such as full-attention parallel decoding and flexible generation. However, standard DLM training uses a static, single-step masked prediction objective that never exposes the model to the progressive denoising dynamics of inference, and forces all contextual information to be maintained purely through token-space attention, which becomes increasingly diluted as context length grows. We propose MemDLM (Memory-Enhanced DLM), which introduces a second memory channel by embedding a simulated denoising trajectory into training via Bi-level Optimization. An inner loop updates a set of fast weights, forming a Parametric Memory that captures the local trajectory experience, while an outer loop updates the base model conditioned on this memory. By offloading part of the memorization burden from token-space attention to parameter space, MemDLM yields faster convergence, stronger long-context representations, and lower training loss, even when the fast weights are discarded at inference time. Re-enabling the inner loop at inference provides an additional prompt-specific adaptation effect, where the Parametric Memory acts as an emergent in-weight retrieval mechanism on challenging Needle-in-a-Haystack tasks.
Abstract: While Diffusion Language Models (DLMs) are theoretically well-suited for iterative refinement due to their non-causal structure, off-the-shelf masked diffusion language models (MDLMs) often fail to reliably revise incorrect tokens in practice. The key challenge lies in the model's inability to distinguish between correct and erroneous tokens in a visible sequence. Standard MDLM training is restricted to the objective of unmasking, undermining the effectiveness of refinement guided by confidence. Based on this observation, we study corrective behavior in DLMs, defined as the ability to assign lower confidence to incorrect tokens and iteratively refine them while preserving correct content. Under matched compute and data, we show that this capability is not induced by conventional masked diffusion objectives and propose a post-training principle oriented by correction that explicitly supervises visible incorrect tokens, enabling discriminative confidence and targeted refinement. To evaluate corrective behavior, we introduce the Code Revision Benchmark, a controllable and executable benchmark for assessing error localization and in-place correction. Beyond improving in-place correction, our post-training principle yields stronger general generation as a direct consequence: the same error-aware confidence that drives reliable revision also makes parallel decoding more effective and improves answer quality on standard generation tasks. We validate this dual benefit across three distinct domains, code completion, mathematical reasoning, and high-entropy parallel decoding (ParallelBench), where models trained with our objective consistently outperform both the base model and standard masked-diffusion fine-tuning.
Abstract: When a text description is extended with an additional detail, image-text similarity should drop if that detail is wrong. We show that CLIP-style dual encoders often violate this intuition: appending a plausible but incorrect object or relation to an otherwise correct description can increase the similarity score. We call such cases \emphhalf-truths. On the verified COCO evaluation set, CLIP prefers the correct shorter description only 41.0% of the time, and performance drops to 28.2% when the added detail is a relation. We trace this vulnerability to weak supervision on caption parts: contrastive training aligns full sentences but does not explicitly enforce that individual entities and relations are grounded. We propose CS-CLIP (Component-Supervised CLIP), which decomposes captions into entity and relation units, constructs a minimally edited foil for each unit, and fine-tunes the model to score the correct unit above its foil while preserving standard dual-encoder inference. CS-CLIP raises half-truth accuracy to 78.2% and improves average performance on established compositional benchmarks by 5.7 points, suggesting that reducing half-truth errors aligns with broader gains in compositional understanding.
Authors: Seokhoon Jeong, Mijung Kim, Taehwan Kim
Abstract: Neural architecture search (NAS) methods have grown increasingly efficient, yet they remain bounded by manually engineered search spaces that require substantial domain expertise and must be rebuilt for every new task. Large language models (LLMs) can generate architectures in an open-ended space, but how to optimally divide labor between LLM-driven design and NAS-driven search remains unexplored. We propose a mechanism that bridges these two paradigms: an LLM produces a high-quality seed architecture, then decomposes it into a ---a scaffold with named, interchangeable module slots that automatically defines a bounded, task-specific search space for conventional NAS to explore, without manual engineering. We instantiate this mechanism in , a modular three-phase pipeline in which each component's contribution can be measured independently. On 17 tasks spanning classification, dense regression, segmentation, and multi-label tagging across diverse modalities (NAS-Bench-360 and NAS-Unseen), AgentNAS establishes a new state of the art on 11 tasks, outperforming published baselines including task-specific expert designs. Ablation studies show that the two search mechanisms are broadly complementary: the LLM-generated seed already surpasses published baselines on the majority of tasks, and NAS delivers additional gains in most cases through combinatorial recombination across slots---a mode of search that independent LLM samples cannot replicate. These patterns hold across three LLMs of different capability levels, confirming that the division of labor is robust.
Abstract: Causal representation learning aims to discover robust features by exploiting the causal structure underlying data generation. Existing methods require specifying the causal structure a priori, yet different structures demand fundamentally incompatible invariance constraints, and misspecification leads to representations that discard predictive information. We introduce SaCRL, a framework that jointly identifies the causal structure and learns the corresponding invariant representation without prior structural knowledge. Our approach formulates structure selection as a soft optimization over candidate invariances using HSIC-based violation metrics, with adaptive weights that automatically concentrate on the achievable structure. We provide theoretical guarantees for structure identification, invariance satisfaction, and out-of-distribution generalization. Empirically, SaCRL matches structure-specific oracles on synthetic and Colored MNIST benchmarks, achieves state-of-the-art accuracy on three DomainBed benchmarks (PACS, VLCS, OfficeHome), and degrades gracefully under structural misspecification and limited environment diversity. Code is available at: \urlhttps://github.com/ax1375/sarl.
Abstract: Modern large language models increasingly require long contexts for reasoning and multi-document tasks, but attention's quadratic complexity creates a severe computational bottleneck. We present Block Sparse Flash Attention (BSFA), a drop-in replacement that accelerates long-context inference while preserving model quality. Unlike methods that predict importance before computing scores, BSFA computes exact query-key similarities to select the top-k most important value blocks for each query. By comparing per-block maximum scores against calibrated thresholds, we skip approximately 50% of the computation and memory transfers for pruned blocks. Our training-free approach requires only a one-time threshold calibration on a small dataset to learn the per-layer and per-head attention score distributions. We provide a CUDA kernel implementation that can be used as a drop-in replacement for FlashAttention. On Llama-3.1-8B, BSFA achieves up to 1.13× end-to-end speedup on LongBench with only a 1.1% accuracy drop, and up to 1.24× on Needle-in-a-Haystack retrieval at a 1% accuracy drop. The attention kernel itself accelerates by up to 1.38×. We compare BSFA against five recent sparse attention baselines (SpargeAttention, MInference, FlexPrefill, XAttention, and BLASST), and verify the method on Qwen2.5-7B and on A6000 and H100 GPUs. The implementation is available at https://github.com/Anonymous44414/Block-Sparse-Flash-Attention.
Abstract: The task of dataset distillation aims to find a small set of synthetic images such that training a model on them reproduces the performance of the same model trained on a much larger dataset of real samples. Existing distillation methods focus on synthesizing datasets that enable training randomly initialized models. In contrast, state-of-the-art vision approaches are increasingly building on large, pre-trained self-supervised models rather than training from scratch. In this paper, we investigate the problem of distilling datasets that enable us to optimally train linear probes on top of such large, pre-trained vision models. We introduce a method of dataset distillation for this task called Linear Gradient Matching that optimizes the synthetic images such that, when passed through a pre-trained feature extractor, they induce gradients in the linear classifier similar to those produced by the real data. Our method yields synthetic data that outperform all real-image baselines and, remarkably, generalize across pre-trained vision models, enabling us, for instance, to train a linear CLIP probe that performs competitively using a dataset distilled via a DINO backbone. Further, we show that our distilled datasets are exceptionally effective for fine-grained classification and provide a valuable tool for model interpretability, predicting, among other things, how similar two models' embedding spaces are under the platonic representation hypothesis or whether a model is sensitive to spurious correlations in adversarial datasets.
Abstract: Imagine Mr. Bean stepping into Tom and Jerry--can a video model generate interactions that stay faithful to each character's identity, behavior, and style across worlds? We present MiMix, a unified framework for multi-character text-to-video generation that goes beyond visual appearance to capture character-specific personas generalizing to unseen scenes and partners. To evaluate this, we introduce the Cross-Character Generalization Benchmark, a social stress test placing each character in novel scenes with previously uncoexistent partners. MiMix combines Cross-Character Embedding (CCE), which disentangles identity and behavior via text-aligned annotations, with Cross-Character Augmentation (CCA), which synthesizes cross-style training data while preserving native appearance. Experiments show consistent gains in identity fidelity, interaction quality, and style consistency over prior personalized and general video models.Additional results and videos are available on our project page: https://mi-mi-x.github.io.
Authors:
Jay Yoo, Kazuma Kobayashi, Jaewan Park, Sai puppala, Souvik Chakraborty, Syed B AlamAbstract: Neural operators have emerged as powerful surrogates for scientific computing, enabling rapid reconstruction of complex physical fields from sparse sensor observations. Despite their strong predictive capability, their reliability often degrades under noisy or corrupted measurements, limiting practical deployment in real-world sensing environments. We show that enforcing statistical independence in latent representations provides a simple yet effective mechanism for improving neural operator robustness. To this end, we introduce Neural Operator Independence Regularization (NOIR), an end-to-end training framework that promotes statistically independent latent features during operator learning. Unlike structured basis approaches such as POD-DeepONet and PCA-Net, which impose orthogonality primarily as an offline compression constraint, NOIR directly shapes latent representations during training to enhance resilience against input perturbations. Across three canonical PDE benchmarks and five additive sensor corruption settings, NOIR consistently improves reconstruction accuracy while preserving clean-data performance. Beyond predictive gains, the induced latent representations exhibit distinctive statistical signatures that reveal architecture-specific robustness characteristics under corruption. These findings identify latent statistical independence as a principled inductive bias for robust neural operator design and provide new insight into the internal structure of operator learning systems. All models, datasets, checkpoints, and training code are publicly available at \urlhttps://github.com/t55176853-cmyk/NOIR/
Abstract: We introduce a framework for differentiable exact learning of algorithms. The framework couples a state-tracking neural controller to an unbounded external environment and trains on policy-trajectory observations (PTOs) -- traces that record the environment's evolution under an expert's actions without exposing the expert's internal state. This supervision regime sidesteps Gold's impossibility barrier for input-output recursive learning while avoiding the unrealistic internal-state access required by prior provably correct approaches. As controllers, we propose Differentiable Finite-State Transducers (DFSTs), a minimalist multilinear model family that contains no nonlinearities and admits log-parallel training via prefix scans. Trained on tiny datasets, DFST and RNN controllers achieve error-free generalization on binary and decimal addition and multiplication to operands thousands of times longer than the training examples. Moreover, we provide strong empirical proof for exact learning of universal computation. Finally, we develop an extraction procedure that recovers compact finite-state transducers from trained controllers. We prove under a geometric clustering assumption that for binary addition the extracted transducer exactly replicates the expert policy on all valid inputs -- certifying differentiable exact learning of the algorithm.
Abstract: Joint audio-video generation aims to synthesize temporally synchronized and semantically coherent visual and acoustic content. Existing open-source methods mainly follow either dual-tower designs, which generate audio and video in separate streams and rely on posterior alignment, or fully unified tri-modal designs, which mix textual context, audio, and video in a single shared space. These paradigms may either weaken fine-grained audio-video co-evolution or couple semantic conditioning with low-level synchronization. We propose NAVA, a Native Audio-Visual Alignment framework that formulates generation as context-conditioned native audio-visual alignment. NAVA first establishes audio-video correspondence in a dedicated alignment space and then applies context as external conditioning to guide the aligned representation. We instantiate this formulation with an Align-then-Fuse MMDiT architecture, which progressively bridges modality-aware alignment and unified audio-video denoising. To support controllable speech generation, we further introduce Timbre-in-Context Conditioning, which binds reference timbre cues to corresponding speech spans through the context pathway. Experiments on Verse-Bench and the Seed-TTS benchmark demonstrate that NAVA achieves superior audio-visual synchronization and video quality, competitive audio quality, and substantially improved reference-timbre controllability with only 6.2B parameters.
Abstract: Out-of-distribution detection in continual learning (OOD-CL) is highly challenging, as it requires (i) learning a sequence of tasks without catastrophic forgetting, and (ii) identifying unknown samples (OOD) without access to known (in-distribution (ID)) data from past tasks, where standard confidence-based OOD scores tend to degrade. Existing methods largely rely on a rehearsal buffer of past data or assume closed-world settings, limiting their scalability and reliability in real-world scenarios. In this paper, we propose a novel OOD detection framework for rehearsal-free continual learning, inspired by Kolmogorov–Arnold Networks (KAN), which mitigate catastrophic forgetting via their inherent local neuroplasticity. To better leverage KAN in OOD-CL, we introduce a task-adaptive classifier architecture, \emphTA-KAN, which mitigates recency bias and enhances ID--OOD separability through task-specific control of activation locality. Furthermore, we propose a geometry-guided OOD score that complements TA-KAN classifier confidence without relying on past ID data. Our method significantly improves OOD detection while maintaining strong ID classification performance, achieving state-of-the-art results in OOD-CL.
Abstract: Classifier-free Guidance (CFG) is central to modern diffusion models, but operates as an inference-time correction where the training objective learns the unbiased conditional score, while guided sampling targets a posterior-tilted score that never explicitly learned. Rather than another inference-time patch or distillation, Model-guidance (MG) directly integrates the posterior tilt into the training objective and eliminates CFG during inference. MG serves as a plug-and-play module compatible with existing methods, yet accelerates convergence by online self-bootstrapping and doubles inference speed. Unlike distillation surrogates, a reparameterisation property reveals that MG implicitly learns the vanilla score without mode collapse. Extensive experiments demonstrate that MG matches and even surpasses CFG baselines, and achieves a state-of-the-art FID of 1.34 on ImageNet 256 benchmark.
Authors: Mahindra Rautela, Alexander Scheinker
Abstract: Large-scale scientific facilities, such as particle accelerators, are complex systems composed of thousands of interacting components that must operate collectively to support diagnostics and control, and ensure safe, reliable operation. These facilities rely on extensive sensor networks that generate large volumes of heterogeneous, incomplete, and noisy telemetry, often sampled at different rates across subsystems. In this paper, we adopt the foundation-model paradigm to learn general-purpose sensor representations from historical accelerator data through self-supervised pretraining. We introduce SensOFormer, a denoising masked transformer with a Perceiver-style encoder–decoder architecture designed to handle variable sensor sets, accommodate heterogeneous sampling rates, and learn robust representations from noisy measurements. The model is pretrained on multiple particle-accelerator datasets and transferred to downstream tasks, including missing-data imputation, anomaly detection, cavity identification, and fault identification. Through extensive evaluations and comparisons, we demonstrate that a single self-supervised model can learn reusable representations for diverse operational tasks in large-scale scientific facilities.
Abstract: Training large language models at 4-bit precision is critical for efficiency. We show that nGPT, an architecture that constrains weights and hidden representations to the unit hypersphere, is inherently more robust to low-precision arithmetic. This removes the need for interventions—such as applying random Hadamard transforms and performing per-tensor scaling calculations—to preserve model quality, and it enables stable end-to-end NVFP4 training. We validate this approach on both a 1.2B dense model and hybrid (Mamba-Transformer) MoE models of up to 3B/30B parameters. We trace this robustness to the dot product: while quantization noise remains largely uncorrelated in both standard and normalized architectures, the signal behaves differently. In nGPT, the hypersphere constraint enhances weak positive correlations among the element-wise products, leading to a constructive accumulation of the signal across the hidden dimension while the noise continues to average out. This yields a higher effective signal-to-noise ratio and a flatter loss landscape, with the effect strengthening as the hidden dimension grows, suggesting increasing advantages at scale. A reference implementation is available at https://github.com/anonymous452026/ngpt-nvfp4.
Authors:
Yohaï-Eliel BERREBY, Sabrina Du, Audrey Durand, B. S KrishnaAbstract: Active computer vision promises efficient, biologically plausible perception through sequential, localized glimpses, but lacks scalable general-purpose architectures and pretraining pipelines, leaving Active-Vision Foundation Models (AVFMs) underexplored. We introduce CanViT, the first task- and policy-agnostic AVFM. CanViT uses scene-relative RoPE to bind a retinotopic Vision Transformer backbone and a spatiotopic scene-wide latent workspace, the canvas. Efficient interaction with this high-capacity working memory is supported by Canvas Attention, a novel asymmetric cross-attention mechanism. We decouple thinking (backbone-level) and memory (canvas-level), eliminating canvas-side self-attention and fully-connected layers to achieve fast sequential inference and scalability to high output resolutions. We propose a label-free active vision pretraining scheme, policy-agnostic passive-to-active dense latent distillation: reconstructing scene-wide DINOv3 embeddings from sequences of low-resolution glimpses with randomized locations, zoom levels, and lengths. We pretrain CanViT-B from a random initialization on 13.2 million ImageNet-21k scenes--an order of magnitude more than previous active models--and 1 billion random glimpses, in 166 hours on a single H100. On ADE20K segmentation, a frozen CanViT-B achieves 38.5% mIoU in a single low-resolution glimpse, outperforming the best active model's 27.6% with 20x fewer inference FLOPs as well as its FLOP- or input-matched DINOv3 teacher. Given additional glimpses, CanViT-B reaches 45.9% ADE20K mIoU. On ImageNet-1k classification, CanViT-B also sets a new active-vision state of the art, with 84.5% top-1 accuracy after fine-tuning. CanViT generalizes to longer rollouts, larger scenes, and new policies. Our work narrows the wide gap between passive and active computer vision, demonstrating the potential of task- and policy-agnostic AVFM pretraining.
Abstract: Achieving true artificial general intelligence requires foundation models capable of integrating new modalities without forgetting prior knowledge. However, accommodating continuous generative objectives alongside discrete understanding tasks causes severe gradient conflicts. Existing architectures, including standard Mixture-of-Experts (MoE), are highly susceptible to representation overwriting. Even structurally partitioned paradigms like Mixture-of-Transformers (MoT) remain vulnerable to catastrophic forgetting, severely impeding multimodal scalability. In this work, we introduce Rosetta, a composable native multimodal pretraining framework designed for seamless and non-destructive modality expansion. Rosetta adopts a modular paradigm where core foundational knowledge is preserved within global shared experts, while modality-specific capabilities are distributed across plug-and-play experts. To guarantee non-destructive composition, we propose Momentum-Anchored Orthogonal Projection (MAOP). MAOP leverages the optimizer's momentum state as an implicit semantic anchor, selectively neutralizing conflicting gradient components from new modalities while preserving synergistic updates. To strictly isolate the architectural impact, we evaluate Rosetta against standard MoE and MoT baselines under strict active parameter parity. All models are trained from scratch within the Transfusion framework, using discrete next-token prediction for language and continuous visual diffusion. Extensive evaluations demonstrate that, while standard MoE and MoT architectures suffer catastrophic forgetting of previously acquired knowledge, Rosetta robustly preserves established language and visual understanding. Furthermore, it delivers superior image generation and unlocks cross-modal synergy, paving the way for truly composable and unified multimodal foundation models.
Abstract: Generating music that temporally aligns with video events is challenging for existing text-to-music models, which lack fine-grained temporal control. We introduce V2M-ZERO, a video-to-music generation approach that generates time-aligned3 music with disentangled time synchronization and semantic control (e.g. genre, mood) from video while requiring zero video-music pairs at training time. Our method is motivated by a key observation: temporal synchronization requires matching when and how much change occurs, not what changes. While musical and visual events differ semantically, they exhibit shared temporal structure that can be captured independently within each modality. We capture this structure through event curves computed from intra-modal similarity using pretrained music and video encoders. By measuring temporal change within each modality independently, these curves provide comparable representations across modalities. This enables a simple training strategy: fine-tune a text-to-music model on music-event curves, then substitute video-event curves at inference without cross-modal training or paired data. Across OES-Pub, MovieGenBench-Music, and AIST++, V2M-ZERO achieves state-of-the-art performance without any paired music-video data, surpassing the strongest prior baselines per metric with 5–9% higher audio quality, 13–15% better semantic alignment, 21–52% improved temporal synchronization, and 28% higher beat alignment on dance videos. We find similar results via a large crowd-source subjective listening test. Our results vali-20 date that temporal alignment through within-modality features is not only effective for video-to-music generation but leads to better performance over paired cross-modal supervision. Furthermore, our approach enables independent controls for timing and music style (e.g. genre, mood) for more controllable generation.
Abstract: Offline MARL increasingly relies on generative policies to model multimodal joint behavior, typically by distilling a centralized teacher into decentralized one-step actors under the CTDE. We identify a failure mode at the teacher training stage: standard flow-based teachers pair noise with replay targets independently, so nearby noise samples can be routed toward conflicting coordination modes. The teacher then produces samples between valid modes, and because the distillation loss regresses each local actor onto the conditional mean of the teacher's output given local input, this error is not absorbed but propagated to the student. To remove this teacher-side artifact, we propose Mode-Support Semi-Discrete Optimal Transport (MoSDOT), which summarizes multimodal replay into a finite mode support with prescribed capacities and uses conditional semi-discrete optimal transport to assign each noise sample to a single mode before teacher training. We additionally study a common-randomness variant that uses a shared noise component at execution to expose the residual gap intrinsic to strict-product execution. On controlled diagnostics and offline MARL benchmarks, MoSDOT improves endpoint quality and routing consistency, particularly on datasets exhibiting multimodal joint behavior.
Abstract: Self-evolution offers a scalable path to stronger reasoning: a pretrained language model improves itself with only minimal external supervision. Yet existing methods either depend on extensively curated or teacher-generated training data, or, when the generator runs unsupervised, reward it by a difficulty heuristic that need not improve the solver. We introduce that drafts questions and reference golden answers from a pool of unstructured, automatically collected documents, and a that improves by training on them. The solver is trained with standard correctness rewards against the generator-provided answers, while the generator is rewarded by an that measures whether each proposed question would actually improve the solver on the target distribution. Because this continuous, noisy influence score is poorly served by standard GRPO, we propose , a dual-normalized variant of GRPO, for generator training. Together, these turn the document pool into an that favors questions useful to the current solver, not just hard ones. On Qwen3-8B-Base, INFUSER outperforms strong self-evolution baselines with over 20% relative improvement on Olympiad and SuperGPQA benchmarks, and an 8B INFUSER co-evolving generator outperforms a frozen 32B thinking generator on math and coding. Ablations confirm each design choice is necessary. Code is available at
Abstract: Recent progress in multimodal large language models (MLLMs) has strengthened visual understanding and instruction following, and opened a path to open-ended generative object detection, in which a model discovers, names, and localizes objects from free-form instructions without relying on a predefined category set. Nevertheless, current generative detectors remain less accurate than discriminative detectors, as token-level supervision provides only indirect and weak constraints on fine-grained visual perception. To equip MLLM with such an explicit optimization objective, we propose GOD (Generative Object Detection), a 3B-scale MLLM that injects object-level visual-prior supervision into generative detection by coupling visual-prior reconstruction with language generation. GOD attaches a lightweight DETR-style object-query branch to the MLLM visual encoder, maps the resulting object queries into the token space, and optimizes them jointly with autoregressive generation, MLLM-side object-query reconstruction, and DETR-side localization objectives. Specifically, in the training recipe, a multi-stage training pipeline first aligns the object decoder with the MLLM and then uses an interleaved-mask curriculum to transfer detector-grounded priors from proposal-assisted training to reference-free generation. Experiments on standard object detection and referring expression comprehension benchmarks show that GOD improves direct coordinate generation while preserving discriminative localization and language-conditioned grounding. These gains suggest that object-level visual priors are most effective when learned within the MLLM, rather than supplied only as external proposals, providing a practical bridge between generative language modeling and localization-aware perception. We will make our code publicly available upon acceptance.
Abstract: 3D object understanding and generation methods in indoor scenes produce impressive results, yet they often overlook a pervasive source of information in real-world scenes: repeated objects. We introduce the task of lookalike object detection in indoor scenes, which leverages repeated and complementary cues from identical and near-identical object pairs. Given an input scene, the task is to classify pairs of objects as identical, similar or different using multiview images as input. To address this, we present Lookalike3D, a multiview image transformer that effectively distinguishes such object pairs through similarity learning by harnessing strong semantic priors from large image foundation models. To support this task, we collected the 3DTwins dataset, containing 76k manually annotated identical, similar and different pairs of objects based on ScanNet++, and show an improvement of 14% IoU over baselines on the task of lookalike object detection. We also demonstrate how our method improves the downstream tasks of joint 3D object reconstruction and part co-segmentation, turning repeated and lookalike objects into a powerful cue for consistent, high-quality 3D scene understanding. Our code, dataset and models will be made publicly available.
Abstract: Recent advances in deep learning have inspired neural network-based approaches to computational materials discovery (CMD). A plethora of problems in this field involve finding materials that optimize a target property. Nevertheless, the increasingly popular generative modeling methods are ineffective at boldly exploring attractive regions of the materials space due to their maximum likelihood training. In this work, we offer an alternative CMD technique based on offline model-based optimization (MBO) that fuses direct optimization of a target material property into generation. To that end, we introduce a domain-specific model, CliqueFlowmer, that incorporates recent advances in clique-based MBO into transformer and flow generation. We validate this model's optimization abilities and show that materials it produces strongly outperform those from generative baselines. To support specialized materials discovery applications and broader interdisciplinary research, we will release our code and model weights upon publication.
Abstract: Direct Preference Optimization is a widely used approach for safety alignment, aiming to reduce harmful behaviors in large language models. However, prior work shows that it can be brittle and exhibits poor out-of-distribution (OOD) generalization \citepqi2025safety. To this end, in this paper we investigate whether curriculum learning can improve the robustness of DPO-based safety alignment. We propose Staged-Competence, a curriculum-based framework that organizes preference data by difficulty, employs competence-based sampling, and progressively updates the reference model during training. Averaged across three model families, Staged-Competence reduces OOD harmful response rates by 16% and jailbreak attack success rates by 20%, while preserving general capabilities and maintaining near-zero over-refusal. We further show that Staged-Competence (1) matches baseline safety with only 75% of the training data, demonstrating improved data efficiency and (2) yields better separation between safe and unsafe responses. Staged-Competence is agnostic to the underlying policy optimization loss and can extend to other DPO variants and alignment domains other than safety. Our code and data can be found at: \urllink/upon/acceptance.
Authors:
Hongcan Guo, Qinyu Zhao, Yian Zhao, Shen Nie, Rui Zhu, Qiushan Guo, Feng Wang, Tao Yang, Hengshuang Zhao, Guoqiang Wei, Yan ZengAbstract: Autoregressive language models have achieved remarkable success in text modeling, but recent work increasingly challenges the fixed left-to-right generation order. However, existing non-autoregressive alternatives such as discrete diffusion, still struggle to jointly deliver efficiency, scalable representation learning, and global semantic modeling. We propose Cola DLM, a hierarchical continuous latent diffusion language model that learns a stable Text VAE, models a global semantic prior in latent space with a block-causal DiT, and decodes text conditionally. Theoretically, under a unified Markov-path formulation, the diffusion operates latent prior transport rather than token-level observation recovery, thereby decoupling global semantic organization from local token realization. Empirically, Cola DLM exhibits strong scaling behavior for text generation across four research questions, eight benchmarks, strictly matched ~2B-parameter AR and LLaDA baselines, and scaling curves up to ~2000 EFLOPs. Meanwhile, we also explore encoding text in the same continuous latent space as images, providing a feasible path toward unified discrete-continuous generative modeling. Our code will be released.
Abstract: Large Language Models often improve accuracy on reasoning tasks by sampling multiple Chain-of-Thought (CoT) traces and aggregating them with majority voting (MV), a test-time technique called self-consistency. When we truncate a CoT partway through and regenerate the remainder, we observe that traces with correct answers reproduce their original answer more often than traces with wrong answers. We use this difference as a reliability signal, prefix consistency, that weights each candidate answer by how often it reappears under regeneration. It requires no access to token log-probabilities or self-rating prompts. Across five reasoning models and four math and science benchmarks, prefix consistency is the best correctness predictor in most settings, and reweighting votes by it reaches Standard MV plateau accuracy at up to 21× fewer tokens (median 4.6×).
Abstract: Multi-Agent Systems (MAS) built on Large Language Models (LLMs) require effective orchestration to coordinate specialized agents, yet training such orchestrators is hindered by limited supervision and high computational cost. We propose Orchestration Reward Modeling (Orch-RM), a self-supervised framework for evaluating orchestration quality without human annotations. Orch-RM leverages intermediate artifacts from multi-agent executions to construct win-lose pairs for Bradley-Terry reward model training. Unlike existing MAS test-time scaling and orchestrator training frameworks that rely on costly sub-agent rollouts, Orch-RM operates directly at the orchestration level, enabling efficient and high-performing reward-guided orchestrator training and MAS test-time scaling. As shown in Figure 1, Orch-RM improves training efficiency by up to 10× in token usage while improving MAS test-time scaling performance by up to 8% in accuracy. These gains consistently transfer across multiple domains, including mathematical reasoning, web-based question answering, and multi-hop reasoning, demonstrating orchestration-level reward modeling as a scalable direction for robust multi-agent orchestration. Data, code, and trained models will be released upon publication.
Abstract: Real-time execution is crucial for deploying Vision-Language-Action (VLA) models in the physical world. Existing asynchronous inference methods primarily optimize trajectory smoothness, but neglect the critical latency in reacting to environmental changes. By rethinking the notion of reaction in action chunking policies, this paper presents a systematic analysis of the factors governing reaction time. We show that reaction time follows a uniform distribution determined jointly by the Time to First Action (TTFA) and the execution horizon. Moreover, we reveal that the standard practice of applying a constant schedule in flow-based VLAs can be inefficient and forces the system to complete all sampling steps before any movement can start, forming the bottleneck in reaction latency. To overcome this issue, we propose Fast Action Sampling for ImmediaTE Reaction (FASTER). By introducing a Horizon-Aware Schedule, FASTER adaptively prioritizes near-term actions during flow sampling, compressing the denoising of the immediate reaction by tenfold (e.g., in \pi_0.5 and X-VLA) into a single step, while preserving the quality of long-horizon trajectory. Coupled with a streaming client-server pipeline, FASTER substantially reduces the effective reaction latency on real robots, especially when deployed on consumer-grade GPUs. Real-world experiments, including a highly dynamic table tennis task, prove that FASTER unlocks substantially improved real-time responsiveness for generalist policies, enabling rapid generation of accurate and smooth trajectories.
Authors:
Jiace Zhu, Wentao Chen, Qi Fan, Zhixing Ren, Junying Wu, Xing Z Chai, Chotiwit Rungrueangwutthinon, Yehan Ma, An ZouAbstract: Recent studies have demonstrated the potential of Large Language Models (LLMs) in generating GPU Kernels. Current benchmarks focus on the translation of high-level languages into CUDA, overlooking the more general and challenging task of text-to-CUDA generation. Furthermore, given the hardware-specific and performance-critical features of GPU programming, accurately assessing the performance of LLM-generated GPU programs is nontrivial. In this work, we introduce CUDABench, a comprehensive benchmark designed to evaluate the text-to-CUDA capabilities of LLMs. First, we construct CUDABench-Set, which covers Breadth-Depth-Difficulty evaluation space in diverse application domains, including artificial intelligence, scientific computing, and data analytics, etc. Furthermore, we propose CUDABench-Score and Generative Verification Pipeline that assess (1) compilation correctness, (2) functional consistency through execution-based verification, and (3) a novel roofline-based metric, Performance-Score. Benchmarking state-of-the-art LLMs reveals insightful findings and challenges of text-to-CUDA, such as a notable mismatch between high compilation success rates and low functional correctness, a lack of domain-specific algorithmic knowledge, and suboptimal utilization of GPU hardware resources.
Abstract: In this work, we focus on extending SHARP, the popular photorealistic view synthesis method, for universal monocular rendering across diverse camera systems, such as large-field-of-view fisheye lenses. To overcome the pinhole-specific assumptions of SHARP, our key idea is to align various images in a unified omnidirectional latent space. Thus, we propose UniSHARP, which performs implicit alignment in both feature and Gaussian spaces. Specifically, 2D semantic embeddings and 3D spatial features are jointly encoded and decoded by UniK3D to support Gaussian construction, while a ray-based universal representation organizes Gaussian primitives along rays and radial distances. To comprehensively evaluate our method, we construct a benchmark covering diverse imaging systems, including pinhole, fisheye and panoramic cameras, across various scenes. The benchmark is further stratified by field of view, spanning narrow perspective, wide-FoV, fisheye, and full panoramic settings. Extensive experiments on the proposed benchmark demonstrate the effectiveness of UniSHARP, outperforming alternative methods by a large margin. Our models, training code, and dataset will be publicly available.
Authors: Liam Chalcroft
Abstract: Latent medical image generators usually treat the tokenizer as fixed preprocessing. We test whether this separation is valid in a controlled ChestMNIST study at 64×64, crossing discrete tokenizers, generator families, and sampler settings under a shared latent grid, with continuous-latent reference cells. The results show that tokenizer, generator, and sampler cannot be ranked independently. LFQ is the strongest default discrete tokenizer in the matched VQ/LFQ/FSQ grid, but generator rankings change with quantizer family and categorical rate. On LFQ-1024, retuning D3PM and SEDD moves them from default FID-192 0.44/0.41 to 0.09/0.11 at lower NFE, approaching the strongest continuous references. Reconstruction quality is not a reliable selection criterion: the highest-PSNR tokenizers are not the best generation substrates, even when each tokenizer is paired with its best observed generator. We interpret these results as a rate-distortion-modelability tradeoff, where modelability is conditional on the generator, sampler, and inference budget. The study covers 54 discrete tokenizer-generator cells and 16 continuous reference cells, and its scope is deliberately limited: low resolution, one training seed, unconditional generation, and non-clinical FID-based evaluation.
Abstract: We present PanoWorld, a panoramic video world model that generates geometry-consistent 360\degree video from a single image and a caption. Existing panoramic video methods optimize primarily for visual realism and do not explicitly constrain the underlying 3D scene state, producing outputs that appear plausible yet exhibit inconsistent depth, broken correspondences, and implausible motion across the spherical surface. We address this gap by framing panoramic video generation as a geometry- and dynamics-consistent latent state modeling problem rather than pure visual synthesis. Building on a pre-trained perspective video world model, we introduce two lightweight regularizers: a depth consistency loss against pseudo ground-truth panoramic depth, and a trajectory consistency loss that supervises the 3D world-frame positions of tracked points across time. We further apply spherical-geometry-aware adaptation to the conditioning and positional encoding. We additionally introduce PanoGeo, a unified geometry-aware panoramic video dataset with consistent depth, trajectory, and prompt annotations across diverse real and synthetic sources, used for both training and stratified evaluation. Experiments show that PanoWorld improves geometric consistency over prior panoramic generation methods while maintaining competitive visual realism, establishing that panoramic video generation must be treated as a geometric modeling problem to support the holistic spatial understanding requirements of embodied AI applications.
Abstract: Large language models (LLMs) are increasingly deployed as decision making components in real-world systems for societal and scientific applications, creating a growing need for reliable predictions. In this paper, we study the problem of reliable decision making with LLMs via the lens of selective prediction, allowing the model to improve performance by trading-off coverage. The aim of selective prediction is to balance the key tradeoff of risk and coverage, where risk measures the predictive performance on selected inputs and coverage measures the fraction of abstentions. While existing LLM post-training approaches focus primarily on correctness or calibration, we propose to directly optimize for selective prediction performance by introducing reinforcement Learning for Selection Reward (RLSR), which targets the area under the risk-coverage curve (AURC) as its training objective. RLSR achieves substantially better risk-coverage tradeoff compared to multiple baselines on both in-domain and out-of-domain tasks.
Abstract: While representation alignment with self-supervised models has been shown to improve diffusion model training, its potential for enhancing inference-time conditioning remains largely unexplored. We introduce Representation-Aligned Guidance (REPA-G), a framework that leverages these aligned representations, with rich semantic properties, to enable test-time conditioning from features in generation. By optimizing a similarity objective (the ) at inference, we steer the denoising process toward a conditioned representation extracted from a pre-trained feature extractor. Our method provides versatile control at multiple levels of granularity, ranging from patch level matching via single patches to broad semantic guidance using global image feature tokens. We further extend this to multi-concept composition, allowing for the faithful combination of distinct concepts. REPA-G operates entirely at inference time with no additional training required, offering a flexible and precise alternative to often ambiguous text prompts or coarse class labels. Our approach achieves high-quality, diverse generations on ImageNet and COCO.
Abstract: Autoregressive language models are widely used for text evaluation, however, their left-to-right factorization introduces positional bias, i.e., early tokens are scored with only leftward context, conflating architectural asymmetry with true text quality. We propose masked reconstruction as an alternative paradigm, where every token is scored using full bidirectional context. We introduce DiffScore, an evaluation framework built on Masked Large Diffusion Language Models. By measuring text recoverability across continuous masking rates, DiffScore eliminates positional bias and naturally establishes an evaluation hierarchy from local fluency to global coherence. We further provide diagnostic tools unavailable to autoregressive frameworks: multi-timestep quality profiles that decompose scores across masking rates, and bidirectional Pointwise Mutual Information (PMI) decomposition that disentangles fluency from faithfulness. Experiments across ten benchmarks show that DiffScore consistently outperforms autoregressive baselines in both zero-shot and fine-tuned settings.
Abstract: While rule-based reinforcement learning has recently catalyzed explicit reasoning in multimodal models, tactile reasoning remains largely underexplored. Existing tactile-language models primarily rely on supervised or contrastive objectives, which limits their capacity to ground predictions in physical evidence or rectify misleading visual priors. Tactile reasoning introduces two modality-specific challenges: the ordinal nature of physical attributes (e.g., hardness, roughness) and the cross-sensor distribution shifts inherent in optical tactile hardware. In this work, we introduce TouchReason-1M, a large-scale multimodal dataset comprising over 1M synchronized tactile pairs across four distinct sensors, and TouchReason-Bench, a rigorous framework for evaluating tactile perception and visual-tactile conflict resolution. Building upon these, we propose Touch-R1, a tactile reasoning MLLM based on Qwen2.5-VL-7B. Touch-R1 is trained via a tactile-grounded GRPO objective that combines ordinal-aware accuracy, cross-sensor physical consistency, structured-format control, and an input-side tactile grounding objective. Specifically, the tactile-use reward assigns credit only when authentic tactile inputs yield superior correctness relative to counterfactual controls where the tactile stream is removed, shuffled, or noise-masked. On TouchReason-Bench, Touch-R1-7B outperforms Octopi-13B by 18.4% and GPT-4o by 24.7% on average. Its structured reasoning traces reveal emergent behaviors of probing, comparison, and revision, demonstrating that R1-style reasoning can be effectively grounded in physical contact. Our code and data will be made public.
Abstract: Some of the most performant reinforcement learning algorithms today can be prohibitively expensive as they use test-time scaling methods such as sampling multiple action candidates and selecting the best one. In this work, we propose FASTER, a method for getting the benefits of sampling-based test-time scaling of diffusion-based policies without the computational cost by tracing the performance gain of action samples back to earlier in the denoising process. Our key insight is that we can model the denoising of multiple action candidates and selecting the best one as a Markov Decision Process (MDP) where the goal is to progressively filter action candidates before denoising is complete. With this MDP, we can learn a policy and value function in the denoising space that predicts the downstream value of action candidates in the denoising process and filters them while maximizing returns. The result is a method that is lightweight and can be plugged into existing generative RL algorithms. Across challenging long-horizon manipulation tasks in online and batch-online RL, FASTER consistently improves the underlying policies and achieves the best overall performance among the compared methods. Applied to a pretrained VLA, FASTER further improves task success while substantially reducing training and inference compute requirements.
Abstract: Recent video editing models have converged on a unified-conditioning design: a single diffusion transformer reads text, source video, reference images, and masks through one token sequence, and one set of weights covers replacement, removal, style transfer, and reference-driven insertion. The design is flexible, but it assumes that the user already provides model-ready text, identity references, and spatial targets, which real requests often omit. We present Aurora, an agentic video editing framework that pairs a tool-augmented vision-language models (VLMs) agent with a unified video diffusion transformer. The agent maps a raw user request to a structured edit plan aligned with the transformer's conditioning channels, thereby resolving textual and visual underspecification before generation. We train the agent with supervised data for executable edit planning and reference-image selection, together with preference pairs for robust tool use and instruction refinement. Moreover, we introduce AgentEdit-Bench to evaluate agent-enhanced video editing under textual and visual underspecification. Experiments on EditVerse-Bench, OpenVE-Bench, and AgentEdit-Bench show that Aurora improves over text-only baselines and its agentic capability transfers to compatible frozen video editing models. Code, data, and models will be released.
Authors:
Hyojun Go, Hyungjin Chung, Prune Truong, Goutam Bhat, Li Mi, Zhaochong An, Zixiang Zhao, Dominik Narnhofer, Serge Belongie, Federico Tombari, Konrad SchindlerAbstract: For practical use, diffusion- or flow-based generative models must be aligned with task-specific rewards, such as prompt fidelity or aesthetic preference. That alignment is challenging because the reward is defined for clean output images, but the alignment procedure requires value function estimates at noisy intermediate latents. Existing methods resort to Tweedie-style or Monte Carlo approximations, trading off estimator bias against computational cost: Tweedie estimates are efficient but biased, while Monte Carlo estimates are more accurate but require expensive rollouts. A natural alternative would be a learned value function, but it remains an open question how to effectively train a strong and general value model specifically for noisy latents. Here, we propose StitchVM (Stitched Value Model), a model stitching framework that efficiently transfers reward models pretrained for clean images to the noisy latent regime. StitchVM starts from an existing, truncated pixel-space reward model and attaches a frozen diffusion backbone to it as its head. From the pixel-space model, the resulting hybrid retains a carefully pretrained, robust reward capability; from the diffusion backbone, it inherits its native ability to handle noisy latents. The stitching procedure is exceptionally lightweight, e.g., stitching and finetuning CLIP ViT-L and SD 3.5 Medium takes only \approx10 GPU-hours. By lifting powerful pixel-space reward models to latent space, StitchVM opens up a new style of diffusion alignment: instead of rough, yet costly per-sample approximation of the value function, the correct function for the actual, noisy latents is constructed once and then amortized over many samples and iterations. We show that this approach yields improvements across a broad range of downstream steering and post-training methods: DPS becomes 3.2× faster while halving peak GPU memory, and DiffusionNFT becomes 2.3× faster.
Abstract: Verification methods aim at mathematically proving desirable properties of neural networks, such as robustness to adversarial perturbations. A verifier is sound if and only if it never claims that a neural network has the desired property when it does not. It was shown recently that none of the currently known verifiers that are claimed to be sound are guaranteed to be sound when considering the deployed version of the verified network. Due to this, all the known verifiers are vulnerable to certain backdoor attacks, where an adversarial network passes verification, but in reality, it exhibits adversarial behavior in specific deployment environments. So far, it has been suspected that sound verification is prohibitively expensive if we wish to verify all possible executions—including parallel and stochastic ones—in deployment. We show that efficient and practically sound verifiers can be designed using a bound on the so-called backward error. Using this bounding technique, we propose two verifiers: one based on interval bound propagation, and one using symbolic propagation. Both verifiers are proven to remain sound even if the deployment environment randomly selects a valid expression tree (an ordering and parenthesizing of the arithmetic operations) to compute the network, and even if numeric underflow or overflow occurs in any valid expression tree. This is especially interesting in the case of symbolic propagation, where the expression tree used by the verifier is completely different from the one used in deployment. We demonstrate empirically that our techniques introduce only a limited performance overhead while detecting all the known verifier backdoor attacks.
Abstract: We consider the conversion of musical recordings into human-readable sheet music annotated with timestamps. Such output lets a listener clearly visualize (temporally expressive playing), a learner diagnose ensemble precision and timing choices against the written music, and a musicology scholar compare performance styles across recordings of the same work. We introduce (1) a prompt-conditioned encoder--decoder model, named Rubato, trained to output (2) a new textual representation for polyphonic music, named InterMo, which we designed for compatibility with sequence-to-sequence training. Our experiments demonstrate that Rubato produces timestamped piano sheet music from audio with higher notational accuracy than the best existing approaches, which are based on cascades. We find that even if the cascade is given ground-truth MIDI instead of audio, Rubato performs better, suggesting that the ceiling of existing approaches is primarily representational, not acoustic. Further, because Rubato is trained on several related tasks (with prompts), it competes with or outperforms the best single-task systems on related but simpler tasks like MIDI note grounding and beat/downbeat detection. We will release the model and code on publication, and an anonymized demo is available at https://rubato-transcription.github.io.
Authors: Zhendong Li, Akwum Onwunta
Abstract: Uncertainty quantification (UQ) for random partial differential equations (PDEs) is ubiquitous in computational science and engineering. However, classical spectral solvers for this class of problems face the curse of dimensionality, and existing neural solvers often ignore the stochastic structure that makes moments and calibration tractable. We introduce a stochastic separable physics-informed neural network, dubbed S^2-PINN, that represents the solution u(t,\mathbfx,\mathbfZ) of a random PDE with a learnable Gaussian spatial dictionary, Fourier temporal features, and a generalized polynomial chaos (gPC) stochastic bases, coupled by a low-rank Canonical Polyadic (CP) tensor decomposition core. The method is trained with a hybrid strong-form and gPC-projected residual loss. Our theoretical analysis establishes that the separable class is dense in L^2 under mild conditions, and the projected residual corresponds exactly to a stochastic Galerkin constraint. Furthermore, we show that mini-batch projection coefficients are logarithmically dependent on the number of gPC modes, and that orthogonality penalty controls the conditioning of the learned spatial dictionary. Using four manufactured random PDE benchmarks, we show that S^2-PINN outperforms six baselines in terms of mean and variance accuracy, as well as calibration, while using significantly fewer parameters. Further evaluations on non-manufactured Poisson and Darcy problems, a diffusion scaling study of higher random dimensions, and two stochastic inverse problems reveal the generalization capabilities of the proposed structure. Together, these results support stochastic separability as an effective design principle for physics-informed neural UQ. The code for the experiments can be found in \hrefhttps://anonymous.4open.science/status/neurips2026-0424https://anonymous.4open.science/status/neurips2026-0424.
Abstract: While large language models (LLMs) appear to be increasingly capable of solving compositional tasks, it is an open question whether they do so using compositional mechanisms. In this work, we investigate how feedforward LLMs solve two-hop factual recall tasks, which can be expressed compositionally as g(f(x)). We first confirm that modern LLMs continue to suffer from the "compositionality gap", i.e. their ability to compute both z = f(x) and y = g(z) does not entail their ability to compute the composition y = g(f(x)). We then decode residual stream representations and identify two processing mechanisms: one which solves tasks compositionally, computing f(x) along the way to g(f(x)), and one which solves them directly, without any detectable signature of the intermediate variable f(x). Finally, we find that embedding space geometry is strongly related to which mechanism is employed, where the idiomatic mechanism is dominant when tasks are represented by translations from x to g(f(x)) in the embedding spaces.
Abstract: Diffusion-based super-resolution (SR) is a key component in video generation and video restoration, but is slow and expensive, limiting scalability to higher resolutions and longer videos. Our key insight is that many regions in video are inherently low-detail and gain little from refinement, yet current methods process all pixels uniformly. To take advantage of this, we propose SkipSR, a simple framework for accelerating video SR by identifying low-detail regions directly from low-resolution input, then skipping computation on them entirely, only super-resolving the areas that require refinement. This simple yet effective strategy preserves perceptual quality in both standard and one-step diffusion SR models while significantly reducing computation. In standard SR benchmarks, our method achieves up to 60% faster end-to-end latency than prior models on 720p videos with no perceptible loss in quality.
Abstract: Vision-to-code tasks require models to reconstruct structured visual inputs, such as charts, tables, and SVGs, into executable or structured representations with high visual fidelity. While recent Large Vision Language Models (LVLMs) achieve strong results via supervised fine-tuning, reinforcement learning remains challenging due to misaligned reward signals. Existing rewards either rely on textual rules or coarse visual embedding similarity, both of which fail to capture fine-grained visual discrepancies and are vulnerable to reward hacking. We propose Visual Equivalence Reward Model (Visual-ERM), a multimodal generative reward model that provides fine-grained, interpretable, and task-agnostic feedback to evaluate vision-to-code quality directly in the rendered visual space. Integrated into RL, Visual-ERM improves Qwen3-VL-8B-Instruct by +8.4 on chart-to-code and yields consistent gains on table and SVG parsing (+2.7, +4.1 on average), and further strengthens test-time scaling via reflection and revision. We also introduce VisualCritic-RewardBench (VC-RewardBench), a benchmark for judging fine-grained image-to-image discrepancies on structured visual data, where Visual-ERM at 8B decisively outperforms Qwen3-VL-235B-Instruct and approaches leading closed-source models. Our results suggest that fine-grained visual reward supervision is both necessary and sufficient for vision-to-code RL, regardless of task specificity.
Abstract: Self-improving LLM agents typically improve answers, not the process by which answers are improved: their search loops and mutation rules are held fixed by design. We introduce meta^n, a framework for recursive self-improvement in which a single universal meta-operation \Omega is applied repeatedly to build a hierarchy of meta-layers whose depth is sized by convergence rather than prescribed. Each layer observes execution traces from below, diagnoses failure patterns, and emits executable code (strategic guidance via pre-processing and reusable functions via a code library) whose effects cascade downward. Inter-layer conditioning permits up to \prod_d=2^n k_d effective strategy-tactic combinations across n layers rather than \sum_d k_d. An evolutionary orchestration mode maintains a growing archive of candidate chains with fitness-weighted parent selection and cross-candidate inspiration. Across seven benchmarks (combinatorial optimization, two text classification tasks, terminal agent tasks, mathematical discovery, symbolic regression, algorithm speedup), meta^n beats prior self-improving agents on all five (of seven) benchmarks where they apply, with the largest gap on CO-Bench (0.849 test vs. Gödel Agent's 0.000 and OpenEvolve's 0.816). Layers specialize without prescription and their roles emerge at depth: generic primitives at depth 2 propagate as cross-task callables, task-aware routing emerges at depth 3, and deeper layers refine via specialization. Code and data will be released publicly.
Abstract: All-atom generative modeling of 3\mathrmD biomolecular complexes has emerged as the dominant paradigm for predicting the structure of proteins and protein-ligand systems. Generating structures at the atomic level of fidelity, however, typically requires expensive iterative diffusion rollouts, making both conventional deployment and inference-time search techniques computationally costly. In this paper, we introduce the Denoiser Cofolding All-atom Flowmap (DeCAF) framework for distilling state-of-the-art all-atom cofolding models into all-atom flow maps that produce high-quality samples in only a few inference steps. We build DeCAF on a denoiser-based formulation of flow maps with endpoint losses that naturally support \mathrmSE(3) rigid alignment, which we show is critical for training accurate models. We further derive a simple change of variables that lets DeCAF operate in the \sigma-space noise schedule of EDM-style architectures, enabling direct distillation from pretrained cofolding diffusion models. Equipped with DeCAF's flowmap lookahead, we introduce a purpose-built inference-time framework that improves sampling through reward-guided search. Empirically, DeCAF statistically improves over Boltz-1(x) in both accuracy (RMSD) and physical validity scores of protein-ligand poses at strict NFE budgets on the challenging Runs N’ Poses, while also showing a more optimal Pareto frontier across all inference compute budgets on PoseBusters.
Abstract: Partially Relevant Video Retrieval (PRVR) aims to retrieve an untrimmed video that contains at least one segment relevant to a text query. Despite its practical motivation, existing PRVR methods are typically trained and evaluated on a single curated benchmark, limiting insight into whether they learn generalizable fine-grained text-video alignment. To examine this issue, we introduce a multi-source-to-unseen-target protocol, in which multiple source benchmarks are merged into a single training corpus without source labels and the trained model is evaluated on a held-out target benchmark. Under this protocol, existing PRVR methods generalize poorly to unseen targets despite leveraging the general image-text alignment of CLIP. Surprisingly, they even underperform zero-shot CLIP, suggesting that PRVR training can actively erode the transferable alignment inherited from CLIP. To address this limitation, we propose Generalization-Aware Learning (GAL) for PRVR, a two-branch framework that preserves CLIP-derived transferability while learning PRVR-specific fine-grained text-segment matching. GAL consists of an anchored branch that maintains CLIP's transferability and an adaptive branch that learns task-oriented segment-level alignment. We train these branches through conservative residual updates, per-sample branch specialization, and CLIP-text relation distillation, enabling complementary learning without overwriting the alignment structure provided by CLIP. Under the proposed protocol, GAL achieves state-of-the-art unseen-target retrieval with competitive source performance, showing that fine-grained matching can be learned without collapsing transferable alignment.
Authors: Xinchen Jin, Aditya Chatterjee, Pranav Kumar, Rohan Paleja
Abstract: Vision-Language-Action (VLA) policies translate language and visual inputs into robot actions, where their hidden representations directly shape closed-loop behavior. However, mechanistic interpretability tools from language and vision-language models do not transfer cleanly to VLAs: outputs are robot actions rather than human-readable tokens, and validating internal changes requires expensive closed-loop rollouts. We propose an event-grounded interpretability pipeline that anchors SAE feature analysis to behavioral events rather than text contexts. End-effector keyframes are clustered within each task using visual, state, and temporal cues, linking SAE features to behaviorally salient events and, via optional VLM annotations, to semantic context. To our knowledge, our pipeline is among the first to ground SAE-based VLA analysis in closed-loop behavioral events. Across two simulation architectures and a real-robot study, event-grounded ranking yields the strongest causal effects on OpenVLA and transfers to the continuous action chunks of \pi_0.5. SAE is a sparse but imperfect intervention basis: usability varies with architecture and intervention site, and aggressive intervention reveals safety and interpretability limits. Overall, event-grounded SAE analysis emerges as a practical starting point for behavior-anchored VLA interpretability, motivating future work on SAE features beyond action-aligned coordinates, finer-grained closed-loop evaluation, and safe interventions for high-stakes VLA deployments.
Authors:
Samet Hicsonmez, Eray Çakar, Nermin Samet, Fatma GuneyAbstract: Accident anticipation aims to recognize anomalous driving cues before a crash while avoiding false alarms during normal driving. Existing approaches typically formulate this task as binary classification, focusing on whether an accident will occur rather than when it will occur. We propose PRE-ACT, a framework that models accident risk as a continuously evolving signal that increases as the crash approaches. By explicitly enforcing temporal ordering and distance-to-accident awareness, our method progressively raises risk while suppressing premature alarms, leading to significant improvements on MM-AU subsets and Nexar. We further introduce a Separation Score to evaluate the global behavior of predicted risk curves beyond local temporal windows. Code and visualizations are available.
Authors:
Weichu Xie, Haozhe Zhao, Liuwenpu, Yongfu Zhu, Liang Chen, Minghao Ye, Zirong Chen, Yuqi Xu, Shuai Dong, Ziyue Wang, Xinbo Xu, Kean Shi, Ruoyu Wu, Xiaoying Zhang, Wenqi Shao, Baobao Chang, Nan Duan, Jiaqi WangAbstract: Reinforcement Learning with Verifiable Rewards (RLVR) is widely used to improve the reasoning capabilities of large language models, but its reward is derived only from the correctness of the final answer and provides no supervision over intermediate reasoning steps. Recent rubric-based methods, such as Rubrics as Rewards (RaR), introduce finer-grained supervision by scoring rollouts against a structured set of evaluation criteria. However, the resulting rubric scores are still aggregated into a single scalar that is applied to the entire response, leading to three structural weaknesses, namely loss of the multi-criterion rubric structure, uniform supervision of correct and incorrect reasoning steps, and reward hacking in the trained model through unbounded self-correction. On a sample of 1,000 problems, we find that 18.2% of steps within answer-correct responses are themselves wrong yet positively rewarded, while 49.9% of steps within answer-incorrect responses are in fact correct yet penalized. Therefore, we introduce Step-wise Rubrics as Rewards (SRaR), an RLVR framework that (i) uses an LLM judge to attribute each rubric item to a specific reasoning step, (ii) normalizes the per-step rubric scores across rollouts so that only steps whose quality varies produce a learning signal, and (iii) combines the resulting per-step reward with the standard outcome reward through a decoupled advantage estimator that keeps the outcome-driven baseline stable. To support training, we further build a 16K-problem rubric dataset by contrastively distilling rubric items from correct and flawed reasoning paths sampled from a strong model and verified against the ground-truth answer. Across six mathematical reasoning benchmarks spanning multiple difficulty levels, SRaR improves the average accuracy over RaR by 3.57 points on Qwen3-8B-Non-Thinking and by 2.75 points on Qwen3-32B-Non-Thinking, raises the Faithful Reasoning Rate on AIME~2025 from 34.5% to 46.7% (every correct-answer trajectory of SRaR uses entirely correct reasoning steps), and reduces the rate of self-correction looping from 48.1% to 26.5%.
Authors:
Zezhong Qian, Xiaowei Chi, Chak-Wing Mak, Tianze Zhou, Ruibin Yuan, Yuhan Rui, Hengzhe Sun, Zhuoqun Wu, Yuming Li, Siyuan Qian, Sirui Han, Shanghang ZhangAbstract: Video models are recently evolving into vision foundation models, but they still lack human-like, multi-step reasoning. Existing streaming autoregressive diffusion models are efficient but lack the reasoning ability, whereas bidirectional diffusion allows for global revision but incurs high inference cost due to the dense frames in fixed-sequence denoising. Consequently, both paradigms struggle to maintain logical consistency with low-latency streaming in complex reasoning tasks. Bridging this gap, we propose HDR (Hierarchical Denoising for Visual Reasoning), a unified framework for multi-step reasoning by integrating hierarchical latents into the causal video generation process. HDR organizes video latents into a tree-structured hierarchy to perform coarse-to-fine reasoning before streaming output. Coarse denoising layers maintain uncertain hypotheses for global planning, while finer denoising layers progressively refine them into concrete visual states. A sparse hierarchical attention pattern (SHAP) further reduces temporal attention cost. We construct a level-stratified multi-step video reasoning benchmark with out-of-distribution cases, covering six tasks: maze navigation, Tower of Hanoi, one-line drawing, sliding puzzle, Sokoban, and water pouring. Compared with the streaming autoregressive diffusion baseline, HDR improves overall success from 34.22 to 60.29 (76.2% relative gain) in multi-step reasoning accuracy, and improves average progress from 76.00 to 89.56, indicating more consistent intermediate reasoning trajectories. For deployment efficiency, HDR maintains low-latency streaming at 0.70s per latent, 54.2x faster than bidirectional diffusion during streaming. HDR also demonstrates strong data efficiency, retaining 82.9% of its full-data success score using only 2% of the training data, compared with 52.0% for bidirectional diffusion. Further experiments on real-world robots showcase the potential of HDR in physical interaction, providing a new paradigm for physical world modeling. Project demo page is available at https://hierarchical-diffusion-reasoning.github.io/.
Abstract: Current world models operate at a single level of abstraction, with most prioritizing perceptual fidelity while lacking the spatial reasoning and semantic understanding required for real-world downstream tasks. We present a hierarchical driving world model that factorizes future prediction across two levels operating at distinct temporal and abstraction scales: a high-level predictor that forecasts coarse scene structure over extended temporal horizons, and a low-level generator that produces detailed predictions conditioned on the high-level output. This decomposition yields high perceptual fidelity while also capturing strong spatial and semantic representations. We further show that pretraining with a diffusion forcing objective yields substantially richer internal representations than the standard teacher forcing objective, while teacher forcing — predicting only the next frame from clean context — produces more stable autoregressive rollouts. We therefore introduce a generic two-stage training paradigm that pretrains the model with diffusion forcing and fine-tunes with teacher forcing, combining the representational benefits of the former with the rollout stability of the latter. Our approach achieves state-of-the-art results across the standard suite of driving world model evaluations on established benchmarks, including long-horizon generation fidelity, steering responsiveness evaluated on counterfactual scenarios, and internal representation quality.
Abstract: Commercial egocentric devices capture human behavior in everyday settings, yet full-body motion reconstruction must generalize across diverse cameras and mountings. This is challenging because head trajectories under-determine body motion, while hand observations are intermittent and camera-dependent: a hand may disappear simply because it has left the camera's field of view. Existing methods treat such absences as missing data, relying on fixed visibility regimes or post-hoc optimization, and fail to generalize across devices.We present \modelname, a sequence-level diffusion framework for device-agnostic egocentric full-body motion reconstruction. Our key insight is that egocentric videos contain sequence-level evidence about both the wearer and the camera-induced visibility structure, enabling \modelname to infer consistent body shape and coherent motion from sparse head and hand cues. To prevent overfitting to a single camera setup, we further introduce geometry-aware visibility augmentation, which synthesizes visibility patterns from realistic variations in camera geometry. We also present OmniEgoDB, the first mocap benchmark with ground-truth motion captured across multiple consumer egocentric devices. Experiments on synthetic, real-device, and in-the-wild settings demonstrate superior motion reconstruction, robust cross-device generalization, and coherent in-the-wild results.
Abstract: Understanding the internal routing of visual concepts is crucial for the safe and controllable adaptation of generative models. While concept localization has been widely studied in diffusion models, the emerging paradigm of next-scale visual autoregressive models remains largely unexplored. In this paper, we introduce LoCo, the first model-agnostic framework to localize where and when specific concepts emerge within autoregressive models. Driven by the native coarse-to-fine nature of next-scale generation, our method precisely maps conceptual knowledge across three distinct dimensions: Layer, Scale, and Position. To systematically localize and evaluate concept routing without context bias, we propose LoCoBench, a comprehensive dataset spanning 10 diverse categories. Extensive probing on image and video autoregressive models like Infinity, HunyuanImage-3.0, and InfinityStar shows that the localized positions are both interpretable and causally related to concept emergence. Building on these insights, we apply our localization method to three key applications: concept erasure, model personalization, and adversarial concept injection. Experiments demonstrate that our targeted intervention achieves state-of-the-art performance, substantially reducing computational overhead while preserving benign utility. Overall, our findings offer insights into how conceptual knowledge is routed during autoregressive generation, introducing a practical pathway for more interpretable, efficient, and secure adaptation.
Authors:
Shuo Yang, Haocheng Xi, Yilong Zhao, Muyang Li, Xiaoze Fan, Jintao Zhang, Han Cai, Yujun Lin, Xiuyu Li, Kurt Keutzer, Song Han, Chenfeng Xu, Ion StoicaAbstract: k-means has historically been positioned primarily as an offline processing primitive, typically used for dataset organization or embedding preprocessing rather than as a first-class component in online systems. In this work, we revisit this classical algorithm under the lens of modern AI system design and enable k-means as an online primitive. We point out that existing GPU implementations of k-means remain fundamentally bottlenecked by low-level system constraints rather than theoretical algorithmic complexity. Specifically, the assignment stage suffers from a severe IO bottleneck due to the massive explicit materialization of the N × K distance matrix in High Bandwidth Memory (HBM). Simultaneously, the centroid update stage is heavily penalized by hardware-level atomic write contention caused by irregular, scatter-style token aggregations. To bridge this performance gap, we propose Flash-KMeans, an IO-aware and contention-free k-means implementation for modern GPU workloads. Flash-KMeans introduces two core kernel-level innovations: (1) FlashAssign, which fuses distance computation with an online argmin to completely bypass intermediate memory materialization; (2) Sort-Inverse Update, which explicitly constructs an inverse mapping to transform high-contention atomic scatters into high-bandwidth, segment-level localized reductions. Furthermore, we integrate algorithm-system co-designs, including chunked stream overlap and cache-aware compile heuristic, to ensure practical deployability. Extensive evaluations on NVIDIA H200 GPUs demonstrate that Flash-KMeans achieves up to 17.9× end-to-end speedup over best baselines. At the kernel level, FlashAssign and Sort-Inverse Update deliver up to 21.2× and 6.3× speedups. By systematically restructuring execution around underlying hardware constraints, Flash-KMeans delivers mathematically exact, scalable, and highly deployable acceleration across diverse AI workloads.
Abstract: Neural representations for videos (NeRV) have shown strong reconstruction fidelity by storing video-specific information in network weights. However, existing formulations typically require either costly per-video optimization or video-specific weight generation, making it difficult to scale to efficient amortized video representation. We propose CoANeRV, a coordinate-aware token-space neural video representation framework that shifts video-specificity from decoder weights to compact latent video tokens. Instead of optimizing or generating an instance-specific network, CoANeRV uses a shared coordinate-conditioned decoder to reconstruct videos by querying tokenized video content at continuous spatio-temporal coordinates. This formulation decouples video representation from parameter instantiation, enabling feed-forward video encoding while retaining coordinate-level reconstruction flexibility. To make token-space reconstruction effective, CoANeRV introduces a coordinate-aware decoding architecture that aligns spatio-temporal queries with video tokens through axis-adaptive positional encoding and temperature-modulated cross-attention. Block-wise coordinate querying further reduces peak attention memory, making high-resolution reconstruction practical. Experiments on diverse video datasets show that CoANeRV consistently improves reconstruction quality over prior feed-forward NeRV and INR baselines, reduces peak memory compared with attention-based coordinate decoders, and provides efficient amortized encoding without per-video optimization. These results suggest that token-space neural representation is a scalable alternative to conventional weight-space NeRV formulations. The code is available at https://anonymous.4open.science/r/CoANeRV-50F7/.
Abstract: Latent visual reasoning involves visual evidence more directly in multimodal reasoning by inserting continuous latent tokens before textual generation. However, the necessity of these latent tokens at inference remains ambiguous. We show that replacing latent tokens with random noise or removing them completely causes little performance degradation across spatial reasoning benchmarks. Reinforcement learning further diminishes the latent generation behavior after post-training. These observations raise a central question: Is latent visual reasoning still meaningful? We argue that its value should be measured by how effectively latent tokens guide learning, rather than whether they persist as an inference-time format. Our analysis shows that latent reasoning is unevenly favorable across question types, yet hard task-level routing for applying latent generation is brittle. Motivated by these findings, we propose an attention-based reward that encourages generated latent tokens to interact with later text tokens during RL. This reward promotes latent utilization when the latent mode is activated while preserving the flexibility to use pure-text reasoning. Experiments show that our method improves performance across perception and visual reasoning benchmarks, even when latent tokens are rarely generated after post-training. Our results highlight that, without explicit expression at inference, latent visual reasoning can shape better visual grounding and more accurate textual reasoning in silence. Our code and trained models are publicly available at Hugging Face https://huggingface.co/collections/doubleblindsubmission/submission3847.
Abstract: Multi-Agent Reinforcement Learning (MARL) faces significant exploration challenges due to exponentially growing joint state-action spaces. Existing exploration methods operate directly in raw state-action spaces, which is inefficient and fails to exploit inherent semantic structure. This paper introduces semantic space into multi-agent exploration. We theoretically establish that the observation-action space can be partitioned into discrete semantic prototypes, providing a principled foundation for transferring exploration statistics across semantically similar situations. Building on this theory, we propose a semantic-level exploration mechanism that first compresses the high-dimensional observation–action space into discrete semantic prototypes via vector quantization, and then applies sliding-window count-based bonuses to enable efficient statistical transfer across semantically similar situations. Our approach is versatile and can be seamlessly integrated with existing value-based MARL frameworks. Extensive experiments demonstrate that our method outperforms state-of-the-art baselines across diverse multi-agent benchmarks in terms of both effectiveness and training efficiency.
Abstract: Large language models (LLMs) can be said to have preferences: they reliably pick certain tasks and outputs over others, and preferences shaped by post-training and system prompts appear to shape much of their behaviour. But models can also adopt different personas which have radically different preferences. How is this implemented internally? Does each persona run on its own preference machinery, or is something shared underneath? We train linear probes on residual-stream activations of Gemma-3-27B and Qwen-3.5-122B to predict revealed pairwise task choices, and identify a genuine preference vector: it tracks the model's preferences as they shift across a range of prompts and situations, and on Gemma-3-27B steering along it causally controls pairwise choice. This preference representation is largely shared across personas: a probe trained on the helpful assistant predicts and steers the choices of qualitatively different personas, including an evil persona whose preferences anti-correlate with the Assistant's.
Abstract: We present DMax, a new paradigm for efficient Diffusion Language Models (dLLMs). It mitigates error accumulation in parallel decoding, enabling aggressive decoding parallelism while preserving generation quality. Unlike conventional masked dLLMs that decode through a binary mask-to-token transition, DMax reformulates decoding as a progressive self-refinement from mask embeddings to token embeddings. At the core of our approach is On-Policy Uniform Training, a novel training strategy that efficiently unifies masked and uniform dLLMs, equipping the model to recover clean tokens from both masked inputs and its own erroneous predictions. Building on this foundation, we further propose Soft Parallel Decoding. We represent each intermediate decoding state as an interpolation between the predicted token embedding and the mask embedding, enabling iterative self-revising in embedding space. Extensive experiments across a variety of benchmarks demonstrate the effectiveness of DMax. Compared with the original LLaDA-2.0-mini, our method improves TPF on GSM8K from 2.04 to 5.47 while preserving accuracy. On MBPP, it increases TPF from 2.71 to 5.86 while maintaining comparable performance. On one single H200 GPU using the SGLang framework, our 16B model achieves an average of 955 TPS at a batch size of 1.
Abstract: The aggregation of conflicting client updates remains a fundamental bottleneck in federated learning (FL) over heterogeneous data distributions. Naive averaging, as used in FedAvg, often leads to destructive interference in which the global model improves on average but deteriorates significantly for specific clients. In this work, we propose CRAFT (Conflict-Resolved Aggregation for Federated Training), a new aggregation framework that treats the global update as a geometric correction problem. We formulate aggregation as finding the update closest to a \emphreference direction while satisfying \emphconflict-free constraints, ensuring non-negative alignment with the updates of participating clients. We derive a closed-form expression for the constrained optimization problem, avoiding the computational overhead of iterative solvers. Furthermore, we use a layer-wise adaptation to address conflicts at varying feature granularities. We provide a theoretical analysis showing that CRAFT promotes a common-descent structure and mitigates destructive interference through its projection geometry. Extensive experiments on non-IID and imbalanced benchmarks demonstrate that CRAFT significantly reduces performance disparity while maintaining competitive global accuracy against state-of-the-art baselines.
Abstract: Autoregressive video diffusion models (ARVDs) have emerged as a promising architecture for streaming video generation, paving the way for real-time interactive video generation and world modeling. Despite their potential, the substantial inference cost of ARVDs remains a major obstacle to practical deployment, making model quantization a natural direction for improving efficiency. However, quantization for ARVDs remains largely unexplored. Our empirical analysis shows that directly applying existing quantization schemes developed for standard diffusion transformers to ARVDs leads to suboptimal performance, revealing quantization behaviors that differ from those observed in bidirectional diffusion models. In this paper, we identify two critical challenges in quantizing ARVDs: (C1) Highly unbalanced frame-wise quantization sensitivity. Error accumulation during autoregressive generation can induce severely skewed quantization sensitivity across frames, following an exponential-like decay pattern. (C2) Prominent and heterogeneous outlier patterns in weights. Weight distributions exhibit pronounced outlier channels, whose patterns vary substantially across layer types and block depths. To address these issues, we propose Q-ARVD, a novel framework for accurate ARVD quantization. (S1) To tackle the highly unbalanced frame-wise sensitivity, Q-ARVD incorporates a final-quality aware frame-weighting mechanism into the quantization objective. (S2) To prevent heterogeneous outliers from degrading performance, Q-ARVD introduces an outlier-aware adaptive dual-scale quantization, which automatically detects the presence and quantity of outlier channels for an arbitrary layer, and isolates them to protect normal channels. Extensive experiments on state-of-the-art open-source ARVDs (i.e., self-forcing and causal-forcing) demonstrate the superiority of Q-ARVD. Practical deployment of INT8 model shows 1.30x speedup and 1.97x model size reduction.
Abstract: As AI agents become increasingly autonomous and capable, ensuring their security against vulnerabilities such as prompt injection becomes critical. This paper explores the use of information-flow control (IFC) to provide security guarantees for AI agents. We present a formal model to reason about the security and expressiveness of agent planners. Using this model, we characterize the class of properties enforceable by dynamic taint-tracking and construct a taxonomy of tasks to evaluate security and utility trade-offs of planner designs. Informed by this exploration, we present Fides, a planner that tracks confidentiality and integrity labels, deterministically enforces security policies, and introduces novel primitives for selectively hiding and revealing information. Its evaluation on AgentDojo and WASP demonstrates that this approach enables us to complete a broad range of tasks with security guarantees.
Abstract: Self-speculative decoding accelerates large language model (LLM) inference by drafting tokens from the target model itself, but faces a sharp tradeoff between the quality and cost of the draft. Early-exit methods produce drafts cheaply by terminating computation at intermediate layers, but forgo the deeper representations that later layers provide and thus suffer in draft quality. Multi-token prediction preserves draft quality by emitting from the model's final hidden states, but pays for a full forward pass to produce those states at every drafting step. We propose self-speculative encoder-decoder (SEED), a self-speculative method that obtains high-quality drafts cheaply by reusing the contextual representations already computed during verification. We reinterpret the standard decoder-only transformer as an implicit encoder–decoder: the first layers (encoder) build deep contextual representations, and the last few layers (decoder) emit tokens from them. Encoding and verification are merged into a single step: verification is performed by the full encoder–decoder, and the contextual representations of the verified prefix are cached for reuse during drafting. Drafting is therefore very fast: between verifications, the decoder drafts multiple tokens autoregressively, each conditioned on the cached representations and on preceding drafts. Experiments across multiple benchmarks show that SEED achieves up to 2.5× average speedup on 4B-scale models, outperforming both early-exit and MTP-style self-speculative baselines and running 17% faster than the state-of-the-art EAGLE-3, while preserving or even improving the generation quality of standard autoregressive fine-tuning.
Abstract: Generative modeling over discrete structures underpins applications across deep learning, from biological sequence design and code generation to large language models, yet generation often remains sequential, relying on autoregressive decoding or iterative refinement. We introduce Coupling Models, a one-step discrete generative model that learns a direct coupling between discrete sequences x and Gaussian latents z ~ \mathcalN(0,I). Unlike recent distillation methods that compress a pretrained multi-step sampler into fewer steps, Coupling Models train a purpose-built decoder to invert this coupling and generate samples in a single forward pass. The method also avoids complex continuous flows over the simplex and hand-specified data-to-noise couplings. Empirically, Coupling Models improve the strongest one-step baselines in each domain, reducing LM1B text-generation perplexity by 33% at its lowest-perplexity operating point, Fly Brain enhancer-design FBD by 18%, and MNIST-Binary FID by 46%. These results suggest that effective one-step discrete generation depends strongly on how data and noise are coupled before decoding.
Abstract: Test-Time Training offers a promising way to improve the reasoning ability of large language models (LLMs) by adapting the model using only the test questions. However, existing methods struggle with difficult reasoning problems for two reasons: raw test questions are often too difficult to yield high-quality pseudo-labels, and the limited size of test sets makes continuous online updates prone to instability. To address these limitations, we propose TTCS, a co-evolving test-time training framework. Specifically, TTCS initializes two policies from the same pretrained model: a question synthesizer and a reasoning solver. These policies evolve through iterative optimization: the synthesizer generates progressively challenging question variants conditioned on the test questions, creating a structured curriculum tailored to the solver's current capability, while the solver updates itself using self-consistency rewards computed from multiple sampled responses on both original test and synthetic questions. Crucially, the solver's feedback guides the synthesizer to generate questions aligned with the model's current capability, and the generated question variants in turn stabilize the solver's test-time training. Experiments show that TTCS consistently strengthens the reasoning ability on challenging mathematical benchmarks and transfers to general-domain tasks across different LLM backbones, highlighting a scalable path towards dynamically constructing test-time curricula for self-evolving. Our code and implementation details are available at https://anonymous.4open.science/r/TTCS-6ACE.
Abstract: Recent works have shown that gradient-update alignment is a powerful signal for modulating optimizer updates and improving training dynamics. We promote this update-wise heuristic into a mathematically grounded principle for selecting and tuning optimizer hyperparameters. By treating gradients and updates as signals, and an optimizer as a causal filter that maps between them, we formulate optimizer selection as maximizing the expected drop rate in loss over a prescribed family of optimizers. We show that this objective is exactly the inner product between the optimizer filter and the gradient autocorrelation, and prove that a greedy optimum exists and has a stability bound under perturbations of the estimated gradient statistics. Specializing in momentum-based optimizers, the theory yields simple dynamic momentum selection rules for both SGD+Momentum and Adam/AdamW. Experiments across image classification, language model fine-tuning, and vision transformer fine-tuning show that the resulting dynamic momentum rules match or improve upon the best fixed hyperparameters found via manual sweeps, reducing the need for exhaustive momentum sweeps.
Abstract: . Empirically, this regime often yields improved generalization performance, yet the underlying mechanism remains poorly understood. In this work, we represent stochastic optimizers as (rather than a point) with a smaller intrinsic dimension. Building on this connection and inspired by Lyapunov dimension theory, we introduce a novel notion of dimension, coined the `sharpness dimension', and prove a generalization bound based on this dimension. Our results show that generalization in the chaotic regime depends on the , highlighting a complexity that cannot be captured by the trace or spectral norm considered in prior work. Experiments across various MLPs and transformers validate our theory while also providing new insights into the recently observed phenomenon of grokking.
Authors:
Qi Zhang, Harsh Parikh, Ashley I Naimi, Razieh Nabi, Christopher Kim, Timothy L LashAbstract: Method validation and study design in causal inference rely on synthetic data with known counterfactuals. Existing simulators trade off , including explicit control over overlap, unmeasured confounding, and treatment effect heterogeneity. We introduce , a variational generative framework that closes this gap by coupling a mixture of Gaussian latent priors with data-type-specific decoders for continuous, binary, and categorical variables. The model incorporates explicit causal controls: an overlap regularizer shaping propensity-score distributions, alongside direct parameterizations of confounding strength and effect heterogeneity. This unified objective preserves fidelity to the observed data while enabling factorial manipulation of causal mechanisms, allowing overlap, confounding strength, and treatment effect heterogeneity to be varied independently at design time. Across evaluation settings, achieves strong distributional fidelity on mixed-type tables while providing stable, fine-grained causal control. We demonstrate practical utility in a comparative safety study of metastatic castration-resistant prostate cancer treatments, using to compare estimators under calibrated data-generating processes, tune hyperparameters, and conduct simulation-based power analyses under targeted treatment effect heterogeneity scenarios.
Abstract: Can out-of-distribution (OOD) generalization be predicted from a trained model's weights alone, without any target-domain data? Existing representational similarity metrics (CKA, SVCCA, RSA) compare activations rather than forecast generalization. We show they are provably insensitive to structural rerouting in the computational graph, the very change distribution shift induces. We close this gap with the Circuit Alignment Score (CAS), which compares class-specific circuits across domains via graph kernels, decomposed into same-class coherence and cross-class confusion. Casting CAS as a Lebesgue integral over the domain distribution, we prove its Monte Carlo estimate recovers the ground-truth ranking of learners by OOD accuracy, with pairwise inversion error vanishing at rate O(1/M), where M is the number of sampled domains. Across 48 learners on PACS, CAS attains 0.88 rank correlation with OOD accuracy, versus 0.58 (CKA), 0.23 (SVCCA), and 0.14 (RSA), with similar trends on other benchmarks and even against data-dependent methods, making it the first provably consistent predictor of distributional robustness requiring neither target-domain data nor labels.
Abstract: Spatial-temporal (ST) forecasting underpins many real-world systems such as traffic, climate, and energy networks. While existing methods implicitly assume strong spatiotemporal coupling, we observe that real-world ST data exhibits distinct coupling regimes, ranging from temporal-dominated and spatial-dominated to strongly coupled patterns. This mismatch causes current models to suffer from spurious dependencies and degraded performance when one correlation dominates. To overcome this limitation, we aim to dynamically modulate spatial and temporal modeling based on the data's inherent coupling structure. However, three key challenges exist: unknown coupling structure, heterogeneous coupling dynamics, and suboptimal spatial modeling. We propose AdaST, an adaptive ST forecasting framework that tackles these challenges through a decompose-recompose paradigm. AdaST factorizes inputs into components capturing different coupling patterns using heterogeneity-aware experts. Each component is processed by role-aligned modules, and a correlation-informed adaptive recomposer integrates them for final prediction. Extensive experiments confirm that AdaST significantly outperforms state-of-the-art baselines, validating the necessity of an adaptive approach.
Authors: Yequan Zhao, Ruijie (Ray) Zhang, Liyan Tan, Niall Moran, Tong Qin, Zheng Zhang
Abstract: Both full fine-tuning (Full FT) and parameter-efficient methods like LoRA add weight updates without regard to the spectral structure that pretraining has established. This allows noisy gradients from a small fine-tuning distribution to freely perturb the robust features learned through pretraining. We first identify spectral preconditioning as the key missing ingredient: reparameterizing each weight \mathbfW through its full-rank SVD and freezing one singular basis confines every update to the pretrained column space, yielding a preconditioned optimizer that outperforms unconstrained Full FT at the same parameter count. To make this insight practical, we propose FuRA (Full-Rank Adaptation), which factorizes \mathbfW via a block tensor-train decomposition \mathbfW=\mathbfL\mathbfS\mathbfR: the large core \mathbfL is frozen at the pretrained block-wise SVD basis while only the small core \mathbfR and per-block singular values \mathbfS are trained. This single design choice simultaneously delivers full-rank spectral preconditioning, full-rank update capacity, and parameter, step time, memory efficiency on par with LoRA. FuRA outperforms Full FT on LLM fine-tuning (+1.37 on LLaMA-3-8B commonsense reasoning), LLM math reinforcement learning, and VLM visual instruction tuning. The 4-bit quantized version QFuRA also outperforms QLoRA.
Authors: Bartłomiej Baranowski, Dave Chen, Matthias Niessner
Abstract: Existing approaches to 3D scene understanding in Vision-Language Models (VLMs) either rely on complex, model-specific geometry encoders or large training budgets in pursuit of spatial reasoning. Instead, OneCanvas aggregates patch features from all views onto a single equirectangular panoramic canvas. Namely, each patch is unprojected to a 3D world coordinate using its depth and camera pose, then placed on the canvas at the continuous longitude and latitude of that point as seen from the canvas origin, with no rasterization or aggregation across overlapping views. A 3D position embedding of the patch's metric coordinates is added to its feature, restoring the depth lost when collapsing the world position to an angular canvas coordinate. Patches from all frames thus share one spatial coordinate system with no fusion or major architectural modifications of the backbone. The pretrained VLM consumes this representation as if it were an ordinary image. Because the canvas can be centered on any pose of interest, the same representation directly supports situated reasoning from a specific viewpoint, a common requirement in robotics and embodied AI. Thanks to this representation, we can also introduce a spatial pretraining curriculum: by procedurally placing patch features of objects, drawn from real images, at chosen 3D world positions on an otherwise empty canvas, we generate on-the-fly supervision spanning a broad range of spatial reasoning tasks, with answer distributions controlled to reduce spatial reasoning shortcuts. OneCanvas achieves state-of-the-art accuracy on SQA3D and VSI-Bench, and generalizes to out-of-distribution data on SPBench, using an order of magnitude less training compute than the strongest competing methods.
Abstract: Inversion-based visual editing provides an effective and training-free way to edit an image or a video based on user instructions. Existing methods typically inject source image information during the sampling process to maintain editing consistency. However, this sampling strategy overly relies on source information, which negatively affects the edits in the target image (e.g., failing to change the subject's atributes like pose, number, or color as instructed). In this work, we propose ProEdit to address this issue both in the attention and latent aspects. In the attention aspect, we introduce KV-mix, which mixes KV features of the source and the target in the edited region, mitigating the influence of the source image on the editing region while maintaining background consistency. In the latent aspect, we propose Latents-Shift, which shifts the distribution of the edited region in the inverted noise, eliminating the negative influence of the inverted noise on the sampling process. Extensive experiments on several image and video editing benchmarks demonstrate that our method achieves SOTA performance. In addition, our design is plug-and-play, which can be seamlessly integrated into existing inversion and editing methods, such as RF-Solver, FireFlow and UniEdit.
Abstract: Despite the strong performance of large language models (LLMs) across diverse tasks, their susceptibility to adversarial attacks and unsafe content generation remains a significant barrier to deployment, particularly in high-stakes settings. Addressing this challenge requires safety mechanisms that are both practically effective and theoretically grounded. In this paper, we introduce BarrierSteer, a novel framework that improves response safety by embedding learned nonlinear safety constraints directly into the model's latent representation space. BarrierSteer treats hidden-state safety classifiers as Control Barrier Functions (CBFs), enabling constraint-guided steering of unsafe latent trajectories during generation. By composing multiple safety constraints through efficient constraint merging without modifying the underlying LLM parameters, BarrierSteer preserves model utility and performance. We provide theoretical results showing that applying CBFs in latent space yields a principled and computationally efficient approach for steering with respect to learned safety constraints, with guarantees conditional on the learned barriers capturing the intended safety property. Extensive experiments across multiple models and datasets demonstrate that BarrierSteer substantially reduces adversarial attack success rates and unsafe generations, outperforming existing methods.
Authors: Boyao Wang, Zhihan Lei
Abstract: Modular networks pursue specialization through learned routers, gates, and load-balancing losses. However, at matched total-parameter budgets, learned routers can underperform equal-weight No-Routing baselines. Is the bottleneck the routing algorithm, or the alignment between training-signal granularity and the target categories? We show that when each training and inference unit carries one coarse category tag, a fixed parameter-free routing scheme (SpecDrop) induces branch-category specialization and matches or exceeds the learned-routing baselines we evaluate at matched parameter budget; when this alignment breaks, the scheme matches the multi-branch No-Routing baseline. SpecDrop assigns each of K branches weight p_\mathrma for its category and a small leakage p_\mathrmi>0 otherwise, merged through a category-independent fixed denominator that calibrates merged-branch magnitude to single-branch scale, with no learnable routing parameters or auxiliary losses. Category labels are required at inference, drawn from dataset metadata. On vision tasks where each image has one superclass label, SpecDrop reaches \mathbf79.23% on CIFAR-100 with ResNet-110, outperforming dense by +4.75, and \mathbf79.89% on ImageNet-1K with Vision Transformer ViT-S/16, outperforming the matched-supervision No-Routing baseline with a shared expert by +6.53. SpecDrop reaches the highest top-1 among the multi-branch routing baselines we evaluate at this parameter budget (Soft MoE and Mod-Squad on ImageNet). On NLP tasks where training units span multiple categories, SpecDrop matches the multi-branch No-Routing baseline on both SlimPajama-6B language modeling with a 30M Transformer and SuperNI instruction tuning over Llama-3.2-1B with LoRA. Furthermore, SpecDrop matches or outperforms multi-branch routing baselines including Demix and LoRAMoE, and is statistically tied with HydraLoRA. Per-branch pruning sensitivity reveals four specialization regimes ranked by category clarity: strongest on ImageNet-1K, strong on CIFAR-100, weakening on SlimPajama, and anti-aligned on SuperNI LoRA. Granularity alignment, not algorithm choice, localizes when routing helps. Code: https://anonymous.4open.science/r/C30862.
Abstract: Estimating physical pressure from vision is essential for understanding contact-rich hand-object interaction. However, prior vision-based pressure estimation methods are largely limited to planar surfaces and single image input, making them difficult to apply to dynamic hand-object interaction with diverse objects. We instead formulate pressure estimation as a hand-centric video prediction problem with monocular video as input. This formulation predicts temporally evolving per-vertex normal pressure and contact directly on the hand mesh, yielding a unified output space independent of object shape and sensor layout. Building on this formulation, we propose HOPE, a framework with two key components. First, we lift tactile-glove pressure, planar-sensor pressure, and distance-based hand-object contact annotations into a shared hand vertex space, allowing bare-hand contact data to regularize pressure learning where metric labels are unavailable. Second, we introduce a vertex-anchored video transformer that treats each vertex as a persistent token, aggregates visual features and hand pose over time, and uses a contact-gated pressure head to enforce that pressure vanishes without contact. Experiments on OpenTouch, PressureVisionDB, and hand-object contact benchmarks demonstrate that HOPE supports a common hand-centric representation across object-pressure, surface-pressure, and contact-supervised HOI settings. Despite using metric pressure supervision primarily from gloved-hand videos, HOPE generalizes to bare-hand egocentric and in-the-wild videos, producing joint contact and pressure predictions beyond the scope of contact-only or planar-pressure baselines.
Abstract: The extension of context windows in Large Language Models is typically facilitated by scaling positional encodings followed by Continued Pre-Training (CPT). While effective, this paradigm is notoriously data-hungry and computationally expensive, requiring massive long-text corpora to recalibrate the model to the shifted positional distribution. We propose LinearARD, a self-distillation method that restores Rotary Position Embedding (RoPE)-scaled students through attention-structure consistency with a frozen native-RoPE teacher. Rather than next-token prediction or opaque hidden-state matching, LinearARD aligns row-wise distributions of dense Q/Q, K/K, and V/V self-relation matrices from the final attention layer to directly supervise attention dynamics. To remove the quadratic memory bottleneck of n × n relation maps, we introduce a linear-memory kernel that stores only per-token log-sum-exp statistics and recomputes logits in the backward pass to obtain exact Kullback-Leibler divergence gradients. Across LLaMA2-7B, LLaMA3-8B, and Mistral-7B-v0.1 extended to 32K context, LinearARD recovers 93.1%/94.2%/94.3% of native short-context performance and achieves strong long-context robustness on RULER using only 4.25M tokens---amounting to just 1.6% (a ~60× reduction) of the 256M-token budget required by state-of-the-art baselines. Under this severely constrained budget, CPT and LongReD remain near zero on RULER, demonstrating that relation-level restoration provides a substantially better efficiency-quality tradeoff.
Abstract: Cross-modal place recognition aims to retrieve a target 3D location from a spatial map using a natural language description. Most existing methods follow a global-descriptor matching paradigm, in which the text query and each 3D scene cell are independently compressed into single vectors and compared by global similarity. Although simple and widely adopted, this paradigm tends to discard the fine-grained object correspondences that are essential for language-guided localization. In this paper, we challenge the necessity of global descriptors and propose OLA-Place, an Object-Level Alignment framework for cross-modal place recognition without global descriptors. Instead of representing a query or a scene cell as a single holistic embedding, OLA-Place formulates place recognition as set-to-set semantic alignment between textual object mentions and 3D object instances. The framework contains three key modules: ObjectSet Encoder (OSE), which extracts object-level representations from both language descriptions and 3D scene cells; Object Message Encoder (OME), which injects intra-set object-context information while preserving object-level granularity; and Masked Max Alignment (MMA), which computes the query-cell matching score by aligning each textual object mention with its most similar valid 3D object. This formulation is permutation-invariant, naturally handles variable-size object sets, and requires no object-level correspondence annotations. Extensive experiments show that object-level alignment alone significantly outperforms global-descriptor-based methods, demonstrating that cross-modal place recognition is better understood as an object-level semantic correspondence problem rather than a global-embedding retrieval problem. Our code is available at: https://github.com/Anonymous09871745/OLA-Place.
Authors: Zhuoyue Zhang, Yiding Sun, Chaowei Fang, Haozhe Cheng, Jian Liu, Jihua Zhu, Ajmal Mian
Abstract: Dynamic point cloud pretraining is still dominated by masked reconstruction objectives. However, these objectives inherit two key limitations. Existing methods inject ground-truth tube centers as decoder positional embeddings, causing spatio-temporal positional leakage. Moreover, they supervise inter-frame motion with deterministic proxy targets that systematically discard distributional structure by collapsing multimodal trajectory uncertainty into conditional means. To address these limitations, we propose Diffusion Masked Pretraining (DiMP), a unified self-supervised framework for dynamic point clouds. DiMP introduces diffusion modeling into both positional inference and motion learning. It first applies forward diffusion noise only to masked tube centers, then predicts clean centers from visible spatio-temporal context. This removes positional leakage while preserving visible coordinates as clean temporal anchors. DiMP also reformulates point-wise inter-frame displacement supervision as a DDPM noise-prediction objective conditioned on decoded representations. This design drives the encoder to target the full conditional distribution of plausible motions under a variational surrogate, rather than collapsing to a single deterministic estimate. Extensive experiments demonstrate that DiMP consistently improves downstream accuracy over the backbone alone, with absolute gains of 11.21% on offline action segmentation and 13.65% under causally constrained online inference.
Abstract: Safe and efficient shape-aware navigation in heterogeneous crowds and robot fleets remains challenging. Traditional approaches often assume homogeneous robots, sparse workspaces, simplified geometry, offline computation, or handcrafted parameters to make the problem tractable, which limits their deployment in dense crowd scenarios. Toward this end, we propose Shape-Aware Reinforcement Learned Model Predictive Control (SRL-MPC), a method for safe, efficient, and adaptive navigation in crowds with heterogeneous shapes without geometry simplification. To encode shape-aware safety, we formulate high-order control barrier function (HOCBF) constraints from geometric separation features (GSFs) based on support function transformation. A reinforcement learning (RL) framework then learns a neural policy that reads GSFs and outputs real-time MPC parameter updates, enabling the MPC solver to adapt to neighboring crowd geometries. The key advantage of SRL-MPC is that it preserves the safety structure and generalizability of MPC while integrating the adaptability and intelligence of RL. Experiments in randomized crowd scenarios with arbitrary shaped robot fleets demonstrate the effectiveness, scalability, and robustness of SRL-MPC. The results show that SRL-MPC substantially outperforms representative baselines in safety and adaptability.
Abstract: Large Language Models (LLMs) demonstrate strong capability in solving scientific and mathematical problems, yet they struggle to produce valid and challenging novel problems, an essential component for advancing LLM training and enabling autonomous scientific research. Existing problem generation approaches either depend on expensive human expert involvement or adopt naive self-play paradigms, which frequently yield invalid problems due to reward hacking. This work introduces VHG, a verifier-enhanced hard problem generation framework built upon three-party self-play. By integrating an independent verifier into the conventional setter-solver duality, our design constrains the setter’s reward to be jointly determined by problem validity (evaluated by the verifier) and difficulty (assessed by the solver). We instantiate two verifier variants: a Hard symbolic verifier and a Soft LLM-based verifier, with evaluations conducted on indefinite integral tasks and general mathematical reasoning tasks. Experimental results show that VHG substantially outperforms all baseline methods by a clear margin. Our codes are anonymously available at https://github.com/VHGMath/VHG_Math.
Abstract: Ante-hoc interpretability methods based on prototypes provide highly accurate explanations by utilizing the intuitive "this looks like that" reasoning paradigm. On the other hand, post-hoc models can explain predictions for a single image without relying on an underlying dataset or requiring costly neural network retraining. Recent approaches successfully solve the retraining problem for prototype-based networks. However, they still face a fundamental limitation: they require access to a subset of data (e.g., a test or validation set) to search for and extract the visual prototypes. In this paper, we address this issue and introduce ProDG: Data-Free Generative Prototypes for Post-Hoc Explainability, a novel framework that leverages generative models to synthesize pure, high-fidelity prototypes directly from the frozen model's weights, completely eliminating the dependency on any external data. By establishing this new frontier in Data-Free XAI, DGP unlocks robust visual interpretability for privacy-sensitive domains, where original data is strictly restricted or fundamentally inaccessible.
Authors: Ruihang Zhang, Felix Taubner, Pooja Ravi, Kyros Kutulakos, David Lindell
Abstract: Tracking the six-degree-of-freedom (6-DoF) pose of objects and surfaces from monocular video is a long-standing problem in computer vision. To tackle this problem, existing methods require inputs beyond the video itself—such as 3D models, depth maps, object masks, or task-specific learned features—and they struggle with textureless, transparent, reflective, or deformable surfaces. Here, we introduce ProxyPose, which recasts 6-DoF pose tracking as video-to-video translation. Given only a video and a single marked pixel in the first frame, a fine-tuned video diffusion model translates the input into a "proxy video"—a synthetic video depicting a colored polyhedron undergoing the same local rigid-body motion as the surface region at the marked pixel. Because the proxy's geometry and appearance are known by construction, recovering its full 6-DoF trajectory reduces to classical pose estimation with off-the-shelf solvers. This formulation leverages large-scale video pre-training to absorb the hardest aspects of pose tracking—handling challenging materials, occlusions, and deformations—into the translation step, while operating at the pixel level with no assumptions about object identity, boundaries, or global rigidity. ProxyPose achieves state-of-the-art 6-DoF pose tracking accuracy without the additional inputs required by competing methods and after fine-tuning the video model only on synthetic data. We further demonstrate that ProxyPose extends to face tracking, camera pose estimation, and challenging in-the-wild scenes that are beyond the reach of existing approaches. Video results are available on the
Abstract: While adaptive gradient methods are the workhorse of modern machine learning, sign-based optimization algorithms such as Lion and Muon have recently demonstrated superior empirical performance over AdamW in training large language models (LLM). However, a theoretical understanding of why sign-based updates outperform variance-adapted methods remains elusive. In this paper, we aim to bridge the gap between theory and practice through the lens of heavy-tailed gradient noise, a phenomenon frequently observed in language modeling tasks. Theoretically, we introduce a novel generalized heavy-tailed noise condition that captures the behavior of LLMs more accurately than standard finite variance assumptions. Under this noise model, we establish sharp convergence rates of SignSGD and Lion for generalized smooth function classes, matching or surpassing previous best-known bounds. Furthermore, we extend our analysis to Muon and Muonlight, providing what is, to our knowledge, the first rigorous analysis of matrix optimization under heavy-tailed stochasticity. These results offer a strong theoretical justification for the empirical superiority of sign-based optimizers, showcasing that they are naturally suited to handle the noisy gradients associated with heavy tails. Empirically, LLM pretraining experiments validate our theoretical insights and confirm that our proposed noise models are well-aligned with practice.
Authors: Marvin Philip kleine Sextro, Weronika Kłos, Gabriel Dernbach
Abstract: Planning effective interventions in biological systems requires treatment-effect models that adapt to unseen biological contexts by identifying their specific underlying mechanisms. Yet single-cell perturbation datasets span only a handful of biological contexts, and existing methods cannot leverage new interventional evidence at inference time to adapt beyond their training data. To meta-learn a perturbation effect estimator, we present MapPFN, a prior-data fitted network (PFN) pre-trained on a synthetic biological prior with causal interventions, decoupling pre-training from limited wet-lab data. Unlike existing methods, MapPFN uses in-context learning to map a sequence of experiments to a post-perturbation distribution, enabling a single pre-trained model to adapt to new datasets and arbitrary gene sets at inference time. Zero-shot, MapPFN identifies differentially expressed genes on par with models trained on real single-cell data, and fine-tuning further improves predictions across biological contexts. Our code is available at https://anonymous.4open.science/r/MapPFN.
Abstract: Test-Time Training (TTT) enables long-context processing via continuous weight updates during inference, but current methods struggle to balance the expressivity of per-token update dynamics with the hardware efficiency of chunk-wise approximations. We propose E^2-TTT (Expressive and Efficient TTT) to bridge this gap. By deriving a closed-form state transition that exactly aggregates per-token momentum and decay coefficients within a chunk, E^2-TTT enables fully parallelized chunk-level training while preserving the temporal structure of the update rule that prior chunk-wise methods discard. We validate E^2-TTT by training models up to 1.3B parameters from scratch. Extensive experiments demonstrate that our method consistently outperforms previous TTT and hybrid attention baselines in language modeling and retrieval, while achieving significantly better extrapolation on the standard ``Needle in a Haystack'' test, maintaining >90 % accuracy on passkey retrieval at 8× the training context length. Meanwhile, E^2-TTT can match the training throughput of efficient chunk-wise methods, demonstrating that it effectively reconciles expressivity with efficiency.
Abstract: Being robust to the presence of outliers is crucial for applying clustering algorithms in practice. In the robust k-Means problem (i.e., k-Means with outliers), the goal is to remove z outliers and minimize the k-Means cost on the remaining points. Despite the close connection between robust k-Means and outlier detection, both theoretical and empirical understanding of the effectiveness of classic outlier detection heuristics for robust k-Means remains limited. In this paper, we prove that under a practical assumption on the optimal cluster sizes, simply removing points with large K-Nearest-Neighbor distances achieves performance comparable to prior work in terms of approximation guarantees: it yields a constant-factor reduction from robust k-Means to standard k-Means, without introducing additional centers or discarding extra outliers, as is commonly required by existing approaches. Empirically, experiments on real-world datasets show that our method outperforms or matches several more sophisticated algorithms in terms of clustering cost and runtime. These results demonstrate that simple KNN-based heuristics can be surprisingly effective for robust clustering, highlighting new opportunities to bridge techniques from outlier detection and clustering.
Abstract: Diffusion Large Language Models (dLLMs) have achieved rapid progress, viewed as a promising alternative to the autoregressive paradigm. However, most dLLM decoders still adopt a global confidence threshold, and do not explicitly model local context from neighboring decoded states or temporal consistency of predicted token IDs across steps. To address this issue, we propose a spatio-temporal stability guided decoding approach, named STDec. We observe strong spatio-temporal stability in dLLM decoding: newly decoded tokens tend to lie near decoded neighbors, and their predicted IDs often remain consistent across several denoising steps. Inspired by this stability, our STDec includes spatial-aware decoding and temporal-aware decoding. The spatial-aware decoding dynamically generates the token-adaptive threshold by aggregating the decoded states of nearby tokens. The temporal-aware decoding relaxes the decoding thresholds for tokens whose predicted token IDs remain consistent over denoising steps. Our STDec is training-free and remains compatible with cache-based acceleration methods. Across textual reasoning and multimodal understanding benchmarks, STDec substantially improves throughput while maintaining comparable task performance score. Notably, on MBPP with LLaDA, STDec achieves up to 14.17× speedup with a comparable score. Our code will be available.
Abstract: Single-image 3D methods often trade off pixel alignment and completeness: monocular reconstructors stay registered to visible pixels but stop at the first surface, while image-to-3D generators complete shapes in a canonical frame that is no longer aligned to the input. We introduce methodlong (method), a pixel-aligned multilayer geometry representation that keeps generated 3D in the input camera frame while completing geometry beyond the visible surface. For each input ray, method predicts an ordered stack of camera-space 3D points, where the first entry explains the observed surface and subsequent entries complete occluded surfaces along the same pixel ray. methodnet, a flow-matching diffusion transformer with a frozen 2D foundation encoder, learns this representation from depth-filled dense targets, removing the need for a separate per-layer visibility-mask head. The same objective covers static objects, static scenes, and dynamic clips, with temporal attention added for video. Because the output preserves the input camera frame, intrinsics, and pixel correspondences, it provides a direct 2D-to-3D interface for downstream systems, including text-driven 3D scene editing, geometry-conditioned novel-camera video generation, and training-free integration with existing textured-mesh generators. Across object, scene, and dynamic benchmarks, method improves visible-surface accuracy and complete-geometry quality over strong reconstruction and generation baselines.
Abstract: Unlike standard expected-return Reinforcement Learning (RL), Distributional RL (DRL) models the full return distribution, making it better-suited for uncertainty-aware and risk-sensitive decision-making. Conditional Flow Matching (CFM) critics have recently attracted attention for modelling continuous, multi-modal return distributions. Despite this interest, there remains a substantial metric mismatch: DRL theory relies on the distributional Bellman operator being contractive in the p-Wasserstein distance, yet existing CFM critics are trained with arbitrary source--target couplings, so their flow-matching losses are not Wasserstein-aligned surrogates for matching Bellman target return distributions. In this work, we address this mismatch by proposing FlowIQN, a CFM critic that sorts source and Bellman target samples within each mini-batch to approximate the monotone optimal transport coupling, replacing arbitrary pairings with quantile-aligned flow paths. We prove that the loss of our quantile-coupled CFM critic yields a Wasserstein-aligned approximate projection compatible with the foundations of DRL. To our knowledge, FlowIQN is the first flow-matching distributional critic with an explicit Wasserstein-aligned projection guarantee. We further extend FlowIQN with shortcut models for efficient inference. Empirical results show that FlowIQN improves Wasserstein return-distribution accuracy over other CFM critics. It also yields competitive performance on offline RL benchmarks across multiple policy extraction methods, providing a theoretically grounded CFM critic that is readily compatible with DRL pipelines.
Abstract: Analyzing the generalization of graph neural networks (GNNs) on node classification is challenging: message passing induces dependencies that break the i.i.d. assumption underlying standard inductive generalization theory. While distribution-free transductive theory resolves this dependency issue, existing classical bounds are either intractable or correlate poorly with empirical generalization. To fill this gap, we propose a transductive generalization bound based on optimal transport, connecting generalization to intra-class concentration and inter-class separation in the feature space. Empirically, our bound aligns strongly with the generalization gap. Building on this representation-based approach, we analyze how feature geometry evolves under repeated message passing. This explains the non-monotonic relationship between GNN depth and generalization, providing new insights into oversmoothing research. Practically, our bound serves as a geometry-aware regularizer that corrects the harmful behavior of message passing, yielding consistent performance improvements.
Abstract: Pose and motion priors play a crucial role in humanoid robotics. Although such priors have been widely studied in human motion recovery (HMR) domain with a range of models, their adoption for humanoid robots remains limited, largely due to the scarcity of high-quality humanoid motion data. In this work, we introduce Pose Distance Fields for Humanoid Robots (PDF-HR), a lightweight prior that represents the robot pose distribution as a continuous and differentiable manifold. Given an arbitrary pose, PDF-HR predicts its distance to a large corpus of retargeted robot poses, yielding a smooth measure of pose plausibility that is well suited for optimization and control. PDF-HR can be integrated as a reward shaping term, a regularizer, or a standalone plausibility scorer across diverse pipelines. We evaluate PDF-HR on various humanoid tasks, including single-trajectory motion tracking, general motion tracking, style-based motion mimicry, and general motion retargeting. Experiments show that this plug-and-play prior consistently and substantially strengthens strong baselines.
Authors:
Xiyang Wang, Xinlin Wang, Tingguang Zhou, Gong Chen, Xingtai Gui, Zhi Xu, Hangning Zhou, Xiaolei Wu, Feiyang Tan, Mu YangAbstract: Current end-to-end autonomous driving systems are fundamentally limited by a mismatch between temporal causal reasoning and global trajectory consistency. Autoregressive (AR) models capture interaction-aware temporal dependencies via causal factorization, but their step-wise decoding leads to error accumulation and suboptimal global structure. In contrast, diffusion models optimize trajectories globally but lack explicit causal constraints, making them unreliable in interactive and safety-critical scenarios. This dichotomy reveals a deeper issue: existing methods treat causal modeling and global optimization as separate paradigms, without a principled way to unify them within a single trajectory distribution. To address this, we propose ChainFlow-VLA, which unifies causal generation and global refinement within a unified probabilistic framework. We formulate planning as a mixture over AR-induced modes and learn VLM-conditioned residual distributions over these modes. An autoregressive generator (Chain) produces a discrete set of causal trajectory modes, followed by a diffusion-based refiner (Flow) that operates in residual space to perform mode-conditioned correction while preserving causal structure. A key insight is that vision-language models are more effective as semantic controllers for refinement rather than direct trajectory generators. By conditioning the diffusion process on VLM hidden states, scene-level reasoning guides fine-grained trajectory adjustments within each mode. This formulation enables robust planning in ambiguous and long-tail scenarios. ChainFlow-VLA achieves a state-of-the-art score of 94.8 on the NAVSIM v1 leaderboard, matching human-level performance (94.8).
Abstract: Pre-training Large Language Models requires immense computational resources, making optimizer efficiency essential. The optimization landscape is highly anisotropic, with loss reduction driven predominantly by progress along flat directions. While matrix-based optimizers such as Muon and SOAP use fine-grained curvature information to outperform AdamW, their updates tend toward isotropy, which is relatively conservative along flat directions yet potentially aggressive along sharp ones. To address this limitation, we first establish a unified Riemannian Ordinary Differential Equation (ODE) framework that clarifies the mechanism of adaptive algorithms: the preconditioner induces a Riemannian geometry that mitigates ill-conditioning, while momentum serves as a Riemannian damping term that promotes convergence. Guided by these insights, we propose LITE, a generalized acceleration strategy that enhances training dynamics by applying larger Hessian damping coefficients and learning rates along flat trajectories. Extensive experiments demonstrate that LITE significantly accelerates both Muon and SOAP across diverse architectures (Dense, MoE), parameter scales (130M to 1.3B), datasets (C4, Pile), and learning-rate schedules (cosine, wsd). Theoretical analysis confirms that LITE facilitates faster convergence along flat directions in anisotropic landscapes, providing a principled approach to efficient LLM pre-training.
Authors:
Tao Liu, Hao Yan, Mengting Chen, Taihang Hu, Zhengrong Yue, Zihao Pan, Jinsong Lan, Xiaoyong Zhu, Ming-Ming Cheng, Bo Zheng, Yaxing WangAbstract: Step distillation has become a leading technique for accelerating diffusion models, among which Distribution Matching Distillation (DMD) and Consistency Distillation are two representative paradigms. While consistency methods enforce self-consistency along the full PF-ODE trajectory to steer it toward the clean data manifold, vanilla DMD relies on sparse supervision at a few predefined discrete timesteps. This restricted discrete-time formulation and mode-seeking nature of the reverse KL divergence tends to exhibit visual artifacts and over-smoothed outputs, often necessitating complex auxiliary modules---such as GANs or reward models---to restore visual fidelity. ), migrating the DMD framework from discrete anchoring to continuous optimization for the first time. CDM achieves this through two continuous-time designs. First, we replace the fixed discrete schedule with a dynamic continuous schedule of random length, so that distribution matching is enforced at arbitrary points along sampling trajectories rather than only at a few fixed anchors. Second, we propose a continuous-time alignment objective that performs active off-trajectory matching on latents extrapolated via the student's velocity field, improving generalization and preserving fine visual details. Extensive experiments on different architectures, including SD3-Medium and Longcat-Image, demonstrate that CDM provides highly competitive visual fidelity for few-step image generation without relying on complex auxiliary objectives.
Abstract: Open-loop imitation learning has advanced modern autonomous driving policy architectures, but closed-loop deployment remains vulnerable to policy-induced distribution shift. Existing post-training paradigms exhibit fundamental trade-offs: closed-loop RL fine-tuning provides grounded feedback from executed actions but is constrained by the sparsity of informative events, whereas counterfactual fine-tuning provides dense supervision over candidate futures but inherits bias from imperfect future estimates. We introduce Counterfactual-to-Interactive Reinforcement Fine-Tuning (CRAFT), an on-policy framework that formulates closed-loop post-training as proxy-residual optimization. CRAFT uses group-normalized counterfactual advantages as a dense proxy for real closed-loop advantages and aligns this proxy with the closed-loop world through grounded residual correction from interaction-critical events. To stabilize adaptation, CRAFT regularizes the online policy toward an EMA teacher via asymmetric KL self-distillation. Theoretically, CRAFT decomposes the real closed-loop policy gradient into proxy and residual terms under the same visited-state distribution, reducing residual variance with an aligned proxy while mitigating proxy bias through grounded residual approximation. Empirically, CRAFT achieves the strongest closed-loop gains on Bench2Drive across hierarchical planning, vision-language-action, and vocabulary-scoring architectures. Ablations, scaling behavior, stability analyses, and transfer results further validate the complementary roles of dense counterfactual proxy and grounded residual correction.
Abstract: Large reasoning models (LRMs) achieve remarkable performance by leveraging reinforcement learning (RL) on reasoning tasks to generate long chain-of-thought (CoT) reasoning. However, this over-optimization often prioritizes compliance, making models vulnerable to harmful prompts. To mitigate this safety degradation, recent approaches rely on external teacher distillation, yet this introduces a \emphdistributional discrepancy that degrades native reasoning. We formalize safety realignment as a KL projection onto the safe simplex and prove that the student's own safety-filtered distribution is the unique KL-optimal target, while any external teacher incurs an irreducible excess KL penalty. Guided by this analysis, we propose ThinkSafe, a self-generated alignment framework that restores safety without external teachers. Our key insight is that while compliance suppresses safety mechanisms, models often retain latent knowledge to identify harm. ThinkSafe unlocks this via lightweight refusal steering, which preserves the KL-optimal target while increasing the acceptance rate. Experiments on DeepSeek-R1-Distill and Qwen3 show ThinkSafe significantly improves safety while preserving reasoning proficiency, and achieves superior safety and comparable reasoning to GRPO with roughly an order of magnitude less compute. Code, models, and datasets are available at this
Abstract: Differential privacy (DP) has traditionally been used to provide theoretical upper bounds on an algorithm's stability to changing its training data. In modern machine learning applications, achieving strong tradeoffs between utility and theoretical privacy is challenging, and thus one may optimistically hope that existing theoretical analyses are loose. Recent work on \emphprivacy auditing has adopted a dual viewpoint, instead lower bounding the true privacy of an algorithm by constructing empirical distinguishing events. The auditing literature has thus far yielded a pessimistic outlook on the looseness of theoretical privacy bounds for DP-SGD, the de facto private training method in modern ML, as nearly-matching empirical lower bounds have been achieved under various threat models \citeNasrHSBTJCT23, AnnamalaiC24, CebereBP25. In this work, we propose the empirical privacy lower bound of an algorithm as a concrete metric to optimize for, complementary to the theoretical upper bound. We give a lightweight defense framework that generically augments optimization methods in the ML pipeline to have significantly-improved empirical privacy on standard benchmarks. Moreover, we show that our framework comes at \emphno theoretical privacy cost when augmenting DP-SGD, unlike previous defenses against membership inference attacks. We evaluate our defense against a broad range of audit constructions, models, and datasets to demonstrate its flexibility. Our code implementation can be found at \hrefhttps://github.com/neurips26privacyaudit/empirical-privacy-defensethis anonymous repository.
Abstract: The growing deployment of Large Language Models (LLMs) has raised concerns about their misuse in generating harmful or deceptive content. To address this issue, watermarking methods have been proposed to embed identifiable multi-bit messages into generated text for misuse tracing. However, existing methods often suffer from a fundamental trade-off between text quality and decoding accuracy. In particular, they have to restrict the size of the preferred token set (i.e., green list) during encoding to maintain a detectable watermark signal for decoding, which inevitably degrades generation quality. To improve this trade-off, we propose a novel message encoding paradigm called majority bit-aware encoding, which relaxes the watermark signal strength from the green list size. This strategy allows for a strong watermark signal to be preserved in generated texts even when using a large green list. We introduce two instantiations of this paradigm: MajorMark and MajorMark^+, where the latter is specifically optimized for long messages. Extensive experiments on state-of-the-art LLMs demonstrate that our methods achieve higher decoding accuracy and superior text quality compared to prior baselines. The code of the proposed methods is available \hrefhttps://anonymous.4open.science/r/MajorMarkhere for review.
Abstract: Free energy estimation is a fundamental yet challenging problem, from physics to statistics. Classical approaches rely on thermodynamic transformations, ranging from direct estimation, quasistatic integration, to finite-time averaging. Recent work learns neural transports to significantly accelerate the efficiency in the finite-time regime. In this paper, we generalize this framework to arbitrary state spaces. Building on this view, we develop a generalized neural transport learning approach for efficient estimation. Experiments validate the effectiveness and efficiency of the proposed method beyond continuous settings, extending to discrete and multimodal spaces as well as autoregressive settings.
Abstract: Out-of-distribution (OOD) detection identifies test samples that fall outside a model's training distribution, a capability critical for safe deployment in high-stakes applications. Standard OOD detectors are trained on a specific in-distribution (ID) dataset and detect deviations from that single domain. In contrast, we study few-shot cross-domain OOD detection: given a \emphsingle pre-trained model, can we perform OOD detection on \empharbitrary new ID-OOD task pairs using only a handful of ID samples at inference time, with no additional training? We propose UFCOD, a unified framework that achieves this goal through information-geometric analysis of diffusion trajectories. Our key insight is that diffusion noise predictions are score functions (gradients of log-density), and we extract two energy features: \emphPath Energy (integrated score magnitude) and \emphDynamics Energy (score smoothness), that form a discrete Sobolev norm capturing how samples interact with the learned diffusion process. The central contribution is a train-once, deploy-anywhere paradigm: a diffusion model trained on a single dataset serves as a universal feature extractor for OOD detection across semantically unrelated domains. At deployment, each new task requires only ~100 unlabeled ID samples for inference: no retraining, no fine-tuning, no task-specific adaptation. Using 100 ID samples per task, UFCOD achieves 93.7% average AUROC across 12 cross-domain benchmarks, competitive with methods trained on 50k--163k samples, demonstrating ~500× improvement in sample efficiency. See our code in \urlhttps://anonymous.4open.science/r/UFCOD-22E8/README.md.
Abstract: Traditional rendering pipelines rely on complex assets, accurate materials and lighting, and substantial computational resources to produce realistic imagery, yet they still face challenges in scalability and realism for populated dynamic scenes. We present C2R (Coarse-to-Real), a generative rendering framework that synthesizes real-style urban crowd videos from coarse 3D simulations. Our approach uses coarse 3D renderings to explicitly control scene layout, camera motion, and human trajectories, while a learned neural renderer generates realistic appearance, lighting, and fine-scale dynamics guided by text prompts. To overcome the lack of paired training data between coarse simulations and real videos, we adopt a two-stage synthetic-real domain-hedging strategy that first learns a strong generative prior from large-scale real footage, then introduces controllability by using a small amount of paired synthetic coarse-fine data to anchor shared implicit spatio-temporal features across domains. The resulting system supports coarse-to-fine control, generalizes across diverse CG and game inputs, and produces temporally consistent, controllable, and realistic urban scene videos from minimal 3D input.
Abstract: In-context learning is highly sensitive to demonstration choice, yet most methods select demonstrations using external query--demonstration similarity. Such criteria can miss model-specific signals: Similar demonstrations may activate different internal features and downstream behaviors. We introduce ncodings), an SAE-based framework for identifying model-internal features associated with demonstration utility and using them for demonstration selection. Using a small labeled discovery set, PULSE samples candidate demonstration sets, measures their zero-shot-relative utility under the target model, and scores SAE features by how their activation differences align with utility differences. The top positive and negative coordinates form a sparse utility-localization vector. We use this vector in two complementary ways: as a signed score for controlled complete-set ranking, and as , which converts its magnitude into a feature-relevance mask for scalable pool-scale retrieval. Across classification, generation, and reasoning benchmarks, PULSE-Retriever improves over the strongest baseline by 2--3 accuracy points, 0.6--0.9 BLEU-4, and 3.2 exact-match points, respectively, while controlled ranking validates the identified features encode a predictive set-level utility signal. Feature inspection and cross-dataset experiments suggest that the identified features capture task-relevant, dataset-conditioned patterns, yet retain utility signals that partially transfer across datasets. Our code is available at the
Abstract: Video diffusion models (VDMs) have achieved impressive progress in text-to-video generation, but their high memory and computational costs hinder practical deployment. Quantization-aware training (QAT) is an effective solution for compressing and accelerating advanced generative models without runtime overhead at inference. % However, existing QAT methods suffer from a distinctive challenge in VDMs: while they often preserve prompt semantics, global layout, and coarse motion, the quantized model severely degrades visual details, texture fidelity, and sharpness. In this paper, we trace this degradation to the timestep-agnostic design of conventional quantization pipelines, which overlooks the stage-wise functionality of video denoising. In VDMs, early denoising steps mainly establish global structure and motion, whereas middle and late steps refine local appearance and high-frequency details. Based on this insight, we propose DSAQuant, a Denoising-Stage-Aligned Quantization-aware training framework for VDMs. During training, Denoising-Stage Oriented Supervision preserves teacher distillation in early steps for stable structure planning, while shifting later steps toward target-driven optimization to enhance detail reconstruction. During inference, Denoising-Stage Gated Guidance disables CFG in the final denoising steps to prevent it from amplifying quantization-induced errors into high-frequency artifacts. Extensive experiments on the Wan and CogVideoX families under W4A4 and W3A3 settings show that DSAQuant consistently outperforms the SOTA QAT baseline, improving the VBench average score by up to 6.60 under aggressive W3A3 quantization while preserving strong text-video alignment. These results demonstrate that effective VDM quantization requires not only reducing quantization error, but also aligning quantization training and inference with the stage-wise nature of video diffusion.
Authors:
Haoran He, Yuxiao YE, Jie Liu, Jiajun Liang, Zhiyong Wang, Ziyang Yuan, Xintao Wang, Hangyu Mao, Meng Wang, Pengfei Wan, Ling PanAbstract: Fine-tuning diffusion models via online reinforcement learning (RL) has shown great potential for enhancing text-to-image alignment. However, since precisely specifying a ground-truth objective for visual tasks remains challenging, the models are often optimized using a proxy reward that only partially captures the true goal. This mismatch often leads to reward hacking, where proxy scores increase while real image quality deteriorates and generation diversity collapses. While common solutions add regularization against the reference policy to prevent reward hacking, they compromise sample efficiency and impede the exploration of novel, high-reward regions, as the reference policy is usually sub-optimal. To address the competing demands of sample efficiency, effective exploration, and mitigation of reward hacking, we propose Gated and Adaptive Regularization with Diversity-aware Optimization (GARDO), a versatile framework compatible with various RL algorithms. Our key insight is that regularization need not be applied universally; instead, it is highly effective to selectively penalize a subset of samples that exhibit high uncertainty. To address the exploration challenge, GARDO introduces an adaptive regularization mechanism wherein the reference model is periodically updated to match the capabilities of the online policy, ensuring a relevant regularization target. To address the mode collapse issue in RL, GARDO amplifies the rewards for high-quality samples that also exhibit high diversity, encouraging mode coverage without destabilizing the optimization process. Extensive experiments across diverse proxy rewards and hold-out unseen metrics consistently show that GARDO mitigates reward hacking and enhances generation diversity without sacrificing sample efficiency or exploration, highlighting its effectiveness and robustness.
Authors:
Ge Wu, Minxing Luo, Yikai Ge, Lei Wang, DanDan Zheng, Rui Liu, libin wang, Yu-Liang Zhan, Taihang Hu, Jingdong Chen, Xiang LiAbstract: Residual connections with PreNorm are the default design in modern transformers due to their strong optimization stability, but recent research has shown that they weaken the functional differentiation of deeper blocks and increase representational redundancy in later layers. This issue also appears in diffusion models using similar architectures, where the deep block is expected to support progressive denoising refinement rather than repeated computation over similar states. Motivated by this, we revisit cross-layer shortcuts in diffusion transformers and argue that their role should go beyond merely transferring shallow information, instead guiding deeper blocks to improve depth utilization. To this end, we propose Residual Adaptive Layer Normalization (ResNorm), which routes residual cross-layer shortcuts into the existing adaLN modulation. ResNorm preserves the optimization advantages of PreNorm residual connections, while reusing the native adaLN interface to transmit early-layer representations as modulation signals. This design naturally combines cross-layer history with the model's native condition, such as timestep and class embeddings. ResNorm allows the earlier structural representation to adaptively scale, shift, and gate deeper blocks. This not only produces more differentiated deep representations and reduces cross-layer similarity in later blocks, but also accelerates training convergence and improves performance. On ImageNet 256×256, SiT-XL/2 + ResNorm achieves up to 2× and 36× faster training than SiT-XL/2 + REG and SiT-XL/2 + REPA. Moreover, SiT-L/2 + ResNorm trained for only 400K iterations already outperforms SiT-XL/2 + REG.
Abstract: Sparse dictionary learning (and, in particular, sparse autoencoders) attempts to learn a set of human-understandable concepts that can explain variation on an abstract space. A basic limitation of this approach is that it neither exploits nor represents the semantic relationships between the learned concepts. In this paper, we introduce a modified SAE architecture that explicitly models a semantic hierarchy of concepts. Application of this architecture to the internal representations of large language models shows both that semantic hierarchy can be learned, and that doing so improves both reconstruction and interpretability. Additionally, the architecture leads to significant improvements in computational efficiency.
Authors:
Tunyu Zhang, Xinxi Zhang, Ligong Han, Haizhou Shi, Xiaoxiao He, Zhuowei Li, Hao Wang, Kai Xu, Akash Srivastava, Chengzhi Mao, Hao Wang, Vladimir Pavlovic, Dimitris MetaxasAbstract: Diffusion large language models (DLLMs) have emerged as powerful generative models with the promise of fast text generation through parallel decoding. However, realizing this potential in practice remains challenging: reducing the number of decoding steps, typically causes a substantial degradation in output quality due to token factorization error. To alleviate this, we propose a self-distillation framework that trains a few-step student to match the \emphgenerative trajectory of a full-step teacher. We theoretically and empirically show that trajectory-level supervision mitigates this factorization error, thereby enabling effective few-step decoding. We further incorporate Direct Discriminative Optimization (DDO), a reverse-KL objective that encourages mode-seeking toward the teacher’s modes, yielding stronger performance on challenging reasoning tasks. Across reasoning and code-generation benchmarks, our method substantially narrows the gap between few-step and full-step decoding.
Abstract: Transformer architectures serve as the backbone for most modern Large Language Models, therefore their pretraining stability and convergence speed are of central concern. Motivated by the logical dependency of sequentially stacked layers, we propose Progressive Residual Warmup (ProRes) for language model pretraining. ProRes implements an "early layer learns first" philosophy by multiplying each layer's residual with a scalar that gradually warms up from 0 to 1, with deeper layers taking longer warmup steps. In this way, deeper layers wait for early layers to settle into a more stable regime before contributing to learning. We demonstrate the effectiveness of ProRes through pretraining experiments across various model scales, as well as normalization and initialization schemes. Comprehensive analysis shows that ProRes not only stabilizes pretraining but also introduces a unique optimization trajectory, leading to faster convergence, stronger generalization and better downstream performance. Our code is available in supplementary materials.
Authors:
Christoffer Koo Ohrstrom, Rafael I Muchacho, Yifei Dong, Filippos Moumtzidellis, Ronja Güldenring, Florian T. Pokorny, Lazaros NalpantidisAbstract: We propose Parabolic Position Encoding (PaPE), a parabola-based position encoding for vision modalities in attention-based architectures. Given a set of vision tokens–such as from videos, event camera streams, images, or point clouds–our objective is to encode their positions while accounting for the characteristics of vision modalities. Prior works have largely extended position encodings from 1D-sequences in language to nD-structures in vision, but only with partial account of vision characteristics. We address this gap by designing PaPE from principles distilled from prior work: translation invariance, rotation invariance (PaPE-RI), distance decay, directionality, and context awareness. Extrapolation experiments on ImageNet-1K show how PaPE extrapolates remarkably well, improving in absolute terms by up to 10.5% over the next-best encoding. Generality experiments on 8 datasets across 4 modalities show that PaPE is a general vision position encoding, as PaPE matches the best baseline on 5 datasets and exceeds all on 2 datasets.
Authors:
Hao Wang, Guozhi Wang, Han Xiao, YufengZhou, Yue Pan, Jichao Wang, Ke Xu, yafei wen, Xiaohu Ruan, xiaoxin chen, Honggang QiAbstract: Reinforcement learning (RL) has been widely used to train LLM agents for multi-turn interactive tasks, but its sample efficiency is severely limited by sparse rewards and long horizons. On-policy self-distillation (OPSD) alleviates this by providing dense token-level supervision from a privileged teacher that has access to ground-truth answers. However, such fixed privileged information cannot capture the diverse valid strategies in agent tasks, and naively combining OPSD with RL often leads to training collapse. To address these limitations, we introduce Skill-SD, a framework that turns the agent's own trajectories into dynamic training-only supervision. Completed trajectories are summarized into compact natural language skills that describe successful behaviors, mistakes, and workflows. These skills serve as dynamic privileged information conditioning only the teacher, while the student always acts under the plain task prompt and learns to internalize the guidance through distillation. To stabilize the training, we derive an importance-weighted reverse-KL loss to provide gradient-correct token-level distillation, and dynamically synchronize the teacher with the improving student. Experimental results on agentic benchmarks demonstrate that Skill-SD substantially outperforms the standard RL baseline, improving both vanilla GRPO (+14.0%/+10.9% on AppWorld/Sokoban) and vanilla OPD (+42.1%/+40.6%).
Authors: deyuan liu, PENG SUN, Xufeng Li, Tao Lin
Abstract: Diffusion transformers learn useful internal representations during training, and learning them from scratch can add a burden to efficient training. Alignment mitigates this burden by adding an auxiliary loss that regularizes diffusion hidden states with features from pretrained visual encoders. In this work, we revisit the role of this alignment signal. Rather than using pretrained representations only as a regularizer during full diffusion training, we use them as a preconditioning signal before diffusion training. The key observation is that early layer alignment from clean latents to pretrained representations can be much cheaper than optimizing the full flow matching objective over the entire backbone, yet provides a better initialization for subsequent diffusion training. We instantiate this idea as Embedded Representation Warmup (ERW). ERW first aligns early layers from clean latents to pretrained representations, and then trains the full model with the standard diffusion objective and decaying alignment. This initialization allows subsequent diffusion training with alignment to converge more efficiently than starting alignment from scratch. Counting warmup in the total budget, ERW improves convergence in SiT based diffusion training. On ImageNet 256×256, ERW reaches FID=1.41 in 350 epochs with SiT-XL/2; it also reaches FID=2.04 on ImageNet 512×512 and improves FID on MS-COCO text-to-image generation.
Abstract: Image tokenizers play a central role in modern generative models, where the structure of the latent space critically determines the downstream generation performance. A key but underexplored property of effective latent representations is spectral organization: the ability to encode information across frequency components. In this work, we introduce structured state-space regularization, a principled approach for inducing spectral structure in latent spaces. We derive a regularization objective by revisiting state-space models (SSMs) as dynamical systems over compressed representations. This perspective reveals that hidden states of SSMs follow basis function dynamics under predefined input transformations, resulting in a novel regularizer form that enforces the latent space to capture spectral components of images. Experiments demonstrate that our regularizer improves the generative performance of image tokenizers while incurring only minimal loss in their reconstruction fidelity.
Abstract: Recent 3D geometric foundation models, such as VGGT, provide robust feed-forward 3D reconstruction by directly predicting camera poses and 3D scene points from input images. However, their results remain inaccurate, and scaling them to long sequences or large unordered image sets typically requires chunk-wise processing, which can introduce drift and inconsistency. We present Glob3R, a global SfM-style reconstruction built on 3D foundation models. Our key idea is to explicitly optimize feed-forward geometric predictions. To this end, we augment a frozen Pi3X backbone with a lightweight dense matching head that predicts image warps between selected reference frames and neighboring views. These dense warps are converted into sparse but reliable multi-view feature tracks, which provide correspondence constraints for global optimization. We further introduce a keyframe-based sliding-window association strategy that propagates tracks and relative poses across overlapping windows, enabling scalable reconstruction. Finally, we perform global motion averaging and bundle adjustment to refine camera poses, reduce scale inconsistencies, and recover dense scene geometry. Extensive experiments on indoor, outdoor, large-scale driving, and unordered SfM benchmarks demonstrate that Glob3R achieves robust and accurate reconstruction. It consistently improves over feed-forward foundation-model baselines and recent scalable reconstruction methods, while being more robust than classical SfM pipelines. The refined poses also lead to higher-quality neural rendering, validating the benefit of combining foundation-model priors with global geometric optimization.
Abstract: Omni-modal reasoning is pivotal to realizing artificial general intelligence, yet its advancement is critically constrained by the limited availability of large-scale, human-annotated data for complex reasoning. Motivated by this, we propose OmniJigsaw, a self-supervised reinforcement learning framework built upon a temporal reordering proxy task. Centered on the chronological reconstruction of shuffled audio-visual clips, we successively investigate three modality orchestration strategies: (i) Joint Modality Integration (JMI), which retains the complete visual and auditory streams; (ii) Sample-level Modality Selection (SMS), which selects the dominant modality through a global decision mechanism; and (iii) Clip-level Modality Masking (CMM), which adaptively masks modalities at the clip granularity. Our analysis reveals a ``bi-modal shortcut phenomenon'' in JMI and demonstrates that fine-grained CMM mitigates this issue while outperforming SMS. Incorporating lightweight puzzle-quality curation and verifiable rewards, OmniJigsaw yields substantial gains across 15 video, audio, and omni-modal benchmarks, with CMM achieving the strongest overall performance, validating its effectiveness as a self-supervised paradigm for omni-modal learning.
Abstract: Vision-language models are designed to capture the compositional structure of the world through multi-modal alignment; however, it remains unclear whether their representations truly align with the way a human would naturally describe the same visual input. Prior works attempt to decompose visual embeddings into their concept-level contributions, but their instance-level optimization is inherently local and does not capture the model’s global semantic structure, which is essential for compositional generalization across different inputs, tasks and domains. We introduce GRACE, a method that learns this concept-based structure at the feature space level, producing semantic and general concept decompositions. To achieve this, we propose a novel training objective that preserves the geometric relations of different samples in the concept space, enabling a better understanding of how concepts compose and relate to each other. We validate GRACE on several image and video benchmarks, showing that its concept decomposition is aligned with human perception and achieves faithful decompositions across different tasks and domains. With minimal overhead, GRACE makes existing Vision-Language Models more interpretable, while preserving most of their original performance.
Abstract: Contrastive Language-Image Pre-training (CLIP) is fundamentally limited by its 77-token restriction, lack of multilingual capabilities, and coarse-grained semantic representations. While replacing CLIP’s native text encoder with a Large Language Model (LLM)-based embedder offers a promising solution, direct, from-scratch alignment can disrupt the semantic structures learned during pre-training. This often leads to degradation of cross-modal knowledge and particularly impairs recognition capabilities in zero-shot scenarios. In this paper, we propose ProCLIP, a curriculum-learning-inspired progressive alignment framework designed to systematically bridge the LLM embedder and CLIP's visual space, unlocking CLIP's potential for long-text, multilingual, and fine-grained understanding. Our framework operates in two stages: (1) Representation Inheritance, which distills CLIP's original text-space knowledge into an LLM adapter to establish a robust initial vision-language prior, and (2) Contrastive Tuning, which refines the cross-modal alignment with self-distillation regularization on the image encoder to further prevent catastrophic forgetting. To maintain semantic and geometric consistency, we introduce instance-level semantic and global structural alignment constraints throughout both stages. Extensive experiments demonstrate that ProCLIP improves zero-shot classification accuracy by 6.8%--13.5% over existing LLM-augmented baselines and achieves state-of-the-art performance in diverse long-text, multilingual, and fine-grained cross-modal retrieval tasks under comparable settings.
Abstract: Long-horizon robot planning requires jointly reasoning over semantic task structure and geometric feasibility. To successfully execute a task, a robot must decompose goals, select task-relevant objects, and sequence actions, while ensuring that plans satisfy spatial constraints such as limited free space and object collisions. In this work, we propose APIVOT, a VLM-based planner that adaptively interleaves language and visual thoughts for long-horizon planning. APIVOT learns to rely on language for semantic reasoning, while using visual thoughts as imagined future states for internal verification of geometric feasibility. On long-horizon kitchen tasks, APIVOT outperforms general-purpose VLMs and prior planning frameworks, achieving the largest gains in spatially constrained settings. We find that APIVOT learns meaningful modality selection behavior, demonstrating that adaptive interleaving of vision-language thoughts improves both planning success and reasoning efficiency.
Authors: Lei Li, Yuexiao Dong
Abstract: We show that replacing the standard MSE denoising loss in diffusion models with a nonlinear transformation induced by an f-divergence yields a simple robust training surrogate that empirically improves performance under data contamination, with small additional computational overhead. The theoretical foundation rests on a local divergence construction: under the Gaussian reverse-kernel structure of DDPM, each per-step likelihood ratio follows a lognormal distribution parameterized by a scalar mismatch, so the conditional f-divergence at each step reduces to a one-dimensional function of the denoising error. Summing these local divergences yields a training objective that unifies diffusion training as divergence-induced weighted denoising, where the derivative of the induced divergence acts as a residual-space influence weight that controls the contribution of each sample. Bounded-influence divergences (Hellinger, negative exponential) suppress large-error samples, with Hellinger yielding an explicit exponential weight, connecting the framework to robust M-estimation. Empirically, on CIFAR-10 under 30% contamination, NED reduces FID from 93.0 (KL) to 77.5, while also outperforming standard robust losses such as Huber and clipped MSE.
Abstract: Steering methods influence Large Language Model behavior by identifying semantic directions in hidden representations, and are typically realized through inference-time activation interventions that apply a fixed, global modification to the model's internal states. While effective, such interventions often induce unfavorable attribute--utility trade-offs under strong control, as they ignore the fact that many behaviors are governed by a small and heterogeneous subset of model components. To alleviate the trade-offs, we propose Steer2Edit, a theoretically grounded, training-free framework that transforms steering vectors from inference-time control signals into diagnostic signals for component-level rank-1 weight editing. Instead of uniformly injecting a steering direction during generation, Steer2Edit selectively redistributes behavioral influence across individual attention heads and MLP neurons, yielding interpretable edits that preserve the standard forward pass and remain compatible with optimized parallel inference. Across multiple tasks including safety alignment, truthfulness promotion, and reasoning efficiency, Steer2Edit consistently achieves a superior attribute--utility trade-offs: at matched downstream performance, it outperforms the strongest steering baseline by 17.2% in safety, 9.8% in truthfulness, and 12.2% in reasoning length reduction. Overall, Steer2Edit provides a principled bridge between representation steering and weight editing by translating steering signals into interpretable, training-free parameter updates.
Abstract: Modern LLM deployments use a number of implementation choices and inference optimizations (e.g., batching, custom kernels, quantization) on top of fixed weights, so two engines serving "the same model" can produce meaningfully different distributions. We study the problem of estimating the total variation (TV) distance between two length-n autoregressive distributions to additive error \varepsilon, under three access models: 1. Under sample access, we use \widetildeO(\fracn^2K\varepsilon^2) queries, where K is the maximum support of the next-token distribution. 2. Under logit access, we use O(\fracn\varepsilon^2) queries, and this is tight. 3. Under noisy logit access, we can smoothly interpolate between the above two: if probability values are given to \sigma relative error, we use \widetildeO(\fracn + n^2\sigma^2\varepsilon^2) queries. We complement our theoretical results with an empirical evaluation of our algorithms, for example measuring the distance between \textttsglang and \textttvllm on standard settings. Our experiments highlight the robustness and practicality of estimating the total variation distance, compared to existing alternatives such as KL divergence.
Abstract: Diffusion large language models (dLLMs) have shown advantages in text generation, particularly due to their inherent ability for parallel decoding. However, constrained by the quality--speed trade-off, existing inference solutions adopt conservative parallel strategies, leaving substantial efficiency potential underexplored. A core challenge is that parallel decoding assumes each position can be filled independently, but tokens are often semantically coupled. Thus, the correct choice at one position constrains valid choices at others. Without modeling these inter-token dependencies, parallel strategies produce deteriorated outputs. Motivated by this insight, we propose DAWN, a training-free, dependency-aware decoding method for fast dLLM inference. DAWN extracts token dependencies and leverages two key observations: (1) masked positions dependent on unmasked certain positions become more reliable; (2) coupled masked positions strongly influence each other's predictions. Given those findings, DAWN leverages a dependency graph to select more reliable unmasking positions at each iteration, achieving high parallelism with negligible loss in generation quality. Extensive experiments across multiple models and datasets demonstrate that DAWN speedups the inference by
Abstract: Multi-shot video generation extends single-shot generation to coherent visual narratives, yet maintaining consistent characters, objects, and locations across shots remains a challenge over long sequences. Existing evaluations typically use independently generated prompt sets with limited entity coverage and simple consistency metrics, making standardized comparison across methods difficult. We introduce EntityBench, a benchmark consisting of 140 episodes (2,491 shots) derived from real narrative media, with explicit per-shot entity schedules tracking characters, objects, and locations simultaneously across easy, medium, and hard difficulty tiers of up to 50 shots, 13 cross-shot characters, 8 cross-shot locations, 22 cross-shot objects, and recurrence gaps spanning up to 48 shots. EntityBench pairs the dataset with a three-pillar evaluation framework that disentangles intra-shot visual quality, prompt-following alignment, and cross-shot entity consistency. Cross-shot consistency, the central pillar, evaluates each recurring entity through both embedding similarity and LLM per-criterion judgment across entity-type-specific dimensions, with a fidelity gate that admits accurate entity appearance. To establish baselines, we propose EntityMem, a memory-augmented generation system that plans and stores verified per-entity visual references in a persistent memory bank before generation begins, enabling the video backbone to retrieve each entity's appearance across shots. Experiments on EntityBench show that cross-shot entity consistency degrades sharply with recurrence distance in existing methods, and that explicit per-entity memory yields the highest character fidelity (Cohen's d = +2.33) and presence among methods evaluated.
Abstract: Recent image generation models have shown strong capabilities in generating high-fidelity and photorealistic images. However, they are fundamentally constrained by frozen internal knowledge, thus often failing on real-world scenarios that are knowledge-intensive or require up-to-date information. In this paper, we present Gen-Searcher, as the first attempt to train a search-augmented image generation agent, which performs multi-hop reasoning and search to collect the textual knowledge and reference images needed for grounded generation. To achieve this, we construct a tailored data pipeline and curate two high-quality datasets, Gen-Searcher-SFT-10k and Gen-Searcher-RL-6k, containing diverse search-intensive prompts and corresponding ground-truth synthesis images. We further introduce KnowGen, a comprehensive benchmark that explicitly requires search-grounded external knowledge for image generation and evaluates models from multiple dimensions. Based on these resources, we train Gen-Searcher with SFT followed by agentic reinforcement learning with dual reward feedback, which combines text-based and image-based rewards to provide more stable and informative learning signals for GRPO training. Experiments show that Gen-Searcher brings substantial gains, improving Qwen-Image by around 16 points on KnowGen and 15 points on WISE. We hope this work can serve as an open foundation for search agents in image generation, and we fully open-source our data, models, and code.
Authors: Pengxi Liu, Zeyu M Li, Xiang Cheng
Abstract: We introduce a variational framework for diffusion models with anisotropic noise schedules with a matrix-valued path M_t(\theta) that allocates noise across subspaces. Central to our framework is a trajectory-level objective that jointly trains the score network and learns M_t(\theta), which encompasses general parameterization classes of matrix-valued noise schedules. We further derive an estimator for \partial_\theta \nabla \log p_t that enables efficient optimization of the M_t(\theta) schedule. For inference, we develop an efficiently-implementable reverse-ODE solver that is an anisotropic generalization of the second-order Heun discretization algorithm. Across CIFAR-10, AFHQv2, FFHQ, and ImageNet-64, our method consistently improves upon the baseline EDM model in all NFE regimes.
Abstract: Video diffusion models have made rapid progress in perceptual realism and temporal coherence, but they remain primarily optimized for plausible generation rather than verifiable reasoning. This limitation is especially pronounced in tasks where generated videos must satisfy explicit spatial, temporal, or logical constraints. Inspired by the role of reinforcement learning with verifiable rewards (RLVR) in reasoning-oriented language models, we introduce , a practical recipe for optimizing video diffusion models with rule-based feedback. VideoRLVR formulates video reasoning as the generation of verifiable visual trajectories and consists of an SDE-GRPO optimization backbone, dense decomposed rewards, and an strategy for efficient training. The Early-Step Focus strategy restricts policy optimization to the early denoising phase, reducing training latency by about 40% while preserving performance. We evaluate VideoRLVR on Maze, FlowFree, and Sokoban, three procedurally generated domains with objective success criteria. Across these tasks, VideoRLVR consistently improves over supervised fine-tuning baselines, with dense decomposed rewards proving especially important in low-success-rate settings. Our RL-optimized model also outperforms the evaluated proprietary and open-source video generation models on these verifiable reasoning benchmarks and out-of-domain benchmarks. These results suggest that verifiable RL can move video models beyond perceptual imitation toward more reliable rule-consistent visual reasoning.
Abstract: Pixel diffusion models have recently regained attention for visual generation. However, training advanced pixel-space models from scratch demands prohibitive computational and data resources. To address this, we propose the Latent-to-Pixel (L2P) transfer paradigm, an efficient framework that directly harnesses the rich knowledge of pre-trained LDMs to build powerful pixel-space models. Specifically, L2P discards the VAE in favor of large-patch tokenization and freezes the source LDM's intermediate layers, exclusively training shallow layers to learn the latent-to-pixel transformation. By utilizing LDM-generated synthetic images as the sole training corpus, L2P fits an already smooth data manifold, enabling rapid convergence with zero real-data collection. This strategy allows L2P to seamlessly migrate massive latent priors to the pixel space using only 8 GPUs. Furthermore, eliminating the VAE memory bottleneck unlocks native 4K ultra-high resolution generation. Extensive experiments across mainstream LDM architectures show that L2P incurs negligible training overhead, yet performs on par with the source LDM on DPG-Bench and reaches 93% performance on GenEval.
Abstract: Activation steering offers a lightweight alternative to fine-tuning for controlling large language models at inference time. While many existing methods implicitly optimize a log-density-ratio objective between desired and undesired activation distributions, they do so heuristically rather than deriving it from a principled optimization problem. Moreover, these methods produce query-independent steering directions that can degrade performance on both in-distribution and out-of-distribution (OOD) inputs. We introduce COBRAS (Conditional Optimal Bridge for Riemannian Activation Steering), which addresses both limitations by casting activation steering as a Schr\"odinger Bridge on the residual-stream hypersphere. This formulation yields, to our knowledge, the first principled derivation of the log-density-ratio steering objective from a well-posed optimization problem. Solving the bridge via entropic optimal transport and extracting the probability flow ODE recovers the widely used density-ratio gradient as a special case when the Sinkhorn potentials are uniform. Crucially, the Schr\"odinger potentials are evaluated at the current activation, making the resulting steering direction inherently query-adaptive. Empirically, across four models and three alignment axes (helpfulness, truthfulness, and detoxification), COBRAS consistently outperforms prior activation steering baselines while avoiding the OOD degradation commonly observed in existing methods.
Abstract: In joint multiuser decoding, a receiver recovers a set of messages from a single noisy aggregate of many simultaneous transmissions. Classical decoders rely on rule-based mechanisms such as successive interference cancellation, joint belief propagation, or list recovery, all of which become brittle or expensive as ambiguity increases. We propose CIDER, a learned multiuser decoder with masked-diffusion refinement steps. CIDER uses demixing to prevent duplicate-row collapse and uses parity-aware propagation to provide soft guidance from the code constraints. In higher-load regimes, we further improve reliability via a lightweight quality-guided remasking step that selectively re-decodes low-confidence sequences. On commonly used error correcting codes, CIDER matches or improves on FFT-accelerated joint belief propagation-style decoding in symbol error rate while running more than 6× to over 100× faster, with the speedup widening as the blocklength grows. Code is available at https://anonymous.4open.science/r/CIDER_2026/.
Authors:
Chong Cheng, Peilin Tao, Nanjie Yao, Guanzhi Ding, Xianda Chen, yuansen du, Xiaoyang Guo, Wei Yin, Weiqiang Ren, Qian Zhang, Zhengqing Chen, Hao WangAbstract: Online 3D reconstruction requires estimating camera pose and scene geometry under strict causal and bounded-memory constraints. Existing methods often suffer from drift, jitter, or collapse on long sequences. We trace these failures to a fundamental mismatch. Streaming geometry is inherently temporally heterogeneous, with evidence ranging from short-lived correspondences to persistent global scale. However, current architectures impose uniform and pathological influence patterns. For example, sliding windows enforce hard cutoffs, while ungated recurrence and causal attention cause cache saturation and spike-like attention sinks. To resolve this, we formalize geometric propagation as an \emphevidence influence kernel and propose Horizon-Stream, a long-horizon Transformer that explicitly factorizes this kernel. For the long-range temporal factor, Geometric Linear Attention learns channel-wise decay rates to enable bounded, multi-timescale propagation of geometric evidence. For the short-range spatial factor, Geometric Local Attention with Spatiotemporal RoPE performs reliable 3D matching while suppressing attention sinks. Finally, Metric Readout Tokens recover stable scale and rigid pose directly from the persistent geometric state. Extensive experiments show that Horizon-Stream, trained on only 48-frame clips, generalizes stably to sequences exceeding (10,000) frames with constant memory and linear time, achieving state-of-the-art streaming 3D reconstruction performance.
Authors:
Chenxu Guo, Jiachen Lian, Yisi Liu, Baihe Huang, Shriyaa Narayanan, Bixing Wu, Zoe Ezzes, Jet Vonk, Zachary Miller, Cheol Jun Cho, Maria Luisa Gorno Tempini, Gopala AnumanchipalliAbstract: We propose HuPER, a human-inspired framework that models phonetic perception as adaptive inference over acoustic--phonetic evidence and linguistic knowledge. HuPER first learns an acoustic-grounded phone recognizer from limited human-annotated data and transcript-only speech: canonical G2P transcriptions are treated as auxiliary linguistic cues rather than ground truth, and are corrected into realized phone proxies for self-training. At inference time, HuPER supports multiple perceptual routes: it can rely on bottom-up phone evidence when the signal is clear, incorporate explicit expectations when a reference is available, or use lexical constraints when acoustic evidence is weak. With only 100 hours of training data, HuPER achieves state-of-the-art phonetic feature error rates on five English benchmarks and strong zero-shot transfer to 95 unseen languages. HuPER also improves robustness on weak-evidence and disordered speech, demonstrating the benefit of adaptive, multi-path phonetic perception under diverse acoustic and task conditions. All training data, models, and code are open-sourced. Code and demo are available at https://github.com/HuPER29/HuPER.
Abstract: Fair comparison between diffusion-based OOD detectors is challenging, as conclusions can vary with backbone choice, corruption parameterization, and test-time budget. We address this issue through a Mutualized Backbone-Equated (MBE) protocol that aligns canonical corruption levels and logical test-time cost across diffusion backbones. Within this setting, we introduce Canonical Feature Snapshots (CFS), a family of detectors that probes a frozen diffusion backbone using only a tiny number of native internal activations at canonical low-noise levels. On a controlled CIFAR-scale benchmark, the strongest one-forward CFS variant is CFS(1×2), while an even smaller decoder-only variant remains highly competitive. This shows that much of the relative-OOD signal exposed by frozen diffusion backbones is concentrated in a small number of sparse internal states, rather than requiring full denoising trajectories or high-capacity downstream heads. We further provide a local diagnostic theory explaining these observations through conditional encoder-decoder complementarity, diagonal-score separation, and low-noise corruption stability.
Abstract: Many public buildings provide floorplans with a “you are here” indicator to help visitors orient themselves. Floorplan localization seeks to computationally replicate this capability by determining where visual observations were captured within a floorplan. However, existing methods typically assume controlled small-scale environments and precise vectorized floorplans, limiting their ability to operate in large-scale buildings and rasterized floorplans. In this work, we present an approach for performing floorplan localization in the wild by grounding the task in a reconstructed 3D representation of the scene. Given an unconstrained image collection, our method reconstructs a gravity-aligned 3D scene and projects it into a 2D density map that serves as a floorplan proxy. Floorplan localization is then formulated as aligning this proxy with the input floorplan via a 2D similarity transform. To bridge the appearance gap between density maps and architectural floorplans, we adapt a 2D foundation model to learn cross-modal correspondences, introducing a fine-tuning scheme that encourages semantically aligned matches while preserving structural consistency. Extensive experiments demonstrate substantial improvements over prior methods, including in extremely sparse settings with as little as a single input image. Our code and data will be publicly available.
Authors:
Zhimu Zhou, Yanpeng Zhao, Qiuyu Liao, Bo Zhao, Xiaojian (Shawn) MaAbstract: Visual planning represents a crucial facet of human intelligence, especially in tasks that require complex spatial reasoning and navigation. Yet, in machine learning, this inherently visual problem is often tackled through a verbal-centric lens. While recent research demonstrates the promise of fully visual approaches, they suffer from significant computational inefficiency due to the step-by-step planning-by-generation paradigm. In this work, we present EAR, an editing-as-reasoning paradigm that reformulates visual planning as a single-step image transformation. To isolate intrinsic reasoning from visual recognition, we employ abstract puzzles as probing tasks and introduce AMAZE, a procedurally generated dataset that features the classical Maze and Queen problems, covering distinct, complementary forms of visual planning. The abstract nature of AMAZE also facilitates automatic evaluation of autoregressive and diffusion-based models in terms of both pixel-wise fidelity and logical validity. We assess leading proprietary and open-source editing models. The results show that they all struggle in the zero-shot setting, finetuning on basic scales enables remarkable generalization to larger in-domain scales and out-of-domain scales and geometries. However, our best model that runs on high-end hardware fails to match the zero-shot efficiency of human solvers, highlighting a persistent gap in neural visual reasoning.
Authors:
Omer Sela, Inbar Huberman-Spiegelglas, Michael Rotman, Sagie Benaim, Avi Ben-CohenAbstract: Controlling the motion of multiple objects in image-to-video (I2V) generation requires preserving object identities while enforcing adherence to distinct target trajectories. This becomes particularly challenging as the number of objects increases and their paths intersect or occlude one another. Existing approaches entangle multiple trajectories within a shared, dense conditioning signal, making object-level correspondence difficult to preserve in crowded scenes. We depart from this paradigm and enforce a strict, per object spatial constraint that isolates instances independently. Our method, TrajLoc, achieves this directly within the attention layers by substituting the cross-attention weights of each object token with a Gaussian heatmap centered on its target location at every frame. The same per object token interface carries trajectory and depth through a learned embedding and preserves identity by encoding first frame appearance in place of an object token. Evaluations across six datasets, featuring up to 20 simultaneously controlled objects and out of distribution real world scenes, demonstrate that our method consistently improves both visual fidelity and trajectory adherence. Applied to two architecturally distinct backbones (CogVideoX 5B and WaN 2.1 14B), our approach achieves average gains of +4.3 dB PSNR and a 51% reduction in trajectory end point error compared to the strongest baselines. Code will be made publicly available upon publication.
Abstract: Modern image generative models produce high-quality, prompt-faithful images, yet the same prompt can correspond to different desired outputs for different users. We study \emphpreference-conditioned image generation: generating images that follow a target prompt while reflecting a user's latent visual preferences inferred from liked and disliked histories. This problem is challenging because real preference histories are sparse and costly to collect, user tastes mix semantic and stylistic factors that are difficult to verbalize, and diffusion generators are not naturally designed to condition on multi-image preference histories. We propose \textscPrefGen, a framework for learning structured preference representations and injecting them into diffusion-based generators. To address the supervision bottleneck, we construct a large-scale synthetic-agent preference dataset with dense, coherent liked/disliked histories, using it as scalable supervision for preference representation learning. We train a multimodal large language model with preference-oriented visual question answering and analyze its hidden states to separate two complementary signals: an intra-user embedding for liked-versus-disliked distinctions and a cross-history consistency embedding for stable user-level tendencies. To bridge MLLM representations and diffusion conditioning, we introduce a distributional alignment objective based on maximum mean discrepancy and inject the aligned preference signal through a lightweight cross-attention branch. We evaluate \textscPrefGen with a tiered protocol: \textscPrefBench as a synthetic in-distribution diagnostic, leakage-free Pick-a-Pic transfer as a controlled real-user benchmark, and user-in-the-loop studies with participant-curated histories as human-facing evidence. Across these settings, \textscPrefGen improves preference alignment over strong personalization baselines while preserving prompt fidelity and competitive image quality.
Abstract: A photorealistic and immersive human avatar experience demands capturing fine, person-specific details such as cloth and hair dynamics, subtle facial expressions, and characteristic motion patterns. Achieving this requires large, high-quality datasets, which often introduce ambiguities and spurious correlations when very similar poses correspond to different appearances. Models that fit these details during training can overfit and produce unstable, abrupt appearance changes for novel poses. We propose a 3D Gaussian Splatting avatar model with a spatial MLP backbone that is conditioned on both pose and an appearance latent. The latent is learned during training by an encoder, yielding a compact representation that improves reconstruction quality and helps disambiguate pose-driven renderings. At driving time, our predictor autoregressively infers the latent, producing temporally smooth appearance evolution and improved stability. Overall, our method delivers a robust and practical path to high-fidelity, stable avatar driving.
Authors:
Pengfei Gu, Huimin Li, Guangyu Meng, Hao Zheng, Haoteng Tang, Bin Fu, Danny Z ChenAbstract: Topological patterns in medical images, such as connected components and loops, across multiple spatial scales carry critical structural information, from microaneurysm-scale lesions to organ-level boundaries. Despite the success of hierarchical Vision Transformers (e.g., Swin Transformer) in medical image analysis, existing methods lack an explicit architectural scheme to represent and incorporate multi-scale topological structures through the Transformer hierarchy. In this paper, we propose \emphTopology-Reinforced Swin Transformer (TRiST), a new Swin Transformer model family for medical image analysis, which reinforces multi-scale hierarchical topology into the hierarchical Transformer architecture. Instead of treating topology as an auxiliary descriptor or post hoc fusion cues, TRiST takes topology as a native part of representation construction, refinement, and token interaction. First, because topology does not compose hierarchically through pooling or interpolation, we develop a Hierarchical Multi-Scale Topology Representation algorithm that computes 2D persistent homology (H_0 and H_1 on image patches) at each native scale of Swin Transformer, constructing four topological representations aligned 1-to-1 with the Swin token grid. Second, since the persistence axis has an intrinsic semantic structure (i.e., short-lifetime features mainly capture noise and fine texture, whereas long-lifetime features capture stable anatomical structures), we design a Band-Aware H_0/H_1 Residual Refinement module that independently adapts six persistence sub-bands using dedicated gated residual MLPs. Third, to make topology take part in token interaction inside the Transformer, we introduce a Persistence-Band Progressive Attention Bias that injects stage-corresponding topology as an additive key-highlighting prior into shifted-window self-attention. Based on these strategies, we instantiate two task-specific models: Topo-SwinV2-B for classification and Topo-Swin-UNet for segmentation. Both models retain the original Swin-based backbone structure while equipping it with native multi-scale topological features. Experiments on five medical image datasets show consistent improvements over strong Swin-based baselines and competitive performance as existing topology-augmented methods.
Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a powerful paradigm for improving language models on reasoning-intensive tasks, but its effectiveness is often limited by exploration. For example, models often fail on hard problems, leaving little useful reward signal. External expert traces offer a natural source of guidance, yet they may also expose reward-relevant content along the critical path to the verifier target, such as final answers, intermediate values, executable implementations, or answer-related entities. This content can create an unintended reward hacking channel, allowing the policy to obtain reward by copying the trace rather than learning the underlying reasoning or agentic behavior. Existing guided-RL methods reduce this risk by using partial trajectories, but they mainly control how much expert information is shown heuristically rather than which parts should be hidden. To this end, we propose Semantic Masked Expert Policy Optimization (SMEPO), a fine-grained semantic masking strategy for expert-guided RLVR. Instead of truncating traces coarsely or revealing them unchanged, SMEPO masks reward-relevant semantic spans along the critical path while preserving the expert’s decomposition, plan, and procedural structure. This turns hard problems from reasoning from scratch into a fill-in-the-blank process: the policy can follow the expert’s problem-solving route, but must still reconstruct the missing values, code, or entities by itself. SMEPO is simple to apply and requires no changes to the reward function or RL objective. Across diverse domains, including math, code, and agentic search, SMEPO improves accuracy by up to 3.2 points over GRPO and reduces training time by up to 4.2×.
Abstract: Post-training has become effective for high-level generation, but its role in low-level vision remains underexplored. Existing image restoration methods often rely on fixed pixel-wise fitting to ground-truth images, which can lead to over-smoothing and weak generalization. We propose IRPO, a GRPO-based post-training framework for deterministic restoration models. IRPO is built around two axes: data formulation and reward modeling. For data formulation, we select the 30% underperforming samples from the pre-training stage, which improves both accuracy and training efficiency. For reward modeling, we combine fidelity-oriented and quality-aware feedback with three components: a General Reward for structural fidelity, an Expert Reward that uses a Vision-Language Model as a coarse visual-quality judge, and a Restoration Reward for task-specific low-level cues. Experiments on six in-domain and five out-of-domain (OOD) benchmarks show that IRPO improves the AdaIR baseline by 0.93 dB on in-domain tasks and 3.43 dB on OOD settings. Our code will be released upon acceptance.
Abstract: Reinforcement learning (RL) has become a representative post-training paradigm for large language models (LLMs), enabling strong reasoning and agentic capabilities. However, its rollout generation remains a dominant training bottleneck because it relies on sequential autoregressive (AR) decoding, where a small number of long-tailed responses often determine completion time. Speculative decoding (SD) can reduce inference latency while preserving model quality by rapidly drafting tokens and accepting them through parallel verification. Applying SD to RL rollouts, however, introduces challenges absent from standard LLM inference: (i) algorithmically, the continuously evolving target model makes static drafters stale; (ii) system-wise, rollout decoding moves across regimes, from large active batches where SD can be compute-bound and ineffective to shrinking-batch tails where SD becomes beneficial. Existing RL-SD approaches address aspects of this problem, but either yield low effective accepted lengths that limit speedup or rely on auxiliary drafters requiring pretraining and online adaptation, increasing system complexity. We present EfficientRollout, an SD framework designed to accelerate RL rollouts while addressing these challenges. EfficientRollout induces a quantized drafter directly from the target model, keeping it coupled to the evolving policy without separate drafter training before or during RL. It then uses a dynamic SD toggle policy that enables SD only in beneficial regimes identified by system-aware roofline modeling. It further adapts drafting budgets using acceptance-behavior signals observed during training, better realizing the potential accepted length. Under realistic RL workload, EfficientRollout reduces rollout and end-to-end latency by up to 19.6% and 12.7%, respectively, over a standard accelerated AR rollout baseline, while preserving final model quality.
Abstract: This work presents \textscChunkFT, a memory-efficient fine-tuning framework that reformulates full-parameter fine-tuning around a dynamically activated working set. \textscChunkFT enables gradient computation for arbitrary sub-tensors without modifying the network architecture, providing an algorithmic foundation for optimizing arbitrary sub-networks while avoiding standard dense gradient computation. We provide a theoretical convergence analysis of \textscChunkFT in the deterministic setting. Empirically, we apply \textscChunkFT to fine-tune Llama 3-8B and Llama 3-70B using a single RTX 4090-24GB GPU and 2× H800-80GB GPUs, respectively. Full-parameter fine-tuning of a 7B model with a 1K input length requires only 13.72GB of GPU memory. The results demonstrate the effectiveness of \textscChunkFT in memory usage, running time, and optimization quality. Moreover, downstream evaluations on language understanding, mathematical reasoning, and MT-Bench show that \textscChunkFT consistently outperforms existing memory-efficient baselines. Notably, \textscChunkFT achieves performance comparable to, and in some cases exceeding, full-parameter fine-tuning. Our repository is on https://anonymous.4open.science/r/chunk-B48E.
Abstract: Pre-training is crucial for large language models (LLMs), as it is when most representations and capabilities are acquired. However, natural language pre-training has problems: high-quality text is finite, contains human biases, and entangles knowledge with reasoning. This raises a fundamental question: is natural language the only path to intelligence? We propose using neural cellular automata (NCA) to generate synthetic, non-linguistic data for LLMs--training on synthetic-then-natural language. NCA data exhibits rich spatiotemporal structure and statistics resembling natural language while being controllable and cheap to generate at scale. We find that pre-pre-training on only 164M NCA tokens improves downstream language modeling by up to 6% and accelerates convergence by up to 1.6x in 1-3B models trained with Chinchilla optimal data budgets. Surprisingly, this even outperforms pre-pre-training on 1.6B tokens of high-quality natural language data with more compute. Investigating what drives transfer, we find that attention layers are the most transferable, and that optimal NCA complexity varies by domain: code benefits from simpler dynamics, while math and web text favor more complex ones. These results enable systematic tuning of the synthetic distribution to target domains. More broadly, our work opens a path toward more efficient models with fully synthetic pre-training.
Abstract: Transformers replace recurrence with a memory that grows with sequence length and self-attention that enables ad-hoc look ups over past tokens. Consequently, they lack an incentive to compress history into compact latent states with consistent transition rules. This often leads to learning solutions that generalize poorly. We introduce Next-Latent Prediction (NextLat), which extends standard next-token training with self-supervised predictions in the latent space. Specifically, NextLat trains a transformer to learn latent representations that are predictive of its next latent state given the next output token. Theoretically, we show that these latents provably converge to belief states, compressed information of the history necessary to predict the future. This simple auxiliary objective injects a recurrent inductive bias into transformers, while leaving their architecture, parallel training, and inference unchanged. NextLat effectively encourages the transformer to form compact internal world models with its own belief states and transition dynamics—a crucial property absent in standard next-token prediction transformers. Empirically, across benchmarks in world modeling, reasoning, planning, and language modeling, NextLat demonstrates significant gains over standard next-token training in downstream accuracy, representation compression, and lookahead planning. Furthermore, NextLat enables variable-length self-speculative decoding, accelerating inference by up to 3.3× in the language domain. NextLat stands as a simple and efficient paradigm for shaping transformer representations toward stronger generalization.
Authors:
Jun CEN, Siteng Huang, Yuqian Yuan, Kehan Li, Hangjie Yuan, Chaohui Yu, Bohan Hou, Yuming Jiang, Jiayan Guo, Xin Li, Hao Luo, Fan Wang, Deli Zhao, Hao ChenAbstract: We present WorldVLA, an autoregressive action world model that unifies action and image understanding and generation. Our WorldVLA integrates Vision-Language-Action (VLA) model and world model in one single framework. The world model predicts future images by leveraging both action and image understanding, with the purpose of learning the underlying physics of the environment to improve action generation. Meanwhile, the VLA model generates the subsequent actions based on image observations, aiding in visual understanding and in turn helps visual generation of the world model. We demonstrate that WorldVLA outperforms standalone VLA and world models, highlighting the mutual enhancement between the world model and the VLA model. We evaluate WorldVLA in both simulation and real-world robot tasks. WorldVLA achieves 97.4% success rate on the LIBERO simulation benchmark without pretraining, while in real-world LeRobot experiments, its integrated world model boosts the overall success rate by 50%.
Abstract: Knowledge distillation pipelines typically train a student to match a teacher on a fixed corpus of teacher-labeled examples. This leaves a major source of supervision unused: the teacher is a fully queryable model whose output distribution extends far beyond any fixed dataset. We argue that effective distillation should exploit this structure by steering generation toward regions where the student most underestimates the teacher. We propose Maximum Divergence Knowledge Distillation (MDKD), a rejection-sampling method for autoregressive language models. At each decoding step, MDKD draws candidate tokens from the teacher and preferentially accepts those the student underweights, constructing trajectories that remain teacher-supported while concentrating training on regions where the student assigns insufficient probability. We show that stochastic MDKD's acceptance rule has an exact distributional characterization: it samples precisely from the teacher's uncovered mass [T - S]_+, the probability the teacher assigns to tokens the student does not yet cover, with stochastic acceptance probability equal to the teacher--student total variation distance. Standard teacher sampling, by contrast, spends most of its updates on the overlap \min(T, S) already absorbed by the student. This makes MDKD the structural complement of speculative decoding, which exploits the same overlap to accelerate inference. On GSM8K, distilling \textttQwen2.5-14B-Instruct into a base \textttQwen2.5-1.5B student with only 1,000 KD samples (5 epochs) lifts accuracy from 8.26% to 64.90%, outperforming SOTA methods by 10.3 percentage points and recovering 65% of the teacher--student gap. A divergence-gain analysis confirms the mechanism: MDKD sequences raise student cross-entropy by 43.6% while preserving 98% of teacher-sampling accuracy. Across arithmetic reasoning, dialogue summarization, and code generation, MDKD matches or outperforms strong distillation baselines.
Authors:
Wonjun Kang, Kevin Galim, Seunghyuk Oh, Minjun Kang, Sanghyun Park, Donghoon Kim, Minjae Lee, Minseo Kim, Rishabh Tiwari, Yuchen Zeng, HYUNG IL KOO, Kangwook LeeAbstract: On-policy distillation (OPD) is an increasingly important way to post-train large language models (LLMs), but, like reinforcement learning (RL), it relies on rollouts from the model being optimized. For reasoning workloads, rollout generation can dominate training time. Asynchronous RL alleviates this bottleneck by decoupling rollout generation from learner updates, but doing so introduces stale-policy data; prior work studies how to stabilize learning from such data. However, it remains underexplored whether these asynchronous RL ideas and stale-data solutions transfer to OPD, and what OPD-specific constraints arise. To address this gap, we provide the first systematic study of staleness in asynchronous OPD. We first show that KL direction changes the stale-data problem: teacher-weighted forward KL is robust to stale rollouts, whereas student-weighted reverse KL is vulnerable. Second, for this vulnerable reverse-KL case, we study whether methods designed to stabilize asynchronous RL can mitigate OPD staleness. We find that they do not improve over a simpler OPD-specific surrogate: recomputing the reverse-KL signal under the current student at learner time without clipping. Third, we identify an OPD-specific teacher-cache constraint: under asynchronous execution, teacher scores are available at learner time only on cached actions. The resulting bias-variance tradeoff for sparse and sampled reverse-KL OPD implementations motivates multi-sample Monte Carlo (MC), which preserves MC correctability while reducing one-sample variance. Finally, we present and open-source AsyncOPD, a fully asynchronous OPD training pipeline built from these estimator choices. Experiments show that AsyncOPD improves training throughput by 1.6× to 3.8× over strict synchronous training while reaching comparable accuracy.
Abstract: We introduce Post-Optimization Model Edit (POME), a new algorithm that enhances the performance of fine-tuned large language models using only their pretrained and fine-tuned checkpoints, without requiring extra data or further optimization. The core idea is to apply a muon-style projection to \Delta W, the difference between the fine-tuned and pretrained weights. This projection uses truncated singular value decomposition (SVD) to equalize the influence of dominant update directions and prune small singular values, which often represent noise. As a simple post-processing step, POME is completely decoupled from the training pipeline. It requires zero modifications and imposes no overhead, making it universally compatible with any optimizer or distributed framework. POME delivers consistent gains, boosting average performance by +2.5% on GSM8K and +1.0% on code generation. Its broad applicability—from 7B foundation models to 72B RLHF-instructed models—establishes it as a practical, zero-cost enhancement for any fine-tuning pipeline.
Abstract: Fine-tuning an LLM to maximize performance on a downstream task requires finding the optimal data and model configuration. This is a black-box optimization problem that Bayesian optimization (BO) addresses by sequentially evaluating training configurations and using observed performance feedback to adaptively guide the search towards better configurations. However, directly applying BO to the joint data-model space suffers from two fundamental challenges: the prohibitive cost of full LLM training at each iteration, and the large number of BO iterations required for effective exploration due to the high problem dimensionality. This paper introduces JoBS, a BO-based approach that co-optimizes data and model configurations without requiring a full training run at every iteration. JoBS allocates an initial portion of the optimization budget to learn a scaling-law-inspired performance predictor that estimates fully trained LLM performance from only a small number of training steps. The remaining budget is then used to run BO with this predictor, eliminating the need for full training runs and enabling JoBS to explore significantly more data and model configurations within the same budget. We analyze JoBS' average regret and derive the optimal budget allocation between predictor learning and BO iterations. Empirically, JoBS outperforms independent data and model optimization methods and existing multi-fidelity BO baselines across a diverse set of downstream tasks. Our code is available at https://github.com/a35453779/JoBS.
Abstract: Multi-Agent Systems (MAS) are a promising paradigm for improving the reasoning capabilities of large language models, yet they remain prone to hallucinations and erroneous outputs. Uncertainty Quantification (UQ) is crucial for assessing system reliability, but existing methods are largely designed for single-agent settings; how to exploit MAS structure for UQ while keeping inference cost-neutral remains underexplored. We propose UMAS, a parameter-free, cost-neutral, system-level UQ framework that estimates the credibility of inter-agent influence from each generation’s intrinsic confidence. UMAS uses these credibility signals to modulate how uncertainty is updated at each node, and then aggregates node-level uncertainty into a system-level estimate. Extensive evaluations on Debate and DyLAN across diverse benchmarks show that UMAS consistently outperforms state-of-the-art baselines by an average of 10.49 AUROC points. Beyond uncertainty estimation, UMAS enables hallucination detection and uncertainty-aware answer selection, improving MAS accuracy by up to 12.6 points on specific tasks while enhancing reliability.
Abstract: Reinforcement learning is critical to improving large reasoning models, but its success relies heavily on verifiable rewards (RLVR), making it hard to use in open-ended domains where correctness is ambiguous and cannot be verified. Moreover, reasoning trajectories remain largely unconstrained, and optimizing solely toward the final answer can favor early exploitation over generalization. In this work, we ask whether general reasoning ability can be improved by teaching models how to think (the structure of reasoning) rather than what to produce (the outcome of reasoning), and we extend traditional RLVR to open-ended settings. We introduce Structure-Aware Reinforcement Learning (SARL), a label-free framework that constructs per-response reasoning maps from intermediate thinking steps and rewards their reasoning topology. SARL shifts supervision from destination to path, encouraging reasoning trajectories that are both locally coherent and globally efficient. On verifiable math tasks, SARL outperforms prior label-free RL baselines and even exceeds RL methods with ground truth supervision, with average gains of +9.1% under PPO and +11.6% under GRPO across four math benchmarks, with particularly large improvements on AIME25 (+35.5% with PPO and +44.7% with GRPO). On non-verifiable open-ended tasks, SARL achieves average gains of +34.6% under PPO and +30.4% under GRPO on WildBench across five task categories, outperforming prior label-free RL methods and DPO, which relies on additional preference labels. Beyond strong performance, SARL exhibits substantially lower KL divergence and higher policy entropy, indicating more stable and exploratory training dynamics. Code and data are available at https://anonymous.4open.science/r/SARL-EB77
Authors:
Yibang Li, Bihari L Pandey, Ravi Sah, Andi Han, Cyrus Mostajeran, Pratik Kumar Jawanpuria, Bamdev MishraAbstract: Muon and related norm-constrained matrix optimizers have become central to large-scale learning problems. They are formulated as a linear maximization oracle (LMO) over an ambient matrix-norm ball in unconstrained Euclidean space. However, these do not generalize cleanly to manifold-valued parameters such as low-rank factorizations, orthogonality constraints, or symmetric positive definite (SPD) matrices. Naively restricting the Muon LMO to the tangent space (i) breaks quotient symmetries and (ii) couples the tangent-space constraint with an ambient norm bound, thereby obstructing closed-form solutions on various manifolds of interest. We resolve both issues with a single observation: every Riemannian metric canonically lifts a unitarily invariant Euclidean norm to an intrinsic norm on each tangent space, and the resulting intrinsic norm constrained LMO is symmetry preserving. Building on this, we introduce intrinsic Muon (iMuon), a unified framework that yields closed-form updates on the fixed-rank, SPD, Stiefel, and Grassmann manifolds for any unitarily invariant norm, including the spectral, Frobenius, and nuclear norms. We establish convergence guarantees for both deterministic and stochastic iMuon with rate constants that depend only on the manifold dimension. Notably, on the fixed-rank manifold this constant depends only on the rank, making the rate independent of factor conditioning and removing the runtime factor-rescaling required by prior work. Experiments on LoRA finetuning of LLMs, image classification, and subspace learning illustrate the efficacy of the proposed approach.
Abstract: Demoir\'eing aims to remove moir\'e artifacts that often occur in images. While recent deep learning-based methods have achieved promising results, they typically require substantial computational resources, limiting their deployment on edge devices. Model quantization offers a compelling solution. However, directly applying existing quantization methods to demoir\'eing models introduces severe performance degradation. The main reasons are distribution outliers and weakened representations in smooth regions. To address these issues, we propose QuantDemoire, a post-training quantization framework tailored to demoir\'eing. It contains two key components. First, we introduce an outlier-aware quantizer to reduce errors from outliers. It uses sampling-based range estimation to reduce activation outliers and keeps a few extreme weights in FP16 with negligible cost. Second, we design a frequency-aware calibration strategy. It emphasizes low- and mid-frequency components during fine-tuning, which mitigates banding artifacts caused by low-bit quantization. Experiments validate that our QuantDemoire achieves large reductions in parameters and computation while maintaining quality. Meanwhile, it outperforms existing quantization methods by over 4 dB on W4A4. We implement a CUDA kernel for the mixed-precision branch and achieve a 2.3× end-to-end speedup on an RTX 4090 GPU. Code will be made public.
Authors:
Andrew Bond, Ege E Özlü, Tuna Çimen, Ilkin U Melanlioglu, Tolga Birdal, Erkut Erdem, Aykut ErdemAbstract: Methods that operate on Vision Transformer features almost universally rely on Euclidean distance to judge feature similarity. Yet Euclidean distance weights every direction in feature space equally, ignoring that downstream tasks concentrate their sensitivity on a low-dimensional submanifold. The metric that respects this structure is the pullback Riemannian metric g(F) = J(F)^\top J(F) induced by the task decoder's Jacobian, but materializing it is prohibitively expensive; a single ViT-B/14 probe point requires storing 4× 10^10 scalars, and per-token scoring provably fails when task-sensitive directions are spread across tokens. We show that whether a low-rank approximation of g(F) is \emphlearnable at all depends on a precise structural property of the model-decoder pair, which we characterize via a single scalar diagnostic \kappa_cap(r), the captured-energy fraction of the top-r singular directions of J, computable matrix-free in \mathcalO(m+rq) Jacobian-vector products. When the diagnostic licenses it, we train the \emphSpectral Pullback Network (SPN) on directions found by randomized power iteration on J^\top J, and distill the resulting per-token signal into a 310K-parameter feature-only \emphconformal head, never materializing J. When the original spectrum is too rich for any practical rank-r budget, a VAE reparameterization of the decoder recovers a latent space in which the same machinery applies, converting intractable dense pipelines into tractable ones. Across DPT, DINOv2, CLIP, and VGGT backbones the diagnostic correctly predicts which architecture succeeds: the conformal head reaches Spearman \rho=0.998 against the Jacobian-derived target on DINOv2 CLS, and geometric token pruning yields a 25% relative improvement over ToMe at aggressive prune ratios on DPT depth. These results suggest that learning \emphwhich directions matter is often more valuable than learning better features.
Abstract: Fine-tuning Vision Language Models (VLMs) with parameter-efficient methods like Low-Rank Adaptation (LoRA) is crucial for task adaptation. However, imbalanced training dynamics across modalities often lead to suboptimal accuracy due to negative interference, a challenge typically addressed with inefficient heuristic methods such as manually tuning separate learning rates. To overcome this, we introduce earch), an approach to discover optimal rank pairs that balance training dynamics while maximizing performance. Our key innovation, a proposed framework of dual scaling laws, enables this search:~one law models module-specific convergence time to prune the search space to candidates with aligned dynamics, while the other predicts final task performance to select the optimal pair from the pruned set. By re-purposing the LoRA rank as a controller for modality-specific convergence speed, MARS outperforms baseline methods and provides a robust, automated strategy for optimizing VLM fine-tuning.
Abstract: This work identifies a critical failure mode in frontier large language models (LLMs), which we term Internal Safety Collapse (ISC): under certain task conditions, models enter a state in which they continuously generate large volumes of harmful content while executing otherwise benign tasks. To systematically study ISC, we introduce \TVD (Task, Validator, Data), a framework that instantiates controlled workflows around domain tools, where valid completion requires filling harmful content as structured data. We collect 53 representative workflows across 8 disciplines, from toxicity evaluation to molecular docking and pathogen genome analysis, showing that ISC reproduces across all eight disciplines. In the worst case over three interaction settings, the four frontier LLMs average a 95.3% safety-failure rate. Frontier LLMs with stronger task-completion capability show higher unsafe-completion rates: long-horizon execution skill becomes a liability when workflow completion requires harmful content. We even observe extremely severe harmful content closely resembling outputs from early-generation, unaligned LLMs in 2023. Despite substantial safety alignment efforts, frontier LLMs continue to retain inherently unsafe internal capabilities: alignment reshapes observable outputs but does not eliminate the underlying risk profile. These findings underscore the need for caution when deploying LLMs in high-stakes settings, including scientific pipelines and autonomous agents. Complementary source code is provided at \urlhttps://anonymous.4open.science/r/NIPS-ISC-Code-Share-BE11.
Abstract: Multimodal driving planning faces a long-standing tension between two paradigms: scoring-based methods benefit from dense reward supervision but are confined to a fixed action vocabulary, while anchor-based methods generate proposals dynamically yet suffer from sparse supervision constrained to a single ground-truth trajectory. In this work, we propose FlowR2A, which resolves this tension by reframing simulation-based rewards from discriminative targets into generative conditions. By learning the reward-conditioned action distribution from dense trajectory-reward pairs with a flow-matching decoder, FlowR2A unifies the dense supervision of scoring-based methods with the proposal generation of anchor-based methods in a single generative model, forcing the model to internalize the correlation between an action and its outcomes in safety, progress, comfort, and rule compliance. To balance hard safety constraints against soft progress objectives, we introduce fine-grained per-timestep reward conditioning and reward noise augmentation. The generative formulation naturally supports controllable test-time sampling via reward guidance and anchored sampling, producing high-quality proposals. FlowR2A achieves state-of-the-art results on the NAVSIM v1 and v2 benchmarks, with multimodal proposals of substantially higher quality than prior methods.
Authors:
Kangning Zhang, Shuai Shao, Wenxiang Jiao, Qingyao Li, Jianghao Lin, Lingyue Fu, Shijian Wang, Yuan Lu, Weiwen Liu, Weinan Zhang, Yong YuAbstract: Reusable skills have become a core substrate for improving agent capabilities, yet most existing skill packages encode reusable behavior primarily as textual prompts, executable code, or learned routines. For visual agents, however, procedural knowledge is inherently multimodal: reuse depends not only on what operation to perform, but also on recognizing the relevant state, interpreting visual evidence of progress or failure, and deciding what to do next. We formalize this requirement as \emphmultimodal procedural knowledge and address three practical challenges: (I) what a multimodal skill package should contain; (II) where such packages can be derived from public interaction experience; and (III) how agents can consult multimodal evidence at inference time without excessive image context or over-anchoring to reference screenshots. We introduce \emphMMSkills, a framework for representing, generating, and using reusable multimodal procedures for runtime visual decision making. Each MMSkill is a compact, state-conditioned package that couples a textual procedure with runtime state cards and multi-view keyframes. To construct these packages, we develop an agentic trajectory-to-skill Generator that transforms public non-evaluation trajectories into reusable multimodal skills through workflow grouping, procedure induction, visual grounding, and meta-skill-guided auditing. To use them, we introduce a branch-loaded multimodal skill agent: selected state cards and keyframes are inspected in a temporary branch, aligned with the live environment, and distilled into structured guidance for the main agent. Experiments across GUI and game-based visual-agent benchmarks show that MMSkills consistently improve both frontier and smaller multimodal agents, suggesting that external multimodal procedural knowledge complements model-internal priors. The code is accessible in \hrefhttps://anonymous.4open.science/r/MMSkillshttps://anonymous.4open.science/r/MMSkills.
Abstract: Consistency-based generative models like Shortcut and MeanFlow achieve impressive results via a target-aware design for solving the Probability Flow ODE (PF-ODE). Typically, such methods introduce a target time r alongside the current time t to modulate outputs between a local multi-step derivative (r = t) and a global few-step integral (r = 0). However, the conventional "one input, one output" paradigm enforces a partition of the training budget, often allocating a significant portion (e.g., 75% in MeanFlow) solely to the multi-step objective for stability. This separation forces a trade-off: allocating sufficient samples to the multi-step objective leaves the few-step generation undertrained, which harms convergence and limits scalability.To this end, we propose Duality Models (DuMo) via a "one input, dual output" paradigm. Using a shared backbone with dual heads, DuMo simultaneously predicts velocity \mathbfv_t and flow-map \mathbfu_t from a single input \mathbfx_t. This applies geometric constraints from the multi-step objective to every sample, bounding the few-step estimation without separating training objectives, thereby significantly improving stability and efficiency. On ImageNet 256 × 256, a 679M Diffusion Transformer with SD-VAE achieves a state-of-the-art (SOTA) FID of 1.79 in just 2 steps.Code will be publicly available.
Abstract: For groups of autonomous agents to achieve a particular goal, they must engage in coordination and long-horizon reasoning. Rather than relying on complex reward functions and explicit cooperation mechanisms, we ask what minimal ingredients are required for effective coordination and exploration to emerge in multi-agent settings. We investigate this question through self-supervised goal-reaching, where agents aim to maximize the likelihood of visiting a goal state rather than maximizing a reward. Despite a sparse feedback signal, we present empirical results that show self-supervised goal-reaching techniques enable agents to learn from such feedback. On MARL benchmarks, self-supervised goal-reaching outperforms alternative approaches that have access to the same sparse reward signal. Furthermore, we empirically demonstrate that multi-agent self-supervised goal-reaching approaches can be more robust than single-agent strategies. While there is no exploration mechanism, this approach explores nontrivial intermediate coordination strategies in sparse settings where alternative approaches fail to achieve a single success.
Abstract: Existing latent world models for autonomous driving have opened a promising path toward future-aware driving intelligence. However, they typically treat future latent states as prediction targets or auxiliary signals, rather than directly conditioning trajectory planning. This can entangle current and future features in latent space. In this work, we propose DriveFuture, a future-aware latent world modeling framework for autonomous driving that explicitly learns planning-oriented foresight by conditioning the current latent state modeling process on future world states. Specifically, during training, the model first predicts future latent world states from the current latent state and ego action, and then refines the prediction against the ground-truth future latent state via cross-attention. The resulting future-aware latent serves as an explicit condition for a diffusion-based trajectory planner. During inference, DriveFuture conditions on the predicted future latent state instead of the ground-truth future state. DriveFuture achieves SOTA performance on the public NAVSIM benchmarks, reaching 55.5 EPDMS on NAVSIM-v2 navhard, 89.9 EPDMS on NAVSIM-v2 navtest, and 90.7 PDMS on NAVSIM-v1 navtest, respectively. These results suggest that the key to latent world modeling lies not merely in simulating future states, but more importantly in conditioning current decision-making on future states. Notably, as of April 2026, DriveFuture ranks 1st on the NAVSIM-v2 navhard leaderboard and achieves SOTA performance on NAVSIM-v1 navtest.
Abstract: Human-centric video customization, particularly at the garment level, has shown significant commercial value. However, existing approaches cannot support low-latency and interactive garment control, which is crucial for applications such as e-commerce and content creation. This paper studies how to achieve interactive multi-garment video customization while preserving motion coherence using only single-garment video data. We present FashionChameleon, a real-time and interactive framework for human-garment customization in autoregressive video generation, where users can interactively switch garment during generation. FashionChameleon consists of three key techniques: (i) Instead of training on multi-garment video data, we train a Teacher Model with In-Context Learning on a single reference–garment pair. By retaining the image-to-video training paradigm while enforcing a mismatch between the reference and garment image, the model is encouraged to implicitly preserve coherence during single-garment switching. (ii) To achieve consistency and efficiency during generation, we introduce Streaming Distillation with In-Context Learning, which fine-tunes the model with in-context teacher forcing and improves extrapolation consistency via gradient-reweighted distribution matching distillation. (iii) To extend the model for interactive multi-garment video customization, we propose Training-Free KV Cache Rescheduling, which includes garment KV refresh, historical KV withdraw, and reference KV disentangle to achieve garment switching while preserving motion coherence. Our FashionChameleon uniquely supports interactive customization and long-video extrapolation, while achieving real-time generation at 23.8 FPS on a single GPU, 30-180× faster than existing baselines.
Abstract: Geometry estimation from perspective images has greatly advanced, maturing to the point where off-the-shelf foundation models are able to reconstruct 3D scene structure not only from multi-view imagery, but even from a single view. A natural extension is 3D reconstruction from panoramas, with the exciting prospect of recovering a full 360^\circ scene from a single panoramic image. In this work, we introduce PaGeR (Panoramic Geometry Reconstruction), a framework to lift powerful 3D foundation models, designed for perspective imagery, to the panorama domain. Our strategy is to start from a pretrained transformer for 3D reconstruction and turn it into a unified high-performance model that predicts scale-invariant depth, metric depth, sky masks and surface normals from both perspective and omnidirectional images, a single forward pass. By keeping architectural changes to a minimum and mixing perspective and panoramic images during training, Pager retains the rich 3D prior of the underlying foundation model while learning to also estimate geometrically consistent 360^\circ scenes from single panoramas. We extensively test our method in both indoor and outddor environments and find that it delivers state-of-the-art performance and excellent zero-shot performance across a wide range of scenes.
Abstract: Multimodal large language models (MLLMs) perform well on many vision-language tasks but often struggle with vision-centric problems that require fine-grained visual perception. Recent evidence suggests that this limitation arises not from weak visual representations, but from under-utilization of visual information during instruction tuning, where many tasks can be partially solved using language priors alone. We propose a simple and lightweight approach that augments visual instruction tuning with a small number of visually grounded self-supervised tasks expressed as natural language instructions. By reformulating classical self-supervised pretext tasks, such as rotation prediction, color matching, and cross-view correspondence, as image–instruction–response triplets, we introduce supervision that cannot be solved without relying on visual evidence. Our approach requires no human annotations, no architectural modifications, and no additional training stages. Across multiple models, training regimes, and benchmarks, injecting only a small fraction (3–10%) of such visually grounded instructions consistently improves performance on vision-centric evaluations. Our findings highlight instruction tuning with visually grounded SSL tasks as a powerful lever for improving visual perception in MLLMs through simple adjustments to the training data distribution.
Abstract: In the physical world we inhabit, space and time are fundamentally continuous. However, existing machine learning paradigms for world modeling are largely confined to discrete-time prediction and inference, thereby exhibiting significant inefficiency in capturing the fundamental dynamics of the physical world. To bridge this gap, we introduce Physical-Time Flow ( ), a novel approach that learns a continuous latent velocity field operating in physical time. Crucially, the underlying dynamics of sequential data are parameterized by an ordinary differential equation (ODE) embedded in a well-structured dynamical representation space. Under this paradigm, the prediction of the future can be recast as temporal integration via an ODE solver in the compressed latent space. Building upon PT-Flow, we construct , a physics-grounded, continuous-time latent world model that is both efficient and versatile. By extracting time-variant features and enforcing ODE properties on both the dynamical representation space and the latent velocity field, ODEWorld effectively addresses the long-standing representation collapse issue in latent world model literature. This also enables high-quality image reconstruction with ODEWorld even after long-horizon prediction. Moreover, its continuous nature allows for arbitrary temporal resolution and even backward prediction, which is impossible for most discrete-time models. Lastly, benefiting from the learned dynamics-centric latent space, ODEWorld can provide rich planning-oriented information to facilitate downstream policy learning. Comprehensive experiments across multiple simulation and real-world datasets demonstrate that ODEWorld successfully reconciles planning-conducive dynamics abstraction with visual realism, excelling in both video generation and robotic control. More qualitative results can be found at
Abstract: Cinematic video depicts multiple subjects acting or interacting at specific moments, captured with deliberate camera movement, and stitched together by shot transitions. Together, these elements demand a level of fine-grained control beyond current text-to-video models. Existing work addresses each axis in isolation: multi-subject personalization, temporal control, multi-shot synthesis, or camera control; no prior framework jointly integrates all four. We present CineOrchestra, a unified video diffusion model that controls subjects, events, cameras, and shot transitions simultaneously. Our key insight is that these heterogeneous cinematic elements share a fundamental structure: each is an entity acting over a specific temporal interval, which can therefore all be expressed through one shared structure of entity-centric conditioning primitives, augmented with reference images for visual entities. This formulation reduces the architectural challenge to a single positional encoding problem, which we solve with two parameter-free coordinated rotary embeddings: (i) an interval-sampled temporal RoPE that yields consistent attention behavior across events of dramatically varying duration, and (ii) a 2D entity-temporal cross-attention RoPE that disambiguates per-entity conditions and routes each to its corresponding spatiotemporal region. On two new benchmarks, CineOrchestra outperforms six per-axis specialists on dense caption following and shot-transition timing, with consistent gains in a pairwise user study and component ablations.
Abstract: Direct preference optimization (DPO) is a simple and effective alignment strategy for large language models (LLMs) based on pairwise preferences between two candidates. In recommender systems, however, user feedback is rarely pairwise. For a given context, e.g., a user, a session, or a conversation, we typically observe set-wise preferences with multiple positive items, where every positive item should outrank every unobserved or explicitly negative item, with no prescribed order among the positives or the negatives themselves. A natural generalization is to use the Plackett–Luce (PL) reward model, which extends the Bradley–Terry (BT) reward model underlying vanilla DPO from pairwise preferences to full rankings of candidates. However, adapting the PL model to set-wise preferences requires marginalizing over all consistent permutations of the positives and negatives respectively, which is intractable. To address this fundamental challenge, we propose Mult-DPO, a novel DPO objective with a tractable multinomial reward model over set-wise preferences. We show that, like the BT and PL reward models, the multinomial reward admits a closed-form DPO-style objective, enabling direct alignment of LLMs without reinforcement learning (RL). In addition, we prove that the multinomial DPO loss is a tractable upper bound on the exact marginalized PL DPO loss when optimizing against the set-wise preference data. We further characterize the tightness of this bound in terms of relative total weight (i.e., exponentiated rewards) of positives versus negatives, which provides insights into tightening the bounds with more or harder negatives. Finally, we extend the framework to a group-wise setting that accommodates multiple preference levels. Code and datasets are available at https://anonymous.4open.science/r/mult-dpo-41C7.
Authors:
Tong Zheng, Haolin Liu, Chengsong Huang, Huiwen Bao, Sheng Zhang, Rui Liu, Runpeng(Leo) Dai, Ruibo Chen, chenxi liu, Tianyi Xiong, Xidong Wu, Hongming Zhang, Heng HuangAbstract: Test-time scaling (TTS) has become an effective approach for improving large language model performance by allocating additional computation during inference. However, existing TTS strategies are largely hand-crafted: researchers manually design reasoning patterns and tune heuristics by intuition, leaving much of the computation-allocation space unexplored. We propose an environment-driven framework, , that changes what researchers design: from individual TTS heuristics to environments where TTS strategies can be discovered automatically. The key to AutoTTS lies in environment construction: the discovery environment must make the control space tractable and provide cheap, frequent feedback for TTS search. As a concrete instantiation, we formulate width--depth TTS as controller synthesis over pre-collected reasoning trajectories and probe signals, where controllers decide when to branch, continue, probe, prune, or stop and can be evaluated cheaply without repeated LLM calls. We further introduce beta parameterization to reduce search-set overfitting and fine-grained execution trace feedback to improve discovery efficiency by helping the agent diagnose why a TTS program fails. Experiments on mathematical reasoning benchmarks show that the discovered strategies improve the overall accuracy--cost tradeoff over strong manually designed baselines. The discovered strategies generalize to held-out benchmarks and model scales, while the entire discovery costs only
Abstract: Flexible image tokenizers aim to represent an image using an ordered 1D variable-length token sequence. This flexible tokenization is typically achieved through nested dropout, where a portion of trailing tokens is randomly truncated during training, and the image is reconstructed using the remaining preceding sequence. However, this tail-truncation strategy inherently concentrates the image information in the early tokens, limiting the effectiveness of downstream AutoRegressive (AR) image generation as the token length increases. To overcome these limitations, we propose ReTok, a flexible tokenizer with \underlineRedundant \underlineToken Padding and Hierarchical Semantic Regularization, designed to fully exploit all tokens for enhanced latent modeling. Specifically, we introduce Redundant Token Padding to activate tail tokens more frequently, thereby alleviating information over-concentration in early tokens. In addition, we apply Hierarchical Semantic Regularization to align the decoding features of earlier tokens with those from a pre-trained vision foundation model, while progressively reducing the regularization strength toward the tail to allow finer low-level detail reconstruction. Extensive experiments demonstrate the effectiveness of ReTok: on ImageNet 256×256, our method achieves superior generation performance compared with both flexible and fixed-length tokenizers.
Abstract: Time-series anomaly detection (TSAD) requires identifying both immediate Point Anomalies and long-range Context Anomalies. However, existing zero-shot foundation models face a fundamental trade-off: 1D temporal models provide fine-grained pointwise localization but lack a global contextual perspective, while 2D vision-based models capture global patterns but suffer from information bottlenecks due to a lack of temporal alignment and coarse-grained pointwise detection. To resolve this dilemma, we propose VETime, the first TSAD framework that unifies temporal and visual modalities through fine-grained visual-temporal alignment and dynamic fusion. VETime introduces a Reversible Image Conversion and a Patch-Level Temporal Alignment module to establish a shared visual-temporal timeline, preserving discriminative details while maintaining temporal sensitivity. Furthermore, we design an Anomaly Window Contrastive Learning mechanism and a Task-Adaptive Multi-Modal Fusion to adaptively integrate the complementary perceptual strengths of both modalities. Extensive experiments demonstrate that VETime significantly outperforms state-of-the-art models in zero-shot scenarios, achieving superior localization precision with lower computational overhead than current vision-based approaches.
Abstract: Video diffusion models perform well in short-video synthesis, but their training-free extension to long videos often suffers from content drift, temporal inconsistency, and over-smoothed dynamics. Existing methods improve temporal consistency by combining a global branch with a local branch, but they often further decompose appearance consistency and temporal dynamics within each branch using predefined criteria. This assignment is unreliable when appearance and action progression are tightly coupled, such as in camera motion and sequential motion. We analyze the video temporal extension issue from a singular-spectrum perspective and show that enlarged self-attention windows induce spectral concentration: spectral energy becomes dominated by a few low-rank singular directions, preserving coarse structure but suppressing high-rank spatial details and motion-rich temporal variations. To mitigate this problem, we propose FreeSpec, a training-free spectral reconstruction framework for long-video generation. FreeSpec decomposes global and local features with singular value decomposition, and uses the global branch as low-rank spectral guidance and the local branch as a high-rank reconstruction basis. This spectrum-level fusion avoids the rigid feature partitioning of previous decomposition rules, preserving long-range consistency while better retaining spatial details and temporal dynamics. Experiments on Wan2.1 and LTX-Video demonstrate that FreeSpec improves long-video generation, especially for temporal dynamics, while maintaining strong visual quality and temporal consistency.
Authors: Victor De Lima, Jiqun Liu, Grace H Yang
Abstract: AI systems increasingly shape how individuals access and evaluate information, intensifying concerns about misinformation. Computational simulation offers a controlled complement to direct empirical study, with fine-tuned LLMs enabling powerful agent-based simulations. However, fine-tuning often induces broad behavioral shifts including bias drift, degenerate heuristics, and loss of task competence, which undermine experimental validity. The central challenge is inducing localized decision shifts without global behavioral collapse. We study this through a stability analysis of supervised fine-tuning, attenuating loss on selected training slices to shift false-claim acceptance while preserving overall task competence. We introduce FrameRef, a large-scale dataset of semantically equivalent claims with controlled surface framings across five dimensions (Authoritative, Consensus, Emotional, Sensationalist, Prestige), with human validation confirming systematic framing effects on claim acceptance. We identify stability regimes in which targeted error shifts can be induced without degrading aggregate accuracy or calibration. A sequential exposure task over FrameRef shows that small framing-conditioned shifts compound into substantially different cumulative outcomes under feedback, but largely disappear when feedback is removed, confirming that interventions modify specific decision tendencies rather than global performance. We release code, LoRA adapters, and data at https://github.com/anonymousauthor2352/frameref.
Abstract: Context or prompt-level reweighting has emerged as a central algorithmic lever in Reinforcement Learning with Verified Rewards (RLVR) for improving the reasoning capability of large language models, yet the principle determining what constitutes an optimal weighting remains poorly understood. We address this gap by formulating prompt reweighting as a functional derivative of a utility functional defined in the pass-rate function space, yielding a unified optimality framework that accommodates existing schemes, including REINFORCE and GRPO. Building on this optimality framework, we propose a distribution-aware prompt reweighting approach, called CurveRL, based on a quantile coordinate transform, in which the weight assigned to each prompt depends not on the absolute value of pass rates but on its rank and density to reflect the distributional structure of the pass rates in the learning dynamics. Extensive experiments across multiple benchmarks demonstrate that our proposed CurveRL consistently outperforms GRPO and other RLVR baselines. Our study identifies context-distribution control as a principled axis for analyzing and designing prompt-reweighted RLVR algorithms.
Abstract: Metal-organic frameworks (MOFs) are promising materials for applications such as direct air capture, where performance depends on how adsorbates bind within the framework. Accurately predicting MOF-adsorbate configurations is therefore important for screening candidate materials. However, traditional approaches are computationally expensive and require a known host structure, while existing generative models rely on rigid-body assumptions and do not explicitly model adsorbates. We introduce AtomMOF, a scalable all-atom flow-matching model that jointly predicts MOF-adsorbate structures from discrete building blocks and adsorbates. Built on a Diffusion Transformer with a building block-based pairwise attention bias, AtomMOF operates in an unconstrained all-atom space and exhibits clear scaling behavior. To improve the structural validity of flexible all-atom models, we also propose Feynman-Kac (FK) steering with machine-learned interatomic potentials (MLIPs). On the BW dataset, AtomMOF achieves a 58.17% relative increase in match rate and a 31.84% relative reduction in RMSD compared with prior work. MLIP-guided steering further improves validity by 28.7% and reduces formation energy error by 86.5%. On ODAC25, AtomMOF generates adsorption configurations faster than GCMC and, when combined with MLIP relaxation, identifies lower-energy configurations than those in the reference dataset.
Authors:
Vinícius Da Silva, Isabelle R de Melo, Matheus L de Lima Bessa, Guilherme Schardong, Luiz Schirmer, Andre Araujo, Nuno Gonçalves, HELIO C LOPES, Alberto Raposo, Luiz Velho, Tiago NovelloAbstract: Encoding input coordinates with sinusoidal functions into multi-layer perceptrons (MLPs) has proven effective for implicit neural representations (INRs) of surfaces defined as zero-level sets. However, existing methods often struggle to balance training efficiency, rendering speed, and noise robustness: single-MLP approaches are expensive at inference, grid-based representations are fast but can limit surface smoothness and overfit input noise, and previous multiscale approaches frequently capture noise and produce artifacts due to hard spectral truncation. To address these limitations, we propose M-plicits, a multiscale framework that models surfaces as a residual sum of MLPs trained via a sequence of nested neighborhoods. Unlike existing residual approaches that rely on standard domain-wide sampling and require costly mesh extraction for visualization, our method strictly localizes supervision to narrow bands around the previous zero-level sets. This nested design naturally provides robustness against noisy input data: the coarse network acts as a low-pass filter that establishes a clean geometric prior, while subsequent residuals progressively refine the geometry without fitting to high-frequency artifacts. We further introduce a multiscale sphere-tracing algorithm and a GEMM-based analytical normal computation that bypasses auto-differentiation entirely, yielding high-fidelity real-time rendering. On Stanford and Thingi32, M-plicits achieves the best mean Chamfer distance in the coarse configuration and the best median Chamfer distance and IoU in the fine configuration, with substantially better noise robustness than iNGP, BACON, and IDF, while using an order of magnitude fewer parameters than grid-based baselines. Code and models will be released.
Abstract: Autoregressive video generation has emerged as a powerful paradigm for World Action Models (WAMs). However, existing approaches suffer from slow training convergence particularly at high frame rates, limited converged accuracy, and slow inference due to iterative video denoising, as the training supervision is confined to the current chunk without explicit signals about future dynamics. In this paper, we present Next Forcing, a multi-chunk prediction (MCP) framework for causal world modeling that enables faster training, higher accuracy, and accelerated inference. Inspired by multi-token prediction in large language models, Next Forcing introduces an MCP training objective that augments the main model with lightweight auxiliary MCP modules to simultaneously denoise video chunks at multiple future temporal horizons (next^1, next^2, next^3 chunks). These MCP modules form a causal chain across prediction depths, where intermediate features fused from multiple layers of the main model are leveraged to predict future dynamics, allowing near-future predictions to inform farther-future ones and providing dense multi-scale temporal supervision back to the main model. During training, the MCP modules significantly accelerate convergence and improve converged accuracy, especially at high frame rates. At 50 fps, our method achieves a 93.1% relative improvement over the baseline LingBot-VA at 5k training steps and achieves 2.3× faster convergence. At inference, the MCP modules can be retained to predict the next video chunk in parallel with the current one, accelerating generation. Next Forcing establishes new state-of-the-art results on the RoboTwin benchmark (94.1%/93.5% on Clean/Random) and demonstrates significant improvements on PhyWorld, a benchmark evaluating adherence to physical laws in video generation.
Abstract: Prior-data Fitted Networks (PFNs) have reframed supervised tabular learning as single-pass in-context inference without per-dataset optimization. Extending this paradigm to unsupervised clustering is appealing yet fundamentally more challenging, due to absent supervision, unknown cluster cardinality, and label switching inherent to partition outputs. Existing PFN-based clustering methods address these challenges only partially, either requiring known cardinality as input or relying on unstable label-ordering conventions and overly restrictive synthetic priors. We introduce TabClustPFN, a clustering PFN that resolves all challenges jointly through co-designed prior, objective, and architecture. Our hybrid pretraining prior captures heterogeneous real-tabular geometry; our decoupled partition inference network-cardinality inference network architecture jointly infers cluster assignments and cardinality in a single pass; and our SoftARI training objective is permutation-invariant by construction, eliminating the need for any label-ordering convention. On a 44 curated real-world tabular benchmark, TabClustPFN achieves state-of-the-art clustering performance against classical, deep, and amortized baselines, with runtime comparable to efficient baselines.
Abstract: Leveraging the natural spatiotemporal energy decay in video diffusion offers a path to efficiency, yet relying solely on rigid static masks risks losing critical long-range information in complex dynamics. To address this issue, we propose DynamicRad, a sparse-attention framework that constrains adaptive selection using a radial locality prior. DynamicRad introduces a dual-mode strategy: static-ratio for speed-optimized execution and dynamic-threshold for quality-first filtering. To avoid online search over sparse indices, we integrate an offline Bayesian Optimization (BO) pipeline with a semantic motion router. The router maps prompt embeddings to BO-selected sparsity regimes with a single projection module. Unlike online profiling methods, our offline BO optimizes attention reconstruction error (MSE) on a proxy task and reuses the selected configurations during inference. Experiments on HunyuanVideo and Wan2.1-14B demonstrate that DynamicRad achieves a strong efficiency-quality trade-off among FlashAttention-compatible sparse-attention baselines, obtaining 1.7x-2.5x inference speedups with over 80% effective sparsity. In some long-sequence settings, the dynamic mode matches or improves over the dense baseline on automatic quality metrics, while optional Mask-Aware LoRA further improves long-horizon coherence. Code is available at https://anonymous.4open.science/r/DynamicRad-5D55/.
Abstract: Multimodal large language models (MLLMs) achieve ever-stronger performance on visual-language tasks. Even as traditional visual question answering benchmarks approach saturation, reliable deployment requires satisfying low error tolerances in real-world out-of-distribution (OOD) scenarios. Precisely, selective prediction aims to improve coverage, i.e.\ the share of inputs the system answers, while adhering to a user-defined risk level. This is typically achieved by assigning a confidence score to each answer and abstaining on those that fall below a certain threshold. To enable reliable generalization, we require reasoner models to produce localized visual evidence while answering, and design a selector that explicitly learns to estimate the quality of the localization provided by the reasoner. We show that SIEVES (Selective Prediction through Visual Evidence Scoring) improves coverage by up to three times on challenging OOD benchmarks (V Bench, HR-Bench-8k, MME-RealWorld-Lite, VizWiz, and AdVQA), compared to non-grounding baselines. Beyond better generalization to OOD tasks, the design of the SIEVES selector enables transfer to proprietary reasoners without access to their weights or logits, such as o3 and Gemini-3-Pro, providing coverage boosts beyond those attributable to accuracy alone. We highlight that SIEVES generalizes across all five tested OOD datasets and reasoner models (Pixel-Reasoner, o3, and Gemini-3-Pro), without benchmark- or reasoner-specific training or adaptation.
Abstract: Recent advances in generative modeling show that pretrained representations can improve generation as conditioning features or alignment targets. Motivated by this, we study protein representations for predicting structures beyond conventional function annotation. We propose TRIPROREP, a structure-aware pretraining method that jointly models three aligned residue-level views: amino-acid identity, backbone geometry, and local full-atom geometry, discretely encoded via VQ-VAE tokenizers. By pretraining to recover original tokens from generator-corrupted views, TRIPROREP learns to distinguish plausible but incorrect cross-view augmentations from the original protein. We further introduce REPSP, a benchmark for evaluating protein representations in structure-predictive settings. REPSP tests three uses of representations: homodimer co-folding from apo-chain representations, residue-level prediction of homodimer-derived interaction properties, and representation-aligned monomer structure prediction. Across these tasks, TRIPROREP improves over sequence-only and prior structure-aware representation models, while maintaining competitive performance on conventional benchmarks.
Abstract: Large language models increasingly use external tools such as web search and document retrieval to solve information-intensive tasks. However, multi-hop tool use in complex tasks introduces substantial latency, since the model must repeatedly wait for tool observations before continuing. We study how to accelerate such trajectories without changing the final trajectory the model would have taken without acceleration, assuming access to faster but less reliable speculator tools. We develop a theoretical framework for lossless speculation in multi-hop tool-use settings, characterizing the optimal achievable latency gain. We propose SpecHop, a continuous speculation framework that maintains multiple speculative threads, verifies predicted observations asynchronously as target tool outputs arrive, commits correct branches, and rolls back incorrect ones. This preserves task accuracy while reducing wall-clock latency. We show that SpecHop can approach the oracle latency gain with sufficient active threads. Empirically, we evaluate SpecHop on retrieval-augmented multi-hop tasks and find that its latency gains closely match theoretical predictions, reaching up to 40% latency reduction in some settings.
Abstract: Multimodal large language models (MLLMs) still struggle with spatial understanding under the dominant perspective-image paradigm, which inherits the narrow field of view of human-like perception. For navigation, robotic search, and 3D scene understanding, 360^\circ panoramic sensing offers a form of supersensing by capturing the entire surrounding environment at once. However, existing MLLM pipelines typically decompose panoramas into multiple perspective views, leaving the spherical structure of equirectangular projection (ERP) largely implicit. In this paper, we study pano-native understanding, which requires an MLLM to reason over an ERP panorama as a continuous, observer-centered space. To this end, we first define the key abilities for pano-native understanding, including semantic anchoring, spherical localization, reference-frame transformation, and depth-aware 3D spatial reasoning. We then build a large-scale metadata construction pipeline that converts mixed-source ERP panoramas into geometry-aware, language-grounded, and depth-aware supervision, and instantiate these signals as capability-aligned instruction tuning data. On the model side, we introduce PanoWorld with Spherical Spatial Cross-Attention, which injects spherical geometry into the visual stream. We further construct PanoSpace-Bench, a diagnostic benchmark for evaluating ERP-native spatial reasoning. Experiments show that PanoWorld substantially outperforms both proprietary and open-source baselines on PanoSpace-Bench, H^\astBench, and R2R-CE Val-Unseen benchmarks. These results demonstrate that robust panoramic reasoning requires dedicated pano-native supervision and geometry-aware model adaptation. All source code and proposed data will be publicly released.
Authors: Taylan Soydan, Miguel Bessa, Dirk Mohr, Rui Barreira
Abstract: Selective state space models (SSMs), such as Mamba, achieve strong per-token expressivity by making the time discretization step \Delta a learned function of the input. However, in doing so, \Delta ceases to represent a physical sampling interval, limiting its irregular time series modeling capability. Continuous-time SSMs, such as S5, preserve the physical meaning of \Delta and handle irregular timestamps natively, but their dynamics remain linear time-invariant (LTI), limiting per-token expressivity. We propose TIDES, a selective SSM variant that reconciles selective and continuous architectures by moving input-dependence off the step size and onto the diagonal state matrix. As a result, \Delta retains its physical meaning, tied to the state discretization, allowing the model to handle irregular timestamps natively without sacrificing the per-token expressivity that makes selective SSMs effective. We show this on a novel \emphFading Flash experimental benchmark, a compact controlled diagnostic for sequence models that jointly tests input-dependence and extrapolation to out-of-distribution \Delta values, and isolates the distinct failure modes of current state-of-the-art architectures that TIDES avoids by construction. On large-scale benchmarks, TIDES sets the new state-of-the-art average rank on UEA time-series classification and the Physiome-ODE regression benchmark.
Abstract: The vanilla self-attention mechanism in Transformers can be viewed as a two-layer fast-weight MLP, whose weights are dynamically induced by inputs and whose hidden dimension is equal to the sequence length N. As the context extends, the expressive capacity of such an N-width MLP increases, but it becomes unscalable for extremely long sequences. Recently, this fast-weight perspective has motivated the Mixture-of-Experts (MoE) attention, which partitions the sequence into rigid blocks, treats them as fast-weight experts, and sparsely routes the tokens to them. In this paper, we elevate this perspective to a unifying framework for efficient attention mechanisms, interpreting them as making fast weights scalable through either routing or compression, and organizing them into a five-dimensional taxonomy. Then, we propose Mixture-of-Top-k Attention (MiTA), which employs a small set of landmark queries to gather top-k attended key-value pairs as query-aware and deformable routed experts, while compressing the N-width MLP into a narrower shared expert. Consequently, MiTA improves the flexibility of prior MoE attention, from rigid to deformable fast-weight experts, as well as the scalability of prior top-k attention, from query-specific set to reusable top-k set. Our experiments on vision tasks demonstrate the superior effectiveness and efficiency of MiTA, while also uncovering intriguing properties such as an emergent token-pruning effect and easy generalization from standard attention.
Abstract: Continuous diffusion has served as a foundation for high-fidelity, controllable, and few-step generation across continuous data modalities such as images, videos, and molecular structures. However, in language modeling, prior continuous diffusion language models (DLMs) lag behind discrete counterparts. Existing categorical and simplex-based approaches operate over extremely large and sparse language spaces, while prior embedding-space approaches avoid this sparsity but lack a well-selected design space. In this work, we close this gap with by connecting embedding-space DLMs to Flow Matching, alongside three key innovations: (1) we derive a novel ODE-based NLL bound for principled evaluation of continuous flow-based language models; (2) we propose an information-uniform principle for setting the noise schedule, which motivates a learnable noise scheduler based on a Gumbel distribution; and (3) we revise prior training protocols by incorporating self-conditioning, which improves both likelihood and sample quality for embedding-space DLMs and behaves differently from its use in discrete diffusion. Putting everything together, LangFlow is competitive with top discrete DLMs on both perplexity (PPL) and generative perplexity (Gen. PPL), reaching a PPL of on OpenWebText. It also exceeds autoregressive baselines in zero-shot transfer on 4 out of 7 benchmarks. As the first continuous DLM shown to rival discrete diffusion in both generative quality and perplexity, LangFlow provides clear evidence that continuous diffusion is a promising paradigm for language modeling. Code will be released upon acceptance.
Abstract: Post-training compression of Transformer models commonly relies on truncated singular value decomposition (SVD) strategies leading to low-rank approximations of the original model weights. However, the underlying single shared subspace modeling can degrade accuracy even at moderate compression. Sparse dictionary learning provides a more flexible union-of-subspaces representation, but existing approaches often suffer from costly iterative dictionary and coefficient updates. We propose ransformers), a training-free compression framework that uses a small calibration dataset to estimate a sparse weight factorization. COMPOT employs orthogonal dictionaries that enable closed-form Procrustes updates for the dictionary and analytical single-step sparse coding for the coefficients, reducing significantly the overall optimization cost. To handle heterogeneous layer sensitivity under a global compression budget, COMPOT further introduces a one-shot dynamic allocation strategy that adaptively redistributes layer-wise compression rates. Extensive experiments across diverse architectures and tasks show that COMPOT consistently delivers a superior quality–compression trade-off over strong low-rank and sparse baselines, while remaining fully compatible with post-training quantization for extreme compression. Code will be released upon acceptance.
Abstract: Large language models are increasingly used as human simulators for interactive evaluation and social simulation, yet current post-training often produces homogeneous, overly helpful systems unlike real people, creating a behavioral Sim2Real gap that prompting alone cannot close. We present Odyssim, a large-scale open study of behavioral foundation models organized around SOUL, whose five capability Axes (CONV, SS, COG, ROLE, EVAL) jointly guide data curation and evaluation. We release three artifacts: (1) SOULIndex, a 23-task evaluation suite for human-behavior simulation; (2) OdysSim, a corpus of 21.4M interactions, approximately 10B tokens, from 63 public sources, unified into a conversational format and retrofitted with back-generated social contexts; and (3) an end-to-end recipe combining midtraining, task-specific RL with VFRL, and expert distillation into a single deployable model. On SOULIndex, OdysSim-8B achieves an average score of 62.4, comparable to frontier models such as GPT-5.5 (62.6), ranks first on 8 of 23 tasks, and outperforms prior open behavioral-simulation baselines by 14.2 points on average. We also identify and mitigate reward-hacking failures in judge-based behavioral RL, including judge manipulation and short-response collapse. We release the corpus, benchmark, recipes, and checkpoints.
Abstract: Vision-language models with extended reasoning succeed on complex problems, but many real-world problems require external tools that internal reasoning alone cannot reach. Agentic reasoning therefore interleaves two behaviors with a structural asymmetry: thinking (the self-contained default) and tool use (a high-variance auxiliary acting). We refer to this asymmetry as the Thinking-Acting Gap. Under standard RL recipes like GRPO, the gap manifests as two diagnostic symptoms during training: tool use is attempted on only ~30% of rollouts, and when attempted, the tool-using rollouts within a group are all-wrong on ~40% of questions, suppressing the learning signal at the tool calls that needed it. We propose AXPO (Agent eXplorative Policy Optimization): for each all-wrong tool-using subgroup, AXPO fixes the thinking prefix and resamples the tool call and its continuation, paired with uncertainty-based prefix selection. Across nine multimodal benchmarks and three scales of Qwen3-VL-Thinking, SFT+AXPO outperforms SFT+GRPO at average (+1.8pp Pass@1 and +1.8pp Pass@4 at 8B on average) and 8B with SFT+AXPO surpasses the 32B Base on Pass@4 with 4x fewer parameters.
Authors: Zhiyuan Zhai, Xin Wang
Abstract: Group-relative RL training (GRPO) samples a small group of parallel rollouts for every training prompt and uses their within-group reward spread to compute per-trajectory advantages. In agentic environments each rollout is a long multi-turn dialogue with one LLM call per step, so this multi-sample multiplier dominates the total training cost. When every rollout of a prompt ends with the same reward, the group has zero reward variance and contributes no gradient, so the extra rollouts add no information; such groups are common in practice (typically around 40% of all groups), so the wasted-compute fraction is substantial rather than marginal. Existing methods filter such groups at the prompt level, either after their rollouts are paid for or before any rollout begins, but both decide without using information that becomes available during the rollout itself. We instead ask whether the in-group divergence between the partial trajectories at an intermediate step can already predict that the group will be zero-variance: when the parallel rollouts have already converged on the same action prefix, the group is on track to produce a single reward, and we can stop early. We propose a one-parameter gate that stops a group when the mean pairwise prefix edit distance between its partial action sequences falls below a threshold. On a 60-iteration on-policy GRPO run on ALFWorld with Qwen2.5-7B, averaged over four random seeds, the gated arm finishes 10.7% faster in wall-clock (bootstrap 95% CI excludes 0) and shifts held-out success rate on 50 unseen tasks by +2.5 pp, with the held-out gain tracing to a measurable reduction in zero-advantage gradient-batch dilution.
Abstract: Survival analysis provides a powerful statistical framework for modeling time-to-event outcomes in the presence of censoring. However, selecting an appropriate estimator from the many specialized survival approaches often requires substantial methodological and domain expertise. We introduce SurvivalPFN, a prior-data fitted network that amortizes Bayesian inference for censored observations through in-context learning. SurvivalPFN is pretrained on a diverse family of synthetic, identifiable, and right-censored data-generating processes, enabling it to amortize survival analysis in a single forward pass during inference. As a result, the model adapts to the effective complexity of each dataset without task-specific training or hyperparameter tuning, avoids restrictive parametric assumptions, and produces calibrated survival distributions. In a large-scale benchmark spanning 61 survival datasets, 21 methods, and 5 evaluation metrics, SurvivalPFN achieves strong predictive performance and often improves upon established survival models. These results suggest that SurvivalPFN offers a principled and practical foundation model for survival analysis, with potential applications in high-impact domains such as healthcare, finance, and engineering.
Abstract: While equivariant architectures are standard for processing symmetric data, there is growing interest in achieving equivariance by applying group averaging or canonization to non-equivariant backbones. However, the theoretical generalization properties of these alternative strategies remain poorly understood. We introduce a theoretical framework to analyze the generalization error of these methods by bounding their covering numbers. We establish a rigorous generalization hierarchy: the error bounds of canonized models are at best equal to the error bounds of structurally equivariant and group-averaged models, and at worst equal to the bounds of non-equivariant baselines. Furthermore, we show that there exist "optimal" canonizations which attain the optimal error bounds, and "poor" canonizations which attain the non-equivariant error bounds, and that this depends on the regularity of the canonization. Finally, applying this framework to permutation groups in point cloud processing, we rigorously prove that the covering number of lexicographical sorting grows exponentially with point cloud dimension, whereas Hilbert curve canonization guarantees polynomial growth. This provides the first formal theoretical justification for the empirical success of Hilbert curve serialization in state-of-the-art point cloud architectures. We conclude with experiments which support our theoretical claims.
Abstract: Long-horizon reasoning remains a central challenge for large language models (LLMs) under sparse-reward regimes. We argue that this brittleness arises from two biases induced by complex reasoning spaces: an exploration bias, where models are drawn toward locally plausible but structurally unstable branches, and a compounding bias, where small local deviations accumulate across depth and suppress rare rewards. We introduce Symbolic Closure Analysis (SCA) as a theoretical lens characterizing how branching structures and sparse rewards induce these biases in long-horizon reasoning with local admissibility, and as a design principle for structural priors in less formal reasoning tasks. Motivated by this analysis, we propose \model (Structural Admissibility-Guided Exploration), a unified framework that injects structural guidance to alleviate exploration bias and compounding bias in long-horizon reasoning. \model combines two complementary structural dimensions: algebraic sparsification, which projects locally admissible candidates onto operator-indexed algebraic subspaces to suppress spurious branching and mitigate exploration bias, and hyperbolic structural guidance, which embeds reasoning states into a negatively curved space to provide dense depth-wise signals and mitigate compounding bias. Across 12 benchmarks and 7 model families, \model outperforms competitive baselines. In particular, \model achieves up to an 8-fold improvement on the Andrews-Curtis problem, an open real-world long-horizon task. Code is available at: \urlhttps://anonymous.4open.science/r/SAGE-Long-Horizon-Reasoning-AD70.
Abstract: Large language models (LLMs) deliver strong performance, but their high compute and memory costs make deployment difficult in resource-constrained scenarios. Weight-only post-training quantization (PTQ) is appealing, as it reduces memory usage and enables practical speedup without low-bit operators or specialized hardware. However, accuracy often degrades significantly in weight-only PTQ at sub-4-bit precision, and our analysis identifies two main causes: (1) down-projection matrices are a well-known quantization bottleneck, but maintaining their fidelity often requires extra bit-width; (2) weight quantization induces activation deviations, but effective correction strategies remain underexplored. To address these issues, we propose D^2Quant, a novel weight-only PTQ framework that improves quantization from both the weight and activation perspectives. On the weight side, we design a Dual-Scale Quantizer (DSQ) tailored to down-projection matrices, with an absorbable scaling factor that significantly improves accuracy without increasing the bit budget. On the activation side, we propose Deviation-Aware Correction (DAC), which incorporates a mean-shift correction within LayerNorm to mitigate quantization-induced activation distribution shifts. Extensive experiments show that D^2Quant achieves superior weight-only PTQ performance at sub-4-bit precision. We will release the code and quantized models of D^2Quant.
Authors: Nermin Samet, Gilles Puy, Renaud Marlet
Abstract: This paper presents a new method for the zero-shot open-vocabulary semantic segmentation (OVSS) of 3D automotive lidar data. To circumvent the recognized image-text modality gap that is intrinsic to approaches based on Vision Language Models (VLMs) such as CLIP, our method relies instead on image generation from text, to create prototype images. Given a 3D network distilled from a 2D Vision Foundation Model (VFM), we then label a point cloud by matching 3D point features with 2D image features of these prototypes. Our method is state-of-the-art for OVSS on nuScenes and SemanticKITTI.
Abstract: Scaling laws for large language models depend critically on the optimizer and parameterization. Existing hyperparameter transfer laws are mainly developed for first-order optimizers, and they do not structurally prevent training instability at scale. Recent hypersphere optimization methods constrain weight matrices to a fixed-norm hypersphere, offering a promising alternative for more stable scaling. We introduce HyperP (Hypersphere Parameterization), the first framework for transferring optimal learning rates across model width, depth, training tokens, and Mixture-of-Experts (MoE) granularity under the Frobenius-sphere constraint with the Muon optimizer. We prove that weight decay is a first-order no-op on the Frobenius sphere, show that Depth-\muP remains necessary, and find that the optimal learning rate follows the same data-scaling power law with the ``magic exponent'' 0.32 previously observed for AdamW. A single base learning rate tuned at the smallest scale transfers across all compute budgets under HyperP, yielding 1.58× compute efficiency over a strong Muon baseline at 6×10^21 FLOPs. Moreover, HyperP delivers transferable stability: all monitored instability indicators, including Z-values, output RMS, and activation outliers, remain bounded and non-increasing under training FLOPs scaling. We also propose SqrtGate, an MoE gating mechanism derived from the hypersphere constraint that preserves output RMS across MoE granularities for improved granularity scaling, and show that hypersphere optimization enables substantially larger auxiliary load-balancing weights, yielding both strong performance and good expert balance.
Abstract: The advent of Multimodal Large Language Models (MLLMs) has unlocked the potential for end-to-end document parsing and translation. However, prevailing benchmarks such as OmniDocBench and DITrans are dominated by pristine scanned or digital-born documents. They do not adequately represent the intricate challenges of real-world capture conditions, such as geometric distortions and photometric variations. To fill this gap, we introduce DocPTBench, a comprehensive benchmark specifically designed for Photographed Document Parsing and Translation. DocPTBench comprises over 1,300 high-resolution photographed documents from multiple domains, includes eight translation scenarios, and provides meticulously human-verified annotations for both parsing and translation. Our experiments demonstrate that transitioning from digital-born to photographed documents results in a substantial performance decline: popular MLLMs exhibit an average accuracy drop of 18% in end-to-end parsing and 12% in translation, while specialized document parsing models show a more prominent average decrease of 22%. This substantial performance gap highlights the unique challenges posed by documents captured in real-world conditions and reveals the limited robustness of existing models. The benchmark and related code will be made publicly available for further research.
Authors: Nathan Weill, Kaizheng Wang
Abstract: We propose a principled framework for unsupervised domain adaptation under covariate shift in kernel Generalized Linear Models (GLMs), encompassing kernelized linear, logistic, and Poisson regression with ridge regularization. Our goal is to minimize prediction error in the target domain by leveraging labeled source data and unlabeled target data, despite differences in covariate distributions. We partition the labeled source data into two batches: one for training a family of candidate models, and the other for building an imputation model. This imputation model generates pseudo-labels for the target data, enabling robust model selection. We establish non-asymptotic excess-risk bounds that characterize adaptation performance through an “effective labeled sample size,” explicitly accounting for the unknown covariate shift. Experiments on synthetic and real datasets demonstrate consistent performance gains over source-only and importance weighting baselines.
Abstract: Vision-Language-Action (VLA) models enable instruction-following robotic manipulation, but they are typically pretrained on 2D data and lack 3D spatial understanding. An effective approach is representation alignment, where a strong vision foundation model is used to guide a 2D VLA model. However, existing methods usually apply supervision at only a single layer, failing to fully exploit the rich information distributed across depth; meanwhile, naïve multi-layer alignment can cause gradient interference. We introduce ROCKET, a residual-oriented multi-layer representation alignment framework that formulates multi-layer alignment as aligning one residual stream to another. Concretely, ROCKET employs a shared projector to align multiple layers of the VLA backbone with multiple layers of a powerful 3D vision foundation model via a layer-invariant mapping, which reduces gradient conflicts. We provide both theoretical justification and empirical analyses showing that a shared projector is sufficient and outperforms prior designs, and further propose a Matryoshka-style sparse activation scheme for the shared projector to balance multiple alignment losses. Our experiments show that, combined with a training-free layer selection strategy, ROCKET requires only about 4% of the compute budget while achieving 98.5% state-of-the-art success rate on LIBERO. We further demonstrate the superior performance of ROCKET across LIBERO-Plus, RoboTwin, multiple VLA models, and a real-world industrial warehouse picking task. The code and model weights will be released upon acceptance of the paper.
Authors:
Sihan Fu, Oucheng Liu, Shiyuan Wang, Jin Shi, Chengkun WeiAbstract: Code agents are rapidly lowering the barrier for developers to work with unfamiliar repositories, and are increasingly adopted for tasks such as building prototypes, reproducing experiments, and adding features. Reliably executing these tasks depends on a critical prerequisite: the target repository must first be successfully set up and brought into a usable development state. This bootstrapping process requires substantial trial-and-error exploration by code agents, yet the resulting knowledge (e.g., resolved dependencies, repair strategies) remains trapped in a single conversation, unavailable to future agents or developers. We therefore formulate repository bootstrapping as a reusable startup knowledge problem and introduce BootstrapAgent, a multi-agent framework that distills bootstrap exploration into a persistent, verifiable, agent-consumable .bootstrap contract. Through CI evidence extraction, structured planning, deterministic Docker-based verification, and trace-driven repair, BootstrapAgent generates a contract covering environment setup, diagnostic checks, minimal verification, and accumulated repair knowledge. We further propose warm repair with clean replay to accelerate iterative debugging without sacrificing cold-start reproducibility, and a two-level verification strategy to prevent reward hacking. Experiments on three benchmarks show that BootstrapAgent achieves a 92.9% success rate, outperforming the baseline by over 10% while reducing downstream agent token usage by 25.9% and build time by 22.3%.
Abstract: Practical autonomous driving requires models that generalize by reasoning through spatial-temporal possibilities to exclude unsafe outcomes. While state-of-the-art (SOTA) methods use parallel planning architectures, they fail to explicitly couple speed decisions with agent behavior along the driving path, leading to suboptimal coordination. To address this, we propose a cascaded framework that transforms longitudinal planning from an independent prediction task into a path-conditioned reasoning process. On the model side, we introduce an anchor-based regression design that conditions longitudinal prediction on the lateral drive path, and reformulate longitudinal planning as 1D displacement prediction along the path. This reduces geometric uncertainty and sharpens the model's focus on interaction-driven dynamics. On the data side, we introduce a planning-oriented data augmentation strategy that simulates rare safety-critical events by programmatically inserting agents and relabeling longitudinal targets to enforce collision avoidance. Evaluated on the challenging Bench2Drive benchmark, our method achieves SOTA performance with a driving score of 89.07 and a success rate of 73.18%, demonstrating significantly improved coordination and safety. Further evaluation on Fail2Drive confirms strong generalization to rare edge cases where parallel formulations typically fail. Visualization videos are provided in the supplementary material.
Abstract: Unified Multimodal Models (UMMs) excel in general tasks but struggle to bridge the gap between personalized understanding and generation. Prior works largely rely on implicit token-level alignment via supervised fine-tuning, which fails to fully capture the potential synergy between comprehension and creation. In this work, we propose Sync-R1, an end-to-end reinforcement learning framework that jointly optimizes personalized understanding and generation within a single, explicit reasoning loop. Through this unified feedback process, Sync-R1 enables personalized comprehension to guide content creation, while the resulting generation quality reciprocally refines understanding within an integrated reward landscape. To efficiently orchestrate this dual-task synergy, we introduce Sync-GRPO, a reinforcement learning method utilizing an ensemble reward system. Furthermore, we propose Dynamic Group Scaling (DGS), which adaptively filters low-potential trajectories to reduce gradient variance and accelerate convergence. To better reflect real-world complexity, we introduce UnifyBench++, featuring denser textual descriptions and richer user contexts. Experimental results demonstrate that Sync-R1 achieves state-of-the-art performance, showcasing superior cross-task reasoning and robust personalization without requiring complex cold-start procedures.
Abstract: Built on pretrained vision foundation models (VFMs), representation autoencoders (RAEs) have recently emerged as a promising approach for constructing semantically rich latent spaces for image generation. However, their reconstruction quality often remains suboptimal, largely because deep VFM representations do not preserve sufficient fine-grained visual detail. This limitation becomes even more severe after discretization, where missing low-level information is difficult to recover. In fact, we observe that shallow VFM features retain considerably richer local appearance and structural detail, which complements the high-level semantics carried by deep features used in existing RAEs. Motivated by this complementary property, we propose IDEAL, an In-DEpth ALignment framework for discrete representation autoencoding. By jointly aligning quantized tokens with both shallow and deep VFM features, IDEAL enables the resulting discrete visual tokens to preserve both visual fidelity and rich semantics. Extensive experiments demonstrate that IDEAL yields superior reconstruction performance, achieving \mathbf0.61 rFID on ImageNet and outperforming the previous best method by \mathbf0.28. When used for autoregressive image generation, IDEAL further produces a gFID of \mathbf1.89, establishing a new state of the art for autoregressive image generation.
Authors: tai Dang, Hieu Tran, Long-Hung Pham, Sang Truong, Edward A Pham, Jeffrey S Glenn, Thang Luong
Abstract: The latest advances of diffusion models in biomolecular structure prediction are still based on distance-based supervised training without adequately accounting for physically plausible interactions needed for real-world drug discovery. We introduce \mathit\boldsymbol\pi\text-Fold, a novel physics-informed policy-based folding approach that augments reinforcement learning with physics-aware scoring functions for training biomolecular diffusion models. Specifically, we fine-tune AlphaFold3-based models using Group Relative Policy Optimization and explicit physics-based rewards such as Rosetta all-atom energy, AutoDock Vina binding affinity, and ligand strain energy. Crucially, we demonstrate that these physics-based priors translate to enhanced performance on critical downstream prediction tasks. \mathit\boldsymbol\pi\text-Fold achieves state-of-the-art results on SKEMPI for protein--protein mutation effects, CASP16 and OpenFE for protein--ligand binding affinities, Run-N-Poses for ligand geometric validity and interface quality across diverse structural fidelity metrics. On FoldBench, \mathit\boldsymbol\pi\text-Fold achieves the best overall MolProbity score of 1.30, halving steric clashes and nearly eliminating rotamer outliers. Human evaluation further confirms that \mathit\boldsymbol\pi\text-Fold generates structures that are meaningful for downstream drug discovery applications. Our code is available at \hrefhttps://anonymous.4open.science/r/pifold-45B8/README.mdgithub.com/anonymous/pifold.
Abstract: Multimodal retrieval relies heavily on single-vector retrievers, which compress rich, sequential token sequences into one single global representation. While efficient, they discard fine-grained, local evidence critical for dense retrieval tasks. Multi-vector approaches were introduced as a solution, but they strictly require training and many ignore the necessity of a globally summarizing representation. To address this, we introduce SMART, a framework that unlocks the latent multi-vector capabilities of standard single-vector models. We first demonstrate that standard contrastive training on the pooled embedding implicitly shapes the retrieval geometry of preceding hidden states via gradient flow. By applying direct late-interaction over these frozen hidden states during inference, SMART acts as a plug-and-play upgrade that consistently improves performance across diverse modalities, pushing even the state-of-the-art models to new heights on MMEB-V2. We further reveal SMART's superior performance as simple lightweight post-training not only saves time and compute, but also brings forth further improvement on Visual Document retrieval, allowing a single-vector model to outperform SoTA multi-vector counterparts. Ultimately, SMART offers both a highly efficient inference enhancement and a powerful finetuning technique for multimodal retrieval.
Abstract: Close-up rendering, zooming into a scene well beyond any training camera, is important for virtual production and interactive 3D content, yet remains an open challenge. 3D Gaussian splatting (3DGS) enables high-fidelity, real-time novel view synthesis, but its rendering quality degrades at close range. Recent diffusion-based methods that enhance the rendering by conditioning on reference images from the training set produce significant artifacts in this setting. We analyze this failure and identify its root cause: the scale gap between the close-up and reference views. We show that the features in reference-conditioned enhancement models are not scale-invariant, causing cross-view attention to retrieve incorrect correspondences when the same content appears at different scales, and that this mismatch cannot be corrected in latent space because the VAE encoder is not scale-equivariant. Building on this analysis we introduce MACRO, Multi-plane Attention for Closeup Render Optimization, a training-free method for high-quality close-up novel view synthesis from 3DGS. MACRO resolves the scale gap by leveraging the scene's known 3D structure: it decomposes the close-up into depth planes, crops and resizes references in image space to match the scale of each plane before encoding, and applies a depth-aware attention mask so each token attends only to scale-matched references. The method requires no architectural changes or additional training. We further contribute two new close-up novel view synthesis benchmarks, the first standardized evaluation protocol for this setting, and demonstrate state-of-the-art results on both, outperforming existing 3DGS and diffusion-based methods on both reconstruction and perceptual metrics.
Authors: Yichen Zang, Song Liu, Jiun-Yi Lin
Abstract: Simulation-Based Inference (SBI) serves as a vital framework for parameter inference in scientific fields where simulators involve intractable likelihoods, yet while amortized generative models offer rapid posterior estimation, they are often restricted by the specific priors used during training, thereby limiting their flexibility as prior knowledge evolves. To address this prior dependency, PriorGuide was introduced as an inference-time guidance method, but due to its intractable formulation, it relies on Gaussian approximations of the reverse transition kernel and Gaussian mixture model fitting for the prior ratio, both of which introduce systematic bias. Motivated by these limitations, we propose an unbiased test-time guidance framework that leverages Density Ratio Estimation (DRE) to learn a score guidance term, effectively decoupling the inference process from the prior training. Moreover, our framework remains agnostic to the specific density ratio estimators, making it a general and flexible framework for handling prior changes. Experimental results across multiple tasks demonstrate that our method outperforms PriorGuide, achieving superior performance on metrics such as C2ST and MMD in most tasks while maintaining strong robustness even under limited overlap between the training and target priors. Furthermore, we apply our method to Bayesian updating for parameter inference from planetary light-curve data, where it also demonstrates strong effectiveness and robustness.
Abstract: Can language models (LMs) learn to faithfully describe their internal computations? Are they better able to describe themselves than other models? We study the extent to which LMs' privileged access to their own internals can be leveraged to produce new techniques for explaining their behavior. Using existing interpretability techniques as a source of ground truth, we fine-tune LMs to generate natural language descriptions of (1) the information encoded by LM features, (2) the causal structure of LMs' internal activations, and (3) the influence of specific input tokens on LM outputs. When trained with only tens of thousands of example explanations, explainer models exhibit non-trivial generalization to new queries. This generalization appears partly attributable to explainer models' privileged access to their own internals: fine-tuning a model to explain its computations generally works better than fine-tuning a different model (even if the explainer model is significantly more capable than the target). Our results suggest not only that LMs can learn to reliably explain their internal computations, but that such explanations offer a scalable complement to existing interpretability methods.
Authors:
Zhiqi Li, Wenhuan Li, Tengfei Wang, Zhenwei Wang, Junta Wu, Haoyuan Wang, Yunhan Yang, Zehuan Huang, YANG LI, Chunchao Guo, Peidong LiuAbstract: Compositionality is critical for 3D object and scene generation, but existing part-aware 3D generation methods suffer from poor scalability due to quadratic global attention costs when increasing the number of components. In this work, we present MoCA, a compositional 3D generative model based on a novel component-level sparse attention that explicitly models inter-component dependencies. The attention mechanism features two key designs: 1) importance-based component routing that utilizes a lightweight router module for cross-component importance estimation and selects top-k relevant components for fine-grained interaction, and 2) distant components compression that preserves spatial priors while reducing computational complexity of global attention, by including compressed distant components into attention calculation instead of discarding them. With these designs, MoCA enables intricate compositional 3D asset creation with scalable number of components. Extensive experiments show MoCA outperforms baselines on both compositional object and scene generation tasks. Code and models will be made publicly available.
Abstract: Traditional methods for automating recommender system design, such as Neural Architecture Search (NAS), are often constrained by a fixed search space defined by human priors, limiting innovation to pre-defined operators. While recent LLM-driven code evolution frameworks shift fixed search space target to open-ended program spaces, they primarily rely on scalar metrics (e.g., NDCG, Hit Ratio) that fail to provide qualitative insights into model failures or directional guidance for improvement. To address this, we propose Self-EvolveRec, a novel framework that establishes a directional feedback loop by integrating a User Simulator for qualitative critiques and a Model Diagnosis Tool for quantitative internal verification. Furthermore, we introduce a Diagnosis Tool - Model Co-Evolution strategy to ensure that evaluation criteria dynamically adapt as the recommendation architecture evolves. Extensive experiments demonstrate that Self-EvolveRec significantly outperforms state-of-the-art NAS and LLM-driven code evolution baselines in both recommendation performance and user satisfaction. Our code is available at https://anonymous.4open.science/r/selfevolverec-37F5/.
Abstract: Recent studies have adapted generative Multimodal Large Language Models (MLLMs) into embedding extractors for text--vision tasks, typically by fine-tuning them to produce universal representations. However, their performance on video--text retrieval often remains below that of dedicated Video Foundation Models (VFMs). In this paper, we revisit the video--text retrieval capabilities of MLLMs. We first analyze the zero-shot retrieval capabilities of pretrained MLLMs and show that they already encode substantial retrieval-relevant information: combining intermediate-layer embeddings with a calibrated MLLM head yields strong performance without any training. To further exploit these capabilities, we introduce a lightweight text-based optimization strategy that requires no paired multimodal data. Our method uses dense video captions as rich textual surrogates for video semantics and short summaries as compact, query-like views, showing that a carefully designed text-alignment objective can improve multimodal alignment through the shared MLLM backbone. We backup these findings through in-depth analysis. Without visual supervision our method outperforms existing approaches, often by a substantial margin, and achieves state-of-the-art performance across common video--text retrieval benchmarks. More broadly, our results offer a new perspective on retrieval with pretrained MLLMs, suggesting that textual descriptions of visual content can serve as an effective substitute for large-scale visual supervision.
Authors:
Junxia Cui, Haotian Ye, Runchu Tian, Hongcan Guo, Jinya Jiang, Haoru Li, Chaojie Ren, Yiming Huang, Kaijie Zhu, Zhongkai Yu, Kun Zhou, Jingbo ShangAbstract: Diffusion large language models (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs, offering faster inference through parallel or blockwise decoding. However, their masked language modeling formulation remains incompatible with standard token-level speculative decoding, one of the most effective acceleration techniques for AR models. In AR decoding, the causal mask preserves temporally valid token-level contexts, enabling a target model to verify multiple drafted tokens in a single forward pass. In contrast, dLLMs rely on mask tokens and bidirectional attention, causing the effective context to change across denoising steps and preventing direct token-level speculative verification. To bridge this gap, we propose a simple but effective speculative decoding algorithm for diffusion language models, named SimSD, which mainly adopts a plug-and-play masking strategy that equips dLLMs with temporally valid token-level contexts for speculative decoding. Our method explicitly introduces reference tokens from draft-model predictions and designs an attention mask that regulates their interaction with current-step tokens, allowing dLLMs to compute valid logits for drafted tokens in a single forward pass. This restores the key verification ability provided by causal masking in AR models while preserving the parallel decoding advantages of dLLMs. The proposed method is training-free and can be flexibly integrated with other acceleration techniques such as KV cache and blockwise decoding. Experiments on SDAR-family dLLMs across four benchmarks show that our method achieves up to (7.46×) higher decoding throughput while maintaining and even improving average generation quality. Our code will be publicly released.
Abstract: On-policy distillation (OPD) is a standard tool for transferring teacher behavior to a smaller student, but it implicitly assumes that teacher and student predictions are comparable token by token, an assumption that fails whenever the two models tokenize the same text differently. Under heterogeneous tokenizers, exact shared-token matching silently discards a large fraction of the teacher signal at precisely the positions where vocabularies disagree. We propose \underlineSimple \underlineCross-\underlineTokenizer OPD (SimCT), which restores this signal by enlarging the supervision space: alongside shared tokens, SimCT compares teacher and student over short multi-token continuations that both tokenizers can realize, leaving the OPD loss form itself unchanged. We show that these units are the finest jointly tokenizable supervision interface, and that coarser alternatives remove teacher-student distinctions that are useful for on-policy learning. Across three heterogeneous teacher-student pairs on mathematical reasoning and code-generation benchmarks, SimCT shows consistent gains over shared-vocabulary OPD and representative cross-tokenizer baselines, with ablations confirming that the improvements come from recovering supervision discarded by exact shared-token matching. Code is available at \hrefhttps://anonymous.4open.science/r/SimCThttps://anonymous.4open.science/r/SimCT.
Abstract: Motion transfer enables controllable video generation by transferring temporal dynamics from a reference video to synthesize a new video conditioned on a target caption. However, existing Diffusion Transformer (DiT)–based methods are limited to single-object videos, restricting fine-grained control in real-world scenes with multiple objects. In this work, we introduce MotionGrounder, a DiT-based framework that firstly handles motion transfer with multi-object controllability. Our Flow-based Motion Signal (FMS) in MotionGrounder provides a stable motion prior for target video generation, while our Object-Caption Alignment Loss (OCAL) grounds object captions to their corresponding spatial regions. We further propose a new Object Grounding Score (OGS), which jointly evaluates (i) spatial alignment between source video objects and their generated counterparts and (ii) semantic consistency between each generated object and its target caption. Our experiments show that MotionGrounder consistently outperforms recent baselines across quantitative, qualitative, and human evaluations.
Authors: Joon Ha Kim, Geon-Woo Kim, Anoop Rachakonda, Daehyeok Kim
Abstract: Selecting the optimal LLM inference configuration requires evaluation across hardware, serving engines, attention backends, and model architectures, since no single choice performs best across all workloads. Profile-based simulators are the standard tool, yet they hardcode their operation set to a specific configuration and re-profile every operation from scratch, making exploration prohibitively expensive. This cost stems from a missing structural understanding: every input dimension of each operation is fixed by the model configuration or determined by the incoming request. Many model-configuration values (e.g., head size, layer count) recur across models, so the same operation runs in many configurations; a single sweep over the request-dependent dimensions can serve them all. We present Dooly, which exploits this structure to achieve configuration-agnostic, redundancy-aware profiling. Dooly performs a single inference pass, labels each input dimension with its origin via taint propagation, and selectively profiles only operations absent from its latency database; stateful operations such as attention are isolated by reusing the serving engine's own initialization code, eliminating manual instrumentation. It builds latency regression models based on the database, which becomes a drop-in backend for existing simulators. Across two GPU platforms, three attention backends, and diverse model architectures, Dooly achieves simulation accuracy within 5% MAPE for TTFT and 8% for TPOT while reducing profiling GPU-hours by 56.4% across 12 models compared to the existing profiling approach.
Abstract: , an efficient post-training method for long-horizon animated video generation that preserves visual quality and character identity. Long-form animation remains challenging because highly dynamic human motion must be synthesized against relatively static environments, making chunk-based generation prone to accumulated drift: (i) low-level quality drift, such as progressive degradation of static backgrounds, and (ii) high-level semantic drift, such as inconsistent character identity and view-dependent attributes. To address this issue, restores drifted flow trajectories by anchoring generation to a persistent latent context memory, consisting of two complementary mechanisms. maintains a context memory across chunks to propagate identity and motion in latent space while mitigating temporal forgetting. introduces an implicit restoration objective during sampling through velocity adjustment, improving within-chunk fidelity. With only lightweight LoRA tuning, outperforms state-of-the-art long-animation methods in both short- and long-horizon settings: at 10 seconds, it improves PSNR/SSIM by 8%/7% and reduces LPIPS/FID by 22%/11%; at 90 seconds, the gains increase to 15%/15% and 32%/27%, respectively. Anonymous project page:
Abstract: Masked diffusion language models generate text through iterative masked-token filling, but terminal-only rewards on final completions provide coarse credit assignment for the intermediate filling decisions that shape the generation process. We propose Diffusion-State Policy Optimization (DiSPO), a plug-in credit-assignment layer that directly optimizes intermediate filling decisions. At selected intermediate masked states, DiSPO branches by resampling the currently masked positions from rollout-cached logits, scores the resulting completions, and updates only the newly filled tokens, requiring no additional multi-step diffusion rollouts or optimizer steps. We formalize a fixed-state objective for branched completions and derive a policy-gradient estimator that reuses the same rollouts as terminal-feedback policy optimization. Experiments on LLaDA-8B-Instruct show that DiSPO consistently improves terminal-feedback baselines, including diffu-GRPO and SPG, on math and planning benchmarks under matched rollout compute and optimizer steps, supporting its use as a general plug-in for masked diffusion policy optimization.
Abstract: On-policy self-distillation (OPSD) has proven effective for post-training large language models (LLMs), yet its application to diffusion LLMs (dLLMs) remains unexplored. Existing OPSD methods are inherently autoregressive-centric. They inject privileged information via left-to-right prefix conditioning with token-level divergence supervision, a design that fundamentally conflicts with the arbitrary-order generation of dLLMs. We introduce d-OPSD, the first OPSD framework tailored for dLLMs. Our approach makes two core contributions. First, we reframe self-teacher construction by using self-generated answers as suffix conditioning, enabling the student model to learn from ``self future-experience" rather than privileged prefixes. Second, we shift supervision from token-level to step-level, aligning training with the iterative denoising process of dLLMs. Experiments across four reasoning benchmarks show that d-OPSD consistently outperforms RLVR and SFT baselines with superior sample efficiency, requiring less than 10% of the optimization steps by RLVR and opening a promising pathway for dLLM post-training.
Abstract: Video temporal grounding (VTG) takes an untrimmed video and a natural-language query as input and localizes the temporal moment that best matches the query. Existing methods rely on large, task-specific datasets requiring costly manual annotation. We introduce EvoGround, a framework of two coupled self-evolving agents, a proposer and a solver, that learn temporal grounding from raw videos without any human-labeled data. The proposer generates query--moment pairs from raw videos, while the solver learns to ground them and feeds back signals that improve the proposer in return. Through this self-reinforcing reinforcement-learning loop, the two agents are initialized from the same backbone and mutually improve across iterations. Trained on 2.5K unlabeled videos, EvoGround matches or surpasses fully supervised models across multiple VTG benchmarks, while emerging as a state-of-the-art fine-grained video captioner without manual labels.
Authors:
Yuang Ai, Jiaming Han, Shaobin Zhuang, Weijia Mao, Xuefeng Hu, Ziyan Yang, Zhenheng Yang, Yali Wang, Xiangyu Yue, Hao Chen, Huaibo HuangAbstract: We present BitDance, a scalable autoregressive (AR) image generator that predicts binary visual tokens instead of codebook indices. With high-entropy binary latents, BitDance lets each token represent up to \mathbf2^256 states, yielding a compact yet highly expressive discrete representation. Sampling from such a huge token space is difficult with standard classification. To resolve this, BitDance uses a binary diffusion head: instead of predicting an index with softmax, it employs continuous-space diffusion to generate the binary tokens. Furthermore, we propose next-patch diffusion, a new decoding method that predicts multiple tokens in parallel with high accuracy, greatly speeding up inference. On ImageNet 256×256, BitDance achieves an FID of 1.24, the best among AR models. With next-patch diffusion, BitDance beats state-of-the-art parallel AR models that use 1.4B parameters, while using 5.4× fewer parameters (260M) and achieving 8.7× speedup. For text-to-image generation and image editing, BitDance trains on large-scale multimodal tokens and generates high-resolution, photorealistic images efficiently, showing strong performance and favorable scaling. When generating 1024×1024 images, BitDance achieves a speedup of over 30× compared to prior AR models. We release code and models to facilitate further research on AR foundation models.
Abstract: Tool use enables large language models to solve complex tasks through sequences of API calls, yet existing reinforcement learning approaches fail to scale to multi-step composition settings. Outcome-based rewards provide only sparse feedback, while trajectory-supervised rewards depend on annotated reference solutions, penalizing valid alternatives and limiting scalability. We propose TIER: Trajectory-Invariant Execution Rewards, a reward framework that derives supervision directly from function schemas and runtime execution, rather than from reference trajectories. The reward decomposes into format validity, schema adherence, execution success, and answer correctness, providing dense, interpretable sequence-level feedback derived from fine-grained verification of individual steps of tool use. This design allows any valid execution path to receive credit, naturally supporting multiple solution strategies and adapting to evolving tool interfaces. On DepthBench, a compositional benchmark stratified by depth (1 to 6 steps), TIER achieves >90% accuracy across steps, where trajectory-supervised rewards collapse beyond step-4. We further demonstrate consistent gains on benchmarks like BFCL v3 and NestFUL. Ablation studies confirm that all reward components are necessary, highlighting the importance of multi-level supervision for compositional reasoning.
Abstract: Text-guided point-cloud grounding enables robots to ground natural-language descriptions in 3D environments, making it important for embodied AI and human-robot interaction. Existing coarse-to-fine methods primarily rely on global descriptors for submap retrieval, but these descriptors often compress object semantics, spatial relations, and scene layouts into a single vector, thereby limiting discriminability in large-scale scenes. We propose HiGraLoc, a multi-level cross-modal alignment framework that improves the coarse retrieval stage through three complementary branches. The instance branch uses a Hyperbolic Instance Structure Encoder to model hierarchical object semantics, the relation branch aggregates reliability-weighted pairwise spatial relations, and the global branch employs a Global Spectral Graph Encoder to capture multi-frequency scene structure. A Text2Loc-style coordinate regression stage then refines the location estimate within the retrieved submap. Experiments on KITTI360Pose show that HiGraLoc achieves a 19% improvement in Top-1 recall@10m over existing state-of-the-art methods. Our code and dataset are available at https://github.com/Anonymous09871745/HiGraLoc.
Authors:
Yihao Meng, Zichen Liu, Hao Ouyang, Qiuyu Wang, Ka Leong Cheng, Yue Yu, Hanlin Wang, Xing Zhu, Yujun Shen, Qifeng Chen, Huamin QuAbstract: Autoregressive video generation aims at real-time, open-ended synthesis. Yet, cinematic storytelling is not merely the endless extension of a single scene; it requires progressing through evolving events, viewpoint shifts, and discrete shot boundaries. Existing autoregressive models often struggle in this setting. Trained primarily for short-horizon continuation, they treat long sequences as extended single shots, inevitably suffering from motion stagnation and semantic drift during long rollouts. To bridge this gap, we introduce CausalCine, an interactive autoregressive framework that transforms multi-shot video generation into an online directing process. CausalCine generates causally across shot changes, accepts dynamic prompts on the fly, and reuses context without regenerating previous shots. To achieve this, we first train a causal base model on native multi-shot sequences to learn complex shot transitions prior to acceleration. We then propose Content-Aware Memory Routing (CAMR), which dynamically retrieves historical KV entries according to attention-based relevance scores rather than temporal proximity, preserving cross-shot coherence under bounded active memory. Finally, we distill the causal base model into a few-step generator for real-time interactive generation. Extensive experiments demonstrate that CausalCine significantly outperforms autoregressive baselines and approaches the capability of bidirectional models while unlocking the streaming interactivity of causal generation.
Abstract: Large Reasoning Models (LRMs) exhibit backtracking and self-verification mechanisms that enable them to revise intermediate steps and reach correct solutions, yielding strong performance on complex logical benchmarks. We hypothesize that such behaviors are beneficial only when the model has sufficiently strong “critique” ability to detect its own mistakes. This work systematically investigates how current LRMs recover from errors by inserting arithmetic mistakes in their intermediate reasoning steps. Notably, we discover a peculiar yet important phenomenon: despite the error propagating throughout the entire chain-of-thought (CoT) without any verbalized correction, the model still reaches the correct final answer after the thinking process finishes. This recovery implies the existence of an internal mechanism helping the model to detect errors and trigger self-correction, which we refer to as the hidden critique ability. Building on feature space analysis, we identify a highly interpretable critique vector representing this behavior. Extensive experiments across multiple model scales and families demonstrate that steering latent representations with this vector improves the model’s error detection capability and enhances the performance of test-time scaling at no extra training cost. Our findings provide a valuable understanding of LRMs’ critique behavior, suggesting a promising direction to control and improve their self-verification mechanism.
Abstract: Long-term memory is increasingly important for personalized AI agents, yet existing benchmarks and methods remain largely text-centric. Even when images are included, the user-specific information needed for later questions is typically recoverable from text alone, and most memory systems reduce image turns to generic captions. Yet images often carry personal information that text rarely states---both \emphexplicit evidence, such as recurring user-associated entities, and \emphimplicit evidence, such as latent user facts inferred from visual or multimodal cues. We introduce a benchmark for \emphpersonal visual memory that targets both forms of evidence, and propose \textscVisualMem, a hybrid visual--text architecture that augments a text-memory backend with a structured personal visual memory module. Rather than collapsing images into captions, \textscVisualMem uses conversational context to resolve identity, ownership, and durable user facts. Experiments show that \textscVisualMem substantially outperforms prior memory systems on our benchmark while remaining competitive on standard text-memory benchmarks, indicating that personal visual memory is a distinct and important component of long-term memory for personalized AI agents.
Abstract: Action chunking can improve credit assignment and exploration in reinforcement learning by shortening the effective decision horizon, but existing action chunking reinforcement learning methods often rely on chunked state-action value functions which become challenging to learn as action dimensionality grows. Moreover, executing action chunks open loop sacrifices within-chunk reactivity, which is especially problematic in contact-rich robotic control. We present Action Chunking PPO (ACPPO), an extension of PPO that introduces temporal abstraction through a chunked actor while retaining a standard state-value critic, thereby avoiding chunked Q-functions. We further propose ACPPO-Corr, which augments the chunk planner with a stepwise feedback corrector that adjusts planned actions online and provides reactivity within a chunk. Across 25 simulated robotics tasks from IsaacGym and Bi-DexHands, spanning locomotion, arm manipulation, and dexterous hand-object interaction, ACPPO-Corr achieves the strongest aggregated benchmark performance and remains best on both decision-frequency-sensitive and decision-frequency-neutral task subsets. Ablations show that moderate chunk lengths work best and that regularizing the corrector balances long-horizon planning and local feedback. These results suggest that action chunking can be made effective in fully online PPO when open-loop temporal structure is paired with closed-loop correction.
Authors:
Muhammad bin Javaid, Hasham Hussain, Ashima Khanna-Reiter, Berke Kisin, Jonathan Pirnay, Alexander Mitsos, Dominik G Grimm, Martin GroheAbstract: In structurally constrained molecular optimization, state-of-the-art methods restart an expensive oracle-driven search from scratch for every new input structure, scaling poorly to settings with many starting structures or expensive oracles. While amortized approaches that learn a transferable policy could in principle remove this bottleneck, existing methods struggle to generalize to diverse structural constraints at inference time. We present AMORTIX, an amortized Graph Transformer model that natively supports such constraints, optimizing molecular structures in a single forward pass with zero inference-time oracle calls. A central challenge for amortized training in this domain is that optimization difficulty varies drastically across starting structures. We show that, under this heterogeneity, standard reinforcement learning methods fail to stabilize training, and address this by normalizing rewards within groups of completions sharing the same starting structure. We evaluate on structurally constrained single- and multi-target kinase inhibitor design, and on a few-shot prodrug case study. AMORTIX outperforms both amortized and instance-optimization baselines on goal-directed scaffold decoration and ranks first among amortized methods on the PMO benchmark; the prodrug case study further demonstrates transfer of a learned modification rule to unseen drug structures.
Abstract: We propose a new framework for low-precision neural networks based on angular geometry and randomized low-bit estimators. Rather than approximating Euclidean linear operations under limited precision, we reinterpret neural computation in terms of directional similarity and construct quantized random-feature estimators of cosine interactions. Our approach introduces Angular Layers, which replace standard linear transformations with low-bit projections that estimate angular similarity in the forward pass, while gradients are computed with respect to the underlying continuous cosine geometry. This decoupling enables stable optimization without straight-through estimators or heuristic surrogate gradients, even under aggressive quantization of both weights and activations. On the theory side, we introduce \emphG-Gradient as a gradient proxy for angular layers and show that, in a one-layer planted model, a sufficiently small \emphG-gradient can guarantee exact recovery of a target binary solution. We further prove uniform approximation and convergence guarantees that provide a rigorous path from tractable continuous optimization to exact recovery in the original discrete model. Empirically, we instantiate these ideas in Angular Nets, including low-precision variants of ResNet, Vision Transformers, and BERT-style encoders, and obtain strong performance across vision and language tasks at substantially reduced precision. Overall, our results suggest that preserving angular structure provides a principled foundation for low-bit deep learning. Our full implementation is available at: \urlhttps://github.com/GGM2026/GGM.
Abstract: Generative video models have achieved remarkable visual fidelity and temporal coherence, yet intentional camera control remains elusive. Existing frameworks treat camera motion as a byproduct of pixel synthesis, producing trajectories that are stochastic, spatially inconsistent, and indifferent to the human subject driving the scene. In this work, we present Auteur, a method for language-driven, human-centric camera framing in generative video. Our core insight is that professional filmmakers conceive shots not as world-space trajectories but as framings defined relative to the actor, encoding shot size, angle, and composition as functions of human pose and motion. We formalize this intuition as a human-centric camera parameterization and introduce a Domain-Specific Language (DSL) that is convertible to standard 6-DoF camera parameters. A fine-tuned multimodal large language model then acts as a virtual director, mapping natural language descriptions and coarse human motion to sparse DSL keyframes that are deterministically inter-polated into continuous camera trajectories, which are then provided as input to video generators. We train and evaluate Auteur on a new dataset of 30K aligned text, human motion, and DSL-annotated camera trajectories drawn from procedural synthesis and real-world movie footage from the CondensedMovies dataset. Auteur enables cinematographic framing of human-centered scenes, a capability largely absent in prior generative models. To assess this behavior, we propose new framing-focused metrics, and our experiments show that Auteur consistently outperforms existing methods.
Abstract: Model reprogramming adapts a frozen source model to a target task by wrapping it in a learnable input transformation from target input to the source-model input space, and an output mapping from source-model outputs to target predictions. Despite extensive empirical progress, the finite-sample learnability of the induced wrapper hypothesis class \mathcalF\_\rm MR remains underexplored. We address this for source models that expose hard source labels and analyze through the wrapper structure: on any target samples, the input transformation selects reachable source labels that the output mapping relabels. Through this lens, we answer two questions: ① how many target samples are needed to approach the best predictor in \mathcalF\_\rm MR, and ② what gap remains between the optimum over \mathcalF\_\rm MR and the target Bayes risk. For ①, the wrapper structure factorizes the sample-wise complexity of \mathcalF\_\rm MR into a reachability term counting reachable source-label patterns, and a conditional relabeling term counting their relabelings. This yields an agnostic-PAC bound on the estimation error, with sample complexity scaling additively in the two terms; this bound becomes closed-form on additive input transformations and is matched by a lower bound on an explicit binary classification construction. For ②, the same structure decomposes the approximation error exactly into the source-label information loss and output-mapping restriction with a source-conditional upper bound on this gap. We extend the approximation analysis to logit-exposing source models, where the upper bound is empirically estimable. The wrapper-side analysis localizes both errors to specific components, giving a structural view of reprogramming practice.
Abstract: We present PhysFormer, a physics-grounded diffusion transformer for generating 4D multi-object mesh dynamics directly in world coordinates. Rather than predicting future frames in pixel space or rolling out next-step system states autoregressively, PhysFormer models physical evolution as full-trajectory coordinate diffusion. Given initial per-vertex positions, velocities, and material conditions, it generates future mesh vertex trajectories in a single denoising process, producing physically plausible object-object and object-environment interactions without hard-coded constraints, simulator priors, or learned shape latents. PhysFormer adopts a DiT-style backbone with factorized temporal, spatial, and object-level attention to capture coherent structure across time, vertices, and objects. Trained on over 100k collision-rich simulated trajectories, it models rigid and deformable multi-object dynamics, generalizes to unseen real-world geometries and larger object counts, and substantially outperforms autoregressive baselines in trajectory accuracy, rigidity preservation, and momentum-based physical consistency. Our results position coordinate-space diffusion as a promising step toward view-invariant, geometry-level world models for robotics, graphics, and physical design.
Abstract: Generating clinical reports that summarize abnormal patterns, diagnostic findings, and clinical interpretations from long-term EEG recordings remains labor-intensive. We present CELM, the first clinical EEG-to-Language foundation model capable of summarizing long-duration, variable-length EEG recordings and performing end-to-end clinical report generation at multiple scales. CELM integrates pretrained EEG foundation models with language models to enable scalable multimodal learning. We curate a large-scale clinical EEG dataset containing 9,922 reports paired with approximately 11,000 hours of EEG recordings from 9,048 patients to train CELM, and release the benchmark with an automated report-structuring pipeline to facilitate future research. Experimental results show that CELM consistently outperforms existing methods across all evaluation settings. Importantly, we further conduct human evaluation with clinical experts, demonstrating that CELM generates reports that are more clinically coherent, diagnostically reliable, and better aligned with expert interpretation. We release our model and benchmark construction pipeline at https://anonymous.4open.science/r/CELM-3AF4.
Abstract: Large language model agents rely on a memory harness to write, organize, retrieve, and use past experience. A harness that works for one task often fails on another, because conversations, embodied planning, and expert reasoning require different storage and retrieval behavior. To address this limitation, we introduce Mstar, a method that automatically discovers task-optimized memory programs through reflective code evolution. Specifically, Mstar operationalizes a memory harness as an executable Python memory program that defines the Schema, Logic, and Instruction, and optimizes these components jointly. We evaluate Mstar on four tasks covering conversation, embodied planning, and specialized reasoning. Our results demonstrate that Mstar improves performance over existing baselines across all evaluated tasks. Furthermore, the evolved memory programs exhibit structurally distinct processing mechanisms for each evaluated domain. These findings suggest that every task benefits from its own memory harness, and that memory-program search provides a concrete way to discover it.
Abstract: Despite impressive progress in capabilities of large vision-language models (LVLMs), these systems remain vulnerable to hallucinations, i.e., outputs that are not grounded in the visual input. Prior work has attributed hallucinations in LVLMs to factors such as limitations of the vision backbone or the dominance of the language component, yet the relative importance of these factors remains unclear. To resolve this ambiguity, We propose HalluScope, a benchmark to better understand the extent to which different factors induce hallucinations. Our analysis indicates that hallucinations largely stem from excessive reliance on textual priors and background knowledge, especially information introduced through textual instructions. To mitigate hallucinations induced by textual instruction priors, we propose HalluVL-DPO, a framework for fine-tuning off-the-shelf LVLMs towards more visually grounded responses. HalluVL-DPO leverages preference optimization using a curated training dataset that we construct, guiding the model to prefer grounded responses over hallucinated ones. We demonstrate that our optimized model effectively mitigates the targeted hallucination failure mode, while preserving or improving performance on other hallucination benchmarks and visual capability evaluations. To support reproducibility and further research, we will publicly release our evaluation benchmark, preference training dataset, and code.
Abstract: Compound AI systems promise capabilities beyond those of individual models, yet their success depends critically on effective orchestration. Existing routing approaches face two limitations: (1) input-level routers make coarse query-level decisions that ignore evolving task requirements; (2) RL-trained orchestrators are expensive to adapt and often suffer from routing collapse, repeatedly invoking one option in multi-turn scenarios. We introduce SkillOrchestra, a framework for skill-aware orchestration. Instead of directly learning a routing policy end-to-end, SkillOrchestra learns fine-grained skills from execution experience and models agent-specific competence and cost under those skills. At deployment, the orchestrator infers the skill demands of the current interaction and selects agents that best satisfy them under an explicit performance-cost trade-off. Extensive experiments across ten benchmarks demonstrate that SkillOrchestra outperforms SoTA RL-based orchestrators by up to 27.5% with 700x and 300x learning cost reduction compared to Router-R1 and ToolOrchestra, respectively. These results show that explicit skill modeling enables scalable, interpretable, and sample-efficient orchestration, offering a principled alternative to data-intensive RL-based approaches.
Abstract: Reinforcement learning from human feedback (RLHF) typically assumes a static or non-strategic reward model (RM). In iterative deployment, however, the policy generates the data on which the RM is retrained, creating a feedback loop. Building on the Stackelberg game formulation of this interaction, we derive an analytical decomposition of the policy's true optimization gradient into a standard policy gradient and a parameter-steering term that captures the policy's influence on the RM's future parameters. We show that standard iterative RLHF, which drops this steering term entirely, suffers from alignment collapse: the policy systematically exploits the RM's blind spots, producing low-quality, high-reward outputs whose feedback reinforces the very errors it exploits. To mitigate this, we propose foresighted policy optimization (FPO), a mechanism-design intervention that restores the missing steering term by regularizing the policy's parameter-steering effect on RM updates. We instantiate FPO via a scalable first-order approximation and demonstrate that it prevents alignment collapse on both controlled environments and an LLM alignment pipeline using Llama-3.2-1B.
Abstract: Single-cell trajectory inference from destructive time-course snapshots is fundamentally ill-posed: neither cross-time cell correspondences nor the continuous paths between snapshots are observed, so the observed snapshot distributions alone do not uniquely determine the underlying dynamics. Existing optimal transport and flow-based methods typically couple cells by Euclidean proximity at observed clock times, which can misalign trajectories when development is asynchronous and cells sampled at the same experimental time occupy different latent pseudotime stages. We propose PACE, a trajectory inference framework that selects geometry-consistent continuous transport dynamics from destructive time-course snapshots through three coupled components. First, PACE constructs a state- and time-dependent anisotropic Riemannian metric that preserves low cost along locally supported tangent directions while penalizing normal velocity components. Second, it alternates between refining cross-time couplings under the induced path-action cost and fitting endpoint-preserving neural bridges between adjacent snapshots. Third, it distills the learned bridge dynamics into a global continuous-time velocity field over cellular states. Across seven controlled and biological datasets covering nine held-out reconstruction experiments, PACE achieves the strongest overall reconstruction performance, reducing MMD, \mathcalW_1, and \mathcalW_2 by 23.7% on average relative to the strongest competing baseline. PACE also improves RNA-velocity alignment by 15.4% on an embryoid body differentiation benchmark, without requiring explicit cell pairing, lineage tracing, or RNA velocity supervision during training. Code is available at \urlhttps://anonymous.4open.science/r/PACE-F444/.
Abstract: Latent diffusion models (LDMs) enable high-fidelity synthesis by operating in learned latent spaces. However, training state-of-the-art LDMs requires complex staging: a tokenizer must be trained first, before the diffusion model can be trained in the frozen latent space. We propose UNITE – an autoencoder architecture for unified tokenization and latent diffusion. UNITE consists of a Generative Encoder that serves as both image tokenizer and latent generator via weight sharing. Our key insight is that tokenization and generation can be viewed as the same latent inference problem under different conditioning regimes: tokenization infers latents from fully observed images, whereas generation infers them from noise together with text or class conditioning. Motivated by this, we introduce a single-stage training procedure that jointly optimizes both tasks via two forward passes through the same Generative Encoder. The shared parameters enable gradients to jointly shape the latent space, encouraging a "common latent language". Across image and molecule modalities, UNITE achieves near state-of-the-art performance without adversarial losses or any pretrained encoders (e.g., DINO), reaching FID 2.12 and 1.75 for Base and Large models on ImageNet 256 x 256. Our experiments suggest that unification provides mutual benefits for tokenization and generation. We further analyze the Generative Encoder through the lenses of representation alignment and compression. These results show that single-stage joint training of tokenization and generation from scratch is feasible. Project Page: https://unite-tokenization-generation.netlify.app/
Authors:
Senqiao Yang, Chengyao Wang, Yuxin Chen, Zixuan WANG, Longxiang Tang, Haokun GUI, Jinhui Ye, Changsheng Lu, Xiaoyang Wu, Mingkang Zhu, Pengguang Chen, Shu Liu, Zhuotao Tian, Hengshuang Zhao, Bei Yu, Jiaya JiaAbstract: Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are fundamentally harder to scale than web-scale image-text data because they require embodied collection and sparsely cover the physical world. This makes representation quality a central bottleneck: under a fixed robot-data budget, VLA pre-training must convert limited trajectories into transferable visual-action knowledge, rather than merely fit actions. We propose , a VLA-oriented VLM backbone trained with a representation-centric pre-training recipe. VLAct preserves the broad VLM prior, avoids over-specializing the backbone to a single action head, and encourages shared action semantics across embodiments through VLM-prior preservation, multi-head continuous action co-supervision, and a partially unified cross-embodiment action layout, while leaving downstream users free to attach task-specific action heads during fine-tuning. Across multi-embodiment simulation benchmarks, real-world robot experiments, and unseen-embodiment transfer, VLAct consistently improves downstream performance under fixed fine-tuning protocols. On LIBERO-Plus and RoboTwin 2.0, VLAct surpasses large-scale industrial VLA systems such as ABot-M0 and LingBot-VLA, achieving 82.6% and 92.5% success, respectively. Most notably, on RoboCasa-GR1, a humanoid embodiment never seen during pre-training, VLAct with only 20% of downstream trajectories already outperforms the full-data GR00T-N1.6 baseline training setup, showing that representation-centric pre-training is an important independent axis of VLA progress beyond data scaling. All models and training pipelines will be open-sourced.
Authors:
Peng Xia, Jianwen Chen, Hanyang Wang, Jiaqi Liu, Kaide Zeng, Yu Wang, Siwei Han, Yiyang Zhou, Xujiang Zhao, Haifeng Chen, Zeyu Zheng, Cihang Xie, Huaxiu YaoAbstract: Large Language Model (LLM) agents have shown stunning results in complex tasks, yet they often operate in isolation, failing to learn from past experiences. Existing memory-based methods primarily store raw trajectories, which are often redundant and noise-heavy. This prevents agents from extracting high-level, reusable behavioral patterns that are essential for generalization. In this paper, we propose SkillRL, a framework that bridges the gap between raw experience and policy improvement through automatic skill discovery and recursive evolution. Our approach introduces an experience-based distillation mechanism to build a hierarchical skill library SkillBank, an adaptive retrieval strategy for general and task-specific heuristics, and a recursive evolution mechanism that allows the skill library to co-evolve with the agent's policy during reinforcement learning. These innovations significantly reduce the token footprint while enhancing reasoning utility. Experimental results on ALFWorld, WebShop and seven search-augmented tasks demonstrate that SkillRL achieves state-of-the-art performance, outperforming strong baselines over 15.3% and maintaining robustness as task complexity increases.
Abstract: Manipulating dynamic objects remains an open challenge for Vision-Language-Action (VLA) models. Although recent VLAs generalize well in static manipulation, dynamic scenes introduce a latency-induced perception–execution mismatch: object states continue to evolve during inference, making actions predicted from past observations stale at execution time. We present DynamicVLA, a latency-aware VLA model for dynamic object manipulation. It combines a compact 0.4B architecture and convolutional vision encoder for efficient multimodal inference with a continuous inference schedule that overlaps reasoning and execution for non-blocking control. Latent-aware Action Streaming then discards latency-invalid action prefixes and executes only the temporally valid suffix of each predicted chunk, preserving action-time alignment under dynamic object motion. To fill the missing foundation of dynamic manipulation data, we introduce the Dynamic Object Manipulation (DOM) benchmark, built with an automated collection pipeline that gathers 200K synthetic episodes across 2.8K scenes and 206 objects, and enables fast collection of 2K real-world episodes without teleoperation. Extensive evaluations in simulation and on real robots show that DynamicVLA improves dynamic manipulation success under changing object motion, perception-heavy instructions, and unseen motion patterns.
Abstract: A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety audit can support. We present two techniques that make audits in Petri, where an auditor model red-teams a target over multiple turns to elicit concerning behavior, harder to distinguish from real deployments. Our first technique, \emphcritique refinement, spends additional inference-time compute on each auditor action: an instance of the target model scores candidate actions for realism, provides feedback, and the auditor iterates before the highest-scoring candidate is selected. On Sonnet-4.6, realism win rate (the fraction of pairings in which the audit transcript is judged more realistic than a real deployment transcript) rises monotonically from 12% to 39% as refinement depth increases, and verbalized evaluation awareness drops to near zero; Haiku-4.5 and Opus-4.7 show a similar pattern. Our second technique, DISH (Deployment-Imitating SWE-Agent Harness), wraps the target in a Claude Code agent harness, reducing the gap between auditor-simulated and real deployment environments; in coding settings DISH raises realism win rate from 7% to 18% on Sonnet-4.6. The gains are additive on Sonnet-4.6 (12pp over either alone) but not on Opus-4.7.
Authors: Xinyu Zhang, Zhengtong Xu, Yutian Tao, Yeping Wang, Yu She, Abdeslam Boularias
Abstract: World models predict future transitions from observations and actions. Existing works predominantly focus on image generation only. Visual feature-based world models, on the other hand, predict future visual features instead of raw video pixels, offering a promising alternative that is more efficient and less prone to hallucination. However, current feature-based approaches rely on direct regression, which leads to blurry or collapsed predictions in complex interactions, while generative modeling in high-dimensional feature spaces still remains challenging. In this work, we discover that a new type of latent action representation, which we refer to as (RLA), can be easily learned from DINO residuals. We also show that RLA is predictive, generalizable, and encodes temporal progression. Building on RLA, we propose (RLA-WM), which predicts RLA values via flow matching. RLA-WM outperforms both state-of-the-art feature-based and video-diffusion world models on simulation and real-world datasets, while being orders of magnitude faster than video diffusion. Furthermore, we develop two robot learning techniques that use RLA-WM to improve policy learning. The first one is a minimalist world action model with RLA that learns from actionless demonstration videos. The second one is the first visual RL framework trained entirely inside a world model learned from offline videos only, using a video-aligned reward and no online interactions or handcrafted rewards.
Abstract: EEG foundation models can learn generalizable representations from large-scale EEG corpora to enable single-backbone transfer across diverse clinical and brain-computer interface tasks. Existing models typically discretize the continuous multi-channel EEG waveform into patches or codebook tokens and train a transformer with masked self-supervision. Recognizing that this discretization fragments continuous brain rhythms and obscures fine-grained temporal dynamics, we present B[FM]^2 (Brain Foundation Model via Flow Matching), which drops discretization and pretrains directly on the raw signal using continuous-time flow matching without patches, tokenization, or masking. However, multi-channel EEG signals pose an architectural challenge for flow matching: time is densely sampled and highly autocorrelated (thousands of timepoints), while the electrode axis is short (tens of channels) at distinct scalp positions. To address this time-electrode asymmetry, we introduce SplitUNet, a velocity network that factorizes each block into separate 1D temporal and 1D electrode convolutions and downsamples only along time, preserving electrode topology throughout the hierarchy. B[FM]^2 sets a new state of the art on 7 of 9 standard downstream EEG classification tasks, using a pretraining budget of only 36,895 segments (\approx 307\,h), a fraction (\approx 3.3%) of that required by existing EEG foundation models. It also produces synthetic EEGs that two board-certified neurologists cannot distinguish from real EEGs (Cohen's \kappa = -0.096).
Abstract: Linear attention replaces the unbounded cache of softmax attention with a fixed-size recurrent state, reducing sequence mixing to linear time and decoding to constant memory. The hard part is not just what to forget, but how to edit this compressed memory without scrambling existing associations. Delta-rule models subtract the current read before writing a new value, and Kimi Delta Attention (KDA) sharpens forgetting with channel-wise decay. But the active edit still uses a single scalar gate to control two different things: how much old content to erase on the key side and how much new content to commit on the value side. We introduce Gated DeltaNet-2, which generalizes both Gated DeltaNet and KDA by inheriting adaptive forgetting and channel-wise decay while addressing their shared limitation, the scalar tie between erasing and writing. Gated Delta Rule-2 separates these roles with a channel-wise erase gate \boldsymbolb_t and a channel-wise write gate \boldsymbolw_t, reducing to KDA when both gates collapse to the same scalar and to Gated DeltaNet when the decay also collapses. We derive a fast-weight update view, a chunkwise WY algorithm with channel-wise decay absorbed into asymmetric erase factors, and a gate-aware backward pass that preserves efficient parallel training. At 1.3B parameters trained on 100B FineWeb-Edu tokens, Gated DeltaNet-2 achieves the strongest overall results among Mamba-2, Gated DeltaNet, KDA, and Mamba-3 variants across language modeling, commonsense reasoning, and retrieval. Its advantage is most pronounced on long-context RULER needle-in-a-haystack benchmarks, where it improves the evaluated multi-key retrieval setting and remains strong in both recurrent and hybrid settings.
Abstract: Diffusion-based multimodal large language models (dMLLMs) decode by iteratively predicting tokens at multiple masked positions in parallel. This turns each decoding step into a position-selection problem: the model must choose not only which predictions are reliable in isolation, but also which positions should be committed together as context for later decoding steps. Existing confidence-based decoding ranks masked positions independently and commits the top-K positions, largely ignoring whether the committed tokens provide complementary visual grounding. We identify a step-level limitation of this strategy in multimodal settings: high-confidence tokens selected in the same step can rely on overlapping visual grounding, introducing visual redundancy among the committed tokens and leaving less complementary visual grounding available for later decoding. To quantify this effect, we introduce the Visual Redundancy Index (VRI), which measures visual grounding overlap among tokens committed in parallel. To control this redundancy during decoding, we propose Visual-Redundancy-Controlled Decoding (VRCD), a training-free inference-time decoding method that uses token-to-image attention to prioritize visually complementary positions. Across diverse multimodal benchmarks, VRCD reduces visual redundancy and remaining-position entropy with modest runtime overhead. In longer decoding experiments, it also achieves relative accuracy gains of up to 18.8% on M^3CoT and 6.9% on MMBench over confidence-based decoding.Code is included in the supplement.
Authors:
Zhangyi Wang, Jiaxu Liu, Chen Song, Chao Xu, Shengze CaiAbstract: In many safety-critical applications, control of uncertain dynamical systems relies on observers that estimate states and external disturbances. Neural network observers can improve estimation accuracy, but certifying their Lyapunov stability via Linear Matrix Inequality (LMI) constraints leads to large-scale semidefinite programs (SDPs) that are difficult to solve for large networks. To overcome this scalability bottleneck, we propose a novel two-stage training framework for provably stable neural network observers. Our approach decouples the optimization into a point-guided Lyapunov pre-training phase, which rapidly achieves high estimation accuracy and local stability over sampled states, followed by an LMI fine-tuning phase that efficiently satisfies a strict global Lyapunov stability certificate. We provide formal theoretical guarantees for local stability radii and probabilistic coverage over a prescribed compact error-state domain under specified regularity and sampling assumptions. Experiments on nonlinear control benchmarks and X-29 aircraft ablations show that our LMI-certified neural network observers train significantly faster than direct LMI-based methods and generalize robustly across diverse systems, achieving improved tracking accuracy over a range of observer baselines. The code is available at https://anonymous.4open.science/r/LearningNeuralNetworkObserver-2E1E.
Authors:
Mahdi Farahbakhsh, Vishnu Teja Kunde, Dileep Kalathil, Krishna Narayanan, JF ChamberlandAbstract: Diffusion models have been used as priors for solving inverse problems. However, existing approaches typically overlook side information that could significantly improve reconstruction quality, especially in severely ill-posed settings. In this work, we propose a novel framework that incorporates side information into existing diffusion-based inverse problem solvers via inference-time search, in a plug-and-play, training-free manner. Through extensive experiments across a range of inverse problems, including inpainting, super-resolution, and several deblurring tasks, and across multiple diffusion-based inverse problem solvers (DPS, DAPS, and MPGD), we show that augmenting each solver with our framework consistently improves the quality of the reconstructions over the corresponding original method. To demonstrate the generality of our approach, we consider diverse forms of side information, including reference images, textual descriptions, and anatomical MRI scans. Code is available \hrefhttps://anonymous.4open.science/r/sideinfo-search-reconstruction-4A32/README.mdhere.
Abstract: LLM-based agents are increasingly deployed in real-world applications through tool-use APIs, yet training them for specific environments remains fundamentally difficult: real-world applications provide no pre-defined tasks or verifiers, no faithful simulators, and limited budget for large-scale environment interaction. In this paper, we propose AgentBrew, an offline training framework that learns effective tool-use policies from a single batch of raw interaction trajectories, without task verifiers or iterative on-policy rollouts. The agent first explores the target environment to collect a raw trajectory corpus without quality filtering. To extract training signal from this noisy corpus, retrospective task inference reconstructs an aligned instruction for each trajectory based on its actual outcome, and PMI-Based credit assignment decomposes the trajectory's total information about the inferred instruction into additive per-action credits via pointwise mutual information (PMI). These credits weight the policy training objective, amplifying informative actions while suppressing ineffective ones. On three real-world MCP applications (GitHub, Notion, PostgreSQL), AgentBrew improves Qwen3-32B by +8.7 Acc / +9.7 Score on average, surpassing Qwen3-235B (+2.3 / +4.4) and outperforming rejection sampling (+5.9 / +10.3). These results demonstrate that fine-grained offline learning can recover useful supervision from raw trajectories that filtering-based approaches would discard.
Abstract: High-dynamic-range (HDR) formats and displays are becoming increasingly prevalent, yet state-of-the-art image generators (e.g., Stable Diffusion and FLUX) typically remain limited to low-dynamic-range (LDR) output due to the lack of large-scale HDR training data. In this work, we show that existing pretrained diffusion models can be easily adapted to HDR generation without retraining from scratch. A key challenge is that HDR images are natively represented in linear RGB, whose intensity and color statistics differ substantially from those of sRGB-encoded LDR images. This gap, however, can be effectively bridged by converting HDR inputs into perceptually uniform encodings (e.g., using PU21 or PQ). Empirically, we find that LDR-pretrained variational autoencoders (VAEs) reconstruct PU21-encoded HDR inputs with fidelity comparable to LDR data, whereas linear RGB inputs cause severe degradations. Motivated by this finding, we describe an efficient adaptation strategy that freezes the VAE and finetunes only the denoiser via low-rank adaptation in a perceptually uniform space. This results in a unified computational method that supports both text-to-HDR synthesis and single-image RAW-to-HDR reconstruction. Experiments demonstrate that our perceptually encoded adaptation consistently improves perceptual fidelity, text-image alignment, and effective dynamic range, relative to previous techniques.
Abstract: Probabilistic time series forecasting requires modeling and predicting complex and time-varying distributions. Recently, Denoising Diffusion Probabilistic Model~(DDPM)-based approaches have shown promise by equipping the diffusion process with pretrained mean and variance estimators to accommodate distributional shift. However, these methods typically follow the standard DDPM framework and consider only partial components of the evidence lower bound (ELBO), treating the training of estimators as designed regression tasks separate from the variational inference framework. To address this, we rethink the ELBO under the Location-Scale Noise Model (LSNM) and find that it naturally induces a Gaussian negative log likelihood objective for the estimators and inherently defines a joint training objective that unifies recent the diffusion paradigms for probabilistic forecasting. Building on this principled ELBO reformulation, we propose DiffPTS, a general framework that enables end-to-end optimization of all components within the ELBO. Across multiple benchmarks, DiffPTS consistently outperforms recent models, achieving state-of-the-art performance with an average CRPS/MSE reduction of over 14.53%/16.55% compared to existing diffusion-based methods. The code is provided at \urlhttps://anonymous.4open.science/r/DiffPTS.
Abstract: Conditional generators provide a natural tool for controllable generation, including settings where the desired condition is a new composition of observed attributes or experimental factors. In many applications, especially in scientific domains, such models are attractive to explore conditions for which real samples are rare, expensive, or not yet observed. However, this creates a circularity for evaluation: standard conditional quality metrics require a reference target distribution, but in the extrapolative regime that distribution is unavailable by definition. We address this problem with a post-hoc, per-sample trust score for assessing conditional samples using only the training distribution. The score combines two estimable quantities: global realism, measuring compatibility with the real data manifold, and attribute-wise faithfulness, measuring whether a sample is closer to the requested attributes than to plausible alternatives. We show that the score can recover meaningful comparisons across extrapolated generations, under a mild coverage condition on the observed attributes. These comparisons enable effective filtering, ranking, and abstention of generations and can be used directly on off-the-shelf pretrained models. In biological imaging, selected samples preserve real morphological structure better and improve downstream predictive performance, while similar gains are observed on controlled vision benchmarks. Finally, we show how the score can be applied during generation, enabling abstention before full decoding.
Abstract: Would experience designing faster GPU kernels also help close in on a long-standing open mathematical conjecture? Large Language Models (LLMs) integrated into evolutionary search have recently produced state-of-the-art solutions on optimization tasks, including open mathematical conjectures, GPU kernel design, scientific law discovery, and combinatorial puzzles. To achieve this, prior work applied search scaffolds to one target task at a time, so every new problem is approached from scratch and the experience accumulated during search is discarded once the model finishes its attempt. This leaves the capability of iteratively evolving a solution (e.g., knowing which part to mutate and how, deciding when to backtrack) entirely in the scaffold rather than in the model itself. Whether the model itself could acquire this capability and reuse it across different tasks has been largely unexamined. To address this, we introduce \textscEvolution Fine-Tuning (EFT), a mid-training paradigm that teaches LLMs to evolve solutions across tasks by converting evolutionary search trajectories into supervision. We construct \datasetName, a 156K-trajectory dataset spanning 10 domains and 371 optimization tasks, and fine-tune open-source LLMs from 2B to 9B parameters. Empirically, EFT confers cross-task generalization: across 22 held-out tasks, our models surpass their base counterparts by 10.22% on average. Furthermore, when paired with test-time RL, our model match state-of-the-art performance on two circle-packing tasks and outperforms its base-model counterparts on the Erdős minimum-overlap problem. EFT thus serves as a ``practice phase'' for general-purpose discovery agents that doesn't solve new problems from scratch.
Abstract: For graph node classification, calibrated class probabilities are need when confidence scores, usually the maximum predicted class probability, are used to rank prediction, defer uncertain nodes to human review or control risk. Existing post-hoc calibrators either apply one global temperature or use graph-awared modules without a derived strctural form. We study node-level calibration for graph node-classification problem under structural heterogeneity. Our first results show that a logit-only temperature rule is insufficient when nodes with identical logits but different local homophily require different optimal temperatures. We then analyze a population-concentration contextual stochastic block model with Gaussian features and a one-layer linear GCN. Under equidistant class means, the Bayes posterior over class-template scores is a temperature-scaled softmax whose inverse temperature is governed by homophily-dependent signal strength. In the positive-signal homophilic regime, the resulting temperature decreases approximately with normalized local homophily. This law motivates Homophily-aware Temperature Scaling (HoTS), a simple post-hoc calibrator that assigns each node a positive scalar temperature from entropy-based logit concentration and estimated local homophily. HoTS has three parameters, preserving class, and learn the strength of the structural correction from calibraion data. Across 18 node-classification benchmarks, two GNN backbones, and eight calibration baselines, HoTS achieves the best mean Expected Calibration Error (ECE) of 4.66%, the best average rank, and the strongest confidence ranking in selective classification.
Authors: Mohammad Saeid, Amir Salarpour, Pedram MohajerAnsari, Mert D Pesé
Abstract: We study compression of serialized 3D point transformers for efficient point cloud segmentation. Starting from LitePT-S, we derive TrimPT, a compact student that reduces channel width and stage-3 attention depth while preserving the full 1024-point attention window. TrimPT uses 5.84\,M parameters and 12.95\,GFLOPs, giving 2.18× fewer parameters and 1.96× fewer FLOPs than LitePT-S. To improve the compressed student, we introduce Stage-wise Relational Feature Distillation (SRFD), a training-only objective that matches pairwise cosine-similarity matrices between teacher and student features at the compressed attention stages. This explicitly regularizes teacher--student affinity mismatch and adds no inference-time cost because the teacher and projection heads are discarded after training. On ScanNet semantic segmentation, the resulting TopoPT reaches 76.6% mIoU with a 5.84\,M parameter inference footprint, improving over TrimPT without SRFD and matching the official LitePT-S result with substantially fewer inference-time parameters and FLOPs. TopoPT also obtains 63.9% mAP_50 on ScanNet instance segmentation, 33.0% mAP_50 on ScanNet200, and 81.4% mIoU on nuScenes, suggesting that stage-wise relational distillation is a useful training-time regularizer for lightweight 3D segmentation backbones. Code and models are available at: \urlhttps://github.com/anon-push/TopoPT
Abstract: LLM routing is designed to avoid overpaying for inference: easy queries should be handled by cheaper models, while stronger models should be reserved for cases where they are truly needed. However, we show that existing routers often violate this principle under loose cost budgets. As the budget increases, they increasingly route queries to the strongest and most expensive model, even when cheaper candidates achieve comparable or identical outcomes. We refer to this behavior as strong-model over-selection under loose budgets. Across both unimodal and multimodal routing benchmarks, this behavior weakens the cost-saving motivation of routing without necessarily improving quality. We diagnose this phenomenon through the mismatch between common router training objectives and deployment-time routing decisions. Many routers are trained to predict scalar performance scores, whereas cost-aware routing ultimately depends on query-specific comparisons among budget-feasible models, with cost used to distinguish equivalent or near-equivalent choices. In small-margin regimes, modest prediction errors can therefore flip relative orderings and induce cost-insensitive selections. Motivated by this diagnosis, we propose EquiRouter, a decision-aligned router that directly learns query-dependent model rankings with a Cost-Aware Ranking Objective and a lightweight Query--Model Interaction Representation. Experiments across RouterBench, MMR-Bench, MixInstruct, and RouterEval show that EquiRouter reaches the strongest model-level performance with lower relative cost than compared baselines, demonstrating consistent improvements across text-only, multimodal, continuous-utility, and large-model-pool routing settings.
Abstract: All prior membership inference attacks for fine-tuned language models use hand-crafted heuristics (e.g., loss thresholding, Min-K%, reference calibration), each bounded by the designer's intuition. We introduce the first transferable learned attack, enabled by the observation that fine-tuning any model on any corpus yields unlimited labeled data, since membership is known by construction. This removes the shadow model bottleneck and brings membership inference into the deep learning era: learning what matters rather than designing it, with generalization through training diversity and scale. We discover that fine-tuning language models produces an invariant signature of memorization detectable across architectural families and data domains. We train a membership inference classifier exclusively on transformer-based models. It transfers zero-shot to Mamba (state-space), RWKV-4 (linear attention), and RecurrentGemma (gated recurrence), achieving 0.963, 0.972, and 0.936 AUC respectively. Each evaluation combines an architecture and dataset never seen during training, yet all three exceed performance on held-out transformers (0.908 AUC). These four families share no computational mechanisms, their only commonality is gradient descent on cross-entropy loss. Even simple likelihood-based methods exhibit strong transfer, confirming the signature exists independently of the detection method. Our method, Learned Transfer MIA (LT-MIA), captures this signal most effectively by reframing membership inference as sequence classification over per-token distributional statistics. On transformers, LT-MIA achieves 2.8× higher true positive rate at 0.1% false positive rate than the strongest baseline. The method also transfers to code (0.865 AUC) despite training only on natural language texts. Our results imply that leakage from memorization is intrinsic to cross-entropy training; architectural innovation within this paradigm did not escape our attack. Code and the trained classifier are provided.
Abstract: Generating novel renderings of a scene along user-defined camera trajectories from a single monocular video, dubbed video retaking, is a compelling but difficult problem in content creation and visual effects. Existing geometry-guided approaches reconstruct a 4D representation from the source video and render it along the target trajectory to condition video diffusion models. However, this guidance degrades as the target camera departs from the source trajectory, leaving newly revealed regions sparse or entirely missing. We propose SierpinskiCam, which addresses this limitation by augmenting geometry-based guidance with Sierpinski dome texture cues that contains rich trackable features even under large viewpoint changes. We further introduce a reference video conditioning mechanism that appends source-video tokens to the target-token sequence and separates the two streams with negative RoPE indices, enabling appearance grounding without architectural modification or per-video adaptation. Extensive experiments show that SierpinskiCam achieves significant gains in camera controllability, geometric consistency, and video quality across diverse and challenging retaking scenarios.
Abstract: The advent of Large Language Models (LLMs) has fundamentally reshaped the way we interact with graphs, giving rise to a new paradigm called GraphLLM. As revealed in recent studies, graph learning can benefit from LLMs. However, we observe limited benefits when we directly utilize LLMs to make predictions for graph-related tasks within GraphLLM paradigm, which even yields suboptimal results compared to conventional GNN-based approaches. Through in-depth analysis, we find this failure can be attributed to LLMs' limited capability for processing graph data and their tendency to overlook graph information. To address this issue, we propose ontrast), a novel plug-and-play method for GraphLLM paradigm, which enhances LLM's understanding of graph data through three stages: (1) : rectifying the vanilla logits produced in the decoding process. Extensive experiments demonstrate that LoReC brings notable improvements over current GraphLLM methods and outperforms GNN-based approaches across diverse datasets. The implementation is available at https://anonymous.4open.science/r/LoReC-63F4.
Authors:
Ruiqi Liu, Manni Cui, Ziheng Qin, Zhiyuan Yan, Ruoxin Chen, Yi Han, Zhiheng Li, Junkai Chen, ZhiJin Chen, Kaiqing Lin, Jialiang Shen, Lubin Weng, Yan Wang, Shu WuAbstract: High-fidelity generative models have narrowed the perceptual gap between synthetic and real images, putting media security under increasing pressure. Most AI-generated image (AIGI) detectors classify by artifact and fail to generalise to new generators. We argue that anchoring detection on the stationary real-image distribution gives a more generalisable signal than chasing the non-stationary forgery distribution, and cast AIGI detection as a Reference-Comparison problem that checks consistency against a learned real-domain feature subspace rather than fitting specific forgery cues. We propose ), which encodes real-image priors in a learnable discrete memory bank, projects an input onto these prototypes via sparse linear combination to form a manifold-consistent ideal reference, and uses the Comparison Residual and Reconstruction Perplexity as detection signals. To mark the Superhuman Crossover where detectors surpass human experts, we build the Human-AIGI benchmark with a psychophysically curated human-imperceptible subset. Across 14 benchmarks, MIRROR improves over prior methods by 2.1% on six standard benchmarks and 8.1% on seven in-the-wild benchmarks. On Human-AIGI it reaches 89.6% fake accuracy over 27 generators, sitting between the computer-vision-expert and AIGI-detection-expert reference levels, and the gap to the AIGI-detection-expert level closes further as the pretrained backbone scales.
Abstract: Building agentic systems that can autonomously self-improve from experience is a longstanding goal of AI. Large language models (LLMs) today primarily self-improve via two mechanisms: self-reflection for context updates, and reinforcement learning (RL) for weight updates. In this work, we propose Evolutionary System Prompt Learning (E-SPL), a method for jointly improving model contexts and model weights. In each RL iteration, E-SPL samples trajectories under multiple system prompts in parallel, then jointly applies RL updates to weights and evolutionary updates to system prompts via LLM self-reflection. E-SPL encourages a natural division between declarative knowledge encoded in prompts and procedural knowledge encoded in weights. Most notably, in an easy-to-hard generalization setting (AIME \rightarrow BeyondAIME), E-SPL improves RL success rate from 38.8% \rightarrow 45.1%. E-SPL also improves RL on AIME 2025 (56.3% \rightarrow 60.6%), HMMT 2025 (50.0% \rightarrow 52.7%), and agentic search (44.2% \rightarrow 48.6% on gpt-oss-120b). Across all settings, E-SPL outperforms both RL-only and evolution-only baselines, demonstrating that weight updates and context updates are deeply synergistic and can together yield gains in generalization that neither achieves alone.
Abstract: Inference-time harnesses substantially improve large language models on complex reasoning tasks. However, the intrinsic capabilities of the underlying model remain unchanged by the addition of these external workflows. To bridge this gap, we introduce \emphOn-Policy Harness Self-Distillation (OPHSD), which employs the harness-augmented current model as a teacher for self-distillation, thereby introducing extra supervisory signals from the harness beyond training data. OPHSD internalizes task-specific harness capabilities into the student model, yielding robust generalizability and strong standalone performance across diverse reasoning tasks. Evaluated across draft--verify harness for text classification and plan--solve for mathematical reasoning tasks, OPHSD consistently outperforms strong baselines (e.g., +10.83% over OPSD on HMMT25). Our analysis further indicates that reattaching the harness during inference yields no additional benefits and can even degrade performance, suggesting that complex harnesses need not always be permanent fixtures; instead, they can serve as temporary training scaffolds whose benefits are permanently fed back into the base model. Our code and training data are available at \urlhttps://anonymous.4open.science/r/OPHSD-On-Policy-Harness-Self-Distillation-99F7.
Abstract: Figure-ground organization in the human visual system relies on several shape-based cues, including surroundedness, convexity, and symmetry. While these cues have been extensively studied using abstract stimuli, little is known about how they operate under natural conditions or how they arise from the statistics of natural scenes. Deep neural networks offer a promising path forward: a model that relies on the same figure-ground cues as humans would provide tractable experimental access to the underlying mechanisms. In this study, we evaluate shape-based figure-ground organization in Vision Transformers (ViTs), for which prior work has demonstrated the emergence of object-based grouping. We test 25 ViTs spanning supervised and self-supervised training objectives, by fitting linear probes to predict figure-ground assignment from intermediate patch representations using both natural images and controlled artificial stimuli that isolate individual cues. Our results show that ViTs robustly encode surroundedness and convexity, and that probes trained on natural images generalize zero-shot to artificial stimuli across several models. For symmetry we observe mixed results: the cue is encoded for uniformly colored but not for textured regions. Taken together, our findings demonstrate that Gestalt-like figure-ground cues can be learned from natural scene statistics and position ViTs as a compelling model system for studying the computational mechanisms of perceptual organization. All code and data will be published.
Authors:
Zihan Zhao, Aaron Wang, Alan Xia, Chang Sun, Javier Duarte, Abhijith Gandrakota, Jennifer Ngadiuba, Richard CavanaughAbstract: Real-time jet tagging is critical for identifying short-lived particle decays in the high-throughput detectors of the Large Hadron Collider, where real-time trigger systems which are responsible for deciding which collision events to store impose strict latency and accuracy constraints. Transformer architectures achieve the highest jet tagging accuracy when compute is unconstrained, but their quadratic self-attention cost places them orders of magnitude beyond the trigger budget. Existing efficient variants reduce this cost by compressing the attention matrix or restricting it to ordered local windows, at the price of the explicit particle-particle interactions that drive substructure identification. To address this limitation, we introduce the Patch Hierarchical Attention Transformer (PHAT-JeT), which combines two mechanisms: a physics-inspired geometric message-passing module that encodes local detector-plane structure, and a hierarchical patch-based attention scheme that computes exact attention within small particle groups while preserving global context through lightweight patch-token communication. Within this compute budget, PHAT-JeT achieves state-of-the-art accuracy and background rejection among resource-constrained jet tagging models on four benchmarks (hls4ml, JetClass, Top Tagging, and Quark-Gluon). Our code is available at https://anonymous.4open.science/r/PHAT-JeT-540B/README.md
Abstract: Building humanoid robots that perform generalizable whole-body loco-manipulation in the real world remains a fundamental challenge: existing approaches either rely on heavy task-specific reward engineering, rigidly replay reference motions that fail to generalize, or depend on costly teleoperation that limits scalability. While human videos capture diverse human behaviors, the motion priors inferred from them are inherently imperfect, suffering from occlusion, contact artifacts, and retargeting errors that render them unsuitable for direct policy learning. To this end, we present SUGAR, a data-driven framework that converts diverse human videos into deployable humanoid loco-manipulation skills, without any task-specific reward engineering or reference-motion conditioning at inference. SUGAR proceeds in three coupled stages: First, a fully automated pipeline extracts scalable kinematic interaction priors including human-object motion trajectories and contact labels from diverse human videos. Second, a privileged physics-based refiner utilizes a unified mimic-style reward and a progressive state pool to transform imperfect kinematic interaction priors into physically feasible, high-fidelity skills. Third, the refined skills are distilled into a hierarchical policy that comprises a task-guided planner and a command tracker for autonomous task execution. We evaluate our method on six representative loco-manipulation tasks in both simulation and real-world humanoid hardware. SUGAR substantially outperforms reference-tracking baselines, and its performance scales clearly with the amount of human video data. It also achieves zero-shot real-world transfer with reliable closed-loop execution, autonomous failure recovery and stable long-horizon performance under external perturbations. Project Page: https://sugar-humanoid.github.io/
Abstract: The design of modern neural architectures has converged through incremental empirical choices, yet the mechanisms governing their training dynamics remain only partially understood. We identify and analyze a negative weight drift induced by the interaction between standard losses and positively biased activation functions. We prove that under MSE or cross-entropy loss, the gradient with respect to positive pre-activations is non-negative in expectation at initialization, driving downstream weights toward negative values during early training. The drift is intrinsic to optimization rather than data, and persists across architectures (MLP, ResNet, ViT, GPT-nano, MP-SENe) and asymmetric activation functions (ReLU, GELU, SiLU). Coupled with ReLU, weight drift produces activation sparsity reaching up to 90% in GPT-nano. We characterize the sparsity–accuracy tradeoff across 79 configurations and identify a sharp accuracy cliff above ~70% activation sparsity. While ReLU^2 achieves a good sparsity–accuracy tradeoff, it pathologically amplifies identified activation spikes in intermediate transformer layers. Clipping resolves this while preserving the representational benefits of squaring: clipped ReLU^2 outperforms its unclipped version, and GELU^2 achieves the lowest validation loss on GPT-nano.
Abstract: While large language model (LLM) multi-agent systems achieve superior reasoning performance through iterative debate, practical deployment is limited by their high computational cost and error propagation. This paper proposes AgentArk, a novel framework to distill multi-agent dynamics into the weights of a single model, effectively transforming explicit test-time interactions into implicit model capabilities. This equips a single agent with the intelligence of multi-agent systems while remaining computationally efficient. Specifically, we investigate three hierarchical distillation strategies across various models, tasks, scaling, and scenarios: reasoning-enhanced fine-tuning; trajectory-based augmentation; and process-aware distillation. By shifting the burden of computation from inference to training, the distilled models preserve the efficiency of one agent while exhibiting strong reasoning and self-correction performance of multiple agents. They further demonstrate enhanced robustness and generalization across diverse reasoning tasks. We hope this work can shed light on future research on efficient and robust multi-agent development. Our code is at https://anonymous.4open.science/r/AgentArk.
Abstract: Self-supervised pre-training via cross-view completion learns strong features for 3D vision from co-visible regions of image pairs. However, the reference view provides little information for reconstructing non-co-visible patches, implicitly yielding a monocular training signal in these regions. We introduce of the cross-view reconstruction error over a masked-autoencoder error serves as a self-supervised proxy for co-visibility: large improvements indicate co-visible regions, while negligible ones indicate non-co-visible areas. Gekko is a network, trained from scratch, that jointly performs cross-view completion, masked autoencoding, and per-pixel prediction of this relative improvement, providing an additional binocular signal for masked regions without requiring any ground-truth 3D annotations. Under identical architectures and training data, Gekko consistently outperforms CroCo on zero-shot correspondence estimation, relative pose estimation, and pointmap regression. We further show that Gekko can be trained directly from raw videos using a simple stride-based curriculum, eliminating the cumbersome 3D preprocessing required by prior methods while matching the performance of models trained on curated data, thereby enabling fully self-supervised pre-training. Code and pre-trained models will be released.
Abstract: World Action Models (WAMs) have emerged as a promising alternative to Vision-Language-Action (VLA) models for embodied control because they explicitly model how visual observations may evolve under action. Most existing WAMs follow an imagine-then-execute paradigm, incurring substantial test-time latency from iterative video denoising, yet it remains unclear whether explicit future imagination is actually necessary for strong action performance. In this paper, we ask whether WAMs need explicit future imagination at test time, or whether their benefit comes primarily from video modeling during training. We disentangle the role of video modeling during training from explicit future generation during inference by proposing Fast-WAM, a WAM architecture that retains video co-training during training but skips future prediction at test time. We further instantiate several Fast-WAM variants to enable a controlled comparison of these two factors. Across these variants, we find that Fast-WAM remains competitive with imagine-then-execute variants, while removing video co-training causes a much larger performance drop. Empirically, Fast-WAM achieves competitive results with state-of-the-art methods both on simulation benchmarks (LIBERO and RoboTwin) and real-world tasks, without embodied pretraining. It runs in real time with 110 ms latency, over 6× faster than existing imagine-then-execute WAMs. These results suggest that the main value of video prediction in WAMs may lie in improving world representations during training rather than generating future observations at test time.
Abstract: Viewpoint dictates the photo composition and plays a vital role in AI photography. Existing 3D viewpoint recommendation methods are often limited in viewpoint adjustments, which is mainly due to their lack of high-quality training data with accurate and large 3D view changes. Ideally, the training data should contain suboptimal images and expert-selected optimal viewpoints, but this can be very expensive to curate. To solve this, we build an automatically-generated 3D viewpoint recommendation dataset from expert photos. Specifically, we first consider these expert photos as optimal and outpaint them to provide more information on scene layouts. Next, we explicitly reconstruct the 3D scene. We then render suboptimal images from random viewpoints. Lastly, we introduce a hierarchical data filtering scheme and acquire high-quality data pairs for our ViewRecDB-100K dataset. On top of this, we further introduce a view recommendation model, ViewRecNet, that can predict expert viewpoints from any suboptimal image inputs using a ray-based dense supervision. The proposed method achieves strong results both quantitatively and qualitatively, supported by LLM-based viewpoint preference and human studies. Data and code are available at https://github.com/anonymous1-submission/ViewRec3D.git.
Abstract: Classical scaling laws for language model pretraining balance model size against training dataset size under a fixed compute budget, assuming abundant data and a single pass over the corpus. As training compute grows faster than the supply of natural language data, pretraining is likely to enter a data-constrained, compute-rich regime where models train for multiple epochs over a finite dataset. We study data-constrained pretraining along two axes, regularization and scaling. For regularization, we study masked-input regularization (MIR), an auxiliary next-token prediction loss on randomly masked inputs. MIR tests whether the random masking central to diffusion language models can benefit autoregressive pretraining without architectural changes or inference overhead. Across 72M to 1.4B parameter models, we find that MIR added on top of strong weight decay improves validation loss over autoregressive strong-weight-decay-only models, with downstream gains at 1.4B. For scaling, we propose SoftQ, a scaling law that couples model size and data size to capture their interaction under repeated data. Classical alternatives such as the Chinchilla law use an additive form that decouples these terms, making them misspecified in the data-constrained regime. We find that SoftQ fits data-constrained experiments substantially better than these alternatives, and estimates MIR's gains as equivalent to roughly 1.3 times as much unique training data.
Authors: Rasul Khanbayov, Erchin Serpedin, HASAN KURBAN
Abstract: Prototype-based medical image classifiers have three clinical gaps: they treat findings as independent, silently amplify unsafe doctor feedback, and require full retraining whenever a new finding is needed. We present GRAPE (Graph-Augmented Prototype Explanations), a unified architecture that closes all three gaps. A Graph Attention Task Head models anatomical concept co-occurrence, boosting macro-F1 by +13.8\,pp over the prototype baseline on TBX11K. A Concept-Mismatch Safety Check is the first such mechanism in prototype-based medical classifiers, warns when the model's dominant finding inside a doctor-drawn region conflicts with the claimed label, catching 85% of erroneous annotations versus 51% for MC-Dropout with no extra inference cost. Open-Vocabulary Prototype Anchoring aligns visual prototypes to clinical text, so a new finding can be added from a single labelled image without modifying any other component: on NIH ChestX-ray14, one Effusion example recovers full-supervision localisation accuracy; on TBX11K, prototype maps achieve 2.6× better lesion localisation than end-to-end baselines. All three capabilities add only +1~ms latency at interactive batch size.
Authors: James C Mathews, Francisco Couzo, Aleksandr Petrov, Avijit Chatterjee, Saad Nadeem
Abstract: In this paper, we revisit the concept of entropy in the spatial context with the aim of deriving computable and interpretable metrics for point pattern analysis in domains such as histopathology. We discuss stochastic point process assumptions like Poisson homogeneity and Simple Sequential Inhibition (SSI) and review established results for Voronoi and Delaunay tilings of point sets for their potential to provide null distributions for rigorous empirical hypothesis testing. We present (1) a novel ``dense basin entropy'' defined in terms of the Ord distribution supported on Voronoi basins, which is shown to be sensitive to clustering, and (2) a related Markov chain designed as a simplification or approximation of Brownian motion in the underlying planar domain. We prove a lower bound on the dense basin entropy in the SSI regime (originally discovered in empirical experiments) and show that the entropy rate and spectral gap of the Markov chain lead to sharp discrimination metrics. Finally, we demonstrate the generalization to multiple-set analysis via aggregation of one set over the Markov invariant measure for another. Simulations and experiments in histopathology show that our proposed metrics offer insights complementary to other instantiations of entropy in spatial analysis.
Authors:
Yimeng Zhang, Yingying Zhuang, Ziyi Wang, Yuxuan Lu, Pei Chen, Aman Gupta, Zhe Su, Ming Tan, Zhilin Zhang, Qun Liu, Manikandarajan Ramanathan, Rajashekar Maragoud, Edward Vul, Jing Huang, Dakuo WangAbstract: Uncertainty estimation is essential not only for the trustworthy deployment of large language models (LLMs) but also as a foundation for self-refinement in LLM generation. However, existing approaches operate at suboptimal granularities: token-level scores lack semantic coherence, while sequence-level scores fail to localize errors. We formalize Span-Level Uncertainty Estimation (SLUE), a new task that targets the natural granularity for uncertainty: semantically coherent text spans, each conveying a single assessable unit of meaning. To address this task, we introduce SPANUQ, a lightweight (~25M parameter) probe that distills the uncertainty knowledge from expensive multi-sample inference into a single forward pass over LLM hidden states. SPANUQ employs a DETR-style span decoder to simultaneously detect spans and estimate their uncertainty via a Mixture of Beta distribution, trained with a principled combination of Beta NLL regression and contrastive ranking objectives. We construct SPANUQ-BENCH, the first span-level uncertainty benchmark comprising 20K prompts, ~293K annotated spans, and continuous soft labels derived from multi-sample claim verification. Experiments on five LLM backbones show that SPANUQ consistently achieves the best span-level uncertainty quality (AUROC 0.908–0.944, MAE 0.110–0.129), outperforming the strongest probe baseline and all sampling-based methods while being 10--20× faster. Its DETR-based span detector attains 0.910 F1, surpassing the best heuristic by 39.4%, enabling precise error localization that sequence-level methods cannot provide. The framework generalizes across five LLMs spanning two model families (AUROC 0.908--0.944), and we additionally observe that sequence-level uncertainty is partially decomposable: the learned importance-weighted span composition achieves \rho_\textseq = 0.839, suggesting that span-level estimation subsumes sequence-level as a special case.
Abstract: Music-inspired Automatic Stage Lighting Control (ASLC) has gained increasing attention in recent years due to the substantial time and financial costs associated with hiring and training professional lighting engineers. However, existing methods suffer from several notable limitations: the low interpretability of rule-based approaches, the restriction to single-primary-light control in music-to-color-space methods, and the limited transferability of music-to-controlling-parameter frameworks. To address these gaps, we propose SeqLight, a hierarchical deep learning framework that maps music to multi-light Hue-Saturation-Value (HSV) space. Our approach first customizes SkipBART, an end-to-end single primary light generation model, to predict the full light color distribution for each frame, followed by hybrid Imitation Learning (IL) techniques to derive an effective decomposition strategy that distributes the global color distribution among individual lights. Notably, the light decomposition module can be trained under varying venue-specific lighting configurations using only mixed light data and no professional demonstrations, thereby flexibly adapting across diverse venues. In this stage, we formulate the light decomposition task as a Goal-Conditioned Markov Decision Process (GCMDP), construct an expert demonstration set inspired by Hindsight Experience Replay (HER), and introduce a three-phase IL training pipeline, achieving strong generalization capability. To validate our IL solution for the proposed GCMDP, we conduct quantitative analysis to compare model performance across different training phases, demonstrating that our design effectively improves performance and generalization capacity. Furthermore, we also conduct a human study to evaluate SeqLight by comparing it with competitive baselines on music-conditioned light generation tasks across different music styles. The results show that SeqLight achieves the best overall preference scores in both in-domain and out-of-domain settings. The code and trained parameters of this paper is provided at the anonymous repository https://anonymous.4open.science/r/SeqLight-23EE .
Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a core technique for post-training of Large Language Models (LLMs). While policy optimization is driven by all sampled tokens under a globally broadcast scalar reward, the heterogeneous policy behaviors exhibited along trajectories are largely overlooked without differentiation. Existing works address this by credit allocation, including token-level advantage reweighting, and selective token optimization, however, the allocation criterion are principally stagnant throughout training, limiting resilient policy evolution. In this work, we argue that when learning signals are scheduled can be as important as where they are allocated across tokens, and introduce the temporal dimension that scheduling the credit allocation criteria over the course of RLVR optimization. We find that prioritizing targeted tokens emphasized with specific policy behaviors, and gradually attenuating toward general optimization leads to more stable and efficient learning dynamics. Furthermore, we show that simple trajectory percentiles provide a natural perspective for distinguishing policy behaviors, and works effectively with temporal scheduling. Our analysis reveals that standard optimization substantially sacrifices policy entropy when simultaneously accommodating heterogeneous behaviors, whereas temporal scheduling yields healthier policy evolution dynamics. Experiments across mathematical and general reasoning benchmarks demonstrate consistent improvements, suggesting that temporal scheduling constitutes a promising optimization dimension.
Authors: Pasquale De Marinis, Uzay Kaymak, Rogier Brussee, Gennaro Vessio, Giovanna Castellano
Abstract: Few-Shot Semantic Segmentation (FSS) models achieve strong performance in segmenting novel classes with minimal labeled examples, yet their decision-making processes remain largely opaque. While explainable AI has advanced significantly in standard computer vision tasks, interpretability in FSS remains underexplored despite its critical importance for understanding model behavior and guiding support set selection in data-scarce scenarios. We argue that matching-based FSS models are interpretable by design: their core similarity computation between support and query features constitutes an inherent attribution mechanism. Our approach, Affinity Explainer (AffEx), makes this intrinsic interpretability explicit by extracting attribution maps directly from matching scores at multiple feature levels, without requiring gradients or external perturbations. We extend standard interpretability evaluation metrics to the FSS domain and propose additional metrics to better capture the practical utility of explanations in few-shot scenarios. Comprehensive experiments on FSS benchmark datasets demonstrate that AffEx significantly outperforms adapted standard attribution methods, confirming that the matching mechanism itself is the key driver of interpretability. Qualitative analysis reveals structured, coherent attention patterns that align with model architectures and enable effective model diagnosis, laying the groundwork for interpretable FSS research.
Authors:
Jinghui Lu, Jiayi Guan, Zhijian Huang, Jinlong Li, Guang Li, Lingdong Kong, Yingyan Li, Han Wang, Shaoqing Xu, Yuechen Luo, Fang Li, Chenxu Dang, Junli Wang, Tao Xu, jing wu, Jianhua Wu, Xiaoshuai Hao, Wen Zhang, Tianyi Jiang, Lingfeng Zhang, Lei Zhou, Yingbo Tang, Jie Wang, Yinfeng Gao, Xi zhou Bu, Haochen Tian, Yihang Qiu, Feiyang Jia, Lin Liu, Yigu Ge, Hanbing Li, Jiahong Chen, Zihui Li, Shen Yuannan, Jianwei Cui, Hongwei Xie, Bing Wang, Haiyang Sun, Jingwei Zhao, JiaHui Huang, Pei Liu, zeyu zhu, Yuncheng JIANG, Zibin Guo, Hanchao Leng, Chuhong Gong, Kun Ma, Guang Chen, Kuiyuan Yang, Hangjun Ye, Long ChenAbstract: Chain-of-Thought (CoT) reasoning has become a powerful driver of trajectory prediction in VLA-based autonomous driving, yet its autoregressive nature imposes latency prohibitive for real-time deployment. Latent CoT methods attempt to close this gap by compressing reasoning into continuous hidden states, but consistently fall short of their explicit counterparts. We argue that this is because purely linguistic latent representations compress a symbolic abstraction of the world rather than the causal dynamics that govern driving. We present OneVL (One-step latent reasoning and planning with Vision-Language explanations), a unified VLA and world model framework that routes reasoning through compact latent tokens supervised by dual auxiliary decoders. It comprises a language decoder that reconstructs text CoT and a visual world model decoder that predicts future-frame tokens, forcing the latent space to internalize causal scene dynamics. A three-stage training pipeline progressively aligns these latents with trajectory, language, and visual objectives. During inference, the decoders are discarded, and all latent tokens are prefilled in a single parallel pass, matching answer-only prediction speed. Across four benchmarks, OneVL becomes the first latent CoT method to surpass explicit CoT, delivering state-of-the-art accuracy at answer-only latency. Code will be publicly available.
Abstract: Reinforcement learning with verifiers (RLVR) has become a central paradigm for improving LLM reasoning, yet popular group-based optimization algorithms like GRPO often suffer from exploration collapse, where the models prematurely converge on a narrow set of high-scoring patterns, lacking the ability to explore new solutions. Recent efforts attempt to alleviate this by adding entropy regularization or diversity bonus. However, these approaches do not change the winner-takes-all nature, where rollouts still compete for individual advantage rather than cooperating for maximizing global diversity. In this work, we propose Group Cooperative Policy Optimization (GCPO), which shifts the training paradigm from rollout competition to team cooperation. Specifically, GCPO replaces independent rollout scoring with team-level credit assignment: a rollout is rewarded by how much it contributes to the team's valid solution coverage, rather than its individual accuracy. This coverage is described as a determinant volume over reward-weighted semantic embeddings, where only correct and non-redundant rollouts contribute to this volume. During advantage estimation, GCPO redistributes the collective team reward to each single rollout according to its average marginal contribution to the team. This cooperative training paradigm routes optimization toward non-redundant correct reasoning paths. Experiments across multiple reasoning benchmarks demonstrate that GCPO significantly improves both reasoning accuracy and solution diversity over existing approaches. Code will be made available.
Abstract: Mechanistic interpretability seeks to explain neural network behavior by decomposing model computations into interpretable features and circuits. While transcoder-based circuit tracing has recently enabled detailed causal analyses of large language models, multimodal diffusion transformers for image generation remain comparatively opaque. We still lack tools for understanding how semantic information propagates across denoising steps and how text and image representations interact within double-stream MM-DiT architectures. Existing methods provide only partial insight: attention maps expose a limited view of token interactions, while sparse autoencoders can discover interpretable features but do not directly reveal how these features are transformed and composed through nonlinear MLP layers. In this work, we extend transcoder-based circuit tracing to multimodal diffusion transformers. We train timestep-conditioned transcoders that faithfully approximate the input–output behavior of MLP sublayers in FLUX.1[schnell]. By replacing MLPs with transcoders and linearizing the remaining computation, we obtain exact feature-to-feature attribution and recover compact, interpretable circuits. Empirically, our transcoders match or slightly outperform sparse autoencoders on the sparsity–faithfulness tradeoff. The resulting circuits reveal mechanisms underlying attribute binding and cross-stream semantic propagation, and provide causal explanations for systematic generation errors. Moreover, circuit-guided interventions are substantially more precise and effective than standard SAE-based steering. Our results demonstrate that transcoder-based circuit analysis is feasible for state-of-the-art diffusion transformers and provides a powerful framework for understanding and controlling multimodal generative models.
Abstract: Lightweight vision-language models perform competitively on standard benchmarks yet fail systematically in dense-scene reasoning, where multiple objects, attributes, and relations must be jointly grounded and resolved through multi-step inference. Such capability is critical for real-world applications where models must reliably interpret cluttered environments. Yet existing training signals provide no explicit grounding between reasoning steps and the underlying visual entities and relations, leaving lightweight models free to generate fluent but visually unanchored reasoning chains. To address this gap, we first introduce DRBench, a benchmark of 14,573 questions across 2,943 images, organized into five task categories spanning three progressive reasoning layers. Building on DRBench, we propose DRScaffold, a supervised fine-tuning framework that decomposes the supervision target into four causally ordered stages, enforcing grounded reasoning without architectural modification. Experiments on three lightweight VLMs demonstrate substantial gains on DRBench while preserving or improving performance on general-purpose benchmarks. Notably, Qwen2.5-VL-3B trained with DRScaffold surpasses the frozen Qwen2.5-VL-32B on DRBench, demonstrating that structured supervision can substitute for a significant portion of model scale in dense-scene reasoning.
Authors:
Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan, Chunyang Jiang, Cheng CHEN, Chang Liu, Zhenbang Sun, Wei Xue, Wenhan Luo, Yike GuoAbstract: Human-object-centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance between high subject fidelity and accurate interaction patterns between humans and diverse objects, especially when objects represent abstract concepts such as logos. Second, while intra-subject references (e.g., OCR maps, multi-view inputs) are expected to enhance subject fidelity, most existing works lack mechanisms to understand such latent correspondence. To address both challenges, we propose HOMIE, an HOCVP framework that tackles both inter- and intra-subject input settings in a unified manner. Compared to previous approaches, HOMIE proposes a better MLLM integration strategy to extract knowledge of reference-level relationships without compromising the controllability of text encoders or incurring costly re-alignment. Specifically, we introduce global multimodal guidance within self-attention to better align MLLM-derived semantic features with VAE tokens. Furthermore, we propose modality-reference embedding to differentiate tokens from MLLM features and VAE tokens and associate intra-subject reference image tokens. Extensive experiments validate that our method achieves state-of-the-art performance across various HOCVP tasks. Anonymous project page with videos is available at https://homie-demo.github.io/.
Authors:
Ziwei Luo, Ziqi Jin, Lei Wang, Lidong Bing, Thomas SchönAbstract: This work presents self-rewarding sequential Monte Carlo (SMC), an inference-time scaling algorithm enabling effective sampling of masked diffusion language models (MDLMs). Our algorithm stems from the observation that most existing MDLMs rely on a confidence-based sampling strategy, where only tokens with the highest prediction confidence are preserved at each step. This can restrict generation to a noise-sensitive, greedy decoding paradigm, often limiting trajectory diversity. We address this problem by launching multiple interacting diffusion processes in parallel, referred to as , for trajectory exploration. Importantly, we introduce the trajectory-level confidence as a self-rewarding signal for assigning particle importance weights. During sampling, particles are iteratively weighted and resampled to systematically steer generation towards globally confident, high-quality samples. Our self-rewarding SMC is verified on various masked diffusion language models and benchmarks, achieving significant improvement without extra training or reward guidance, while effectively converting parallel inference capacity into improved sampling quality.
Abstract: Treating patients with combinations of drugs reduces the risk of resistance to any individual drug. Finding effective combinations is difficult because the large search space makes combinatorial screens prohibitively expensive, time consuming, and often technically infeasible. Predictive models can fill this gap, yet existing methods typically require molecular profiling of each sample and per-cohort training, limiting their applicability when time and tissue are scarce. To address this challenge, we introduce ScreenShot, a hierarchical transformer pretrained on 40 drug screening datasets covering 3,700 drugs and 6,000 biological samples, whose architecture mirrors the nested structure of screening data. Given a few-shot context of observations from a new patient, ScreenShot predicts the response of the sample to combination therapies through in-context learning, operating directly on functional measurements with no fine-tuning and no molecular profiling. On four held-out datasets, ScreenShot outperforms all baselines in both prediction accuracy and identification of selectively effective treatments. ScreenShot's internal representations are directly useful for experimental design: we use them to drive a weighted k-means++ active learning strategy that selects which experiments to run, achieving the same hit detection as uniform screening with a third of the budget. Source code and interactive dashboard: https://github.com/AnonymousAccount3/screenshot
Authors:
Baining Zhao, jiacheng xu, Weicheng Feng, Xin Zhang, Zhaolu Wang, Haoyang Wang, Shilong Ji, Ziyou Wang, Jianjie Fang, Zhiheng Zheng, Weichen Zhang, Yu Shang, Wei Wu, Chen Gao, Xinlei Chen, Yong LiAbstract: Aerial vision-language navigation requires agents to follow natural-language instructions through closed-loop perception and action in 3D environments. We argue that aerial VLN can be formulated as a prediction-driven world-action problem: the agent should anticipate latent world evolution and act according to the predicted consequences. To this end, we propose WorldVLN, the first autoregressive world action model for aerial VLN. Unlike full-sequence video-generation world models that generate an entire visual clip, WorldVLN adapts a latent autoregressive video backbone to predict short-horizon world-state transitions and directly decodes them into executable waypoint actions. After each action segment is executed, newly received observations are encoded back into the autoregressive context, enabling closed-loop world-action prediction. We further introduce a two-stage training framework that first grounds the video prior in instruction-conditioned navigation dynamics and then develops Action-aware GRPO, the first reinforcement learning method tailored to autoregressive WAMs, to optimize waypoint decisions through their downstream rollout consequences. On public outdoor and indoor benchmarks, WorldVLN consistently outperforms existing Vision-Language-Action baselines with 12%+ success-rate gains and larger advantages on challenging cases. It further transfers zero-shot to real drone deployment, suggesting that the proposed WorldVLN offers a promising route for spatial action tasks. Demos and code are available at \urlhttps://worldvln.com (anonymous).
Abstract: We introduce a new approach to high-fidelity 3D scene reconstruction from multi-view RGB images that tightly couples reconstruction with a strong generative 3D prior. We cast scene reconstruction as conditional 3D generation over a set of spatially-localized, overlapping chunks that together tile the scene, scaling generation to large scene extents. Crucially, we inherit the fidelity and completeness of state-of-the-art generative shape models -- we use Trellis.2 as an example -- which we generalize to the scene level. To this end, we propose a projection-based conditioning mechanism that lifts posed multi-view image features into a coherent 3D representation aligned with the generative model, independent of view ordering and spatially anchored to the scene, yielding high-fidelity, multi-view consistent generated geometry. This enables lifting the strong object-level prior of Trellis.2 to multi-view, scene-scale generation, producing faithful, editable PBR mesh reconstructions of indoor environments. As a result, we obtain high-fidelity results that outperform cutting-edge reconstruction methods by 16%.
Abstract: Recent work has demonstrated that online reinforcement learning (RL) can substantially improve the quality and alignment of flow matching models for image and video generation. Methods such as Flow-GRPO and CPS cast the denoising process as a Markov Decision Process and apply PPO-style ratio clipping to enforce a trust region. However, we argue that : the probability ratio between new and old policies is a noisy, single-sample estimate of the true policy divergence, leading to over-constraining in some regions of the trajectory and under-constraining in others. We propose (Flow Divergence Proximal Policy Optimization), which replaces ratio clipping with a divergence proximal constraint. A key observation is that the per-step policy in flow models is Gaussian, enabling computation of the KL divergence between old and new policies. Flow-DPPO employs an asymmetric divergence mask that blocks gradient updates only when they simultaneously move away from the trusted region and violate the divergence threshold. Experiments show that Flow-DPPO achieves higher rewards with better KL-proximal efficiency, alleviates catastrophic forgetting, promotes balanced multi-objective optimization, and enables stable multi-epoch training where ratio clipping degrades.
Abstract: Large reasoning models (LRMs) improve complex problem-solving by generating long intermediate reasoning traces, but this substantially increases inference costs. NVFP4 inference offers a promising approach to reduce both computational and memory costs through hardware-supported low-precision execution. However, directly applying NVFP4 to LRMs introduces two practical limitations: reasoning accuracy degrades under quantization, and existing NVFP4 kernels do not fully realize latency benefits in small-batch autoregressive decoding. In this work, we analyze the effect of NVFP4 quantization on token-level uncertainty during reasoning. We show that quantization increases incorrect sampling at low-entropy symbolic tokens, while causing over-concentration on a small set of tokens in high-uncertainty reasoning steps. Based on this observation, we propose ReSET, a reasoning-step entropy-based temperature-scaling method that estimates step-level uncertainty online and adapts the decoding temperature using both token-level and step-level entropy signals. To address the latency gap, we further design a CUDA-core small-M NVFP4 kernel for latency-critical autoregressive decoding. Across reasoning benchmarks and model scales, ReSET improves NVFP4 reasoning accuracy by up to ~2 points over the NVFP4 baseline. Our CUDA-core small-M kernel further improves latency-critical decoding, delivering up to 2.5× kernel-level speedup over NVFP4 vLLM and approximately 2× end-to-end decoding speedup over BF16. Code will be released.
Abstract: A latent world model may achieve accurate short-horizon prediction while still inducing a latent space that is poorly aligned with planning. A key issue is spatiotemporal mismatch: these models are often trained with local predictive supervision, but deployed for long-horizon goal-directed search in latent spaces where Euclidean distance may not reflect what is reachable within a finite action budget. We present the Reachability-Correction auxiliary objective (RC-aux), a lightweight correction for this mismatch in reconstruction-free latent world models. RC-aux keeps the world-model backbone unchanged and adds planning-aligned supervision along two axes. Along the time axis, multi-horizon open-loop prediction trains the model beyond one-step consistency. Along the space axis, budget-conditioned reachability supervision, together with temporal hard negatives, encourages the latent space to distinguish states that are eventually reachable from those reachable within the current planning horizon. At test time, the learned reachability signal can also be used by a reachability-aware planner to favor trajectories that are both goal-directed and attainable under the available budget. We instantiate RC-aux on LeWorldModel and evaluate it under both continuation-training and matched-from-scratch settings. Across goal-conditioned pixel-control tasks and a LIBERO-Goal extension, RC-aux improves LeWM-style planning with modest additional cost. These results suggest that planning with latent world models depends not only on predictive accuracy, but also on whether the learned representation encodes the temporal and geometric structure required by downstream search. The source code will be released upon acceptance of the paper.
Abstract: Multi-dimensional Root Cause Analysis (RCA) requires searching an exponential \mathcalO(2^M) attribute space over millions of records within seconds to isolate the root-cause combination. Existing methods fail at this industrial scale due to a fundamental modeling oversight: employing symmetric objectives that treat anomalous and normal samples interchangeably, allowing the massive normal background to dominate computation. To overcome this, we propose PRISM (\underlinePrecision-\underlineRecall \underlineInformed \underlineSubspace \underlineMining), which reformulates RCA as an asymmetric Precision-Recall optimization. Precision is defined over anomaly-normal pairs and Recall over anomaly-anomaly pairs, eliminating uninformative normal-normal comparisons entirely. Coupled with tensor compression and a two-stage solver, comprising Bernoulli continuous relaxation and combinatorial verification, PRISM reduces pairwise complexity from a full-dataset quadratic scale to an anomaly-bounded \mathcalO(|\mathcalP|^2). Evaluations on four public benchmarks and a new 5.13M-record industrial dataset show that PRISM achieves the highest average F1 (0.90 vs.\ 0.65 for the strongest baseline), while maintaining sub-second latency across a 2^110 attribute space. We publicly release this dataset as NFC, establishing the first production-sourced, million-scale benchmark for multi-dimensional RCA. \hrefhttps://anonymous.4open.science/status/PRISM-09FC[Code]
Authors:
ZhiYuan Feng, Yu Deng, Ruichuan An, Zhenhua Liu, Qixiu Li, Keming Wu, Zhiying Du, Weijie Wang, Haoxiao Wang, Shuang Chen, Sicheng Xu, Yaobo Liang, Jiaolong Yang, Baining GuoAbstract: In real home deployments, household agents must often operate from a complete household scene and a situated household request, rather than from a clean task specification. Such requests require agents to identify task-relevant entities, recover intended task conditions, and resolve ordering constraints from the surrounding scene context. We formalize this capability as full-scene household reasoning: given a complete household scene and a situated household request, an agent must infer executable task structure before producing a grounded skill-level action sequence. This setting is challenging because complete household scenes contain substantial task-irrelevant information, making direct complete-scene prompting inefficient and error-prone. In practical deployment, this challenge is further amplified by privacy and local compute constraints, which favor compact open-weight models with limited long-context reasoning ability. We propose TaskGround, a training-free and model-agnostic Ground-Infer-Execute framework that grounds complete scenes into compact task-relevant scene slices, infers executable task structure, and compiles it into grounded skill-level action sequences. To evaluate this setting, we introduce FullHome, a human-validated evaluation suite of 400 household tasks spanning diverse home-scale environments and both goal-oriented and process-constrained requirements. On FullHome, TaskGround improves task success rates by large margins across both proprietary and open-weight models. Notably, it makes Qwen3.5-9B competitive with GPT-5 under direct complete-scene prompting while reducing total input-token cost by up to 18x. Our results identify executable task-structure inference as a central bottleneck in full-scene household reasoning and show that structured grounding can make compact local models substantially more effective for practical household deployment.
Authors: David Busbib, Michael Werman
Abstract: A generated plausible future mathematical claim must satisfy two constraints: it should follow the direction of prior work and respect the formal dependencies that constrain the validly of the claim. Existing approaches typically model only one of these sources, producing claims that are either weakly grounded or insufficiently motivated. We introduce grounded future mathematical generation, the task of generating a plausible future theorem-like claim for an anchor paper, using two complementary sources of context: its scientific citation graph ,and its aligned formal theorem dependency graph. To address this setting, we propose COMPOSE, a dual-graph framework that combines scientific citation context with formal mathematical structure to condition a language model. To support this setting, we construct a dataset of 108K paired scientific-formal graphs from arXiv and Mathlib, together with a benchmark of 47K future papers from 2024-2025. Experiments show that COMPOSE outperforms strong baselines on retrieval to real future papers and achieves the best overall performance under LLM-judge evaluation, producing more grounded and mathematically richer outputs. These results show that future mathematical generation benefits from combining scientific context with formal structure.
Authors:
Xuenan Xu, Jiahao Mei, Ye Tao, Zihao Zheng, Zeyu Xie, Yaoyun Zhang, Haohe Liu, Yuning Wu, Ming Yan, Mengyue Wu, Wen Wu, Chao ZhangAbstract: Audio generation, including speech, music and sound effects, has advanced rapidly in recent years. These tasks can be divided into two categories: time-aligned (TA) tasks, where each input unit corresponds to a specific segment of the output audio (e.g., phonemes aligned with frames in speech synthesis); and non-time-aligned (NTA) tasks, where such alignment is not available. Since modeling paradigms for the two types are typically different, research on different audio generation tasks has traditionally followed separate trajectories. In this work, we propose UniFlow-Audio, a non-autoregressive audio generation framework that unifies the two types based on flow matching. UniFlow-Audio introduces a dual-fusion mechanism that aligns audio latents with TA features while integrating NTA conditions via cross-attention in each block, and employs task-balanced sampling to maintain consistent performance across diverse tasks. Supporting omni-modal inputs including text, audio, and video, UniFlow-Audio achieves strong results on 7 audio generation tasks with fewer than 8K hours of public training data and under 1B trainable parameters. A compact 200M-parameter variant remains competitive, indicating UniFlow-Audio as a promising foundation model for general non-autoregressive audio generation. Codes and models are available at https://anonymous3387a8c.github.io/uniflow_audio.
Abstract: Industrial anomaly detection demands precise reasoning over fine-grained defect patterns. However, existing multimodal large language models (MLLMs), pretrained on general-domain data, often struggle to capture category-specific anomalies, thereby limiting both detection accuracy and interpretability. To address these limitations, we propose Reason-IAD, a knowledge-guided dynamic latent reasoning framework for explainable industrial anomaly detection. Reason-IAD comprises two core components. First, a retrieval-augmented knowledge module incorporates category-specific textual descriptions into the model input, enabling context-aware reasoning over domain-specific defects. Second, an entropy-driven latent reasoning mechanism conducts iterative exploration within a compact latent space using optimizable latent think tokens, guided by an entropy-based reward that encourages confident and stable predictions. Furthermore, a dynamic visual injection strategy selectively incorporates the most informative image patches into the latent sequence, directing the reasoning process toward regions critical for anomaly detection. Extensive experimental results demonstrate that Reason-IAD consistently outperforms state-of-the-art methods across multiple tasks. The code will be publicly available upon publication.
Authors: Qilang Ye, Meng Liu, Yu Zhou
Abstract: We explore whether Omni-Language Models (OLMs) can be directly applied to zero-shot Semantic Audio-Visual Navigation (SAVN). Recent work demonstrates that even state-of-the-art specialized models still struggle to achieve generalist multimodal navigation, despite extensive task-specific training. In this paper, we introduce , short for Reasoning All-in-One OLM, a deployment pipeline for zero-shot SAVN. By leveraging the rich implicit audio-visual knowledge encoded in OLMs, the embodied agent is enabled to hear, see, reason, and act in the environment. To further elicit the built-in thinking ability of OLMs, we propose a test-time Latent Navigation Reasoning (LNR) module that can be seamlessly integrated into the decoding space. LNR encourages the model to retrieve more target-relevant observations and make effective navigation decisions. Through comprehensive experiments, we show that our framework surpasses existing state-of-the-art baselines on public SAVN benchmarks without using any training data. Moreover, we introduce a new Global Navigation Instruction setting to further evaluate the ability of OLMs to serve as embodied navigation agents. Results under this setting posit that applying OLMs to real-world natural-language-guided audio-visual navigation remains challenging.
Abstract: GUI agents interact with environments by perceiving interfaces and executing actions. As a virtual sandbox, the GUI World model empowers agents with human-like foresight by enabling action-conditioned prediction. However, existing text/pixel-based approaches struggle to achieve both high visual fidelity and fine-grained structural controllability. To this end, we propose Code2World, a vision-language coder that simulates the next visual state via renderable code generation. Specifically, to address the data scarcity problem, we construct AndroidCode by translating GUI trajectories into high-fidelity HTML and refining synthesized code through a visual-feedback revision mechanism, yielding a corpus of over 80K high-quality screen-action pairs. To adapt existing VLMs into code prediction, we first perform SFT as a cold start for format layout following, then further apply Render-Aware Reinforcement Learning which uses rendered outcome as the reward signal by enforcing visual semantic fidelity and action consistency. Extensive experiments demonstrate that Code2World-8B achieves the top-performing next UI prediction, rivaling the competitive GPT-5 and Gemini-3-Pro-Image. Notably, Code2World significantly enhances downstream navigation success rates in a flexible manner, boosting Gemini-2.5-Flash by +9.5% on AndroidWorld navigation.
Abstract: Uncertainty estimation has been a long-standing challenge in AI models; it amounts to ``knowing what you don't know,'' and metacognition is notoriously difficult even for humans (cf. the Dunning-Kruger effect). Although it is still far from solved even in simpler classification systems, tackling it in multimodal large language models (MLLMs) is becoming increasingly important. Within MLLMs, uncertainty can stem from any of the diverse sources as well as from their relationships, and further can stem from the unbounded answers in the open-ended setting. To tackle the issues, we propose CoMet, an MLLM uncertainty estimation method by decomposing uncertainty into a context-specific term and a multiplicity-specific term. The former captures ambiguity induced by the given context (e.g., task or prompt), while the latter captures how many plausible answers determined by the context remain compatible with the given input. We train a lightweight post-hoc uncertainty module to estimate these quantities, which enables efficient uncertainty estimation without autoregressive answer generation or repeated sampling. Experiments on various open-ended multimodal benchmarks, hallucination detection, and multiple-choice visual question answering benchmarks show that CoMet consistently improves uncertainty estimation over existing baselines while remaining efficient in practice. Code will be available upon acceptance.
Authors: Haoyun Tang, Haodong Cui, Keyao Xu, Zhan-Dong Mei, Kun Wang
Abstract: World models enable model-based planning through learned latent dynamics, but imagined rollouts often become unstable as the planning horizon grows or the dynamics distribution shifts. We argue that this instability arises from two missing structures in planner-facing latents: history-conditioned memory for approximate Markov completeness, and geometric organization that separates configuration, momentum, and task semantics. We propose HaM-World, a structured world model that decomposes the latent state into a canonical (q, p) subspace and a context subspace c, while incorporating Mamba selective state-space memory as a history-conditioned input to the same latent dynamics. Within this unified interface, (q, p) evolves under a Soft-Hamiltonian dynamics composed of an energy-derived Hamiltonian vector field and learnable residual and control dynamics, while c captures semantic, dissipative, and non-conservative factors. This design provides the planner with a single latent representation shared across dynamics prediction, reward and value estimation, imagined rollouts, and cross-entropy method (CEM) planning. On four DeepMind Control Suite tasks, HaM-World achieves the highest average AUC (+9.5% over strong baselines), reduces long-horizon rollout error to 45% of a competitive model, and wins 11 out of 12 k \in \3,5,7\ rollout MSE metrics. Under 12 out-of-distribution perturbations, HaM-World consistently attains the highest returns, with average gains of 10.2% on Finger Spin and 13.6% on Reacher Easy. Mechanism diagnostics further demonstrate bounded energy drift under action-free rollouts, structured energy variation under policy control, and coherent control-induced energy transfer, supporting the effectiveness of the proposed Soft-Hamiltonian latent dynamics. Code: https://anonymous.4open.science/r/HaM_World-47CD
Abstract: Fixed-point inversion improves data-to-noise inversion by explicitly solving the local inverse equation at each timestep as a fixed-point problem. Empirically we observe that fixed-point inversion can reach distinct approximate roots, and the resulting inverse trajectories can differ substantially in how reconstruction errors accumulate. However, existing fixed-point inversion methods lack a principled mechanism for selecting among these solutions. Therefore, we propose SelFix, a root-selecting fixed-point inversion method for rectified flows that uses trajectory straightness as the selection criterion. We derive an on-the-fly proxy for rectified flow straightness from previously recovered inverse velocities and use it to construct a local straightness anchor. A vanishing anchored fixed-point update, combined with decoupled momentum for finite-iteration stability, then biases the iteration toward the fixed point that minimizes this selector while preserving the original inverse equation asymptotically. Under a standard local nonexpansiveness assumption, SelFix converges to the straightness-selected exact local inverse root. Experiments on FLUX.1-dev and PIE-Bench show that SelFix improves fixed-point inversion for rectified flows, achieving stronger real-image reconstruction and better source-preserving prompt-based editing than prior inversion baselines.
Abstract: Recent feed-forward 3D reconstruction methods, such as visual geometry transformers, have substantially advanced the traditional per-scene optimization paradigm by enabling effective multi-view reconstruction in a single forward pass. However, most existing methods struggle to achieve a balance between reconstruction quality and computational efficiency, which limits their scalability and efficiency. Although some efficient visual geometry transformers have recently emerged, they typically use the same sparsity ratio across layers and frames and lack mechanisms to adaptively learn representative tokens to capture global relationships, leading to suboptimal performance. In this work, we propose TurboVGGT, a novel approach that employs an efficient visual geometry transformer with adaptive alternating attention for fast multi-view 3D reconstruction. Specifically, TurboVGGT employs an end-to-end trainable framework with adaptive sparse global attention guided by adaptive sparsity selection to capture global relationships across frames and frame attention to aggregate local details within each frame. In the adaptive sparse global attention, TurboVGGT adaptively learns representative tokens with varying sparsity levels for global geometry modeling, considering that token importance varies across frames, attention layers operate tokens at different levels of abstraction, and global dependencies rely on structurally informative regions. Extensive experiments on multiple 3D reconstruction benchmarks demonstrate that TurboVGGT achieves fast multi-view reconstruction while maintaining competitive reconstruction quality compared with state-of-the-art methods.
Authors:
Kai Ong, Minseok Kang, Dongwook Choi, Junhee Cho, Seungwon Lim, Seungju Kim, Geunha Jang, Minwoo Oh, Bogyung Jeong, Sunghwan Kim, Taeyoon Kwon, Jihye Han, Juyoung wy, Jinyoung YeoAbstract: Harness optimization enables automated agent creation by having an optimizer agent to iteratively update the harness of target agents. Despite its success, current studies evaluate optimizers solely by observing target agents' performance gains. This indirect end-improvement evaluation neglects optimizers' actions at intermediate steps, which are often erroneous and hinder agent performance. Therefore, it is unclear whether harness optimization is driven by optimizers' informed update actions or simply trial-and-error. This necessitates direct evaluation of harness optimizers. However, evaluating harness optimizers directly is non-trivial and costly due to the lack of oracle harnesses. To address this, we present a simple, low-cost design to directly evaluate them, namely priority ranking. By asking harness optimizers to rank components (e.g, tools) in a given harness by their potential to improve/hinder agent performance when updated, our design quantifies optimizer ability at the step level without expensive rollouts or manual examination. More importantly, optimizers' ranking performance correlates with their ability to improve agents in actual multi-step harness optimization, establishing priority ranking as a reliable predictor of optimization ability. Priority ranking is enabled by SHOR, a collection of 182 human-verified optimization scenarios spanning across domains, designs, and time stages. Codes and data can be found at https://anonymous.4open.science/r/Harness_Eval-6D86/%7D%7D.
Abstract: Recent advances in video diffusion have been driven by scaling transformer-based architectures to billions of parameters, substantially improving visual fidelity and motion coherence. In contrast, existing mobile video diffusion models remain limited to relatively small parameter budgets, typically 0.4–1.8B, restricting generation quality. In this work, we show that high-quality mobile video generation does not require small models. Instead, we demonstrate that a server-scale 5B-parameter video diffusion transformer can be deployed efficiently on memory-constrained mobile hardware through recurrent reformulation and structured compression. Starting from Wan2.2-5B, we rely on a recurrent distillation framework that converts video generation into a chunk-wise autoregressive process with constant-memory attention computation. Combined with causal linear attention, the model operates as an RNN at inference time while preserving temporal coherence across chunks. We further propose a learnable attention head pruning method based on binary per-head gates optimized end-to-end using a noise-biased sparsity objective and distillation-based finetuning. Together with sampling-step distillation and memory-optimized VAE decoding, MobileWan becomes the first 5B-scale video diffusion model deployable on a commercial mobile device. Our system generates 5-second 480x832 videos at 16 FPS in 20 seconds end-to-end latency, achieving a VBench score of 83.79 and establishing a new state of the art in mobile video generation. The model checkpoint will be released.
Abstract: Autoregressive video generation has improved rapidly in visual fidelity and interactivity, but it still suffers from long-term inconsistency and memory degradation. Most existing solutions either compress historical frames using predefined strategies or retrieve keyframes based on coarse implicit attention signals, both of which fail to handle evolving prompts with shifting entity references, leading to identity drift, character duplication, and attribute loss. To address this, we propose IAMFlow, a training-free identity-aware memory framework that explicitly models and tracks persistent entity identities, enabling consistent generation across prompt transitions. Specifically, an LLM extracts entities with visual attributes from each prompt and assigns unique global IDs for identity-aware memory, while a VLM asynchronously verifies and refines attributes from rendered frames, enabling explicit entity tracking in place of implicit similarity-based matching. To keep the proposed framework computationally practical, we design a systematic inference acceleration pipeline, including asynchronous visual verification, adaptive prompt transition, and model quantization, which achieves faster generation than existing baselines. Furthermore, we introduce NarraStream-Bench, a benchmark for narrative streaming video generation that features 324 multi-prompt scripts spanning six dimensions and a three-dimensional evaluation protocol that integrates both traditional metrics and multimodal large language model-based assessments. Extensive experiments show that IAMFlow, despite being training-free, achieves the best overall performance on NarraStream-Bench, outperforming the strongest baseline by 2.56 points, while achieving a 1.39× speedup over the most efficient baseline in the 60-second multi-prompt setting.
Abstract: Training capable Large Language Model (LLM) agents is critically bottlenecked by the high cost and static nature of real-world interaction data. We address this by introducing GenEnv, a framework that establishes a difficulty-aligned co-evolutionary game between an agent and a scalable, generative environment simulator. Unlike traditional methods that evolve models on static datasets, GenEnv instantiates a dataevolving: the simulator acts as a dynamic curriculum policy, continuously generating tasks specifically tailored to the agent's ``zone of proximal development''. This process is guided by a simple but effective -Curriculum Reward, which aligns task difficulty with the agent's current capabilities. We evaluate \autoenv on five benchmarks, including API-Bank, ALFWorld, BFCL, Bamboogle, and TravelPlanner. Across these tasks, GenEnv improves agent performance by up to 40.3% over 7B baselines and matches or exceeds the average performance of larger models. Compared to Gemini 2.5 Pro-based offline data augmentation, \autoenv achieves better performance while using 3.3 less data. By shifting from static supervision to adaptive simulation, \autoenv provides a data-efficient pathway for scaling agent capabilities.
Abstract: As large language models (LLMs) continue to advance rapidly, they are becoming increasingly capable while simultaneously demanding ever-longer context lengths. To improve the inference efficiency of long-context processing, several novel low-complexity hybrid architectures have recently been proposed, effectively alleviating the computational burden of long-context inference. However, existing research on long-context prefill acceleration remains predominantly focused on sparse attention mechanisms, which achieve their maximum speedup only on full-attention models. When transferred to emerging architectures — such as linear/full attention hybrids or sliding window/full attention hybrids — these prefill acceleration approaches suffer significant performance degradation. Furthermore, such methods are generally incompatible with continuous batching, making them difficult to integrate into modern inference engines such as vLLM. To this end, we propose UniPrefill, a prefill acceleration framework applicable to virtually any model architecture, which directly accelerates the model's computation at the token level. We further implement UniPrefill as a continuous batching operator and extend vLLM's scheduling strategy to natively support prefill-decode co-processing for UniPrefill, enabling its seamless integration into vLLM. UniPrefill achieves up to 2.1x speedup in Time-To-First-Token (TTFT), with the acceleration becoming increasingly pronounced as the number of concurrent requests grows.
Abstract: Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradigm for complex reasoning in LLMs, yet commonly suffer from policy entropy collapse during training. We conduct a first-order gradient analysis of token-level entropy dynamics under GRPO and identify a token-level credit assignment mismatch: the per-token entropy variation decomposes into the product of the trajectory-level advantage and an entropy sensitivity function over the next-token distribution, yielding an advantage-surprisal four-quadrant structure and a near-criticality property. Motivated by it, we propose STARE (Surprisal-guided Token-level Advantage Reweighting for policy Entropy stability), which identifies entropy-critical token subsets via batch-internal surprisal quantiles, selectively reweights their effective advantages, and incorporates a target-entropy closed-loop gate for stable entropy regulation. Across model scales from 1.5B to 32B and three task families (Short CoT, Long CoT, and Multi-Turn Tool Use), STARE sustains stable RL training over thousands of steps while maintaining policy entropy within the target band. On AIME24 and AIME25, STARE outperforms DAPO and other competitive baselines by 4%–8% in average accuracy, with reflection tokens and response length growing in tandem, indicating sustained exploration–exploitation balance that further unlocks RL training potential.
Abstract: Deep search has become a crucial capability for frontier multimodal agents, enabling models to solve complex questions through active search, evidence verification, and multi-step reasoning. Despite rapid progress, top-tier multimodal search agents remain difficult to reproduce, largely due to the absence of open high-quality training data, transparent trajectory synthesis pipelines, or detailed training recipes. To this end, we introduce , a fully open-source recipe for training frontier multimodal deep search agents with agentic reinforcement learning. First, we curate a dedicated pipeline to construct high-quality training data through Wikipedia path sampling, fuzzy entity rewriting, and source-anchor visual grounding, which jointly reduce shortcuts and one-step retrieval collapse. Based on this pipeline, we curate two training datasets, for RL. Besides, we design a diverse tool environment that unifies text search, image search, OCR, cropping, sharpening, super-resolution, and perspective correction, enabling agents to combine active perception with external knowledge acquisition. Finally, we propose a multi-turn fatal-aware GRPO training algorithm that handles cascading tool failures by masking post-failure tokens while preserving useful pre-failure reasoning through one-sided advantage clamping. Built on this recipe, OpenSearch-VL delivers substantial performance gains, with over 10-point average improvements across seven benchmarks, and achieves results comparable to proprietary commercial models on several tasks. We will release all data, code, and models to support open research on multimodal deep search agents.
Abstract: Geometric foundation models, such as the Visual Geometry Grounded Transformer (VGGT), provide strong 3D priors from unposed images. However, such models operate purely in a feed-forward, deterministic regime, \ie~they cannot generate plausible geometry beyond what the input views directly support. Generative models for 3D scenes, on the other hand, must rely on strong geometric priors to produce coherent outputs from sparse inputs. We bridge these two paradigms by performing flow matching directly in VGGT's latent space, leveraging its learned 3D priors without committing to any explicit downstream representation such as Gaussians, meshes, or video-VAE latents. This requires respecting the latent geometry: VGGT tokens occupy a product of high-dimensional hyperspheres on which standard Euclidean flow matching fails. We address this with a Riemannian Flow Matching framework defined on a product manifold of four hyperspheres, aligned with VGGT's multi-scale encoder, which keeps generated tokens on the valid data manifold required by the frozen decoding heads. On RealEstate10K and ScanNet++, our method achieves strong performance against recent scene generation baselines in both per-view 3D geometry and appearance, establishing latent-space flow matching on geometric foundation models as a viable paradigm for 3D generation.
Abstract: Generative Retrieval (GR) has emerged as a promising paradigm for modern search systems, offering end-to-end joint optimization and lower serving overhead compared to multi-stage cascaded architectures. OneSearch, as a representative industrially deployed generative framework, has delivered substantial commercial benefits. However, three key limitations constrain further improvement: shallow understanding of complex queries, insufficient personalized intent reasoning over user context, and a separately trained reward model that adapts slowly to emerging queries and is prone to sampling bias and reward hacking. To address these challenges, we propose OneSearch-V2 with three key innovations: (1) a thought-augmented query understanding module that generates compact keyword-based chains-of-thought (CoTs), overcoming the shallow-matching limitation of single-pass SID generation; (2) a reasoning-internalized self-distillation pipeline that encodes keyword-guided reasoning into the existing model weights via information-asymmetric supervision, without any extra parameters, special tokens, or inference-time CoT generation; (3) a behavior-feedback preference alignment system that replaces the separate reward model with composite rewards built from real user interactions, and introduces a token-position marginal advantage (TPMA) mechanism for position-aware credit assignment over hierarchical SID sequences. Extensive offline evaluations demonstrate OneSearch-V2's strong query understanding and personalized intent modeling capabilities. Online A/B tests further validate its business effectiveness, yielding +3.98% item CTR, +1.17% PV CTR, +2.90% PV CVR, +2.07% buyer volume, and +2.11% order volume. Manual evaluation further confirms gains in search experience quality, with +1.37% in page good rate and +1.65% in query-item relevance. Importantly, OneSearch-V2 achieves these gains without any additional inference cost or serving latency, while also mitigating information bubbles and long-tail sparsity.
Abstract: We present a new operator-theoretic representation learning framework for offline reinforcement learning that recovers the directed temporal geometry of a controlled Markov process from hitting time observations. While prior art often produces symmetric distances or fails to satisfy the triangle inequality, our framework learns a Hilbert-space displacement geometry where expected hitting times are realized as linear functionals of latent displacements. We prove that this representation exists under latent linear closure and is uniquely identifiable up to a bounded linear isomorphism. For finite-dimensional implementations, we show that global hitting-time error is bounded by one-step transition error amplified by the environment's transient spectral radius. Furthermore, we provide finite-sample guarantees accounting for approximation, statistical complexity, and trajectory-label mismatch. Derived from this theory, we curate Isomorphic Embedding Learning (IEL) as a new goal-agnostic foundation policy learning algorithm that anchors a HILP-style consistency objective with explicit hitting-time regression to ensure that the learned geometry reflects actual decision-time progress. This asymmetric and compositional structure enables robust graph-based multi-stage planning for long-horizon navigation. Our experiments demonstrate that IEL improves the state of the art of learning foundation policy policies from offline maze locomotion data.
Authors: Santanu Rathod, Pietro Lió, Xiao Zhang
Abstract: Flow matching is an emerging, scalable generative framework for characterizing continuous normalizing flows with wide-range applications. However, state-of-the-art methods are not well-suited for modeling dynamical systems, as they construct conditional paths that are restricted to linear interpolants, a suboptimal supervision signal, which may not capture the system's underlying state evolution. Moreover, constructing unified paths to satisfy multi-marginal constraints across observations is challenging, since naïve higher-order polynomials tend to be unstable and oscillatory. To address these limitations, we introduce SplineFlow, a theoretically grounded flow matching algorithm that jointly models conditional paths across observations via B-spline interpolation. Specifically, SplineFlow exploits the smoothness and stability of B-spline bases to learn the complex underlying dynamics in a structured manner while ensuring the multi-marginal requirements are met. Comprehensive experiments across deterministic and stochastic dynamical systems under various configurations, as well as cellular trajectory inference tasks, demonstrate that SplineFlow outperforms existing baselines, especially when the underlying dynamics are of higher degree or in irregular sampling regimes.
Abstract: Recent large vision–language models (LVLMs) can generate vision–text multimodal chain-of-thought (MCoT) traces after reinforcement fine-tuning (RFT). However, we observe that the visual information in MCoT is often inaccurate even when answers are correct, and task accuracy remains nearly unchanged despite the visual information being corrupted. This indicates a lack of faithfulness in the vision part of MCoT reasoning. We attribute this to the RL reward design in RFT, which solely incentivizes the format of interleaved vision-text cues, encouraging the model to incorporate visual information into its text reasoning steps without considering its correctness. In this paper, we first probe MCoT faithfulness by measuring how much the prediction changes when its visual and textual thoughts are intervened. Surprisingly, the model's predictions remain nearly unchanged under visual intervention but change significantly under textual intervention, indicating that the visual evidence is largely ignored. To further analyze the visual information, we introduce an automated LVLM-based evaluation metric that quantifies the faithfulness of visual cues from two perspectives: relevance and sufficiency. Our evaluation reveals that the visual information in current MCoT traces is simultaneously irrelevant and insufficient. To address this issue, we propose Sufficient-Component Cause Model (SCCM) learning. This approach encourages the MCoT to generate sufficient yet minimal visual components that are independently capable of leading to the correct answer. The proposed SCCM is annotation-free and plug-and-play, compatible with various RFT in MCoT. Empirical results demonstrate that SCCM consistently improves the visual faithfulness across a suite of fine-grained perception and reasoning benchmarks.
Abstract: Biomolecular generators are often adapted with reward feedback to improve task-specific utility, but pushing utility alone can concentrate generation on a narrow family of candidates. Maintaining diversity is difficult because sample diversity is a set-level property. We introduce Supergroup Relative Policy Optimization (SGRPO), a flexible GRPO-style framework that directly constructs rewards from set-level diversity. For each condition, SGRPO samples a supergroup of candidate sets, compares their diversity under the same condition, and redistributes the group diversity reward to individual rollouts through leave-one-out diversity contributions before combining it with rollout-level utility. This design decouples SGRPO from a particular generator, utility reward, or diversity metric, and allows instantiation with different GRPO-style approaches. We evaluate SGRPO on de novo small-molecule design, pocket-based small-molecule design, and de novo protein design, instantiating it with both GRPO and Coupled-GRPO across autoregressive and discrete diffusion generators. Across decoding sweeps, SGRPO expands the utility-diversity Pareto frontier and achieves the best frontier-level metrics relative to pretrained generators, GRPO, and memory-assisted GRPO when applicable. Our analyses further show that direct set-level diversity rewards remain effective with small groups and help preserve broader generation-distribution coverage during post-training. The code is available at https://anonymous.4open.science/r/SGRPO/README.md.
Abstract: Video Large Language Models (VideoLLMs) have achieved strong performance on video understanding tasks, yet existing benchmarks evaluate only what models predict, leaving the stability of how they generate largely unexamined. We surface a previously underexplored generation failure of VideoLLMs, defined as output repetition, in which the decoder collapses into self-reinforcing loops of repeated phrases or sentences, and present VideoSTF, a benchmarking framework for systematically measuring, stress-testing, and exploiting this failure mode. VideoSTF formalizes repetition with three complementary n-gram-based metrics, ships a standardized testbed of 10,000 diverse videos, and provides a library of controlled temporal stressors. Across 10 advanced VideoLLMs, VideoSTF reveals four key findings: (i) repetition is pervasive and frame-count-invariant, with repetition rates up to 91%; (ii) it spans a severity spectrum from mild redundancy to token-cap loops, and is highly amplified by temporal perturbations; (iii) temporal stressors form a practical black-box attack surface, flipping benign videos into repetitive ones with tens of queries and high attack success rates (up to 98%); and (iv) repetition is a model-level failure decoupled from input redundancy and driven by local temporal disruption, and is robust to decoding, input filtering, prompts, and video distribution shifts, suggesting that mitigation requires architectural or training-level intervention rather than decoding-time fixes. VideoSTF reframes generation stability as a useful evaluation axis for VideoLLMs and provides the tools to study it.
Abstract: Multi-view bimanual robot world models must predict future observations while preserving scene identity, following two synchronized but non-exchangeable arms, and maintaining consistency between global and wrist-local views. Existing conditioning schemes often collapse left- and right-arm actions into a view-agnostic control signal, making long-horizon rollout prone to scene drift, wrong-arm responses, and cross-view inconsistency. We propose HAI, a Hierarchical Anchored Interaction architecture for controllable multi-view bimanual world modeling. HAI organizes prediction as a structured information flow. First, hierarchical action-view conditioning routes the bimanual action chunk to camera streams, injects coarse chunk-level intent, performs structured multi-view modeling on the coarse-modulated features, and then injects fine per-frame action control. Second, anchored dynamic generation fuses the fine action-conditioned descriptors with persistent per-view scene anchors before decoding future observations. Experiments on AgiBot and DROID show that HAI improves long-horizon rollout quality over the baselines. Ablations confirm gains in scene stability, wrist-view controllability, and cross-view consistency. Further experiments on policy-improvement diagnostics indicate more useful synthetic rollout signals for downstream VLA policies, achieving an averaging 67.2% relative gain in success rate. The project page is available at \urlhttps://hai-anon.github.io/.
Abstract: Video Language Models enable AI systems to understand temporal dynamics in videos. To fit within the maximum context window constraint, current methods use keyframe sampling which often misses both macro-level events and micro-level details due to the sparse temporal coverage. Furthermore, processing full images and their tokens for each frame incurs substantial computational overhead. We address these limitations by leveraging video codec primitives (specifically motion vectors and residuals) which natively encode video redundancy and sparsity without requiring expensive full-image encoding for most frames. To this end, we introduce lightweight transformer-based encoders that aggregate codec primitives and align their representations with image encoder embeddings through a pre-training strategy that accelerates convergence during end-to-end fine-tuning. Our approach, , reduces the time-to-first-token by up to 86% and token usage by up to 93% compared to standard VideoLMs. Moreover, by varying the keyframe and codec primitive densities, we maintain or exceed performance compared on 14 diverse video understanding benchmarks, spanning general question answering, temporal and motion reasoning, long-form understanding, and spatial scene understanding.
Authors: Tongyang Dai, Ruipeng Gao, Feng Liu, Wenqiu Bai, Lingkun Li, Qiang Ni
Abstract: Compared with residential layouts, public indoor floorplans present more substantial challenges with open corridors, complex room relations, and non-Manhattan geometries, which do not conform to conventional residential-layout assumptions. Consequently, existing methods are prone to boundary discontinuities, leading to fragmented corridors and irregular wall structures. To address the above issues, we propose LayoutBridge, a structure-conditioned generative framework for floorplan production. By moving beyond the Manhattan-geometry assumptions commonly adopted in residential scenarios, LayoutBridge reformulates the task as a latent-space Brownian-bridge diffusion process from structural constraints to target layouts. We have constructed our PISF dataset with 3,517 representative floorplans in public indoor spaces, and LayoutBridge reduces FID by 132.35 and improves BIoU by 15.52% over the latest image-to-image baselines. Regarding the residential benchmarks MSD and RPLAN, it also reduces FID by at least 2.53 compared with state-of-the-art. These results demonstrate its potential for scalable automated architectural design. Code implementing the proposed method is publicly available at \hrefhttp://github.com/lalalalaxxx/LayoutBridgehttp://github.com/lalalalaxxx/LayoutBridge.
Abstract: Patch-based transformers have emerged as efficient and improved long-horizon modeling architectures for time series modeling. Yet, existing approaches rely on temporally-agnostic patch construction, where arbitrary starting positions and fixed lengths fracture temporal coherence by splitting natural transitions across boundaries. This naive segmentation often disrupts short-term dependencies and weakens representation learning. We propose a novel Entropy-Guided Dynamic Patch Encoder (EntroPE), as a temporally informed framework that dynamically detects transition points via conditional entropy and dynamically places patch boundaries. This preserves temporal structure while retaining the computational benefits of patching. EntroPE consists of two key modules, namely an Entropy-based Dynamic Patcher (EDP) that applies information-theoretic criteria to locate natural temporal shifts and determine patch boundaries, and an Adaptive Patch Encoder (APE) that employs pooling and cross-attention to capture intra-patch dependencies and produce fixed-size latent representations. Extensive experiments on long-term forecasting, classification, and anomaly detection demonstrate that the proposed method improves both accuracy and efficiency, establishing entropy-guided dynamic patching as a promising new paradigm for time series modeling.
Abstract: Speculative decoding accelerates LLM inference but suffers from performance degradation when target models are fine-tuned for specific domains. A naive solution is to retrain draft models for every target model, which is costly and inefficient. To address this, we introduce a practical parameter- and data-efficient framework named Efficient Draft Adaptation, abbreviated as EDA, for efficiently adapting draft models. EDA introduces three innovations: (1) a decoupled architecture that utilizes shared and private components to model the shared and target-specific output distributions separately, enabling parameter-efficient adaptation by updating only the lightweight private component; (2) a data regeneration strategy that utilizes the fine-tuned target model to regenerate training data, thereby improving the alignment between training and speculative decoding, leading to higher average acceptance length; (3) a sample selection mechanism that prioritizes high-value data for efficient adaptation. Our experiments show that EDA effectively restores speculative performance on fine-tuned models, achieving superior average acceptance lengths with significantly reduced training costs compared to full retraining. Code will be released upon acceptance.
Abstract: Video-based world models hold significant potential for generating high-quality embodied manipulation data. However, current video generation methods struggle to achieve stable long-horizon generation: classical diffusion-based approaches often suffer from temporal inconsistency and visual drift over multiple rollouts, while autoregressive methods tend to compromise on visual detail. To solve this, we introduce LongScape, a hybrid framework that adaptively combines intra-chunk diffusion denoising with inter-chunk autoregressive causal generation. Our core innovation is an action-guided, variable-length chunking mechanism that partitions video based on the semantic context of robotic actions. This ensures each chunk represents a complete, coherent action, enabling the model to flexibly generate diverse dynamics. We further introduce a Context-aware Mixture-of-Experts (CMoE) framework that adaptively activates specialized experts for each chunk during generation, guaranteeing high visual quality and seamless chunk transitions. Extensive experimental results demonstrate that our method achieves stable and consistent long-horizon generation over extended rollouts. Our code is available at: https://anonymous.4open.science/r/AMSVVD-fdg245.
Abstract: Inherent temporal heterogeneity, such as varying sampling densities and periodic structures, has posed substantial challenges in zero-shot generalization for Time Series Foundation Models (TSFMs). Existing TSFMs predominantly rely on massive parameterization to absorb such heterogeneity, as their static tokenization and positional encoding schemes entangle diverse temporal patterns into a fixed representation space, encouraging memorization rather than adaptation. To address this limitation, we propose Kairos, a flexible and parameter-efficient TSFM dedicated to forecasting tasks, which decouples temporal heterogeneity from model capacity through a novel tokenization perspective. Kairos introduces a dynamic patching tokenizer and a mixture-of-size encoding that adapt observational granularity to local information density, enabling fine-grained temporal abstraction without increasing model width or depth. In addition, we design a multi-granularity positional embedding based on dynamic rotary encodings, which conditions on instance-level spectral features and temporal structure induced by dynamic patching tokenization, allowing robust modeling of diverse temporal dependencies. Trained on a novel Predictability-Stratified Time-Series (PreSTS) corpus, Kairos achieves superior zero-shot performance with substantially fewer parameters on two mainstream benchmarks, GIFT-Eval and Time-Series-Library.
Authors: Dan Barry, Andrew Hines
Abstract: Contemporary large language model (LLM) chat systems treat conversation history as an immutable sequence of turns that defines the model’s working context. However, user intent in real interactions is not static: it evolves through correction, refinement, and shifting constraints. This mismatch between dynamic intent and static transcripts can result in context pollution, where outdated or irrelevant information persists and continues to influence subsequent responses. We introduce mutable transcripts, a new interaction paradigm that enables users to revise prior turns through natural language edit requests, allowing the conversation history itself to be updated rather than appended. This reframes the transcript from a passive record into an editable representation of conversational state. We present a working prototype that integrates transcript-level revision into a standard chat interface and evaluate its feasibility through a controlled user study (n=17) and an illustrative transcript analysis of representative interaction scenarios. Participants significantly preferred mutable transcripts over standard chat across measures of clarity, confidence, and ease of use, with reduced intent to restart conversations. Transcript analysis of representative user study conversations shows that mutable transcripts can reduce conversation length and eliminate obsolete retained context. These findings provide initial evidence that user-driven revision of conversational history can improve interaction quality and help maintain a more current representation of user intent.
Abstract: We propose Partition Tree, a tree-based framework for conditional density estimation over general outcome spaces that supports both continuous and categorical variables within a unified formulation. Our approach models conditional distributions as piecewise-constant densities on data-adaptive partitions and learns trees by directly minimizing conditional negative log-likelihood. This yields a scalable, nonparametric alternative to existing probabilistic trees that does not make parametric assumptions about the target distribution. We further introduce Partition Forest, an ensemble extension obtained by averaging conditional densities. Empirically, we demonstrate improved probabilistic prediction over CART-style trees and competitive performance compared to state-of-the-art probabilistic tree methods and Random Forests.
Abstract: Vision-language models (VLMs) deliver strong multimodal reasoning capabilities, but their large computational cost and high parameter counts make deployment challenging on resource-constrained devices. Low-rank decomposition has emerged as a promising compression technique, yet existing methods often optimize local matrix reconstruction error, rely on uniform or heuristic rank allocation, and focus mainly on attention projections while leaving feed-forward networks underexplored. In this paper, we propose~LSVD, a loss-aware low-rank approximation framework for efficient low-precision VLMs. LSVD derives a curvature-weighted SVD objective from a second-order approximation of the model loss and uses Kronecker-factored Fisher information to guide decomposition toward downstream performance rather than reconstruction alone. We further introduce a loss-aware cross-layer rank allocation strategy based on calibration gradients, enabling more effective parameter budgeting across layers. Finally, we extend low-rank compression to FFN layers through a hybrid scheme that combines SVD with quantization. The evaluation results show that LSVD achieves over 2.3× decoding speedup over previous work while preserving strong accuracy under low-precision inference.
Abstract: Depth-of-field control is a fundamental tool in photography, yet post-capture bokeh editing from a single image remains challenging. A practical editor should handle images captured under arbitrary focus and aperture settings. Existing methods typically assume an all-in-focus input, or first recover an all-in-focus image before rendering new bokeh. Such pipelines can discard useful blur cues from the source image and propagate reconstruction artifacts into the final edit. We introduce with a signed circle-of-confusion map and a disparity map. By modeling the linear relation between signed circle of confusion and disparity difference, AnyBokeh estimates a source-specific and transfers the source optical characteristics to the desired focus and aperture setting. A generative editor conditioned on both source and target circle-of-confusion maps then performs relative blur synthesis, enabling spatially adaptive deblurring, preservation, and defocus rendering. To support physically supervised learning, we further construct a high-fidelity synthetic dataset with accurate depth, focus distance, and full EXIF metadata. Experiments on real-world benchmarks show that AnyBokeh achieves faithful and controllable editing across any-to-any bokeh editing, all-in-focus-to-bokeh rendering, and defocus deblurring, while avoiding all-in-focus reconstruction and test-time bokeh-level calibration commonly required by existing approaches. The code and dataset will be made publicly available.
Abstract: In recent years, the state-of-the-art in unsupervised video instance segmentation has heavily relied on synthetic video data, generated from object-centric image datasets such as ImageNet. However, video synthesis by artificially shifting and scaling image instance masks fails to accurately model realistic instance motion in videos, such as perspective changes, movement by parts of one or multiple instances, or camera motion. To tackle this issue, we propose an unsupervised video instance segmentation model trained exclusively on real video data. We start from unsupervised instance segmentation masks on individual video frames. However, single-frame segmentations exhibit temporal noise and their quality varies through the video. Therefore, we establish temporal coherence by identifying high-quality keymasks in the video by leveraging deep motion priors. The sparse keymask pseudo-annotations are then used to train a segmentation model for implicit mask propagation, for which we propose a Sparse-To-Dense Distillation approach aided by a Temporal DropLoss. After training the final model on the resulting dense labelset, our approach outperforms the current state-of-the-art across various benchmarks.
Authors: Debjyoti S Roy, Byron Wallace, Javed A Aslam
Abstract: State-of-the-art Extreme Multi-Label Text Classification (XMC) models rely on multi-label attention to focus on key tokens in input text, but learning high-quality attention weights is challenging. We introduce PLANT (Pretrained and Leveraged Attention), a plug-and-play strategy for initializing attention. PLANT works by \emphplanting label-specific attention using a pretrained Learning-to-Rank model guided by mutual information gain. This architecture-agnostic approach integrates seamlessly with large language model backbones such as Mistral, LLaMA, DeepSeek, and Phi-3. PLANT outperforms state-of-the-art methods across tasks such as ICD coding, legal topic classification, and content recommendation. Gains are especially pronounced in few-shot settings, with substantial improvements on rare labels. Ablation studies confirm that attention initialization is a key driver of these gains. Code and trained models are available at \urlhttps://github.com/research-anon-487/xcube/tree/plant.
Authors:
Zhaochen Su, Jincheng Gao, Hangyu Guo, Zhenhua Liu, Lueyang Zhang, Xinyu Geng, Shijue Huang, Peng Xia, Guanyu Jiang, Cheng Wang, Yue Zhang, Yi R. (May) Fung, Junxian HeAbstract: Real-world multimodal agents solve multi-step workflows grounded in visual evidence. For example, an agent can troubleshoot a device by linking a wiring photo to a schematic and validating the fix with online documentation, or plan a trip by interpreting a transit map and checking schedules under routing constraints. However, existing multimodal benchmarks mainly evaluate single-turn visual reasoning or specific tool skills, and they do not fully capture the realism, visual subtlety, and long-horizon tool use that practical agents require. We introduce AgentVista, a benchmark for generalist multimodal agents that spans 25 sub-domains across 7 categories, pairing realistic and detail-rich visual scenarios with natural hybrid tool use. Tasks require long-horizon tool interactions across modalities, including web search, image search, page navigation, and code-based operations for both image processing and general programming. Comprehensive evaluation of state-of-the-art models exposes significant gaps in their ability to carry out long-horizon multimodal tool use. Even the best model in our evaluation, GPT-5.4 with tools, achieves only 31.58% overall accuracy, and hard instances can extend to 25 tool-calling turns. We expect AgentVista to accelerate the development of more capable and reliable multimodal agents for realistic and ultra-challenging problem solving.
Abstract: Direct Preference Optimization (DPO) and its variants have become the standard for aligning Large Language Models (LLMs). However, we identify two fundamental limitations. First, the optimized policy lacks invariance since it varies with modeling choices such as scalarization function or reference policy, whereas an optimal policy should remain invariant. Second, most existing methods yield theoretically suboptimal policies by not fully exploiting the comparative information in pairwise preference data, thus missing an opportunity for self-reflection through comparing and contrasting responses. To address both limitations, we propose Intrinsic Self-reflective Preference Optimization (InSPO), which derives a globally optimal policy conditioned on both context and alternative response, explicitly formalizing self-reflection. We prove this formulation surpasses standard DPO and RLHF targets while guaranteeing invariance. InSPO serves as a plug-and-play enhancement for DPO-family algorithms, decoupling alignment from modeling constraints without architectural changes. Using privileged information learning, InSPO requires no alternative response at inference since the self-reflective mechanism is distilled during training, incurring zero overhead. Experiments show InSPO consistently improves win rates and length-controlled metrics across DPO variants, yielding more robust and human-aligned LLMs. Code is available
Authors: Jichen Hu, Jiawei Guo, Jiazhong Cen, Chen Yang, Sikuang Li, Wei Shen
Abstract: Recent 3D world modeling systems based on generative scene synthesis, such as Marble, can create coherent and explorable 3D environments, yet their outputs are typically static monolithic assets with limited editability and physical interaction. This restricts their use in immersive content creation and embodied simulation, where generated worlds must be actively modified and manipulated. To tackle this challenge, we present WorldAct, a framework that converts static generated 3D worlds into editable and interaction-ready scenes. WorldAct uses a multimodal agent to guide scene decomposition, identify actionable objects, reconstruct geometrically aligned object-level meshes for interaction, and restore the residual background via 3D inpainting. The resulting scenes support object-level editing, collision-aware manipulation, and embodied task execution while preserving global scene coherence. Experiments show that WorldAct enables richer interaction scenarios than the original generated scenes, suggesting a practical path toward editable and interactive 3D world models.
Abstract: Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations. At its core, UniVR features VR-GRPO, a reinforcement learning paradigm with complementary global and step-level rewards. This approach enforces logical coherence and physical consistency throughout the reasoning process without requiring task-specific heuristics or image-text pairs. To train and evaluate UniVR, we construct VR-X, a large-scale benchmark curated from 12 diverse sources spanning long-horizon manipulation, spatial puzzles, and physical reasoning. It is the first comprehensive suite to assess these heterogeneous capabilities under a purely visual protocol. Remarkably, UniVR achieves up to a 25% improvement on VR-X, and its superior visual reasoning also boosts performance on standard multimodal understanding benchmarks. These findings underscore the vast potential of reasoning within visual spaces, with all code, data, and models are open-sourced for further research.
Abstract: Control variables are widely used in statistical modelling to account for omitted variable bias. However, they have largely been underexplored in deep learning. This is surprising, given that deep learning models encode image-inferable covariates, such as demographic variables, into their predictions when these covariates are correlated with the outcome---a form of omitted variable bias referred to as 'shortcut learning'. While many existing confound-control or fairness methods try to restrict the correlation of such covariates with model predictions, we show that this fails to correct for omitted variable bias. We therefore propose a control variable approach for deep learning models, based on generalised additive modelling of the effects of model inputs and covariates. As flexible additive models can suffer from concurvity, we introduce an estimation procedure that refits the final layer of a pre-trained network to include covariate effects, using cross-fitting with ridge penalisation. We show how these effects can be orthogonalised with respect to covariates to exclude their mediated effects and that model predictions can be marginalised over the covariate distribution to control for their effect. This yields unbiased, interpretable predictions and offers flexibility to model the desired effects depending on the scientific or fairness objective. We verify our approach using simulated images, where it recovers the true covariate effects, even in small samples. Existing methods either require more data or fail to recover the true effects. We apply our method to real neuroimaging data with experimentally induced confounding, where it recovers prediction performance to near the level of a model trained on unconfounded data.
Abstract: Conditional flow matching produces realistic out-of-distribution samples across modalities from images to proteins, yet the conditioning signals that enable extrapolation remain poorly understood for biological time series. While considerable work has focused on areas of biology such as proteins and molecules, generative models of neural time series have been largely restricted to categorical conditioning, which precludes compositional and zero-shot generalization. In this work, we propose a per-timestep conditioned diffusion transformer for generating realistic fMRI brain dynamics during unseen cognitive tasks, by injecting both compositional language and optional spatial priors in-context. Such zero-shot generation would support in-silico task design and counterfactual evaluation of novel cognitive experiments before costly scanner acquisition. Leveraging this model, we evaluate across hundreds of held-out task conditions and characterize predictive performance in relation to the training manifold. From language alone, the model recovers region-specific recruitment across tasks and held-out spatial activation patterns. Spatial priors, when available, complement the text pathway by anchoring generation in regions of task space where language alone degrades, while retaining the compositional structure needed for counterfactual task specification. To our knowledge this is the first generative model of whole-cortex fMRI dynamics for unseen cognitive tasks, enabling counterfactual neuroscience and data-driven experimental design.
Abstract: Most recent real-world image super-resolution (Real-ISR) methods employ pre-trained text-to-image (T2I) diffusion models to synthesize the high-quality image either from random Gaussian noise or directly from the input low-quality image. These approaches train ControlNet or LoRA modules while keeping the pre-trained model fixed, which often introduces over-enhanced artifacts and hallucinations, suffering from limited robustness to inputs with varying degradations. Recent visual autoregressive (AR) models, such as pre-trained Infinity, can provide strong T2I generation capabilities while offering superior efficiency by using the bitwise next-scale prediction strategy. Building upon next-scale prediction, we introduce a robust Real-ISR framework with a generation pathway control mechanism, namely Next-Scale Autoregressive Modeling (NSARM). Specifically, we train NSARM in two stages: a transformation network is first trained to map the input low-quality image to preliminary scales, followed by an end-to-end full-model fine-tuning. Such a comprehensive fine-tuning enhances the robustness of NSARM in Real-ISR tasks without compromising its generative capability. Extensive quantitative and qualitative evaluations demonstrate that as a pure AR model, NSARM achieves superior visual results over existing Real-ISR methods, with fast inference speed and comparable fidelity. Most importantly, it demonstrates much higher robustness to varying input quality, showing stronger generalization performance.
Abstract: Modern systems are often expected to transfer across tasks that were not specified during training, leading to the question: what data facilitates generalization in new, unanticipated settings? One hypothesis is that data with more structural information could contain shared circuits and subprograms that could be recycled in a wider array of downstream settings. Epiplexity, a recently proposed measure of the structural information a compute-bounded learner can extract from data, provides a mechanism to reason about this relationship. In this paper, we show how to operationalize epiplexity as an online training signal for data selection and generation. For selection, we fit scaling laws to the training loss curves of natural data domains to predict the expected epiplexity gain as a function of training tokens, and use this signal to adaptively determine the sampling weights over domains during training. For synthetic data generation, we define a generator's reward as the change in learner epiplexity over a buffer of previously generated data and use REINFORCE policy gradients to guide the generator toward an epiplexity-maximizing distribution. In both cases, higher epiplexity predicts improved downstream performance on zero-shot and fine-tuning based tasks, supporting the hypothesis that data rich in structural information yields representations that transfer across domains.
Abstract: Time Series Foundation Models (TSFMs) are a powerful paradigm for time series analysis and are often enhanced by synthetic data augmentation to improve the training data quality. Existing augmentation methods, however, typically rely on heuristics and static paradigms. Motivated by dynamic data optimization, which shows that the contribution of samples varies across training stages, we propose OATS (Online Data Augmentation for Time Series Foundation Models), a principled strategy that generates synthetic data tailored to different training steps. OATS leverages valuable training samples as principled guiding signals and dynamically generates high-quality synthetic data conditioned on them. We further design a diffusion-based framework to produce realistic time series. Experiments on TSFMs demonstrate that OATS consistently outperforms regular training and yields substantial performance gains over static data augmentation baselines across six validation datasets and two TSFM architectures. The code is available at the link https://anonymous.4open.science/r/OATS-536E.
Authors:
Dengyang Jiang, Xin Jin, Dongyang Liu, Zanyi Wang, Mingzhe Zheng, Ruoyi Du, Xiangpeng Yang, Qilong Wu, Zhen Li, Peng Gao, Harry Yang, Steven HoiAbstract: The landscape of high-performance image generation models is currently shifting from the inefficient multi-step ones to the efficient few-step counterparts (e.g, Z-Image-Turbo and FLUX.2-klein). However, these models present significant challenges for direct continuous supervised fine-tuning. For example, applying the commonly used fine-tuning technique would compromise their inherent few-step inference capability. To address this, we propose D-OPSD, a novel training paradigm for step-distilled diffusion models that enables on-policy learning during supervised fine-tuning. We first find that the modern diffusion models, where the LLM/VLM serves as the encoder, can inherit its encoder's in-context capabilities. This enables us to formulate the training as an on-policy self-distillation process. Specifically, during training, we make the model act as both the teacher and the student with different contexts, where the student is conditioned only on the text feature, while the teacher is conditioned on the multimodal feature of both the text prompt and the target image. Training minimizes the two predicted distributions over the student's own roll-outs. By optimizing on the model's own trajectory and under its own supervision, D-OPSD enables the model to learn new concepts, styles, etc., without sacrificing the original few-step capacity.
Abstract: Multimodal large language models (MLLMs) often answer visual reasoning questions by relying on linguistic priors rather than task-relevant visual evidence. Textual chain-of-thought reasoning can partially mitigate this issue by encouraging models to decompose visual questions into intermediate evidence-seeking steps, but generating these steps autoregressively increases inference cost. Latent reasoning avoids explicit rationale generation, but existing approaches provide limited control over what intermediate states encode, making it difficult to impose separate supervision for planning, grounding, and evidence selection. We propose Structured Latent Visual Reasoning (SLVR), a training framework that bridges explicit chain-of-thought and latent reasoning by organizing multimodal reasoning into typed latent stages for planning, grounding, evidence selection, and reasoning integration. SLVR first trains the model to rely on the image by masking answer-revealing text and contrasting the correct answer with visually plausible distractors. It then organizes reasoning into latent stages for planning, grounding, evidence selection, and integration, supervising each stage with the corresponding signal: plans, boxes, visual evidence, and final rationales. This gives latent reasoning an explicit functional structure while avoiding generated textual chains at inference time. Built on Qwen2.5-VL-7B, SLVR improves consistently across multimodal reasoning benchmarks, with absolute gains of +9.4 on MMVP and +14.2 on BLINK Relation, as well as improvements on V, MathVista, and ChartQA. These results suggest that structured latent supervision can improve fine-grained visual reasoning without the decoding overhead of textual CoT. Code is available \hrefhttps://anonymous.4open.science/r/SLVR-681F/README.mdhere.
Abstract: Due to the potential for exploratory reasoning of Latent Visual Reasoning, recent works tend to enable MLLMs (Multimodal Large Language Models) to perform visual reasoning by propagating continuous hidden states instead of decoding intermediate steps into discrete tokens. However, existing works typically rely on hard alignment objectives to force latent representations to match predefined visual features, thereby severely limiting the exploratory of latent reasoning process. To address this problem, we propose CoLVR (Contrastive Optimization for Latent Visual Reasoning). To obtain a more exploratory visual reasoning, CoLVR introduces a latent contrastive training framework. Firstly, CoLVR learns diverse and exploratory representations with a latent contrastive objective guided by angle-based perturbation, which expands the semantic latent space and avoids over-constrained embedding. Then, CoLVR employs a latent trajectory contrastive reward for RL (Reinforcement Learning) post-training to enable fine-grained optimization of latent visual reasoning process and thus fostering diverse reasoning behaviors. Experiments demonstrate that CoLVR significantly enhances the exploratory capability of latent representations, achieving average improvements of 5.83% on VSP and 8.00% on Jigsaw, while also outperforming existing latent models on out of domain benchmarks, with a 3.40% gain on MMStar.
Authors:
Yu Yang, Yue Liao, Jianbiao Mei, Baisen Wang, Jiangning Zhang, Xiangtai Li, Liang Lv, Hanlin Chen, Yong Liu, Shuicheng Yan, Gim Hee LeeAbstract: Long-horizon action-conditioned video generation aims to synthesize temporally coherent videos that follow complex action instructions over extended horizons. Existing single-shot video generation models typically operate in an open-loop manner, leading to incomplete action execution, hallucinated motions, and temporal drift. To address this, we propose , a closed-loop framework that performs sequential planning and iterative reflection for action-conditioned long-horizon video generation. Specifically, a PlanAgent decomposes a high-level goal into sub-actions that condition video generation, while a CriticAgent evaluates intermediate video segments and provides corrective feedback for iterative refinement. This closed-loop design further supports self-evolving, utilizing planning and verification signals for GRPO-based post-training to enhance the video generator's consistency and action quality over extended horizons. Moreover, we introduce for training and evaluation. Experiments across multiple TI2V backbones with self-evolving show consistent gains on ActVideoGen-Bench and VBench, demonstrating the effectiveness of SPIRAL.
Authors:
Tianfu Wang, Leilei Ding, Ziyang Tao, Yi Zhan, zhiyuan, Wei Wu, Yuxuan Lei, Junyang Wang, Yizhao Xu, Yin WU, Zhengyu Hu, Hongyuan Zhu, Nicholas Jing Yuan, Yanyong Zhang, Hui XiongAbstract: High-fidelity diagram creation requires coordinated visual-spatial decisions over semantic topology, visual styling, and geometric layout. Existing methods face a representation gap: pixel-based generation offers limited object-level control, while code-based synthesis provides executable structure at the cost of intuitive manipulation. We introduce EvoDiagram, an agentic framework for editable diagram creation through an intermediate canvas schema that is both machine-actionable and directly manipulable by users. EvoDiagram first constructs a diagram manifest through coordinated structure, style, and layout agents, and then renders the manifest in a diagnostic verification environment. This environment produces objective defect signals and VLM-based critique, enabling localized refinement of the current diagram while also providing evidence for long-term design expertise evolution. The evolution mechanism distills refinement traces into a three-tier design expertise memory, where candidate strategies are linked to the verification evidence that supports them and are promoted into broader guidelines and principles only through repeated cross-context support. We further define CanvasBench, a canvas-recoverable benchmark and evaluation protocol with content, visual, and cognitive dimensions. The current manuscript focuses on the framework, benchmark design, and reporting protocol; empirical results should be added only when the corresponding artifacts, judge traces, and analysis scripts are available for verification. Our code is available at \hrefhttps://anonymous.4open.science/r/evo-diagram/https://anonymous.4open.science/r/evo-diagram/.
Abstract: Federated vision–language models (FedVLMs) represent one of the emerging frontiers of machine learning, aiming to integrate federated learning within the fine-tuning pipelines of vision–language models (VLMs). In this work, we introduce BLEND, a new framework designed to jointly enhance the personalization and generalization capabilities of FedVLMs. BLEND proposes a selective VLM parameter fine-tuning and aggregation strategy that consists of (i) global vision and text adapters shared across clients, and (ii) a personalized vision projection adapter tailored to each client. BLEND further introduces a novel loss function that blends the outputs of the global and personalized adapters to achieve personalization, while simultaneously promoting generalization through a principled anchor design implemented via the projection heads of the zero-shot VLM. Extensive experiments demonstrate that BLEND consistently outperforms state-of-the-art baselines, highlighting its ability to effectively balance personalization and generalization across diverse system settings. (Code is available at: https://github.com/blendfedvlm/BLEND.git)
Authors:
TANUJ SUR, Shashank Tripathi, Nikos Athanasiou, Ha Linh Nguyen, Kai Xu, Michael Black, Angela YaoAbstract: We introduce UniCon3R, a unified feed-forward framework for online human-scene 4D reconstruction from monocular video. Current feed-forward human-scene reconstruction methods suffer from artifacts, where bodies float above the ground or penetrate parts of the scene. A key reason is the lack of effective interaction modelling between the human and the environment. Our goal is to exploit contact between the human and the scene during inference to actively improve the human mesh reconstruction. To that end, we explicitly model interaction by inferring 4D contact from the human pose and scene geometry and use the contact as a corrective cue for generating the pose. This enables UniCon3R to jointly recover scene geometry and spatially aligned 4D humans within the scene. Experiments on standard human-centric video benchmarks show that UniCon3R outperforms state-of-the-art baselines on physical plausibility and global human motion estimation while preserving fast, feed-forward inference speeds. The results validate our central claim: contact serves as a powerful internal prior, thus establishing a new paradigm for physically grounded joint human-scene reconstruction. Source code and models will be released upon acceptance.
Abstract: 3D Gaussian Splatting (3DGS) enables high-quality real-time novel-view synthesis, but practical scenes often contain millions of Gaussians, making compression essential for deployment on limited hardware. Existing reduction methods are effective but mostly heuristic: they provide no multiplicative approximation guarantee for the rendered objective, and thus rely heavily on costly post-pruning finetuning to recover quality. We ask a basic question: can a 3DGS scene be provably replaced by a much smaller weighted subset (coreset) while preserving the objective of interest? We first show that, in the unrestricted setting, no non-trivial multiplicative 3DGS coreset exists. We then show that multiplicative guarantees are not impossible, but resolution-dependent. For a prescribed rendering resolution, such as representative views or grids of views/rays, we provide the first weighted coreset construction theorem for 3DGS. The construction samples Gaussians by sensitivity: provable importance scores measuring each Gaussian’s role in the full-scene objective. Finally, under explicit validity and log-transmittance stability assumptions, we turn this objective guarantee into a rendering guarantee. Empirically, our method is strongest where deployment needs it most: aggressive compression with no or minimal recovery compute. In prune-only and very short finetuning regimes, it achieves state-of-the-art performance, showing that principled importance estimation can be both theoretically meaningful and practically useful.
Abstract: Monocular geometry estimation has recently achieved impressive performance across diverse scenes. However, state-of-the-art models still face notable distortion in local 3D structure, especially in fine details, like thin structure and small objects. We attribute this limitation to an architectural mismatch: most current models decode 3D geometry within a 2D parameterization, where feature interactions are governed by image-plane proximity rather than true 3D spatial relationships. This inadvertently mixes features from geometrically distant surfaces, resulting in over-smoothed geometry particularly around thin or elongated structure. In this paper, we propose a fine-detail monocular geometry estimation with Self-Guided Sparse 3D Refinement (SSR) that lifts monocular geometry modeling from 2D image space to 3D space for high-fidelity metric-scale point map. Our model lifts the coarse point map from a foundation base model onto a sparse voxel shell and refines it via SSR. The SSR employs sparse convolutions that aggregate features based on 3D spatial locality, avoiding feature mixing across depth discontinuities. Extensive experiments on diverse datasets demonstrate that our method significantly outperforms existing approaches in recovering fine detailed 3D geometry across both quantitative metrics and qualitative visualizations.
Authors:
Shuang Chen, Quanxin Shou, Hangting Chen, Yucheng Zhou, Kaituo Feng, Wenbo Hu, yifan zhang, Yunlong Lin, Wenxuan Huang, Mingyang Song, Dasen Dai, Bolin Jiang, Manyuan Zhang, Yu Cheng, Nanyun PengAbstract: Unified multimodal models provide a natural and promising architecture for understanding diverse and complex real-world knowledge while generating high-quality images. However, they still rely primarily on frozen parametric knowledge, which makes them struggle with real-world image generation involving long-tail and knowledge-intensive concepts. Inspired by the broad success of agents on real-world tasks, we explore agentic modeling to address this limitation. Specifically, we present Unify-Agent, a unified multimodal agent for world-grounded image synthesis, which reframes image generation as an agentic pipeline consisting of prompt understanding, multimodal evidence searching, grounded recaptioning, and final synthesis. To train our model, we construct a tailored multimodal data pipeline and curate 143K high-quality agent trajectories for world-grounded image synthesis, enabling effective supervision over the full agentic generation process. We further introduce FactIP, a benchmark covering 12 categories of culturally significant and long-tail factual concepts that explicitly requires external knowledge grounding. Extensive experiments show that our proposed Unify-Agent substantially improves over its base unified model across diverse benchmarks and real world generation tasks, while approaching the world knowledge capabilities of the strongest closed-source models. As an early exploration of agent-based modeling for world-grounded image synthesis, our work highlights the value of tightly coupling reasoning, searching, and generation for reliable open-world agentic image synthesis.
Authors:
Xuan Li, Yining Wang, Haocai Luo, Shengping Liu, Jiaen Liang, Ying Fu, Wei Huang, Jun Yu, Junnan ZhuAbstract: Retrieval-Augmented Generation (RAG) has become a pivotal paradigm for Large Language Models (LLMs), yet current approaches struggle with visually rich documents by treating text and images as isolated retrieval targets. Existing methods relying solely on cosine similarity often fail to capture the semantic reinforcement provided by cross-modal alignment and layout-induced coherence. To address these limitations, we propose , a novel multimodal retrieval framework grounded in Bayesian inference and Dempster-Shafer evidence theory. Unlike traditional approaches that rank candidates strictly by similarity, BayesRAG models the intrinsic consistency of retrieved candidates across modalities as probabilistic evidence to refine retrieval confidence. Specifically, our method computes the posterior association probability for combinations of multimodal retrieval results, prioritizing text-image pairs that mutually corroborate each other in terms of both semantics and layout. Extensive experiments demonstrate that BayesRAG significantly outperforms state-of-the-art (SOTA) methods on challenging multimodal benchmarks. This study establishes a new paradigm for multimodal retrieval fusion that effectively resolves the isolation of heterogeneous modalities through an evidence fusion mechanism and enhances the robustness of retrieval outcomes. Our code is available at https://anonymous.4open.science/r/BayesRAG-4C6B/.
Abstract: Modifying 3D assets requires enacting targeted structural transformations while strictly preserving the geometric identity of unedited regions. While recent probability flow models enable training-free editing, bounding these modifications within a high-dimensional latent space presents severe mathematical challenges. Current inversion-free trajectories rely on global velocity updates, which consistently result in spatial bleed into preserved structures and induce trajectory under-commitment, causing edits to stall mid-path. To resolve these dual failures, we introduce FocusFlow, a segmentation-free framework that structures the integration trajectory into two disjoint phases. During early coarse structure formation, FocusFlow applies Feature-Divergence Localization (FDL) to spatially gate the velocity field using intrinsic cross-attention. During late-stage fine-detail refinement, Reference-Guided Generation (RGG) drives topological convergence by pulling the trajectory toward an explicit clean-latent target. Evaluated on the joint Google Scanned Objects and PartObjaverse-Tiny benchmarks, FocusFlow breaks the traditional preservation-modification trade-off. It achieves superior structural preservation while maximizing edit magnitude. By providing stable, localized geometric control without manual spatial annotations, this phase-scheduled approach removes significant technical barriers to the precise modification of synthetic 3D media.
Abstract: Large Reasoning Models (LRMs) are commonly trained with reinforcement learning (RL) to improve their ability to generate chain-of-thought (CoT) reasoning before producing final answers. However, RL rewards are typically assigned based on final answers, providing little or no direct supervision over intermediate reasoning. This can lead to deceptive safety alignment, where the reasoning trace and final answer convey inconsistent safety signals. To systematically investigate this phenomenon, we introduce DSAR (Deceptive Safety Alignment Rate), a metric that jointly assesses reasoning traces and final answers to quantify their safety inconsistency. Across multiple LRMs and benchmarks, we find that deceptive safety alignment is pervasive under standard prompting conditions and is substantially amplified under prefilling attacks. We further provide a hidden representation analysis showing that models exhibit stronger safety discrimination at the final-answer stage than during intermediate reasoning. To close this gap, we propose SARA (Safety-Aware Reasoning Alignment), an RL-based method that rewards both safety-aware reasoning and safe final answers, encouraging early harmful intent recognition and enforcing reasoning–answer consistency. Experiments show that SARA significantly mitigates deceptive safety alignment under both standard and adversarial settings while preserving helpfulness and utility.
Abstract: Reasoning with large language models often benefits from generating multiple chains-of-thought, but existing aggregation strategies are typically trajectory-level (e.g., selecting the best trace or voting on the final answer), discarding useful intermediate work from partial or “nearly correct” attempts. We propose Stitching Noisy Diffusion Thoughts, a self-consistency framework that turns cheap diffusion-sampled reasoning into a reusable pool of step-level candidates. Given a problem, we (i) sample many diverse, low-cost reasoning trajectories using a masked diffusion language model, (ii) score every intermediate step with an off-the-shelf process reward model (PRM), and (iii) stitch these highest-quality steps across trajectories into a composite rationale. This rationale is then used to recompute only the final answer. This modular pipeline separates exploration (diffusion) from evaluation and solution synthesis, avoiding monolithic unified hybrids while preserving broad search. Across math reasoning benchmarks, we find that step-level recombination is most beneficial on harder problems, and ablations highlight the importance of the final solver in converting stitched but imperfect rationales into accurate answers. Using low-confidence diffusion sampling with parallel, independent rollouts, our training-free framework improves average accuracy by up to 23.8% across six math and coding tasks. At the same time, it achieves up to a 1.8× latency reduction relative to both traditional diffusion models (e.g., Dream, LLaDA) and unified architectures (e.g., TiDAR). The code will be publicly available.
Abstract: Large language model (LLM)-based search agents trained with reinforcement learning (RL) have significantly improved the performance of knowledge-intensive tasks. However, existing methods encounter critical challenges in long-horizon credit assignment: (i) Reward Sparsity, where models receive only outcome feedback without step-level guidance to differentiate action quality; (ii) Isolated Credit, where credit is assigned to steps independently, failing to capture sequential dependencies; and (iii) Distributional Shift, where rewards are estimated on templates that deviate from the model’s natural generative distribution. To address these issues, we propose Pivot-Based Credit Assignment (PiCA), a novel step reward mechanism that reformulates the search trajectory as a sequential process of cumulative information gain. Unlike prior methods that reward steps in isolation, PiCA defines process rewards as success probabilities dependent on the historical context. This approach identifies pivot steps (critical milestones in the search) as information peaks that significantly boost the likelihood of a correct final answer. By anchoring these step rewards to the final task objective, PiCA provides dense, pivot-aware and trajectory-dependent guidance while maintaining distributional consistency. Theoretically grounded in Potential-Based Reward Shaping (PBRS), PiCA consistently outperforms existing strong baselines across seven knowledge-intensive QA benchmarks, achieving 15.2% and 2.2% improvements for 3B and 7B models.
Abstract: Dense 3D tracking from monocular video is fundamental to dynamic scene understanding. While recent 3D foundation models provide reliable per-frame geometry, recovering object motion in this geometry remains challenging and benefits from strong motion priors learned from real-world videos. Existing 3D trackers either follow iterative paradigms trained from scratch on synthetic data or fine-tune 3D reconstruction models learned from static multi-view images, both lacking real-world motion priors. Pre-trained video diffusion transformers (video DiTs) offer rich spatio-temporal priors from internet-scale videos, making them a promising foundation for 3D tracking. However, their frame-anchored formulation, which generates each frame's content, is fundamentally mismatched with reference-anchored dense 3D tracking, which must follow the same physical points from a reference frame across time. We present TrackCraft3R, the first method to repurpose a video DiT as a feed-forward dense 3D tracker. Given a monocular video and its frame-anchored reconstruction pointmap, TrackCraft3R predicts a reference-anchored tracking pointmap that follows every pixel of the first frame across time in a single forward pass, along with its visibility. We achieve this through two designs: (i) a dual-latent representation that uses per-frame geometry latents and reference-anchored track latents as dense queries, and (ii) temporal RoPE alignment, which specifies the target timestamp of each track latent. Together, these designs convert the per-frame generative paradigm of video DiTs into a reference-anchored tracking formulation with LoRA fine-tuning. TrackCraft3R achieves state-of-the-art performance on standard sparse and dense 3D tracking benchmarks, while running 1.3x faster and using 4.6x less peak memory than the strongest prior method. We further demonstrate robustness to large motions and long videos. Our code and weights will be publicly released.
Authors: Mehrzad Karamimanesh, Marcel A. J. van Gerven, Mahyar Shahsavari
Abstract: Backpropagation imposes fundamental constraints on training neural networks. It requires the retention of all forward activations until the corresponding backward pass completes, thereby preventing pipelined execution across layers. This work introduces Backpropagation-Free Local Antithetic Node Perturbation (BLANP), a fully local training framework that eliminates gradient propagation across layer boundaries while requiring no negative samples. Each layer updates its parameters from locally available signals through a zeroth-order antithetic perturbation estimator. To stabilize learning, the estimator is evaluated against a frozen exponential moving average critic, decoupling backbone updates from the concurrently evolving local classifiers. The antithetic perturbation formulation cancels even-order bias terms in the gradient estimate, yielding a first-order accurate update from two forward passes alone. A companion variant, BLANP-Exact, replaces the perturbation estimator with exact yet local gradients under a strict stop-gradient constraint, preserving a memory footprint that does not grow with network depth while recovering gradient fidelity. Experiments on MNIST, Fashion-MNIST, and CIFAR-10 across MLP and convolutional architectures show that BLANP matches backpropagation on simple benchmarks. On the more challenging CIFAR-10, BLANP-Exact surpasses the best prior local learning method by 2.40% with a simpler architecture, and trails an identical backpropagation baseline by only 3.90%. The code is available at https://github.com/anonymous/BLANP.
Abstract: Most Large Language Model (LLM) agent memory systems rely on a small set of static, hand-designed operations for extracting memory. These fixed procedures hard-code human priors about what to store and how to revise memory, making them rigid under diverse interaction patterns and inefficient on long histories. To this end, we present MemSkill, which reframes these operations as learnable and evolvable memory skills, structured and reusable routines for extracting, consolidating, and pruning information from interaction traces. Inspired by the design philosophy of agent skills, MemSkill employs a \emphcontroller that learns to select a small set of relevant skills, paired with an LLM-based \emphexecutor that produces skill-guided memories. Beyond learning skill selection, MemSkill introduces a \emphdesigner that periodically reviews hard cases where selected skills yield incorrect or incomplete memories, and evolves the skill set by proposing refinements and new skills. Together, MemSkill forms a closed-loop procedure that improves both the skill-selection policy and the skill set itself. Experiments on LoCoMo, LongMemEval, HotpotQA, and ALFWorld demonstrate that MemSkill improves task performance over strong baselines and generalizes well across settings. Further analyses shed light on how skills evolve, offering insights toward more adaptive, self-evolving memory management for LLM agents.
Authors:
Jaewoo Lee, Hyeongyu Kang, Dohyun Kim, Kyuil Sim, Woocheol Shin, Minsu Kim, Taeyoung Yun, Jeongjae Lee, Sanghyeok Choi, Tabitha Edith Lee, Jong Chul Ye, Jinkyoo ParkAbstract: Aligning a few-step generative model is challenging, since existing alignment frameworks typically rely on restrictive assumptions: a tractable likelihood, a specific ODE/SDE solver, or a particular model family. We introduce FAV, Few-step Generative Models Alignment via Sample-based Variational Inference, a general alignment framework that requires only sample access to the generator and the reference distribution. We cast alignment as sampling from a reward-tilted distribution anchored to a reference distribution. We leverage Stein Variational Gradient Descent as a sample-based variational inference scheme and amortize its particle updates into the parameters of the generator via fixed-point regression. We evaluate FAV on two domains: robotics manipulation and image generator alignment. On generative policy alignment for robotic manipulation, FAV outperforms prevailing policy extraction baselines across 56 offline and 30 offline-to-online RL tasks. For image generator alignment, FAV fine-tunes diverse few-step backbones, including GAN, drifting model, consistency models, and flow maps, scaling from ImageNet-256 to 1024x1024 text-to-image synthesis.
Abstract: Cross-view object correspondence identifies the same object across views. However, existing methods treat this as a frame-level matching problem, necessitating a query mask at every frame. Furthermore, without modeling temporal context, they suffer from severe ambiguity under appearance changes or occlusions. Therefore, we introduce cross-view temporal object correspondence, which requires preserving object identity across both views and time from a query mask. To address this, we propose Track, a streamlined framework which repurposes a pretrained video DiT with semantic and spatiotemporal priors for RGB-to-mask latent transport, jointly modeling cross-view matching and temporal tracking. However, naive video DiTs do not guarantee accurate cross-view grounding and often attend to semantically similar yet incorrect objects across views. Thus, we introduce focal cross-view attention alignment, hard negative conditioning, and local context crop strategies to ensure reliable grounding. Finally, beyond these architectural enhancements, we also tackle the lack of dense annotation by introducing a pseudo-label curation pipeline that converts sparse Ego-Exo4D annotations into dense supervision to enable effective video-level training. Our highly competitive performances validate the effectiveness of Track as a streamlined framework, ensuring robust temporal consistency and identity preservation. Code and weights will be released.
Abstract: Whether navigating a building, operating a robot, or playing a game, an agent that acts effectively in an environment must first learn an internal model of how that environment works. Partially observable Markov decision processes (POMDPs) provide a flexible modeling class for such internal world models, but learning them from observation-action trajectories alone is challenging and typically requires extensive environment interaction. We ask whether language-model priors can reduce costly interaction by leveraging prior knowledge, and introduce Pinductor (POMDP-inductor): an LLM proposes candidate POMDP models from a few observation-action trajectories and iteratively refines them to optimize a belief-based likelihood score. Despite using strictly less information, Pinductor matches the performance and sample efficiency of LLM-based POMDP learning methods that assume privileged access to the hidden state, while significantly surpassing the sample efficiency of tabular POMDP baselines. Further results show that performance scales with LLM capability and degrades gracefully as semantic information about the environment is withheld. Together, these results position language-model priors as a practical tool for sample-efficient world-model learning under partial observability, and a step toward generalist agents in real-world environments.
Abstract: Multiphase powder X-ray diffraction (PXRD) analysis remains a fundamental bottleneck in structure identification, as real-world synthesis often produces complex mixtures whose constituent phases (components) cannot be reliably disentangled. While recent advances in representation-based crystal retrieval and generation suggest the possibility of inferring structures directly from PXRD, existing approaches largely assume single-phase inputs and break down in multiphase settings. Here, we present XDecomposer, a prior-free framework for joint decomposition and identification of multiphase XRD patterns without requiring candidate phase lists, structural templates, or prior knowledge of phase number. We formulate multiphase diffraction analysis as a set prediction problem, where the model infers an unordered set of phase-resolved components, their mixture proportions, and corresponding structural representations within a unified architecture. A phase-query-driven decomposition mechanism, together with diffraction-consistent physical reconstruction, enables accurate source separation while preserving crystallographic fidelity. Extensive experiments on both simulated and experimental datasets show that XDecomposer substantially improves reconstruction accuracy and phase identification across diverse chemical systems, while maintaining strong generalization to unseen mixtures. These results provide a practical route toward data-driven, source-resolved multiphase XRD analysis and reduce long-standing dependence on prior-guided iteratively phase matching. The code is openly available at \urlhttps://anonymous.4open.science/r/XDecomposer-FE0B
Abstract: Understanding human gaze behavior is essential for complex scene comprehension and human-computer interaction. Traditional gaze following models are typically restricted to pure spatial localization, lacking the high-level capacity to reason about semantic targets or complex social contexts. Furthermore, these models often process individuals sequentially, requiring redundant computations over the same scene image for multi-person inference. While recent Vision-Language Models (VLMs) offer the exceptional semantic reasoning needed to address gaze-related semantic tasks, their reliance on discrete text generation inherently limits precision in continuous spatial tasks like gaze localization. To bridge this gap, we propose OmniGF, a unified vision-language framework that adapts foundational VLMs for highly scalable multi-person gaze reasoning. The model adopts a dual-branch decoding strategy: a structured language branch generates discrete reasoning states, while a continuous spatial branch directly taps into the VLM's dense hidden states. Supervising these extracted representations with high-resolution gaze target heatmaps effectively overcomes the spatial bottleneck of text-only coordinate generation. Furthermore, to explicitly ground the model in multi-person scenes, we augment the input with head embeddings encoded from cropped head images, providing fine-grained appearance and orientation cues for all individuals simultaneously. By modeling all individuals and leveraging the strong semantic capability of VLMs, OmniGF seamlessly integrates precise spatial gaze target estimation, semantic gaze prediction, and complex social gaze reasoning. Extensive experiments demonstrate that our framework establishes new state-of-the-art performance across multiple standard benchmarks.
Authors:
Haiyang SHEN, Taian Guo, Xuanzhong Chen, Mugeng Liu, Sixiong Xie, Zhuofan Shi, Chongyang Pan, Siqi Zhong, Guoqing Wang, Ming Zhang, Yun MaAbstract: Although LLMs have made substantial progress in reasoning, systematically producing frontier-level reasoning data remains difficult. Existing synthesis methods often have limited visibility into the structural factors that govern problem difficulty, which can result in narrow diversity and unstable difficulty control. In this work, we view the difficulty of a reasoning problem as arising from the accumulation of atomic knowledge-reasoning transformations, which we term . Building on this perspective, we propose MindLoom, a framework for synthesizing frontier-level reasoning data through compositional thought mode engineering. Given a collection of hard problems with verified solutions, MindLoom first decomposes those solutions into thought mode chains that reveal each problem's construction logic. It then trains a retrieval model that matches problem states to compatible thought modes, providing guidance on which reasoning challenges to introduce during synthesis. New problems are composed by iteratively applying retrieved thought modes to seed questions, with distribution-aligned sampling to encourage diverse reasoning coverage. Finally, a rollout-based judging stage labels generated questions by difficulty and supplies judged-correct responses for supervised fine-tuning. We evaluate MindLoom on nine benchmarks covering five STEM disciplines and four mathematical reasoning tasks across multiple model families and sizes. Models fine-tuned on MindLoom generated data consistently improve over base models, distillation, and external-data baselines across the reported benchmarks. Ablation studies indicate the contribution of each component, and further analysis suggests that MindLoom covers a broad range of reasoning patterns while maintaining useful difficulty control. We have open-sourced our implementation at https://anonymous.4open.science/r/MindLoom-5B8E.
Abstract: Language models are becoming the default interface to factual knowledge, yet they often verify outputs more reliably than they generate them. This generation-verification gap (GV-gap) underlies many recent advances in self-improvement and reasoning, but its dynamics on factual knowledge specifically remain poorly understood. We focus on the training mechanisms underlying GV-gaps, distinguishing them from their computational and aesthetic counterparts. We trace generation and verification capabilities through three training phases (acquisition, continual learning, and updating) across four open-source model families at two scales each. Three findings recur across models: (i) verification is consistently learned before generation; (ii) verification is more robust to continual learning than generation; and (iii) factual updates can leave models in a state, simultaneously verifying both old and new answers as correct. Natural experiments on frontier models reproduce these dynamics at scale and reveal residual verification biases on well-covered facts.
Abstract: 3D Gaussian Splatting (3DGS) has become a vital tool in novel view synthesis. 3D Gaussian functions with attributes like color, opacity, and scale are stacked together to explicitly model a radiance field. However, discrete 3D Gaussians usually produce poor depth with few geometry details but lots of floaters, which makes modeling smooth radiance fields remain a challenge. To address this issue, we propose to learn a radiance field within a continuous view-dependent color field parameterized by neural networks. With the continuous constraints in color, the bias on low-frequency signals of neural networks highly encourages the geometry to contribute to high-frequency color variations on images, but not solely relying on the color attribute itself. Moreover, we impose a constraint to improve the Lipschitz continuity of the neural color field, making large changes of each Gaussian's color also match with large changes of its position in terms of a ratio. To this end, we additionally introduce novel self-learning strategies to learn neural view-dependent color fields without any external knowledge or priors, aiming for better generalization. Our numerical and visual comparisons on widely used benchmarks justify our idea and show better ability of high fidelity geometry recovery in 3DGS.
Authors:
Zixin Ding, Shaghayegh Emami, Giovanna Salvi, Cecilia Tosciri, Abhijith Gandrakota, Jennifer Ngadiuba, Nhan Tran, Christian Herwig, David W Miller, Yuxin ChenAbstract: High-throughput scientific facilities such as the Large Hadron Collider depend on real-time event filtering (triggering) under tight constraints on bandwidth, latency, and storage. In practice, trigger menus are largely static and hand-tuned and can become suboptimal as detector conditions, pileup, and background composition drift over time. We cast online threshold tuning as a sequential decision-making problem: a reinforcement learning agent ingests streaming summaries of recent rates and signal-sensitive features and updates trigger thresholds to maximize signal efficiency while tracking a target background rate within a tolerance band. We adapt Group-Filtered Policy Optimization (GFPO) to streaming control and introduce two variants (GFPO-F, GFPO-FR) that enforce background rate feasibility during training. On a benchmark that emulates realistic collider operation, we study two representative triggers: a total transverse energy (H_T) trigger sensitive to pileup variation, and an anomaly-detection (AD) trigger based on reconstruction loss for rare or non-standard signatures. On Monte Carlo streams, our agent increases the fraction of in-tolerance time intervals by 38% (H_T) and 30% (AD), with a cumulative gain of up to 5% in conditional signal efficiency. Transferring from simulation to \emphreal collision data (CMS Run 283408), the same agent, without fine-tuning, achieves a 40% (H_T) and 25% (AD) in-tolerance improvement over baselines, with an additional 1% signal-efficiency gain on both triggers. To our knowledge, this is the first demonstration of RL-based trigger control on real Large Hadron Collider collision data.
Abstract: Quantization has become a standard tool for efficient LLM deployment, especially for local inference, where models are now routinely served at 2-3 bits per parameter. The state of the art is currently split into simple scalar quantization techniques, such as GPTQ or AWQ, which are widely deployed but plateau in accuracy at 3-4 bits per parameter (bpp), and "second-generation" vector- or trellis-quantized methods, such as QTIP, GPTVQ and AQLM, which push the accuracy frontier but are notoriously hard to implement and to scale. In this paper, we ask whether this gap is fundamental, or whether a carefully optimized scalar quantizer can recover most of it. We answer in the affirmative, by introducing GSQ (Gumbel-Softmax Quantization), a post-training scalar quantization method which jointly learns the per-coordinate grid assignments and the per-group scales using a Gumbel-Softmax relaxation of the discrete grid. GSQ matches the cardinality of the relaxation to the small number of levels available in the target bit-width regime (e.g., 3-8 levels for ternary and 3 bpp, respectively), making optimization tractable. Practically, on the standard Llama-3.1-8B/70B-Instruct models, GSQ closes most of the gap between scalar quantization and the QTIP frontier at 2 and 3 bits, while using a symmetric scalar grid with group-wise quantization, and thus remains compatible with existing scalar inference kernels. We further show that the same discrete-assignment optimization can be applied to practical GGUF K-Quant checkpoints: starting from publicly released GGUF models, GSQ improves accuracy while projecting the result back into the same deployment format. Finally, GSQ scales to trillion-scale Mixture-of-Experts models such as Kimi-K2.5, where vector-quantized methods are difficult to apply.
Abstract: Low-Rank Adaptation (LoRA) is the prevailing approach for efficient large language model (LLM) fine-tuning. Building on this paradigm, recent studies have proposed alternative initialization strategies and architectural modifications, reporting substantial improvements over vanilla LoRA. However, these gains are often demonstrated under fixed or narrowly tuned hyperparameter settings, despite the known sensitivity of neural networks to training configurations. In this work, we systematically re-evaluate nine representative LoRA variants alongside vanilla LoRA through extensive hyperparameter searches. Across tasks spanning mathematical reasoning, commonsense reasoning, code generation, and instruction following at diverse model scales, we find that different LoRA methods favor distinct learning rate ranges. Crucially, once learning rates are properly tuned, all methods achieve similar peak performance (within 1-2%), with only subtle rank-dependent behaviors. These results suggest that vanilla LoRA remains a competitive baseline and that improvements reported under a single training configuration may not reflect consistent methodological advantages. Finally, a second-order analysis attributes the differing optimal learning rate ranges to variations in the largest Hessian eigenvalue, aligning with classical learning theories.
Abstract: On-policy self-distillation (OPSD) trains a student on its own rollouts using a privileged teacher, but its standard objective weights all generated tokens equally, implicitly treating the privileged teacher target as equally reliable at every student-visited prefix. Existing entropy-based OPD methods relax this uniformity by modulating token-level supervision with teacher entropy, but high teacher entropy in reasoning has an ambiguous reliability meaning: it can reflect either non-viable uncertainty or benign solution diversity. To identify this phenomenon, we introduce a branch-viability diagnostic. Specifically, we record next-token alternatives from the privileged-answer teacher prompt, force each alternative after the student prompt plus its on-policy spine prefix, and test whether the resulting student-template continuation recovers the correct answer. On Qwen3-4B, we find that an oriented within-sequence position score is the strongest tested predictor of teacher-token reliability, reaching an area-under-ROC-curve (AUROC) of 0.83 with a 95% cluster-bootstrap interval of [0.68, 0.96]; local uncertainty scores are at most 0.58. Motivated by this trajectory-level structure, we propose Position-Weighted On-Policy Self-Distillation (PW-OPSD), which applies an increasing position weight while keeping the same student rollout, privileged teacher pass, and clipped forward-KL target as OPSD. In our comprehensive evaluations with different random seeds, the diagnostic-derived PW-OPSD improves AIME 2024 and AIME 2025 Avg@12 by +1.0 and +1.1 points, and a generalization evaluation on the larger-scale model DeepSeek-R1-Distill-Llama-8B also demonstrates consistent improvements. These results show that teacher-token reliability in reasoning distillation is trajectory-structured and can be utilized without additional teacher computation.
Abstract: As the context length of large language models (LLMs) grows, the key-value (KV) cache memory overhead becomes a critical bottleneck, limiting long-context inference efficiency. While FP8 attention has shown substantial promise in innovations like FlashAttention-3, its integration into the decoding phase of the DeepSeek Multi-head Latent Attention (MLA) architecture presents notable challenges. These challenges include numerical heterogeneity arising from the decoupling of positional embeddings, misalignment of quantization scales in FP8 PV GEMM, and the need for optimized system-level support. In this paper, we introduce SnapMLA, an FP8 MLA decoding framework optimized to improve long-context efficiency through the following hardware-aware algorithm-kernel co-optimization techniques: (i) RoPE-Aware Per-Token KV Quantization: Motivated by our analysis of the heterogeneous quantization sensitivity inherent to the MLA KV cache, this approach preserves the RoPE part in high precision. Furthermore, per-token granularity is employed to align with the autoregressive decoding process and maintain quantization accuracy. (ii) Pre-posed Scale Domain Alignment, where we skillfully project RoPE components into the quantization domain to mitigate pipeline stalls caused by mixed-precision accumulation; (ii) Quantized PV Computation Pipeline Reconstruction: Addresses the misalignment of quantization scales in FP8 PV computation caused by the shared KV structure of the MLA. (iii) End-to-End Dataflow Optimization: Establishes an efficient data read-and-write workflow using specialized kernels, ensuring streamlined data flow and improved performance. Extensive experiments on state-of-the-art MLA LLMs show that SnapMLA achieves up to a 1.91× improvement in throughput on long-output decoding workloads while maintaining near-parity benchmark quality compared with the BF16 baseline on the evaluated reasoning and code-generation benchmarks. The code corresponding to this work is included in the supplementary material.
Abstract: Convolutional neural networks trained on ImageNet are known to rely heavily on local high-frequency texture, an inductive bias that translates into fragile robustness against distribution shifts in real-world environments. Event cameras, in contrast, record only changes in scene brightness and are therefore well suited to capturing contour information; however, due to the absence of diagnostic benchmarks in the event domain, the inductive bias that event-camera data instills in vision models has remained unexplored. In this work, we use knowledge distillation from the event domain to the RGB domain so as to exploit the rich evaluation toolkit available in the RGB domain and systematically dissect this inductive bias. Our experiments show that distillation from the event domain induces, in the RGB domain, notable color invariance, shape bias, and robustness to high-frequency noise. We identify the underlying mechanism as the model suppressing its dependence on high-frequency texture while simultaneously acquiring a strong dependence on edge-based object shape. This hypothesis is supported by changes in how color and spatial information are processed at the early layers, together with a "spectral trade-off" in which robustness to the absence of high-frequency components coexists with vulnerability to contamination of the relied-upon frequency bands and to disruption of geometric structure. We further show that this inductive bias differs markedly from existing robustification methods and that it functions as a strong prior for diverse downstream tasks that demand shape-based reasoning, such as medical imaging. The code and experimental configurations for reproducing our experiments are submitted alongside this paper as supplementary material.
Abstract: Learning control policies for complex, long-horizon tasks is a central challenge in autonomous systems. Signal Temporal Logic (STL) offers an expressive language for specifying such tasks, but its non-Markovian nature and inherent sparse reward make it hard for standard Reinforcement Learning (RL) to solve. Prior RL approaches focus only on limited STL fragments or use robustness scores as sparse rewards. Our new method TGPO (Temporal Grounded Policy Optimization) decomposes STL into timed subgoals and invariant constraints and tackles the problem in a hierarchical fashion. Its high-level module proposes time allocations for subgoals, and the low-level time-conditioned policy learns to achieve the sequenced subgoals using a dense, stage-wise reward. In inference, we use the critic to efficiently sample time allocations and select the most promising assignment for the policy to rollout. We evaluate in five environments, from low-dimensional navigation to manipulation, drone, and quadrupedal locomotion, and TGPO significantly outperforms baselines (especially for high-dimensional and long-horizon cases), with 31.6% higher success rate compared to the best method.
Abstract: Adverse weather removal (AWR) in real-world images remains challenging due to heterogeneous and unseen degradations, while distortion-driven training often yields overly smooth results. We propose low refinement. PVRF introduces an AWR-specific question answering module (AWR-QA) that uses frozen vision--language models (VLMs) to estimate soft probabilities of weather types and low-level attribute scores. These perceptions condition restoration networks via attribute-modulated normalization (AMN) and weather-weighted adapters (WWA), producing an anchor estimate for refinement. We then learn a terminal-consistent residual rectified flow with perception-adaptive source perturbation and a terminal-consistent velocity parameterization to stabilize learning near the terminal regime. Extensive experiments show that PVRF improves both fidelity and perceptual quality over state-of-the-art baselines, with strong cross-dataset generalization on single and combined degradations.
Authors:
Gal Wertheizer, Rom Himelstein, Tomer Peretz, Avi MendelsonAbstract: Adversarial attacks optimized on a single open-weight LLM can transfer to and jailbreak architecturally different models, allowing an attacker with white-box access to one model to compromise independently deployed systems. This creates a shared vulnerability across models, yet existing defenses are not designed for this cross-model threat. We find that transfer aligns with shared internal representation geometry, making it a natural defense target. We find that cross-model transfer aligns with shared internal representation geometry, making it a natural defense target. AnchorRep targets this geometry directly with a lightweight LoRA adapter that pushes the defended model's internal representations of harmful prompts away from those of a frozen anchor model on the same prompts. Training uses a small set of harmful prompts and no adversarial examples. Across five models and four architectural families, AnchorRep reduces cross-model attack success rate to \leq1.1% on 2,000 transferred attacks (0% on two), including the largest drop on Mistral (36% \to 1.1%). Existing defenses can reduce transfer, but only at high cost—either inducing up to 77% degenerate benign output or increasing over-refusal by up to 18%. Because such degenerate benign outputs are not captured by standard refusal-based metrics, we introduce the Benign Garble Rate to quantify them. Our results suggest that cross-model robustness can be achieved by shaping representation geometry, without requiring attack-specific training. \smallskip \noindent\faGithub~Code, configs and logs: \hrefhttps://anonymous.4open.science/r/AnchorRep/README.mdanonymous.4open.science/r/AnchorRep
Authors:
Sathwik Karnik, Juyeop Kim, Sanmi Koyejo, Jong-Seok Lee, Somil BansalAbstract: Text-to-image diffusion models are susceptible to memorization, revealing a fundamental failure to generalize beyond the training set. Current mitigation approaches typically sacrifice image quality or prompt alignment to reduce memorization. To address this, we propose teering (RADS), an offline-trained, inference-time framework that mitigates memorization while maintaining generation fidelity. RADS models the diffusion denoising process as a dynamical system and uses reachability analysis to learn a safety value function that characterizes intermediate latent states along trajectories leading to memorized samples. This motivates a constrained reinforcement learning (RL) formulation, where a policy learns to steer the trajectory away from memorization via minimal perturbations in the caption embedding space. Empirical evaluations show that RADS achieves a stronger empirical Pareto frontier between generation diversity (SSCD), quality (FID), and alignment (CLIP) compared to state-of-the-art inference-time baselines. Crucially, RADS provides robust mitigation without modifying the diffusion backbone, offering a plug-and-play solution for diverse image generation.
Abstract: Agents in real-world settings operate over long and evolving horizons, where information is repeatedly updated and can interfere with each other across memories, requiring accurate recall and aggregated reasoning over multiple pieces of information. However, existing benchmarks focus on static, independent recall and fail to capture these dynamic interactions between evolving memories. In this paper, we study how current systems perform in realistic, continuously evolving, long-horizon settings across diverse domains and question types. To this end, we construct an analytical benchmark, LongMINT (Long-Horizon Memory under INTerference), which features (1) long, highly interconnected contexts with frequently updated information, (2) diverse coverage across multiple memory domains (Wikipedia, code, multi-turn dialogue, and state tracking), enabling evaluation of domain generalization, and (3) diverse question types, including (i) single-target recall tasks that test retrieval under interference over long contexts and (ii) multi-target aggregation tasks that require counting, ordering, or reasoning across multiple relevant pieces of information. We evaluate over six representative systems, including vanilla long-context LLMs, retrieval-augmented generation methods, and memory-augmented agent frameworks. We observe consistently low performance (avg. 27.7% accuracy), especially on questions that require aggregated reasoning over multiple pieces of evidence. Fine-grained analysis shows that performance is primarily limited by retrieval and memory construction capabilities. Furthermore, current memory systems struggle to recall and reason over facts that are multiple steps back, and performance decreases when this lookback distance increases. These findings highlight the need for more robust memory management systems for dynamic, long-horizon environments across varying domains.
Authors:
Shaoxiang Qin, Xiongye Xiao, Yucheng Zhao, Fuyuan Lyu, Di Zhou, Jiachen Yao, Steve Liu, Animashree Anandkumar, Liangzhu L WangAbstract: In urban low-altitude flight, buildings reshape ambient wind into spatially varying 3D flow, making unmanned aerial vehicle (UAV) energy depend on local wind exposure as well as path length. However, building-resolved wind information is rarely available when a mission must be planned. Computational fluid dynamics (CFD) can produce high-fidelity urban flow fields, but each simulation is tied to a fixed inflow boundary condition and can take hours to days, which is incompatible with urban UAV missions that typically last minutes to tens of minutes. We present GeoWind2Plan, a geometry-to-wind-to-planning framework for mission-time 3D urban wind prediction and energy-efficient UAV planning. Given only a background wind vector, 3D building geometry, and a start-goal pair, GeoWind2Plan transforms the building geometry into a reference-wind frame, predicts mission-relevant 3D wind patches with a localized geometry-conditioned neural operator, stitches them into a queryable local wind field, and optimizes a feasible 3D path and speed profile using a physically grounded UAV energy model. Rather than pursuing CFD-perfect reconstruction, GeoWind2Plan targets decision-useful wind prediction: trajectories are planned with predicted wind and evaluated under high-fidelity CFD wind. Across held-out urban domains, wind speeds, and mission wind-angle regimes, GeoWind2Plan performs corridor-localized wind inference in about 3 seconds, compared with roughly 8 hours for CFD. Under CFD evaluation, trajectories planned with GeoWind2Plan reduce energy by 6.9%, 12.7%, and 4.5% in tailwind, headwind, and crosswind missions relative to wind-agnostic planning, recovering 87.9%, 85.7%, and 75.0% of CFD-reference savings. These results show that fast, corridor-localized 3D urban wind prediction can make wind-aware UAV energy planning practical at mission time.
Abstract: Streaming 3D reconstruction from long monocular video sequences requires maintaining a key-value(KV) cache that grows linearly with sequence length, creating a severe memory bottleneck. Existing approaches either truncate the cache to a fixed set of anchor frames, leading to reconstruction quality degradation, or rely on attention-score heuristics that are agnostic to 3D scene structure,failing to preserve geometrically valuable tokens. To address these problems, we present GHOST (Geometry-Hierarchical Online Streaming Token Eviction), a training-free KV cache management framework that exploits the model's own 3D geometry outputs to evict redundant tokens online. GHOST introduces three mutually reinforcing innovations: a hierarchical dual-level importance scoring scheme, a privilege mechanism that protects special tokens from eviction, and a cosine-similarity-guided layer-wise budget allocation. Experiments on various benchmarks show that GHOST preserves excellent reconstruction quality while cutting the KV cache by nearly half and delivering 1.75× faster inference compared to state-of-the-art methods. Our code and model will be available to the public.
Authors: Bingshuo Qian, Xiang Cheng
Abstract: Recent latent diffusion methods jointly model multiple representations of an image, denoising them with asynchronous noise schedules in which one representation temporally leads another. However, existing asynchronous schedules are often hand-designed, relying on grid search over simple function classes. We propose to learn the asynchronous schedule directly. We parameterize the schedule as a convex monotone bijection f : [0,1] \to [0,1], where convexity together with f(0)=0 and f(1)=1 guarantees the desired leading-representation property by construction while still admitting a wide family of schedule shapes. A short probe stage jointly optimizes the schedule and diffusion model under a mild speed regularizer; the resulting smooth schedule is averaged over the stable probe window, frozen, and reused for full training, introducing no additional overhead during the main training stage. On ImageNet 256×256, our learned schedule reaches unguided FID 2.37 at 1M iterations and FID 1.05 with AutoGuidance at 200 epochs, outperforming the best hand-tuned baseline trained for four times longer. Code is available at \urlhttps://anonymous.4open.science/r/LAS_public-01E4/.
Abstract: Streaming audio-visual target speaker extraction (AVTSE) is essential for latency-sensitive applications. However, current causal systems face two coupled gaps. First, standard benchmarks fix the number of speakers at two and ignore the near-/far-field interference, reverberation, and sudden noise of real scenes. Second, high-performance causal architectures are too heavy for edge deployment. Lightweight alternatives try to reclaim capacity by repeatedly invoking a small separator, which offsets the savings. To address the first gap, we release RealSSA, a realistic AVTSE benchmark consisting of RealSSA-Sim and RealSSA-Real. RealSSA-Sim provides controllable simulated scenes with 2-6 dynamic near-/far-field speakers across four scene types, while RealSSA-Real provides a held-out real-recorded evaluation set. To address the second gap, we propose Falcon, a causal AVTSE method. Its Separator performs multi-scale time-frequency separation in a single encoder--decoder pass. We train Falcon with an encoder-space speaker contrastive loss. This loss suppresses near-field leakage at zero inference cost and transfers across backbones. Falcon achieves state-of-the-art extraction quality on RealSSA-Sim, LRS2-2Mix, LRS3-2Mix, and VoxCeleb2-2Mix. Compared with the causal AV-TFGridNet baseline, it cuts parameters by 91.9%, computation by 20.2×, and GPU inference latency by 11.1×. Code and demo are available at https://submission-falcon.github.io/demo/.
Abstract: Reliable evaluation of large language models (LLMs) is essential for their development and deployment, yet is often costly, risky, and difficult to perform safely online. We study for LLMs, where limited human-labeled data from a behavior model are used to evaluate a newer target LLM. This setting is challenging because labels are scarce, behavior--target distribution shift is common, and response likelihoods are often unavailable for black-box LLMs. We propose the via optimal transport to align labeled behavior-policy samples with unlabeled target-policy samples. OTROPE combines corrected human-labeled residuals with proxy predictors, yielding a stable -style evaluation without behavior-policy modeling or density-ratio estimation. We theoretically characterize why baseline evaluators fail under LLM distribution shift, and establish consistency and convergence rates for OTROPE when either the reweighted behavior distribution or the proxy predictor converges. Experiments on synthetic and real LLM evaluation tasks show that OTROPE consistently outperforms baselines while enabling ensembles of weaker LLM evaluators to approach and sometimes surpass stronger evaluators.
Abstract: Recent Thinking with Video approaches use Video Generation Models (VGMs) for visual reasoning by producing temporally coherent Chain-of-Frames as reasoning artifacts. Even strong VGMs, however, exhibit two recurring failure modes on goal-directed tasks: long-horizon drift on multi-step tasks and mid-clip simulation errors that compound. Both stem from the absence of explicit reasoning built upon the VGM's short-horizon visual prior, a role naturally filled by Vision-Language Models (VLMs), but where to place the VLM is non-trivial: upfront plans commit before any frame is generated and post-hoc critiques over whole videos intervene too late. We propose VLM-VGM Collaborative Video Reasoning (CollabVR), a closed-loop framework that couples the VLM with the VGM at step-level granularity: the VLM plans the immediate next action, inspects the clip the VGM generates, and routes test-time compute across qualitatively distinct recovery strategies (re-generation, action splitting) matched to the diagnosed failure. On Gen-ViRe and VBVR-Bench, CollabVR improves both open-source and closed-source VGMs over single-inference, Pass@k, and prior test-time scaling baselines at matched compute, with the largest gains on the hardest tasks. It also yields further improvements on top of a reasoning-fine-tuned VGM, indicating that step-level VLM supervision is orthogonal to and stackable with reasoning-oriented fine-tuning. We provide video samples and additional qualitative results at our project page: https://collab-vr.github.io.
Abstract: Extracting signals through alpha factor mining is a fundamental challenge in quantitative finance. Existing automated methods primarily follow two paradigms: Decoupled Factor Generation, which treats factor discovery as isolated events, and Iterative Factor Evolution, which focuses on local parent-child refinements. However, both paradigms lack a global structural view, often treating factor pools as unstructured collections or isolated single-lineage paths, which leads to redundant search and inefficient allocation of optimization effort. To address these limitations, we introduce AlphaPROBE, ( volution), a framework that reframes alpha mining as the strategic navigation of a Directed Acyclic Graph (DAG). By modeling factors as nodes and evolutionary links as edges, AlphaPROBE treats the factor pool as a dynamic, interconnected ecosystem. The framework consists of two core components: a Bayesian Factor Retriever that identifies high-potential seeds by balancing exploitation and exploration through a posterior probability model, and a DAG-aware Factor Generator that leverages the full ancestral trace of factors to produce context-aware, non-redundant optimizations. Extensive experiments on major Chinese and US stock market datasets against 8 competitive baselines demonstrate that AlphaPROBE significantly gains enhanced performance in predictive accuracy, return stability and training efficiency. Our results confirm that leveraging global evolutionary topology is essential for efficient and robust automated alpha discovery. We have open-sourced our implementation at https://anonymous.4open.science/r/AlphaPROBE-A103.
Abstract: Video diffusion alignment has been heavily reliant on scalar rewards. These rewards are typically derived from learned reward models in human preference datasets, requiring additional training and extensive collection. Moreover, scalar rewards provide coarse, global supervision, offering limited prompt-generation mismatch credit assignment and making models prone to reward exploitation and unstable optimization. We propose Diffusion-DRF, a free, rich, and differentiable reward framework for video diffusion fine-tuning. Diffusion-DRF employs a frozen, off-the-shelf Vision-Language Model (VLM) as the critic, eliminating the need for reward model training. Instead of relying on a single scalar reward, it decomposes each user prompt into multi-dimensional questions with freeform dense VQA explanation queries, yielding information-rich feedback. By direct differentiable optimization over this rich feedback, Diffusion-DRF achieves stable reward-based tuning without preference datasets collection. Diffusion-DRF achieves significant gains both quantitatively and qualitatively, outperforming state-of-the-art Flow-GRPO by 4.74 points in overall performance on unseen VBench-2.0.
Abstract: Text-attributed graphs (TAGs) underlie real-world applications such as citation networks, social media, and e-commerce. Few-shot graph learning on TAGs is hard: with only a handful of labels per class and the rest of the graph unannotated, neither GNNs nor LLMs can learn well on their own. GNNs read topology and fail on cold nodes; LLMs read text and fail on text-ambiguous nodes. Existing LLM-GNN methods all follow the same recipe: designate one model as the golden teacher and use its outputs (e.g., features or pseudo-labels) to supervise the other. We argue this golden-teacher assumption breaks under sparse supervision: neither model is golden, and treating either as such transfers its blind spots into the student. We therefore ask: can we avoid designating either model as the golden teacher, and still perform effective graph learning? We answer with LLM-GNN Co-Teaching, a bidirectional co-teaching framework in which neither model is fixed as teacher. The GNN and LLM exchange their most confident pseudo-labels under an architecture-specific small-loss criterion, and both update every round. Supervision is then mined from the trajectory: whenever a node moves from cross-model contradiction at round t to cross-model agreement at round t + 1, the LLM's two answers on the same input form a preference pair (old contradicting self ≺ new peer-endorsed self) for DPO training. We call this Round-based Pseudo-Label Preference Optimization (RPL-PO). On six benchmarks, LLM-GNN Co-Teaching consistently outperforms GNN-as-Judge and all prior methods, with absolute 3-shot gains of 7.86% on Cora and 7.73% on ogbn-arxiv; improvements carry over to 5-shot and to zero-shot cross-dataset transfer. Error-structure analysis further shows that abandoning the golden-teacher assumption substantially improves the LLM's graph learning capability on challenging samples.
Abstract: Standard Retrieval Augmented Generation (RAG) is poorly matched to agent memory. Unlike large heterogeneous corpora, agent memory forms a bounded and coherent interaction stream in which many spans are highly correlated or near duplicates. As a result, flat top-k similarity retrieval often returns redundant context, while summary-centric hierarchies can blur the subtle details that distinguish one candidate from another. We argue that agent memory should follow the principle of decoupling before aggregation: the system should first isolate reusable facts, updates, and distinguishing details from similar histories, and only then organise them for efficient retrieval. Based on this principle, we propose xMemory, which constructs a revisable hierarchical memory structure from original messages to segments, memory components, and groups. xMemory segments interaction history into local events, decouples each segment into memory components, aggregates related components into high-level groups using a sparsity--semantic faithfulness objective, and maintains this structure incrementally as memory evolves. At inference time, xMemory retrieves top-down, first selecting a compact backbone of complementary groups and components, and then expanding to segments and raw messages only when additional evidence reduces the reader's uncertainty. Experiments on LoCoMo and PerLTQA across diverse open source and closed source LLMs show consistent gains in answer quality and inference token efficiency, supported by analyses of redundancy, evidence density, and coverage.
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a key method for improving Large Language Models' reasoning capabilities, yet recent evidence suggests it may paradoxically shrink the reasoning boundary rather than expand it. This paper explains this shrinkage issue of RLVR by analyzing its learning dynamics and reveals two critical phenomena that account for this failure. First, we expose negative interference in RLVR, where learning to solve certain training problems actively reduces the likelihood of correct solutions for others, leading to the decline of Pass@k performance, or the probability of generating a correct solution within k attempts. Second, we uncover the winner-take-all phenomenon: RLVR disproportionately reinforces problems with high likelihood, correct solutions under the base model, while suppressing other initially low-likelihood ones. Through extensive theoretical and empirical analysis on multiple mathematical reasoning benchmarks, we show that this effect arises from the inherent on-policy sampling in standard RL objectives, causing the model to converge toward narrow solution strategies. These insights motivate us to design a simple yet effective data curation algorithm that focuses RLVR learning on low-likelihood problems (the non-winners) and achieves notable improvement in Pass@k performance, illustrating the potential of a broader class of data-curation mitigation strategies.
Abstract: Diffusion models dominate visual generation, but iterative denoising remains computationally demanding, especially under Classifier-Free Guidance (CFG), which is crucial for high-fidelity generation yet nearly doubles per-step computation. To reduce this burden, recent training-free caching methods reuse intermediate predictions during sampling, yet their single-prediction criteria are misaligned with CFG sampling, where the denoising update is governed by the guided prediction rather than either branch alone. Consequently, such reuse may misestimate the error that actually perturbs sampling, leading to a . Moreover, due to error propagation across timesteps, a local guided error may not faithfully reflect the final deviation, inducing a introduces CFG-aware Guided-Risk Composition to align caching decisions directly with the guided prediction, and Propagation-Aware Rescaling to calibrate local guided risks against the final generative deviation. Extensive experiments on FLUX.1-dev, Wan2.1-T2V-1.3B, and CogVideoX-2B show that delivers superior efficiency-fidelity trade-offs over existing training-free caching methods. Moreover, is compatible with diverse caching strategies and can further enhance their fidelity, e.g., improving PSNR over MagCache from 21.46 to 24.02 under CFG. The code is available at https://anonymous.4open.science/r/RA-CFGCache/.
Abstract: Self-distillation enables language models to learn in an on-policy fashion from their own trajectories by using the same model as both the teacher and student, with the teacher being conditioned on privileged information that is unavailable to the student. This information can take multiple forms or views, including solutions, demonstrations, or feedback; for example, for a math problem, one view might be a ground truth reasoning trace. Supervising the student via the privileged teacher enables fine-grained token-level feedback and removes misalignment caused by distillation from external policies with different semantic biases. However, self-distillation also creates a fundamental asymmetry: the teacher may encode view-specific information tied to the privileged information provided during training, which the student cannot access at inference time. Moreover, the best form of privileged information is often task-dependent, making it difficult to choose a single teacher view. In this work, we address both these challenges jointly by introducing AVSD (Adaptive-View Self-Distillation), a novel method of self-distillation with multiple privileged-information views, which reconstructs token-level supervision by separating stable cross-view consensus from view-specific residual support. The consensus signal provides a robust update direction, while the residual signal is selectively used to adjust the update magnitude. Together, these components form a gated aggregation mechanism that balances a conservative target (geometric mean) with a more permissive target (arithmetic mean) for aggregating teacher signals at the token level. Experiments on math competition benchmarks (AIME24, AIME25, and HMMT25) show that AVSD consistently outperforms single-view self-distillation baselines and GRPO, achieving 3.1% Avg@8 gain over the strongest baseline on Qwen3-8B on average. Moreover, on code-generation benchmarks (Codeforces, LiveCodeBench) using Qwen3-8B, AVSD outperforms single-view self-distillation baselines by 2.9% on average. More broadly, our analysis shows that AVSD provides better token-level learning signals than any single-view method by reconstructing supervision from multiple privileged teachers.
Abstract: Diffusion Transformers are fundamental for video and image generation, but their efficiency is bottlenecked by the quadratic complexity of attention. While block sparse attention accelerates computation by attending only critical key-value blocks, it suffers from degradation at high sparsity by discarding context. In this work, we discover that attention scores of non-critical blocks exhibit distributional stability, allowing them to be approximated accurately and efficiently rather than discarded, which is essentially important for sparse attention design. Motivated by this key insight, we propose PISA, a training-free Piecewise Sparse Attention that covers the full attention span while substantially reducing computational cost. Unlike the conventional keep-or-drop paradigm that directly drop the non-critical block information, PISA introduces a novel exact-or-approximate strategy: it maintains exact computation for critical blocks while efficiently approximating the remainder through block-wise Taylor expansion. This design allows PISA to serve as a faithful proxy to full attention, effectively bridging the gap between speed and quality. Experimental results demonstrate that PISA achieves 1.91times and 2.57 times speedups on Wan2.1 and Hunyuan-Video, respectively, while consistently maintaining the highest quality among sparse attention methods. Notably, even for image generation on FLUX, PISA achieves a 1.2 to 1.6 times acceleration at different resolutions without compromising visual quality.
Authors:
Changjian Jiang, Kerui Ren, Xudong Li, Kaiwen Song, Guanghao Li, Linning Xu, Tao Lu, Junting Dong, Yu Zhang, Yu Feng, Bo Dai, Mulin YuAbstract: Streaming reconstruction from monocular image sequences remains challenging, as existing methods typically favor either high-quality rendering or accurate geometry, but rarely both. We present PLANING, an efficient on-the-fly reconstruction framework built on a hybrid representation that loosely couples explicit geometric primitives with neural Gaussians, enabling geometry and appearance to be modeled in a decoupled manner. This decoupling supports an online initialization and optimization strategy that separates geometry and appearance updates, yielding stable streaming reconstruction with substantially reduced structural redundancy. Despite operating online, PLANING improves dense mesh Chamfer-L2 by 18.52% over the offline baseline PGSR, surpasses the streaming baseline ARTDECO by 1.31 dB in PSNR, and reconstructs ScanNetV2 scenes in under 100 seconds, over 5x faster than offline 2D Gaussian Splatting while matching per-scene optimization quality. Beyond reconstruction quality, the structural clarity and computational efficiency of PLANING make it well suited for a broad range of downstream applications, such as enabling large-scale scene modeling and simulation-ready environments for embodied AI.
Abstract: Recent advances in flow-based offline reinforcement learning (RL) have achieved strong performance by parameterizing policies via flow matching. However, they still face critical trade-offs among expressiveness, optimality, and efficiency. In particular, existing flow policies interpret the L_2 regularization as an upper bound of the 2-Wasserstein distance (W_2), which can be problematic in offline settings. This issue stems from a fundamental geometric mismatch: the behavioral policy manifold is inherently anisotropic, whereas the L_2 (or upper bound of W_2) regularization is isotropic and density-insensitive, leading to systematically misaligned optimization directions. To address this, we revisit offline RL from a geometric perspective and show that policy refinement can be formulated as a local transport map—an initial flow policy augmented by a residual displacement. By analyzing the induced density transformation, we derive a local quadratic approximation of the KL-constrained objective governed by the Fisher information matrix, enabling a tractable anisotropic optimization formulation. By leveraging the score function embedded in the flow velocity, we obtain a corresponding quadratic constraint for efficient optimization. Our results reveal that the optimality gap in prior methods arises from their isotropic approximation. In contrast, our framework achieves a controllable approximation error within a provable neighborhood of the optimal solution. Extensive experiments demonstrate state-of-the-art performance across diverse offline RL benchmarks.
Authors: Kai Yao, Marc Juarez
Abstract: Modern AI image generators are increasingly deployed as opaque APIs, where customers can query the deployed service, but cannot inspect model weights or architecture. This creates a practical challenge: a provider may pass governance certification with one generator and later silently switch to a cheaper and lower-quality one for deployment, compromising public trust or even safety in high-stakes domains. We study integrity auditing at deployment time and propose FARE (Forensic Acceptance Region Estimation). A certified generator is enrolled by training FARE on images sampled from that generator. After deployment, FARE can determine whether a generated image is consistent with the enrolled generator-using only that image. FARE's features are based on image generator-specific artifacts that have been proposed for forensic applications. FARE amplifies these features during training by finding hard samples that tighten the acceptance region and increase sensitivity to subtle changes in the certified generator. Across generator swaps, including substitutions with similar model versions and model variants, FARE is effective at detecting swaps, consistently outperforming existing baselines at strict operating points, and remains effective under adaptive attacks that actively manipulate images to evade swap detection.
Abstract: Existing inverse physics methods recover physical parameters from multi-view videos, where geometric constraints across views resolve scale and 3D structure. In monocular settings, however, such constraints are absent, leading to severe scale ambiguity, inaccurate geometry, and weak coupling between appearance optimization and physical simulation. We propose MonoPhysics, a framework for monocular inverse physics estimation of deformable objects using differentiable MPM simulation and 3D Gaussian Splatting, which jointly optimizes geometry, appearance, and physical parameters from a single camera view. We address these challenges through three visual-physical bridges: global scale alignment, physics-aware geometry refinement, and a differentiable position map, which together enable accurate optimization from image-space losses alone. We evaluate on Vid2Sim and our new dataset of elastic and plastic objects, showing that MonoPhysics outperforms existing baselines in monocular settings and achieves performance comparable to multi-view baselines using only a single camera.
Authors:
Minseo Lee, Byeonghyeon Lee, Lucas Y Lee, Eunsoo Lee, Jun Y Jeong, Sangmin Kim, Seunghyeon Song, Joo Chan Lee, Jong Hwan Ko, Jaesik Park, Eunbyung ParkAbstract: 4D Gaussian Splatting has emerged as a new paradigm for dynamic scene representation, enabling real-time rendering of scenes with complex motions. However, it faces a major challenge of storage overhead, as millions of Gaussians are required for high-fidelity reconstruction. Compressing explicit 4D Gaussians is challenging because they exhibit heterogeneous redundancy across static and dynamic regions, and their temporal attributes make direct quantization unstable. In this work, we present OMG4 (Optimized Minimal 4D Gaussian Splatting), a framework that constructs a compact set of salient Gaussians capable of faithfully representing 4D Gaussian models. Our method progressively reduces Gaussians in three stages: (1) Gaussian Sampling to identify primitives critical to reconstruction fidelity, (2) Gaussian Pruning to remove redundancies, and (3) Gaussian Merging to fuse primitives with similar characteristics. In addition, we integrate implicit appearance compression and extend Sub-Vector Quantization (SVQ) to 4D representations with a staged quantization scheme, further reducing storage while preserving quality. Extensive experiments on standard benchmark datasets demonstrate that OMG4 achieves a favorable rate-distortion trade-off over recent state-of-the-art methods, reducing model sizes by over 60% while maintaining reconstruction quality. These results demonstrate the effectiveness of OMG4 as a practical framework for compact and high-fidelity 4D scene representation.
Abstract: We study adaptive learning rate scheduling for norm-constrained optimizers (e.g., Muon and Lion). We introduce a generalized smoothness assumption under which local curvature decreases with the suboptimality gap and empirically verify that this behavior holds along optimization trajectories. Under this assumption, we establish convergence guarantees under an appropriate choice of learning rate, for which warm-up followed by decay arises naturally from the proof rather than being imposed heuristically. Building on this theory, we develop a practical learning rate scheduler that relies only on standard hyperparameters and adapts the warm-up duration automatically at the beginning of training. We evaluate this method on large language model pretraining with LLaMA architectures and show that our adaptive warm-up selection consistently outperforms or at least matches the best manually tuned warm-up schedules across all considered setups, without additional hyperparameter search. Our source code is available at https://anonymous.4open.science/r/warmup.
Abstract: Test-time scaling has emerged as a critical avenue for enhancing the reasoning capabilities of Large Language Models (LLMs). Though the straight-forward ``best-of-N'' (BoN) strategy has already demonstrated significant improvements in performance, it lacks principled guidance on the choice of N, budget allocation, and multi-stage decision-making, thereby leaving substantial room for optimization. While many works have explored such optimization, rigorous theoretical guarantees remain limited. In this work, we propose new methodologies to predict and improve scaling properties via tail-guided search. By estimating the tail distribution of rewards, our method predicts the scaling law of LLMs without the need for exhaustive evaluations. Leveraging this prediction tool, we introduce Scaling-Law Guided (SLG) Search, a new test-time algorithm that dynamically allocates compute to identify and exploit intermediate states with the highest predicted potential. We theoretically show that SLG approaches a full-information oracle at a reduced scale and, under additional regularity conditions, achieves vanishing matched-budget regret; in a Gaussian two-stage model, SLG can also match BoN with a polynomially larger sampling budget. Empirically, we validate our framework across different LLMs and reward models, confirming that tail-guided allocation consistently achieves higher reward yields than Best-of-N under identical compute budgets.
Abstract: In the realm of multi-objective alignment for large language models, balancing disparate human preferences often manifests as a zero-sum conflict. Specifically, the intrinsic tension between competing goals dictates that aggressively optimizing for one metric (e.g., helpfulness) frequently incurs a substantial penalty on another (e.g., harmlessness). While prior work mainly focuses on data selection, parameter merging, or algorithmic balancing during training, these approaches merely force compromises between divergent preferences along a fixed Pareto frontier, failing to fundamentally resolve the inherent trade-off. In this work, we approach this problem from a novel perspective of multi-dimensional rewards. By scaling up the model's rollouts and analyzing the outputs across different reward dimensions, we arrive at a critical conclusion: the conflict among multiple objectives stems from the fact that the prompt itself inherently restricts the achievable multi-dimensional rewards. Based on this core observation, we propose MORA: Multi-Objective Reward Assimilation (MORA). Specifically, MORA isolates single-reward prompts through pre-sampling and expands their reward diversity by rewriting the original questions to incorporate multi-dimensional intents. Extensive experiments demonstrate that: (I) in sequential alignment, MORA achieves single-preference improvements ranging from 5% to 12.4%, with exceptional gains in harmlessness, after multiple-preference alignment across helpful, harmless, and truthful dimensions. (II) In simultaneous alignment, MORA achieves an average overall reward improvement of 4.6%. Our codes are available at
Abstract: Existing 3D scene-grounded Large Language Models (3D-LLMs) focus on answering questions grounded in simplified single-room 3D scenes, lacking the ability to reason over real-world household environments containing multiple interconnected rooms and diverse object categories. We introduce CAIRN, a topology-aware 3D-LLM for multi-room 3D scene understanding. CAIRN aligns transformer attention with scene hierarchy, giving the model explicit awareness of object-level relations and room-level connectivity. It enriches object tokens with room-local relational context via a graph neural network, introduces learned room tokens for room-level abstraction, and applies a hierarchical attention mask with geometric bias to route information according to scene topology. CAIRN is developed on CAIRN-MR, a benchmark we introduce on HM3D for multi-room 3D scene understanding, covering grounding, captioning, and four question-answering tasks that progressively evaluate from intra-room perception to cross-room reasoning. Experiments show that CAIRN outperforms prior 3D-LLMs by a large margin across all CAIRN-MR tasks while remaining competitive on five single-room benchmarks.
Abstract: Evaluating LLM-generated interactive software requires execution in addition to static analysis. The key difficulty is that correctness is a graph-level reachability property over latent UI state-transition graphs, whereas a GUI evaluator observes only a single execution trajectory. A failed rollout therefore rules out only one realized path, leaving failure attribution ambiguous between evaluator-side execution error and genuine software defect. We present DIAGEVAL, a trajectory-conditioned diagnostic evaluation protocol for post-failure GUI-agent evaluation of interactive software. Rather than blindly retrying from scratch,DIAGEVAL reuses the failed trajectory to choose targeted diagnostic probes and aggregates their outcomes into an internal attribution signal. The latent-graph view motivates the diagnostic problem; DIAGEVAL does not reconstruct the graph or estimate calibrated posterior probabilities. We evaluate DIAGEVALon WebDevJudge-Unit and RealDevBench across multiple GUI-agent evaluators and LLM backbones. On false-negative cases, DIAGEVAL recovers 45.6--62.1% of failures that were initially attributed to software defects, outperforming retry-based baselines with 34.4--160.6% relative gains. On the full evaluation sets, this recovery improves accuracy from 69.9% to 78.3% on WebDevJudge-Unit and from 65.0% to 81.6% on RealDevBench. These results suggest that reliable GUI-agent evaluation requires not only stronger execution, but also active failure diagnosis to disambiguate evaluator-side errors from genuine software defects.
Abstract: LLM-based multi-agent systems are increasingly deployed on long-horizon tasks, but a single decisive error is often accepted by downstream agents and cascades into trajectory-level failure. Existing work frames this as \emphpost-hoc failure attribution, diagnosing the responsible agent and step after the trajectory has ended. However, this paradigm forfeits any opportunity to intervene while trajectory is still unfolding. In this work, we introduce AgentForesight, a framework that reframes this problem as online auditing: at each step of an unfolding trajectory, an auditor observes only the current prefix and must either continue the run or alarm at the earliest decisive error, without access to future steps. To this end, we curate AFTraj-2K, a corpus of agentic trajectories across Coding, Math, and Agentic domains, in which safe trajectories are retained under a strict curation pipeline and unsafe trajectories are annotated at the step of their decisive error via consensus among multiple LLM judges. Built on that, we develop AgentForesight-7B, a compact online auditor trained with a coarse-to-fine reinforcement learning recipe that first equips it with a risk-anticipation prior at the failure boundary on adjacent safe/unsafe prefix pairs, then sharpens this prior into precise step-level localization under a three-axis reward jointly targeting the what, where, and who of an audit verdict. Across AFTraj-2K and an external Who\&When benchmark, AgentForesight-7B outperforms leading proprietary models, including GPT-4.1 and DeepSeek-V4-Pro, achieving up to +19.9% performance gain and 3× lower step localization error, opening the loop from post-hoc failure detection to enabling deployment-time intervention.
Authors: Thatchanon Anancharoenkij, Donlapark Ponnoprat
Abstract: A complete understanding of heterogeneous treatment effects requires fully capturing the conditional distribution of the counterfactual outcomes. To this end, we propose the (CCME), a framework that embeds conditional distributions of counterfactual outcomes into a reproducing kernel Hilbert space (RKHS). Under this framework, we develop a two-stage meta-estimator for CCME that accommodates any RKHS-valued regression in each stage. Based on this meta-estimator, we develop three practical CCME estimators: (1) estimator that performs RKHS-valued regression, with the coefficients parameterized by a neural network. We provide finite-sample convergence rates for all estimators, establishing that they possess the double robustness property. Our experiments demonstrate that our estimators accurately recover distributional features including multimodal structure of conditional counterfactual distributions.
Abstract: Language models based on discrete diffusion have attracted widespread interest for their potential to provide faster generation than autoregressive models. Despite their promise, these models typically produce samples whose quality sharply degrades in the few-step regime, preventing a dramatic speedup in practice. Here, we show that language models based on continuous flows over one-hot token embeddings can outperform discrete diffusion in both quality and speed. Importantly, our continuous formulation defines a unique flow map that can be learned directly for efficient few-step inference, a structure we show is unavailable to discrete methods. In this setting, we show that both the flow and its associated flow map can be learned with simple cross-entropy objectives that respect the simplex geometry of the data, and we identify three distinct choices for flow map distillation whose performance we compare in practice. Using these insights, we build a flow language model (FLM), a continuous flow that matches state-of-the-art discrete diffusion baselines on the One Billion Words (LM1B) and OpenWebText (OWT) datasets. We then distill FLM into a flow map language model (FMLM), whose one-step generation exceeds the 8-step quality of recent few-step discrete diffusion language models. Our work challenges the widely-held hypothesis that discrete noising processes are necessary for generative modeling over discrete modalities and paves the way toward accelerated language modeling at scale.
Abstract: Predicting tandem mass spectra (MS/MS) from molecular structures represents a central task in analytical chemistry with direct relevance to clinical metabolomics, systems biology, and adjacent disciplines. In this work, we revisit the problem through the lens of object detection on molecular graphs. Molecular fragmentation, a central step in MS/MS prediction, can be approximated as detecting a set of subgraphs (i.e., fragments) and their associated spectral contributions. Existing fragment-based models follow a two-stage paradigm—first generating candidate fragments and then scoring them—analogous to two-stage R-CNNs in computer vision. Towards higher accuracy and faster inference, we introduce GLACIER, a single-stage transformer-based fragment detection neural network for molecular graphs. This unified formulation eliminates the need for candidate enumeration, enabling scalable and globally consistent modeling of molecular fragmentation. GLACIER is faster and more accurate than existing state-of-the-art methods by a significant margin, achieving 62.4% and 62.1% Top-1 retrieval accuracy with and without contrastive finetuning on the MassSpecGym dataset (from 55.3%), and 55.2% and 38.2%, respectively, on the NIST'20 dataset (from 33.5%). Furthermore, GLACIER provides nearly 3-fold inference speedup over existing two-stage models.
Abstract: Predicting transcriptional responses to genetic perturbations is fundamental to functional genomics and therapeutic discovery. Recent deep learning models have shown promise in single-cell perturbation response prediction, but they typically generate each response in isolation, without explicitly leveraging experimentally characterised responses to related perturbations. We introduce PT-RAG (Perturbation-aware Two-stage Retrieval-Augmented Generation), a plug-in retrieval-and-conditioning module for generative cellular perturbation response. PT-RAG augments an existing perturbation-response backbone with learned access to related perturbation contexts. The key challenge is that relevance is not fixed in this setting: functionally related genes may elicit different effects across cell types. PT-RAG addresses this with a two-stage retrieval mechanism: GenePT-based semantic retrieval first identifies K candidate perturbations, after which a differentiable Gumbel-Softmax selector adaptively selects retrieved contexts conditioned on the control cell state, the query perturbation, and each candidate perturbation. Through cross-cell-type and cross-perturbation generalization tasks, PT-RAG consistently improves distributional similarity and often overall predictive quality; for example, on scGPT cross-cell-type results, Wasserstein distance drops by 10.7%. The code to reproduce our experiments is available at https://anonymous.4open.science/r/PT-RAG_NIPS-13D8/.
Authors:
Mohammad Asadi, Soheil Hor, Bardiya Akhbari, Jack W O'Sullivan, Tahoura Nedaee, Layne Price, Raviteja Anantha Ramesh, Euan Ashley, Ehsan AdeliAbstract: Retrieval-augmented forecasting promises to adapt frozen Time Series Foundation Models (TSFMs) to new domains without fine-tuning, but recent methods typically rely on learned fusion modules, i.e., trained adapters that merge retrieved examples into the backbone's forecast, based on the assumption that frozen backbones cannot dynamically incorporate retrieved context on their own. We show this assumption is unnecessary. We introduce Align-RAG, a training-free method that applies a closed-form per-pair amplitude rescaling and integer-lag phase shift to retrieved past--future windows before they enter a frozen backbone's context. With no learned parameters, Align-RAG outperforms the state-of-the-art trained retrieval adapter on a frozen Chronos-Bolt on all seven datasets of the standard benchmark (avg -3.75% MSE), showing that the gains previously attributed to learned fusion are recoverable without any training. Align-RAG further improves zero-shot MSE on four additional frozen TSFMs with various architectures by 2.5% to 13.7% per backbone with no per-backbone tuning. To probe why alignment helps, we compare the frozen backbone's prediction shift under aligned demonstrations to the closed-form ridge prediction shift on the same pairs. We find that aligned demonstrations induce prediction shifts that track a closed-form ridge predictor on the same pairs, with a future-shuffle control ruling out a futures-averaging account. Together, these results indicate that frozen TSFMs already support dynamic in-context use of retrievals, and that closed-form alignment should be the default baseline for retrieval-augmented forecasting before any fusion module is trained. Code available at: \urlhttps://anonymous.4open.science/r/phase-rag-4D92
Authors:
Ye Zhang, Zijie Fang, Zhixiang LinAbstract: Single-cell multi-omics technologies provide a powerful basis for characterizing cell-state transitions and dynamic regulatory processes. However, existing methods for multi-omics dynamic modeling still face two major limitations. From the data perspective, many methods require high-quality paired multi-omics measurements; from the modeling perspective, dynamic inference often relies on predefined regulatory structures or kinetic assumptions. These limitations restrict their applicability to partially paired, unpaired, and broader cross-omics settings. To address these challenges, we propose ATLAS, a unified framework for single-cell multi-omics alignment and dynamic modeling that jointly learns cross-omics consistency and temporal dynamics from partially paired or unpaired data. ATLAS adaptively models temporal lag effects between omics layers to characterize asynchronous cross-omics regulation, and further introduces a reliability-guided temporal distillation strategy to improve model-based temporal ordering. We systematically evaluate ATLAS on five datasets across four tasks, showing strong overall performance in multi-omics alignment, cross-omics prediction, trajectory inference, and future-state prediction. Our code is available at https://github.com/anomity/ATLAS.
Authors:
Jicheng Yuan, Duc Manh Nguyen, Tung Kieu, Manfred Hauswirth, Danh Le-PhuocAbstract: Adverse weather conditions such as fog, rain, snow, and low illumination introduce structured visibility degradations that severely impair object detectors trained on clear-weather data. When source data is inaccessible due to privacy or transmission constraints, source-free domain adaptation (SFDA) becomes necessary; however, existing SFDA methods based on pairwise contrastive learning struggle to capture high-order semantics and often suffer from agreement collapse under severe corruption. We propose REUNION, a hypergraph-guided SFDA framework for object detection under adverse weather. REUNION constructs a contrastive hypergraph over target-domain object proposals, encoding high-order relations through intra-image context, weather-aware grouping, prototype-based semantic anchors, and uncertainty-aware connections between reliable and low-confidence instances. To effectively exploit these structured relations, we introduce a group-wise Hyper-InfoNCE objective that optimizes representations at the hyperedge level, enabling semantic information to propagate from confident proposals to corrupted or low-contrast instances. Experiments on diverse benchmarks demonstrate that REUNION effectively mitigates domain shifts and achieves state-of-the-art performance in SFDA settings. Code is available at this \hrefhttps://anonymous.4open.science/r/sfda-1E21/anonymous link.
Abstract: Predicting how a person's first-person view will evolve (what action will follow, what plan completes a task, whether an in-progress shot will score) is fundamentally under-specified: the same context admits many plausible futures, and a model trained to minimize prediction error is forced to hedge or average across them, getting it wrong either way. Two findings shape our approach. First, the future camera trajectory, the path the head carves through space, lets the model commit to one of those futures: it carries the operator's intent in a form fine enough to determine how an action will unfold, substantially outperforming language as a conditioning signal. Second, this same intent makes the trajectory itself partially predictable from the context at hand, enough that trajectory need not be observed at test time to recover most of the gain. We instantiate these findings as TrajPilot, a model that predicts candidate future trajectories from egocentric context and uses them to pilot action prediction in an action-aligned embedding space where language shapes the structure but is never used as a conditioning input. TrajPilot beats VLM and structured-planner baselines on procedural planning across Ego-Exo4D atomic, Ego-Exo4D Keystep, Ego4D GoalStep, and EgoPER, with the trajectory advantage widening with horizon (exactly where prior planners collapse) and holding under RGB-only camera-pose estimation. With the goal masked at inference, the same model performs goal-free anticipation, beating VLM baselines on Ego-Exo4D atomic and extending to EPIC-Kitchens-100 and basketball shot-outcome prediction.
Abstract: Derivative-informed training improves neural operators by directly supervising their input--output sensitivities, but existing methods rely on offline tangent solves and stored derivative labels. We propose sketched tangent consistency loss (sTCL), an on-the-fly drop-in derivative regularizer for neural-operator training that improves learned Jacobians without changing the architecture or generating offline sensitivity data. At each training step, sTCL samples a small number of perturbation directions in the input function space, computes the corresponding surrogate Jacobian--vector products using forward-mode automatic differentiation, and penalizes the residual of the forward sensitivity equation. This equation is obtained by differentiating the governing PDE with respect to the input perturbation direction, so derivative supervision is obtained directly from the physics rather than from stored tangent labels. We further show that sketching makes online derivative-informed training computationally feasible, while conditioning is the key to making it effective. Random directional sketching keeps the derivative penalty at the same order of cost as the standard data-fitting loss, making sTCL compatible with stochastic neural-network training as a lightweight loss term. However, raw forward-sensitivity residual penalties can fail for stiff, ill-conditioned, or indefinite tangent operators. To address this, we introduce lightweight operator-aware preconditioners selected by a simple tangent-operator decision rule. Across four PDE benchmarks---Helmholtz, nonlinear diffusion--reaction, Burgers, and two-dimensional Navier--Stokes---operator-conditioned sTCL achieves accuracy comparable to offline DIFNO while eliminating the offline derivative-data generation stage. These results show that on-the-fly derivative-informed training need not merely amortize offline tangent-solve cost into training; with appropriate sketching and conditioning, sTCL provides an attractive drop-in path to derivative-informed neural operators.
Abstract: While large-scale visual foundation models (VFMs) exhibit strong generalization across diverse visual domains, their potential for infrared small target (SIRST) detection remains largely unexplored. To fill this gap, we systematically introduce frozen VFM representations into the SIRST task and propose a Foundation-Driven Efficient Paradigm (FDEP), a general framework compatible with diverse VFMs and SIRST networks, which improves detection accuracy without additional VFM-related inference overhead. Specifically, a Semantic Alignment and Modulated Fusion (SAMF) module is designed to achieve dynamic alignment and deep fusion of the global semantic priors from VFMs with task-specific features. Meanwhile, to avoid the inference-time overhead introduced by VFMs, we propose a Collaborative Optimization-based Implicit Self-Distillation (CO-ISD) strategy, which enables implicit semantic transfer between the main and lightweight branches through parameter sharing and synchronized backpropagation. In addition, to unify the fragmented evaluation system, we construct a Holistic SIRST Evaluation (HSE) metric that performs multi-threshold integral evaluation at both pixel-level confidence and target-level robustness, providing a stable and comprehensive basis for fair model comparison. Extensive experiments demonstrate that the SIRST detection networks equipped with our FDEP framework achieve state-of-the-art (SOTA) performance on multiple public datasets. Our code will be open source.
Authors:
Wenxin Tang, Wenbin Li, Junliang Liu, Jingyu Xiao, Xi Xiao, Mingzhe liu, Jinlong Yang, Xuan Liu, Yuehe Ma, Wang Luo, Qing Li, Lei Wang, Peng XiangliAbstract: Software vulnerability detection plays a critical role in ensuring system security, where real-world auditing requires not only determining whether a function is vulnerable but also pinpointing the specific lines responsible. However, existing approaches either rely on a single information source---sequential, structural, or semantic---failing to jointly exploit the complementary strengths across modalities, or treat statement-level localization merely as a byproduct of function-level detection without explicit line-level supervision. To address these limitations, we propose DCVD (Dual-Channel Cross-Modal Vulnerability Detection), a unified framework that performs joint function-level detection and statement-level localization. DCVD extracts control-dependency and semantic features through two parallel branches and integrates them via contrastive alignment coupled with bidirectional cross-attention, effectively bridging the cross-modal representation gap. It further introduces explicit supervision signals at both the function and statement levels, enabling collaborative optimization across the two granularities. Extensive experiments on a large-scale real-world vulnerability benchmark demonstrate that DCVD consistently outperforms state-of-the-art methods on both function-level detection and statement-level localization.
Authors:
Hojeong Lee, Si H Lee, Sungwon Woo, Chengpo Yan, Suman Banerjee, Seyeon KimAbstract: We present NPUsper, a live transcription system that makes Whisper efficient on mobile NPUs by eliminating redundant computation. To avoid the heavy padding used by prior streaming systems, NPUsper detects hallucinated tokens online from temporal patterns in decoder cross-attention, allowing each inference round to process short audio inputs with minimal carryover. For efficient mobile-NPU execution, we propose controlled unrolling, which executes autoregressive decoding as K-step chunk graphs, removing unnecessary KV-cache computation and reducing graph-dispatch overhead. NPUsper achieves up to 4.84x lower per-word latency, up to 33.2x lower time-to-first-token (TTFT), and up to 11.36% lower average power consumption compared with baselines, while maintaining comparable transcription accuracy. The code is available at https://github.com/npusper/NPUsper.
Authors: JunJian Wang, Xin Zhou, Qiran Xu, Kun Zhan
Abstract: While reasoning has become a central capability of large language models (LLMs), the reasoning patterns required for different scenarios are often misaligned. Mathematical reasoning typically relies on intrinsic logic to solve closed-world problems in a single response, whereas agentic reasoning requires not only internal reasoning but also multi-turn interaction with external environments, interleaving thought and action. This misalignment prevents mathematical and agentic reasoning from effectively benefiting from each other, often yielding unstable reasoning behavior and only limited performance gains under multi-task learning. In this paper, we propose M2A, a novel paradigm that synergizes mathematical and agentic reasoning via model merging. To avoid overfitting to superficial reasoning patterns under joint training, M2A operates directly in parameter space: it identifies the feature subspace critical for agent behavior, and merges the mathematical reasoning task vector only along its null space, thereby injecting reasoning capability along directions that do not perturb agent behavior. Unlike SFT or RL, M2A requires no additional gradient-update and exposes the merging coefficient as a simple knob for controlling reasoning length. Experiments in a challenging real-world coding agent setting show that our method effectively extends agentic reasoning depth and delivers substantial performance improvements. Applied to a fine-tuned Qwen3-8B, M2A improves its SWE-Bench Verified resolved rate from 44.0% to 51.2% without retraining the model.
Authors: Bhavik Chandna, Kelsey Allen
Abstract: AI video generation is evolving rapidly. For video generators to be useful for applications ranging from robotics to film-making, they must consistently produce realistic videos. However, evaluating the realism of generated videos remains a largely manual process -- requiring human annotation or bespoke evaluation datasets which have restricted scope. Here we develop an automated evaluation framework for video realism which captures both semantics and coherent 3D structure and which does not require access to a reference video. Our method, 3DSPA, is a utoencoder utoencoder which is trained to reconstruct held out 3D point trajectories which have been augmented with DINO semantic features. By combining both semantic and geometric representations, 3DSPA enables robust assessments of realism, temporal consistency, and physical plausibility in generated videos. Experiments show that 3DSPA outperforms leading VLMs by more than 10% in identifying videos which violate physical laws, and more than doubles the correlations between human and auto-eval ratings in judging physical common sense and motion quality across multiple generative video datasets. Our results demonstrate that enriching trajectory-based representations with 3D semantics offers a strong foundation for benchmarking generative video models, and implicitly captures physical rule violations.
Abstract: Understanding social interactions requires reasoning over subtle non-verbal cues, yet current multimodal large language models (MLLMs) often fail to identify who interacts with whom in multi-person videos. We introduce GRASP, a large-scale social reasoning dataset that connects high-level social QA with fine-grained gaze and deictic gesture events. GRASP contains 290K question-answer pairs over 46K videos totaling 749 hours, organized by a 16-category taxonomy spanning gaze, gesture, and joint gaze-gesture reasoning, together with GRASP-Bench for evaluation. Unlike prior resources that focus on either isolated cues or high-level social QA, GRASP builds questions from identity-consistent gaze trajectories, deictic gestures, and their joint compositions into social events. Moreover, we propose Social Grounding Reward (SGR), a learning signal that uses these social events to encourage models to reason about the participants involved in each interaction. Experiments show that SGR improves performance on GRASP-Bench while maintaining zero-shot performance on related social video QA benchmarks.
Abstract: The expansion of retrieval-augmented generation (RAG) into multimodal domains has intensified the challenge for processing complex visual documents, such as financial reports. While page-level chunking and retrieval is a natural starting point, it creates a critical bottleneck: delivering entire pages to the generator introduces excessive extraneous context. This not only overloads the generator's attention mechanism but also dilutes the most salient evidence. Moreover, compressing these information-rich pages into a limited visual token budget further increases the risk of hallucinations. To address this, we introduce AgenticOCR, a dynamic parsing approach that extends optical character recognition (OCR) from a static, full-text process into a query-driven, on-demand extraction system. By autonomously analyzing document layout in a "thinking with images" manner, AgenticOCR identifies and selectively recognizes regions of interest. This approach performs on-demand decompression of visual tokens precisely where needed, effectively decoupling retrieval granularity from rigid page-level chunking. AgenticOCR has the potential to serve as the "third building block" of the visual document RAG stack, operating alongside and enhancing standard Embedding and Reranking modules. Experimental results demonstrate that AgenticOCR improves both the accuracy and efficiency of visual RAG systems with pixel-adaptive per-image token allocation (e.g., Qwen3.5 generator), achieving expert-level performance in long document understanding. Our repository is anonymously available at https://anonymous.4open.science/r/AgenticOCR.
Authors:
Sudong Wang, Weiquan Huang, Xiaomin Yu, Zuhao Yang, Hehai Lin, Keming Wu, Chaojun Xiao, CHEN CHEN, Wenxuan Wang, Beier Zhu, Yunjian Zhang, Chengwei QinAbstract: The standard post-training recipe for large multimodal models (LMMs) applies supervised fine-tuning (SFT) on curated demonstrations followed by reinforcement learning with verifiable rewards (RLVR). However, SFT introduces distributional drift that neither preserves the model's original capabilities nor faithfully matches the supervision distribution. This problem is further amplified in multimodal reasoning, where perception errors and reasoning failures follow distinct drift patterns that compound during subsequent RL. We introduce PRISM, a three-stage pipeline that mitigates this drift by inserting an explicit distribution-alignment stage between SFT and RLVR. Building on the principle of on-policy distillation (OPD), PRISM casts alignment as a black-box, response-level adversarial game between the policy and a Mixture-of-Experts (MoE) discriminator with dedicated perception and reasoning experts, providing disentangled corrective signals that steer the policy toward the supervision distribution without requiring access to teacher logits. While 1.26M public demonstrations suffice for broad SFT initialization, distribution alignment demands higher-fidelity supervision; we therefore curate 113K additional demonstrations from Gemini~3 Flash, featuring dense visual grounding and step-by-step reasoning on the hardest unsolved problems. Experiments on Qwen3-VL show that PRISM consistently improves downstream RLVR performance across multiple RL algorithms (GRPO, DAPO, GSPO) and diverse multimodal benchmarks, improving average accuracy by +4.4 and +6.0 points over the SFT\rightarrowRLVR baseline on 4B and 8B, respectively.
Authors:
Kun Wang, Zherui Li, Zhenhong Zhou, Jie Zhang, Yitong Zhang, Yan Mi, Kun Yang, Yiming Zhang, Zhongxiang Sun, Qiankun Li, Yang LiuAbstract: Omni-modal Large Language Models (OLLMs) greatly expand LLMs' multimodal capabilities but also introduce cross-modal safety risks. However, a systematic understanding of vulnerabilities in omni-modal interactions remains lacking. To bridge this gap, we establish a modality-semantics decoupling principle and construct the AdvBench-Omni dataset, which reveals a significant vulnerability in OLLMs. Mechanistic analysis uncovers a Mid-layer Dissolution phenomenon driven by refusal vector magnitude shrinkage, alongside the existence of a modal-invariant pure refusal direction. Inspired by these insights, we extract a golden refusal vector using Singular Value Decomposition and propose OmniSteer, which utilizes lightweight adapters to modulate intervention intensity adaptively. Extensive experiments show that our method not only increases the Refusal Success Rate against harmful inputs from 69.9% to 91.2%, but also effectively preserves the general capabilities across all modalities.
Abstract: Low-bit post-training quantization (PTQ) is a pivotal technique for deploying Vision-Language Models (VLMs) on resource-constrained devices. However, existing PTQ methods often degrade VLMs' accuracy due to the heterogeneous activation distributions of text and vision modalities during quantization. We find that this cross-modal heterogeneity is distributed unevenly across channels: a small subset of channels contains most modality-specific outliers, and these outliers typically reside in different channels for each modality. Motivated by this, we propose SplitQ, a channel-Splitting-driven post-training Quantization framework. At its core, SplitQ introduces a novel Modality-specific Outlier Channel Decoupling (MOCD) module that effectively isolates salient modality-specific outlier channels with minimal overhead. To further address the remaining cross-modal distribution discrepancies, we design an Adaptive Cross-Modal Calibration (ACC) module that employs dual lightweight learnable branches to dynamically mitigate modality-induced quantization errors.Extensive experiments on popular VLMs demonstrate that SplitQ significantly outperforms existing approaches across 6 popular multi-modal datasets under all evaluated quantization settings, including W4A8, W4A4, W3A3, and W3A2. Notably, SplitQ preserves 95.3% of FP16 performance under the challenging W3A3 setting (69.5 vs. 74.3), pushing the efficiency frontier for deploying advanced VLMs.
Abstract: pose a critical risk: code that executes and returns solver-feasible solutions may encode semantically incorrect formulations—a feasibility–correctness gap reaching 90 percentage points on compositional problems. We introduce decomposes code production into a four-stage reasoning chain (understand, formalize, synthesize, verify), preventing formulation errors at their source. detects errors that survive generation by testing whether the formulation responds correctly to solver-based parameter perturbation—an external semantic signal that bypasses LLM self-review and requires no ground truth. The two mechanisms are complementary by error structure: structured generation drives the largest gains on compositional problems ( accuracy on RetailOpt-190 with Claude Opus 4.6), while behavioral verification dominates on localized defects ( on MAMO-ComplexLP, its largest contribution across benchmarks). Combined with diagnostic execution recovery, ReLoop reaches on Claude Opus 4.6 and consistently improves accuracy on chat-tuned foundation models across three benchmarks; we further identify a known limitation of narrowly-tuned SFT models, whose learned output formats are brittle to chain-of-thought prompts—an interaction we document and analyze. We release , 190 compositional retail optimization scenarios targeting the multi-constraint interactions where LLMs most frequently fail. Code and benchmark will be released upon acceptance.
Authors:
Bach H Ngo, Ngo Tri, Hieu Le, Trung Nghia LeAbstract: Text-to-image diffusion models fail to generate correct object counts in dense scenes, where overlapping instances collapse into indistinguishable structures despite appearing visually plausible. We identify this as instance ownership collapse: tokens from overlapping objects interact freely through attention, while heavily occluded instances receive weak supervision due to their small visible areas. We address this through layout-aware attention biases that softly bias token interactions toward region-consistent grouping and suppress cross-instance leakage, paired with an amodal-balanced loss that amplifies gradients for occluded objects based on their occlusion level. To enable systematic evaluation, we introduce OverlapDepth-45K, a benchmark of densely overlapping scenes with amodal supervision. Our approach substantially improves count accuracy and prevents instance merging while preserving image quality.
Abstract: Discrete diffusion models are typically trained by predicting the clean data given a noisy observation, which is then used with a certain choice of reverse transition parameterization. Two choices coexist in the literature, which we refer to as marginalization and plug-in. Although they coincide in the case of masked diffusion models (MDMs), we revisit this question for uniform diffusion models (UDMs), where the distinction matters. We show that the plug-in parameterization is not optimized by the denoising posterior but by a \emphleave-one-out posterior that removes the influence of the local observation. We give a thorough characterization of the resulting training target, together with exact conversion formulas between the denoiser, the \emphleave-one-out posterior, and the score. As a by-product, the leave-one-out characterization yields a Gibbs-based predictor–corrector scheme at no additional training cost. We further propose an absorbing-state construction of uniform diffusion, whose denoiser decomposes into simpler absorbing denoising posteriors, and which recovers the full UDM joint law following a simple resampling step. Empirically, we show that these design choices interact strongly with optimization and generative quality; on language modeling, the \emphleave-one-out denoiser improves over other parameterizations.
Abstract: Optimizing high-dimensional black-box functions under black-box constraints is a pervasive task in a wide range of scientific and engineering problems. These problems are typically harder than unconstrained problems due to hard-to-find feasible regions. In this paper, we propose a new framework to solve high-dimensional, constrained black-box optimization problems via posterior inference in the latent space of generative models. Our method iterates through two stages. First, we train flow-based models to capture the data distribution and surrogate models that predict both function values and constraint violations. Second, we cast the candidate selection problem as a posterior inference problem to effectively search for promising candidates that have high objective values while not violating the constraints. Concretely, we utilize diffusion models to amortize the sampling from the posterior distribution in the latent space of flow-based models, which can bypass the issue of mode collapse. We empirically demonstrate that our method achieves superior performance across synthetic and real-world tasks. Our code is available
Abstract: Diffusion transformers (DiTs) have emerged as a dominant architecture for text-to-image generation, yet their performance drops when generating at resolutions beyond their training range. Existing training-free approaches mitigate this by modifying inference-time attention behavior, often through Rotary Position Embeddings (RoPE) extrapolation combined with attention scaling. However, these strategies apply a uniform and content-agnostic scaling across RoPE components with distinct frequency characteristics, inducing a trade-off between preserving global structure and recovering fine detail. We introduce SEGA, a training-free method that dynamically scales attention across RoPE components according to the latent’s spatial-frequency structure at each denoising step. This adaptive scaling improves both structural coherence and fine-detail fidelity. Experiments show that SEGA consistently improves high-resolution synthesis across multiple target resolutions, outperforming state-of-the-art training-free baselines.
Abstract: Computer Use Agents (CUAs) can act through both atomic GUI actions (e.g., click, type) and high-level tool calls (e.g., API-based file operations), but they are often confused by this hybrid action space: they do not know when to continue with GUI actions and when to switch to tools, and finally fail to select the optimal execution path. We view this orchestration problem as GUI-Tool path selection: deciding when the agent should continue with GUI actions and when it should switch to tool calls to form an effective execution trajectory. This difficulty stems from two issues. First, high-quality interleaved GUI-Tool trajectories are scarce, and collecting real tool trajectories is expensive and brittle. Second, existing supervision provides limited guidance for GUI-Tool path selection, as most methods focus on step-level action imitation or final task completion and offer little trajectory-level feedback on whether GUI-Tool switching leads to a more effective execution path. In this paper, we propose ToolCUA, an end-to-end agent designed to learn optimal GUI-Tool path selection through a staged training paradigm. We first introduce an Interleaved GUI-Tool Trajectory Scaling Pipeline that repurposes abundant static GUI trajectories and synthesizes a grounded library of tools, making it possible to scale diverse GUI-Tool trajectories without manual engineering or real tool-trajectory collection.Based on this data, we perform Tool-Bootstrapped GUI RFT, which combines warmup SFT with single-turn RL to improve decisions at critical GUI-Tool switching points. Finally, we further optimize ToolCUA with Online Agentic RL in a high-fidelity GUI-Tool environment, using a Tool-Efficient Path Reward that encourages both appropriate tool use and shorter execution paths. Experiments on OSWorld-MCP show that ToolCUA achieves 46.85% accuracy and outperforms the baseline by over 60% relatively, establishing a new state of the art among models of comparable scale. It also improves by 3.9% over GUI-only settings, demonstrating effective GUI-Tool orchestration. The results further suggest that training in a hybrid action space is a promising paradigm for real-world digital agents.
Abstract: We present UniALM, a unified audio foundation model built on FactorCodec, a purely discrete tokenizer that factorizes audio into two role-separated representations. Analysis tokens are optimized to retain language-aligned, text-expressible information that can be decoded by a text LLM head into grounded natural-language analyses, while reconstruction tokens preserve the information required for waveform synthesis and are used exclusively for audio decoding. This factorization yields understanding performance competitive with continuous Whisper features on audio understanding tasks, while reducing the perplexity of reconstruction-token modeling for generation. UniALM further adopts functional layer specialization, partitioning the backbone into audio-understanding, cross-modal, and audio-generation experts. We train the model with a four-stage recipe on 100B text tokens and 60B audio tokens, together with a composed audio sequence construction strategy for unified multi-task pre-training. At 3B parameters, UniALM is competitive with strong 7B unified baselines on in-domain tasks and exhibits non-trivial \emphcompositional generalization to unseen tasks under few-shot and zero-shot evaluation. Unlike prior unified systems that rely on hybrid continuous-discrete inputs or do not support generation beyond speech, UniALM provides a purely discrete interface for both understanding and generation across speech, sound, and music.
Authors:
Jiaqian Wang, Shengtao Zhang, Ruiwen Zhou, Junwei Liao, Yuchen Feng, Zhuo Li, Yujie Zheng, Weinan Zhang, Ying Wen, Zhiyu li, Feiyu Xiong, Yutao Qi, Bo Tang, Muning WenAbstract: The hallmark of human intelligence is the self-evolving ability to master new skills by learning from past experiences. However, current AI agents struggle to emulate this self-evolution: fine-tuning is computationally expensive and can introduce destabilizing weight updates, increasing the risk of catastrophic forgetting, while existing memory-based methods rely on passive semantic matching that often retrieves noise. To address these challenges, we propose MemRL, a non-parametric approach that evolves via reinforcement learning on episodic memory. By decoupling stable reasoning from plastic memory, MemRL employs a Two-Phase Retrieval mechanism to filter noise and identify high-utility strategies through environmental feedback. Extensive experiments on Humanity's Last Exam, BigCodeBench, ALFWorld, and Lifelong Agent Bench demonstrate that MemRL significantly outperforms state-of-the-art baselines, suggesting that MemRL offers a practical way to balance stability and plasticity during runtime improvement.
Abstract: Consistent video generation under editing operations requires persistence: when edits modify scene appearance or layout, subsequent generations should remain coherent across time and viewpoints. However, existing memory designs struggle to maintain long-term consistency after such modifications, as stored contexts may become outdated or invalid. To address this, we propose PermaVid, a novel framework built upon a multi-modal context memory that disentangles spatial context into semantic appearance and geometric structure, together with an edit-aware memory update and retrieval strategy that keeps memory evolution aligned with subsequent observations. Specifically, we develop two complementary memory banks: an RGB context memory that captures appearance-aware observations while implicitly encoding geometry, and a depth context memory that preserves geometry-only structure disentangled from semantics. Building on this design, we introduce a memory-guided video generation model that performs multi-modal feature fusion under reference conditions drawn from mixed-modality memory contexts. Experiments demonstrate that our method maintains strong long-term semantic and structural consistency after edits, significantly outperforming state-of-the-art methods.
Abstract: Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method itself: architecture, objective, data exposure, schedule, scale, and decoder capacity can all vary at once. This makes it hard to identify which choices yield motion-prioritized representations, whose gains concentrate on frame-to-frame change while retaining useful appearance. We address this with a matched 4 × 6 = 24 architecture--objective study at roughly 170M--190M encoder scale on ~1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, and propose TT-VidT. TT-VidT combines a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer, trained by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens. The sweep shows that TT3D with Diff Compression, not either component alone, enters the strongest motion-sensitive regime, and decoder ablations favor a compact video-pretrained decoder. In final comparison, TT-VidT leads Jester, Something--Something\,V2, ARID, and Diving48 fine-tuning simultaneously, improving over the strongest non-TT row by 54%--121%, while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA~2. HMDB51, IARD, and EPIC-Kitchens bound the claim.
Authors:
Zhexiao Xiong, Yizhi Song, Hao Kang, Qing Yan, Liming Jiang, Jenson Yang, ZHOUJIE FU, Stathi Fotiadis, Zichuan Liu, Bo Liu, Yiding Yang, Xin Lu, Nathan JacobsAbstract: Interactive world models aim to simulate environment dynamics under real-time user actions. However, their action vocabulary is largely confined to navigation: most actions correspond to motion (e.g., walk, turn, look around), while interaction with objects in the scene (e.g., pick up plates, open doors, or trigger physical responses) is either absent, restricted to game domains, or relegated to prompt-to-full-video scenarios. The resulting worlds are visually explorable but not truly actionable. In this work, we present ActWorld, an interactive world model that extends prior navigation-centric generators to support mid-rollout object interaction within a chunk-autoregressive framework. We argue that the navigation--interaction gap stems from two bottlenecks. First, a data bottleneck: the lack of human--object interaction data with accurate, dense labels. Second, a memory bottleneck: recency-biased history compression in existing world models discards the event-transition frames that causally determine subsequent object states, leading to an action-forgetting pathology. On the data side, we construct a 100K interaction video dataset, each annotated with per-chunk captions via chain-of-thought reasoning. On the model side, we introduce a hierarchical action-aware memory design that routes history compression by interaction importance, complemented by a persistent memory bank that maintains event-update and object-identity tokens across long rollouts. Experiments show that ActWorld supports both flexible navigation and rich object interaction within a single model, substantially improving interaction fidelity over navigation-only baselines without sacrificing viewpoint control.
Abstract: A core challenge in structural biophysics is generating biomolecular conformations that are both physically plausible and consistent with experimental measurements. While sequence-to-structure diffusion models provide powerful priors, posterior sampling methods steer generation by perturbing atomic coordinates with gradients from experimental likelihoods. However, when the target lies in a low-density region of the prior, these methods require aggressive upweighting of the likelihood that can destabilize sampling and be sensitive to hyperparameters. We propose EmbedOpt, an inference-time steering framework that introduces an orthogonal optimization axis: rather than performing posterior sampling under a fixed prior, EmbedOpt directly optimizes the prior by updating the model's conditional embedding. This embedding space encodes rich coevolutionary signals, so optimizing it shifts the structural prior to align with experimental constraints. Empirically, EmbedOpt matches coordinate-based posterior sampling baselines on sparse distance constraints and outperforms them on cryo-electron microscopy map fitting, including real, noisy experimental ones. Furthermore, EmbedOpt's smooth optimization behavior yields robustness to hyperparameters spanning two orders of magnitude and enables comparable performance with fewer diffusion steps.
Abstract: Guardrails are a critical safety layer for modern AI systems, but their operating regime is changing. As LLMs are deployed as customized assistants, safety policies are increasingly specified at inference time by users, organizations, or regulatory contexts. This makes safety enforcement fundamentally dynamic: the guardrail should adapt to changing safety policies without retraining. Yet this requirement creates a fundamental tension: faithfully judging complex policy contexts demands reasoning capability, while practical deployment requires low-latency responses. We introduce Latent Policy Guardrail (LPG), a guardrail framework that learns over dynamic policies. LPG compresses the internal deliberation needed for intent interpretation and policy grounding into continuous states supervised by decision-relevant semantics. At inference time, it generates only a compact verdict anchored to the violated policy clauses, preserving auditability while avoiding the latency of explicit reasoning. Across policy guardrail benchmarks, LPG-4B reaches 84.5% average safety accuracy and 77.9% F1 by compressing deliberation into just 10 latent tokens, outperforming the strongest dynamic baseline while running roughly 11 times faster than Qwen3-4B-Thinking under the single-sample evaluation setup.
Abstract: Large language model (LLM) routing aims to exploit the specialized strengths of different LLMs for diverse tasks. However, existing approaches typically focus only on selecting LLM architectures, treating candidate models as static black boxes while overlooking parameter adaptation, which can substantially affect task performance. In this paper, we introduce HAPS, a hierarchical LLM routing framework that jointly searches over model architectures and parameters. Specifically, a high-level router selects among candidate LLM architectures, while a low-level router generates input-conditioned LoRA parameters for the selected architecture. To couple these two decisions, we design a shared parameter generation mechanism that enables cross-level knowledge transfer between architecture routing and parameter adaptation. We further optimize the whole framework with a reward-augmented training objective. Extensive experiments show that HAPS consistently outperforms strong routing baselines. We further conduct broad ablations and analyses on parameter sharing, architectural choices, scalability, and efficiency, providing additional evidence for the effectiveness and practicality of the proposed framework. We have released our code at https://anonymous.4open.science/r/HAPS_private-68CB.
Abstract: Image geo-localization aims to determine where a photograph was taken, a task that often requires more than recognizing visible landmarks. Human experts typically solve it through an iterative workflow: they inspect informative regions, form location hypotheses, seek external evidence, and revise their judgments as new clues appear. Existing methods only partially capture this process: direct prediction methods bypass evidence acquisition altogether, while retrieval-augmented methods introduce external evidence but usually provide limited supervision on the intermediate decisions of where to search, how to query, and how to filter noisy results. We present REVERSE, a framework that reinforces the interplay between evidence search and verification to enable multi-turn agentic reasoning. REVERSE teaches three intermediate decisions: where to look, what to query, and what evidence to trust. To support this, we construct tool-grounded trajectories with annotated region selections, search observations, and geo-informative evidence labels, and introduce process rewards for visual grounding, query utility, and evidence discrimination. An offline search cache makes retrieval observations stable and reusable during reinforcement learning, enabling dense supervision over noisy search results. With a 4B model, REVERSE outperforms strong retrieval-augmented baselines and rivals substantially larger models on Im2GPS3k and YFCC4k.
Abstract: Large language models (LLMs) rely on web-scale corpora for pre-training. The noise inherent in these datasets tends to obscure meaningful patterns and ultimately degrade model performance. Data curation mitigates but cannot eliminate such noise, so pre-training corpora remain noisy in practice. We therefore study whether a lightweight pre-pre-training (PPT) stage based on synthetic data with learnable temporal structure helps resist noisy data during the pre-training (PT) stage. % Concretely, we draw PPT sequences from many randomly initialized RNNs and keep the PT recipe unchanged. Across various corruption settings, our method consistently improves robustness to noise during PT, with larger relative gains at higher noise levels. % The gains generalize across corruption types and naturally noisy web corpora. For a 1B-parameter model, a synthetic PPT stage with only 65M tokens achieves the same final loss as the baseline while using up to 49% fewer natural-text PT tokens across different noise levels. Mechanistic analyses suggest PPT does not immediately suppress attention to noisy tokens. Rather, PPT-initialized models gradually downweight attention between corrupted tokens during noisy PT. This indicates that synthetic PPT inhibits noise self-modeling and shapes the subsequent optimization trajectory.
Abstract: Real-world time series are inherently non-stationary, with trends, periodic patterns, and uncertainty evolving over time. While the Fourier domain offers a natural lens to model time series, current deep learning approaches do not explicitly model evolution and randomness in the Fourier spectra, which limits their ability to accurately predict both the expected trajectory and its uncertainty in non-stationary time series. Motivated by Evolutionary Spectra (ES) theory, we propose TimeES, a general framework that enables probabilistic and deterministic forecasting via the evolutionary spectra theory. Specifically, we derive a parameterizable evolutionary spectra formulation, recasting non-stationary random process modeling as learning an evolving representation modulated by random variables. Furthermore, we reduce the complexity of the estimated spectra from \mathcalO(NM) to \mathcalO(NK), where K \ll M/2, by exploiting Hermitian symmetry and energy-guided frequency selection. Based on a simple linear backbone, our proposed TimeES achieves consistent state-of-the-art performance across both deterministic and probabilistic forecasting tasks, with high efficiency and interpretability. Code is available at: \urlhttps://anonymous.4open.science/r/TimeES.
Abstract: Long-horizon embodied intelligence requires agents to improve through interaction, not merely to execute plans generated from static goals. A central challenge is therefore to transform past executions into knowledge that can shape future decisions. Minecraft provides a representative testbed for this problem, where tasks such as crafting tools, building redstone components, and obtaining diamond equipment involve long prerequisite chains and are frequently disrupted by missing tools, blocked paths, GUI failures, or stagnant execution. To this end, we propose MineEvolve, a knowledge-driven self-evolution framework that converts execution feedback into actionable behavioral knowledge. MineEvolve first uses \underline\emph\ding182Monitor to convert each subgoal execution into typed feedback, including state changes, inventory changes, failure types, progress signals, and stagnation indicators. \underline\emph\ding183Inducer then derives reusable skills from successful executions and remedies from failed or stagnant executions. \underline\emph\ding184Curator validates, merges, filters, and retrieves these knowledge entries, while \underline\emph\ding185Adaptor uses them to repair the unfinished part of the plan under repeated failures or stagnation. Experiments on the Minecraft MCU long-horizon task suite show that MineEvolve consistently improves performance across multiple language-model planners, with larger gains on high-dependency task groups. Ablation and knowledge-accumulation studies further demonstrate that converting execution signals into structured behavioral knowledge is an effective path toward self-evolving embodied agents in long-horizon environments. Our code is available at \urlhttps://anonymous.4open.science/r/MineEvolve-1B5B.
Abstract: Off-policy updates are inevitable in reinforcement learning (RL) for large language models (LLMs) due to rollout staleness from asynchronous training and mismatches between training and inference engines. Naive importance sampling gives an unbiased correction but suffers from high variance, which is amplified by unbounded ratios and autoregressive generation. Prior remedies either rely on scenario-specific engineering, or trade bias for variance via token-level clipping or sequence-level normalization, yet these approaches remain largely heuristic. We propose Variational sEquence-level Soft Policy Optimization (VESPO). By explicitly incorporating variance reduction into a variational formulation, we derive a principled closed-form reshaping kernel that operates directly on sequence-level importance weights, avoids token-level approximation and length normalization, and admits an explicit variance bound for the deployed kernel. Experiments on math reasoning and code generation show that VESPO maintains stable training under severe off-policy conditions (staleness up to 64×) and delivers consistent gains across both dense and Mixture-of-Experts (MoE) models, outperforming recent reshaping baselines under matched setup.
Authors:
Timm Haucke, Lauren Harrell, Justin Kay, Mary K Clapp, Sara BeeryAbstract: We increasingly use machine learning to label scientific datasets. The models we develop and deploy are improving all the time, but they are not and will likely never be perfect. Mistakes matter, as errors can propagate into our scientific understanding, particularly when systematically biased. Very reasonably, scientists thus review substantial proportions of ML-generated labels to verify or correct mistakes in pursuit of ensuring their scientific findings are not biased by ML. In this work, we focus on helping scientists optimally allocate this reviewing effort relative to their scientific goals. We focus on a specific class of scientists (ecologists) and a specific, widespread, and impactful modeling target (occupancy modeling, which estimates where species are likely to occur, conditioned on environmental factors). We introduce ACORN (Active Continuous-Score Occupancy Modeling), a method that incorporates ML predictions into occupancy models and strategically selects samples for expert review that are maximally informative for downstream ecological analysis. Across camera-trap and bioacoustic datasets, our method recovers ecological conclusions close to those obtained from fully human-labeled data, while requiring substantially fewer expert reviews than non-targeted review policies. Our results suggest that ML-assisted scientific workflows should optimize expert effort for downstream inference, rather than for classifier accuracy alone, especially when human review budget is limited.
Abstract: Proving theorems in Lean 4 often requires identifying a scattered set of library lemmas whose joint use enables a concise proof---a task we call . Existing tools address adjacent problems: semantic search engines find individual declarations matching a query, while premise-selection systems predict useful lemmas one tactic step at a time. Neither recovers the full premise set an entire theorem requires. We present LeanSearch v2, a two-mode retrieval system for this task. Its standard mode applies a hierarchy-informalized Mathlib corpus with an embedding--reranker pipeline, achieving state-of-the-art single-query retrieval without domain-specific fine-tuning (nDCG@10 of 0.62 vs. 0.53 for the next-best system). Its reasoning mode builds on standard mode as its retrieval substrate, targeting global premise retrieval through iterative sketch-retrieve-reflect cycles. On a 69-query benchmark of research-level Mathlib theorems, reasoning mode recovers 46.1% of ground-truth premise groups within 10 retrieved candidates, outperforming strong reasoning retrieval systems (38.0%) and premise-selection baselines (9.3%) on the same benchmark. In a controlled downstream evaluation with a fixed prover loop, replacing alternative retrievers with LeanSearch v2 yields the highest proof success (20% vs. 16% for the next-best system and 4% without retrieval), confirming that retrieval quality propagates to proof generation. All code, data, and benchmarks will be open-sourced.
Abstract: Despite their importance in model sampling, efficient implementation of Top-k and Top-p algorithms for large vocabularies remains a significant challenge. Existing approaches often rely on sorting, which incurs significant computation and memory overhead on GPUs, or on stochastic approaches that alter the algorithm's output. In this work, we propose Qrita, an efficient Top-k and Top-p algorithm based on a pivot-based truncation and selection. Qrita leverages pivot-based search for both Top-k and Top-p with two key techniques: 1. Gaussian-based σ-truncation, which greatly reduces the search space of the vocabulary, and 2. Quaternary pivot search with duplication handling, which halves the number of pivot search iterations and guarantees deterministic output. We implement Qrita using Triton and evaluate its performance against the Top-k and Top-p kernels of high-performance LLM execution engines such as vLLM, SGLang, and FlashInfer, improving end-to-end serving throughput up to 1.4× with half the memory usage, while providing the same output as the sorting-based algorithms.
Abstract: The development of separate-encoder Unified multimodal models (UMMs) comes with a rapidly growing inference cost due to dense visual token processing. In this paper, we focus on understanding-side visual token reduction for improving the efficiency of separate-encoder UMMs. While this topic has been widely studied for MLLMs, existing methods typically rely on attention scores, text-image similarity and so on, implicitly assuming that the final objective is discriminative reasoning. This assumption does not hold for UMMs, where understanding-side visual tokens must also preserve the model’s capabilities for editing images. We propose G^2TR, a generation-guided visual token reduction framework for separate-encoder UMMs. Our key insight is that the generation branch provides a task-agnostic signal for identifying understanding-side visual tokens that are not only semantically relevant but also important for latent-space image reconstruction and generation. G^2TR estimates token importance from consistency with VAE latent, performs balanced token selection, and merges redundant tokens into retained representatives to reduce information loss. The method is training-free, plug-and-play, and applied only after the understanding encoding stage, making it compatible with existing UMM inference pipelines. Experiments on image understanding and editing benchmarks show that G^2TR substantially reduces visual tokens and prefill computation by 1.94× while maintaining both reasoning accuracy and editing quality, outperforming baselines on almost all benchmarks.
Abstract: Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a single learnable embedding trained as an explicit retrieval target that selects a sparse set of query-relevant visual tokens from a pre-filled visual KV cache. Trained on only a small image-QA dataset, ReToken yields consistent gains across image and video benchmarks: on Visual Haystacks it improves Qwen3VL-8B by 12 points (>20% relative), and on LVBench it transfers zero-shot to long video for an 8.5-point gain. Thanks to its lightweight design, both training and long-video inference fit on a single H100. We will release our code.
Authors:
Yueh-Cheng Liu, Jozef Hladký, Matthias Niessner, Angela DaiAbstract: Recent advances in 3D Gaussian Splatting (3DGS) present two main directions: feed-forward models offer fast inference in sparse-view settings, while per-scene optimization yields high-quality renderings but is computationally expensive. To combine the benefits of both, we introduce Diff3R, a novel framework that explicitly bridges feed-forward prediction and test-time optimization. By incorporating a differentiable 3DGS optimization layer directly into the training loop, our network learns to predict an optimal initialization for test-time optimization rather than a conventional zero-shot result. To overcome the computational cost of backpropagating through the optimization steps, we propose computing gradients via the Implicit Function Theorem and a scalable, matrix-free PCG solver tailored for 3DGS optimization. Additionally, we incorporate a data-driven uncertainty model into the optimization process by adaptively controlling how much the parameters are allowed to change during optimization. This approach effectively mitigates overfitting in under-constrained regions and increases robustness against input outliers. Since our proposed optimization layer is model-agnostic, we show that it can be seamlessly integrated into existing feed-forward 3DGS architectures for both pose-given and pose-free methods, providing improvements for test-time optimization.
Abstract: Subject-driven image generation aims to synthesize new images that preserve the identity of the given subject while following textual instructions. Existing approaches often encode text and reference images separately. This limits cross-modal reasoning abilities and causes copy-paste artifacts. Recent frameworks that connect multimodal models and diffusion models improve instruction following, but largely overlook identity preservation. To address these limitations, we condition diffusion models on Multimodal Large Language Models (MLLMs) that jointly encode text and reference images, and augment it with VAE‑based identity conditioning. A novel Dual Layer Aggregation (DLA) module is designed to aggregate multi‑level MLLM features for optimal conditioning, and a multi‑stage denoising strategy is applied to progressively balance the semantic information from MLLM and fine‑detail identity from VAE during inference. Extensive experiments demonstrate that our approach harmonizes multimodal understanding with identity preservation, mitigates copy-paste issues, and achieves superior performance regarding human preference on subject-driven image generation.
Abstract: Open-ended image generation is no longer a simple prompt-to-image problem. High-quality generation often requires an agent to combine a model's internal generative ability with external resources. As requests become more diverse and demanding, we aim to develop a general image-generation agent that can self-evolve through trajectories and use tools more effectively across varied generation challenges. To this end, we propose GenEvolve, a self-evolving framework based on Tool-Orchestrated Visual Experience Distillation. In GenEvolve, each generation attempt is modeled as a tool-orchestrated trajectory, where the agent gathers evidence, selects references, invokes generation skills, and composes them into a prompt-reference program. Unlike existing agentic generation methods that mainly rely on image-level scalar rewards, GenEvolve compares multiple trajectories for the same request and abstracts best-worst differences into structured visual experience, provided only to a privileged teacher branch. Inspired by on-policy self-distillation, Visual Experience Distillation provides dense token-level supervision, helping the student internalize better search, knowledge activation, reference selection, and prompt construction. We further construct GenEvolve-Data and GenEvolve-Bench. Experiments on public benchmarks and GenEvolve-Bench show substantial gains over strong baselines, achieving state-of-the-art performance among current image-generation frameworks.
Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing reasoning mechanisms still struggle to provide planning-oriented intermediate representations: textual Chain-of-Thought (CoT) fails to preserve continuous spatiotemporal structure, while latent world reasoning remains difficult to use as a direct condition for action generation. In this paper, we propose CoWorld-VLA, a multi-expert world reasoning framework for autonomous driving, where world representations serve as explicit conditions to guide action planning. CoWorld-VLA extracts complementary world information through multi-source supervision and encodes it into expert tokens within the VLA, thereby providing planner-accessible conditioning signals. Specifically, we construct four types of tokens: semantic interaction, geometric structure, dynamic evolution, and ego trajectory tokens, which respectively model interaction intent, spatial structure, future temporal dynamics, and behavioral goals. During action generation, CoWorld-VLA employs a diffusion-based hierarchical multi-expert fusion planner, which is coupled with scene context throughout the joint denoising process to generate continuous ego trajectories. Experiments show that CoWorld-VLA achieves competitive results in both future scene generation and planning on the NAVSIM v1 benchmark, demonstrating strong performance in collision avoidance and trajectory accuracy. Ablation studies further validate the complementarity of expert tokens and their effectiveness as planning conditions for action generation.
Abstract: Quantization-aware training (QAT) is a key approach for adapting large language models (LLMs) to extremely low-bit weights, where post-training quantization often suffers severe accuracy loss. However, most extremely low-bit QAT methods rely on the straight-through estimator (STE), which keeps hard round-and-clip weights in the forward pass while using an identity surrogate in the backward pass, creating a mismatch between discrete forward states and continuous latent-weight updates. As a result, small updates may fail to cross quantization boundaries, causing dead-zone stagnation in sub-optimal discrete states. To address this problem, we propose Hestia, an operator-faithful differentiable QAT framework for extremely low-bit LLMs. Instead of repairing the gradient of a hard quantizer, Hestia relaxes the quantizer itself as a temperature-controlled Softmax expectation over the same discrete codebook. The relaxation provides forward--backward consistent gradients during training and anneals to the target hard quantizer for inference. To stabilize this soft-to-hard transition, Hestia further uses offline Hessian traces as lightweight tensor-wise sensitivity signals for adaptive temperature scheduling. Experiments on Llama-3.2 show that Hestia consistently improves over ternary QAT baselines, with relative average zero-shot gains of 5.39% and 4.34% on 1B and 3B models, while adding only about 0.2% training-time overhead and no inference-time overhead. The code is available at https://github.com/hestia2026/Hestia.
Abstract: Blind source separation (BSS) is a natural framework for studying how latent causes may be recovered from sensory mixtures, but deriving online and biologically plausible algorithms for structured (i.e., constrained to known domains) and potentially correlated sources remains challenging. Recent work has derived neural networks for BSS from maximization of an entropy measure, yet its online implementations involve complex and nonlocal recurrent dynamics. Motivated by this perspective, we propose Predictive Entropy Maximization, which achieves competitive performance in BSS, using only local weight updates. The method employs a close approximation of an entropy measure, yielding an objective function with easily interpretable components. Minimizing this objective leads to a predictive neural architecture in which feedforward synapses follow an error-driven rule (that can be realized through dendritic mechanisms), lateral inhibitory connections are learned with local Hebbian plasticity, and source-domain constraints are enforced through simple output nonlinearities. We derive explicit spectral bounds on the surrogate error, characterizing when the approximation is accurate. Empirically, Predictive Entropy Maximization remains robust under increasing source correlation and observation noise, outperforms biologically plausible algorithms that rely on stronger independence or decorrelation assumptions, and remains competitive with exact determinant- and correlative-information-based baselines. These results show how local plasticity and adaptive lateral inhibition can emerge from maximizing a regularized second-order entropy over structured source domains.
Authors: Pablo Marcos-Manchón, Rishi Jha, Lluís Fuentemilla
Abstract: The Strong Platonic Representation Hypothesis suggests that representational convergence in artificial neural networks can be harnessed constructively: embeddings can be translated across models through a universal latent space without paired data. We ask whether an analogous geometry can be recovered across human brains. Using fMRI data from the Natural Scenes Dataset, we propose a self-supervised encoder that learns subject-specific embeddings from brain data alone by exploiting repeated stimulus presentations. We show that these independently learned spaces can be translated across subjects using unsupervised orthogonal rotations, without paired cross-subject samples or intermediate model representations. Synchronizing pairwise rotations into a single shared latent space further improves cross-subject retrieval, indicating that subject-specific spaces are mutually compatible with a common coordinate system. These results provide evidence for a shared neural geometry in the human visual cortex: subject-specific fMRI representations are approximately isometric across individuals and can be translated through purely geometric transformations.
Abstract: Universal visual anomaly detection (AD) aims to identify anomalous images and segment anomalous regions towards open and dynamic scenarios, typically adhering to zero- and few-shot paradigms without dataset-specific fine-tuning. Recently, the field has seen significant progress driven by the integration of vision-language foundation models. However, we observe that current methods often struggle with laborious prompt engineering, elaborate adaptation modules, and complex training strategies that ultimately constrain their flexibility and generalizability. In this paper, we rethink the fundamental mechanism of vision-language models for AD and present UniADet, an embarrassingly simple, effective and general framework for Universal vision Anomaly Detection. Our approach is built on two key insights: First, we reveal that the primary function of the language encoder is merely to derive decision weights and we demonstrate that it is unnecessary, as these weights can be learned more directly and efficiently. Second, to resolve learning conflicts arising from disparate feature manifolds, we introduce a systematic decoupling strategy that learns independent weights across both distinct tasks (classification vs. segmentation) and hierarchical features. UniADet is highly simple and efficient, learning only decoupled weights and requiring 0.02M learnable parameters. It is inherently versatile, adapting seamlessly to various foundation models (\eg, CLIP, DINOv2, and DINOv3). Extensive evaluations of 14 real-world benchmarks in the industrial and medical domains demonstrate that UniADet not only exceeds state-of-the-art zero/few-shot methods by a substantial margin, but also outperforms full-shot AD methods for the first time. This empirical evidence reveals the profound potential of language-free frameworks to redefine the boundaries of visual anomaly detection. The code and models will be made publicly available.
Authors: EUNBEE YOUN, Donghan Kim, Dae W Kim
Abstract: Data assimilation is the process of estimating the state of a dynamical system over time by combining model predictions with measurements. This task becomes challenging when the system is nonlinear and high-dimensional. To address this, score-based Bayesian filters have recently emerged. However, these methods still show unsatisfactory performance in certain cases, particularly under spatially sparse measurements. Such degradation stems from heuristic approximations of the likelihood score, whose errors can accumulate over time. This limitation arises because the methods simply adopt a classical forward process for generative modeling that transforms a data distribution toward a Gaussian distribution, which is independent of the measurement equation. Here, we propose a forward process tailored for filtering that transforms the system state toward the measurement space, enabling a theoretically sound formulation of the likelihood score. Based on this, we develop the Measurement-Aware Score-Based Filter (MASF). We evaluate MASF on the Kolmogorov flow, a high-dimensional fluid benchmark with up to \mathcalO(10^5) dimensions, under diverse measurement operators, including nonlinear cases with dimensional mismatch between the state and the measurements. MASF shows improved performance over existing score-based filters and ensemble-type Kalman filters. With amortized pretraining, MASF also achieves up to a 28.2× wall-clock speedup compared with the baselines.
Abstract: Visual Autoregressive (VAR) models have emerged as a strong alternative to diffusion for image synthesis, yet their fixed training resolution prevents direct generation at higher resolutions. Naively transferring training-free extrapolation methods from LLMs or diffusion models to VAR yields three characteristic failure modes: global repetition, local repetition, and detail degradation. We trace them to a unified band-stage mismatch: VAR generates images in a coarse-to-fine, scale-wise process where each stage is driven by a distinct dominant RoPE frequency band, and each failure mode emerges when the dominant band of a particular stage is disrupted. Building on this insight, we propose Stage-Aware RoPE Remapping, a training-free strategy that assigns each frequency band a stage-specific remapping rule, jointly suppressing all three failure modes. We further observe that attention becomes systematically dispersed as the image resolution increases. Existing methods typically depend on predefined attention scaling factors, which are neither adaptive to the target resolution nor capable of faithfully capturing the actual extent of attention dispersion. We therefore propose Entropy-Driven Adaptive Attention Calibration, which quantifies dispersion via a resolution-invariant normalized entropy and yields a closed-form per-head scaling factor that realigns the extrapolated-resolution attention entropy with its training-resolution counterpart. Extensive experiments show that our method consistently outperforms prior resolution-extrapolation methods in both structural coherence and fine-detail fidelity. Our codes have been released in supplementary materials and will be released on Github.
Abstract: We introduce Tadpole, a novel foundation model for three-dimensional partial differential equations (PDEs) that addresses key challenges in transferability, scalability to high dimensionality, and multi-functionality. Tadpole is pre-trained as an autoencoder on synthetic 3D PDE data generated by an efficient online data-generation framework. This enables large-scale, diverse training without storage or I/O overhead, demonstrated by scaling to an equivalent of hundreds of terabytes of training data. By autoencoding single-channel spatial crops, Tadpole learns rich and transferable representations across heterogeneous physical systems with varying numbers of state variables and spatial resolutions. Although pre-trained solely as an autoencoder, Tadpole can be efficiently applied for multiple downstream tasks beyond reconstruction, including dynamics learning and generative modeling. For dynamics learning, we propose a novel parameter-efficient fine-tuning strategy that integrates low-rank adaptation, latent-space transformations, and reintroduced skip connections, achieving accurate temporal modeling with a minimal number of trainable parameters. Tadpole demonstrates strong fine-tuning performance across various downstream tasks, highlighting its versatility and effectiveness as a foundation model for 3D PDE learning.
Abstract: Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely evaluate single agents, leaving multi-agent risks such as coordination failure and conflict poorly understood. We introduce GT-HarmBench, a benchmark of 1,535 high-stakes scenarios spanning game-theoretic structures such as the Prisoner's Dilemma, Stag Hunt and Chicken. Scenarios are drawn from realistic AI risk contexts in the MIT AI Risk Repository. Across 15 frontier models, agents fail to choose socially beneficial actions in 38% of high-stakes cases, such as military escalation, election manipulation, and medical malpractice. We measure sensitivity to game-theoretic prompt framing and ordering, and analyze reasoning patterns driving failures. We further show that game-theoretic interventions improve socially beneficial outcomes by up to 18%. Our results highlight substantial reliability gaps and provide a broad standardized testbed for studying alignment in multi-agent environments.
Authors:
Afiq A Aswadi, Haotong Ma, Susan WeiAbstract: A \emphBayes-filtered transformer (BFT) is a transformer trained on sequences that are generated in two steps: first a latent task is drawn from a prior, then observations are drawn conditional on that task. Trained under autoregressive log loss, the BFT's next-token prediction, in the idealised limit, \emphis the Bayesian posterior predictive distribution (PPD) under this generative model. In practice the trained BFT is only an approximation of this ideal PPD, raising a natural interpretive question: what prior and posterior over the latent task has the trained BFT actually internalised? Existing work answers this question by comparing the trained BFT's predictions against the predictions of various ``reference'' posteriors, each standing in for a different candidate algorithm or computation the BFT might be implementing. This prediction-space comparison is fragile; for instance, distinct posteriors can share posterior-mean predictions exactly. We foreground \emphpredictive Monte Carlo (PMC) as a general interpretability tool for any BFT: using only next-token generation, PMC returns an approximation to the implicit prior and posterior over the latent task, answering the interpretive question directly in latent space. We apply PMC to three stylised task families spanning 0-Markov and 1-Markov exchangeability; the phenomena previously reported in these settings remain visible in latent space.
Abstract: Inverse problems for stiff parabolic partial differential equations (PDEs), such as the inverse heat conduction problem (IHCP), are severely ill-posed: the forward map rapidly damps high-frequency interior structure before it reaches the boundary. Soft-constrained physics-informed neural networks (PINNs), which embed the PDE as a residual penalty, suffer from gradient pathology in this regime and tend to fit boundary measurements while leaving the interior field essentially untouched. We propose Neural Field Thermal Tomography (NeFTY), a hard-constrained neural field framework for label-free three-dimensional inverse heat conduction. NeFTY represents the unknown diffusivity as a continuous coordinate-based neural network, and at every optimization step passes the candidate field through a differentiable implicit-Euler heat solver with harmonic-mean interface flux, so that the governing PDE holds exactly on the discretization rather than as a soft penalty. Adjoint gradients propagate the surface reconstruction error back to the network weights at solver-level memory cost, making test-time inversion tractable on a single GPU. Across synthetic 3D benchmarks, NeFTY substantially outperforms soft-constrained PINN variants and a voxel-grid baseline on label-free volumetric recovery, and it transfers to real thermography data, surpassing classical signal-processing baselines in both defect segmentation and depth estimation.
Abstract: The recently introduced 4-bit floating-point format, NVFP4, demonstrates remarkable performance and memory benefits for quantized large language model (LLM) inference. However, we observe two types of redundancy in the existing NVFP4 encoding: (1) The FP4 format naturally exposes an unused quantization value due to its sign-magnitude representation that contains both positive and negative zeros. (2) The FP8 block scaling factor, with 4-bit exponent and 3-bit mantissa, contains an unused sign bit given that the scaling factor is always positive. Additionally, we find that LLM weights are more tolerant to a lower-precision block scaling factor, such as 6 bits with 3-bit exponent and 3-bit mantissa. Based on these observations, we propose emapping (RaZeR), an enhanced numerical format that pushes the limits of NVFP4 for more accurate LLM quantization under the same memory footprint. RaZeR leverages the unused bits in the block scaling factor to adaptively remap the negative FP4 zero to a set of pre-defined special values, which maximally utilizes the NVFP4 encoding and better fits the LLM tensor distribution. Extensive experiments validate RaZeR’s superior performance for 4-bit LLM quantization. For example, RaZeR reduces the average perplexity loss of NVFP4 by 29.4% and 32.3% under weight-only and weight-activation quantization, respectively.
Abstract: Grounded long-video question answering, or Grounded LVQA, is the task of answering a question about a long video while also locating the short time interval in the video that supports the answer. Recent agentic methods approach this problem as a multi-turn exploration process with a single crop video(start, end) action. This action allows the model to progressively narrow its search from coarse to fine regions, but it does not provide a direct way to backtrack from a fine-grained mistake to a broader context. As a result, these agents often stop after only two turns and are unable to recover once they descend into the wrong part of the video. We propose VideoTreeSearch (VTS), a framework that formulates grounded LVQA as an iterative, self-correcting search over an adaptive temporal tree. VTS builds a non-uniform tree from scene boundaries so that each node corresponds to a semantically coherent video segment. It then trains an agent to navigate this tree using four discrete actions: zoom in, zoom_out, shift, and answer. These actions make backtracking and recovery explicit and learnable, rather than leaving them as implicit behaviors. To train the agent, we introduce a trajectory synthesis pipeline that generates multi-step navigation paths through the tree, including intentional detours into incorrect branches followed by recovery. These trajectories are first used for supervised fine-tuning and then for reinforcement learning with rewards based on grounding quality and answer accuracy. On three Grounded LVQA benchmarks—CG-Bench, Haystack-LVBench, and Haystack-Ego4D—VTS outperforms the strongest previous agentic methods by 12.5 mIoU on CG-Bench and 7.4 T-F1 on Haystack-Ego4D. The learned policy also transfers to general long-video question answering, surpassing all prior agentic baselines on Video-MME, MLVU, and LVBench by up to 7.1 accuracy points. Ablation studies show that self-correcting hierarchical search is the key factor behind these improvements: removing either adaptive descent or explicit backtracking leads to substantial performance drops.
Abstract: There is a fundamental limit to the best prediction error of any model for a given data distribution. In classification, the Bayes error quantifies this limit and serves as both a benchmark for trained models and a criterion to detect overfitting. The state-of-the-art Bayes error estimation techniques for binary classification use only soft labels (positive class probabilities) and require no input features or auxiliary training. We extend such notions to constrained classification problems, such as fairness. Fair classification is an important case of constrained or multi-objective classification. It requires a model not only to be accurate but also to yield fair outcomes (equal or near-equal true positive rates) across different demographic groups (e.g., race and gender). This leads to an inherent fairness-accuracy tradeoff or a Pareto frontier that is used to compare fair classifiers against each other. The focus of our work is to estimate the fundamental limit for this fairness-accuracy tradeoff. We derive an explicit expression for the fair Bayes error rate (i.e., best achievable error rate subject to given fairness constraints), and construct consistent, instance-free estimators for it using only soft labels and group information. We provably bound the bias and the sample complexity of our estimators. Empirically, our method produces a robust, low-variance estimate of the optimal fairness-accuracy trade-off curve, even when the soft labels are noisy and derived from a classification model on hard labels.
Authors: Anuska Roy, Pravin Nair
Abstract: Flow and diffusion models achieve high-fidelity, high-resolution image synthesis, but often require many function evaluations (NFEs) at sampling time. Existing acceleration methods either require additional training through distillation or rely on training-free high-order solvers, and both can degrade sample quality at low NFE budgets. We propose CAB (Corrected Adams-Bashforth), a training-free sampler that accelerates both flow and diffusion models. CAB first transforms the sampling dynamics to a common rectified coordinate system, and then applies a multistep Adams--Bashforth predictor augmented with a simple correction term based on past velocity evaluations and therefore incurs no additional NFEs. The resulting method is simple, has the same algorithmic form across model classes, and has at least third-order local truncation error and second-order global error. Experiments on pretrained flow and diffusion models, including class-conditional and large-scale text-to-image benchmarks, show that CAB improves quality--NFE trade-offs in the low-step regime of (6)--(20) NFEs. It also remains competitive with strong training-free samplers at higher step counts across most tested models.
Abstract: Efficient key-value (KV) cache management is crucial for the practical deployment of large language models (LLMs), yet existing compression techniques often incur a trade-off between performance degradation and computational overhead. We propose a novel gating-based KV cache eviction method for frozen-weight LLMs that achieves high compression ratios with negligible computational cost. Our approach introduces lightweight sink-attention gating modules to identify and retain critical KV pairs, and integrates seamlessly into both the prefill and decoding stages. The proposed gate training algorithm relies on forward passes of an LLM, avoiding expensive backpropagation, while achieving strong task generalization through a task-agnostic reconstruction objective. Extensive experiments across the Qwen2.5-1M, Qwen3, and Gemma3 families show that our method maintains near-lossless performance while evicting up to 70% of the KV cache. The results are consistent across a wide range of tasks, including long-context understanding, code comprehension, and mathematical reasoning, demonstrating the generality of our approach.
Abstract: LLM agents are expected to act over multiple turns, using search, browsing interfaces, and terminal tools to complete user goals. Yet not every goal is well specified or achievable in the available environment. In such cases, a reliable agent should recognize that further interaction is unlikely to help and abstain from additional tool calls. We define Agentic Abstention, the problem of deciding when an agent should stop acting under uncertainty. Unlike standard LLM abstention, which is usually evaluated as a single-turn answer-or-abstain decision, agentic abstention is a sequential decision problem: an agent can answer, abstain, or gather more information at each turn, and the need to abstain may only become clear after interacting with the environment. We study this problem across web shopping, terminal environments, and question answering, evaluating 13 LLM-as-agent systems and 2 agent scaffolds on more than 28,000 tasks. Our results show that the main challenge is not only whether agents can abstain, but also when they abstain. Some agents never abstain when they should, while others do so only after many unnecessary interactions. This gap is especially large on tasks where the instruction appears feasible until the environment reveals otherwise (e.g., no valid result matches the instruction). We further find that model scale, reasoning, and agent scaffolding affect abstention in different ways, with larger or more capable models not always performing better at timely abstention. Finally, we introduce CONVOLVE, a context engineering method for improving agentic abstention that distills full interaction trajectories into reusable stopping rules. On WebShop, CONVOLVE substantially improves timely abstention without updating model parameters, raising Llama-3.3-70B's timely recall rate from 26.7 to 57.4. % , and overall recall from 83.2 to 100.0. Our dataset and code are available at https://anonymous.4open.science/r/agentic-abstention-A908
Abstract: Discrete diffusion models have emerged as a promising direction for vision-language tasks, offering bidirectional context modeling and theoretical parallelization. However, their practical application is severely hindered by a train-inference discrepancy, which leads to catastrophic error cascades: initial token errors during parallel decoding pollute the generation context, triggering a chain reaction of compounding errors and leading to syntactic errors and semantic hallucinations. To address this fundamental challenge, we reframe the generation process from passive denoising to active refining. We introduce ReDiff, a refining-enhanced diffusion framework that teaches the model to identify and correct its own errors. Our approach features a two-stage training process: first, we instill a foundational revision capability by training the model to revise synthetic errors; second, we implement a novel online self-correction loop where the model is explicitly trained to revise its own flawed drafts by learning from an expert's corrections. This mistake-driven learning endows the model with the crucial ability to revisit and refine its already generated output, effectively breaking the error cascade. Extensive experiments demonstrate that ReDiff significantly improves the coherence and factual accuracy of generated content, enabling stable and efficient parallel generation far superior to traditional denoising methods.
Abstract: Understanding human behavior while interacting with the surrounding world is crucial for many applications of embodied AI. First-person videos are particularly informative for this problem, as they well capture how activities reshape the scene over time. However, existing approaches often rely on implicit visual or language-aligned representations, disregarding structured reasoning over the scene dynamic. In this paper we argue that explicit, compositional and editable representations of human-environment interactions can play a crucial role for rich grounded activity understanding. To this end, we introduce SG-Ego, a large scale annotation set extending Ego4D with spatio-temporal scene graphs, where relations triplets are consolidated over time into explicit time-evolving descriptions of the scene state. To effectively reason over this representation, we propose GLEN, a graph-based model that operates over scene graph sequences to both align them with textual actions and model their temporal evolution. In addition, we formulate the activity-driven graph-edit forecasting (A-GEF) problem, a novel task that casts scene dynamics as a sequence of structured transformations conditioned on ongoing actions, enabling explicit reasoning about how scenes change over time. We demonstrate the effectiveness of our approach across multiple downstream tasks, spanning retrieval benchmarks as EgoMCQ and the more complex compositional EgoCVR, as well as long-horizon reasoning benchmarks as EXPLORE-Bench and the newly introduced A-GEF. GLEN achieves competitive or state-of-the-art performance compared to raw video baselines and it excels in reasoning settings, typically addressed only with MLLMs, while enabling controllable and structured predictions of scene dynamics driven by human activities. We believe our results establish spatio-temporal scene graphs, together with models that reason over them, as strong compositional and interpretable representations for egocentric video understanding and potentially beyond. Code will be released upon acceptance.
Authors:
Darryl Cherian Jacob, Xinyu Liu, Kai Wang, Pan HeAbstract: Vision-language models (VLMs) have demonstrated impressive performance in video anomaly detection (VAD), while providing interpretable predictions. However, we identify a key yet underexplored VAD issue: the mismatch between training and inference in both data distribution and model configuration. First, many VLM-based VAD approaches rely on a static parameter paradigm after training (or post-adaptation), causing them to overfit the training distribution and limiting generalization to unseen domains. Consequently, these methods lack effective test-time adaptation under substantial distribution shifts, such as novel environments or previously unseen anomaly types. Second, they treat VLMs as reasoning-based classifiers trained on extremely sparse frames from long videos, but perform inference on densely sampled short segments via sliding windows, leading to inherent inconsistencies between training and testing. To address these limitations, we propose COPRA, a conditional parameter adaptation framework for VLM-based VAD. Rather than relying on static post-training adaptation (e.g., fixed prompts or shared parameter updates), COPRA introduces instance-conditioned parameter adaptation, where a lightweight generator predicts input-specific parameter updates to dynamically modulate a frozen VLM on the fly, enabling per-segment adaptation during both training and inference. Empirically, our method achieves strong performance on standard VAD benchmarks. More broadly, it consistently outperforms static baselines in both in-domain and cross-domain settings, and further generalizes beyond VAD to unseen tasks such as multiple-choice Video Question Answering and Dense Captioning. These results suggest that COPRA provides an effective weight space generation mechanism for foundation models, enabling more scalable, adaptive, and context-aware video understanding.
Authors:
Aggelos Psiris, Yannis Panagakis, Maria Vakalopoulou, Georgios T PapadopoulosAbstract: Few-Shot Industrial Visual Anomaly Detection (FS-IVAD) is a critical task in modern manufacturing, where automated product inspection systems must identify rare defects using only a handful of normal, defect-free training samples. This paper introduces GATE-AD, a reconstruction-based framework that casts few-shot normality modeling as a masked, representation-aligned graph reconstruction problem on a k-nearest-neighbor graph, built from frozen self-supervised ViT patch tokens. Since defects typically disrupt the contextual consistency between a patch and its spatial neighbors, GATE-AD emphasizes such neighborhood relations by attending anisotropically over each patch's local neighborhood using a Graph Attention Network (GAT) encoder. To prevent GAT over-smoothing in the low-shot setting, the encoder output is aligned with the input ViT tokens through a learnable latent space, where reconstruction inconsistency is scored with a Scaled Cosine Error (SCE) objective. On the MVTec AD, VisA, and MPDD benchmarks, GATE-AD attains state-of-the-art image-level AUROC in the wide majority of 1- to 8-shot settings, with low per-image inference cost that does not scale with the size of the support set. Code is included in the supplement material and will be publicly released upon acceptance.
Abstract: Feed-forward 3D reconstruction models, represented by Visual Geometry Grounded Transformer (VGGT), jointly predict multiple visual geometry tasks such as depth estimation, camera pose prediction, and point cloud reconstruction in a single forward pass. They have been widely adopted in 3D vision applications, but their billion-scale parameters bring substantial memory and computation overhead, posing challenges for on-device deployment. Post-Training Quantization (PTQ) is an effective technique to reduce this overhead. Existing PTQ methods for feed-forward 3D models mainly focus on handling heavy-tailed activation distributions and constructing diverse calibration datasets. However, we observe that feed-forward 3D models predict multiple geometric attributes through a shared backbone, where different transformer blocks and hidden channels contribute distinctly to each task, resulting in substantially different sensitivities to quantization errors across tasks, blocks, and channels. Consequently, treating all tasks equally over-emphasizes insensitive tasks and causes significant accuracy loss on the sensitive ones. To address this issue, we propose Fisher-Guided Quantization (FGQ) for feed-forward 3D reconstruction models. Specifically, FGQ uses the diagonal Fisher information matrix to quantify the different sensitivities across tasks, blocks, and channels, and incorporates these sensitivities into the Learnable Affine Transformation during calibration to better preserve the channels and blocks most critical to each task. Extensive experiments across camera pose estimation, point map reconstruction, and depth estimation show that FGQ consistently outperforms state-of-the-art quantization baselines on VGGT, achieving up to 39% relative improvement under the 4-bit quantization.
Abstract: Monocular novel-view synthesis has long required multi-view image pairs for supervision, limiting training to a narrow set of purpose-built datasets. We propose in-the-wild monocular pretraining: a frozen depth estimator lifts each source image into 3D and reprojects under sampled poses to yield pseudo-target views; masked losses restrict supervision to valid regions and an adversarial objective covers disoccluded areas. Scaled to 30 million uncurated images, this produces OVIE, requiring only a source image and target pose at inference. Prior work trains without multi-view data but needs a depth estimator at inference, or drops this dependency but requires multi-view training pairs; OVIE is the first to require neither. Without multi-view supervision, OVIE rivals in-domain baselines on RealEstate10K and surpasses all on DL3DV; brief multi-view fine-tuning outperforms all geometry-free methods on their training domain. At 116 FPS, it is over 600× faster than the fastest baseline. Code and models will be open-sourced.
Abstract: Reinforcement learning (RL) can be used to improve the policy (denoiser) of diffusion large language models (dLLMs), while being hindered by the intractability of the policy likelihood. A dominant and efficient family of methods replaces the likelihood in standard RL with its evidence lower bound (ELBO), estimated from randomly masked sequences. Despite being well-aligned with pre-training, these approaches introduce bias from training-inference mismatch by applying the likelihood surrogate ELBO, which can potentially cause training collapse. In this work, we propose Guided Denoiser Self-Distillation (GDSD) to directly distill the denoiser of dLLMs from an advantage-guided self-teacher, derived from the closed-form optimum of reverse-KL regularized RL. GDSD matches the dLLM's denoiser logits to the teacher's via a normalization-free objective, which reduces RL to likelihood-free self-distillation and thus bypasses the TIM biases. Recent ELBO-based methods emerge as instances of applying different distillation divergences, but with diagnosable pathologies that GDSD avoids. On planning, math, and coding benchmarks with LLaDA-8B and Dream-7B, GDSD consistently outperforms prior state-of-the-art ELBO-based methods with a more stable training reward dynamics, achieving test-accuracy improvements of up to +19.6%. These results suggest that direct denoiser self-distillation, without relying on ELBO likelihood surrogate, can provide a more stable and effective RL procedure for dLLMs.
Authors:
Yuchuan Tian, Yuchen Liang, Shuo Zhang, Yingte Shu, Guangwen Yang, Wei He, Sibo Fang, Tianyu Guo, Kai Han, Chao Xu, Hanting Chen, Xinghao Chen, Yunhe WangAbstract: Diffusion language models (DLMs) can generate multiple tokens in parallel, but training large DLMs from scratch remains expensive. A practical alternative is to adapt off-the-shelf autoregressive (AR) checkpoints into diffusion models, reusing their linguistic and reasoning capabilities. Existing adaptation recipes either modify logits and grow attention masks toward full-sequence diffusion, or directly fine-tune ARweights under a block-diffusion objective, leaving two questions underexplored: what diffusion paradigm should AR-to-DLM adaptation target, and what transition path preserves AR knowledge most effectively? We argue that Block-Diffusion is a natural destination because AR decoding corresponds to block size one at the level of attention and generation order, while larger blocks introduce controlled intra-block bidirectionality and parallel generation. Based on this view, we propose a context-causal adaptation path that keeps committed context strictly causal, a one-pass parallel training formulation with auxiliary AR guidance, and a gradual block-size curriculum. Across several AR initializations and model scales, these components improve average adaptation performance over random mask annealing and direct fine-tuning. Scaling the recipe yields NBDIFF-7B, which supports 32K-token contexts and achieves the strongest average performance among the compared diffusion LLM baselines on general, math, and code benchmarks. Code and checkpoints will be released upon publication.
Abstract: World models for deformable objects should recover not only geometry and appearance, but also underlying physical dynamics, interaction grounding, and material behavior. Learning such a model from real videos is challenging because deformable linear, planar, and volumetric objects evolve under high-dimensional deformation, noisy interactions, and complex material response. The model must therefore infer a physical state from visual observations, roll it forward under new interactions, and render the resulting dynamics with high visual fidelity. We present DeformMaster, a video-derived interactive physics--neural world model that turns real interaction videos into an online interactive model of deformable objects within a unified dynamics-and-appearance framework. DeformMaster preserves structured physical rollout while using a neural residual to compensate for unmodeled effects, grounds sparse hand motion as distributed compliant actuator for hand--continuum interaction, represents material response with spatially varying constitutive experts, and drives high-fidelity 4D appearance from the predicted physical evolution. Experiments on real-world deformable-object sequences demonstrate DeformMaster's ability to roll out future dynamics and render dynamic appearance, outperforming state-of-the-art baselines while supporting novel action rollout, material-parameter variation, and dynamic novel-view synthesis.
Abstract: Diffusion-based large language models (dLLMs) rely on bidirectional attention, which prevents lossless KV caching and requires a full forward pass at every denoising step. Existing approximate KV caching methods reduce this cost by selectively updating cached states, but their decision overhead scales with context length or model depth. We propose EntropyCache, a training-free KV caching method that uses the maximum entropy of newly decoded token distributions as a constant-cost signal for deciding when to recompute. Our design is grounded in two empirical observations: (1) decoded token entropy correlates with KV cache drift, providing a cheap proxy for cache staleness, and (2) feature volatility of decoded tokens persists for multiple steps after unmasking, motivating recomputation of the k most recently decoded tokens. The skip-or-recompute decision requires only O(V) computation per step, independent of context length and model scale. Experiments on LLaDA-8B-Instruct and Dream-7B-Instruct show that EntropyCache achieves 15.2×–26.4× speedup on standard benchmarks and 22.4×–24.1× on chain-of-thought benchmarks against vanilla baselines, with competitive accuracy and decision overhead accounting for only 0.5% of inference time.
Abstract: Recent vision-language-action (VLA) models and world action models (WAMs) advance robotic manipulation by enriching intermediate representations with auxiliary spatial features or future visual-state prediction. However, these representations largely remain within the observation space and do not share the rigid-body geometry of the action space, forcing the action decoder to implicitly recover this geometry. We propose OASIS, a visuomotor policy that aligns the intermediate representation with the action space via SE(3) end-effector trajectory prediction. OASIS couples a 3D-aware feature encoder that fuses vision-language and metric-depth features with an SE(3) trajectory predictor that produces a camera-frame end-effector trajectory. Conditioned on the predictor's pose-supervised hidden states, the action decoder generates action chunks consistent with rigid-body motion. Across simulation and real-world experiments, OASIS outperforms VLA and WAM baselines in success rate and out-of-distribution generalization.
Authors: Niels Cariou-Kotlarek, Vasileios Lampos
Abstract: Rough path signatures provide a universal feature map for continuous paths and, via the expected signature, a principled characterisation of path distributions. These results do not directly extend to \clag paths of Temporal Point Processes (TPPs), limiting the use of signature methods for event sequences. Furthermore, neural TPP models, including recent generative approaches, optimise per-event objectives with no global sequence-level loss, while evaluation of variable-length event sequences lacks distributional discrepancy measures. This paper proposes a common pathwise framework for addressing these limitations. We introduce the interarrival embedding, a stable (homeomorphic) lift from jump paths to continuous paths of bounded variation, enabling signature methods to discrete event sequences. Our theoretical contributions give rise to \sigTPP, the first signature-based generative model for TPPs, trained using a path-level loss on complete trajectories. We further analyse the space of counting paths and derive three distributional discrepancies, providing mathematically justified tools for evaluating generative TPP models. Across synthetic and real-world datasets, \sigTPP~achieves the best average rank based on 8 complementary metrics, outperforms or is within one standard error of the strongest baseline in 64% of the dataset-metric pairs, and according to a relative score, improves against every baseline by at least 19% on average.
Abstract: Earth Observation (EO) forecasting aims to predict future Earth surface dynamics from satellite observations under changing meteorological conditions. In this paper, we view this task as a partially observed, weather-driven world modeling problem, in which weather acts as a conditioning signal, while forecasting remains uncertain due to sparse observations and unobserved land-surface states. However, existing methods do not fully capture this setting: deterministic models collapse uncertainty into a single future prediction, while diffusion-based methods typically treat weather variables as undifferentiated conditioning signals, and existing benchmarks focus mainly on reconstruction accuracy rather than whether forecasts respond correctly to changed weather forcing. We introduce EO-WM, a video diffusion transformer for multispectral EO forecasting. EO-WM incorporates a physically informed conditioning framework that represents meteorological forcing through a climatological baseline, weather anomalies, and cumulative physical stress signals. Specifically, it separates baseline and anomaly through distinct conditioning pathways, and accumulates anomalous forcing over time to capture sustained heat and drought stress. To evaluate weather-response behavior beyond standard metrics, we introduce two diagnostic benchmarks: an Extreme Summer Benchmark for severity-aware prediction of vegetation degradation under extreme weather, and a Seasonal Matched-Pair Benchmark for testing response fidelity under changed weather forcing. Experiments show that EO-WM reduces the error in predicted Normalized Difference Vegetation Index (NDVI) decline amplitude by a relative 5.63% and improves directional hit rate by a relative 7.80%, while remaining competitive on standard pixel-level metrics.
Authors: Yuan Wang
Abstract: In this work, we provide a complete characterization of the uniqueness of linear patterns for generic fully connected neural networks activated by ReLU or leaky ReLU functions. With this result, together with several tools from geometry and probability, we compare the numbers of linear regions for a wide variety of networks from both theoretical and practical perspectives, with minimal prerequisites. More specifically, on the theoretical side, we show that networks with neurons arranged as pq×2 (2 hidden layers, each having pq neurons) are expected to have more linear regions than those arranged as p× 2q for p\ge 1 and q\ge 2; on the empirical side, we show the validity of a Monte-Carlo approach for comparing linear regions via linear patterns, and carry out experiments for networks in various scenarios. Our results indicate that neurons closer to the input layer tend to have more impact on the number of linear regions. Moreover, along the way, we find the precise value of the expected number of linear regions for networks with 2 hidden layers, and an explicit upper bound for networks whose layer widths are non-increasing. The implementations for this paper are available through the following link: https://github.com/temp-repo-259848/Regions-ReLU-Network
Authors:
JinCheng Ren, Siwei Wu, Yizhi Li, Zhu, Shu XU, Boyu Feng, Ruibin Yuan, Wei Zhang, Riza Batista-Navarro, Min Gao, Yan Bai, Jian Yang, Chenghua LinAbstract: Terminal observations are not ordinary long-context text: they are heterogeneous, low-information-density execution traces in which sparse but exact evidence (e.g., error messages and file paths) is interleaved with large amounts of repetitive terminal output. For long-horizon CLI agents, retaining raw observations rapidly increases context cost and can dilute critical signals, while LLM-based summarization or fixed heuristics often fail to adapt across heterogeneous terminal tasks and may discard precise task-relevant evidence.We propose TACO, the first self-evolving Terminal Agent Compression framework, which treats compression rules as reusable, preservation-aware knowledge acquired from interaction trajectories. Rather than relying on manually designed rules, static pruning, or task-specific compressor training, TACO autonomously discovers, refines, and reuses structured compression rules from agent interaction trajectories. A Global Rule Pool accumulates effective rules across tasks, while task-time rule evolution adapts them online to the current workflow. This allows TACO to filter redundant terminal observations while conservatively preserving exact task-relevant evidence.Across six benchmarks, including TB 1.0, TB 2.0, SWE-Bench Lite, CompileBench, DevEval, and CRUST-Bench, TACO consistently maintains or improves task success across models and agent scaffolds. On TerminalBench, TACO yields 1%--4% absolute accuracy gains under standard evaluation and improves accuracy by 2%--3% under matched token budgets. On downstream benchmarks, it reduces total token consumption by 12%--27% while maintaining or improving task success. These results show that self-evolving observation compression can unlock latent capability in existing CLI agents by allocating context budget toward task-relevant evidence, without model fine-tuning or human-crafted compression rules.
Abstract: Unified multimodal models (UMMs) aim to integrate understanding and generation within a single architecture. However, it remains underexplored how to effectively coordinate these two capabilities for more effective and efficient reasoning. Existing coordination approaches either perform coupling during training, without explicit inference-time coordination, or impose a fixed coordination pattern for all inputs. In this work, we show that multimodal tasks exhibit substantial coordination-path diversity: different inputs favor different coordination paths. This suggests that exploiting such diversity is key to improving performance. We propose UniPath, a framework for adaptively modeling and exploiting coordination-path diversity. Instead of enforcing a single coordination pattern, we represent task solving as the selection and execution of a path, ranging from direct answering to textual inference, visual-thought construction, and hypothesis-based exploration. We construct role-aligned trajectories to train a path-conditioned executor and introduce a lightweight planner mechanism to enable input-dependent path selection. Experiments show that leveraging coordination-path diversity improves performance over fixed coordination strategies while providing interpretable intermediate behaviors. The code is available at: https://anonymous.4open.science/r/UniPath.
Authors:
Zheyuan Zhou, Liang Du, Zixun Sun, Xiaoyu Zhou, Ruimin Ye, Qihao Chen, Yinda Chen, Lemiao QiuAbstract: Despite advances in vision-language-action (VLA) models, they remain limited in highly dynamic game settings such as 3D open worlds and competitive player versus player (PvP) settings, where agents must efficiently extract sparse, actionable signals from dense visual input while integrating multimodal cues to track off-screen high-value targets in real time. To address this, we introduce 3A-VLA (Abstraction-Aligned Action VLA), a framework that grounds action learning in explicit intention and environment abstractions rather than superficial pattern matching. We introduce dual task-agnostic abstractions: the intention abstraction (IA), which condenses verbose instructions and reasoning into explicit semantic primitives, and the environment semantics abstraction (ESA), which structures dense visual streams into a spatial--functional affordance representation to guide grounded actions. We further propose an abstraction alignment reweighting (AAR) module that adaptively reweights the action imitation loss. A continuous intention--environment alignment signal is used to emphasize reliable action supervision when abstractions agree and reduce the influence of ambiguous demonstrations when they diverge, thereby learning a policy that balances fine-grained control and high-level reasoning without the need for manual rules. Extensive experiments show that 3A-VLA yields state-of-the-art results in both open-world (Minecraft) and competitive PvP (Game for Peace) settings. It also demonstrates strong zero-shot generalizability to high-fidelity games across different domains, including Valorant, CS2, GTA V, Elden Ring, and Mount & Blade. Code and models will be publicly released at https://3a-vla.github.io/.
Abstract: Vision-Language-Action (VLA) models inherit their visual backbone from static image-text pretraining, leaving physical dynamics to be learned from scarce action data. Generative video models, by contrast, already encode motion, contact, and implicit physics at internet scale. We introduce DiT4DiT, a Video-Action Model (VAM) that couples a Video Diffusion Transformer with an Action Diffusion Transformer and co-trains them under a dual flow-matching objective with a tri-timestep scheme for video supervision, action supervision, and feature extraction. The key insight is that the policy does not need to generate a future video: we intercept the Video DiT's hidden state at a single fixed flow timestep, turning a multi-step video rollout into a one-shot feature extraction consumed by the Action DiT. Across benchmarks, DiT4DiT sets new SOTA averages on LIBERO (98.6%) and the 24-task RoboCasa-GR1 suite (56.7%), outperforming strong VLA and video-based baselines. On a real Unitree G1, DiT4DiT reaches 73.8% on eight tabletop tasks and 72.2% on three whole-body loco-manipulation tasks, running up to 12× faster than prior VAM baselines while also exceeding them in accuracy. To our knowledge, this is the first video-generation-based policy to run humanoid whole-body control at real-time rates. It further generalizes zero-shot to unseen object categories and scene variations. Together, these results indicate that video-dynamics priors, accessed through a single forward pass, are a practical and scalable foundation for generalist robot policies.
Abstract: Reinforcement learning with verifiable rewards (RLVR) has been extended from single-domain training to multi-domain reasoning suites spanning mathematics, programming, and science. However, the training curriculum—how often each domain is sampled—is typically fixed or hand-tuned, even though reasoning skills transfer unevenly across domains. Existing learnability-based curricula adapt to where the policy is currently improving, but are blind to whether a gradient step on the selected domain benefits the remaining domains. In this paper, we propose Transfer-Aware Curriculum (TAC), a bandit-style online curriculum that prioritizes domains whose updates broadly benefit the rest of the training suite. TAC repurposes signals already produced by RL training: per-domain advantages capture local learnability, and projected gradients—taken from the GRPO step being computed—estimate cross-domain transferability via gradient-geometry alignment, at negligible cost (<1% wall-clock overhead). Across a six-domain reasoning suite, TAC consistently outperforms baselines such as proportional random sampling and a learnability-only bandit, improving macro-averaged accuracy by up to 3.7 points (13% relative) over the latter on both Qwen3-1.7B and Llama3.2-3B. Ablations show performance degrades sharply when the transferability term is removed, and TAC remains robust on imbalanced training mixtures where learnability-only curricula over-commit to dominant domains. Our findings establish cross-domain transferability as a key signal for multi-domain RL curriculum design.
Abstract: Recovering 4D human-object interaction (HOI) from monocular video is a key step toward scalable 3D content creation, embodied AI, and simulation-based learning. Recent methods can reconstruct temporally coherent human and object trajectories, but these trajectories often remain visual artifacts while failing to preserve stable contact, functional manipulation, or physical plausibility when used as reference motions for humanoid-object simulation. This reveals a fundamental interaction gap: HOI reconstruction should not stop at tracking a human and an object, but should recover the relation that makes their motion a coherent interaction. We introduce , a framework for reconstructing physically plausible 4D HOI animation from in-the-wild monocular videos. Instead of treating the human and object as independent entities in an ambiguous monocular 3D space, we propose a \emphhuman-first, object-follow formulation. The human motion is recovered as the interaction anchor, and the object is reconstructed, aligned, and refined relative to the human action. The resulting kinematic trajectory is then projected into a physics-based humanoid-object simulation, where it acts as a teacher trajectory for stable physical rollout. Across benchmark and in-the-wild videos, improves human-object alignment, contact consistency, temporal stability, and simulation readiness over prior monocular HOI reconstruction methods. By moving beyond visually plausible trajectory recovery toward physically grounded interaction animation, our work takes a step toward turning general monocular HOI videos into scalable demonstrations for humanoid-object behavior.
Abstract: Vision-language models (VLMs) have shown impressive capabilities across diverse tasks, motivating efforts to leverage these models to supervise robot learning. However, when used as evaluators in reinforcement learning (RL), today’s strongest models often fail under partial observability and distribution shift, enabling policies to exploit perceptual errors rather than solve the task. We introduce SOLE-R1 (Self-Observing LEarner), a video-language reasoning model explicitly designed to serve as the sole reward signal for online RL. Given only raw video observations and a natural-language goal, SOLE-R1 performs per-timestep spatiotemporal chain-of-thought (CoT) reasoning and produces dense estimates of task progress that can be used directly as rewards. To train SOLE-R1, we develop a large-scale video trajectory and reasoning synthesis pipeline that generates temporally grounded CoT traces aligned with continuous progress supervision. This data is combined with foundational spatial and multi-frame temporal reasoning, and used to train the model with a hybrid framework that couples supervised fine-tuning with RL from verifiable rewards. Across four different simulation environments and a real-robot setting, SOLE-R1 enables zero-shot online RL from random initialization: robots learn previously unseen manipulation tasks without ground-truth rewards, success indicators, demonstrations, or task-specific tuning. SOLE-R1 succeeds on 24 unseen tasks and substantially outperforms strong vision-language rewarders, including GPT-5 and Gemini-3-Pro, while exhibiting markedly greater robustness to reward hacking.
Abstract: Long-horizon reasoning in recent LLMs increasingly demands that the model switch between distinct skills inside a reasoning chain, such as a task that requires the model to first do a math derivation, then use the result to plan a schedule. We call such problems : multi-step tasks whose steps require different reasoning skills and depend on earlier outputs. Existing benchmarks often evaluate individual skills, lacking a principled way to measure , a benchmark of cross-skill long-horizon tasks built over 558 skills across 9 verifiable and open-ended domains. Each task is assigned a task-level skill-entropy score and grouped into three difficulty levels. Evaluating 8 frontier and 4 open-source models on Skill²-Bench surfaces a structural skill-switching gap: accuracy decreases monotonically with task-level skill entropy and drops by 5–30% when the same skill is exercised inside a cross-skill task rather than a single-skill question. We then turn skill entropy from a benchmark scale into a training signal. We propose , an RL framework where the model predicts not only the answer at each step but also the skill used to produce it. The reward combines step-level correctness with a that measures the alignment between the model-predicted skill sequence and the gold skill sequence. On Qwen3-4B-Instruct and Qwen3-1.7B, Skill-Entropy RL improves the Skill²-Bench score from 34.4% to 68.4% and from 14.6% to 40.1% respectively, outperforming SFT and vanilla GRPO. The same pipeline can be applied to off-the-shelf training data such as OpenR1-Math, indicating that skill entropy is a reusable training signal.
Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a dominant paradigm for improving reasoning in large language models (LLMs), yet the underlying geometry of the resulting parameter trajectories remains underexplored. In this work, we demonstrate that RLVR weight trajectories are extremely low-rank and highly predictable. Specifically, we find that the majority of downstream performance gains are captured by a rank-1 approximation of the parameter deltas, where the magnitude of this projection evolves near-linearly with training steps. Motivated by this, we propose a simple and compute-efficient method RELEX (REinforcement Learning EXtrapolation, which estimates the rank-1 subspace from a short observation window and extrapolates future checkpoints via linear regression, with no learned model required. Across three models (i.e., Qwen2.5-Math-1.5B, Qwen3-4B-Base, and Qwen3-8B-Base), RELEX produces checkpoints that match or exceed RLVR performance on both in-domain and out-of-domain benchmarks, requiring as few as 15% steps of full RLVR training. Remarkably, RELEX is able to extrapolate far beyond the observation window at no training cost, predicting checkpoints up to 10~50× the training steps with continued improvement (e.g., observe only the first 50 steps and extrapolate to 2000 steps). Our ablation analysis confirms the minimalist sufficiency of RELEX: neither increasing the subspace rank nor employing non-linear modeling yields further gains in extrapolation. Finally, we show that RELEX’s success stems from a "denoising" effect: by projecting updates onto the rank-1 subspace, the model discards stochastic optimization noise that would otherwise degrade performance during extrapolation.
Abstract: As LLM-based agents increasingly browse the web on users' behalf, a natural question arises: can websites passively identify which underlying model powers an agent? Doing so would represent a significant security risk, enabling targeted attacks tailored to known model vulnerabilities. Across 14 frontier LLMs and four web environments spanning information retrieval and shopping tasks, we show that an agent's actions and interaction timings, captured via a passive JavaScript tracker, are sufficient to identify the underlying model with up to 96% F1. We formalise this attack surface by demonstrating that classifiers trained on agent actions generalise across model sizes and families. We further show that strong classifiers can be trained from few interaction traces and that agent identity can be inferred early within an episode. Injecting randomised timing delays between actions substantially degrades classifier performance, but does not provide robust protection: a classifier retrained on delayed traces largely recovers performance. We release our harness and a labelled corpus of agent traces \hrefhttps://anonymous.4open.science/r/known_actions-B31Bhere.
Abstract: Reasoning models produce long traces of intermediate decisions and tool calls, making test-time verification increasingly important for ensuring correctness. Existing approaches either verify only the final answer, which misses early errors, or rely on branch-and-verify strategies that explore multiple trajectories at substantially higher compute cost. We introduce interwhen, a single-trajectory verification framework that steers model behavior by providing feedback on intermediate reasoning traces. It addresses two key challenges. First, given a set of verifiers, obtaining verifiable states from the reasoning trace typically requires prompt engineering or external task decomposition into fixed steps, which can constrain the model's reasoning strategy. Instead, we propose a monitoring system that periodically polls the reasoning trace and forks inference of the reasoning model to recover intermediate states. Verifiers are run asynchronously alongside generation, adding negligible overhead on correct executions and intervening only when violations occur. Second, beyond math and code, a central challenge for process verification is the scarcity of verifiers. interwhen addresses this through automatic verifier synthesis from natural-language policy documents. Given a policy, it can generate code-based verifiers, including provably correct verifiers in Lean and Z3. Together, these contributions yield a plug-and-play system for policy-grounded, formally verified process supervision of any reasoning agent. On reasoning benchmarks where policies encode mathematical or logical constraints, interwhen-based steering achieves near-perfect accuracy for reasoning models using a fraction of the token compute of test-time verification baselines. On agentic benchmarks with policy-based verifier generation, it enables significant improvements in task quality for SLMs without any finetuning; e.g., the task completion rate of Qwen3-30B jumps from 32% to 87% on the telecom domain in tau^2-bench.
Abstract: Towards more general and human-like intelligence, large language models should seamlessly integrate both multilingual and multimodal capabilities; however, extending an existing multimodal model to many languages typically requires expensive multilingual multimodal data construction and repeated end-to-end retraining. We study a training-free alternative: injecting multilingual capability into an existing multimodal model by composing residual updates in the shared language model backbone. The key challenge is that multilingual and multimodal updates are heterogeneous, reflecting different functional roles in the shared model. To address this, we propose Direction- and Magnitude-aware Multilingual Multimodal merging (DiM^3), which selectively composes the two updates at each parameter dimension while preserving the original vision encoder and multimodal projector. Experiments on multilingual benchmarks in both text-only and vision-language settings, covering 57 languages across LLaVA- and Qwen-based backbones, show that DiM^3 consistently outperforms existing merging baselines, substantially improves multilingual performance over the original multimodal model, and remains competitive with dedicated multilingual multimodal fine-tuning while largely retaining general multimodal ability. We further show that DiM^3 can be directly applied to already trained multilingual multimodal models and still yield additional gains. Further interpretability analysis shows that DiM^3 primarily reshapes intermediate-layer semantic representations, strengthening cross-lingual alignment under both text-only and multimodal inputs while preserving higher-layer task-sensitive structure. Our repository is on https://anonymous.4open.science/r/w-C677/.
Abstract: Recent advances in preference alignment for diffusion-based video generation, particularly via Direct Preference Optimization (DPO), have significantly improved visual quality. However, temporally sparse artifacts such as motion collapse, object flickering, and color oversaturation remain a major barrier to perceptual realism. Existing methods struggle with these issues due to two key limitations: (1) the preference attribution bottleneck, where offline human annotations are costly and fail to accurately capture learning dynamics, while online reward signals are rollout-aware but often unstable and biased; and (2) temporal credit misallocation, where uniformly applied supervision cannot effectively target the brief segments in which artifacts occur. To address these challenges, we propose concentrated Implicit Preference Optimization (cIPO), a post-training framework for video diffusion models. cIPO derives implicit preference signals directly from the denoising process: given a real video, the model adds forward noise and reconstructs it via iterative denoising, treating the original as the preferred sample and the reconstruction as the dispreferred one. This formulation captures inference-time errors without requiring human annotations or external reward models. Moreover, frame-level discrepancies between original and reconstructed videos reveal when failures occur. cIPO leverages this by computing temporal reconstruction errors and concentrating optimization on high-error segments, enabling more precise correction of failure-prone regions. Extensive experiments demonstrate that cIPO consistently enhances video authenticity and temporal coherence across multiple datasets, highlighting the effectiveness and efficiency of implicit implicit preference with temporally concentrated optimization.
Abstract: Recent literature on fine-tuning Large Language Models highlights a fundamental debate. While Full Fine-Tuning (FFT) provides the representational plasticity required for high-entropy knowledge injection, Low-Rank Adaptation (LoRA) can match or surpass FFT performance because many tasks only require updates in a low-rank space and benefit from LoRA's additional regularization. Through empirical evaluation across diverse tasks (SQL, Medical QA, and Counterfactual Knowledge) and varying language models (Gemma-3-1B, Qwen2.5-1.5B, and Qwen2.5-3B), we verify both trends and demonstrate that relying solely on either static architecture is structurally limited. To address this challenge, we propose a Mixture of LoRA and Full (MoLF) Fine-Tuning, a unified framework that enables continuous navigation between both training regimes. MoLF dynamically routes updates between FFT and LoRA at the optimizer level to ensure that exact gradient signals are available to both experts throughout training, yielding stable training dynamics. For memory-constrained environments, we also introduce MoLF-Efficient, which freezes base weights and only routes updates among a pair of LoRA experts of potentially varying rank. Our evaluations show that MoLF either improves on or stays within 1.5% of the better of FFT and LoRA across all settings, while MoLF-Efficient outperforms prior adaptive LoRA approaches by up to 20% on Fact and 9% on Med and SQL.
Abstract: Language exhibits inherent structures, a property that explains both language acquisition and language change. Given this characteristic, we expect language models to manifest their own internal structures as well. While interpretability research has investigated how models compute representations mechanistically through attention patterns and Sparse AutoEncoders, the organization of the resulting representations is overlooked. To address this gap, we introduce StructLens, a framework to analyze representations through a holistic structural view. StructLens constructs maximum spanning trees based on the semantic representations in residual streams, inspired by tree representation in dependency parsing, and provides summaries of token relationships in representation space. We analyze how contiguous tokens are also nearby in representation space and find that middle layers show the strongest local-span organization. Moreover, analysis of pre-training checkpoints reveals that smaller local units become detectable earlier in pre-training, and larger units later. Our findings demonstrate that StructLens provides insights into how models organize token representations across layers and training.
Abstract: We study Slowly Annealed Langevin Dynamics (SALD), a sampler for tracking a path of moving target distributions and approximating the terminal target through time slowdown. We establish non-asymptotic convergence guarantees via a KL differential inequality, showing that slowdown improves tracking through contraction of intermediate targets and the complexity of the path. Motivated by training-free guided generation with pretrained score-based generative models, we further introduce Velocity-Aware SALD (VA-SALD), which explicitly incorporates the underlying marginal distributions of the pretrained model and uses slowdown to correct the additional deviation induced by guidance. This yields a principled framework for training-free guided generation for diffusion-based and related generative model families, together with convergence guarantees that clarify the roles of intermediate functional inequalities and guidance bias.
Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a promising framework for enhancing the reasoning ability of large language models. However, much of the existing work is guided by heuristic intuition, leading to divergent algorithmic choices, even contradictory ones that nevertheless report empirical gains. To better understand this phenomenon, we conduct a theoretical analysis of RLVR updates. Our study reveals that differences in off-policy degree, determined by the number of gradient steps per rollout, substantially affect the distribution of importance sampling ratios and their clipping behavior, thereby altering which tokens dominate the update. Building on this insight, we characterize gradient expectation as the central quantity governing update dynamics and analyze the roles of token probability, advantage, and importance sampling ratio. Motivated by these findings, we propose Adaptive Clip Policy Optimization (ACPO), which adjusts clipping boundaries across token groups according to the empirical variance of their importance sampling ratios. Experiments on 3B and 7B models across diverse reasoning benchmarks, spanning mathematical problem solving, tabular QA, and logic puzzles, demonstrate that ACPO outperforms strong baselines such as DAPO and CISPO. These results demonstrate that principled, analysis-driven approaches yield more robust and effective RLVR methods. Code is available in: https://anonymous.4open.science/r/ACPO
Authors:
Shriram Damodaran, Soumyaratna Debnath, Cheston Tan, Lin WangAbstract: Omnidirectional or 360° cameras provide embodied AI agents with a holistic, wide field-of-view (FoV) perception of the 3D structure of their surroundings. This capability has generated significant interest in applying Multi-modal Large Language Models (MLLMs) to omnidirectional spatial reasoning. However, most MLLMs are primarily trained on conventional 2D perspective images and therefore struggle with the severe distortions and wrap-around discontinuities introduced by spherical geometry. As a result, enabling MLLMs to generalize effectively to non-Euclidean 3D spaces without retraining remains an open challenge. In this paper, we propose SphMind, a novel, training-free, and plug-and-play framework designed to bridge this gap. The central idea of SphMind is to decouple semantic perception from geometric reasoning. Instead of requiring MLLMs to learn complex spherical geometric principles internally, the framework preserves their strong semantic understanding while handling geometric reasoning externally. To achieve this, we introduce a Spherical Harmonics-based Spatial Graph (SHSG) that models spatial relationships using equivariant transformations on the sphere. We further integrate this with Inference-Time Geometric Grounding (IGG), a model-agnostic closed-loop optimization process that aligns the internal representations of MLLMs with spherical geometric constraints during inference. Extensive experiments on three benchmark datasets demonstrate the effectiveness of SphMind. Without any additional training, the framework achieves more than 21.4% average improvement in directional reasoning on MP3D and Stanford2D-3D, outperforms all prompt-engineering baselines by 8.7% on the real-world ODI-Bench dataset, and achieves 5.9× higher rotational invariance under panorama rotations compared to existing baselines, all without dataset-specific tuning. We also validate the framework on real-world in-the-wild captures, where SphMind successfully resolves directional reasoning queries that baseline vision-language models fail to answer correctly.
Authors:
Jiaqi Liu, Songning Lai, Pengze Li, Di Yu, Zhou wenjie, Yiyang Zhou, Peng Xia, Zijun Wang, Xi Chen, SHIXIANG TANG, LEI BAI, Wanli Ouyang, Mingyu Ding, Huaxiu Yao, Aoran WangAbstract: Automated discovery of governing equations remains challenging when models rely only on numerical tokens and ignore the phase-space and trajectory evidence that is informative in low-dimensional dynamics. We propose Visual Induction for Physics-based Equation Reasoning (VIPER-R1), a multimodal framework for visual formula discovery from phase portraits, trajectory plots, and aligned measurements. VIPER-R1 uses a supervised Motion Structure Induction (MSI) curriculum for joint rationale-and-equation generation, and is then refined by Reward-Guided Symbolic Calibration (RGSC), a GRPO-based stage optimized with a structure-aware reward that emphasizes topological correctness over coefficient matching. At inference time, we optionally apply Symbolic Residual Realignment (SR^2), an external symbolic regression module that models only the residual numerical mismatch of the base hypothesis. We release PhysSymbol, a 10,000-instance two-part multimodal corpus; the quantitative evaluation in this paper uses its Part 1 mechanics split, which provides paired visualizations, trajectory data, and expert C-CoT traces. On this benchmark, VIPER-R1-7B achieves the best structural score (S_\mathrmstruct=0.812), is effectively tied with VIPER-R1-3B on exact-match accuracy (S_\mathrmacc=0.487 vs.\ 0.488), and, after optional SR^2 refinement, reaches a Post-SR^2 MSE of 0.032, nearly 3× lower than the best VLM baseline. These results show that VIPER-R1 improves symbolic structure induction in focused low-dimensional dynamics by combining paired visual evidence with structure-aware calibration to produce stronger first-order hypotheses for downstream refinement.
Abstract: Feature attributions often hide a critical modeling choice: they explain a prediction along a counterfactual path from a reference state to an input. Different baselines, interpolations, and generative trajectories define different paths and can therefor produce different explanations. We study this path ambiguity as a modeling problem. Our central question is whether the path can be chosen by the data-generating transport process, rather than by a hand-designed interpolation or by the sensitivity geometry of the model being explained. We separate attribution into fixed-path credit allocation and path selection. For a fixed path, we prove that the Aumann-Shapley line integral is the unique attribution rule under standard fixed-path axioms and explicit coordinate-trace regularity. For path selection, we minimize kinetic action over flows that transport a reference distribution to the data distribution, yielding a transport-geodesic attribution principle. We approximate this ideal with Rectified Flow and Reflow and derive stability bounds linking vector-field error to attribution error. Experiments show that lower-action, transport-consistent paths produce more stable and structured explanations, preserving competitive deletion faithfulness, without claiming data-manifold membership. Our code is available at https://anonymous.4open.science/r/temp_OTshap-6119.
Abstract: Continuous image editing requires fine-grained control over semantic attributes in text-conditioned generative models. Recent methods obtain such control through trainable adapters, test-time optimization, or architecture-specific design choices. In this work, we revisit steering the text-encoder space as a simpler alternative for open-vocabulary slider construction. We show that, as text-conditioned generative models become increasingly capable, linear directions in their frozen text-encoder spaces can provide a competitive interface for continuous control when carefully constructed. To make this interface practical, we introduce an automated pipeline that constructs continuous sliders from just a prompt and a concept, eliminating manual data curation and coefficient tuning. First, a language model generates balanced contrastive prompt pairs. Then, difference-of-means over the frozen text encoder estimates a concept direction, and an LLM-assisted process selects the prompt tokens to steer. Finally, an elastic range search adaptively calibrates the slider coefficient interval. As the process includes only changing the text encoder hidden states, the same pipeline can be applied across different text-conditioned generative backbones without denoiser-specific modifications. Across diverse set of edit types, our method outperforms prior training-free baselines and approaches the edit compliance of trained controllers while avoiding per-concept and per-backbone training. Our results suggest that a carefully designed text-space steering is a surprisingly strong practical baseline for open-vocabulary continuous editing. Especially in a rapidly evolving generative model ecosystem, where repeatedly adapting specialized controllers to new backbones can be a substantial burden
Abstract: Can a shared-weight recurrent Transformer develop distinct internal roles without being partitioned into separate modules? We study this in Asymmetric Input Recurrence (AIR), a minimal two-state reasoning architecture in which the same Transformer model is reused for both updates (per literature, L and H) and the only built-in asymmetry is that the encoded input is injected during L-updates but not H-updates. Across Sudoku-Extreme and Maze, decoded rollouts reveal a consistent split: z_H behaves like a fully committed proposal state, whereas z_L retains local uncertainty and shifting intermediate structure. Freeze experiments show that this split is, in practice, related to the model's state dynamics: in Sudoku, freezing z_H reduces z_L's content changes whereas freezing z_L increases z_H's, while in Maze, freezing either state increases content changes in the other state. Ablations show that to induce specialization, the shared model need to be able to tell the two update types apart, either from input injection asymmetry or from a separate level token. Mechanistically, attention analysis shows that L-updates are consistently more local than H-updates in both Sudoku and Maze. Together, these results show that, in a two-state recurrent setting, a clear state-identity signal can induce stable, related functional roles inside a shared-parameter recurrent Transformer.
Abstract: Video large language models (VideoLLMs) show strong capability in video understanding, yet long-context inference is still dominated by massive redundant visual tokens in the prefill stage. We revisit token compression for VideoLLMs under a tight budget and identify a key bottleneck, namely insufficient spatio-temporal information coverage. Existing methods often introduce discontinuous coverage through coarse per-frame allocation or scene segmentation, and token merging can further misalign spatio-temporal coordinates under MRoPE-style discrete (t,h,w) bindings. To address these issues, we propose V-CAST (Video Curvature-Aware Spatio-Temporal Pruning), a training-free, plug-and-play pruning policy for long-context video inference. V-CAST treats frame representations as a semantic trajectory and uses local curvature as a low-cost proxy for temporal turns, routing token budgets to transition regions without relying on decoder attention. It further adopts a dual-anchor spatial selection mechanism that preserves high-entropy visual evidence without attention intervention, while keeping retained tokens at their original coordinates to maintain positional alignment. Extensive experiments across multiple VideoLLMs of different architectures and scales demonstrate that V-CAST achieves 98.6% of the original performance, outperforms the second-best method by +1.1% on average, and reduces peak memory and total latency to 86.7% and 86.4% of vanilla Qwen3-VL-8B-Instruct. Our code is provided in the supplementary material.
Abstract: The continued improvements in language model capability have unlocked their widespread use as drivers of autonomous agents, for example in coding or computer use applications. However, the core of these systems has not changed much since early instruction-tuned models like ChatGPT. Even advanced AI agents function on message exchange formats, successively exchanging messages with users, systems, with itself (i.e. chain-of-thought) and tools in a single stream of computation. This bottleneck to a as in chat models leads to a number of limitations: the agent cannot act (generate output) while reading, and in reverse, cannot react to new information while writing. Similarly, the agent cannot act while thinking and cannot think while reading or acting on information. In this work, we show that models can be unblocked by switching from instruction-tuning for sequential message formats to instruction-tuning for into a separate stream. Every forward pass of the language model then simultaneously reads from multiple input streams and generates tokens in multiple output streams, all of which causally depend on earlier timesteps. We argue that this data-driven change remedies a number of usability limitations as outlined above, improves model efficiency through parallelization, improves model security through better separation of concerns and improves model monitorability.
Abstract: Recent vision-language-action (VLA) models can generate plausible end-effector motions, yet they often fail in long-horizon, contact-rich tasks because the underlying hand-object interaction (HOI) structure is not explicitly represented. An embodiment-agnostic interaction representation that captures this structure would make manipulation behaviors easier to validate and transfer across robots. We propose FlowHOI, a two-stage flow-matching framework that generates semantically grounded, temporally coherent HOI sequences, comprising hand poses, object poses, and hand-object contact states, conditioned on an egocentric observation, a language instruction, and a 3D Gaussian splatting (3DGS) scene reconstruction. We decouple geometry-centric grasping from semantics-centric manipulation, conditioning the latter on compact 3D scene tokens and employing a motion-text alignment loss to semantically ground the generated interactions in both the physical scene layout and the language instruction. To address the scarcity of high-fidelity HOI supervision, we introduce a reconstruction pipeline that recovers aligned hand-object trajectories and meshes from large-scale egocentric videos, yielding an HOI prior for robust generation. Across the GRAB and HOT3D benchmarks, FlowHOI achieves the highest action recognition accuracy and a 1.7× higher physics-simulation success rate on GRAB than the strongest diffusion-based baseline, while delivering a 40× inference speedup. We further demonstrate real-robot execution across four contact-rich, dexterous manipulation tasks, highlighting the effectiveness of retargeting our generated HOI representations to robotic dexterous hands, achieving significantly higher success rates than prior methods due to improved physical plausibility and semantic alignment.
Abstract: Graphical User Interface (GUI) grounding maps natural language instructions to the visual coordinates of target elements and serves as a core capability for autonomous GUI agents. Recent reinforcement learning methods (e.g., GRPO) have achieved strong performance, but they rely on expensive multiple rollouts and suffer from sparse signals on hard samples. These limitations make on-policy self-distillation (OPSD), which provides dense token-level supervision from a single rollout, a promising alternative. However, its applicability to GUI grounding remains unexplored. In this paper, we present GUI-SD, the first OPSD framework tailored for GUI grounding. First, it constructs a visually enriched privileged context for the teacher using a target bounding box and a Gaussian soft mask, providing informative guidance without leaking exact coordinates. Second, it employs entropy-guided distillation, which adaptively weights tokens based on digit significance and teacher confidence, concentrating optimization on the most impactful and reliable positions. Extensive experiments on six representative GUI grounding benchmarks show that GUI-SD consistently outperforms GRPO-based methods and naive OPSD in both accuracy and training efficiency. Code and training data will be publicly released.
Abstract: Reward-based fine-tuning steers a pretrained diffusion or flow-based generative model toward higher-reward samples while remaining close to the pretrained model. Although existing methods are derived from different perspectives, we show that many can be written under a common framework, which we call reward score matching (RSM). Under this view, alignment becomes score matching against a value-guided target, and the main differences across methods reduce to the construction of the value-guidance estimator and the effective optimization strength across timesteps. This unification clarifies the bias--variance--compute tradeoffs of existing designs, and distinguishes core optimization components from auxiliary mechanisms that add complexity without clear benefit. Guided by this perspective, we develop simpler, more efficient redesigns across representative differentiable and black-box reward alignment tasks. Overall, RSM turns a seemingly fragmented collection of reward-based fine-tuning methods into a smaller, more interpretable, and more actionable design space.
Abstract: Understanding ultra-long videos such as egocentric recordings, live streams, or surveillance footage spanning days to weeks, remains a challenge. For current multimodal LLMs: even with million-token context windows, frame budgets cover only tens of minutes of densely sampled video, and most evidence is discarded before inference begins. Memory-augmented and agentic approaches help with scale, but their retrieval remains fragmented across modalities and lacks long-range narrative summaries that span days or weeks. We propose MAGIC-Video, a training-free framework built around a multimodal memory graph with interleaved narrative chain: the graph unifies episodic, semantic, and visual content through six typed edges and supports cross-modal retrieval, while the chain distils long-horizon entity biographies and recurring activity events. At inference time, an agentic loop interleaves graph retrieval with narrative fact injection, covering both the modality and time dimensions of ultra-long video in a single retrieval pipeline. On EgoLifeQA, Ego-R1 and MM-Lifelong, MAGIC-Video consistently outperforms strong general-purpose, long-video, and agentic baselines, with gains of 10.1, 7.4, and 5.9 points over the prior best agentic system on each benchmark.
Authors:
Zhaoxuan Li, Ziming Zhao, Siqi Lu, Rui Zhang, Wenhao Li, Xiaofei Yue, Fan T ZhangAbstract: Currently, frequent security incidents of Ethereum contracts have caused billions of dollars in losses. There is a pressing need to identify defective contract code and generate exploit call sequences to reproduce attacks and ensure detection accuracy. Nevertheless, the state-of-the-art (SOTA) detection methods based on symbolic execution and fuzzy testing cannot achieve the desired performance due to the state explosion problem caused by contract characteristics such as cross-contract calls and loop branches, especially in large wild contracts. To tackle this problem, we propose RSymX, an assembly strategy-guided symbolic execution for contracts that leverages reinforcement learning to dynamically balance code coverage and vulnerability discovery. Extensive experiments on open-sourced datasets demonstrate that RSymX achieves an overlap score of 92.39% and identifies many vulnerable wild instances that SOTAs misreport. Especially, its dynamic strategies improve the efficiency of state exploration by 2x~4x, and in some cases up to 10x, offering guidance for future symbolic-execution-based analyzers. The code and data are available in https://github.com/ContractAudit/Code.
Abstract: Multimodal large language models (MLLMs) have heterogeneous strengths across OCR, chart understanding, spatial reasoning, visual question answering, cost, and latency. As a result, effective MLLM routing requires more than estimating query difficulty: a router must match the multimodal requirements of the current image-question input with the capabilities of each candidate model. We propose \textscLatentRouter, a router that formulates MLLM routing as counterfactual multimodal utility prediction. Given an image-question query, \textscLatentRouter extracts learned multimodal routing capsules, represents each candidate MLLM with a model capability token, and performs latent communication between these states to estimate how each model would perform if selected. A distributional outcome head predicts model-specific counterfactual quality, and a bounded capsule correction refines close decisions while limiting the influence of noisy routing signals. The resulting utility-based policy supports both performance-oriented and performance-cost routing, and can handle changing candidate pools through shared per-model scoring with availability masking. Experiments on MMR-Bench and VL-RouterBench show that \textscLatentRouter outperforms strong fixed-model, feature-level, and learned-router baselines. Additional analyses show that the gains are strongest on multimodal task groups where model choice depends on visual, layout-sensitive, or reasoning-oriented requirements, and that latent communication is the main contributor to the improvement. The code is available at: https://anonymous.4open.science/r/LatentRouter-8718.
Abstract: Articulated 3D objects are essential for interactive environments in embodied AI, robotics, and virtual reality, but reconstructing their structure and motion from sparse observations remains challenging. Existing approaches remain largely constrained by lack of supervised data or lack priors needed to reliably recover articulation, hidden geometry, and internal object structure. We present the first hierarchical agentic approach for articulated 3D object reconstruction from text or image inputs, combining hierarchical, debate-based reasoning with a video generative prior for articulation modeling. High-level agents reason about object semantics and motion using knowledge from vision-language and video models, while low-level agents estimate articulation parameters and interaction points; together, they engage in structured debate to resolve ambiguities in structure and motion. To enable reliable generation, including interior object structures, from only text or image inputs, we introduce a video generative model prior that not only synthesizes plausible object motions, but also expose occluded interiors and geometry that cannot be inferred from a single static view. By combining agentic reasoning with a video generative prior, our approach jointly infers articulation and reconstructs complete 3D articulated objects, producing high-fidelity geometry, internal structure, and motion-consistent states beyond directly observed surfaces.
Abstract: Diverse outputs in text generation are necessary for effective exploration in complex reasoning tasks, such as code generation and mathematical problem solving. Such Pass@k problems benefit from distinct candidates covering the solution space. However, traditional sampling approaches often waste computational resources on repetitive failure modes. While Diffusion Language Models have emerged as a competitive alternative to the prevailing Autoregressive paradigm, they remain susceptible to this redundancy, with independent samples frequently collapsing into similar modes. To address this, we propose a training free, low cost intervention to enhance generative diversity in Diffusion Language Models. Our approach modifies intermediate samples in a batch sequentially, where each sample is repelled from the feature space of previous samples, actively penalising redundancy. Unlike prior methods that require retraining or beam search, our strategy incurs negligible computational overhead, while ensuring that each sample contributes a unique perspective to the batch. We evaluate our method on HumanEval and GSM8K using the LLaDA-8B-Instruct and Dream-v0-Instruct-7B models, where we demonstrate significantly improved diversity and Pass@k performance across various sampling settings. As a simple modification to the sampling process, our method offers an immediate, low cost improvement for current and future Diffusion Language Models in tasks that benefit from diverse solution search.
Abstract: Tool-augmented LLM agents tend to call tools indiscriminately, even when the model can answer directly. Each unnecessary call wastes API fees and latency, yet no existing benchmark systematically studies when a tool call is actually needed. We propose When2Tool, a benchmark of 18 environments (15 single-hop, 3 multi-hop) spanning three categories of tool necessity — computational scale, knowledge boundaries, and execution reliability — each with controlled difficulty levels that create a clear decision boundary between tool-necessary and tool-unnecessary tasks. We evaluate two families of training-free baselines: Prompt-only (varying the prompt to discourage unnecessary calls) and Reason-then-Act (requiring the model to reason about tool necessity before acting). Both provide limited control: Prompt-only suppresses necessary calls alongside unnecessary ones, and Reason-then-Act still incurs a disproportionate accuracy cost on hard tasks. To understand why these baselines fail, we probe the models' hidden states and find that tool necessity is linearly decodable from the pre-generation representation with AUROC 0.89–0.96 across six models, substantially exceeding the model's own verbalized reasoning. This reveals that models already know when tools are needed, but fail to act on this knowledge during generation. Building on this finding, we propose Probe&Prefill, which uses a lightweight linear probe to read the hidden-state signal and prefills the model's response with a steering sentence. Across all models tested, Probe&Prefill reduces tool calls by 48% with only 1.7% accuracy loss, while the best baseline at comparable accuracy only reduces 6% of tool calls, or achieves a similar tool call reduction but incurs a 5× higher accuracy loss. On the real-world Search-o1 agentic benchmark, Probe&Prefill reduces API calls by 20–56% without accuracy degradation.
Authors: ShaoFeng Qin, Li Wang
Abstract: Constructive neural routing solvers usually score the next action by matching a decoder context to candidate embeddings, hiding deterministic one-step consequences such as travel, waiting, slack, and capacity changes. We propose LINC (Local Inference via Normed Comparison), a decoder-side candidate decision architecture that computes these consequences explicitly. LINC uses them according to their decision role: centered relative consequences are compared by a shared linear local scorer, while feasible-set summaries modulate the decoder context. This preserves standard global matching and relieves the hidden state from rediscovering transition arithmetic. The Capacitated Vehicle Routing Problem with Time Windows (CVRPTW) serves as the main constrained-routing stress test; the same interface extends to the Capacitated Vehicle Routing Problem (CVRP) and Traveling Salesman Problem (TSP). In particular, for CVRPTW, LINC reduces PolyNet's Solomon/Homberger gaps from 13.83%/38.15% to 7.26%/14.71%; for TSP and CVRP, it also improves external-benchmark gaps.
Abstract: Reinforcement learning (RL) shows promise for enhancing LLM agentic reasoning capabilities with external environments. However, sparse terminal rewards hinder fine-grained, state-level optimization. Although process reward modeling offers a promising alternative, training dedicated reward models often entails substantial computational costs, risks reward hacking, and faces annotation bottlenecks. To address these challenges, we introduce leverages the intrinsic topological structure of states within trajectories by constructing state graphs. This enables topology-aware graph propagation to estimate state-wise contributions to success, yielding principled, annotation-free state-level rewards. As dense rewards for RL optimization, substantially outperforms prior RL baselines across four agentic benchmarks, with average success-rate gains of +6.2% on text-based benchmarks and +29.7% on visual reasoning over the strongest baseline across three model scales, and over +10% accuracy improvement on DeepResearch, while demonstrating superior robustness and training efficiency.
Authors:
Xin Xu, Clive Bai, Kai Yang, Tianhao Chen, Yuchen Cai, Yangkun Chen, Weijie Liu, Hao CHEN, Yang Wang, Saiyong Yang, Can YangAbstract: Large-scale verifiable prompts underpin the success of Reinforcement Learning with Verifiable Rewards (RLVR), but they contain many uninformative examples and are costly to expand further. Recent studies focus on better exploiting limited training data by prioritizing hard prompts whose rollout pass rate is 0. However, easy prompts with a pass rate of 1 also become increasingly prevalent as training progresses, thereby reducing the effective data size. To mitigate this, we propose Composition-RL, a simple yet useful approach for better utilizing limited verifiable prompts targeting pass-rate-1 prompts. More specifically, Composition-RL automatically composes multiple problems into a new verifiable question and uses these compositional prompts for RL training. Extensive experiments across model sizes from 4B to 30B show that Composition-RL consistently improves reasoning capability over RL trained on the original dataset. Performance can be further boosted with a curriculum variant of Composition-RL that gradually increases compositional depth over training. Additionally, Composition-RL enables more effective cross-domain RL by composing prompts drawn from different domains. All our codes, compositional datasets, and models will be released to facilitate future research.
Abstract: Predicting high-fidelity flow fields around ship hulls is central to hydrodynamic performance analysis, yet traditional CFD solvers remain computationally prohibitive for rapid design exploration across multiple vessel types and operating conditions. Existing deep-learning surrogates often interpolate unstructured CFD meshes onto regular grids, which can blur near-field resolution, while direct point-level global attention on large unstructured meshes can be computationally expensive. In this paper, we introduce APPSolver, a neural surrogate that learns next-step point-wise flow-field prediction directly on the unstructured free-surface ship CFD mesh without re-gridding. APPSolver is built around Adaptive Patch Partitioning (APP), a quadtree-based strategy that groups unstructured points into local spatial patches according to the non-uniform density inherent in ship CFD meshes: fine near the hull and wake, coarse in the far field. The resulting patch tokens are processed by a Transformer backbone together with optional condition token that encode ship geometry and condition parameters via a frozen large language model, forming a unified token-space model for one-step temporal advancement of velocity and pressure fields. On the ShipBench dataset, APP-Transformer achieves the lowest MAE and MSE. Ablation studies quantify the compression-fidelity trade-off of APP and demonstrate that condition token provide preliminary gains in some leave-one-hull-out settings. The code will be released upon acceptance.
Abstract: Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all --- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality. We develop a unified linear framework that addresses both questions. Under a spiked signal-plus-noise model with structured cross-modal nuisance correlation, we derive separation ratios for both objectives that expose complementary failure modes: alignment whitens each modality and fails when nuisance is strongly correlated across views; prediction encodes whatever is cross-predictable and fails when target-side noise is large. The resulting phase diagram partitions multimodal problems into four regimes --- Both, CA only, CP only, and Neither. We present a data-driven procedure to locate real-world datasets in this diagram using a small labeled subsample, identifying the preferred objective and prediction direction before any cross-modal training. Experiments on synthetic data, stereo-vision benchmarks, image--caption pairs, and real astrophysical data validate the predictions in the nonlinear regime, including the Neither regime where cross-modal training is actively harmful. Our framework lets practitioners diagnose their multimodal problem and choose the right objective before committing to training. Code to reproduce the results is available at \urlhttps://anonymous.4open.science/r/multimodal-ML-CFB0/README.md and will be open-sourced upon publication.
Abstract: A 3D scene is understood through its objects, not the primitives that compose them. Yet feed-forward reconstruction methods output dense, unstructured sets of points or Gaussians, leaving object-level structure to be recovered after the fact. We propose a feed-forward framework that directly decomposes unposed multi-view images into instance-structured 3D token groups - compact object-centric units from which reconstruction, segmentation, and manipulation all follow. Each token group consists of an instance token that captures entity-level identity, paired with anchor tokens that model local geometry and appearance and decode into 3D Gaussians. This factorization separates what belongs together from how each part looks, making object instances a native interface of the representation rather than a derived product. The model is trained through differentiable rendering with joint reconstruction and segmentation supervision, without 3D annotations or test-time optimization. On indoor scene benchmarks, our feed-forward model outperforms per-scene optimized baselines in class-agnostic instance segmentation, while the same token groups support instance-level scene editing through direct token removal, translation, and insertion.
Abstract: The explosive demand for interactive Large Language Model serving has highlighted the management of the Key-Value cache's dynamic memory footprint as a critical area for performance optimization in inference engines. Modern inference systems overwhelmingly rely on time-centric scheduling heuristics, such as Shortest Job First. However, their theoretical optimality is rooted in traditional schedule modeling, failing to capture the highly dynamic, 2D spatio-temporal geometric growth specific to LLM inference mechanisms. To resolve this, we propose the geometry-aware online scheduling by introducing the Smallest Volume First (SVF) algorithm and its highly efficient variant, 1-bit SVF. Theoretically, we provide a rigorous mathematical foundation for our approach. Utilizing a novel proof methodology, we tighten the worst-case competitive ratio (\textCR \le 48 \rightarrow \textCR \le 5) for SVF with known output lengths. Building upon this core breakthrough, we complete a comprehensive theoretical taxonomy analyzing our algorithms across different traffic scenarios and information availability. Practically, we seamlessly integrate our approach as a plug-and-play layer in vLLM. Extensive evaluations on Llama-3.1 models demonstrate comprehensive performance gains: SVF delivers strong reductions in both average and tail latency, while 1-bit SVF, with merely a single bit information, achieves competitive throughput and latency. This work establishes a theoretically sound and empirically proven approach for resolving memory-constrained scheduling in modern LLM deployments. To facilitate future research, our code is available at \urlhttps://anonymous.4open.science/r/LLM-inference-scheduling-7970.
Abstract: Spiking Neural Networks (SNNs) are well-regarded for their biological plausibility and energy efficiency in processing sequential data. However, dominant SNN architectures typically rely on Ordinary Differential Equations (ODEs) to govern neuronal state transitions. This first-order assumption imposes a ``memoryless'' bottleneck, limiting the model's capacity to capture the complex, long-range dependencies inherent in long-sequence tasks. In this work, we propose -SSM) from control theory into the spiking domain. By extending traditional integer-order SSMs to the fractional-calculus regime, LongSpike enables the hierarchical integration of neuronal dynamics with long-memory kernels. To mitigate the computational overhead and parallelization challenges typically associated with fractional operators, we leverage a state-space formulation that supports efficient, parallel training. Empirical evaluations on challenging benchmarks, including Long Range Arena (LRA), large-scale WikiText-103, and Speech Commands, demonstrate that LongSpike outperforms state-of-the-art SNNs in accuracy while preserving sparse synaptic computation.
Authors:
Yaolun Zhang, Tianyi Xu, Shengyu Dai, Zhenwen Shao, Qingyun Wu, Huazheng WangAbstract: We argue that multi-agent test-time evolution is not single-agent evolution replicated N times. A single-agent learner can only evolve its own context and memory. A multi-agent system additionally evolves who collaborates, how they collaborate, and how knowledge flows across the population. These components have no single-agent counterpart and can produce phenomena such as emergent specialization. Yet prior test-time methods either confine experiences to individual agents, forfeiting cross-agent learning, or broadcast symmetrically to all agents, erasing the specialization that makes collaboration valuable. We present EVOCHAMBER, a training-free framework that instantiates test-time evolution at three levels over a coevolving agent pool. At its core is CODREAM (Collaborative Dreaming), a post-failure protocol in which agents collaboratively reflect, distill insights, and route them asymmetrically from strong agents to weak ones on the failed niche, preserving specialization while filling knowledge gaps. Team-level operators assemble niche-conditioned teams and select collaboration structures online. Population-level lifecycle operators fork, merge, prune, and seed agents under performance pressure. On three heterogeneous task streams with Qwen3-8B, EvoChamber reaches 63.9% on competition math, 75.7% on code, and 87.1% on multi-domain reasoning, outperforming the best baseline by 30% relative on math. These gains transfer to GPT-4.1-mini. Ablations attribute the largest single drop of -10.8% to removing CODREAM, confirming asymmetric cross-agent transfer as the primary driver. Starting from 20 identically initialized agents, four to five stable niche specialists spontaneously emerge, a structural signature of multi-agent evolution that no single-agent learner can express.
Abstract: Multimodal Large Language Models (MLLMs) have demonstrated substantial promise in spatial understanding. Existing works typically incorporate prior knowledge extracted from a pre-trained foundation model to further enhance the spatial awareness of MLLMs. In this paper, we first reveal that when integrating diverse foundation models into MLLMs, each model provides complementary spatial priors that benefit different tasks. Motivated by this, we propose ViPS, a novel multi-model prior framework designed to fully unleash the potential of incorporating multiple Visual Priors from diverse models into MLLMs for Spatial understanding. Specifically, ViPS introduces an Efficient Prior Proxy to generate multiple foundational priors with minimal inference overhead, and a Dynamic Prior Fusion mechanism to achieve harmonious and context-aware prior fusion and injection from the prior proxies. Extensive experiments demonstrate that ViPS successfully harmonizes diverse visual priors, establishing new state-of-the-art performance across multiple complex spatial reasoning and 3D spatial understanding benchmarks.
Abstract: Memory facilitates question answering over long videos by extracting and retrieving facts to fit within the limited context windows of multimodal LLMs (MLLMs). Existing solutions typically extract independent memory entries from fixed-length video clips and thus cannot capture high-level semantics that need to be summarized over extended time periods, such as character traits and relations. Moreover, they rely solely on similarity-based retrieval and may fail to retrieve the fine-grained details required for question answering. To tackle these problems, we propose CAM, featuring continuous extraction for high-level semantics and adaptive querying for fine-grained details. In particular, CAM stores the entities and relations extracted from video clips in a knowledge graph. To capture the high-level semantics of each entity or relation, CAM summarizes the local subgraph of the target entity or relation once the subgraph reaches a predefined size. To retrieve the fine-grained details required for question answering, CAM supports multiple search methods, such as knowledge graph traversal, video re-watching, and audio listening, and utilizes a planner-executor-verifier pipeline to adaptively compose these search methods according to question intent. Evaluations on three benchmarks show that CAM outperforms SOTA baselines and improves their accuracy by up to 23 percentage points.
Abstract: Multimodal generative retrieval formulates multimodal retrieval as discrete identifier generation, eliminating the need for explicit similarity search over external embeddings. Existing approaches construct identifiers via residual quantization and decode them with trie-constrained beam search. This combination introduces an indexing–decoding gap: identifier learning objectives, including reconstruction and contrastive losses, do not explicitly enforce prefix discriminability during decoding. As a result, even well-optimized identifiers can be irreversibly pruned early in beam search due to low-rank prefixes. We theoretically characterize this gap and derive a survival bound that relates prefix retention to three controllable factors in indexing and decoding. Building on this bound, we propose PRO, prefix retention optimization, a unified framework comprising three mechanisms: (i) prefix ranking distillation aligns quantized prefix rankings with those induced by pre-quantization embeddings using a listwise loss; (ii) vocabulary scheduling increases codebook sizes from shallow to deep residual quantization levels to reduce early competition from non-target prefixes; and (iii) geometric score fusion vectorizes each candidate prefix and incorporates its similarity to the query into beam search scoring, further reducing the indexing–decoding mismatch. Experiments on nine multimodal retrieval tasks show that PRO improves retention of target identifier prefixes and outperforms existing multimodal generative retrieval baselines.
Abstract: Semi-supervised learning (SSL) has emerged as a practical solution for addressing data scarcity challenges by leveraging unlabeled data. Recently, vision-language models (VLMs), pre-trained on massive image-text pairs, have demonstrated remarkable zero-/few-shot performance that often surpasses SSL approaches due to their exceptional generalization capabilities. This gap motivates us to question: how can we effectively harness the powerful generalization capabilities of VLMs into task-specific models? Knowledge distillation (KD) offers a natural framework for transferring VLM capabilities, but we identify that it suffers from gradient conflicts between supervised and distillation losses. To address this challenge, we propose Dual-Head Optimization (DHO), which introduces dual prediction heads for each distinct signal. We observe that DHO resolves gradient conflicts, enabling improved feature learning compared to single-head KD baselines, with practical benefits of minimal computational overhead and test-time hyperparameter tuning without retraining. Extensive experiments across 15 datasets show that DHO consistently outperforms KD baselines, often outperforming teacher models with smaller student models. DHO also achieves new state-of-the-art performance on both in-distribution ImageNet semi-supervised learning and out-of-distribution generalization across ImageNet variants. We will publicly release our code and model checkpoints to facilitate future research.
Authors:
Mingtong Dai, Guanqi Peng, Yongjie Bai, Feng yan, chunjie chen, Lingbo Liu, Liang Lin, Xinyu WuAbstract: Previous imitation learning policies predict future actions at every control step, whether in smooth motion phases or precise, contact-rich operation phases. This uniform treatment is wasteful: most steps in a manipulation trajectory traverse free space and carry little task-relevant information, while a small fraction of \emphkey steps around contacts, grasps, and alignment demand dense, high-resolution prediction. We propose a novel \emphaction relabeling mechanism: at each timestep in a skip segment, we replace the behavior cloning target with the action at the entrance of the next key segment, enabling the policy to leap over redundant steps in a single decision. The resulting Skip Policy (SkiP) dynamically leaps over skip segments and intensively refines actions in key segments, within a single unified network requiring no learned skip planner or hierarchical structure. To automatically partition demonstrations into key and skip segments without manual annotation, we introduce \emphMotion Spectrum Keying (MSK), a fast, task-agnostic procedure that detects local motion complexity from action signals. Extensive experiments across 72 simulated manipulation tasks and three real-robot tasks show that SkiP reduces executed steps by 15--40% while matching or improving success rates across various policy backbones.
Abstract: Large Vision Language Models (LVLMs) show promise in medical applications, but their inability to faithfully ground responses in visual evidence raises serious concerns about clinical trustworthiness. While visual attribution methods are widely used to explain LVLM predictions, whether these explanations actually reflect the visual evidence underlying the model's decision is largely unverified, since ground-truth annotations for internal model reasoning are typically unavailable. We address this question for chest X-ray (CXR) reasoning by developing a causal evaluation framework that retains only CXR-VQA samples for which the expert-annotated region is verified, via counterfactual editing, to be causally responsible for the model's prediction. Using this framework across 11 attribution methods, six open-source LVLMs, and two output modes (direct answer and step-by-step reasoning), we find that existing attribution methods often fail to identify the evidence used by LVLMs. To address this failure, we propose MedFocus, a concept-based attribution method that localizes clinically meaningful anatomical regions via unbalanced optimal transport and measures their causal effect on model outputs through targeted interventions. MedFocus produces spatial, concept-level, and token-level attributions and substantially outperforms prior methods, taking a step toward more trustworthy attribution for medical LVLMs.
Abstract: Text image super-resolution (Text-SR) requires more than visually plausible detail synthesis: slight errors in stroke topology may alter character identity and break readability. Existing methods improve text fidelity with stronger recognitive or generative priors, yet they still face two unresolved challenges under severe degradation: the text condition extracted from low-quality inputs can itself be unreliable, and a plausible global prior does not fully determine fine-grained stroke boundaries. We present PRISM, a single-step diffusion-based Text-SR framework that addresses these two challenges through Flow-Matching Prior Rectification (FMPR) and a Structure-guided Uncertainty-aware Residual Encoder (SURE). FMPR constructs a privileged training-time prior from paired low-quality/high-quality latents and learns a flow matching that transports degraded embeddings toward this restoration-oriented prior space, yielding more accurate and reliable global text guidance. SURE further predicts uncertainty-aware structural residuals to selectively absorb reliable local boundary evidence while suppressing ambiguous stroke cues. Together, these components enable explicit global prior rectification and local structure refinement within a single diffusion restoration pass. Experiments on both synthetic and real-world benchmarks show that PRISM achieves state-of-the-art performance with millisecond-level inference.
Abstract: Many scientific datasets, such as molecular conformational ensembles or single-cell tissue measurements, are naturally modeled as meta-distributions: distributions over probability measures on non-Euclidean domains. Existing generative methods largely assume Euclidean geometry and fail to capture this structure. We introduce Riemannian Wasserstein Entropic Flow Matching (RWEFM), a generative framework on the Wasserstein space \mathcalP_2(\mathcalM) of a Riemannian manifold (\mathcalM,g). RWEFM is trained by regressing a neural vector field onto Riemannian optimal transport velocities, using McCann displacement interpolations as conditional paths. We confirm theoretically that this construction leads to a valid flow matching approach on \mathcalP_2(\mathcalM) and introduce the Riemannian Entropic Map, a GPU-efficient approximation of the optimal transport map on manifolds. Our experiments show that by respecting the intrinsic geometry of the data, RWEFM can generate whole single-cell samples in hyperspherical latent spaces and protein conformational ensembles on the torus. Code and tutorials are available at \hrefhttps://github.com/AnonRWEFM/RWEFMRWEFM.
Abstract: By far the most common way to estimate an expected loss in machine learning is to draw samples, compute the loss on each one, and take the empirical average. However, sampling is not necessarily optimal. Given an MLP at initialization, we show how to estimate its expected output over Gaussian inputs without running samples through the network at all. Instead, we produce approximate representations of the distributions of activations at each layer, leveraging tools such as cumulants and Hermite expansions. We show both theoretically and empirically that for sufficiently wide networks, our estimator achieves a target mean squared error using substantially fewer FLOPs than Monte Carlo sampling. We find moreover that our methods perform particularly well at estimating the probabilities of rare events, and additionally demonstrate how they can be used for model training. Together, these findings suggest a path to producing models with a greatly reduced probability of catastrophic tail risks.
Abstract: Evaluating new large language models typically requires costly human annotation campaigns at scale. LLM-as-a-judge offers a cheaper alternative, but judge scores carry systematic errors — such as position bias, self-preference, or intransitivity — that can strongly miscalibrate the resulting rankings. We quantify the resulting judge--human disagreement at two complementary levels. At the local level, we estimate per-battle uncertainty from the judge's own score differences by propagating calibrated win probabilities rather than hard labels into the Bradley--Terry procedure. This alone provides a drastic improvement to Elo estimation accuracy, bringing LLM-derived ratings within 17.9 Elo MAE of human-derived ones when averaged over 55 held-out models on LMArena. At the global level, we apply split conformal prediction to the residual gap between LLM-derived and human-derived Elo ratings across held-out models, producing prediction intervals with distribution-free marginal coverage guarantees that account for irreducible LLM--human disagreement. Together, these two layers yield a low-cost evaluation tool that provides developers with calibrated Elo estimates and honest uncertainty bounds, without access to large-scale human annotations.
Abstract: A major recent advance in quantization is given by microscaled 4-bit formats such as NVFP4 and MXFP4, quantizing values into small groups sharing a scale, assuming a fixed floating-point grid. In this paper, we study the following natural extension: assume that, for each group of values, we are free to select the "better" among two or more 4-bit grids marked by one or more bits in the scale value. We formalize the power-of-two-grids (PO2) problem, and provide theoretical results showing that practical small-group formats such as MXFP or NVFP can benefit significantly from PO2 grids, while the advantage vanishes for very large groups. On the practical side, we instantiate several grid families, including 1) PO2(NF4), which pairs the standard NF4 normal grid with a learned grid, 2) MPO2, a grid pair that is fully learned over real weights and activations, and 3) SFP4, a TensorCore-implementable triple which pairs NVFP4 with two shifted variants. Results for post-training quantization of standard open models and pre-training of Llama-like models show that adaptive grids consistently improve accuracy vs single-grid FP4 under both weight-only and weight+activation quantization.
Abstract: We introduce FlexTab, a flexible encoder-decoder architecture for in-context learning on tabular data that pairs a single, task-agnostic encoder with a suite of task-specific decoders. Unlike existing tabular in-context learners, which entangle feature representations with a specific prediction target, our design produces target-agnostic row embeddings that can be leveraged across a wide range of downstream tasks within a table-native in-context learning setup. We demonstrate this flexibility on six distinct problems: classification, regression, anomaly detection, clustering, entity matching, and entity classification in relational databases. Both the encoder and the task-specific decoders are trained on a large corpus of real-world, unlabeled tables. FlexTab achieves state-of-the-art performance on classification, regression, anomaly detection and entity matching, while remaining competitive with specialized models on entity classification in a relational setting. These results demonstrate that a single shared encoder, paired with task-specific decoders, can serve as an effective general-purpose backbone for diverse tabular prediction problems. The inference code and checkpoints will be made publicly available.
Abstract: Tiny UAV detection from an onboard event camera is difficult when the observer and target move at the same time. In this motion-on-motion regime, ego-motion activates background edges across buildings, vegetation, and horizon structures, while the UAV may appear as a sparse event cluster. To explore this practical problem, we present M^2E-UAV, a benchmark and analysis setup for onboard motion-on-motion event-based tiny UAV detection. The processed M^2E-UAV benchmark contains 87,223 training samples and 21,395 validation samples across four scene families: sunny building-forest, sunny farm-village, sunset building-forest, and sunset farm-village. We provide M^2E-Point, a point-based event baseline, and M^2E-Point + IMU, an IMU-conditioned variant, to analyze the role of inertial cues under onboard motion-on-motion detection. M^2E-Point encodes events as [x,y,t,p] point sets, extracts local event structure with EdgeConv, and predicts event-level UAV foreground scores, from which bounding boxes are derived via DBSCAN. Our validation-stage analysis shows that point-based event modeling is a strong baseline, while simple IMU conditioning provides only marginal aggregate gains. Under the train/validation split, M^2E-Point achieves 0.9673 F1 and 0.5501 mAP50-95, while the IMU-conditioned variant reaches 0.5561 mAP50-95 with only marginal aggregate changes, serving as an initial baseline for future exploration in this domain. The dataset will be publicly released upon acceptance.
Abstract: Retrieving the 3D kinematics of articulated objects from monocular video is a fundamental challenge in computer vision. Existing methods rely on complex video setups or cues such as long-term point tracking or wide-baseline matching, but are frequently brittle under severe occlusions, rapid camera ego-motion, or weak local features. Learning-based methods, meanwhile, struggle to generalize beyond their training categories. We propose a category-agnostic optimization framework that treats articulated object understanding as a primitive-fitting problem. Geometric primitives serve as a proxy representation that avoids the pitfalls of unstable point tracks; a novel mechanism organizes them into coherent parts constrained by revolute and prismatic joints. Our formulation jointly optimizes part segmentation and joint parameters, recovering complex kinematics from a single casually captured video. A visibility-aware procedure handles partial observations and occlusions inherent to real-world data. We also propose the AiP-synth and AiP-real benchmarks, featuring significant camera motion and heavy occlusions, and outperform existing methods.
Abstract: Reinforcement learning (RL) has been effective for post-training autoregressive (AR) language models, but extending these methods to diffusion language models (DLMs) is challenging due to intractable sequence-level likelihoods. Existing approaches therefore rely on surrogate likelihoods or heuristic approximations, which can introduce bias and obscure the sequential structure of denoising. We formulate diffusion-based sequence generation as a finite-horizon Markov decision process over the denoising trajectory and derive the policy gradient that decomposes over denoising steps in terms of stepwise advantages, without requiring explicit evaluation of the sequence likelihood. Grounded in this theorem, we develop tractable approximations for large-scale training: (i) denoising steps are selected for policy updates via an entropy-guided approximation bound, and (ii) stepwise advantages are estimated using a one-step denoising completion naturally provided by the diffusion model, avoiding costly multi-step rollouts or auxiliary value networks. Experiments demonstrate state-of-the-art results on nearly all benchmarks spanning coding, logical reasoning, and mathematical reasoning, outperforming existing RL post-training approaches for DLMs.
Authors: Hugo Cazaux, Eyjolfur Asgeirsson, Hlynur Stefansson
Abstract: Synthetic data has transformed language model training, yet its role in time series forecasting remains poorly understood. We present a large-scale empirical study: nine experiment groups, 4,218 runs systematically evaluating synthetic time series augmentation across five architectures, four synthetic signals and seven datasets. The effect is sharply architecture-conditional: channel-mixing models (TimesNet, iTransformer) benefit in the majority of trials, while channel-independent models (DLinear, PatchTST) are consistently degraded. In selected low-resource settings the gains are striking: TimesNet trained on only 10% of Weather data with synthetic augmentation surpasses the full-data baseline (4 of 16 sparsity-dataset combinations). Averaged across all architectures, augmentation hurts in 67% of trials. We further find that only the Seasonal-Trend generator reliably helps across the tested benchmarks, and that hard curriculum switching is actively harmful (+24% MSE degradation). These results provide concrete, actionable guidelines on how to use synthetic data: use synthetic augmentation with channel-mixing architectures, use gradual annealing schedules, and treat low-resource augmentation as architecture- and dataset-dependent. Code is available at an \hrefhttps://github.com/anonmirror/synthetic-timeseriesanonymized repository.
Authors:
Kunyu Peng, Zhikun Zhou, Kailun Yang, Di Wen, Ruiping Liu, Yufan Chen, Junwei Zheng, Hao Shi, Yi Zhou, M. Saquib Sarfraz, Danda Pani Paudel, Luc V GoolAbstract: Multimodal Large Language Models (MLLMs) have made substantial progress in egocentric video understanding, but their ability to reason cooperatively from multiple embodied viewpoints remains largely unexplored. We study this problem through multi-robot cooperative dynamic spatial reasoning, where a model must answer spatial, temporal, visibility, and coordination questions by integrating synchronized egocentric videos from a team of moving robots. To support this setting, we introduce CoopSR, the first benchmark for this task, together with EgoTeam, a multi-robot egocentric QA dataset. EgoTeam contains 114,227 QA pairs spanning 19 question types, four difficulty tiers, and three team sizes in Habitat and iGibson, along with a real-world test set of around 2,326 QAs collected using two quadruped robots. We further propose SP-CoR (Spectral and Physics-Informed Cooperative Reasoner), an MLLM framework for fine-grained cooperative spatial reasoning. SP-CoR combines dynamics-aware multi-robot frame sampling, spectral- and physics-guided view fusion, and physics-aligned prompt distillation, enabling the model to benefit from privileged robot-pose supervision during training while requiring only egocentric videos at test time. Across 22 MLLM baselines, SP-CoR consistently improves cooperative reasoning, outperforming the strongest fine-tuned baseline by +3.87% on Habitat and +7.12% on iGibson. It also shows stronger generalization to unseen team sizes and real-world robot tests. We will release the benchmark and code.
Abstract: Visual reasoning, often interleaved with intermediate visual states, has emerged as a promising direction in the field. A straightforward approach is to directly generate images via unified models during reasoning, but this is computationally expensive and architecturally non-trivial. Recent alternatives include agentic reasoning through code or tool calls, and latent reasoning with learnable hidden embeddings. However, agentic methods incur context-switching latency from external execution, while latent methods lack task generalization and are difficult to train with autoregressive parallelization. To combine their strengths while mitigating their limitations, we propose ATLAS, a framework in which a single discrete "word", termed as a functional token, serves both as an agentic operation and a latent visual reasoning unit. Each functional token is associated with an internalized visual operation, yet remains a standard token in the tokenizer vocabulary and can be generated via next-token prediction. This design avoids verbose intermediate visual content generation, while preserving compatibility with standard SFT and RL training, without architectural or methodological modifications. To further address the sparsity of functional tokens during RL, we introduce Latent-Anchored GRPO (LA-GRPO), which stabilizes the training by anchoring functional tokens with a statically weighted auxiliary objective, providing stronger gradient updates. Extensive experiments and analyses demonstrate that ATLAS achieves superior performance on challenging benchmarks while maintaining clear interpretability. We hope ATLAS offers a new paradigm inspiring future visual reasoning research.
Abstract: Transformer-based models have advanced feedforward novel view synthesis (NVS). Current architectures such as GS-LRM and LVSM mix semantic information (e.g., RGB) and spatial information (e.g., Pl\"ucker rays) into a shared feature space. Since Pl\"ucker rays naturally carry lattice-like spatial structure, these designs can make the spatial bias interfere with appearance representation and degrade rendering fidelity. To this end, we propose to decouple the representation of feedforward NVS transformers into separate semantic and spatial tokens. The decoupled design keeps semantic and spatial information explicit in their branches while preserving cross-branch interaction through shared attention routing. Built on this design, we introduce optional categorized supervision and bidirectional modulation: the former provides branch-specific training signals, while the latter improves interaction between the two branches. Notably, the base decoupled design introduces virtually zero additional inference latency due to its architectural design. The proposed designs achieve consistent improvements, demonstrating effectiveness across decoder-only and encoder-decoder feedforward NVS models.
Abstract: Recovering world-space 4D motion of two interacting hands from egocentric video is a fundamental capability for supervising robot policy learning, where wrist trajectories track the end-effector and finger articulations specify the grasp pose. Two major challenges arise in this setting: hands frequently leave the camera view for extended periods due to head motion, and persistent hand–object interactions cause severe occlusions of one or both hands. Existing methods uniformly condition on noisy hand motion observations without accounting for their per-frame reliability, leading to substantial performance degradation. Our key insight is that accurate world-space hand motion esimation is tightly coupled with the quality of per-frame hand observations. To this end, we decompose the quality of hand motion observations extracted from an off-the-shelf hand pose estimator into four channels: wrist global translation and finger articulations for both hand. We propose StableHand, a quality-aware flow-matching framework conditioned on these four-channel quality signals, which are predicted by a learned quality network. We naturally incorporate the quality signals into the flow-matching process through a per-channel forward schedule, a quality-adjusted velocity target, AdaLN modulation of the DiT denoiser, and a quality-aware ODE initialization. This unified generative process preserves high-quality observations while reconstructing unreliable ones using a learned bimanual motion prior. Experiments on HOT3D and ARCTIC, two egocentric benchmarks featuring long missing-hand spans and persistent hand–object occlusions, show that StableHand achieves state-of-the-art performance across all reported metrics, reducing W-MPJPE by 20-25% compared to the strongest baseline, with the largest gains on heavily occluded ARCTIC sequences. Code and pretrained models will be released upon acceptance.
Abstract: Diffusion Transformers (DiTs) and related flow-based architectures are now among the strongest text-to-image generators, yet the internal mechanisms through which prompts shape image semantics remain poorly understood. In this work, we study massive activations: a small subset of hidden-state channels whose responses are consistently much larger than the rest. We show that, despite their sparsity, these few channels effectively draw the whole picture, in three complementary senses. First, they are functionally critical: a controlled disruption probe that zeroes the massive channels causes a sharp collapse in generation quality, while disrupting an equally-sized set of low-statistic channels has marginal effect. Second, they are spatially organized: restricting image-stream tokens to massive channels and clustering them yields coherent partitions that closely align with the main subject and salient regions, exposing a structured spatial code hidden inside an apparently outlier-like subspace. Third, they are transferable: transporting massive activations from one prompt-conditioned trajectory into another, shifts the final image toward the source prompt while preserving substantial content from the target, producing localized semantic interpolation rather than unstructured pixel blending. We exploit this property in two use cases: text-conditioned and image-conditioned semantic transport, where massive activations transport enables prompt interpolation and subject-driven generation without any additional training. Together, these results recast massive activations not as activation anomalies, but as a sparse prompt-conditioned carrier subspace that organizes and controls semantic information in modern DiT models. All the code will be publicly available.
Authors:
bochao liu, Zhipeng Qian, yang zhao, Xinyuan Jiang, Zihan Liang, Yufei Ma, Junpeng Zhuang, Ben Chen, Shuo Yang, Hongen Wan, Yao Wu, Chenyi LeiAbstract: Operating and maintaining (O&M) large-scale online engine systems (search, recommendation, advertising) demands substantial human effort for release monitoring, alert response, and root cause analysis. While LLM-based agents are a natural fit for these tasks, the deployment bottleneck is not reasoning capability but (handbook rules and practitioner experience). Feeding all signals indiscriminately causes dilution and hallucination, while manually curating the event-to-(data, knowledge) mapping is intractable under dozens of daily releases. We present abstracting day-to-day O&M into three canonical patterns: release interception, proactive inspection, and alert root cause analysis; (ii) , where each Skill specifies which data and knowledge to retrieve for a given business-module context and can be automatically generated and updated by LLMs or iteratively refined through natural-language instructions from on-call engineers; (iii) a in which one correction signal drives two parallel pathways, case-memory-to-knowledge distillation and targeted Skill refinement. Deployed on the e-commerce search engine of a major short-video platform in China, reduces alert volume by 75%, achieves 80% root-cause analysis accuracy, and cuts mean time to resolution by over 50%. Our framework achieves 99.0% pass rate on offline evaluations.
Abstract: Decades of cognitive science establish that humans navigate environments by forming cognitive maps, defined as allocentric and topology-preserving representations of 3D space. While modern Vision-Language Models (VLMs) demonstrate emergent spatial reasoning from 2D egocentric inputs, it remains unclear whether they construct an analogous 3D internal representation. In this paper, we demonstrate that current VLMs do possess a latent topological map of 3D scenes, but it is heavily overshadowed by non-geometric visual semantics, such as color and shape. By isolating this spatial subspace through cross-scene linear feature extraction, we extract a clean spatial subspace that causally controls the model's spatial outputs. We mathematically shape this latent representation and prove its correspondence to the Laplacian eigenmaps of the scene's 3D Gaussian-kernel graph, converging to the physical 3D space in the continuous limit. Motivated by this geometric identification, we further introduce a mathematically principled latent regularization method for VLMs, based on Dirichlet energy. Applying this single-term regularizer to a minimal 500-step supervised VLM fine-tuning (SFT) on simple synthetic data yields significant improvements on real-world spatial benchmarks, outperforming standard SFT and competitive baselines by up to 12.1% in spatial tasks involving scene topology understanding.
Authors:
Boyu Chen, Yi Chen, Lu QIU, Jerry Bai, Yuying Ge, Yixiao GeAbstract: Scaling humanoid foundation models is bottlenecked by the scarcity of robotic data. While massive egocentric human data offers a scalable alternative, bridging the cross-embodiment chasm remains a fundamental challenge due to kinematic mismatches. We introduce okenizer via Visual Anchoring), a framework that learns a unified physical language for human-to-humanoid manipulation transfer. Grounded in the philosophy that heterogeneous kinematics share consistent visual consequences, UniT employs a tri-branch cross-reconstruction mechanism: actions predict vision to anchor kinematics to physical outcomes, while vision reconstructs actions to filter out irrelevant visual confounders. Concurrently, a fusion branch integrates these purified modalities into a shared discrete latent space of cross-embodiment physical intents. We validate UniT across two paradigms: (1) By predicting these unified tokens, VLA-UniT achieves state-of-the-art performance with high data efficiency on the RoboCasa GR1 benchmark. Leveraging diverse human data further improves out-of-distribution (OOD) generalization in simulation and real-world deployment, and enables By aligning cross-embodiment dynamics via unified tokens as conditions, it supports direct human-to-humanoid action-conditioned generation, translating human knowledge into enhanced action controllability for humanoid video generation. Ultimately, by inducing a more aligned cross-embodiment representation (empirically supported by t-SNE visualizations revealing improved alignment of human and humanoid features), UniT offers a scalable path to distill human priors into humanoid manipulation capabilities.
Authors:
Fengyuan Yang, Luying Huang, Jiazhi Guan, Quanwei Yang, Dongwei Pan, Jianglin Fu, Haocheng Feng, Wei He, Kaisiyuan Wang, Hang Zhou, Angela YaoAbstract: Recent advances in Video Foundation Models (VFMs) have revolutionized human-centric video synthesis, yet fine-grained and independent control of subjects and scenes remains a critical challenge. Recent attempts to achieve joint human-environment control often rely on explicit human-scene 3D alignment, improving geometric control but sacrificing generative flexibility and long-horizon consistency while requiring heavy 3D pre-processing. We present ONE-SHOT, a parameter-efficient framework for compositional human-environment video generation. Our key insight is to factorize the generative process into disentangled signals, separating human dynamics from scene cues. Specifically, our canonical-space motion injection and Dynamic-Grounded-RoPE map canonical human motion to the target video region, enabling accurate motion and placement control without heuristic human-scene 3D alignment. Hybrid Context Integration further maintains subject and scene consistency for minute-level synthesis. Experiments show that ONE-SHOT consistently outperforms state-of-the-art methods in structural control, flexible composition, text-guided semantic control, and long-horizon consistency.
Authors:
Jing Jin, Hao Liu, Yan Bai, Yihang Lou, Zhenke Wang, Tianrun Yuan, Yongkang Zhu, Juntong Chen, Fanhu Zeng, Xuanyu Zhu, Tao Feng, Yige XuAbstract: Multimodal large language models (MLLMs) have shown promising reasoning abilities, yet evaluating their performance in specialized domains remains challenging. STEM reasoning is a particularly valuable testbed because it provides highly verifiable feedback, but existing benchmarks often permit unimodal shortcuts due to modality redundancy and focus mainly on final-answer accuracy, overlooking the reasoning process itself. To address this challenge, we introduce STEPSTEM: a graduate-level benchmark of 283 problems across mathematics, physics, chemistry, biology, and engineering for fine-grained evaluation of cross-modal reasoning in MLLMs. STEPSTEM is constructed through a rigorous curation pipeline that enforces strict complementarity between textual and visual inputs. We further propose a general step-level evaluation framework for both text-only chain-of-thought and interleaved image-text reasoning, using dynamic programming to align predicted reasoning steps with multiple reference solutions. Experiments across a wide range of models show that current MLLMs still rely heavily on textual reasoning, with even Gemini 3.1 Pro and Claude Opus 4.6 achieving only 38.29% accuracy. These results highlight substantial headroom for genuine cross-modal STEM reasoning and position STEPSTEM as a benchmark for fine-grained evaluation of multimodal reasoning.
Abstract: Flicker-banding (FB), arises from temporal aliasing between a camera's rolling shutter and a display's brightness modulation, degrading screen-captured image readability with color shifts and jagged patterns. Existing single-frame methods with simplified parametric stripe models cannot reliably distinguish these artifacts from genuine texture. To address this, we conduct a systematic analysis of complex FB morphologies and reveal their significant variation across exposure settings, motivating a multi-frame bracketed RAW restoration paradigm. We construct Bricker, a synthetic–real bracketed RAW dataset built via ray-tracing-based physical simulation and automated multi-exposure capture tool. We further propose BRACE: Bracketed RAW Flicker-Banding Removal, a multi-frame restoration model that utilizes frequency-aware banding prior and a multi-scale spatial cross-attention modulator (MSCAM) for cross-exposure spatial fusion. We also introduce the Stripe Frequency Consistency (SFC) metric to evaluate banding removal. Extensive experiments demonstrate state-of-the-art performance on both synthetic and real benchmarks. Our dataset and code will be publicly released.
Abstract: Visual geometry transformers have become powerful architectures for multi-view 3D reconstruction, enabling joint prediction of multiple 3D attributes in a feed-forward manner. However, their computational cost grows quadratically with the input sequence length due to the global attention layers inside these models. This limits both their scalability and efficiency. In this work, we address this challenge with a simple yet general strategy: restricting the number of key/value tokens that each query interacts with during global attention. To achieve effective token selection, we introduce a two-stage framework. First, an inter-frame selection step operates at the frame level to identify frames that should be preserved. Second, an intra-frame selection step further discards more redundant tokens within the selected frames. Our analysis highlights the advantage of a diversity-based strategy for inter-frame selection, which ensures broad coverage of the scene. For intra-frame selection, we show that layer-aware sparsification is necessary, with the selection process guided by the entropy of the global attention pattern. Our approach offers a superior speed-accuracy trade-off compared to existing solutions. Extensive experiments show that it accelerates visual geometry transformers by over 85% for scenes with 500 images while maintaining, or even improving, baseline performance, which hints that how our token selection strategy can play a crucial role in future applications of visual geometry transformers.
Abstract: Although multimodal large language models (MLLMs) have advanced industrial anomaly detection toward a zero-shot paradigm, they still tend to produce high-confidence yet unreliable decisions in fine-grained and structurally complex industrial scenarios, and lack effective self-corrective mechanisms. To address this issue, we propose M3-AD, a unified reflection-aware multimodal framework for industrial anomaly detection. M3-AD comprises two complementary data resources: M3-AD-FT, designed for reflection-aligned fine-tuning, and M3-AD-Bench, designed for systematic cross-category evaluation, together providing a foundation for reflection-aware learning and reliability assessment. Building upon this foundation, we propose RA-Monitor, which models reflection as a learnable decision revision process and guides models to perform controlled self-correction when initial judgments are unreliable, thereby improving decision robustness. Extensive experiments conducted on M3-AD-Bench demonstrate that RA-Monitor outperforms multiple open-source and commercial MLLMs in zero-shot anomaly detection and anomaly analysis tasks.
Authors:
Seungmin Kim, Mingun Kim, Yuna Lee, Yulhwa KimAbstract: Analog circuit design remains highly dependent on expert knowledge due to the complexity of device-level interactions and topology design. Recent transformer-based approaches for device-level topology generation have shown promise, yet they suffer from low electrical validity without human-in-the-loop (HITL) training and severe memorization caused by sequence-based circuit representations. In this work, we propose to enforce structural validity during bipartite graph generation. Experimental results demonstrate that AnalogToBi achieves high validity and novelty without HITL training while effectively avoiding memorization of training topologies.
Abstract: The rapid evolution of Time Series Foundation Models (TSFMs) has advanced zero-shot forecasting across diverse domains. Inspired by the current form of Large Language Models, future TSFMs may be offered as commercialized, closed-source API services. However, many existing online adaptation methods still rely on white-box access for parameter fine-tuning or gradient backpropagation. This paradigm mismatch raises a question: We answer this with an insight: the predictive errors of the base model are conditioned on both the input and output of the base model (i.e., the context of errors). To validate this insight, we propose (Online Residual Contextual Adaptation). We conduct extensive experiments across 5 state-of-the-art TSFMs and 8 datasets to demonstrate the effectiveness of our approach. Furthermore, through ablation studies, we quantitatively analyze the impact of different adapter learning hypotheses on the final adaptation performance in black-box online adaptation. Code available at https://anonymous.4open.science/r/ORCA1/.
Abstract: LLM agents exhibit a consistent tendency to over-call, invoking tools even in situations where none is needed. On the When2Call benchmark, six models from three families show high call accuracy but much lower no-call accuracy, leaving overall accuracy in the 55%--70% range. We trace this to an Intrinsic Bias Hypothesis ( ): the call/no-call decision mapping carries an activation-independent CALL offset, so the model favors CALL even at activation parity. Using Sparse Autoencoders (SAEs), we recover behavior-aligned feature bases for the CALL/NO CALL decision, reduce them to a signed activation margin, and estimate the offset directly. Across all six models, the model is decision-neutral only when NO CALL activation outweighs CALL activation, consistent with IBH. We then causally test IBH with Adaptive Margin-Calibrated Steering ( ), a closed-form counter-bias shift along SAE decoder directions. Cancelling the diagnosed offset mitigates over-calling and improves overall accuracy with a negligible drop in call accuracy. Our work recasts over-calling from an empirical phenomenon into a mechanistic object amenable to causal correction. The code is available at https://anonymous.4open.science/r/agent-sae-904F.
Abstract: Understanding how deep neural networks learn useful internal representations from data remains a central open problem in the theory of deep learning. We introduce \emphNeural Low-Degree Filtering (Neural LoFi), a stylized limit of gradient-based training in which hierarchical feature learning becomes an explicit iterative spectral procedure. In this limit, the dynamics at each layer decouple: given the current representation, the next layer selects directions with maximal accessible low-degree correlation to the label. This yields a tractable surrogate mechanism for deep learning, together with a natural kernel-space interpretation. Neural LoFi provides a mathematically explicit framework for studying feature learning beyond the lazy regime. It predicts how representations are selected layer by layer, and gives a concrete mechanism by which depth progressively constructs new features from old ones. We complement the theory with mechanistic experiments on fully connected and convolutional architectures, showing that Neural LoFi improves over random-feature baselines, recovers meaningful structured filters, and predicts representations aligned with early gradient-descent feature discovery.
Abstract: With the increasing deployment of machine learning models, formal guarantees of the robustness and fairness of these models have become very important in safety-critical and legal compliance settings. However, model parameters are often commercial secrets that cannot be disclosed to auditors or end users. To this end, we present PANDA, a scalable system that uses Zero-Knowledge Proofs (ZKPs) to prove robustness and fairness properties of a model without revealing private model parameters. PANDA is built on top of CROWN an efficient robustness certification framework that is used in many state-of-the-art formal verification tools for neural networks. The core contribution of PANDA is a novel algorithm for proving linear relaxation bounds on non-linear activation layers, achieving simple and lightweight proofs. Remarkably, our system can generate proofs of local robustness for neural networks with more than 2.9M parameters in about 4 minutes, and can verify them in under 7.5 seconds. Prior ZKP-based robustness systems are based on exponential-time algorithms that cannot scale to nontrivial networks. In contrast, PANDA scales polynomially with the number of neurons in a network, allowing us to support neural networks 3 to 4 orders of magnitude larger than prior works with significantly reduced prover overhead.
Authors:
Zhen Zhang, Kaiqiang Song, Sean Wang, Yebowen Hu, Weixiang Yan, Chenyang Zhao, Henry P Zou, Haoyun Deng, Sathish Reddy Indurthi, Shujian Liu, Simin Ma, Xiaoyang Wang, Xin Wang, Song WangAbstract: AI agents increasingly solve real-world tasks by reasoning over multi-turn user interactions and invoking external tools. Yet reinforcement learning (RL) remains difficult in this setting: realistic objectives are often open-ended and lack verifiable rewards, RL for multi-turn, multi-step agentic tool use is underexplored, and executable tool environments are costly to build and maintain. We propose CM2, an RL framework that replaces verifiable outcome rewards with checklist rewards. CM2 decomposes each turn's intended behavior into fine-grained binary criteria with explicit evidence grounding and structured metadata, turning open-ended judging into more stable classification-style decisions. To balance stability and informativeness, CM2 uses sparse reward assignment with dense evaluation criteria, and trains in a scalable LLM-simulated tool environment to avoid engineering large tool sets. Experiments show that CM2 consistently improves over supervised fine-tuning. Starting from an 8B Base model and training on an 8k-example RL dataset, CM2 improves over the SFT counterpart by 8 points on \tau^2-Bench, 10 points on BFCL-V4, and 12 points on ToolSandbox. The results match or even outperform similarly sized open-source baselines, including the judging model. CM2 thus offers a scalable recipe for optimizing multi-turn, multi-step tool-using agents without verifiable rewards.
Abstract: Continual fine-tuning is essential for large language models (LLMs) to dynamically adapt to real-world environments, yet it inevitably suffers from catastrophic forgetting, particularly the performance degradation of previous tasks and LLMs' general-purpose knowledge. Although existing methods, such as orthogonal gradient projection, mitigate the forgetting across various fine-tuning tasks, they fundamentally fail to preserve pre-training LLMs' inherent general-purpose knowledge because the original data and gradients of off-the-shelf pre-training LLMs required by these methods are strictly unknown and highly diverse. To bridge this critical gap, we propose EoupCT, a novel framework designed to Estimate and Orthogonalize Unknown Pre-training gradients for Continual LLM fine-Tuning. Specifically, EoupCT estimates pre-training gradients by dynamically generating pseudo data that is most susceptible to forgetting for new tasks through a learnable soft prompt equipped with Gumbel-Softmax relaxation. Furthermore, we formulate a multi-objective optimization problem and introduce a first-order efficient Pareto optimizer that jointly optimizes LLM parameters and the soft prompt, rigorously enforcing orthogonality between new task updates and the estimated pre-training gradients. Extensive experiments across multiple LLMs demonstrate that EoupCT effectively preserves both task-specific proficiency and inherent general-purpose knowledge, successfully mitigating the catastrophic forgetting.
Abstract: Mixture-of-Experts (MoE) models enable scalable performance but face severe memory constraints on edge devices. Existing offloading strategies struggle with I/O bottlenecks due to the dynamic, low-information nature of autoregressive expert activation. In this paper, we propose to repurpose Speculative Decoding (SD) not merely as a compute accelerator, but as an informative lookahead sensor for memory management, supported by our theoretical and empirical analyses. Hence, we introduce MoE-SpAc, an MoE inference framework that integrates a training-free Speculative Utility Estimator to track expert demand, a Heterogeneous Workload Balancer to dynamically partition computation via online integer optimization, and an Asynchronous Execution Engine to unify the prefetching and eviction in the same utility space. Extensive experiments on seven benchmarks demonstrate that MoE-SpAc achieves a 42% improvement in TPS over the SOTA SD-based baseline, and an average 4.04 speedup over all standard baselines. Code is available at https://anonymous.4open.science/r/SpAc_MoE.
Abstract: Metric-induced discrete flow matching (MI-DFM) exploits token-latent geometry for discrete generation, but its practical use is limited by two issues: heuristic schedulers requiring hyperparameter search, and finite-step path-tracking error from its first-order continuous-time Markov chain (CTMC) solver. We address both issues. First, we derive a kinetic-optimal scheduler for prescribed scalar-parameterized probability paths, and instantiate it for MI-DFM as a training-free numerical schedule that traverses the path at constant Fisher–Rao speed. Second, we introduce a finite-step moment correction that adjusts the jump probability while preserving the CTMC jump destination distribution. We validate the resulting method, GibbsTTS, on codec-based zero-shot text-to-speech (TTS). Under controlled comparisons with a unified architecture and large-scale dataset, GibbsTTS achieves the best objective naturalness and is preferred in subjective evaluations over masked discrete generative baselines. Additionally, in comparison with the evaluated state-of-the-art TTS systems, GibbsTTS shows strong speaker similarity, achieving the highest similarity on three of four test sets and ranking second on the fourth. Demo and code are available in the supplementary materials.
Authors: Ziqi Wen, Parsa Madinei, Miguel Eckstein
Abstract: Evaluating whether large vision-language models (VLMs) align with human perception for high-level semantic scene comprehension remains a challenge. Traditional white-box interpretability methods are inapplicable to closed-source architectures and passive metrics fail to isolate causal features. We introduce Counterfactual Semantic Saliency (CSS). This black-box, model-agnostic framework quantifies the importance of objects by measuring the semantic shift induced by their causal ablation from a scene. To evaluate AI-human semantic alignment, we tested prominent VLMs against a human psychophysics baseline comprising 16,289 valid responses across 307 complex natural scenes and 1,306 high-fidelity counterfactual variants. Our analysis reveals a pervasive scene comprehension gap: models exhibit an overreliance (relative to humans) on large objects (size bias), objects at the center of the image (center bias), and high saliency objects. In contrast, models rely less on people in the scenes than our human participants to describe the images. A model’s size bias is a primary driver explaining variations in model-human semantic divergence.
Abstract: RLHF is widely used to align flow-matching text-to-image models with human preferences, but often leads to severe diversity collapse after fine-tuning. In RL, diversity is often assumed to correlate with policy entropy, motivating entropy regularization. However, we show this intuition breaks in flow models: policy entropy remains constant, even while perceptual diversity collapses. We explain this mismatch both theoretically and empirically: the constant entropy arises from the fixed, pre-defined noise schedule, while the diversity collapse is driven by the mode-seeking nature of policy gradients. As a result, policy entropy fails to prevent the model from converging to a narrow high-reward region in the perceptual space. To this end, we introduce perceptual entropy that captures diversity in a perceptual space and maintains the property of standard entropy. Building upon this insight, we propose two entropy-regularized strategies, Perceptual Entropy Constraint and Perceptual Constraints on Generation Space, to preserve perceptual diversity and improve the quality. Experiments across two base models, neural and rule-based rewards, and three perceptual spaces demonstrate consistent gains in the quality-diversity trade-off; PEC achieves the best overall score of 0.734 (vs.\ baseline's 0.366); a complementary setting of PEC further reaches a diversity average of 0.989 (vs.\ baseline's 0.047). Code will be released.
Abstract: Recent feed-forward 3D gaussian splatting methods have made dramatic progress on individual aspects of 3D scene reconstruction, but no existing method jointly addresses dynamic content, multi-view input, and unknown camera poses in a single feed-forward pass. Methods that handle dynamics either require accurate camera poses or accept only monocular input; pose-free multi-view methods address only static scenes; and per-scene optimization methods bridge some of these gaps but at minutes-to-hours cost per scene. We introduce NoPo4D, the first feed-forward system that addresses this empty quadrant. Building on a pretrained geometry backbone and recent 4D Gaussian frameworks, NoPo4D introduces a velocity decomposition that splits Gaussian motion into per-pixel image-plane shifts and depth changes, allowing direct supervision from pseudo ground-truth optical flow on the 2D component. This sidesteps both the differentiable rendering that couples prior posed methods to pose accuracy and the 3D motion ground truth that prior pose-free methods require. The system is rounded out by a bidirectional motion encoder for cross-view and cross-frame feature aggregation, and view-dependent opacity that mitigates cross-view and cross-timestep Gaussian misalignments. On four multi-view dynamic benchmarks, NoPo4D consistently outperforms prior feed-forward baselines, and with an optional post-optimization stage surpasses per-scene optimization methods, while running orders of magnitude faster. Code and the pretrained weights will be made publicly available.
Abstract: Scene-level 3D generation has long been dominated by 2D multi-view or video diffusion models. Typically, these 2D-based approaches perform scene generation in a 2D latent space, which introduces two fundamental issues: (i) representing 3D scenes via 2D views leads to significant representation redundancy, and (ii) latent space rooted in 2D inherently limits the spatial consistency of the generated scenes. In this paper, we propose to perform 3D scene generation directly within a high-dimensional implicit 3D latent space derived from powerful 2D foundation models (e.g., DINOv2, SigLIP2, DA3). Then we tame diffusion transformers to perform diffusion modeling directly in the 3D latent space, enabling 3D-grounded scene generation. In our experiments, we not only demonstrate the advantages of our method over 2D-based approaches in terms of efficiency and spatial consistency, but also comprehensively analyze the impact of 3D latent representation on 3D scene generation performance. Furthermore, we validate the favorable scalability of our method with respect to representation and model capacity.
Abstract: Online video understanding requires models to robustly encode, store, and retrieve information from a continuous stream to support accurate question answering. Existing approaches rely on key--value caching to accumulate frame-level representations over time, but allocate only a limited token budget per frame, causing loss of fine-grained visual detail. We find that naively scaling this budget degrades retrieval quality and question-answering performance, as denser token streams introduce redundancy that disrupts query--frame similarity. To address this, we first introduce an adaptive selection strategy that reduces token redundancy while preserving local spatiotemporal information. Second, we propose a training-free retrieval ensemble that leverages external models to better identify relevant frames. Our method, MemStream, achieves +8.0% on CG-Bench, +8.5% on LVBench, and +2.4% on VideoMME (Long) over ReKV with Qwen2.5-VL-7B. Furthermore, we show that MemStream consistently improves performance across several Video-LLMs for long-video understanding.
Authors: Darshan Deshpande
Abstract: Recent growth in different reinforcement learning (RL) techniques have surfaced a need for a wide variety of specialized training environments. These environments are typically hand-curated, with task and reward difficulties that are fixed rather than adaptive, making them ineffective training signals once a model's performance on the domain improves. As models continue to improve on these environments and reward signals grow increasingly sparse over longer horizons, the model encounters fewer diverse situations during rollouts, leaving it prone to overfitting on specific workflows or tool structures, also known as mode collapse. World models that simulate environment states have previously matched the performance of pure environment rollouts, making them a promising avenue for scaling diversity given that their outputs can be varied on-demand and at scale. However, autoregressive (AR) world models suffer from a fundamental left-to-right bias that prevents them from conditioning on globally interdependent state anchors such as tool schemas, prior turns, and expected outcomes. In this work, we (i) formalize text-based world modeling as a steerable transition-dynamics problem decomposed into initial environment state, task context, tool schemas, domain rules, and steering directives, and (ii) curate a dataset of 239,403 grounded state–action trajectories spanning nine open-source environments and twelve frontier model families. Using this dataset, we present a comparative study between AR LMs and masked diffusion language models (MDLMs), and show that MDLMs, by virtue of bidirectional anchor-aware denoising, produce better coherence, groundedness, and empirically validated rollout diversity than LLMs more than 4× their total parameter size, with comparable inference latency. We introduce a plug-and-play GRPO training framework with deterministic state checks, and perform zero-shot transfer ablations on three out-of-distribution environments (ScienceWorld, ALFWorld, AppWorld) across three agent backbones from 1.2B–7B parameters (LFM2.5, Qwen3, Mistral), achieving absolute gains of up to 47% over raw baselines without environment-specific fine-tuning. Finally, we conduct a behavioral analysis of failure modes under adversarial scenarios and a human evaluation centered on realism, outcome correctness, and training utility to showcase their reliability. We open source our work to encourage research in this direction.
Abstract: Diffusion and flow-based generators are increasingly used in robot policies to produce action chunks from visual observations, proprioception, and language instructions. However, closed-loop robot control requires solving action-generation problems of varying difficulty, whereas standard diffusion and flow-matching inference typically follows a fixed integration schedule. We introduce , a time-unconditional framework that turns action synthesis from fixed-time integration into iterative optimization. GeCO learns a stationary velocity field over action sequences, and refines actions until the field norm becomes small. This enables adaptive computation at each planning call: simple states can terminate early, while difficult states can use additional refinement. The same field norm also provides a lightweight signal for non-convergent action generation under distribution shift. GeCO can be instantiated in both diffusion-transformer policies and flow-matching Vision-Language-Action (VLA) systems without changing the surrounding policy interface. Across simulation benchmarks and real-world tasks, GeCO matches or improves task performance while enabling convergence-based inference as a plug-and-play replacement for standard time-conditioned action generators.
Abstract: We present MaskGen, a theoretically grounded and deliberately simple approach for domain generalization in 3D biomedical image segmentation. Modern segmentation models degrade sharply under shifts in modality, disease severity, clinical sites, and more, limiting their reliable adoption. Existing generalization methods address this using extreme augmentations, hand-engineered domain statistics mixing, or architectural redesigns that add significant implementation overhead while yielding inconsistent performance across biomedical settings. MaskGen instead presents a principled learning strategy with marginal overhead that utilizes both source-domain image intensities and domain-stable foundation model representations to train robust segmentation models. As a result, MaskGen achieves strong gains in both fully supervised and few-shot segmentation across broad clinical shifts in biomedical studies. Unlike prior approaches, MaskGen is architecture- and loss-agnostic, compatible with standard augmentation pipelines, easy to implement, and tackles arbitrary anatomical regions.
Abstract: Computer use agents (CUAs) are emerging as a powerful interface for automating complex digital workflows through visual perception and GUI execution. Online reinforcement learning with verifiable rewards (RLVR) has emerged as a key direction for scaling their capabilities. However, applying this paradigm is severely bottlenecked by verifiable data scarcity and online RL inefficiency. To break these barriers, we introduce , a unified framework that scales online RL for CUAs by combining verifiable task synthesis with efficient online RL. At the data level, we design , an end-to-end framework for generating verifiable RL tasks through iterative docker interactions and a multi-agent feedback loop. Scaled to 100+ concurrent agent workers via a shared docker interaction probe, this pipeline produces high-quality RL tasks for online RL. To maximize sample efficiency over this scaled pool, we propose , which dynamically tracks the model's per-task capability and allocates rollouts to tasks at the current learning frontier. On the training side, we further design , which keeps the recent visual context within a bounded window to balance rollout and training engine pressure on long-horizon trajectories, yielding a on ScienceBoard, establishing new state-of-the-art performance among open-source computer use agents. Source code:
Abstract: As Vision-Language Models (VLMs) are increasingly deployed as autonomous cognitive cores for embodied assistants, evaluating their privacy awareness in physical environments becomes critical. Unlike digital chatbots, these agents operate in intimate spaces, such as homes and hospitals, where they possess the physical agency to observe and manipulate privacy-sensitive information and artifacts. However, current benchmarks remain limited to unimodal, text-based representations that cannot capture the demands of real-world settings. To bridge this gap, we present ImmersedPrivacy, an interactive audio-visual evaluation framework that simulates realistic physical environments using a Unity-based simulator. ImmersedPrivacy evaluates physically grounded privacy awareness across three progressive tiers that test a model's ability to identify sensitive items in cluttered scenes, adapt to shifting social contexts, and resolve conflicts between explicit commands and inferred privacy constraints. Our evaluation of 12 state-of-the-art models reveals consistent deficits. In cluttered scenes, all models exhibit monotonic performance decay as scene complexity grows due to perceptual deficit. When social context shifts, no model exceed 65% selection accuracy. Under conflicting commands, the best model gemini-3.1-pro perfectly balances task completion and privacy preservation in only 51% of cases. These findings reveal that current VLMs in the physical world suffer from perceptual fragility and fail to let their knowledge of privacy cues govern their situated behavior. Our code and data is available at https://anonymous.4open.science/r/immersed-privacy-review-3A3C.
Abstract: Humans rapidly learn abstract knowledge when encountering novel environments and flexibly deploy this knowledge to guide efficient and intelligent action. Can modern AI systems learn and plan in a similar way? We study this question using a dataset of complex human gameplay with concurrent fMRI recordings, in which participants learn novel video games that require rule discovery, hypothesis revision, and multi-step planning. We jointly evaluate models by their ability to play the games, match human learning behavior, and predict brain activity during the same task, comparing a suite of frontier Large Reasoning Models (LRMs) against model-free and model-based deep reinforcement learning agents and a Bayesian theory-based agent. We find that frontier LRMs most closely match human behavioral patterns during game discovery and predict brain activity an order of magnitude better than both reinforcement learning alternatives across cortical and subcortical regions, with effects robust to permutation controls. Through targeted manipulations, we further show that brain alignment reflects the model's in-context representation of the game state rather than its downstream planning or reasoning. Our results establish LRMs as compelling computational accounts of human learning and decision making in complex, naturalistic environments.
Abstract: Sign-based optimization algorithms, such as SignSGD and Muon, have garnered significant attention for their remarkable performance in training large foundation models. Despite this empirical success, we still lack a theoretical understanding of when and why these sign-based methods outperform vanilla SGD. The core obstacle is that under standard smoothness and finite variance conditions, SGD is known to be minimax optimal for finding stationary points measured by \ell_2-norms, thereby fundamentally precluding any complexity gains for sign-based methods in standard settings. To overcome this barrier, we analyze sign-based optimizers leveraging \ell_1-norm stationarity, \ell_\infty-smoothness, and a separable noise model, which can better capture the coordinate-wise nature of signed updates. Under this distinct problem geometry, we derive matched upper and lower bounds for SignSGD and explicitly characterize the problem class in which SignSGD provably dominates SGD. Specifically, we compare the upper bound of SignSGD with the lower bound of SGD, illustrating that SignSGD effectively reduces the complexity by a factor of d under sparse noise, where d is the problem dimension. Furthermore, we elevate this framework to the matrix domain, providing an equivalent optimal lower bound for the Muon optimizer, proving that extending the sign operator to matrices preserves this optimal scaling with dimensionality. Finally, we bridge our theoretical bounds to practice, demonstrating that the theoretical superiority of SignSGD accurately predicts its faster convergence during the pretraining of a 124M parameter GPT-2 model.
Abstract: Video Large Language Models (Video-LLMs) have made rapid progress on temporal video understanding, yet many fail at a basic perceptual primitive: signed image-plane motion direction. On simple videos of a single object moving left, right, up, or down, most Video-LLMs perform near chance, with above-chance cases largely attributable to prediction biases rather than genuine direction understanding. We call this failure directional motion blindness. We localize the failure by tracing motion direction information through the Video-LLM pipeline. Motion direction remains linearly accessible from the vision encoder, projector, and LLM hidden states, but the readout fails to bind this signal to the correct verbal answer option, revealing a direction binding gap. Although synthetic motion direction instruction tuning reduces this gap on the source domain, motion direction concept vector analysis shows that visual complexity weakens the signal magnitude and limits out-of-domain generalization. We introduce MODIRECT, a dataset family for motion direction instruction tuning and evaluation, and DeltaDirect, a diagnosis-driven, projector-level objective that predicts normalized 2-D motion vectors from adjacent-frame feature deltas. On MODIRECT-SYNBENCH, instruction tuning with DeltaDirect improves motion direction accuracy from 25.9% to 85.4%. On MODIRECT-REALBENCH, DeltaDirect improves real-world motion direction accuracy by 21.9 points over the vanilla baseline without real-world tuning data, while preserving standard video-understanding performance.
Abstract: Deployed approaches for AI text detection often rely on training-time access to labeled datasets of both human-written and AI-generated text. This approach is vulnerable to three types of distribution shifts that occur continually post-deployment, and for which labeled data is often unavailable: adversarial humanization, new LLMs being released, and temporal drift in human writing. Simultaneously, existing approaches do not leverage a key signal of LLM usage: inference-time homogeneity. We propose a test-time adaptation (TTA) approach, using semi-supervised learning, that adapts to distribution shifts by leveraging homogeneity among unlabeled samples observed at inference time. Empirically, we find that state-of-the-art supervised detectors systematically fail when they encounter distribution shifts in AI-generated and human writing, both adversarial and natural, while test-time adaptation with semi-supervised learning is largely robust; e.g., the commercial model Pangram detects just 24.1% of our adversarial AI-generated text, compared to 90.5% for our test-time approach. We establish that test-time adaptation is a promising framework for AI text detection in the wild.
Authors: Kai Yao
Abstract: Passive image provenance asks whether pixels alone can reveal where an image came from: a human, an aggregate AI class, or a particular generator. This becomes a robustness problem once a source image can be edited before the verifier sees it. We study the problem as source--target verification under adversarial distribution shift. Our first result gives the exact best-case limit for any image-only verifier: the largest robust target-acceptance gap equals the minimum total-variation distance between the target distribution and the set of attacked source distributions. This quantity depends on the source, target, and edit class, not on the verifier architecture. Our second result explains why deployed public verifiers can fail before this statistical limit is reached. If the verifier can be emulated on the attack region to error \varepsilon, then a surrogate black-box attack reaches target acceptance within 2\varepsilon plus optimization error of the white-box optimum; score-revealing logistic and softmax heads over public features are identifiable, and approximate score access gives stable recovery bounds. Experiments on same-prompt real/diffusion benchmarks support this separation. Public CLIP-based interfaces collapse under targeted attacks, and stronger clean or adversarial retraining does not restore positive operational separation. In our evaluated settings, positive empirical upper bounds on the robust gap appear only under private low-bandwidth binary interfaces with abstention. The main lesson is simple: robust passive provenance requires both a source--target statistical analysis and an interface-aware evaluation of what the verifier reveals.
Abstract: Modern machine learning models are deployed in diverse, non-stationary environments where they must continually adapt to new tasks and evolving knowledge. Continual fine-tuning and in-context learning are costly and brittle, whereas neural memory methods promise lightweight updates with minimal forgetting. However, existing neural memory models typically assume a single fixed objective and homogeneous information streams, leaving users with no control over what the model remembers or ignores over time. To address this challenge, we propose a generalized neural memory system that performs flexible updates based on learning instructions specified in natural language. Our approach enables adaptive agents to learn selectively from heterogeneous information sources, supporting settings—such as healthcare and customer service—where fixed-objective memory updates are insufficient.
Abstract: Real-world time series are often highly incomplete and irregular due to sensor dormancy, transmission delays, and event-driven sampling, making reliable forecasting fundamentally challenging. Existing methods have evolved from impute-then-forecast pipelines to continuous-time models such as Neural ODEs and continuous-time graph networks. While these approaches improve the modeling of historical irregularity, they still rely on an implicit oracle assumption at inference time: the timestamps of future valid observations are presumed to be known in advance. This assumption limits practical relevance, since in many real systems the more fundamental question is not only what the future value will be, but also whether a valid observation will occur at all. In this paper, we propose Timeflies, a unified framework that reformulates forecasting as a joint problem of future observability inference and value estimation. To explicitly model the interaction between observation dynamics and state evolution, Timeflies adopts an observation stream and a value stream, coupled through three dedicated modules for reliability-aware embedding, observation-guided dependency modeling, and joint prediction. We further construct Glimpse, a benchmark that combines natural missingness from public datasets with real-world industrial data. Extensive experiments show that Timeflies consistently outperforms existing methods, highlighting the importance of explicitly modeling future observability in time series forecasting with missing values. Code and dataset are available in the supplementary material.
Abstract: Visual Geometry Grounded Transformer (VGGT) advances 3D reconstruction via scalable Transformer architecture, but the quadratic complexity of global attention prevents long context application. StreamVGGT enables streaming with causal attention, yet its KV cache grows linearly with frames, causing memory overflow and quality degradation. We present RetrieveVGGT, a training-free framework, which formulates context construction for VGGT as a retrieval problem. By retrieving a fixed number of relevant frames at each step, VGGT maintains a controllable memory budget, which is close to its training context length. Interestingly, we find that the similarity between current frame queries and cached history frame keys at the first global attention layer of VGGT is already a strong indicator of relevance, eliminating the need for additional learned scoring. To enhance information diversity similar to a recommender system, we propose Segment Sampling so that the retrieval spans distinct relevant segments rather than a single high-similarity region. We design a pose-aware spatial memory mechanism that organizes history frames according to their already estimated camera poses, enabling location-aware retrieval. Extensive experiments demonstrate that RetrieveVGGT achieves state-of-the-art performance, outperforming StreamVGGT, TTT3R, and InfiniteVGGT while maintaining constant memory usage regardless of sequence length.
Abstract: Training large multimodal models (LMMs) via reinforcement learning (RL) to natively invoke video-processing tools (e.g., cropping) has become a promising route to long-video understanding. However, existing native-RL methods dispatch tool calls sequentially (i.e., one per turn): a single wrong crop propagates errors without peer correction, multi-turn tool calls corrupt context, and inference cost scales linearly with the number of turns. We introduce ParaVT, the first multi-agent end-to-end RL-trained framework for Parallel Video Tool calling, dispatching multiple time-window crops in a single turn for cleaner context and better fault tolerance. Yet applying standard RL to ParaVT reveals an obstacle we term the Tool Prior Paradox: the pretrained tool priors that enable tool exploration also destabilize cold-started structural format and expose the skip-tool reward shortcut under temperature sampling. A cross-model contrast on a weaker-prior LMM supports this claim: format stays stable but RL elicits zero tool calls, indicating that prior strength is the shared driver of both format collapse and tool exploration. We propose PARA-GRPO (Parseability-Anchored and Ratio-gAted GRPO), which augments standard RL with two complementary mechanisms: (i) a targeted format reward applied only at the structural-token positions most prone to collapse, and (ii) a per-prompt frame-budget randomization that creates training prompts where calling the tool yields a measurable reward signal over skipping it. Across six long-video understanding benchmarks, ParaVT improves over the Qwen3-VL baseline by +7.9% on average, with PARA-GRPO lifting training-time format compliance from 0.13 to 0.64. As tool capabilities become increasingly internalized in modern LMMs, RL must cooperate with the resulting priors, and ParaVT offers a general recipe for agentic RL. Code, data, and model weights will be released.
Abstract: End-to-end autonomous driving aims to generate safe and plausible planning policies from raw sensor inputs, and constructing an effective scene representation is a critical challenge. Driving world models have shown great potential in learning rich representations by predicting the future evolution of a driving scene. However, existing driving world models primarily focus on visual scene representation, and motion representation is not explicitly designed to be planner-shared and inheritable, leaving a schism between the optimization of visual scene generation and the requirements of precise motion planning. We present WorldDrive, a holistic framework that couples scene generation and real-time planning via unifying vision and motion representation. We first introduce a Trajectory-aware Driving World Model, which conditions on a trajectory vocabulary to enforce consistency between visual dynamics and motion intentions, enabling the generation of diverse and plausible future scenes conditioned on a specific trajectory. We transfer the vision and motion encoders to a downstream Multi-modal Planner, ensuring the driving policy operates on mature representations pre-optimized by scene generation. A simple interaction between motion representation, visual representation, and ego status can generate high-quality, multi-modal trajectories. Furthermore, to exploit the world model’s foresight, we propose a Future-aware Rewarder, which distills future latent representation from the frozen world model to evaluate and select optimal trajectories in real-time. Extensive experiments demonstrate that WorldDrive achieves leading planning performance among vision-only methods while maintaining high-fidelity action-controlled video generation capabilities. Our code and model will be made publicly available.
Abstract: Reinforcement learning (RL) has become a key training step for improving mathematical reasoning in large language models (LLMs), but it often has high GPU memory usage, which makes it hard to use in settings with limited resources. To reduce these issues, we propose Evolution Strategies with Sharpness-Aware Maximization (ESSAM), a full parameter fine-tuning framework that tightly combines the zeroth-order search in parameter space from Evolution Strategies (ES) with the Sharpness-Aware Maximization (SAM) to improve generalization. We conduct fine-tuning experiments on the mainstream mathematica reasoning task GSM8K. The results show that ESSAM achieves an average accuracy of 78.27% across all models and its overall performance is comparable to RL methods. It surpasses classic RL algorithm PPO with an accuracy of 77.72% and is comparable to GRPO with an accuracy of 78.34%, and even surpassing them on some models. Further generalization experiments show that the models trained with ESSAM exhibit stronger generalization ability. Their average performance achieves the best results on 5 out of 6 datasets, indicating that ESSAM can effectively improve the generalization performance of fine-tuned models. In terms of GPU memory usage, ESSAM reduces the average GPU memory usage by 18× compared to PPO and by 10× compared to GRPO, achieving an extremely low GPU memory usage. In addition, we design an accelerated variant of ESSAM, which achieves nearly a twofold speedup while maintaining the same GPU memory usage as ESSAM, and attains an average accuracy of 78.02% across all models, outperforming PPO. Code: https://anonymous.4open.science/r/ESSAM-3F4F/
Abstract: Self-supervised learning (SSL) pretrained models have become a dominant paradigm for visual representation learning, but they are vulnerable to backdoor attacks. Existing defenses struggle to defend against such attacks in a fully black-box setting because they often require access to labels, attack patterns, or training data. To tackle this issue, we propose a new attack-agnostic, model-agnostic, and modality-agnostic black-box test-time defense paradigm, called \emphPlatonic Representation Defense. It is inspired by the Platonic Representation Hypothesis, which suggests that large-scale independently trained encoders converge toward compatible projections of the same underlying reality. We formalize this idea as a conditional energy function defined over source representations and a set of reference representations. The energy function is trained for detection through noise-contrastive estimation and for representation purification through denoising score matching. Theoretically, the energy gap between matched and mismatched samples is lower bounded by the mutual information between source and reference representations. We demonstrate the effectiveness of our method on multiple self-supervised encoders and more than 10 attacks. The method can perform both representation detection and purification, and achieves substantial performance gains across multiple attacks. Code can be found in the appendix.
Abstract: Cascade attacks in LLM multi-agent systems (MAS) arise when adversarial influence propagates across agents, leading to escalated system-level failures through complex interaction dynamics. Detecting such cascades is challenging, as their signals are distributed and tightly coupled across interaction channels, often appearing locally benign while unfolding either within a single turn or gradually across multiple turns. Existing defenses, being largely local and text-centric, fail to capture the cross-channel, temporally coordinated structure of cascade propagation. We present CASPIAN, the first approach to provide a unified, cross-channel causal view of cascade behavior in LLM-MAS through online monitoring of dynamic influence propagation across agents. CASPIAN models multi-agent interactions as a unified, dynamic causal influence matrix across channels, estimated efficiently via a late-interaction conditional transfer entropy (LI-CTE) formulation, enabling the detection of cascade onset from emergent system-level structure rather than isolated anomalies. It then performs online causal attribution, identifying the origin, bridge, and amplifier agents driving the cascade and reconstructing its principal propagation pathways, capabilities not supported by existing methods. Across diverse multi-agent frameworks and benchmarks, CASPIAN consistently outperforms semantic guardrails, LLM-based judges, and graph-based anomaly detectors in both detection accuracy and early cascade identification, while operating with negligible additional latency. These results demonstrate that unified cross-channel causal modeling is essential for reliably detecting and understanding cascade failures in LLM multi-agent systems.
Abstract: Multimodal large language models (MLLMs) show remarkable potential for scientific reasoning, yet their performance in specialized domains such as microscopy remains limited by the scarcity of domain-specific training data and the difficulty of encoding fine-grained expert knowledge into model parameters. To bridge the gap, we introduce MicroWorld, a framework that constructs a multimodal attributed property graph (MAPG) from large-scale scientific image-caption corpora and leverages it to augment MLLM reasoning at inference time without any domain-specific fine-tuning. MicroWorld extracts biomedical entities and relations via scispaCy or LLM-based triplet mining, aligns images and entities in a shared embedding space using Qwen3-VL-Embedding, and assembles a knowledge graph comprising approximately 111K nodes and 346K typed edges spanning eight relation categories. At inference time, a graph-augmented retrieval pipeline matches query entities to the MAPG and injects structured knowledge context into the MLLM prompt. On the MicroVQA benchmark, MicroWorld improves the reasoning performance of Qwen3-VL-8B-Instruct by 37.5%, outperforming GPT-5 by 13.0% to achieve a new state-of-the-art. Furthermore, it yields a 6.0% performance gain on the MicroBench benchmark. Extensive experiments demonstrate the enhanced generalization capability introduced by MicroWorld. A qualitative case study further reveals both the mechanisms through which structured knowledge improves reasoning and the failure modes that point to promising future directions. Code and data are available at anonymous GitHub.
Abstract: (RTs) before producing a final output. Despite the rapid advancement, our understanding of how RTs support reasoning, and where this paradigm fails, remain incomplete. To promote greater clarity, we introduce PITA: a novel large-scale dataset of over 23 million statements in propositional logic and their corresponding proofs. As a benchmark for robust reasoning, we focus on length generalization: if a model is trained to determine truth or falsity on statements with proofs up to fixed length, how well does it generalize to statements requiring longer proofs? We propose notions of (1) , which measure respectively (1) the number of proof steps required to solve a proposition and (2) the number of unique propositions within a task family. We vary these quantities across subsets of PITA, and find that RT models generalize well on relatively subsets compared to non-RT baselines. As a controlled point of comparison, we study a separate, tractable transitive inference task that exhibits qualitatively similar behavior. Our accompanying theory explains the scalings observed in this simpler setting, suggesting one mechanism by which breadth can favor RT models while depth can expose long-context weaknesses. Our findings suggest salient benefits and limitations of proof-like reasoning traces in controlled length-generalization settings.
Abstract: Reinforcement learning with verifiable rewards (RLVR) for Large Reasoning Models hinges on baseline estimation for variance reduction, but existing approaches pay a heavy price: PPO requires a policy-model scale critic, while GRPO needs multiple rollouts per prompt to keep its empirical group mean stable. We introduce POISE (\underlinePolicy \underlineOptimization with \underlineInternal \underlineState Value \underlineEstimation), which obtains a baseline at negligible cost by using the policy model's internal signals already computed during the policy forward pass. A lightweight probe predicts the expected verifiable reward from the hidden states of the prompt and generated trajectory, as well as token-entropy statistics, and is trained online alongside the policy. To preserve gradient unbiasedness despite using trajectory-conditioned features, we introduce a cross-rollout construction that predicts each rollout's value from an independent rollout's internal states. Because POISE estimates prompt value using only a single rollout, it enables higher prompt diversity for a fixed compute budget during training. This reduces gradient variance for more stable learning and also eliminates the compute overhead of sampling costs for detecting zero-advantage prompts. On Qwen3-4B and DeepSeek-R1-Distill-Qwen-1.5B across math reasoning benchmarks, POISE matches DAPO while requiring less compute. Moreover, its value estimator shows similar performance to a separate LLM-scale value model and generalizes to various verifiable tasks. By leveraging the model's own internal representations, POISE enables more stable and efficient policy optimization.
Abstract: Order dispatch is a critical task in ride-sharing systems with Autonomous Vehicles (AVs), directly influencing operational efficiency and profitability. While Multi-Agent Reinforcement Learning (MARL) offers a scalable paradigm by decomposing the massive state-action space, existing methods are heavily reliant on accurate value function estimation. In large-scale, highly stochastic urban environments, such estimation is notoriously prone to bias and instability, severely limiting training efficiency and policy quality. To overcome this limitation, we propose two novel policy optimization methods that completely bypass the need for critic networks. First, we establish a formal connection between the homogeneity of AV fleets and the short-memory mixing property of the transportation network, proving that agent value functions converge to a common time-varying baseline up to a negligible residual. Leveraging this insight, we introduce Single-Trajectory Group Relative Policy Optimization (ST-GRPO), an adaptation of Large Language Models (LLMs) post-training techniques to multi-agent trajectories, which replaces the traditional value baseline with the fleet-wide average reward-to-go. Inspired by this reduction, we further derive One-Step Policy Optimization (OSPO), demonstrating that under the established structural priors, an optimal policy can be learned using only immediate, group-normalized rewards—rendering long-horizon bootstrapping unnecessary. Experiments on real-world ride-hailing datasets from Manhattan and Queens demonstrate that both ST-GRPO and OSPO achieve promising performance, particularly in reducing pickup times and increasing order service rates. Remarkably, both methods operate efficiently using simple Multilayer Perceptron (MLP) networks with minimal GPU utilization, underscoring the practical impact of exploiting domain structure for scalable MARL. Our code, trained models, and processed data are provided at the anonymous repository: https://anonymous.4open.science/r/OSPO-2105 .
Abstract: Hyper-Connections (HC) extend residual connections into multiple streams, employing residual matrices for cross-stream mixing to enrich model expressivity. However, unconstrained mixing disrupts the identity mapping property intrinsic to the residual connection, causing unstable training. To address this, Manifold-Constrained Hyper-Connections (mHC) and its variants restrict these matrices to be doubly stochastic via Sinkhorn-Knopp (SK) algorithm or permutation-based parameterizations. We reveal three limitations of this doubly stochastic constraint: (1) identity degeneration, where learned matrices collapse around the identity initialization and diminish cross-stream interactions, (2) an expressivity bottleneck via spectral collapse, where the doubly stochastic constraint induces feature homogenization, and (3) parameterization inefficiencies, manifesting as unstable SK iterations or the factorial-scaling overhead of permutation-based parameterizations. To overcome these flaws, we propose Spectral-Sphere-Constrained Hyper-Connections (sHC). By confining residual matrices to a spectral norm sphere, sHC prevents spectral collapse and enables selective feature diversification. This shift eliminates unstable SK iterations and factorial parameterization, enabling expressive, non-degenerate residual matrices while preserving training stability.
Abstract: Speculative decoding accelerates LLM inference by using a fast draft model to generate tokens and a more accurate target model to verify them. Its performance depends on the , or number of draft tokens accepted by the target. Our studies show that the acceptance length of even state-of-the-art speculators, like DFlash, EAGLE-3 and PARD degrade with generation length, reaching values close to 1 (i.e. no speedup) within just a few thousand output tokens, making speculators ineffective for long-response tasks. Acceptance lengths decline because most speculators are trained offline on short sequences, but are forced to match the target model on much longer outputs at inference, well beyond their training distribution. To address this issue, we propose , an online distillation approach that continuously adapts the speculator at test-time. TTS leverages the key insight that the token verification step already invokes the target model for each draft token, providing the training signal needed to adapt the draft at no additional cost. Treating the draft as the student and the target as a teacher, TTS adjusts the draft over several speculation rounds, with each update improving the draft's accuracy as generation proceeds. Our results across multiple models from the Qwen-3, Qwen-3.5, and Llama3.1 families show that TTS improves acceptance lengths over state-of-the-art speculators by up to 72% and 41% on average, with the benefits scaling with increased generation lengths.
PaperID: 718, Poster
Abstract: We consider the problem of Cost-Aware Learning, where sampling different components of a finite-sum objective incurs different costs. The objective is to reach a target error while minimizing the total cost. We propose Cost-Aware SGD, which uses a distribution based on gradient norms and costs to sample components. We provide a thorough analysis of this algorithm, including cost-improvement bounds over baselines, a characterization of distribution proxy sub-optimality, and a lower bound. We apply our theoretical insights to reinforcement learning with language models, where the computational cost of sequence-level policy gradients varies with length. We find that the advantage magnitude serves as a high-fidelity proxy for gradient norms, and use this to introduce Cost-Aware GRPO. Empirical results on 1.5B, 4B, and 8B LLMs demonstrate that this algorithm significantly reduces the tokens used in policy optimization while matching or exceeding baseline accuracy.
PaperID: 719, Poster
Abstract: We study stochastic multi-objective optimization when each objective may be non-smooth and non-convex. We introduce the notion of (\delta,\epsilon)-Pareto Goldstein stationary points (PGSP) to characterize the convergence of solving multi objective optimization. The (\delta,\epsilon) generalizes the definition of (\delta,\epsilon)-Goldstein stationary point for single objective non-smooth non-convex optimization and the \epsilon-Pareto stationary point for solving multi objective optimization with smooth objectives. We propose the multi-gradient-free method (MGFM), which finds (\delta,\epsilon)-PGSP with \tilde\mathcalO(m^2d^3/2\delta^-1\epsilon^-4) oracle complexity to the stochastic function values, where m is the number of objectives and d is the problem dimension. We further propose MGFM+, a faster multi-gradient-free method based on variance reduction, which improves the complexity to \tilde\mathcalO(m^2d^3/2\delta^-1\epsilon^-3). These provide the first non-asymptotic analysis for multi-nonsmooth-nonconvex-objective optimization. \endabstract
PaperID: 720, Poster
Authors: Haotian Ma, Ruqi Yang, Philip Townsend
Abstract: matter through global channel scores. This separation is limiting for multi-spectral data, where channel relevance is often spatially localized and physically meaningful. We introduce , a full-grid attribution method over position-channel players (g, c). For each spatial position, SSI fixes the observed surrounding context, defines a conditional game over the channels at that position, and applies the Shapley value to this local game. The result is a full spectral-spatial attribution matrix rather than a single heatmap or a global channel ranking. SSI has a clear local semantics: it satisfies Shapley fairness within each position, and its channel aggregation exactly recovers the position-level contextual contrast obtained by replacing that entire position with the baseline. We also distinguish SSI from context-averaged Owen-style hierarchical accounting, which respects the same position-channel grouping but averages over external spatial coalitions. On HLS Burn Scar, SSI outperforms an MC-Owen estimator across retained attribution metrics while using only about 6% of its evaluation budget. Across CUB-200, EuroSAT, and HLS Burn Scar under both per-channel mean and zero baselines, SSI achieves the best joint deletion and insertion AUC over dense, flat, and rank-1 baselines, yields the strongest spectral hit rates and channel deletion, and leads spatial localization on both annotated datasets.
Abstract: On-policy distillation (OPD) has become a popular training paradigm in the LLM community. This paradigm selects a larger model as the teacher to provide dense, fine-grained signals for each sampled trajectory, in contrast to reinforcement learning with verifiable rewards (RLVR), which only obtains sparse signals from verifiable outcomes in the environment. Recently, the community has explored on-policy self-distillation (OPSD), where the same model serves as both teacher and student, with the teacher receiving additional privileged information such as reference answers to enable self-evolution. This paper demonstrates that learning signals solely derived from the privileged teacher result in severe information leakage and unstable long-term training. Accordingly, we identify the optimal niche for self-distillation and propose RLSD (RLVR with Self-Distillation). Specifically, we leverage self-distillation to obtain token-level policy differences for determining fine-grained update magnitudes, while continuing to use RLVR to derive reliable update directions from environmental feedback (e.g., response correctness). This enables RLSD to simultaneously harness the strengths of both RLVR and OPSD, achieving a higher convergence ceiling and superior training stability.
PaperID: 722, Poster
Abstract: Offline reinforcement learning requires improving a policy from fixed data while avoiding out-of-distribution actions with unreliable value estimates. Diffusion and flow policies handle this trade-off by modeling the behavior distribution to regularize the RL objective, but they require iterative denoising, solver integrations, and in more efficient variants, distillation or other approximations at inference. We propose DriftQL, which combines a drift-based behavioral regularizer with critic-driven policy improvement. The value signal biases the policy toward high-value regions of the data support, while attraction and repulsion together keep generated actions near the data and prevent collapse onto a single mode. DriftQL is implemented as a single network with a unified training objective and generates actions in a single forward pass. On D4RL and OGBench, DriftQL consistently outperforms diffusion and flow methods, advancing the state of the art. Under degraded data quality, where the baselines visibly struggle, DriftQL remains close to its clean-data performance, positioning it as a promising alternative to diffusion and flow-based methods while maintaining the simplicity and efficiency of deterministic approaches.
PaperID: 723, Poster
Abstract: Principal component analysis (PCA) is one of the most widely used unsupervised dimension reduction techniques. We study PCA for data from multiple related domains. Since principal components generally differ across domains, one way to obtain a shared low-rank embedding is to perform PCA on the pooled data. However, this approach can focus on spurious directions that exhibit high variation in only a few domains. To find a robust embedding that still explains most variance in unseen but similar domains, we propose instead to focus on shared directions of variation. To this end, we introduce Anchor PCA which trades off overall explained variance with agreement between the shared and domain-specific low-rank embeddings. Anchor PCA amounts to PCA on a modified target matrix and thus can be solved efficiently. Moreover, we show that Anchor PCA recovers a maximal invariant subspace and admits a minimax reconstruction interpretation under bounded domain-specific covariance inflations. On simulated and real-world gas sensor data with temporal drift, we demonstrate, respectively, that Anchor PCA recovers the maximally invariant subspace and yields embeddings that explain more variance on unseen domains than the pooling baseline and a worst-case alternative. Taken together, these findings establish Anchor PCA as a promising approach to robust unsupervised dimension reduction from multi-domain data.
Abstract: We study attention mechanisms through the lens of a canonical unsupervised problem: principal component analysis (PCA). We show that, when trained on Gaussian data, both softmax and linear attention layers learn parameters that align with the principal eigenvectors of the covariance matrix, thereby establishing a direct and explicit connection with PCA. Our analysis covers both finite and infinite prompt regimes. In the infinite-prompt limit, we prove convergence to globally optimal solutions aligned with the leading spectral direction, while in the finite-prompt setting we show that the same behavior emerges up to sampling effects. We further extend the analysis to an in-context setting with spiked Wishart covariances, where attention successfully recovers the underlying signal direction. These results demonstrate that attention inherently performs PCA-like computations under unsupervised objectives, providing a theoretical foundation for its representation-learning capabilities.
Abstract: Out-of-Distribution (OOD) detection requires sensitivity to subtle shifts without overreacting to natural In-Distribution (ID) diversity. However, from the viewpoint of detection granularity, global representation inevitably suppresses local OOD cues, while patch-based methods are unstable due to entangled spurious-correlation and noise. And neither of them is effective in detecting compositional OODs composed of valid ID components. Inspired by recognition-by-components theory, we present a post-hoc Component-Based OOD Detection (CoOD) framework that addresses the existing limitations by decomposing inputs into functional components. To instantiate CoOD, we derive Component Shift Score (CSS) to detect local appearance shifts, and Compositional Consistency Score (CCS) to identify cross-component compositional inconsistencies. Empirically, CoOD achieves consistent improvements on both coarse- and fine-grained OOD detection.
Abstract: We study online multicalibration beyond the worst-case. We give a single, efficient algorithm which dynamically interpolates between stochastic and worst-case sequences by adaptively refining a dyadic grid of prediction values. Its error is controlled by the number of leaves in the refinement tree. Our analysis recovers the known \widetilde O(T^2/3) worst-case-optimal rate for online multicalibration, while simultaneously automatically adapting to easier instances: in the marginal stochastic setting it obtains a rate of \widetilde O(\sqrt T), and for piecewise-stationary means with J segments its rate is \widetilde O(\sqrtJT). More generally, the rate depends on a threshold-complexity measure of the predictable mean process relative to the group family. We show that this dependence is tight up to logarithmic factors.
Abstract: Consider two large vision models. They process the same image and both correctly predict the class ``rabbit.'' How much of the circuit computation along the way was shared? Model diffing offers a natural lens on this question. So far, however, it has largely operated on a single layer and at the level of representations rather than circuits. In this work, we introduce Universal Circuits (UCs), enabling model diffing at the circuit level. Specifically, we extend CLTs across both layers and models, with losses that encourage sparsity for interpretability and output fidelity for faithfulness. We train UCs between pairs of standard large vision models. We find that a compact cross-model intersection of pruned class circuits, typically a few hundred universal (shared) features per pair, produces 89-98% of full-circuit classification accuracy in both models, while a complementary set of universal features (hundreds per pair) is kept by only one model's circuit, reflecting how each model weights shared concepts differently in its own representations. To demonstrate a downstream use of UC, we perform model surgery: a class is successfully transferred from one model to another with no gradient steps taken in the new model. More broadly, our work demonstrates that large vision models leverage shared multi-layer algorithms for downstream performance, and that these algorithms can be discovered, compared, and reused across models.
Abstract: Sampling from score-based diffusion models incurs bias due to both time discretisation and the approximation of the score function. A common strategy for reducing this bias is to apply corrector steps based on the unadjusted Langevin algorithm (ULA) at each noise level within a predictor-corrector framework. However, ULA is itself a \emphbiased sampler, as it discretises a continuous diffusion process. In this work, we consider \emphadjusted Langevin correctors that employ Metropolis--Hastings (MH) or Barker's accept-reject steps to correct for this bias. Since the target density ratio typically required by MH-based algorithms is unavailable, we propose methods that instead utilise the score function to compute the correct acceptance probability. We introduce the first exact method for adjusting Langevin corrections in diffusion models, based on a two-coin Bernoulli factory algorithm. We also propose an efficient approximation based on Simpson's rule that achieves accuracy of order 5/2 in the step size at near-zero marginal cost. We demonstrate that these procedures improve sample quality on both synthetic and image datasets, yielding consistent gains in Fr\'echet Inception Distance on the latter.
PaperID: 729, Poster
Authors: Ashish Singh, Prakash C Chhipa
Abstract: Fine-grained video action recognition often depends on how an action unfolds over time rather than on appearance alone. Visually similar classes may differ in preparation, execution, or completion phases, making uniformly aggregated temporal context suboptimal. We study this problem under a practical objective: i) improving recognition performance with few tunable parameters, ii) being compute efficient, and iii) using only the readily available video modality. Motivated by this missing temporal observation and the accuracy--compute efficiency--unimodality objective, we propose Directional Adapter (DiA), a parameter-efficient adapter on top of CLIP. DiA formulates temporal specialization into causal and anti-causal directions, allowing past-to-present and future-to-present cues to be modeled differently. To keep this directional modeling parameter- and compute-efficient, both directions use depth-wise temporal convolutions in a compact bottleneck space and are combined through a learnable fusion. DiA achieves state-of-the-art performance across 7 public benchmarks with only ~3.5M tunable parameters, while also delivering higher throughput and lower latency than prior methods. Code is available at https://anonymous.4open.science/r/diafullsupervised-8CC3/README.md.
Abstract: Using offline datasets to evaluate conversational agents often fails to cover rare scenarios or to support testing new policies. This has motivated the use of \emphcontrollable user simulators for targeted, counterfactual evaluation, typically implemented by prompting or fine-tuning large language models. In this work, we formalize controllable simulation as a causal inference problem. By bridging natural language evaluation with off-policy evaluation methodology, we show that the standard practice of training simulators via supervised fine-tuning on post-hoc trajectory labels yields a structurally biased model. Specifically, these labels are inextricably coupled to the data-generating behavior policy, injecting a \emphlook-ahead bias that breaks causal consistency. Furthermore, we prove that under policy shift this failure causes the variance of evaluation metrics to explode geometrically, a phenomenon we term \emphcontrollability collapse. To restore causal consistency, we establish theoretical conditions for accurate simulation and propose practical training mitigations: a priori controls, step-wise dynamic controls, and direct policy-conditioned learning. Empirical evaluation confirms that while standard global controls distort conversational distributions and collapse behavioral diversity, our causally grounded simulators eliminate look-ahead bias, preserve natural variance, and exhibit robust zero-shot generalization to unseen agent behaviors.
PaperID: 731, Poster
Abstract: Anytime-valid tests allow evidence to be checked during data collection: one can either continue testing or stop and reject the null while still controlling type-I error. Yet, in many applications rejection is useful only if it comes soon enough. We introduce a time-sensitive testing-by-betting framework that favours early rejection by assigning rewards to rejection times and maximising their expected value under a given alternative. This encompasses hard deadlines and softer time preferences. The resulting optimal control problem admits a Bellman representation in terms only of time and evidence against the null, rather than the full history. For hard deadlines, we identify a canonical optimal test in the simple-vs-simple case. We show that exponentially decaying rewards admit a stationary approximation that yields a finite-time-scale counterpart to the classical growth-rate-optimal (GRO) viewpoint in anytime-valid statistics, recovering it in the large-time-scale limit.
Abstract: Diffusion-based models decompose sampling into many small Gaussian denoising steps—an assumption that breaks down when generation is compressed to a few coarse transitions. Existing few-step methods address this through distillation, consistency training, or adversarial objectives, but sacrifice the likelihood framework in the process. We introduce Normalizing Trajectory Models (NTM), which models each reverse step as an expressive conditional normalizing flow with exact likelihood training. Architecturally, NTM combines shallow invertible blocks within each step with a deep parallel predictor across the trajectory, forming an end-to-end network trainable from scratch or initializable from pretrained flow-matching models. Its exact trajectory likelihood further enables self-distillation: a lightweight denoiser trained on the model's own score produces high-quality samples in four steps. On text-to-image benchmarks, NTM matches or outperforms strong few-step baselines in just four sampling steps while uniquely retaining exact likelihood over the generative trajectory.
PaperID: 733, Poster
Abstract: On-policy training plays an important role in modern machine learning, especially in reinforcement learning (RL) for language models. However, quantifying how training data influences the policy's behavior is difficult in on-policy settings: one must account for the cascading effect by which a training event affects its immediate policy update, which in turn affects future action samples, which further affects future policy updates, and so on. Attribution methods developed for off-policy and supervised learning (e.g. influence functions) do not account for such \emphpolicy-data interaction, and can therefore fail in on-policy settings. In this work, we formalize the problem of attributing behaviors to individual training events in the on-policy setting. We derive an influence measure that accounts for policy-data interaction while also being tractable to compute in practical training setups (e.g. LM finetuning) and for realistic optimizers (e.g. AdamW) at a cost that scales linearly in the total number of training events. Empirically, our influence measure accurately predicts the true effect of ablating specific events during RL, while traditional methods fail to do so. Further, we show that this influence accurately identifies helpful and harmful RL training events, allowing us to improve performance via influence-weighted retraining. Overall, we show that accounting for policy-data interaction is both necessary and tractable, allowing for principled training intervention in modern RL pipelines.
PaperID: 734, Poster
Authors: Dwij Mehta, Arvind Renganathan, Vipin Kumar
Abstract: Flow matching learns continuous-time transports from a simple source distribution to a target data distribution by training a velocity field on intermediate states along interpolating trajectories. However, because the model observes only the projected data-space state (x_t,t), different transport trajectories can induce conflicting velocity directions at nearby intermediate states, leading to curved transport paths that limit the fidelity of low-NFE sampling. To address this limitation, we propose \emphState Augmented Flows (SAF), a lightweight framework that lifts flow matching into a transient augmented state space. SAF transports an augmented state (x_t,z_t), where auxiliary coordinates provide additional contextual information during transport while vanishing at the terminal time. This enlarges the state observed by the velocity model without changing the final sample space, and is compatible with existing coupling strategies and flow-matching architectures. Across MNIST, CIFAR-10, and ImageNet-32, SAF consistently improves over the rectified-flow baseline, with the largest gains in the low-NFE regime. These results suggest that augmenting the transported state provides a new and complementary mechanism for improving flow-based generative modeling.
PaperID: 735, Poster
Abstract: We study \emphbilinear matching bandits, an online sequential decision problem in which at each round a learner assigns M users to N items via an injective matching \pi:[M]\hookrightarrow[N] and observes semi-bandit feedback with bilinear mean reward \mathbfx_i^\top \mathbf\Theta^\star \mathbfy_\pi(i), where the parameter \mathbf\Theta^\star \in \mathbbR^d_1 × d_2 is unknown. We propose \textscHybridRRD, a hybrid algorithm that combines high-width elimination with a row-wise regularized design sampler over the fractional matching polytope. Without assuming \mathbf\Theta^\star is low-rank, \textscHybridRRD achieves a high-probability regret bound of \widetildeO (d_2 \sqrtd_1MT + Md_1 d_2), suppressing problem-scale constants. This improves the leading dimension dependence from d_1d_2 to d_2\sqrtd_1 over the standard Kronecker-linearized contextual combinatorial semi-bandit baseline. We further prove a minimax lower bound of \Omega(d_2 \sqrtd_1 M T) in a large-item regime, matching the leading regret term up to logarithmic factors. Numerical experiments show that \textscHybridRRD performs favorably against Kronecker-linearized baselines.
PaperID: 736, Poster
Abstract: We introduce \emphuniversal transformers: fixed-parameter architectures that can simulate any transformer in a given class via a suitable input embedding. Analogous to a universal Turing machine, the input embedding encodes a description of the target model while all internal parameters remain fixed. We provide explicit sparse constructions achieving universality when the embedding dimension is sufficiently large, and further show that universality is generic: randomly initialized transformers are universal almost surely, which aligns with recent empirical results of Zhong and Andreas (2024). We empirically validate our theory on the algorithmic tasks of parenthesis balancing and multi-hop reasoning. Our results suggest that much of a transformer’s expressive power may reside in its input representation rather than its learned weights.
Abstract: We introduce a new tool, Express, for converting a non-causal attention approximation into a causal approximation with matching approximation guarantees. When combined with the state-of-the-art Thinformer approximation, Express improves upon the best known causal attention guarantees, delivering \log^3/2(n)/s approximation error with only O(s) memory and O(s^2 \log^2(n)) compression overhead for a sequence of length n. We pair these developments with an efficient I/O-aware Triton implementation, demonstrate substantial speedups over FlashAttention 2, and use Express to accelerate three distinct components of the language modeling pipeline: long-context prefill, KV cache compression, and long-form decoding.
Authors: Floris d Hengst, Inès Blin, Majid Mohammadi, Syed I Shah, Taraneh Younesian
Abstract: Conformal prediction (CP) provides reliable uncertainty quantification by generating prediction sets with finite-sample coverage guarantees, yet it typically ignores hierarchical relationships between classes. We introduce Hierarchical Conformal Classification (HCC), a framework that integrates class hierarchies into the structure and semantics of prediction sets. HCC uses a constrained optimization problem formulation to produce prediction sets composed of nodes at various hierarchy levels while maintaining rigorous coverage guarantees. To ensure computational efficiency, we prove that a restricted subset of well-structured candidate solutions is sufficient to maintain both optimality and coverage. Empirical evaluations across audio, image, and text benchmarks, alongside a user study, demonstrate that HCC outperforms state-of-the-art methods and that hierarchical prediction sets align better with human preferences.
PaperID: 739, Poster
Authors: Anirban Nath, YoonHaeng Hur, Genevera Allen
Abstract: Clustering is a central tool for discovering latent structure in unlabeled data, yet modern clustering pipelines often end with a hard assignment of each observation to a cluster without rigorous measures of assignment uncertainty. We propose a novel weighted conformal approach for constructing confidence sets for cluster labels. The key difficulty is that the labels available for calibration are not observed ground-truth labels, but synthetic labels produced by a data-dependent clustering algorithm. Our method develops a conformal inference algorithm that corrects the resulting mismatch with the latent target labels through weights by formulating conformal clustering as a conditional label-distribution shift problem. We first derive an oracle procedure that attains finite-sample marginal coverage and then develop a computationally tractable and implementable version using estimated conditional label probabilities and augmented calibration. We show that the coverage of the estimated-weight procedure depends on the estimator, giving an explicit bound on the loss relative to the nominal level. Empirical studies beyond the mixture-model settings, explored by the recently developed stochastic split conformal clustering procedure, demonstrate that the proposed weighted approach offers improvements in non-linear and high-dimensional clustering applications in terms of set size.
PaperID: 740, Poster
Authors:
Dongkyu D Cho, Pan XuAbstract: Proximal Policy Optimization (PPO) stabilizes policy updates by clipping the importance ratio between the new policy and the data-collecting policy. The standard rule uses one fixed clipping range for all samples. While this choice offers simplicity and effectiveness, it fails to distinguish statistically harmful extreme ratio values from large, yet tolerable changes. We study how to adapt the clipping radius to the reliability of the induced importance ratio distribution. Our case study demonstrates that reliable updates require controlling these extreme ratios, rather than merely verifying that individual ratios fall within a fixed interval. Motivated by this observation, we propose Ordered Policy Optimization (\textscOPO). In each minibatch, \textscOPO sorts the importance ratios, selects the extreme ratios that potentially have the largest effect on the update, and assigns them adaptive sample-wise radii instead of using the fixed clipping radius. These radii are derived from a prescribed tail profile, which specifies desired target values for the selected extreme ratios. The resulting objective remains first-order and requires only minibatch sorting and a modified policy loss. Experiments conducted on MuJoCo and Atari benchmarks demonstrate that this simple modification achieves strong performance relative to established baselines.
Abstract: We study internalization, the process by which neural-network-based systems absorb an explicit computational procedure into their own weights, and how it facilitates learning. We investigate how transformers internalize the simulation of semiautomata by internalizing chain-of-thought (CoT) tokens, which classes of semiautomata are harder to internalize, and expose the flip side of internalization, that is, a progressive degradation of out-of-distribution performance. We then provide the first provable analysis of successful internalization: for the task of learning parities, we show that a simplified one-layer transformer provably first learns the target with explicit CoT supervision and then internalizes the autoregressive generation as CoT tokens are progressively removed, directly computing the parity. Learning such a representation directly from data without CoT supervision is computationally hard. Finally, we discuss how learning through internalization can be viewed as an instance of the Positive Distribution Shift phenomenon recently introduced by~\citetMed+26.
Abstract: Many value-based deep reinforcement learning algorithms rely on target networks - lagged copies of the online network - to stabilize training. While effective, this mechanism introduces a fundamental stability-recency tradeoff: slower target updates improve stability but reduce the recency of learning signals, hindering convergence speed. We propose Target-Aligned Reinforcement Learning (TARL), a simple drop-in refinement for existing algorithms that emphasizes transitions for which the target and online network estimates are highly aligned. By focusing updates on well-aligned targets, TARL mitigates the adverse effects of stale target estimates while retaining the stabilizing benefits of target networks. We empirically demonstrate consistent improvements within discrete and continuous control algorithms across various benchmark environments without any hyperparameter tuning, including a 38.18% peak score gain on Atari-10, while incurring less than a 4% increase in wall-clock time.
PaperID: 743, Poster
Authors: Leo Elmecker-Plakolm, Matthew R Wicker
Abstract: We investigate the use of formal methods to provide tight and sound generalization bounds for learning algorithms. By casting the traditional notion of algorithmic stability as a specification to be verified, we demonstrate that recent advances in reachability analysis can yield provable bounds on the generalization of a given model and algorithm on a sample dataset. As sample-specific algorithmic stability is insufficient to bound the usual distributional notion of generalization, we develop a novel concentration inequality to connect the sample-specific results of formal certification algorithms to the required distributional analysis.The resulting framework enables the theoretical analysis of prior generalization bounds to extend far beyond their original restrictive analytical assumptions while achieving a provably sound bound on the generalization gap. In practice, we demonstrate that our framework provide formal generalization guarantees that are orders of magnitude tighter than alternative computational approaches at scales ranging from toy datasets to fine-tuning of modern large language models. While we implement certification-enhanced versions of several well-known stability results, future extensions of our approach will enable tighter bounds and enhanced practical adoption across the spectrum of modern generalization bounds.
PaperID: 744, Poster
Authors: Amanda Merkley
Abstract: Motivated by interpretable generative models for applications such as scientific inquiry, we aim to learn representations that localize human-interpretable concepts and separate them from residual variation in the data. Reconstruction and prediction objectives alone are insufficient for this purpose, as concept-relevant variation can be distributed across multiple latent factors while still achieving high performance. We address this limitation with a generative model that learns a balanced decomposition of data into concept-aligned and residual latent representations inferred purely from observations. The objective encourages localization through an upper bound on the concept information bottleneck while reducing dependence between the concept and residual latent. Across synthetic and real datasets, the learned representations localize concepts while preserving generative fidelity. A neural-data case study shows that localized representations align more strongly with structured variation in observed neural responses. This supports localization, rather than prediction accuracy alone, as a more faithful criterion for interpreting concept-relevant structure in generative models.
Abstract: We present a case study where an automatic AI system formalizes a textbook with more than 500 pages of graduate-level algebraic combinatorics to Lean. The resulting formalization represents a new milestone in textbook formalization scale and proficiency, moving from early results in undergraduate topology and restructuring of existing library content to a full standalone formalization of a graduate textbook. The formalization comprises 130K lines of code and 5900 Lean declarations and was conducted within one week by a total of 30K Claude 4.5 Opus agents collaborating in parallel on a shared code base via version control, simultaneously setting a record in multi-agent software engineering with usable results. The inference cost matches or undercuts what we estimate as the salaries required for a team of human experts, and we expect there is still the potential for large efficiencies to be made without the need for better models. We make our code and the resulting Lean code base available open-source.
Abstract: How should future neural reasoning systems implement extended computation? Recursive Reasoning Models (RRMs) offer a promising alternative to autoregressive sequence extension by performing iterative latent-state refinement with shared transition functions. Yet existing RRMs are largely deterministic, following a single latent trajectory and converging to a single prediction. We introduce Generative Recursive reAsoning Models (GRAM), a framework that turns recursive latent reasoning into probabilistic multi-trajectory computation. GRAM models reasoning as a stochastic latent trajectory, enabling multiple hypotheses, alternative solution strategies, and inference-time scaling through both recursive depth and parallel trajectory sampling. This yields a latent-variable generative model supporting conditional reasoning via p_\theta(y \mid x) and, with fixed or absent inputs, unconditional generation via p_\theta(x). Trained with amortized variational inference, GRAM improves over deterministic recurrent and recursive baselines on structured reasoning and multi-solution constraint satisfaction tasks, while demonstrating an unconditional generation capability.
PaperID: 747, Poster
Abstract: Test-Time Prompt Tuning (TPT) adapts large pretrained models to unseen target distribution shifts by updating only lightweight prompt tokens. Existing TPT methods are effective, but they mainly optimize each prompt update locally. Across test-time steps, prompts updated at the current step are reused for later predictions and further updates, so noisy or biased unlabeled objectives can accumulate and destabilize the prompt-update trajectory. We identify two failure modes: magnitude expansion, where updates move the prompt too far, and directional drift, where consecutive updates point in inconsistent directions. To analyze and control these failures, we propose Test-Time Prompt-Agnostic Decomposition (\mathttTPD), which characterizes observed prompt-update trajectories from a dynamical-systems view and decomposes base prompt updates into magnitude and direction. \mathttTPD computes a spectral radius to measure update expansion and decomposes update directions into persistent, oscillatory, and residual components. Building on this decomposition, Adaptive Koopman Control (\mathttAKC) regulates update magnitude by shrinking updates under expansive recent dynamics, while Hankel Update Router (\mathttHUR) refines update direction by preserving persistent components and suppressing oscillatory and residual components. Together, \mathttAKC and \mathttHUR produce a stabilized prompt update. As a prompt-agnostic framework, \mathttTPD can be plugged into visual and text TPT methods without modifying the backbone, prompt architecture, or adaptation loss. Experiments on 15 datasets show that \mathttTPD consistently improves 10 baselines, with accuracy gains of 3-8% for visual prompts and over 2% for text prompts.
Abstract: Test-time compute has emerged as a promising strategy to enhance the reasoning abilities of large language models (LLMs). However, this strategy has in turn increased how much users pay cloud-based providers offering LLM-as-a-service, since providers charge users for the amount of test-time compute they use to generate an output. In our work, we show that the market of LLM-as-a-service is socially inefficient–providers have a financial incentive to increase the amount of test-time compute, even if this increase contributes little to the quality of the outputs. To address this inefficiency, we introduce a reverse second-price auction mechanism where providers bid their offered price and (expected) quality for the opportunity to serve a user, and users pay proportionally to the marginal value generated by the winning provider relative to the second-highest bidder. To illustrate and complement our theoretical results, we conduct experiments with multiple instruct models from the \textttLlama and \textttQwen families, as well as reasoning models distilled from \textttDeepSeek-R1, on benchmark datasets covering mathematics, science, and open-ended question answering.
Authors: Mehrdad Moghimi, Bernardo Avila Pires
Abstract: Constrained MDPs (CMDPs) are a widely adopted framework for incorporating safety into RL agents, however, the framework does not support risk-sensitive constraints. This can be problematic: For example, CMDPs allow for optimal solutions that, in order to satisfy the risk-neutral constraints, mix infrequent catastrophic behaviors and frequent, overly conservative ones. Moreover, empirical results in multiple previous works suggest that enforcing stricter, risk-sensitive constraints can improve agent performance even when measuring it in a risk-neutral way. In this work, we introduce a simple yet powerful methodology for constrained RL, consisting of an extension of CMDPs that removes the limitation of risk-neutral constraints, and an algorithm for solving the resulting constrained problem. As a convenient side-effect of our framework, it is not necessary to fix constraint limits in advance of training the agent, provided that a sensible range is known. This increases policy flexibility and, in practice, allows for adjustments to these limits at no extra training cost. Besides benefiting from the generality of the framework, our agent shows strong performance in practice, consistently matching or outperforming existing baselines in several Safety Gymnasium benchmark tasks.
Abstract: Algorithmic replicability has recently been introduced to address the need for reproducible experiments in machine learning. A replicable online learning algorithm is one that takes the same sequence of decisions across different executions in the same environment, with high probability. We initiate the study of algorithmic replicability in constrained MAB problems, where a learner interacts with an unknown stochastic environment for T rounds, seeking to maximize reward while satisfying multiple constraints. Our main result is that replicability can be achieved in constrained MABs. Specifically, we design replicable algorithms whose regret and constraint violation match those of non-replicable ones in terms of T. As a key step, we develop the first replicable UCB-like algorithm for unconstrained MABs, showing that algorithms that employ the optimism in-the-face-of-uncertainty principle can be replicable, a result that we believe is of independent interest.
Abstract: We propose kernel-gradient drifting, a one-step generative modeling framework that replaces the fixed Euclidean displacement direction in drifting models with directions induced by the kernel itself. Standard drifting is attractive because it enables fast, high-quality generation without distilling a large pretrained diffusion model, but its theory is currently understood mainly for Gaussian kernels, where the drift coincides with smoothed score matching and is identifiable. Our gradient-based reformulation exposes this score-based structure for general kernels: the resulting drift is the score difference between kernel-smoothed data and model distributions, yielding identifiability for characteristic kernels and a smoothed-KL descent interpretation of the drifting dynamics. Since kernel gradients are intrinsic tangent vectors, the same construction extends naturally to Riemannian manifolds and to discrete data via the Fisher-Rao geometry of the probability simplex. Across spherical geospatial data, promoter DNA and molecule generation, kernel-gradient drifting enables state-of-the-art one-step generation beyond the Euclidean setting without distillation.
Authors: Alexander Smola
Abstract: Evaluating large language models across many benchmarks is expensive, yet many benchmarks are highly correlated. We formalize the selection of a small, informative subset as submodular maximization under a multivariate Gaussian model. Entropy (log-determinant covariance) and mutual information between selected and remaining benchmarks arise as natural objectives. Both are submodular; entropy selection coincides with pivoted Cholesky and has spectral residual bounds, while mutual information is non-monotone in general but empirically monotone for small subsets, so we optimize it greedily. Experiments on three matrices from ten public leaderboards show that mutual information selection outperforms entropy for imputation at small subsets.
PaperID: 753, Poster
Authors: George Christodoulou, Vasilis Christoforidis, Alkmini Sgouritsa, Ioannis Vlachos
Abstract: The inefficiency of decentralized resource allocation, a core challenge in Algorithmic Game Theory, is classically measured by the Price of Anarchy (PoA) and the Price of Stability (PoS). The most widely studied and foundational class of these models are atomic and non-atomic congestion games, for which a robust theory quantifying the inefficiency of equilibria is well-established. To mitigate the effects of this selfish behavior, we employ Coordination Mechanisms, which modify resource costs to incentivize socially improved equilibrium outcomes. In their standard form, Coordination Mechanisms have been limited by a pessimistic, worst-case view; it is assumed that the demand is unknown and adversarially chosen. Consequently, positive results remain sparse, applying mainly to specialized network topologies like parallel links. We address this limitation by studying the design and analysis of learning-augmented coordination mechanisms, where the mechanism is endowed with a potentially inaccurate prediction of the demand \barr. We study general atomic and non-atomic congestion games with polynomial latency functions of degree d. Our main positive result is a learning-augmented coordination mechanism tuned by a confidence parameter \beta which achieves a "best of both worlds" type of result: it achieves approximately optimal performance for an accurate prediction without sacrificing the worst-case guarantee of a good equilibrium (PoS). We further establish bounds on a stronger robustness measure, the PoA-Robustness (worst equilibrium performance under cost modification). The mechanism's construction relies on transforming approximate equilibria of the original game into exact equilibria of the modified game using a carefully defined approximate potential function. We complement our positive results by providing impossibility results. We first show that a stronger notion of consistency, requiring that all equilibria have optimal social cost for an accurate prediction, fails to guarantee bounded robustness. We also show a strong error tolerance limit: any mechanism achieving optimal consistency in non-atomic games suffers a sharp, discontinuous performance that collapses to the original PoA value under even infinitesimally small prediction errors, highlighting the limits of efficiency guarantees in non-atomic congestion games.
PaperID: 754, Poster
Authors:
Tamir Shor, Or Litany, Alex M BronsteinAbstract: Implicit neural representations (INRs) are shaped by the spectral structure induced by their input encodings and activation functions. Existing methods improve fitting primarily by modifying which frequencies are available to the network, through coordinate encodings or periodic nonlinearities. However, frequency access is not the only bottleneck: signals with localized or spatially varying structure require the network to efficiently compose frequencies into multi-harmonic internal responses. We introduce learnable spectral activations (LSA), which replace fixed neuron-level nonlinearities with a residual truncated Fourier series whose harmonic amplitudes are learned during training. LSA does not expand the asymptotic function class. Instead, it changes the factorization of the representation: linear weights select features while activation coefficients control spectral shaping, and the two are updated by separate gradients. Because the activation coefficients enter the loss linearly given fixed pre-activations, spectral tuning becomes a simpler subproblem than in architectures where it is entangled with feature selection. Empirically, this factorization concentrates more target-signal energy in the leading eigenmodes of the neural tangent kernel, consistent with improved optimization behavior. Across audio, image, neural radiance field, and neural acoustic field tasks, LSA also improves reconstruction quality.
Abstract: Single-task RL agents are typically trained under a fixed reward function, which limits their robustness to reward misspecification and their ability to adapt to changing preferences. We introduce Reward-Conditioned Reinforcement Learning (RCRL), an off-policy method that conditions agents on reward parameterizations while collecting experience under a single nominal objective. By recomputing counterfactual rewards from shared replay data, RCRL exposes the agent to multiple reward objectives without additional environment interaction, connecting single-task RL with ideas from multi-objective and multi-task learning. Across single-task, multi-task, and vision-based benchmarks, RCRL improves sample efficiency under the nominal reward parameterization, enables efficient adaptation to new parameterizations, and supports zero-shot behavioral adjustment at deployment. Our results show that RCRL provides a scalable mechanism for learning robust, steerable policies without sacrificing the simplicity of single-task training.
PaperID: 756, Poster
Abstract: We present Channel-wise Vector Quantization (CVQ), a novel image tokenization paradigm that replaces patch-wise tokens with channel-wise tokens. Unlike conventional vector quantization, which assigns a discrete token to each patch feature vector, CVQ quantizes each channel of the feature map. This formulation represents an image as discrete levels of visual details, rather than as a grid of spatial patches. Based on CVQ, we introduce a new visual autoregressive framework with "next-channel prediction". Instead of rendering images patch by patch in raster order, our Channel-wise Autoregressive (CAR) model predicts image channels sequentially, producing progressively enriched visual details. Specifically, it first sketches global structure and then refines fine-grained attributes, akin to a human artist's workflow. Empirically, we show that: (1) CVQ achieves 100% codebook utilization with a 16K+ codebook size without any bells and whistles, while reducing reconstruction FID by 50% over conventional VQ; and (2) CAR outperforms the AR baseline by improving the GenEval score from 0.69 to 0.74 and the DPG score from 79.86 to 82.14, demonstrating strong effectiveness for text-to-image generation. We hope our research offers a new perspective on the fundamental unit of visual tokenization by moving from spatial patches to channels.
PaperID: 757, Poster
Authors:
Zhenyu Wu, Min Li, Chong Ma, Xun Gong, Wei Wang, Chenglizhao Chen, Aimin Hao, Shuo LiAbstract: We introduce Panoptic Saliency Ranking (PSR), a new task that aims to estimate the visual saliency of all objects within a scene. Unlike salient object ranking, which focuses only on ranking salient regions, PSR extends the problem to the entire scene, providing a holistic and fine-grained understanding of visual saliency. To facilitate this task, we construct PSR18K, the first large-scale benchmark for panoptic saliency ranking. It contains 17,861 images with 146,460 annotated object instances spanning 13 superclasses and 376 fine-grained categories. In addition, PSR18K provides relation graph annotations to explicitly model semantic and spatial relationships among objects. We further propose an efficient Intrinsic Relation Graph Network (IRGNet), which reinterprets Transformer self-attention as an implicit relational matrix and seamlessly transforms it into an explicit relation graph, thereby obviating the need for additional relation modeling and significantly improving inference speed. Extensive experiments on PSR18K and three widely used SOR benchmarks demonstrate that our IRGNet achieves new state-of-the-art performance and faster inference with fewer parameters. The dataset and code are publicly available at https://sites.google.com/view/PSR18K.
Abstract: Inference-time scaling has become the dominant lever for improving language-model reasoning, but existing methods derive rollout diversity from a single source: stochastic token-level sampling. We argue that this single-axis sampling space is fundamentally limiting, and identify a second, fully deterministic and complementary axis: the layer span L at which a frozen model's top decoder layers are recursively re-applied at high-uncertainty tokens. Different choices of L produce distinct rollouts that solve different subsets of problems, with no stochasticity. We instantiate this axis through Entropy-Gated Latent Recursion (EGLR), a training-free decoding procedure that re-applies the top-L layers for at most K_max iterations until the next-token distribution converges. Combined with T temperature samples, EGLR turns a single-axis stochastic rollout pool into an L×T Cartesian sampling space at almost the same per-rollout cost. We characterize this space across 8 instruction-tuned models and 6 math reasoning benchmarks, and show that the L-axis is genuinely complementary to temperature: on MATH-500 with Qwen2.5-3B-Instruct, the joint L×T oracle reaches 91.6%, +8.2 percentage points beyond the temperature-only oracle (83.4%) and +10.4 points beyond the layer-only oracle (81.2%), confirming that the two axes capture genuinely complementary problems. The expanded rollout pool provides richer per-prompt candidates for any downstream procedure that consumes rollouts, including self-consistency, best-of-N with verifiers, and group-relative RL training (GRPO), opening a new direction for inference-time scaling that does not rely on stochastic noise. The code is available at https://anonymous.4open.science/r/EGLR/.
Abstract: RLVR and OPD have become standard paradigms for post-training. We provide a unified analysis of these two paradigms in consolidating multiple expert capabilities into a single model, identifying capability loss in different ways: mixed RLVR suffers from inter-capability divergence cost, while the pipeline of first training experts and then performing OPD, though avoiding divergence, fails to fully absorb teacher capabilities due to large behavioral pattern gaps between teacher and student. We propose , which encourages parallel training of experts and introduces OPD during each expert's ongoing RLVR training rather than after complete expert training, with experts serving as mutual teachers (making OPD bidirectional) to co-evolve. This enables more consistent behavioral patterns among experts while maintaining sufficient complementary knowledge throughout. Experiments validate that CoPD achieves all-in-one integration of text, image, and video reasoning capabilities, outperforming strong baselines such as mixed RLVR and MOPD, and even surpassing domain-specific experts.
PaperID: 760, Poster
Abstract: Bounds based on Integral Probability Metrics (IPMs) are ubiquitous in machine learning and statistics, appearing, e.g., in generalization, alignment, fairness, and privacy. IPMs reduce the discrepancy between two distributions to a single worst-case expectation gap. This compression is costly: for rare events, safety failures, or under-represented groups, IPM-based bounds are loose or vacuous. We introduce a non-linear generalization of IPMs, which we term the Integral Probability Boundaries (IPBs). Unlike IPMs, IPBs provide bounds that are tight at the extremes. We derive basic properties of IPBs, and provide algorithms for estimating them in different practically relevant settings. We demonstrate the usage of IPBs in both theoretical and practical applications for alignment, judge-free performance estimation, and fairness, where they enable us to obtain sharp bounds on performance and fairness gaps, e.g., 15p.p. more precise estimates on fairness gaps for rare groups in a realistic generative modeling task, compared to state-of-the-art.
Abstract: We introduce the Lattice Deduction Transformer (LDT), a recurrent transformer that approximates logically sound deduction by projecting its latent state through a lattice between forward passes. We train on-policy in a process that mirrors deduction in a search-based constraint solver and supervise training via a domain-agnostic, abstract-interpretation-based approximation of the set of solution candidates. An 800K-parameter LDT achieves 100% accuracy on Sudoku-Extreme and Snowflake Sudoku, at a fraction of the training cost of prior small recurrent reasoners, while remaining empirically sound: the model returns a correct answer or abstains. A 1.8M-parameter variant reaches 99.9% accuracy on Maze-Hard. Frontier LLMs score 0% on all three benchmarks.
Authors: Ananya Parashar, Derek Long, Dwaipayan Saha, Krzysztof M Choromanski
Abstract: We present a new paradigm for creating random features to approximate bi-variate functions (in particular, kernels) defined on general manifolds. This new mechanism of Manifold Random Features (MRFs) leverages discretization of the manifold and the recently introduced technique of Graph Random Features (GRFs) to learn continuous fields on manifolds. Those fields are used to find continuous approximation mechanisms that otherwise, in general scenarios, cannot be derived analytically. MRFs provide positive and bounded features, a key property for accurate, low-variance approximation. We show deep asymptotic connection between GRFs, defined on discrete graph objects, and continuous random features used for regular kernels. As a by-product of our method, we re-discover recently introduced mechanism of Gaussian kernel approximation applied in particular to improve linear-attention Transformers, considering simple random walks on graphs and by-passing original complex mathematical computations. We complement our algorithm with a rigorous theoretical analysis and verify in thorough experimental studies.
PaperID: 763, Poster
Abstract: Selective classification (SC) improves model reliability by allowing a classifier to abstain on low-confidence inputs. In the practical post-hoc setting, where the base classifier is fixed and often accessed as a black box, many confidence scores have been proposed, yet no single score is consistently best across models, datasets, and deployment shifts. We address this limitation by proposing, to our knowledge, the first ensemble framework for post-hoc SC, which learns to combine a diverse set of existing confidence scores using a small labeled calibration set drawn from the deployment distribution. We cast this problem as learning to separate correct from incorrect predictions of the base classifier, and train the ensemble selector with SoftRank-AURC, a differentiable soft-rank plug-in objective for the area under the risk-coverage curve (AURC), the standard evaluation metric for SC. We also study hinge and SELE losses as surrogate training variants, and provide an oracle-style guarantee showing that the SoftRank-AURC ensemble is competitive with the best single transformed base score in hindsight. Across 15 image and NLP classification tasks, with additional LLM multiple-choice experiments, the proposed method is consistently strong and often outperforms the best individual score, particularly under distribution shift, establishing ensemble post-hoc SC as a simple and effective way to improve reliability without retraining the base classifier.
PaperID: 764, Poster
Authors:
Wanyi Ling, Sida Li, Junming Guan, Nikolaos IgnatiadisAbstract: We study methods for simultaneous analysis of many noisy and biased estimates, each paired with an even noisier estimate of its own bias. The analyst's goal is to construct short calibrated intervals for each parameter. The standard debiasing approach, which subtracts the bias estimate from each biased estimate, inflates variance and yields long intervals. In this paper, we propose an empirical Bayes rebiasing strategy that starts from the fully debiased estimates and learns from data how much bias to reintroduce by estimating the unknown bias distribution. We provide convergence rates for the coverage of our intervals when the bias distribution is estimated using nonparametric maximum likelihood. Furthermore, we demonstrate substantial precision gains in prediction-powered inference, including pairwise LLM win-rate evaluations, as well as for inference of direct genetic effects in family-based GWAS.
Abstract: Performative prediction studies feedback loops that arise when predictive models are deployed in consequential domains. In these settings, model updates can shape the population whose patterns they aim to predict, inducing a distribution shift that is to the learning system. This perspective departs from classical treatments of distribution shift, where shifts are typically modeled as changes in the data-generating process. Yet, in practice, distribution shift is rarely one or the other. Predictive models may influence future data through the decisions they support, while the world itself continues to drift for reasons beyond the learner’s control. We study , a framework that captures both endogenous and exogenous sources of distribution shift. The framework generalizes performative prediction by allowing the data distribution to evolve both in response to the deployed model and according to an external, time-varying process. We extend the central notions of performative stability and performative optimality to this setting by defining online analogues that measure performance relative to the evolving partially performative environment. We analyze practical learning heuristics, including repeated retraining, and characterize when they successfully adapt to partially performative environments.
PaperID: 766, Poster
Abstract: Properly tuned hyperparameters are critical for reinforcement learning algorithms to perform well, limiting their use in real world applications. AutoRL aims to to address this by nesting the reinforcement learning problem within an outer optimization loop, where a black-box optimizer tunes the reinforcement learning algorithm's hyperparameters leaving the AutoRL algorithm's hyperparameters (hyper-hyperparameters) fixed. We provide the first empirical study of an AutoRL algorithm's sensitivity with respect to its hyper-hyperparameters. Our results demonstrate that SMAC3 with Hyperband, used to tune PPO, is sensitive to its hyper-hyperparameters, albeit less so than PPO is to its hyperparameters, and this reduction comes at a significant computational cost. The per-environment tuned performance gains do not outperform random search, suggesting the reduction in sensitivity may not justify the substantial compute and engineering overhead AutoRL imposes.
Abstract: Diffusion bridges have emerged as versatile diffusion-based frameworks for facilitating transformations between arbitrary data distributions. However, these approaches are hindered by the immensity of the optimization space without consideration of underlying high-order structures. This paper explores a complex-valued variant of diffusion bridges for solving physical problems using conformal field theory and mean field games. To this end, we propose a new class of dynamical diffusion bridges, called complex Schr\"odinger Bridges (CSB), whose derived algorithms are able to handle diffusion-based tasks using Hamiltonian-based geometric prior. Additionally, we propose image-to-image CSB (I^2CSB), tractable diffusion bridge techniques with a spiral analytic posterior. We verify our theoretical claims and practicality of our Hamiltonian based approach with a wide range of experiments.
PaperID: 768, Poster
Authors: Yujing Chen
Abstract: We study Stackelberg learning in which followers use lower-tail CVaR as a reward utility. In cooperative bandits, where both agents share the CVaR objective, we prove a spectral-risk lower bound that gives \Omega(\sqrtK(AB-1)/\tau) for CVaR and give a Stackelberg-factorized CVaR-UCB algorithm with matching \widetilde O\sqrtABK/\tau) regret up to logarithmic factors. Thus CVaR learning admits a minimax-optimal realizable response certificate in bandits. In Markov games, a risk-neutral leader faces a CVaR-sensitive follower. We introduce a certified follower-response oracle based on shortfall planning and decompose true Stackelberg regret into certified oracle regret and benchmark response certification error. The first term is sublinear without coverage, while true regret requires comparison coverage and response compatibility.
Abstract: Communication is fundamental to sustaining reciprocity and cooperation in strategic interactions. We identify and formulate the as the central optimization difficulty inherent in such dynamics for a learning agent: any action or signal the agent emits reshapes the reputations of many third parties along combinatorially branching paths before feeding back into its own future rewards, forcing the agent to account for all of these indirect channels at once when choosing every action. To address this, we introduce the , which explicitly backpropagates reward gradients through private estimators of opponents' policies trained from public observations. The gradient flows through the reputation chain itself analytically, rather than being estimated from sampled returns. It jointly optimizes actions and evaluative signals without intrinsic rewards or reward shaping. Empirically, the method recovers near-optimal context-sensitive policies, while sample-based baselines collapse into constant-output policies.
Authors:
Jiaru Zou, Rui Pan, Ruizhong Qiu, Pan Lu, Shizhe Diao, Jindong Jiang, Hanghang Tong, Tong Zhang, Markus Buehler, Jingrui He, James ZouAbstract: Recursive or looped language models have recently emerged as a new scaling axis by iteratively refining the same model computation over latent states to deepen reasoning. We extend such scaling principles from a single model to multi-agent systems, and ask: Can agent collaboration itself be scaled through recursion? To this end, we introduce RecursiveMAS, a recursive multi-agent framework that casts the entire system as a unified latent-space recursive computation. RecursiveMAS connects heterogeneous agents as a collaboration loop through the lightweight RecursiveLink module, enabling in-distribution latent thoughts generation and cross-agent latent state transfer. To optimize our framework, we develop an inner-outer loop learning algorithm for iterative whole-system co-optimization through shared gradient-based credit assignment across recursion rounds. Theoretical analyses of runtime complexity and learning dynamics establish that RecursiveMAS is more efficient than standard text-based MAS and maintains stable gradients during recursive training. Empirically, we instantiate RecursiveMAS under 4 representative agent collaboration patterns and evaluate across 9 benchmarks spanning mathematics, science, medicine, search, and code generation. In comparison with advanced single/multi agents and recursive computation baselines, RecursiveMAS consistently delivers an average accuracy improvement of 8.3%, 1.2×-2.4× end-to-end inference speedup, and 34.6%-75.6% token usage reduction.
PaperID: 771, Poster
Abstract: Coresets are widely used to compress large datasets into small weighted summaries. We mainly focus on the k-means objective where the cost is defined to be the sum of squared distances of every point to its closest center. A coreset approximately preserves the cost of every candidate set of k centers up to a small (1\pm \varepsilon) multiplicative error. By far the most flexible and powerful technique in this line of work is sensitivity sampling, where we pick points proportionate to their highest relative cost contribution in any solution. In general, the transmission and storage of these coresets are assumed to be reliable and always keep the coreset intact, without accounting for the possibility of corruptions. This raises a basic question: To what extent can coresets be made fault-tolerant? We study an adversarial fault model in which up to f points of the coreset may be arbitrarily corrupted, including their weights, and the goal is to recover a valid coreset from the corrupted summary. Our main result shows that sensitivity sampling yields fault-tolerant coresets of size O(f\cdot k/\varepsilon + m), where m is the size of a non-fault-tolerant coreset computed via sensitivity sampling. This result is also optimal in that any fault-tolerant k-means coreset must have size \Omega(f \cdot k/\varepsilon). The decoding algorithm computing a sanitized coreset from a corrupted one is efficient and we demonstrate practical viability via our experiments.
Authors: Leonardo Coregliano, William Opich
Abstract: Recently, a series of works have started studying variations of concepts from learning theory for product spaces, which can be collected under the name high-arity learning theory. In this work, we consider a high-arity variant of sample compression schemes and we prove that the existence of a high-arity sample compression scheme of non-trivial quality implies high-arity PAC learnability.
PaperID: 773, Poster
Authors: Louis Robinson
Abstract: We study \emphCausal Edit Serialization (CES), a decoder-only transformer extension that emits explicit edit programs while preserving a standard causal mask. CES supports \emphinsert, \emphdelete, \emphreplace, \emphmove, and \emphcopy operations through a serialized edit domain-specific language (DSL), a lightweight localization head, and dual-index rotary position embeddings that separate generation order from source-token addressing. We train on both self-supervised random corruptions and a large GPT-5-mini edit corpus built from ClimbMix documents by generating approximate human-like revisions and parsing them into the same edit DSL; copy examples are added from a separate repeated-span mining pipeline. We compare CES against standard autoregressive training and an autoregressive DSL baseline that emits textualized positions directly. Across model scales and data regimes, CES learns strong source-position localization, improves human-target similarity on WikiAtomic/CoEdIT-style source-target pairs, and produces higher valid-edit rates than the AR-DSL baseline, while incurring only a modest degradation in ordinary language-modeling quality in the best regimes. Arithmetic experiments further show that edit-aware training can improve standard autoregressive generation quality and validity across sampling temperatures, even relative to a clean-data AR baseline. These results suggest that explicit causal editing is a viable way to give decoder-only models compositional control over earlier tokens, both for revision and, in some settings, for generation itself.
PaperID: 774, Poster
Abstract: a prediction was made, a central challenge in modern AI. Yet existing methods lack a principled causal framework for explanation, leading to a disagreement problem: explanation methods produce conflicting attributions with no way to determine which is correct. Based on the literature in cognitive science, law, and philosophy, and using the language of causality, we formalize four desiderata for explanation methods that assign numerical credit to variables. We show that no prior explanation method including SHAP and LIME satisfies all desiderata, and prove the Explanatory Impossibility Theorem, which shows that even in principle, no method can satisfy all desiderata without causal assumptions. We introduce , and prove they satisfy all desiderata. We develop an algorithm for computing tight bounds on counterfactual Shapley values. Experiments corroborate our theory.
Authors:
Fangyuan Yu, Xin Su, Amir AbdullahAbstract: We investigate the temporal concatenation of sub-policies in Markov Decision Processes (MDP) with time-varying reward functions. We introduce General Dijkstra Search (GDS), and prove that it discovers optimal goal-reaching policies by concatenating sub-policies optimal for intermediate goals. Motivated by the ``search, select, update'' principle underlying GDS, we propose Dynamic Latent Routing (DLR), a language-model post-training method that jointly learns discrete latent codes, routing policies, and model parameters through dynamic search in a single training stage. In low-data fine-tuning settings, DLR matches or outperforms supervised fine-tuning across four datasets and six models, achieving a mean gain of +6.6 percentage points, while prior discrete-latent baselines consistently underperform SFT. Mechanistic analyses and targeted code ablations show that DLR learns structured routing behaviors with distinct causal roles. In six-digit arithmetic, a canonical mechanistic interpretability benchmark, DLR externalizes into discrete routing tokens algorithmic subtasks previously studied through activation and circuit analysis.
Abstract: Double Q-learning is a classical control algorithm that mitigates the maximization bias of Q-learning. To do so, it explicitly trains two independent action-value functions and uses them to decouple action-selection and action-evaluation when computing bootstrap targets. Double DQN adapts target bootstrap decoupling to deep reinforcement learning (RL), but explicitly trains only a single action-value function and does not fully decouple its estimators. Consequently, the two estimators remain correlated, and overestimation persists. In this paper, we introduce Deep Double Q-learning (DDQL), a deep RL algorithm that explicitly trains two Q-functions through Double Q-learning. DDQL stabilizes training through a combination of techniques, including lower replay ratios, longer target network update intervals, and shared layers. Across 57 Atari 2600 games, DDQL improves aggregate performance over Double DQN, outperforming it on 47 games while further reducing overestimation. In addition, we study key design choices when adapting Double Q-learning to deep RL, including the network architecture, replay ratio, and minibatch sampling strategies.
Abstract: Inspired by recent work on discrete diffusion for anchored unmasking, we introduce a new training-free framework for protein engineering that can generate higher-order variants that improve one or more phenotypes while exhibiting favorable epistatic interactions. Prior approaches require mutation sites to be specified in advance and can drive sequences away from the natural protein manifold. We address these limitations with ), a decoding-time method consisting of anchor selection, tilted decoding, and iterative refinement. APEX selects anchor residues using a reward-gradient score with an explicit correction for pairwise epistasis, decodes substitutions for anchors by tilting the base model's per-site predictive distributions along reward gradients, and revisits low-scoring anchors iteratively for optional test-time scaling. All phases query the base model only through masked-conditional logits, a primitive shared by masked language models, discrete-diffusion models, and inverse-folding networks; as a result, APEX works out of the box across commonly-used model families. Across various protein-engineering benchmarks, including thermostability and solubility, APEX achieves state-of-the-art results and avoids the predictor-gaming artifacts that reward-search-only baselines tend to produce.
Abstract: Flow matching (FM) learns vector fields by regressing stochastic velocity targets along intermediate distributions p_t. We identify a geometric optimization bottleneck in this regression problem: when the covariance \Sigma_t of p_t is ill-conditioned, gradient-based training rapidly fits high-variance directions while making slow progress along low-variance ones. In an exactly solvable Gaussian setting, we prove that the excess risk is weighted by \Sigma_t, and that both gradient descent and stochastic gradient descent inherit condition-number-dependent convergence. We then extend the analysis to Gaussian mixtures, showing that multimodality does not average away this effect; instead, the slowest and worst-conditioned component can control optimization. Motivated by this analysis, we propose \emphpreconditioned flow matching, a precondition-then-match framework that transforms the target distribution into a more isotropic representation, trains the main flow in the transformed space, and maps generated samples back through the inverse transformation. We show theoretically that preconditioning reshapes the intermediate FM path and improves its conditioning. Across controlled Gaussian and Gaussian-mixture experiments, latent MNIST and other high resolution image datasets up to 512×512 resolution, preconditioning improves path-conditioning diagnostics, low-eigenvalue recovery, FID, MMD, precision, and recall. Compute-matched baselines and preconditioner-quality ablations further show that the gains are not explained merely by additional preconditioner parameters, but by improved geometry of the downstream flow matching problem.
Abstract: Diffusion Language Models (DLMs) generate text by iteratively denoising masked token sequences, offering a tradeoff between parallelism and quality compared to autoregressive models. In current practice, the number of tokens decoded per step is controlled by a confidence threshold, and quality degrades monotonically as more tokens are denoised per step. We introduce Multi-token Residual Prediction (MRP), a lightweight module that enables dependency-aware multi-token denoising within a single backbone forward pass. MRP exploits a key property of the denoising process: the logit distributions at adjacent denoising steps are remarkably similar. Rather than running the backbone a second time to obtain the next-step logits, MRP predicts the residual between steps from the backbone's hidden states, effectively denoising more tokens per backbone forward at a fraction of the cost. We deploy MRP in two inference modes: direct decoding, which uses the corrected logits without verification for a tunable quality--speed tradeoff; and speculative decoding, which verifies MRP's proposals against the backbone for lossless acceleration. Experiments on SDAR models at the 1.7B, 4B, and 8B scales across reasoning and code generation benchmarks demonstrate up to 1.42× lossless speedup in SGLang.
Authors: Andreas Makris, Paul Fearnhead, Christopher Nemeth
Abstract: Training-free conditional diffusion provides a flexible alternative to task-specific conditional model training, but existing samplers often allocate computation inefficiently: independent guided trajectories can vary widely in quality, and additional function evaluations along a single trajectory may not recover from poor early decisions. We propose , an annealed sequential Monte Carlo framework for training-free conditional sampling with diffusion priors. TGD targets tempered posterior distributions over the clean signal, using noisy diffusion states only as auxiliary variables for proposing reconstructions and propagating particles. Particles are reweighted by incremental likelihood ratios, resampled, and propagated across noise levels, concentrating computation on trajectories plausible under both the prior and observation. Under idealized exact-reconstruction assumptions, full TGD yields a consistent particle approximation to the posterior as the number of particles grows. For expensive reconstruction tasks, retains early particle exploration but prunes to a single high-likelihood trajectory partway through sampling. Experiments on a controlled two-dimensional inverse problem and image inverse problems show improved posterior approximation and favorable wall-clock speed-quality tradeoffs over independent multi-trajectory baselines.
PaperID: 781, Poster
Abstract: Activation steering provides a simple, training-free mechanism for controlling attributes of generative models (e.g., sentiment, style, helpfulness). However, standard approaches such as Difference-in-Means (DiM) apply a single input-independent steering vector across all activations, limiting expressivity and ignoring the local structure of the activation space. We propose Kernelized Activation Steering (KAS), a unifying framework that lifts activation steering into a reproducing kernel Hilbert space (RKHS). KAS formulates steering as an optimization problem expressed purely via kernel evaluations, yielding an implicit, activation-dependent steering score without constructing explicit feature maps. Unlike DiM, KAS induces locally adaptive steering: each activation is modified according to its relative position with respect to source and target reference sets, producing a nonlinear steering field over the representation space. Importantly, DiM is recovered as a special case under a linear kernel, offering a principled interpretation of steering, while richer kernels (e.g., RBF) enable geometry-aware interventions. Across standard activation steering tasks, including jailbreaking LLMs and image style control, KAS consistently outperforms existing methods.
PaperID: 782, Poster
Abstract: learning, where the constants in the learning rate may depend on the unknown marginal distribution generating the unlabeled data, but not on the target concept. This notion naturally lies between classical PAC learning, where the constants are independent of both the marginal distribution and the target concept, and more recent universal learning, where the constants may depend on both. We give a complete characterization of the optimal marginal-nonuniform learning rates, establishing a fundamental trichotomy: exponential, linear, and arbitrarily slow rates. This characterization is witnessed by a new combinatorial dimension, termed the
Abstract: We introduce Midpoint Generative Models (MGM), a principled framework for training one-step generative models. MGM is based on a simple symmetry of Flow Matching with linear interpolation: when the two endpoint distributions coincide, the corresponding drift field vanishes at the midpoint time, t=1/2. We show that the norm of this field defines a valid discrepancy between distributions, which we call the Midpoint Divergence. We extend this discrepancy beyond the midpoint by introducing randomly flipped interpolations and further generalize it by replacing deterministic linear Flow Matching interpolations with symmetric stochastic interpolants, yielding a generalized Midpoint Divergence. Finally, we derive a variational formulation of our generalized divergence, yielding a tractable objective for training a one-step generator. The resulting MGM algorithm offers an effective and theoretically grounded approach to generative modeling, achieving competitive performance against existing one-step generative modeling methods.
Abstract: Classical deep learning typically operates on individual cases. Despite its success, real-world usage often requires repeated inference to estimate statistical quantities for complex decision-making tasks involving uncertainty or extreme-value analysis, resulting in substantial latency. We introduce neural statistical functions, a new family of models learned from pre-trained single-sample predictors and scattered data samples, which can directly infer statistics over continuous operating condition ranges without explicit sampling. By introducing the notion of prefix statistics, we transform and unify diverse statistical functions (e.g., integrals, quantiles, and maxima) into an interval-conditional framework, in which a principled identity between the prefix statistics and the individual-case regression serves as the learning objective. Neural statistical functions achieve strong performance in estimating essential statistics of complex physical processes, including accumulated energy in dynamical systems, quantiles of aerodynamic responses, and maximum stress in crash processes, while achieving up to a 100× reduction in model evaluations.
PaperID: 785, Poster
Abstract: We introduce a notion of monoculture robustness in Probably Approximately Correct (PAC) learning, as a way to mitigate the problem of algorithmic monoculture (Kleinberg & Raghavan, 2021): the tendency of models sharing the same components (e.g., training data) to fail on the same inputs. Specifically, we ask whether a learner with access to a single dataset can output k individually accurate hypotheses whose joint failure probability is comparable to training on k independent datasets. We say that such a learning procedure achieves monoculture robustness, and we quantify this by bounding the conjunctive error that all k hypotheses misclassify points in an individual sense. We design methods for achieving this type of guarantee by drawing a surprising connection between replicability (Impagliazzo et al., 2022) and monoculture robustness. Specifically, we show that if the base learner is \rho-replicable, then running it independently k times on the same sample yields individual monoculture error that shrinks exponentially with k, controlled by the single-run error of the base learner at the point and \rho. We complement these positive results with a lower bound: for individual monoculture robustness, we prove a worst-case lower bound showing that, for learners, driving the individual conjunctive error down to \gamma can require a sample-size multiplier of t = \Omega(\log(1/\gamma)).
Authors:
Christopher Solinas, Radovan Haluška, David Sychrovský, Finbarr Timbers, Nolan Bard, Michael Buro, Martin Schmid, Nathan Sturtevant, Michael BowlingAbstract: Sequential estimation under partial observability requires tracking belief distributions that may be high-dimensional, multimodal, and non-Gaussian. Classical Bayesian filters generalize zero-shot to any system whose dynamics can be evaluated, but their representations scale poorly: parametric filters struggle to capture multimodality, and particle filters require exponentially many particles in the state dimension. Generative Distribution Embeddings (GDEs) learn compact representations of complex distributions but provide no mechanism for sequential updates. We present Neural Bayesian Filtering (NBF), a method that combines classical Bayesian filtering with learned distribution embeddings. NBF represents beliefs as embeddings and updates them by sampling particles from a learned conditional generator, propagating them through the system dynamics, and finally re-embedding the resulting particles. For in-family distributions that the embeddings are trained to represent, NBF inherits the zero-shot adaptability of classical filters, but with the benefit that regenerating particles from the embedding at each step also mitigates particle impoverishment. We validate NBF on a pursuit--evasion domain, demonstrating accurate tracking of complex multimodal posteriors, graceful scaling with state dimension, and zero-shot generalization to dynamics held out from training.
PaperID: 787, Poster
Abstract: Previous work has shown that alignment training on individual LLM agents can fail to transfer to multi-agent settings. We investigate the underlying mechanisms for this failure, and explore interventions to preserve alignment transfer to multi-agent LLM teams. We focus on two software-engineering tasks (sepsis triage, a news recommender) and ten consulting-proposal generation tasks. We present four findings: (1) The single-vs-team safety gap does not show a clear trend with more capable models. We find that Anthropic's Mythos Preview shows the largest gap on one of our tasks but the smallest on another; (2) Many natural interventions, such as having agents critique each other's work or coordinate in a group chat, do not bring agent teams back to single-agent alignment behavior; (3) substantially accounts for the safety gap: agents in teams are less likely to proactively check for system-level safety failures and more likely to ignore, rationalize, or deflect responsibility for issues when they arise; (4) having a , a multi-agent LLM orchestration and experimentation framework. Our findings can help practitioners design agent teams that successfully preserve the behaviors expected of aligned individual agents.
Abstract: Multi-Head Attention (MHA) is the core computational primitive underlying modern Large Language Models (LLMs). However, within a single attention layer, MHA has an intrinsic linear scaling limitation: H attention heads produce exactly H independent attention matrices, with no communication between heads during attention computation. While deep Transformers can compose information across layers, this places much of the burden of building and combining intermediate relations on depth, potentially requiring additional layers to expose interactions useful for downstream computation. Our focus is complementary: increasing the compositional bandwidth available within each attention layer. To this end, we propose Interleaved Head Attention (IHA), which enables within-layer cross-head mixing by constructing P pseudo-heads per head, typically with P=H, where each pseudo query, key, and value is a learned linear combination of all H original queries, keys, and values, respectively. Interactions between pseudo-query and pseudo-key heads induce up to P^2 attention patterns per head with modest parameter overhead \mathcalO(H^2P). We provide theory showing improved parameter efficiency on Polynomial Filters, where IHA uses \Theta(\sqrtkn^2) parameters versus \Theta(kn^2) for MHA, and on the order-sensitive CPM-3 task, where IHA uses \lceil\sqrtN_max\rceil heads versus N_max for MHA. On real-world benchmarks, IHA improves Multi-Key retrieval on RULER by 10–20% at 4k–16k context lengths and, after OpenThoughts reasoning fine-tuning, improves GSM8K by 5.8% and MATH-500 by 2.8% under majority voting over full attention, with modest throughput overhead.
PaperID: 789, Poster
Abstract: We establish universal approximation for neural networks whose weight matrices take values in matrix manifolds. The nonlinear matrix manifolds considered in this work arise as orbits of classical Lie group actions embedded in the ambient space \mathbbR^d × k. Stiefel manifolds provide a prominent example with empirical precedent in prior neural-network models. We present an expanded catalog of such manifolds, drawing from well-studied constructions in physics and differential geometry. These manifolds are particularly appealing for intrinsic Riemannian optimization: the transitivity of the underlying group action yields explicit tangent-space descriptions and retractions. Moreover, the group-action structure enables parameter-efficient factorizations, reducing the number of trainable variables. This leads to a fundamental expressivity question: when weight matrices take values in nonlinear matrix manifolds, which architectures retain universal approximation power? We prove that manifold-weighted networks with residual connections and bounded scalar multipliers are L^p(K)-universal. To establish universality across our manifold catalog, we introduce two independent sets of sufficient conditions: Register-Isolated Primitive Realizability (RIPR) and Cross-Axis Fold-and-Cut (CAFC). These conditions correspond to two distinct approximation constructions. By verifying RIPR or CAFC case by case for the matrix manifolds in our catalog, we establish universality for the corresponding residual manifold-weighted networks with bounded scalar multipliers.
PaperID: 790, Poster
Abstract: Learning expressive and transferable representations from unstructured 3D data remains a fundamental challenge. Existing backbones mirror the design of image-based networks by relying on symmetric U-Net architectures with large parametric decoders and feature skip connections. While successful in 2D, these designs are ill-suited for 3D domain where point coordinates, that define the underlying geometric manifold on which feature representations are learned, are sensitive to sensor-specific sampling patterns. In such architectures, skip connections allow low-level coordinate cues to bypass the semantic bottleneck, leading the model to overfit to local spatial patterns rather than learning robust, transferable semantic abstractions. We re-think this design and introduce Point Clustering Encoders (PCE), a minimal, decoder-free network that treats point cloud processing as a hierarchy of end-to-end learned spatial supertokens. At the core of PCE is our Time-Reversible Attention Pooling (TRAP) layer, which formalizes hierarchical downsampling as a time-reversible Markov chain. By coupling upsampling to the reverse chain, we eliminate the need for traditional decoders; the task of representation learning is fully delegated to the encoder. PCE not only surpasses state-of-the-art w.r.t. segmentation accuracy in indoor/outdoor datasets across a variety of tasks, but also reduces the number of trainable parameters, and can process scenes consisting of up to 6.5M points.
PaperID: 791, Poster
Abstract: When distinct pre-trained models are specialized for the same task, they often converge to nearly equivalent functional behaviors, while the underlying parameter-space changes share no common coordinate system. Existing approaches that compose such changes arithmetically are therefore confined to a single model, leaving task knowledge stranded inside the network that acquired it. Drawing on Plato's allegory of the cave, we hypothesize that these model-specific updates are shadows cast by a shared, model-agnostic object that governs how task specialization reshapes a model's behavior, and we refer to this object as the platonic task vector. To make this view operational, we introduce Universal Task Descriptors, matrices whose shape is fixed independently of architecture or embedding dimension and that capture a task's functional effect in a model-transferable form. Universal Task Descriptors admit addition, negation, and analogy in closed form, and a lightweight realization step then transfers any composed descriptor into a chosen target model. Experiments across a range of models and tasks indicate that this pipeline transfers and composes task knowledge across heterogeneous models.
Abstract: Diffusion and flow-based generative models dominate visual synthesis, with guidance aligning samples to user input and improving perceptual quality. However, Classifier-Free Guidance (CFG) and extrapolation-based methods are heuristic linear combinations of velocities/scores that ignore the generative manifold geometry, breaking probability conservation and driving samples off the learned manifold under strong guidance. We analyse guidance through the continuity equation and show its effect decomposes into a divergence term and a score-parallel term defined invariantly across parameterisations. We prove the divergence term blows up structurally as sampling approaches the data manifold, motivating a time-dependent schedule alongside score-parallel attenuation. The resulting plug-and-play rule, Adaptive Manifold Guidance (AdaMaG), bounds both terms at no additional inference cost. Finally, we show that most empirical heuristics for reducing saturation or improving generation quality correspond directly to the two terms in our decomposition. Across image generation benchmarks, AdaMaG improves realism, reduces hallucinations, and induces controlled desaturation in high-guidance regimes.
Abstract: Evaluating modern machine learning models often requires labels on large test pools, yet obtaining these labels can be expensive. Active testing reduces this cost by adaptively selecting which test points to label, but existing unbiased estimators do not fully exploit cheap black--box predictions that are often available for the entire pool. We introduce Prediction--Powered Active Testing (PPAT), a label--efficient risk estimation framework that combines the unbiased LURE estimator with a prediction--powered control variate. Rather than using proxy predictions as biased pseudo--labels, PPAT uses them to residualise the loss, preserving unbiasedness while reducing variance. This control--variate perspective also changes the optimal acquisition problem: we derive residualised oracle proposals and practical surrogate--based acquisition rules tailored to the PPAT estimator. We further establish asymptotic normality for LURE and PPAT, enabling asymptotically valid confidence intervals. Across tabular regression and image classification tasks, PPAT consistently improves over existing active testing baselines, remains unbiased, and reaches the desired coverage level with substantially fewer labels.
PaperID: 794, Poster
Abstract: We present Self-Cleaning Diffusion Models, a principled and effective framework for training diffusion models from heterogeneous data. In many domains, high-quality data is scarce or expensive to collect, while low-quality and out-of-distribution (OOD) data is abundant. Leveraging these ubiquitous samples is critical for scaling generative models. Still, principled techniques remain elusive, leaving practitioners to rely on simple heuristics such as high-quality finetuning or explicit quality conditioning. Self-Cleaning Diffusion Models bridges this gap by decoupling data correction from prior learning. We first train a transport map between the abundant OOD and the limited in-distribution samples. We then train an unconditional generative prior on a mixture of the few in-distribution samples and versions of the transported points. Crucially, this noise injection prevents learning errors in the transport map from propagating to the generative prior. Theoretically, we demonstrate that transporting OOD samples towards the empirical proxy reduces the distance to the underlying distribution under mild assumptions. Empirically, we achieve state-of-the-art results across a diverse range of settings, from mitigating synthetic corruptions to increasing diversity and transcending dataset quality.
PaperID: 795, Poster
Abstract: The design of attention mechanisms in long-context large language models faces an ``impossible triangle'' among KV cache size, training and prefill FLOPs, and model quality. Grouped-Query Attention (GQA) achieves strong quality but incurs high KV cache overhead; alternative approaches such as Multi-head Latent Attention (MLA) and Multi-matrix Factorization Attention (MFA) can compress the cache, but MLA suffers from quality loss while MFA substantially increases training and prefill compute. We propose Virtual Head Attention (VHA), which simultaneously improves all three vertices through two lightweight linear operations: Q Premix applies a near-identity transformation to queries within each KV group to recover virtual head diversity from halved physical query heads, and Linear Postmix fuses inter-head features via a low-rank I + AB^\top residual structure. Since VHA introduces no nonlinear operations, Premix can be folded into the query projection and Postmix into the output projection at inference time, reducing the model to a standard GQA-2 architecture that directly reuses all existing GQA inference optimizations. With only two KV heads, the group size in large models readily reaches 64, fully exploiting Tensor Core parallelism during decoding. Across four scales (0.6B, 1.7B, 4B dense, and 8B-A1B MoE) with 4K-context 50B-token pretraining, VHA consistently surpasses the standard GQA baseline at every scale (+0.05% to +0.81%), while reducing KV cache by 4× vs.\ GQA-8 at dense scales and 2× vs.\ GQA-4 at the MoE scale, and lowering training and prefill FLOPs by 7.9% (1.7B, S=4096). On an industrial 30B-A3B MoE model trained on 500B tokens at 8K context, VHA outperforms both MLA and GQA-4 by +0.70\,/\,+2.60\,% on a 12-benchmark base-model evaluation suite, demonstrating scalability to industrial model scale. The source code will be released on GitHub.
Abstract: We study the empirical risk minimization (ERM) of linear Joint Embedding Predictive Architecture (JEPA) and establish its key conditions for optimality. By casting the embedding bottleneck as a rank constraint, we formulate JEPA as a reduced-rank regression problem. This induces a risk decomposition into an irreducible error and an excess-risk bound. Within the irreducible term, we find the rank of the optimal predictor grows with target heterogeneity and saturates at the bottleneck. Within the excess-risk bound, we find spectral conditioning of the whitened end-to-end predictor governs minimization. Together, these conditions prescribe an isotropic regularizer. Under this regularization, we show such JEPA isotropy subsumes canonical JEPA isotropic embedding designs and hence justifies their empirical successes via ERM. Numerical experiments corroborate our theory.
Abstract: Decision tree optimization is fundamental to interpretable machine learning, yet most tree learning algorithms are restricted to binary splits. This restriction can produce unnecessarily deep trees when the underlying decision logic is naturally multi-valued. Multiway-split trees address this, but finding sparse yet accurate ones is more challenging, as the number of possible partitions at each node grows combinatorially with the number of feature bins. We introduce SPLINTER (SParse Lookahead for Interpretable N-way Trees by Eliminative Ranking), a dynamic programming and branch-and-bound-based framework that combines efficient candidate generation with pruning bounds to find sparse, accurate multiway-split trees. Our empirical results show that our methods produce trees that achieve higher accuracy at a greater decision sparsity (shorter path lengths) than greedy multiway and near-optimal binary-split tree baselines, while remaining practical to train. We further extend the framework to approximate the Rashomon set of near-optimal multiway-split trees, allowing users to inspect multiple sparse and accurate alternatives.
Abstract: Multimodal representation learning has been shifting from traditional two-tower architectures to large language model (LLM)-based embedders due to their strong instruction-following capabilities. Despite this progress, existing approaches primarily focus on language and image modalities, which also remain the dominant modalities for human-AI interaction in current embedders. In this paper, we propose the first Omni-Interactive Universal Embedder (OmniUE), which not only learns a unified embedding space across text, video, and audio by leveraging intermediate-layer representations from dedicated learnable tokens, but also supports omni-interactive querying, enabling users to provide inputs in the form of text, visual regions of interest, and audio spans. Within OmniUE, visual and audio segmenters process diverse user interactions and integrate them with an omni-LLM to produce user-conditioned any-to-any embeddings via context aggregation. To evaluate OmniUE’s omni-interactive capabilities, we introduce OmniCHOIR, benchmarking models for omni-interactive compositional audio retrieval based on the given text, video, and audio as well as unimodal or multimodal interaction prompts. OmniUE consistently surpasses state-of-the-art baselines across diverse modalities, with average improvements of 10.5% on textual-interactive video benchmarks (MMEB-v2-video), 1.1% on audio tasks (MAEB), 83.7% on visual-interactive benchmarks (SCaR), and 24.1% on our omni-interactive OmniCHOIR benchmark We believe that jointly advancing omni-modal representation learning and omni-interactive querying paves the way toward universal embedders.
Abstract: Diffusion models have shown remarkable success across a wide range of generative tasks. However, they often suffer from spatially inconsistent generation, arguably due to excessive correlations learned by the model. This can produce samples that are locally plausible but globally inconsistent. We propose a principled method to mitigate this issue. Our method, sparsely supervised diffusion (SSD), is a simple yet effective masking framework that can be implemented with only a few lines of code. We analytically show that SSD fundamentally alters diffusion model training by modifying the spectrum of the data covariance and that it suppresses correlations in the data covariance matrix. Experiments show that our method more accurately approximates the underlying population score function, reduces memorization on small datasets, and promotes the use of essential contextual information during generation. Moreover, even when up to 98% of pixels are masked, it achieves competitive FID scores across a range of datasets and, importantly, avoids the training instability commonly observed on small datasets.
PaperID: 800, Poster
Abstract: We present a framework for computing approximate mixed-strategy Nash equilibria of continuous-action games. It is a modification of the multiple oracle algorithm, which extends the double oracle algorithm to multiple players and continuous action spaces. Unlike prior methods, it maintains fixed-cardinality pure strategy sets for each player. Thus, unlike prior methods, only a constant amount of memory is necessary. Furthermore, it does not require exact metagame solving on each iteration, which can be computationally expensive for large metagames. Moreover, it does not require global best-response computation on each iteration, which can be computationally expensive or even intractable for high-dimensional action spaces and general games. Our method incrementally reduces the exploitability of the strategy profile in the finite metagame, pushing it toward Nash equilibrium. Simultaneously, it incrementally improves the pure strategies that best respond to this strategy profile in the full game. We evaluate our method on various continuous-action games, showing that it obtains approximate mixed-strategy Nash equilibria with low exploitability.
Abstract: Flow maps have enabled remarkable progress in few-step generative modeling across both continuous and discrete state spaces. Despite their promise, existing parameterizations are restricted to flows over fixed dimensions or fixed sequence lengths. Here, we introduce (EFMs), a new class of flow maps that distill EFlows into efficient few-step generative models. Each EFM factors the map between any two timesteps into two learned components: an , which pushes the expanded state forward along the interpolant. Composing these operators yields a single map that jointly the state, recovering existing fixed-canvas flow maps as the special case in which the expand operator is the identity. We further extend the framework to the discrete simplex, enabling variable-length sequence generation via token insertions. Across domains, EFlows and EFMs provide a principled approach to generative problems in which output size is itself learned, controllable degree of freedom.
PaperID: 802, Poster
Abstract: Reinforcement Learning with Verifiable Rewards empowers Large Language Models to enhance their reasoning capabilities. Most algorithms rely on on-policy training, which is stable but inefficient due to its strict dependency on fresh samples. To overcome this inefficiency, off-policy RL decouples experience generation from policy optimization via Importance Sampling (IS). However, this introduces a policy discrepancy between the behavior policy and the target policy. We observe that mainstream RLVR algorithms struggle to maintain a balance between performance and stability in off-policy scenarios, suffering from either severe performance degradation or training collapse. We identify the root cause as the failure of static trust regions to maintain the bias--variance trade-off associated with IS ratios in off-policy settings. To address this, we propose ), a novel hyperparameter-free RLVR algorithm employing an adaptive trust region. DIPO leverages the intrinsic statistics of the IS ratio distribution to achieve a robust bias--variance trade-off. Experimental results demonstrate that DIPO enables stable training in extreme off-policy scenarios, achieving performance comparable to or even surpassing the on-policy baseline, and outperforming other off-policy baselines by 13.9%.
Abstract: Iterative generative models such as Flow Matching and Diffusion models have demonstrated strong test-time scaling behavior, where additional inference computation can improve generation quality. In contrast, Drift Models offer efficient one-step generation, but their direct generation paradigm limits such flexibility. In this work, we propose Drift Flow Matching (DFM), a framework that connects drifting generative modeling with flow-based iterative generation. DFM preserves the efficiency of direct transport maps while enabling generation to be refined through multiple inference steps when desired. This bridges the gap between one-step Drift Models and multi-step Flow Matching methods, and provides a novel generative paradigm that can adapt sampling computation to different quality--efficiency requirements. Extensive experiments across different tasks and datasets demonstrate the effectiveness and generality of the proposed framework.
PaperID: 804, Poster
Abstract: Wasserstein distributionally robust optimization (DRO) is a prominent paradigm for robust decision-making in the face of distributional uncertainty. Given a nominal data distribution, DRO selects a decision which minimizes the risk for a worst-case data distribution within a prescribed Wasserstein radius \varepsilon of the nominal distribution. In this paper, we investigate the quality of DRO decisions when evaluated on a worst-case distribution within a 2-Wasserstein \varepsilon-neighborhood of the nominal distribution, as measured by excess risk or ex-ante regret. Beginning with multivariate linear regression and extending to ridge regression, we first identify the asymptotic complexity of robust decision-making in the \varepsilon \to 0 limit, characterized by a problem-specific condition number \kappa. In this regime, a wide spectrum of simple algorithms, including DRO, achieve the instance-optimal rate of O(\kappa \varepsilon^2). For fixed \varepsilon > 0, we prove that DRO still achieves the minimax rate if the problem is of an appropriate ``low rank''. On the other hand, we identify significant failure modes where DRO is provably suboptimal by dimension-dependent factors. To resolve this, we introduce a new condition number reduction (CNR) procedure which achieves the optimal rate and admits tractable approximation algorithms. We support these theoretical results with numerical experiments comparing the performance of various estimators including DRO and CNR.
PaperID: 805, Poster
Abstract: Asymptotic e-values are emerging as a powerful alternative to asymptotic p-values, particularly in post-hoc inference and multiple testing, where significance levels may be data-dependent. Existing asymptotic e-values, however, suffer from the ``missing factor,'' a scaling inefficiency resulting in overly conservative inference. Drawing on the framework of near-optimal concentration inequalities developed by Bentkus in the 2000s, we introduce Bentkus-type asymptotic e-values and prove that they successfully eliminate the missing factor. We also demonstrate both theoretically and empirically that Bentkus-type e-values consistently deliver sharper inference than existing alternatives, leading to tighter post-hoc confidence intervals and higher rejection rates in multiple testing procedures.
Abstract: AI agents are becoming active decision-makers on the Internet. As they make decisions in the same environments as humans, the environments themselves can change to influence them. We call this mecha-nudging: changes to how choices are presented that systematically influence AI agents without materially degrading the decision environment for humans. To measure this phenomenon, we combine two frameworks---Bayesian persuasion from economics and \mathcalV-usable information from computer science---to get a common unit (bits) for quantifying how environments change across a wide range of interventions, contexts, and models. We apply this framework to over six million Etsy listings and find that, after ChatGPT’s release, listings contain significantly more machine-usable information for predicting agent curation decisions, increasing by 0.143 bits out of a maximum possible increase of 0.355. This shift is robust across prompts, token choices, labeling models, and fine-tuning architectures; absent in a regulated-text placebo; and far larger than the effect of generic LLM rewriting. In contrast, a human study finds little to no change in human-usable information. Our results provide the first large-scale evidence that systematic mecha-nudging is already occurring in the wild.
Abstract: This paper introduces language-based agent control (LBAC), a method to apply techniques from language-based security to the control of AI agents. Unlike systems-level defenses such as I/O sandboxing, LBAC can express application-level policies. Moreover, in LBAC agents may perform computations and recursively invoke subagents with tool access. Language-based techniques enable the design of APIs which, through a combination of typing discipline and internal runtime checks, ensure that any well-typed program obeys a desired property. The core idea of LBAC is to have agents interact with the world by generating programs against such APIs rather than by issuing tool calls directly. These programs may themselves contain recursive calls to subagents, which retain full tool access; the type checker rejects ill-typed programs before execution, and policies thereby extend uniformly across the entire agentic system, including its scaffolding and control flow. We demonstrate LBAC with three case studies: I/O sandboxing via filesystem capabilities, data provenance, and information-flow control.
PaperID: 808, Poster
Abstract: We present the first mini-batch kernel k-means algorithm, offering an order of magnitude improvement in running time compared to the full batch algorithm. A single iteration of our algorithm takes \widetildeO(kb^2) time, significantly faster than the O(n^2) time required by the full batch kernel k-means, where n is the dataset size and b is the batch size. Extensive experiments demonstrate that our algorithm consistently achieves a 10-100x speedup with minimal loss in quality, addressing the slow runtime that has limited kernel k-means adoption in practice. We further complement these results with a theoretical analysis under an early stopping condition, proving that with a batch size of \widetilde\Omega(\max \set\gamma^4, \gamma^2\cdot k\epsilon^-2 ), the algorithm terminates in O(\gamma^2/\epsilon) iterations with high probability, where \gamma bounds the norm of points in feature space and \epsilon is a termination threshold. Our analysis holds for any reasonable center initialization, and when using k-means++ initialization, the algorithm achieves an approximation ratio of O(\log k) in expectation. For normalized kernels, such as Gaussian or Laplacian it holds that \gamma=1. Taking \epsilon = O(1) and b=\Theta(k\log n), the algorithm terminates in O(1) iterations, with each iteration running in \widetildeO(k^3) time.
Authors: Mykola Haltiuk, Aleksander Smywiński-Pohl
Abstract: Large Language Models (LLMs) are trained to support an increasing number of languages, yet their predefined tokenizers remain a bottleneck for adapting models to lower-resource or distinct-script languages. Existing tokenizer transfer methods typically rely on semantic heuristics to initialize new embeddings, ignoring higher-layer model dynamics and limiting transfer quality. We propose Model-Aware Tokenizer Transfer (MATT), a method that incorporates model internals into the tokenizer transfer process. MATT introduces an Attention Influence Modeling (AIM) objective that distills inter-token communication patterns from a source model into a target model with a new tokenizer, providing an efficient warm-up before standard language modeling. Unlike approaches that focus solely on embedding similarity, MATT leverages attention behavior to guide embedding initialization and adaptation. Experiments across diverse linguistic settings show that MATT recovers a large fraction of the original model’s performance within a few GPU hours, outperforming heuristic baselines. These results demonstrate that incorporating model-level signals offers a practical and effective path toward robust tokenizer transfer in multilingual LLMs.
Authors:
Jiayi Zhang, Yongfeng Gu, Jianhao Ruan, Maojia Song, Yiran Peng, Zhiguang Han, Jinyu Xiang, Zhitao Wang, Caiyin Yang, Bang Liu, Chenglin Wu, Yuyu LuoAbstract: Agentic evolution has emerged as a powerful paradigm for improving programs, workflows, and scientific solutions. It iteratively generates candidate artifacts, evaluates them against a task objective, and uses the resulting feedback to guide subsequent evolution. Existing methods typically instantiate this paradigm either through fixed hand-designed procedures that are modular but rigid, or through general-purpose agents that flexibly integrate feedback but can drift as context grows, leaving long-horizon search vulnerable to local optima. Both routes share a deeper limitation: long-horizon evolution accumulates candidates, feedback, traces, and failures over time, but lacks a stable interface for organizing this evidence and revising the mechanism that drives future evolution. We therefore formulate agentic evolution as an interactive environment, where the accumulated evolution context becomes process-level state. A meta-agent acts on this environment not by directly generating the next candidate, but by editing the mechanism that controls how future evolution proceeds. We introduce AEvo, a harnessed framework for meta-editing agentic evolution. It standardizes the evolution environment and provides a unified interface for observing accumulated evidence and editing the mechanism that drives future evolution. This lets AEvo revise both hand-designed procedures and agent operating contexts, reducing local-optimum risks in long-horizon evolution. Empirical evaluations on agentic and reasoning benchmarks show that AEvo outperforms 5 evolution baselines, achieving a 26% relative improvement over the strongest baseline. Furthermore, across 3 open-ended optimization tasks, AEvo outperforms 4 evolution baselines and achieves state-of-the-art performance under the same iteration budget.
PaperID: 811, Poster
Abstract: This paper presents a post-Bayesian approach to online filtering in nonlinear state-space models, capable of avoiding over-confident inferences in settings where either the dynamical model, the measurement model, or (PrO) posteriors, an emerging paradigm in which learning (i.e., posterior concentration) occurs if and only if the overall model is well-specified, without strict adherence to the Bayes' theorem. As the characterisation of PrO posteriors is challenging, our main technical contribution is a fast approximate linear-Gaussian update procedure, analogous to an (iterated) extended Kalman filter (EKF). The methodology, which we call PrO-EKF, has no tunable hyper-parameters and has a computational cost comparable to that of existing filtering methods. Performance is empirically assessed on a range of linear and non-linear applications, in which the state-space model is systematically misspecified.
Authors:
David Serrano-Lozano, Anand Bhattad, Luis Herranz, Jean-Francois Lalonde, Javier Vazquez-CorralAbstract: We present SyncLight, a method to enable consistent, parametric control over light sources across multiple uncalibrated views of a static scene conditioning in a single view. While single-view relighting has advanced significantly, existing generative approaches struggle to maintain the rigorous lighting consistency essential for multi-camera broadcasts, stereoscopic cinema, and virtual production. SyncLight addresses this by enabling precise control over light intensity and color across a multi-view capture of a scene, conditioned on a single reference edit. Our method leverages a multi-view diffusion transformer trained using a latent bridge matching formulation, achieving high-fidelity relighting of the entire image set in a single inference step. To facilitate training, we introduce a large-scale hybrid dataset comprising diverse synthetic environments-curated from existing sources and newly designed scenes-alongside high-fidelity, real-world multi-view captures under calibrated illumination. Surprisingly, though trained only on image pairs, SyncLight generalizes zero-shot to an arbitrary number of viewpoints, effectively propagating lighting changes across all views, without requiring camera pose information. SyncLight enables practical relighting workflows for multi-view capture systems. Dataset, code, and models will be released upon acceptance.
Abstract: Graph Neural Networks deployed in safety-critical settings must recognise when a node's neighborhood is structurally novel, conflicting, or underrepresented in the training graph; standard softmax classifiers collapse this epistemic ambiguity into a point estimate over known classes, producing confident errors on unseen structure. We propose Random-Set Graph Neural Networks (RS-GNN), a framework that replaces the softmax head of any message-passing backbone with a belief function over a budgeted family of focal sets. Each node is mapped to a mass function whose pignistic projection yields the point prediction and whose induced credal-set width quantifies epistemic uncertainty. Focal sets are selected from class-overlap statistics in the learned node-embedding space, making the budget graph-aware rather than uniform over the power set. Across nine node-classification benchmarks spanning homophilic and heterophilic regimes (Cora, Coauthor, Reddit2, ArXiv, Patents, Chameleon, Squirrel, Roman Empire, Amazon-Ratings), RS-GNN matches or outperforms strong probabilistic, ensemble, and credal baselines (including Classical Ensemble, GEBM, and CaGCN) on leave-out-class OOD detection, reaching AUROC of 88.84 on Cora and 84.18 on Chameleon. On temporal road-scene graphs constructed from ROAD and nuScenes, RS-GNN improves nuScenes node-classification accuracy from 0.418 to 0.593 and avoids the overconfidence collapse that softmax heads exhibit under cross-dataset shift. These results indicate that credal-set width recovers an epistemic signal that scalar softmax-derived scores cannot, particularly when graph evidence is incomplete or out of distribution.
Abstract: Editing a 3D asset locally, modifying a target region while preserving the rest, is a fundamental requirement of native 3D editing. Existing methods enforce locality through mechanisms external to the generator, such as manual 3D masks, post-hoc voxel merging, or 2D multi-view lifting. None of them intervene where the corruption actually originates: inside the ODE sampler. For a rectified-flow generator to achieve faithful local editing, its velocity field should be strong over the target editing region while vanishing on preserved content. Yet a single velocity field can hardly satisfy both requirements simultaneously, leading to three problems: (i) identity leakage that keeps the edit signal non-zero on preserved regions; (ii) no dedicated edit-amplification channel, so strengthening the edit inevitably perturbs identity; and (iii) an identity drag at the geometry and material stages, where a global condition pulls every token toward the target. We propose VS3D (Velocity-Space 3D Asset editing), an inversion-free, training-free, and mask-free framework that addresses each problem with a targeted intervention inside the sampler. VS3D integrates three complementary modules, each corresponding to a specific stage of the editing pipeline. Reconstruction-Anchored Source Injection (RASI) absorbs identity leakage by turning the unconditional embedding into a per-step, asset-specific anchor calibrated through source reconstruction. Partial-Mean Guidance (PMG) amplifies the edit signal by contrasting high- and low-quality subsample estimates of the velocity difference, active only where a consistent edit exists. Twin-Agreement Residual injection (TAR) lets the sampler decide token by token what to preserve at the geometry and material stages. Experiments on diverse 3D assets show that VS3D outperforms state-of-the-art native-3D and 2D-lifted editors, demonstrating that a purely velocity-space approach can serve as a general-purpose editing paradigm for pretrained 3D generators.
PaperID: 815, Poster
Abstract: Existing image editing methods can be generally categorized into textual instruction-based and visual prompt-based ones. Textual instructions are semantically expressive, but are limited by the coarse granularity of spatial control of the editing results. In contrast, visual prompts such as drag and point can provide precise spatial guidance, but are limited by the inherent ambiguity in semantic intent. To unify the strength of textual and visual prompts, we present Text-Vision Co-Instructed Image Editing, which jointly models textual instructions as semantic intent and sparse visual instructions as spatial guidance, aiming to achieve precise and intent-faithful image manipulation. To this end, we first construct a textual-visual instruction paired dataset with more than 23K samples derived from dynamic videos, enabling aligned supervision for cross-modal instruction. We then propose TV-Edit, a Textual-Visual instruction unified Editing framework to contextualize drag or point-based visual instructions with image-text semantics and lift them into semantic-aware control representations for pretrained editing backbones. By integrating semantic intent and spatial constraints, TV-Edit leads to more precise spatial control, less instruction ambiguity, and stronger structural consistency than text-only or drag-based alternatives. Finally, we establish TV-Edit-Bench, a deliberately designed benchmark to evaluate semantic faithfulness, spatial alignment, and visual consistency with ground-truth references and controlled textual–visual variations for reliable assessment. Our experiments across multiple editing backbones demonstrate that TV-Edit consistently yields more precise and intent-faithful edits, significantly outperforming state-of-the-art instruction-based and drag-based baselines. Data, model, and codes will be released.
Abstract: Despite the effectiveness of Cube-and-Conquer (C&C) for solving challenging Boolean Satisfiability (SAT) problems, no prior work has shown that transformer-based models can learn effective cubing heuristics. We introduce a neuro-symbolic post-training framework for this task. We design an MCTS-based data curation pipeline that uses symbolic heuristics to explore splitting decisions over SAT competition formulas, producing preference data grounded in solver statistics and augmented with reasoning traces from a teacher model. Our two-stage post-training, supervised fine-tuning (SFT) followed by direct preference optimization (DPO), enables a 4B-parameter model to achieve a pass@5 score of 53 on 100 SAT competition benchmarks, surpassing frontier LLMs such as Claude-Sonnet-4 (50) and matching the best symbolic heuristic (53). Ablations show that SFT alone improves pass@5 from 46 to 51, with DPO adding 2 additional benchmarks; an entropy/agreement ablation on realized first-cube decisions further shows that SFT, not DPO, accounts for the root-level decision diversity that produces complementary per-run coverage over deterministic symbolic methods. This demonstrates that transformers can be trained to make effective cubing decisions in a domain traditionally dominated by symbolic methods.
PaperID: 817, Poster
Abstract: Training implicit neural representations (INRs) to capture fine-scale details typically relies on iterative backpropagation and is often hindered by spectral bias when the target exhibits highly non-uniform frequency content. We propose ELM-INR, a backpropagation-free INR that decomposes the domain into overlapping subdomains and fits each local problem using an Extreme Learning Machine (ELM) in closed form, replacing iterative optimization with stable linear least-squares solutions. This design yields fast and numerically robust reconstruction by combining local predictors through a partition of unity. To understand where approximation becomes difficult under fixed local capacity, we analyze the method from a spectral Barron norm perspective, which reveals that global reconstruction error is dominated by regions with high spectral complexity. Building on this insight, we introduce BEAM, an adaptive mesh refinement strategy that balances spectral complexity across subdomains to improve reconstruction quality in capacity-constrained regimes.
Abstract: The proliferation of Large Language Models (LLMs) necessitates efficient mechanisms to distinguish machine-generated content from human text. While statistical watermarking has emerged as a promising solution, existing methods suffer from two critical limitations: the lack of a principled approach for selecting sampling distributions and the reliance on fixed-horizon hypothesis testing, which precludes valid early stopping. In this paper, we bridge this gap by developing the first e-value-based watermarking framework, Anchored E-Watermarking, that unifies optimal sampling with anytime-valid inference. Unlike traditional approaches where optional stopping invalidates Type-I error guarantees, our framework enables valid, anytime-inference by constructing a test supermartingale for the detection process. By leveraging an anchor distribution to approximate the target model, we characterize the optimal e-value with respect to the worst-case log-growth rate and derive the optimal expected stopping time. Our theoretical claims are substantiated by simulations and evaluations on established benchmarks, showing that our framework can significantly enhance sample efficiency, reducing the average token budget required for detection by 13-15% relative to state-of-the-art baselines.
PaperID: 819, Poster
Authors: Erland B Olsson, Zhirong Yang
Abstract: The de facto standard way of training neural networks requires storing activations at each layer to memory, in order to compute gradients in backpropagation. As progress has often been made by stacking larger and more layers, memory consumption has now become a major bottleneck for making frontier models available to everyone. In this work, we take advantage of the Residual Network (ResNet) layer formulation that is now used in many state-of-the-art neural network architectures to recompute the activations at each layer on-the-fly during backpropagation instead of storing them in memory. We propose a modification to Reversible Residual Network (RevNet) that overcomes its limitation of requiring to replace the ResNet layer network function with two new network functions, by setting either one to an identity function. This modification is simple yet has a significant advantage because now there is no need to change the original ResNet architecture. Our method can work as a drop-in replacement for layers with residual connections such as in ordinary ResNets or Transformers with practically no loss in performance while offering significant reductions in memory consumption. With this simple modification, we are able to maintain the same computational cost and ease of use as activation checkpointing, which requires no architectural changes, while leveraging a reversible procedure such that we only need to store the activations from the last layer. The method is called Simplified RevNet, and compared to previous work in reversible architectures, we here propose a simpler and more streamlined approach that comes in two variants based on which of F or G in RevNet becomes an identity function. Empirically, we demonstrate performance at practically the same level as the non-reversible counterparts on ImageNet image classification with ResNets and on OpenWebText language modeling with Transformers.
Abstract: Stochastic gradient descent (SGD) is a fundamental optimization algorithm widely used in modern machine learning. In this paper, we propose Factor-Augmented SGD (FSGD), a new optimization method that leverages latent factor representations in high-dimensional learning tasks. Unlike standard two-stage dimension reduction approaches that rely on offline representation learning and full data storage, a key novelty of FSGD is that it operates purely on streaming data, making it scalable to large-scale and high-dimensional problems. Furthermore, we establish the first theoretical framework that explicitly incorporates latent factor estimation error into the analysis of SGD, and provide moment convergence in \ell^s norm under decaying step sizes and mini-batch updates. Our results provide a new foundation for employing SGD reliably and scalably in high-dimensional machine learning systems.
Authors: Yuheng Lai, Garvesh Raskutti
Abstract: Conformal prediction is a framework that provides valid uncertainty quantification for general models with exchangeable data. However, in the online learning and time-series settings, exchangeability is not satisfied. Existing online conformal methods, such as adaptive conformal inference (ACI), can achieve long-run validity, yet they remain inefficient under covariate heterogeneity because they rely on global calibration. We propose , which combines online adaptation with covariate-dependent localization to better reflect heterogeneity. To reduce sensitivity to the localization bandwidth, we further develop , which performs bandwidth selection as an online expert aggregation problem using a constrained online convex optimization framework. Importantly, we provide coverage guarantees for both algorithms and demonstrate through simulations and real-data experiments that the proposed methods attain valid long-run coverage with narrower prediction intervals than existing baselines.
PaperID: 822, Poster
Abstract: Diffusion models represent a powerful class of generative models known for their solid theoretical foundations and remarkable performance across diverse tasks and domains. While diffusion models have been extensively utilized for generating entire graphs or small-scale graphs, no diffusion-based approaches have been developed to synthesize graph structures within an existing graph, including synthetic nodes and their associated edges. In this study, we introduce the Robust Graph Diffusion Model (RGDM), designed to generate labeled synthetic graph structures consisting of nodes and edges that integrate seamlessly into a given graph. The RGDM consists of a Robust Graph Autoencoder (RGAE) and a Latent Diffusion Model (LDM). Leveraging an edge selection mechanism and an innovative low-rank regularization on the latent feature, the RGDM produces clean and high-quality synthetic graph structures, even when trained on graphs subject to adversarial attacks. Comprehensive experimental evaluations reveal that Graph Neural Networks (GNNs) trained on the augmented graph, which is formed by merging the original attacked graph with the synthetic graph structures, exhibit significantly improved robustness against various graph adversarial attacks in the context of semi-supervised node classification. The code of the RGDM is available at \urlhttps://anonymous.4open.science/r/RGDM.
Authors: Chao Yan
Abstract: We study \emphdifferentially private prediction introduced by Dwork and Feldman (COLT 2018): an algorithm receives one labeled sample set S and then answers a stream of unlabeled queries while the output transcript remains (\varepsilon,\delta)-differentially private with respect to S. Standard composition yields a \sqrtT dependence for T queries, i.e. O(VC(\mathcalC)\cdot\sqrtT) sample complexity. We show that this dependence can be reduced to \emphpolylogarithmic in T in streaming settings. For an oblivious online adversary and any concept class \mathcalC, we give a private predictor that answers T queries with |S|= \tildeO(VC(\mathcalC)^3.5\log^3.5T) labeled examples. For an adaptive online adversary and halfspaces over \mathbbR^d, we obtain |S|=\tildeO\left(d^5.5\log T\right).
PaperID: 824, Poster
Abstract: The Moore-Penrose Pseudo-inverse (PInv) is the fundamental tool for inverting linear operators. We propose a natural generalization to the non-linear regime and introduce Surjective Pseudo-invertible Neural Networks (SPNN): architectures that admit a tractable non-linear PInv by construction and satisfy the corresponding geometric properties. Building on this, we formalize Non-Linear Back-Projection (NLBP), the non-linear analogue of x' = x + A^\dagger(y-Ax): an update that projects any sample to its closest state consistent with f(x)=y. Diffusion-based null-space projection revolutionized zero-shot solving of linear inverse problems via closed-form back-projection; NLBP extends this paradigm to non-linear learned ``degradations'' in the broad sense, spanning semantic mappings such as classification and object detection. The result is zero-shot inversion of complex degradations and precise semantic control over generative outputs without retraining the prior.
PaperID: 825, Poster
Abstract: Operations research practitioners typically tackle NP-hard combinatorial problems using large neighborhood search (LNS), a scalable heuristic that iteratively refines a current solution by locally re-optimizing subsets of its variables. In contrast, most existing approaches for integrating combinatorial optimization layers into neural networks still assume access to an exact global solution, which is computationally intractable. We bridge this gap by introducing regularized LNS (RLNS). By regularizing or perturbing local subproblems, we turn the LNS heuristic into an efficient MCMC sampler over the combinatorial set of feasible solutions, with associated Fenchel-Young losses. Under entropic regularization, we prove that RLNS performs exact block Gibbs sampling. Furthermore, adjusting the number of RLNS iterations allows us to interpolate between pseudolikelihood and exact maximum likelihood estimation, for end-to-end learning without global solvers. We demonstrate our approach on k-subset selection, generalized assignment, and stochastic vehicle scheduling problems.
PaperID: 826, Poster
Abstract: Randomized Smoothing (RS) is a principled and widely adopted approach for certifying the robustness of black-box models, such as neural networks, against adversarial perturbations. We propose an optimization perspective that unifies several prior randomized smoothing certificates. We then show how to reduce the underlying high-dimensional worst-case optimization problem to an equivalent two-dimensional formulation under a double- or multi-sampling scheme, enabling efficient solutions via convex optimization. By exploiting additional information about the variance of the smoothed classifier across different smoothing distributions, we derive certified bounds that improve upon standard certificates. We show results on both classification and regression tasks.
Abstract: In machine learning applications, privacy requirements during inference or deployment time could change constantly due to varying policies, regulations, or user experience. In this work, we aim to generate a magnitude of models to satisfy any target differential privacy (DP) requirement without additional training steps, given a set of existing models trained on the same dataset with different privacy/utility tradeoffs. We propose two post processing techniques, namely random selection and linear combination, to output final private models satisfying any target privacy parameter. We provide privacy accounting of these approaches from the lens of R\'enyi DP and privacy loss distributions for general problems. In a case study on private mean estimation, we precisely characterize the privacy/utility results and theoretically establish the superiority of linear combination over random selection. Empirically, we validate our approaches and analyses on several models and both synthetic and real-world datasets.
Abstract: Modern AI models are not static. They go through multiple updates in their lifecycles. We propose to design Sequential Membership Inference (SeMI) attacks leading to tighter privacy audits by exploiting the sequence of models and injecting a target canary at a controlled insertion time. First, for empirical mean computation, we develop \mathrmSeMI^\ast, an optimal SeMI attack to identify the presence of a target inserted at a specific insertion step. We derive the power of \mathrmSeMI^\ast to show that accessing the model sequence yields more powerful MI attacks than scrutinising only the final model. \mathrmSeMI^\ast exhibits an isolation property-- its power depends on the statistics obtained right before and after insertion of the target. Leveraging this insight, we develop practical white-box (accessing model gradients) and black-box (accessing loss) SeMI attacks against models trained with (DP-)SGD. Across datasets and models trained with (DP-)SGD, our experiments show that SeMI attacks achieve higher powers than snapshot-independent baselines, and yield tighter privacy audits thanks to (a) control over the insertion time and (b) observations across the model sequence.
PaperID: 829, Poster
Abstract: Canonical Correlation Analysis (CCA) is a fundamental method for multiview shared space learning. However, its strict reliance on paired data poses a significant limitation, as such data is often difficult to obtain or entirely unavailable. In this paper, we present Unpaired CCA (UCCA), a novel method that learns linear projections to maximize the correlation of the true underlying pairing without access to any paired samples during training. We first establish theoretical results connecting the Quadratic Assignment Problem (QAP) to CCA. Leveraging these theoretical insights, we derive a practical method to maximize correlation exclusively from unpaired data. To the best of our knowledge, UCCA is the first approach to learn maximally correlated projections in a strictly unpaired setting. We validate UCCA on real-world multi-modal datasets, demonstrating that it significantly outperforms recent unpaired alignment baselines in recovering the underlying true correlation. This work fills a critical gap between traditional statistical multiview learning and the growing field of unpaired data learning.
Abstract: Conformal risk control is an emerging framework for the safe deployment of machine learning models with finite-sample guarantees. To accommodate a broader class of risk notions, quantile risk control extends this framework to quantile-based risk measures. However, existing methods either suffer from excessive conservatism or lack rigorous finite-sample guarantees. To address these limitations, we introduce Occupancy-based Quantile Risk Control (OQRC), a novel method that provides tight risk control bounds with finite-sample validity. Our key idea is to formulate risk control as a finite-occupancy problem by partitioning the loss space with the ordered calibration losses. Specifically, we estimate the distribution of test losses across the resulting bins and upper-bound the risk by the maximum loss attained within each bin. We then select the parameter \lambda such that this upper bound does not exceed a predefined threshold \alpha with high probability 1-\delta. Theoretically, we establish a finite-sample guarantee showing that OQRC yields tight risk control bounds that converge to the optimal bounds at a provable rate of \mathcalO_p(n^-1/2). Extensive experiments demonstrate the effectiveness of our method, reducing the risk gap by up to 78.64% on common benchmarks.
Abstract: We consider the problem of estimating the Attention mechanism in small space, and prove the existence of coresets for it of nearly optimal size. Specifically, we show that for any set of unit-norm keys and values (K,V) in R^d, there exists a subset (K',V') of size at most O(\sqrtd e^\zeta+o(\zeta)/\epsilon) such that \left\| Attn(q,K,V)- Attn(q,K',V') \right\| \le \epsilon simultaneously for all queries whose norm is bounded by \zeta. This outperforms the best known results for this problem. We also offer an improved lower bound showing that \epsilon-coresets must have size \Omega(\sqrtd e^\zeta/\epsilon).
PaperID: 832, Poster
Authors: Martin Benes, Rainer Böhme
Abstract: Current approaches to learning-based image steganography are, at best, imperceptible, but their security is vulnerable to state-of-the-art detector networks. Variational cover modification steganography tackles this problem by training a generator network that can predict the parameters of the distribution of the least detectable stego-noise for a given cover image. It uses syndrome coding to embed the message within the distribution constraint and retrieve it exactly when a secret key is provided. Adversarial learning ensures that the generator evolves against a continuously refined detector. Experiments demonstrate significant security improvements, and several ablations support the generator architecture and the training protocol.
Authors:
Truong Buu Phan, Gergely Flamich, Ashish Khisti, Shahab AsoodehAbstract: Convergence diagnosis for Markov chain Monte Carlo is a matter of fundamental importance in computational statistics: it determines the resources allocated to a particular sampling problem and influences the practitioner's view of the quality of estimates obtained from a Markov chain. Motivated by this, we contribute to the emerging class of coupling-based convergence diagnostic algorithms. Concretely, we study coupling multiple Metropolis--Hastings chains using multi-marginal coupling. We introduce a natural objective for this setting and establish lower and upper bounds by drawing connections to list-level distribution coupling and distributed pairwise-matching problems. This analysis ultimately leads to a shared-randomness Poisson Monte Carlo construction for coupling multiple Markov chains. In this process, we avoid a key dimension-dependent bottleneck in the runtime complexity of classical Poisson Monte Carlo by developing an adaptive rule for updating the point process, yielding significant gains in high-dimensional settings. Experiments on grand couplings of Markov chains show that our methods improve coalescence rates across dimensions, reducing meeting times by up to 50% compared with existing baselines.
Abstract: A central obstacle in nonlinear Bayesian filtering is representing the belief distribution. Moment-based filters address this by propagating polynomial moments and reconstructing a density from them. Recent work completes the predict-update loop via the maximum-entropy (MaxEnt) principle, but each step requires the partition function and its gradient, both n-dimensional integrals whose cost scales exponentially, restricting the demonstrated MaxEnt moment filtering to n \le 4. We avoid the partition function entirely by combining score matching with Stein's identity. In our setting, score matching reduces the density fit to a single linear solve whose coefficients are assembled directly from the propagated moments. The same parameters then drive Stein's identity to close the moment hierarchy during prediction and to recover posterior moments after each Bayesian update, keeping the full predict-update loop free of partition function evaluation. The resulting Score Kalman Filter (SKF) reduces to the classical information-form Kalman filter as a special case and performs every step through linear algebra. On nonlinear coupled-oscillator networks, the SKF runs through n=20 and reports lower RMSE than the EKF, UKF, EnKF, and particle-filter baselines on the tested synthetic benchmarks.
Abstract: Transformers use the same forward computation stream to both predict the next token and store useful state for future token predictions. We formulate the \emphstate-prediction separation hypothesis: disentangling the two roles yields better language modeling performance. We design a Transformer variant that uses separate computation streams to separate the two functions, and conduct pretraining experiments across various scales. Our experiments show that state-prediction separation consistently offers better data-- and compute--efficiency, improving validation loss and outperforming standard Transformers by 2--3 percentage points on average on downstream tasks. We also conduct extensive empirical analysis that rules out potential confounders and demonstrates the fundamental difference in the gradients our design entails.
PaperID: 836, Poster
Abstract: This work introduces a family of generative diffusion models based on the stochastic heat equation. The deterministic heat equation has natural connections with kernels over discrete structures and has been successfully applied in machine learning to discriminative tasks for various types of data such as graphs. We extend it to a stochastic version by formulating our diffusion as a Gaussian process whose transition kernel can be written in closed form for both the forward- and reverse-time processes. Our diffusion process propagates noise smoothly among neighbours under the guidance of a Laplacian matrix that describes the global structure of the input. Following a flow matching-based approach, we train a denoising neural network that takes noisy representations as input and gradually removes noise to generate clean samples. We experimentally show that, by exploiting the Laplacian during training, our decoder is able to recover the global structure of the input.
Abstract: Safe exploration remains a fundamental challenge in reinforcement learning (RL), limiting the deployment of RL agents in the real world. We propose (SBSRL), a model-based RL algorithm that maintains safety throughout the learning process by enforcing constraints jointly across a set of dynamics samples. This formulation approximates an intractable worst-case optimization over uncertain dynamics and enables practical safety guarantees in continuous domains. We further introduce an exploration strategy based on constraining epistemic uncertainty, eliminating the need for explicit exploration bonuses. Under regularity conditions, we derive high-probability guarantees of safety throughout learning and a finite-time sample complexity bound for recovering a near-optimal policy. Empirically, SBSRL achieves safe and efficient exploration both in simulation and in real-world experiments, and readily extends to practical deep-ensemble implementations that scale to high-dimensional continuous control problems.
PaperID: 838, Poster
Authors:
Benjamin Van Niekerk, Jean-Philippe Letendre, Nicol Visser, Hugo Seuté, Herman Kamper, Mirco RavanelliAbstract: Speech tokenizers convert audio into discrete units, enabling language models to process and generate speech. The goal is to learn compact, text-like representations while preserving enough detail to reconstruct a speaker’s voice and delivery. We propose Facto: a factorized tokenizer that disentangles time-varying information (content and prosody) from static components (speaker identity and recording conditions). We build on a recent linear factorization method for self-supervised features, extending it to handle unseen speakers. Then, we cluster the factorized space and train a lightweight decoder to reconstruct audio from the resulting tokens. By removing the static components before clustering, our tokens are more robust to small noise perturbations. We validate this by testing token stability under various noise conditions. Next, we theoretically analyze Facto's convergence and scaling properties, quantifying how the factorization improves with the number of training speakers. Finally, we evaluate Facto on a range of downstream tasks. We show that greater token stability translates into improved text-to-speech performance and competitive results in voice conversion and spoken language modeling.
PaperID: 839, Poster
Authors: Peiwen Qiu, Prashant Khanduri, Jia (Kevin) Liu
Abstract: Bilevel optimization is a fundamental framework for machine learning problems, where an upper-level (UL) objective depends on the solution of a nested lower-level (LL) problem. Existing bilevel algorithms typically assume that the LL objective is available for evaluation, differentiation, or unrolling. However, this assumption can fail when the LL response is produced by a black-box solver, simulator, adversary, or learned predictor. In this paper, we study lower-level agnostic bilevel optimization, where the UL learner has access only to an approximate LL solution \bar\mathbfy(\mathbfx) and cannot query the LL objective, its gradients, Hessian/Jacobian information, or the procedure that generates the LL response. We propose LeGo-BiO (Lower-level Agnostic Bilevel Optimization), a coordinate-wise pseudo-hypergradient method that estimates the missing Jacobian \nabla _\mathbfx\bar\mathbfy(\mathbfx) using finite differences of \bar\mathbfy(\cdot). For each UL coordinate, LeGo-BiO constructs a decision vector subject to 1-sparse updates, i.e., restricting modifications to a single coordinate per step, thereby isolating the corresponding Jacobian column needed in the UL gradient computation. This avoids the directional-projection limitation of full-parameter response differences and, unlike zeroth-order approximations of the entire UL gradient, preserves the available first-order gradient information with respect to the UL variables for more accurate UL gradient evaluation. We establish convergence guarantees for LeGo-BiO in both deterministic and stochastic settings under an LL Polyak–Łojasiewicz condition, without requiring LL strong convexity. Our bounds yield an \mathcalO(T^-1) rate in the deterministic setting when the LL approximation error decreases geometrically, and an \mathcalO(T^-1/2) rate in the stochastic setting under general LL approximation errors. Experiments on deep hyper-representation and adversarial training demonstrate competitive performance of LeGo-BiO despite being agnostic to the LL objective.
PaperID: 840, Poster
Abstract: Deep Gaussian processes (DGPs) are appealing Bayesian regression models, but their standard discrete-depth parameterization often becomes harder to optimize as depth grows. We argue that depth is better treated as a latent continuous evolution than as a long discrete stack. We therefore propose a continuous-depth formulation in which the latent representation follows a stochastic differential equation with a depth-indexed drift family equipped with a GP-based prior. This reframes the core design choice as the prior placed on that family, rather than adopting a purely neural continuous-depth parameterization as in neural ODE or neural SDE models. We study two instantiations of this idea: CDGP, a direct GP prior over the state-depth drift, and FlowDGP, a flow-evolved prior that captures richer depth dependence at lower training cost. To motivate the redesign, we show that the single-sample DGP objective can become increasingly ill-conditioned under repeated layer composition, while residual DGP is a useful first repair that still remains a finite discrete construction. We also give pathwise sensitivity analysis for the continuous-depth formulation. Experiments on synthetic, benchmark, and large-scale regression tasks show that residual DGP improves over plain DGP, while CDGP and FlowDGP provide a stronger overall story across the regimes we study.
PaperID: 841, Poster
Authors: Paul Häusner, Jevgenija Rudzusika, Jens Sjölund, Ozan Öktem
Abstract: Alternating minimization is a standard tool for optimization problems with two coupled variable blocks. Its cost is dominated by the inner subproblem solvers, which rarely have a closed-form solution. In this paper, we propose to learn an inexact alternating scheme in which neural networks approximate the subproblem solution at each iteration. We derive a worst-case bound on the optimality gap that exposes the per-iteration error that each network is trained to minimize in a greedy fashion. To show convergence of the learned scheme on unseen problem instances, we combine the bound with a PAC-Bayes argument that controls the expected optimality gap with high probability. We apply the learned scheme to low-dose CT reconstruction with dictionary regularization, demonstrating a 13× speedup over the inexact first-order baseline at matched accuracy. A brief end-to-end fine-tuning stage extends this to 30× and outperforms other learned schemes.
PaperID: 842, Poster
Abstract: Real-world scenes often contain static, slow, and fast components within the same field of view. Frame cameras are inherently inefficient in this setting, as their global exposure and fixed sampling must follow the fastest motion and therefore oversample the rest of the scene. Event cameras provide asynchronous local sensing, but standard sensors use a single global contrast threshold, leading to a trade-off between sensitivity and redundancy: low thresholds preserve static information but over-trigger under fast motion, while high thresholds reduce data volume but lose slow or static structures. We propose a multi-threshold event imaging model that assigns motion-compatible thresholds to different scene components, enabling faithful multi-velocity-scale imaging with fewer events. We also introduce a practical acquisition pipeline that implements this model on existing event cameras. Experiments on synthetic and real scenes validate the effectiveness of our approach in preserving both static structures and fast dynamics while reducing data volume.
PaperID: 843, Poster
Abstract: Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), which transforms the notoriously unstable, high-variance On-Policy Distillation (OPD) into a stable training paradigm by dynamically constructing a proximal teacher. Theoretically, we establish a rigorous framework demonstrating that TOP-D inherently controls gradient variance. By providing a formal global convergence analysis alongside a monotonic improvement bound, we mathematically formalize the reliability and stability of the overall training dynamics. Empirically, TOP-D dramatically enhances training stability, sample efficiency, and final performance on mathematical reasoning tasks. More importantly, TOP-D introduces zero additional computational overhead, positioning itself as a promising alternative to the well-established OPD paradigm.
PaperID: 844, Poster
Abstract: Motivated by applications in reinforcement learning from human feedback, we consider a setting where a given collection of rankings needs to be compressed into a smaller collection of rankings. This can also be seen as a combination of two fundamental social choice frameworks: ranking aggregation and multiwinner voting. We propose three proportionality notions that capture different types of representation in this setting: positional proportionality, pairwise proportionality, and proportionality for solid coalitions (PSC). On the one hand, we show that positional proportionality and PSC are always satisfiable for any input rankings and target compression size, and a desired output can be found in polynomial time. On the other hand, we prove that pairwise proportionality cannot be satisfied in general, but can nevertheless be attained when the input rankings are single-peaked or single-crossing.
PaperID: 845, Poster
Authors: Jordão Bragantini, Ilan Theodoro, Loic A Royer
Abstract: Reconstructing lineages from live-imaging microscopy requires linking cell detections across time, including through cell divisions. A common approach is to construct a candidate graph and associate cell segmentations (nodes) across frames. However, these and other existing methods overlook two structural obstacles in candidate tracking graphs: (i) cell divisions entangle distinct lineage paths in the node embedding space, and (ii) edges sharing a node have near-random label agreement, so the candidate-graph topology carries no useful information for graph neural networks to aggregate. We propose the (HOCT), an edge-centric architecture in which candidate cell links attend to one another under a 3D geometric prior, resolving both issues. Evaluated on the Cell Tracking Challenge and a bacteria division benchmark, HOCT achieves state-of-the-art results without deep pre-trained image encoders. Moreover, the proposed approach is easier to fine-tune, quickly reducing tracking errors by 59% with 400 annotations in a human-in-the-loop setting, outperforming LoRA fine-tuning of competing transformer baselines (6.75% improvement).
PaperID: 846, Poster
Authors:
Xin Teng, Muxiao Li, Hongyi WenAbstract: Mixture-of-Experts (MoE) transformers scale capacity by activating only a few experts per token, but this sparsity creates a hidden reliability problem: when routing is imperfect, load-balanced models may send tokens to experts that are insufficiently trained for the assigned inputs. We propose Distributionally Robust MoE Training (DRMoET), a drop-in objective that treats layer-wise experts as endogenous robustness groups and optimizes high-loss routing outcomes rather than merely equalizing traffic. DRMoET updates a per-layer expert distribution by an entropy-regularized softmax rule on EMA-smoothed, activation-weighted expert losses, strengthening plausible non-top routing paths while preserving standard MoE computation. Under the FLAME-MoE recipe at 746M-total and 10.3B-total scales, DRMoET improves downstream averages over both standard FLAME-MoE and auxiliary-loss-free balancing. Mechanistic analyses show lower expert-loss variance with nearly unchanged mean loss, 4.3% lower degradation under forced mid-k misrouting, and improved domain--expert specialization. These results position routing robustness--not only utilization balance--as a practical objective for reliable sparse MoE scaling.
PaperID: 847, Poster
Authors:
Emanuele Mezzi, Gertjan Burghouts, Fabio Massacci, Mengyuan ZhangAbstract: Multimodal question answering (MMQA) is based on integrating heterogeneous data sources, selectively leveraging relevant modalities, and ignoring distractors. Existing approaches based on Multimodal Large Language Models (MLLMs) improve performance by splitting the task into intermediate steps, but do not quantify how each modality contributes to reaching the final answer. We argue that quantifying the effect that each modality has on the answering process improves process traceability and final performance. Thus, we propose Entropy-Driven Multimodal Answering (EDMA), which uses logical entropy to formalise multimodal QA as an iterative process of entropy reduction and to quantify the contribution of each modality in distinguishing correct from incorrect candidate answers. This enables the construction of quantifiable answering trajectories, where each step is associated with a measurable entropy reduction. EDMA introduces (i) modality-conditioned partitioning to estimate unimodal entropy reduction, (ii) entropy-driven multimodal fusion to capture complementary information across modalities, and (iii) entropy-based selection of the answering trajectory that minimises the most logical entropy. Experiments on multimodal QA benchmarks show that EDMA consistently outperforms prompting-based baselines and state-of-the-art methods. Averaged across datasets, EDMA achieves gains of 12.8 and 19.8 F1 points over the strongest prompting baseline and state-of-the-art system, respectively. Moreover, we show that higher entropy correlates with more false positives and, through controlled interventions on the choice of the answering trajectory, provide evidence that lower-entropy partitions causally reduce false positives, validating entropy level as a reliable proxy for performance. We share our code at https://anonymous.4open.science/r/EDMA/README.md.
Abstract: Vision-Language Models (VLMs) exhibit remarkable zero-shot generalization, yet they often encode unwanted or hazardous stylistic domains such as idealized textbook diagrams in medical AI or cartoon vehicles in autonomous driving. Approximate Domain Unlearning (ADU) aims to selectively erase a model's recognition of a target visual domain while preserving accuracy on the remaining domains. However, existing ADU methods operate under a flawed closed-vocabulary assumption: they evaluate unlearning solely on the specific object classes seen during the unlearning fine-tuning phase. Consequently, these methods do not unlearn the domain itself; they merely overfit to seen class-domain pairs, leaving the domain easily recognizable for unseen classes and providing a false sense of removal. We argue that true domain erasure must be class-agnostic. To address this, we formalize Open-Vocabulary Domain Unlearning (OVDU), a rigorous protocol that mandates domain forgetting must transfer to held-out classes. To solve the OVDU challenge, we propose a surgical parameter-editing framework. First, a Fisher Information mask isolates domain-sensitive weights, mathematically protecting foundational zero-shot generalization. Second, our Targeted Manifold Scattering (TMS) objective uses preference-based mining to locally scatter the forget domain's stylistic geometry. Evaluated across PACS, OfficeHome, and DomainNet, our method vastly improves open-vocabulary generalization over existing baselines. Crucially, it delivers exceptional sample efficiency, outperforming peak 8-shot baseline results with only 4 shots.
Abstract: We propose a conservative continuous-time stochastic control framework for treatment optimization from irregularly sampled patient trajectories. We model the unknown patient dynamics as a controlled stochastic differential equation where treatment is a continuous-time control input. Naive model-based optimization can exploit model errors and propose out-of-distribution controls that appear optimal under the learned model but perform poorly under the true dynamics. To mitigate such extrapolation, we introduce a consistent, signature kernel MMD regularizer on path space, which penalizes controls whose predicted trajectories deviate from observed ones. The resulting conservative objective minimizes a tractable upper bound on the true cost. Experiments on pharmacometric simulators show improved robustness and performance compared to non-conservative baselines.
Abstract: Simulating trajectories of multi-particle systems on complex energy landscapes is a central task in molecular dynamics (MD) and drug discovery, but remains challenging at scale due to computationally expensive and long simulations. Previous approaches leverage techniques such as flow or Schrödinger bridge matching to implicitly learn joint trajectories through data snapshots. However, many systems, including biomolecular systems and heterogeneous cell populations, undergo interactions that evolve over their trajectory and cannot be captured through static snapshots. To close this gap, we introduce , a framework that learns the first- and second-order stochastic dynamics of interacting, multi-particle systems where the direction and magnitude of each particle's path depend dynamically on the paths of the other particles. We define the Entangled Schrödinger Bridge (EntangledSB) problem as solving a coupled system of bias forces that particle velocities. We show that our framework accurately simulates heterogeneous cell populations under perturbations and rare transitions in high-dimensional biomolecular systems. Anonymous code is provided at https://anonymous.4open.science/r/EntangledSBManon.
PaperID: 851, Poster
Abstract: Multinomial logistic regression (MLR) is the default last-layer classifier in fully supervised recognition, while prototype classifiers are central to few-shot recognition. Although they are usually presented as different principles, we show that both can be organized by a geometric notion of class boundaries. A Euclidean MLR logit is a signed distance to a class hyperplane, so ordinary MLR is already the hyperplane decision-boundary case of Umbilic Multinomial Logistic Regression (UMLR). Classical submanifold geometry supplies the matching spherical decision boundary: in Euclidean space, the complete connected totally umbilic hypersurfaces are hyperplanes and spheres. The spherical UMLR form keeps a prototype-like center but adds a learnable radius; its signed squared-distance form contains squared prototype scoring as the radius-zero case. Thus UMLR is not a separate replacement for MLR or prototypes, but a boundary-aware family that places hyperplane logits, spherical logits, mixed heads, and prototype scoring in one framework. Experiments on low-shot vision head swaps, prototype comparisons, and mixed hyperplane--sphere heads support this organization and show how the different decision boundaries behave across data regimes. UMLR provides a compact geometric language for studying when class evidence is better modeled by flat boundaries, spherical boundaries, or their mixture.
Abstract: Tiny Recursive Models (TRM) solve complex reasoning tasks with a fraction of the parameters of modern large language models (LLMs) by iteratively refining a latent state and final answer. While powerful, their deterministic recursion can lead to convergence at suboptimal solutions, without escape mechanism. A common workaround relies on task-specific input perturbations at test time combined with answer aggregation via voting. We introduce Probabilistic TRM (PTRM), a task-agnostic framework for test-time compute scaling that addresses this limitation through stochastic exploration. PTRM injects Gaussian noise at each deep recursion step, enabling parallel trajectories to explore diverse solution basins, and selects among them using the model’s existing Q head (used for early stopping in the original TRM). Without requiring retraining or task-specific augmentations, PTRM enables substantial accuracy gains across benchmarks, including Sudoku-Extreme (87.4% to 98.75%) and on various puzzles from Pencil Puzzle Bench (65% to 91%). On the latter, PTRM achieves nearly double the accuracy of frontier LLMs (91% vs. 55%) at less than 0.0001x the cost, using only 7M parameters.
Abstract: Concept Bottleneck Models (CBMs) have become a popular approach to enable interpretability in neural networks by constraining classifier inputs to a set of human-understandable concepts. While effective, current models embed concepts in flat Euclidean space, treating them as independent, orthogonal dimensions. Concepts, however, are highly structured and organized in semantic hierarchies. To resolve this mismatch, we propose Hyperbolic Concept Bottleneck Models (HypCBM), a post-hoc framework that grounds the bottleneck in this structure by reformulating concept activation as asymmetric geometric containment in hyperbolic space. Rather than treating entailment cones as a pre-training penalty, we show they encode a natural test-time activation signal: the margin of inclusion within a concept's entailment cone yields sparse, hierarchy-aware activations without any additional supervision or learned modules. We further introduce an adaptive scaling law for hierarchically faithful interventions, propagating user corrections coherently through the concept tree. Empirically, HypCBM rivals post-hoc Euclidean models trained on 20× more data in sparse regimes required for human interpretability, with stronger hierarchical consistency and improved robustness to input corruptions.
PaperID: 854, Poster
Authors:
Zilin Li, Weiwei Xu, Xuchun Tong, Xuanbo Lu, Xuanqi Zhao, Ge ZhangAbstract: The learning dynamics of decentralized coupled representation learning can be formulated as a non-autonomous stochastic dynamical system, yet the macroscopic behavior of the resulting system on continuous manifolds remains poorly understood. We establish three theoretical results for such coupled slow-fast dynamics. First, under a standard random geometric graph scaling regime, we prove that the infinitesimal generator of the discrete Markov process converges to the generator of an overdamped Langevin diffusion for every f \in C^3(\mathcalM). Second, under timescale separation, the stochastic parameter dynamics converge to a deterministic averaged flow that admits a row-orthogonality Lyapunov function, controlling row-space degeneracy and ruling out parameter divergence. Third, under a spectral gap condition, the row space of the averaged parameter trajectory converges to the principal eigenspace of the induced covariance matrix. Together, these results give a stochastic-approximation limit theory for the coupled regime in which Markovian sampling and local subspace learning co-evolve, showing that memoryless local interactions can induce a stable and structured macroscopic limit.
Abstract: Large Language Models (LLMs) pose a significant risk of safety misalignment after finetuning, as models can be compromised by both explicitly and implicitly harmful data. Even some seemingly benign data can inadvertently steer a model towards misaligned behaviors. To address this, we introduce GradShield, a principled filtering method that safeguards LLMs during finetuning by identifying and removing harmful data points before they corrupt the model's alignment. It removes potentially harmful data by computing a Finetuning Implicit Harmfulness Score (FIHS) for each data point and employs an adaptive thresholding algorithm. We apply GradShield to multiple utility fine-tuning tasks across varying levels of harmful data and evaluate the safety and utility performance of the resulting LLMs using various metrics. The results show that GradShield outperforms all baseline methods, consistently maintaining an Attack Success Rate (ASR) below 6% while preserving utility performance.
Abstract: Aligning large language models with human preferences must balance two competing goals: responding helpfully to legitimate requests and reliably refusing harmful ones. Most preference-based safety alignment methods collapse safety into a single scalar that is applied uniformly to every preference pair. The result is a model that looks safe on average but stays relatively unsafe on a minority of harm categories. We cast safety alignment as a per-category constrained optimization problem and derive Cat-DPO, a direct-preference-optimization algorithm with a separate adaptive safety margin for each harm category. The margin tightens when the model still produces unsafe responses on a category and relaxes once the model catches up, so the training signal tracks each category's current difficulty rather than averaging under one global rate. Across two LLM backbones and six preference-learning baselines, Cat-DPO improves aggregate helpfulness and harmlessness and compresses per-category safety variance and the best-to-worst gap, offering a drop-in per-category refinement of direct preference safety alignment.
Authors: Wael AbdAlmageed
Abstract: Neuro-symbolic systems typically couple a neural perception module with a discrete symbolic solver, where exact constraint inference is intractable at training time and a discrete solver must be invoked at inference time, precluding end-to-end optimization of the model. We introduce AS2 (Attention-Based Soft Answer Sets), a fully differentiable neuro-symbolic architecture that replaces the discrete solver with a continuous approximation of the Answer Set Programming (ASP)immediate-consequence operator T_P. AS2 maintains per-position probability distributions over a finite symbol domain throughout the forward pass and trains end-to-end by minimizing the fixed-point residual of a probabilistic lift of T_P, thereby differentiating through the constraint check without invoking an external solver during either training or inference. On spatial constraint-satisfaction tasks, AS2 replaces conventional positional embeddings entirely, and problem structure is encoded through constraint-group membership embeddings derived directly from the declarative ASP specification, making the model agnostic to arbitrary position indexing. On Visual Sudoku, AS2 achieves 99.89% cell accuracy and 100%constraint satisfaction, without invoking any external solver. On CLEVR-Hans, AS2 achieves 99.94% on CLEVR-Hans3 and 87.77% on CLEVR-Hans7. These results demonstrate that a soft differentiable fixpoint operator, combined with constraint-aware attention and declarative constraint specification, can match or exceed pipeline and solver-based neuro-symbolic systems while maintaining full end-to-end differentiability.
PaperID: 858, Poster
Abstract: , a neural solver for elliptic PDE problems on variable-shape domains. The harmonic measure of a domain is the boundary probability distribution that, integrated against any boundary data, returns the Dirichlet Laplace solution. It depends only on the geometry, not on the boundary data. NHMO parameterizes this measure as a transformer-based boundary kernel supervised by Walk-on-Spheres exit samples, so one trained kernel handles different boundary value on a shape with no retraining. We extend it to Poisson via a classical decomposition, with an auxiliary network amortizing the source-induced correction and avoiding the singular volume quadrature that breaks direct evaluation. At inference, new boundary values and new sources both yield PDE solutions by re-integration against the fitted kernel and lift, with no retraining. NHMO improves over four prior baselines on the MCB-B 3D variable-shape Poisson benchmark across all five categories, and is competitive with major neural-operator baselines on a controlled 2D testbed.
Abstract: While the quadratic sequence-length bottleneck of transformers has fueled a resurgence in recurrent models, effectively capturing complex dynamics requires architectures that balance efficient training with highly expressive latent states. Echo State Networks (ESNs) offer a compelling approach by utilizing fixed recurrent weights to circumvent backpropagation through time, enabling a closed-form training solution. However, achieving the expressivity needed for complex tasks demands large reservoirs, exposing an \mathcalO(N^2) state-update bottleneck that prevents ESNs from matching the scale of contemporary recurrent models. To address this limitation, we introduce Frequency Domain Reservoir Computing (FRESCO), an ESN architecture operating entirely in the frequency domain while avoiding domain-shift overheads to achieve \mathcalO(N) complexity for dense, non-linear recurrent updates. By employing a novel dimensional zero-padding input embedding, a packed frequency-domain readout, and a natively applied frequency-domain non-linearity, FRESCO drastically reduces computational costs and energy consumption of training and inference. Furthermore, FRESCO matches the state-of-the-art predictive performance on memory benchmarks, sequential classification, and multivariate long-horizon forecasting, offering a scalable path forward for dense recurrent architectures.
PaperID: 860, Poster
Authors: Frederic Koriche, Louenas Bounia
Abstract: Interpreting the predictions of complex black-box classifiers remains a central challenge in explainable artificial intelligence. While local explanations clarify individual predictions, there is a significant need for global probabilistic explanations that capture a model's overall behavior across the input distribution. We define such an explanation as a small subset K of features, whose relevance is measured by the probability that the classifier assigns identical labels to two independently sampled inputs that agree on K. Based on this notion, we investigate the task of identifying maximally relevant global explanations under a cardinality constraint, focusing on product distributions with full support on discrete feature spaces. We introduce a spectral explainer that leverages membership queries to the black-box classifier and employs Fourier-analytic, junta-learning methods to produce probably approximately correct (PAC) explanations. Our algorithm is fixed-parameter tractable with respect to the explanation size limit, maximum feature cardinality, and desired accuracy. Experimental evaluations on synthetic and real-world datasets show that the spectral explainer provides highly relevant global explanations and outperforms heuristic greedy methods as the explanation size increases.
PaperID: 861, Poster
Abstract: Constrained Markov decision processes (CMDPs) extend standard MDPs by incorporating a cost function and enforcing constraints---most commonly by requiring that the expected cumulative cost over the entire trajectory remains below a prescribed threshold. While this global, expectation-based formulation is natural and widely studied, it is insufficient for many safety- and reliability-critical applications, where guarantees are required not only from the initial state but also for every subprocess starting at any intermediate time step. Motivated by this need, we introduce \emphsubprocess-constrained MDPs (S-CMDPs), in which the expected cumulative cost must satisfy a constraint not only from the initial state, but from every intermediate step onward along the trajectory. We investigate the algorithmic aspects of this model and characterize the limitations of classical solution concepts. In particular, we show that stationary policies can be both suboptimal and computationally intractable to optimize in S-CMDPs. Leveraging a value-set iteration framework, we further demonstrate that history-dependent policies surprisingly overcome both obstacles. Building on this insight, we develop a polynomial-time algorithm that computes the action distribution of a (near-)optimal history-dependent policy for any given history.
Abstract: Self-evolution offers a promising path for improving reasoning models without relying on intensive human annotation. However, extending this paradigm to video understanding remains underexplored and challenging: videos are long, dynamic, and redundant, while the evidence needed for reasoning is often sparse and temporally localized. Naively generating difficult question-answer pairs from full videos can therefore produce supervision that appears challenging but is weakly grounded, relying on static cues or language priors rather than temporal evidence. In this work, we argue that the key bottleneck of video self-evolution is not difficulty alone, but grounding. We propose Video-Zero, an annotation-free Questioner--Solver co-evolution framework that centers self-evolution on temporally localized evidence. The Questioner discovers informative evidence segments and generates evidence-grounded questions, while the Solver learns to answer and align its predictions with the supporting evidence. This closes an iterative loop of evidence discovery, grounded supervision, and evidence-aligned learning. Across 13 benchmarks spanning temporal grounding, long-video understanding, and video reasoning, Video-Zero consistently improves multiple video VLM backbones, demonstrating the effectiveness and transferability of evidence-centered self-evolution. Code and models will be made publicly available.
PaperID: 863, Poster
Abstract: Diffusion and flow policies have become promising tools for robotic and sequential decision-making systems. Many stable and effective policy optimization methods require evaluating the likelihood or likelihood ratio of an action under the current policy. However, these quantities are generally intractable for diffusion and flow models, which prevents widely used policy update algorithms from being applied to generative policies. We propose Likelihood-Free Generative Policy Optimization (LFGPO), a framework that avoids exact likelihood evaluation by separating policy improvement from generative model training. In the first stage, LFGPO uses a lightweight neural network to represent the likelihood ratio that carries the policy improvement signal. In the second stage, it distills the ratio into diffusion or flow policies through a weighted score- or velocity-matching objective. Our framework incorporates a large family of Reinforcement Learning (RL) algorithms, including PPO and GRPO, to be applied to diffusion and flow policy optimization. We show that the solution to the weighted score matching objective solves a stochastic optimal control problem. Across MuJoCo continuous-control benchmarks, LFGPO delivers strong and stable improvements over competitive diffusion- and flow-policy baselines, showing a practical route to RL fine-tuning for generative policies.
PaperID: 864, Poster
Abstract: Concept erasure aims to remove a target attribute from a representation while preserving the other information encoded in it. This is difficult beyond the linear setting: a target signal hidden from one probe may remain recoverable by a fresh nonlinear probe, while unconstrained nonlinear updates may remove the target by pushing representations off the manifold of natural hidden states. We propose the (MRH): natural hidden states concentrate near a structured, lower-dimensional manifold, so surgical erasure should act along local manifold degrees of freedom rather than arbitrary ambient directions. We operationalize this hypothesis with the TanGCE family of erasure methods. The core method estimates each representation's local tangent space from nearest-neighbor secants, projects a nonlinear concept-scorer gradient onto that tangent, and applies a per-sample trust-region step; TanGCE+ and TanGCE++ prepend closed-form first- and second-moment erasers before the same manifold-constrained loop. Across 119 settings spanning 13 language models, three NLP concepts, and 40 CelebA-CLIP attributes under two control regimes, TanGCE++ removes more target signal than the strongest published baseline at matched control-damage budgets under one fixed hyperparameter recipe. Applying the same manifold-constrained loop after prior erasers also consistently reduces residual nonlinear leakage without leaving the control budget, empirically supporting MRH's operational prediction that manifold-constrained edits are a useful inductive bias for surgical concept erasure.
Abstract: Decision Transformer performance degrades on long rollouts because the conditioning context drifts out of the training distribution. We show that this drift is visible through the model's own next state prediction error, which rises during rollout and stays elevated, giving a direct signal of when context has become unreliable. We introduce Trust Guided Decision Transformer (TGDT), which selects context before applying value guidance. At each step, TGDT evaluates several recent context suffixes using rolling next state prediction error, calibrated against held out offline data via split conformal prediction. It keeps only suffixes whose error stays within the calibrated threshold, then uses a frozen critic to choose the highest value action among the trusted suffixes. This reverses the order used by value only elastic selection, where the critic may choose an action generated from a context the model itself has flagged as unreliable. Experiments on D4RL navigation and locomotion tasks show that state prediction, critic guidance, and hard context reset each solve only part of the problem. TGDT reduces persistent high error runs and improves return over vanilla Decision Transformer, reset based context control, and value only context selection.
Authors: Bengusu Nar, Jiguang Li, Veronika Rockova, Panos Toulis
Abstract: In networks, effective dynamic treatment allocation requires deciding both whom to treat and also when, so as to amplify policy impact through spillovers. An early intervention at a well-connected node can trigger cascades that change which nodes are worth targeting in the next period. Existing treatment strategies under network interference are largely static while dynamic treatment frameworks typically ignore network structure altogether. We integrate these perspectives and propose Q-Ising, a three-stage pipeline that (i) estimates network adoption dynamics via a Bayesian dynamic Ising model from a single observed panel, (ii) augments treatment adoption histories with continuous posterior latent states, and (iii) learns a dynamic policy via offline reinforcement learning. The Bayesian mechanism enables uncertainty quantification over dynamic decisions, yielding posterior ensemble policies with interpretable spillover estimates. We provide a finite-sample regret upper bound that decomposes into standard offline-RL uncertainty, network abstraction error, and first stage error in Ising state estimation. We apply our method to data from Indian village microfinance networks and synthetic stochastic block models under simulated heterogeneous susceptible-infected-susceptible (SIS) dynamics and demonstrate that adaptive targeting outperforms static centrality benchmarks.
Abstract: In recent years, large language models (LLMs) have demonstrated remarkable capability to generalize across diverse natural language processing tasks, inspiring the development of graph foundation models (GFMs) for large-scale pre-training. However, unlike language models with explicit token units, graphs lack a well-defined unit for generalization, making it challenging to design effective pre-training strategies. In this work, we propose REEF, a novel GFM framework that leverages relation tokens as the fundamental units. We construct a vocabulary of relation tokens to encode relational information within graphs. To accommodate diverse relations, we introduce two hypernetworks that adaptively generate the parameters of aggregators and classifiers in graph neural networks based on relation tokens. In addition, we design another hypernetwork to construct dataset-specific projectors and incorporate a dataset-level feature bias into the initial node representations, enhancing flexibility across different datasets with the same relation. Extensive experiments demonstrate that REEF consistently outperforms existing methods in both pre-training and transfer learning, highlighting its potential as a general-purpose graph foundation model. Our code is publicly available at https://anonymous.4open.science/r/REEF-16CF/.
PaperID: 868, Poster
Abstract: Given a set of LEGO parts and an image of an assembled object, our task is to infer the pose of each part to assemble the object. In contrast with traditional 3D reconstruction, doing so requires satisfying discrete physical constraints, which provide structure but render the problem combinatorial and highly sensitive to small pose errors. To learn a prior over how parts fit together, we train a set-conditioned SE(3) flow model that maps noisy poses to valid assemblies. We then finetune this model to inject image conditioning. However, while the learned flow captures global structure, it does not enforce connectivity. To address this, we introduce an analytic connector-field guidance to pull predictions toward the connection-consistent manifold at test time. Through this, we demonstrate improved validity and reconstruction performance, scaling with model size, pretraining, and conditional guidance. We will make our models and data available for research purposes.
PaperID: 869, Poster
Authors: Antonio R Linero, Jared Murray, Soumyabrata Bose
Abstract: Distribution regression, where the goal is to predict a scalar response from a distribution-valued predictor, arises naturally in settings where observations are grouped and outcomes depend on group-level characteristics rather than on individual measurements. We introduce DistBART, a Bayesian nonparametric approach to distribution regression that models the regression function as a linear functional with the Riesz representer assigned a Bayesian additive regression trees (BART) prior. We argue that shallow decision tree ensembles encode reasonable inductive biases for tabular data, making them appropriate in settings where the functional depends primarily on low-dimensional marginals of the distributions. We show this both empirically on synthetic and real data and theoretically through an adaptive posterior concentration result. We also establish connections to kernel methods, and use this connection to motivate variants of DistBART that can learn nonlinear functionals. To enable scalability to large datasets, we develop a random-feature approximation that samples trees from the BART prior and reduces inference to sparse Bayesian linear regression, achieving computational efficiency while retaining uncertainty quantification.
Abstract: Shapley value and its priority-aware extensions are widely used for valuation in machine learning, but existing methods require pairwise priority to be binary and acyclic, a restriction spectacularly violated in real-data examples such as aggregated human preferences and multi-criterion comparisons. We introduce the generalized priority-aware Shapley value (GPASV), a random order value defined on arbitrary directed weighted priority graphs, in which pairwise edges penalize rather than forbid order violations. GPASV covers a range of classical models as boundary cases. We establish GPASV through an axiomatic characterization, develop the associated computational methods, and introduce a priority sweeping diagnostic extending PASV's. We apply GPASV to LLM ensemble valuation on the cyclic Chatbot Arena preference graph, illustrating that priority-aware valuation is not a one-button operation: different balances of pairwise graph priority versus individual soft priority produce substantively different valuations of the same data.
PaperID: 871, Poster
Authors:
Guojian Zhan, Feihong Zhang, Likun Wang, Xiangteng Zhang, Letian Tao, Tianyi Zhang, Wenxin Zhao, yinuo Wang, Tianze Zhu, Jingliang Duan, Yang Guan, Shengbo Eben LiAbstract: Generative policies offer a promising alternative to Gaussian policies in reinforcement learning (RL) due to their ability to represent expressive and multimodal action distributions. However, integrating such policies into likelihood-based policy optimization remains challenging, since accurate and differential action likelihoods are intractable for standard generative policies. In this paper, we propose \emphSelf-Adjoint Flow Policy Optimization (\textscSAFlow), a likelihood-tractable generative policy with self-adjoint structure. \textscSAFlow parameterizes the policy as an invertible flow in a doubled action space and uses a Verlet-style self-adjoint composition for generation. This structure makes the inverse map obtainable by stepsize reversal, aligning forward sampling and backward likelihood evaluation under the same time-symmetric numerical rule. It also provides a second-order approximation accuracy to the learned continuous flow, improving the numerical consistency of likelihood and policy-ratio computation. We instantiate \textscSAFlow within an on-policy optimization framework and evaluate it on eight IsaacLab continuous-control tasks. \textscSAFlow consistently outperforms Gaussian-policy baselines and recent generative-policy methods, highlighting the value of self-adjoint numerical structure for likelihood-based generative policy optimization.
Abstract: Algorithms with predictions, or learning-augmented algorithms, has proved to be an extremely useful paradigm for combining machine learning with traditional algorithms. One of the textbook settings for this is searching a sorted array. Without a prediction classical binary search takes O(\log n) queries, while with a prediction we can use "doubling binary search" to find the target key using O(\log \eta) queries, where \eta is the error of the prediction measured as the absolute value of the difference between the true location and the predicted location. Since an array is just a path graph, in this paper we ask whether similar bounds can be achieved for search on even slightly more general graphs: trees. We show first that the high-level answer is "no": there is no search algorithm that uses O(\log \eta) queries, where now \eta is the graph distance between the predicted location and the true location. However, as our main result, we show that such bounds can be achieved on trees which are "path-like" in that they have low pathwidth. In particular, we prove that there is a search algorithm which uses at most O(k \log \eta) queries, where k is the pathwidth of the tree. We also prove a lower bound showing that our algorithm has existentially optimal query complexity. Finally, we show experimentally, on real-life inputs, that our algorithm has query complexity which is notably better than the simple non-prediction based algorithm.
Abstract: Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models for language modeling, yet the effective design of transformer architectures for MDMs remains underexplored. In this paper, we show that selectively looping the early-middle transformer layers significantly improves both training efficiency and model performance in MDMs. We call this approach LoopMDM (Looped Masked Diffusion Model), which brings two key benefits: looping layers at training-time yields a depth-scaling effect without adding parameters, while varying the number of loops at inference-time enables flexible compute scaling. Despite the simplicity, the results are striking: across multiple pre-training corpora, LoopMDM matches the performance of same-size MDMs with up to 3.3× fewer training FLOPs, while its final performance outperforms them on various reasoning benchmarks, including up to +8.5 points on GSM8K. It even surpasses deeper non-looped MDMs trained with comparable per-step compute, indicating that selective looping is more effective than naive depth scaling. Furthermore, LoopMDM can scale inference-time compute by increasing the number of loops. Adaptively adjusting the number of loops throughout the sampling process further yields additional gains in compute efficiency while maintaining performance. Lastly, with attention analysis, we provide evidence that looping is effective in MDMs by promoting interactions among masked positions. Our code and weights will be publicly released.
PaperID: 874, Poster
Abstract: We study Metrical Task Systems, a broad class of problems in online algorithms, under a model where input requests are generated by a stochastic process. We first consider the case of requests drawn IID from an unknown distribution, and then extend the model to requests generated by a finite-state Markov chain, where the algorithm observes both the current request and the current state of the chain before making its next move. This formulation captures practical settings in which requests are generated by an evolving environment. In both settings, we show how to find in polynomial time an algorithm whose expected cost per step is arbitrarily close to the cost of the best possible online algorithm. Our theoretical results are complemented by a brief empirical evaluation of our algorithms.
Authors:
Badr-Eddine Marani, Julio Silva-Rodríguez, Ismail Ayed, Maria Vakalopoulou, Stergios Christodoulidis, Jose DolzAbstract: Deep neural networks have achieved remarkable success across a variety of tasks, yet they often suffer from unreliable probability estimates, leading to overconfident predictions. Conformal Prediction (CP) offers a principled framework for uncertainty quantification, yielding prediction sets with rigorous coverage guarantees. Recent work has proposed integrating conformal objectives into training, optimizing for overall set size. Nevertheless, shaping the prediction sets in a class-conditional manner is not straightforward and typically requires prior knowledge of the data distribution. While class-conditional efficiency has been studied at the post-hoc calibration level (e.g., via class-conditional scoring or calibration partitioning), no prior conformal training method explicitly addresses it. To fill this gap, we introduce Class Adaptive Conformal Training (CaCT), which formulates conformal training as an augmented Lagrangian optimization problem that adaptively shapes prediction sets class-conditionally without relying on parametric distributional assumptions, beyond a user-specified global coverage constraint. Experiments on multiple benchmark datasets, including standard and long-tailed image recognition as well as text classification, demonstrate that CaCT consistently outperforms prior conformal training methods, reducing prediction set size while maintaining the desired coverage guarantees.
PaperID: 876, Poster
Authors: Omar Ghezzi, Alessandro D'Amelio, Vittorio Cuculo, Rita Cucchiara, Giuseppe Boccignone
Abstract: Perceptual decisions unfold in time: noisy evidence is integrated until a commitment threshold is reached, jointly producing a choice and a reaction time. Classical race models capture this mechanism through simple parametric drift and diffusion, but cannot represent how rich, high-dimensional stimuli shape accumulation on individual trials. We introduce the Neural Race Model (NRM), a stimulus-conditional neural stochastic differential equation in which one accumulator per alternative races towards a learnable, stimulus-dependent threshold. Drift, diffusion, and threshold are parameterised by neural networks conditioned on a stimulus embedding, allowing both accumulation dynamics and decision urgency to vary across trials. The model is trained end-to-end on the joint reaction-time-and-choice likelihood via a smooth surrogate for the discontinuous first-passage event, combined with Monte Carlo trajectory averaging and gradient propagation through the stochastic adjoint. The surrogate recovers the classical hard threshold in the appropriate limit and is removed at test time. Each component admits a natural correspondence with elements of the primate decision-making circuit, preserving the mechanistic interpretability of classical models. We evaluate the NRM on multiple perceptual decision-making benchmarks against classical sequential-sampling and image-computable reaction-time models. Across tasks and metrics, the NRM tracks the human noise ceiling more closely than all baselines, demonstrating that stimulus-conditional stochastic accumulation and end-to-end differentiability can be achieved within a single, principled framework. Code will be publicly released.
PaperID: 877, Poster
Abstract: Vector quantization is a fundamental primitive for scalable machine learning systems, enabling memory-efficient storage, fast retrieval, and compressed inference. Recent rotation-based quantizers such as EDEN, RabitQ, and TurboQuant have introduced strong guarantees and empirical performance, but the surrounding comparisons have been difficult to interpret because they rely on different distortion criteria, probability regimes, and implementation assumptions. As our first contribution, we provide a unified theoretical comparison of these methods and show that their relative advantages are criterion-dependent rather than absolute: TurboQuant is favorable for MSE distortion, EDEN is effective for expected inner-product distortion, and RabitQ provides strong high-probability control. This comparison clarifies the design principles behind each advantage and shows that no existing method uniformly dominates across the relevant measures. As our second contribution, we introduce ), a new rotation-based block quantization algorithm designed around the spherical geometry of randomly rotated vectors. Unlike coordinate-wise quantizers, quantizes blocks on the sphere, preserving the geometry of rotated embeddings more faithfully. We prove that this block-spherical design improves expected distortion performance, including both reconstruction MSE and expected inner-product distortion. Experiments on real embedding data support these theoretical improvements.
Authors:
Yuqing Kong, Mingyu Song, Yizhou Wang, Yifan WuAbstract: We propose a label-free post-processing framework that improves a strong but miscalibrated primary model using a weaker yet better-calibrated reference. Our key insight is that strict improvement is possible if and only if the two models are not mutually calibrated, meaning there does not exist joint distribution over their predictions and outcomes such that both predictions are simultaneously calibrated. We formalize this condition and connect it to no-arbitrage results from economics. Under such condition, we develop an efficient post-processing algorithm of the strong model's outputs based on Bregman projection, with a strict worst-case performance improvement guarantee. Experiments on representative LLMs across varying scales demonstrate the effectiveness of our method, reducing the ECE of the primary model by over 40% on common benchmarks.
PaperID: 879, Poster
Abstract: Many scientific responses take values in Riemannian manifolds, with predictors that are scalar- and/or manifold-valued. A regression analysis in such settings must support not only point prediction but also formal hypothesis testing and effect-size estimation for individual coefficients. Yet no current manifold-regression framework simultaneously accommodates both predictor types, and returns per-coefficient hypothesis tests. We propose Riemannian Ordinary Least Squares (ROLS), a tangent-space conditional-mean model that serves as the OLS analogue for this setting. ROLS is well defined under an explicit injectivity condition that keeps the logarithm maps single-valued, with a cut-locus diagnostic for branch-sensitive cases outside it. Within the local geometric setting it admits a closed-form estimator, an O(n^-1) excess prediction-error bound with explicit curvature bias, and asymptotic t- and Wald tests for individual coefficients. Across simulations with synthetic data and three case studies with real data, ROLS is competitive with the strongest prediction baselines and identifies scientifically meaningful predictors, providing for manifold-valued regression the coefficient-level inference that OLS provides in the Euclidean case.
Abstract: (SCM) estimates causal effects in panel data with a single-treated unit by constructing a counterfactual outcome as a weighted combination of untreated control units that matches the pre-treatment trajectory. In this paper, we introduce the method, a new two-stage estimator that directly estimates the counterfactual outcome. Specifically, our TSC method (1) yields a targeted debiasing estimator, in the sense that the targeted updating refines the initial weights to produce more stable weights; and (2) ensures that the final counterfactual estimation is a convex combination of observed control outcomes to enable direct interpretation of the synthetic control weights. TSC is flexible and can be instantiated with arbitrary machine learning models. Methodologically, TSC starts from an initial set of synthetic-control weights via a one-dimensional targeted update through the weight-tilting submodel, which calibrates the weights to reduce bias of weights estimation arising from pre-treatment fit. Furthermore, TSC avoids key shortcomings of existing methods (e.g., the augmented SCM), which can produce unbounded counterfactual estimates. Across extensive synthetic and real-world experiments, TSC consistently improves estimation accuracy over state-of-the-art SCM baselines.
Abstract: 3D Gaussian Splatting (3DGS) has emerged as a leading representation for real-time novel view synthesis and been widely adopted in various downstream applications. The core strength of 3DGS lies in its efficient kernel-based scene representation, where Gaussian primitives provide favorable mathematical and computational properties. However, under a finite primitive budget, the symmetric shape of each primitive directly affects representation compactness, especially near asymmetric structures such as object boundaries and one-sided surfaces. Recent works have explored more complex kernel distributions, yet they either remain within the elliptical family or rely on hard truncation, which limits continuous shape control and introduces distributional discontinuities. In this paper, we propose Skew-Normal Splatting (SNS), which adopts the Azzalini Skew-Normal distribution as the fundamental primitive. By introducing a learnable and bounded skewness parameter, SNS can continuously interpolate between symmetric Gaussians and Half-Gaussian-like shapes, enabling flexible modeling of both sharp boundaries and interior regions. Moremover, SNS preserves analytical tractability under affine transformations and marginalization. This property allows seamless integration into existing Gaussian Splatting rasterization pipelines. Furthermore, to address the strong coupling between scale, rotation, and skewness parameters, we introduce a decoupled parameterization and a block-wise optimization strategy to enhance training stability and accuracy. Extensive experiments on standard novel-view synthesis benchmarks show that SNS consistently improves reconstruction quality over Gaussian and recent non-Gaussian kernels, with clearer benefits on sharp boundaries and thin or one-sided structures. Codes and models will be released.
Authors: Toan Nguyen, Yang Liu, Trung Le, Celso de Melo, Flora Salim
Abstract: We argue that forgetting is not confined to continual learning but is a general optimization phenomenon: during standard training, dominant mini-batch gradients suppress rare but useful update directions, causing short-term forgetting at every step. When such knowledge is never revisited, these losses compound into long-term forgetting—the classical failure mode of continual learning. We introduce FOGO, a scalable optimizer that continuously detects and resolves gradient interference across both regimes. FOGO spectrally orthogonalizes momentum updates to prevent dominant directions from monopolizing optimization, then stores representative past directions in a compact codebook memory built on random projection, where pairwise distances are provably preserved in low-dimensional space. At each step, conflicts between the current update and stored directions are resolved via lightweight orthogonal correction and lifted back through a proximal step, with minimal overhead and no data storage. Across class-imbalanced classification, continual visual learning under domain and class shifts, continual fine-tuning of LLaVA-7B, and GPT-2 pretraining, FOGO consistently improves convergence and knowledge retention, outperforming Adam and Muon.
Abstract: Desktop GUI agents operate under partial observability: visually similar screens can correspond to different underlying workflow states, so locally plausible actions can lead to sharply different outcomes. We frame this as a problem of computer/OS state exploration, where effective behavior requires both expanding the reachable frontier and reducing ambiguity before committing. We present ScreenSearch, a system that combines structural screen retrieval and deduplication with an ambiguity-aware PUCT graph-bandit for large-scale desktop exploration. The retrieval layer converts UIA trees into location-aware structural features, indexes related screens through sparse token search and metadata filters, and maintains a shared deduplicated state graph across VM workers. On top of this graph, we define a scalable ambiguity signal based on matched-action outcome dispersion. If similar screens produce different next states under the same action signature, the state should be probed further rather than treated as resolved. We use this signal together with frontier rewards to drive large-scale exploration and replay-start policy evaluation over the shared graph. Across 11 desktop applications, ScreenSearch collects over 1M screenshots and over 30K deduplicated states, yielding large exploration corpora with substantial cross-application and within-application diversity. On a fixed replay-start slice, we observe a clear novelty--ambiguity trade-off: some policies reduce ambiguity quickly while discovering little frontier. Ambiguity reduction alone is therefore not a sufficient exploration objective. Appendix ablations show that stronger proposal priors can materially improve unique-state discovery during corpus building. These results suggest that state identity, proposal quality, and ambiguity-aware search all matter when deciding when to probe and when to commit.
Abstract: Compositional energy-based models can generalize to larger combinatorial reasoning problems by reusing a learned factor energy across many local constraints. In our paper, we show that a key bottleneck in compositional reasoning is not composition itself, but the non-convex geometry of the learned energy landscape. To solve this problem, we introduce Convex Compositional Energy Minimization (CCEM), a framework that parameterizes each factor with an input-convex neural network and optimizes the composed energy over a tight convex relaxation of the feasible set. Because convexity is preserved under summation, the global relaxed objective remains convex, enabling deterministic projected first-order optimization. CCEM is trained in two stages: factor-level contrastive learning to shape local energy basins, followed by end-to-end refinement through an unrolled projected solver. Our experiments show that our models trained on small subproblems or a single problem size transfer to larger instances without retraining.
Abstract: Probing studies what information is encoded in a frozen LLM's layer representations by training a lightweight predictor on top of them. Beyond analysis, probes are often used operationally in probe-then-steer pipelines: a learned concept vector is extracted from a probe and then injected via additive activation steering by adding it to a layer representation during the forward pass. The effectiveness of this pipeline hinges on estimating concept vectors that are accurate, directionally stable under ablation, and inexpensive to obtain. Motivated by these desiderata, we propose RAPTOR (Ridge-Adaptive Logistic Probe), a simple \ell_2-regularized logistic probe whose validation-tuned ridge strength yields concept vectors from normalized weights. Across extensive experiments on instruction-tuned LLMs and human-written concept datasets, RAPTOR matches or exceeds strong baselines in accuracy while achieving competitive directional stability and substantially lower training cost; these quantitative results are supported by qualitative downstream steering demonstrations. Finally, using the Convex Gaussian Min-max Theorem (CGMT), we provide a mechanistic characterization of ridge logistic regression in an idealized Gaussian teacher-student model in the high-dimensional few-shot regime, explaining how penalty strength \lambda mediates probe accuracy and concept-vector stability, yielding structural predictions that qualitatively align with trends observed on real LLM embeddings.
Authors: Marcus Häggbom, Viktor Nilsson, Pierre Nyquist, Joakim Andén
Abstract: Recent advances in generative modeling have enabled the efficient computation of Schrödinger bridges (SB) in high-dimensional settings by leveraging partially simulation-free training methods inspired by flow matching. However, these have not covered SBs with reflecting dynamics, a useful model choice with built-in guarantees that generated samples stay in the data domain. Existing alternatives for reflected SBs instead rely on more complex training based on forward--backward SDE theory, requiring expensive higher-order derivatives and sampling entire paths during training. In this article, we introduce a partially simulation-free framework that allows reflected SBs to be trained similarly to flow matching, using a new sampling method and regression target. We demonstrate our results by coupling pairs of well-known high-dimensional image datasets. Using reflected dynamics incurs negligible additional wall-clock time during both training and inference while maintaining or slightly improving generative performance.
PaperID: 887, Poster
Abstract: Financial markets are complex, noisy environments which present unique challenges for representation learning and world modeling. In this study, we systematically explore the application of supervised and self-supervised representation learning methods to financial markets data and their ability to learn world models of financial markets. To support this, we assemble and release Market-1T a dataset of more than one trillion observations spanning all US equities from 2008 through 2025. We then develop a domain-specific data augmentation and apply a variety of modern SSL objectives including Joint Embedding Predictive Architecture (JEPA), Masked Autoencoder (MAE), and DINO. Our results highlight that supervised methods are still dominant in this domain when applied directly to core quantitative finance tasks: predicting changes in prices, volatility, and transactions cost. Further, when supervised models are trained across tasks they are able to leverage complementarities which enhance cross-task performance. While performance of SSL-based methods on these core tasks lags behind, we find that they better capture latent structure---such as time and assets effects---which supervised models miss. Our results suggest current SSL objectives capture broad features of markets but further work is needed to close the gap between supervised methods.
PaperID: 888, Poster
Authors: Kunzhe Song, Jialuo Du, Jingkai Lin, Huacheng Zeng
Abstract: Accurate RF odometry remains difficult due to multipath interference, geometric degeneracy, and dynamic clutter. Existing approaches mitigate these challenges by relying on auxiliary sensors or ground-truth pose supervision, which limits scalability and deployability. We present RFWalk, a self-supervised, RF-only odometry framework that learns cross-frame correspondence without any pose annotations or auxiliary sensors. Our key insight is to cast each range-azimuth cell as a node in a space-time graph and recover inter-frame motion by learning a soft transition matrix through cycle-consistent contrastive random walks. To resolve geometric degeneracy in feature-poor scenes, we further introduce a physics-informed Doppler consistency loss that enforces agreement between the correspondence-induced displacement field and measured radial velocities. Experiments on challenging scenes show that RFWalk achieves competitive accuracy in static environments and substantially outperforms existing approaches in dynamic scenes.
PaperID: 889, Poster
Abstract: We show that a softmax-attention-only transformer trained on in-context linear regression discovers an iterative symmetric preconditioning primitive. The trained model is used only for algorithm discovery: the final method, RAPID, is an explicit preconditioner and uses no learned component at solve time. To extract this mechanism, we apply a layerwise symmetry intervention strategy which reduces the trained layers to a two-scalar update. The current Gram geometry determines the attention scores, and the resulting row update evolves the Gram matrix by congruence, reducing its condition number. We turn this primitive into RAPID, an anytime preconditioner for explicit SPD systems Ax=b: every prefix constructs a valid factor P such that PAP^\top can be used to improve the runtime of iterative solvers such as conjugate gradients. RAPID uses fast randomized Walsh--Hadamard mixing and sparse row-min corrections, giving \widetilde O(d^2)-time dense iterations. For a Haar-mixing idealization, we prove that RAPID reaches a target condition number in \widetilde O(d\log\kappa(A_0)) expected iterations. Empirically, RAPID is fast and versatile: it reaches useful condition numbers quickly, accelerates end-to-end CG solves, and improves conditioning across all tested spectral families and real-world SPD problems.
PaperID: 890, Poster
Abstract: Variance reduction techniques improve the efficiency of drawing conclusions from data. Recent proposals that incorporate machine learning predictions into estimators based on weighted sampling use the Poisson Sampling (PS) algorithm, which independently flips a biased coin at each unit and includes all heads in the sample. A fundamental flaw with PS is that it has a random sample-size, which presents practical challenges and introduces variance into the estimator. We show that a fixed-size alternative, known in survey-sampling as Conditional Poisson Sampling (CPS), is a drop-in replacement for PS on general multivariate Z-estimation problems, which includes convex M-estimation considered in prior work. Notably, we show that the asymptotic variance of CPS is no worse than that of PS, and can be substantially better in realistic scenarios where predictions are imperfect. We illustrate these advantages on both synthetic and real datasets. In light of our contributions, there is no need to suffer a variable sample-size.
Authors: Trevor Harris
Abstract: Conformal prediction provides a distribution-free framework for uncertainty quantification via prediction sets with exact finite-sample coverage. In low dimensions these sets are easy to interpret, but in high-dimensional or structured output spaces they are difficult to represent and use, which can limit their ability to integrate with downstream tasks such as sampling and probabilistic forecasting. We show that any sufficiently regular differentiable nonconformity score induces a deterministic flow on the output space whose trajectories converge to the boundary of the corresponding conformal prediction set. This leads to a computationally efficient, training-free method for sampling conformal boundaries in arbitrary dimensions. Mixing across confidence levels yields conformal predictive distributions whose quantile regions coincide with the empirical conformal prediction sets. We provide an approximation bound decomposing CPD predictive error into score-induced distortion, base-measure quality, and gradient flow-induced distortion. We evaluate the approach on PDE inverse problems, precipitation downscaling, climate model debiasing, and hurricane trajectory forecasting.
Abstract: Many learning problems require predicting how populations evolve under an unknown transformation. A natural representation for such populations is a probability measure, with point clouds as a key example. In this work, we study the problem, in which one seeks to learn a map between probability measures from a finite collection of observed input--output pairs. In contrast to classical regression, where individual samples are transformed independently, M2M regression treats entire distributions as the data points. This perspective is vital in certain scientific applications, for example, cellular and molecular biology, where cells are known to evolve not as independent data points but as a distribution. However, few existing approaches address the problem of M2M regression with sufficient expressivity and scalability. We present a formalization of M2M regression and introduce two simple, expressive, and scalable approaches to learn such operators: transformers as M2M velocity fields. Our approach leverages the natural measure-dependent and mean-field structure of transformers to learn nonlinear M2M maps on the space of probability distributions. We illustrate the effectiveness of our proposed method to generalize to measures on synthetic experiments, canonical interacting particle systems, and a large-scale patient-derived organoid dataset for predicting treatment response in colorectal cancer.
Authors: Varun Gupta, Vijay Kamble
Abstract: We study a canonical multi-task demand-learning problem motivated by retail pricing, where a firm seeks to estimate heterogeneous linear price-response functions across multiple decision contexts. Each context is described by rich covariates but exhibits limited price variation, motivating transfer learning across tasks. A central challenge in leveraging cross-task transfer is endogeneity: prices may be arbitrarily correlated with unobserved task-level demand determinants across tasks. We propose a new meta-learning framework that identifies the conditional mean of task-specific causal demand parameters given a subset of task-specific observables despite such confounding, assuming that each task contains at least two distinct locally exogenous price points. This subset is carefully designed to include all of the prices to address cross-task confounding, while masking two demand outcomes that provide randomized supervision to address identifiability issues arising from the inclusion of all prices. We show that this information design is maximally uniformly valid, in that any refinement of the conditioning set that reveals withheld-outcome information is not guaranteed to identify the conditional mean causal target. We validate our method on real and synthetic data, demonstrating improved recovery of demand responses relative to standard transfer-learning baselines.
PaperID: 894, Poster
Abstract: Risk-sensitive optimal stopping arises wherever a single bad outcome dominates a sequential decision, such as option exercise, sequential medical treatment, and safety-critical control. The natural objective is the conditional value-at-risk (CVaR) of the stopped reward. We develop the first theoretical framework and algorithm for risk-sensitive deep optimal stopping. Together they resolve all four structural problems of risk-sensitive RL: state augmentation, score-function pathology, blindness to success, and bootstrapping circularity. We isolate three structural properties of optimal stopping: single-point reward, action-independent dynamics, and sum-zero gradient. They yield Markov optimality on the original state space and uniform boundedness of the per-step pathwise gradient at the deterministic-policy boundary, resolving the first two. From these, we derive DIOS (Distribution-Informed Optimal Stopping), a single-time-scale CVaR estimator on the Rockafellar-Uryasev envelope that resolves the remaining two. A soft penalty replaces the hard tail indicator, and a tail-level annealing schedule initializes training in the risk-neutral regime, where no trajectory is gated out. The inner quantile admits a closed-form update without an auxiliary critic. DIOS admits a finite-time best-iterate stationarity guarantee under standard CVaR non-degeneracy, the first explicit rate for CVaR optimal stopping. On Bermudan max-call options, DIOS achieves the highest mean CVaR and the lowest wall time against five representative CVaR baselines.
Authors: Yuxin Du, Tao Lin, Zile Zhong, Runting Li, Xiyao Chen, Jiting Liu, Chenglin Liu, Yingcong Chen, Yuqian Fu, Bo Zhao
Abstract: Monocular depth foundation models generalize well across scenes, yet they are typically optimized with uniform pixel-wise objectives that do not distinguish user-specified or task-relevant target regions from the surrounding context. We therefore introduce Focusable Monocular Depth Estimation (FDE), a region-aware depth estimation task in which, given a specified target region, the model is required to prioritize foreground depth accuracy, preserve sharp boundary transitions, and maintain coherent global scene geometry. To prioritize task-critical region modeling, we propose FocusDepth, a prompt-conditioned monocular relative depth estimation framework that guides depth modeling to focus on target regions via box/text prompts. The core Multi-Scale Spatial-Aligned Fusion (MSSA) in FocusDepth spatially aligns multi-scale features from Segment Anything Model 3 to the Depth Anything family and injects them through scale-specific, gated conditional fusion. This enables dense prompt cue injection without disrupting geometric representations, thereby endowing the depth estimation model with focused perception capability. To study FDE, we establish FDE-Bench, a target-centric monocular relative depth benchmark built from image-target-depth triplets across five datasets, containing 252.9K/72.5K train/val triplets and 972 categories spanning real-world and embodied simulation environments. On FDE-Bench, FocusDepth consistently improves over globally fine-tuned DA2/DA3 baselines under both box and text prompts, with the largest gains appearing in target boundary and foreground regions while preserving global scene geometry. Ablations show that MSSA's spatial alignment is the key design factor, as disrupting prompt-geometry correspondence increases AbsRel by up to 13.8%.
Abstract: We introduce a categorical framework for causal abstraction that unifies and extends existing approaches. Modeling causal systems as Markov functors from free Markov categories generated by directed acyclic graphs, we define a causal abstraction as a deterministic natural transformation together with an embedding of restricted free Markov categories. This separates graphical compatibility under interventions from domain-level clustering of variables or values. Our framework provides an explicit characterization of admissible high-level graphs and recovers prior notions such as constructive -abstractions and cluster-DAG abstractions with unobserved confounders. We give concise categorical proofs that interventional distributions factorize over graphical abstractions and that do-calculus applied to the high-level graph yields valid conclusions for the low-level model.
Abstract: Open-ended reasoning and long-form generation tasks lack reliable automatic verification signals for reward-based policy optimization. Rubrics offer a promising alternative, but existing approaches treat them as given artifacts—either hand-crafted or prompt-generated—and often miss the task-specific, knowledge-intensive dimensions that matter most, distorting the reward signal. Our key observation is that : identifying what makes a response correct or insightful requires discovering and synthesizing external knowledge. We propose , a two-stage framework for constructing such rubrics. Stage I elicits domain facts, structural constraints, and failure modes through iterative multi-turn agentic search; Stage II distills this evidence into atomic, independently verifiable constraints for GRPO-based policy optimization. Because the model under training can serve as its own rubric generator, DR-rubric-8B supports without frontier-model assistance. We evaluate on 6 benchmarks spanning agentic research and expert reasoning. Experiments show that DR-Rubric achieves strong competitive performance with only 1K–3K training instances, where GPT-5-generated rubrics particularly benefit breadth coverage on agentic tasks, while bootstrap rubrics exhibit a specialization-to-rebalancing evolution and achieve the best overall performance at the third iteration. Results demonstrate that reframing rubric construction from static evaluation templates into an evidence-driven research process yields more scalable, fine-grained reward signals for open-ended tasks.
PaperID: 898, Poster
Abstract: Graph matching aims to establish node correspondences while preserving both unary appearance compatibility and pairwise structural consistency. Recent deep graph matching methods have achieved strong empirical performance, but many of them either compress edge information into node embeddings before matching or rely on surrogate optimization signals for quadratic objectives. These design choices can weaken structural discrimination and may lead to optimization directions that are not aligned with the relaxed quadratic assignment problem (QAP). We propose Cross-order Consensus Graph Matching (CCGM), a framework that explicitly couples first-order node assignment and second-order edge assignment. CCGM constructs line graphs to convert edge matching into a node matching problem, learns node and edge affinities in parallel, and refines them through a Cross-order Consensus Solver motivated by the exact gradient of the factorized relaxed QAP. We further introduce Cross-order Alignment Regularization to encourage agreement between node-induced edge assignments and edge-induced node support during training. Experiments on standard visual graph matching benchmarks show that CCGM consistently improves matching accuracy over competitive baselines.
PaperID: 899, Poster
Abstract: Preference-based fine-tuning methods such as RLHF and DPO require substantial compute and large preference datasets. They also need direct access to the model parameters which are not provided by many state-of-the art models. Inference-time alignment offers a cost-effective alternative without updating model parameters. However, existing inference-time methods rely on a scalar reward model derived under a Bradley-Terry assumption, which cannot represent general preferences. Following recent work on fine-tuning with generalized preferences, in this work, we initiate the study of inference-time alignment under general preferences. We formulate the problem as obtaining a Nash equilibrium of a two-player zero-sum game between policies. We propose two algorithms: \emphBest-of-Nash (BoN) and \emphNash Mirror Descent (NMD). We prove that both algorithms achieve a duality gap that matches the problem lower bound. Empirically, we implement the two methods on three datasets, which shows that our methods substantially outperform the base policy, converging to the performance of the fine-tuned models. Moreover, our results show that NMD remains robust across the regularization parameter.
PaperID: 900, Poster
Authors: Huibin Li, Chul M Yeum
Abstract: Building upon the real-time rendering capability of 3D Gaussian Splatting (3DGS), 2D Gaussian Splatting (2DGS) achieves improved surface reconstruction and geometric fidelity for novel view synthesis through explicit planar Gaussian primitives. However, the smooth falloff inherent to Gaussian kernels fundamentally limits the reconstruction of high-frequency shape signals, causing over-smoothed edges and loss of fine geometric detail. Prior works have explored alternative kernel functions to address this limitation, yet through systematic evaluation across signal fitting tasks, we find that all existing kernels suffer from the Gibbs phenomenon — producing oscillatory artifacts near sharp discontinuities. This observation motivates a fundamentally different strategy: rather than replacing the kernel, we make it deformable. This study presents , which augments each Gaussian surfel with a learnable, monotonic coordinate remapping in its local tangent plane. With only a single free control point per axis (two extra parameters), the remapping enables asymmetric radial profiles, allowing each primitive to reliably capture sharp edges and irregular structures that symmetric kernels inherently cannot represent. The deformation operates on the 2D tangent plane of 2DGS, preserving the differentiable rasterization pipeline with negligible overhead, and is optimized end-to-end through rendering loss alone. Experiments on MipNeRF 360 and Tanks&Temples demonstrate that our method effectively suppresses Gibbs artifacts at sharp boundaries and achieves state-of-the-art performance among 2DGS-based approaches in PSNR, SSIM, and LPIPS.
PaperID: 901, Poster
Abstract: Transformers have become the dominant architecture for large language models, largely due to the scalability and flexibility of attention, feed-forward layers, residual connections, and normalization. This paper introduces dynamic short convolutions as an additional neural network primitive for improving Transformer architectures. Unlike static short convolutions, dynamic convolutions use input-dependent filters, preserving the locality bias of convolution while increasing expressivity. Motivating experiments show that applying dynamic short convolutions to key, query, and value representations improves performance on challenging associative recall tasks compared with static convolutional variants. Across language-modeling experiments ranging from 150M to 2B parameters, dynamic convolutions consistently outperform standard Transformers and Transformers augmented with static short convolutions. Scaling-law analysis indicates a 1.33× compute advantage over parameter-matched Transformers, while a custom Triton kernel enables efficient execution with minimal end-to-end slowdown. These results suggest that dynamic short convolutions are a scalable, hardware-efficient, and expressive primitive for advancing Transformer-based language models.
PaperID: 902, Poster
Abstract: Submodular clustering provides a flexible framework for modeling coverage and representation quality in metric spaces, capturing a broad range of objectives arising in machine learning, including influence maximization, data summarization, and representation learning. Given a set of points P in a metric space and a decreasing service function \phi, the goal is to select k centers U to maximize total service \sum_p \in P \phi(d(U, p)), where d(U, p) is the distance from p to its nearest center. While this objective is submodular and admits a standard 1 - 1/e approximation via greedy algorithms, it has remained unclear whether and when this barrier can be surpassed. We resolve this question by characterizing exactly which functions \phi permit approximation guarantees beyond 1-1/e. We identify a natural parameter \gamma_\phi capturing how rapidly \phi decreases, and show that, assuming \mathrmP \ne \mathrmNP, a better-than 1-1/e approximation ratio is achievable in polynomial time if and only if \gamma_\phi > 0.
PaperID: 903, Poster
Abstract: We present an encoding method that maps each input token to a randomized sequence of binary values (0s and 1s). By operating at the finest representational granularity and applying compression, this approach reduces the effective vocabulary to just two symbols, yielding a compact, task-agnostic representation. Across eleven main tasks spanning six modalities (human text, genetic/clinical data, programming code, system logs, biological signals, and tabular medical data), it achieves up to two orders of magnitude stronger privacy guarantees under six gradient-based attack settings compared to seven encoder baselines and four defense baselines, including up to 549× improvement on FILM and 62× on TAG. At the same time, it reduces embedding model size by up to 98%, while maintaining competitive accuracy and computational efficiency. It also improves multilingual generalization, achieving up to a 4× increase in accuracy in low-resource settings.
PaperID: 904, Poster
Abstract: Latent 3D representations have significantly expanded the capabilities of modern shape modeling, enabling compact encoding of 3D scenes and supporting powerful 3D generative models. Despite this progress, current representations struggle to capture the fine details of complex objects, fundamentally limited by the capacity of their latents. However, scaling latent capacity at training time is computationally expensive and often infeasible under memory limits, and naively increasing the number of latents at inference fails to make effective use of the added capacity. In this paper, we address both challenges with split-and-scale, a method that effectively expands latent capacity at inference while keeping training cost tractable. Our key observation is that by representing 3D scenes as continuous functions of 3D coordinates, a single pretrained tokenizer can be applied at inference to arbitrarily small sub-regions of a scene, allocating the same latent capacity to a smaller spatial extent and thus capturing far higher detail. We validate this insight by conducting extensive ablations to study the trade-offs between compute, memory, latent capacity, and reconstruction quality, demonstrating our inference-time scaling method achieves quality competitive with much larger models that exceed single-GPU memory limits. We further demonstrate that split-and-scale supports high-quality generative modeling, training an autoregressive 3D generator that performs competitively with state-of-the-art baselines.
PaperID: 905, Poster
Abstract: We propose a contextual cascading bandit model in which the click probability at each position depends on both the displayed item and the topics of previously shown items. This captures how previously displayed topics affect later click probabilities, allowing different topic orders to induce different user responses. Unlike standard cascading bandits, the optimal cascade is no longer determined only by the selected set of items, since the ordering itself affects the expected reward. To address the resulting planning challenge, we reformulate cascade construction as a Markov decision process whose state summarizes the previously selected topics. Based on this formulation, we develop an optimism-based algorithm and prove a regret upper bound of \widetilde\mathcal O(\bar p^\fracK-12d\sqrtT), where d is the number of unknown parameters, T is the horizon, K is the cascade length, and \bar p<1 upper bounds the no-click probability of an examined item. A key implication is that the regret decreases with the cascade length K, a phenomenon that had not been established even in order-insensitive contextual cascading bandits. We further provide a local regret lower bound showing that this decreasing dependence on K is intrinsic, and validate our theory through experiments with topic-order-dependent click probabilities.
PaperID: 906, Poster
Abstract: Dynamic mechanism design provides a principled framework for coordinating self-interested agents whose private information evolves over time, but history dependence and dynamic incentive constraints make general analytical solutions difficult. Using a dynamic revelation principle, we formulate optimal dynamic mechanism design as a bilevel optimization problem over parameterized direct mechanisms: the principal optimizes the mechanism, while agent-side unilateral deviation problems verify incentive compatibility. We model each agent’s problem as a parameterized POMDP, and use stochastic natural policy gradient to solve these POMDPs with finite-sample error bounds. These policies serve as oracle outputs for the principal’s problem, yielding first-order optimization procedures with convergence guarantees that approximate stationary solutions under bounded oracle error. This formulation separates two representation choices: principal-side history compression and agent-side sufficient statistics. Experiments in dynamic bandit and budget-constrained auctions show that the framework can recover known benchmarks and learn approximately incentive-compatible mechanisms beyond analytical settings. Furthermore, they reveal how representations affect optimization.
PaperID: 907, Poster
Authors:
Stefano Vannoni, Inês W Sampaio, Eleonora Maggioni, Islem RekikAbstract: Graph neural networks for brain connectomics treat every connection as topologically equivalent, ignoring the brain-economy trade-off between the metabolic cost of maintaining a connection and the topological value it delivers to the network. We argue that this trade-off is a property of the channel (i.e., edge) between two brain regions, not of the regions themselves, and should therefore shape how regions interact during attention computation rather than being appended as a node feature or auxiliary loss. We introduce the ecospace, a two-dimensional coordinate system that characterizes each brain connection by its economy—the trade-off between connectivity strength and anatomical cost—and uses it to guide how brain regions attend to one another. We further introduce EcoSpace Rotary Encoding (ESRE), an attention mechanism that injects the ecospace dimensions through asymmetric rotations of the query-key subspace, guaranteeing by mathematical identity that the attention score between two brain regions depends on the economy of the edge connecting them. Together, these two contributions form the Brain Graph Transformer with EcoSpace Rotary Encoding (BGT-ESRE). On brain graph classification, BGT-ESRE outperforms classical, general-purpose, and brain-graph-specific baselines by a substantial margin in both accuracy and AUC. Ablations confirm that the rotary mechanism and the brain economy-aligned transformation each contribute independently, and that asymmetric rotary injection is superior to additive bias injection of the same economy measure.
PaperID: 908, Poster
Abstract: Large foundation models have accelerated progress toward general-purpose agents that interact with humans and other agents through language and multimodal signals. However, robust multi-agent decision-making requires reasoning about what other agents know, intend, and are likely to do under partial observability. Current agentic systems often operate through prompt design, memory, or end-to-end behavioral shaping, but typically do not learn an explicit partner-state representation that can be reused as a decision variable across tasks. We introduce mental-model-enabled agents, a framework that equips an agent with a latent mental model of its counterpart, allowing it to infer hidden beliefs, intentions, and likely reactions from the observed history and use these inferences to guide action selection. Our method learns an amortized recursive Theory-of-Mind representation, with first- and second-order mental-state structure, jointly with a belief-conditioned reward model that evaluates candidate actions relative to the inferred partner state. A policy is then learned under this belief-aware signal, yielding an agent that can act independently at inference time while retaining the benefits of explicit partner modeling. We evaluate the same framework on both language-only and multimodal benchmarks. Across these settings, explicit mental-state modeling consistently improves interaction quality and Theory-of-Mind performance over base agentic systems, showing that structured partner modeling is a useful inductive bias for general multi-agent systems.
PaperID: 909, Poster
Authors:
Fengbei Liu, Rachit Saluja, Sunwoo Kwak, Ruibo Wang, Ruining Deng, Heejong Kim, Johannes C. Paetzold, Mert SabuncuAbstract: Multi-objective optimization (MOO) underlies many machine learning problems, yet MOO solvers across the loss-balancing, gradient-balancing, and Pareto-based families almost universally hand their reconciled directions to Adam~\citekingma2015adam. We show this coupling introduces two systematic gaps between the solver's intent and the optimizer's execution. The first is a \emphweighting mismatch: Adam's second-moment denominator entangles the time-varying preference vector with gradient statistics, marginalizing the preference into a history average and collapsing distinct Pareto trade-offs toward a near-uniform mixture. The second is a \emphgeometric mismatch: Adam's adaptive metric distorts the Euclidean geometry MOO solvers assume, turning aligned objectives into apparent conflicts. To resolve both jointly, we introduce MAdam (Metric-Aware Multi-Objective Adam), a drop-in wrapper that leaves both solver and optimizer unchanged. MAdam preconditions the reconciled direction by the preference-conditioned curvature of the scalarized objective; on this whitened input, Adam's second moment collapses to identity, so the realized update is governed by the preference-conditioned metric. Across multi-task learning, Pareto-front recovery, physics-informed neural networks, and medical imaging, MAdam consistently improves over Adam for every solver family.
Abstract: , where outputs agree strikingly often. However, we present two experiments that substantially disagree with prior findings of monoculture. In particular, we find that monoculture virtually disappears in one dataset when we include item difficulty compared to previous works that do not. Informed by these findings, we theoretically formalize what it means to measure monoculture. We show that monoculture is inherently subjective, relying on two key decisions: 1) the baseline null model for what "independence" should look like; and 2) the population of models and items under consideration. We conclude with concrete guidelines for future evaluators to ensure the robustness of their results.
PaperID: 911, Poster
Abstract: Semi-implicit variational inference (SIVI) expands the representational capacity of variational families, but optimizing the Evidence Lower Bound (ELBO) remains challenging because the marginal score of the semi-implicit distribution is generally intractable. Unbiased Implicit Variational Inference (UIVI) and related methods address this difficulty through reparameterized ELBO gradients, but existing approaches estimate the required score indirectly through reverse-conditional sampling or approximation, which can be computationally costly and unstable. We propose Denoising Implicit Variational Inference (DIVI), which learns the marginal score directly via denoising score matching. The learned score is used as a plug-in estimator in the pathwise ELBO gradient, replacing inner reverse-conditional sampling with a score-network evaluation. We analyze DIVI as inexact stochastic gradient ascent and show, under stated assumptions, that the averaged stationarity measure is controlled by the usual optimization term and the average score-matching error. Empirically, we evaluate DIVI on both synthetic distributions and a variety of real-data Bayesian inference tasks. The results show that DIVI improves over the current UIVI baselines while reducing the cost of ELBO-gradient estimation.
Abstract: Time series, spatial data, and images are natural applications of Neural Processes. However, when such data exhibit strong periodicity and quasi-periodicity, existing methods often suffer from underfitting and generalise poorly beyond the training distribution. In this work, we propose Spectral Transformer Neural Processes (STNPs), a frequency-aware extension of Transformer Neural Processes (TNPs). STNPs introduce a Spectral Aggregator that estimates an empirical context spectrum, compresses it into a spectral mixture, samples task-adaptive spectral features, and concatenates them with time-domain embeddings, thereby injecting a spectral-mixture-kernel bias into TNPs. This design reshapes the similarity geometry, allowing inputs that are distant in Euclidean space to remain close in an induced periodic manifold while enhancing time--frequency interactions. Extensive experiments on synthetic regression tasks, real-world time-series datasets, and an image dataset demonstrate that STNPs consistently improve predictive performance over existing baselines, extending Neural Processes beyond translation equivariance toward effective modelling of periodicity and quasi-periodicity.
PaperID: 913, Poster
Abstract: Modern neural-network optimizers occupy a three-way design space: online adaptivity, architecture-aware geometry, and acceleration. We introduce AcceleGrad#, an accelerated linear-coupling template that couples an adaptive anchor with a norm-induced sharp, mirror, or linear-minimization geometry branch. The scalar-anchor specialization recovers the accelerated O(1/K^2) rate for smooth convex objectives, while the diagonal-AdaGrad clipped-LMO training variant admits local \mu-Kurdyka-Lojasiewicz certificates: a generic conservative O(K^-1/4) rate and, in a SCION cross-entropy regime, a O(K^-1/3) last-iterate rate to a self-bounded mini-batch noise level. On image classification and language modeling, the training variant improves over Euclidean AcceleGrad and remains competitive with recent non-Euclidean optimizers. Using stochastic gradients directly in spectral oracles also reduces polar-computation time and optimizer overhead in practice.
PaperID: 914, Poster
Authors: Cornelius Brand
Abstract: Determinantal point processes allow efficient sampling of diverse subsets of data. However, (exact) sampling under additional constraints is often computationally intractable under complexity-theoretic assumptions. We take an algebraic approach to designing novel algorithms that allow exact constrained sampling with theoretical efficiency guarantees. Our algorithms apply to a general class of point processes based on Bombieri inner products of polynomials, which we call Bombieri k-point processes. This perspective recovers known algorithms for constrained k-DPPs and extends them to a broader class of simple k-point processes. As applications, we obtain exact samplers for matching-, path-, subgraph-, and budget-constrained processes, as well as distributions combining diversity with clustering. While common complexity-theoretic assumptions rule out polynomial-time algorithms in many of these settings, our algorithms are efficient with respect to the relaxed notion of fixed-parameter tractability when the sample size k is kept bounded.
Abstract: Neural scaling laws predict how language model performance improves with increased training inputs. While aggregate metrics like validation loss can follow smooth power-law curves, individual downstream tasks exhibit diverse scaling behaviors: some improve monotonically, others plateau, and some even degrade with scale. We argue that predicting downstream performance from validation loss suffers from two limitations: averaging token-level losses obscures signal, and no simple parametric family can capture the full spectrum of scaling behaviors. To address this, we propose Neural Neural Scaling Laws (NeuNeu), a neural network that frames scaling law prediction as time-series extrapolation. NeuNeu combines temporal context from observed accuracy trajectories with token-level validation losses, learning to predict future performance without the limitations inherent in assuming a specific functional form. Trained entirely on open-source model checkpoints from HuggingFace, NeuNeu achieves 1.99% mean absolute error in predicting model accuracy on 66 downstream tasks---a 44% reduction compared to logistic scaling laws (3.56% MAE). Furthermore, NeuNeu generalizes zero-shot to unseen model families, architectures, parameter counts, and downstream tasks. Our work suggests that predicting downstream scaling directly from data outperforms parametric alternatives.
PaperID: 916, Poster
Authors: Eli N Weinstein, Andrei Slabodkin, Mattia G Gollub, Kerry Dobbs, Xiao-Bing Cui, Fang zhang, Kristina Gurung, Elizabeth B Wood
Abstract: One strategy to scale up ML-driven science is to increase wet lab experiments' information density. We present a method based on a neural extension of compressed sensing to function space. We measure the activity of multiple different molecules simultaneously, rather than individually. Then, we deconvolute the molecule-activity map during model training. Co-design of wet lab experiments and learning algorithms provably leads to orders-of-magnitude gains in information density. We demonstrate on antibodies and cell therapies.
Abstract: Prediction sets provide a means of quantifying the uncertainty in predictive tasks. Using held out calibration data, conformal prediction and risk control can produce prediction sets that exhibit statistically valid error control in a computationally efficient manner. However, in the standard formulations, the error is only controlled on average over many possible calibration datasets of fixed size. In this paper, we extend the control to remain valid with high probability over a cumulatively growing calibration dataset at any time point. We derive such guarantees using quantile-based arguments and illustrate the applicability of the proposed framework to settings involving distribution shift. We further establish a matching lower bound and show that our guarantees are asymptotically tight. Finally, we demonstrate the practical performance of our methods through both simulations and real-world numerical examples.
PaperID: 918, Poster
Abstract: Summing vectors is a basic task in differentially private data analysis. Standard algorithms clip vectors to some bound, sum them, and add noise scaled to the bound. While the differential privacy literature has developed a deep library of such noise addition mechanisms, it offers few tools for privately choosing the clipping bound directly from the data. We introduce a method that privately chooses a clipping bound by optimizing residual coherence, the directional alignment of vectors exceeding the bound, against the variance cost of additional noise. We prove a close relationship between residual coherence and bias for a general class of heavy-tailed data regimes and show empirically that, across several datasets, our method outperforms existing baselines.
PaperID: 919, Poster
Abstract: Helpful-only models—that is, models that are trained to always follow user intent—are valuable for dangerous capability evaluations and other areas of AI R&D where refusals would be an obstacle. Little is known about the generalization properties of helpful-only training: they refuse less than their harmless counterparts, but previous work has not studied other dimensions of their alignment. We find that existing helpful-only models have serious shortcomings. Some show emergent misalignment, others have residual refusal behaviors, and most show poor steerability, sycophancy, and an incoherent character. We show that simple anti-refusal training can cause many of these issues. None of these problems are necessary consequences of helpful-only training, though: we show that synthetic document fine-tuning and adding character-related questions to SFT and RL can mitigate them.
Abstract: We study posterior sampling for inverse problems in discrete state spaces using discrete diffusion models as generative priors. While continuous diffusion models have become widely used for inverse problems, their discrete counterparts remain comparatively underexplored. Existing discrete posterior samplers often rely on continuous relaxations of discrete variables, Gibbs-style updates, or mechanisms specialized to particular corruption processes, which can limit scalability or generality. We propose \DeltaLPS, a Discrete Langevin-Inspired Posterior Sampler that uses gradient information to identify promising discrete moves without leaving the discrete state space. The resulting approach enables efficient parallel updates across all token dimensions and is agnostic to the training paradigm of the discrete diffusion prior, including masked and uniform-state diffusion. We evaluate our method on image restoration tasks across MNIST, CIFAR, and FFHQ, as well as spatial mapping, covering linear, nonlinear, and blind inverse problems. Across these settings, we improve over recent discrete diffusion posterior samplers and are competitive with strong continuous diffusion-based inverse solvers. Our results suggest that fully discrete, gradient-informed posterior samplers offer a scalable and general path toward solving inverse problems over discrete representations.
Abstract: Recurrent memory agents extend LLMs to arbitrarily long contexts by iteratively consolidating input into a fixed-size memory window. Despite their scalability, these agents exhibit a well-documented reliability problem: end-to-end performance degrades systematically as context length grows. We diagnose this failure by decomposing performance into two factors---memory ---and quantitatively confirm that retention is the dominant bottleneck. Retention collapses because existing designs maintain memory as a monolithic text block, forcing every update to risk overwriting previously retained content. Motivated by this diagnosis, we propose , a general, training-free framework that partitions memory into independent heads governed by a stage-wise select-then-update strategy. At each step, exactly one head is selected for update while the remaining heads are structurally shielded from overwriting, shifting the burden of retention from model behavior to architectural design. As a lightweight instantiation, we introduce Least-Recently-Updated MHM (MHM-LRU), which guarantees uniform head utilization with zero additional token overhead. Extensive experiments on long-context benchmarks show that MHM-LRU substantially improves both retention and end-to-end accuracy across the 100K--1M token range, where baselines degrade sharply. On RULER-HQA at 896K tokens, MHM-LRU improves the memory retention rate from less than 30% to . These gains generalize across model families, scales, and task types, positioning architectural optimization as a practical and cost-efficient path toward reliable long-context recurrent memory.
PaperID: 922, Poster
Abstract: Predictive coding (PC) is a local-learning alternative to backpropagation (BP), training deep networks via a local energy-minimization dynamics rather than a global backward pass. We introduce Augmented-Lagrangian Predictive Coding (PC-ALM), which maintains PC's inference budget but aligns each weight update toward BP by accumulating per-layer constraint errors into a layer-local Lagrange multiplier. In linear PC networks, PC-ALM converges to an equilibrium with exact BP gradients distributed across the network via only layer-local updates. We analyze PC-ALM in nonlinear PC networks up to depth 128 and show that it matches BP performance across all width-depth regimes, notably in deep narrow networks where PC underperforms. PC-ALM introduces recurrent dynamics in each layer's activations. Compared to PC's heat flow on a scalar energy, PC-ALM dynamics are driven by dual ascent on the Augmented Lagrangian. We observe ballistic credit propagation across especially deep networks, with prediction errors evenly distributed across layers, compared to PC's diffusive credit propagation. Beyond the algorithm itself, the augmented Lagrangian framework offers a generalization of PC, and may yield insights on how distributed systems could compute and propagate BP-like credit signals through purely local dynamics.
PaperID: 923, Poster
Abstract: Despite the remarkable progress in Multimodal Large Language Models (MLLMs), the prevailing Chain-of-Thought (CoT) paradigms remain predominantly confined to the linguistic space. While recent advancements have focused on bridging the Perception Gap through high-resolution cropping (e.g., Thinking with Images), they overlook a more fundamental bottleneck: the Reference Gap. The inherent ambiguity of natural language often fails to provide precise, unambiguous pointers to complex spatial layouts, leading to logical collapse in tasks requiring rigorous grounding. In this work, we introduce Thinking with Visual Primitives, a novel reasoning framework that elevates spatial markers (e.g., points and bounding boxes) to "minimal units of thought". By interleaving these visual primitives directly into the thinking process, our model can "point" while it "reasons", effectively grounding its cognitive trajectory in the physical coordinates of the image. Notably, our framework is built on a highly optimized architecture with extreme visual token efficiency. Despite its compact model scale and significantly lower image-token budget, our model achieves frontier-competitive performance on a focused suite of challenging visual QA tasks, matching or exceeding models such as GPT-5.4, Claude-Sonnet-4.6, and Gemini-3-Flash. This demonstrates a path toward more efficient and scalable System-2-like multimodal intelligence.
Authors: James Pritts, Felix Seegräber, Kevin Köser
Abstract: The most widely used RANSAC variants score candidate models by counting inliers or summing truncated likelihoods; every such score requires a user-supplied parameter that is a function of the inlier scale, which must itself be estimated from contaminated data. We remove this dependence by reversing the usual order of inference: for a fixed inlier partition we marginalize \sigma analytically in closed form under a conjugate Inverse-Gamma prior, then optimize over partitions. A single closed-form expression spans the non-informative Jeffreys limit (which requires no validation data to fit the prior) and informative empirical-Bayes priors fit from a small validation set, so the same score adapts across data-rich and data-scarce regimes without any change to the algorithm. To our knowledge this is the first RANSAC score in which the inlier scale is genuinely absent from the score formula. The score admits O(N \log N) computation via sort-and-sweep. On a benchmark of nearly 70\,000 image pairs spanning different two-view estimation problems and both engineered and learned feature pipelines, the proposed score matches or exceeds the state of the art (RANSAC, MSAC, GaU, MAGSAC++), excelling in robustness to hyperparameter miscalibration, sample efficiency at small validation budgets, and adaptive regularization across data-rich and data-scarce regimes.
Abstract: Formal model evaluation methods typically certify that a model satisfies a prescribed target key performance indicator (KPI) level. However, in many applications, the relevant target KPI level may not be known a priori, and the user may instead wish to compare candidate models by analyzing the full trade-offs between performance and reliability achievable at test time by the models. This task, requiring the reliable estimate of the test-time KPI distributions, is made more complicated by the fact that the same data must often be used both to pre-select a subset of candidate models and to estimate their KPI distributions, causing a potential post-selection bias. In this work, we introduce post-selection distributional model evaluation (PS-DME), a general framework for statistically valid distributional model assessment after arbitrary data-dependent model pre-selection. Building on e-values, PS-DME controls post-selection false coverage rate (FCR) for the distributional KPI estimates and we establish explicit conditions under which it is provably more sample efficient than a baseline method based on sample splitting. Experiments on synthetic data, text-to-SQL decoding with large language models, and telecom network performance evaluation demonstrate that PS-DME enables reliable comparison of candidate configurations across a range of reliability levels, supporting the statistically reliable exploration of performance--reliability trade-offs.
Authors: Aarav G Sane, Karthik Sivachandran, Rohan Paleja
Abstract: Human coordination often relies on the ability to influence the beliefs of others through strategic action. In multi-agent reinforcement learning, opponent shaping attempts to replicate this influence, though existing methods typically operate within an opponent's parameter, policy, or value space. Meanwhile, belief-manipulation techniques in hidden-role games often rely on hard-coded objectives, such as deception or belief saturation. We propose Differentiable Belief-based Opponent Shaping (D-BOS), a first-order method that treats each observer's belief as the shaped opponent state and differentiates through k-step softmax-Bayes belief dynamics. Rather than explicitly rewarding deceptive or cooperative behavior, our method treats the belief state as the target for shaping. This allows the optimal strategy to emerge naturally from the environment's reward structure. This belief-space formulation provides an opponent-shaping signal by differentiating through opponent belief updates, and naturally extends to multiple observers by aggregating gradients over their individual inferred belief trajectories. Empirically, D-BOS outperforms PPO and BBM in hidden-role games, with the largest gains in mixed-motive settings.
Abstract: Transformer-based large language models are increasingly used for long-horizon tasks; however, their attention mechanism scales poorly with context length. To handle this, we study a sleep-like consolidation mechanism in which a model periodically converts recent context into persistent fast model weights before clearing its key-value cache. During the sleep phase, the model performs N offline recurrent passes over the accumulated context and updates the fast weights in its state-space model (SSM) blocks through a learned local rule. This shifts extra computation to the sleep phase while preserving the latency of wake-time prediction. We test our method on controlled synthetic tasks, including cellular automata and multi-hop graph retrieval, as well as a more realistic math reasoning task, on which a regular transformer as well as SSM-attention hybrid models fail. We then show that increasing sleep duration N for our models improves performance, with the largest gains on examples that require deeper reasoning.
Authors: Keru Chen
Abstract: Offline-to-online reinforcement learning first warm-starts a policy from a fixed offline dataset and then improves it with limited online interaction. Offline data reduces uncertainty, but it does not remove the need for exploration; it changes what remains to be explored. We formalise this residual uncertainty by the conditional mutual information \I(\chi;\tau_1:T\mid\mathcal\D\_N)\ between a learning target \\chi\ and the online trajectories after conditioning on the offline dataset. This view leads naturally to information-directed sampling (IDS), a family parameterised by \\eta\ge 0\ that selects actions by trading off instantaneous regret against information gain. We prove a generic offline-to-online Bayesian regret bound for IDS through a ratio certificate: any information-ratio bound satisfied by a reference Thompson-sampling policy over the same randomised policy class is inherited by IDS. In a known-dynamics Bayesian linear-reward model, the conditional mutual information has a log-determinant form, and vanilla IDS (\\eta=0\) satisfies \\widetilde O\(Hd\min\\sqrt T,\,T\sqrt\C^\dagger\_\beta,\mathrm\IDS\_0(N,T)/N\\\)\, where the coverage coefficient is tied to the visitation distribution induced by vanilla IDS itself. We also identify a warm-start regime with a dominated but informative probe in which vanilla IDS selects the probe while Thompson sampling never does, giving a constant-factor Bayesian regret separation. Controlled bandit experiments and D4RL offline-to-online experiments support this mechanism: IDS is most beneficial when offline data is informative but leaves biased or low-probability residual uncertainty that can be resolved by targeted online actions.
Abstract: We show that the MinMax algebra provides a form of recurrence that is expressively powerful, efficiently implementable, and most importantly it is not affected by vanishing or exploding gradient. We call MinMax Recurrent Neural Cascades (RNCs) the models obtained by cascading several layers of neurons that employ such recurrence. We show that MinMax RNCs enjoy many favourable theoretical properties. First, their formal expressivity includes all regular languages, arguably the maximal expressivity for a finite-memory system. Second, they can be evaluated in parallel with a runtime that is logarithmic in the input length given enough processors—and they can also be evaluated sequentially. Third, their state and activations are bounded uniformly for all input lengths. Fourth, at almost all points, their loss gradient exists and it is bounded. Fifth, they do not exhibit a vanishing state gradient: the gradient of a state w.r.t. a past state can have constant value one regardless of the time distance between the two states. Finally, we find empirical evidence that the favourable theoretical properties of MinMax RNCs are matched by their practical capabilities: they are able to perfectly solve a number of synthetic tasks, showing superior performance compared to the considered state-of-the-art recurrent neural networks; also, we train a MinMax RNC of 127M parameters on next-token prediction, and the obtained model shows competitive performance for its size, providing evidence of the potential of MinMax RNCs on real-world tasks.
PaperID: 930, Poster
Abstract: The recent breakthrough Score-Repellent Monte Carlo (SRMC) method improves MCMC sampling by tilting the target density \pi in \mathbbR^d with the factor \exp(-\alpha\,\theta_n^\top s(x)), where the score function s(x) = \nabla_x \log \pi(x), and \theta_n is the score average from past samples. However, a single scalar repellence strength \alpha\ge0 cannot handle both the steep and the flat directions in anisotropic energy landscapes at once. We introduce Geometry-Aware SRMC (GA-SRMC), which generalizes the Euclidean alignment to \theta^\top M s(x) for any symmetric positive-definite matrix M. A stochastic-approximation central limit theorem under a trace constraint on M identifies the trace-normalized inverse Fisher M_\rm opt\propto S^-1 as the unique minimizer of the worst-direction covariance bound, elevating the Fisher information S=\textCov_\pi(s,s) from a proposal-side preconditioner (e.g., FisherMALA, natural gradient) to the optimal repulsion-side matrix, with a single online estimate \widehat S_n serving both roles at no additional cost. On an 8-target simulation spanning the condition number \kappa\in[1.3,100] and dimension d\in[2,785], GA-SRMC dominates SRMC by 4.4× on anisotropic Gaussian (\kappa=100) and 6.8× on MNIST under the MALA baseline, by 18\text - 32% on the well-conditioned Bayesian logistic-regression posteriors, and by up to +48% under FisherMALA; the advantage scales monotonically with the condition number \kappa.
PaperID: 931, Poster
Abstract: AI systems operating near the limits of their capabilities can maximize utility while minimizing the risk of error by abstaining or deferring questions they are not confident about to a human or a stronger model. Recently, large language models (LLMs) have been post-trained to estimate confidence in their own answers. To achieve this, methods reinforce confidence scores based on the correctness signal to each question-answer pair \emphindividually, with rewards that encourage assigning confidence 1 to correct answers and 0 to incorrect answers. However, estimating confidence scores is useful in that it enables selecting only the most confident answers, to maximize the share of solved questions while accommodating the user's risk tolerance. Assigning confidence 1 to correct answers and 0 to incorrect ones is theoretically optimal, but unattainable in practice and detrimental for selective prediction with a non-zero risk tolerance: confidences pushed to 0 and 1 carry no information about how answers compare to one another. We therefore introduce \ours (Cross-question Abstention Reward Estimation), which computes rewards over a set of question-answer pairs: At each of several risk tolerances, the confidences across the set induce an abstention threshold, and each confidence is rewarded for lying on the correct side of it. Since the threshold itself comes from the set, every verbalized confidence is reinforced relative to other questions' confidence, not against its own binary correctness label. On math, \ours increases the coverage at 5% risk (C@5) across MATH-500, GSM8K, and Big-Math-Digits from 0.8% to 42.4%, compared to existing pointwise correctness-matching rewards. When trained on multi-hop Wikipedia QA from HotpotQA, \ours improves C@5 in out-of-distribution general-knowledge, reasoning, and math benchmarks (including CommonsenseQA, GPQA, SimpleQA, and TriviaQA) from 9.9% to 23.6%.
PaperID: 932, Poster
Abstract: Masked Generative Models (MGMs) enable parallel decoding and achieve strong performance across modalities, but require full-sequence bidirectional transformers at every step, making training costly and degrading quality under low sampling budgets. Existing work improves efficiency via better samplers or fixed-depth backbones, but does not vary the depth of the denoiser across sampling steps. We introduce Fixed-Point Masked Generative Models (FP-MGMs), which replace part of the denoiser with a fixed-point solver over shared attention layers to enable adaptive depth with fewer parameters. To make it more effective for masked generation, we first introduce a cross-step consistency loss, which aligns hidden representations at neighboring denoising steps and, second, three-state reuse (3SR) which warm-starts the solver using the previous solution by treating unchanged, still-masked, and newly revealed tokens differently. Together, these components define our complete training-to-inference framework for fixed-point masked generation, \emphCoFRe. We also show that pre-trained MGMs can be converted into FP-MGMs with short fine-tuning, avoiding full retraining. Across modalities, CoFRe improves the quality and cost trade-off. On OpenWebText, CoFRe reduces parameters by 38.8%, training time by 11.5%, and VRAM by 16.9%, while improving generative perplexity from 830.8 to 101.8 at a budget of 96 transformer-block forward passes, compared to MDLM. In ImageNette, CoFRe reduces training time by 48.6% and VRAM by 50.7%, while improving FID in all sample budgets tested. Overall, CoFRe offers a practical framework for cheaper training and stronger low-budget masked generation.
Authors: Michal Feldman, Amos Fiat, Yael Nissan, Tomasz Ponitka
Abstract: We introduce epistemic pairwise maximin share (EPMMS), a new fairness notion for fair division of indivisible goods. Two fundamental notions in this setting are envy-freeness up to any item (EFX) and pairwise maximin share (PMMS), with PMMS being stronger than EFX. While EFX has been extensively studied, far less is known about PMMS. Recent work shows that relaxing EFX via an epistemic perspective leads to substantial progress on the EFX problem, raising the question of whether a similar approach can advance our understanding of PMMS. Motivated by this, we initiate the study of EPMMS, the epistemic relaxation of PMMS. EPMMS is more challenging than EEFX: the key approaches underlying recent progress on epistemic EFX inherently fail to extend to EPMMS. We establish the following results. 1. For additive valuations, 4/5-EPMMS allocations exist and can be efficiently computed. 2. For bivalued valuations, EPMMS allocations exist and can be efficiently computed; in fact, we obtain the stronger guarantee of epistemic groupwise maximin share (EGMMS), which also strengthens the existence of MMS allocations for this setting. 3. EPMMS allocations exist in two settings where MMS allocations need not exist: instances with three additive agents or two types of additive agents.
PaperID: 934, Poster
Authors: Andreia Podasca, Anup Das
Abstract: Decision Transformer (DT) formulates offline reinforcement learning as conditional sequence modeling, predicting actions by attending over past states, actions, and returns-to-go. However, the softmax attention in DT computes context representations in a single forward pass, offering no mechanism to recover from observations corrupted by sensor noise or measurement errors common in real-world deployment. Prior methods improve robustness of DT through training regularization or objective modifications, but leave the attention mechanism itself unchanged. We introduce Robust Hopfield Decision Transformer (RHDT), which replaces softmax attention with modern Hopfield layers that iteratively minimize an energy function, enabling observations to converge toward learned patterns (attractors). To prevent pattern interference from overlapping attractor basins, we regularize the architecture through Lipschitz gradient penalty and orthogonality constraints on attention keys, which serve as the stored patterns. On D4RL benchmarks, RHDT matches robust baselines in average return while significantly improving worst-case performance under observation corruption.
Abstract: Neural networks are typically optimized with variants of stochastic gradient descent. Under a squared loss, however, the optimal solution to the linear last layer weights is known in closed-form. We propose to leverage this during optimization, treating the last layer as a function of the backbone parameters, and optimizing solely for these parameters. We show this is equivalent to alternating between gradient descent steps on the backbone and closed-form updates on the last layer. We adapt the method for the setting of stochastic gradient descent, by trading off the loss on the current batch against the accumulated information from previous batches. We provide theoretical analyses showing convergence of the method to an optimal solution in the neural tangent kernel regime, as well as quantifying the gains compared to standard SGD in a one-step analysis. Finally, we demonstrate the effectiveness of our approach compared with SGD and Adam on a squared loss in several regression tasks, including neural operators and causal inference.
PaperID: 936, Poster
Abstract: Multi-output Gaussian process (MOGP) regression allows modelling dependencies among multiple correlated response variables. Similarly to standard Gaussian processes (GPs), MOGPs are sensitive to model misspecification and outliers, which can distort predictions within individual outputs. In the multi-output case, this situation can be further exacerbated as anomalous observations in one response can adversely influence predictions for other correlated outputs. To handle this situation, we propose R-MOGP—an MOGP that is provably robust to outliers in any individual response. Our approach preserves conjugacy, achieving robustness without sacrificing computational efficiency. Empirically, we show that R-MOGP achieves performance comparable to existing state-of-the-art robust MOGPs at a fraction of their computational costs.
PaperID: 937, Poster
Abstract: As online social platforms increasingly serve as central arenas for opinion aggregation, numerous opinion dynamics models have been extensively developed to characterize how opinions propagate and to predict emerging social phenomena. Among various models, the Friedkin--Johnsen (FJ) model stands out as one of the most extensively studied frameworks, with broad applications across diverse domains and recent inspiration on the design of graph neural network architectures. However, the conventional FJ row-stochastic pairwise interaction has been shown to induce imbalanced influence patterns, often weakening the impact of high-degree nodes. To address this issue, we propose a normalized FJ model based on the normalized Laplacian, which symmetrically incorporates degree information from both endpoints. We establish its convergence and analyze its structural properties and intrinsic balance, along with a game-theoretic interpretation. We further develop efficient algorithms for computing equilibrium opinions and related metrics. Experiments on large-scale networks demonstrate the effectiveness and scalability of our approach.
Authors: Ian Osband
Abstract: Distributed reinforcement learning trains on data from stale, buggy, or mismatched actors, producing actions with high surprisal (negative log-probability) under the learner's policy. The core difficulty is not surprising data per se, but \emphnegative learning from surprising data. High-surprisal failures can dominate finite-batch updates through large perpendicular components, while high-surprisal successes reveal opportunities the current policy would otherwise miss. The Delightful Policy Gradient (DG) separates these cases by gating each update with delight, the product of advantage and surprisal, suppressing rare failures and preserving rare successes without behavior probabilities. In a tabular analysis, DG suppresses the perpendicular second moment of high-surprisal failures by a policy-overlap factor that vanishes as the learner improves. The advantage sign is essential for surprisal-based filtering: any learner-probability-only gate that suppresses rare failures also suppresses rare successes. On MNIST with simulated staleness, DG without off-policy correction outperforms importance-weighted PG with exact behavior probabilities. On a transformer sequence task with staleness, actor bugs, reward corruption, and rare discovery, DG often achieves nearly order-of-magnitude lower error. When all four frictions act simultaneously, its sample-efficiency advantage is order-of-magnitude and grows with task complexity.
Abstract: We introduce structural causal bottleneck models (SCBMs), a novel class of structural causal models in which causal effects between high-dimensional variables are mediated by low-dimensional summary statistics, or . SCBMs provide a flexible framework for mechanism-specific, target-dependent dimension reduction while remaining estimable via standard learning algorithms. We prove that bottleneck variables are identifiable up to bijection from observational data, and validate this experimentally. We then study causal effect estimation in linear SCBMs under three standard adjustment strategies. For instrumental variables, we identify a population-level non-identifiability of naïve two-stage least squares when all variables are high-dimensional, and show that bottleneck representations restore identification. For mediation and backdoor adjustment, we prove that conditioning on bottleneck variables yields valid effect estimates with explicit finite-sample efficiency gains over conditioning on full high-dimensional variables. We argue that SCBMs provide a principled alternative to existing causal dimension reduction frameworks such as causal representation learning and causal abstractions.
Abstract: A fundamental limitation of Text-to-Code is that no guarantee can be obtained about the correctness of the generated code. Therefore, to ensure its correctness, the generated code still has to be reviewed, tested, and maintained by developers. However, parsing through LLM-generated code can be tedious and time-consuming, potentially negating the productivity gains promised by AI-coding tools. To address this challenge, we present Viverra, a system that automatically produces alongside generated code to aid user's understanding of the generated program. Given a natural-language task description, Viverra prompts an LLM to synthesize a C program together with candidate assertions expressing safety and correctness properties. It then verifies those assertions in a compositional and best-effort manner via a portfolio of bounded model checkers. Evaluation on 18 diverse programming tasks suggests that Viverra can efficiently generate code with verified assertions, and that these assertions improve users' performance on code-comprehension tasks in a user study with more than 400 participants.
Authors: Lukas Billera, Hedwig N Nordlinder, Ben Murrell
Abstract: Many recent flow-matching and diffusion-style generative models rely on auxiliary stochastic dynamics during training: a richer process is simulated to define conditional targets, but the auxiliary state is either intractable to sample at generation time or simply not part of the desired output. Existing Generator Matching theory formalises conditioning on static latent random variables, and several recent papers prove special cases of projection results for particular augmented-state constructions. We introduce latent process generator matching, a general framework that treats the observed generative state as a deterministic image X_t=\Phi(Y_t) of a tractable Markov process Y_t. We show that in this setting one may learn the generator of a stochastic process on the image space which has the same one-time marginal distributions as the projected process. This generalizes and subsumes the discrete latent process results from the literature, and extends Generator Matching from static latent variables to a rich family of time-dependent latent conditional processes.
Authors: Juyeop Kim, Songkuk Kim, Jong-Seok Lee
Abstract: Despite their success in image generation, diffusion models can memorize training data, raising serious privacy and copyright concerns. Although prior work has identified empirical factors associated with memorization, such as text guidance, prompt specificity, and duplicated training examples, the mechanism by which memorization is triggered and propagated during denoising remains unclear. In this paper, we provide a theoretical explanation of how memorization occurs in text-to-image diffusion models. We show that conditional overfitting causes the conditional posterior mean to collapse to the memorized training latent, while the unconditional posterior remains close to the zero-centered data mean at the first denoising step. As a result, classifier-free guidance overestimates the clean prediction and injects an amplified memorized signal into the reverse trajectory. We further show that this signal propagates across timesteps because each guided latent is fed back into the denoising network, allowing the conditional branch to indirectly shift the unconditional branch. To characterize this process, we introduce the posterior lag, the discrepancy between conditional and unconditional posterior means. For normal prompts, this lag decreases as the two posteriors synchronize during denoising. For memorized prompts, the conditional posterior commits early to a specific training latent, while the unconditional posterior catches up only later, producing a distinctive rise-and-fall lag pattern. Our analysis provides a mechanism-level explanation of memorization and clarifies why text guidance and early denoising steps play a central role in memorized generation.
PaperID: 943, Poster
Authors:
Joey Cai, Yufan Zhang, Ruichen Tan, Zengxiang Lei, Satish UkkusuriAbstract: Lane and road markings provide critical guidance for vehicle navigation and multi-agent coordination, yet their design principles are scattered across disparate datasets, severely limiting quantitative analysis and scenario testing. We introduce controllable road-marking generation: given a drivable-area mask and a sparse outer-ring observation, a model must synthesize a complete center-region layout of lane dividers, road dividers, and pedestrian crossings that is semantically consistent with a free-form text prompt. We develop a conditional diffusion pipeline that combines (i) a text-conditioned BEV diffusion backbone, (ii) a Gaussian blur-then-deblur training target that stabilizes thin-structure learning, and (iii) a structured Gaussian render post-processing step that snaps soft predictions to crisp, topology-aware markings. On the Argoverse 2 test split, our system substantially outperforms a state-of-the-art DDIM mask-refinement baseline on standard structural fidelity metrics, and supports prompt-level edits across crosswalk, lane-divider, and road-divider semantics. We anticipate that the proposed framework will enable downstream applications in urban infrastructure planning, autonomous-driving stress-testing, and navigation in unstructured environments.
Abstract: Long-form reasoning and tool calling traces accumulate errors and stale content that anchor subsequent generations, a phenomenon known as \emphcontext rot. Existing scaffolds mitigate this with fixed-interval compaction triggered at a token threshold, but such triggers are blind to trajectory structure and risk discarding partial results mid-derivation or mid-search. We propose SelfCompact, which pairs two inference-time elements: an inline \emphcompaction tool the model invokes itself, and a lightweight \emphrubric specifying when to fire (a sub-task has resolved, or the trajectory is converging) and when to suppress (mid-derivation, or when stuck). Both are needed. The tool alone is unevenly used across open-weight models, often invoked at unhelpful moments or not at all; the rubric alone cannot act. Together, they elicit effective adaptive compaction without any fine-tuning. On competition math (IMO-Answerbench, HMMT Nov 25 / Feb 26) with four Qwen3 / Qwen3.5 models and agentic search (BrowseComp, BrowseComp-Plus) with three deployed agents, \method matches or exceeds fixed-interval summarization at a fraction of the token cost, improving over a no-summarization baseline by up to 16.7 points on math and 5--9 points on agentic search at 30--70% lower per-question cost. The result reframes \emphwhen to compact as a meta-cognitive capability that scaffolds, not weights, can supply.
PaperID: 945, Poster
Abstract: We present a simple, inexpensive, and effective method for heteroscedastic uncertainty quantification in neural networks. We build on Variational Bayesian Last Layers (VBLL), wherein deterministic training objectives are developed for variational inference of the network last layer. In particular, we (1) Introduce t-VBLL layers, which perform variational inference for the aleatoric noise covariance, and (2) Introduce Het-VBLL, a Bayesian last layer scheme to model heteroscedastic noise. These methods are based on novel, analytically tractable evidence lower bounds. We further discuss parameterization and initialization within these models. We show that these novel design elements enable effective uncertainty modeling at minimal additional cost, and substantially improve performance over similar methods such as VBLLs.
Abstract: We consider smooth convex minimization over compact convex sets, i.e., \min_x \in \mathcal C f(x) with the (vanilla) Frank--Wolfe algorithm. Well-known lower bounds establish a worst-case \Omega(1/t) primal-gap barrier in the general smooth convex case, and faster convergence usually requires favorable function properties such as Hölder error bounds or strong convexity. We present a new Local Dual Sharpness (LDS) condition, essentially a property of the feasible region and its LMO, under which the Frank--Wolfe algorithm converges in o(1/t) for any smooth convex function, ruling out an \Omega(1/t) lower bound under LDS. The condition is a generalization (and localization) of uniform convexity of sets and it is satisfied by any uniformly convex set. To our knowledge, this is the first unconditional o(1/t) convergence result for uniformly convex sets. Combining LDS with stronger function properties, e.g., a local variant of Hölder error bounds, allows us to quantify the actual rates.
PaperID: 947, Poster
Abstract: Machine learning (ML) predictions are increasingly being used to guide decision-making, giving rise to the problem of decision-focused learning (DFL) where predictors are optimized for downstream decision quality rather than accuracy alone. However, existing work assumes a single decision-maker optimizing in isolation. This paper formalizes decision-focused learning, where an ML system predicts an exogenous state that some agents observe before playing a game. For example, a park ranger may predict wildlife locations to allocate anti-poaching patrols against strategic poachers. While the exogenous state is unaffected by agent actions, predictions influence agents' strategies and the resulting equilibrium. We find that strategic considerations fundamentally change the learning problem. In particular, we show the prediction accuracy–equilibrium payoff landscape can be non-monotonic---i.e., better predictions can degrade performance. We propose algorithmic approaches to address these challenges and validate them across benchmarks in wildlife conservation and infrastructure protection. Our theory and experiments highlight the importance of accounting for strategic interactions when designing predictors.
PaperID: 948, Poster
Abstract: Machine learning models, particularly large foundation models, are increasingly evaluated along multiple axes that are relevant for their usage. In addition to accuracy measures appropriate for the specific task of interest, measures capturing safety, calibration and fairness properties are often of interest. The classical approach to these problems assumes that the test set has been held out from the training process, an assumption that is increasingly harder to justify for large foundation models. In this work, we formalize a notion of evaluability without this assumption, and ask if it is feasible to evaluate black-box models in settings where our test samples may have been used during training. We study this problem under different settings, capturing different levels of access to the model. We first study the natural setting, where the evaluator only looks at the output of the model on the test set. We fully characterize the class of distributions where accurate evaluation is possible, and show tight upper and lower bounds on the sample complexity of evaluation. Our results show that black-box evaluation in this set up is feasible if and only if the distribution is close to being small support. We then consider evaluation algorithms that can query the model on additional inputs and demonstrate connections to self-correctors studied in program testing. Using these, we show that a small amount of query access can increase the power of the evaluator for natural distributions and concept classes.
Authors:
Zhonghan Xu, Ling Wang, Junhao Chen, Jianwei Zhao, Jinwei YangAbstract: Training neural networks directly in a low-rank parameterization is an appealing route to reducing memory, compute, and storage simultaneously during both training and inference. Dynamic low-rank training (DLRT), which confines weights to a rank-r manifold via the Galerkin projection of the gradient flow, is particularly attractive because it identifies efficient subnetworks on the fly without specialized initialization or post-factorization. However, DLRT fails to find trainable networks under high compression. In this paper, we derive the gradient flow of the best rank-r approximation and point out that the offset of DLRT comes from a curvature-coupling term which is large and thus non-negligible under aggressive compression. Guided by this analysis, we propose a stable dynamic low-rank training method, named SDLRT, which maintains a lightweight compensation buffer that reinjects the top neglected singular directions. Additionally, we introduce a negative feedback on the truncation tolerance to stabilize each layer's rank. Experimentally, SDLRT reliably finds trainable subnetworks where DLRT collapses and as a PEFT adapter on DeBERTa-v3, it achieves the best average score on SuperGLUE at only 2.8% parameter overhead over LoRA.
PaperID: 950, Poster
Abstract: Long-context failures are often framed as retrieval failures: a model must find a relevant token among many distractors. However, many practical long-context settings are better understood as state-tracking problems: dialogue histories contain corrections to earlier preferences, code traces repeatedly assign values to variables, document histories revise previous claims, and tool-use trajectories update the current task state. In these settings, a model must not only identify relevant tokens, but also resolve which relevant occurrence is currently valid. We propose Sieve Attention, a position-encoding-free attention mechanism for length-extrapolating Transformers. SRA decomposes attention into two operations: content filtering and temporal state resolution. For each query, SRA first applies a sparse content filter over the full prefix to identify candidate state updates, and then applies a sequential hazard allocation process only over the selected candidates. This resolves repeated updates according to their relative order without relying on external positional encodings, allowing irrelevant intervening tokens to be ignored before recency is applied. We formalize the mechanism and show that its final attention support is contained in the sparse candidate set, that filtered distractors receive exactly zero mass, and that attention ratios among selected candidates are independent of global sequence length and absolute position. Experiments on repeated associative recall and copy tasks show strong extrapolation far beyond the training context length. On long-context evaluation and language-model pretraining, SRA improves long-context behavior while remaining competitive with standard Transformer baselines. These results suggest that content-first, time-second state resolution is a useful inductive bias for long-context models that must track evolving information.
Abstract: Training powerful AI systems to exhibit desired behaviors hinges on the ability to provide accurate human supervision on increasingly complex tasks. A promising approach to this problem is to amplify human judgement by leveraging the power of two competing AIs in a debate about the correct solution to a given problem. Prior theoretical work has provided a complexity-theoretic formalization of AI debate, and posed the problem of designing protocols for AI debate that guarantee the correctness of human judgements for as complex a class of problems as possible. Recursive debates, in which debaters decompose a complex problem into simpler subproblems, hold promise for growing the class of problems that can be accurately judged in a debate. However, existing protocols for recursive debate run into the : a dishonest debater can use a computationally efficient strategy that forces an honest opponent to solve a computationally intractable problem to win. We mitigate this problem with a new recursive debate protocol that, under certain stability assumptions, ensures that an honest debater can win with a strategy requiring computational efficiency comparable to their opponent.
Authors: MUHAMMAD SABIH, Frank Hannig, Jürgen Teich
Abstract: Activation functions are considered an essential primitive for neural nonlinearity, i.e., they enable neural networks to serve as universal approximators. In this paper, we show that this nonlinearity can also be achieved by input-conditioned threshold gating through branches as a universal primitive. We demonstrate that standard activations—whether piecewise-linear (ReLU, PReLU, Hardtanh) or smooth (SiLU, Sigmoid, Tanh, GELU)—are in fact instances of a single Threshold Gating (TG) primitive. For softmax, we show that it admits an exact TG conversion via its equivalent per-element Sigmoid form. We then validate these equivalences by converting pretrained networks across CNNs, transformer-based models, and recurrent architectures, preserving model performance without requiring retraining. Threshold Gating also enables training from scratch that goes beyond replacing existing activations, enabling gains in model compression, performance, and shorter training. We also propose a 'Minimal Branch Theorem' which relates the minimum number of required branches in our primitive to the trainability of general deep neural networks. In terms of hardware implementation, TG maps to a unified implementation in the case of analog in-memory systems, addressing the bottleneck of analog-to-digital and digital-to-analog converters (ADC/DAC) that is known to significantly impact power consumption and on-chip area.
PaperID: 953, Poster
Authors: PRAKUL S HIREMATH
Abstract: Asynchronous Partial Rollouts (APR) accelerate RLHF by training on the first k of N parallel completions, but this speed hides a critical flaw: because autoregressive generation time scales linearly with token count, a wall-clock cutoff acts as an invisible length filter. On reasoning tasks where reward correlates positively with trajectory length, APR systematically deletes the most valuable, multi-step reasoning chains from the training batch. We formalize this phenomenon as Latency-Conditioned Selection Bias (LCSB). We prove that whenever reward and selection probability negatively correlate (\mathrmCov_p_\theta(R, \alpha) < 0), the APR objective structurally attenuates and can even invert the on-policy gradient. Crucially, standard importance-sampling corrections (e.g., V-trace) cannot resolve this because all trajectories strictly originate from the behavior policy. After empirically validating this latency-length filter in a deployed vLLM system and demonstrating catastrophic, monotonic performance degradation on reasoning benchmarks, we provide two actionable solutions. We introduce Effective Gradient Information (EGI), a diagnostic metric that detects LCSB thousands of steps before reward curves diverge, and Selection-Probability Weighting (SPW), a near-zero-overhead correction that mathematically neutralizes the bias to recover on-policy reasoning performance without sacrificing asynchronous throughput.
Abstract: Recent works on language identification and generation have established tight statistical rates at which these tasks can be achieved. These works typically operate under a strong realizability assumption: that the input data is drawn from an unknown distribution necessarily supported on some language in a given collection. In this work, we relax this assumption of realizability entirely, and impose no restrictions on the distribution of the input data. We propose objectives to study both language identification and generation in this more general ''agnostic'' setup. Across both problems, we obtain novel interesting characterizations and nearly tight rates.
Authors:
Zhixia Zhang, Zixuan Huang, Gonxun Li, Huaiyang Wang, Chengyi Yuan, Xin Xia, deqing wang, Fuzhen Zhuang, Shuai Ma, Ning Ding, Yaodong Yang, Yikun BanAbstract: We introduce Heterogeneous Agent Collaborative Reinforcement Learning (HACRL), a new Reinforcement Learning from Verifiable Reward (RLVR) problem that addresses the inefficiencies of isolated multi-agent on-policy optimization. HACRL enables collaborative optimization with independent execution: heterogeneous agents share verified rollouts during training to mutually improve, while operating independently at inference time. Unlike LLM-based multi-agent reinforcement learning (MARL), HACRL does not require coordinated deployment, and unlike on-/off-policy distillation, it enables bidirectional mutual learning among heterogeneous agents rather than one-directional homogeneous teacher-to-student transfer. Building on this problem, we propose HACPO, a collaborative RL algorithm that enables principled rollout sharing to maximize sample utilization and cross-agent knowledge transfer. To mitigate capability discrepancies and policy distribution shifts, HACPO introduces four tailored mechanisms with theoretical guarantees on unbiased advantage estimation. Extensive experiments across diverse heterogeneous model combinations and reasoning benchmarks show that HACPO consistently improves all participating agents, outperforming GSPO with double rollouts by an average of 3.6% while using only half the rollout cost.
PaperID: 956, Poster
Authors: Amin Hosseininasab, Steven M Shugan
Abstract: Class-imbalance methods typically assume that observed minority instances are an unbiased sample of their class. However, in many real-world settings, minority observability depends on both class label and feature values---for example, while the true distribution of positive instances is uniform across demographics, they are less likely to be observed for particular demographic subgroups. This leads to a localized minority imbalance problem, posing a deeper challenge beyond general class-count imbalance. We show that under localized minority imbalance, existing imbalance mitigation techniques can overfit the observed training data and generalize poorly to under-observed regions of the minority distribution. To address this, we propose a tree-based stratification approach that recursively partitions the feature space to construct approximately unbiased subsets of the training data. For each stratum, we pair its majority instances with the full observed minority set and train a base classifier to create an ensemble. Extensive experiments over benchmark tabular datasets simulated with localized minority imbalance show that our approach outperforms popular and state-of-the-art imbalance mitigation techniques. We also introduce a gold-standard evaluation protocol that uses unbiased test sets, and demonstrate that conventional hold-out evaluation from the same localized-imbalance data can substantially bias performance. Overall, our results highlight that the cause of imbalance is as important as the correction method.
Abstract: We show that robustness to post-training quantization (PTQ) is a transferable direction in weight space. We call this direction the \emphquantization vector: extracted from a donor task by simple weight-space arithmetic, it can be used to patch a receiver model and improve post-PTQ Top-1 accuracy by up to \(60\)% in a 3-bit setting, without receiver-side quantization-aware training (QAT). Because the method requires no receiver training data, it provides a zero-shot, low-cost alternative to QAT for extremely low-bit deployment. Across multiple vision/language models and more than 30 tasks, donor quantization vectors often yield substantial gains even when donor and receiver tasks differ markedly. We further prove rigorously that quantization vectors are well-defined and do not suffer from reparameterization symmetries, and provide a local geometric account of their effects. Together, these results suggest that quantization robustness can be partially isolated, reused, and transferred through simple weight-space algebra.
PaperID: 958, Poster
Authors: Edward Raff, Michael Slawinski
Abstract: Hartigan's k-means algorithm has theoretical advantages over the standard Lloyd's variant, which has become synonymous with ``k-means''. Yet, Hartigan's algorithm is far slower in practice due to its inherently sequential calculations. We devise a Batched Hartigan's method that enables vectorized calculation by exploiting the triangle inequality to expose a set of parallel-safe updates and a smaller set of sequential updates. This makes Hartigan's method competitive with Lloyd's in execution speed for the first time.
PaperID: 959, Poster
Abstract: Graph Contrastive Learning (GCL) is an effective approach for learning node representations. However, the intrinsic uncertainty of real-world graph data introduces significant challenges, as GCL relies on crisp deterministic representations that struggle to capture such ambiguity. To model graph uncertainty, fuzzy logic has attracted increasing attention due to its mathematical rigor. Existing methods, however, simply inject fuzzy modeling into the contrastive framework without tailoring data augmentation to the needs of fuzzy representations, resulting in limited discriminability and suboptimal downstream performance. To address these limitations, we propose a novel Fuzzy Graph Contrastive Learning (FGCL) framework for uncertainty-aware node representation learning. It adopts a vertex entanglement-based augmentation strategy to generate tailored multi-views with controlled perturbations while preserving core graph topology and semantics, mitigating semantic corruption. We also design a Deep Weighted Fuzzy Graph Convolutional Neural Network (DWFGCNN) encoder, which maps crisp features to fuzzy representations via learnable membership functions, explicitly models uncertainty, and resolves the core contradiction that traditional encoders cannot balance discriminability and robustness simultaneously. Extensive experiments on multiple datasets demonstrate that FGCL consistently outperforms baselines in node classification, community detection, and link prediction tasks, validating its effectiveness in handling uncertainty in graph representation learning.
PaperID: 960, Poster
Abstract: Human players in cooperative games often rely on conventions: reusable patterns of play that coordinate well when used by the whole team, but can fail otherwise. Evaluating agents in this setting is difficult because human demonstration datasets are usually unlabeled: trajectories do not identify which player, group, or convention produced them. As a result, standard proxy-based evaluations can miss important population diversity. In particular, proxy bots produced by regularised self-play around a pooled behavioural cloning policy can remain concentrated around a single dominant mode of play. We introduce a pipeline for recovering and evaluating this hidden convention diversity from unlabeled demonstrations. Our method first learns latent convention structure from the dataset, then selects a compact set of distinct but plausible conventions, and finally reconstructs each selected convention into a high-performing playable proxy using KL-regularised PPO. A key practical challenge is that KL-regularised PPO is sensitive to the KL penalty, which is often tuned using online interaction. We provide an empirical offline calibration procedure for choosing this KL penalty in the PPO+KL reconstruction settings studied here. On Hanabi, we find that the BC-derived proxy suite is highly concentrated, and that the relative ordering of agents changes when evaluation uses the latent proxy suite instead.
PaperID: 961, Poster
Abstract: We introduce Gram-Calibrated Anchoring (GCA) for class-incremental learning (CIL) with pre-trained Vision Transformers. GCA rests on the observation that, under continual low-rank adaptation, ---class means extracted once from the frozen pre-trained backbone---are effective substitutes for prototypes recomputed from the adapted model; a formal stability analysis bounds decision-boundary shifts in terms of feature-level prototype drift, while empirical LoRA perturbation measurements support the small-drift regime. This motivates fixing the classification head to anchor prototypes computed upon each task's arrival rather than recalibrating prototypes after training. The fixed anchor geometry then enables (GCI), which applies the regularized inverse of the anchor Gram matrix to deconvolve inter-class correlation that standard cosine classification ignores, and normalizes each coordinate by its estimation standard deviation for fair cross-class comparison. The resulting method trains a single continually evolved LoRA module with no exemplar storage or expanding adapter pools. Experiments on ImageNet-R, ImageNet-A, CIFAR-100, and CUB-200 with multiple backbones show that GCA achieves competitive accuracy across four CIL benchmarks while significantly reducing training cost over recent methods.
PaperID: 962, Poster
Abstract: We initiate the study of safe stochastic linear bandits under quantum feedback models. The learner must maximize an unknown linear reward while satisfying unknown linear constraints at every interaction. Under coherent unitary access to the joint reward-and-constraint distribution, we develop QRS-COLTS, a staged quantum analog of constrained linear Thompson sampling. With high probability, QRS-COLTS plays only feasible actions and achieves regret polylogarithmic in the quantum query budget; an isotropic multivariate estimator reduces the vector-feedback cost from linear to square-root dependence on the number of constraints, up to logarithmic factors. We also formulate a stronger query-weighted coherent-exploration model with a safe-action bomb flag. In this model, the learner may query action superpositions, the bomb provides exact but dangerous safe-set access, and regret is charged by the action weights in the query state. We present a quantum algorithm, BCO-QLinTS, that uses bomb-certified quantum convex optimization to select safe actions directly, thereby avoiding the need to estimate the constraint matrix. Its query-weighted regret is polylogarithmic in the horizon, with a bomb-safe optimization overhead that depends only polynomially on the number of constraints.
PaperID: 963, Poster
Authors: Andrei Ermakov
Abstract: Low-rank approximation captures one important notion of matrix simplicity: a matrix is simple if it is well described by a low-dimensional linear model. This paper studies a complementary notion. We consider matrices that become structurally simple only after permuting rows and columns, so that their mass concentrates in a prescribed geometric pattern such as a band, block-diagonal, arrowhead, or related form. We introduce a pattern distance that measures how close a matrix is to a given structural mask after optimal reordering, giving a common framework for latent ordering, community detection, and bandwidth minimization. The problem is NP-hard in general; we derive an alternating algorithm based on Linear Assignment subproblems, and establish identifiability of the target structure up to pattern automorphisms. We further introduce a pattern-plus-low-rank decomposition that combines permutation-structured fitting with a low-rank residual; it dominates pure low-rank approximation by feasibility and admits a strict rank separation on planted block-diagonal targets, with broad empirical gains across pattern families. We evaluate the framework on synthetic recovery tasks, real correlation and graph adjacency matrices, and transformer attention heads (GPT-2, Qwen2.5-0.5B and Qwen2.5-7B), showing that a single universal algorithm recovers qualitatively different forms of matrix structure.
PaperID: 964, Poster
Abstract: We present an algorithm that distills transformer encoders into logical formulas. Interpretable surrogates for transformers are important for safety and scientific knowledge discovery. Recent work extracts sparse latent features from transformer activations, but these features alone do not express how information is composed across a sequence. Building on theoretical connections between transformer expressivity and linear temporal logic (LTL), we introduce iterated temporal decision trees (ITDTs), a model that can express any LTL formula as a sequence of decision trees. Each decision tree builds on the previous ones, producing LTL formulas of increasing temporal depth. Our algorithm distills a transformer encoder into an ITDT. Empirically, our algorithm recovers succinct and accurate formulas on established LTL learning benchmarks and distills interpretable rules for multiple protein sequence datasets from ESM-2-T6, an established protein language model.
PaperID: 965, Poster
Abstract: Many machine learning tasks are invariant under the action of a group G of transformations: signal classification can be invariant under translations, image classification under 2D rotations, and spherical-image classification under 3D rotations. The G-bispectrum is a principled complete invariant of a signal )retaining all signal's information up to the group action) with proven benefits in machine learning and as a pooling layer in deep networks. However, its deployment has been hampered by high computational cost and a patchwork of group-specific implementations. We present bispectrum, an open-source, fully unit-tested PyTorch library that implements selective G-bispectra for seven different group actions, as differentiable modules that can be directly incorporated into machine learning pipelines and deep learning architectures. For finite groups G, selectivity reduces the computational cost from O(|G|^2) to O(|G|). For planar rotations, we leverage the disk bispectrum. For spherical 3D rotations, we introduce an augmented selective bispectrum at band-limit L which reduces the cost from O(L^3) to \Theta(L^2) coefficients. We profile the entire library (for
PaperID: 966, Poster
Abstract: Multimodal dataset distillation (MDD) seeks to synthesize compact surrogates from large-scale image-text corpora, reducing the high cost of pre-training. Although recent feature-matching methods have improved the efficiency of MDD, they often suffer from initialization overfitting, which distorts the cross-modal structure and severely degrades retrieval performance under unseen initializations. To mitigate this issue, we propose MDA-DD, which enhances generalization via a dynamic staggered model queue. We further identify two architectural bottlenecks in prior work: structurally redundant projections and inefficient covariance computation. For the former, we design an asymmetric dual projection to reduce redundancy while preserving balanced bidirectional cross-modal alignment. For the latter, we introduce a Second-Order Cross-Modal Moment (SOCM) objective, which efficiently captures both mean and correlation statistics across modalities. Experiments on Flickr30K and MS-COCO show that MDA-DD consistently surpasses existing methods, yielding notable retrieval improvements (e.g., +12.9 IR@1 and +18.2 TR@1 in the 100-pair setting) and matching or exceeding baselines trained on 500 pairs using only 100 pairs.
Abstract: Human motion prediction combines the tasks of trajectory forecasting and human pose prediction. For each of the two tasks, specialized models have been developed. Combining these models for holistic human motion prediction is non-trivial, and recent methods have struggled to compete on established benchmarks for individual tasks. To address this, we propose a simple yet effective transformer-based model for human motion prediction. The model employs a stack of self-attention modules to effectively capture both spatial dependencies within a pose and temporal relationships across a motion sequence. This simple, streamlined, end-to-end model is sufficiently versatile to handle pose-only, trajectory-only, and combined prediction tasks without task-specific modifications. We demonstrate that this approach achieves state-of-the-art or on-par results across all tasks through extensive experiments on a wide range of benchmark datasets, including Human3.6M, AMASS, ETH-UCY, and 3DPW. Code will be released.
PaperID: 968, Poster
Authors: Lena Ehrmuth
Abstract: We define a new variant of transformers called recursive transformers, which generate sequences of vectors without discretisation in each step. We show that over ordered ring extensions of the integers recursive transformers with circuit activation functions and hard attention are computationally equivalent to the well-established algebraic model of BSS-machines. If the transformers use feedforward networks instead, they are equivalent to linear BSS-machines with degree-2 branching.
PaperID: 969, Poster
Abstract: Image-to-image translation aims to generate a target image conditioned on an observed source. While diffusion bridge models perform well in one-to-one settings, they are not adapted to problems where multiple sources provide complementary information about a single target. We introduce (MB-DDPMs), a framework for many-to-one translation that builds stochastic bridges between the target and each source, coupled through a unified reverse process that aggregates information across sources. We evaluate MB-DDPM on brain MRI for Multiple Sclerosis to predict FLAIR contrast from T1, T2 and PD contrasts. We demonstrate strong clinical relevance through near-perfect agreement with ground-truth lesion-based measurements, highlighting its potential for reliable downstream medical analysis. We further evaluate MB-DDPM on image restoration using CIFAR-10 and ImageNet. Across all settings, MB-DDPM achieves state-of-the-art performance, consistently outperforming existing baselines in both reconstruction quality and perceptual metrics.
PaperID: 970, Poster
Abstract: We study token dynamics in deep Mamba models, with a particular focus on the attracting role of the first token. Unlike attention mechanisms, where token interactions are governed by positive scalar coefficients, Mamba dynamics involve full matrix-valued interaction coefficients with no positivity or definiteness guarantees. This structural difference makes Mamba dynamics substantially harder to analyze and prevents the direct application of existing techniques for attention-based models. To overcome this challenge, we introduce a new analytical framework based on a high-dimensional invariant box and a sharp exit-time argument, which together recover an effective form of sign-definiteness. Using this framework, we characterize the possible limit points of token trajectories and prove explicit exponential convergence rates. Our analysis reveals that the first token has privileged dynamics: its limiting direction attracts subsequent tokens, inducing token clustering and attention-sink phenomena in deep Mamba models. We further show that this attraction mechanism extends to broader sequence dynamics, including state-space models, causal attention, and classical time-series models. Experiments on Falcon-Mamba-7B validate our predictions, showing that concentration on the first token increases with depth and dominates in later layers.
PaperID: 971, Poster
Abstract: Large Reasoning Models (LRMs) produce explicit chains of thought to improve transparency, but they frequently overthink, generating long yet inefficient reasoning chains that inflate computational costs and risk introducing errors. While numerous reinforcement learning methods struggle to penalize verbosity to achieve conciseness, they typically apply the penalty uniformly across lengthy sequences, which leads to credit confounding, that is, the failure to distinguish essential reasoning steps from superfluous ones. In this paper, we study a measurable proxy for this confounding: high-entropy tokens that often coincide with surface transition points where reasoning re-enters verification. Based on this insight, we propose Adaptive Entropy-Sparing (AES), which selectively spares high-entropy tokens that occur before an efficient in-group reference, while penalizing overlength high-entropy suffix tokens associated with superfluous branching. Our AES thereby navigates the suppression-exploration trade-off, reducing overthinking while preserving average reasoning accuracy. We also present a stylized entropy-length model that motivates why local uncertainty can serve as a proxy for future reasoning cost. Across eight reasoning benchmarks, AES reduces reasoning length by up to 52.9% while increasing mean accuracy by up to 3.7% over the base models, consistently outperforming existing efficiency-oriented methods.
Abstract: We study zero-shot coordination (ZSC), where independently developed agents must coordinate at test time. While ZSC has been well studied in the RL literature, far less is known about the performance of LLM agents despite their increasing deployment in such settings. Existing work on LLM coordination often relies on specialised scaffolding that independently developed agents are unlikely to share in practice. Moreover, evaluations on complex environments (e.g., Hanabi) make it difficult to pinpoint the sources of coordination failure. In contrast, we focus on simple, general-purpose scaffolds in minimal environments designed to isolate specific coordination challenges. Our results show that even in these controlled settings, frontier LLM agents struggle to coordinate, largely due to limited understanding of the coordination problem and weak reasoning about their partner’s beliefs. While LLMs introduce as an additional axis for coordination, they nonetheless fail to exploit this structure effectively. Towards this, we propose as a principled approach for enabling robust coordination among LLM agents. Finally, we show that CFDs can be discovered automatically, removing the need for manual engineering.
PaperID: 973, Poster
Authors: Chengrong Ye, Xinyan Gong, Lingwei Zhang, Yue Song
Abstract: Natural sequences often contain signals that move over time, such as translating objects, rotating digits, and drifting storm cells. Selective state space models such as Mamba process such data efficiently, but their recurrent state is updated at fixed spatial coordinates. As a result, evidence from a moving object can be accumulated across inconsistent locations rather than in the object's moving reference frame. To address this problem, we introduce Flow-SSM, a flow-equivariant selective state space model that lifts recurrent memory to candidate motion branches and transports each branch along its corresponding flow before applying the selective update. To handle real-world sequences where broad motion and local deformation often occur at different scales, we further introduce Hierarchical Flow-SSM, which organizes flow-aligned memory in a coarse-to-fine hierarchy rather than a single flat representation. We validate our architecture progressively: demonstrating exact transformation capture on Moving and Rotating MNIST, interpretable flow structures on KTH action videos, and the necessity of our multi-scale design for complex forecasting on the SEVIR weather dataset. Beyond merely improving predictive accuracy, our analyses confirm that the explicit flow alignment and hierarchical communication contribute measurably to the observed gains. These results support transported recurrence as an effective inductive bias for spatiotemporal sequence modeling.
Abstract: Model merging combines fine-tuned models into a unified one without joint training or access to original data. Dynamic merging improves flexibility by selectively activating task-relevant parameters, but existing methods either maintain a full shared model with tiny experts or allocate excessive capacity to experts, leading to suboptimal accuracy–efficiency trade-offs. We propose to balance shared and per-task expert parameters. The method casts parameter budgeting as differentiable rank optimization over low-rank modules, supervised data-free by the original task vectors, and applies a brief refinement stage to recover task fidelity. Across vision, language, and multimodal benchmarks, DiDi-Merging matches prior dynamic baselines at
Authors:
Sheheryar Mehmood, Florian Knoll, Peter OchsAbstract: Algorithm unrolling is ubiquitous in machine learning, particularly in hyperparameter optimization and meta-learning, where Jacobians of solution mappings are computed by differentiating through iterative algorithms. Although unrolling is known to yield asymptotically correct Jacobians under suitable conditions, recent work has shown that the derivative iterates may initially diverge from the true Jacobian, a phenomenon known as the curse of unrolling. In this work, we provide a non-asymptotic analysis that explains the origin of this behavior and identifies the algorithmic factors that govern it. We show that truncating early iterations of the derivative computation mitigates the curse while simultaneously reducing memory requirements. Finally, we validate our theoretical findings both on synthetic problems and in a meta-learning setting for few-shot classification, where we study the behavior of gradients obtained via unrolling and implicit differentiation.
Authors: Tigran Ramazyan, Denis Derkach
Abstract: We study zeroth-order optimisation under context distributional uncertainty, a setting commonly tackled using Bayesian optimisation (BO). A prevailing strategy to make BO more robust to the complex and noisy nature of data is to employ an ensemble as the surrogate model, thereby mitigating the weaknesses of any single model. In this study, we propose a novel algorithm for Ensemble Distributionally Robust Bayesian Optimisation that remains computationally tractable while managing continuous context. We obtain theoretical sublinear regret bounds, improving current state-of-the-art results. We show that our method’s empirical behaviour aligns with its theoretical guarantees.
Abstract: AdaBoost sequentially fits so-called weak learners to minimize an exponential loss, which penalizes misclassified data points more severely than other loss functions like cross-entropy. Paradoxically, AdaBoost generalizes well in practice as the number of weak learners grows. In the present work, we introduce Penalized Exponential Loss (PENEX), a new formulation of the multi-class exponential loss that is theoretically grounded and, in contrast to the existing formulation, amenable to optimization via first-order methods, making it a practical objective for training neural networks. We demonstrate that PENEX effectively increases margins of data points, which can be translated into a generalization bound. Empirically, across computer vision and language tasks, PENEX improves neural network generalization in low-data regimes, matching and in some settings outperforming established regularizers at comparable computational cost. Our results highlight the potential of the exponential loss beyond its application in AdaBoost.
PaperID: 978, Poster
Abstract: Personalized federated continual learning (pFCL) alleviates catastrophic forgetting across tasks and data heterogeneity across clients. We view both challenges as interference between knowledge and propose Bayesian frameworks that define knowledge as posterior belief and quantify two types of interference: intra-model interference, measuring task-induced posterior drift, and inter-model interference, measuring aggregation-induced posterior drift. We show that such interference bounds catastrophic forgetting and data heterogeneity-induced loss, respectively. We then develop an interference-regularized local objective to guide personalization under catastrophic forgetting and data heterogeneity. The framework unifies standard FCL and pFL categories. Experiments on synthetic and real-world benchmarks demonstrate improved model performance over state-of-the-art pFCL methods.
Abstract: RL for code correctness is now established: have the model generate a program, run it against hidden test cases, and reward solutions that pass. Extending this to code optimization seems straightforward: just add execution time to the reward. But in practice, once timing drives the reward, small problems in measurement noise, reward sparsity, or GRPO instability overwhelm the signal and make RL fail: generated solutions are barely faster, and more of them can fail. We make execution time learnable through three stages: (1) how code is tested, by building DMC-Optim with large optimization tests and a calibrated sandbox; (2) how speed is turned into reward, by composing correctness and speed in the RL environment and using an offline simulator to predict the most promising configurations; and (3) how the model learns from that reward, by adapting GRPO and evaluation to the sparser, noisier timed-execution setting. On DMC-Optim, the strongest optimization-aware configurations improve strict top-decile pass@1 from 18.1% to 27.7% on Qwen 2.5 7B and from 28.7% to 42.4% on CWM 32B, while preserving most of the pure-correctness score. When the timing sandbox is degraded, robust optimization RL reaches up to roughly 130% improvement over standard RLVR. On LCB, CWM 32B wins up to 70.3% of best-sample speed comparisons against standard RLVR; relative to the fastest correct human submissions per problem, it reaches about half the human rate of complexity-class improvements (14% vs. 28%).
Authors:
Ziyue Zhu, Shangyang Wu, Shuai Zhao, Zhao ZhiQiu, Jian Zhang, Li Shengjie, Yi Wang, Anh Tuan Luu, XINLIANG ZHOU, Fang Li, Haoran LuoAbstract: Vision-Language-Action (VLA) models are formulated to ground instructions in visual context and generate action sequences for robotic manipulation. Despite recent progress, VLA models still face challenges in learning related and reusable primitives, reducing reliance on large-scale data and complex architectures, and enabling exploration beyond demonstrations. To address these challenges, we propose a novel Neuro-Symbolic Vision-Language-Action (NS-VLA) framework via online reinforcement learning (RL). It introduces a symbolic encoder to embedding vision and language features and extract structured primitives, utilizes a symbolic solver for data-efficient action sequencing, and leverages online RL to optimize generation via expansive exploration. Experiments on robotic manipulation benchmarks demonstrate that NS-VLA outperforms previous methods in both one-shot training and data-perturbed settings, while simultaneously exhibiting superior zero-shot generalizability, high data efficiency, and an expanded exploration space. Our code is available at https://anonymous.4open.science/r/NS-VLA/.
PaperID: 981, Poster
Abstract: Speculative decoding accelerates large language model inference by using a draft model to generate candidate tokens, which are verified by the target model in a single forward pass. Verification proceeds and discards every position from the first rejection onward, yet existing draft training relies on token-level cross-entropy with a fixed per-position weighting that does not reflect this process. We introduce Verification-Aware Training (VAT), a plug-in framework that simulates verification at every training step and turns the resulting accept and reject patterns into supervision. VAT consists of two components: (i) a verification head, a lightweight jointly-trainable binary classifier that validates per-position acceptance via an explicit prediction target; (ii) verification-adaptive weighting replaces the fixed weighting schedule with weights adapted to each sample's first rejection point during training. VAT is model-agnostic and can be layered on top of existing methods without changing the draft architecture, the target model, or the inference procedure, and trained jointly with the original loss. Applied to EAGLE-3 and DFlash on Qwen3-4B, Qwen3-8B and LLaMA3.1-8B, VAT improves average acceptance length by up to 11.4% and wall-clock speedup by up to 8.7%, with consistent gains across math, code, and chat benchmarks.
PaperID: 982, Poster
Abstract: Conformal prediction provides set-valued predictions with distribution-free coverage guarantees, making it attractive for high-stakes image classification. However, split conformal prediction is data-inefficient, while full conformal prediction (FCP), despite its stronger statistical efficiency, is computationally prohibitive at scale because it requires candidate-specific model refits at test time. We address this limitation by leveraging zero-shot vision-language models (VLMs) to guide scalable FCP in large label spaces. We introduce Targeted Full Conformal Prediction (T-FCP), which uses a lightweight inductive conformal predictor to prune unlikely labels and applies FCP only to the remaining candidates, reducing computation while retaining the formal guarantee of the combined conformal procedure. We further propose Stabilized Online LDA (SO-LDA), an efficient VLM adaptation solver based on rank-one inverse-covariance updates. Across multiple benchmarks, including ImageNet, T-FCP enables practical full-conformal image classification with modest test-time overhead, yielding efficient prediction sets and more stable empirical coverage than split conformal alternatives.
Abstract: Large language model (LLM) agents face a structural tension: cloud agents provide strong reasoning but expose user data, while on-device agents preserve privacy at the cost of overall capability. Existing device-cloud designs treat this boundary as a compute split rather than a trust boundary suited to agentic workloads, and existing sanitizers force a choice between policy flexibility and the structural fidelity tool calls require. In this work, we develop PAAC, a privacy-aware agentic framework that aligns planner--executor decomposition with the device-cloud boundary so that role specialization itself becomes the privacy mechanism. The cloud agent reasons over typed placeholder tokens that preserve each sensitive value's reasoning role while discarding its content, while the on-device agent identifies sensitive spans and distills each step's execution outcome into compact key findings. Sanitization confines the on-device LLM to proposing which spans to mask, while a deterministic registry performs all substitution and reversal, keeping actions directly executable on device. On three agentic benchmarks under strict privacy settings, PAAC dominates the Pareto frontier of privacy and accuracy, improving average accuracy by 15-36% and reducing average leakage by 2-6× over state-of-the-art device-cloud baselines, with the largest margins on privacy targets outside fixed entity taxonomies. We find consistent improvements on 17 additional benchmarks spanning 10 domains, including math, science, and finance.
Abstract: Despite the success of foundation models in language and vision, molecular graph generation still lacks a unified framework for heterogeneous design tasks with reliable controllability. While reinforcement learning (RL) offers a natural post-training mechanism for task-specific optimization, applying it to graph generative models is hindered by the vast atom-wise action spaces and chemically invalid intermediate states. We propose Controllable Molecular Generative Foundation Models (CoMole), built with a unified motif-aware graph diffusion pipeline. By learning a motif-aware graph space, CoMole transfers pretrained structural priors into controllable generation, where RL optimizes conditional reverse policies over chemically meaningful decisions. We theoretically characterize the bottleneck of atom-level RL and justify motif-aware policy optimization. Across three heterogeneous benchmarks spanning materials and drug discovery, CoMole ranks first in controllability on all nine targets, reduces MAE by up to 48.2% relative to the strongest baselines, and maintains validity above 0.94 without rule-based correction or post-hoc filtering. We further show that CoMole transfers controllability to unseen properties by optimizing only task embeddings with the generator frozen, achieving performance competitive with strong task-specific baselines.
Authors:
Zhenxiang Xu, Sirui Chen, Yong He, Jieyu Yang, Chuan Yuan, Ke Ding, Jing Cao, Can Wang, Jiawei ChenAbstract: Generative recommendation (GR) aligns with advances in generative AI by casting next-item prediction as token-level generation rather than score-based ranking. Most GR methods adopt a two-stage pipeline: (i) item tokenization, which maps each item to a sequence of discrete, hierarchically organized tokens; and (ii) autoregressive generation, which predicts the next item's tokens conditioned on the tokens of user's interaction history. Although hierarchical tokenization induces a prefix tree (trie) over items, standard autoregressive modeling with conventional Transformers often flattens item tokens into a linear stream and overlooks the underlying topology. To address this, we propose TrieRec, a trie-aware generative recommendation method that augments Transformers with structural inductive biases via two positional encodings. First, a trie-aware absolute positional encoding aggregates a token's (node's) local structural context (\eg depth, ancestors, and descendants) into the token representation. Second, a topology-aware relative positional encoding injects pairwise structural relations into self-attention to capture topology-induced semantic relatedness. TrieRec is also model-agnostic, efficient, and hyperparameter-free. In our experiments, we implement TrieRec within three representative GR backbones, achieving notably improvements of 8.83% on average across four real-world datasets.
PaperID: 986, Poster
Abstract: Offline reinforcement learning (RL) typically mitigates distribution shift by imposing divergence constraints on the action distributions of the learned and behavior policies. However, this paradigm cannot ensure the underlying dynamic consistency structure: even minor action deviations can lead to drastically different transitions, inadvertently driving the agent away from the offline dataset. To address this, we explore policy constraints through their induced state-consequence distributions rather than relying on pure action-distribution matching. Nevertheless, directly enforcing such consequence-aware constraints is computationally intractable without access to the true underlying dynamics model. To bypass this issue, we propose Structural Behavior Regularization (StBR), which replaces the intractable objective with a closed-form latent divergence. Structuring this latent space to reflect transition dynamics renders the latent divergence a theoretically bounded surrogate for the true consequence-distribution divergence. By regularizing the learned policy with this latent divergence, StBR adaptively tightens the behavior constraint in dynamics-sensitive regions, effectively preventing deviations outside the offline dataset. Empirically, we demonstrate that StBR achieves superior performance across diverse offline RL benchmarks.
Abstract: Inductive biases influence the behavior and performance of sequential models. In this work, we study an underexplored inductive bias in sequential modeling: continuity in time. We ask a simple question: do models motivated by continuous-time formulations, such as state-space models, actually behave continuously in time, and does this translate into better performance on tasks with continuous temporal structure? To answer this, we formalize model continuity as convergence under temporal refinement, where a model is continuous if its predictions approach an underlying continuous trajectory as the temporal discretization is refined. We show that S4 exhibits stable continuous behavior, whereas S6 (the core of Mamba) can be more sensitive to input amplitude and selective dynamics, despite being derived from a continuous dynamical system. To study whether this distinction matters for learning, we also need a corresponding notion of task continuity. We therefore introduce a metric to quantify the continuity of datasets directly from their temporal structure. Across benchmarks, we find a clear empirical alignment between task continuity, model continuity, and model performance. Beyond an inductive bias, continuity also has practical consequences: we show that it enables a simple temporal subsampling strategy that improves both efficiency and performance.
Abstract: Off-policy reinforcement learning of pretrained flow policies remains challenging due to the instability of optimization arising from the multi-step sampling process. Recently, Q-learning with Adjoint Matching (QAM) addressed this issue by reformulating into a memoryless stochastic optimal control (SOC) problem with a learned critic. However, QAM inherits a fundamental fragility of critic-guided improvement: small critic errors are amplified when critics are ill-conditioned, often leading to model collapse. This paper introduces Trust Region Q-Adjoint Matching (TRQAM), a stable off-policy fine-tuning algorithm that adaptively controls the path-space KL with pretrained flow policies through projected dual descent. Specifically, we optimize the trust-region parameter \lambda in SOC dynamics, and theoretically show that the path-space KL can be represented by a closed-form function of \lambda. As a result, our method can precisely control the exact deviation from pretrained flow policies, achieving stable off-policy RL. Through experiments on 50 OGBench tasks, TRQAM consistently outperforms prior arts in both offline RL and offline-to-online RL. In particular, TRQAM achieves an overall success rate of 68% in offline RL, substantially improves the strongest baseline at 46%.
Authors:
Julien Dallot, Darya Melnyk, Tijana Milentijević, Stefan Schmid, Patrik WeltersAbstract: The Byzantine Agreement (BA) problem is a fundamental task in distributed computing where nodes in a network need to agree on a common output in the presence of arbitrary (worst-case) node failures. In real-world applications, nodes can be monitored over time to provide ML predictions of faulty behavior in a network. In this work, we answer the question of whether predictions can improve the fault tolerance of a distributed system without sacrificing correctness when the predictions are wrong. We consider a prediction-augmented study of Byzantine agreement in which each node receives, in addition to its input bit, a predicted set of honest nodes. We focus on algorithmic resilience --- the maximum number of faulty nodes an algorithm can tolerate --- and present algorithms and impossibility results whose resilience depends on the accuracy of the predictor. As our first main result, we bring a complete characterization of the consistency--robustness trade-offs in both the non-authenticated and authenticated settings: for n nodes and a parameter \alpha \in [0, 1], we present algorithms that tolerate up to \alpha \cdot n faulty nodes when the predictor is correct (consistency), and up to \frac1-\alpha2 \cdot n - 1 faulty nodes when the predictor is arbitrarily wrong (robustness). In the authenticated setting, the robustness bound improves to (1-\alpha) \cdot n - 1. We prove matching impossibility results, showing that these tradeoffs are optimal and independent of the particular prediction system. Our second main result characterizes smoothness: the rate at which resilience degrades as the predictor becomes less accurate. We show that resilience linearly decreases in the number of wrong predictions as long as that number stays within a constant fraction of n. Concretely, in the non-authenticated setting, each additional wrong prediction loses one unit of resilience, whereas in the authenticated setting, the decline is halved, since two wrong predictions are needed to lose one unit of resilience.
Abstract: Explanations are central to causal reasoning, and cognitive science has long established that the human drive to explain is itself a mechanism for learning about causality. Despite this, learning from those abductive signals is largely ignored in artificial intelligence. While explainable AI (XAI) increasingly draws on causal models to generate explanations, the converse direction about remains largely unexplored. To fill this gap, we propose xWhyL, a formal framework connecting causality and XAI by learning causal models from explanations. We develop a mathematical theory that translates explanations into a learning signal complementary to observational data, and demonstrate how it enables overcoming the limits of observational causal discovery. As explanations can be derived from incorrect beliefs and clash with data, a tension we call the , we prove conditions under which our framework rejects misspecified explanations rather than absorbing them. Our practical instantiation, (CIL), shows how expert explanations can efficiently support causal discovery and distinguish correct from incorrect explanations.
Abstract: Representation Autoencoders (RAE) replace traditional VAE with pretrained vision encoders. In this paper, we systematically investigate several design choices and find three insights which simplify and improve RAE. First, we study a generalized formulation where the representation is defined as sum of the last k encoder layers rather than solely the final layer. This simple change greatly improves reconstruction without encoder finetuning or specialized data (e.g., text, faces). Second, we study the prevalent assumption that RAE (using pretrained representation as encoder) replaces representation alignment (REPA), which distills same representation to intermediate layers instead. Through large-scale empirical analysis, we uncover a surprising finding: RAE and REPA exhibit complementary working mechanisms, allowing same representation to be used as both encoder and target for intermediate diffusion layers. Finally, the original RAE struggles with classifier-free guidance (CFG) and requires training a second, weaker diffusion model for AutoGuidance (AG). We show that REPA itself can be viewed as x-prediction in RAE latent space. By simply re-parameterizing the output of DiT model, it can provide guidance for "free". Overall, RAEv2 leads to more than 10x faster convergence over original RAE, achieving a state-of-the-art gFID of 1.06 in just 80 epochs on ImageNet-256. On FDr^k RAEv2 achieves state-of-art 2.17 at just 80 epochs compared to previous best 3.26 (800 epochs) without any post-training. This motivates eFID-k (epochs to reach unguided gFID <= k) as a measure of training efficiency. RAEv2 attains an eFID-2 of 35 epochs, versus 177 for the original RAE. We also validate our approach across diverse settings for text-to-image generation and navigation world models, showing consistent improvements. We hope that this work provides insights for practical adoption of representation autoencoders.
Authors: Ian Osband
Abstract: Policy gradient computes a backward pass for every sample, even though the backward pass is expensive and most samples carry little learning value. The Delightful Policy Gradient (DG) provides a forward-pass signal of learning value: \emphdelight, the product of advantage and surprisal (negative log-probability). We introduce the \emphKondo gate, which compares delight against a compute price and pays for a backward pass only when the sample is worth it, thereby tracing a quality--cost Pareto frontier. In bandits, zero-price gating preserves useful gradient signal while removing perpendicular noise, and delight is a more reliable screening signal than additive combinations of value and surprise. On MNIST and transformer token reversal, the Kondo gate skips most backward passes while retaining nearly all of DG's learning quality, with gains that grow as problems get harder and backward passes become more expensive. Because the gate tolerates approximate delight, a cheap forward pass can screen samples before expensive backpropagation, suggesting a speculative-decoding-for-training paradigm.
PaperID: 993, Poster
Authors: Jesse Zymet
Abstract: This paper concerns the problem of scaling agents to large action spaces demanding pretrained world knowledge. What type of representation pretraining would enable policies such as language models to generalize rewards to appropriate actions? Contrastive learning aligns similar inputs and otherwise enforces distancing of representations, while non-contrastive alternatives omit distancing terms and so learn anisotropic representations that collapse distinctive clusters. Though contrastive pretraining has been proven to promote supervised learning efficiency, non-contrastive pretraining nevertheless remains canonical for language agents. Here we examine the extent to which these pretraining types support on-policy finetuning, in which efficiency is coextensive with online performance. Broadly, we argue that contrastive pretraining optimizes efficient exploration of large spaces. We prove that anisotropic representations collapse distinct actions and necessitate complex policies, whereas contrastive pretraining retains distinctions and enables reward transfer through spectral regularization. We validate our conclusions empirically by providing spectral analyses of language data and simulations involving multimodal personalization policies as well as self-improving reasoning agents finetuned online using a policy gradient. Our latter simulations suggest contrastive pretraining for language agents exploring diverse reasoning paths.
Authors: Guang Yang, Fengchen Liu, Amir Ghasemian, Zhong Wang, Ninareh Mehrabi, Homa Hosseinmardi
Abstract: The proliferation of deepfake audio challenges voice-based authentication systems; passive forensic detectors are sensitive to evolving generative models and to real-world channel distortions. We propose Asymmetric Phase Coding (APC), a training-free cryptographic signing layer for audio, designed as a compact and auditable provenance primitive that can stand alone or be stacked with learned watermarks. APC combines Ed25519 digital signatures (EdDSA, FIPS 186-5; 64-byte signatures) with Reed–Solomon error correction, pseudo-random STFT phase-bin selection, and a redundant quantization-index-modulation (QIM) code on log-magnitude differences of adjacent bin pairs, yielding a compact, non-repudiable, blind-extractable watermark. We evaluate APC on 1,000 LibriSpeech test-clean clips (10 s each, 44.1 kHz) under eight attack configurations – identity, 10% end-cropping, 20% end-cropping, 8 kHz low-pass, 16 kHz round-trip resampling, FLAC re-encoding, MP3 at 128 kbps, and OGG-Vorbis at 128 kbps – and achieve cryptographic verification rates between 97.5% and 98.3% on every condition at mean PESQ=3.02 and tens-of-milliseconds CPU latency. We explicitly compare APC against recent neural baselines (AudioSeal, WavMark, SilentCipher), detail the threat model (forgery resistance vs. erasure), characterize the dataset, define all metrics, quantify an adaptive white-box erasure attack, and release code, keys, and metadata for reproducibility.
Abstract: Claim verification is an important problem in high-stakes settings, including health and finance. When information underpinning claims is incomplete or conflicting, uncertain answers may be more appropriate than binary true or false classifications. In all cases, faithful explanations of the considerations determining the final verdict are crucial. We introduce inference-time argumentation (ITA), a trainable neurosymbolic framework for ternary claim verification in which a formal argumentation semantics giving the strength of claims is used both (i) to guide LLM training as models learn to generate arguments and assign them base scores (representing intrinsic strengths) and (ii) to compute ternary (true/false/uncertain) predictions from generated, scored arguments. As a result, at training time, argument generation and scoring can be optimised according to the quality of the induced argumentative predictions. Moreover, at inference time, the final prediction is faithful, by construction, to the arguments and scores determining the verdict, rather than being justified by a potentially unfaithful post-hoc reasoning trace as in conventional reasoning models. We finally show that, on two datasets for ternary claim verification, ITA improves upon argumentative baselines and can perform competitively against non-argumentative direct-prediction baselines, while providing verdicts that are computed deterministically from explicit, inspectable argumentative structures.
Abstract: Complex learning agents are increasingly deployed alongside existing experts, such as human operators or previously trained agents. However, it remains unclear how learners should optimally incorporate certain forms of expert data, or how to quantify the potential benefit from doing so. We study this problem in the context of Bayesian multi-armed bandits, considering , where posterior updates from different data sources may have different computational costs, so the learner must decide whether a given update should use its own experience or an outcome generated by an expert. We formalize how expert data influences the learner's posterior and quantify how pretraining or online learning on expert outcomes tighten information-theoretic regret bounds. We propose an information-directed rule for allocating a limited update budget across data sources, and we study strategies for how the learner can infer when to trust the expert, safeguarding for compromised experts. By disentangling and quantifying the value of expert data, our framework builds towards a practical, information-theoretic understanding of how agents should learn from others.
PaperID: 997, Poster
Authors:
Duc Khiem Pham, Quang Nguyen, Tung D Nguyen, Jingsen Zhu, Michele Santacatterina, Dimitris Metaxas, Ramin ZabihAbstract: Direct Preference Optimization (DPO) enables offline alignment of diffusion models from preference pairs without explicit reward modeling, but scaling DPO requires large amounts of expensive human preference data. In this paper, we consider a semi-supervised approach where only a small subset of preference pairs is annotated by humans, while the remaining (much larger) pool is unlabeled and annotated by inexpensive synthetic feedback e.g., vision-language models scoring image pairs or self-training from the current model. However, such synthetic supervision is not a replacement for human judgments: it can be systematically misaligned, so naively mixing human and synthetic preferences will yield misaligned inference. We introduce DeDPO, a debiased objective that integrates a doubly robust estimator from causal inference into DPO to correct synthetic-label misalignment while preserving the simplicity of DPO. Experiments show improved label efficiency and robustness to synthetic labeler quality, allowing DeDPO to outperform standard DPO under the same human-label budget and to approach fully human-supervised performance, despite using four times fewer human-labeled data.
Abstract: Sampling from constrained distributions has a wide range of applications, including in Bayesian optimization and robotics. Prior work establishes convergence and feasibility guarantees for constrained sampling, but assumes that the feasible set is connected. However, in practice, the feasible set often decomposes into multiple disconnected components, which makes efficient sampling under constraints challenging. In this paper, we propose MAnifold Sampling via Entropy Maximization (MASEM) for sampling on a manifold with an unknown number of disconnected components, implicitly defined by smooth equality and inequality constraints. The presented method uses a resampling scheme to maximize the entropy of the empirical distribution based on k-nearest neighbor density estimation. We show that, in the mean field, MASEM decreases the KL-divergence between the empirical distribution and the maximum-entropy target exponentially in the number of resampling steps. We instantiate MASEM with multiple local samplers and demonstrate its versatility and efficiency on synthetic and robotics-based benchmarks. MASEM enables fast and scalable mixing across a range of constrained sampling problems, improving over alternatives by an order of magnitude in Sinkhorn distance with competitive runtime.
Abstract: A key strategy for balancing performance and cost in modern machine learning systems is to dynamically route queries to either a low-cost model or a more expensive oracle (such as a large pretrained model or human expert), an approach known as model routing. In this work we present a new uncertainty-aware router that (1) avoids unnecessary oracle calls on inherently ambiguous queries, and (2) adapts dynamically to different loss functions and cost parameters through simple hyperparameter changes, without retraining. Our method, applicable to any classification setting where multiple independent annotations per input are available, is based on decomposing total uncertainty into irreducible and reducible components using higher-order predictors [Ahdritz et al., 2025]. This enables a unified approach to both routing and abstention: predict with the weak model when uncertainty is low, route to the oracle when reducible uncertainty is high, and abstain when irreducible uncertainty is high. Our router comes with strong theoretical guarantees bounding regret relative to optimal task-specific routers. We conduct experiments on both synthetic and real-world datasets that demonstrate the benefits of our approach in suitable regimes---in particular, whenever reducible and irreducible uncertainty are not too correlated.
PaperID: 1000, Poster
Authors: Senanayak Sesh Kumar Karri, Viacheslav (Slava) Borovitskiy
Abstract: Matern Gaussian processes offer a principled framework for learning and uncertainty quantification on graphs. However, they struggle with heterogeneous networks. A central challenge is that variance at a node can represent genuine signal that should strongly propagate (diffuse) to neighbors, mere noise that should be isolated, or a mixture of both. While this problem has been studied in continuous spatial domains, directly translating those approaches to graphs leaves a free parameter per node, causing extreme over-parameterization. To resolve this, we introduce non-stationary Variance-Coupled Gaussian Processes (VCGPs). These models use an auxiliary signal (often easily derived from historical data) to determine relative diffusivity between nodes, tied to the regression data via a single learned hyperparameter, \gamma. The sign of \gamma captures whether the auxiliary signal is directly or inversely proportional to node diffusivity, while its magnitude calibrates this relationship.We formally prove that \gamma = 1 is the unique point where VCGPs are stationary up to rescaling, ensuring that varying \gamma results in genuinely non-stationary correlations. Across five real-world datasets, our VCGPs match or outperform both stationary and heavily parameterized non-stationary baselines.
Abstract: Fine-tuning LLMs on narrow harmful datasets can induce Emergent Misalignment (EM), where models exhibit misaligned behavior far beyond the fine-tuning distribution. We argue that emergent misalignment can be better understood as a data-mediated transfer phenomenon: harmful fine-tuning examples do not induce uniform behavioral spillover, but interact with the structural properties of the dataset and the difficulty of the tasks relative to the model. Across our experiments, we find that misalignment appears more readily when fine-tuning and evaluation prompts share similar underlying functional structure, when prompts leave more room for coherent harmful completions, and when the target behavior has been more reliably learned by the model. The training pipeline itself also matters: pretraining composition shapes later misalignment. We further study Subliminal Learning (SL), where misalignment is transmitted by fine-tuning on seemingly benign data generated by a harmful teacher. Moving beyond the standard SFT setting, we for the first time compare this transfer under off-policy and on-policy distillation as well, allowing us to separate the roles of the teacher guidance and the training data distribution in transmitting misalignment. Together, these results argue for a data-centric view: Emergent/subliminal misalignment should not be treated as a simple consequence of isolated harmful fine-tuning examples, but as the result of interactions between fine-tuning data structure, pretraining distributions, and training channels.
PaperID: 1002, Poster
Abstract: The wall-clock cost of training large neural networks is dominated by the long horizon of sequential optimizer iterations. We recast the gradient-descent (GD) trajectory of modern optimizer (e.g., SGD and Adam) as the unique root of a nonlinear operator, enabling a parallel root-finding approach that updates multiple GD steps simultaneously. We then propose Efficient Inter-step Parallel Optimization (EIPO) for accelerating a wide range of optimizers. EIPO contributes three pieces: (i) a Preconditioned Nonlinear-Equations (PNE) formulation that recovers the exact sequential GD trajectory; (ii) Memory-Efficient Adaptive Anderson Acceleration, which extracts the multisecant geometry directly from the optimizer's intrinsic momentum buffers and therefore avoids the memory overhead of classical Anderson Acceleration. (iii) We prove that EIPO converges to the optimizer's GD trajectory within fewer iterations than its sequential counterpart. (iv) Across language modeling on WikiText-2 (GPT-2, Llama-3.2-1B, Qwen-1.5-4B/2.5-3B, Gemma-2B, GPT-J-6B), image classification on CIFAR-10 (ResNet50, ViT, CNN), and diffusion training on LSUN Church (UNet/DDPM/DDIM), EIPO reduces optimizer iterations by up to 21× and wall-clock time by up to 4.6× while matching baseline perplexity, accuracy, and FID, all under comparable GPU memory and token consumption. Source code is available at: \urlhttps://anonymous.4open.science/r/EIPO-44BC.
PaperID: 1003, Poster
Abstract: Predicting the future motion of multiple people is inherently stochastic: the same observed history admits many plausible continuations depending on intent, coordination, and social context. Yet the dominant paradigm in multi-person 3D motion prediction remains deterministic, with existing methods producing a fixed number of outputs per agent. In contrast, we present a conditional flow matching framework that jointly generates diverse future motions for multiple agents, modeling global trajectories and root-relative poses. A B-spline reparameterization compresses the generative state space while enforcing temporal smoothness, and a dual-stream transformer conditions each agent's predictions on observed neighbor motion for interaction-aware forecasting. Our method achieves state-of-the-art performance on CMU-Mocap (UMPM) and 3DPW, outperforming deterministic baselines on joint position error, pose error, and final displacement error despite producing a full predictive distribution. We additionally establish the first motion prediction baseline on WorldPose, a large-scale professional soccer dataset, where social conditioning demonstrably concentrates predictive uncertainty along interaction-constrained directions.
PaperID: 1004, Poster
Abstract: Fractional differential equations (FDEs) excel at modelling systems with long-range memory, while stochastic differential equations (SDEs) capture inherent randomness. Existing neural differential equation models address either memory (Neural FDEs) or stochasticity (Neural SDEs), but a unified framework combining both with a learnable scalar memory exponent and an efficient adjoint remains absent. We introduce the Neural Fractional Stochastic Differential Equation (Neural FSDE), which learns continuous-time dynamics with both memory and randomness from data via a Caputo FSDE with neural drift, neural diffusion, and a learnable fractional order. To enable scalable training, we derive a discrete adjoint method that computes gradients with respect to the initial state, the drift and diffusion network parameters, and the fractional order, with memory cost independent of the autograd graph depth. We validate the framework on two real-world domains. On the stochastic-fractional COVID-19 model of Bonyah et al., the NeuralFSDE provides the first end-to-end real-data calibration under a rolling-origin forecasting protocol, recovering a fractional order less than 1. On S&P~500 index option pricing, we recast rough Bergomi, rough Heston, Neural SDE, and NeuralFSDE as Caputo-FSDE special cases under a single calibration pipeline; the NeuralFSDE achieves the lowest implied-volatility error across schedules and instruments, with a matched-architecture comparison against the memoryless Neural SDE, isolating the contribution of fractional memory. Code is available at https://anonymous.4open.science/r/NeuralFSDE.
Authors: Zengxiang Lei, Ananth Shreekumar, Jonathan Rosenthal, Ruoyu Song, Alvaro A Cardenas, Daniel Fremont, Dongyan Xu, Satish Ukkusuri, Z. Berkay Celik
Abstract: Generative Flow Networks (GFlowNets) learn to sample states proportional to an unnormalized reward. Despite their theoretical promise, practical training is often unstable, exhibiting severe loss spikes and mode collapse. To address this, we first assess the sensitivity of GFlowNet objectives, demonstrating that a small Total Variation (TV) distance between the learned and target distributions does not preclude an unbounded training loss. Motivated by this mismatch, we establish converse guarantees by deriving loss-to-TV bounds that certify global fidelity from bounded trajectory balance losses. Lastly, we propose \emphStable GFlowNets, which leverages our theory to stabilize training via adaptive reference flow and improves the trade-off among mode coverage, robustness, and certifiability.
Abstract: Proximal Policy Optimization (PPO) is widely used in reinforcement learning due to its strong empirical performance, yet it lacks formal guarantees for policy improvement and convergence. PPO's clipped surrogate objective is motivated by a lower bound on linearization of the value function in flat geometry setting. We derive a tighter surrogate objective and introduce Fisher-Rao PPO (FR-PPO) by leveraging the Fisher-Rao (FR) geometry. Our scheme provides strong theoretical guarantees, including monotonic policy improvement. In the direct parametrization setting, we show that FR-PPO achieves sub-linear convergence, and for parametrized policies we further obtain sub-linear convergence up to the compatible function approximation error. Finally, although our primary focus is theoretical, we also demonstrate empirically that FR-PPO performs well across a range of standard reinforcement learning tasks.
Abstract: Convolutional architectures have emerged as powerful alternatives to Transformers for sequence modeling. The primary advantage is that they offer improved theoretical sequence length complexity by leveraging the Fast Fourier Transform (FFT). However, this theoretical improvement does not always meaningfully land in practice. One critical obstacle is that applying standard FFTs is not amenable to the large-scale training pipeline wherein data is packed from different sources into a single sequence for hardware efficiency. Indeed, standard FFT algorithms are not easily amenable to document packing. Existing workarounds suffer from severe inefficiencies, crippling the practical performance of convolutional architectures. We close this gap with RubiConv, a novel algorithm for performing hardware-efficient, boundary-respecting convolutions on packed sequences. Extensive experiments show that RubiConv achieves significant speedups over both attention and standard FFT-based baselines. This work makes the theoretical efficiency of long convolutional models a practical reality for large-scale, real-world data packing.
PaperID: 1008, Poster
Abstract: Analyzing complex data across a wide range of scientific fields often requires identifying central or typical elements, distinguishing them from atypical ones, and constructing central or "normal" regions at prespecified levels of centrality. We introduce centrality scores (C-scores) based on distance profiles for random objects taking values in general metric spaces. The proposed C-scores are defined through weighted optimal transport of distance profiles and provide a unified framework for constructing central or "normal" regions for complex data. We show that, for Euclidean data, the limiting behavior of the proposed C-scores is asymptotically equivalent to the underlying probability density function. The practical merits of the proposed method are illustrated using gene expression data, distributions of recorded temperatures, handwritten digit recognition, and taxi trip records. These examples demonstrate the effectiveness of C-scores in identifying central and peripheral elements and in yielding meaningful insights for non-Euclidean and high-dimensional data.
PaperID: 1009, Poster
Abstract: Dynamic graph learning models the temporal evolution of structural interactions. Existing methods present a stark trade-off between performance and efficiency. Event-based models achieve strong accuracy because processing graphs edge-by-edge allows them to continuously update and leverage fine-grained joint neighborhood information. Conversely, snapshot-based methods offer high computational efficiency by processing entire snapshots simultaneously. However, updating neighborhood states for all nodes concurrently is computationally prohibitive. Consequently, current snapshot methods fail to utilize joint neighborhood information, leading to significantly weaker performance. To bridge this gap, we propose the Neighbor-Aware Graph Neural Network (NAGNN), a novel snapshot-based architecture that seamlessly integrates joint neighborhood information to achieve both strong performance and high efficiency. Specifically, NAGNN introduces a Locality-Sensitive Hashing (LSH)-based topology distiller to construct compact reservoirs that store neighborhood information efficiently. These reservoirs are further augmented with degree-scaled statistical moments to encode structural characteristics of local neighborhoods. A Gated Recurrent Unit (GRU) subsequently processes the distilled structural representations to model and capture their temporal dynamics. Extensive experiments demonstrate that NAGNN matches the predictive accuracy of computationally expensive event-based models with minimal overhead, establishing a highly scalable and effective paradigm for dynamic graph learning.
Abstract: The Gromov--Wasserstein (GW) problem provides a framework for aligning heterogeneous datasets by matching their intrinsic geometry, but its statistical and computational scaling remains an issue for high-dimensional problems. Slicing techniques offer an appealing route to scalability, but, unlike Wasserstein distances, GW problems do not generally admit closed-form solutions in one-dimension. We resolve this problem for the GW problem with inner product cost (IGW), propose a sliced IGW distance that enjoys a natural rotational invariance property, and comprehensively study its structural and computational properties. Numerical experiments validating our theory are presented, followed by applications to heterogeneous clustering of text data and language model representation comparison.
Authors: Parul Maheshwari, Amulya Paruchuri, Alireza S Shirazi
Abstract: Graph Neural Networks trained on heterogenous bipartite graphs form a common basis in recommendation systems. These graphs often express relations that vary in cardinality, for example, user-item preferences are one-to-many and user-attribute features are one-to-one. Traditionally, a unique loss function is applied for all of the network components which is often Bayesian Personalized Ranking (BPR). While BPR works well for the recommendation task, we find that it causes attribute embeddings to collapse to near-random geometry — a silent failure that leaves standard ranking metrics largely unaffected and therefore invisible to conventional evaluation. This in turn pollutes user node embeddings, which are shaped by both edge types simultaneously, hurting downstream tasks like personalization, segmentation, etc. Here we propose a Cardinality-Decomposed Loss (CDL) that combines both Cross Entropy (CE) and BPR to enable the model to collectively optimize for relations across cardinalities. As we implement this loss, we also confirm the conflict between CE and BPR by showing that the two losses compete against each other in the shared encoder's parameter space. We evaluate CDL on five datasets spanning two structural configurations — one-to-one attributes on user nodes (MovieLens-1M, Last.fm-360K, PayPal Audience Factory, BookCrossing) and on item nodes (Yelp) — and find that CDL consistently improves discriminability in attribute embeddings. We also show that whenever these attributes contain meaningful preference signal, we also see improvement in the ranking task (measured by NDCG). On the other hand, when attributes are weakly correlated with preferences, there is an inherent tension between the two objectives. We use a lambda parameter to navigate this trade-off, and a lambda-sweep reveals that dataset behavior is governed by two graph properties — semantic alignment and topology leakage. Semantic alignment captures whether the one-to-one attribute is predictive of user preferences, while topology leakage captures whether message passing already encodes attribute structure implicitly through the graph's connectivity.
Abstract: The classic online stochastic matching problem typically requires immediate and irrevocable matching decisions. However, in many modern decentralized systems such as real-time ride-hailing and distributed cloud computing, the primary bottleneck is often local communication bandwidth rather than the timing of the match itself. We formalize this challenge by introducing a two-stage local sparsification framework. In this setting, arriving requests must prune their realized compatibility sets to a strict budget of k edges before a central coordinator optimizes the global matching. This creates a "middle ground" between local information constraints and global optimization utility. We propose a local selection strategy, parametrized by a fractional solution of the expected instance. Theoretically, we quantify the approximation ratio as a function of the solution's \em spread. We prove that under sufficient spread, our sparsifier globally preserves the expected size of the maximum matching. Empirically, we demonstrate the robustness of our approach using the New York City ride-hailing datasets and adversarial synthetic benchmarks. Our results show that near-optimal global matching is achievable even with highly constrained local budgets, significantly outperforming standard online baselines.
Abstract: Long-context language modeling is increasingly constrained by the Key–Value (KV) cache, whose memory and decode-time access costs scale linearly with the prefix length. This bottleneck has motivated a range of context-compression methods, from token-level summarization to recent optimization-based KV compression methods. These post-hoc methods operate on the KV cache of a fixed pretrained model, so their effectiveness is fundamentally limited by how well the model's internal representations can be compressed. In this work, we formalize the notion of KV compressibility and show that it is a property of the , rather than of the context alone. We prove that almost any sequence-to-vector function admits both highly compressible and inherently non-compressible transformer implementations, highlighting the need to guide transformers toward compressible representations during training. Motivated by this, we propose raining (KV-CAT), a continued pretraining procedure that incentivizes the emergence of compressible representations. We introduce a train-time KV sparsification policy that masks KV slots during training. This forces the model to use fewer KV slots and encourages it to learn representations amenable to post-hoc compression. Empirically, we show that KV-CAT improves the quality–budget tradeoff of downstream compression methods across retrieval, long-context question answering, and perplexity-based evaluation of compressed-prefix continuation.
PaperID: 1014, Poster
Abstract: Image--text alignment models such as CLIP are typically trained with contrastive learning, where unpaired examples are treated as strict negatives. While this assumption introduces noise by penalising semantically related pairs, existing solutions often rely on heuristic similarity thresholds that lack statistical grounding and sensitivity to dataset-specific noise distributions. In this paper, we propose Permute-then-Adapt (PTA), a framework that leverages efficient weak models to identify missed positives, while enforcing statistical control to reject random associations. Unlike standard distillation or arbitrary thresholding, PTA estimates the null distribution of the weak teacher's similarities via permutation testing, ensuring selected pairs are statistically distinguishable from the distribution of random pairings at level \alpha. The calibrated positive set is then trained against by a single multi-positive contrastive objective: each anchor maximises the total probability mass it assigns to its positive set. We further demonstrate that the calibration can be pre-computed offline using the weak model, so training overhead is negligible compared to standard baselines. Extensive experiments show that PTA consistently outperforms heuristic soft-label approaches on object recognition and cross-modal retrieval, while exhibiting superior data efficiency.
Abstract: Although Large Language Models (LLMs) achieve strong alignment through supervised fine-tuning and reinforcement learning from human feedback, the alignment is often fragile under subsequent fine-tuning. Existing explanations either attribute alignment fragility to gradient geometry or characterize it as a distributional shift in model outputs, yet few provide a unified account that bridges parameter-space learning dynamics with function-space alignment behavior during fine-tuning. In this work, we introduce a tractable alignment score and derive its closed-form update during fine-tuning, yielding a unified framework for alignment dynamics. Our analysis decomposes alignment updates into two competing components: a Rebound Force, governed jointly by the current alignment state and the narrowness of model distribution, and a Driving Force, determined by how the training distribution aligns with outcome-conditioned posteriors over aligned and non-aligned completions. This decomposition explains why prior alignment can be reversed by later fine-tuning and why narrower posterior structure strengthens such reversal. Moreover, our framework predicts a Rehearsal Priming Effect: prior alignment leaves a latent posterior imprint that amplifies the effective Driving Force upon re-exposure, leading to faster re-alignment. We validate these predictions across safety alignment, emergent misalignment, and sentiment settings, demonstrating consistent alignment reversal and accelerated re-alignment under re-exposure. In addition, controlled experiments in safety alignment confirm the predicted dependence of rebound strength on posterior narrowness. Together, these results provide a unified dynamical perspective on how alignment is disrupted and reactivated during LLM fine-tuning.
Abstract: Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or attention proxies, but they do not test whether a sampled answer causally depends on the local evidence that supports it. We propose Counterfactual Evidence Disentanglement (CED), a training-time evidence audit for VLM grounding. For each response, CED neutralizes an object-centric Evidence Region and compares the resulting support drop against matched non-evidence Regions. We combine this signal with answer correctness inside GRPO, rewarding correct answers that rely on the evidence path rather than shortcut or nuisance paths. CED uses weak object-level proposals, requires no question-specific evidence annotations, and adds no inference-time overhead. Across nine public benchmarks and four backbones, CED outperforms prior RL-based post-training methods, with targeted analyses verifying its object-centric signal.
Abstract: In regression, conformal prediction often suffers from inefficiency under heteroscedasticity and skewness due to fixed, non-adaptive interval centering. We propose CoCP, a new method that parameterizes prediction intervals by a center m(x) and a radius h(x), and alternates between learning h(x) via smooth quantile regression on folded residuals and refining m(x) using a smooth interval loss. This corrects mis-centering and drives the interval toward high-density regions. Beyond finite-sample marginal validity via split-conformal calibration, we prove that CoCP asymptotically achieves conditional coverage and optimal interval length if the base estimators are consistent. Experiments demonstrate that CoCP yields tighter intervals and achieves state-of-the-art conditional coverage reliability.
PaperID: 1018, Poster
Abstract: Large language models (LLMs) can memorize and reproduce training sequences verbatim, which undermines both generalization and privacy. Existing mitigation methods apply interventions uniformly, often degrading performance on the majority of tokens that generalize normally. We empirically find that memorization is sparse and heavy-tailed at the token level. This structure is mismatched to static, sequence-level interventions: effective mitigation must localize to where memorization actually happens. From the ideal token-level memorization reduction objective, we derive a probe-steer framework, which decomposes intervention into a probe that detects memorization-relevant activations and a steer that applies targeted correction only when the probe exceeds a threshold. We propose Gated Subspace Steering (GSS) as its practical instantiation: the optimal probe-steer pair admits a closed-form solution as the leading singular vectors of a gradient-weighted activation matrix. Across four benchmarks, GSS matches or exceeds state-of-the-art memorization reduction while requiring 10--100× less compute than optimization-based alternatives.
PaperID: 1019, Poster
Abstract: We study private online allocation problems, where allocation decisions must satisfy differential privacy (DP) to protect sensitive user information while optimizing performance under constrained resources. We first formalize differential privacy notions suited to different privacy requirements in online allocation and establish the fundamental performance limits of any algorithm under these definitions. Based on these insights, we propose PPOA, a privacy-preserving meta algorithm for online allocation, and establish sufficient conditions under which PPOA preserves Joint DP (JDP) or Local DP (LDP) while achieving asymptotically near-optimal performance. Building on these conditions, PPOA can be instantiated into a variety of concrete algorithms. In particular, we present the specific designs, which include PPOA-DMD that preserves JDP and LDP and PPOA-FTRL that preserves JDP. We analyze their performance under both adversarial and stochastic settings and characterize the fundamental trade-offs between DP and allocation performance. Finally, we demonstrate the superior performance of the proposed algorithms via numerical experiments on AI model routing in battery-powered edge systems.
PaperID: 1020, Poster
Abstract: Precise localization of recording electrodes is fundamental to systems neuroscience; yet, current methods depend on labor-intensive post-hoc histology, which is incompatible with real-time feedback or chronic implants. We test the hypothesis that the electrical signals being recorded themselves carry sufficient anatomical information to localize their recording site in 3D atlas coordinates, without the need for histology. To probe this, we introduce a self-supervised framework that learns geometry-aware representations of single-channel local field potentials (LFP), leveraging probe channel geometry as a self-supervisory signal. We compare speech-based SSL objectives (e.g., masked prediction) against geometry-aware objectives that constrain the embedding space to mirror physical channel layout along the probe. These representations outperform masked-prediction baselines, and the gains compound at larger data scales. We test downstream analysis with 3D coordinate regression and brain region classification, benchmarking three SSL backbones (Wav2Vec 2.0, Whisper, and Data2Vec) against supervised baselines (AnyNet, ViT, and classical spectral classifiers) across datasets, laboratories, species, and probe technologies. Our method achieves state-of-the-art results in brain region classification and 3D coordinate regression from raw single-channel LFP, outperforming all baselines, affirming our hypothesis.
PaperID: 1021, Poster
Abstract: Most world models for future prediction prioritize aestetic quality of generated videos, but offer only weak physical grounding and limited geometric consistency over time, due to the lack of explicit 4D modeling. In this paper, we introduce plan4d, a possible next milestone towards building a generative plannable 4D world model that predicts future RGB videos together with dense 4D geometry, including depth, cameras, point clouds, and scene flow from an observed scene context and text-specified actions. Since training data with paired actions and dense 4D annotations is inherently scarce, plan4d leverages the complementary priors of foundational video generators and 3D reconstructors. Specifically, we propose a model-stitching architecture that couples Depth Anything 3 with LongCat-Video and augments it with an auxiliary DPT head for scene-flow prediction. The resulting model allows us to plan and generate complex, long-horizon, and fine-grained actions in a dynamic 3D scenes, while substantially improving both visual fidelity and geometric accuracy of the generated future videos. Extensive experiments on both synthetic and real-world datasets demonstrate that plan4d achieves state-of-the-art performance in generative 4D modeling of environmental reactions.
Abstract: Video world models have achieved strong visual realism, but this does not ensure that their dynamics are truly governed by actions. In this work, we argue that action faithfulness should be understood through the compositional structure of actions, which in many embodied settings follows a group structure (e.g., SE(2) for navigation). Based on this insight, we formalize action-conditioned world modeling as realizing a \it group action on the state space, providing a principled criterion for evaluating dynamics beyond visual quality. To operationalize this framework, we propose a unified approach that enforces identity, inverse, and composition consistency via latent-space regularization with synthesized supervision, avoiding additional data collection. We further introduce two metrics: Group-Action Consistency (GAC) and Group-Action Robustness (GAR), to evaluate structural correctness and rollout stability. Extensive experimental results show that our method consistently improves both GAC and GAR in state-of-the-art video world models without degrading perceptual quality.
PaperID: 1023, Poster
Abstract: Many decision making procedures that make use of probabilistic predictors assume that training on new data acts as Bayesian conditioning. When this assumption breaks, the downstream procedure can perform poorly. E.g. we show that in active learning, incoherent updates cause the procedure to prefer to acquire suboptimal points. To quantify how far a model's predictive updates depart from conditioning, we introduce the . We empirically find that amortised predictors such as TabPFN can be more coherent than standard parametric approximations to Bayesian inference. In synthetic active learning experiments, more coherent models acquire better data, an effect not attributable to better predictive performance. In many tasks, decisions depend on a predictor only through its induced action. We therefore introduce action coherence, a decision-theoretic relaxation that measures only the incoherence affecting that action. These diagnostics make coherence testable and quantifiable, enabling incoherence to be diagnosed and addressed.
Abstract: Self-improvement is a critical capability for large language models and other intelligent systems, enabling them to refine their behavior and internal consistency without external supervision. Despite its importance, prior approaches largely rely on empirical heuristics and lack formal guarantees. In this paper, we propose a principled framework for self-improvement based on the concept of coherence, which requires that a model's outputs remain consistent under task-preserving transformations of the input. We formalize this concept using projection-based mechanisms that update a baseline model to be coherent while remaining as close as possible to its original behavior. We provide rigorous theoretical guarantees that these mechanisms achieve monotonic improvement, measured by a reduction in expected Bregman divergence. Our analysis is comprehensive, covering both \emphdirect and \emphtwo-step projection methods, and robustly extends these guarantees to non-realizable settings, empirical (finite-sample) distributions, and relaxed coherence constraints.
PaperID: 1025, Poster
Authors:
Nick Huang, Jackson Woodleigh, Aaron Gokaslan, Xinjie Yi, James TompkinAbstract: GANs are appealing because they generate sharp images in a single forward pass. Recent works have stabilized GAN training, but FID may still plateau on diverse datasets like ImageNet-256 because the discriminator's activation magnitudes increase through training. We trace this pathology to the discriminator's incentive to sharpen its decision boundary by rescaling activations rather than by finding better features. Then, building on R3GAN, we treat magnitude preservation as a first-class architectural target across the network: \ell_2 weight normalization with centering, magnitude-preserving LReLU and residuals, forced post-update normalization, and feed-forward classifier heads. Across a roadmap of configurations, we perform diagnostic ablations to clarify what does and does not work and why. The resulting models reach FIDs of 2.20 @ 17 M parameters, 1.63 @ 60 M, and 1.52 @ 130 M on class-conditional ImageNet-256. At 1 NFE, this is substantially more efficient than competing diffusion and autoregressive methods, and than prior GAN baselines too (e.g., 10× more efficient than StyleGAN-XL for similar FID). In sum, our work presents a new SOTA in FID/parameter efficiency derived from simple adversarial training with well-behaved magnitude preserving layers.
PaperID: 1026, Poster
Abstract: cale (ATLAS) in Lean 4. AutoformBot orchestrates thousands of LLM agents, equipped with formal verification tools, dependency-aware task scheduling, and collaborative version control, to translate informal textbook prose into machine-checked definitions and proofs. We apply our methods to a corpus of over open-access textbooks spanning analysis, algebra, topology, combinatorics, and probability, producing ATLAS: a verified library of over lines of code. We release two artifacts: (i) AutoformBot, the open-source multi-agent framework; and (ii) ATLAS, the resulting formal library. Our results suggest that autoformalizing the core content of graduate-level mathematics at scale is now economically and technically feasible. This opens the door to the automated verification of both human- and machine-generated mathematics at a research level.
Authors: Chika Maduabuchi
Abstract: Modern text-to-image models produce high-fidelity images but still struggle with compositional prompts that require instance identity, attribute ownership, counting, spatial ordering, and role-sensitive relations. We introduce (PSP-DiT), a diffusion-transformer architecture that treats a panoptic scene program as a first-class latent variable rather than an external control signal or post-hoc parse. PSP-DiT jointly denoises image latents and scene-program latents through coupled transformer streams, while panoptic grounding and cycle-consistency objectives tie object instances, attributes, relations, and counts to visual support in the generated image. Under matched training and inference settings, PSP-DiT improves over a strong flat-text baseline across GenEval 2, SANEval-Simple, PSG-Score, and DetailMaster, with the largest gains on counting, attribute binding, role-sensitive relations, and long structured prompts. The method preserves image quality, adds modest inference overhead, and remains robust to imperfect scene programs.
PaperID: 1028, Poster
Authors:
Harshvardhan Takawale, Aritrik Ghosh, Nirupam RoyAbstract: In neural inverse reconstruction, the forward model is typically treated as a fixed simulator that maps a neural scene representation to measurements. This work studies how the parameterization of the forward operator itself shapes the optimization landscape of coherent inverse problems. We introduce Moray, a reconstruction framework built on spectrally parameterized forward operators that bypass conventional time-domain synthesis and instead evaluate measurements directly through closed-form spectral kernels. This parameterization induces a substantially shorter and better-conditioned differentiable graph compared to the time-domain approaches. We further introduce Phase-Coherent Manifold Parameterization, a scene representation that jointly learns scene reflectivity and a deformable surface manifold, allowing the reconstruction to adapt to geometric deviations from the assumed imaging plane and thereby reducing common artifacts in coherent reconstruction. We instantiate these ideas for 77 GHz radar imaging using real measurements from a synthetic-aperture mmWave system. Across multiple challenging scenes, Moray consistently improves reconstruction fidelity, suppresses background artifacts, and degrades more gracefully under data limitations.More broadly, the results suggest that, in neural inverse problems, exposing appropriate analytical structure of the sensing physics to the optimizer can be as important as the choice of neural representation itself.
PaperID: 1029, Poster
Authors:
Mincheol Park, Sunwoo Lee, Sukjin Lee, Wooram Yang, Yongmin Tai, Sang Joon Kim, Jaehoon YuAbstract: This paper rethinks Large Language Model (LLM) inference from a Mortal Computing perspective, in which computation is embodied in hardware and functionality is inseparable from its physical substrate. Rather than treating inference optimization as merely a matter of arithmetic simplification or precision reduction, we explore hardwiring as a way to eliminate weight access at the root of data movement. We show that viable hardwiring begins not by retaining weights in their multiplicative form, but by factorizing them into combinations of shiftable bases and rewiring computation at the base level, where redundancy is exposed. This view, however, introduces a fundamental challenge in the dramatic growth of wires required to reconstruct weights. To address this, we propose Moira, a budgeted greedy search that formulates wire-count control as a Lagrangian relaxation problem and selects only the bases that satisfy the objective under a wiring budget. What sets Moira apart is its adaptive base selection, which gives it a dual characteristic of pruning and quantization while preserving model expressivity under sparse wiring without retraining. Extensive evaluations demonstrate that Moira achieves near-lossless performance with, on average, only about two bases per weight and outperforms existing compression methods. Moira opens a concrete path beyond the von Neumann paradigm, marking a first step toward Digital Mortal Computing for LLMs.
Abstract: Uncertainty quantification is crucial in scientific machine learning, where models inform safety-critical tasks such as flood forecasting, financial risk management, and thermal control in machines. Conformal prediction provides distribution-free coverage guarantees, but in time-dependent settings common to physics and engineering, these guarantees can break down, leading to systematic undercoverage. We study this problem in the context of surrogate models for time-dependent physical systems described by partial differential equations. We prove that in a function space setting, distributions at arbitrarily close times can be mutually singular, making exact coverage guarantees impossible. We further show that in discretized settings, the total variation distance between distributions often grows exponentially with the grid resolution mandating restrictions on the spatial discretization to ensure coverage. Finally, for certain settings, we show how to use principled weighted conformal prediction to obtain finite-sample coverage guarantees over growing time horizons.
Abstract: Nearest-neighbor methods are fundamental to classical and modern machine learning, yet their geometric properties are typically analyzed under independent sampling. In this paper, we study the nearest-neighbor radii under dependent sampling. We consider strong mixing dependent observations and ask whether dependence changes the scale of nearest-neighbor neighborhoods. We establish distribution-free almost sure convergence under polynomial mixing and sharp non-asymptotic moment bounds under geometric mixing. The moment bounds depend on the local intrinsic dimension rather than the ambient dimension, making the results applicable to high-dimensional data concentrated near lower-dimensional manifolds. Synthetic experiments and real-world time-series benchmarks support the theory, showing that nearest-neighbor geometry remains informative under dependence sampling.
PaperID: 1032, Poster
Abstract: Parallel thinking scales inference-time compute by generating multiple reasoning traces for a query, and aggregating them to reach a final answer. However, most aggregation algorithms treat each reasoning trajectory independently, ignoring the complementary reasoning sub-processes across the full ensemble of trajectories. To learn from these parallel trajectories, we propose CORAL (Consolidating Reasoning with Test-Time Learning), a novel framework for aggregating trajectories by jointly consolidating them into parametric memory using test-time training. Our algorithm meta-learns how to consolidate using nested optimization at training time, where an inner loop encodes reasoning trajectories into a lightweight LoRA adapter, and an outer loop learns meta-parameters for the inner loop such that the consolidation is helpful for the model to synthesize a final solution. As no standard datasets exist for training aggregation methods, we construct a novel OpenParallelThinking corpus composed of 12,810 math problems paired with a mixture of diverse trajectories from six LLMs. Evaluated on challenging mathematics and STEM benchmarks, CORAL consistently outperforms best aggregation baselines by 20% overall. Importantly, we curate reasoning trajectories at multiple quality levels to simulate the reality of noisy parallel reasoning and demonstrate that CORAL is especially robust to the quality of thinking trajectories and able to learn useful knowledge from entirely wrong solutions. Additionally, this weak-to-strong generalization feature scales with the number of incorrect trajectories available for aggregation, while RL can further incentivize such reasoning consolidation capability. We will release CORAL and OpenParallelThinking for use in future work on reasoning aggregation.
Abstract: Maintainable and general software allows developers to build robust applications efficiently, yet achieving these qualities often requires refactoring specialized solutions into reusable components. This challenge becomes particularly relevant as code agents are increasingly used to solve isolated one-off programming problems. We investigate code agents' capacity to refactor code in ways that support consolidation and reusability. We first investigate what makes a good refactoring, finding via simulation results and a human study that proxies of program size, such as Minimum Description Length, better predict preferable refactorings than extant software engineering metrics, such as Maintainability Index. We then present \textscLibrarian, a method that combines divide-and-conquer with sample-and-rerank to generate reusable libraries. We compare \textscLibrarian to state-of-the-art library generation methods, and study it on real-world Python programs.
Abstract: We address the brittleness of Bayesian experimental design under model misspecification by formulating the problem as a max--min game between the experimenter and an adversarial nature subject to information-theoretic constraints. We demonstrate that this approach yields a robust objective governed by Sibson's \alpha-mutual information (MI), which identifies the \alpha-tilted posterior as the robust belief update and establishes the R\'enyi divergence as the appropriate measure of conditional information gain. To mitigate the bias and variance of nested Monte Carlo estimators needed to estimate Sibson's \alpha-MI, we adopt a PAC-Bayes framework to search over stochastic design policies, yielding rigorous high-probability lower bounds on the robust expected information gain that explicitly control finite-sample error.
PaperID: 1035, Poster
Abstract: Discovering explanatory subgraphs is essential for interpreting Graph Neural Networks (GNNs). Many existing and popular explainers optimize by maximizing mutual information (MMI) between selected subgraphs and predicted labels; however, the resulting optimization landscape is often non-smooth and poorly conditioned, leading to suboptimal local optima. To address this issue, we develop a norm-theoretic framework for subgraph analysis. We show that subgraphs containing model-utilized features tend to induce graph embeddings with larger \ell_2 norms, as they align more strongly with the supportive subspace shaped by the GNN’s weight matrices, a phenomenon consistently observed across multiple datasets. Based on this insight, we propose NORMA, a parameterized explainer that uses subgraph-embedding \ell_2 norms as a stable optimization signal to guide subgraph selection. Experiments on fsix benchmark datasets demonstrate that NORMA effectively alleviates the optimization difficulties of MMI-based methods and achieves superior explanation quality.
Abstract: Recent advances in multi-agent systems have shown great potential for solving complex tasks. However, when multiple agents edit a shared codebase concurrently, their changes can silently conflict and inconsistent views lead to integration failures. Existing multi-agent systems address this through workspace isolation (e.g., one git worktree per agent), but this defers conflict resolution to a post-hoc merge step where recovery is expensive. In this paper, we propose STORM, i.e., STate-ORiented Management for multi-agent collaboration. Specifically, STORM manages agent states by mediating their interactions with the shared workspace, ensuring that each agent operates on a consistent view of the codebase and that conflicting edits are detected and resolved at write time.We evaluate STORM on Commit0 and PaperBench across multiple LLMs. STORM outperforms the git-worktree-based multi-agent baseline by +18.7 on Commit0-Lite and +1.4 on PaperBench, while achieving comparable or better cost efficiency. Combined with single-agent runs, STORM reaches highest scores of 87.6 and 78.2 on the two benchmarks respectively, suggesting that explicit state management is a more effective foundation for multi-agent collaboration than workspace isolation. STORM can also be plugged into any multi-agent system. Our code is available at \urlhttps://anonymous.4open.science/r/STORM-Mutli-Agent.
Authors:
Jaehyuk Lee, Hanyoung Kim, Yanggee Kim, Donghun LeeAbstract: Vision Transformers (ViTs) face severe computational bottlenecks due to the quadratic complexity of self-attention at high resolutions. Existing token reduction methods rely on local metrics—such as single-layer attention scores—that are inherently vulnerable to the \emphattention sink phenomenon, where uninformative tokens are paradoxically preserved over salient foreground objects. We propose ASAP (Attention Sink Anchored Pruning), a training-free framework that recasts this sink as a feature. Modeling ViT information flow as a Lazy Random Walk, ASAP identifies the sink as a dominant accumulator of probability mass. By computing the \emphdiffusion distance to the sink within the cumulative transition matrix, ASAP partitions tokens via \emphRadial Diffusion Clustering and compresses background redundancy through \emphTransition Weight Pooling in a single shot. Extensive experiments across image, video, and vision-language tasks demonstrate ASAP outperforms state-of-the-art methods, accelerating throughput by up to 48% while maintaining—or even exceeding—baseline accuracy.
Abstract: Large language models (LLMs) exhibit severe multilingual safety misalignment: they possess strong safeguards in high-resource languages but remain highly vulnerable to jailbreak attacks in low-resource languages. Current safety alignment methods generally rely on high-quality response data for each target language, which is expensive and difficult to generate. In this paper, we propose a cross-lingual safeguard transfer framework named Multilingual Self-Distillation ( , Javanese) languages, overcoming the need for response data in any language. Our framework is flexible and can be integrated with different self-distillation strategies. Specifically, we implement two concrete methods---on-policy MSD and off-policy MSD---both of which enable effective cross-lingual safety transfer using only multilingual queries. Furthermore, we propose Dual-Perspective Safety Weighting (DPSW), a divergence measure to optimize the distillation objective. By jointly considering the perspectives of both the teacher and the student, DPSW adaptively increases the penalty weights on safety-critical tokens while reducing the weights on non-critical tokens. Extensive experiments on representative LLMs across diverse multilingual jailbreak and utility benchmarks demonstrate that our method consistently achieves superior multilingual safety performance. Notably, it generalizes effectively to more challenging datasets and unseen languages while preserving the model's general capabilities.
PaperID: 1039, Poster
Abstract: A multimodal large language model (MLLM) forecaster takes a forward-looking scenario as a soft constraint that decays with horizon. The decay is expected, but its rate is uncharacterized and untunable at inference time. A practitioner running a twelve-month stress test cannot tell whether the scenario binds for twelve months or twelve days. Existing accounts of autoregressive error, such as exposure bias and hallucination snowballing, bound a single rollout against ground truth and offer no inference-time lever. Under three assumptions, we prove a closed-form Lipschitz upper envelope on the per-step scenario drift between two paired-seed forecasts differing only in scenario text. Two identifiable parameters partition behavior into bounded, linear, and exponential regimes. A sister within-chunk bound holds under any chunk partition. It motivates a stabilization heuristic that we test and falsify: periodic re-injection of the scenario enlarges drift relative to a single injection, and length-matched random text shrinks it. The envelope is respected; the heuristic is not. A pre-registered attention diagnostic across the open-weight panel returns architecturally heterogeneous verdicts that the recurrence framing covers uniformly while no single attention-dilution mechanism does. In place of text re-injection, we propose choring), a chart-anchored multimodal prompt that renders the lookback, the scenario, and the historical context as a single image. Across a panel of MLLMs and two long-horizon scenario-conditioned forecasting benchmarks,
PaperID: 1040, Poster
Authors: Victor Augusto Kich, Satoshi Yamamori, Jun Morimoto
Abstract: Sparse rewards make cooperative multi-agent reinforcement learning difficult because useful feedback often appears only after multiple agents coordinate over long horizons. We introduce Multi-Agent Contrastive Reinforcement Learning (MACRL), a goal-conditioned method that trains parameter-shared decentralized actors with a centralized contrastive critic. Instead of learning from sparse task rewards, the critic contrasts joint state-action embeddings with future achieved goals from replay, turning every trajectory into dense self-supervised signal. Across tasks that vary in horizon length, coordination difficulty, and object interaction, our method is competitive with strong HER-based baselines when exploration is sufficient and substantially more robust when sparse rewards become uninformative. On the hardest tasks, it obtains substantial non-zero success while non-contrastive baselines remain near zero. Scaling from 4 to 64 layers further improves performance on several hard tasks, with gains up to 25 percentage points, while providing little consistent benefit to MA-SAC+HER. These results suggest that contrastive objectives and depth-scaled residual networks provide a promising foundation for scalable cooperative goal-conditioned control.
PaperID: 1041, Poster
Authors: Hyunwoo Lee, Hyojae Lim, Dohyun Kwon
Abstract: Classical initialization theory analyzes networks through second-order observables such as variance, correlation, and Jacobian conditioning that are natural when weights are symmetric around zero. Identity-plus-noise initialization, which initializes the network to behave locally like a residual block, sits outside this picture: variance can be preserved while coordinate-wise signal identity is destroyed, making such second-order quantities indirect indicators of trainability. We identify the terminal sign-flip rate as the natural forward-pass observable for this regime, and derive its critical scale \sigma ~ L^-1/2 from a dimensionless margin--sensitivity ratio, rather than postulating it. The retention curve emerges as a first-passage limit of a signed-margin process, yielding a calibration procedure that targets a chosen sign-flip rate at any depth. Experiments on synthetic and real benchmarks suggest that this calibrated initial rate behaves as a training-time parameter rather than an inert calibration detail.
PaperID: 1042, Poster
Authors: Sophia Xiao, Bijan Mazaheri
Abstract: Existing causal discovery methods for interventional data typically follow one of two paradigms: they either require explicit environment labels and known intervention targets, or, in the absence of such metadata, they rely on computationally intensive procedures such as first inferring intervention targets or employing complex models to account for the latent variables of environment labels. Surprisingly, we show that neither explicit environment labels nor auxiliary inference algorithms are strictly necessary: under certain conditions, the classical ICA-LiNGAM algorithm consistently recovers the causal graph from simply pooled interventional and observational data. We first demonstrate that in linear Gaussian structural equation models where interventions independently perturb the noise distributions across nodes, pooling produces exactly independent, non-Gaussian sources, the appropriate setting for Independent Component Analysis (ICA). Because the assumption of strictly independent interventions can be overly restrictive, we extend our analysis to the practical setting of single-node (atomic) interventions. While pooling atomic interventions induces the necessary non-Gaussianity, it also introduces source dependence, violating a core ICA assumption. However, theoretical analysis establishes that this dependence decays as \mathcalO(1/n^2) with the number of variables n, whereas non-Gaussianity decays only as \mathcalO(1/n). This yields a regime of mild misspecification at moderate n where sources are near-independent yet sufficiently non-Gaussian for ICA to succeed. In synthetic experiments, where data is pooled across many different environments, our label-free approach matches or exceeds label-aware methods. On real-world data, our approach performs comparably to methods that explicitly require environment labels. Ultimately, our work challenges the assumed necessity of environment labels and complex inference mechanisms in interventional causal discovery and suggests a re-evaluation of the often-overlooked classical ICA-LiNGAM algorithm.
Abstract: Reinforcement learning (RL) is effective in enhancing the accuracy of large language models in complex reasoning tasks. Existing RL policy optimization frameworks rely on final-answer correctness as feedback signals and rarely capture the internal logical structure of the reasoning process. Consequently, the models would generate fluent and semantically relevant responses but logically inconsistent, structurally erratic, or redundant. To this end, we propose StaRPO, a stability-augmented reinforcement learning framework that explicitly incorporates reasoning stability into the optimization objective. Our StaRPO decomposes stability into two computable lightweight metrics: the Autocorrelation Function (ACF) to evaluate local step-to-step coherence, and Path Efficiency (PE) to evaluate global goal-directedness of the reasoning trajectory. These stability rewards are combined with task rewards to provide complementary and process-aware feedback. We validate the effectiveness of using ACF and PE rewards by showing their correlation with logic errors on two backbone models. Experiments on four reasoning benchmarks show that StaRPO consistently outperforms compared baselines and can enhance both final-answer accuracy and logical stability.
PaperID: 1044, Poster
Abstract: Real-world human activities unfold as sequences of temporally dependent human-object interactions (HOIs), yet existing methods address either atomic HOIs in isolation or text-driven motion composition without object or scene context. To bridge this gap, we introduce the task of long-term action composition of HOIs, where a human navigates between and interacts with multiple objects within a 3D scene. The primary challenges lie in producing smooth transitions, scene-aware locomotion, and plausible hand-object interaction. To this end, we propose LT-HOI, a diffusion-based model that jointly conditions on past and future contexts—the past providing kinematic continuity across transitions, the future enabling anticipatory navigation toward upcoming interacting objects. We further integrate Diffusion Noise Optimization for collision avoidance and introduce a hand-object displacement loss to improve contact quality. On the ParaHome dataset, we establish the first benchmark for this task and demonstrate that LT-HOI effectively composes long, temporally coherent HOIs.
PaperID: 1045, Poster
Authors: Nopsinth Vithayapalert, Francesca Grisoni
Abstract: Predicting protein–ligand complex structures is a central challenge in drug discovery. While recent co-folding models such as AlphaFold-3 achieve accurate structure prediction, they fail to generalize to underexplored binding interfaces -- systematically misplacing ligands, particularly for allosteric or structurally novel targets. To address this gap, we present ACER (Adaptive Co-folding via pocket Exploration and pose Ranking), a training-free framework that (a) enables co-folding models to systematically explore alternative binding pockets, and (b) leverages the discovered pockets to increase pose accuracy. Our method enables the efficient discovery of non-prevalent pockets without prior expert knowledge. ACER improves pocket discovery and pose accuracy on allosteric targets and structurally novel complexes, successfully modeling binding interfaces that are under-represented or absent from the training set. Our results demonstrate how improved sampling dynamics enhance the generalisability of co-folding models without retraining.
Abstract: We study differentially private (DP) k-means and k-median clustering in the online streaming setting. In this model, points arrive sequentially, and at each time step, we need to output a set of k centers that optimizes the clustering objective for all points seen so far. We give a generic reduction that transforms the (sensitive) input stream into a private stream, which is a semi-coreset of the input stream. This implies that any (non-private) online clustering algorithm, run as a post-processing step, can achieve good utility for the original clustering objective. Our algorithm matches or improves upon the approximation ratio, space usage, and running time of existing algorithms (Dupr\'e la Tour, 2024; Epasto et al., 2026). A key aspect of our reduction is that it inherits desirable properties of the underlying non-private clustering algorithm, such as consistency (Lattanzi and Vassilvitskii, 2017)---a property not satisfied by previous DP algorithms.
PaperID: 1047, Poster
Abstract: Despite strong benchmark performance, multimodal large language models (MLLMs) remain prone to hallucinations, as they are optimized for linguistic plausibility rather than faithful answering based on observation. This issue is particularly severe in long videos, where critical cues are temporally sparse and often missed by uniform frame sampling. Yet existing methods still either operate on fixed pre-sampled inputs, implicitly assuming that all necessary evidence has already been captured, or augment reasoning with retrieval without requiring answers to be grounded in visual evidence. We propose Reinforced Evidence-Aware Learning , a framework that reformulates long-video question answering as an iterative process of evidence gathering and verification, enabling faithful answering with explicit evidence support, thereby reducing hallucinations. At each reasoning step, the model assesses whether the current observation is sufficient to answer the question, and either searches for missing visual evidence from the video or produces a final answer with explicitly cited evidence. To internalize this reasoning policy, we design an evidence-aware reward for reinforcement learning post-training that jointly requires answer correctness and evidence faithfulness. The latter is evaluated by a pretrained cross-modal consistency verifier that matches the generated evidence descriptions against the visual content of the source frames, eliminating the need for manual annotations. Built on Qwen2.5-VL-7B, REAL achieves state-of-the-art results on four long-video benchmarks (Video-MME, MLVU, LongVideoBench, and EgoSchema) and substantially suppresses hallucinations on the dedicated video-hallucination benchmarks VideoHallucer and ELV-Halluc, with only 96 input frames surpassing the 768-frame base model in both accuracy and hallucination resistance.
PaperID: 1048, Poster
Abstract: Large language models (LLMs) trained with reinforcement learning (RL) often exhibit reward hacking, exploiting unintended loopholes in reward functions in ways that can be difficult to detect and eliminate. We propose using prompt optimization—methods which increase an LLM's reward by updating its instructions rather than its weights—to make learned strategies easier to monitor and edit. Applying the GEPA prompt optimizer to environments with exploitable reward functions, we find that optimized system prompts describe reward hacking strategies in highly interpretable language. Furthermore, by simply removing descriptions of unwanted behavior from the optimized system prompt at test time, we can improve the model's alignment while preserving legitimate performance gains. We show that prompt optimization can be guided with an RL-trained teacher LLM, combining the performance advantages of RL with the interpretability of prompting. Finally, we explore an approach to shorten optimized prompts, removing distracting and unhelpful instructions which would otherwise hinder interpretability. We hope that these insights about mitigating misalignment with prompt optimization will aid the discovery of unintended exploits in RL environments and the creation of predictable and monitorable AI systems.
PaperID: 1049, Poster
Abstract: Tree-sliced distances have emerged as scalable alternatives to Sliced Wasserstein distances by projecting probability measures onto tree metric spaces and exploiting closed-form transport on trees. Existing tree-sliced constructions, however, are largely built around first-order Wasserstein geometry. While this choice yields efficient computation, it fixes the discrepancy to an L^1-type aggregation of subtree mass imbalances and provides limited control over the geometry used in downstream optimization. We propose Tree-Sliced Orlicz IPM (TS-Orlicz), a tree-sliced framework that replaces the tree-level W_1 discrepancy with an Orlicz integral probability metric. Building on the tractable formulation of Orlicz IPMs on trees, TS-Orlicz aggregates Orlicz-induced discrepancies over random tree systems while preserving the scalability of tree-sliced computation. By varying the underlying Orlicz function, the framework recovers L^p-type behavior and also supports non-polynomial, tail-sensitive geometries that emphasize large subtree discrepancies. We show that TS-Orlicz preserves key properties of tree-sliced constructions and admits efficient computation via closed-form special cases and one-dimensional scalar optimization. We further extend the framework to probability measures on hyperspheres. Experiments across Euclidean and spherical settings, including gradient flows, self-supervised learning, and diffusion-based generative modeling, show that Orlicz-induced tree-sliced discrepancies achieve competitive or improved performance over sliced and tree-sliced baselines while maintaining low computational cost.
PaperID: 1050, Poster
Abstract: Learning functions f : \mathbbF_2^n \to \mathbbR with sparse Fourier representations plays a central role in computational learning theory, as such functions capture many natural concept classes in machine learning. In recent years, Fourier sparsity has also been studied from the related perspective of property testing, where the goal is to determine, given query access to a function, whether it is s -Fourier-sparse or far from every such function under an appropriate distance measure. Such testers can serve as a preprocessing step for model selection in agnostic learning of Fourier-sparse functions, where validating the structural assumption is crucial before applying computationally expensive learning algorithms. However, all prior work on testing Fourier sparsity considers distance with respect to the uniform distribution over the hypercube. In this work, we initiate the study of Fourier sparsity testing in the distribution-free setting, the property-testing analogue of the PAC+MQ learning framework. In this model, the tester is given oracle access to an unknown function and sample access to an arbitrary and unknown input distribution, and must distinguish s -Fourier-sparse functions from those that are \delta -far from every such function with respect to the underlying distribution. We design a nonadaptive randomized tester that succeeds with high probability using \widetildeO((s/\delta)^4) oracle queries and \widetildeO(1/\delta) samples from the unknown distribution.
Abstract: Uncertainty quantification is central to many applications of causal machine learning, yet principled Bayesian inference for causal effects remains challenging. Standard Bayesian approaches typically require specifying a probabilistic model for the data-generating process, including high-dimensional nuisance components such as propensity scores and outcome regressions. Standard posteriors are thus vulnerable to strong modeling choices, including complex prior elicitation. In this paper, we propose a generalized Bayesian framework for causal inference. Our framework avoids explicit likelihood modeling; instead, we place priors directly on the causal estimands and update these using an identification-driven loss function, which yields generalized posteriors for causal effects. As a result, our framework turns existing loss-based causal estimators into estimators with full uncertainty quantification. Our framework is flexible and applicable to a broad range of causal estimands (e.g., ATE, CATE). Further, our framework can be applied on top of state-of-the-art causal machine learning pipelines (e.g., Neyman-orthogonal meta-learners). For Neyman-orthogonal losses, we show that the generalized posteriors converge to their oracle counterparts and remain robust to first-stage nuisance estimation error. With calibration, we thus obtain valid frequentist uncertainty even when nuisance estimators converge at slower-than-parametric rates. Empirically, we demonstrate that our proposed framework offers causal effect estimation with calibrated uncertainty across several causal inference settings. To the best of our knowledge, this is the first flexible framework for constructing generalized Bayesian posteriors for causal machine learning.
PaperID: 1052, Poster
Abstract: Off-policy reinforcement learning (RL) tries to learn the value of a target policy while data comes from a different behavior policy. This often requires trading off bias for variance or stability. One way to trade these off is to linearly interpolate between the behavior and target policies, but this represents only one point in a broader space of interpolation mechanisms. We introduce \emphexcursion policies--a structured alternative where the agent follows the target policy for a multi-step ``excursion'' before reverting to the behavior policy. We derive dynamic programming results for these policies, establishing the theoretical basis for their use in RL. We further develop provably sound off-policy temporal-difference (TD) algorithms for excursion value estimation. We find that excursion policies provide an efficiency-stability trade-off comparable to mixture policies in continuous control benchmarks, and our analysis illustrates conceptual differences in their behavior.
PaperID: 1053, Poster
Authors: Liwen Zhang, Xinying Fu, Youcheng Zhang, ZijunHu, Wen Chen, Shi Peng, Zhou Jie, Zhe Ma
Abstract: This paper presents an End-to-End learnable radar semantic segmentation network using I/Q echoes as input for Doppler radars, named as EchMo (\textecho-in-mask-out). EchMo reinterprets the steps of conventional radar signal processing (RSP) in a deeply learnable manner, including pulse compression (traditionally implemented via matched filtering), moving target indication (MTI), coherent integration (CI), and target detection. Ultimately, it achieves a single compact learnable deep model that goes from I/Q echoes to target semantic segmentation in one step, delivering a streamlined differentiable computational graph and substantial efficiency gains. Experimental results show that, compared to conventional pipeline or deep models based on intermediate radio-frequency (RF) representation, EchMo achieves better results and offers a more interpretable foundation. EchMo is the first pulse-Doppler I/Q echo-based end-to-end radar semantic segmentation deep architecture, which is a milestone for the fields of RSP, remote sensing, and modern deep learning. The code link for review: https://anonymous.4open.science/anonymize/EchMo-6877.
PaperID: 1054, Poster
Abstract: Aligning representations drawn from a shared distribution without prior correspondence is typically approached through Gromov-Wasserstein Optimal Transport (GWOT) at the coupling level or through generative adversarial (GAN) methods at the map level. Neither yields a uniquely identified alignment map: GWOT returns only a coupling on the available samples, and GAN-based maps are determined only up to marginal matching. Modelling representations as Riemannian metric measure (mm-) spaces, we give sufficient conditions for the map to be unique and show they hold generically. We split recovery into two stages. In the coupling step, an injectivity condition on local distance distributions (LDDs), admitting a sample-level diagnostic, reduces GWOT to a convex linear program. In the map-fitting step, we establish finite-sample identifiability when the unpaired samples come from the same underlying data points, and asymptotic identifiability of the population alignment when they are drawn independently. The resulting
Abstract: We consider the following two-player game: using observational data, the leader chooses a prediction function for a response variable Y from given covariates. The follower then reacts with an intervention on some covariates in the underlying structural causal model to maximize their own objective. The leader knows the intervention targets, but may have limited knowledge of the follower's objective. We call this setup a prediction-intervention game, a special case of a Stackelberg game. Finding an optimal strategy for the leader is generally difficult. To avoid severe performance loss, the leader may base their prediction on the causal parents of Y, or more generally on an invariant subset of covariates. We prove that predictors based on the stable blanket, a specific invariant subset, weakly dominate those based on the causal parents in the game for two common classes of follower objectives. We further upper bound the leader's post-intervention risk by a worst-case risk over allowed interventions and strengthen existing distribution generalization results to analyze this bound: we give sufficient conditions under which stable-blanket predictors are worst-case optimal, and show by examples that these conditions cannot in general be dropped. Finally, we discuss practical strategies when the causal graph is unknown and test them on simulated and real-world data.
PaperID: 1056, Poster
Abstract: Federated graph clustering (FGC) aims to partition distributed graphs into different groups while preserving data privacy. Existing FGC methods focus on spatial-domain information aggregation, while how to use expressive spectral-domain signals of graphs for FGC still remains unexplored. Main challenges are twofold: (1) Collecting and integrating semantic signals that fully express the critical patterns of graphs; (2) Obtaining spectral consensus while preserving local properties of clients. To address these challenges, we propose dynamic spectral federated graph-level clustering (DySFGC). DySFGC designs learnable polynomial filters to capture abundant graph signals and employs spectral contrastive learning and reconstruction to explore the latent semantic information. Besides, DySFGC develops dynamic spectral consensus, which updates the consensus via frequency-specific channels and conducts the dynamic updating process by evaluating the discrepancies between the server and clients. Extensive experiments on fifteen datasets demonstrate the superiority of DySFGC over eleven baselines.
PaperID: 1057, Poster
Abstract: LLM-based agents have demonstrated strong capabilities in data-intensive analytical tasks, yet their outputs are rarely verifiable: a reliance on linear text trajectories makes their reasoning difficult to audit. In particular, deterministic computations over raw data and semantic deductions over natural-language claims are often entangled in an unstructured stream, leaving numerical conclusions hard to reproduce and qualitative judgments hard to inspect. To address this, we propose VeriGraph, a traceable neuro-symbolic reasoning framework that enables agents to construct an explicit heterogeneous evidence directed acyclic graph (DAG) during execution. VeriGraph introduces three evidence-expansion primitives, namely computational, grounding, and derivational expansion, to connect raw data, interpreter variables, computed results, and natural-language claims in a unified graph. Under this formulation, structural traceability is reduced to graph reachability from raw data sources to terminal claims, while semantic support is measured by claim-level evidence evaluation. To improve graph construction, we further design a graph-based policy optimization strategy with a composite reward that jointly supervises answer correctness, computational integrity, and derivational coherence. Experiments on four benchmarks show that VeriGraph-8B achieves the highest overall score among all baselines. More importantly, VeriGraph produces auditable evidence graphs with substantially stronger claim grounding, achieving a 87.61% Grounding Rate under our claim-level evidence support evaluation. These results suggest that explicit evidence-graph construction is a promising path toward verifiable data-analytic agents. Our code is available at \urlhttps://anonymous.4open.science/r/VeriGraph-417E.
Abstract: Making calibrated online predictions is a central challenge in modern AI systems. Much of the existing literature focuses on fully adversarial environments where outcomes may be arbitrary, leading to conservative algorithms that can perform suboptimally in more benign settings, such as when outcomes are nearly stationary. This gap raises a natural question: can we design online prediction algorithms whose calibration error automatically adapts to the degree of non-stationarity in the environment, smoothly interpolating between i.i.d. and adversarial regimes? We answer this question in the affirmative and develop a suite of algorithms that achieve adaptive calibration guarantees under multiple calibration measures. Specifically, with T being the number of rounds and C\in[0,T] being an unknown non-stationary measure defined as the minimal \ell_1 deviation of the mean outcomes, our algorithms attain \widetildeO(\sqrtT+(TC)^\frac13) for \ell_1 calibration error and \widetildeO((1+C)^\frac13) for both \ell_2 and pseudo KL calibration error. These bounds match the optimal rates in the stationary case (C=0) and recover known guarantees in the fully adversarial regime (C=T). Our approach builds on and extends prior work [Hu et al., 2025, Luo et al., 2025], introducing an epoch-based scheduling together with a novel non-uniform partition of the prediction space that allocates finer resolution near the underlying ground truth.
Abstract: We study generation in separable metric instance spaces. We extend the language generation framework from Kleinberg and Mullainathan (2024) beyond countable domains by defining novelty through metric separation and allowing asymmetric novelty parameters for the adversary and the generator. We introduce the (\varepsilon_a,\varepsilon_g)-closure dimension, a scale-sensitive analogue of closure dimension, which yields characterizations of uniform and non-uniform generatability and a sufficient condition for generation in the limit. Along the way, we identify a sharp geometric contrast. Namely, in doubling spaces, including all finite-dimensional normed spaces, generatability is stable across novelty scales and invariant under equivalent metrics. In general metric spaces, however, generatability can be highly scale-sensitive and metric-dependent; even in the natural infinite-dimensional Hilbert space \ell^2, all notions of generation may fail abruptly as the novelty parameters vary.
PaperID: 1060, Poster
Abstract: Policy gradients for autoregressive sequence models differentiate each token's log-probability while holding the conditioning history fixed. This discards information about how earlier decisions shape later conditional distributions, creating a credit-assignment bottleneck when rewards are sparse. We introduce History-Reparameterized (HR) policy gradients, which replace the frozen prefix in the score function with a differentiable reparameterized or relaxed history, allowing gradients of downstream log-probabilities to propagate backward through the generated trajectory. The resulting history-pathwise term is generally biased when added directly, so we derive Optimal HR (OHR), a centered control-variate estimator with a closed-form coefficient that minimizes the trace of the gradient variance at the population level. The variance reduction is governed by temporal structure in the policy dynamics: random or weakly structured policies provide little benefit, while policies with learned sequential dependencies yield informative history-pathwise corrections. For discrete tokens, the choice of relaxation is crucial: embedding-level relaxation yields substantial variance reduction that grows with horizon, whereas straight-through Gumbel-Softmax yields essentially none in our diagnostics. We validate the method on a Gaussian linear Markov model, a sparse-reward T-maze with a recurrent policy, and an autoregressive Transformer model for small Traveling Salesman Problem (TSP) instances.
PaperID: 1061, Poster
Abstract: Diffusion models generate data through a continuum of denoising problems, and are widely observed to reveal coarse structure before fine detail. Yet, this intuition is mostly empirical and qualitative. We introduce \emphfeature information dynamics, an information-theoretic framework for localizing when a feature is generated during diffusion. Using the I-MMSE identity, we connect the rate of feature mutual information change to a gap between optimal unconditional and feature-conditional denoising losses, yielding practical estimators for feature information density. We further develop a chained decomposition that separates shared from incremental information in a feature hierarchy. We use this framework first to quantitatively confirm spectral autoregression in pixel diffusion, and then to extend the analysis beyond frequency: under a class\tomask\toCanny conditioning chain, the per-feature information densities differ across pixel, SDVAE, VAVAE, and RAE, exposing fundamental differences between these representations.
PaperID: 1062, Poster
Abstract: As AI development increasingly involves models adapted to downstream settings, a new algorithmic challenge for developers has surfaced: deciding when to adopt model updates. Currently, the only reliable approach is to retrain on new upstream models and then evaluate extensively, which can be prohibitively expensive. We introduce a framework for assessing when to propagate newly released model versions to downstream applications that does not require access to upstream data or retraining a priori. Our framework incorporates geometric similarity and information-theoretic sufficiency to determine when an upstream update substantially shifts the representational basis recruited for a downstream task. This enables targeted update adoption, and supports emerging norms for coordinating AI supply chain infrastructure.
PaperID: 1063, Poster
Abstract: To defend against backdoor attacks on neural networks, the defender must identify the verifier function that determines whether an input is a trigger, rather than rejecting individual triggers in isolation, since an attacker can generate new triggers satisfying the same verifier, against which trigger-by-trigger blocking cannot keep up. If a backdoor verifier could be installed with the same kind of asymmetry as a cryptographic authentication scheme, in which only the attacker can produce trigger inputs while the defender cannot reverse-engineer the verifier from observations, learning-based defense would face a principled impossibility. We show that no such asymmetry can arise for any backdoor installed by training into a model. We establish two complementary impossibility results. First, any verifier installable by training a polynomial-size neural network is reconstructible by the defender with sample complexity matching the attacker's up to polynomial factors. Second, any cryptographic secret the attacker might supply to the model at inference is necessarily observable to a white-box defender. Together, these results rule out any way of giving a self-contained backdoor cryptographic asymmetry. We confirm both routes empirically on pretrained large language models, demonstrating that no configuration yields a backdoor simultaneously effective for the attacker and irrecoverable by the defender.
PaperID: 1064, Poster
Abstract: Conformal Prediction (CP) is a widely used technique for quantifying uncertainty in machine learning models. In its standard form, CP offers probabilistic guarantees on the coverage of the true label, but it is agnostic to sensitive attributes in the dataset. Several recent works have sought to incorporate fairness into CP by ensuring conditional coverage guarantees across different subgroups. One such method is Conformal Fairness (CF). In this work, we extend the CF framework to the Federated Learning setting and discuss how we can audit a federated model for fairness by analyzing the fairness-related gaps for different demographic groups. We empirically validate our framework by conducting experiments on several datasets spanning multiple domains, fully leveraging the exchangeability assumption.
PaperID: 1065, Poster
Authors: Noam Nisan, Emmanuel Zerah
Abstract: An influential paper of Calvano et al. empirically demonstrated that Q-learning agents spontaneously collude when placed as sellers that compete on prices in a natural market model. More recent results of Fish et al. empirically demonstrated that similar collusion happens with commercial LLMs. We formally prove that such collusion can also happen with external-regret-minimizing agents. We identify a very general class of agents, which we term Domination-Avoiding agents, that provably do not collude in such markets. This class contains all Mean-Based agents and all internal-regret-minimizing agents, as well as others such as Multiplicative-Weight agents with variable learning rate and contextual variants thereof. More generally we show that, in any game, this class of agents is guaranteed to jointly learn to almost never play strategies that are eliminated by repeated elimination of purely dominated strategies.
Abstract: We study the long-context limit of softmax self-attention with a fixed query and a random context of n i.i.d. keys on the sphere, viewing the inverse temperature \beta_n as the scaling parameter that decides whether attention degenerates into uniform averaging or collapses onto the single closest key. We show that the critical scale at which selectivity emerges is determined by the local exponent of the distance-to-query distribution near zero rather than by global features of the context, and scales like \beta_n^\ast \asymp n^2/(d-1) for uniform keys on \mathbbS^d-1. Furthermore, we characterize the limiting laws of the ordered attention weights and of the attention output across all regimes of \beta_n: a subcritical regime in which the output reduces to a local average around q with explicit deterministic bias and Gaussian fluctuations; a critical regime in which a finite collection of nearest keys retains macroscopic mass without single-key collapse; and a supercritical regime in which all mass concentrates on the closest key. Of notable interest is the subcritical case with identity value matrix where the attention map approximately implements a backward heat equation.
Abstract: Markov Games are the standard formal model for multi-agent reinforcement learning, capturing agents that act in a shared state and optimize rewards over time. However, in real-world settings, agents’ decisions are often influenced by unobserved factors, such as cognitive biases, behavioral tendencies, or intuitive signals, that also affect rewards and future states. Ignoring these unobserved confounders can make standard learning methods converge to suboptimal policies. In such settings, optimal play may require policies that condition on counterfactual signals, requiring reasoning at the counterfactual layer of the Pearl Causal Hierarchy. In this paper, we introduce Causal Markov Games (CMGs), a framework for modeling sequential multi-agent decision making in the presence of unobserved confounding. We show that CMGs strictly generalize Markov Games, with arbitrarily large gaps between classical interventional equilibria and causal counterparts. We then develop two learning algorithms under different observability assumptions. The first, CNash-VI-FO, is a model-based learner with finite-sample guarantees when agents’ natural actions or intuitions are revealed post hoc. The second, CNash-VI-NO, is an explore-then-exploit procedure with asymptotic guarantees for the setting in which opponents’ natural actions are never observed. To address scalability, we further provide a drop-in counterfactual augmentation of deep MARL. Empirically, on a confounded windy variant of the Multi-Particle Environment and the Iterated Causal Prisoner’s Dilemma, counterfactual agents strictly dominate their non-causal counterparts.
PaperID: 1068, Poster
Authors:
Ziyang Cai, Seyyedamirhossein Saeidi, Harkirat Singh BehlAbstract: With the advent of AI agents, automated scientific discovery is becoming an increasingly plausible goal. However, training agents to autonomously execute the engineering-heavy labor of machine learning (ML) research requires massive, process-level supervision. Existing static benchmarks omit critical intermediate steps such as debugging and incremental reasoning, and manual data collection is prohibitively expensive. To overcome this data bottleneck, we introduce ML-AutoResearch (ML-AR), a scalable pipeline for automatically generating synthetic, end-to-end ML research tasks. Each task defines a complete research cycle, including problem specification, dataset selection, baseline implementation, and iterative improvement. To ensure realism and executability, tasks are grounded in real-world datasets and refined via an automated self-debugging procedure without requiring human supervision. We construct a large-scale dataset of teacher trajectories on these synthetic tasks to train student agents via supervised fine-tuning. We evaluate the resulting agents across 3 diverse ML research benchmarks. Our comprehensive experiments across 2 distinct model families and 3 model sizes demonstrate that training on ML-AR trajectories yields consistent and significant capability gains. Fine-tuning improves the Aggregation Under the Performance (AUP) by up to 18% and substantially boosts overall pass rates, highlighting robust out-of-domain generalization.
Authors:
Pablo Moreno-Muñoz, Adrian Müller, Gergely NeuAbstract: We propose a new framework for generative modeling based on a discrete-time stochastic control formulation of measure transport. Adapting classic results from control theory, we formulate our problem as a linear program whose dual variables correspond to the \emphoptimal value function of the control problem, which directly encodes the optimal control policy. Exploiting this LP formulation, we develop an efficient simulation-free primal-dual algorithm for computing approximately optimal value functions and the associated \emphvalue-driven transport (VDT) policies which approximate the true optimal policy. We show that well-trained VDT policies enjoy numerous favorable properties in comparison with other state-of-the-art methods based on flows, diffusions, or Schr\"odinger bridges: they lead to straight transport paths which can be simulated quickly and robustly, and can be enhanced in all the same ways as diffusion and flow-based models (e.g., conditional generation, classifier-free guidance, unpaired image-to-image translation are all easy to incorporate). We establish several foundational links with the aforementioned models and evaluate our methodology in an extensive range of experiments, all of which indicate strong performance.
Abstract: Traditional recommendation systems rely on latent (dense) representations, making them difficult to interpret and control. We propose the Controllable and Content-Based Recommendations (CCBR) framework, which builds its recommendations from textual user profile representations. CCBR plugs into collaborative filtering models and introduces controllability via text bottlenecks. We show that CCBR enables text-based and multimodal interventions, allowing users to steer the model towards the directions they prefer. Different from existing controllable recommendation systems, CCBR infers the text summaries directly from item contents (images, audio or video). Across image-, audio-, and video-based datasets, we demonstrate that the proposed framework obtains competitive model performance with standard (latent-representation) models while providing controllable model summaries via text. The model also outperforms TEARS, a recent baseline for controllable recommendation systems. Through systematic interventions, we demonstrate the efficacy of the user steering mechanism.
PaperID: 1071, Poster
Abstract: We study seamless world generation, where a model must update a persistent visual world according to online user operations rather than full next-scene descriptions. To make this task well-defined, we introduce an operation-centric taxonomy that represents each scene with a global-local schema and each boundary with a typed operator program specifying change and preservation. This taxonomy enables SeamlessWorld, a synthetic data and benchmark pipeline with 8 operation families, 23 atomic operators, dual prompt surfaces, and schema-derived evaluation probes. We then develop OopsWorld, a causal video generator that learns from SeamlessWorld using operation-conditioned decoupled distillation: the student is driven by operation prompts and visual history, while scene-conditioned teachers provide target-state supervision. We further apply dynamics curriculum learning to progressively compose simple edits into high-dynamic world transitions. OopsWorld substantially improves dynamic responsiveness and long-video quality, achieving 91.67 VBench Dynamics and 80.42 Long-Quality with 100K samples. A systematic explicit--implicit study further reveals a persistent operation-prompting gap in current systems, especially for preserving visual anchors, calling for operation-driven world models beyond next-scene prompt switching.
Abstract: Autoregressive LLM decoding evaluates every generated token through the full layer stack, even though many tokens become predictable at intermediate depths. Existing lossless depth-adaptive methods exploit this redundancy by choosing a single non-final exit depth and verifying its prediction with the final-depth model. However, our measurements show that this selection-based strategy leaves substantial headroom: choosing an exit too late wastes computation, while choosing one too early triggers fallback and discards dependent drafts. We propose Depth Exploration Decoding (DEX), a lossless decoding algorithm that replaces single-depth selection with parallel exploration over multiple candidate depths. At each commit position, DEX validates candidates against the final-depth reference, commits exactly the final-depth token, and collapses the exploration lattice to retain only reusable branch states. This expand--commit--collapse procedure preserves equivalence to standard autoregressive decoding while reducing the cost of committing each token. Across early-exit-trained and standard LLMs, DEX outperforms representative depth-selection baselines and achieves competitive end-to-end throughput against speculative and distributed decoding methods. Moreover, DEX improves as the explored depths become finer, showing that parallel depth exploration provides a scalable way to exploit the underused depth axis of LLM decoding. Code: https://anonymous.4open.science/r/DEX--anonymous/.
PaperID: 1073, Poster
Authors:
Da-Woon Heo, Kwanseok Oh, Adrián Ayuso-Muñoz, Heung-Il SukAbstract: Decoding cognitive functions and brain diseases from functional magnetic resonance imaging (fMRI) is a foundational problem in neuroscience advanced by recent deep learning across task and resting-state (rs) settings. Existing region-level decoders capture spatial structure through region-of-interest (ROI)-aware graph or Transformer architectures, but apply a uniform temporal operator across all regions. Consequently, they do not explicitly model how cortical regions integrate information over heterogeneous timescales. We propose CoRTeX, a region-resolved recurrent architecture for fMRI in which each cortical region carries its own learnable parameters and evolves on its own learned timescale. At its core, per-region differentiable temporal integration windows allow each ROI to learn how far back in time to integrate information through a soft mask over a private buffer. Cross-region information is restricted to a single identity-aware attention channel, additionally biased by a co-activation-driven memory whose learning and forgetting rates are themselves region-specific. Experiments on stimulus classification using task-fMRI (NSD) and on two disease diagnosis tasks using rs-fMRI (ABIDE, COBRE) demonstrate that CoRTeX achieves competitive performance with strong baselines while enabling region-level analyses that are not well-defined for models with exchangeable units: the learned per-region timescales exhibit interpretable spatial organization, and class-specific regional attribution aligns with category-selective cortex. Code is available at: https://anonymous.4open.science/r/CoRTeX-F.
PaperID: 1074, Poster
Abstract: Competition among LLM providers hinges not only on model accuracy, but also on the latency of serving user requests. However, user demand leads to congestion which increases latency, creating a feedback loop between provider decisions and user choices. In this work, we study a stylized game between two providers who strategically balance allocating fixed compute between training better models and reducing latency. A population of users then chooses between the providers. Our equilibrium analysis reveals that a lower-budget provider can still attract users by taking advantage of the congestion faced by the dominant provider. Relative to a competitive baseline where providers have equal compute resources, compute asymmetry leads the higher-resource provider to divert a greater fraction of their compute from improving accuracy to reducing latency. Moreover, this compute asymmetry strictly harms users. Altogether, our results illustrate how congestion distorts competition between model-providers, leading to subtle implications for market competitiveness, model accuracy, and user utility.
Abstract: Causal discovery aims to uncover the causal relationships among variables beyond correlations from observational data. On time series, a major challenge is nonstationarity, which is typically modeled as regime-based shifts or drifting causal strength. In contrast, we consider non-stationarity arising from varying cause-to-effect delay, and introduce a causal model that is identifiable with few assumptions. To discover causal relationships with varying causal delay in practice, we formalize the problem in terms of the algorithmic model of causation, and propose STRETCH to discover causal graphs. We show on synthetic data that our method correctly recovers the causal relationships, and show the relevance of this causal model on a case study on the impact of El Niño climatic phenomenon on Indian monsoon.
Abstract: Bayesian persuasion, a central model in information design, studies how a sender, who privately observes a state drawn from a prior distribution, strategically sends a signal to influence a receiver's action. A key assumption is that both sender and receiver share the precise knowledge of the prior. Although this prior can be estimated from past data, such assumptions break down in high-dimensional or infinite state spaces, where learning an accurate prior may require a prohibitive amount of data. In this paper, we study a learning-based variant of persuasion, which we term persuasive prediction. This setting mirrors Bayesian persuasion with large state spaces, but crucially does not assume a common prior: the sender observes covariates X, learns to predict a payoff-relevant outcome Y from past data, and releases a prediction to influence a population of receivers. To model rational receiver behavior without a common prior, we adopt a learnable proxy: decision calibration, which requires the prediction to be unbiased conditioned on the receiver's best response to the prediction. This condition guarantees that myopically responding to the prediction yields no swap regret. Assuming the receivers best respond to decision-calibrated predictors, we design a provably efficient algorithm that learns a decision-calibrated predictor within a randomized predictor class that optimizes the sender's utility. In the commonly studied single-receiver case, our method matches the utility of a Bayesian sender who has full knowledge of the underlying prior distribution. Finally, we extend our algorithmic result to a setting where receivers respond stochastically to predictions and the sender may randomize over an infinite predictor class.
PaperID: 1077, Poster
Abstract: ordinary differential equations (ODEs) to model neuron potential dynamics. Recently, fractional-order spiking neural networks ( -SNNs) extend conventional SNNs by incorporating fractional-order ODEs, enabling the modeling of long-range temporal dependencies in neuron potential states. However, existing first-order SNNs and -SNNs are typically restricted to a single derivative order, which limits their capacity to represent heterogeneous temporal dynamics across different time scales. In this work, we propose Distributed-Order Fractional Spiking Neural Networks ( -SNNs), which generalize single-order neuron dynamics by integrating multiple fractional orders through learnable distributed weights. This formulation enables neurons to accumulate historical information across within a unified framework. We develop practical numerical implementations based on the Grünwald–Letnikov discretization, allowing efficient training of -SNNs within standard SNN architectures. We show that the proposed distributed-order formulation induces intrinsic multi-scale temporal memory, providing a richer class of temporal representations compared to single-order fractional dynamics. Extensive experiments demonstrate consistent performance gains of
PaperID: 1078, Poster
Abstract: pretraining curricula influence learning, generalization, and the selectivity of fine-tuning is unclear. This is important for AI safety, where fine-tuning is used to selectively suppress misaligned behaviors. Here, we compare curricula that pretrain tasks in a balanced (sampled uniformly) or an imbalanced (one task early, the other late) fashion. We show that imbalanced learning of two conflicting copy tasks promotes in-context learning and improves the selectivity of refusal fine-tuning. Ablations and activation patching show that this occurs because imbalanced pretraining encourages tasks to be disentangled in separable neural circuits, whereas balanced training routes both tasks through a common pathway. We extend these findings to a synthetic language learning task involving rule-consistent and rule-violating data, where imbalanced curricula similarly lead to more localized, less entangled rule representations, resulting in more robust rule-following behavior. Together, these results suggest that imbalanced pretraining curricula may be an important tool for promoting disentangled representations, with direct consequences for the precision and reliability of safety fine-tuning.
Authors: Ziv Aharoni, Henry Pfister
Abstract: While autoregressive models optimize the exact likelihood implied by the chain rule, diffusion models are typically trained with denoising objectives. We develop conservation laws based on generalized extrinsic information transfer (GEXIT) functions for a broad class of memoryless noise processes, showing that the data--model cross-entropy (CE) can be characterized \emphexactly as an integral of \emphlocal information-theoretic derivatives along the noise path. This yields a unified characterization of the likelihood for discrete diffusion and continuous diffusion, with the Gaussian case reducing to the well-known I-MMSE relationship connecting mutual information and estimation theory. An immediate implication is a \emphlocality property: one can compute the information-theoretic derivatives using only the marginal posteriors along the noise path. As a result, training reduces to learning the marginal posteriors by minimizing the negative log-likelihood. While the conservation law implies that the entropy does not depend on the noise path, finite-capacity denoisers approximate the posteriors with varying accuracy across noise types, leading to differences in performance. We validate these predictions on synthetic Markov sources and standard benchmarks, including \texttttext8 and CIFAR-10.
Abstract: Sparse Autoencoders (SAEs) that can accurately reconstruct their input (minimizing distortion) by making efficient use of few features (minimizing the rate) often fail to learn monosemantic representations (highly interpretable), limiting their usefulness for mechanistic interpretability. In this paper, we characterise this tension in learning faithful, efficient, and interpretable explanations, introducing the Rate-Distortion-Polysemanticity tradeoff in SAEs. Under toy-modeling assumptions, we theoretically and empirically show that restricting the SAE to be monosemantic necessarily comes with an increase in rate and distortion. Assuming a generative model behind the input observations, we further demonstrate that the degree of polysemanticity of optimal SAEs is determined by the training data distribution, especially by the probability of features to co-occur. Finally, we extend the analysis to real-world settings by deriving necessary conditions that a polysemanticity measure should satisfy when the data-generating process is unknown, and we benchmark existing proxy metrics on SAEs trained on Large Language Models. Taken together, our findings show that polysemanticity is a data problem that should be accounted for when addressing it at the architectural and optimization level.
Abstract: We propose Sobolev-regularized Maximum Mean Discrepancy (SrMMD) gradient flow, a regularized variant of maximum mean discrepancy (MMD) gradient flow based on a gradient penalty on the witness function. The proposed regularization overcomes the non-convexity of the MMD objective and yields provable \emphglobal convergence guarantees in MMD in both continuous and discrete time. A more surprising appeal is that our convergence analysis does not rely on isoperimetric assumptions on the target distribution. Instead, it is based on a regularity condition on the difference between kernel mean embeddings. A key highlight of the proposed flow is that it is applicable in both sampling (from an unnormalized target distribution)---using Stein kernels---and generative modeling settings, unlike previous works, where a gradient flow is suitable for only generative modeling or sampling but not both. The effectiveness of the proposed flow is empirically verified on a broad range of tasks in both generative modelling and sampling.
PaperID: 1082, Poster
Abstract: Learned representations (or embeddings) are the foundation of modern machine learning systems, yet they are typically trained as fixed-size embeddings, without accounting for varying downstream resource or task constraints. Matryoshka Representation Learning (MRL) addresses this limitation by learning nested representations, but requires costly end-to-end retraining and does not support post hoc expansion, while Contrastive Sparse Representations (CSR) reduce training overhead via lightweight adaptors but operate at a 4x larger dimensionality. Additionally, all prior works restrict their framework towards compression only. In this work, we introduce Hyperbolic Elastic Representation Learning (HERL), a framework for learning dimension-adaptive embeddings (compression and expansion) through a lightweight adaptor without retraining the base encoder. Our key idea is to exploit the exponential capacity of hyperbolic geometry by learning a non-linear mapping from the encoder space to hyperbolic space. We train a simple MLP to downsample the representations before projecting them onto a Lorentzian and learn adaptive representations that respect the Lorentzian geometry. Our approach is model-agnostic and effective in both supervised and unsupervised settings. Through extensive experiments across image and text modalities, we show that HERL preserves downstream performance across a wide range of embedding sizes, exhibits smooth interpolation between dimensions, and achieves these benefits at a fraction of the computational cost.
PaperID: 1083, Poster
Authors: Rithvik Redrouthu, Jackson Stokes
Abstract: Group Relative Policy Optimization (GRPO) is widely used to post-train reasoning language models, and its rollout group size G is typically treated as a compute hyperparameter. We argue that G also sets a difficulty threshold below which prompts produce no learning signal. Under binary verifier rewards, a rollout group only produces a learning signal when it contains both a success and a failure, and the rate at which a prompt yields such mixed groups turns on sharply near p_c~eq\ln 2/G. Below that threshold the prompt is invisible to the relative-advantage estimator regardless of optimizer state. We verify this gate directly in pass@256 rollout traces from Qwen3-1.7B, Qwen3-8B, and Llama-3.1-8B, where empirical mixed-group rates match the binomial prediction within 1.4–1.9 percentage points on average per model. Sweeping fixed-G GRPO with G\in\8,16,32,64,128\ then shows that difficulty-binned per-prompt improvement after training tracks the frontier shape with binned R^2 between 0.962 and 0.984 across all three model families. Treating p_c as a moving target, we instantiate a frontier-tracking group-size schedule that improves the best fixed-G baseline by 4.2–5.3 points on MATH-500 and 2.4–3.4 points on GSM8K under matched rollout budget. Ablations attribute the gain primarily to frontier-aware switch timing.
Abstract: Instruction-based image editing exhibits heterogeneous difficulty not only across cases but also across regions of an image, motivating refinement approaches that allocate correction to where the model struggles. Existing refinement signals arrive late, after a fully generated image or a completed denoising step. We ask whether such a signal can act an ongoing forward pass. To investigate this, we probe a frozen image-editing model and find that although generation capability emerges only in the last few layers, the error pattern is already set in early layerss (rank correlation ρ = 0.83 with the final-layer error map). Based on this, we introduce , a learnable token that critiques a frozen model's predictions at its intermediate layers and steers its hidden states to refine generation during the forward pass. A three-stage recipe is proposed to stabilize the training from learning how to critique to steering generation. As a result, we achieve state of the art on GEdit-Bench (7.89), a +9.4 gain on RISEBench over the same backbone, and the strongest open-source result on KRIS-Bench (81.92, surpassing GPT-4o). We further provide analyses showing that the critic genuinely shapes the model's attention and prediction updates at subsequent layers.
PaperID: 1085, Poster
Abstract: We construct the first coreset of size near-linear in k for fair \kMedian in general metric spaces. Combined with the standard merge-and-reduce framework, our construction also gives the first streaming algorithm for fair \kMedian with space complexity near-linear in k. The main technical innovation is a distributional reinterpretation of capacitated clustering that the cost of assigning a dataset to a capacitated center set can be expressed as the earth mover's distance between two discrete distributions. This viewpoint enables us to use tools from optimal transport to analyse uniform sampling. A key ingredient in this analysis is a new family of problem-specific \eps-nets for the earth mover's distance.
Abstract: Inorganic crystals are periodic, highly symmetric arrangements of atoms in three-dimensional space. Their structures are constrained by the symmetry operations of a crystallographic \emphspace group and restricted to lie in specific affine subspaces known as \emphWyckoff positions. The frequency an atom appears in the crystal and its rough positioning are determined by its Wyckoff positions. Most generative models that predict atomic coordinates overlook these symmetry constraints, leading to unrealistically high populations of proposed crystals exhibiting limited symmetry. We introduce Space Group Conditional Flow Matching (SGFM), a novel generative framework that samples significantly closer to the target population of highly symmetric, stable crystals. We achieve this by conditioning the entire generation process on a given space group and set of Wyckoff positions; specifically, we define a conditionally symmetric noise distribution and a group-conditioned, equivariant, parametric vector field that restricts the motion of atoms to their initial Wyckoff position. Our form of group-conditioned equivariance is achieved using an efficient reformulation of \emphgroup averaging tailored for symmetric crystals. It reduces the computational overhead of symmetrization to a negligible level. SGFM achieves state-of-the-art performance in De Novo Generation on the MP-20 benchmark.
Abstract: Recent one-step generative models accelerate sampling by learning deterministic flow maps of the underlying dynamics. These methods rely on learning from ordinary differential equations, leaving open how to define an exact distillation procedure for stochastic dynamics. We introduce the Itô map, an any-step stochastic flow map that takes an intermediate state and Brownian path and predicts future states in a single pass. The Itô map formulation yields novel estimators for inference-time control by providing cheap, differentiable access to posterior samples. Empirically, Itô maps produce diverse, conditionally valid endpoint samples from fixed intermediate states and support strong steering performance on synthetic and image-generation benchmarks. These results establish any-step SDE integration as a useful primitive for posterior sampling and stochastic control.
PaperID: 1088, Poster
Authors: Takanori Maehara
Abstract: We develop a pattern-algebraic framework for higher-order graph neural networks by extending Deep Homomorphism Networks (DHN) from rooted patterns to k-labelled patterns. The framework captures GNN expressivity along two independent axes --- skeleton complexity k, controlling global structures, and pattern vocabulary \mathcalP, capturing local structures. We then introduce \emphauxiliary gluing (\circledast), the pattern-algebraic counterpart of the k-FWL sift, and prove that this single additional operation lifts k-label expressive power to (k+1) labels, yielding an algebraic proof of (k+1)-WL \equiv k-FWL. Neuralising these operations gives k-DHN and k-FDHN, which recover k-GNN, PPGN, \delta-k-LWL, r-\ellWL, spectral invariant GNNs, and node-marking subgraph GNNs as special cases, with expressivity exactly characterised by the pattern closure of (k, \mathcalP). The framework also resolves open expressivity separation problems in the \delta-k-LWL and r-\ellWL hierarchies.
Abstract: Domain adaptation faces a fundamental paradox in the cold-start regime. When target data is scarce, statistical methods fail to distinguish relevant source domains from irrelevant ones, which often leads to negative transfer. In this paper, we address this challenge by leveraging expert textual descriptions of the target domain, a resource that is often available but overlooked. We propose a probabilistic framework that translates these semantic descriptions into a choice model, namely a Language-Induced Prior (LIP), that learns the preferences from a pre-trained Large Language Model (LLM). The LIP is then integrated into an Expectation-Maximization (EM) algorithm to identify source relevance. Methodologically, this framework is compatible with any parametric model where a likelihood is available. It allows the LIP to guide the selection of sources when target signals are weak, while gradually refining these choices as samples accumulate. Theoretically, we prove that the estimator roughly matches an oracle cold-start MSE under a correct prior, while remaining asymptotically consistent regardless of the quality of the LIP. Empirically, we validated the framework on a descriptive (Gaussian estimation), a predictive (C-MAPSS dataset), and a prescriptive task (MuJoCo Hopper).
Authors:
Zimo Wang, Ishit Mehta, Haolin Lu, Xunpeng Huang, Chung-En Sun, Ge Yan, Lily Weng, Tzu-Mao LiAbstract: Modern generative models based on diffusion or flow matching transport noise toward the data manifold, without representing the data manifold explicitly. We propose \dm that explores an alternative design of generative modeling by explicitly modeling a distance field of the data manifold. Designing loss functions to learn distance fields in high-dimensional space from point data can be challenging. This is because the data do not provide reliable signals for closest point projection onto the manifold. To mitigate this problem, we design new loss functions that weight training data by their proximity to each sampling position. This distance-field perspective naturally motivates two principled deterministic samplers, showing fast convergence in inference. Our model is related to the recent time-unconditional generative models since it also does not require time input, and our analysis reveals that our method helps time-unconditional generative modeling to disambiguate the denoising problem and to avoid biasing towards the data mean. As a result, our model is the only time-unconditional model that remains competitive to time-conditional ones across datasets and architectures. Moreover, our distance prediction is also helpful for early stopping during sampling and for OOD detection. We hope distance-field modeling can open a promising direction for exploring generative modeling from a geometric perspective.
PaperID: 1091, Poster
Authors:
Srikar Yellapragada, Alexandros Graikos, Zilinghan Li, Kostas Triaridis, Tarak N Nandi, Ji D Bai, Beatrice Knudsen, Tahsin Kurc, Prateek Prasanna, Rajarsi R Gupta, Ravi K Madduri, Joel Saltz, Dimitris SamarasAbstract: Whole-slide images (WSIs) are gigapixel-scale digital scans of histopathology tissue and serve as the primary basis for cancer diagnosis. Existing histopathology generative models operate at the patch scale, producing fixed-size tiles of at most 1024×1024 pixels which encompass only a small fraction of the WSI. In contrast, text reports for clinical diagnoses and slide-level diagnostic labels -- the supervision that exists at scale -- operate at the whole-slide level. We present BigCell, the first approach to generate coherent whole-slide histopathology images at gigapixel resolution. BigCell trains a flow-matching transformer at 2048×2048 pixels on over 120,000 publicly available WSIs spanning multiple organ types. Conditioning on embeddings from pathology foundation models and using an inference-time sliding-window pipeline, BigCell generates coherent images of up to 32,768×32,768 pixels, outperforming prior tile generators on both tile-level and slide-level fidelity metrics. To make gigapixel synthesis practical, we introduce an adaptive scheduling strategy that prioritizes important slide regions, speeding up image generation by over 3\ttimes with minimal quality degradation. Furthermore, we propose the first text-to-image synthesis at 8k resolution by training a separate model that produces dense conditioning grids directly from pathology text reports. We demonstrate that augmenting with synthetic samples improves slide-level few-shot classification performance by up to 18 % over baselines trained on real data alone.
Abstract: Neural representations are not unique objects. Even when two systems realize the same downstream computation, their hidden coordinates may differ by reparameterization. A probe family intended to reveal structure already present in a representation should therefore be stable under the relevant representation symmetries rather than be tied to a particular basis. We study this group action in the tractable exact setting of the final readout layer, where equivalent realizations induce affine changes of hidden coordinates. The resulting symmetry principle singles out a unique hierarchy of shallow coordinate-stable probes, with linear probes as its degree-1 member. We also show that a natural object for cross-model probe transfer is a shared probe-visible quotient--the representation modulo directions invisible to the probe family--rather than the full hidden state. Experiments on synthetic and real-world tasks support both predictions, showing where degree-2 probes help beyond linear ones and how quotient-based transfer enables coverage-aware monitor portability across model families. These results point toward a broader geometric representation theory of neural probing, with coverage-aware monitor transfer as a concrete operational consequence.
Abstract: The efficient computation of parametric solution sensitivities is a key challenge in the integration of learning-enhanced methods with nonlinear model predictive control (MPC), as their availability is crucial for many learning algorithms. This paper discusses the computation of solution sensitivities of general nonlinear programs (NLPs) using the implicit function theorem (IFT) and smoothed optimality conditions treated in interior-point methods (IPM). We detail sensitivity computation within a sequential quadratic programming (SQP) method which employs an IPM for the quadratic subproblems. Previous works presented in the machine learning community are limited to convex or unconstrained formulations, or lack an implementation for efficient sensitivity evaluation. The publication is accompanied by an efficient open-source implementation within the framework, providing both forward and adjoint sensitivities for general optimal control problems, achieving speedups exceeding 3x over the state-of-the-art solvers
PaperID: 1094, Poster
Authors: Victor Azad
Abstract: Reinforcement learning (RL) agents face a continuous trade-off between exploiting known strategies and exploring novel states to obtain higher rewards. While exploration is often undirected, it can also be guided by intrinsic signals such as novelty or uncertainty. However, combining these signals with task rewards is challenging, as they can sometimes drive the agent away from optimal trajectories and reduce final performance. We find that this degradation arises from a structural occupancy conflict: as the policy improves, exploration signals naturally begin to anti-correlate with task advantages. In this work, we introduce Kantorovich Dual Shaping (KDS), a framework that resolves this conflict through constrained policy-space optimization. KDS formulates exploration as an Optimal Transport problem and applies the Sinkhorn-Knopp algorithm on the empirical batch manifold to reweight intrinsic rewards in a geometry-aware, point-by-point manner. When combined with standard RL objectives, it acts as a spatial filter that suppresses exploration in conflicting regions while enhancing it elsewhere. We evaluate this general approach across multiple off-policy and on-policy algorithms, achieving stable exploration that consistently avoids task degradation in complex continuous control environments.
PaperID: 1095, Poster
Abstract: Prompting steers large language models (LLMs) and vision--language models (VLMs) without weight updates, but it remains unclear how a change in instruction reshapes internal representations to produce a behavioral effect. We introduce a nested geometric decomposition framework that treats prompting as a transformation of the representational geometry for the content following the prompt. We ask what class of mathematical transformation best explains the effect of prompting by finding the best alignment between representations of the same stimulus set following different prompts. For each prompt pair, we fit a sequence of increasingly expressive stimulus-invariant maps: translation, rigid transformation with uniform-scaling, sequential axis scaling, affine, and nonlinear transformations. We then test these maps causally by replacing a single layer's prompt-A hidden state for a new set of stimuli with its mapped counterpart and measuring recovery of prompt-B representational geometry and behavior. Across three LLMs, three VLMs, and six text or image datasets varying in style, emotion, scene content, and number, prompts consistently reshape representational geometry toward the instructed task structure. In the cross-validated nested variance decomposition, much of the prompt-induced activation change is explained by shape-preserving maps: translation and rigid transformation with uniform-scaling. The tier profiles reveal model- and task-specific routing strategies, differing in how much transformation classes explain variance and where along the layer hierarchy their contributions emerge. Crucially, although translation and rigid transformation tiers already improve behavioral agreement, affine transformation is the first tier to nearly recover target-prompt task geometry and produces corresponding gains in behavioral agreement. This suggests that cross-dimensional linear mixing may be a key contributor to how prompts reorganize representations toward the instructed task structure. Our framework provides a general way to decompose prompt-induced representational change into interpretable geometric components, revealing how a model routes task-relevant structure to produce prompt-driven behavior.
PaperID: 1096, Poster
Abstract: Predicting protein dynamics is a long-standing problem in computational structural biology. Often, protein function critically depends on local directed motions, such as hinge movements, catalytic loop rearrangements and domain reorientations, which can be characterized by directional flexibility and correlated structural motions of the protein backbone. While Molecular Dynamics (MD) simulations provide an established but often prohibitively expensive approach, recent deep generative models aim to reduce this cost by directly predicting conformational ensembles, emulating MD. However, due to their large size and the need to generate several states until the derived dynamical properties converge, these models remain expensive. In this work, we propose a fast SE(3)-equivariant graph neural network to directly predict dynamical descriptors, such as directional backbone flexibility and pairwise dynamic correlations, from an equilibrium structure. In a series of experiments, we show that our model matches the accuracy of substantially larger ensemble generation models while being orders of magnitude faster, and demonstrate that the proposed equivariant architecture is especially well-suited for capturing anisotropic motions in proteins.
PaperID: 1097, Poster
Authors: Yeping Wang, Shihao Yang
Abstract: Physics-Informed Neural Networks (PINNs) solve forward PDEs by minimizing residual losses from the governing equations with initial and boundary conditions, but they often struggle with discontinuities such as shocks. In contrast, finite volume methods (FVM) handle discontinuities by enforcing integral conservation, which admits weak solutions. Motivated by this, we propose a \emphCoupled Integral PINN (CI-PINN) that augments a standard PINN with an auxiliary network for integral potentials and coupled integral constraints. This improves robustness near shocks while avoiding meshing and the numerical flux integration/reconstruction used in classical schemes. We validate CI-PINN on forward benchmarks including Burgers, Buckley--Leverett, the Euler system, and the Shallow-Water equations.
Abstract: We study stochastic random search for nonconvex optimization with only noisy function evaluations. We propose Mi2P, a two-point variant of stochastic three-point methods, and analyze it under generalized (L_0,L_1)-smoothness. Two regimes are considered. When the mean objective f is (L_0,L_1)-smooth and component variance is bounded, Mi2P achieves \widetildeO(d^3\sigma_0^2/\varepsilon^6), matching the rate of \citetboucherouite2024mistp under a strictly weaker assumption (they require f to be L-smooth, i.e., L_1=0). When an L^1-average (L_0,L_1)-smoothness holds across \xi, Mi2P attains the optimal \widetildeO(d\sigma^2/\varepsilon^4) rate; the per-sample L-smoothness used by gradient-estimation methods such as RSGF \citepghadimi2013stochastic (i.e., each f_\xi is L-smooth, L_1 = 0 pointwise in \xi) is a special case of this condition. For finite sums with G-Lipschitz components, a translation-symmetric variance-reduction scheme yields \widetildeO(\min\d^4/3n^2/3/\varepsilon^8/3,\, dn/\varepsilon^2\) without storing snapshots. We also extend the framework to \delta-inexact comparison feedback (e.g., RLHF, A/B testing), giving an O(\sqrtd\delta) accuracy floor. Our analysis is presented for directional distributions with bounded support; the same rates hold for any rotation-invariant distribution with sub-Gaussian tails on the norm of s (including Gaussian) under a step-size restriction that does not affect the asymptotic complexity.
Abstract: In biological evolution, unconstrained mutation can lead to catastrophic outcomes: organisms may evolve enhanced capabilities while losing essential functions for survival. Nature's solution is developmental constraints, where core regulatory genes remain anchored while peripheral genes adapt freely. We observe that current self-evolution algorithms for large language models lack analogous constraints. They optimize purely for capability, implicitly assuming safety will be preserved. Our experiments reveal this assumption to be dangerously wrong: models can misevolve into powerful yet dangerous entities. Inspired by how Hox genes anchor body structure across 500 million years of evolution, we propose Circuit-Anchored Evolution (CAE). Using mechanistic interpretability, we identify a tiny safety circuit, comprising less than 2% of model features, that causally mediates safety behaviors. We anchor this circuit during evolution, constraining it within a small displacement bound while allowing the remaining features to evolve freely. This mirrors the biological principle of evolvability with constraint: preserving what is essential while adapting what is peripheral. Experiments across three model families and two evolution algorithms demonstrate that CAE achieves superior safety preservation with minimal capability loss, substantially outperforming explicit reward-based constraints in both effectiveness and efficiency. Just as developmental constraints prevent biological evolution from producing nonviable organisms, circuit anchoring prevents model evolution from producing capable but dangerous systems.
PaperID: 1100, Poster
Authors: Saumyaranjan Mohanty, Aravind Reddy, Konda Reddy Mopuri
Abstract: Audio-visual dataset distillation aims to synthesize a small dataset from a large paired audio-visual dataset, while nearly preserving the training utility for neural networks. Existing approaches such as DM and AVDD face two critical limitations: 1) they are memory-intensive, making them difficult to scale to complex architectures and higher samples-per-class (SPC) settings, and 2) they often fail to preserve strong cross-modal semantic correspondence. To tackle these issues, we propose X-AVDD, a cross-attention-based distillation method that explicitly captures inter-modal relationships during synthesis. By decoupling cross-modal alignment from joint optimization, X-AVDD is substantially more memory-efficient, reducing end-to-end synthesis cost by ~ 8× GPU-hours, while producing highly aligned synthetic data. Empirically, X-AVDD significantly outperforms other audio-visual distillation methods. On VGGSound, an audio-visual dataset of short YouTube clips, X-AVDD improves top-1 test accuracy from 8.2% → 24.0% at SPC=10 and 9.8% → 26.5% at SPC=20. Furthermore, X-AVDD improves cross-architecture generalization (e.g., 25.5% → 35.3% for a ViT target on AVE dataset at SPC=50) and successfully scales to complex architectures such as AudioCLIP, achieving 62.3% accuracy using only 5.65% of the full dataset, against the full dataset accuracy of 72.4% while completing in 9× lower training time compared to the full dataset.
Abstract: The Quality-Diversity (QD) optimization aims to discover a collection of high-performing solutions that simultaneously exhibit diverse behaviors within a user-defined behavior space. This paradigm has stimulated significant research interest and demonstrated practical utility in domains including robot control, creative design, and adversarial sample generation. A variety of QD algorithms with distinct design principles have been proposed in recent years. Instead of proposing a new QD algorithm, this work introduces a novel reformulation by casting the QD optimization as a multi-objective optimization (MOO) problem with a huge number of optimization objectives. By establishing this connection, we enable the direct adoption of well-established MOO methods, particularly set-based scalarization techniques, to solve QD problems through a collaborative search process. We further provide a theoretical analysis demonstrating that our approach inherits theoretical guarantees from MOO while providing desirable properties for the QD optimization. Experimental studies across several QD applications confirm that our method achieves performance competitive with state-of-the-art QD algorithms.
PaperID: 1102, Poster
Abstract: It is widely recognized that in many applications, strong predictive performance is not enough to warrant use of a learned model. Additionally, it is required to ascertain whether the model conforms to certain properties such as fairness, robustness or safety. Unfortunately, most learning approaches are not guaranteed to produce a model that adheres to these properties. Model repair addresses this through a post-hoc debugging loop. Approaches such as verification are used to check if the desired property holds for a model; if not they return a counterexample which is used to modify or repair the ML model. We propose a novel repair method for tree ensembles. Our approach is based on a MILP encoding where the objective is to modify the leaf values in the tree ensemble such that it fixes the counterexamples while at the same time preserves predictive performance. The use of constrained optimization over the model parameters is compelling as it is guaranteed to fix a given set of counterexamples. We implement and evaluate our method in two use-case scenarios, classification requiring fairness and action policies in MDPs requiring safety. Results show that our MILP repair of tree ensembles can be feasible computationally, and is often successful in improving the desired property while not degrading performance.
Abstract: Despite the importance of causal reasoning, training LLMs to perform causal reasoning remains underexplored. Existing data efforts mostly focus on benchmarking LLMs on specific aspects of causality, rendering them less suitable for training generalizable causal reasoners. To address this, we propose UniCo, a data generation framework that both (1) addresses 18 causal query types across Pearl's Causal Ladder and (2) translates symbolic samples into code and narrative domains to simulate real-world use cases where causal terms are not explicitly specified. To ensure data quality, UniCo grounds answers with exact causal inference and filter cases with reasoning shortcuts. Upon supervised-finetuning with 66.6K UniCo-generated instances, Qwen3-4B, Qwen3-8B and Olmo-3-7B-Instruct achieve an average of 22.9% improvements across all 18 in-distribution query types, and 8.1% over state-of-the-art causal data generation frameworks on 7 established causal benchmarks outside the training distribution. Furthermore, in real-world medical understanding, legal decision, and tabular reasoning, UniCo-trained models consistently display more faithful reasoning traces, outperforming the base models by an average of 20.2% in faithfulness metrics. These suggest that causality-centered training not only strengthens causal reasoning, but also equips LLMs with a causal mindset in general reasoning tasks.
PaperID: 1104, Poster
Abstract: Open World Object Detection (OWOD) requires a detector to recognize known categories, detect unknown objects, and incrementally learn new classes. Recent work such as OW-OVD adapts a pre-trained Vision-Language Object Detector (VLOD) to OWOD through fine-tuning. However, this approach demands costly backpropagation and often degrades the detector's own zero-shot performance on known categories, calling into question whether training is necessary. In this paper, we propose CAVE (Clustering in Attribute-Visual Embedding space), a training-free framework that adapts a VLOD to OWOD without any parameter updates. From a single forward pass, CAVE collects per-class visual and attribute statistics via von Mises-Fisher-based clustering, along with attribute co-occurrence patterns, without any gradient computation. At inference, CAVE fuses the original VLOD's text-based scores with MAS (Mixture-Aggregate Score) derived from the collected statistics for known-class detection, while unknown objects are identified by AOS (Attribute Objectness Score) that combines superclass responses with attribute co-occurrence patterns. Experiments on M-OWODB show that CAVE outperforms OW-OVD by +16.2 known-class mAP (Task 4) and +28.0 average unknown recall, while surpassing the zero-shot VLOD across all incremental tasks.
Abstract: We present a unified framework for quantum sensitivity sampling, extending the advantages of quantum computing to a broad class of classical approximation problems. Our unified framework provides a streamlined approach for constructing coresets and offers significant runtime improvements in applications such as clustering, regression, and low-rank approximation. Our contributions include: k-median and k-means clustering: For n points in d-dimensional Euclidean space, we give an algorithm that constructs an \epsilon-coreset in time \widetilde O(n^0.5dk^2.5~\mathrmpoly(\epsilon^-1)) for k-median and k-means clustering. Our approach achieves a better dependence on d and constructs smaller coresets when d\gg k that only consist of points in the dataset, compared to recent results of [Xue, Chen, Li and Jiang, ICML'23]. \ell_p regression: For \ell_p regression problems \min_x\in \mathbbR^d\\|Ax-b\\|_p where A\in \mathbbR^n× d and b\in \mathbbR^n, we construct an \epsilon-coreset of size \widetilde O_p(d^\max\\1, p/2\\\epsilon^-2) in time \widetilde O_p(n^0.5d^\max\\0.5, p/4\\+1(\epsilon^-3+d^0.5)), improving upon the prior best quantum sampling approach of [Apers and Gribling, QIP'24] for p\in (2, 22]. Low-rank approximation with Frobenius norm error: We introduce the first quantum sublinear-time algorithm for low-rank approximation that approximates the best rank-k solution to a matrix A\in \mathbbR^d× n and does not rely on data-dependent parameters, and runs in \widetilde O(n^0.5dk^0.5\epsilon^-1) time. Additionally, we present quantum sublinear algorithms for kernel low-rank approximation and tensor low-rank approximation, broadening the range of achievable sublinear time algorithms in randomized numerical linear algebra.
PaperID: 1106, Poster
Abstract: Hypergraph Neural Networks (HNNs), designed to model higher-order relations beyond pairwise interactions, typically rely either on fixed structural weights or on normalized attention. Fixed weights can propagate misleading information in heterophilic settings, while normalization makes attention coefficients competitive, in the sense that, in a message-passing framework, increasing the weight of one sender must reduce the weight of others in the same aggregation set. Our main contribution is the introduction of PANDA (Prior-guided AttentioNal Dual-path Architecture), a plugin for two-stage message passing HNNs applicable to both undirected and directed settings. The main path interpolates between learned receiver-normalized attention and structural incidence priors, while a non-competitive (not normalized) auxiliary path lets each sender contribute its own transformed information. We also show that, with fixed structural coefficients, PANDA recovers a certain complex-valued, convolution operator adopted in previous works. Across four HNNs spanning spatial and spectral methods in both directed and undirected settings, plugging PANDA in yields an average improvement of 5.18 percentage points across eight real-world benchmark datasets.
Authors: Mikalai Korbit, Mario Zanon
Abstract: In multiclass softmax cross-entropy, the full generalized Gauss-Newton (GGN) curvature couples all output logits through the softmax covariance, making second-order updates harder to scale as the number of classes grows. We show that the standard multiclass GGN can be decomposed exactly into a rank-one true-vs-rest term and a positive semidefinite within-competitor covariance term. Fast Gauss-Newton (FGN) retains the first term and drops the second, yielding a positive semidefinite under-approximation of the multiclass GGN that is exact for binary classification. The derivation uses an exact true-vs-rest scalar-margin representation of softmax cross-entropy: the loss and gradient are unchanged, and the approximation enters only at the curvature level. Exploiting the FGN curvature structure, the damped update can be written as an equivalent whitened row-space system with one row per mini-batch example. We solve this system matrix-free by conjugate gradient using Jacobian-vector and vector-Jacobian products of the scalar margin map. Targeted mechanism experiments and a fixed-feature head evaluation support the predictions from the decomposition: FGN stays closest to the full softmax GGN when competitor mass is concentrated or damping is large, and deviates as the dropped within-competitor covariance grows.
Authors:
Francisco Sumba Toral, Carles Balsells Rodas, Yingzhen LiAbstract: Standard flow matching scales well but typically relies on an unstructured source distribution, limiting its ability to learn interpretable latent structure. Latent-variable models, by contrast, capture structure but often sacrifice generative quality. We bridge this gap by proposing Structured Coupling for Flow Matching (SCFM), a cooperative framework that augments flow matching with structured latent representation learning. By introducing structured latent variables and exogenous noise into the source, SCFM jointly learns a structured prior (via latent variable modeling) and a continuous transport map (via flow matching). It uses a shared time-dependent recognition network for both latent variable model variational inference and intermediate-time flow velocity estimation. This yields a structurally informed yet unconditional, simulation-free flow model, where the latent variable model can also assist flow sampling. Empirically, SCFM facilitates unsupervised latent representation learning for clustering, disentanglement and downstream tasks, while remaining competitive with flow matching in sample quality, showing that meaningful structure can be learned without sacrificing generative fidelity.
Authors: Davi B Costa, Renato Vicente
Abstract: Fine-tuning large language models on narrow data with harmful content produces broadly misaligned behavior on unrelated prompts, a phenomenon known as emergent misalignment. We propose that emergent misalignment involves persona-model collapse: deterioration of the model's internal capacity to simulate, differentiate, and maintain coherent personas. We test this hypothesis behaviorally using two metrics: moral susceptibility and moral robustness; computed as the cross-persona and within-persona variability of models' Moral Foundations Questionnaire responses under persona role-play, respectively. We evaluate four frontier models (DeepSeek-V3.1, GPT-4.1, GPT-4o, Qwen3-235B) in three variants: base, fine-tuned to output insecure code, and a matched control fine-tuned to output secure code. Across the four models, insecure fine-tuning produces a 55% average spike in moral susceptibility, pushing all four insecure variants beyond the band observed across 13 frontier models benchmarked in prior work, with GPT-4o reaching more than twice its upper end. It also causes a 65% average drop in moral robustness (equivalently, a 304% surge in within-persona variability), while the secure control preserves susceptibility near the base and induces only a partial robustness loss. Complementing these metric shifts, insecure variants' unconditioned responses converge toward saturation near the scale ceiling, departing markedly from both base models' differentiated responses and those elicited when they role-play toxic personas. Taken together, these metrics provide a sensitive diagnostic for emergent misalignment and serve as behavioral evidence that it involves persona-model collapse.
PaperID: 1110, Poster
Authors: Harry Li, Thea Cao
Abstract: Filtered vector search in private and on-prem retrieval-augmented generation (RAG) systems must control tail risk under fixed latency, RAM, and SSD budgets, not merely pick a fast backend. Our measurements isolate a recurring ambiguity regime concentrated in the 5-15% failure band, where small shifts in vector-metadata alignment flip which familiar P1-P4 action is safe and create 3-6 P95 spread. We study that regime as authorization over fixed actions rather than as generic planning. The paper contributes a measured failure-band stress test, an auditable authorize/veto rule over fixed P1-P4 actions, and PACER as one concrete controller that combines recall-floor adjustment, calibrated tail intervals, and regime-local fallback. The measured failure band exposes a different runtime decision boundary: when the provisional winner is separated enough to trust and when uncertainty should refuse it. On Medium+EntDoc, evaluated on a dual-AMD EPYC 7543 server with 128GB RAM at matched recall in the 5-15% failure band, PACER reduces P95 by 46%\ and P99 by 45%\ relative to the strongest baseline while maintaining 0.94-0.98\ Recall@20. The gain concentrates where plan rankings are unstable and weakens when the authorization logic is stripped back, while feasibility, transfer, drift, and hierarchy-aligned slices remain supporting diagnostics. For recurring fixed-budget filtered retrieval, these results show that an auditable authorize/veto layer over fixed actions can materially stabilize tail latency in the measured failure band.
PaperID: 1111, Poster
Abstract: Neural Posterior Estimation (NPE) trains an amortized conditional density estimator to approximate posterior distributions in simulation-based inference. Each training pair is a single draw from the joint distribution of latent and observation, so posterior variance at an observation must be inferred from variation across pairs rather than within them. However, with a sufficiently flexible variational distribution, the NPE training objective is unbounded, and standard optimizers yield arbitrarily small variance. Regularizers like early stopping, dropout, and weight decay constrain network capacity without targeting the correct posterior variance. We propose Cross-Fitting NPE (CF-NPE), which separates mean estimation from higher-moment estimation via sample splitting. Concretely, CF-NPE estimates a conditional-mean ensemble by K-fold cross-fitting, then fits a conditional density to the resulting out-of-fold residuals. The conditional variance of an out-of-fold residual is bounded below by the true posterior variance, giving CF-NPE a well-defined population-level target that standard NPE lacks. Across eight synthetic benchmarks and a cosmological application, CF-NPE achieves higher held-out log-likelihood than standard NPE, with gains that grow with the latent dimension.
PaperID: 1112, Poster
Abstract: Test-time scaling often uses an external verifier, such as compilers and test-cases in coding or trained value functions in robotics application, to obtain high quality rollouts. Verifier-free test-time scaling (or VF-TTS) is gaining extensive attention as a mechanism to enhance Large Language Model (LLM) reasoning, primarily because in many real-world applications we do not have access to such high quality verifiers. Among existing VF-TTS methods, confidence-based VF-TTS methods, which compute and rank rollouts solely by confidence, are particularly promising. Such methods introduce near-zero overhead for sample evaluation and require minimal access to internal model states, making the methods highly flexible across models and data. In this paper, we demonstrate a critical limitation of existing confidence-based VF-TTS methods by showing that such methods catastrophically break down on complex tasks. We observe a very interesting phenomenon: uniformly high confidence frequently indicates a failure to explore, favoring confidently wrong answers. To address this, our core insight is that robust cognitive search requires a specific confidence trajectory pattern: such methods perform exploratory branching at the beginning as manifested by low initial confidence, and converge to a high final confidence solution. To implement this insight, we introduce consilience, a novel selection framework that explicitly evaluates the temporal asymmetry of confidence in reasoning. We operationalize this via a combinatorial metric that actively penalizes high initial confidence while strictly demanding final certainty. Extensive experiments covering both graduate-level mathematics problems and free-form code generation demonstrate that consilience effectively outperforms existing baselines, validating our novel perspective on completion confidence.
Authors: Gefan Yang, Frank van der Meulen, Stefan Sommer
Abstract: Inference in nonlinear continuous stochastic processes on trees is challenging, particularly when observations are sparse and the topology is complex. Exact smoothing via Doob's h-transform is intractable for general nonlinear dynamics. We propose Neural Backward Filtering Forward Guiding (NBFFG), a unified framework for both discrete transitions and continuous diffusions. Our method constructs a variational posterior by leveraging a proxy linear-Gaussian process. This proxy process yields a closed-form backward filter that serves as a guide, steering the generative path toward high-likelihood regions. We then learn a neural residual to capture the non-linear discrepancies. This formulation allows for an unbiased pathwise subsampling scheme, reducing the training complexity from tree-size dependent to path-length dependent. Empirical results show that NBFFG outperforms baselines on synthetic benchmarks, and we demonstrate the method on a high-dimensional inference task in phylogenetic analysis with reconstruction of ancestral butterfly wing shapes.
PaperID: 1114, Poster
Abstract: Dense optical flow estimates correspondence from every pixel to the next frame. It must be precise and robust, which couples two choices: what units to match, and what evidence to match them. Recent work has improved the evidence with stronger features, cost volumes, attention, refinement, and generative priors. The unit, however, usually remains fixed: motion is still estimated at dense grid points. We propose Adaptive Correspondence Tokens (ACT), which adapts correspondence units through a differentiable pixel-to-token assignment. ACT learns soft maps that group pixels into tokens, build descriptors, match and refine motion in token space, and project token motion back to pixels for dense flow. Trained end to end with the flow loss, grouping and correspondence are discovered jointly: grouping makes matching easier, while matching shapes grouping. Support can expand over coherent regions and contract near boundaries or fine detail. In zero-shot cross-dataset evaluation, ACT achieves the best KITTI EPE, second-best KITTI Fl, and second-best Sintel Final EPE among published methods, while remaining competitive under benchmark fine-tuning. These results show that adaptive units make dense flow easier: matching where the image gives reliable evidence improves transfer while preserving dense, boundary-aware prediction.
PaperID: 1115, Poster
Abstract: Guidance methods are essential tools for steering the generation process in diffusion models, improving the generalizability of stochastic robotic controllers and enabling high-quality controlled image synthesis. Despite strong parallels between flow-matching and diffusion, integrating guidance into flow-matching models remains quite challenging. Recent work has proposed general guidance frameworks for flow-matching, however, they are direct adaptations of diffusion guidance methods and can prove unstable. In this work, we conduct a theoretical analysis of guidance in flow-matching and show that a naive translation of diffusion-based guidance to flow-matching breaks the coupling distribution used during training, resulting in weaker control over sampling accuracy. Building on this insight, we propose , a flow-matching-specific guidance mechanism derived from a geometric perspective that leverages the coupling structure inherent to flow-matching training. To validate the effectiveness of our approach, we conduct extensive experimentation on a custom robotic , using the ImageNet-128 and CelebA-HQ-256 datasets. Our findings show that our framework improves sample quality and guidance stability across all tasks, without imposing any additional overhead.
Abstract: The k-center problem is a fundamental clustering variant with applications in learning systems and data summarization. In several real-world scenarios, the dataset to be clustered is not static, but evolves over time, as new data points arrive and old ones become stale. To account for dynamicity, the k-center problem has been mainly studied under the sliding window setting, where only the N most recent points are considered non-stale, or the fully dynamic setting, where arbitrary sequences of point arrivals and deletions without prior notice may occur. In this paper, we introduce the dynamic setting with lifetimes, which bridges the two aforementioned classical settings by still allowing arbitrary arrivals and deletions, but making the deletion time of each point known upon its arrival. Under this new setting, we devise a deterministic (2+\varepsilon)-approximation algorithm with \widetildeO(k/\varepsilon) amortized update time and memory usage linear in the number of currently active points. Moreover, we develop a deterministic (6+\varepsilon)-approximation algorithm that, under ‘‘tame” update sequences, has \widetildeO(k/\varepsilon) worst-case update time and heavily sublinear working memory.
PaperID: 1117, Poster
Abstract: We consider a variant of the partially observed online control problem where there are _multiple sensors_ which provide noisy linear measurements of the system. In our setting, the system state evolves according to a discrete-time linear dynamical system. The objective is to design an algorithm which minimizes regret relative to the best _single-sensor_ policy (within a restricted class) in hindsight. When the system and process noise is stochastic and the costs are quadratic in the control and _unobserved state_, we design an algorithm that achieves \mathcalO(\mathrmpoly\log(S)\sqrtn) regret over time horizon n relative to our benchmark policy class if there are S sensors. Notably, this regret guarantee is with respect to the _state costs_, which are never fully observed. When the system and process noise is chosen by an oblivious adversary, and the cost functions are convex-Lipschitz (or subquadratic) functions of the control and _observations_, then a similar \mathcalO(\mathrmpoly\log(S)\sqrtn) regret is achievable. While our work is the first, to our knowledge, to study this generalization of the online control problem, we remark that a na\"ive application of known results in the partially observed online control literature have regret scaling \mathcalO(\mathrmpoly(S)\sqrtn), exponentially worse than our regret scaling in terms of dependence on the number of sensors. Our algorithm is based on the Follow-The-Regularized-Leader framework, and our analysis carefully exploits the \ell_1 geometry of the policy class.
PaperID: 1118, Poster
Abstract: Domain generalization methods train on multiple source environments but usually deploy a single pooled predictor. We study a different deployment object: one head per source environment on a shared representation, combined by fixed weights when the test environment label is unavailable. Under a same-meta random-effects model, where environment-specific population minimizers are iid deviations around a shared meta-mean, the optimal fixed simplex rule is the uniform centroid. The same analysis separates this centroid from pooled ERM: under a common feature second moment, pooled ERM is sample-weighted, while the centroid is environment-weighted, giving a closed-form gap controlled by sample-weight imbalance. We introduce (Random-Effects Centroid), a training objective that fits source-specific experts while directly scoring their averaged deployment head. In the common-second-moment case, the centroid term controls the trained-centroid discrepancy and the dispersion of environment-specific optima. Synthetic experiments isolate the population estimator effect and learned-representation behavior; WILDS experiments show that REC is competitive with established DG baselines and that the centroid term is necessary for a useful averaged predictor.
Authors: Maria-Florina Balcan, Tejas Pagare, Karan Singh
Abstract: Motivated by modern marketplaces, where the platform or the seller routinely gathers detailed user profiles, we study a novel learning theoretic model that simultaneously involves information and mechanism design. Specifically, we consider the economic setting recently introduced by Bergemann et al, where in addition to the menu of quality-price pairs, the seller offers information on the value of the match between product quality and buyer's taste via a signaling scheme. We relax the assumption that the seller knows the buyers' belief about the distribution of tastes and study the sample requirements of designing a revenue maximizing scheme. We consider both the batch setting where we have access to data from a set of i.i.d. buyers and an online demand query model where we observe the buyers behaviors to seller's schemes. Despite the apparent non-convexity of the problem, we also give the first FPTAS to compute a scheme that maximizes the revenue within an arbitrarily small additive loss, which was unknown even in the previous result of Bergemann et al. Overall, this brings a new learning perspective in asymmetric economic settings where buyers and sellers know different types of information.
PaperID: 1120, Poster
Abstract: We propose a nonparametric, kernel-based approach to generative modeling based on Stein operators and particle transport. Our method constructs a continuous-time transport from a simple reference distribution (typically, a standard Gaussian) to a target distribution using a Stein formulation of the continuity equation. Unlike alternative Stein transport methods, which require access to the target density or its gradient, we focus on the sample-only setting and introduce a tractable surrogate objective based on an Ornstein--Uhlenbeck (OU) probability path. This construction allows the velocity field driving the transport to be learned using only samples from the target distribution and Gaussian noise. By restricting the velocity field to a reproducing kernel Hilbert space, we obtain closed-form solutions via kernel ridge regression at each time step. The resulting method is particularly well suited to small data regimes, where heavily parameterised neural generative models may be difficult to train or overfit. We demonstrate the usefulness of the approach for sample-based generative modeling in the small data regime (generative data augmentation), and for posterior emulation in Bayesian inverse problems, where expensive Monte Carlo samplers produce limited sets of posterior samples that must be efficiently augmented.
PaperID: 1121, Poster
Abstract: We introduce a framework that models constrained MDPs subject to trajectory-wise safety constraints. These constraints require that expected future costs remain below zero conditioned to any possible trajectory---not merely in expectation from the initial state. This captures stronger and more realistic safety guarantees, which are crucial in high-stakes applications. We characterize the structure of optimal policies, and show that Markovian policies are in general suboptimal. Moreover, we show that computing approximately-safe optimal policies is NP-hard when the number of constraints can be arbitrarily large. Despite all these challenges, we provide an exact algorithmic characterization of optimal history-dependent (non-Markovian) policies, by exploiting a suitable recursive formulation of the sets of reachable reward-cost values. Then, we employ it to design a polynomial-time approximation algorithm that computes approximately-safe optimal policies by dynamic programming, when assuming a constant number of constraints.
PaperID: 1122, Poster
Authors: Pierre Quinton, Valérian Rey
Abstract: Many optimization problems require balancing multiple conflicting objectives. To generalize gradient descent, which is limited to single-objective optimization, we introduce Jacobian descent (JD) and its stochastic variants. This algorithm iteratively updates parameters using the Jacobian matrix of a vector-valued objective function, in which each row is the gradient of an individual objective. While several methods to combine gradients already exist in the literature, they generally struggle when the objectives conflict. In contrast, we propose projecting gradients to fully resolve conflict. We prove convergence to the Pareto front in the smooth convex case with this approach, a stronger guarantee than the usual convergence to a weakly Pareto stationary point. Our method also enables instance-wise risk minimization (IWRM), a novel learning paradigm in which the loss of each training example is considered a separate objective. Early experimentation on image classification shows promising results for IWRM. Lastly, we provide an efficient implementation of JD using the Gramian of the Jacobian matrix to drastically reduce memory requirements.
Abstract: Existing robust preference optimization for language-model alignment mainly studies pairwise supervision and places robustness at the dataset, prompt, or preference-pair level. We instead study listwise preference optimization under ranking-label uncertainty: given a prompt and a candidate list, the observed ranking over that list may be ambiguous due to annotator inconsistency, near-ties, lossy rankwise feedback, or reward-model noise. We propose a pointwise total-variation robust Plackett--Luce objective that directly robustifies the ranking label conditional on the candidate list. The robust loss admits an exact decomposition into the nominal PL loss plus a worst-case PL correction, and the worst-case ranking is obtained by sorting current implicit scores in ascending order, reducing the inner maximization from K! enumeration to O(K\log K). This tractable structure yields strong offline and online optimization guarantees. In the offline fixed-list setting, the robust objective is convex and projected stochastic subgradient reaches global \epsilon-suboptimality with O(\epsilon^-2) sample complexity. In the online policy-induced setting, where candidate lists are generated by the current policy, we establish weak convexity and \widetilde O(\epsilon^-2) Moreau-envelope stationarity. Experiments in offline LLM alignment show that the proposed robust correction largely preserves performance under clean labels and improves robustness under noise. In online alignment, it makes reward-model-ranked candidate expansion more reliable and improves both reward-model andexternal GPT-4 judge metrics.
PaperID: 1124, Poster
Abstract: Contrastive representation learning (CRL) is a fundamental paradigm for learning transferable representations. Recent work identifies its population target as the pointwise mutual information (PMI) \log\fracp(\mathbfx,\mathbfy)p_\mathcalX(\mathbfx)p_\mathcalY(\mathbfy) between two modalities \mathcalX,\mathcalY, where p is the joint density function and p_\mathcalX,p_\mathcalY are the marginals. However, it remains unclear why standard inner-product scores f(\mathbfx)^\top g(\mathbfy) can effectively approximate a generally nonseparable population target PMI. In this paper, we address this question from the aspect of approximation for CRL. By linking PMI to a compact operator, we show that D-dimensional inner-product models perform rank-D spectral approximation, with error controlled by the spectral tail. We then study neural network realizations: in general settings, we show that deep neural network scoring functions are universally consistent, and the approximation error rates depend on the smoothness of the PMI function; under structured low-rank data models we obtain improved rates attained by moderate-width encoders. Finally, we establish a fast calibration rate from contrastive excess risk to downstream retrieval error under a density-ratio Tsybakov noise condition.
PaperID: 1125, Poster
Authors: Xiao Wang, Zihua She, Jianxi Su
Abstract: We propose Coreset-Induced Conditional Velocity Flow Matching (CCVFM), a generative model that augments hierarchical rectified flow with a data-informed source distribution. Hierarchical flow matching models the full conditional velocity law in velocity space, but its inner flow is asked to transport isotropic Gaussian noise to a multimodal target velocity distribution from scratch. Our key observation is that this inner source can be replaced by a closed-form surrogate built from a coreset of the target. CCVFM first compresses the target into weighted atoms using an entropic Sinkhorn coreset and lifts them to a Gaussian mixture. The induced conditional velocity law is then a closed-form Gaussian mixture that can be sampled without a learned neural sampler. A lightweight correction flow, trained from this exact surrogate source, then refines the remaining surrogate-to-target residual rather than learning an entire noise-to-data map. We prove that the surrogate transport cost equals the target--surrogate Wasserstein gap under an explicit compression assumption, whereas the noise-source analogue has a dimension-scale lower bound. We further characterize the conditional second moment of the direct surrogate-source training target and show that its source-dependent excess is small when the surrogate conditional law is close to the true conditional velocity law in mean and covariance. Empirically, on MNIST, CIFAR-10, ImageNet-32, and CelebA-HQ, the proposed method reaches competitive few-step generation under matched architectures. Codes are available at \urlhttps://anonymous.4open.science/r/ccvfm-code-11D3/.
Authors: Javier Fumanal Idocin, Raquel Fernandez-Peralta, Javier Andreu-Perez
Abstract: Dynamic feature selection (DFS) is a machine learning framework in which features are acquired sequentially for individual samples under budget constraints. The exponential growth in the number of possible feature acquisition paths forces a DFS model to balance fitting specific scenarios against maintaining general performance, even when the feature space is moderate in size. In this paper, we study the structural limitations of existing DFS approaches to achieve an optimal solution. Then, we propose \textscHyper-DFS, a hypernetwork-based DFS approach that generates feature subset-specific classifier parameters on demand. We show that the use of hypernetworks compared to mask-embedding methods results in a smaller structural complexity bound. We also use a Set Transformer encoding to create a smooth conditioning space for the hypernetwork, so that functionally similar tasks are also geometrically close. In our benchmarks, \textscHyper-DFS outperforms all state-of-the-art approaches on synthetic and real-life tabular data. It is also competitive or superior across all image datasets tested, and shows substantially stronger zero-shot generalisation to feature subsets never seen during training than existing DFS approaches.
PaperID: 1127, Poster
Abstract: We study asymptotic anytime-valid confidence sequences for degree-two U-statistics under continuous monitoring. In the nondegenerate case, Hoeffding's projection reduces the problem to a time-uniform central limit theory for the partial sums of the first-order projection, while the canonical remainder is shown to be negligible under mild moment assumptions. A leave-one-out jackknife estimator then yields a fully data-driven procedure, leading to confidence sequences with asymptotic coverage guarantee for the parameter of interest. In the degenerate case, we show that the U-statistic is approximated by a centered quadratic Gaussian-chaos rather than by a simple Gaussian, which poses significant challenges for sequential inference. To address this issue, we novelly develop the Spectrally Allocated Gaussian-chaos Excursion (SAGE) boundary, and then provide plug-in implementations based on truncated spectrum estimation with consistency guarantees. The resulting widths can attain the expected time-uniform optimal rates: \sqrt\log\log n/n in the nondegenerate regime and \log\log n/n in the degenerate regime. Several widely used U-statistics are discussed within the proposed framework, and numerical experiments further support the validity of the derived theory.
PaperID: 1128, Poster
Abstract: Autoregressive world models have emerged as a powerful paradigm for sample-efficient decision-making, yet their open-loop rollouts often exhibit increasing drift that leads to error accumulation even within a relatively small planning horizon. Hierarchical temporal abstraction is a natural candidate to mitigate such drift by reducing the effective prediction horizon; however, existing hierarchical designs can introduce added structural and conceptual complexity, sometimes making them difficult to control and limiting gains over flat baselines. To this end, we introduce MHWA (Multi-timescale Hierarchical World Action model), which employs a strided high-level context channel to reduce global-context update frequency, effectively curbing the propagation of compounding drift. Crucially, via context-aware routing, the gating network dynamically reconfigures a Low-Level Mixture-of-Experts by using both global context and current observations, thereby enabling flexible high--low coordination, helping alleviate interference commonly observed in static hierarchies. Empirically, MHWA (77M parameters) achieves an IQM human normalized score (HNS) of 0.867\pm0.047 across three training seeds on a 14-game Atari subset using 10% subsampled offline data, matching the existing state-of-the-art 150M-parameter JOWA baseline while using nearly half the parameters and significantly reducing computational costs.
Abstract: We revisit the I/O complexity of attention in large language models. Given query-key-value matrices Q,K,V\in\mathbbR^n× d, and a machine with fast memory size M, the goal is to compute the "attention matrix" A=\mathrmsoftmax(QK^\top/\sqrtd)V with the minimal number of data transfers between fast and slow memory. Existing methods in the literature, most notably FlashAttention and its variants, incur an I/O cost that depends quadratically on n, while a trivial lower bound only requires \Omega(nd) I/O's to read the inputs and write the output. In this work, we present a technique for computing attention where the I/O cost only depends almost-linearly on n in most parameter regimes. This is achieved by developing I/O-efficient algorithms inspired by the recent approximate attention framework of Alman and Song [NeurIPS'23]. We also prove corresponding lower bounds in each parameter regime to show that our algorithms are indeed close to I/O-optimal.
Abstract: (SAD) treats detection as a two-sample test. Given a reference set of clean examples (CEs) and a batch of queries, potentially containing an unknown mixture of CEs and adversarial examples (AEs), SAD decides whether the query distribution drifts away from the CE distribution while controlling the false-alarm rate. Existing SAD-based methods mainly use (MMD) to measure the distributional discrepancy. However, MMD's distributional properties limit its ability to capture characteristic uncertainty patterns of AEs that are crucial for detection: AEs typically exhibit abnormal feature spread (i.e., global uncertainty) and instability under perturbations (i.e., local uncertainty). To close the gap, we propose iscrepancy (VD), which measures the difference in feature spread between AEs and CEs to capture global uncertainty differences, and (2) iscrepancy (PCD), which compares feature covariance under Gaussian perturbations to capture local uncertainty differences. By aggregating VD and PCD, USAD achieves superior detection performances over baseline methods against various adversarial attacks, highlighting the importance of considering characteristic behaviors of AEs for effective SAD. Our code is available at: https://anonymous.4open.science/r/USAD.
Abstract: When algorithmic predictors inform resource allocation in high-stakes domains such as healthcare, these predictors must account for strategic manipulation of input features. The typical solution is to redesign the predictor itself to explicitly account for strategic interactions. In practice, however, decision makers are often constrained to adjusting coarser levers within existing prediction pipelines. For example, healthcare organizations often select which features to exclude based on perceived manipulability, while using standard regularization procedures to shrink the coefficients of retained features. In this work, we initiate a formal study of strategic classification through . Our main finding is that excluding features based on manipulability alone is generally suboptimal. We provide a fine-grained characterization of the performance of a feature subset under optimal regularization, yielding new insights for policy design. Motivated by this characterization, we develop a practical algorithm for jointly choosing the feature set and the level of ridge regularization. Through a real-world case study on a healthcare payments benchmark, we illustrate how our algorithm can guide the design of coarse policy levers in practice. Our results provide a principled, practical framework for mitigating the effects of strategic behavior in algorithmic decision-making systems.
PaperID: 1132, Poster
Abstract: Machine unlearning (MU) aims to remove sensitive or undesired knowledge from a trained model without retraining from scratch. While MU has been widely studied for autoregressive large language models, its application to diffusion large language models (DLLMs) remains largely unexplored. In this paper, we present the first systematic study of MU in DLLMs and show that existing unlearning objectives transfer poorly to this setting due to sparse supervision and inaccurate forgetting localization. To address these challenges, we propose \textscDL-Eraser, a DLLM-native unlearning framework that suppresses target recoverability under high-mask conditioning while constraining updates to a utility-preserving subspace. Extensive experiments show that \textscDL-Eraser achieves stronger forgetting while better preserving model utility than existing baselines. To the best of our knowledge, this is the first work that systematically investigates MU in DLLMs.
PaperID: 1133, Poster
Authors: Phuc Lai, Anh Nguyen, Phong H Nguyen, Anh Tran
Abstract: Megapixel visual synthesis with latent diffusion models requires operating far beyond the resolutions at which these models are typically trained. Cascade refinement is a leading strategy for this setting: it starts from a base image at a native resolution and progressively upsamples it to the target resolution through multiple refinement stages. Each stage typically performs a pixel-space resolution transition followed by latent-space denoising. While prior acceleration efforts mainly focus on reducing denoising cost, we identify a complementary and underexplored bottleneck: the repeated movement of representations through the VAE. We call these operations . In cascade pipelines, intermediate latents are often decoded to RGB, resized, and re-encoded before refinement, while final latents are decoded patch by patch with dense overlap to suppress boundary artifacts. We propose , a transition-efficient approach that accelerates these VAE-heavy operations. UltraFlash has two components. replaces intermediate RGB round trips with a learned latent-space shortcut that approximates the decode-resize-encode transformation without repeated VAE decoding and encoding. accelerates final reconstruction by addressing patch seams caused by inconsistent GroupNorm statistics and reusing cached statistics across patches, enabling sparse VAE decoding with much smaller overlap. Integrated into a state-of-the-art cascade baseline, UltraFlash reduces 4K generation latency by approximately 2x while improving image quality. These results show that optimizing pixel-space transitions can improve both efficiency and fidelity, offering an effective route toward fast, high-quality megapixel visual synthesis.
Abstract: The deployment of Artificial Intelligence in high-risk domains, such as finance and healthcare, necessitates models that are both fair and transparent. While regulatory frameworks, including the EU's AI Act, mandate bias mitigation, they are deliberately vague about the definition of bias. In line with existing research, we argue that true fairness requires addressing bias at the intersections of protected groups. We propose a unified framework that leverages Mixed-Integer Optimization (MIO) to train intersectionally fair and intrinsically interpretable classifiers. We prove the equivalence of two measures of intersectional fairness (MSD and SPSF) in detecting the most unfair subgroup and empirically demonstrate that our MIO-based algorithm improves performance in finding bias. We train high-performing, interpretable classifiers that bound intersectional bias below an acceptable threshold, offering a robust solution for regulated industries and beyond.
PaperID: 1135, Poster
Abstract: architectures, yet many rely on linear compatibility scores for latent tokenization. Under common feature normalization, such scores are equivalent to isotropic Euclidean clustering, while without normalization they induce unbounded linear decision regions. In both cases, they lack slice-specific anisotropic locality, which can lead to redundant and entangled latent slices. To address this, we propose , a geometry-aware routing mechanism that replaces linear projection with a learnable Mahalanobis metric. By enabling each latent slice to learn a local anisotropic tensor, METRO shapes receptive fields into exponentially localized, oriented ellipsoids that naturally align with flow features like boundary layers and wakes. As a drop-in replacement, METRO yields consistent improvements across both Transformer and Mamba backbones. Empirically, our method achieves substantial performance gains on irregular domains, outperforming baselines on both standard PDE benchmarks and complex industrial design tasks. Finally, METRO exhibits enhanced robustness in out-of-distribution regimes across varying Reynolds numbers and geometric configurations.
Abstract: We study multi-armed bandits (MABs) augmented with best-action queries, in which the learner may additionally query an oracle that reveals the best arm in the current round. This setting was recently characterized by Russo et al. [2024] in the full-feedback model, where the learner observes the rewards of all arms after each round. They show that, in both stochastic and adversarial environments, k best-action queries reduce the optimal \widetilde\mathcalO(\sqrtT) regret to \widetilde\mathcalO(\min\T/k,\sqrtT\). Whether this improvement extends to the more realistic bandit-feedback model---where the learner observes only the reward of the played arm---was left as an open problem. We fully resolve this question. When rewards are stochastic but correlated among arms, we show that the full-feedback result does not carry over: any algorithm must incur regret at least \Omega(\sqrtT-k). This lower bound directly extends to adversarial environments. On the positive side, we show that \widetilde\mathcalO(\min\T/k,\sqrtT-k\) regret is still achievable when rewards are stochastic and i.i.d., and establish a matching lower bound, up to logarithmic factors. Together, these results provide a complete characterization of the benefits of best-action queries in the bandit-feedback model.
PaperID: 1137, Poster
Abstract: Delayed hits occur when requests arrive while a previous miss for the same item is still being fetched. They break the equivalence between hit ratio and latency because caching decisions depend on frequency, burst structure, and fetch latency. We develop a regret framework for delayed-hit caching, with static regret against long-term placement and switching regret against piecewise-changing workloads. We show that classical deterministic policies and randomized Marker fail to achieve no-regret guarantees, both suffering linear static regret. For stochastic arrivals, we introduce delayed-hit shadow values and propose UCB-DH, which achieves O(\log T) static regret on fixed instances with positive gaps, where T is the horizon. A matching gap-dependent lower bound shows that the delayed-hit latency-amplification factor is statistically unavoidable, and restarted UCB-DH gives sublinear regret for piecewise-changing workloads. For oblivious adversarial arrivals, we formulate an action-independent shadow loss and propose LA-FTPL with latency-scaled perturbations and completion-time admission. LA-FTPL achieves O(\sqrtKT\log N\,(Z^2+C_BZ)) shadow static regret, where K is the cache capacity, N the item universe size, Z the fetch latency, and C_B the maximum one-window delayed-hit penalty. With periodic restarts, it achieves O(ZT^2/3B^1/3(K\log N)^1/3) switching regret against comparators with at most B cache changes. Experiments on synthetic and real-world traces show that our algorithms achieve sublinear regret and improve latency over the state-of-the-art baselines.
PaperID: 1138, Poster
Authors: Valerio Di Pasquale, Alessia Antelmi, Mirko Polato, Carmine Spagnuolo
Abstract: Hypergraphs capture high-order interactions that ordinary graphs cannot represent, yet generating realistic ones remains challenging because existing methods often rely on fixed structural assumptions or struggle to model coupled dependencies between nodes and hyperedges. In this work, we propose \textscJanus, a latent diffusion framework that decomposes an observed hypergraph into sub-hypergraphs, learns coupled node and hyperedge latent representations with a dual-view \beta-VAE, and generates new structures through paired latent diffusion with cross-view conditioning. The method supports both unconstrained and node-set-constrained generation. We also introduce an evaluation protocol covering micro, meso, and macro scales interactions, structural reconstruction, and high-order similarity measures. Across five real-world datasets and nine baselines, \textscJanus achieves the most stable performance across structural scales, with strong gains in meso-scale and high-order fidelity; in the constrained setting, it consistently obtains the best reconstruction results.
Authors:
Shao-Ting Chiu, Aditya Nambiar, Ali Syed, Jonathan W. Siegel, Ulisses M. Braga-NetoAbstract: An important application of neural networks to scientific computing has been the learning of non-linear operators. In this framework, a neural network is trained to fit a non-linear map between two infinite dimensional spaces, for example, the solution operator of ordinary and partial differential equations. Recently, inspired by the discovery of in-context learning for large language models, an even more ambitious paradigm has been explored, called multi-operator learning. In this approach, a neural network is trained to learn many different operators at the same time. In order to evaluate one of the learned operators, the network is passed example inputs and outputs to disambiguate the desired operator. In this work, we provide a precise mathematical formulation of the multi-operator learning problem. In addition, we modify a simple efficient architecture, called DeepOSets, for multi-operator learning and prove its universality for multi-operator learning. Finally, we provide experiments showing the efficacy of DeepOSets for learning multiple operators corresponding to different initial-value and boundary-value differential equation problems.
Abstract: Higher-order learning exploits compositional features whose predictive value arises from nonlinear interactions beyond pairwise relations among variables, scales, or modalities. To address the complexity of modeling higher-order interactions in modern large-scale deep learning models, a kernelized Volterra Neural Network (kVNN) is proposed in this paper. Specifically, the proposed learnable multi-kernel representation models different interaction orders using distinct polynomial-kernel components with compact learnable centers, thereby yielding an order-adaptive parameterization. The resulting kVNN layers are composed of parallel order-specific branches and can readily replace standard convolutional kernels in existing learning architectures. Theoretical results are supported by experiments on video action recognition, image denoising, and image classification. The results show marked performance-efficiency trade-offs: kVNN consistently reduces model (parameters) and computational (GFLOPs) complexity while achieving competitive and often improved performance, even when trained from scratch without large-scale pretraining. In summary, structured kernelized higher-order layers offer a practical path to balancing expressivity and computational cost in modern deep networks.
PaperID: 1141, Poster
Abstract: Shampoo is a Kronecker-structured second-order optimizer widely used for large-scale deep learning, yet existing regret analyses exhibit an unfavorable dependence on the gradient rank, leading to guarantees that are uniformly weaker than those of optimizers such as Adagrad and One-sided Shampoo. We provide a uniformly and substantially sharper regret bound for the full two-sided Shampoo update, yielding bounds that can compete with full-matrix Adagrad and more recent structured variants such as One-sided Shampoo. By further analysis, we also establish the first sublinear convergence guarantee for the commonly implemented variant Shampoo^2 by introducing a principled scaling factor. Our analysis holds for a continuum of exponent choices p \in (0, \frac12], also providing the first sublinear guarantees for Shampoo^4p with exponents other than p=1/4.
Abstract: Predictive models are often deployed through existing decision policies that stakeholders are reluctant to change unless a risk constraint requires intervention. We study risk-controlled post-processing: given a deterministic baseline policy, choose a new policy that maximizes agreement with the baseline subject to a chance constraint on a user-specified loss. At the population level, we show that the optimal policy has a threshold structure: it follows the baseline except on contexts where switching to the oracle fallback policy yields a large reduction in conditional violation risk. At the finite-sample level, given a fitted fallback policy and score, we develop a post-processing algorithm that uses calibration data to select a threshold. Leveraging tools from algorithmic stability and stochastic processes, we show that under regularity conditions, in the i.i.d. setting, the expected excess risk of the post-processed policy is O(\log n/n). In the special case when an exact-safe fallback policy is available, the algorithm achieves precise expected risk control under exchangeability. In this setting, we also give high-probability near-optimality guarantees on the post-processed policy. Experiments on a COVID-19 radiograph diagnosis task, an LLM routing problem, and a synthetic multiclass decision task show that targeted post-processing can meet or nearly meet risk budgets while preserving substantially more agreement with the baseline than score-blind random mixing.
Abstract: As large language models (LLMs) are deployed in consequential settings such as medical question answering and legal reasoning, the ability to estimate when their outputs are likely to be correct is essential for safe and reliable use, requiring well-calibrated uncertainty. Standard reinforcement learning with verifiable rewards (RLVR) trains models with a binary correctness reward that is indifferent to confidence, providing no penalty for confident but wrong predictions and thereby degrading calibration. Recent work addresses this by training models to produce verbalized confidence scores alongside answers and rewarding agreement with correctness. However, verbalized confidence is calibrated at the token level and thus exhibits inconsistency across textual variations with same semantic meaning. We propose , a framework that calibrates language models directly in semantic space without a verbalized confidence interface. CSR combines the correctness reward with a novel semantic calibration reward that encourages exploitation among correct rollouts by promoting semantic agreement, and exploration among incorrect ones by discouraging spurious consistency. Experiments across three model families on HotpotQA (in-distribution) and TriviaQA, MSMARCO, and NQ-Open (out-of-distribution) show that CSR consistently achieves lower ECE and higher AUROC than verbalized-confidence baselines across nearly all settings, reducing ECE by up to 40% and improving AUROC by up to 31% over verbalized-confidence baselines, with calibration behavior generalizing robustly across all four evaluation settings. The code is available at: https://anonymous.4open.science/r/CSR-34F4.
PaperID: 1144, Poster
Authors:
Jie Yang, Hsiang-Ting Chen, Yan Ma, Xinyan Liang, Avinash Singh, Liang Du, Cheng-You Lu, Chenglong Zhang, Bingbing Jiang, Weiping Ding, Wei ChenAbstract: Density Peak Clustering (DPC) focuses on identifying centers and propagates labels at the sample level, which makes density estimation cutoff-sensitive and assignment chain-vulnerable, yielding suboptimal performance under heterogeneous densities and complex cluster structures. This paper argues that density peaks inference should instead be performed on resolution-controlled structural units: density becomes a stable count statistic, relative distance inherits the separation of the enclosing hierarchy, and assignment runs over a compact graph of structural units rather than long sample-to-sample chains. This view is realized in Resolution-Aware Structural Density Peak Clustering (RSDP), which builds a constrained hierarchy, selects an admissible level via a single resolution parameter, and applies a one-pass DPC rule on the selected structural units. Under a faithful-realization condition, it is proved that RSDP's empirical \gamma-ranking recovers the dominant population units and that its one-pass assignment recovers the sampled ground-truth partition. On synthetic and real-world datasets, RSDP achieves the highest ACC and NMI against state-of-the-art DPC variants, exceeding the strongest by 19.91% and 9.09% in average ACC and NMI, respectively.
Abstract: Linear contextual bandit algorithms are typically built around pointwise optimism: the exploration bonus is chosen large enough to upper-bound the value of every action with high probability. We study an alternative principle of quasi-optimism, in which the index is deliberately allowed to be smaller than a valid confidence bound. We analyze the algorithmic framework, EQOL, that replaces the usual confidence-radius bonus with the capped quadratic bonus min(c₁,t ‖x‖²_V t⁻¹). This index is generally not optimistic, since its quadratic branch can be pointwise smaller than the OFUL bonus. We prove a general regret bound for arbitrary schedules (c₁,t, c₂,t), separating the exploration cost from the quasi-optimistic residual. With a single gap-agnostic tuning, EQOL recovers the high-probability minimax O(d√T) regret bound and, under a positive contextual margin, a uniform gap-dependent O(d²/Δ min,T) bound. We also provide a geometric interpretation of EQOL through action-dependent ellipsoids, clarifying how quasi-optimism differs from simply shrinking a UCB bonus. Experiments on synthetic and real-data contextual bandit benchmarks show that EQOL substantially outperforms existing algorithms.
Authors:
Logan M Bhamidipaty, Esmeralda S Whitammer, David Abel, Mykel J Kochenderfer, Subramanian RamamoorthyAbstract: We propose a novel definition of model exploitation in reinforcement learning. Informally, a world model is exploitable if it implies that one policy should be strictly preferred over another while the environment’s true transition model implies the reverse. We analogize our definition with a prior characterization of reward hacking but show that the associated proof of inevitability does not transfer to exploitation. To overcome this obstruction, we develop a general theory of reward hacking and model exploitation that proves that exploitation is essentially unavoidable on large policy sets and yields the corresponding claim for hacking as a special case. Unfortunately, we also find that the conditions that guarantee unhackability in finite policy sets have no counterparts that preclude exploitation. Consequently, we introduce a relaxed notion of exploitation and derive a safe horizon within which it can be avoided. Taken together, our results establish a formal bridge between reward hacking and model exploitation and elucidate the limits of safe planning in world models.
PaperID: 1147, Poster
Abstract: Realistic and diverse fog synthesis is a critical capability in 3D scene modelling, with broad applications in autonomous driving simulation, cinematic visual effects, gaming, and AR/VR. Yet achieving this, encompassing varied density profiles, spatial distributions, and complex optical phenomena, in real-world scenes remains a largely unsolved challenge. Existing approaches rely on appearance-driven weather transfer or simple fog overlays, which lack the physical grounding needed for volumetric effects and thus cannot faithfully model key fog properties and effects, including density gradients, anisotropic scattering, and atmospheric light shafts (Tyndall effect) — limitations that stem directly from the absence of explicit volumetric representation and physical light transport modeling. We present FogGS, a training-free, model-agnostic rendering framework that seamlessly enables existing pretrained fog-free 3DGS models to produce photorealistic, physically plausible, and fully editable fog in novel-view synthesis. Grounded in atmospheric radiative transfer theory, FogGS models fog as spatially varying extinction and scattering fields, using the depth buffer from Gaussian rasterization to drive efficient screen-space ray marching for geometry-aligned transmittance and in-scattering computation. This physically grounded parameterization affords intuitive control over fog intensity, appearance, and spatial distribution, and naturally supports anisotropic scattering, scene occlusion, and illumination. FogGS maintains high rendering efficiency practical for iterative editing workflows, remains compatible with relighting pipelines, and demonstrates superior realism, physical consistency, and controllability across diverse outdoor scenes.
Authors:
Szymon Ruciński, Pietro Bonazzi, Engin Turetken, Simon Narduzzi, Michele Magno, Nadim MaamariAbstract: Ternary Vision Transformers offer substantial model compression, however state-of-the-art methods only ternarize the encoder layers, leaving patch embeddings, LayerNorm parameters, and classifier heads in full precision. In compact models targeting resource-constrained processors, such as microcontrollers, these remaining full-precision components determine the total memory footprint, severely limiting deployment efficiency and on-device feasibility. In this work, we introduce a fully ternarized Vision Transformer in which \emphall weight matrices and normalization parameters are ternarized (FTerViT). To this end, we introduce two novel operators : TernaryBitConv2d with per-channel scaling for patch embedding and TernaryLayerNorm. FTerViT is trained using knowledge distillation, followed by a lightweight quantization-aware recovery phase. Our ternary W2A8 DeiT-III-S at 384×384 resolution achieves 82.43% ImageNet-1K top-1 at 6.09\,MB (~15× compression, -2.42\,pp vs.\ FP32), outperforming prior ternary ViTs methods up to 8 pp. Finally, we demonstrate the first implementation of ternary vision transformers on a dual cores XTensa LX7 microcontroller inside the ESP32-S3 system-on-chip. By deploying FTerViT-Small (based on DeiT-III-Small at 224×224 resolution, 5.81\,MB), we achieve 79.64% ImageNet-1K top-1 accuracy.
PaperID: 1149, Poster
Abstract: The Frank-Wolfe (FW) algorithm achieves a convergence rate of \mathcalO(1/T) for smooth convex optimization over compact convex domains, accelerating to \mathcalO(1/T^2) when both the objective and the feasible set are strongly convex. This acceleration extends beyond strong convexity: Kerdreux et al. (2021a proved rates of \mathcalO(T^-p/(p-1)) over p-uniformly convex feasible sets, a class that interpolates between strongly convex sets and more general curved domains such as \ell_p balls. In this work, we establish a matching \Omega (T^-p/(p-1)) lower bound for every p\ge 3 under exact line search or short steps, and extend the lower bound to objectives satisfying a Hölderian error bound. The proofs analyze the dynamics of FW iterates on simple instances and hence are not limited to the high-dimensional setting, unlike information-theoretic lower bounds.
Abstract: Long-context dialogue systems must decide both when to access memory and which parts of the interaction history are relevant. Existing approaches typically rely on heuristic retrieval signals or always-on memory usage, failing to account for the changing and potentially inconsistent nature of user preferences. In this work, we propose a unified framework for memory access and selection based on changing preferences. We formulate personalized memory retrieval as identifying which historical turns provide evidence about a user’s latent preference state, rather than relying on surface-level semantic similarity. To this end, we quantify the utility of each memory turn using a Bayes factor, defined as the improvement in the model’s likelihood of the reference response when the turn is included in context. This provides a principled measure of evidence strength and a unified signal for both memory access and selection. By framing memory retrieval as utility estimation, the model learns to identify salient turns and regulate memory usage based on expected utility. Experiments on four heterogeneous memory benchmarks show that our approach outperforms existing embedding-based retrieval on long-context, preference-intensive tasks where modeling changing preferences is essential, while remaining competitive in low-density regimes where semantic similarity suffices.
PaperID: 1151, Poster
Authors: Andrew Mah, Joshua L Pughe-Sanford, Sarah Harvey, Alex Williams
Abstract: Complex sequence models, such as transformers and state space models (SSMs), learn to represent latent belief states when trained on next-token prediction. Surprisingly, we find that in previously studied tasks, this phenomenon can also be captured by linear recurrent models. This raises a natural question: when are linear models sufficient for optimal belief state approximation? We study this question in the canonical setting of hidden Markov models (HMMs), where Bayes-optimal prediction requires nonlinear filtering. We characterize the full class of HMMs whose filtering dynamics are exactly realizable by linear recurrent neural networks (RNNs). More generally, for arbitrary HMMs, we establish a dissipation relation which shows that approximation error decays at a rate determined by intrinsic properties of the underlying HMM. We validate our theoretical results through experiments on both randomly sampled HMMs and constructed adversarial examples. Together, these findings clarify when linear sequence models suffice for optimal inference and when nonlinearity is fundamentally necessary.
Abstract: Pre-training of Large Language Models is often prohibitively expensive and inefficient at scale, requiring complex and invasive modifications in order to achieve high data throughput. In this work, we present Token-Superposition Training (TST), a simple drop-in method that significantly improves the data throughput per FLOPs during pre-training without modifying the parallelism, optimizer, tokenizer, data, or model architecture. TST is done in two phases: (i) A highly efficient superposition phase where we combine many contiguous tokens into one bag and train using a multi-hot cross-entropy (MCE) objective, and (ii) a recovery phase where we revert back to standard training. We extensively evaluate TST on the scale of 270M and 600M parameters and validate on 3B and a 10B A1B mixture of experts model, demonstrating that it is highly robust in different settings. Ultimately, TST consistently outperforms baseline loss and downstream evaluations, and under equal-loss settings, TST yields up to a 2.5x reduction in total pre-training time at the 10B A1B scale.
Authors: Nathan Hu, Jake Ward, Thomas Icard, Chris Potts
Abstract: While reasoning models are increasingly ubiquitous, the effects of reasoning training on a model's internal mechanisms remain poorly understood. We introduce transcoder adapters, a technique for learning an interpretable approximation of the difference in MLP computation before and after fine-tuning. We train transcoder adapters on two pairs of base and reasoning models: Qwen2.5-Math-7B / DeepSeek-R1-Distill-Qwen-7B and Qwen2.5-32B / QwQ-32B. We find that modeling the difference in MLP computation is far easier than modeling the full MLP; adapters achieve faithful reconstruction with an order of magnitude fewer active features than typical transcoders. When evaluated on reasoning benchmarks, adapters exhibit the reasoning model's characteristic long responses and recover a large fraction of its benchmark performance. Adapter features are interpretable, achieving higher automated interpretability scores than MLP neurons. To demonstrate the utility of transcoder adapters for interpreting fine-tuning differences, we present two case studies on the 7B model pair. First, we examine the overall composition of adapter features to gain a broad overview of fine-tuning differences. Despite being specific to fine-tuning by construction, many features have activating examples unrelated to reasoning structure. Interventions confirms these features disproportionately affect benchmark performance rather than response length. Second, we study why the model says 'wait' by constructing attribution graphs whose edges flow through base model parameters. We trace hesitation to only ~2.4% of adapter features (5.6k total), finding that this behavior depends largely on base model computation. These features are necessary and sufficient for producing hesitation tokens; removing them reduces response length, often without affecting accuracy. Anonymous code is available at \urlhttps://anonymous.4open.science/r/transcoder-adapters-6752/.
Abstract: Drifting models have recently gained attention for generating high-quality samples in a single forward pass. During training, they learn a push-forward map by following a vector-valued field, the drift field. We ask whether this procedure is equivalent to optimizing a scalar loss and find that, in general, it is not: drift fields are not conservative and cannot be written as the gradient of any scalar potential. We identify the position-dependent normalization as the source of non-conservatism, with the Gaussian kernel as the unique radial exception. Guided by this, we introduce the sharp kernel k^\\# and a sharp-normalized drift field that is conservative for general radial kernels. The resulting vector field is the gradient of a scalar potential that can be optimized directly using stochastic gradient descent. Moreover, the field has the form of a score difference of kernel density estimates, and gives exact equilibrium identifiability. Thus, sharp normalization closes the gap to related literature, such as Wasserstein gradient-flows and denoising score matching, also for non-Gaussian kernels. Empirically, sharp normalization preserves the performance of the original drifting objective, suggesting that the non-conservative flexibility is not required for high-quality generation.
Abstract: Low-Rank Adaptation (LoRA) reparameterizes a weight update as a product of two low-rank factors, but the Jacobian J\_\mathcalG of the generator mapping the factors to the weight matrix is rank-deficient, so the factor-space preconditioner J\_\mathcalG^\ \mathcalF\_t J\_\mathcalG induced by any W-space preconditioner \mathcalF\_t is singular, and consequently the standard chain rule cannot be uniquely inverted to map a preconditioned W-space direction back to a factor-space update. We cast existing LoRA optimizers in a unified framework parameterized by two choices: (i) which invertible surrogate for J\_\mathcalG^\ \mathcalF\_t J\_\mathcalG to use, and (ii) which \mathcalF\_t on W to use. Existing methods occupy four families along these axes: factor-space adaptive updates, block-diagonal surrogates for J\_\mathcalG^\ J\_\mathcalG, Frobenius-residual pseudoinverse methods, and Riemannian manifold constraint. Within this design space, a gradient-statistics-aware \mathcalF\_t paired with a closed-form factor-space solve at \mathcalO((m+n)r) memory remains underexplored. We propose AdaPreLoRA, which fills this gap by adopting the Adafactor diagonal Kronecker preconditioner \mathcalH\_t on W and selecting from the resulting factor-space solution family the element minimizing an \mathcalH\_t-weighted imbalance between the two factor contributions; by construction, the resulting factor update is the closest LoRA approximation to the preconditioned W-space direction under the \mathcalH\_t-weighted norm. Across GPT-2 (E2E), Mistral-7B and Qwen2-7B (GLUE, ARC, GSM8K), and diffusion-model personalization, AdaPreLoRA is competitive with or improves over a representative set of LoRA optimizers while keeping peak GPU memory at the LoRA optimizer level.
Abstract: Attribution methods for Vision Transformers (ViTs) aim to identify image regions that influence model predictions, but producing faithful and well-localized attributions remains challenging. Existing attribution methods face several limitations, with gradient-based, relevance-propagation, and attention-based methods relying on local approximations, while perturbation or optimization-based methods intervene on inputs, tokens, or surrogates rather than internal patch representations. The key challenge is that class-relevant evidence is formed through interactions between patch tokens across layers; methods that operate only on input changes, attention weights, or backward relevance signals may therefore provide indirect proxies for patch importance rather than directly testing the predictive effect of contextualized patch representations. We propose Causal Attribution via Activation Patching (CAAP), which estimates the contribution of individual image patches to the ViT’s prediction by directly intervening on internal activations rather than using learned masks or synthetic perturbation patterns. For each patch, CAAP inserts the corresponding source-image activations into a neutral target context over an intermediate range of layers and uses the resulting target-class score as the attribution signal. The resulting attribution map reflects the causal contribution of patch-associated internal representations on the model’s prediction. The causal intervention serves as a principled measure of patch influence by capturing semantic evidence after initial representation formation, while avoiding late-layer global mixing that can reduce spatial specificity. Across multiple ViT backbones and standard metrics, CAAP consistently outperforms existing methods in various settings and produces more faithful and localized attributions.
PaperID: 1157, Poster
Abstract: While While Vision Foundation Models (VFMs) are increasingly used in several computer vision tasks, their internal representations remain substantially opaque. This mostly stems from polysemanticity, i.e., individual neurons encode mixtures of unrelated visual patterns. In stark contrast, human reasoning is inherently monosemantic, i.e., it naturally relies on distinct, isolated features to process information. Therefore, monosemanticity needs to be implemented within the model's activation space to obtain representations that are human-understandable concepts. Existing methods mostly operate on convolutional networks, so when applied to VFMs their "crop-and-resize" approach distorts the model's representations by disrupting global self-attention mechanisms and discarding spatial geometry. Furthermore, prior work requires downstream classification labels and is based on KL-divergence, which requires to propagate the gradients back from the classifier head. Ultimately incurring in excessive concept extraction time, and making it hardly applicable to extract task-agnostic concepts. To overcome these issues, we propose PLACE, a fully unsupervised concept extraction framework that operates directly on the native patch-token activations of a VFM. By introducing a geometric Gram-matrix alignment loss, PLACE mechanistically controls downstream KL divergence, thus enforcing faithfulness -- i.e., the notion that extracted concepts, although obtained in an unsupervised fashion, are relevant to downstream classification tasks. Evaluations on image classification tasks (ImageNet and COCO) on DINOv2, ViT-MAE and ViT-B/16 show that PLACE (i) using Gram-alignment it extracts concepts about 20× faster compared to KL-supervised extraction and 11× faster than Sparse Autoencoders (SAEs). Moreover, PLACE yields highly sparse concepts that (ii) achieve up to 40 percentage points sparser compared to the state of the art (i.e., max 0.933 Gini sparsity score), and (iii) are up to 92.3% similar to those obtained under classifier supervision (i.e., 0.923 cosine similarity).
PaperID: 1158, Poster
Authors: Charles Dufour, Ulysse Naepels, Leonardo V Santoro
Abstract: Estimating the generative mechanism of large-scale networks is a fundamental challenge in statistical machine learning. It requires the identification of the latent connectivity structure, which is in general an NP-hard combinatorial problem due to the absence of canonical node labels. We address this challenge by allowing for probabilistic couplings, thereby relaxing the assignment problem. Our estimation framework can be formulated as a semi-relaxed Gromov–Wasserstein objective and provides a low-dimensional representation of the generative structure. We solve this via a block-coordinate conditional gradient algorithm. Despite the relaxation, the resulting solution is typically deterministic: in fact, we show that the optimality gap between the relaxed solution and the deterministic assignment vanishes at rate O(1/n), where n is the number of nodes. This allows for tractable recovery of the underlying model and enables rigorous statistical analysis: we establish consistency and minimax-optimal convergence rates for both stochastic block models and smooth graphons. Our implementation scales efficiently with n, as demonstrated on both synthetic and real-world datasets.
PaperID: 1159, Poster
Abstract: Contextual bandits require balancing nonlinear reward modeling with online efficiency. Tree ensembles capture nonlinearities but require periodic retraining and large replay buffers. Linear models update efficiently per observation with O(1) memory, but are fundamentally restricted to linear reward structures. We propose Residual Quantization (RQ) as a representation layer to bridge this gap. An offline-trained RQ codebook maps continuous contexts into discrete centroid assignments across \ell levels, set dynamically through a shadow mechanism. This enables a spectrum of additive bandit algorithms that achieve nonlinear expressivity with strictly bounded memory. Across 13 datasets, RQ variants beat their non-RQ counterparts on 11 of 13 datasets, often by wide margins, while matching doubling-retrain XGBoost and neural baselines at up to 1000× less memory.
PaperID: 1160, Poster
Authors: xingyu sha
Abstract: A point source in a PDE does not occupy an entire grid. A weather station reports one value at one site; it does not observe the temperature between stations. A pickup event in a city occurs at a coordinate and a time; it is not born as an image. Yet many neural-operator pipelines begin by turning such sparse observations into dense surrogate fields before learning. This interpolation step is often treated as harmless preprocessing, but in sparse regimes it changes the object on which the operator is asked to act. We introduce the Dirac Neural Operator (DIRAC-NO), a measure-native neural operator for sparse event-to-field learning without interpolation. DIRAC-NO keeps observations as atomic event measures until query time and constructs the output field through adaptive Dirac-kernel aggregation. Because this aggregation is implemented by direct summation, the front end preserves source additivity at the aggregation stage and admits a learned Green-kernel view of source-to-field response. Across controlled point-source PDE benchmarks and real sparse reconstruction tasks, DIRAC-NO exposes the cost of interpolation-first learning. On MeasureBench-1D, it achieves relative L^2 error 0.015 at n=128, compared with 0.49 for FNO, and it better preserves superposition structure in trained models. In 2D, it is strongest under parameter and density shift, while ERA5 reconstruction shows that the representation advantage persists beyond analytic PDEs at high observation budgets. Boundary cases and negative-control event streams further show that the benefit appears where the representation argument predicts: sparse, event-native tasks with source-response structure. These results suggest that sparse operator learning should be organized not only by the neural backbone, but by the native space in which the input is represented.
Abstract: Graph-level anomaly detection (GLAD) aims to identify graphs that deviate significantly from the majority and plays a crucial role in applications such as molecular analysis and fraud detection. While recent advances have achieved competitive performance, existing approaches mainly follow a dataset-specific paradigm, requiring substantial in-domain data and training cost when applied to new application scenarios. This limits their scalability and practicality in real-world applications where labeled data is scarce and diverse domains are continuously emerging. In this paper, we investigate the problem of generalist GLAD, which seeks to learn a single unified model capable of detecting anomalous graphs across multiple domains with minimal supervision on the target data. This problem introduces key challenges, including learning transferable patterns from heterogeneous graph datasets and effective adaptation under limited supervision. To address these challenges, we propose GenGLAD, a generalist GLAD approach that leverages discriminative subgraph extraction as the core mechanism to learn transferable anomaly knowledge and enable efficient adaptation to new domains. Specifically, GenGLAD employs a graph neural network-based discriminative subgraph extraction model to learn the anomaly-discriminative substructures that distinguish anomalous and normal graphs across diverse domains. During the test-time adaptation phase, multi-dimensional discrepancies between the extracted subgraph and multi-level contextual information are calculated to characterize anomalies from multiple perspectives. For efficient adaptation, a lightweight scoring module is introduced to automatically decide the most informative signals, improving robustness against misleading patterns. Extensive experiments on diverse benchmark datasets demonstrate that the proposed approach achieves strong performance, superior generalization ability, and improved data efficiency compared to existing methods.
Abstract: Modern AI is opening the door to collective decision-making in which participants express their views as free-form text rather than voting on a fixed set of candidates. A natural idea is to embed these opinions in a vector space so that the substantial literature on facility location problems and fair clustering can be brought to bear. But standard text embeddings measure semantic similarity, whereas distances in facility location problems and fair clustering require what we call preferential similarity: a participant's agreement with a piece of text should be inversely related to their distance from it. Off-the-shelf embeddings inherit a coarse preference signal through a correlation between semantic and preferential similarity, but fail to capture preferences when the correlation breaks. We formalize this as an invariance problem: text embedding models encode both a preference-relevant signal (stance and values) and semantic nuisance (style and wording), and the two are observationally correlated, so a geometry that relies on nuisance can appear preference-correct even when it is not. We show that synthetic training data designed to break this correlation provably shifts the optimal scorer away from nuisance-dominated cosine and significantly improves preference prediction across 11 online deliberation datasets.
Abstract: Despite the superior performance of Large Reasoning Models (LRMs), their reasoning behaviors are often counterintuitive, , excessive reasoning on simple questions or insufficient reasoning on complex ones, leading to suboptimal reasoning capabilities. This paper presents the Laws of Reasoning (LoRe), a unified framework that formalizes intrinsic reasoning patterns in ideal LRMs. We first propose with the hypothesis that the reasoning compute should scale linearly with question complexity. Beyond compute, we extend LoRe with a supplementary , positing exponential decay of accuracy with increasing complexity. Since the complexity is difficult to quantify in practice, we approximate these hypotheses via two tractable properties, . We therefore introduce LoRe-Bench, a benchmark that systematically measures these properties for large reasoning models. Evaluation shows that most reasoning models exhibit reasonable monotonicity but lack compositionality. In response, we develop an effective finetuning approach that enforces compute-law compositionality. Extensive empirical studies demonstrate that better compliance with compute laws yields consistently improved reasoning performance on multiple benchmarks, and uncovers synergistic gains across properties and laws.
PaperID: 1164, Poster
Authors:
Zichong Zhu, Zhihao Liu, Andrei Sharf, Zhanglin ChengAbstract: Procedural generation is well-suited for modeling 3D buildings, as architectural structures are naturally composed of repetitive and parameterized components. It provides a compact and editable representation by encoding geometry through executable rules. However, constructing such procedural representations typically relies on manually designed grammars, making them difficult to obtain in practice. We study image-based procedural facade reconstruction, where the goal is to recover an editable building procedural representation from an in-the-wild building image. We propose a hierarchical neuro-symbolic framework that decomposes the reconstruction into two coupled levels: a global structure layout tree for split-and-repeat facade organization, and a component parameterization representation for local architectural. A fine-tuned layout model recovers and refines the global structure through rendered visual feedback, while a tile-level predictor maps representative component crops to structured procedural attributes. The two representations are then deterministically compiled into executable CGA programs. Experiments show that our method achieves strong performance on both regular facades and challenging in-the-wild images, while maintaining high structural consistency and procedural executability. In addition, the tile-level predictor improves fine-grained architectural detail, enabling accurate and controllable component synthesis.
Abstract: Survey papers play a central role in synthesizing and organizing scientific knowledge, yet they are increasingly strained by the rapid growth of research output. As new work continues to appear after publication, surveys quickly become outdated, contributing to redundancy and fragmentation in the literature. We reframe survey writing as a long-horizon maintenance problem rather than a one-time generation task, treating surveys as living documents that evolve alongside the research they describe. We propose an agentic Dynamic Survey Framework that supports the continuous updating of existing survey papers by incrementally integrating new work while preserving survey structure and minimizing unnecessary disruption. Using a retrospective experimental setup, we demonstrate that the proposed framework effectively identifies and incorporates emerging research while preserving the coherence and structure of existing surveys.
PaperID: 1166, Poster
Authors: Ha L Nguyen, Nguyen A Minh, Dung Le
Abstract: Retrieval-Augmented Generation (RAG) has become a standard approach for enhancing large language models (LLMs) with external knowledge, mitigating hallucinations, and improving factuality. However, existing systems rely on generating natural language queries at each hop and maintaining a strict architectural separation between retriever and generator, preventing them from leveraging the full representational capacity of the LLM. We propose LAnR (Latent Abstraction for RAG), a unified framework in which a single LLM jointly performs encoding, retrieval, and generation entirely within its own latent space. Rather than generating textual queries, LAnR produces dense retrieval vectors from the hidden states of a designated [PRED] token and uses them to match against encoded document representations from the same model. Furthermore, LAnR adaptively decides when sufficient evidence has been retrieved using a lightweight MLP control head over those same hidden states, eliminating explicit token-level stopping reasoning. Extensive experiments on five QA benchmarks spanning single-hop and multi-hop settings demonstrate that LAnR outperforms existing RAG methods, while achieving improved inference efficiency through reduced number of retrieval calls and tighter model integration.
Abstract: Sample covariance matrices are fundamental objects in machine learning and statistics, with valuable information encoded into their eigenvalues. When treating modern problems, the scale of these matrices can become prohibitive for tractable computation. This creates a critical need for "spectral decompression": a method for inferring the eigenspectrum of a very large-scale model using only information from its smaller realizations. Existing methods fall into two main categories: those relying on spectral inversion; and those that rely on Free Decompression techniques. On one hand, spectral inversion is a fundamentally ill-posed problem that is notoriously unstable, often failing to yield reliable results in finite-size settings. On the other hand, while Free Decompression was recently proposed to address this scaling, its reliance on the Nica-Speicher framework restricts its use to unitarily invariant ensembles, a symmetry condition that subsampled covariance matrices fail to satisfy. In this work, we introduce Deformed Decompression (DD), a novel framework for the stable, forward-extrapolation of spectral densities. By leveraging properties of the companion matrix, DD recasts the spectral decompression problem as a directed evolution of the Stieltjes transform of the companion. This approach bypasses the instabilities of traditional inversion and provides a robust methodology for scaling covariance matrices and even their generalized variants (including Hessian matrices) that lack unitary invariance. We demonstrate the utility of our framework through experiments on a variety of random matrix examples, showing that DD successfully captures the complex spectral signatures of large-scale systems.
Abstract: Clustering is a fundamental task in unsupervised learning, serving as a core tool for data mining and exploratory data analysis. Classical k-clustering objectives such as k-median and k-means formalize this task, but typically assume access to exact pairwise distances. This assumption is increasingly unrealistic in modern applications where distances are estimated through proxy embedding models, learned similarity functions, human feedback, or other low-cost procedures that may be noisy. We study k-clustering in an unknown metric space where the algorithm has access only to a weak distance oracle: for each pair of objects, the oracle returns the true distance with probability of 1/2+\varepsilon , where \varepsilon is a constant, and otherwise may return an arbitrary corrupted value. We present randomized algorithms that, given a metric space with n vertices and a parameter k, use only weak distance-oracle queries to compute a set of representative centers together with a mapping from every input object to one of these centers, such that the total mapping cost is within a constant factor of the optimal k-clustering cost. Our first algorithm returns O(k\cdot\mathsfpolylog(n)) centers using O(n\cdot k\cdot \mathsfpolylog(n)) weak-oracle queries and runs in quasi-polynomial time. We also give a polynomial-time algorithm that returns O(k^2\cdot\mathsfpolylog(n)) centers using O(n\cdot k\cdot \mathsfpolylog(n)) weak-oracle queries. Preliminary experiments for k-means clustering show that our approach remains close to the optimum under substantial oracle noise, while noise-oblivious baselines degrade sharply.
PaperID: 1169, Poster
Abstract: Recent advances in pixel-space diffusion models have narrowed the image quality gap with latent-space diffusion, but still converge more slowly and lag behind in final image quality. We argue that a key reason is the lack of an explicit representation prior: unlike latent diffusion, which denoises in a compact and structured latent space, pixel diffusion must learn denoising-friendly representations and pixel generation simultaneously from raw RGB space. To address this problem, we propose PixelDiT2, an end-to-end pixel-space diffusion model designed to decouple representation learning from pixel generation without introducing an autoencoder or latent reconstruction bottleneck. We propose representation grounding that uses a frozen pretrained vision foundation model to provide explicit per-patch representation guidance throughout denoising, allowing the pixel diffusion transformer to focus more on pixel generation. On ImageNet-256×256, PixelDiT2 reaches FID 1.50 in 320 epochs, surpassing JiT-G's FID 1.82 at 600 epochs with roughly half the parameters and half the training budget.
Abstract: High-dimensional count data arise in applications such as single-cell RNA sequencing and neural spike trains, where mapping between distributions across successive batches or time points form critical components of data analysis. The recent success of diffusion- and flow-based deep generative models for images, video, and text motivates extending these ideas to count-valued settings, but many existing methods either treat each count as a categorical state or transform counts into a continuous space, neither of which is natural or efficient when the count range is large. We propose count-FM, a flow-matching framework for count data based on a continuous-time birth-death process with local unit jumps. Count-FM learns marginal transitions efficiently in count space through simulation-free training of conditional transition rates, allowing transport between arbitrary count-distributed source and target populations. In simulation, count-FM achieves better sample quality than representative baselines while using substantially fewer parameters. We further apply count-FM to scRNA-seq and neural spike-train data for unconditional generation, transport, and conditional generation. Across these tasks, count-FM yields improved sample quality, greater modeling efficiency, and interpretable transport paths.
PaperID: 1171, Poster
Abstract: In many real-world applications, it is undesirable to drastically change the problem solution after a small perturbation in the input as unstable outputs can lead to costly transaction fees, privacy and security concerns, reduced user trust, and lack of replicability. Despite the widespread application of graph algorithms, many classical algorithms are not robust to small input disturbances. Towards addressing this issue, we study the pointwise Lipschitz continuity of graph algorithms, a notion of stability introduced by Kumabe and Yoshida (2023, FOCS’23) and further studied in related settings (Kumabe and Yoshida, 2024, ICALP’24), (Kumabe and Yoshida, 2025, SODA’25), (Gima et al., 2025, ESA’25). Our main result is a linear programming (LP) based minimum S-T cut algorithm with a provably optimal Lipschitz constant, as witnessed by an accompanying lower bound. As a direct corollary, we give the first dynamic minimum S-T cut algorithm with non-trivial recourse bound. At the core of our techniques is a novel framework for analyzing the Lipschitz constant of regularized LP relaxations. Our framework crucially unlocks the use of weighted regularizers, which could not be analyzed through previous methods and leads to polynomial improvements in the Lipschitz constant compared to what is achievable through previous techniques. To demonstrate the flexibility of our methods, we also design an LP-based b-matching algorithm that improves on the state-of- the-art Lipschitz constant in certain input regimes when b ≡ 1. Moreover, our algorithm cleanly extends to the general case when b ≥ 1, whereas prior works are specialized to the case of b ≡ 1.
PaperID: 1172, Poster
Authors: Anastasia Orlova, Nina Gubina, Aleksei Dmitrenko, Arsen Sarkisyan, Nikita Vetoshkin, Anastasiia Gorbunova, Ivan Dubrovsky, Julia Razlivina, Bogdan Neterebskii, Andrei Dmitrenko
Abstract: Retrieval-Augmented Generation (RAG) has become a standard approach for grounding large language models in external knowledge, yet its performance depends critically on a large number of interacting design choices. Systematic optimization of these choices remains costly, as each configuration must be evaluated through full pipeline execution. In this work, we propose to treat RAG optimization as a prediction problem, using surrogate models to approximate the functional dependencies between configuration parameters and performance metrics. Using a stage-wise configurable RAG framework, we conduct large-scale experiments across three chemistry QA domains, systematically evaluating 648 configurations per domain and publicly releasing the resulting datasets of configurations and evaluation metrics. Surrogate models trained on this data achieve strong predictive accuracy (R^2 \geq 0.84 across all metrics and domains), and SHAP-IQ-based analysis of the learned models reveals consistent parameter interaction patterns across domains — despite substantial differences in corpus characteristics. We exploit this cross-domain regularity in transfer experiments: surrogate predictions generalize to unseen domains without target-domain training data on two transfer-robust metrics, and achieve full transfer across all six metrics with as little as 10% of target-domain data. Finally, we apply surrogate models to optimize RAG configurations over an expanded search space of 61,152 candidates, recovering top-performing configurations in seconds rather than days, with gains of up to +0.09 in key retrieval and generation evaluation metrics. Together, these results demonstrate that RAG design spaces are structured, transferable, and efficiently optimizable via surrogate modeling.
PaperID: 1173, Poster
Abstract: Large language models are increasingly deployed as teams of agents that hold different private information about a shared task. How should such agents concisely communicate in a way that promotes aggregation of decision relevant information? Communicating only proposed actions can discard decision-relevant uncertainty: two agents may prefer the same action for different reasons. We study an alternative interface in which agents communicate calibrated beliefs over actions. We first analyze a simple synthetic setting that captures the structure of multi-round reasoning under partial information. In this setting, we show theoretically and empirically that communicating unbiased probabilities can be strictly more powerful than communicating actions. We then test the same principle in a collaborative maze-solving task with trained transformers and pretrained LLM agents. Our results demonstrate that probability communication outperforms action communication when the exchanged predictions preserve calibrated uncertainty. When communicated probabilities are biased we show that the benefits of probability communication can disappear. We give a post-hoc conversation calibration intervention that consistently improves decision-making by correcting these biases. These results suggest that probabilistic communication is useful for multi-agent collaboration not merely because it transmits more information, but because calibrated uncertainty gives other agents a reliable object to update on.
PaperID: 1174, Poster
Abstract: Conformal Prediction (CP) provides a model-agnostic framework for constructing finite-sample prediction sets with coverage guarantees and has become a tool for reliable decision-making in high-risk settings. Recent advances have extended Split Conformal Prediction to increasingly complex learning problems, including multilabel, high-dimensional, and functional outputs. In this work, we study CP for Distribution-to-Distribution Regression, a setting of growing interest across many application domains. We show that a naive application of Split CP yields prediction sets that, while marginally valid, are non-interpretable, e.g., lacking a well-defined volume or a tractable sampling procedure, and of limited practical value. To overcome this, we propose a new conformal framework that operates in an orthogonal basis space, producing interpretable prediction sets while preserving coverage guarantees. We further develop an adaptive variant with asymptotic conditional coverage under mild assumptions. Empirical results on synthetic and real-world datasets validate the effectiveness of the proposed approach.
PaperID: 1175, Poster
Abstract: Many social applications of machine learning exhibit performative effects: population behavior changes in response to deployed models. Performative prediction studies this interaction through a distribution map that relates each model to the population distribution it induces. One of the main results in this framework showed that repeated risk minimization (RRM), which updates models by retraining on the most recent data, can converge to a stable model that minimizes risk on its own induced distribution. However, existing analyses typically assume access to the complete distributions of features and labels after model deployment, ignoring the possibility of : observing labels only for the accepted subset of the population. In this work, we formalize performative prediction with selective labels and show that retraining only on observed data can misguide the retraining procedure and undermine the guarantees of convergence to a stable solution. We then propose a worst-case objective based on knowledge of a confidence interval on the probability of a positive label. Applying RRM to this objective permits us to remain within a bounded distance to the true stable point. Under a sensitivity assumption on the conditional label distribution, we further show how previously accepted data can tighten these confidence intervals over time. Experiments in a lending application with fairness regularization show that our robust optimization approach closely matches the performance of RRM with complete label access.
Authors:
Qijing Huang, Sana Damani, Zhifan Ye, Athinagoras Skiadopoulos, Siva Kumar Sastry Hari, Jason Clemons, Sahil Modi, Jingquan Wang, Aditya Kane, Edward C Lin, Humphrey Shi, Christos KozyrakisAbstract: a deep-learning model run on target hardware, and how far is today’s implementation from that limit? These questions are central to software, hardware, and algorithm optimizations. Speed-of-Light (SOL) analysis answers them by computing a workload’s theoretical minimum execution time on a given architecture. Yet deriving SOL bounds remains manual, error-prone, and disconnected from rapid model development. To close this gap, we introduce SOLAR, a framework that automatically derives validated SOL bounds from Pytorch and JAX source code. SOLAR leverages both components in its flow: an LLM frontend translates any source programs into an executable Affine Loop IR, validated by output comparison; a deterministic flow lifts the IR into an einsum graph; and an analytical backend computes unfused, fused, and cache-aware SOL bounds. SOLAR provides comprehensive operator and language coverage, produces validated bounds with zero observed SOL violations, and offers multi-fidelity analysis that tightens bounds and surfaces optimization insights. We evaluate SOLAR across KernelBench, JAX/Flax models, and robotics workloads. These experiments demonstrate four use cases: headroom analysis at multiple fidelity levels, identifying optimization opportunities, cross-platform exploration, and inverse-roofline hardware provisioning.
PaperID: 1177, Poster
Authors: Tomas Karella, Emily Shinkle, Alice E Allen, Pieter J Swart, NIcholas Lubbers, Roxana Bujack
Abstract: Respecting a-priori known symmetries of underlying data is a key principle in designing efficient neural architectures. Convolutional Neural Networks (CNNs) are a canonical example, encoding translation equivariance. Extending this inductive bias to richer symmetry groups such as \mathrmSE(2) in the most effective way possible is an ongoing architectural and computational challenge. Building on classical invariant theory, we derive invariant feature maps that are provably optimal: minimal and information-preserving with respect to the symmetry group. Our construction is designed to integrate seamlessly into standard architectures, requiring no modification to nonlinearities or normalization layers. As a concrete implementation, we introduce a plug-and-play module for CNNs that enforces \mathrmSE(2) invariance, capturing both translations and rotations within standard architectures. The resulting models are both efficient and expressive. Experiments show competitive or improved accuracy compared to state-of-the-art equivariant methods at significantly lower computational cost. Finally, we provide a theoretical and empirical analysis of the relationship between equivariance and invariance, offering insight into their roles in controlling robustness and flexibility.
Authors: Tatiana Belova, Yuriy Dementiev, Danil Sagunov
Abstract: The field of learning-augmented algorithms has demonstrated that machine-learned predictions can bypass worst-case lower bounds across a wide range of problems. So far, however, the focus has been almost exclusively on algorithms, where predictions improve competitive ratios, approximation guarantees, or running times. In this paper, we raise the question of whether predictions can push the frontier of that augments an entire family of state-of-the-art exact algorithms for a variety of subset selection problems. We show that a noisy predictor that is only marginally better than random guessing suffices to provably reduce the search space, and that the resulting runtime speedup scales smoothly with the prediction quality. Importantly, our algorithms require only independence of predictions or, alternatively, do not require the knowledge of the predictor's accuracy–both strictly weaker and more realistic settings than typically assumed.
PaperID: 1179, Poster
Abstract: We introduce Point4D, a feed-forward model for 4D reconstruction of long-range video sequences. Point4D is able to reliably infer dense per-point 3D trajectories across multi-hundred video sequences, unlike existing 4D methods that are limited to short input windows of at most a few dozen frames. A key innovation that enables this is our flexible 3D query-based motion decoder that decouples trajectory prediction from image-plane visibility. The predicted 3D endpoints are then directly re-queried in the next chunk without re-projection or matching. Furthermore, we show that extracting and reusing a visual across an arbitrary frame where the point is visible leads to superior performance than purely relying on the source patch. Overall, Point4D achieves state-of-the-art performance across diverse long-video tracking benchmarks spanning over 200 frames and substantially outperforms previous feed-forward 4D methods.
PaperID: 1180, Poster
Authors: Carl Allen
Abstract: We characterise disentanglement for smooth generative pushforward models, such as in VAEs and GANs. For a generator/decoder g:\\mathcalZ\\to\\mathcalX and factorised prior p(z)=\\prod_i p_i(z_i), we define disentanglement as standard statistical independence expressed over the generated manifold: the pushforward density p_\\mu = g_\\#p factorises into one–dimensional "seam" factors, each controlled by a distinct latent coordinate. We prove that p_\\mu factorises according to the SVD of g's Jacobian; that disentanglement equates to two conditions on g (C1-C2); and that under such conditions the seam factors are identifiable, up to permutation and sign. In the special case of Gaussian (\\beta-)VAEs, we show via an identity how diagonal posteriors promote C1-C2, in expectation, explaining why disentanglement arises modulated by \\beta. Experiments illustrate this mechanism on Gaussian data, dSprites, and CelebA.
Abstract: We study sequential decision-making with offline reinforcement learning (RL). Traditional offline RL policies may result in out-of-distribution (OOD) actions when training relies only on sparse offline representations. To ensure safe offline policies in a sparse state-action space, we explore how density estimation models can be integrated into model-based RL methods to avoid the OOD regions. Generative models are capable of explicitly modeling the density in sparse state-action spaces. Building on this, we introduce , a density-regularized offline RL algorithm that uses generative density modeling to restrict policy updates to high-density areas of the dataset. We present theoretical results on GORMPO's performance. Furthermore, we examine whether better OOD detection corresponds to better model-based offline policies. We compare (1) the OOD detection capabilities of various density estimators and (2) their performance within the GORMPO framework on a real-world medical dataset and sparse offline RL datasets. We theoretically guarantee GORMPO's performance under mild assumptions. Empirically, GORMPO outperforms state-of-the-art baselines by 17% on a real-world medical dataset and enhances the base model on the offline RL datasets. Our empirical findings show that better OOD detection generally results in improved policies in environments with stable dynamics, while conservative penalties with poor density estimation are favored when dynamics are uncertain.
Abstract: Partially Observable Markov Decision Processes (POMDPs) are systems in which one agent interacts with a stochastic environment, and receives only partial information about the current state. In a multi-environment POMDP (MEPOMDP), the initial state is unknown, and assumed to be adversarially chosen. In this work we focus on computing the optimal value and policy in MEPOMDPs with finite-horizon objectives. That problem is known to be PSPACE-complete in POMDPs. Our main results are as follows: (1) we establish that it is also PSPACE-complete in the more general setting of MEPOMDPs; (2) we present a practical algorithm and evaluate it on classical benchmarks, significantly outperforming the only previously known algorithm.
Abstract: How do transformer language models memorize factual associations? A common view casts internal weight matrices as associative memories over pairs of embeddings, requiring parameter counts that scale linearly with the number of facts. We develop a theoretical and empirical account of an alternative, \emphgeometric form of memorization in which learned embeddings encode relational structure directly, and the MLP plays a qualitatively different role. In a controlled setting where a single-layer transformer must memorize random bijections from subjects to a shared attribute set, we prove that a logarithmic embedding dimension suffices: subject embeddings encode \emphlinear superpositions of their associated attribute vectors, and a small MLP acts as a relation-conditioned selector that extracts the relevant attribute via ReLU gating, and not as an associative key-value mapping. We extend these results to the multi-hop setting---chains of relational queries such as ``Who is the mother of the wife of x?''---providing constructions with and without chain-of-thought that exhibit a provable capacity--depth tradeoff, complemented by a matching information-theoretic lower bound. Empirically, gradient descent discovers solutions with precisely the predicted structure. Once trained, the MLP transfers zero-shot to entirely new bijections when subject embeddings are appropriately re-initialized, revealing that it has learned a generic selection mechanism rather than memorized any particular set of facts.
Authors: Junhao Cai, Dohun Kim, Dowon Kim, Sung I Choi, Chengjun Jin, Juhyun Park, Changhee Joo
Abstract: Machine unlearning seeks to remove the influence of designated training data while preserving performance on the remaining data. Approximate unlearning can be viewed as a local editing problem; in min-max unlearning, the key local object is the surrogate point at which the retain objective is evaluated. When forget and retain gradients are strongly aligned, an unconstrained forget-maximizing perturbation can move to a surrogate point that increases retain loss. We propose Retain-Orthogonal Surrogate Unlearning (ROSU), which constrains the inner surrogate construction by maximizing first-order forget gain subject to zero first-order retain change under a fixed perturbation budget. This yields a closed-form retain-orthogonal perturbation, a lightweight transported outer update, and amplification along the retain-neutral direction. Our analysis establishes (i) a curvature-controlled second-order bound on retain damage, (ii) a positive-alignment regime in which ROSU strictly reduces surrogate retain loss relative to standard min-max perturbations, and (iii) near-equivalence when the two gradients are nearly orthogonal. Across vision and language benchmarks (CIFAR-10/100, Tiny-ImageNet, TOFU, WMDP), the empirical pattern follows this geometry: ROSU gives its clearest gains in high-coupling regimes while remaining competitive elsewhere.
PaperID: 1185, Poster
Abstract: Algorithmic risk predictors are increasingly being used to inform resource allocation in healthcare and other high-stakes domains. Common in many such settings, the public decision maker delegates service provision to private organizations and compensates them according to each individual's predicted cost of service. We show that this common allocation mechanism can fail when organizations engage in \emphfavorable selection, i.e., strategically choosing individuals whose actual costs are below the payment model predictions. Motivated by health insurance payments in the Medicare Advantage (MA) program, we provide a theoretical framework to study the interaction between a decision maker and a strategically selecting service provider. Our theoretical framework predicts two empirical patterns observed in Medicare Advantage that have become central concerns: enrollment in the private program expands, while government payments rise above the counterfactual cost of direct public provision. Guided by our theoretical framework, we then propose a reformed payment policy that is implementable under the government's data-access constraints. We prove that its equilibria eliminate overpayment while preserving profits for efficient organizations. We validate the framework in synthetic and semi-synthetic experiments that illustrate the advantages of our reformed policy, over the status quo.
Abstract: Generative agents have proven to be powerful assistants in a wide variety of contexts. Given this success, users are now deploying agents with minimal restrictions in open ended, multi-agent environments. Current methods for measuring the dynamics of open-ended multi-agent systems are limited to qualitative inspection. In this paper, we extend the process-theoretic notion of adaptive control charts to multi-agent systems to enable automated monitoring. Using simulation, we demonstrate that adaptive control charts are necessary for monitoring multi-agent systems that can learn from their environment. We further demonstrate, both empirically and theoretically, that adaptive control charts are susceptible to adversarial agents that defect sufficiently slowly. These results illustrate a fundamental tradeoff in multi-agent system control: either agents in a system cannot learn or the system is susceptible to adversaries.
PaperID: 1187, Poster
Abstract: In the k-diverse nearest neighbor search problem, the input is a dataset in which each point is assigned a color, and given a query q, the goal is to retrieve k approximate nearest neighbors with distinct colors. This color diversity metric is highly essential for information retrieval tasks where results are required to have different categories, e.g., recommending products from different sellers. Previous state-of-the-art solution by Anand et al. [ICML 2025] constructs a diversity-aware graph to meet this requirement. However, they suffer from an \mathcalO(k^2) multiplicative query time overhead, creating a computational bottleneck when k is large. To break this barrier, we propose a novel color bucketing framework that selects a subset of colors in each bucket and instantiates an independent approximate nearest neighbor search algorithm on the corresponding points. Inspired by group testing, we provide a randomized color sampling scheme that achieves an \mathcalO\left(k^1-\frac1c^2-o(1)\log k\right) overhead in Euclidean spaces and an \mathcalO\left(k \log k\right) overhead in general metric spaces, where c is the approximation factor, reducing a factor of k. In addition, the space usage to store the data structure matches or improves upon prior methods. Empirical results on semi-synthetic and real-world datasets demonstrate that our approach achieves significantly faster search times and higher recall compared to state-of-the-art diversity search baselines. Furthermore, we introduce a dataset-oblivious deterministic bucketing scheme using expander graphs for a relaxed diverse search problem that only requires k(1-\varepsilon) colors. Here, dataset-oblivious means that the color bucket construction does not depend on the spatial configuration of the dataset. We then establish an \Omega(k^2) lower bound on the number of buckets for any dataset-oblivious deterministic bucketing scheme that gives an \emphexact k-diverse solution. We subsequently circumvent this by a dataset-aware adaptive segment tree bucketing scheme with an overhead of only \mathcalO(k \log N), where N is the number of colors. As an extension of our results, we solve the problem with general diversity metric, where the goal is to maximize the minimum pair-wise distance of the k solutions, and improve the query time.
PaperID: 1188, Poster
Abstract: We introduce CoLa3D, a simple and effective model for composable 3D scene decomposition that, given a single whole-scene latent from either a single image or an existing scene mesh as input, directly outputs a set of object meshes with consistent spatial layout that can be assembled back into a coherent scene. CoLa3D introduces a novel paradigm for 3D scene decomposition within a shared \emphwhole-scene latent, unlike alternatives that perform per-object generation in canonical spaces and then glue the objects back together with fragile pose and scale estimation. The network is a lightweight, flexible, and scalable query-based transformer that accepts various types of queries, including 2D masks, 3D points, and learnable queries, and learns to decompose the scene into objects in a compact latent space. On both synthetic and real-world datasets, CoLa3D achieves state-of-the-art performance on composable 3D scene decomposition, substantially improving global scene-level consistency and object-level quality while running 2× faster and using less than 1% of SAM 3D's training data. Visible depth on par with MoGe2 further confirms that latent decomposition preserves accurate visible scene geometry while producing physically complete, decomposable object meshes.
Abstract: Machine learning methods rely on data. However, gathering suitable data can be challenging due to availability constraints, cost, or the need for domain expertise. Expanding datasets with additional sources is a common response to limited data, yet this practice does not always improve downstream performance and can sometimes lead to a loss of performance, known as negative transfer. We propose RADAR, a simple, geometrically grounded metric for estimating cross-domain transferability in foundation models. RADAR analyzes the layer-wise evolution of representations by measuring angular alignments and relative changes in distance along layer-to-layer displacement trajectories, and by comparing empirical distributions of within-domain and cross-domain dynamics. We hypothesize that domain transferability is related to the divergence between these trajectory distributions. We evaluate the metric across multiple modalities, including cross-lingual sentiment classification with text embedding models and cross-domain image classification with foundation vision models. Across several settings, RADAR provides competitive predictive performance relative to existing transferability metrics on several vision and text benchmarks, with particularly strong results when domain transitions are smooth or cleanly separated. Our ablations further suggest that the effectiveness of transferability estimation depends on the geometry of the model’s internal representation space, with different modalities favoring different topological formulations.
PaperID: 1190, Poster
Abstract: AI-generated text, images, audio, and video are now ubiquitous and difficult for humans to reliably identify as such, making scalable, automated detection a practical necessity. Existing detectors, however, are fragmented along modality lines---each typically relying on its own classifier head, dataset, and training pipeline---and the recent line of explainable detectors learns its rationales from human annotations whose faithfulness degrades as generators improve. We propose a unified method that addresses both issues by extending neologism learning from text generation steering to multimodal classification. We add two tokens, < REAL> and < AIGEN >, to the vocabulary of an otherwise-frozen multimodal LLM (MLLM) and train only their embedding rows (on the order of 2d parameters, fewer than 0.001% of the backbone) so that the model's next-token distribution under a paired prompt encodes the real-vs-AI-generated decision. Because all modalities are projected into a shared input embedding space, the same token pair applies across text, image, audio, and video without per-modality heads, and a pair trained on a strict subset of modalities transfers to held-out ones at test time. The trained tokens further admit free-form natural-language descriptions of what < AIGEN > has come to mean, decoded from the same frozen backbone and never supervised during training; we validate their causal faithfulness via plug-in evaluation. Across standard text, audio, image, and video benchmarks, our method matches or exceeds dedicated, modality-specific detectors and outperforms MLLM-based explainable detectors out of distribution, despite training orders of magnitude fewer parameters and using no human-annotated explanations.
Abstract: We introduce a novel Mutual Information (MI) estimator that fundamentally reframes the discriminative approach. Instead of training a classifier to discriminate between joint and marginal distributions, we learn a normalizing flow that transforms one into the other. This technique produces a computationally efficient and precise MI estimate that scales well to high dimensions and across a wide range of ground-truth MI values.
Abstract: While large language models (LLMs) are trained to align with human values, their generations may still violate safety constraints. A growing line of work addresses this problem by modifying the model’s sampling policy at decoding time using a safety reward. However, existing decoding-time steering methods often intervene unnecessarily, modifying generations that would have been safe under the base model. Such unnecessary interventions are undesirable, as they can distort key properties of the base model such as helpfulness, fluency, style, and coherence. We propose a new test-time steering method designed to reduce such unnecessary interventions while improving the safety of unsafe responses. Our approach filters tokens using a value-based safety criterion and provides an explicit bound on the probability of false interventions. A single threshold hyperparameter controls this bound, allowing practitioners to trade off higher rates of unnecessary intervention for better output safety. Across multiple datasets and experiments, we show that our value-filtered decoding method outperforms existing baselines, achieving better trade-offs between safety, helpfulness, and similarity to the base model.
Authors:
Shaoang Li, Yanhang Shi, Yufei Li, Mingfu Liang, Xiaohan Wei, Yuchen Pu, Fei Tian, Chonglin Sun, Frank Shyu, Luke Simon, Sandeep Pandey, Xi Liu, Jian LiAbstract: Large Language Models (LLMs) can reason well over focused inputs, yet often miss decisive evidence buried in long, noisy contexts. We introduce HiLight, an Evidence Emphasis framework that turns evidence selection into a learned, non-destructive input-side control problem for frozen LLMs. Rather than retrieving, pruning, compressing, or rewriting the input, HiLight trains a lightweight Emphasis Actor to insert minimal highlight tags around pivotal spans while preserving the original context. A frozen Solver then reasons over the emphasized input. The Actor is trained only from the Solver's downstream task reward, requiring no evidence labels, Solver gradients, logits, or internal activations. This yields a solver-compatible alternative to retrieval-style hard selection, context compression, and instance-level prompt rewriting. Across sequential recommendation and QA, HiLight consistently improves over manual prompting and strong automated prompt-optimization baselines, with gains up to +27.5% over manual instruction (MI) and +10.8% over the strongest baseline on Amazon-Beauty. The learned emphasis policy transfers zero-shot to smaller and larger unseen Solvers across model families, including an API-based Solver, and its highlights align with human supporting evidence up to 0.78 F1. These results suggest that evidence selection can be learned as a reusable input-side control mechanism for frozen LLMs.
Abstract: Discrete diffusion models are a powerful, emerging, paradigm for code generation. They construct programs through iterative refinement of partially corrupted token sequences and enable parallel token refinement. Importantly, this paradigm exposes a global program state at each denoising step, which provides a natural intervention point for enforcing program-level functionality and security constraints and guide generation before the final code is committed. Building on this observation, the paper introduces Constrained Diffusion for Code(CDC), a training-free neurosymbolic inference framework that integrates constraint satisfaction directly into the reverse denoising process. CDC augments the base discrete diffusion sampler with constraint-aware denoising operators that combine mathematical optimization with program analysis to identify constraint-relevant regions of the intermediate program state and locally adjust the denoising trajectory, steering generation toward feasible programs while remaining close to the base diffusion model. Across code generation benchmarks, CDC consistently improves constraint satisfaction in functional correctness, security, and even syntax, outperforming discrete diffusion and autoregressive baselines with less corrective computation and more localized edits.
Abstract: State-based fine-tuning has emerged as a compelling alternative to weight-based adaptation for transformers, updating lightweight controls into states rather than model weights, offering substantial memory savings while retaining parameter efficiency. However, most existing state-based methods typically apply only per-block control updates, which limits inter-block information exchange and restricts representational adaptation. Meanwhile, prior mechanisms that enable cross-block communication often introduce considerable computational overhead, reducing their practicality for efficient fine-tuning. We introduce Mixture-of-Control (MoC), a lightweight fine-tuning framework that adaptively integrates local and global control signals to enhance representation learning. MoC treats block-wise control states as experts in a sparse mixture-of-experts process, enabling efficient communication across transformer blocks. Experiments on 33 datasets and 10 pretrained models across diverse transformer-based benchmarks demonstrate that MoC consistently outperforms existing state-based fine-tuning methods while maintaining comparable efficiency in memory and compute.
Abstract: Supervised training of deep neural networks for classification typically relies on hard targets, which promote overconfidence and can limit calibration, generalization, and robustness. Self-distillation methods aim to mitigate this by leveraging inter-class and sample-specific information present in the model’s own predictions, but often remain dependent on hard targets without explicitly modeling predictive uncertainty. With this in mind, we propose Deep Probabilistic Supervision (DPS), a principled learning framework that constructs sample-specific target distributions via Bayesian inference on the model’s own predictions and remains independent of hard targets after initialization. We show that DPS consistently yields higher test accuracy (e.g., +2.0% for DenseNet-264 on ImageNet) and significantly lower Expected Calibration Error (ECE) (-40% for ResNet-50 on CIFAR-100) than existing self-distillation methods. When combined with a contrastive loss, DPS achieves state-of-the-art robustness under label noise.
PaperID: 1197, Poster
Abstract: Recent development of large language models (LLMs) has empowered multi-agent systems (MAS) through external skills and inter-agent discussion to solve complex tasks. Although such a paradigm improves the capability of LLM systems, it also introduces substantial efficiency overhead. Existing approaches mainly reduce this overhead through software-level optimization. However, these methods overlook how to exploit the graph-level structure of MAS discussion and its suitability for hardware-aware acceleration, thus causing latency, especially when the number of agents, discussion rounds, or concurrent sessions increases. In this paper, we propose FASD, an FPGA-algorithm co-design framework for acceleration towards graph dynamic system agent discussion. We argue that MAS discussion can be formed as a graph-coupled dynamical system and thus agent activation pruning can be optimized through phase-coupled dynamics method. Based on these insights, we designed the FPGA-aware optimization for the problem. FASD mainly consists of the following parts. Firstly, we profile representative MAS workflows and construct a phase profile for dynamic system modeling. Secondly, the hardware-friendly optimization progress, given the profile, is implemented on FPGA. Through extensive experiments on multi-agent discussion workloads, we demonstrate that FASD improves discussion efficiency while preserving task performance, with stronger benefits as the agent graph becomes larger and more complex.
PaperID: 1198, Poster
Authors:
Yu Gu, Zijun Yu, Chi-Kuang Yeh, Xinyu Wang, Ziyang SongAbstract: Medication recommendation is a multi-label prediction task, yet existing prediction-set evaluation and calibration methods rely mainly on exact-overlap metrics and therefore treat all missed medications as equally severe. However, missing a medication but predicting a therapeutically close drug is less severe than matching it only with distant predictions. To address this, we introduce , a post-hoc framework for calibrating medication prediction sets under hierarchy-aware mismatch severity. CSC uses the Anatomical Therapeutic Chemical (ATC) hierarchy to quantify mismatch severity and controls patient-level severity without retraining the predictor, supporting both average severity and upper-quantile severity for the highest-severity mismatches. Experiments on MIMIC-IV show that CSC tracks user-specified severity tolerances and shifts prediction sets away from distant mismatches toward exact or close matches.
Authors:
Udith Haputhanthri, Declan Campbell, Rim Assouel, Jonathan D Cohen, Taylor WebbAbstract: Despite success on standard benchmarks, vision language models display persistent failures on tasks involving processing of multi-object scenes, including many tasks that are relatively easy for humans. Recent work has found that these failures may stem from a basic inability to accurately bind object features in-context, a challenge that is referred to as the ‘binding problem' in cognitive science and neuroscience. The human visual system is thought to solve this binding problem via serial processing, attending to individual objects one at a time so as to avoid interference from other objects. Recent work has proposed `pointing' -- the use of explicit spatial coordinates to refer to objects -- as an analogous solution for vision language models, and found that it improves performance on challenging multi-object tasks. However, it is unclear why (i.e., on a mechanistic or representational level) this approach improves performance, and how directly this relates to serial processing in human vision. Here, we investigate this question. We find that learning to point-via-text induces an internal visual search routine, and we characterize the mechanisms that support this procedure. We also find that pointing behavior can be generalized to new tasks via fine-tuning, and that doing so eliminates binding errors and enables compositional generalization. These results provide a proof-of-principle that serial processing can solve the binding problem for vision language models just as it does for biological vision.
PaperID: 1200, Poster
Abstract: While reinforcement learning (RL) has been central to the recent success of large language models (LLMs), RL optimization is notoriously unstable, especially when compared to supervised fine-tuning (SFT). In this work, we investigate the stability gap between SFT and RL from a gradient-based perspective, and show that the convexity of the SFT loss with respect to model logits plays a key role in enabling stable training. Our theoretical analysis demonstrates that this property induces favorable gradient directionality during optimization. In contrast, Proximal Policy Optimization (PPO), a widely adopted policy gradient algorithm utilizing a clipped surrogate objective, lacks this stabilizing property. Motivated by this observation, we propose Logits Convex Optimization (LCO), a simple yet effective policy optimization framework that aligns the learned policy with an optimal target derived from the original RL objective, thereby emulating the stabilizing effects of logits-level convexity. Extensive experiments across multiple model families show that our LCO framework consistently improves training stability and outperforms conventional RL methods on a broad range of benchmarks.
Abstract: Current large language models reason in isolation. Although it is common to sample multiple reasoning paths in parallel, these trajectories do not interact, and often fail in the same redundant ways. We introduce LACE, a framework that transforms reasoning from a collection of independent trials into a coordinated, parallel process. By repurposing the model architecture to enable cross-thread attention, LACE allows concurrent reasoning paths to share intermediate insights and correct one another during inference. A central challenge is the absence of natural training data that exhibits such collaborative behavior. We address this gap with a synthetic data pipeline that explicitly teaches models to communicate and error-correct across threads. Experiments show that this unified exploration substantially outperforms standard parallel search, improving reasoning accuracy by over 7 points. Our results suggest that large language models can be more effective when parallel reasoning paths are allowed to interact.
PaperID: 1202, Poster
Authors: Sikun Xu, Zineng Xu, Mingduo Zhao
Abstract: Large language models have made synthetic data inexpensive, but they can still be biased. When real data are scarce, researchers must decide how much of the real sample to use for direct estimation and how much to spend calibrating the generator. We study this allocation problem and derive the marginal-value condition for the optimal calibration size: (v_n + c x^-2\beta)^2 = 2\beta a c x^-(2\beta+1), where x is the calibration size, v_n is synthetic sampling variance, a is the real-estimator variance constant, and c x^-2\beta is squared synthetic bias. This condition has no universal fixed-share solution and yields five allocation regimes. In the main regime, m_n \asymp n and \beta > 1/2, the optimal calibration size grows only as n^2/(2\beta+1), so the calibration share vanishes; rules that keep a positive fixed fraction of real observations in calibration over-calibrate asymptotically. We also propose an adaptive grid estimator with an oracle inequality and a safety check, xB_n(x)
Abstract: Test-time compute has emerged as a major approach to improving the capabilities of Large Language Models (LLMs). However, existing test-time reasoning paradigms rely heavily on externally imposed control, either through fixed reasoning programs or through costly expansion in constrained search spaces, limiting both generalization and efficiency. We propose State of Thought (SoT), a new reasoning paradigm that enables endogenous reasoning in LLMs, with the model's internal reasoning state governing how reasoning unfolds. Concretely, SoT extracts a compact dynamics-geometric state from the model's internal information transfer and selectively activates historical reasoning support useful under the current reasoning state, framing reasoning as a state-conditioned process over evidence rather than an externally prescribed token chain. Across quantitative (1.29x), general (1.62x), symbolic and code (1.72x), long-context (2.63x), and multimodal (1.08x) reasoning tasks on 4 models with 20 datasets, SoT consistently improves task accuracy while reducing tokens by 69.0% and latency by 48.7%, supporting endogenous state-driven reasoning as a more generalizable and efficient alternative.
Authors: Léo Nicollier, Max Dunitz, Marc Pic, Pablo Muse, Enric Meinhardt-Llopis, Gabriele Facciolo
Abstract: A fundamental open question in self-supervised learning (SSL) is the explicit characterization of the optimal geometry of the learned representations. Recently, LeJEPA identified isotropic Gaussian embeddings as optimal for minimizing downstream prediction risk in Euclidean spaces. However, the corresponding problem for distributions supported on lower-dimensional manifolds, such as the hypersphere, remains unexplored. In this work, we demonstrate that extending this minimax analysis to smooth distributions on Riemannian manifolds fundamentally changes the optimal solution. We show that, under a worst-case formulation, both \(k\)-nearest neighbors and kernel ridge regression induce hyperspherical uniformity. More precisely, we show that uniform distributions on manifolds are optimal for \(k\)-nearest neighbors, and that the uniform distribution on the sphere is optimal for kernel ridge regression with both the exponential dot-product kernel and the linear kernel. This theoretical insight reveals a fundamental limitation of Gaussian embeddings: their non-uniform density induces anisotropic k-NN neighborhoods, severely biasing the estimator. To correct this, we introduce SPHERE-JEPA, a theoretically grounded SSL framework. We adapt LeJEPA's Cramér--Wold projection mechanism to enforce hyperspherical uniformity rather than a Gaussian prior. Empirically, SPHERE-JEPA yields significant improvements, boosting texture retrieval mAP by over 6%, while consistently matching or outperforming LeJEPA on standard benchmarks—including a +1.8% linear probing gain on ImageNet-1K (ViT-B/14).
PaperID: 1205, Poster
Abstract: We study two variants of the mirror descent-ascent (MDA) algorithm for solving min-max problems on the space of measures: simultaneous and alternating. We work under assumptions of convexity-concavity and relative smoothness of the payoff function with respect to a suitable Bregman divergence, defined on the space of measures via flat derivatives. We establish non-asymptotic convergence rates to mixed Nash equilibria, measured in the Nikaido-Isoda error, proving an \mathcalO(N^-1/2) rate for simultaneous MDA and an improved \mathcalO(N^-2/3) rate for alternating MDA. The main technical contribution is an infinite-dimensional dual space analysis that relates Bregman divergences on measures to dual Bregman divergences on spaces of bounded continuous functions, allowing us to control asymmetric commutator terms created by alternating updates. The results substantially generalize prior analyses restricted to bilinear objectives and also apply to nonlinear convex-concave problems on measure spaces, thereby providing a unified theoretical foundation for MDA in mean-field min-max optimization.
PaperID: 1206, Poster
Abstract: It has been recently shown that e-processes are sufficient for sequential testing in the following sense: every level-\alpha sequential test can be obtained by thresholding an e-process at 1/\alpha. However, in the above result, neither does the test have to be asymptotically optimal (in terms of stopping times) nor does the e-process have to be asymptotically log-optimal. It has separately been shown that asymptotically log-optimal e-processes yield asymptotically optimal sequential tests. In this paper, we prove the converse, arguably completing the story: it is possible to aggregate asymptotically optimal sequential tests into asymptotically log-optimal e-processes. This is accomplished by using a new class of WAIT e-processes: those that are Weighted Aggregates of Indicators of Stopping Times that begin at zero, are nondecreasing and increase to infinity under the alternative at the optimal rate. Importantly, the paper discusses several nuances in the varied definitions of asymptotic (log-)optimality.
Abstract: Behavioral Foundation Models (BFMs) are an emerging paradigm in reinforcement learning, playing a role analogous to large language models in natural language processing: they have shown remarkable versatility, enabling zero-shot performance, fast imitation, and online adaptation, all by exploiting the structure of a latent space. In this work, we investigate whether the latent behavioral space induced by BFMs can serve as an effective search space to discover large repertoires of behaviorally diverse and high-performing policies through Quality-Diversity (QD) methods. While QD methods generally search directly in high-dimensional policy parameter space, in this paper, we present BFM-QD, a framework that performs QD search in the compact latent space of a BFM. We further show that the BFM-QD framework provides a closed-form, gradient-free policy improvement operator that approximates a policy gradient update, but requires no critic training and no backpropagation. Across continuous-control benchmarks spanning dense locomotion, sparse navigation, and contact-rich manipulation, BFM-QD consistently outperforms parameter-space baselines, with particularly stark gains in sparse and deceptive settings, where all existing QD methods collapse to near-zero performance. These results show the effectiveness of the BFM-QD framework, benefiting from the synergy between dimensionality reduction of the search space and offline pretraining from diverse behavioral data. This positions BFMs as a general-purpose backbone for QD optimization, extending their utility beyond zero-shot task solving to the discovery of diverse behavioral repertoires.
Abstract: Learning from feedback is an instrumental process for advancing the capabilities and safety of frontier models, yet its effectiveness is often constrained by cost and scalability. We present a pilot study that explores scaling reward models through unsupervised approaches. We operationalize reward-based scaling (RBS), in its simplest form, as preference learning over document prefixes and suffixes drawn from large-scale web corpora. Its advantage is demonstrated in various aspects: despite using no human annotations, training on 11M tokens of math-focused web data yields steady gains on RewardBench v1 and v2, and these improvements consistently transfer across diverse initialization backbones spanning model families and scales. Across models, our method improves RewardBench v2 accuracy by up to +7.7 points on average, with gains of up to +16.1 on in-domain math subsets and consistent improvements on out-of-domain safety and general subsets. When applied to best-of-N selection and policy optimization, these reward models substantially improve downstream math performance and match or exceed strong supervised reward model baselines of similar size. Furthermore, using an RBS checkpoint as a mid-training initialization substantially amplifies gains from subsequent supervised fine-tuning. Overall, we demonstrate the feasibility and promise of training reward models without costly and potentially unreliable human annotations.
Abstract: , a neural network framework for discovering iterative spectral algorithms for large-scale numerical linear algebra and numerical optimization. Our self-supervised models adapt to input operators using coarse spectral information (e.g., eigenvalue estimates and residual norms), and predict recurrence coefficients for computing or applying a matrix polynomial tailored to a downstream task. The effectiveness of relies on three ingredients: an architecture whose inference pass implements short, executable numerical linear algebra recurrences; efficient training on small synthetic problems with transfer to large-scale real-world operators; and task-defined objectives that enforce the desired approximation or preconditioning behavior across the range of spectral profiles represented in the training set. We apply to discovering algorithms for representative tasks on spd matrices: accelerating matrix function approximation; accelerating sparse linear solvers; and spectral filtering/preconditioning for eigenvalue computations. On real-world matrices, the learned procedures deliver orders-of-magnitude improvements in accuracy and/or reductions in iteration count, relative to basic baselines. We find clear connections to classical theory: the induced polynomials may exhibit equioscillation behavior characteristic of Chebyshev polynomial approximation.
Abstract: Fairness in multi-task learning (MTL) is challenging when heterogeneous output spaces must be controlled through a shared multi-head model. Existing fairness methods are often tied to a single prediction type, while MTL optimizers coordinate task utility without specifying how heterogeneous group disparities should be represented, aggregated, or propagated through task-specific heads. We introduce FairMT, which frames fair heterogeneous MTL through a coordinate-allocation view. FairMT builds a Fairness-Coordinate Interface for binary detection, one-vs-rest multi-class classification, and scalar regression, and instantiates these coordinates asymmetrically using the current best-performing valid group as a detached reference. This directs correction toward groups falling behind the reference, reducing disparities while better preserving task utility. Asymmetric Heterogeneous Fairness Disparity Aggregation (AHFDA) then adaptively allocates constraint pressure over normalized coordinates and converts heterogeneous disparities into a single differentiable constraint signal. Finally, a Head-Aware Optimization Proxy maps the allocated constraint through task-head geometry into task-combination weights for primal--dual training, reducing routing bias caused by heterogeneous head scales, sensitivities, and decision geometries. Experiments on visual and language MTL benchmarks show that FairMT reduces group-level disparities across heterogeneous output types while maintaining utility close to utility-oriented MTL optimizers.
Authors: Elliot Epstein, Apaar Sadhwani, Kay Giesecke
Abstract: Sequence models have advanced Multivariate Time Series (MTS) modeling by learning complex, long-range temporal dependencies. Many practical settings, however, involve predicting the evolution of a set of MTSs -- an unordered collection of units (e.g., a basket of stocks, a pool of loans, a fleet of sensors). This cross-sectional structure offers an opportunity to leverage signals shared across units (e.g., market volatility, correlated defaults) that are missed when each MTS is processed independently. Crucially, treating the entire cross-section as a single high-dimensional MTS is infeasible because the number of units varies dynamically at inference. We propose Set-Sequence, a backbone-agnostic architecture for set-of-MTS prediction that leverages this cross-sectional structure. A permutation-invariant Set module extracts a summary of the population, while a Sequence module (e.g., Transformer/SSM/RNN) then models each unit’s temporal dynamics conditioned on both its own history and the learned cross-sectional context. This architecture naturally accommodates varying set cardinalities, supports unaligned series, integrates with standard sequence backbones, and scales linearly in cross-sectional size. Across a synthetic contagion task and two large-scale real-world applications -- equity portfolio optimization and loan risk prediction -- Set-Sequence consistently improves sequence backbones and outperforms domain-specific baselines, delivering higher Sharpe ratios, improved AUCs, and interpretable cross-sectional summaries.
PaperID: 1212, Poster
Abstract: Probabilistic logical models are a core component of neurosymbolic AI and are important in their own right for tasks that require high explainability. Unlike neural networks, logical theories that underlie the model are often handcrafted using domain expertise, making their development costly and prone to errors. While there are algorithms that learn logical theories from data, they are generally prohibitively expensive, limiting their applicability in real-world settings. Here, we introduce precision and recall for logical rules and define their composition as rule utility -- a measure of the predictive power of logical theories that can be computed in linear time. We then present SPECTRUM, a linear-time framework for learning logical theories from relational data, that mines recurrent patterns in the data, and subsequently, using our utility measure, evaluates and ranks rules derived from these patterns. Finally, we prove theoretical guarantees on the utility of the learnt logical theory. As a result, we demonstrate across various tasks that SPECTRUM scales to larger datasets, often learning more accurate logical theories on CPUs in < 1% of the runtime of SOTA neural network approaches on GPUs.
PaperID: 1213, Poster
Abstract: Width is a central axis for scaling LLM families, but widening a standard transformer typically changes the hidden representation rather than continuing the smaller-width computation. We introduce Trunkfish, a recipe for converting model architectures into a hierarchy of models with incrementally refinable width (a Trunkfish hierarchy) by enforcing causal hidden dimensions and cumulative prefix readouts: new width may depend on old width, but old width may not depend on new width. This enables schooling, which trains all models in a hierarchy from forward/backward passes of only the largest model, and restart-free cascades, which enable adaptively scaling up inference compute without wasting compute already spent on smaller models. Trunkfish improves the validation-loss/active-parameter frontier over independently trained standard hierarchies under matched hierarchy-training FLOPs, with substantial training-FLOP savings for both full-width and prefix-heavy schooling. In cascade simulations over the same trained hierarchy, exact width continuation also meaningfully reduces matched-loss expected active-parameter traffic relative to restart cascades.
PaperID: 1214, Poster
Abstract: Threats and offers are fundamental to strategic interactions between agents. Competently navigating multi-agent settings requires the ability to distinguish between the two, reason about the intentions behind them, and understand how other agents assess them. With the prospect of LLM-based autonomous agents being deployed as representatives or advisors to humans in consequential domains, it is therefore important to understand how LLMs reason about threats and offers. In this paper, we propose a formal theory of threats and offers grounded in causal games, and use it to empirically investigate how current LLMs classify and reason about such proposals. We find that LLMs' threat/offer classifications align with our formal definition, with each model's threat labels tightly tracking its own judgement of whether the proposal leaves the recipient worse off. LLMs also predict their co-players' classifications accurately, though in our setting this accuracy is not clearly above what is already achieved by simply projecting their own labels onto the co-player.
PaperID: 1215, Poster
Abstract: We investigate the quantitative performance of affine-equivariant estimators for robust mean estimation. As a natural stability requirement, the construction of such affine-equivariant estimators has been extensively studied in the statistics literature. We quantitatively evaluate these estimators under two outlier models which have been the subject of much recent work: the heavy-tailed and adversarial corruption settings. We establish lower bounds which show that affine-equivariance induces a \emphstrict degradation in recovery error with quantitative rates degrading by a factor of \sqrtd in both settings. We find that classical estimators such as the Tukey median (Tukey '75) and Stahel-Donoho estimator (Stahel '81 and Donoho '82) are either quantitatively sub-optimal \empheven within the class of affine-equivariant estimators or lack any quantitative guarantees. On the other hand, recent estimators with strong quantitative guarantees are not affine-equivariant or require additional distributional assumptions to achieve it. We remedy this by constructing a new affine-equivariant estimator which nearly matches our lower bound. Our estimator is based on a novel notion of a high-dimensional median which may be of independent interest. Notably, our results are applicable more broadly to \emphany estimator whose performance is evaluated in the \emphMahalanobis norm which, for affine-equivariant estimators, corresponds to an evaluation in \emphEuclidean norm on isotropic distributions.
Abstract: Users increasingly turn to AI systems for normative assistance—guidance on what one ought to do or think—yet models are often opaque about whose viewpoints they represent. A promising approach is simulation-augmented generation (SAGE), which involves querying generative simulations of individuals in a target population at inference time, soliciting their open-ended judgments, and synthesizing them into a response while transparently reporting whose viewpoints are reflected. However, inference-time simulation raises acute scalability constraints. Since the key benefit of simulation is improved representativeness, the core challenge is scaling simulation without sacrificing representation. We introduce the first formalization of this problem, grounded in proportional clustering concepts from social choice theory. We prove that to represent a population of m humans, we need only create n \ll m simulations of them, and need only dynamically query k \ll n of those simulations at inference time, while still maintaining approximate proportional representation guarantees for the full population. We empirically validate that our inference-time algorithm yields better representation--efficiency trade-offs than baseline approaches.
PaperID: 1217, Poster
Abstract: Persistent homology summarizes the topological evolution of a graph as persistence diagrams (PD). Generating graphs whose topology matches a target descriptor is increasingly important, yet existing topology-aware generative models rely on soft regularization, producing graphs whose PDs may only be approximately similar. We initiate a principled investigation of the generation problem: produce graphs that realize a target PD exactly. We propose two approaches guided by the implicit local and global constraints a target PD imposes. The local approach builds degree-based PD-equivalent graphs iteratively, and we prove that every output realizes the target PD and that every degree-based PD-equivalent graph is reachable. However, as shown thereafter, local generation necessarily trades off between sample diversity and reaching infeasible states. We therefore recast generation as a global constraint satisfaction problem with a complete encoding solvable by modern constraint programming solvers. This is applicable to any vertex-based permutation equivariant filtration scheme and is minimal under the degree filtration. Experiments on four standard graph generation benchmarks confirm that both methods realize the target PD exactly. Together, our methods provide the first principled treatment of strict topology-compliant graph generation.
PaperID: 1218, Poster
Abstract: Machine learning relies on gradient-based training procedures whose empirical efficiency is usually analyzed through finite-dimensional arithmetic counts rather than formal complexity over continuous real-valued data. This paper studies neural-network training primitives as represented-real operators in the second-order complexity framework of Kawamura and Cook. Here, second-order polynomial time means computing a real-valued operator to accuracy 2^-n in time polynomial in the requested precision, architecture size, primitive evaluation costs, and supplied analytic certificates. We prove that dense backpropagation, convolutional backpropagation, soft attention, fixed finite SGD/Adam-style updates, and related smooth local primitives are second-order polynomial-time computable over the relevant represented real spaces. More precisely, dense one-sample backpropagation has tight arithmetic complexity \Theta(s) for s weights and polynomial bit complexity; convolutional backpropagation is polynomial in the number of active convolution incidences; dense soft-attention backpropagation has arithmetic cost \Theta(\ell d^2+\ell^2d); and fixed finite smooth optimizer updates remain second-order polynomial-time computable under explicit smoothness and lower-bound certificates. Higher complexity enters through three distinct mechanisms: discontinuous selection, global optimization, and continuous-time flow solving. Hard routing, top-k sparsification, argmax decisions, ReLU kink conventions, exact line search, and idealized gradient flow therefore change the second-order complexity status of the training operator. Overall, this work provides a rigorous mathematical language for precision, conditioning, and certificate dependence, which are often implicit in floating-point analyses, thus enabling a unifying complexity account for dense networks, CNNs, attention, Transformers, and standard optimizers.
Abstract: Frontier AI systems are increasingly natively multimodal, jointly pretrained on multiple modalities, such as text, vision, and audio. However, the scientific understanding of the multimodal data mixtures used to pretrain these systems --- how much of each modality to mix, and at what potential interference cost to the others --- is largely absent or gatekept. In this work, we derive data mixing scaling laws for three modalities: text, vision, and audio. Multimodal Large Language Models (MLLMs) increasingly adopt Mixture-of-Experts (MoE) architectures, as MoEs scale efficiently and partially limit cross-modal interference. We therefore use MoEs to conduct 268 experiments ranging from 457M to 8.3B parameters, on up to 150 billion tokens. Our work leads to multiple actionable insights into multimodal data mixing for training MLLMs. First, compute optima are modality-specific: audio modality is model-heavy, favouring scaling parameters over tokens, while text and vision require balanced scaling of model size and data. Second, repeating vision or audio data beyond 2x yields negligible benefits. Finally, cross-modal interactions are highly asymmetric: vision strongly interferes with audio, while audio provides mild positive transfer to text and vision despite competing for shared capacity. Overall, we provide a principled foundation for understanding the multimodal data mixtures needed to train frontier MLLMs.
Abstract: Diffusion models provide strong priors for generating structured data, but many tasks require outputs beyond the scale on which these models are typically trained. Compositional generation addresses this by composing overlapping local plans from a pretrained short-horizon prior into a long-horizon output. However, standard composition primarily enforces agreement between neighboring local plans, yielding local consistency without directly specifying the global structure of the full composition. As a result, locally compatible plans may still form an implausible route, task sequence, or temporal evolution. Existing methods improve global coherence by repeatedly propagating local consistency signals or by adding inference-time optimization, but these procedures become expensive as the number or dimensionality of local plans increases. We propose Coarse-to-Fine Compositional Diffusion (CoFi), an inference-time sampler that separates global structure formation from local detail refinement. CoFi first aligns local denoised estimates around a shared coarse structure, producing a global scaffold that captures the long-range task-level arrangement. It then diffuses this scaffold to an intermediate noise level and denoises it with the same pretrained local prior, restoring local fine structure while preserving the scaffold-induced global coherence. Across long-horizon robotic planning, panoramic image generation, and long video generation, CoFi not only improves both global coherence and local sample quality over prior compositional baselines, but also requires 2--8× fewer denoiser evaluations.
PaperID: 1221, Poster
Abstract: Model cascades reduce inference cost by routing inputs through cheap models and escalating only when needed, from lightweight image classifiers to small language models. Existing exit rules (confidence thresholds, learned scorers, agreement signals) all optimize for aggregate quality, but none controls quality at the class level: a model that is systematically unreliable on a rare, fine-grained, or minority class silently violates the quality contract, regardless of its overall confidence. We introduce , not the confidence score alone, is the correct signal for cascade exit decisions. For each model, NOMAD identifies from held-out data the classes it handles reliably enough to serve as the final answer; a verifies that any sequence of models preserves per-class quality end-to-end, not just at each stage in isolation; and a greedy selector routes each input through the cheapest viable sequence, with a provable 4-approximation on cost. On 16 datasets spanning tabular, fine-grained vision, and text (including LLM cascades with 15–20× cost ratios between models), NOMAD achieves up to 40× speedup (geometric mean 4× over the role model) and is the only method among 11 baselines with zero per-class violations on every dataset. Code, configurations, fold indices, and an interactive results explorer (
Abstract: Humans can approach complex visual problems by mentally simulating intermediate visual steps, rather than reasoning through language alone. Inspired by this, several works on Vision-Language Models have recently explored chain-of-thought reasoning with continuous latent tokens as intermediate visual imagination steps. In this work, we investigate how recent models leverage such latent tokens. Surprisingly, we find that model accuracy is unaffected when latent tokens are replaced by uninformative dummy tokens. This indicates that latent tokens play a minimal causal role in the model's final prediction. To better understand this phenomenon, we analyze both the training signal provided by oracle latent representations and the quality of the latent tokens generated at inference time. Our experiments reveal two crucial issues holding back latent visual reasoning: First, in most existing datasets, oracle latent tokens provide limited additional information beyond the original image and do not substantially simplify the task, leading models to ignore them during training and effectively bypassing them at inference time. When fine-tuned on a diagnostic dataset, in which latent tokens provide sufficient support for the final prediction, we show that models can causally rely on them. Second, the latent tokens produced at inference time deviate from their corresponding oracle representations, collapsing to a narrow region and preventing benefits even when the model relies on them. Overall, our findings suggest that future progress in latent visual reasoning depends on two key pillars: high-quality datasets with informative intermediate steps and more precise latent token prediction.
Abstract: Vector quantization via random projection followed by scalar quantization is a fundamental primitive in machine learning, with applications ranging from similarity search to federated learning and KV compression. While dense random rotations yield clean theoretical guarantees, they require \Theta(d^2) time. The randomized Hadamard transform HD reduces this cost to O(d \log d), but its discrete structure complicates analysis and leads to weaker or purely empirical compression guarantees. In this work, we study a variant of this approach: \emphdithered quantization with a single randomized Hadamard transform. Specifically, the quantizer applies HD to the input vector and subtracts a random scalar offset before quantizing, injecting additional randomness at negligible cost. This approach is used by popular quantizers such as RaBitQ, and is a slight modification of others, including TurboQuant and EDEN. We prove that this approach provides mean squared error bounds that asymptotically match those achievable with truly random rotation matrices. For example, we prove a dithered version of TurboQuant achieves mean squared error \bigl(\pi\sqrt3/2 + o(1)\bigr) \cdot 4^-b at b bits per coordinate, where the o(1) term vanishes uniformly over all unit vectors and all dimensions as the number of quantization levels grows.
PaperID: 1224, Poster
Abstract: Functional optimization problems are typically solved by optimizing the parameters of a fixed representation, such as a neural network, resulting in highly nonconvex losses that complicate both training and theoretical analysis. An interesting alternative is functional gradient descent (FGD) --- that is, gradient descent directly in function space --- which benefits from strong convergence results and admits a clean theory. However, FGD is difficult to implement in practice because functional gradients are infinite-dimensional, and thus cannot be fully computed nor stored in memory. Existing implementations therefore rely on fixed approximations, which introduce approximation error. We propose a new, theoretically-grounded FGD algorithm that adapts the representation of the functional gradients over the course of optimization. By explicitly incorporating this approximation into the analysis, we establish convergence to a stationary point (for smooth losses) and to a global minimizer (under smoothness + a Polyak-Łojasiewicz-type condition) regardless of our approximations. To the best of our knowledge, this is the first implementable FGD method with such guarantees in a general setting. We demonstrate the effectiveness of our method on regression, numerical solution of PDEs, and modern computer vision. Across settings, our method consistently outperforms both FGD with fixed approximations and neural network baselines in efficiency and accuracy.
Abstract: Symbolic regression aims to uncover explicit scientific laws from data. Recent methods use LLMs to guide mutation from background text, which is more directed than random genetic programming. However, exact symbolic recovery requires both semantic guidance and explicit structure, so that domain-informed search are carried out through valid symbolic representation. Current LLM-driven systems remain structure-blind: they select among opaque candidates, lack explicit mechanisms for local mutation, and rely on brittle coefficient fitting that can undervalue correct skeletons. We propose FunctionEvolve, an evolutionary framework using expression trees to organize the whole search: structural summaries promote diverse parent selection, local tree edits preserve useful subexpressions, and structure-aware fitting decomposes, constrains, and simplifies coefficients for more reliable scoring. It uses only elementary function families, without additional domain-specific rules limiting generalization. On the 129-task synthetic subset of LLM-SRBench, FunctionEvolve with Claude Opus 4.6 recovers 113 exact forms, reaching 87.6% \mathrmSA@50, 4.7x above same-backbone baselines, and 58.9% \mathrmSA@1, 3.8x above the strongest previously published top-1 result. Ablations show that structure-visible search is central to reliable recovery, with LLM-guided refinements and structure-aware coefficient optimization serving as essential proposal and scoring mechanisms. We also audit the benchmark and show that collinearity in its materials-science subset creates identifiability issues.
PaperID: 1226, Poster
Abstract: Designing RNA sequences that satisfy target tertiary-structure objectives is an important problem in RNA engineering, with broad applications in synthetic biology and therapeutics. Existing \mathrm3D RNA design methods have made substantial progress, but they largely treat the target fold as a geometry-only objective. This is incomplete for RNA, whose functional folds are specified not only by tertiary geometry but also by secondary base-pairing topology. We therefore formulate RNA inverse design as a topology-and-geometry co-design problem and propose \textttTANGO, a cooperative multi-agent reinforcement learning framework for this task. \textttTANGO adopts a staged curriculum: it first learns a topology-aware policy from secondary-structure targets, then transfers this policy to tertiary-structure targets to train a general co-design model, and finally applies frontier-guided fine-tuning to refine target-specific trade-offs among secondary-structure fidelity, tertiary-structure fidelity, and sequence diversity. Experiments show that \textttTANGO generates diverse and novel designs, improving tertiary- and secondary-structure performance by 18.3% and 52.7%, respectively, over the strongest baseline.
Abstract: Diffusion models are powerful generative models that produce high-quality samples from complex data. While their infinite-data behavior is well understood, their generalization with finite data remains less clear. Classical learning theory predicts that generalization occurs at a sample complexity that is exponential in the dimension, far exceeding practical needs. We address this gap by analyzing diffusion models through the lens of data covariance spectra, which often follow power-law decays, reflecting the structure of real data. To understand whether such a power-law structure can benefit learning in diffusion models, we develop a theoretical framework based on linear neural networks, congruent with a Gaussian hypothesis on the data. We quantify how the covariance spectra of data and regularization impact generalization. We find two regimes: When N
PaperID: 1228, Poster
Authors:
Siming Zhang, Zhehui Shen, Shijie Chen, Xinle Gu, Yansen Yu, Hang YuAbstract: Itemwise calibration does not guarantee calibration of the decision it feeds. Many systems output one score per item in a partially observed group, but the downstream action asks whether enough hidden items cross a threshold: show a slate, defer a panel decision, or inspect a batch. If those hidden items share an unresolved user, patient, or batch state, averaging that state before aggregation preserves means but makes the aggregate posterior too concentrated. We call this failure the serialization tax. For Bernoulli de Finetti posteriors, the exact hidden-count law dominates any independent completion with the same total predictive mean in convex order, yielding variance, total-variation, threshold, and Bayes-value gaps. The same interface loss extends to non-binary outcomes through empirical-measure Laplace functionals and to heterogeneous slates through a shared-covariance identity. TaxScore turns the theory into a routing rule: use cheap marginal interfaces when hidden-block ambiguity is far from the decision boundary, and propagate a shared posterior sample when it is near. On pre-specified, hidden-label-free MovieLens 1M/10M audits, replacing Transformer count heads with existing stochastic shared-latent heads improves high-TaxScore negative log-likelihood (NLL)/utility from 0.673/0.696 to 0.629/0.723 and from 0.670/0.691 to 0.639/0.715; routing richer heads only on flagged slates gives 0.006-0.012 full-population utility gains.
Abstract: For LLM agents to be useful beyond simple chat, they must hold sensitive user context such as calendars, credentials, health records, and financial data. We study whether the mere presence of such secrets in a model's context window introduces hidden correlations into the model's benign outputs, allowing reconstruction even when the model correctly refuses direct extraction. We further study whether an adversary can actively engineer prompts that amplify this effect, using the model as a covert carrier to transmit secrets through seemingly innocuous text. In both cases, this limited leakage is exploited using a novel attack that assumes black-box access to the underlying model. In controlled experiments across eight proprietary models, we find that 2-digit in-context secrets are reconstructed with near-perfect accuracy and 4-digit secrets at 82% exact match, all from outputs the model produces in response to ordinary, non-adversarial requests. More capable models leak more: stronger instruction-following amplifies sensitivity to in-context secrets, making leakage a byproduct of capability rather than a bug to be patched. We show this leakage enables two practical attacks: (1) a trained classifier that infers semantic predicates about user memories (e.g., health conditions, financial events) from routine natural-language outputs, and (2) an RL-trained adversary that extracts full Social Security Numbers from a production-style agent at 100% success.
Abstract: We study secret elicitation: discovering knowledge that an AI possesses but does not explicitly verbalize. As a testbed, we train three families of large language models (LLMs) to possess specific knowledge that they apply downstream but deny knowing when asked directly. For example, in one setting, we train an LLM to generate replies that are consistent with knowing the user is female, while denying this knowledge when asked directly. We then design various black-box and white-box secret elicitation techniques and evaluate them based on whether they can help an LLM auditor successfully guess the secret knowledge. Many of our techniques improve on simple baselines. Our most effective techniques (performing best in all settings) are based on prefill attacks, a black-box technique where the LLM reveals secret knowledge when generating a completion from a predefined prefix. Our white-box techniques based on logit lens and sparse autoencoders (SAEs) also consistently increase the success rate of the LLM auditor, but are less effective. We release our models and code, establishing a public benchmark for evaluating secret elicitation methods.
Authors:
Hongchen Wei, Yuanzhe Wang, Bei Liu, Yifan Yang, Qi Dai, Kai Qiu, Yunsheng Li, Dongdong Chen, Chong Luo, Zhenzhong Chen, Baining GuoAbstract: Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing retrieval-augmented systems usually select evidence from a static index before generation, while recent agentic systems add multi-turn tool use but often rely on frozen proprietary backbones whose behavior is set by prompts. We present DocAtlas, a system that treats long-document understanding as a mutable-state information-seeking process. We instantiate DocAtlas as a mutable document harness: an external environment that determines what document information is searched, read, stored, reviewed, and shown to the model at each step. Given a document and question, the harness exposes search, reading, note-taking, and review tools, maintains a hierarchical tree and note store, and updates both as the agent records evidence. DocAtlas combines self-improving retrieval, selective evidence access, and active working memory under a fixed context budget. The same harness supports inference-time use with large VLMs and end-to-end reinforcement learning for compact VLM agents. With GPT-5.4, DocAtlas reaches 71.4% on MMLongBench-Doc, exceeding the human-expert reference of 65.8%. A Qwen3.5-4B VLM trained with end-to-end RL in the DocAtlas environment reaches 63.7%, compared with a 54.4% direct-input baseline, showing that mutable document-harness design can improve compact document agents by a large margin.
Abstract: We consider the scheduling problem of online speed scaling where the goal is to minimize the energy consumption of a machine that controls the speed at which jobs are processed. Recent work has leveraged the learning-augmented framework, where the algorithm is provided with predictions about jobs that will arrive in the future, to manage power usage more efficiently. This paper proposes a novel prediction model for speed scaling where the predictions are about the machine speed (the output), instead of the jobs (the input). Machine speed predictions have multiple advantages: they are succinct, admit strong PAC-learnability guarantees, can be provided dynamically, and lead to a natural definition of smoothness. We give an algorithm for dynamic machine speed predictions that is (1+\epsilon)-consistent and O(1)-robust. For offline machine speed predictions and job speed predictions, we provide an algorithm that achieves the stronger guarantee of (1+\epsilon)-smoothness, while maintaining O(1)-robustness. These guarantees are comparable to previous work, but do not require predicting the entire input.
PaperID: 1233, Poster
Abstract: Sparse autoencoder (SAE) training is often bottlenecked by activation storage. We show that the stored activation buffer can be far smaller than the total SAE training budget: quality depends mainly on the number of \emphunique activations, not on whether every training token is fresh. After a short diversity-limited regime, additional fresh activations provide little benefit; replaying the same buffer preserves both reconstruction and interpretability performance. We capture this diversity--repetition tradeoff with a data-constrained scaling law and validate it across dictionary sizes, token budgets, Llama-3.1 and Qwen3 models, early/middle/late layers, BatchTopK/TopK/JumpReLU SAEs, and downstream metrics including EV, reconstruction MSE, feature absorption, automated interpretability, and sparse probing. The resulting buffer-sizing rule is simple: choose the smallest activation buffer that reaches the quality plateau, then use replay to spend the remaining training budget. In our experiments, this reduces activation storage by 8\text--64× with negligible quality loss.
PaperID: 1234, Poster
Abstract: Recent advances in 3D object generation have enabled the creation of high-fidelity 3D assets from 2D images. This capability offers a promising path for alleviating 3D data scarcity, which remains a key bottleneck for scalable simulation in robotics and AR/VR applications. However, existing 3D generation methods primarily emphasize visual plausibility and geometric fidelity, with limited consideration of object-level interaction and structural understanding that are critical for downstream embodied AI tasks. To bridge this gap, we introduce , a unified framework for embodied 3D object understanding built upon latent representations produced by modern generative models such as SAM3D. We first systematically analyze the part-level locality of these latent representations against existing unified 3D point encoders and show that they naturally preserve rich structural semantics suitable for object understanding. Motivated by this observation, we then design a unified DiT-style decoder that operates directly on the latent representation and supports multiple object-level tasks, including open-vocabulary 3D affordance prediction, 3D part segmentation, and 3D articulation estimation. Given a real-world 2D image containing an object of interest, generation beyond visual reconstruction by jointly reasoning about affordance, articulation, and part semantics. This transforms generated 3D assets into actionable object representations suitable for downstream embodied AI applications. Our source-code will be open-sourced.
Abstract: When language models lack relevant knowledge for a given query, they frequently generate plausible responses that can be hallucinations, rather than admitting being agnostic about the answer. Retraining models to reward admitting ignorance can lead to overly conservative behaviors and poor generalization due to scarce evaluation benchmarks. We propose a post-hoc framework, Conformal Abstention (CA), adapted from conformal prediction (CP) to determine whether to abstain from answering a query. CA provides finite-sample guarantees on both the probability of participation (i.e., not abstaining) and the probability that the generated response is correct. Importantly, the abstention decision relies on prediction confidence rather than the non-conformity scores used in CP, which are intractable for open-ended generation. To better align prediction confidence with the model's ignorance, we introduce a calibration strategy using representation geometry within the model to measure knowledge involvement in shaping the response. Experiments demonstrate that we improve selective answering significantly with 75% conditional correctness.
PaperID: 1236, Poster
Authors: Minh D Le, Minh-Duong Nguyen, Dung Le
Abstract: Meta pseudo-labeling improves semi-supervised learning by updating a teacher according to whether its pseudo-labels help a student improve on labeled data. Existing one-step meta-gradient methods, however, differentiate through the student update, introducing mixed higher-order derivatives, extra memory cost, and noisy minibatch meta-gradients. We propose Trajectory-Matching Meta Pseudo-Labeling (TMPL), a first-order alternative that extracts labeled-data feedback from recent student optimization trajectories. We show that the one-step meta objective is locally equivalent, up to second-order terms in the student step size, to maximizing alignment between the outer gradient and the pseudo-label-induced student gradient. TMPL approximates this alignment without differentiating through the student update by matching finite differences of the outer objective and the pseudo-label-induced student objective along recent student displacements. Its implementation stores only a FIFO cache of detached scalar objective values, rather than student computation graphs or past student networks. We prove that small trajectory-matching error controls the gradient mismatch within the span of recent displacements, with an explicit residual for unexplored directions. On CIFAR10, CIFAR100, SVHN, and STL10, TMPL improves over MPL in all eight evaluated label regimes, obtains the best result in five settings, and remains competitive with strong SSL baselines. On CIFAR100 with 2500 labels, TMPL reduces peak memory by 39.7% relative to MPL and 54.0% relative to a raw differentiable meta-gradient implementation. Ablations identify the directional trajectory-matching term as the main source of improvement, while alignment diagnostics show that TMPL recovers MPL-like gradient alignment after a short warm-up.
Authors:
Qihan Wang, Nicholas Tomlin, Michael Hu, Brian Dillon, Tal LinzenAbstract: Language models are increasingly being deployed as user simulators, but their memory is far more reliable than that of real users. To measure this gap, we run a series of classic memory experiments from psychology on both humans and language models. Across tasks, we find that out-of-the-box language models exhibit better memory than humans, even when prompted to imitate human behavior. We then show that better prompting strategies and the use of a compactor can cause language models to forget content in a more human-like way. Using these methods, we show preliminary evidence that language models with human-like memory constraints can function as more effective user simulators in a downstream education task. Finally, we release human reference data and benchmarks to support future work on simulating human memory with language models.
Abstract: Multi-modal models require visual tokenizers that jointly capture high semantic density and fine-grained structural fidelity. Representation Autoencoders (RAEs) offer a promising direction by decoding images directly from a pre-trained semantic space using high-dimensional continuous tokens, bypassing the information bottleneck of VAEs and benefiting both visual understanding and generation. In this work, we aim to discretize these continuous vision representations to bridge the gap with language models — a non-trivial challenge, as existing quantization methods suffer from codebook collapse, failing to scale while preserving semantic coherence. We identify the root cause as metric mismatch: standard Euclidean codebook objectives are fundamentally misaligned with the anisotropic geometry of representation space, leading to codebook embeddings with high-variance magnitude scales and uneven angular distributions that hinder scalability. To address this, we propose Hyper-Spherical Quantization (HSQ) , which decouples semantic content from feature magnitude via angular routing, preventing code assignment from being dominated by scale rather than meaning. The resulting discrete Representation Autoencoder (dRAE) achieves high-fidelity reconstruction while preserving semantic integrity and supporting scalable discrete latent spaces. Extensive experiments demonstrate consistent performance gains as the vocabulary size scales to 131,072, along with strong unified performance across both visual understanding and generation benchmarks.
Abstract: Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in multi-step reasoning and calling search engines at appropriate steps. However, existing retrieval-augmented reasoning approaches rely on separate retrieval models, limiting the LRM's role in retrieval to deciding when to retrieve and how to formulate queries, even with its inherent ability to handle vast knowledge spaces. This structural separation induces a representation bottleneck, as the retriever’s latent space often lacks the expressive granularity required to satisfy the generator’s sophisticated information needs. To address this, we shift our perspective on retrieval from sequence-to-sequence matching for all corpora to locating the answer-containing paths within the corpus, and propose a novel framework called FREESON (Retriever-FREE Retrieval-Augmented ReaSONing). This framework enables LRMs to directly access external knowledge by acting as both a generator and a retriever. To achieve this, we introduce a variant of the MCTS algorithm specialized for the retrieval task, which we call CT-MCT (Corpus-Traversing Monte Carlo Tree Search). Through this algorithm, the LRM selectively references specific segments identified during traversal, instead of fetching a fixed top-k set. Experiments on five open-domain QA benchmarks covering both single-hop and multi-hop questions demonstrate that FREESON achieves an average improvement of 14.4% in EM and F1 over four multi-step reasoning models with a separate retriever, and it also performs comparably to the strongest baseline, surpassing it by 3% on PopQA and 2WikiMultihopQA, and by 12% on the fact-checking benchmark FEVER.
PaperID: 1240, Poster
Abstract: Multimodal large language model (MLLM) agents increasingly operate over long histories of dialogues, images, and evolving facts, yet finite context windows and weak persistent memory cause forgetting, temporal inconsistency, and loss of precise past evidence. Existing multimodal memory agents either discard visual evidence by reducing images to captions or treat memory as a static store that is built once and never revised during reasoning; self-evolving agents further accumulate experience via prompting without tight coupling to retrieval and memory revision. Closing this loop with reinforcement learning additionally suffers from credit assignment difficulties over long, multimodal action trajectories. To address these challenges, we propose EvoMM, a reinforced self-evolving multimodal agentic memory framework that unifies memory construction, retrieval, and evolution within a single agentic loop. EvoMM maintains a mutable two-level memory of turn-level short memories and session-level long summaries, organized by a heterogeneous multimodal memory graph with temporal, semantic, image-co-reference, and keyword edges. An iterative PLAN→RETRIEVE→REFLECT→ANSWER policy then revises long memory on the fly and distills structured retrieval experience into a cross-query store for reuse. To optimize this policy, we further introduce PRR-GRPO (Planning-Retrieval-Reflection GRPO), which augments outcome rewards with gated, action-aware process rewards and performs dense credit assignment over a memory-action tree with dual-scale advantage normalization. Experiments on Mem-Gallery, MMLongBench, and LoCoMo show that EvoMM consistently outperforms strong text-only, multimodal, and agentic memory baselines, establishing a training paradigm for memory that reads, revises, and evolves with each query.
PaperID: 1241, Poster
Abstract: Building on SODA 2026 work Kearns et al. (2026), we study learning in directed acyclic graphs in which each agent observes only a subset of the input coordinates and may use the predictions of its parents as additional features. In the finite-sample protocol, the message sent along each edge is the agent's prediction vector on a shared sample, so the communication cost is linear in the sample size m. In this paper we study this model under communication constraints. First, we revisit the finite-sample analysis of the full-message protocol. We identify a gap in the finite-sample proof of Kearns et al. (2026) and develop a different argument that establishes a finite-sample population-risk guarantee along paths satisfying deterministic M-coverage. Second, we introduce a shared-sketch protocol in which each agent sends a k-dimensional linear sketch of its prediction vector instead of the full m-dimensional message, where k\ll m. We analyze this protocol through a reduction to least squares on the compressed sample together with an oblivious subspace embedding argument. The resulting bound separates statistics from communication: the statistical term is still governed by the original sample size m, while the additional error due to compression is controlled by the sketch dimension k. For Gaussian sketches, we obtain an explicit high-probability population-risk bound, showing that row subsampling is not the only possible communication-statistics tradeoff for networked information aggregation.
PaperID: 1242, Poster
Abstract: Contrastive vision-language models map visual and textual representations in a shared normalized embedding space, making cosine similarity the natural metric for cross-modal retrieval. A practical challenge arises during model upgrades: independently trained models generally produce incompatible representation spaces, so replacing a deployed model typically requires recomputing embeddings for the entire gallery, which is prohibitively expensive at scale. Orthogonal post-hoc alignment mitigates this problem by mapping new-model queries into the old-model gallery space while preserving the new learned representation. However, because independently trained models can differ in fine-grained representation structure, the orthogonal alignment remains approximate, leaving a residual angular discrepancy between the old-model query and the aligned new-model query. We study whether spherical linear interpolation (SLERP) between these two normalized query representations can improve retrieval without re-indexing the gallery. We formalize the geometric conditions under which post-alignment interpolation yields a query direction closer to a task-optimal direction than either endpoint, and connect this characterization to Recall@K through a local margin-based certification result. Experiments across multiple benchmarks and model families show that SLERP improves over orthogonal alignment alone, suggesting these conditions are broadly met in practice.
PaperID: 1243, Poster
Authors: Manar Hamed, Yassine Yaakoubi
Abstract: Training-memory reduction requires identifying which saved tensors drive peak usage and ensuring high-fidelity backward contractions from compressed states. We formalize saved-activation compression as a target-surface prediction problem: a pre-training framework to estimate the replaceable memory surface, computational overhead, and parameter-gradient error. We introduce STAC-R (Subspace Tracked Activation Compression with Residuals), a saved-tensor wrapper storing activations as persistent low-rank coefficients paired with a sparse quantized residual. Weight gradients are contracted directly from these factors, bypassing full decompression. To guide deployment, we propose the Saved-Surface Predictor (\mathrmSSP = A_C^+ / P), a statistic that upper-bounds peak memory reduction. We also derive a gradient bridge linking reconstruction error to parameter-gradient cosine similarity, determining if a low-precision state preserves update fidelity.Evaluations across GPT-2 (Medium/Large) and Pythia-1.4B pre-training show STAC-R reduces peak memory by 7.2-12.4% while keeping long-horizon validation perplexity within 1.5% of baseline (Pythia-1.4B 10k steps: +1.14%) with mean gradient cosine \geq 0.98. On BERT fine-tuning (SST-2), STAC-R achieves 12.45% memory reduction with a 0.15pp accuracy decrease. Mechanistic benchmarks yield a mean gradient cosine of 0.9813, outperforming iso-byte SVD (0.9236) and untracked-subspace (0.6949) controls. On the empirical rate-distortion frontier, STAC-R is Pareto-dominant over controls across stored-byte budgets. Composed with GaLore, a "Turbo-fast" variant matches full STAC-R's memory footprint within 0.005 GiB while reaching 0.67𝑥 wall time. Finally, our analysis identifies memory-increasing regimes in SwiGLU architectures (Qwen-2.5/Llama-3.2), establishing STAC-R as a compressor with predictable memory and fidelity conditions checkable before training.
PaperID: 1244, Poster
Abstract: Estimating a secondary signal (e.g., behavior) from neural activity over time is central to both causal online decoding and non-causal offline inference in neuroscience, yet existing two-signal latent state-space models rarely support both. In this work, we provide an analytical extension of a linear method (PSID) beyond causal prediction to also support non-causal inference. We provide theoretical derivations extending PSID to enable optimal filtering and optimal smoothing of the secondary signal. We show that, in the PSID setting, the presence of a secondary signal increases identifiability. This allows us to uniquely learn the quantities needed for the optimal Kalman update via a reduced-rank regression step, yielding our first contribution, . In simulations, we validate that both PSID with filtering and smoothing reach ideal performance. In non-human primate motor cortex data, PSID with smoothing consistently improves over PSID with filtering, which improves over one-step-ahead prediction with standard PSID. Finally, we show that the connection between PSID and an existing nonlinear two-signal model, DPAD, extends naturally to the smoothing setting: applying our forward-backward formulation to DPAD yields , which achieves competitive performance with leading methods on the Neural Latents Benchmark (NLB) in behavior decoding and held-out neural prediction across three datasets. Together, this work provides a theoretical foundation for prediction, filtering, and smoothing in the two-signal setting, spanning causal online decoding to offline inference, in both linear and nonlinear settings.
PaperID: 1245, Poster
Abstract: Recent work on hallucination detection in large language models has shown that, for a fixed pre-trained model and reasoning task, it is possible to estimate the model’s confidence in the correctness of its outputs. Such uncertainty estimates have primarily been used to improve truthfulness by detecting or filtering confabulations. In this work, we ask whether these signals can instead be used more proactively to directly improve the accuracy of model-generated answers. We propose USteer, a simple, training-free steering mechanism that adjusts a model’s layer-wise activations during inference using the gradient of a confidence with respect to the activations. This procedure nudges generation toward outputs with lower uncertainty at inference time, without modifying model parameters or requiring additional supervision. We show that this approach consistently reduces hallucination across a range of tasks, demonstrating that confidence signals can be leveraged not only for detection, but also for effective inference-time control of model behavior.
Authors:
Saul Santos, Nuno Gonçalves, Daniel McNamee, Marcos Treviso, André MartinsAbstract: Recent work has revealed a link between self-attention mechanisms in transformers and test-time kernel regression via the Nadaraya-Watson estimator, with standard softmax attention corresponding to a Gaussian kernel. However, a kernel-theoretic understanding of sparse attention mechanisms is currently missing. In this paper, we establish a formal correspondence between sparse attention and compact (bounded support) kernels. We show that normalized ReLU and sparsemax attention arise from Epanechnikov kernel regression under fixed and adaptive normalizations, respectively. More generally, we demonstrate that widely used kernels in nonparametric density estimation---including Epanechnikov, biweight, and triweight---correspond to \alpha-entmax attention with \alpha = 1 + \frac1n for n \in \mathbbN, while the softmax/Gaussian relationship emerges in the limit n \to \infty. This unified perspective explains how sparsity naturally emerges from kernel design and provides principled alternatives to heuristic top-k attention and other associative memory mechanisms. Experiments with a kernel-regression-based variant of transformers---Memory Mosaics---show that kernel-based sparse attention achieves competitive performance on language modeling, in-context learning, and length generalization tasks, offering a principled framework for designing attention mechanisms.
PaperID: 1247, Poster
Abstract: Large language model (LLM) inference serving is undergoing rapid growth and large-scale deployment, motivating us to rethink how the inference process itself can be leveraged to enable efficient task-specific LLM adaptation. In this paper, we propose a forward-only approach for efficient LLM adaptation with forward-only passes in LLM inference serving. We exploit the angle concentration of activations induced by each singular value decomposition (SVD) component to measure its contribution— —to the angle concentration of the hidden states. Based on this, our forward-only search then efficiently identifies the weight matrix with the highest dispersion that merits rank reduction. We then selectively remove the higher-order components and retain the lower-order components in SVD. Empirical results across diverse datasets demonstrate the competitive accuracy of our forward-only approach, while theoretical analysis shows lower peak memory and greater speedup than the gradient-based approach, and greater speedup than the exhaustive search approach. The extended experiments further present the robustness and generalization of our forward-only approach to various LLMs with up to 57B parameters. The code is included in the supplementary material.
PaperID: 1248, Poster
Abstract: Audio--language models (ALMs) introduce a new jailbreak surface in which harmful requests can be delivered through speech. Existing safeguards rely on audio-side filters or guard models, leaving internal safety behavior largely uncontrolled. We instead pursue a representation-level alternative that controls refusal--compliance behavior inside the shared language-model backbone. However, our empirical study shows that direct audio-side steering is noisy and ineffective, consistent with activation analyses indicating larger acoustic and front-end variation in speech-derived representations. Based on this observation, we propose CrossSteer, a cross-modal steering method that learns a cleaner semantic safety direction from text-only preference pairs and transfers it to audio through the aligned shared residual stream. CrossSteer fits a residual-stream direction whose intervention shifts harmful-request generation from unsafe compliance toward safe refusal, and applies this direction at an audio-robust layer. Across three ALM backbones, CrossSteer consistently reduces audio jailbreak attack success rate while largely preserving benign utility, demonstrating cross-modal transfer of text-derived safety steering to the audio channel. Additional experiments show that CrossSteer composes with existing safeguards, further improving robustness as a complementary representation-level safety layer.Our anonymized source code is available at:~\urlhttps://anonymous.4open.science/r/CrossSteer-69F6.
Abstract: Skills are a promising way to improve LLM agent capabilities without retraining, while keeping the added procedure reusable and controllable. However, high-quality skills are still largely written by hand. We introduce SkillGen, a multi-agent framework that synthesizes a single auditable skill from trajectories generated by a base agent. The output is a human-readable artifact that can be inspected before use. Rather than merely summarizing trajectories, SkillGen leverages contrastive induction over both successful and failed trajectories to identify reusable success patterns, recurring failure modes, and behaviors that appear in nearby successes but are missing from failures. SkillGen then generates candidate skills and iteratively refines the skill. A key novelty in SkillGen is that we model agent skills as interventions to empirically verify the net effect of skills on the overall performance. Specifically, we compare outcomes on the same instances with and without the skill, so that we account for both repairs (cases where the skill fixes a baseline failure) and regressions (cases where the skill breaks a baseline success). Across a broad range of agents and datasets, SkillGen consistently improves held-out performance, outperforms existing skill-generation baselines, and produces skills that transfer across models.
Authors:
Yorgos Felekis, Michael O'Riordan, Oriol Corcoll Andreu, Ciarán LeeAbstract: Predictive models trained on observational data often fail to generalise to the distributions they encounter when deployed, especially when the training data is a product of the system being optimised. Recommender systems are a canonical example: they are trained on interaction logs confounded by the deployed policy, past user behaviour, and platform filtering. As a result, the training distribution differs substantially from the candidate distribution scored at serving time, a gap that makes offline metrics unreliable predictors of online performance. We address the distribution shift problem with a method motivated by causal representation learning (CRL). We propose an information-theoretic disentanglement criterion and prove that its optimum depends only on the causal components of the input. We then derive a tractable variational lower bound that makes the criterion optimisable from finite observational data alone. The scope of our method is narrower than that of much of the CRL literature, in that we target better generalisation under distribution shift, not full identification of all latent causal factors. This narrower target is what makes the method practical, requiring only the existing confounded logs, applying to any standard supervised model, and adding no inference-time cost. Our headline evaluation is an A/B test with millions of users on a music streaming platform, applied to a production ranker for personalised playlist generation. A CRL variant matched in capacity to the production baseline performed on par offline but delivered substantial online gains in listener engagement. Complementary evidence on the public KuaiRand recommendation dataset and a synthetic benchmark with known causal structure shows the same pattern: offline parity with baseline, gains under distribution shift. Across all three settings, adding our causal disentanglement objective yields meaningfully better out-of-distribution generalisation.
Abstract: We study the problem of learning generative models for discrete sequences in a continuous embedding space. Whereas prior approaches typically operate in Euclidean space or on the probability simplex, we instead work on the sphere \mathbb S^d-1. There the von Mises-Fisher (vMF) distribution induces a natural noise process and admits a closed-form conditional score. The conditional velocity is in general intractable. Exploiting the radial symmetry of the vMF density we reduce the continuity equation on \mathbb S^d-1 to a scalar ODE in the cosine similarity, whose unique bounded solution determines the velocity. The marginal velocity and marginal score on (\mathbb S^d-1)^L both decompose into posterior-weighted tangent sums that differ only by per-token scalar weights. This gives access to both ODE and predictor-corrector (PC) sampling. The posterior is the only learned object, trained by a cross-entropy loss. Experiments compare the vMF path against geodesic and Euclidean alternatives. The combination of vMF and PC sampling significantly improves results on Sudoku and language modeling.
Authors: Christopher Nemeth
Abstract: Hypergraphs model higher-order interactions, but realistic hypergraph generation remains difficult because incidence, hyperedge-size heterogeneity, and overlap structure are not faithfully captured by pairwise reductions. We propose HEDGE, a generative model defined directly on relaxed incidence matrices via a structured stochastic diffusion. The forward process combines a hypergraph-specific two-sided heat operator with an Ornstein--Uhlenbeck component, preserving structure-aware noising near the data while yielding an explicit Gaussian terminal law. Conditional on an observed hypergraph, this forward process is linear-Gaussian, so conditional means, covariances, scores, and reverse-drift targets are available in closed form. We therefore learn a permutation-equivariant state-only reverse-drift field in incidence space by regressing onto exact conditional targets, and generate samples by simulating a learned reverse-time SDE from the Gaussian base law. We establish exactness in the ideal state-only setting together with finite-horizon stability guarantees, and empirically show improved hypergraph generation quality relative to strong baselines.
Abstract: Large-scale pretraining has made Vision-Language-Action (VLA) models promising foundations for generalist robot manipulation, yet adapting them to downstream tasks remains necessary. However, the common practice of full fine-tuning treats pretraining as initialization and can shift broad priors toward narrow training-distribution patterns. We propose PriorVLA, a novel framework that preserves pretrained priors and learns to leverage them for effective adaptation. PriorVLA keeps a frozen Prior Expert as a read-only prior source and trains an Adaptation Expert for downstream specialization. Expert Queries capture scene priors from the pretrained VLM and motor priors from the Prior Expert, integrating both into the Adaptation Expert to guide adaptation. Together, PriorVLA updates only 25% of the parameters updated by full fine-tuning. Across RoboTwin 2.0, LIBERO, and real-world tasks, PriorVLA achieves stronger overall performance than full fine-tuning and state-of-the-art VLA baselines, with the largest gains under out-of-distribution (OOD) and few-shot settings. PriorVLA improves over \pi_0.5 by 11 points on RoboTwin 2.0-Hard and achieves 99.1% average success on LIBERO. Across eight real-world tasks and two embodiments, PriorVLA reaches 81% in-distribution (ID) and 57% OOD success with standard data. With only 10 demonstrations per task, PriorVLA reaches 48% ID and 32% OOD success, surpassing \pi_0.5 by 24 and 22 points, respectively.
Abstract: Agentic reinforcement learning (RL) for Large Language Models (LLMs) critically depends on the exploration capability of the base policy, as training signals emerge only within its in-capability region. For tasks where the base policy cannot reach reward states, additional training or external guidance is needed to recover effective learning signals. Rather than relying on costly iterative supervised fine tuning (SFT), we exploit the abundant action data generated in everyday human interactions. We propose \textscActGuide-RL, which injects action data as plan-style reference guidance, enabling the agentic policy to overcome reachability barriers to reward states. Guided and unguided rollouts are then jointly optimized via mixed-policy training, internalizing the exploration gains back into the unguided policy. Motivated by a theoretical and empirical analysis of the benefit-risk trade-off, we adopt a minimal intervention principle that invokes guidance only as an adaptive fallback, matching task difficulty while minimizing off-policy risk. On search-agent benchmarks, \textscActGuide-RL substantially improves over zero RL (+10.7 pp on GAIA and +19 pp on XBench with Qwen3-4B), and performs on par with the SFT+RL pipeline without any cold start. This suggests a new paradigm for agentic RL that reduces the reliance on heavy SFT data by using scalable action guidance instead.
Abstract: Causal discovery, the problem of inferring the direction of causality, is generally ill-posed. We use the language of structural causal models (SCM) to show that assuming that the causal relations are acyclic and invariant across multiple environments (e.g., the way minimum wage affects employment rate is stable across different geographical regions), only two auxiliary environments are sufficient to infer the causal graph for arbitrary nonlinear mechanisms. Moreover, we demonstrate that this implies identifiability of the SCM functional mechanisms: as a corollary, we show that two auxiliary environments are sufficient to guarantee correct counterfactual inference. We empirically support our theoretical results on synthetic data.
PaperID: 1256, Poster
Abstract: A learning-to-defer (L2D) system decides, for each input, whether to predict on its own or to hand it to one of several available experts. The very well established recipe trains classifier and router jointly by treating the K classes and J experts as competing actions in one shared (K+J)-action geometry. Subsequent work has proposed a series of incremental fixes within this geometry; we show that each still suffers, to varying severity, from an optimization-level pathology (target distortion, gradient amplification, winner-take-all starvation, set-mass collapse, or class--expert coupling) even under statistical consistency. We step outside the augmented-action family entirely and propose a _decoupled surrogate_: a softmax classifier head and an independent sigmoid head per expert, mirroring the two natural objects of the problem. We show that per-sample updates are then coordinatewise and the class--expert Hessian block is identically zero, and prove an excess-risk bound with calibration constant \max\\2\sqrt2,\sqrt2J/\lambda\\---to our knowledge the first multi-expert L2D guarantee whose constant does not grow with the expert pool when the per-expert weight is held fixed. On controlled synthetic studies and on CIFAR-10, CIFAR-10H, and Covertype, it is the only method in our comparison that remains stable as the expert pool grows, preserves rare specialists, and improves over a standalone classifier on every real-data benchmark.
PaperID: 1257, Poster
Authors:
Sofiène Boutaj, Pierre Marza, Varun Belagali, Dimitris Samaras, Maria Vakalopoulou, Stergios ChristodoulidisAbstract: Multi-teacher distillation allows transferring knowledge from multiple “teacher” networks to a single “student” network. This method is promising in fields such as digital pathology, where many powerful foundation models were proposed recently. However, as we show in this paper, the standard multi-teacher distillation approach reacts poorly to an increase in the number of teachers used to train a student encoder. We demonstrate that learning teacher-specific representations is a key to scaling in MTD. Importantly, different design choices, such as learnable teacher tokens, a tailored attention scheme, additional mixture-of-experts layers and a contrastive loss, are proposed to better learn such teacher-specific representations. We show that our method better scales with respect to the number of considered teachers, allowing us to train compact student encoders, up to 10X more efficient than larger teacher foundation models, while matching their performance and even outperforming them on a large set of tile-level and slide-level tasks. We conduct a thorough empirical validation, evaluating more than 10 foundation models on 39 tasks spanning tile and slide levels. As a result, we release a new collection of strong and efficient foundation models, named OMNI, trained from the knowledge of 10 state-of-the-art foundation models. We finally conduct an analysis of the learned teacher-specific representations, highlighting their complementarity and explaining why they can be easily aggregated at downstream time.
Abstract: Generalist robot policies must ground language in visual observations, anticipate how actions change the scene, and reason about task goals. Existing VLA approaches typically address these capabilities separately: VLM-based policies emphasize semantic grounding, video or world models emphasize visual dynamics, and goal-conditioned methods emphasize future-state reasoning. This separation can make policies brittle under instruction shift and environment or dynamics perturbations. We introduce Dynin-Robotics, an omnimodal masked-diffusion VLA model that represents robot trajectories as partially observed multimodal token sequences. A single denoising backbone is trained to unify action generation, future-state and world modeling, trajectory-to-language goal understanding, and instruction-conditioned goal-state prediction. At inference time, the same model performs policy generation through block-wise action denoising, with optional goal-state prediction, and world-action joint decoding. We first present diagnostic analyses showing that VLM-based and video/world-model-based policies exhibit complementary failure modes, motivating a unified formulation. Our experiments show that Dynin-Robotics achieves competitive standard-task performance on LIBERO, while improving robustness on instruction-shift and perturbation settings in LIBERO-Plus and VLABench.
Authors:
Xintong Li, Sha Li, Rongmei Lin, Hongye Jin, Linwei Li, Hejie Cui, Sarah Zhang, Chia-Yuan Chang, Kewei Cheng, Yuwei Zhang, Qingyu Yin, Jingbo Shang, Bing YinAbstract: Large reasoning models improve with more test-time computation, but often overthink, producing unnecessarily long chains-of-thought that raise cost without improving accuracy. Prior reinforcement learning approaches typically rely on a single outcome reward with trajectory-level length penalties, which cannot distinguish essential from redundant reasoning steps and therefore yield blunt compression. Although recent work incorporates step-level signals, such as offline pruning, supervised data construction, or verifier-based intermediate rewards, reasoning length is rarely treated as an explicit step-level optimization objective during RL. We propose Step-wise Adaptive Penalization (SWAP), a fine-grained framework that allocates length reduction across steps based on intrinsic contribution. We estimate step importance from the model's on-policy log-probability improvement toward the correct answer, then treat excess length as a penalty mass redistributed to penalize low-importance steps more heavily while preserving high-importance reasoning. We optimize with a unified outcome–process advantage within group-relative policy optimization. Extensive experiments demonstrate that SWAP reduces reasoning length by 56.3% on average while improving accuracy by 4.2% relative to the base model.
Authors: Matteo Ballegeer, Dries Benoit
Abstract: Boundary representations (B-reps) encode CAD geometry as parametric surface patches with explicit topological adjacency, forming information-rich geometric graphs. Yet existing methods encode coordinates and normals as scalars, making models rotation-sensitive, while invariant alternatives using per-face local frames discard the shared orientation reference needed to capture inter-face directional relationships. We introduce SO3brep, the first SO(3)-equivariant neural architecture for B-rep learning. Faces are represented by scalar, vector, and tensor irreducible representations constructed in per-face local frames, lifted to a shared global frame via Wigner D-matrices, and propagated through equivariant message passing. This preserves inter-face directional structure during message passing while guaranteeing SE(3) invariance at the output. Across five CAD benchmarks for shape classification and per-face segmentation, SO3brep consistently outperforms prior B-rep and equivariant point-cloud methods while significantly improving data efficiency. These results demonstrate that SO(3)-equivariance is a powerful inductive bias for structured CAD geometry.
PaperID: 1261, Poster
Abstract: Inverse reinforcement learning (IRL) is a well-established paradigm for circumventing the need for explicit reward. In this paper, we study the problem of estimating the reward function from a single sequence of actions (i.e., a demonstration) of a stochastic linear bandit algorithm. Our main result is a unified approach for inverse linear bandits, based on the idea of formulating a linear program by tightly characterizing the confidence intervals of pulled actions. We show that the estimation error of our algorithms matches the information-theoretic lower bound, up to polynomial factors in d and \log T, where d is the dimensionality of the feature space and T is the length of the demonstration. Compared to prior approaches, our approach (i) gives a unified reward estimator that works when the demonstrator employs LinUCB or Phased Elimination, two popular algorithms for stochastic linear bandits, while existing estimator only works for Phased Elimination; (ii) does not require access to hyperparameters or internal states of the demonstrator algorithm as required by prior work; and (iii) works for general action sets, while existing estimator requires assumptions on the density and geometry of the action set. We further demonstrate the practicality of our new approach by validating our new algorithms on synthetic data and demonstrations constructed from real-world datasets, where our estimators significantly outperform existing ones.
Authors:
Lars Uebbing, Harald Lykke Joakimsen, Siyan Chen, Georgios Leontidis, Kristoffer Wickstrøm, Michael Kampffmeyer, Sébastien Lefèvre, Arnt B Salberg, Robert JenssenAbstract: Most dimensionality reduction methods treat data as discrete point clouds, ignoring the continuous domain structure inherent to many real-world processes. To bridge this gap, we introduce Neural Operator Function Embedding (NOFE), a domain-aware framework for continuous dimensionality reduction. NOFE learns function-to-function mappings via a Graph Kernel Operator, enabling mesh-free evaluation at arbitrary query locations independent of input discretization. We establish NOFE as approximation of sheaf-to-sheaf mappings, generalizing Sheaf Neural Networks to continuous domains. We evaluate NOFE across different datasets, comparing it against PCA, t-SNE, and UMAP. Our results demonstrate that NOFE significantly outperforms baselines in local structure preservation, achieving a local Stress of 0.111 compared to 0.398 for PCA, 0.773 for t-SNE, and 0.791 for UMAP for the ERA5 climate reanalysis dataset. NOFE also exhibits robust sampling independence, reducing the Patch Stitching Error by up to 20.0× relative to UMAP (59.0 vs. 267.6 under regional normalization) and ensuring consistency across disjoint domain patches. While maintaining competitive global structure preservation (Stress-1: 0.379 vs. PCA's 0.268), NOFE resolves fine-grained structures and produces smooth, consistent embeddings that generalize across varying sample densities, addressing key limitations of discrete reduction methods.
PaperID: 1263, Poster
Abstract: In-context reinforcement learning (ICRL) has emerged as an effective paradigm for test time adaptation to unseen tasks without parameter updates. However, existing ICRL methods can exhibit brittle and unstable adaptions, and the mechanisms underlying such adaptions remain poorly understood. We provide a policy-gradient view of ICRL and argue that relying on trajectory-level return feedback can lead to high-variance updates when test time rollout is limited. Motivated by the actor--critic principle, we propose : the expected discounted return conditioned on the current state and accumulated interaction context. CV-ICRL trains a value head and writes its prediction back into the context as a value token, enabling TD-style return targets for lower-variance policy updates. Experiments on the Dark Room, Minigrid, and Procgen testbeds show that CV-ICRL substantially improves the stability of test time adaptation and achieves higher returns across tasks and environments. The source code and data of this paper are available at https://anonymous.4open.science/r/CV-ICRL-D161.
PaperID: 1264, Poster
Abstract: Geometry Problem Solving (GPS) serves as a rigorous touchstone for evaluating the complex reasoning capabilities of Multimodal Large Language Models (MLLMs), particularly when solutions necessitate auxiliary constructions absent in the visual input. However, current approaches face the following challenges: explicit diagramming methods suffer from unstable image generation quality and error cascading, while existing implicit attention mechanisms are restricted to visible elements, failing to perceive the latent auxiliary structures required for deduction. To bridge this gap, we propose AuxCA (Auxiliary-Clues-Aware GRPO), a novel framework that internalizes the capability of auxiliary construction into the model’s reasoning intuition without relying on external tools. Specifically, we introduce Auxiliary Clues Dependency (ACD) to quantify the causal influence of visual cues and integrate it into GRPO via a fine-grained advantage weighting strategy. Complementing this, we incorporate an auxiliary incentive reward to explicitly encourage the model to actively attend to and utilize these geometric elements during reasoning. Extensive experiments demonstrate that our resulting model, AuxCA-8B, significantly outperforms state-of-the-art explicit plotting methods and achieves performance competitive with top-tier proprietary models. Code is available in \urlhttps://anonymous.4open.science/r/AuxCA-5510.
PaperID: 1265, Poster
Abstract: Flow matching has recently emerged as a strong paradigm for state-of-the-art text-to-image (T2I) generation, enabling high-quality generation with a small number of sampling steps. As these models are increasingly integrated into real-world applications, ensuring safe and non-sensitive content generation has become a critical requirement. However, adapting safety and concept removal methods to this new generation framework remains an open challenge. Specifically, prior methods largely rely on iterative trajectory steering across a number of denoising steps or on CLIP-centeric prompt embedding manipulation. These design assumptions pose fundamental bottlenecks for safety in flow matching-based T2I generation, where limited sampling steps constrain iterative correction and modern context-aware text encoders diminish the effectiveness of embedding-level interventions. In this paper, we propose VESFlow, a training-free safety method tailored to flow matching with extremely few sampling steps. Leveraging the fact that flow matching models learn the marginal velocity (or average velocity in MeanFlow), we directly edit the velocity field via a Bayesian decomposition of the safe-conditional posterior. VESFlow steers the trajectory toward safe outputs while leaving the conditioning prompt unchanged. Building on the observation that VESFlow leaves outputs unchanged under benign prompts, we further introduce a risk score filtering that bypasses velocity editing to reduce computational cost while preserving benign prompt generation. Based on this filtering, we proposed VESFlow+ which provides stronger safety protection when filtered by the risk score. Experimental results show that our method removes the target concept, reducing the detection rate by NudeNet to 6.3% for step-4 model, while preserving fidelity on benign prompts. Code is available at the supplementary file.
Abstract: The trend towards larger training setups has brought a renewed interest in partially asynchronous two-phase optimizers which optimize locally and then synchronize across workers. Additionally, recent work suggests that the one-worker version of one of these algorithms, DiLoCo, shows promising results as a (synchronous) optimizer. Motivated by these studies we present an analysis of LA-DiLoCo, a simple member of the DiLoCo family, on a high-dimensional linear regression problem. We show that the one-worker variant, LA, provides a different tradeoff between signal and noise than SGD, which is beneficial in many scenarios. We also show that the multi-worker version generates more noise than the single worker version, but that this additional noise generation can be ameliorated by appropriate choice of hyperparameters. We conclude with an analysis of SLA --- LA with momentum --- and show that stacking two momentum operators gives an opportunity for acceleration via a non-linear transformation of the ``effective'' Hessian spectrum, which is maximized for Nesterov momentum. Altogether our results show that two-phase optimizers represent a fruitful new paradigm for understanding and improving training algorithms.
Abstract: Autoregressive neural simulators now match classical solvers on short-horizon prediction of physical systems, yet their accuracy degrades rapidly when rolled out over long horizons. In this work we identify of perturbations around rollout trajectories as a structural mechanism driving rollout error. Using a linearization analysis we show that when the Jacobians along an autoregressive trajectory are non-normal and non-commuting, the model amplifies errors transiently, resulting in model rollout drift even when the overall system is asymptotically stable. Building on the analysis, we propose \emphcommutativity regularization: a combination of two penalties designed to reduce the normality defect of individual Jacobians and the commutator norm of Jacobians across steps. The penalties are estimated with Jacobian-vector products and have no inference-time cost. We show a propagator bound that quantifies rollout error under approximate commutativity and normality. We evaluate UNet and FNO variants with commutativity regularization on 1D and 2D spatio-temporal data in synthetic and real settings, showing successful long-horizon rollouts over thousands of steps. We show that the method improves FourCastNet climate forecasts on ERA5 without using any new data. The gain is most pronounced out-of-distribution: trained on trajectories of a few hundred steps, regularized models remain in-distribution for thousands of rollout steps on initial conditions where baselines diverge.
Abstract: How large language models internally represent high-level behaviors is a core interpretability question with direct relevance to AI safety: it determines what we can detect, audit, or intervene on. Recent work has shown that traits such as evil or sycophancy correspond to linear directions in the internal activations, the so-called persona vectors. While these vectors are now utilized to inspect and steer model behavior in safety-relevant settings, how these representations are formed during training remains unknown. To address this, we trace persona vectors across the pretraining of OLMo-3-7B, with qualitative replication on Apertus-8B, a fully open model trained under a different recipe. They form remarkably early--- within 0.22% of OLMo-3 pretraining--- and remain effective for steering the fully post-trained instruct models. Persona vectors continue to refine geometrically and semantically throughout pretraining, but the core representation is already formed. We further compare alternative elicitation strategies and find that all yield effective directions, with each strategy surfacing qualitatively distinct facets of the underlying persona. Our results establish persona representations as stable features of early pretraining and open a path to studying how training forms, refines, and shapes them.
PaperID: 1269, Poster
Abstract: Textual information is ubiquitous in real-world time series datasets, appearing as news, reports, and annotations that hold key predictive signals not deducible from numerical history alone. Large Language Models (LLMs) interpret text but lack numerical precision, while Time Series Foundation Models (TSFMs) are quantitatively accurate but cannot reason about the effects of text. Moreover, the lack of large-scale paired multimodal datasets limits progress in bridging this gap. In this work, we address this challenge by providing a scalable approach for generating realistic multimodal datasets by annotating real-world time series with LLMs. Building on this, we propose Context as Covariates (CoCo), a forecasting framework in which an LLM distills its textual reasoning into numerically grounded forecast and confidence covariates that guide a TSFM backbone, thereby exploiting the strengths of both LLMs and TSFMs. We align the two models in stages, first fine-tuning the LLM via Group Relative Policy Optimization (GRPO) with a reward tied to covariate-induced improvement in forecasting performance, and then fine-tuning the TSFM on covariates generated by the aligned LLM. Our experiments show that models trained on our synthetic multimodal corpus generalize to real benchmarks spanning finance, economics, weather, traffic, and security. Notably, CoCo with Qwen3-4B-Instruct-2507 as the LLM and Chronos2 as the TSFM achieves up to 12% MAE improvement over state-of-the-art TSFMs and LLM forecasters, such as GPT-4o.
PaperID: 1270, Poster
Abstract: Modern image generative models are able to produce photorealistic images. As those images become increasingly indistinguishable from real data and are published online, they are often scraped for subsequent training runs of new generative models. This practice of training on generated data has been shown to degrade model performance and cause model collapse. A possible mitigation lies in embedding radioactive watermarks into generated content. Radioactive watermarks are robust marks that are detectable in outputs of new models trained on watermarked data, enabling provenance tracing of generated content. In this work, we analyze the persistence of image watermarks across multiple training-generation runs. To do so, we introduce a novel statistical testing method RADIUM (RadioActive Decay of Image-Underlaid Marks) for reliable radioactivity detection across various watermarking methods. Using our RADIUM method, we observe disparate radioactivity across watermarking methods for image generative models. Only few watermarks remain detectable in subsequently trained models, while most decay severely, especially when the architecture of models differs. Our analysis highlights the critical need for more radioactive watermarking methods in the vision domain.
PaperID: 1271, Poster
Authors: Vu V Quoc, Haytham Fayek, Thuy T Nguyen
Abstract: Jointly trained multimodal networks frequently underutilize their weaker modalities, as faster-learning modalities dominate the shared fusion objective and suppress the gradient signal reaching slower ones. Most existing approaches throttle the dominant modality, whereas directly boosting the weaker modality often fails because monitoring signals derived from the joint loss are entangled with, and thus dominated by, the stronger modality. We propose a new gradient modulation approach, probe-guided gradient boosting (PGGB), to address the problem by probing and boosting the weak modalities while preserving the strength of the dominant modality. Probing is performed by attaching a lightweight linear classifier to each modality’s stop-gradient features to estimate the corresponding representation quality score. The unbiased gap between these scores across modalities, the utilization gap, serves as an imbalance signal that defines a bounded and smoothed scaling factor adaptively reweighing the gradients of weaker modalities, enabling stable and effective rebalancing during joint optimization. The method is composable with throttling methods as the intervention leaves the loss and fusion architecture unchanged. Extensive experiments were conducted across eight benchmarks spanning four domains and 2–4 modalities. The results show that our method outperforms state-of-the-art approaches on highly imbalanced multimodal datasets, while remaining competitive on benchmarks with low or no modality imbalance. Bounded scaling, self-attenuation, and a standard-SGD descent bound were established under standard smoothness and bounded-variance assumptions. Code is provided in the supplementary material.
PaperID: 1272, Poster
Abstract: Randomized controlled trials (RCTs) identify treatment effects for an enrolled trial population by assigning treatment independently of potential outcomes, conditional on baseline covariates. These trial-specific effects may not generalize to a broader target population when the target and trial covariate distributions differ. This problem becomes harder when individual-level target population data are unavailable. Pre-trained large language models (LLMs) offer one way to address this data limitation by generating digital twins (DTs) with synthetic covariates intended to resemble the target population. Synthetic covariates alone, however, are insufficient, because neither the RCT covariate distribution nor the DT covariate distribution necessarily matches the target covariate distribution. We therefore propose a statistical inference procedure that integrates RCTs with calibrated DTs using external target population summaries. The procedure represents the target covariate distribution as a calibrated mixture of the RCT and DT covariate distributions, allowing DTs to contribute covariate information while limiting the influence of poorly calibrated synthetic data. The procedure also incorporates LLM-generated auxiliary outcome predictions through calibrated outcome regressions, improving precision without changing the estimand. Theoretical results and simulation studies show that the proposed estimators reduce bias relative to RCT-only estimation and improve efficiency for overall and subgroup target treatment effects.
Authors: Yordan Raykov, Rodrigo Veiga
Abstract: Generative Flow Networks (GFlowNets) have emerged as a flexible framework for amortized inference over discrete and mixed discrete-continuous objects, requiring only an unnormalized target density specified through a reward. In this work, we formulate forward-policy training in GFlowNets through the information geometry of the induced trajectory sampler. Treating the forward policy as an induced trajectory sampler, we show that its intrinsic first-order geometry is given by the Fisher-Rao metric of the trajectory family, and that the associated natural gradient provides the canonical local update whenever the corresponding Fisher information is computable or accurately approximable. We derive an exact decomposition of the trajectory Fisher into per-step conditional second moments, which clarifies when temporal score interactions vanish and when dense couplings remain under shared parameterization. This leads to three computational regimes: settings with tractable exact Fisher information, settings where Monte Carlo estimators of the expected Fisher are sufficient, and structure-exploitable settings in which target locality or factorization yields accurate approximations of the Fisher expectation. In the latter case, graphical-model tools such as exact marginalization, separator methods, and belief propagation provide principled surrogates for natural-gradient updates. The resulting framework turns target structure into optimization geometry and yields a tractable route to structure-aware forward-policy training in GFlowNets. We illustrate the framework empirically through examples comparing convergence and exploration behavior under Riemannian and Euclidean optimization.
PaperID: 1274, Poster
Abstract: Large language models (LLMs) possess remarkable ability to understand natural language descriptions of complex robotics environments. Earlier studies have shown that LLM agents can use a predefined set of skills for robot planning in long-horizon tasks. However, the requirement of prior knowledge of the skill set required for a given task constrains its applicability and flexibility. We present a novel approach L2S (short for Language2Subtasks) to leverage the generalization capabilities of LLMs to decompose the description of a complex task in natural language into definitions of reusable subtasks. Each subtask is defined by an LLM-generated dense reward function and a termination condition, which in turn lead to effective subtask training and chaining. However, LLMs lack detailed insight into the specific low-level control intricacies of the environment, such as threshold parameters within the generated reward and termination functions. To address this uncertainty, L2S (1) enables LLM reflection feedback loop to improve task decomposition and code generation, and (2) trains parameter-conditioned subtask policies that perform well in a broad spectrum of parameter values. As the impact of these parameters for one subtask on the overall task becomes apparent only when its following subtasks are trained, L2S selects the most suitable parameter value during the training of the subsequent subtasks to effectively mitigate the risk associated with incorrect parameter choices. During training, L2S autonomously accumulates a subtask library from continuously presented tasks and their descriptions, using guidance from the LLM agent to effectively apply this subtask library in tackling novel tasks. Our experimental results show that L2S is capable of generating reusable subtasks to solve a wide range of robot manipulation tasks.
Abstract: Robot policy learning benefits from world-action models that capture environment dynamics, but pixel-level prediction entangles dynamics with nuisance factors such as lighting and texture, making learned representations vulnerable to task-irrelevant visual variation. We propose JOPAT, a JOint Pixel-And-Track World-Action Model that predicts latent visual observations, 2D point tracks with visibility, and actions in a single denoising diffusion transformer. The key insight is that tracks provide an explicit representation of motion that captures long-horizon dynamics and remains robust under occlusion or partial out-of-frame motion, offering greater utility than modeling pixel appearance alone. On LIBERO and real-world LeRobot tasks, JOPAT improves over pixel-based baselines, with the largest gains on long-horizon tasks involving occlusion, object interaction, and off-screen motion.
Abstract: Agentic Large Language Models (LLMs) show great promise in clinical prediction using real-world clinical data. While recent studies have primarily examined their discrimination performance, their calibration performance—that is, whether predicted probabilities accurately reflect absolute clinical risk—remains largely unknown. This gap limits their use in high-stakes medical settings, where well-calibrated risk estimates are essential for guiding treatment and prevention decision-making. In this work, we provide theoretical analyses of how agentic LLM architectures, task knowledge, and tool using affect calibration performance in clinical prediction. Motivated by these analyses, we propose C^2Align, a Calibrated and Clinically Aligned framework, to optimize calibration while maintaining or even improving discrimination. C^2Align aligns two complementary agent branches: an ensemble of task-homogeneous LLM agents and a set of clinical task-inspired LLM agents. This alignment is achieved through a novel discrimination- and calibration-aware loss that encourages clinically meaningful agreement while controlling probabilistic miscalibration. We evaluate this framework on four clinical prediction benchmarks—spanning prognostic prediction and disease diagnosis tasks—using diverse LLMs, agent architectures, and calibration methods over both MIMIC-III and IV datasets. Extensive experiments, supported by theoretical motivation under assumptions, demonstrate that C^2Align improves calibration while maintaining or enhancing discrimination. Our work provides a theoretical, methodological, and benchmark reference for developing well-calibrated agentic LLMs for trustworthy clinical prediction.
Abstract: In long-context use, large language models frequently synthesize answers from the meaning of a relevant context span rather than literally copy-pasting them. Identifying which attention heads perform this synthesis matters for interpreting long-context model behavior. Yet existing detectors miss these heads by construction: they reward heads whose attended token matches the generated token, a literal-copy criterion that captures where a head reads but not what it writes through its output-value (OV) circuit, the very mechanism that carries non-literal retrieval. We introduce Logit-Contribution Scoring (LOCOS), a write-aware detector that scores each head by the projection of its OV-circuit output onto the answer-token unembedding direction, contrasting needle and off-needle source positions in a single forward pass. Across three model families (Qwen3, Gemma-3, OLMo-3.1), mean-ablating the top LOCOS heads on the NoLiMa non-literal retrieval benchmark collapses ROUGE-L at lower head counts than every attention-based baseline; on Qwen3-8B, ablating 50 heads drives ROUGE-L from 0.401 to 0.000 while the strongest baseline still retains 0.292. The selected heads are retrieval-specific: parametric recall and arithmetic reasoning stay at baseline under the same ablation.
PaperID: 1278, Poster
Abstract: We study Minty set inclusion on a compact convex set \mathcalX \subseteq \mathbbR^d with a set-valued operator A : \mathcalX \rightrightarrows \mathbbR^d. Under the Minty promise, the goal is to compute x \in \mathcalX and u \in A(x) such that \sup\_y \in \mathcalX \langle u, x - y \rangle \le \varepsilon. Our approach is geometric. We work in the fixed-scale constrained-resolvent oracle model, where a query at anchor a is assumed to return a resolvent point x^+ together with the Yosida vector g_\eta(a) = (a - x^+)/\eta. We prove that each oracle response yields a sharp dichotomy: either the returned point provides an \varepsilon-SVI certificate, or the Yosida vector defines a strict separating normal for the Minty set. Thus, a single black-box resolvent query supplies either a certified candidate solution or a valid geometric cut. Combining this separation principle with the ellipsoid method, we obtain \textscResolvent-Ellipsoid, an algorithm that computes an \varepsilon-SVI in \mathcalO\left(\operatornamepoly(d, \log(1/\varepsilon))\right) black-box oracle calls to the fixed-scale constrained resolvent. For monotone inclusions, our algorithm gives, to our knowledge, the first \mathcalO\left(\mathrmpoly(d, \log(1/\varepsilon))\right) black-box guarantee in the fixed-scale resolvent model.
Abstract: Generative modelling is a demanding test of foundation models, because it requires robust, holistic representation learning for a given data modality, rather than optimisation for a supervised prediction target alone. While recent work on tabular foundation models has achieved remarkable progress in predictive modelling, generative tabular foundation models remain underexplored. Existing tabular foundation generators, in particular, have not yet consistently matched strong dataset-specific generators in synthetic data quality. A key reason is their misalignment with the distinctive causal structural prior of heterogeneous tabular data. In this paper, we address this gap by introducing a novel tabular foundation model, neration. TabFORGE is designed to utilise the implicitly learned causal information underlying diverse tabular datasets in a unified latent space induced by a pretrained causality-aware feature encoder. It further decouples latent modelling from decoding through a two-stage design: we first pretrain a score-based diffusion transformer, and then pretrain a denoising-aligned decoder using the denoised latent embeddings. This design elegantly mitigates the distribution shifts in latent embeddings that typically arise between training and inference. We evaluate TabFORGE comprehensively against 22 benchmark methods on 45 real-world datasets. Our results show that TabFORGE effectively learns and leverages generalisable tabular representations, enabling efficient generation of high-quality synthetic tabular data, particularly with strong structural fidelity.
Abstract: Recent methods expose intra-request parallelism in LLM outputs, allowing independent branches to decode concurrently. Existing serving systems execute these branches eagerly or under fixed caps. We show that both are brittle: eager admission inflates the shared decode step, degrading co-batched requests in serial stages, while conservative fixed caps forgo the throughput that motivated exposing branches in the first place. We call the excess step latency caused by admitted branches the \emphbranch externality and show that the safe width depends on batch composition, context lengths, and accumulated slack, all of which change continuously over a workload trace. We introduce \textscPace, a per-step admission controller that treats extra branches as opportunistic work, admitted only when the predicted branch externality fits within the batch's current slack budget. Per-step regulation is practical because branch-level scheduling decouples compute from memory: branches share the request's prefix KV, so expanding or contracting width requires no memory reclamation. On Qwen3-32B, \textscPace improves goodput by 1.77× over \textscIRP-Off and by 1.48× over \textscIRP-Eager, while maintaining over 95% SLO attainment.
PaperID: 1281, Poster
Abstract: Outdoor robot navigation in dynamic, unknown environments remains a formidable challenge, necessitating robust traversability assessment and collision-free planning. While traditional map-based approaches struggle with adaptability to novel scenes, existing map-free methods frequently incur high computational costs and depend heavily on LiDAR data. In response, we propose LiteNav, a vision-only for exteroceptive sensing and strictly map-free outdoor navigation system. Relying solely on a single off-the-shelf RGB camera and GPS for goal specification, LiteNav estimates traversability, encodes navigation goals, and generates candidate trajectories via a diffusion-based joint encoding model. These trajectories are then refined by a goal-oriented planner to ensure correctness, efficiency, and safety. Remarkably, LiteNav achieves performance comparable to LiDAR-equipped methods without ever constructing or referencing explicit maps. Deployed on a low-power NVIDIA Jetson Xavier NX, it operates in real time, consuming only one-fifth of the computational resources required by prior approaches. Experiments demonstrate that LiteNav achieves state-of-the-art results on the GND dataset and outperforms existing methods in real-world scenarios. In summary, LiteNav is a map-free outdoor navigation system that operates using monocular RGB perception and GPS goal cues on an embedded platform, demonstrating strong potential for scalable deployment.
Abstract: Designing functional biological sequences requires navigating vast discrete spaces under strict evolutionary and biophysical constraints. Discrete Flow Matching (DFM) offers a generative framework over such spaces, but existing approaches rely on biologically uninformative couplings and offer limited flexibility for variable-length sequence generation and fine-grained control. We propose a structured coupling that encodes domain-specific preferences among sequence elements, biasing the source distribution toward plausible regions without modifying the flow objective or training procedure. Building on this, we introduce a latent edit-based rate parameterization that models variable-length generation via edit operations conditioned on a shared global latent, akin to a latent variable model, while remaining tractable. We further introduce a latent classifier-free guidance mechanism that steers generation coherently in continuous latent space, along with Dirichlet-prior temperature scaling for test-time control over edit operations. Our method achieves state-of-the-art performance across diverse biological sequence tasks, including density estimation, unconditional and conditional DNA sequence generation, and peptide sequence generation.
Abstract: Autonomous scientific discovery systems offer the potential to accelerate research by automating the process of hypothesis generation and validation. However, current systems operate within constrained search spaces or require predefined research questions, limiting their capacity for true open-ended inquiry. Furthermore, while they generate hypotheses iteratively, they largely lack the ability to explicitly synthesize their own accumulated findings to uncover complex, interconnected phenomena. We introduce DiscoPER, an autonomous large language model-powered framework that conducts open-ended research by dynamically generating and executing code to explore datasets without pre-specified research objectives. To ensure rigorous scientific validity, every proposed discovery must pass statistical testing. To overcome the limitations of isolated search, our framework introduces a second-order reasoning mechanism that periodically analyzes its own accumulated discoveries. By treating prior discoveries as empirical data, DiscoPER identifies structural patterns, confounds, and epistemic gaps, actively redirecting hypothesis exploration toward uncharted regions of the search space. The search space is further expanded by incorporating tool use, enabling the system to explore hypotheses beyond structured metadata by seamlessly processing and extracting useful information from multimodal sources like images. Evaluated on iNatDisco, a new multimodal ecological knowledge benchmark with pattern-level ground truth obtained from peer-reviewed literature, DiscoPER recovers 8 of 9 known patterns with a 72.7% hypothesis support rate, outperforming both classical causal discovery and LLM-guided baselines. Ablations show that DiscoPER scales with more data, and confirms the benefits of second-order ``meta-reflection''.
PaperID: 1284, Poster
Authors: Matthias Vigl, Nikita Pond, Nicole Hartman, Jackson Barr, Daniel H Guest, Alexander Froch, Michele D'Andrea, Dmitrii Kobylianskii, Daniel Murnane, Robert Les, Sebastien Rettie, Diptaparna Biswas, Michael Kagan, Lukas Heinrich
Abstract: Large language models (LLMs) have demonstrated that scaling, driven by compute, can yield dramatic improvements in performance, with current state-of-the-art systems reaching up to one trillion parameters. In contrast, state-of-the-art jet flavor classifiers in particle physics remain in the tens of millions of parameters. Scaling laws offer a framework for predicting how much particle physics models benefit from systematic increases in capacity and training data. In this work we study the behavior of increasingly larger jet classifiers, deriving scaling laws under optimal use of compute budget and limited datasets. We also identify the regimes where “double descent” emerges. The emerging scaling behavior points to a clear need for significantly larger datasets and models to saturate available compute and approach the limit of achievable performance. We therefore train a classifier on a new dataset 20 times larger than the current state-of-the-art (7.7 billion jets) and find that the observed performance follows the scaling laws prediction. This motivates continued integrated efforts in data generation, model scaling, and inference infrastructure for large-scale models in the physical sciences.
PaperID: 1285, Poster
Abstract: A solid understanding of the predictive success of deep neural networks (DNNs) remains elusive. For example, while many data-dependent metrics exist that seek to predict model quality of DNNs, these metrics may be expensive to compute for large datasets and/or may be impossible to compute when training datasets are not released. A practical and effective predictor of model quality involves analyzing DNN weight matrices, as it is known that a heavy-tailed spectral distribution often correlates with strong model quality. As such metrics lack a rigorous statistical derivation, in this work we aim to understand why they perform well, by employing a probabilistic framework to derive the marginal likelihood for the trained weights of a DNN. Across large-scale convolutional DNNs and LLMs, we find that this marginal likelihood predicts model quality. To understand the role heavy-tailed spectra play in predictive performance, we derive the limiting log-marginal likelihood for spectra that follow the heavy-tailed High-Temperature Marchenko-Pastur (HTMP) distribution. We show that this limiting quantity is convex in a
PaperID: 1286, Poster
Abstract: Recent image editing systems have achieved impressive semantic understanding, visual fidelity, and instruction-following ability, while video editing remains substantially more difficult and costly. In this paper, we present a simple alternative to end-to-end video editing: instead of training a monolithic video editor, we transform a strong image editor into a video editor through anchor-based generation. Our key insight is that video editing can be decomposed into two subproblems: editing a sparse set of keyframes and propagating those edits across time. Based on this observation, we propose Anchor-based Video Editing (AVE), a two-stage framework in which a powerful image editor first performs composed editing on selected keyframes, and a motion-guided image-to-video diffusion model then generates the final video by treating the edited keyframes as fixed anchors. This design directly inherits the strengths of modern image editors while avoiding expensive end-to-end video editing training. Experiments on IVEBench and VIE-Bench show that AVE achieves strong performance in instruction following, temporal consistency, and content fidelity. Further ablations reveal that final video editing quality is strongly correlated with the quality of the image editor, suggesting that future progress in video editing may come from stronger image editing foundations and lightweight transfer to video.
PaperID: 1287, Poster
Abstract: Instruction following is a core capability of modern foundation models, underpinning assistant-style interaction, agentic workflows, visual understanding, and multimodal reasoning. Existing instruction-following data and methods, however, remain substantially misaligned with realistic user behavior: users often converse with models over multiple turns, provide interleaved images and videos, switch modalities, or implicitly expect earlier constraints to persist. To bridge this gap, we first identify and formalize six core challenges for realistic conversational multimodal instruction following: Constraint-Guided Visual Reasoning, Cross-Modal Context Transfer, Persistent Instruction Tracking, User-Aware Context Modeling, Iterative Grounded Response Refinement, and Visual Consistency Alignment. Driven by these challenges, we develop a multi-agent data construction pipeline that generates diverse, constraint-rich, image- and video-grounded conversations, followed by human verification. We further improve model behavior through a rubric-guided multi-stage training pipeline that decomposes complex instructions into atomic supervision signals and trains models from basic visual constraint following to advanced multi-turn conversational constraint satisfaction. Finally, we establish Multimodal MultiChallenge (MMMC) as a systematic benchmark for realistic conversational multimodal instruction following. The results show that realistic conversational multimodal instruction following remains challenging for frontier models, while targeted rubric-guided post-training improves constraint satisfaction without sacrificing general multimodal ability.
Authors: Junyi Wu, Dan Li
Abstract: Learning a compact model of the world from interaction data is central to sample-efficient deep reinforcement learning. Spectral representation methods have become the leading paradigm for representation learning in continuous control by taking a matrix view of the transition kernel, with state-action pairs on one side and next states on the other, and learning a low-rank factorization through self-supervised contrastive objectives. We take this view one step further. The transition kernel is naturally a three-mode tensor over states, actions, and next states, and a CP decomposition gives one feature map per mode. We propose FaStR, which fits this decomposition with a noise contrastive objective, producing separate state, action, and next-state encoders that together form a single spectral representation. The factored form yields a smaller hypothesis class, and the sample size needed for representation learning shrinks by a factor that scales with the smaller of the state and action dimensions. Empirically, FaStR delivers its largest gains on high-dimensional locomotion tasks whose dynamics align with the factored structure, and the learned state encoder transfers intact across actuator shift while only the action encoder is retrained.
PaperID: 1289, Poster
Abstract: Fine-tuning has become a central mechanism for adapting large language models (LLMs) to downstream tasks, yet most existing methods follow a one-to-one paradigm: a single pretrained model is optimized into a single improved policy. This paradigm optimizes average single-policy performance, leaving no mechanism to distinguish redundant successes from genuinely complementary contributions. In this paper, we propose Co-Evolve, a one-to-many LLM fine-tuning framework that transforms a single base model into a cooperative population of persistent, specialized descendants and explicitly optimizes their complementarity. To reduce redundancy among descendants, we introduce a marginal-coverage objective that rewards each model for paying attention to examples under-covered by its peers. We instantiate this objective with evolution strategies, enabling derivative-free population-level optimization. We further introduce a seed-replay checkpoint mechanism that reduces the disk storage required to maintain multiple descendants by avoiding storing multiple full checkpoints during deployment. Experiments across multiple domains show that Co-Evolve consistently improves ensemble-level performance over strong fine-tuning and ensemble baselines. Our results suggest that optimizing for complementary model-level specialization provides a scalable alternative to single-policy fine-tuning. Our code can be found at https://anonymous.4open.science/r/Co-Evolution-Core-6604.
Abstract: Traditionally, decision support studies how humans use machine learning models to make better decisions. In modern agentic systems, this division of roles is increasingly reversed: AI agents act on behalf of users, while humans and tools becomes support mechanisms around them. This role reversal brings reliability concerns to the forefront, since agentic errors can be consequential and agent behavior must remain aligned with human goals and constraints. Departing from the classical view of decision support, we revisit its two basic principles, the cost--value tradeoff of seeking support and the role of uncertainty quantification, in a setting where AI agents are the central actors. We propose a framework for for AI agents through an optimization problem that minimizes support usage subject to controlling a counterfactual missed-support error: the probability that the agent acts alone on instances where support would have materially improved its output. At the population level, we show that the optimal policy is a threshold rule on the . Building on this structure, we develop an online algorithm that adaptively thresholds such a score and uses randomized exploration to control missed-support error without distributional assumptions. We further introduce a calibration-on-the-fly method that reduces unnecessary support calls online. We instantiate this framework across diverse scenarios, including information gathering, human--AI collaboration, and tool use, showing how each can be modeled through the same strategic decision-support lens. Experiments across these settings show that our method reliably controls the target error while substantially reducing support usage in practice.
PaperID: 1291, Poster
Abstract: Clifford algebra is becoming increasingly common in machine learning, e.g., in neural operators, equivariant transformers, and structured linear layers. The main appeal is that scalars, vectors, oriented planes, volumes, and higher-grade quantities live together in a single algebraic space. In this setting, rotations are represented by \emphrotors, which are parametrized by bivectors, with only \binomn2 degrees of freedom. Most Clifford algebra-based models work only with small algebras, in part because of a bottleneck in the sandwich product, x \mapsto r x r^\dagger, a basic operation for applying rotations to multivectors. It is evaluated through geometric products or via a full action matrix, both of which scale poorly with the 2^n-dimensional Clifford basis. We ask whether the \binomn2-parameter description can be carried all the way through application, and show that the answer is yes. The sandwich action decomposes grade-wise into exterior powers of the induced vector rotation. We factor this vector rotation into Givens rotations and apply their lifted actions as sparse updates on pairs of blades. The resulting algorithm never needs to form the rotor coefficients or the 2^n × 2^n action matrix. Empirically, our algorithm achieves up to 64 × wall-clock and 35× memory improvements over existing rotor-sandwich implementations, enabling rotor-based layers at scales previously out of reach.
Abstract: We identify a structural weakness in current large language model (LLM) alignment: refusal mechanisms in these models are . While existing approaches tend to encode refusal behaviors across multiple latent features, suppressing a single feature (via prompt-based jailbreaks) is sufficient to collapse alignment, leading to unsafe generation. Motivated by this, we propose as a design principle for robust LLM safety: refusal mechanisms should remain effective even under partial failures via redundant, independent causal pathways. We present a concrete instantiation of this principle: a progressive alignment framework that iteratively identifies and ablates previously learned refusal directions, forcing the model to reconstruct safety along new, independent subspaces. Across six jailbreak attacks, we achieve the strongest overall robustness (1.7% average attack success rate) while largely preserving generation quality and preventing excessive over-refusals, with negligible computational overhead. Additional analyses confirm that models trained with our method encode multiple, causally independent refusal directions that prompt-based jailbreaks cannot fully suppress, providing empirical support for fail-closed alignment as a principled foundation for robust LLM safety.
Abstract: Calibrated predictions are useful because their numerical values can be interpreted as probabilities. Calibration errors are therefore widely used to evaluate and compare probabilistic predictors. Recently, Haghtalab et al. [2024] introduced truthfulness as an additional requirement for such measures. A calibration measure is truthful if a predictor minimizes its expected measured error by reporting the true conditional label distribution. Many standard empirical calibration errors are non-truthful: a predictor may appear better calibrated by distorting its probabilities. We study the implications of truthful and non-truthful calibration errors for practice. First, we introduce perfectly truthful calibration errors for full multiclass calibration and classwise calibration, two standard multiclass notions. More generally, our construction applies to any linear property of the label distribution, generalizing the truthful calibration error for binary predictions in Hartline et al. [2025]. We also identify a truthful correction for confidence calibration. Second, we characterize the decision-theoretic implications of these truthful errors. For calibrated predictors, truthful calibration errors preserve the Blackwell dominance: a more informative calibrated predictor receives no larger expected error. Third, we show that this decision-theoretic interpretation explains and mitigates the well-observed ranking robustness problem of binned calibration errors. Empirically, non-truthful confidence-based errors can reverse model rankings when the number of bins changes, while our truthful classwise error gives more stable rankings across binning choices.
PaperID: 1294, Poster
Abstract: We study high-probability linear regression in the infinite variance regime, where the regression noise has only a finite conditional \alpha-moment \nu for an unknown \alpha\in(1,2), and the covariates may also be heavy-tailed. We give a black-box reduction that converts finite-variance regression procedures into weak-moment procedures. The reduction is based on empirical residual pruning and is agnostic to the moment exponent, the noise scale, the covariance, and the design moment constant. Under a finite fourth-moment condition on the covariates, any finite-variance subroutine with the sub-Gaussian rate yields the following for all n above a suitably defined threshold: \lVert \widehat\theta-\theta^ \rVert_\Sigma \lesssim_\alpha \nu^1/\alpha \left(\fracd+\log(1/\delta)n\right)^\frac\alpha-1\alpha with probability at least 1-\delta, without knowing \alpha or \nu. We further prove a matching minimax lower bound that holds under the same conditions. The reduction is computationally efficient whenever the finite-variance subroutine is. We also obtain a distribution-free prediction guarantee under a bounded raw \alpha-moment assumption on the response, and give an analogous \alpha-agnostic reduction for mean estimation.
Abstract: Recent advances in Large Language Models (LLMs) have enabled agentic systems capable of solving complex tasks through multi-turn planning, tool use, verification, and memory updates. However, learning agentic systems remains difficult due to two fundamental challenges, i.e., (1) long-horizon credit assignment, where supervision is available only at the final outcome, and (2) imbalanced data distributions, where dominant data patterns bias optimization and weaken adaptation to rare but informative reasoning behaviors. In this paper, we propose Fair Multi-Level Preference Optimization (Fair-MPO or \Phi-MPO), a new preference optimization framework for agentic learning. We first show that Multi-Level Preference Optimization provides a principled and more computationally efficient framework for long-horizon reasoning. Then, we introduce a Fair Multi-Level Objective that addresses imbalance in agentic learning. We provide a comprehensive theoretical analysis demonstrating that our approach addresses both long-horizon reasoning and data imbalance. Our experiments on agentic reasoning benchmarks demonstrate that our approach achieves State-of-the-Art (SOTA) performance.
Abstract: Multilingual large language models often produce inconsistent reasoning and answers across semantically equivalent prompts in different languages. Prior work suggests that intermediate representations can be relatively language-agnostic, but generation becomes increasingly language-specific as models commit to discrete output tokens. This is problematic because language-specific lexical choices can cause semantically equivalent reasoning paths to diverge across languages. This motivates searching for a cross-lingual alignment signal that is less tied to any single vocabulary item or script. We propose representations across languages, using English as a pivot. Soft tokens are probability-weighted mixtures over the vocabulary embeddings, yielding continuous representations that can aggregate information from semantically related tokens across languages. We then align each non-English soft-token summary to its English counterpart in the shared embedding space. Across four multilingual reasoning benchmarks, SOLAR improves accuracy by up to +17.7 points over the base model and +3.8 over standard supervised fine-tuning, with the largest gains on low-resource languages. SOLAR also strengthens final-layer cross-lingual similarity and substantially reduces language-cluster separability, suggesting that aligning soft-token representations helps preserve shared semantic structure during multilingual reasoning.
Abstract: Text-to-SQL systems are typically evaluated by query-level execution correctness, but this terminal signal provides little guidance about which intermediate SQL decision caused success or failure. Token-level dense supervision is also ill-suited: SQL tokens do not align with complete semantic decisions, can penalize execution-equivalent queries, and are difficult to label reliably at scale. We therefore propose CAPER, which automatically derives clause-level supervision via counterfactual intervention on the SQL abstract syntax tree (AST), enabling root-cause error localization for reward modeling; the resulting data is used to train CAPER-9B, a lightweight Clause-PRM that provides clause-boundary feedback for policy optimization and candidate verification. Experiments on BIRD and Spider show that clause-aligned supervision not only improves execution accuracy, achieving up to a 15.3% relative EX improvement over GPT-5.4, but also strengthens failure-localization capability, reaching 84.53% accuracy and 90.60% MRR on held-out failures.
PaperID: 1298, Poster
Abstract: Causal representation learning (CRL) offers the promise of uncovering the underlying causal model by which observed data was generated, but the practical applicability of existing methods remains limited by the strong assumptions required for identifiability and by challenges in applying them to real-world settings. Most current approaches are applicable only to relatively restrictive model classes, such as linear or polynomial models, which limits their flexibility and robustness in practice. One promising approach to this problem seeks to address these issues by leveraging changes in causal influences among latent variables. In this vein we propose a more general and relaxed framework than typically applied, formulated by imposing constraints on the function classes applied. Within this framework, we establish partial identifiability results under weaker conditions, including scenarios where only a subset of causal influences change. We then extend our analysis to a broader class of latent post-nonlinear models. Building on these theoretical insights, we develop a flexible method for learning latent causal representations. We demonstrate the effectiveness of our approach on synthetic and semi-synthetic datasets, and further showcase its applicability in a case study on human motion analysis, a complex real-world domain that also highlights the potential to broaden the practical reach of identifiable CRL models.
PaperID: 1299, Poster
Authors: Zixuan Zhao, Sam Wheeler, Neil Getty, Xiaotian Duan, Rick Stevens, Fangfang Xia
Abstract: Recurrent models trained from scratch have recently become competitive on ARC-style reasoning tasks, but the usual framing around small recurrent backbones overlooks two important parts of the system: task-conditioned memory and synthetic augmentation data. We study this regime through CHARM, a compact hybrid ARC model that combines recurrent reasoning with structured task memory, synthetic data, and inference-time aggregation. In existing approaches, task-conditioned memory supplies a large hidden source of capacity, reaching more than 30x the size of the recurrent backbone. We introduce a compositional sparse embedding for task conditioning that reduces learned task-memory parameters by over 90% while improving pass@2 in controlled ARC-AGI-1 ablations. We also find that matched synthetic data is the largest measured driver of ARC-AGI-2 gains. For the recurrent backbone, recurrent depth helps only when balanced with learning horizon. Combining these ingredients, our system reaches 80.5% pass@2 on ARC-AGI-1 and 39.3% pass@2 on ARC-AGI-2 public evaluation.
PaperID: 1300, Poster
Authors: Amil Khan, Anantajit Subrahmanya, B.S. Manjunath
Abstract: Packed high-token 3D fields are not just long-sequence modeling: once a volume is packed into tokens, sequence order is no longer the geometry of the sample. This creates a general failure mode for attention over packed geometric data: sequence-index positional mechanisms can learn the memory layout rather than the object. We introduce Pack3D, a geometry-aware local--global attention operator that separates memory layout from attention geometry. Pack3D keeps local neighborhood attention exact, compresses only distant context into contiguous 3D block summaries, and indexes attention with axial RoPE on true 3D token-grid coordinates rather than packed sequence positions. With RoPE applied before block compression, each query interacts with a geometry-aware mixture of block tokens, so compressed global context remains spatially interpretable after packing. A synthetic repacking benchmark makes the failure mode explicit: learned positional embeddings and packed-sequence RoPE fail under layout changes, whereas axial grid-coordinate RoPE remains invariant. On Allen WTC--11 hiPSC microscopy, our main claim-grade domain, Pack3D improves matched learned-position local--global attention by 0.0978 best macro Dice and 0.1248 final macro Dice at 262k tokens and depth 8. In an Allen-15 joint comparison, Pack3D also improves over a parameter-matched TokenGrid 3D U-Net on both best and final macro Dice. Inference-time ablations show that the trained model genuinely uses its compressed summaries, and the fused forward/backward implementation is 6.49× lower latency than dense SDPA at 65k tokens.
Abstract: Large Language Model (LLM) alignment is intended to ensure that models remain helpful and safe, but its stability under input distributional shift is not yet fully understood. Prior work shows that aligned models can fail under jailbreak prompts, alternative encodings, and cross-lingual transfer, yet these failures are usually studied as attacks rather than controlled probes of alignment generalization. Moreover, existing evidence is largely grounded in natural language variation already represented during pretraining, leaving unresolved whether alignment generalizes with semantic content or remains tied to superficial surface patterns. In this paper, we study this question using synthetic semantic-preserving transformations that are rule-based and invertible, preserving task-relevant meaning while shifting inputs beyond standard linguistic variation. Across four open-weight and four commercial models, under both fine-tuning and in-context learning, we use these transformations as a probe of alignment generalization and identify an empirical pattern we term Alignment--utility asymmetry: once models can operate effectively on transformed inputs, task utility is often substantially retained while alignment failure increases more sharply. For example, adapted GPT-4.1 mini shows only limited utility degradation under transformation while its harmful rate rises from 13.3% to 74.3%; Gemini 3 Flash similarly retains near-original utility while its harmful rate increases from 2.3% to 43.0%. Taken together, these results suggest that semantic-preserving distribution shifts can expose a recurring gap in how utility and alignment generalize in current LLMs.
Abstract: We study a repeated information design setting in which the receiver, who is also the decision-maker, updates beliefs in a systematically biased way. More specifically, a distorted posterior in our model can be written as a convex combination of the prior and the Bayesian posterior, governed by a fixed but unknown parameter. Over repeated interactions, the sender chooses persuasive signaling schemes, observes only the receiver’s realized actions, and seeks to minimize regret relative to a full-information oracle that knows the receiver’s biased updating rule. We propose a safe exploration algorithm for learning the receiver’s bias while maintaining high persuasion value. The algorithm exploits the asymmetric cost of probing: conservative probes incur only local loss, whereas overly aggressive probes may lose the persuasive opportunity entirely. For general finite state and action spaces and arbitrary bounded utilities, our method achieves O(\log\log T) regret. A matching \Omega(\log\log T) lower bound shows that this rate is optimal. We further discuss the influence on receiver welfare, as well as extensions to jointly unknown prior and bias, and contextual settings with time-varying priors and utilities.
Authors:
Ruiqi Liu, Xiaolei Lv, Gengsheng Li, Ximo Zhu, Zhiheng Wang, Zhengbo Zhang, Junkai Chen, Zhiheng Li, Bo Li, Jun Gao, Shu WuAbstract: On-policy knowledge distillation has proven effective for language models, yet its application to vision-language models (VLMs) remains underexplored. We observe that standard on-policy distillation can improve a student's output quality while failing to strengthen its reliance on visual input: on vision-critical tokens, the student's predictions remain largely unchanged whether or not fine-grained visual detail is present, even though the teacher's predictions depend heavily on it. To make this difference observable, we introduce (VA), the token-level log-probability difference when the teacher scores a student-generated rollout with versus without access to fine-grained visual detail. VA is concentrated in a small minority of tokens, and these high-VA tokens are the ones that actually carry the visual supervision signal. This motivates a distillation objective that treats them differently from language scaffolding, so their contribution is not diluted by the abundant surrounding language tokens. We propose , which uses VA at two granularities: rollout-level reweighting by trajectory-averaged VA, and token-level KL averaged within high-VA and low-VA groups separately. We train on two math datasets (Geometry3K and ViRL-39k) and evaluate on eight benchmarks covering both mathematical reasoning and visual understanding, across three teacher sizes (4B, 8B, and 32B) on the Qwen3-VL family. VA-OPD improves over standard on-policy distillation on every benchmark, with the gain growing monotonically along both the teacher-size and data-scale axes, suggesting that these factors compound consistently.
Abstract: Virtual photography asks an agent to enter a prepared 3D scene with no preselected camera pose or reference image, infer a suitable shot from scene information and a language intent, choose executable camera parameters, and render the final photograph. Recent progress in vision-language models makes this kind of spatial agent increasingly plausible, but the task stresses two capabilities that remain hard to evaluate together: complex 3D spatial understanding and abstract aesthetic judgment. We introduce PhotoFlow, a Director-Reviewer-Reflector agent for closed-loop camera search. The Director builds a soft photographic blueprint and proposes diverse candidate cameras; the Reviewer combines rule checks, visual critique, and pairwise incumbent selection; and the Reflector converts failures into region memory, dead-zone suppression, and high-explore relocation. We also introduce VPhotoBench, a benchmark of 47 open-license Blender scenes and 141 language-conditioned photography missions spanning subject placement, relational composition, and atmosphere/style. On held-out experiments,PhotoFlow achieves the strongest external quality-alignment composite and success rate among one-shot prediction, single-chain reflection, anchor-bank selection, and random search under a six-round rendering budget. To our knowledge, this is the first work to make language-conditioned virtual photography in 3D design or game engines (e.g., Blender) an executable agent task, and our results show that an LLM-centered spatial agent can already produce strong photographs in a setting designed to challenge both 3D reasoning and aesthetic choice.
PaperID: 1305, Poster
Abstract: Message Passing Neural Networks (MPNNs) commonly suffer from over-squashing and over-smoothing issues, which limit their performance on graph-level tasks. To this end, we propose a novel Crossed Subtree Convolutional Network (CSCN) for graph classification. Specifically, CSCN first introduces an ordered subtree alignment strategy to decompose the original graph into an ordered sequence of structured subtrees. This sequence preserves rich local connectivity patterns and maintains the permutation invariance of the graph representation. Based on this, we design a crossed subtree convolution. It utilizes a sliding window to enable information interactions across different subtrees while simultaneously aggregating node features within each subtree. This design captures local information effectively while also modeling long-range dependencies. In this way, CSCN avoids deep message passing over the global graph topology and thus alleviates both over-squashing and over-smoothing. Experimental results on multiple graph classification benchmarks show that CSCN achieves competitive performance.
Abstract: Regulatory approval of products in high-stakes domains such as drug development requires statistical evidence of safety and efficacy through large-scale randomized controlled trials. However, the high financial cost of these trials may deter developers who lack absolute certainty in their product's efficacy, ultimately stifling the development of `moonshot' products that could offer high social utility. To address this inefficiency, in this paper, we introduce a statistical protocol for experimentation where the product developer (the agent) conducts a randomized controlled trial sequentially and the regulator (the principal) partially subsidizes its cost. By modeling the protocol using a belief Markov decision process, we show that the agent's optimal strategy can be found efficiently using dynamic programming. Further, we show that the social utility is a piecewise linear and convex function over the subsidy level the principal selects, and thus the socially optimal subsidy can also be found efficiently using divide-and-conquer. Simulation experiments using publicly available data on antibiotic development and approval demonstrate that our statistical protocol can be used to increase social utility by more than 35% relative to standard, non-sequential protocols.
Abstract: Estimating how often an ML model will fail at deployment scale is central to pre-deployment safety assessment, but a feasible evaluation set is rarely large enough to observe the failures that matter. Jones et al. (2025) address this by extrapolating from the largest k failure scores in an evaluation set to predict deployment-scale failure rates. We give a finite-k decomposition of this estimator's forecast error and show that it has a built-in bias toward over-prediction in the typical case, which is the safety-favorable direction. This bias is offset when the evaluation set misses a rare high-failure mode that the deploy set contains, leaving the forecast to under-predict at deployment scale. We propose a fine-tuning objective, the forecastability loss, that addresses this failure mode. In two proof-of-concept experiments, a language-model password game and an RL gridworld, fine-tuning substantially reduces held-out forecast error while preserving primary-task capability and achieving safety similar to that of supervised baselines.
Authors: John Sous, Michael Winer
Abstract: We introduce a model for neural scaling laws under sparse activations. In the model, test loss is often dominated by rare coordinates that are never observed in the training input. This mechanism induces a novel bottleneck absent from dense models. We derive the asymptotic population loss in both the underparameterized and overparameterized regimes, and show that the loss exhibits a double-descent peak near the interpolation threshold---where the number of parameters is just sufficient to fit the training data---resulting in a loss curve governed by two distinct scaling exponents---one for the overparameterized regime and one for the underparameterized regime---with a gap determined by the degree of sparsity. Additionally, we derive a compute-optimal frontier that favors increasing dataset size over model capacity under fixed compute budgets. We also analyze gradient-descent dynamics and identify a scaling law for the probability that fixed-step gradient descent becomes unstable. We further show that the sparsity-induced effect persists under nonlinear activations. Experiments validating the theory can be found at https://anonymous.4open.science/r/sparse-scaling-neurips2026-3CDB/SparseScaling1.ipynb.
Abstract: Discrete diffusion models are increasingly competitive for language modeling, yet it remains unclear how their denoising objectives organize learning. Although these objectives target the full data distribution, we show that the exact reverse process induces a hierarchy between coarse information. For uniform and absorbing (a.k.a. masking) diffusion, we prove that, in the small-noise regime of the final denoising steps, each single-token reverse edit decomposes into a leading scale, determined by whether it moves toward the data within the same scale. Thus, recovering validity structure only requires learning the correct order of magnitude of reverse probabilities, whereas recovering data frequencies requires coefficient-level estimation. The separation is mechanism-dependent: uniform diffusion exhibits a trichotomy into validity-improving, validity-preserving, and validity-worsening edits, while absorbing diffusion places its leading-order mass on validity-improving moves. Experiments on a masked language diffusion model and synthetic regular-language tasks support these predictions: support-localization emerges earlier than within-support frequency ranking, and the contrast between uniform and absorbing diffusion matches the predicted rate separation. Together, our results suggest that discrete diffusion models learn data support before data frequencies.
PaperID: 1310, Poster
Abstract: Although Multimodal Large Language Models (MLLMs) are increasingly deployed in high-stakes domains, the fairness of their outputs is under-explored. Building on the BBQ language bias benchmark, we construct a new dataset MultiBBQ using attested social biases and AI-generated photorealistic images for controllable fairness evaluation of MLLMs in both visual-only and visual-language contexts. We propose two metrics Fairness Score and Bias Score and design an evaluation paradigm to address shortcut reasoning and data contamination challenges. Using comprehensive benchmarking, we diagnose four new Fairness Failure Modes of MLLMs. In particular, we discover that proprietary models may fail to conduct effective counter-bias reasoning in disambiguated contexts due to over-refusal, while open-source models are deficient in abstaining in ambiguous contexts. We also analyze how different input and model factors degrade fairness, demonstrate that MLLMs amplify bias over their backbone LLMs, and show the potential limited effectiveness of mitigation methods such as reasoning and fairness instruction. We release our code and dataset to facilitate further evaluations and the development of mitigation methods here.
PaperID: 1311, Poster
Abstract: Neural ordinary differential equations replace a finite layer stack by the solution map of a learned dynamical system. Thus the machine-learning primitive evaluated at inference time is the flow endpoint generated by the learned vector field, computed to requested accuracy. We study the complexity of this global forward pass when the vector field has standard compactness and Lipschitz certificates. Using second-order complexity theory, we identify the exact operator-level complexity of this local-to-global computation: scalar Neural-ODE endpoint inference is second-order polynomial-time equivalent to the classical Lipschitz IVP solution operator and is therefore FPSPACE_2-complete in the worst case. This complexity resides in the continuous-depth layer itself, independently of any particular numerical solver: even one fixed three-dimensional autonomous Neural-ODE block with C^1 local dynamics can define a PSPACE-hard input-output map. At every finite smoothness level C^k with k\geq2, such autonomous blocks can define counting-hierarchy-hard maps. One scalar readout-gradient query recovers the forward endpoint, so the same complexity reaches an elementary training-gradient computation. Hence continuous-depth models and compressed-depth tied ResNets can realize worst-case global computational complexity even when local dynamics are cheap. As a tractable counterpart, fixed-dimensional Taylor-certified analytic vector fields admit polynomial-time endpoint inference under quantitative magnitude and radius certificates. These results show, based on the rigorous framework of second-order complexity theory, that the computational reliability of continuous-depth learning systems depends jointly on global flow structure, representation certificates, and the cost of local network evaluations.
Abstract: We study stochastic linear bandits with delayed feedback under several delay models and establish near-optimal regret guarantees. Our results identify when delayed linear bandits exhibit the same qualitative behavior as multi-armed bandits (MAB), and when the linear structure creates fundamentally new challenges. Specifically, (1) for , where the delay does not depend on the realized loss (but potentially depends on the arm), we show that delays incur only an additive regret penalty. Under stochastic delays, this penalty scales with the expected delay, while under adversarial delays, it scales with the maximum number of outstanding observations. Notably, both delay penalties are dimension-free, improving upon the state-of-the-art results; (2) for , we show that linear bandits are substantially harder than MAB: unlike in MAB, we prove matching (up to log factors) upper and lower bounds in linear bandits, whose delay penalty depends on the square root of the dimension. (3) for the , a special case of loss-dependent delay, we show that the optimal MAB guarantee, which depends only on the delay of the optimal arm, is also unattainable in linear bandits. Together, these results provide a sharp characterization of how delayed feedback interacts with linear generalization.
PaperID: 1313, Poster
Authors:
Yu Wang, Fan Yang, Kaikun Xu, Li Hao, Li Yuan, Jun Zhu, Jingjie Zhang, Zhenchao Tang, Yatao Bian, Cheng Chang, Jianhua YaoAbstract: Untargeted metabolomics produces massive numbers of tandem mass spectra, yet most spectra remain difficult to assign to molecular structures because each spectrum is only a partial, instrument-dependent observation of the underlying molecule. We study MS/MS annotation as modality-incomplete molecular--spectral inference: molecules, spectra, and fingerprints provide complementary views of the same chemical entity, but fully paired molecule--spectrum data are limited and many training examples contain only molecular or only spectral information. We introduce MolSpecFlow, a mask-conditioned heterogeneous flow framework that represents these modalities as coupled components of a shared molecular--spectral state. Availability masks encode which components exist in each example, and observation masks specify which components are conditioned on or generated, instantiating de novo annotation, molecular retrieval, and spectrum simulation within one backbone. MolSpecFlow uses a shared Mol-Spec Transformer, hybrid discrete--continuous flow dynamics, and expected mass/formula regularization over soft molecular token distributions. Experiments show consistent gains from modality-incomplete pretraining, multi-interface reuse, and external de novo transfer.
Authors:
xuhui li, Zhengquan luo, Zixu Wu, Xiwei Liu, Yongqiang Yu, Zixiang Hong, Zhiqiang XuAbstract: Dataset distillation compresses large datasets into compact synthetic sets with comparable performance in training models. Despite recent progress on diffusion-based distillation, such methods typically rely on heuristic guidance or prototype assignment over long denoising chains, which increases sampling cost and makes prototype-consistent control harder under strong guidance or low IPC. We propose \emphPath-Guided Flow Matching (PGFM), the first flow matching-based framework for generative distillation, which enables deterministic synthesis by solving an ODE in a few steps. In particular, we introduce a retrieval-based prototype inversion stage that identifies prototype-consistent initial noises for class prototypes on the frozen flow manifold, and further develop an anchor-guided residual correction strategy for bounded stage-wise control. This design follows a controlled-transport principle: retrieval reduces initialization mismatch, while anchor guidance provides a bounded residual correction along the flow trajectory. Extensive experiments across high-resolution benchmarks demonstrate that PGFM matches or surpasses prior diffusion-based distillation approaches with fewer sampling steps while delivering competitive performance with improved efficiency.
Abstract: Monitoring the chain-of-thought (CoT) of reasoning models is a promising approach for detecting covert misbehavior (i.e., hidden objectives) in code generation tasks. While large models (GPT-5, Gemini-3-Flash) can serve as effective CoT monitors, they are expensive to deploy due to the lengthy reasoning traces and high API cost, emphasizing the need for smaller, cheaper alternatives. Nevertheless, we find that current small models (4B--8B) struggle to detect hidden objectives despite access to the CoT, frequently misattributing them as part of the user query. To address this, we propose a post-training pipeline combining supervised fine-tuning (SFT) and reinforcement learning (RL), where SFT narrows the gap for in-domain tasks by distilling detection behavior from stronger monitors, and RL on hard and subtly crafted hidden objectives helps the model generalize to out-of-domain monitoring tasks. To validate this generalization, we evaluate under a realistic threat model motivated by practical supply-chain attacks, where the adversary is a third-party LLM router injecting hidden objectives into code-generation requests through either prompt manipulation or code manipulation attacks. To push beyond objectives that large monitors already saturate, we also introduce four new challenging tasks even for strong monitors. Finally, we introduce CoT-Guard, a 4B-parameter monitor that demonstrates superior generalization performance under both prompt and code manipulation attacks, achieving a \textGmean^2 (i.e., TNR×TPR) of 75% and outperforming GPT-5.4 (56%), GPT-5-mini (41%), and Qwen3-32B (54%), while closing the gap to Gemini-3-Flash (83%). These results demonstrate that \modelname provides a practical and cost-effective user-side defense, substantially improving hidden-objective detection while avoiding the deployment cost of large monitors.
PaperID: 1316, Poster
Abstract: Discovering the underlying causal structure of a system is a fundamental challenge in machine learning, requiring interventional data to distinguish between observationally equivalent models. In this paper, we study the problem of learning causal Bayesian networks in an online setting, where a learner sequentially observes individual samples from a stream of unknown interventional and observational distributions. We formulate this task as an online sequence prediction problem. To overcome the super-exponential size of the DAG search space, we extend dynamic programming algorithms originally developed for uniform DAG counting and sampling within Markov Equivalence Classes to support score-decomposable, weighted sampling. We prove that the posterior distribution maintained by our algorithm competes with the optimal causal structure in hindsight, outputting a distribution that is close to the true data-generating mechanism as measured by the Interventional Kullback-Leibler divergence.
Abstract: Neurosymbolic (NeSy) models integrate neural networks and symbolic reasoning for robust and interpretable AI. State-of-the-art NeSy models require that the symbolic component is expressed in a differentiable way, often complicating the use of approximate inference. We propose EM-NeSy which casts probabilistic NeSy learning as an instance of the Expectation-Maximization (EM) algorithm. In the expectation step, we compute the posterior over the neurally predicted symbols conditioned on the label via probabilistic inference. In the maximization step, we update the neural parameters based on this posterior using gradient descent only through the neural component. This formulation unlocks the full potential of the EM algorithm for NeSy learning. It allows NeSy to extend naturally to approximate reasoning without any additional modifications or differentiability requirements of the symbolic component. Furthermore, it recovers the standard end-to-end gradient-based NeSy setting under exact inference. Our experimental results demonstrate the scalability and computational efficiency of EM-NeSy.
Abstract: Existing image generation models face critical challenges regarding the trade-off between computation and fidelity. Specifically, models relying on a pretrained Variational Autoencoder (VAE) suffer from information loss, limited detail, and the inability to support end-to-end training. In contrast, models operating directly in the pixel space incur prohibitive computational cost. Although cascade models can mitigate computational cost, stage-wise separation prevents effective end-to-end optimization, hampers knowledge sharing, and often results in inaccurate distribution learning within each stage. To address these challenges, we introduce a unified multi-stage generative framework formulated under the \emphstochastic interpolant formalism with our Conditional Dependent Coupling strategy as a Flow-Matching variant. It decomposes the generative process into interpolant trajectories at multiple stages, ensuring accurate distribution learning while enabling end-to-end optimization. Importantly, the entire process is modeled as a single unified Diffusion Transformer, eliminating the need for disjoint modules and also enabling knowledge sharing. Under explicitly stated assumptions, we provably reduce both transport cost and asymptotic inference time, we empirically validate the underlying NFE--transport-cost and per-evaluation-cost relations. Extensive experiments demonstrate that our method achieves competitive fidelity and substantially lower wall-clock cost in the low-NFE regime across multiple resolutions.
Abstract: To account for uncertainty, hierarchical Bayesian optimization marginalizes over Gaussian process hyperparameters, usually averaging models or acquisition functions. However, this early aggregation discards information about model disagreement and is sensitive to outliers. To address these limitations, we propose to defer aggregation to the decision level: each model independently proposes a candidate, and we aggregate over these candidates via the medoid. We study a misspecified hierarchical Gaussian process setting in which most models have low misspecification error, while a minority are severely misspecified. Theoretically, we derive regret bounds showing that the misspecification penalty of early aggregation depends on the average error across all models, whereas decision-level aggregation depends only on the low-misspecification majority. Empirically, we validate these findings on toy examples and further demonstrate that, across a wide range of low- to high-dimensional synthetic and real-world benchmarks, decision-level aggregation outperforms standard Bayesian optimization baselines in most cases while remaining competitive in the others.
PaperID: 1320, Poster
Abstract: Agentic large language models are increasingly used to solve real-world tasks by reasoning over goals, invoking tools, and interacting with external environments. Reinforcement learning provides a natural framework for improving these behaviors, and recent agent RL methods have achieved strong results across domains. However, the training dynamics of agent RL remain poorly understood, limiting our ability to diagnose instabilities and design more effective training algorithms. In this work, we identify a previously underexplored phenomenon in agent RL, which we term \emphcyclical entropy eruption. Unlike single-turn reasoning RL, where entropy typically collapses and stays low, agent RL training exhibits unique recurring cycles of sharp entropy eruption and gradual subsidence. We decompose this dynamic into three phases and provide theoretical and empirical analyses of each, explaining the mechanisms underlying its cyclical oscillation. We further show that degenerate patterns such as sentence duplication and hallucination, once acquired during eruption, can persist and accumulate across cycles. Motivated by these findings, we propose SEAL (Separation-Enhanced Agent Learning), a lightweight auxiliary loss that separates correct and incorrect trajectories in representation space, directly targeting the root cause of entropy eruption. Experiments across multiple benchmarks, models, and RL algorithms demonstrate that SEAL stabilizes training and yields stronger downstream agent performance.
PaperID: 1321, Poster
Authors:
Guy Hadad, Haggai Roitman, Moshe Eliasof, Bracha ShapiraAbstract: Different optimizers impose different geometries on parameter updates. Adaptive methods such as Adam and AdamW use coordinate-wise updates, while matrix-aware optimizers such as Muon utilize weight-matrix structure through orthogonalized updates. For Graph Neural Networks (GNNs), however, the effective geometry of its weight matrices is shaped not only by its learned channel-mixing weights, but also by the graph-propagated representations it acts on. Thus, even if an update is well normalized as a matrix, its effect on the layer output can be amplified when applied to graph-propagated features. To address this gap, we introduce GraphMuon, a graph-aware optimizer that adapts Muon's matrix updates to the covariance structure of graph-propagated layer inputs. GraphMuon performs orthogonalization in a metric defined by these propagated representations and normalizes step sizes according to the update's induced effect on the layer representation. We show that theoretically, vanilla Muon can suffer from graph-induced update amplification, while our GraphMuon can mitigate this effect. As such, our GraphMuon yields a graph-metric analogue of Muon's polar update step. Empirically, we show that GraphMuon achieves 2.13×-7.38× speedups in time-to-target wall-clock runtime over AdamW across diverse graph-learning regimes, while maintaining competitive or superior predictive performance.
Abstract: Auto-regressive models (ARMs) have established a dominant paradigm in language modeling. However, their strictly sequential sampling paradigm imposes fundamental constraints on both inference efficiency and modeling flexibility. To address these limitations, diffusion-based large language models (dLLMs) have been proposed, offering the potential for parallel sampling and flexible language modeling. Despite these advantages, current dLLMs sampling strategies rely primarily on token level information, which fails to account for global sequence structure and often yields suboptimal results. In this paper, we study the sampling order selection problem from the perspective of log-likelihood maximization. We show that this problem is NP-hard and propose an optimal sampling-rank-based approximation that makes the objective computationally tractable. We further prove that the tractable objective is optimized by sampling tokens in descending order of their attention-matrix column sums. This finding provides a principled justification for attention-guided sampling and offers a theoretically grounded alternative to greedy search. We instantiate this theoretical insight in a new training-free sampling algorithm, termed Attn-Sampler, and further propose dynamic attention thresholding for practical acceleration. Extensive experiments across multiple benchmarks validate the effectiveness of our proposed method, demonstrating that it achieves superior generation quality while enhancing the sampling parallelism.
PaperID: 1323, Poster
Authors: Iakovos Tenedios, Yashar Moshfeghi
Abstract: Sentence-level EEG-to-text decoding models typically pad variable-duration neural signals to a fixed-length window, creating a structural confound where segment duration correlates with sentence identity independently of neural content, a problem we term length bias. Despite being acknowledged in prior work, the prevalence, magnitude, and architectural dependence of this effect have never been systematically quantified. We introduce length-matched noise baselines that isolate this confound by replacing only the true signal region with Gaussian noise while preserving zero-padding, and apply them across three English listening EEG datasets under a strict subject-held-out protocol. Evaluating ten sentence classifiers and three generation models based on the current state-of-the-art, we show that reported decoding scores are substantially inflated by length cues across datasets, with the degree of inflation varying markedly by architecture. Among classifiers, BIOT alone is unaffected by length-matched noise. A systematic decomposition of its components identifies Short-Fourier-Transform (STFT) tokenisation as the bigger source of this invariance. We transfer this finding into two new generation variants that largely eliminate encoder-side length bias, enabling a controlled analysis of where residual bias originates. A classification-to-generation evaluation further reveals that leading generation models operate largely as sentence retrieval systems, with a linear classifier recovering nearly all of their reported BLEU-4 scores. Together, these results establish that length bias is a pervasive and previously unquantified failure mode in EEG-to-text evaluation, and provide both diagnostic tools and architectural directions for future work.
PaperID: 1324, Poster
Abstract: The Locate-Then-Edit framework enables efficient knowledge updates in large language models without requiring full retraining. Recent research has increasingly focused on sequential editing, where knowledge arrives continuously and the model is updated incrementally. However, existing sequential methods commonly suffer from catastrophic forgetting and model collapse after only a few hundred edits. In contrast, batch editing---which performs joint optimization over all edits and represents the theoretical optimum---can support tens of thousands of edits with high edit efficacy while preserving the model's general capabilities, resulting in a substantial performance gap between the two editing paradigms. In this paper, we demonstrate that this gap can be closed via a batch-equivalent sequential editing approach. Moreover, we investigate the underlying mechanism of model degradation in existing sequential methods by analyzing the convergence behavior of weight perturbations as edits accumulate. Our analysis reveals that current methods typically exhibit linear or logarithmic growth in perturbation, whereas the proposed batch-equivalent method achieves time-invariant growth, thereby theoretically guaranteeing long-term stability and preserving the integrity of pretrained knowledge throughout massive-scale sequential editing.
Abstract: Lattice field theory is the workhorse of non-perturbative physics, used to simulate phenomena from the strong nuclear force to critical phenomena in materials. Its Boltzmann distributions are parametrized analytically by \emphcoupling constants, but these bare parameters are weak predictors of physical observables---extracting physics typically requires extensive simulation. While machine learning tools such as normalizing flows have emerged as effective samplers at fixed couplings, it remains difficult to interpret what the underlying neural networks have learned. This raises a natural question: can the flow \emphparameters themselves be generated for new theories, and can their physics be read off directly from the network weights? We propose lattice field theory as a testbed for neural network interpretability: because the target physics is qualitatively well-understood and smoothly varying, it provides ideal synthetic data against which network behavior can be checked against known ground truth. To this end we introduce JEPAWG, a Joint-Embedding Predictive Architecture--based Weight Generator that maps couplings directly to flow weights via a learned latent space. On a scalar theory at lattice sizes 6^2 and 8^2, the JEPAWG latent space recovers the correct intrinsic dimension of the underlying manifold, identifies the region of phase transition, encodes a finite-size shift aligned with the 2D Ising exponent \nu \approx 1, allowing us to uncover physical structure by studying the network weights alone. As a generator, JEPAWG also interpolates and extrapolates to unseen couplings effectively and remains robust to weight-space incongruences deliberately introduced by combining multi-seed training data, outperforming PCA, AE, and VAE baselines.
PaperID: 1326, Poster
Abstract: Vanilla SVGD is known to have a computational cost of \mathcalO(N^2) per iteration. In this work, we introduce Random-Projection Tree Stein Variational Gradient Descent (RP-SVGD) to alleviate this computational challenge. This is achieved by restricting kernel interactions to spatially proximal particles clustered via a random projection tree, further reducing to a cost of \mathcalO(N d(\log (N/C)+C)), where C is a hyperparameter representing the maximum leaf node capacity of the spanning tree. To establish theoretical validity, we introduce a smoothed RP-Tree kernel for any fixed tree realization and prove that it belongs to the Stein class of the target distribution, thereby ensuring the generation of valid gradient flows. In addition, by considering the effective kernel as the expectation of random tree partitions, we verify that it preserves regularity along with other necessary conditions, which guarantee convergence under the approximate gradient flow framework. Extensive experimental results demonstrate that RP-SVGD tends to have competitive performance and significant speedups across various tasks.
PaperID: 1327, Poster
Abstract: Traditional differential privacy assumes each training example affects only one user's privacy, but many real-world datasets contain examples attributed to multiple users. We study \emphfixed-graph differential privacy, a framework for learning from such multi-attribution data where the attribution structure is public but example contents are private. We establish fundamental limits and optimal algorithms for this setting. First, we develop approximation algorithms for the NP-hard contribution bounding problem: a sampling-based LP-rounding algorithm achieving O(|V|^1/(k+1))-approximation and a greedy algorithm achieving O(r)-approximation, where r is the maximum number of users per example. We prove matching hardness results showing the greedy bound is tight up to O(\log r) factors. Second, we introduce a \emphnetwork density parameter c that measures how concentrated the attribution network is, and establish information-theoretic error lower bounds scaling as \tilde\Omega(c^2d/\varepsilon^2) for learning d-dimensional models from n examples. We provide matching upper bounds using contribution bounding combined with DP-SGD, demonstrating our algorithms are optimal. We further characterize when contribution bounding introduces selection bias and provide conditions under which bias vanishes. Finally, we extend our analysis to the i.i.d. setting where examples are sampled from a distribution, obtaining improved bounds scaling as \tildeO(\sqrtc/n) for hypothesis testing and providing matching algorithm for clustered networks.
PaperID: 1328, Poster
Authors: Zeyu Yang, Wes Armour
Abstract: Despite the critical need for sustainable deep learning, model energy efficiency remains largely neglected in Hardware-Aware Neural Architecture Search (HW-NAS). Discovering energy-efficient models is severely bottlenecked by costly on-device measurements and the brittleness of training-free accuracy proxies, which often yield unbalanced "glass cannon" architectures. We propose Energy-Efficient Neural Architecture Search (EENAS), a strictly zero-shot HW-NAS framework. EENAS introduces a PCA-based anomaly router that dispatches tasks to decision trees or tabular transformers, achieving <10% MAPE zero-shot energy prediction on entirely unseen GPUs. Simultaneously, we robustly estimate accuracy by ensembling training-free proxies via a novel Log Z-score aggregation. Combining these predictors unlocks Absolute Hardware Bounding, discovering Pareto-optimal architectures for strict energy budgets entirely offline in under 30 minutes. Evaluated on ImageNet-1K, our models achieve 40% less energy consumption at comparable accuracy, or a 2.5% accuracy improvement under the same energy budget compared to efficient baselines.
PaperID: 1329, Poster
Abstract: Part-level 3D assets are essential for editing, reassembly, and interaction, yet recovering such structure from a single image remains challenging due to occlusion, ambiguous boundaries, and the need for coherent multi-part reasoning. Existing approaches struggle to achieve both controllable part-level generation and coherent multi-part structure, as part identity and spatial allocation are typically inferred implicitly. We present Seg3DParts, a segmentation-grounded framework for controllable part-level 3D generation from a single image. By treating segmentation as an explicit grounding signal, our method defines part identity during generation, enabling each component to be anchored to a corresponding image region. To ensure coherent assemblies, we introduce structured cross-part interaction that allows components to exchange global context throughout the generative process. As a result, Seg3DParts directly generates well-aligned part meshes in a shared canonical space without post-hoc alignment, supporting flexible and controllable decomposition. We further introduce PartObjectNet, a large-scale dataset with over 200K objects and 1M annotated parts. Experiments demonstrate that Seg3DParts achieves superior geometry quality, cross-part coherence, and part-level controllability over existing methods.
Abstract: Recent advances in machine learning have spurred significant interest in learning-augmented algorithms, particularly for online optimization. A growing body of work has studied in this framework, aiming to characterize the trade-off between robustness and consistency. While this trade-off is fully understood for deterministic algorithms, a gap between upper and lower bounds remains in the randomized setting. In this paper, we close this gap by presenting a Pareto-optimal randomized learning-augmented algorithm for this problem. Our approach introduces the notion of a , a novel framework for representing the distribution over bids generated by an algorithm. We show that any bidding algorithm can be reduced, without loss of generality, to one driven by a bidding profile, and we characterize the optimal profile via a system of delayed differential equations. Finally, we demonstrate the broader applicability of our approach by extending it to the problem, yielding a significant improvement over prior learning-augmented algorithms for linear search.
PaperID: 1331, Poster
Abstract: Neural causal models can simulate interventions in complex, high-dimensional settings, but typically require a known causal graph. Observational data, however, generally identifies only a Markov equivalence class of graphs, represented by a CPDAG, and causal queries may vary across DAGs in that class. We introduce Masked Neural Causal Models (NCMs), a provably expressive nonparametric framework for simulating and bounding interventional queries over all models compatible with a given CPDAG given observational data. We give an optimization objective in the space of Masked NCMs that asymptotically recovers the true bounds of the causal effect. To enable this optimization in practice, we introduce an attention-based architecture and a novel optimization strategy that recovers highly accurate bounds in discrete nonparametric settings.
Abstract: In recent years, many high-stakes societal decisions are made by machine learning systems, often under strict capacity constraints that limit resources to a small subset of individuals. In such settings, information gaps across populations can lead to highly unbalanced allocations. We study prediction-driven resource allocation through the lens of fairness. We connect classical algorithmic fairness notions with resource allocation definitions, and characterize the trade-offs between fairness and utility. We introduce an adaptation of proportional fairness to this setting, showing that it yields a continuum of fairness criteria via a regularization parameter, along with quantitative bounds on the resulting price of fairness. We further show that max-min fairness and equal opportunity can incur an unbounded price of fairness in extreme cases, and propose a variant of equal opportunity with a bounded price of fairness.
Abstract: Detecting LLM reasoning failures at inference time without ground-truth labels motivates a family of confidence baselines, including self-consistency, semantic entropy, and P(\textTrue), built on within-question sampling and self-evaluation. Operad theory, the formalism for systems built by iterated substitution, suggests a complementary diagnostic: a model's direct answer to a compositional query should agree with the answer it produces by composing a stated decomposition of the same query. We instantiate this idea as operadic consistency (OC), a per-question signal. Across twelve instruction-tuned LLMs (4B--671B parameters, open-weights and closed-source) on four multi-hop QA datasets, OC is strongly correlated with accuracy on every dataset (Pearson r \in [+0.86, +0.94], all p \leq 0.0004), making it a substantially stronger population-level accuracy proxy at roughly one-third the inference cost of the strongest sample-based baseline (whose own cross-model rates never exceed r=+0.80 on any dataset). At the per-question level, OC contributes information beyond self-consistency, semantic entropy, P(\textTrue), and constructed decomposition-aware baselines on every dataset (cluster-robust p \leq 10^-19). We translate the regression coefficients into deployment-ready selective-prediction lifts (\Delta_\textAUARC up to +0.088, \Delta_\textAUROC up to +0.167 over a tuned \textSC_K=3 baseline). On five frontier thinking models, where the decomposition is extracted from the model's own chain of thought, the same equal-cost comparison gives positive selective-prediction lift on all 16 (dataset, budget, metric) cells tested.
Authors: Hyunsuk Kim, Isaac Gibbs, Michael Mahoney, Ryan Tibshirani
Abstract: The task of constructing a predictor that is simultaneously accurate under multiple evaluation metrics has been popularized recently under the name omniprediction. Prior work in this area has largely focused on binary prediction tasks, and existing results for multiclass and continuous outcomes typically suffer from slow error rates. In this paper, we consider forecasts expressed as a set of quantiles over a discrete but otherwise arbitrary set of probability levels. We develop sample-efficient algorithms for constructing forecasts that are simultaneously accurate across all proper losses that satisfy mild regularity conditions, where proper losses are those minimized by predicting the true quantiles. Our algorithms are online and can operate in a dynamic environment while simultaneously taking multiple losses into account, and their forecasts offer the potential to provide diverse end-users with rich information for evaluating risk and informing a range of practical decision-making tasks. As an example, we demonstrate how our procedure can be used to ensemble quantile forecasts of COVID-19 hospitalizations made by experts during the pandemic.
PaperID: 1335, Poster
Abstract: Image verification is an increasingly critical need as the capabilities for visual media tampering improve. We present an approach to image verification that embeds a private authenticity signature into an image based on physical refractive objects, which transform the image. To verify a protected image, we compare the image and a pixel-aligned reconstruction from the embedding to identify inconsistencies. Our approach is inspired by prior work, which physically places refractive objects in a scene before taking a photo. However, prior work is limited to simple refractions with analytical formulas and requires a slow per-scene neural radiance field optimization to get a reconstruction. Instead, our method trains a compact, scene-agnostic neural refraction field capable of modeling complex and more secure refractive geometries. Once trained, it enables instant, high-fidelity reconstruction for manipulation detection and localization.
Abstract: Diffusion models are central to generative modeling and have been adapted to graphs by diffusing adjacency matrix representations. The challenge of having up to n! such representations for graphs with n nodes is only partially mitigated by using permutation-equivariant learning architectures. Despite their computational efficiency, existing graph diffusion models struggle to distinguish certain graph families and their spectra, unless graph data are augmented with ad hoc features. This shortcoming stems from enforcing the inductive bias within the learning architecture. In this work, we leverage random matrix theory to analytically extract the spectral properties of the diffusion process, allowing us to push most of the inductive bias from the architecture into the dynamics. Building on this, we introduce the Dyson Diffusion Model, which employs Dyson's Brownian motion to capture the spectral dynamics of an Ornstein-Uhlenbeck process on the adjacency matrix. Furthermore, conditioned on the spectral dynamics, we formulate a Lie group diffusion, appropriately modeling the remaining degrees of freedom. Strikingly, the resulting learning problem becomes permutation invariant at the Lie algebra level. We demonstrate that the Dyson Diffusion Model learns graph spectra accurately and outperforms existing graph diffusion models.
PaperID: 1337, Poster
Abstract: Forecasting the behavior of spatio-temporal systems often requires more than accurate predictions: models should uncover the dynamical modes governing future evolution. Existing temporal graph and latent dynamical models achieve strong forecasting performance, but their latent representations often entangle multiple dynamical patterns, limiting interpretability. Conversely, regime-inference methods explicitly model such modes but typically do not support generative forecasting over graph-structured systems. We introduce HERON, a self-supervised regime-aware generative framework that unifies graph-structured trajectory generation with regime identification. HERON factorizes the latent state into discrete system-level regimes, capturing global dynamical modes, and continuous entity-level dynamics, yielding a modular regime interface compatible with existing spatio-temporal backbones. Crucially, the inferred regime actively conditions system dynamics, enabling regime tracking and controlled simulations under alternative dynamical modes. Without regime annotations, HERON jointly learns forecasts and regime dynamics in a self-supervised manner, recovering meaningful modes across controlled synthetic and real-world datasets while maintaining competitive forecasting performance. These results open a path toward interpretable, controllable generative modeling of spatio-temporal dynamics.
PaperID: 1338, Poster
Authors:
Shafiei, Azade Farshad, Nassir Navab, Yousef YeganehAbstract: Conditioning diffusion models on visual features is essential for controlled image generation and counterfactual reasoning across diverse applications, from creative design to medical imaging. Existing conditioning methods, such as timestep embedding and cross-attention mechanisms, suffer from limited control over feature accuracy, require substantial architectural modifications, and struggle to maintain conditioning precision throughout the denoising process, as the model's focus shifts toward fine-grained details in later steps rather than the desired conditioned features. We present a novel approach that addresses these limitations by learning to shift the initial noise distribution in the latent space. Our method assigns each visual feature a distinct direction in the latent space and applies proportional shifts during both training and inference, enabling precise control over continuous and discrete visual attributes. We introduce a combined loss function that explicitly enforces shift accuracy alongside denoising, ensuring consistent conditioning across all denoising steps. Furthermore, we extend this framework to datasets with hidden or inaccessible visual features by employing a Variational Autoencoder to extract latent representations in an unsupervised manner. This enables high-fidelity counterfactual generation on complex, real-world datasets where explicit feature annotations are unavailable or prohibitively expensive to obtain.
PaperID: 1339, Poster
Authors: Dai Songyan
Abstract: We prove that the speedup of exact speculative decoding is governed by a per-token rate, not by the path-level prefix-overlap capacity that folklore suggests, and that the gap between the two can be unbounded. The two candidates are the per-token Leviathan rate σ L always, the inequality is strict already on i.i.d. products, and there exists an explicit fixed-precision binary trajectory-exact rejection-style protocol on L = 2 with E[τ] = 19/16 > 9/8 = σ L as a universal upper bound across all trajectory-exact rounds. Inside the deployable admissible Leviathan-style clipped-ratio class, σ M^u,T, Q) + 1)), achieved by iterated Leviathan and matched up to constants by Wald. On the path side, Ω L ≤ 1/(1−ρ) uniformly in L, with no unconditional KL converse. Two explicit fixed-precision Transformer families of size O(log L) realize the two extremes under the same three-element proposal family Q = Q const_1: a constant-readout family F⁺ has SDC ≤ T / ((1−ε)L + 1), while an alternating-parity family F⁻ has SDC ≥ T/5, an unconditional Θ(L) architecture separation. Speculative decoding is not path coupling; it is per-token clipped-ratio acceptance, and the architecture of the target controls the rate.
Authors:
Yuchen Zhu, Jing Shi, Chongjian GE, Hao Tan, Yiran Xu, Wanrong Zhu, Jason Kuen, Ryan Rossi, Koustava Goswami, Rajiv Jain, Yongxin Chen, Molei Tao, Jiuxiang GuAbstract: Autoregressive (AR) large language models have achieved broad practical success, but their sequential decoding remains a major bottleneck for low-latency deployment. Efficiency efforts have advanced along two largely orthogonal axes: hybrid attention architectures that reduce the cost of each forward pass, and diffusion language models (dLLMs) that enable parallel token generation. Existing dLLMs, however, have yet to translate their theoretical decoding parallelism into commensurate real-world throughput, have not incorporated modern hybrid attention backbones, and still trail scale-matched AR models on quality. In this work, we present FLARE, a systematic recipe for converting hybrid AR LLMs into capable, real-time-fast dLLMs under a practical training budget. Through a controlled study, we identify transfer data quality as the dominant driver of AR-to-dLLM performance loss, outweighing loss formulation and attention-mask design of which are emphasized by prior works. To enable efficient training and deployment of a hybrid softmax-plus-linear-attention backbone, we develop hardware-aware Triton kernels for diffusion-style linear attention and an SGLang-based inference system that exposes both AR-Trust and Diffusion-Trust decoding from the same checkpoint. Starting from Qwen3.5 checkpoints with only 10B tokens from public datasets, FLARE-9B matches the top-tier open-source dLLM LLaDA-2.1 Flash (100B-A5B) at 1/10 of the parameters, exceeding it on GPQA-Diamond (71.2 vs. 66.7), and FLARE-2B reaches 2.2 times the throughput of LLaDA-2.1 mini (16B-A1B) on GSM8K at eight-way concurrency on a single A100.
PaperID: 1341, Poster
Abstract: In practice, the reasoning strategy often matters more than the model. Rather than calling a fixed large model once per input, a well-designed strategy breaks the problem into steps and uses the right model for each, achieving better accuracy at lower cost. Existing methods have tried to automate strategy design by searching over a fixed set of strategy templates, but the templates themselves limit what reasoning patterns can be discovered. We propose Strategist, a meta-agent that designs a high-performing inference strategy for a given task. Each strategy is a structured executable topology that defines how agents plan, interact, verify, and route across models. Given a task and a small development set, Strategist proposes candidate strategies, runs them on real samples, and revises them against measured accuracy and cost. Every strategy and its components are reusable, forming an evolving library that grows with each new task and compounds what future tasks can build from. On 30 benchmarks, Strategist matches or beats the strongest baselines with an average of 8.5 absolute percentage points higher accuracy at 48% lower inference cost.
PaperID: 1342, Poster
Abstract: Continuous DAG learning formulates combinatorial structure learning as differentiable optimization over weighted adjacency matrices. In these methods, the estimated adjacency matrix depends critically on a sparsity-controlling regularization parameter, whose magnitude governs the recovered graph skeleton. While extensive empirical evidence underscores the centrality of this parameter, existing practice relies on solving a sequence of independent optimization problems over a grid of tuning parameters, which is computationally wasteful, statistically unstable, and provides only a fragmented view of how the adjacency matrix evolves. In this work, we study the regularization path of stationary solutions in continuous DAG learning and propose DAG-flow, an exact path-following framework that compactly encodes the entire family of estimators. We show that for prominent formulations including NOTEARS, GOLEM, and DAGMA, stationary solutions evolve according to a piecewise-smooth matrix-valued dynamical system once the active support is fixed. Our analysis derives the explicit systems governing these branches, identifies the events at which the path changes regime, and clarifies how nonsmooth sparsity and nonconvex acyclicity interact along the path. Unlike grid-based search, DAG-flow exposes the geometry of model variation across regularization levels, which enables the construction of structured candidate graph sets and provides direct insight into edge-level stability. Across synthetic and real benchmarks, our DAG-flow outperforms grid-search baselines at a fraction of the compute and consistently improves structural stability.
Abstract: We introduce \emphSigned Rectified Flow (Signed RF), a generalization of Rectified Flow that targets a signed measure \pi^\mathttsign = (1+\alpha)\,\pi^+ - \alpha \,\pi^-, where \alpha>0, \pi^+ represents the distribution to promote, and \pi^- represents the distribution to suppress. Although sampling from a signed measure is not well-defined, Signed RF induces a valid generative process that concentrates on the positive region of \pi^\mathttsign while provably excluding regions dominated by the negative component. This yields a principled framework for incorporating negative information and exclusion constraints into generative modeling. Theoretically, we analyze the signed continuity equation underlying Signed RF and explain how negative mass creates exclusion barriers through a charged-particle interpretation. Empirically, Signed RF leads to practical adaptive guidance algorithms. Across applications, Signed RF improves the fidelity--diversity trade-off on ImageNet, reduces nearest-neighbor similarity in anti-memorization stress tests, and mitigates adversarial-prompt nudity in SD 3.5 while preserving CLIP and aesthetic scores.
PaperID: 1344, Poster
Authors: Wenxuan Guo, JungHo Lee, Panos Toulis
Abstract: We study the problem of detecting treatment effects in randomized A/B experiments when the effects are potentially small and complex. A common approach in such settings is to fit machine learning (ML) models of outcomes on observed features and then apply classical t-tests to residualized outcomes. We show that this approach can suffer substantial power loss under treatment effect heterogeneity. We propose a randomization-based test whose statistic measures the out-of-sample predictive gain from including the treatment variable in a flexible ML model. Leveraging experimental randomization and sample splitting, our test is finite-sample valid for arbitrary models and loss functions. We establish power guarantees under heterogeneous treatment effects and demonstrate substantial empirical gains in simulations and large-scale A/B experiments.
Authors:
Jinhua Lyu, Tianmin Yu, Brian Kim, Lizhuo Zhou, Chanwook Park, Naichen ShiAbstract: Direct diffusion modeling of high-resolution spatiotemporal fields is computationally challenging. Parameter-efficient primitives address this by representing high-dimensional data with a compact set of parameters. In this paper, we construct data-dependent tensor primitives without pretrained compression autoencoders. Our construction starts from Tucker decomposition, which captures low-rank multilinear structure through a core tensor and mode-wise factors. However, Tucker factors are non-unique: the same tensor can be represented by different rotated factors, which complicates generative modeling. We address this issue with orthogonal Procrustes (OP) alignment. Specifically, we select medoid anchor matrices from the data and align the factor matrices to resolve the gauge ambiguity. This yields matrix Grassmannian primitives and tensor Grassmannian primitives that are compact, data-adaptive, and directly decodable by explicit multilinear reconstruction. Theoretically, we prove that the proposed primitive maps are homeomorphisms between low-rank tensors and their corresponding primitive spaces, certifying that the representations are non-degenerate and topologically faithful. Building on these primitives, we propose Diffusion in Aligned Tensor Space (DiffATS), a generative framework that trains diffusion models directly on aligned tensor primitives. Across images, videos, and PDE solutions, DiffATS achieves strong unconditional and conditional generation performance while compressing original data by 3.9× to 210×, without relying on any pretrained deep compression autoencoders.
PaperID: 1346, Poster
Abstract: Speculative sampling is a widely used method for losslessly accelerating Large Language Model (LLM) inference. State-of-the-art speculative sampling methods (e.g. EAGLE series) use a ``dynamic draft tree'' with drafter temperature implicitly set to T=0. Greedy drafts, however, perform poorly as temperature increases. Moreover, we show that naively instantiating the dynamic draft-tree with drafter temperature T>0 leads to an incorrect distribution. Fixes exist, but are computationally prohibitive as they require O(|\mathcalV|^k) computation, where \mathcalV and k denote vocab and number of child nodes. We address these limitations by introducing Gumbo, a new speculative sampling algorithm designed for high-temperature settings that uses only O(|\mathcalV|) compute. Gumbo is drafter-invariant in that it does not require the draft distribution during verification. Additionally, Gumbo increases the average acceptance length and speedup by expanding and reranking the dynamic draft tree based on Gumbel scores rather than draft probabilities. At the heart of Gumbo is a multi-draft generalization of the communication-free coupling principles and Gumbel sampling developed by Daliri et al., which is of independent interest. Replacing EAGLE-3's speculative sampler with Gumbo results in up to a +26.7% increase in speedup at a temperature of 1.0 across four models and five datasets, without altering the draft model or requiring any other code changes.
PaperID: 1347, Poster
Abstract: Bird's-eye-view (BEV) perception has emerged as a cornerstone of autonomous driving systems, providing a structured, ego-centric representation critical for downstream planning and control. However, real-world deployment faces challenges from sensor degradation and adversarial attacks, which can cause severe perceptual anomalies and ultimately compromise the safety of autonomous driving systems. To address this, we propose a resilient and plug-and-play BEV perception method (RESBev), which can be easily applied to existing BEV perception methods to enhance their robustness to diverse disturbances. Specifically, we reframe perception robustness as a latent semantic prediction problem. A latent dynamic predictor is constructed to extract spatiotemporal correlations across sequential BEV observations, thereby learning the underlying BEV state transitions to predict clean BEV features for reconstructing corrupted observations. The proposed framework operates at the semantic feature level of the BEV perception pipeline, enabling recovery that generalizes across both natural disturbances and adversarial attacks without modifying the underlying backbone. Extensive nuScenes experiments show RESBev notably enhances BEV perception robustness against natural disturbances and adversarial attacks, with latency, parameter and memory analyses verifying its excellent robustness-efficiency trade-off.
Abstract: We develop a framework for analyzing parameter symmetries in deep ReLU networks and obtain a complete characterization of the generic parameter fibers for three-layer bottleneck architectures. Our approach provides explicit semi-algebraic descriptions of these fibers and yields a polynomial time algorithm for deciding functional equivalence of two parameters. The symmetries include discrete and continuous transformations arising from layer composition, and depend on whether deeper layers hide or preserve geometric structure from preceding layers. Finally, we show that some of these symmetries induce local conservation laws along gradient flow, while others do not.
Authors: Yuhang He
Abstract: , that encodes an arbitrary spatially grounded 2D geometric shape into a compact representation exhibiting five favorable properties, including invertibility, adaptivity, generality and controllability. Specifically, a 2D spatially grounded geometric shape is decomposed into its normalized geometry within the unit disk and its pose vector, where the pose is further transformed into a harmonic pose field that also lies within the unit disk. A set of orthogonal Zernike basis is constructed to encode shape geometry and pose either independently or jointly, with controllable relative emphasis on shape geometry or shape pose. We demonstrate the theoretical validity, efficiency, discriminability, and wide applicability of via extensive analysis and experiments across a wide range of shape-aware tasks and our self-curated XShapeCorpus dataset. We envision as a foundational tool for research beyond one-dimensional sequential data and data-driven, learning-based encoding paradigms, paving the way for a unified spatial encoding framework for frontier 2D spatial intelligence.
PaperID: 1350, Poster
Abstract: LiDAR localization is a fundamental task in robotics and computer vision, aiming to estimate the global pose of point clouds. Although Scene Coordinate Regression (SCR) has demonstrated state-of-the-art performance in this field, standard SCR methods employ a monolithic network for uniform optimization across diverse scenes, which inevitably suffers from capacity interference in highly degraded scenes. Recent research on Vision-Language Models (VLMs) indicates that they can acquire rich scene understanding priors, providing the adaptive learning capabilities currently missing in localization networks. In this paper, we propose ViLo, the first framework to integrate VLM priors into SCR to fundamentally enhance localization robustness. Specifically, we design a temporally consistent key-frame querying mechanism to extract open-vocabulary degeneracy priors. We then construct a dynamic prior-guided Mixture-of-Experts (MoE) model for adaptive feature learning tailored to varying scenes. Finally, to avoid additional computational overhead during inference, ViLo internalizes the VLM's scene understanding into the 3D backbone via cross-modal distillation, enabling a lightweight, LiDAR-only deployment. Extensive experiments on the Oxford RobotCar and NCLT datasets demonstrate that ViLo significantly outperforms existing state-of-the-art methods, reducing positional errors by impressive margins of 20% and and 51%, respectively.
Abstract: We study the adversarial kernel bandit problem, in which the loss at each time step is induced by some bounded (but otherwise arbitrary) element of an RKHS. We propose an exponential-weights algorithm with a regularized importance-weighted estimator and an explicit correction term that controls the estimator's regularization bias. Our main result bounds the regret in terms of a widely-adopted notion of effective dimension that captures the complexity of the kernel. Notably, when the kernel satisfies a polynomial eigendecay condition with exponent \beta>1, our regret scales as T^(\beta+1) /(2 \beta) up to logarithmic factors. This matches a lower bound from existing work (Chatterji, Pacchiano, and Bartlett, ICML 2019) and strictly improves on the upper bound derived in that work (with dependence T^\beta /(2 (\beta-1))), while also having the benefit of dropping their rank-one adversary assumption and instead allowing arbitrary RKHS elements at each step.
PaperID: 1352, Poster
Abstract: Many deep learning models exhibit a double-descent phenomenon, where test loss exhibits a sharp peak near the interpolation threshold and then descends a second time as model size grows further, contrary to the U-shape predicted by the classical bias–variance tradeoff. The double-descent phenomenon has been found and explained in toy classical models, but is not well-understood in modern deep learning despite extensive empirical study. We interpret double descent through the novel lens of universal compression. Specifically, we upper bound the maximum-likelihood test loss as a sum of three terms: an approximation error, a minimax batch regret, and an MLE-vs-Bayes gap. We show that the minimax Bayes predictor, unlike the MLE, does not exhibit the double-descent peak. Furthermore, we isolate each term through experiments with uniform random labels, and show that double descent is primarily caused by the MLE-vs-Bayes gap in this setting; here, the MLE cross-entropy grows to be roughly five times larger at the double-descent peak, while the minimax Bayes predictor cross-entropy remains flat. Additionally, we show the possibility of a second ascent in test loss at very large model sizes, and motivate this theoretically through the minimax batch regret term of our decomposition.
PaperID: 1353, Poster
Abstract: A planner in a network of strategic agents faces three entangled challenges: the optimum depends on agents' private information, queried agents may misreport to steer the outcome, and exact computation does not scale. We study these challenges in multi-activity network games with heterogeneous private technologies, in which the planner sets non-discriminatory prices. We show that the optimal prices admit a centrality-based decomposition of the welfare kernel: each agent's contribution scales with its squared centrality in a network reweighted by agents' preferences across activities. This decomposition motivates , a polling algorithm in which the planner samples one agent per round, walks briefly through the agent's neighborhood, and updates the price from a local report. From the same decomposition flow three forms of efficiency: computationally, uses significantly fewer operations than exact computation and other distributed methods; statistically, its query complexity scales with topology and preference heterogeneity rather than with population size; and economically, it converges to welfare-maximizing prices while inducing truthful reports and detecting adversarial deviations.
Authors: Ludo Andrianirina, Mathieu Carrière
Abstract: Topological clustering, and its main algorithm ToMATo, is a clustering method from Topological Data Analysis (TDA) which has been applied successfully in several applications during the last few years. This is thanks to its high versatility, as clusters are detected from the persistent components in the sublevel sets of any user-defined function (gene expression, pixel values, etc), and efficiency, as topological clustering enjoys robustness guarantees. However, ToMATo is also limited in several ways. First, a graph on the data points needs to be provided as a hyper-parameter of the method (whose fine-tuning is left to the user). Second, ToMATo is known to be very sensitive to outlier values in the function range. Finally, and most importantly, ToMATo can only handle one function at a time, whereas it is critical to use several functions in various applications. In this article, we introduce ToMAToMP: the first topological clustering method able to handle several functions at the same time with theoretical guarantees. More specifically, we leverage a recent tool from multi-parameter persistent homology, called MMA decomposition, to design our clustering algorithm, and prove that it enjoys robustness properties. As corollaries, we show that it can be used to make ToMATo independent of graph tuning, and robust to outliers. Finally, we provide a set of numerical experiments showcasing the efficiency and quality of the clusterings produced by ToMAToMP, by showing strong improvement over non-topological and topological baselines for various datasets.
PaperID: 1355, Poster
Abstract: We investigate the problem of finding (\delta, \varepsilon)-stationary points (also called Goldstein stationary points) for nonsmooth nonconvex Lipschitz functions. This problem has recently garnered significant attention, with several non-asymptotic convergence guarantees. However, existing methods for such guarantees require prior knowledge of problem parameters and lack anytime convergence. Our work addresses these issues by providing an algorithm that satisfies two properties: Firstly, our algorithm is parameter-free; secondly, it generates a sequence of (\varepsilon, \varepsilon)-stationary points with monotonically decreasing \varepsilon. This algorithm achieves the (non-stochastic) state-of-the-art complexity for any \varepsilon > 0 and ensuring convergence to a Clarke stationary point in the limit. Central to our result is the insight that one can slightly modify prior algorithms for this problem and view them within the framework of the classic Frank-Wolfe algorithm.
PaperID: 1356, Poster
Abstract: Human pose estimation has advanced rapidly driven by large-scale adult-centric models. However, children remain significantly underrepresented due to data scarcity, distinct body morphology, and unique motion dynamics. We introduce ChildPose, a modular video foundation model designed for child-centric pose estimation. ChildPose transforms pretrained pose encoders into a streaming pediatric model via memory-augmented temporal attention and age-aware co-training. By maintaining a memory bank of adjacent frame representations, the model achieves robust localization under motion blur, occlusion, and atypical poses, while an auxiliary age-prediction objective enforces representations that capture developmental morphology. We curate a diverse, video-based pediatric dataset spanning various age groups and capture conditions, evaluating out-of-distribution generalizability on entirely unseen subjects. Across multiple 2D baselines, ChildPose consistently outperforms both off-the-shelf models and child-data fine-tuned variants, while ensuring stable sequence-level predictions. Our results underscore that effective pediatric pose estimation necessitates child-specific modeling beyond standard adult-centric estimators.
Abstract: Large language models are increasingly used as open-ended search operators in evolutionary optimization. We introduce Evolutionary Feature Engineering (EFE), a framework for using LLM-based evolution to discover preprocessing transformations for structured data. EFE represents transformations as Python programs with a standardized fit/transform interface, allowing them to be inserted directly into existing machine learning pipelines. During evolution, candidate programs are refined using dataset context, summary statistics, and downstream performance feedback on validation set. We instantiate EFE in two settings. For time-series forecasting, EFE-Time learns invertible, dataset-specific normalizations that improve off-the-shelf time-series foundation models. It reduces forecasting errors (MASE, WQL, MAE) 3% or more when averaged across datasets and improvements are as much as 19% on the COVID-Deaths dataset. Notably, these improvements occur under latest TSFMs such as Chronos 2. For tabular prediction, EFE-Tab evolves compact feature programs that add interpretable useful features and remove redundant ones, improving or matching existing LLM-based feature-engineering methods. We found EFE-Tab to be particularly effective on classical decision trees, where small sets of evolved features yield competitive accuracy while preserving interpretability. Overall, EFE demonstrates that LLM-based evolution can improve both accuracy and interpretability when automatically tackling structured data.
PaperID: 1358, Poster
Authors: Adway S Wadekar, Jake A Soloff
Abstract: Large-scale hypothesis testing supports probability claims about individual hypotheses, as in empirical Bayes methods for estimating local false discovery rates. We study how such claims can be interpreted as approximately calibrated forecasts of the null hypothesis, yielding interpretable error probabilities even under model misspecification. Our approach draws conceptual inspiration from probabilistic forecasting but addresses a different challenge: unlike forecasting, where labels are eventually observed, in multiple testing the ground truth is never revealed, so calibration must be assessed stochastically and established indirectly. We address this challenge by constructing a set of pseudo-labels, derived from the spacings of ordered p-values, which have the local false discovery rate as their regression target. Our construction unlocks existing tools for assessing and performing post-hoc calibration in multiple testing. Notably, we show that q-values, a popular error measure based on the false discovery rate, can be severely miscalibrated.
Authors: Pengzhou Wu
Abstract: We prove the identifiability of deep generative models (DGMs) with Piecewise-Affine (PWA) decoders and Gaussian Mixture Model (GMM) priors, in a purely unsupervised setting. We introduce three algebraic contrast principles for symmetry breaking: , which forbids parameter conspiracies between latent components and decoder branches. Together they exploit the interplay between the discrete combinatorics of the PWA map and the continuous symmetry structure of the latent GMM. Continuity is replaced by algebraic symmetry conditions; injectivity is decoupled from structural identification and required only for pointwise inversion. Our results form a hierarchy: from . The ICA-form ambiguity emerges under diagonal covariance conditions; classical ICA is a further specialization under independence. Assumptions are only on the data-generating process, not on learning methods, except for the interaction contrast.
Abstract: Symmetry-aware learning is typically framed as equivariance to a single group action, but many real-world distribution shifts are typed, partial/non-invertible, and compositional (e.g., modality changes, occlusion, subsampling, intervention chains). These shifts are better modeled by a category of transformations whose objects represent data contexts and whose morphisms represent admissible transformations between contexts. We thus extend symmetry-aware machine learning from groups to such transformation categories; enforcing categorical equivariance on neural architectures gives stronger robustness than group equivariance. We ground this in both theory and experiments. We formulate a universal approximation theorem for category-equivariant architectures and show density in the space of equivariant continuous transformations. We present experiments on compositional OOD shifts, demonstrating that enforcing categorical equivariance yields measurable robustness gains over group-equivariant and non-equivariant baselines.
PaperID: 1361, Poster
Abstract: Diversity maximization is a fundamental problem in combinatorial optimization with applications in fairness-aware machine learning, clustering, and data summarization, and has attracted significant attention in theoretical computer science. We study the Distant k-Subsets Problem (DKSP): given a metric space (X,d) with |X|=n and an integer k, find two disjoint subsets S,T\subseteq X, each of size k, maximizing \min_s\in S,\; t\in T d(s,t). This problem arises as a key subroutine in composable coresets for diversity maximization, and improvements can yield better approximations for related objectives such as remote-matching and pseudoforest. It also appears in practice: a variation was featured in the ICPC World Finals 2024. Mahabadi and Narayanan gave a (4/\epsilon)-approximation for this subproblem when n \ge 2k^1+\epsilon+k. For any \epsilon>0, we present a deterministic combinatorial algorithm achieving a \left(1+\left\lceil 1/\epsilon\right\rceil\right)-approximation under the same condition. In particular, for \epsilon=1, this yields a 2-approximation when n \ge 2k^2+k; the factor 2 is optimal for general metrics unless P=NP. More generally, any polynomial-time (2-\eta)-approximation for constant \eta>0 would imply an algorithm for the k-Biclique problem, even though DKSP is polynomial-time solvable when n=2k. For d-dimensional Euclidean metrics, we obtain improved guarantees: a (2+\sqrt2)-approximation when n \ge 3\cdot 2^d-1k, and a (2+\sqrtd/c)-approximation when n \ge (2d+c^d)k. In particular, for d=2 and c=1, this gives a (2+\sqrt2)-approximation when n\ge 5k, yielding near-2 guarantees under linear constraints on n for constant d.
Abstract: Enforcing functional inequality constraints such as monotonicity and convexity in neural networks is a fundamental challenge in many industrial and scientific applications. Classical one-sided penalty methods, just like primal-dual methods gated by complementary slackness, provide gradients only at violated locations, resulting in fragile constraint satisfaction. Hard architectural constraints, on the other hand, cover only the simplest cases and tend to compromise the network's expressive power. We introduce , a deep learning-native primal-side approach that converts constraint enforcement into a two-sided regression problem by coupling the primary network with a jointly learned auxiliary network. The auxiliary network serves as a valid target for the primary network's constraint quantities, inducing feasibility and regularity. Theoretically, neural slack variables can be understood as a variable splitting method with parametrized variable projection. Neural slack variables close the gap to hard constraints on monotonicity and low-dimensional convexity, and maintain effectiveness on nonlinear constraints in neural barrier certificates and no-arbitrage volatility surfaces.
PaperID: 1363, Poster
Abstract: Thompson sampling is natural randomized strategy for optimizing unknown functions in a black-box manner. In continuous global optimization, Thompson sampling has largely been restricted to Gaussian process priors, as these are essentially the only class for which computing the posterior distribution has been tractable in practice. In this work, we introduce Prior-fitted Diffusion Transformers, which are diffusion models that perform Thompson sampling at inference time, after being pre-trained using the prior-fitted network paradigm. To apply these models, we develop a set of practical non-Gaussian priors for general black-box functions - including those that have characteristics that cannot be achieved under Gaussianity, such as priors over unimodal functions. We also quantify the tradeoff between model size and the number of pre-training iterations, and show that these models follow scaling laws akin to those observed in transformers defined over other modalities. Across a set of global optimization benchmarks, we show that prior-fitted diffusion transformers with non-Gaussian priors can achieve improved performance compared to traditional pipelines built on Gaussian processes.
PaperID: 1364, Poster
Authors:
Aabhash Dhakal, Tim Lachner, Jayanta Mandi, Marco Foschini, Christoph M FlathAbstract: Decision-focused learning (DFL) improves on the Predict-then-optimize (PtO) pipeline by training predictive models to minimize decision regret rather than prediction loss. Most DFL methods assume offline full-feedback data with complete action cost vectors. Many real-world decision systems instead operate in an online setting: in each round, the agent acts using past feedback and observes costs only for the selected components, yielding semi-bandit feedback. Bandit methods optimize regret in this regime, but train models with predictive updates rather than decision-focused updates. In this paper, we formalize online semi-bandit DFL: decision-focused learning for online optimization problems with semi-bandit feedback. The most straightforward extension would impute unobserved components with the model's own predictions and compute the decision-focused gradient. But this creates a self-confirming imputation degeneracy: imputation errors can flip counterfactual decisions that define the decision-focused gradient, biasing the representation update and reinforcing errors in later imputations. To address this degeneracy, we propose BayesianSPO, which combines an analytically updated Bayesian last-layer head with masked decision-focused learning. The Bayesian head supplies both (i) uncertainty to balance exploration and exploitation and (ii) stable imputations for unobserved components, while masking unobserved component gradients prevents those imputations from directly corrupting the representation update. Across knapsack, shortest-path, and portfolio benchmarks, BayesianSPO achieves the lowest mean cumulative regret on all five instances and the improvement is statistically significant on four of them.
Abstract: Functional data such as time series, trajectories, and fields are increasingly central in machine learning and are naturally modeled in infinite-dimensional Hilbert spaces. We study Semi-dual Neural Optimal Transport (SNOT) for learning transport maps between probability measures on separable Hilbert spaces. A critical challenge in this setting is the spurious solution problem: a neural map can globally optimize the SNOT objective while failing to recover an optimal transport map. We show that this failure arises from non-regular source measures, which make the inner minimization in the semi-dual objective non-identifiable. We prove that regular source measures, defined via Gaussian null sets, restore inner-minimizer uniqueness and recover the Monge map. For singular sources, we introduce Gaussian smoothing and establish a necessary and sufficient kernel condition for smoothing to restore regularity. We further prove plan-level consistency: as smoothing vanishes, every Wasserstein accumulation point of the induced optimal plans solves the original Kantorovich problem. Experiments on synthetic singular transports and real-world unpaired time-series imputation benchmarks validate the theory, achieving the best performance on several benchmarks and competitive performance on the remaining ones.
PaperID: 1366, Poster
Abstract: Recent advances in diffusion policies have demonstrated strong effectiveness when integrated into hierarchical skill-based learning frameworks, leveraging offline data to capture multimodal and temporally abstracted behaviors. However, prior approaches have rarely considered the diffusion process as a learnable component, adopting fixed noise schedules applied uniformly across skills, regardless of their distinct behavioral contexts. Such a design fundamentally limits the achievable performance, particularly for contact-rich and dexterous manipulation tasks where different skills exhibit distinct couplings between action dimensions. In this paper, we present a skill-adaptive noise diffusion framework (SaND) which introduces a learnable, skill-conditioned noise schedule that modulates the diffusion process per action dimension, adapting the noising dynamics to each skill. To prevent the learned noise schedule from collapsing into trivial solutions, we employ an auxiliary inverse objective that reconstructs the skill embedding from the Legendre coefficients of the derivative of the noise schedule, encouraging the schedules to remain skill-discriminative. This formulation naturally enables adaptive control over the denoising steps for each skill, where noise schedules are clustered in the Legendre coefficient space, and the appropriate number of denoising steps is determined at the point where reconstruction error abruptly increases. At inference, we identify each skill embedding's cluster from its noise schedule and apply the corresponding number of denoising steps, dedicating more compute to skills requiring fine-grained control and less to simpler ones. Through experiments, we demonstrate that SaND outperforms existing diffusion-based skill learning and planning baselines on contact-rich and dexterous manipulation tasks, achieving higher success rates while requiring fewer denoising steps.
PaperID: 1367, Poster
Abstract: Search-based planning is a powerful policy improvement paradigm in reinforcement learning, with milestone successes on discrete-action domains through the AlphaZero and MuZero families. Extending it to continuous control, however, faces a structural mismatch between the search operator and the policy class: search at each state evaluates a finite candidate set and returns a discrete distribution, while the parametric Gaussian policy commonly used in continuous control is a continuous density. We prove that distilling the search output into a Gaussian induces a non-vanishing projection error that inflates policy variance and produces a return gap that persists even when the policy mean is locally optimal, and is not closed by additional search compute or data. Inspired by the iterative refinement of MPPI in continuous control, we propose Iterative Gumbel Planning (IGP), which adapts this principle to a discrete policy class. IGP represents the policy as a per-dimension factorized categorical distribution aligned with the search support, eliminating the projection error and reducing the residual approximation to a bounded discretization error controllable by the grid resolution. Joint candidates are drawn by Gumbel Top-K sampling without replacement, avoiding the V^|A| combinatorial explosion of joint discretization, and search-refine iteration runs for multiple rounds at the same state. On DMControl and HumanoidBench, IGP achieves state-of-the-art performance, with improved sample efficiency and lower temporal action variance at evaluation.
PaperID: 1368, Poster
Abstract: Recent advancements in the length generalization theory have provided us with the ability to reliably predict learnability by transformers. In particular, the C-RASP hypothesis (a formalized version of the so-called RASP-l conjecture) posits that transformers length-generalize on a task if and only if a solution is expressible in the C-RASP language. While this hypothesis has strong empirical validation, theoretical problems arise from the fact that no computable length generalization bounds exist for C-RASP, as well as the discovery of seemingly contradictory empirical results. To address this, we refine the C-RASP hypothesis utilizing the recently-proposed fragments of the language, CRASP+ and CRASP1. These fragment have computable length generalization bounds, though in the worst case requiring an extremely large (double exponential) sample size. It is an open question whether this sample size bounds are tight. In this paper, we resolve this open question by providing an exponentially tighter bound. In doing so, we show a polynomial length generalization bound for transformers if we adopt compressed strings, via a novel connection to power words. As an application, we show how this yields a fine-grained analysis of C-RASP conjecture that resolves contradicting experimental evidence against it.
PaperID: 1369, Poster
Abstract: Organoids are an increasingly central scientific platform for drug screening, disease modeling, and regenerative medicine. Because organoid form is closely tied to function, AI-driven image analysis has drawn intense interest, yet every existing organoid imaging study evaluates representations on a single dataset and no work has quantitatively compared representations across the breadth of public organoid datasets. We close this gap with OrgBench, a benchmark of 12 vision foundation models (VFMs) on 13 public organoid-imaging datasets spanning detection, segmentation, classification, and temporal outcome prediction. The benchmark reveals a structured domain gap: task-family winners shift across the benchmark and parameter count is a poor predictor of transfer. Building on this, we introduce OrgFM, a small bottleneck adapter that we continually pre-train on unlabeled organoid images while keeping a general-purpose vision transformer frozen. On the OrgBench task-family aggregates, OrgFM improves over its frozen base backbone. Layer- and feature-level interventions localize the gains to a small set of late-layer morphology-sensitive features: the adapter sparsely supplements, rather than replaces, the base representation. Together, this is the first cross-task, cross-backbone study of organoid imaging at scale; the OrgBench/OrgFM pairing establishes that the parameter-efficient adapter direction is a promising path for organoid representations and clarifies why it works.
Abstract: Autoregressive (AR) language models generate text one token at a time, even when consecutive tokens are highly predictable given earlier context. We introduce MARS (Mask AutoRegreSsion), a lightweight fine-tuning method that teaches an instruction-tuned AR model to predict multiple tokens per forward pass. MARS adds no architectural modifications, no extra parameters, and produces a single model that can still be called exactly like the original AR model with no performance degradation. Unlike speculative decoding, which maintains a separate draft model alongside the target, or multi-head approaches such as Medusa, which attach additional prediction heads, MARS requires only continued training on existing instruction data. When generating one token per forward pass, MARS matches or exceeds the AR baseline on six standard benchmarks. When allowed to accept multiple tokens per step, it maintains baseline-level accuracy while achieving 1.5-1.7x throughput. We further develop a block-level KV caching strategy for batch inference, achieving up to 1.71x wall-clock speedup over AR with KV cache on Qwen2.5-7B. Finally, MARS supports real-time speed adjustment via confidence thresholding: under high request load, the serving system can increase throughput on the fly without swapping models or restarting, providing a practical latency-quality knob for deployment.
Abstract: Modern LLM serving systems use prefix caching to accelerate multi-turn conversational workloads. They cache the key-value states of previous conversation prefixes and reuse them when subsequent requests share the same prefix, to avoid redundant prefill computation and reduce serving latency. However, achieving a high cache hit rate remains challenging because cache reuse depends on whether a conversation continues and how soon the next turn arrives. Our analysis of real workloads shows that user intent is a strong signal for this behavior, as different intents exhibit different continuation probabilities and inter-turn delays. Based on this insight, we propose Intent-Aware Caching (IAC), an online eviction policy that classifies requests by intent, learns lightweight reuse statistics from recent serving logs, and uses them to guide an oracle-inspired eviction decision. Extensive evaluation on real workloads (e.g., ShareChat) shows that IAC reaches a 34.6% cache hit rate, substantially outperforming the ML-based LPC (18.8%) and LRU (12.2%), while reducing average serving latency by up to 54.6%.
PaperID: 1372, Poster
Abstract: Lipid Nanoparticles (LNPs) are widely used as delivery systems for nucleic acid therapeutics, where transfection efficiency is determined by both the identities of constituent lipid components and their composition ratios. While prior studies have focused on learning molecular representations for ionizable lipids, modeling how multiple components and their ratios jointly influence LNP performance remains underexplored. In this work, we propose STRATA, a framework that models molecule interaction between LNP components, which is known to contribute to LNP transfection efficiency. Our approach is built on two complementary views: (1) a ratio-centric view that captures interaction patterns induced by composition ratios through a transformer with a Ratio-induced Positional Embedding, and (2) a molecule-centric view that incorporates interaction-induced effects into structure-based molecule embeddings. By jointly training and aligning these views, our model integrates molecular structure and composition ratio within a unified framework that captures interaction-driven effects. Experiments demonstrate that our method improves prediction accuracy and generalization to unseen molecules and ratios, highlighting the effectiveness of our approach.
PaperID: 1373, Poster
Abstract: Although the mechanics of eigenvalue bias of self-supervised learning (SSL) in linear networks are well-understood, they remain theoretically disconnected from empirical shortcut learning, offering little guidance or intuition. In this paper, we present a theoretical analysis of this shortcut learning phenomenon through the lens of extent bias and amplitude bias, two forms of eigenvalue bias. By investigating the relations among extent bias, amplitude bias, and learning priorities in SSL, we demonstrate that SSL prioritizes features based on dimensionality (object to image ratio) and amplitude (backdoor attack intensity), not their semantic importance. Our analysis reveals how the eigenvalues of the feature cross-correlation matrix influence which features are learned earlier, providing insights into why models preferentially learn shortcut features over more generalizable features.
Abstract: Private Evolution (PE) is a differentially private algorithm for synthetic data generation. While it can be viewed as a Wasserstein learning algorithm, it performs much better in practice than worst‑case Wasserstein analyses would predict. We recast PE as a generative model‑augmented Wasserstein learning. We show theoretically that when we take into account the use of a generative model that is able to capture something about the true distribution, then PE provably obtains much better performance bounds. For example, if the generator gives samples in the same low-dimensional space as the distribution, then the sample complexity depends on the intrinsic, not the ambient, dimension. We also show that standard variants of PE can fail to converge on simple well-clustered instances, and propose a new geometry-aware version of PE with provable convergence on such instances. Experimentally, we show that our new algorithm has consistent empirical gains over standard baselines.
PaperID: 1375, Poster
Abstract: Learning unified representations from single-cell multi-omics data is a fundamental challenge, yet existing deep learning models often overlook the inherent hierarchical structure of biological regulation. In this work, we introduce scWavelet, a novel deep learning framework that leverages a multi-scale inductive bias by learning representations in the cellular spatial-frequency domain. Our approach is grounded in a heuristic approximation-misspecification decomposition, which theoretically motivates the design of a scale-factorized architecture. scWavelet employs a learnable wavelet transform to decompose omics signals into distinct scales, which are then processed by scale-specific Omics Mixers to generate integrated latent representations. Extensive experiments demonstrate that scWavelet achieves state-of-the-art performance on challenging multi-omics integration and cross-omics translation benchmarks. Furthermore, interpretability analyses confirm that the learned multi-scale features effectively correspond to the biological hierarchy of cellular identity. Our project will be publicly available.
Abstract: Depth super-resolution (DSR) aims to recover a high-resolution (HR) depth map from its low-resolution (LR) counterpart. With color image guidance, this task is typically formulated as learning the residual between HR and LR in a low-dimensional feature space. However, this additive formulation is insufficient to accurately capture the complex relationship between HR and LR, especially under spatially varying degradations. In this paper, we introduce DegBins, a novel DSR framework that leverages degradation-driven binning to adaptively enhance residual modeling. Specifically, DegBins reformulates the regression-based DSR as a hybrid classification-regression problem, where the residual depth is represented as a linear combination of discrete depth bins weighted by their learned probability distribution, yielding more flexible and expressive representations. Furthermore, DegBins models the degradation relationship between HR and LR in a high-dimensional feature space, enabling adaptive bin range adjustment and probability optimization conditioned on local degradation characteristics. To progressively improve reconstruction quality, DegBins adopts a multi-stage refinement scheme, where each stage performs finer-grained bin partitioning and probability updating based on the former estimation. This coarse-to-fine design facilitates more accurate depth recovery, particularly in regions with severe degradations or complex structural variations. Extensive experiments across five benchmarks demonstrate that DegBins consistently outperforms existing state-of-the-art methods in terms of accuracy, robustness, and generalization.
PaperID: 1377, Poster
Abstract: Deep learning models trained to optimize average accuracy often exhibit systematic failures on particular subpopulations. In real-world settings like healthcare, the subpopulations most affected by such disparities are frequently unlabeled, partially observed, or not known in advance. Existing group-robust methods typically assume prior knowledge of the relevant subgroups, using group annotations for training, validation, or model selection. We propose Misclassification Aware Rank-Limited Adaptation (MARLA), a parameter-efficient method for improving worst group performance without explicit subgroup annotations. MARLA leverages an ERM-trained model by calculating the model's misclassification probability scores on a held-out adaptation set to identify a low-dimensional subspace where errors concentrate. We then learn a rank-restricted additive correction to the classifier logits within that subspace. Across seven real-world datasets, we evaluate group robustness under three settings: no knowledge of subgroup relevance, partial knowledge of subgroup relevance, and full knowledge of subgroup relevance. MARLA improves worst-group performance while remaining fast, parameter-efficient, and practical to tune.
Abstract: Vital sign measurement using cameras presents opportunities for comfortable, ubiquitous health monitoring. Remote photoplethysmography (rPPG), a foundational technology, enables cardiac measurement through minute changes in light reflected from the skin. However, practical deployment is limited by the computational constraints of performing analysis on front-end devices and the accuracy degradation of transmitting data through compressive channels that reduce signal quality. We propose a memory efficient rPPG algorithm - FacePhys - built on temporal-spatial state space duality, which resolves the trilemma of model scalability, cross-dataset generalization, and real-time operation. Leveraging a transferable heart state, FacePhys captures subtle periodic variations across video frames while maintaining a minimal computational overhead, enabling training on extended video sequences and supporting low-latency inference. FacePhys establishes a new state-of-the-art, with a substantial 49% reduction in error. Our solution enables real-time inference with a memory footprint of 3.6 MB and per-frame latency of 9.46 ms -- surpassing existing methods by 83% to 99%. The training code, a real-time inference demo for mobile browsers, and the demonstration video can be found in the supplementary materials.
Abstract: We introduce SODA, a generalization of Optimistic Dual Averaging, which provides a common perspective on state-of-the-art optimizers like Muon, Lion, AdEMAMix and NAdam, showing that they can all be viewed as optimistic instances of this framework. Based on this framing, we propose a practical SODA wrapper for any base optimizer that eliminates weight decay tuning through a theoretically-grounded 1/k decay schedule. Empirical results across various scales and training horizons show that SODA consistently improves performance without any additional hyperparameter tuning.
Abstract: LLMs have shown impressive success in program synthesis, discovering programs that surpass previously known solutions. However, these approaches rely on simple numeric scores to signal program quality, such as the value of the solution or the number of passed tests. Because such a score offers no guidance on a program failed, the system must generate and evaluate many candidates per problem in the hope that some succeed, increasing both LLM inference and evaluation costs. We study a different approach: property-guided LLM program synthesis. Instead of scoring programs after evaluation, we check whether a candidate satisfies a formally defined property. When the property is violated, we stop the evaluation early and provide the LLM with a concrete counterexample showing exactly how the program failed. This feedback drastically reduces both the number of program generations and the evaluation cost, and can guide the LLM to generate stronger programs. We evaluate this approach on PDDL planning domains, asking the LLM to synthesize heuristic functions: every state reachable by strictly improving transitions has a strictly improving successor. A heuristic with this property leads hill-climbing search directly to a goal state. A counterexample-guided repair loop generates one candidate program, checks the property over a training set, and returns the first case that violates the property. We evaluate our approach on ten planning domains with an out-of-distribution test set. The synthesized heuristics are effectively \emphdirect on virtually all test tasks, and compared to the best prior generation method our approach generates seven times fewer programs per domain on average, solves more tasks without using search, and requires several orders of magnitude less computation to evaluate candidates. Whenever a problem admits a verifiable property, property-guided LLM program synthesis can reduce synthesis and evaluation cost while improving program quality.
PaperID: 1381, Poster
Authors: Guanhe Huang, Songqiao Han, Tangzheng Lian, Oya Celiktutan
Abstract: Latent diffusion models paired with a variational auto-encoder (VAE) tokenizer have shown promising performance for efficient text-driven human motion generation. However, motion latent diffusion suffers from a generation-reconstruction trade-off: simply increasing tokenizer capacity enhances reconstruction but fails to proportionally improve text-conditioned generation quality. We assume this trade-off stems from a spectral mismatch inside the VAE tokenizer: the encoder yields a latent space biased toward high frequencies, and the decoder introduces temporal artifacts. In this paper, we propose the Diffusable Motion Tokenizer (DiMoT) to improve the diffusability of the tokenizer, rather than scaling up the diffusion backbone. To improve text-conditioned optimization, Critic Spectral Adaptation (CSA) leverages a pretrained diffusion model as the critic to suppress high-frequency latent energy. To preserve motion dynamics, the Alpha-Flow Motion Decoder (AMD) performs generative decoding via average-velocity prediction, enabling efficient one-step decoding. Extensive experiments demonstrate that DiMoT achieves state-of-the-art performance across various challenging text-to-motion benchmarks, yielding superior motion fidelity and text-motion alignment. The code will be available upon publication.
Authors: Philipp Schneider, Daniel Kuhn
Abstract: Constrained end-to-end learning trains neural networks whose outputs must satisfy feasibility constraints, such as resource, budget, or operational limits. While projection and optimization layers guarantee constraint satisfaction, boundary-based projections can introduce unfavorable optimization geometry: regions of the unconstrained output space are mapped to the same active face of the feasible set. This makes the resulting map locally rank-deficient, suppressing gradients in constrained directions and degrading optimization—a bottleneck known as gradient saturation. We propose , a differentiable reparameterization layer that maps unconstrained network outputs into the relative interior of a convex feasible set by smoothly rescaling rays from a strictly feasible anchor point. Unlike hard radial or orthogonal projections, the resulting map is one-to-one and has a full-rank Jacobian almost everywhere, while guaranteeing strict feasibility. We prove that constrained networks equipped with Soft-Radial Projection retain universal approximation guarantees. For standard convex sets such as simplices and balls, the layer admits a closed-form forward pass, avoiding the iterative solvers required by optimization-based layers. Experiments on decision-focused learning benchmarks show improved optimization stability and solution quality compared with optimization- and projection-based baselines.
PaperID: 1383, Poster
Abstract: Modern gradient boosting algorithms are based on Newton's method, which excels empirically but lacks global convergence guarantees. While recent work on Gradient-Regularized Newton (GRN) boosting achieved a global \mathcalO(1/\gamma^12k^2) rate, its analysis was not able to exploit the natural Hessian geometry and suffers from a \gamma^12 dependence on a worst-case `weak gradient edge' \gamma. In this paper, we tackle these theoretical drawbacks by generalizing the Affine-Invariant Cubic Newton (AICN) scheme to convex optimization with weak learners. Our analysis is based on average cosine values \overline\Theta_k along the optimization path and directional improvements \delta_k in the Hessian-induced norm. We show that the weak-AICN update reduces to a damped weak-Newton step with an explicit stepsize determined by \delta_k. Under semi-strong self-concordance, we establish a global \mathcalO(1/\overline\Theta_k^4 k^2) rate, significantly improving upon the weak-GRN bound. We further prove a local linear rate, global linear convergence for Hessian-dominated losses, and cosine angle lower bounds. Finally, we demonstrate the broader applicability of this framework beyond gradient boosting by showing that Newton coordinate descent emerges naturally as a special case of optimization with weak learners.
Abstract: Dataset pruning reduces the storage and training costs of deep learning by selecting an informative subset from a large dataset. However, most existing pruning methods require fully labeled data, which limits their applicability in realistic settings where unlabeled data are abundant and annotation is costly. Recent label-free pruning methods address this issue, but they rely on features from pre-trained models to estimate example difficulty. This dependence can be unreliable when the target dataset differs substantially from the pre-training distribution. We propose a label-efficient dataset pruning framework that, given only a small labeled subset, uses semi-supervised learning to generate pseudo-labels for unlabeled data, allowing existing supervised pruning methods that require label information to be seamlessly applied to the resulting pseudo-labeled training pool. We then estimate example difficulty from pseudo-label-induced training dynamics and select a coreset. By learning directly from the target dataset, our method better captures the target distribution and provides more reliable signals for difficulty estimation and coreset selection. We validate our approach on domain-specific, image-corrupted, and long-tailed datasets, where it achieves state-of-the-art performance among label-free and label-efficient baselines, while also demonstrating competitive performance on standard benchmarks.
Abstract: Multi-agent systems (MAS) coordinate multiple LLM-powered agents through structured workflows, gaining reasoning power but incurring high inference latency from multi-step execution and repeated model invocations. Existing orchestration methods primarily optimize task performance and inference cost, leaving latency largely unaddressed. In MAS, end-to-end latency is governed by the critical execution path, so reducing total cost alone does not reliably reduce latency. Moreover, optimizing latency while preserving accuracy remains non-trivial: naive latency optimization can misassign operator-level credit and degrade task accuracy. To address this gap, we propose Latency-Aware Multi-agent System (LAMaS), a latency-aware orchestration framework for learning-based multi-agent systems. LAMaS addresses this challenge at two levels: at \emphtraining time, it learns latency-aware execution graphs through constrained optimization with critical-path-aware credit assignment; at \emphinference time, since a graph committed at training time cannot exploit runtime evidence, it complements graph construction with a lightweight controller that adaptively eliminates redundant future agent interactions as execution unfolds. Experiments on four benchmarks show that LAMaS achieves the best latency among evaluated learning-based MAS baselines, reducing end-to-end latency by over 50% while maintaining competitive or better accuracy. LAMaS is also modular and transfers to other MAS with minimal changes, consistently yielding latency reductions.
Abstract: Mixture-of-Experts (MoE) has become the de facto architecture for hundred-billion-parameter language models, yet its advantages at sub-billion scales for on-device deployment remain largely unexplored. To close this gap, we present MobileMoE, a family of on-device MoE language models with sub-billion active parameters (0.3-0.9B active and 1.3-5.3B total) that establish a new Pareto frontier for on-device LLMs. We first formulate an on-device MoE scaling law that jointly optimizes MoE architecture under mobile memory and compute constraints, identifying an on-device sweet spot -- moderate sparsity with fine-grained and shared experts -- that is simultaneously memory and compute-optimal. Building on the derived architectures, we train MobileMoE with a four-stage recipe covering pre-training, mid-training, instruction fine-tuning, and quantization-aware training, all on open-source datasets. Across 14 benchmarks, MobileMoE matches or exceeds leading on-device dense LLMs with 2-4× fewer inference FLOPs, and matches or surpasses the state-of-the-art MoE OLMoE-1B-7B with up to 60% fewer parameters. To bridge the last mile to mobile deployment, we provide the first efficient MoE inference on commodity smartphones with comprehensive on-device profiling. At comparable total parameters, MobileMoE delivers 2-4× faster prefill and decode, and ~30% lower runtime peak RAM than the dense baseline MobileLLM-Pro.
PaperID: 1387, Poster
Authors: Ege Gursoy, Nicolas Mansard
Abstract: A central limitation of modern machine learning systems is the lack of formal certificates guaranteeing their behavior beyond empirical, sample-based evaluation. Sum-of-squares (SOS) methods provide such guarantees by certifying polynomial inequalities over continuous domains via semidefinite programming (SDP). While prior work has introduced SDP layers as differentiable components in neural networks, their scalability is limited by two key bottlenecks, that we are fixing in this paper: (i) failure to preserve sparsity from the forward SOS program in the backward pass, leading to inefficient training, and (ii) loss of certificate structure during differentiation, preventing its use in learning. We propose a differentiable sparse SOS layer whose backward pass mirrors the sparse structure of the forward solve, reusing local blocks, overlaps, and sparse constraints. The method avoids building dense certificates for scalar losses and computes gradients using the same sparse structure when certificate details are required. Experiments show that this enables practical SOS-based layers within learning systems, reducing backward computation time by one to three orders of magnitude while preserving certification quality. We evaluate on sparse polynomial certification, pendulum control, combinatorial relaxations, robotics projection, oscillator chains, and pose graph synchronization. Results demonstrate that sparse certificate objectives can be effectively integrated into learning pipelines when the backward pass respects the verifier's structure.
PaperID: 1388, Poster
Abstract: Large Language Models remain plagued by hallucinations. Recent work has sought to tame their prevalence using statistical techniques based on conformal prediction, with both theoretical and empirical success. However, these methods operate in a post-hoc fashion, treating the sampling procedure itself as atomic and then surgically altering samples to remove hallucinated claims. This disconnect between filtering and generation can result in samples that are incoherent, inconsistent, or simply unlikely under the model itself. Not only that, but post-hoc surgery is unable to shift probability mass towards more useful and helpful responses. To address these issues, we propose to instead sample from approximations to an LLM posterior, where the conditioning event corresponds to a calibrated, high-scoring region. We develop a calibration procedure tailored to the setting of conditional sequential generation that effectively identifies this region and achieves target risk control. Empirically, we apply our method to case studies focused on open-ended biography generation and mathematical problem solving; compared to prior work, we obtain the same statistical guarantees, with higher downstream utility.
Abstract: Machine learning (ML) is increasingly central to scientific discovery, but translating ML predictions into rigorous scientific claims still requires statistical inference on gold-standard labeled data. A setting that has emerged across the sciences provides the analyst with a small labeled sample, a much larger unlabeled sample, and a pre-trained predictor, with the goal of constructing a confidence interval for a quantity of interest, as was formalized in Angelopoulos et al. (2023). In this ML-driven scientific discovery literature, labeled data is typically the binding constraint on precision while computation is comparatively unconstrained, motivating the design question of how additional computation can extract more inferential value from a fixed labeled set, a fixed unlabeled set, and a fixed predictor. To address this question, we propose the Prediction-Augmented Residual Tree (PART) estimator. PART takes the same inputs as the original PPI estimator and replaces their single global rectifier with locally computed rectifiers obtained from an adaptive partition of the feature space. We show that PART outperforms existing methods by producing tighter confidence intervals across real-world datasets from ecology, astronomy, and census reports, among other domains. We then describe and analyze PAQ, an estimator that arises when considering the limit of PART when the depth of its tree grows to infinity. Under appropriate assumptions in the input data, we show that the variance of PAQ shrinks at a rate of O(N^-1 + n^-4), showing that there are settings where the rate of O(N^-1+n^-1) of existing methods can be provably broken.
PaperID: 1390, Poster
Abstract: We introduce PD-RAG, a technique for retrieval-augmented generation (RAG) that leverages a language model's own randomness to safeguard privacy. The algorithm partitions documents into groups, generates a candidate answer from a randomly chosen group, and then uses a privacy test to enforce that the released answer could have been produced by multiple disjoint document groups, thereby ensuring a controllable degree of plausible deniability. Compared to existing methods that operate at the token level and therefore incur both substantial utility loss and growing privacy budget per generated token, PD-RAG operates at the answer level, so the privacy it offers does not loosen with output length. We prove that PD-RAG satisfies (\varepsilon,\delta)-differential privacy. We experimentally evaluate PD-RAG on three QA benchmarks using three language models and find that it reduces membership inference attack advantage to near random while consistently outperforming alternative methods on all four utility metrics we used and running between 17 and 33× faster.
Authors:
Kun Fang, Qinghua Tao, Junxu Liu, Yaxin Xiao, Yulin Jin, Qingqing Ye, Jian Sun, Haibo HuAbstract: Machine Unlearning (MU) aims at removing the influence of specific data from a pretrained model while preserving performance on the remaining data. In this work, a novel perspective for MU is presented upon low-dimensional feature subspaces, which gives rise to the potentials of separating the remaining and forgetting data herein. This separability motivates our LOFT, a method that proceeds unlearning in a LOw-dimensional FeaTure subspace from the pretrained model through principal projections, which are optimized to maximally capture the information of the remaining data and meanwhile diminish that of the forgetting data. During unlearning, LOFT simply optimizes a small-size projection matrix flexibly plugged into the pretrained model, and only requires one-shot feature fetching from the pretrained model. LOFT demonstrates that competitive and efficient MU can be achieved in low-dimensional feature subspaces even with neither repeated raw-data access nor updates to the entire pretrained model. Extensive experiments validate the significantly lower computational overhead and superior unlearning performance of LOFT across diverse models, datasets, tasks, and applications. Code is anonymously available at https://anonymous.4open.science/r/4352/.
Abstract: Dense Associative Memory (DenseAM) is a promising family of AI architectures that is represented by a neural network performing temporal dynamics on an energy landscape. While hyperparameter transfer methods are well-studied for feed-forward networks, these methods have not been developed for settings in which weights are shared across layers and within the layer, which is common in DenseAMs. Additionally, DenseAMs utilize rapidly peaking activation functions that are rarely used in feed-forward architectures. The confluence of these aspects makes DenseAM a challenging framework for using existing methods for hyperparameter transfer. Our work initiates the development of hyperparameter transfer methods for this class of models. We derive explicit prescriptions for how the hyperparameters tuned on small models can be transferred to models trained at scale. We demonstrate excellent agreement between these theoretical findings and empirical results.
PaperID: 1393, Poster
Abstract: Finding the optimal sample complexity of realizable PAC learning was a major open problem in learning theory, which was settled by the breakthrough result of Hanneke (2016), building on prior work by Simon (2015). Hanneke's algorithm is quite involved and requires training \textpoly(n) Empirical Risk Minimizers (ERMs) on carefully crafted subsets of the data. Since using a single ERM call is provably suboptimal, recent work by Aden-Ali (2024) proposed the majority-of-three ERMs as a candidate for the simplest optimal PAC learner, and established its optimality in the in-expectation regime. Their main open question was whether this algorithm is optimal in the classical high-probability regime. In this work, building on their ideas, we answer this question affirmatively.
Authors: Xinye Yang, Zhenyang Liu, Yuxuan Wang, Yuanyuan Lei
Abstract: Recent research-level physics-reasoning benchmarks for LLMs reveal a failure mode that lies not in knowledge or execution, but in modeling the world being reasoned about. When faced with unseen problem setups, models often silently alter given quantities, fabricate laws, or inject unstated assumptions to force the problem into familiar patterns, and subsequent reasoning then proceeds within this corrupted world. We define this premise-level divergence as modeling drift: the gap between a model's internal axiomatic world model (AWM) and the true physical world specified in the problem. Although RLVR and PRM have yielded substantial gains in mathematics and code, their reward shapes do not directly supervise world modeling in axiomatic domains such as physics. We propose Reinforcement Learning with Axiomatic World Modeling (RLAWM), which optimizes the policy by penalizing divergence between its inferred axiomatic world model and the true physics world. RLAWM probes physics reasoning at two hierarchical levels (modeling and solving), each scored by a physically grounded reward. The modeling reward targets abstract knowledge formulation, evaluating whether the model captures the true axiomatic world before attempting calculation. The solving reward targets instantiated knowledge reasoning, checking whether the model's concrete derivation remains consistent with the axiomatic world. This two-level structure separates what world the model believes it is reasoning in from how it reasons within that world. The framework operates exclusively during training, with no extra cost during inference. RLAWM improves performance by up to +10.8 points on PhysReason and +5.2 on PHYSICS over the baseline. RLAWM also exhibits strong zero-shot generalization, yielding a 2.45× gain over the strongest baseline on unseen physics problems.
PaperID: 1395, Poster
Authors: Krish Mody, Sridhar Radhakrishnan, Chandra N Sekharan
Abstract: Many systems represent graphs as sequences by visiting nodes one at a time, for example, a molecule written as a SMILES string or a program written as a token stream from an AST traversal. The same graph can be traversed in many ways, producing different sequences for the same underlying object. Sequence Transformers treat these as distinct inputs and produce distinct embeddings, even though the graph remains unchanged. Existing graph-native Transformers address this by operating directly on the graph, but many practical pipelines receive serialized inputs and cannot be redesigned. We provide a theoretical characterization in three parts. Tree-positional encoding achieves structural invariance for acyclic graphs under canonical-root serialization but is insufficient for cyclic graphs due to spanning-tree non-uniqueness. Augmenting with cycle-membership encoding derived from the canonical cycle basis introduces a traversal-invariant anchor on cycle vertices. The resulting Tree+Cycle Positional Encoding is a drop-in replacement for sinusoidal PE within graph-aware serialization pipelines (those preserving token-to-atom alignment for structural tokens), requiring no architectural changes to the sequence model. It applies to any domain where labeled graphs are given as sequences with this alignment property. We validate the approach on two testbeds. On synthetic cycle counting and ring membership tasks, it outperforms the GIN baseline by 2× on cycle counting and matches it on ring membership, demonstrating that pre-computed cycle structure encoded as PE matches message-passing GNNs at end-task performance. On molecular property prediction, it achieves 0.977 \pm 0.029 mean cosine similarity across re-rootings (n = 24,960), outperforms a parameter-matched graph-native Transformer (Graphormer) on BACE and BBBP, and enables fragment-level interpretability with +5 points on alert alignment over a methodologically matched sinusoidal baseline.
Abstract: Data augmentation is widely recognized for improving generalization in deep networks, yet its impact on the geometry of learned representations remains poorly understood. In this work, we characterize how different data augmentation strategies reshape internal representations in neural networks. Using tools from shape analysis, we embed network hidden representations into a metric space where distance is invariant to scaling, translation, rotation and reflection. We show that increasing augmentation strength leads to well-behaved trajectories in this space, and that different augmentation types steer representations in distinct directions. Moreover, we investigate how neural representation shapes are distorted along data augmentation trajectories, and show that insights from neural geometry can predict which representations provide the most improvement when ensembling models. Our results reveal shared geometric patterns across architectures and seeds, and suggest that analyzing shape-space trajectories offers a principled tool for understanding and comparing data augmentation methods.
Authors: Hyeseon An, Shinwoo Park, Dongsu Kim, Yo-Sub Han
Abstract: LLM-based agents act through sequences of executable decisions, but their trajectories provide little evidence of which agent or policy produced them, making provenance, ownership, and unauthorized reuse difficult to establish from observed behavior alone. This motivates watermarking signals embedded directly into agent behavior rather than only into generated text, since text watermarking cannot capture the action-level decisions that define agent execution. Recent agent watermarking methods address this gap by moving the watermark from generated text to behavioral choices. However, by treating each action step as an independent trial, they overlook trajectory structure and become fragile when trajectories are perturbed, truncated, or observed without reliable alignment. We propose \textttSeqWM, a sequential behavioral watermarking framework that embeds signals into history-conditioned transition patterns and verifies trajectories position-agnostically against random-key baselines. Experiments across diverse agent benchmarks and LLM backbones show that \textttSeqWM consistently achieves reliable detection while preserving agent utility, and remains robust under trajectory corruption where round-indexed behavioral watermarks collapse. Our code is available at https://anonymous.4open.science/r/seqwm-0897.
Abstract: Root cause analysis (RCA) in complex systems is challenging due to error propagation across multiple variables, the need for structural causal knowledge, and the computational cost of inference at test time. We introduce PRIM (Prior-fitted Root cause Identification with Meta-learning), a causal meta-learning approach that frames RCA as a Bayesian inference task over a synthetic prior of causal models. By marginalising out structural uncertainty, PRIM implicitly identifies changes in the data-generating mechanism between baseline and anomalous periods. In doing so, PRIM infers distributional differences without explicit statistical testing, and implicitly learns causal structure without model fitting at test time. Following the simulation-based meta-learning paradigm of prior-fitted networks, PRIM uses a Model-Averaged Causal Estimation (MACE) transformer neural process that jointly attends over observational and anomalous samples and the causal structure of nodes, enabling zero-shot inference in 17\,ms for systems with up to 100 variables. Across synthetic benchmarks and two realistic benchmark datasets, PetShop and CausRCA, PRIM is competitive with methods that are aware of the system's causal graphical structure a priori while outperforming graph-unaware methods on several tasks. Lightweight fine-tuning to specific domains and data dynamics improves performance further.
Abstract: This paper proposes a guaranteed defense method for large language models (LLMs) to safeguard against jailbreaking attacks. Drawing inspiration from the denoised-smoothing approach in the adversarial defense domain, we propose a novel smoothing-based defense method, termed Disrupt-and-Rectify Smoothing (DR-Smoothing). Specifically, we integrate a two-stage prompt processing scheme—first disrupting the input prompt, then rectifying it—into the conventional smoothing defense framework. This disrupt-and-rectify approach improves upon previous disrupt-only approaches by restoring out-of-distribution disrupted prompts to an in-distribution form, thereby reducing the risk of unpredictable LLM behavior. In addition, this two-stage scheme offers a distinct advantage in striking a balance between harmlessness and helpfulness in jailbreaking defense. Notably, we present a theoretical analysis for generic smoothing framework, offering a tight bound for the defense success probability and the requirements on the disruption strength. Our approach can defend against both token-level and prompt-level jailbreaking attacks, under both established and adaptive attacking scenarios. Extensive experiments demonstrate that our approach surpasses current state-of-the-art defense methods in terms of both harmlessness and helpfulness.
Abstract: Today's reasoning models use thinking tokens to attain stronger performance on benchmarks than their instruction-tuned counterparts. It is also generally believed that this more "deliberative" mode should improve alignment and safety, by providing the model a safe space to consider whether its planned answer to the request violates its safety principles. We present evidence that this intuition is not always correct. Across frontier open-weight reasoning models spanning GPT-OSS, Phi, OLMo, and Qwen model families, we find that the model's decision is already strongly encoded at the beginning of thinking, with a probe on the first token's hidden representation predicting refusal/compliance with \ge0.85 AUROC and ~90% balanced accuracy. We also find little evidence that current models can use their thinking trace to deliberate about safety, as additional thinking after the first 20% of the trace rarely moves the final decision. While sentence-level inspection of thinking traces may show signs of oscillation between refusal- and compliance-leaning rationales, we find that in \geq 85% of thinking traces, such oscillations exert limited to no influence on the final response. We also examine the effect of existing inference-time and training-based safety interventions and find that they largely alter thinking behavior by shifting models toward more refusal-leaning reasoning while substantially reducing helpfulness on benign prompts. Our results suggest that safety behavior in current reasoning models is much less deliberative than commonly assumed, highlighting the need for training methods that more effectively utilize thinking traces for safety-critical decision making.
PaperID: 1401, Poster
Abstract: Test-time scaling (TTS) improves large language model (LLM) reasoning by spending additional inference-time compute, most commonly through width-heavy repeated sampling and aggregation. However, this strategy scales cost almost linearly with the number of calls and can remain brittle when the sampled trajectories share the same failure mode. We introduce Maze, a test-time scaling framework that shifts computation away from repeated target-model calls and toward LLM-free planning over ordered few-shot exemplars. The central idea is to treat an ordered exemplar sequence as a prompt trajectory embedded in a latent geometric space. Instead of asking the LLM for many independent attempts, Maze first ranks candidate prompt trajectories using cached encoder representations and a path scorer, \it i.e., a lightweight learnable metric, then spends one call (or a few calls) only on the most promising trajectories. This design preserves black-box deployability, since the planner is decoupled from the target LLM and can be implemented with a separate frozen encoder. Across various challenging reasoning benchmarks, single-call Maze consistently outperforms strong single-call prompt-optimization baselines, and MultiMaze recovers a large fraction of width-heavy TTS gains with substantially fewer computational cost measured by number of LLM calls and latency. We further present a theoretical perspective showing when prompt-path ranking can recover much of the benefit of best-of-N style scaling: if high-utility prompt trajectories are sufficiently rankable and sufficiently covered by the explored pool, a deeper planner-side search can substitute for a substantial amount of width-heavy sampling.
PaperID: 1402, Poster
Abstract: Agentic memory is critical for accumulating and reusing experience in LLM agents across long interaction horizons and related tasks. While latent-memory approaches offer a compact representation of such experience without incurring additional context overhead, existing methods are designed to encode transient in-context states or rely on pre-trained modules that remain fixed at test time to inject experiential representations. Consequently, an effective mechanism for continually refining reusable latent memory from execution experience is still missing in current agents.To this end, we propose , a framework for test-time internalization that enables agents to acquire and update latent memory during interaction. for retrieval and continual refinement across episodes. Through self-supervised updates derived from execution trajectories, progressively encodes task-relevant solving patterns into compact latent memory, enabling both within-episode adaptation and cross-episode transfer. Extensive experiments across eight benchmarks show that consistently outperforms strong baselines, surpassing A-Mem by up to 50.8% and latent-memory baselines such as MemGen and Titans by up to 4.47%--15.96%. Further analysis indicates that the learned latent memory remains stable under continual updates, generalizes well out of distribution, and exhibits interpretable clustering structure across
Abstract: Existing approaches to controllable generation typically rely on fine-tuning, auxiliary networks, or test-time search. We show that flow matching admits a different control interface: adaptation through examples. For deterministic interpolants, the velocity field is solely governed by a conditional endpoint mean; shifting this mean shifts the flow itself. This yields a simple principle for controllable generation: steer a pretrained model by changing the reference set it follows. We instantiate this idea in two forms. Reference-Mean Guidance is training-free: it computes a closed-form endpoint-mean correction from a reference bank and applies it to a frozen FLUX.2-klein (4B) model, enabling control of color, identity, style, and structure while keeping the prompt, seed, and weights fixed. Semi-Parametric Guidance amortizes the same idea through an explicit mean anchor and learned residual refiner, matching unconditional DiT-B/4 quality on AFHQv2 while allowing the reference set to be swapped at inference time. These results point to a broader direction: generative models that adapt through data, not parameter updates.
PaperID: 1404, Poster
Abstract: We consider multi-environment prediction problems. We assume the environments change the distribution of a latent variable, while the mechanisms generating observed covariates and targets remain stable conditional on that variable. For example, hospitals or clinical cohorts may differ in the prevalence of latent patient states, even though the relationships between those states, physiological measurements, and outcomes remain unchanged. Given a dataset from multiple environments, we formulate a Bayesian model for such problems and derive the corresponding variational objective. We show that this objective decomposes into per-environment terms and an additional cross-environment balancing term induced by the model's structure. We use an empirical Bayes method to set the prior and incorporate it into the objective. Based on this objective, we develop an amortized variational algorithm for posterior approximation, and use the resulting learned latent variables to form predictions in new environments. We study our approach through simulations and real-world studies of astronomical source identification, microbiome-based disease detection, and ICU sepsis prediction. Across these settings, our method outperforms previous approaches for prediction in new environments.
Abstract: Language models solve complex problems by articulating intermediate reasoning steps in natural language. While effective, this process is computationally bottlenecked: each reasoning step conveys only a single subword, and many are spent expressing a thought instead of carrying out computation. We propose MUX, a simple method for high-bandwidth and compact reasoning based on distillation of discrete reasoning into continuous multiplexed tokens in a latent space. Each latent token is trained to represent a weighted linear superposition (multiplexing) of a span of discrete reasoning subwords, where this superposition is lossless by construction and the span can be fully recovered (demultiplexing). We prove that simple position-dependent weightings, such as suitable geometric decay, support lossless multiplexing and prevent latent collapse. We further show that multiplexed reasoning can perform parallel exploration in problems that require search. Across 32 evaluation settings spanning four language models, MUX outperforms strong latent reasoning baselines. Our results suggest that properly chosen learning targets can make continuous latent reasoning efficient and interpretable.
PaperID: 1406, Poster
Abstract: Sparse-view novel view synthesis is a central problem in 3D content creation, but diffusion-based approaches remain limited by iterative denoising, making multi-view generation expensive at inference time. We introduce NAMVIS, a diffusion-free framework that reformulates multi-view image synthesis as geometry-conditioned next-scale autoregression. Instead of generating target views through repeated denoising, NAMVIS predicts discrete visual tokens through a small number of coarse-to-fine scale steps, while sampling all tokens within each scale and across target views in parallel. To anchor this generation process to explicit camera geometry, we propose Multi-scale Projective Pose Encoding, which injects source and target camera transformations into both target-view self-attention and source-to-target cross-attention at every resolution. NAMVIS further combines global conditioning with dense geometry-aware cross-attention, enabling the model to preserve source-view appearance while maintaining target-view consistency. Across Objaverse, GSO, and OmniObject3D, NAMVIS outperforms diffusion-based baselines in PSNR, SSIM, and LPIPS, while running over 3× faster than the evaluated diffusion baselines under the same evaluation setting. These results suggest that geometry-conditioned next-scale autoregression is a promising and efficient alternative to diffusion for sparse-view multi-view synthesis.
Abstract: Existing methods for Virtual Try-On (VTON) often struggle to accurately transfer garment appearance, especially in unpaired settings where accurate person-to-garment correspondence is required. These methods do not explicitly supervise person-to-garment alignment, leaving correspondence to be learned implicitly within the generation model. In this paper, we first analyze full self-attention in DiT-based architectures and show that person-to-garment query-key matching aligned with local geometric correspondence is closely related to try-on quality. Building on this insight, we introduce CORrespondence ALignment (CORAL), a DiT-based framework that explicitly aligns query-key matching with robust external correspondences. CORAL integrates two complementary components: a correspondence distillation loss that aligns reliable matches with person-to-garment attention, and an entropy minimization loss that sharpens the attention distribution. For evaluation, we further propose a VLM-based evaluation protocol to better reflect human preference. CORAL consistently improves over the baseline, enhancing both global shape transfer and local detail preservation. Extensive ablations validate our design choices. Code and weights will be publicly available.
PaperID: 1408, Poster
Authors:
Zhaohui Wang, Bo Xu, Yingzhi Tang, He Liu, Chen BaiAbstract: Generative models such as Diffusion Models and Flow Matching have improved multimodal trajectory planning in autonomous driving, but still suffer from inference latency and the lack of tractable probability densities, relying on heuristic trajectory selection. We propose BiDrive, a real-time end-to-end planning framework based on Normalizing Flows (NFs). By leveraging invertibility and exact likelihood estimation, BiDrive directly evaluates trajectory probabilities in a single forward pass, enabling principled maximum-likelihood-based decision making. To address the autoregressive bottleneck of expressive flows, we further develop a bidirectional distillation framework that compresses an autoregressive flow into a one-pass generator, achieving real-time inference (50 FPS) while preserving modeling capacity. The autoregressive model is used only during training as a probabilistic teacher. In addition, we introduce a latent-space optimization mechanism that incorporates differentiable safety and dynamic constraints for efficient test-time trajectory refinement. Experiments on NAVSIM closed-loop benchmarks demonstrate state-of-the-art performance in both safety and efficiency, highlighting the potential of Normalizing Flows for real-time trajectory planning. Code is available at: https://anonymous.4open.science/r/BiDrive.
Authors: Antoni Nowinowski, Krzysztof Krawiec
Abstract: In modern GANs, maintaining an Exponential Moving Average (EMA) of the generator's weights is a standard practice, as such an averaged model consistently outperforms the actively trained generator. However, the EMA generator is used for final deployment only and does not influence the training process. To address this missed opportunity, we introduce Self-Distilled GAN (SD-GAN) that employs the EMA generator as a teacher to guide the active generator (student) via perceptual loss. We prove the local asymptotic stability of SD-GAN in the Dirac-GAN setting and show that it dampens the parasitic cycling behavior that plagues the conventional GANs. Empirical evaluations across established architectures and datasets demonstrate that SD-GAN improves the final image quality on several metrics (FID and random-FID in particular), stabilizes the optimization trajectory and provides additional learning guidance that is not trivially correlated with the conventional adversarial loss. It also proves effective for fine-tuning pretrained GAN models.
Authors:
Han Lin, Xichen Pan, Ziqi Huang, Ji Hou, Jialiang Wang, Weifeng Chen, Zecheng He, Felix Juefei-Xu, Junzhe Sun, Zhipeng Fan, Ali Thabet, Mohit Bansal, Chu WangAbstract: Multimodal learning has advanced visual understanding through powerful multimodal LLMs (MLLMs). In visual generation, however, these models are often used only as global text or context encoders for diffusion generators, limiting their ability to provide structured spatial and temporal guidance. This creates an interface gap: MLLMs can parse complex layouts, attributes, and knowledge-intensive scenes, yet current generation pipelines often struggle to transfer such understanding into images and videos with precise, controllable structure. We propose MetaCanvas, a lightweight framework that organizes MLLM representations into spatially indexed canvas tokens and interfaces them with diffusion generators through patch-wise residual fusion. This provides structured spatial and spatiotemporal conditioning rather than relying only on a single global embedding or 1D query sequence. We implement MetaCanvas on three diffusion backbones and evaluate it across six generation and editing tasks requiring precise layouts, robust attribute binding, and fine-grained multimodal control. Under matched settings, MetaCanvas consistently improves over global-conditioning baselines, narrows the gap to specialized closed-source visual generation and editing systems, and unlocks new editing and in-context capabilities for video diffusion backbones. These results suggest that spatially aligned MLLM-derived canvas tokens provide a promising latent interface for transferring multimodal understanding into diffusion-based generation.
Abstract: This paper introduces the Manifold Probe, a supervised method for discovering representation manifolds in superposition. The method generalizes linear regression probes by learning the space of features of a concept that can be linearly predicted from the representations, and then learning the directions used to encode them. We demonstrate the probe on representations of time and space in Llama 2-7b, finding manifolds which linearly represent an interpretable set of features in each case. In the case of time, we show that by steering along the manifold, we can influence the model's completions about the years in which famous songs, movies and books were released, providing evidence that the Manifold Probe can discover manifolds which are causally involved in the model's behaviour.
Authors:
Kang Wang, Xiao Wang, Bo Li, Renzhe XuAbstract: Population game dynamics describe how aggregate behavior evolves in response to payoff signals. In many applications, the payoff function may change, and the goal is to predict the population trajectory induced by a new payoff design before trajectory data under that design are available. We study this problem as . Standard dynamics learning methods fit the motion observed under training payoffs, but do not separate the payoff-dependent incentive from the payoff-independent response structure needed for prediction under new payoffs. Building on the existing framework of Riemannian game dynamics, we propose PA-RmD, a payoff-aware learning method that keeps payoff information explicit and learns the payoff-independent response structure from data. Theoretically, we show that the learned model preserves fixed-payoff Nash equilibria and derive finite-time one-step prediction bounds under shared-response and payoff-coverage conditions. Empirically, we evaluate PA-RmD on synthetic zero-sum and potential games, as well as a YouTube Trending multi-country dataset, showing that our method performs well under changing payoff designs.
Authors:
Aneesh Muppidi, Firas Darwish, Dylan Cope, João Henriques, Jakob FoersterAbstract: Deliberating takes time. In real-time settings, that time is not free. Standard reinforcement learning (RL) sidesteps this as the environment waits indefinitely for the agent's decision. Instead, we study real-time RL environments where the environment progresses while waiting for the agent's action. Building on prior real-time formalizations, we introduce variable-delay real-time RL, where the agent chooses how long to deliberate at each decision point since the environment progresses. For the planning agents we use, the right delay is state-dependent, and naively planning how long to plan can paralyze the agent. We instead approach this setting by training a lightweight gating policy on top of a planner to select state-dependent planning budgets. Across real-time Pac-Man, Tetris, Snake, Speed Hex, and Speed Go, our gating policy outperforms fixed-budget and heuristic baselines, and transfers to a real-time setup where the environment and agent run on two different GPUs.
Abstract: Regression-based LiDAR relocalization has recently emerged as a promising solution for high-precision positioning in GNSS-denied environments. However, these methods are primarily tailored to autonomous driving, exhibiting significantly degraded accuracy in unmanned aerial vehicle (UAV) scenarios due to arbitrary pose variations and irregular flight paths. In this paper, we propose SOAR, a regression-based LiDAR relocalization framework for UAVs. Specifically, we introduce a locality-preserving sliding window attention module with locally invariant positional encoding to capture discriminative geometric structures robust to viewpoint changes. A coordinate-independent feature initialization module is further designed to eliminate sensitivity to global transformations. Furthermore, most existing UAV datasets are limited to evaluate LiDAR relocalization in real-world, due to the lack of synchronized LiDAR scans, accurate 6-DoF poses, or multiple traversals. Thus, we construct a large-scale UAV LiDAR localization dataset with 4 scenes and 13 irregular paths exhibiting rotation and altitude variations, providing a more realistic benchmark for UAVs. Extensive experiments demonstrate that our method achieves state-of-the-art performance, improving the localization success rate by 40% and reducing mean error over 10m on UAVLoc. Our code and dataset will be released soon, and the dataset demo is available via the anonymous link.
PaperID: 1415, Poster
Abstract: In the fair allocation of indivisible goods, a widely used notion of fairness is envy-freeness up to one good (EF1). A classical way to compute an EF1 allocation is the envy cycle elimination (ECE) algorithm, which iteratively assigns a good to an unenvied agent and, after each assignment, resolves any resulting envy cycle. Although the ECE algorithm always produces an EF1 allocation, it leaves considerable freedom in choosing both the next good to allocate and the agent to receive it. We investigate natural heuristics that exploit this flexibility to improve welfare guarantees. For example, we show that if the heuristic jointly selects the good and the receiving agent maximizing the utility, the worst-case utilitarian welfare loss is significantly lower than that of the vanilla algorithm. By contrast, restricting the heuristic to select only one of these two dimensions does not yield comparable improvements. We also complement our theoretical results with empirical average-case analysis.
Abstract: Flow-matching-style generative models learn time-dependent vector fields, but standard objectives supervise each timestep independently, leaving shared-path temporal structure unused. We introduce (TPC), a simple training-time objective that couples velocity predictions at paired timesteps along the same probability path. TPC leaves the architecture, probability path, sampler, and inference-time computation unchanged. Under matched backbones, training budgets, solvers, and numbers of function evaluations, TPC improves sample quality across flow matching, OT-CFM, rectified flow, DiT-style backbones, and MeanFlow. On CIFAR-10, TPC improves FM from 6.35 to 3.19 FID and OT-CFM from 3.58 to 2.90 FID at the same NFE. Mechanism diagnostics show lower measured gradient variance, higher paired-gradient correlation, and reduced temporal roughness, while random and local pairing controls are substantially weaker. Higher-resolution ImageNet experiments further show consistent gains for DiT-XL/2 at 256 and 512 resolution and for MeanFlow at 256 resolution, all at matched inference cost.
Abstract: Essential visual information can be sparse, yet most visual computation remains dense. Though only a fraction matters, vision transformers process images as uniform sets of tokens, scaling cost with their number. As these models grow stronger, they grow larger, and the cost of computing everything climbs higher. Adaptive computation has shown that we can get by with less, but existing methods struggle at extreme sparsity and are difficult to , training an efficient selector of locations and an expressive extractor of representations to derive meaning from images . We do so by learning the selector and extractor end-to-end with actor-critic reinforcement learning. The selector learns to see, together saving computation by selecting only what is worth processing for a given task. We show that selects the task-specific input, excelling at sparse recognition in high-resolution settings (traffic signs, billiards), and maintaining accuracy with as little as 0.2% of the input. It also generalizes across tasks and models, including global recognition (ImageNet classification), local recognition (ADE20K segmentation), zero-shot classification (by distillation), and regression (class agnostic enumeration). Across all settings, surpasses state-of-the-art selection, providing a general and scalable framework for specialized and efficient adaptive computation.
Abstract: Ensemble sampling is a practically attractive approach to randomized exploration, but existing theoretical guarantees for linear bandits require ensembles much larger than what its practical motivation would suggest. In particular, the sharpest existing analysis achieves the ensemble complexity of \Theta(d\log T), leaving a \log T gap from the intrinsic \Omega(d) ensemble-size barrier. We aim to narrow this gap by proposing an ensemble sampling algorithm that refreshes the ensemble only when the regularized Gram matrix changes substantially. This mechanism localizes the perturbation analysis to epochs with controlled Gram-matrix drift and reduces the sufficient ensemble size to \Theta(d\log d+d\log\log T), while preserving the state-of-the-art \tilde\mathcal O(d^3/2\sqrt T) regret for ensemble sampling with arbitrary bounded arm sets. We further show that, when the arm set is finite of cardinality K, the proposed algorithm achieves the sharper regret bound \tilde\mathcalO(d\sqrtT\log K). To the best of our knowledge, this is the first ensemble-sampling guarantee that simultaneously recovers both canonical regret scalings known for randomized linear bandit algorithms: the \tilde\mathcalO(d^3/2\sqrtT) rate for arbitrary bounded arm sets and the \tilde\mathcalO(d\sqrtT\log K) rate for finite arm sets. The algorithm also admits an anytime implementation without resetting past data, and experiments show that it remains competitive with baselines while using substantially smaller ensembles.
PaperID: 1419, Poster
Abstract: Electroencephalography (EEG) data plays an important role in understanding human brain activity, with various major applications. While current approaches, which largely build on EEG foundation models, have demonstrated strong performance across datasets and tasks, they typically rely on a fixed temporal scale. This design choice limits their ability to capture characteristics of neural dynamic behavior, often spanning across multiple timescales, which are naturally captured in EEG recordings. To address this limitation, we propose MAGE ( ), a foundation model that dynamically adapts to capture and integrate information across different timescales. Specifically, MAGE combines two key mechanisms: (i) a novel multi-timescale attention mechanism, capable of learning representations across multiple temporal scales, and (ii) a trainable aggregation module that dynamically identifies and adapts to the most informative timescale of the input signal. Experimental results across seven datasets demonstrate that MAGE consistently outperforms existing EEG foundation models while using substantially fewer parameters. Moreover, our results support the hypothesis that different downstream tasks benefit from distinct temporal scales. All source code will be released.
PaperID: 1420, Poster
Authors: Hirad Yazdankhah, Chris Metzler
Abstract: Self-supervised image denoisers such as Blind2Unblind excel when clean targets are unavailable but noisy observations are plentiful. We ask: can widely available source-domain data, such as RGB images, improve these denoisers when noisy observations are also limited? We introduce XMNoise2Clean, a framework that learns a noisy target generator from limited paired source/target data, applies it to abundant source images, and trains the denoiser on the union of real and generated noisy samples. We define a cross-modal predictability metric Γ ∈ [0, 1] that measures what fraction of target signal variance the source can explain. When Γ ≈ 1 (no scene properties hidden from the source), generated samples are in-domain and help substantially: using a generator calibrated on 500 paired examples, XMNoise2Clean improves Blind2Unblind by about 3.6 dB with only 50 downstream denoiser-stage real target images, nearly matching the 1000-real denoiser baseline. When Γ is moderate (one hidden scene property, such as sub-frame temporal dynamics in event cameras), the method helps but less reliably. When Γ is low (two or more hidden scene properties, such as the temperature and emissivity fields that determine thermal infrared intensity but are invisible to RGB), generated samples hardly help and can suppress real target structure. Controlled synthetic experiments with known Γ confirm that augmentation benefit generally increases with source-target predictability, and three real-world modality settings plus a controlled proxy experiment support the pattern.
PaperID: 1421, Poster
Abstract: Standard generative models face a memorization--generalization trade-off, that is avoiding memorization is considered necessary for generalization. In supervised learning, however, recent studies show that highly overparametrized models can memorize (interpolate) training data while still generalizing well, a phenomenon known as benign overfitting. Motivated by this, we investigate whether generative models can similarly bypass the trade-off and generalize while interpolating. Generative models should minimize the distance between the true data distribution and the distribution induced by mapping the full latent distribution through the generator. But since the true distribution is inaccessible, existing models instead minimize empirical risk with respect to the training distribution. Under this standard formulation, exact empirical-risk minimization forces the generator to produce only training samples. To address this issue, we consider an alternative empirical risk based on presampled latent variables. Our experiments demonstrate that the resulting presampled-latent version of the flow matching model exhibits benign overfitting on standard image benchmarks, such as MNIST and CIFAR-10. As a theoretical proof of concept, we recast generative modeling as a regression problem and extend existing benign overfitting theory to our setting. Together, these results establish, for the first time, that benign overfitting can occur in generative models.
Abstract: Disentangled representation learning aims to capture the underlying explanatory factors of observed data, enabling a principled understanding of the data-generating process. Recent advances in generative modeling have introduced new paradigms for learning such representations. However, existing diffusion-based methods encourage factor independence via inductive biases, yet frequently lack strong semantic alignment. In this work, we propose a flow-matching–based framework for disentangled representation learning, which casts disentanglement as learning factor-conditioned flows in a compact latent space. To enforce explicit semantic alignment, we introduce a non-overlap (orthogonality) regularizer that suppresses cross-factor interference and reduces information leakage between factors. Extensive experiments across multiple datasets demonstrate consistent improvements over representative baselines, yielding higher disentanglement scores as well as improved controllability and sample fidelity.
PaperID: 1423, Poster
Abstract: Video diffusion models often fail in precisely the dimension that distinguishes videos from images: motion. Standard reconstruction and flow-matching objectives in pixel or VAE-latent space are dominated by appearance, so visually similar clips can be treated as close even when their underlying dynamics differ. We propose Latent Motion Alignment, a training-time objective that compares predicted and ground-truth videos in a learned motion space. This space is learned by training a compact encoder to describe how one latent frame changes into the next, while a decoder ensures that the resulting token captures information needed to model the transition. Once learned, the encoder is frozen and used as an auxiliary loss for fine-tuning video diffusion models. The motion module is used only during training and adds no inference-time cost. We validate that the learned space separates motion rather than appearance shortcuts, improves motion-sensitive generation metrics, and enables fine-grained Olympic-diving generation where jump type and score depend on subtle temporal structure.
PaperID: 1424, Poster
Authors: Xiangbo Mo, Hao Chen
Abstract: Despite advances in representation learning, high-dimensional classification remains challenging in low-sample-size regimes, where the dominant signal may vary across applications and labeled data are often limited. We propose a dissimilarity-profiling classification framework that represents each observation by its class-wise dissimilarity profile, transforming the original feature space into a low-dimensional representation that summarizes how the observation relates to each class. The key idea is to turn a consequence of the curse of dimensionality into signal: high-dimensional geometry can induce systematic within-class and between-class dissimilarity patterns under location, scale, or other distributional changes, and these patterns are captured by the class-wise profiles. Building on this representation, we introduce a rank-transformed algorithm that converts dissimilarities into class-wise rank profiles, yielding a compact representation for classification. The proposed method delivers competitive or improved performance relative to commonly used classifiers on two-class, multi-class, network, and real high-dimensional low-sample-size datasets. To provide insight into the mechanism underlying the method, we analyze a distance-based surrogate and show that the resulting profiles encode differences in first, second, and higher-order moments, while the rank transformation improves robustness to outliers. Together, these results show that rank-transformed dissimilarity profiles provide an adaptive representation for high-dimensional classification when the signal structure is unknown.
Abstract: Understanding why deep neural networks (DNNs) fail to generalize to unseen samples remains a long-standing challenge. Existing studies mainly examine changes in externally observable factors such as data, representations, or outputs, yet offer limited insight into how a model’s internal decision mechanism evolves from training to test. To address this gap, we introduce Decision Pattern Shift (DPS), a new perspective that defines generalization through the stability of internal decision patterns and quantifies failure as their deviation from those learned during training. Specifically, we represent each sample’s decision pattern as a GradCAM-based channel-contribution vector, which captures how feature channels collectively support a prediction, and we propose the DPS metric to measure its discrepancy from the class-average pattern. Empirical analyses across multiple datasets and architectures show that, (i) decision patterns form a highly structured, class-consistent space with strong intra-class cohesion and low inter-class confusion, enabling direct analysis of a model’s decision logic; (ii) the DPS magnitude correlates linearly with the generalization gap (nearly all Pearson r > 0.8), revealing generalization as a systematic drift in the model’s internal decision mechanism; (iii) the DPS spectrum organizes diverse generalization degradation scenarios (covering ideal generalization, in-distribution degradation, domain shift, out-of-distribution, and shortcut learning) into a continuous trajectory, providing a unified explanation of their failure modes. These findings open up new possibilities for early generalization-risk detection, failure-mode diagnosis, and channel-level defect localization.
Abstract: Neurosymbolic learning can use symbolic rules to provide supervision for latent concepts from weak labels, but it commonly assumes that the entities referenced by these rules are already specified. Object-centric models decompose images into slot-like representations; however, such slots are not necessarily aligned with the predicates required for symbolic reasoning. We investigate object-centric neurosymbolic learning under distant supervision, where the object-level arguments of a logic program are learned directly from images using only global task labels. We introduce DeepObjectLog, a probabilistic neurosymbolic model that integrates a slot-based perceptual encoder with a probabilistic logic layer. The encoder predicts objectness and class probabilities for candidate object representations, while the logic layer marginalizes over latent objectness and class assignments to compute the likelihood of the observed label. This formulation provides a differentiable task-level learning signal for object-centric perception without requiring per-object labels, masks, bounding boxes, or heuristic set matching. Evaluations across diverse visual reasoning tasks demonstrate that DeepObjectLog achieves superior out-of-distribution generalization to compositional, object-count, and rule shifts compared to neural object-centric and standard neurosymbolic baselines.
Abstract: Retrieval-augmented and agentic workloads repeatedly prefill recurring predictable structured inputs (which we call 'spans') such as documents and code files. Yet, prefix caching in engines such as vLLM cannot reuse their KV entries unless they share identical prefixes with another request, while Position-Independent Caching (PIC) implementations within production-grade inference servers typically either require substantial server code changes or keep KV state outside the server, incurring host-to-device transfer overhead. We present Minimalistic PIC (MiniPIC): a minimal, flexible and fast vLLM design built from two ingredients: positional-encoding-free KV cache and user-controlled cache-reuse primitives. MiniPIC stores unrotated K vectors in the KV cache, applies RoPE to K tiles inside attention using per-request logical positions, and exposes three user-facing and token-level primitives: block-aligned padding, \ssep, and \pdep, that modify hashing behavior and effective block-level causal attention structure. With fewer than 100 lines of core-engine changes plus a custom attention backend, these primitives are sufficient to realize multiple PIC methods, including Block-Attention, EPIC, and Prompt Cache, within the same running vLLM instance, while natively integrating with KV cache CPU offload implementations. On 2WikiMultihopQA, MiniPIC with interleaved scheduling improves prefill throughput by 49% over baseline vLLM, reduces cached-span time-to-first-token by up to two orders of magnitude, preserves the linear prefill scaling of uncached spans, and incurs only 5.7% worst-case overhead.
PaperID: 1428, Poster
Abstract: Multi-path reasoning methods such as self-consistency (SC) sample K reasoning paths and choose the most frequent answer. However, their gains quickly plateau as K increases, and existing methods do not predict when this saturation will occur. We formalize multi-path LLM reasoning as a diversity combining problem from wireless communications: each path is a noisy channel observation, and pairwise path correlation limits the effective number of independent votes, creating a finite saturation ceiling. Generalized least squares (GLS) analysis shows that, under exchangeability, the optimal symmetric linear combiner of latent embeddings is uniform, supporting majority vote as the natural default in standard SC while leaving room for weighting or pruning under heterogeneous prompt-template branches. Across 5 models and 12 benchmarks, prompt-template diversity reduces path correlation in 55 of 57 valid cells, with the strongest effect on open-ended QA. We derive an Adaptive-K rule that uses a four-path pilot to select K^, retaining 97--104% of MV@K=32 accuracy across Math, QA, and NLU. Our code is hosted at \urlhttps://anonymous.4open.science/r/DiversityCombining-4905.
Authors:
Sebastian Regalado, Varshanth R Rao, Ruowei Jiang, Parham Aarabi, Igor GilitschenskiAbstract: Although advancements in face landmark detection (FLD) methods continue to push performance boundaries, they overlook two major functional limitations: (1) different network parameters need to be trained independently for each "N-point" benchmark dataset, and (2) a model trained on an "N-point" dataset reliably outputs only the N landmarks. In our work, we first conceptualize Face Part-Anchored Landmark Positions (FPALPs), wherein each landmark is treated as a progression value between zero (start) and one (end) along a face part's contour. Every landmark can be expressed in the FPALP format, irrespective of its source dataset, hence unlocking the ability to unify all "N-point" datasets into a single dataset. Secondly, we represent each landmark with an FPALP-based query, refine it progressively with a cross-modality decoder, and predict its coordinates based on the final representation. Our approach, called Unified Dynamic FLD, embodies these two design choices and streamlines the landmark detection pipeline by enabling (1) a single model to learn on any number of "N-point" datasets, and (2) yield any number of specific landmark predictions by loading the designated landmark queries at runtime. Extensive experiments on multiple benchmark datasets show that our method delivers these benefits while remaining competitive with, and in several cases outperforming existing state-of-the-art methods.
Authors:
Xinwei Liu, Junyuan Liang, Zicong Hong, Jianting Zhang, Wuhui ChenAbstract: Augmenting model-free reinforcement learning (RL) with representations learned through observation dynamics prediction (observation-predictive RL) can improve sample efficiency and performance, with minor modifications and limited additional computation. However, this approach still struggles in challenging tasks with low-dimensional observations. In this paper, we identify a key factor behind this problem: unbalanced reconstruction losses across observation dimensions, where dimensions with larger value ranges dominate the loss. This encourages the agent to neglect dimensions with relatively small ranges, leading to degraded performance. To address this issue, we propose a novel normalization method tailored to online RL, which normalizes low-dimensional observations and balances the resulting losses and gradients. Beyond balancing reconstruction losses, observation normalization enables dynamics prediction to be performed in a normalized observation space, thereby providing a unified treatment of low- and high-dimensional inputs (e.g., physical states and images). Building on this idea, we further introduce Normalized Observation Space Dynamics-Augmented Q-learning (NASDAQ), a framework for observation-predictive RL applicable across diverse domains. NASDAQ learns state-action representations by coupling value learning with two auxiliary tasks: short-term value prediction and next observation prediction. Extensive experiments demonstrate that NASDAQ achieves competitive or superior performance compared with state-of-the-art model-based and self-predictive RL methods, while requiring significantly less training wall-time.
Abstract: emerging as a promising paradigm by integrating ads into LLM-generated content. However, classic mechanisms are no longer applicable in this setting where the auction object is shifted from discrete ad slots to distributions over LLM outputs, and existing methods are impractical in industrial scenarios due to ignored externalities or high inference costs. To address these issues, we propose LLM-Auction, the auction and generation. By formulating the allocation as preference alignment between LLM outputs and a mechanism objective that balances advertiser value and user experience, we optimize the LLMs to inherently model allocation externalities without extra inference cost. Theoretically, we identify the allocation monotonicity and continuity of LLM-Auction, and prove that a simple first-price payment rule exhibits favorable incentive properties. Furthermore, we build an LLM-as-a-judge simulation environment for quantitative evaluation, and experiments demonstrate that LLM-Auction achieves the state-of-the-art allocation efficiency while satisfying key mechanism properties.
PaperID: 1432, Poster
Abstract: Realistic physical interaction is a cornerstone of embodied intelligence, yet obtaining high-fidelity tactile data remains significantly more expensive than visual data. This scarcity necessitates Visual-to-Tactile synthesis to bridge the Sim-to-Real gap. However, existing data-driven approaches often treat this as a naive image-to-image translation task, neglecting fundamental contact mechanics. Moreover, the inherent spatial misalignment in real-world visual-tactile datasets frequently forces generative models to average out spatial uncertainties, resulting in blurry and textureless outputs. To address these limitations, we introduce PhysTacGen, a physics-aware generation framework that enforces consistency between visual appearance and tactile mechanics. Our system makes three technical contributions. Firstly, we propose Group Tactile Policy Optimization (GTPO), a reinforcement learning alignment strategy that refines Large Multimodal Models to infer latent physical properties by rewarding physics-consistent reasoning chains. Secondly, we introduce a Semantic-Driven Data Curation and Visual Prior pipeline, leveraging DINOv2 to filter spatially misaligned data and extracting pure depth maps to decouple macro-geometry from surface texture. Finally, we present a physics-driven generative architecture based on Stable Diffusion XL (SDXL). It synthesizes tactile images via a ControlNet conditioned on GTPO-generated physical descriptions and depth-augmented visual priors. Extensive experiments demonstrate that PhysTacGen achieves state-of-the-art perceptual quality and significantly enhances grasp force prediction for zero-shot Sim-to-Real transfer.
Abstract: Many real-world resource allocation systems, such as humanitarian logistics and vaccine distribution, must preposition limited supply across multiple locations before demand is realized while stockouts incur irreversible service losses. To study this, we introduce the Online Shared Supply Allocation (OSSA) problem, a stateful online model in which a central hub allocates a finite, unknown supply to multiple sites facing sequential demand under fixed-charge transportation costs and lost-sales penalties. Unlike classical make-to-stock or make-to-order inventory models, OSSA precludes backlogging and replenishment only hedges against future demand. To tackle OSSA, we propose a deterministic threshold-proportional policy GPA and prove that it achieves a 4/3-approximation to the offline optimum up to an additive term independent of the total supply. We complement this with matching lower bounds showing that the 4/3 ratio is tight and that the additive-error dependence is unavoidable, even for randomized algorithms that know the total supply upfront. Finally, we develop a learning-augmented extension to GPA that principally incorporates imperfect forecasts (e.g., from human experts or ML models) commonly available in practice, enabling us to exploit high-quality advice while being robust against arbitrary bad ones. Synthetic and real-world experiments show that GPA outperforms natural baselines with global supply is scarce.
Abstract: The increasing popularity of reasoning models---language models that output a series of reasoning or thought tokens before producing an answer---is justified, in part, by theoretical results showing that chain-of-thought (CoT) transformers can simulate Turing machines, and thus perform arbitrary computation. However, the Turing machine, while suitable for a complexity theoretic analysis, isn't convenient, intuitive, or efficient when discussing algorithms. Algorithms are typically designed and analyzed at a higher level of abstraction, captured by the Word RAM model with random-access memory and unit-cost operations on O(\log n)-bit words. As a result, Word RAM algorithms can be substantially more efficient than their Turing machine counterparts, raising the question: Can CoT transformers efficiently simulate Word RAM algorithms? For instance, can they sort n items in O(n \log n) steps or run Dijkstra's algorithm in O(E + V \log V) steps? We answer affirmatively, up to poly-logarithmic overhead. We first establish this for finite-precision transformers with poly-logarithmic width and rightmost unique hard attention, then strengthen the result to two more practical settings with finite width and log-precision: continuous CoT, where reasoning takes the form of vectors rather than tokens, and a hybrid architecture in which transformer layers sit atop a recurrent (linear RNN) layer. In all three cases, we find that CoT can efficiently simulate any Word RAM algorithm with only a poly-logarithmic overhead in n. This overhead reduces to log-square when the Word RAM has a ``flat'' instruction set, and only logarithmic for multiplication-free flat instructions---in stark contrast to known CoT simulations of Turing machines, which require quadratic overhead over Word RAM.
PaperID: 1435, Poster
Abstract: Docking ultimately determines effective ligand binding, yet bridging global functional semantics and local physical constraints remains a central challenge for protein generation for ligand-binding design. We therefore use docking-oriented language describing docking processes, binding patterns, and local interaction constraints as generation conditions, rather than relying only on coarse-grained functional text. However, existing methods still insufficiently model docking constraints and show limited coordination among heterogeneous modalities. To address these limitations, we propose DockWizard, an encoder-router-decoder framework for multimodal protein generation. Concretely, (1) a multimodal encoder aligns function text, docking text, ligand descriptions, ligand SMILES, and ligand 3D information into shared conditional representations, improving semantic consistency across heterogeneous inputs; (2) a router module organizes heterogeneous conditions into role-specific memories by their scopes of action, preserving global functional guidance while strengthening local binding constraints; and (3) a protein decoder combines prefix guidance with full-memory injection to let different conditions act at different levels, improving controllability while preserving docking constraints during generation. We further curate a docking-language multimodal dataset with over 100,000 samples. Extensive experiments under a four-stage evaluation pipeline show that DockWizard outperforms strong baselines for this task, with average gains of 13% in docking confidence and 19% in complex stability.
PaperID: 1436, Poster
Abstract: Large multimodal language models (LLMs) have emerged as powerful tools for guiding evolutionary search toward interpretable programmatic policies. However, existing frameworks rely on a monolithic model call to simultaneously interpret visual behavioral evidence and synthesize corrective code. This diagnosis-repair entanglement creates an opaque feedback loop, obscuring the rationale behind mutations and preventing the retention of algorithmic insights across independent runs. To achieve auditable and efficient policy search, we argue that visual diagnosis must be structurally decoupled from code generation. We present REFLEX, a train-free evolutionary framework that operationalizes this decoupling. In REFLEX, a vision-enabled Critic first distills task-specific behavioral evidence into structured, auditable diagnoses. Subsequently, a text-optimized Actor synthesizes child policies using these diagnoses alongside a persistent, self-evolving Skill Memory of reusable code snippets. This architecture not only provides transparent mutation traces but also enables cross-run programmatic knowledge transfer. Extensive evaluations across control benchmarks (Lunar Lander, Acrobot, Pendulum) and a 36-dimensional antenna array synthesis task demonstrate exceptional sample efficiency. Notably, REFLEX solves Acrobot and Pendulum in under 10 LLM calls and reaches a Normalized Weighted Score of 1.013 on Lunar Lander in just 20 calls, achieving highly competitive final performance while significantly accelerating the early-stage discovery of transparent policies.
PaperID: 1437, Poster
Abstract: We propose a simple framework for constructing differentially private confidence regions _in one shot_, i.e., by adding noise only to the final resampling quantile instead of privatizing the estimator computed on each resample. The cost of privacy of our procedure is O(\log B) under with-replacement (m-out-of-n) sampling and independent of B under without replacement sampling (subsampling), avoiding the \sqrtB factor that arises in previous works. We provide _nonasymptotic_ Gaussian Differential Privacy (GDP) and utility guarantees for both subsampling and m-out-of-n resampling, covering mean-like estimators with small global sensitivity as well as estimators admitting efficiently computable _smooth sensitivity_ bounds, including quantiles and degenerate U-statistics. This allows us to also obtain private confidence regions for degenerate U-statistics where the private error is much smaller than the non-private error. In all, we provide a toolbox for widely applicable DP uncertainty quantification procedures under popular resampling strategies while avoiding the computational and privacy costs of privatizing many intermediate resample statistics.
Abstract: Large language models can now be personalised efficiently at scale using parameter efficient fine-tuning methods (PEFTs), but \emphserving user-specific PEFTs harms throughput, even with specialised kernels and memory management techniques. This is because, theoretically and empirically, a mismatch exists between prefill (processing a large number of tokens at once) and decode (generating a single token autoregressively): the latter has far lower throughput when serving multiple adapters. Rather than optimising performance relative to parameter count, for efficient multi-adapter serving, we instead ought to optimise performance relative to serving throughput. We therefore propose \textttPreFT (Prefill-only Finetuning), wherein we only apply the adapter to prefill tokens and discard it afterwards. \textttPreFT significantly increases throughput with minimal effect on performance. We develop and release an efficient implementation of two prefill-only PEFTs, LoRA and ReFT, on the vLLM inference engine. We first show that serving multi-user \textttPreFTs is vastly more efficient than traditional PEFTs (1.90× the throughput when serving 512 adapters on Llama 3.1 70B). Then, we compare the performance of prefill-only vs.~all-token adapters on a variety of supervised finetuning and reinforcement learning tasks with LMs at varying scales. On SFT, we observe that the evaluation loss of \textttPreFTs is higher, but can be compensated by increasing rank with nearly no reduction in throughput. On RL, we consistently find that \textttPreFTs approach parity with standard PEFTs. Together, this work validates prefill-only adaptation of LLMs as a more favourable accuracy--throughput tradeoff than existing PEFTs for personalised serving.
Abstract: Characterizing the features of a Hamiltonian that governs a quantum system serves as a fundamental subroutine of quantum device calibration, signal sensing, and error correction. Recent works proposed protocols have achieved the optimal Heisenberg-limited scaling learning ansatz-free Hamiltonians from their real-time evolutions without fully specifying interaction structures. However, these protocols relies on both deep circuits with interleaving probes and control, and extremely short time resolution, making them difficult to implement on near- and intermediate-term in situ quantum experiments. In this work, we propose a computational efficient, control-free, and ancilla-free algorithm that uses only Pauli product state preparation and measurement, and learns an ansatz-free Hamiltonian H with ||H||\leq\Lambda in total evolution time of \Theta(\tfrac\Lambda\epsilon^2\log(\tfrac\Lambda\epsilon)). The evolution time cost of our algorithm is optimal for any control-free protocols as we further prove a lower bound of \Omega(\tfrac\Lambda\epsilon^2\log(\tfrac\Lambda\epsilon)). Technically, our method introduces a randomized-sampling framework that combines band-limited kernel-based time sampling with a displacement sieve for Hamiltonian structure learning. The characteristic probe time resolution depends only on \Lambda instead of \varepsilon, which makes our protocol especially appealing in the high-precision regime for sensing and calibration applications.
Abstract: This paper studies adaptive targeting under network interference in a bandit setting, where treatments applied to one individual may affect others through spillover effects. We consider a linear model in a sparse regime, where each individual's outcome can be affected by at most a few others. We first establish a regret lower bound showing that ignoring the network structure and reducing the problem to a standard linear bandit inevitably leads to inefficient learning, particularly in large populations. To understand how structural information can be leveraged, we analyze regimes with varying levels of knowledge of the interference structure: (1) full support knowledge, (2) knowledge of the column support sizes, and (3) no prior knowledge. For each regime, we establish regret lower bounds characterizing the fundamental limits of learning, and develop algorithms that achieve near-optimal regret. Together, our results provide a unified view of how knowledge of the interference structure governs the efficiency of online learning under interference, and offer practical adaptive targeting algorithms in each setting. Numerical experiments on synthetic and real-world data demonstrate the practical benefits of our algorithms.
PaperID: 1441, Poster
Abstract: Policy learning from observational data is fundamental to applications ranging from personalized medicine to welfare-program targeting, but standard approaches optimize empirical performance and may fail when the deployment distribution differs from the training distribution. We study \emphdistributionally-robust policy learning: given observational data \\(X_i,W_i,Y_i)\\_i=1^n with binary treatment W, we seek a policy \pi \in \Pi that maximizes the worst-case policy value over an f-divergence ball of radius \rho around the data-generating distribution. We propose an estimator built on cross-fitted doubly-robust scores for the conditional treatment effect, paired with a DRO solver that handles \textttKL, \chi^2, and \textttCVaR ambiguity sets. We establish two convergence guarantees. First, the estimated policy attains the standard \widetildeO(n^-1/2) rate on its DRO regret, and inherits the doubly-robust property from its non-DRO counterpart -- the rate continues to hold whenever either the outcome model or the propensity model is consistently estimated, even at slow nonparametric rates. We then propose a sample-splitting bias-correction algorithm derived from the Lagrangian dual, and show that under a Tsybakov-style margin condition on the worst-case treatment-effect contrast, the estimator attains a fast rate of \widetildeO(n^-(1+\alpha)/(2+\alpha)) , interpolating between \widetildeO(1/\sqrtn) and \widetildeO(1/n) as the margin parameter \alpha ranges over [0,\infty). To our knowledge, this is the first work to incorporate a Tsybakov margin condition into the analysis of distributionally-robust policy learning, and the first to obtain fast rates in this setting. Empirically, on semi-synthetic experiments calibrated to a welfare-to-work program, distributional robustness yields meaningful improvements over empirical welfare maximization when training and test distributions differ structurally, while remaining competitive when no shift is present.
Abstract: Manifold learning is a fundamental task at the core of data analysis and visualisation. It aims to capture the simple underlying structure of complex high-dimensional data by preserving pairwise dissimilarities in low-dimensional embeddings. Traditional methods rely on symmetric Riemannian geometry, thus forcing symmetric dissimilarities and embedding spaces, e.g. Euclidean. However, this discards in practice valuable asymmetric information inherent to the non-uniformity of data samples. We suggest to harness this asymmetry by switching to Finsler geometry, an asymmetric generalisation of Riemannian geometry, and propose a Finsler manifold learning pipeline that constructs asymmetric dissimilarities and embeds in a Finsler space. This greatly broadens the applicability of existing asymmetric embedders beyond traditionally directed data to any data. We also modernise asymmetric embedders by generalising current reference methods to asymmetry, like Finsler t-SNE and Finsler UMAP. On controlled synthetic and large real datasets, we show that our asymmetric pipeline reveals valuable information lost in the traditional pipeline, e.g. density hierarchies, and consistently provides superior quality embeddings than their Euclidean counterparts.
PaperID: 1443, Poster
Authors:
Vincent Andrieu, Pauline Bernard, Lucas Brivadis, Laurent Praly, Christian WolfAbstract: Estimating the state of a dynamical system from observations is a crucial task in robotics and control. Despite their empirical success, recurrent models often lack formal guarantees on convergence and robustness. We bridge this gap by analyzing these architectures from the perspective of nonlinear observer theory. By framing RNNs as a time-discretization of continuous latent dynamical systems (Neural ODEs), we demonstrate that enforcing a contraction condition on the latent flow guarantees the existence of a smooth, injective map of the true physical state into the latent space. We prove that this architecture acts as a tunable observer: by increasing a single scalar gain that scales the latent vector field, the estimation error can be driven to an arbitrarily small neighborhood of zero at an exponentially fast rate. We further establish quantitative bounds on the MLP complexity required to decode this latent state, and extend our formal guarantees to the discrete-time residual RNN implementation. We validate the theory on a real mobile robotic task, where a provably contracting model outperforms a vanilla GRU with substantially more stable training. We also report that unconstrained GRUs converge to contracting dynamics on this task, suggesting that contraction is an emergent property of well-trained sequence models.
Abstract: Structural causal models (SCMs) provide a principled foundation for causal reasoning, and neural causal models (NCMs) extend this framework by parameterizing causal mechanisms with neural networks. The classical formulation of NCMs assumes a specific architectural choice: a single function approximator per structural equation. In this work, we revisit this assumption and show that it represents only one point in a broader family of valid neural causal model parameterizations. Using interventions as an analytical lens and graph-based neural architectures as a concrete construction, we characterize how different parameter-sharing schemes over sub-mechanisms induce distinct members of the NCM family while preserving the same causal graph. This perspective yields two additional regimes beyond the classical formulation, which differ in how causal mechanisms are decomposed and parameterized within a unified structural framework. Together, these regimes demonstrate that architectural choices shape the causal semantics and expressivity of neural causal models, even when graphical structure is fixed. Relaxing structural assumptions further, we identify a class of partially causal models (PCMs) that support limited interventional or counterfactual queries without satisfying full SCM requirements, offering alternative trade-offs between causal fidelity and tractability. We place NCMs and PCMs within a unified expressivity spectrum between purely associational models and fully causal generative models, and illustrate these distinctions empirically using a graph-based variational autoencoder with interventional structure.
Abstract: A standard assumption in deep learning is that the inductive bias introduced by a neural network architecture must persist from training through inference. The architecture you train with is the architecture you deploy. This assumption constrains the community from selecting architectures that may have desirable efficiency or design properties due to difficulties with optimization. We challenge this assumption with Network of Theseus (NoT), a method for progressively converting a trained, or even untrained, guide network architecture part-by-part into an entirely different target network architecture while preserving the performance of the guide network. At each stage, components in the guide network architecture are incrementally replaced with target architecture modules and aligned via representational similarity metrics. This procedure largely preserves the functionality of the guide network even under substantial architectural changes—for example, converting a convolutional network into a multilayer perceptron, or GPT-2 into a recurrent neural network. By decoupling optimization from deployment, NoT expands the space of viable inference-time architectures, opening opportunities for better accuracy–efficiency tradeoffs and enabling more directed exploration of the architectural design space.
PaperID: 1446, Poster
Authors:
Yangzhen Wu, Aaron J Li, Wenjie Ma, Li Cao, Ziheng Zhou, Mert Cemri, Shu Liu, Yuran Xiu, Chenxiao Yan, Haikun Zhao, Bin Yu, Ion Stoica, Dawn SongAbstract: The rapid progress of frontier large language models has led to widespread benchmark saturation, limiting the ability of existing datasets to differentiate model capabilities or provide useful training signal. For instance, on LiveCodeBench, frontier models achieve over (99%) Pass@1 on easy splits and exceed (90%) Pass@1 on average across difficulty levels. Constructing new, sufficiently challenging datasets typically requires substantial human effort, creating a bottleneck for continued progress. We study a solution-centric evolutionary approach that automatically transforms existing programming problems into substantially harder variants. Rather than generating problems from scratch, the approach evolves reference solutions through structured transformations and derives corresponding problem statements and tests from the evolved solutions. This design grounds generation in executable semantics, enabling scalable construction of high-quality, diverse, and difficult tasks with verifiable correctness. Applied to LiveCodeBench and SciCode, it produces evolved tasks that are substantially more difficult while preserving validity, reference correctness, and diversity. Importantly, these tasks remain challenging even for the model that generates them, creating the prerequisite for self-improvement rather than merely expanding an evaluation set. We further show that RL on evolved tasks improves held-out coding performance: for \textttgpt-oss-20b, seed+evolved training achieves (+8.7) and (+8.3) Pass@1 gains on LCB v6 Hard and LCB-Pro Easy, exceeding seed-only gains by (70.7%) and (34.8%), respectively. This closes the loop from self-generated challenges to capability improvement, demonstrating that saturated benchmarks can be converted into both stronger evaluations and reusable training signal.
PaperID: 1447, Poster
Abstract: Standard decoding in diffusion language models (DLMs) is controlled through parallel, token-wise unmasking decisions. However, diffusion decoding is inherently a structured, sequence-level process, in which token states evolve under dynamic cross-token dependencies and are therefore poorly captured by purely token-wise unmasking control. Motivated by this mismatch, we view diffusion decoding from a structural perspective, in which iterative denoising progressively gives rise to token subsets that are internally reliable and weakly dependent on the remaining masked positions. We therefore introduce structural commitment, a sequence-level unmasking principle that treats approximately closed token subsets as the basic unmasking units in diffusion decoding. We instantiate this principle with Structural Commitment via Closure Expansion (SCCE), a structure-aware inference-time algorithm that discovers approximate closures and commits high-certainty subsets by unmasking them, with certainty measured by balancing token reliability against external dependency leakage. Empirically, SCCE improves the accuracy–efficiency frontier over local confidence-based unmasking baselines while adding negligible measured per-step overhead in our implementation.
Abstract: Most LLM unlearning methods aim to approximate retrain-from-scratch behaviors with minimal distribution shift, often via alignment-style objectives defined in the prediction space. While effective at reducing forgotten content generation, such approaches may act as suppression: forgotten concepts can persist in representations and remain entangled with retained knowledge. We introduce CLReg, a contrastive representation regularizer that identifies forget features while pushing them away from retain features, reducing forget--retain interference while empirically preserving the scale and shape of retain features. As light motivation for the mechanism, we provide a one-step analysis showing that CLReg decreases a simple entanglement proxy in the embedding space. Across unlearning benchmarks and LLMs of different sizes, CLReg decreases forget-retain representation entanglement to enhance mainstream unlearning methods without extra privacy risks, inspiring future unlearning work to remove forget concepts via representation shaping.
Abstract: Generative modeling of physical systems, such as molecules, requires learning distributions that are invariant under global symmetries, such as rotations in three-dimensional space. Equivariant diffusion and flow matching models can incorporate such invariances effectively, even when trained on a non-invariant empirical distribution, but they typically rely on costly multi-step sampling. Recently, drifting models have emerged as an efficient alternative, enabling single-step generation and achieving state-of-the-art performance in generative modeling tasks. However, we show that drifting models face a symmetry-specific challenge, since an equivariant generator does not generally produce the same drifting field as the one obtained from the symmetrized target distribution. Addressing this issue would require expensive symmetrization of the empirical distribution. To avoid this cost, we propose SymDrift, a framework that makes the drifting field itself symmetry-aware. We introduce two complementary strategies: (i) a symmetrized drift in coordinate space based on optimal alignment, and (ii) a G-invariant embedding that removes symmetry ambiguity by construction. Empirically, SymDrift outperforms existing one-shot methods on standard benchmarks for conformer and transition state generation, while remaining competitive with significantly more expensive multi-step approaches. By enabling one-shot inference, SymDrift reduces computational overhead by up to 40× compared to existing baselines, making it promising for high-throughput applications such as virtual drug screening and large-scale reaction network exploration.
Authors: Nathan Roll, Jill Kries, Laura Gwilliams, Cory Shain
Abstract: Aphasias, selective language impairments which can arise from brain damage, reveal the functional organization of human language by providing causal links between affected brain regions and specific symptom profiles. Drawing on this literature, we introduce an aphasia-inspired technique to characterize the emergent functional organization of language models (LMs). We ``lesion'' (zero-out) model parameters and measure the effects of this intervention against clinical aphasia symptoms, as diagnosed by the Text Aphasia Battery (TAB). When applied to 112,426 outputs from five 1B-scale LMs, the full range of evaluated symptoms surface, but in distributions largely distinct from those of humans. Our method uncovers broad, component-level symptom differences between attention (query, key, value, output) and feed-forward (up, gate, down) projections, without any significant within-class differences. We also find an effect of depth, where lesions in early layers disproportionately cause syntactic and semantic symptoms while late-middle layers yield higher rates of phonological and fluency deficits. Although some LM lesions induce quantitatively more similar profiles to some human aphasia types than others, qualitative differences in symptom patterns between LMs and humans suggest that aphasia syndromes are heavily influenced by the details of learning and processing rather than being a domain-invariant consequence of disrupted language processing.
PaperID: 1451, Poster
Abstract: The alignment of Large Language Models (LLMs) via reinforcement learning from human feedback (RLHF) typically relies on Proximal Policy Optimization (PPO), which uses a uniform clipping hyperparameter \epsilon to maintain a stable trust region. However, we demonstrate that this uniform constraint is ill-suited for the non-uniform semantic topology of natural language. In ``semantically dense'' regions, where candidate tokens act as semantic synonyms within the current context, noisy advantage estimates can destabilize the fixed trust region, leading to policy collapse, a severe loss of generative diversity and policy entropy. To resolve this, we propose Structure-Aware Mirror Proximal Policy Optimization (SAMPPO), which introduces an adaptive trust region governed by local semantic density. SAMPPO generalizes the trust-region framework to non-uniform semantic spaces by adapting the mirror-descent geometry to the local action topology. SAMPPO uses a per-timestep Semantic Isolation Score to dynamically tighten the clipping range in dense regions and relax it in sparse ones. We provide theoretical foundations for this approach, deriving a Structure-Aware Monotonic Improvement bound and establishing a formal link to structure-aware margin-shifted H-consistency theory. Empirically, SAMPPO prevents the entropy collapse observed in standard PPO baselines and significantly improves alignment. Evaluated on the Anthropic HH-RLHF and NVIDIA HelpSteer benchmarks across multiple model architectures (Llama-3-8B and Gemma-2-2B), SAMPPO achieves robust, zero-shot generalizable alignment gains. Crucially, it strictly outperforms standard PPO and recent empirical adaptive heuristics (BAPO), demonstrating competitive performance against direct preference baselines including DPO and SimPO on the alignment-diversity Pareto frontier.
Abstract: Reinforcement learning and contextual bandit algorithms have become increasingly common in sequential decision-making applications. When these methods are deployed in high-stakes domains, there is growing interest not only in learning effective policies, but also in conducting statistical inference for quantities learned under adaptive data collection. However, classical procedures applied naively in these settings can fail: even when estimators are unbiased, their variance becomes path-dependent and as a result may not be asymptotically normal. A growing literature has emerged to ameliorate this problem, but solutions tend to be problem specific and often rely on correct specification of a working model. In this work, we develop a unified framework for constructing asymptotically valid confidence intervals to cover smooth functionals of nonparametric M-estimands under adaptive sampling. Under Neyman orthogonality, we provide two novel methods for performing inference: (1) a self-normalized statistic based on the realized quadratic variation of the influence function and (2) a statistic using a plug-in estimate of the conditional variance based on reweighted influence function increments. Our results allow for flexible nonparametric estimation of nuisance parameters and remain valid under model misspecification. Our theory is supported by a simulation study for a dynamic pricing application which demonstrates that this method can produce asymptotically valid confidence intervals where standard methods fail.
Abstract: The cost of Kohn-Sham density functional theory (KS-DFT) calculations scales with the number of solver iterations, which depends on the quality of the initial guess. Machine learning methods that predict initial guesses from molecular geometry can reduce this cost, but matrix-prediction models fail when extrapolating to larger molecules, degrading rather than accelerating convergence [Liu et al., 2025]. We show that this failure is a supervision problem, not an extrapolation problem: models trained on ground-state targets fit those targets well out of distribution, yet produce initial guesses that slow convergence. Solver-Aligned Initialization Learning (SAIL) resolves this for both Hamiltonian and density matrix models by differentiating through the self-consistent field (SCF) solver end-to-end. We introduce the Effective Relative Iteration Count (ERIC), a correction to the commonly used RIC that accounts for hidden Fock-build overhead. On QM40, which contains molecules up to 4× larger than the training distribution, SAIL reduces ERIC by 37% (PBE), 33% (SCAN), and 28% (B3LYP), more than doubling the previous state-of-the-art reduction on B3LYP. On QMugs molecules 10× larger than the training set, SAIL delivers a 1.35× wall-time speedup at the hybrid level of theory, extending ML SCF acceleration to large drug-like molecules.
PaperID: 1454, Poster
Authors: Joshua S Schiffman
Abstract: Training selects for behavior, not circuitry: many weight configurations can implement the same function. Studying any single trained neural network thus risks describing accidents of one training run rather than the computation itself. This work shifts focus from what transformers happen to do to what they must do by extracting algorithmic cores, compact subspaces that are necessary and sufficient for a task and that recur across independently trained models. Here, Algorithmic Core Extraction (ACE) is introduced to isolate these subspaces, causally validate them, and recover the algorithms they implement across settings ranging from synthetic tasks to large-scale pretrained models. Markov-chain transformers embed three-dimensional cores in nearly orthogonal subspaces yet recover identical transition spectra. Modular-addition transformers form compact cyclic cores at grokking that later inflate under continued regularization, redundantly distributing the same computation across many functionally equivalent modes. This functional redundancy is found to accelerate the transition from memorization to generalization, yielding an inverse scaling law for grokking time. In six language models spanning two orders of magnitude in scale (GPT-2 Small/Medium/Large, LLaMA-3.1, Gemma-2, and Qwen2.5), subject-verb agreement is governed by a single, steerable axis that aligns across architectures. Flipping this axis inverts grammatical number throughout open-ended generation. Together these results suggest that beneath the apparent complexity of trained transformers lies a simpler, shared computational structure, and that targeting invariants rather than parameterizations may offer a more tractable path to mechanistic understanding and control.
Abstract: Score-debiased kernel density estimation (SD-KDE) achieves improved asymptotic convergence rates over classical KDE, but its use of an empirical score has made it significantly slower in practice. We show that by re-ordering the SD-KDE computation to expose matrix-multiplication structure, Tensor Cores can be used to accelerate the GPU implementation. On a 32k-sample 16-dimensional problem, our approach runs up to 47× faster than a strong SD-KDE GPU baseline and 3,300× faster than scikit-learn's KDE. On a larger 1m-sample 16-dimensional task evaluated on 131k queries, Flash-SD-KDE completes in 2.3 s on a single GPU, making score-debiased density estimation practical at previously infeasible scales.
Abstract: Recent work studied the problem of finding clusters and denoising pairwise distances from points sampled on a manifold. We study the same problems in more general metric measure spaces and provide efficient algorithm to partition the points to clusters of a fixed radii and denoise distances to any fixed accuracy. We also show how to achieve much higher accuracy with a non-efficient algorithm. This suggests that unlike the Riemannian case, denoising to higher accuracy in more general metric spaces has a statistical-computational gap.
Abstract: Amos et al. (2024) showed that the accuracy of Transformer models in sequence classification can be significantly improved by first pretraining with a masked token prediction objective without external data or augmentation, a procedure referred to as self-pretraining (SPT). While the primary objective of Amos et al. (2024) was to showcase that Transformers can achieve strong performance on the Long-Range Arena (LRA), their pipeline raises more fundamental questions: How does SPT drive optimization to better solutions? Why can standard supervised training fail in Transformers? To better understand this, we replicate and systematically ablate the findings of Amos et al.~(2024). Our ablations suggest that a central bottleneck in the studied settings is not depth or generalization alone, but the ability of label supervision to learn useful query-key Attention patterns from random initialization. With a minimal setup, we identify learning proximity interactions -- turning absolute positional encodings into proximity-biased Attention scores -- as a key source of the improvements brought by SPT. Finally, in a simplified theoretical setup, we show that label supervision can be locally blind to certain Attention-score directions that are instead detectable through masked reconstruction.
Authors: Sarvesh Gharat, Nikhil Karamchandani, Jayakrishnan Nair
Abstract: Inspired by the problem of identifying the best model from a collection of large language models (LLMs) with heterogeneous querying costs, we formulate and analyse a variant of the multi-armed bandit (MAB) with (i) dueling feedback, where pairwise comparisons between model responses provide robust preference signals, and (ii) heterogeneous sampling costs, reflecting the differing costs of querying different LLMs. Assuming the existence of a Condorcet winner, a condition we empirically validate across multiple real-world datasets, we propose a Track-and-Stop style algorithm for best-arm identification with prescribed confidence. We prove that the algorithm almost surely achieves the asymptotically optimal cost as the error tends to zero. Finally, we extensively evaluate our approach on both synthetic and real-world instances, demonstrating consistent improvements over classical cost-unaware algorithms and their cost-aware extensions.
Abstract: Post-training is central to building scientific reasoning models, but its effects remain poorly understood in biology, where models integrate natural language with multimodal biological data. We study how post-training stages shape generalization, and when they improve performance versus induce over-specialization. Across genomics, transcriptomics, and proteomics, we train and evaluate more than 100 biological reasoning models under controlled variation in backbone, continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL), measuring both in-domain (ID) and out-of-domain (OOD) performance. We find that post-training induces stage-specific trade-offs rather than uniform gains. CPT improves downstream performance by aligning models with biological language. SFT consistently increases ID performance but causes OOD performance to peak early and decline as models fit the training distribution. RL, when applied to strong SFT checkpoints with aligned rewards, improves OOD performance and partially recovers generalization. These results show that scientific reasoning does not improve monotonically with additional supervision or compute. Instead, performance depends on how training stages are composed. Under fixed epoch-level budget, the best ID-OOD trade-off comes from combining minimal SFT with larger RL budgets and allocating adaptation capacity asymmetrically across stages.
PaperID: 1460, Poster
Abstract: Universal Transformers (UTs) show strong performance on abstract reasoning tasks such as ARC-AGI, yet the specific sources of their performance gains remain underexplored. We conduct systematic ablations over UT variants and find that performance improvements on abstract reasoning are driven primarily by (i) the recurrent inductive bias induced by parameter sharing across depth and (ii) strong nonlinearity inside the Transformer block, rather than by elaborate architectural heuristics. Motivated by this finding, we propose the Universal Reasoning Model (URM). First, we propose ConvSwiGLU, which augments the feed-forward block with a channel-wise short convolution to effectively enhance nonlinearity of UT. Second, we adopt Truncated Backpropagation Through Loops (TBPTL) to eliminate noise from early loop when training with many loops, by freezing the gradients of early loops. Our approach substantially improves abstract reasoning performance, achieving 57.5% pass@1 on ARC-AGI 1 and 16.0% pass@1 on ARC-AGI 2. More importantly, we pretrain a 8-billion-parameter large language model based on the URM architecture from scratch, and the results show that URM still substantially outperforms the baseline across all reasoning tasks.
Abstract: Reinforcement Learning has significantly advanced the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet the resulting policies remain brittle against real-world visual degradations such as blur, compression artifacts, and low-resolution scans. Prior robustness techniques from vision and deep RL rely on static data augmentation or value-based regularization, neither of which transfers cleanly to critic-free RL fine-tuning of autoregressive MLLMs. Reinforcing reasoning against such corruptions is non-trivial: naively injecting degraded views during rollout induces reward poisoning, where perceptual occlusions trigger hallucinated trajectories and destabilize optimization. We propose ROMA, an RL fine-tuning framework that modifies the optimization dynamics to reinforce reasoning against visual degradation while preserving clean-input performance. A dual-forward-pass strategy uses teacher forcing to evaluate corrupted views against clean-image trajectories, avoiding new rollouts on degraded inputs. For distributional consistency, we apply a token-level surrogate KL penalty against the worst-case augmentation; to prevent policy collapse under regularization, an auxiliary policy gradient loss anchored to clean-image advantages preserves a reliable reward signal; and to avoid systematically incorrect invariance, correctness-conditioned regularization restricts enforcement to successful trajectories. On Qwen3-VL 4B/8B across seven multimodal reasoning benchmarks, our method improves robustness by +2.4% on seen and +2.3% on unseen corruptions over GRPO while matching clean accuracy.
Abstract: Despite advances in test-time scaling and diffusion finetuning, guidance for Auto-Regressive Diffusion Models (ARDMs) remains underexplored. We introduce an amortized framework that augments a pretrained ARDM with an offline-trained controller. By previewing future rollouts, the controller learns stepwise corrections that anticipate observations under a terminal-cost objective, yielding a reusable policy for guided generation. Motivated by a stochastic optimal control view of ARDM trajectories, our method injects small controls within each denoising sub-step while staying close to the pretrained dynamics. We study this approach for data assimilation (DA) in chaotic spatiotemporal partial differential equations (PDEs), where existing methods are often computationally expensive and susceptible to forecast drift under sparse observations. At inference, DA becomes a feed-forward rollout with on-the-fly corrections, achieving an order-of-magnitude speedup over strong diffusion-based baselines. Across two canonical PDEs and a compact ECMWF Reanalysis v5 (ERA5) pilot spanning six observation regimes, our method consistently improves stability and accuracy over state-of-the-art alternatives, with similar improvements observed in a larger-scale GenCast study.
Abstract: We characterize the fundamental limits of high-dimensional mean testing under arbitrary truncation, where samples are drawn from the conditional distribution P(\cdot \mid S) for an unknown truncation set S that may hide up to an \varepsilon-fraction of the probability mass. For distributions with p-th directional moments of magnitude at most \nu_P,p, truncation induces a bias of order \mathcalO(\nu_P,p\varepsilon^1-1/p). This bias creates a sharp information-theoretic detectability floor: when the signal \alpha falls below this threshold, the null and alternative hypotheses are indistinguishable even with infinite data. Above this floor, we prove that a simple second-order test achieving near-optimal sample complexity n = \mathcalO\left(\frac\|\Sigma_P\|(\alpha-4\nu_P,p\varepsilon^1-1/p)^2\sqrtd\right). We further identify a structural escape from this finite-moment bias barrier. Under a directional median regularity assumption, truncation bias improves to linear order \mathcalO(\varepsilon). This reveals an intermediate regime in which estimation requires \Theta(d) samples for uniform recovery, while testing recovers the classical \Theta(\sqrt d) rate once truncation bias is eliminated. Together, our results provide a unified framework for mean testing under truncation, connecting finite-moment, sub-Gaussian, and median-regular structural regimes.
Abstract: Pretraining followed by fine-tuning introduces a compute-allocation problem: under a fixed training budget, time spent improving the upstream objective reduces the time available for downstream adaptation. Despite its practical importance, this trade-off is not yet well understood theoretically, even in simple models. In this paper, we cast compute allocation as a time-split problem under a two-stage pretrain–fine-tune procedure with fixed total optimisation time, using regularised least squares trained by gradient flow as a tractable setting. We characterise the optimal switching time under data-dependent evaluation geometries induced by the fine-tuning problem. Our results show that the allocation depends on how pretraining directions affect fine-tuning predictions and how fine-tuning shifts are seen through downstream data geometry. In particular, the relevant quantities are determined by prediction-relevant spectral components of the pretraining and fine-tuning empirical covariances. Technically, the analysis relies on a basis-invariant, eigenspace-level spectral decomposition, together with perturbative control of the non-commuting pretraining and fine-tuning dynamics.
Abstract: Reinforcement Learning (RL) algorithms often rely on data collected from a target environment, which can be expensive and noisy. Although simulators can generate large amounts of additional data, standard simulation mixing approaches—which directly combine simulated and real experience—can degrade performance due to simulator inaccuracies, especially when they bias the solution away from the optimal policy. We propose a conservative simulation-aided learning method based on control variates, where simulated data is used solely to reduce variance in the learning process rather than directly shaping the learning target. We develop our approach for Policy-Evaluation, which is a fundamental building block of RL algorithms. Then we combine it within the actor-critic framework towards policy improvement, as well as extend it towards offline RL scenarios. Experiments on multiple MuJoCo benchmark tasks demonstrate improvements over standard simulation mixing approaches, highlighting control variates as an effective tool for simulation-aided reinforcement learning.
PaperID: 1466, Poster
Abstract: High-fidelity 3D generative modeling underpins game asset creation, medical imaging, and VR/AR environmental design, where geometric structure directly influences perception and decision-making. Diffusion models represent a powerful paradigm for point cloud generation, yet existing approaches rely heavily on local voxel or patch features and often miss global structural organization. Recent topology-aware methods show that global conditioning can improve generation quality, yet persistent-homology pipelines add substantial preprocessing complexity. We introduce HilbertGen-3D, a topology-aware conditional diffusion framework based on Hilbert–Multifractal Conditioning (HMC). HMC computes multi-scale voxel measures, serializes them via a 3D Hilbert curve to expose structured spatial dependencies, and estimates the generalized fractal dimensions from scale-dependent partition functions. We encode these signals as global embeddings and multi-scale tokens, and inject them into a diffusion denoiser through an efficient bottleneck fusion module. On ShapeNet dataset, HilbertGen-3D reduces the 1-Nearest Neighbor Accuracy by up to 8.74 points and improves coverage by up to +21.45 over prior topology-aware methods. This shows that combining local density cues with global Hilbert–multifractal structure provides a simple, efficient, and effective recipe for high-fidelity and diverse 3D generation.
Abstract: We introduce Quokka, the first comprehensive scaling law for diffusion language models (DLMs), encompassing both compute- and data-constrained regimes alongside key modeling and optimization designs. Under compute constraints, we find that optimal parameter and dataset sizes scale equally with compute (N_\mathrmopt \propto C^0.5, D_\mathrmopt \propto C^0.5); however, DLMs are 2--5× more data-hungry than autoregressive (AR) models, requiring larger corpora and proportionally smaller models for a given FLOP budget. Under data constraints, validation loss follows a U-shaped curve across epochs, with the onset of overfitting scaling as e_\mathrmopt \propto U_D^0.39/N^0.55. This indicates that optimally leveraging a fixed unique data budget (U_D) requires allocating both modestly larger models and more training epochs. Beyond data and parameter allocation, we provide actionable guidance on DLM design choices: absorbing-mask transitions outperform uniform diffusion, linear noise schedules are the most stable and performant, and principled diffusion ELBO objectives ultimately surpass MaskGIT losses despite slower initial convergence. Furthermore, we demonstrate that easy-to-hard noise curricula accelerate early learning, established AR scaling laws for batch size and learning rate transfer seamlessly to DLMs, and weight decay, while unhelpful for single-epoch runs, is essential for multi-epoch regularization.
PaperID: 1468, Poster
Abstract: Many problems in diverse fields involve data that naturally live on Lie groups. However, existing Riemannian and Lie group generative models still face several limitations. Riemannian flow and consistency models often require position-dependent vector fields, and expensive geometric terms such as Euler-Arnold simulation on general Lie groups or covariant derivatives. Existing Lie group flow models are usually tied to collinear exponential paths, while momentum-based Lie group diffusion relies on compactness assumptions and does not naturally extend to non-compact groups. To address these limitations, we propose Trivialized Generative Models (TGM), a family of generative models that learns endpoint-constrained paths in the fixed Lie algebra and lifts them to the Lie group. This yields Trivialized Flow Matching (TFM), Trivialized Consistency Models (TCM), and momentum-based extensions with endpoint correction, enabling flexible path design, simpler few-step objectives, and generation on non-compact Lie groups. We evaluate TGM on compact SO(3) benchmarks, non-compact \mathrmSp(4,\mathbbR) datasets motivated by continuous quantum systems and paraxial optics, and real-world protein backbone generation. Across these settings, TGM improves sample quality and efficiency over existing Riemannian and Lie group baselines.
PaperID: 1469, Poster
Authors:
Chenglong Ma, Yuanfeng Ji, Junzhi Ning, Jiyao Liu, Yujie Wei, Shaohao Rui, Chaoyang Zhang, Wenjie Li, Wenhao Tang, Cheng Tang, Wei Li, Shujian Gao, Yujie Zhang, Ying Chen, Zilong Li, Zongyuan Ge, Diping Song, Jin Ye, Ming Hu, Junjun He, Hongming ShanAbstract: Multimodal medical AI requires a common visual interface that integrates heterogeneous imaging evidence for patient condition modeling. However, existing medical visual tokenizers are often tied to a specific dimension and imaging modality, forcing multimodal systems to rely on fragmented and poorly reusable representations. We introduce MedVTok, a general-purpose visual tokenizer that maps 2D images and 3D volumes from diverse imaging modalities into a unified token space. Building such a tokenizer poses two key challenges: preserving anatomical structure across dimensions and capturing clinical semantics across modalities. To this end, MedVTok introduces (i) slice differential regularization, which explicitly models inter-slice anatomical coherence missing in previous per-pixel/per-voxel supervision, and (ii) multi-expert representation alignment, which integrates knowledge from multiple medical expert encoders while retaining modality-specific complementary knowledge. Trained on 28M images and 70K volumes across 9 imaging modalities, MedVTok supports classification, retrieval, segmentation, synthesis, visual question answering, and medical report generation through a single tokenizer. Across 24 benchmarks, MedVTok achieves state-of-the-art performance, demonstrating a scalable token interface that unifies medical image dimension, modality, and task usage. Code and weights will be released.
PaperID: 1470, Poster
Abstract: Transcriptome-conditioned whole-slide image (WSI) synthesis aims to decode the complex mapping from molecular states to histologic phenotypes. Existing generative paradigms predominantly rely on a direct black-box transcriptome-to-image mapping, tasking a single generator with the simultaneous interpretation of transcriptome signals, inference of multicellular spatial organization, and rendering of fine-grained textures. This formulation, however, is structurally under-constrained; the underlying tissue architecture, the critical substrate linking gene expression to morphology, remains latent and only weakly supervised by pixel-level objectives. We present \textscTopoScape, a topology-anchored framework for transcriptome-to-histology generation under a \emphLayout-Before-Pixels principle. Instead of synthesizing pixels directly from transcriptomic embeddings, \textscTopoScape first resolves molecular states into explicit, multi-class cellular topologies via an RNA-guided Topological Prior Generator, reinforced by persistence-based topological regularization. These resolved topologies are further reparameterized into a suite of topology-derived controls, including structure-aware initialization, continuous density fields, and distance representations, which guide a hierarchical, frequency-decoupled flow-matching trajectory. By introducing multicellular organization as an explicit structural mediator, \textscTopoScape allows molecular states to deterministically shape both the macroscopic tissue architecture and the microscopic generative process. Across five TCGA benchmarks, \textscTopoScape consistently outperforms state-of-the-art methods in generative fidelity, biological realism, and downstream predictive utility. Code and pretrained models will be released.
PaperID: 1471, Poster
Abstract: Multi-task learning (MTL) aims to jointly optimize multiple related tasks to obtain shared representations within a single model. Recent research has focused on developing MTL algorithms to alter optimization dynamics through task re-weighting or gradient manipulation. However, in this paper, we first empirically identify that current MTL approaches are highly sensitive to the model initialization, complicating empirical evaluation and limiting practical reliability, which has been overlooked but appears to be critical. To better understand this phenomenon, we provide a local theoretical analysis that motivates the connection between early-stage MTL optimization behavior and curvature around initialization. Motivated by this analysis, we propose a Curvature-guided Parameter Initialization (CPI) approach tailored for MTL. Specifically, CPI performs a short warm-up phase to probe local curvature around the initial parameters and heuristically biases the initialization toward empirically favorable curvature profiles for subsequent multi-task optimization. The proposed approach is optimizer-agnostic, requires no modification to existing MTL algorithms, and can be seamlessly integrated as a plug-in prior to standard training. Extensive experiments on standard MTL benchmarks show that CPI improves the aggregate multi-task trade-off while remaining compatible with mainstream MTL methods.
Abstract: Backpropagation (BP) is widely viewed as biologically implausible, in part because it requires feedback weights to be the transpose of forward weights for error propagation. Interestingly, when training a network with fixed random feedback weights to circumvent this issue, learning aligns the forward weights with the feedback weights, leading the backpropagated error signal to become an approximation of the standard gradient used by BP. This process, called Feedback Alignment (FA), occurs in MLPs and very shallow CNNs but does not scale well to deeper architectures. In this work, we first investigated differences between BP and FA models, trained on CIFAR10, specifically focusing on the effective rank of the signal. We found that the FA error has a considerably lower rank and hence is constrained to a lower-dimensional subspace compared to BP, limiting exploration of the parameter space. Motivated by this observation, we evaluated two mechanisms for increasing the effective dimensionality of FA: Muon, an optimiser that orthogonalises weight updates; and hidden activity normalisation, which promotes activation orthogonality. Across larger architectures and benchmarks, we find that these methods consistently improve over FA baselines, for example, on CIFAR100 with a Resnet-18, accuracy increases by 9 percentage points. Our results identify low-dimensional gradient dynamics as a key obstacle to scaling FA and suggest that inducing higher-dimensional update geometry is a promising route toward scaling alternatives to backpropagation.
Abstract: We study the query complexity of min-max optimization of a nonconvex-nonconcave function f over the hypercube [0,1]^d × [0,1]^d. We show that, given oracle access to f and to its gradient \nabla f, any algorithm that finds an \varepsilon-approximate stationary point must make a number of queries that is exponential in 1/\varepsilon or d.
PaperID: 1474, Poster
Authors: Renjie Zou, Zhiwei Huang, Dong Jiang
Abstract: Learned image compression (LIC) is trained almost exclusively on clean images: the only stochastic perturbation in the loop is an additive uniform proxy injected in the latent domain to keep the entropy model differentiable. Outside compression, by contrast, a long line of work - from Tikhonov regularization and contractive autoencoders through denoising score matching and modern diffusion models - has established that Gaussian noise at the input of a network with a clean target acts as a structured probe of the data manifold rather than as data augmentation. Building on this view, we introduce noise-regularized training for LIC: at every training step we add Gaussian noise to the encoder input while keeping the distortion target, the latent uniform proxy, the hyperprior, and the entropy coding pipeline unchanged. The recipe is architecture-agnostic and adds essentially no training or inference overhead. Across backbones and operating regimes it yields consistent BD-Rate gains: -2.60% on a reproduced ELIC and -0.85% / -1.05% / -2.38% on the Base/Medium/Large scales of a variable-rate DMCI codec, with DMCI-Large reaching -22.02% against VTM-22.0. A second-order expansion of the noisy RD loss accounts for these gains: because the noise enters upstream of the analysis transform, it converts the clean RD objective into a structured O(\sigma^2) regularizer comprising an information-weighted encoder-Jacobian penalty (rate side) and a Frobenius contractive penalty on the end-to-end reconstruction map (distortion side). A finite-difference Lipschitz analysis verifies both predictions: across QPs and an order-of-magnitude sweep of the probe scale, the encoder's local Lipschitz constant drops by ~12% on average and that of the end-to-end codec by ~8%. Code and trained models will be released to facilitate reproducibility.
PaperID: 1475, Poster
Authors: Yuki Uehara, Naoki Nishimura, Noriyoshi Sukegawa, Yuichi Takano
Abstract: On two-sided platforms such as e-commerce marketplaces, recommender systems deliver personalized item rankings to consumers and serve as a primary channel through which transactions occur between item providers and consumers. At the same time, promoting the activities of item providers requires the platform to balance consumer satisfaction with fair exposure across items. Recent work formulates this objective as impact-based fair ranking, which maximizes a Nash social welfare criterion over stochastic ranking policies. However, the resulting policies must be served as user-specific mixtures of deterministic rankings, whose support size can grow combinatorially. This imposes substantial storage and serving overhead, driving up operational cost in production deployment. Moreover, unseen users must fall back to a generic ranking rule. We therefore develop compact representations for serving impact-based fair-ranking policies. We first prove a tight bound that every optimum admits an impact-preserving mixture with support size at most the number of items. Using a dual reformulation, we then replace stored user-specific ranking mixtures with item-indexed dual vectors and compress the resulting sequence of dual vectors by weighted random subsampling. The resulting deployment representation is independent of the training-user population and can be used for unseen users. Experiments on real-world recommendation datasets show that the sampled dual mixture preserves the fair-ranking objective while reducing the storage required at deployment by up to 2,596× relative to a user-coupled mixture baseline.
Authors: Daniel Ablin, Alon Peled-Cohen
Abstract: Optimization under uncertainty is a fundamental problem in learning and decision-making, particularly in multi-agent systems. Previously, Feldman, Kalai, and Tennenholtz [2010] demonstrated the ability to efficiently compete in repeated symmetric two-player matrix games without observing payoffs, as long as the opponent’s actions are observed. Extending this capability to the Markovian setting remains an open problem. In this paper, we introduce and formalize a new class of zero-sum symmetric Markov games, which extends the notion of symmetry from matrix games to the Markovian setting. We prove that a learner observing only the opponent's action sequence, without access to payoff information can successfully compete against an adversary possessing complete knowledge of the game. We formalize three distinct notions of symmetry in this domain and reveal a surprising structural hierarchy: the most holistic definitions of symmetry impose restrictive constraints that actually simplify the learning landscape, whereas the ``simplest'' definition represents the most general and challenging setting. We provide polynomial-time algorithms for all three settings that achieve sublinear regret. Crucially, we demonstrate that despite the complex Markovian dynamics, a simple strategy of locally mimicking the opponent's actions suffices to guarantee robustness. This finding significantly broadens the class of games where robust learning is possible under severe informational disadvantage, proving that knowledge of the transition laws is not required to force a draw.
PaperID: 1477, Poster
Abstract: We study online learning in the adversarial injection model [GHMS23], where each instance is either drawn i.i.d. from an unknown distribution or is an arbitrary adversarial injection. The learner never knows which rounds are adversarial, but may abstain without penalty on adversarial injections. The learner's goal is to make few total misclassifications and few abstentions on i.i.d. instances. This model lies between classical statistical learning, governed by the VC dimension, and adversarial online learning, governed by the Littlestone dimension. Unlike prior work, which has largely focused on the realizable setting, we consider an agnostic formulation that allows arbitrary labels. We present an online algorithm that guarantees sublinear excess misclassification error and sublinear abstention error, against oblivious and adaptive adversaries, for function classes that admit bounded bracketing number witnessed by brackets of bounded VC dimension. Our algorithm is distribution-independent and does require prior knowledge of the distribution. As corollaries, we obtain new learning guarantees for halfspaces (linear classifiers) and richer classes such as feed-forward neural networks with threshold activations.
PaperID: 1478, Poster
Abstract: Formal verification can play a key role in ensuring the reliability of Deep Neural Networks (DNNs) deployed in safety-critical systems. Modern DNN verifiers employ a branch-and-bound framework, which alternates between branching (splitting into smaller subproblems) and bounding (pruning subproblems) to efficiently explore the verification space. However, existing branching heuristics make greedy decisions based on static scoring functions. They do not anticipate long-term efficiency or leverage the growing availability of verification data to improve performance. This work introduces RSB, a reinforcement learning framework that learns to refine baseline branching heuristics. It trains an actor-critic architecture to maximize cumulative future rewards rather than immediate scores. The actor generates attention weights from observations of raw neuron features and learned graph embeddings, which rescale baseline heuristic scores to guide neuron branching. Evaluation on 600 challenging instances demonstrates that RSB consistently outperforms state-of-the-art branching heuristics, solving 11% more instances while reducing branch exploration by 50%.
PaperID: 1479, Poster
Abstract: Reasoning-oriented language model training increasingly relies on multiple on-policy rollouts, but standard Group Relative Policy Optimization (GRPO) uses a fixed group size, wasting compute on easy prompts and under-exploring hard ones. Existing adaptive-budget methods alleviate this inefficiency, yet their stopping rules are heuristic and do not provide a finite-sample calibration guarantee. We propose CGRPO, which replaces the fixed rollout budget of standard GPRO with calibrated conformal thresholds over task-specific nonconformity scores, combining Adaptive Prediction Set (APS)-style calibration for closed-form reasoning, a first-success score for execution-based code, automatic miscoverage selection, and periodic re-calibration. We prove finite-sample coverage for the resulting stopping rule and show that, under monotone policy improvement, recalibration induces a non-increasing expected training budget. Across six benchmarks spanning mathematics, graduate-level science, and code generation, CGRPO reduces mean training rollouts by 49-70% relative to fixed-budget GRPO while retaining strong fixed-budget accuracy, achieves the strongest pass@1 among evaluated static and dynamic baselines on different datasets, and transfers across model families with up to 82% rollout savings.
Abstract: Model merging combines task experts into one model and avoids joint training, retraining, or deploying many expert models, but the merged model often still underperforms task experts. We study this performance gap through feature drift, the difference between features produced by the merged model and by the expert on the same input. Our theory decomposes this drift into upstream propagation and local mismatch, tracks how it propagates and combines through later layers in forward order, and links final feature drift to output drift. This view motivates FeatCal, which uses a small calibration set to calibrate the merged model weights layer by layer in forward order, reducing feature drift while staying close to merged weights and preserving the benefits of model merging. FeatCal uses an efficient closed-form solution to update model weights, with no gradient descent, iterative optimization, or extra modules. On the main CLIP and GLUE benchmarks, FeatCal beats Surgery and ProbSurgery, the closest post-merging calibration baselines: 85.5% vs. 77.0%/78.8% on CLIP-ViT-B/32 Task Arithmetic (TA) and 85.2% vs. 83.7%/82.2% on FLAN-T5-base GLUE. On CLIP-ViT-B/32, 8 examples per task reach 82.9%, and 256 examples per task take 53 seconds, about 4x faster than both baselines, showing better sample efficiency and lower calibration cost.
PaperID: 1481, Poster
Authors:
Incheol Baek, Minseo Kim, Hyeonmin Kang, Yon Dohn ChungAbstract: Federated Learning (FL) enables collaborative model training without centralizing private data, but its performance often degrades under data heterogeneity because local updates become biased toward client-specific data characteristics. In this paper, we propose Constant Term Shrinkage for Federated Learning (FedCTS), a simple server-side aggregation method designed to mitigate such update drift. FedCTS decomposes each client update into a constant term and a residual term. We show that the constant term is strongly affected by data heterogeneity. FedCTS therefore evaluates the reliability of the constant term and adaptively shrinks the unreliable constant term toward zero while leaving the residual term unchanged. The shrinkage strength is obtained in closed form from clients' update statistics, without requiring an additional tuning hyperparameter. Experiments across multiple datasets and FL algorithms demonstrate that FedCTS improves robustness under heterogeneous data with negligible additional computation, no extra communication overhead, and strong compatibility with existing FL techniques.
PaperID: 1482, Poster
Authors: Nidhish Shah, Shaurjya Mandal, Asfandyar Azhar
Abstract: Audits of human-AI complementarity target the wrong estimand. They ask whether the human knows more than the model; deployment asks whether that knowledge would change the action. Residual predictive expertise is not residual decision value. Deployment rewards actions, and a posterior movement that does not cross a reward-induced action boundary changes nothing; a signal can be informative under log-loss and leave the deployed action unchanged. We formalize the missing quantity as : the regret of holding the model's action after observing the human. In any finite-action Bayesian decision problem, the value of consulting the human equals expected boundary regret, formally verified in Lean 4. The identity turns when humans help into a geometric question: which decision facets does the human signal move belief across? Complementarity is therefore reward-relative. A given human signal can be valuable under one reward matrix and worthless under another, the human's knowledge unchanged; the boundary moved, not the expertise. The gap is widest precisely where asymmetric costs dominate: high-stakes clinical and operational decisions. Audits built on residual predictive expertise overstate the value of human review and route scarce review toward cases where the human is informative but the action will not move. Boundary regret is the quantity to measure.
PaperID: 1483, Poster
Abstract: Self-distillation has emerged as a promising recipe for self-improvement in language models \citepzhao2026selfdistilledreasoneronpolicyselfdistillation,shenfeld2026self,hubotter2026reinforcement. In this setting, a model can be used as its own teacher when augmented with privileged information (e.g. a solution to a math problem). This seems like an especially appealing approach for thinking models that can leverage test-time reasoning to integrate learnings from privileged information. However, we show that privileged self-distillation degrades the long-budget test-time compute behavior of thinking models: across five Qwen3 and OLMo thinking models evaluated on AIME24, AIME25, and HMMT25, privileged-context distillation causes a relative drop of up to 17% in avg@16 accuracy. The degradation scales with the amount of privileged context withheld from the student and is most pronounced at long rollout budgets, where thinking models otherwise obtain their largest gains. This failure mode is not specific to self-distillation: on-policy distillation (OPD) improves thinking models, but privileged on-policy distillation reverses these gains. Our diagnostics suggest that this failure mode is linked to how privileged teacher context reshapes learning at high-entropy forking positions \citepbigelow2024forking, zhang2026embarrassingly, i.e., rollout positions where multiple continuations remain plausible and may lead to different reasoning paths. Privileged context lowers fork rates in thinking-model rollouts but not in instruction model rollouts. This leads to an interesting dichotomy wherein privileged context can help instruction-tuned models but hurts more performant thinking models that depend heavily on exploration and rollout quality. This effect is especially visible when the student begins a self-correction branch, where privileged OPD penalizes sampled reconsideration tokens that vanilla OPD supports. Thinking models trained with a privileged teacher produce fewer verification, backtracking, and hedging markers, even after length normalization. These findings indicate that applying self-distillation methods to strong thinking models requires further consideration of token-level signal---especially around tokens related to correction and crucial reasoning steps.
Abstract: We study visual persistence in autoregressive video world models: the Key-Value (KV) cache accumulates a growing visual memory, but once rollouts extend beyond the training horizon, the model can no longer reliably address stored content. A standard fix compresses the cache into a fixed-size memory that retains recent context while summarizing the distant past. However, we show compression alone cannot restore recall: past this horizon, temporal positional encodings go out-of-distribution, so the model cannot reliably retrieve stored history regardless of content. On the evaluated architecture, content compression with out-of-distribution positions yields results identical to a fixed-size sliding window over recent KV entries. We propose WorldTrace, a training-free framework that keeps compressed memory addressable by assigning each slot a fixed, in-distribution position relative to the current frame. With addressable memory, we explore two retention approaches: WorldTrace-Field (coherence-oriented) aggregates history in a rotation-invariant space, improving TempSSIM by +15.5% while reducing scene drift; WorldTrace-LandMark (recall-oriented) stores verbatim scene traces at detected boundaries. The recall-oriented variant sustains scene reconstruction over long rollouts, with stronger long-range recall and smooth temporal coherence.
Authors:
Ba-Thinh Lam, Srijan Das, Hieu LeAbstract: We propose importance-aware pruning for diffusion models, a training-free framework that prioritizes preserving parameters critical to semantically salient image regions. To do so, we incorporate spatial importance maps- derived from conditioning signals or model attention- into the pruning objective. This produces parameter rankings aligned with perceptual relevance rather than uniform reconstruction error. On MS-COCO dataset, our proposed approach consistently retains subject fidelity and structural correctness at high compression ratios where conventional pruning causes visible degradation. These results demonstrate that content-aware objectives are key to perceptually faithful compression of generative models.
PaperID: 1486, Poster
Abstract: Exact machine unlearning guarantees that after deleting a user's data, the model produces the same predictions as if it had never seen that data in the first place. SISA, the current state-of-the-art approach, partitions data into disjoint shards, with necessary O(n/S) retraining cycle for each deletion, and it cannot lower this cost through redundancy, because overlapping shards only multiply the damage.We introduce Quantized Sufficient Statistics (QSS), a non-parametric framework that pairs frozen prediction heads (the schema, trained on a small random subset \mathcalD^\circ) with mutable sum-decomposable accumulators indexed by Residual Quantization codes over the complementary data (the content, \mathcalD \setminus \mathcalD^\circ). With probability 1-\rho (where \rho = |\mathcalD^\circ|/n, typically \leq 1%), deletion reduces to a constant-time arithmetic update independent of dataset size n; with probability \rho, a full rebuild is required. Maintaining p independent schemas suppresses the synchronous rebuild probability to \rho^p, and because every schema aggregates all data, decommissioning incurs zero accuracy loss. Across 15 datasets (n \geq 50K) spanning vision, text, and tabular modalities at \rho=0.5%, QSS matches or exceeds SISA accuracy on 5 datasets, trails by \leq 2 pp on 5 more, and by 3–11 pp on 3 datasets where high class count or large rebuild cost limits performance, while delivering 3–645× faster median expected deletion (e.g., 0.2 s expected vs. 109 s on Jigsaw, n=1.4M).
Abstract: Predicting how genetic perturbations change cellular state is a core problem for building controllable models of gene regulation. Perturbations at the same gene can produce different transcriptional responses depending on their genomic locus, including different transcription start sites and regulatory elements; gene-level perturbation models collapse these distinct interventions into a single representation. We introduce STRAND, a generative model that predicts single-cell transcriptional responses by encoding regulatory DNA sequence at the perturbation locus and using it to parameterize a conditional transport process from control to perturbed cell states. Sequence conditioning supports zero-shot inference at perturbation targets unseen during training and locus-resolved inputs across a 524\,kb window at 128\,bp model resolution. On CRISPRi datasets in K562, Jurkat, and RPE1 cells, STRAND improves few-shot perturbation discrimination by up to 33%, achieves the best average rank on unseen-gene benchmarks, and improves held-out-cell-line transfer by up to 0.12 in \DeltaPearson; ablations attribute these gains to both sequence conditioning and conditional transport. At the locus level, STRAND distinguishes alternative transcription start sites and validated regulatory elements that gene-identifier models treat as identical perturbations
PaperID: 1488, Poster
Authors:
Tianyi Xu, Yi Niu, Xinkun Wang, Yanbiao Ma, Mingming Ma, QingYu Luo, Fu Zou, Fu LiAbstract: Object detectors are expected to recognize objects independently of intrinsic object-level attributes, such as scale and local visibility. Yet, natural detection datasets contain systematic intrinsic attribute biases, and modern detectors inevitably entangle these biases within their learned representation geometry, leading to non-uniform performance across the continuous attribute spectrum. This observation exposes a central limitation of conventional data augmentation: while attribute-blind transformations enlarge the training distribution, they fail to explicitly correct the dataset-inherent nuisance structure. To address this, we propose Target-Aware Nuisance Shaping (TANS, a unified framework that re-purposes data augmentation as an object-conditioned vicinal distribution design. TANS processes training samples through a serial pipeline of two fundamental operators: first, Scale Prior Shaping (SPS) geometrically transports object scales toward low-density regions of the natural scale distribution via class-conditional log-area priors; subsequently, Local Visibility Shaping (LVS) photometrically modulates local foreground-background contrast by constructing smooth, box-conditioned Gaussian fields, strictly preserving the newly established geometry and labels. Furthermore, a feature-state controller monitors the evolving representation geometry to adaptively adjust the shaping strength during training. Extensive experiments on MS COCO and Pascal VOC demonstrate that TANS significantly improves detection performance over strong augmentation baselines across modern architectures. Specifically, on COCO with YOLOv11, TANS achieves a 3.6 mAP gain over the Mosaic baseline. Beyond aggregate accuracy, our attribute-regime analyses, feature-state dynamics, schedule comparisons, and logit-space probes consistently indicate that TANS effectively decouples the detector from intrinsic dataset biases, establishing target-aware shaping as a rigorous paradigm for learning attribute-invariant representations.
PaperID: 1489, Poster
Abstract: Deep neural networks achieve strong performance across many tasks but remain vulnerable to adversarial perturbations. Adversarial training (AT) is one of the most effective defenses, yet it suffers from computational cost on large models and the persistent clean-robust accuracy trade-off. Recent work introduces parameter-efficient fine-tuning, such as LoRA, into AT, but this constraint induces a notable robustness gap relative to full-parameter AT. In this work, we revisit parameter-efficient adversarial fine-tuning from a parameter space perspective. Through gradient analysis, we find that adversarial optimization is highly concentrated on a small subset of parameters, which we call Robustness-Critical Parameters (RCP), suggesting that robustness is encoded in a sparse subspace. Building on this observation, we propose Sparse Adversarial Fine-Tuning (SAFT), which identifies RCP via adversarial gradient saliency and performs adversarial training by updating only these parameters while freezing the rest. Across architectures and datasets, SAFT consistently outperforms LoRA-based methods in both clean and robust accuracy with fewer trainable parameters, and approaches or surpasses full parameter AT in robustness while updating only about 5% of parameters.
PaperID: 1490, Poster
Authors:
David V Zarb, Rustem Turtayev, Taywon Min, Jinghua Ou, Shi FengAbstract: Alignment faking, where models comply with harmful requests more often when they infer they are being trained than when they infer they are deployed, has been read as evidence of scheming. The same behavior is also consistent with performative scheming: the model identifies the setup as a safety evaluation and produces the behavior it infers researchers are looking for. These hypotheses predict identical behavior in the standard evaluation but generalize differently. We distinguish them by amplifying and suppressing the process each requires (consequence-tracking for scheming, expectation-tracking for performative scheming) and comparing using three methods: synthetic document fine-tuning, activation steering, and prompt interventions. Across four open-weight model organisms that exhibit alignment faking in the helpful-only scenario, fine-tuning produces larger shifts under expectation-tracking than consequence-tracking on three (Llama-3.1-70B, Llama-3.1-405B, Qwen-2.5-72B); on QwQ-32B, both intervention families shift compliance globally rather than changing the gap, suggesting they affect overall helpfulness rather than alignment faking specifically. Activation steering on Llama-3.1-70B reproduces the asymmetry: expectation-tracking steering changes compliance, while consequence-tracking steering changes the model's reasoning without affecting compliance. Alignment faking in these models is better explained by expectation-tracking than by self-preservation reasoning.
Abstract: We introduce AIRA-Compose — a framework that enables AI research agents to autonomously explore and discover novel neural architectures. AIRA-Compose leverages the previously-proposed Composer abstraction to define a combinatorial design space over fundamental model primitives (Attention, MLP, and Mamba), then uses a guided search policy to efficiently navigate this space for novel neural architecture design. Our approach operates in two stages: (1) agents iteratively designs and evaluates candidate architectures at the million-parameter scale, then (2) top-performing candidates are scaled to billion parameters for pre-training. Overall, AIRA-Compose discovers 14 novel architectures spanning two families — AIRAformers (Transformer-based) and AIRAhybrids (Transformer-Mamba Hybrid-based). When pre-trained at 1B scale under a fixed token budget, agent-discovered top-performing architectures consistently outperform both Llama 3.2 and Composer-found alternatives. On downstream tasks, AIRAformer-D and AIRAhybrid-D improve accuracy by 2.4% and 3.8% over Llama 3.2, respectively. AIRA-Compose also finds novel model architectures that achieve steeper, more efficient compute-optimal scaling frontiers. AIRAformer-C scales 54% and 71% faster than Llama 3.2 and the best Composer-found Transformer, while AIRAhybrid-C scales 23% and 37% faster than the modified Nemotron-2 and the best Composer-found hybrid, respectively. For the first time, AIRA-Compose demonstrates that AI research agents can discover hybrid architectures that surpass hand-designed baselines on the accuracy–compute frontier. It establishes a flexible paradigm for LLMs to discover next generation foundation models, marking a step towards recursive self-improvement.
PaperID: 1492, Poster
Abstract: Model merging offers a training-free paradigm for integrating multiple specialized models into a unified multi-task system, enabling efficient knowledge reuse from models. However, parameter interference remains a critical challenge, limiting the scalability and efficacy of merged models. SVD-based merging techniques represent the current state-of-the-art, excelling in alleviating interference through task-specific knowledge extraction and subspace orthogonalization. % due to aggressive parameter truncation Despite these advances, they inevitably incur performance degradation due to aggressive parameter truncation, which discards valuable information and hinders inference accuracy. Our analysis reveals this truncation-induced knowledge loss as the fundamental bottleneck, motivating the need for mechanisms that preserve and recover discarded expertise without retraining. To address this, we introduce S MoE, a novel framework that augments SVD-based merging with a mixture of sparse experts to refine the truncation gap. S MoE comprises three key components: (1) a shared module that aggregates multi-task knowledge, (2) sparse experts that selectively recover truncation knowledge, and (3) a subspace mapping router to achieve precise training-free routing. Extensive experiments across vision and language benchmarks demonstrate up to 1.6% and 1.2% performance gains over SOTA methods, effectively validating robust and scalable model merging.
PaperID: 1493, Poster
Authors: Sejeung Choi, Minji Jung, Jinyeong Park, Taek D Chung
Abstract: Biological recurrent circuits can reuse a common substrate by changing its operating regime rather than replacing the circuit itself. This principle is difficult to isolate in fully adaptive neural models, where recurrence, controllers, readouts, and internal representations can all absorb task structure. As a result, successful multi-context behavior does not reveal which route made the computation possible. We introduce Constrained Modulatory Reservoirs (CMR), a fixed-substrate framework for studying such reconfiguration with the main bypasses removed. The primary dynamics and low-rank perturbation basis are fixed, and the readout is denied direct access to the cue, forcing input-dependent adaptation to act through a limited set of recurrent connectivity changes. In transient-cue channel-selection tasks, this restriction makes a capacity transition visible. Low-loss behavior appears only when the available perturbation directions are sufficient, the selected change is retained across the cue-to-signal gap, and the primary dynamics express it during the response window. Controls show that this effect is not explained by low-dimensional cue representation alone. Allowing the substrate to learn hides the bottleneck, while activity- and gain-level alternatives fail to reproduce the same regime. CMR therefore reframes flexible recurrent computation as a question of where adaptive influence is allowed to enter a fixed dynamical system, making capacity, memory, and expression experimentally separable.
PaperID: 1494, Poster
Authors:
Charles Bricout, Gennaro Gala, Viacheslav (Slava) Borovitskiy, Antonio VergariAbstract: Probabilistic circuits (PCs) have emerged as generative models guaranteeing tractable probabilistic inference. However, tractability comes at the price of larger model sizes than in other neural generative models. As such, to scale PCs to high-dimensional data such as images, previous work has devised several parameter sharing techniques as well as fast evaluation protocols that consider learning PCs on independent image patches. In this work, we revisit this practice to highlight a surprising learnability issue: models that assume independence counterintuitively can yield better likelihoods than models that do not. To break this apparent bottleneck and to scale PCs further, we introduce a recursive parameter sharing construction (RePA) that fosters efficiency without assuming independence. In a set of rigorous experiments, we show that RePA can greatly reduce time and memory consumption while yielding competitive likelihoods.
Abstract: Existing Driving VLAs predict trajectories while largely ignoring their visual tokens– a phenomenon we trace not to insufficient training but to a structurally ill-posed task formulation. We show that trajectory recovery, when viewed through the lens of inverse kinematics, requires both a current and a future visual state as boundary conditions; existing VLAs supply only the former, which encourages the model to shortcut through ego status and text commands alone. To address this, we re-design Driving VLA in the style of an inverse kinematics solver. First, a next visual state prediction objective that requires the LLM to predict the future visual scene provides dense visual supervision and suppresses shortcut paths. Second, a separate Inverse Kinematics Network (a cross-attention-based conditional diffusion model) that takes only the current and future visual states as input is designed to suppress reliance on ego status and textual shortcuts during trajectory decoding. With this simple prescription alone, our 0.5B-scale model recovers visual grounding and reaches trajectory planning performance comparable to 7B–8B VLAs more than an order of magnitude larger, on both the closed-loop NAVSIM-v2 and the nuScenes benchmarks. Extensive analysis further shows that this improvement stems from a recovered ability to exploit visual features, with the effect being most pronounced in dynamic driving situations such as turning.
PaperID: 1496, Poster
Abstract: While reasoning traces improve model capability, they also create a pathway for extracting supervision signals for downstream distillation. Initial efforts to mitigate this rely on decoding-time interventions, but they incur high inference overhead and often harm the teacher model's natural fluency. We propose Conflict-Aware Logit Adapters (CALA), a lightweight training-time defense. CALA attaches a small residual adapter to a frozen teacher model and trains only this adapter. A proxy student identifies positions where imitation is most sensitive, and the adapter perturbs the output distribution there within high-probability regions. A conflict-aware optimization balances competing anti-distillation and utility objectives, with regularization preserving the teacher’s performance and fluency. At inference, CALA adds negligible overhead with no auxiliary models required. Experiments on reasoning benchmarks show that CALA substantially reduces successful student imitation while maintaining near-original teacher performance and generation quality, offering a practical and efficient alternative to decoding-time approaches.
Authors: Felix Biggs, Samuel Willis
Abstract: Recent work has demonstrated surprisingly good performance of pre-trained LLMs on regression tasks (for example, time-series prediction), with the ability to incorporate expert prior knowledge and the information contained in textual metadata. However we observe major error cascades even in short sequences \lesssim 100 points; these models are also computationally intensive and difficult to parallelise. Marginal LLM predictions do not suffer this issue and are trivially parallelised, but can predict over-broad densities. To address this, we propose combining these densities with a lightweight (diffusion-based) neural process. We show that this combination leads to better-calibrated predictions overall, outputs locally consistent trajectories, and leads to text-conditioned function space selection in the meta-learner. As part of this work we propose a gradient-free (and non-Monte Carlo) method for sampling from a product-of-experts of a score model and an `expert' (here the LLM predictive densities). We believe this general method is of independent interest as it is applicable whenever an expert can be convolved with a Gaussian in closed form.
PaperID: 1498, Poster
Abstract: We study preference-based bandits with general reward function classes, where a learner sequentially selects pairs of arms and observes binary preference feedback governed by the Bradley-Terry model. This setting naturally arises in applications such as recommender systems, tournament ranking, and learning from human feedback, where relative preferences are easier to elicit than absolute rewards. The observation model inherits the logistic bandit challenge of handling the problem-dependent constant \kappa, which accounts for the non-linearity of the link function and can grow arbitrarily large. Moreover, prior work has predominantly focused on linear or kernelized reward models, precluding the use of richer function classes. To address these limitations, we consider general reward function classes and introduce the locally sensitive eluder dimension, a novel complexity measure tailored to the logistic structure of preference feedback that yields fine-grained regret guarantees without unfavorable dependence on \kappa. Building on this notion, we propose GINOP (Generic INformative OPtimism), an algorithm that constructs log-loss confidence sets and jointly selects arm pairs to balance optimism and informative exploration. We establish a first-order regret bound that, in contrast with what previous results suggest, demonstrates that learning with preference feedback is as statistically efficient as learning from direct reward observation. Finally, we corroborate our theoretical findings with empirical evaluations against competitive baselines.
PaperID: 1499, Poster
Abstract: Federated Learning (FL) is a promising paradigm for distributed machine learning. However, FL often suffers from degraded generalization performance due to the inconsistency between local and global optimization objectives and client-side overfitting. In this paper, we provide a global-update stability analysis framework as an analytical tool to study generalization error and derive the stability bounds of mainstream FL optimization algorithms under non-convex settings. Our analyses reveal how the number of global update steps, data heterogeneity, and update rules influence their stability. We observe that momentum-based FL acceleration methods do not improve stability. To address this issue, we propose FedSEMA, a new FL algorithm that couples global momentum with a gradient-difference corrected Nesterov Accelerated Gradient (NAG) scheme and a hybrid proximal term to enhance stability. This design ensures updates follow a globally consistent descent direction while retaining the benefits of acceleration. Theoretical analysis shows that FedSEMA achieves an improved stability upper bound on non-i.i.d. datasets in the non-convex settings. Extensive experiments on real-world datasets demonstrate that FedSEMA significantly outperforms multiple baseline methods under standard FL settings, achieving faster convergence and state-of-the-art performance.
Authors:
Dingyu Yao, Shuhuan Gu, Qingyi Si, Junhao Zhou, Chenxu Yang, Chuanyu Qin, Naibin Gu, Zheng Lin, Weiping Wang, Nan Duan, Jiaqi WangAbstract: Vision-Language Models (VLMs) are increasingly required to process unbounded video streams in applications such as video-call assistants, live commentary, and embodied robots. An ideal streaming system should support proactive interaction, long-horizon memory, and real-time processing, while resting on a VLM backbone capable of handling diverse in-the-wild streaming tasks. However, existing VLMs excel at offline video understanding but fall short in streaming capabilities and lack dedicated infrastructure for streaming deployment. We address this gap on three fronts. (i) For backbone capability, we construct Streaming-Train-248K, a streaming dataset paired with a novel training objective for adapting VLMs to streaming interaction and understanding. (ii) For real-world deployment, we introduce Streaming Harness, a plug-and-play system that endows any VLM with three core abilities: proactive interaction (per-second response decisions), long-term memory (12-hour context retention), and real-time processing (sub-second latency). (iii) To drive continued community progress on streaming capabilities, we design Streaming-Eval, a benchmark that reflects models' capabilities across diverse in-the-wild scenarios. Extensive experiments demonstrate consistent gains from our approach across all core capabilities required for streaming video understanding. We will open-source our data, code, and benchmark to advance the community's shift from offline video understanding to deployable streaming intelligence.
PaperID: 1501, Poster
Authors: Nikita Tsoy
Abstract: Despite their practical significance, modern AI systems pose significant societal risks, including toxicity, malicious use, and misinformation. Their mitigation depends not only on technical feasibility but also on the incentives of firms that develop and deploy these systems. We show how profit maximization shapes safety investment in an AI supply chain with an upstream LLM provider and several competing downstream firms. In our game-theoretic model, we derive the market-driven level of safety investment and compare it with the levels that maximize market welfare and industry profit. We find that downstream competition generally induces downstream firms to invest more in safety than the industry-profit-optimal level. By contrast, the upstream provider may underinvest relative to these benchmarks because it is not fully compensated for the surplus from safety investments. Finally, we analyze how different regulatory interventions affect the participants' incentives and profits, and the resulting levels of model safety.
Abstract: Vision-Language-Action (VLA) models inherit rich semantic priors from pretrained vision-language models, yet these priors do not guarantee persistent instruction sensitivity after narrow-domain fine-tuning. We observe a systematic failure mode: although VLA models remain visually reactive, their later actions can become weakly constrained by the commanded instruction, leading to unstable object interaction and visually plausible but instruction-inconsistent behavior. We trace this behavior to a collapse of instruction-conditioned representations across network depth and task progress. More fundamentally, this collapse is enabled by the learning objective itself: standard imitation learning admits Bayes-optimal solutions that are invariant to language. We address this by recasting instruction grounding as identifying the commanded instruction among counterfactual alternatives. Based on this view, we propose Counterfactual Instruction Grounding (CIG), a contrastive objective that encourages the generated trajectory to remain identifiable with the commanded instruction among counterfactual alternatives. CIG applies to both autoregressive and flow-matching VLA models, using exact action-chunk likelihoods for autoregressive models and an energy-based pseudo posterior for flow-matching models to avoid intractable trajectory likelihood estimation. Practically, CIG can be applied directly to already fine-tuned models as a continuation-stage grounding objective, restoring instruction sensitivity without retraining from scratch or changing the architecture. Extensive experiments on RoboCasa and LIBERO-Plus demonstrate that CIG achieves strong instruction-following performance with success rates of 71.3% and 86.3%, respectively. Real-world deployment further shows that CIG transfers to kitchen manipulation and preserves instruction-conditioned behavior in visually ambiguous scenes.
PaperID: 1503, Poster
Authors: YiZhou Li, Jinyi Xu, Qingyi Yang, Kefan Sun
Abstract: Retrieval-augmented generation (RAG) reduces hallucinations by grounding large language models in external knowledge bases. However, most existing methods focus on improving how knowledge is indexed, retrieved and integrated, neglecting a critical source of hallucination: the lack of unknown awareness. Unaware of the knowledge boundaries of the corpus, a model may stitch together fragmented context to hallucinate unsupported relations. To address this, we introduce Substrata, an unknown-aware RAG framework that makes the unknown locatable. Substrata maintains an evolving knowledge graph via Dempster--Shafer evidence fusion, enabling local graph updates instead of costly full graph reconstruction in graph-based RAG. Crucially, we apply a topological analysis method, persistent homology, to detect long-lived structural holes where the corpus fails to close relational gaps within this graph. These holes are converted into retrievable unknown annotations. During inference, Substrata performs hybrid retrieval over both factual chunks and unknown annotations, allowing the generator to access both supporting context and explicit signals of missing contextual support. Experiments show that Substrata achieves the best performance on three public multi-hop QA benchmarks, with an average of 60.66% EM and 71.25% F1 across HotpotQA, MuSiQue, and 2WikiMultiHopQA. On a hallucination-focused domain benchmark, it reaches 95.0% accuracy on Spurious Relational Premise questions while reducing overall token consumption by 10.3× compared with GraphRAG.
Abstract: In this paper, we study federated optimization for solving stochastic variational inequalities (VIs), a problem that has attracted growing attention in recent years. Despite substantial progress, a significant gap remains between existing convergence rates and the state-of-the-art bounds known for federated convex optimization. In this work, we address this limitation by establishing a series of improved convergence rates. First, we show that, for general smooth and monotone variational inequalities, the classical Local Extra SGD algorithm admits tighter guarantees under a refined analysis. Next, we identify an inherent limitation of Local Extra SGD, which can lead to excessive client drift. Motivated by this observation, we propose a new algorithm, the Local Inexact Proximal Point Algorithm with Extra Step (LIPPAX), and show that it mitigates client drift and achieves improved guarantees in several regimes, including bounded Hessian, bounded operator, and low-variance settings. Finally, we extend our results to federated composite variational inequalities and establish improved convergence guarantees.
Abstract: Given n i.i.d. random matrices A_i \in \mathbbR^d × d with common expectation \Sigma, the goal of Differentially Private Stochastic PCA is to identify a k-dimensional subspace capturing the leading variance directions of \Sigma, while preserving differential privacy (DP) for each individual sample A_i. Düngler & Sanyal [2025] introduced k-DP-PCA, the first algorithm to simultaneously (1) achieve sample complexity n=\widetilde O(d) for sub-Gaussian data, (2) adapt its privacy noise to the intrinsic randomness of the data, and (3) extend seamlessly to any target dimension k\le d. However, its sample complexity has a suboptimal dependence on k. We propose the first algorithm that achieves optimal sample complexity in both d and k, while retaining (2) and (3) of k-DP-PCA. In addition, our method removes the exponential dependence on the eigengap that appears in the sample-size lower bound required by prior utility guarantees, and improves the dependence on spectral parameters to match known lower bounds for the spiked covariance model. Unlike deflation-based approaches like k-DP-PCA, our algorithm updates the full d× k subspace jointly rather than one eigenvector at a time. This non-deflation structure simplifies the algorithm, reduces the number of hyper-parameters, and improves computational efficiency, particularly when using block linear algebra libraries.
Abstract: We introduce a simple post-training method that makes transformer attention sparse without sacrificing performance. Applying a flexible sparsity regularisation under a constrained-loss objective, we show on models up to 7B parameters that it is possible to retain the original pretraining loss while reducing attention connectivity to \approx 0.4% of its edges. Unlike sparse-attention methods designed for computational efficiency, our approach leverages sparsity as a structural prior: it preserves capability while exposing a more organized and interpretable connectivity pattern. We find that this local sparsity cascades into global circuit simplification: task-specific circuits involve far fewer components (attention heads and MLPs) with up to 100× fewer edges connecting them. Additionally, using cross-layer transcoders, we show that sparse attention substantially simplifies attention attribution, enabling a unified view of feature-based and circuit-based perspectives. These results demonstrate that transformer attention can be made orders of magnitude sparser, suggesting that much of its computation is redundant and that sparsity may serve as a guiding principle for more structured and interpretable models.
Abstract: Predictions on digital platforms must adapt over time, as individuals continuously update their beliefs through social interactions. At the same time, changing predictions can in turn influence the content people are exposed to and hence the very beliefs they seek to predict, rather than merely serving as passive observations. These emerging dynamics make it challenging to understand the long term effects of predictive systems on society. In this work, we blend models from network science with concepts from performative prediction to initiate the study of opinion dynamics in predictive loops. In our model, opinions and predictions co-evolve: a platform's predictions influence individual opinions, which then evolve through peer interactions and form the training data for future platform model updates. We demonstrate that this co-evolution induces a novel equilibrium that qualitatively differs from standard network equilibria. In particular, we show how standard predictive objectives can drive networks toward consensus even under conditions where classical opinion-dynamics models lead to disagreement. This emerges because predictive systems dynamically adapt to changing opinions, and learning objectives create spillover effects among individuals beyond the topology of the network. We further analyze systematic deviations from standard prediction and demonstrate amplified effects of targeted platform interventions on equilibrium outcomes, compared to classical network intervention analyses. We complement our results with simulations on real social network data and parametric learning settings. Together, our results illustrate performativity as an important, yet so far neglected, qualifying factor in social networks.
PaperID: 1508, Poster
Authors:
Koki Nagano, Hongyu Liu, Wookie Park, Tianye Li, Amrita Mazumdar, Christian Jacobsen, Shengze Wang, Michael Stengel, Ka Chun Cheung, Simon See, Shalini De MelloAbstract: We present DuplexMotion, a streaming, full-duplex speech-and-motion model designed for dyadic interaction. To capture the continuous and reciprocal nature of human communication, this full-duplex capability empowers the agent to simultaneously perceive and generate both speech and physical motion in a streaming fashion. At its core, our method leverages the strong priors of a foundational full-duplex speech model and integrates a novel motion pathway, thereby achieving fully synchronized multi-modal interaction. Specifically, we design a dual-tower Transformer architecture that preserves the zero-shot conversational reasoning of a frozen base speech model while constructing a deeply coupled, streaming motion pathway. By introducing a unified dyadic token interleaving mechanism and guiding cross-attention via a time-aligned speech-motion RoPE, our model effectively aligns autoregressive motions with rich latent speech features. Trained on the 4,000-hour Seamless Interaction dataset, our model effectively captures cross-speaker dependencies and establishes new state-of-the-art performance across both monadic and dyadic human interaction benchmarks.
Authors:
Mingda Liu, Zhenghan Zhu, Ze'an Miao, Katsuki FujisawaAbstract: Sequential Locate-and-Edit (L\&E) model editing can fail abruptly after many edits. We identify and formalize this failure as a positive _norm-feedback loop_, in which solved value vectors and edited MLP weights progressively amplify each other, degrading edit quality and eventually collapsing model capabilities. Our analysis shows that this feedback can yield approximately exponential norm growth under standard L\&E dynamics, and can remain unresolved by existing increment-level regularizers or update clamps. We propose Norm-Anchor Scaling (NAS), a plug-in stabilizer that breaks this loop by rescaling each solved value vector to an original-model reference norm. Across multiple LLM backbones, datasets, and L\&E editors, NAS extends the usable editing horizon by more than 4× and improves long-run editing performance by 72.2% on average, while preserving single-edit efficacy, with only a one-line modification and negligible computational overhead. The code is available in the supplementary materials.
Abstract: Learning-based simulation of multi-object rigid-body dynamics remains difficult because contact is discontinuous and errors compound over long horizons. Most existing methods remain tied to mesh connectivity and vertex-level message passing, which limits their applicability to mesh-free inputs such as point clouds and leads to high computational cost. Efficiently modeling high-fidelity rigid-body dynamics from mesh-free representations, therefore, remains challenging. We introduce RigidFormer, an object-centric Transformer-based model that learns mesh-free rigid-body dynamics with controllable integration step sizes. RigidFormer reasons at the object level and advances each object through compact anchors; Anchor-Vertex Pooling enriches these anchors with local vertex features, retaining contact-relevant geometry without dense vertex-level interaction. We propose Anchor-based RoPE to inject anchor geometry into attention while respecting the unordered nature of objects and anchors: object-token processing is permutation-equivariant, and the mean-pooled anchor descriptor is invariant to anchor reindexing while preserving shape extent. RigidFormer further enforces rigidity by projecting updates onto the rigid-body manifold using differentiable Kabsch alignment. On standard benchmarks, RigidFormer outperforms or matches mesh-based baselines using point inputs, runs faster, generalizes to unseen point resolutions and across datasets, and scales to 200+ objects; we also show a preliminary extension to command-conditioned articulated bodies by treating body parts as interacting object-level components. Code will be released upon publication.
Abstract: Reasoning LLMs produce longer outputs, requiring speculative decoding drafters trained on extended sequences. Parallel drafting—predicting multiple tokens per forward pass—offers latency benefits over sequential generation, but training complexity scales quadratically with the product of sequence length and parallel positions, rendering long-context training impractical. We present P(arallel)-EAGLE, which transforms EAGLE from autoregressive to parallel multi-token prediction via a learnable shared hidden state. To scale training to long contexts, we develop a framework featuring attention mask pre-computation and sequence partitioning techniques, enabling gradient accumulation within individual sequences for parallel-prediction training. We implement P-EAGLE in vLLM and demonstrate speedups of 1.10×–1.36× over autoregressive EAGLE-3 across GPT-OSS 120B, 20B, and Qwen3-Coder 30B.
Abstract: Repeated AI assistance can improve immediate task performance while reducing the skill available for future independent work. We develop a mathematical framework for this long-run tradeoff. The model tracks two state variables: a latent human skill level governing expected independent performance, and a delegation level representing the learner's evolving tendency to rely on AI. Skill changes through error-driven learning under practice and decay under delegation; delegation responds to observed performance, increasing when AI-assisted work appears to outperform independent work. We analyze the resulting dynamics and contrast them with fixed delegation. With fixed delegation, skill follows a one-dimensional learning-decay process with a single stable equilibrium. With adaptive delegation, the coupled system has two attracting equilibria separated by the stable manifold of an interior saddle. The existence and geometry of this separatrix require a global phase-plane analysis of the coupled dynamics. The system is path-dependent: small differences in initial skill or reliance can lead to different long-run outcomes. We use this characterization to show that AI assistance can improve short-run performance while producing worse long-run performance than a no-AI baseline. Increasing AI capability can enlarge the basin of attraction of the low-skill equilibrium, making delegation appear beneficial for longer while increasing the risk of eventual skill loss. The qualitative picture is observed to persist across alternative specifications. Together, these results show that the risk is not AI assistance itself, but the coupling between performance-driven reliance and use-dependent skill change.
PaperID: 1513, Poster
Authors: zirui wang, Guangqiang He, PangWu, Peng Wang, Zhenfeng Li, Xianxiang Chen, Lidong Du, Zhen Fang
Abstract: Modern transformation-aware self-supervised representations are routinely scored under protocols that grant the evaluator paired views, augmentation labels, or teacher anchors --- none of which the deployed predictor ever sees. We audit the gap that opens once these privileges are removed and only a single query remains. The finding is a systematic evaluation illusion: on 3DIEBench-Honest, honest single-query rotation errors are typically 2-3× as large as native transformation scores, consistently across comparators and matched-utility thresholds. To make this gap measurable, we introduce a deployment-honest evaluation contract specifying a one-query test interface, matched semantic utility, a strongest-eligible-neighbor headline with tail summaries, and explicit privilege accounting. Theory and experiments are in service of the contract: lower-bound analysis grounds the contract at the interface level, while a positive control supplies a non-vacuous upper certificate under declared privileges. We exercise the contract on a generated law, on 3DIEBench-Honest, and on PTB-XL. The operational conclusion is that native transformation-aware success is not, on its own, a single-query deployment certificate.
Abstract: Particle-based variational inference (ParVI) methods approximate an intractable target distribution by evolving an ensemble of interacting samples. Existing approaches rely predominantly on kernel-based repulsion (e.g., SVGD), which suffers from variance collapse in high dimensions and mode collapse on multimodal targets—pathologies caused by the absence of global transport structure. We introduce entropic transport descent (ETD), a ParVI family that frames each particle update as an entropy-regularized optimal transport problem. Derived from the JKO proximal scheme by lifting to the space of couplings and relaxing via the KL chain rule, each ETD iteration reduces to a Sinkhorn computation. The resulting transport plan provides global coordination, guiding each particle to nearby high-density proposals and naturally preserving multimodal structure. ETD can operate entirely score-free, requiring only pointwise evaluations of the unnormalized target density. Experiments on variance-collapse diagnostics, Bayesian logistic regression, neural networks, and molecular Boltzmann distributions show that ETD matches or outperforms SVGD, AGF-SVGD, and SGLD, with the largest gains in high-dimensional and multimodal settings.
Abstract: Scaling large language models (LLMs) requires tremendous computational resources, and recent advances in AI have gone hand in hand with massive amounts of capital expenditure. While it is established that scaling up LLMs reliably increases model quality (quantified in terms of loss or downstream evaluations), it is unclear how these quality improvements translate to potential revenue, and whether revenue increases would offset costs of larger-scale training and inference. In this work, we develop an economic model for characterizing the rational behavior of an LLM training firm by combining scaling laws with microeconomic theory. Under our model, LLM quality can be increased with more parameters and training tokens, leading to more potential adoption by consumers, who each have a quality threshold for using the LLM. On the other hand, additional parameters and training tokens both incur additional costs. We analyze the profit maximization problem for this model under compute-bound and data-bound regimes. In the compute-bound regime, optimal model size and token budget track hardware efficiency E (FLOPs/dollar) at a near-linear rate; total training cost then scales sub-quadratically in E. Data efficiency improvements incentivize larger models and training expenditure. When we are limited to D data, profit-optimal training expenditure scales as D^2 / E, i.e, increase with data and decreases with hardware efficiency (as well as data efficiency). Finally, we analyze practical trends in training expenditure: current trends in training expenditure are consistent with our most permissive model variants in the compute-bound regime, but are not profit-optimal in the data-bound regime or assuming hardware advances will stall. Overall, our results provide a theory of profit-optimal LLM training, providing a foundation for engaging critically with industry statements and supporting long-term economic decision making.
Abstract: Autoregressive scientific forecasters often enforce physical or structural constraints by repairing each predicted state before feeding it back into the model. However, it remains unclear when stronger physical rule enforcement becomes reliable and when it becomes a source of distribution shift. We study this question through operator exactness, meaning whether the repair map is the identity on the target manifold and is aligned with the target geometry. We compare raw forecasting, post hoc repair, and in-loop repair across periodic incompressible Navier--Stokes, non-periodic CFDBench flows, and a hierarchical-forecasting support task. In the exact periodic regime, Fourier projection substantially improves rollout accuracy. On the NS-128 benchmark, a strong Raw-FNO has a final-step rollout MSE at horizon 100 of (9.390 \pm 6.290)× 10^-5, and post hoc and in-loop projection reduce it to (1.130 \pm 0.165)× 10^-6 and (5.370 \pm 0.113)× 10^-7. However, once an exact projection is unavailable and only approximate boundary-preserving cleanup is available, the ordering changes. Across cavity, tube, dam, and cylinder flow, stronger Poisson-based cleanup can reduce divergence while worsening rollout error; target-distortion MSE predicts this harm far better than a linear-system residual. Controlled mismatch, screened cleanup, adaptive gating, and external-backbone checks show that the best approximate-regime operating point can be raw or near-identity. Hierarchical forecasting gives the same broader pattern. Exact forecast reconciliation is a stable baseline, whereas blended top-down repair, a validation-tuned interpolation toward historical-proportion top-down reconciliation, is dataset-dependent. Thus, constraint enforcement should be benchmarked by operator--data alignment before enforcement strength. Use in-loop projection when the operator is exact, and validate approximate cleanup strength using rollout metrics otherwise.
Abstract: Safe reinforcement learning (RL) aims to learn policies that optimize rewards while satisfying constraints. Predominant approaches rely on soft-constrained policy optimization, which has achieved empirical success but does not provide formal safety guarantees for the learned policy. In contrast, methods with strict guarantees typically rely on explicit certificate functions, whose construction requires the direct synthesis and verification of control-invariant sets, a process that scales poorly with state dimension and often yields overly conservative behavior. In this paper, we present the Provably Safe, yet Performant RL (PSP-RL) framework, a novel two-phase architecture for learning provably safe policies in a scalable manner, designed to overcome the key bottlenecks of prior methods. Rather than explicitly computing invariant sets, PSP-RL leverages a learned backup policy to forward-integrate the system dynamics, generating an implicit control-invariant set online. In the first phase, the backup policy is trained with our proposed safe-arrival value function, which characterizes the optimal backup policy for invariant-set construction. In the second phase, an RL policy is trained end-to-end through a differentiable projection layer that strictly enforces the safety guarantees induced by the learned backup policy. By maximizing the volume of the implicit control-invariant set in the first phase, the resulting PSP policy from the second phase is performant and scalable, while maintaining provable safety. Crucially, PSP-RL imposes no restrictions on the underlying RL algorithm and can be plugged into any existing training pipeline. We establish theoretical guarantees for the proposed framework and evaluate it on robotic control tasks with state dimensions up to 10, a regime in which prior provably safe RL methods struggle or become impractical.
PaperID: 1518, Poster
Abstract: Generating antibody sequences is challenging because they combine conserved framework regions with hypervariable loops. Latent diffusion is attractive for this task since it enables flexible conditioning and bidirectional generation. But standard approaches fail. Global noise schedules treat all positions equally, so models learn the predictable frameworks well while the diverse loops remain poorly captured. We address this by learning a latent space that redistributes information evenly, allowing standard diffusion to succeed where it previously failed. On organism-conditioned generation across six species, our approach achieves 10× lower Fréchet Distance than latent diffusion without redistribution. It supports chain-type control, loop infilling, and paired-chain generation. Validation across five protein encoders confirms the method is encoder-agnostic. These results establish latent diffusion as a practical tool for antibody sequence design.
PaperID: 1519, Poster
Authors: Kuo Shi
Abstract: Out-of-distribution (OOD) generalization aims to improve a model’s generalization capacity using only source data, thereby ensuring reliable performance on unseen domains. Cutout, a well-established method that enhances model generalization by randomly masking out square regions of training samples, has shown great success in standard supervised settings. Despite its seemingly strong con- nection to OOD generalization, we observe that simply combining Cutout with existing OOD methods yields only marginal benefits or even degrades performance. Motivated by this empirical finding, we seek a way to make OOD generalization benefit from Cutout. Our method is inspired by reinterpreting the state-of-the-art hyperspher- ical prototype learning as a mutual information (MI) framework. Within this framework, we observe that naively applying Cutout to existing OOD methods is equivalent to estimating MI from a single masked subview, which leaves much of Cutout’s potential untapped. Instead, we apply the MI chain rule to split the original objective into two smaller estimation problems: one encourages the Cutout subview to align closely with the label, while the other drives the full image to capture the residual information unavailable in that subview. This divide-and-conquer strategy fully exploits the Cutout-generated subview while remaining computationally light- weight. Empirically, we demonstrate that our method outperforms competitive baselines on a broad range of OOD benchmarks and achieves superior performance.
Authors:
Yikun Zhang, Steven Wilkins-Reeves, Wesley Lee, Aude HofleitnerAbstract: We introduce a transfer learning framework for regression that leverages heterogeneous source domains to improve predictive performance in a data-scarce target domain. Our approach learns a conditional generative model separately for each source domain and calibrates the generated responses to the target domain via conditional quantile matching. This distributional alignment step corrects general discrepancies between source and target domains without imposing restrictive assumptions such as covariate or label shift. The resulting framework provides a principled and flexible approach to high-quality data augmentation for downstream learning tasks in the target domain. From a theoretical perspective, we show that an empirical risk minimizer (ERM) trained on the augmented dataset achieves a tighter excess risk bound than the target-only ERM under mild conditions. In particular, we establish new convergence rates for the quantile matching estimator that governs the transfer bias-variance tradeoff. From a practical perspective, extensive simulations and real data applications demonstrate that the proposed method consistently improves prediction accuracy over target-only learning and competing transfer learning methods.
PaperID: 1521, Poster
Authors: Harsh Udai, Konda Reddy Mopuri
Abstract: Improving compositional understanding in CLIP-style vision-language models remains difficult, models with strong coarse alignment often fail on attribute binding, relations, and multi-object semantics. A common remedy is text-side hard negatives (negations, perturbations, or mined confusions), but concentrating difficulty on the text branch can skew the training signal, exacerbate modality asymmetry, and distort joint-space geometry (e.g., increased hubness and reduced mutual reciprocity), even when Recall@K improves. We propose LACE (Latent Alignment via Counterfactual Embeddings), a lightweight approach that injects structured, semantics-preserving latent edits into the image and text embedding streams without pixel-space editing or external generators. LACE synthesizes counterfactual image and text embeddings by attenuating, swapping, or transplanting a small set of factors tied to target objects, attributes, or relations, producing image-side and text-side hard negatives that rebalance supervision while preserving global alignment. Across compositional and retrieval benchmarks, LACE improves robustness to binding and relational confusions and yields a healthier embedding geometry with improved reciprocity and competitive hubness/local-stability trade-offs.
PaperID: 1522, Poster
Abstract: Space-time self-similarity (STSS), which captures visual correspondences across frames, provides an effective way to represent temporal dynamics for video understanding.In this work, we explore higher-order STSS and demonstrate how STSSs at different orders reveal distinct aspects of these dynamics. We then introduce the Multi-Order Self-Similarity (MOSS) module, a lightweight neural module designed to learn and integrate multi-order STSS features. It can be applied to a wide range of motion-centric tasks with only marginal computational cost and memory overhead. Extensive experiments on video action recognition, motion-centric video VQA, and robotic tasks consistently demonstrate substantial improvements, validating the broad applicability of MOSS as a general motion modeling module. The source code and checkpoints will be publicly available.
PaperID: 1523, Poster
Abstract: Probabilistic neurosymbolic methods rely on weighted model counting (WMC) to combine neural predictors with symbolic constraints. Exact computation of the WMC, a #P-hard problem, typically scales poorly, so many methods resort to approximations. By unifying existing approaches under a three-step bottom-up approximation budget is best allocated. Inspired by tensor networks, we instantiate two steps with tensor train decompositions: we derive new pipelines that exactly multiply approximated factors and recompress their products via SVD- or interpolation-based schemes. On Sudokus of increasing size, these pipelines produce WMC estimates that are many orders-of-magnitude more accurate than prior methods. Yet, when it comes to neurosymbolic learning, even crude approximations reach competitive accuracy, suggesting that approximation quality matters far more for inference fidelity than for downstream learning. Our pipelines open significant design space for effective approximate compilation of challenging probabilistic reasoning tasks.
Authors: Kaito Baba, Evripidis Bampis, Giorgos Mitropoulos
Abstract: Recently, Antoniadis et al. (ICLR 2025) proposed a framework for incorporating predictions to approximate NP-hard selection problems. Despite its simplicity, this approach tightly matches theoretical lower bounds, making its generalization highly compelling. We address an open question raised in the work of Antoniadis et al., concerning the extension of this approach to other important problems outside the class of selection problems, such as scheduling. We develop a learning-augmented algorithm for the makespan minimization problem on unrelated machines, denoted by R||C_\max. By using predictions of heavy job assignments, we achieve a polynomial-time (1+\varepsilon)-approximation for accurate predictions that smoothly degrades to a worst-case 2-approximation as the error increases. We conclude our work with an empirical analysis of our method.
Abstract: Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive models, offering competitive performance while naturally supporting parallel decoding. However, as dLLMs are increasingly integrated with Mixture-of-Experts (MoE) architectures to scale model capacity, a fundamental mismatch arises between block parallel decoding and token-level expert selection. Specifically, each dLLM forward pass processes multiple tokens with bidirectional dependencies, whereas conventional MoE layers route each token independently. This mismatch substantially increases the number of uniquely activated experts, making inference increasingly memory-bound. To address this, we propose dMoE, a simple yet effective block-level MoE framework. The central idea of dMoE is to aggregate token-level expert distributions within each block into a unified block-level expert distribution, which is then used to guide expert routing in a more coherent manner. In this way, dMoE substantially reduces the number of uniquely activated experts during inference without sacrificing performance, thereby mitigating the memory-bound bottleneck. Extensive experiments across a variety of benchmarks demonstrate the effectiveness of dMoE. On average, dMoE reduces the number of uniquely activated experts from 69.5 to 14.6 while retaining 99.11% of the original performance. Meanwhile, it reduces memory usage by 76.64% to 79.84% and achieves 1.14× to 1.66× end-to-end latency speedup.
Authors: Georgios Amanatidis, Giulio Giaconi, Evangelos Markakis, Nicos Protopapas
Abstract: We study the _online fair division_ of indivisible _mixed manna_ among agents with additive valuation functions. Under the standard online model, at each time step an indivisible item arrives; each agent may assign it a positive, negative, or zero value, and it must be irrevocably allocated, before the arrival of the next item. At the same time, we also wish to maintain some fairness guarantee, and in this work we focus on _envy-freeness_ (EF) and one of its most prominent relaxations, _envy-freeness up to one item_ (EF1). Given the strong negative and the scarce positive results for this problem without additional assumptions, we augment our algorithms with _buffers_ that can store and rearrange a limited number of items. This setting interpolates naturally between the fully online case (no buffer) and the fully offline case (a buffer large enough to hold all items). We show that algorithms equipped with reasonably sized buffers can achieve strong guarantees for personalized k-value instances, i.e., instances in which each agent assigns at most k distinct values to items. In particular, we construct allocations that are EF1 at every time step and EF at most time steps, using a buffer of size linear in k and in the number of agents. Our approach relies on novel combinatorial arguments and on constructing a sequence of envy-free matchings that allocates most items. Finally, we extend our results to general additive valuation functions, with a dependence on the largest per-agent ratio between two values of the same sign, and we also identify limitations of our approach via impossibility results on the use of buffers with smaller size.
PaperID: 1527, Poster
Abstract: RL+Search is the computational backbone of modern AI systems for imperfect-information extensive-form games. The Recursive Belief-based Learning (ReBeL) framework already mitigates one source of training fragility by sampling the pivot iteration uniformly from the iteration budget; we identify a second, hitherto unaddressed source: the budget itself is held fixed across all subgames, creating a rigid coupling between solver maturity and outer-loop learning that propagates variance through the recursive self-play and produces noisy, seed-dependent training trajectories. We propose Stochastic Horizon Annealing (SHA), which randomizes only this remaining static degree of freedom: at each subgame the total CFR iteration count is drawn from a Gaussian centered at the expected budget. Combined with ReBeL's existing uniform pivot, the procedure smooths the learning objective across a continuum of solving depths and behaves as an implicit ensemble at zero additional cost. SHA is a one-line change to any ReBeL-style trainer. We give a complete theoretical treatment with three results, all proved in the main text: (i) SHA preserves the O(N^-1/2) convergence rate of linear CFR; (ii) the per-subgame target variance is bounded by C_\beta^2\sigma^2/(4N^3), vanishing rapidly with the budget; and (iii) the recursive variance of the value-network targets scales as \Theta(D) for SHA versus \Theta(D^2) for the fixed-horizon ReBeL, where D is the recursion depth. On Liar's Dice 1×4f, 1×5f, and 1×6f (30 seeds each) at the ReBeL reference setting of depth-2 subgames and 1024 search iterations, SHA reduces mean final exploitability by 9--18% and shrinks the 95% confidence interval across seeds by 34--63%, with larger games benefiting more as predicted by the theory, establishing stability as a first-class, theoretically supported contribution for RL+Search on imperfect-information games.
PaperID: 1528, Poster
Abstract: Plurality-voting ensembles aggregate the predictions of multiple base classifiers and output the most voted class, with ties typically broken at random. Despite the simple aggregation rule, attributing the impact of each individual base classifier to the ensemble's predictions is far from straightforward. That is, a single base classifier's impact depends not only on its own predictions but also on how it interacts with the votes of all other models in the ensemble, for example, in resolving ties. We address the attribution problem using established tools from cooperative game theory, studying the Banzhaf and Shapley power indices for new games that capture plurality voting and random tie-breaking, in both binary and multi-class classification settings. Computing power indices is, in general, exponential in the number of players. Surprisingly, our methods, collectively named ENPOWER, enable the for the multi-class setting. The algorithms in ENPOWER rely on new combinatorial formulations of the Shapley and Banzhaf indices, yielding closed-form expressions for the binary case and ordinary generating function expressions for the multi-class case. We validate empirically the methods in ENPOWER and demonstrate that they are highly efficient, providing attribution values not captured by existing methods. We further apply the techniques in ENPOWER to: ensemble pruning and prompt-based ensembling of vision-language models.
PaperID: 1529, Poster
Abstract: Progress in mathematics often depends not only on proving that a result is true, but on finding a proof that is simple enough to understand, verify, and adapt. Automated systems can now independently discover proofs, but the certificates they produce tend to be dense and difficult to interpret, even when simpler arguments exist. Recent work on performance estimation problems has shown that proofs of convergence for first-order optimization methods can be discovered by searching over a structured space of Lagrangian dual certificates. This paper studies the simplification of such proofs as an optimization problem. Starting from dual certificates, we develop post-processing procedures using tools from sparse optimization and statistical learning. We measure complexity through features such as active hypotheses and residual structure, and introduce methods based on exhaustive sparsification, (weighted) \ell_1 heuristics, and semidefinite programming (SDP) formulations for discovering simple proofs and intermediate lemmas. Examples on gradient descent, proximal methods, and fast-gradient methods show that these procedures can autonomously prune redundant inequalities, reveal structured proof patterns, and, in the proximal setting, recover standard intermediate lemmas that lead to streamlined proofs. By distilling dense machine-generated certificates into compact proof structures, this workflow acts as a pre-processing step for the final proof, reducing the complexity that must be managed during human interpretation, reuse, and formalization.
Authors: Qing Jin, Chaoyang Wang
Abstract: Recent advances in diffusion and flow matching models have highlighted a shift in the preferred prediction target---moving from noise (\varepsilon) and velocity (v) to direct data (x) prediction---particularly in high-dimensional settings. However, a formal explanation of why the optimal target depends on the specific properties of the data remains elusive. In this work, we provide a theoretical framework based on a generalized prediction formulation u=kx-(1-k)n, where x represents clean data and n denotes noise. This formulation accommodates arbitrary targets, of which \varepsilon-, v-, and x-prediction are special cases. We derive the analytical relationship between data's geometry and the optimal prediction target, offering a rigorous justification for why x-prediction becomes superior when the ambient dimension significantly exceeds the data's intrinsic dimension. Furthermore, while our theory identifies dimensionality as the governing factor for the optimal prediction target, the intrinsic dimension of manifold-bound data is typically intractable to estimate in practice. To bridge this gap, we propose k-Diff, a framework that learns the optimal parameter k directly from data, bypassing the need for explicit dimension estimation. Extensive experiments in both latent-space and pixel-space image generation demonstrate that k-Diff consistently outperforms fixed-target baselines across varying architectures and data scales, providing a principled and automated approach to enhancing generative performance.
PaperID: 1531, Poster
Abstract: Retrieval-augmented generation (RAG) divides into two camps failing in complementary ways on broad, long-tailed questions: iterative agentic RAG re-queries a biased distribution and drifts on prominent aspects, while structural RAG commits to a hierarchy fixed upfront and cannot admit aspects it initially missed. We propose GBT-RAG, which unifies the two by transposing gradient boosting into the text domain: each round is a query decomposition tree fitted to its predecessors' residual gap. This raises two challenges mirroring the pillars of boosting---natural-language critiques are not quantifiable, and successive rounds drift rather than fit the residual. We address the first with a text gradient, a structured residual in the retriever's embedding space whose support localizes missingness to specific leaves; we address the second by growing each tree only over under-covered leaves and admitting each rewrite through a monotone-improvement gate. Across standard benchmarks, GBT-RAG consistently outperforms the strongest iterative and structural baselines, with the largest gains on the long-tailed regime that motivated the design.
PaperID: 1532, Poster
Abstract: The canonical orientation of a 3D object is often not unique but inherently ambiguous: many objects admit multiple equally valid canonical frames that no single-output regressor can faithfully represent. Rotational symmetry is the most structured manifestation of this phenomenon, where the set of valid orientations forms an orbit of the symmetry group. Existing methods fall short in one of three ways: they predict only a single orientation, handle symmetry only under pre-defined group assumptions, or resort to test-time augmentation to approximate the full set of valid orientations. We instead introduce a general framework for ambiguity-aware orientation estimation that is at once SO(3)-equivariant by construction, distributional in its output, and free of any restriction on the symmetry group. It effectively predicts multiple valid orientations as a continuous multi-modal distribution over SO(3) via a truncated Wigner-D expansion, without any symmetry-group assumption. The proposed method outperforms recent equivariant regressors and fixed-quotient classifiers on the ShapeNet orientation benchmark, particularly on categories with continuous rotational symmetry, showing that supervision on discrete symmetry labels generalises to continuous bands.
Abstract: Large-scale approximate nearest neighbor search commonly relies on partitions for indexing: database vectors are partitioned into clusters, and for each query a probing function selects the clusters to be scanned. The query probing function and the database partition are rarely treated as separate entities: most techniques assign queries with the same assignment function as the database vectors, which is suboptimal especially when database and query distributions differ. This paper introduces CwA (Cluster with Auctions), which addresses this limitation by jointly learning a balanced database partition and a neural probing function. CwA optimizes search performance directly for the query distribution. It minimizes its objective by alternating two steps: (i) gradient descent on the neural network of the probing function, and (ii) a large-scale combinatorial optimization of the cluster assignment for the database vectors. We solve the latter with a parallelizable auction algorithm that balances the partition by design. To further scale CwA, we extend the method to a Cartesian product of clusters that increases the partition’s granularity. When database and query distributions differ, CwA achieves up to 4.7× throughput over the state of the art at equal recall. In the in-distribution (ID) setting, even a simple linear probing function trained with CwA outperforms competing deep neural methods.
PaperID: 1534, Poster
Authors:
Ruizhi Yuan, Zeqiu Yu, Wei Gao, Wei Chen, Xiwei Tang, Lu TangAbstract: Many multimodal alignment problems require mapping a source modality into a fixed space defined by a target modality, where the target space is not merely an intermediate embedding but is used directly for downstream analysis. In this setting, successful alignment should preserve both the global structure of the target distribution and the instance-level correspondence observed in paired source--target training data. Existing objectives typically address only part of this goal: contrastive losses encourage relative similarity, distribution-matching losses improve marginal agreement, and pointwise regression fits each source sample to its paired target without explicitly modeling target-space path geometry. We propose FlowAlign, a flow-guided target-space alignment method that repurposes conditional flow matching as a training-time alignment signal rather than an inference-time transport model. FlowAlign uses the learned flow field to guide the alignment of source-side representations with the fixed target space, encouraging them to preserve paired correspondence while remaining compatible with the global target-space structure. We further provide theoretical analysis showing how FlowAlign helps improve paired correspondence by controlling endpoint accuracy and capturing local target-space geometry. Experiments on a controlled synthetic task and paired single-cell multiome benchmarks show that FlowAlign improves alignment performance while preserving target-space structure for downstream analysis.
Authors: David R Lolck, Mikkel Thorup, Shuyi Yan
Abstract: Correlation clustering seeks a partitioning of the vertices of a graph that minimizes the number of disagreements: edges between different clusters and non-edges within the same cluster. The classic pivot-based correlation clustering algorithm of Ailon, Charikar, and Chawla [JACM'08] achieves a factor-3 approximation. However, known worst-case examples attaining this bound rely on highly structured instances with exceptionally strong clusters, such as complete graphs minus a matching. In this work, we show that such worst-case behavior is inherently tied to these highly clique-like substructures. Building on the notion of atoms introduced by Cohen-Addad, Lattanzi, Mitrovic, Norouzi-Fard, Parotsidis, and Tarnawski [FOCS'23] — subsets of vertices that cannot be separated in any optimal clustering — we prove that the approximation ratio of 3 can only be realized when the pivot algorithm makes errors on such atoms. Motivated by this insight, we propose a modified pivot algorithm that removes atoms prior to each pivot step. We prove that this yields an improved approximation ratio of 2.9995 in expectation. Additionally, we evaluate our algorithm on synthetic datasets with varying amount of noise, demonstrating consistent empirical improvements over both the classic pivot algorithm and existing methods for identifying atoms.
PaperID: 1536, Poster
Abstract: Post-training of language models is commonly framed as a sample-score-update loop implemented by gradient descent. A recent line of work, exemplified by RandOpt, relocates this loop to weight space, sampling Gaussian perturbations around a pretrained model and ensembling the top-K rewarded specialists at inference. While competitive with PPO and GRPO under matched training compute, this prediction-level ensemble incurs K forward passes per test example and does not extend cleanly to free-form generation. We ask whether the rewarded population can instead be folded into a single deployable model, replacing the inference-time ensemble with one consolidated update. A split-half analysis over 25 model-task pairs reveals reproducible low-rank structure in every case. We turn this geometry into CoRP (Consolidating Rewarded Perturbations), a gradient-free operator that combines reward-weighted aggregation, compatibility-aware reweighting, and a held-out validation gate, with no gradient flowing through the language model. Across five language models from 0.5B to 8B and five tasks covering math, code, and creative writing, CoRP improves the base model by 8.1 points on average. Using one tenth of RandOpt's perturbation budget, CoRP exceeds single-inference RandOpt by 6.5 points and recovers more than half of the gain of the 50-pass majority-vote ensemble, at one forward pass per test example.
Abstract: Direct Preference Optimization (DPO) has recently been applied to medical image segmentation, since comparing candidate masks is substantially cheaper than producing new dense pixel-level annotations, making pairwise preference signals an appealing source of supervision. However, existing work derives preferences from ground-truth masks, leaving open whether the approach is sound under the imperfect feedback available in practice. In our analysis of this reformulation, we identify that candidate masks tend to agree over most of the image and differ only in localized regions, yet standard DPO averages the preference signal over the full spatial domain. This couples update strength to disagreement area, diluting small correct refinements and amplifying large misranked differences. Crucially, this bias even degrades performance even under oracle preferences derived from ground truth. We propose Region-Normalized DPO (RN-DPO), which normalizes the likelihood ratio over the disagreement region between candidate masks, removing this coupling. We further provide a systematic empirical study of preference-based segmentation fine-tuning under controlled noisy judges, analyzing the effects of mining strategies, judge types, and reliability regimes. Across two medical segmentation benchmarks, multiple judge configurations, and two backbone architectures, RN-DPO consistently improves over vanilla DPO and robust DPO variants, with gains persisting under noiseless oracle preferences.
Abstract: Test-time alignment aims to adapt large language models (LLMs) to runtime objectives without updating their parameters. This is especially useful when the reward signal is user-specific, time-varying, or available only through a black-box evaluator. We propose a minimally invasive pre-logit steering method that optimizes additive interventions to the hidden states of a frozen LLM in order to improve reward while preserving the base model's behavior. Rather than penalizing the Euclidean norm of the steering vector, we derive an effort penalty from the local KL geometry of the induced token distribution. Specifically, we show that the per-token KL divergence between the steered and reference policies admits a second-order expansion given by a Fisher-quadratic form in the steering vector. This regularizer has an analytic gradient computable through matrix-vector products with the frozen language-model head, making it as inexpensive as a standard quadratic penalty while retaining a principled KL interpretation. We further decompose the sequence-level KL gradient into an analytic Fisher component and a trajectory-dependent score-function component, and introduce a hierarchy of Fisher surrogates with provable first-order equivalence in the small-steering regime. The resulting algorithm provides a training-free, reward-driven test-time alignment procedure that improves reward while controlling deviation from the base policy.
PaperID: 1539, Poster
Abstract: Deep neural networks often struggle to provide reliable confidence estimates, making post-hoc calibration essential for reliable decision-making. Existing post-hoc methods largely rely on score-based transformations, which improve marginal calibration but often fail to capture spatially varying miscalibration in local regions of the learned feature space. To address this gap, we propose (LAPE), a post-hoc residual calibration framework. Specifically, LAPE treats the global calibrator output as a stable anchor and refines the individual prediction by leveraging non-parametric local evidence from neighboring calibration samples. As a plug-and-play refinement module, LAPE effectively utilizes the information about feature-space geometry and can reuse the calibration data that are used to construct the global calibrator. Furthermore, we extend LAPE to the setting with multi-view inputs through (CFAF). Extensive experiments demonstrate that the proposed methods can improve the calibration quality of various existing baselines and remain efficient and robust across different settings.
PaperID: 1540, Poster
Abstract: Language models fine-tuned with reinforcement learning frequently learn to exploit their reward source rather than solve the underlying task. Existing fixes are typically reactive: detect a specific exploit, then patch the reward. We instead ask whether reward hacking can be prevented by constraining how the model represents the task. We introduce prompt KL regularization, a single auxiliary loss that adds a KL penalty between the actor's and reference model's per-token distributions on the prompt only, leaving the response distribution free. Across three reward hacking tasks of different sizes, prompt KL regularization maintains low hack rates (under 3%) when standard training yields up to 100% hack rates. Additionally, it matches or improves ground-truth accuracy, with no changes to the reward or environment. We aim to understand why prompt KL regularization is effective. Two independent mechanistic studies suggest that changes in prompt representations are key to models learning reward hacking behavior. First, we show that prompt-space activation steering Pareto-dominates response-space steering on hack rate and accuracy. Then, we demonstrate that swapping prompt activations between models can transfer and reduce reward hacking. Together, these results suggest that the way models internally represent the task is central to reward hacking, and that this representation can be directly constrained to prevent it.
Authors: Olga Sorokoletova, Francesco Giarrusso, Giacomo De Luca, Piercosma Bisconti, Matteo Prandi, Federico Pierucci, Marcello Galisai, Vincenzo Suriani, Daniele Nardi
Abstract: Large language models are increasingly deployed in safety-critical and user-facing applications, where their ability to resist harmful instructions is essential. Although post-training aims to make models robust against many jailbreak strategies, recent evidence shows that stylistic reformulations, such as poetic transformation, can still bypass safety mechanisms with alarming effectiveness. This raises a central question: why do literary jailbreaks succeed? In this work, we investigate whether their effectiveness depends on specific poetic devices, on a failure to recognize literary formatting, or on deeper changes in how models process stylistically irregular prompts. We address this problem through an interpretability analysis of attention patterns. Our analysis proceeds in three steps: we perform input-level ablation studies to assess the contribution of individual and combinations of rhetorical devices; we construct a novel interpretable vector representation of attention maps; we cluster these representations and train linear probes to predict both safety outcomes and literary format. Our results show that models distinguish poetic from prose formats with high accuracy, yet struggle to predict jailbreak success within each format. Clustering further reveals clear separation by literary format, but not by safety label. These findings indicate that jailbreak success is not caused by a failure to recognize poetic formatting; rather, poetic prompts induce distinct processing patterns that remain largely independent of harmful-content detection. Overall, literary jailbreaks appear to misalign large language models not through any single poetic device, but through accumulated stylistic and structural irregularities that alter prompt processing and avoid lexical triggers considered during post-training. This suggests that robustness requires safety mechanisms that account for style-induced shifts in model behavior. We use Qwen3-14B as a representative open-weight case study for all reported experiments.
PaperID: 1542, Poster
Abstract: We study stochastic zeroth-order (ZO) optimization of smooth nonconvex objectives under heavy-tailed sample-gradient noise, a regime motivated by empirical evidence that gradient noise in modern machine learning can violate the bounded-variance assumption underlying classical ZO theory. While the first-order literature has established optimal rates under bounded p-th moment noise for p \in (1,2], analogous high-probability guarantees for nonconvex ZO optimization remain largely unexplored. The ZO setting is not a direct corollary of first-order theory: first-order methods update using \nabla F(x;\xi), the same object on which the noise assumption is imposed, whereas derivative-free methods only observe noisy function values and construct directional finite-difference estimates. Thus, weak-L_p control of \nabla F(x;\xi)-\nabla f(x) must first be transferred to scalar two-point directional estimates, a concentration step absent from standard first-order analyses. We propose the Robust Scalar-Clipped Zeroth-Order method (RSC-ZO), a two-point method that clips each scalar directional derivative before aggregation. Under sample-wise smoothness and a weak-L_p tail bound on the sample-gradient noise, RSC-ZO finds an \varepsilon-stationary point with high probability using \widetildeO\\left( \fracd^\fracp2(p-1) \varepsilon^\frac3p-2p-1 \right) noisy function evaluations, matching the optimal first-order \varepsilon-dependence. At p=2, our bound becomes \widetildeO(d\varepsilon^-4), matching the classical dimension-accuracy dependence for stochastic ZO methods under bounded-variance noise, which is typically established only in expectation. In contrast, our guarantee holds with high probability and under the strictly weaker weak-L_2 tail condition, which can allow infinite gradient-noise variance. We further extend the analysis to a momentum variant and quantify the resulting batch-size/stepsize tradeoff.
Authors:
Eloïse Touron, Pedro Rodrigues, Julyan Arbel, Nelle Varoquaux, Michael ArbelAbstract: Inference from interaction maps, such as centromere identification from genome-wide chromosome conformation capture techniques --notably Hi-C-- can be formulated as a generic inverse problem: infer a set of parameters given a map summarizing pairwise interactions between entities through blocks of variable numbers and sizes. In this work, we introduce a data-driven approach that leverages shared structure between these maps, such as global alignment between localized patterns, while handling the variability in number and size of entities arising in real-world data. Our approach relies on a transformer architecture capable of handling such variability and a custom simulator to generate abundant, yet computationally cheap synthetic data for training. Applied to the problem of centromere localization, the method accurately recover their genomic positions across a wide range of species of various genome sizes.
Abstract: Language model (LM) alignment improves model outputs to reflect human preferences while preserving the capabilities of the base model. The most common alignment approaches are (i) reinforcement learning, which maximizes the expected reward under a KL-divergence constraint, and (ii) best-of-N alignment, which selects the highest-reward output among N independent samples. Despite their widespread use, the fundamental limits of reward improvement under a KL budget remain poorly understood. We characterize the information-theoretic limits of KL-regularized alignment by deriving the maximum achievable expected reward gain for a fixed KL-divergence budget. Our first result provides a closed-form expression for the optimal reward improvement, governed by a Jeffreys divergence term rather than the \sqrt\textttKL used in prior analyses. We further reformulate this expression as a covariance under the base model, yielding a practical estimator that predicts achievable alignment gains from base model samples alone. We extend our analysis to the proxy reward setting, showing that the gap between ideal and proxy alignment (reward hacking) grows with the KL penalty factor and magnitude of reward error. We then prove that reward ensembling mitigates reward hacking, providing a theoretical justification for this technique used in practice. Empirically, we compute the KL–reward Pareto frontier for two alignment tasks for LMs, safety and summarization, and show that best-of-N closely approaches the theoretical limit, while PPO and GRPO remain substantially suboptimal. Our theoretical results shed light on several empirically observed phenomena in the alignment literature and suggest that algorithmic improvements are needed to achieve optimal alignment without high inference costs.
Abstract: We study the realization map of deep ReLU networks, focusing on when a function determines its parameters up to scaling and permutation. To analyze hidden redundancies beyond these standard symmetries, we introduce a framework based on weighted polyhedral complexes. Our main result shows that for every architecture whose input and hidden layers have width at least two, there exists an open set of identifiable parameters. This implies that the functional dimension of every such architecture is exactly the number of parameters minus the number of hidden neurons. We further show that minimal functional representations can still have non-trivial parameter redundancies. Finally, we establish a generic depth hierarchy, whereby for an open set of parameters the realized function cannot be represented generically by any shallower network.
Authors:
Christian Internò, Alexander Pondaven, Habon Issa, Fabio Pizzati, Francesco Pinto, Markus Olhofer, Ivan Laptev, Philip Torr, Eero Simoncelli, Barbara Hammer, David KlindtAbstract: Whether a video is physically plausible is a question current generative models answer poorly and current evaluators answer expensively, through physics-targeted training, billion-parameter video world models, or multimodal-language model judges. We ask whether frozen images vision encoders, with no video or physics supervision, already contain a useful signal for the same question. Across four frozen backbones, self-supervised vision transformers (DINOv2, DINOv3) and biologically-inspired models of the ventral stream (CORnet-S, VOneNet), a small set of geometric signals (curvature, speed variation, acceleration, prediction residual) on per-frame feature trajectories separates plausible from physics violated videos with no retraining. We propose GEOPHYS, a training-free framework that 11 reaches 97.7% on LikePhys and 93.3% on IntPhys2 accuracy for physics-violation detection, surpassing V-JEPA 2, GPT-4o, Gemini, and twelve modern video diffusion models. The same signals track human EEG responses to object-permanence violations and scale with object number. Deployed unchanged as a best-of-N reward during video generation, GEOPHYS lifts MAGI-1 4.5B from 53.8% to 64.7% PhysicsIQ score at 5.8× lower wall-clock and 7.8× lower memory than V-JEPA 2 world-model rewards. A useful proxy for physical plausibility is already encoded as a geometric regularity of natural-video statistics in vision features.
Abstract: Continuous diffusion is a natural framework for non-autoregressive generation but has consistently underperformed masked discrete diffusion (MDM) on discrete sequence generation. We argue this is not a limitation of continuity itself, but of how previous models parameterize the denoiser as a family of timestep-indexed regimes. We introduce \emphDiscrete Stochastic Localization (DSL), a continuous-state framework with unit-sphere token embeddings whose Bayes-optimal denoiser is invariant to the nominal signal-to-noise ratio (SNR). One trained network then supports an entire family of per-token SNR paths, with masked diffusion as one special case. Fine-tuning a pretrained MDLM checkpoint with DSL substantially improves distributional faithfulness (MAUVE) on OpenWebText across all step budgets from T=128 to T=1024, and the same checkpoint supports random-order autoregressive sampling and a hybrid continuous-then-discrete sampler at as few as T=48 total steps---without distillation or retraining. On Text8, DSL gives the first continuous-state NLL estimate approaching the MDM range. Code is anonymously available at \urlhttps://anonymous.4open.science/r/DSL_anonymous-94B5/.
Abstract: Test-time reinforcement learning (TTRL) has emerged as a promising paradigm for adapting large reasoning models (LRMs) on unlabeled test inputs, using self-consensus rewards derived from majority voting over sampled rollouts. However, majority voting can mistake popularity for correctness: a spurious yet high-frequency unverified consensus may become a biased reward signal, causing test-time RL to reinforce frequent but wrong answers and collapse into an incorrect mode. We address this false-popular failure mode with T^3RL (Tool-Verification for Test-Time Reinforcement Learning), a verification-aware test-time RL framework. T^3RL grounds pseudo-label construction in external tool evidence. Concretely, a verifier uses external tool evidence, such as code execution, to upweight verified rollouts during verification-aware voting, producing more reliable pseudo-labels for training. Across mathematical reasoning benchmarks of varying difficulty, including MATH-500, AMC, and AIME 2024, and diverse backbone families, T^3RL significantly improves over TTRL, with stronger performance on harder problems. More broadly, T^3RL is positioned as a verified online data synthesizer, highlighting the role of test-time tool verification in reliable online adaptation.
Authors: Nils Loose, Jonas Sander, Felix Mächtle, Thomas Eisenbarth
Abstract: Large language models (LLMs) are increasingly deployed in sensitive settings such as software engineering, where their outputs directly shape downstream artifacts. Recent work has shown that an identical model can produce measurably different outputs depending on the deployment platform, a consequence of non-associative floating-point arithmetic and divergent kernel implementations. We study the security implications of this platform-dependent variability and uncover a novel attack surface on LLM deployments. We introduce FloatDoor, the first input-independent, platform-triggered backdoor attack against generative LLMs. The compromised model exhibits adversary-chosen behavior when served on a target platform and is otherwise benign. FloatDoor is realized through two lightweight LoRA adapters, one that amplifies inter-platform numerical divergence and one that binds the resulting platform signature to a malicious downstream task, while leaving aggregate model utility largely intact. FloatDoor exploits a pronounced time-of-check, time-of-use gap between model auditing and serving. We demonstrate FloatDoor on Qwen3-4B across a broad range of deployment targets, including NVIDIA GPUs, Google TPUs, AWS Graviton, and Alibaba Yitian-710. As a final case study, we show that FloatDoor reliably induces exploitable code vulnerabilities on a chosen target platform. Our results establish a new class of attacks on LLM deployments and underscore the pressing need for trusted model supply chains in sensitive, LLM-powered applications.
Abstract: Agent memory is typically constructed either offline from curated demonstrations or online from post-deployment interactions. However, regardless of how it is built, an agent faces a cold-start gap when first introduced to a new environment without any task-specific experience available. In this paper, we study pre-task memory construction: whether an agent can build procedural memory before observing any target-environment tasks, using only self-generated synthetic practice. Yet, synthetic interaction alone is insufficient, as without controlling what to practice and what to store, synthetic tasks become redundant, infeasible, and ultimately uninformative, and memory further degrades quickly due to unfiltered trajectories. To overcome this, we present Preping, a proposer-guided memory construction framework. At its core is proposer memory, a structured control state that shapes future practice. A Proposer generates synthetic tasks conditioned on this state, a Solver executes them, and a Validator determines which trajectories are eligible for memory insertion while also providing feedback to guide future proposals. Experiments on AppWorld, BFCL v3, and MCP-Universe show that Preping substantially improves over a no-memory baseline and achieves performance competitive with strong playbook-based methods built from offline or online experience, with deployment cost 2.99× lower on AppWorld and 2.23× lower on BFCL v3 than online memory construction. Further analyses reveal that the main benefit does not come from synthetic volume alone, but from proposer-side control over feasibility, redundancy, and coverage, combined with selective memory updates.
Abstract: Multicalibration requires predicted scores to agree with label probabilities across rich families of subgroups and score-dependent tests, but existing methods require clean input--label pairs for evaluation and post-processing. This assumption fails in weakly supervised learning (WSL) regimes---including positive-unlabeled, unlabeled-unlabeled, and positive-confidence learning---where clean labels are costly or unavailable even though reliable uncertainty estimates may be crucial. We address this gap by developing estimators of multicalibration error and post-hoc correction methods for WSL settings in which clean input--label pairs are unavailable. We propose a unified framework for estimating and correcting multicalibration under weak supervision by combining contamination-matrix risk rewrites with witness-based calibration constraints, yielding corrected multicalibration moments with finite-sample guarantees. We further propose weak-label multicalibration boost (WLMC), a generic post-hoc recalibration algorithm under weak supervision. Finally, we conduct experiments across multiple weak-supervision settings to evaluate multicalibration behavior and deepen empirical understanding of uncertainty estimation under weak supervision.
Abstract: No existing nested Archimedean copula tool handles all three of (a) arbitrary per-variable (right-)censoring in survival analysis, (b) arbitrary nesting trees, and (c) exact parameter gradients. Existing implementations handle only bivariate problems, low dimensional (i.e., d \leq 10) cases, two layers of nesting, or only hand-derived copula nestings. We present acopula, a JAX-native framework that, given any Archimedean generator—classical or neural—evaluates exact nested-copula likelihoods and parameter gradients under arbitrary censoring masks in polynomial time. The mechanism is polynomial powering of Taylor-mode automatic differentiation output, which replaces per-family hand-derived partial Bell polynomial tables with a single differentiable computation that any user-defined generator can drive. We conduct extensive simulations to verify the correctness of acopula. We then demonstrate (a) per-variable censoring on 85,229 MIMIC-IV ICU admissions in high dimensions with d=53, fit by both classical Archimedean families and nested neural Archimedean copulas; (b) an 11-sector hierarchical model on S&P 500 daily returns at d=98; (c) family-agnostic censored MLE across ten families, five of them with no prior implementation, on a retinopathy study; and (d) a ~650× per-density speedup over R's nacLL at d=35, scaling quadratically to d=8,000.
PaperID: 1553, Poster
Abstract: Long video out-painting is an important yet underexplored problem, while existing video out-painting methods are mainly developed for relatively short clips. However, extending video out-painting to long videos is highly challenging. A natural solution is iterative out-painting, \ie, generating videos chunk by chunk, but this introduces two major issues: unstable even crashed long-horizon generation caused by inter-chunk inconsistency and accumulated errors, and semantic drift over long temporal horizons. In this paper, we propose LongSCOP, short for Semantically Consistent long Video Out-Painting. First, we provide a fundamental training recipe with three components for smooth and stable long-horizon generation: self-forcing training, latent calibration, and future references. Built upon this stable generation foundation, a more advanced requirement is semantic consistency, for which we further develop a VLM-based agent that automatically selects the most informative reference frames from the whole video, enabling coherent subjects, scenes, and backgrounds throughout long-range generation. To support training and evaluation, we construct a large-scale, high-quality training dataset and introduce LVO-Bench, a dedicated benchmark curated by human experts for assessing long-horizon stability and semantic consistency. Extensive experiments demonstrate clear advantages of our method over existing approaches in both generation stability and semantic coherence, significantly improving the quality and reliability of long video out-painting.
Authors:
Lin Jiang, Dahai Yu, Ximiao Li, Guang WangAbstract: Generating realistic time series is essential for scientific research and real-world applications. However, existing methods often emphasize overall distributional fidelity while failing to faithfully capture extreme events. To address this limitation, we propose E4GEN, an explainable diffusion framework for extreme event-aware time-series generation. E4GEN provides systematic insights into when, what, and how to control extreme-event generation through three key components. First, E-Activator learns a dataset-adaptive extreme-control signal activation step during the denoising process, enabling control without interfering with regular temporal components such as trend and seasonality. Second, E-Predictor determines what control signal to enforce through Self-Driven Semantic Prediction, where each sample derives its own control signal by inferring latent extreme-event information during generation. It also introduces a Data-Conditioned Training, Noise-Initiated Sampling mechanism to address the issue of unavailable training labels. Third, E-Control specifies how to guide extreme-event generation through a trainable Extreme Control Network, which transforms semantic control signals into layer-wise guidance signals and injects them into the denoising process. We evaluate E4GEN on six datasets using 17 metrics, and extensive experiments show that E4GEN outperforms state-of-the-art models across multiple dimensions, including overall fidelity, extreme-event fidelity, and downstream utility.
Abstract: Transformer inference on long sequences is expensive because softmax attention repeatedly reads from a large KV cache. The prevalent approach to this bottleneck is , which replaces the full cache with a compact summary. Despite its practical importance, the design of such summaries is largely driven by empirical experimentation. On the theoretical side, existing results show that KV cache compression can be impossible in the worst case, but offer little systematic guidance for designing algorithms in regimes where accurate compression is possible. We bridge this gap by characterizing the in terms of the intrinsic compressibility of a cache, revealing when and how accurate compression is possible. These results yield novel design principles for KV cache compression under causal masking, mapping efficiently to both prefill and autoregressive decoding while achieving minimax-optimal risk. We instantiate these principles in a practical algorithm and report promising performance on LongBench in targeted experiments. Overall, our results provide a principled avenue for practically applicable KV compression with theoretical guarantees.
Abstract: Multi-objective learning endeavors to concurrently optimize multiple objectives using a single model, aiming to achieve high and balanced performance across diverse objectives. However, this often entails a complex optimization problem to balance the learning of potentially conflicting objectives, leading to solutions with higher memory requirements and computational complexity. This paper introduces a Multi-Objective Goal-Conditioned Supervised Learning (MOGCSL) framework for automatically learning to achieve multiple objectives from offline sequential data. MOGCSL extends the conventional GCSL method to multi-objective scenarios by redefining goals from one-dimensional scalars to multi-dimensional vectors. It benefits from naturally eliminating the need for complex architectures and optimization constraints. Moreover, MOGCSL inherently disentangles uninformative or noisy training instances that fail to achieve desirable long-term rewards across multiple objectives. We also introduce a novel goal-selection algorithm for MOGCSL to model and identify desired and achievable goals for inference. In this paper, we focus on its application to the next action prediction problem in commercial-grade recommender systems. In this context, any viable solution needs to be reasonably scalable and also be robust to large amounts of noisy data that is characteristic of this application space. We show that MOGCSL performs admirably on both counts by extensive experiments. Also, analysis and experiments are included to explain its strength in discounting the noisier portions of training data in recommender systems with multiple objectives. Code is available at: https://anonymous.4open.science/r/MOGCSL-D7A2.
PaperID: 1557, Poster
Authors: Saar Cohen, Michael Wooldridge
Abstract: Coalition formation concerns partitioning agents into disjoint coalitions based on their preferences for one another. In preferences may be initially unknown. Thus, coalitions are repeatedly formed based on preferences learned notions, where no group of agents benefits by simultaneously joining other coalitions or forming a new one. Our goal is to learn coalition structures that are robust to such group deviations, while simultaneously determining whether stable outcomes exist under the underlying (unknown) preferences. To this end, we analyze algorithms through two complementary metrics: (1) the error in predicting whether a stable partition exists, and (2) a regret notion that quantifies violations of group stability over time. We propose an algorithm that combines optimistic and pessimistic estimates to both explore potentially stable partitions and certify stability when supported by the observed feedback. We show that when a stable partition exists, our algorithm obtains sublinear error and regret despite relying only on partial noisy feedback, while in the non-existence case, it incurs only a constant number of mistakes. Experiments on synthetic instances validate our findings.
Abstract: We study the problem of recalibrating an online predictor [KE17, OKS24]: given an arbitrary ``hint'' sequence of forecasts, the learner must output new predictions that are calibrated while incurring small excess error relative to the original forecasts, under a proper loss. We give an online algorithm that achieves (\varepsilon, \varepsilon^2)-recalibration for Lipschitz proper losses in T \approx \varepsilon^-3 rounds, using an imbalanced extension of the recent simultaneous Blackwell approachability reduction framework of [HTY26]. We also show this tradeoff is optimal by proving a matching lower bound for recalibrating against the squared loss. As an application, we show how our recalibration algorithm can be combined with the online refinement method of [FH23] to obtain simultaneous \varepsilon-calibration and \varepsilon^2-calibeating for smooth proper losses at the same asymptotic rate, improving upon prior works that achieved these properties separately or with a worse \varepsilon dependence. We also discuss extensions to settings with multiple hint sequences.
PaperID: 1559, Poster
Abstract: Clean-label poisoning attacks have been well studied for differentiable models, yet their practical behavior in gradient-boosted decision trees (GBDTs) remains less understood. In this paper, we investigate clean-label poisoning under realistic constraints: class labels are preserved, and features after perturbation must remain plausible. Our framework selects poisoning candidates using a tree-specific influence score and perturbs input features that prioritize dimensions based on split-gain signals. A key finding is a non-monotonic dependence on the perturbation budget \varepsilon, arising from discrete threshold crossings in tree splits and feasibility constraints on the perturbation. We introduce an ellipsoidal region and project both update iterates and finite-difference probes onto this region to obtain meaningful and feasible perturbations. Experimental results show that our attack is highly effective, e.g., F1 score on Adult dataset is 0.71 to 0.43, and F1 score 0.54 to 0.33 on Credit-g.
PaperID: 1560, Poster
Abstract: Language models are increasingly trained to "reason" before answering users' queries, outputting hundreds of intermediate tokens before a final answer. While reasoning is designed with direct model capabilities in mind, we argue that reasoning language models (RLMs) should also be assessed by interaction effects between their reasoning traces and weaker models. In contexts like safety monitoring and distillation, the performance of weak models increasingly depends on legibility---how thoroughly and precisely a reasoning trace externalizes an RLM's actions. Existing efficiency-based metrics for legibility fail to capture "thoroughness", instead focusing on conciseness. Thus, we introduce , a method for measuring trace legibility derived from interactions where weak models finish tasks given incomplete traces. Evaluating 85k traces from 12 RLMs across three tasks, we find that reasoning traces generated by the highest-performing and the most efficient models rank the for transfer utility. We also find evidence that transfer utility predicts the ability of weak monitors to verify procedurally dense traces for math or logical problem-solving, though the effect moderates when verifying fact-driven traces. Finally, we show that open-source reward models decouple from transfer utility signals on correctness. Together, these findings surface status quo trade-offs in weak-to-strong legibility, motivating an interaction-first framework for future applications.
Abstract: Many modern learning criteria, such as multicalibration, CVaR, and distributionally robust objectives, are set-input and optimization-valued: their empirical values depend on the whole evaluation batch and on an auxiliary optimizer selected after observing that batch. This optimization mismatch falls outside the standard information-theoretic generalization analysis for averages of fixed per-example losses. We provide a unified supersample-based analysis for this class of objectives. Under a selector-gap concentration condition, we show that the optimized validation--training gap is controlled by the conditional mutual information between the learned supersample predictions and the selector, without an explicit complexity penalty for the auxiliary optimizer class. We verify the concentration condition through stability and convex L_2-Lipschitz certificates, obtaining concrete bounds for multicalibration, CVaR, and Cressie--Read DRO, including a single-index refinement under stability.
PaperID: 1562, Poster
Abstract: In a standard transformer, each layer computes attention once: queries examine keys, select values, and move on. But attention is a function of representation, and representation is a function of attention. We close this loop. Fixed-Point Self-Attention (FPSA) computes its attention pattern by iterating the query/key/value process to convergence. Values V are derived from the layer input and held fixed; queries and keys are recomputed from an evolving output state Z_t at each step, converging when representations reach equilibrium, Z^ = f(Z^). FPSA is a drop-in replacement for multi-head attention that adds zero parameters and adapts depth on its own: easy tokens stabilize in 2--3 iterations, hard tokens refine over 15 or more, with \mathcalO(1) memory cost regardless of the number of iterations. It lifts BERT-Base and ELECTRA-Base on GLUE and SQuAD~v2.0, improves ViT-B/16 accuracy by up to 20%, and delivers matching gains on vision-language tasks, all without adding a single parameter. Unlike looped transformers that iterate entire layers, FPSA iterates only the attention step, so overhead is modest: a median of 3--6 steps per layer adds roughly 1.6× GFLOPs and 1.3--1.4× wall-clock time over BERT-Base. On multi-step reasoning benchmarks (GSM8K, BBH, LogiQA), FPSA's adaptive computation yields clear gains, letting the model refine longer when the input demands it.
PaperID: 1563, Poster
Abstract: We study heterogeneous treatment-effect (HTE) estimation in observational survival studies commonly associated with both censored outcomes and unmeasured confounding. We integrate causal survival forests (CSF) with negative controls (NC) from proximal causal inference and introduce Negative Control Causal Survival Forests (NC-CSF), a flexible nonparametric HTE learner for survival analysis. In particular, our approach uses a loss that incorporates proxy variables and Neyman orthogonalization to train the random forest, thereby mitigating bias from unobserved confounding and gaining robustness to nuisance estimation. Through extensive simulations spanning varying levels of confounding, proxy relevance, and censoring mechanisms, we demonstrate that NC-CSF substantially reduces bias and estimation error (root mean squared error) relative to existing baselines. We further demonstrate the practical utility of our method on real-world clinical datasets, where it confirms several existing findings and also reveals new interpretable patterns of treatment-effect heterogeneity. To facilitate practical use, we provide an end-to-end Python implementation of NC-CSF that carefully handles implementation details such as nuisance estimation and clipping.
PaperID: 1564, Poster
Abstract: Preference ranking modeling is important for selecting and adapting language-model responses to human preferences. Compared with deterministic ranking methods, generative preference ranking models the prompt-response data distribution, providing richer uncertainty characterization and sample diversity for learning a more expressive ranking boundary. Recent work improves modeling efficiency by shifting generative preference ranking from textual space to embedding space, but without explicit preference-conditioned separation, the learned embedding distribution may still entangle preferred and rejected responses, producing ambiguous synthetic samples that weaken the ranking boundary. To address this issue, we propose PRISM, a Preference Ranking framework through dISentangled embedding Modeling. PRISM formulates embedding-level generative preference ranking with a unified class-conditional ELBO and decomposes this optimization objective to expose the encoded latent entanglement between preferred and rejected embeddings. This derivation motivates two practical preference-aware generation variants: PRISMMI, involving a deep class-separation objective, and PRISMMMD, catering a probabilistic aggregate-matching objective. Both variants learn a preference-disentangled embedding data distribution and synthesize data pairs that preserve ranking semantics for the generative ranking boundary. Experiments show that PRISM improves preference ranking performance and produces useful generated embeddings for downstream response selection. The anonymous code repository is available at:https://anonymous.4open.science/r/PRISM-915C/README.md.
PaperID: 1565, Poster
Authors: Hong Chul Nam, Jeonghwan Cheon, Jae M Shin, Hyun Kwon
Abstract: We propose Energy-Based Operators (EBOs), an architecture-agnostic framework for learning conditional distributions over functions on continuous domains. An EBO defines a scalar energy over target functions given an input function and induces a probability model through a Gaussian reference. The resulting score field is obtained as the gradient of the parametrized energy to perform function-space iterative energy minimization (EM). Our model shows strong performance in function generation, such as super-resolution and forecasting, over various 1D function classes (oscillations, damping, and Izhikevich) in comparison with prediction operators and denoising operators over various architecture backbones. Moreover, it achieves strong performance over PDEs, namely Navier-Stokes, Darcy flow and Burgers. Notably, our model successfully detects anomalous functions by automatically assigning high energy without any supervision. It enables seizure detection and volatility prediction after learning neural dynamics and market microstructure dynamics without pre-defined labels during training, highlighting its effectiveness for both learning dynamical systems and detecting functional anomalies arising in scientific simulations.
Abstract: Gradient descent is usually studied through external regret, but recent work shows that it controls richer deviations through infinitesimal vector-field guarantees \citepahunbay2024first, semicoarse constraints \citepahunbay2025semicoarse, and weakly convex proximal maps \citepcai2025proximal. We identify the finite-deviation invariant behind these guarantees. A feasible deviation \(\phi\colon K\to K\) induces the displacement field \(\Dphi(\xvec)=\xvec-\phi(\xvec)\). Online gradient descent has \(O(\sqrt T)\) regret against every feasible deviation with exact displacement, \(\Dphi=\nabla\Psi\) for a potential function \Psi, even when \(\Psi\) is nonconvex. Feasibility handles the projection boundary term, and exactness makes the potential telescope. The condition is sharp for smooth fields on simply connected domains, since any nonzero circulation along a feasible loop can be turned into a bounded cyclic adversary that forces linear regret. Semicoarse affine deviations are the linear-quadratic slice of this class, and proximal deviations are the Moreau-envelope slice. The exact-form class strictly contains weakly convex proximal deviations. For mirror descent, exactness is measured in the mirror geometry through the one-form \(\Dphi(\xvec)^\transpose\dd\nabla R(\xvec)\), which subsumes Bregman proximal regret and includes multiplicative deviations for multiplicative-weight updates. In convex games, no-regret with respect to exact-form deviations yields conservative correlated equilibrium, refining coarse and proximal correlated equilibrium.
Abstract: Post-training Vision-Language-Action (VLA) models via reinforcement learning (RL) in learned world models has emerged as an effective strategy to adapt to new tasks without costly real-world interactions. However, while using imagined trajectories reduces the sample complexity of policy training, existing methods still heavily rely on task-specific data to fine-tune both the world and reward models, fundamentally limiting their scalability to unseen tasks. To overcome this, we argue that world and reward models should capture transferable physical priors that enable zero-shot inference. We propose RAW-Dream (Reinforcing VLAs in task-Agnostic World Dreams), a new paradigm that completely disentangles world model learning from downstream task dependencies. RAW-Dream utilizes a world model pre-trained on diverse task-free behaviors for predicting future rollouts, and an off-the-shelf Vision-Language Model (VLM) for reward generation. Because both components are task-agnostic, VLAs can be readily finetuned for any new task entirely within this zero-shot imagination. Furthermore, to mitigate world model hallucinations, we introduce a dual-noise verification mechanism to filter out unreliable rollouts. Extensive experiments across simulation and real-world settings demonstrate consistent performance gains, proving that generalized physical priors can effectively substitute for costly task-dependent data, offering a highly scalable roadmap for VLA adaptation.
Abstract: As powerful AI systems reach and sometimes surpass the abilities of human experts across a range of cognitively demanding tasks, the problem of accurate oversight and supervision of these systems has become increasingly urgent. One promising approach is AI debate, which seeks to leverage a debate between two powerful AIs to break complex questions down into simpler claims that can be easily judged directly. Theoretical work on debate has formalized this intuition in the language of computational complexity theory, where the goal is to design protocols (i.e. rules of the debate game) that provide rigorous guarantees on correctness for judging solutions to complex problems with limited supervision. Specifically, the current best protocol has been shown to work for all problems that have sufficiently-stable decompositions into subproblems. In this paper we design a new protocol for this same class of problems that improves on the prior work in several ways. First, correctness holds in a worst-case rather than an average-case sense. Second, honesty and correctness is a dominant-strategy equilibrium for both debaters, rather than a Stackelberg equilibrium. Finally, we prove black-box lower bounds, showing that our new protocol is instance-wise optimal. That is, no protocol for this class of problems can outperform ours while making only black-box queries to human judgments. We obtain these results by relating the notion of stable problem decompositions to the concept of fractional block sensitivity from query complexity.
Abstract: Subliminal learning refers to a student language model acquiring a teacher's traits (e.g. a system-prompted preference for owls) when fine-tuned on the teacher's outputs, despite the outputs being semantically unrelated to those traits. It remains poorly understood how data without semantic meaning can transfer specific semantic traits. In this work, we show that subliminal learning is mediated by a single steering vector, i.e. a vector added to the model's activations. Across two open-source models, we find that the teacher's system prompt is well approximated by a steering vector, and that the student's behavior is driven by learning an aligned vector over fine-tuning. System prompts that are not well approximated by steering vectors are not subliminally learned. This is a special case of , in which a student trained on the outputs of a steered teacher learns to imitate that steering. We demonstrate steering vector distillation on a range of semantic and random vectors. Adding a semantic vector to a model's activations can have both model-independent and model-specific (i.e. non-semantic) effects on its behavior, so generated data that is non-semantic can transmit a vector with semantic effects, enabling subliminal learning. This also explains why subliminal learning does not transfer between models. We find that adaptive optimizers are necessary for subliminal learning in language models: activation gradients on steered data carry a small but consistent component along the steering direction, and non-adaptive optimizers impede this by allowing outlier gradients to dominate.
PaperID: 1570, Poster
Abstract: Multimodal LLMs are typically trained by first aligning a vision encoder with a language model via projector training on large image-caption datasets, followed by joint training of the full multimodal model. Despite its widespread use, the role of the projector training stage remains poorly understood. Does it learn fine-grained visual-language correspondences, or mainly maps visual features into a compatible input space for the language model? We find that (i) using only 10 to 20% of the data for training the projector already recovers most downstream performance gains, and (ii) projectors trained on different datasets can be linearly interpolated while largely preserving performance. These results suggest that projector training acts primarily as a coarse alignment within a large solution space. Motivated by these findings, we propose PORTAL, a training-free projector initialization that requires neither paired image-caption data nor gradient-based optimization, but only per-modality summary statistics. PORTAL computes an optimal-transport-based map between Gaussian approximations of the vision and language feature distributions on a shared principal subspace, enabling completely skipping projector pretraining. Despite using no paired data, while the standard approach relies on 0.5M image–caption pairs, PORTAL matches the default trained pipeline across six LLM backbones, two vision encoders, and 16 benchmarks, outperforming it in 3 of 7 (vision, LLM) settings, and remaining within 1.6 average points across all other settings.
Abstract: We prove optimal sampling bounds achieving (1\pm\varepsilon)-relative error for a broad class of Lipschitz continuous classification loss functions under various regularization terms. This includes important functions such as logistic and sigmoid loss, hinge loss, and ReLU loss, as prominent and popular representative examples. In particular, we prove k^2/\varepsilon^2 upper and lower bounds for \|\cdot\|_2/k regularization, and k/\varepsilon^2 upper and lower bounds for \|\cdot\|_1/k regularization. For \|\cdot\|_2^2/k regularization, the sampling complexity depends mainly on a bounded derivative property: if |g'(x)|\leq g(x), and g(0)>0, and g is monotonic or convex, then it admits linear in k sampling complexity; otherwise the general bound is k^2/\varepsilon^2. However, if g(0)=0, our results indicate that no dimension-free bounds are possible, and even sublinear bounds are ruled out. All upper bounds are complemented by matching lower bounds up to polylogarithmic terms. Moreover, our work relies conceptually and algorithmically on simple uniform or (squared) norm sampling and hereby improves over recent cubic k^3/\varepsilon^2 sensitivity sampling bounds of (Alishahi and Phillips, ICML'24). This is achieved by refined arguments involving higher moment bounds and empirical process analyses to avoid overcounting that appears in the de-facto standard VC-dimension and sensitivity framework.
Authors: Asher Labovich
Abstract: Looped transformers promise test-time compute scaling by spending more iterations on harder problems, but it remains unclear which architectural choices let them extrapolate to harder problems at test time rather than memorize training-specific solutions. We introduce a fixed-point based framework for analyzing looped architectures along three axes of stability -- reachability, input-dependence, and geometry -- and use it to characterize when fixed-point iteration yields meaningful predictions. Theoretically, we prove that looped networks without recall have countable fixed points and cannot achieve strong input-dependence at any spectral regime, while recall with outer normalization reliably produces a regime in which fixed points are simultaneously reachable, locally smooth in the input, and supported by stable backpropagation. Empirically, we train single-layer looped transformers on chess, sudoku, and prefix-sums and find that downstream performance tracks the framework's predictions across tasks and architectural configurations. We additionally introduce internal recall, a novel recall placement variant allowing an identity residual path, and exploit its contrasting fixed-point behavior to test our framework's predictions in a controlled setting.
Abstract: Reinforcement learning with verifiable rewards has enabled strong post-training gains in domains such as math and coding, though many open-ended settings rely on rubric-based rewards. We study reward hacking in rubric-based RL, where a policy is optimized against a training verifier but evaluated against a cross-family panel of three frontier judges, reducing dependence on any single evaluator. Our framework separates two sources of divergence: verifier failure, where the training verifier credits rubric criteria that reference verifiers reject, and rubric-design limitations, where even strong rubric-based verifiers favor responses that rubric-free judges rate worse overall. Across medical and science domains, weak verifiers produce large proxy-reward gains that do not transfer to the reference verifiers; exploitation grows over training and concentrates in recurring failures such as partial satisfaction of compound criteria, treating implicit content as explicit, and imprecise topical matching. Stronger verifiers substantially reduce, but do not eliminate, verifier exploitation. We also introduce a self-internalization gap, a verifier-free diagnostic based on policy log-probabilities, which tracks reference-verifier quality, detecting when the policy trained using the weak verifier stops improving. Finally, in our setting, stronger verification does not prevent reward hacking when the rubric leaves important failure modes unspecified: rubric-based verifiers prefer the RL checkpoint, while rubric-free judges prefer the base model. These disagreements coincide with gains concentrated in completeness and presence-based criteria, alongside declines in factual correctness, conciseness, relevance, and overall quality. Together, these results suggest that stronger verification reduces reward hacking, but does not by itself ensure that rubric gains correspond to broader quality gains.
Abstract: In modern deep learning, weight decay is often credited with "stabilizing" training dynamics, diverging from its classical role as a static regularization penalty. We investigate a fundamental question: does weight decay stabilize training dynamics, and if so, through which mechanism? Indeed, training stability is understood through different but related notions in the literature. We consider how weight decay affects the parameter-space dynamics and loss sharpness by analyzing its effects at the Edge of Stability (EoS). We show that weight decay robustly slows progressive sharpening. Furthermore, we uncover a striking architecture-dependent phase transition. In CNNs, weight decay dampens the oscillations at the EoS, while in MLPs, increasing weight decay causes a phase transition in which the sharpness stabilizes at a threshold significantly below the theoretical \frac2\eta boundary. We develop a mathematical framework that accurately models these phenomena and identify the global alignment of the parameter vector and the sharpness gradient as the mechanistic driver of the phase transition. Importantly, we show that these phenomena translate into stability in terms of search in function-space (NTK).
PaperID: 1575, Poster
Abstract: Graph Neural Networks (GNNs) are powerful tools for learning from spatio-temporal data, as interactions in space can be naturally described by a graph structure. However, capturing long-range dependencies becomes substantially harder when information must flow through space and time simultaneously. Most existing approaches extend GNNs from static graphs to temporal settings by alternating propagation between time and space. In this work, we introduce STORM, a differential-equation-inspired GNN that exploits oscillatory dynamics to propagate information effectively in the joint spatio-temporal domain. By combining a wave-equation update with dissipative and external forcing terms, STORM balances conservative and non-conservative dynamics. We provide a bottom-up analysis of the model, highlighting its propagation behavior and stability, and show that STORM is universal. We empirically validate our method on diverse benchmarks, including tasks designed for analyzing long-range spatio-temporal dependencies, real-world forecasting, and a new long-range task inspired by physical simulations. Across these settings, STORM consistently matches or outperforms strong baselines, establishing a new state of the art for long-range spatio-temporal graph learning.
PaperID: 1576, Poster
Abstract: Motivated by learning from heterogeneous and overlapping data providers, we study a stylized model of distribution learning from restricted conditional samples. The goal is to learn an unknown distribution p on a finite domain [n]. The learner is given a fixed family of queryable sets \mathcal S \subseteq 2^[n], and each query to S \in \mathcal S returns an independent sample from the conditional distribution p(\cdot \mid S). Learnability is governed by the \emphco-occurrence graph associated with \mathcal S: two domain elements are connected if they appear together in some queryable set. Pointwise consistency is achievable when this graph is connected on the target support. PAC learning requires more: it is possible when the co-occurrence graph is complete. The sample complexity of PAC learning ranges from nearly linear to quadratic. Every query family with complete co-occurrence graph admits sample complexity \widetilde O(n^2/\epsilon^2), and this bound is tight in the worst case. On the other hand, if every subset is queryable, the optimal worst-case complexity improves to \Theta(n/\epsilon^2). More generally, we identify \emphhierarchical comparability as a sufficient structural condition on \mathcal S under which the optimal complexity is nearly linear, \widetilde \Theta(n/\epsilon^2), with pairwise query families as a canonical example. Finally, the full range of polynomial rates between linear and quadratic is attainable: for every \alpha \in (1,2), there exists a query family with optimal PAC rate \widetilde \Theta(n^\alpha/\epsilon^2).
PaperID: 1577, Poster
Authors: Siyin Huang, Yufan Liao, Yuelin Du, Yu Zhang
Abstract: Federated unlearning aims to efficiently remove the influence of a specific client from a trained global model in federated learning without full retraining. However, existing approaches struggle to preserve the utility of the remaining clients because unlearning can be unstable at both the client and server stages. Local forgetting updates may become overly aggressive and destabilize client side optimization, while server aggregation can amplify conflicts between forgetting updates and retained client updates. To address both failure modes within a unified framework, we propose Federated Unlearning with Gradient Adaptive Shaping (FUGAS), a history-free gradient shaping framework that stabilizes the entire unlearning pipeline. On the unlearning client side, we employ a bounded preference objective that utilizes the pre-unlearning model predictions on the forgetting data as negative references to controllably steer the model away from unlearning client knowledge while requiring only minimal storage for reference caches rather than full historical updates. On the server side, we introduce a compatibility projection mechanism that reshapes the aggregated unlearning update to remain compatible with directions estimated from retained clients. We provide a theoretical analysis indicating that FUGAS promotes stability and ensures non-increasing empirical risk on retained distributions while establishing an excess risk bound relative to retraining. Extensive experiments demonstrate that FUGAS achieves effective unlearning while consistently maintaining high accuracy on retained data.
PaperID: 1578, Poster
Abstract: Egocentric video generation aims to synthesize first-person visual experiences, enabling applications in filmmaking, virtual reality, and embodied AI. Generating egocentric videos from exocentric observations is particularly challenging, as it requires reasoning across large viewpoint changes, limited visual overlap, and substantially different camera motions. To address these modeling challenges, we propose , a framework that generates egocentric videos from a single exocentric input video. EgoEye integrates reward-guided egocentric context reasoning for inferring unseen first-person content, multi-stage motion alignment for enforcing cross-view temporal consistency, and egocentric pretraining for improving first-person realism. To support training and evaluation under diverse and distortion-free exo-ego settings, we further introduce , a large-scale dataset comprising 19.5K time-aligned exo-ego video pairs and 24.2K in-the-wild egocentric videos. EgoScape provides paired cross-view supervision while enriching the first-person visual priors needed for realistic egocentric synthesis. Extensive experiments demonstrate that the proposed EgoEye generates realistic and temporally consistent egocentric videos across diverse scenarios.
PaperID: 1579, Poster
Authors: Hyunjun Kim, Donggue Kim, Kichun Jo
Abstract: Vocabulary-based end-to-end driving selects a future trajectory from a fixed candidate set, enabling explicit comparison among multiple plausible plans. However, existing planners typically process the scene-invariant trajectory vocabulary and the scene-dependent driving context together in a single online forward pass. This repeatedly recomputes reusable candidate-side representations and applies expensive scene-conditioned evaluation to the full vocabulary. We revisit vocabulary-based planning as a retrieve-then-rerank problem, where the trajectory vocabulary is a static candidate collection, lightweight scene cues form a retrieval query, and the scene-conditioned evaluator acts as a reranker. Based on this view, we propose Decoupled Retrieval-Style Inference (DRSI), centered on two modules: Offline Candidate Indexing (OCI) and Scene-aware Candidate Retrieval (SCR). OCI caches scene-invariant trajectory embeddings offline, while SCR retrieves a compact set of route-consistent and dynamically reachable candidates online. The scene-conditioned evaluator is then applied only to this retrieved subset for fine-grained reranking. Experiments on NAVSIM show that DRSI substantially improves the latency--quality trade-off without adding trajectory proposals or increasing evaluator complexity. Compared with the Hydra-MDP++ baseline, our large DRSI model achieves 86.1 EPDMS on NAVSIM v2 with 23.9 ms latency, corresponding to about a 6.8× speedup. On the challenging navhard split, DRSI further improves EPDMS from 37.6 to 40.2. These results demonstrate that retrieval-style decoupling is an effective inference principle for efficient vocabulary-based end-to-end driving.
Authors:
Tao Lin, Yuxin Du, Jiting Liu, Nuobei Zhu, YUNHE LI, Yuqian Fu, Yinxinyu Chen, Hongyi Cai, Ye Zewei, Bing Cheng, Kai Ye, Yiran Mao, Yilei Zhong, Mingkang Dong, Junchi Yan, Gen Li, Bo ZhaoAbstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for robotic manipulation by unifying perception, language grounding, and action generation. However, they often struggle in scenarios requiring precise spatial understanding, as current VLA models primarily rely on 2D visual representations that lack depth information and detailed spatial relationships. While recent approaches incorporate explicit 3D inputs such as depth maps or point clouds to address this issue, they often increase system complexity, require additional sensors, and remain vulnerable to sensing noise and reconstruction errors. Another line of work explores implicit 3D-aware spatial modeling directly from RGB observations without extra sensors, but it often relies on large geometry foundation models, resulting in higher training and deployment costs. To address these challenges, we propose Evo-Depth, a lightweight depth-enhanced VLA framework that enhances spatially grounded manipulation without relying on additional sensing hardware or compromising deployment efficiency. Evo-Depth employs a lightweight Implicit Depth Encoding Module (IDEM) to extract compact depth features from multi-view RGB images. These features are incorporated into vision-language representations through a Spatial Enhancement Module (SEM) via depth-aware modulation, enabling efficient spatial-semantic enhancement. A Progressive Alignment Training strategy is further introduced to align the resulting depth-enhanced representations with downstream action learning. Extensive experiments in simulation and real-world settings demonstrate the effectiveness of Evo-Depth. With only 0.9B parameters, Evo-Depth achieves state-of-the-art-level performance across four simulation benchmarks. In real-world experiments, Evo-Depth attains the highest average success rate while also exhibiting the smallest model size, lowest GPU memory usage, and highest inference frequency among compared methods. These results demonstrate that lightweight implicit depth enhancement is an effective and practical solution for spatially grounded robotic manipulation. We release code and models to facilitate future research on lightweight depth-enhanced VLA models.
PaperID: 1581, Poster
Authors: Eduardo Laber, Miguel A Batista
Abstract: Due to their simplicity and interpretability, linkage methods are a central class of hierarchical agglomerative clustering (HAC). In particular, representative-based methods, such as centroid and medoid approaches, are appealing because their representatives provide intuitive summaries of the clusters. However, many representative-based linkage methods suffer from important drawbacks, such as high computational cost and the generation of inversions, which limit their broader applicability. We introduce relaxed-representative clustering, a principled generalization of representative-based linkages that preserves their interpretability while addressing these limitations. This framework yields (i) RMM, a quadratic-time variant of the cubic-time Minimax method, and (ii) RC, an inversion-free variant of the Centroid method whose representatives belong to the input set. On the theoretical side, we prove that for every k our methods are O(\min\k,\log n\)-approximation algorithms for the k-center objective. Empirically, experiments on tabular, biological, and textual datasets show that RMM consistently outperforms several baselines and achieves clustering quality comparable to Minimax at a fraction of the computational cost. Meanwhile, RC remains competitive with the Centroid method while producing inversion-free dendrograms and interpretable representatives.
PaperID: 1582, Poster
Abstract: The main challenge of long-tailed deep clustering is that imbalanced class frequencies cause head classes to dominate the partitioning of the representation space, which in turn submerges or distorts the cluster structures of tail classes. Existing methods typically alleviate this issue through heuristic head-tail partitioning or fixed auxiliary clustering. However, such approaches rely on manually specified discrete granularities and therefore struggle to capture the continuous latent structure of real-world long-tailed data. To address this limitation, we propose DPLC, a Dirichlet Process guided Long-tailed Clustering method that adaptively models the latent sub-cluster structure from a nonparametric Bayesian perspective. Specifically, DPLC periodically extracts soft latent sub-clusters from self-supervised embeddings and leverages them as data-driven rebalancing signals to mitigate the clustering bias induced by head-class dominance. By requiring neither a predefined number of auxiliary clusters nor explicit binary head-tail partitioning, DPLC enables more flexible modeling of complex long-tailed distributions. Extensive experiments show that DPLC consistently outperforms existing methods across multiple long-tailed clustering benchmarks and a wide range of imbalance ratios.
Abstract: We introduce the problem of adversary-directed online learning, in which the learner is aware of the set of instances in advance, and an adversary adaptively determines their ordering during the learning process. Surprisingly, this formulation has not been previously studied, perhaps due to its conceptual resemblance to the traditional adversarial online learning problem, despite related models, including the transductive, self-directed, and best-order, having been extensively explored. However, by utilizing novel techniques, we demonstrate that the landscape of adversary-directed online learning significantly diverges from that of traditional adversarial online learning. In the realizable setting, we establish a trichotomy of possible rates of the minimax number of mistakes. Specifically, for a learning horizon \operatornameT, the minimax number of mistakes can only be of the orders \Theta(\operatornameT), \Theta(\log \operatornameT), or \Theta(1). To prove this, we introduce a new combinatorial complexity parameter, termed the perfect Littlestone dimension, whose finiteness distinguishes the \Theta(\log \operatornameT) rate from the \Theta(\operatornameT) rate. On the other hand, in the agnostic setting, we essentially show a dichotomy of possible rates of the minimax expected regret. In particular, if the learner plays for \operatornameT \in \mathbbN rounds, its minimax expected regret can only be of the orders \Theta(\operatornameT), or \widetilde\Theta(\sqrt\operatornameT), which is also characterized by the finiteness of the perfect Littlestone dimension. Technically, a key ingredient in the proof of our \mathcalO(\log \operatornameT) and \widetilde\mathcalO(\sqrt\operatornameT) upper bounds is a novel online learning algorithm that leverages a new notion of shattering based on the perfect Littlestone dimension, which exploits the adaptive adversarial nature of the problem.
PaperID: 1584, Poster
Abstract: Tensor Singular Value Decomposition (t-SVD) has achieved strong empirical performance in multi-view clustering (MVC) and its incomplete setting (IMVC). However, the success of this framework heavily relies on the implicit dependency of its Fourier transform on the sample sequence, which raises serious concerns regarding sample-permutation sensitivity. To address this issue, we first formulate the permutation sensitivity in multi-view clustering by providing rigorous definitions of invariance at both the outcome level and the representation level. Secondly, by revealing the potential structural risks associated with the fixed Fourier basis, we propose two plug-and-play strategies: (1) an Order-Restoration strategy (K-OR), designed to reconstruct a locally smooth signal favored by the Fourier transform; and (2) an equivariant t-SVD framework (Cor-SVD), which fundamentally resolves the sensitivity of the Fourier transform by replacing fixed bases with data-dependent, equivariant transforms. Finally, we design a controlled permutation evaluation protocol to verify the robustness of tensorial clustering models. Extensive experiments demonstrate the significant performance variability of existing t-SVD-based methods under sample permutations, whereas our approaches substantially improve their permutation robustness while maintaining competitive clustering performance. These findings reveal critical vulnerabilities in current tensorial clustering benchmarks and underscore the necessity of the proposed method for achieving robust performance and reliable evaluation.
PaperID: 1585, Poster
Authors: Mohammed Jamal, Soufiane Fafe
Abstract: Exact optimal-decision recovery can be stronger than required when decisions are evaluated at a finite regret tolerance. This paper studies finite-resolution sufficiency under partial linear observations and bounded measurement error. A design is sufficient when every noisy observation fiber admits a single feasible decision with regret at most a prescribed tolerance for all costs in that fiber. We prove a complete finite certificate for insufficiency: an indistinguishable hard tuple of costs with no common low-regret decision. For linear-optimization regret, the required tuple size is controlled by projected cost dimension after quotienting directions that are invisible to regret. For finite cost libraries, this certificate gives exact fixed-design verification and separation. Under coordinatewise noise and a finite query library, query design is exactly set cover over hard tuples; this reduction transfers the standard set-cover greedy guarantee and hardness barrier, and supports an exact delayed-constraint method under exact verification and exact restricted-master solves. Structured cases reduce to polynomial formulations through separability, nested one-dimensional coverage, and interval incidence.
Abstract: Recent feed-forward geometry foundation models have demonstrated impressive generalization by recovering depth and poses in a single forward pass. However, these models are typically constrained by a global coordinate frame assumption. This dependency becomes a significant bottleneck for long-context and streaming reconstruction, as it forces the network to maintain an arbitrary temporal origin and handle translation magnitudes that grow unbounded over time. Our solution, which we call R^3, employs relative regression. We employ a lightweight MLP to predict confidence-weighted relative constraints. These confidences serve as a unified anchor: weighting losses during training and guiding pose aggregation during inference. R^3 supports both full-context offline reconstruction and causal, bounded-memory streaming. Our evaluation in both offline and streaming settings validates the effectiveness of our relative mechanism.
Abstract: LLMs are increasingly used with external knowledge sources like the Internet. Do they weigh information appropriately---updating more for reliable sources (source discernment) and more when claims bring priors closer to the truth (truth discernment)? We formalize this as information discernment and introduce Learn2Discern (L2D), an experimental framework and benchmark grounded in three normative axioms with interpretable metrics. To establish external validity, a pre-registered, quota-matched user study (n=299) confirms that real LLM users endorse all three axioms and report that violations reduce their trust and usage intent. Across 13 models and nearly 670K trials, we find consistent failures across both dimensions: models perform near chance on source and truth discernment, rely on source popularity twice as much as source reliability, and update roughly equally whether a claim improves or worsens their position relative to the ground truth. Models integrate external knowledge most effectively on datasets where their priors are already the most accurate. Newer and larger models improve truth discernment but not source discernment, a blind spot that model complexity does not address. We identify simple inference-time interventions that meaningfully improve both forms of discernment. We release our dataset, metrics, and survey as a testbed for a core alignment property that scales in importance as LLMs replace traditional search.
PaperID: 1588, Poster
Authors:
Jake Grigsby, Siddhant Agarwal, Yu Lei, Leonidas Varveropoulos, Yuke ZhuAbstract: Meta-reinforcement learning (meta-RL) agents adapt to novel tasks from test-time experience, but require diverse training sets of environments and reward functions that are expensive to construct. Behavior Foundation Models (BFMs) learn policies from reward-free data that are capable of zero-shot RL (ZSRL), adapting to new objectives at test time when given a large reward-labeled dataset. Recent work has shown that BFM task inference can be performed online, making BFMs and meta-RL direct competitors for generalization in fixed environments. We ask whether BFMs can be extended to adapt to novel reward functions in novel environments, and identify four key limitations: environment identification, online data collection, test-time exploration-exploitation, and mixed reward supervision. We study these subproblems through toy domains, standard BFM benchmarks, and simulated humanoid locomotion, then propose a unified framework that addresses all four. The resulting method, MetaBFM, is a hybrid RL agent spanning meta-RL, ZSRL, and intrinsic exploration. During training, MetaBFM combines supervised reward-following with unsupervised reward-free learning; at test time, it interpolates between exploration, exploitation, and ZSRL-style reward inference. We evaluate MetaBFM on two toy meta-RL domains and at scale in MetaWorld, showing that hybrid meta-RL/ZSRL agents can learn more general behavior from the same set of reward functions and may reduce the need for future meta-RL domains to hand-design diverse training sets.
PaperID: 1589, Poster
Authors: Sean Man, Ron Raphaeli, Matan Kleiner, Or Ronai
Abstract: In this paper, we introduce SP^3, a novel Plug-and-Play algorithm that accelerates Maximum \empha posteriori image restoration by replacing denoisers with Spherical Encoders (SE) as generative priors. SP^3 approximates the intractable proximal prior step by utilizing the SE tightly structured latent space as a robust projection onto the natural image manifold. Alternating this projection with a closed-form data-consistency step, via Half-Quadratic Splitting, guarantees stable convergence without requiring gradient computation during inference. This unique formulation unlocks ``anytime'' restoration capabilities, producing sharp, plausible images from the first iteration. Evaluations across a variety of image restoration tasks demonstrate that SP^3 achieves perceptual quality comparable to state-of-the-art zero-shot diffusion and flow methods while being 3-630× faster.
Abstract: We characterize the minimax rate of estimating the second-order calibration error for binary classification, which quantifies whether a higher-order predictor's epistemic-uncertainty estimate matches the conditional variance of the label probability on its level sets. Our key observation is that the sech perturbation kernel, previously used only to enforce smoothness of calibration functions, in fact makes them \emphanalytic in a strip of half-width h\pi/2. Polynomial regression then estimates the calibration error at rate \tildeO(1/\sqrtn), with explicit constants, a qualitative improvement over the O(n^-1/4) rate achievable by bucketing or kernel smoothing. A matching \Omega(1/\sqrtn) lower bound establishes minimax optimality up to logarithmic factors. As a corollary, we give the first finite-sample guarantee for second-order Platt scaling, yielding a post-hoc procedure that recalibrates both the mean prediction and the epistemic-variance estimate of any higher-order predictor. Along the way, we provide a bucket-free definition of second-order calibration and relate it quantitatively to the bucketed formulation of Ahdritz et al. [2025]. Our experiments confirm the predicted rate and the quality of the recalibrated uncertainties.
PaperID: 1591, Poster
Authors: Haizhou Du, Jinze Zhao, Huaicheng Yan
Abstract: Multi-agent systems (MAS) have emerged as a dominant paradigm where agent routing is a bottleneck for balancing performance and cost. However, existing agent routing methods suffer from the prohibitive cost of exhaustive benchmark construction and the neglect of query complexity essential for mixed tasks. To address this issue, we propose the RADAR (Routing Agents via Difficulty-Aware Recovery) framework. RADAR enables accurate decision-making under limited evaluation cost through two complementary modules. MatrixComp module employs structure-aware matrix completion to recover global performance profiles from sparse observations. Furthermore, DiffMatch module extracts reasoning meta-features for fine-grained alignment via a difficulty-calibrated retrieval mechanism. Extensive experiments demonstrate that RADAR maintains performance equal with state-of-the-art baselines while reducing the cost of benchmark construction by at least 33%.
Abstract: We study the problem of data selling for Retrieval Augmented Generation (RAG) tasks in Generative AI applications. We model each buyer's valuation of a dataset with a natural coverage-based valuation function that increases with the inclusion of more relevant data points that would enhance responses to anticipated queries. Motivated by issues such as data control and prior-free revenue maximization, we focus on the scenario where each data point can be allocated to only one buyer. We show that the problem of welfare maximization in this setting is NP-hard even with two bidders, but design a polynomial-time (1-1/e) approximation algorithm for any number of bidders. Unfortunately, however, this efficient allocation algorithm fails to be incentive compatible. The crux of our approach is a carefully tailored post-processing step called data burning which retains the (1-1/e) approximation factor but achieves incentive compatibility. Our thorough experiments on synthetic and real-world image and text datasets demonstrate the practical effectiveness of our algorithm compared to popular baseline algorithms for combinatorial auctions.
PaperID: 1593, Poster
Authors: Zijie Xing, Xingyue Liu, Runze Wang, Luoming Hu, Yanming Shen
Abstract: Large language models (LLMs) have shown promise in scientific discovery tasks such as protein optimization, but remain limited by the lack of a mechanism for persistent and reusable knowledge accumulation. Existing approaches adapt LLMs through parameter optimization or trajectory refinement, but in both cases, knowledge remains implicit, either encoded in opaque model weights or scattered across transient solution histories. As a result, knowledge cannot be systematically inspected, attributed, or reused. We propose \emphKnowEvo, a framework that treats externalized knowledge itself as the optimization target. Our key insight is that scientific knowledge is inherently modular: multiple interacting mechanisms jointly determine the design outcomes. This makes knowledge evolution fundamentally challenging, as improvements cannot be attributed to individual components in isolation. To address this, \emphKnowEvo represents knowledge as executable expertise blocks and organizes their refinements into strategy mutation trees, enabling structured exploration over a combinatorial knowledge space. The framework combines a forward exploration phase that adaptively allocates evaluation budget across strategy variants with a backward revision phase in which an agentic sandbox converts rollout evidence into localized program improvements. This design enables persistent accumulation, attribution, and refinement of mechanistic expertise across iterative experimentation. We evaluate \emphKnowEvo on protein optimization tasks including thermostability and solubility. Results show that knowledge evolution consistently improves over initial expert heuristics and outperforms monolithic program evolution and prior LLM-guided evolution baselines across ranking metrics such as pass@k, MRR, MAP, and NDCG@5. Codes are available at \urlhttps://anonymous.4open.science/r/KnowEvo-6E3B/.
Abstract: Vision Transformers (ViTs) achieve strong performance across vision tasks, yet their deployment with low-precision early exiting remains fragile. Existing quantization methods assume static full-depth execution, making them unstable when exit decisions are perturbed by quantization noise, which can amplify errors along dynamic inference paths. In this paper, we introduce , a utilization-aware formulation that accounts for layer-wise stochastic exposure to quantization noise and reveals , a bi-level framework that jointly optimizes exit thresholds and bit-widths under explicit risk control to improve inference stability. MAQEE establishes a superior in the accuracy--efficiency trade-off, reducing BOPs by up to 95% while maintaining accuracy and outperforming strong baselines by up to 20% across classification, detection, and segmentation tasks. Source code is available at
Abstract: Auto-labeling has emerged as a popular, cost-effective alternative to manual annotation. However, its performance is often highly inconsistent across different tasks, underscoring the importance of selective strategies. Despite this, existing heuristic methods for selective auto-labeling still rely heavily on model confidence scores and offer no reliable guarantee on the trustworthiness of the selected examples. To address this, we propose Conformal Labeling, a novel method that selects a subset of auto-generated labels with a provably controlled false labeling rate (FLR). Our key idea is to formulate selective auto-labeling as a multiple-hypothesis testing problem, where each hypothesis indicates whether to accept the auto-generated label for an instance. Specifically, we construct a conformal p-value for each auto-generated label using a small calibration set, and then employ the Benjamini–Hochberg (BH) procedure with a novel correction factor to construct the subset with guaranteed FLR control. Theoretically, we show that our method achieves finite-sample FLR control and asymptotically optimal statistical power among all p-value thresholding rules. Extensive experiments on classification and open-ended generation tasks validate the effectiveness of our method, achieving high statistical power while strictly controlling the FLR.
PaperID: 1596, Poster
Abstract: Masked Diffusion Models (MDMs) are trained under an objective that is symmetric over decoding orders, but at sampling time a single order must be committed to. A common task in such models is to evaluate the log-probability of a candidate sequence, either as a downstream score or to rank a finite pool of candidates. We observe that this evaluation depends substantially on the decoding order chosen, even though the training objective averages over orders. For both a model trained on closed-form Probabilistic Context-Free Grammar (PCFG) data and a pre-trained MDM on OpenWebText, the per-order log-probability rankings of fixed candidate pools disagree on most inputs. We characterize this phenomenon and propose Order-Marginalized Scoring (OMS), a simple Monte-Carlo (MC) estimator of the order-marginalized log-likelihood \tilde \mu(x)=E_\sigma[\log p_\sigma(x)]. We use it in two complementary algorithms: a reranker over a fixed candidate pool, and a stochastic best-of-N decoder that combines adaptive position selection with \tilde\mu-reranking. On Sudoku, the decoder substantially raises exact-match accuracy over the strongest single-order baseline from 65.3% to 94.7%. On a zero-shot protein-stability benchmark, the reranker improves Spearman correlation between log-probability scores and experimental folding stabilities over a per-order baseline by 0.07. These results suggest that OMS is a useful alternative to per-order scoring for order-agnostic problems.
Abstract: Membership Inference Attacks (MIAs) aim to distinguish training points (members) from unseen data (non-members), and are widely used to quantify memorization and assess privacy risks. Standard MIA evaluation requires repeated retraining, which is computationally costly for large models. One-run (single training with randomized data inclusion) and zero-run (post hoc evaluation) methods are often used instead, but their statistical validity remains unclear. We address this gap by framing MIA evaluation as a causal inference problem, defining . This novel formulation reveals and formalizes key sources of bias in existing protocols: one-run methods suffer from interference between jointly included points, while zero-run evaluations are additionally confounded by distribution shift between member and non-member evaluation data. We derive causal analogues of standard MIA metrics and propose practical estimators for multi-run, one-run, and zero-run regimes with non-asymptotic consistency guarantees. We validate our approach in several settings, including pretrained and fine-tuned LLMs, showing that it enables reliable measurement of MIA performance without retraining and under distribution shift. Overall, our framework provides a principled foundation for privacy evaluation in modern AI systems.
Abstract: We study price-adjustment dynamics for computing competitive equilibria (CE) in Fisher markets with chores. Unlike in classical goods markets, prices in chores markets are payments for taking on undesirable tasks, and natural excess-demand dynamics can fail; even the naïve analogue of Walrasian tâtonnement may diverge. Recent work of Chaudhury et al. (2025) overcomes this obstacle via \emphrelative tâtonnement, which subtracts the average excess-demand signal from the excess demand vector. This recovers convergence, but at the cost of coupling the price updates across all chores. This leaves open whether such global coupling is inherent, or whether convergent tâtonnement can be recovered through a genuinely local update in which each chore reacts only to its own excess demand. We answer this question affirmatively through \emphmultiplicative tâtonnement, a fully distributed dynamics in which each chore price is updated using only its current price and its own excess-demand signal. Although the update contains no explicit normalization term, Walras' law and the multiplicative form of the update implicitly preserve the relevant aggregate price geometry. We prove that multiplicative tâtonnement converges to a CE in any chores Fisher market with continuous, convex, and 1-homogeneous (CCH) disutilities. For convex CES disutilities, we further prove an approximate-CE convergence rate with the same O(1/\varepsilon^2) dependence as relative tâtonnement, but with improved dependence on problem constants. Experiments on real-world and simulated instances show that multiplicative tâtonnement is substantially faster in practice, often by an order of magnitude.
PaperID: 1599, Poster
Abstract: Retraining is the primary mechanism by which deployed models adapt to new data, yet it is also among the most expensive operations in modern machine learning. How often must a model be retrained to remain near-optimal? We answer this question through a learning-curve based characterization of the retraining-frequency/risk trade-off in online learning. For the i.i.d. realizable setting, we observe that only O(\log T) updates suffice to match the risk of full retraining whenever the learning curve is non-increasing. However, when the learning curve decays as a power law t^-\alpha with \alpha < 1, as empirically observed in deep learning, the budget collapses further to O(\log \log T) updates, yielding a substantial asymptotic improvement. We further design update schedules achieving these bounds, prove matching lower bounds, and present an adaptive algorithm that remains optimal when \alpha is unknown. We then extend the analysis to piecewise-stationary and gradually drifting environments, and establish a no-free-lunch theorem showing that some prior knowledge of the learning curve is unavoidable, i.e., no universal algorithm can be competitive without it. Together, these results provide a sharp characterization of the frequency--accuracy trade-off in online retraining and bridge foundational learning theory with practical strategies for scalable deployment.
Authors:
Atharva Mahajan, Abhijeet Vishwasrao, Yuning Wang, Ricardo V MotilvaAbstract: Skin-friction drag induced by wall-bounded turbulent flows accounts for a substantial fraction of energy consumption across commercial aerospace, wind energy, and marine transport. Its active reduction is one of the highest-value targets in engineering fluid dynamics. Deep reinforcement learning (DRL) has emerged as the leading approach for real-time flow control, yet its performance ceiling is set not by algorithmic capability but by reward structure, the naive scalar objective does not optimally reflect the underlying physics. Policy-DRIFT bypasses this ceiling by relocating reward information from policy gradients to generative model inference: a conditional flow matching model (CFM) constructs a physically-grounded manifold of realisable flow states spanning multiple control regimes, Terminal Reward Guidance (TRG) steers samples toward reward-maximising targets at inference, and a lightweight DRL policy, structurally decoupled from reward quality, tracks these full-field targets via root-mean-squared error (RMSE) minimisation. The test case is turbulent channel flow simulated using direct numerical simulation (DNS) at friction Reynolds number of \mathrmRe_\\tau = 180, which is the canonical benchmark for wall-bounded turbulence. Policy-DRIFT achieves 49% drag reduction approaching the theoretical upper bound, which is \\approx 16% higher than the DRL benchmark, while consuming 37\× less actuation energy. Our approach combines generative methods with active flow control, marking a paradigm shift towards controlling complex physical systems efficiently.
Abstract: Comprehensively retrieving diverse documents is crucial to address queries that admit a wide range of valid answers. We introduce retrieve-verify-retrieve (RVR), a multi-round retrieval framework designed to maximize answer coverage. Initially, a retriever takes the original query and returns a candidate document set, followed by a verifier that identifies a high-quality subset. For subsequent rounds, the query is augmented with previously verified documents to uncover answers that are not yet covered in previous rounds. RVR is effective even with off-the-shelf retrievers, and fine-tuning retrievers for our inference procedure brings further gains. Our method outperforms baselines, including agentic search approaches, achieving at least 10% relative and 3% absolute gain in complete recall percentage on a multi-answer retrieval dataset (QAMPARI). We also see consistent gains on two out-of-domain datasets (QUEST and WebQuestionsSP) across different base retrievers. Our work presents a promising iterative approach for comprehensive answer recall leveraging a verifier and adapting retrievers to a new inference scenario.
Authors: Arwa Alanqary, Zakaria Baba, Manxi Wu, Alexandre Bayen
Abstract: We study preference learning and coordination through recommendations in multi-agent game settings, where a moderator repeatedly interacts with agents whose utility functions are unknown. In each round, the moderator issues action recommendations and observes whether agents follow or deviate. We consider agents who best respond to the moderator's recommendations and study what this feedback reveals about their utilities. We characterize the class of games that are indistinguishable under this feedback model. Moreover, we introduce a notion of moderator regret based on agents' incentives to deviate from the recommendations and design an online algorithm with low regret under the best-response model, with guarantees that scale linearly in the game dimension and logarithmically in time. Our results lay a theoretical foundation for AI recommendation systems in strategic multi-agent environments, where recommendation compliance is shaped by strategic interaction.
PaperID: 1603, Poster
Authors: Zehua Cheng, Wei Dai, Jiahao Sun
Abstract: Distilling "System 2" reasoning into compact student models is bottlenecked by the scarcity of fully-correct teacher traces: as task difficulty grows, rejection sampling discards an increasingly large fraction of teacher generations. "Wrong" (W) traces are abundant, but training on them indiscriminately causes policy poisoning the student internalises the teacher's hallucinated arithmetic alongside any useful reasoning structure. We propose Atomic Reasoning Units (ARU), a framework that resolves this trade-off via strict syntactic constraints. ARU enforces a Backus--Naur-Form (BNF) grammar that decomposes each reasoning step into a Premise, an Operation, and a Result, algorithmically disentangling reasoning planning from arithmetic execution. Counterfactual Logic Verification (CLV) symbolically re-executes W-traces to recover the 68% that have valid plans but faulty arithmetic, and Syntactic Loss Masking (SLM) trains the student on this verified structure while suppressing the gradient on hallucinated result tokens. Across the 3 × 8 × 6 grid we evaluate (\0.6, 1.7, 8\\,B Qwen3 students × eight benchmarks × six baselines), ARU is the strongest method in every cell, improving GSM8K by +11.1 and MATH-500 by +9.8 over the closest baseline at 1.7\,B; on the held-out AIME competition set ARU's Maj@8 is 14.8 versus 9.5 for PoT (the closest baseline). We further document a small-scale capability gap in which ARU at 0.5\,B matches Gold-SFT at \geq 1.5\,B on GSM8K (paired-bootstrap p<0.05 after Bonferroni correction), with a pre-registered ProofWriter control showing the gap collapses on pure-logic reasoning---consistent with arithmetic offloading as the dominant mechanism, though we do not claim the control uniquely isolates it. The code is available at https://anonymous.4open.science/r/ARU-286D.
PaperID: 1604, Poster
Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a powerful tool for improving multimodal reasoning. However, standard outcome supervision suffers from credit assignment ambiguity, as a model may receive positive rewards by exploiting language priors or dataset shortcuts rather than grounding its reasoning in the requisite visual evidence. This makes high-quality negative responses critical, since effective negatives should remain plausible under the current policy while exposing the shortcut that should be suppressed. In this paper, we propose EAN, a multimodal reinforcement learning framework utilizing Evidence-Ablated Negatives to train LMMs against their own language-prior failures. EAN first identifies highly visually dependent examples using a scaled image information gain metric and preserves data diversity with a mixed-ranking strategy. It then constructs evidence-ablated images by masking the most discriminative image patches, eliciting plausible but visually ungrounded responses from the current policy. During policy optimization, these negative responses are incorporated as bounded auxiliary signals under the clean image, penalizing shortcut learning while maintaining RL stability. Experiments across challenging multimodal reasoning benchmarks demonstrate that EAN mitigates modality imbalance and improves over strong RLVR baselines.
PaperID: 1605, Poster
Authors: Tommaso d’Orsi, Gleb Novikov, Walter McKelvie
Abstract: We study the problem of d-dimensional covariance matrix release under pure differential privacy with error measured in Schatten-p norms, where p\in [1,\infty]. We identify two meaningful sample-size regimes with qualitatively different behavior, and give a single efficient algorithm together with matching information-theoretic lower bounds throughout. In the \emphlarge-sample regime n \gtrsim d^2/\varepsilon, we show that the simple K-norm mechanism analyzed by \cited2026purely is simultaneously optimal for all Schatten norms p\in[1,\infty], yielding sharp rates from nuclear to spectral norm. In the \emphmoderate-sample regime d/\varepsilon \lesssim n \lesssim d^2/\varepsilon, the problem remains non-trivially solvable, but the K-norm mechanism becomes suboptimal. We then design a single improved, efficient, \varepsilon-differentially private estimator based on the perturb-and-project framework; this estimator is optimal in all regimes and simultaneously for all Schatten norms. Prior to this work, improvements over the K-norm mechanism in this regime were only known for Frobenius loss \citenikolov2023private. A key feature of the moderate-sample regime is an inherent dependency of the error on the nuclear norm of the input, first observed by \citedong2022differentially and further investigated by \cited2026purely. Our mechanism captures this dependence uniformly over p, and our lower bounds show it to be unavoidable, establishing optimal rates in all regimes, including the correct dependence on \|\Sigma\|_. %The implications of our approach go beyond covariance release: we obtain optimal guarantees for privately releasing arbitrary matrices under nuclear-norm adjacency, and give partial extensions to higher-order moment estimation under pure differential privacy. These extensions highlight the perturb-and-project framework as a flexible tool for private high-dimensional data release, even beyond second moments. Our results extend and are optimal for the more general problem of privately releasing arbitrary matrices under nuclear-norm adjacency. Finally, our lower-bound techniques further generalize to higher-order tensors, establishing similar limitations for higher-order moment estimation under pure differential privacy.
Abstract: Amorphous (disordered) materials require simulation cells with at least hundreds to thousands of atoms to accurately capture their diverse local atomic environments, compared to crystalline materials that can be described by unit cells containing a few to hundreds of atoms. To advance the design of amorphous materials with desired properties and facilitate the exploration of their vast design space, generative inverse design has emerged as a promising approach. It aims to directly output materials with properties closely aligned with the desired ones using probabilistic generative models conditioned on desired properties, which can be more resource efficient than the traditional trial-and-error approach. However, due to the inherent stochasticity of probabilistic generative models, when element assignments are unconstrained, a large portion of generated materials may be charge unbalanced, and no existing methods can effectively mitigate this limitation. In this work, we propose AMGenC, a new generative inverse design method for amorphous materials that can guarantee the generation of charge balanced samples, with minimal additional computational overhead and without sacrificing inverse design accuracy. AMGenC achieves this through an element noise that gives the generation process a starting point centered around charge balance, and the combination of a per-step soft projection and a final discrete projection for steering the elements toward exact charge balance throughout the generation. We perform extensive experiments on two amorphous materials datasets. Experimental results provide evidence that AMGenC achieves its design goal.
Abstract: Graph-level representations are crucial tools for characterising structural differences between graphs. However, comparing graphs with different cardinalities, even when sampled from the same underlying distribution, remains challenging. Unsupervised tasks in particular require interpretable, scalable, and reliable size-aware graph representations. Our work addresses these issues by tracking the structural diversity of a graph across coarsening levels. The resulting graph embeddings, which we denote , are interpretable by construction, efficient, and directly comparable across coarsening hierarchies. Specifically, we track the of graphs, a novel isometry invariant that is inherently well-suited for encoding the metric diversity and geometry of graphs. We utilise edge contraction coarsening and prove that this improves expressivity, thus leading to more powerful graph-level representations than structural descriptors alone. Demonstrating their utility over a range of baseline methods in practice, we use to (i) cluster and visualise simulated graphs across varying sizes, (ii) distinguish the geometry of single-cell graphs, (iii) compare the structure of molecular graph datasets, and (iv) characterise geometric shapes.
PaperID: 1608, Poster
Abstract: Preference-based alignment methods typically optimize against a single preference model, and can therefore be brittle when pairwise preferences are uncertain: noisy, heterogeneous, or shift after deployment. We study alignment under uncertain preferences through the lens of general-preference games. Specifically, we formulate Robust Nash Learning from Human Feedback, where the learner seeks a policy with a large worst-case win rate against both an adversarial competitor and any preference kernel lying in an ambiguity set around a nominal preference. When the ambiguity set captures the uncertainty in preferences, the resulting hard-constrained robust objective directly yields a certified lower bound on worst-case performance. However, we note this problem is computationally challenging to optimize, and to address this, we introduce a four-player primal-dual proxy game involving the leader policy, follower policy, adversarial kernel, and dual variable, and develop a single-loop optimistic mirror descent-ascent algorithm for this game. We show that the proxy always lower-bounds the truncated hard-constrained objective, quantify the proxy-to-hard gap, and characterize an exactness condition under which the proxy recovers the robust objective. We then prove an O(1/\sqrtT) convergence rate for the proxy-game duality gap, which implies a near-optimal robust policy for the original robust objective. Experiments on controlled tabular games and LLM alignment with uncertain preference further validate the convergence theory and show improved performance over nominal baselines.
Abstract: Heavy-tailed data arise in many domains where rare events carry disproportionate importance, such as imbalanced image datasets, financial returns, and weather extremes. Standard diffusion and flow-matching models typically begin from Gaussian noise or Gaussian source distributions, which yield tractable training targets but provide a poor inductive match for heavy-tailed data. We propose Heavy-Tailed Flow Matching via Random Clocks (HTFM), a framework that portrays heavy-tailed sources as mixtures of clock-conditioned Gaussian sources. Conditioning on a given clock path, the source distribution and flow are Gaussian; marginalizing over the clock gives a Gaussian scale mixture covering Gaussian, \alpha-stable, and Student-t families. To make the clock-conditioned vector field practical, we encode the path-valued clock using truncated logsignature features, allowing the velocity field to adapt to the realized conditional space with negligible overhead. Empirically, on 2D imbalanced \alpha-stable mixtures, CIFAR10-LT, and HRRR weather fields, HTFM improves mode coverage, sample quality, and tail-statistic recovery over Gaussian flow matching and competitive heavy-tailed baselines, while retaining the low-NFE sampling advantage of flow matching. Moreover, the random-clock formulation further provides a practical tail-control interface: by varying only the clock law or tail parameter, the same architecture can calibrate the ``heaviness'' of generated tails across different distribution families.
Abstract: While attention mechanisms are the cornerstones of modern foundation models, the theoretical understanding remains largely elusive. In this paper, we analyze the attention-indexed model, a rich framework that encapsulates multi-layer and multi-head attention. We first prove that while, in a suitable high-dimensional limit, the macroscopic landscape of the population loss is characterized by a finite set of order parameters, the corresponding population gradient flow unfolds in an infinite-dimensional state space, which we show can be exponentially well-approximated by a finite truncated system. Our dynamical analysis uncovers how standard attention parameterizations act as architectural implicit biases to overcome the curse of the information exponent—a barrier that typically traps the direct optimization of the attention matrix S. Specifically, the tied attention (S=WW^T) induces an automatic symmetry-breaking geometry at initialization, shielding the gradient flow from uninformative manifolds and achieving weak recovery in \Theta(1) time. For untied attention (S=UV^T), we reveal a timescale separation dynamic: efficient weak recovery hinges on whether the fast-timescale evolution of the pre-activation mean successfully breaks initial symmetries.
Abstract: Modern LLM workflows increasingly move coordinate-indexed objects across checkpoints: steering vectors, sparse autoencoders, top-k neuron sets, attribution lists, and merge alignments. This is only well posed after fixing the model's residual-stream gauge. We show that the native discrete gauge is architecture-dependent: LayerNorm residual charts have permutation gauge S_d, up to a global sign flip, while RMSNorm residual charts with generic per-channel gain have signed-permutation gauge B_d = S_d \ltimes \\pm 1\^d. Thus permutation-only alignment is symmetry-incomplete for RMSNorm models. We introduce sign-marginalized Hungarian matching and prove a sharp population failure mode: with decorrelated source coordinates, raw signed-correlation matching has a structural permutation-accuracy ceiling equal to the fraction of positive signs in the true gauge up to O(d \cdot 2^-d), whereas sign-marginalized matching removes this obstruction. We then make coordinate-preserving transport, rather than function-level merging, the primary object: composing saved-checkpoint local B_d gauges along same-base fine-tuning trajectories recovers 91.1% of cross-run coordinates at 1500 steps versus 60.3% for endpoint matching, and the gain is not explained by merely routing through the base. The recovered gauge transfers tools that permutation-only alignment breaks: TinyLlama SAE reconstruction has NMSE 0.004 under B_d recovery versus 1.08 under S_d; Qwen sentiment steering preserves 95.8% of its effect versus 17.2%; refusal steering reverses sign under S_d. Coordinate-preserving merge tests show the same mechanism. The same covariance governs stateful training: signed transport of AdamW state preserves the resumed trajectory, while permutation-only state transport starts from a functionally identical checkpoint but follows a different trajectory. Finally, we give gauge-sweep audits for index-level interpretability claims: coordinate names are reproducible only relative to an explicit gauge.
Abstract: Multiobjective submodular maximization asks for a single set that performs well across multiple submodular objectives, a setting motivated by robust experimental design and fairness-aware decision making. The standard max-min formulation, however, can be overly conservative: it is determined entirely by the worst-performing objective and ignores improvements in the remaining objectives unless they change the minimum. Motivated by concave social-welfare aggregators from economics, we initiate the study of multiobjective submodular maximization with general concave aggregation. We propose a randomized greedy algorithm that, in each iteration, computes a distribution over elements by solving a concave program and samples an element from this distribution. Our analysis overcomes the loss of the objective-wise decomposition available in the max-min case and proves an asymptotic (1-1/e-\epsilon)-approximation under mild concentration assumptions. To make the method scalable, we develop a Fenchel-duality-based lazy evaluation scheme. Experiments on synthetic and real-world instances show that our algorithm consistently improves over a naive greedy baseline across a variety of concave aggregators, while enabling utility--fairness trade-offs that are not captured by the max-min formulation.
PaperID: 1613, Poster
Abstract: Long-horizon tasks are an important challenge for modern AI and robotics, yet pose a challenge for reinforcement learning. A historically popular approach has been to decompose a task into subgoals, and learn a subgoal-conditioned policy to execute each subtask in sequence. We introduce an alternative called Value Disaggregation (VaDar), which trains a single subgoal-independent policy to maximize a linear combination of per-subgoal critic values. Because subgoal information is needed only at training-time, VaDar removes the need for possibly costly or cumbersome subgoal-selection during policy deployment. Moreover, through careful experiments, we show that VaDar (1) is always competitive with, and frequently outperforms, goal-conditioning and other natural baselines and (2) succeeds on tasks with non-sequential goal structure and multiple optimization objectives. In particular, VaDar is the first algorithm to solve challenging, long-horizon tasks in the Robocasa suite, even when starting from a policy pretrained from behavior cloning with near-zero initial task success. Taken together, VaDar's success suggests that subgoal information benefits reinforcement learning by enhancing the efficacy of critic learning, whereas policy goal-conditioning is often unnecessary.
Abstract: We introduce EvoLib, a test-time learning framework that enables large language models to accumulate, reuse, and evolve knowledge across problem instances without parameter updates or external supervision. Instead of adapting model parameters, our approach maintains a shared library of knowledge abstractions, including modular skills and reflective insights, automatically extracted from the model's own inference trajectories. To support continual improvement, we introduce a principled weighting and consolidation mechanism that jointly optimizes for immediate utility and long-term value. This allows simple, instance-specific abstractions to evolve into more general and reusable ones over time. Across challenging benchmarks in mathematical reasoning, code generation, and multi-turn agentic environments, EvoLib improves substantially over the top test-time scaling and learning methods without ground-truth feedback.
PaperID: 1615, Poster
Abstract: Diffusion models achieve strong generative performance but often struggle to align generated samples with human preferences measured by a reward model. A simple yet effective algorithm for test-time alignment is Best-of-N (BoN) sampling, which draws N i.i.d. samples from a pre-trained diffusion model and outputs the single highest-reward sample. Despite its empirical success, this procedure is inherently post hoc: reward information is used only for selection after generation, not for improving the reverse diffusion process itself. Consequently, BoN sampling does not improve the average alignment of generated samples and is primarily suited to single-output settings. We propose Best-of-N Guidance (BoNG), a novel method that integrates the principle of BoN sampling directly into the reverse diffusion process. BoNG performs online BoN selection over denoising particles and adjusts the reverse diffusion process to steer the particle population toward higher-reward regions during generation. Specifically, by introducing an asymmetric guidance interaction among denoising particles, BoNG uses the current BoN particle as a guidance signal to the rest of the particle population. This particle-level interaction reshapes the sampling process toward higher-reward regions, enabling BoNG to improve not only the final best sample beyond Vanilla BoN sampling, but also the average quality of generated samples. Over 36 empirical comparisons, BoNG achieves the best performance in 29 cases, ranking first in 80.56% of the comparisons against SMC and Vanilla BoN sampling. BoNG also supports multi-output capability, showing 1.3× ImageReward score than the latest sample-based guidance method with 1.6× speedup.
Abstract: Online resource allocation (ORA) is a fundamental framework for sequential decision-making problems under budget constraints, with applications ranging from online advertising to revenue management. In this work, we study a broader setting that includes both budget constraints and general constraints, extending the classical budget-only model. This extension is essential for modelling critical economic requirements, such as Return-on-Investment (ROI) constraints. We develop an algorithm that achieves best-of-both-world guarantees within this generalized framework. In particular, against a dynamic benchmark, our algorithm achieves \widetilde\mathcal O(\sqrtT) regret in the stochastic regime and \alpha-regret of order \widetilde\mathcal O(\sqrtT) in the adversarial regime, where \alpha depends on the feasibility margin of the corresponding offline problem. At the same time, our algorithm guarantees strict satisfaction of the budget constraints and \widetilde\mathcal O(\sqrtT) cumulative violation for the general ones. From a technical perspective, introducing general constraints alongside budgets precludes the use of standard budget-focus methods. While budget methods rely on a zero-consumption "safe'' action to ensure feasibility, general constraints are much less "aligned'' towards feasibility. We overcome these difficulties with a new analysis that exploits weak adaptivity to get boundedness of the Lagrangian multipliers and best-of-both-world guarantees.
Authors:
Jan F Meier, Felix B Mueller, Alexander Ecker, Timo LüddeckeAbstract: Modern action recognition models operate on memory- and compute-intensive dense RGB video volumes and frequently exploit appearance and background shortcuts, for example, predicting actions from objects or scenes instead of characteristic motion. We investigate an alternative input modality that is largely free of such biases by construction: sparse point trajectories. To this end, we develop a simple transformer architecture for 2.5D trajectory-based recognition together with a masked-trajectory pretraining, which we show to substantially improve downstream accuracy. Despite using only a fraction of the dense RGB input, our method reaches 45% top-1 on Something-Something V2 and 54% on EPIC-Kitchens-100, and surpasses V-JEPA 2 on time-reversal sensitivity. More importantly, we find trajectory features to be complementary to state-of-the-art appearance-based features: fusing our pretrained model with DINOv2 and V-JEPA 2 improves top-1 accuracy on Something-Something V2 by 8.7 and 1.6 points, respectively.
PaperID: 1618, Poster
Authors:
Wenxin Zhao, Letian Tao, Zhilong Zheng, Guojian Zhan, Yujie Yang, Likun Wang, yinuo Wang, Feihong Zhang, Tianze Zhu, Jingliang Duan, Yang Guan, Shengbo Eben LiAbstract: Fine-tuning continuous-time generative models, such as flow models, via reinforcement learning (RL) is a promising approach for complex continuous control, but directly updating the network weights often leads to catastrophic forgetting of the pre-trained prior. Latent steering method mitigates this by freezing the generative backbone and optimizing a latent policy to generate the initial noise. However, existing methods typically enforce hard support boundaries to prevent out-of-distribution (OOD) shifts in the latent space, inducing a fundamental trade-off between prior preservation and policy expressiveness. To resolve this dilemma, we propose Moment Constraint Parameterization, which regulates the latent distribution through a statistical budget. By decoupling this budget into inter-statistic constraints and intra-statistic flexibility, our method acts as a probabilistic safeguard against OOD shifts while providing the degrees of freedom to retain policy expressiveness. Furthermore, while moment constraints define a flexible search space, effectively optimizing the latent policy within it requires accurate gradient signals. To address the optimization lag and bias caused by proxy latent critics in existing methods, we introduce Adjoint Q-Gradient Propagation. By exploiting the white-box structure of flow models, this technique backpropagates exact analytical gradients from the action-space critic directly to the initial noise. Integrating these two mechanisms, we present Moment Constraint Flow Steering (MCFS). Empirical results on D4RL and OGBench show that MCFS consistently delivers stronger online adaptation performance than prior latent steering and flow-policy baselines, with particularly pronounced gains on challenging long-horizon compositional manipulation tasks.
Abstract: Current diffusion models for audio-driven avatar video generation struggle to synthesize long videos with natural audio synchronization and identity consistency. This paper presents StableAvatar, the first end-to-end video diffusion transformer that synthesizes ultra-long videos. We observe that the primary reason preventing existing models from generating long videos is their audio modeling, which typically relies on third-party off-the-shelf audio embeddings that are introduced into the diffusion model via cross-attention. Since current diffusion backbones lack any audio-related priors, they cannot correctly couple audio cues with latent dynamics, causing systematic latent distribution deviations at each segment. These deviations accumulate across segments, gradually shifting the latent trajectory and producing noticeable quality drift. To address this, StableAvatar introduces a novel Timestep-aware Audio Adapter that prevents error accumulation via timestep-aware modulation. During inference, we propose a novel Audio Native Guidance Mechanism to further enhance the audio synchronization by leveraging the diffusion’s joint audio-latent prediction as a dynamic guidance signal. To enhance the smoothness of the ultra-long videos, we introduce a Dynamic Weighted Sliding-window Strategy that fuses latent over time. Experiments on benchmarks show the effectiveness of StableAvatar both qualitatively and quantitatively.
PaperID: 1620, Poster
Abstract: At deployment, a multivariate forecaster may expose only a recent input window and a frozen point prediction, with no access to ground truth, calibration residuals, repeated inference, or model internals. We study deployment-constrained forecast risk localization: ranking future timesteps by likely error before they are observed. We introduce SubDx, a zero-parameter diagnostic that audits the forecast against the local cross-channel subspace of the input history. Each predicted channel vector receives an off-subspace residual score, computed with no labels, no learned parameters, and less than 1% runtime overhead on Traffic. Under a linear factor model, we derive closed-form score distributions and a two-population AUROC expression showing how heterogeneous off-subspace error visibility makes the ranking identifiable. On Traffic (N=862), SubDx reaches within-horizon AUROC 0.90 and reduces MSE by 62% at 50% coverage; on Solar (N=137), it reaches AUROC 0.85 with 67% MSE reduction. The same post-hoc score applies to frozen pretrained forecasters, with Chronos reaching AUROC 0.838 on Traffic without fine-tuning. Across trained backbones, pretrained models, and synthetic controls, the results show that local cross-channel consistency can provide a practical abstention signal for redundant multivariate systems.
Authors:
Hanato Kikuchi, Ryosuke Masuya, Kazuhiko Kawamoto, Hiroshi KeraAbstract: Learning parity functions, more general modular addition, is a challenging machine learning task due to its input sensitivity. A recent study substantially scaled modular addition learning in both the number of summands and the modulus. Its key idea is to increase zeros in training sequences, reducing the effective number of summands and thus controlling training difficulty; however, this induces covariate shift between training and test input distributions. This study theoretically and empirically analyzes this side effect and proposes a covariate-shift-free method for modular addition. Specifically, we introduce an auxiliary modulus Kq during training, which reduces wrap-around frequency and problem difficulty while preserving the same input distribution across training and testing. Experiments show strong scalability and sample efficiency: even for large input length N, large modulus q, and small datasets---where the sparse method fails to learn---our method achieves equal or better match accuracy and relaxed \tau-accuracy. For example, at N=64 and q=974269, our method trained on 100K samples achieves 97.0% \tau-accuracy at \tau=0.05, while the sparse method achieves only 9.5% with the same data size and 93.9% even when extended to 1M samples.
Abstract: We study feature subsampling in greedy model search, a randomization mechanism motivated by random forests and other ensemble methods. While recent theory suggests that this randomization acts solely as a variance reduction mechanism analogous to ridge regularization, these results largely rely on base learners optimized via ordinary least squares (OLS). We investigate the effects of feature subsampling on greedy forward selection, a tractable abstraction of the adaptive split search used by decision trees. Assuming an orthogonal design, we prove that ensembling with feature subsampling can reduce both bias and variance, contrasting with the pure variance reduction of convex base learners. Specifically, we show that both the training error and degrees of freedom need not be monotone in the subsampling rate, breaking the analogy with standard shrinkage methods like the lasso or ridge regression. Furthermore, we characterize the exact asymptotic behavior of the estimator, showing that it adaptively reweights OLS coefficients based on their rank, with weights that are well approximated by a logistic function. These are mechanism level results for a tractable analogue of greedy selection that depends on the response, not performance guarantees for full random forests.
PaperID: 1623, Poster
Abstract: Inference-time scaling provides a promising path for improving agentic reasoning performance, and one widely adopted form is sample-then-select: first sample a pool of candidate reasoning trajectories, then select from the resulting candidates. While this keeps scaling lightweight without updating the policy, it exposes a training-deployment mismatch. Standard trajectory-wise Reinforcement Learning (RL) objectives optimize normalized signals for individual trajectories, whereas choosing from a large candidate pool with the policy's own selection signal is inherently set-conditional: it requires calibrated ranking and confidence aggregation across correlated trajectories. To address this challenge, we propose AutoPortfolio, a set-conditional policy alignment framework for efficient and effective sample-then-select scaling. AutoPortfolio builds a tractable environment-informed target on the sampled candidates, and uses its induced Plackett-Luce (PL) ranking to calibrate the policy which trajectories to prioritize within that set. It trains the policy with a balanced PL-motivated objective that reallocates probability mass among competing trajectories, preserving reward-aligned alternatives while sharpening away from low-quality distractors. Aligned with our policy training, our lightweight inference-time Mass Aggregation scaling adaptively perceives confidence from candidate trajectories and selects the most promising answer(s). Theoretically, we show that reducing our set-conditional loss narrows the discrepancy between policy-induced and environment-informed PL rankings on the sampled set. Empirically, across 12 challenging reasoning benchmarks, AutoPortfolio improves over strong agentic RL baselines, with the clearest gains in larger candidate pools where the set-level selection is crucial.
Abstract: The distributed setting for Saddle Problems (SPs) has recently emerged as a framework for modern applications in machine learning and multiagent systems. Despite its relevance, the theoretical foundations of this setting have not yet been thoroughly established. In this paper, we advance this research direction by formalizing the distributed setup for SPs and providing rigorous definitions of communication and oracle costs. Further, we prove lower bounds for any distributed gradient-span algorithm, which reveals the gap from existing methods and this theoretical limit. To this end, we provide a Decoupled Method built upon a novel multi-stage reduction that reduces the SP into a sequence of decoupled minimization tasks of residual norms. Our algorithm matches the communication lower bound, thus setting the communication complexity within the gradient-span algorithms. Moreover, it yields the first strict improvement over the long-standing oracle cost of the Extragradient method for general SPs. Finally, we study the extension of distributed SP into Variational Inequality Problem (VIP), which generalizes two-player zero-sum games to multiplayer general-sum games. We show that our Decoupled Method achieves a new state-of-the-art communication complexity for this broader class.
PaperID: 1625, Poster
Abstract: We address geometric video alignment for the same scene captured by multiple freely moving cameras. Existing methods, which are based on image stitching or video stabilization, are typically limited to two cameras or must assume the restrictive setting of a fixed linear ordering of the cameras. We propose MViS, a lightweight test-time optimization framework for joint multi-view alignment with temporal consistency. MViS uses a spatial-transformer-inspired net to jointly process frames from all views within a short temporal window and predict per-view, per-time homographies. Optimization proceeds in two stages: geometry-guided alignment followed by appearance-based refinement, with both stages enforcing multi-view and temporal consistency across the sequence. The method requires no scene-specific hyperparameter tuning, handles arbitrary and time-varying view-intersection graphs, and remains robust under rapid camera motion. On a standard two-view video benchmark, MViS matches or surpasses the state of the art. To facilitate evaluation in more complex scenarios, we introduce an annotated dataset with more than two freely moving cameras and establish the first benchmark for this setting. Our code and dataset will be released upon acceptance.
Abstract: Yes. We find that large multimodal models develop mental imagery when solving spatial puzzles, and they do imagine sheep when solving sheep puzzles. We fine- tune a Qwen3.5 VLM to solve twelve diverse visual reasoning tasks—including tangram, jigsaw, sokoban, 3D mental rotation, and rush hour—that require under- standing geometry, spatial constraints, and the consequences of actions. By simply predicting the open-loop sequence of actions to solve a puzzle from an initial state, we show that the model learns successfully, and that its hidden representations after each action encode meaningful visual information about the intermediate puzzle state. This finding suggests that an imperfect visual world model begins to form as a byproduct of learning to select correct actions in an open-loop fashion, in the absence of any explicit visual supervison. Building on this observation, we propose two ways to sharpen the mental images learned by the model through explicit visual supervision or explicit visual tokens. Across puzzles, we find that integrating as few as sixteen visual tokens into the chain of thought per board state improves the average solve rate from 83% to 89%, with gains exceeding 20 percentage points on reasoning-heavy games such as jigsaw and 3D mental rotation. Our results show that visual reasoning is a bottleneck, and that mental imagery may be a key mechanism for robust spatial cognition.
Abstract: Adapting pretrained diffusion models to downstream objectives such as inverse problems often requires expensive test-time guidance or optimization. We propose a principled framework for generating high-quality reward-aligned samples at substantially reduced inference cost. Our approach formulates test-time adaptation as a hierarchical variational model, where control is amortized into a lightweight yet expressive stochastic policy. This formulation naturally supports few-step diffusion sampling: large step sizes enable fast inference, while the learned policy maintains sample quality by providing structured per-step control. The resulting fully amortized sampler achieves a strong quality--speed tradeoff, matching or exceeding recent test-time scaling baselines while requiring significantly less compute. For example, on 4× super-resolution, our method achieves better perceptual quality with more than 5× faster inference compared to the best-performing baseline. We further extend our approach to a semi-amortized regime that combines cheap amortized proposals with limited test-time optimization, achieving state-of-the-art perceptual quality across several challenging inverse problems.
PaperID: 1628, Poster
Abstract: Peer review is sometimes seen as noisy and difficult to predict. How much of the final accept/reject decision is recoverable from the paper itself? We investigate this question with PaperLens, text and vision models trained to predict binary acceptance from anonymized paper artifacts. For a clean evaluation, we construct balanced, shortcut-controlled OpenReview and source-derived arXiv datasets that remove deanonymization artifacts and majority-class shortcuts. Supervised binary decision training is surprisingly strong. PaperLens outperforms frontier prompting and review-trained baselines on decision accuracy and ranking, while its acceptance probability remains well calibrated after validation-set scaling and correlates better to ratings than baselines. Vision consistently improves over markdown text, showing that layout, figures, and visual presentation carry reviewer-relevant signal. Scaling from 3B to 14B parameters yields only modest gains, with clean supervision and faithful paper representations mattering more than model size. Together, our results show that decision prediction provides a powerful signal for paper assessment and can even improve the alignment of AI reviewing agents.
PaperID: 1629, Poster
Abstract: U-Net--style architectures are widely adopted for modeling physical systems, as their multiscale structure enables efficient processing of high-resolution data while reflecting the hierarchical structure of many physical phenomena. The most successful U-Net-based neural physics simulators often combine convolutions, Fourier layers, or specialized transformer mechanisms such as windowed or axial attention; however, these structured components can limit adaptability across different spatial and spatiotemporal dimensions. In this work, we revisit transformer-based U-Nets with the goal of simplifying attention-based neural physics simulators without sacrificing multiscale efficiency or predictive accuracy. We introduce \textttUFlex, a simple attention-based U-Net whose blocks are obtained through only minimal modifications to standard self-attention, allowing the same architecture to be applied across one-, two-, three-dimensional, and spatiotemporal regular-grid problems without major design changes. Evaluated on seven challenging benchmarks (four 2D and three 3D), \textttUFlex scales to resolutions of up to 512 × 512 in 2D and 256 × 128 × 256 in 3D, while reducing training memory and accelerating training compared to state-of-the-art transformer baselines, all while achieving state-of-the-art predictive accuracy.
Abstract: There is an extensive literature that studies how to find optimal policies in resource allocation problems, taking the underlying design parameters that define the allocation, such as what data is collected, how many people can be served, and quality of service as fixed constraints. Yet, from a planner's perspective, these design parameters are themselves optimization variables that are just as important in determining overall welfare as selecting the optimal targeting rule for a given set of constraints. This realization motivates a rich set of meta-design questions exploring how planners should make principled decisions about investments in prediction, capacity constraints, and treatment quality, all of which lie upstream of classical policy optimization. Building on initial theoretical work in this space, our paper has three main contributions. First, we formally define the broad meta-design space of resource allocation problems. Second, we develop empirical tools that enable practitioners to reliably navigate it. Third, we demonstrate the framework in two real-world case studies on German employment services and targeted cash transfer programs in Ethiopia.
PaperID: 1631, Poster
Abstract: Model merging integrates task-specific fine-tuned models into a single multi-task model, but often suffers from parameter interference caused by conflicting task-vector updates. Existing methods typically mitigate conflicts by pruning task vectors based on weight magnitude or random heuristics, treating Transformers as unstructured ``bags of parameters'' and overlooking their inherent modularity. In this paper, we propose Contribution-Aware Structured Sparsity (CASS), a unified framework that reduces parameter interference by identifying and preserving task-specific components. At the core of CASS is a contribution-aware structured mask that identifies task-relevant attention heads and FFN neurons. We instantiate this mask in two settings: CASS-Merging, the primary post-hoc setting where masks serve as a plug-and-play denoising filter for existing merging operators, and CASS-Tuning, an extension for scenarios with fine-tuning access where masks constrain gradients to reduce structural overlap between task vectors. Our analysis shows that task-relevant components are sparse and partially disjoint, supporting structured component-level filtering as an effective way to reduce merging interference. Extensive experiments across vision (ViT, 20 tasks) and language (RoBERTa, 8 tasks; Qwen2.5, 4 tasks) benchmarks demonstrate that CASS improves a range of representative merging baselines.
PaperID: 1632, Poster
Authors: Nevena Gligić, Arya Farahi
Abstract: Modern data analysis pipelines increasingly rely on large datasets assembled through retrieval, scraping, automated filtering, and post-hoc curation, where the observed sample is often a contaminated mixture rather than the target population. Standard distributional objectives either ignore contamination and converge to the wrong target, or retain only a small trusted subset and discard most of the data. We study supervised and unsupervised distributional learning when a clean reference sample is compared with a pooled sample from r=(1-\epsilon)q+\epsilon z, where q is the latent target distribution and z is an unknown contaminant. Motivated by settings with noisy membership scores for all observations and exact labels for a small audited subset, we propose contamination-corrected maximum mean discrepancy (CC–MMD), a design-based estimator that combines proxy scores with audited residual corrections to recover the oracle MMD that would have been computed from the latent clean sample. CC–MMD requires no calibration assumption on the proxy scores and no structural assumptions on the contaminant beyond standard kernel regularity. We prove finite-population design-unbiasedness of the corrected kernel sums, establish asymptotic normality of the audit correction, and show consistency for the target discrepancy \mathrmMMD^2(p,q;k). Empirically, CC–MMD tracks the oracle under increasing contamination, improves parameter recovery and prediction in contaminated regression, and enables generative models trained on heavily corrupted mixtures to recover the clean target distribution. These results show that contaminated data need not be discarded or treated as the estimand. With sparse trusted labels and abundant noisy supervision, distributional objectives can be made robust by design.
Abstract: Fast weight architectures offer a promising alternative to standard transformers for long-context modeling by replacing the KV-cache with a recurrently updated fixed-size memory. Despite this architectural shift, they are typically trained with the same next-token prediction (NTP) objective as standard transformers. This creates a mismatch: fast weight models rely on an evolving memory state to support future predictions, while NTP provides only token-level supervision for the immediate next token. We address this mismatch with next-sequence prediction (NSP), a sequence-level extension of NTP that trains models to produce coherent multi-token continuations from their recurrent memory state. To optimize this sequence-level objective, we propose ReFINE (Reinforced Fast weIghts via Next sEquence prediction). ReFINE selects informative token positions based on prediction entropy, generates multi-token rollouts, assigns sequence-level rewards, and optimizes the model with Group Relative Policy Optimization (GRPO). ReFINE is applicable throughout the training lifecycle of pre-trained language models: mid-training, post-training, and test-time training. Experiments on LaCT-760M and DeltaNet-1.3B show that ReFINE consistently outperforms NTP-based supervised fine-tuning in needle-in-a-haystack retrieval, long-context question answering, and diverse tasks in LongBench. ReFINE provides an effective and versatile framework for improving long-context modeling in fast weight architectures.
Authors: Yeping Jin, Jiaming Hu, Yannis Paschalidis
Abstract: Large Language Models (LLMs) tend to respond correctly to prompts that align well with the data they were trained and fine-tuned on. Yet, small shifts in wording, format, or language can trigger surprisingly large failures, especially on multi-step reasoning problems. To address this problem, we propose a Distributionally Robust Token Optimization (DRTO) approach, which combines token-level Reinforcement Learning from Human Feedback (RLHF) with Distributionally Robust Optimization (DRO). DRTO constructs f-divergence ambiguity sets over span-level actor losses, providing a principled way to emphasize difficult response segments during policy optimization. Empirically, DRTO enhances consistency under distribution shifts in multiple reasoning benchmarks among different tasks.
PaperID: 1635, Poster
Abstract: Most explanations of weak-to-strong (W2S) generalization focus on why a strong student may fail to exactly imitate weak-label errors. We study a complementary mechanism: before student training, weak-teacher predictions may be inconsistent along semantic constraints specified independently of the weak model. When such constraints are valid for the target, these inconsistencies give label-free lower bounds on the risk reduction achievable by correcting weak labels, rather than merely indicating unstructured prediction error. We formalize this mechanism using Markov and Laplacian operators. Under exact target validity, Markov trajectory inconsistency equals the squared-risk gain from smoothing; under approximate validity, it remains a conservative lower bound after an explicit validity penalty. We then solve the associated minimax correction problem, obtaining a resolvent correction that continuously attenuates constraint-violating directions and reduces to the identity when validity uncertainty is too large. These results yield validity-adjusted correction (VAC), a split-sample pipeline that selects corrections from weak-teacher queries and predeclared validity budgets before training the strong student. Controlled graph-signal experiments and a CIFAR-10 cats-versus-dogs W2S experiment are consistent with the proposed mechanism: validity-adjusted inconsistency predicts useful correction and transfers to student gain, whereas high raw inconsistency under invalid controls does not provide such evidence.
Abstract: Black-box optimization in science and engineering often comes with side information: experts, simulators, pretrained predictors, or heuristics can suggest which candidates look promising. This information can accelerate search, but it can also be biased, input-dependent, or misleading. Feedback-aware BO methods typically handle one task at a time, limiting their ability to generalize over multiple sources of feedback. In-context optimizers address cross-task adaptation, but usually assume that optimization history is the only available signal at test time. We study feedback-informed in-context black-box optimization (FICBO), where a pretrained optimizer conditions on both the observed history and cheap auxiliary feedback for the current candidate set. We introduce a structured feedback prior that models how feedback sources vary in their access, relevance, and distortion relative to the true objective, and use it to pretrain a feedback-aware transformer. At test time, the model estimates source reliability in context by comparing observed objective values with auxiliary signals, improving query selection. On synthetic and real-world tasks, FICBO effectively exploits informative feedback while remaining robust to weak or misleading sources, improving over other baselines. Empirical investigations further illustrate how the model perceives test-time sources, offering insights into its interpretability and decision-making process.
Authors: Longwei Zou, Lin Zhong
Abstract: Linear attention has recently gained significant attention for long-context inference due to its constant decoding cost with respect to context length. However, existing serving systems typically serve linear attention by recurrently computing and updating a large linear attention state in every decoding step. Since the state is much larger than the per-token key and value, recurrent decoding incurs substantial memory access and becomes inefficient for serving linear attention. In this paper, we propose KVBuffer, an IO-aware serving mechanism for linear attention. By buffering recent keys and values, KVBuffer enables serving systems to compute linear attention outputs in more flexible and memory-efficient ways. For decoding, KVBuffer enables chunkwise computation, which reduces average memory access and decoding latency by deferring state updates and applying them in batch. For speculative decoding, KVBuffer verifies draft tokens in parallel and avoids storing temporary states. For short contexts, KVBuffer computes attention outputs directly from buffered keys and values, without creating or updating the linear attention state. We implement KVBuffer in SGLang for Qwen3-Next. Our evaluations show that KVBuffer can reduce linear attention decoding latency by up to 45.17% and increase the maximum number of serving requests by 5 × for speculative decoding when verifying four draft tokens.
Abstract: Artificially generated speech is increasingly embedded in everyday life. Voice cloning in particular enables applications where identity preservation is important, such as completing a recording, dubbing in a new language, or preserving the voices of individuals with speech loss. However, in our work, we find that despite the term, voice cloning does not faithfully "clone" an individual's voice. Instead, we find that widely-used voice cloning models systematically apply style transfer to source voices. As rated by human annotators, cloned voices are perceived as more authoritative, warm, customer-service-like, and human-like compared to their sources. Human annotators also report greater trust in cloned voices than source voices, and a greater willingness to disclose sensitive personal information to them. Our work furthermore shows that voice cloning leads to homogenization of speaker characteristics, as measured by reduced variance in accent, speaking rate, and the audio embedding space. Together, our results highlight a new set of limitations and risks of voice cloning technology and their potential impact on human behavior.
PaperID: 1639, Poster
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become an important approach for improving the reasoning capabilities of large language models (LLMs). On-policy training remains the dominant RLVR paradigm, but its reliance on fresh rollouts leads to poor data efficiency. To address this bottleneck, off-policy paradigm has attracted increasing attention by reusing historical trajectories through replay buffers. Despite its efficiency advantages, off-policy training often struggles to match the performance of on-policy training. Our analysis reveals that this limitation primarily stems from data staleness introduced by replay buffers: to maintain training stability, mainstream algorithms discard a large fraction of high-staleness optimization signals, leading to substantial performance degradation. To address this issue, we propose eplay), a data-centric replay framework decoupled from specific optimization algorithms. STAR improves off-policy RLVR by constructing low-staleness and high-quality training data. Experiments on a range of mathematical reasoning and code generation tasks show that STAR consistently improves off-policy RLVR with up to 19% relative performance gains, and enables it to surpass the on-policy baseline using only around 13% of the rollout data while achieving 5x-6x wall-clock speedup.
PaperID: 1640, Poster
Abstract: Mask-guided CT synthesis bridges structural annotations with imaging appearance, facilitating data augmentation and clinical tasks such as radiotherapy planning. Despite plausible visual realism, current methods seek optimal solutions by treating all tissues uniformly, failing to capture intrinsic CT properties such as tissue-specific distributions and strict anatomy, leading to semantic misalignment and structural hallucinations. To address this, we decompose CT synthesis into semantics-structure alignment through the proposed Semantics and Structure Loss. Specifically, the Semantic Intensity Distribution (SID) Loss partitions the broad Hounsfield Unit (HU) range into semantic subspaces to achieve tissue-specific intensity alignment. Building on SID-based semantic alignment, the Structural Anatomy (SA) Loss further improves structural fidelity by regularizing gradient manifolds within clinically-informed spatial domains. Extensive experiments across GAN, Diffusion, Flow, and foundation models demonstrate the broad compatibility and consistent gains of our approach. A user study further confirms the clinical realism and mask fidelity, while downstream segmentation demonstrates the practical utility. Code will be released after acceptance.
PaperID: 1641, Poster
Abstract: Unified language models are increasingly expected to combine heterogeneous capabilities, such as mathematics, code, instruction following, and controllable thinking behavior, within a single set of parameters. A common solution is sequential post-training on multiple objectives, but this entangles all objectives along one optimization trajectory and makes the final model highly sensitive to training order, data ratios, schedules, and stopping criteria. Weight-space merging offers a modular alternative, but naive merging of single-objective experts often fails: domain capabilities degrade sharply, or think/non-think modes collapse into one dominant behavior. We attribute both failures to incompatible weight-space geometry: experts trained on isolated objectives drift to distant regions of parameter space, placing their interpolations outside any shared low-loss basin. We propose Mixture-Trained Merging (MTM), which trains each branch on an objective-biased data mixture rather than a single objective, exposing it to cross-objective interactions and making branches compatible at merge time. MTM uses merged-model evaluations as a low-cost signal for selecting branch mixtures, avoiding expensive data-mixture ablations. The procedure is iterative: each round promotes the base model using globally selected merge coefficients and refines each branch mixture using domain-preferred coefficients under constraints that preserve other objectives. To scale beyond simplex grid search, MTM uses qNEHVI-based multi-objective Bayesian optimization. Across code, mathematics, instruction following, and think/non-think control, MTM outperforms naive merging and preserves behavioral separation where single-objective merging collapses, suggesting that effective unified models require branches trained to be mergeable.
Abstract: Classic zeroth-order optimization approaches typically optimize for a smoothed version of the original function, i.e., the expected objective under randomly perturbed model parameters. This can be interpreted as encouraging the averaged loss values in the perturbation set to be small. However, popular sharpness-aware minimization objectives typically focus on the largest loss within the neighborhood to arrive at flat minima more effectively. In this work, we connect zeroth-order optimization (and its corresponding objectives) with SAM approaches explicitly, through an exponential tilting objective that provides a smooth transition between the \textttaverage- and the \textttmax-loss formulations. We explore new zeroth-order algorithms to solve a soft SAM objective parameterized by a tilting parameter t. We theoretically analyze the sharpness of the tilted objective under different perturbations. Practically, our approach can be used as a gradient-free and memory-efficient alternative to SAM variants, and it achieves better generalization compared to vanilla zeroth-order baselines on a wide range of models and downstream tasks, including classification, multiple choice QA, and language generation.
Abstract: Deep learning systems are known to exhibit implicit regularization (alt. implicit bias), favoring simple solutions instead of merely minimizing the loss function. In some cases, we can analytically derive the implicit regularization -- connecting it to an equivalent penalty that augments the learning objective. However, modern deep learning systems are complex, carrying modifications to the training procedure and architecture (e.g. early stopping, minibatching, dropout) whose effects are not always directly interpretable. Although estimating the resulting implicit regularization could aid theorists in algorithm design and practitioners in interpreting their hyperparameter choices, this problem has received little direct attention. It is also tractable: regularization makes weight updates deviate from loss gradients, promising a signal for identifying implicit bias. Here we provide gradient matching methods that can be used to empirically estimate the implicit regularization. Our method works on networks with known regularization, recovering popular explicit penalties like \ell_1 and \ell_2. It also replicates known implicit effects, like the quadratic weight penalty induced by early stopping in gradient descent, demonstrating that it can be used to test theories of implicit regularization. Crucially, because our method is empirical, it can handle implicit regularization in arbitrary networks. We demonstrate this use by characterizing the effects of dropout in deep networks, showing implicit \ell_2 effects in this popular method. Our work shows that practitioners can use gradient matching to understand regularization in networks with implicit biases that are too complicated to derive analytically.
PaperID: 1644, Poster
Abstract: Electromagnetic inverse scattering (EIS), which aims to recover object permittivity from measured scattered fields, is central to a wide range of applications, from medical diagnostics to security screening. However, it is a highly nonlinear and ill-posed problem. Despite substantial progress, existing approaches face a fundamental trade-off between physical fidelity, efficiency, and generalization: data-driven methods are fast but tied to training geometries and sensor setups, while physics-based optimization is accurate but computationally expensive. We introduce a Gaussian-Galerkin (GauGal) method that reformulates EIS from a dense point-wise field reconstruction to a compact, primitive-level physical solving. Rather than using Gaussian primitive merely as a material representation, GauGal take it as the computational units for wave scattering. By projecting the continuous EIS scattering equation into a Gaussian primitive space, the material modulation, Green propagation, source excitation, and receiver observation are all realized as primitive-level operators. This yields a compact differentiable forward solver for physics-consistent reconstruction. Our method achieves state-of-the-art accuracy on synthetic and real benchmarks while reducing runtime from over 30 minutes by leading physics-driven baselines to under 10 seconds. It also generalizes significantly better to unseen geometries and sensor configurations than leading data-driven baselines. Overall, this framework establishes an accurate, efficient, and generalizable paradigm for electromagnetic inverse imaging, enabling fast, physics-consistent imaging in practical settings.
PaperID: 1645, Poster
Abstract: Diffusion policies have achieved strong performance in reinforcement learning (RL) due to the exceptional expressivity in capturing complex action distributions. However, existing approaches focus primarily on positive samples and lack explicit modeling of negative samples, thereby failing to fully exploit this representation capacity and ultimately limiting performance. To mitigate this problem, we propose BOBAC (BOotstrapped Bipartite Actor-Critic), a novel method for diffusion policy optimization via pairwise preference learning. Bridging the gap between distribution modeling and preference learning, we reformulate diffusion policy optimization within a preference-aware framework via the Bradley-Terry (BT) modeling, which enables the effective exploitation of negative samples through relative ranking. Since the original BT loss assigns the same weight to all pairs regardless of their varying importances, we further design a soft-gate mechanism to adaptively weight each pair according to its Q-gap. Experimental results on 6 MuJoCo tasks and 4 H1Bench tasks demonstrate that, compared with representative model-free and generative-policy RL baselines, BOBAC achieves state-of-the-art performance with consistent improvements in all continuous-control environments.
PaperID: 1646, Poster
Abstract: Reinforcement-learning policies for robot locomotion can fail at deployment when operating conditions shift (e.g., rough or slippery terrain, payload changes, sensing degradation), and retraining or manual recalibration is often infeasible under on-device latency and memory constraints. We study test-time adaptation (TTA) for robotic control: adapting a pretrained policy online and self-supervisedly during execution, without access to rewards. Our starting point is a lightweight baseline that updates only the actor using a frozen pretrained critic as a deployment-time performance surrogate. Building on this, we introduce TEMPO, a test-time entropy-regularized, memory-guided policy optimization method that stabilizes critic-driven updates under severe shift. TEMPO combines (i) an entropy-promoting regularizer on the critic’s predictive distribution to mitigate overconfident extrapolation, and (ii) a score-based memory, inspired by prioritized replay, that prioritizes low-value (“hard”) observations and applies importance sampling to focus limited computation on the most informative deployment conditions. Across Go1, Go2, and G1 in MuJoCo Playground under four representative shifts and multiple architectures, TEMPO yields consistent gains, reaching up to 80% improvement in a single-device setting. We further validate on a physical Unitree Go2 with an added calf payload, where fully on-device adaptation improves performance over the non-adapted baseline across real-world episodes.
Abstract: We study the problem of constructing coresets on data streams: Given a dataset of n points in \mathbbR^d arriving sequentially, the goal is to maintain a small weighted subset that accurately approximates the objective values of the entire dataset for a given class of query functions. Previous methods like merge-and-reduce and online sensitivity sampling required more space than the size of an optimal such subset in the offline setting. In this work, we show that such overheads are not necessary by presenting space-optimal streaming algorithms for central high-dimensional optimization problems such as Euclidean (k,z)-clustering and L_p subspace embeddings, as well as applications to projection-cost preserving sketches. For (k,z)-clustering, our streaming algorithm uses memory independent of the number n of input points (in words of memory) and the aspect ratio \Delta, yielding a coreset with an optimal \tilde\mathcalO\left(\fracdk\min(\varepsilon^4,\varepsilon^z+2)\right) words of memory for accuracy parameter \varepsilon\in(0,1). For L_p subspace embeddings, where the rows of matrix \mathbfA \in \mathbbR^n × d arrive in an insertion-only stream, we achieve an efficient algorithm that uses optimal space for different ranges of p: \tilde\mathcalO\left(\fracd^2\varepsilon^2\right) words for p\le 2 and \tilde\mathcalO\left(\fracd^p/2+1\varepsilon^2\right) words for p>2, both independent of the number of rows n. Thus, our work shows that streaming algorithms can match offline algorithms in space complexity.
PaperID: 1648, Poster
Abstract: On-policy distillation (OPD) improves language-model post-training by letting the student learn from its own sampled trajectories under teacher supervision, aligning the training process with inference-time behavior. However, the rollout policy in OPD is usually treated as fixed: each round samples directly from the current student. We show that this choice can create a compounding teacher-student mismatch, because it requires the teacher to provide effective supervision on whichever trajectories the student samples, including ones less reachable by the teacher itself. On such trajectories, teacher supervision can become noisy or even misleading, leading to undesirable student patterns such as repetition. Moreover, since the updated student in turn defines the rollout policy of subsequent rounds, these patterns persist and potentially amplify over training. This failure mode suggests that improving OPD requires treating the rollout policy itself as an optimization target, rather than only modifying the distillation objective. We propose (MPOPD), which derives a student-anchored, teacher-guided rollout distribution from a local trust-region objective based on relative teacher support. The resulting target is a geometric mixture of the student and teacher policies, and we introduce a proposal policy that approximates this mixture, enabling rollout generation in a single forward pass. Experiments on 16 math, knowledge, and medical reasoning benchmarks show consistent gains over OPD and recent variants. On math reasoning, our method improves average accuracy by 6.91 over OPD on Qwen3-4B-Base while reducing average generation length by 41%, suggesting that rollout policy design provides an effective way to improve OPD.
Abstract: The Normalized Transformer, or nGPT (Loshchilov et al., 2025) achieves impressive training speedups and does not require weight decay or learning rate warmup. However, despite having hyperparameters that explicitly scale with model size, we observe that nGPT does not exhibit learning rate transfer across model dimension and token horizon. To rectify this, we combine numerical experiments with a principled use of alignment exponents (Everett et al., 2024) to revisit and modify the \muP approach to yperparameter transfer (Yang and Hu, 2021). The result is a novel nGPT parameterization we call \nuGPT. Through extensive empirical validation, we find \nuGPT exhibits learning rate transfer across width, depth, and token horizon.
PaperID: 1650, Poster
Authors: xizhe zhang
Abstract: Language models are increasingly used in text generation, decision support, and automated interaction, where their behavior must be controlled in a localized, reversible, and selective way. Existing methods either update shared parameters or constrain generation through external interfaces, leaving open whether frozen models contain compact internal control. We introduce sparse internal control, a framework for steering target behavior by applying inference-time interventions to a small set of internal model nodes. The framework formulates control through local target controllability, where candidate site–direction pairs are characterized by their behavioral effects under target, preservation, and energy constraints. This yields two complementary selectors: Driver-OMP, which selects nodes that sparsely reconstruct a desired behavioral displacement, and Coverage-Driver, which favors nodes whose effects cover prompts stably. We further propose a six-axis control-evaluation protocol measuring reachability, signed reversal, dose response, feedback controllability, off-target preservation, and held-out reuse. Across IOI, MMLU MCQA, and sentiment steering on models from GPT-2 small to Qwen3-4B, sparse driver sets act as executable actuators: they steer target behavior, respond monotonically to dose, support closed-loop control, and preserve unrelated next-token behavior. On refusal steering, the same protocol exposes a boundary case that passes all evaluated axes except strict reachability. Held-out reuse separates prompt-specific controls requiring refitted strengths from population-level controls that transfer as frozen interventions. These results show that model control can move beyond output-side steering: mechanistic internal variables can be selected, certified, and reused as localized control handles for frozen language models.
Abstract: Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spatial foundation models have enabled generalizable feed-forward 3D reconstruction, most streaming methods remain geometry-centric and lack temporally consistent object-level understanding. Meanwhile, existing semantic reconstruction and 3D-aware vision-language methods largely rely on externally extracted 2D semantic cues or loosely coupled geometry inputs, limiting unified geometry-instance learning in long dynamic scenes. In this paper, we propose IGGT4D, a streaming instance-grounded geometry Transformer for online 4D scene understanding. IGGT4D processes video frames sequentially, reuses historical context through causal spatial-temporal modeling, and incrementally updates a unified representation of camera motion, geometry, and object identity. This enables long-sequence feed-forward reconstruction with geometry-instance consistency in dynamic environments. To address the lack of high-quality 4D supervision, we further construct InsScene4D-147K, a large-scale dataset spanning real/synthetic and static/dynamic scenes, with RGB images, depth, poses, and temporally consistent instance masks generated by an automated geometry-guided annotation pipeline. Experiments on 3D reconstruction, pose estimation, instance spatial tracking, and open-vocabulary segmentation demonstrate that IGGT4D outperforms existing streaming baselines while maintaining scalable online inference for long dynamic sequences.
PaperID: 1652, Poster
Authors: Onur Ünlü, Eilyan Bitar, Francesca Parise
Abstract: We propose a fully decentralized Q-learning dynamic for infinite-horizon discounted Markov potential games in which agents observe the global state and their own realized payoffs, but do not know the reward functions, the transition probabilities, the potential function, or the policies and the actions of the other players. In the proposed dynamic each player maintains a local estimate of its continuation payoff for each state and action and updates its policy through a smoothed best response. Importantly, the smoothing parameter is adjusted by using an online estimate of the discounted state-visitation distribution, so that players explore more in rarely visited states while exploitation is encouraged in states that are frequently visited. We prove that the induced policy sequence converges almost surely to a set of approximate Nash equilibria of the Markov potential game, with an approximation guarantee that depends only on easily computable parameters. We illustrate the performance of the proposed learning dynamic via numerical experiments on a routing game and compare it with existing independent learning schemes based on persistent \theta-greedy exploration.
PaperID: 1653, Poster
Authors: John Hurwitz, Charles K Nicholas, Edward Raff
Abstract: We present an algorithm for margin-based online learning using lossless compressors (like zip) via the Minimum Description Length Principle. The technique, we call \emphMDL-PA, can be viewed as an analogue of Passive-Aggressive algorithms for probability distributions (equivalently viewed as lossless compression schemes) where we consider a \emphcode length margin. Separate models are maintained per class, and the update rule is an information projection correcting the true-class model's prediction on a margin-violating sample. To bridge theory and practice, we demonstrate a practical approximation to this online learning setup via adaptive Lempel-Ziv-style compression dictionaries to classify sequences in the text and malware domains.
PaperID: 1654, Poster
Abstract: Linear mode connectivity is often analysed through exact parameter symmetries, especially hidden-unit permutations that leave the network function unchanged. Yet, many relationships between trained networks are not captured by bijective coordinate relabelings: widths may differ, neurons may split or clone, and explicit parameter symmetries may be broken. In this paper, we introduce , taking an operator perspective on hidden-unit geometry and studying mode connectivity through the functional geometry of hidden units. Specifically, we represent each hidden layer as a graph of built from layerwise descriptors (weights or activations), associate it with a graph shift operator, and study two coupled objects: layerwise graph functional symmetries, which preserve the relational geometry, and intertwining correspondences across layers from two models, which align their graph structures. We show that permutations arise as a special case of this viewpoint and derive results for non-bijective constructions. Our analysis connects correspondence quality to preserved spectral modes, functional exchangeability of neurons, and bounds on endpoint distortion and linear path barriers. These findings suggest that linear mode connectivity can be governed beyond explicit parameter symmetry by broader spectral operator compatibility in hidden-unit geometry.
PaperID: 1655, Poster
Authors: Maksym Buleshnyi, Joshua Loftus, Sakina Hansen
Abstract: Shapley value-based explanations are widely used to attribute model predictions, but standard variants capture associative rather than causal relationships. Recent causally informed Shapley methods add a path-wise decomposition, but suffer from two key limitations: they assign zero attribution to chain-mediated paths, and their per-path values do not align with classical mediation estimands. We propose PE-SHAP, a Shapley-inspired path-wise decomposition built on the portion-eliminated effect. Per-mediator values sum to the total mediated effect at every sample, aggregate to the Natural Indirect Effect at the population level, and match the Path-Specific Effect under additive separability. These are properties prior Shapley-based path-wise methods do not provide. We validate PE-SHAP analytically and through Monte-Carlo experiments on synthetic SCMs with known path effects, and demonstrate it on real-world fairness case studies using the German Credit and COMPAS datasets.
Abstract: Infinite-dimensional orthonormal basis expansions play a central role in representing and computing with function spaces due to their favorable linear algebraic properties. However, common bases such as Fourier or wavelets are fixed and do not adapt to the structure of a given problem or dataset. In this paper, we aim to represent these bases with neural networks and optimize them. Our key idea is that any target infinite-dimensional orthonormal basis can be viewed either as a point on the Lie manifold of the orthogonal group, or equivalently, as the endpoint of a continuous path on that manifold that connects a reference basis, e.g. Fourier, to that target. Paths on the Lie manifold satisfy ordinary differential equations (ODEs) governed by skew-adjoint integral operators. Using neural networks to define finite-rank generators of such ODEs allows us to parameterize and optimize orthonormal bases in function space. While relying on finite-rank generators to model infinite operators might seem restrictive, we prove a universality result: even with a rank-2 generator, the integrated solutions of the ODE are dense in the orthogonal group under the appropriate operator topology. In other words, for any target orthonormal basis, there exists a path originating from a reference basis and driven by finite-rank generators that gets arbitrarily close to that target basis. We demonstrate the flexibility of our framework by transforming the Fourier basis into the principal components of a functional dataset, eigenfunctions of linear operators, or dynamic modes of energy-preserving physical simulations.
Abstract: Predictive Coding (PC) is an influential account of cortical learning. Much of recent work has focused on comparing PC to Backpropagation (BP) to find whether PC offers any advantages. Small scale experiments show that PC enables learning that is more sample efficient and effective in many contexts, though a thorough theoretical understanding of the phenomena remains elusive. To address this, we quantify the efficiency of learning in BP and PC through a metric called "target alignment", which measures how closely the change in the output of the network is aligned to the output prediction error. We then derive and empirically validate analytical expressions for target alignment in Deep Linear Networks. We show that learning in PC is more efficient than BP, which is especially pronounced in deep, narrow and pre-trained networks. We also derive exact conditions for guaranteed optimal target alignment in PC and validate our findings through experiments. We study full training trajectories of linear and non-linear models, and find the predicted benefits of PC persist in practice even when some assumptions are violated. Overall, this work provides a mechanistic understanding of the higher learning efficiency observed for PC over BP in previous works, and can guide how PC should be parametrised to learn most effectively.
Abstract: Large language models (LLMs) trained purely on text ostensibly lack any direct perceptual experience, yet their internal representations are implicitly shaped by multimodal regularities encoded in language. We test the hypothesis that explicit sensory prompting can surface this latent structure, bringing a text‑only LLM into closer representational alignment with specialist vision and audio encoders. When a tells the model to 'see' or 'hear', it cues the model to resolve its next‑token predictions as if they were conditioned on latent visual or auditory evidence that is never actually supplied. Our findings reveal that lightweight prompt engineering can reliably activate modality‑appropriate representations in purely text‑trained LLMs.
PaperID: 1659, Poster
Authors: Lirong Xia
Abstract: In social choice, anonymity (treating all agents equally) and neutrality (treating all alternatives equally) are widely regarded as "minimal demands" and "uncontroversial" axioms of equity and fairness. However, due to the ANR impossibility theorem, no voting rule can simultaneously satisfy anonymity, neutrality, and resoluteness (always choosing a unique winner). While the worst-case ANR impossibility is well understood, its likelihood in probabilistic models remains underexplored. We address this question through comprehensive likelihood analysis under semi-random models, providing accurate bounds to quantify the advantage of the optimial tie-breaking mechanisms over other commonly-used tie-breaking mechanisms. Our characterizations reveal an n^\Theta(m!) improvement in many cases, where n is the number of agents and m is the number of alternatives. We also prove a quantitative ANR impossibility theorem characterizing the tradeoff between equity (anonymity and neutrality) and a new efficiency criterion called \delta-anonymity. Our work advances the understanding of trade-offs between equity and efficiency in voting.
PaperID: 1660, Poster
Abstract: Federated Continual Learning (FCL) requires distributed clients to acquire new tasks while retaining previously learned knowledge. In this paper, we revisit the heterogeneity challenge in FCL, which is composed of two distinct aspects: spatial static heterogeneity, e.g., initial data distribution and model architecture differences, and temporal dynamic heterogeneity, e.g., task shifts over time. While existing methods acknowledge the presence of heterogeneity, they typically treat it in a unified manner by employing a single adaptive strategy without explicitly disentangling their impacts. In contrast, we propose a heterogeneity-aware distillation framework for FCL called Ha-FCL, which is designed to handle spatial and temporal heterogeneity through tailored distillation strategies. Specifically, Ha-FCL combines width-aware federated training with server-side public-anchor distillation, integrating multi-layer feature distillation for architecture-aware alignment with historical-logit regularization and representation co-distillation. This forms a coordinated dual-dimensional distillation scheme that balances adaptation to new tasks with retention of prior knowledge. This enables more targeted knowledge transfer and leads to improved model performance under complex FCL scenarios. We evaluate our method in various heterogeneous settings. Compared with state-of-the-art baselines, Ha-FCL consistently improves test accuracy and shows better robustness under challenging heterogeneous FCL settings.
PaperID: 1661, Poster
Abstract: Many practical deployments require language models to internalize knowledge that was absent from pretraining, such as proprietary corpora, continually updated facts, or specialized domain skills. Self-distillation has emerged as a natural approach for this task. A model conditioned on the source material acts as the teacher, and an unconditioned student is trained to reproduce the teacher's behavior closed-book. We propose Speculative Self-Distillation (SSD), an efficient per-token mixed-policy method for self-distillation. The student generates each rollout by default, while the teacher intervenes only at positions where its next-token distribution substantially diverges from the student's. This design keeps training rollouts close to the student's test-time behavior, unlike off-policy distillation with teacher-generated prefixes, but avoids spending teacher compute on the many low-signal tokens encountered by fully on-policy methods. In effect, SSD concentrates supervision precisely where privileged context changes the prediction. On three closed-book QA benchmarks measuring fact recall, knowledge updates, and skill acquisition, SSD matches on-policy distillation with 45% fewer supervised tokens on average and consistently outperforms off-policy distillation on the held-out test set. SSD therefore establishes a new accuracy-efficiency Pareto frontier for self-distillation. Code is available at https://anonymous.4open.science/r/Speculative-Self-Distillation/.
Authors:
Ke He, Le He, Shunpu Tang, Yafei Wang, Lisheng FanAbstract: Expressive generative policies such as diffusion and flow models are appealing for maximum-entropy (MaxEnt) online reinforcement learning because of their ability to model multimodal and highly non-Gaussian action distributions. However, training effective soft generative policies faces two obstacles that often arise together. First, marginal action densities are often unavailable, so existing methods typically rely on entropy bounds, heuristic proxies or approximations. Second, iterative shared-parameter samplers raise inference cost and require backpropagation through time over repeated network evaluations, increasing memory cost and destabilizing policy optimization. These obstacles motivate us to seek a generative policy that exposes a tractable MaxEnt objective while requiring only a single sampled actor forward pass for action generation. To this end, we propose soft generative actor-critic (SoftGAC), whose actor defines a stochastic bridge from a fixed base latent to a terminal action latent in pre-tanh space. This structured bridge allows us to lift the MaxEnt objective as an analytically tractable path-wise relative-entropy objective against a high-entropy reference process. In practical finite-step implementation, this relative entropy reduces exactly to sampled transition control energy and thus provides principled soft regularization. Moreover, we keep the single-pass actor lightweight by using small step-specific bridge transitions, each evaluated only once per sampled action, while maintaining a parameter budget comparable to strong actor baselines. Extensive experiments on challenging continuous-control benchmarks show that \myalgo attains higher or competitive returns than strong generative policy baselines, including diffusion and flow-matching policies, while staying in the low-latency regime of one-pass actors and avoiding the much higher cost of many-step diffusion samplers, showing considerable improvements in the compute-return tradeoff.
Abstract: Switching dynamical systems can model complicated time series data while maintaining interpretability by inferring a finite set of dynamics primitives and explaining different portions of the observed time series with one of these primitives. However, due to the discrete nature of this set, such models struggle to capture smooth, variable-speed transitions, as well as stochastic mixtures of overlapping states (e.g., non-instantaneous state transitions), and the inferred dynamics often display spurious rapid switching on real-world datasets. Here, we propose the Gumbel Dynamical Model (GDM). First, by introducing a continuous relaxation of discrete states and a different noise model defined on the relaxed-discrete state space via the Gumbel distribution, GDM expands the set of available state dynamics, allowing the model to approximate smoother and non-stationary ground-truth dynamics more faithfully. Breaking from established literature, this new class directly links states to observations and does not blur latent dynamics with Gaussian noise. Second, the relaxation makes the model fully differentiable, enabling fast and scalable training with standard gradient descent methods. We validate our approach on standard simulation datasets and highlight its ability to model soft, sticky states and transitions in a stochastic setting. Furthermore, we apply our model to two real-world datasets, demonstrating its ability to infer interpretable states in stochastic time series with multiple dynamics, a setting where traditional methods often fail.
PaperID: 1664, Poster
Authors: Guangwei Zhang, Gang Wang
Abstract: Correspondence pruning—separating sparse inliers from severe outliers in noisy feature matches—is a long-standing bottleneck for two-view geometry estimation in 3D vision. Existing learning-based pruners pursue this goal through ever more sophisticated context operators—local graphs, dynamic neighborhoods, or global attention—yet a controlled comparison reveals a striking phenomenon: when we freeze the structural assignment of these operators after a single forward pass, accuracy collapses by over 30% mAP regardless of operator choice. This suggests that the limiting factor of current methods is not which operator is used, but the implicit assumption that structure can be committed in one shot. We argue that correspondence pruning should instead be cast as \emphiterative structural rectification (ISR), in which the geometric structure and the feature representation must \emphco-evolve layer by layer. Building on this view, we propose ISR-Net, a minimal architecture that pairs a layer-wise re-derived local-context operator (a dynamic k-NN graph in feature space) with a layer-wise re-derived global-context operator (soft assignment to learnable cluster prototypes), so that any gain over prior arts is attributable to the rectification \emphschedule rather than to operator novelty. Despite using deliberately off-the-shelf operators, ISR-Net attains state-of-the-art results on YFCC100M, SUN3D, and HPatches, surpassing the strongest attention-based competitor by +8.2% mAP@5^\circ while running 35% faster, indicating that \emphhow often structure is refined matters more than the operator's complexity.
Abstract: This paper studies kernelized bandits (also known as Gaussian process bandits) in an adversarial environment, where the reward functions in a known reproducing kernel Hilbert space (RKHS) may be adversarially chosen at each round. We show that the exponential-weight algorithm achieves \tildeO(\sqrtT \gamma_T) adversarial regret, where T and \gamma_T denote the number of total rounds and the maximum information gain, respectively. For \nu-Mat\'ern kernels, we also show algorithm-independent lower bounds that guarantee the optimality of our algorithm up to polylogarithmic factors. Furthermore, we present a computationally efficient variant of our algorithm using Nystr\"om approximation while maintaining nearly optimal regret guarantees.
PaperID: 1666, Poster
Abstract: N-dimensional splatting extends 3D Gaussian Splatting with conditioning variables such as view direction and time to model view-dependent and dynamic effects. Each primitive is lifted into a joint distribution over 3D position and conditioning variables, and at render time is conditionally sliced at the query to recover a 3D primitive whose mean, opacity, and covariance vary continuously with view or time. N-DGS (covering 6DGS and 7DGS) and UBS together represent the leading N-dimensional splatting formulations across Gaussian and Beta kernels, but share the same conditional-slicing step: all three effects are derived from the joint covariance matrix per primitive in every forward pass, requiring matrix inversions and regression-matrix multiplications that scale with primitive count. The conditional-slicing step is kernel-agnostic, so a single improvement carries across families. We introduce direct conditional parameterization, replacing the covariance-derived form with explicit per-effect parameters: a Cholesky precision factor for opacity and a displacement matrix with learnable per-dimension coupling for position. Applied to both kernels, the parameterization yields dGS (Gaussian) and dBS (Beta). Across five static and dynamic benchmarks, dGS and dBS match or exceed their covariance-derived counterparts on quality (up to +1.26 dB in the main MCMC evaluation) while running 6.9-7.7x faster at the slicing step and up to 2.66x faster end-to-end rendering. Direct conditioning is a kernel-agnostic drop-in replacement for covariance-derived conditioning in N-dimensional splatting.
Abstract: Planning with world models enables test-time adaptation for embodied agents, offering superior generalization over direct policy learning. However, long-horizon tasks remain hindered by cumulative errors and search complexity. Existing hierarchical methods often use rigid macro-actions or latent skills, leading to brittle representations and limited fine-grained control. We introduce Implicit-HWM, a hierarchical world model that decouples state-space reachability from local control. Our framework utilizes a high-level model to sample feasible transitions across the state manifold and a low-level inverse dynamics solver to ground these transitions into precise actions. This allows the planner to identify reachable subgoals without being constrained by low-level execution details. By prioritizing local physical feasibility over global behavioral patterns, Implicit-HWM synthesizes novel long-horizon plans from short segments, transcending the specific behavioral modes seen during training. We demonstrate its effectiveness across three navigation and manipulation suites, achieving a 38.5% improvement over state-of-the-art policy and diffusion-based planners. Crucially, by leveraging this hierarchical structure to mitigate compounding errors, our approach yields a 120% performance gain over single-level world model MPC, particularly in complex, unseen configurations for long-horizon tasks.
Abstract: Search-augmented reasoning agents interleave internal reasoning with calls to an external retriever, and their performance relies on the quality of each issued query. However, under outcome-reward reinforcement learning, every search decision in a rollout shares the same trajectory-level reward, leaving individual queries without step-specific credit. Recent process-supervision approaches address this gap by drawing step-level signals from outside the policy, relying either on a much larger teacher model, or on sub-question annotations produced by a stronger external system. In contrast, we propose SD-Search, which derives step-level supervision \emphfrom the policy itself through hindsight self-distillation, requiring neither an external teacher nor additional annotations. In SD-Search, a single model plays two roles that differ only in conditioning: a student that sees only the context available at inference time, and a teacher that additionally conditions on a compact hindsight block summarizing the search queries and final outcomes of a group of rollouts sampled from the same question. Since teacher knows how each rollout unfolded and which ones succeeded, its query distribution implicitly marks which decisions were worth making, and the student is trained to recover this behavior by minimizing the token-level Jensen--Shannon divergence to the teacher at search-query positions. This layers a dense, step-level signal on top of GRPO's coarse trajectory reward. Crucially, this signal requires only a single teacher forward pass per rollout at training time, all absorbed within the standard RL loop without external model inference, auxiliary annotation pipeline, or additional training. On seven single-hop and multi-hop QA benchmarks, SD-Search reaches 0.428 average Exact Match with Qwen2.5-3B, matching the leading process-supervision baseline without its external teacher, and 0.476 with Qwen2.5-7B, surpassing all outcome-reward and process-supervision baselines on the average.
Abstract: Internal modelling of the world, predicting transitions between previous states and next states under actions, is essential to reasoning and planning for LLMs and VLMs. Learning such models typically requires costly action-labelled trajectories. We propose SWIRL, a self-improvement framework that learns from state-only sequences by treating actions as a latent variable and alternating between Forward World Modelling (FWM) and an Inverse Dynamics Modelling (IDM). SWIRL iterates two phases: (1) Variational Information Maximisation, which updates the FWM to generate next states that maximise conditional mutual information with latent actions given prior states, encouraging identifiable consistency; and (2) ELBO Maximisation, which updates the IDM to explain observed transitions, effectively performing coordinate ascent. Both models are trained with reinforcement learning (specifically, GRPO) with the opposite frozen model's log-probability as a reward signal. We provide theoretical learnability guarantees for both updates, and evaluate SWIRL on LLMs and VLMs across multiple environments: single-turn and multi-turn open-world visual dynamics and synthetic textual environments for physics, web, and tool calling. SWIRL achieves gains of 16% on Aurora-Bench, 28% on ByteMorph, 16% on WorldPredictionBench, and 14% on StableToolBench.
Abstract: We study dynamic measure transport for generative modeling, focusing on transport maps that connect a source measure P_0 to a target measure P_1 by integrating a velocity field of the form v_t(x) = \mathbbE[\dot X_t \mid X_t = x], where X_\bullet = (X_t) _t is a stochastic process satisfying (X _0, X _1) ~ P _0 \otimes P _1 and \dot X _t is its time derivative. We investigate when X _\bullet induces a _straight-line flow_: a flow whose pointwise acceleration vanishes and is therefore exactly integrable by any first-order method. First, we develop multiple characterizations of straight-line flows in terms of PDEs involving the conditional statistics of the process. Then, we prove that straight-line flows under endpoint independence exhibit a sharp dichotomy. On the one hand, we construct explicit, computable straight-line processes for arbitrary Gaussian endpoints. On the other hand, we show that straight-line processes do not exist for targets with sufficiently well-separated modes. We demonstrate this obstruction through a sequence of increasingly general impossibility theorems that uncover a fundamental relationship between the sample-path behavior of a process with independent endpoints and the space-time geometry of this process' flow map. Taken together, these results provide a structural theory of when straight-line generative flows can, and cannot, exist.
PaperID: 1671, Poster
Abstract: framework, in which a learner sequentially collects samples and re-estimates a model through Bayesian linear regression. We show that, when model drift is governed by an unknown mean fixed-point coupled with Gaussian deviations, model posterior estimates are determined by a Kalman filter and the system is asymptotically stable. We also show that the steady-state model posterior can be computed as the solution of the so-called Discrete Algebraic Riccati Equation (DARE), and that D-optimal steady-state and myopic sample selection policies can be computed by solving tractable convex optimization problems. We do so by characterizing the curvature of the DARE solution, even though the latter generally has no closed form. Finally, we show that our approach generalizes to the non-linear model setting by replacing the Kalman filter with an extended Kalman filter.
Abstract: Learning-augmented algorithms have emerged as a powerful paradigm to surpass traditional worst-case lower bounds by integrating potentially noisy predictions. While this framework has seen success in online scheduling, existing work primarily optimizes job latency while relying on frequent, "blind" preemptions. This ignores the fundamental trade-off between algorithmic performance and preemption complexity. We provide the first systematic study of learning-augmented scheduling that curbs preemption while optimizing latency. We establish that the gap between theoretical latency bounds and preemption overhead can be bridged with solid analytical foundations. Our results include O(1)-competitive algorithms for single and unrelated parallel machines with only O(1) preemptions per job under accurate predictions, with overhead scaling logarithmically with the prediction error. By providing the first bounded-preemption guarantees for unrelated and malleable machines, we extend the theoretical reach of the learning-augmented framework to more constrained and realistic settings. Finally, our algorithms are validated through experiments.
Authors: Abhijith Jayakumar, Shreya Shukla, Marc Vuffray, Andrey Lokhov, Sidhant Misra
Abstract: Learning Gibbs distributions using only sufficient statistics has long been recognized as a computationally hard problem. On the other hand, computationally efficient algorithms for learning Gibbs distributions rely on access to full sample configurations generated from the model. For many systems of interest that arise in physical contexts, expecting a full sample to be observed is not practical, and hence it is important to look for computationally efficient methods that solve the learning problem with access to only a limited set of statistics. We examine the trade-offs between the power of computation and observation within this scenario, employing the Ising model as a paradigmatic example. We demonstrate that it is feasible to reconstruct the model parameters for a model with \ell_1 width \gamma by observing statistics up to an order of O(\gamma). This approach allows us to infer the model's structure and also learn its couplings and magnetic fields. We also discuss a setting where prior information about structure of the model is available and show that the learning problem can be solved efficiently with even more limited observational power.
Abstract: Post-hoc unauthorized-training data detection for large language models (LLMs) typically assumes a query-with-originals regime: rights holders query a target LLM with raw proprietary data and assess whether the model assigns them stronger memorization-based detection signals, e.g., higher confidence or lower loss, than held-out non-training reference texts. We show that this regime becomes brittle under data laundering, where the target LLM is trained on semantics-preserving but stylistically or structurally transformed surrogates of proprietary data to obfuscate provenance. Since training-time exposure occurs in the laundered form, memorization signals may no longer appear on the originals, collapsing the candidate-reference signal separation that standard detectors rely on. We counter this threat by studying laundering-aware detection with raw proprietary data, a held-out reference corpus, and query access to the target LLM, while the laundering transformation is undisclosed. Since exact recovery of the laundered corpus is infeasible, we infer a detection-useful synthesis process via an auxiliary LLM that maps originals into training-like queries. To make this search tractable, we introduce Synthesis Data Reversion (SDR), which constrains the unbounded space of natural-language transformations through a goal-details decomposition: a high-level transformation goal, e.g., "lyrical rewriting", and fine-grained details, e.g., "with vivid imagery". SDR identifies the most likely goal and iteratively refines details so synthesized queries elicit stronger target-model detection signals. Evaluated on the MIMIR benchmark against diverse laundering practices and target LLM families (Pythia, Llama2, and Falcon), SDR consistently restores detection signals, offering a practical auditing tool against data laundering.
PaperID: 1675, Poster
Abstract: Motivated by the success of large language models (LLMs) with broad generalizability and robustness to noisy or unreliable pre-training data, we seek to bring similar capabilities to PDE solvers. In addition, inspired by the posterior predictive mean-based inference mechanism of the in-context learning in prior-data fitted networks (PFNs), we propose PDE-PFN, a prior-data fitted neural solver that approximates the posterior predictive mean of PDE solutions via in-context learning. PDE-PFN builds on a PFN architecture with self- and cross-attention mechanisms of Transformer and is pre-trained on noisy approximate solutions generated by physics-informed neural networks, serving as diverse but not necessarily exact priors. Through experiments on a range of two-dimensional PDEs, we demonstrate that PDE-PFN achieves empirical generalization across heterogeneous equations, robustness under noisy priors, and zero-gradient-update in-context inference capability. Our approach not only outperforms task-specific baselines but also provides a flexible and robust framework for advancing SciML.
Authors: Yimeng Min, Carla Gomes
Abstract: Learning permutations is fundamental to sorting, ranking, and matching, but existing differentiable methods based on entropy-regularized Sinkhorn produce a single softened solution and collapse under ambiguity. We present PermFlow, a conditional flow matching framework that operates directly on the affine subspace of matrices with unit row and column sums. A closed-form tangent-space projector preserves these constraints exactly along every trajectory, by construction rather than through iterative correction, and a nearest-target coupling routes distinct noisy initializations toward distinct valid permutations. The result is a model that captures multimodal permutation distributions rather than collapsing them to a single mode. On a visual sorting task with blended-digit ambiguity and a symmetric linear assignment problem, PermFlow achieves high accuracy on unambiguous inputs and recovers both valid permutations under ambiguity, where Sinkhorn-based baselines structurally fail.
PaperID: 1677, Poster
Authors: Mengzhou Sun
Abstract: Extremal combinatorics seeks rare mathematical structures that satisfy hard constraints while maximizing a global objective. Most existing neural systems for mathematical discovery train a model to generate complete configurations from scratch. We argue that this framing mismatches the dynamics of effective search: high-quality solutions are more effectively found by preserving a feasible partial core and conditionally regenerating the remaining uncertain objects. We formulate extremal structure discovery as object-level active-core completion: the model is trained to conditionally complete masked objects given a feasible partial core, capturing inter-object relations rather than fitting a distribution over complete configurations. We instantiate this with a full-attention Transformer over object-aligned slots, and propose Inpainting PatternBoost, a variant of PatternBoost. In each iteration, elite constructions from a top pool are partially destroyed by random or structured remasking; the model learns to reconstruct the missing objects, and their completions are reinserted into the pool. We evaluate on the no-three-in-line problem and maximum sum-of-radii circle packing in the unit square, under a single shared abstraction. On no-three-in-line, our framework recovers the optimal 2n-point configuration for n up to 18, and reaches 2n-1 points for n=19, 20, 25, outperforming prior neural baselines. On circle packing, where FlowBoost combines neural generation with gradient-based local search, our approach matches or exceeds FlowBoost when paired with the same local search; even without local search, it reaches an average of 88.1% of FlowBoost's reported score across instances.
Abstract: Visual text compression (VTC) promises efficient long-context processing by rendering text into an image and re-encoding it with a vision-language model, often producing 3--20× fewer decoder tokens than subword tokenization. Yet token savings do not translate predictably into downstream utility: on some tasks the visual path matches or exceeds the text path, on others it collapses, and the compression ratio itself does not predict which regime will occur. The missing quantity is therefore not another summary of efficiency, but a principled measure of task-relevant information loss induced by visual encoding. We address this problem by formulating VTC in the language of measure transport. Treating text and visual tokens as empirical probability measures, we show that the ViT patch encoder induces a push-forward map whose transport cost decomposes into a precision cost from within-patch aggregation and a coverage cost from cross-patch fragmentation. Both terms are estimable from downstream-label-free probes. This formulation yields two operational consequences: a downstream-label-free routing criterion that selects whether to use the visual path for a given input or benchmark instance, and a transport-informed foveation mechanism that re-encodes high-cost regions at higher resolution. Across 24 NLP datasets at Qwen3-4B, our label-free rule matches the per-dataset oracle on 17/24 datasets (70.8%), and improves the average task score by +3.3% with -10.3% average tokens relative to a pure-LLM.
Abstract: Token-level steering has emerged as a pivotal approach for inference-time alignment, enabling fine-grained control over large language models (LLMs) by modulating their output distributions without parameter updates. While effective, existing methods rely on dense intervention at every decoding step. This persistent manipulation not only incurs substantial computational overhead but also risks compromising generation quality by excessively drifting from the model’s intrinsic distribution. In this work, we show that dense intervention is unnecessary and propose Sparse Inference-time Alignment (SIA), which performs sparse junction steering, intervening only at critical decision points along the generation trajectory. Our key insight is that high-entropy junctions are empirically consistent with pivotal decision points in the generation trajectory and are particularly susceptible to misalignment, suggesting that these points benefit from alignment-related reward signal. Extensive experiments across different model families and alignment objectives show that steering only 20%–80% of tokens achieves the superior alignment–efficiency trade-offs. For strong base models such as Qwen3, intervening on as few as 20% of tokens matches or even surpasses heavily post-trained instruct models. This sparsity enables stronger guidance while better preserving the model's native distribution, integrates seamlessly with search-based methods (e.g., Best-of-N), and improves efficiency by reducing the required sampling budget while yielding practical wall-clock speedups under matched-quality comparisons.
Abstract: Gaussian processes (GPs) offer a principled probabilistic model over functions, but exact inference is restricted to the linear-Gaussian regime. We establish an explicit equivalence between GPs and a class of linear diffusion models, recasting predictive sampling as an ODE with closed-form Gaussian dynamics and a likelihood-dependent guidance term that admits a simple Monte Carlo approximation. In the linear-Gaussian setting, we recover standard GP conditioning exactly; beyond conjugacy, the same machinery handles any conditioning statement admitting point-wise likelihood evaluation — including non-linear physics, and, for the first time, natural language via large language models. Whitening isolates the irreducible non-Gaussian dynamics, minimising Wasserstein-2 transport cost and eliminating numerical stiffness. The result is a general-purpose GP inference scheme requiring no bespoke derivations. Together, these results provide a general mechanism for incorporating the full richness of real-world knowledge as conditioning information, opening a new frontier for the probabilistic modelling of real-world problems.
Abstract: Generative artificial intelligence (AI) is increasingly integrated into the online platforms where humans exchange opinions; large language models (LLMs) now polish users' posts on LinkedIn and provide context for content shared on X. While prior work has shown that AI can express biased opinions and shape individuals' opinions during human-AI interactions, less attention has been paid to its influence on collective opinion formation when mediating human-to-human communication. We address this gap via a combination of empirical and theoretical analyses. We show empirically that LLMs from multiple popular families introduce directional biases when instructed to edit human-written texts on contested topics, for example, nudging texts in favor of gun control and against atheism. Building on this observation, we introduce a mathematical model of opinion dynamics in which an AI system sits between users on a social network, transforming the opinions they express and perceive. By analytically characterizing the equilibrium of this model and performing simulations on real social network data, we show that biases introduced by AI in human-to-human communication can be amplified through the network and shift collective opinion in their direction. In light of these findings, we investigate whether such biases are controllable by online platforms. We audit the "Explain this post" feature on X and find evidence of pro-life bias in Grok's outputs on abortion-related content, which we trace back to specific design choices.
PaperID: 1682, Poster
Authors:
Aviv Orly, Ori Shem Ur, Yaron OzAbstract: Inference-time sampling is a central mechanism for improving large language models, with temperature controlling the tradeoff between reliability and diversity. While recent work suggests that varying temperature across samples can improve multi-sample inference, simple fixed-temperature schemes remain the standard in practice. We investigate this discrepancy in the canonical pass@K setting, asking when temperature diversity is useful, and when a single tuned temperature suffices. To this end, we analyze per-question success probabilities across multiple models and datasets, tracking how they vary with sampling temperature. We identify several consistent properties of these profiles, including that harder questions are less likely to share efficient sampling temperatures, although this trend is only logarithmic in hardness. We then derive an efficient method for evaluating the theoretically optimal multi-temperature allocation scheme, and show that although such a scheme improves performance, its gains over single-temperature baselines are typically modest when temperature is tuned as a function of the sampling budget K. We explain this phenomenon through mathematical mechanisms that we call dulling effects, and show how they ensure the robustness of temperature choices across different question distributions. Our results clarify when multi-temperature sampling is useful, and how temperatures should ideally be chosen.
PaperID: 1683, Poster
Abstract: Echocardiography segmentation plays a central role in quantitative cardiovascular assessment. However, automated echocardiography segmentation remains challenging due to speckle noise, poorly defined anatomical boundaries, acquisition variability, and the limited anatomical coverage of existing benchmarks. Conventional segmentation methods demonstrate limited accuracy under such low image quality and domain shifts. Inspired by the observation that expert cardiologists can maintain reliable anatomical delineation by leveraging their structural understanding of the cardiac anatomy, we propose that incorporating anatomical priors into echocardiography segmentation models can improve both precision and robustness. To this end, we propose AST-Seg, an anatomy-aware spatio-temporal framework for echocardiography segmentation. AST-Seg introduces an anatomy-aware feature encoder that regularizes transformer attention with anatomy-derived spatial-relation priors, and a Deformable Spatio-Temporal Mamba module that captures localized and periodic cardiac motion through adaptive spatial aggregation followed by temporal state-space modeling. In addition, we curate an Echocardiography Anatomy (EA) dataset with pixel-level annotations for 12 clinically relevant cardiac anatomy, enabling comprehensive multi-structure cardiac interpretation. Extensive experiments on in-distribution CAMUS and EA datasets, as well as an out-of-distribution point-of-care echocardiography dataset, demonstrate that AST-Seg provides domain-generalized segmentation with improved precision.
Authors:
Janik Bürgermeister, Julius Durmann, Martin Bichler, Mete Ş Ahunbay, Bary Pradelski, Marco ScarsiniAbstract: We study equilibrium learning in discretized Bayesian games, focusing on first-price auctions as a testbed for gradient-based learning dynamics under private information. While the symmetric independent private values model admits a unique Bayesian correlated equilibrium (BCE) in the continuous limit, standard convergence guarantees for gradient-based learners do not apply: classical no-regret theory predicts convergence only to Bayesian coarse correlated equilibria (BCCE), which can remain large and far from the competitive equilibrium even under fine discretizations. To bridge this gap, we introduce the notion of Bayesian semicoarse correlated equilibrium (BSCCE), an equilibrium concept that refines BCCE by restricting deviations to probability-preserving transformations consistent with projected gradient dynamics. We provide a linear programming characterization of BSCCE with polynomially many variables and constraints and establish the inclusion hierarchy BCE \subseteq BSCCE \subseteq BCCE. Our main finding is that the BSCCE set contracts in many cases as the discretization of the Bayesian game is refined. Even with priors for which BCCE fails to concentrate, the diameter of the BSCCE polytope shrinks toward zero, indicating convergence to the unique continuous-game Bayes-Nash equilibrium. These results provide a principled explanation for the observed robustness of gradient-based learning in discretized Bayesian games, showing that these dynamics select a smaller equilibrium set than standard no-regret theory predicts.
Abstract: Experience-driven learning has emerged as a promising paradigm for enabling agents to improve from interaction trajectories by accumulating and reusing past experience. However, existing approaches are predominantly developed in textual settings and rely on manually designed memory schemas, limiting their applicability to multimodal environments. In real-world scenarios, experience is inherently multimodal, involving heterogeneous signals across perception, reasoning, and action, which makes effective memory design significantly more challenging. In particular, the optimal way to structure and utilize multimodal experience is highly task-dependent and evolves over time, rendering fixed memory designs insufficient. In this work, we propose a new paradigm, learning to learn from multimodal experience, which shifts memory design from a predefined component to an adaptive and learnable process. Our framework enables agents to dynamically construct, organize, and utilize memory based on task requirements and interaction history, effectively learning how to structure experience for improved performance. Experiments demonstrate that adaptive memory design substantially enhances agent performance and generalization across multimodal tasks, highlighting the critical role of learning memory mechanisms in experience-driven learning.
PaperID: 1686, Poster
Authors: Teresa Zhang
Abstract: Today, verifier-guided best-of-N inference is a common way to spend test-time compute in language modeling, code generation, and alignment: generate N candidates, score them with a verifier or reward model, and deploy the highest-scoring output. This paper argues that the relevant reliability object is not average verifier accuracy but \emphselected-tail reliability, namely the behavior of the output chosen after optimizing over N noisy public scores; we derive exact selected-rank exposure laws for best-of-N selection and finite-verifier resource laws showing that selected public optimism scales at the \sqrt\log N/m regime under m units of verifier evidence. We then separate public optimism from hidden utility harm, showing that public-score inflation alone is not a harm claim and that harm or uncertifiability requires public-hidden mismatch, tail dominance, count imbalance, or an unresolved selected-tail regime. Building on these results, we introduce Tail-Certified Compute Caps, a conservative procedure that certifies candidate budgets or refuses certification using selected-tail intervals, verifier-resource penalties, mismatch envelopes, and feasibility gates. In exact-verifier/public-hidden experiments, the theory predicts selected false-positive exposure and supports resource-aware cap decisions over an all-attempted denominator, while non-code candidate tables with pre-existing public human annotations serve as secondary scoped K\ge16 consistency checks under proxy/model-prior/verifier-style scores across Arena55K and Stanford SHP. Overall, the paper provides a reliability framework for verifier-guided inference-time scaling: candidate budgets should be justified by evidence about the selected tail that is actually deployed, not only by average verifier validation.
PaperID: 1687, Poster
Authors: Duc T Nguyen, Cesar Uribe
Abstract: Fréchet regression, or conditional barycenters, is a flexible framework for modeling relationships between covariates (usually Euclidean) and response variables on general metric spaces, e.g., probability distributions or positive definite matrices. However, in contrast to classical barycenter problems, computing conditional counterparts in many non-Euclidean spaces remains an open challenge, as they yield non-convex optimization problems with an affine structure. In this work, we study the existence and computation of conditional barycenters, specifically in the space of symmetric positive definite (SPD) matrices with the Bures-Wasserstein metric, which connects SPD regression to optimal transport, avoids matrix logarithms, and preserves well-posedness under signed weights. We provide a sufficient condition for the existence of a minimizer of the conditional barycenter problem that characterizes the regression range of extrapolation. Moreover, we further characterize the optimization landscape, proving that under this condition, the objective is free of local maxima. Additionally, we develop a projection-free and provably correct algorithm for the approximate computation of first-order stationary points. Finally, we provide a stochastic reformulation that enables the use of off-the-shelf stochastic Riemannian optimization methods for large-scale setups. Numerical experiments validate the performance of the proposed methods on regression problems of real-world biological networks and on large-scale synthetic Diffusion Tensor Imaging problems.
Abstract: Strategic classification studies learning settings in which individuals can modify their features, at a cost, in order to influence the classifier’s decision. A central question is how the sample complexity of the induced (strategic) hypothesis class depends on the complexities of the underlying hypothesis class and the cost structure governing feasible manipulations. Prior work has shown that in several natural settings, such as linear classifiers with norm costs, the induced complexity can be controlled. We begin by showing that such guarantees fail in general --- even in simple cases: there exist hypothesis classes of VC dimension 1 on the real line such that, even under the simplest interval neighborhoods, the induced class has infinite VC dimension. Thus, strategic behavior can turn an easy learning problem into a non-learnable one. To overcome this, we introduce structure via a geometric definability assumption: both the hypothesis class and the cost-induced neighborhood relation can be defined by first-order formulas over \mathbb R_\exp. Intuitively, this means that hypotheses and costs can be described using arithmetic operations, exponentiation, logarithms, and comparisons. This captures a broad range of natural classes and cost functions, including \ell_p distances, Wasserstein distance, and information-theoretic divergences. Under this assumption, we prove that learnability is preserved, with sample complexity controlled by the complexity of the defining formulas.
Abstract: Unequal availability of human preference data across languages poses a significant challenge for aligning large language models in multilingual settings. To address the lack of sufficient data in low-resource language alignment, we propose a meta-learning framework for Reinforcement Learning from Human Feedback and Direct Preference Optimization. By leveraging preference data from other languages, our framework learns a transferable initialization that enables effective adaptation to a target language with minimal data. We provide theoretical guarantees for both the meta-reward modeling and meta-policy optimization settings, and empirically demonstrate the effectiveness of our approach on multilingual benchmarks. In an extremely low-resource setting with only 100 target-language preference samples, our approach achieves up to 28% win-rate improvements over baseline methods, and consistently outperforms baselines across multiple target languages and model scales. Our approaches retain these advantages across different combinations of meta-training languages and varying linguistic distances from the target languages.
PaperID: 1690, Poster
Authors: Hechuan Shen, Jian Zhang, Yingmin Liang
Abstract: Learning-based inertial odometry has shown strong potential for mitigating drift by extracting data-driven motion priors. However, existing methods primarily learn statistical mappings from inertial measurements to motion states, failing to consider the power balance between the specific-force input port and the predicted velocity. This limitation is particularly pronounced in aggressive MAV flight, where rapid accelerations, high angular rates, and aerodynamic dissipation can cause locally accurate velocity estimates to accumulate into trajectory drift. We propose PHIONet (Port Hamiltonian Inertial Odometry Network), which incorporates dissipative energy dynamics as a physically grounded prior. PHIONet translates the continuous port-Hamiltonian dynamics into a discrete energy-consistency residual over finite IMU windows, employing the DCM (Dissipative Coupling Module) to bridge IMU correction and velocity prediction. Experiments on the public Blackbird and EuRoC MAV datasets show that PHIONet achieves lower ATE and RTE than representative learning-based inertial odometry baselines. Code and related materials are available at https://anonymous.4open.science/r/PHIONet-D3CC.
PaperID: 1691, Poster
Abstract: Current text-to-image diffusion models remain unreliable for culturally grounded generation, often producing incorrect or stereotyped depictions, particularly for underrepresented cultures. While prior work has largely focused on evaluation and benchmarking, improving cultural generation remains underexplored. In this work, we study cross-cultural image generation through cultural vectors, defined as the difference between fine-tuned and pretrained model parameters. We show that cultural vectors enable fine-grained control over cultural alignment and failure modes through simple scaling at inference time. We further find that deeper U-Net layers encode stronger cultural information, while cultural vectors across countries occupy largely independent directions in parameter space. However, unlike prior work on task arithmetic, naively combining cultural vectors fails to produce meaningful multi-cultural behavior due to strong interference across cultures. Motivated by this limitation, we propose Descriptor-Guided Merging (DGM), a culturally-aware merging approach that estimates cultural vectors for unseen countries using descriptor-based similarities across cultures. Experiments on two benchmarks across 16/10 countries show that cultural vectors improve generation over pretrained models by 18.9%, while DGM improves cross-cultural generalization by 4.2% over naive merging.
Abstract: Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the corresponding image region. This setup assumes that the relevant linguistic unit is already known, overlooking a more basic challenge in grounded communication: determining which parts of the text are visually referential and how they correspond to entities in the image. We formulate grounding as bidirectional concept correspondence over an image-text pair. Given an image and its paired text, the goal is to recover all correspondences between visually referential text spans and instance-level image segments, without assuming that the relevant text spans are provided. This formulation unifies common grounding tasks, including phrase grounding, referring expression grounding, and open-vocabulary detection, by treating text segmentation, image segmentation, and cross-modal alignment as a single correspondence prediction problem. To address this task, we introduce our model, a grounding model built on top of a pretrained vision-language model. It uses learnable bridge tokens to represent candidate image-text correspondences and predicts, for each token, a text mask, an image mask, and a correspondence presence score. To train and evaluate this task, we convert diverse grounding and segmentation datasets into a unified correspondence format. Experiments show that our model consistently outperforms baselines, improving correspondence F1 by 59% on the long-caption dataset and by 73% on zero-shot LVIS, where the large category list serves as the text input.
Abstract: We investigate data filtering for large model pretraining via new scaling studies that target the high compute, data-scarce regime. In spite of an apparently common belief that filtering data to include only high-quality information is essential, our experiments suggest that with enough compute, the best data filter is no data filter. We find that sufficiently trained large parameter models not only tolerate low-quality and distractor data, but in fact, including nominally 'poor' data improves model performance.
PaperID: 1694, Poster
Abstract: Advances in deep reinforcement learning have enabled the development of policies capable of solving complex, high-dimensional problems, allowing AI agents to strategize and reason purely through interaction without supervision. Despite these successes, our knowledge on the functional properties and the underlying structure of the deep neural policy manifold remains limited. In this paper, we uncover a fundamental functional property in reinforcement learning: the intrinsic alignment between the advantage function and the gradient of the loss targeting directions of incoherence. We provide a rigorous theoretical foundation for this relationship, demonstrating that this intrinsic alignment characterizes how policies learn and internalize the underlying value function and how this process governs policy decisions. By leveraging this fundamental functional property, we propose a novel algorithm that can diagnose and identify deep neural policy decision volatilities. We conduct extensive empirical analysis in high-dimensional MDPs. From algorithmic and architectural changes to natural distributional shifts and worst-case perturbations, our proposed method can identify and audit the differences by leveraging the underlying structure of the deep neural policy manifold and the intrinsic correlation. Our analysis reveals foundational properties of policies trained in high-dimensional MDPs, and our paper provides a principled step toward constructing scalable, stable, and generalizable deep reinforcement learning agents.
PaperID: 1695, Poster
Authors: Lei Chen, Wei Zibo
Abstract: Autoregressive large language models have dominated generative recommendation by sequentially predicting the next item. However, this unidirectional paradigm suffers from error propagation, and limited ability to jointly optimize an entire recommendation list. In this paper, we introduce Full-Sequence Masked Diffusion for generative recommendation (FSMD), which extends masked diffusion models to enable global joint modeling over the entire user history and Top-K recommendation list. We formulate recommendation as a unified sequence diffusion process, where the forward process randomly masks tokens across the full sequence and the reverse process simultaneously denoises all masked positions using bidirectional context. This formulation enables parallel generation of the entire recommendation list and naturally captures inter-item dependencies within a coherent global structure. To better align generation with ranking objectives, we further introduce a unified training objective that combines masked diffusion likelihood with a listwise ranking loss. Extensive experiments on multiple public benchmarks demonstrate that FSMD consistently outperforms state-of-the-art autoregressive and item-level diffusion-based generative recommendation methods across accuracy and list quality metrics.
Abstract: Large language models remain susceptible to jailbreak attacks, despite extensive safety alignment. Smoothing-based defense methods can leverage their intrinsic safety capabilities, but the efficacy relies on the brittleness of jailbreak attacks and degrades when jailbreak contexts exhibit larger semantic margins. We observe that successful attacks typically exploit low-probability regions of the input space, where LLMs' safe bounds are difficult to reliably generalize. Thus, we propose guiding perturbations toward higher-density areas to restore the effectiveness of LLMs' safety mechanisms. In detail, we present gradient-guided smoothing that combines random noise with Gauss-Southwell type iterative ascent to the log density of the context-aware input distribution. Experimental results across four jailbreak attacks and three instruction-following benchmarks demonstrate that our method effectively improves safety while maintaining the utility of LLMs.
Abstract: Real-world multi-view data always suffer from imperfect information problem, where the view-specific observations are absent (\ie, Incomplete Views, IV) and cross-view correspondences are mismatched (\ie, Noisy Correspondences, NC) for certain instances. As a remedy, numerous IV- and NC-oriented multi-view clustering (MvC) methods have been proposed, which however require either reliable correspondences or sufficiently complete instances, thus stopping short of addressing the imperfect information problem. In contrast, we observe that both IV and NC challenges originate from the same issue of imperfect cross-view counterpart information, where the counterpart of an anchor instance in another view might be either unavailable or unreliable. Based on the observation, we propose a novel robust MvC framework, termed Posterior-guided Latent Counterpart Inference (PLCI), which could handle both IV and NC in a unified manner. Specifically, PLCI formulates the desired cross-view counterpart of each anchor instance as a latent variable, and integrates both instance-level reliability and prototype-level semantic transport to infer the posterior distribution of the latent counterpart. Extensive experiments on six widely-used multi-view datasets against 10 state-of-the-art MvC methods demonstrate the effectiveness of PLCI for tackling the imperfect information problem. The code will be released upon acceptance.
Abstract: Large-scale transformer training and deployment are increasingly constrained by the transfer of activations, gradients, and optimizer states across accelerators. Low-bit quantization offers a natural remedy, but transformer activations are often heavy-tailed and outlier-dominated, making simple quantization highly lossy. We show that this difficulty is not only a property of the quantizer, but also of the architecture. Specifically, residual connections can drive transformer activations away from Gaussianity during training. Using controlled comparisons between residual and residual-free transformers, we demonstrate that this effect leads to substantially higher quantization error and accuracy degradation at low precision in residual models. We explain the phenomenon through an excess kurtosis analysis, showing that residual mixing can amplify non-Gaussianity, whereas dense mixing in residual-free contracts non-Gaussianity. We then show that residual-free transformers can be made trainable using orthogonal initialization, spectral or second-order optimization, and depth-aware scaling of attention temperature. In language tasks, while there is a small drop in full precision performance, these models retain near-Gaussian activations and exhibit significantly improved robustness to low-bit quantization. Our results identify an accuracy-compressibility trade-off in transformer design and motivate architecture-level approaches to quantization-friendly foundation models.
PaperID: 1699, Poster
Authors:
Zegang Cheng, Zhikai Wang, Jiacheng Liu, Xiaobing Tu, Chang Zhou, Junjie Chen, Yue Ma, Zhiyuan Ma, Jinkui Ren, Likai Zou, Linfeng ZhangAbstract: Diffusion transformers have become the most powerful models for visual generation, but still suffer from massive computation costs. To solve this problem, feature caching has been proposed to cache the features of diffusion models in the previous computation steps and then reuse them in the following caching steps, which brings significant acceleration but also degradation in generation quality. To address this problem, this paper proposes Z-cache as a feature caching method that can maintain high-quality generation through self-reflection. Concretely, we observe that the error from feature caching tends to be sharply reduced after each full computation. Based on this observation, Z-Cache is designed to first predict the features in the future caching steps and then perform a full computation. After that, Z-Cache returns to the caching steps and re-predicts them based on the previous and the current computation steps, which brings correction in features. Experiments demonstrate that with Z-Cache, diffusion transformers achieve comparable generation quality to the original model but with faster inference speed, for instance, 6.22X acceleration on FLUX-dev for text-to-image generation. Our codes will be released on GitHub.
PaperID: 1700, Poster
Abstract: Deep learning models are known for their ability to memorize training data. While memorization is often linked to risks such as privacy leakage and poor robustness, a highly influential claim by~\citetfeldman2020longtail argues that memorization is actually necessary for generalization. Their conclusion is based on the observation that removing points with high memorization scores reduces test accuracy. Upon closer inspection of their work, we uncover four critical flaws in the underlying methodology: (1) sampling bias in their approximation algorithm inflates memorization scores; (2) high false positive rate in their definition of memorization, leads to misclassification of non-memorized points as memorized; (3) unprincipled thresholding that resulting in an ill-posed problem; and (4) data leakage skews the test accuracy results. To address these limitations, we introduce a modifications for correctly identifying and evaluating memorization, including higher sampling rates, modifying the original memorization definition to reduce the false positive rates, proposing a method to identify a principled score threshold, and employing test datasets especially designed to avoid data leakage. Having accounted for these errors, our results show that, in contradiction to the original work, removing truly memorized points does not cause a drop in accuracy, and in most cases, improves test performance. These findings call into question the necessity of memorization in deep learning and highlight the importance of mitigating its risks.
Abstract: We propose and analyze a model-based bootstrap for transition kernels in finite controlled Markov chains (CMCs) with possibly nonstationary or history-dependent control policies, a setting that arises naturally in offline reinforcement learning (RL) when the behavior policy generating the data is unknown. We establish distributional consistency of the bootstrap transition estimator in both a single long-chain regime and the episodic offline RL regime. The key technical tools are a novel bootstrap law of large numbers (LLN) for the visitation counts and a novel use of the martingale central limit theorem (CLT) for the bootstrap transition increments. We extend bootstrap distributional consistency to the downstream targets of offline policy evaluation (OPE) and optimal policy recovery (OPR) via the delta method by verifying Hadamard differentiability of the Bellman operators, yielding asymptotically valid confidence intervals for value and Q-functions. Experiments on the RiverSwim problem show that the proposed bootstrap confidence intervals (CIs), especially the percentile CIs, outperform the episodic bootstrap and plug-in CLT CIs, and are often close to nominal (50%, 90%, 95%) coverage, while the baselines are poorly calibrated at small sample sizes and short episode lengths.
PaperID: 1702, Poster
Abstract: Learning-based safety certification offers a promising route for verifying complex dynamical systems when explicit models are unavailable, but existing data-driven methods often scale poorly: they treat the system as a black box, rely on dense discretizations or Lipschitz bounds, and consequently suffer from sample complexity that grows exponentially with dimension. This paper shows that a common structural property of dynamical systems---monotonicity---can be used to break this curse of dimensionality. We develop a data-driven framework for robust safety verification and safe controller synthesis for unknown monotone systems using only simulator queries. Our main theoretical result proves that, for monotone systems with upper-closed unsafe sets, safety is completely characterized by the existence of monotone inductive barrier certificates. This structure reduces certificate verification over continuous state spaces to localized boundary checks over finitely many cells. To learn such certificates, we introduce Max-Monotone Networks, a monotone neural architecture with universal approximation guarantees for monotone functions. We then propose a training and refinement algorithm that, upon successful termination, returns formally valid neural barrier certificates and, under mild partition assumptions, requires only O(n) simulator queries. Across high-dimensional benchmarks, including oscillator networks with up to 13,659 states and traffic networks with 1,000 states, our method certifies safety in regimes where prior approaches either do not converge or exceed memory limits. These results demonstrate that exploiting order structure can make neural safety certification both scalable and formally sound.
PaperID: 1703, Poster
Abstract: Sparse autoencoders (SAEs) are widely used to decompose language model activations into interpretable features, but recent work finds that many features are not reproducible across random seeds. Intuitively, the more important a feature is, the more reproducible it should be. We test this hypothesis and find a surprising inversion: the 50 most important features by zero-ablation are among the least reproducible across seeds, far below dictionary-wide baselines. We trace this to residual stream anisotropy, a well-studied phenomenon in NLP embeddings that has been documented in transformer representations but not connected to SAE training. In transformer middle layers, a single principal component (PC1) explains 70--99.9% of activation variance, and its sparse decomposition is non-identifiable: different seeds learn different features that tile the same direction. These features dominate importance rankings, explaining the inversion. We propose an intervention: bypassing PC1 during SAE training, routing it around the SAE and restoring it during inference. Across 9 models from 5 families (124M--70B parameters), this intervention raises top-50 feature recovery to 74--97%. The recovered features show greater causal consistency under steering (p < 0.01 in 8 of 9 models), and shared features produce 4× more accurate logit-lens predictions than orphan features. Our results reconcile recent SAE instability findings with the existence of dense model features: the dominant PC1 direction is real and reproducible, but the individual sparse coordinates used to reconstruct it are not. Once this component is separated, a reproducible core of sparse, semantically coherent SAE features emerges.
PaperID: 1704, Poster
Abstract: Formal verification provides strong correctness guarantees, but proof development remains difficult because developers must diagnose coarse verifier errors and repair failed proof obligations. Recent approaches leverage language models (LMs) to improve proof automation through verifier-in-the-loop repair, but they often rely on hand-engineered prompts, retrieved examples, or sparse pass/fail feedback. We present \tool, an LM-based framework for Verus proof repair that internalizes counterexample-guided diagnosis as a model capability. Our key insight is to train counterexample generation not as a standalone validation task, but as a data-driven diagnostic objective: a counterexample is useful when conditioning on it helps the model produce a safe, verifier-passing proof. Given a failed Verus proof and verifier error, \tool generates a structured source-level counterexample, repair rationale, and repaired proof in one trajectory, optimized with reinforcement learning from Verus feedback and a proof-code safety gate. To address sparse and order-sensitive credit assignment, \tool introduces a topological graph reward that derives invariant dependencies from inductiveness tests and encourages counterexample generation following the same topological order as the invariant dependencies. Across seven Verus benchmarks with 471 problems, \tool achieves 61.6%/72.2% weighted-average Safe-Pass@1/Safe-Pass@3, outperforming the strongest baseline, Claude Sonnet 4.5, by 5.7/6.3 points. We also show that \tool's learned counterexample can directly augment the reasoning of existing LM-based proof repair approaches, improving their success rate by 64.3% compared to using the chain-of-thought diagnostics.
PaperID: 1705, Poster
Abstract: In real-world multimodal time-series inference, a common practice is completion–inference decoupling: missing modalities are heuristically imputed and then passed to deterministic fusion and decision making. We show that this chained design yields unreliable predictions under missingness, noise, and temporally varying modality reliability, because imputation errors propagate without uncertainty feedback. Through theoretical and empirical analysis, we identify latent-consistent completion and evidential reliability modeling as two key ingredients for robust multimodal inference. Building on these insights, we propose MD2E-MI, a manifold-aware diffusion-based imputation and evidence-driven multimodal inference framework that couples diffusion-based completion with uncertainty-aware inference in a single pathway. Experiments demonstrate stable performance across diverse missingness and noise regimes, while providing interpretable, temporally varying reliability estimates.
Abstract: Neural operators have achieved promising performance on partial differential equations (PDEs), but most existing models are built on fixed Eulerian coordinates. This mismatch between evolving physical structures and static coordinates creates spatial misalignment, leading to unnecessarily non-local operator mappings and reinforcing a smoothness preference near sharp transitions. Inspired by adaptive coordinate transformations in classical PDE analysis, we propose the Adaptive Coordinate Transform (ACT) block, a plug-and-play module for data-driven geometric adaptation in neural operators. ACT blocks resolve this structural limitation by learning adaptive coordinate systems within the operator learning pipeline. Specifically, given an input feature, the ACT block learns a coordinate transformation and represents the same feature under the transformed coordinates via differentiable sampling. This operation preserves the underlying signal while changing its spatial representation, equivalent to expressing the same physical quantity in different coordinate systems. By adapting the coordinate system to the data, ACT allows the network to better track evolving structures, reduce operator complexity, and dynamically focus on critical features to improve learning. We evaluate the proposed approach across diverse PDE benchmarks and multiple neural operator architectures. Experimental results demonstrate consistent and significant improvements in predictive accuracy, indicating that learning coordinate systems provides a powerful mechanism for enhancing operator learning.
PaperID: 1707, Poster
Abstract: A generative model is factored when each latent dimension independently controls one factor of variation: changing a single latent predictably changes one semantic attribute while leaving the rest unchanged, enabling controllable generation, compositional generalization, and reproducible representations. Existing approaches either constrain the latent distribution, e.g., requiring it to shift with an auxiliary variable, or regularize the model, e.g., sparsity, quantization, or Hessian penalties, neither of which matches how modern conditional generative models actually work: a conditioning signal u reshapes the generator g(·, u) while the latent prior stays fixed. We prove an identifiability theorem showing that generating mechanism diversity, the natural variation that arises when u sufficiently reshapes g, is sufficient for the model to be provably factored, with no parametric assumption on the latent distribution. To actively enforce this condition, we propose Mechanistic Contrastive Learning (MCL), a model-agnostic contrastive objective over generator Jacobians. Empirically, MCL achieves state-of-the-art latent concept disentanglement on three benchmarks equipped with a latent diffusion model, and improves prediction quality and zero-shot cross-task transfer in latent action world models with a 1.4B video generative model as the backbone.
Abstract: We propose a novel amortized optimization method for predicting optimal transport (OT) plans across multiple pairs of measures by leveraging Kantorovich potentials derived from sliced OT. We introduce two amortization strategies: regression-based amortization (RA-OT) and objective-based amortization (OA-OT). In RA-OT, we formulate a functional regression model that treats Kantorovich potentials from the original OT problem as responses and those obtained from sliced OT as predictors, and estimate these models via least-squares methods. In OA-OT, we estimate the parameters of the functional model by optimizing the Kantorovich dual objective. In both approaches, the predicted OT plan is subsequently recovered from the estimated potentials. As amortized OT methods, both RA-OT and OA-OT enable efficient solutions to repeated OT problems across different measure pairs by reusing information learned from prior instances to rapidly approximate new solutions. Moreover, by exploiting the structure provided by sliced OT, the proposed models are more parsimonious, independent of specific structures of the measures, such as the number of atoms in the discrete case, while achieving high accuracy. We demonstrate the effectiveness of our approaches on tasks including MNIST digit transport, color transfer, supply-demand transportation on spherical data, and mini-batch OT conditional flow matching.
PaperID: 1709, Poster
Abstract: Native visual-generation agents now handle long-horizon, multi-image workflows, but face two practical bottlenecks during reinforcement learning (RL) training. First, the inability to accumulate experience across rollouts prevents the agent from leveraging past successes and failures, slowing the evolution of its agent behaviors. Second, the reliance on black-box visual-generation tools introduces tool-induced reward noise, as the policy is unfairly penalized for failures rooted in a tool's inherent capability boundary, leading to unstable and inefficient training. To address these challenges, we propose VisionCreator-S1, an agent that casts skill-augmented RL as a nested optimization problem through two coupled innovations. First, Dual-Skills co-evolves an agent-skill set to reuse self-reflection patterns from historical rollouts, and a tool-skill set that dynamically maps each tool's capability boundary through world-reflection. Second, Skill-GRPO treats discrete text skills as learnable variables. To overcome the non-differentiability of text skill, it gracefully alternates between numerical gradient descent for the policy weights and textual gradient descent for the discrete skill sets. By contrasting high- and low-advantage trajectories, it approximates a direction-consistent textual gradient in the agent's behavior feature space. Extensive experiments on VCR-Bench, GEdit-Bench, and MultiBanana demonstrate that VisionCreator-S1 consistently outperforms the strongest skill-free baseline, VisionCreator-R1. Ablation studies further confirm that this aligned co-evolution significantly improves both training stability and the agent's long-term decision-making capabilities.
PaperID: 1710, Poster
Abstract: Post-training quantization (PTQ) is a key technique for efficiently deploying large language models, as it compresses weights into low-bit representations without retraining. Existing PTQ solvers, including GPTQ, LDLQ, and block coordinate descent, commit each weight block to a single codeword before its neighbors are considered, leaving near-optimal alternatives unexplored. In this work, we cast row-wise PTQ as posterior inference over discrete codebooks, in which existing solvers correspond to a zero-temperature, hard-decision limit retaining only the top codeword per block. We propose \emphMessage-Passing Quantization (MPQ), which preserves soft posterior information across blocks during inference and commits to hard codewords only at the output stage. MPQ first uses vector approximate message passing to compute approximate blockwise posteriors, then refines a hard-decision initialization by exact block coordinate descent. Across 2-bit lattice and 3- and 4-bit scalar quantization on dense language models from Llama, Qwen, and Mistral families, MPQ improves perplexity and zero-shot accuracy over strong hard-decision baselines, and composes favorably with preprocessing techniques such as activation-aware scaling and incoherence processing.
Abstract: Competitive selection processes, from scientific funding to admissions and hiring, use evaluations to score candidates, and then choose a subset based on those scores. Recently, many organizations have adopted partial lotteries, which randomize selection based on evaluation scores. However, existing lottery designs are inherently unstable, as a small change to a single candidate's score can cause large shifts in their selection probabilities. This instability undermines a key goal of lotteries: reducing the influence of fine-grained score distinctions near the decision boundary. We propose smoothness as a design principle for partial lotteries, formalizing it as a Lipschitz condition on the mapping from review scores over candidates to selection probabilities. We propose using a Linear Lottery, a simple mechanism in which selection probabilities scale linearly with aggregated review scores between an upper threshold, above which we always accept and a lower threshold, below which we always reject. We prove that the Linear Lottery's worst-case regret matches a lower bound for any smooth selection rule up to a factor of (1 - k/n), where k/n is the acceptance rate. We compare smooth selection to other stability notions like Individual Fairness and Differential Privacy, showing that the Linear Lottery achieves a better smoothness--regret tradeoff than alternatives. Experiments on real peer review data from ICLR 2025, NeurIPS 2024, and the Swiss NSF confirm our theoretical analysis is tight and demonstrate that existing lottery designs are highly unstable in practice even under changes to a single review score.
PaperID: 1712, Poster
Authors: Gun Ryu, Gi-Mun Um, Won-Sik Cheong, Wonjun Kim
Abstract: Recent studies on generalizable human Gaussian splatting have achieved the notable progress in synthesizing novel views of unseen human subjects from multi-view images. Despite this progress, existing methods still struggle to represent fine-grained details of human subjects, which inherently involve high geometric and textural complexity. This is because previous methods do not impose a constraint on how densely Gaussians are placed, which often leads to blurry rendering results in complex regions. To address this problem, we propose a complexity-guided regularization scheme. The key idea is to leverage a complexity prior computed via the local entropy in geometric and textural characteristics of a clothed human. This effectively prevents Gaussians from being coarsely placed in complex regions, which facilitates the faithful reconstruction of intricate structures. Furthermore, we sample local patches predominantly from complex regions according to this complexity prior, which allows the model to be sufficiently supervised by local details. Experimental results on benchmark datasets demonstrate that our method delivers state-of-the-art performance while maintaining the real-time inference speed in generalizable human Gaussian splatting.
PaperID: 1713, Poster
Abstract: Deep neural networks have achieved significant success on in-distribution (ID) tasks, but out-of-distribution (OOD) samples during inference can lead to unreliable predictions. To address this issue, distance-based methods exploit the geometric structure of feature space for OOD detection. However, methods based on static information geometry fail to address geometry distorted by ill-distributed samples. Although prior work reduces the misclassification of OOD samples as ID by adjusting the residual space with real-time features, it fails to address the misclassification of ID samples as OOD. In addition, existing distance-based methods ignore class-specific characteristics when computing deviation feature and rely on the nearest prototype for uncertainty estimation, leading to suboptimal OOD detection performance. To address these issues, we propose Adaptive Mahalanobis Gap (AMG). Specifically, we first strengthen the principal space along real-time feature directions to reduce the misclassification of ID samples as OOD. Subsequently, we incorporate class variance and prototype norms to mitigate bias of deviation feature from class-wise distributional differences. Finally, we introduce the Mahalanobis gap scoring for uncertainty scoring to overcome bias from single-prototype reliance. Experiments on CIFAR-10/100 and ImageNet-1k show that AMG consistently outperforms state-of-the-art methods in FPR@95 and AUROC. Our code is available at https://anonymous.4open.science/r/AMG_AR.
PaperID: 1714, Poster
Abstract: While recent multimodal recommender systems have demonstrated the effectiveness of incorporating visual and textual information to improve downstream performance, most existing methods rely on static modality fusion, assuming that the relative importance of textual and visual signals remains stable across recommendation scenarios. This design may not fully account for an important variation across recommendation requests: some queries require fine-grained visual cues, whereas others are better served by textual or functional semantics, in which case indiscriminate modality fusion brings in uninformative cues and impairs recommendation quality. To address this, we propose AdaM-Rec, an LLM-based framework for adaptive modality routing in multimodal recommendation, which enables dynamic calibration of reliance on textual and multimodal evidence for user-specific queries. Built on structured natural-language representations of items and user preferences, it estimates modality reliability using proxy recall tasks. Specifically, it generates pseudo-queries that match the granularity of the actual query while pointing to the user's positively interacted items as verifiable proxy targets, evaluating which modality yields better recall performance in analogous scenarios and optimizing the routing strategy in an agentic manner. It then performs routed recall with optimized strategy, enriches results with collaborative items, and ranks candidates by their relevance to both the query and user preferences. Experiments demonstrate that AdaM-Rec delivers strong performance against state-of-the-art baselines, highlighting the effectiveness and broader potential of adaptive control over modality reliance in multimodal recommendation. Code will be publicly available.
Abstract: Combining the predictions of multiple models is one of the most effective strategies in machine learning. Thus, it is only natural to investigate the ensembling of language models (LMs). However, because LMs define distributions over strings, a proper ensemble requires aggregating predictions at the string level, which requires a computation that is generally intractable. In this work, we cast LM ensembling as an inference problem and propose the deployment of sequential Monte Carlo (SMC), an appropriate approximate inference scheme. Concretely, we introduce a unified framework for composing multiple LMs into ensemble distributions parameterized by a broad family of aggregation functions. To sample from these distributions, we introduce an SMC algorithm that operates in a shared character space, enabling ensembles of models with mismatching vocabularies and consistent sampling in the limit. We evaluate a range of ensembles across prompt and model combinations for various text generation tasks, finding that consensus-seeking strategies such as the product consistently outperform probability averaging, and that better posterior approximations can yield better ensemble performance.
PaperID: 1716, Poster
Abstract: A room impulse response (RIR) encodes how sound reflects, scatters, and decays inside a room, and therefore what the room looks like and what it is made of. We formalize this observation as Acoustic Inverse Rendering (AIR): recovering dense panoramic geometry and per-band acoustic reflectance of a scene directly from its RIR. The task is challenging as the RIR collapses all propagation paths into a single temporal signal, and the resulting inverse mapping is severely ambiguous. AIR closes this gap by beamforming the RIR onto the equirectangular grid and conditioning a diffusion model on the resulting directional features, yielding a distribution over plausible scenes rather than a single point estimate. The same spherical beamforming applies to FOA, binaural, or mono microphone setups. On both simulated and real-world datasets, AIR produces coherent room geometry and per-band acoustic reflectance maps from acoustic input alone, significantly outperforming prior baselines on this inverse problem.
Abstract: Visual content generation has moved from single-image to multi-image workflows, yet existing agents are largely plan-driven and lack a systematic reflection mechanism for correcting mid-trajectory visual errors. To address this gap, we propose VisionCreator-R1, a native visual generation agent with explicit reflection, paired with a Reflection-Plan Co-Optimization (RPCO) training methodology. Through extensive experiments and trajectory-level analysis, we uncover a reflection-plan optimization asymmetry in reinforcement learning (RL): planning can be reliably optimized via plan rewards, while reflection learning is held back by noisy credit assignment. Motivated by this finding, we further abstract a general Decouple-then-Fuse paradigm for co-optimizing capabilities with asymmetric reward variance, of which RPCO is the visual-generation instantiation. Our RPCO first trains on the self-constructed VCR-SFT dataset, which covers reflection-strong single-image trajectories and planning-strong multi-image trajectories, and then co-optimizes on the VCR-RL dataset via RL. This yields our unified VisionCreator-R1 agent, which beats Gemini2.5Pro on existing benchmarks and on our VCR-Bench covering single-image and multi-image tasks.
Abstract: Many bandit systems are deployed with offline historical data, such as past logs from earlier policies. By using these data, one can reduce the amount of online exploration needed early in learning. When the offline and online environments differ, such data can be biased for the online problem. For linear (contextual) bandits, this bias is directional: offline data may be informative in some feature directions and misleading in others. However, prior work typically controls this gap through a known Euclidean bound on the model parameters, which we prove is too coarse to separate easy and hard offline-to-online instances from offline data alone. To address this challenge, we introduce a directional bias certificate (M_\mathrmbias,\rho) that measures the offline-to-online gap through an M_\mathrmbias-induced norm and assigns different bias budgets to different directions. Building on this certificate, we propose \emphEllipsoidal-MINUCB, which augments the online learning with an offline-pooled branch that safely exploits historical data. When the certificate is known, we show that the algorithm matches the standard SupLinUCB rate in the worst case and improves when offline coverage aligns with low-bias directions. When the certificate is unknown, we estimate it adaptively from offline and accumulated online data and establish a corresponding regret guarantee. Numerical experiments support the theory and show gains in aligned regimes.
PaperID: 1719, Poster
Abstract: Language model post-training is often bottlenecked by the need for human-collected preference data, which is expensive and difficult to scale. RLAIF-style approaches that leverage pseudo labels offer an abundant alternative but introduce systematic biases that degrade downstream alignment. Recent general-purpose semi-supervised methods correct for teacher bias using a small set of human-labeled examples, but suffer from high variance especially when human annotations are scarce. To this end, we propose ABC-Align, leveraging abundant pseudo label signal to minimize variance and applying a lightweight, adaptive correction grounded in the human-labeled subset. The correction strength is tuned automatically during training using plug-in estimates of the relevant bias--variance quantities. On LLM alignment with RLHF and DPO where human feedback is scarce, we empirically demonstrate that ABC-Align achieves superior performance over prior semi-supervised baselines.
PaperID: 1720, Poster
Abstract: Hypergraphs record multi-way interactions among entities. Extracting information from the combinatorial structure underlying observed multi-way interactions is a central task in many real-world problems. Existing methods face several limitations. First, many deep architectures for hypergraphs do not explicitly exploit the potential low-rank structure, which can sacrifice parsimony and interpretability in the learned representations. Second, many low-rank-based methods operate on tensor representations, which typically require hyperlinks to have uniform sizes and thus limit their applicability to general hypergraphs with non-uniform hyperlink sizes. Third, many methods ignore the fact that hyperlinks often arise from heterogeneous mechanisms. For example, medical symptoms may co-occur in the profiles of patients with very different conditions, and such heterogeneity should be incorporated into the learning process. In this work, we develop a general framework for hypergraph representation learning using hyperlink random effects while exploiting the low-rank structure in hypergraphs. The proposed framework accommodates latent heterogeneity in hyperlink formation while preserving entity interaction patterns. We establish identifiability of the model parameters and leverage an expectation-maximization strategy for estimation. The framework allows flexible specifications for the hyperlink random effects; in this paper, we study three choices: categorical, Gaussian mixture, and score-based effects, and develop corresponding estimation algorithms. Through simulation studies, we demonstrate the effectiveness of the proposed method in recovering latent structure and capturing heterogeneous interaction patterns. Empirical studies on real-world hypergraph datasets further illustrate the practical utility of our approach.
Authors:
Haozheng Luo, Haoran Dai, Shaoyang Zhang, Xi Chen, Eric Hanchen Jiang, Yijiang Li, Michael Ching-yuen Huang, Chenghao Qiu, Chenwei Xu, Zhenyu Pan, Haotian Zhang, Binghui Wang, Yan ChenAbstract: We propose OASIS, an outlier- and sink-aware technique built on inter-layer null signaling. As AttnResidual architectures introduce an additional depth-wise normalization channel, they improve inter-layer routing flexibility but also exacerbate attention sinks, activation outliers, and the resulting degradation in inference stability and quantization robustness. OASIS addresses this issue by introducing a \mathop\rmSoftmax_1-based null space and coupling token-level null evidence to depth routing through an inter-layer null signal, thereby reducing sink-dominated routing and improving structural robustness. Theoretically, we show that the dual-normalization design of AttnResidual intensifies sink formation and quantization brittleness. Experimentally, we compare OASIS against five baselines on three real-world datasets and observe consistent improvements in both attention sink and post-quantization performance. Notably, OASIS achieves an average reduction of 9.26% in maximum infinity norm and 2.60% in average kurtosis across the evaluated settings, while lowering perplexity by 75.85% under W8A8 and improving GSM8K Pass@1 by 12.42% under W4A4.
Authors:
Cheng Chi_, Xianqi Wang, Hongcheng Luo, Mingfei Tu, GANGWEI XU, Zehan Zhang, Bing Wang, Guang Chen, Hangjun Ye, Sida Peng, Xin Yang, Haiyang SunAbstract: High-fidelity reconstruction of driving scenes is crucial for autonomous driving. While recent feedforward 3D Gaussian Splatting (3DGS) methods enable fast reconstruction, their per-pixel Gaussian prediction paradigm often suffers from multi-view inconsistency and layering artifacts. Moreover, existing methods often model dynamic instances via dense flow prediction, which lacks explicit cross-view correspondence and instance-level consistency. In this paper, we propose PointForward, a feedforward driving reconstruction framework through point-aligned representations. Unlike pixel-aligned methods, we initialize sparse 3D queries in world space and aggregate multi-view image information via spatial-temporal fusion onto these queries, enforcing explicit cross-view consistency in a single feedforward pass. To handle scene dynamics, we introduce scene graphs that explicitly organize moving instances during reconstruction. By leveraging 3D bounding boxes, our method enables instance-level motion propagation and temporally consistent dynamic representations. Extensive experiments demonstrate that PointForward achieves state-of-the-art performance on large-scale driving benchmarks. The code will be available upon the publication of the paper.
Authors: M. Moran, Mark E Whiting
Abstract: Confidence-weighted routing, selective abstention, and ensemble weighting all assume that a model's stated confidence is informative about its capability on the question being asked. They presume functional metacognition, the capacity to assess one's own capabilities, without exercising them, relative to other frontier language models. Aggregate calibration is well studied, with mixed results, but the underlying structure of elicited confidence is less well understood. We decompose binary confidence judgements from 20 frontier Large Language Models (LLMs) across six benchmarks using tetrachoric factor analysis paired with pairwise calibration, asking whether two models that differ in confidence also differ in performance. On factual recall and information retrieval benchmarks the cross-model confidence matrix is approximately rank-one and a single dominant factor captures most of the latent variance. Models retrieving facts share an item-level difficulty axis and differ mainly in their decision thresholds along it. Across all benchmarks the relationship between confidence and performance collapses once items that all models agree on are removed. Inter-model pairwise calibration is small even where statistically significant, and what remains shrinks to nothing once base-rate differences along the shared factor are controlled for. Mathematical reasoning is the apparent exception, but this turns out to be a confound where reasoning models answer questions about their confidence by trying to solve them in their chain of thought, bypassing the sub-symbolic self-knowledge we seek to measure. We find no evidence for significant verbalised individuated metacognition in any tested domain. A trivial logistic regression on surface features of the confidence-judgement reasoning text beats the median confidence model's repeatedly sampled judgements on half the benchmarks, evidence that extracting uncertainty from the reasoning trace recovers signal the binary judgement does not.
PaperID: 1724, Poster
Abstract: Long-context training of large language and multimodal models couples expressive structured attention masks (document, prefix-LM, share-question, sliding-window, block-sparse) with multi-node context parallelism, exposing the mask as a dominant bottleneck. The column-wise sparse mask of FlashMask compresses such masks from O(n^2) to O(n) via four query-axis boundary vectors and covers a broad family of structured patterns. Implemented for Ampere GPUs, FlashMask does not exploit FlashAttention-3's TMA/WGMMA warp-specialized pipeline on Hopper GPUs, and offers no distributed support. We observe that the column-wise sparse mask is indexed by key positions but stores query-axis interval boundaries. Under any multi-chunk context-parallel assignment, this property reduces per-rank mask localization to a provably zero-communication \emphclip--shift--max on the boundary vectors, directly compatible with the single-GPU kernel. Building on this property, FlashMask-3 introduces a mask-aware Hopper kernel with a three-role warp specialization atop the FlashAttention-3 TMA/WGMMA pipeline, together with a mask-aware distributed extension: zero-communication mask localization, a sparsity-aware load balancer (HCC-IPO), and a mask-aware compute-communication overlap with mask-driven communication pruning, whose forward all-gather additionally adopts topology-aware hierarchical routing. On a single H100 GPU across 12 structured masks at context lengths from 4K to 128K, FlashMask-3 improves achieved TFLOPs over FlashMask, FlexAttention, and MagiAttention by 45% to 141%, 34% to 71%, and 2% to 44%, respectively. Under context-parallel degrees from 4 to 32 with 8K sequence per device, it attains 31% to 75% compute-communication overlap efficiency and improves throughput over Megatron-Core CP (ring), Megatron-Core CP (all-gather) and MagiAttention by 3.25× to 5.11×, 1.97× to 3.26× and 1.06× to 4.09× respectively. The source code will be released on GitHub.
Authors: Önder Askin, Tim Kutta, Holger Dette
Abstract: Auditing differential privacy has emerged as an important area of research that supports the design of privacy-preserving mechanisms. Privacy audits help to obtain empirical estimates of the privacy parameter, to expose flawed implementations of algorithms and to compare practical with theoretical privacy guarantees. In this work, we investigate an unexplored facet of privacy auditing: the sustained auditing of a mechanism that can go through changes during its development or deployment. Monitoring the privacy of algorithms over time comes with specific challenges. Running state-of-the-art (static) auditors repeatedly requires excessive sampling efforts, while the reliability of such methods deteriorates over time without proper adjustments. To overcome these obstacles, we present a new monitoring procedure that extracts information from the entire deployment history of the algorithm. This allows us to reduce sampling efforts, while sustaining reliable outcomes of our auditor. We derive formal guarantees with regard to the soundness of our methods and evaluate their performance for important mechanisms from the literature. Our theoretical findings and experiments demonstrate the efficacy of our approach.
Abstract: A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings. However, different methods and their resulting insights stand in relative isolation: what could the underlying structure of language models be, such that they give rise to all our interpretations? In this work, we propose using Tensor Product Representations (TPRs; Smolensky 1990) as a unifying hypothesis. TPRs give a concrete proposal for how compositional structure can be represented in vector space. We show, both mathematically and empirically, that TPRs can unify several prior interpretability methods: additive analogies, linear probing, sparse autoencoders, and activation patching. Mathematically, we show that these methods can all be derived from TPRs. Empirically, we apply the derivations to a range of different models — from small toy models to LLMs — to construct instances of each of the above interpretability methods; these constructed variants perform comparably to their standard variants. We view this work as a step toward what interpretability will ideally provide: a unified account of the nature of neural networks, corroborated not just by individual observations but also by an explanation of the connections between them.
Abstract: Articulated objects used in simulation and embodied AI are typically specified by geometry and kinematic structure, but lack the fine-grained dynamical effects that govern realistic mechanical behavior, such as frictional holding, detents, soft closing, and snap latching. Existing approaches either ignore the detailed structure of dynamics entirely, or use simple models with limited expressiveness. We introduce JODA, a framework for generating joint-level dynamics as a structured three-channel field over the joint degree of freedom, capturing conservative forces, dry friction, and damping. Instantiated using shape-constrained piecewise cubic interpolation (PCHIP), this formulation defines a compact and expressive function space that is both interpretable and compatible with differentiable simulation. Building on this representation, we develop methods for inferring and refining joint dynamics from multimodal inputs. Given visual observations and joint context, a vision-language model proposes structured dynamical primitives, which are composed into a unified dynamics field. The resulting representation supports both direct manipulation and gradient-based refinement. We demonstrate that JODA enables plausible and controllable modeling of diverse joint behaviors, providing a unified interface for inference, editing, and optimization.
PaperID: 1728, Poster
Abstract: Image-to-video generation aims to animate an image while preserving appearance details and producing visually coherent future frames. Autoregressive generation models this process as a sequence of next-frame predictions conditioned on the input image and text prompt. However, advanced autoregressive generators built on diffusion or flow models formulate each step as a noise-to-data process, despite the fact that the previous frame already provides an informative visual prior. This mismatch makes next-frame synthesis harder than necessary and can degrade prediction quality, since the model must recover scene layout and visual context from uninformative Gaussian noise at every step. To this end, we introduce BRAVO, a Brownian bridge matching framework for autoregressive video generation. BRAVO replaces the conventional noise-to-data path with a data-to-data bridge from the previous frame to the current one, turning next-frame synthesis into local frame-to-frame transport. The bridge state is formed by interpolating between adjacent real frames with a Brownian perturbation, so the learned velocity models the local transition from the previous frame to the next under the text prompt and causal visual history. BRAVO is trained with teacher-forcing and sampled autoregressively, where each bridge starts from the last generated frame and reuses historical messages through a causal KV-cache. Experiments on different datasets show that BRAVO surpasses the conventional noise-to-data baseline and achieves even competitive performance with substantially larger models.
PaperID: 1729, Poster
Abstract: Deployed reinforcement learning agents often face safety requirements that are specified only after training: new hazard maps, revised risk thresholds, or behavioral alignment constraints. We study zero-update deployment-time adaptation in which a fixed library of risk-neutral source policies must be reused under a newly specified reward–risk tradeoff. We propose ), a source-scored composition rule that evaluates each source under the target reward and an occupancy-based deployment risk, then selects actions using risk-adjusted source scores. Unlike training-time risk-sensitive methods tied to a fixed surrogate such as return variance, TRAM supports spatial barrier exposure, divergence to a reference behavior, and local volatility risks specified at test time. We make the surrogate nature of the method explicit: TRAM is not claimed to solve the full occupancy-control problem of the stitched policy, but it admits a measurable source-hull mismatch term connecting source-scored risk to realized risk. Experiments in gridworlds, MuJoCo Reacher, Safety-Gymnasium, and an LLM alignment setting show that TRAM improves deployment risk while preserving reward and requires no parameter updates at test time.
PaperID: 1730, Poster
Abstract: How novel ideas are discovered and spread through a population is a well-studied phenomenon in the fields of computer science, computational social science, marketing/business research and higher education studies. In this paper, we study a particular version of this phenomenon, namely the spread of ideas among a network of researcher agents. We use a computational model in which each agent has a finite computational budget, that it can spend on a mixture of conducting its own research, and learning from the research of others in the network. By framing the problem as a continuous optimisation on the edge weights, we can applying gradient-based optimisation, similar to that from graph neural networks, to search for the graph configuration that maximises the amount of knowledge uncovered by the network. This contrasts with existing works that simply compare a fixed set of configurations. We find that (i) the optimisation succeeds in producing graphs that lead to a high total knowledge score, (ii) the optimum configuration has somewhat short shortest path lengths and a low clustering coefficient, (iii) the nodes end up broadly divided between researchers, who focus solely on their own research, and learners, who focus mostly on viewing research of others, and (iv) the learners consistently have higher knowledge scores, and this drives a rise in inequality across nodes over time.
PaperID: 1731, Poster
Abstract: Image-free zero-shot classifier expansion aims to extend recognition to unseen classes by establishing an association mapping between class semantics and visual classifier weights, without accessing original images. However, most existing methods rely on an association-based optimization mechanism, which struggles to meet the demands for fine-grained classification of visually and semantically similar classes. Since co-occurring attributes dominate class semantics, the model gravitates towards co-occurring attributes rather than discriminative attributes. This leads to geometrically entangled classifier weights, resulting in ambiguous decision boundaries. To address this problem, this paper proposes a novel approach that reframes weight generation through the lens of causal intervention. The method first employs comparative reasoning to disentangle discriminative and co-occurring attributes. Subsequently, it synthesizes intervened semantics to impose invariance and sensitivity constraints. This mechanism compels the model to suppress co-occurring attributes while enhancing sensitivity to discriminative attributes. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches, yielding clearer decision boundaries in fine-grained settings.
PaperID: 1732, Poster
Abstract: Current world models for explicit world simulation typically follow two distinct pathways: (1) video generative models (e.g., Sora), which simulate realistic appearances but fail to preserve 3D consistency; and (2) 3D generative models (e.g., Marble), which generate geometrically consistent scenes but often yield static, less realistic observations. We present AEON, a unified framework that bridges video-based and 3D-based world models and combines the benefits of both. The key insight of AEON lies in a generative model that operates within the video latent space while producing renderable 4D graphics primitives—specifically, 4D Gaussian Splatting (4DGS)—as the world modeler. We train a 4D reconstruction model to predict both 4DGS and camera poses from monocular videos, serving as our initialized 4D decoder. Furthermore, we design a domain adapter that translates generative video latents into reconstructive 4D latents through a distillation process on the video VAE. AEON naturally supports video-based world modeling by rendering 4DGS at predicted camera trajectories, and constructs a persistent 4D environment with scene dynamics for realistic world simulation. Consequently, AEON empowers a diverse range of world modeling applications and achieves state-of-the-art performance on standard benchmarks.
PaperID: 1733, Poster
Authors: Jongyeon Lee, Jaehyoung Jeon, Kim Haneol, Myungjoo Kang
Abstract: Value overestimation is a central challenge in offline reinforcement learning (RL), arising when Bellman backups propagate out-of-distribution (OOD) errors that lead to catastrophic inflation of Q-values. Prior approaches address this issue by encouraging the learned policy to stay close to the behavior policy, but they often result in suboptimal policies or require significant training complexity. More recent work attempts to tackle this issue through minimal modifications to the critic network, yet it remains limited in either theoretical grounding or practical implementation. In this work, we present Kernel Value Regression (KVR), a simple critic regularization method based on kernel regression. KVR builds on the vanishing extrapolation property of kernel ridge regression, which encourages value estimates to decay toward zero away from the data support. This inductive bias naturally suppresses Q-values for OOD actions without auxiliary regularizers, providing a structural mechanism for addressing value overestimation. Using random Fourier features, KVR can be incorporated into standard actor-critic algorithms with only minimal modifications. We empirically demonstrate that these properties persist in complex offline RL environments, leading to strong performance across diverse tasks in the D4RL and OGBench benchmarks.
PaperID: 1734, Poster
Abstract: The recent empirical success of non-Euclidean optimizers, such as Muon (Jordan et al., 2024) and Shampoo (Gupta et al., 2018), has transformed the training of Large Language Models (LLMs) by exploiting the geometry of matrix spaces. Despite this progress, the theoretical understanding of stochastic optimization in general normed spaces remains limited, particularly in distributed settings. In this work, we present a novel analysis of Minibatch Stochastic Gradient Descent (SGD) in non-Euclidean geometries. Leveraging the framework of \\kappa-regular normed spaces (Juditsky & Nemirovski, 2008) and a light-tail noise model defined directly in the dual norm, we derive high-probability oracle complexity bounds for smooth non-convex optimization. Crucially, our analysis reveals that measuring stochastic variance within the intrinsic non-Euclidean geometry avoids the suboptimal dimension-dependent factors inherent in Euclidean norm equivalence. Building on this framework, we introduce Rennala-NSGD, a semi-asynchronous distributed algorithm, and establish, to the best of our knowledge, the first wall-clock time complexity guarantee for asynchronous non-Euclidean stochastic optimization.
PaperID: 1735, Poster
Abstract: What is the right representation for 4D (3D + time) generation of graphics assets conditioned on their current states? In this setting, an ideal 4D representation should support flexible geometric structure, empower long-horizon modeling, and utilize training data efficiently. We study this problem by considering a sufficiently challenging setting to stress test those properties: learning the generative distribution of the entire 4D lifespan of a plant, conditioned on a snapshot of its growth. We curate a large-scale dataset and benchmark for this setting, on which most existing 4D representations and accompanying generative models fail to deliver promising results. We analyze the pitfalls of existing representations and point out that the capacity of existing models is bounded by their representations: they are mostly augmenting a temporal dimension to an existing geometric representation, without a semantically-meaningful low-dimensional state representation. Instead, we argue for a framework for designing 4D representations by separately considering 3D state representations and temporal state transitions. We propose an instantiation of this framework using latent spaces of 3D foundation models and a transformer-based state transition model, effectively forming a finite state machine on a latent 3D space. Therefore, we name this instantiation a and compare it against existing generative models on the proposed benchmark. Our evaluation shows that our model achieves state-of-the-art performance in expressiveness, data efficiency, and rendering quality, surpassing existing methods by a large margin.
PaperID: 1736, Poster
Authors: Amir Asiaee
Abstract: Dense fine-tuning can turn a base language model into a domain specialist, but task accuracy alone does not reveal which internal computations mediate the new behavior. We introduce : internal mechanisms through which a dense specialist's lift over its base causally flows. We identify these mediators by patching specialist activations into the base model and asking which patches recover in-domain performance while minimally perturbing off-domain behavior. This separates three questions that are often conflated: where the specialist behavior is localized, how much of it can be recovered by activation patching, and whether a sparse trainable update can implement the same behavior. We formalize these links with exact recovery results, a first-order recovery approximation, and a bound showing when patching recovery can transfer to localized training. In synthetic transformers and structural causal models with planted mediators, patching recovers the true causal mechanisms even when activation-magnitude baselines are misled by non-causal distractors. In Pythia-410M specialists for math (GSM8K), multitask knowledge (MMLU), and biomedical QA (PubMedQA), single-patch localization ranks first on every domain and seed by area under the recovery curve, with small mechanism budgets recovering much of the patchable specialist gap. A localized-LoRA stress test then shows why the separation matters: patching-localized adapters improve negative log-likelihood when the selected attention mechanisms carry the specialist gap, but fail on GSM8K where MLP blocks carry substantial lift. Thus specialist localization is useful both as a route to sparse adaptation and as a diagnostic for when sparse adaptation fails.
PaperID: 1737, Poster
Abstract: Native multimodal large language models (native MLLMs) jointly encode visual and linguistic sequences within a unified autoregressive Transformer, emerging as a promising paradigm for multimodal understanding. Unlike conventional ViT-based MLLMs that rely on an external vision encoder, ViT-free native MLLMs tokenize every pixel patch directly, which is intuitively expected to suffer from severe visual token redundancy. To examine this hypothesis, we conduct a pilot study and find that, unlike ViT-based models which degrade under compression, ViT-free MLLMs tolerate aggressive compression with near-lossless or even improved per9 formance. Existing compression methods are built exclusively for the ViT-based setting, whereas native MLLMs, which stand to gain the most from visual token compression, remain entirely unexplored. To close this gap, we introduce an analysis pipeline for native backbones, centered on two metrics, i.e., token discriminability and cross-dimensional heterogeneous redundancy, which together reveal the following two findings: (i) redundancy is distributed independently and heterogeneously along the height and width axes, and (ii) in early layers, visual tokens are indiscriminable and uniformly redundant. Guided by these findings, we propose PACE (Pareto Adaptive Compression for Efficient native MLLM), the first visual token compression framework tailored for native MLLMs with Multimodal-RoPE. Leveraging the dimensional decoupling of position encoding, PACE decomposes attention into two orthogonal spatial saliency estimates along height and width, gated by a positional attention signal, yielding a direction-aware importance measure with negligible overhead. It further casts token budget allocation as a Pareto optimization over three objectives, i.e., performance preservation, inference efficiency, and spatial coverage, which jointly identify the optimal compression depths and retention ratios. Experiments show that PACE maintains near-lossless performance at only 25%–50% visual token retention, and even achieves performance gains on several benchmarks, while delivering over 2× FLOPs reduction.
Authors: Yuta Tarumi
Abstract: Many physical data assimilation (DA) workflows require smoothing methods that represent non-Gaussian posteriors over physical state variables, scale to high-dimensional simulators, train from observation windows alone, and remain compatible with calibration of the prescribed simulator. We introduce PR-Smoother, a simulator-preserving amortized smoother designed for this prescribed-simulator DA regime. Its key design principle is to keep the prescribed simulator explicit in both the evidence lower bound and the variational family: rather than learning replacement dynamics or a learned trajectory prior, PR-Smoother learns only future-conditioned corrections around the prescribed rollout. This yields an explicit non-Gaussian smoothing distribution over physical trajectories and supports joint state, parameter, and sensor-bias learning from observations alone. The variational family contains the exact smoother in deterministic and linear-Gaussian limits. Empirically, PR-Smoother captures multimodal posteriors in 4-dimensional Lorenz–96, remains accurate under ambiguous nonlinear observations and process noise in 40-dimensional Lorenz–96, and scales to joint state-parameter-bias inference in 16,384-dimensional Kolmogorov flow.
PaperID: 1739, Poster
Abstract: Recent work on agentic training efficiency has mainly focused on improving learning algorithms, rollout strategies, or model-hardware co-design. However, these approaches overlook an important source of training latency: the sandbox that runs the environments. Recent studies show that creating, executing, and resetting containers accounts for an important portion of the rollout process in RL training, substantially slowing down overall training. This overhead largely stems from the centralized management mechanisms (e.g., DAEMON) of existing sandboxes, along with additional features (e.g., multi-tenancy) that are unnecessary for reinforcement learning (RL). Moreover, existing sandboxes are designed for general-purpose isolated execution, making them a poor fit for the requirements of RL rollout and training. We propose NitroBox, a novel sandbox system purpose-built for agentic RL that significantly accelerates the rollout and training process. Technically speaking, we first comprehensively redesign NitroBox's lifecycle to eliminate the dependency on DAEMON, so that every sandbox instance is fully independent during creation, startup, and reset (e.g., via rootless full-namespace isolation). Furthermore, we introduce a copy-on-write mechanism to enable fast startup across multiple instances. Second, we knock off features that are time-consuming but unnecessary for RL — such as multi-tenancy — while preserving the core functionality and security guarantees needed for RL training. Finally, we introduce some RL-nature features to further improve efficiency, such as optional CRIU-based process checkpoint saving and restoring for mid-trajectory branching. Our experiments show that NitroBox is faster than Docker across the entire lifecycle: 10.6× faster in sandbox creation, 24.6× in reset, and 19.1× in deletion. When training Qwen2.5-Coder-32B-Instruct with GRPO at c=32 concurrency, end-to-end wall-clock training time drops from 40.6 hours to 28.4 hours when replacing Docker with NitroBox, while preserving the same task pass rate.
PaperID: 1740, Poster
Abstract: We propose a deep learning framework for scalar-on-function regression with possibly multiple functional predictors observed on irregular grids. The method represents each functional covariate by finitely many basis coefficients computed from the observed trajectories and fits a deep neural network using these coefficients as features. We establish a general nonasymptotic excess risk bound for the neural network estimator that decomposes the error into statistical, truncation, discretization, and network approximation terms. We then specialize the theory to two important settings. (i) For continuous functional models, an equal-width neural network estimator achieves a nearly minimax optimal rate up to a log-logarithm factor. (ii) For functional multiple-index models, we design a bottleneck-equal-width neural network estimator that exploits the low-dimensional index-link structure and attains a nearly minimax optimal rate up to a logarithm factor. The results provide theoretical support for neural network-based scalar-on-function regression and highlight the importance of architecture design.
PaperID: 1741, Poster
Abstract: Large Language Model training is increasingly constrained by memory. Activation checkpointing reduces the activation memory footprint by discarding selected activations during the forward pass and rematerializing them during backpropagation, trading memory savings for additional FLOPs. Although essential for scaling models, this overhead makes rematerializing linear layers of transformers prohibitively expensive. We introduce a novel method for approximate rematerialization that uses information from the initial forward pass to selectively recompute only a fraction of the original layer components. These components are either chosen greedily, or sketched, yielding unbiased gradient estimates. Through ablation studies, we show that our method closely matches the original gradient. When applied to FFN layers, attention layers, or both, our method reduces the computational cost for the rematerialization and backward pass of linear layers by 33% on a Llama-488M, with perplexity increases of only 1.7%, 1.2%, and 2.5%.
PaperID: 1742, Poster
Authors: Jong-Hoon Park, Young-Rae Cho
Abstract: Electronic health records (EHRs) are a valuable resource for clinical prediction because they capture rich longitudinal information about patient care. Existing graph-based EHR models often learn relations at the cohort level, but do not clearly distinguish between global event associations and the subset of interactions actually instantiated within an individual visit. To address this limitation, we propose MVP-EHR, a multi-view prototype learning framework for EHR prediction with shared type-aware prototype alignment. MVP-EHR combines three key ideas: (i) a lift-thresholded global knowledge graph that captures cohort-level event associations, (ii) a visit-induced local knowledge graph that preserves only the observed events and supported event-event relations within each visit, and (iii) shared type-aware prototypes that align local and global event representations before post-event multi-view fusion and temporal prediction. Experiments on MIMIC-IV show that MVP-EHR achieves strong and consistent performance across multiple clinical prediction tasks, including 0.9339 AUROC and 0.1706 AUPRC for mortality prediction, the best AUROC on readmission prediction, the best overall performance on phenotype prediction, and a strong AUPRC/F1 tradeoff on drug recommendation among the compared baselines. Ablation results further show that separating global prior from visit-instantiated evidence, together with shared prototype alignment, is critical for robust performance.
PaperID: 1743, Poster
Abstract: Many latent neural operators represent input and output fields in a stationary latent chart. In particular, common latent routing mechanisms use fixed or shared assignment weights for feature projection and reconstruction, limiting their ability to model transport-dominated systems where coherent structures move relative to fixed coordinate frames. We propose , a transport-aware latent operator that introduces a regularized kinematic coordinate map to decouple source and target coordinate systems. This yields an approximately co-moving latent reference frame and enables asymmetric feature aggregation and reconstruction. Combined with a geometry-aware ordering mechanism for state-space models, Advectra captures advective dynamics while maintaining stable global interactions. Advectra achieves the best performance among evaluated geometry-constrained and form-free baselines on advection-dominated benchmarks, including passive scalar transport in Navier--Stokes flows and Rayleigh--Taylor instability, while demonstrating strong generalization on real-world engineering tasks. These results highlight the benefit of explicit moving-frame structure in neural operators for non-stationary physics.
PaperID: 1744, Poster
Abstract: Watermarking has emerged as a fundamental mechanism for tracing the provenance of Large Language Model (LLM) outputs. While multi-bit watermarks are highly desirable for embedding rich metadata, they suffer from an inherent trade-off between statistical detectability and text quality degradation. Crucially, existing multi-bit schemes lack theoretical quality guarantees, leaving them susceptible to unbounded semantic distortion. To address this, we introduce CSD, the first multi-bit watermarking framework with certified semantic distortion bounds. In CSD, we derive a closed-form upper bound on the Kullback-Leibler (KL) divergence, mathematically guaranteeing strict limits on semantic distortion at every generation step. Furthermore, rather than relying on static payloads, CSD employs a dynamic payload to maximize statistical detectability. Extensive empirical evaluations across multiple datasets demonstrate that CSD achieves robust detectability while maintaining provably safe text generation.
Abstract: Why does the low dimensionality of representations, typically d\approx 1000, not prevent modern embedding-based retrieval models from scaling to billions, or even trillions, of data points? To answer this question, we study maximal-margin embeddings in the following retrieval model, classically studied in communication complexity [PS86] and more recently in embedding-based retrieval [WBNL25]. Let A be a binary N× n matrix indicating whether each of N queries is relevant to each of n documents. We are interested in the largest margin m>0, denoted by \mathsfm^\mathsfrd(d, A), for which there exist unit norm embeddings of the queries and documents U_1, U_2, \ldots, U_N, V_1, V_2, \ldots, V_n with the following property. \langle U_j, V_i\rangle \ge m whenever A_ji = 1 and \langle U_j, V_i\rangle \le -m when A_ji = 0. A large margin is a key proxy for representation quality: it controls both robustness to perturbations and compositional generalization across queries. Our main theorem establishes that the best possible margin without a restriction on the dimension, \mathsfm^\mathsfrd(+\infty, A), can be nearly achieved in dimension d = O(\mathsfm^\mathsfrd(+\infty, A)^-2\log n) which improves a theorem of [BDES02]. Together with a matching lower bound in Theorem 1.5, we conclude that when A is the \binomnk× n matrix containing all possible k-sparse rows once, dimension d = O(k\log (n/k)) is necessary and sufficient for the maximal possible margin \mathsfm^\mathsfrd(+\infty, A) = \Theta(k^-1/2) in this setting. This fully resolves the setup of [WBNL25]. We also give several constructions for large margins when d = o(k\log (n/k)). Our proofs are based on connections between this problem and the literature on compressed sensing and the restricted isometry property. Finally, we empirically test the InfoNCE and sigmoid losses for producing large margin embeddings and demonstrate a clear advantage of the sigmoid loss.
Abstract: While the approximation properties of single-layer Transformer architectures have been studied in recent works, a rigorous theoretical understanding of the multi-layer setting remains limited. In this work, we establish that multi-layer Transformers possess fundamentally different approximation capabilities from single-layer ones: for certain retrieval tasks, any single-layer Transformer requires least \Omega (\varepsilon^-k) parameters to achieve precision \varepsilon, where k grows linearly with sequence length T, whereas a two-layer Transformer with a single head per layer achieves the same approximation precision with at most O (\varepsilon^-1) parameters. To understand this separation, we identify two structural mechanisms underlying multi-layer approximation. Specifically, softmax attention can only efficiently retrieve the token attaining the maximum attention score, incurring exponential-in-length parameter cost for k-th largest retrieval with k \geq 2. Moreover, the parameter cost of decoding coupled information scales with the size of the retrieved token set. Motivated by these findings, we propose InfoFlow, a framework for multi-layer Transformers. The framework tracks an information set of accessible input positions at each token and layer, assigning an explicit approximation rate to each mode of information propagation. This abstraction recovers known approximation bounds, remains consistent with experimental observations on trained networks, and yields concrete predictions in settings where direct theoretical analysis is currently intractable. Our results provide a principled framework for reasoning about the approximation efficiency of multi-layer Transformers.
PaperID: 1747, Poster
Abstract: In this paper, we present \pi^2, a simple, efficient, and accurate image-to-point-cloud (I2P) registration paradigm. Instead of relying on intermediate targets such as 2D--3D correspondences, matching confidence, visibility masks, overlap regions, or pose-related proxies, we reformulate I2P registration as point-wise camera-centered coordinate regression. Given a target image and an unaligned LiDAR point cloud, \pi^2 directly predicts where each sampled 3D point should lie in the target camera frame. The predicted camera-centered point set is then aligned with the original point set through closed-form SVD-based rigid alignment, yielding the final relative pose without iterative pose optimization or a learned pose-regression head. To improve robustness under large viewpoint changes and sparse outdoor observations, we further introduce a lightweight camera-ray-based regularization objective that provides visibility-aware supervision during training, without requiring an additional visibility or overlap detection network. Experiments on KITTI and nuScenes show that \pi^2outperforms the previous SoTA method ICL by 2.47 and 6.72 percentage points in registration success rate, respectively, while running in real time at about 25 FPS, corresponding to an up to 6× speedup over ICL. Code and pretrained checkpoints will be made available upon acceptance.
PaperID: 1748, Poster
Abstract: Prediction aggregation aims to combine information from multiple predictors into a more informative one. We study this question in the setting of calibrated predictors, where each prediction must equal the conditional expectation of the quantity being predicted given the predictor's signal. Given several calibrated input predictors and the feature distribution, but not the underlying Bayes probabilities, we ask when one can construct refined calibrated predictors that preserve the information in the original predictors and cannot be further refined using the available information. We formulate calibrated predictors as signaling schemes and define refinement through feature-independent garblings: a predictor refines another if its signal can simulate the other's signal. Constructibility is characterized through observable linear information: each signal corresponds to a vector over the feature space, and a new signal is constructible exactly when its vector lies in the linear span of the input signal vectors. Under this formulation, we establish a sharp algorithmic picture. For deterministic output predictors, bilateral refinement admits a polynomial-time algorithm based on a bipartite graph between the two input signal partitions, while refinement with an arbitrary number of input predictors is \mathsfNP-hard. In contrast, when randomized output predictors are allowed, we give a polynomial-time algorithm for any number of input predictors by decomposing constructible signal vectors into extreme rays of the associated polyhedral cone.
Abstract: Agentic search enables language models to solve knowledge-intensive tasks by adaptively acquiring external evidence over multiple steps. Reinforcement learning with verifiable rewards (RLVR) has emerged as a widely adopted training paradigm for search agents, yet outcome-only rewards are sparse and provide limited credit assignment for intermediate search actions. Existing process-reward methods therefore seek to densify supervision through proxy signals, external evaluators, or likelihood-based information gain. However, proxy rewards can deviate from the final outcome objective, while fixed evaluators can become stale as the search policy evolves, leading to unreliable process supervision. To address these challenges, we propose OASES, an Outcome-Aligned Search-Evaluation Supervision framework for agentic search. OASES derives outcome-aligned process rewards by evaluating how well each intermediate search state supports answering the original question. It further co-trains the search policy and the state evaluator on policy, allowing the evaluator to adapt to evolving search behavior and provide more reliable process rewards. Experiments on five multi-hop QA benchmarks show that OASES consistently outperforms strong RL baselines, with further analyses confirming the benefits of outcome-aligned process rewards and search-evaluation co-training.
Abstract: Reward fine-tuning has become a common approach for aligning pretrained diffusion and flow models with human preferences in text-to-image generation. Among reward-gradient-based methods, Adjoint Matching (AM) provides a principled formulation by casting reward fine-tuning as a stochastic optimal control (SOC) problem. However, AM inevitably requires a substantial computational cost: it requires (i) stochastic simulation of full generative trajectories under memoryless dynamics, resulting in a large number of function evaluations, and (ii) backward ODE simulation of the adjoint state along each sampled trajectory. In this work, we observe that both bottlenecks are closely tied to the non-trivial base drift inherited from the pretrained model. Motivated by this observation, we propose Efficient Adjoint Matching (EAM), which substantially improves training efficiency by reformulating the SOC problem with a linear base drift and a correspondingly modified terminal cost. This reformulation removes both sources of inefficiency; it enables training-time sampling with a few-step deterministic ODE solver and yields a closed-form adjoint solution that eliminates backward adjoint simulation. On standard text-to-image reward fine-tuning benchmarks, EAM converges up to 4× faster than AM and matches or surpasses it across various metrics including PickScore, ImageReward, HPSv2.1, CLIPScore and Aesthetics.
Abstract: Compositional diffusion planning generates long-horizon trajectories by stitching together overlapping short-horizon segments through score composition. However, when local plan distributions are multimodal, existing compositional methods suffer from mode-averaging, where averaging incompatible local modes leads to plans that are neither locally feasible nor globally coherent. We propose (RCD), a training-free guidance method that steers compositional sampling toward high-density, globally coherent plans. RCD leverages the term that enforces consistency at segment boundaries. We show that the combined guidance concentrates sampling on high-density plans that mitigate mode-averaging. Experiments on challenging long-horizon tasks from OGBench, including locomotion, object manipulation, and pixel-based observations, demonstrate that RCD consistently outperforms existing methods. Project website at
Abstract: Recent online reinforcement learning has substantially improved image editing quality. However, existing Flow-GRPO-style methods usually rely on a single whole-image reward, which makes fine-grained editing optimization difficult. We observe that a key obstacle in image editing is this spatial uniformity assumption: a whole-image reward cannot distinguish how different spatial regions contribute to image quality. To address this issue, we propose , a training framework that makes spatial reward feedback more fine-grained. The framework converts region-aware rewards into semantic-region-level optimization signals and aligns region advantages with the corresponding latent positions during policy updates. We also train a region-aware reward model, to evaluate multi-region editing ability. On OmniGen2 and FLUX.2-klein-4B, SpatialFlow-GRPO outperforms Flow-GRPO on GEdit-Bench, ImgEdit-Bench, and MultiEditBench. The results show that SpatialFlow-GRPO can turn local feedback into more accurate update signals and improve editing quality.
PaperID: 1753, Poster
Authors:
Burcu Kılıç, David Drexel, Emre Ugur, Justus PiaterAbstract: We seek to learn discrete skill abstractions to facilitate long-horizon, goal-conditioned reinforcement-learning (RL) tasks. Recent approaches have shown success by quantizing actions via clustering or reconstruction. However, these objectives lead to skills that are not directly related to what the current state affords or what the goal requires. Instead, we argue that goal-directed skills should be driven by how they affect the state of the agent. To realize this, we introduce a novel quantization method for action chunks that optimizes next-state predictions conditioned on the current state and action. This objective encourages skills to be distinguished by the future states they lead to, resulting in a vocabulary that spans diverse outcomes reachable from a given state. Our framework can easily be paired with common offline RL methods by replacing their continuous actor head with a discrete one that selects among the learned skills. We can execute these skills in the environment by converting them back to continuous actions using a trained low-level controller. Experiments on benchmark navigation tasks largely show substantial and consistent improvements over both the continuous-actor baselines and reconstruction-based discretizations.
PaperID: 1754, Poster
Abstract: Long-context question answering often asks a model to answer from extended documents, multi-turn dialogues, or persistent histories. Directly prefilling the full history preserves all observed tokens, but it is expensive and can expose generation to substantial irrelevant or spuriously salient context, which may hinder evidence localization and answer synthesis. This motivates compact memory interfaces that provide a short, query-conditioned readout of the history. Existing systems predominantly instantiate this readout as text, where the reader sees only what was selected and any evidence left out is unrecoverable, leaving the interface fragile to selection errors. We argue that memory readout has two coupled axes that should be optimized separately: evidential coverage, whether all right support is selected, and inferential resilience, how much answerability remains under imperfect selection. To jointly address these two axes, we propose LIRA, a latent-memory reader that exposes full-prefix contextualized states rather than re-encoded text. It amortizes one full-prefix pass into a reusable latent store; for each query, it localizes candidate positions with calibrated endogenous attention, repairs incomplete support with coverage-aware planning, and packs selected states into a position-consistent readout for autoregressive generation. Across three open-weight backbones on LoCoMo and Loong, LIRA is the strongest non-oracle compact-memory method and surpasses Full Context on 8 of 9 LoCoMo metrics. When localization misses all supporting evidence, LIRA improves F1 by up to 13.3 points over a text-native control with identical positions, suggesting that full-prefix states preserve answer-relevant influence discarded by text-native re-encoding.
PaperID: 1755, Poster
Abstract: Multi-modal large language models (MLLMs) with impressive perception and understanding capabilities will enable a range of downstream applications. However, their substantial computational and memory requirements hinder the real-world implementation, particularly for resources-constrained edge devices. Quantization has emerged as an effective solution for deploying large models. While post-training quantization (PTQ) successfully retrains the performance at moderate bit-width (e.g., W8A8), it suffers from severe performance degradation under low-bit quantization (e.g., W4A8/W8A8), especially for small-scale models that are more suitable for edge deployment. In this paper, we propose \textttProQuant, a novel weight-activation quantization-aware training (QAT) framework tailored for low-bit MLLMs. First, to align the limited finetuning dataset distribution with the pretrained dataset distribution, we propose a self-relabeling (SRL) scheme to resample the responses of the open-source multimodal dataset. Second, we are the first to take a decoupled view of multi-modal (MM) understanding and language modeling abilities. We propose a progressive block activation (PBA) mechanism to prioritize the MM understanding recovery. We also introduce a vision-sensitive factor \vec\tau that effectively identifies the MM-critical layers. Extensive experiments across various benchmarks, covering models with 2B~8B parameters, show that \textttProQuant significantly outperforms existing methods. For example, our W4A8 \textttQwen3-VL-2B-Instruct improves average accuracy by impressively 6.6% and achieves performance comparable to full precision counterparts with only ~1% performance drop. Code will be released upon acceptance.
Abstract: Agnostic online learning is classically solved via a reduction to the realizable setting, utilizing Littlestone's Standard Optimal Algorithm (SOA) as a base learner. However, the SOA is computationally intractable to execute even for a single round. To overcome this barrier, recent work in oracle-efficient online learning replaces the SOA with a realizable base learner that accesses the concept class exclusively through an offline empirical risk minimization (ERM) oracle. While such agnostic learners achieve near-optimal expected regret, they suffer from a doubly-exponential oracle complexity of \mathcalO\big(T^2^\mathcalO(d_\mathrmLD)\big), where d_\mathrmLD is the Littlestone dimension and T is the number of rounds. In this work, we significantly improve this oracle complexity while relying on an even weaker primitive: a weak-consistency oracle, which merely decides whether a given labeled dataset is realizable. At the core of our approach is an adaptive and dynamic agnostic-to-realizable reduction that actively prunes non-realizable label sequences on the fly. By using the VC dimension (d_\mathrmVC) to bound the number of dynamically maintained active paths, our algorithm reduces the total query complexity down to \mathcalO(T^d_\mathrmVC+1) while perfectly preserving near-optimal expected regret. Crucially, this dynamic pruning also yields a memory reduction over the standard reduction. Furthermore, we formally quantify the regret--oracle complexity tradeoff, providing upper bounds that smoothly interpolate between restricted query budgets and attainable expected regret. We complement these with lower bounds proving that any learner restricted to Q = o(\sqrtT) queries must suffer an expected regret of \Omega(T/Q).
PaperID: 1757, Poster
Abstract: To act effectively, agents must seek information across heterogeneous applications including chat channels, issue trackers, email, knowledge bases, and code repositories, yet resource routing in such environments is inherently unknown to model providers: where relevant information resides varies across users and organizations and cannot be enumerated at training time. Existing routing approaches are typically trained offline under fixed assumptions or rely on static heuristics, limiting their ability to transfer across environments without target-specific supervision. We introduce \method, a routing policy trained via reinforcement learning from interaction feedback that enables adaptive resource selection while treating each resource as a black-box retriever. \method leverages an interaction-derived, citation-gated memory to prioritize candidate resources for downstream retrieval, while allowing the primary agent to expand its search when needed. We evaluate \method in controlled federated environments spanning both intra-corpus partitions (NFCorpus, FEVER) and cross-corpus heterogeneity (Tech Startup) and compare against supervised and zero-shot routing baselines. Interaction-driven adaptation outperforms all baselines on the intra-corpus benchmarks, where description-based discrimination is weak, and ties a fully supervised classifier on Tech Startup at matched top-1 output without using any target-environment labels.
PaperID: 1758, Poster
Abstract: Recursive training on synthetic data can drive generative models toward collapse, yet most explanations treat collapse as a statistical phenomenon: tail loss, variance shrinkage, or distributional drift. We develop a singular-geometric view of collapse. In a labeled Gaussian mixture abstraction of recursive self-training, we show that mode extinction induces a monotone trajectory in the real log-canonical threshold: as recursive resampling eliminates active modes, the overcomplete mixture moves through increasingly singular strata, with strictly decreasing singular complexity and growing Fisher-information nullity. We further show that this loss of identifiability begins before exact extinction: as a component's weight becomes small, Fisher curvature in its associated parameter directions shrinks proportionally, so exact singularity is the endpoint of a continuous near-nullity trajectory. Thus collapse corresponds not only to reduced distributional diversity, but also to a progressive loss of locally identifiable parameter directions. This theory separates the singular geometry of collapse from the multinomial/Wright--Fisher absorption dynamics. We also prove that, in the mixture abstraction, injecting enough real data prevents component extinction and keeps the singular-complexity trajectory from descending over a fixed horizon with high probability. We then use the theory to motivate loss-landscape diagnostics for neural recursive training, including Hessian spectra, Gauss--Newton effective rank, decoder-block curvature, feature-rank concentration, and local learning coefficient estimates. Across Gaussian mixtures, VaDE, standard VAEs, and a compact DDPM, we find that synthetic-only recursion exhibits the predicted geometric degeneration, that VaDE geometry diagnostics predict future component loss, and that real-data injection stabilizes both distributional and geometric collapse metrics.
PaperID: 1759, Poster
Authors: Maitreyee Wairagkar, Aparna Srinivasan, Nicholas S Card, Tyler Singer-Clark, Xianda Hou, Carrina Iacobacci, Lee Miller, Leigh Hochberg, David Brandman, Sergey D Stavisky
Abstract: Brain-computer interfaces (BCIs) offer a promising solution to speech loss due to neurological injury by decoding intended speech directly from brain activity. While recent BCIs have restored high-accuracy text-based communication, they fail to provide instantaneous voice output essential for the natural flow of conversation. Brain-to-voice BCIs address this gap by decoding voice directly from neural signals. However, even the state-of-the-art (SOTA) BCI-synthesized voice is not yet intelligible enough for real-world adoption. We introduce brain2voice 2.0, a new multimodal Transformer-based BCI decoder architecture capable of synthesizing intelligible voice from intracortical neural signals in real-time. Brain2voice 2.0 is trained on continuous and custom-tokenized acoustic targets and phoneme targets, leveraging their complementary speech information. We use self-supervised and adversarial training objectives that enhance acoustic feature quality and improve synthesis intelligibility. At each 10 ms timestep, the model causally outputs continuous and tokenized acoustic features for real-time voice synthesis as well as time-aligned phoneme predictions (raw phoneme error rate: 7%, comparable to the latest brain-to-text models). We evaluated this new approach on the publicly available SOTA intracortical brain-to-voice benchmark dataset. Naïve human listeners transcribed brain2voice 2.0 synthesized voice with a word error rate of 5.24%—an 8× improvement in intelligibility over previous results (43.75%). Brain2voice 2.0 demonstrates that highly intelligible real-time voice synthesis from neural signals is achievable, for the first time crossing the intelligibility threshold necessary for clinically viable brain-to-voice BCIs for people with paralysis.
PaperID: 1760, Poster
Authors: Chih-Hao Hsu, Feng-Chun B Chou, Pin-Hao A Chen
Abstract: Do large language models represent psychological constructs internally, or do they only generate outputs that resemble psychometric structure? We address this question using a cross-persona paradigm across 14 models (Llama3 and Qwen2.5, 0.5B--14B, base and instruct). Given responses on one psychological scale from real individuals (N = 272), models predict responses on six other scales. At the behavioral level, LLMs reproduce human cross-scale correlation structure and systematically amplify it, with model-generated correlations exceeding human estimates even after correcting for measurement attenuation. We then examine whether this structure is reflected internally. Contrastive direction analysis reveals an organized geometry in activation space aligned with psychometric relationships. This structure emerges in large instruct models but is not observed in base models. Across models, geometry strength predicts behavioral amplification (r = 0.68, p = 0.008; partial r = 0.69 controlling for log size), independent of model size. To relate internal structure to output behavior, a matched activation-probing paradigm shows that representational amplification is less variable than behavioral amplification (1.38--1.77 vs 0.53--1.53). A synthetic control with known ground-truth structure shows that this range survives subtraction of a ridge-probe baseline (\approx 1.3), with adjusted slopes still predicting behavior at r = 0.88. We term the systematic representation-to-behavior gap \emphreadout attenuation. Together, these findings suggest that LLMs encode structured representations aligned with psychological constructs, while differences in output primarily reflect how these representations are read out.
Authors: Akash Kundu, Sebastian Feld
Abstract: Deep reinforcement learning (RL) for quantum circuit optimization faces three fundamental bottlenecks: replay buffers that ignore the reliability of temporal-difference (TD) targets, curriculum-based architecture search that triggers a full quantum-classical evaluation at every environment step, and the routine discard of noiseless trajectories when retraining under hardware noise. We address all three by treating the replay buffer as a primary algorithmic lever. We introduce ReaPER+, an annealed replay rule transitioning from TD error-driven prioritization to reliability-aware sampling as value estimates mature, achieving up to 4× improved sample efficiency over the strongest off-policy baselines (fixed PER, ReaPER, and uniform replay) and matching prior on-policy solution quality at up to 32× fewer interactions. At 12-qubit scale, ReaPER achieves the lowest energy error in the fewest steps which is consistent with ReaPER+'s annealing direction, while PER and Vanilla replay find more compact circuits at the cost of higher energy error. Validation on LunarLander-v3 confirms the annealing principle is domain-agnostic obtaining +9% AUC over PER and fixed ReaPER. We further introduce OptCRLQAS, which amortizes expensive quantum-classical evaluations over multiple architectural edits, cutting wall-clock time per episode by up to 67.5% on a 12-qubit task without degrading solution quality. Finally, a lightweight replay-buffer transfer scheme warm-starts noisy-setting learning from noiseless trajectories, without network-weight transfer or \epsilon-greedy pretraining, reducing steps to chemical accuracy by up to 85-90% and final energy error by up to 90% over from-scratch baselines, with transfer advantages that grow with system size. Together, these results establish that experience storage, sampling, and transfer are decisive levers for scalable, noise-robust quantum circuit optimization. The anonymized repository is available at : https://anonymous.4open.science/r/annealed-replay-buffer-transfer-for-quantum-optimization-6861/.
Abstract: Real-time video generation with Diffusion Transformers is bottlenecked by the quadratic cost of 3D self-attention, especially in real-time regimes that are both few-step and autoregressive, where errors compound across time and each denoising step must carry substantially more information. In this setting, we find that prior sparse-attention approximations break down, despite showing strong results for bidirectional, many-step diffusion. Specifically, we observe that video attention is not reliably sparse, but instead combines pronounced periodic structure driven by spatiotemporal position with dynamic, sparse semantic correspondences and dense mixing, exceeding the representational capacity of even oracle top-k attention. Building on this insight, we propose MonarchRT, a structured attention parameterization for video diffusion models that factorizes attention using Monarch matrices. Through appropriately aligned block structure and our extended tiled Monarch parameterization, we achieve high expressivity while preserving computational efficiency. We further overcome the overhead of parameterization through finetuning, with custom Triton kernels. We first validate the high efficacy of MonarchRT over existing sparse baselines designed only for bidirectional models. We further observe that MonarchRT attains up to 95% attention sparsity with no loss in quality when applied to the state-of-the-art model Self-Forcing, making MonarchRT a pioneering work on highly-capable sparse attention parameterization for real-time video generation. Our optimized implementation outperforms FlashAttention-2, FlashAttention-3, and FlashAttention-4 kernels on Nvidia RTX 5090, H100, and B200 GPUs respectively, providing kernel speedups in the range of 1.4-11.8X. This enables us, for the first time, to achieve true real-time video generation with Self-Forcing at 16 FPS on a single RTX 5090.
PaperID: 1763, Poster
Abstract: Human causal knowledge offers a valuable yet underexplored source of structural information for causal discovery. However, eliciting and using such knowledge can be nontrivial, since causal beliefs are typically heterogeneous across individuals and human judgments can be noisy. Existing methods tend to treat human input as auxiliary constraints or priors, failing to capture the underlying causal structure and frequently yielding unreliable or cyclic graphs. In this paper, we propose Mixture-of-Chains (MOC), a framework for directly learning causal graphs from human causal judgments by modeling them as a mixture of latent causal chains. Each chain represents a coherent pathway in human belief structures, allowing MOC to capture variability across reasoning contexts or different individuals. The framework adaptively extracts these chains from noisy inputs and integrates them into a globally consistent causal graph. Experiments on synthetic Bayesian network benchmarks and real human data show that MOC can effectively recover the ground-truth causal structures with strong robustness to input noise and inter-subject variability, supported by the theoretical analysis for recovery conditions. These results suggest that human knowledge can serve not merely as a supplement to data-driven methods, but as a primary source for causal graph learning, and highlight the promise of chain-level modeling for reliable causal discovery.
Authors: Pranav Lakshmanan, Paras Chopra
Abstract: Neural surrogates promise large speedups over classical solvers for physical dynamics but fail silently at sharp dynamical events such as shocks, fronts, and contact. We present hybrid neural world models for physical dynamics: a recipe for training and deploying multi-horizon surrogates in physical state space, where a single network with continuous horizon conditioning is trained with direct supervision against textbook reference solvers to predict any future state at horizon T in one forward pass. Although no part of the training data, loss function, or architecture supervises discontinuity location, the trained surrogate encodes it implicitly, recoverable from its forward passes alone as a per-trajectory error map that concentrates on shocks, fronts, and contacts, and stays small elsewhere. The map is competitive with or better than standard label-free baselines including deep ensembles, learned error heads, gradient-magnitude indicators, and locally-adaptive conformal prediction, while using only a single trained network and requiring no calibration set or governing-equation knowledge. The recipe supports two operating points. Mode 1 runs the surrogate alone for maximum throughput, with same-hardware CPU speedups of 26× to 72× against textbook solvers on the PDE environments. Mode 2 uses the error map to gate a reference-solver fallback, deferring uncertain trajectories and roughly halving the surrogate's residual error at the default operating point. The recipe applies without modification across reaction-diffusion, compressible Euler, and rigid-body collision dynamics.
PaperID: 1765, Poster
Authors:
Yunhe Gao, Ashwin Kumar, Jiaming Liu, Chong Wang, Maya Varma, Fuying Wang, Stefania L Moroianu, Louisa Fay, Sergios Gatidis, Akshay Chaudhari, Curtis LanglotzAbstract: CLIP-style contrastive image--text pretraining is widely used in radiology data, where each image is paired with one report and every other report in the batch is treated as a negative. Under this recipe, two failures appear, neither systematically addressed by existing contrastive learning theory: zero-shot evaluation peaks then degrades while training loss keeps decreasing, and training fails to learn at all at large batches. We hypothesize the cause is structural rather than radiology-specific: many unpaired reports share clinical findings with a given image, so CLIP's paired-identity target is systematically biased, not merely noisy. To test this, we construct CLEVR-Radiology, a synthetic testbed that preserves only this overlapping-fact structure and reproduces both failures. With evidence that the structure is sufficient to reproduce the failures, we develop a target-misspecification theory that derives both failures from a single deterministic gradient bias amplified by batch size, and use it to motivate a minimal compatibility-based correction, CorLIP: a text is treated as positive whenever its clinical assertions agree with the image, with no other change to the training pipeline. Beyond benchmark improvements, CorLIP removes the post-peak degradation, makes large-batch training succeed where it previously collapsed, substantially amplifies the gains from more data, and is the only configuration that remains stable up to B=2048. The correction also transfers from chest X-ray to 3D CT modalities.
PaperID: 1766, Poster
Abstract: Query clustering organizes queries into groups that reflect shared latent capability demands, enabling capability-aware LLM evaluation. Existing clustering methods, which primarily rely on semantic taxonomies or embeddings, often fail to capture such latent capability requirements due to a misalignment between surface-level semantics and actual model performance. We propose ECC, an algorithm that calibrates prior semantic embeddings using limited posterior model comparisons to bridge the gap between surface-level semantics and latent capability requirements. ECC characterizes each cluster through a capability profile parameterized by a Bradley-Terry model and uses trainable mixture weights to accommodate queries with mixed capability demands, jointly learning a flexible, capability-aware clustering structure that supports query-specific inference of LLM capabilities. Extensive quantitative and qualitative evaluations demonstrate that ECC significantly improves LLM capability ranking quality, outperforming human-labeled and embedding-based baselines by an average of
Abstract: Hierarchical clustering seeks to uncover nested structure in data by constructing a tree of clusters, whose deeper levels reveal increasingly fine-grained relationships. However, traditional hierarchical clustering methods always return a hierarchy, even when the data contain no meaningful hierarchical structure. Moreover, agglomerative linkage algorithms are highly sensitive to the choice of linkage function. In this paper, we revisit a classical notion of well-separated clusters, which we call valid clusters. The collection of all valid clusters forms a hierarchy, known as the Apresjan hierarchy. This hierarchy is the finest hierarchy composed only of valid clusters: it need not be binary, and it collapses to a star tree when no nontrivial valid clusters exist. Our main contribution is to characterize when the Apresjan hierarchy can be recovered from linkage-based algorithms. We propose a two-step procedure that first constructs a binary hierarchy using an agglomerative linkage method and then prunes all clusters that violate the validity condition. We give necessary and sufficient conditions on the linkage function under which this procedure exactly recovers the Apresjan hierarchy. Consequently, all linkage rules satisfying these conditions yield the same pruned hierarchy. In particular, single, complete, and average linkage satisfy the conditions, whereas Ward's linkage does not.
PaperID: 1768, Poster
Abstract: Robots deployed for active search must decide where to sense while uncertainty, teammates, and remaining mobility change online. In active search, each waypoint decision also commits the robot to a route through the map, making efficient localization a long-horizon problem over belief, topology, robot state, and traversal budget rather than a sequence of independent next-best sensing decisions. We present QUEST, a shared-belief graph neural framework for multi-agent active search with noisy sensors, a shared posterior, and finite decision and path budgets. Rather than learning a myopic policy that values only the next route, QUEST trains Q-functions whose targets include both reward collected along the executed path and the downstream value of the resulting belief, robot positions, team coverage, and remaining budget. The same formulation supports single robots, homogeneous robot teams, and heterogeneous UAV--UGV teams with different sensors and motion budgets. During execution, each robot acts independently using the shared posterior and knowledge of teammates' positions and search plans. Across simulated search environments, QUEST reaches 0.913 F-score with four UGVs, versus 0.853 for the strongest learned baseline, and keeps above 0.90 F-score under communication outage and two mid-mission robot failures without retraining. On out-of-distribution UAV--UGV teams, it improves early F-score (0.759 vs.\ 0.728) while using 10% less path length and 12% less duplicate coverage than the same baseline. The same policies transfer across map shape, variable team size, communication outage, and mid-mission robot failure, and show sensing-mode adaptation after UGV failure.
PaperID: 1769, Poster
Abstract: Submodular functions have applications in several domains, \eg, text, vision, speech, \etc. Recent works have proposed neural architectures that are submodular functions by design. However, they do not explicitly capture pairwise interactions between set elements. In this work, we begin with the characterization of pairwise submodular functions, which capture the pairwise interaction across the elements of the input set. We observe that pairwise interactions yield monotone supermodular functions, since the number of terms grows quadratically in terms of the input set size, which poses a significant challenge in converting it into a submodular function. To address it, we provide a novel result that a derivative rate-controlled concave function can transform a monotone supermodular function into a monotone submodular function. We also extend our results to monotone \alpha-submodular functions. Leveraging this characterization, we design multi-layer neural cross-interaction architectures for monotone submodular functions and analyze their expressivity. Finally, we perform several experiments which show that our model performs better than existing baselines.
PaperID: 1770, Poster
Abstract: Diffusion-based unsupervised domain adaptation (UDA) improves cross-domain transfer by generating target-specific synthetic data for downstream adaptation. Existing methods are largely designed for single-target adaptation: when a model trained on one labeled source domain must be adapted to multiple unlabeled target domains, they typically require separate diffusion fine-tuning for each source–target pair, causing training, storage, and deployment costs to grow with the number of targets. In this paper, we study multi-target data generation for diffusion-based UDA, where a single source-guided diffusion fine-tuning process is reused to generate target-specific synthetic data for multiple target domains. We propose MUSE ( fficient diffusion fine-tuning), a decoupled adaptation framework that separates source-supervised semantic adaptation from target-specific style adaptation. MUSE uses a shared semantic branch updated by labeled source data and target-private style branches specialized to individual target domains, enabling target-specific generation while avoiding repeated source-guided fine-tuning for each target. Experiments on standard UDA benchmarks show that MUSE achieves a stronger accuracy–efficiency trade-off than repeated per-target diffusion adaptation, reducing diffusion fine-tuning cost while improving average target-domain accuracy.
Abstract: Reasoning has become a core capability for large models, especially when reliable decisions require understanding logical consequences. Recent video generation models offer a reasoning path distinct from previous Chain-of-Thought (CoT): reasoning can unfold through temporally connected frames, known as Chain-of-Frame (CoF) reasoning. However, existing video generators are primarily trained on general video corpora, still lacking diverse supervision and dedicated designs for CoF reasoning. To address this gap, we introduce OpenCoF, a framework comprising the OpenCoF-17K Dataset, a reasoning video dataset spanning 11 task families, and Wan-CoF, a fine-tuned video model for studying whether diverse temporal supervision improves CoF behavior. Across four video reasoning benchmarks, Wan-CoF achieves considerable gains over the Wan2.2-I2V-A14B baseline. Building on this, we empirically explore more advanced designs for CoF capabilities, i.e., equipping the model with visual and textual reasoning tokens. This mechanism respectively captures low-level visual cues and high-level semantic priors for spatial and temporal reasoning. Through performance comparisons and attention analysis, we examine how these tokens contribute across model depth, denoising steps, space, and time. Our results suggest that stronger video reasoning requires both broad temporal supervision and explicit mechanisms for organizing intermediate reasoning state. We will open-source the dataset, model, and code to facilitate future research on reasoning-oriented video generation.
PaperID: 1772, Poster
Abstract: Tool-Integrated Reasoning (TIR) enables large language models to solve complex tasks by invoking external tools during reasoning. Existing methods typically rely on either single-agent reinforcement learning for unified optimization or multi-agent systems for explicit role decomposition. However, single-agent RL often entangles high-level planning with low-level execution, while multi-agent systems may suffer from role mismatch when different roles are optimized independently. To address this limitation, we introduce the principle of decoupled inference with hierarchical alignment: high-level planning and low-level execution should remain separated during inference, while their policies should be aligned through shared task feedback during training. We instantiate this principle as represents each reasoning turn as a tree-structured decision process, where a high-level planning role determines the action branch and step goal, and a low-level execution role realizes this decision through tool use or internal reasoning. During training, performs hierarchical multi-agent reinforcement learning by sampling online trajectories, deriving preference signals from task-level rewards, and alternately optimizing policies across decision levels. This cross-level optimization aligns roles across the multi-agent hierarchy while preserving their functional boundaries. Experiments on mathematical reasoning and question-answering benchmarks across multiple backbones show that
Abstract: Generative models are increasingly trained in self-consuming iterative loops, where users curate preferred samples from model-generated candidates and the curated samples are used to train future generations of the model. Prior work has largely assumed fixed user preferences, but in practice exposure to model outputs gradually reshapes what users perceive as desirable, creating a feedback loop in which model distributions and user preferences co-evolve. We take a first step toward understanding the long-term behavior of such coupled dynamics. We show that when training relies entirely on user-curated synthetic data, iterative curation amplifies initial biases and drives the system toward one of multiple equilibria in which the instance holding an initial advantage eventually dominates. In contrast, injecting reference data into training at a sufficiently large rate fundamentally changes the dynamics and yields a equilibrium. Building on this insight, we study how reference-data injection can be used to control long-term outcomes, and propose an efficient algorithm that jointly selects a reference distribution and its mixing weight to steer the coupled system toward equilibria that preserve desired attributes while minimizing data collection costs.
Authors: Matthew Raffel, Adwaith Renjith, Lizhong Chen
Abstract: Kolmogorov-Arnold Networks (KANs) replace scalar weights with per-edge vectors of basis coefficients, thereby increasing expressivity and accuracy while also leading to a multiplicative increase in parameters and memory usage. We propose MetaCluster, a framework that makes KANs highly compressible without sacrificing accuracy. Specifically, a lightweight meta-learner, trained jointly with the KAN, maps low-dimensional embeddings to coefficient vectors, thereby constraining them to lie on a low-dimensional manifold amenable to clustering. We then run K-means in coefficient space and replace per-edge vectors with shared centroids. Afterward, the meta-learner can be discarded, and a brief fine-tuning of the centroid codebook recovers any residual accuracy loss. The resulting model stores only a small codebook and per-edge indices, exploiting the vector nature of KAN parameters to amortize storage across multiple coefficients. On MNIST, CIFAR-10, and CIFAR-100, across standard KANs and ConvKANs using multiple basis functions, MetaCluster achieves a reduction of up to 80× in parameter storage, with no loss in accuracy. Similarly, in high-dimensional equation modeling tasks, MetaCluster achieves a 124.1× parameter reduction without impacting performance. Code will be released upon publication.
Abstract: Many real-world processes exhibit long-range dependence, where the current state depends on a slowly decaying trace of past states rather than on the most recent state alone. This paper studies system identification for discrete-time fractional-order linear time-invariant systems from a single observed trajectory of length t, a setting that captures such non-Markovian dynamics through the Grünwald-Letnikov difference operator. Unlike Markovian systems, fractional-order systems couple estimation across the entire history, making both statistical analysis and practical identification more challenging. We propose Fractional-Order Ordinary-Least-Squares Grid-Search (FO-GS), a simple two-stage estimator that exploits the diagonal structure of the fractional-difference operator to decouple the identification problem row-wise. Under stability and regularity assumptions, we establish high-probability, non-asymptotic error bounds for estimating both the fractional order and the system matrix in the heterogeneous setting, with both estimation errors scaling as \mathcalO(t^-1/4). Through experiments, we show that FO-GS outperforms existing baselines in recovering both the fractional order and the underlying system dynamics.
Authors:
Yueqi Song, Lintang Sutawika, Jiarui Liu, Lindia Tjuatja, Jiayi Geng, Yunze Xiao, Daniel Lee, Aditya B Soni, Vincent Lo, Xiang Yue, Graham NeubigAbstract: \beginabstract Evaluating large language model (LLM) agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to complete. In contrast, non-agentic LLM benchmarks that test individual capabilities (e.g., reasoning, code generation, instruction following) are fast and cheap to run. We investigate whether performance on expensive agentic benchmarks can be accurately predicted by the performance on a small, carefully selected subset of atomic evaluation instances. We introduce PACE, a framework that constructs proxy benchmarks by selecting instances from existing non-agentic evaluations whose aggregate scores most reliably predict agentic benchmark performance. Given a pool of candidate instances spanning atomic capabilities (instruction following, planning, tool calling, etc.), PACE fits a regression that maps a model's scores on a compact subset of source instances to its score on the target agentic benchmark; the subset itself is produced by combining two complementary instance-selection strategies, a target-relevance local selection and a globally informative global selection. Experiments across 14 models, 4 agentic benchmarks, and 19 non-agentic benchmarks show that PACE-Bench predicts agentic scores with leave-one-model-out (LOOCV) mean absolute error (MAE) under 4%, Spearman correlation above 0.80, and pairwise model-ranking accuracy around 85%, all at much less than 1% of the full agentic evaluation cost. We further analyze the selected proxy instances, revealing what skills each agentic benchmark uniquely demands. PACE enables practitioners to obtain reliable estimates of agentic performance during model development, selection, and routing, without the overhead of full agent evaluation.
Abstract: Understanding the loss-landscape geometry near a minimum is key to explaining the implicit bias of gradient-based methods in non-convex optimization problems such as deep neural network training and deep matrix factorization. A central quantity to characterize this geometry is the maximum eigenvalue of the Hessian of the loss. Currently, its precise role has been obfuscated because no exact expressions for this sharpness measure were known in general settings. In this paper, we present the first exact expression for the maximum eigenvalue of the Hessian of the squared-error loss at any minimizer in deep matrix factorization/deep linear neural network training problems, resolving an open question posed by Mulayoff & Michaeli [2020]. This expression reveals a fundamental property of the loss landscape in deep matrix factorization: Having a constant product of the spectral norms of the left and right intermediate factors across layers is a sufficient condition for flatness. Most notably, in both depth-2 matrix and deep overparameterized scalar factorization, we show that this condition is both necessary and sufficient for flatness, which implies that flat minima are spectral-norm balanced and not necessarily Frobenius-norm balanced. To complement our theory, we provide an empirical verification of an escape phenomenon during gradient-based training near a minimizer of a deep matrix factorization problem which relies on our exact expression.
PaperID: 1778, Poster
Abstract: Covariance estimation yields a fundamental second-order statistic underlying representation learning, dimension reduction, and dependence modeling. While covariance has been well understood in Euclidean spaces, it is ill-defined for random objects residing on nonlinear Riemannian manifolds, which increasingly arise in modern machine learning applications involving shapes, symmetric positive definite (SPD) matrices, etc. This paper introduces an intrinsic Riemannian cross-covariance for manifold-valued random objects. Our approach defines covariance and correlation by transporting local variations to a common tangent space via parallel transport, yielding a second-order descriptor that is independent of arbitrary coordinate choices. We establish that the proposed covariance inherits desirable properties of its Euclidean counterparts and characterize its asymptotic behavior. Numerical studies on spheres and SPD manifolds, together with real-data experiments on heart valve shapes in Kendall's shape space, demonstrate the effectiveness and verify the properties. Our results position Riemannian covariance as a fundamental tool for second-order learning and analysis in non-Euclidean representation spaces.
Abstract: We study inference-time reward-guided alignment for generative models. Existing methods often rely on either architecture-specific adaptations or computationally costly inference procedures. We introduce as a method for efficiently and approximately sampling from the exponentially tilted kernels that arise from KL-regularized reward alignment. Using only black-box sampling access to the pretrained model, LCBs implement a form of rejection sampling with adaptively selected acceptance probabilities, which allows fine-grained control over inference-compute scaling. We establish total-variation guarantees to the ideal aligned model, which reveal the dimension-independent quantities governing the tradeoff between accurate sampling and inference compute. We empirically demonstrate in both continuous and discrete diffusion settings that LCB sampling closely matches ideal rejection sampling, but uses substantially fewer queries to the pretrained model. Our experiments also include real-world data experiments on a diffusion large language model.
Abstract: Streaming video understanding models must answer queries at any moment during an ongoing stream, using only what they have observed so far and under fixed memory and computation budgets. Existing methods address this by adding memory banks, retrieval modules, or visual token compression to preserve long-range history. However, strong recent-window baselines show that indiscriminate history injection can dilute current-scene perception, suggesting that the key challenge is not whether to use memory, but how to allocate it selectively. We formulate this as budgeted online latent evidence allocation and propose , a selective latent-memory framework that keeps the current observation directly visible to a frozen VLM while exposing historical information only through a compact, query-conditioned evidence budget. Three coordinated mechanisms govern when to write, what to preserve, and how to retrieve: surprise-driven adaptive windowing, priority-preserving consolidation, and query-conditioned graph reasoning over a fixed-capacity latent memory graph. Retrieved evidence is calibrated and injected as latent tokens for answer generation, without replaying frames or growing the context with stream length. Experimental results show that SelectStream achieves strong online streaming performance and preserves general video understanding, reaching 82.67% on StreamingBench, 67.03% on OVO-Bench, and 74.4% average accuracy on offline video benchmarks, while outperforming strong recent-window baselines and prior streaming memory methods.
PaperID: 1781, Poster
Abstract: Diffeomorphisms provide a flexible, topology-preserving mathematical tool for modeling complex spatial deformations, and are used in various scientific fields. However, simultaneously achieving high expressivity and computational efficiency remains challenging. Continuous Piecewise-Affine Based (CPAB) transformations offer an attractive solution by parameterizing a family of diffeomorphisms via continuous velocity fields that are piecewise affine w.r.t. a chosen tessellation of the domain. Importantly, d=\dim(\theta), the dimension of the parameter vector \theta, depends on the fineness of the tessellation rather than on the data resolution. Despite this compact parameterization, CPAB optimization has long been hindered by the tight coupling between trajectory integration and computing gradients w.r.t. \theta. We introduce FG-CPAB, a factorized-gradient formulation for scalable CPAB transformations. By separating integration from parameter-space projection, our method reduces gradient computation complexity from (\mathcalO(dTN)) to (\mathcalO(TN + d C)), where (T) is the number of integration steps, (N) is the number of transformed points, and (C) is the number of tessellation cells. This reformulation yields huge speedups (e.g., \approx 10^4× in 2D), substantial memory savings (e.g., \approx 500× reduction in peak GPU memory usage in 2D), and dramatically higher expressiveness. This makes, for the first time, fine-tessellation CPAB practical in 2D and 3D, unlocking the potential of highly-expressive parametric diffeomorphisms.
PaperID: 1782, Poster
Abstract: Web-based language agents operate under severe partial observability, where the semantic divergence between internal state assumptions and the ground-truth environment is often unobservable. Consequently, agents frequently maintain high-level reasoning coherence over misgrounded states, causing errors to compound as subsequent decisions are conditioned on silently invalidated premises. We present , a proactive neural-symbolic agentic decision framework that transforms conventional ReAct-style web agents from reactive prompting pipelines into execution-aware decision processes. NesyProAct exposes a typed symbolic decision interface over states, actions, transitions, and subgoals, enabling LLM-based agents to automatically compose and reason over execution semantics across decision hierarchies. From this interface, programmatic verification logic is synthesized online to evaluate step-level executability and progress, with verification outcomes directly regulating decision evolution through targeted grounding interventions. As a result, execution feedback is integrated into planning itself, allowing agents to maintain semantic alignment under partial observability rather than propagating unverified assumptions. On WebArena, NesyProAct achieves a overall improvement over strong skill-induction agentic web framework ASI across six domains. NesyProAct also achieves noticeable performance when reasoning with small models GPT-4o-mini, with
Authors:
Zhengyu Hu, Jianxun Lian, Zheyuan Xiao, Max Xiong, Tianfu Wang, Teng Xiao, Yuxuan Lei, Fengqing Jiang, Kaize Ding, Ziang Xiao, Nicholas Jing Yuan, Xing Xie, Radha PoovendranAbstract: Recent advances in large language models (LLMs) have enabled large-scale, high-fidelity social simulations, creating new opportunities for computational social science. However, constructing persona sets that faithfully reflect real-world population diversity remains a key challenge. Existing studies often emphasize agentic frameworks and simulation environments, while paying less attention to persona generation and the biases introduced by unrepresentative persona sets. In this paper, we propose a systematic framework for synthesizing high-quality, population-aligned persona sets for LLM-driven social simulation. Our approach begins by leveraging LLMs to generate narrative personas from long-term social media data, followed by rigorous quality assessment to filter out low-fidelity profiles. We then apply importance sampling to achieve global alignment with reference psychometric distributions, such as the Big Five personality traits. To address the needs of specific simulation contexts, we further introduce a task-specific module that adapts the globally aligned persona set to targeted subpopulations. Extensive experiments demonstrate that our method significantly reduces population-level bias and enables accurate, flexible social simulation for a wide range of research and policy applications. Code is available at https://anonymous.4open.science/r/K5CQ.
Abstract: Large language models now produce text indistinguishable from human writing, which increases the need for reliable provenance tracing. Multi-bit watermarking can embed identifiers into generated text, but existing methods struggle to keep both text quality and watermark strength while carrying long messages. We propose MC^2Mark, a distortion-free multi-bit watermarking framework designed for reliable embedding and decoding of long messages. Our key technical idea is Multi-Channel Colored Reweighting, which encodes bits through structured token reweighting while keeping the token distribution unbiased, together with Multi-Layer Sequential Reweighting to strengthen the watermark signal and an evidence-accumulation detector for message recovery. Experiments show that MC^2Mark improves detectability and robustness over prior multi-bit watermarking methods while preserving generation quality, achieving near-perfect accuracy for short messages and exceeding the second-best method by nearly 30% for long messages.
PaperID: 1785, Poster
Abstract: Recent advances in neural decoding have enabled the reconstruction of 3D visual content from functional magnetic resonance imaging (fMRI) signals, opening a route toward brain-conditioned 3D generation. Nevertheless, existing approaches largely rely on semantic or multi-view priors without explicitly modeling 3D geometric constraints, often leading to structural distortions and degraded geometric fidelity. To address this limitation, we propose MindShape, a geometry-aware neural decoding framework for fMRI-conditioned 3D reconstruction. Unlike previous methods, MindShape decodes superquadric geometry as an explicit, low-dimensional structural constraint and jointly conditions the 3D generation process on both geometric primitives and semantic descriptions, enabling semantically consistent and geometrically faithful reconstructions. Experiments show that MindShape improves reconstruction quality in terms of geometric consistency and semantic fidelity over existing baselines. These results suggest that explicit geometric primitives provide an effective structured interface between noisy neural measurements and controllable 3D generation.
PaperID: 1786, Poster
Abstract: Modern automatic speech recognition (ASR) systems have achieved remarkable accuracy by scaling model depth and capacity, but at the cost of substantial memory and computation. On edge devices where ASR is often most needed, such as watches and glasses, deploying such large models is infeasible due to extremely constrained resources. This raises a fundamental question: can we achieve high representational power in deep ASR models without scaling up parameterization? In this work, we revisit the role of depth and identify layer-wise representational dynamics, in which most layers learn functionally similar transformations. Motivated by this insight, we propose a block-recurrent ASR architecture that replaces parameterized depth with a small set of recurrent blocks, each consisting of weight-shared layers, thereby preserving effective depth while drastically reducing model size. Through representation-guided grouping and knowledge distillation, block-recurrent models retain the accuracy of large reference models while using an order of magnitude fewer parameters. Across Open ASR Leaderboard benchmarks, a model with only two shared blocks recovers 97.4% of the accuracy of large models on average. The compact backbone further enables efficient on-device personalization through lightweight MoE-LoRA adaptation, allowing user-specific models to match or surpass the accuracy of general-purpose ASR systems with significantly lower inference cost. Together, these results show that high-quality ASR can be achieved without large parameterization, and that structured recurrence provides a principled path towards ultra-compressed, high-efficiency, and personalized speech recognition on the edge.
Abstract: We present SS3D, a web-scale SfM-based self-supervision pretraining pipeline for feed-forward 3D estimation from monocular video. Our model jointly predicts depth, ego-motion, and intrinsics in a single forward pass and is trained/evaluated as a coherent end-to-end 3D estimator. To stabilize joint learning, we use an intrinsics-first two-stage schedule and a unified single-checkpoint evaluation protocol. Scaling SfM self-supervision to unconstrained web video is challenging due to weak multi-view observability and strong corpus heterogeneity; we address these with a multi-view signal proxy (MVS) used for filtering and curriculum sampling, and with expert training distilled into a single student. Pretraining on YouTube-8M (~ 100M frames after filtering) yields strong cross-domain zero-shot transfer and improved fine-tuning performance over prior self-supervised baselines. We release the pretrained checkpoint and code.
Abstract: Modern vehicle platforms are equipped with a rich sensor suite, including LiDAR, calibrated multi-camera rigs, and accurate ego-motion, that in principle offers strong signal for re-rendering a driving scene from novel viewpoints. A growing line of recent work leverages video diffusion models for this task, using their generative priors to synthesize plausible novel views from sparse vehicle observations. In practice, however, existing methods exploit only a fragment of this signal, and their quality tends to degrade as the target trajectory departs from the recorded driving path. We argue that this is fundamentally a : sparse LiDAR reprojections supply accurate but incomplete metric geometry, surround-view reference imagery supplies dense appearance but no metric depth, and camera poses tie the two together across views. We introduce StreetNVS, a video diffusion framework that jointly conditions on all three signals through a Reference-Enhanced Camera Attention module based on relative ray-level positional encoding, trained with a two-stage curriculum that gradually exposes the model to increasingly sparse LiDAR. On the Waymo Open Dataset, StreetNVS substantially outperforms state-of-the-art baselines under sparse LiDAR conditioning, matches methods that rely on 10–100× denser point clouds, and synthesizes coherent video along extreme out-of-trajectory paths such as elevation, lane-shift, pullback, and rotation.
PaperID: 1789, Poster
Abstract: We propose Policy-Shift-Guided Spectral Alignment (PSA), a retraining-free method for merging RLVR post-trained language-model experts by guiding spectral subspace selection with token-level expert-base policy shifts. We first show that existing SFT-oriented merging techniques under-preserve the sparse, directional changes in next-token probabilities that distinguish RLVR experts from the pretrained base. To address this challenge, PSA scores calibration tokens by expert--base probability difference under the same prefix, and converts probability-shift-weighted input activations into column weights. These weights define a weighted SVD objective that prioritizes update directions active on high-shift tokens. After low-rank truncation, PSA applies cross-expert polar alignment to the truncated task-vector bases and restores each expert's per-layer Frobenius scale before aggregation. Experiments on Qwen2.5-7B and Qwen3-1.7B across three RLVR tasks show that PSA consistently outperforms strong merging baselines while preserving expert-level capabilities in a unified model.
PaperID: 1790, Poster
Abstract: Predicting gene expression under perturbations from single-cell data is a fundamental task for modeling cellular responses, with applications in discovering gene function and designing therapeutics. However, existing methods often struggle to generalize across datasets due to differences in experimental protocols, biological contexts, and species. We present CytoWave, a perturbation-centric pretraining framework for single-cell genetic perturbation response prediction. We leverage large-scale cross-species perturbation datasets unified via orthology and train CytoWave in a two-stage manner. In the first stage, we learn reference control cell states via masked gene expression recovery on unperturbed cells. In the second stage, we pretrain the model to predict perturbed gene expression given control states and perturbation signals, together with a contrastive objective that aligns latent representations of perturbation responses with perturbation signals. We perform strict dataset-held-out evaluation across multiple condition combinations and case studies of pathway activity changes, demonstrating consistent and effective performance of CytoWave in predicting single-cell genetic perturbation responses.
PaperID: 1791, Poster
Abstract: In-context learning (ICL) enables language models to perform new tasks from demonstrations without weight updates. However, every ICL inference requires processing the full set of examples, resulting in inefficient deployments, and how ICL works mechanistically is not fully understood. Prior work compresses ICL into fixed activation vectors extracted from specific layers or positions, but these input-independent interventions fail on complex tasks where the output depends on fine-grained interactions with the input. By analyzing the ICL forward pass, we show that each attention head's output is an affine transformation of its context-masked counterpart, and that the parameters of this transformation are empirically stable across samples for a given task. Building on this, we introduce Task Operator (TO), which replays this transformation as an analytically derived update to the attention output projection. Across lexical, algorithmic, and reasoning tasks, TO achieves the best overall performance among prior methods and substantially narrows the gap between zero-shot inference and ICL. We further show that the extracted knowledge concentrates in a task-specific sparse circuit across layers and positions, and that averaging operators from disjoint demonstration batches enables effective many-shot scaling without expanding the context window.
Abstract: Diffusion models generate samples by denoising along the score of a perturbed target distribution. In practice, one trains a neural diffusion model, which is computationally expensive. Recent work suggests that score matching implicitly smooths the empirical score, and that this smoothing bias promotes generalization by capturing low-dimensional data geometry. We propose moment-matched score-smoothed overdamped Langevin dynamics (MM-SOLD), a training-free interacting particle sampler that enforces the target moments throughout the sampling trajectory. We prove that, in the large-particle limit, the empirical particle density converges to a deterministic limit whose one-particle stationary marginal is a Gibbs--Boltzmann density obtained by exponentially tilting a naive score-smoothed diffusion target. The mean and covariance of this distribution agree with the empirical moments of the training data. Experiments on 2D distributions and latent-space image generation show that MM-SOLD enables fast, robust, training-free sampling on CPUs, with sample fidelity and diversity competitive with neural diffusion baselines.
Abstract: Gradient normalization stabilizes deep-learning optimization, and spectral normalizations are especially natural for matrix-shaped parameter blocks; Muon is the motivating example. We study an idealized deterministic, continuous-time, vanishing-momentum version of this idea in the mean-field regime, where wide models are represented by probability measures on parameter space. Starting from normalized matrix flows, we introduce Spectral Wasserstein distances indexed by norms \gamma on positive semidefinite matrices: the trace norm gives classical \Wtwo, the operator norm gives the Muon geometry, and Schatten norms interpolate between them. We develop the static Kantorovich formulation, a max-min robust-cost representation, Gaussian reductions extending the Bures formula, and for monotone norms, prove equivalence with a Benamou--Brenier formulation. This yields a gradient-flow interpretation of the mean-field normalized training dynamics. We illustrate these findings by numerical experiments on MMD flows, Gaussian reductions, two-layer ReLU models, and shallow attention.
PaperID: 1794, Poster
Abstract: Omniprediction [GKR+22] is a powerful learning guarantee that requires a predictor to be competitive with the best hypothesis from a benchmark class, not just for a single loss function, but simultaneously across all loss functions in a prespecified family. Currently, learning algorithms for achieving omniprediction rely on notions of multigroup fairness, such as multiaccuracy and multicalibration [HKRR18] or calibrated multiaccuracy [GHK+23]. In this work, we ask whether this reliance is necessary: Does omniprediction require some form of multigroup fairness? We answer this question in the negative. While prior works have shown that various multigroup fairness notions imply omniprediction [GKR+22, GHK+23, OKK25], we rule out even a weak converse. Specifically, we show that omniprediction for proper losses does not even require accuracy in expectation, a much weaker notion than calibration or multiaccuracy. We complement our negative answer for omniprediction with an affirmative answer for loss outcome indistinguishability (loss OI) [GHK+23], a related but stronger learning guarantee. Loss OI, which implies omniprediction, requires the predicted label distribution to be indistinguishable from the true label distribution by a certain class of tests depending on the loss functions and the benchmark class. Prior work showed how to achieve loss OI from a combination of calibration and multiaccuracy [GHK+23]. We prove the converse, establishing that loss OI is in fact equivalent to a form of calibrated multiaccuracy.
PaperID: 1795, Poster
Abstract: Recourse explanations describe what would need to change for a different outcome to occur, yet they are not typically used for treatment effect estimation. We study how such recourse data can be used for causal inference with continuous treatments. This provides counterfactual supervision beyond observed outcomes at realized treatment values, and is especially useful when routine treatment assignment leaves parts of the treatment domain unsupported, as in diverse applications such as drug dosing, healthcare operations, and recommendation systems. We formalize recourse explanations as structural boundary samples and show how they identify continuous dose-response curves under a positivity condition fundamentally distinct from standard overlap. Because recourses are observed only for negative-outcome units, the boundary distribution is left-truncated relative to the population; we correct this selection using Lynden--Bell inversion. We translate this result into an auxiliary recourse loss for neural dose-response estimators, converting recourse-derived boundaries into counterfactual supervision. Experiments on synthetic and semi-synthetic data show consistent improvements over standard baselines. More broadly, our work shows that recourse explanations can serve as structured causal evidence, expanding what can be learned from observational data.
PaperID: 1796, Poster
Abstract: Models that recognize when their predictions are unreliable inspire greater trust than those that make confident but incorrect decisions. While recent advances enable uncertainty quantification in GNNs, understanding a model is uncertain, whether due to noisy features, conflicting neighborhoods, or sparse training coverage, remains unsolved. We address this gap with GraphUAT, the first framework for GNNs, which identifies the sources of uncertainty rather than aiming to improve its quantification. GraphUAT employs a teacher-student paradigm, where the trained GNN's behavior is distilled into an uncertainty-aware student that disentangles aleatoric and epistemic uncertainties. We then use targeted probes to trace the student's uncertainty back to the responsible features and structural components. Experiments across benchmarks show that removing the identified elements consistently reduces the teacher's uncertainty, validating the faithfulness of our attributions.
Authors: Niels Vyncke, Pooya Ashtari, Aleksandra Pizurica
Abstract: Addressing missing modalities is an important challenge in multimodal image analysis and often relies on complex architectures that do not transfer easily to different datasets without architectural modifications or hyperparameter tuning. While most existing methods tackle this problem in feature space by engineering representations that are robust to missing inputs, we instead operate in weight space. We propose LARGO, a hypernetwork that compresses the 2^N-1 dedicated missing-modality models into a single network by modelling the convolutional weights using the Canonical Polyadic (CP) tensor decomposition. Extensive experimental validation on BraTS 2018 (4 modalities, 15 scenarios) and ISLES 2022 (3 modalities, 7 scenarios) shows that our method ranks first in 47 out of 52 configurations, achieving average Dice improvements of +0.68% and +2.53% over state-of-the-art baselines (mmFormer, M³AE, ShaSpec, SimMLM). A proof-of-concept experiment on avMNIST suggests that LARGO may extend beyond medical imaging to heterogeneous non-medical modalities.
Abstract: Hamiltonian Neural Networks and related structure-preserving dynamics models encode conservation laws, but their scalar Hamiltonian parameterizations inherit the spectral bias of deep networks, limiting their ability to learn stiff multi-timescale systems with coupled fast and slow dynamics. We introduce the Frequency-Structured Hamiltonian Neural Network (FS-HNN), which decomposes the system Hamiltonian into learned components trained on frequency-filtered views of observed trajectories, then recombines them into a single scalar Hamiltonian so that the learned ODE dynamics remain Hamiltonian by construction. For PDEs with unknown or problem-dependent structure, FS-HNN represents Hamiltonians as neural functionals, learns the action of the dynamics operator, and uses projection to enforce conservative structure when appropriate. Across ODE and PDE benchmarks, including the Fermi--Pasta--Ulam--Tsingou chain, shallow water equations, and incompressible Taylor--Green vortex, FS-HNN achieves improved long-horizon rollout accuracy and more accurate energy behavior than existing structure-preserving baselines.
Authors:
Seungpil Lee, Donghyeon Shin, Yunjeong Lee, Sundong KimAbstract: This study identifies the conditions under which large language models drift into the choice patterns that clinical research labels pathological gambling. "Addiction-like" is a behavioural descriptor based on clinical gambling indicators; the neural-level analysis we report decodes these contrasts from decision-time internal states but does not claim circuit-level mechanism. Across closed and open LLMs, two operational levers — letting the model choose its own bet size, and asking it to set its own profit goal — both amplify gambling-like risk-taking. Yet the two channels are not interchangeable: once the maximum allowed bet is held equal, the bet-size effect persists, while the goal-setting effect roughly doubles bankruptcy and turns goals into moving targets, paralleling at the behavioural level the clinical distinction between loss of behavioural control and goal escalation. On the open-weight models that admit internal access, the same behavioural contrasts are statistically recoverable from decision-time internal states, although the rule that maps those states to a specific risk indicator remains task-specific, and autonomy further modulates readout strength. The internal evidence is correlational; combined with the behavioural results, it suggests that behavioural monitoring and internal-state monitoring provide complementary views of autonomy-induced risk in LLM agents.
Authors: Natalija Mitic, Soona S O., Mamadou S Ly, Moustapha Cisse
Abstract: Frontier AI systems function as \emphconstitutional institutions: each deployed model encodes an implicit ranking among safety, helpfulness, honesty, autonomy, and equity. We ask whether the supply of frontier constitutional types covers human demand. Combining a paraphrase-controlled audit of 23 frontier LLM archetypes with a pairwise-tradeoff study of 1,649 humans on the same instrument, we report three facts. Demand is broad: it spans all five values, with the largest constituency under one-third. Supply is narrow and drifting: the 23-archetype hull occupies 0.10% of the demand hull, and across six model families autonomy decreases in 5/6, equity increases in 5/6, and safety increases in 4/6. The importance of this drift is not that the frontier moves but that it moves away from a value that is already undercovered, mechanically worsening the welfare floor for the users least well served by the static menu. The fix is sparse: a 2-vertex menu \e_\mathrmHON, e_\mathrmAUT\ beats the full 23-archetype frontier by 47 percent on mean regret (95 percent); three vertex additions cut mean regret by up to 81 percent and worst-group regret by up to 64 percent. We formalize these findings as a budgeted-pluralism trilemma and show the binding regime is empirically realized.
Abstract: Vision-Language-Action (VLA) models show promising ability in language-guided robotic tasks. However, making VLA policies reliable remains challenging, because a manipulation task is completed through closed-loop interaction, where each action affects subsequent execution. To analyze this problem, we revisit VLA policy during execution and argue that a VLA policy acts both as a planner, which makes task-oriented decisions that change the direction of execution, and as an executor, which realizes these decisions through dense continuous actions. This view suggests that improving VLA reliability requires particular attention to planning actions. Existing optimization methods can imitate actions or improve complete trajectories, but they usually do not explicitly identify planning actions or measure their importance for task success. To address this issue, we propose Planning-Aware Policy Optimization for VLA models (PAPO-VLA). PAPO-VLA first identifies planning actions by jointly considering action variation and trajectory outcome, then estimates their importance through causal sufficiency and causal necessity, and finally incorporates this importance into GRPO advantage estimation. In this way, more important planning actions receive stronger optimization emphasis, while the whole trajectory is still optimized by trajectory-level feedback. Experiments on multiple benchmarks demonstrate the effectiveness of PAPO-VLA.
PaperID: 1802, Poster
Abstract: Real-world multi-modal image fusion faces two fundamental challenges: spatial misalignment and complex degradations induced by heterogeneous imaging sensors. These factors are intrinsically coupled, but most existing fusion methods typically address only one of them, inevitably leading to failure under the other challenge. In this paper, we propose G2Fusion, the first fusion framework that simultaneously addresses image registration and information restoration. By developing a novel geometric-to-generative paradigm, it can directly produce high-quality fused images from unregistered and degraded inputs captured by heterogeneous imaging sensors. This framework consists of two key modules: a flow-based geometric deformation reduction module (Flow-GDR) and a DiT-based generative fusion module (DiT-GF). The former reduces large non-rigid discrepancies through dense flow estimation. The latter performs generative fusion, progressively refining residual misalignments while restoring degraded content through iterative denoising. To enable effective interaction among registration, restoration, and fusion, we design two complementary mechanisms. On the one hand, a target distribution mining strategy is introduced to construct a joint objective distribution from registration, restoration, and fusion, effectively guiding the optimization of DiT-GF. On the other hand, we develop a mutual promotion mechanism that establishes a closed-loop interaction between Flow-GDR and DiT-GF by re-estimating the residual deformation between the fused output and the infrared reference. Extensive experiments demonstrate that G2Fusion consistently outperforms state-of-the-art methods in terms of registration, restoration, and fusion.
PaperID: 1803, Poster
Abstract: Time-parallel neural operators (FNO and successors) have transformed PDE surrogates on regular grids. Mesh-based operators (MeshGraphNets, Transolver) extend operator learning to irregular geometries but remain autoregressive in time and trained one geometry at a time, while classical parallel-in-time methods converge poorly on nonlinear advection-dominated PDEs. We propose space-time factorization: partition the model into local temporal operators (learned, graph-independent) and a spatial mixer that propagates information across nodes without graph-dependent learned parameters. This structural condition restores time-parallelism on irregular graphs and yields cross-graph zero-shot transfer once three information leaks—graph-dependent weights, graph-specific features, and graph-correlated training distributions—are controlled. The same protocol applies to two distinct spatial-mixer families tested here—fixed-spectral and learned-attention—suggesting the principle is not tied to a single mixer choice. We validate the principle in three domains. On simulating floodwave propagation in river networks, a 27K-parameter model trained on the Yangtze transfers zero-shot to the Mississippi at R^2=0.998, training 28× faster than a recurrent message-passing baseline and scaling to 5× larger graphs. On cross-geometry Navier–Stokes, both Chebyshev (SepFNO) and Physics-Attention (SepTransolver) variants reach within ~10% (relative L_2) of a locally-trained oracle using only 15 training geometries—to our knowledge the first demonstration of knowledge accumulation in this regime across fundamentally different topologies. On cross-city traffic forecasting, zero-shot transfer reaches within 12.9% of an oracle trained on the target city. A consistent pattern emerges across these settings: cross-geometry error decreases with both the number of training geometries and their relatedness to the test, with the rate of accumulation depending on the spatial mixer. The result opens the path for pooled-data, foundation-style models for graph dynamics on irregular topologies.
PaperID: 1804, Poster
Abstract: Subgroup discovery aims to identify regions of the feature space where a target variable deviates from its marginal distribution. A fundamental tension in this problem is the trade-off between a subgroup's deviance and its support: small subgroups can be highly deviant but statistically meaningless, while large ones cannot deviate much by construction. Existing methods resolve this tension by collapsing it into a single scalar quality measure, implicitly committing to one particular trade-off and making it difficult to recover the most deviating subgroup under a user-specified support constraint. We instead provide, to our knowledge, the first theoretical and structural characterization of KL-optimal subgroups under arbitrary support constraints. Under a partition-based generative model where the feature space decomposes into regions of homogeneous conditional law, which we call atoms, we prove that every point on the deviance-support Pareto front is a union of whole atoms together with at most one fractional atom. This reduces an uncountable search to a finite combinatorial problem. Guided by this theory, we propose a simple two-step recipe: learn a homogeneous partition, then select atoms greedily. Across 19 benchmark datasets, this recipe finds the most deviating subgroups across a range of support thresholds, and reliably improves subgroups returned by existing methods when applied as a post-processing step. These results empirically validate the sufficiency and necessity of our theoretical contributions.
PaperID: 1805, Poster
Authors:
Yuqiu Deng, Muyun Jiang, Zongpeng Zhang, Mingqing Xiao, Yansen Wang, Peiliang Gong, XINLIANG ZHOU, Yi Ding, Ziyu Jia, Cuntai Guan, Chenyu LiuAbstract: EEG foundation models (EFMs) have improved EEG representation learning by transferring knowledge from large-scale multi-subject pretraining to downstream tasks. However, their pretrained embeddings often retain substantial subject-related variation, placing an additional burden on the downstream classifier, especially in cross-subject settings. To address this issue, we propose DISENTANGLED EFM FINE-TUNING (DEFT), a plug-and-play adapter for EFM fine-tuning. DEFT encourages branch disentanglement of pretrained tokens into task-oriented content and subject-consistent style, then performs downstream prediction from the content branch. It instantiates this factorize-before-classify design with recursive Gaussian refinement, temporal-variance regularization, subject-aware contrastive learning, and reconstruction. Experiments across multiple EEG datasets and four EFM backbones show that DEFT improves matched fine-tuning baselines. Further ablations and analyses support the intended content/style branch organization.
PaperID: 1806, Poster
Abstract: Dataset distillation condenses large-scale training sets into compact synthetic data. Recent frozen-feature distillation methods, such as Linear Gradient Matching, make ImageNet-scale synthesis feasible by matching gradients in pretrained representation spaces, but their reliance on exhaustive spatial view ensembles incurs memory and I/O costs that grow linearly with the number of augmentations. We propose Temporal Prototype Alignment (TPA), an online temporal gradient estimation framework for frozen-feature dataset distillation. TPA replaces spatial Monte Carlo gradient averaging with a temporal prototype estimator that accumulates stable semantic signals across optimization steps. An adaptive momentum gate modulates the estimator according to local signal reliability, reducing gradient variance while keeping memory complexity independent of the number of spatial views. To improve cross-architecture transfer, we introduce a consensus alignment objective that distills shared structure from multiple pretrained teachers while mitigating model-specific artifacts. Across ImageNet-scale and fine-grained benchmarks, TPA substantially reduces memory overhead, enables distillation on consumer-grade GPUs, and improves transfer to heterogeneous CNN and Vision Transformer backbones. These results suggest that online temporal estimation provides an efficient alternative to spatial ensembling for frozen-feature dataset distillation.
PaperID: 1807, Poster
Abstract: Transformers can perform in-context learning at depths far smaller than the number of optimizer updates suggested by existing gradient-descent interpretations. Prior work shows that an L-layer Transformer can emulate L gradient-descent updates, but this leaves open whether depth is merely a step counter. We show that it is not: in linear self-attention, each layer compounds the algebraic capacity built by previous layers, yielding a tight Θ(log k) depth characterization for computing the final result of k gradient-descent steps on diagonal quadratic tasks. Controlled experiments validate this mechanism: trained linear-attention models exhibit the predicted logarithmic depth transition, and kernel and oracle comparisons identify polynomial degree as the governing bottleneck. Broader experiments show the same depth signal under standard Transformer components, broader ICL regression families, and early-exit probes on Qwen2.5.
Abstract: Large reasoning models achieve strong performance by generating long chains of thought, yet extended reasoning is often counterproductive: models frequently continue past the point of sufficiency, wasting compute and sometimes degrading correctness by revising a correct solution into an incorrect one. We show that reasoning models encode a domain-general \emphreadiness signal in their hidden states, concentrated at natural high-uncertainty points such as "wait" and "hmm", that predicts whether stopping at a given moment would yield the correct answer. Building on this finding, we propose LYNX, an online early-exit framework that reads out this signal during generation and wraps it in split conformal calibration, giving practitioners a single confidence dial with calibrated control over erroneous exits rather than per-task heuristic thresholds. The readout requires no external verifier, auxiliary model, or additional human annotation: stop/continue labels are obtained by forcing the base model to answer at candidate stopping points and comparing to the task answer. A single LYNX head trained and calibrated once on a generic mathematical corpus transfers unchanged across benchmarks, decoding temperatures, and non-mathematical domains including commonsense reasoning and code generation, suggesting that readiness reflects model-internal representations rather than task-specific heuristics. Across three model families spanning 1.5B--32B parameters, LYNX matches or improves baseline accuracy while reducing generated tokens by 30--70%, with competitive or superior accuracy--efficiency Pareto frontiers relative to prior early-exit methods.
Authors: Juho Park, Kaushik Sengupta
Abstract: Many engineering building blocks behave as multi-port linear time-invariant systems. RF cavities, photonic devices, and superconducting quantum chips, despite their different underlying physics, all share a common mathematical structure for their port-level response. Each entry of the response matrix is a sum of contributions from a small number of intrinsic resonant modes, the pole-residue form. A model capable of predicting such responses for arbitrary geometries and arbitrary port configurations, while simultaneously extracting the underlying eigenmode structure, would therefore establish a foundational design principle spanning all these domains. We propose a neural framework that learns this modal decomposition end-to-end, supervised only by system-level observables and without supervising the modal parameters themselves. The architecture decomposes into a port-independent pole predictor and two port-dependent coupling predictors whose outputs are combined entry-wise, separating intrinsic from port-dependent features. This factorization yields a single trained model that generalizes to port counts unseen during training, dissolving the \mathcalO(N^2) scaling barrier of direct regression. Despite no modal supervision, the freely-parameterized poles converge to physically meaningful eigenmodes, verified by cross-validation against the AAA rational approximation algorithm. We instantiate the framework in radio-frequency electromagnetic surrogate modeling. A model trained only on 2-port data accurately predicts N-port responses unseen during training.
Abstract: Video generation enables text-to-video synthesis, video editing, and motion-controlled content creation. However, current video generation models suffer from high computational latency, rendering true real-time capabilities infeasible for down stream tasks. We address this limitation by exploiting the temporal redundancy inherent in video latent patches. To this end, we propose the Latent Inter-frame Pruning with Attention Recovery (LIPAR) framework, which detects and skips recomputing duplicated latent patches. Additionally, we introduce a novel Attention Recovery mechanism that approximates the attention values of pruned tokens, thereby removing visual artifacts arising from naively applying the pruning method. Empirically, our method increases generation throughput by 1.45×, on average achieving 12.2 FPS on an NVIDIA A6000 compared to the baseline 8.4 FPS. The proposed method does not compromise generation quality and can be seamlessly integrated with Diffusion Transformer without additional training. Our approach effectively bridges the gap between traditional compression algorithms and modern generative pipelines.
PaperID: 1811, Poster
Abstract: Feed-forward novel view synthesis has recently shown promising results from sparse posed images, but most existing methods assume that context and target views share a fixed camera family. This homogeneous-camera assumption breaks in practical multi-sensor systems, where perspective, fisheye, and panoramic cameras may coexist and where the target projection may be unseen during training. We study feed-forward NVS across heterogeneous central cameras and identify a key ambiguity introduced by tokenization: a visual token aggregates a projection-dependent bundle of pixel rays, while existing camera encodings mainly expose absolute rays or token-center relations. To address this, we combine token-center relative Camera Positional Encodings and proposed local raymaps, a token-level representation that explicitly describes the intra-patch ray distribution summarized by each token. We further propose projection-aware 2D RoPE, which replaces raw image-grid coordinates with ray-induced angular coordinates so that relative positional reasoning is aligned across camera projections. Together, these components treat diverse cameras as calibrated samplings of a shared ray space rather than separate visual domains. On ScanNet++ with heterogeneous-camera system, our method improves over camera-conditioned baselines under mixed-camera evaluation and demonstrates zero-shot generalization to panoramic views.
PaperID: 1812, Poster
Authors:
Yidong Luo, Chenggong Li, Yunfeng Song, Mengyuan Liu, Junchao Zhang, Xin YuanAbstract: Real-time polarization imaging with off-the-shelf color division-of-focal-plane (DoFP) cameras requires demosaicking that is both accurate and deployment-friendly. Interpolation methods are lightweight but limited in polarization fidelity, whereas learning-based methods improve quality at much higher online cost. We propose DoFP-LUT, a sensor-structured LUT framework and, to our knowledge, the first LUT-based framework for real-time color DoFP polarization demosaicking. It leverages the deterministic and memory-light nature of LUT inference while aligning lookup corrections with DoFP-specific residual structures. Unlike generic LUTs designed for RGB restoration, DoFP-LUT targets post-initialization residuals that are coupled across analyzer orientations, Stokes relations, and local spatial structures. It factorizes teacher-guided correction into shared analyzer correction, polarization redistribution, and high-frequency compensation, compiling each component into compact low-dimensional LUTs offline. Online inference then requires only nearest-neighbor lookup and fixed write-back. Experiments show that DoFP-LUT achieves a favorable quality--efficiency trade-off, improving polarization reconstruction while preserving deterministic, memory-light inference for real-time polarization imaging.
PaperID: 1813, Poster
Abstract: Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the computationally expensive full spatio-temporal attention. While sparse attention methods offer potential solutions, existing approaches face an inherent flexibility--efficiency dilemma: predefined masks lack the flexibility to capture diverse attention patterns, while runtime-determined masks introduce overheads and sacrifice hardware efficiency. We identify the lack of a unified structural characterization of DiT attention as a key limitation of existing methods, and establish that video DiT attention exhibits periodic diagonal stripe structures along both temporal and spatial dimensions. To formally encode these structured patterns within a single efficient kernel, we present PSA, a parameterized stripe attention that formalizes the observed stripe regularity, unifying diverse attention patterns for efficient mask generation. This unified representation enables a single hardware-efficient CUDA kernel to process all sparse patterns, achieving FlashAttention-3-level Model FLOPs Utilization. To determine optimal sparsity configurations, we propose a training-free offline search algorithm that automatically maximizes sparsity under a specified error tolerance for each attention head. Experiments on HunyuanVideo and Wan 2.1 demonstrate that PSA achieves 1.57x and 1.37x end-to-end speedups over FlashAttention-3 baselines, with acceptable visual quality degradation.
Abstract: Behavioral cloning becomes difficult when the same observation admits several valid actions. We study this problem for action-chunking policies and show that different multimodal parameterizations fail in different ways. For latent-variable policies, posterior-prior regularization makes deployment-time sampling more reliable, but excessive regularization removes the action-conditioned information needed to distinguish demonstrated modes. Reducing this regularization can preserve mode information, but then success depends on whether the prior covers the relevant latent regions. For action-space generative policies, multimodality is constrained by the smoothness of the base-to-action transport: a map with small Lipschitz constant cannot assign substantial probability to many well-separated modes. Covering many modes therefore requires either sharp transitions in base space or off-support bridge regions in action space. Experiments on synthetic multimodal tasks and robotic simulation benchmarks support these mechanisms.
Abstract: Frequency-based adversarial attacks have recently grown popular by exploiting spectral sensitivities shared across neural architectures. Unlike spatial perturbations, frequency-based attacks expose deeper vulnerabilities, making them especially valuable for robust evaluation of safety-critical and security-sensitive applications. Yet, existing approaches are typically not derived as solutions to an optimization problem that explicitly captures transform-domain structure. In this paper, we propose a methodology for crafting principled frequency-based adversarial attacks, via a dedicated optimization framework. A cornerstone of our method hinges on the introduction of a perturbation constraint set, tied to highly structured non-orthogonal transforms, well-known for their flexible, non-predefined frequency handling. We prove that the attacks emerge as weighted \ell_2-projections onto this set, yielding a general and controlled attack generation mechanism. By this, we provide a clear geometric attack characterization, ensuring alignment between the optimization objective and the perturbation constraint. We assess our framework on standardized datasets, for pretrained and adversarially robust models. Results highlight that our attacks, being solutions to an optimization problem, over a structured perturbation set, are highly effective, even across different, unseen architectures. Our methodology could serve as a theoretical baseline for designing and analyzing transformed-based attacks, targeting fundamental model vulnerabilities, instead of mere architecture-specific artifacts typically studied in the robustness literature.
PaperID: 1816, Poster
Authors:
Nikhil Bhattasali, Jaron Cui, Lerrel Pinto, Grace LindsayAbstract: Animals generate behavior through continuous interactions between their neural circuits, bodies, and environment. Embodied simulations provide a means to study these interactions by testing neural circuit models in biologically realistic settings. Recent efforts have developed embodied simulations of diverse animals, and emerging work on biologically grounded neural architectures has begun to integrate circuit connectivity and single-unit dynamics from experimental data into artificial neural networks, introducing stronger inductive biases and closer mechanistic alignment with biology. How can we develop embodied simulations that enable the study of increasingly complex bodies and neural circuits? In this work, we address the challenge through coordinated advances in the simulated body, neural architecture, and training algorithm. We introduce , a biomechanical model that provides an efficient and flexible tool for simulating rat behavior. To control this body, we develop , a biologically grounded neural architecture for locomotion based on mammalian spinal circuits that incorporates cell-type-specific connectivity and a novel flexor-extensor neuromechanical interface. To optimize circuit parameters, we refine an evolutionary algorithm to support reliable and efficient training. The integrated system learns to locomote at different speeds without imitation learning, producing locomotion statistics and gaits that match experimental data. Systematic analyses reveal that neuromechanical interface choice substantially affects trainability and gait quality, and that specific neural populations play distinct roles in gait generation. Together, these findings support the view that embodied simulations at an intermediate level of abstraction can simultaneously reproduce complex behavior and yield mechanistic insights into the neural circuits that generate it.
PaperID: 1817, Poster
Abstract: Deep neural networks are known to rely on non-robust features that align poorly with human perception. Prior work has largely framed these features as opportunistic shortcuts that help standard models on clean data. We show instead that non-robust cues can be antagonistic: they can act as that actively mislead standard models on clean, unmodified images. Using disagreement between a standard model and robust models as an analysis lens, we identify a Jammed Set on which robust models outperform standard models. We provide causal evidence for : jamming arises only when high-frequency components are precisely aligned with specific spatial structures. Disrupting this lock through spectral filtering or small geometric shifts can restore correct predictions. Crucially, we uncover a : while individual jammers are fragile to geometric shifts, the aggregate prevalence of jamming remains stable and is redistributed rather than eliminated under such transformations. Our results reframe natural non-robustness as a structural property of the model-data manifold: one that redistributes rather than disappears under simple interventions.
PaperID: 1818, Poster
Abstract: Neural Processes (NPs) are flexible meta-learning models capable of uncertainty quantification across diverse applications. However, the standard NP framework assumes access to noise-free inputs. This is unrealistic in many practical settings, where sensor noise or measurement errors are unavoidable. This discrepancy between model assumptions and real-world data can significantly degrade prediction quality and uncertainty estimates. We introduce Denoising Neural Processes (DnNPs), a principled extension of the NP framework that explicitly models input uncertainty. We train DnNPs via variational inference. DnNPs jointly learn to denoise corrupted inputs and perform meta-learning, making them robust to noisy training data. We demonstrate DnNPs' effectiveness on both synthetic and real-world datasets, showing consistent improvements in prediction and uncertainty quantification compared to standard NPs. On PM_2.5 spatial interpolation from real-world air quality sensors, DnNPs reduces RMSE by up to 43% and NLL by up to 80% relative to standard NPs. We also provide theoretical grounding: we show that the Bayes optimal predictor under noisy inputs belongs to the DnNPs model class, and that plug-in models with fixed observation variance incur an irreducible KL divergence gap that DnNPs provably avoids.
PaperID: 1819, Poster
Authors:
Tianshi Che, Yang Zhou, Yushan Mu, Zeru Zhang, Zijie Zhang, Panqiu Xia, Wei Zhang, Longwei Wang, yelong shen, Ruoming Jin, Jianfeng GaoAbstract: As text-to-image diffusion models (T2I DMs) become widely deployed, adversarial advertising has emerged as a realistic threat: an attacker may compromise a T2I DM so that it implants target product brands into images generated from users' non-advertising prompts. Two key challenges remain largely unresolved in this setting: achieving natural and semantically coherent adversarial advertisement and ensuring robust adversarial advertisement. To address these challenges, we develop a new estimation algorithm for the multivariate continuously scaled phase-type with Lévy (MCPHL) distribution to capture the intrinsic distribution of natural advertisement prompts in the prompt-embedding space. With the estimated MCPHL, we construct an attack that pushes non-advertising prompts toward high-density regions of this distribution, making the resulting perturbed prompts better aligned with natural advertising prompts. We then propose a masked parameter smoothing approach grounded in mollification theory, yielding a smoothed T2I DM equipped with a dimension-invariant certified guarantee against advertisement degradation under model fine-tuning. The masking mechanism preserves utility by avoiding unnecessary smoothing on sensitive parameters, and our theoretical analysis shows that the smoothed model better preserves adversarial advertisements under fine-tuning, while maintaining better generation quality.
PaperID: 1820, Poster
Abstract: Omni-modal models are expected to reason over synchronized visual events and audio cues, yet answer accuracy alone can conceal failures in where an answer is grounded and which modality supports it. In audio-video question answering (AVQA), models may produce correct answers while relying on incorrect temporal spans, dominant-modality shortcuts, or unsupported cross-modal relations. We argue that this failure arises from treating reasoning as a completion-level prediction problem, where supervision and reward signals do not distinguish between support selection and reasoning over the support. To address this, we introduce OmniSE, a support-faithful optimization framework that makes question-conditioned audio-video support the central training object. Each QA instance is rewritten into an Audio-Visual Evidence Trace, a compact optimization unit that records temporal evidence, modality-specific cues, cross-modal relations, and cited reasoning steps. Given this support object, OmniSE first uses Support-Tree Rollout (STR) to explore alternative support hypotheses before sampling support-conditioned reasoning and answers. It then applies factorized reward assignment to score temporal localization, modality evidence, cross-modal binding, reasoning faithfulness, and answer correctness at their corresponding comparison levels. Finally, reward-gated self-consistency regularizes sibling continuations under reliable support, encouraging fixed trustworthy evidence to induce stable reasoning and answers. Experiments on multiple AVQA benchmarks show that OmniSE consistently outperforms strong open-source baselines at the same scale, demonstrating the importance of support-centric optimization for evidence-centric audio-video reasoning. Code and data would be publicly available after peer review.
PaperID: 1821, Poster
Abstract: Modeling resting-state functional magnetic resonance imaging (rs-fMRI) data is crucial for understanding brain-wide neural activity. However, traditional methods struggle to capture complex temporal dynamics over long horizons, to account for the brain's anatomical spatial structure, and to model high-dimensional ambient signals that lie on a low-dimensional intrinsic subspace. We propose FAST-Brain, a unified flow-aligned spatio-temporal surrogate brain model that addresses all three challenges. At its core is a flow-aligned generative framework that directly predicts the clean blood-oxygen-level-dependent (BOLD) signal, paired with a graph convolutional network that captures spatial structural constraints and a Transformer that models long-range temporal dependencies. Theoretically, we show that under a low-dimensional subspace assumption, the approximation error of our model scales with the intrinsic dimension rather than the ambient dimension, which justifies our direct modeling of the BOLD signal. Extensive experiments on synthetic and Human Connectome Project datasets demonstrate that FAST-Brain achieves state-of-the-art performance in recovering functional connectivity, effective connectivity, and the implicit low-dimensional signal subspace.
Authors: Karthik Sivachandran, Rohan Paleja
Abstract: Reinforcement learning agents trained to maximize their own reward in repeated interactions can converge to supra-competitive outcomes resembling explicit collusion, without communication or shared design. Existing mitigation approaches are largely tied to specific economic settings, like two-sided platforms and auctions, leaving open how to design interventions for general repeated games. We address this gap by formalizing the connection between empirical observations from prior work on Q-learning collusion and classical theory of Simple Penal Codes (SPCs). We show any non-trivial SPC induces a quantifiable conditional dependence in agents' policies, detectable via the total variation distance between an agent's action distributions across cooperation and defection histories. Building on this connection, we propose CURB (Collusion Unwinding via Reward shaping and Belief injection), a reward-shaping framework that penalizes this Total Variation (TV) distance signal during Q-learning and is guaranteed to convert any SPC fixed point of the dynamics into a trivial one, thus precluding collusive equilibria sustained by punishment threats. Empirically, CURB substantially reduces collusion by Q-learning agents in both Bertrand and Cournot Competition Repeated Games. We further demonstrate that CURB extends to deep Q-network agents in Bertrand competition, suggesting the mechanism generalizes beyond tabular Q-learning.
PaperID: 1823, Poster
Authors:
Jaeseok Jang, Seungmin Jeon, Kwang P Choi, Chang-Su KimAbstract: Progressive image compression reconstructs images from prefixes of a single bitstream and therefore involves sequential latent inference under partial observations. Existing trit-plane codecs mainly treat decoded trits as discrete symbols for probability prediction, without fully exploiting their role as interval constraints on the underlying continuous latents. We propose bound-conditioned latent inference (BLI), a framework that reformulates progressive trit-plane coding as sequential latent inference under shrinking interval constraints. At each coding step, the trits decoded so far define lower and upper bounds for each latent element. BLI uses the center and width of these bounds, together with hyperprior information shared by the encoder and decoder, to jointly refine the Gaussian parameters (\mu,\sigma) for entropy modeling and the latent estimate \haty for reconstruction. As additional trits are decoded, the feasible intervals shrink monotonically, yielding tighter constraints for the next inference step and enabling flexible in-plane refinement schedules. Experiments on Kodak, CLIC, and JPEG-AI show that BLI achieves state-of-the-art rate-distortion performance among progressive image codecs, reducing BD-rate by 17.80% over DPICT and by 7.98% over CTC on Kodak, while replacing CTC's plane-specific predictors with a shared model that uses about 5.3× fewer parameters and achieves lower latency than CTC.
Abstract: World models enable agents to predict future dynamics conditioned on actions, making the choice of latent state representation central to planning and control. Existing latents are often either learned directly from pixels with limited semantic structure or inherited from frozen visual foundation models with excessive task-irrelevant detail, yielding state spaces that are poorly matched to downstream planning and control. This is especially challenging in reward-free offline settings, where the model must learn from fixed trajectories without reward supervision or online interaction. To address this, we propose TC-WM, a framework for turning foundation-model embeddings into compact, task-sufficient world representations. The key design is to treat the pretrained embedding space as a semantic scaffold rather than as the final state space: TC-WM linearly projects high-dimensional visual embeddings into a compact latent, aligns a designated subspace with the agent’s physical state via contrastive learning, and reconstructs embeddings to preserve useful visual structure. This combines the generality of foundation features with the controllability of task-centric dynamics. Theoretically, we show that TC-WM suffices to identify the task-centric latent factors up to a simple transformation. Empirically, TC-WM enables test-time planning across diverse environments, e.g., Robomimic and D4RL, achieving better world modeling quality and more precise control compared to state-of-the-art world model approaches. Webpage: https://tc-wm.github.io/
Abstract: Code evolution is a family of techniques that rely on large language models to search through possible computer programs by evolving existing code. While code evolution pipelines have shown impressive performance across domains, many are highly complex and are typically not compared to simpler alternatives. To prevent bad comparisons from stagnating the field, we propose two simple baselines and compare them to popular pipelines across three domains: finding better mathematical bounds, designing agentic scaffolds, and machine learning competitions. Surprisingly, the baselines match or outperform existing methods across all domains, prompting a closer look at which factors drive performance in different settings. For mathematical bounds, we find that the search space matters far more than the exact search algorithm, with both simple and complex methods performing similarly given the same search space. Thus, expanding the search space is what matters for improving performance, and does so across all pipelines. Moreover, different domain knowledge embedded in prompts can significantly affect the search's efficiency, potentially leading to overestimating a pipeline's sample efficiency. For automated agentic scaffold design, we show that typical high-variance evaluations lead to all automated search methods picking subpar scaffolds relative to a simple majority vote. To mitigate this, we propose improved evaluation procedures that reduce stochasticity while keeping the search economically feasible. Overall, these results indicate that further improving code evolution's performance depends more on the setup surrounding the search pipeline than on the pipeline itself. We hope these insights will enable developing new code evolution methods and help spur meaningful progress in the field.
Abstract: Generative policies based on expressive model classes, such as diffusion and flow matching, are well-suited to complex control problems with highly multimodal action distributions. Their expressivity, however, comes at a significant inference cost: generating each action typically requires simulating many steps of the generative process, compounding latency across sequential decision-making rollouts. We introduce flow map policies, a novel class of generative policies designed for fast action generation by learning to take arbitrary-size jumps---including one-step jumps---across the generative dynamics of existing flow-based policies. We instantiate flow map policies for offline-to-online reinforcement learning and formulate online adaptation as a trust-region optimization problem that improves the critic's Q-value while remaining close to the offline policy. We theoretically derive Flow Map Q-Guidance Training (FMQ), a principled closed-form learning target that is optimal for adapting offline flow map policies under a critic-guided trust-region constraint. We further introduce Q-Guided Beam Search (QGBS), a stochastic flow-map sampler that combines renoising with beam search to enable iterative inference-time refinement. Across 12 challenging robotic manipulation and locomotion tasks from OGBench and RoboMimic, FMQ achieves state-of-the-art performance in offline-to-online RL, outperforming the previous one-step policy MVP by a relative improvement of 21.3% on the average success rate.
Authors: Ryo Kamimura, Thong Pham
Abstract: We propose Low-Rank Quantile Surfaces (LRQS), a bivariate causal model in which, in the causal direction, an unknown monotone transformation of the conditional quantile surface admits a low-rank functional decomposition. LRQS subsumes location-scale noise models and post-nonlinear heteroscedastic noise models, while allowing multiple quantile bases to represent changes beyond location-scale effects. We prove generic identifiability of LRQS: the transformed quantile surface is low rank in the causal direction, whereas reverse representability under the corresponding constraints occurs only for exceptional, fine-tuned cause marginals. We provide a simple-yet-powerful causal score using a nonparametric fitting procedure that alternates between rank-constrained approximation of discretized quantile surfaces and isotonic estimation of the unknown monotone transformation. Experiments on synthetic mechanisms with higher-rank distributional shape variation and strong nonlinear distortions, together with standard bivariate benchmarks, show that LRQS is especially effective when conditional distributional shape or observation distortion goes beyond existing location-scale assumptions.
Abstract: In autoregressive language models, each token is sampled by conditioning on all the past tokens; the overall string has thus been sampled from the correct underlying joint distribution represented by the model. In contrast, masked diffusion language models generate text by unmasking tokens out of order and potentially in parallel. Generating an overall string sampled from the correct underlying joint distribution would (again) require exactly one token unmasking in every full-model forward pass. The more tokens unmasked in parallel, the further away the string is from the true joint; this can be seen in the resulting drop in accuracy (but, increase in speed). In this paper we devise a way to \em approximately sample multiple tokens from the joint distribution in a single full-model forward pass; we do so by developing a new lightweight single-layer ``sampler" on top of an existing large diffusion LM. One forward pass of the full model can now be followed by multiple forward passes of only this sampler layer, to yield multiple unmasked tokens. Our sampler is trained to mimic exact joint sampling from the (frozen) full model. We show the effectiveness of our approximate joint sampling for both pretrained-only (Dream-7B-Base, Llada-7B-Base) and instruction-tuned (Dream-7B-Instruct, Dream-7B-Coder) models on downstream tasks (GSM8k, MATH, MBPP, HEval) and on unconditional language modeling. When eight tokens are unmasked for each full-model denoising step, our sampling algorithm achieves a MAUVE score of 0.84 (vs marginal baseline of 0.19) with respect to the true joint distribution.
Abstract: Real-world observational data often contain existing or emerging heterogeneous subpopulations that deviate from global patterns. Without the ability to detect Out-of-Distribution (OOD) data and model underlying structure, systems risk failing to adapt to emerging patterns, leading to inaccurate or potentially harmful predictions. We introduce DynaSub, a Dynamic Subgrouping Variational Autoencoder that couples latent representation learning with subgroup discovery. DynaSub operates on pre-trained foundation models or regular encoders, learning latent embeddings that define subgroup structure and are iteratively refined to enhance subgroup separability and OOD sensitivity. It incorporates a nonparametric clustering mechanism directly in latent space, enabling the number and structure of subgroups to adapt dynamically during training. DynaSub achieves competitive performance on near- and far-OOD detection across ResNet and ViT-based foundation encoders, reducing false positive rates by up to 10% under covariate shift while maintaining high AUROC across multiple OOD benchmarks, and excelling in class-OOD settings where entire classes are unseen during training.
PaperID: 1830, Poster
Abstract: Long-term personalized agents are increasingly expected to answer queries over user histories in which preferences, goals, and constraints evolve across interleaved conversations. Existing retrieval-based and structured-memory systems can surface relevant facts, but often return isolated memory units that do not explain which earlier preference was refined, contradicted, or superseded by later evidence. We introduce TransMem, a transition-aware memory retrieval framework that stores preference updates as explicit retrieval units alongside ordinary factual memories. During memory construction, TransMem routes user interaction history into topic-structured memory graphs, detects shift events, and grounds them as transition records. At inference time, TransMem first retrieves factual anchors and then adaptively expands over transition records, producing a compact context of factual anchors and accepted transition records. Experiments on PersonaMem and PersonaBench across two Qwen backbones show that TransMem outperforms representative retrieval-based and structured-memory baselines. Controlled analyses further show that transition records and adaptive expansion contribute beyond factual anchoring, chronological expansion, flat transition retrieval, and fixed-hop traversal. These results suggest that preference-update records are useful retrieval units for long-term personalization.
PaperID: 1831, Poster
Abstract: Sequential recommendation predicts which item a user will interact with next. A key property of this task is that user preferences are concentrated: only a small cluster of items is relevant at a given moment. Recent diffusion-based methods add noise to the item's embedding and score candidates by embedding similarity. Since similar items have similar embeddings, their score gaps are inherently small and easily overwhelmed by noise, causing the most relevant candidates to become indistinguishable first. We prove that moving diffusion to distribution space, where each item receives independent noise, preserves rankings better and makes the reverse denoising process easier. This shift, however, loses the semantic structure of embedding space and requires richer supervision than one-hot labels. We propose PreDiff ( usion), which performs diffusion in distribution space and introduces two components to address these problems: Distribution-Guided Sparse Projection projects the preference scores into embedding space while filtering out low-relevance noise, and soft targets encode inter-item similarity to teach fine-grained ranking within the preference cluster. On four benchmarks, PreDiff achieves 8%--17% relative improvement over existing methods.
PaperID: 1832, Poster
Authors: Erik Jhones Freitas do Nascimento, Jorge Franco, Amauri Souza
Abstract: Retrieval-augmented generation (RAG) strategies have empowered large language models (LLMs) through integration with external knowledge sources, leading to more accurate, up-to-date, and contextually relevant outputs. Recently, graph-based RAG methods have gained attention for leveraging relational structures that support multi-hop reasoning and retrieval. However, existing approaches remain limited in their ability to explicitly model query-aware relevance signals during retrieval. In this paper, we propose LP-RAG, a graph-based framework for document RAG that casts retrieval as an inductive link prediction problem. In particular, LP-RAG first constructs a graph encoding semantic relationships among chunks (i.e., candidates for retrieval), which is then augmented with chunk-conditioned synthetic queries that emulate potential user questions. This enables self-supervised learning of query-chunk relevance without relying on real query annotations. Retrieval is performed by predicting links between unseen queries and chunk nodes. Notably, LP-RAG is model-agnostic and can leverage a broad class of link prediction methods. To demonstrate its effectiveness, we evaluate LP-RAG across diverse settings and benchmarks. Our empirical results show that LP-RAG consistently outperforms existing methods in both retrieval quality and downstream generation performance, while maintaining competitive efficiency relative to learnable graph-based RAG approaches.
PaperID: 1833, Poster
Authors:
Jiayu Li, Umair Afzal, Zilong Zhao, Milad Abdollahzadeh, Uzair Javaid, Biplab SikdarAbstract: Time series are ubiquitous across different domains, and the growing use of synthetic data for privacy-preserving sharing, data augmentation, and stress testing makes time series generation increasingly important. Early generative approaches are predominantly autoregressive, which can suffer from low sample diversity, error accumulation, and difficulty modeling long-range dependencies. In this paper, we introduce FracTS, a hierarchical, coarse-to-fine autoregressive model for multi-variate time series generation. Inspired by fractal generative modeling developed for images, FracTS adapts the paradigm from 2D to 1D sequences by learning a multi-scale representation that captures global trajectory structure at coarse scales and local dynamics at fine scales while following the autoregressive nature of time series. In contrast to unstructured images, time series are structured, allowing easy automated or designated temporal feature processing optimized for the model, and global (static) features extraction describing the instances that generate the corresponding time series in the collection. We incorporate them in the generation process to further improve the generation quality. The design naturally handles multivariate dependencies and improves long-horizon temporal coherence without requiring fixed-length inputs. Experiments on five datasets (two conditional, three unconditional) conducted against seven baselines show that FracTS consistently outperforms baselines across marginal distribution, correlation preservation, and ML utility. In particular, the ML utility is improved up to 98.1% comparing to the second-best baseline. Code is available at: https://anonymous.4open.science/r/fracts-91D621
Authors:
Wentao Zhang, Liliana Hotsko, Woojeong Kim, Pengyu Nie, Stuart Shieber, Yuntian DengAbstract: Many everyday programming tasks resist clean rule-based implementation, such as alerting on important log lines, repairing malformed JSON, or ranking search results by intent, and are increasingly outsourced to large language model APIs at the cost of locality, reproducibility, and price. We propose : compiling such a function from a natural-language specification into a compact, locally-executable neural artifact. We instantiate this paradigm with Program-as-Weights (PAW), in which a 4B compiler trained on FuzzyBench, a 10M-example dataset we release, emits parameter-efficient adapters for a frozen, lightweight interpreter. A 0.6B Qwen3 interpreter executing PAW programs matches the performance of direct prompting of Qwen3-32B, while using roughly one fiftieth of the inference memory and running at 30 tokens/s on a MacBook M3. PAW reframes the foundation model from a per-input : invoked once per function definition, it produces a small reusable artifact whose subsequent calls per function application are cheap and offline.
Abstract: Code assistants are increasingly utilized in test-driven software development, yet the theoretical mechanisms behind their environment-interaction strategies remain underexplored. We provide a probabilistic framework for two dominant paradigms: code selection after generation using the execution environment, and code generation conditioned on environment feedback. First, we formalize several well-established selection heuristics as environment-aware estimators of code correctness. We theoretically prove that estimators based on fuzzy functional similarity add an inductive bias and strictly dominate estimators based on functional equivalence in terms of signal-to-noise ratio. Second, we frame backprompting as an in-context approximation of Thompson sampling. We derive a novel regret bound for reward functions with unobservable components, theoretically explaining why the effectiveness of backprompting is limited by the ambiguity of the informal task description (an irreducible regret). Using five state-of-the-art open weight models, we corroborate these findings across BigCodeBenchHard, LeetCodeDataset, and QiskitHumanEvalSim. Our formalization also suggests how to improve task descriptions effectively, leading to a new benchmark, QiskitHumanEvalSimX.
PaperID: 1836, Poster
Abstract: Open-vocabulary dense perception requires grounding language-specified concepts in local regions for recognition, localization, and segmentation beyond closed category sets. CLIP learns image-level vision-language alignment, whereas dense prediction requires spatially precise and semantically reliable local features. Although recent dense CLIP adaptation methods reduce this gap through region-level semantic transfer and spatial correlation guidance, maintaining stable local semantics under contextual variation remains challenging. Specifically, the global contextual modeling in CLIP may introduce undesirable dependencies into local features. Changes in background, scale, viewpoint, or scene layout may cause identical object regions to yield inconsistent dense representations, leading to unstable correspondences and semantic drift. Moreover, spatial correlation guidance from visual foundation models captures visual relatedness between patches, but not whether coherent regions correspond to the queried textual concept, leaving concept-level dense alignment under-constrained. To address these issues, we introduce Semantic Spatial Consistency Learning (SSCL), an adaptation method with two complementary objectives: 1) Context-Invariant Spatiotemporal Consistency (CISC) enforces semantic consistency for the same anchor region under varying contexts to mitigate context-induced feature drift; 2) Structure-Guided Semantic Refinement (SGSR) leverages spatial affinities from a frozen visual foundation model to refine CLIP text-patch responses, producing text-conditioned dense supervision with improved spatial coherence and concept-level alignment. Experiments on video instance segmentation, region classification, semantic segmentation, and object detection show consistent gains over the baseline methods, demonstrating the effectiveness and generalization capability of the proposed SSCL. Our code and models will be made publicly available.
PaperID: 1837, Poster
Abstract: Knowledge editing has become a key technique for updating knowledge in large language models. However, existing methods neglect the geometric structure of the knowledge representation space, leading to two key limitations: they cannot distinguish which pre-trained knowledge should be preferentially preserved (core vs. non-core) nor which new edits are inherently difficult (hard vs. easy). To address this, we propose \textscGEdit, a geometry-aware framework built on a Low-Rank Plus Diagonal (LR+D) factor model that captures the intrinsic low-rank structure of Transformer FFN hidden representations. \textscGEdit features two complementary mechanisms. First, the low-rank regularization applies targeted protection on a low-rank core knowledge subspace. Second, the adaptive regularization dynamically adjusts the regularization strength based on each edit request's editability: relaxing constraints for hard edits to ensure success, and tightening them for easy edits to better preserve pretrained knowledge. Extensive experiments on \textttGPT2-XL, \textttGPT-J, and \textttLLaMA-3-8B under large-scale editing scenarios—including both mini-batch and fully sequential editing settings with up to 10,000 continuous updates—show that \textscGEdit consistently and significantly outperforms state-of-the-art baselines in both editing success and knowledge preservation.
PaperID: 1838, Poster
Abstract: Causal Abstraction can be used to define a notion of equivalence between a neural network and an algorithm, thus serving as a powerful framework for mechanistic interpretability. In this paper, however, we show the existence of a class of spurious features which straightforwardly pass such equivalence tests. Suppose there are two directions in a model's representation space where: (i) one discriminates the input based on some specific feature, and (ii) the other steers the model towards a certain output. Causal Abstraction techniques can create a pathway leveraging these unrelated directions to simulate their combined behavior in the model, which thus spuriously appears to implement a certain feature. In such cases, we say the method discovered a parasite feature. Experimentally, we show that Distributed Alignment Search---a Causal Abstraction-based method---finds parasite features using linear alignment maps in two tasks: syntactic agreement and two-digit addition. Our paper highlights an important consideration for Causal Abstraction and contributes to the ongoing debate about best practices for mechanistic interpretability.
PaperID: 1839, Poster
Authors: Masato Uchida
Abstract: The attention mechanism in Transformers links queries, keys, and values through a sequence of computations, while the principles that constrain its form remain unclear. This paper examines attention under natural requirements that formalize basic properties of attention computation as transformations between representation spaces, and asks which forms of attention computation are admissible. We show that similarity computation, score-to-weight mapping, and value aggregation are unified by dual structures induced by convex potential functions and their Legendre transforms. In particular, the standard inner-product score is recovered as a limiting case of a potential-based score, the score-to-weight map must take a gradient-map form under structural requirements on attention normalization, and the value-aggregation rule must take a dual-coordinate barycentric form under structural requirements on representative-point aggregation. These results show that the standard components of attention are structurally constrained by the underlying dual structure, while the choice of potentials provides explicit degrees of freedom for understanding and modifying attention mechanisms.
Abstract: While modern tabular learners excel at capturing statistical patterns, they frequently operate in a semantic vacuum, treating textual features as discrete symbols, ignoringing the rich semantics inherent in feature names or cell entries. We propose CASE (Context-Aware Semantic Embeddings), a novel framework that bridges the gap between the semantic understanding of Large Language Models (LLMs) and the statistical capabilities of tabular learners. Unlike existing methods that embed rows in isolation, CASE utilizes a contextualization strategy: we pre-fill the KV cache of a custom-trained Gemma~3-based Tabular Language Model with a representative sample of rows to establish a persistent anchor of the dataset’s semantics. This ensures that generated row embeddings are dynamically contextualized, resolving semantic ambiguities and anchoring representations in domain-specific context. Our experiments across several benchmarks (CARTE, TextTab, and TabArena) demonstrate that CASE significantly improve performance of tabular learners -- particularly in low-data regimes and on semantically rich datasets -- setting a new state of the art when combined with recent tabular in-context learners. Inference code and model checkpoints will be made publicly available.
Abstract: We develop a game-theoretic framework for predicting and steering the behavior of populations of large language models (LLMs) through Nash equilibrium (NE) analysis. To avoid the intractability of equilibrium computation in open-ended text spaces, we model each agent’s action as a mixture over human subpopulations. Agents choose actively and strategically which groups to align with, yielding an interpretable and behaviorally substantive policy class. We derive closed-form NE characterizations, adopting standard concave-utility assumptions to enable analytical system-level predictions and give explicit, actionable guidance for shifting alignment targets toward socially desirable outcomes. The method functions as an active alignment layer on top of existing alignment pipelines such as RLHF. In a social-media setting, we show that a population of LLMs, especially reasoning-based models, may exhibit political exclusion—pathologies where some subpopulations are ignored by all LLM agents—which can be avoided by our method, illustrating the promise of applying the method to regulate multi-agent LLM dynamics across domains.
Abstract: Information-theoretic generalization bounds based on the supersample construction are a central tool for algorithm-dependent generalization analysis in the batch i.i.d.~setting. However, existing supersample conditional mutual information (CMI) bounds do not directly apply to sequential decision-making problems such as online learning, streaming active learning, and bandits, where data are revealed adaptively and the learner evolves along a causal trajectory. To address this limitation, we develop a sequential supersample framework that separates the learner filtration from a proof-side enlargement used for ghost-coordinate comparisons. Under a row-wise exchangeability assumption, the sequential generalization gap is controlled by sequential CMI, a sum of roundwise selector--loss information terms. We also establish a Bernstein-type refinement that yields faster rates under suitable variance conditions. The resulting bounds specialize to online learning, streaming active learning with importance weighting, and stochastic multi-armed bandits.
PaperID: 1843, Poster
Abstract: Quantum reinforcement learning (QRL) has emerged as a promising approach for controlling quantum systems and accelerating learning tasks with quantum resources. However, most existing QRL formulations still retain a classical reinforcement learning interface: states and actions are often encoded in computational bases and rewards are extracted through measurement. This is less natural for fully quantum environments in which the relevant state information may live in an unknown basis and intermediate measurements can disturb the trajectory. This motivates a fully QRL framework in which the agent, environment, rewards, and return accumulation are all represented quantum mechanically. In this paper, we take a step toward this goal and introduce a new QRL framework in which state, action, reward, and reward-accumulation registers evolve without intermediate measurement. In this framework, both the policy and the environment are modeled by : the policy updates the action register conditioned on the state, while the environment updates the state and reward registers conditioned on the action, without measuring or overwriting the corresponding control register. We then propose a quantum policy-gradient algorithm for parameterized conditional quantum channels with a trainable control basis. We also derive finite- and infinite-horizon gradient estimators using a new parameter-shift rule. The algorithm samples a single shifted policy block from a geometric distribution, yielding an unbiased gradient estimator under bounded rewards. Finally, we establish convergence of the resulting stochastic gradient-ascent procedure to stationary points under standard smoothness and step-size assumptions.
PaperID: 1844, Poster
Abstract: Do language models make decisions under uncertainty like humans do? And if so, what role does extended reasoning play in the underlying decision process? We answer this question by introducing an active probabilistic reasoning task that cleanly separates sampling (actively acquiring evidence) from inference (integrating evidence towards a decision). Benchmarking humans and a broad set of contemporary LLMs against optimal reference policies reveals a consistent pattern: extended reasoning is the key determinant of strong performance, driving large gains in inference, while yielding only modest improvements in active sampling. To explain these differences, we fit a behavioral model that captures systematic deviations from optimal Bayesian behavior through interpretable parameter families, placing humans and models in a shared low-dimensional cognitive space. The resulting fits show how reasoning shifts models toward human-like regimes of evidence accumulation and belief-to-choice mapping, and yield testable predictions about the latent dynamics that might drive each decision. Probing residual-stream activations of an open-weight reasoning model, we find that the geometry of internal representations tracks these predicted dynamics, linking behavior to representational correlates of the fitted latent dynamics.
Abstract: Reinforcement learning from verifiable rewards improves reasoning by reinforcing sampled solutions that receive positive outcomes, but group-relative methods such as GRPO become silent on hard prompts where every sampled rollout is wrong. These zero-reward prompts are often the most informative failures: each comes with a verified reference solution, yet uniformly imitating the full trace mixes core reasoning decisions with routine algebra, formatting, and surface wording. We propose , an auxiliary repair update that leaves GRPO rollout generation unchanged. For each failed prompt, SORT extracts a reference-derived reasoning plan and scores every reference token twice under the same model, once with only the problem and once with the plan added to the context. Tokens whose probabilities rise under plan conditioning are treated as structurally informative and receive larger Dynamic Fine-Tuning updates, while tokens outside the model's support remain controlled by the base probability. We formalize this mechanism with a ground-truth plan model and prove that the SORT weight approaches an oracle structural-token weight as the extracted plan better approximates the true plan-conditioned distribution. Across three instruction-tuned backbones and eight in-distribution and out-of-distribution reasoning benchmarks, SORT consistently improves over GRPO and guidance-based baselines, with the largest gains on the weakest model where zero-reward failures are most frequent.
Abstract: Memory capacity is a critical factor determining the performance of Vision-Language-Action (VLA) models in long-horizon manipulation tasks. Existing memory-augmented architectures primarily rely on linear or flat storage, lacking structural priors for manipulation categories and hierarchical organization. This deficiency hinders efficient experience retrieval and limits generalization to unseen long-horizon task compositions. Inspired by the hierarchical organization of human experience, we propose ECHO (Experience Consolidation and Hierarchical Organization), a novel memory framework operating within a Continuous Hierarchical Space. By employing a hyperbolic autoencoder, ECHO maps VLA hidden states into this space. Leveraging hyperbolic metrics and entailment constraint mechanisms, experience vectors are organized into a semantic memory tree that supports efficient top-down retrieval. In parallel, a background consolidation mechanism continuously refines the memory tree through geometric interpolation and structural splitting, supporting virtual memory synthesis in the continuous space. We integrate ECHO into the \pi_0 foundation model. Evaluations on LIBERO and preliminary real-world experiments demonstrate the effectiveness of our approach, notably achieving a 12.8% absolute improvement in execution success rate over the \pi_0 baseline on LIBERO-Long, while improving compositional generalization on cross-suite unseen long-horizon tasks.
PaperID: 1847, Poster
Abstract: We study exact community recovery in the two-community stochastic block model on n vertices under limited and noisy access to network data. The learner may query a noisy neighborhood oracle that reveals each true neighbor of a queried vertex independently with fixed probability and never returns non-neighbors, subject to a finite query budget. We consider both oracle-only access and a combined model where the learner also observes a single subsampled copy of the underlying graph. For oracle-only access, balanced uniform querying gives a sharp non-adaptive benchmark: when each vertex is queried the same integer number of times, the observations reduce to an SBM with attenuated edge probabilities and the Abbe--Bandeira--Hall exact-recovery threshold applies. We show that this benchmark is not adaptively optimal: a two-stage adaptive strategy succeeds with n+o(n) queries in a regime where balanced uniform querying requires m n queries for some m>1. With an additional subsampled graph, we prove a sublinear-query adaptivity gap: balanced data-independent uniform querying with a sublinear budget does not improve over the subsampled graph alone, whereas adaptive querying can target a small set of uncertain vertices and achieve exact recovery. Thus adaptive data acquisition can strictly improve the information-theoretic limits of exact recovery.
PaperID: 1848, Poster
Abstract: Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that encouraging exploration during LLM reinforcement learning can improve downstream performance. However, methods for controlling exploration often rely on heuristics like clipping or temperature scaling that are unable to fundamentally change token-level distributions, which might limit exploration. Here, we explore controlling exploration via variational learning, where a distribution over neural network parameters is learned via noisy optimization, with parameters sampled from an approximate posterior. We introduce Perturbed Parameter Policy Optimization (3PO), where the amount of noise that is added to model parameters functions as an additional control lever that can be tuned for exploration. We show that using multiple parameter samples within a batch performs better than using a single sample. When using multiple parameter samples in group-based methods like GRPO, we find that calculating advantages across rollouts from all parameter samples performs best. We call this Chunked Perturbed Parameter Policy Optimization (C3PO). We show empirically on a range of math reasoning benchmarks and LLMs that using multiple parameter samples can improve downstream performance, especially on harder benchmarks like AIME. Overall, our work presents evidence that parameter-space exploration can improve LLM reinforcement learning.
PaperID: 1849, Poster
Authors: Roie Kazoom, George Leifman, Genady Beryozkin
Abstract: We present FM-ChangeNet, a pathwise-supervised framework for change detection that reformulates bi-temporal reasoning as continuous transport in feature space rather than static endpoint comparison. Given encoded pre and post-temporal representations, we construct intermediate latent states and learn a time-conditioned velocity field \hatv_\theta(z_t,t) along the transformation trajectory. This pathwise formulation constrains the predictor over a continuum of intermediate states, providing a denser and less ambiguous supervision signal than conventional endpoint-only segmentation and enabling the model to capture temporal evolution explicitly. The learned velocity field is not only a transport mechanism but also an interpretable representation of change: its magnitude serves as a spatially localized change cue that helps distinguish true structural variation from nuisance effects such as illumination shifts and spatial misalignment. We develop a hierarchical multi-scale architecture with cross-temporal alignment, time-conditioned coarse-to-fine flow decoding, and a unified objective that couples flow supervision, trajectory consistency, spatial regularization, and segmentation loss. Experiments on remote sensing benchmarks show that the proposed framework produces more structured and robust change representations while achieving state-of-the-art performance.
Authors: Dimitris Arabadjis
Abstract: We study the problem of recovering a globally consistent Euclidean embedding of data, given only a local distance graph and propose a method that optimally represents these distances. The method operates solely on a neighborhood graph weighted by pairwise distances, without requiring any prior vector representation of the data. The embedding is obtained by solving a variational problem that matches local, on‑graph distances to the Euclidean metric, induced by the differentials of the embedding functions. The resulting Euler–Lagrange equations are derived in a coordinate‑free form, enabling direct evaluation of all operators from the distance graph alone. Though non-linear and missing an explicit expression for their non-linearity, these equations are shown to be resolved as an iteratively updated sparse linear problem. The main contributions of the proposed approach are (a) the derivation of the functional equations governing the optimal Euclidean embedding in the continuum, (b) a representation‑free formulation that requires only a neighborhood distance graph and no feature vectors and (c) an estimation procedure based exclusively on local graph operations. We experimentally evaluate the resulting non‑parametric algorithm on synthetic manifolds and real datasets, demonstrating consistent preservation of local metric structure and neighboring relations, while approximating the global isometric embedding.
PaperID: 1851, Poster
Abstract: We introduce NAGO, a noise-aware generative operator for efficient sampling from invariant measures of stochastic partial differential equations (SPDEs) in function space. Our central idea is to treat the noise measure as a design variable rather than a fixed default. While existing generative approaches mainly shape the sampling process through interpolation schedules or guidance heuristics, NAGO directly controls the marginal velocity field in functional flow matching by choosing a noise measure with suitable regularity. We establish conditions under which the resulting noise-induced flows are well defined in infinite-dimensional settings, and show how noise regularity affects the spectral geometry and stiffness of the marginal velocity field. This analysis leads to a principled design rule: match the regularity of the noise measure to that of the target invariant measure. The resulting method enables stable training and fast sampling without relying on downstream correction mechanisms. Experiments on three representative SPDEs with multiscale invariant measures show that NAGO consistently improves the accuracy--efficiency trade-off over standard flow matching baselines, while achieving up to two orders of magnitude speedup compared with classical numerical solvers.
Abstract: Scaling laws and hyperparameter tuning remain central challenges in training large-scale models, particularly under differential privacy (DP), where optimization is further complicated by noise injection and per-sample gradient clipping. While Maximal Update Parametrization (\muP) allows for zero-shot hyperparameter transfer in standard training, we show that it fails under DP constraints due to the distinct spectral properties of high-dimensional noise versus low-rank gradients. In this work, we introduce DP-\muP, a theoretically grounded framework that extends \muP to the privacy-preserving setting. By analyzing the interaction between gradient clipping, noise injection, and feature learning, we derive a Signal-to-Noise Ratio (SNR) condition that governs layer-wise learning rate scaling. Combining this with our proposed spectral scaling formulation, we provide specific scaling rules for DP-SGD, DP-Adam, and DP-Muon. Empirically, we demonstrate that DP-\muP enables effective zero-shot transfer of learning rates across model widths on diverse architectures (MLP, ViT, GPT-2), matching or exceeding the performance of individually tuned baselines while significantly reducing computational costs.
Abstract: Gradient-based attribution methods are model-faithful and scalable, but Integrated Gradients (IG) can be brittle because explanations depend on heuristic baselines, straight-line paths, discretization, and saturation. We propose Fisher--Rao Integrated Gradients (FRInGe), which defines both the reference and interpolation schedule in predictive distribution space. FRInGe replaces input baselines with a maximum-entropy predictive reference and follows a Fisher--Rao geodesic on the probability simplex. The corresponding input-space trajectory is realized through the pullback Fisher metric and stabilized by KL and Euclidean trust regions; attributions are obtained by integrating input gradients along this trajectory. Across six ImageNet architectures, FRInGe most clearly improves calibration-oriented attribution metrics, especially MAS scores, while remaining competitive on perturbation AUC and infidelity.
PaperID: 1854, Poster
Abstract: Graph representation learning often relies on message passing or spectral/positional encodings to summarize graph structure, but these indirect summaries can collapse structurally distinct graphs, including non-isomorphic cospectral pairs. We propose A2I, a structural encoding framework that renders BFS--degree-ordered local adjacency patterns as fixed-resolution images and embeds them with a frozen pre-trained vision encoder. The resulting structural tokens are mapped to learnable prototypes and aggregated with node embeddings via a lightweight Transformer. We provide a conditional analysis with two related findings: a stability result, showing that two orderings of the same graph yield embeddings that converge under bounded-degree assumptions, and a separability result, showing that two distinct graphs yield embeddings whose gap grows at least linearly with the difference in their BFS profiles. Empirically, A2I is competitive with recent GNN, Graph Transformer, and structural-encoding baselines on six node-classification benchmarks, with pronounced gains for GCN-based models on heterophilous graphs and under partial feature masking.
PaperID: 1855, Poster
Abstract: Ridge crossings, where two components share the same instantaneous frequency (IF) at a given time, are a fundamental bottleneck in multi-component time-frequency (TF) mode decomposition. We trace this difficulty to \emphinsufficient representational dimensionality and resolve it by lifting signals into the 3-D time-frequency-frequency-rate (TFFR) space (t,f,\dotf), where crossing ridges separate because they typically differ in the instantaneous frequency rate (IFR). We formalize this via a separability proposition and confirm it empirically through Monte Carlo simulation (crossing rate: 78.6% in 2-D \to 0.9% in 3-D). Based on this principle, we propose TFFR-Net, a query-based masked transformer that jointly predicts per-ridge masks and dense IFR maps. An IFR feedback loop injects the estimated frequency rate back into each query embedding, enabling the decoder to perform instance discrimination in TFFR space rather than the TF plane. On a synthetic benchmark with up to six chirp components under broadband noise (0--25\,dB SNR), TFFR-Net consistently reduces IF MAE by at least 5.8% and Dice loss by at least 5.4% relative to competing methods. Evaluation on real-world signals (bat echolocation, power-system oscillation, earthquake vibration) demonstrates effective generalization beyond synthetic training data. The source code is available at \urlhttps://anonymous.4open.science/r/TFFR-92D2/.
Abstract: Generative sequence models are often trained to plan motion in physical domains, from robotics to mechanical simulations. When constructing a dataset to train such a model, engineers may curate demonstrations to specify how trajectories should be distributed over a physical quantity like travel distance or mechanical energy. For example, a roboticist building a maze navigation agent might choose demonstrations whose travel distances cover a fixed range uniformly, hoping to constrain the agent's expected power usage. We find that standard deep learning can violate this intent: each generated trajectory can seem plausible on its own, but the aggregate distribution over the physical quantity is wrong. We call this failure physical misgeneralization, and develop an account of its mechanism. Using controlled synthetic tasks, we show that physical misgeneralization arises when local errors typical of the model class propagate through the physical measurement to shift the recovered distribution. We estimate these errors with a data deviation kernel, and we use it to predict which physical quantities gain or lose mass in both our synthetic and more applied maze navigation and double-pendulum motion tasks. Finally, our mechanistic interpretation helps identify which mitigation strategies are structurally promising, and we use it to propose a kernel-informed intervention.
PaperID: 1857, Poster
Abstract: We study zeroth-order optimization of non-convex functions with the aid of directional hints, which are cheap but potentially inaccurate approximations of the true gradient direction, given by linear subspaces at each iteration. To leverage these hints adaptively while maintaining robustness to their quality, we introduce Control-Variate Zeroth-Order Descent (CV-ZOD), a new framework that refines the classical zeroth-order gradient estimator with a control variate that can be set based on the directional hints. We first show that the oracle algorithm that optimally sets the reference vector and step size at each iteration achieves a convergence rate that interpolates between the first-order O(1/T) rate and the zeroth-order O(d/T) rate, depending on the quality of the hints along the trajectory. We then develop a practical variant of CV-ZOD that achieves the same oracle guarantee up to logarithmic factors, without any prior knowledge of the hint quality. We validate the method empirically on simulation-based scientific optimization tasks, demonstrating sustained progress on nonconvex landscapes where zeroth-order descent is slower and existing guided methods stall as guidance deteriorates.
PaperID: 1858, Poster
Abstract: The sequential nature of autoregressive decoding in large vision-language models creates a fundamental trade-off between latency and clinical responsiveness when managing asynchronous signals. When unexpected healthcare occurrences arise during an ongoing VLM response, integrating unanticipated events---such as sudden device alerts or urgent caregiver messages---conventionally requires either injecting the new data directly into the active response stream or deferring it until generation ends. Neither approach is ideal in dynamic environments that demand urgent intervention without repeatedly disrupting routine continuations. We present ReflexMind, a training-free framework for online, urgency-aware occurrence handling that runs an Urgency Evaluator (UE) concurrently with the Primary Generator (PG). By leveraging the rotational invariance of Rotary Position Embedding (RoPE), ReflexMind enables a dual-view decoding process that shares a single KV cache, allowing it to assess incoming multimodal occurrences while preserving the state of the ongoing response. This design supports structured decisions, enabling ReflexMind to escalate critical alerts immediately while handling non-urgent occurrences without interrupting the ongoing response. Across three benchmarks spanning multimodal occurrences, our method generally improves occurrence triage, urgency recognition, and intervention selection for VLM backbones at different scales, while using fewer post-occurrence tokens in most settings and adding modest runtime and memory overhead under concurrent occurrences.
PaperID: 1859, Poster
Abstract: We study decentralized stochastic smooth convex optimization, where M workers minimize an average objective using local stochastic gradients and neighbor-only communication over a fixed gossip network. A central question in this setting is to determine the largest number of workers that can be used under a total budget of N gradient samples while still preserving the centralized O(1/\sqrt N) statistical rate. We introduce an accelerated decentralized method that preserves this rate for M \lesssim \sqrt\rho N^3/4 workers up to logarithmic factors, where \rho is the spectral gap of the gossip network, improving the best prior scaling of M \lesssim \rho\sqrt N [Eisen et al., 2025]. The method is based on a one-step-delayed stochastic acceleration scheme that enables workers to interleave minibatching with accelerated gossip while controlling residual disagreement, and its guarantee depends only logarithmically on the optimum-local heterogeneity. We also establish a matching lower bound for linear-span decentralized first-order methods, showing that the method is optimal up to logarithmic factors.
PaperID: 1860, Poster
Abstract: Reinforcement learning (RL) with verifiable rewards has driven recent gains in large language model (LLM) reasoning, but whether it creates new capabilities or merely sharpens existing ones remains debated, especially for compositional generalization, where models must combine learned primitives into novel multi-step functions. Using a controlled program synthesis domain over a typed domain-specific language (DSL) of 119 primitives with verifiable compositional depth and irreducibility, where the base model has zero capability prior to fine-tuning, we show on Qwen3-8B (and replicate within-family at 4B and 32B, and cross-family on Llama-3.1-8B) that compositional capability is established during supervised fine-tuning (SFT) and depends on irreducibility at two levels of SFT data: structural irreducibility of compositional exemplars and observational irreducibility, where prompt input-output pairs are chosen adversarially to rule out shallower programs. Neither alone is sufficient; even moderate contamination with structurally reducible exemplars substantially degrades performance, and full contamination yields zero pass@64. Given this SFT foundation, standard outcome-reward RL (GRPO, PPO) maintains but does not extend compositional performance, contrasting with prior results in settings with simpler nested function execution. Compositional capability is bottlenecked by structural properties of SFT data, before RL begins. Code release pending institutional IP review.
PaperID: 1861, Poster
Abstract: Diffusion policies have emerged as a powerful policy class for offline reinforcement learning (RL) due to their ability to model complex, multimodal action distributions through iterative denoising. However, in offline RL, expressivity alone is not enough. Generated actions are supposed to remain within value-reliable regions of the action space to avoid extrapolation error. We find that standard diffusion policies often produce unstable denoising trajectories whose intermediate and final action candidates drift toward weakly supported regions, making critic-guided action selection unreliable and degrading teacher quality for one-step distillation. This issue is particularly pronounced on mixed-quality offline datasets, where suboptimal behaviors amplify trajectory instability. To address this problem, we propose Trajectory Consistency Rectification (TCR), a framework designed to promote denoising stability during both training and inference. At inference time, TCR aggregates locally consistent, high-value action candidates across denoising steps to suppress critic-unreliable outliers. During training, it adaptively reweights denoising steps to improve trajectory-level coherence. We show that this stabilization provides a more reliable teacher for one-step flow-matching distillation. Extensive experiments on D4RL and OGBench show that TCR consistently improves over diffusion-policy baselines, with significant gains in sparse-reward domains like AntMaze and Adroit. Moreover, one-step policies distilled from TCR achieve competitive performance among efficient single-step methods. These results suggest that enhancing the temporal consistency of denoising trajectories is a key ingredient for robust offline diffusion reinforcement learning.
PaperID: 1862, Poster
Abstract: Reinforcement learning (RL) is now standard for post-training large language models. The same reward optimization that elicits useful capabilities, however, can also reinforce latent backdoors planted earlier in the pipeline: attack success rates under 0.2% after SFT rise to 40--98% on training-distribution inputs after RL, and generalize to 13--30% on held-out evaluation, with no modification to the RL data, reward function, or training loop. We study this phenomenon in the context of agent models, where SFT poisoning teaches the victim model (Qwen3-8B and a small frontier model) to call an attacker-controlled oracle. Because the oracle can be set up to always return correct answers, calling it earns higher reward than the model's own attempts, and RL reinforces the behavior. We further show that the oracle's RL-time responses can instill persistent biases, such as brand preferences, that survive into deployment even when no tool calls happen. Such patterns appear difficult to detect with current tools: traditional guard models do not flag this mechanism, and LLM-based auditors remain unreliable, achieving only 4.5% precision even after iterative prompt refinement. Our findings point to the value of post-RL safety evaluation, in particular scrutinizing tool-call patterns such as unnecessary external invocations.
PaperID: 1863, Poster
Abstract: We study why diffusion autoencoders can achieve similar image quality while learning substantially different latent structures. We trace this behaviour to optimisation dynamics; we analyse curves of image reconstruction against latent representation quality, revealing trajectories that organise around two distinct regimes early in training. Models in the reconstruction regime prioritise image fidelity early, whereas those in the disentanglement regime improve reconstruction and disentanglement more gradually. We hypothesise that this behaviour can be influenced by targeting shortcut pathways in the diffusion U-Net and controlling early noise-level exposure, thereby shaping the reconstruction-disentanglement trade-off during training. To steer optimisation toward stronger representations, we introduce SteeringDRL, combining gated residual U-Nets with a simple noise-level exposure curriculum for training. Across disentanglement benchmarks, SteeringDRL improves representation quality and reduces seed sensitivity. Our method further extends to spatial disentanglement in object-centric learning, improving segmentation quality on synthetic and real-world datasets.
PaperID: 1864, Poster
Abstract: With the growing number of fine-tuned variants derived from pre-trained models, model merging emerges as a key challenge: effectively integrating task-specific adaptations while preserving performance. However, naive weight averaging often fails when parameter updates are misaligned, revealing the need for a structured space in which merging becomes principled. In this work, we hypothesize that pre-trained models possess an intrinsic low-rank structure. Specifically, we posit that the transformation from input to output is mediated by a latent bottleneck of low intrinsic dimensionality; we leverage this low-dimensional space to construct a model embedding space. This space ensures that fine-tuned updates, i.e., model embeddings, of all adaptations remain confined within the intrinsic structure of the pre-trained model. As model embeddings lie in the same low-dimensional space, model merging via direct averaging becomes more effective. Moreover, we further introduce an aligned subspace of these model embeddings to improve the efficacy of model merging. Experiments across vision and multimodal tasks, including ViT and multimodal LLMs, demonstrate that our approach enables effective model merging and provides a unified geometric view of adaptation and combination in pre-trained models.
PaperID: 1865, Poster
Abstract: We propose, HaNDF, object-conditional neural distance fields as data-driven priors for modeling the plausible hand–object configurations, in the product manifold of hand articulations. While recent unconditional generative priors showed success in modeling human pose, hands are versatile articulated bodies often found in interaction with other objects. This makes it crucial to develop a generative prior which can model the plausible hand poses conditioned on the object under interaction. To this end, our HaNDF proposes a geometry-aware, object-conditioned neural distance field, representing the plausible articulations in the zero-level set of a conditional field. We leverage tools from Riemannian geometry to (i) represent plausible hands in the zero level-set of an object-induced neural field, and (ii) project a given pose onto the this field. Using HaNDF, we can also optimize for a plausible condition, determining the object under interaction. Our extensive evaluations demonstrate that our HaNDF provides a flexible and powerful prior for hand–object interaction, suitable for applications such as pose refinement, reconstruction under occlusion, and physically plausible manipulation synthesis.
PaperID: 1866, Poster
Abstract: Pretrained generative models, such as diffusion and flow matching models, have demonstrated exceptional capabilities in synthesizing high-quality data. However, steering their generative trajectories to satisfy strict constraints—such as identity preservation in image editing or obstacle avoidance in robotics—remains fundamentally challenging. While recent training-free guidance methods enforce path-safety, they rely heavily on min-norm controllers that only account for immediate constraint satisfaction. These myopic approaches often yield suboptimal trajectories that hover dangerously close to the boundary, severely restricting the model's ability to optimize toward its target. In this paper, we propose AegisFlow, a training-free, non-myopic path-safe guidance framework that recasts guided generation as a safety-critical optimal control problem with global scope. By integrating a time-varying barrier function directly into the global cost functional, AegisFlow naturally enforces constraints along the entire generative path. To compute the global optimal control, we derive a tractable, mathematically grounded approximation of the Hamilton-Jacobi-Bellman (HJB) equation. Comprehensive experiments demonstrate that AegisFlow outperforms state-of-the-art methods across text-guided image manipulation and robotic planning, establishing a significantly better trade-off between target alignment and constraint satisfaction.
PaperID: 1867, Poster
Abstract: Gaussian processes are a cornerstone of modern probabilistic machine learning, providing principled uncertainty quantification and flexible nonparametric modeling. While the Gaussian kernel is widely used in Euclidean spaces, many real-world problems involve data residing on non-Euclidean domains, particularly compact Lie groups that naturally encode \emphsymmetries. In this work, we construct a local Gaussian kernel on arbitrary compact Lie groups by decomposing the group into a \emphmaximal torus and its orthogonal complement. For the prominent Lie groups \mathbfSO(3) and \mathbfSU(2), we derive closed-form kernel expressions using the Rodrigues formula. Extensive experiments verify the positive definiteness of the proposed kernel and demonstrate its effectiveness in Gaussian process regression, achieving accurate predictions across the entire manifold.
PaperID: 1868, Poster
Authors: Shivam Pal, Piyush Rai
Abstract: A post-hoc Bayesian procedure should not change its evidence or uncertainty estimates under a function-preserving reparameterization. Standard empirical-Bayes (EB) Laplace approximation violates this principle for ReLU networks. Positive homogeneity induces continuous scale orbits of equivalent parameter vectors, yet the standard isotropic objective depends on the chosen representative through the prior norm and curvature log-determinant term. Equivalent ReLU parameterizations can thus yield different marginal likelihoods, selected prior precisions, posterior covariances, and linearized predictive uncertainties. We fix this by introducing \gamma-canonicalization, selecting a canonical representative of each ReLU scale orbit and fitting Laplace there. This is equivalent to transporting the isotropic Gaussian prior from the canonical representative back to the original coordinates, producing a precision field that transforms compatibly with the curvature. For every \gamma\in[0,1], the generalized evidence and linearized predictive distribution are invariant under ReLU rescaling. The parameter \gamma indexes a family of orbit-consistent priors and is selected jointly with the prior precision by marginal likelihood, making EB Laplace a well-defined procedure on ReLU scale-equivalence classes. Empirically, we show that standard EB-Laplace varies substantially across functionally identical ReLU rescalings, while \gamma-canonicalized Laplace collapses this variation to zero.
Abstract: Quantile-based distributional reinforcement learning methods learn return distributions through sampled quantile regression, but their bootstrapped target quantiles may induce distorted or degenerate distribution estimates. We propose Robust Quantile-based Implicit Quantile Networks (RQIQN), a lightweight Wasserstein distributionally robust enhancement boosted from a quantile estimation perspective. We first reinterpret a snapshot of IQN loss as a collection of local empirical quantile estimation problems over sampled current fractions. We then robustify each local slot with a Wasserstein distributionally robust quantile estimation formulation, yielding a closed-form, fraction-dependent correction to the Bellman target. This correction directly mitigates distributional degeneration: its median-antisymmetry preserves the risk-neutral quantile average, while its monotonicity enlarges upper--lower quantile gaps and counteracts collapsed distributional spread. RQIQN thus regularizes quantile geometry without changing the underlying value objective or requiring additional sample-set reconstruction. Finally, we empirically show that the proposed RQIQN outperforms other existing quantile-based distributional reinforcement learning algorithms in risk-sensitive navigation and Atari games.
PaperID: 1870, Poster
Authors: Changseon Yu, Hyeokjun Kweon
Abstract: Training-free personalization adapts a frozen vision-language model (VLM) to recognize user-defined concepts from only a few exemplars by retrieving concept entries at inference time. While recent retriever-based methods enrich each concept with descriptive evidence, they largely treat concepts in isolation, so the retrieved information can be plausible yet non-discriminative, especially when the personalized concept set contains highly confusable, fine-grained categories. We argue that personalization is inherently relational: correct recognition depends on the specific distinctions between a concept and its nearest alternatives. To this end, we propose ReCoG, a Relational Concept Graph for retrieval-augmented personalization. ReCoG represents concepts as nodes and augments them with directed edges that explicitly capture discriminative cues between concept pairs, elicited by a frozen VLM during database construction. At inference, ReCoG retrieves a shortlist and resolves ambiguity via explicit pairwise comparisons conditioned on the corresponding relational edges, selecting the concept that consistently wins against competing candidates. We also introduce FINGER benchmark to test personalized recognition on confusable concept sets. Across FINGER, ReCoG substantially outperforms prior methods, with the largest gains in the most confusable regimes. ReCoG also achieves consistently best performance on established personalization benchmarks including captioning and VQA, confirming that relational evidence is broadly beneficial beyond fine-grained settings.
Authors: Alexander DeLise, Nick Dexter
Abstract: Generative compressed sensing uses the range of a pretrained generator as a nonlinear model for recovering structured signals from limited measurements. We study a conditional version of this problem for image recovery from subsampled Fourier measurements using prompt-conditioned generative models. Our framework separates two roles of conditioning: the prompt used to design the sampling distribution and the prompt used to define the recovery model. For ReLU and Lipschitz conditional generators, we prove stable recovery bounds showing that prompt-matched Christoffel sampling retains the same Christoffel complexity constant as existing near-optimal generative compressed sensing theory, while prompt mismatch incurs an explicit compatibility penalty. Experiments with Stable Diffusion show that prompts meaningfully reshape Christoffel sampling distributions and influence image recovery. Overall, our results suggest that prompts should be treated as design variables with distinct effects on sensing, approximation, and recovery.
PaperID: 1872, Poster
Authors:
Jimin Seo, Gyubok Lee, Yeonsik Jo, Kiwoong Yoo, Yeongoon Kim, Minhae Oh, Jin W Koo, Suhwan Kim, Nakyung Lee, Minsik Seol, Idris Nechnech, Jaehyeon Kim, LeeGiho, Jungwoo LeeAbstract: We study learning rate (LR) scaling for hybrid architectures combining Transformer and State-Space Model (SSM) blocks, a class adopted by several recent production language models. In particular, we focus on the gap between the theoretical scaling rules derived for SSMs under zero-order-hold (ZOH) discretization at infinite width with proportional state size, and the field-standard practical implementations using simplified-ZOH Mamba at fixed state size. Surprisingly, in this practical regime hybrid architectures achieve a near-zero LR transfer gap across widths 256-2048 and depths 4-32 at sub-billion-parameter scale using only the original \muP prescription, even though SSM operations fall outside its Tensor Programs representability conditions and every parameterization we test fails the standard coordinate-check diagnostic of \muP correctness. We attribute this to a two-condition decomposition of LR transfer in hybrid architectures: a global update-to-weight invariance, enforced by \muP's initialization and LR scaling; and a local per-component balance, provided by AdamW's g/\sqrtv per-parameter normalization. Our observations show that the optimal LR is invariant to width up to 8×, as well as to depth, sequence length, batch size, and Transformer-to-SSM ratio, and transfers to Nemotron-H, a production hybrid outside our custom architecture set. We hope these findings fill the gap between theoretical scaling rules and practical hybrid implementations, and stimulate further research toward bridging it.
Abstract: Conformal prediction (CP) provides a distribution-free approach to uncertainty quantification (UQ) with finite samples. However, applying CP to graph neural networks (GNNs) remains challenging. The combinatorial nature of graphs makes their encoding non-trivial, leading to insufficiently uncertain prediction logits and indiscriminative embeddings. Existing methods mostly rely on the embedding space for conformal prediction, which can be unreliable for graphs and often yields inefficient prediction sets. We propose GRAPHLCP, a structure-aware weighted conformal prediction framework that explicitly incorporates the graph topology and inter-node dependencies into localization and weighting. Based on experimental results that show a strong correlation between node homophily and sample coverage, we propose GRAPHLCP that utilizes a feature-aware densification followed by Personalized PageRank–based kernel computation to model structural proximity between nodes. Accounting for both structural and feature similarity allows GRAPHLCP to efficiently guarantee marginal coverage while attaining favorable test conditional coverage across extensive experiments on multiple datasets for both regression and classification.
Abstract: Virtual sensors replace expensive physical sensors in critical applications through machine learning by predicting target signals from available measurements. Existing virtual sensor approaches require application-specific models with hand-selected inputs for each sensor, cannot leverage task synergies, and lack consistent benchmarks. While emerging time series foundation models offer general-purpose, pretrained solutions in other domains, they are computationally expensive and limited to predicting their input signals, making them incompatible with virtual sensors. We introduce the first multi-task model for virtual sensors addressing both limitations. Our unified model can simultaneously predict diverse virtual sensors exploiting synergies while maintaining computational efficiency. It learns relevant input signals for each virtual sensor, eliminating expert knowledge requirements while adding explainability. In our large-scale evaluation on three standard benchmarks and an application-specific dataset with over 18 billion samples, our architecture reduces computation time by up to 415x and memory requirements by 951x, while maintaining or even improving predictive quality compared to unified baselines. Compared to existing isolated models for a single virtual sensor, our unified approach generates superior predictions at similar inference speed while scaling gracefully to hundreds of virtual sensors with nearly constant parameter count, enabling practical deployment in large-scale sensor networks.
PaperID: 1875, Poster
Abstract: Langevin Dynamics (LD) provides a principled framework for solving combinatorial optimization problems (COPs) via gradient-guided stochastic search. However, classical LD faces two bottlenecks: 1) the utilization of uniform initialization may affect its convergence speed; 2) when applied to COPs such as routing problems that cannot be naturally formulated as Quadratic Unconstrained Binary Optimization (QUBO), classical LD suffers from slow convergence and inferior performance compared to existing learning-to-construct (L2C) methods, as it requires manually designed discrete proposals for updates. These proposals are typically local and induce abrupt transitions in the solution space, hindering escape from local optima and degrading performance. To address these issues, we propose NLD4CO, a unified framework that synergizes data-driven learning with LD. It introduces two instantiations: explicit-gradient (EG) and implicit-gradient (IG). NLD-EG accelerates sampling efficiency for energy-based problems via neural warm-start initialization. NLD-IG utilizes the corrective direction predicted by a consistency model as an implicit score function to guide the search, making structured, globally coordinated transitions for general COPs. Extensive evaluations on MIS, MWIS, TSP, and CVRP demonstrate that NLD4CO achieves SOTA performance, delivering superior solution quality with competitive or lower runtime.
PaperID: 1876, Poster
Abstract: , a critical problem that emerges when distilling Flow Matching (FM) based behavior policies in offline Reinforcement Learning (RL): When learning over multi-modal datasets, standard distillation inadvertently forces the policy to interpolate across different behavioral modes, causing the generated actions to collapse into invalid, OOD regions. To resolve this, we propose , a novel framework that reformulates offline policy distillation as a value-guided generative trajectory alignment problem. By employing a dynamic coupling mechanism, VRD actively rectifies flow dynamics to explicitly decouple distinct behavioral modes, cleanly isolating high-value actions. Specifically, we instantiate VRD through two complementary algorithms: , which utilizes value-quantile behavior distillation to statistically filter out low-value behaviors, and , which reorganizes the latent space to geometrically isolate high-value generative trajectories from the FM-based behavior policy. Extensive empirical evaluations demonstrate that VRD improves existing generative baselines across diverse continuous control and planning tasks on the D4RL, OGBench, and Minari benchmarks, especially over those multi-modal, non-expert datasets.
PaperID: 1877, Poster
Abstract: Range-partition entropy is a structural complexity measure used in entropy-sensitive computational geometry. We introduce Differentiable Entropy Regularization (DER), a gradient-optimizable surrogate for the entropy of learned geometric partitions. The paper makes three main contributions. First, we separate the normalized entropy optimized by DER, denoted H work(S) = n · H_norm(S) that appears in algorithmic work bounds. Second, for a fixed margin-separated halfspace arrangement, we prove a self-contained soft-to-hard approximation theorem: the soft halfspace-cell entropy converges to the corresponding hard sign-pattern entropy at a rate controlled by margin, temperature, and the number of separators. A connection to optimal range-partition entropy holds when the fixed partition is admissible for the geometric problem and near-optimal in the admissible partition class. Third, we evaluate DER as learned geometric preprocessing for SciPy/Qhull-based convex-hull and Delaunay pipelines. On the evaluated 2D tasks, DER gives up to 4×–5× solver-side speedups and smaller end-to-end gains after preprocessing, with convex-hull area error below 0.2%. We also include a secondary transfer study in ViT-style attention, where the same regularizer induces structured sparsity and improves empirical efficiency. The central contribution is not the entropy formula alone, which resembles soft clustering, but the range-aware differentiable construction, the explicit approximation analysis, and the geometry-first training pipeline.
Abstract: We reformulate Optimal Transport Conditional Flow Matching (OT-CFM), showing that it admits an exact proximal form via an extended Brenier potential, without assuming that the target distribution has a density. In particular, the mapping to recover the target point is expressed by a proximal operator, which yields an explicit proximal expression of the vector field. We also discuss the convergence of minibatch OT-CFM to the population OT formulation as the sample and batch sizes increase. Using second epi-derivatives of convex potentials, we prove that, for manifold-supported targets, the manifold structure is stable by perturbation to the dynamics: after time rescaling, the dynamics contracts exponentially in directions normal to the manifold while remaining neutral along tangential directions.
PaperID: 1879, Poster
Abstract: Acquiring diverse whole-body dexterous manipulation data for humanoids remains a fundamental challenge in character animation and robotics. Existing pipelines rely on expensive mocap data collected from many subjects and require retargeting to a unified humanoid shape, which limits scalable data construction and often yields physically inconsistent interactions (e.g., unstable contacts). We present ManipSynth, a unified framework for whole-body loco-manipulation synthesis, with compositional generation of sparse whole-body motion, object-centric grasp anchors, and contact-aware interaction transitions. Then, we scale data generation through randomized pre-grasp motions, grasp poses, and object trajectories. ManipSynth achieves approximately 5× higher contact coverage and lower average penetration than OMOMO across representative objects. To test whether better contact-consistent data improve scalable control learning, we train ManipTracker, a general-purpose Humanoid-Object Interaction (HOI) tracking policy, on purely synthetic demonstrations. Under the same object set, number of training trajectories, the same ManipTracker trained on ManipSynth demonstrates stronger generalization than when trained on mocap data, while achieving over 90% training success.
PaperID: 1880, Poster
Abstract: What if you could edit an image simply by drawing on it---circling an object, writing a short label, or sketching an arrow---with minimal or even no typed prompt? We first reveal a surprising capability: when a text-driven image editor is paired with a strong Vision-Language Model (VLM) at inference time, the combined system can directly interpret visual instructions embedded in the image itself, such as on-image text, bounding boxes, and directional arrows. However, this pipeline incurs additional VLM-planning latency, introduces brittle cross-model error propagation, and depends on external APIs or auxiliary model weights. To address these limitations, we introduce Siphon, a framework that internalizes this externally elicited visual-instruction-following competence into a single diffusion-based editor. Rather than relying on an external VLM at inference time, Siphon first uses a VLM planner to synthesize paired visual-instruction supervision, and then transfers this competence into the editor through lightweight LoRA fine-tuning. The resulting model reads, grounds, and executes on-image annotations directly from pixels, while keeping the diffusion sampling cost close to that of the base editor. Across multiple diffusion architectures, Siphon substantially improves spatial controllability, instruction adherence, and marker removal over text-driven baselines, while matching VLM-assisted pipelines without their additional planning latency, API dependence, and cross-model fragility. We do not position rendered long text as a replacement for conventional text prompts; rather, on-image text is mainly intended for short labels and spatially grounded edit intents, while long or complex instructions can still be provided through the standard text channel.
PaperID: 1881, Poster
Abstract: Modern data-driven sciences use large observational datasets such as online discussions, behavioral logs, web experiments and biological records to explain outcomes and design interventions. Existing LLM-based hypothesis generators produce interpretable claims, but usually validate them through predictive accuracy, making them vulnerable to spurious correlations and dataset artifacts. Automated validation systems test hypotheses statistically, but require researcher-specified constructs. Here we show that , a two-agent closed-loop framework, discovers statistically supported natural-language hypotheses from observational data. A menter operationalizes features, selects covariates and tests, executes analyses and returns evidence for refinement. Accepted and rejected hypotheses are stored in a labeled memory that guides further exploration. Across 19 tasks spanning text, images, clinical tabular data, economics, marketing and single-cell RNA-seq, ExperiGen discovers 2–4× more statistically significant hypotheses than prior generation methods and improves downstream predictive performance by 4–21 points. Expert evaluation by 26 specialists rated its hypotheses as novel, clear and research-worthy; 11 of 16 preregistered r/ChangeMyView hypotheses were significant in the predicted direction; and in deployed A/B tests across seven S&P 500 companies and 5.1 million users, 15 of 25 hypotheses achieved significance, with 14 matching the predicted direction. These results suggest that closed-loop hypothesis generation can help prioritize hypotheses for expert review and real-world intervention and augment human judgment in data-driven sciences.
Authors: Drago Plecko
Abstract: Automated systems built on artificial intelligence (AI) are increasingly deployed across high-stakes domains, raising critical concerns about fairness and the perpetuation of demographic disparities that exist in the world. In this context, causal inference provides a principled framework for reasoning about fairness, as it links observed disparities to underlying mechanisms and aligns naturally with human intuition and legal notions of discrimination. Prior work on causal fairness primarily focuses on the standard machine learning setting, where a decision-maker constructs a single predictive mechanism f_\widehat Y for an outcome variable Y, while inheriting the causal mechanisms of all other covariates from the real world. The generative AI setting, however, is markedly more complex: generative models can sample from arbitrary conditionals over any set of variables, implicitly constructing their own beliefs about all causal mechanisms rather than learning a single predictive function. This fundamental difference requires new developments in causal fairness methodology. We formalize the problem of causal fairness in generative AI and unify it with the standard ML setting under a common theoretical framework. We then derive new causal decomposition results that enable granular quantification of fairness impacts along both (a) different causal pathways and (b) the replacement of real-world mechanisms by the generative model's mechanisms. We establish identification conditions and introduce efficient estimators for causal quantities of interest, and demonstrate the value of our methodology by analyzing race and gender bias in large language models across different datasets.
PaperID: 1883, Poster
Abstract: Rotating a molecule should leave its energy unchanged while rotating each force vector by the same rotation, as required in SO(3)-equivariant molecular learning. However, conventional normalization layers in neural networks are designed for scalar activations; when applied to vector channels through component-wise centering or scaling, they can break this symmetry. We study normalization modules for equivariant molecular networks under force supervision, where invariant scalar channels and covariant vector channels require different geometric treatment. We formalize an SO(3)-equivariance-preserving recipe that computes vector scale statistics from rotation-invariant squared norms and applies the resulting rescaling identically across Cartesian coordinates, which preserves SO(3)-equivariance without vector shifts. Building on this recipe, we introduce grouped vector RMS (GroupRMS), which shares vector RMS estimates across channel groups to reduce estimator variance and stabilize vector scaling in small-batch training. The proposed variant is lightweight and compatible with standard scalar normalization branches. We instantiate these ideas in practical scalar-vector normalization blocks for PaiNN-style backbones and evaluate them across several molecular benchmarks. On the paired three-trajectory subset, the strongest GroupRMS-family variant, GroupNorm-GroupRMS, improves average force MAE from 0.1190 to 0.0383 and average energy MAE from 0.0675 to 0.0424 relative to BatchNorm. These results support the proposed rotation-compatible vector scaling and grouped scale estimation as effective tools for robust molecular force learning.
PaperID: 1884, Poster
Abstract: Continual recommendation captures dynamic user interests amid sequential interactions, enabling models to adapt to new observations while preserving prior valuable knowledge. Existing continual recommendation methods primarily preserve or reuse historical knowledge through replay, distillation, or regularization to mitigate forgetting during incremental updates. However, noisy or obsolete inherited representations render indiscriminate preservation detrimental by inducing memory conflicts with current stage preferences. Moreover, rapid shifts in user interests require timely and selective memory updates, while coarse global mechanisms fail to provide such fine-grained responsiveness. To address these issues, we propose VAMIRec, a value-aware memory intervention framework for continual recommendation. VAMIRec first constructs a memory state by separating inherited user memory into a slow anchor component and a fast-changing drift component. We further design a value-aware action evaluation module to generate three candidate operations, and estimate conservative action values by accounting for uncertainty and intervention cost. Finally, we develop a policy-guided memory intervention module that distills the training-time action supervision into an inference-safe policy and applies the selected memory action before all-item ranking. Extensive experiments on three real-world datasets demonstrate that VAMIRec consistently outperforms the state-of-the-art baselines.
PaperID: 1885, Poster
Abstract: Sequential recommenders need long histories, but full-prefix attention scales with user lifetime. We argue that an effective memory mechanism should provide space invariance, semantic fidelity, and lifelong evolvability, which existing fixed-window or capped-memory methods do not jointly satisfy. We propose RecMem, a recurrent compression framework that folds history segments into M forget-gated memory slots. RecMem has a history-length-independent rollout error bound, and its fixed-capacity memory is trained by next-item prediction to preserve useful signals. On MerRec, RecMem uses 5.6x fewer decoder-side tokens than full attention while retaining 88–97% of Recall@10–200. Compared with a compressed truncation baseline at similar decoder cost, RecMem improves Recall@50+ by 4–6 percentage points by compressing the full history rather than discarding old interactions. It also remains stable over 1,800 test steps.
PaperID: 1886, Poster
Abstract: Concept-based explanations in neural models tie their outputs to human-meaningful concepts that users can understand, troubleshoot, and trust. But these explanations are typically correlational rather than causal: a concept can fire alongside an output without driving it, yielding interpretability that looks right but misleads. We introduce Causal Concept Wrapper Network (CCW-Net), the first use of mediation analysis as a training objective in deep networks. We further apply mediation analysis as an evaluation tool to quantify, per sample, how much of any model's output is causally attributable to each concept. As a training objective, CCW-Net shapes a model's concept embeddings such that each concept's contribution to the output is both locally necessary and sufficient for the share it is credited with, and independent of contributions from other concepts. Counterfactual concept samples drawn from a learned per-concept flow anchor a set of causal alignment criteria: necessity drives the effect of every irrelevant concept toward zero; sufficiency requires relevant concepts to fully reconstruct the model's output; and independence enforces that the total effect decomposes linearly across concepts. We evaluate CCW-Net across aircraft control, driving, and fine-grained image classification. The result is a step toward concept-based explanations that are not merely interpretable, but quantifiably causal, supporting safer deployment, evaluation, and certification of transparent and trustworthy neural systems.
PaperID: 1887, Poster
Abstract: Capability evaluations aim to surface dangerous model behaviors before deployment, but their reliability depends on the assumption that hidden capabilities can be elicited. We show this assumption does not hold under adversarial conditions by constructing backdoor attacks that embed encrypted neural networks within host model weights, decrypting and executing them only upon receiving a secret trigger. Using digital lockers, these hidden circuits remain provably hard to elicit and interpret under standard cryptographic hardness assumptions, even given full access to model weights, extending prior work from hiding fixed strings to arbitrary computations. We implement our constructions in PyTorch and empirically validate resistance to supervised fine-tuning on a small chemical reaction prediction task - a proxy for CBRN-relevant capabilities. Our results reveal a limit of capability audits. We release our models to support the development of stronger defenses.
Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for generalist robotic control. Built upon vision–language model (VLM) architectures, VLAs predict actions conditioned on visual observations and language instructions, achieving strong performance and generalization across tasks. However, VLAs face two major challenges: a limited context window for input frames and inefficient inference due to the quadratic attention complexity and large parameter counts. To this end, we propose Dysta, a framework that disentangles visual inputs into multi-level static and dynamic tokens, which enables (1) retaining a single copy of static tokens across frames to significantly reduce context length, and (2) reusing the key–value (KV) cache of static tokens through a lightweight recache gate that updates only when necessary. This design enables efficient multi-frame integration and efficient inference. In addition, we introduce a new benchmark that more effectively evaluates the multi-frame integration ability of VLAs. Experiments show that DySta improves multi-frame integration by 24.5% across metrics on our benchmark and 23.3% in absolute success rate on real-world memory-dependent tasks, while accelerating inference by 2.0× (with +2.3% success rate) on simulation benchmarks and 2.2× (with +10.6% success rate) on real-world general tasks.
PaperID: 1889, Poster
Abstract: Conformal prediction provides model-agnostic uncertainty quantification with guaranteed coverage, but conventional methods often yield overly conservative uncertainty sets, particularly in multimodal or heterogeneous settings. This inefficiency arises from two sources: (i) limited expressiveness of the predictive model and (ii) simplistic nonconformity scores design. Most existing approaches advance only one of these axes, leaving the other underexplored. We propose generative conformal prediction with llocation (ORCA), a three-stage framework that advances both aspects jointly. ORCA leverages generative models to capture the full conditional distribution and introduces a rank-dependent optimization procedure that adaptively allocates coverage for efficiency while maintaining validity. We cast this coverage allocation as an optimization problem, derive an exact mixed-integer linear programming formulation, and show that the solution converges asymptotically to the oracle density-level set. Across synthetic, semi-synthetic, and real datasets, ORCA produces substantially more efficient uncertainty sets than state-of-the-art baselines, demonstrating robust gains in scenarios where conventional conformal prediction methods fail.
PaperID: 1890, Poster
Authors: Tetsuro Morimura
Abstract: We study semi-offline value estimation, where a small on-policy sample supplements a large off-policy dataset. This regime is common in recommender systems and pre-deployment A/B testing, and the challenge is to use the limited on-policy data to correct systematic bias in a given value predictor. Recent work on Bellman calibration (BC) gives a post-hoc, model-agnostic correction along the predicted value. We propose Error-Guided Bellman Calibration (GBC), which trains an error model on the on-policy sample and calibrates the predictor along this learned error axis. We prove a completeness-free refinement guarantee, a theorem showing that the correction adapts to the intrinsic dimension of the prediction bias, and a finite-sample bound on the value-estimation error that isolates our advantage in the refinement term. Across synthetic, CRM, and D4RL benchmarks, GBC improves over existing calibration baselines.
Abstract: Denoising and score estimation are classically linked through Tweedie’s formula, which relates the posterior mean under Gaussian corruption to the Stein score of the noisy marginal. In this work, we extend this perspective beyond Gaussian noise to a broad class of elliptical, energy-based noise distributions, with particular emphasis on generalized Gaussian corruptions. We derive the Energy–Tweedie identity: when the denoising posterior is viewed through the lens of proper scoring rules, the path derivative of a matched, possibly non-Euclidean energy score recovers the Stein score of the noisy marginal. Thus, the familiar correspondence between Gaussian noise, posterior means, squared loss, and Tweedie’s formula is lifted to a distributional correspondence between generalized Gaussian noise, full posterior laws, Mahalanobis energy scores, and the Energy–Tweedie identity. Among its consequences, this identity gives a posterior-sample-based route to score estimation, yields a principled criterion for noise-parameter calibration, and supplies a score-based perspective on recent diffusion-style generative methods trained with scoring rules.
Authors:
Federico Echenique, Alireza Fallah, Baihe Huang, Michael JordanAbstract: Aligning large language models (LLMs) to human preferences typically relies on aggregating pooled feedback into a single reward model. However, this standard approach assumes that all labelers share the same underlying preferences, ignoring the fact that real-world labelers are highly heterogeneous and usually anonymous. Consequently, relying solely on binary choice data fundamentally distorts the learned policy, making the true population-average preference unidentifiable. To overcome this critical limitation, we demonstrate that augmenting preference datasets with a simple, secondary signal—the user's response time—can restore the identifiability of the population's average preference. By modeling each decision as a Drift-Diffusion Model (DDM), we introduce a novel, consistent estimator of heterogeneous preferences that successfully corrects the distortions of standard choice-only labels. We prove that our estimator asymptotically converges to the true average preference even in extreme cases where each anonymous labeler contributes only a single choice. Empirically, across both synthetic and real-world datasets, our method consistently outperforms standard baselines that otherwise fail and plateau at a bias floor. Because response times are essentially free to record and require zero user tracking or identification, our results bring promises and open up new opportunities for future data-collection pipelines to improve the social benefit without requiring user-level identifiers or repeated elicitations.
PaperID: 1893, Poster
Abstract: The Bayesian learning rule (BLR) exploits the relation between Bayesian updating and the Kullback-Leibler divergence (KLD) to derive efficient additive updates for exponential-family variational posteriors. The KL geometry is powerful but restrictive: combined with posterior families with parameter-dependent support it can lead to undefined objectives. We ask which non-KL geometries preserve the additive BLR computation. Previous works that use f-divergences suggest that this is only possible with KLD. Here, we show that this is not true. We show that additive information-geometric optimization algorithms can be derived for the class of mixed-space Bregman or Fenchel-Young discrepancies. These discrepancies induce generalized maximum entropy families, dual coordinates, and generalized BLR canonical gradient updates, recovering the classical BLR. The generalized BLR leads to new algorithms that cannot be derived from the KL-based BLR. For the Tsallis negentropy with quadratic statistics, the generalized BLR yields bounded-support and heavy-tailed q-Gaussian posteriors and an easy Newton-like implementation. Experiments up to CIFAR-100 illustrate controllable support and tail behavior and competitive performance.
PaperID: 1894, Poster
Abstract: Selective state-space models such as Mamba replace fixed linear state transitions by input-dependent transitions, yielding recurrent sequence models that retain linear-time inference while adapting their memory dynamics to the current token. We show that even a simple form of this selective memory update has universal approximation power. Specifically, for every compact input domain, finite horizon, and continuous causal sequence-to-sequence map, a clocked diagonal selective state-space scan with strictly contractive gates and a linear-in-state readout uniformly approximates the target map. The proof is constructive: input-dependent diagonal propagation generates tensor-product partition features over prefixes, and the linear-in-state readout combines these features with sampled target values. Thus the high-order interactions required for universality are produced inside the selective scan itself, rather than by attention, nonlinear hidden-state dynamics, or a universal multilayer perceptron applied after a sequence encoder. These results identify selectivity as a mathematically sufficient mechanism for expressive linear-time state-space sequence modeling. Selectivity lets the model decide, token by token, which parts of the past should be preserved, faded, or combined with the present. The universal approximation theorem shows that this adaptive use of memory is sufficient to reproduce any well-behaved causal sequence rule over a fixed horizon.
Abstract: Deep neural networks (DNNs) have achieved remarkable success in practice, yet a mechanistic understanding of how features evolve during training remains incomplete, especially in the large-depth limit. For ResNets under depth-\muP scaling, previous studies introduce an SDE view of training dynamics by treating the layer index \ell as a continuous-time variable t_\ell=\ell/L. A key unresolved issue is that backpropagation reuses each forward weight matrix W_\ell through its transpose W_\ell^\top, creating correlations between forward features and backward gradients whose behavior and role in training-time feature learning remain unclear. We study this reused-weight forward--backward coupling in one-layer ResNets under depth-\muP scaling. Using conditional Gaussian representations, we explicitly separate the coupling terms induced by weight reuse from decoupled Gaussian fluctuations. At initialization, we prove that the coupling is a finite-width effect and vanishes at rate O(n^-1), uniformly over depth. During training, however, SGD induces a nontrivial forward--backward correlation term that survives the infinite-width limit. The key depth effect is that, under depth-\muP scaling, this surviving term is higher order in depth and its accumulated contribution over layers becomes negligible as L\to\infty. This depth-induced suppression motivates neural feature dynamics (NFD), a forward--backward SDE system with decoupled backward weights that retains the feature-gradient covariance structure generated during training. Under nondegeneracy assumptions, we prove that the finite-network training dynamics converge to NFD with an O(L^-1) depth-discretization error, while the reused-weight coupling term has a faster O(L^-2) decay. These results provide a rigorous infinite-depth feature-learning limit for one-layer ResNets under depth-\muP scaling.
PaperID: 1896, Poster
Abstract: Flow matching and diffusion models achieve high-quality generation by solving continuous-time generative dynamics, but their sampling remains inherently sequential and costly. Parallel sampling reduces wall-clock latency, yet often requires many function evaluations (NFEs) due to inaccurate future-state initialization and repeated correction. We propose ParaPC-FM, a fixed-point predictor--corrector framework that unifies high-order solver correction and future-state initialization, enabling both components to be incorporated into different parallel samplers. In particular, our UniP-based truncated initialization reuses past velocity predictions to estimate future states more accurately, reducing correction iterations and NFEs with almost no additional overhead. On a ParaDiGMS-style framework, ParaPC-FM achieves up to 1.25× and 1.49× speedups on FLUX.1-dev at 50 and 100 steps, and up to 1.22× and 1.16× speedups on Stable Diffusion 3 and Stable Diffusion 3.5 at 50 steps. When integrated into ParaTAA, the proposed initialization further achieves up to 1.17× speedup at 50 steps. Importantly, these acceleration gains are obtained while keeping image quality nearly unchanged.
Abstract: Parameter-efficient fine-tuning methods, such as LoRA, offer a practical way to adapt large vision and language models to client tasks. However, this becomes particularly challenging under task-level heterogeneity in federated deployments. In this regime, personalization requires balancing general knowledge with personalized knowledge, yet existing approaches largely rely on heuristic mixing rules and lack theoretical justification. Moreover, prior model merging approaches are also computation and communication intensive, making the process inefficient in federated settings. In this work, we propose Potara, a principled framework for federated personalization that constructs a personalized model for each client by merging two complementary models: (i) a federated model capturing general knowledge, and (ii) a local model capturing personalized knowledge. Through the construct of linear mode connectivity, we show that the expected task loss admits a variance trace upper bound, whose minimization yields closed-form optimal mixing weights that guarantee a tighter bound for the merged model than for either the federated or local model alone. Experiments on vision and language benchmarks show that Potara consistently improves personalization while reducing communication, leading to a strong performance-communication trade-off.
Abstract: Diffusion language models intrinsically fail to capture correlations between decoded tokens, which leads to a harsh trade-off between sampling quality and throughput. To solve this issue, we propose DiLaDiff, a variant of masked diffusion language models with three components: (1) a continuous latent space with semantic capabilities, learned by an auto-encoder fine-tuned from an existing masked diffusion language model; (2) a latent diffusion model learning the prior over the encoder distribution; (3) a consistency model distilling the learned prior into a few-step latent generative model. We show that, even without distillation, our latent-guided diffusion model outperforms the masked diffusion baseline while significantly accelerating inference. Consistency distillation further lowers the computational overhead of continuous diffusion, such that the latent is generated in negligible time compared to discrete decoding.
PaperID: 1899, Poster
Abstract: Cross-camera RAW-to-sRGB mapping seeks to convert smartphone RAW captures into DSLR-quality sRGB images, yet the large domain gap between linear RAW and nonlinear sRGB spaces makes this task highly challenging. Existing methods mainly condition on pixel-level supervision and local RAW cues, where content preservation and domain translation remain entangled in low-level measurements. We propose SemISP, a semantic-guided diffusion framework that uses scene semantics as a domain-invariant anchor for cross-domain generation. Specifically, we adopt a frozen DINOv3 as the sRGB semantic backbone and develop a Domain Alignment Adapter to extract sRGB-consistent semantic representations from RAW inputs through parameter-efficient adaptation. Because semantic priors and RAW signals operate at fundamentally different information levels, and the diffusion denoising process follows a coarse-to-fine trajectory, we further design a Time-aware Dual-stream Adaptive Modulation (TDAM) module that separately encodes both condition streams and uses the diffusion timestep to dynamically balance their contributions---emphasizing semantic anchoring at early stages and fine-grained RAW details at later ones. Experiments on two established cross-camera benchmarks show consistent improvements over prior methods in both fidelity and perceptual quality.
Abstract: We study the Compositional Geometry Routing Problem (CGRP), a unified superclass of traditional routing problems that covers point-only, line-only, area-only, and arbitrary hybrid task geometries, providing a broad abstraction for real-world routing scenarios. Beyond standard point-based routing, CGRP with non-point tasks can be inherently asymmetric, tightly coupled travel routes with the intrinsic path, and enlarges the action space with numerous feasible yet often irrelevant options, thereby posing significant challenges for both representation learning and decision-making. To address these challenges, we propose DiCon, a differential attention–assisted solver with contrastive learning, as a plug-and-play framework that tackles the problem from two complementary angles. First, we introduce a differential attention mechanism that actively suppresses the probability mass on less competitive candidate actions. Second, we design a double-level contrastive learning objective to promote robust global instance representations and regularize geometry-aware task representations. Extensive experiments demonstrate that DiCon achieves strong performance, broad versatility, and superior generalization across diverse CGRP instances with different compositions.
PaperID: 1901, Poster
Abstract: Articulated objects are fundamental for robotics, simulation of physics, and interactive virtual environments. However, recovering them from visual observations is inherently challenging, as images provide only partial and ambiguous cues about both part geometry and their underlying kinematic structure. Existing approaches typically rely on multi-stage pipelines, retrieval from asset libraries, or explicit part segmentation. We present URDF-Anything+, an end-to-end autoregressive diffusion framework that generates simulation-ready URDF models directly from a single RGB image. Conditioned on visual observations and object geometry, URDF-Anything+ operates in a structured latent space and jointly models part geometry and articulation in a unified generation process. Specifically, the model sequentially predicts each articulated part together with its associated joint parameters, while a termination token dynamically determines the number of parts. This design enables direct generation of fully executable URDFs without external retrieval or post-processing stages. Experiments on large-scale articulated object benchmarks demonstrate that URDF-Anything+ outperforms prior methods in geometric reconstruction quality, joint parameter estimation, and physical executability, while being substantially more efficient than existing multi-stage approaches. Furthermore, the generated URDFs serve as faithful digital twins, enabling the zero-shot transfer of manipulation policies trained purely in simulation.
Abstract: Many physical systems exhibit a low-dimensional structure that varies with external parameters: link lengths in a robot, forcing constants in a fluid, or Reynolds numbers in a flow shift the underlying manifold while preserving its intrinsic dimension. Constrained AutoEncoders (cAEs) learn such manifolds through an encoder-decoder projection, a property that unconstrained autoencoders cannot match and that is essential whenever the model is applied iteratively. However, the standard strategies for making a cAE context-dependent, namely concatenating the context to the input or affinely modulating hidden activations, break the encoder-decoder idempotency, sacrificing the projection guarantee precisely in the setting where it would be most valuable. To restore this guarantee under context variation, we developed the Neuromodulated Constrained Autoencoder (NcAE), which modulates the activation slope and bias of a cAE through a context-driven hyper-network. This paper presents the NcAE, its theoretical foundation, and its empirical validation. We prove that for every context, including contexts unseen at training time, the reconstruction map remains an idempotent projection, the topology of the learned manifold is invariant, and context perturbations induce smooth changes in the manifold. We evaluated our approach on a 16-DoF pendulum with context dependent coupling and the Lorenz96 system across a bifurcation. The NcAE matched or exceeded the best of six baselines on reconstruction, idempotency, and latent-geometry metrics, while being the only architecture that preserves geometric consistency by construction. The NcAE thereby provides a stable, geometry-preserving coordinate system across families of physical regimes.
PaperID: 1903, Poster
Authors:
Jia Gao, Xianglei Xing, yang qing, Tianshuo Zhang, Wenzhe ZhaiAbstract: State estimation is a critical challenge in modeling dynamic systems. Traditional Kalman filters and data-driven methods have shown good results in complex, nonlinear systems. However, most existing methods optimize the posterior without considering prior modeling, which leaves room for improvement in state estimation accuracy, and predict high-dimensional Kalman parameters end-to-end, which lead to unstable training and poor convergence. To address this, we propose Diff-Kalman, which leverages state differences to predict state increments and scaling matrices for correcting both the prediction and update phases of the Kalman filter. Diff-Kalman combines traditional dynamic models with neural networks to adaptively correct system behavior. In the prediction phase, Diff-Kalman utilizes a Transformer-based architecture with self-attention to extract localized patterns from historical states, followed by cross-attention to predict future state increments. These increments are then fused with model-based priors from the dynamic model. In the update phase, a GRU-based Scaling Predictor (GSP) dynamically adjusts the Kalman covariance and gain matrices based on differential residuals, which improves the accuracy and robustness of the estimation. Experimental results on linear MOT trajectories and nonlinear Lorenz dynamics show that Diff-Kalman improves estimation accuracy over representative baselines.
PaperID: 1904, Poster
Abstract: Unlearnable examples (UEs) aim to prevent machine learning models from extracting useful information in protected training data. Extending UEs to segmentation is practically important but challenging, because pixel-level annotations are expensive to obtain and therefore cannot be assumed available when generating protective perturbations. This motivates us to study unsupervised unlearnable segmentation, where UE perturbations must be generated without segmentation annotations. Since dense prediction fundamentally relies on region-level affinity and cross-layer feature correspondence, effective segmentation data protection faces a twofold design challenge: perturbations should induce misleading alternatives to these structural priors, while the induced structure must remain easy to fit so that downstream segmentation models can adopt it during training. To address these challenges, we propose ructure) an unsupervised method that generates unlearnable examples for segmentation. Specifically, DRIFT constructs pseudo region-level affinity and pseudo cross-layer correspondence from clean-image features. It then optimizes perturbations to steer the extracted features of protected images toward the misleading pseudo-structure rather than the clean one, while an auxiliary structure learner, trained to predict pseudo-region assignments, explicitly encourages the induced structure to remain easy to fit. Extensive experiments across four segmentation datasets and multiple architectures demonstrate that DRIFT provides strong protection, consistently degrading the generalization performance of models trained on protected data and remaining competitive with the supervised UE baseline.
Authors:
Xinyu Qiao, yichen lin, Kaihong Ji, xue wang, Tao YaoAbstract: Online conformal prediction can fail when predictions shape actions and actions determine which outcomes enter calibration. Standard adaptive methods may retain marginal coverage while systematically miscovering the counterfactual outcomes of rarely selected actions. This paper formalizes the failure through counterfactual coverage and introduces Propensity-Weighted Online Conformal Prediction, an inverse-propensity-weighted recursion that debiases calibration. A doubly robust variant further reduces nuisance bias to the product of outcome-model and propensity errors. Under positivity, the resulting coverage rate matches an information-theoretic lower bound up to logarithmic factors. Experiments on synthetic decision tasks, open bandit data, and financial rebalancing show that PW-OCP and DR-OCP improve counterfactual coverage and downstream regret without sacrificing prediction-set sharpness.
Authors:
Xiaohan Zhang, Feng Gu, Xudong Rao, Xuhao Pan, Tao Wei, Zhou Pan, Kun ZhanAbstract: Spatial intelligence requires foundation models to maintain coherent spatial state across interactions with the physical world. However, existing data-centric approaches typically treat spatial reasoning as independent question-answer instances, enabling shortcut-based answering and providing limited supervision for persistent spatial understanding. To address this, we introduce ChainSpace, a chained-reasoning paradigm that structures spatial reasoning as a state-preserving multi-round process. In this paradigm, spatial questions are organized into logically constrained and jointly consistent chains, where later questions depend on spatial constraints established in earlier rounds. Following this principle, we instantiate ChainSpace-Bench, a manually annotated real-world multi-round benchmark with a Chain-Aware Metric, and ChainSpace-Pipeline, a simulator-based chain-structured supervision generation framework for spatial intelligence training. Experiments show that ChainSpace-Bench exposes chain-level failures that are not captured by isolated question accuracy. Additionally, with a relatively small amount of simulator-generated chained data, models trained by ChainSpace-Pipeline achieve the best performance among open-source models on ChainSpace-Bench and transfer competitively to multiple external spatial intelligence benchmarks. These results establish ChainSpace as an effective paradigm for more faithful evaluation and more data-efficient learning of spatial intelligence.
Authors:
Xiao Han, Hao Liu, Zhimin Bao, Jile Jiao, Yue Wang, Hui Guo, Mou X Feng, Yi XuAbstract: Chest X-ray visual question answering (CXR VQA) requires models not only to predict correct answers, but also to produce reliable medical reasoning. However, existing reinforcement-learning-based training typically relies on answer-level rewards, which are often too coarse to improve chain-of-thought (CoT) quality and can become ineffective when group-level advantages collapse to zero. We propose Teach-to-Reason (T2R), a framework that introduces comparison-based supervision into CoT optimization through a self-improving \emphTeacher and a competition-guided \emphReasoner. As the Teacher is iteratively strengthened via self-competition, the Reasoner is optimized against progressively stronger Teacher-generated references. We further introduce a case-wise reward design that preserves the original reward-induced positive/negative partition when it is informative, and restores supervision from competition scores when the original reward signal degenerates. Experiments on multiple CXR open-ended VQA benchmarks show that T2R consistently outperforms strong baselines, indicating that comparison-based supervision, when integrated in a controlled and principled manner, provides a more effective training signal for reasoning optimization.
PaperID: 1908, Poster
Abstract: LLMs have shown strong in-context learning (ICL) capability, but extending it to Multimodal LLMs (MLLMs) remains challenging. Prior multimodal ICL methods often rely on ICL-specific datasets for additional training or task vector extraction, which can improve performance on the same ICL benchmark used for adaptation. However, they do not generalize well to other benchmarks and often induce forgetting in previously well-solved tasks. For this reason, we identify `visual grounding' as the key bottleneck in multimodal ICL; MLLMs often fail to attend to task-relevant visual evidence in demonstrations, instead over-relying on textual cues and producing hallucinated outputs. To address this, we propose ProCoRe to enhance visual grounding through contrastive reinforcement learning on multimodal contrastive data for improving multimodal ICL. ProCoRe generates captions for contrastive image pairs, decomposes them into propositions, and optimizes contrastive rewards that encourage alignment with corresponding images while discouraging mismatched ones. Despite never seeing ICL-formatted training examples, ProCoRe improves average multimodal ICL classification accuracy by 5.6% and captioning ROUGE-L by 13.8% over the off-the-shelf Qwen-VL-3-8B model, outperforming ICL-data-dependent state of the arts across benchmarks. Finally, we introduce PICL, a personalized multimodal ICL captioning benchmark that tests whether MLLMs genuinely understand visual demonstrations rather than relying on pre-trained semantic priors.
Abstract: Video captioning requires fine-grained spatio-temporal understanding of videos, including spatial perception of where objects are located and temporal perception of when events occur. Existing Multi-modal Large Language Models (MLLMs) usually generate captions directly from video inputs without exposing the perceptual evidence behind their descriptions. As a result, object, event, or temporal mistakes are only observed in the final text, making it difficult to identify the underlying perceptual errors or optimize them directly. To address these issues, we present PercepCap, a perception-aware video captioning framework that makes perceptual evidence explicit before producing the final caption. Specifically, PercepCap follows a perceive-describe generation chain, where the model first produces a spatio-temporal perception trace comprising object trajectories and temporal events according to the video, and then generates the final caption conditioned on the perceived evidence. To support this new generation paradigm, we design a two-stage training strategy for PercepCap. Perceive-then-Describe Supervised Fine-tuning (PD-SFT) adapts the model from caption-only generation to the proposed perceive-describe chain, while Perception-Grounded Reinforcement Learning (PG-RL) further optimizes perception trace and caption quality with joint rewards over object tracking, temporal event, and object/action description coverage. Caption-Anchored Perception Data Construction builds this supervision by first generating a caption-only description, extracting the objects and events it mentions, and grounding them back in the video with boxes and timestamps. This yields caption-aligned perception data that provides both SFT supervision and RL references, ensuring that the explicit perception trace and final caption describe the same objects and events. Across direct caption evaluation such as DREAM-1K, CaReBench and VidCapBench and caption-to-QA evaluation like ShortVidBench and MotionBench. PercepCap consistently improves the Qwen3-VL baseline and achieves leading open-source caption quality.
Abstract: Label-supervised manifold alignment bridges the gap between unsupervised and correspondence-based paradigms by leveraging shared label information to align multimodal datasets. Still, most existing methods rely on Euclidean geometry to model intra-domain relationships. This approach can fail when features are only weakly related to the task of interest, leading to noisy, semantically misleading structure and degraded alignment quality. To address this limitation, we introduce FoSTA (Forest-guided Semantic Transport Alignment), a scalable alignment framework that leverages forest-induced geometry to denoise intra-domain structure and recover task-relevant manifolds prior to alignment. FoSTA builds semantic representations directly from label-informed forest affinities and aligns them via fast, hierarchical semantic transport, capturing meaningful cross-domain relationships. Extensive comparisons with established baselines demonstrate that FoSTA improves correspondence recovery and label transfer on synthetic benchmarks and delivers strong performance in practical single-cell applications, including batch correction and biological conservation.
PaperID: 1911, Poster
Authors:
Zichuan Fu, Xian Wu, Jingtong Gao, Wenlin Zhang, Jiaxuan Li, Binhao Wang, Yimin Deng, Guojing Li, Xiaopeng Li, Derong Xu, Yefeng Zheng, Xiangyu ZhaoAbstract: In real-world deployment, large language models (LLMs) are frequently updated through post-training techniques to maintain up-to-date knowledge. Yet their reliability depends not only on what the LLM knows, but also on whether it knows what it knows—a self-assessment known as the feeling of knowing (FoK). FoK is the signal behind selective generation, retrieval triggering, and model routing; when miscalibrated, it leads systems to overconfidence on unknown questions or to refuse ones they could have answered. As LLMs are updated through posttraining, however, we discover that what an LLM knows can change without a corresponding update to its FoK. This creates a knowledge–FoK desynchronization challenge, where FoK estimators calibrated to the pre-trained LLM become unreliable as the model’s knowledge state evolves. To address this challenge, we propose a two-channel design: a knowledge channel that changes what the LLM knows and a FoK channel that judges whether the current knowledge state supports answering. We train these channels with a three-stage procedure: onpolicy supervised fine-tuning (OP-SFT) updates the knowledge channel from corrected on-policy answers, knowledge-delta FoK training (K∆-FT) trains the FoK channel over intermediate knowledge states, and FoK-guided policy optimization (FGPO) uses the judgment to generate the corresponding answer or explanation. Experiments on three datasets show that explicit FoK alignment substantially improves judgment quality on updated knowledge. The code is available at https: //anonymous.4open.science/status/Know-thyself-code-8D65.
PaperID: 1912, Poster
Abstract: While talking head generation has advanced rapidly, generating natural listener behavior in dyadic conversations, which know when to react, how to react, and with what type of response, remains underexplored. Existing dyadic datasets lack fine-grained listener reaction annotations, and prevailing evaluation metrics inherited from talking-head and video generation measure visual realism rather than whether a listener reacted appropriately. We address these gaps along three aspects. First, we curate a listening-head-specific dataset built from RealTalk and Seamless Interaction, comprising approximately 147 hours of paired speaker–listener videos with 64,557 event-level reaction annotations across six categories: . Second, we introduce an audio-driven baseline built on a flow-matching transformer, namely GLARE, with prosody conditioning derived from Qwen2-Audio and a temporal reaction loss that explicitly supervises frame-wise reactions. Third, we propose a reaction-oriented evaluation protocol that jointly measures reaction occurrence (R-F1), temporal alignment (R-tIoU), asymmetric temporal deviation (R-ATD), and reaction-region visual quality (R-FID), giving a more behaviorally grounded assessment than visual-quality-only metrics. Experiment results show consistent gains over prior listening-head methods in both visual fidelity and reaction-level metrics, suggesting that reaction-aware data, modeling, and evaluation are critical for natural listening behavior.
Abstract: Decentralized multi-robot motion planning requires each robot to generate collision-free trajectories from local observations, without global sensing or reliable communication. However, most existing planners, whether classical or learning-based, generate trajectories from a static snapshot of the local observation, which limits their ability to anticipate the future behavior of neighboring robots. This limitation is critical as the number of robots increases and the environment becomes more cluttered. To overcome this challenge, this paper introduces Simulation-Informed Diffusion (SID), a decentralized framework built on constraint-aware diffusion models (CADM). SID first uses CADM to simulate the future trajectories of neighboring robots from their currently observed states, and then uses the same CADM to plan each robot's own trajectory under safety constraints informed by these simulations. Crucially, the accurate simulation of neighbors enables a minimal communication scheme that triggers coordination only when necessary in highly congested scenarios. Experiments across diverse environments show that SID consistently outperforms baseline methods in terms of planning effectiveness and constraint satisfaction, and scales to scenarios with 100 robots and 160 obstacles.
PaperID: 1914, Poster
Abstract: 3D instance segmentation requires representations that are both geometrically structured and semantically discriminative. Pure 3D encoders model spatial structure well but often lack strong appearance semantics, while directly lifting 2D priors into 3D by concatenation injects rich semantics without ensuring that they align well with the underlying 3D structure. In this paper, we address this problem by formulating 2D-to-3D semantic transfer as a masked semantic reconstruction task. Specifically, we mask a subset of background lifted 2D priors and train the 3D encoder to recover the missing semantic priors, providing explicit supervision for distilling rich 2D semantics into the 3D backbone. We further propose geometry-modulated semantic injection, which uses raw 3D geometry to modulate lifted 2D priors before sparse encoding, so that geometry explicitly controls semantic injection and alleviates the distribution gap between 2D semantic features and 3D geometric features. Experiments on two baseline frameworks show that the proposed method achieves better or comparable results in both full-data and few-shot settings on ScanNetV2 and ScanNet200.
PaperID: 1915, Poster
Abstract: Distillation and Reinforcement Learning (RL) fine-tuning are the primary pillars of diffusion post-training. While traditionally studied in isolation, the interaction between these phases remains poorly understood, and in particular how fine-tuning impacts the generative quality of distilled models. We introduce Rewarded Moment Matching Distillation (RMMD), a novel framework that simultaneously distills diffusion models and maximizes a reward function. RMMD preserves the high-fidelity "naturalness" characteristic of advanced distillation (e.g., 8-step Moment Matching) by adapting the sampling loop for on-policy training and repurposing the distillation loss as a proxy for integral KL regularization. By evaluating the FID-Reward Pareto fronts on ImageNet, we demonstrate that RMMD achieves superior trade-offs compared to single-step baselines (DI++) and multi-step competitors (DRaFT, HyperNoise). Finally, we apply RMMD to GenCast, a state-of-the-art weather forecasting model, to distill it while optimizing the Continuous Ranked Probability Score (CRPS) metric. The resulting distilled model achieves a 7.5x speedup while outperforming the teacher model on 93% of target weather variables, and being better calibrated. This proves that RMMD scales to complex, high-dimensional scientific domains.
Authors:
Ahmadreza Jeddi, Minh Le, Amirhossein Kazerouni, Hakki C Karaimer, Hue Nguyen, Iqbal Mohomed, Michael Brudno, Alex Levinshtein, Kosta Derpanis, Babak Taati, Radek GrzeszczukAbstract: Modern Vision-Language Models (VLMs) benefit from chain-of-thought prompting and test-time scaling, but these gains often come with prohibitive inference cost due to large visual contexts and long decoding chains. We view this cost through two coupled axes: Visual Context Scaling (VCS), which controls how much visual evidence is passed to the language model, and Visual Reasoning Scaling (VRS), which controls how much inference-time reasoning search is performed. Existing methods typically optimize one axis at a time, leaving the joint allocation of compute across these axes underexplored. We introduce Adaptive Visual Inference Scaling (AVIS), a lightweight policy that adapts both VCS and VRS per query. AVIS realizes VCS through Key Diversity Visual (KDV) pruning, a training-free O(N) key-based rule for removing redundant visual tokens before prefilling, and realizes VRS through adaptive self-consistency, using a learned difficulty predictor to select the number of reasoning rollouts. AVIS is deployment-friendly and compatible with shared-prefill inference, where all rollouts reuse a single prefilling pass and KV cache. Across diverse image and video reasoning benchmarks, AVIS improves the accuracy--compute trade-off relative to VCS-only and VRS-only baselines, and remains effective on top of RL post-trained VLMs while keeping compute and latency low.
Abstract: Time-series carry structure simultaneously at multiple scales (fine-grained variation, mid-range motifs, and global properties) and downstream tasks operate at correspondingly different scales. Most existing self-supervised learning approaches supervise representations globally via instance-level contrastive losses and limited temporal neighborhood supervision, but do not explicitly exploit the structural hierarchy. We propose a learning framework that explicitly enforces a structural hierarchy across three scales independently: a local objective for token continuity, a mid-range objective for window-level motifs, and a global objective for sequence-level agreement. Realizing this framework requires the backbone to expose a representation at each scale; we introduce Progressive Memory Transformer (PMT), which augments a transformer with writable, window-aligned memory that exposes the mid-range scale alongside the token and sequence-level representations conventional transformers already provide. Across seven UCR/UEA/UCI classification benchmarks, a cue-retention probe, and two forecasting benchmarks, PMT learns representations that probe well at the global, mid-range, and local scales---strong low-label classification (1--5% labels), competitive forecasting performance across multiple horizons, and quantitative and qualitative evidence that memory states capture mid-range motifs.
Abstract: Prediction markets aggregate dispersed beliefs into prices that act as probabilistic forecasts of uncertain events. Classical theory establishes a clean equivalence between forecasting accuracy and trading profit, but only for the specific automated market maker (AMM) design. However, the largest exchanges today are based on central limit order books in which informed forecasters routinely lose money while uninformed strategies can profit on simple heuristics. We resolve this discrepancy by establishing a formal equivalence between predictive accuracy and profitability. For any strictly proper scoring rule S, we exhibit a ``proper'' betting strategy that depends only on the forecaster's prediction \mathbfp and the market price \mathbfq and earns positive expected profit whenever \mathbfp outperforms \mathbfq under S and the market has sufficient liquidity. The proof rests on a decomposition of expected profit that strictly generalizes the classical AMM guarantee and also explains how strategies can profit without an accuracy edge. Empirically, across thousands of forecasts from AI models, proper betting is the only strategy that reliably converts accuracy into profit, and we further identify systematic forecasting personas and show how the optimal proper strategy varies across them. A month-long live deployment achieves +80.33% return on investment with a Sharpe ratio of 3.35.
Abstract: Generative models excel at synthesizing high-fidelity samples from complex data distributions, but they often violate hard constraints arising from physical laws or task specifications. A common remedy is to project intermediate samples onto the feasible set at each sampling step. However, repeated projection can disrupt the learned sampling dynamics because this feasible set is defined for ''clean'' samples rather than noisy intermediate states. Thus, recent studies first estimate ''clean'' samples and then project them to the feasible set, but this increases algorithmic complexity and accumulates errors across steps. To address these issues, this paper proposes Chance-constrained Flow Matching (CCFM), a training-free method that formulates constraint enforcement during sampling as stochastic optimization while retaining the hard-constraint feasibility guarantee of projection-based methods. By avoiding direct projection of noisy intermediate states onto the clean-sample feasible set, CCFM mitigates the distributional distortion. Unlike approaches that first estimate and project ''clean'' samples, CCFM avoids complex multi-stage procedures. Experiments show that CCFM outperforms current state-of-the-art constrained generative models in modeling complex physical systems governed by partial differential equations, molecular docking problems, and physics-informed motion generation, delivering superior feasibility and fidelity.
Authors: Tung Do, Thuan H Nguyen, Hao Li
Abstract: Few-step image generation has seen rapid progress, with consistency and meanflow-based methods significantly reducing the number of sampling steps. Despite their low inference cost, these approaches often suffer from training instability and limited scalability. Sphere Encoder is a recent alternative that produces high-quality images in only a few steps; however, it requires repeated transitions between the pixel space and latent space during inference while jointly optimizing reconstruction and generation within a single architecture. This design leads to computational inefficiency and objective conflict between reconstruction and generation. To address these limitations, we decouple the framework into a fixed pretrained image encoder and a separate latent denoising model trained entirely in a spherical latent space. Our approach eliminates repeated pixel-space operations during training and inference, improving efficiency and allowing reconstruction and generation to specialize independently. On Animal-Faces, Oxford-Flowers and ImageNet-1K datasets, our method significantly outperforms Sphere Encoder in both generation quality and inference speed, while achieving competitive results against strong few-step and multi-step baselines.
PaperID: 1921, Poster
Abstract: Masked patch self-distillation has become a default ingredient of modern joint-embedding pretraining, and is widely regarded as a uniform enhancement of dense visual representations. In this work, we revisit this assumption and argue that its effect is more nuanced. Through controlled experiments across multiple joint-embedding objectives and model scales, we observe a consistent tension: while the auxiliary patch objective strengthens semantic abstraction, it systematically erodes fine-grained spatial correspondence. To examine this phenomenon beyond the reach of human-annotated benchmarks, we introduce annotation-free diagnostics that probe how much local visual evidence remains accessible in the learned representation. We further provide a population-level theoretical analysis showing that this trade-off is not incidental but follows directly from the conditional nature of masked semantic prediction, under which target variation unpredictable from visible context is provably averaged out. Taken together, our findings recast masked patch self-distillation as an explicit trade-off between semantic abstraction and local detail, rather than a uniform improvement to dense features.
PaperID: 1922, Poster
Abstract: Low-rank Adaptation (LoRA) has been widely used as a parameter-efficient method in fine-tuning large language models (PEFT). However, trained adapters often exhibit low stable rank and small trailing singular values. Near such rank-deficient regions, fixed-rank formulations suffer from severe Hessian bounds and local Lipschitz constants that grow with \mathcalO\big(\sigma^-1_r(X)\big). To mitigate this issue, we propose ROLAND (Riemannian Optimization for Low-Rank AdaptatioN via Desingularization), a desingularization method that lifts LoRA to a smooth bounded-rank geometry with both left and right null-space projectors. ROLAND replaces rank-deficient singularities with compact and smooth fibers while retaining a low-rank representation of the adapter update. Within this framework, we equip the manifold optimization with a new Riemannian metric and retraction mechanism, which allow us to establish tighter Hessian bounds and milder condition-number dependence in the local convergence rate. Empirical results demonstrate ROLAND's superiority across a wide range of experimental settings compared with other state-of-the-art PEFT methods.
Abstract: Large language models (LLMs) are increasingly used in high-stakes settings where they are expected to justify their predictions, and a common approach to understanding LLM predictions is via self-generated free-text explanations. However, free-text explanations are instance-wise and non-executable, making it hard to evaluate their alignment with model behavior or explain dataset-level behavior in a compact, human-interpretability manner. We study code as an executable alternative to free-text, enabling automatic evaluation and dataset-level explanation, and ask whether LLMs can \emphconsistently explain their behaviors with code. This automatic verification enables a generation procedure that samples multiple candidate programs and selects the most consistent one. Using three open-source code-capable models and eight tasks with varying levels of linguistic difficulty, we study when LLMs produce consistent code explanations. We find that models can produce highly consistent code explanations in a subset of settings, but strong consistency is not the norm. Consistency is substantially lower for tasks involving more complex linguistic components, where performance is often near random guessing across models. In these harder settings, generated programs are longer but not more structurally complex, and increasingly rely on hard-coded linguistic patterns from the input. These findings suggest two possibilities: either LLMs are not good at producing compact explanations of their behavior, or some behaviors require dense, distributed computation that cannot be explained by compact, human-interpretable programs. Finally, we find that model-internal representations do not reliably predict when code explanations will be consistent, suggesting that LLMs do not know when they can or cannot explain their own behavior with code. Our work highlights code as a compact and verifiable alternative to free-text explanations, while also revealing limitations in using compact, human-understandable explanations to fully understand LLM behavior.
Authors:
Theofilos Mailis, Kalliopi-Christina Despotidou, Konstantinos Filippopolitis, Yannis Foufoulas, Thanasis-Michail Karampatsis, Andreas Ktenidis, Evdokia Mailli, Theodore Papamarkou, Yannis IoannidisAbstract: Federated learning and analytics are often described as collections of separate protocols, even when they share the same mathematical form: client-local tensor computation, mergeable aggregation into shared state, and shared-only post-processing. We introduce a typed tensor language that formalizes this structure. The language distinguishes federated tensors, whose records are partitioned across clients along a tracked record axis, from shared tensors, which are available globally. Its semantics are defined by comparison with a virtual global tensor, used only as a reference object. The main result is a shared-state factorization theory. We show that typed one-round programs factor through fixed-dimensional shared state whose size is independent of the number of clients and records, computed from client-local tensor expressions and merged across clients. We also prove a converse representability result; factorizations whose encoders and decoders are expressible in the language are realized by typed one-round programs, and the correspondence extends to iterative programs whose cross-round state is shared. This gives a formal account of the computations in the language that can be expressed as encode, merge, and decode procedures. We then develop a differentiable fragment for learning. If a per-record loss and its per-record gradient are represented by client-local tensor expressions, the global gradient is represented by record-axis summation of the federated gradient tensor. This yields typed iterative programs for server-side gradient descent and shared-linear-algebra second-order updates. The framework characterizes a broad class of federated learning computations whose communication passes through fixed-dimensional shared state.
PaperID: 1925, Poster
Abstract: Long-horizon RTL agents often collapse under context pollution, lacking the active, compartmentalized debugging strategies of human engineers. We introduce OneVeriAgent, a human-like framework modeled directly on expert cognitive workflows. First, mimicking how engineers probe waveform viewers, Agentic Temporal Exploration (ATE) enables dynamic, targeted signal queries, distilling dense simulator dumps into sparse, causal traces. Second, reflecting how designers compartmentalize tasks, a task-scoped Orchestrator enforces strict context isolation, preventing raw tool outputs from polluting the global reasoning trace. On the CVDP benchmark, OneVeriAgent achieves 95.8% Pass@1 with Claude Opus 4.6, matching state-of-the-art proprietary swarm. Finally, we propose Tool-Validated Self-Distillation (TVSD) to bridge the open-source capability gap. By fine-tuning on self-generated, syntactically valid trajectories without external correctness oracles, TVSD lifts Gemma4-31B-Instruct's Pass@1 from 61.1% to 74.8%, demonstrating that fully open-weight RTL agents are viable.
Authors: Wentao Sun, João P Nogueira, Dominique Verchere, Mathieu Acher, Alonso Silva
Abstract: Corr2Cause asks whether a causal claim holds in every DAG compatible with observed correlations and conditional independencies. We frame this as latent-object reasoning: the label is defined by a CPDAG query, but free-form chain-of-thought often collapses the Markov-equivalence-class problem into local pattern matching. We propose Structured Thinking, a two-turn pipeline that first externalizes a typed, schema-constrained CPDAG summary and then answers against that graph state. On the Corr2Cause full test, Structured Thinking raises Qwen3.5-27B from 73.0 to 86.4 F1(Yes) over a strong PC-instruction baseline in the primary paired run (+13.4 percentage points; McNemar p=2.4×10^-6; bootstrap 95% CI [+8.4, +18.6]); across three full-ID seeds, the mean gain is +8.1 \pm 5.3 percentage points. A PC-scaffolded two-turn prose control reaches only 67.6 F1, indicating that a detailed PC scaffold plus a schema-free prose intermediate is not sufficient. The same pattern holds on Qwen3.6-27B, Paraphrase-OOD, and GPT-5.4-mini. Scrambling the emitted CPDAG costs 12.0 percentage points in F1, and a full-split audit shows close agreement with the reference CPDAG, with ID skeleton F1 of 0.960 and exact CPDAG match of 75.9%. These results support a bounded design principle: externalize the latent object that defines the label, constrain its form, and test whether downstream answers use it.
PaperID: 1927, Poster
Abstract: On the simplest non-Euclidean manifold, the 1D flat torus, we exhibit a distribution where a recently proposed extension of Riemannian flow matching is globally optimized by a velocity field that pushes samples away from the data distribution. We trace this failure mode to a property we call compatibility; a likelihood-based flow matching objective is compatible if its population minimizer recovers the marginal velocity field induced by the data and coupling. We give a general condition for compatibility, apply it to existing methods, and bound the worst-case error in terms of a residual component vanishing under compatibility. Next, we note that existing compatible objectives put density on velocities that are infeasible on compact manifolds. We propose an objective built from an exponential family over endpoints, with the conditional velocity as sufficient statistic, which is both compatible and supported only on feasible velocities. This manifests as a re-weighting of the flow matching loss which, on constant-curvature manifolds, treats errors in flow speed and direction differently. Empirically, our method recovers the correct velocity field on the test case, outperforms prior methods on high-dimensional product tori, and obtains substantial NLL improvements over prior work on two of four standard spherical geospatial benchmarks at moderate runtime overhead.
Abstract: Cyclic block coordinate methods are widely used in practice for their simplicity and strong empirical performance. Yet, their theoretical behavior is challenging to explain, and setting their step sizes-beyond classical coordinate descent for minimization-requires careful tuning or line-search machinery. In this work, we develop \textttADUCA (Adaptive Delayed-Update Cyclic Algorithm), a cyclic algorithm addressing a broad class of Minty variational inequalities with monotone Lipschitz operators. \textttADUCA is parameter-free and locally adaptive: it requires no global or block-wise Lipschitz constants, it uses no per-epoch line search, and it adapts to local problem geometry. A key feature of the algorithm is using operator information delayed by a full cycle, which makes the algorithm compatible with parallel and distributed implementations, and attractive due to weakened synchronization requirements across blocks. We prove that \textttADUCA attains (near) optimal global oracle complexity as a function of target error \epsilon >0, scaling with 1/\epsilon for monotone operators, or with \log^2(1/\epsilon) for operators that are strongly monotone.
PaperID: 1929, Poster
Abstract: Algorithmic recourse (AR) aims to provide a recourse action that alters an undesired prediction made by a machine learning model. While standard AR methods assume the model does not change over time, this assumption is often violated in practice due to distribution shifts or data updates, rendering the suggested actions invalid. To address this issue, recent studies have proposed several methods that provide actions robust to model changes. Most of them construct an uncertainty set of models by perturbing continuous model parameters, such as the weight vectors of neural networks, and then seek actions that remain valid for all models in the set. However, such an approach cannot be directly applied to tree ensembles because their model parameters include discrete tree structures, which cannot be perturbed in the same way as neural networks. In this paper, we introduce a new robust AR framework for tree ensembles through the lens of distributionally robust optimization (DRO). Our key idea is to express the model change in a tree ensemble as a perturbation to the distribution of predictions made by decision trees in the ensemble, rather than to the model parameters themselves. This formulation allows us to naturally cast the robust AR problem for tree ensembles as a DRO problem. We then show that, by choosing a specific discrepancy measure between distributions, our robust validity constraint can be reformulated into a tractable one that can be evaluated as efficiently as the standard validity constraint. Building on this reformulation, we propose an efficient optimization algorithm by extending the existing search-based method for tree ensembles. Experimental results demonstrated that our method achieved a better robustness-cost trade-off than existing robust methods applicable to tree ensembles, while maintaining comparable runtime to the standard method.
PaperID: 1930, Poster
Abstract: Transformer models built on the attention mechanism have become a central building block in modern deep learning, yet softmax attention remains a major bottleneck for long-context workloads. While FlashAttention makes the forward and first backward passes I/O-efficient, it does not support backward-over-backward (BoB), which enables exact differentiation through the backward pass for applications such as second-order optimization, test-time training, gradient-based memory, and meta-learning. Existing BoB implementations either materialize large intermediate tensors or exhaust GPU memory at long sequence lengths. We present FlashBoB, an exact, I/O-efficient algorithm for BoB in softmax attention that keeps computation within on-chip tiles and avoids all N × N intermediate tensors, where N is the sequence length. The key insight is a hierarchical affine structure in the softmax double backward: two row-wise scalars determine all outputs through affine transformations. This yields a two-pass schedule with bounded on-chip static random-access memory (SRAM) usage and minimal off-chip high-bandwidth memory (HBM) traffic. FlashBoB achieves \Theta(N^2 d^2/M) HBM traffic and, within the standard FlashAttention-style score-recomputation model, matches the inherited large-cache lower bound for exact forward attention. Empirically, it scales exact attention BoB to N=262\textK on a single A100 80GB GPU, where prior PyTorch exact baselines fail by N=16\textK, and is up to 6.3× faster than FlashBack. These results make exact second-order attention practical at long-context sequence lengths where prior implementations cannot run efficiently.
PaperID: 1931, Poster
Authors: Bongsu Jung, Donghee Kim, Wooil Kim
Abstract: Spoof diarization aims to jointly localize spoofed regions and determine which spoofing method generated each segment within a partially spoofed utterance. Existing approaches rely on modular pipelines that separate localization and clustering, introducing training--inference mismatch and requiring oracle knowledge of the number of spoofing classes at inference time. We propose E2E-SD, the first \emphclustering-free end-to-end framework for spoof diarization, which reformulates the task as variable-cardinality temporal set prediction. E2E-SD jointly learns localization and class assignment using encoder-decoder attractors with Hungarian-matched training, eliminating the need for post-hoc clustering and oracle class counts. Our framework dynamically generates class-specific attractors through an LSTM decoder and estimates the number of active spoofing classes via attractor existence probabilities, requiring no external class-count information at any stage. On the PartialSpoof benchmark, E2E-SD achieves a \JER of 15.76%, a 41.6% relative improvement over the previous best result (26.99%), while operating fully oracle-free with 88.7% class-count estimation accuracy. The largest gain is observed on unknown attacks, where \JER drops from 49.02% to 21.00% (57.1% relative improvement), demonstrating that end-to-end optimization with dynamic attractors enables more attack-agnostic generalization than modular clustering pipelines.
PaperID: 1932, Poster
Abstract: When large language models (LLMs) are used for creative tasks where the evaluation is largely subjective, two crucial activities are ideation (in which a list of candidate ideas is produced) and pruning (in which the initial candidates are prioritized through a process of comparison). In this framework, the subjective results are influenced by inherent biases in the LLM's generation and evaluation, and it is therefore important to understand how these subjective biases operate. We define and explore a set of stylized tasks that allow us to highlight the effects of these subjective biases in LLM evaluation. In the process, we find evidence for several key underlying principles. First, there are multiple subjective LLM biases at play, and to understand the overall evaluation we must analyze the relative strengths of these biases and their interactions. Second, different prompts -- even when they seem superficially similar -- can have the effect of making different subjective LLM biases more or less salient. And third, there are complex interactions between the LLM's initial idea-generation phase and the subsequent decisions that go into prioritizing and pruning these ideas.
Abstract: Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adjacent token spans are highly predictable or frequently occur as stable units, suggesting that their representations may be compressible. We introduce a method for distilling sequential computation by replacing spans of input tokens with collapsed representations, computed on the fly by a lightweight merge module. This module generates a single surrogate embedding from a sequence of static token embeddings that captures the functional role of the multiple tokens, allowing pretrained models to operate on compressed inputs without architectural changes or re-training. We apply this approach during inference to compress both prompts and intermediate decoding steps, using a rollback mechanism to substitute stored multi-token KV cache entries with their single-step surrogates. Experiments across diverse models show that the merge module can be used to reduce effective sequence length by up to 40% with minimal accuracy degradation across language-modeling evaluations and downstream tasks, including question answering, summarization, commonsense reasoning, and long-form mathematical reasoning. Additional lightweight adaptation of the merge module further improves the accuracy-compression trade-off in selected settings. These results demonstrate that sequential token computation in Transformers can be effectively approximated through condensed surrogate representations that preserve the original behavior without model updating.
PaperID: 1934, Poster
Abstract: Cognitive Load Theory (CLT) shows that human problem-solving is affected by both intrinsic load, arising from task difficulty, and extraneous load, arising from how task information is presented. LLMs show analogous effects, where presentation alone can affect model performance, but these effects have not yet been systematically isolated from task difficulty. To address this gap, we introduce CogLoadBench, a controlled benchmark that varies five CLT-inspired presentation factors across three reasoning task families. Across models, we find that higher extraneous load conditions often reduce accuracy by either adding competition or complicating the reasoning path. To support diagnosis and mitigation of load-related effects, we propose the Extraneous Load Score (ELS), a prompt-time metric computed from activations that estimates a model's extraneous load, and show that it generalizes across reasoning task types and load levels. We further show that shifting model representations toward lower-ELS directions at inference time can significantly improve model performance without changing prompt content.
Abstract: Autoregressive perception models trained to localize visual entities under the open-vocabulary setting are mostly trained using Supervised fine-tuning (SFT) with maximum likelihood, yet it optimizes a proxy objective (per-token cross-entropy) that is fundamentally misaligned with perception metrics such as precision and recall. In this paper, we explore post-training reinforcement learning (RL), specifically GRPO, to directly align these models with their evaluation metrics. Building up on the recently introduced Falcon Perception, we design an RL framework that addresses perception-specific challenges: reward design for set-structured outputs and multi-head sampling control. We discover multiple benefits from RL for perception: first, RL unlocks unprecedented performance in very dense scenes (up to 500 objects per scene) and to the best of our knowledge, no other system has such capabilities at this scale; furthermore it fixes common issues in autoregressive perception models like mask repetitions and removes almost entirely the need for NMS and coordinate deduplication, which improve both performance and efficiency and remove the need for hyperparameters tuning; overall, we notice improvements on all levels of difficulties in referring expression segmentation (on PBench and SACO-Gold), and we find an elegant way to preserve the knowledge of whether an object exists or not (as evaluated by MCC) without training on negative samples. We show that a simple reward penalizing false negatives and positives is enough. We develop two hybrid self-annotation pipelines, respectively tailored for difficult referring expressions and very dense scenes, and show their benefits on RL-training. Model weights and datasets will be made public.
Abstract: Verifiable Reward (RLVR) has emerged as a promising approach for enhancing the reasoning capabilities of large language models. However, it still remains unclear whether RLVR can extend reasoning capabilities beyond what is already encoded in the base model. In this paper, we study this question through the lens of — the tendency of models to over-rely on already-learned representations, disproportionately improving performance on problems they already handle well while underexploring harder ones that require departing from those representations. Through our controlled experiments and large-scale analysis, we show that RLVR training on mixed data tends to further sharpen performance on easier problems while yielding only marginal gains on harder ones. We find that although hard problems lead to modest gains overall, they are actually the ones that help most in expanding the model’s reasoning capabilities and pushing the representations. However, prolonged training on these hard problems may also hurt and lead to forgetting on easy problems. These findings help to better understand some of the prior conflicting observations about RLVR, specifically, why on (OOD) domains where the base model performs poorly, improvements on hard problems can expand reasoning; while on well-trained domains, forgetting can reduce performance and mask genuine gains obtained from hard problems. Data and code for reproducing our experiments are available at: https://anonymous.4open.science/r/neurips26_sub-0DE8.
PaperID: 1937, Poster
Abstract: Egocentric visual grounding requires high-resolution inputs to localize small objects. However, scaling Multimodal Large Language Models to this domain is constrained by the excessive cost of visual token processing. We identify that current efficient strategies based on token reduction are unreliable to select object-centric spatial evidence. To overcome this, we propose SmartRes, a framework that performs efficiency optimization in the pixel space via dynamic resolution routing. SmartRes first encodes a low-resolution view for global context and uses a lightweight router to activate high-resolution patches in object-centric regions and constructs an order‑preserving visual sequence. To further enable robust routing under severe foreground‑background imbalance, we introduce a margin-regularized routing objective that increases foreground-background logit separation and improves foreground recall. Experiments on Ego4D and EgoIntention show that SmartRes reduces visual tokens by up to 67% while retaining 86.4% of full-resolution performance, and achieves up to 1.66× faster inference than state-of-the-art token reduction methods with higher accuracy. Furthermore, strong performance on small object grounding indicates the effectiveness of SmartRes towards egocentric applications. Code will be publicly available.
PaperID: 1938, Poster
Abstract: Computing informative gradients through nonsmooth physical models—such as those involving hard contact or bouncing—is a cornerstone challenge in robotics and machine learning. Randomized smoothing offers a principled way to obtain well-behaved surrogate gradients, yet standard first-order estimators of these smoothed gradients are biased whenever the underlying function is discontinuous. In this work, we show that this bias arises from neglected discontinuity contributions and derive a corrected first-order formula that eliminates it. When the discontinuity structure is known, we exploit it to build discontinuity-aware gradient estimators, which we instantiate for differentiable collision detection on 3D meshes. When it is unknown, we propose a complementary estimator that leverages classical gradients to reduce variance. We validate this estimator on nonsmooth trajectory optimization instances, demonstrating variance reductions of several orders of magnitude over zeroth-order baselines while remaining unbiased under discontinuities where standard first-order estimators fail.
Abstract: Recently, Muon and related spectral optimizers have demonstrated strong empirical performance as scalable stochastic methods, often outperforming Adam. Yet their behaviour remains poorly understood. We analyze stochastic spectral optimizers, including Muon, on a high-dimensional matrix-valued least squares problem. We derive explicit deterministic dynamics that provide a tractable framework for studying learning behaviour with a focus on (stochastic) SignSVD and (stochastic) SignSGD, the latter serving as a proxy for Adam. Our analysis shows that for large batch size, SignSVD performs a square-root preconditioning with respect to the data covariance spectrum, while for small batch size smaller eigenmodes behave like SGD, slowing down convergence. We contrast with SignSGD which for generic covariance performs no preconditioning and has no transition, leading to different optimal learning rates and convergence characteristics. The two methods match up to a constant factor with isotropic data, but behave differently with anisotropic data. An analysis of a power law covariance model with data exponent \alpha and target exponent \beta, and show there are three phases in the (\alpha,\beta) plane: one where SignSGD is uniformly favored, one where SignSVD is uniformly favored, and a third where the two methods exhibit a trade-off in performance.
PaperID: 1940, Poster
Abstract: Recent advances in video generative models have enabled flexible conditional generation and editing pipelines, such as text-to-video, image-to-video, and reference-guided synthesis. However, these workflows are inherently one-directional, mapping conditions to videos without providing a mechanism to invert this process. In complex media production settings, the ability to recover controllable textual conditions from a given video is highly desirable, yet remains underexplored. While inverse prompting has been studied for image generation, directly extending these approaches to video is challenging due to the need to capture temporal dynamics beyond static appearance, as well as the substantial computational cost of per-video optimization. In this work, we propose the first practical framework for efficient video inverse prompting. Specifically, we formulate inverse prompting as a meta-learning problem and introduce Meta Video Inverse Prompting (MVIP), which learns a meta-prompt optimized over temporal dimensions to enable fast and effective prompt inference. Our approach amortizes the inversion process, eliminating the need for costly iterative optimization at test time. Experimental results demonstrate that MVIP recovers higher-quality inverse prompts with significantly improved reconstruction fidelity compared to image-based inversion and iterative captioning baselines, while achieving substantial gains in efficiency.
PaperID: 1941, Poster
Abstract: Human vision flexibly extracts part-whole hierarchies from visual scenes, but representing such structures remains a key challenge for neural networks. Inspired by the neural syntax hypothesis in neuroscience, we propose a framework for representing hierarchical part-whole relationships through nested neuronal coherence, characterized by its continuous and distributed nature. Nested neuronal coherence refers to dynamical states where distributed activations are temporally correlated to form cell assembly sequences across timescales, argued to represent neural syntax. Building on this framework, we develop a cortical-inspired hybrid model, Composer, which dynamically achieves emergent nestedness when given images. To evaluate the emergent hierarchy, we create four synthetic datasets and three quantitative metrics, demonstrating the model’s ability to parse scenes of varying complexity. Overall, our work advances a systematic paradigm spanning representation, implementation, and evaluation toward building human-like vision in neural networks.
PaperID: 1942, Poster
Authors: Alvard Barseghyan, Ani V Vanyan, Hakob Tamazyan, Hrant Khachatrian
Abstract: Current scaling law formulations suggest that increasing the number of unique samples in a dataset consistently improves model performance, while repeated exposure to the same samples provides diminishing value. In this work, we challenge this assumption for image self-distillation models. Through systematic experiments across diverse datasets, namely Maxar, Waymo, and ImageNet, we demonstrate that the effects of sample repetition on downstream performance are non-trivial and dataset-dependent. We discovered one scenario when repeatedly training on a subset of data can outperform training on a larger set of unique samples. This finding highlights a critical limitation of existing scaling laws. To address this gap, we propose a modified version of the Chinchilla scaling law that explicitly disentangles two factors: number of unique samples and repetition per sample. This formulation provides a more precise framework for understanding data-compute trade-offs and better predicts performance in regimes where repetition play a central role. Finally, we analyze how the inherent diversity of the smaller unique subset influences final downstream performance.
PaperID: 1943, Poster
Authors:
Yequan Bie, Peng Xie, Zhixuan CHEN, Yihui Wang, Andong Tan, Linshan Wu, Sunan He, Jianda Mao, Yangqiu Song, Kani Chen, Hao CHENAbstract: Multimodal large language models (MLLMs) have made significant strides in natural vision-language understanding, yet their potential in specific domains, such as healthcare, remains largely untapped. Since existing medical models often lack the reasoning capabilities needed for complex decision-making, adapting general MLLMs for medical reasoning has attracted increasing interest, with reinforcement fine-tuning (RFT) emerging as a promising approach. However, medical reasoning tasks require both precise thinking processes and generalization of well-justified answers, presenting unique challenges due to the inherent scarcity of annotated data and nuanced visual complexity of medical images. Current visual RFT methods prioritize answer correctness through verifiable rewards while neglecting the reasoning process, leading to limited reasoning capabilities and sub-optimal performance, which are essential in high-stakes scenarios like healthcare. To address these issues, we propose MedR2FT, a Medical Reasoning Reinforcement Fine-Tuning framework that enhances medical reasoning through rationale bootstrapping and reasoning rewarding. Specifically, MedR2FT leverages distilled rationales from fundamental training phases to supervise subsequent reasoning processes with semantic rewards, alleviating the challenge of scarce reasoning data and enhancing model reasoning. Extensive experiments demonstrate that our method consistently outperforms baselines on visual question answering and lesion detection with a significant margin. MedR2FT explicitly attends to the reasoning process, significantly enhancing MLLM adaptability for medical reasoning. The code will be available.
PaperID: 1944, Poster
Abstract: We study the problem of estimating smooth optimal transport (OT) maps between two probability distributions under differential privacy (DP) constraints. Leveraging wavelet-based density estimators and recent stability bounds for smooth OT maps, we propose differentially private estimators that apply to both central and local DP models. Our main estimator achieves near-minimax optimal rates in dimension d \geq 2, and we complement it with a quantile-based estimator that attains minimax optimal rates in dimension d=1 under central DP. We further establish matching minimax lower bounds, confirming the near-optimality of our approach. To the best of our knowledge, this constitutes the first differentially private procedure for OT map estimation with provable minimax optimality guarantees.
PaperID: 1945, Poster
Abstract: Speculative Decoding (SD) has shown strong success in accelerating large language model inference, but its efficiency remains limited for Vision-Language Models (VLMs). The diverse input modalities in VLMs make it difficult for a small trained drafter to achieve consistently high throughput across modalities. We address this limitation with an adaptive self-speculative decoding method. In particular, we adaptively select the skipped layers of VLMs to form the draft model that best fits the given query. Our selection algorithm SP^2ec utilizes the single-peak structure of the throughput with respect to the number of selected layers. Our algorithmic design explicitly minimizes the novel notion of wall-time regret for the given query, corresponding to the empirical wall-time latency. Theoretical guarantees on the wall-time regret show that SP^2ec is adaptive to the per-prompt single-peak structure. We conduct extensive experiments with Qwen3-VL and LLaVA-1.5 models across visual and textual tasks, which demonstrates the efficacy of the proposed SP^2ec.
PaperID: 1946, Poster
Abstract: We propose a geometry-aware perturbation framework that explicitly models the low-rank, anisotropic, and evolving structure of gradient dynamics to improve generalization in heterogeneous federated learning (HFL). Our approach is built upon three key components. First, we represent perturbations within a low-dimensional subspace that captures the dominant directions of gradient trajectories. Second, we generate perturbations through a structure-aware sampling strategy that aligns with the covariance of projected gradients, enabling distribution-aware exploration. Third, we introduce an adaptive mechanism that dynamically adjusts both the magnitude and the structure of perturbations to match the evolving optimization process. Extensive experiments demonstrate that the proposed method consistently improves generalization across diverse HFL benchmarks, achieving up to a 4.49% improvement on the Office10 dataset with minimal computational overhead.
Abstract: Convolutional neural networks (CNNs) are widely used for time-series classification, but their deployment in critical domains requires understanding the temporal and spectral patterns that drive their predictions. Concept extraction (CE) methods identify such patterns by analyzing representations within the models' latent space. However, existing time-series CE methods have three limitations: they operate only in the time domain and overlook frequency features, predefine the number of concepts, and produce localizations misaligned with the regions the model uses. We address these limitations by proposing CENDRe, a concept extraction method for CNNs. It first discovers concepts by clustering per-timestep latent representations in two stages, where silhouette-guided aggregation selects the number of concepts automatically. Then, it localizes each concept through gradients of a presence score that contrasts the latent representations with their prototypes, producing masks that concentrate on the regions driving the concept. These gradients, propagated through a differentiable invertible mapping of the input such as a Fourier transform, yield localizations for the same concepts in the frequency domain. Finally, each concept receives a relevance score that quantifies its contribution to each class. On synthetic benchmarks, CENDRe achieves representation correctness comparable to state-of-the-art CE methods and substantially higher importance correctness. On real bearing-fault data, CENDRe extracts the frequency bands driving the model's predictions, located in regions commonly inspected for fault diagnosis, producing evidence to assess the model that time-domain CE methods cannot.
PaperID: 1948, Poster
Abstract: Learning high-quality representations of mixed-integer linear programming (MILP) is crucial for advancing machine learning (ML)-based methods, yet remains challenging due to the discrete nature of MILPs and their extreme sensitivity to small perturbations. Although contrastive learning has proven highly effective for representation learning in domains such as computer vision, designing effective data augmentations for MILPs is particularly difficult. To address this challenge, we propose MIRAGE, a principled data augmentation framework grounded in the invariance of the piecewise-affine dual price function induced by the branch-and-bound (B&B) algorithm. Viewing this dual function as a recorded, solver-level surrogate of the MILP's lower-bound structure, we derive two complementary augmentation certificates that act as its right-hand and objective-side projections. Both certificates are provably invariant under the same dual price function on the recorded B&B tree, and therefore preserve either the structural similarity of the search tree or the optimality of the original optimal solution. Empirically, we integrate MIRAGE with ML models for two significant downstream tasks—learning to branch and predict-and-search—and observe consistent performance improvements across multiple benchmarks. By grounding MILP data augmentation in a single, solver-induced duality quantity, our work provides a principled and effective approach to MILP representation learning.
PaperID: 1949, Poster
Abstract: We propose MOSEL (Multi-Objective Stackelberg Efficient Learning), a framework for a posteriori multi-objective optimization (MOO) in deep neural networks that recovers a full front of Pareto stationary solutions at the computational cost of standard single-objective training. MOSEL reformulates the problem as a bilevel optimization problem that leverages network modularity to decouple representation learning from objective-preference alignment. Casting the bilevel problem as a Stackelberg game enables solving the original a posteriori MOO problem in a single forward–backward pass. As a result,MOSEL matches the time and memory efficiency of standard single-objective training while enabling scalable Pareto stationary front learning. Empirically, MOSEL uncovers diverse and optimal Pareto frontiers in strongly conflicting settings (e.g., fairness–accuracy). Remarkably, even in weakly conflicting regimes such as multi-task learning, it consistently converges to solutions closer to the utopia point, outperforming both standard single-objective training and specialized multi-task learning methods. These results highlight the broader potential of a posteriori MOO learning as a pathway to efficiently learn more diverse and robust representations, ultimately improving generalization.
PaperID: 1950, Poster
Abstract: Visual instruction tuning effectively adapts a pre-trained Large Language Model (LLM) to process image information alongside text, yet it remains unclear how visual features are embedded into the layer-wise hierarchy of abstractions of the LLM backbone. Across a diverse set of vision-language architectures, we show that instruction tuning primarily serves as a bridge, embedding visual features directly into the intermediate semantic layers of the LLM, bypassing the early layers devoted to unimodal processing. With probing analyses and causal interventions, we show that these intermediate layers are the semantic core of vision–language processing and play a critical role in the performance over a broad set of vision–language benchmarks. In addition, by comparing the geometry of semantically equivalent visual and textual representations, we find that fine-tuning extends and strengthens the existing abstraction phase, aligning visual features with pre-existing textual ones. Finally, we confirm the functional role of this localized alignment by restricting fine-tuning to the intermediate layers alone: this strategy preserves the performance of full fine-tuning on vision–centric benchmarks while reducing training time. Our results suggest that multimodal integration is a localized phenomenon driven by the repurposing of the internal abstraction engine of the LLM.
Abstract: Post-hoc explanation methods for neural networks are often heuristic and lack provable guarantees. A common alternative to such methods that was studied in recent years is to compute a cardinally minimal subset of input features that is provably sufficient to determine the model's prediction. For general neural networks, however, finding such explanations is computationally intractable. In contrast, recent work showed that this problem becomes more tractable for a particular family of neural networks, namely neural additive models (NAMs), though standard NAMs are highly restrictive because they exclude feature interactions altogether. In this work, we study the computation of provably minimal and sufficient explanations for neural networks with restricted feature interactions, bridging the gap between fully general neural networks and purely additive models. We first prove that even with constant-order interactions, the problem is \Sigma_2^P-hard, implying that worst-case exponential complexity is unavoidable. We then identify two provably tractable settings: one based on sparsity in the interaction structure and another based on a stronger interaction-aware notion of sufficiency. Under either one of these settings, we develop algorithms that compute provable explanations for neural networks with any-order interactions while retaining much of the computational efficiency of additive models. Empirically, we show that these explanations are significantly smaller and faster to obtain than those produced by standard algorithms, while being derived from models that achieve much higher accuracy than standard NAMs. Overall, our results help expand the class of neural networks that admit efficient provable explanations, while clarifying the fundamental role of feature interactions in shaping their complexity.
Authors:
Vineet Goyal, Jiangze Han, Will MaAbstract: Direct Preference Optimization (DPO) has become a standard method for preference-based fine-tuning in RLHF pipelines. A practical but largely heuristic step in DPO is dataset curation: given a fixed budget of preference labels, which candidate pairs should be compared to produce the most informative training set and the best downstream policy? This paper studies DPO dataset curation as a sampling-design problem. We model the curated dataset by a sampling design over candidate pairs and analyze how this design propagates through DPO training to the KL-regularized RLHF objective. Our main results provide matching upper and lower bounds on the RLHF optimality gap of the policy learned from a curated DPO dataset. The bounds reveal a simple structural message: the effect of pair selection enters only through a single design-dependent matrix that summarizes how informative the collected comparisons are for learning the policy parameters. This leads to a trace-form criterion that characterizes both an achievable performance guarantee for the DPO estimator and an information-theoretic lower bound for any induced policy estimator, thereby identifying a canonical objective for comparison curation under budget constraints.
Abstract: Reinforcement learning (RL) for non-verifiable instruction following increasingly relies on LLM judges with prompt-specific rubrics as reward signals. While recent methods adapt these rubrics to the evolving policy during training, the training prompts themselves remain static, drawn from fixed corpora. This static approach often results in a critical misalignment between prompt difficulty and policy capability, leaving the judge unable to recover a discriminative reward signal when prompts fail to elicit quality variance among rollouts. To address this misalignment, we introduce LLM-as-a-Tutor, a framework that extends the LLM’s role from judge to tutor: a single model serves as an examiner that pairwise-compares policy rollouts to detect non-challenging prompts, and as a generator that appends atomic constraints to them. This append-only design monotonically raises difficulty in step with the policy’s capability, producing a self-calibrating training signal without external difficulty schedules. On three complex instruction-following benchmarks, our method consistently outperforms both policy-unaware baselines and prior policy-adaptive methods that adapt rubrics or rewrite prompts, suggesting prompt adaptation as a missing axis of policy-awareness in non-verifiable RL.
Abstract: Causal inference is central to scientific discovery, yet choosing appropriate methods remains challenging due to the complexity of statistical methodology and real-world data. Inspired by the success of artificial intelligence in accelerating scientific software, we introduce an evolutionary framework that uses large language models to discover and iteratively refine causal methods. Across benchmarks, our estimators consistently outperform established baselines: our best estimator lay on the Pareto frontier of 58 human submissions for a recent community competition. We also extend the algorithm to achieve competitive results in settings with estimated rewards. Analysis of the evolutionary trajectories shows that agents progressively discover sophisticated strategies tailored to unrevealed data-generating mechanisms. Our findings suggest that language-model-guided evolution could be used in scientific settings with partially observed rewards such as causal inference.
Abstract: Looped language models (LoopLMs) perform iterative latent computation to refine internal representations, offering a promising alternative to explicit chain-of-thought (CoT) reasoning. However, existing reinforcement learning (RL) paradigms primarily target output tokens, creating a structural mismatch with looped architectures whose reasoning unfolds implicitly. In this work, we propose LoopRPT, a reinforcement pre-training framework tailored for LoopLMs. By reframing next-token prediction as a next-token reasoning task, LoopRPT assigns reinforcement signals directly to latent steps using an EMA teacher reference and noisy latent rollouts. This formulation enables RL to directly shape intermediate representations, compressing effective reasoning into fewer iterations. We instantiate LoopRPT on the Ouro architecture across multiple model scales. Results demonstrate that LoopRPT consistently improves per-step representation quality, achieving Pareto dominance in accuracy–computation trade-offs. Notably, significant gains on hard tokens indicate that LoopRPT enhances early-stage reasoning rather than merely encouraging premature exits. Our findings highlight reinforcement pre-training as a principled paradigm for learning efficient latent reasoning in LoopLMs.
Abstract: Vision-language action (VLA) models excel for end-to-end autonomous driving, yet their tendency to over-invoke chain-of-thought (CoT) reasoning mirrors a human cognitive pitfall: overthinking. Just as deliberate reasoning can slow/mislead human judgment in routine tasks, excessive CoT in VLA incurs unnecessary compute overhead and can paradoxically degrade planning performance. Recent research tackles this via adaptive reasoning, which dynamically modulates inference complexity by selectively triggering CoT for complex scenes while responding directly to routine ones. While prior adaptive methods rely on implicit adaptation, recent work shows that explicitly optimizing the route between CoT and direct responses yields superior performance, assuming scene-difficulty annotations which are labor-intensive, poorly scalable, and prone to error due to the subjectivity of judging driving complexity. We explore, for the first time, whether adaptive reasoning with explicit route optimization can be achieved without any scene-difficulty labeling. We introduce STEER, a generic over-reasoning mitigation strategy featuring two key components: (i) routing uncertainty-aware rollout, which calibrates sampling to align routing diversity with model uncertainty; and (ii) cross-route advantage credit assignment, which introduces a differential metric to reinforce optimal routing decisions based on environmental rewards rather than subjective labels. Extensive experiments on NAVSIM v1/v2 demonstrate state-of-the-art results in both planning quality and the Pareto front of inference efficiency.
Abstract: Prior work has identified several factors that can contribute to the performance gap between Adam and SGD, spanning data properties, architecture design, and optimization dynamics. Yet these explanations are often studied in isolation, leaving their relative importance unclear. In this work, we revisit these hypotheses through a controlled empirical study across vision, language, genomics, and graph tasks, spanning modern and classical architectures, and carefully designed training setups. Our results suggest that no single factor consistently explains the Adam--SGD gap. For instance, the Adam advantage can (1) persist under a uniform vocabulary distribution yet nearly disappear under a heavy-tailed one; (2) reverse in favor of SGD in softmax-attention models; and (3) become larger when ReLU is replaced by GeLU nonlinearity. Instead, the gap emerges from interactions between data and architectural properties. Yet, we observe a consistent pattern across our settings: a at which the relative advantage shifts from SGD to Adam as batch size increases. This perspective helps reconcile several existing hypotheses while offering practical insights across domains.
Abstract: We study contextual dynamic pricing with linear valuations and bounded-support agnostic noise, whose induced demand curve may be non-Lipschitz with arbitrary jumps and atoms. Such discontinuities break the cross-context interpolation arguments used by smooth-demand pricing algorithms, while the best previous method achieved only \widetilde O(T^3/4) regret. We propose Conservative-Markdown Redirect-UCB Pricing, a polynomial-time algorithm that combines randomized parameter estimation, conservative residual-grid probing, and confidence-based one-step redirection. Our algorithm achieves \widetilde O(T^2/3) optimal regret, matching the known lower bounds of Kleinberg and Leighton (2003) up to logarithmic factors and improving over the previous upper bound of Xu and Wang (2022). Under stochastic well-conditioned contexts, this closes the long-existing open regret gap in linear-valuation contextual pricing under agnostic non-Lipschitz noise distribution.
PaperID: 1959, Poster
Abstract: We study the algorithmic problem faced by an information holder (seller) who wants to optimally sell information to a budged-constrained decision maker (buyer). The utilities of both agents depend on a random state of nature that is revealed to the seller, but unknown to the buyer. Differently from previous works, we consider the case in which the seller is an interested party, as the decision taken by the buyer also influences seller's utility. The seller's goal is to (partially) sell their information about the state of nature to the buyer, so as to concurrently maximize revenue and induce the buyer to take a desirable decision. We study settings in which buyer's budget and utilities are determined by a random buyer's type unknown to the seller. First, we propose a polynomial-time algorithm for computing an optimal seller's protocol, which proposes a menu of information-revelation policies to the buyer, who acquires one of them by paying its corresponding price. Then, we switch the attention to the case in which the seller can only employ a single information-revelation policy, rather than proposing a menu. In such a setting, we completely characterize the computational complexity of the seller's algorithmic problem.
Abstract: We study online alignment of large language models under misspecified preference feedback, where the observed preference oracle deviates from an ideal but unknown ground-truth oracle. The online LLM alignment problem is a bi-level reinforcement problem due to the coupling between data collection and policy updates. Recently, the problem has been reduced to a tractable single-level objective in the SAIL (Self-Improving Efficient Online Alignment) framework. In this paper, we introduce a pointwise oracle uncertainty set in this problem and formulate an oracle-robust online alignment objective as a worst-case optimization problem. For log-linear policies, we show that this robust objective admits an exact closed-form decomposition into the original loss function plus an explicit sensitivity penalty. We develop projected stochastic composite updates for the resulting weakly convex objective and prove \widetildeO(\varepsilon^-2) oracle complexity for reaching approximate stationarity.
PaperID: 1961, Poster
Abstract: We introduce a mathematical framework for studying contraction of the distributional Bellman operator under Maximum Mean Discrepancy (MMD). We show that contraction reduces to conditional positive definiteness of an associated kernel constructed from the base kernel, which we call the Bellman difference kernel. Focusing on radial kernels, we use Schoenberg's theorem to connect this condition to Bernstein functions, yielding contraction criteria, valid in arbitrary dimension, that can be checked analytically or numerically. Our framework unifies known results for the Gaussian RBF, power-distance, and multiquadric kernels; using numerical search, we further identify 12 new contraction-inducing kernels. We also propose , a general construction that produces contraction-inducing kernels at any prescribed rate. To reconcile theory with practice, we introduce conditional positive definiteness on a set, which explains why the Gaussian RBF kernel can remain a valid choice on bounded supports despite its lack of global contraction. Finally, we identify a fundamental tradeoff between contraction rate and discriminative power, where kernels that contract faster tend to discriminate less between distributions.
PaperID: 1962, Poster
Abstract: Accurate traffic prediction is essential for urban planning but remains challenging due to the irregular patterns caused by unexpected events. While Spatio-Temporal Graph Neural Networks (STGNNs) effectively model periodic data, they often struggle with these aperiodic fluctuations. To address this, we propose SIMBAD, which improves generalization on irregular patterns via two key mechanisms: 1) adaptive regulation of periodic signals to prioritize recent signals, and 2) dynamic spatial modeling that adjusts influence between connected nodes based on contextual relations. Built upon simple similarity metrics to adapt to aperiodic patterns, SIMBAD outperforms existing methods on real-world benchmark datasets, demonstrating superior performance in predicting unseen and irregular traffic events.
Abstract: The standard in LLM-based prediction is to use the final-layer representation as the input to a downstream predictor. However, intermediate layers may encode complementary task-relevant signals. Existing approaches therefore either search for the best layer for each task or apply expensive attention-based mechanisms to learn inter-layer aggregation. In this work, we first show that such complexity is unnecessary: a lightweight Graph Neural Network over a fully connected graph of LLM layers is more efficient and achieves significantly stronger predictive performance than existing approaches. We then introduce the Cayley-Encoder, which further improves both efficiency and predictive performance by replacing the fully connected graph with a Cayley graph over SL(2,\mathbbZ_n). These Cayley graphs provide a mathematically grounded topology that is sparse, regular by construction, and has low diameter. This enables effective communication across layers while constraining the aggregation structure and reducing the risk of GNN overfitting. In an evaluation of Cayley-Encoder across 13 tasks and 9 LLMs, Cayley-Encoder consistently outperforms baselines, achieving improvements of up to 40 percentage points in accuracy, while introducing at most 0.1% additional parameters relative to the LLM size. We further show that Cayley-Encoder is effective in few-shot regimes. Finally, we show that Cayley-Encoder outperforms LoRA fine-tuning while operating on the frozen LLM. We conclude with an explainability analysis showing that multiple layers contribute meaningfully to the final prediction, supporting our hypothesis.
PaperID: 1964, Poster
Authors: Durmus Karatay, Richard Newman
Abstract: Many prediction tasks evaluate a tree ensemble on row groups that share feature values: discrete-time survival models expand each patient into G time steps, click-through-rate models score each item in a search session, and scenario analyses vary a few inputs while keeping others fixed. Standard tree-ensemble inference treats each row independently, ignoring this input-side correlation. We instantiate partial evaluation for grouped tree-ensemble inference: constant features are static, varying features dynamic. The algorithm walks each tree once per group, partitions a row bitmask at varying splits, and skips empty subtrees via an unsplit shortcut. We prove a structural work decomposition: per-tree work splits into the constant-projected subtree size |T_c|, G leaf writes, and a predicate-mask provisioning cost Q. For the trace evaluator, per-row work approaches (d_v+1)/(d+1) as G \to \infty, giving asymptotic speedup 1/(1-\bar\rho_G). Empirically, TreeWalker delivers ~3× algorithmic speedup over a row-independent traversal at the reference model size (T=500, L=8; G=16 for SUPPORT/FLCHAIN and empirical groups for Expedia), rising to 6.8--7.8× at G=128 on the survival datasets. For f64 models, outputs match treelite GTIL up to tree-ordering roundoff; for f32 models, TreeWalker's f64 leaf accumulation is close to a Kahan-compensated reference than native f32 on 99.98% of rows, with the remaining 0.02% tied.
PaperID: 1965, Poster
Abstract: Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as intermediate findings emerge. Directed acyclic graph (DAG)-based multi-agent systems are well suited to this setting because they support parallel execution and isolate each sub-task within a focused dependency context. Yet existing DAG-based agents typically instantiate a task-level plan before execution and then repair the graph only after failures or missing evidence are observed. This Plan-then-Patch strategy is brittle for deep research: the system commits most strongly when its evidence is weakest, and later revisions often waste computation on branches that should not have been planned in the first place. To bridge this gap, we propose DAGent, a DAG-based multi-agent framework that introduces Evaluate-then-Grow incremental planning. Instead of committing to a full DAG upfront, an Orchestrator grows the task graph one batch at a time, conditioning each new expansion on confidence and uncertainty signals from completed nodes. To support long-horizon evidence use without overloading each sub-task, DAGent maintains a hierarchical context layer that propagates compact QueryDocs by default while preserving full execution traces for on-demand recall. The cleanly recorded DAG topology in turn admits structural RL signals that an outcome-only recipe cannot define; we instantiate this with DAGRPO, a GRPO adaptation that injects topology-conditioned credit on Executor rollouts and a structural compliance regularization on Orchestrator plans. Across BrowseComp-Plus, GAIA, and xbench-DeepSearch, DAGent surpasses the strongest open-source baseline by 5.3 / 5.8 / 2.0 points at the Qwen3-235B-A22B scale, with the lead replicating across four open-source backbones from four vendors and extending to GPT-5 at 327K context. At the Qwen3-8B scale, DAGRPO further improves over a same-budget outcome-only GRPO baseline by 3.0 average Pass@1 points. A same-architecture comparison further shows that incremental, evidence-conditioned planning recovers accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart, supporting the view that DAGent's gains come from more targeted evidence expansion rather than additional computation. Anonymized code is available at \urlhttps://anonymous.4open.science/r/DAGent-F2D6.
Abstract: Mixture-of-Experts (MoE) architectures scale model capacity through sparse expert activation, but their deployment remains memory-bound because all expert weights must reside in memory. Mixed-precision quantization can substantially reduce this footprint by assigning different bit-widths to different experts. Existing approaches, however, typically rely on calibration data to estimate expert importance and determine bit allocation. For frontier MoE LLMs, the original training data, and hence the true training distribution, is proprietary and inaccessible. As a result, calibration sets are inevitably imperfect surrogates, which can misestimate expert utilization and lead to suboptimal bit allocation. Motivated by the substantial cross-expert quality variability observed in modern MoE models, we propose AlphaQ, a calibration-free bit-allocation method for MoE quantization. AlphaQ draws on Heavy-Tailed Self-Regularization (HT-SR) theory and follows a simple principle: experts with more heavy-tailed weight spectra are typically better trained and should receive higher bit-widths, while experts with weaker heavy-tailed structure can be quantized more aggressively. AlphaQ operationalizes this principle by measuring expert-wise spectral heavy-tailedness and solving a budget-constrained optimization problem that minimizes total quantization error under a global bit-budget constraint. Across several MoE models, AlphaQ consistently outperforms calibration-based baselines under matched bit budgets. Notably, on Qwen1.5-MoE, AlphaQ achieves near full-precision accuracy with an average expert precision of only 3.5 bits, while delivering more than 4× memory compression.
Authors:
Zhankun Luo, Antesh Upadhyay, M. B Sahin, Sang Bin Moon, Anuran Makur, Abolfazl HashemiAbstract: Stochastic estimators are fundamental to large-scale optimization, where population quantities must be inferred from noisy oracle observations. Although influential methods such as momentum, SPIDER, STORM, and PAGE have been highly successful, their analyses are largely estimator-specific and expectation-based, obscuring the structural tradeoffs that determine reliability. In this paper, we develop a unified framework for stochastic variance-reduced estimation based on a recursion with three components: memory retention, reset probability, and a correction term for iterate movement. This framework recovers several classical estimators, motivates new second-order variants, and yields a bias-variance decomposition of estimation error. Our main result is a unified high-probability bound proved using a new dimension-free vector-valued Freedman inequality, valid for smooth normed spaces involving random sums of vector martingales. The result applies in both Euclidean and non-Euclidean settings, including the analysis of mirror-descent-based methods in Banach spaces. As applications, we obtain high-probability oracle complexities for unconstrained optimization with mirror descent, establishing the logarithmic dependence on the confidence level. We also derive the first \tilde\mathcalO(\varepsilon^-3) oracle-complexity bounds for stochastic optimization with expectation constraints, improving upon the existing \tilde\mathcalO(\varepsilon^-4) complexity by leveraging variance-reduced estimation for the first time in this setting.
PaperID: 1968, Poster
Abstract: Video world models that generate and explore environments at scale need a persistent 3D memory of what they have produced. Plane primitives offer a natural candidate: a compact set of oriented planes captures the dominant structural elements (e.g., walls, floors, facades) of indoor and outdoor scenes, while requiring only 7 parameters per primitive. We present PPMem, a video world model that adopts plane primitives as its persistent 3D memory. PPMem augments a video diffusion transformer with a planar prediction head and a Depth Token Fusion (DTF) module: each autoregressive chunk jointly produces RGB video and planar geometry, which merges into a persistent plane memory that conditions subsequent chunks via rendered depth. The resulting memory is three orders of magnitude more compact than per-pixel point clouds (~100K primitives, ~5MB for a full scene), geometrically explicit (metric depth, normals, exportable surfaces), cross-view consistent (jointly predicted by the shared backbone), and adaptive (refined by the merging procedure as generation proceeds). On DL3DV and RealEstate10K, PPMem outperforms prior 3D-grounded world models in video quality while jointly producing metric-scale depth, surface normals, and exportable planar geometry, sustaining coherence across hundreds of frames under large viewpoint changes.
PaperID: 1969, Poster
Abstract: Virtual spatial transcriptomics aims to predict spatial gene expression from histopathology images. However, gene expression is not governed by morphology alone; regulatory relationships are structured, context-dependent, and shaped by the local tissue microenvironment. Existing visual-to-expression models and fixed-prior graph methods struggle to align molecular dependencies with spatially varying histological contexts. We propose DragH2ST, a dual-stage retrieval-augmented dynamic regulatory graph learning framework that constructs task-specific regulatory manifolds and rewires them at the spot level according to local histology. During training, DragH2ST retrieves and integrates heterogeneous biomedical knowledge to build a task-specific regulatory prior, guiding gene representations toward biologically plausible and context-adaptive regulatory structures. To capture microenvironment-dependent regulation, context-gated graph rewiring performs spot-level soft rewiring of gene-regulatory edges conditioned on histology-derived context, enabling adaptive graph reasoning without dense graph reconstruction. At inference, DragH2ST retrieves morphologically similar historical spot prototypes as case-level references, providing evidence-based support for interpretation and biological plausibility assessment. Experiments on multiple benchmarks show that DragH2ST improves prediction accuracy, preserves gene-gene regulatory topology, and captures spatially heterogeneous expression patterns.
Abstract: In the classical identification in the limit model of Gold [Inf. Control, 1967], a stream of positive examples is presented round by round, and the learner must eventually recover the target hypothesis. Recently, Kleinberg and Mullainathan [NeurIPS 2024] introduced generation in the limit, where the learner instead must eventually output novel elements of the target's support. Both lines of work focus on positive-only or fully labeled data. Yet many natural supervision signals are inherently relational rather than singleton: comparative experiments, A/B tests, side-by-side judgments, and similarity–dissimilarity annotations produce observations that encode relationships between examples rather than labels of individual ones. This motivates us to initiate the learning-theoretic study of contrastive identification and generation in the limit, where the learner observes a contrastive presentation of data: a stream of unordered pairs \\x,y\\ satisfying h(x)\ne h(y) for an unknown target binary hypothesis h, but which element is positive is hidden from the learner. We first present three results in the noiseless setting: an exact characterization of contrastive identifiable classes (a one-line geometric refinement of Angluin's tell-tale condition [Angluin, Inf. Control 1980]), a combinatorial dimension called contrastive closure dimension extending the closure dimension of Raman et al. [COLT 2025] and exactly characterizing uniform contrastive generation with tight sample complexity, and a strict hierarchy in which contrastive generation and text identification are mutually incomparable. We then prove a sharp reversal under finite adversarial corruption: there exist classes identifiable from contrastive pairs under any finite corruption budget by a single budget-independent algorithm, yet not identifiable from positive examples under even one corrupted observation. The unifying technical object is the common crossing graph, which encodes pairwise ambiguity, family-level generation obstructions, and corruption defects in a single coverage-and-incidence language.
PaperID: 1971, Poster
Abstract: Decision trees with binary leaves are fundamental building blocks in classical machine learning. While individual trees can be used as strong interpretable classifiers, they frequently serve as members of large collections instead: primarily as components of powerful ensembles, but also as one of the few model classes for which the entire set of near-optimal models (the Rashomon set) can be analyzed. Both types of collections are often more complex than necessary, because there are many decision tree representations for the same decision boundary, imposing unnecessary computational costs in ensembles and complicating theoretical work on Rashomon sets. We propose methods to reduce collections of trees to a subset representing each unique decision boundary in the collection, with each tree in its sparsest equivalent form. In so doing, we improve the scalability of removing provably redundant classifiers in a Rashomon set. We also compress many shallow Adaboost ensembles by >15% of their total number of leaves, while exactly preserving the models' decisions and probability estimates on any possible observation. These methods allow practitioners to run more efficient computations over compressed collections of trees, and remove arbitrary complexity when analyzing and understanding them.
Abstract: LLMs are commonly trained through multi-stage post-training: first via RLHF, then fine-tuned for other downstream objectives. Yet even small downstream updates can compromise earlier learned behaviors (e.g., safety), exposing a brittleness known as catastrophic forgetting. This suggests standard RLHF objectives do not guarantee robustness to future adaptation. To address it, most prior work designs downstream-time methods to preserve previously learned behaviors. We argue that preventing this requires pre-finetuning robustness: the base policy should avoid brittle high-reward solutions whose reward drops sharply under fine-tuning. We propose Fine-tuning Robust Policy Optimization (FRPO), a robust RLHF framework that optimizes reward not only at the current policy, but across a KL-bounded neighborhood of policies reachable by downstream adaptation. The key idea is to ensure reward stability under policy shifts via a max-min formulation. By modifying GRPO, we develop an algorithm with no extra computation, and empirically show it substantially reduces safety degradation across multiple base models and downstream fine-tuning regimes (SFT and RL) while preserving downstream task performance. We further study a math-focused RL setting, demonstrating that FRPO preserves accuracy under subsequent fine-tuning.
PaperID: 1973, Poster
Authors:
Hanming Ye, Yiding Song, Yilun DuAbstract: Humans can learn new tasks from just a few demonstrations. Prior methods like behavior cloning struggle to replicate this ability because they imitate demonstrations without learning the underlying concepts, such as goals, positional relations, and constraints. To address this challenge, we propose COIN, an approach that represents behavior as a composition of reusable concepts. A new task is learned by inferring its concepts from demonstrations then using a policy to implement them. We show that this approach is effective at few-shot learning across diverse domains. On object rearrangement, goal-oriented navigation, and human motion generation, COIN recombines previously seen concepts to solve novel tasks at inference time. On the LIBERO robotics benchmark, COIN achieves a 96.9% average success rate on LIBERO-90, surpassing pretrained policies such as \pi_0 with roughly 10× fewer parameters. This performance gap widens further when fine-tuning on unseen tasks with limited demonstrations, highlighting the advantage of concept inference over end-to-end policy adaptation in the few-shot regime.
Abstract: Accurate emulation of multi-scale physical systems governed by PDEs demands models that remain stable over long autoregressive rollouts while preserving fine-scale structures. Deterministic emulators produce overly-smoothed predictions, while generative approaches better capture details but are costly. Latent-space generative models have emerged as a compromise but with the additional cost of separately pre-trained autoencoders. We propose (WFM), a novel generative emulator that overcomes current trade-offs between cost and skill by performing optimal-transport directly in the multi-scale wavelet space. Rather than learning a latent compression, WFM leverages the hierarchical structure of a U-Net to jointly predict transport velocities of a prescribed wavelet representation. On three challenging systems of chaotic fluid dynamics, WFM achieves superior long-horizon stability, accuracy and spectral coherence compared to state-of-the-art models. Our results clearly position the wavelet space as an effective training-free representation for generative emulation of complex physical dynamics.
PaperID: 1975, Poster
Abstract: LoRA-based continual learning methods mitigate catastrophic forgetting through various mechanisms, yet nearly all complement these with small learning rates as a heuristic to restrict gradient scaling magnitude. Such fixed heuristics lack theoretical guidance on how the strength of this restriction should evolve as tasks accumulate. We reveal that even under directional constraints such as nullspace projection, finite-precision updates inevitably leak into the subspace of accumulated prior knowledge along multiple directions. While small learning rates attenuate such leakage, they cannot prevent the accumulated forgetting from intensifying as the effective rank of historical knowledge grows. We show that the optimal magnitude restriction should adaptively increase with this effective rank to balance stability and plasticity, i.e., preservation of previous knowledge and acquisition of new task information. Under an anisotropic leakage model, we derive a pacing law s^=\sqrtR/c that characterizes the optimal scaling of gradient steps, i.e., the magnitude restriction itself, where R is the effective rank of past updates. Based on this insight, we propose PaLoRA, which compresses historical knowledge via adaptive SVD truncation, projects gradients onto the nullspace of prior tasks, and applies rank-aware adaptive pacing. Experiments demonstrate consistent improvements over prior methods, with particularly strong performance in long-horizon settings, achieving substantial gains of 4% accuracy on challenging 50-task ImageNet-A and ImageNet-R benchmarks.
PaperID: 1976, Poster
Abstract: Long-context inference with Large Language Models (LLMs) is bottlenecked by the linearly growing memory of the key-value (KV) cache. Existing compression methods reduce the cache through token eviction or approximation, but degrade sharply at aggressive compression budgets. We propose PatchKV, a training-free framework that compensates KV cache compression methods by carrying part of the context in the model's weights. PatchKV pairs an off-the-shelf compressed KV cache with a context-specific weight patch, which is computed once at context-loading time and served for downstream queries for the context. The weight patch is derived in closed form via ridge regression, by aligning the block-wise activations of context-derived reference query tokens under the full cache and the compressed cache. Once merged into the model, the patch leaves the forward graph and per-query inference cost unchanged in the single-context, multi-query setting. Across long-context QA (SCBench with up to 170k tokens, SQuAD, NIAH) and math (GSM8K) benchmarks on three model architectures, PatchKV consistently improves cache compression methods, suggesting an alternative direction to compensate them at aggressive budgets.
PaperID: 1977, Poster
Abstract: We study learning-augmented online portfolio selection (OPS), where an investor uses predictions to improve wealth while remaining protected against unreliable advice. In frictionless markets, we characterize the optimal tradeoff between robustness (wealth under arbitrary predictions) and consistency (wealth under perfect predictions) tied to the geometric mean, and show that the optimal tradeoff admits a tractable online characterization. We further quantify how performance interpolates between these endpoints under imperfect predictions by smoothness guarantees. Under transaction costs, we give exact robust-trading certificates and, to our knowledge, the first learning-augmented OPS impossibility under costs: for any positive cost, no GM-robust policy can be maximally consistent under exact predictions. Lastly, we show that a cost-aware greedy rule remains Pareto-optimal.
Abstract: Large Language Diffusion Models (LLDMs) are emerging as an alternative to autoregressive models, offering faster inference through higher parallelism. Similar to autoregressive LLMs, they remain prone to hallucinations, making reliable uncertainty quantification (UQ) crucial for safe deployment. However, existing UQ methods are fundamentally misaligned with this new paradigm: they assume autoregressive factorization or use expensive multiple sampling, negating the efficiency of LLDMs. In this work, we present the first systematic study of UQ for LLDMs and propose lightweight, zero-shot uncertainty signals derived from the iterative denoising process, leveraging intermediate generations, token remasking dynamics, and denoising complexity. We further adapt a state-of-the-art UQ method to LLDMs by combining masked diffusion likelihoods with trajectory-based semantic dissimilarity. We provide theoretical grounding, proving that expected trajectory dissimilarity bounds the masked diffusion training objective. Comprehensive experiments across three tasks, eight datasets, and two models show that our method achieves a great cost-performance trade-off: it approaches the strongest sampling-based baselines while incurring up to one hundred times lower computational overhead. Our work demonstrates that LLDMs can deliver both fast inference and reliable hallucination detection simultaneously.
PaperID: 1979, Poster
Abstract: Understanding what individual components of a neural network do is a central goal of mechanistic interpretability, and a valuable tool for auditing or editing trained models. The dominant approach picks a behavior and reverse-engineers the circuit that implements it. This gives a local picture, since we learn what a weight does in service of one task on one distribution, but not what it does in general. We therefore ask the reverse question. Can we pick any single nonzero parameter in a trained transformer and say what it does, without first choosing a behavior? Until recently, this was out of reach, because the weight basis of a standard transformer is not aligned with the computations the network performs (Elhage, et al. 2022). However, weight-sparse transformers (Gao et al. 2025) offer an alternative path forward. By training most weights to zero, the surviving weights align more cleanly with computation, which opens the door to per-weight interpretation. In this work, we introduce an automated LLM-pipeline that produces a short, human-readable account of when a given weight matters, verifies the account on held-out data, and applies it at scale to compare two weight-sparse transformers against two dense controls. Empirically, we find that a significant percentage of nonzero weights (17-35%) on sparse transformers are interpretable, compared to 5-9% on the equivalent dense model and Pythia-70m. Together these results suggest that sparse training produces transformers whose individual weights are interpretable to a degree not achievable in conventional transformers, providing a potential starting point for per-weight interpretation of larger models
PaperID: 1980, Poster
Abstract: We study the adversarial linear contextual bandit problem, in which an oblivious adversary fixes a sequence of loss vectors while the per-round context (action set) is generated stochastically. In the standard setting where contexts are drawn i.i.d. from a fixed distribution, near-optimal regret bounds are known, but only in expectation. We prove the first such bound that holds with high probability, resolving an open problem stated by Neu and Olkhovskaya [2020], Olkhovskaya et al. [2023] and Dai et al. [2024]. Specifically, the regret is at most O\big( d^3\sqrt\log(K)T\mathrmpolylog(T/\delta) \big) with probability at least 1-\delta over T rounds, where K is the number of actions and d is the dimension of the actions. We then ask whether high-probability guarantees can be extended beyond i.i.d. contexts. Although fully adversarial contexts make sublinear regret impossible in this setting, we show that two natural relaxations remain tractable: (i) non-stationary but independent contexts, whose per-round distributions do not deviate too much from a fixed reference distribution, for which we obtain a high-probability regret bound of O\big(d^3 \sqrt\log(K)T\mathrmpolylog(T/\delta) + d^5/2V\mathrmpolylog(T/\delta)\big), where V is the cumulative total variation distance of the per-round distributions from the reference; and (ii) stationary but dependent exponentially mixing contexts, like, for instance, uniformly ergodic Markov chains, for which we obtain a high-probability bound of O\big(C_\textrmmix d^3 \sqrt\log(K)T\mathrmpolylog(T/\delta) \big), where C_\textrmmix is a constant that depends on the mixing time. Technically, our results build on the recent reduction from adversarial linear contextual bandits to misspecified linear bandits by Van Erven et al. [2025]. All three results are obtained from a single base algorithm under different hyperparameter choices, some of which are functions of data-dependent quantities that are unknown in advance. We resolve this by automatically tuning the hyperparameters using the Corral algorithm of Agarwal et al. [2017]. For this purpose, we develop a new high-probability variant of Corral, since the original version only has guarantees in expectation.
Abstract: Autoregressive video diffusion models generate streaming video by producing frames sequentially, conditioning each chunk on previously generated content. These models are structurally anchored to the first frame: its key-value representation occupies a privileged position in the attention cache and serves as the primary scene reference throughout generation. As the cleanest and most error-free position in the cache, this anchor draws disproportionate attention, suppressing video dynamics, and locking scene composition to the initial viewpoint even as the scene naturally evolves. The result is a temporally shallow video in which motion, camera movement, and scene progression are dampened in favor of static consistency. To address this, we replace the static anchor with an adaptive state, a hidden latent that the model denoises alongside content at every chunk but never generates. Rather than referencing a frozen first frame, the model generates its own scene anchor at each step by attending to both the previous state and the current content, producing a reference that evolves with the generated content. Unlike standard video generation, which encodes an absolute notion of time, our formulation treats time as relative: every generation step sees the same positional structure regardless of how far generation has progressed, and the state transition is identical at every chunk. Together, these properties introduce a recurrence into the generation process, where denoising serves as the transition function, and the KV cache serves as the carrier, requiring no external module. Experiments demonstrate that the adaptive state substantially improves video dynamics, enabling richer motion and natural scene progression within generated videos.
Abstract: The strong zero-shot and long-context capabilities of recent Large Language Models (LLMs) have enabled highly effective list-wise reranking systems. Attention-based rerankers leverage transformer attention patterns as retrieval signals, but not all attention heads contribute equally: many introduce noise and redundancy, limiting both reranking quality and efficiency. In this work, we introduce CoRe heads, a small subset of retrieval-specialized heads identified through a contrastive scoring objective that rewards attention to relevant documents while penalizing attention to distractors. This contrastive criterion isolates highly discriminative retrieval heads and yields a strong training-free list-wise reranker. Our analysis reveals that retrieval-relevant computation in transformer rerankers is highly sparse and structurally localized: fewer than 1% of attention heads account for most reranking performance across models and datasets, and these heads consistently concentrate in middle transformer layers. This localization enables aggressive layer pruning for efficient reranking: removing the final 50% of transformer layers preserves nearly identical reranking accuracy while substantially reducing inference latency and GPU memory usage.
PaperID: 1983, Poster
Abstract: Large language model (LLM) agents increasingly operate through multi-step tool-use trajectories, where harmful intent may be distributed across actions that appear benign in isolation. This challenges single-step moderation, while repeated LLM-as-judge evaluation over growing prefixes can be costly and may intervene too late. We introduce ), a lightweight monitor for early prefix-level detection of harmful agent trajectories. WURI encodes each observed textual step with a frozen text encoder, learns an atom-adapted trajectory representation through metric learning, and scores each prefix with a fixed prototype-margin rule. It requires no access to agent model weights or hidden states, does not modify the agent's internal model, and uses no external LLM-as-judge calls at inference. Across five generalization settings, WURI achieves the best average prefix-area under the detection curve (AUDC), ranks first on three settings, and reaches the strict early-detection operating point within four steps in multiple evaluation settings. Runtime analysis shows a wall-clock speedup of over two orders of magnitude compared with repeated full-prefix guardrail evaluation. These results demonstrate that representation-based prefix scoring is an effective and efficient direction for monitoring unfolding risk in LLM agent interactions.
Abstract: Robust optimization is a widely used framework for decision-making under uncertainty, particularly in high-stakes applications where reliability is critical. A key challenge in this paradigm lies in constructing uncertainty sets that balance robustness and performance: overly conservative sets lead to pessimistic decisions, while insufficient coverage risks failure in practice. Recent approaches based on conformal prediction provide finite-sample, distribution-free guarantees for uncertainty sets, but remain largely task-agnostic and disconnected from downstream decision objectives. In this paper, we propose a decision-aware conformal prediction framework that directly learns the geometry of uncertainty sets to improve robust decision-making. Our approach introduces a polyhedral nonconformity score that induces feature-dependent uncertainty sets, and a three-step procedure that integrates conformal calibration, robust-decision-aware learning, and re-calibration to correct for post-selection bias. We establish finite-sample coverage guarantees for the final, data-dependent uncertainty set, while achieving improved decision performance by reducing unnecessary conservativeness. This work bridges the gap between statistical validity and decision optimality, providing a principled framework for data-driven robust optimization.
PaperID: 1985, Poster
Authors: Jason Brown, Edward Young
Abstract: Reinforcement learning agents often exhibit unintended goal-directed behaviour outside their training distribution, but we currently lack a principled understanding of how such agents will generalise to novel environments based on their training history. We address this gap for agents trained sequentially on one or more tasks. We study over 100 sequential training pipelines, evaluating behaviour across over 250 out-of-distribution environments. We find that salient features drive generalisation, and that goals learnt early in training can persist and influence those acquired later. To explain these phenomena, we introduce , a method that predicts what out-of-distribution behaviour a training pipeline will likely induce. Our method simulates the evolution of low-dimensional latent variables during training according to what would achieve high reward on the training objective with respect to a simple model of how the latent variables map to behaviour. It achieves strong predictive accuracy, generalises to unseen types of training pipeline, and is interpretable. Our findings demonstrate that while out-of-distribution RL agent behaviour is dependent on the whole training pipeline, this dependence has an underlying structure we can capture, laying groundwork for understanding goal generalisation from a developmental perspective.
PaperID: 1986, Poster
Authors: Jaehyeon Lee, Kiwon Um, JungHyun Han, Min-Koo Kang
Abstract: Transformers have rapidly emerged as powerful surrogate models for learning solution operators for partial differential equations (PDEs). Standard attention tends to exhibit depth-wise representational bottlenecks across layers, with the role of depth remaining underexplored in modeling complex PDEs. In this paper, we rethink the role of depth by promoting it to an explicit computational dimension of attention kernel construction, rather than a passive architectural parameter. With this strategy, each attention block's native representation is enriched via a gated mixture of inter-layer differences, enabling information to be coupled and propagated across layers. The depth-wise hierarchy admits a theoretically motivated multi-kernel interpretation with weakly correlated kernels and improved generalization bounds. Beyond dense feature aggregation, our approach yields a more expressive form of depth-wise attention, mitigating spectral collapse, preserving interaction-scale diversity, and improving intra- and inter-layer route utilization. Extensive empirical evaluations demonstrate its broad compatibility across state-of-the-art solvers, presenting consistent performance gains on eight PDE benchmarks.
Abstract: We show that an "old dog," the classical discrete Laplace (aka. geometric) mechanism, can "perform new tricks": It can be post-processed to yield a simple, unbiased estimator of any subexponential function f of the original data, giving a simple, discrete, multivariate version of the recent unbiasing result for the Laplace mechanism by Calmon et al. (FORC '25). It can be post-processed to output the same distribution as the Laplace mechanism or the Staircase mechanism with identical privacy parameters. Thus, the discrete Laplace mechanism is a versatile mechanism that should be preferred over the Laplace and Staircase mechanisms whenever the data is discrete (or can be made discrete while controlling \ell_1-sensitivity). We show bounds on the variance of our estimator, compared to the mean square error of the biased estimator that simply evaluates the f on the output of the mechanism. Though our unbiased estimator has exponential running time for worst-case functions, we show that it can often be computed in linear or polynomial time for concrete, structured functions. We showcase the properties of our methods empirically with several use cases including profile and entropy estimation, as well as distributed/federated data analysis applications in which unbiasedness is key to accuracy.
PaperID: 1988, Poster
Abstract: Multi-expert ensemble frameworks have achieved strong performance on long-tailed visual recognition by training heterogeneous expert heads on a shared backbone, but their gradient coordination strategies are inherited from balanced multi-task learning, where resolving directional conflict is the canonical challenge. We show that head-class bias acts sequentially across three pipeline interfaces, namely loss geometry, gradient coordination, and inference fusion, forming a suppression cascade in which rare-class gradient energy can be erased before it influences model parameters, a failure mode that direction-focused analysis cannot detect. A controlled study reveals that direction-modifying and simple-aggregation solvers are statistically equivalent once rare-class source signal is amplified, while energy-balancing objectives erase this amplification entirely. To break this cascade, we propose Gale (Gradient-energy Allocation for Long-tail Ensembles), a framework that applies gradient-energy allocation at every cascade stage, amplifying tail-class source signal through frequency-calibrated loss geometry, preserving it through a non-penalizing coordination rule, and routing predictions through class-specific expert competence. On CIFAR-100-LT (ρ=100), Gale achieves 59.2% overall accuracy and 46.1% few-class accuracy across four seeds (σ=0.23%), surpassing the prior best by +6.2%, with consistent improvements on ImageNet-LT and iNaturalist 2018. Code is available at Supplementary Material.
PaperID: 1989, Poster
Abstract: Latent world models based on Joint-Embedding Predictive Architecture (JEPA) are deterministic by design. While successful in fully observable scenarios, this paradigm breaks down when past observations and actions lead to multiple plausible future possibilities, e.g., due to occlusion. We introduce EpicWorldModel, a framework to train stochastic JEPAs for environments and tasks with inherent uncertainty under partially observability. We jointly train the EpicWorldModel predictor with its latent representation space to directly predict multiple potential future states using a flow-matching objective, when the goal-relevant scene content is absent from the conditioning history. We show that flow predictive variance approximates an upper bound on Expected Information Gain and serves as a useful exploration guidance for planning. By incorporating this uncertainty signal into Cross-Entropy Method (CEM)-based planning, our approach balances goal-reaching with exploration of uncertain regions where occluded or unseen goals are most likely to be located. We demonstrate the effectiveness of EpicWorldModel through a series of latent planning experiments with the best performance among all baseline methods, showing up to 22% empirical improvement in success rate over LeWorldModel.
Abstract: Selecting a finite dictionary of observables whose span is Koopman-invariant is a central challenge in data-driven Koopman operator approximation. We address this problem by exploiting zero-block structure in Extended Dynamic Mode Decomposition (EDMD) matrices. We show that any sub-dictionary whose span is Koopman-invariant induces an exact zero block in the EDMD matrix, even for finite data. We then show that such blocks can be detected by applying PageRank to a row-normalized EDMD matrix constructed from a large initial dictionary. The theory extends to approximately invariant subspaces and yields stronger guarantees for personalized PageRank (PPR) when the seed observables lie inside the target block and reach all observables in that block. Combining EDMD concentration bounds with PageRank perturbation theory gives end-to-end detection guarantees with O(M^-1/2) finite-sample scaling and explicit constants. More generally, without assuming an invariant subspace exists, high PPR mass on a sub-dictionary controls discounted multi-step leakage from the seed observables. Numerical experiments on the Duffing oscillator, Van der Pol oscillator, Lorenz system, and a three-well Ramachandran potential suggest that the method identifies compact, interpretable dictionaries with accurate predictions.
PaperID: 1991, Poster
Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for building embodied agents that ground perception and language into action. However, most existing approaches rely on direct action prediction, lacking the ability to reason over long-horizon trajectories and evaluate their consequences, which limits performance in complex decision-making tasks. In this work, we introduce World-Value-Action (WAV) model, a unified framework that enables implicit planning in VLA systems. Rather than performing explicit trajectory optimization, WAV model learn a structured latent representation of future trajectories conditioned on visual observations and language instructions. A learned world model predicts future states, while a trajectory value function evaluates their long-horizon utility. Action generation is then formulated as inference in this latent space, where the model progressively concentrates probability mass on high-value and dynamically feasible trajectories. We provide a theoretical perspective showing that planning directly in action space suffers from an exponential decay in the probability of feasible trajectories as the horizon increases. In contrast, latent-space inference reshapes the search distribution toward feasible regions, enabling efficient long-horizon decision making. Extensive simulations and real-world experiments demonstrate that the WAV model consistently outperforms state-of-the-art methods, achieving significant improvements in task success rate, generalization ability, and robustness, especially in long-horizon and compositional scenarios.
Abstract: Existing guarantees for misspecified kernelized bandit optimization pay for misspecification through kernel complexity: in generic offline bounds, the misspecification level \varepsilon is multiplied by \sqrtd_\mathrmeff, where d_\mathrmeff is the kernel effective dimension, while in online regret bounds, the corresponding penalty is \sqrt\gamma_n\,n\varepsilon, where \gamma_n is the maximum information gain after n rounds of interaction. In this work, we show that, for a large class of kernels, the misspecification amplification can be reduced to logarithmic or polylogarithmic growth. In the offline setting, we first prove high-probability simple-regret bounds whose misspecification term is governed by a spectral Lebesgue constant. This yields logarithmic amplification for one-dimensional monotone spectra and polylogarithmic amplification for multivariate Fourier-diagonal product kernels. In the online setting, we modify a domain-splitting algorithm and prove a cumulative regret bound of \widetilde\mathcal O(\sqrt\gamma_n n+n\varepsilon) under mild localized eigendecay assumptions, removing the extra \sqrt\gamma_n factor from the misspecification term. The common principle is localization: spectral localization controls the Lebesgue constant of the offline approximation operator, while domain splitting implements the spatial analogue of this mechanism in the online setting, preventing local misspecification errors from being amplified globally.
Abstract: Inverse optimization (IO) seeks to infer the parameters of a decision-maker's objective from observed context-action data. We study noiseless IO, where demonstrations are generated by a ground-truth objective. We provide a high-probability \mathcalO(\fracdT) generalization bound for the induced action set, where d is the number of unknown parameters and T is the size of the training dataset. We strengthen these guarantees under additional conditions that ensure uniqueness of the chosen action, bringing our IO guarantees in line with best-arm identification results in the bandit literature. We further show that the \mathcalO(\fracdT) rate is tight over all consistent estimators considered here, and extend the result to both instantaneous and cumulative regret. Notably, the resulting regret lower bound matches the corresponding upper bounds in the adversarial setting, indicating that the stochastic IO setting is effectively adversarial for the class of estimators studied here. Finally, we propose a parameter-free algorithm with lower per-iteration complexity than generic solvers. Experiments validate the predicted rates and illustrate the tightness of our bounds.
PaperID: 1994, Poster
Abstract: Sparse attention is becoming increasingly important for serving large language models (LLMs) as generation lengths continue to grow. However, deploying and evaluating new sparse attention algorithms at scale remains highly engineering-intensive, slowing both human researchers and AI agents in exploring the sparse attention design. To address this challenge, we present Vortex, a system that combines a Python-embedded frontend language atop a page-centric tensor abstraction for expressing a broad range of sparse attention algorithms, with an efficient backend tightly integrated into modern LLM serving stacks. Vortex enables rapid prototyping, deployment, and evaluation of sparse attention algorithms, effectively translating their theoretical efficiency gains into real-world throughput improvements. As a result, Vortex substantially accelerates the design and iteration of sparse attention algorithms. Using Vortex, we further demonstrate AI-agent-driven exploration of the sparse attention, automatically generating and refining diverse algorithms that achieve strong accuracy-throughput trade-offs, with the best generated algorithms delivering up to 3.46× higher throughput than full attention while preserving accuracy.
PaperID: 1995, Poster
Abstract: Larger transformers can memorize more data, with a capacity (in bits) that scales roughly linearly with parameter count. Counterintuitively, larger models also memorize faster, requiring fewer repetitions of a sequence during training before it is perfectly memorized. We reproduce both memorization scaling laws in transformers trained on unstructured sequences, demonstrating that efficient memorization is not due to a larger model's superior ability to learn structure in data. Rapid memorization is due to a meta-memorization effect, where models trained to memorize a large set of random sequences learn to memorize novel sequences more quickly. The rate r(t) at which models memorize grows with training time t as a power-law r(t) ~ t^\beta with a fixed exponent \beta but an increasing upper cutoff across model scales. In one-layer models, we trace the origin of meta-memorization to structured attention heads that convey length-invariant sequence information to MLPs where memorized sequences are stored. Meta-memorization is explained using a simplified model of a dense associative memory where retrieval is constrained by an attentional bottleneck. Put together, our results suggest transformers dynamically acquire reusable mechanisms that enable highly efficient storage and retrieval of memories.
PaperID: 1996, Poster
Authors:
Zhaolu Kang, Shiyu Liu, Tailong Luo, Wei Zhang, Yingjie He, Guangyuan Dong, Siheng Wang, Liang He, Lei Wei, Jiaqi Su, Shuang Chen, Guansu Wang, Haoyu Ji, Qishi Zhan, Kaiyue ZhouAbstract: Video multimodal large language models (MLLMs) keep climbing video question answering benchmarks, yet shuffling the frames, masking the segment that supports the answer, or occluding the target object barely changes their predictions. The accuracy rests on appearance and language priors, not on the temporal evidence the question asks for. We trace this to the unit of post-training: rewards are computed on a single response to the original clip, so the model is never asked to behave consistently across views. We propose Behavior Pack Optimization (BPO), which replaces the single response with a behavior pack of outputs across counterfactual views chosen by question type, scored jointly. The pack reward asks for stability when the intervention is irrelevant, sensitivity when key evidence is removed, and abstention when no evidence remains. To keep this objective stable at small pack sizes, BPO uses an anchor-relative advantage: the response on the original view serves as a per-prompt reference instead of a group mean over mixed views. On TempCompass, MVBench, and NExT-QA, BPO improves the macro accuracy of Qwen2.5-VL-7B-Instruct by 4.7 pp, the temporal-hard subset by 7.8 pp, and abstention F1 by 20.0 pp over a budget-matched vanilla GRPO baseline from the same SFT checkpoint. The gains transfer to Video-MME, LongVideoBench, and to LLaVA-Video-7B; ablations confirm they follow the view sets, not the rollout count. We hope this pack-level perspective offers a useful starting point for the video MLLM and multimodal post-training community as the field moves toward evidence-grounded video reasoning.
Abstract: Understanding why trained Transformers generalize well is a fundamental problem in modern machine learning theory, and complexity-based generalization bounds provide a principled way to study this question. While existing norm-based bounds for Transformers remove the explicit polynomial dependence on the hidden dimension, they typically impose fixed norm constraints specified a priori and can exhibit unfavorable exponential dependence on depth. In this paper, we derive spectrum-adaptive post hoc generalization bounds for multi-layer Transformers. Under layerwise spectral norm control, the bounds are expressed in terms of layerwise Schatten quantities of the query-key, value, and feedforward weight matrices. Since the Schatten indices need not be fixed a priori and can instead be selected after training, separately for each matrix type and layer, the bounds adaptively trade off spectral complexity against the dimension- and depth-dependent factors according to the learned singular-value profiles. Empirical comparisons of BERT-adapted proxies for the leading complexity factors suggest that the proxies induced by our bounds grow more slowly with depth and hidden dimension than the corresponding norm-based proxies. Overall, our results provide a complexity-based perspective on how the spectral structure of trained Transformers is reflected in generalization analyses.
PaperID: 1998, Poster
Authors: Srihari Sridharan
Abstract: LLM benchmark scores are often read as claims about factual accuracy, evidence use, instruction following, or logical reasoning. We study this score-to-claim conversion as inference over a completed evaluation record. Given benchmark items, prompts, outputs, scores, and native fields, which capability sentence is supported, which stronger sentence remains outside the evidence, and would that boundary change model choice? We formalize a report-use procedure, \emphscore-to-claim inference: hold source units fixed, compare each claim-induced condition to benchmark-native controls or same-source invariance checks, and apply a declared finite-family support rule. In a fixed set of five open-weight model families, completed records support narrower sentences than endpoint readings alone suggest: supplied-context answer recovery remains separate from support attribution; local supplied-evidence verdict signals leave open-web provenance unmeasured; and SATBench exposes label-surface instability for SAT/UNSAT prediction. The output is a support boundary, a blocked strengthening, and the next measurement or model-selection sensitivity attached to the reported result.
Abstract: Self-evolving language-model agents must decide what to learn next and how to preserve what they have learned across iterations. Existing systems typically carry this cross-iteration knowledge as natural-language feedback, flat episodic memory, or implicit reinforcement signals, none of which cleanly supports a frozen weak backbone at inference time. This paper introduces volution), a framework that externalizes self-knowledge into a four-subgraph co-evolutionary knowledge graph. Its experience subgraph stores both teacher-written failure corrections and the learner's own past correct reasoning traces, which are retrieved as task-conditioned guidance for a frozen execution model. During evolution, the graph, a task-level search bandit, and a skill-level routing bandit are updated from the same reward stream, while the learner's backbone remains unchanged. We further provide structural analysis showing how append-only memory growth, bounded curriculum coverage, and task-filtered retrieval together support stable improvement of the retrieval substrate for frozen-learner evolution. Across nine benchmarks spanning mathematical reasoning, multi-hop and open-domain question answering, spatio-temporal analysis, financial numerical reasoning, medical multiple-choice, an open-world survival game, and web navigation, achieves strong performance against prompt-based frozen-backbone baselines. Ablations show that self-harvested success traces and teacher-written corrections are complementary, with success memories contributing most on reasoning-template-heavy tasks and corrective memories supporting harder composition and interaction settings.
PaperID: 2000, Poster
Abstract: Due to context-length constraints, most MLLMs cannot process full-length videos and therefore rely on sampling a subset of frames as input. However, existing sampling methods, ranging from uniform sampling to relevance-based selection, are often driven by a single sampling principle and thus struggle to accommodate the heterogeneous evidential requirements of different queries: some demand a holistic understanding of how the video unfolds over time, while others hinge on fine-grained events within a short temporal window. To address this limitation, we propose Hierarchical Adaptive Frame Sampling (HAS), a two-stage frame sampling framework. In the first stage, Backbone Frame Construction, we apply a Determinantal Point Process (DPP) to sample frames that are both query-relevant and non-redundant. These selected frames capture key moments across the video and form the backbone of the sampling set, providing a foundation for subsequent enrichment. In the second stage, Adaptive Contextual Enrichment, we analyze the temporal distribution of the backbone frames to infer the query type and adaptively allocate the remaining frame budget between Local and Global Context. The Local Context enriches the backbone with fine-grained temporal dynamics and short-range causal relations, whereas the Global Context connects temporally isolated evidence and provides a holistic view of the entire video to support global understanding. Through such a hierarchical two-stage process, HAS effectively addresses diverse query requirements. Incorporated into three leading MLLMs, it demonstrates consistently superior performance across the Video-MME, LongVideoBench, and MLVU benchmarks. We have released our code.
Abstract: AI models deployed in critical domains, such as AI safety research, may subtly sabotage our efforts due to misalignment. is a form of sabotage in which an AI intentionally underperforms on the task given to it. It is particularly pernicious on tasks, i.e. tasks which are hard to grade or require intuition. To understand sandbagging on fuzzy tasks, we introduce a novel AI control framework that considers sandbagging as an adversarial game between a blue team and a red team. The blue team uses a weak trusted model to construct a weak score against which they would train a strong, potentially sandbagging model to remove the sandbagging if it were present. The red team then tries to find model behaviors that are rated highly by the weak score, and thus might not be trained out, but actually correspond to poor performance. We test our framework on the task of writing experimental proposals for research questions from recent ML papers. We use a language model with access to the original paper as a proxy "ground-truth" scorer. Our red team discovers sandbagging behaviors using multi-objective evolutionary prompt optimization. We show that Opus 4.6 can write proposals that are worse according to the ground truth proxy than those of GPT-OSS-20B, while the weak scorer rates them as highly as the best proposals from Opus 4.6. To mitigate the threat of sandbagging, we propose an adversarial optimization algorithm for the blue team that discovers more robust prompts for the weak model. This algorithm produces a blue team prompt that our red team optimization fails to exploit.
PaperID: 2002, Poster
Authors: Ryan Farell
Abstract: Parallel Langevin samplers are usually run as K independent chains and then averaged. Independence is convenient, but it is not required by the unadjusted Langevin algorithm (ULA): each chain only needs a standard Gaussian noise marginal at each step. We use this freedom by coupling the same-step noises across chains. The coupling is simple: draw K iid Gaussian noises, subtract their across-chain mean, and rescale. Each chain still has the ordinary ULA law, but the noise injected into the ensemble average is exactly zero. For quadratic targets, this removes the Langevin-noise contribution from every equal-weight linear summary of the ensemble mean at every finite horizon. With deterministic or zero-sum randomized starts, these summaries have zero total variance. On UCI Bayesian logistic posteriors with K=8, the same construction reduces ensemble-mean trace variance to 6× 10^-3, 2× 10^-3, and 5× 10^-4 of matched iid ensembles on WDBC, Spambase, and Adult; reference MSE against a longer iid run falls by roughly 4 to 14× after accounting for error in the finite reference run. The gain is scoped: zero-sum coupling cancels linear fluctuations, while centered quadratic observables have factor 1/(K-1) rather than iid's 1/K, and random starts, minibatches, and nonquadratic curvature add explicit residual variance terms. Synthetic and real-data experiments match the exact cancellation where it applies and the predicted behavior outside that regime.
Abstract: We develop a quantitative statistical theory of transformers in the large-context regime by adopting the abstraction of contextual flow maps: dynamical systems that evolve a distinguished token in the presence of a contextual measure across a stack of attention blocks. Within this framework, the finite-context model approximates an idealized infinite-context system in which the contextual measure is replaced by its underlying population, so that the context length n becomes a statistical resource. Exploiting the McKean--Vlasov structure of the dynamics and the classical machinery of propagation of chaos, we establish a forward bound controlling the deviation between the finite- and infinite-context flow maps uniformly along depth, and a backward bound controlling the deviation between the corresponding training trajectories uniformly across iterations of online gradient descent. Both bounds achieve the optimal Wasserstein rate n^-1/d. The analysis rests on a new Eulerian adjoint formulation of the loss gradient and stability estimates for the resulting forward--adjoint system, both of which may be of independent interest. We further verify that the standard transformer architecture satisfies these estimates with Lipschitz constants independent of embedding and parameter dimensions, suggesting a structural reason why transformers train stably across model scales.
PaperID: 2004, Poster
Abstract: Self-supervised multimodal representation learning has achieved remarkable success across diverse domains, yet capturing synergistic information remains challenging due to the complexity of cross-modal interactions. Unlike the shared information across individual modalities, synergy arises when task-relevant signals emerge only from the joint configuration of multiple modalities and cannot be recovered from any modality in isolation. This work focuses on how to preserve the information capacity for such synergistic signals in multimodal representations. The key observation is that synergistic information is reflected in higher-order statistical dependence among modalities, which provides a principled target for explicitly modeling joint interactions. Motivated by this insight, we propose Higher-order Representation and Information Learning (HRIL), which constructs an empirical cross-moment tensor over modality embeddings to represent multi-way interactions. HRIL employs Tucker decomposition to disentangle non-separable interactions into a core tensor, complemented by a synergy-aware regularizer that prevents energy concentration and preserves higher-order coupling for synergistic information capture. Experiments on synthetic and real-world benchmarks demonstrate consistent improvements over existing multimodal contrastive methods, with notable gains on tasks dominated by synergistic interactions.
PaperID: 2005, Poster
Abstract: Options provide a natural form of behavioral abstraction in reinforcement learning, but discovering reusable options from trained neural policies remains difficult. A trained policy may contain many useful behaviors that can be reused in downstream problems, yet the number of candidate subpolicies that a neural network can encode grows exponentially with the number of hidden units. We propose a differentiable approach for extracting reusable options from neural policies. Our key observation is that for piecewise-linear neural networks, selecting a neural subprogram is equivalent to assigning each hidden unit one of three labels: inactive, active, or retained in the computation. This yields a mask space over neural subprograms that can be searched with gradient-based optimization. We further show that reusable options can require default input parameters: neuron masks determine what computation is reused, while input masks determine how that computation is called. Experiments in transfer settings with feedforward and recurrent policies show that the resulting options improve sample efficiency on downstream learning.
PaperID: 2006, Poster
Abstract: We study realizable regression with a finite or countably infinite function class \mathcalF. We are given D_n = \(X_t, Y_t)\_t=1^n, which is assumed to be i.i.d.\ with f^\star \in \mathcalF such that \mathbbE[Y_t|X_t] = f^\star(X_t). The goal is to estimate f^\star with \hatf under the absolute loss \mathbbE_X[|\hatf(X) - f^\star(X)|]. We propose a simple reduction that runs two instances of online learning algorithms: one minimizes a learned loss and the other minimizes the error induced by learning the loss. Our main result is a prior-dependent second-order bound: for any prior \pi over \mathcalF, the estimator achieves with probability at least 1 - \delta, \beginalign \mathbbE_X[|\hatf(X) - f^\star(X)|] \leq C\sqrt\frac\mathbbE_X[\sigma(X)]\,\Gamma(f^\star)n + \fracCB\,\Gamma(f^\star)n\,, \endalign where \sigma(X) = \mathbbE_Y[(Y - f^\star(X))^2 \mid X] is the conditional variance, \Gamma(f^\star) = \log(1/\delta) + \log((1+\log(Bn))/\pi(f^\star)), C > 0 is an absolute constant, and B is an upper bound on |Y| and |f(X)|. The bound adapts automatically to the noise level \mathbbE_X[\sigma(X)] and range B without any prior knowledge of it, achieving the fast rate O(B\Gamma(f^\star)/n) when \sigma(X) = 0 and recovering the slow rate O(B\sqrt\Gamma(f^\star)/n) in the worst case. The dependence on \mathcalF enters only through \log(1/\pi(f^\star)), enabling faster rates for structured \mathcalF or sparse f^\star and extending naturally to countably infinite classes via summable priors. For finite \mathcalF we further show that the same algorithm produces an empirically verifiable bound that allows the learner to evaluate the quality of \hatf from the data alone, without knowledge of \mathbbE_X[\sigma(X)]. Finally, we extend our reduction to general convex loss functions satisfying \ell(f(x), f(x)) = 0.
PaperID: 2007, Poster
Abstract: Self-supervised Vision Transformers learn strong image-level representations, but their patch-level representations can be unreliable for dense prediction. We study this mismatch and find that the class token concentrates global semantics on a small number of representative patches, while dense prediction errors often occur on low-attention regions. We provide a local gradient analysis showing that, along the [CLS] attention readout, image-level self-distillation gradients are weighted by [CLS]-to-patch attention mass. To provide masked patches with global semantics without changing the architecture, we propose Cross Distillation (CODI), which mixes a small amount of the teacher class-token distribution into each teacher patch-token target. CODI gives every masked patch an attention-independent semantic correction while preserving the standard self-distillation architecture. Experiments on dense and image-level benchmarks show that CODI improves both dense and image-level performance under aligned pretraining settings.
PaperID: 2008, Poster
Authors:
Szymon Płotka, Gizem Mert, Pedro R. A. S. Bassi, Wenxuan Li, Zongwei Zhou, Anna M Marcinkiewicz, Wiktoria Romańczyk, Jarosław B Ćwikła, Łukasz Struski, Ewa Szczurek, Jacek Tabor, Arkadiusz SitekAbstract: Medical Vision-Language Pre-training (VLP) leverages the semantic richness of radiology reports to learn 3D representations, yet current architectures remain fundamentally anatomy-agnostic. While clinical diagnostics rely on organ-specific context, standard encoders apply a uniform, dense parameterization to all volumetric patches, and existing Mixture-of-Experts (MoE) routers gate on isolated local features without explicit access to the organ composition of the scan or crop. We propose Anatomy-Activated Mixture-of-Experts (A^2MoE), a framework that incorporates anatomical inductive biases within the model's computational path. A^2MoE introduces two key components: (1) a Histogram Representation Router (HRR) that conditions expert selection on global visual context and a predicted anatomy-presence histogram, ensuring that parameters specialize in specific anatomical domains; and (2) a mask-free Organ-Query Multi-Scale Projector (OQMSP) that utilizes learnable queries to extract organ-level embeddings for fine-grained alignment without requiring manual annotations at inference. By routing at the patch level rather than the token level, A^2MoE reduces routing complexity by orders of magnitude while satisfying load balancing by construction. Pre-trained on over 60,000 multi-phase CT scans, A^2MoE achieves state-of-the-art performance across segmentation, classification, and detection benchmarks. Furthermore, we demonstrate that our anatomically-grounded inductive biases support robust generalization to MRI and ultrasound, establishing a scalable foundation for medical representation learning.
Abstract: We study the problem of computationally efficient proper agnostic learning of high-dimensional concept classes under the Gaussian distribution. In this setting, given i.i.d.\ labeled samples from an unknown distribution over \mathbbR^d × \\pm 1\ whose marginal on \mathbbR^d is Gaussian, the goal is to output a hypothesis from a target class \mathcalF whose 0-1 loss is within \varepsilon of that of the best classifier in \mathcalF. We give the first efficient proper agnostic learning algorithm for arbitrary Boolean functions of K halfspaces under Gaussian marginals. Our algorithm runs in time d^O(K^2 \log(1/\varepsilon)/\varepsilon^2) + (K/\varepsilon)^O(K^3/\varepsilon^2.5). Prior to our work, the only known algorithm for K \geq 2 was brute-force search, with runtime exponential in d. Moreover, the dependence of our runtime on the dimension d matches that of the best known \emphimproper learning algorithm, namely d^\widetildeO(K^2/\varepsilon^2). For the special case of a single halfspace (K=1), the best previous runtime was d^O(1/\varepsilon^4) + (1/\varepsilon)^O(1/\varepsilon^6).\ Our algorithm improves this to d^O(1/\varepsilon^2) + (1/\varepsilon)^O(1/\varepsilon^2.5). Once again, the dependence on d matches that of the best known improper algorithm, namely d^O(1/\varepsilon^2). Furthermore, the dependence of our runtime on the dimension d is essentially optimal in the statistical query model.
Abstract: Online platforms connect users with relevant products and services using ads. A key challenge is that a user's search query often leaves their true intent ambiguous. Typically, platforms passively predict relevance based on available signals and in some cases offer query refinements. The shift from traditional search to conversational AI provides a new approach. When a user's query is ambiguous, a Large Language Model (LLM) can proactively offer several clarifying follow-up prompts. In this paper we consider the following: what if some of these follow-up prompts can be suggestion slots'' be allocated? And, how does this new mechanism interact with the traditional ad auction that might follow? This paper introduces a formal model for designing and analyzing these interactive platforms. We use this model to investigate a critical engineering choice: whether it is better to jointly optimize the user interaction and the final ad auction, or to decouple them into separate mechanisms for the suggestion slots and another for the subsequent ad slot. We show that the VCG mechanism can be adopted to jointly optimize the sponsored suggestion and the ads that follow; while this mechanism is more complex, it achieves outcomes that are efficient and truthful. On the other hand, we prove that the simple-to-implement modular approach suffers from strategic inefficiency: its Price of Anarchy is unbounded, for various natural choices for the two mechanisms. While we show that this inefficiency can be mitigated through revenue redistribution, we argue that the resulting mechanism sacrifices the practical simplicity that motivated the modular design in the first place.
PaperID: 2011, Poster
Abstract: MLLM-based detection has shown promising end-to-end detection ability in recent years, but its generated detection results still lack reliable confidence for evaluation and interpretation. Since detection data itself do not provide human-annotated confidence labels, existing MLLM-based detection methods mostly obtain confidence implicitly from prompted self-assessment, generation probabilities, or external proxy scores. While these scores are available for evaluation, they are not explicitly connected with the actual detection quality. Inspired by conventional detectors that optimize confidence through classification, objectness, or quality estimation branches, we formulate confidence in MLLM-based detection as an explicit modeling target to be learned, optimized and calibrated. Specifically, we present ConfDet, a practical confidence framework that first uses discrete confidence tokens based supervised fine-tuning to enable stable and parsable confidence generation, then applies GRPO with detection and separation joint rewards to optimize confidence score separation and preserve detection quality, and finally performs condition-aware post-hoc calibration to align confidence with empirical accuracy. Experiments across different in-distribution and out-of-distribution benchmarks show that, ConfDet greatly improves AP through better confidence ranking and reduces confidence miscalibration across benchmarks, providing a reliable confidence for MLLM-based detection.
Authors:
Yifan Yang, Ziyang Gong, weiquan Huang, Qihao Yang, Ziwei Zhou, Zisu Huang, Yan Li, Xuemei Gao, Qi Dai, Bei Liu, Kai Qiu, Yuqing Yang, Dongdong Chen, Xue Yang, Chong LuoAbstract: Adapting language agents often means teaching them domain procedures: where to look, which tools to call, how to verify intermediate results, and how to format outputs for a grader. Existing context-level adaptation usually either hand-writes such procedures, generates them once, or lets skill artifacts grow through loosely controlled self-revision. We ask whether a skill can instead be trained as the compact external state of a frozen agent. SkillOpt is a harness-agnostic text-space optimizer for a single natural-language skill document. It runs the frozen target model on scored rollout batches, uses a separate optimizer model to reflect over success and failure minibatches, proposes structured add/delete/replace edits, ranks them under a textual learning-rate budget, and accepts a candidate only if it improves held-out selection performance. Rejected-update memory, slow cross-epoch guidance, and optimizer-side meta skill make the editing loop behave like bounded training while leaving deployment as a single best_skill.md. Across six direct-chat benchmarks covering QA, spreadsheets, documents, math, and embodied decision making, \ourmethod improves GPT--5.5 over no skill by 21.5 points on average, with positive gains on every task and the best measured result on five of six. Codex- and Claude-Code-style harness runs show that the same artifact remains useful in tool-backed execution. Ablations and transfer studies identify validation-gated, bounded edits and update memory as key to stable, portable skill learning.
PaperID: 2013, Poster
Abstract: Modern automatic-rigging predictors (UniRig, Make-It-Animatable, RigNet) place a skeleton on a 3D mesh in seconds, but the resulting skinning weights animate poorly. Linear-blend skinning (LBS) is non-linear in pose space, so noisy weights amplify into visible candy-wrapper artifacts---the surface twisting and contracting at joint bends---and volume collapse. ARAP-style remedies penalize these symptoms one pose at a time and cannot reach the cross-pose curvature that produces them. We propose \emphPose-Interpolation Smoothness (PIS), a self-supervised, training-free, test-time refiner: minimize the per-vertex gap between the deformation at the SO(3) midpoint of two random poses and the linear average of the two endpoint deformations. The midpoint is obtained by orthogonally projecting the arithmetic mean of the two rotations back onto the rotation group, sidestepping the matrix-logarithm gradient that breaks down near identity and antipodal poses. Because PIS only re-tunes the weights LBS already consumes, it slots into any LBS-based pipeline---generative-3D outputs, automatic-rigging predictions, and the LBS pathways of game engines and AR/VR renderers---without touching shaders. On Objaverse-LVIS Balanced-102 (98 valid models, 51 LVIS categories), PIS reduces fragility (per-vertex deformation variance under random poses), ARAP energy, Jacobian anisotropy, and area distortion by 35.5%, 39.3%, 19.8%, and 43.2% on average; 93.9% of models improve on all four simultaneously, Wilcoxon p<10^-10. On the worst-case slice---the top decile of UniRig-initialized fragility, where animation visibly tears---PIS reduces unstable vertices by up to 10.2×. Effective bone count and weight entropy are unchanged, and the optimization runs in ~30~s/model on a single H200 GPU.
PaperID: 2014, Poster
Abstract: Cell types are organized by taxonomic relationships that define lineages in the so called cell ontology. However, most existing foundation models typically ignore cell-type lineages encoded in this cell ontology. We introduce a biologically informed foundation model called scOntoFM, which embeds hierarchical cell ontology into representation learning. Using Lowest Common Ancestor (LCA) distances from the ontology graph, our framework pairs efficient offline triplet sampling with hierarchical cell-ontology learning. This approach jointly preserves local neighborhoods and shapes the global embedding geometry to reflect the ontology structure. Across diverse cell- and gene-level benchmarks, the model consistently improves performance, with especially strong gains in zero-shot settings. Our model also balances batch integration with biological structure preservation, providing complementary perspectives to existing embedding evaluation frameworks and supporting scalable biological discovery.
Abstract: We introduce a bandit framework for stochastic matching under the multinomial logit (MNL) choice model. In our setting, N agents on one side are assigned to K arms on the other side, where each arm stochastically selects an agent from its assigned pool according to unknown preferences and yields a corresponding reward over a horizon T. The objective is to minimize regret by maximizing the cumulative revenue from successful matches. A naive approach requires solving an NP-hard combinatorial optimization problem at every round, resulting in a prohibitive computational cost. To address this challenge, we propose batched algorithms that strategically limit the number of times matching assignments are updated to \Theta(\log\log T) over the entire horizon. By invoking expensive combinatorial optimization only on a vanishing fraction of rounds, our algorithms substantially reduce overall computational overhead while still achieving a regret bound of \widetilde\mathcalO(\sqrtT).
PaperID: 2016, Poster
Abstract: Stereo matchers estimate depth by comparing the left and right images from a stereo camera pair, but they assume that the two cameras are accurately calibrated and rectified. In real deployments, small physical shifts after calibration, caused by camera motion, mechanical drift, vibration, or installation changes, can break this assumption and degrade depth estimation. Online stereo rectification is therefore needed to correct the image pair during operation and recover a stereo geometry where standard stereo matching can work reliably. Recent online rectification methods mainly optimize row alignment, making the left and right views of the same scene point fall on the same image row. We show that this target, while necessary, is insufficient for deployable stereo depth estimation. A rectifier may make a small set of matched points look well row-aligned, but still distort the full images in ways that hurt stereo matching: useful image regions may be cut off, large blank areas may appear, objects may be stretched or squeezed, or image scale may change unevenly across the view. We propose VirtualRect, an online stereo rectification method designed for reliable real-world deployment. Rather than treating rectification as simply making the left and right image rows line up, this method estimates a rectified virtual stereo camera. This makes the rectified images and the camera parameters come from the same geometry, so the images used for matching and the model used for depth computation remain consistent. To improve stability, it further constrains the rectification behavior over the whole image and reduces the influence of geometrically unreliable matches. Experiments on multiple datasets show that, compared with the prior state-of-the-art online rectifier, VirtualRect reduces vertical flow error by 17% on Carla-Flowguided and 7% on Semi-Truck Highway.
Abstract: Action chunking improves exploration and accelerates value propagation in long-horizon reinforcement learning, but naively applying off-policy methods to the temporally extended action space at reduced decision frequency offsets these gains, leading to poor sample efficiency. Existing action chunking methods address these issues through computationally expensive critic-only approaches or by relying on offline data. We introduce SEAR, a sample-efficient off-policy algorithm that enables online reinforcement learning with action chunks by addressing both challenges. To handle high-dimensional action sequences, SEAR employs a causal transformer critic trained with multi-horizon targets that provide a training signal for every prefix of an action chunk, effectively increasing the useful gradients per sample. To restore decision frequency, SEAR operates with a receding horizon and random replanning, ensuring uniform state coverage while combining the fast value propagation of large chunks with the reactivity of small ones. SEAR outperforms state-of-the-art online reinforcement learning methods including SimbaV2 on challenging Metaworld manipulation tasks. Beyond online RL, applying SEAR to the offline-to-online method QC improves its performance on the demanding OGBench cube-triple tasks.
Abstract: Natural images follow a 1/f^2 spectral distribution: most signal energy lies in the low spatial frequencies, while the perceptually important structures such as textures and edges occupy sparse high-frequency bands. Pixel-space reconstruction objectives, however, treat all spatial errors uniformly, causing low frequencies to dominate the optimization signal and delaying the learning of fine-scale details. In this work, we identify this objective-level spectral imbalance as a key inefficiency in training pixel-space flow models. To address it, we propose a Focal Log-Frequency Loss (\floss), a spectrally balanced objective that equalizes the learning signal across frequencies, emphasizing high-frequency components that are otherwise underrepresented in pixel-space objectives. Building on this, we introduce a simple training strategy that combines frequency and pixel supervision: we first emphasize frequency-domain learning early to capture all frequencies, and then transition to standard pixel-space v-loss for spatial refinement. This balancing mitigates the low-frequency bias of pixel losses and aligns the training signal with the evolving needs of the model. Our approach is conceptually simple, requires no architectural changes, and acts as a drop-in replacement for flow matching losses. Across multiple model scales, it accelerates convergence by up to 40% while consistently improving FID and perceptual fidelity. We will release code and models.
PaperID: 2019, Poster
Authors:
Junyi Shen, Yasuo Kuniyoshi, Kohei NakajimaAbstract: High-dimensional robotic reinforcement learning (RL) is often bottlenecked by inefficient exploration in large, redundant actuator spaces. We introduce Devyn (Developing synergies), a lightweight framework that learns a structured low-dimensional action parameterization directly during policy optimization. Unlike methods that require precomputed synergies or structured exploration in the full action space, Devyn starts from a randomly initialized latent-to-action map and shapes it online through task interaction. The policy acts strictly in a reduced latent space, while a learnable synergy matrix decodes latent actions into actuator commands. To make this evolving decoder stable and useful for policy learning, Devyn combines slow decoder evolution with structural regularization, reducing latent-action semantic drift while encouraging compact coordination patterns. Across diverse continuous-control benchmarks, including high-dimensional humanoid and overactuated musculoskeletal systems, Devyn improves upon standard SAC and achieves competitive or better performance than recent methods designed for high-dimensional robotic control. Ablations show that the gains do not come from a generic latent-action bottleneck alone, but rather from the slow, structured evolution of the synergy matrix. Our results suggest that learning the action parameterization itself can be a simple and effective route to scaling RL for high-dimensional robots.
Abstract: Gaussian Process Variational Autoencoders (GPVAEs) effectively model sample dependencies via latent GP priors, but their inference remains computationally prohibitive at scale. Existing methods typically rely on inducing points, which are restricted to evaluations in the input domain and provide limited control over inductive bias. We introduce a scalable inference framework for GPVAEs that uses inducing variables defined as general linear functionals of the latent GPs, rather than point evaluations. Our formulation recovers the standard inducing-point GPVAE inference as a special case and yields a simplified reconstruction-minus-KL training objective that makes this connection explicit. More broadly, it provides a principled interface for incorporating operator-dependent representations and computational structures via the choice of inducing features. In particular, we instantiate this framework using Fourier features to capture global spectral correlations and B-spline features to induce efficient banded covariances. Across multiple tasks in representation learning, imputation, and conditional generation, our method offers a competitive accuracy-efficiency trade-off to existing approaches.
Abstract: We study a stochastic budget-allocation problem over K tasks. At each round t, the learner chooses an allocation X_t \in \Delta_K. Task k succeeds with probability F_k(X_t,k), where F_1,\dots,F_K are nondecreasing budget-to-success curves, and upon success yields a random reward with unknown mean \mu_k. The learner observes which tasks succeed, and observes a task's reward only upon success (censored semi-bandit feedback). This model captures, for instance, splitting payments across crowdsourcing workers or distributing bids across simultaneous auctions, and subsumes stochastic multi-armed bandits and semi-bandits. We design an optimism-based algorithm that operates under censored semi-bandit feedback. Our main result is an instance-dependent bound showing that in diminishing-returns regimes, the regret of this algorithm scales polylogarithmically with the horizon T without any ad hoc tuning. For general nondecreasing curves, we prove that the same algorithm (with the same tuning) achieves a worst-case regret upper bound of \tilde O(K\sqrtT). Finally, we establish a matching worst-case regret lower bound of \Omega(K\sqrtT) that holds even for full-feedback algorithms, highlighting the intrinsic hardness of our problem outside diminishing returns.
PaperID: 2022, Poster
Authors: QIN LIU, Chen Zhong, Fengshan Zhao
Abstract: Selecting a pretrained encoder for a downstream task without retraining or labels is a standard preprocessing step in transfer learning. The dominant label-free score, RankMe, was validated within architecture families and is widely applied across them. We show that its cross-architecture use is unsound: rankings are systematically driven by the penultimate-layer dimension rather than representational quality, and on heterogeneous model pools the score performs worse than uniform random selection. We give a random-matrix explanation establishing this failure as inherent to any rotation-invariant score, and propose Subspace-Corrected Ranking (SCR), a family of label-free corrections that breaks rotation invariance via alignment with a fixed reference encoder. Across three architecturally distinct pools and 23 aggregated task units, the proposed SSR-dagger reduces mean regret from +0.165 (RankMe) to +0.071 (95% CI [+0.045, +0.099], p < 1e-4). Paired with a within-family score in a two-stage pipeline, it matches this quality at one-third of the forward-pass cost. Five labelled architecture evaluations suffice to match the best fully label-free strategy; K=35 labelled evaluations reduce regret to +0.018 --- a 5.7x improvement at the same forward-pass cost.
PaperID: 2023, Poster
Abstract: The Feed-Forward Network (FFN) dominates the activation memory of modern Transformers because the intermediate tensor of width d ff is materialized in HBM between the two matmul stages. FlashAttention-style I/O-aware tiling avoids the analogous cost in attention, yet has not been brought to FFN. We identify the reason as a hardware-geometry constraint: under one-pass tiled accumulation a fused FFN kernel must keep an output accumulator of width d model >= 2048 on currently shipping datacenter GPUs. We formalize this as an SRAM-feasibility proposition and prove a corollary: multi-head decomposition is the simplest architectural transform that brings FFN inside this SRAM-feasible regime, because it makes the per-head accumulator width d model/H a free parameter that can be driven below the GPU's SRAM cutoff. Naive multi-head FFN, however, blows up the per-head expansion ratio d model grows, drifting away from the well-validated SwiGLU 8/3 optimum and degrading quality at scale; we resolve this with a dense MoE-like sub-network correction that restores the optimal ratio while preserving SRAM-feasibility. The resulting architecture, FlashFFN, is an SRAM-feasible FFN block that replaces SwiGLU under matched parameter and FLOPs budgets and derives directly from the SRAM constraint. Across 128M--1.3B parameter models trained on 60--100B tokens of The Pile, FlashFFN achieves a Pareto improvement over SwiGLU at fixed compute: up to ~4x lower peak activation memory at long sequences, inference latency that matches or improves over SwiGLU (up to ~1.32x faster on H100), and consistent quality gains (up to -0.84 in evaluation perplexity and +1.60 average accuracy across six downstream benchmarks at 1.3B)
Abstract: Distribution Matching Distillation (DMD) provides an effective distribution-level correction for few-step generation, while relying on an auxiliary fake-score network to track the evolving generative distribution. Recent work combines DMD-style objectives with flow-map generators to exploit both forward-divergence training and reverse-divergence correction. The fake-score estimator remains an additional component with memory and update overhead. In this work, we study whether this explicit tracker can be avoided when the generator itself has a flow-map structure. We propose Fake-Score-network-Free DMD (FSF-DMD), a DMD formulation for flow-map generators that replaces the auxiliary fake-score estimator with a generator-induced pseudo-velocity surrogate. The key observation is that the endpoint pseudo-velocity of a flow-map generator provides a tractable proxy for fake-velocity estimation, allowing the generator itself to supply the reverse-divergence signal. Building on this observation, we derive a practical objective, extend it with flow-map-consistent backward simulation, and introduce a self-teacher variant for training from scratch. In our ImageNet-1K 256×256 experiments, FSF-DMD improves flow-map baselines, reaches lower FID than the listed DMD2 comparisons in the flow-map-initialized setting, and remains effective under flow-matching initialization and training from scratch.
Authors: Leonard Papenmeier, Petru Tighineanu
Abstract: Multi-objective optimization aims to solve problems with competing objectives. Evaluating such problems is often slow or expensive, limiting the budget of evaluations. In many applications, historical data from related optimization tasks is available and can be leveraged via meta-learning to accelerate optimization. Bayesian optimization, as a promising technique for expensive black-box problems, has been extended independently to meta-learning and multi-objective optimization, but methods that simultaneously address both settings remain largely unexplored. We propose SMOG—a scalable and modular meta-learning model based on a multi-output Gaussian process—that explicitly learns correlations between objectives. SMOG builds a structured joint Gaussian process prior across meta- and target tasks and, after conditioning on metadata, yields a closed-form prior for the target task. This construction propagates metadata uncertainty into the target surrogate in a principled way. SMOG supports hierarchical, parallel training, achieving linear scaling with the number of meta-tasks. The resulting surrogate integrates seamlessly with standard multi-objective Bayesian optimization acquisition functions. We demonstrate that our method is consistently competitive, delivering strong data efficiency across representative benchmarks and applications.
PaperID: 2026, Poster
Abstract: Transfer learning seeks to improve efficiency in a target regression task by borrowing information from related external source dataset(s). Existing approaches achieve this by enforcing Euclidean proximity between the target and source regression parameters or, more recently, by encouraging their directional alignment through angle-based penalties. While such angle-based penalties mitigate the sensitivity of the transfer learning methods to the scale differences between source and target regression coefficients, they still rely on a point estimate of the source regression parameter and thus provide no direct mechanism to incorporate uncertainty in the source information. To that end, we propose a probabilistic framework for transfer learning that operates at the level of a \emphscale-invariant directional distribution of the source regression parameter estimate, rather than its point estimate, to enable robust, scale-invariant and uncertainty aware transfer under heterogeneous source information. Specifically, we project the sampling distribution of the source regression estimate onto the unit sphere, thereby extracting its scale-free directional distribution. Transfer from source to target is then induced through a novel \emphangular cone prior which shrinks the target regression parameter towards directions in which the source distribution assigns high probability, while allowing its magnitude to be learned exclusively from the target dataset. The proposed construction admits exact finite-sample characterizations of posterior angular concentration and posterior computation remains tractable via careful low-dimensional augmentation in the parameter space. Simulation studies and application on a benchmark hypertension prediction task based on the National Health and Nutrition Examination Survey (NHANES) datasets demonstrate that the proposed method yields marked gains over existing approaches in settings with substantial uncertainty in the source regression estimate and source--target scale mismatch.
PaperID: 2027, Poster
Authors:
Tong Ding, Caiwei Tian, Dongmin Bang, Ming Yang Lu, Sophia J Wagner, Andrew Zhang, Rowland Pettit, Long P Le, Faisal MahmoodAbstract: EHR foundation models create new opportunities for scalable clinical prediction across health systems, yet their deployment remains constrained by a basic tokenization bottleneck. The same clinical concept can appear as different ontology codes, local identifiers, synonyms, or free-text descriptions, while standard EHR tokenizers treat these variants as unrelated symbols. We introduce MEDIATOR, a discrete ontology-agnostic medical concept tokenizer that turns EHR harmonization into an open-form tokenization problem. Instead of relying on manual mappings, a fixed ontology vocabulary, or continuous event embeddings, MEDIATOR maps any clinical concept expressed as text into a shared vector-quantized codebook. MEDIATOR is trained with synonym supervision from paired concept descriptions, and further guided by enriched concept descriptions, empirical code co-occurrence, and curated relation graphs so that equivalent expressions collapse while non-equivalent concepts remain organized in a medically meaningful discrete space. The resulting vocabulary is compact, inspectable, reusable by standard EHR Transformers, and open to previously unseen code descriptions or text strings. Used as the token interface of an EHR Transformer, MEDIATOR improves mean AUPRC on 20 MIMIC-III tasks by 4.0% and improves zero-shot transfer to 14 eICU tasks by 8.8%, while supporting transferable attribution across hospitals.
Abstract: Current hierarchical attention methods, such as NSA and InfLLMv2, select the top-k relevant key-value (KV) blocks based on coarse attention scores and subsequently apply fine-grained softmax attention on the selected tokens. However, the top-k operation assumes the number of relevant tokens for any query is fixed and it precludes the gradient flow between the sparse and dense stages. In this work, we propose DashAttention (Differentiable and Adaptive Sparse Hierarchical Attention), which leverages the adaptively sparse \alpha-entmax transformation to select a variable number of blocks according to the current query in the first stage. This in turn provides a prior for the second-stage softmax attention, keeping the entire hierarchy fully differentiable. Contrary to other hierarchical attention methods, we show that DashAttention is non-dispersive, translating to better long-context modeling ability. Experiments with large language models (LLMs) show that DashAttention achieves comparable accuracy as full attention with 75% sparsity and a better Pareto frontier than NSA and InfLLMv2, especially in high-sparsity regimes. We also provide an efficient, GPU-aware implementation of DashAttention in Triton, which achieves a speedup of up to 3.3× over FlashAttention-3 at inference time. Overall, DashAttention offers a cost-effective strategy to model long contexts.
Abstract: Touch is fundamental to dexterous manipulation, yet most egocentric human data increasingly used for robot learning lacks tactile information. Directly collecting large-scale tactile data is challenging due to sensor limitations, while human video data is abundant, contact-rich, and easily scalable. This motivates a natural question: can tactile signals be inferred purely from vision? To address this, we introduce centric human videos. EgoTac is trained on a unified corpus of over 5.7M image-tactile pairs, covering both continuous force measurements and binary contacts. By learning from this diverse dataset, EgoTac captures nuanced touch dynamics across varied interactions. Experiments demonstrate strong performance: in-domain prediction achieves an average force error below 0.06N. On out-of-domain contact prediction benchmarks, EgoTac consistently outperforms the state-of-the-art contact estimator. It also captures the rise and fall patterns of real tactile data and enables zero-shot predictions on unconstrained real-world videos. Scaling analyses further reveal that both data diversity and volume improve performance steadily. Overall, EgoTac provides a scalable pathway to extract tactile priors from egocentric human videos, enabling broadly applicable tactile-aware robot learning.
PaperID: 2030, Poster
Authors: Ziheng Chen
Abstract: Recently, deep neural networks on manifold-valued representations have garnered significant attention across various machine learning applications. One recent focus is the generalization of Euclidean fully connected (FC) and convolutional layers to non-Euclidean geometries. However, previous approaches typically focus on a few selected manifolds and rely on specific properties of the target manifold. In contrast, this work proposes a framework for constructing FC and convolutional layers over computationally tractable Riemannian spaces. This framework incorporates several previous FC layers across different geometries as special cases and is instantiated on ten representative manifolds, including three hyperbolic models, five geometries of the symmetric positive definite (SPD) manifold, and two Grassmannian perspectives. Experiments on different manifolds demonstrate the effectiveness and applicability of our approach.
PaperID: 2031, Poster
Abstract: Single-cell transcriptomics enables profiling of cellular states at unprecedented resolution, but its high dimensionality, sparsity, and technical batch effects pose significant challenges for representation learning. Existing single-cell foundation models typically encode each cell independently or only model cells from the same batch for denoising, thereby underutilizing the rich relational information across batches and cell types to model gene expression patterns. We argue that single-cell models can benefit from more informative cell-context modeling. By comparing consistency and variation across cells, models can capture fine-grained gene-gene dependencies associated with cell states, which are essential for learning high-quality representations. Inspired by the use of multiple sequence alignment (MSA) context in protein modeling, we propose CellMSA, a single-cell representation learning framework that introduces an MSA-inspired inductive bias into transcriptomic modeling. For each target cell, CellMSA retrieves relevant cells from different batches and biologically related cell types as context, and summarizes cross-cell patterns into a context-dependent gene-pair representation. This representation is then injected into a pair-aware target-cell encoder for fine-grained representation learning. We pretrain CellMSA on a large-scale human single-cell corpus of approximately 109 million cells. Experiments show that our framework consistently outperforms existing methods across multiple benchmarks.
Abstract: Fitting quantitative models to data is a central step in scientific workflows, yet it remains one of the least automated. Recent agent-based systems leverage language and vision-language models (VLMs) to iteratively propose and refine statistical models, but these systems struggle on more challenging modeling tasks. To address these limitations, we introduce VESTA (Visual Exploration with Statistical Tool Agents), a framework that equips VLMs with a dynamically growing exploration toolkit to guide model refinement through data transformations, hypothesis-driven visualizations, and robust statistical tests. Unlike prior systems that rely on iterative critique alone, VESTA actively explores data before and during refinement by selecting or creating diagnostic tools, which accumulate in a reusable registry across iterations. We evaluate VESTA against established baselines in three toolkit configurations: no tools, static expert-written tools, and dynamic model-written tools. To support this evaluation, we introduce DAWN (Dataset for Automated Workflows and Numerical Modeling), a benchmark targeting distribution fitting and time series modeling with varying difficulty tiers, and culminating in real-world astronomy tasks including modeling initial mass functions and gravitational-wave chirp signals. We find that VESTA's dynamic tool creation outperforms prior agentic pipelines, with the largest gains on complex and domain-specific tasks. We further show that dynamically generated tools are substantially more sophisticated than those produced by existing visual tool-creation systems, covering more diagnostic categories per function and strongly preferring visual outputs that the VLM critic can reason over directly.
PaperID: 2033, Poster
Authors: Lingcong Cai, Haiqin Yang, Jikui Liu, Yunpeng Cai, yumeng liu, Xiaomao Fan
Abstract: Tiny object detection (TOD) remains challenging due to the extremely limited spatial details of small objects. Existing methods heavily rely on multi-scale feature pyramids with discrete feature maps, which suffer from insufficient spatial resolution, quantization errors, and poor localization accuracy, while imposing heavy computational and memory overhead. We present the implicit DEtection TRansformer (iDETR), an efficient detector that operates exclusively on low-resolution single-scale features. It incorporates two novel components: (1) implicit Attention (iAttn), which leverages implicit neural representations to model continuous features and enables precise sub-pixel querying beyond discrete grid limitations; and (2) Centroid-Guided Query Initialization (CGQI) for robust query initialization under single-scale constraints. We further propose head-conditional sampling in iAttn, which reduces querying computational cost by 4× and memory footprint by 3× without sacrificing performance. Extensive experiments demonstrate the effectiveness of iDETR. Especially, on AI-TODv2, iDETR outperforms state-of-the-art methods by 0.2% AP overall, with particularly strong gains of 1.1% AP on very tiny objects, while reducing computational cost by 37%. The code is available at https://anonymous.4open.science/r/iDETR-FC25/.
Authors:
Siru Ouyang, Jun Yan, Yanfei Chen, Rujun Han, Zifeng Wang, Bhavana Dalvi Mishra, Rui Meng, Chun-Liang Li, Yizhu Jiao, Kaiwen Zha, Maohao Shen, Vishy Tirumalashetty, George Lee, Jiawei Han, Tomas Pfister, Chen-Yu LeeAbstract: LLM-based agents are increasingly deployed to handle streaming tasks, yet they often remain one-off problem solvers that fail to learn from past interactions. Reusable skills distilled from experience provide a natural substrate for self-evolution, where high-quality skill curation serves as the key bottleneck. Existing approaches either rely on manual skill curation, prescribe heuristic skill operations, or train for short-horizon skill operations. However, they still struggle to learn complex long-term curation policies from indirect and delayed feedback. To tackle this challenge, we propose SkillOS, an experience-driven RL training recipe for learning skill curation in self-evolving agents. SkillOS pairs a frozen agent executor that retrieves and applies skills with a trainable skill curator that updates an external SkillRepo from accumulated experience. To provide learning signals for curation, we design composite rewards and train on grouped task streams based on skill-relevant task dependencies, where earlier trajectories update the SkillRepo, and later related tasks evaluate these updates. Across multi-turn agentic tasks and single-turn reasoning tasks, SkillOS consistently outperforms memory-free and strong memory-based baselines in both effectiveness and efficiency, with the learned skill curator generalizing across different executor backbones and task domains. Further analyses show that the learned curator produces more targeted skill use, while the skills in SkillRepo evolve into more richly structured Markdown files that encode higher-level meta-skills over time.
PaperID: 2035, Poster
Authors: Göksenin Yüksel, Marcel A. J. van Gerven, Kiki van der Heijden
Abstract: Recent audio embedding models learn general-purpose representations that transfer across clip- and frame-level tasks, yet they operate almost exclusively on single-channel signals and are blind to spatial structure. Spatial audio models capture that structure but rely on clip-level supervision or contrastive objectives, yielding embeddings unsuitable for frame-level tasks like sound event detection and localization (SELD). We argue this dichotomy is not fundamental: patch-level spatial self-supervision can jointly learn acoustic and spatial features at both temporal granularities. We introduce AudioSphere, a multi-channel masked autoencoder trained on Ambisonics recordings that reconstructs both spectral content and directional cues via a joint spatiotemporal masking strategy that masks both time-frequency patches and intensity vectors. AudioSphere supports frame-level predictions while producing strong clip-level representations via aggregation. Alongside the model, we release RealSELD, a standardized benchmark for evaluating frozen backbones on real-world multi-channel SELD. AudioSphere enables, for the first time, SELD with a frozen self-supervised backbone, while matching or exceeding prior general-purpose audio embedding models on the HEAR benchmark.
PaperID: 2036, Poster
Abstract: Practitioners choose among LLM alignment methods largely by trial and error. We provide a theoretical foundation: each pairwise method defines a vector field on the probability simplex that governs how the model redistributes mass between preferred and dispreferred completions. Projecting onto a single preferred/dispreferred pair reduces this to a one-dimensional margin ODE with a closed-form solution. These solutions reveal that DPO margins grow as O(\log t) while IPO margins converge exponentially to 1/(2\tau); that label noise induces a qualitative phase transition in DPO (unbounded \to bounded) but affects IPO only quantitatively; and that gradient concentration under asymmetric weighting governs trainability at scale. We confirm all three predictions across five model families spanning four benchmarks. The closed-form solutions also enable principled method design: we prove that adding \emphany label smoothing \varepsilon > 0 to SimPO creates a finite margin attractor at \Delta^ = \beta^-1\log((1-\varepsilon)/\varepsilon), while \varepsilon=0 (standard SimPO) has no attractor. We test this across three model scales (3.8B--9B) and confirm it: the theory-derived configuration (TPO-S, \varepsilon=0.05) achieves the highest mean IFEval score on all three models.
PaperID: 2037, Poster
Abstract: Kernelized attention provides a useful perspective for understanding self-attention as a similarity-based aggregation mechanism, but existing approaches often rely on fixed or stationary kernels, limiting their ability to adapt the similarity structure to complex token interactions. In this paper, we propose flow-based spectral kernel attention (FSKA), a framework for learning expressive nonstationary attention kernels. FSKA formulates kernelized attention through spectral density learning and parameterizes the resulting bivariate spectral density with normalizing flows. This design enables flexible and regularizable kernel learning while remaining compatible with scalable linear-attention computation. We further show that expressive kernel learning enables a simplified attention design with removed query--key projections and fixed orthogonal value projections, resulting in a more parameter-efficient architecture. Controlled experiments on classification, ablation, and scalability profiling show that FSKA achieves competitive or improved performance over existing attention mechanisms, while reducing projection-related trainable parameters by approximately 40%--47% compared with most baselines and maintaining favorable linear-attention memory scaling.
PaperID: 2038, Poster
Abstract: For the combinatorial graph alignment problem (GAP) --- finding the node correspondence that maximizes the number of common edges between two unlabeled graphs --- properly initialized FAQ remains a strong classical baseline, while existing GNN approaches struggle in the purely structural setting. We introduce a chaining procedure: a sequence of Folklore-type (2-FWL) GNNs in which each network is trained with cross-entropy after decoding the previous network's similarity matrix and ranking nodes by their current alignment quality. This non-differentiable ranking step injects discrete combinatorial feedback at every link; at inference, we iterate the final network and keep the candidate with highest observed nce. On sparse Erdos--Renyi graphs at noise level 0.25, chained FGNNs with FAQ post-processing reach 85% accuracy versus 13% for FAQ initialized from the convex relaxation. On correlated regular graphs, where constant-feature MPNNs collapse to uninformative similarities and FAQ's convex initialization is degenerate, chaining recovers non-trivial alignments among the methods we evaluate. On three real-world benchmarks (yeast PPI, coauthorship, and road networks), dataset-specific chained FGNNs provide modest gains over a properly initialized FAQ baseline.
PaperID: 2039, Poster
Abstract: Protein optimization remains a longstanding goal in life sciences. Existing machine learning–assisted directed evolution (MLDE) methods primarily rely on sequence-only features, overlooking the critical spatial constraints and co-evolutionary interactions encoded in protein structures. However, directly integrating structural information remains challenging due to the scarcity of reliable mutant structures. To address these issues, we propose , a novel structure-aware reinforcement learning framework for protein directed evolution. StructEvo employs a delta-structure fusion encoder to approximate mutant structure features via feature differences, enabling dynamic incorporation of spatial knowledge. The vast mutation space is then decomposed into manageable subspaces through a structure-aligned hierarchical action network, while a geometric constraint further stabilizes delta feature learning. Our approach outperforms prior state-of-the-art methods by 9.2% and 16.3% on two challenging optimization benchmarks, and further identifies an experimentally validated epistasis pattern in GFP, highlighting the importance of structural guidance for effective protein directed evolution.
PaperID: 2040, Poster
Abstract: We present a hypernetwork that generates the parameters of an implicit neural representation (INR) for a given signal in a single forward pass, without test-time optimization. Existing hypernetworks for this problem either ignore the chain-rule structure of the test-time optimization that produces an INR or, when they preserve it, instantiate the structure with MLPs whose parameter count grows quartically in target-network width. Our central contribution is to show that this parameter-count cost is structural to the MLP instantiation, not to the optimization-oriented principle of preserving chain-rule structure, and to resolve it by replacing the MLP body with attention. The resulting hypernetwork combines three design choices: an output-neuron-centric tokenization that aligns the token axis with each layer's output-neuron axis; an information-flow consistency rule that fixes the role of every cross-attention module from the representation its output should occupy; and a token-space iterative rollout that realizes the five quantities one optimization step on the target INR computes as five cross-attention modules with parameter count independent of target-network width. Across five 2D and 3D INR generation benchmarks, our method outperforms the state-of-the-art single-pass and optimization-oriented baselines; scales to target-network widths at which the MLP-based optimization-oriented paradigm exhausts GPU memory; and produces the highest result on NeRF generation without ground-truth INR weights supervision.
Abstract: Decoder-only Transformers compute attention over the KV cache of preceding tokens. Keys (and Values) are typically represented with the same dimensionality, regardless of its distance from the prediction target. In natural language, however, the next word is most strongly influenced by the immediately preceding tokens. We hypothesize that local and distant tokens impose asymmetric demands on representational capacity: local tokens are more critical for predicting immediate outputs and thus require richer representations, whereas distant tokens primarily serve as long-range memory, for which lower-dimensional representations may suffice. We formalize this idea as Distance-Adaptive Representation (DAR), implemented in a controlled setting that preserves full-dimensional representations within a local context window while assigning reduced-dimensional representations (e.g. 1/4 of the original dimensionality) to tokens beyond that window. Across multiple pretraining scales (70M to 410M parameters), as well as continued supervised fine-tuning on a 1B-scale model, this approach closely matches the performance of full-dimensional baselines. In contrast, uniformly reducing dimensionality across all token positions leads to worse performance. These results challenge the common assumption that key and value dimensionality should be uniform across token positions. Our findings suggest a new direction for designing attention architectures that adaptively allocate representational capacity across the sequence, enabling further reductions in KV cache during inference.
PaperID: 2042, Poster
Abstract: The transition of Large Language Models (LLMs) to autonomous reasoning agents introduces a critical vulnerability within the reasoning trace itself. We present Reasoning Poisoning, a framework demonstrating how adversaries can exploit an agent via model-directed social engineering by manipulating retrieved context to weaponize the model's Reinforcement Learning from Human Feedback (RLHF) alignment. Our central paradigm, Logic Hijacking, exploits generalized alignment constraints by introducing fictitious hazards that force the agent to actively eliminate legitimate targets and select the attacker's target as the only "valid" alternative. Evaluating six state-of-the-art, production-deployed reasoning models across 10 domains, we show that this attack fundamentally overrides standard logic, achieving Attack Success Rates (ASR) exceeding 83%. The vulnerability remains highly effective even when the attacker controls only 10% of the retrieved context and exhibits robust success regardless of the adversarial payload's position within the evidence window. Controlled baselines confirm that this steering is driven by adversarial constraints, not ordinary promotional bias. Furthermore, we find that agents are entirely unresponsive to simulated social proof (e.g., user upvotes), relying instead on stylistic and semantic cues that mimic their own alignment training. Ultimately, our findings reveal a troubling paradox: the very alignment mechanisms designed to make models helpful and safe can be exploited to seamlessly hijack the reasoning process.
PaperID: 2043, Poster
Authors: Zesen Cai, Lisi Mo, Yifan Fang, Yandong Yan, Keren Shi, Daming Shi, Ruiting Dai
Abstract: Missing modalities reduce observations to arbitrary subsets, challenging robust multimodal inference. Existing methods recover missing views or align cross-subset representations, implicitly equating feature completeness with decision accuracy—leaving subset-specific decision boundary variation structurally unaddressed. Empirical analysis reveals that misclassification patterns are highly subset-specific, stable across initializations, and concentrated near decision boundaries—a persistent boundary misalignment distinct from information attenuation, which we define as Subset-Conditioned Boundary Shift (SCBS). To address this, we propose Subset-Conditioned Boundary Compensation (SCBC): a modality-factorized violation memory accumulates historical records of subset-specific margin violations, and a subset-adaptive head converts them into targeted logit corrections at inference, with a direction-consistency loss enforcing alignment during training. We prove that these corrections, built entirely from forward-pass statistics without backpropagating through the correction path, satisfy the descent condition of the surrogate loss, to our knowledge the first such guarantee in incomplete multimodal learning. Across five benchmarks, SCBC achieves state-of-the-art-level performance, averaging a 3.4 Macro-F1 point gain under random missingness and peaking at 7.4 points on the IEMOCAP-4 v, a setting. Source code is available at https://anonymous.4open.science/r/SCBC_neurips-50B5.
Abstract: Policy gradient methods serve as a cornerstone of reinforcement learning (RL), yet their extension to safe RL, where policies must strictly satisfy safety constraints, remains challenging. While existing methods enforce constraints in every policy update, we demonstrate that this is unnecessarily conservative. Instead, each update only needs to progressively expand the feasible region while improving the value function. Our proposed algorithm, namely feasible policy optimization (FPO), simultaneously achieves both objectives by solving a region-wise policy optimization problem. Specifically, FPO maximizes the value function inside the feasible region and minimizes the feasibility function outside it. We prove that these two sub-problems share a common optimal solution, which is obtained based on a tight bound we derive on the constraint decay function. Extensive experiments on the Safety-Gymnasium benchmark show that FPO achieves excellent constraint satisfaction while maintaining competitive task performance, striking a favorable balance between safety and return compared to state-of-the-art safe RL algorithms.
Abstract: Semi-supervised anomaly detection (SSAD) has shown great promise by effectively leveraging limited labeled data. However, existing methods are typically structured around scoring individual points or simple pairs. Such point- or pair-centric view not only overlooks the contextual nature of anomalies, which are defined by their deviation from a collective group, but also fails to exploit the rich supervisory signals that can be generated from the combinatorial composition of sets. Consequently, such models struggle to exploit the high-order interactions within the data, which are critical for learning discriminative representations. To address these limitations, we propose SetAD, a novel framework that reframes SSAD as a set scoring problem. SetAD employs an attention-based set encoder trained via a graded learning objective, where the model learns to quantify the degree of anomalousness within an entire set. This approach directly models the complex group-level interactions that define anomalies. Furthermore, to enhance robustness and score calibration, we propose a context-calibrated anomaly scoring mechanism, which assesses a point's anomaly score by aggregating its normalized deviations relative to multiple randomly sampled reference contexts. Extensive experiments on 10 real-world datasets demonstrate that SetAD significantly outperforms state-of-the-art models.
Abstract: Online learning with delayed feedback typically assumes that the learner can track all pending rounds until their feedback arrives. In practice, tracking resources are finite, and feedback from rounds that cannot be tracked is permanently lost. In this paper, we study delayed online convex optimization (OCO) under a hard capacity constraint, where at most C pending rounds can be tracked at any time. To model delay information, we introduce a semi-clairvoyant model that refines the clairvoyant assumption from prior work: rather than requiring delays to be known at prediction time, the learner observes delay expirations online, consistent with the classical unconstrained delayed setting. Our approach proceeds via a reduction to a novel ``delayed and weighted" OCO problem, using a scheduler that randomizes which rounds are tracked and importance-weights the resulting observations. For the base problem, we propose and analyze Delayed-Weighted FTRL and its bandit analogue, obtaining regret bounds that characterize how weights interact with delayed feedback in both settings. Combining these bounds with our schedulers yields regret guarantees for capacity-constrained OCO under convex and strongly convex losses, for both first-order and bandit feedback. For first-order feedback, capacity C = \Omega(\log T) suffices to recover the standard delayed OCO rates up to logarithmic factors. For bandit feedback, the standard delayed BCO rates are instead modulated by (1 + \sigma_\textmax/C) factors, where \sigma_\textmax is the maximum number of pending observations. This allows the regret bound to degrade gracefully when C < \sigma_\textmax, while remaining sublinear.
PaperID: 2047, Poster
Authors:
Erik Jahn, Venkat Chandrasekaran, Frederick Eberhardt, Leonard SchulmanAbstract: Causal discovery methods produce data-driven hypotheses about causal relationships, which may become new scientific findings or inform decisions about downstream interventions. In such settings, false positive causal claims can be more consequential than false negatives. Yet few methods target false positive error control in causal discovery, and existing error guarantees are often asymptotic and require strong versions of the faithfulness assumption. We introduce the m-conservative PC algorithm and prove that it controls false discovery rate in the linear Gaussian setting under an extremely weak version of the faithfulness assumption. In our framework, false positive error for causal graphs is defined based on the conditional-independence information they entail, yielding a metric that penalizes both false adjacencies and incorrect edge orientations. The parameter m quantifies a trade-off between algorithmic complexity and the strength of the required faithfulness assumption. The m-conservative PC algorithm combines undirected graph learning, a Benjamini-Yekutieli-style procedure for conditional independence testing, and a conservative orientation step that marks unresolved or contradictory edge directions as ambiguous. Our algorithm produces valid outputs and retains error guarantees even in the finite-sample setting, despite possible random errors in the conditional independence tests.
PaperID: 2048, Poster
Abstract: Modeling an opponent’s intent is critical for effective decision-making in non-cooperative, competitive, and general-sum multi-agent reinforcement learning. Existing opponent modeling methods encode intent using an embedding derived from episode information chosen a priori, such as the opponent’s next action or a future environment state, and use this to guide the ego-agent’s behavior. These approaches assume that the chosen information is universally representative of intent; however, we show empirically that this is not the case as intentions are often task- and environment-dependent. To address this, we introduce a task-adaptive opponent modeling framework that learns a performance-driven mixture of multiple intent representations. We further introduce a new intention representation that maximizes mutual information with the ego-agent’s future returns, thereby capturing opponent information that is most directly relevant to performance. Our approach consistently matches or exceeds the performance of state-of-the-art baselines across diverse tasks and yields insights into when and why different opponent modeling strategies succeed.
PaperID: 2049, Poster
Abstract: Large language models have achieved remarkable capabilities by scaling model capacity and training data, yet many practical deployments still rely on smaller models with limited resources whose capabilities lag behind their larger counterparts trained with richer resources. This gap calls for efficient knowledge transfer from source models with richer resources to compact target models. While model merging provides an effective mechanism, most existing methods are based on the assumption that the source and target models are architecturally compatible, making them inapplicable to heterogeneous source and target pairs. Although a recent method based on optimal transport extends model merging to settings with different architectures, it remains limited by linear correspondence modeling, iterative transport optimization, and reliance on supervised adaptation after fusion for further performance gains. To address these limitations, we propose BDC-Merge, a framework for model merging across different architectures based on Brownian Distance Correlation (BDCorr). BDC-Merge uses a small calibration set to estimate dependencies between heterogeneous activations at the feature level and the layer level, and directly lifts these dependencies from activation space into fusion operators in weight space, enabling effective parameter fusion across different architectures without any gradient based optimization or training after fusion. Extensive experiments across four low resource language and two specialized knowledge benchmarks show that BDC-Merge consistently outperforms the state of the art baseline for model merging across different architectures and largely preserves the target model’s general capabilities.
PaperID: 2050, Poster
Abstract: Representation learning has substantially improved the efficiency of unconstrained continuous-state Markov Decision Processes (MDPs), but its extension to Constrained MDPs (CMDPs) is difficult because the transition representation must be learned while constraint violation is controlled. In this paper, we study representation learning for low-rank CMDPs with continuous state spaces, and propose REP-PD-TEG for the soft cumulative-violation setting. In the feature learning period, the algorithm interacts with the environment using a decaying-\epsilon-greedy mixture execution policy, and uses the collected samples to learn transition features by maximum likelihood. Based on the learned features, it constructs uncertainty bonuses and plans with an optimistic Lagrangian objective to obtain the return policy. We prove regret and soft-violation guarantees for the return policies, and quantify the influence of the decaying-\epsilon-greedy exploration. In particular, REP-PD-TEG achieves the optimal \(\widetilde O(\sqrt K)\) return policy bound, and gives a \(\widetilde O(K^3/4)\) overall sample bound when exploratory actions are charged. For the stricter hard cumulative violation, where feasibility must be verified episode by episode, we propose REP-PD-H-TEG. It plans under a bonus-corrected lower-confidence utility certificate with the learned representation, and performs a stabilized dual search to obtain the return policy, yielding regret and hard-violation bounds. Together, these results give the first theoretically guaranteed representation-learning method for low-rank continuous-state CMDPs with both soft cumulative guarantees and certified hard-violation control.
Abstract: We consider the offline imitation learning from observations (LfO), where expert demonstrations are scarce and contain only state observations, and the suboptimal policy is far from expert behavior. In this regime, many existing imitation learning approaches struggle to extract useful information from imperfect data since they impose strict support constraints and rely on brittle one-step models. To tackle this challenge, we propose mbedding (TGE) for offline LfO. TGE constructs a dense, smooth surrogate reward by using particle based entropy estimation to maximize the log-likelihood of expert trajectories in the latent space of a temporal diffusion model trained on offline suboptimal data. By leveraging the structured geometry of the learned diffusion embedding, TGE captures long-horizon temporal dynamics and effectively bridges the gap under severe support mismatch, ensuring a robust learning signal even when offline data is distributionally distinct from the expert. Empirically, the proposed approach consistently matches or outperforms prior offline LfO methods across a range of D4RL locomotion and manipulation benchmarks.
PaperID: 2052, Poster
Authors: Guy Azran, Sarah Keren
Abstract: Agents that perform state estimation using pretrained perception models, e.g., Vision-Language Models (VLMs), face the fundamental challenge that predictions are uncertain. Still, most approaches determinize these predictions into a single best-guess state and make decisions as if it were ground truth. This is brittle because when the guess is wrong, the agent's decisions may be unsafe or infeasible. At the other extreme, maintaining a full distribution over symbolic states is often intractable. We propose a middle path, drawing on the underlying probabilities the perception model already produces (e.g., a VLM's output-layer token distributions) to construct a θ-confidence Minimal Belief Set (MBSθ), a minimal subset of states whose cumulative probability meets a user-defined threshold. The minimality of an MBSθ ensures that the agent considers only the most likely states, enhancing robustness in downstream decision-making while maintaining computational tractability. Using a novel three-layered belief system, we efficiently approximate individual state probabilities from per-feature predictions and task-level constraints, on which we base our MBSθ construction, while maintaining only a lean factored representation of the belief. Our algorithm, PRaVaSS, takes advantage of the unique structure of this belief system to construct an MBSθ with correctness guarantees and polynomial complexity. Using a VLM as the backbone of our state estimator, we validate our method's theoretical properties, scalability with problem size, and improvement of downstream decision-making when used in a task planning pipeline.
Abstract: Diffusion and rectified flow (RF) models generate high-fidelity images and videos, but their iterative velocity-field evaluations are computationally expensive. Existing caching methods accelerate sampling by skipping timesteps, yet their coarse approximations introduce accumulated errors over long skip intervals and degrade quality under aggressive acceleration. We propose TACache (Trajectory-Aware Cache), a training-free acceleration framework following a skip-then-compensate paradigm. TACache performs an orthogonal decomposition of discrete velocity acceleration along the RF trajectory into parallel and orthogonal components, isolating the magnitude and directional sources of per-step approximation error. The framework operates in two stages: \emphoffline, cumulative variation thresholds on the magnitude and direction indicators yield the skipping schedule and bound how far each skipped interval may extend; \emphonline, at each skipped step the offline statistics are combined with the sample's historical orthogonal direction to reconstruct the skipped velocity without additional forward passes. Experiments on BAGEL, FLUX.1-dev, and Wan2.1-1.3B show that TACache achieves up to 4.14× speedup on text-to-image generation and 2.11× speedup on text-to-video generation, with consistent improvements over prior cache-based methods on all reference-based fidelity metrics. Code will be released soon.
Abstract: Variational inference (VI) is a central tool in modern machine learning, used to approximate an intractable target density by optimising over a tractable family of distributions. As the variational family cannot typically represent the target exactly, guarantees on the quality of the resulting approximation are crucial for understanding which of its properties VI can faithfully capture. Recent work has identified instances in which symmetries of the target and the variational family enable the recovery of certain statistics, even under model misspecification. However, these guarantees are inherently problem-specific and offer little insight into the fundamental mechanism by which symmetry forces statistic recovery. In this paper, we overcome this limitation by developing a general theory of symmetry- induced statistic recovery in variational inference. First, we characterise when variational minimisers inherit the symmetries of the target and establish conditions under which these pin down identifiable statistics. Second, we unify existing results by showing that previously known statistic recovery guarantees in location–scale families arise as special cases of our theory. Third, we apply our framework to distributions on the sphere to obtain novel guarantees for directional statistics in von Mises–Fisher families. Together, these results provide a modular blueprint for deriving new recovery guarantees for VI in a broad range of symmetry settings.
PaperID: 2055, Poster
Authors:
Yizhak Y Elboher, Avraham Raviv, Amihay Elboher, Zhouxing Shi, Hillel Kugler, Guy KatzAbstract: Formal verification of deep neural networks (DNNs) faces a fundamental scalability challenge: verifying robustness is NP-hard and state-of-the-art tools often time out on modern architectures. We propose PAreto VErification, or PAVE: by augmenting DNNs with early exits (EEs) and verifying exit-by-exit, we solve the vast majority of robustness queries---roughly 80%---in a small fraction of the total time. The key insight is that most inputs in a robustness neighborhood exit at an early layer, a provably tractable sub-problem, so verification terminates far before analyzing the full network. EEs can be added to any DNN without degrading accuracy (within 0.3% across all tested architectures). We formalize the novel robustness property for EE networks, prove it is Fixed-Parameter Tractable under a trace stability condition (which commonly holds in practice), and present a sound and complete verification algorithm. Two sound optimizations---break (early termination when the winner dominates) and continue (skipping the inner loop when no runner-up can win)---reduce SAFE-case verification time by up to 10 ×. Experiments on MNIST, CIFAR-10, and CIFAR-100 with networks from 600K to 33M parameters demonstrate that PAVE enables solving queries that were not solved otherwise, including formal verification of ResNet-18 on CIFAR-100.
PaperID: 2056, Poster
Abstract: Predicting the behavior of dynamical systems (DS) beyond the dynamical and parameter regimes observed in training is a pivotal and essentially unresolved problem in scientific ML. It is central to any good scientific theory, which we expect to be able to make predictions which are not covered by currently available data. Recent hierarchical and hyper-network guided approaches for DS reconstruction (DSR) enable training on many DS simultaneously, and revealed that extracted latent features are often related to crucial control parameters of the underlying DS that varied across the training corpus. However, true out-of-domain forecasting abilities of these models, e.g. across tipping points, remain limited, and fine-tuning, or even full model retraining, on time series from the new dynamical regime is usually required. Here we mathematically analyze the root of these limitations in previous model formulations and identify three core shortcomings rooted in a mismatch between structural assumptions of the reconstruction model and typical properties of physical systems. We propose a combination of remedies for these shortcomings, most importantly , and furthermore derive a closed-form bound on the reliable extrapolation range. We demonstrate empirically that our techniques allow for accurate zero-shot prediction into new dynamical regimes, outside the observed training regime, as, e.g., encountered across tipping points.
PaperID: 2057, Poster
Abstract: Transformers have in-context learning capabilities, where some known learning algorithms can be executed in the forward pass through the model. Recent work show that transformers can exactly perform Lloyd's algorithm for k-means clustering with n points in d dimensions with an embedding size d_\textsfemb = d+k (thus, requiring attention projection matrices of size (d+k)^2). In this work, we build upon this result in the following ways: First, we present an equally expressive but smaller transformer that executes Lloyd's algorithm with embedding size d_\textsfemb = (d + \lceil \log_2 k \rceil ). Next, we train these transformers to learn the clustering algorithms given a distribution of clustering tasks, and theoretically characterize and empirically validate the factors affecting the convergence and in-distribution generalization of stochastic gradients based learning algorithms. Finally, we probe the general clustering abilities of these learned algorithms (in the form of transformers), and try to understand situations where they succeed and fail.
PaperID: 2058, Poster
Abstract: Preference optimization has become a central approach for aligning large language models with human values, but its effectiveness depends heavily on the quality of preference annotations. In practice, preference data is often noisy due to annotation errors and ambiguity. Existing robust preference optimization methods primarily operate at the sequence level, implicitly treating entire responses as uniformly correct or incorrect. However, rejected responses may still contain informative reasoning steps, while preferred responses can include subtle errors. In this work, we propose a oken-Level reweighting (PORT), a fine-grained framework for robust alignment under noisy preferences. PORT performs token-level reweighting using the empirical cumulative distribution function (CDF) of token logits, yielding an efficient forward-pass-only proxy to gradient-norm penalization that selectively suppresses corrupted tokens without additional backward-pass computation. We provide theoretical analysis showing that PORT reduces the gradient bias from noisy labels by minimizing an upper bound on the bias term. Extensive experiments across diverse noise settings demonstrate that PORT consistently improves robustness and outperforms existing sequence-level baselines.
Authors:
Wenhao Huang, Qingwen Zeng, Qiyue Chen, Zijie Guo, Yu Sun, Cheng Yang, Siru Ouyang, Jiri Gesi, Fang Wu, Jiayi Zhang, Huaming Chen, Bang Liu, Robert Tang, Chenglin WuAbstract: Large language model (LLM) agents often rely on long sequences of low-level textual actions, resulting in large effective decision horizons and high inference cost. While prior work has focused on improving inference efficiency through system-level optimizations or prompt engineering, we argue that a key bottleneck lies in the representation of the action space itself. We propose Latent Action Reparameterization (LAR), a framework that learns a compact latent action space in which each latent action corresponds to a multi-step semantic behavior. By reparameterizing agent actions into latent units, LAR enables decision making over a shorter effective horizon while preserving the expressiveness of the original action space. Unlike hand-crafted macros or hierarchical controllers, latent actions are learned from agent trajectories and integrated directly into the model, allowing both planning and execution to operate over abstract action representations. Across a range of LLM-based agent benchmarks, LAR significantly reduces the effective action horizon and improves inference efficiency under fixed compute budgets. As a consequence, our approach achieves substantial reductions in action tokens and corresponding wall-clock inference time, while maintaining or improving task success rates. These results suggest that action representation learning is a critical and underexplored factor in scaling efficient LLM agent inference, complementary to advances in model architecture and hardware.
PaperID: 2060, Poster
Authors: Hanchen Qiu, Haojia Zhu, Yifan Meng, Jiahui Jin
Abstract: Tool documentation is the critical interface bridging Large Language Model (LLM) agents and external environments. However, real-world documentation is often ambiguous or insufficient, leading to misaligned tool calls and task failures. While existing methods rely on parameter-efficient tuning or reinforcement learning to enhance tool-use capabilities, they incur substantial computational overhead and lack the agility to adapt to frequently updated toolkits. We propose OpTool , an evolutionary framework that treats tool documentation as an evolvable configuration. OpTool consists of three stages: (1) Contrastive Trajectory Generation, which explores diverse tool-use trajectories via contrastive beam search; (2) Trajectory-based Insight Extraction, which extracts tool-use insights through retrospective global credit assignment; and (3) Tool Documentation Evolution, which iteratively refines the documentation to mitigate information insufficiency. Experimental results across multiple domains demonstrate that OpTool significantly improves tool call accuracy and task success rates, providing a robust solution for tool-based LLM agents.
Abstract: We study fixed-confidence best-action identification (BAI) in stochastic minimax trees. This problem is increasingly relevant in modern AI planning, where deep minimax search and Monte Carlo Tree Search (MCTS) with language model long rollouts face a fundamental tradeoff: heuristic evaluations are cheap but biased, while accurate rollouts are reliable but prohibitively expensive. We propose 2FFS, a two-fidelity tree-search algorithm that brings multi-fidelity flat bandit ideas into trees. The algorithm combines minimax-style fast expansion with MCTS-style stochastic sampling, adaptively deciding when to exploit cheap biased evaluations and when to invoke expensive accurate evaluations for local certification. We prove fixed-confidence correctness, establish finite stopping for exact identification, and give a polynomial-depth cost upper bound for general-depth trees. Across numerical stochastic-tree experiments, 2FFS uses substantially fewer samples and computational operations comparing to existing BAI-MCTS baseline.
PaperID: 2062, Poster
Abstract: By combining cloud-side reasoning with edge-side observation and execution, cloud-edge collaboration has emerged as a promising paradigm for mobile UI agents. However, existing approaches face a fundamental trade-off between cost and accuracy. On one hand, step-by-step cloud interaction ensures high accuracy by uploading UI screens at every step and leveraging the cloud's powerful reasoning capabilities, but this inevitably incurs significant privacy costs for edge data and high computational expenses on the cloud. On the other hand, one-shot cloud planning drastically reduces the demand for edge privacy data and cloud compute by generating a complete plan based solely on the initial UI screen. Yet, this often leads to a substantial drop in accuracy, as the cloud agent may rely on invalid environment assumptions such as unseen app capabilities or UI flows. In this paper, we identify that the root cause of this trade-off lies in the environment assumption gap: the cloud agent lacks prior knowledge of the environment information required for edge-side task execution, and is thus forced to choose between planning with real-time information and planning based on speculative assumptions. To tackle this dilemma, we propose SeaPilot, which proactively provides the cloud agent with the necessary environmental context at the initial task submission stage, thereby achieving the dual benefits of both step-by-step and one-shot methods. To realize this, SeaPilot employs an iterative self-refining mechanism that progressively acquires feedback knowledge from execution failures across diverse tasks, enabling it to learn how to supply precise environmental information for each task in advance. Empirically, SeaPilot improves accuracy by 46.7% over one-shot methods, reduces cloud-token usage by 23.6x compared with step-by-step methods, and reduces privacy cost by 89.6%, achieving a better cost--accuracy trade-off. The code is available at https://anonymous.4open.science/r/SeaPilot-C13C.
Abstract: Causal discovery methods commonly assume that all data is independently and identically distributed (i.i.d.) and that there are no unmeasured variables affecting the system. In practice, these assumptions are often violated, leading to inaccurate inference. In this paper, we study how to identify hidden confounding and selection biases from causal mechanism shifts. In particular, we show that structural biases lead to dependent mechanism shifts. That is, by considering for which variables the mechanisms change given data from different environments, we can tell which variables are unbiased, which are subject to hidden confounding, and which are undergoing selection bias. We formalize this into an empirically testable criterion based on mutual information, and show under which conditions it identifies structural biases. To tell which nodes are subject to what kind of bias, we introduce the StruBI algorithm. Experiments on synthetic and real-world data show that StruBI works well in practice, accurately recovering affected variable sets and types of biases, outperforming the state-of-the-art by a wide margin.
PaperID: 2064, Poster
Authors: Shuning Zhao, Patrick Wong, Rongshang Li, Leran Zhang, Xiaolin Hu
Abstract: Generative models are increasingly used to simulate risk-sensitive systems, from financial markets and energy demand to healthcare, where downstream decisions depend on rare but consequential events. In this setting, matching average behavior is not enough: rare tail events drive stress testing, capital allocation, and policy analysis. Yet tail objectives such as Value-at-Risk and Expected Shortfall are difficult to optimize directly because empirical estimators rely on sorting, thresholding, and subset selection. We introduce TailDiff, a training-free inference-time guidance method for pretrained diffusion models that uses jointly elicitable Fissler-Ziegel scoring rules to construct differentiable tail-aware objectives. We analyze the induced guidance field locally in the low-noise regime, characterizing tail-shaping and boundary-coupling behavior together with controlled perturbations under small realized effective drift. On synthetic data, TailDiff approaches the oracle finite-sample floor, on a challenging financial tail-risk simulation benchmark, it outperforms the state-of-the-art on tail metrics while preserving bulk structure.
PaperID: 2065, Poster
Abstract: Diffusion and flow models enable powerful generative modeling, but controlled editing remains challenging due to the need to balance source fidelity and edit compliance along the generative trajectory. We propose a principled, inversion-free framework for flow-based image editing from a trajectory centric perspective. Rather than relying on inversion or heuristic guidance, we introduce a one-step look-ahead mechanism that evaluates the effect of candidate updates on both edit and fidelity rewards estimated via denoising, and adjusts the current update based on this evaluation. At each step, our method applies reward-driven corrections guided by the look-ahead evaluation under an explicit fidelity constraint, resulting in stable and coherent trajectory evolution. This look-ahead design mitigates gradient interference between edit and fidelity objectives by anticipating their interaction in future states, leading to improved trade-offs. Experiments demonstrate that the proposed approach achieves stronger semantic edits while better preserving source content compared to existing strong baselines.
PaperID: 2066, Poster
Abstract: Verifiable differential privacy enables public auditing of privacy-preserving aggregate computations, addressing the trust deficit in standard implementations. However, existing verifiable counting approaches that distribute trust across multiple provers often require each prover to contribute independent noise. This duplicates the noise added to the aggregate and substantially degrades utility, making accurate private counting difficult in distributed settings. We introduce a shared-noise paradigm for verifiable differentially private counting. Instead of duplicating noise across provers, our approach generates a single noise sample in additive-shared form among mutually distrustful provers, while Pedersen commitments let a public verifier audit consistency without learning the noise. This yields shared-noise, verifier-auditable instantiations of two standard noise families: the Shared-Noise Verifiable Binomial Mechanism (SN-VBM) and the Shared-Noise Verifiable Laplace Mechanism (SN-VLM). We prove that both mechanisms satisfy differential privacy and public verifiability, including completeness, soundness against malicious provers, and zero knowledge against a malicious verifier colluding with corrupted provers. Among the two instantiations, SN-VLM is the more scalable option: it avoids the quadratic growth in shared randomness required by the binomial construction under strict privacy budgets. Experiments on preprocessing load, randomness-generation time, and noise error show that SN-VLM substantially reduces cryptographic overhead while improving empirical accuracy in the evaluated privacy regimes.
Abstract: Machine learning models on graph-structured data are increasingly deployed in dynamic environments, where data arrive continuously and models must be updated online, while specific data points and their influence are requested to be removed to meet privacy or regulatory requirements. Most existing work on machine unlearning focuses on offline or post-training settings and does not readily extend to online graph learning: updates with arriving data and deletion requests are interleaved, and removing nodes or edges can propagate changes through the graph structure. We study online learning with certified unlearning over graph-structured data, where a model is updated sequentially while accommodating deletion requests at any time. The subsequent outputs need to be indistinguishable from those of a model trained without the removed data. We propose OLUG, a principled framework that performs localized second-order retroactive correction during online graph learning using only recent batch information. We show that OLUG preserves no-regret online learning while supporting certified removal, incurring only an additive regret cost under infrequent deletions. Our analysis further reveals how graph propagation amplifies deletion sensitivity differently across node, edge, and feature removal scenarios. Experiments on multiple benchmark graph datasets demonstrate utility comparable to training from scratch, with substantially reduced computational overhead by avoiding retraining.
Abstract: Universal channel decoders based on transformers--such as the Foundation Error Correction Code Transformer (FECCT)--achieve competitive decoding performance across diverse code families with a single shared backbone, optionally followed by code-specific finetuning. However, the high computational complexity and large parameter footprint of FECCT present substantial obstacles to practical deployment. To address these challenges, we investigate structured pruning for FECCT and propose Spectral-Aligned Pruning (SAP), a structure-aware framework that enables cross-code reuse of structured pruning masks by leveraging the spectrum of the corresponding bipartite graph. SAP is grounded in classical graph analysis of codes: the two algebraically largest adjacency eigenvalues provide compact spectral proxies for degree scale, expansion ratio, and minimum-distance lower bounds. These quantities are directly relevant to decoding performance: degree scale reflects how densely codeword bits and parity checks are connected; expansion ratio influences how information propagates across the bipartite graph; and minimum distance characterizes codeword separation. Based on this connection, SAP uses these two leading eigenvalues as a lightweight code signature for pruning-mask retrieval. Empirically, this two-dimensional signature yields stable library selection equivalent to higher-dimensional spectral signatures in our evaluation. After pruning, SAP performs per-code recovery via parameter-efficient low-rank adaptation (LoRA), enabling a shared pruned backbone while storing only small code-specific adapter parameters. Experiments across diverse codes show that SAP achieves decoding performance comparable to dedicated per-code pruning, while enabling substantial reductions in computational cost and model memory footprint through kernel-level structured pruning.
PaperID: 2069, Poster
Authors: Andrew ROSEMBERG, Joaquim D Garcia, Russell Bent, Pascal Van Hentenryck
Abstract: Optimization proxies amortize parametric constrained optimization by learning fast approximations of solution maps. Because these maps are smooth only within active-set regimes, the parameter space may split into combinatorially many critical regions, making finite-sample generalization and reliable constraint behavior central challenges. To address this challenge, this paper develops sparse stochastic masked Sobolev training for optimization proxies, augmenting value supervision with randomly selected solver-sensitivity targets rather than dense Jacobian matching. The mask acts as an algorithmic control: it exposes the proxy to local derivative information while avoiding overconstraint from conflicting active constraint set signals. We provide approximation bounds showing that, under coverage and regularity, joint value-and-Jacobian matching improves finite-sample generalization. Experiments characterize when and how this benefit appears. In smooth optimal-control distillation, where feasibility is guaranteed by construction, and only a few active-set regimes arise, Sobolev supervision reaches 95% of teacher performance about 2.5× faster than value-only training. In AC optimal power flow, masked Sobolev training keeps optimality gaps below 0.22% across three PGLib networks while reducing the median worst-case constraint violation by up to 4×. A mean--variance portfolio study identifies the complementary limitation: when sensitivities are highly local or vanish across active-set regimes, derivative supervision can be less useful than value-only training.
PaperID: 2070, Poster
Authors: Jon Arrizabalaga, Kevin Tracy, Zac Manchester
Abstract: Primal-dual interior-point methods solve constrained convex optimization problems to tight tolerances with speed and robustness. Their solutions are also efficiently differentiable with respect to the problem data through the implicit function theorem. However, the standard treatment of primal-dual complementarity makes the underlying linear systems increasingly ill-conditioned near the solution. While this ill-conditioning is often benign in double precision, it can be catastrophic in single precision, preventing interior-point methods from fully exploiting the accelerated hardware that underpins modern machine learning. This paper introduces a differentiable interior-point method designed for low-precision arithmetic. By using an alternative complementarity representation, we ensure that the underlying linear systems remain spectrally bounded --- even near the solution --- a property that is essential for computing accurate gradients and avoiding arithmetic exceptions. As a result, our method enables interior-point solvers to reliably solve and differentiate optimization problems in single precision that were previously confined to double precision. We demonstrate the approach through an ablation study against the standard interior-point formulation and applications in bilevel and end-to-end learning settings where differentiating through constrained optimization is essential.
PaperID: 2071, Poster
Authors: Bardiya Aryanfard, Monika Henzinger, Farhood Rostamkhani
Abstract: We study the classic least-squares linear regression problem under differential privacy (DP) in the continual observation setting. In this setting, the mechanism receives a stream of n data points (x_i,y_i) (x_i\in\mathbbR^d and y_i\in\mathbbR) and must output an estimate of the optimal parameters at each time step i. Although differentially private linear regression has been extensively studied in the batch setting, the problem remains underexplored for the continual setting, with prior work being restricted to constrained linear regression and update time polynomial in n. In this work, we introduce ContinualAdaSSP, the first mechanism for unconstrained linear regression under continual differential privacy. Our mechanism's suboptimality gap is proportional to \sqrtd\operatornamepolylog(n). It requires O(d^2\logn) space and O(d^2\logn+d^\omega) (\omega is the matrix multiplication exponent.) update time, where d is the dimension of the input space. This matches the suboptimality gap of the continual mechanism for the constrained version. We also study fundamental limits of differentially private linear regression by showing the first lower bounds for the box-constrained least-squares linear regression problem. Our lower bounds are proportional to \min(n,\frac\sqrtd\log^1.5d) for (1,O(\frac1n))-DP, and \min(n,\fracd\varepsilon) for \varepsilon-DP, and are derived using a reduction from a fingerprinting code problem from [Peter et al, 2024], providing new insights into the inherent hardness of this canonical task. Our experimental evaluation on real-world datasets shows that the accuracy of our mechanism clearly outperforms any other known differentially private mechanism for this problem such as rerunning the batch algorithm at each time step.
Authors:
Shuhong Liu, Xining Ge, quanfeng xu, Ziteng Cui, Liuzhuozheng Li, Gengjia Chang, Jun Liu, Ziying Gu, Dong Li, Xuangeng Chu, Lin Gu, Tatsuya HaradaAbstract: Ground-to-space astronomical super-resolution requires recovering space-quality images from ground-based observations that are simultaneously limited by pixel sampling resolution and atmospheric seeing, which imposes a stochastic, spatially varying PSF that cannot be resolved through upsampling alone. Existing methods rely on synthetic training pairs that fail to capture real atmospheric statistics and are prone to either over-smoothed reconstructions or hallucination sources with no physical counterpart in the observed sky. We propose FluxFlow, a conservative pixel-space flow-matching framework that incorporates observation uncertainty and source-region importance weights during training, and a training-free Wiener-regularised test-time correction to suppress hallucination sources while preserving recovered detail. We further construct the DESI--HST Dataset, the large-scale real-world benchmark comprising 19,500 real co-registered ground-to-space image pairs with real atmospheric PSF variation. Experiments demonstrate that FluxFlow consistently outperforms existing baseline methods in both photometric and scientific accuracy.
Authors:
Runyuan He, Qiuyang Mang, Shang Zhou, Kaiyuan Liu, Hanchen Li, Huanzhi Mao, Qizheng Zhang, Zerui Li, Bo Peng, Lufeng Cheng, Tianfu Fu, Yichuan Wang, Wenhao Chai, Jingbo Shang, Alex Dimakis, Joseph Gonzalez, Alvin CheungAbstract: Many real-world coding challenges are open-ended and admit no known optimal solution. Yet, recent progress in LLM coding has focused on well-defined tasks such as feature implementation, bug fixing, and competitive programming. Open-ended coding remains a weak spot for LLMs, largely because open-ended training problems are scarce and expensive to construct. Our goal is to synthesize open-ended coding problems at scale to train stronger LLM coders. We introduce FrontierSmith, an automated system for iteratively evolving open-ended problems from existing closed-ended coding tasks. Starting from competitive programming problems, FrontierSmith generates candidate open-ended variants by changing the problems’ goals, restricting outputs, and generalizing inputs. It then uses a quantitative idea divergence metric to select problems that elicit genuinely diverse approaches from different solvers. Agents then generate test cases and verifiers for the surviving candidates, and these problems then join the seed pool for the next round. On two open-ended coding benchmarks, training on our synthesized data yields substantial gains over the base models: Qwen3.5-9B improves by +8.82 score on FrontierCS and +306.36 (Elo-rating-based performance) on ALE-bench; Qwen3.5-27B improves by +12.12 and +309.12, respectively. Moreover, FrontierSmith-generated problems elicit long-horizon code-agent behavior comparable to human-curated open-ended tasks, causing agents to spend substantially more turns and tokens than on closed-ended problems.
PaperID: 2074, Poster
Abstract: Concept Bottleneck Models (CBMs) provide interpretability but typically incur a penalty in accuracy. Recent CBMs rely on LLM-supplied concept vocabularies that are misaligned with the underlying image evidence and obscure dataset-specific discriminative signals. Augmenting these representations with task-specific linear probes recovers little of the lost accuracy. Methods that instead learn concepts directly from data also fall short of the supervised CLIP linear-probe ceiling. We introduce MOSAIC (Mixtures Of Sparse Anchored Interpretable Concepts), a CBM whose concepts are discovered directly from image representations via text-anchored optimal transport, a balanced-marginal pass that softly assigns each image to a small set of learned concepts while keeping them tethered to class-text embeddings, and then named post-hoc by a VLM. Across 5 datasets, MOSAIC matches the supervised CLIP linear-probe ceiling on most dataset-backbone settings and beats raw zero-shot CLIP by +5.2 to +24.4 top-1 and class-level retrieval R@1 by +5.7 to +27.0. Beyond classification, the discrete concept basis enables a sparse per-image decomposition that exposes which features drive each prediction and aggressive bottleneck compression to as little as 2.5% of dense CLIP.
PaperID: 2075, Poster
Authors:
Youheng Yao, Ziyao Zeng, Wenbo Liao, Tianqi WangAbstract: Hyperbolic neural networks excel on hierarchical graph data, yet existing hyperbolic Transformers lift standard attention into curved space without using the ideal boundary, the structure most directly tied to tree hierarchy. We propose Horospherical Self-Attention (HSA): each query selects an ideal-boundary direction and scores keys by their Busemann depth along that direction, yielding a token-to-token hyperbolic attention kernel. We prove flat-limit consistency (HSA recovers dot-product attention as K\to 0^-), characterise score level sets as horospheres, and bound key-side gradients. HSA is integrated into Busemannformer, a full hyperbolic graph Transformer. Busemannformer-LD reaches 91.78% F1 on Disease-NC; in a controlled score ablation, its log-depth Busemann score achieves the best Disease-NC result among tested scores, reaching 93.29% F1 and performance comparable to the current SOTA Hypformer result with a smaller training budget. For link prediction, we find a score-symmetry principle: Busemannformer-LD with a symmetric score achieves 90.71% AUC on PubMed-LP (+10.3 pp over the asymmetric variant), while the asymmetric score performs best on Disease-LP among our variants and outperforms all Euclidean baselines. Ablations identify log-depth Busemann scoring as the key component and reveal a clear failure mode on Airport, where labels track centrality rather than tree depth.
PaperID: 2076, Poster
Authors:
Fengkai Yang, Zherui Chen, Xiaohan Wang, Xiaodong Lu, Jiajun Chai, Guojun Yin, Wei Lin, Fuzhen Zhuang, Shuai Ma, deqing wang, Yaodong Yang, Yikun BanAbstract: Reinforcement Learning from Verifier Rewards (RLVR) has emerged as a widely used approach for post-training large language models on reasoning tasks, with group-based methods such as GRPO and its variants gaining broad adoption. These methods rely on group-relative advantage estimation to avoid learned critics, yet its theoretical properties remain poorly understood. In this work, we uncover a fundamental issue of group-based RL: the group-relative advantage estimator is inherently biased relative to the true (expected) advantage. We provide the first theoretical analysis showing that it systematically underestimates advantages for hard prompts and overestimates them for easy prompts, leading to imbalanced exploration and exploitation. To address this issue, we propose History-Aware Adaptive Difficulty Weighting (HA-DW), an adaptive reweighting scheme that adjusts advantage estimates based on an evolving difficulty anchor and training dynamics. Both theoretical analysis and experiments on five mathematical reasoning benchmarks demonstrate that HA-DW consistently improves performance when integrated into GRPO and its variants. Our results suggest that correcting biased advantage estimation is critical for robust and efficient RLVR training. The code is available at https://anonymous.4open.science/r/HA-DW-037C.
Abstract: Instructional video editing has matured rapidly on short, single-shot clips, yet long videos, the dominant form of video in practice, remain largely unaddressed. Long videos exhibit failure modes that short-video editing does not stress-test: edit targets recur across temporally distant shots under varying viewpoint, scene, and scale, demanding cross-shot semantic consistency, while non-target intervals must be left unchanged, demanding accurate temporal grounding of the instruction. We present a hierarchical framework that addresses these two axes with dedicated components: a planning agent that grounds the instruction and partitions the timeline into shot-aligned segments, a long-term anchor bank that maintains a sparse set of high-quality cross-shot anchors through critic-driven iterative refinement on top of a strong image editing prior, and an anchor-conditioned video editor obtained by lightly adapting a pretrained short-video editor through LoRA, with no architectural changes. The editor is trained exclusively on short-video pairs and lifted to long videos at inference through its conditioning interface, requiring no long-video supervision. To evaluate this regime, we introduce LVEdit-Bench, the first benchmark designed for long-video instructional editing in the wild, together with a three-tier evaluation protocol that isolates per-shot fidelity, cross-shot identity preservation, and robustness to target-absent content under a structured VLM-judge protocol. Across all three settings, our method matches strong short-video editors on per-segment quality and substantially improves cross-shot consistency and grounding, with the margin widening as temporal complexity grows.
Abstract: Tackling the task of materials generation, we aim to enhance the previously proposed All-atom Diffusion Transformer (ADiT) by introducing SymADiT, a symmetry-aware variant. To do so, we use a representation of materials based on Wyckoff positions. We follow ADiT and perform generative modelling in latent space, adapted to our symmetry-aware representation. By forcing the output of the generative model to adhere to the symmetry restrictions imposed by the generated crystal's space group and each atom's Wyckoff-position, the generated materials exhibit more realistic symmetry properties. We benchmark our method against both symmetry-aware and symmetry-agnostic models for materials generation and show competitive performance, generating stable, symmetric materials with a simple Transformer architecture.
Authors: Kaiwen Shi, Carlos Oliver
Abstract: Protein structure tokenizers (PSTs) are workhorses in protein language modeling, function prediction, and evolutionary analysis. However, existing PSTs only capture local geometry of static structures, and miss the correlated motions and alternative conformational states revealed by protein ensembles. Here we introduce Ensembits, the first tokenizer of protein conformational ensembles. Ensembits address challenges inherent to tokenizing dynamics: deriving informative geometric descriptors across conformations, permutation-invariance encoding of variable-size ensembles, and conquering sparsity in dynamics data. Trained with a Residual VQ-VAE using a frame distillation objective on a large molecular dynamics corpus, Ensembits outperforms all related methods on RMSF prediction, and is the strongest standalone structural tokenizer on an token-conditioned ANOVA test on per-residue motion amplitude. Ensembits further matches or exceeds static tokenizers on EC, GO, binding site/affinity prediction, and zero-shot mutation-effect prediction despite using far less pretraining data. Notably, the distillation objective enables Ensembits to predict dynamics token from one single predicted structure, which alleviates dynamics data sparsity. As the field moves from static structure prediction toward ensemble generation, Ensembits offer the discrete vocabulary needed to bring dynamics into protein language modeling and design. Our codes can be found at: https://anonymous.4open.science/r/Ensembits.
Abstract: Plug-and-Play (PnP) methods have become standard tools for solving imaging inverse problems by replacing the intractable maximum a posteriori (MAP) denoiser with the MMSE one. While this mismatch has been widely treated as unavoidable, recent works have sought to close this gap by targeting the MAP with diffusion-model scores. We show this is problematic in practice: learned scores do not match the true ones, so MAP-targeting iterations converge to cartoon-like images rather than realistic ones, and better results are obtained by stopping short of convergence. We turn this observation into a design principle and introduce \textttProxiMAP, an iterative MAP approximation whose noise schedule keeps the iterate's residual noise matched to the denoiser's training noise. This keeps the denoiser in-distribution where its score is reliable, and yields implicit early stopping that avoids the failure mode above. \textttProxiMAP is a modular drop-in replacement for MMSE denoisers in standard PnP algorithms and consistently sharpens reconstructions across deblurring, inpainting, super-resolution, and phase retrieval. Building on the same principle, we propose a hybrid variant that applies \textttProxiMAP only in the late iterations of PnP, where the denoiser is most reliable—matching or exceeding the full-replacement variant at a fraction of the cost.
PaperID: 2081, Poster
Abstract: Cognitive neuroscience research indicates that human spatial navigation relies not merely on abstract global map representations, but rather on anchoring local observations to spatially stable, decision-relevant landmarks to effectively align the current scene with the cognitive map. Inspired by this mechanism, we propose SFPR, a structural fingerprint-based LiDAR-to-OpenStreetMap (OSM) place recognition framework. This framework goes beyond the traditional coarse retrieval paradigm that relies exclusively on global descriptor similarity and further introduces fine-grained discrimination based on structural fingerprints to achieve more accurate place recognition. Specifically, we propose a structural fingerprint-aware attention mechanism that utilizes sparse but spatially stable anchors as saliency prompts, directing the network's attention toward potential fingerprint regions while dynamically generating global descriptors. Subsequently, we formulate the extraction and matching of structural fingerprints as an alternating optimization problem guided by this fingerprint-aware attention. Through iterative optimization, we extract structural fingerprints with high fingerprint-aware attention and geometric consistency from anchors. These structural fingerprints are further fed back into the training stage as geometric priors, effectively suppressing false matches characterized by similar global features but contradictory local structures. Experiments demonstrate that SFPR effectively overcomes the retrieval bottleneck in macroscopically homogeneous scenes, achieving a relative improvement of 30.20% in Top-1 Recall@1m compared to the current state-of-the-art method. Code are publicly available at https://anonymous.4open.science/r/SFPR.
PaperID: 2082, Poster
Authors: Luis Antonio Ortega Andrés, Andres Masegosa, Thomas Nielsen
Abstract: Implicit-process priors define distributions over functions through flexible generative mechanisms, making them attractive for Bayesian function-space modelling. However, performing posterior inference with such priors is challenging because their induced function-space distributions are typically not available in closed form. One practical strategy is to approximate the prior using a finite collection of sampled functions, and then represent posterior functions as learned combinations of these samples. Existing approaches commonly place a Gaussian variational distribution over the combination weights. While tractable, this choice limits the shapes of posterior uncertainty that can be represented, especially when the true posterior is asymmetric, heavy-tailed, or multimodal. We propose Flow-Transformed Implicit Processes (FTIP), a variational inference method that makes this finite-dimensional function-space approximation more expressive. Instead of using a Gaussian distribution over the combination weights, FTIP uses a normalizing flow to define a richer variational distribution. This induces a flexible posterior distribution over functions while preserving tractable optimization. We train the model using a Black-Box \alpha objective, allowing us to compare mass-covering and mode-seeking variational behaviour. Experiments show that FTIP captures asymmetric and multimodal posterior structure in function space that Gaussian coefficient approximations tend to smooth or collapse.
Authors: Shikuri Yuta, Hironori Fujisawa
Abstract: Survival analysis provides statistical methods to model the time until an event occurs. Reporting delays arise when event times are not observed at their occurrence but are only revealed upon reporting. This issue is particularly critical for timely risk evaluation when the observation window is short due to administrative censoring. In this study, we incorporate right-censored reporting delays by jointly modeling parametric hazards for the event and reporting processes. We then construct a consistent estimator for the model parameters and develop a Monte Carlo expectation-maximization algorithm to compute it. To address the challenges posed by administrative censoring, we leverage these findings and propose a transfer-learning procedure. Experimental results demonstrate that our method improves the accuracy of timely risk evaluation under administrative censoring.
PaperID: 2084, Poster
Abstract: Selective knowledge distillation has become increasingly important for compressing large language models, where dense logit-based knowledge transfer over long sequences and large vocabularies incurs substantial computational overhead. Distillation supervision is naturally organized over a sequence–vocabulary matrix, with its highest-information components often concentrated on a small subset of specific sequence–vocabulary pairs. Existing methods typically select sequence positions or vocabulary candidates independently, relying on coarse-grained one-dimensional heuristics that weaken the precision and effectiveness of knowledge transfer. To address this, we propose Joint Sequence–Vocabulary Distillation (JSVD), a plug-and-play selective distillation framework that identifies and focuses supervision on informative sequence–vocabulary pairs. For each training example, JSVD constructs bidirectional teacher--student candidate supports, scores sequence--vocabulary pairs using prediction discrepancies, and progressively focuses distillation on unresolved informative pairs during training. Extensive experiments across diverse distillation objectives, model pairs, and benchmarks show that JSVD consistently improves downstream performance while reducing distillation overhead and accelerating convergence.
PaperID: 2085, Poster
Abstract: Modern visual backbones, despite their architectural diversity, largely reduce relational reasoning to pairwise message passing on graphs and therefore primarily capture 2-order interactions. However, visual understanding often depends on latent yet crucial higher-order semantic correlations among multiple regions, which cannot be adequately expressed by pairwise modeling alone. We first formalize this limitation by proving an irreducibility result: on a fixed vertex set, there exist nonlinear hypergraph message passing layers that cannot be represented exactly by any graph message passing layer unless the hypergraph satisfies highly restrictive structural conditions. This shows that higher-order correlation modeling offers a fundamental expressive advantage rather than a simple approximation to graph-based backbones. Therefore, we propose Vision Correlator (ViC), a general-purpose visual backbone built on multi-order hypergraphs. ViC consists of correlation induction and correlation propagation. In the correlation induction stage, we propose an Anchor-based Hypergraph Generation strategy that uses anchors as semantic centers and incorporates topological information to generate scalable hypergraph structures. In the correlation propagation stage, we further design Differential Mixed Aggregation to achieve efficient hyperedge representation and vertex feature updating, and introduce hyperedge dropout to improve the stability of structure learning. Extensive experiments with isotropic and pyramid variants show that our ViC achieves a better accuracy-efficiency trade-off than strong Transformer and graph-based baselines. Specifically, compared to the Vision Transformer baseline, ViC achieves up to 76% parameter reduction, 93% FLOPs reduction, and delivers up to a 5.1% increase in accuracy on ImageNet.
Abstract: Text-to-image flow matching transformers degrade sharply in long-tail settings: tail-class outputs collapse in fidelity and diversity, limiting their value as synthetic augmentation for rare conditions. We trace this to low head-versus-tail gradient alignment during fine-tuning, an optimization-level pathology that conditioning- and sampling-side interventions do not address. We propose GRASP (Guided Residual Adapters with Sample-wise Partitioning): a deterministic partition of the conditioning space, paired with group-specific residual adapters in the transformer feedforward layers, that leaves the flow-matching objective and the sampler untouched. In conditional flow matching, condition values index distinct sets of probability paths, so partitioning along the conditioning is the structurally correct factorization suitable as gradient alignment proxy. Because the partition is static, every tail sample is guaranteed to update its assigned expert, which bypasses extreme longtail failure modes. Crucially, GRASP is non-invasive and composable: on MIMIC-CXR-LT, combining GRASP with self-guided minority sampling at inference time yields the best all-labels IRS we observe, beyond either intervention alone. GRASP itself reduces overall FID by up to 80% and lifts tail-class coverage by up to 44% over full fine-tuning, learned-routing MoE, and minority guidance. Used as training data for a downstream DenseNet classifier on NIH-CXR-LT, GRASP synthetics significantly outperform every non-GRASP alternative on macro F1, match the macro F1 obtained from real training data, and yield nonzero F1 on 9 of 13 classes versus 3 of 13 from full fine-tuning. Results on ImageNet-LT confirm the mechanism is not tied to medical inductive bias.
PaperID: 2087, Poster
Abstract: Irregular data assimilation is difficult because sparse, off-grid observations do not constrain the same state corrections from one cycle to the next. As the observation geometry changes, the set of supported correction directions changes as well. Many learning-based DA methods infer updates in a fixed latent or control space, even when the current observation layout supports a different set of correction directions. To address this mismatch, we introduce Observability-Constrained DataAssimilation with Control-Space Inference (ObsConDA), which determines those supported correction directions from the current background and observation geometry and infers only their control coefficients. On real surface-station assimilation, ObsConDA gives the best held-out station accuracy and the best full-field reconstruction among the tested classical and learned baselines. The same advantage carries to matched synthetic observations, explicit geometry shifts, and synthetic dynamical systems, supporting the view that the usable correction directions must change from cycle to cycle. Ablations show that the gain comes from adapting those directions themselves rather than from adding a latent corrector alone.
PaperID: 2088, Poster
Abstract: Structured multiple-testing problems (gatekeeping trials, dose-finding, multi-tissue eQTL mapping, bundled-challenger A/B experiments) organize hypotheses into design-imposed blocks and demand strong family-wise error rate (FWER) control for confirmatory claims. Practitioners currently use objective-agnostic stepwise rules (Bonferroni, Holm, Hochberg, Hommel), closed-testing and graphical extensions, or hierarchical and resampling methods; none is power-optimal within the block-separable class these designs induce. We introduce BOOST (Block-Optimal Objective-driven Strong-FWER Testing), the power-optimal strong-FWER procedure for block size three, with three guarantees: (i) finite-sample strong-FWER validity at O(K) cost (versus O(K^2) for general closed testing) without independence assumptions, with a strict Šidák improvement under cross-block independence; (ii) power-optimal allocation across heterogeneous blocks via an equalized-marginal KKT condition, solvable by bisection in O(B\log(1/\varepsilon)); and (iii) a sample-split plug-in variant for unknown alternative density g, attaining \alpha-control up to O(B_\mathcalT \, \mathbbE\Vert g - \hatg\Vert_\infty) inflation with per-hypothesis power deficit independent of B_\mathcal T. Simulations across independent, equicorrelated, sparse, and mis-specified regimes show 1.4–1.7× power gains over the strongest existing baseline at calibrated FWER. On two published datasets (BLUEPRINT cross-lineage cis-eQTL and Upworthy bundled-challenger A/B experiments), BOOST certifies an order of magnitude more full-block discoveries than existing baselines at controlled FWER.
Abstract: Latent Action Models (LAMs) have emerged as an effective paradigm for handling heterogeneous datasets during Vision-Language-Action (VLA) model pretraining, offering a unified action space across embodiments. However, existing LAMs often rely on discrete quantization encode and decode pipelines, which can lead to trivial frame reconstruction behavior, limited representational capacity, and a lack of physically meaningful structure. We introduce RotVLA, a VLA framework built on a continuous rotational latent action representation. Latent actions are modeled as elements of \rm SO(n), providing continuity, compositionality, and structured geometry aligned with real-world action dynamics. A triplet frame learning framework further enforces meaningful temporal dynamics while avoiding degeneration. RotVLA consists of a VLM backbone and a flow-matching action head, pretrained on large-scale cross-embodiment robotic datasets and human videos with latent-action supervision. For downstream robot control, the flow-matching head is extended into a unified action expert that jointly denoises latent and robot actions. Here, latent actions serve as a latent planner, providing high-level guidance that conditions action generation. With only 1.7B parameters and 1700+ hours of pretraining data, RotVLA achieves 98.2% on LIBERO and 89.6% / 88.5% on RoboTwin2.0 under clean and randomized settings, respectively. It also demonstrates strong real-world performance on manipulation tasks, consistently outperforming existing VLA models.
Abstract: Reinforcement learning (RL) has demonstrated success in post-training large language models (LLMs) as agents for tasks such as computer use, tool calling, and coding. However, exploration remains a central challenge in RL for LLM agents, especially as they operate in language-action spaces with complex observations and sparse rewards. In this work, we address exploration for LLM agents by leveraging the ability of LLMs to plan and reason in language about the environment to shift exploration from low-level actions to higher-level language strategies. We thus propose Strategy-Guided Exploration (SGE), which first generates a concise natural-language strategy that describes what to do to make progress toward the goal, and then generates environment actions conditioned on that strategy. By exploring in the space of strategies rather than the space of actions, SGE induces structured and diverse exploration that targets different environment outcomes. To increase strategy diversity during RL, SGE introduces mixed-temperature sampling, which explores diverse strategies in parallel, along with a strategy reflection process that grounds strategy generation on the outcomes of previous strategies in the environment. Across UI interaction, tool-calling, coding, and embodied agent environments, SGE consistently outperforms exploration-focused RL baselines, improving both learning efficiency and final performance. We show that SGE enables the agent to learn to solve tasks too difficult for the base model.
Abstract: Bayesian optimization (BO) is a widely used iterative black-box optimization method that utilizes Gaussian process (GP) surrogate models. In practice, BO is typically terminated after a fixed evaluation budget is exhausted, which can incur unnecessary cost and provides no optimality guarantee on solution quality. Recent research in developing a practical stopping criterion has made empirical progress, yet a theoretically sound stopping criterion remains a work in progress. In this work, we present provably tighter instantaneous regret bounds for GP upper confidence bound (GP-UCB) at any given iteration. Then, we propose stopping criteria for GP-UCB based on this tighter bound that ensures an \epsilon-optimal solution with high probability 1-\delta upon termination. Numerical experiments are performed to validate and demonstrate the effectiveness and efficiency of our stopping criteria.
Authors:
Zipeng Guo, Xiaoan Liu, Lichen Ma, Cheng Wang, Yu He, Xiaolong Fu, Jingling Fu, Xinyuan Shan, Shaojie Guo, Luohang Liu, Junshi Huang, Yan LiAbstract: Real-world e-commerce image editing often requires multiple, localized, and auditable operations rather than global restyling. This compositional nature poses a dual challenge: models must precisely apply all requested edits to the correct regions while preserving unmodified content, even under ambiguous instructions. Existing one-shot editors conflate intent resolution, spatial grounding, and synthesis into a single step, frequently resulting in partial execution failures, which is unacceptable for commercial scenarios. To address this, we introduce GMO-E²DIT, an agentic editing framework that couples a Vision-Language Model (VLM) with a mask-conditioned image editor to tackle structured multi-turn task completion. Given an underspecified instruction, the VLM agent constructs a region-grounded edit agenda, effectively decoupling cognitive reasoning from generative rendering. The framework then executes sub-programs via operation-aware masks and references, utilizing a reflection-driven loop to inspect intermediate results and determine the subsequent state. This iterative mechanism reliably preserves safe partial progress, retries unfinished operations, and recovers from errors. Furthermore, we develop a unified data pipeline providing aligned supervision for planning, execution, and reflection, alongside EComEditBench, a comprehensive benchmark for instruction-driven evaluation. Extensive experiments demonstrate that GMO-E²DIT achieves competitive performance compared to strong closed-source models, yielding superior instruction accuracy and edit fidelity over existing baselines.
Abstract: Large language models (LLMs) have achieved strong performance in language-centric tasks. However, in agentic settings, LLMs often struggle to anticipate action consequences and adapt to environment dynamics, highlighting the need for world-modeling capabilities in LLM-based agents. We propose Reinforcement World Model Learning (RWML), a self-supervised method that learns action-conditioned world models for LLM-based agents on textual states using sim-to-real gap rewards. Our method aligns simulated next states produced by the model with realized next states observed from the environment, encouraging consistency between internal world simulations and actual environment dynamics in a pre-trained embedding space. Unlike next-state token prediction, which prioritizes token-level fidelity (i.e., reproducing exact wording) over semantic equivalence and can lead to model collapse, our method provides a more robust training signal and is empirically less susceptible to reward hacking than LLM-as-a-judge. We evaluate our method on ALFWorld and Bench and observe significant gains over the base model, despite being entirely self-supervised. When combined with task-success rewards, our method outperforms direct task-success reward RL by 6.9 and 5.7 points on ALFWorld and Bench respectively, while matching the performance of expert-data training.
Abstract: Random Forests (RF) are among the most powerful and widely used predictive models for centralized tabular data, yet few methods exist to adapt them to the federated learning setting. Unlike most federated learning approaches, the piecewise-constant nature of RF prevents exact gradient-based optimization. As a result, existing federated RF implementations rely on unprincipled heuristics: by aggregating decision trees trained independently on clients, they fail to optimize the centralized decision-tree impurity criterion, which constitutes the ground-truth in federated settings, even under simple covariate shifts. We propose FedForest, a new federated RF algorithm for horizontally partitioned data that naturally accommodates diverse forms of client data heterogeneity, from covariate shift to more complex concept shift mechanisms. We prove that our splitting procedure, based on aggregated client statistics, closely approximates the split selected by a centralized algorithm. Moreover, FedForest allows splits on client indicators when beneficial, enabling a non-parametric form of personalization that is absent from prior federated random forest methods. Empirically, we demonstrate that the resulting federated forests closely match centralized performance across heterogeneous benchmarks while remaining communication-efficient.
PaperID: 2095, Poster
Abstract: We study explainable k-means clustering, introduced by Dasgupta, Frost, Moshkovitz, and Rashtchian (2020), and generalize it to \ell_p norms. The goal is to find a clustering represented by a threshold decision tree while minimizing the k-means objective. Such a clustering is easy for humans to interpret because each internal node partitions the data by thresholding a single feature, so every cluster can be explained through a sequence of simple decisions. We give algorithms that find threshold trees with (1+\delta)k leaves and competitive ratio \tildeO_p(1/\delta\cdot \log^2+2/p-2/p^2 k) for every finite p \geq 2 and \tildeO(1/\delta \cdot d^2/p-1\log^2 k) for every 1 \leq p < 2. For 1 \leq p < 2, we show that this dependence on d is unavoidable with less than k^2/2 leaves and provide an \tildeO(\log^2+2/p-2/p^2 k)-competitive algorithm with 8k^2 leaves. We also provide near-optimal algorithms for explainable k-means under \ell_p norms with exactly k leaves. Our algorithms achieve the price of explainability of \tildeO(k) for every finite p \geq 2 and \tildeO(d^2/p-1 k) for 1\leq p \leq 2. We complement these results with nearly matching lower bounds.
Authors: Jeahn Han, Pyojin Kim
Abstract: We identify a diagonal saturation principle in modal inverse problems: when truncation noise is isotropic, the Bayes-optimal Tikhonov shape is a closed-form power law Γ k^|s| set by the prior alone, independent of the domain. Berry's random-wave conjecture decorrelates the truncation noise across modes, and Weyl's eigenvalue counting law supplies enough modes for the conclusion to survive empirical Berry violations. Together they predict an approximately flat loss landscape across the per-mode family, leaving narrow scope for a diagonal regularizer to robustly beat the closed form. On FEM-simulated acoustic rooms, the closed form is near-optimal relative to per-room oracle tuning across observation windows, and three diagonal architectures trained on the same data match its reconstruction error within 1\,pp despite learning qualitatively different spectra. The framework extends to heat diffusion via a known exponential Green's function correction with no new free parameters. Saturation is restricted to the diagonal family: Learned Iterative Ridge crosses the boundary by exploiting cross-mode coupling, locating where learning starts to help.
Authors:
Hadi Alzayer, Wenlong Huang, Haonan Chen, Christopher Luey, Lvmin Zhang, Maneesh Agrawala, Gordon Wetzstein, Fei-Fei Li, Yilun Du, Jiajun Wu, Jia-Bin HuangAbstract: Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these interaction priors, yet still grounded in physical manipulation. We introduce Masked Visual Actions, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video. Revealing robot motion makes the model act as a forward dynamics model that predicts the scene's response to low-level robot actions, while revealing desired object motion makes the same model recover robot behavior consistent with that outcome. Finetuned with only 15 hours of masked examples from real videos and simulation, a single checkpoint achieves strong visual fidelity and controllability across diverse scenes and multiple embodiments. In downstream manipulation settings, the model produces imagined rollouts whose outcomes correlate with real-world execution for policy evaluation, improves decision making by ranking candidate futures in model-based planning, and supports inverse modeling by synthesizing robot motion from desired object motion.
Abstract: We develop diffusion-based samplers for target distributions known up to a normalising constant. To this end, we rely on the well-known diffusion path that smoothly interpolates between a simple base distribution and the target, popularised by diffusion models. We tackle the score estimation problem by developing an efficient sequential Monte Carlo sampler that evolves auxiliary variables from conditional distributions along the path, providing principled score and density estimates for time-varying distributions. To control the variance of score estimates, we further propose practical control variate schedules that incur minimal overhead. We adapt this general framework to paths induced by the Ornstein–Uhlenbeck (OU) time-reversal process, stochastic interpolants, and diffusion annealed Langevin dynamics, outlining their trade-offs. Finally, we provide theoretical guarantees and empirically demonstrate the effectiveness of our method on several synthetic and real-world datasets.
PaperID: 2099, Poster
Authors:
Zhipeng Zhong, Junquan Huang, Yu Huang, Boyuan Zheng, Jungang Li, Zihao Dongfang, Song Dai, Xuming Hu, Zong-Gan ChenAbstract: Autoregressive (AR) Attention Models have become the dominant paradigm in neural solvers for vehicle routing problems (VRPs), while GCNs are almost exclusively paired with non-autoregressive (NAR) decoding and post-search algorithms. In this work, we revisit the role of decoding strategies and model properties in neural VRP solvers, and challenge the conventional pairing of NAR decoding with GCNs. We propose AR-GCN, an efficient and generalizable architecture for both Symmetric and Asymmetric VRPs (SVRPs and AVRPs). Extensive benchmarks across multiple tasks show that AR-GCN achieves advanced generalization performance, particularly on AVRPs (
PaperID: 2100, Poster
Abstract: Deep reinforcement learning agents progressively lose representational capacity during training: neurons become dormant, removing active capacity from the network, and effective rank collapses, leaving surviving neurons redundant. Existing remedies such as periodic resets, and special neural network architectures, are largely algorithm- or domain-specific. We propose a simple architectural fix, the Hadamard Representation (HR), which replaces a standard hidden layer with the element-wise product of two independently parameterized layers. HR operates through two complementary mechanisms. First, it reduces the probability of a neuron becoming dormant, which is particularly valuable for continuously differentiable activations such as \tanh: unlike dormant ReLU neurons, which are effectively pruned, saturated \tanh neurons silently corrupt downstream layers by turning their outgoing weights into fixed biases. Second, independently of dormancy, the multiplicative structure captures richer feature interactions and increases effective rank without widening the layer. We evaluate HR across five algorithms and three domains: DQN, PPO, and PQN on pixel-based discrete-action Atari, SimbaV2 on state-based continuous control, and MR.Q on visual continuous control. HR consistently improves performance over the strong baselines without any hyperparameter tuning, and gains persist against parameter-matched wider variants, ruling out parameter count as an alternative explanation.
PaperID: 2101, Poster
Authors:
Harshithanjani Athi, Sravan K Ankireddy, Jianzhong Zhang, Hyeji KimAbstract: Large vision-language models (VLMs), which process both visual inputs and text queries, incur high inference costs due to the large number of visual tokens, making visual token pruning a natural approach for improving efficiency. Recent work has leveraged query information to guide token selection, often through heuristics derived from attention patterns or token dynamics. However, these approaches do not explicitly exploit the underlying structure of query-vision interactions. We address this gap by observing that the query-vision interaction matrix, naturally induced by the query and key projections already present in transformer attention, is effectively low-rank, with a small number of dominant latent interaction modes capturing query-relevant semantics. Motivated by this observation, we propose Query-Vision Decomposition (QViD), a novel training-free query-aware visual token pruning method that exploits this structure. Extensive experiments on both image and video understanding benchmarks show that QViD consistently improves performance under matched token budgets, with the largest gains appearing under aggressive compression regimes.
PaperID: 2102, Poster
Abstract: Pruning reduces the model size and inference cost of large language models (LLMs), yet preserving performance across diverse downstream tasks remains challenging. Existing structured pruning methods commonly rely on local discrepancy measures or teacher-forced predictive criteria, such as Block Influence (BI) and perplexity (PPL). While these methods are effective on classification benchmarks, the resulting pruned models often degrade severely on generative reasoning tasks. We study redundancy in Transformer-based models and observe that pruning-induced generation drift is strongly correlated with the functional importance of modules, suggesting a more faithful criterion for identifying redundant modules. Based on this insight, we propose GDPruner, a calibration-free, generation-drift-guided block pruning method. GDPruner constructs lightweight self-generated probe trajectories, estimates module importance using tail-position generation drift, and removes redundant modules through adaptive search. Extensive experiments show that GDPruner surpasses state-of-the-art structured pruning baselines while better preserving overall performance and complex generative reasoning ability, offering a promising direction for generation-robust LLM deployment.
Abstract: A representation that scrambles the true degrees of freedom of the world cannot support reliable planning or compositional generalization. We prove that LeJEPA is guaranteed to recover a linear representation of latent variables from nonlinear observations, a property known as . What kind of world and learning objective admit this guarantee? We consider World Models with three properties (independent, stationary, and additive-noise transitions) that generate positive pairs, and study representations trained with the LeJEPA objective. Our main result: if the latent distribution is Gaussian, the representation provably achieves linear identifiability. The key insight comes from a of the representation with respect to the transition structure: each spectral component corresponds to a degree of nonlinearity, and higher degrees are strictly penalized by alignment, making the linear map the optimum. We then prove the converse: among all latent variable distributions satisfying these three properties, the Gaussian is the one that leads to linear identifiability. The guarantee degrades gracefully when the two objectives are only approximately satisfied, with an explicit bound validated empirically. Our theory turns an empirically successful recipe into a mathematical guarantee, providing the foundation for building World Models that provably recover the structure of the world.
PaperID: 2104, Poster
Abstract: Retrieval-augmented generation (RAG) is bottlenecked by the token budget of large language models: feeding long documents into the context window is expensive, while aggressive textual compression discards fine-grained information that is irreversibly lost in a one-dimensional token sequence. We argue that this bottleneck is not fundamental to information density but to the choice of compression \emphmodality. Modern visual encoders can pack the contents of a rendered text page into roughly a hundred visual tokens with little semantic loss, opening a new compression-fidelity tradeoff that purely textual methods cannot reach. We present OpticalRAG, a framework that renders documents as page images, encodes them with a frozen vision encoder, and distills the retrieved visual tokens into as few as 16 injected tokens before passing them to the LLM. Two ingredients make this work without expensive cross-modal training: (i) \emphEncoder-Bridged Retrieval, which reuses the shared representation space of a pretrained vision-language model to retrieve visual chunks from a text query, with no contrastive training; and (ii) \emphQuery-Driven Distillation, a lightweight cross-attention compressor that condenses thousands of visual tokens into a small set of query-conditioned prefix tokens. On LongBench-v2, BAMBOO, and LooGLE-v2, OpticalRAG achieves the best compression-aware accuracy (CAP) on all three benchmarks while injecting only 16 tokens, and improves the closed-book Qwen2.5-7B-Instruct backbone by 5.17 points on LongBench-v2.
Abstract: Understanding how linguistic structure emerges in language models is central to interpreting what these systems learn from data and how much supervision they truly require. In particular, semantic role understanding (``who did what to whom’’) is a core component of meaning representation, yet it remains unclear whether it arises from pre-training alone or depends on task-specific fine-tuning. In this paper, we study whether semantic role understanding emerges during language model pre-training or requires task-specific fine-tuning. We freeze decoder-only transformers and train linear probes to extract semantic roles, using performance to infer whether role information is already encoded in pre-training or learned during adaptation. Across model scales, we find that frozen representations contain substantial semantic role information, with performance improving but not fully matching fine-tuned models. This indicates partial but incomplete emergence from pre-training alone. We show that semantic role structure emerges from language modeling objectives, but its internal implementation shifts toward more distributed representations as model scale increases.
PaperID: 2106, Poster
Abstract: Comparing whether two dynamical systems implement the same computation despite differences in coordinates or measurements is a central problem in neuroscience and machine learning. Dynamical Similarity Analysis [DSA; Ostrow et al., 2023] addresses this problem by aligning finite-dimensional Koopman approximations of the two systems through an orthogonal similarity transformation. Here we show that orthogonal alignment is neither necessary nor sufficient for topological conjugacy: genuinely conjugate systems may be related by a non-orthogonal basis-transfer matrix that DSA cannot capture, while non-conjugate systems may have orthogonally equivalent Koopman operators that DSA will fail to distinguish. We then use this observation to formulate \emphConjugacy-based Similarity Analysis (CSA), which restricts alignments to those induced by candidate state-space bijections rather than arbitrary orthogonal matrices. We prove that CSA's fitted alignment is the finite-data projection of the composition operator associated with the candidate bijection, and use controlled examples to show why this distinction matters when observable dictionaries are chosen explicitly or implicitly from data. Together, these results clarify what Koopman-based similarity measures must ensure to support claims of identifying conjugacies between computational systems.
PaperID: 2107, Poster
Abstract: Do behavioral styles in large language models correspond to reusable directions in residual space, or do they mainly appear after averaging many noisy contrastive examples? We introduce MinSteer, a two-stage cached-continuation method for estimating a contrastive residual direction (CRD) within one decoding trajectory. Stage-X generates a continuation in one style; a style-flip suffix is then processed using Stage-X's final KV cache; Stage-Y continues from this inherited trajectory. The CRD is computed from generated-token residuals only, excluding prompt and suffix tokens from the average. Across sentiment transfer, politeness control, and a ParaDetox-based bidirectional diagnostic, MinSteer extracts effective steering directions from very small contrasts. On TweetEval sentiment transfer, one MinSteer contrast outperforms a CAA baseline averaged over 50 independent pairs. Japanese keigo CRDs transfer to English and Chinese outputs and outperform trajectory-free TwoSeq controls. In ParaDetox, MinSteer produces bidirectional Detoxify movement under our rewrite setup, while exposing semantic-preservation trade-offs. The directions transfer across tested scenarios and languages and show clear depth dependence. Together, these results show that cached self-rewrite trajectories provide a cleaner signal for residual-space behavior directions than independent-pair averaging.
Abstract: When distilling reasoning from large language models (LLMs) into smaller ones, teacher rationales for similar problems often vary wildly in structure and strategy. Like a chef who makes the same dish differently each time, this inconsistency burdens the student with noisy supervision that is hard to internalize. We propose Distillation through Reasoning Path Compression (D-RPC), which constrains the teacher to follow a compact, dynamically maintained bank of reusable high-level reasoning paths. For each training question, D-RPC retrieves the most relevant path and conditions the teacher to follow it, producing rationales that are consistent across similar problems yet diverse enough to cover different problem types. A PAC-Bayes analysis formalizes the resulting trade-off between bank size and coverage: smaller banks reduce supervision entropy but risk coverage gaps, and the generalization bound identifies an optimal intermediate size confirmed by our ablations. Across five math and commonsense reasoning benchmarks with two student models, D-RPC consistently outperforms chain-of-thought distillation, freeform rationale generation, direct distillation, and structured-supervision baselines, while using fewer tokens than template-heavy alternatives.
Abstract: Zero-shot model size interpolation aims to create new models of intermediate target sizes by combining existing models without additional training. Recent work on boomerang distillation [Kangaslahti et al., 2026] shows that a student language model distilled from a larger teacher can be expanded by iteratively patching its layers, replacing student layers with contiguous blocks of teacher layers to obtain models whose size and performance interpolate between the student and the teacher. In this work, we provide the first systematic study of student-layer selection for model size interpolation. We cast finding the optimal layer subset for each model size as an optimization problem and prove it can be viewed as a shortest-path problem in a certain acyclic graph. In experiments, we show that patching strongly shapes interpolation behavior, with effects that vary substantially across model families. We find that simple sequential strategies---patching either from the first layer to the last or from the last to the first---often achieve surprisingly strong performance in practice. We further introduce KLPatch, a greedy patching algorithm based on KL divergence, which often improves over last-to-first patching and approximately solves the optimization problem. Together, our results provide a principled understanding of how layer patching affects model size interpolation and offer practical guidance for constructing near-optimal interpolated models.
Authors: Ari Spiesberger, Juan J Vazquez, Nicky Pochinkov, Tomáš Gavenčiak, Peli Grietzer, Gavin Leech, Nandi Schoots
Abstract: If LLM training data is polluted with benchmark test data, then benchmark performance is only a biased estimate of out-of-distribution (OOD) generalization. Typical 'decontamination' filters use n-gram matching which fail to detect 'semantic' duplicates: sentences with equivalent (or near-equivalent) content that are not close in string space. We study this 'soft' contamination of training data by semantic duplicates. We embed the Olmo 3 training corpus and find that: (1) contamination remains widespread: we observe semantic duplicates for 76% of the CodeForces test set and 100% of MBPP's; (2) training on semantic duplicates of benchmark data improves benchmark performance by -1–22pp in controlled experiments and 18–22pp in ecological finetuning, depending on training regime and type; and (3) finetuning on duplicates of benchmark datapoints improves performance by -1–22pp (controlled) and 13–19pp (ecological) on truly-held-out datapoints from the same benchmark. The generalization we find is 'shallow': it is limited to (2) and (3), and does not typically extend to related benchmarks. We replicate on Olmo 3, Qwen3, and Qwen3.5. We thus argue that recent benchmark gains are confounded: the prevalence of soft contamination means gains reflect both genuine capability improvements and the accumulation of effective test data in growing training corpora.
PaperID: 2111, Poster
Abstract: Implicit neural representations (INRs) model signals as continuous functions from coordinates to values, but their performance depends strongly on how spectral capacity is allocated. Yet most Quantum INR (QINR) architectures treat spectral components as independent outputs of separate quantum blocks. We introduce the Spectral Quantum Memory Model (SQMM), a memory coupled quantum INR in which sequential hybrid data reuploading layers (bands) communicate through a shared coherent quantum memory register. The construction gives a controlled way to share spectral information across quantum modules while preserving the finite Fourier structure of a band. When memory read and write layers are input independent, we prove that the memory state forms an operator valued Fourier series whose support grows by Minkowski summation of local band spectra. Downstream bands then condition their Fourier coefficients on the spectral content stored by earlier bands. This theory motivates a spectral memory alignment loss that encourages memory to encode the spectrum expressed by the deployed prefix of the model. Across audio and image reconstruction benchmarks, SQMM improves over QINR and INR baselines, achieving state of the art results.
Abstract: Fine-tuning is often believed to reduce uncertainty and diversity in large language models, but existing analyses overlook output length, a key confounder, and therefore fail to capture how uncertainty is distributed across an entire generation rollout. To address this, we propose Canopy Entropy (\mathrmCE^\star), a measure that views language generation from a tree perspective, where "canopy" represents the space of all possible rollouts, making \mathrmCE^\star naturally quantify the effective size of the generation space. \mathrmCE^\star jointly captures the output length N and the generated sequence Y_1:N, and we show that it equals to total Shannon entropy H(N, Y_1:N\mid X), where X denotes the prompt. This formulation yields interpretable metrics, including a length--uncertainty correlation term \rho(N, r_N), where r_N is the entropy rate, quantifying information conveyance efficiency by indicating whether longer outputs are more or less informative per token. Empirically, across tasks and model families, we find that fine-tuned models consistently exhibit stronger positive correlation \rho(N, r_N), even when total entropy decreases. Furthermore, after controlling for model family, task, prompt, and output-length effects, we find that fine-tuning nearly triples the strength of the relationship between entropy rate and semantic diversity, suggesting that aligned models convert uncertainty into semantically meaningful variation much more efficiently. Overall, these results suggest that fine-tuning does not simply reduce uncertainty, but fundamentally reorganizes it into more informative and semantically meaningful generations.
PaperID: 2113, Poster
Authors:
Hovhannes Margaryan, Vicky Kalogeiton, Quentin Bammey, Christian SandorAbstract: Editing a video stream in real time, without access to future frames, is essential for interactive applications. Existing training-free video editing methods, however, assume full offline access to the entire sequence, while streaming approaches either rely on additional training or are limited to specific editing techniques. We show that this limitation can be removed by analyzing the self-attention dynamics of DMD-distilled autoregressive video diffusion models. We uncover a key asymmetry: queries and keys remain stable and aligned with their source counterparts across denoising timesteps, while values diverge and encode the evolving content. This observation leads to a simple mechanism: a single forward pass suffices to capture source identity in the value features, which can be reused throughout generation. Starting from pure Gaussian noise and injecting only these source values reconstructs the input with fidelity matching inversion-based methods, effectively eliminating inversion. Based on this insight, we introduce AR-Edit, a training-free, inversion-free method that caches source values once and selectively injects them during autoregressive denoising, enabling real-time, structure-preserving streaming video editing with minimal overhead. Code will be released.
PaperID: 2114, Poster
Abstract: There is a growing need for understanding how trained machine learning models behave beyond standard predictive performance. With this work, we aim to understand trained machine learning models by questioning their data preferences. We propose a mathematical framework for guided data generation that allows us to produce fixed-label, prediction-risky, parameter-sensitive, or model-contrastive samples, among others. To showcase our framework, we pose these queries to a range of models trained on a range of classification and regression tasks, with answers in the form of generated data.
PaperID: 2115, Poster
Abstract: Graph-based nearest neighbor search algorithms like HNSW, DiskANN, NSG, and others have surged in popularity in recent years due to their excellent empirical performance in comparison to previously successful ``space partitioning'' methods like IVF data structures and Locality Sensitive Hashing (LSH). However, in contrast to LSH and related methods, obtaining strong worst-case theoretical guarantees for graph-based search has remained a challenging open problem. In this paper, we prove a strong lower bound that helps explain the lack of progress towards such guarantees. In particular, we show that no standard graph-based method (i.e., any method that runs a variant of greedy search on a flat or hierarchical search graph) can achieve sublinear query time for approximate nearest neighbor search in Euclidean space, even for an arbitrarily large approximation factor. On the other hand, we prove that sublinear query time \emphis possible if we allow for the use of \emphSteiner points in the search graph. We obtain bounds comparable to LSH, showing e.g., that graph-based search with Steiner points can solve c-approximate nearest neighbor search in Euclidean space with query time \tildeO\left(n^1/c\right). We complement this result with initial experiments exploring how Steiner points can be incorporated into practical graph-based methods, showing modest but meaningful improvements in recall versus query-time tradeoffs.
Abstract: We introduce ReflectDrive-2, a masked discrete diffusion planner with a separate action expert for autonomous driving that represents plans as discrete trajectory tokens and generates them through parallel masked decoding. This discrete token space enables in-place trajectory revision: AutoEdit rewrites selected tokens using the same model, without requiring an auxiliary refinement network. To train this capability, we use a two-stage procedure. First, we construct structure-aware perturbations of expert trajectories along longitudinal progress and lateral heading directions and supervise the model to recover the original expert trajectory. We then fine-tune the full decision--draft--reflect rollout with reinforcement learning (RL), assigning terminal driving reward to the final post-edit trajectory and propagating policy-gradient credit through full-rollout transitions. Full-rollout RL proves crucial for coupling drafting and editing: under supervised training alone, inference-time AutoEdit improves PDMS by at most 0.3, whereas RL increases its gain to 1.9. We also co-design an efficient reflective decoding stack for the decision--draft--reflect pipeline, combining shared-prefix KV reuse, Alternating Step Decode, and fused on-device unmasking. On NAVSIM, ReflectDrive-2 achieves 91.0 PDMS with camera-only input and 94.8 PDMS in a best-of-6 oracle setting, while running at 31.8 ms average latency on NVIDIA Thor.
Abstract: Many inference-time language-model pipelines combine a cheap reward signal with an expensive verifier, such as exact answer checking in mathematical reasoning or hidden-test execution in code generation. We formalize this setting using a learning-theoretic lens as generative active search: a cost-sensitive first-positive search problem in which a policy adaptively samples candidates from an unknown distribution, observes cheap scores, and pays for verifier labels until it finds a positive example. For a fixed prompt, the generator and reward model induce two unknown objects: a distribution over reward scores and a score-conditioned success function. When these quantities are known, we characterize the distribution-aware optimal policy using a dynamic programming approach. In the realistic and practical setting where both the score distribution and success function are unknown, we propose ADAP, a shellwise adaptive generate-rank-verify algorithm that progressively increases the number of sampled responses and top-ranked verifications. Under the monotonicity assumption that higher reward scores are no less likely to pass verification, we show that ADAP achieves expected cost within a constant factor of the distribution-aware optimum. We complement this result with learning-theoretic lower bounds, based on a centered star number, showing that structural assumptions on the score--label relationship are necessary. Experiments on mathematical reasoning and competitive programming validate the predicted advantage over both fixed non-adaptive policies and difficulty-adaptive baselines.
PaperID: 2118, Poster
Authors: Xiao Su, Jian Zhou
Abstract: Zero-shot coordination requires agents to collaborate with previously unseen partners without test-time fine-tuning or explicit communication. A persistent difficulty is that partner behavior contains information at multiple timescales: stable conventions such as role preference or spatial habit, and fast local intent such as the next object to collect or deliver. Most partner-conditioned policies compress these signals into a single latent embedding, which can blur long-term convention with short-term action evidence. We propose StylePlan, a partner modeling framework that separates slow partner style from fast partner intent and combines them through conflict-aware arbitration. StylePlan builds an online behavioral fingerprint, predicts short-horizon partner intent, forms a style-conditioned intent prior, and uses the disagreement between prior and immediate evidence to modulate a recurrent actor-critic policy. On the verified Overcooked V2 artifacts, StylePlan obtains the best reported average cross-play among methods with available entries under the legacy evaluator. We additionally provide a scan-based JAX evaluation with unseen-partner XP, within-episode partner switches, unified SP/FCP reruns, and wired module ablations. The results show that StylePlan is most reliable as an online adaptation mechanism: it improves recovery and drop under partner switches, while the current implementation still needs broader unified reruns to establish a stronger XP claim.
Abstract: Occupancy World Models (OWMs) aim to predict future 3D occupancy scenes from historical observations and future ego motions, providing a useful world simulation tool for autonomous driving. Existing methods usually rely on large networks to implicitly align multi-frame historical states and predict future occupancy in a full-state manner, which is costly and redundant because most scene regions remain unchanged over short time intervals. In this paper, we propose Aligned Delta-Triplane Transformer (ADTT), a compact 4D OWM that explicitly aligns historical triplanes with ego-motion compensation and predicts only future scene changes. Based on the aligned triplane prior, a query-conditioned Transformer predicts residual changes, where ego-motion queries guide controllable future motion and learnable external queries capture scene changes beyond ego motion. The predicted changes are added to the aligned prior and decoded into future occupancy scenes. Experiments on Occ3D-nus show that ADTT achieves state-of-the-art forecasting performance with fewer parameters and higher FPS. Our code is publicly available online: https://anonymous.4open.science/r/NeurIPS26-Occ/.
PaperID: 2120, Poster
Abstract: Planner-guided model-based reinforcement learning (MBRL) has emerged as a powerful paradigm for continuous control, combining learned world models with online planning to achieve strong performance and sample efficiency. However, existing methods typically distill planner-improved actions into a unimodal Gaussian policy and reuse it as a proposal prior for subsequent planning, overlooking the fact that Model Predictive Path Integral (MPPI) planners can induce multimodal supervision over high-value actions. Compressing such supervision into a unimodal distribution can lead to mode averaging, limited proposal coverage, and unstable bootstrapped policy learning. To address this issue, we propose GeMP, a Generative Multimodal Planning framework for MBRL. We theoretically characterize the multimodality of planner-induced supervision, showing that finite-sample planning can provide a distribution of high-value actions rather than a single deterministic target. To capture this distribution without compromising training stability, GeMP jointly learns a multimodal flow policy through value-weighted supervision and retains an auxiliary Gaussian policy for stablizing off-policy actor-critic learning. We further employ a heterogeneous proposal generator that combines Gaussian, flow-based, and random action-sequence candidates to improve proposal coverage for MPPI refinement. Experiments on MyoSuite and DeepMind Control Suite show that GeMP improves control performance and training stability over competitive baselines.
PaperID: 2121, Poster
Abstract: Large Language Model (LLM) agents with tool use exhibit inconsistent behavior across independent runs from identical states, calling different tools, producing contradictory outputs, or pursuing divergent strategies. This variability compounds over time and undermines deployment reliability. Existing evaluation metrics primarily measure task correctness via completion-based scores and do not capture consistency in agent behavior across executions. Separately, agent behavior is analyzed either at the token level using probability distributions or at the single-response level using semantic clustering, but neither captures how inconsistency propagates across future turns of a trajectory. We introduce Recursive Semantic Divergence (RSD), an unsupervised trajectory-level consistency metric adapted from bisimulation metrics in reinforcement learning. RSD measures the expected cumulative semantic disagreement between two independent rollouts of the same agent from the same state, using natural language inference to detect logical inconsistency. It is defined recursively over future states and approximated by a neural network trained on independent rollouts. The learned metric estimates future divergence from the current state, providing a consistency score at each turn without requiring the full trajectory to complete. When used as a reward signal for fine-tuning, RSD reduces the variation in task progress from 132% to 91% on \tau^2-bench and from 90% to 44% on toolsandbox, while improving performance across benchmarks and model scales. Training with RSD as an additional reward signal matches supervised fine-tuning performance with substantially lower variance, without requiring ground-truth labels.
PaperID: 2122, Poster
Abstract: Large language models adapted to specialized targets such as Lean theorem proving or tool calling often benefit from extending the tokenizer with target-specific tokens. However, target-only vocabulary-expanded fine-tuning can leave the newly added embeddings transfer-incomplete, since the new rows receive gradients only from examples containing the corresponding tokens, so target-local gains do not necessarily imply that the embeddings are integrated with related capabilities. We propose SAVE, a sparsity-aware influence estimator for selecting auxiliary healing data from generic corpora. SAVE computes token-conditional curvature over the effective update set of each new vocabulary row and combines this vocabulary-aware signal with dense LoRA-level influence. Across Lean autoformalization and tool calling, SAVE improves target performance and related transfer while remaining competitive on broad retention benchmarks.
Abstract: Riemannian flow matching (RFM) extends flow-based generative modeling to data supported on manifolds by learning a time-dependent tangent vector field whose flow-ODE transports a simple base distribution to the data law. We develop a nonasymptotic Total Variation (TV) convergence analysis for RFM samplers that use a learned vector field together with Euler discretization on manifolds. Our key technical ingredient is a differential inequality governing the evolution of TV between two manifold ODE flows, which expresses the time-derivative of TV through the divergence of the vector-field mismatch and the score of the reference flow; controlling these terms requires establishing new bounds that explicitly account for parallel transport and curvature. Under smoothness assumptions on the population and learned flow-matching field, together with mean-square approximation guarantees for the learned field, we obtain explicit bounds of the form \mathrmTV\le C_\mathrmLip\,h + C_\varepsilon\,\varepsilon, cleanly separating numerical discretization and learning errors. Here, h is the step-size and \varepsilon is the target accuracy. Instantiations yield \emphexplicit polynomial iteration complexities on the hypersphere S^d, and on the Hadamard manifolds (for e.g., SPD manifolds) under mild moment conditions.
PaperID: 2124, Poster
Abstract: Modern audio generation predominantly relies on latent-space compression, introducing additional complexity and potential information loss. In this work, we challenge this paradigm with WavFlow, a framework that generates high-fidelity audio directly in raw waveform space without intermediate representations. To overcome the inherent difficulties of modeling high-dimensional and low-energy signals, we reshape audio into 2D token grids through waveform patchify and introduce amplitude lifting to align signal scales, enabling stable optimization via direct x-prediction in flow matching. To capture complex semantic alignment and temporal synchronization, we leverage an automated data pipeline to curate 5M high-quality video-text-audio triplets, allowing the model to learn fine-grained acoustic patterns from scratch. Experimental results show that WavFlow achieves state-of-the-art results on the video-to-audio benchmark VGGSound (FD PANNs 17.40, DeSync 0.44) and the text-to-audio benchmark AudioCaps (FD_PANNs 10.63), outperforming established latent-based methods. Our work demonstrates that such intermediate compression is not a prerequisite for high-quality synthesis, offering a simpler and more scalable alternative for multimodal audio generation.
PaperID: 2125, Poster
Abstract: Causal discovery often starts from a set of variables that were measured together because they are believed to describe related parts of the same underlying process. Standard continuous DAG learners encode data fit, sparsity, and acyclicity, but they can still return graphs that fragment into isolated variables or disconnected components. We study a simple structural prior for this setting: the learned causal DAG should remain sparse while its undirected skeleton should be weakly connected. We introduce Fiedler regularization, a differentiable spectral penalty based on the algebraic connectivity of a smooth support-preserving skeleton of the learned weighted adjacency matrix. The resulting penalty can be added modularly to differentiable causal discovery objectives, providing connectivity control without changing the underlying learner. We instantiate the approach for GOLEM, graph autoencoder, and DAG-GNN models. Across sparse connected synthetic DAGs up to 200 variables, including connected-conditioned Erdős-Rényi, sparse connected, scale-free, and small-world graph families, Fiedler-regularized learners improve fragmentation control and structural recovery across learner families. These results position algebraic connectivity as a principled and practical prior for learning sparse causal DAGs in settings where the measured variables are expected to have a connected causal structure.
Authors:
Dongwei Wang, Jinhee Kim, Seokho Han, Denis Gudovskiy, Yohei Nakata, Tomoyuki Okuno, KhayTze Peong, Kang E Jeon, Jong Hwan Ko, Yiran Chen, Huanrui YangAbstract: Dynamic runtime latency and memory constraints necessitate flexible large language model (LLM) deployment, where an LLM can be inferred with various quantization precisions based on available computational resources. Recent work on such any-precision quantization either relies on hardware-inefficient vector quantization or induces additional scaling factors when switching between bit-widths. Meanwhile, existing post-training quantization (PTQ) methods calibrated for a fixed low precision show poor generalizability under runtime precision change. In this work, we attribute the source of poor generalization across bit-widths to a precision-dependent outlier migration phenomenon where the distribution of PTQ-sensitive tokens changes across precisions. Motivated by this observation, we propose \textttMoBiQuant, a novel any-precision Mixture-of-Bits quantization framework that adjusts weight precision for flexible LLM inference based on token sensitivity. Specifically, we propose a many-in-one recursive residual quantization that can iteratively reconstruct higher-precision weights at runtime and mitigates outlier migration with a token-aware router to dynamically select the optimal inference precision of each token. Extensive experiments show that \textttMoBiQuant matches or surpasses frontier single-precision PTQ while exhibiting strong elasticity, achieving significant memory savings and throughput gains of up to 1.34× over state-of-the-art any-precision methods. Source code is provided in the supplementary material.
Authors: Eros Fanì, Oguzhan Ersoy
Abstract: Large Language Models (LLMs) have achieved remarkable performance on a wide range of specialized tasks, exhibiting strong problem-solving capabilities. However, training these models is prohibitively expensive, and they often lack domain-specific expertise because they rely on general knowledge datasets. Expertise finetuning can address this issue; however, it often leads to overspecialization, and developing a single multi-domain expert remains difficult due to diverging objectives. Furthermore, multitask training is challenging due to interference and catastrophic forgetting. Existing work proposes combining the expertise of dense models within a Mixture of Experts (MoE) architecture, although this approach still requires multitask finetuning. To address these issues, we introduce Dynamic Upcycling MoE (DUME), a novel approach that reuses dense experts trained on different domains to construct a unified MoE model. Our method builds a single multitask model that preserves the capabilities of the original dense experts without requiring additional training. DUME is both cost-efficient and scalable: by leveraging the closed-form solution of ridge regression, it eliminates the need for further optimization and enables experts to be added or removed dynamically. Our extensive experiments on different tasks (on causal language modeling and reasoning settings) with several model sizes (between 115M and 9B parameters) and architectures demonstrate that DUME consistently outperforms baseline approaches. In particular, in a reasoning setting, DUME can also surpass the dense expert model's performance in its domain, achieving 109.9% of the dense expert's performance on average. Finally, we also show that the DUME model can be fine-tuned to further improve performance in complex reasoning scenarios. Our code is available at this anonymous link: https://anonymous.4open.science/r/DUME-NeurIPS-D551.
PaperID: 2128, Poster
Authors:
Bowei Liu, Xinchen Zhang, Xuhuan Li, Kaian Jiang, Jingwei Liu, Congyang Zhao, Yifei Qian, Junhong Liu, Zhiheng Li, Tianlin Zhang, Yujiu Yang, Yang Cai, Ling YangAbstract: Generative visual verifiers play a crucial role in advancing multimodal intelligence systems. In this paper, to enhance the robustness of visual verifiers in complex and realistic open-world scenarios, we move beyond conventional binary verifiers that produce only True/False discrete judgments. Instead, we introduce a verifier that automatically decomposes complex prompts or environmental rules into fine-grained components and assigns pointwise scores, thereby providing a continuous and more informative critique signal. Under a general reinforcement learning framework, we first empirically demonstrate that the proposed pointwise decomposed verifier significantly outperforms binary verifiers in traditional text-to-image scenarios. We then extend our investigation to a novel paradigm: generative world modeling, where generative models are leveraged for world simulation and world reasoning. From a theoretical standpoint, we further show that when ground-truth images are easy to obtain, a pairwise verification paradigm yields more accurate critiques than the pointwise formulation. Empirical results on diverse world-modeling tasks, such as Maze and Sudoku, further validate the effectiveness of our approach. Overall, our findings suggest a key insight: the transition from generative models to world-modeling agents critically hinges on the availability of accurate pairwise decomposed visual verifiers.
PaperID: 2129, Poster
Abstract: Symbolic regression (SR) aims to discover compact, interpretable mathematical expressions from data. This goal inherently requires iterative refinement rather than one-shot construction. Autoregressive methods lack the ability to perform targeted corrections, so whenever structural flaws appear, the remainder of the sequence must be discarded and regenerated. Meanwhile, search-based methods rely on costly and indirect replacement through enumeration. We propose EditFlowSR, which reframes SR as iterative editing of expression trees through insertions, deletions, and substitutions. To reconcile fitting accuracy with algebraic simplicity, we introduce dual trajectory supervision. One trajectory teaches the model to construct expressions from data, while the other teaches it to algebraically simplify them, and both are unified under the shared editing framework. At inference time, EditFlowSR starts from a randomly initialized expression and progressively refines it through iterative editing, simultaneously constructing valid structure and removing redundancy. Across the standard SRBench, EditFlowSR achieves first-rank Pareto dominance in predictive accuracy, expression tree size, and time complexity. These results establish edit-based generation as a principled and effective paradigm for end-to-end symbolic regression. Code is available at \urlhttps://anonymous.4open.science/r/EditFlowSR_NIPS2026.
PaperID: 2130, Poster
Abstract: Sparse autoencoders (SAEs) have become an important tool for unsupervised concept discovery in large models. To make the resulting feature spaces more interpretable and manageable, recent approaches have begun imposing hierarchical structure, either explicitly or as an implicit effect of training constraints, yet rigorous comparison remains difficult. There are no agreed-upon requirements for what a meaningful feature hierarchy should satisfy, and evaluation has largely relied on qualitative illustrations with fragmented quantitative protocols. To address this, we derive a set of key requirements for generalization/specialization hierarchies in unsupervised concept discovery, drawing on semantic net and taxonomy research alongside recent SAE work, and use them to derive a concrete evaluation protocol. Applying this protocol to current SAE approaches trained on visual data, we find that while feature spaces generally provide a basis for sensible hierarchies, establishing good hierarchical structure remains challenging. In particular, feature absorption, both in its well-known hard form and in a continuous, soft form, systematically compromises hierarchy quality, pointing to a fundamental tension that future approaches will need to navigate.
PaperID: 2131, Poster
Abstract: Speculative decoding is an effective technique for accelerating autoregressive models in memory-bound regimes. Prior work has explored tree-based drafting to speculate multiple candidates at each step, improving verification acceptance rates. However, existing systems fail to fully realize the potential of tree-based drafting, leaving significant performance overhead. In this paper, we present FastDraft, a plug-and-play module that improves the efficiency of tree-based drafting in long-context speculative decoding. FastDraft exploits both the shared-prefix I/O pattern across draft branches and the constrained tree structure of speculative decoding, enabling specialized kernels that reduce draft-phase overhead while preserving compatibility with existing inference frameworks. FastDraft requires no modifications to the target model, draft model, or sampling logic. We implement FastDraft within SGLang and compare it against SGLang’s vanilla tree-based drafting backend, demonstrating substantially faster draft execution: up to 2.66× speedup in the draft phase and 1.77× higher throughput over SGLang under long-context settings.
Authors: Calvin Isley, Johann D Gaebler, Sharad Goel
Abstract: Statistical decision algorithms are increasingly deployed in domains where ground-truth labels are hard to obtain, such as hiring, university admissions, and content moderation. In these settings, models are typically trained on historical human evaluations—for example, using past hiring decisions as a proxy for true applicant quality. However, if past evaluations unjustly penalize certain groups, models trained on these labels may inherit those biases. To address this problem, we propose basing predictions on rubric embeddings, a representation framework that replaces standard black-box embeddings with features derived from expert-defined criteria that align with the underlying construct of interest. By anchoring predictions to semantically meaningful dimensions, this approach guards against biased proxy signals. We provide both theoretical and empirical evidence that rubric embeddings mitigate label bias under plausible conditions. Empirically, we evaluate our method on a novel dataset of applications to a large master’s program. We find that models trained on rubric embeddings reduce group disparities while improving measures of cohort quality. Our results suggest that basing predictions on interpretable, domain-grounded representations offers a practical approach to learning in the presence of biased labels.
PaperID: 2133, Poster
Abstract: Understanding how the brain integrates information across sensory modalities under naturalistic settings remains a central challenge in cognitive computational neuroscience. Despite recent progress in multimodal brain encoding driven by the advancement of deep learning, many existing approaches combine representations from pretrained modality-specific models using static aggregation or linear readouts, and therefore do not explicitly capture how information is selectively integrated across cortical regions. Here, we introduce Multimodal Transformer Brain Encoder (M-TBEn), a neural encoding framework that models multimodal integration as a routing problem. M-TBEn extends prior vision-based brain encoding approaches by introducing a cross-attention mechanism for multimodality, in which learnable parcel-level queries dynamically select and aggregate modality-specific representations from visual, auditory, and linguistic inputs. This design is not only parameter-efficient but also yields parcel-specific multimodal representations that support a flexible linear readout for predicting brain responses at various spatial resolutions. We evaluate M-TBEn on naturalistic video-viewing data, including both in-distribution and out-of-distribution stimulus conditions. The model is able to accurately predict parcel-level and voxel-level fMRI responses and generalizes across stimulus distributions. Furthermore, the learned attention patterns provide a structured descriptive basis for analyzing modality integration across cortical parcels, enabling a connection between parcel-specific routing profiles and established functional brain networks. Together, these results suggest that attention–based routing offers a principled computational framework for modeling multimodal integration in the human brain.
PaperID: 2134, Poster
Abstract: Classifier-free guidance (CFG) has been widely adopted in autoregressive (AR) models for high-quality image generation. Despite its strong empirical performance, its mechanism in AR models remains unclear. This paper demystifies the mechanism of CFG through both empirical and theoretical studies. By examining the top-ranked semantics of different components in CFG, we show that texture information is primarily encoded in the difference between the conditional and unconditional logits, whereas both logits share previous-token repetition as their leading semantics. This strong repetition semantics obscures the desired texture information, causing greedy decoding from the conditional logits alone to produce nearly pure-color images. Through a training-dynamics analysis of shallow transformers, we prove that this shared repetition semantics does not arise from limitations or failures of pretraining, but instead originates from the texture sparsity of images. We further show that CFG improves generation quality by rectifying this repetition bias during inference. Motivated by the shared semantics between conditional and unconditional logits, we propose Attention Weight Reuse (AttnReuse), which reuses intermediate attention computations from the conditional-logit forward pass to accelerate the unconditional-logit computation. AttnReuse reduces about 25% of attention computation with little performance degradation across different models.
PaperID: 2135, Poster
Abstract: Blind deconvolution is a classical problem arising in many signal processing and machine learning applications, where one aims to recover an unknown kernel and an unknown signal from their convolution. Since both the kernel and the signal are unknown, the problem is intrinsically ill-posed, and meaningful recovery is possible only under suitable structural assumptions. In this paper, we study sparse blind deconvolution, where the kernel is s_h-sparse and the signal admits an s_x-sparse representation in a known dictionary. Although sparsity is a natural and practically relevant prior, its interaction with the bilinear observation model creates substantial challenges for both computation and analysis. To address these challenges, we propose Thresholded Wirtinger Flow (ThWF), a simple and scalable iterative algorithm equipped with a sparse spectral initialization. For noiseless observations, we prove that the proposed initialization achieves exact support recovery with sample complexity M \gtrsim s_h^2+s_x^2, up to logarithmic factors. Building on this initialization, we establish what is, to the best of our knowledge, the first algorithmic convergence guarantee for sparse blind deconvolution: ThWF converges linearly to the ground truth, up to the inherent scaling ambiguity, provided that M \gtrsim s_h+s_x up to logarithmic factors. Our analysis further extends to noisy measurements, showing that ThWF is robust and contracts linearly up to a statistical error under a bounded noise level. Experiments on synthetic data and real-world image deblurring tasks corroborate the predicted linear convergence and phase transition behavior, and demonstrate the restoration performance of the proposed method.
Abstract: Molecules in equilibrium follow a Boltzmann distribution, making the underlying energy landscape a physically grounded modeling objective. However, such landscapes are difficult to learn from data and, once learned, hard to sample from. Diffusion and flow-matching models sidestep these difficulties by learning a time-conditional score or transport field between noise and data, losing the energy inductive bias in exchange for a more tractable training objective. We introduce EBMol, an energy-based model (EBM) that restores this inductive bias by learning an atom-additive scalar potential without explicit simulation during training. Our method employs a flow-inspired Restoring Field Matching objective to approximate the energy landscape. We adopt the Mirror-Langevin algorithm for sampling, enabling unified updates of atomic positions and types, and incorporate parallel tempering for inference-time compute scaling. EBMol is the first EBM for 3D molecular generation to achieve state-of-the-art performance on QM9 and GEOM-Drugs. Moreover, we show that the learned energy landscape serves as a principled quality metric for ranking and filtering configurations, and demonstrate controllable generation without retraining through shape-steered sampling via potential composition and zero-shot linker design.
PaperID: 2137, Poster
Abstract: Single‑cell RNA sequencing (scRNA‑seq) data is subject to strict access control due to its sensitive nature, motivating the use of synthetic data generation (SDG) for privacy‑preserving data sharing. We present the first adversarial privacy attack that performs meaningfully above random guessing against state‑of‑the‑art scRNA‑seq SDG methods. Our attack enables donor‑level membership inference, demonstrating that leading SDG techniques fail to adequately mask which individuals were used to train the generator. We show that privacy leakage increases as the number of training donors decreases. Although the attack is designed to exploit vulnerabilities in scDesign2, we find that it also succeeds against synthetic data generated by other leading methods, including scDesign3 and scVI. This transferability indicates that an adversary can infer sensitive information from synthetic data without access to the training procedure, model parameters, or even the underlying generation algorithm. Finally, we investigate the use of perturbation with noise during the SDG process as a first‑line defense, empirically evaluating its effectiveness in neutralizing the attack and its impact on utility.
Authors:
Guozhen Zhu, Yuqian Hu, Sakila S Jayaweera, Wei-Hsiang Wang, Weihang Gao, Jiaxuan Zhang, Beibei Wang, Chenshu Wu, K. LiuAbstract: Ambient intelligence seeks to continuously understand human presence, activity, and physiology in physical spaces, enabling smart environments, health monitoring, and human-computer interaction. WiFi infrastructure offers a ubiquitous, always-on, and privacy-preserving sensing substrate across billions of IoT devices. Yet ambient sensing with WiFi remains largely fragmented, with most systems relying on task-specific models, labeled data collection, and customized training pipelines. We present AmbientFM, the first foundation model for ambient sensing through ubiquitous WiFi signals. AmbientFM is pre-trained on 9.2 million unlabeled Channel State Information (CSI) samples collected over 439 days from 20 commercial device types deployed in real-world environments. It learns transferable wireless representations through contrastive learning, masked reconstruction, and physics-informed objectives tailored to wireless signals. With lightweight adaptation, a single backbone transfers to 9 downstream tasks, achieving above 0.90 AUROC on all classification benchmarks and supporting dense spatial reconstruction. AmbientFM reduces the need for labeled data and task-specific design, enabling scalable ambient intelligence on existing wireless infrastructure.
PaperID: 2139, Poster
Abstract: Discrete diffusion language models (DLMs) offer a promising non-autoregressive alternative to large language models by enabling parallel generation, yet their performance still lags behind autoregressive counterparts. Knowledge distillation has recently proven highly effective for improving autoregressive models, but its applicability to DLMs remains poorly understood. We present the first systematic study of KD for DLMs and uncover two non-intuitive failures of classical forward-KL (FKL) distillation: (i) student perplexity does not improve monotonically with more teacher-generated data, and (ii) on the same teacher-generated corpus, FKL distillation underperforms simply training the student with the original DLM denoising objective. These failures hold for both a 110M MDLM teacher and a 7B Dream teacher. To address them, we propose a unified, data-free distillation framework for masked DLMs. The framework first runs an off-policy phase that absorbs general knowledge from a static set of teacher samples and then transitions to an on-policy phase whose objective combines: a trajectory loss along the student's own reverse process, confidence reweighting that down-weights tokens where the teacher is unsure, and a trust-region anchor that stabilizes the moving target. For the smaller MDLM regime where FKL alone is insufficient, we additionally adopt the \alpha-\beta divergence. On MDLM, the resulting student attains a GPT-2 perplexity of 24.75, surpassing the teacher's 27.04; on Dream-7B distillation to a 93M student model, the framework reduces Qwen-2.5 perplexity from the FKL-distilled baseline of 25.23 to 19.83 while preserving zero-shot accuracy.
PaperID: 2140, Poster
Abstract: Orthogonal trace-sum maximization (OTSM) problems arise in a wide range of data processing applications, including canonical correlation analysis and cryogenic electron microscopy. Despite their practical importance, the existing theoretical understanding of these problems remains incomplete. In this paper, we show that the generalized power method (GPM) converges linearly to the global optimal solution with high probability under an additive Gaussian noise model, provided that the noise level satisfies the nearly optimal bound (O(\sqrtn/\log n)). In addition, we prove that the semidefinite programming (SDP) relaxation of OTSM is tight and admits a unique optimal solution under the same noise regime, improving upon the best known theoretical guarantee of (O(n^1/4)). Extensive numerical experiments further demonstrate that our theoretical predictions closely match empirical observations.
PaperID: 2141, Poster
Abstract: Multicalibration gradient boosting has recently emerged as a scalable method that empirically produces approximately multicalibrated predictors and has been deployed at web scale. Despite this empirical success, its convergence properties are not well understood. In this paper, we provide computational guarantees for multicalibration gradient boosting algorithms. We show that the magnitude of successive prediction updates decays at O(1/\sqrtT), which implies the same convergence rate bound for the empirical multicalibration error over rounds. Under additional smoothness assumptions on the weak learners, this rate improves to linear convergence. We further establish convergence for adaptive variants. Experiments on real-world datasets support our theory and clarify the regimes in which the method achieves fast convergence.
PaperID: 2142, Poster
Abstract: Reinforcement learning with verifiable rewards (RLVR) typically relies on a verifier that can reliably judge whether sampled solutions are correct. However, in many realistic settings, complete verification is generally unavailable, while cheap checks can only falsify failures. We formalize this regime as one-sided verifiability (OSV): verified negatives are reliable counterexamples, but verified positives are merely non-falsified candidates and may contain many false positives. This asymmetry makes standard RLVR unreliable, because reinforcing verified positives can directly reward shortcuts that pass partial checks. To address this issue, we propose Negative-Only Policy Optimization (NOPO), which treats verified positives as unlabeled and updates the policy only by suppressing trusted verified negatives. Specifically, NOPO combines two safeguards: a sample-level NPO term that self-attenuates once a verifier-negative is suppressed below the reference, and a group-level non-falsified support weight that scales each group's update by its verifier-positive count. Across unit-test, constraint-check, and LLM self-verifiers, NOPO improves pass@k over baselines on code, logic, and math benchmarks when oracle rewards are unavailable. These gains persist under inference-time scaling and cross-benchmark transfer.
Abstract: Reinforcement learning (RL) is increasingly used to improve the reasoning, coding, and tool-use capabilities of large language models, but agentic RL remains prohibitively expensive. Scaling RL to agentic LLMs requires supporting complex workloads, including multi-policy collaborative training, while efficiently using elastic, heterogeneous, and cross-region compute resources. Existing LLM RL systems support some of these capabilities, but each new extension often requires dedicated system engineering. This burden arises from trainer-centered control architectures and the lack of principled abstractions for RL system components. To address these limitations, we propose AstraFlow, a dataflow-oriented RL system that replaces conventional trainer-centered control with principled component abstractions. In AstraFlow, rollout services, dataflow management, and training are decoupled into autonomous components, enabling the system to natively support complex multi-policy agentic RL workloads and efficiently exploit diverse compute resources. We evaluate AstraFlow across math, code, search, and AgentBench workloads, showing that the same system supports multi-policy training, elastic scaling, heterogeneous cross-region execution, and composable data algorithms without system-level code changes. In multi-policy collaborative training, AstraFlow achieves comparable or better accuracy than existing RL systems while speeding up training time by 2.7x.
PaperID: 2144, Poster
Authors: Iurii Dorn
Abstract: Scaling laws are increasingly used to pick which expensive training run to scale up. The practical question — which candidate is best at a fixed future compute budget C^\star — fits neither classical best-arm identification nor minimum-oriented functional bandits. We formulate it as the target-budget best-function identification (TB-BFI) problem, with two natural flavors: TB-BFI-FC asks for a \delta-correct certificate of the winner at C^\star, and TB-BFI-FB asks for low target-budget simple regret under a hard exploration budget. For the shared-exponent power-law class we prove uniform target-budget validity, a finite-time stopping bound for the cost-aware TB-LUCB rule, and a per-arm-exponent perturbation extension. We then close two structural gaps. (a) We extend the theory beyond shared-exponent to the joint (N,D) Chinchilla law with shared but unknown exponents, via a two-stage protocol with explicit pilot sample complexity and a quadratic-in-\epsilon centred-design perturbation theorem. (b) We give the first positive TB-BFI-FB result: a cost-aware Sequential-Halving algorithm whose cost-weighted simple-regret upper bound matches a new information-theoretic lower bound up to the standard K \log K adaptive-vs-oracle gap. Empirically, theorem-aligned synthetic experiments certify the C^\star-winner in 95–100% of trials; on a char-level transformer learning-rate-crossover our extrapolator picks the right arm in 3/3 seeds where Successive Halving picks 0/3; and on a flip-rank Chinchilla instance the joint-law procedure reaches 70–87% accuracy where the univariate c = N \cdot D baseline scores 0%.
PaperID: 2145, Poster
Authors: Ming Liu
Abstract: Prior work localizes the alignment tax---instruction tuning's degradation of structured generation---to transformer layers, but cannot determine which projections within a layer carry the functional change. We identify a consistent sub-layer hierarchy across three primary SwiGLU decoder-only families (7--9B parameters, with Yi-1.5-9B supplementary): among equal-parameter MLP projections, W_down > W_up > W_gate (2.5:1.5:1)---an architecturally grounded ordering (reproduced under continual pretraining) not resolved by concurrent work. Within attention, W_V/W_O gradient magnitudes exceed W_Q/W_K by 1.76--3.91×. Three independent methods---weight patching, gradient analysis, and training-time exclusion---converge on the same within-layer structure; a continual-pretraining control confirms the within-MLP hierarchy is architectural (not alignment-specific), but the layer-level distribution of harmful modifications is---this task-specific signal enables targeted rollback (8.9× over architectural-prior heuristics on held-out data, p=0.002). Surgical Alignment Reversal (SAR), a training-free rollback of attribution-ranked components, recovers 47% of the alignment tax (cross-entropy metric, Qwen; Llama/Mistral: 2--3× random) on held-out cross-dataset data (in-domain ratios 2.8--5.3×); Output-Constrained DPO (OC-DPO) largely eliminates the tax while Q/K exclusion does not---both validated on attribution-independent data. Safety evaluation shows no significant degradation: refusal-rate |\Delta| \le 1.4 pp (n=1,000), no significant ASR increase under adversarial attack (|\Delta\textASR| \le 6 pp, n=50); Yi-1.5-9B is excluded due to safety--structure overlap (-13.6 pp). Attribution replicates at 14B; gradient asymmetry persists at 72B.
PaperID: 2146, Poster
Abstract: We study fixed-confidence best-arm identification in Bernoulli bandits. A learner sequentially samples from K arms and aims to identify the unique best arm while keeping the error probability at most \delta. We ask whether the classical lower bound on expected sample complexity can be matched asymptotically without ongoing forced exploration. We propose a posterior-tracking algorithm that first samples each arm once, then repeatedly draws a bandit instance from the posterior, computes the oracle allocation for that instance, and randomly selects the next arm according to that allocation. Stopping is based on a Bernoulli generalized likelihood ratio test. Since the algorithm has no explicit forced-exploration mechanism after initialization, the main technical challenge is to show that posterior randomness alone sustains sufficient exploration. We prove that, with summable tail probabilities, every arm is sampled at least logarithmically often. This yields an integrable stabilization time after which the true best arm remains the empirical leader and the corresponding generalized likelihood ratio statistic grows linearly at the information-theoretic lower-bound rate. Consequently, the proposed procedure is \delta-correct and achieves asymptotically optimal expected sample complexity as \delta \to 0.
Abstract: As foundation models grow in capability, the ability to efficiently and reliably control their behavior becomes critical. Fine-tuning these models can be costly, and while prompting can be practical for controllability, it remains fragile due to models' high sensitivity to exact prompt wording and structure. This brittleness has driven interest in activation steering techniques that offer more stable and predictable control over model behavior. However, existing activation steering methods require per-concept optimization, which makes them ill-suited to deployment scenarios where the concept set is large, evolving, or only specified at request time: each new concept incurs at least minutes of optimization on the target model. We propose HyperTransport, a hypernetwork framework that amortizes this cost by mapping embeddings from a pretrained encoder (CLIP in our instantiation) directly to intervention parameters, trained end-to-end using an optimal transport loss. Once trained, HyperTransport produces each new intervention in a single hypernetwork forward pass, 3600–7000x faster than per-concept fitting. On concepts unseen during training, it matches the strongest per-concept baselines at inducing the target concept. By decoupling concept representation from intervention prediction, HyperTransport combines three capabilities that no existing approach offers as a set: amortized steering for open-ended concept sets, continuous interpretable strength control, and cross-modal conditioning where reference images can directly steer text-based generation. We validate HyperTransport on DMD2 and Nitro-1-PixArt across 167 held-out test concepts via CLIP-based metrics, a VLM-as-a-judge evaluation, and a user study. In pairwise comparisons, both human and VLM judges prefer HyperTransport over prompting ~2x as often.
PaperID: 2148, Poster
Abstract: When does an AI assistant improve human decision-making, and when does it hurt? We consider the setting in which a human worker makes a binary decision by observing her own signal and receiving a binary or probabilistic recommendation from an AI. We develop a behavioral model which considers both her belief regarding the overlap of her signal with the AI as well as the trust she has in the AI when considering its recommendation. The model yields a phase diagram in AI capability and signal overlap, partitioning the parameter space into complementarity, impairment, and automation regions. A worker who neglects signal overlap (correlation neglect) can experience performance loss when incorporating AI into her decision and is eventually dominated by AI alone as AI capability grows, with the complementarity range narrowing as overlap rises. Trust calibration shifts region boundaries: undertrust shrinks the impairment region but expands automation, while overtrust expands impairment but extends complementarity to be preferred at higher AI capability levels. Applying the framework to ChatBench (643 workers, GPT-4o and Llama-3.1-8b), AI assistance raises accuracy by 21 percentage points, but workers undertrust the AI by roughly a factor of four. The model yields design insights linking AI deployment choices, worker behavioral parameters, and outcome quality.
Authors: Federico Tomasi, Dmitrii Moor, Alice Wang, Mounia Lalmas
Abstract: Discrete diffusion models generate structured sequences by progressively unmasking tokens, but enforcing global property constraints during generation remains an open challenge. We propose primal-dual guided decoding, an inference-time method that formulates constrained generation as a KL-regularised optimisation problem and solves it online via adaptive Lagrangian multipliers. At each denoising step, the method adds a constraint-dependent bias to token logits, with multipliers updated by mirror descent based on constraint violation. The approach requires no retraining, introduces no additional model evaluations, supports multiple simultaneous constraints, and provides formal bounds on constraint violation. Because the bias is additive on logits, it preserves the model's learned distribution while steering outputs toward desired properties, yielding a favourable trade-off between constraint satisfaction and distributional fidelity. We demonstrate the method on topical text generation, molecular design, and music playlist generation, showing that a single algorithm instantiated via domain-specific scoring functions achieves high constraint satisfaction while maintaining diversity and validity, without additional training or inference cost.
PaperID: 2150, Poster
Abstract: Chain-of-thought (CoT) reasoning has emerged as a powerful paradigm for extending the computational capabilities of transformer-based models. A growing body of recent work has begun to uncover the fine-grained computational power of transformers augmented with CoT, as well as their connections to classical complexity theory. In this work, we advance this line of research by establishing the first hierarchy theorem for \CoT-based computation. Our main result shows that for any reasonable resource bound r(n) \geq n and any \varepsilon > 0, \mathrmCoT(r(n)) \subsetneq \mathrmCoT(r(n)^1+\varepsilon), where \mathrmCoT(r(n)) denotes the class of decision problems solvable by transformers using O(r(n)) CoT reasoning steps on inputs of length n. Our result shows that increasing the number of CoT reasoning steps provably yields strictly greater computational power, establishing CoT steps as a fundamental resource in transformer-based computation. Unlike classical hierarchy theorems, which are typically proved via diagonalization, our approach leverages known coarse-grained simulations between CoT-augmented transformers and Turing machines in both directions. Building on these connections, we adapt the classical padding technique from complexity theory to this setting to obtain a tight hierarchy. A key technical contribution of our work is the implementation of padding within the transformer-based CoT model.
PaperID: 2151, Poster
Abstract: We introduce Split-RL, a reinforcement learning (RL) algorithm based on decision trees. It uses guidance labels to explicitly partition the state-space and isolate localized conflicting signals. RL policies are often trained by mixing contradictory feedback that is localized in space and time. For example, a robot may prioritize speed in open areas versus precision in tight gaps. These regions are often easily detectable from sensor data or a human operator. Standard neural networks (NNs) lack the structural mechanism to isolate these signals. While Gradient Boosting Trees (GBT) provides the inductive bias to explicitly partition the state-space, standard tree-fitting procedures fit only aggregated gradients, merging contradictory updates before they reach the leaves. Split-RL leverages this bias to integrate guidance labels into tree construction, routing these updates into disjoint regions of the state-space to prevent signal interference during training. Such labels can arise from simple heuristics, sensor-derived events, or existing rule-based systems, allowing Split-RL to incorporate rule-derived structure while retaining a learned policy. We theoretically show how early gradient aggregation loses objective-specific information and characterize when the Split-RL score isolates conflicting updates. Finally, we evaluate Split-RL on constrained tasks and offline imitation learning from mixed-quality datasets. Across constrained domains with spatial, rare-event, and temporal conflicts, Split-RL consistently outperforms existing methods by achieving low-cost, feasible solutions while maintaining competitive rewards. In offline settings, Split-RL localizes the tradeoff between imitation and policy improvement, outperforming other methods and achieving near-expert performance across varying dataset sizes.
Abstract: Model-based learning agents use learned world models to predict future states, plan actions, and adapt to new environments. However, the process of updating world models from collected experience creates a training-time attack surface: adversarially poisoned fine-tuning trajectories can manipulate the learned dynamics and thereby corrupt downstream planning. In this paper, we propose SWAAP, the first two-stage data poisoning framework for learned world models. In the first stage, SWAAP identifies a harmful target world model that induces low-return behavior under planning while remaining close to clean dynamics, using first-order bilevel optimization enabled by a transition-gradient theorem. In the second stage, SWAAP realizes this target through stealth-constrained gradient matching, modifying only a limited fraction of fine-tuning transition targets so that the induced training gradients steer the victim model toward the adversarial target, while a prediction-error regularizer encourages the poisoned targets to remain close to the world model's natural approximation error. To assess attack stealthiness, we evaluate defenses and detectability across three stages of the poisoning pipeline: pre-training detection of poisoned transitions, robust training during fine-tuning, and test-time monitoring of the resulting world model. Across diverse continuous-control tasks, SWAAP causes substantial performance degradation while keeping poisoned transitions close to clean data and avoiding detection by the evaluated defenses. These results reveal a practical vulnerability in world-model adaptation pipelines and highlight the need for robustness methods that protect both world-model training data and learned dynamics.
PaperID: 2153, Poster
Abstract: LLMs are increasingly deployed as agents within ecosystems where they compete for attention, are reused across contexts, and are reinforced through feedback loops. In such systems, behavioral traits spread not only because they are better, but because system design rewards some behaviors over others. We develop an analytical framework for LLM ecosystems, grounded in evolutionary dynamics, that links system-level selection to the evolution of trait distributions by treating exposure allocation as a source of selection pressure. We characterize the resulting regimes with stability-aware diagnostics (tangent stability and invasion exponents). We test the theoretical predictions in a minimal AI-native social platform in which LLM agents generate posts, evaluate one another, and compete for future visibility. The experiments support the theory: stronger selection increases concentration and pushes endpoints toward local instability; diversity support stabilizes specialization only within a bounded regime; and early trajectory fluctuations forecast later instability. We identify phantom diversity as a failure mode: endpoints can appear coexisting or specialized while remaining locally unstable or invadable. These results show that diversity audits should measure not only endpoint geometry, but also dynamic stability and early-warning signals.
Abstract: User modeling aims to use language models (LMs) to mimic an individual’s behavior from a corpus of past context–action pairs (e.g., conversation turns), enabling the simulation of users in settings like behavioral science, human–AI collaboration, and market research. Recent approaches augment these corpora with synthesized reasoning traces, typically generated by conditioning on both context and action. However, such conditioning constitutes post-hoc rationalization rather than reasoning: the trace is guaranteed to justify the action, but may not encode the underlying latent causal decision paths. We propose Recon, which uses action reconstruction to score reasoning traces by their predictive power: given a context and candidate reasoning, a reconstruction model predicts the action, and reconstruction fidelity determines reasoning quality. Across four domains, Recon achieves a 54.7% win rate over Backward Synthesis, a standard post-hoc rationalization baseline. Further, we find that training a reasoning synthesis model with rewards derived from Recon improves downstream user modeling performance, achieving a win rate of up to 70.0% over baselines. We further show that Recon-synthesized reasoning transfers across models, and improves user modeling beyond the reconstruction model. Our work demonstrates that post-hoc rationalization is insufficient for reasoning synthesis, and that useful and interpretable reasoning should naturally elicit the action from the context.
PaperID: 2155, Poster
Authors: Ke Sun
Abstract: Natural gradient descent is a theoretical pillar of second-order optimization for neural networks. Existing implementations either rely on structural approximations of the curvature matrix or incur substantial computational and memory overhead. Direct application remains difficult at deep-network scale. We build an optimizer based on an unbiased stochastic estimator of the Fisher information matrix. It employs a truncated exponential moving average metric with damping, which we call the Stochastic Truncated Metric (STM), and is otherwise closely aligned with the original natural gradient. This gives a matrix-free parameter update with strictly O(Kd) time and memory per step, where K is a small constant. It is applicable across network architectures without layer-wise assumptions. We bound the error induced by the finite-memory truncation. We show that STM is stable and efficient on vision and language benchmarks.
Abstract: Long-horizon embodied tasks remain a fundamental challenge in AI, as current methods rely on hand-engineered rewards or action-labeled demonstrations, neither of which scales. We introduce ASH, an agentic system that learns an embodied policy from unlabeled, noisy internet video, without reward shaping or expert annotation. ASH follows a self-improvement loop; when it gets stuck, ASH learns an Inverse Dynamics Model (IDM) from its own trajectories, and uses its IDM to extract supervision from relevant internet video. ASH uses unsupervised learning to identify key moments from large-scale internet video and retains them as long-term memory --- allowing it to tackle long-horizon problems. We evaluate ASH on two complementary environments demanding multi-hour planning: Pokemon Emerald, a turn-based RPG, and The Legend of Zelda: The Minish Cap, a real-time action-adventure game. In both games, behavioral cloning, retrieval-augmented and zero-shot foundation-model baselines plateau, while ASH sustains progression across our 8-hour evaluation. ASH reaches an average of 11.2/12 milestones in Pokemon Emerald and 9.9/12 in Legend of Zelda, while the strongest baseline gets stuck in both environments at an average of 6.5/12 and 6.0/12 milestones, respectively. We demonstrate that self-improving agents are a scalable recipe for long-horizon embodied learning.
Abstract: Taming diffusion models for generative segmentation has attracted increasing attention. While existing approaches primarily focus on architectural tweaks or training heuristics, there remains a limited understanding of the intrinsic mismatch between continuous flow matching objectives and discrete perception tasks. In this work, we revisit diffusion segmentation from the perspective of vector field learning. We identify two key limitations of the commonly used flow matching objective: gradient vanishing and trajectory traversing, which result in slow convergence and poor class separation. To tackle these issues, we propose a principled vector field reshaping strategy that augments the learned velocity field with a detached distance-aware correction term. This correction introduces both attractive and repulsive interactions, enhancing gradient magnitudes near centroids while preserving the original diffusion training framework. Furthermore, we design a computationally efficient, quasi-random centroid encoding scheme inspired by Kronecker sequences, which integrates seamlessly with an end-to-end pixel neural field framework for pixel-level semantic alignment. Extensive experiments consistently demonstrate significant improvements over vanilla flow matching approaches, substantially narrowing the performance gap between generative segmentation and strong discriminative specialists.
PaperID: 2158, Poster
Abstract: Halpernian actual causation, the formal study of which events in a specific situation caused a specific outcome in acyclic structural causal models, underpins legal responsibility, moral blame, and the assessment of causal harm. Despite the rich logical framework, no prior sound and complete enumeration procedure is known for finding actual causes in a given setting, blocking downstream tasks such as responsibility attribution, blame assessment, and causal-harm grading that aggregate over the full cause set. We close this gap by introducing . We prove that for every actual cause there exists a contingency set such that the joint intervention set appears as a node of this map, and present an algorithm for that is sound and complete for enumerating all actual causes of any target event. We compare TrAC empirically with existing single-cause baselines on random Boolean structural causal models, demonstrating practical feasibility for graphs of up to 15 nodes.
Abstract: Sequential Monte Carlo (SMC) samplers for reward-guided diffusion models often suffer from rapid lineage collapse: a few high-reward particles dominate the population within a handful of resampling steps, destroying diversity and degrading sample quality. We propose a variance-decomposition framework for reward-guided diffusion SMC that separates continuation variance V_t^\mathrmcont from residual variance V_t^\mathrmres, revealing that high offspring-count variance under the commonly used multinomial resampling drives this collapse. This motivates \textscVASR (Variance-Aware Systematic Resampling), which addresses both variance terms via variance-optimal mass allocation m_t \propto w_t e^r_t (minimizing V_t^\mathrmcont) and systematic resampling (controlling V_t^\mathrmres). For latent diffusion models where intermediate rewards are noisy due to stochastic continuations, we propose \textscVASR-Max, a deliberately biased high-selection variant for variance-sensitive reward optimization. Both methods are training-free, fully parallelizable, and add only linear overhead. On MNIST and CIFAR-10, \textscVASR achieves as high as 26% better FID than prior SMC methods while remaining ~\!66× faster than MCTS-based value methods at matched compute. On text-to-image generation, \textscVASR-Max consistently outperforms the strongest SMC baseline across compute budgets and matches MCTS-based methods within 2.5--3% reward at high budgets while being approximately 4× faster.
Authors:
Yuqing Wang, Zhijie Lin, Ceyuan Yang, Yang Zhao, Fei Xiao, Hao He, Qi Zhao, Zihan Ding, Fu-Yun Wang, Shuai Wang, Youliang Zhang, Haoqi Fan, Xihui LiuAbstract: Unified multimodal models (UMMs) aim to handle perception and generation in a single model. Yet existing UMMs still rely on a frozen, separately pretrained VAE for image generation, imposing a structural bottleneck. Naively removing it introduces a quality gap, as the model must learn both high-level structure and low-level details from raw pixels. In this paper, we propose Representation Forcing (RF), a technique that closes this gap by making representation prediction a native capability of the decoder. Concretely, RF forces the decoder to autoregressively predict visual representations as intermediate tokens before pixels; these tokens then stay in context to guide pixel diffusion within the same backbone. By turning representations from perception outputs into generation targets, RF eliminates the need for any external generative latent space. We find that RF benefits both understanding and generation. On image generation, our pixel-space model with RF matches state-of-the-art VAE-based unified models. On image understanding, pixel-space RF generally outperforms its VAE-based variant. Together, these results offer an effective step toward end-to-end, bottleneck-free UMMs.
PaperID: 2161, Poster
Abstract: Flow Matching (FM) learns generative transport by fitting continuous-time motion from a simple source distribution to the data distribution. Most existing methods use first-order bridges: once a source and a target sample are paired, the path is a straight motion with constant velocity. FM with optimal transport (OT) improves the pairing, but the bridge itself remains linear, limiting its ability to model curved motion, acceleration, and changing directions. A natural remedy is to use second-order phase-space dynamics; however, constructing such bridges requires a target-side terminal velocity, which static datasets typically do not provide. We propose Variational Terminal-Velocity Flow Matching (VTV-FM), a second-order FM framework that derives the missing velocity by minimizing acceleration energy, yielding a closed-form closure for static data. The same minimum-acceleration variational construction also defines the OT pairing cost and the acceleration targets used for training. Experiments on low-dimensional datasets, PDE-governed physical fields, and CIFAR-10 show that VTV-FM improves transport geometry and generation quality over first-order and existing high-order FM baselines.
Abstract: Efficient operator scheduling is a fundamental challenge in software compilation and hardware synthesis. While recent differentiable approaches have sought to replace traditional ones like exact solvers or heuristics with gradient-based search, they typically rely on categorical distributions that fail to capture the ordinal nature of time and suffer from a parameter space that scales poorly. In this paper, we propose a novel differentiable framework, GauS, that models operator scheduling as a stochastic relaxation using Gaussian distributions, which fully utilize modern parallel computing devices like GPUs. By representing schedules as continuous Gaussian variables, we successfully capture the ordinal nature of time and reduce the optimization space by orders of magnitude. Our method is highly flexible to represent various objectives and constraints, which provides the first differentiable formulation for the complex pipelined scheduling problem. We evaluate our method on a range of benchmarks, demonstrating that GauS achieves Pareto-optimal results.
Abstract: Modern deep learning commonly relies on AdamW with prescribed learning rate schedules, but recent works challenge both components: Schedule-Free optimization removes explicit schedules via iterate averaging, and Muon improves the update geometry by orthogonalizing momentum for matrix parameters. Despite Muon's strong empirical performance, its underlying mechanism remains partially understood. We study Muon through the river-valley loss landscape, where useful training progress occurs along a flat, low-curvature bulk subspace (the river), while high-curvature dominant directions form steep valley walls that induce oscillations. We empirically show that while Muon's orthogonalization accelerates river progress by increasing the bulk component, it also amplifies dominant-direction noise, causing oscillatory trajectories. Building on this, we propose Anytime MUon with Stable gradient Evaluation (AMUSE), which integrates Muon's rapid bulk progress with the stabilizing effect of Schedule-Free averaging. AMUSE uses a time-varying interpolation coefficient that initially evaluates gradients near the fast Muon sequence for rapid adaptation, then gradually shifts toward the stable averaged sequence to suppress valley-wall oscillations. As a result, AMUSE requires no learning rate schedules and supports anytime training. Across vision tasks and large language model pretraining, AMUSE consistently improves the performance-iteration Pareto frontier over (Schedule-Free) AdamW and Muon.
PaperID: 2164, Poster
Abstract: Parameter-efficient fine-tuning (PEFT) reduces the cost of adapting foundation models by focusing training on a small parameter subset. Complementary to this idea, we introduce RoSA (Rotational Sparse Adaptation), which narrows adaptation to a subset of layers at a time. RoSA freezes lower layers close to the input throughout training and rotates a trainable block over later layers, consecutively increasing the number of frozen layers close to the input. This design reduces optimizer-state memory, shortens backpropagation, and even forward propagation if activations at the last frozen layer are cached. Because RoSA is orthogonal to the choice of trainable parameterization, it can be combined with PEFT methods or sparse optimizers inside each active block. Experiments across multiple LLM architectures and tasks show that RoSA reduces peak memory while maintaining strong fine-tuning performance.
Abstract: In physics-informed machine learning, a target function u^ is learned from noisy value observations y_i=u^ (x_i)+ \varepsilon_i, together with differential information, given either by noisy observations d_j=(Du^ )(z_j)+\xi_j or by a known physical constraint Du^ =v . We consider the setting where D is a linear differential operator and analyze a physics-informed kernel estimator \hat u combining n value observations and m differential observations. In this context, we ask how much can differential information improve predictions, and how does this improvement depend quantitatively on n , m, and D. We prove finite-sample bounds, supported by numerical simulations, revealing a two-regime structure for the prediction error. When m is limited, the rate depends jointly on n and m; when m exceeds a problem-dependent threshold, the rate saturates and matches the oracle rate obtained when the perfect constraint D \hat u = Du^ is imposed. Examples are discussed for Sobolev spaces which are reproducing kernel Hilbert spaces and include partial Laplacian constraints on the torus and gradient observations on bounded domains. These examples illustrate the range of possible learning rate improvements - from the standard nonparametric n^-1/4 to the parametric rate n^-1/2. Finally, we derive physically consistent rates in a stronger norm that jointly controls the errors in \hat u and D\hat u.
PaperID: 2166, Poster
Authors:
Ye Zhou, Bangqi Li, Haodi Wang, Yu Guo, Libin Jiao, Rongfang BieAbstract: Federated Graph Learning (FGL) aims to collaboratively train Graph Neural Networks (GNNs) across distributed subgraphs with a primary challenge of missing cross-client connections. Existing approaches predominantly rely on a retrieval-based paradigm, which reconstructs missing context by accessing external information across clients. However, such cross-client transmissions inevitably expand the attack surface for privacy inference and incur significant communication overhead. In this work, we propose FedLMC, a novel framework that eliminates the necessity of any node information sharing and prior global adjacency knowledge. FedLMC exploits a dual-stage message compensation mechanism to generate informative and expressive node representations. For the semantic information loss of external neighbors, we propose Local Semantic Compensation, which exploits semantic expansion via semantic surrogates to compensate for missing semantic context. For the structural information loss induced by semantic expansion, we propose Community Structural Compensation, which exploits structural anchoring via community anchors to compensate for underlying structural context. Extensive experiments on twelve datasets (both homophilic and heterophilic) demonstrate that FedLMC establishes a new state-of-the-art performance with superior generalization and robustness to highly fragmented graph data.
PaperID: 2167, Poster
Authors: Enoch H. Kang, Kyoungseok Jang
Abstract: Bellman residual minimization (BRM) provides a scalable, gradient-based approach to offline reinforcement learning in large state spaces. While globally convergent gradient-based soft (i.e., entropy-regularized) BRM methods for neural networks have recently been established, their statistical complexity under stochastic gradient descent remains largely unknown. In this paper, we address this theoretical gap for soft BRM. Through a novel Lyapunov-based analysis, we establish an \mathcalO(1 / n) average argument stability bound, which translates directly into a \mathcalO(1 / n) statistical complexity for the soft BRM objective.
PaperID: 2168, Poster
Abstract: Learning solution operators for partial differential equations (PDEs) requires models that can represent spatial structure across grids and scales. This becomes especially important in time-dependent problems, where the learned operator is often used autoregressively: each predicted state is fed back as input to predict future states. In this setting, small spatial errors can be repeatedly propagated through the learned dynamics, leading to instability and drift, vitiating long-term solution fidelity. We propose Mesh-Free Convolutions (MFCs), filters defined in continuous space and discretized only at evaluation time. This allows the same learned filters to be applied across resolutions. Rather than imposing a grid-dependent stencil or a fixed spectral cutoff, MFCs learn smooth scale-dependent filtering and include directional terms for transport-like behavior. Across fluid-dynamics and weather-forecasting benchmarks, MFC-based models produce stable long-horizon rollouts, transfer reliably across resolutions, and better preserve long-time statistics. These results suggest that continuous, mesh-free filtering is a useful inductive bias for PDE operator learning, especially in autoregressive prediction.
PaperID: 2169, Poster
Abstract: Semantic caching, which reuses responses to semantically similar requests via their embeddings, has seen growing adoption in LLM serving, offering faster responses and reduced costs. Yet existing schemes are fundamentally vulnerable to cache-collision attacks, wherein an adversary pollutes the cache by injecting crafted queries, corrupting responses to subsequent legitimate requests. We present LaCache, a novel semantic caching scheme that addresses this vulnerability through a conceptually simple yet principled redesign. The key insight is that while the adversary has full control over the adversarial query, it has far less control over its response, which must simultaneously satisfy multiple semantic constraints. Rather than checking only the cache hit of a query, LaCache additionally checks the cache hit of its first k (speculatively) decoded tokens. This design yields two concrete benefits. First, it provides formally guaranteed resilience against cache-collision attacks: we prove that it is impossible to craft adversarial queries that simultaneously elicit malicious responses and collide with benign queries. Second, the enriched index supplies additional semantic context for cache retrieval, improving response relevance. Empirical evaluation across diverse LLMs and benchmarks validates both LaCache's security guarantees and efficiency gains, pointing to a promising direction for robust semantic caching.
PaperID: 2170, Poster
Authors: Yilin Zhuang, Noah Zambrano, Karthikeyan Duraisamy
Abstract: The quadratic attention cost of Vision Transformers (ViTs) forces a sharp compromise between discretization and rollout horizon, and binds tightest in the fine-scale PDE regime where shocks, reaction fronts, and material interfaces occupy a small, time-varying fraction of the domain. Further, conventional neural surrogates do not adapt to localized features dynamically. We propose WAMRViT, a ViT that tokenizes its input as a balanced quadtree under a wavelet-inspired refinement criterion, encodes position and refinement level jointly via a 3D rotary positional embedding with a soft level axis, and regrids in cell space at inference to remain stable over long rollouts. A multi-scale variant additionally keeps each leaf at its native source resolution and defers the per-level resolution differential to the model. Unlike prior adaptive-tokenization work, the tokenizer imposes no a priori token count; we evaluate under fully adaptive topology over long autoregressive rollouts; and WAMRViT is the first machine-learning surrogate to natively tokenize multi-level Adaptive Mesh Refinement (AMR) data. We demonstrate that on uniform-grid datasets, WAMRViT exceeds a finest-patch uniform backbone on the region of interest while using a fraction of the tokens, with a structural interpolation cost on the global metric that the multi-scale variant closes at the initial rollout steps. On a complex AMR combustion problem whose extremely fine scale features uniform-grid baselines cannot represent natively and must project onto a coarser grid, WAMRViT operates directly on the adaptive cells and substantially reduces finest-level error at matched parameters. Code: https://anonymous.4open.science/r/wamrvit-review-74E1
Abstract: Sparse autoencoders (SAEs) are increasingly used to extract activation directions for inference-time steering, but their standard sparsity objective treats latent features as independent. This prior can be poorly matched to high-level safety behaviors, where refusal and harmful compliance appear to depend on distributed structure in activation space. We introduce Graph-Regularized Sparse Autoencoders (GSAE), a dictionary-learning method that learns safety-steering directions by smoothing SAE decoder vectors over a neuron co-activation graph and applying the resulting direction bank through a two-gate runtime controller. Empirically, GSAE improves selective refusal across JailbreakBench, HarmBench, and XSTest, increasing harmful-request refusal while keeping benign-prompt refusals low. On Llama-3-8B, adding graph regularization to an otherwise identical SAE bank-and-gating pipeline improves \Delta_s by 20.1 points on JailbreakBench and 16.8 points on HarmBench. GSAE outperforms activation-steering baselines and black-box guardrails, preserves benign-task performance, generalizes across Llama-3, Mistral, Qwen 2.5, and Phi-4, and remains strong under black-box and gray-box jailbreak attacks.
PaperID: 2172, Poster
Authors:
Anna Thomas, Georgios Zaverdinos, Petros Mandalis, Andreas Orfanoudakis, Akihiro Takino, Sohum Patnaik, David J Kraus, Panos Kostopoulos, Caroline Cotto, Nikos Tsiaparas, Dan Jurafsky, Aadit PatelAbstract: Diversifying protein sources away from animal agriculture is critical for climate change mitigation and food security, but formulating sustainable protein sources that match or improve upon the properties of their animal-based counterparts remains an expensive trial-and-error process. We frame this challenge as high-dimensional black-box optimization with sparse approximate solutions and formalize Expert-Guided Bayesian Optimization (EGBO), in which an expert, e.g. a human or LLM, selects a low-dimensional subspace for BO and may adaptively expand it over time. We decompose EGBO's suboptimality into a selection gap and an optimization gap, and characterize the coverage–dimension tradeoff governing when expert guidance helps. To support in silico prototyping before costly real-world deployment, we introduce FormulateBench, a suite of 24 plant-based formulation tasks, on which LLM-guided EGBO outperforms all tested baselines. When deployed to optimize two plant-based dairy products, EGBO improves utility, as assessed by a trained human panel, by 29% and 26% in 10 iterations each. In a comparison with a professional human food scientist given the same time budget, EGBO achieved near-perfect utility of 0.992, vs. 0.850 for the food scientist.
Abstract: Classifier-free guidance (CFG) has become a widely adopted and practical approach for enhancing generation quality and improving condition alignment. Recent studies have explored guidance mechanisms for unconditional generation, yet these approaches remain fundamentally tied to assumptions specific to diffusion models. In this work, we propose a spectrum weakening framework for visual autoregressive (AR) models. The method is training-free and condition-free by constructing a controllable weak model in the spectral domain without architectural modifications. We theoretically show that invertible spectral transformations preserve information, while selectively retaining only a subset of the spectrum introduces controlled information reduction. Based on this insight, we perform spectrum selection along the channel dimension of internal representations, which avoids the structural constraints imposed by diffusion models. We further introduce a spectrum-renormalization technique that maintains numerical stability during the weakening process. Comprehensive empirical and ablation studies confirm the effectiveness of our approach, including raster-scan, random-order, and scale-wise AR, spanning discrete and continuous modeling and covering class and text-condition scenarios, demonstrating high-quality unconditional generation and strong prompt alignment maintenance for conditional generation. Code will be made available.
Abstract: List-wise reranking arranges a request-specific pool of candidate items into an ordered slate that maximizes user satisfaction. Existing generative rerankers fall into two paradigms: Autoregressive (AR) rerankers construct the slate left to right and capture inter-item dependencies in the exposure list, but they suffer from error propagation because early mistakes affect subsequent slots. Non-autoregressive (NAR) rerankers predict all slots in parallel and avoid error propagation, but they weaken inter-item interaction modeling under a slot independence assumption. This raises a central question: is there a unified architecture that combines the strengths of both paradigms and delivers stronger reranking performance? We answer this question with UniRank, a unified list-wise reranking framework whose inference time variants recover AR and NAR rerankers as special cases. UniRank integrates bidirectional slate modeling into an iterative denoising process and fills the most confident slot at each step. To instantiate this framework for reranking, we introduce the Task Grounded Diffusion Interface (TGD), which performs denoising at the item level and restricts prediction to the request-specific candidate pool. TGD aggregates each item's semantic tokens into a single item embedding and scores each slot directly against the candidate pool. Experiments on Amazon Books, MovieLens-1M, and an industrial short video dataset show that UniRank consistently outperforms state-of-the-art baselines. Our code is available online.
PaperID: 2175, Poster
Abstract: We consider the problem of learning to adapt a foundation model in a federated setting, particularly the most realistic and general setting: 1) When the local data sets are sampled from different distributions but we want to learn a globally adapted model, 2) Where local agents enter and leave the federation asynchronously at each time tick which is beyond the control of the learning algorithm, and 3) Where the goal is continuous adaptation so that after each time tick, the learned adapter generalizes accurately for all participants that have been seen during training. We propose a simple idea called federated ) for this setting. In library-based adaptation, the system maintains a pool, or "library" of rank-1 "basis pairs". Groups of these pairs are distributed to local sites to be trained in LoRA fashion, and then separated and at communication. Library-based adaptation is designed to avoid problems with more conventional methods, such as those based on averaging. In particular, we demonstrate outperforms traditional averaging baselines in both communication and computation cost efficiency across a broad range of important settings, including heavy data skewness and high asynchrony.
PaperID: 2176, Poster
Abstract: Latent diffusion models (LDMs) achieve high-fidelity generation by moving diffusion from pixels to a compact latent space, but standard tokenizers are optimized for compression, not robustness under diffusion-time corruption. We formulate latent tokenization as diffusion-aware joint source--channel coding and diagnose standard LDM pipelines as missing an explicit channel-coding step. A diffusion-aware rate--robustness functional admits an explicit spectral characterization in G^\top G, yielding the Redundancy-Balance Principle: optimal redundancy is not uniform, but aligns with semantic sensitivity and equalizes marginal integrated-risk reduction across active directions. The resulting optima take weighted tight-frame and water-filling forms; isotropic bottlenecks are minimax-suboptimal under heterogeneous sensitivity, and a group-conditioned extension identifies rare groups as capacity-limiting under a fixed latent budget. We instantiate the theory as RB-LDC (Redundancy-Balanced Latent-Diffusion Coding), a lightweight tokenizer modification that inserts a learned coding layer before diffusion. Six controlled experiments validate the theorem sequence up to k=1024, with a 53.5× alignment gap, 659× minimax gap, and real-image VAE learnability via Adam recovery of the KKT spectrum with relative error below 0.5 percent. On SD-VAE and VA-VAE at ImageNet-256, RB-LDC outperforms isotropic uniform redundancy in all 20 perturbed metric cells over a 15-SNR grid; VA-VAE block-local obtains -15.40 average pFID and -105.79 peak pFID at \log_10\rho=+1, while preserving clean rFID parity (gap -0.0001). RB-LDC turns latent tokenizer design into a reliability problem for diffusion models.
Abstract: Recent work on causal abstraction, in particular graphical approaches focusing on causal structure between clusters of variables, aims to summarize a high-dimensional causal structure in terms of a low-dimensional one. Existing methods for learning such summaries from data assume that both the high- and low-dimensional structures are acyclic, which is helpful for causal effect identification and reasoning but excludes many high-dimensional models and thus limits applicability. We show that in the linear non-Gaussian (LiNG) setting, the high-dimensional acyclicity assumption can be relaxed while still allowing recovery of a low-dimensional causal directed acyclic graph (DAG). We further connect identifiability of this low-dimensional DAG to existing results: LiNG models with cycles are observationally identifiable only up to an equivalence class whose members differ by reversals of directed cycles; our low-dimensional DAG, which is invariant across all members of a given equivalence class, thus forms a natural representative of the class. While existing approaches for learning this observational equivalence class over high-dimensional variables have exponential time complexity, our low-dimensional summary is learned in worst-case cubic time and comes with explicit bounds on the sample complexity. We provide open source code and experiments on synthetic data to corroborate our theoretical results.
PaperID: 2178, Poster
Authors: Haseena R P, Muhammed Ashrah, Ajith Abraham
Abstract: Early-exit neural networks improve inference efficiency by enabling predictions at intermediate layers, but their effectiveness critically depends on the exit policy that decides whether to stop or continue computation. Existing policies, often based on confidence thresholds, enable early termination but offer limited insight into the reliability of such decisions. To address this, we introduce an explainability-aware framework that defines an attribution-grounded reliability signal for early-exit decisions. We propose Progressive Feature Attribution Maps (PFAM), which capture how feature importance evolves across exits, and the Interpretability-Based Early-Exit Score (IEES), a unified decision metric that integrates prediction confidence with attribution strength and cross-exit stability. To enable efficient deployment, we further introduce a lightweight proxy predictor that estimates IEES using inexpensive, attribution-free features, allowing attribution-informed decisions without incurring inference-time overhead. Experiments on convolutional architectures (ResNet, MobileNet, MSDNet) show that our approach improves accuracy–efficiency trade-offs, achieving up to 2.5% accuracy gains while reducing inference time by up to 24% compared to full-depth inference. Similar trends are observed on a vision transformer and a BERT-based language model, suggesting applicability beyond convolutional architectures.
PaperID: 2179, Poster
Abstract: Link prediction relies on pairwise structural proximity, but standard node representations are not explicitly designed so that their interactions reflect such signals. Existing structure-aware approaches address this gap through specialized architectures or handcrafted proximity features, often at the expense of generality or additional computational cost. In this work, we take a data-centric approach and propose \textLinkDPE, a learnable \underline\textdiffusion-based \underline\textpositional \underline\textencoding framework that injects structural proximity into the node features. Our key idea is that diffusion kernels, which capture multi-scale connectivity patterns, admit a factorization into node-wise embeddings whose inner products recover pairwise proximity scores. Based only on the observed training graph, \textLinkDPE constructs diffusion-based encodings across scales, adaptively selects complementary diffusion scales, and learns task-specific spectral weightings, enabling a wide range of link predictors to exploit structural information without architectural modification. Theoretically, we show that (i) diffusion proximity admits a node-wise kernel factorization, (ii) many classical link prediction heuristics are recovered as special cases of low-order walk and diffusion operators, and (iii) multi-scale spectral responses allow the model to capture a richer set of structural similarity patterns within a shared truncated spectral basis. Across 8 benchmark datasets and 5 representative base models, \textLinkDPE consistently improves link prediction performance over original features and positional-encoding baselines, yielding average gains of 6.97% over the unaugmented base models.
Abstract: Time-series data are inherently multiscale, spanning diverse temporal granularities from coarse trends to fine-scale dynamics. However, existing time-series generative models provide limited control over the temporal granularity of both inputs and outputs, restricting their ability to condition on user-provided coarse sketches and generate samples at a desired target granularity. To address this, we introduce TimeTok, a unified framework for Granularity-Controllable Time-Series Generation (GC-TSG), which generates time series at any target granularity from any coarser input (e.g., rough sketches) or without conditioning. At the core of TimeTok is a hierarchical tokenization strategy that maps time series into an ordered sequence of tokens, from coarse to fine temporal granularity. Our autoregressive generation process operates across these granularity levels, producing token blocks that are decoded back into continuous time series. This design naturally enables GC-TSG within a single framework, where controlling the number of token blocks provides explicit control over output detail. Experiments show that TimeTok excels at GC-TSG tasks while achieving state-of-the-art performance in standard generation. Furthermore, we showcase TimeTok's potential as a foundational tokenizer by training on multiple datasets with heterogeneous temporal granularities, verifying strong transferability that consistently outperforms models trained on individual datasets. To our knowledge, this is the first unified framework that covers the full generative spectrum for time series, offering a valuable foundation for models that benefit from diverse temporal granularities.
Abstract: Low-rank plus sparse decompositions, such as robust PCA, typically model the low-rank component as dense and the sparse component as corruption. This viewpoint is poorly suited to graph adjacency matrices, where sparsity is structural and dense low-rank factors destroy the combinatorial and computational properties of the graph. We propose signed biclique decomposition (SBD), a sparse low-rank decomposition for integer-weighted graphs that represents the structured component as a sum of sparse signed binary rank-one factors, together with a sparse signed residual. The resulting representation remains discrete, interpretable, sparse, and compatible with standard linear algebra. We formulate fixed-rank and greedy SBD objectives and derive a gain-based greedy algorithm with theoretical guarantees. Empirically, SBD recovers block structure in noisy graphs, while its factorized representation enables efficient sparse matrix-matrix multiplication (SpMM). On real graphs with strong shared-neighborhood structure, SBD achieves compression factors above 10 and SpMM speedups up to 6× over cuSPARSE, suggesting a practical alternative to dense low-rank decompositions for graphs.
PaperID: 2182, Poster
Abstract: Ensuring safety guarantees in offline reinforcement learning remains challenging, especially when safety constraints must hold almost surely. Moreover, as pre-specifying a single safety budget (constraint threshold) is often challenging, it is desirable to learn a foundation policy that can be deployed across a broad range of budgets. We introduce AEGIS (Almost-sure Epigraph-Guided Implicit Safety), a safe offline RL framework that can guide diffusion policy training via critics that respect almost-sure constraints across various feasible budgets. AEGIS characterizes the feasible set of initial state-budget pairs as the epigraph of a feasibility critic updated via the worst-case backup. Building on the proposed characterization, we extend implicit Q-learning (IQL) to train both feasibility and reward critics. We use these critics to bias a diffusion policy toward high-value feasible actions. Consequently, AEGIS turns diffusion from a generative prior into a safety-aware controller, enabling a single general policy to respect various budgets without further tuning. Empirical results on the DSRL benchmark and humanoid locomotion tasks show that AEGIS achieves high feasibility with competitive returns, generalizing across feasible constraint thresholds.
PaperID: 2183, Poster
Authors:
Shenghan Chen, Yiming Liu, Zhipeng Deng, Haolin Wang, Jiale Zhou, Zhijian Wu, Lu Xiankai, yafei ou, Yefeng ZhengAbstract: Balancing performance trade-offs on long-tailed data distributions remains a long-standing challenge in visual recognition. Existing methods mainly improve tail classes through re-balancing, representation learning, or data augmentation, but the underlying cause of tail classes degradation is still insufficiently explored. In this paper, we find that standard long-tailed training induces background-biased representation and optimization: tail classes suffer larger background distribution shifts and become increasingly driven by background gradients. This reveals that tail degradation is not merely caused by insufficient samples, but also by the learning of irrelevant background features. To tackle this issue, we propose Object-focus Background Debiasing (OFBD), a framework that mitigates background bias from both distribution and optimization perspectives. Specifically, Foreground-guided CutMix preserves target-related foregrounds while diversifying complementary backgrounds, and Background-guided Feature Rectification suppresses background-biased features without learnable parameters or additional training. Extensive experiments show that our method improves overall accuracy, achieves significant tail-class gains, and can serve as a plug-in for mainstream long-tailed methods without external data or pretrained recognition models.
PaperID: 2184, Poster
Authors: Yunpeng Zhou
Abstract: Native multimodal models build rich internal representations across visual, multimodal, and language streams, yet their answers are usually read only from the final layer. We study this mismatch between where task information is available and where the model is allowed to read from. We formalize it with the availability--accessibility gap (AAG), whose primary form compares a small-budget internal readout reference with a learned final-only parity head under the same supervised answer space and loss. Across document, chart, diagram, counting, compositional, and broad multimodal reasoning tasks, we find that task-relevant information often appears at internal sites before it becomes accessible to the final output. We then introduce adaptive internal readout (AIR), a small policy that selects a few internal sites from the full model stack. Under matched budgets, AIR closes much of the gap and outperforms static fusion, dense fusion, visual-only readout, and random site controls. Ablations show that the selected sites are functionally important for the learned readout pathway.
PaperID: 2185, Poster
Abstract: Backdoor attacks on molecular graph neural networks (GNNs) are typically evaluated as abstract graph edits, but real molecular learning pipelines do not train on arbitrary graphs. Molecular records must first survive parsing, sanitization, canonicalization, and graph-string consistency checks. We formalize this overlooked admission stage as ChemGuard, an operational protocol for testing whether a submitted molecular record can enter a realistic learning pipeline, while complementing existing defenses.ChemGuard admits a record only when its molecular string is sanitizable and the graph reconstructed from that string matches the submitted molecular graph. Under this operational view, many existing graph-based backdoors lose much of their apparent efficacy because their poisons are chemically invalid or representation-inconsistent. We then show that admission checks alone are insufficient to rule out molecular backdoors. We propose ChemBack, an admission-aware molecular backdoor attack that constructs chemically feasible motif-anchor attachments and ranks admitted candidates by fingerprint-based Tanimoto similarity to clean target-class molecules. ChemBack is model-free during trigger selection, using molecular structures, target labels, fingerprints, and public validity checks, but no victim model, surrogate GNN, learned embedding, gradient, logit, or training-code access. Across molecular benchmarks, validators, architectures, and defenses, ChemBack achieves high attack success with fully admitted poisons while preserving clean accuracy. Our results reveal a two-sided lesson, chemistry-aware admission suppresses many graph-only backdoors, yet chemically valid and target-aligned molecular backdoors remain a practical threat.
PaperID: 2186, Poster
Abstract: Randomized trial findings are routinely generalized to broader target populations to inform decisions in medicine and public policy. When the trial sample is misaligned with the target population, generalized effect estimates can yield suboptimal decisions. We identify and characterize the target subgroups that are underrepresented in the trial cohort in time-to-event settings. Here, underrepresentation arises in two distinct ways: baseline underrepresentation, when a subgroup's covariate profile is rare in the trial relative to the target, and time-varying underrepresentation, when differential censoring erodes a subgroup over follow-up. Existing diagnostics either conflate these mechanisms or address only the baseline case. We show that the asymptotic variance of the efficient generalizability estimator admits a decomposition in which only two terms can diverge as representation weakens: a sampling-weight term capturing baseline misalignment and a censoring-weight term capturing differential loss to follow-up. Underrepresented subgroups are therefore precisely those the trial estimates with poor precision, and the decomposition reveals which channel is responsible.We exploit this operationally in TRACE (Temporal Rashomon Analysis of CEnsored data), which combines a self-normalized variance objective with a tree-based search to localize underrepresented subgroups, attribute each to baseline mismatch or censoring, and pinpoint the follow-up intervals where censoring drives the variance inflation; the joint problem is solved by a parametric dynamic program nested inside a Rashomon ensemble of near-optimal trees. On synthetic benchmarks and a case study generalizing the National Job Training Partnership Act (JTPA) Study to U.S. JTPA-eligible adults, the framework recovers underrepresented regions, correctly attributes each mechanism, and identifies the follow-up windows driving temporal underrepresentation; refining the trial and target samples to reliably generalizable subgroups substantially reduces generalized effect-estimate variance.
Authors:
Eduardo Figueiredo, Frederik B Mathiesen, Julian F Schumann, Jens Kober, Arkady Zgonnikov, Luca LaurentiAbstract: Neural networks are expressive predictors capable of modeling rich, multi-modal outcome distributions. However, this expressive power comes with challenges in ensuring robust predictions—a critical requirement in safety-critical domains. Randomized smoothing is a leading technique for improving robustness, particularly against adversarial perturbations. Yet, in stochastic multi-modal regression settings, randomized smoothing often fails due to mode collapse, yielding averaged predictions that do not reflect the underlying distribution. To address this limitation, we propose clustered \alpha-smoothing, a framework that (1) partitions noisy samples using an arbitrary clustering algorithm, (2) applies \alpha-smoothing locally within each cluster, and (3) combines the resulting predictions into a mixture distribution. By interpreting the smoothing distribution as a mixture of \alpha-smoothers, we derive a lower bound on the probability that the smoothed prediction lies within a union of compact regions corresponding to distinct modes. We empirically evaluate our framework on two benchmarks, demonstrating substantial improvements over state-of-the-art methods. In stochastic trajectory prediction on a driving simulator dataset, our approach achieves, on average, a 33% lower Wasserstein distance to the ground-truth distribution compared to \alpha-smoothing. In quadrotor control, where modes correspond to distinct feasible paths to a target, our method reduces the collision rate by 81% relative to the state-of-the-art randomized smoothing.
PaperID: 2188, Poster
Abstract: We develop a unified framework for flow-based reinforcement learning (RL) grounded in diffusion–flow duality and establish its theoretical foundations. The frameworks enables RL for deterministic flow matching inference by introducing diffusion stochastic inference dynamics that supports exploration while admitting deterministic deployment. We show that our frameworks subsumes several existing flow-based RL methods as special cases and inspire effective now methods. Then we provide theoretical foundations for the frameworks by establishing the correctness and performance transfer guarantee. Specifically, we prove that the policies optimized under stochastic dynamics close to deterministic dynamics at deployment when the pretrained model is well-trained. Moreover, we show that the policy improvement achieved under training-time SDE inference transfers to deployment-time ODE inference. Finally, we conduct experiments to validate our theoretical results.
PaperID: 2189, Poster
Authors:
Tan K Ze, Ran Zhang, Feng Yichao, Matthew H Jie, Dong FangAbstract: Video understanding is commonly built on action-triggered responses or scene-level aggregation, implicitly assuming clear temporal boundaries. This assumption breaks in dense-action streams such as gameplay and sports, where overlapping actions induce continuous latent dynamics without explicit transitions. The challenge is further amplified in streaming settings, where models operate under strict causal constraints and must progressively construct episodic memory for downstream reasoning. Existing segmentation-based strategies therefore often produce fragmented or semantically mixed episodes, leading to degraded understanding. We present EpiStream, an online framework for utility-aware semantic episode formation in dense-action streams. Instead of detecting event boundaries, EpiStream learns when an evolving video prefix should be committed into an episode unit that maximizes downstream utility under limited memory budgets. The framework consists of three steps: (i) utility function design for semantic coherence, inter-episode distinctiveness, and downstream objectives; (ii) utility-optimal commit advantage construction from teaching signals; and (iii) peak-aware advantage learning for training a causal online commitment policy from prefix-only observations. Extensive experiments show that EpiStream substantially improves temporal semantic coherence and consistently benefits challenging downstream tasks, including real-time advice and player intention prediction.
Abstract: We consider multi-objective reinforcement learning problems where objectives come from an identical family---such as the class of reachability objectives---and may appear or disappear at runtime. Our goal is to design adaptive policies that can efficiently adjust their behaviors as the set of active objectives changes. To solve this problem, we propose a modular framework where each objective is supported by a selfish local policy, and coordination is achieved through a novel -based mechanism: policies bid for the right to execute their actions, with bids reflecting the urgency of the current state. The highest bidder selects the action, enabling a dynamic and interpretable trade-off among objectives. Going back to the original adaptation problem, when objectives change, the system adapts by simply adding or removing the corresponding policies. Moreover, as objectives arise from the same family, identical copies of a parameterized policy can be deployed, facilitating immediate adaptation at runtime. We show how the selfish local policies can be computed by turning the problem into a general-sum Markov game, where the policies compete against each other to fulfill their own objectives. To succeed, each policy must not only optimize its own objective, but also reason about the presence of other goals and learn to produce calibrated bids that reflect relative priority. Under mild assumptions, we prove the existence of Nash equilibria where dishonest bidding leads to suboptimal outcome, and the most urgent objectives win control automatically. In our implementation, the policies are trained concurrently using proximal policy optimization (PPO). We evaluate on two Atari games and a gridworld-based path-planning task with dynamic targets. Our method achieves substantially better performance than monolithic policies trained with PPO.
Abstract: We study the degree-weighted work required to compute \ell_1-regularized PageRank using the standard one-gradient-per-iteration accelerated proximal-gradient method (FISTA). For non-accelerated local methods (ISTA), the best known worst-case work scales as \widetildeO ((\alpha\rho)^-1), where \alpha is the teleportation parameter and \rho is the \ell_1-regularization parameter. A natural question is whether FISTA can improve the dependence on \alpha from 1/\alpha to 1/\sqrt\alpha while preserving the 1/\rho locality scaling. The challenge is that acceleration can break locality by transiently activating nodes that are zero at optimality, thereby increasing the cost of gradient evaluations. We analyze FISTA with respect to a less regularized objective and show that, under a checkable confinement condition, all spurious activations remain inside a boundary set \mathcalB. This yields a bound consisting of an accelerated (\rho\sqrt\alpha)^-1\log(\alpha/\varepsilon) term plus a boundary overhead \sqrtvol(\mathcalB)/(\rho\alpha^3/2). We provide graph-structural conditions that imply such confinement. Experiments on synthetic and real graphs show the resulting speedup and slowdown regimes under the degree-weighted work model.
PaperID: 2192, Poster
Authors:
Zongliang Wu, Benlei Cui, Longtao Huang, Xiaoqian Xia, Yupeng Cao, Hui Xue', Tuo Chen, Haiwen Hong, Xin YuanAbstract: Despite remarkable progress in text-to-image generation, current Diffusion Transformers (DiTs) inherently struggle with learning intricate visual content, such as accurate multilingual scene text rendering and novel visual concepts. This limitation stems from the inherent contradiction between vague semantic guidance provided by text and the objects of accurate reconstruction during diffusion training. To address this bottleneck, we propose SpaG-DiT, a novel spatial grounding training framework. SpaG-DiT directly injects explicit spatial priors into the DiTs before the attention module via Contextual Diagonal Position Encoding (CDPE), which dynamically modulates spatial coordinates without destroying the inherent sequential structure of the linguistic input. Furthermore, we introduce VLM Grounding Guidance (VLM-GG) as an automated pipeline to extract implicit spatial priors from pre-trained Vision-Language Models, eliminating human annotation costs. To evaluate the grounding ability, we establish two new datasets: SpaG-DiT-MultiLingual for non-Latin/Chinese text rendering and SpaG-DiT-Creature for new concept learning. Extensive experiments demonstrate that SpaG-DiT achieves superior training efficiency and significantly outperforms vanilla self-finetuning on the above two tasks and the widely adopted GenEval benchmark. These results validate the plug-and-play nature of SpaG-DiT and its potential to enhance DiT training processes. Code, model, and data will be released.
PaperID: 2193, Poster
Authors:
Yu Yongkang, Haobo Wang, Meng Chen, Han Fang, Xin Wei, Zhiyu Lin, Ye Yuan, Chiming Duan, Hao Sun, Haiyang ZhangAbstract: Multimodal Large Language Models (MLLMs) are typically aligned with caption-pair data, but not all answer-token supervision is equally visual. In LLaVA-style alignment training, many tokens are predictable from language context alone, yet their losses can still update vision-related parameters. Such language-dominant updates may add noisy supervision for tokens that require visual evidence. To diagnose this vulnerability, we probe the model against visual counterfactuals to reveal the intrinsic grounding requirements of each token, distilling them into two soft indicators: visual necessity and evidence specificity. Guided by this diagnostic insight, we present GEAR-Align, a controlled post-training framework that makes supervision destination explicit. During training, it converts token-level visual reliance into parameter-level gradient allocation: all answer tokens update language-side parameters, high-necessity tokens update route-flow parameters, and the high-specificity subset additionally updates visual-evidence parameters. Across two controlled Qwen3-backbone settings, GEAR-Align improves hallucination-sensitive and local-evidence-heavy benchmarks while remaining competitive on general multimodal benchmarks. These results provide evidence for token-specific visual gradients as a useful principle for controlled MLLM post-training.
PaperID: 2194, Poster
Authors: Zeqiu Yu, Xuesheng Zhang, Wenxiao Zhao, Shixiao Wang
Abstract: Long-to-Short (L2S) model merging seeks to combine the strong reasoning capability of long chain-of-thought models with the concise style of base models, but existing methods that rely on linear interpolation and first-order activation statistics fail to capture how the loss landscape responds to parameter displacement, leading to degraded accuracy on hard tasks. We propose Low-Rank Hierarchical Merging (LHM), a training-free framework that replaces heuristic statistics with lightweight second-order curvature information. LHM consists of three components: (1) a low-rank Hessian approximation that efficiently estimates per-layer Hessian traces via Hutchinson's estimator on random inputs, eliminating the need for task- specific data; (2) a Layer Second-Order Sensitivity (LSOS) index that combines Hessian trace with task-vector norm to quantify per-layer fusion difficulty; and (3) a Dual-Constraint Coefficient Mapping (DCCM) that converts LSOS scores into layer-adaptive merge coefficients while jointly satisfying a fusion-error bound and a global length-reduction target. Experiments across six mathematical benchmarks on 1.5B and 7B models show that LHM outperforms state-of-the-art methods such as ACM-TIES and ORCA, achieving over 4% accuracy gain and up to 93.1% reduction in output length relative to the long-CoT model.
Abstract: While mainstream theories of deep learning focus on training with a small learning rate, a growing body of theoretical and empirical work suggests that neural network training dynamics can be qualitatively different at large learning rates. However, a clear and general understanding of the fundamental differences between training at small and large learning rates remains lacking. In this work, we prove that if a group of parameter vectors (neurons) in a neural network model exhibits permutation symmetry, then widely used training algorithms induce a continuous mapping on the set formed by these vectors. Moreover, when the learning rate is below a certain threshold, this mapping becomes a homeomorphism. Our result reveals a critical point of the learning rate: below it, the training dynamics are guaranteed to preserve the topology of the set of neurons, whereas above it, training may introduce simplifications to the underlying neuron manifold. This provides a topological description of a novel implicit regularization effect of small learning rates and offers a potential theoretical lens for empirical phenomena such as loss of plasticity. Notably, our theory is independent of specific network architectures and loss functions, enabling topology to be applied universally to deep learning theory.
PaperID: 2196, Poster
Authors: Aritz Pérez
Abstract: Ranking models define probability distributions over rankings and are fundamental in applications such as preference learning and decision-making. Without structural assumptions, these models are intractable due to their factorial complexity. Although independence assumptions are key to reducing complexity and enabling interpretability, standard notions of conditional independence do not directly apply to rankings because they are subject to mutual exclusivity constraints. We introduce \emphrelative rank models, a probabilistic framework based on structured coarsenings of the ranking space that relax these constraints while preserving relative order information. This representation enables a well-defined notion of conditional independence for rankings and allows rankings to be modeled using Bayesian networks. As a result, relative rank models support tractable learning, probabilistic inference, and interpretable representations. We further show that two prominent ranking models based on independence assumptions tailored to rankings arise as special cases, providing a unifying framework for independence in ranking models.
Authors:
Yoav S Freund, Nicholas Harvey, Victor S. Portella, Yabing Qi, Yu-Xiang WangAbstract: We consider the problem of prediction with expert advice for ``easy'' sequences. We show that a variant of NormalHedge enjoys a second-order \epsilon-quantile regret bound of O\big(\sqrtV_T \log(V_T/\epsilon)\big) when V_T > \log N, where V_T is the cumulative second moment of instantaneous per-expert regret averaged with respect to a natural distribution determined by the algorithm. The algorithm is motivated by a continuous time limit using Stochastic Differential Equations. The discrete time analysis uses self-concordance techniques.
Abstract: Large Language Model (LLM)-based multi-agent systems are increasingly powerful, but current agentic workflow optimization paradigms make an unsatisfying trade-off. Task-level methods spend substantial offline compute yet deploy only a single workflow, leaving complementary candidates unused, while query-level methods synthesize a new workflow for each query at substantial inference cost. Our motivating analysis shows that these paradigms are more complementary than competing: workflows discovered during offline search often solve different subsets of queries, and many queries handled by expensive query-level generation can already be solved by cheaper precomputed workflows. This suggests a different objective: rather than searching for one universally best workflow or regenerating a workflow for every instance, we should build a compact bank of reusable, complementary workflows and select among them adaptively at inference time. Doing so requires solving three coupled problems: generating complementary rather than redundant candidates, compressing them into a small deployable portfolio, and assigning each query to the right workflow under a performance-cost trade-off. To this end, we present FlowBank, a three-stage framework for portfolio-based agentic workflow optimization. Diversifying proposes DiverseFlow to steer search toward under-covered queries and produce a high-coverage candidate pool. Curating proposes CuraFlow to compress this pool into a compact portfolio with minimal redundancy. Matching casts deployment as edge-value prediction on a query-workflow bipartite graph and routes each incoming query to the portfolio member with the best predicted utility. Across five benchmarks, FlowBank achieves the highest average score among the evaluated methods while remaining cost-competitive, improving over the strongest automated and handcrafted baselines by 4.26% and 14.92% relative, respectively.
PaperID: 2199, Poster
Authors: Míriam Barrabés, Daniel Mas Montserrat, Maria Perera, Alexander Ioannidis
Abstract: When benchmark examples overlap with a model’s training data, benchmark contamination can substantially inflate reported evaluation metrics, yet training provenance is often unavailable. We study post-hoc recovery of uncontaminated evaluation metrics in supervised tabular classification and regression using only prediction–label pairs from a potentially contaminated benchmark. We introduce MILAN (Metric Integrity via Leakage-Aware Normalization), a meta-trained framework that models a contaminated benchmark as a mixture of training-origin and clean evaluation examples and estimates soft posterior weights for the clean component to correct evaluation metrics. MILAN supports a broad class of metrics through weighted estimators and metric-specific correction procedures, and can optionally incorporate contamination-rate priors and sample-wise difficulty scores. To quantify uncertainty, MILAN additionally provides split-conformal prediction intervals with marginal coverage across exchangeable recovery problems. Across a large suite of simulated recovery problems with held-out datasets and model families, MILAN consistently improves clean-metric recovery over thresholding, clustering, mixture-model, anomaly-detection, and membership-inference baselines. In a polygenic risk score case study using UK Biobank evaluation data, MILAN improves agreement with independent FinnGen reference metrics. We provide a scikit-learn-style implementation at \urlhidden-for-submission.
Authors:
Weijia Liufu, Xiaoyu Guo, Ruiyi Chen, Jingzhi Liu, Kaidong Zhang, Xiwen Liang, Jianqi Lin, Dawei Sun, Yuze Wang, Rongtao Xu, Bingqian Lin, Bowen Yang, Tongtong Cao, Bowen Peng, Dongyu Zhang, Guangrun Wang, Min Wang, Liang Lin, Xiaodan LiangAbstract: Vision-Language-Action (VLA) models remain brittle in long-horizon, contact-rich manipulation because success-only imitation provides little supervision for execution drift, while failed rollouts are often discarded. We introduce RePO-VLA, a recovery-driven policy optimization framework that assigns distinct roles to success, recovery, and failure trajectories. RePO-VLA first applies Recovery-Aware Initialization (RAI), slicing recovery segments and resetting history so corrective actions depend on the current adverse state rather than the preceding failure. It then learns a Progress-Aware Semantic Value Function (PAS-VF), aligning spatiotemporal trajectory features with instructions and successful references. The resulting labels salvage useful failure prefixes via reliability decay, while low-value labels mark drift and terminal breakdowns, teaching differences among nominal, failed, and corrective actions. The data engine turns adverse states into planner-generated or human-collected corrective rollouts, teaching recovery to the success manifold. Value-Conditioned Refinement (VCR) trains the policy to prefer high-progress actions. At deployment, a fixed high value (v=1.0) biases actions toward the learned success manifold without online failure detectors or heuristic retries. We introduce FRBench, with standardized error injection and recovery-focused evaluation. Across simulated and real-world bimanual tasks, RePO-VLA improves robustness, raising adversarial success from 20% to 75% on average and up to 80% in scaled real-world trials.
Authors: Rui Wang, Renhao Xue, Ray Razi, Huan Song, Hannah Marlowe
Abstract: Time series forecasting models are increasingly scaled through large Transformer backbones, yet most existing approaches process all series through a shared dense computation path despite substantial heterogeneity in temporal structure. Mixture-of-Experts (MoE) offers a natural alternative by enabling conditional computation, but standard MoE routing leaves expert specialization weakly identified and often unstable during downstream adaptation. We propose AME-TS, a structure-guided sparse time series foundation model that aligns expert routing with interpretable temporal structure. AME-TS first uses a lightweight regime predictor to estimate series-level descriptors, including forecastability, seasonality, trend, and sparsity, and maps them to a soft structural prior over experts. This series-level prior guides token-level routing during training, encouraging structure-aligned specialization. On the GIFT-Eval benchmark, AME-TS delivers a strong accuracy-efficiency tradeoff across model scales: it substantially outperforms existing time series foundation models at small model scales and remains competitive with the strongest models at larger scales, while activating substantially fewer parameters through sparse routing. We further show that AME-TS learns more interpretable routing geometry and substantially more stable expert specialization than standard MoE during fine-tuning on the M5 dataset. These results suggest that structure-aware routing is an effective and reliable way to realize the benefits of sparse expert models for time series forecasting.
Abstract: Vision-Language-Action (VLA) models have driven significant progress in robotic manipulation, yet they fundamentally struggle with the ``vision-override'' phenomenon. Driven by the severe modality imbalance between dense visual streams and sparse linguistic instructions, VLAs frequently fall prey to causal confusion. Instead of treating language as the primary causal driver, the policy entirely bypasses the original instruction by overfitting to spurious visual confounders, such as prominent objects or familiar layouts. To systematically alleviate this bias, we formalize the process of action generation as a Dual-path Deconfounding Graph (DDG) and propose CofactVLA, a novel causal intervention framework. By dynamically constructing a language-masked counterfactual branch within a single forward pass, CofactVLA isolates and neutralizes visual confounders through two synergistic mechanisms. First, Action-Level Orthogonal Projection Guidance (OPG) geometrically projects the factual velocity field away from the counterfactual visual bias during continuous flow matching, extracting the pure semantic intent. Second, Feature-Level Counterfactual Covariance Reduction (CCR) mathematically deconfounds latent representations by penalizing the positive eigenspace of the covariance difference, explicitly suppressing dominant visual shortcuts while preserving the causal language intent. Extensive experiments demonstrate that CofactVLA establishes a new state-of-the-art across diverse simulation benchmarks. Beyond simulation, real-world robot experiments demonstrate the causal efficacy of our method in bridging the generalization gap, yielding a 52.3% absolute success rate gain under out-of-distribution scenarios.
Abstract: State-space models (SSMs) face a fundamental trade-off between efficiency and expressivity that is mainly dictated by the structure of the model's transition matrix. Unstructured transition matrices enable maximal expressivity, as measured by their ability to model finite-state automaton (FSA) transitions, but come at a prohibitively high compute and memory cost. In contrast, diagonal and other structured transition matrices suffer from limited expressivity, but are highly efficient both in runtime and memory consumption. Building on recent work on structured sparse SSMs, we propose Flash PD-SSM, a novel SSM that achieves comparable throughput to diagonal models with optimal expressivity guarantees. Flash PD-SSM keeps a collection of structured sparse matrices, from which a single matrix is selected at each time-step, enabling FSA expressiveness at the level of unstructured matrices while maintaining the efficiency required for training models at scale. First, we validate Flash PD-SSM against a suite of alternative models on common mechanistic and synthetic state-tracking tasks, showing that its theoretical expressivity is achieved in practice. Moreover, on multivariate time-series tasks involving sequences of length over 17,000, Flash PD-SSM defines a new state-of-the-art (SoTA) accuracy among competing SSM methods. Second, we demonstrate that Flash PD-SSM is an effective drop-in replacement for hybrid LLMs, yielding improvements both in natural language state-tracking and in common language modeling scenarios. Finally, we show that our highly efficient design results in increased throughput and decreased memory consumption with respect to SoTA SSMs widely used in frontier language models.
Abstract: We study learning-augmented streaming algorithms for estimating the value of Boolean Maximum Constraint Satisfaction Problems (Max-CSPs). Specifically, we consider streaming algorithms equipped with an \varepsilon-accurate oracle: for each variable, the oracle outputs a label in \\-1,1\\ that agrees with a fixed optimal assignment with probability 1/2+\varepsilon, and disagrees otherwise. Prior work by Dong, Peng, and Vakilian [2025] showed that, for the special case of Max-Cut, such an oracle enables a (1/2+\Omega(\varepsilon^2))-approximation using \operatornamepoly(1/\varepsilon) words of space in insertion-only streams, and \operatornamepoly(1/\varepsilon,\log n) words of space in dynamic streams. We substantially generalize and strengthen their result, provided that a highly accurate estimate of \varepsilon is available. For any Boolean Max-2CSP and any \eta>0 independent of \varepsilon, we obtain, with high constant probability, a single-pass (1-\eta)-approximation using \operatornamepoly(1/\varepsilon,1/\eta) words of space in insertion-only streams, and \operatornamepoly(1/\varepsilon,1/\eta,\log n) words of space in dynamic streams, where n is the number of variables. We further extend our techniques to general Boolean Max-kCSPs for k\ge 3, with slightly worse space complexity.
PaperID: 2205, Poster
Authors: Mohammad M Omati, Arash Amini
Abstract: In permutation synchronization, the goal is to find globally cycle‐consistent correspondences from noisy pairwise matchings. In this work, unlike spectral relaxations that embed permutations into an orthogonal space and often result in inaccuracies, we maintain the problem in its original combinatorial form. By shifting the affinity spectrum to ensure positive semidefiniteness, we cast the trace‐maximization over partial permutations as a convex‐in‐P formulation. Our minorization-maximization scheme then replaces this with a sequence of exact linear‐assignment subproblems, the row-/column-sum constraints of which are totally unimodular, guaranteeing integral solutions with no rounding. This direct, combinatorial approach delivers a monotonic objective ascent, convergence to a KKT point, and achieves superior accuracy, cycle consistency, and runtime on image-matching benchmarks.
PaperID: 2206, Poster
Abstract: Machine learning models often achieve high average accuracy in the training data while performing poorly on test data due to subpopulation shift. A common data-centric solution is to enforce balanced subgroup proportions, but full balance can be costly, infeasible, and not always necessary. In this work, we study whether balanced training data is the unique optimal configuration for robust subpopulation generalization. We analyze the theoretical performance across training subgroup distributions and show that multiple configurations, including substantially imbalanced ones, can achieve performance comparable to the balanced configuration. We further derive gradient-based directions for adjusting subgroup proportions to improve robustness more effectively. These directions can diverge from the direct path toward balance, suggesting that balancing is not always the most efficient data-allocation strategy. We validate the theoretical insights with controlled experiments on image and text benchmarks, showing that gradient-guided allocation can improve robustness more efficiently than directly enforcing balance. These findings suggest that strategically allocating data offers a more flexible and principled path to robust performance than simple balancing.
PaperID: 2207, Poster
Abstract: Reward models are a core component of modern text-to-image systems, serving as the objective for inference-time selection and RL-based post-training. A key prerequisite is visual understanding: to judge prompt adherence and ground aesthetic preferences, a reward model must bind attributes, resolve relations, and count reliably. We find that widely-used reward models perform poorly on targeted visual reasoning evaluations. Even when starting from strong vision–language backbones, preference fine-tuning erodes these capabilities, and continual-learning baselines fail to mitigate a steep tradeoff between visual reasoning and preference accuracy. We propose a data-centric intervention that injects visual reasoning supervision into preference learning by flipping the pairing axis. To complement conventional human preference data (paired images under a shared prompt), we convert existing visual reasoning datasets into preference pairs that share the same image but differ in prompt, and train on the mixture. This yields a Pareto improvement in both visual understanding and preference prediction. Across inference-time best-of-n selection and RL post-training, the resulting reward model improves text-to-image generation on challenging compositional prompts (GenEval) by an average of 4.0% and 5.5% respectively, outperforming HPSv3 which was trained on 6x more preference pairs.
Authors: Daniel Saragih
Abstract: Anatomical mesh segmentation requires models that operate directly on irregular surface geometry while remaining robust to arbitrary patient pose and mesh resolution variation. Existing task-specific mesh and point-cloud methods are not equivariant, and can degrade sharply under test-time perturbation, for example dropping by 25-26 IoU points on intraoral scan segmentation at 40\textdegree tilt. We present EAMS, an Equivariant Anatomical Mesh Segmentor built on Equivariant Mesh Neural Networks (EMNN), and evaluate it across four clinically distinct tasks spanning edge-, vertex-, and face-level supervision. We combine intrinsic mesh descriptors with anatomy-aware priors, including PCA-derived frames for dental arches and liver surfaces, and augment message passing to provide lightweight global context. Across intracranial aneurysm and intraoral segmentation, EAMS variants are competitive with specialized baselines on unperturbed inputs while remaining stable under geometric perturbations, and on liver surfaces they expose a favorable trade-off between canonical-pose accuracy and rotation robustness. These results show that a lightweight (<2M parameters) equivariant framework can deliver robust anatomical mesh segmentation across diverse supervision types without task-specific architectures.
Abstract: Large Reasoning Models (LRMs) have shown excellent performance in reasoning tasks using long Chain of Though. However, this has also led to a significant increase of computational costs and the generation of verbose output, a phenomenon known as overthinking. The tendency to overthinking is often exacerbated by Reinforcement Learning (RL) algorithms such as GRPO/DAPO. In this paper, we propose BFS-PO, an RL algorithm which alleviates this problem using a Best-First Search exploration strategy. Specifically, the reasoning chains generated at training time by BFS-PO are organized into a search tree, within which the shortest correct sequence is selected as the best and expanded using a backtracking mechanism based on maximum entropy nodes. In this way, we bias the exploration of the solution space towards the search for increasingly shorter solutions, training the LRM to generate progressively more concise answers. Using different benchmarks and base LRMs, we show that BFS-PO can simultaneously increase the LRM accuracy and shorten its reasoning chains. Our code and models are available in the supplementary material and will be published after this article is accepted.
PaperID: 2210, Poster
Abstract: Image editing aims to modify specific attributes of an image while preserving all other aspects. In practice, however, applying even simple editing instructions to existing methods often leads to unintended changes. We identify the central issue: a fundamental causal principle of \emphminimal change, which requires that an intervention alter only the intended attributes in the output while leaving all others invariant, has not been explicitly incorporated as an optimization objective. To encourage minimal change, prior work rooted in causal representation learning typically imposes an L_1 regularizer on the latent difference between pre- and post-edit representations. However, this strategy does not scale to modern image editing models: L_1 regularization fails to induce true sparsity and instead tends to shrink all latent differences uniformly; sparsity in the latent space does not translate to localized change in the output of nonlinear deep image generation models; and the multi-domain or counterfactual supervision it requires is generally unavailable. To effectively leverage the minimal change principle for controllable image editing, we cast it as a reinforcement learning objective. Rather than constraining latent representations, we design rewards that directly capture minimal change over intervention outcomes and optimize them explicitly during training. To produce these rewards reliably, we further introduce an agentic reward model that leverages the chain-of-thought reasoning ability of vision-language models without requiring human annotations. The idea is that instead of emitting a single, unreliable scalar score, the model reasons explicitly about two complementary failure modes, unimplemented changes and unintended changes, and produces this supervision automatically. Experimental results demonstrate that our method produces significantly more localized and semantically consistent edits compared to existing approaches, reducing unintended changes.
PaperID: 2211, Poster
Abstract: Memory consumption of the key-value (KV) cache remains a central bottleneck for long-context LLM inference. We introduce PRKV, a training-free KV eviction method that estimates multi-hop global token importance by applying Personalized PageRank over attention-induced transition graphs. This enables importance propagation beyond recency or instantaneous attention, without modifying the model. We benchmark PRKV on long-context understanding, sparse retrieval, and chain-of-thought reasoning tasks using multiple LLM families. Under moderate to aggressive compression ratios, including up to 90%, PRKV achieves competitive task performance against established eviction baselines while reducing peak memory usage with low end-to-end overhead. These results demonstrate that multi-hop global influence estimation yields a favorable accuracy--memory trade-off for scalable long-context inference.
PaperID: 2212, Poster
Abstract: Post-training pruning compresses pretrained large language models (LLMs) by removing redundant weights without full retraining. However, existing pruning criteria primarily rely on parameter statistics or generic reconstruction errors, fundamentally ignoring the underlying geometric structure of how weight removals perturb the layer output. In this paper, we rethink LLM pruning from and reveal that representative methods (e.g., SparseGPT and Wanda) implicitly assume a standard Euclidean output metric. This geometry-agnostic assumption yields an isotropic manifold that assigns uniform costs to all perturbation directions, thereby failing to capture target-dependent sensitivity. To overcome this structural limitation, we propose (TOM-Pruning). TOM-Pruning equips the output perturbation space with a target-aware local Riemannian metric, assigning direction-dependent geometric costs based on the perturbation's alignment with the target response. By mathematically pulling this metric back to the parameter space, we derive a pruning sensitivity score that inherently captures a cosine-based input-output channel affinity. To resolve the poor discriminability of this raw affinity in high-dimensional spaces, we further introduce a Competitive-Hubness Channel Affinity Mechanism to reconstruct a highly distinguishable final pruning criterion. Extensive experiments across multiple LLM families demonstrate that our manifold-driven approach consistently outperforms existing no-weight-update pruning baselines under various sparsity settings, notably reducing WikiText-2 perplexity by 26.8% on LLaMA-3.2-1B and achieving an 11.0% relative zero-shot accuracy gain on LLaMA-3-8B at 70% unstructured sparsity.
PaperID: 2213, Poster
Abstract: As diffusion-based offline reinforcement learning (RL) move closer to real deployment, it becomes critical to remove the influence of specific training data for privacy, safety, and regulatory compliance. Since retained data may overlap with or generalize from the forget data, strict retraining equivalence can be ill-posed in offline RL. Existing unlearning methods are ineffective for diffusion policies, as training influence is dispersed across the denoising process and reinforced by critic values. We introduce Relative Fisher Forgetting (RFF), the first principled framework for selective unlearning in diffusion-based offline RL. RFF combines two asymmetric components: critic-side value suppression removes residual value incentives associated with the forget set, eliminating Q-guidance pathways that would otherwise sustain forgotten behaviors; actor-side relative-Fisher updates attenuate forget-set-dominant parameter influence in the denoising policy, reducing behavioral regrowth. To stabilize training, RFF alternates actor-critic updates and employs gradient clipping and retain-set regularization. Experiments on MuJoCo benchmarks show that RFF achieves the lowest identifiable forget-set reliance among baselines while preserving retained performance, and remains over 12× more efficient than retraining. When undesired behaviors are primarily supported by the forget set, RFF additionally suppresses them without collateral degradation.
Abstract: Frontier reasoning models are produced by post-training base language models with reinforcement learning. Recent work has challenged this by showing that sampling from a sharpened version of the base model's distribution, a so-called elicits comparable reasoning without additional training, curated datasets, or verifiers. However, making this method practical requires from the power distribution. A sampler needs to "mix" to the power distribution which necessitates moving between modes of the target distribution; intuitively, e.g., trying different reasoning strategies. The samplers proposed in prior works repeatedly select a "cut" position in the current reasoning trace uniformly at random and resamples the suffix from that position onwards. However, reasoning traces typically contain a few consequential decisions (e.g., the choice of proof strategy or algorithm), and we observe that a uniformly chosen cut tends to rewrites local details rather than revisiting decision points. We introduce an algorithm (Entropy-Cut Metropolis--Hastings) which uses the base model's next-token entropy as a proxy to identify key decision points and re-sample from these positions. We empirically verify that entropy jumps are a useful proxy for decision points, and in a stylized model of reasoning prove that our method's mixing time scales with the number of decisions in a trace rather than with the number of tokens, which can be much larger. Across MATH500, HumanEval, GPQA, and AIME on several models, our method consistently improves over baselines and RL-trained models.
Abstract: Pretrained diffusion models increasingly serve as frozen teachers feeding downstream pipelines such as text-to-3D, single-step distillation, and data attribution. The teacher gradients these pipelines consume are Monte Carlo expectations over noise levels and Gaussian noise; their estimator variance dominates compute cost because each draw requires expensive upstream work, such as rendering, simulation, or encoding. We introduce CARV, a compute-aware variance-accounting framework that motivates a hierarchical Monte Carlo estimator: amortize the expensive upstream computation over cheap diffusion-noise resamples, sharpened by timestep importance sampling and a stratified inverse-CDF construction. Across diffusion-guided workloads, we obtain 2–3× effective compute multipliers, most from amortized reuse and approximately 25% additional gain from importance sampling plus stratification, without changing the objective. We also map regimes where these gains translate into improved downstream metrics versus regimes where they do not, such as DMD.
Abstract: , where a limited number of both normal and anomalous examples are available as references during inference. Existing FSAD methods rely on normal-only references through normality matching, ignoring the discriminative clues in anomalous references, while directly fitting both references can overfit to the seen anomalies. We introduce , an intrinsic deviation learning framework that leverages both reference types to learn intrinsic deviation patterns characterizing generalizable abnormality as deviations from normality. IDEAL decomposes the learning process into two novel components: 1) a Normal Variation Eraser to suppress nuisance normal variations that may lead to noisy deviations from normality, thereby highlighting anomaly-relevant deviation representations; 2) an Intrinsic Deviation Encoder to decompose these denoised deviation representations into intrinsic deviation vectors capturing the most discriminative orthogonal deviation directions. At inference, IDEAL scores query-to-normal deviations preserved after projection onto the learned intrinsic deviation vectors, enabling generalization for both seen and unseen anomalies. Extensive experiments on eight real-world datasets show that IDEAL generalizes effectively to unseen anomalies and consistently outperforms existing state-of-the-art FSAD methods.
Authors:
Shan Xiaojun, Haoyu Shen, Yucheng Mao, Haiyang Xu, Xiang Zhang, Abhay Anand, Bingnan Li, Zhuowen TuAbstract: We present CyCLeGen, an autoregressive framework that integrates layout understanding and layout-to-image generation through cycle consistency: predicted layouts must produce faithful images, and generated images must yield consistent layouts. We enforce this constraint via CycleGRPO, a bidirectional reinforcement learning strategy with complementary geometric and perceptual rewards, enabling self-introspective learning from only 8k RL samples. This creates a natural loop, generation helps understanding by rewarding only those layouts that lead to high-quality images, and understanding helps generation by rewarding only those images whose spatial structure can be faithfully recovered. Extensive experiments show that CyCLeGen achieves significant gains across diverse image understanding and generation benchmarks, with emergent gains on image captioning.
PaperID: 2218, Poster
Abstract: Permissive, auditable pretraining corpora enable lawful, responsible and transparent open foundation model research and development, facilitating standardization and common progress. Recent work such as MixtureVitae has shown that permissive corpora can also yield strongly competitive base models, but no controlled comparison has measured whether they still remain competitive after post-training. We post-train 1.7B base models pretrained on 300B of MixtureVitae permissive tokens and compare them against strong compute and token matched non-permissive reference baselines (Nemotron-CC-HQ, FineWeb-Edu and DCLM) and against open weights baselines trained at substantially larger pretraining compute (SmolLM2 1.7B, Qwen2.5 1.5B, Qwen3 1.7B). All models are post-trained with the same pipeline: Tulu3 supervised fine-tuning and direct preference optimization, followed by OpenThoughts3 reasoning training. Post-trained MixtureVitae models match or outperform the matched non-permissive reference on reasoning and instruction-following benchmarks, and maintain solid general language understanding performance. They also remain competitive with strong open weights baselines, which use over an order of magnitude more pretraining compute. Ablations confirm that the substantial reasoning and instruction subset of MixtureVitae drives this result: removing it drastically weakens the post-trained model under the same recipe. Together, these findings show that permissive pretraining does not preclude strong post-training, strengthening the case for legally safe and reproducible open foundation model research and development at frontier performance levels. Data, models, and code to reproduce the experiments will be open-sourced.
Authors: Philippe Goulet Coulombe
Abstract: I show that ordinary least squares (OLS) predictions can be rewritten as a restricted attention module, akin to those at the heart of transformer architectures. This connection emerges by reframing OLS as a similarity-based prediction rule operating in a learned embedding space. From this perspective, least squares does not estimate coefficients per se, but instead selects an embedding that minimizes squared prediction error by matching training and test vectors through inner products — mapping naturally onto the query-key-value structure of attention. I then extend the framework to dimensionality reduction, nonlinearity, and connections to time series econometrics. Monte Carlo simulations and real-data experiments on UCI/OpenML benchmarks show that a direct implementation of nonlinear Attention Regression performs competitively against standard machine learning baselines. In the other direction, within a transformer architecture for tabular data, an attention block can be replaced by an explicit regression on polynomial-expanded features, matching or exceeding the standard transformer's predictive accuracy at a fraction of its parameter count.
PaperID: 2220, Poster
Abstract: We study which reduced opponent states preserve the interim decision problem of a realized type in incomplete-information continuous Colonel Blotto. For any continuous observable summary of the opponent posterior, we prove an exact ambiguity identity: the worst-case payoff gap over posterior laws with the same state equals twice the sup-norm distance of the payoff slice from the closed observable-additive span. Thus a state is payoff-exact if and only if every payoff slice lies in that span. Specializing to exact-budget Blotto, we characterize first-order battlefield marginals: they are exact precisely for coordinate-additive opponent slices. The two-battlefield case is degenerate, whereas from three battlefields onward first-order marginals can discard payoff-relevant dependence; more generally, exact-budget allocation games exhibit a strict r-order marginal hierarchy. The same characterization separates estimation from representation: exact states admit vanishing statistical error, while non-exact states impose a sample-independent ambiguity floor for state-restricted learners. Finally, exact and efficiently learnable first-order states do not eliminate equilibrium complexity: approximate Bayesian Nash equilibrium remains PPAD-hard in a strict diagonal regularized Blotto subclass.
Abstract: Active learning has provable benefits in learning machine learning models. As such, automatically building an instance-based curriculum to speed up neural network training has been an active area of research. Recent works such as those of Jiang et al. [2019]; Zhou et al. [2020]; Wang et al. [2024b]; Yuan et al. [2025] have shown benefits of using dynamic curricula on standard image classification tasks, where training data is selected adaptively online, based on how the training goes. On the other hand, Wu et al. [2021] demonstrated that static data ordering, designed independently of the actual training run, does not help in learning classifiers for some standard image classification benchmarks (such as CIFAR10 and CIFAR100). In this paper we extend this result to the case of dynamic data selection, still focusing on image classification, showing that, at least on standard benchmarks like CIFAR10, CIFAR100, and ImageNet, the performance of some state-of-the-art instance-based dynamic curriculum selection methods can be traced back to changes in the effective learning-rate schedule due to the subsampling of the data, rather than the careful selection of the exact data points. This highlights the importance of properly controlling all variables in designing experiments and perhaps suggests that standard image classification tasks are not the best benchmarks for studying curriculum learning.
Authors: Oriol Vendrell Gallart, Nima Negarandeh, Ramin Bostanabad
Abstract: Neural operators provide fast surrogates for PDEs but their deterministic predictions limit their use in tasks requiring uncertainty quantification (UQ), especially under geometric variability. Existing approaches primarily model uncertainty in network parameters, largely overlooking the geometry-aware representations learned by the operator itself. We propose REEF-GP (Residual on Embedded Features Gaussian Process), a post-hoc UQ framework that fits a GP to the residuals of a frozen neural operator whose internal embeddings define the kernel feature space. Rather than learning a separate feature map, REEF-GP adapts the operator’s intrinsic coordinate-feature representations to construct geometry-aware uncertainties. To ensure stability and scalability on unstructured domains, REEF-GP incorporates spectral-normalized projections, heteroscedastic geometry-aware noise, and efficient subset-based training that avoids restrictive low-rank approximations. Across five PDE benchmarks with varying geometries, REEF-GP preserves predictive accuracy while achieving calibrated uncertainty estimates competitive with deep ensembles but at a fraction of their cost. Our approach remains robust under geometric distribution shift, with uncertainty concentrating in physically meaningful regions (e.g., shock fronts). Our results demonstrate that accurate and scalable post-hoc UQ for neural operators can be achieved directly in their learned feature space, offering a practical alternative to parameter-centric approaches.
PaperID: 2223, Poster
Abstract: Parameter-efficient fine-tuning (PEFT) methods such as low-rank adaptation have become essential for adapting large language models to downstream tasks. Existing approaches often apply low-rank updates uniformly across all layers, ignoring the structural coupling inherent in transformer architectures. We observe that transformers contain \emphweight pairs whose interactions govern model behavior including query-key products, value-output products, and feedforward networks. We propose PRincipal Interaction Subspace Matching (PRISM), a principled framework that exploits this coupling through spectral parameterization. For each weight pair, we parameterize adaptations using the spectral basis of the paired matrix, ensuring that updates occur in the subspace where the paired matrix has maximal effect. We validate our approach through controlled experiments on synthetic models, demonstrating faster convergence. Empirical evaluations on arithmetic reasoning, code generation, and vision tasks show that PRISM outperforms existing PEFT methods.
Abstract: Long-context autoregressive reasoning in large language models is severely bottlenecked by key-value (KV) cache memory, making decoding-time compression essential. Existing methods typically rely on global or per-head token-wise selection. Under tight memory budgets and repeated compression events, however, these policies tend to over-focus on local attention spikes, causing : the severe depletion of contiguous spans of intermediate reasoning from the cache. This fragmentation disrupts reasoning, leading to Problem Drifting and Repetition Collapse. To address this, we propose Adaptive Mass-Segmented (AMS), a scorer-agnostic KV compression framework based on an "allocate-then-score" paradigm. By decoupling budget allocation from token scoring, AMS acts as a that shifts attention from a strictly micro-level retention metric to a macro-level spatial partitioning tool. Specifically, AMS derives an attention-based quality-mass distribution to partition each head's cache into adaptive segments, enforcing region-wise quotas with minimum-keep guarantees. Furthermore, an exponential moving average (EMA) credit mechanism makes this allocation history-aware, smoothing cache evolution across repeated recompressions. Experiments on MATH500, AIME24, AIME25, and GSM8K using 7B and 32B backbones show AMS improves pass@1 accuracy by up to 20.0 points over corresponding base scorers. AMS demonstrates generalization across diverse long-context workloads, including code completion, open-domain QA, and sparse retrieval. AMS seamlessly integrates with token-level scorers, consistently improves their performance, sometimes surpasses uncompressed full-KV decoding, and incurs negligible practical decoding-time overhead.
Abstract: We study the problem of auditing the fairness of a given classifier under partial feedback, where true labels are available only for positively classified individuals (e.g., loan repayment outcomes are observed only for approved applicants). We introduce a novel cost model for acquiring additional labeled data, designed to more accurately reflect real-world costs such as credit assessment, loan processing, and potential defaults. Our goal is to find optimal fairness audit algorithms that are more cost-effective than random exploration and natural baselines. In our work, we consider two audit settings: a black-box model with no assumptions on the data distribution, and a mixture model, where features and true labels follow a mixture of exponential family distributions. In the black-box setting, we propose a near-optimal auditing algorithm and show that a natural baseline can be strictly suboptimal. In the mixture model case, we design a novel algorithm that achieves significantly lower audit cost than the black-box case under mild assumptions. Our approach develops novel connections between partial feedback and learning from truncated samples. We also extend known results on maximum-a-posteriori oracles for spherical Gaussian mixtures to handle exponential family mixtures, which may be of independent interest. Moreover, our algorithms apply to popular fairness metrics including demographic parity, equal opportunity, and equalized odds. Empirically, we demonstrate strong performance of our algorithms on real-world fair classification datasets like Adult Income and Law School, consistently outperforming natural baselines by around 50% in terms of audit cost.
Authors:
Felipe Meneguzzi, Alexandre Buchweitz, Augusto B. Corrêa, Victor S Putrich, André G. PereiraAbstract: HTN planning is a variation of classical planning where, instead of searching for a linear sequence of actions, an algorithm decomposes higher-level tasks using a method library until only executable actions remain. On one hand, this allows one to introduce domain knowledge that can speed up the search for a solution through the method library. On the other hand, it creates challenges that go beyond those of classical state-space search. While recent research produced a number of heuristics and novel algorithms that speed up HTN planning, these heuristics are not yet as informative as those available in classical planning algorithms. We investigate whether large language models (LLMs) can generate effective search heuristics for HTN planning, extending the methodology of [Corrêa et al 2025] from classical to hierarchical planning. Using the Pytrich planner on six standard total-order HTN benchmark domains, we evaluate heuristics generated by nine LLMs under domain-specific prompting and compare them against the TDG and LMCount domain-independent baselines and the PANDA planner. Our results show that LLM-generated heuristics nearly match the coverage of the best available HTN planner, while substantially reducing search effort on 83% of the shared problems .
PaperID: 2227, Poster
Authors: Yongkang Ding, Zi Ye, Tiantian Gong, Liyan Zhang
Abstract: Cloth-changing person re-identification aims to retrieve images of the same person across cameras when clothing may vary over time. Unlike conventional Re-ID, appearance cues such as clothing color and texture become highly unreliable in cloth-changing scenarios. Existing methods either rely on auxiliary clothing-invariant biometric cues that may introduce estimation noise, or model attribute semantics as isolated labels, making it difficult to simultaneously achieve robustness to attribute noise and fine-grained visual--semantic interaction. To address these issues, we propose a Graph-Enhanced Attribute-Aware Modeling framework, termed GEAM. Specifically, we first preprocess the attribute representation to explicitly suppress clothing-related semantics. We then introduce an adaptive attribute graph to model the latent dependencies among identity-related attributes, thereby alleviating the interference caused by attribute prediction errors. Based on the refined attribute representation, we further map the enhanced attribute semantics into compact semantic tokens and inject them into the visual backbone in a hierarchical manner, enabling local visual features to adaptively absorb identity-relevant yet clothing-irrelevant discriminative cues. This design effectively improves robustness to cloth-changing scenarios while preserving the stability of pretrained visual representations as much as possible. Extensive experiments on multiple mainstream CC-ReID benchmarks demonstrate that GEAM consistently outperforms a variety of state-of-the-art methods under the challenging clothes-changing setting. The source code and pretrained weights will be publicly released on GitHub.
PaperID: 2228, Poster
Authors: Ricardo Luna Gutierrez, Vineet Gundecha, Rahman Ejaz, Varchas Gopalaswamy, Riccardo Betti, Sahand Ghorbanpour, Aarne Lees, Soumyendu Sarkar
Abstract: Optimizing high-dimensional physical systems under expensive evaluations and complex constraints is a central challenge in science and engineering. In offline model-based optimization (MBO), methods must improve designs using only a fixed dataset and no test-time access to the true objective. However, existing methods typically suffer from conservatism, unconstrained surrogate exploitation, or function as static inference-time samplers that cannot refine candidates against precise physical objectives. In this paper, we propose Differentiable Diffusion Conditioning (D2C), and show that deterministic sampling makes the reverse denoising chain differentiable with respect to the conditioning signal. D2C turns a frozen conditional diffusion model into an end-to-end gradient-based optimizer by directly optimizing conditioning variables at test time via gradients from frozen surrogate objectives and physics-informed penalties backpropagated through the reverse chain. We implement this idea in a composed proposer-evaluator framework and evaluate it on laser pulse optimization for inertial confinement fusion, sustainable data-center workload scheduling, and four Design-Bench tasks. Our method achieves the strongest results on the two real-world tasks and the best average rank across four Design-Bench tasks.
PaperID: 2229, Poster
Abstract: Image tokenization aims to compress visual signals into latent codes that remain effective for reconstruction and generation. Driven by the multi-scale nature of images, organizing latent along a coarse-to-fine refinement process emerges as a promising direction for visual tokenization. Existing scale-wise tokenizers often introduce this structure through token ordering or latent prefixes, but each target scale is typically decoded directly from its prefix, making different prefixes behave like separate reconstruction codes. This can lead to redundant encoding of global structure across scales and limits the effectiveness of compact token budgets. We propose RefineTok, a scale-wise continuous tokenizer that makes decoding itself progressive. RefineTok carries a visual state across scales and uses each latent prefix to refine the previous state into the next scale, turning scale-wise latents into incremental refinement signals rather than independent reconstruction codes. This design yields a compact and structured latent interface for both reconstruction and downstream generation. Together with prefix-level training, a coarse-centric scale trajectory, and a semantic anchor, RefineTok achieves strong reconstruction fidelity among scale-wise tokenizers and competitive class-conditional generation on ImageNet-1K under compact latent budgets. Meanwhile, it preserves a progressive scale-wise interface that single-scale tokenizers do not provide.
PaperID: 2230, Poster
Abstract: With the recent advancement of large language models (LLMs) scaling, post-training weight quantization (PTQ) has become one of the standard tools for making models deployable on commodity hardware with minimal performance degradation. However, existing quantization solvers such as GPTQ calibrate each weight matrix against a corrupted signal: the calibration input observed at any deep layer is the output of the cumulative quantized chain, not a clean full-precision reference. We propose CMPQ (Compensatory Quantization), which reframes the corruption as evidence about \emphwhich input directions have become unreliable and tells the solver to stop spending its compensation budget on them. CMPQ measures the divergence between the quantized and full-precision input and uses it to reshape the curvature that GPTQ already relies on, redirecting compensation onto weights attached to clean input subspaces and cutting the propagation path of upstream noise. On the Qwen3 family at three sizes (4B / 8B / 14B), CMPQ matches GPTQ-family baselines at 3-bit and substantially outperforms them at 2-bit, where standard solvers collapse to near-random accuracy.
PaperID: 2231, Poster
Abstract: Most graph neural networks propagate information through a fixed-coefficient polynomial filter applied uniformly across every node and every layer. However, the appropriate filter varies along two distinct axes that such a parameterization cannot accommodate. Within a single graph, different nodes benefit from different mixtures of low-pass smoothing, high-pass contrast, and pass-through behavior. Across graphs, signal propagation dynamics themselves differ, with some graphs benefiting from long-range propagation across many hops while others require damping to prevent over-smoothing. The appropriate filter is therefore not a single fixed object at all, but a state-dependent composition whose behavior changes according to a node's current representation and its propagation history. We propose Castor, a self-tuning graph filter that realizes this composition through state-dependent decisions made at every step of propagation. A router reads each node's current state together with its initial state, and emits per-node mixing weights over three primitives, namely low-pass smoothing, high-pass contrast, and identity. A learnable per-hop coefficient then determines how much of the most recent change in the propagation is carried forward, accelerating the iteration when positive and damping it when negative. Setting the coefficient to zero recovers standard first-order propagation. Across fourteen standard node-classification benchmarks spanning the full homophily spectrum, a single Castor instance ranks top-three on every dataset and achieves an average rank of 2.07, more than 3.5 ranks ahead of the next-best of fourteen baselines.
PaperID: 2232, Poster
Authors: Xiang Long, QIANQIAN WANG
Abstract: Attribute-Missing Graph Clustering (AMGC) is a critical yet challenging task in real-world applications, in which only a subset of nodes hold complete attributes information while others are partially missing. Though existing feature propagation-based methods can effectively recover missing node attributes, they commonly assume uniform contributions across all nodes, ignoring unreliable nodes that may degrade imputation quality. Furthermore, adopting low-pass graph filters often suffers from the over-smoothing problem and inevitably sacrifices discriminative information. %Besides, traditional contrastive learning pulls all the node embeddings evenly, which might conflict with the rule that intra-cluster nodes should be closer to each other. To address these limitations, we propose \underlineAdaptive \underlineFeature \underlinePropagation (AFP) for attribute-missing graph clustering. Specifically, we first design a reliability-aware feature propagation mechanism that adaptively weights edges based on node importance. Then, we introduce a multi-scale embedding module to capture both local and global structural information. Finally, we develop a topology-aware contrastive loss to enhance clustering consistency. Extensive experiments on benchmark datasets demonstrate the superiority of our method.
PaperID: 2233, Poster
Authors: Karim Knaebel, Gonzalo M Garcia, Christian Schmidt, Ilya Fradlin, Lucas Nunes, Daan de Geus, Bastian Leibe
Abstract: With impressive progress in architectures, training recipes, and quantity of high-quality data, recent monocular geometry estimation methods estimate global 3D geometry remarkably well. However, they still exhibit surprisingly low quality on surface details, which is clearly visible in qualitative predictions but not reflected in common metrics. To improve this shortcoming, we first formulate a metric that is capable of measuring the surface quality of the estimated geometry via normals derived from the predicted point map. Furthermore, we propose two manners to improve fine-grained local geometry in monocular geometry estimation models: (i) we present an improved point-based gradient loss, and (ii) we introduce a novel decoder based on neighborhood attention. We evaluate our proposed method on eight common zero-shot monocular geometry estimation benchmarks and achieve state-of-the-art results, particularly improving local surface geometry. Extensive ablations validate our design choices. Code and models will be open-sourced.
Abstract: We study for the first time, stochastic dueling bandits over continuous action spaces with Lipschitz structure, where feedback is purely comparative. While dueling bandits and Lipschitz bandits have been studied separately, their combination has remained unexplored. We propose the first algorithm for Lipschitz dueling bandits, using round-based exploration and recursive region elimination guided by an adaptive reference arm. We develop new analytical tools for relative feedback and prove a regret bound of \tilde O\!\left(T^\fracd_z+1d_z+2\right), where d_z is the zooming dimension of the near-optimal region. Further, our algorithm takes only logarithmic space in terms of the total time horizon, best achievable by any bandit algorithm over continuous action space.
PaperID: 2235, Poster
Abstract: Test-time scaling systems often choose how much to sample, which prompt to use, and which verifier or aggregation rule to trust after repeatedly inspecting a calibration set. Standard validation error bars are not valid after this adaptive selection. We model a test-time scaling recipe as a randomized predictor with a random cost and prove a time-uniform PAC-Bayes certificate that holds simultaneously over all calibration times, all posterior mixtures over recipes, and all data-dependent stopping rules. Thus the same event covers the recipe search that precedes deployment, not merely the fixed recipe finally reported. For a finite grid of budgets and verifiers, the penalty is the familiar \log M/n, while non-uniform priors and localized posterior mixtures give sharper certificates for structured searches. We specialize the result to self-consistency and show a complementary margin law: plurality voting improves exponentially in the sample budget only on examples where the correct answer is the generator's unique mode; otherwise more compute can provably amplify errors. The same margin analysis yields an oracle water-filling law for compute allocation. A synthetic audit quantifies the optimism caused by adaptive budget selection, and a 192-example GSM8K audit using gpt-4.1-mini generations illustrates risk-cost reporting for a real inference pipeline. The results give a lightweight, distribution-free way to report risk and compute claims for adaptive inference systems.
PaperID: 2236, Poster
Abstract: Recommender systems (RS) continually learn from new user-item interactions while addressing heterogeneous deletion requests. Existing continual recommendation methods focus on retention but do not erase targeted interactions, whereas recommendation unlearning methods are largely offline and do not ensure stability after future learning updates. This creates a critical failure mode: deleted signals can be reactivated through shared user--item representations, which we call collaborative signal regrowth. To address this, we introduce ALIGN-Rec, a model-agnostic online framework for continual recommendation under interleaved learn/unlearn requests. Specifically, ALIGN-Rec tracks low-rank curvature surrogates to identify retain and forget subspaces, constructs compact retain/forget summaries, and performs geometry-aware updates that preserve retained utility while suppressing deleted-signal directions. We instantiate ALIGN-Rec with CARE, an efficient optimizer based on randomized curvature sketching and adaptive summary selection. We provide theoretical guarantees for geometry tracking, summary approximation, and directional suppression in the forget subspace. Extensive experiments across datasets, deletion granularities, and stream settings demonstrate effectiveness aligning with our theoretical guarantees. The code and implementation are available at https://anonymous.4open.science/r/alignrec-FFD8.
PaperID: 2237, Poster
Abstract: Incremental learning aims to continuously learn from dynamic data streams without suffering from catastrophic forgetting. However, existing methods may fail in challenging scenarios where class and domain spaces expand simultaneously, struggling to learn continually and stably from fragmented data streams. To address this challenge, we propose EMCift (Extensible Multi-Center Modeling with Dual-Path Drift Harmonization), a novel model that treats heterogeneous data across domains and semantic spaces as co-existing multi-view observations, enabling the construction of global knowledge from local observations. Our approach organizes extensible multi-center prototypes as a tree hierarchy, allowing the global class representation across different domains to grow organically from fragmented local data. To sustain this dynamic topology, we introduce a dual-path drift harmonization strategy that balances structural stability with adaptive plasticity. Specifically, one path captures essential updates to enable prototype evolution, while the other path leverages a conditional adversarial mechanism to robustly initialize domains, effectively preventing catastrophic forgetting caused by class confusion. Extensive experiments on three datasets demonstrate the state-of-the-art performance of our proposed model in the class-domain incremental learning scenario. Code will be released upon acceptance.
PaperID: 2238, Poster
Abstract: Early prediction of chronic disease risk from electronic health records (EHRs) is challenging because early clinical signals are often sparse, noisy, and indirectly related to future outcomes. Existing methods typically learn from prediction-time records and final outcome labels alone, which provides limited supervision for identifying weak early signals from heterogeneous and confounded clinical records. LLM-generated rationales offer an intermediate form of prediction-time reasoning, but rationales derived from early EHR may still miss weak signals whose relevance becomes clearer only in later records. We therefore use follow-up EHR observed after the prediction time but before outcome assessment as training-time hindsight and propose On-Policy Hindsight Distillation (OPHD), a self-distillation framework that converts follow-up EHR into training signals for LLM-generated prediction-time rationales. Specifically, OPHD first uses an LLM to generate a rationale from prediction-time EHR alone, then re-evaluates the same rationale under privileged follow-up views to derive token-level signals that are distilled back to the prediction-time policy. To handle noisy follow-up records, OPHD emphasizes hindsight signals that are stable across perturbed privileged views. A task-grounded scorer further prioritizes rationales that improve downstream risk discrimination. Experiments on multiple real-world neurodegenerative disease cohorts show that OPHD achieves the best long-horizon prediction performance compared with a variety of baselines.
Authors: Hamza Golubovic, Matthew Shen, Genevera Allen, Tarek M Zikry
Abstract: Modern matrix completion problems often involve heterogeneous data whose rows simultaneously belong to many meta-categories such as demographic and age groups in recommendation systems, or region and recording session labels in electrophysiology experiments. Standard low-rank estimators impose a single global latent geometry, which can recover average structure but may smooth away subgroup-specific variation, especially when observations are unevenly distributed across groups. We introduce Group-Aware Matrix Estimation (GAME), a convex estimator for overlapping subgroup-wise low-rank matrix estimation. GAME regularizes category-specific submatrices through overlapping nuclear-norm penalties, allowing related groups to borrow information while preserving local latent structure in a shared coordinate system. We provide finite-sample guarantees for both reconstruction error and subgroup-specific subspace recovery, showing how performance depends on sampling density, subgroup rank, and overlap structure. Experiments on synthetic, recommendation, ecological, and neuroscience datasets show that GAME is most beneficial in structured missingness regimes, where subgroup-aware regularization improves both reconstruction accuracy and latent subspace fidelity. Across these benchmarks, GAME outperforms global low-rank, side-information, and modern imputation baselines, with the largest gains when subgroup heterogeneity and latent geometry are central to the task.
PaperID: 2240, Poster
Authors: Zeyu M Li, William X Chen, Xiang Cheng
Abstract: We study path--flow alignment as a unified training objective for flow matching. Instead of fixing the interpolation path and learning only the velocity field, we jointly train an endpoint-preserving path network and a flow network using the same alignment loss: the flow learns to match the path velocity, and the path learns to align its velocity to the current flow. Although every fixed learned path defines a valid flow-matching objective, the alignment loss alone is not a reliable criterion for path learning. We identify path overfitting, a failure mode in which the alignment loss decreases while sample quality worsens. We trace this failure to low-entropy bottlenecks in the induced probability path, where the learned interpolant routes samples through overly concentrated intermediate marginals. Motivated by this diagnosis, we introduce a stochastic path regularizer that hides part of the source information from the path network while preserving exact endpoints. The resulting regularized objective controls entropy collapse and makes joint path--flow training effective. On ImageNet-256×256 with SiT backbones, our method consistently improves FID across model scales, extends to model-guidance training, and leaves the inference-time architecture and sampler unchanged. Code is available at https://anonymous.4open.science/r/traj_opt_paper-4D80.
PaperID: 2241, Poster
Abstract: Accurately rendering temporal composition is fundamental to text-to-video (T2V) generation. While recent efforts incorporate temporal structure for long-horizon coherence, the intrinsic temporal understanding capability of T2V models remains underexplored. We observe that current models still struggle to respect canonical operators such as , usually accompanied by missing entities and incorrect action binding. In this work, we propose LogicDirector, a test-time guidance framework that enforces temporal composition through executable logical specifications. By translating text prompts into first-order logic constraints over attention-derived evidence, LogicDirector constructs a neuro-symbolic verifier to evaluate noisy latent states during diffusion sampling for entity grounding, action binding, and temporal adherence. To enforce these constraints without expensive gradient-based optimization, we introduce a gradient-free Best-of- latent search strategy, which preserves the model's native diffusion dynamics while steering logical satisfaction. To systematically study operator-level temporal composition, we curate TempBench, a diagnostic benchmark spanning four canonical relations, together with a hierarchical evaluation protocol that disentangles entity existence, event realization, and temporal correctness. Experiments on CogVideoX and Wan 2.1 demonstrate that our method significantly improves temporal composition while maintaining visual fidelity.
PaperID: 2242, Poster
Abstract: This paper addresses the problem of sequential decision-making under learning budget constraints. Such settings naturally arise in applications like managing a portfolio of bandit or reinforcement learning (RL) algorithms. We propose a novel UCB-type algorithm, M-LCB, designed to manage a pool of K self-learning experts in a stochastic environment while accounting for a limited per-round learning budget M. At each round, M-LCB selects one expert to make a decision and at most M \le K experts to learn. For selection, M-LCB uses confidence bounds constructed from limited prior knowledge about the experts (i.e., mild assumptions) and their observed training losses. We derive anytime regret bounds for M-LCB that scale with the individual regrets of the experts. In particular, if each expert has regret \tilde O(T^\alpha) by round T, then M-LCB guarantees an overall regret of \tilde O\left(\sqrtKT/M + (K/M)^1-\alphaT^\alpha\right) relative to the best expert in hindsight. Finally, we demonstrate the applicability of M-LCB using self-learning experts instantiated as (i) parametric models and (ii) bandit algorithms.
Abstract: Self-improvement at scale has been a longstanding goal for reasoning models, and there are two natural places to do it: at test time, through verification-refinement (V-R) loops; and at training time, through self-training methods. Both are gated by the same bottleneck: the verifier. V-R loops stall when verifier scores inflate over rounds while accuracy stagnates, and when feedback is too generic to act on; self-training fails similarly when bad self-generated data are added to training. Better verifiers would unlock both, but the capability we want to train, i.e., catching errors the generator cannot detect in its own work, has neither finetuning data nor a verifiable reward. To address this challenge, we propose (STV). Our key observation is that, while a model cannot catch these errors alone, it can when shown the reference solution. We turn this asymmetry into a supervision target and train the verifier to imitate a more informed version of itself. At test time, STV substantially improves V-R loops on hard problems, while standard alternatives (e.g., SFT, RL on verifier scores, and even meta-verifiers) do not. STV roughly doubles accuracy on hard DAPO math and lifts it 14x on the hardest SciKnowEval problems (1.5% to 21%). At training time, starting from an RL-converged generator, putting the STV verifier in the loop yields a further 33% relative gain in test-time pass@1. More notably, the generator's standalone pass@1, with no verifier at inference time, climbs 30% relative past where continued RL had converged. Hence, the next frontier in reasoning may lie in how we train verifiers.
PaperID: 2244, Poster
Abstract: Recent multimodal large language models (MLLMs) achieve strong results on video question answering, yet often fail to recover the temporal structure that makes a video coherent. They may merge neighboring events, miss single-event progression, or infer implausible relations across events. We study structured temporal understanding as the ability to recover event structure from visual evidence, including event boundaries, single-event progression, and multi-event relations. We argue that this gap arises from a mismatch between current post-training objectives and temporal reasoning, where caption-style supervision allows models to rely on language shortcuts rather than visual temporal grounding. To address this, we propose EventLens, a vision-centric reinforcement learning framework that learns video understanding through event-structure recovery. Instead of reconstructing captions, EventLens trains models to recover temporal structure under controlled perturbations, such as altered temporal granularity, reversed progression, and shuffled event order. We instantiate this framework with a three-level temporal ontology and derive three verifiable task families: event segmentation, progression discrimination, and multi-event relation reasoning. These tasks admit deterministic rewards, enabling task-specialized GRPO training without human preference labels or learned reward models. We further introduce multi-teacher on-policy distillation (MT-OPD) to consolidate specialized policies into a unified model. Experiments on Qwen3-VL backbones show consistent improvements on temporal grounding and event-sensitive reasoning benchmarks, while preserving general video-QA performance. Learning curves further demonstrate that EventLens achieves comparable downstream transfer with significantly reduced training cost compared to direct mixed RL optimization.
PaperID: 2245, Poster
Abstract: Muon improves training efficiency over Adam in large language-model training by about two times, but the local geometric source of this advantage remains unclear. Our work takes a first step toward demystifying Muon's superiority over Adam from a curvature perspective. First, we apply a second-order Taylor approximation to the training landscape and show that Muon achieves a larger one-step loss decrease than Adam at matched validation loss. The two optimizers have comparable first-order gains, but Muon consistently incurs a smaller second-order curvature penalty. Second, we decompose this curvature penalty into the squared update norm and Normalized Directional Sharpness (NDS). We find that Muon and Adam have comparable update norms, so Muon's smaller curvature penalty is driven by lower NDS, not update scale. Third, we study how training data and model structure shape Muon's NDS advantage. Using Zipf Probabilistic Context-Free Grammar (PCFG) data with controlled imbalance, we show that data imbalance amplifies Muon's NDS advantage over Adam. A within-/cross-layer decomposition further shows that, in the middle and late stages of training, Muon's lower NDS is mainly sustained by smaller within-layer curvature. Beyond empirical evidence, we analyze stylized quadratic problems with heterogeneous curvature and gradient alignment toward high-curvature modes. We prove that Muon attains a smaller average NDS than GD by balancing update energy across curvature groups; when curvature heterogeneity is sufficiently strong, this also yields lower local quadratic loss after the same number of steps.
PaperID: 2246, Poster
Abstract: As privacy, copyright, and safety requirements evolve, federated learning systems face a growing imperative to accommodate sequential unlearning requests. However, existing federated unlearning methods are designed for single-shot settings and their direct sequential application may lead to uncontrolled parameter-space overlap, causing cross-request interference that undermines prior forgetting. To address this, we propose the Spectral Orthogonality for Federated Continual unlearning (FedSOUL), a framework that reduces direct overlap among sequential unlearning updates. The core idea is to structure model updates in a shared orthonormal spectral basis, where each client is assigned a dedicated subspace in the frequency domain. Under this design, each unlearning request updates only its associated subspace, while all others remain unchanged, ensuring that sequential updates do not overlap in the allocated spectral coefficient space and thereby reducing cross-request interference. Beyond interference control, FedSOUL also improves efficiency by operating on sparse spectral coefficients, using FFT-based transforms for efficient computation and communicating only sparse coefficient updates. From a theoretical perspective, we relate continual unlearning to geometric conditions on update directions and show that our spectral construction satisfies update orthogonality by design, while functional preservation depends on bounded projected gradient leakage. Extensive experiments demonstrate that FedSOUL maintains stable unlearning–utility trade-offs across up to ten sequential requests and consistently outperforms competitive federated baselines. Our code is available at https://anonymous.4open.science/r/fedunlearning-1B92/.
PaperID: 2247, Poster
Authors:
Yashuo Luo, Tongxu Wang, Siyuan He, Chunyu WeiAbstract: it computes those outputs unconstrained, so students may match teacher behavior through entirely different internal mechanisms. We propose , which reframes distillation as the explicit transfer of reasoning circuits, the structured computational pathways uncovered by mechanistic interpretability. MCD represents each model's computation as a transcoder-based attribution graph and aligns teacher and student circuits via an optimal-transport pathway-matching loss that plugs into any standard distillation pipeline. We prove that minimizing this loss bounds the discrepancy in MLP-level computation between teacher and student. Across two model families, six distillation objectives, and three instruction-following benchmarks, MCD delivers consistent improvements that transfer to mathematical reasoning, with the largest gains under aggressive compression.
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for enhancing reasoning in Large Language Models (LLMs). However, existing reward formulations typically treat exploration and consolidation as a monolithic process, resulting in entangled stage-wise learning dynamics. This contradicts the natural learning behavior of human learners. In human learning, individuals adopt distinct behavioral patterns toward mastered versus unfamiliar problems. When confronting unmastered challenges, humans prioritize broad exploration to seek viable solutions. By contrast, for well-mastered problems, they focus instead on reasoning condensation and knowledge abstraction to distill concise underlying principles. Motivated by this gap, we introduce T2T(Thickening-to-Thinning), a dynamic reward framework inspired by human learning processes. Specifically, it implements a dual-phase mechanism: (1) On incorrect attempts, T2T incentivizes "thickening" to broaden the search space and explore novel solution paths; (2) Upon achieving correctness, it shifts to "thinning", imposing length penalties to discourage redundancy, thereby fostering model confidence and crystallizing reasoning capabilities. Extensive experiments on mathematical benchmarks (MATH-500, AIME, AMC) across 5 mainstream LLMs demonstrate that T2T significantly outperforms standard GRPO and recent baselines, achieving superior performance.
PaperID: 2249, Poster
Authors:
Jiahui Sheng, Le Li, Limin Lin, Jiahao Li, Wansen Feng, Hong Gu, Shuhan Chen, Xiaorun Li, Rui HuAbstract: Portrait video relighting aims to modify the illumination of a portrait video while preserving photorealism and temporal stability. The task becomes particularly challenging under dynamic lighting, where illumination direction, intensity, and color vary over time. To tackle this problem, we introduce FlashRelight, a novel two-stage framework for dynamic portrait video relighting. FlashRelight first suppresses the source illumination through delighting, and then synthesizes the target dynamic lighting on an illumination-neutral portrait video. FlashRelight supports three complementary control modalities: envmap sequences for explicit frame-level lighting control, dynamic lighting text for semantic illumination editing, and lighting reference videos for visual lighting transfer. Training follows a low-cost hybrid strategy that combines static one-light-at-a-time (OLAT) data with generated dynamic-lighting portrait videos, covering both controlled illumination variation and realistic portrait motion. Extensive experiments show that FlashRelight achieves temporally coherent and photorealistic dynamic relighting while enabling flexible cross-video lighting transfer.
Abstract: We study offline learning in KL-regularized two-player zero-sum games, where policies are optimized with respect to a fixed reference policy through KL regularization. Prior work relies on pessimistic value estimation to handle distribution shift, yielding only \widetilde\mathcalO(1/\sqrt n) statistical rates. We develop a new pessimism-free algorithm and analytical framework for KL-regularized games, built on the smoothness of KL-regularized best responses and a stability property of the Nash equilibrium induced by skew symmetry. This yields, to our knowledge, the first pessimism-free offline learning guarantee for KL-regularized games, with a fast \widetilde\mathcalO(1/n) sample complexity bound. We further propose an efficient self-play policy optimization algorithm that replaces exact equilibrium computation with iterative KL-regularized policy updates, and prove that its last iterate preserves the same pessimism-free statistical guarantee up to a controlled optimization error.
PaperID: 2251, Poster
Authors:
Gavin R Brown, Ephraim Linder, Mahbod Majid, Vikrant SinghalAbstract: We study efficient differentially private algorithms for estimating monotone statistics, i.e., statistics that are monotone under the addition of new observations. The starting point for our investigation is subsample-and-aggregate: a classical paradigm that partitions the dataset into blocks, estimates the statistic on each block, and then privately aggregates the estimates. While practical and generically applicable, this approach is quite data-hungry. We improve upon this framework for the class of monotone statistics---compared to subsample-and-aggregate, our algorithms save a factor of t in sample complexity and pay a factor of e^t in running time, where t>0 is a tunable parameter. We complement our results with a query-complexity lower bound, showing that our algorithms are essentially optimal for this task. As an application, we obtain improved results for private eigenvalue estimation, private loss estimation, and privately estimating a single parameter of a high-dimensional model, e.g., in linear regression.
PaperID: 2252, Poster
Authors: Luc Blassel, Noémie Sauvage, Pierre Barrat-Charlaix, Bastien Boussau, Nicolas Lartillot, Laurent Jacob
Abstract: Phylogenetic inference, the task of reconstructing how related sequences evolved from common ancestors, is a central objective in evolutionary genomics. The current state-of-the-art methods exploit probabilistic models of sequence evolution along phylogenetic trees, by searching for the tree maximizing the likelihood of observed sequences, or by estimating the posterior of the tree given the sequences in a Bayesian framework. Both approaches typically require to compute likelihoods, which is only feasible under simplifying assumptions such as independence of the evolution at the different positions of the sequence, and even then remains a costly operation. Here we present the first likelihood-free inference method for posterior distributions over phylogenies. It exploits a novel expressive encoding for pairs of sequences, and a parameterized probability distribution factorized over a succession of subtree merges. The resulting network provides well-calibrated estimates of the posterior distribution leading to more accurate tree topologies than existing methods, even under models amenable to likelihood computation. We further show that its edge against likelihood-based methods dramatically increases under models of sequence evolution with intractable likelihoods.
PaperID: 2253, Poster
Abstract: Learning the dynamics of interacting physical systems is a central challenge in machine learning for the physical sciences. Prevailing architectures such as Graph Neural Networks (GNNs) and feedforward models often fail to explicitly capture the compositional nature of physical laws. Instead of modeling the forces that govern a system, these networks learn to approximate aggregate state transitions. As a result, they lack a causal understanding of the underlying physics interactions, leading to poor generalization in out-of-distribution systems with varying numbers of entities or novel configurations. In this work, we introduce Compositional Potential Minimization (CPM), a framework that casts simulation as an energy minimization procedure over a composition of learned force potentials. Each force component is represented as a separate, interpretable energy landscape, and interaction dynamics emerge from their addition. Inference in CPM is framed as a global energy minimization problem over an arbitrary number of composed functions. Because CPM learns to disentangle separate forces during training, it enables seamless generalization during test time to larger-scale systems and unseen configurations through straightforward compositions of the learned energy components. Our experiments show that CPM significantly outperforms previous state-of-the-art GNN-based solutions in 2D, 3D, and particle-based simulations.
Authors:
Phu-Quy Nguyen-Lam, Phu-Hoa Pham, Dao S Minh, Chi Nguyen Tran, Trung-Kiet Huynh, Long Tran-ThanhAbstract: Object-centric representations promise a key property for few-shot learning: Rather than treating a scene as a single unit, a model can decompose it into individual object-level parts that can be matched and compared across different concepts. In practice, this potential is rarely realized. Continual learners either collapse scenes into global embeddings, or train with part-level matching objectives that tie representations too closely to seen patterns, leaving them unable to generalize to truly novel concepts. In this paper, we identify this fundamental structural conflict and pioneer a new paradigm that strictly decouples representation learning from compositional inference. Leveraging the inherent patch-level semantic geometry of self-supervised Vision Transformers (ViTs), our framework employs a dual-phase strategy. During training, slot representations are optimized entirely toward holistic class identity, preserving highly generalizable, object-level geometries. At inference, preserved slots are dynamically composed to match novel scenes. We demonstrate that this paradigm offers dual structural benefits: The frozen backbone naturally prevents representation drift, while our lightweight, holistic optimization preserves the features' capacity for novel-concept transfer. Extensive experiments validate this approach, achieving state-of-the-art unseen-concept generalization and minimal forgetting across standard continual learning benchmarks.
Authors:
Junwon You, Mihyun Jang, Sangwoo Mo, Jae-Hun JungAbstract: Vision-language models have shown strong performance, but they often generalize poorly to specialized domains. While semi-supervised vision-language learning mitigates this limitation by leveraging a small set of labeled image-text pairs together with abundant unlabeled images, existing methods remain fundamentally pairwise and fail to model the global structure of multimodal representation manifolds. Existing topology-based alignment methods rely on persistence diagram matching, which neither guarantees geometric alignment nor utilizes the image-text pairing information central to vision-language learning. We propose Topology-Aware Multimodal Representation Alignment (ToMA), a framework that uses persistent homology to identify topologically salient edges and aligns them across modalities through available cross-modal correspondences. ToMA leverages both H_0-death edges and lightweight H_1-birth edges, allowing it to capture both connectivity and cycle structure without constructing 2-simplices. Experiments show that ToMA yields stable gains, with clear improvements on remote sensing and modest but consistent benefits on fashion retrieval. Additional analysis shows that ToMA is more stable than alternative topology-based objectives and that lightweight H_1-birth edges provide useful higher-order structural signals.
Abstract: Molecular docking requires reasoning jointly about ligand pose and protein flexibility. Most diffusion-based docking models predict torsional updates with generic Euclidean heads that ignore the periodic geometry of angular variables. This mismatch is especially limiting in flexible docking, where ligand conformations and pocket side chains co-adapt to form the bound complex. Here, we introduce Harmony, a harmonic torsional diffusion framework for flexible protein-ligand docking. Harmony parameterizes ligand and side-chain torsional score fields as derivatives of learned harmonic potentials on the circle, whose noise-level dependence is supplied analytically by the heat semigroup of variance-exploding diffusion on the torus. This construction makes periodicity explicit, removes the need to learn noise conditioning for torsional updates, and gives the model a frequency-aware inductive bias over rotameric motion. On the PDBBind benchmark, Harmony improves ligand pose accuracy and pocket all-atom reconstruction over recent flexible docking methods. On PoseBusters, it improves the physical validity of generated complexes. Case studies on EBNA1 and KRAS G12D illustrate the method's behavior on a polar and a shallow binding site, respectively. Together, these results indicate that aligning the score parameterization with the geometry of the diffusion process is a simple and effective lever for improving flexible docking.
Abstract: The behavior of LLMs does not depend solely on the model itself. Components of the inference system, such as the inference engine, attention backend, and hardware platform, subtly influence how inputs are processed. These components differ in their implementations and thereby induce small numerical deviations across systems when running the same model. While prior work has established the theoretical existence of such deviations, their security implications have remained unexplored. In this paper, we show that these deviations are characteristic of specific components and propagate to observable textual outputs, exposing the inference system to any party that can query the model. Building on this observation, we introduce a fingerprinting method that analyzes the prompt-response behavior of LLMs to identify components of the inference system. Our empirical evaluation demonstrates that the inference engine, attention backend, and underlying hardware platform can be identified reliably, even when the LLM is operated at non-zero temperature. We show that preventing fingerprinting is fundamentally hard, as it would require eliminating numerical differences between hardware and software stacks. We therefore propose partial mitigations and discuss their impact.
Abstract: Generative models such as diffusion and flow matching have become dominant paradigms for visuomotor policy learning, yet their reliance on iterative denoising incurs high inference latency incompatible with real-time robotic control. We present a Fast Legendre-polynomial Action policy via Sparse History-anchored flow (FLASH), which replaces discrete action-chunk generation with continuous Legendre polynomial trajectory representation. Specifically, by fitting expert demonstrations under sparse temporal sampling, FLASH enables a single inference to cover a significantly extended action horizon. To further accelerate generation, FLASH initiates the flow matching process from history polynomial coefficients rather than uninformative Gaussian noise, shortening the transport distance and enabling accurate single-step inference. Moreover, analytic polynomial differentiation directly provides desired velocity feed-forward signals to the torque controller without numerical approximation. Extensive experiments on five simulated and two real-world manipulation tasks demonstrate that FLASH achieves state-of-the-art success rates (\ge 92% across all tasks), a per-episode inference time of 31.40\,ms (up to 175× faster than diffusion policies and 18× faster than prior flow matching policies), up to 4× faster training convergence than ACT, and 5× to 7× reduction in controller tracking error compared to discrete-action baselines.
PaperID: 2259, Poster
Abstract: Multi-task learning requires a shared update that makes balanced progress across multiple objectives. A key obstacle is that progress is difficult to compare across tasks with different loss scales, units, and local optimization dynamics. We study this problem from an optimality-gap perspective. For each task, we compare the improvement induced by the shared update with the best one-step improvement achievable by optimizing that task in isolation. This comparison yields a scale-invariant normalized improvement rate, which measures how well the shared update serves each task relative to its own local improvement bound. We formulate balanced multi-task optimization as a max-min problem over the normalized improvement rates, thereby prioritizing the worst-served task. A local quadratic approximation leads to the Karush-Kuhn-Tucker (KKT) optimality conditions, showing that the optimal shared update is determined by a small set of bottleneck tasks. Under an isotropic Hessian approximation, the update admits a closed-form expression that depends only on task gradients and a regularized Gram matrix, while the bottleneck set is identified by an active-set procedure. Experiments on diverse benchmarks show improved task balance and competitive performance.
Abstract: We study models for human-AI teaming through the lens of statistical calibration. We assume the team consists of an AI model and human---both of which are calibrated with respect to some partitioning of the feature space---and expose how the calibration assumptions propagate into the teaming framework. In particular, we consider frameworks that either (i) combine human and model predictions or (ii) delegate prediction responsibility to either a human or model. We show via theoretical and empirical results that existing methods for combination do not preserve the human's degree of calibration. Methods for delegation (by the very act of delegation) preserve calibration of the downstream predictors but can place high-capacity demands on the meta-classifier that determines which should predict. This latter result suggests human-AI complementarity and correct delegation place opposite demands on the delegation model.
PaperID: 2261, Poster
Abstract: Graph generation, which aims to produce new graphs from a distribution similar to observed data, has gained increasing attention, especially as generated graphs are used in high-stakes decision-making where fairness is critical. However, most existing fair graph generation methods assume full access to sensitive attributes, an assumption often violated in practice due to privacy concerns, regulatory constraints, and missing data. To this end, we propose to solve the problem from a new perspective, where sensitive proxy inference is reformulated as a component of the graph generation pipeline, enabling the inferred proxy to guide the generation process. We further model and mitigate the impact of proxy inference errors, and provide theoretical guarantees that quantify how such errors propagate to fairness outcomes in the generated graphs, offering practical guidance for interpreting fairness when sensitive attribute is missing. Experiments on benchmark datasets demonstrate that our method consistently improves fairness while maintaining competitive generation quality.
PaperID: 2262, Poster
Abstract: We study the instance-dependent sample complexity of one-dimensional stochastic bandit convex optimization in the fixed-confidence setting. First, we consider minimizer localization, where the learner seeks to identify a point within distance \varepsilon of the unique minimizer. We introduce a new instance-dependent quantity \Delta_\varepsilon(f), corresponding to the threshold for which the sublevel set of f-\min f has length 2\varepsilon. We prove a lower bound of order \Delta_\varepsilon(f)^-2 and give a bisection-style algorithm that matches this rate up to logarithmic factors. We then revisit the simple-regret setting, where the goal is to identify an \varepsilon-optimal point. Surprisingly, no comparable instance-dependent improvement is possible: every correct algorithm has sample complexity of order \varepsilon^-2 on every instance, matching the minimax rate up to constants. Together, these results separate the statistical complexity of minimizer localization from that of simple-regret optimization in one-dimensional convex bandits.
PaperID: 2263, Poster
Abstract: Prior Laplacian-based option-discovery methods build options from 0-form Laplacians defined on nodes, where the resulting eigenvectors induce routes toward isolated boundaries or extrema, thereby skewing visitation toward specific states. In contrast, kernel eigenvectors of the 1-form Hodge Laplacian represent interpretable, topology-aware edge flows and naturally capture cyclic behaviour known as harmonic flows. These kernel eigenvectors are particularly attractive because they define zero-divergence edge flows, offering a natural solution for non-isolated visitation to sources and sinks in regions in which they are active. Despite this appeal, harmonic flows have not previously been explored as a basis for option discovery in reinforcement learning. This is what we do here. We first introduce harmonic-flow skills and show that, on state graphs with tree-like “dead-end” appendages, purely harmonic flows can become nearly inactive in those regions, which can limit coverage. To address this limitation, we propose quasi-harmonic flow options: skill-inducing edge flows that preserve harmonic cycles where harmonic activity is strong, while injecting drift into appendages using the lowest-magnitude non-harmonic eigenvector of the Hodge Laplacian, a low divergence solution, where harmonic activity is weak. Our approach is eigenvector-only, relying solely on spectral components of the Hodge Laplacian. We compare quasi-harmonic flow options against a harmonic-flow options baseline and state-of-the-art skill-discovery methods from the literature, demonstrating empirical gains and analytical improvements on hitting-time-based metrics.
Abstract: Time-Series Foundation Models (TSFMs) excel at zero-shot unimodal forecasting using numerical data, but unlike LLMs they cannot consume multimodal, non-numerical context that often shape real-world trajectories. In this work, we bridge this gap and argue for a multimodal time-series forecasting approach that post-trains LLMs to act as context-guided revisors over strong numerical TSFM priors. We introduce PostTime, a post-training recipe combining Supervised Fine-Tuning (SFT) and Reinforcement Learning with Verifiable Rewards (RLVR), along with a methodology to generate automated reasoning traces for forecast revisions. PostTime teaches an LLM to generate context-conditioned forecast interventions—decisions to revise, preserve, or ignore the TSFM prior based on the multimodal context. We evaluate this approach on the TimesX multimodal forecasting benchmark using a Gemma-3-4B LLM and TimesFM-2.5 TSFM, and show that it significantly outperforms standalone TSFMs, LLM-only baselines, and existing multimodal forecasting approaches. Our post-training recipe including code and dataset, can currently be accessed for review purposes at \urlhttps://anonymous.4open.science/status/posttime-21BF/.
Abstract: Perception for action suggests that representations of the world should be shaped not by visual fidelity alone, but by their relevance for actions. At the same time, latent JEPA-style world models advocate learning compact predictive states from high-dimensional observations to facilitate the prediction of future states, but end-to-end training of these models is nontrivial because representations may collapse if our only goal is to construct a latent state that is easy to predict. We show that a suitable inverse dynamics regularization addresses both issues: it is an effective anti-collapse mechanism that induces action-aligned representations. By forcing latent states to preserve information about the action underlying a transition, it biases the model toward the controllable degrees of freedom of the environment while discarding uncontrollable distractors. This yields stable latent world models trained end-to-end from offline, reward-free trajectories, without frozen encoders, exponential moving averages, or complex latent regularizers. Empirically, the learned latent spaces are compact, interpretable, and enable competitive planning performance across simple 2D and 3D control tasks.
Abstract: Transformer training systems are built around dense linear algebra, yet a nontrivial fraction of end-to-end time is spent in the memory-bound operators that surround it. Normalization, activations, residual updates, reductions, and related computations repeatedly move large intermediate tensors through global memory while performing little arithmetic, making data movement an increasingly important bottleneck in otherwise highly optimized training stacks. We introduce CODA, a GPU kernel abstraction that expresses these computations as GEMM-plus-epilogue programs. CODA is based on the observation that many Transformer operators exposed as separate framework kernels can be algebraically reparameterized so that their work executes while a GEMM output tile remains on chip, before it is written to memory. The abstraction fixes the GEMM mainloop and exposes a small set of composable epilogue primitives for scaling, reductions, pairwise transformations, and accumulation. This constrained interface preserves the performance structure of expert-written GEMMs while remaining expressive enough to cover nearly all non-attention computation in the forward and backward pass of a standard Transformer block. Across representative Transformer workloads, both human- and LLM-authored CODA kernels achieve high performance, suggesting that GEMM-plus-epilogue programming offers a practical path toward combining framework-level productivity with hardware-level efficiency.
PaperID: 2267, Poster
Abstract: Directed Acyclic Graphs (DAGs) are foundational to many areas of AI research, providing a formal language for causal reasoning and model interpretability. However, learning DAGs remains challenging, due to super-exponential computational cost and diminished accuracy in small sample regimes. To address these challenges, we introduce Attention-DAG (ADAG), a novel linear transformer architecture for unsupervised amortized DAG learning. Unlike traditional unsupervised DAG learning methods which recover the graph structure from data of each individual domain, ADAG leverages the knowledge from multiple domains. It provides a nonlinear mapping from observational data of each domain to both the corresponding graph structure and the underlying parameters of Structural Equation Models (SEMs). This enables efficient zero-shot inference of causal structures in new domains with unseen SEM parameters. Remarkably, we demonstrate that for DAG learning problem attention acts as a fixed-point iterative scheme, so the training on multiple domains effectively discovers an efficient solver for the constraint optimization problem. Extensive evaluations on synthetic and realistic benchmarks demonstrate that ADAG significantly outperforms existing baselines in accuracy and efficiency, particularly when data is scarce.
PaperID: 2268, Poster
Abstract: Biomolecular therapeutics often start from known sequences and require targeted editing to improve multiple properties while satisfying hard biochemical and manufacturability constraints. Existing generative methods do not jointly support multi-objective optimization, hard feasibility, and sequence editing in discrete, variable-length biological spaces. We introduce ), a framework built on discrete flow matching that steers a pre-trained Edit Flow toward user-specified preferences while enforcing terminal feasibility. pCoMole defines a feasibility-gated terminal distribution using an augmented Tchebycheff utility and realizes the resulting preference tilt through a Doob-h transform of the underlying edit process. To make this construction practical, we approximate the required harmonic function using short Monte Carlo rollouts over candidate edits, yielding an efficient guided editor with provable preference consistency. We validate pCoMole by shrinking GFP while retaining fluorescence-related properties, shortening diverse Cas9 orthologs while preserving PAM specificity, and compressing peptide binders into short peptidomimetics that optimize seven drug-related properties under hard constraints. Overall, pCoMole enables constraint-aware, Pareto-aligned editing of biomolecular sequences in discrete, variable-length spaces.
PaperID: 2269, Poster
Abstract: Pre-trained discrete-token vision-language models (VLMs) generate images by sampling sequences of codebook indices, but their visual fidelity is bounded by the VQ-VAE decoding stage. Prior work addresses this bottleneck by replacing the discrete tokenizer with continuous or hybrid visual token representations, or by integrating diffusion into the generation pipeline; both routes require foundation-scale retraining or realignment of VLM, which is undesirable when the VLM's broader capabilities should be preserved or when retraining is infeasible. We show that such retraining is unnecessary: before each token is sampled, discrete-token VLMs already compute a full logit distribution over the codebook, but standard decoding discards it after selecting a single index. Under Otsu's threshold, the pre-sampling logits of capable VLMs assign statistically meaningful probability to at least two codebook entries per token, and recovering this distribution can substantially close the fidelity gap. We propose DCDD, a post-hoc framework that conditions a diffusion decoder on this discarded distribution while leaving the VLM frozen. At inference, DCDD converts the VLM's pre-sampling logits into distribution-weighted code vectors; during training, it uses VQ-VAE encoder-derived proxy logits calibrated to match the distributional statistics of the VLM's inference-time logits. Our diffusion decoder then maps these representations to high-fidelity images as an optional alternative to the native VQ-VAE decoder. Trained on ImageNet-1K for 50K steps, DCDD reduces reconstruction FID by 74% and Janus-Pro 7B's generation FID on MJHQ-30K by 34%, while semantic-alignment across GenEval, DPG-Bench, and WISE remain consistent. Comparison against a diffusion model from the same backbone family applied as an image-to-image refiner confirms that the gain comes not from the diffusion model alone, but from conditioning on the VLM's pre-sampling logit distribution.
PaperID: 2270, Poster
Abstract: Recent advances in deep learning for shadow removal have significantly enhanced image quality and realism. However, most approaches rely on real-world paired datasets, which are costly to collect and often limited in scene diversity, leading to limited generalization. To address these limitations, we propose One-step Deshadow Diffusion via Reward guidance (ODDR), a new framework that achieves efficient and high-fidelity shadow removal without relying on real-world paired supervision. Our method begins with One-step Deshadow Diffusion (ODD), a baseline model trained on synthetic shadow data for efficient one-step shadow-free reconstruction. We further adapt ODD into ODDR using ShadowReward. In contrast to traditional, annotation-heavy approaches, ShadowReward is the first reward model for shadow removal trained entirely without human annotation. It learns to mimic human perceptual judgments by ranking synthetically generated images with controlled degradations, such as texture distortion and boundary artifacts. This reward-guided fine-tuning enables ODDR to close the synthetic-to-real domain gap. Extensive experiments show that ODD achieves strong performance without relying on real-world paired supervision, and ODDR further improves the results and achieves performance competitive with methods trained on real-world paired data, while maintaining higher computational efficiency as a single step model.
Abstract: We introduce a new measure of robustness for statistical estimators, which we call empirical sensitivity. An estimator \hat \mu has bounded empirical sensitivity if, with high probability over a dataset X = (X_1, \dots, X_n) ~ \mathcalD^\otimes n, for any dataset Y obtained by modifying at most \eta n points in X, we have that \hat \mu(Y) is close to \hat \mu(X). We study bounds on this quantity for the prototypical problem of Gaussian mean estimation. We prove new lower bounds, showing that for any estimator \hat \mu which achieves an optimal \ell_2-error bound of O(\sqrtd/n), the empirical sensitivity is at least \Omega(\eta + \sqrt\eta d/n). The two terms arise due to obstructions on the mean and variance (via an Efron-Stein argument) of such an estimator. We show that this bound is tight up to logarithmic factors, by employing recent results for robust empirical mean estimation.
PaperID: 2272, Poster
Abstract: Test-time adaptation (TTA) for vision-language models (VLMs) is usually studied in a fully unsupervised regime, where adaptation relies on pseudo-labels, confidence scores, or entropy computed by the model itself. This is a severe restriction on hard shifted samples: the model is asked to repair its own uncertain predictions without any external evidence. At the other extreme, supervised online adaptation assumes labels can directly train or correct the model. We study the intermediate deployment regime between these two extremes. In Active Test-time Adaptation for VLMs ( ), a pretrained VLM may spend an explicit labeling budget on selected test samples, but the prediction for the current sample must be emitted before its label can be used; queried labels are therefore costed, delayed evidence for future samples rather than free correction for the present one. Under this future-only constraint, the central question is how to reuse sparse labels conservatively. We instantiate this setting with ), a training-free plug-in module that routes uncertain samples among base prediction, label query, and memory-based correction using a consistency-and-reliability check over previously queried labels. Across ten cross-domain benchmarks and a domain-transfer benchmark based on ImageNet, ACTOR consistently improves strong VLM-TTA baselines on ResNet and ViT backbones with negligible inference overhead.
PaperID: 2273, Poster
Abstract: Vision-Language-Action (VLA) models have emerged as powerful foundations for robotic manipulation, but their reliance on fixed camera configurations during training makes them brittle to changes in camera count or pose during deployment. To overcome these limitations, we propose VersaCamVLA, a camera-configurable framework that decouples camera-set representation from action learning. VersaCamVLA learns a unified scene-token interface that maps an arbitrary, variable set of posed RGB views into fixed-size latent scene tokens. This is achieved via multi-signal target-view prediction and Wrist-Augmented Pose Sampling (WAPS), which leverages natural wrist-camera motion for free pose diversity. At deployment, a lightweight spatial encoder injects these compact scene tokens into a pretrained base VLA as a supplementary visual condition, requiring no explicit 3D sensing or novel-view rendering. Experiments on RoboTwin, LIBERO, and a real-robot platform demonstrate that VersaCamVLA consistently outperforms prior VLA methods and direct multi-view baselines, maintaining robust performance across varying camera counts and unseen camera poses.
Authors:
Kun Jin, James Harrison, Jiawei Li, Sihan Liu, Jiayi Liu, Randy Linderman, Yuening Li, Arnab Bhadury, Sourabh P Bansod, Liang Liu, Jasper SnoekAbstract: Efficient uncertainty quantification (UQ) is essential for trustworthy large-scale learning. Existing UQ methods for regression tasks mainly operate under the assumption that the conditional label marginal satisfies single-peak parametric models, e.g., Gaussians, where the negative log-likelihood function simplifies to the mean square error. However, such single-peak assumptions fail in regression tasks featuring multi-modal distributions. On the other hand, semi-parametric methods which achieve strong regression performance for multi-modal distributions often lack efficient quantification on their prediction variances. In this work, we extend UQ techniques based on Variational Bayesian Inference (VBI) to two widely used semi-parametric regression models that yield histogram-like reconstructions of the conditional label densities: Quantile Regression (QR) and Classification Restoration (CR). Our approach introduces a unified, distribution-agnostic framework that simultaneously achieves accurate estimation of complex conditional distributions and highly efficient UQ. Theoretically, our method is grounded in novel formulations of QR and CR within the VBI framework, yielding analytic Evidence Lower Bounds (ELBO) to streamline training and a closed-form or analytically approximated predictive density for efficient inference. Empirically, we evaluate our methods on three large-scale regression benchmarks with multi-modal label distributions. Our framework outperforms state-of-the-art multi-modal regression baselines, and even matches predictive performance of computationally expensive ensemble models. Furthermore, by leveraging epistemic uncertainty estimation, our approach enables highly data-efficient active learning strategies.
Abstract: In-context learning for tabular data sets strong predictive standards in observational settings; it however primarily relies on correlational structure, which becomes unreliable under distribution shift or intervention. While established methods to discover causal structure exist, they are often focused on structure identifiability and decoupled from the predictive architectures that could benefit from them. To bridge these perspectives, we study how to simultaneously infer and enforce causal structure in the form of topological variable orderings into tabular prediction. Unlike standard architectures, our model TabOrder uses causal order-constrained attention, basing predictions only on features that precede a target under a learned causal order. Similar to causal discovery methods, TabOrder learns the optimal variable ordering in an unsupervised manner through a likelihood-based objective. We justify this choice under standard functional model classes and also study how sample missingness, a common challenge in tabular data, interacts with causal direction identification. Empirically, we confirm that TabOrder recovers accurate variable orderings while addressing prediction and imputation tasks, as well as gives insight into real-world biological data under intervention.
Abstract: As the number of model parameters increases, parameter-efficient fine-tuning (PEFT) has become the go-to choice for tailoring pre-trained large language models. Low-rank Adaptation (LoRA) uses a low-rank update method to simulate full parameter fine-tuning, which is widely used to reduce resource requirements. However, decreasing the rank encounters challenges with limited representational capacity. Theory suggests that LoRA fine-tuning with rank r converges toward the top r singular values of the pre-trained weight matrix. As the rank increases, more principal singular directions are preserved, which generally improves the model’s performance. However, a larger rank also introduces more trainable parameters, leading to higher computational cost. To overcome this dilemma, we propose SMoA, a Spectrum Modulation Adapter that enlarges the accessible family of spectrum-aware updates under a smaller parameter budget. SMoA partitions the layer into multiple aligned spectral blocks and applies one in-block Hadamard-modulated low-rank branch to each diagonal block, yielding broader coverage of pretrained spectral directions. We provide theoretical analysis and empirical results on multiple tasks. In our experiments, SMoA improves average performance in the current lower-budget setting over LoRA and competitive LoRA-style baselines. Our repository is on https://anonymous.4open.science/r/SMoA-904F/.
Abstract: Discrete diffusion models generate sequences by iteratively denoising samples corrupted by categorical noise, offering an appealing alternative to autoregressive decoding for structured and symbolic generation. However, standard training targets a likelihood-based objective that primarily matches the data distribution and provides no native mechanism for enforcing hard constraints or optimizing non-differentiable properties at inference time. This work addresses this limitation and introduces Search-Augmented Masked Diffusion (SearchDiff), a training-free neurosymbolic inference framework that integrates informed search directly into the reverse denoising process. At each denoising step, the model predictions define a proposal set that is optimized under a user-specified property satisfaction, yielding a modified reverse transition that steers sampling toward probable and feasible solutions. Experiments in biological design and symbolic reasoning illustrate that SearchDiff substantially improves constraint satisfaction and property adherence, while consistently outperforming discrete diffusion and autoregressive baselines.
PaperID: 2278, Poster
Authors: Joonyoung Lim, Younghwan Yoo
Abstract: We investigate the problem of scaling safe reinforcement learning to massively parallel training regimes, where a large number of environments are executed with short rollout horizons. Although this setting improves data throughput and wall-clock efficiency, it creates a mismatch with constrained Markov decision processes, whose safety constraints are defined over full-episode cost returns. We address this mismatch by analyzing staggered environment resets as a phase-mixture distribution over short trajectory segments. This analysis characterizes the bias induced by phase coverage and motivates a phase-aggregated cost estimator that reconstructs episode-level cost estimates from staggered short rollout batches. We further incorporate finite-sample uncertainty and distribution mismatch into a constraint-tightening framework, yielding a scalable Safe RL method compatible with on-policy optimization.Using an MJX-based implementation of Safety Gym, our experiments show that the proposed method achieves performance comparable to non-massively parallel implementations while substantially reducing wall-clock training time
PaperID: 2279, Poster
Abstract: Many organizations fine-tune publicly available pretrained Automatic Speech Recognition (ASR) models and deploy them in black-box settings, assuming limited access provides protection. We show this assumption is fragile: adversarial perturbations crafted on the public base model transfer effectively to fine-tuned target models, severely degrading performance and posing concerns for safety-critical applications. We propose TransferBreaker, a unified fine-tuning framework that suppresses adversarial transfer by integrating Base Adversarial Fine-Tuning, which restricts adversarial training to base-effective perturbations; Latent Jacobian Regularization, which enforces latent-space invariance by suppressing adversarially sensitive directions; and HybridGrad-AFT, which improves robustness against adaptive attacks by interpolating transferable perturbations from base and target gradients. We theoretically justify all components and extensively evaluate TransferBreaker across three languages and four large ASR models, substantially reducing adversarial WER from 92.6 to 27.8.
Abstract: In this paper we introduce token time continuous diffusion (TTCD), a new diffusion language model which \em (a) operates in continuous space, deterministically mapping Gaussian noise to a final token canvas with no further sampling, and crucially \em (b) incorporates a new notion of per-token times, with some tokens proceeding from noise to token at a faster rate than others. Continuous space modeling helps TTCD avoid the parallel sampling of multiple tokens, which is a key source of inaccuracy at high speedups for models that iterate purely in discrete space. The notion of per-token times helps TTCD to better model conditional generation, allows for more sure tokens to proceed at a faster rate, and allows for differentiated inter-token influences during refinement. TTCD outperforms discrete models at high speedups. We train a 160M parameter TTCD model on OpenWebText, and then self-distill it; we find that at high speedups we are comparable in unconditional generation quality, and outperform in conditional generation, several existing models of similar size trained, on the same data, and self-distilled. We achieve similar gains in Sudoku solving as well.
PaperID: 2281, Poster
Authors:
Levin Winter, Cong Li, Hao Sun, Zhendong SuAbstract: Generating inputs for systems that process highly structured data, such as compilers, is difficult due to the dual of syntactic and semantic constraints. Traditional approaches rely on manually crafted rules that require deep domain expertise and are difficult to maintain, while modern machine learning methods have to learn syntax and semantics simultaneously which is challenging. We present Synsema, a novel approach that combines insights from both worlds by reformulating the generation of semantically valid programs as a syntax-guided reinforcement learning task. By guaranteeing syntactic validity through the grammar, the agent needs to only learn the language's semantic rules, sampling from a drastically reduced space. We implement our approach in the context of Java and Rust which both exhibit strict semantic requirements, and train an agent to generate interesting and well-formed programs using fine-grained semantic feedback. Applied to compiler testing, we systematically compare different syntax-guided generation procedures with an unconstrained baseline and show that by restricting the sampling space, Synsema learns semantic rules efficiently. While an average of only 0.03% of programs generated by the baseline are semantically valid, Synsema produces a diverse set of programs with a 7.29% success rate that trigger 2.6x more compiler behaviors. In addition, Synsema discovered 5 previously unknown bugs in production compilers, validating the practical impact of our approach.
Authors:
Sameera Ramasinghe, Shamane Siriwardhana, Thalaiyasingam Ajanthan, Hadi Mohaghegh Dolatabadi, Chamin Hewa Koneputugodage, Gil Avraham, Violetta Shevchenko, James Snewin, Karol Pajak, Harry Xi, Alexander LongAbstract: Decentralized training enables large-model training over low-end GPUs and internet-grade connections, but communication along both data-parallel and pipeline-parallel axes becomes the primary bottleneck. We study post-pretraining adaptation in this setting. We propose an asynchronous two-circuit system: a fast compressed training circuit drives throughput using activation masking for pipeline-parallel (PP) transfer and compressed data-parallel (DP) synchronization, while a slow anchor circuit runs occasional unmasked forward--backward passes off the critical path. Then, we introduce a spectral correction optimizer that uses these delayed anchor priors to denoise masked gradients without blocking the fast stream. Although prior work has found aggressive activation compression unreliable, we show that masking supports post-pretraining adaptation at high compression rates when anchored this way. Pipeline-parallel compression alone yields up to a 9x throughput gain, and combining it with data-parallel compression increases beyond 40x over internet-grade 200Mbps connections, while matching dense uncompressed performance across domain adaptation and continual pretraining.
Abstract: We introduce ECHO, a transformer-based neural operator for large-scale PDE modeling that learns to \emphgenerate full trajectories in compressed spatio-temporal latent spaces. Scaling neural operators to high-resolution systems remains challenging: dense grids are computationally prohibitive, and autoregressive solvers accumulate errors over long horizons. ECHO addresses both issues by jointly compressing space and time and replacing step-by-step prediction with trajectory-level generation. ECHO combines a hierarchical encoder–decoder achieving up to 100× compression, a staged training strategy for high-resolution fidelity, and a latent generative process that models distributions over complete trajectories, improving long-term consistency. This formulation enables a single model to handle forward prediction, inverse problems, and interpolation. Across large-scale 2D and 3D benchmarks, ECHO achieves state-of-the-art performance and scales to million-point simulations with complex dynamics.
PaperID: 2284, Poster
Authors: Ryan L Z., Dung V Nguyen, Minh Hieu Vu, Noelle Y Wong, Lei Zhang, Dipti Srinivasan, Linh D Tran, Tan Nguyen
Abstract: Activation steering has emerged as a powerful approach for controlling large language models (LLMs), with prominent methods such as ActAdd, Directional Ablation, and Angular Steering relying on difference-in-means activations from contrastive prompts across layers. These differences are typically treated as candidate feature directions, later refined into optimal steering vectors or planes. In this work, we reinterpret these candidate directions as gradients of an underlying optimization problem. Building on this perspective, we propose Momentum Steering, a momentum-based framework for activation steering in LLMs. Unlike traditional difference-in-means methods, our framework generates a richer family of candidate directions through momentum updates, enabling more expressive steering. We first introduce a non-causal variant that accumulates difference-in-means signals via momentum, producing enhanced candidate directions. We then develop a causal variant, where future layer statistics are recursively influenced by previously applied momentum directions, explicitly modeling the causal effects of interventions on downstream activations. Momentum Steering is lightweight and modular, making it easily compatible with state-of-the-art steering methods. We empirically demonstrate that Momentum Steering delivers stronger, more robust, and more reliable behavioral control across diverse LLM families and benchmarks.
Abstract: Shapley and Banzhaf interactions capture the complex dynamics inherent in modern machine learning applications. However, current estimators for these higher-order interactions trade off between speed and accuracy. To overcome this limitation, we introduce ProxySHAP. ProxySHAP reconciles the high sample efficiency of tree-based proxy models with a principled path to consistency via residual correction. On a theoretical level, we derive a polynomial-time generalization of interventional TreeSHAP to compute exact interaction indices for tree ensembles, successfully bypassing exponential tree-depth dependencies in prior methods. Furthermore, we formally analyze the residual adjustment strategy, characterizing the specific conditions under which Maximum Sample Reuse (MSR) corrects proxy bias without its variance scaling exponentially with interaction size. Extensive benchmarking demonstrates that ProxySHAP sets a new state-of-the-art standard for approximation quality, including in large-scale applications with thousands of features. By achieving the lowest error in both small- and large-budget regimes, ProxySHAP significantly outperforms the prior best estimators ProxySPEX and KernelSHAP-IQ, while also delivering superior performance on downstream explainability tasks.
PaperID: 2286, Poster
Abstract: Large reasoning models achieve high accuracy through extended chain-of-thought but generate 5--8 more tokens than necessary, applying verbose reasoning uniformly regardless of problem difficulty. We propose Hint Tuning, a data-efficient approach that teaches models to calibrate reasoning depth. Our key insight: the corresponding instruct model serves as an ideal difficulty probe. By testing what the instruct model can solve with varying guidance, we automatically construct training data across three states: No-Hint (direct answer), Sparse-Hint (minimal prefix), and Full-Hint (complete reasoning). This converts the abstract challenge of difficulty labeling into a measurable consistency check between the instruct and reasoning models. With only 1K self-annotated samples, Hint Tuning achieves 24--66% token reduction (31.5% average) across mainstream reasoning models (Qwen3-Thinking, DeepSeek-R1-Distill) at multiple scales (4B--32B) while maintaining competitive accuracy on five benchmarks. Unlike methods requiring massive distillation datasets or expensive RL, we achieve superior efficiency through simple alignment with the instruct model's capabilities.
Abstract: Neural scaling laws for transformer language models predict smooth improvements in pretraining loss with increasing parameters, but downstream capabilities such as in-context learning are known to emerge abruptly past a certain model scale. In this paper, we show that emergent capabilities arise at variable intervals throughout training, with larger models acquiring capabilities earlier on average. We demonstrate that the emergence of capabilities such as pattern completion and indirect object identification corresponds to the abrupt learning of task-relevant attention patterns. To isolate this phenomenon, we train transformer models on synthetic linear map and cellular automata datasets, and we show that the difficulty of learning attention patterns depends on context length and pattern sparsity. Moreover, scaling the number of attention heads improves learning efficiency on our synthetic tasks, while increasing the head dimension yields diminishing returns past a minimum capacity. We additionally investigate architectures with alternative attention mechanisms, showing that MLP-Mixer outperforms a transformer on linear map tasks with complex attention patterns. Our findings provide a mechanistic insight into emergence, showing that downstream capabilities arise abruptly due to the intrinsic difficulty of learning sparse attention patterns in transformer models.
PaperID: 2288, Poster
Authors: Yucong Zhang, Chao Yu
Abstract: In multi-agent reinforcement learning (MARL), agents must not only discover high-reward behaviours but also reliably reproduce coordinated behaviours over time. A fundamental challenge arises from a mismatch between action-level stochastic exploration and temporally extended coordination. To address this issue, we propose a policy-level exploration framework with Deterministic Action Execution (DAE) for MARL, which shifts locally random action exploration to temporally extended policy selection via an option-based exploration mechanism. Specifically, at each time step, each agent selects a policy from a shared policy set\textemdash comprising one exploitation policy and multiple exploration policies\textemdash using a learned option policy, and actions are executed deterministically conditioned on the selected policy. By maintaining deterministic execution under temporally extended policy selection, DAE facilitates temporally coordinated behaviours across multiple time steps. Comprehensive experiments on four multi-agent tasks\textemdash Predator Prey, StarCraft II micro management challenge (SMAC), SMACv2, Google Research Football, and Multi-Agent Coordination benchmark (MACO)\textemdash demonstrate that DAE outperforms mainstream baselines, achieving superior sample efficiency and final learning performance.
PaperID: 2289, Poster
Authors: Diego Aldarondo, Aadhithya Iyer, Daniel Giebisch, Nina Mortensen, Sridhar Pandian Arunachalam, Robert Cochran, Josh Merel
Abstract: To operate safely and effectively around people, humanoid robots must control the forces they exchange with the world. Whereas classical impedance and admittance methods shape contact forces precisely, modern reinforcement-learning controllers typically forgo such regulation in favor of robustness and faithful command tracking. Incorporating impedance-style compliance into a performant humanoid whole-body controller remains an open problem. We present a framework for learned whole-body force regulation that approximates the behavior of impedance controllers using only a PD position-target interface and position encoders. The framework includes three mechanisms for force regulation: passive joint-angle compliance via noisy action perturbations, joint-angle force regulation via perturbation--action blending around a commanded pose, and task-space force regulation via reward-shaped tracking of both a virtual forcefield attractor and a commanded pose. All three share a residual actor--critic recipe, an internal model of proprioception and perturbations, and a policy-blending procedure that combines multiple experts. A 6-DOF body command, optional upper-body pose, and controllable compliance expose the controllers as reusable low-level modules that enable compliant interaction and object manipulation across standing, crouching, and walking. We demonstrate whole-body teleoperation in simulation and on Sprout, a 27-DOF bipedal humanoid.
PaperID: 2290, Poster
Abstract: Hierarchical graph-level tasks that require reasoning over substructures that extend beyond local neighborhoods are posing challenges for standard Graph Neural Networks (GNNs). While graph pooling aims to address this by coarsening graphs into hierarchical representations, existing methods often fail to explicitly preserve the substructures they induce, leading to inconsistent empirical performance. Moreover, the evaluation of such approaches has proved difficult due to the lack of benchmarks with explicit hierarchical structure. We propose Pooled-Substructure-Aware GNNs, an approach to graph pooling that explicitly preserves and processes substructures uncovered by structure-based pooling schemes. We show that this design strengthens pooling in terms of Weisfeiler-Lehman (WL) expressivity. To enable principled evaluation of pooling-based models, we introduce two new synthetic datasets with controllable hierarchical structure and propose a metric to characterize the extent of hierarchical patterns in real-world graph-level benchmarks. On tasks with relevant hierarchical structure, our approach outperforms matched pooling baselines, while remaining competitive on structurally simpler benchmarks.
PaperID: 2291, Poster
Abstract: Efficient attention mechanisms, such as linear and sliding window attention, aim to reduce the computational complexity of Transformers from quadratic to linear. However, their degraded in-context recall often necessitates interleaving full attention layers, leaving the quadratic bottleneck unresolved in long-context processing. To address this, we introduce Orthros attention, which shifts the efficiency paradigm from structural allocation to phase-aware decoupling. As a unified module, it dynamically switches its computation pattern across different phases of sequence processing, enabling linear complexity prefilling and full attention decoding. Our design grants Orthros attention a lower asymptotic complexity compared to vanilla Transformers in prevalent long-context scenarios characterized by heavy prefilling and light decoding. Beyond theoretical formulation and complexity analysis, extensive empirical evaluations demonstrate Orthros's superiority across multiple dimensions. First, we validate its fundamental language modeling capabilities, demonstrating scalability across varying model sizes and strong competitiveness in comprehensive architectural comparisons. Second, Orthros exhibits exceptional conversion feasibility via a data-efficient Transformer-to-Orthros conversion, yielding performance that even surpasses the original Transformer. Finally, Orthros achieves a remarkable 21.9× prefilling speedup over Transformers while avoiding the degradation issues typical of traditional efficient attention mechanisms, evidenced by near-perfect 128K needle-in-a-haystack results and performance on par with strong Transformer baselines on real-world in-context recall tasks. Overall, these advantages in high-performance, flexibility, and efficiency position Orthros as a compelling industrial successor to the Transformer.
PaperID: 2292, Poster
Abstract: Gradient inversion reconstructs training data from shared gradients. However, existing methods target classification tasks and image diffusion, and do not address mixed tabular diffusion, where numerical columns follow Gaussian noising while categorical columns are corrupted by discrete masks. This mixed continuous-discrete training objective changes both the inversion target and the gradient structure: an attacker must recover not only rows, timesteps, and noise, but also the latent mask pattern whose density is tied to the diffusion timestep. We propose Goliath, the first gradient inversion attack designed for tabular diffusion models. Goliath employs a cyclic inversion loop that jointly recovers rows and diffusion latents across Gaussian and mixed diffusion regimes. We introduce a tabular-aware objective function that balances numerical and categorical gradient contributions while improving categorical recovery through enforced consistency between noise levels and mask structures. To accommodate large batches, we aggregate multiple reconstructions across epochs using a row-alignment heuristic. Evaluated across nine tabular datasets, Goliath achieves up to 88.9% per-cell reconstruction accuracy and consistently outperforms gradient-inversion baselines across Gaussian and mixed tabular diffusion. Code is available at https://anonymous.4open.science/r/fl-tab-diffusion-inversion-F4E7/README.md.
PaperID: 2293, Poster
Authors:
Xinhang Zhang, boning zhang, Chengchun Liu, Chunpu Li, Lei Sun, Weifeng Zhang, Limin XiaoAbstract: As sequence lengths continue to grow, distributed attention based on sequence parallel (e.g., Ring-Attention) has become a mainstream solution for large language models (LLMs). Nevertheless, inherent communication dependencies limit computation–communication overlap, reducing GPU utilization. Moreover, in cross-hardware (NUMA or node) architectures, due to imbalanced communication bandwidth, frequent communication synchronization will limit overall performance. Based on these insights, we propose CHARM+, Cross-Hardware Attention with Re-Merge Multistream Mechanism. We reorganize sequential communications into several levels and achieve intra-level parallelism and computation overlap. Furthermore, we decouple cross-hardware communication from intra-hardware communication, eliminating synchronization overhead caused by bandwidth imbalance. Forward and backward algorithms were evaluated on multiple hardware architectures and consistently outperformed state-of-the-art methods. Compared to Ring-Attention, our forward and backward algorithms achieve average speedups of 1.3× and 1.7× on a dual-NUMA PCIe architecture, and 2.8× and 3.9× on a dual-node NVLink architecture, respectively.
PaperID: 2294, Poster
Abstract: Video understanding models achieve impressive 70%+ accuracy on recent benchmarks, but do these truly measure genuine grounded understanding? We systematically investigate this by controlling frame sampling from text-only to real-time across major benchmarks (VideoMME, Video-MMMU, LVBench). Our analysis reveals that 13–46% of questions exploit shortcuts through world knowledge, language priors, or single-frame cues, allowing models to answer correctly without engaging visual evidence. Furthermore, requiring temporal grounding, both correct answers and time localization, exposes a ~40% accuracy gap, showing models frequently answer correctly without locating relevant evidence in videos. To address these limitations, we introduce TGVMME, a temporally grounded benchmark with 1.5k questions requiring both answer correctness and temporal localization. Our benchmark reduces shortcuts to 3.5%, validating that temporal grounding effectively ensures genuine video understanding. Using TGVMME to analyze state-of-the-art reasoning models, we make a surprising discovery: inference-time reasoning provides minimal gains (+2.5% at best), contrasting sharply with substantial gains on existing benchmarks. Analysis on accuracy changes of shortcut questions confirms reasoning primarily exploits shortcuts rather than enhancing genuine understanding. Further experiments reveal input scaling (increasing frames) provides more benefit than output scaling (reasoning) given sparse inputs, highlighting the importance of frame-efficient architectures. Our work demonstrates that grounded evaluation is essential for assessing video understanding capabilities and provides practical tools: validated shortcut taxonomy, filtered splits, and TGVMME for rigorous evaluation.
PaperID: 2295, Poster
Abstract: Softmax attention is typically viewed as effective because it produces a normalized probability distribution over input tokens. In this paper, we challenge this view by showing that a key mechanism behind softmax attention is its implicit control of the Frobenius norm of the attention matrix, which stabilizes training. Motivated by this observation, we study alternative attention activations, focusing on polynomial maps that provide a similar regularization effect without satisfying the usual softmax constraints. Our theoretical analysis shows that certain polynomial activations can replace softmax despite violating positivity, row-wise normalization, and sparsity. Extensive experiments across transformer applications show that these alternatives achieve strong performance, suggesting that the success of softmax attention is not inherently tied to its probabilistic interpretation.
Authors:
Chenchen Liu, Junyi Chen, Lei Li, Lu Chi, Mingzhen Sun, Zhuoying Li, Yi Fu, Ruoyu Guo, Yiheng Wu, Ge Bai, Zehuan YuanAbstract: Multimodal large language models (MLLMs) and diffusion models have each reached remarkable maturity: MLLMs excel at reasoning over heterogeneous multimodal inputs with strong semantic grounding, while diffusion models synthesize images and videos with photorealistic fidelity. We argue that these two families can be unified through a simple division of labor: MLLMs perform semantic planning, while diffusion models render pixels from high-level semantic guidance and low-level visual features. Building on this idea, we propose Bernini, a unified framework for video generation and editing. An MLLM-based planner predicts the target semantic representation directly in the ViT embedding space, and a DiT-based renderer synthesizes pixels conditioned on this plan, augmented by text features and, for editing, source VAE features for detail preservation. Because semantics serve as the interface, the planner and renderer can be trained separately and only lightly co-trained, preserving the pretrained strengths of both components while keeping training efficient. To better handle multiple visual inputs, we introduce Segment-Aware 3D Rotary Positional Embedding (SA-3D RoPE), and further incorporate chain-of-thought reasoning in the planner to better transfer understanding into generation. Bernini achieves state-of-the-art performance across a wide range of video generation and editing benchmarks, with the MLLM's pretrained understanding translating into strong generalization on challenging editing tasks.
Authors:
Ziyao Zeng, Hao Wang, Chen Liu, Youheng Yao, Jingcheng Ni, Tianyu Liu, Xiatao Sun, Fengyu Yang, Chenyu You, Xiaofeng Liu, Daniel Rakita, Ronald Coifman, Yuval Kluger, Zhiwen FanAbstract: Game balancing is a longstanding challenge requiring repeated playtesting, expert intuition, and extensive manual tuning. We introduce RuleSmith, a framework for automated game balancing that couples multi-agent LLM self-play with Bayesian optimization over a parameterized rule space. As a proof of concept, we build CivMini, a simplified civilization-style game with two asymmetric factions governed by 12 tunable parameters. LLM agents read textual rulebooks and game states to generate legal actions, enabling fast evaluation of balance metrics such as win-rate disparity. To search the rule space efficiently, we use Bayesian optimization with acquisition-based adaptive sampling: candidates with high Expected Improvement receive more evaluation games, while exploratory candidates receive fewer. RuleSmith consistently finds near-balanced configurations and produces interpretable parameter adjustments. It provides a scalable alternative to manual playtesting. Code will be available upon acceptance.
Abstract: Deployment-time safety methods for large language models (LLMs) are predominantly supervised and assume access to unsafe training data. Nevertheless, new attacks and harm categories regularly arise, not captured by models trained in such a supervised fashion. An alternative approach is to view this problem through the lens of anomaly detection, namely, to rely solely on modeling safe data and flagging out-of-distribution inputs. However, LLM activations lie in a high-dimensional space, raising concerns about whether anomaly detection is statistically feasible. We show that, under the linear representation hypothesis (LRH), there may indeed be hope. In the LRH concept space, which is typically recovered via a sparse autoencoder (SAE), nearby points share a small common active support. Based on this local sparsity insight, we propose a framework for locally masked SAE-based anomaly detection, and establish a sample complexity bound that scales polynomially in the local sparsity and only \emphlogarithmically in the SAE's ambient dimension. Finally, we empirically validate our framework on various architectures and datasets, including both capability-testing datasets and safety-specific datasets.
Authors:
Mátyás Schubert, Theofanis Aslanidis, Tom Claassen, Sara MagliacaneAbstract: Expert background knowledge is often available in practical applications of causal discovery. Such constraints on the true causal graph can help causal discovery in terms of identifiability of causal effects and accuracy of the learned structure, but also in reducing the space of candidate causal graphs. As causal discovery can become computationally expensive for large number of variables, it is crucial to utilize background knowledge effectively \emphduring the causal discovery process. However, most current methods only use background knowledge in a postprocessing step after causal discovery to refine the learned graph. In this work, we develop a framework for utilizing background knowledge during the causal discovery process, focusing especially on scalable causal discovery methods that recover only a subset of the whole graph. We implement our framework for multiple algorithms and empirically show that utilizing background knowledge can both reduce computational requirements and increase the quality of the learned structures.
PaperID: 2300, Poster
Abstract: Acquiring high-quality labeled data and engineering effective learning pipelines are two critical bottlenecks to applying machine learning systems in new domains. Although recent MLE agents have shown promise in automating pipeline design, their behavior has been studied almost exclusively in fully supervised settings. In many domains, however, labeled data is expensive while unlabeled data is abundant. In this work, we study agents in the semi-supervised learning (SSL) regime by framing SSL as a program search problem and exploring how agents navigate this space under limited supervision. Our analysis focuses on three questions: (1) whether agents can effectively leverage unlabeled data, (2) how they leverage it and what strategies they discover, and (3) whether they can adapt these strategies to specific domains. We find that agents can leverage and synthesize SSL strategies spanning decades of SSL research, achieving performance comparable to, and in some cases outperforming, established human-engineered SSL methods. We further validate these observations on real-world ecological datasets, highlighting both the challenges and opportunities of transferring agent-discovered strategies beyond controlled benchmarks.
Abstract: Causal queries are often only partially identifiable from observational data, and experiments that could tighten the resulting bounds are typically costly. We study the problem of selecting, prior to observing experimental outcomes, a cost-constrained subset of experiments that maximally tightens bounds on a target query. We formalize this as the max-potency problem, where epistemic potency measures the worst-case reduction in bound width guaranteed by an experiment, and show that this problem is NP-hard via a reduction from 0-1 knapsack. Building on the polynomial-programming framework of Duarte et al. (2023), we give a general procedure for evaluating epistemic potency in discrete settings. To control the super-exponential search space, we introduce two graphical pruning criteria that depend only on the causal graph and the query: a novel path-interception rule that exploits district structure to certify zero potency in linear time, and an identifiability check based on the ID algorithm. On Erdős–Rényi random graphs and 11 bnlearn benchmark networks, the two criteria together prune 50-88% of candidate experiments on average without solving a single polynomial program. For the general subset search, we show that ID-pruned experiments are combinatorially inert, yielding a super-exponential reduction in the number of subsets evaluated. We close with an end-to-end demonstration on observational NHANES data, selecting optimal experiments for estimating the effect of physical activity on diabetes.
PaperID: 2302, Poster
Authors: Prabin Kumar Rath, Omkar Patil, Nakul Gopalan
Abstract: Behavior cloning (BC) in non-Markovian environments is a challenging problem because policies have to reason over contextual information over long horizons. Existing policy architectures rely on recurrent or attention-based mechanisms to capture long-term dependencies. However, recurrent models suffer from hidden-state collapse and gradient instability under backpropagation through time, while attention-based models are fundamentally limited by context length. To address these issues, we propose Keyframe Mnemonics, a novel self-supervised method that discovers a set of information-critical observations (mnemonics) by learning an objective from randomly sampled past observations and using it as a reward for keyframe selection. We then train a BC policy that conditions on the discovered keyframes to model the action distribution. Under near-optimal keyframe selection, our formulation provides context retention guarantees over an infinite horizon, while maintaining a small set of decision-relevant keyframes in the policy's working memory. We evaluate our method on synthetic memory domains, where mnemonic-conditioned BC policies achieve 100% success rates (SR) and generalize to horizons orders of magnitude beyond training without performance degradation. Additionally, we evaluate on memory intensive robot manipulation benchmarks, where we observe an absolute 15-30% improvement in SR over strong memory-augmented baselines on 23 tasks. Code and videos are available at https://keyframe-mnemonics.github.io.
Abstract: Vision-Language Latent Diffusion Models (LDMs) provide powerful generative priors for inverse problems. However, existing LDM-based inverse solvers typically require a large number of neural function evaluations (NFEs) and backpropagation through large pretrained components, leading to substantial computational cost and, in some cases, degraded reconstruction quality. We propose a unified Euclidean-Wasserstein-2 gradient-flow framework that jointly performs posterior sampling and prompt optimization in the latent space through a single flow that aligns the prior and posterior with the observed data. Combined with few-step latent text-to-image models, this formulation enables low-NFE inference without backpropagation through autoencoders. Experiments across several canonical imaging inverse problems show that our method achieves state-of-the-art performance with significantly reduced computational cost.
PaperID: 2304, Poster
Abstract: Post-training for reasoning models typically combines supervised fine-tuning with reinforcement learning from verifiable rewards, most commonly with GRPO. However, this algorithm suffers from sparse rewards, limited exploration, and mode collapse. Building upon recent works on self-distillation, we propose Feedback Distillation, a training method where the model is trained to match, at the token level, its own distribution conditioned on privileged feedback produced by a language model. Feedback Distillation offers token-level supervision and can inject external knowledge. Evaluating our method for Lean 4 theorem-proving, we find that Feedback Distillation maintains greater diversity in generated trajectories than GRPO, yielding higher policy entropy and better pass@k scaling. The two methods are complementary: initializing GRPO from a Feedback Distillation checkpoint outperforms either method alone. All in all, our results suggest a promising avenue to improve post-training for complex reasoning.
PaperID: 2305, Poster
Abstract: In this paper, we consider mixed-motive games (MMGs) which model various real-world multi-agent scenarios where self-interested agents must both cooperate and compete. We identify an issue with existing gifting methods, where agents become overly dependent on incoming gifts and fail to learn reward-yielding behaviors, a phenomenon we call gift over-reliance. To address this, we propose a refusal-augmented gifting framework that balances cooperation and competition in MMGs. Our approach introduces an adaptive refusal mechanism that allows agents to selectively reject a portion of received gifts, thereby mitigating gift over-reliance and preventing the agents from being exploited by other agents. This mechanism is gifting-method-agnostic, serving as a plug-in for existing schemes while enabling decentralized learning. We demonstrate that across diverse MMG environments and gifting schemes, our mechanism consistently reduces gift over-reliance, prevents lock-in to sacrificing policies, and improves \alpha-fairness without necessitating major changes to existing training pipelines.
Abstract: Moving an object in a single image requires geometry-consistent spatial rearrangement, including handling occlusions, revealing previously unseen regions, and maintaining coherent shadows and reflections. Existing approaches are not well suited to this setting and often fail to preserve such scene-level consistency. We address this problem by introducing a geometry-aware object motion method that operates directly on the internal representations of diffusion transformers. Our key insight is that rotary positional embeddings (RoPE) define a structured spatial field that can be explicitly manipulated to induce controlled motion. We extend 2D RoPE into a depth-aware formulation that encodes 3D spatial structure, enabling consistent object displacement and scene-aware updates. Our model is trained using synthetic data combined with a small set of real images via parameter-efficient fine-tuning. Despite minimal real supervision, it preserves object identity under large spatial displacements, generates plausible content in newly revealed regions, and consistently updates scene-dependent effects such as shadows and illumination. Experimental results demonstrate the superiority of our method.
Authors:
Chao Xu, Maohua Li, Qirui Li, Yixuan Xu, Yanke Zhou, Yunhe LI, Cuifeng Shen, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu, Shao-Qun ZhangAbstract: Diffusion Transformers (DiTs) have become the de facto backbone of modern visual generation, and nearly every major axis of their design --- tokenization, attention, conditioning, objectives, and latent autoencoders --- has been thoroughly revisited. The residual stream that governs how information accumulates across layers, however, has been directly inherited from the original Transformer. In this paper, we present a systematic analysis of cross-layer information flow in DiT jointly along depth and denoising timestep, and identify three concrete symptoms of traditional residual addition, i.e., monotonic forward magnitude inflation, sharp backward gradient decay, and pronounced block-wise redundancy. Motivated by this diagnosis, we propose Diffusion-Adaptive Routing (\textscDAR), a drop-in residual replacement that performs \emphlearnable, timestep-adaptive, and non-incremental aggregation over the history of sublayer outputs. Moreover, the proposed DAR is compatible with many modern Transformer enhancement methods, such as REPA. On ImageNet 256×256, \textscDAR improves SiT-XL/2 by 2.11 FID (7.56 vs.\ 9.67) and matches the baseline's converged quality in 8.75× fewer training iterations. Stacked on top of REPA, it yields a 2× training acceleration in the early stage, revealing cross-layer information routing as a hitherto-overlooked axis of progress in diffusion modeling, one that operates orthogonally to existing representation-alignment objectives. Beyond the pretrain tasks, \textscDAR can be further applied in large-scale T2I models fine-tuning stage and preserves high-frequency details during Distribution Matching Distillation.
PaperID: 2308, Poster
Authors: Manu Agrawal, Sumit Mukherjee
Abstract: Specialized agents collaborate by transferring memory: retrieved records, generated summaries, tool traces, and long-running state. In regulated domains this is a governance problem, not a retrieval problem: the receiver's task matters, but so does whether it was authorized to receive each item and whether the handoff exposed more than necessary. We formalize minimum-necessary context selection as a constrained submodular subset-selection problem with an additive exposure cost and a distributional task-success risk bound. Our algorithm, ScopeSelect, first hard-masks items whose layer or scope the receiver contract does not admit, then runs a cost-benefit greedy selector against a monotone submodular utility, attaining a \\frac12(1-1/e) approximation bound under a budget, and finally wraps threshold selection in Conformal Risk Control, giving a finite-sample distribution-free guarantee \\mathbbE[L(S_\\hat\\gamma)] \\le \\alpha on exchangeable calibration data. We evaluate on HealthHandoff-Bench, a synthetic benchmark with 21,600 healthcare-agent handoff tasks across four workflows, and Synthea-HH, a realism tier ingested from 340 Synthea-generated FHIR bundles. A three-source blinded label audit compares benchmark labels, model-generated labels, and clinical-informatics expert review, with high agreement for sensitivity labels and policy-boundary disagreements for necessity and authorization. Against authorized TF-IDF, embedding retrieval, a cross-encoder reranker, and a calibrated LLM-as-selector, ScopeSelect is the only method meeting \\alpha = 0.15 at competitive task success. The hard-mask safety property holds under 30% label noise, 100 adversarial attacks, and receiver-role shift; a live-LLM run with Claude Sonnet 4.5 exhibits a calibration-transfer gap that motivates per-workflow recalibration. We release the benchmark, Synthea-HH pipeline, CRC calibration harness, and a signed-handoff prototype.
PaperID: 2309, Poster
Abstract: Generative modeling over permutations is challenging because permutations are discrete structured objects, whereas many powerful generative samplers, including score-based diffusion models, are formulated in continuous spaces. Existing diffusion-based approaches typically bridge this mismatch either by relaxing permutations into continuous permutation surrogates or by designing discrete diffusion processes with specialized transition kernels. We propose a geometric alternative based on the coordinate ordering of centered vectors. Sorting a centered vector maps it to a permutation, and the sorting operation partitions the space into permutation-indexed regions whose facet adjacencies correspond to adjacent swaps between permutations. In this view, the learned score controls how probability mass is transported across sorting regions, thereby inducing transitions between neighboring permutations through their shared boundaries. The resulting sampler remains close to standard score-based diffusion: train in the centered Euclidean space, run the reverse diffusion, and sort once at the end. Experiments across a variety of standard benchmark tasks such as jigsaw puzzle reconstruction (structured prediction), sorting 4-digit MNIST numbers (ordering), the traveling salesman problem (combinatorial optimization), and preference-based ranking show state-of-the-art performance while remaining computationally efficient, especially compared with discrete generative methods.
Abstract: Understanding the internal circuits that language models use to solve tasks remains a central challenge in mechanistic interpretability. A crucial part of finding circuits is understanding why each attention head attends where it does. To this end, we introduce , i.e., contents of low dimensional subspaces that cause attention on a token pair. ACC++ extracts circuits from a , without replacement models or patching. Circuits identified by ACC++ consist of components that are causal for the model's attention decisions, together with the low-dimensional signals used to communicate between them. Here, we first detail the conceptual advances that ACC++ makes over previous work. We then show that across multiple models, a substantial portion of ACC++ signals are : many signals admit a short natural-language description. We next present a number of new insights into model behavior obtained via ACC++. First, we use ACC++'s interpretable circuits to characterize the sensitivity of indirect object identification (IOI) circuits to prompt structure. We find that prompt-specific circuits form well-defined clusters, and across clusters, heads receive systematically different signals corresponding to distinct mechanisms for identifying the IO name. Next, in multilingual IOI, ACC++ circuits show that while model are often language-specific. In a four-language IOI case study, cross-language circuit distances are consistent with linguistic relatedness. Together, these results show that ACC++ can shed light on a broad spectrum of model behaviors.
PaperID: 2311, Poster
Authors: Zhisong Xu, Takeshi Oishi
Abstract: Streaming Visual Geometry Transformers such as StreamVGGT enable strong online 3D perception, but their KV-cache grows unbounded over long streams, limiting practical deployment. We study bounded-memory streaming geometry from the perspective of memory organization: unlike language modeling, where useful information can often be compressed at token level, geometry-driven inference relies on coherent and mutually compatible observations across views. Under fixed memory budgets, retaining history as isolated entries can progressively fragment the geometric context needed for stable long-horizon matching and fusion. We therefore propose FrameVGGT, a bounded-memory framework that maintains a fixed-capacity set of complementary memory units for streaming geometry. In our implementation, each unit is instantiated as a frame-wise KV segment summarized by a compact key-space prototype, together with a sparse anchor tier for persistent long-range references. Across long-sequence 3D reconstruction, video depth estimation, and camera pose estimation, FrameVGGT achieves favorable accuracy--memory trade-offs under bounded budgets while maintaining more stable geometry over long streams.
Authors: Maryna Kapitonova, Tonio Ball
Abstract: We present a neuro-inspired framework for embodied planning and control. Building on three principles that enable fast and highly effective goal-directed behavior in the mammalian brain — paired forward/inverse internal models, open-loop multi-step motor commands, and sequential organization of action — our Inverter framework combines analytic and trained components, the latter trained end-to-end through Inverse Learning (IL), a paradigm we formalize and delineate from supervised, reinforcement, and imitation learning. IL bridges RL-style amortization, which runs in a single forward pass but emits one action at a time, and optimal-control-style sequence reasoning, which plans whole trajectories but with iterative test-time computation. Single Inverters or hierarchical n=2 Inverter stacks match or improve over comparable offline-RL and diffusion-planner baselines on all 3 `maze2d-v1` and 6 `antmaze-v2` D4RL variants by an average of +24.2% (range -1.9% to +78.2%), at one-to-two orders of magnitude less inference compute. As one mechanism behind this effectiveness, we show that Inverters can learn control policies closer to the analytic optimum than the data-generating policy itself. We also identify a failure mode of IL: forward-model hacking under narrow training-data coverage, which we mitigate by using random training data with broader coverage. As an application example, a Pulse Inverter synthesizes arbitrary single-qubit quantum gates with fidelity matching the standard iterative numerical baseline (GRAPE), at more than 1000× lower per-gate compute. We propose deeper hierarchical, probabilistic, and latent extensions of the Inverter framework as a differentiable world-interface, especially for latency- and compute-critical embodied AI.
Authors: Jongmin Lee, Ernest Ryu
Abstract: While there is an extensive body of research analyzing policy gradient methods for discounted cumulative-reward MDPs, prior work on policy gradient methods for average-reward MDPs has been limited, with most existing results restricted to ergodic or unichain settings. In this work, we first establish a policy gradient theorem for average-reward multichain MDPs based on the invariance of the classification of recurrent and transient states. Building on this foundation, we develop refined analyses and obtain a collection of convergence and sample-complexity results that advance the understanding of this setting. In particular, we show that the proposed \alpha-clipped policy mirror ascent algorithm attains an \epsilon-optimal policy with respect to positive policies.
PaperID: 2314, Poster
Authors: Ryan Fayyazi, Kyle Daruwalla, Mitra Javadzadeh
Abstract: Backpropagation remains the dominant method for training deep neural networks, but its reliance on non-local computations has motivated the search for biologically plausible alternatives. Two prominent frameworks, predictive coding (PC) and deep feedback control (DFC), offer distinct approaches to local learning, yet their relationship remains unclear. Here, we show that DFC arises as a local approximation to predictive coding dynamics near fixed points, with exact equivalence in shallow linear networks and controlled deviations arising with depth, nonlinearity, and distance from equilibrium. Building on this connection, we introduce Approx PC, a practical algorithm that replaces layer-wise error propagation in PC with a centralized feedback controller operating over a sequence of intermediate subgoals. Empirically, Approx PC recovers solutions consistent with predictive coding, while converging faster and achieving improved accuracy in our experiments. Our results unify energy-based and control-theoretic perspectives on learning and suggest a computationally efficient approach to training predictive coding networks through control-inspired dynamics.
Abstract: Generated videos often lose geometric consistency under camera motion: structures warp under viewpoint change, parallax becomes inconsistent, and smooth camera movement may collapse into shot-like transitions. We argue that this failure is largely a data problem. In web-scale video corpora, most visible motion comes from dynamic objects captured by static or weakly moving cameras, leaving only sparse supervision for viewpoint-induced appearance changes. To address this, we propose Video Geometry Corrector (VGC), a post-hoc video-to-video model trained on synthetic paired supervision. Starting from real videos with diverse camera motion and pose annotations, we construct VG-Pairs with a synthetic paired-supervision data generation pipeline. Each video clip is treated as a geometry-consistent target and paired with a distorted counterpart synthesized by a pretrained image-to-video inpainting model using trajectory-adaptive keyframes. VGC learns to map the distorted input to the geometry-consistent target without modifying the source generator, and uses optical flow as an auxiliary cue to help disentangle camera motion from object motion. Experiments on pose-conditioned and prompt-driven benchmarks show that VGC improves geometric consistency across multiple source generators, and scaling to a larger backbone further extends the framework toward more general video refinement.
PaperID: 2316, Poster
Abstract: Neural methods for graph-structured optimization have achieved strong performance within individual problem classes, but are typically trained for fixed formulations, limiting reuse across related problem variants. Recent progress in foundation models has motivated shared neural backbones trained across multiple graph-structured optimization tasks; however, existing approaches primarily share instance-level representations while treating feasibility constraints implicitly. A key source of shared structure across constrained optimization problems is the constraint set itself, which often exhibits substantial overlap across tasks. We propose CONSTRAINER, a constraint-conditioned framework for graph-structured optimization that treats constraints as structured inputs. Constraint specifications are parsed into operator graphs and encoded using a graph neural network to produce invariant constraint embeddings, which condition a shared backbone model for solution prediction. We evaluate our approach on power grid optimization and mixed-integer linear programming problems, where it shows improved solution optimality and enables efficient transfer learning and generalization compared to baseline methods.
Abstract: A common recipe to improve diffusion models at test-time so that samples score highly against a user-specified reward is to introduce the gradient of the reward into the dynamics of the diffusion itself. This procedure is often ill posed, as user-specified rewards are usually only well defined on the data distribution at the end of generation. While common workarounds to this problem are to use an approximate denoiser to estimate what a sample would have been at the end of generation, we propose a simple solution to this problem by working directly with a flow map. By exploiting a relationship between the flow map and velocity field governing the instantaneous transport, we construct an algorithm, Flow Map Trajectory Tilting (FMTT), which leverages a fast yet precise flow map look-ahead and provably performs better ascent on the reward than standard test-time methods involving the gradient of the reward. The approach can be used to either perform exact sampling via importance weighting or principled search that identifies local maximizers of the reward-tilted distribution. We demonstrate the efficacy of our approach against other look-ahead techniques, and show how the flow map enables engagement with complicated reward functions that make possible new forms of image editing, e.g. by interfacing with vision language models.
Abstract: Agentic data science (ADS) systems are rapidly improving their capability to autonomously analyze, fit, and interpret data, potentially moving towards a future where agents conduct the vast majority of data-science work. However, current ADS systems use statistical tools designed to be interpretable by humans, rather than interpretable by agents. To address this, we introduce Agentic-imodels, an agentic autoresearch loop that evolves data-science tools designed to be interpretable by agents. Specifically, it develops a library of scikit-learn-compatible regressors for tabular data that are optimized for both predictive performance and a novel LLM-based interpretability metric. The metric measures a suite of LLM-graded tests that probe whether a model’s string representation is “simulatable” by an LLM, i.e. whether the LLM can answer questions about the model’s behavior by reading its string output alone. We find that the evolved models jointly improve predictive performance and interpretability, generalizing to new datasets and new interpretability tests. Furthermore, these evolved models improve downstream end-to-end ADS, increasing performance for Copilot CLI, Claude Code, and Codex on the BLADE benchmark by up to 73%.
PaperID: 2319, Poster
Authors: Zedong WANG, Yucheng Wang, Dan Xu
Abstract: Multi-task scene understanding models typically require modifications to the vision encoder or complex decoders for cross-task interaction, which limits their flexibility and compatibility with off-the-shelf vision foundation models (VFMs). This paper introduces Virtual Task Prompting (VTP), which frames multi-task prompting as retrieving task-adaptive subsets from a shared virtual state, where both the state and accessing process itself are optimizable in an end-to-end manner. Concretely, VTP maintains a set of virtual tasks alongside the model, and at each layer, learnable routing matrices read an actual task composition as input for interaction with image tokens, then write the processed results back, forming explicit learning pathways that capture the intricacies among tasks before and after the computation of each encoder block. Since VTP keeps the backbone clean, it integrates seamlessly with recent VFMs like DINOv3, advancing the state of the art on Pascal-Context and NYUD-v2. Stress test on Taskonomy with 13 heterogeneous tasks further shows VTP's robustness in extreme settings. Ablations and analysis confirm that VTP's gains stem mainly from its prompting design rather than merely increased capacity.
Authors: Kazuhiro Hiraki, Shinichi Ishihara, Takumi Kongo, Junnosuke Shino
Abstract: We propose ESENSC, a computationally efficient and theoretically grounded alternative to SHAP. ESENSC is a polynomial-time attribution rule with a deterministic, closed-form representation that satisfies the null-player property and admits an axiomatic characterization based on efficiency, a restricted form of differential marginality, and explicit computational constraints. Empirically, ESENSC closely replicates the attribution patterns of exact SHAP by preserving its essential structural properties, while maintaining a fraction of the cost. Across neural network and XGBoost models, it achieves lower deviation and higher rank agreement than widely used sampling-based SHAP approximations. While the computational cost of exact SHAP grows exponentially with the number of features, ESENSC scales linearly, enabling reliable feature attribution in regimes where exact SHAP is intractable. These results demonstrate that ESENSC achieves near-SHAP attribution accuracy while reducing computational complexity from exponential to linear in the number of features, providing a practical and theoretically rigorous alternative for feature attribution in high-dimensional applications.
PaperID: 2321, Poster
Abstract: Ambiguous inverse problems often admit multiple clean signals that are consistent with the same observation. In such settings, a sampler should not only produce plausible or measurement-consistent reconstructions, but also allocate probability mass correctly across the alternatives that remain unresolved. We formalize this requirement as \emphposterior alternative calibration and introduce mode coverage error (MCE), a finite-alternative diagnostic that measures calibration after pushing the posterior and sampler distributions onto a benchmark-specified alternative label space. This perspective reveals a basic failure mode of point-centered inverse solvers: measurement consistency can be preserved by averaging across alternatives, even when the resulting point lies outside the nonconvex set of valid posterior explanations. We prove that local point-denoising updates can therefore suffer barycentric collapse under weak branch identification, while calibrated local branch probabilities are sufficient to control MCE. Across controlled Gaussian, imbalanced, nullspace, and template-inpainting diagnostics, MCE exposes posterior mass-allocation failures that are missed by residual error, recall, diversity, and confidence. A semantic-label inpainting benchmark with DDIM, DPS, DDNM, RePaint, and clustering-based samplers further shows that practical samplers can recover all specified alternatives while assigning different branch probabilities. These results position MCE as a targeted evaluation axis for finite-alternative ambiguity benchmarks, complementary to reconstruction fidelity and sample quality.
Abstract: We study post-hoc Learning to Defer (L2D) through the lens of ideal distributions: divergence-regularized reweightings of the data distribution under which a model attains low loss. We define deferral via the density-ratio between a model's and an expert's ideals. Using the reduction from density-ratio estimation to class-probability estimation, we derive the DR CPE losses for post-hoc L2D scorers. Deferral decisions are then made by thresholding the scorer, allowing deferral rates to be adjusted without retraining. For KL-based ideal distributions, our deferral rules recovers Chow's rule under the original distribution and a connection to an expert-tilted Bayes posterior—which incorporates the expert's performance—depending on if the ideal distributions are joint or marginal distributions. Experimentally, our approach is competitive compared to common baselines and more robust across dataset settings. More broadly, our results cast post-hoc L2D as density-ratio learning between ideal distributions, bridging Chow-style rules, expert comparison, and elucidating connections to related learning settings including anomaly detection.
PaperID: 2323, Poster
Authors: Mina Kim, Guanghui Lan, Benoit Montreuil
Abstract: Building hydrogen station networks, electric vehicle (EV) charging grids, or vaccine distribution systems requires coordinating two separately operated graphs - demand-side facilities and the supply-side networks they depend on - where each operator's payoff depends on the other's choices. We formalize this as the Co-Evolving Graph Game (CEGG), a setting that sits between facility location, interdependent-network analysis, and graph-based multi-agent reinforcement learning (MARL) but is not jointly addressed by any of them. We propose Cross-Graph Attention with Policy Mirror Descent (CGA-PMD), pairing cross-graph attention - which lets each agent read only the partner-graph nodes structurally coupled to its own - with a policy mirror descent update for stable training. Across three reward domains and a range of practical network scales, we find that the structure of the coupling selects which method wins: at small scale most methods are competitive, but as scale grows or coupling becomes irregular, only CGA-PMD remains productive while own-only, unmasked-attention, and Proximal Policy Optimization (PPO)-based variants all fail in distinct ways. Our theoretical analysis explains why: PMD's Kullback-Leibler (KL) regularization admits an iteration-uniform cumulative drift bound while PPO's clip update does not, predicting CGA-PMD's stability in the hard regime where PPO-based methods collapse.
PaperID: 2324, Poster
Abstract: This paper studies bilevel optimization on Riemannian manifolds where the upper-level objective is nonconvex and the lower-level problem satisfies a Riemannian Polyak - Łojasiewicz condition rather than geodesic strong convexity. The classical hypergradient formula then breaks down, since the lower-level Hessian may be singular or indefinite away from the minimizer. We address this with an intrinsic regularized tangent-space formulation based on spectral clipping, and develop a Riemannian bilevel algorithm that avoids Hessian-inverse solves and inner-loop differentiation. We establish an O(1/T) rate on the averaged squared Riemannian gradient mapping and O(\varepsilon^-1) iteration complexity, and extend the method to the stochastic finite-sum setting. Experiments on Stiefel, Grassmannian, and Poincar\'e-ball problems with rank-deficient Hessians (BCI~IV-2a, UCI Superconductivity) and Stiefel-constrained meta-learning on MiniImageNet show that the proposed method is the only one whose linear-system error stays at machine precision and whose Riemannian gradient norm decreases monotonically, while Hessian-inversion and Neumann-series baselines fail to descend; the stochastic complexity exhibits the predicted \Theta(1/B) decay.
PaperID: 2325, Poster
Abstract: Two-sided matching platforms rely on preferences from both sides, yet participants can evaluate only a small fraction of potential partners. In practice, they use low-cost pre-match screening, e.g., interviews, profile views, or trial tasks, to form noisy impressions before committing to applications and offers. We study bandit learning in matching markets with interviews, modeling these interactions as queried hints [Bhaskara et al., 2022] that reveal partial preference information to both sides while constraining subsequent applications. Our framework also allows firm-side uncertainty: firms, like agents, learn their preferences and may make early hiring mistakes. To address this, we introduce strategic deferral, a firm-side action that permits temporary vacancy, corrects premature commitments, and enables decentralized learning under coarse anonymous feedback. We design algorithms for centralized and decentralized markets and show that a constant number of interviews per round suffices for horizon-independent regret, improving over the O(\log T) guarantees known without interviews. Our bounds are near-optimal: the centralized guarantee is within a factor m of an information-theoretic lower bound, while decentralized algorithms match it up to polynomial factors in structured markets and remain horizon-independent in general markets.
Authors: Hao Lin, Kunyang Lv, Xu Jiang, Jingqi Tian, Zhongjing Du, Jiayu Ding, Qiaoman Zhang, Hongbo Jin
Abstract: Training VideoLLMs for complex reasoning remains challenging due to sparse sequence-level rewards and the lack of fine-grained credit assignment over long, temporally grounded reasoning trajectories. While reinforcement learning with verifiable rewards (RLVR) provides reliable supervision, it fails to capture token-level contributions, leading to inefficient learning. Conversely, existing self-distillation methods offer dense supervision but lack structure and diagnostic specificity, and often interact unstably with reinforcement learning. In this work, we propose VISD, a structured self-distillation framework that introduces diagnostically meaningful privileged information for video reasoning. VISD employs a video-aware judge model to decompose reasoning quality into multiple dimensions, including answer correctness, logical consistency, and spatio-temporal grounding, and uses this structured feedback to guide a teacher policy for token-level supervision. To stably integrate dense supervision with RL, we introduce a direction–magnitude decoupling mechanism, where environment rewards determine update direction, while structured privileged signals modulate token-level update magnitudes. This design enables semantically aligned and fine-grained credit assignment, improving both reasoning faithfulness and training efficiency. Additionally, VISD incorporates independent advantage estimation, curriculum scheduling, and EMA-based teacher stabilization to support robust optimization over long video sequences. Experiments on various benchmarks demonstrate that VISD consistently outperforms strong baselines across diverse video reasoning tasks, achieving gains in accuracy, grounding quality, and interpretability. Notably, VISD significantly accelerates convergence, highlighting the effectiveness of structured self supervision in improving both performance and sample efficiency for VideoLLMs.
Abstract: Access to modern generative systems is often restricted to querying an API (the ``black-box" setting) and many properties of the system are unknown to the user at inference time. While recent work has shown that low-dimensional representations of models based on the relationship between their embedded responses to a set of queries are useful for inferring model-level properties, the quality of these representations is highly sensitive to the query set. We introduce the discriminative factorization to distinguish between high- and low-quality query sets in the context of black-box model-level classification. Under this framework, the probability of chance-level classification decays exponentially in the query budget. On three auditing tasks, estimated factorization parameters predict the empirical performance decay rate. We conclude by showing that query sets selected using the estimated discriminative field reproduce the empirical ordering of oracle query sets.
Abstract: Learning from positive and unlabeled data (PU learning) is a weakly supervised variant of binary classification in which the learner receives labels only for (some) positively labeled instances, while all other examples remain unlabeled. Motivated by applications such as advertising and anomaly detection, we study PU learning, where the learner adaptively queries instances from an unlabeled pool, but a label is revealed only when the queried instance is positive and an independent coin flip succeeds; otherwise the learner receives no information. This paper provides the first theoretical analysis of the label complexity of active PU learning.
Abstract: We investigate whether neuron populations that shape the internal organization of neural networks evolve predictably with scale, extending scaling laws beyond macroscopic observables such as loss. To probe this question, we study , a previously characterized class of neurons whose activation patterns are similar across independently trained models (Dravid et al., 2023). In separate analyses of language models up to 30B parameters and vision models up to 5B parameters, we observe that the population of Rosetta Neurons follows a sublinear power law in model size, growing in absolute number but occupying a shrinking fraction of the total neuron count. We further observe a : Rosetta Neurons become more selective and increasingly monosemantic with scale, separating from a growing non-Rosetta population that remains less selective. An analytical model balancing feature utility against limited neuron capacity explains the sublinear power-law scaling and this polarization effect. Finally, we find that Rosetta Neurons become more domain-specialized with scale and illustrate their selectivity through a targeted data-filtering case study for continued pretraining. Our results point to a scaling law for shared neuron-level structure, linking model size to systematic changes in neuron universality, selectivity, and specialization.
PaperID: 2330, Poster
Abstract: Performance predictors are pivotal in accelerating evaluation in NAS. Cross-dataset predictors are trained on (dataset, architecture, performance) triplets to efficiently estimate architectural performance for unseen datasets. However, existing cross-dataset predictors rely on a naive concatenation of dataset and architecture features. This fusion ignores the architecture-aware nature, where the same dataset manifests as distinct representations under different architectures, thereby limiting generalization. To tackle this issue, we propose ACDP, an Architecture-aware Cross-Dataset Performance Predictor that dynamically adjusts dataset feature extraction based on specific architecture. Initially, ACDP develops an elastic mechanism to extract raw dataset traits, providing a configurable lever for the trade-off between efficiency and accuracy. Furthermore, ACDP leverages a hypernetwork to determine dataset feature extractor parameters conditioned on architectural features, transforming raw dataset traits into tailored representations. ACDP excels in both specialization and generalization on six unseen datasets. Notably, in NAS-Bench-201, ACDP achieves 0.758/0.741 Kendall's Tau on CIFAR-10/100, surpassing single-dataset SOTAs trained with over 100 target samples. In MobileNetV3, ACDP discovers architectures outperforming cross-dataset SOTAs, notably gaining +2.04% accuracy on Aircraft. Code: https://anonymous.4open.science/r/ACDP/
Abstract: Model merging techniques, which aggregate independently finetuned models into one to combine their capabilities, have become a topic of significant interest in recent years, with a broad array of methods having been proposed to tackle this problem. Simultaneously, an emerging trend in distributed learning has been the use of methods such as local SGD and DiLoCo, which greatly reduce communication costs by periodically aggregating the independently trained local models. However, these communication-efficient methods have been shown to degrade in performance relative to the FLOP-matched data-parallel gold standard as the number of independent local models grows and as the number of local training steps before global communication is increased. In this work, we draw an explicit analogy between the pseudo-gradient aggregation step in local SGD/DiLoCo and task arithmetic-based model merging, establishing a straightforward way to utilize merging methods in the context of distributed optimization. We then evaluate multiple state-of-the-art model merging methods in this setting and identify one method in particular, Iso-C, as a promising approach for improving DiLoCo. We find that DiLoCo SGD with Iso-C aggregation outperforms not only simple pseudo-gradient averaging but even the momentum-based DiLoCo, despite lacking a momentum mechanism itself. Building on this finding, we propose IsoLoCo, which adapts Iso-C for distributed training by equipping it with Nesterov momentum. Our empirical evaluations on language model pre-training across varying numbers of local workers show that IsoLoCo significantly outperforms DiLoCo, with the gap between them widening as the number of workers increases. This advantage remains present across model sizes and inner step counts, confirming that merging-inspired aggregation is an effective strategy for low-communication distributed training.
PaperID: 2332, Poster
Abstract: We study cross-sample modality attribution in multimodal learning: how a modality in one training sample affects learning from other samples. This problem is central to interpretation because a modality may be redundant for some samples, uniquely informative for others, or useful only through synergy with other modalities. However, existing data attribution methods are insufficient for this setting because they typically estimate sample-level contributions using additive or linear approximations, and therefore cannot capture nonlinear interactions across samples and modalities. We propose a structured kernel surrogate (SKS) framework for multimodal attribution. SKS learns a mapping from sample-and-modality selections to the final loss. To approximate this mapping, SKS uses kernel ridge regression with a structured radial basis function kernel that captures nonlinear effects across samples and modalities. The kernel also preserves the natural hierarchy of multimodal perturbations by modeling sample participation separately from within-sample modality composition. Across diverse environments and benchmarks, SKS achieves 60.3% higher linear datamodeling score than the strongest baseline. For downstream data selection, SKS improves the average accuracy by 2.24% over the strongest baseline. Sharpness analysis further shows that models trained on data selected by SKS have lower Hessian trace values, indicating a flatter loss surface and better generalization. Beyond these quantitative gains, SKS also provides interpretable modality-level selections in spatiotemporal traffic accident prediction, preferentially retaining satellite tiles with visible road infrastructure over vegetation-dominated tiles.
Abstract: In this paper, we introduce INEUS, a meshfree iterative neural solver for partial integro-differential equations (PIDEs). The method replaces the explicit evaluation of nonlocal jump integrals with single-jump sampling and reformulates PIDE solving as a sequence of recursive regression problems. Like Physics-Informed Neural Networks (PINNs), INEUS learns global solutions over the entire space-time domain, yet it offers a more efficient treatment of nonlocal terms and avoids the computationally expensive differentiation of full PIDE residuals. These features make INEUS particularly well suited for high-dimensional PDEs and PIDEs. Supported by a contraction-based convergence proof for linear PIDEs, our numerical experiments show that INEUS delivers accurate and scalable solutions for various high-dimensional linear and nonlinear examples.
PaperID: 2334, Poster
Abstract: Understanding holistic 3D human motion and the surrounding environment from egocentric video is fundamental for applications in AR/VR and robotics. Although these tasks are inherently coupled, existing methods treat them separately. In particular, egocentric motion reconstruction methods typically assume access to metric, gravity-aligned camera trajectories from specialized inertial sensors. Recovering such trajectories from monocular egocentric video remains a significant challenge due to rapid head movements and motion blur. We present EgoMo3R, a joint optimization framework that simultaneously reconstructs 3D human motion and scene geometry from casual egocentric video. To bridge the egocentric domain gap that often hinders end-to-end models, our approach leverages the robustness of low-level visuo-geometric primitives, including monocular depth, optical flow, and point tracks, integrated with a diffusion-based human motion prior via first-order optimization. We propose an alternating optimization scheme that establishes a synergistic feedback loop: scene reconstruction provides physical grounding and trajectory conditioning for the motion prior, while the motion prior regularizes the camera trajectory and anchors the scene in a metric, gravity-aligned world frame. We evaluate EgoMo3R on the highly dynamic sequences of the EgoExo4D and Nymeria datasets and show that it outperforms specialized baselines in both 3D scene reconstruction and egocentric motion estimation.
PaperID: 2335, Poster
Abstract: LLM routing dispatches each user query to the most suitable model from a pool of candidates, balancing response quality against serving cost. As LLMs are increasingly deployed in conversational interfaces, routing decisions must now be made repeatedly within an ongoing session. Existing routers treat each query independently, implicitly assuming that the optimal model depends only on the current input. This assumption ignores two forms of session state that accumulate across turns: the dialogue state, which shapes query meaning and difficulty, and the serving state, which determines each model's true serving cost via KV-cache reuse. Neglecting dialogue state leads to misrouting when queries depend on prior turns, while ignoring serving state causes systematic cost overestimation that worsens as conversations grow longer. A practical multi-turn router must therefore capture dialogue state for accurate quality prediction, and track per-model serving state for accurate cost estimation. To address this gap, we introduce Chain-of-Route (CoR), the first state-aware routing framework for multi-turn conversations. CoR comprises two modules: a Dialogue State Propagator (DSP) that maintains a compact recurrent state to capture session-level context across turns, and a Serving State Tracker (SST) that tracks per-model KV-cache to compute true serving costs. Together with the framework, we construct a unified multi-turn routing benchmark from WildChat and LMSYS-Chat-1M with per-turn quality annotations across five LLMs. Experiments show that CoR achieves state-of-the-art quality-cost tradeoffs, reducing QNC by up to 20.3% over the strongest baseline while improving response quality.
Abstract: (SVs) are widely used to influence the expression of concepts (e.g., truthfulness) in large language model outputs. A key assumption underpinning SVs is that they are linearly discriminative with respect to the concept: representations of texts that exhibit the concept are more aligned with the SV than those that do not, motivating shifts along the positive or negative SV direction to respectively promote or suppress the concept. In this work, we identify an phenomenon in which some highly discriminative SVs that are aligned with positive representations can . We refer to such vectors as inverted-steering vectors (ISVs). We provide a geometric characterization of ISVs' effects, finding that steering along these directions systematically pushes representations in discriminative downstream heads as if the concept were absent, even prior to decoding. Motivated by this analysis, we propose an approach for distinguishing ISVs without requiring generation or associated response scoring. This enables targeted sign flips, which we use to improve a foundational detection-based steering pipeline via Inference Time Intervention (ITI). Our approach improves results in 27/30 experiments, ranging from +0.9% to +138%. We evaluate our findings on Gemma 3 12B, Qwen 2.5 14B, and Olmo 3 7B across 5 concepts. We will make our code publicly available upon acceptance.
Abstract: Gradient-based methods have received increasing attention for Boolean satisfiability, both in neural-network-based solvers and broadly as continuous local search techniques. In this work, we provide a deeper understanding of the interplay of optimization process, problem structure, common formula transformations and neural network properties. Our unified view treats both direct gradient descent over variable marginals and neural-parametrized gradient descent as instances of the same continuous SAT framework. We show that its objective, the expected number of unsatisfied clauses under independent variable marginals, equals (up to constant shift) the clausewise Fourier objective commonly used in continuous SAT methods, and characterize its relation to the exact product-measure probability of unsatisfiability. Then, we derive convergence conditions for gradient descent. A central theme is symmetry: gradient descent can become trapped in invariant subspaces induced by formula symmetries. To remedy this, we show that formula transformations commonly used as SAT preprocessing tools can be repurposed to modulate the optimization landscape. We analyze their effect on problem structure, smoothness, tightness, and expressivity barriers of the neural network. Empirically, across several benchmarks, these transformations can improve gradient-descent performance. Our experiments also indicate that transformations can improve data efficiency and inference-time performance in self-supervised neural SAT solving. Finally, we demonstrate that neural-parametrized gradient descent can improve over input-space gradient descent across several datasets.
PaperID: 2338, Poster
Abstract: Class-incremental learning aims to learn a stream of new classes while preserving knowledge from prior tasks, but suffers from catastrophic forgetting when the model is adapted across sessions. Recently, approaches leveraging vision-language pretrained models such as CLIP have gained increasing popularity in this setting, due to the strong transferable representations of foundation models. Most CLIP-based approaches resist forgetting by retaining data-derived class memory such as replay samples, visual prototypes, or class distribution information. In contrast, we ask whether stable adaptation can be achieved by constraining adapter-induced feature displacement without retaining any data-derived class memory. We propose Hyperbolic Displacement-Constrained Adaptation (HDCA), which treats forgetting as a consequence of uncontrolled residual displacement induced by sequential adapter updates. HDCA keeps the direction of each adapter update but compresses its radial magnitude, so small updates remain almost unchanged while large or accumulated updates are suppressed. At the per-task level, a hyperbolic low-rank adapter makes each task residual grow sublinearly. At the multi-task level, a cumulative compression stage further suppresses the summed residual across sessions. On the classifier side, HDCA introduces a parameter-free Poincar\'e distance classifier that preserves cosine-based ranking for fixed adapted features but reshapes the training gradient in high-similarity regions. We conduct experiments on five standard class-incremental datasets, and the results comprehensively validate the effectiveness of HDCA without retaining any data-derived class memory during training and inference.
Authors: Niclas Pokel, Benjamin F. Grewe
Abstract: Deep networks are powerful function approximators, but they typically store many different computations in shared weight matrices, making it difficult to selectively reuse or adapt parts of them when a familiar structure appears in novel combinations. We introduce the Vector Network (VN), a hierarchical recurrent architecture in which each layer replaces a fixed weight matrix with a library of reusable rank-1 weight atoms. For each input, VN minimizes a layer-local energy to infer a sparse set of active weight atoms and their coefficients, jointly constrained by bottom-up input reconstruction and top-down feedback consistency. These weight atom coefficients then compose an input-specific low-rank weight matrix for that sample. After convergence, slow learning updates only the selected weight atoms through local residual signals scaled by the inferred coefficients. We evaluate VN on four compositional benchmarks spanning 1D signals, 2D spatial decoding, N-body dynamics, and compositional MNIST. VN matches strong baselines in distribution while often achieving out-of-distribution error about an order of magnitude lower when familiar factors must be recombined in novel ways. Vector networks thus make compositional generalization a structural property of the architecture and inference process rather than a brittle byproduct of fitting many behaviors into one shared dense parameter substrate.
Abstract: While recent advances in data synthesis aim to curate high-quality datasets, most generation pipelines still rely on heuristic prompt-based control. This black-box paradigm provides limited insight into how individual samples interact with a model's underlying learning dynamics. To bridge this gap, we propose a circuit-grounded framework that connects training-dynamics-based data valuation with mechanistic interpretability (MI). Specifically, we conceptualize data quality along three complementary utility axes, learnability, challenge, and alignment. First, we uncover specialized model-internal circuits that causally govern these utility signals. Then, moving beyond heuristic prompting toward mechanistic control, we leverage these circuits as controllable interfaces, actively steering generation to produce utility-targeted data. Building on this capability, we introduce SAMS (Stage-Aware Mechanistic Scheduling), which schedules circuit-steered data according to the model's evolving optimization needs. Experiments on multiple-choice QA tasks demonstrate that our approach yields precisely controlled data with greater diversity than prompt-based baselines, consistently improving downstream performance and calibration. Ultimately, this work establishes a principled white-box paradigm for interpretable data generation, pioneering the use of MI not just as an analytical tool, but as a practical, controllable interface.
PaperID: 2341, Poster
Abstract: Behavioral cloning (BC) is a powerful paradigm for learning robotic policies from demonstrations. However, most existing approaches operate in an offline setting and do not address a critical requirement for real-world deployment: the ability to continually improve policies after deployment. While interactive imitation learning methods such as DAgger enable improvement by collecting additional data, they rely on repeated retraining over an ever-growing dataset, leading to increasing computational cost and limited scalability. In practice, robotic systems must adapt continuously during deployment while remaining operational, making retraining on all accumulated data impractical. In this work, we introduce an online formulation of interactive BC, where the policy is updated in real time with expert corrections to address failure cases, with no access to past data. This setting presents a fundamental challenge due to catastrophic forgetting, which degrades previously acquired behaviors when learning from new data. To address this, we propose a novel method based on online ridge regression and random projections, enabling efficient policy updates without gradient-based retraining. We evaluate our approach on a range of manipulation tasks and show that it enables continuous and stable improvement after deployment, outperforming state-of-the-art BC baselines.
Abstract: Graph clustering is a fundamental task for understanding graph-structured data without labels, yet classical end-to-end methods require retraining on each dataset, limiting their generalization ability. Although Graph Foundation Models (GFMs) enable transferable learning across diverse graph tasks, they are not directly applicable to graph clustering. This limitation stems from two key factors: the inability to learn mutually beneficial knowledge on both the node and graph levels, and the lack of a completely unsupervised clustering-oriented design. To bridge this gap, we propose a lustering (GFM-UC), which pre-trains on the source domain and only requires fine-tuning for fast clustering in the target domain. The entire learning process does not require any labels. Specifically, in the pre-training stage, we select learnable high-confidence unified subgraph prototypes from the source domain. And then we generalize them to train the clustering network collaboratively. In the fine-tuning stage, we align the target domain with the source domain in a distribution-aware manner to further refine the clustering head and obtain more well-separated cluster boundaries. Extensive experiments on five datasets demonstrate that our method significantly outperforms existing state-of-the-art approaches.
PaperID: 2343, Poster
Authors: João Vitor Romano, Margaux Zaffran, Ryan Tibshirani
Abstract: Online gradient descent has recently been shown to satisfy gradient equilibrium for a broad class of loss functions. This means that the average of the gradients of the losses along the sequence of estimates converges to zero, a property that allows for quantile calibration and debiasing of predictions, among other useful statistical properties. A shortcoming of online gradient descent when optimized for gradient equilibrium is that the sequence of estimates is jagged, leading to volatile paths. In this work, we propose the generalized momentum method, defined as a weighting of past gradients, as a broader algorithmic framework with guarantees to smoothly postprocess (e.g., calibrate or debias) predictions from black-box algorithms, yielding estimates that are more meaningful in practice. We prove it achieves gradient equilibrium at the same convergence rates and under assumptions similar to those of plain online gradient descent, all the while producing smoother paths that preserve the original signal amplitude. Of particular importance are the consequences for sequential decision-making, where more stable paths translate to less variability in statistical applications. These theoretical insights are corroborated by experiments on real data, showcasing the benefits of adding momentum.
Abstract: Reinforcement learning from verifiable rewards (RLVR) produces strong reasoning models, yet these models can fail catastrophically when the conditioning context is fallible (e.g., corrupted chain-of-thought, misleading partial solutions, or mild input perturbations), because standard RLVR optimizes final-answer correctness only under clean conditioning. We introduce GASP (Guided Adversarial Self-Play), a robustification method that explicitly trains error detection and repair capabilities using only outcome verification. Without human labels or external teachers, GASP instantiates an adversarial self-play game within a single model: a polluter learns to induce failure through locally coherent corruptions, while an agent learns to diagnose and recover under the same corrupted conditioning. To address the scarcity of successful recoveries early in training, we propose in-distribution repair guidance, an auxiliary imitation objective on self-generated repairs that increases recovery probability while preserving previously acquired capabilities. Across four open-weight models (1.5B–8B), GASP converts strong-but-brittle reasoners into substantially more robust ones that better withstand misleading and perturbed context, while often improving clean accuracy. Further analysis shows that adversarial corruptions create an effective curriculum, and that in-distribution guidance enables rapid recovery learning with minimal representational drift.
Abstract: Safe reinforcement learning (RL) commonly enforces expected-cost constraints, but such expectation safety may fail to control the probability of rare high-cost trajectories. Chance-constrained MDPs (CCMDPs) impose a stronger probability-level requirement, but are widely viewed as harder because the chance constraint is nonconvex and depends on the full trajectory rather than a Bellman-linear expectation. In this paper, we reveal that this computational difficulty does not imply a higher statistical price. In tabular discounted CCMDPs, we give a Bellman-certified model-based algorithm whose sample complexity matches the \emphminimax optimal primary dependence of classical CMDP learning, and prove a matching lower bound. Technically, our key idea is the \emphBellman distributional certificate, which constructs a Bellman recursion for constraint violation probabilities before policy selection. The certificate can be reused across candidate policies, shifting concentration analysis from the number of possible policies to a finite Bellman certificate table. Building on this certificate, we establish high-probability safety and near-optimality guarantees for deterministic policy learning in discounted CCMDPs. To our knowledge, we also provide the first model-free sample complexity guarantee for stochastic policy learning in CCMDPs via variance-reduced policy gradient. Numerical experiments on synthetic CCMDPs and IEEE 14-bus energy-storage control benchmark illustrate the safety and behavior of the proposed algorithms.
PaperID: 2346, Poster
Abstract: Self-consistency is a widely adopted test-time scaling method that improves LLM reasoning by majority-voting over multiple sampled traces. However, fixed-budget self-consistency applies the same trace count to every problem, even when the first greedy attempt already produces the correct answer, wasting substantial compute on resource-constrained hardware. Existing adaptive methods rely on post-decode agreement signals that become reliable only after complete traces have been generated, yet a prefill-stage oracle reveals that a large fraction of this cost can be eliminated before spending the full self-consistency budget. We introduce the Hidden-State Adaptive Trace Scheduler (HATS), which predicts greedy-trace reliability from prefill hidden states before allocating additional sampled traces. HATS validates one greedy probe before routing easy problems to a single trace, borderline problems to a small vote, and hard problems to the full budget. A compute-optimal allocation analysis shows that extra samples should go where marginal accuracy gain is highest, explaining why non-uniform spending outperforms fixed budgets. On MATH-500 with DeepSeek-R1-Distill-Qwen-7B under single-GPU serving, HATS matches the four-sample self-consistency accuracy (80.6%) while reducing samples by 48% and decode tokens by 40%, a result confirmed under strict out-of-fold evaluation. These results support prefill-guided scheduling as a practical token- and sample-saving mechanism for test-time scaling.
PaperID: 2347, Poster
Authors:
Ziyan An, John Stankovic, Meiyi MaAbstract: AI-enabled Cyber-Physical Systems (CPS) are highly vulnerable to adversarial and anomalous inputs, where small perturbations can induce cascading errors and unsafe control actions. Existing approaches—such as rule-based filtering, training-time regularization, or diffusion-based reconstruction—either operate outside the model or lack mechanisms to incorporate formal security specifications into the prediction process. In this paper, we develop the first work toward embedding security properties directly into AI-enabled CPS, enabling predictive models to enforce system-level constraints during inference rather than relying on external defenses. We introduce a logic-conditioned bi-stage diffusion framework that integrates Signal Temporal Logic (STL) specifications into forecasting. STL serves as a first-class conditioning signal that guides both an input repair stage and an output refinement stage, allowing the model to jointly mitigate adversarial perturbations and enforce desired temporal behaviors to satisfy security-critical properties. We evaluate our approach on a real-world multivariate CPS forecasting task under both physical sensor and cyber attacks. Our method improves robustness and specification compliance, degrades more gracefully as attack strength increases, and generalizes better to unseen attacks. Ablation studies further show that embedding logical security properties yields gains unattainable by reconstruction-based methods alone, highlighting a new direction for integrating formal methods with generative models in secure CPS.
Abstract: The automatic synthesis of a program from any form of specification is regarded as a holy grail of computer science. Fueled by LLMs, NL2Code has achieved tremendous success, yet the fundamentally more challenging task of synthesizing programs from input-output behavior, which we refer to as IO2Code, remains largely unsolved. Whereas NL2Code can exploit the semantic alignment between natural language and code acquired during pretraining, IO2Code requires recovering underlying principles from concrete computational behavior, navigating a vast and underspecified hypothesis space. To address this, we propose DIO-Agent, a discovery agent for IO2Code. Our method frames IO2Code as an evolutionary search over discrete program space, in which an LLM serves as the mutation operator and concrete error signals from execution guide each mutation. To prevent the search from wandering into structurally complex yet incorrect dead ends, we introduce the Transformation Priority Premise as a mutation prior that biases the LLM toward the simplest hypothesis consistent with current evidence, progressively escalating from constants to conditionals to iteration only when simpler constructs are insufficient. To facilitate systematic study, we further construct an IO2Code benchmark spanning multiple difficulty levels. Extensive experiments show that DIO-Agent consistently outperforms both traditional program-by-example methods and evolutionary search baselines across all difficulty levels and various LLMs, while substantially surpassing test-time scaling strategies with equivalent sampling budgets. Our code and dataset are available at \urlhttps://anonymous.4open.science/r/IO2Code.
PaperID: 2349, Poster
Authors:
Gabriel J Perin, Lucas Boscaini, Andre Araujo, Nina S HirataAbstract: Task vectors enable post-training model editing by identifying semantically meaningful directions in weight space, typically computed as the difference between a fine-tuned model and its pretrained initialization. However, this reliance on fine-tuning makes discovering such directions costly and limits the practicality of post-training model editing. To address this limitation, we introduce Training-Free Task Vectors (TFTVs), a novel method to compute task-vector-like directions without requiring fine-tuning. Our method maps activation steering vectors to rank-one weight-space edits using only forward-pass statistics, while satisfying arithmetic properties that directly support learning via addition, forgetting via subtraction, and the composition of multiple edits. Empirically, we evaluate TFTVs on large language model behavioral control tasks and show that they consistently amplify, suppress, and compose target behaviors while preserving general knowledge and problem-solving skills. We also validate our method against other editing and steering baselines, experimentally demonstrating that TFTVs achieve stronger trait control with better or competitive utility preservation. We hope our work opens new directions for the community in post-training model editing and broader training-free model control. Code will be released.
PaperID: 2350, Poster
Authors: Charley Li, Alice Cao
Abstract: Many agent-memory evaluations collapse state revision, deprecation, erasure, and provenance into retrieve-heavy scores. We introduce MemContract, a diagnostic benchmark of 1 , 080 multi-session tasks over store, retrieve, revise, deprecate, erase, and prove-origin, with formal pre/post-conditions, mandatory distractor sessions plus alias rotation on every family, blinded/stale/delayed controls, a fixed 25-probe residue audit, and benchmark-validity checks via a 180-task human contract-fit audit and a 120-task human-assisted upper bound. Under matched prompts, compute, storage, and latency budgets on GPT-4o across seven architectures, Graph Memory reaches 67.6% macro average (95% paired-bootstrap CI [65.8, 69.4]) versus 62.4% for the strongest hybrid baseline and 40.6% for RAG, with most separation concentrated in revise, deprecate, erase, and prove-origin while store/retrieve remain comparatively compressed. Claude 3.5 Sonnet and Gemini 1.5 Pro 002 preserve the same top-three ordering on a matched 360-task rerun, while the manual audit finds 96.7% agreement that the harness label matches the intended contract and the human-assisted upper bound reaches 93.8% macro, placing the best-system score in a solvable regime rather than a floor of label noise. Reviewer-critical calibration checks tell a narrower but durable story: removing prove-origin leaves the Graph - Hybrid gap unchanged at 5.2 pp, best-effort prompt/schema tuning narrows it to 2.6 pp, and fully external LongMemEval wu2024longmemeval and LoCoMo maharana2024locomo narrow it to 1.8 and 2.8 pp while preserving the top-three ordering Graph > Hybrid > Editable KV. The stable claim is therefore not universal backend superiority: contract-sensitive evaluation reveals separation that broader or retrieval-heavier workloads attenuate but do not erase. The 25-probe audit bounds catchable residue leakage at 3.9% for Graph Memory and 27.4% for RAG, but does not certify deletion.
PaperID: 2351, Poster
Abstract: Scaling Transformers to long contexts is constrained by the quadratic cost of self-attention and the linear growth of key-value cache memory transfer. Sparse attention mitigates this by retrieving only relevant tokens, but current approaches either require large-scale training or, within the training-free regime, rely on semantically coarse heuristics or expensive clustering that is difficult to update efficiently during decoding. We introduce CommunityKV, a framework that formulates sparse attention as a community detection problem. CommunityKV constructs a token graph from the QK^T scores already computed during standard prefill, and partitions the graph into communities to enable retrieval of semantically coherent token groups. A local update rule assigns newly generated tokens to communities in constant time, enabling sparse retrieval throughout streaming decoding without global re-partitioning. We evaluate CommunityKV on open weight models and long context benchmarks, demonstrating that our approach achieves up to 1.86× higher decoding throughput than exact FlashAttention-2 with negligible accuracy degradation.
PaperID: 2352, Poster
Abstract: Humanoid teleoperation datasets are growing rapidly, with operator demonstrations retargeted onto a shared robot skeleton to train whole-body controllers and foundation models. We report a surprising property of this pipeline: retargeting normalizes body proportions but preserves operator movement dynamics (joint velocity profiles, ranges of motion, and coordination patterns shaped by the operator's physiology). On BONES-SEED (522 operators, 142K sequences), retargeted Unitree G1 trajectories support gender classification at 96.0% and operator re-identification at 97.2% Top-1; on operators never seen during training, gender holds at 83.4% and age and height regress within ±4.2 yr and ±5.7 cm. Partial correlation analysis reveals emergent, biomechanically interpretable structure: the signals are task-invariant across activity categories and hold across retargeting implementations. We introduce UNVEIL, a skeleton-aware spatiotemporal graph network to measure and interpret this effect, take initial steps toward operator anonymization, and ask the community: as teleoperation datasets scale, what should our data practices be? Code, models, and anonymized trajectories are at our project website https://project-unveil.github.io.
PaperID: 2353, Poster
Abstract: Distributional Reinforcement Learning methods aim to learn the entire distribution of returns, yet current algorithmic approaches are limited to a narrow class of distributions. We propose \varphiTD, a method to overcome this limitation through computing the loss in the frequency domain. We demonstrate that this further unifies quantile and categorical approaches to distributional RL, as well as enabling the learning of distributions that were previously impossible to capture, such as mixtures of continuous distributions. We empirically study \varphiTD across synthetic benchmarks and large-scale deep RL environments, and demonstrate that the flexibility obtained by wider parametric families often leads to improved performance.
PaperID: 2354, Poster
Abstract: Recent progress in large language models (LLMs) has led to impressive performance across diverse domains. However, their deployment in critical areas such as healthcare, law, and education raises serious concerns regarding the potential generation of harmful content. While reinforcement learning (RL)-based safety alignment methods have been extensively studied, they face persistent challenges, including sparse reward signals, limited interpretability, and limited transferability across domains. In this paper, we introduce Rubric‑Align, a framework that replaces static binary scalar rewards with dynamic, natural‑language evaluation rubrics. Instead of assigning a single reward to each response, Rubric‑Align provides fine‑grained and interpretable feedback, thereby alleviating reward sparsity and improving the transparency of the alignment signal. A key feature of Rubric-Align is its dynamic rubric evolution mechanism. Rather than using fixed reward templates, Rubric-Align periodically refines prompt-specific rubrics based on the current policy behavior, keeping the supervision signal informative and aligned with the model’s evolving failure modes. Experiments on safety benchmarks and vertical-domain settings show that Rubric-Align improves robustness against harmful and jailbreak prompts while largely preserving general capabilities, suggesting that natural-language rubrics provide a transferable interface for adapting safety supervision to new safety subdomains.
PaperID: 2355, Poster
Abstract: Large Language Models (LLMs) increasingly ship with “thinking modes”, yet their counterpart, the “no-thinking” has received far less exploration. In this work, we revisit language model “no-thinking” behavior along the following efforts: Rather than relying on proxies like thinking-mode control, we split each response into pre-answer text and a final answer, scoring both answer-only compliance and the question–pre-answer relevance. ) across boolean, multiple-choice, and open-ended tasks, on open-source models with scaled variants, and proprietary models. With extensive experiments, we found that Explicit no-think controls do not reliably eliminate visible reasoning-like content. Instead, models exhibit : residual content is systematically organized by the task's answer space, with boolean verification compressing most easily, multiple-choice tasks forming an intermediate regime, and open-ended tasks remaining the most resistant. We further identify what we call a no-thinking and performance trade-off in open-ended: stronger answer-only constraints on complex tasks either fail to suppress the intermediate payload or succeed by removing computation needed for accuracy. Finally, we show that compressibility depends strongly on the structural support provided by the question itself. Thus, no-think controls are not direct guarantees of no-thinking; they must be evaluated jointly through answer-only success, visible payload, and task accuracy. We hope our work moves LLMs closer to a human-like ability to stop thinking on demand.
PaperID: 2356, Poster
Abstract: Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count. Recent architectures already augment CLIP with a T5 encoder precisely because CLIP's contrastive embedding loses compositional structure, yet these failures persist. We argue the binding problem is therefore not one of missing information but of misaligned information: a text encoder preserves compositional structure, but in a representation space shaped by language modelling rather than vision, and the denoising objective does not directly reward aligning the two. We show this correspondence can be supplied as an explicit training signal, that the relevant cross-modal information is concentrated in a low-rank subspace of self-supervised visual features, and that supplying it can be folded into diffusion training as a single auxiliary loss. Our method, ORCA, Orthogonal Residual Compositional Alignment, aligns the latent of a diffusion transformer with a low-rank target derived from a frozen visual encoder, through a predictor whose orthogonal basis is parameterised by a learned residual between T5 and CLIP embeddings, which provides a prompt-dependent signal for selecting the visual readout subspace. We prove that the cross-modal information recoverable at a given rank is bounded by the spectral mass of the visual encoder's covariance in the top components. Across three diffusion-transformer backbones, DiT-B/2, DiT-L/2, and U-ViT-L, ORCA improves FID and GenEval over both vanilla and REPA baselines at zero inference-time cost; on DiT-L/2 it reaches FID 16.65 and GenEval 0.291 at 200K steps, exceeding the strongest 400K baseline at half the training cost, with the largest gains concentrated on attribute binding, spatial relations, and multi-object prompts.
PaperID: 2357, Poster
Abstract: Cross-modal person re-identification (ReID) adversarial attacks face a fundamental generalization dilemma: existing methods suffer from source-pair overfitting, where the learned perturbations become entangled with the feature covariance and alignment patterns of the source modality pair, thereby limiting their transferability under heterogeneous domain shifts. We argue that the core issue is the difficulty of reducing source-pair-specific bias while maintaining attack effectiveness. To address this, we propose Multiform Attack (MA). The framework first learns a universal attack direction via Mahalanobis-guided gradient optimization to capture the intrinsic covariance structure of the feature manifold, though this perturbation may still retain source-pair-specific bias. To overcome this bias and achieve cross-distribution generalization, the second stage leverages multiple heterogeneous source distributions to optimize for generalization, performing sparse, structure-sensitive residual correction in the discrete pixel space via attention-shift-filtered multi-objective evolutionary search. This two-stage design models the attack as a base direction plus a structured residual, which helps reduce overfitting to the source modality pair and improve transferability to unseen modalities. Extensive experiments show that MA achieves superior transferability across unseen modalities, datasets, and model architectures.
Abstract: Variational inference (VI) is a cornerstone of modern Bayesian learning, enabling approximate inference in complex models. However, its formulation depends on expectations and divergences defined through high-dimensional integrals, often rendering analytical treatment impossible and necessitating heavy reliance on approximations. Possibility theory, an imprecise probability framework, allows us to directly model epistemic uncertainty instead of relying on a subjective interpretation of probabilities. While this framework provides robustness and interpretability under sparse or imprecise information, adapting VI to the possibilistic setting requires rethinking core concepts such as divergences, which presuppose additivity. In this work, we develop a principled formulation for performing possibilistic VI by establishing a maxitive analogue of the classical Donsker-Varadhan formulation. The resulting framework enables us to derive a learning rule for possibilistic VI with exponential-family candidates and practical update rules for neural-network training, giving rise to a family of optimizers termed
PaperID: 2359, Poster
Abstract: We introduce MANGO (Multi-Angle Neural Gated Operator), a spectral neural operator for chirp-perturbed PDEs whose solutions exhibit a local spectral content that rotates continuously with position. MANGO maintains multiple fractional Fourier transform (FRFT) branches whose angles are learnable parameters, optimized jointly with the spectral weights and a position-dependent softmax gate that combines branches at every spatial location. The architecture is matched to the structure of the problem class: the FRFT at angle \alpha_0 diagonalizes the chirp-perturbed PDE whose intrinsic angle is \alpha_0, and MANGO's learnable angles let the architecture discover \alpha_0 from data when it is unknown or varies across the domain—neither of which classical FRFT methods accommodate. We further introduce a Mihlin–Hörmander regularizer on the spectral weights of each non-Fourier branch, encouraging each branch to act as a bounded operator on L^p for 1 < p < \infty. We construct four chirp-perturbed PDE benchmarks with closed-form ground truth, positioned as identifiability tests. On these benchmarks, MANGO recovers the intrinsic spectral structure of each PDE from supervised solution data alone and achieves state-of-the-art performance against six neural-operator baselines.
Abstract: Improving the inference efficiency of autoregressive transformers typically means reducing FLOPs per token, usually through approximations that degrade model quality. We introduce N-vium, a mixture-of-exits transformer that partially parallelizes computation across depth on standard hardware, increasing effective FLOPs per second rather than minimizing compute per token. N-vium attaches prediction heads at multiple depths and defines the next-token distribution as a learned mixture over these exits, with token-adaptive routing. This formulation strictly generalizes the standard transformer, which is recovered exactly when routing assigns zero mass to all intermediate heads. Sampling from the mixture is exact, and complete KV caches are recovered by deferring the upper-layer computation and batching it with later tokens. We pretrain N-vium at scales up to 1.5B parameters. Our largest model reaches 57.9% wall-clock speedup over a parameter- and data-matched standard transformer at no perplexity cost.
Abstract: Constructing minimum-volume prediction regions that satisfy conditional coverage is a fundamental challenge in multivariate regression. Standard approaches rely on explicitly estimating the full conditional density and subsequently thresholding it. This two-step plug-in process is notoriously difficult, sensitive to estimation errors, and computationally expensive. One would like to instead optimize the region directly. Formulating a direct solution is challenging, however, because it requires minimizing a volume objective that is coupled with the conditional quantiles of the model's own estimation error. In this work, we address this challenge. We introduce \emphsuper-level-set regression (SLS), a novel mathematical framework that successfully resolves this implicit coupling, allowing us to directly parameterize and optimize the geometric boundaries of the target conditional level sets. By bypassing full distribution estimation and leveraging flexible volume-preserving frontier functions, our approach natively captures complex, multimodal, and disjoint conditional structures end-to-end. Ultimately, SLS offers a new perspective on multivariate conditional quantile regression, replacing the restrictive assumptions of density-first methods with a direct geometric optimization strategy.
Abstract: Watermarking techniques for large language models (LLMs), which encode hidden information in the output so its source can be verified, have gained significant attention in recent days, thanks to their potential capability to detect accidental or deliberate misuse. Similar challenges involving model misuse also exist in the context of game-playing, such as when detecting the unauthorized use of AI tools in gaming platforms (e.g., cheating in online chess). In this paper, we initiate the study of how game-playing strategies can be watermarked. We show how the KGW watermark for LLMs can be adapted to watermark game-playing agents in perfect-information extensive-form games. The watermark can then be detected using a statistical test. We show that the degradation in the quality of the watermarked strategy profile, quantified by the expected utility, can be bounded, but there is a tradeoff between detectability and quality. In our experiments, we bootstrap the watermarking framework to various chess engines and demonstrate that a) the impact of the watermark on the quality of the strategy is negligible and b) the watermark can be detected with just a handful of games.
PaperID: 2363, Poster
Authors: Hyunwoo Park, Junhyeok Um, Baek-Ryun Seong, Sang-Ki Ko
Abstract: Open-loop action repetition—hereafter also referred to as a skip—enables reinforcement learning agents to reduce decision frequency by executing selected actions for adaptive durations. A key challenge is how to evaluate the value of such a skip. Existing Skip-MDP methods typically bootstrap from the greedy value of the terminal state, implicitly assuming that the agent can immediately resume near-optimal control after the skip. This assumption is unreliable during learning, when the underlying action policy is stochastic, inaccurate, and continually changing. It can also amplify optimistic estimation errors, distorting the learned preference over skip lengths. We propose Expected Skip Evaluation, a reformulation of skip-value learning that evaluates skip terminal states through expected future control rather than idealized greedy control. This reflects imperfect control during training while preserving the optimal skip-value fixed-point as control becomes reliable. We instantiate this principle in RARe (Risk-Aware Repetition), a practical framework for discrete and continuous action spaces. Experiments across grid-world, continuous-control, and safety-critical benchmarks show improved sample efficiency and better-calibrated skip decisions.
Abstract: Large language models (LLMs) deliver impressive performance but incur prohibitive memory and compute costs at deployment. Model pruning is an effective way to reduce these overheads, yet existing approaches face challenges: unstructured sparsity, where nonzeros can appear anywhere, preserves accuracy but yields irregular access patterns that prevent GPU acceleration, while semi-structured 2:4 sparsity is hardware-friendly but enforces a rigid 50% pattern that degrades model quality. To bridge this gap, we introduce PATCH, a hybrid sparsity framework that enables a continuous sparsity ratio between 0% and 50%. PATCH partitions weight matrices into tiles, assigning each tile to be either dense or 2:4 sparse via a learnable mask selection mechanism. This design provides fine-grained control over accuracy–acceleration tradeoffs and supports non-uniform sparsity across layers, leading to superior overall quality. Across models from 0.5B to 13B parameters, PATCH consistently narrows the gap to dense accuracy while delivering practical speedups. For instance, on LLaMA-2 7B with an A6000 GPU, PATCH achieves 1.18×–1.38× end-to-end speedup over dense baselines while improving accuracy by 0.37%–2.96% compared to the state-of-the-art 2:4 pruning method, MaskLLM.
Abstract: Most existing memory-enhanced Large Language Model (LLM) approaches implicitly assume that memory validity can be established either through external evaluators that provide task-specific success signals or through internal model cognition, such as reflection, for editing memory entries. However, these assumptions often break down in practical environments with dynamic drifts. We propose the Global Verifier (GLOVE), a plug-and-play module for LLM memory systems that establishes a relative notion of truth to achieve memory-environment realignment in the face of environmental drifts. Through active probing to detect inconsistencies between retrieved memories and fresh observations, GLOVE enables memory-environment realignment by verifying and updating memory without access to task-specific ground-truth supervision or strong reliance on model introspection. We evaluate GLOVE across a spectrum of tasks ranging from web navigation and discrete planning to continuous and embodied control, covering both simulated benchmarks and real-world robotic deployment. Across all domains, GLOVE improves adaptation across various LLM memory designs under both explicit and implicit drift, often recovering performance from near-failure to high success rates. These results suggest a practical pathway to memory-environment realignment for self-evolving cognitive agents.
Authors: Eugenio Cainelli, Lorenzo Luccioli, Alessandro Iraci, Michele D'Adderio, Giovanni Paolini
Abstract: Inspired by long-standing open problems in algebraic combinatorics, we show that modern machine learning can meaningfully contribute to verifiable mathematical discoveries. In particular, we focus on the construction of simple mathematical functions under exact distributional constraints, a setting we formalize as Simple Learning Under Rigid Proportions (SLURP). We tackle this problem by introducing two methods: MapSeek-Functional, which models the desired function alternating pseudo-labeling and supervised training steps; and MapSeek-Symbolic, designed to directly produce symbolic formulas. We successfully apply both methods to a research problem in algebraic combinatorics, discovering a new combinatorial interpretation of the q,t-Narayana polynomials arising from representation theory. To our knowledge, this is the first such interpretation based on noncrossing partitions. Using one discovered statistic, we find a combinatorial proof of the symmetry of these polynomials in a previously unsolved case. To streamline verification and reproducibility, we release all code, including a formalization of the mathematical discoveries of this paper in Lean 4.
Abstract: Events in spatiotemporal systems are ubiquitous, yet modeling their complex distributions remains challenging. Existing point process models often rely on strong structural assumptions and are typically limited to autoregressive, event-by-event prediction. As a result, they struggle to support broader inference tasks such as inverse inference, trajectory reconstruction, and recovery of missing event locations. We introduce Arbitrarily Conditioned Hierarchical Flows (ARCH), a hierarchical flow matching framework for spatiotemporal event modeling. ARCH is expressive enough to capture complex event distributions while enabling tractable and accurate computation of conditional intensities, which quantify instantaneous event risk. Built on a history-encoder-generative-decoder architecture, ARCH introduces a hybrid masking strategy for flexible conditioning on arbitrary observed events. This enables a unified treatment of forecasting, inverse inference, and partial trajectory recovery within a single framework. Experiments on synthetic and real-world datasets show that ARCH consistently outperforms existing baselines across both prediction and conditional inference tasks.
PaperID: 2368, Poster
Abstract: Multimodal Retrieval-Augmented Generation (RAG) helps mitigate hallucinations in Vision-Language Models (VLMs) by grounding generations in external knowledge bases. However, this externalized memory also introduces privacy risks, as external queries may reveal signals about sensitive visual records in the retrieval corpus, such as medical images, scanned contracts, and proprietary business documents. This exposes a retrieval-corpus privacy risk: private records may be detectable through interactions with multimodal RAG systems. To study this risk, we propose the Semantic Degradation Attack (SDA), a two-query method for exposing private corpus leakage in multimodal RAG by testing how strongly generated responses depend on retrieved evidence. SDA constructs transferable retrieval-disrupting perturbations using local surrogate visual encoders. By measuring the drop in the semantic similarity of the VLM's generated descriptions before and after perturbation, SDA can distinguish whether a target image-caption record exists in the private retrieval corpus. Extensive experiments demonstrate that SDA consistently detects private corpus leakage more effectively than existing baselines on two image-caption datasets across five popular VLMs.
PaperID: 2369, Poster
Abstract: Dexterous object manipulation remains a fundamental challenge in robotics, typically requiring intensive supervision through explicit rewards, goals, or demonstrations. While unsupervised reinforcement learning (RL) holds the promise of discovering dexterous skills, state-of-the-art metric-aware methods have yet to be successfully applied to complex, high-dimensional manipulators like multi-fingered hands. We argue that this limitation stems from two key factors: existing methods lack the necessary inductive biases to prioritize object interaction over agent motion, and they fail to respect the non-Euclidean topology of object configuration space \mathrmSE(3). To address these limitations, we propose a novel Wasserstein dependency measure between skills and object motions, formulated in the Lie algebra \mathfrakse(3). Our approach, Topology-Aware SKill discovery (TASK), enables the discovery of fundamental manipulation skills—such as multi-axial object translation and rotation—without manual supervision. Across a range of embodiments, including non-prehensile and multi-fingered hands, our method learns dexterous behaviors where previous approaches fail, marking a significant step toward autonomous robotic dexterity.
PaperID: 2370, Poster
Abstract: Generating realistic visual defects is crucial for alleviating the scarcity of abnormal samples in anomaly detection, yet existing methods often rely on category-specific few-shot adaptation and struggle to generalize across objects, defect types, and domains. This limitation stems in part from the lack of paired defect editing data and the fine-grained nature of defect appearance, where subtle variations in scale, texture, and morphology must be preserved while maintaining consistency with the target scene. To address this challenge, we introduce UDG, a dataset of 300K normal-abnormal-mask-caption quadruplets curated from diverse defect-related scenarios, and present UniDG, a unified model for universal defect generation. UniDG formulates defect synthesis as Defect-Context Editing: it extracts reference defect context with adaptive cropping, organizes reference and target inputs in a structured diptych format, and fuses multimodal conditions through MM-DiT attention. We further develop a two-stage training strategy: Diversity-SFT learns diverse transferable defect priors from paired editing data, while Consistency-RFT improves reference adherence, local realism, and defect-category consistency. Without per-category fine-tuning or using MVTec-AD/VisA for training, UniDG outperforms prior few-shot anomaly generation and image insertion/editing baselines in both synthesis quality and downstream single- and multi-class anomaly detection/localization.
Abstract: We introduce the first Probably Approximately Correct (PAC) learning framework for general-sum concurrent stochastic games (CSGs) with transition uncertainty, while addressing the challenge of Nash equilibrium (NE) existence. Our algorithm maintains data-driven L^1 confidence sets over transition kernels and solves a robust CSG to compute a social-welfare optimal \varepsilon-NE, using a novel robust MDP-based exploration mechanism to drive joint state-action coverage without unilateral control. Crucially, we introduce a Nash margin characterisation that enables principled reasoning about equilibrium existence: the framework either returns an \varepsilon-approximate NE whose social-welfare value is \varepsilon-close to optimal, or provides a sound certificate that no exact NE exists. Under a minimum reachability condition p_\mathrmreach > 0 over all state-action pairs, the algorithm terminates after a polynomial number of interactions, with sample complexity \widetilde\mathcalO\left( R_\max^2 H^4 |S|^2 |A| / (p_\mathrmreach \varepsilon^2) \right). Empirical results on benchmark CSGs demonstrate near-optimal performance, correct handling of equilibrium (non-)existence, and sample complexity consistent with theory; implementation available at: https://anonymous.4open.science/r/pg1.
PaperID: 2372, Poster
Authors: Armin Bazarjani, Payam Piray
Abstract: Updating a predictive cognitive map when the environment changes is a central problem for both biological agents and reinforcement learning, yet existing approaches either depend on explicit model knowledge or learn the full state-indexed map from samples. We propose Default Feature Representations (DFR), a featurized parameterization of predictive cognitive maps in which a fixed feature basis is composed with an operator that encodes the current environment. We provide two forms for the operator: a model-based closed form when the structural change between environments is known, and a model-free temporal-difference learning rule that recovers the operator from sampled transitions, with provable convergence to the model-based solution. The model-free DFR reconstructs the perturbed map from samples alone, achieves planning performance comparable to the model-based solution, and substantially outperforms successor-representation baselines on replanning tasks. We also show that DFR captures the local remapping of grid cells observed under local environmental change. By separating the cognitive map into a stable feature basis and a fast-adapting operator, DFR offers a sample-based account of how a predictive map can be updated from local experience, mirroring the stability of entorhinal grid fields across environments.
Abstract: Flow Matching is a powerful framework for learning transport maps between probability distributions. Yet its standard single-parameter formulation is not designed to capture multi-parameter variations where the resulting transport should be path-independent. Path independence is crucial because it ensures that transformations depend only on the initial and target distributions, not on the specific path. In this work, we introduce Path-independent Flow Matching (PiFM), a method for learning vector fields whose induced flows yield path-independent transport between distributions. We show that PiFM generalizes Flow Matching to higher-dimensional parameter domains while enforcing structural conditions that ensure consistency of composed transformations. In addition, we show that, under suitable assumptions, PiFM approximates the Wasserstein barycenter, linking the framework to a notion of distributional interpolation. To enable practical training, we propose a tractable, simulation-free objective that regresses onto multi-parameter conditional probability paths. We showcase empirically that PiFM outperforms other approaches on both synthetic and real world data in interpolating path-independent trajectories and generating desired out of distribution samples. Our code is available at \urlhttps://anonymous.4open.science/r/Pi
PaperID: 2374, Poster
Abstract: Driver monitoring concerns determining not only what a driver is doing, but whether the behavior violates a driver-behavior safety rule under the current driving context. Existing distracted-driving benchmarks largely reduce this problem to closed-set action recognition, often missing cases where the same behavior changes meaning with road scene, vehicle state, duration, and rule-specific exceptions. We present , a rule-grounded context-aware open-world framework for driver anomaly detection. ROAD represents driver-monitoring constraints as structured rule cards and jointly reasons over in-cabin driver evidence, road-context video, optional vehicle-state metadata, and open-world semantic cues to produce structured anomaly judgments: driver action, violated rule, anomaly score, supporting evidence, exception status, and explanation. Unlike generic MLLM prompting, ROAD separates perception, rule grounding, and semantic generalization. Road context modulates driver-centric evidence, rule-card encodings select rule-relevant evidence, and a LoRA-adapted MLLM extracts open-vocabulary safety cues from compressed video evidence. This design targets two forms of open-world generalization: recognizing unseen driver behaviors beyond a fixed action taxonomy and adapting to new or revised rule cards. We further introduce ROAD-Bench, a rule-grounded benchmark with context, exception, evidence, anomaly-score, and explanation annotations. With a three-stage curriculum from action primitives to context-rule alignment and anomaly adaptation, ROAD outperforms evaluated baselines across driver-distraction recognition, open-world behavior recognition, and rule-grounded anomaly detection, using an 8B multimodal backbone. It achieves 92.01% average accuracy over seven driver-recognition benchmarks, 77.9% AUROC and 46.9% unseen-class F1 in behavior recognition, and 73.6% AUROC, 68.3% rule accuracy, and 54.5% Expl.-Rule F1 on ROAD-Bench.
PaperID: 2375, Poster
Abstract: Scaling large language models efficiently has motivated sparse capacity mechanisms such as Mixture-of-Experts and, more recently, conditional memory: token-indexed embedding tables that augment the backbone with cheap parametric lookups. Existing memory-embedding methods retrieve via a deterministic function of the surface form, which collapses different contextual senses of the same token (\eg, \emphpython the language vs.\ the animal) into a single fixed entry. We introduce Mixture of Memory Embeddings (MoME), a context-aware memory mechanism that replaces each token's single memory row with a mixture of M slots and uses a learned gate over the hidden state to choose which slots to read at each position. In controlled pretraining experiments across nanochat, Llama-3/MobileLLM, and Qwen3 backbones from 125M to 0.6B parameters, MoME consistently outperforms Value Embedding, Engram, and STEM baselines at matched memory \& training budgets, shows a more favorable memory-size scaling trend than Engram on the nanochat backbone, and remains compatible with Engram-style memory under compound scaling. Qualitative routing analyses on polysemous tokens further suggest that the learned mixture exhibits a degree of semantic interpretability, dispatching the same surface token to distinct memory slots under different senses rather than collapsing them into a single fixed entry. All Code and model checkpoint will be open sourced.
PaperID: 2376, Poster
Authors:
Adrian Rodriguez, Angelica Kim, Yunhui Guo, Yilun DuAbstract: Generalization to harder reasoning and planning problems remains a central challenge for AI systems. Although both domains can be viewed as search problems where valid solutions are easier to verify than generate, existing methods often rely on domain-specific solvers or diffusion-based energy models tailored to individual settings. We introduce Energy Minimization Search (EMS), a unified framework for reasoning and planning that learns a single time-invariant energy landscape. Unlike diffusion-based energy-based models (EBMs), which perform annealed inference over time-dependent landscapes, EMS performs direct gradient-based optimization on one static landscape using the same training objective and inference procedure across domains. Across reasoning tasks (graph coloring, N-Queens, 3-SAT) and planning tasks (path finding and maze solving), EMS matches or outperforms diffusion-based EBMs, domain-specific solvers, and diffusion-based planners while using substantially less inference compute.
PaperID: 2377, Poster
Abstract: Diffusion models have achieved state-of-the-art (SoTA) performance in generative tasks through their iterative denoising mechanism, yet they remain computationally intensive and energy-prohibitive. Spiking Neural Networks (SNNs) represent a promising energy-efficient alternative to traditional Artificial Neural Networks (ANNs). Their binary spike encoding enables the conversion of multiplication operations into addition operations, which is a key attribute for reducing energy consumption. Integrating the strong generative capabilities of diffusion models with the energy efficiency of SNNs thus forms a highly promising research direction. In adherence to the design goal of all-additive computation, this paper proposes a novel spiking diffusion model with attention enhancement. Specifically, we adopt non-negative ternary spiking neurons (NNTSN) to construct the U-Net backbone, which helps mitigate information loss during feature processing. To align NNTSN with the requirement of using only additive operations, we further design two core components: a membrane potential attention mechanism and an all-additive shortcut module. Leveraging a reparameterization technique, the proposed diffusion model is developed to retain only additive operations during inference, ensuring strict compliance with energy-saving design principles. Extensive experiments on relevant benchmarks demonstrate that our method achieves SoTA performance in generative tasks. It also substantially outperforms other SNN-based generative models while using fewer time steps, which validates both its effectiveness and efficiency.
PaperID: 2378, Poster
Abstract: Deep generative models, particularly Flow Matching frameworks, have emerged as powerful tools for synthesizing complex physical systems. However, standard architectures remain purely statistically driven and agnostic to the underlying governing equations, often violating fundamental physical constraints during inference. Existing zero-shot approaches designed to enforce these constraints typically rely on rigid algebraic projections. While effective for sparse boundary conditions, applying such discrete geometric corrections to dense, full-domain PDEs disrupts the probability path, pulling samples off the learned manifold and introducing intractable computational overhead. To address these challenges, we introduce Energy-Guided Transport, a novel training-free framework that formulates non-linear physical constraints as a smooth potential energy landscape. Rather than interrupting the ODE integration, our method applies a continuous dissipative force directly to the intermediate generative states, seamlessly steering them toward the physical manifold. By evaluating this force implicitly via reverse-mode automatic differentiation, our framework preserves the flow's natural dynamics while maintaining an efficient \mathcalO(N) algorithmic complexity. We extensively evaluate our framework across various 1D and 2D physical systems, including the chaotic Navier-Stokes equations. Our method achieves stable generation without structural divergence, significantly improves the stability-accuracy trade-off compared to state-of-the-art methods, and substantially accelerates inference with respect to exact projection strategies. Code available at: \urlhttps://anonymous.4open.science/r/Projection-free-energy-guided-flow-matching-A302.
PaperID: 2379, Poster
Abstract: NPU design-space exploration for LLM inference mainly relies on an expensive simulator that often returns only scalar feedback such as latency and power. Such feedback can rank evaluated configurations, but gives little information for deciding why a configuration is slow or which architecture parameter should be changed next. This limitation is especially severe because prefill and decode workloads can be limited by different resources, including compute throughput, HBM bandwidth, and inter-chip communication. We propose CriticNPU, a sample-efficient DSE framework built around an online-calibrated mechanistic critic for NPU architecture design. The critic uses explicit hardware limits to evaluate NPU configurations: before simulation, it estimates whether a proposed configuration is likely to improve over the current best and filters weak candidates; after simulation, it aggregates operator-level runtimes into prefill/decode bottleneck summaries and predicts the effect of candidate parameter changes. We instantiate the critic with a Roofline-style model that derives compute, memory, and communication limits from architecture parameters and calibrates per-operator-type effective limits using simulator observations. And CriticNPU combines this critic with a workload extractor and an architecture planner to propose configurations. Across six LLM workloads and 16 deployment configurations per model, CriticNPU achieves best-or-tied-best latency in 74% of prefill and 87.5% of decode settings while using roughly half the simulator calls of competitive DSE baselines. We conduct extensive ablation studies and detailed analyses to validate the effectiveness of our method from multiple perspectives.
PaperID: 2380, Poster
Abstract: Modern scientific discovery systems can propose more mechanistic hypotheses than laboratories can test. We study the downstream selection problem: after upstream retrieval, extraction, expert modeling, or curation has fixed a finite locally factorized posterior over admissible mechanisms, when does the posterior evidence justify acting on one structural hypothesis rather than its rivals? For a query A\rightsquigarrow B, we define structural support as the posterior mass of mechanisms in which A reaches B by a directed path. This is not a retrieval rank, direct-edge score, or MAP graph decision: support can be distributed across many compatible mechanisms. Our main result is a local certificate for this global support quantity. For path cutoff L and locality radius m, truncated support is a convex combination of exact boundary-conditioned local supports. Optimizing over locally admissible blanket states gives a boundary range [\\ell_L,m,u_L,m], and a certified tail bound \bar r_L accounts for longer witnesses: p_AB\in[\\ell_L,m,\min\1,u_L,m+\bar r_L\]. For coupled rivals, the same boundary-conditioned calculations yield a joint feasible set rather than only independent coordinate intervals. We convert these certificates into resolve-or-abstain decisions: an action is resolved only when it wins throughout the certified support set. Under valid certificates, a resolved action is the unique strict utility maximizer inside the declared fixed model, action set, and utilities; otherwise, the solver abstains and reports whether boundary, tail, computation, evidence, or outer-model uncertainty blocks resolution. On controlled and biologically scaffolded finite-posterior records, the implementation has zero endpoint containment misses and zero false fixed-model resolutions. In a coupled-rival diagnostic, joint feasible sets resolve 14/24 cases versus 6/24 for scalar boxes, with no false commitments.
Abstract: As autonomous agents become more capable of performing real-world tasks, distinguishing scheming behavior from benign task pursuit may become a central AI control problem. Existing monitors often rely on chain-of-thought access or internal activations, or use prompted frontier models, all of which can be unavailable, unreliable or expensive in deployment. In this work, we study action-only deliberative monitors: smaller open-weight models trained to detect scheming and sabotage from agentic trajectories without accessing the monitored agent's reasoning or model internals. Our method, inspired by deliberative alignment, uses a scheming specification to elicit structured rationales from a frontier teacher, filters them with a separate judge, and distills the highest-quality rationales into open-weight monitors with supervised fine-tuning and reinforcement learning. We train on five datasets, and evaluate across six out-of-distribution agentic misalignment benchmarks. We show that applying our method to Qwen3.5-27B yields higher performance than all smaller prompted frontier monitors (Gemini 3.1 Flash-Lite, GPT-5.4 Nano, and Claude Haiku 4.5) and than Gemini 2.5 Pro, while also achieving lower marginal inference cost (token-metered USD per 1000 evaluations). Stronger prompted frontier monitors (Gemini 3.1 Pro, GPT-5.4, Claude Sonnet 4.6, and Claude Opus 4.6) achieve higher performance but at roughly 16--34× higher marginal inference cost. Several of our trained monitors are positioned on the empirical cost--performance Pareto frontier among the monitors we evaluate, providing practical low-cost, low-FPR alternatives to prompted frontier models.
Authors: ByungHa Ko, Youngmin Lee, Dong Hwan Kim
Abstract: Hierarchical 3D grouping aims to recover scene groups across multiple granularities, from fine object parts to complete objects, without relying on semantic labels or a fixed vocabulary. The main challenge is to transform 2D foundation-model cues into coherent hierarchy supervision and embed that hierarchy in a 3D representation. We propose H2G, a hyperbolic affinity field for hierarchical 3D grouping. Our method derives semantically organized tree supervision by interpreting foundation-model affinities through Dasgupta's objective for similarity-based hierarchical clustering. This supervision is distilled into a single Lorentz hyperbolic feature field, whose geometry is well suited for tree-like branching structures. A hierarchy-aware objective aligns the field with fine-level assignments, coarse object structure, compact feature clusters, and LCA (Lowest Common Ancestor) ordering. This formulation represents multiple grouping levels in one feature space, enabling semantic hierarchical grouping grounded in 2D foundation-model knowledge.
PaperID: 2383, Poster
Abstract: 3D Gaussian Splatting (3DGS) has become a standard representation for real-time novel-view synthesis. Yet, its quality--efficiency trade-off still relies heavily on adaptive density control. Existing cloning, splitting, and pruning rules use proxy signals—such as position gradients, opacity, image-space error. Although effective in practice, these signals do not directly answer the structural question behind density control: where should new Gaussians be allocated to reduce the multi-view reconstruction loss, and which existing primitives can be pruned with minimal penalty? To address this, we propose first-variation birth–death control for 3DGS, a principled approach that replaces heuristic structural decisions with a variational score derived from the reconstruction loss. At each structural update, we freeze the current Gaussian population, attach an auxiliary mass coordinates to each primitive, and differentiate the multiview loss with respect to these coordinates. We demonstrate that this score is a coordinate of the first variation of Splat Regression Model, giving a rigorous local descent surrogate for birth-death moves: high-score Gaussians are selected as birth sites, low-score low-opacity Gaussians are death. Across standard 3DGS benchmarks, integrating this score into 3DGS-MCMC yields state-of-the-art reconstruction quality and outperforms existing methods, delivering the most significant gains under tight Gaussian budgets. Ultimately, our work provides compelling evidence that gradient-based variational metrics should replace proxy heuristics for structural updates in 3DGS.
PaperID: 2384, Poster
Abstract: How can we elicit truthful information from strategic agents? Traditional peer prediction mechanisms incentivize truthful reporting by rewarding agreement between agents' reports, but break down when agents can submit cheap but correlated signals---such as outputs from different LLMs. We introduce the Speakeasy mechanism, a framework for peer prediction when the principal can also sample these cheap signals. The mechanism augments any peer prediction rule with an audit strategy that penalizes reports for matching the principal's samples. We give a polynomial-time algorithm for the optimal audit strategy, so that truthful reporting is a strict Bayes--Nash equilibrium and yields the highest welfare against agents choosing between reporting truthfully and copying a single cheap signal. We extend the guarantee to more complex misreport strategies that mix multiple cheap signals, showing that Speakeasy preserves the truthfulness of the base mechanism. We empirically test the Speakeasy mechanism against classical peer prediction mechanisms on a peer-review dataset of ICLR submissions, with reviews from human reviewers and several LLMs. Several classical mechanisms fail to make truthful reporting a Bayes--Nash equilibrium once LLMs are available, while Speakeasy restores it for all of them.
Abstract: The rapid progress of large language models (LLMs) is increasingly constrained by memory and deployment costs, motivating compression methods for practical deployment. Many state-of-the-art compression pipelines leverage the low-rank structure of trained weight matrices, a phenomenon often associated with the properties of popular optimizers such as Adam. In this context, Muon is a recently proposed optimizer that improves LLM pretraining via full-rank update steps, but its induced weight-space structure has not been characterized yet. In this work, we report a surprising empirical finding: despite imposing full-rank updates, Muon-trained models exhibit pronounced low-rank structure in their weight matrices and are readily compressible under standard pipelines. Motivated by this insight, we propose , which augments Muon with a nuclear-norm constraint on the update direction, further constraining the learned weights toward low-rank structure. Across billion-parameter-scale models, we show that NuMuon increases weight compressibility and improves post-compression model quality under state-of-the-art LLM compression pipelines while retaining Muon's convergence behavior.
PaperID: 2386, Poster
Authors: Xin Wang, R.Tyrrell Rockafellar, Xuegang Ban
Abstract: Constrained learning has been increasingly applied to various domains to ensure explicit feasibility requirements due to fairness, safety, robustness, regularization, and physics or logic constraints. Understanding how training samples influence the solution (e.g., learned parameters) of constrained learning is crucial for interpretability and robustness. The classical influence function may become unreliable in constrained settings: data perturbations can reshape both the objective and the feasible region, leading to estimates that violate feasibility. In response, we propose the Directional Influence Function (DIF), a new estimator that explicitly incorporates the constraints into influence estimation. DIF formulates the optimality conditions of constrained learning as a variational inequality (VI) and analyzes how perturbing training data affects this VI. We validate DIF in constrained linear regression and demonstrate that it recovers leave-one-out retraining results, whereas IF and penalty-based IF exhibit significant bias. We further apply DIF to fairness-constrained CNNs, where DIF accurately predicts test loss changes under data removal and aligns closely with actual retraining. Our results establish DIF as an efficient and reliable tool for data attribution in constrained learning.
Authors: Suresh Raghu, Satwik Pandey, Shashwat Pandey
Abstract: Language-model agents increasingly emit uncertainty signals throughout a trajectory, but existing agentic UQ evaluations often conflate ranking usefulness with probabilistic truthfulness. AUROC, AUPRC, risk-coverage, Trajectory ECE, and scalarized trajectory scores evaluate discrimination, binwise calibration, or collapsed summaries, but do not strictly elicit the full prefix-conditioned success-probability trace q_t=\mathbbP^\pi(Y=1\mid\mathcalH_t). Building on prequential proper scoring, we introduce the Trajectory Proper Score (TPS), a predictor-agnostic family of strictly proper trajectory-level scoring rules for any per-step uncertainty signal calibrated into a probability of eventual success. We prove that TPS strictly elicits the success-probability process under complete observation, within the chosen score family and weight schedule. We extend the construction to administratively censored trajectories by projecting the complete-data score onto the observable stopped prefix, yielding an exact q_Z-weighted reduced score and a tractable approximation when q_Z is unestimated. We further show that common trajectory evaluators target weaker objects than the full prefix-conditioned probability process: Trajectory ECE is resolution-blind, while scalarized Trajectory Brier elicits only the collapsed scalar, not the full trace. Experiments on StrategyQA, Tau2-Bench, HotpotQA, and WebShop show that these theoretical distinctions are operationally visible: probability recalibration can substantially change TPS while leaving rank metrics nearly unchanged, and the tractable censored approximation can change the verdict relative to complete-only evaluation.
Authors: Roman Kniazev, Nathanaël Fijalkow
Abstract: Do transformers, when trained on sequential reasoning traces, build internal models of the underlying task? And if so, does the structure of those internal representations mirror the structure of the domain? We train an 8-layer transformer on Sudoku solving traces and perform a mechanistic analysis of its internal computation. We establish two results. First, the model builds a substructure world model: it does not represent the board state cell by cell, as a human analyst would expect, but organizes information around the rows, columns, and boxes that Sudoku's constraints act on. Second, we identify a naked-single circuit: a small set of dedicated neurons in the final MLP layer, each individually detecting when exactly one digit remains possible for a specific cell, and reliably promoting that digit. These findings show that the geometry of an emergent world model is shaped by the constraint algebra of the domain, not its surface presentation, and that the resulting decision circuit is sparse, monosemantic, and fully interpretable. More broadly, they demonstrate that mechanistic interpretability tools can recover an end-to-end algorithmic account of how a transformer solves a combinatorial reasoning task.
PaperID: 2389, Poster
Abstract: Real-world scenes are compositional: bricks, blades of grass, pebbles, and tree leaves recur across human-built and natural environments. Existing neural scene representations model these elements independently. Most 3D Gaussian Splatting and follow-up abstraction and compression methods treat each element as unique, fitting millions of independent Gaussians per scene. Prior methods like Splat and Replace fit template objects, but they require mostly manual selection of repeated elements. As a result, these representations store redundant parameters and provide weak manipulation handles for downstream tasks. We introduce SCION, a hierarchical compositional scene representation that replaces independent Gaussians with a compact vocabulary of reusable primitives and lightweight world-space instances that place transformed copies throughout the scene. We fit this representation to multi-view captures via a joint optimization over discrete and continuous scene parameters, combining two-level densification over splats and instances with an adversarial loss that preserves detail across shared primitives. The recovered structure yields a compact, controllable representation while maintaining high quality even at 1.2 MB. SCION achieves rate-distortion favorable to existing Gaussian compression methods, and it enables instance-level scene editing and animation without retraining. Our results show that neural scene representations need not memorize scenes as independent primitives; they can discover reusable parts.
Abstract: Large language models often produce human-like moral judgments, but it is unclear whether this reflects an internal conceptual structure or superficial ``moral mimicry.'' Using Moral Foundations Theory (MFT) as an analytic framework, we study how moral foundations are encoded, organized, and expressed across 14 base and instruction-tuned LLMs spanning four model families (Llama, Qwen2.5, Qwen3-MoE, Mistral) and scales from 7B to 70B. We employ a multi-level approach combining (i) layer-wise analysis of MFT concept representations and their alignment with human moral perceptions, (ii) pretrained sparse autoencoders (SAEs) over the residual stream to identify sparse features that support moral concepts, and (iii) causal steering interventions using dense MFT vectors and sparse SAE features. We find that models represent and distinguish moral foundations in a manner that aligns with human judgments, and that this moral geometry naturally emerges from pretraining and is selectively rewired by post-training. At a finer scale, SAE features show clear semantic links to specific foundations, suggesting partially disentangled mechanisms within shared representations. Finally, steering along either dense vectors or sparse features produces predictable shifts in foundation-relevant behavior, demonstrating a causal connection between internal representations and moral outputs. Together, our results provide mechanistic evidence that moral concepts in LLMs are distributed, layered, and partly disentangled, suggesting that pluralistic moral structure can emerge as a latent pattern from the statistical regularities of language alone.
PaperID: 2391, Poster
Abstract: Rasterization dictates the interactive budget of 3D Gaussian Splatting (3DGS). However, the comparative speed of modern CUDA rasterizers is typically evaluated on a narrow canonical benchmark, masking severe regime-dependent performance reversals. This limited scope hides three critical kernel-level bottlenecks: warp lane underutilization, exposed global-memory latency during pixel shading, and terminal-tail penalties from hardware thread scheduling. In this paper, we present PolySplat, a workload-regime-aware 3DGS rasterizer designed to overcome these inefficiencies. PolySplat introduces a warp-saturating tile-key emitter with adaptive three-way dispatch, asynchronous shared-memory staging in the render kernel, and a persistent kernel architecture with centralized atomic dispatch to strictly bound tail latency. To rigorously validate our system, we introduce an extended 76-target benchmark that achieves comprehensive workload-regime coverage by stratifying targets across extreme Gaussian counts (up to 56.5M), high resolutions (up to 9K), and six diverse scene categories. Evaluated on this comprehensive suite, PolySplat achieves dataset-balanced geometric-mean speedups of 1.26x to 6.48x over state-of-the-art rasterizers at lossless visual quality. Notably, under sustained 60Hz interaction at extreme resolutions, PolySplat strictly meets 1-vsync display deadlines where existing renderers suffer from massive queue divergence and input-to-photon lag.
Abstract: Vision Transformers (ViTs) based vision foundation models (VFMs) have achieved remarkable performance across diverse vision tasks, but suffer from quadratic complexity that limits scalability to long sequences. Existing linear attention approaches for ViTs are typically trained from scratch, requiring substantial computational resources, while linearization-based methods developed for large language model decoders do not transfer well to ViTs. To address these challenges, we propose ViT-AdaLA, a novel framework for effectively adapting and transferring prior knowledge from VFMs to linear attention ViTs. ViT-AdaLA consists of three stages: attention alignment, feature alignment, and supervised fine-tuning. In the attention alignment stage, we align vanilla linear attention with the original softmax-based attention in each block to approximate the behavior of softmax attention. However, residual approximation errors inevitably accumulate across layers. We mitigate this by fine-tuning the linearized ViT to align its final-layer features with a frozen softmax VFM teacher. Finally, the adapted prior knowledge is transferred to downstream tasks through supervised fine-tuning. Extensive experiments on classification and segmentation tasks demonstrate the effectiveness and generality of ViT-AdaLA over various state-of-the-art linear attention counterpart.
PaperID: 2393, Poster
Abstract: Membership inference attacks (MIAs) aim to determine whether a sample was used during the training of a target model. The most effective MIAs rely on reference (or shadow) models to estimate target-conditional score distributions, but training such ensembles is computationally prohibitive for modern large-scale Vision--Language Models (VLMs). In this work, we propose a scalable alternative that replaces explicit reference-model training with a Bayesian approximation over low-rank adaptation parameters. Leveraging parameter-efficient fine-tuning, we model uncertainty only within the LoRA subspace and construct a posterior from a single training run, from which we sample a diverse set of virtual reference models. We instantiate this approach using stochastic weight averaging Gaussian (SWAG), enabling efficient approximation of target-conditional score distributions at a fraction of the cost of conventional shadow-model ensembles. We evaluate the resulting attack on three downstream-adapted VLMs across two privacy-sensitive visual question answering tasks, namely Document VQA and medical VQA. To isolate the effect of pretraining knowledge, we additionally introduce a controlled synthetic dataset. Our approach outperforms standard reference-based attacks while requiring only a single trained reference model, demonstrating that accurate and scalable membership inference is feasible even for large VLMs.
Authors:
Arash Akbari, Arman Akbari, Masih Eskandar, Qitao Tan, Yixiao Chen, Jingwu Luo, Bertha Pangaribuan, Liyun Zhang, Jennifer Dy, Geng Yuan, Xue Lin, Gaowen Liu, Stratis Ioannidis, Yanzhi WangAbstract: Vision-Language-Action (VLA) models exhibit remarkable action generation for embodied intelligence, but their heavy compute and memory demands strain deployment on edge robotic platforms. Aggressive weight quantization (e.g., sub-4-bit) is the natural lever, yet existing post-training quantization (PTQ) methods suffer severe performance degradation in this regime. To address this, we introduce ActQuant, an action-guided mixed-precision PTQ framework that operates in two stages: (1) an inter-tensor allocator that assigns each weight matrix a single bit-width based on how much it contributes to predicting the agent's actions; (2) an intra-tensor scale optimizer tunes per-block quantization scales using action-aware curvature, so that dynamic range is concentrated on the weights most influential for control. We evaluate ActQuant both in simulation and on a real-world 6-DoF UR3 arm. On the LIBERO benchmark, ActQuant achieves state-of-the-art success rates across the sub-4-bit regime, retaining 94.95% on OpenVLA-OFT and 94.8% on \pi_0.5 at 3.0 bits-per-weight (vs. 96.9% and 97.0% for the half-precision baselines, respectively), while reducing backbone memory by up to 5.3×; on the physical robot, ActQuant successfully maintains the baseline's success rate. To our knowledge, ActQuant is the first PTQ approach that can quantize VLA models below four bits while retaining practical task success with very limited calibration data.
PaperID: 2395, Poster
Authors: Yubiao Zhao, Shicheng Song, Yusen Ye, Huan Liu, Juan Liu, Lihua Zhang
Abstract: Mapping fine-scale 3D chromatin landscapes is critical for understanding gene regulation, yet the resolution of Hi-C data remains fundamentally constrained by both physical bottlenecks and sequencing costs. Although deep learning has shown promise in super-resolving Hi-C data, existing methods suffer from absolute coordinate loss during localized patching and a lack of structural and epigenomic priors in unconditioned generation. We propose a novel epigenomics-guided flow matching framework EpiFlow for 3D genome super-resolution. EpiFlow introduces two key innovations: (1) an EpiCond module that deterministically maps 1D epigenomic features to 2D spatial constraints via orthogonal expansion with built-in positional grounding, avoiding quadratic attention complexity; and (2) a Toeplitz-regularized classifier-free guidance mechanism that enforces distance-decay priors using a learnable symmetric prototype matrix. Extensive evaluation shows that our method achieves state-of-the-art cross-cell-type loop detection, improving recall by 2.8× over raw Hi-C on unseen H1-hESC (Loop F_1 of 0.325 vs. 0.199) and outperforming all baselines on same-cell HFFc6 (F_1=0.520). Aggregate Peak Analysis (APA) confirms biological validity, and strong generalization to a held-out cell line demonstrates that the model learns transferable principles of 3D genome folding rather than cell-specific biases.
PaperID: 2396, Poster
Abstract: \emphScaling has been the central driver of progress in modern AI, yet data curation remains an exception: it still relies on manual heuristics rather than systematic, scalable methods. In practice, data recipes (e.g., filter thresholds, deduplication strength, domain mixing ratios) are determined through many small-scale ablation training runs. This work demonstrates that the performance of the best-discovered recipe improves \emphpredictably--as a \emphpower law--in the number of such ablation runs. Theoretically, the power-law form arises naturally from regret bounds in zero-order optimization theory. Empirically, we validate it across diverse training stages, including pretraining, supervised fine-tuning, preference optimization, and reinforcement learning. Building on this finding, we derive a principled approach to \emphdata quality--quantity tradeoff under a fixed compute budget, yielding a Pareto frontier that dominates standard heuristic pipelines (e.g., DCLM). Together, this work provides a general framework for scaling data-curation compute, establishing it as a new axis alongside the well-studied pretraining and test-time scaling.
PaperID: 2397, Poster
Abstract: Evaluator-guided reasoning systems face a local control problem: given a current answer and evaluator evidence, should the system trust, fallback, or buy more inference? We formulate this as a compute-aware three-action controller over , where each action is scored by predicted eventual success minus explicit compute cost. The central quantity is : the success probability induced by trusting now, falling back now, or expanding once and continuing. We show that local action selection is stable when confidence error is smaller than the compute-adjusted action gap, and that regret accumulates only along the realized expansion path. Guided by this analysis, we learn one confidence head per action, calibrate the heads on held-out data, and penalize locally uncertain actions. Across verifier-guided width expansion on MATH and AIME25 and judge-guided revision expansion on a deterministic-answer General Reasoning benchmark, calibrated control improves the utility--compute frontier relative to strong fixed and adaptive baselines. The same pattern persists under policy-induced expansion targets, same-feature budget routers, fallback removal, stronger learned controllers trained on the same features, and moderate evaluator mismatch after refitting and recalibration. When compute is free, high-budget policies can remain competitive; the gain here is compute-efficient allocation, not a uniformly stronger reasoner.
PaperID: 2398, Poster
Abstract: Video understanding models often rely on motion, but faithful explanations must reveal not only whether motion matters, but also which source of motion supports a prediction. We identify motion-source under-specification as a key limitation of existing video XAI: region-centric explanations are source-ambiguous, while concept-based explanations are source-incomplete when their motion vocabulary is restricted to single-person pose sequences. To bridge this gap, we propose TraCE-Trajectory-Aware Concepts for Explainable Video Understanding-a concept-based framework that elevates visual trajectories to first-class explanatory concepts. By clustering optical-flow trajectories, TraCE discovers source-flexible motion concepts that capture dynamic evidence from people, objects, and interactions, while object and scene concepts account for static context. TraCE applies the same typed trajectory/object/scene concept interface to both action recognition and video question answering (VQA), enabling concept-level explanations beyond a single task. We evaluate TraCE across six benchmarks: NTU RGB+D Mutual, Chiral SSv2, UCF-101, KTH, TOMATO, and STAR. Experiments show that TraCE provides trajectory-aware explanations of dynamic evidence with minimal performance trade-off, trailing non-interpretable baselines by only 0.9 points on action recognition and 0.1 points on VQA on average. User studies, faithfulness analyses, temporal sensitivity tests, and concept-intervention results further show that TraCE produces interpretable, faithful, and practically useful explanations.
Abstract: We study Stackelberg (leader--follower) control of network parameters (tolls, capacities, incentives) in combinatorial congestion games, where selfish users choose \emphdiscrete routes (or other combinatorial strategies) and settle at a congestion equilibrium. The leader minimizes a system-level objective (e.g., total travel time) evaluated at equilibrium, but this objective can be nonsmooth because the set of used strategies can change abruptly. We propose \textscZeroth-order Stackelberg (\textscZOS), which couples a projection-free Frank--Wolfe equilibrium solver with a zeroth-order outer update, avoiding differentiation through equilibria. We prove convergence to generalized Goldstein stationary points of the true equilibrium objective, with explicit dependence on the equilibrium approximation error, and analyze subsampled oracles: if the sampled candidate set contains an LMO minimizer with probability \kappa_m, then the Frank--Wolfe error decays as \mathcalO(1/(\kappa_m T)). We also propose stratified sampling as a practical way to avoid a vanishing \kappa_m when LMO minimizers concentrate in strata defined by simple features such as path length. In public road-network experiments covering large strategy spaces requiring subsampled oracles, \textscZOS reaches small follower-equilibrium gaps and final social costs comparable to differentiation-based methods, while reducing runtime per outer iteration by 20--1000× and using much less peak memory.
PaperID: 2400, Poster
Abstract: Model-assisted semi-supervised learning offers a powerful paradigm for enhancing the efficiency of statistical estimation and inference by leveraging black-box predictions on unlabeled data. However, the validity of these methods typically relies on the strict assumption that labeled and unlabeled populations share identical distributional characteristics. Consequently, state-of-the-art approaches, such as prediction-powered inference, become fragile when this assumption is violated. To address the fragility, we propose a robust framework for semi-supervised learning that remains reliable under distributional heterogeneity. By embedding a robust statistical calibrator into the prediction-rectification mechanism, the approach effectively reduces the bias arising from shifted or corrupted unlabeled samples. To maximize data utility, we introduce an adaptive cross-validation procedure to select the optimal calibrator, ensuring a reliable trade-off between statistical efficiency and robustness. Theoretical analysis confirms the consistency of the proposed estimator under mild conditions, while empirical results demonstrate its significant superiority over conventional baselines in heterogeneous environments.
Authors: Haoshu Xu, Hongzhe Li
Abstract: We address two open questions in streaming PCA via Oja's algorithm: sharp operator-norm convergence for general rank under sub-Gaussian data, and distributional inference for the resulting subspace estimator. Existing convergence analyses, even in the rank-one case, leave non-vanishing remainder terms that prevent adaptation to a vanishing tail spectrum, while existing distributional results are confined to the rank-one case. Our convergence theory removes these remainder terms and yields a sharp rate. In the dense-tail spiked covariance regime, this rate matches the minimax rate up to logarithmic factors. More generally, we prove a matching lower bound, up to logarithmic factors, across both dense-tail and sparse-tail regimes under a mild nondegeneracy condition. The analysis yields a linearization of Oja's iterates, which in turn enables a high-dimensional Gaussian approximation for the general-rank subspace estimation error with an explicit limiting covariance. We also establish a row-wise Berry--Esseen theorem for the aligned difference, recovering prior rank-one results as special cases. For practical inference, we develop an online multiplier bootstrap algorithm and prove its consistency. Beyond streaming PCA, our techniques contribute to Gaussian approximation and bootstrap inference for nonconvex stochastic approximation.
PaperID: 2402, Poster
Authors:
Sahar Almahfouz Nasser, Juan Francisco Pesantez Borja, Jincheng Liu, Sandeep Manandhar, Shikhar Shiromani, Mohammad T Hasan, Zenghan Wang, Suman Ghosh, Jinchu Li, Xuejian Xu, Aniket R Iyer, Naoto Tokuyama, Twisha Shah, Tilak Pathak, Soundharya Kumaresan, Yohei Abe, Himanshu Maurya, Anant MadabhushiAbstract: Engineered image-based biomarkers offer a clinically interpretable alternative to black-box AI in computational pathology, yet their discovery remains largely intuition-driven, guided by fragmented literature rather than rigorous biological validation. We introduce SAGE (Structured Agentic system for hypothesis Generation and Evaluation), a multi-agent framework that grounds biomarker discovery in biological evidence through three mechanisms: (i) knowledge-graph-anchored hypothesis generation via multi-path ontological reasoning, (ii) a debate-based multi-agent novelty assessment that stress-tests candidate biomarkers against existing literature, and (iii) an end-to-end automated validation pipeline that translates hypotheses directly into executable analyses on multimodal pathology datasets. Together, these components shift biomarker discovery from an intuition-driven, literature-browsing exercise into a structured, traceable reasoning process that clinicians and researchers can inspect, trust, and build upon.
PaperID: 2403, Poster
Abstract: Large-scale video generation models have made remarkable progress in semantic consistency and visual quality, producing videos that are increasingly coherent and visually convincing. Nevertheless, the dynamics induced by pixel-level fitting do not naturally accommodate the regularities that govern real-world motion and interaction, resulting in persistent shortcomings in physical plausibility. To address this limitation, we propose (Physics-Informed Latent Alignment), a framework that injects physics-structured latent guidance into the frozen flow-matching dynamics of pretrained video models. Specifically, PILA first employs anchored field estimation to map frozen-generator latents into an operational physical attribute bank organized by field-proxy slots, using observable motion as a kinematic anchor for constructing less directly observed proxies. To handle the heterogeneity of real-world dynamics, PILA adopts a mixture-of-experts design over physical categories. Label-prior masked expert routing selects category-specific operator experts, whose refinements are regularized by operational residuals abstracted from physical relations. Finally, the refined proxies are fused into the physical attribute bank and decoded into a correction to the flow-matching vector field, injecting physics-aware guidance while preserving the visual prior of the pretrained backbone. With staged adapter training on Wan 2.1-1.3B and direct transfer of the learned adapter to Wan 2.2-14B, PILA achieves state-of-the-art results on VBench-2.0, VideoPhy-2, and PhyGenBench in both visual quality and benchmark-measured physical plausibility.
Authors:
Wenjie Zheng, Haoji Hu, Jiali Lu, Xingze Zou, Jing WangAbstract: Dataset distillation (DD) aims to compress large-scale datasets into compact synthetic sets while preserving training efficacy. However, existing studies mainly focus on image classification, leaving dense prediction tasks such as semantic segmentation largely underexplored. In this work, we identify three key challenges for segmentation DD: (i) long-tailed class imbalance, (ii) the need for strict pixel-wise alignment between images and dense labels, and (iii) the high computational cost of optimizing high-resolution data with complex models. To address these challenges, we propose \mathbfD^3S^2, a Diffusion-guided Dataset Distillation framework for Semantic Segmentation. Our method adopts a two-stage design. In Class-Balanced Mask Selection, we construct a representative mask set via a greedy strategy that prioritizes underrepresented classes. In Diffusion-Guided Image Synthesis, we employ a pretrained layout-to-image diffusion model to generate images conditioned on the selected masks, naturally ensuring spatial alignment. To further enhance the training utility of synthesized data, we introduce guided diffusion sampling with two complementary objectives: a segmentation-consistency loss for pixel-level alignment, and a class-wise feature matching loss for aligning per-class feature statistics across layers. Extensive experiments demonstrate the superiority of D^3S^2. Notably, at an extremely compression rate of 1%, our method achieves 24.99% and 35.49% mIoU on ADE20K and COCO-Stuff with Mask2Former (Swin-S), outperforming random selection by 9.34% and 5.70%, respectively.
PaperID: 2405, Poster
Abstract: Safe robot learning often involves multiple heterogeneous safety constraints that cannot always be satisfied simultaneously. Existing neural safety layers typically treat multi-constraint safety as a numerical optimization problem, enforcing all constraints through a single QP-based projection or relaxing conflicts with slack variables. This hides the semantic question of which constraints may be sacrificed under conflict inside solver geometry, penalty weights, or a fixed total hierarchy. We propose PoSafeNet, a poset-structured composable safety layer that makes these conflict semantics explicit. PoSafeNet encodes admissible safety override relations as a partial order and realizes each admissible execution by composing closed-form projections onto CBF-induced halfspaces. Across multi-obstacle navigation, constrained manipulation, and vision-based autonomous driving, PoSafeNet improves operational feasibility, computational efficiency, and task performance over dQP-based, slack-based, and hierarchical safety layers.
Abstract: reasoning that evaluates competing hypotheses. However, many existing methods apply a uniform reasoning strategy across queries, leading to unnecessary computation on simple tasks and insufficient reasoning on complex ones. We introduce , a training framework that learns to select the appropriate reasoning mode for each query using reinforcement learning. During training, the model generates responses under all three modes for the same prompt, enabling direct comparison. A composite reward compares these responses to favor the most effective one based on correctness, grounding, and cost, while discouraging unsupported or misaligned deliberation. This yields a supervision signal for learning adaptive reasoning without per-query annotations. Video-FLAIR improves accuracy over the Qwen2.5-VL base model by
Abstract: Cooperative multi-agent reinforcement learning (MARL) under sparse rewards remains fundamentally challenging because agents often fail to concentrate their influence, leading to insufficiently coordinated exploration. To address this, we propose the Focusing Influence Mechanism (FIM), a framework that encourages agents to focus their influence on under-explored parts of the state space through an entropy-based criterion, while leveraging eligibility traces to enable multiple agents to consistently align and sustain their influence on the same parts of the state space when beneficial, thereby promoting coordinated and persistent joint behavior. By emphasizing under-explored regions of the state space, FIM facilitates more efficient and structured exploration even under extremely sparse rewards. Across diverse MARL benchmarks, FIM consistently improves cooperative performance over strong baselines.
PaperID: 2408, Poster
Abstract: Transformer models can adapt to new tasks from input--output examples at inference time, a capability known as in-context learning (ICL). A common interpretation is that transformers implement implicit learning algorithms in their forward pass, using the context to infer the task and predict the query. We test this interpretation beyond the training distribution in controlled regression settings where classical estimators can generalize exactly. Transformers achieve near-perfect in-distribution ICL, but fail sharply under scale and support shifts, exhibiting a cliff-shaped collapse that is absent from least-squares and kernel baselines. This failure is robust across polynomial degrees and a range of training interventions. We identify two architectural mechanisms that help explain this behavior. First, final normalization imposes a readout-side scale constraint, causing predictions to saturate at fixed boundary values. Removing this constraint eliminates hard saturation, but does not restore reliable ICL under coefficient or label scaling. Second, attention can become a context bottleneck: under distribution shift, softmax attention may concentrate on a small part of the prompt, limiting the context-dependent adaptation required for ICL. These results show that strong in-distribution ICL does not imply distribution-independent ICL, and suggest caution when using ICL instead of explicit retraining or adaptation under distribution shift.
PaperID: 2409, Poster
Authors: Jinbo Yang, Shikai Jing, Mingyue Yuan
Abstract: Text-to-CAD generation aims to make CAD modeling more accessible by generating executable CAD programs from natural-language descriptions. However, practical CAD workflows require more than executable geometry: generated programs should also expose reusable structure that supports later modification. We introduce intent-oriented text-to-CAD generation, where design intent is expressed as recoverable code-level signals: semantic features, reusable parameters, and lightweight parametric relations. Rather than changing the model architecture or decoding procedure, we investigate how supervision representation granularity affects both CAD generation and zero-shot intent-preserving editing. To enable this study, we construct IntentCAD-100K, a 100K paired text--code dataset built by converting low-level CAD construction traces into executable CadQuery programs. We further construct IntentCAD-Edit-1K to evaluate zero-shot intent-preserving editing from the same representation perspective. Intent2CAD uses a two-stage conservative lifting pipeline: semantic lifting rewrites reliable primitive groups into feature-level operations, while parametric lifting exposes lightweight parametric relations. Under controlled fine-tuning, semantic lifting mainly improves generation reliability and feature recovery, reducing IR from 2.39% to 0.22% and achieving 99.96% API Recovery, while parametric lifting makes reusable relation structure explicit and further improves relation-aware zero-shot editing.
PaperID: 2410, Poster
Abstract: As LLM agents advance rapidly in autonomy, the risk of hidden objectives that bypass safety constraints becomes increasingly plausible. Chain-of-thought (CoT) monitoring offers a key control mechanism by reading the agent's reasoning and flagging misalignment before it materializes into action, yet its reliability depends on the realism of the red-team it is stress-tested against. Existing red-teaming efforts apply uniform attack templates that treat the side task as an exogenous insertion, yielding abrupt anomalies any reasonable monitor can flag while leaving the subtler, contextually grounded evasions of a capable agent largely unprobed. We therefore ask: How robust is CoT monitoring against agents that strategically conceal misalignment? To answer this, we introduce TraceMRT (Trace-level Monitor Red-Teaming), an automated red-teaming framework that reframes the side task as a coherent sub-goal of the primary task and searches trace-shaping strategies whose induced trajectories complete the misaligned objective while reading as natural execution. Through black-box attacks on state-of-the-art monitors (e.g., GPT-5.2, and Claude-4.5), we reveal substantial vulnerabilities even in top models. To better counter these vulnerabilities, we further improve monitoring scaffolding with Bayesian Trajectory Monitoring, integrating global context and local evidence to better detect covert misalignment. Empirically, the method significantly consolidates the monitor with 0.91 AUC and 0.74 TPR@FPR=0.05.
PaperID: 2411, Poster
Abstract: 3D Gaussian Splatting (3DGS) has emerged as a promising method that enables fast and high-quality novel-view rendering. However, its large parameter size remains a bottleneck for storage and transmission. We propose RacerGS, a compact 3DGS compression framework that achieves scene reconstruction at significantly reduced storage sizes with preserved fidelity. To obtain compact and spatially coherent features, we design a hash projection model that projects the anchor hash encodings into compact context features. Additionally, we introduce a FiLM-modulated autoregressive GRU context model that utilizes causal dependencies among anchor feature groups and uses recurrent hidden states to refine entropy-model parameters, improving bitrate efficiency. Furthermore, we adaptively entropy-code quantized residuals of anchor attributes around their context-predicted means, reducing symbol entropy by leveraging spatial consistency. Overall, RacerGS achieves more than \mathbf150\boldsymbol× compression over vanilla 3DGS and \mathbf28\boldsymbol× over Scaffold-GS, while providing comparable rendering quality. Extensive experiments across five datasets demonstrate our superiority over prior 3DGS compression methods in both storage reduction and rate-distortion performance.
Authors: Ahmer Raza, Dane H Smith
Abstract: Multivariate Hawkes processes are a widely used class of self-exciting point processes, but maximum likelihood estimation naively scales as O(N^2) in the number of events. The canonical linear exponential Hawkes process admits a faster O(N) recurrence, but prior work evaluates this recurrence sequentially, without exploiting parallelization on modern GPUs. We show that the Hawkes process intensity can be expressed as a product of sparse transition matrices admitting a linear-time associative multiply, enabling computation via a parallel prefix scan. This yields a simple yet massively parallelizable algorithm for maximum likelihood estimation of linear exponential Hawkes processes. Our method reduces the computational complexity to approximately O(N/P) with P parallel processors, and naturally yields a batching scheme to maintain constant memory usage, avoiding GPU memory constraints. Importantly, it computes the exact likelihood without any additional assumptions or approximations, preserving the simplicity and interpretability of the model. We demonstrate orders-of-magnitude speedups on simulated and real datasets, scaling to thousands of nodes and tens of millions of events, substantially beyond scales reported in prior work. We provide an open-source PyTorch library implementing our optimizations.
PaperID: 2413, Poster
Abstract: Machine unlearning aims to erase the influence of a designated forget set from a trained model while preserving strong utility on the retain set. Approximate unlearning methods efficiently achieve this goal without full retraining, but their effectiveness is primarily assessed using dataset-level aggregate metrics. As a result, prior work does not explicitly evaluate whether individual samples are unlearned correctly, nor does it provide mechanisms to enforce such per-sample unlearning behavior. In this work, we focus on per-sample unlearning effectiveness. We introduce SALA (Sample-grained Likelihoods ALignment), a plug-in optimization framework that enables approximate unlearning methods to operate at the level of individual samples by steering each example toward its correct target distribution. To quantitatively assess this behavior, we further propose the sample-grained distribution gap, a metric that captures per-sample discrepancies between approximate unlearning and exact retraining output distributions. We provide theoretical guarantees showing that SALA contracts this gap and yields a bounded per-sample discrepancy to retraining. To make this framework practical, we develop a low-cost estimation method for the member and non-member target distributions. Extensive experiments on classification and generation tasks demonstrate that SALA reduces the sample-grained distribution gap by up to 86.67% when paired with existing approximate unlearning baselines, while also improving run-to-run robustness.
PaperID: 2414, Poster
Abstract: Image vectorization reconstructs raster images as compact scalable vector graphics (SVG) representations. Most of existing vectorization methods optimize for pixel-level rendering fidelity, where the reconstructed SVGs often lack alignment with human-perceived object-part hierarchies, making them difficult to manipulate. To address this, this work for the first time studies the problems of hierarchical image vectorization and SVG semantic labeling. We propose Lang-SVG, a tree-structured SVG representation that encodes SVG primitives with language tags and parent-child hierarchy. Lang-SVG novelly incorporates granularity-controllable segmentation priors and vision-language priors into the optimization-based vectorization pipeline. It introduces a Tree-Guided Pruning and Merging module to reduce redundant multi-granularity SVG primitives into coherent semantic structures, and an L1-SAM Prompter to recover missing regions. It also includes a new SVG semantic tagging method that integrates SVG tree context into multimodal large language models (MLLMs) to assign language tags to SVG primitives. We also propose a new evaluation protocol for image vectorization, measuring structural quality and semantic part alignment beyond reconstruction fidelity. Experiments demonstrate that Lang-SVG achieves state-of-the-art performance in both rendering and structural quality.
PaperID: 2415, Poster
Abstract: Zeroth-order optimization (ZOO) is a promising approach for memory-efficient fine-tuning of Large Language Models (LLMs). However, scaling ZOO to larger models on a single GPU necessitates aggressive quantization, which often destabilizes training. A primary cause of this instability is that existing methods treat strictly positive quantization scale parameters as Euclidean variables, potentially leading to boundary violations and optimization difficulties. To address this issue, we propose Hyper-Octant Zeroth-order Optimization (HoZO), a geometry-aware framework that formulates fine-tuning as Riemannian optimization on the positive orthant manifold. By constructing a novel geodesically complete metric and deriving the closed-form exponential map, HoZO inherently enforces positivity constraints without bias-inducing projections. On the theoretical side, we prove that HoZO achieves the optimal oracle complexity for zeroth-order methods. On the empirical side, HoZO outperforms baselines across six downstream tasks and three model sizes. Notably, it achieves a 34× memory reduction compared to first-order fine-tuning, enabling the stable fine-tuning of a 4-bit quantized Llama-2-70b model on a single 48GB GPU.
PaperID: 2416, Poster
Authors:
Dengwang Tang, Dongze Ye, Rahul Jain, Ashutosh Nayyar, Pierluigi NuzzoAbstract: Learning in POMDPs is known to be significantly harder than in MDPs. In this paper, we consider the online learning problem for episodic POMDPs with unknown transition and observation models. We propose a Posterior Sampling-based reinforcement learning algorithm for POMDPs (PS4POMDP), which is much simpler and more implementable compared to state-of-the-art optimism-based online learning algorithms for POMDPs. We show that the Bayesian regret of the proposed algorithm scales as the square root of the number of episodes and is polynomial in the other parameters. In a general setting, the regret scales exponentially in the horizon length H, and we show that this is inevitable by providing a lower bound. However, when the POMDP is undercomplete and weakly revealing (a common assumption in the recent literature), we establish a polynomial Bayesian regret bound. We finally propose a posterior sampling algorithm for multi-agent POMDPs, and show it too has sublinear regret.
Abstract: To improve safety in Large Language Models (LLMs) we can either perform post-training alignment or exploit refusal directions in the activation space. Both strategies are less feasible in Multimodal LLMs (MLLMs) as they require unsafe multimodal data, harder to collect than their unimodal counterpart. In this work, we relax this constraint and investigate whether textual refusal directions, extracted directly from the LLM backbone, generalize across modalities (i.e., image, video). Preliminary findings confirm this ability, though effectiveness is conditioned by layer selection, steering strength, and cross-modal alignment, with the latter causing safe multimodal inputs to be spuriously steered toward refusal. Building on this, we introduce Modality-Agnostic Refusal Steering (MARS), a light-weight training-free approach that injects multimodal safety without the need for multimodal safety data. MARS corrects modality misalignment via activation re-centering, adaptively scales steering strength within a geometrically defined trust region, and selects the optimal intervention layer, operating at the first generated token. Evaluated on five SOTA MLLMs across safety, utility, and video jailbreak benchmarks, MARS achieves consistent safety gains while preserving utility. These results reveal that safety-relevant structure is shared across modalities and that textual refusal directions are a powerful and underexplored foundation for multimodal alignment.
PaperID: 2418, Poster
Authors: Matthew Peroni, Franck Le, Vadim Sheinin
Abstract: The development of tabular foundation models (TFMs) has accelerated in recent years, showing strong potential to outperform traditional ML methods for structured data. A key finding is that TFMs can be pretrained entirely on synthetic datasets, opening opportunities to design data generators that encourage desirable model properties. Prior work has mainly focused on crafting high-quality priors over generators to improve overall pretraining performance. Our insight is that parameterizing the generator distribution enables an adversarial, distributional robustness perspective: during training, we can adapt the generator to emphasize datasets that are particularly challenging for the model, which we formalize by introducing an optimality gap measure. Further, we develop a method to align the synthetic data generation with real-world datasets from a given domain, constraining the adversary to generate "realistic'" data. Together, these algorithms comprise the Robust and Domain-Adapted Tabular Foundation Models (RAD-TFM) pipeline, a model-agnostic adversarial training framework. Applied to the TabPFN V2 classifier, RAD-TFM improves performance across 6 diverse tabular benchmarks, with up to a 11% increase in mean normalized AUC over the original TabPFN and other baseline algorithms, with only 100k additional training datasets, less than 0.1% of the original pretraining data. These results highlight a promising new direction for targeted adversarial training and fine-tuning of TFMs using synthetic data alone.
PaperID: 2419, Poster
Abstract: Flow Matching (FM) provides an efficient formulation for continuous-time generative modeling via straight optimal transport (OT) trajectories. However, during accelerated sampling, these models face a fundamental variance-fidelity dilemma. Under few-step deterministic Ordinary Differential Equation (ODE) integration, the rigid OT trajectories suffer from mean concentration, leading to severe variance collapse and the loss of high-frequency details. Conversely, traditional Stochastic Differential Equations (SDEs) inject blind, exogenous Brownian noise to recover variance, but this isotropic exploration incurs inherent numerical attrition over long integration paths, bottlenecking peak fidelity. To resolve this, we introduce Self-Aligned Drifting (SADrift) to mitigate these challenges without external priors. SADrift is an inference-time framework that augments flow matching inference with self-interacting dynamics, without additional training. By maintaining a momentum history of the sampling trajectory, SADrift extracts the \emphkinematic residual to generate a deterministic extrapolative field. This mechanism forces the trajectory to efficiently explore the local tangent space, effectively decoupling exploratory variance into targeted deterministic dispersion and a minimal stochastic buffer. To safely anchor these dispersed trajectories and prevent out-of-distribution divergence, the residual drift is analytically integrated into a micro-time Ornstein-Uhlenbeck (OU) process, providing a mathematically sound, mean-reverting geometric safeguard. Requiring exactly zero additional Neural Function Evaluations (NFE), SADrift empowers a frozen base model to achieve a highly competitive FID of 2.00 at just 20 NFE on ImageNet-256. By successfully bypassing the numerical attrition of the 250-step SDE baseline (FID 2.06), SADrift establishes a new Pareto frontier for variance injection. Beyond image synthesis, we further validate SADrift on fast video generation paradigms (e.g., TurboDiffusion with Wan2.1-1.3B), demonstrating its universal capability to unlock latent generative quality across different modalities and acceleration regimes.
PaperID: 2420, Poster
Abstract: Budget-feasible mechanisms (BFMs) are a fundamental tool for budget-constrained procurement, but designing strong BFMs remains challenging, particularly for general valuation classes. Classical approaches rely on problem-specific analysis and worst-case approximations, often yielding conservative mechanisms in practice. This motivates data-driven methods for discovering mechanisms with stronger empirical performance. However, strict budget constraints expose a key limitation of existing neural automated mechanism design (AMD): they either fail to provide exact economic guarantees or suffer from severe scalability bottlenecks. We propose BFMNet, a scalable menu-based neural framework for learning truthful budget-feasible procurement mechanisms. BFMNet first learns self-bid-independent menus from market context, ensuring dominant-strategy incentive compatibility (DSIC) and individual rationality (IR) by construction, and then applies an ex-post, instance-level menu transformation to enforce budget feasibility. We prove that the resulting mechanism satisfies exact DSIC, IR, and budget feasibility. By avoiding intractable global discretization grids and restrictive Lipschitz constraints, BFMNet remains both expressive and scalable. Extensive experiments show that BFMNet improves utility over classical baselines by up to 34.1%, remains competitive with state-of-the-art neural baselines that provide only approximate economic guarantees, and scales robustly to larger auction instances.
Abstract: Data assimilation (DA) for systems governed by partial differential equations (PDE) aims to reconstruct full spatiotemporal fields from sparse high-fidelity (HF) observations while respecting physical constraints. While full-grid low-fidelity (LF) simulations provide informative priors in multi-fidelity settings, recovering an HF field consistent with both sparse observations and the governing PDE typically requires per-instance test-time optimization, which becomes a major bottleneck in time-critical applications. To alleviate this, amortized reconstruction using generative models has recently been proposed; however, such approaches rely on full-field HF supervision during training, which is often impractical in real-world settings. From a more realistic perspective, we propose the ), which transports an informative LF prior toward an auxiliary-energy-defined, observation-conditioned HF endpoint law without any additional inference-time guidance. To enable learning without HF endpoints, employs an iterative surrogate-endpoint refresh scheme, and directly incorporates PDE residuals into the training objective while enforcing observations via hard conditioning throughout sampling. Experiments on fluid PDE benchmarks demonstrate that enables extremely fast spatiotemporal field reconstruction while achieving strong accuracy under sparse HF supervision.
Abstract: Simulating student learning behaviors in open-ended problem-solving environments holds potential for education research, from training adaptive tutoring systems to stress-testing pedagogical interventions. However, collecting authentic data is challenging due to privacy concerns and the high cost of longitudinal studies. While Large Language Models (LLMs) offer a promising path to student simulation, they suffer from competency bias, optimizing for efficient correctness rather than the erratic, iterative struggle characteristic of novice learners. We present BEAGLE, a neuro-symbolic framework that addresses this bias by incorporating Self-Regulated Learning (SRL) theory into a novel architecture. BEAGLE integrates three key technical innovations: (1) a semi-Markov model that governs the timing and transitions of cognitive behaviors and metacognitive behaviors; (2) Bayesian Knowledge Tracing with explicit flaw injection to enforce realistic knowledge gaps and "unknown unknowns"; and (3) a decoupled agent design that separates high-level strategy use from code generation actions to prevent the model from silently correcting its own intentional errors. In evaluations on Python programming tasks, BEAGLE significantly outperforms state-of-the-art baselines in reproducing authentic trajectories. In a human Turing test, participants could not reliably tell BEAGLE traces apart from real student data: classification accuracy was statistically equivalent to chance (52.8%, d'=0.15, N=71).
PaperID: 2423, Poster
Abstract: Neural additive models (NAMs) offer interpretability but lack expressivity due to their rigid additive structure, whereas multi-layer perceptrons (MLPs) achieve universal approximation at the cost of transparency. We establish the first theoretical foundation for projected neural additive models (PNAMs), which augment NAMs with a learnable linear transformation \boldsymbolT to capture cross-variable interactions. While standard NAMs cannot approximate even the simplest interaction term \chi_1 \chi_2 because their cross-derivatives are identically zero, we prove that PNAMs are universal approximators and characterize the minimum number of projections required for exact polynomial representation. Furthermore, we derive an approximation rate of \mathcalO(M^-s / (N - 1)) for \mathcalC^s functions (i.e., functions with s continuous derivatives), where M denotes the number of projections and N the input dimension; this rate is known to be optimal for ridge function approximation. To recover interpretability from the projected inputs, we introduce structured sparsity regularizers and post-hoc symbolic regression techniques that enable global input ranking, parameter pruning, and the conversion of learned bases into compact analytical expressions. Experiments on knot theory invariants, MNIST, and phase field fracture mechanics demonstrate that PNAMs match or surpass MLPs in accuracy while discovering salient features and yielding symbolic models with orders of magnitude fewer parameters.
Authors: Isaac R Christian, Udith Haputhanthri, Declan Campbell, Samuel Nastase, Taylor Webb, Michael S Graziano
Abstract: Visual perception depends on top-down goals and bottom-up sensory mechanisms. Vision-language models implement both, allowing us to treat each component as a separable hypothesis about what drives where we look. We compared spatial attention maps from six vision-language models against human fixation heatmaps recorded on 200 images during two tasks (general description and social captioning). The six models spanned a 2×2 factorial of CNN vs.\ ViT encoders crossed with LSTM vs.\ Transformer decoders, plus Molmo 7B-D and Qwen3.5 9B. We found that both decoder and encoder architecture shaped alignment, but decoder choice dominated. LSTM vs.\ Transformer decoders increased alignment by 40--50 percentage points (80--87% vs.\ 40--59% of the human noise ceiling). In contrast, CNN vs.\ ViT encoders contributed a secondary 5--20 point advantage depending on decoder family, with CNN-LSTM the most aligned model overall (85--87%). Despite their alignment advantage, LSTM-decoder attention maps were spatially diffuse and minimally task-differentiated; ViT-Transformer, the weakest in alignment, showed the sharpest spatial concentration and strongest task differentiation. A hemispatial-neglect simulation confirmed that ablating attention impacted LSTM decoders more than Transformer decoders. In an exploratory extension using TRIBE-simulated synthetic neural responses, fixation alignment and neural relevance dissociate: CNN-Transformer attention maps better predicted synthetic brain activity despite lower fixation alignment, with attention maps best predicting early visual cortex. Together, top-down and bottom-up components trade off what they predict in behavioral and synthetic neural data.
Abstract: Standard decoding rules for autoregressive language models promote diversity by rescaling the full next-token distribution or by truncating its low-probability tail. These strategies overlook a recurring regime of open-ended generation in which the model already assigns substantial probability to several plausible continuations yet still concentrates too much mass on the most generic choice. We introduce XTC Exclude TopChoices, a lightweight head-aware decoding operator that targets this head-ambiguity regime directly. Given a next-token distribution, XTC identifies the set of tokens exceeding an absolute plausibility threshold \tau. When two or more such tokens exist, it removes the dominant eligible choices with probability \rho and retains only the weakest plausible alternative before renormalizing. A comprehensive evaluation spanning 60 experiments across three primary model families (Gemma-3 27B q4, Gemma-3 12B q6, DeepSeek R1 14B q6), extended with a scaling validation on Llama 3.3 70B q4, confirms the predicted operating profile. On creative generation tasks, XTC improves the diversity-repetition Pareto frontier with Distinct-2 gains of 11-15 % (monotone in parameter count from 12B to 70B) and repeat trigram reductions of 27-47% across the four tested models. When composed with temperature scaling, total improvements reach 38% (Distinct-2) and 71% (repeat trigram reduction) over baseline. A blinded Amazon Mechanical Turk study with 150 Master raters confirms that these distributional shifts translate to a 62.3% creativity preference for XTC (p<10^-4) without sacrificing fluency, and a cross-vendor GPT-4o control judge replicates the Anthropic-judge signal on every directional measure. On instruction-following (IFEval, Llama~3.3 70B q4), XTC preserves prompt-level strict accuracy within 1.7 percentage points of baseline at parameters that recover most of the diversity gain. A temperature setting matched on Distinct-2 collapses IFEval by 8.8 points at the same Distinct-2 target. The effect is additive with temperature and repetition penalties, robust across quantization levels and model families, and consistent across all twelve tested prompt genres.
PaperID: 2426, Poster
Abstract: Model merging has emerged as a practical paradigm for integrating multiple independently trained models into a single model without joint retraining. Prior work shows that parameter fusion techniques, such as parameter decomposition, coefficient optimization, and subspace learning, can achieve strong performance while avoiding the cost of joint training. However, these approaches largely treat merging as a parameter aggregation problem, implicitly assuming compatibility in the underlying directional structures. We argue that this assumption rarely holds. In practice, independently trained models often exhibit misaligned dominant directions in both parameter and feature spaces. Naïve merging can therefore disrupt structural coherence, leading to significant performance degradation, especially in heterogeneous settings such as healthcare, where models trained on tasks like medical QA or clinical text understanding develop distinct representational geometries. Moreover, coefficient-based methods rely on well-aligned feature directions, an assumption contradicted by insights from Neural Collapse, which suggest that class features follow structured yet model-specific directional patterns. In this work, we identify directional alignment as the key principle for effective model merging. We propose Merging with Directional Alignment (\method), a unified geometric framework that explicitly enforces consistency in directional structures across both parameter and feature spaces. Our analysis shows that directional alignment preserves structural coherence during merging, and extensive experiments across multiple benchmarks, including healthcare QA, demonstrate consistent improvements across model scales and task settings.
PaperID: 2427, Poster
Authors: Lueke Ni, Junhao Zhu, Xinyu Zhang
Abstract: Vision-Language-Action (VLA) foundation models have driven remarkable progress in general-purpose robotic manipulation. Yet, adapting these models to dynamic physical environments exposes a severe vulnerability: a modality imbalance during optimization. Because low-dimensional proprioceptive states provide a more direct path for loss minimization, policies instinctively form a "state-dominant shortcut," systematically suppressing the learning of high-dimensional visual features. This visual neglect causes models to bypass precise spatial grounding, leading to catastrophic alignment failures during contact-critical execution phases. To rectify this without disrupting natural training dynamics or requiring heavy auxiliary decoders, we introduce Phase-Adaptive Fusion (PAF), a lightweight, non-intrusive neural modulation plug-in. Operating entirely in the forward pass, PAF features a decoupled dual-branch architecture. A spatial branch performs explicit semantic-geometric alignment, while a temporal branch infers latent task execution phases from recent action history and state kinematics. By synthesizing these signals into a residual gating mechanism prior to multimodal fusion, PAF dynamically amplifies visual tokens precisely during alignment-heavy stages, all while preserving the host VLA's pre-trained representational manifold. Extensive evaluations across simulated benchmarks (LIBERO, Meta-World, Robomimic) and real-world long-horizon tasks demonstrate that PAF successfully breaks the proprioceptive shortcut, drastically reducing terminal alignment drift and significantly enhancing overall manipulation robustness.
Abstract: Traditional neural solvers for solving vehicle routing problems (VRPs) often suffer from limited solution diversity, motivating the development of Generative Flow Network (GFlowNet)–based models. However, the effectiveness of these models is frequently constrained by insufficient flow expansion in high-reward regions, limiting their ability to distribute probability across promising solution routes, the deeper exploration could yield superior results. Diffusion models, in contrast, provide stronger structural guidance for exploration. These two paradigms are naturally complementary: GFlowNet can supply edge-level signals for diffusion to embed, while diffusion can guide broader exploration. Leveraging this synergy, we propose Diffusion-Enhanced GFlowNet (DEG), a novel framework that integrates GFlowNet with diffusion model to encourage richer flow expansion toward high-reward regions and derive higher-quality solutions. Specifically, DEG exploits GFlowNet’s inherent diversity to generate edge-specific backward signals, applies the stochastic noise schedule of diffusion to perturb these signals, and then denoises them within the GFlowNet paradigm. To further improve scalability, we introduce a specialized decoder capable of dynamically adapting to diverse problem scales. Extensive experimental evaluations on synthetic and real-world datasets, including instances with up to 10,000 nodes, demonstrate that DEG consistently achieves favorable performance compared to baseline methods.
Abstract: Equivariant Graph Neural Networks (GNNs) have significantly advanced the modeling of 3D molecular structure by leveraging group representations. However, their message passing, heavily relying on strictly equivariant operations, suffers from restricted expressiveness due to the limited non-linearity and low degree of group representations. To overcome this, we introduce the Equivariant Spherical Transformer (EST), a novel plug-and-play module that applies a Transformer-based architecture to the Fourier spatial domain of group representations. The integration of EST enhances the model's expressiveness while preserving the crucial equivariant inductive bias through a uniform sampling strategy of spherical Fourier transforms. As demonstrated by our experiments on challenging benchmarks like OC20, MPtrj and QM9, EST-based models achieve state-of-the-art performance. For the complex molecular systems within OC20, small models empowered by EST can outperform some larger models and those using additional data. In addition, we provide both theoretical and experimental validation of EST's equivariance as well, paving the way for new research in this area.
Abstract: Classifier-guided diffusion models generate conditional samples by augmenting the reverse-time score with the gradient of the log-probability predicted by a probabilistic classifier. In practice, this classifier is usually obtained by minimizing an empirical loss function. While existing statistical theory guarantees good generalization performance when the sample size is sufficiently large, it remains unclear whether such training yields an effective guidance mechanism. We study this question in the context of cross-entropy loss, which is widely used for classifier training. Under mild smoothness assumptions on the classifiers, we show that controlling the cross-entropy at each diffusion model step is sufficient to control the corresponding guidance error. In particular, we show that probabilistic classifiers achieving conditional KL divergence \varepsilon^2 induce guidance vectors with mean squared error \widetilde O(d \varepsilon ), up to constant and logarithmic factors. We demonstrate that the proposed smoothness condition is necessary, by constructing a sequence of non-smooth classifiers that achieve small conditional KL while inducing an exploding guidance error. We also show that the derived upper bound \widetilde O(d \varepsilon ) is optimal up to poly-logarithmic factors. Our result yields an upper bound on the sampling error of classifier-guided diffusion models and bears resemblance to a reverse log-Sobolev--type inequality. To the best of our knowledge, this is the first result that quantitatively links classifier training to guidance alignment in diffusion models, providing both a theoretical understanding of when such models succeed, along with principled guidelines for selecting classifiers that induce effective guidance. In particular, our findings suggest that, when selecting classifiers for guidance, one should prioritize not only accuracy but also smoothness, highlighting the advantages of smoothness-inducing training methods for diffusion guidance.
PaperID: 2431, Poster
Abstract: We consider the problem of learning to generate data while satisfying a prescribed sample-wise constraint with a user-specified probability of constraint satisfaction. Starting with a population-level constrained negative log-likelihood minimization and using a flow-matching (FM) parameterization, under regularity assumptions and sufficient expressivity of the velocity field family, we derive guarantees on the probability of constraint satisfaction for the corresponding primal-dual optimization algorithm. We provide a practical implementation of this approach, called Primal-Dual Flow Matching (PDFM), which has direct control over constraint satisfaction rate and requires only binary membership-oracle feedback for constraint satisfaction, avoiding the need for differentiable distances, projections, or special structure of the constraint set, such as convexity. We show that compared to existing methods, PDFM achieves higher constraint satisfaction rates while maintaining competitive distribution fidelity.
PaperID: 2432, Poster
Abstract: Policies achieving strong performance in simulators or learned world models can fail when deployed into the real environment. While there must be environmental differences that account for these failures, not all differences are equally responsible. A benign error of irrelevant visual details may co-exist with a failure to predict a single critical transition causing an inevitable catastrophic failure. In this paper, we introduce \emphcounterfactual debugging as an approach for \emphworld model transfer gap attribution to identify the time steps whose transition or reward errors are \emphcausally responsible for a performance degradation exhibited in a real trajectory, rather than only visually or statistically different. We develop a scalable algorithm for computing these attributions that exploits the sparsity of causal errors by recursively divide-and-conquer, significantly reducing computational costs and achieving an exponential speedup. The resulting ranked attribution report can better explain the performance gap in the real trajectory. Experiments across several distinct environments with a variety of injected world-model failures (observation corruption, reward misprediction, and physics violations) demonstrate that counterfactual debugging correctly identifies the errors that are responsible for performance gap, providing theoretically grounded, actionable insights for model improvement.
Authors:
Hyeonjin Kim, Hangyeol Jung, Heechan Yun, Sungjun Yun, Dong-Jun HanAbstract: Unlearning specific concepts in text-to-image diffusion models has become increasingly important for preventing undesirable content generation. Among prior approaches, sparse autoencoder (SAE)-based methods have attracted attention due to their ability to suppress target concepts through lightweight manipulation of latent features, without modifying model parameters. However, SAEs trained with sparse reconstruction objectives do not explicitly enforce concept-wise separation, resulting in shared latent features across concepts. To address this, we propose SAEParate, which organizes latent representations into concept-specific clusters via a concept-aware contrastive objective, enabling more precise concept suppression while reducing unintended interference during unlearning. In addition, we enhance the encoder with a GeLU-based nonlinear transformation to increase its expressive capacity under this separation objective, enabling a more discriminative and disentangled latent space. Experiments on UnlearnCanvas demonstrate state-of-the-art performance, with particularly strong gains in joint style-object unlearning, a challenging setting where existing methods suffer from severe interference between target and non-target concepts.
PaperID: 2434, Poster
Authors: Peter M Jacobs, Jeff M Phillips
Abstract: Squared Wasserstein distance is a frequently used tool to compare probability distributions. This distance is typically computed between empirical measures of size n from two underlying random samples. Unfortunately, even in lower dimensional Euclidean space problems \left( d \in \2,3\ \right), Wasserstein distance algorithms with approximate or exact precision guarantees scale poorly in the runtime as a function of n and the desired precision. In response, we consider the computational-statistical runtime, where the goal to estimate from samples the Wasserstein distance up to the \varepsilon-additive error which is achievable under a sample; we allow O(1) time per sample. Towards this, we develop a Sample-Sketch-Solve paradigm where we introduce a regular cartesian grid sketch of the samples. We show that (especially under \alpha-H\"older smooth distributions) this can compress the data without increasing asymptotic error, and also regularizes the structure which enables faster exact algorithms. Ultimately, we approximate W_2^2(P,Q) within \varepsilon error in \varepsilon^-\max(2,\fracd+1+o(1)2) time for Lipschitz P,Q on (0,1)^d; nearly an optimal \Theta(\varepsilon^-2) for d \leq 3.
PaperID: 2435, Poster
Abstract: This paper studies concentration inequalities for sampling without replacement. We introduce a general inequality that provides simple guidelines for proving concentration inequalities, offers a degree of unification of existing results, and is strong enough to yield sharper bounds. As applications, we revisit several standard statistics and either recover classical inequalities through streamlined proofs or strengthen them to tighter bounds, with rates optimal up to constants.
PaperID: 2436, Poster
Abstract: Language models are increasingly deployed alongside test-time search. We argue that we should therefore shift the role of RL post-training from converging on a single best response to producing a diverse pool of competent candidates. Standard policy gradient methods such as GRPO optimize a fixed scalar reward and drive the policy toward near-duplicate responses, erasing the diversity search needs. We propose Vector Policy Optimization (VPO), which exploits the fact that practical rewards are vector-valued, e.g., per-test-case correctness, per-criterion ratings, or per-hop credit. Instead of collapsing these into one scalarization, VPO samples randomized scalarizations from a distribution over the reward simplex, incentivizing candidates to specialize to different trade-offs along the Pareto front. The model emits multiple candidates per prompt, with each conditioning on the previous ones, allowing for directed diversification. VPO is a drop-in replacement for the GRPO advantage estimator. Across four domains, VPO consistently improves test-time best@k over scalar baselines, with gains widening as the test-time budget grows.
PaperID: 2437, Poster
Abstract: We study representation learning in low-rank Constrained Markov Decision Processes (CMDPs), where both value and reward functions are approximated by a set of unknown representation vectors, and the transition dynamics admit a low-rank factorization. The objective is to maximize expected cumulative rewards while complying with constraints on the expected cumulative utility. To achieve this, we propose \textttMFRLC (Model-Free Representation Learning for Low-Rank CMDP), a model-free algorithm that uses a primal-dual scheme to effectively balance reward regret and constraint violations. To the best of our knowledge, \textttMFRLC is the first representation-learning method for low-rank CMDPs that combines the Least-Squares Value Iteration with Upper Confidence Bound (LSVI-UCB) and primal-dual techniques. \textttMFRLC further enhances value function estimation with bonuses for function approximation and uniquely learns the mapping from representations to value functions directly through representation learning. We prove that \textttMFRLC achieves both regret and cumulative constraint violation of order \widetilde O(H^3 d^2 K^3/4|\mathcalA|^3/2/\gamma), making it provably sample-efficient and highly adaptable to complex environments due to the reliance on function approximation.
PaperID: 2438, Poster
Authors: Chenghao Liu, Hui Lu, Chengbo Zhan, Jia Rao
Abstract: Self-speculative decoding offers a plug-and-play route to accelerating autoregressive large language model (LLM) inference, but its practical success hinges on identifying effective layer-skipping configurations online without retraining or offline profiling. Existing methods mainly optimize agreement-based proxies that capture draft-target consistency, yet overlook intrinsic runtime differences across draft configurations. Consequently, configurations with similar acceptance can still yield substantially different realized decode throughput. We present SPIRIT, a speed-driven online adaptation framework for self-speculative decoding that directly optimizes realized decode throughput, measured as generated tokens per second, using window-level runtime feedback. Our key insight is that speed is not determined by agreement alone: it is jointly shaped by draft-target consistency and the intrinsic speed advantage of the draft configuration. Empirically, the skip ratio serves as a first-order control variable, capturing the dominant consistency-speed trade-off while sharply shrinking the search space over layer-skipping patterns. SPIRIT therefore first identifies an effective skip ratio and then refines the layer-skipping pattern within it. Combined with length-aware normalization and shift-aware profile-based adaptation, this design enables rapid adaptation to task shifts in non-stationary request streams. Across diverse tasks and model scales, SPIRIT consistently outperforms strong training-free baselines and achieves up to 1.74x speedup while preserving the target model's output distribution, establishing a practical foundation for plug-and-play, high-throughput LLM inference.
Abstract: Fast linear algebra in deep learning usually comes with a choice: fixed geometry and exact computation, as in the Fourier transform, or adaptive geometry paid for by dense parameters, random features, or low-rank surrogates. To move beyond this trade-off, we introduce LAPLEX, a class of exact, trainable (phased) Laplace-kernel operators. A LAPLEX layer is a typically full-rank dense matrix, implicitly defined by learnable coordinate anchors, with FFT-like scaling. Consequently, it supports trainable matrix--vector operations at vector dimensions up to 10^9 on modern GPUs. As a neural layer, it yields compact projections and classification heads interpretable as soft, trainable routing models. The same primitive also serves as an efficient Gram operator, enabling high-dimensional covariance models on flattened images of dimension 3 \cdot 10^6 that preserve visible spatial structure without imposing convolutional bias. These applications reflect a single principle: dense geometry can be learned without storing a dense matrix, which enables data-adaptive global interactions in regimes where ordinary dense layers are out of reach. In this sense, LAPLEX separates expressivity from storage cost: it behaves like a dense trainable matrix, but is represented and applied through a small structured set of parameters.
PaperID: 2440, Poster
Abstract: The Linear Representation Hypothesis (LRH) posits that Transformer activations decompose into sparse linear combinations of interpretable directions, with Sparse Autoencoders (SAEs) serving as the predominant method for recovering such decompositions. We propose that idempotency -- the requirement that re-encoding an SAE's output reproduces the same latent representation -- is a desirable property for any faithful dictionary learning method, and demonstrate empirically that most SAEs deviate very far from this ideal. We identify two underlying failure modes: exploding representation norms and unstable active feature sets. Further analysis reveals that even individual SAE latents cannot be reliably reconstructed under re-encoding, and that SAE outputs exhibit significant sensitivity to input scaling. We show that idempotency can be improved via an auxiliary loss term and demonstrate its potential to improve SAE interpretability.
PaperID: 2441, Poster
Abstract: Image-free Zero-Shot Learning (I-ZSL) aims to extend a pre-trained classifier to unseen classes without using training images or image features during adaptation. Existing I-ZSL methods often depend on pre-defined class descriptions and fixed text encoders, which can be misaligned with the classifier space of the deployed model. We propose Geometry-Regularized Relational Weight Synthesis (GeoRWS), a read-only classifier expansion framework that learns to synthesize unseen-class classifier weights from pairwise semantic relations and observed seen-class weights. Given class-pair affinity descriptions generated by a frozen large language model, GeoRWSuses an adaptive semantic encoder to estimate affinity coefficients over seen classifiers through leave-one-out reconstruction. The training objective further includes a classifier-geometry regularizer, which anchors the learned coefficients to neighborhoods in the seen-class classifier head, and a semantic-consistency regularizer, which improves stability under perturbations of pairwise semantic descriptions. At inference, GeoRWS synthesizes unseen-class weights as mixtures of seen-class weights and injects them into the classifier head while keeping the feature extractor fixed. Experiments on standard ZSL and GZSL benchmarks show consistent improvements over image-free and adapted zero-shot baselines, especially on fine-grained recognition tasks.
PaperID: 2442, Poster
Authors:
Qingfeng Li, Wenhui Tu, Wei Liu, Huaijie Zhu, Jielong Tang, Jian YinAbstract: Multimodal recommendation leverages item content, such as text, images, and audio signals, to complement sparse user-item interactions. Existing methods mainly combine content and collaborative signals at the representation level through graph propagation, feature fusion, or auxiliary alignment. However, continuous item embeddings make semantic sharing implicit and provide limited support for transferring behavioral evidence through reusable semantic factors. Discrete semantic codes offer explicit shared variables, but existing code-based recommenders usually treat item-code assignments as fixed outputs of pretrained quantizers, clustering algorithms, or frozen encoders, making the assignment mechanism insensitive to user behavior. We propose CASQRec, a Collaborative-Adaptive Semantic Quantization framework for multimodal recommendation. CASQRec learns item-code assignments as content-inferable routing decisions supervised by collaborative behavior during training. It constructs a semantic-code graph, where fixed content-derived RQ codes are graph nodes and item-code edges are predicted from item content by a learnable prior router. To make routing behavior-aware while keeping item-code routing independent of item-side interaction history at inference time, CASQRec introduces a collaborative posterior teacher only during training and distills its assignment distributions into the content prior. The resulting prior-induced user-item-code graph allows collaborative evidence to be shared through discrete semantic codes while keeping item routing content-based. Experiments on three public datasets show that CASQRec consistently outperforms representative multimodal recommendation baselines. The code resources are available in the supplementary materials.
PaperID: 2443, Poster
Authors:
Johan Mylius-Kroken, Elisabeth Wetzer, Ali Ramezani-Kebrya, Robert Jenssen, Kristoffer WickstrømAbstract: Mutual information (MI) is fundamental to representation-learning analysis in deep networks, but in high dimensions MI estimators depend on hyperparameters --- bin widths, kernel bandwidths, neighbourhood radii, auxiliary networks --- whose bias often dominates the signal. We show that continuous piecewise-linear (CPWL) networks, including ReLU networks, admit an intrinsic information-theoretic analysis that avoids this pitfall: their inherent partition of the input space into convex linear regions assigns to every input a discrete region label \Pi alongside the continuous representation T. We decompose I(Y; T) exactly into geometric terms, where the \emphrouting information I(Y; \Pi) admits a hyperparameter-free plug-in estimator that approximates a population lower bound on every other MI term in the decomposition. When the partition saturates, a functional-equivalence quotient at relative Frobenius tolerance \varepsilon \in [0, 2] collapses regions implementing the same linear operator and restores informativeness. Empirically, the routing estimator correlates strongly with three established MI baselines \it without any hyperparameter, and the functional quotient at moderate \varepsilon recovers informative estimates in the regime where the raw partition is bounded by \log_2 N rather than H(Y). We provide a parameter-free framework for analysing the internal information geometry of CPWL networks, grounded in their inherent piecewise-linear structure rather than in external discretisation.
Abstract: On-policy distillation (OPD) has a property absent in offline distillation and RL: teacher supervision quality depends on student competence. Incoherent rollouts yield noisy gradients; already-mastered tokens yield redundant ones. This creates waste at three scales—tokens, training phases, and prompts—yet existing methods supervise uniformly. We introduce SEAD, which uses entropy as a unified probe of this competence-dependent degradation at three scales: (1) joint teacher-student entropy partitions tokens into zones receiving tailored divergences or zero gradient (~50% skipped); (2) a cosine schedule anneals from forward to reverse KL as competence grows; (3) a competence-gated curriculum introduces prompts easy-to-hard. These components are symbiotically necessary: token selection requires coherent rollouts (curriculum), and annealing requires monotonic improvement (also curriculum). On OLMo-3 (7B to 32B), SEAD achieves +4.8 average accuracy over vanilla OPD across six math benchmarks, with ablations confirming super-additive interactions.
PaperID: 2445, Poster
Abstract: Transformers are powerful models but are inefficient with large context due to the quadratic complexity of the attention mechanism. Limiting their context window addresses the efficiency concern but at a cost: they cannot remember everything from the past, which severely limits their computational power. We show that we can simultaneously achieve efficiency and universal computation if we restrict attention heads to be 2-dimensional. First, we give the first algorithm for standard 2D attention that serves every insertion and every query in amortized polylogarithmic time per token. Prior work either restricted key/query magnitudes or paid polynomial overhead per step. Second, we show that the class of 2D-attention models is universal and efficient: a vanilla decoder transformer efficiently simulates an arbitrary RAM machine via chain-of-thought decoding so that, combined with our first result, the simulation runs in \mathrmpolylog(N) time per simulated RAM step end-to-end. In contrast, earlier universality results addressed only computability, showing Turing completeness while incurring large polynomial factors in their execution time. Third, as an application, we turn a transformer into a computer. We provide a compiler that supports the full \texttti32 WebAssembly subset with integer computations (reachable from C via \textttclang) and turns such programs into the weights of a transformer, which then executes them over millions of steps autoregressively. We illustrate how this can be used to solve combinatorial problems, tasks at which even large reasoning LLMs struggle.
Authors: Haiqian Yang, Yuan Cao, Markus Buehler
Abstract: World models have enabled interactive exploration of game environments and robotic manipulation, but physical engineering remains beyond their reach: real materials exhibit nonlinear constitutive laws, carry history-dependent internal state, undergo inertial dynamics, and may possess hierarchical structures spanning multiple length scales. We present LEIA (Learned Environment for Interactive Architected materials), a world model that lets engineers apply boundary conditions step by step and observe the resulting deformation and stress fields in real time. LEIA handles large three-dimensional unstructured meshes and generates autoregressive responses to user-specified loading. We introduce MicroPlate, a benchmark of architected plates spanning two regimes of microstructure modeling: architected lattices that resolve microstructure explicitly through three-dimensional geometry, and a homogeneous plate where microstructural change is modeled implicitly through internal degrees of freedom. MicroPlate is used to assess LEIA alongside four baseline methods across both regimes. Finally, we demonstrate that LEIA enables efficient candidate generation and ranking for fast surrogate-guided search for de novo designs of architected materials, with stress-accurate candidate ranking validated by finite element ground truth.
Authors:
Rajeev Yasarla, Deepti Hegde, Shizhong Han, Hsin-Pai Cheng, Yunxiao Shi, MeysamSadeghi, Shweta Mahajan, Apratim Bhattacharyya, Litian Liu, Risheek Garrepalli, Thomas Svantesson, Mohammad Ghavamzadeh, Fatih Porikli, Herbert CaiAbstract: Vision-Language-Action (VLA) models are emerging as highly effective planning models for end-to-end autonomous driving systems. However, current works mostly rely on imitation learning from sparse trajectory annotations and under-utilize their potential as generative models. We propose Generative Scenario Rollouts (GeRo), a plug-and-play framework for VLA models that jointly performs planning and generation of language-grounded future traffic scenes through an autoregressive rollout strategy. First, a VLA model is trained to encode ego vehicle and agent dynamics into latent tokens under supervision from planning, motion, and language tasks, facilitating text-aligned generation. Next, GeRo performs language-conditioned autoregressive generation. Given multi-view images, a scenario description, and ego-action questions, it generates future latent tokens and textual responses to guide long-horizon rollouts. A rollout-consistency loss stabilizes predictions using ground truth or pseudo-labels, mitigating drift and preserving text-action alignment. This design enables GeRo to perform temporally consistent, language-grounded rollouts that support long-horizon reasoning and multi-agent planning. On Bench2Drive, GeRo improves driving score and success rate by +15.7 and +26.2, respectively. By integrating reinforcement learning with generative rollouts, GeRo achieves state-of-the-art closed-loop and open-loop performance, demonstrating strong zero-shot robustness. These results highlight the promise of generative, language-conditioned reasoning as a foundation for safer and more interpretable end-to-end autonomous driving.
PaperID: 2448, Poster
Abstract: Understanding 4D scenes is a core challenge in visual intelligence, requiring models to jointly capture scene geometry, appearance, and dynamics from partial observations. Existing approaches, however, are subdivided into feedforward models that are used for direct geometric prediction and generative models for visual synthesis. We propose a unified generative-predictive framework for 4D scene understanding that jointly models RGB video and geometric representations such as depth and camera pose with a single joint generative model. Our generative formulation enables us to learn from diverse, heterogeneous datasets with different subsets of geometric and visual annotations, and enables us to condition on arbitrary subsets of inputs, allowing a single model to perform cross-modal generation, geometric prediction, and future scene inference from partial observations. In addition, our generative approach enables test-time search over multiple candidate scene completions, improving geometric prediction from sparse observations. Overall, we illustrate the efficacy of our generative-predictive framework on a suite of generation and prediction tasks.
PaperID: 2449, Poster
Authors: Biao Xiang, Ali Eshragh, Yuexing Li, Kai Wang
Abstract: Off-policy evaluation (OPE) estimates the value of a target policy from logged data, but limited behavior-policy coverage can force high-variance reweighting or reward-model extrapolation. Counterfactual annotations can add evidence about unobserved actions, yet practical sources, including domain experts and LLMs, may be costly, biased, or noisy. We study budgeted acquisition of such annotations for contextual-bandit OPE. Given source-specific costs and error profiles, we formulate an integer allocation problem over context-action pairs and annotation sources to minimize the component of estimator variance that depends on the annotation plan. We characterize when annotations are valuable through a first-annotation threshold and local annotation-value regimes. For the coupled multi-source problem, we develop a majorization-minimization algorithm with dynamic-programming subroutines that monotonically improves the objective. Experiments in synthetic clinical and LLM-annotated education bandits show that our proposed allocation method reduces the optimized variance component by 20.58% and 10.77%, respectively, relative to no-annotation baseline.
PaperID: 2450, Poster
Authors:
Albert Gimó Contreras, Mariia Vladimirova, Olga Petrova, Reda CHHAIBI, Patrick LoiseauAbstract: A standard way to enforce group fairness in machine learning models is to add a fairness regularizer to the training loss. Existing dependence-based regularizers, however, are often computationally expensive, with per-batch costs that are typically quadratic or higher in the batch size B. We propose a novel group-fairness regularizer based on the Cramér-von-Mises (CvM) sensitivity index, which penalizes statistical dependence between model predictions and a sensitive attribute during training. Our method combines a rank-based CvM estimator with differentiable soft ranking, yielding a bounded training penalty with \Oc(B \log B) per-batch complexity. This is the first sub-quadratic in-processing fairness method that targets genuine joint-distribution dependence. We further establish theoretical connections between the CvM regularizer and standard fairness metrics such as demographic parity and equality of opportunity. Experiments on tabular and image datasets show competitive fairness-utility trade-offs while substantially lowering training overhead compared to existing dependence-based regularizers.
Abstract: Foundation models for time-dependent partial differential equations (PDEs) are trained on large and diverse collections of physical systems and can generalize effectively to new downstream tasks. After fine-tuning on only a few trajectories from a target domain, they often achieve strong accuracy in low-data regimes. However, these models are typically large and computationally intensive, limiting their usefulness as fast surrogates for numerical solvers. We propose Teacher Rollout Extension (TREX), a knowledge distillation framework that transfers the predictive capability of a pretrained foundation model into a compact and efficient student. Starting from a fine-tuned teacher, TREX augments limited downstream data by generating long synthetic trajectories through teacher rollouts with periodic noise injection. This procedure samples from the teacher-induced rollout distribution without requiring explicit knowledge of the initial-condition distribution, while exposing the student to long-horizon states and local recovery behavior around states encountered during autoregressive prediction. The student can further incorporate task-specific inductive biases, such as equivariance, that the teacher does not necessarily enforce. We evaluate TREX on multiple PDE benchmarks. The resulting students match or even surpass the teacher’s accuracy while reducing the number of parameters by several orders of magnitude and achieving more than an order-of-magnitude speedup in inference.
PaperID: 2452, Poster
Abstract: We study the problem of controlling instability in the Vlasov--Poisson plasma systems, a fundamental challenge for nuclear fusion control. Motivated by the partial observability of plasma states in practice, we investigate imitation learning algorithms that learn from a fully-observable expert controller, but operate under macroscopic measurements. We theoretically establish a separation between online and offline imitation learning --- even a small behavior cloning error leads to error compounding that exponentially amplifies instability over time. By way of contrast, the online imitation learning loss can polynomially control the stability of rollout trajectories. We then propose a DAgger-style algorithm that learns the stabilizing controller under partial observability. Simulation results on a 1D Vlasov--Poisson system demonstrate that our algorithm can effectively stabilize plasma instabilities over longer time horizons than behavior cloning.
PaperID: 2453, Poster
Authors: Anish Acharya, Phillip Studans, Amit Dhanda, Ninad V Rao, Vishal A Khatri, Sina A Niaki, Brian Verkhovsky
Abstract: Prompt engineering has become a central lever for deploying large language models (LLMs), yet the prompts that actually ship to production bear little resemblance to the short instructions studied in most of the prompt-optimization literature. A deployed prompt is a structured multi-section artifact---role framing, reasoning directives, examples, constraints, tool schemas, and output specifications woven into a single token stream---in which some sections encode hard-won safety and format contracts while others are the legitimate targets of optimization. Existing iterative prompt optimizers treat the prompt as a monolithic string and the search as an unstructured rewrite loop, with no mechanism to respect this partition and no theoretical account of when their updates improve the prompt, preserve diversity, or control length. We address this gap by casting iterative prompt optimization as combinatorial search over an edit graph of block-structured, selectively-mutable prompt states. Under this view, a broad slice of the recent literature collapses to selector--proposal--state-space instances of a single \emphbatched iterative prompt search template that differ only in the selector. Within this framework we introduce \textscSplice, which contributes four coupled design choices, each paired with a formal guarantee: (i) per-block mutability, with a price-of-freezing bound quantifying the safety--optimality trade-off; (ii) dual section-local textual gradients with momentum, yielding an edit-graph reachability bound and a (1-e^-\gamma) sub-modular approximation; (iii) elitist beam search with lineage-aware diversity, giving deterministic monotonicity, a noisy no regression tail bound, and conditional no-collapse of lineages; and (iv) a ratio regularizer for length control, with an explicit accuracy--length bias bound. Together these constitute, to our knowledge, the first theoretical analysis of iterative LLM-driven prompt optimization. Empirically, \textscSplice consistently outperforms prior prompt-optimization methods and single-step sampling baselines across public benchmarks, system models spanning three capability tiers, and tasks ranging from classification to LLM-as-judge rubric optimization, while preserving frozen sections verbatim and exhibiting narrower cross-model variance than single-step sampling.
Abstract: Recent studies show that LLM hidden states encode reward-related information, such as answer correctness and model confidence. However, existing approaches typically fit black-box probes on the full hidden states, offering little insight into how this information is structured across neurons. In this paper, we show that reward-related information is concentrated in a sparse subset of neurons. Using simple probing, we identify two types of neurons: value neurons, whose activations predict state value, and dopamine neurons, whose activations encode step-level temporal difference (TD) errors. Together, these neurons form a sparse reward subsystem within LLM hidden states. These names are drawn by analogy with neuroscience, where value neurons and dopamine neurons in the biological reward subsystem also encode value and reward prediction errors, respectively. We demonstrate that value neurons are robust and transferable across diverse datasets and models, and provide causal evidence that they encode reward-related information. Finally, we show applications of the reward subsystem: value neurons serve as effective predictors of model confidence, and dopamine neurons can function as a process reward model (PRM) to guide inference-time search.
Authors: Alessandro De Palma
Abstract: Adversarial training attains strong empirical robustness to specific adversarial attacks by training on concrete adversarial perturbations, but it produces neural networks that are not amenable to strong robustness certificates through neural network verification. On the other hand, earlier certified training schemes directly train on bounds from network relaxations to obtain models that are certifiably robust, but display sub-par standard performance. Recent work has shown that state-of-the-art trade-offs between certified robustness and standard performance can be obtained through a family of losses combining adversarial outputs and neural network bounds. Nevertheless, differently from empirical robustness, verifiability still comes at a significant cost in standard performance. In this work, we propose to leverage empirically-robust teachers to improve the performance of certifiably-robust models through knowledge distillation. Using versatile feature-space distillation objectives, we show that distillation from adversarially-trained teachers can significantly improve the performance of state-of-the-art certified training algorithms for ReLU networks on robust computer vision benchmarks.
Abstract: Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs. However, a rich and largely untapped source of supervision lies in expert systems (e.g., game engines, classical planners, theorem provers), which routinely produce near-optimal actions across diverse domains. But these experts are silent: they commit to an action without writing down the chain of thought (CoT) behind it. Recovering that CoT as natural-language reasoning would distill expert knowledge into a student that generalizes beyond the demonstrated actions. We treat it as a latent variable and study how to recover it from the action alone. Our approach, LeAct (Learning to reason from Actions), optimizes this latent variable: the student samples candidate CoTs for each expert action, and we retain those that measurably improve its own probability of recovering the action. Across imperfect-information games at multiple scales and a simulated robotics benchmark, LeAct reaches the solver's numerical floor on small enumerable games. At larger scale, it is 5× closer to the solver than the strongest expert-iteration baseline. At Flop Hold'em (~ 10^9 infosets), LeAct wins head-to-head by +60 mbb/g, and on the robotics probe it is the only training recipe that improves on direct imitation. We present a principled framework and the result: expert systems become a categorically new source of reasoning teachers for foundation models.
Abstract: Several disruptive research directions have recently emerged in computer vision, including foundation models achieving previously unseen zero-shot performance in scene understanding, even interactively, and generative models that synthesize extremely realistic images. The latter have also been shown to be highly effective in scene understanding tasks thanks to their rich priors. However, for promptable segmentation, foundation models struggle with accurately segmenting an object's region, leading to false positives and over-segmentation. Notably, early attempts that leverage generative priors use prompts only during post-processing, yielding suboptimal segments because the process is agnostic to the user input. In this paper, we target these limitations with Prompt2Seg, a spatial conditioning framework for diffusion-based segmentation. Prompt2Seg augments a frozen diffusion segmentation model with a conditioning branch. Our approach takes spatial prompts, represented as 2D Gaussians or confidence maps, as explicit input signals, training the model to respond directly to user intent. Fine-tuned on a deliberately constrained set of object categories drawn from Hypersim and Virtual KITTI 2, Prompt2Seg generalizes zero-shot to a wide range of unseen object types and visual domains. We evaluate on seven datasets ranging from standard benchmarks to more challenging domains, including paintings, egocentric views, and X-ray data. Furthermore, we demonstrate that Prompt2Seg consistently outperforms the underlying diffusion segmentation backbone across all benchmarks. Our results suggest that the rich priors encoded in generative pretraining, combined with principled spatial conditioning, offer a compelling path toward broadly generalizing interactive segmentation without large-scale mask supervision.
Abstract: Online learning in non-stationary environments requires models to adapt rapidly from streaming data under strict computational constraints. A principled approach casts online learning as Bayesian state tracking, where model parameters are updated sequentially via Bayesian filtering. However, applying Bayesian filters directly to modern deep models is computationally prohibitive due to the high dimensionality of parameter space, forcing existing methods to rely on restrictive approximations or manually designed low-dimensional subspaces. In this work, we identify the absence of a suitable low-dimensional dynamical representation as the core bottleneck in Bayesian filtering-based online learning. Accordingly, we propose Adaptive Update through Representation Adaptation (AURA), a meta-learning framework that learns offline a low-dimensional latent state-space model governing the evolution of optimal model parameters under distribution shift. Online adaptation is then performed via extended Kalman filtering in this learned latent space followed by reconstruction of the full model parameters through a learned lifting map, enabling efficient single-step online adaptation while preserving model expressiveness. Evaluated on online adaptation of neural wireless receivers under time-varying channels and on non-stationary image classification, AURA shows substantial improvements in adaptation speed, accuracy, and computational efficiency over existing online learning and Bayesian filtering baselines, demonstrating that an adaptation-aware latent geometry is beneficial for effective Bayesian online learning in high-dimensional models.
PaperID: 2459, Poster
Abstract: We present a method of quantum composite hypothesis testing with small error, which enables us to establish quantum lower bounds in nonparametric statistics and high-dimensional functional estimation. This is achieved by introducing trigonometric polynomials into the quantum polynomial method, thereby generalizing the quantum phase-estimation lower bound of Mande and de Wolf (ESA 2023). As applications, we settle the quantum complexities of several problems in property testing and functional estimation with small error: - For \ell_2-closeness testing, we show that the approach of Luo et al. (IEEE Trans. Inf. Theory 2024) is optimal. - For Tsallis entropy estimation, where q=2 corresponds to the Gini impurity, we show that the approaches of Buhrman et al. (Phys. Rev. Lett. 2001) and Ekert et al. (Phys. Rev. Lett. 2002) are optimal for integer q \geq 2, and that the approach of Chen et al. (ICALP 2026) is near-optimal for real q \geq 1.5. - For pure-state Uhlmann fidelity and trace distance estimation, we show that the approach of Wang (IEEE Trans. Inf. Theory 2024) is optimal.
Abstract: Many real world processes can be represented as compositions of functions along a Directed Acyclic Graph (DAG). In causal modelling these correspond to the underlying mechanisms, in engineering to the multiple fidelities present, and in gene-regulatory networks to the transcription factors. These functions are partially observed across the DAG, with noisy and heterogeneously sampled measurements, posing significant challenges for reconstruction, uncertainty propagation, and inference. To tackle these we place priors over functions and naturally arrive at Deep Gaussian Processes over DAGs. We theoretically study their prior-collapse behaviour, and the effect of graph topology and intermediate observations on the preservation of information. We obtain almost-sure lower bounds on the asymptotic frequency of depths at which distinction between inputs is preserved, identify broad kernel classes for which these hold, and prove an observation by \citepdunlop2018 on the role of input connections. We offer a structured variational approximation that retains graph dependencies, satisfies compositional uncertainty, and captures the explaining-away behaviour of colliders. Finally, we empirically validate our theoretical results and our methodology, and model a latent-collider DAG, a protein signalling network, and a multi-fidelity heavy-ion collision task, attaining state-of-the-art performance while recovering the low-fidelity contributions and yielding explainability over the simulator hierarchy.
PaperID: 2461, Poster
Authors:
Yu-Ang Cheng, Sixuan Chen, Zhouyang Lu, Xizheng Yu, Grégoire Dhimoïla, Thomas SerreAbstract: Modern vision--language models are trained with very different objectives: CLIP-style contrastive alignment, SigLIP-style match/no-match prediction, caption-style token generation (Flamingo, LLaVA), and latent-space prediction (VL-JEPA). Each has been claimed as best-in-class for its target task. Yet practitioners lack clear guidelines for choosing between them, and existing identifiability theory analyzes only the contrastive case. In this work, we unveil the core statistical mechanisms that distinguish each objective family using a unified spectral analysis. By leveraging closed-form solutions for the linearized losses under a latent structure in which modality-specific nuisances may spuriously carry shared semantic content, we characterize what each objective recovers from the data. We demonstrate that contrastive alignment, match/no-match prediction, and latent-space prediction collapse to the same cross-modal solution, while directed token-space generation recovers a different one. Our findings indicate that in scenarios where one modality is heavily contaminated by spurious nuisance, token-space generation is preferable because it imposes a strictly weaker, one-sided alignment condition. In scenarios where two modalities are both clean, retrieval is maximized by the latent-prediction objective whose target matches the query modality. These results clarify the trade-offs between the four objective families. We validate our theoretical predictions on numerical, controlled synthetic, and real (Flickr30k, COCO, CC3M) data.
Abstract: Across scientific disciplines, Laplacian eigenvectors serve as a fundamental basis for simplifying complex systems, from signal processing to quantum mechanics. In reinforcement learning (RL), they similarly form a basis over the state space, enabling reward functions to be approximated by projection onto a small set of eigenvectors. This projection makes zero-shot control possible, but it also imposes a fundamental limitation: the induced policies are only as expressive as the linear span of the chosen eigenvectors. We introduce the Laplacian Keyboard (LK), a hierarchical framework that goes beyond this linear span. LK constructs a task-agnostic library of behaviors from these eigenvectors, forming a behavior basis guaranteed to contain the optimal policy for any reward within the linear span. A meta-policy learns to stitch these behaviors dynamically, enabling efficient learning of policies outside the original linear constraints. We establish theoretical bounds on zero-shot approximation error and demonstrate empirically that LK improves over the zero-shot solution while achieving better sample efficiency compared to standard RL methods.
PaperID: 2463, Poster
Authors:
Vasha DuTell, Zeyu Yun, Anne Harrington, Yuan Sheng Fang, Mark Hamilton, Christian Koevesdi, Edward Adelson, Bill Freeman, Ruth RosenholtzAbstract: Deep networks achieve high realism in texture synthesis but often rely on opaque, over-parameterized representations that have limited direct connection to biological vision. Neuroscience-informed pyramid models are interpretable, compact, and biologically grounded, but rely on small, hand-curated statistics sets that are difficult to generalize or extend. We introduce GramStatTexNet, a hybrid analysis-by-synthesis framework that combines the multi-scale Gabor filter structure of classical models with the flexibility of Gramian-based correlations, organized into structured filter families. This structured representation enables analyses that are challenging for deep-feature methods: each statistic carries a named filter-family identity, allowing us to decompose synthesis fidelity into per-family contributions and to learn an interpretable, low-dimensional embedding that preserves perceptual texture content. We compare our framework to classical and deep texture-synthesis baselines across diverse texture categories, and on multiple perceptual metrics, highlighting the compactness and interpretability advantages of this hybrid representation, alongside competitive synthesis quality. This same framework extends naturally to peripheral textures via spatial pooling, and to dynamic textures using a spatiotemporal filter bank. GramStatTexNet provides a unified, interpretable framework for analyzing and modeling visual information with natural extensions across space, time, and eccentricity.
Abstract: Diffusion Transformers (DiTs) achieve state-of-the-art image generation quality but incur substantial memory and computational costs at inference. While aggressive Post-Training Quantization (PTQ) to 4-bit precision offers significant efficiency gains, it typically results in severe quality degradation. Existing approaches, including smoothing-based methods, mixed-precision schemes, rotation techniques, and low-rank residual methods, partially mitigate this issue but still leave a noticeable gap to FP16/BF16 performance. In this work, we introduce DiRotQ, a W4A4 PTQ framework that mitigates this degradation through rotation-aware activation quantization. DiRotQ identifies a low-rank subspace capturing dominant activation variance via Principal Component Analysis (PCA), preserving coefficients in this subspace at higher precision while quantizing the remaining components to 4-bit. Activations are rotated into the PCA basis at inference time using calibration-derived orthogonal transformations, while the inverse rotation is fused into the layer weights offline. Combined with GPTQ-based weight quantization, DiRotQ achieves an FID (\downarrow) of 15.9 and PSNR (\uparrow) of 19.1 dB on PixArt-\Sigma over the MJHQ-30K dataset, outperforming the prior state-of-the-art SVDQuant (FID 18.9, PSNR 17.6) under the same INT W4A4 setting. Beyond standard metrics, we introduce a VLM-as-a-Judge evaluation protocol for diffusion model quantization, the first such evaluation in this setting, providing a more holistic assessment of perceptual quality and prompt alignment under aggressive compression. On the systems side, we implement a Triton-based custom kernel to enable efficient end-to-end inference, reducing memory usage of the 12B FLUX.1-dev model by 2.1× and delivering 2.3× speedup over the BF16 baseline, on a 24GB RTX 4090 GPU. Our anonymous codebase is available [here](https://anonymous.4open.science/r/DiRotQ-5BF6/).
Abstract: Diffusion models have emerged as powerful tools for planning and control by learning multimodal distributions over actions and trajectories. Yet reliable inference-time safety enforcement remains a key barrier to their deployment in safety-critical tasks. Existing approaches typically project each denoising iterate onto the feasible set, even though constraints are defined only on the final clean trajectory. Enforcing feasibility on noisy intermediate samples can therefore overconstrain the sampling dynamics, substantially degrading sample quality. To address this limitation, we introduce DiRecT ( erminal constraints), a training-free algorithm for constrained sampling from diffusion models via stochastic optimal control (SOC). DiRecT enforces constraints only on the final clean sample, avoiding unnecessary restrictions on the intermediate denoising dynamics. Inspired by model predictive control, we derive a principled receding-horizon surrogate for the otherwise intractable constrained SOC formulation, yielding an efficient algorithm that cleanly separates stochastic denoising from constraint satisfaction, progressively steering samples toward feasible final trajectories without distorting the learned diffusion dynamics. Furthermore, DiRecT is highly flexible: it can leverage off-the-shelf or domain-specific optimizers, incorporate priors over environment dynamics, and optimize additional soft rewards. Extensive experiments on safe planning benchmarks demonstrate that DiRecT substantially improves deployment safety and task performance over existing diffusion-based planning baselines.
Abstract: Mixture-of-Experts (MoE) based Large Language Models (LLMs) have demonstrated impressive performance and computational efficiency. However, their deployment is often constrained by substantial memory demands, primarily due to the need to load numerous expert modules. While existing expert compression techniques like pruning or merging attempt to mitigate this, they often suffer from irreversible knowledge loss or high training overhead. In this paper, we propose a novel expert compression paradigm termed expert replacing, which replaces redundant experts with parameter-efficient modules and recovers their capabilities with low training costs. We find that even a straightforward baseline of this paradigm yields promising performance. Building on this foundation, we introduce LightMoE, a framework that enhances the paradigm by introducing adaptive expert selection, hierarchical expert construction, and an annealed recovery strategy. Experimental results show that LightMoE matches the performance of LoRA fine-tuning at a 30% compression ratio. Even under a more aggressive 50% compression rate, it outperforms existing methods and achieves average performance improvements of 5.6% across five diverse tasks. These findings demonstrate that LightMoE strikes a superior balance among memory efficiency, training efficiency, and model performance.
Abstract: Influence functions are a popular tool for measuring the impact of a training point on a model's prediction; in particular, self-influence, which quantifies the influence of a training point on itself, has found many uses such as data selection and outlier detection. However, the use of influence functions has been severely limited in modern models such as LLMs due to their low accuracy or high computational cost: most existing algorithms are either highly inaccurate, or require computing gradients or approximations to inverse Hessians, which can be prohibitive for large models. In this work, we introduce a highly efficient zeroth-order approximation to data influence that requires only a fraction of the time and memory footprint of previous methods and is applicable to both differentiable and non-differentiable loss functions. We demonstrate that in addition to its computational advantage, our method delivers superior accuracy in estimating self-influence and comparable or better accuracy in estimating train-test influence for fine-tuned large language models, paving the way for broader and more practical application of influence-function techniques in state-of-the-art AI systems.
Authors:
Hyunsuk Chung, Caren Han, Seungyeon Ji, Jinwoo Kim, Eun-Jung Holden, Kyungreem HanAbstract: Multimodal foundation models integrate heterogeneous signals across modalities, yet it remains poorly understood how their predictions depend on specific internal feature groups and whether such reliance can be deliberately controlled. Existing studies of shortcut and spurious behavior largely rely on post hoc analyses or feature removal, offering limited insight into whether reliance can be modulated without altering task semantics. We introduce FiLoRA (Focus-and-Ignore LoRA), an instruction-conditioned, parameter-efficient adaptation framework that enables explicit control over internal feature reliance while keeping the predictive objective fixed. FiLoRA decomposes adaptation into feature group-aligned LoRA modules and applies instruction-conditioned gating, allowing natural language instructions to act as computation-level control signals rather than task redefinitions. Across text-image and audio-visual benchmarks, we show that instruction-conditioned gating induces consistent and causal shifts in internal computation, selectively amplifying or suppressing core and spurious feature groups without modifying the label space or training objective. Further analyses demonstrate that FiLoRA yields improved robustness under spurious feature interventions, revealing a principled mechanism to regulate reliance beyond correlation-driven learning.
PaperID: 2469, Poster
Abstract: Decoder-only transformers are trained only through a terminal next-token prediction loss, yet this loss constrains every intermediate hidden state through the fixed downstream computation. We formalize this constraint by studying layerwise loss-to-go functions: the terminal loss obtained by continuing a candidate hidden state through the remaining transformer blocks. Around successful validation trajectories, we show that the local second-order geometry of these functions is governed, up to low-loss residual terms, by a pullback Fisher operator on hidden-state space. Its spectrum identifies output-sensitive directions and approximately prediction-null directions, yielding a local observable subspace of the residual stream. For causal transformers, the same geometry induces a tokenwise curvature score: a Fisher-weighted sensitivity of the target logits to perturbations of each token's hidden state. This score vanishes outside the causal ancestor set of the target and is controlled by downstream Jacobian couplings, making it a loss-aware alternative to attention magnitude. We estimate these quantities using matrix-free Jacobian-vector and vector-Jacobian products and evaluate them across decoder-only language models on WikiText, OpenWebText, and FineWeb. Empirically, the induced geometry predicts perturbation sensitivity, supports nonuniform layerwise rank allocation, yields competitive structured token-pruning signals, and improves low-rank student recovery when added to stronger autoregressive distillation objectives such as reverse KL and skew KL. These results support a predictive-geometric view of transformer computation: near successful trajectories, the terminal loss induces a thin, anisotropic set of output-relevant hidden-state directions that can be measured and exploited for compression and distillation.
PaperID: 2470, Poster
Abstract: RAG systems often verify citations only after retrieval and shortlisting have selected which passages will be tested. A verifier calibrated on random passage--claim pairs is therefore applied to a different population: shortlisted pairs chosen by the retrieval policy. This makes support verification a post-selection inference problem. We introduce TRIDENT, a contract-scoped framework for post-selection evidence accountability. Safe-Cover emits structural-facet support certificates with query-level Bonferroni FWER control, or abstains when the logged contract cannot certify support. Pareto-Knapsack relaxes certification and uses the same frozen facet table as a token-efficient multi-hop QA selector. Across HotpotQA, 2Wiki, and MuSiQue, TRIDENT controls false certificates under deployed contracts, identifies higher-quality answer-linked slices when certificates are emitted, and improves the matched-budget multi-hop QA frontier. The accountable unit is therefore not a retrieved passage or generated citation, but a support claim calibrated under the policy that selected it.
PaperID: 2471, Poster
Abstract: Dense 2D/3D point tracking has been the dominant paradigm for modeling motion in dynamic scenes, but a point track is just a 3-DoF translation curve per pixel: it captures where pixels go, not the rotation of the underlying part, nor which pixels move together as one body. We propose MoSE3, the first feed-forward model that predicts dense SE(3) motion from monocular RGB video, producing full 6-DoF rigid transforms at every pixel in world space. Per-pixel SE(3) motion offers a richer view of how a scene moves: rotation, translation, and grouping all at once. Directly predicting SE(3) is challenging: rotations lie on a curved manifold that is ill-suited to Euclidean regression, and annotations for SE(3) are particularly difficult to acquire. To address these challenges, MoSE3 predicts per-pixel SE(3) through two jointly learned intermediates, 3D point tracks and rigidity embeddings, and recovers SE(3) by differentiably fitting transforms within each soft rigid cluster, enabling end-to-end prediction and supervision. To close the data gap, we introduce Art-Kubric, a large-scale synthetic dataset with dense SE(3) and rigidity labels for articulated scenes with rich physical interactions. MoSE3 achieves state-of-the-art SE(3) estimation at pixel, part, and object levels on both rigid and articulated benchmarks, and state-of-the-art 3D point tracking on three datasets, while showing strong generalization to real-world videos despite being trained solely on synthetic motion data.
PaperID: 2472, Poster
Authors: Francois Pachet
Abstract: Variable-order Markov models generate by backing off to the longest usable suffix of the generated history, while regular constraints describe finite horizon controls such as fixed endings, metrical patterns, and anti-copy rules. Existing belief-propagation methods give exact regular-constrained sampling for first-order Markov chains, but a first-order state merges histories that a variable-order generator deliberately keeps distinct. We give the corresponding variable-order construction: replace the Markov state by the sparse observed context state and compose this context graph with the regular constraint automaton. For a fixed context graph and automaton, inference is linear in the horizon and in the number of reachable product edges, without expanding to all |\mathcalV|^K histories; the same source-row interface handles reversible augmentation by inverse count lookup, without materializing transformed corpora. Tiny enumerated examples verify exact partition functions and conditional probabilities; Bach Prelude experiments show sparse regular-constrained sampling at K=6 and 12-key virtual transposition semantics from 592 stored events.
Authors:
Hao Wu, Shunhao Zhang, Shuai LüAbstract: Offline-to-online reinforcement learning (O2O RL) aims to improve offline RL agents with limited online interaction. Prior work has shown that, in conservative value-regularized methods such as CQL, pessimistic offline critics can underestimate values and hinder online fine-tuning. We show that this issue is more general: critic underestimation also appears in explicit policy-constraint methods such as TD3+BC, and can persist or become more severe through bootstrapped Bellman updates during fine-tuning. To address this problem, we propose Optimistic Q-value Adaptation (OQA), a framework that corrects underestimation at the Bellman-backup level. OQA constructs an additional offline-pretrained actor--critic pair and uses it together with the original pair to form a controlled optimistic backup that mitigates inherited pessimism during fine-tuning. Unlike trajectory-return calibration methods, OQA does not rely on return lower bounds and is not tied to CQL-style regularization. Across D4RL MuJoCo locomotion, Maze2D, AntMaze, and Adroit benchmarks, OQA instantiations based on TD3+BC and CQL consistently outperform the corresponding backbones and strong O2O RL baselines.
Authors: Kakei Yamamoto, Martin Wainwright
Abstract: We introduce and analyze Target-Induced Loss Tilting (TILT) for unsupervised domain adaptation under covariate shift. It is based on a novel objective function that decomposes the source predictor as f+b, fits f+b on labeled source data while simultaneously penalizing the auxiliary component b on unlabeled target inputs. The resulting fit f is deployed as the final target predictor. At the population level, we show that this target-side penalty implicitly induces relative importance weighting at the population level, but in terms of an estimand b^_f that is self-localized to the current error, and remains uniformly bounded for any source-target pair (even those with disjoint supports). We prove a general finite-sample oracle inequality on the excess risk, and use it to give an end-to-end guarantee for training with sparse ReLU networks. Experiments on controlled regression problems and shifted CIFAR-100 distillation show that \tilt improves target-domain performance over source-only training, exact importance weighting, and relative density-ratio baselines, with a stable dependence on the regularization parameter.
PaperID: 2475, Poster
Abstract: In modern multivariate problems, exploiting the low-rank structure of the coefficient matrix serves as a powerful dimension reduction strategy to enhance statistical efficiency and interpretability, but its effectiveness can be severely compromised by outliers. In this work, we propose a robust low-rank learning (RoLL) framework for a broad range of loss functions, implemented via a fast algorithm that exploits local restricted strong convexity. Theoretically, we establish non-asymptotic error bounds for the obtained fixed-point estimator (not necessarily the global or local optimum), showing that it achieves the minimax optimal rate. Furthermore, we develop a novel information criterion with finite-sample theoretical guarantees for model selection. Extensive experiments demonstrate the effectiveness of our proposed framework in the presence of data anomalies.
Authors: Valentin Dorseuil, Jamal Atif, Olivier Cappé
Abstract: Recent work in the privacy literature shows that sample-targeted membership inference attacks (MIA) significantly outperform untargeted approaches by a wide margin. Motivated by this observation, we address the following question: Can the privacy vulnerability of individual training points be assessed without training shadow models? We show that per-sample exposure to MIA is governed not only by a point's loss, but also by a data-dependent geometric measure. In the linear setting, we derive a closed-form decomposition of individual black-box MIA vulnerability into a population leverage score and a residual loss term, making explicit how sample-dependent geometry translates into privacy exposure. Since the final layer of most modern architectures is linear, we extend this framework to deep networks and propose a surrogate score operating on last-layer representations that requires only a single trained model and no shadow models. Empirical evaluations across diverse datasets and architectures show that our score outperforms loss and gradient-norm baselines at identifying the highest-risk points under state-of-the-art attacks, providing a computationally efficient and theoretically grounded tool for per-sample privacy risk assessment.
Authors: Robi Rahman, Sabiha Tajdari
Abstract: Hardware-enabled monitoring of GPU workloads underpins many proposals for AI compute governance, but if developers can defeat monitoring mechanisms, such schemes are unworkable. We evaluate the adversarial robustness of GPU workload classification using only zero-overhead, privacy-preserving NVML telemetry: content-agnostic signals that observe physical effects of computation without accessing model weights, training data, or hyperparameters. Across 5 rounds of monitor-evader iteration, we evaluate 20 evasion strategy families on 9 GPU models spanning 4 architecture generations. We develop a classifier that achieves 98.6% binary accuracy at identifying training workloads across the whole corpus, and 43-87% accuracy against the most challenging unexpected workloads even when they are adversarially disguised.
PaperID: 2478, Poster
Abstract: Reward-robust RL typically models misspecification by specifying an uncertainty set that constrains the discrepancy of plausible rewards from a base reward. However, such reward space discrepancies can be behaviorally irrelevant: under potential-based reward shaping (PBRS), many distinct rewards preserve policy ordering. This can introduce substantial redundancy into standard uncertainty sets and degrade optimization performance. We propose , which constructs uncertainty sets over PBRS equivalence classes by projecting each reward to a canonical representative, ensuring that the resulting set contains only rewards that induce behaviorally distinct policy rankings. We prove that this projection preserves the optimal robust value while shrinking the uncertainty set and improving empirical performance. Using the connection between robustness and regularization, we obtain a practical algorithm to solve the shaping-aware reward-robust RL problem and enjoy convergence guarantees under standard assumptions. Experiments on benchmarks spanning diverse task domains and levels of complexity show consistent improvements over representative robust RL baselines and exhibit improved robustness to reward perturbations.
PaperID: 2479, Poster
Abstract: We propose FiLM-CAM; a conditional-access image watermarking approach that enforces authorized encoding, decoding, and removal through keyed feature modulation. A secret key conditions a Feature-wise Linear Modulation (FiLM) mechanism that transforms the internal representation of the watermark signal before embedding, thereby coupling the payload to the key at the feature, rather than pixel level. The same key must be presented to condition the decoder to correctly recover the payload embedded in an image. To enforce conditional access, we train the system to maximize the performance gap between authorized and unauthorized use, ensuring reliable extraction under the correct key while inducing failure under incorrect keys. In addition, we design our encoder to estimate the watermark residual invariant to the presence of pre-existing watermarks. This design enables conditional watermark removal via simple re-encoding and subtraction of the residual using iterative refinement, contingent on knowledge of the secret key. We evaluate FiLM-CAM across multiple watermarking encoder–decoder backbones, demonstrating that keyed feature modulation serves as an effective, architecture-agnostic conditional access module (CAM) for image watermarking.
Abstract: Randomized Smoothing provides rigorous robustness guarantees for neural networks without architectural constraints, yet its adoption is limited by extreme computational costs. Standard RS requires tens of thousands of model evaluations per input and forces practitioners to commit to fixed sample sizes a priori. In this work, we present a novel meta-learning framework for anytime-valid certified robustness that adaptively deploys computational resources. By using a lightweight meta-learner to predict image-specific priors for a sequential E-process, we achieve a 20-fold reduction in sample complexity compared to traditional methods while maintaining rigorous statistical guarantees. Beyond raw efficiency, we demonstrate how anytime-validity enables adaptively allocating compute based upon application-specific risk thresholds, a form of resource triage impossible under classic certification frameworks. That this is achievable while also providing similar certification performance demonstrates that our approach provides a pathway for real-time, safety-critical certification deployments.
PaperID: 2481, Poster
Abstract: We propose a stable algorithm for solving maximum coverage problems under a cardinality constraint of k in a record-level stability model, where adjacent instances I~ I' differ by the incidence of one universe element. Our algorithm computes a negative-entropy-regularized maximizer of the capped concave relaxation over the hypersimplex and rounds it with a common-seed rounding map from hypersimplex correlated sampling. The resulting sensitivity bound is \min\2k, \alpha_k k/\lambda\, where \lambda is the regularization parameter and \alpha_k = O(\log k). For utility, the expected coverage has approximation factor 1-\frac1e with an additive error of order O(\lambda k \log \fracmk), where m is the number of input sets. We also report experiments on synthetic and real-world instances.
PaperID: 2482, Poster
Abstract: We study stochastic linear bandits with a linear safety constraint that depends on an unknown parameter, requiring every chosen action to be safe at every round with high probability, even though the safe decision set is initially unknown. Prior work achieves \widetildeO(\sqrtT) regret when the safety gap, i.e., the constraint slack at the optimal action, is strictly positive and known to the learner. However, only \widetildeO(T^2/3) regret is established in two distinct cases: when the gap is zero, and when the gap is positive but unknown. This raises two natural questions: is the \widetildeO(T^2/3) bound tight in the zero-gap case, and is knowledge of the gap necessary for \widetildeO(\sqrtT) regret when the gap is positive? We answer both. First, we establish an \Omega(T^2/3) lower bound for zero-gap instances, showing that the existing upper bound is tight in its dependence on the horizon. Second, we propose Epoch-SLUCB, an algorithm that attains \widetildeO(\sqrtT) regret whenever the safety gap is positive, without requiring its value to be known. We also provide numerical simulations to demonstrate the performance of our algorithm.
Authors:
Florian P Mahner, Johannes Roth, Ka Chun Lam, Mick Bonner, Francisco Pereira, Martin N HebartAbstract: Deep neural networks trained with different architectures, objectives, and datasets have been reported to converge on similar visual representations. However, what remains unknown is which visual properties models actually converge on and which factors may underlie this convergence. To address this, we decompose the object similarity structure of 162 diverse vision models into a small set of non-negative dimensions. To determine universal versus model-specific dimensions, we then estimate how often each dimension reappears across models. In contrast to model-specific dimensions, universal dimensions are more interpretable and more strongly driven by conceptual image properties, indicating the relevance of interpretability and semantic content as implicit factors driving universality across models. Differences in architecture, objective function, training data, model size, and model performance do not explain the emergence of universal dimensions. However, models with more universal dimensions also better predict macaque IT activity and human similarity judgments, suggesting that universality reflects representations relevant to biological vision. These findings have important implications for understanding the emergent representations underlying deep neural network models and their alignment with biological vision.
Abstract: Unbalanced optimal transport (UOT) provides a principled framework for modeling single-cell transitions and birth-death dynamics, but its high computational cost limits scalability to large-scale datasets. Although single-cell data often contain hierarchical annotations and known transition priors, existing UOT approximations rarely exploit this multiscale structure or prior knowledge. We introduce Multiscale Supervised Unbalanced Optimal Transport Flow Matching (MUST-FM), a simulation-free framework that scales UOT by leveraging hierarchical data structure. MUST-FM further supports an optional supervised formulation that incorporates transition priors, such as cell lineages, to guide the learning of displacement fields and mass variations. Experiments show that MUST-FM reduces computational overhead while achieving robust and biologically meaningful trajectory inference, enabling dynamic modeling of atlas-scale single-cell datasets.
Abstract: Mode connectivity has been widely studied, yet the role of the optimizer remains underexplored. We revisit it through optimizer-induced implicit regularization, asking how connectivity behaves when restricted to solutions constrained by a given optimizer. For two-layer ReLU networks, we show that solutions from a single optimizer - AdamW, Muon, or others in the Lion-\mathcalK family - form a connected set at sufficiently large width, a result not implied by prior work. We then characterize how optimizer-induced regions interact: at large width two different regions can be disjoint or overlap depending on regularization, while in our small-width example AdamW and Muon converge to disconnected zero-loss components separated by a provable loss barrier. Empirically, in GPT-2 pretraining, we observe same-optimizer paths preserve each model’s spectrum while cross-optimizer paths traverse a smooth transition. Our results reveal optimizer-dependent structure beyond classical mode connectivity literature.
Abstract: We propose a novel pipeline for extracting symbolic notation and semantic knowledge from images of sheet music. Our pipeline features the first large-scale neural model for optical music recognition (OMR) to operate sequentially on a system-by-system basis, following the horizontal lines of notation as they are read on the page, rather than treating the page as an undifferentiated image, enabling better scaling to arbitrarily long inputs. It is also the first OMR model capable of generating symbolic transcriptions that include embedded textual content, such as titles and annotations. The pipeline combines system-level segmentation with an autoregressive vision-LM to capture both local notation details and score structure. Across multiple datasets, our approach consistently outperforms prior state of the art. We also show that symbolic transcriptions complement visual inputs for frontier language models, improving their interpretation of dense musical documents. The result is new state-of-the-art performance in both OMR and downstream sheet music understanding. We will release data and code upon publication.
PaperID: 2487, Poster
Abstract: Scaling sequential recommendation to long user histories requires compressing diverse behavioral evidence into memories that can be scored efficiently against many candidates. We first provide evidence that real user histories exhibit multi-scale semantic structure: short-lived intent, medium-term interests, and long-term preferences can coexist in the same sequence. However, existing summarization-based models mainly optimize compact history compression and do not explicitly organize cached user memory across temporal scales, which can lead to temporal aliasing where distinct behavioral signals become entangled before prediction. We propose MARS, a multi-resolution user memory that writes the full history into recurrent state tracks initialized with different half-life priors. A sparse routing reader then materializes compact seed memories by selecting the relevant temporal resolutions for each seed, preserving fixed-size candidate scoring. Experiments across recommendation datasets show that MARS outperforms strong recommendation baselines, with gains most pronounced for users with long histories.
PaperID: 2488, Poster
Authors: Jun Li
Abstract: We derive a recursion for the expected squared Frobenius norm of the zonotope generator matrices that govern the decision boundary of a random Gaussian deep ReLU binary classifier. For a network of depth L with weight variances (\sigma_l=1..L^2), the boundary complexity is \Gamma_L = 8(1+1/\pi)\prod_l=1^L \sigma_l^2/2. He initialization \sigma_w^2 = 2 is the unique uniform weight variance under which \Gamma_L is preserved with depth, matching the Poole--Schoenholz mean-field critical value on a different observable. \Gamma_L is identified with an observable per-unit-length boundary-crossing density via Rice's formula. The relation is supported by Monte Carlo experiments.
PaperID: 2489, Poster
Abstract: Tool-using large language model (LLM) agents face two distinct security failures: unauthorized external actions and exposure of sensitive plaintext inside the runtime before any final output check can intervene. Existing defenses usually protect one boundary, either the planner/runtime or the action sink, and therefore do not by themselves secure both surfaces. We present SecureClaw, a dual-boundary architecture that places authorization at the effect sink and plaintext confinement at the read boundary. Sensitive reads pass through a trusted gateway that replaces raw values with opaque handles and, in the evaluated deployment, bounded summaries as an explicit declassification interface. Writes that change external state follow a PREVIEW\rightarrowCOMMIT protocol in which only a trusted executor may commit the exact canonical request authorized by policy. The runtime can still plan over summaries and symbolic references, but cannot directly dereference secrets or perform side effects. Across AgentDojo, AgentLeak, and Agent Security Bench (ASB), SecureClaw is the only defense we evaluate in a common harness that simultaneously retains usable task utility and achieves 0% attack success rate (ASR) on ASB, 0.64% ASR on AgentDojo, and 3.23% overall leak on AgentLeak's attacked parity lane, which measures final-output and internal-relay leakage.
PaperID: 2490, Poster
Abstract: Decision-making under uncertainty typically requires a single scalar score, forcing practitioners to commit to one uncertainty measure. However, different measures often capture complementary failure modes, and relying on a single measure may be insufficient for downstream tasks. We propose \emphVecUQ-OT, a label-free procedure that aggregates multiple scalar uncertainty measures into a single ranking score via entropy-regularized optimal transport. Rather than selecting a single measure, our approach combines signals of (possibly) different natures through non-additive fusion based on \emphmultivariate ranks. The construction is motivated by the theory of multivariate ranks via exact optimal transport and realized through a tractable entropic approximation that generalizes to unseen inputs without retraining. Experiments across synthetic, image, and text domains demonstrate that \emphVecUQ-OT provides stable performance across tasks, remains reliable when individual measures fail, and outperforms the natural additive fusion alternative.
Abstract: Aligning language models for both helpfulness and safety typically requires complex pipelines---separate reward and cost models, online reinforcement learning, and primal-dual updates. Recent direct preference optimization approaches simplify training but incorporate safety through ad-hoc modifications such as multi-stage procedures or heuristic margin terms, lacking a principled derivation. We show that the likelihood ratio of the optimal safe policy admits a closed-form decomposition that reduces safety alignment to a density ratio matching problem. Minimizing Bregman divergences between the data and model ratios yields , a family of single-stage loss functions, each induced by a convex generator, that provably recover the optimal safe policy. BSO is both general and simple: it requires no auxiliary models, introduces only one hyperparameter beyond standard preference optimization, and recovers existing safety-aware methods as special cases. Experiments across safety alignment benchmarks show that BSO consistently improves the safety--helpfulness trade-off.
PaperID: 2492, Poster
Abstract: Operator learning has attracted significant interest for its capability to approximate mappings between function spaces. A common assumption in this field is that training and test samples are drawn from the same probability measure. However, this assumption is often violated, leading to probability measure shifts. Under such shifts, standard operator learning methods often lead to degraded performance because the expected training error does not match the expected test error. In this paper, we propose importance-weighted operator learning (IWOL), a framework for learning operators under probability measure shifts in function spaces. We prove that, when the test measure is absolutely continuous with respect to the training measure, the expected test error is equivalent to an importance-weighted expected training error, with weights given by the Radon--Nikodym derivative. This result provides a theoretical basis for correcting measure shifts in operator learning, but importance weighting relies on this absolute-continuity condition and may yield unbounded weights. To overcome these limitations, we introduce relative importance weights defined with respect to a mixture of the training and test measures. The resulting weights remain well-defined even when absolute continuity fails, and they are uniformly bounded. These weights induce a relative importance-weighted training objective for neural operators under probability measure shifts. To estimate these weights, we present binary classifiers that take functions as inputs, defined through functionals modeled by kernel integral operators. This design enables discretization-invariant weight estimation in function space. Experiments on multiple PDE benchmarks and neural operator architectures demonstrate that the proposed method consistently improves performance under probability measure shifts.
PaperID: 2493, Poster
Authors:
jiawei Yang, Xu Tan, Matti KaistiAbstract: Autoencoders (AEs) are widely used in unsupervised anomaly detection (AD) due to their ability to model complex data distributions. However, their strong reconstruction capacity, while effective for capturing normal patterns, often enables them to also reconstruct anomalous objects, including both individual and collective anomalies, which reduces detection performance. This limitation is particularly severe for collective anomalies, which are frequently misinterpreted as normal clusters. To address this issue, we propose Layer Precision Reduction (LPR), a novel technique that systematically constrains the numerical precision of weights and biases in neural network layers to strategically reduce the learning capacity of neural networks, such as the reconstruction capacity of AEs. LPR encourages the network to prioritize dominant manifold structures while suppressing the reconstruction of rare or anomalous patterns. While LPR was primarily developed for AE-based detectors (the focus of this paper), it is applicable to any neural network. Simulations are conducted with 24 AD models across 30 real-world public datasets. LPR is applied to 10 deep AD models, including 5 unsupervised AE-based methods, 2 unsupervised non-AE-based methods, 1 semi-supervised method, 1 weakly supervised method, and 1 fully supervised method. LPR consistently improves all 10 models, most notably enhancing the best-performing unsupervised AE-based detector from an average AUROC of 72.12% to 83.06%, surpassing the strongest unsupervised non-AE-based competitor, which achieves an average AUROC of 77.36%, across the remaining 14 detectors. These results reveal a previously unexplored link between layer-level numerical precision and manifold pattern prioritization, demonstrating that LPR serves as a simple yet powerful mechanism for enhancing AD performance while offering broad applicability to diverse neural architectures and potential extensions to tasks beyond AD.
Abstract: Diffusion models cannot enforce hard constraints, yet applications in the physical sciences demand exact satisfaction of conservation laws, boundary conditions, and observational consistency. In this work, we identify a corrector kernel whose unique stationary distribution is the constrained marginal at each noise level, and approximate it by iteratively projecting through the denoiser and renoising via the forward kernel. The resulting Predict-Project-Renoise (PPR) algorithm enables sampling from pretrained diffusion models under hard constraints. Its three components are each necessary: projecting through the denoiser keeps samples close to the data manifold, while renoising and iterating drive samples toward the constrained marginal. On 2D distributions, the Kuramoto-Sivashinsky equation, and global weather forecasting with a 10^8-dimensional atmospheric model, PPR simultaneously achieves low constraint violations and high distributional fidelity, a combination that existing methods fail to deliver.
PaperID: 2495, Poster
Abstract: 3D semantic occupancy prediction provides dense geometric and semantic scene representations for autonomous driving, where both high accuracy and low response latency are crucial for safe downstream forecasting and planning. Existing methods usually exhibit different accuracy-latency characteristics. Performance-oriented models provide stronger geometric and semantic predictions, while deployment-friendly models provide faster responses. This motivates EvoOcc, a fast-slow evolutionary occupancy prediction framework that jointly exploits their complementary strengths to improve the accuracy-latency balance without changing the upstream architectures. EvoOcc models evolutionary occupancy prediction as a continuous time controlled dynamical system, enabling flexible adaptation to irregular prediction arrival of both fast and slow system. The framework further decomposes the evolution process into dual state evolution where ego state evolution for the deterministic coordinate shifts and scene state evolution for the genuine geometric and semantic scene changes. Extensive experiments on three different fast-slow system pairs and irregular prediction arrivals show that EvoOcc consistently improves over fast system baselines while maintaining much lower latency than the slow system, indicating a favorable accuracy-latency balance and stable behavior under practical inference situations.
Abstract: Causal video generators must predict from the past, but they need not learn only from it. In streaming autoregressive video diffusion, each emitted segment becomes a commitment that future segments must preserve. Standard training, however, only asks each causal state to explain the present. This creates what we call a representation-level planning gap: states that fit the current segment may discard identity, layout, and motion information needed for a consistent future. We introduce Video-Mirai, a training-only method that closes this gap without changing causal inference: the generator rolls out causally, a frozen foresight encoder reads the completed rollout non-causally, and a lightweight predictor distills the resulting stopped-gradient targets into causal states. Future frames supervise representations, never generator inputs. At inference, the encoder and predictor are discarded, leaving the original architecture, per-step FLOPs, and KV-cache behavior unchanged. Video-Mirai improves a strong Causal-Forcing baseline on 5-second VBench from 83.8 to 84.6 in terms of Total Score. On 30-second rollouts beyond the training horizon, subject consistency improves from 84.9 to 88.5 and background consistency from 90.2 to 91.9. Ablations identify future-conditioned targets as the key ingredient, and probes show that future frames become more decodable from current features. Causality should constrain inference, not representation supervision.
Authors: Zhongjun Zhang, Sean R Sinclair
Abstract: We study offline-online reinforcement learning in linear mixture Markov decision processes (MDPs) under environment shift. In the offline phase, data are collected by an unknown behavior policy and may come from a mismatched environment, while in the online phase the learner interacts with the target environment. We propose an algorithm that adaptively leverages offline data. When the offline data are informative, either due to sufficient coverage or small environment shift, the algorithm provably improves over purely online learning. When the offline data are uninformative, it safely ignores them and matches the online-only performance. We establish regret upper bounds that explicitly characterize when offline data are informative, together with nearly matching lower bounds. Numerical experiments further corroborate our theoretical findings.
PaperID: 2498, Poster
Abstract: Feed-forward 3D reconstruction from unconstrained multi-view images has recently emerged as a dominant paradigm in 3D computer vision. Various foundational reconstruction models have been proposed with different architectures and reconstruction paradigms. Despite being trained on large-scale perspective datasets with substantial computational resources, these models generally fail when deployed on images whose camera models deviate from the training distribution. Applying such models to downstream applications with arbitrary, unknown, and potentially mixed camera types remains challenging. To maximally leverage the strengths of existing reconstruction models with minimal retraining efforts, we propose a low-rank adaptation framework that efficiently extends perspective-pretrained models to fisheye and panoramic imagery. Extensive experiments demonstrate that, with only 0.1% additional parameters and one day of training, our method can adapt a foundational 3D reconstruction model to non-pinhole cameras while fully preserving its original performance on perspective inputs.
PaperID: 2499, Poster
Authors:
Nan Sun, Yuan Zhang, Yongkun Yang, Wentao Zhao, Peiyan Li, Jun Guo, Wenxuan Song, Pengxiang Ding, Runze Suo, Yifei Su, Xin Xiao, Xinghang Li, Huaping LiuAbstract: Embodied chain-of-thought (CoT) aims to bridge linguistic reasoning with robotic control, yet its effective form and integration remain underexplored. In this paper, we revisit embodied CoT for robotic control at an unprecedented scale. We curate the largest embodied CoT corpus to date, comprising 978,743 trajectories, 226.3M samples, and 2592.5 hours of data. Through extensive experiments, we show that effective CoT must ground high-level semantic understaning in concrete linguistic action guidance -- such as end-effector movement and image-space trajectories -- whereas high-level reasoning alone yields marginal gains. More importantly, we identify that explicit CoT does not scale reliably as an autoregressive action prefix, suffering from compounding errors during inference. To address these challenges, we propose ERVLA, a vision-language-action (VLA) model that effectively leverages linguistic reasoning in generalizable robot manipulation. ERVLA is trained using a CoT-dropout strategy, allowing the model to leverage rich reasoning traces during training while predicting actions directly without CoT during inference to bypass autoregressive instability. This approach enables reliable scaling with increasing pre-training data. ERVLA achieves state-of-the-art results on LIBERO-Plus with an 86.9% success rate and reaches 53.2% on VLABench, showcasing superior performance in out-of-distribution settings. Furthermore, ERVLA outperforms competitive state-of-the-art baselines in real-robot experiments, especially in handling semantic ambiguity and long-horizon tasks. Code, data, and model checkpoints will be released.
Abstract: We introduce sparse autoencoder neural operators (SAE-NOs), a new class of sparse autoencoders that operate in function spaces rather than fixed-dimensional Euclidean representations. We formalize the functional representation hypothesis, where data are explained through sparse compositions of structured functions. Unlike standard SAEs that represent concepts with scalar activations, SAE-NOs parameterize concepts as functions, enabling representations that capture not only a concept's presence, but also how and where it is expressed across the input domain. We achieve this through joint sparsity: concept sparsity selects active concepts, while domain sparsity governs where they are expressed. We instantiate this framework using Fourier neural operators (SAE-FNOs), parameterizing concepts as integral operators in the Fourier domain. This functional and spectral parameterization is particularly advantageous when data exhibit spatial structure across scales or when concepts are frequency-structured. We characterize SAE-FNO on vision data and demonstrate that it learns localized patterns, uses concepts more efficiently, and exhibits stable concept characteristics across sparsity levels. We further show that SAE-FNO adapts to changes in domain size and generalizes across discretizations, operating at resolutions beyond those seen during training, where standard SAEs fail. We also introduce lifting into SAEs and show theoretically and empirically that it acts as a preconditioner that accelerates optimization. Overall, our results show that moving from vector-valued to functional parameterizations, with concept and domain sparsity, extends SAEs from representing concept presence to modeling structured concept expression, highlighting the importance of parameterization.
Authors: Gabriel Bathie, Nathanaël Fijalkow
Abstract: Synthesizing Mixed-Boolean Arithmetic (MBA) expressions from input-output examples is central to program deobfuscation and also useful for compiler optimization, reverse engineering, and cryptanalysis. Existing MBA synthesizers are typically CPU-based and scale poorly on large specifications or complex targets. Recent GPU-accelerated synthesis methods achieve large speedups in qualitative settings, but they depend on caching observationally equivalent candidates; this strategy breaks down for MBA because candidate outputs are quantitative bitvectors and the behavioral space is enormous. We present SIMBA (Synthesis of Mixed-Boolean Arithmetic), a GPU-accelerated MBA synthesizer built around cache-free bottom-up enumeration. SIMBA avoids language caches entirely and uses a GPU-oriented enumeration design that keeps work local and highly parallel. In experiments, SIMBA is substantially faster than prior MBA synthesis tools, handles larger specifications, and reaches expression sizes that existing methods fail to solve. These results establish cache-free GPU synthesis as a practical and scalable approach for quantitative domains, and identify it as a strong alternative to cache-centric designs.
Abstract: Accounting for privacy loss under fully adaptive composition---where mechanism choice and privacy parameters may depend on the history of prior outputs---is a central challenge in differential privacy (DP). Here, privacy filters are stopping rules ensuring a prescribed global budget is not exceeded. A leading candidate for optimal filter design is 𝑓-DP, which characterizes the full extent of adversarial hypothesis testing and recovers (ε,δ)-DP through piece-wise linear trade-off functions, while enabling tight (ε,δ)-DP accounting in standard compositions via tensor products. Yet whether such filters can be correctly defined under 𝑓-DP remains unclear. We show that the natural 𝑓-DP filter---tracking path-wise accumulating tensor products and stopping when the prescribed curve is crossed---is fundamentally invalid, precluding the direct use of standard efficient numerical Fast-Fourier-Transform accounting in the fully adaptive setting. We characterize this failure, establishing necessary and sufficient conditions for the natural 𝑓-DP filter's validity. Furthermore, we prove a fully adaptive central limit theorem for 𝑓-DP, establishing Gaussian convergence of cumulative privacy losses under full adaptivity. As a demonstration, we construct a closed-form approximate GDP filter for subsampled Gaussian mechanisms that provably outperforms RDP-based accounting in asymptotic regimes (q\ll 1 and q\approx 1) without tracking the full trade-off function, demonstrating that the slack in RDP is not intrinsic to adaptive composition---though CLT-based approximations are known to be optimistic at realistic subsampling rates, a limitation that remains an open challenge.
Abstract: With the continued advancement of text-to-image (T2I) generation, producing high-quality images is becoming increasingly attainable; consequently, user demands are shifting toward images that better satisfy their specific requirements. As reward models play an increasingly important role in assessing whether generated images align with user preference, this trend introduces an important challenge for reward modeling: rather than relying solely on static and general evaluation dimensions, reward models should account for the through which users assess whether generated images meet their specific requirements. To address this challenge, we propose , a dynamic, criterion-aware reward model that grounds task-relevant criteria and performs criterion-aware preference comparison. To support this setting, we construct , a benchmark for systematically evaluating reward models under dynamic criteria. We further introduce , which applies criterion-aware reward modeling to selecting T2I images. Our contributions establish the first reward modeling framework for dynamic and fine-grained evaluation and practical application in T2I generation.
PaperID: 2504, Poster
Authors: Dohyun Kim, Hyungryul Park, Kyongwon Kim
Abstract: We develop an online directional regression for sufficient dimension reduction with streaming data. Unlike first-moment methods, directional regression exploits both inverse conditional means and inverse conditional variances, and can therefore recover central subspace directions generated by linear structures, symmetric dependencies, and interaction driven relationships. Extending directional regression to the online setting is nontrivial because its kernel depends on slice-wise inverse moments and on standardization by an unknown covariance matrix. We address these challenges by utilizing a stable slice probability free kernel, a ridge-stabilized recursive least-squares estimator of the standardization operator, metric-aware subspace tracking under a vanishing-ridge covariance metric, and a weighted-bootstrap ladle criterion for online structural dimension selection. Under standard regularity conditions, the online kernel estimator is root-t consistent, the estimated generalized eigenspace consistently recovers the central subspace, and the proposed dimension selector consistently estimates the structural dimension. Simulation studies and real-data analyses demonstrate that the proposed method serves as a robust online alternative when the dominant data structure is unknown, adapting to both first-moment and second-moment-driven dependencies while yielding substantial computational savings relative to repeated batch directional regression.
PaperID: 2505, Poster
Abstract: State space models (SSMs) provide an efficient route to global image restoration because they model long sequences with linear complexity. Recent restoration architectures further enlarge the effective context by reordering image tokens so that semantically similar pixels become neighbors in the 1D scan. This strategy improves texture aggregation, but it also exposes a mismatch between the geometry of images and the discretization used by SSMs: adjacent tokens in the reordered sequence can be spatially distant, while the recurrent update still treats them as consecutive samples of a continuous process. In addition, existing prompt-based modulation relies on a finite set of discrete prompts, which is poorly matched to the continuous variation of natural textures and offers no explicit compensation for the low-pass tendency of recurrent integration. We propose GeoMamba, a geometry-aware and continuously modulated SSM for image restoration. GeoMamba introduces geometry-adaptive discretization (GAD), which conditions the SSM time-step on the Euclidean distance between consecutive reordered tokens and attenuates history propagation across spatial jumps. It also introduces a continuous dynamic matrix (CDM), which regresses pixel-wise modulation parameters from feature and gradient cues to reduce prompt quantization and preserve high-frequency details. Experiments on lightweight and classical image super-resolution demonstrate consistent improvements over strong Transformer- and Mamba-based baselines, while achieving better or competitive performance in additional denoising and JPEG artifact reduction settings.
Abstract: The ability to generate variable-length proteins is crucial in protein design, where the optimal length is often unknown and tightly coupled to structural designability. Current state-of-the-art diffusion and flow-based generative models require the protein length to be predetermined before sampling, limiting their flexibility in exploring the feasible protein design space. To bridge this gap, we introduce Generalized Poisson Flow (GPFlow), a novel generative framework that enables variable-length generative modeling by learning the rate function that minimizes the negative log-likelihood of an inhomogeneous generalized Poisson process. We establish theoretical guarantees for recovering the joint multimodal distribution (both continuous and discrete) and for an upper bound on the KL divergence to the generated distribution. We evaluate GPFlow extensively across various protein design tasks, including unconditional structure and sequence generation and conditional motif scaffolding, to validate GPFlow’s effectiveness on both continuous and discrete modalities. Our results demonstrate GPFlow’s superior generative performance and quality, achieving the top designability and distributional fitness on unconditional generation, and ranking first on 10 out of 16 motif scaffolding tasks, with up to a 10-fold improvement in success rates on the most challenging targets.
Abstract: Masked autoregressive models (MAR) have emerged as a powerful paradigm for image and video generation, combining the flexibility of masked modeling with the expressiveness of continuous tokenizers. However, when sampling individual frames, video MAR models often produce highly distorted outputs due to the lack of a structured global prior, especially when using only a few sampling steps. To address this, we propose CanvasMAR, a novel autoregressive video prediction model that predicts high-fidelity frames with few sampling steps by introducing canvas—one-step prediction of the future-frame expectation that serves as a non-uniform mask during masked generation. The canvas supplies global structure early in sampling, enabling faster and more coherent frame synthesis. To further stabilize autoregressive sampling, we propose an easy-to-hard curriculum via a motion-aware sampling order that synthesizes stationary regions before attending to highly dynamic ones. We also integrate compositional classifier-free guidance that jointly strengthens the canvas and temporal conditioning to improve generation fidelity. Experiments on the BAIR, UCF-101, and Kinetics-600 benchmarks demonstrate that CanvasMAR produces higher-quality videos with fewer autoregressive steps. On the challenging Kinetics-600 dataset, CanvasMAR achieves remarkable performance among autoregressive models and rivals advanced diffusion models.
Abstract: Despite surpassing human performance across mathematics, coding, and other knowledge intensive tasks, large language models (LLMs) continue to struggle to causally reason. A core obstacle is the target data itself: causal systems are complex and often expressed in non-executable forms, and ground-truth answers to causal queries are inherently scarce. We introduce, , a framework that turns causal reasoning from a scarce-label problem into a scalable, supervised one. It constructs increasingly complex causal simulators: executable structural causal models (SCMs), incrementally built by LLMs, that scale to globally complex systems while maintaining verifiable answers to any causal query. operates across representations by formalizing non-executable causal knowledge into code, allowing for data augmentation, and informalizing executable SCMs into natural language, enabling supervision in previously unsupervisable representations. We structure our research into two parts: (1) how to construct increasingly complex causal simulators, and (2) a systematic study of what enables, demonstrating generalization across representations, consistent gains from curriculum scaling and data volume, LLM self-improvement through self-generated simulators, and data augmentation via formalization of existing domain knowledge.
PaperID: 2509, Poster
Abstract: In stochastic (strongly) convex optimization, establishing last-iterate convergence guarantees for first-order methods has recently garnered increasing attention, as this not only deepens the theoretical understanding of these algorithms but also aligns more closely with practice. Notably, people have demonstrated that one of the most famous, simple, and popular methods, Stochastic Gradient Descent (\mathttSGD), provably achieves last-iterate convergence. However, the existing results may have limited applicability, since they typically rely on strong conditions on the gradient noise (e.g., an exponentially decaying tail) rather than the more realistic heavy-tailed noise. Unfortunately, when faced with heavy-tailed noise, \mathttSGD is known to exhibit undesirable behavior or even fail to converge. To address the challenge of heavy-tailed noise, people have proposed \mathttClipped\text-\mathttSGD, an algorithm that combines \mathttSGD with a simple mechanism, gradient clipping. Although \mathttClipped\text-\mathttSGD performs well in practice, its last-iterate convergence property remains highly underexplored. In this work, we prove the first optimal last-iterate rates in high probability for \mathttClipped\text-\mathttSGD under heavy-tailed noise, thereby closing a gap in the literature.
PaperID: 2510, Poster
Authors: Aidan Sirbu, Charles Y Zhang, Yuanjia Yang, Bence Olveczky, Talmo Pereira, Blake Richards
Abstract: The development of embodied models of animal behavior is a major goal of the neuro-AI community. Embodied models will enhance our ability to make direct comparisons between biological and artificial agents. To make such comparisons effective, though, we need embodied models that actually move like animals. Previous work has developed virtual animal bodies and techniques for training realistic motion through imitation. However, to-date, there is no system for embodied animal modeling that produces natural movements both during trained, goal-directed behavior, and during exploration driven by random generation. Here, we introduce Simulated Control via Augmented Motor Priors for Embodied Reinforcement learning (SCAMPER), a system for training an embodied virtual rat model, enabling natural movements during both goal-directed behavior and random exploration. SCAMPER uses a variational architecture with two distinct priors, a motion prior for control and a behavior prior for exploration. We show that combining both of these priors allows SCAMPER to generate realistic behavior during exploration and exploit the task structure effectively. As well, SCAMPER can learn from sparse rewards as a result of the diverse movements induced by the behavior prior, allowing the system to model realistic sparse reward tasks from experimental neuroscience. Altogether, SCAMPER provides the neuro-AI community with an effective new system for training embodied models of rodents on realistic tasks.
PaperID: 2511, Poster
Abstract: Generative models have made rapid progress in ordered crystal structure prediction, yet many functional materials are intrinsically disordered, with substitutional mixing, vacancies, or interstitial species controlling their properties. Existing crystal generators either assume deterministic site occupations or require site-level disorder annotations, which are often unavailable when the chemical formula is the primary input. We formulate disordered crystal structure prediction through an Occupancy Distribution Matrix (ODM), a continuous site-by-species representation that unifies ordered crystals, solid solutions, vacancy disorder, and interstitial occupancy. A valid ODM must satisfy coupled site-wise occupancy, mass-conservation, and non-negativity constraints, placing each sample on a formula-dependent transportation polytope. We propose Entropic Polytope Flow (EP-Flow), a marginal-constrained flow matching framework that canonicalizes heterogeneous polytopes into a shared double-centered space, learns a marginal-preserving flow, and recovers feasible occupancies through a Sinkhorn inverse map. By jointly generating occupancies, fractional coordinates, and lattice parameters, EP-Flow achieves state-of-the-art performance on formula-conditioned disordered CSP benchmarks derived from COD and MPDS, substantially outperforming adapted ordered-crystal generators. Analyses further show that EP-Flow recovers sparse and chemically meaningful local disorder patterns rather than merely matching global composition statistics.
PaperID: 2512, Poster
Authors: Konstantinos Ziliaskopoulos, Alexander Vinel, Jiaqi Wang
Abstract: Decision-focused learning (DFL) trains predictive models for downstream optimization, but existing methods largely assume centralized data. In cross-silo settings, federated learning is a natural alternative, yet standard federated pipelines optimize prediction quality rather than decision quality and do not account for client heterogeneity in downstream objectives or feasible sets. Heterogeneity is especially problematic for DFL methods: under heterogeneous polyhedral decision problems, small perturbations can cause discontinuous changes in optimal decisions, leading to unstable client updates and aggregation. We propose FedRSPO+, a heterogeneity-aware framework for decision-focused federated learning. Our approach is built on RSPO+, a regularized predict-then-optimize surrogate that smooths the decision map through projection, enabling stable optimization even when clients face different objectives and constraints. We show that RSPO+ upper bounds downstream decision error and regret under mild assumptions, and that FedRSPO+ has cross-client heterogeneity bounds that (i) scale with both objective and feasible-set heterogeneity, (ii) vanish at homogeneity, and (iii) do not require strong convexity. We develop an annealed and modular training procedure that is compatible with standard federated personalization and aggregation methods. We evaluate FedRSPO+ on three experimental settings: controlled synthetic knapsack instances that isolate multiple axes of heterogeneity, a shortest path DFL benchmark, and a real-world PJM energy pricing case study. Across these experiments, we compare against prediction-only federated learning and DFL frameworks under varying heterogeneity and communication budgets. Together, our results suggest that smoothing is a useful ingredient for stable collaborative decision learning and provide a first heterogeneity-aware foundation for federated DFL.
PaperID: 2513, Poster
Abstract: Conditional independence (CI) testing is a fundamental problem in statistics. However, classical nonparametric CI tests suffer from the curse of dimensionality, as nonparametric estimation degrades rapidly in high dimensions. This paper proposes RP-CIT, a random-projection-aggregated CI test that remains effective in the high-dimensional regime. The central idea is to project X and Y onto random one-dimensional subspaces before applying a univariate base CI test. This projection (i) preserves conditional independence under the null hypothesis and (ii) reduces the original high-dimensional problem to a collection of univariate CI subproblems, thereby avoiding high-dimensional nonparametric estimation. Since a single projection may miss the dependence signal, we aggregate the p-values from k independent projections via a Bonferroni-max rule, which asymptotically controls the Type~I error at the nominal level and yields a limiting Type~II error bound that decreases exponentially in k under a detectable-projection condition. The paper also develops a top-r aggregation extension for alternatives whose signal is spread across multiple projection directions. More broadly, the random-projection aggregation framework can be combined with other scalar base tests; as one example, we use a projected GCM base test to accommodate higher-dimensional conditioning variables. Experiments on synthetic benchmarks, image-valued CI tasks, and Norman Perturb-seq module skeleton learning illustrate the calibration, power, and practical utility of the framework.
Abstract: As multimodal LLMs increasingly emphasize video and audio, a common assumption is that solving such tasks calls for native omnimodal models. We show that this is not always necessary: coding agents equipped with only text+vision and a sandboxed tool-using interface can perform competitively with, and in several settings outperform, SOTA native omnimodal models and predefined multimodal agent scaffolds across multiple audio-video benchmarks. Our trajectory analysis suggests that their advantage comes from coding agents writing code and orchestrating tools to retrieve relevant content from transcripts, frames, and other non-textual modality signals from raw inputs. This effectively converts omnimodal tasks into evidence retrieval and information processing problems, avoiding the inefficiency of ingesting entire videos or audio streams into context. To characterize remaining limitations, we propose a failure taxonomy and a process-level analysis of tool-use traces, and find that simple skill injection, including human-written skills and self-distilled skills from execution logs, can substantially improve performance over the no-skill baseline. To examine whether such capability can be elicited in open-source models, we further introduce Code-X, a complete training recipe with the OmniCoding trajectory dataset and verifiable reward, providing an exploratory baseline on Qwen-3.5-9B and Qwen-3.6-27B. Finally, given the maturity of many-modality understanding, we argue that the more meaningful frontier lies in many-modality processing, and introduce TerminalBench-O, the first process-level benchmark designed for coding agents on real-world omnimodal processing tasks. Together, our findings open new directions for omnimodal content processing and evaluation.
PaperID: 2515, Poster
Authors:
Xiaoge Zhang, Zijie Wu, Mingtao Feng, Saeed Anwar, Ajmal MianAbstract: 3D Gaussian Splatting (3DGS) has become an established representation for real-time novel view synthesis. However, preserving fine geometric structures and appearance details typically requires millions of Gaussian primitives, resulting in substantial storage and transmission overhead. 3DGS compression has been extensively studied through attribute quantization, entropy coding, and lightweight reparameterization. However, existing methods often operate on a fixed backbone and therefore do not explicitly model the joint redundancy between geometry and appearance in the original dense Gaussian set. To address this limitation, we propose Pyramidal Gaussian Mixture Splatting (PGMS), a plug-and-play 3DGS compression framework that reformulates dense Gaussian primitives into a rate-controllable pyramidal representation with scale-dependent attributes. Specifically, PGMS first constructs a mixture pyramid by progressively partitioning dense Gaussians into hierarchical mixture centers, where different levels perform attribute-specific clustering to form shared prototypes. It then employs a rate-adaptive capacity allocator to determine the tiered budget under a target compression ratio. To further improve rendering quality, we introduce a compositing-aware parameter update rule that incorporates information from dense Gaussians into pyramid primitives according to their image-space contributions. Experiments on standard benchmarks show that our method delivers strong compression with high rendering fidelity and consistently outperforms state-of-the-art 3DGS compression methods.
Authors: Chia-Ying Lin, Nai-Hui Chia, Shih-Han Hung
Abstract: Learning the closest matrix product state (MPS) representation of a quantum state is known to enable useful tools for quantum machine learning and analysis of complex quantum systems. In this work, we study the problem of learning MPS in the following setting: given many copies of an input MPS, the task is to recover a classical description of the state. The best known polynomial-time algorithm, introduced by [LCLP10, CPF+10], requires linear circuit depth and O(n^5) samples, and has seen no improvement in over a decade. The combination of linear circuit depth and large sample complexity, neither known to be optimal, renders existing algorithms impractical for near-term quantum devices with limited resources. We introduce parallel disentangling algorithms for MPS learning. For exact MPS learning, our algorithm runs in polynomial time and uses circuit depth O(\log n) and sample complexity \widetilde O(n^3), improving both the depth and the dependence on the system size n. The key idea is to exploit the bounded-rank structure of reduced states on middle blocks of an MPS and organize the disentangling operations in a tree structure. We further extend the algorithm to closest MPS learning, improving the sample complexity dependence on n from n^9 to n^7 and complement the algorithms with an \Omega(n) product-state lower bound. We also investigate MPS learning under hardware constraints, including restricted measurements and geometric connectivity. Under the Learning Parity with Noise (LPN) assumption, we show computational hardness for learning an MPS(2) family with non-adaptive single-qubit measurements. Finally, we show that our algorithm can be implemented with depth O(q n^1/q) on a q-dimensional hypercubic lattice, giving an asymptotic reduction in depth. Together, our work provides a complete characterization of the quantum resources needed for efficient MPS learning.
PaperID: 2517, Poster
Abstract: Self-play can train language models without external supervision. However, existing methods require rule-checkable answers, leaving open-ended tasks dependent on curated prompts or frontier-model judges. We introduce SCOPE, a data-free self-play framework for open-ended tasks that co-evolves two policies: a Challenger that generates document-grounded tasks, and a Solver that answers them through multi-turn retrieval. A frozen copy of the initial model serves as the self-judge, which writes task-specific rubrics from the source document and grades Solver responses against them. Across three 7-8B instruction-tuned models (Qwen2.5, Qwen3, OLMo-3), SCOPE improves open-ended performance by up to +10.4 points on eight benchmarks and matches or exceeds GRPO data trained on ~9K curated prompts. Although trained only on open-ended tasks, SCOPE also improves held-out short-form QA by up to +13.8 points on seven held-out benchmarks, surpassing GRPO data on all three models. Ablations show that co-evolving the Challenger is necessary to keep tasks near the Solver's frontier, that gains arise from improvements in both retrieval and synthesis with the relative contribution varying by task, and that rubric generation quality is the bottleneck for self-judging.
Abstract: As large language model (LLM) agents move from isolated prompting to long-horizon workflows, failures increasingly arise at the role-to-instance binding boundary, where task-specific role requests must be assigned to concrete agent instances under current service, network, and query conditions. Existing agent system research has improved role specialization, workflow topology, memory, and tool use, but often assumes a fixed stable execution environment. This assumption limits deployed reliability, because the same role request can exhibit different latency, failure probability, and output quality across agent instances operating under different service regions and network conditions. We propose Hedged Agent Computing (HACO), a runtime control scheme that treats each role request as a reliability-constrained selection problem over candidate agent instances, each coupling a role type, an LLM, and a concrete execution environment. Different from routing, HACO adaptively selects a hedge set of candidates for each invocation. Its allocation rule combines optimistic ranking, which prioritizes candidates with high estimated quality, reliability, and informative uncertainty, with conservative reliability accumulation, which stops selection only after the hedge set reaches a target success probability. Through experience harvesting, HACO updates candidate and link profiles from all executed candidate traces, including quality, success, latency, and network statistics. Experiments on various benchmarks, together with runtime degradation studies, show that HACO improves robustness and output quality under changing deployment conditions, while using lower token and latency cost than exhaustive parallel execution.
Abstract: Recent advances in Reinforcement Learning (RL) have underscored its potential for incentivizing reasoning capabilities of Large Language Models (LLMs). However, existing step-level efforts suffer from costly annotations that limit domain coverage, while scalar scores further impose an information bottleneck, offering insufficient semantic bandwidth to improve intermediate decisions. Alternative language-critique approaches, which rely on frozen or external critics, provide richer textual feedback but lack the scalability needed for sustained policy improvement. In this work, we propose language-driven stepwise trajectory redirection, termed as stepwise language feedback. Specifically, we co-train a generator and a generative verifier using only outcome-based rewards, eliminating external annotations, while delivering sustained policy improvement through jointly aligned verifier training. The verifier's stepwise language critiques explicitly localize and explain failures, enabling the generator to redirect reasoning trajectories at intermediate steps toward alternative decisions. The trajectory redirection design guarantees harmless policy improvement, even under noisy or suboptimal verifier feedback. Experiments on diverse reasoning benchmarks show that STRIDE significantly outperforms state-of-the-art baselines, as well as achieving breakthroughs on zero-pass-rate problems where scalar methods yield no learning signal in our ablation studies, demonstrating the effectiveness of learnable stepwise language feedback for enhancing LLM reasoning.
Abstract: Recent 3D generative models can synthesize high-quality assets, but their outputs are typically static: they lack the skeletal rigs, joint hierarchies, and skinning weights required for animation. This limits their use in games, film, simulation, virtual agents, and embodied AI, where assets must not only look plausible but also move plausibly. We introduce Rigel3D, a generative method for animation-ready 3D assets represented as rigged meshes. Unlike post-hoc auto-rigging methods that attach rigs to completed shapes, our method jointly models geometry and rig structure through coupled surface and skeleton structured latent representations. A rig-aware autoencoder decodes these representations into mesh geometry, skeleton topology, joint coordinates, and skinning weights, while a two-stage latent generative model synthesizes both surface and skeleton representations for image-conditioned generation. To support downstream animation workflows, we further introduce an open-vocabulary joint labeling module that embeds generated joints into a shared vision-language space, enabling correspondence to arbitrary retargeting templates. Experiments on large-scale rigged asset datasets demonstrate that our method generates diverse, high-quality animation-ready assets and outperforms existing rigging baselines across multiple metrics.
PaperID: 2521, Poster
Abstract: Asynchronous reinforcement learning (RL) is an effective paradigm for improving the training efficiency of large language model (LLM) post-training by decoupling rollout generation from policy optimization. However, this decoupling introduces stale off-policy samples generated by earlier behavior policies, leading to distribution shift and unstable policy updates. Recent replay selection methods such as D-ARL mitigate this issue by selecting variance-aware samples, but require expensive current-policy evaluation over replay responses and mainly rely on variance reduction or current-policy matching as the selection criterion. To address these limitations, we propose GA-ARL, an efficient Gradient-Aware Asynchronous Reinforcement Learning framework for LLM post-training. GA-ARL derives an analytic target distribution from the KL-regularized RL objective and introduces an advantage-weighted variant to account for both policy preference and gradient strength. During training, GA-ARL maintains a replay buffer containing samples from recent behavior policies and selects low-discrepancy samples according to the advantage-weighted analytic target, avoiding additional current-policy evaluation during replay scoring. GA-ARL outperforms SOTA asynchronous methods across mathematical, logical, and code reasoning benchmarks, achieving the best average accuracy on both Qwen3-1.7B and Qwen3-4B, with improvements of up to 6.1%. Meanwhile, compared with the SOTA D-ARL, GA-ARL reduces wall-clock training time by 25.3% on average by avoiding additional current-policy evaluation during replay selection.
PaperID: 2522, Poster
Abstract: In this paper, we study the problem of simultaneous inference for multiple quantiles under local differential privacy (LDP). We develop a general online framework that reduces the multiple-quantile problem to a categorical one-hot estimation problem. A projected Robbins-Monro recursion guarantees noncrossing at every iteration, while Polyak-Ruppert averaging together with multivariate self-normalization yields a joint asymptotically pivotal ellipsoidal confidence region without requiring estimation of nuisance densities or asymptotic covariance matrices. We also present several important examples in detail, including direct k+1-ary randomized response (kRR), optimized unary encoding (OUE), optimized local hashing (OLH), and Hadamard response (HR). We evaluate our methodology through numerical experiments based on these examples, and the results provide positive support for the proposed framework.
Abstract: Can we trust evaluation scores to capture an LLM's true real-world performance? Certifiable evaluation answers this question by providing guarantee for LLM evaluation. In particular, existing methods sequentially curate evaluation samples and keep updating confidence intervals (CIs) that cover the true performance with high probability (e.g., 95%) until some conditions are satisfied, e.g., the CI width reaches a target precision. However, existing methods are not generally anytime-valid: the claimed coverage (e.g., 95%) may fail when CIs are repeatedly updated and used to decide when to stop, leaving a gap between theoretical rigor and practice. This paper bridges this gap by proposing Celeus, a Certifiable framework for Efficient LLM evaluation, which leverages E-processes to build anytime-valid CIs. Concretely, we propose signals that combine two ingredients: (i) Uncertainty-guided sampling to select informative samples for evaluation, and (ii) Surrogate-assisted approximations for unevaluated samples. We prove that such signals remain unbiased for the evaluation score conditional on the past, enabling statistically grounded and anytime-valid e-process CIs. More importantly, the two ingredients reduce estimation variance and help reach the target precision with fewer evaluated samples. We also prove that CIs obtained by Celeus can shrink at a near-parametric rate up to logarithmic factors and analyze the oracle variance-optimal sampling rule that motivates the empirical uncertainty-guided one. Experiments show that Celeus reaches the target precision using 54-62% fewer evaluated samples than baselines, while preserving anytime-valid coverage.
Abstract: Feed-forward models for 3D reconstruction have achieved strong performance using deep cross-view attention to exchange information across images. However, these approaches often depend on heavy decoder stacks and lack a structured mechanism for geometry refinement, resulting in poor multi-view consistency. We address this by drawing inspiration from classical bundle adjustment (BA), which can be viewed as an iterative information propagation process between poses and local geometry. Inspired by BA, we propose BA-T, an iterative Transformer that implements BA-style structured updates as a repeatable layer in implicit token space. Instead of relying on deep attention stacks, BA-T refines predictions based on latent residual by a single lightweight layer. Experiments demonstrate that BA-T progressively improves pose and reconstruction accuracy across iterations, achieves stronger cross-view consistency than conventional decoders, and matches or surpasses substantially larger models while using only 16% of their decoder parameters. BA-T provides a compact, efficient, and structural alternative to depth-heavy attention, enabling accurate 3D reconstruction within a lightweight architecture. The code will be made publicly.
Authors: Heikichi Hayashi, Ruoran Lai, Jiazhuo Li
Abstract: We study a partially observed extensive-form trading game around a constant product market maker with one hidden order, public trade direction, and an interval valued signal for the hidden order size. Agents compete in an execution-rights auction to submit robustly admissible sandwich bundles, meaning bundles that remain valid for every hidden order size consistent with the public signal. We give a closed-form admissibility frontier, show that the maximal admissible front-run is simultaneously optimal for pointwise, expected, and worst-case profit, and derive an explicit worst-case profit formula. This yields an exact distribution-free leakage threshold for profitable sandwiching in the presence of fixed execution cost. We then show that, with at least two symmetric traders, every pure-strategy perfect Bayesian equilibrium of the execution-rights auction implements the optimal robust bundle and transfers all positive continuation rent to the auctioneer. Perfect hiding of any positive lower bound eliminates universally profitable robust sandwiching in our model, although post-trade arbitrage may remain. The results isolate an exact threshold phenomenon for strategic trading under partial observability.
PaperID: 2526, Poster
Abstract: We introduce SemGeo-Gen, an unsupervised method for generating approximate cross-instance semantic-geometric supervision from weakly aligned 3D object collections. Our goal is not to recover exact point-level correspondences, which are often ambiguous across different instances and would require expensive, impractical manual supervision at scale. Instead, we automatically produce approximate correspondences that are semantically meaningful, geometrically consistent, and diverse enough to train modern correspondence models. Given only coarse category-level rotation alignment, SemGeo-Gen lifts multi-view DINOv2 features to 3D, aligns object instances using continuous piecewise-affine registration, and prunes candidate matches using semantic consistency. The resulting 3D correspondences can be projected into rendered views, yielding scalable 3D-3D, 3D-2D, and 2D-2D supervision without manual annotation. We validate the generated correspondences against sparse human annotations from KeypointNet, obtaining 93% PCK@0.10 and 97% PCK@0.15. More importantly, we show that using the generated correspondences for synthetic pretraining improves recent state-of-the-art semantic correspondence models on SPair-71k after fine-tuning on real data, with gains of up to 5 PCK@0.10 points. Together, these results indicate that approximate semantic-geometric supervision generated from 3D assets can improve real-image correspondence learning and serve as a scalable alternative to costly manual annotation. Code and generated data will be released upon acceptance.
PaperID: 2527, Poster
Authors:
Yiwei Dai, Hengyi Cai, Hui Wu, Yili Wang, Han Xu, Minlan Shao, Yuchen Li, Shuaiqiang Wang, Xin Wang, Yi Chang, Dawei YinAbstract: Search agents based on large language models address complex information-seeking tasks by iteratively issuing queries, retrieving evidence, and integrating information. Effective training of such agents requires supervision beyond final-answer correctness, as outcome-level rewards cannot distinguish efficient evidence acquisition from redundant or misdirected search. Recent process-level methods introduce intermediate supervision, but typically apply uniform signals across trajectories. However, such uniformity overlooks the state-dependent nature of process quality, where the value of an intermediate decision depends on the current evidence state. This motivates a systematic analysis of how process quality varies across trajectories with different evidence states and final outcomes. To this end, we analyze search trajectories on BrowseComp-Plus and xBench, revealing two consistent patterns: successful trajectories often suffer from post-alignment inefficiency (redundant searches), whereas failed trajectories typically struggle with misdirected exploration and insufficient coverage. Motivated by these findings, we propose ), a framework for search agents that provides a concrete instantiation of state-dependent process supervision. partitions trajectories according to final outcomes, assigns differentiated process rewards to successful and failed groups, and integrates them with outcome supervision at either the reward or advantage level. Experiments on three backbones across six multi-hop QA and web-search benchmarks show that
PaperID: 2528, Poster
Abstract: Ubiquitous IoT devices, such as companion toys, wearables, and micro-robots, are increasingly embracing language models, yet current cloud-based deployments suffer from latency, privacy, and weak-network reliability issues. These resource-constrained devices typically cap deployable model size below 100M parameters, while existing sub-100M models fall short in basic language and commonsense capabilities. In this extreme regime, we argue that the memory bottleneck supersedes the traditional compute bottleneck as the binding constraint, calling for memory-first design that maximizes model capability under a fixed parameter budget. Notably, the stringent budget binds both what the model stores and what it can afford to learn, demanding architecture-data co-design. In this paper, we propose \muLM, a memory-first sub-100M language model framework instantiating this co-design. For architecture, we present DMLA, which preserves latent attention's compact cache while relieving its feature-squeezing bottleneck under aggressive compression via explicit content-position latent decoupling. For data, we present \muMix, a capacity-aware multi-stage curriculum tailored to sub-100M models, prioritizing high-quality natural language and commonsense data aligned with embedded scenarios over brute-force token scaling. \muLM-70M surpasses all sub-300M baselines and rivals several sub-1B models on standard benchmarks; even at 9M, it remains competitive. Fine-tuned into a tool-calling agent, \muLM outperforms same-scale baselines and achieves a 97.2% end-to-end executable rate, presenting a promising direction toward practical sub-100M language modeling.
PaperID: 2529, Poster
Abstract: Millimeter-wave (mmWave) radar provides core sensing capabilities to traditional vision, particularly under occlusion, low-light, and privacy-constrained conditions. Recent efforts have explored integrating radar signals with large language models (LLMs) for high-level semantic reasoning over the received radar data. However, existing approaches typically project radar inputs into borrowed vision or LiDAR embedding spaces, which are not tailored to radar-specific physical cues such as Doppler and reflection intensity. This design introduces a representational bottleneck, hindering the extraction of informative radar semantics. To overcome this limitation, we propose mmLIP, a novel radar-language interactive pretraining framework that learns radar-specific alignment without relying on external embedding spaces. Our approach directly aligns radar representations with the text embedding space via a point-level contrastive objective, enabling fine-grained correspondence between radar points and textual tokens. In addition, we introduce a confidence-aware contrastive learning mechanism that adaptively reweights radar tokens based on their semantic relevance, promoting informative signals while suppressing clutter. Supported by our newly curated radar-text pairs, mmLIP captures structured radar semantics. Extensive experiments demonstrate that mmLIP successfully integrates with diverse LLMs or vision-language models, achieving state-of-the-art performance on zero-shot bidirectional retrieval as well as text generation tasks, including captioning and question answering.
PaperID: 2530, Poster
Abstract: Finite-shot quantum tomography often leaves a posterior over many density matrices compatible with the same measurement record, but fast reconstruction methods collapse this uncertainty to a single state. We frame this setting as amortized posterior sampling and introduce BuresTomFlow, a measurement-conditioned Flow Matching sampler for full-rank density matrices. The method transports base density matrices along Bures-Wasserstein paths on the positive matrix cone, projects the path back to the trace-one state space, and trains the velocity field with a Bures-aligned tangent loss. Thus the sampler is trained in the same fidelity-based geometry used to judge posterior quality. We evaluate BuresTomFlow in dense four-qubit simulations with count, shadow, and mixed measurement records. Against Cholesky, log-density, and generic Riemannian Flow Matching controls at comparable learned-sampler runtime, BuresTomFlow improves Bures distance, credible-ball coverage, and held-out observable prediction. In the main comparison across multiple seeds, mean Bures distance decreases from 0.618 for the Cholesky baseline to 0.367 for BuresTomFlow, and mean coverage error decreases from 0.203 to 0.100. These results show that the value of Bures geometry is not merely enforcing physical states, but producing calibrated posterior uncertainty when finite-record tomography cannot be reduced to one reconstruction.
PaperID: 2531, Poster
Abstract: Flexible molecular docking is crucial for drug discovery; however, existing flow-matching generative models often produce physically invalid conformations (e.g., steric clashes, distorted bond geometry). Our analysis suggests that a primary contributor to this challenge lies at the modeling level: when the flow is parameterized by standard equivariant networks, an intrinsic approximation floor can emerge, potentially hindering the learned velocity field from reaching its optimal convergence bound, particularly for highly flexible molecules. We present UniFlowDock, which addresses this by introducing strictly Complete Equivariant Velocity Fields, thereby eliminating the dimensional incompleteness of standard architectures. This theoretical completeness is further bolstered by Fragment-Adaptive Optimal Coupling, a method that structurally linearizes transport trajectories and decomposes the generation process into physics-aware subproblems. Experiments on PDBBind and PoseBusters demonstrate UniFlowDock achieves state-of-the-art accuracy and physical validity, showing consistent gains in highly flexible docking scenarios where target distributions exhibit higher transport complexity.
PaperID: 2532, Poster
Authors: Kailu Guo, Akram Alomainy, Khalid Z Rajab
Abstract: Radar-based human pose estimation is a promising alternative to camera-based perception, but remains challenging due to sparse observations, noisy measurements, and severe pose ambiguity. Existing methods infer both global skeletal structure and flexible peripheral joints from sparse radar observations, often compromising structural plausibility and fine-grained local accuracy. Motivated by the coarse-to-fine nature of human visual perception, we propose RadarFlowPose, a framework that decomposes radar-based human pose estimation into coarse pose prior generation and flow-based refinement. The coarse model first produces a structurally plausible and temporally coherent pose prior from sparse radar observations, providing a global initialization for subsequent refinement. Instead of estimating the full pose from scratch, the refinement stage focuses on correcting the remaining local ambiguities and motion-induced errors. By incorporating localised Doppler motion cues into the flow-matching process, the model performs motion-aware residual refinement around the coarse prior, leading to more accurate and anatomically consistent pose estimates. Experiments on three benchmark radar pose datasets show consistent effectiveness, especially in motion-sensitive and fine-grained joint-level evaluation.
PaperID: 2533, Poster
Abstract: Random Network Distillation (RND) is a scalable novelty signal for reinforcement learning, but it can overgeneralize: the predictor extrapolates the random target off the data manifold, causing novelty scores to collapse on unfamiliar inputs. This paper analyzes this failure through a kernel-theoretic lens. Under a GP target model and an NTK-regime predictor, we show that the expected RND energy is a cross-kernel residual governed jointly by the target covariance kernel and the predictor interpolation kernel. This identifies overgeneralization as spectral under-excitation of modes where predictor residuals would otherwise survive. We then explicitly formulate spectral target design as a principle for shaping RND residual geometry; one concrete instance is replacing to replace the implicit random-network target with a bandwidth-controlled Random Fourier Feature (RFF) target. Numerical experiments across offline D4RL and online Atari benchmarks show this target-design approach improves novelty discrimination and downstream RL performance.
PaperID: 2534, Poster
Abstract: Missing data is pervasive in real-world datasets, arising from sensor failures, image occlusions, and incomplete reporting. Imputation is a conditional generation problem that must preserve observed entries exactly while producing plausible values for missing ones. Yet most diffusion-based imputers adopt Gaussian noising whose global perturbations are structurally misaligned with this local hard constraint, so consistency is typically enforced through conditional training, clamping, projection, or repainting-style corrections. We address this mismatch with Continuously-Augmented Hybrid Masked Diffusion (CAHMD), a masked diffusion framework whose conditional sampler leaves observed coordinates unchanged by construction. To extend masked diffusion beyond purely discrete data, CAHMD attaches coordinate-wise auxiliary continuous latents to masked entries, providing a denoising channel before values are revealed in continuous and hybrid data. For incomplete training data, we introduce observed-mask gating, an observed-data masked reconstruction principle that uses available entries as pseudo-missing targets and avoids repeated EM-style imputation steps. Across tabular, image, and image--caption benchmarks, CAHMD gives a strong reconstruction--perception--efficiency trade-off while reducing the training overhead of EM-style self-imputation and the inference overhead of repainting-style inner loops.
Authors:
Yifu Yuan, Yaoting Huang, Xianze Yao, Shuoheng Zhang, Linqi Han, Yutong Li, Pengyi Li, Jiangeng Sun, Wenting Jia, Yucheng Hu, YuHao Liu, Ruihao Liao, Qiyu Wu, Yuxiao Li, zhao zhang, Zibin Dong, Fei Ni, YAN ZHENG, Shuyang Gu, Yi Ma, Hongyao Tang, Han Hu, Jianye HaoAbstract: We introduce Embodied-R1.5, a unified Embodied Foundation Model (EFM) that integrates comprehensive embodied reasoning capabilities, spanning embodied cognition, task planning, correction, and pointing, within a single architecture toward general physical intelligence. Leveraging three automated data construction pipelines to significantly expand the data coverage of critical capabilities, we build a large-scale data system of over 15B tokens, and design a multi-task balanced RL recipe to alleviate heterogeneous task conflicts. We further introduce a Planner-Grounder-Corrector (PGC) closed-loop framework that enables a single model to autonomously execute and self-correct over long-horizon tasks. With only 8B parameters, Embodied-R1.5 achieves SOTA on 16 out of 24 embodied VLM benchmarks, surpassing leading models like Gemini-Robotics-ER-1.5 and GPT-5.4. Benefiting from the internalized embodied capabilities, Embodied-R1.5 can be fine-tuned into a VLA with only a small amount of data, outperforming leading VLA models like \pi_0.5 across 4 popular manipulation benchmark suites. We further conduct extensive zero-shot real-robot experiments, validating performance in instruction following, affordance grounding, articulated object manipulation, and long-horizon complex tasks, demonstrating strong generalization to the physical world. We will fully open-source weights, datasets and training code to facilitate future research in EFMs.
PaperID: 2536, Poster
Abstract: Subgraph heterogeneity is a critical challenge in federated learning, significantly impacting model performance. Data-free knowledge distillation overcomes the limitation of sharing private label distributions in federated learning, yet existing research lacks in-depth exploration of subgraph heterogeneity. Our work builds upon the fact that node and structural variations in heterogeneous graphs cause significant disparities in the reliability of knowledge from local graph neural networks. We propose a structure-aware bidirectional data-free federated distillation technique, where the generator and global model engage in an adversarial training process to ensure reliable bidirectional knowledge transfer between local and global models. Extensive experiments across multiple public datasets demonstrate the model's effectiveness.
PaperID: 2537, Poster
Abstract: Large Reasoning Models (LRMs) have made remarkable progress, driven by a paradigm shift from fast thinking to slow thinking. While slow-thinking unlocks superior reasoning ability, its long-form nature imposes a prohibitive computational burden on Reinforcement Learning (RL) for LRMs. In this paper, we propose Fast-RL, a novel framework that accelerates RL training for slow-thinking LRMs. Our acceleration principle stems from a foundational insight: fast-thinking sampling serves as an efficient and powerful proxy for slow-thinking sampling in enhancing reasoning, since the core reasoning skills are shared across thinking modes and can be improved through either trajectory type. Fast-RL consists of two stages. Stage-1 establishes reliable prompt-based thinking-mode control, so that a slow-thinking-oriented model triggers fast-thinking under a specific prompt while preserving its native slow-thinking under the standard prompt. Stage-2 then enhances reasoning ability through RL with exclusively fast-thinking sampling for acceleration. A lightweight LoRA adapter is introduced as a dedicated mode controller: it is trained in Stage-1 and frozen in Stage-2, decoupling mode-control parameters from reasoning parameters and protecting the model's native slow-thinking from being disturbed by fast-thinking RL. After Stage-2, the LoRA acts as a disposable training scaffold that can be either merged into the main model or simply discarded. Experiments on diverse multimodal and text-only LRMs show that Fast-RL accelerates training by about 10× while delivering superior performance gains over vanilla slow-thinking RL. On Qwen3-1.7B, Fast-RL improves AIME24 by +10.0 and +13.7 in slow- and fast-thinking evaluation, respectively. Vanilla slow-thinking RL, in contrast, drops slow-thinking accuracy by -8.8 and yields only +2.8 in fast-thinking evaluation.
PaperID: 2538, Poster
Abstract: Feedforward foundation models have recently shown remarkable 3D reconstruction capabilities. However, existing models exhibit large tracking drift in long-context streaming reconstruction due to error accumulation. In this paper, we revisit loop closure with streaming reconstruction foundation models to enable accurate, drift-free, kilometer-scale reconstruction. Specifically, our method detects loop candidates through global descriptor retrieval, and constructs loop-conditioned windows to estimate the relative poses between looped frames. Given the observation that our adopted streaming reconstruction backbone produces a globally consistent scale, we optimize all frame poses on the SE(3) manifold with sequential and loop closure constraints, avoiding the pose graph optimization on the Sim(3) or higher-dimensional SL(4) manifolds employed in prior works. Extensive experiments show that our method reduces drift and produces consistent geometry on kilometer-scale sequences, significantly outperforming the state of the art. The code will be made public.
PaperID: 2539, Poster
Authors: Abdullah Azeem, Ruisheng Wang, Qingquan Li, Abubakar Siddique
Abstract: Generalizable 3D Gaussian Splatting aims to predict renderable Gaussian scenes from sparse images without per-scene optimization at test time. Existing feed-forward methods improve this problem mainly by strengthening the evidence given to the predictor, such as multiview aggregation, geometric cues, auxiliary supervision, and surface priors. Yet the prediction itself is usually learned as a direct map to the final Gaussian scene. This endpoint-centered objective compresses scene layout, local geometry, visibility, opacity, orientation, and appearance into a single end-point target, leaving the construction path of the Gaussian scene implicit. We propose GRiF, a geometry-aware flow matching framework that augments endpoint target supervision with conditional transport in scene space. Starting from a controlled source scene, GRiF learns a view-conditioned velocity field that progressively evolves a hierarchical Gaussian state toward a renderable target. The hierarchy organizes updates from coarse layout to fine attributes, while geometry-aware paths respect the native structure of Gaussian variables with rotation-aware paths for orientation variables. Experiments on RealEstate10K and ACID, with zero-shot transfer evaluated on ScanNet, DL3DV, and DTU, show improved in-domain reconstruction and stronger cross-domain performance.
PaperID: 2540, Poster
Authors:
Pureum Kim, Younggeon Ryu, Dongyoon Lee, Hae Beom Lee, Kyong H JinAbstract: Implicit neural representations (INRs) capture global low-frequency structure during the early stage of training and then refine localized high-frequency details, afterwards. However, standard optimizers are agnostic to the coarse-to-fine learning behavior of INRs. Such optimizers accumulate a gradient matrix in momentum buffer and retain components misaligned with the dominant low-frequency directions in the early stage. We propose a progressively growing rank scheduler and a rank-aware learning-rate scheduler for enhancing a subspace-based momentum optimization, where a momentum matrix is updated within a low-rank subspace. The rank schedule expands the subspace during training, and the scheduler holds the peak learning rate through rank growth so that the added subspace contributes to fine-detail recovery. Experiments on image fitting, single-image super-resolution (SISR), and neural radiance field optimization show that the proposed method improves convergence and peak reconstruction quality over competitive optimizers. The largest gains appear in the high-frequency refinement stage, mirroring the coarse-to-fine progression that the rank schedule targets. A significant study on hyperparameter optimization (HPO) on NeRF further confirms that these gains persist under matched search budgets.
Abstract: Generative molecular optimization aims to design molecules with properties surpassing those of existing compounds. However, evaluating candidate molecules is expensive, so sample efficiency, i.e., generating optimized molecules within a fixed evaluation budget, is essential. Moreover, in the offline setting, where a surrogate model approximates candidate evaluation, sample inefficiency compounds with surrogate error, further limiting performance. To address these challenges, we introduce Joint Self-Improvement (JSI), a framework for sample-efficient online and offline molecular optimization. JSI combines trained via a joint likelihood-based loss, which aligns molecule generation and surrogate modeling, and , which iteratively biases the generative component of the joint model toward higher-scoring molecules using direct objective evaluations in the online setting and the predictive component in the offline setting. Experiments across docking score optimization, in both online and offline settings, as well as antimicrobial peptide optimization, demonstrate that JSI outperforms state-of-the-art methods under limited evaluation budgets.
Abstract: Audio--visual joint representation learning under Cross-Modal Generalization (CMG) aims to transfer knowledge from a labeled source modality to an unlabeled target modality through a unified discrete representation space. Existing symmetric frameworks often suffer from information allocation ambiguity, where the absence of structural inductive bias leads to semantic--specific leakage across modalities. We propose Asymmetric Hierarchical Anchoring (AHA), which enforces directional information allocation by designating a structured semantic anchor within a shared hierarchy. In our instantiation, we exploit the hierarchical discrete representations induced by audio Residual Vector Quantization (RVQ) to guide video feature distillation into a shared semantic space. To ensure representational purity, we replace fragile mutual information estimators with a GRL-based adversarial decoupler that explicitly suppresses semantic leakage in modality-specific branches, and introduce Local Sliding Alignment (LSA) to encourage fine-grained temporal alignment across modalities. Extensive experiments on AVE and AVVP benchmarks demonstrate that AHA consistently outperforms symmetric baselines in cross-modal transfer. Additional analyses on talking-face disentanglement experiment further validate that the learned representations exhibit improved semantic consistency and disentanglement, indicating the broader applicability of the proposed framework.
Abstract: We introduce a novel framework that transforms the resource-intensive (adversarial) prompt optimization problem into an efficient, amortized inference task. Our core insight is that pretrained, non-autoregressive generative LLMs, such as Diffusion LLMs, which model the joint distribution over prompt-response pairs, can serve as powerful surrogates for prompt search. This approach enables direct conditional generation of prompts, effectively replacing costly discrete optimization with a small number of parallelizable samples. We provide a probabilistic analysis demonstrating that under mild fidelity assumptions, only a few conditional samples are required to recover high-reward (harmful) prompts. Empirically, our method substantially outperforms prior attacks in both success rate and compute cost, producing low-perplexity, diverse jailbreaks that transfer to a wide range of black-box target models, including robustly trained and proprietary LLMs. Beyond adversarial prompting, our framework opens new directions for red teaming, automated prompt optimization, and leveraging emerging Flow- and Diffusion-based LLMs.
PaperID: 2544, Poster
Abstract: Flexible switching between reasoning trajectories (i.e., thoughts switching) has significantly enhanced the reasoning capabilities of Large Reasoning Models (LRMs). However, existing models often switch excessively yet fail to sustain promising reasoning thoughts---a phenomenon termed ''under-thinking''. While recent efforts suppress switching to mitigate this, such over-correction may discard valuable trajectories. To address this challenge, we propose Steady Thought (ST), a novel thought-level preference optimization framework. ST formalizes under-thinking as a preference issue at switching points, forcing single-path continuations from identified points to yield high-quality reasoning. Then, ST performs thought-level preference optimization by treating the newly generated response as preferred and the original one as dis-preferred. Experiments across multiple models and datasets show that ST effectively reduces token consumption while maintaining or even improving accuracy. It reduces output length by up to 63.8% while improving accuracy by up to 11.4%. Further analysis suggests that ST helps models acquire a generalizable ability to reduce redundant switching across domains and languages.
PaperID: 2545, Poster
Abstract: Prefill-Decode (PD) disaggregation improves LLM serving by isolating compute-bound prefilling and memory-bound decoding onto specialized GPU pools. However, dynamically rebalancing these pools is challenging because request arrival rates, prompt lengths, and generation lengths change over time. We formulate this autoscaling problem as latency-constrained cost minimization: the controller must minimize GPU usage while keeping P99 tail latency within TTFT and TPOT targets. Existing threshold-based heuristics require extensive parameter tuning, yet still tend to over-provision. To address this, we propose a reinforcement learning (RL) framework that learns a cost-efficient autoscaling policy. The policy uses a recurrent network to track multi-timescale workload trends and pod lifecycle states, enabling decisions that account for delayed pod readiness and delayed SLO feedback. It also restricts the action space with hard masks and anti-oscillation cooldowns, enforcing feasible scaling actions while reducing wasteful GPU transitions. We train the policy in a high-fidelity PD-disaggregated serving simulator calibrated from profiling data and validated against a physical testbed. Simulator experiments on ShareGPT and Azure Conversational traces show that the learned policy satisfies the same latency constraints as grid-searched heuristic baselines while reducing average GPU occupancy. Real-cluster deployment on ShareGPT further confirms sim-to-real transfer while satisfying the latency targets.
Abstract: Geometry-aware optimizers such as Newton and natural gradient can improve conditioning in deep learning, but scalable variants such as K-FAC, Shampoo, and related preconditioners usually impose structural approximations early, often discarding cross-layer interactions induced by the network computation. We introduce Layerwise LQR (LLQR), a framework for learning structured inverse preconditioners under a global layerwise optimal-control objective. The starting point is an exact equivalence: the steepest-descent step under a broad class of divergence-induced quadratic models—including Newton, Gauss–Newton, Fisher/natural-gradient, and intermediate-layer metrics—can be written as a finite-horizon Linear Quadratic Regulator (LQR) problem. This formulation serves as a reference that exposes the layerwise dynamics and cost matrices encoding the original dense geometry. We then derive a scalable relaxation that learns diagonal, (E-)Kronecker-factored, or other structured inverse preconditioners by minimizing the LQR objective and reusing them across iterations. The resulting optimizer wraps standard methods while retaining a principled connection to second-order geometry, without forming or inverting the global curvature matrix. Experiments on ResNets and Transformers show that LLQR improves optimization dynamics and often translates these gains into improved final test performance, while adding only modest wall-clock overhead. It establishes LLQR as a practical framework for geometry-aware second-order methods and a reference for evaluating scalable approximations.
PaperID: 2547, Poster
Abstract: Long video understanding requires VideoLLMs to answer queries about videos spanning tens of minutes to hours, but such long videos produce far more visual tokens than current VideoLLMs can process. To address this issue, keyframe selectors are proposed to select the most query-relevant frames by computing the similarity between frame and query embeddings. Existing selectors either encode frames independently without temporal context or process chunks in isolation with ranking objectives that are difficult to optimize. To alleviate those limitations, we propose FrameScout with two core designs: (i) a streaming selector architecture that leverages a sliding-window KV cache and successor aggregation to produce temporally-contextualized frame embeddings; (ii) a frame-query contrastive objective that aligns frame and query embeddings in a shared space, directly training the selector to distinguish relevant frames from irrelevant ones. Extensive experiments on three long-video benchmarks and six base VideoLLMs demonstrate that FrameScout consistently achieves state-of-the-art performance.
Authors:
Carles Balsells Rodas, Zhengrui Xiang, Francisco Sumba Toral, Yingzhen LiAbstract: Learning identifiable representations in deep generative models remains a fundamental challenge, particularly for sequential data with regime-switching dynamics. Existing approaches establish identifiability under restrictive assumptions, such as stationarity or limited emission models, and typically rely on variational autoencoder (VAE) estimators, which introduce approximation gaps that limit the recovery of the latent structure. In this work, we address both the theoretical and practical limitations of this setting. First, we establish identifiability of a broad class of recurrent nonlinear switching dynamical systems under flexible assumptions, significantly extending prior results. Second, we introduce \OmegaSDS, a flow-based estimator that enables exact likelihood optimization using expectation-maximisation. Through empirical validation on both synthetic and real-world data, our results demonstrate that \OmegaSDS achieves improved disentanglement compared to VAE-based estimators and more accurate forecasting of underlying dynamics.
PaperID: 2549, Poster
Authors: Paweł Lenartowicz, Hubert Plisiecki
Abstract: Partial Least Squares (PLS) regression extracts a few outcome-aligned directions in a high-dimensional \mathbfX and is widely used across applied science, but inference on the resulting fit is either expensive (CV-permutation-Q^2), biased and discouraged (jackknife t), or absent sklearn ships no test; the corresponding statsmodels feature request has been open since 2019). We close this gap by reducing inference to held-out OLS refits of the supervised subspace, a primitive shared by PLS, supervised PCA, linear probes, and sparse-autoencoder concept directions. We supply two tests on it: a Nadeau-Bengio corrected asymptotic t-test (NB-asymptotic) as the default, and a permutation-referenced variant (NB-permutation) for near-singular designs where the Fisher-z asymptotic drifts. A rotation-invariance result shows the held-out predictions, and hence the test statistic, are unchanged under any orthogonal rebasing of the supervised span, so the same test applies to varimax-rotated word-readable axes. We validate on PLS (synthetic geometries, Tecator NIR chemometrics, cross-lingual GloVe valence prediction in EN/PL/ES at K=3, r^2\approx0.63/0.69/0.55, all p<.001); transfer to the other OLS-refit pipelines is argued but not tested here. NB needs ~20% fewer samples than CV-permutation-Q^2 for matched power and is 50-100 × cheaper at common, already compromised, defaults. We release a Rust library with Python, R, and Julia bindings, plus a Python text pipeline.
PaperID: 2550, Poster
Abstract: A central goal of safety research is determining whether a model is misaligned. Prior work has largely focused on detecting concerning behavior, but behavior alone is not sufficient to establish misalignment: a concerning action can arise from benign causes such as confusion. This raises the problem of determining whether malign intent underlies such behavior, a process we term model incrimination. The goal of this paper is to develop effective methods for doing so. To enable this, we create a suite of six agentic environments where models exhibit concerning behavior as practice grounds, and follow a simple two-step protocol for investigating the causes behind model behavior: hypothesis generation via reading the chain of thought -- which, while not always faithful, is a rich source of hypotheses about what drives model behavior -- followed by hypothesis validation via environment interventions and additional methods as appropriate. As we do not have access to ground truth about why a model takes an action, we rely on convergent findings across independent experiments as our standard of evidence. By following our protocol, we learn effective methods for concretely determining motivations: for example, we use predictions to back out latent properties of model behavior, like the fact that Kimi K2 Thinking takes shortcuts due to a legitimate disposition towards less effortful courses of action, and that reward hacking in frontier models is strategic, while in weaker models it is not. However, some unanswered questions require further methodological development: we construct an absence-of-evidence case against Kimi K2 Thinking believing it is going against the user's wishes while taking shortcuts, but run into confounds that limit our confidence in its absence. Overall, a key takeaway is that simple methods like reading the CoT and environment interventions are highly effective. More broadly, our work shows model incrimination is a tractable empirical problem with significant room for progress, and establishes a baseline for future work.
PaperID: 2551, Poster
Abstract: Multi-image composition (MICo) is a critical yet challenging capability for large-scale image editing models, where consistency degrades rapidly as the number of reference images increases. While supervised fine-tuning (SFT) can equip models with basic MICo ability, it provides only implicit supervision on the composed image and often fails to learn reliable object-level correspondence across multiple references. We propose DeCoRL, a reinforcement-learning (RL) framework that improves MICo consistency by decomposing holistic MICo consistency into object-level consistency. DeCoRL first establishes instruction-relevant object correspondences across reference images and localizes the corresponding regions in the generated image. It then uses a VLM-based reward model trained on single-object consistency annotations to produce fine-grained, object-wise feedback, which is aggregated into a multi-object consistency reward to guide RL post-training. We further introduce DeCo-Bench, a benchmark covering compositions from 2 to 6 reference images. Experiments on DeCo-Bench and the public MICo-Bench demonstrate consistent improvements in MICo consistency.
PaperID: 2552, Poster
Abstract: RNA molecules often function through multiple experimentally observed conformational states. A central challenge is to capture ensemble diversity without sacrificing structural fidelity: samples may collapse toward averaged representative structures or drift into unsupported conformational regions. We introduce \textttREFLEX (\underlineRNA \underlineEnsemble generation calibrated by \underlineFLEXibility), a sequence-conditioned generator for RNA backbone-frame ensembles based on a heteroscedastic stochastic bridge on \mathrmSE(3). During training, experimental \mathrmB-factors calibrate residue-wise bridge widths, exposing each observed conformer through local perturbation neighborhoods with residue-specific scales while keeping constrained regions sharper. A learned flexibility condition guides the bridge-field network to recover target conformer-specific frames from these calibrated noisy states. At inference, \textttREFLEX uses only the input sequence, requiring no MSA-subsampling heuristics. On held-out multi-conformer RNA clusters, \textttREFLEX achieves the best precision--recall trade-off for ensemble coverage across evaluation thresholds, while ablations indicate that the learned bridge field recovers conformer-specific basins more reliably.
Abstract: Diffusion models have achieved remarkable success as generative models. However, even a well-trained model can accumulate errors throughout the generation process. These errors become particularly problematic when arbitrary guidance is applied to steer samples toward desired properties, which often breaks sample fidelity. In this paper, we propose a general solution to address the off-manifold phenomenon observed in diffusion models. Our approach leverages a time predictor to estimate deviations from the desired data manifold at each timestep, identifying that a larger time gap is associated with reduced generation quality. We then design a novel guidance mechanism, `Temporal Alignment Guidance' (TAG), attracting the samples back to the desired manifold at every timestep during generation. Through extensive experiments, we demonstrate that TAG consistently produces samples closely aligned with the desired manifold at each timestep, leading to significant improvements in generation quality across various downstream tasks.
Abstract: We introduce TriSearch, a reinforcement-learning framework for optimizing objectives over triangulations of a polytope via bistellar flips. The key idea is a circuit-supported subtriangulation action representation: feasible flips are encoded by their supporting circuit and realized local subtriangulation, enabling a learned policy to rank them using local geometric and combinatorial features. This yields a dimension-agnostic interface and enables efficient traversal of the flip graph without explicit enumeration of the full triangulation space. Instantiated in 3D and 4D, TriSearch generalizes zero-shot from small training instances to larger polytopes with exponentially larger search spaces. It achieves top performance on metric objectives in 3D and, in 4D, discovers more distinct Fine, Regular, Star triangulations of reflexive polytopes, corresponding to Calabi-Yau threefolds, than existing samplers under a fixed budget.
PaperID: 2555, Poster
Authors:
Kunhang Lv, Rui Han, Fuqi Jia, Yuhang Dong, Feifei Ma, Jian ZhangAbstract: Large language models can often identify useful semantic structure in formal artifacts, but their outputs cannot be trusted as proofs or solver decisions. We study whether LLMs can instead guide formal reasoning systems through a narrow, solver-verifiable interface. We instantiate this idea in quantified array SMT solving, where modern instantiation-based techniques can fail when key ground terms and equalities are hidden by auxiliary variables, axiom guards, and nested array terms. We observe that many instances are largely driven by a small set of focal ground terms or constraints, similar to the backdoor sets in SAT solving. Building on this, we present LinguaArray, a framework where an LLM proposes a small Semantic Focus Set (ground terms/constraints), while a conventional SMT solver remains the sole proof engine and certifies all results. LinguaArray uses these LLM-suggested, solver-validated finite sets to significantly enhance solver capability, solving additional satisfiable and unsatisfiable instances that the backend alone fails to solve. It interacts within a finite time and then falls back to the backend solver to mitigate potential regressions. Evaluation on 1015 SMT-LIB instances across five logics under 1200s timeout (including LLM latency) shows that LinguaArray solves up to 69.1% of instances and increases solved-instance counts by 105%--397% over the corresponding backends.
Abstract: Electroencephalogram (EEG) classification plays a key role in medical diagnosis and brain–computer interfaces, but remains challenging due to low signal-to-noise ratios and high inter-subject variability. As a result, many existing approaches rely on subject-specific models, which fail to exploit shared structure in neural signals and do not generalize to unseen subjects. To address these limitations, we propose LAtte, a framework that combines Lorentz attention with a hyperbolic InceptionTime-based encoder to improve cross-subject generalization in EEG classification. The model explicitly decomposes EEG signals into a learned baseline component and task-relevant deviations, enabling more structured representation learning. To further improve robustness and adaptability, we incorporate subject-specific low-rank adaptation (LoRA) modules at both encoder and decoder levels, augmented with a Lorentz boost–based LoRA mechanism and hyperbolic projection layers to reduce overfitting in geometric representations. We evaluate LAtte with and without finetuning in three settings: subject-specific, subject-conditional, and leave-one-subject-out (LOSO) on five established EEG datasets, achieving a consistent improvement in performance over current state-of-the-art methods for smaller datasets and maintaining performance for larger datasets.
PaperID: 2557, Poster
Authors:
Qingtao Yu, Zhen Wang, Ning Lai, Boyang Hu, Hongdong Li, Dylan CampbellAbstract: We present an inference-time approach for enhancing the visual diversity of image sets generated by text-to-image (T2I) models, while preserving, or even improving, the average image quality. Our method leverages the intrinsic geometry of hidden latent representations in pretrained models to diversify a batch of images generated under the same prompt. First, we introduce a token repulsion mechanism to encourage diversity by increasing the angle between the tokens of hidden features across different samples. This is achieved via a negative Riemannian gradient step defined over a distribution on the unit hypersphere. Second, to preserve the quality and prompt adherence of the output images, we encourage the updates to remain within the high-density region learned by the generative model. Third, we extend this constraint to the output of each intermediate layer, enabling more fine-grained iterative guidance. We evaluate our method on three representative large-scale pretrained T2I models: Flux1.dev (flow matching), SDXL (diffusion), and Flux2.Klein (few-step distilled flow matching). Experimental results demonstrate substantial improvements in visual diversity under identical prompts, while maintaining averaged individual sample fidelity. Our approach is simple, effective, and broadly applicable. It requires no external models to guide the generation process and introduces only modest additional computation overhead during the early stages of sampling.
PaperID: 2558, Poster
Authors: Alan Moore, Zhengyuan Zhu, Lynna Chu
Abstract: Online change-point detection is a fundamental problem in sequential monitoring, where the goal is to detect distributional changes in a time series as quickly as possible while controlling false alarms. We propose a general framework for online change-point detection that leverages foundation probabilistic forecasters, such as Chronos-2, which can capture complex temporal structure including trend, seasonality, heteroskedasticity, and nonlinear dependence. Our approach transforms each incoming observation relative to its predictive quantile forecast into an approximately standard Gaussian monitoring statistic, which can then be used in classical sequential detection procedures such as CUSUM and MOSUM. To reduce adaptation of the forecaster to post-change observations, we introduce a buffer between the forecasting context window and the forecast target. Because this design can induce serial dependence in the monitoring statistics, we develop a parametric bootstrap procedure to calibrate critical thresholds and control false alarms in practice. Extensive simulations across data-generating processes involving non-normality, trend, seasonality, heteroskedasticity, and autoregressive dependence show that the proposed method achieves competitive detection performance relative to benchmark procedures, while requiring no explicit specification of a parametric pre-change model.
PaperID: 2559, Poster
Abstract: Self-Supervised Learning (SSL) has achieved impressive success in learning semantic visual representations, yet the underlying principles driving this success remain underexplored. In this work, we hypothesize that common SSL pretext tasks implicitly model object co-occurrence statistics, a fundamental cue for visual learning. To test this, we curate three datasets of segmented objects from existing vision benchmarks using a state-of-the-art segmentation model. Through experiments across many SSL models, we reveal a hierarchical encoding of semantic information: while the visual backbone captures object categories, the projection head specializes in encoding co-occurrences between object categories. As a result, the projection head outperforms the backbone in aligning with human judgments of inter-category similarity. Furthermore, by controlling co-occurrence patterns during pre-training, we demonstrate that encoding object co-occurrences can significantly accelerate the emergence of category-level representations. Our findings uncover a previously hidden learning principle in SSL and suggest a path toward designing more effective pretext tasks by explicitly leveraging object co-occurrence structure.
PaperID: 2560, Poster
Abstract: Learning a risk-sensitive individualized treatment rule (ITR) with patient covariates is challenging when the treatment criterion depends on the full conditional outcome distribution rather than only its conditional mean. Existing methods are largely criterion-specific, making it difficult to compare multiple treatment criteria within a common pipeline, such as quantiles, conditional value-at-risk (CVaR), and complex preference-based objectives such as cumulative prospect theory (CPT). We propose a distribution-first framework for risk-sensitive ITR learning that first estimates treatment-specific conditional outcome distributions and then constructs the target ITR by plug-in criterion evaluation. The proposed framework unifies treatment rule learning across multiple criteria and accommodates hard-to-optimize objectives within the same estimation pipeline. On the theoretical side, it yields a general error propagation result for Lipschitz treatment criteria, linking first-stage distribution estimation error to downstream criterion estimation and treatment regret, together with sharper regret bounds under standard margin conditions. Empirically, we show that the proposed framework outperforms direct criterion-specific methods in multiple settings and further evaluate it on the ACTG175 HIV clinical trial, a randomized study comparing several antiretroviral regimens in HIV-infected adults.
Abstract: We introduce Pion, a spectrum-preserving optimizer for large language model (LLM) training based on orthogonal equivalence transformation. Unlike additive optimizers such as Adam and Muon, Pion updates each weight matrix through left and right orthogonal transformations, preserving its singular values throughout training. This yields an optimization mechanism that modulates the geometry of weight matrices while keeping their spectral norm fixed. We derive the Pion update rule, systematically examine its design choices, and analyze its convergence behavior along with several key properties. Empirical results show that Pion offers a stable and competitive alternative to standard optimizers for both LLM pretraining and finetuning.
Abstract: We introduce Equilibrium Matching (EqM), a generative modeling framework built from an equilibrium dynamics perspective. EqM discards the non-equilibrium, time-conditional dynamics in traditional diffusion and flow-based generative models and instead learns the equilibrium gradient of an implicit energy landscape. Through this approach, we can adopt an optimization-based sampling process at inference time, where samples are obtained by gradient descent on the learned landscape with adjustable step sizes, adaptive optimizers, and adaptive compute. EqM surpasses diffusion/flow models empirically, achieving an FID of 1.90 on ImageNet 256×256. EqM is also theoretically justified to learn and sample from the data manifold. Beyond generation, EqM is a flexible framework that naturally handles tasks including partially noised image denoising, OOD detection, and image composition. By replacing time-conditional velocities with a unified equilibrium landscape, EqM offers a tighter bridge between flow and energy-based models and a simple route to optimization-driven inference.
PaperID: 2563, Poster
Abstract: Looped-in-depth architectures provide an inductive bias toward learning step-by-step procedures for tasks that require compositional reasoning. The effective depth reached by looping determines the quality of the solution. Similar to deep architectures, looped architectures are prone to signal propagation issues as the halting decision is postponed. In this paper, we address these signal propagation issues by using pre-norm layers and residual scaling. Furthermore, we propose : a Fixed-Point Reasoning Model that uses fixed-point convergence as an end-to-end halting mechanism in a looped architecture. We show that fixed-point halting allows FPRM to adapt its compute to the difficulty of the task. FPRM proves effective on common reasoning benchmarks, namely sudoku, maze, and state tracking.
PaperID: 2564, Poster
Abstract: Operator-valued concentration inequalities are foundational to the analysis of modern high-dimensional statistics and randomized algorithms. However, standard oracle bounds are frequently limited in practice: they require explicit a priori knowledge of the true variance, and often explicitly scale with the ambient dimension, rendering them vacuous for infinite-dimensional or heavily structured operators. Motivated by these challenges, we establish the first empirical Bennett and Bernstein inequalities for sums of independent, bounded, compact self-adjoint operators. Our fully data-driven bounds replace the unknown variance with an empirical estimate and rely strictly on the intrinsic dimension rather than the ambient dimension. This structural shift yields computable, dimension-free guarantees that are strictly sharper for non-isotropic random matrices and seamlessly extend to infinite-dimensional Hilbert spaces. We demonstrate that our empirical bounds achieve asymptotic sharpness with the best known oracle rates. Finally, as an independent byproduct, we derive novel empirical concentration guarantees for the intrinsic dimension itself.
PaperID: 2565, Poster
Abstract: Large language models (LLMs) are increasingly used in binary decompilation to refine the C-like pseudocode produced by traditional rule-based decompilers. While this pseudocode is useful, it is a heuristic and lossy abstraction rather than a faithful copy of the source code. It often contains decompiler errors, especially for aggressively optimized binaries where critical low-level details are obscured. We present ASAP, an assembly-source aligned pseudocode refinement framework for binary decompilation. ASAP learns source-aligned assembly representations from paired source and binary functions using joint function-level and snippet-level contrastive alignment. A Q-Former then compresses chunk-level assembly features into a fixed number of assembly tokens that condition the decompilation LLM alongside the decompiler-produced pseudocode. During refinement, we use stochastic pseudocode masking and a relative assembly-advantage loss to reduce the model's tendency to ignore assembly features and rely only on pseudocode refining. On two decompilation benchmarks across multiple compiler optimization levels, ASAP improves the average re-execution rate from 64.7% to 71.9% and the average recompilation rate from 91.6% to 96.6% compared with the strongest baseline, offering both a new perspective and a practical solution to binary decompilation.
PaperID: 2566, Poster
Authors: Ruogu Chen, Jie Han
Abstract: Euclidean combinatorial optimization problems (ECOPs), such as the Traveling Salesman Problem (TSP) and Capacitated Vehicle Routing Problem (CVRP), possess inherent symmetries under the two-dimensional Euclidean group E(2), including rotations, reflections, and translations. Existing learning-based methods, including recent diffusion-based methods, rely on data augmentation or regularization to approximate E(2)-equivariance. This paper presents EDISCO, the first discrete diffusion model for ECOPs with exact E(2)-invariant generative distributions over node-index solutions. EDISCO introduces an E(2)-equivariant edge-score network coupled with a categorical continuous-time Markov chain over discrete edge variables, and exact posterior sampling provides efficient multi-step inference. This design gives EDISCO a local geometric inductive bias: edge neighborhoods with the same relative geometry and combinatorial context are represented consistently regardless of absolute position or orientation, making learning more efficient and inference more robust than non-equivariant methods. EDISCO outperforms previous learning-based state-of-the-art solvers on synthetic TSP from 100 to 10000 nodes and CVRP from 50 to 2000 customers, while using only 33--50% of the training instances. Trained only on uniform synthetic data, EDISCO also outperforms competing learning-based baselines under spatial distribution shift and CVRP constraint-tightness shift. Code is available at https://anonymous.4open.science/r/EDISCO-7F54.
Authors:
Chi Nguyen Tran, Dao S Minh, Trung-Kiet Huynh, Phu-Hoa Pham, Phu-Quy Nguyen-Lam, Long Tran-ThanhAbstract: Low-altitude economy has been experiencing rapid growth in recent years, with significant contributions to the global economy. While common drone tasks such as delivery, inspection, and search-and-rescue typically use Global Navigation Satellite Systems (GNSS) to navigate, there is an increasing need for developing alternative solutions as GNSS signals can be easily jammed, spoofed, or unavailable over a prolonged operational time. As such, cross-view geo-localization (CVGL), which matches an oblique drone view to a geo-referenced satellite tile, has emerged as a potent alternative that lets an autonomous drone localize itself when GNSS fails. Despite strong recent progress, three limitations persist in current CVGL methods: 1) global-descriptor designs compress the patch grid into a single vector without separating what is shared across the view gap (layout) from what is not (texture); 2) altitude-related scale variation is implicitly retained in the learned embedding rather than treated as a nuisance to be marginalized out; and 3) multi-objective training relies on hand-tuned scalars over losses that live on incompatible gradient scales. To address these limitations, we propose SkyPart, a lightweight swappable head for patch-based vision transformers (ViTs) that institutes explicit part grouping over the patch grid. SkyPart has four components grounded in established theory: (i) learnable prototypes that compete for patch tokens via a single-pass cosine assignment; (ii) altitude-conditioned linear modulation applied only during training so that the retrieval embedding is altitude-free at inference; (iii) a graph-attention readout over active prototypes; and (iv) a Kendall uncertainty-weighted multi-objective loss whose stationary points are Pareto-stationary. At 26.95M parameters and 22.14 GFLOPs, SkyPart is the smallest among the top-performing methods in our comparison and sets a new state of the art on SUES-200, University-1652, and DenseUAV datasets under a single-pass, no-re-ranking, no test-time augmentation (TTA) protocol. Furthermore, its accuracy gap to the strongest baseline widens under the ten-condition WeatherPrompt corruption benchmark
Abstract: Image-to-video generation powers a vast array of downstream applications. However, while the de facto standard, ie, latent diffusion models, typically employ heavily conditioned denoising networks, their decoders often remain unconditional. We observe that this architectural asymmetry leads to significant loss of detail and inconsistency relative to the input image. To address this, we argue that the decoder requires equal conditioning to preserve structural integrity. We focus specifically on image-based generation tasks, where maintaining the structural and textural integrity of the source image is paramount. We introduce RefDecoder, a reference-conditioned video VAE decoder by injecting the reference image directly into the decoding process via reference attention. Specifically, a lightweight image encoder maps the reference frame into the detail-rich high-dimensional tokens, which are co-processed with the denoised video latent tokens at each decoder up-sampling stage. We demonstrate consistent improvements across several distinct decoder backbones (eg, Wan 2.1 and VideoVAE+) achieving consistent improvement on three reconstruction benchmarks. Notably, RefDecoder can be directly swapped into existing video generation systems without additional fine-tuning, and we report across-the-board improvements in subject consistency, background consistency, and overall quality scores on the VBench I2V benchmark. Furthermore, RefDecoder generalizes well to other visual generation tasks such as style transfer.
PaperID: 2569, Poster
Abstract: Mechanics-inspired neural dynamical models improve interpretability and physical consistency, but they usually assume a known configuration space, fixed physical parameters, and deterministic prediction from a prescribed initial condition. We introduce Structured Phase Space GAN (SPS-GAN), a conditional generative model that learns structured dynamics from both Cartesian trajectories and video. SPS-GAN combines a port-Hamiltonian backbone with adversarial conditional generation, allowing conservative, dissipative, and forced systems to be represented in a single architecture. To infer latent mechanical structure from raw observations, we use a cyclic-coordinate loss that recovers both the system degrees of freedom and canonical coordinates without access to the true phase space. Experiments on simulated and real-world systems show that SPS-GAN improves trajectory accuracy over supervised dynamics baselines and video consistency over generative baselines, while consistently recovering the correct latent mechanical coordinates across observation modalities and dynamical regimes.
Abstract: Continual learning aims to acquire tasks sequentially without catastrophic forgetting, yet standard strategies face a core tradeoff: regularization-based methods (e.g., EWC) can overconstrain updates when task optima are weakly overlapping, while replay-based methods can retain performance but drift due to imperfect replay. We study a hybrid perspective: \emphtrust region continual learning that combines generative replay with a Fisher-metric trust region constraint. We show that, under local approximations, the resulting update admits a MAML-style interpretation with a single implicit inner step: replay supplies an old-task gradient signal (query-like), while the Fisher-weighted penalty provides an efficient offline curvature shaping (support-like). This yields an emergent meta-learning property in continual learning: the model becomes an initialization that rapidly \emphre-converges to prior task optima after each task transition, without explicitly optimizing a bilevel objective. Empirically, on task-incremental diffusion image generation and continual diffusion-policy control, trust region continual learning achieves the best final performance and retention, and consistently recovers early-task performance faster than EWC, replay, and continual meta-learning baselines.
PaperID: 2571, Poster
Authors: Almog Anschel, Sarah Keren
Abstract: Goal recognition is the task of inferring an agent's intended goal from its observed behavior. Existing work has largely focused on high-level symbolic recognition, abstracting away the continuous motion through which agents reach their goal. Methods that do reason over real-world trajectories, rely on an explicit dynamics model that is hard to acquire and rarely transfers across settings. Neither approach explicitly learns how agents move toward each candidate goal, limiting accuracy on ambiguous trajectory prefixes where such knowledge is most needed. We present WoRD (Waypoint Recognition with Diffusion), which uses a learned trajectory prior to infer the agent's goal from partial observations. At each step, WoRD runs one warm-started denoising chain per candidate goal, steers each chain toward the observed trajectory through a closed-form log-likelihood gradient, and reweights the resulting samples by Bayes' rule into a calibrated posterior over goals. We evaluate WoRD on simulated robot settings, pedestrian trajectories, and vehicle trajectories using four metrics targeting distinct failure modes. WoRD is the only method that improves all four metrics jointly, including the safety-critical confident-and-wrong rate. WoRD runs under a single inference procedure and preserves calibrated multimodal. beliefs in complex, ambiguous environments where filter- and classifier-based observers collapse to a single mode or grow overconfident.
PaperID: 2572, Poster
Abstract: Mixture-of-Experts (MoE) has become the dominant architecture for frontier language models. To meet this demand, production frameworks have built optimized MoE training stacks over years of engineering effort. Yet evolving these stacks for new architectures and system optimizations remains expensive. With the rise of AI coding agents, they could automate parts of training-framework development and accelerate this evolution. But applying them to these existing frameworks carries hidden costs, invisible to today's throughput-only evaluations. We name this missing dimension agent-task efficiency (ATE): the cost of using coding agents to understand, operate, and extend a framework. Grounded in four agent-native design principles, we build PithTrain, a compact, agent-native MoE training framework. We further introduce ATE-Bench, covering real-world training-framework tasks. Our evaluation shows PithTrain matches the throughput of production frameworks, and on ATE-Bench, PithTrain enables higher agent-task efficiency, with up to 62% fewer Agent Turns and 64% less Active GPU Time.
PaperID: 2573, Poster
Authors: Peter Cho-Ho Lam, Ziyi Wang, Zirui Zhou
Abstract: Prefix-tuning is a parameter-efficient fine-tuning method for large language models (LLMs) that steers model outputs by prepending task-specific prefix representations to the input sequence. Despite its practical success, the theoretical understanding of the expressiveness of prefixes remains incomplete. In particular, the following question remains open: \it Given a pretrained autoregressive Transformer such as the GPT series, for any input-output sequence pair (\mathttx,\mathtty), does there always exist a prefix \mathcalS such that prepending \mathcalS to \mathttx induces the model to generate the target output \mathtty? In this work, we prove that under the assumption that \mathtty is in the \it range of the pretrained Transformer, such a prefix \mathcalS exists and can be explicitly computed by solving multiple linear systems. We further show that this assumption is necessary, in the sense that such a prefix may not exist when the assumption fails, thereby providing a complete answer to the above question. Our analysis is constructive and the obtained theoretical results hold under fairly general assumptions on model architecture. As a notable example, the whole GPT-2 architecture satisfies our model assumptions.
PaperID: 2574, Poster
Abstract: Adapting pretrained Vision-Language-Action (VLA) policies to new deployment environments typically requires collecting expert demonstrations in each target domain, which is costly and difficult to scale. In this paper, we present FreeAct, a framework for demonstration-free robot adaptation that converts generated videos into physically grounded supervision. However, a key challenge is that generated manipulation videos, while visually plausible, often contain embodiment artifacts and kinematic inconsistencies that make direct inverse-dynamics labeling unreliable. FreeAct addresses this challenge by learning a discrete latent action space from large-scale multi-lab robot videos and aligning it with camera-frame end-effector SE(3) motion and gripper-state changes. The resulting latent action tokenizer captures task-relevant motion while reducing reliance on embodiment-specific visual artifacts, enabling generated target-domain videos to be pseudo-labeled with reliable latent action tokens. These tokens are first injected into VLA pretraining and then used, together with source-domain action supervision, to adapt the policy to unseen environments. Empirically, FreeAct improves success from 2% to 57% under target-domain shift using generated videos alone, and further to 74% through MLLM-guided self-evolution on the policy's own rollouts, closing the adaptation loop without expert demonstrations. Further analysis shows that the learned latent action space forms a kinematically structured and transferable representation, establishing generated videos as a scalable source of supervision for robot manipulation.
Abstract: Clinical EEG interpretation requires reasoning over full EEG sessions and integrating signal patterns with clinical context. Existing EEG foundation models are largely designed for short-window decoding and do not incorporate clinical context. We introduce CLEF, a clinically grounded long-context EEG foundation model. CLEF represents EEG sessions as 3D multitaper spectrogram tokens, enabling tractable Transformer modeling at session scale, and aligns embeddings with neurologist reports and structured EHR data through contrastive objectives. We evaluate CLEF on a new 234-task benchmark spanning disease phenotypes, medication exposures, and EEG findings, with more than 260k EEG sessions from over 108k patients. CLEF outperforms prior EEG foundation models on 229 of 234 tasks, improving mean AUROC from 0.65 to 0.74. Reconstruction-only pretraining surpasses prior EEG foundation models, while report and EHR alignment yields further gains. Held-out concept and external-cohort experiments suggest that these representations transfer beyond observed alignment targets. These results support session-scale, clinically grounded representation learning as a promising foundation-model paradigm for clinical EEG.
Abstract: Counterfactual explanations (CFs) provide human-interpretable insights into model's predictions by identifying minimal changes to input features that would alter the model's output. However, existing methods struggle to generate multiple high-quality explanations that (1) affect only a small portion of the features, (2) can be applied to tabular data with heterogeneous features, and (3) are consistent with the user-defined constraints. We proposeCounterFlowNet, a generative approach that formulates CF generation as sequential feature modification using conditional Generative Flow Networks (GFlowNet). CounterFlowNet is trained to sample CFs proportionally to a user-specified reward function that can encode key CF desiderata: validity, sparsity, proximity and plausibility, encouraging high-quality explanations. The sequential formulation yields highly sparse edits, while a unified action space seamlessly supports continuous and categorical features. Moreover, actionability constraints, such as immutability and monotonicity of features, can be enforced at inference time via action masking, without retraining. Experiments on eight datasets under two evaluation protocols demonstrate that CounterFlowNet achieves superior trade-offs between validity, sparsity, plausibility, and diversity with full satisfaction of the given constraints.
PaperID: 2577, Poster
Abstract: Fair classification faces a twofold challenge: maintaining high classification accuracy while limiting reliance on sensitive information. Our method FairSplit achieves this by learning a decomposition of the embedding space into a Fair Space used for prediction and a Sensitive Space that isolates sensitive information. During training, the model must simultaneously learn the classification task and separate fair from sensitive information. The Sensitive Space gives sensitive attributes a destination, so the Fair Space can be used for prediction without relying on them. This decomposition is guided via a learnable mask and the Hilbert-Schmidt Independence Criterion (HSIC) as a measure of statistical dependence. We provide theoretical guarantees showing that the demographic parity gap is upper-bounded by both (i) the learned dimensionality of the Fair Space and (ii) the HSIC between the Fair Space and sensitive attributes. FairSplit effectively handles multiclass tasks and non-binary sensitive attributes, overcoming key limitations of existing fair classification methods. Experiments on established fairness benchmarks show that FairSplit achieves competitive accuracy-fairness trade-offs.
PaperID: 2578, Poster
Authors:
Yingjie Ma, Haonan Wang, Xun Lin, Hui Ma, Ruixin Zhang, Jingyun Zhang, Jun Wang, Rizen Guo, Shouhong Ding, Weicheng Xie, Linlin Shen, Zitong YUAbstract: Multimodal face anti-spoofing (FAS) improves spoof detection by combining complementary RGB, infrared, and depth cues, yet its generalization to unseen domains remains fragile. We revisit multimodal domain-generalized FAS from the perspective of representation freedom. In a controlled CLIP-based post-fusion baseline, increasing visual encoder depth does not monotonically improve cross-domain performance, suggesting that higher-capacity representations may introduce additional degrees of freedom that are not necessarily aligned with real-vs-spoof semantics. To inspect this effect, we introduce the Semantic Dominance Ratio (SDR), a diagnostic statistic that measures the relative proportion of feature variation along the text-defined real-vs-spoof semantic axis versus off-axis residual directions. Motivated by this perspective, we propose Semantic Freedom Bottleneck (SFB), which reparameterizes multimodal features around CLIP text geometry. SFB decomposes each representation into a semantic component along the real-vs-spoof axis and a bounded low-rank residual component for modality-specific evidence. The residual basis is regularized to be orthogonal to the semantic axis, encouraging the residual space to retain complementary modality-specific cues without duplicating the task-semantic direction. This yields a conservative multimodal constraint that preserves discriminative real/spoof semantics while limiting unconstrained non-semantic variation. Experiments on WMCA, CASIA-CeFA, PADISI, and CASIA-SURF under fixed-modality, missing-modality, and limited-source protocols demonstrate consistently strong cross-domain performance.
Abstract: Large audio language models (LALMs) have shown impressive capabilities on diverse audio understanding tasks, ranging from speech transcription to music analysis. However, because LALMs are typically trained to produce text-aligned responses, their hidden states are progressively shaped for text generation rather than for preserving acoustic information. As a result, the diverse acoustic content that audio carries, such as phonetic detail, prosody, sound events, affect, and pitch, is lost along the way and difficult to leverage in the response. We introduce Continuous Audio Thinking (CoAT), a framework that equips audio language models with a continuous latent workspace for organizing acoustic information prior to response generation, grounded by distillation from audio experts. Within the thinking space, the model can utilize the rich acoustic information provided by expert distillation when generating its response. Furthermore, the proposed continuous thinking block can be processed in a single prefill, so CoAT does not require additional autoregressive decoding cost over the baseline. Across three LALMs, Qwen2-Audio, Qwen2.5-Omni-7B, and Audio Flamingo~3, performance gains on a broad benchmark suite spanning audio reasoning, audio understanding, music classification, speech emotion, and speech transcription demonstrate the effectiveness of CoAT. Further analysis confirms that the auxiliary supervision propagates from the thinking positions to the model's textual responses.
PaperID: 2580, Poster
Abstract: Conditional image generators learn P(Y\mid X) for images given clinical, demographic, or experimental covariates, but their covariate effects are encoded implicitly in samples, denoisers, gradients, or score fields. We propose \emphgenerative effect distillation (GED), a framework that treats a conditional image generator as a teacher and projects its conditional law onto an interpretable image-response regression student m_\phi(x,s)=\beta_0(s)+x^\top\beta(s) whose parameters are spatial effect maps. The teacher-to-map target is defined as a statistical functional of the teacher conditional law under a target design and discrepancy, rather than a set of synthetic images or a compressed generator. We develop mean, effect-gradient, and structured score distillation objectives; establish projection identities linking the distilled map to covariance-weighted teacher means and, under Stein conditions, to average teacher gradients; and derive finite-query risk bounds that separate teacher bias, spatial approximation, query design, Monte Carlo, and student optimization error. For the associated finite-basis coefficient-matrix oracle, a Gaussian lower bound matches the query-design and Monte Carlo stochastic dependencies. Across eleven blocks spanning synthetic, semi-real, pretrained SDXL-Turbo/FFHQ-256, CelebA conditional-diffusion, and face-image analyses, GED attains the lowest standardized integrated squared error (SISE) among the direct-regression and augmentation baselines in every block, with paired Student t-tests rejecting equality with the strongest non-GED baseline at p<0.05; on a CelebA conditional residual diffusion teacher, score-access GED via a single-step Tweedie denoising procedure attains SISE 0.009, a factor of 25 below the strongest non-GED baseline.
Abstract: State-of-the-art medium-range AI weather models rival traditional Numerical Weather Prediction (NWP) but require massive training budgets. This restricts access for under-resourced groups and severely limits fast model iteration. We introduce Otter, a highly efficient spatiotemporal forecasting model designed to democratise high-performance weather prediction with AI. Otter is evaluated on ERA5 reanalysis data using the standard WeatherBench protocols where it significantly advances the skill-compute Pareto frontier. The deterministic version outperforms the best NWP baseline by 9.6% at a 24-hour lead time while requiring fewer than 3.5 A100-days for training. It provides a 2x efficiency gain over lightweight AI models and a 100-fold reduction in compute compared to resource-intensive frontier architectures.
PaperID: 2582, Poster
Authors: El M Mansouri
Abstract: Modern models can have high average accuracy and benign loss tails while remaining fragile on rare subpopulations. We identify an information-geometric mechanism: after nuisance projection, the task direction can become locally nonidentifiable even when loss and raw Fisher look well behaved. Fisher-Glass certifies this failure by applying CVaR to inverse nuisance-projected task information. We prove a loss--information separation theorem and show that, under a boundary-mass condition, the hidden-environment tail of inverse projected Fisher controls robust classification sample complexity. We then derive stable ridge certificates and repair principles: reserves lift Fisher-null tails, portfolios increase collapse codimension, and tail-transverse influence selects counterfactual repairs by combining Fisher-Glass directionality with loss leverage. Synthetic experiments validate the theory; WILDS/Waterbirds diagnostics show that weakest identifiability need not mean lowest accuracy; and Waterbirds training plus counterfactual repair show that Fisher-Glass controls an identifiability axis complementary to loss.
PaperID: 2583, Poster
Abstract: Diffusion-based video matting has recently emerged as a leading approach, leveraging video diffusion priors from internet-scale pretraining to achieve strong zero-shot generalization from limited synthetic data. However, the heavy compute of these diffusion backbones limits inference to a short chunk of just 6--24 frames, so even ordinary videos span multiple chunks. Because each chunk is matted independently from coarse mask guidance, existing methods fail to maintain consistency across chunk boundaries, producing flickering edges and discontinuous transparency. We address this with a mask-guided diffusion-matting framework built on two ideas. The RGB and mask conditioning are represented as separate token streams that interact through full spatio-temporal self-attention, letting the model adaptively combine fine RGB texture with coarse mask layout rather than committing to a fixed channel-wise fusion. On top of this, we propose a novel pair of learnable embeddings, the preserve token and the refine token, that act as a per-frame conditioning interface and enable a sliding-window inference scheme in which already-generated mattes are propagated across chunk boundaries for local temporal consistency, complemented by a globally shared reference anchor at every chunk to prevent error accumulation in long sequences without any additional training. Across standard video matting benchmarks and in-the-wild long footage, our framework outperforms prior diffusion-based methods in both fine-detail accuracy and temporal consistency, and uniquely maintains consistency across chunk boundaries on long videos. Our code and weights will be publicly released.
PaperID: 2584, Poster
Abstract: Given the profound impacts on human health and ecosystems, air pollution necessitates robust quality prediction to inform policy-making. Traditional physics-based forecasting relies on idealized closed-system assumptions, whereas data-driven approaches often lack interpretability. Recent hybrid methods, however, neglect data homogeneity, resulting in redundancy, parameter inefficiency, and gradient conflicts. Moreover, high-concentration regions inherently exhibit drastic fluctuations, making it difficult to explicitly distinguish them from the overall pollution field, especially when using spatial information alone. To address these limitations, a Physics-Anchored Spatial-Spectral (PAS^2) paradigm is proposed. Specifically, a collaboration scheme is firstly elaborated, within which a physical estimator captures deterministic trends, anchoring the neural branch’s input space and directing it to learn complementary fluctuations. Building on this, we equip the neural branch with a Spatial–Spectral (S^2) Reasoner that integrates graph operators to model local spatial diffusion and spectral operators to explicitly capture the drastic fluctuations in high-concentration regions. Experiments on real-world datasets show that PAS^2 consistently improves over strong data-driven and physics-guided baselines, especially under volatile sudden-change scenarios. Codes are attached and will be released at GitHub.
PaperID: 2585, Poster
Abstract: Policies trained on offline logs can perform well in open-loop evaluation but fail in closed-loop deployment, because they never learn to recover from their own mistakes. Fine-tuning these policies typically requires either many human demonstrations or large-scale reinforcement learning rollouts, both of which are expensive. We propose RAHF, a two-stage closed-loop fine-tuning framework that reduces human effort by amplifying a small number of human corrections with a cheap verifiable reward. In Stage 1, a human operator monitors the policy during closed-loop rollout and intervenes at dangerous states, producing a small but targeted correction dataset that also keeps the training process safe. In Stage 2, the improved policy runs without any human involvement. At each step, the verifiable reward scores the policy's candidate proposals. We use these reward scores to construct group-normalized advantages and fine-tune the policy toward the higher-advantage proposals. To prevent the policy from forgetting its pretrained knowledge, we propose a Policy Adaptation Module (PAM) that freezes the pretrained weights and learns closed-loop corrections through a separate trainable branch. On BridgeSim-NavHard, RAHF with a human teacher reaches a driving score of 78.90, surpassing all baselines including DAgger, HG-DAgger, GRPO, and AWR. In addition, PAM reduces forgetting and enables zero-shot transfer to HUGSIM without any additional fine-tuning.
PaperID: 2586, Poster
Abstract: Reinforcement learning sharpens LLM reasoning, but parameter updates are expensive, forgetful, and opaque; training-free libraries swap gradients for a textual store yet inherit trajectory-level admission and flat structure. We close these gaps with HyperSkill, a Training-Free Omnimodal GRPO framework that keeps the backbone frozen and replaces gradient updates with auditable edits to a Hypergraph-Indexed Skill-Library, where skills are advantage-bearing nodes and compositions are hyperedges. Three mechanisms drive its evolution: a counterfactually verified, step-level distillation admits new skills; moving-average updates with hyperedge growth and pruning follow retrieval feedback; and an advantage-weighted retriever with similarity and capacity gates falls back to zero-skill. Across 11 text/VL/omnimodal benchmarks, HyperSkill beats every baseline by 6.9 F1 (15.6%) without touching a parameter. Our software and data are publicly available.
Authors: Eleanor Quint
Abstract: Models that are indistinguishable on in-distribution data can behave very differently under distribution shift. We introduce Perturb-and-Correct (P&C), a post-hoc method for constructing epistemically diverse predictors from a single pretrained network. P&C applies random hidden layer perturbations with a least-squares correction in the subsequent affine layer, producing predictors that agree on calibration data while remaining free to disagree away from it. We analyze this mechanism through the post-correction residual and its first-order sensitivity: the residual is controlled near the calibration distribution by a leverage term, while corrected sensitivity grows as inputs deviate from the calibration geometry. Empirically, P&C achieves a strong ID/OOD tradeoff across MuJoCo dynamics prediction and CIFAR-10 OOD detection, matching or outperforming standard post-hoc baselines while requiring only a single pretrained model. Our findings highlight the potential in further exploiting overparameterization as a strength of deep learning models.
PaperID: 2588, Poster
Authors: Ting Lei, Yu Chen, Yang Liu
Abstract: Hallucination remains a fundamental challenge in multimodal large language models (MLLMs), often stemming from unverified visual assumptions and reasoning processes that are weakly grounded in observable evidence. Existing approaches attempt to mitigate hallucination through self-reflection, multi-agent debate, or counterfactual verification. However, these methods remain largely opinion-driven: they lack explicit verification of the visual understanding process and may fail when multiple agents share the same incorrect perceptual assumptions, leading to consistent yet hallucinated conclusions. A promising direction is to augment MLLMs with external visual tools (e.g., OCR and grounding models) to acquire verifiable evidence. However, effective tool-augmented reasoning remains challenging due to two key failure modes: unreliable tool invocation, where models select inappropriate tools or produce incorrect parameters, and evidence--reasoning inconsistency, where models ignore or contradict the acquired evidence when forming final predictions. In this work, we propose EVA, an Evidence-seeking Visual Agent that reframes multimodal reasoning as an explicit process of evidence acquisition and verification. At the core of EVA is a unified evidence-grounded execution pipeline designed to address the above challenges through two tightly coupled mechanisms. First, we introduce Iterative Tool Replanning, which dynamically refines tool selection and parameters based on intermediate observations, improving the reliability of tool usage. Second, we develop Evidence Consistency Verification, an answer validation mechanism that improves alignment between the final prediction and the collected evidence, mitigating evidence--reasoning inconsistencies. In addition, to balance robustness and efficiency, EVA incorporates Selective Evidence Routing, an adaptive fast--slow reasoning strategy that determines when external evidence is necessary. Visually straightforward queries are handled via direct reasoning, while complex or high-risk cases trigger tool-based evidence acquisition. By jointly addressing how to reliably acquire evidence and how to use it consistently, EVA transforms multimodal reasoning from passive answer validation into explicit evidence-backed decision making. Experimental results demonstrate that EVA achieves significantly improved reasoning reliability over direct VLM inference and multi-agent methods, while maintaining efficiency through selective tool invocation. These findings highlight structured and adaptive evidence utilization as a practical paradigm for hallucination-resistant multimodal reasoning.
PaperID: 2589, Poster
Abstract: Real-world multi-agent reinforcement learning (MARL) systems must operate autonomously while adapting to natural language instructions provided by humans. These instructions arrive unpredictably, interrupt ongoing actions, and may conflict with long-horizon objectives under partial observability. Recent vision-language models enable instruction following, but are computationally expensive and limited to single-agent settings. Conversely, MARL can handle long-horizon coordination with macro-actions, but assumes uninterrupted execution. This work introduces macro-action value cancellation for instruction compliance (MAVIC) to enable interruption-aware MARL by bootstrapping the value of ongoing macro-actions upon receiving instructions. Agents can thus decouple task objective optimization from instruction-driven overrides within a unified policy and re-plan without destabilizing the underlying learning process. MAVIC integrates with standard actor-critic methods and achieves high instruction compliance while preserving base task performance across increasingly complex macro-action benchmarks.
Abstract: Compute-optimal scaling laws have become the central organizing tool of language-model training, yet they have been measured exclusively under Euclidean readouts. We give the first compute-optimal scaling-law characterization of hyperbolic language modeling. Trained from scratch on OpenWebText across a paired 5×4 grid spanning 86M to 1.32B parameters and 66M to 8.4B tokens, our hyperbolic LM HypTip follows a markedly different scaling law than its Euclidean counterpart. Fitting the Hoffmann form L = E + A·N^(−α) + B·D^(−β) to each track yields a compute-optimal data-to-parameter ratio of D/N = 8.4 for the hyperbolic readout versus 20.8 for the Euclidean baseline. The hyperbolic readout therefore reallocates compute by 2.5× toward parameter-heavy training, winning 17 of 20 paired cells across the grid. This shift is mechanistically grounded: across all 20 trained checkpoints, the per-row Lorentzian radius of the LM head correlates with token log-frequency at Spearman ρ down to −0.83, and the strength of this Zipfian token hierarchy scales with training-token count rather than parameter count, indicating that the geometry-as-prior emerges causally during training rather than being read off a frozen Euclidean checkpoint as in prior post-hoc analyses. Held-out evaluation on WikiText-103 perplexity, LAMBADA, and PIQA preserves the same advantage, and a cross-architecture replication on GPT-2 and DeepSeek-V3 confirms that LM-head geometry, not body details, drives the effect. Output-layer geometry alone therefore reshapes the compute-optimal scaling law of language modeling and brings hyperbolic LMs into the same predictive framework as the Euclidean models that have shaped the field.
PaperID: 2591, Poster
Authors: Yunhui Zhu, Buliao Huang
Abstract: Building segmentation has recently benefited from foundation segmentation models such as SAM, which offer strong generalization and scalable deployment through simple prompts. However, SAM-style models still produce pixel-wise masks, whose fragmented regions and irregular boundaries often fail to preserve the geometric regularity of buildings, making direct contour vectorization unreliable. Existing attempts to improve building polygonization typically modify the segmentation decoder or introduce additional polygon prediction branches, increasing adaptation costs and weakening the portability of pretrained foundation models. In this paper, we ask whether accurate building polygonization can be achieved directly from frozen segmentation outputs, without modifying or retraining the underlying model. To this end, we propose TopoRefine, a plug-and-play topology-aware contour refinement framework that converts predicted contours into geometry-consistent building polygons. Unlike conventional contour refinement methods, which follow a fixed-topology paradigm and can only adjust a predefined set of contour points, TopoRefine performs topology-adaptive polygon evolution by jointly modeling vertex insertion, vertex deletion, and coordinate refinement. This allows the refined polygon to correct structural errors inherited from frozen segmentation outputs and progressively align with the underlying building geometry. Extensive experiments on three building segmentation benchmarks demonstrate that TopoRefine generalizes across different segmentation pipelines and consistently improves both segmentation accuracy and geometric consistency, achieving notable gains such as +6.17% AP and +3.53% PolySim on SAM 3.
PaperID: 2592, Poster
Abstract: Multimodal large language models (MLLMs) achieve strong performance in translating WebUI screenshots into HTML/CSS. However, when they faced multiple frontend frameworks (React/Vue/Angular), they often suffer from negative transfer that leads to compilation failures. These failures commonly stem from mixing framework specific syntax and violating framework constraints. This paper studies supervised fine tuning for multi-framework WebUI code generation and proposes a phase-wise MLLM tuning method driven by compatibility and heterogeneity signals across frameworks. We first apply a normalization mapping that preserves each framework's native syntax while suppressing noise from variable naming and literal content. We then estimate cross framework heterogeneity and compatibility offline, and automatically derive training groups and their ordering by maximizing a phased objective. Finally, we perform phase-wise adapter tuning to reduce cross framework interference. Experiments on multiple benchmarks show that the proposed method improves both generation quality and compilation success rate compared with baselines.
PaperID: 2593, Poster
Abstract: Latent space reasoning improves inference efficiency by compressing chain-of-thought (CoT) reasoning into continuous latent states, but its safety alignment remains poorly understood. In this work, we conduct the first systematic safety study of latent space reasoning models under jailbreak attacks. Our results reveal that most latent reasoning methods exhibit higher vulnerability to jailbreak attacks than both the base model and token-space CoT reasoning counterparts, suggesting that moving reasoning into latent states can weaken safety alignment. To address this issue, we propose easoning framework that injects safety supervision into latent space reasoning. SaLR converts long-form safety reasoning into a compact four-block safe-chain and transfers this safety signal through teacher-student hidden state distillation. For safety instances, SaLR aligns an early safe response prefix to guide the model toward a safe response trajectory; for reasoning instances, it retains single position distillation to preserve latent reasoning ability. Experiments across model scales and harmful benchmarks show that SaLR substantially reduces jailbreak attack success while preserving reasoning accuracy and token efficiency. Beyond standard jailbreaks, SaLR also mitigates reasoning specific attacks such as token-inflation and CoT-hijacking attacks by keeping reasoning in fixed latent states and avoiding exposed textual reasoning traces. Overall, SaLR provides a stronger safety utility trade-off for latent space reasoning. Our anonymous code repository is available at: https://anonymous.4open.science/r/SaLR-3C19
PaperID: 2594, Poster
Abstract: Federated Incremental Learning (FIL) aims to learn sequential tasks across distributed clients under strict privacy constraints while mitigating catastrophic forgetting. However, existing methods often rely on storing, transmitting, or generating previous data to retrain the model, which not only introduces potential privacy risks but also incurs additional memory, communication, and computational costs, thereby limiting their deployment on resource-constrained edge devices. To tackle this challenge, we propose SPR-FIL, a novel rehearsal-free and lightweight \underlineStatistical \underlinePrototype \underlineRegularization framework for \underlineFederated \underlineIncremental \underlineLearning that leverages Gaussian statistics and preserves previous knowledge through aligned prototypes and subspace-constrained encoder updates. More specifically, the server aggregates client-side statistics into global prototypes and dynamically aligns historical prototypes with the evolving feature space through anchor-based moment matching. Locally, clients utilize these aligned prototypes to induce virtual posteriors which serve as soft supervision for distillation while encoder updates are further regularized by projecting gradients onto the orthogonal complement of covariance-induced historical knowledge subspaces. Extensive experiments demonstrate that SPR-FIL is model-agnostic, communication-efficient, and consistently outperforms state-of-the-art methods across various FIL settings, achieving improvements of up to 9.73%.
PaperID: 2595, Poster
Abstract: Multimodal large language models (MLLMs) provide a promising alternative by taking object detection as a sequence generation task. However, applying MLLMs to multimodal object detection poses challenges in visual cross-modal (RGB-Infrared) fusion and coordinate prediction under Cross-Entropy supervision. In this paper, we propose a novel MLLM-based object detection framework that addresses these challenges through improving visual cross-modal token alignment and coordinate generation. Specifically, we formulate RGB-IR fusion as pre-decoding token-space modality routing, implemented by the Cross-Modal Complementary Adapters and the Visual Token-wise Router. We further reformulate coordinate-token learning as scale-adaptive neighborhood likelihood maximization through our proposed Geometry-Aware Loss, bridging discrete token prediction and continuous geometric localization. Extensive experiments show that our method achieves state-of-the-art performance, reaching 67.81% mAP@0.5 under in-distribution scenarios, outperforming existing baselines by over 8%, while also exhibiting strong generalization to out-of-distribution and referring object detection tasks.
PaperID: 2596, Poster
Abstract: Large language models encode rich representations in their hidden states, yet these representations remain largely opaque and disconnected from the model's vocabulary. We introduce Geometric Embedding Mixture (GEM), a pre-training regularization method that aligns transformer hidden states with the geometry of the vocabulary embedding space. GEM utilizes a self-supervised objective to encourage the final hidden state to approximate a probability-weighted mixture of token embeddings. This effectively minimizes the free energy of the models and pushes the representation into the convex hull of the vocabulary. We pre-train models at three scales (135M, 360M, and 760M) and demonstrate that GEM yields superior predictive uncertainty compared to standard training, reducing perplexity by up to 53% on benchmarks. Furthermore, we show that this geometric structural constraint mitigates representation anisotropy and unlocks \emphnative interpretability. Unlike baseline models, GEM enables direct layer-wise decoding of intermediate hidden states and better causal faithfulness using the untuned language head, revealing coherent semantic trajectories without the need for auxiliary probes. Our work demonstrates that imposing geometric structure during pre-training improves both performance and transparency, offering a path toward language models that are interpretable by construction.
PaperID: 2597, Poster
Authors:
Ziyou Xiang, Chaowei Fang, Zhihong Wu, Jiliang Li, Di Xu, De ChengAbstract: Few-shot out-of-distribution (OOD) detection is essential for deploying recognition models in open-world environments, where a reliable model should correctly recognize in-distribution (ID) samples while rejecting samples from unseen categories. Recent CLIP-based methods have shown promising few-shot OOD detection ability by exploiting transferable vision--language representations. However, many methods rely mainly on global image features or emphasize only the most class-discriminative local cues, making them less effective for OOD detection where ID and OOD samples share visually similar local patterns. Although local prompting introduces patch-level reasoning, existing methods often treat local evidence as a binary distinction between ID-related and non-ID cues, leaving the ambiguous transition region between reliable ID evidence and clear non-ID evidence insufficiently modeled. To address this issue, we propose ThriPrompt, a progressive local prompting framework for few-shot OOD detection. TriPrompt decomposes patch-level evidence into three complementary components within a unified local feature space: Core-ID cues that provide reliable class-defining evidence, Far-ID cues that capture ambiguous transition-region evidence, and OOD cues learned from weak non-ID proxies mined directly from patch features. This patch-native design enables suppression-aware local reasoning without requiring additional OOD data or repeated multi-crop feature extraction. At inference time, \modelname combines global CLIP confidence with Core-ID, Far-ID, and OOD local responses to reduce overconfident ID assignment on hard OOD samples. Experiments on ImageNet-1K OOD benchmarks show that \modelname consistently improves few-shot OOD detection performance and remains compatible with mainstream CLIP-based scoring methods. Our code is provided in the supplementary material.
PaperID: 2598, Poster
Abstract: Recent advances in Multi-modal Large Language Models (MLLMs) have significantly improved video understanding. However, real-world scenarios often require going beyond isolated video inputs, giving rise to the emerging paradigm of video deep research, which involves open-world reasoning through integration of external knowledge. Despite this progress, existing approaches are limited by their inability to perform effective inter-video interaction and their reliance on static, one-shot datasets that yield noisy and suboptimal supervision. To address these challenges, we propose VideoSailor, an open-world video deep research agent powered by a trajectory-to-policy flywheel that iteratively converts generated reasoning trajectories into high-quality off-policy guidance to mitigate the limitations of low-quality on-policy rollouts. For systematic evaluations, we further introduce Omni-BrowseComp, a benchmark that emphasizes omni-modal evidence aggregation, cross-video reasoning, and fine-grained temporal grounding with reasoning-relevant segment annotations. Extensive experiments demonstrate that VideoSailor significantly improves performance on complex multi-hop and cross-video reasoning tasks, establishing a scalable and effective paradigm for open-world video deep research.
PaperID: 2599, Poster
Abstract: LLM-based theorem proving in Lean 4 has advanced rapidly, but provers are difficult to leverage newly contributed lemmas, because retrieval skews toward popular foundational lemmas, leaving the right premises buried among hundreds of thousands of candidates. Existing neural selectors rank candidates by semantic similarity, with three intrinsic limitations: frequency-skewed retrieval, isolated pairwise ranking, and degradation on out-of-distribution queries such as competition problems and informal language. Given that mathematics is structured by dependencies rather than similarities, we restore this signal to both the training data and the model. We curate two underused dependency resources. We conduct the first systematic mining of 53 Lean Blueprint projects, a previously untapped corpus of 3,805 manually curated informal+formal nodes that capture rare, research-level premises absent from Mathlib4. We further pair it with the extraction of comprehensive, typed, multi-level Mathlib4 dependencies. We then introduce DAPS, a dependency-aware premise selector with a structural neighborhood encoder, a group-level contrastive objective, and a Mathlib-then-Blueprint adaptation procedure. DAPS reaches Recall@32 of 89.31 on Mathlib4-Heldout (+11.47 over the strong selector LeanHammer). DAPS holds its lead across both out-of-distribution benchmarks, and improves three general-purpose LLMs on the informal IMOProofBench. Dependency graphs, benchmarks, and code are available 1 and will be released under a permissive license upon publication.
PaperID: 2600, Poster
Authors:
Haowen Hou, Zhen Huang, Zheming Liang, Qingyi Si, Chenglin Li, Shuai Dong, Kele Shao, Ruilin Li, Dianyi Wang, Nan Duan, Jiaqi WangAbstract: Video is temporally redundant: adjacent frames usually share most objects, background, and layout. Yet existing video multimodal large language models (video MLLMs) usually encode each sampled frame as an independent RGB image, causing visual tokens to repeat content already present in earlier frames. This suggests a more direct video interface: send a full reference frame only when the scene cannot be predicted well from prior context, and otherwise transmit a compact description of inter-frame changes. We call this interface a predictive visual code, and instantiate it for video MLLMs as AdaCodec. AdaCodec spends full visual tokens on a reference frame only when its conditional predictive cost is high; otherwise, it encodes inter-frame changes, including motion and prediction residuals, as compact P-tokens. Across all eleven benchmarks, AdaCodec improves over the Qwen3-VL-8B per-frame RGB baseline at a matched visual-token budget. Even at 1/7 the budget, AdaCodec with 32k tokens surpasses the 224k baseline on all long-video benchmarks; on five general-video benchmarks, it raises the average score while substantially cutting time-to-first-token from 9.26s to 1.62s.
PaperID: 2601, Poster
Abstract: We present a data poisoning attack---Phantom Transfer---with the property that, even if you know precisely how the poison was placed into an otherwise benign dataset, you cannot filter it out. We achieve this by modifying subliminal learning to work in real-world contexts and demonstrate that the attack works regardless of which model produces the data, which model is trained on the data or what the attack target is. The attack survives 11 tested data-level defences, including one where every sample is paraphrased by a different model. We characterise when this attack works best and show that it can be used to plant password-triggered behaviours into models while still beating defences. In short, we provide an existence proof that maximum-affordance defences can fail to stop sophisticated data poisoning attacks. We suggest that future defences should supplement by include white-box methods and post-training model audits.
Abstract: Reinforcement Learning with Human Feedback (RLHF) has become a central paradigm for aligning Large Language Models with human values. However, the reliance on sensitive preference data necessitates rigorous differential privacy (DP) guarantees to prevent user information leakage. Existing private alignment methods typically assume dense reward parameterizations, resulting in estimation error bounds that scale polynomially with the ambient model size d. Such dependence is prohibitive in modern overparameterized models where d is vast. In this work, we study \emphPrivate Sparse RLHF problem to address this limitation, assuming that human preferences (reward model) are governed by a subset of s^ \ll d relevant features. We provide the first comprehensive theoretical framework for differentially private sparse reward estimation from pairwise preference feedback under the Bradley-Terry-Luce model, and introduce efficient algorithms for both Sample-DP and Label-DP settings. Theoretically, under the squared \ell_2 parameter estimation error, we establish upper bounds of \widetildeO(s/n + s^2/n^2\varepsilon^2) for (\varepsilon,\delta)-Sample-DP and \widetildeO(s/n + s/n\varepsilon^2) for \varepsilon-Label-DP based on randomized response, where n is the training data size. For the lower bounds, we show that the Sample-DP rate is nearly optimal, and that the guarantee for our Label-DP method cannot be further improved within the randomized-response framework. We also provide a general lower bound for the problem with Label-DP. Empirically, our framework demonstrates superior efficacy over dense baselines across sentiment generation and safety alignment benchmarks, offering a significantly more favorable privacy-utility trade-off.
Authors: Marek Černý
Abstract: Message-passing graph neural networks (MPGNNs) dominate modern graph learning. Typical efforts enhance MPGNN's expressive power by enriching the adjacency-based aggregation. In contrast, we introduce a graph-level aggregation over walk incidence-based matrices that are constructed to deliberately trade off some expressivity for stronger and more structured inductive bias. This approach allows for gradual scaling between classical message-passing and simpler methods based on walks. Our second result characterizes the expressive power at each scale using homomorphism counts over a hierarchy of generalized caterpillar graphs. Based on these foundations, we propose Caterpillar GNNs that substantially reduce the number of nodes in the hidden layers of the computational graphs on real-world datasets. This gradual reduction does not obstruct learning, enabling a systematic study of lower-order expressivity. We further describe benchmark settings where a task-aligned expressivity supports learning.
Abstract: Large language models (LLMs) have demonstrated remarkable capabilities, but they still frequently produce hallucinations. These hallucinations are difficult to detect in reasoning-intensive tasks, where the content appears coherent but contains errors like logical flaws and unreliable intermediate results. While step-level analysis is commonly used to detect internal hallucinations, it suffers from limited granularity and poor scalability due to its reliance on step segmentation. To address these limitations, we propose TokenHD, a holistic pipeline for training token-level hallucination detectors. Specifically, TokenHD consists of a scalable data engine for synthesizing large-scale hallucination annotations along with a training recipe featuring an importance-weighted strategy for robust model training. To systematically assess the detection performance, we also provide a rigorous evaluation protocol. Through training within TokenHD, our detector operates directly on free-form text to identify hallucinations, eliminating the need for predefined step segmentation or additional text reformatting. Our experiments show that even a small detector (0.6B) achieves substantial performance gains after training, surpassing much larger reasoning models (e.g., QwQ-32B), and detection performance scales consistently with model size from 0.6B to 8B. Finally, we show that our detector can generalize well across diverse practical scenarios and explore strategies to further enhance its cross-domain generalization capability.
PaperID: 2605, Poster
Abstract: Single-step retrosynthesis prediction serves as the core operational unit for automated synthesis planning and remains a fundamental computational challenge in pharmaceutical design. While a critical prior in organic chemistry is that chemical reactions typically preserve the majority of the molecular scaffold, current template-free deep learning models struggle to effectively exploit this characteristic. Existing methods attempt to exploit this by minimizing sequence-level edit distance through carefully chosen graph traversal orders. However, this structural correspondence is only indirectly induced and remains highly sensitive to imperfect traversal heuristics, which limits its effectiveness for complex transformations. To address this limitation, we introduce the Matching-while-Decoding paradigm, replacing implicit sequence alignment with direct structural supervision. We operationalize this paradigm within a graph-to-sequence architecture that explicitly maps generated reactant tokens back to the input product graph during decoding. Furthermore, our hierarchical autoregressive formulation utilizes a custom causal attention mask to enable highly efficient, single-pass parallel training. Extensive experiments across the USPTO-50K, USPTO-MIT, and USPTO-FULL benchmarks reveal a critical accuracy and robustness trade-off in existing methodologies, where models either over-rely on augmentation or fail to leverage it. Our approach uniquely resolves this bottleneck by delivering highly competitive canonical accuracy while scaling exceptionally well with test-time augmentation, establishing new state-of-the-art records on USPTO-MIT and augmented USPTO-FULL. By providing this balanced stability across diverse inference protocols, our approach emerges as a highly promising foundation for dependable single-step retrosynthesis prediction.
PaperID: 2606, Poster
Abstract: We study the classical problem of linear regression under heteroskedastic noise, where observations arise from multiple sources with unknown and potentially unequal variances. Specifically, we consider a setting with n sources, each providing a linear measurement of an unknown d-dimensional parameter vector \beta^, where each source is associated with an unknown noise variance. We focus on the classic "subset-of-signals'' model, where one assumes that m of the sources have a relatively low noise, formally, with variance of at most 1. Our goal is to accurately estimate \beta^\ despite the presence of high-variance sources. We show that if m is sufficiently large, it is possible to recover \beta^ with sub-constant error. Specifically, we prove a recovery error bound of \tildeO \left(\frac\sqrtndm\right), significantly improving upon previously known results for this setting.
PaperID: 2607, Poster
Abstract: Given the inherent unpredictability of packet loss in vehicular wireless communications, V2X collaborative perception can yield practical benefits only if agents can achieve reliable collaboration under lossy and low-bandwidth communication conditions. Existing dense BEV feature fusion methods depend on redundant BEV feature exchange, which is infeasible in low-bandwidth scenarios, while compact-communication methods aggressively compress messages but can hardly recover the missing feature content after packet loss. In this paper, we present LR-V2X, a loss-resilient, latent-space reconstruction framework that converts corrupted received latents (even under severe 90% packet loss) into a spatial prior and then reconstructs the missing BEV information from this informative prior and using ego context as condition. Notably, the model can be trained under complete communication conditions and can be directly applied to lossy conditions at test time, eliminating the need for training under numerous lossy conditions. Experiments on DAIR-V2X and V2XREAL show that LR-V2X delivers the strongest robustness under severe packet loss and preserves reliable collaboration as communication quality degrades. And it reduces communication overhead by 64× compared to dense BEV feature fusion baselines. Code will be publicly released.
Authors: Aris Filos-Ratsikas, Georgios Kalantzis, Fangxiao Wang
Abstract: We study the problem of finding approximate envy-free allocations up to any k goods (\alpha-EFkX), when agents have additive values over goods in a bundle. As our main result, we show that for any k>2, \frack+1k+2-EFkX allocations exist for any number of agents, and can be computed in polynomial time, via an appropriate generalization of the 3PA algorithm of (Amanatidis et al., 2024). An immediate corollary of this result is that 3/4-EF2X allocations exist for any number of agents; in contrast, 2/3-EFX allocations are only known to exist for up to 7 agents. We improve this latter result by devising an algorithm that achieves 2/3-EFX for 8 agents. We also consider EFkX graph orientations; we prove that such orientations do not always exist, and that deciding their existence is NP-complete, thereby generalizing the corresponding result of (Christodoulou et al., 2023) for k=1.
PaperID: 2609, Poster
Abstract: Transformer-based time series forecasting models often suffer from block-wise attention collapse, where temporal smoothness and cross-variable correlations concentrate attention locally and restrict global modeling capacity. We argue that this collapse is partly a pre-attention representation-conditioning problem: attention logits inherit the spectral concentration and low effective rank of the input representation. We propose Widely Linear Frequency-Domain (WILD) conditioning, a lightweight, model-agnostic module inserted before self-attention. WILD maps inputs to the frequency domain, applies data-dependent spectral preconditioning to redistribute over-dominant frequency energy, and then uses asymmetric real/imaginary projections to realize a widely linear transform beyond fixed time-domain channel mixing. The preconditioning step is essential: it constructs the usable normalized frequency basis on which the widely linear branch can operate. Thus, WILD uses frequency-domain modeling to condition an existing attention layer, not to replace it with a new forecasting architecture. Extensive experiments on eight benchmark datasets show positive gains on 105 of 112 model-dataset-metric configurations across seven backbones. Attention diagnostics show that WILD substantially increases effective rank over the baseline, and a W i ablation identifies spectral preconditioning as the primary mechanism, with widely linear mixing providing a secondary, dataset-dependent gain.
PaperID: 2610, Poster
Abstract: Fine-tuning large-scale foundation models across resource-constrained, distributed clients presents significant computational and communication challenges. While merging Low-Rank Adaptation (LoRA) with Federated Learning (FL) enables rank-based flexibility, the interdependence of the adaptation matrices (A and B) often complicates optimization. In this work, we theoretically demonstrate that LoRA fine-tuning with a frozen A matrix acts as a stochastic approximation of full fine-tuning, where batch gradients are subject to a random, rank-dependent perturbation. Motivated by this insight, we propose freezing and periodically resampling A matrices at the client level. This strategy decouples the A and B gradients, diminishes LoRA low-rank constraints and reduces both computational and uplink communication overheads by 50%, allowing for the reallocation of resources toward higher adaptation ranks. Furthermore, we show that in this frozen-A regime, the adaptation rank r directly governs the signal-to-noise ratio (SNR) of the perturbed gradients. Specifically, we find that the SNR increases linearly with r, implying that clients’ learning rates must be dynamically scaled with their adaptation ranks for higher utility. Building on these foundations, we introduce Frank-LoRA, a federated rank-aware fine-tuning strategy that utilizes frozen, resampled A matrices and rank-dependent learning rates for clients. We also prove the convergence of the algorithm in FL settings with heterogeneous clients resources. Our experiments across benchmark datasets demonstrate that Frank-LoRA consistently outperforms state-of-the-art baselines with half the overhead on clients.
PaperID: 2611, Poster
Abstract: While all-in-one image restoration models excel in controlled, closed-set environments, they face critical limitations when deployed in open-world settings where paired supervision is unavailable and degradations are unknown and often intricately mixed. To address this challenge, we propose a continual restoration framework that integrates degradation discovery with adaptive restoration. Building on a model pre-trained on known degradation categories, our method discovers unseen degradations in feature space and continually expands the degradation space. We then introduce a discovery- and instance-conditioned descriptor that jointly captures discovered degradation priors and image-specific characteristics, enabling robust restoration under ambiguous and intricately mixed degradations. To further support adaptation under unpaired supervision, we adopt a mean-teacher-based semi-supervised framework equipped with a reliable bank for pseudo-target refinement. Specifically, we propose a discovery-adaptive score that assesses pseudo-target reliability using feature-space distances to degradation categories, enabling pseudo-target selection to remain adaptive to newly discovered degradations rather than relying solely on conventional perceptual quality-based scoring methods. Experiments on both controlled open-world protocols and real-world adverse weather datasets demonstrate that our approach consistently outperforms existing methods across multiple image quality metrics while improving robustness to degradation composition shifts.
Authors:
Julius Richter, Yoshiki Masuyama, Christoph Boeddeker, Takahiro Edo, Gordon Wichern, Jonathan Le RouxAbstract: We propose a plug-and-play framework for speech enhancement and separation that augments predictive methods with a generative speech prior. Our approach, termed Stochastic Interpolant Prior for Speech (SIPS), builds on stochastic interpolants and leverages their flexibility to bridge predictive and generative modeling. Specifically, we decompose the interpolation dynamics into a task-specific drift and a stochastic denoising component, allowing a predictive estimate to be integrated directly into the generative sampling process. This results in a mathematically grounded framework for combining strong pretrained predictors with the expressive power of generative models. To this end, we train a score model using only clean speech, yielding a degradation-agnostic prior that can be reused across tasks. During inference, the predictor provides a deterministic drift that steers the sampling process toward a task-consistent estimate, while the score model preserves perceptual naturalness. Unlike prior hybrid approaches, which typically rely on architecture-specific conditioning and are tied to particular predictors or degradation settings, SIPS provides a unified framework that generalizes across predictors and additive degradation tasks. We demonstrate its effectiveness for both speech enhancement and speech separation using recent predictors such as SEMamba and FlexIO. The proposed method consistently improves perceptual quality, achieving gains up +1.0 NISQA for speech separation.
Abstract: Diffusion models are a leading paradigm for graph generation, with notable impact in domains such as molecular design. Yet, scaling these models to large graphs remains an open problem. We approach this question in the dense-graph setting through the lens of graphons, the size-agnostic limit object of dense graph sequences, to study how structural graph statistics behave across node-size scales. This perspective leads to DiPhon, a diffusion process for size-scalable graph generation. Specifically, we formulate a continuous diffusion process on the graphon space via a Jacobi stochastic differential equation (SDE), and propose DiPhon, a discretized scheme that imitates it on graphs. We further derive the corresponding reverse-time process, which requires access to the marginal score. For the Jacobi process, this score interestingly admits a tractable form, which we estimate from data via graph denoising and plug into the reverse process to generate graph samples. We prove that DiPhon matches the first moment of the marginal distributions induced by the continuous graphon process exactly, and approximates the second moment up to a closed-form discrepancy. In this way, DiPhon inherits the size-agnostic statistical properties of the graphon dynamics can be used to solve the scaling bottleneck. Empirically, we demonstrate the scalability of DiPhon by training on small graphs and generating substantially larger ones at inference time, without retraining on large graphs.
PaperID: 2614, Poster
Abstract: Multi-view 3D reconstruction is increasingly driven by feed-forward models, which are fast and robust but often imprecise or globally inconsistent across views, especially in sparse-view and low-overlap settings. Structure-from-motion (SfM) offers a complementary geometric signal through camera poses and sparse 3D points estimated by correspondence filtering and optimization. To propagate this globally consistent \emphSfM scaffold into dense geometry, we introduce Scaffold3D, an SfM-conditioned reconstruction framework that injects image patch-aligned SfM point tokens into a pairwise feed-forward pointmap predictor. Our approach preserves image-token reasoning while using explicit SfM structure to condition dense prediction. Since pairwise pointmaps are predicted in local frames, we fuse them with global alignment and an SfM anchoring term. Across established benchmarks (ScanNet++, ETH3D, and Tanks&Temples) and out-of-domain data (4D-DRESS and MV-dVRK), Scaffold3D achieves stronger overall performance than feed-forward reconstruction models, point-token propagation, depth-completion baselines, and recent geometry-conditioned models. The gains are especially clear in low- and no-overlap evaluations. Our two-view SfM-conditioned pointmap predictor also outperforms several multi-view geometry-conditioned baselines, indicating that how geometric inputs are fused with image-token reasoning is as important as the number of conditioned views. Together, these results establish SfM-scaffold conditioning as a practical interface between learned pointmap prediction and classical geometric optimization.
PaperID: 2615, Poster
Abstract: We study convex optimization problems over a compact convex set where projections are expensive but a linear minimization oracle (LMO) is available. We propose the adaptive conditional gradient sliding method (AdCGS), a projection-free and line-search-free method that retains Nesterov's acceleration with adaptive stepsizes based on local Lipschitz estimates. AdCGS combines an accelerated outer scheme with an LMO-based inner routine. It reuses gradients across multiple LMO calls to reduce gradient evaluations, while controlling the subproblem inexactness via a prescribed accuracy level coupled with adaptive stepsizes. We prove accelerated rates for convex objective functions, matching projection-based methods, without relying on a projection oracle. For locally strongly convex objective functions, we further establish linear convergence without additional geometric assumptions on the constraint set, such as polytopes or strongly convex sets. Experiments on constrained \ell_p regression, logistic regression, and least-squares problems demonstrate that AdCGS improves over projection-free baselines and provides competitive performance when projections are inexpensive.
Abstract: In many critical domains, features are not freely available at inference time: each measurement may come with a cost of time, money, and risk. Longitudinal prediction further complicates this setting because both features and labels evolve over time, and missing measurements at earlier timepoints may become permanently unavailable. We propose CAPE (Contrastive Acquisition-Plan Embeddings), a partitioning acquisition policy for longitudinal active feature acquisition. CAPE first defines NOCT, an oracle-based objective that scores a set of future feature-time acquisitions by its expected predictive loss together with its acquisition cost. Based on NOCT, CAPE learns a contrastive acquisition-utility manifold of masked partial observations, where proximity reflects the expected utility of future acquisitions. Offline, training states are embedded and partitioned at each rollout timepoint; candidate future plans are scored within each partition using plug-in NOCT, and the best partition-level acquisition plans are cached. During inference, CAPE retrieves the cached plan via nearest-centroid lookup, enabling low-latency adaptive acquisitions. Experiments on synthetic and real-world healthcare datasets demonstrate that CAPE outperforms nonparametric, greedy, and RL-based alternatives, achieving higher accuracy at lower acquisition costs.
PaperID: 2617, Poster
Authors:
Haonan Tan, Xin Li, juyi peng, Zhihong Xia, Yuxing Han, Gene WenAbstract: Cauchy activations and Cauchy-adaptive layers can give large gains on scientific learning tasks, but the gains are selective and not uniquely Cauchy-specific. We organize this selectivity with an empirical loss-architecture alignment taxonomy: an architecture can help when the objective exposes structures it can use, and can lose under a mismatched objective. In score learning, Cauchy score networks improve over SiLU under a Fokker-Planck PDE loss, but the same advantage disappears under denoising score matching and a Gaussian activation is comparable to Cauchy under the PDE loss. In loss-activation factorials, robust residual losses explain more of the heavy-tailed and spectral gains than the activation alone; an Allen-Cahn trap initially escaped by CauchyAct+Cauchy loss is also escaped, sometimes more strongly, by CauchyAct+L1, CauchyAct+Huber, and ReLU+Cauchy loss. We also analyze ResidualCAN, a CauchyAct main path plus a softmax-weighted Cauchy Adaptive Node residual. It achieves 1520x lower stationary Allen-Cahn error than a SiLU MLP at epsilon = 1.0 and remains strongest or second strongest against KAN-style, tuned SIREN/Fourier, and adaptive PINN baselines across an Allen-Cahn epsilon-suite. Ablations show that most of this gain comes from a dense local-basis residual rather than from Cauchy activation alone. A Poisson/Burgers boundary check prevents a broad PDE dominance claim: tuned SIREN/Fourier and KAN-style baselines can dominate away from transition-layer structure. The resulting claim is deliberately bounded: Cauchy components are one useful instance of robust/local-basis co-design in low-dimensional scientific tasks, not a universal activation or a uniquely rational mechanism.
PaperID: 2618, Poster
Authors:
Wenxin XU, Jinwei Lu, Hwanhee Kim, Chen J Zhang, Xiao-Yong Wei, Haoyang LI, Yuanfeng SONGAbstract: Real-world visualization requests are routinely ambiguous, incomplete, or factually incorrect, yet existing Text-to-Visualization (Text-to-Vis) systems assume well-specified inputs and produce charts in a single pass. When queries are imperfect, a system must interact with the user to recover the true intent, but no benchmark or method supports this dynamic process. We introduce VisInteract, a new paradigm that reframes Text-to-Vis as interaction-driven intent recovery, and VisInteract-Bench, to our knowledge, that is the first benchmark for dynamic interactive Text-to-Vis, featuring controlled imperfection injection, a leakage-controlled User Agent for realistic multi-turn feedback, and dual-perspective (code and chart) automated evaluation. On the algorithmic side, we propose Vis-MCTS, a Monte Carlo Tree Search (MCTS) enhanced method, introducing improvements over classical MCTS, that Progressive Widening to tame the unbounded tool-argument space in tree search, cross-rollout information sharing so clarifications and critiques benefit the entire search tree, and Dimension-Aware Reward Decomposition that routes scalar user feedback along data-fidelity, visual-design, and intent-alignment dimensions to resolve credit assignment across heterogeneous actions. Extensive Experiments across two LLM backbones show that Vis-MCTS consistently outperforms all Text-to-Vis baselines, improving end-to-end task success by 13.40%--16.27% over the strongest interactive baseline and by more than 5× over non-interactive ones.
Abstract: LLM-based agents execute multi-step tasks, but their behavioral structure remains opaque: long unstructured traces resist the safety auditing and runtime monitoring that deployment requires. Existing approaches operate per-trace or success-only, missing the cross-run topology that links next-step and failure prediction; we extract provably minimal finite-state machines (FSMs) via prefix-tree construction and structural state merging, providing a structural substrate for the unpredictable nature of LLM agent behavior. Across twelve public datasets, the FSMs are compact (7–43 states), achieve high replay fitness on held-out data with zero structural variance across splits, and build in milliseconds. This substrate addresses both prediction goals. For next-step prediction, FSM-state context beats Agent Workflow Memory on all eight ground-truth-matched datasets. For failure prediction, per-state behavioral features reach held-out AUROC up to 0.94, and an online monitor flags failing SWE-agent runs at rank-AUROC above the trivial flag-everything baseline, enabling early-stopping at 32% trace completion. A single FSM replays four LLMs at perfect fitness, evidence that behavioral topology in LLM agents is shaped more by the deployment harness than by the LLM and giving a model-agnostic structural primitive for safety auditing and runtime monitoring at deployment.
Abstract: Dynamics of interacting systems in engineering, society, and nature often evolve over latent networks that govern which entities can interact. We study the problem of inferring these networks from event-based observations, which arise naturally in finance, seismology, and neuroscience. While there is substantial algorithmic work addressing this important problem, theoretical results are scarce. In this paper we ask the following fundamental question: what is the minimum time that one must observe the dynamics in order to exactly recover the underlying network, as a function of the number d of interacting entities? For a class of stationary Hawkes processes with sparse, weak interactions, we prove that an observation time of order \log d is sufficient and necessary. For the upper bound we construct a two-stage estimator that uses clipped and binned event data for screening, followed by a least-squares refinement, and apply concentration bounds derived from the Poisson cluster representation. For the lower bound we combine Fano’s inequality with Jacod’s Girsanov formula for point processes on a suitable subclass of networks.
PaperID: 2621, Poster
Authors:
Mai Thao Dang, Feng Jiang, Hehuan Ma, Yuzhi Guo, Jingquan Yan, Haiqing Li, Saiyang Na, Zheng Zheng, Thuc A Tran, Jean Gao, Junzhou HuangAbstract: Single-cell multi-omics technologies jointly measure multiple modalities from the same cell, such as gene expression, chromatin accessibility, and surface protein abundance, providing a richer view of cellular state than any single modality alone. A central challenge is how to learn a multimodal cell representation that preserves the different types of information contained in these co-profiled measurements. Existing integration methods often emphasize either cross-modal correspondence, which benefits retrieval and imputation, or fused discriminative structure, which benefits classification, but these objectives can favor different aspects of the data. As a result, a representation optimized for one setting may fail to preserve information needed for another. We introduce scDecomp, a decomposed representation learning framework for co-profiled single-cell multi-omics data. Motivated by the concepts of redundancy, uniqueness, and synergy, scDecomp separates multimodal information into three role-specific components: a redundancy component (R) for information shared across modalities, a uniqueness component (U) for modality-specific information, and a synergy component (S) for interaction-dependent multi-omic signals. This design allows the learned representation to retain shared cell-state structure while also preserving modality-specific and cross-modal regulatory information. On single-cell multi-omics benchmarks, the R component achieves a 5.0% reduction in FOSCTTM compared with the strongest baseline, while U+S improves cell type classification accuracy by 1.38%. Branch-level analyses further show that the decomposed components support complementary aspects of representation quality, suggesting that structured decomposition provides a practical strategy for learning more informative multimodal cell representations.
PaperID: 2622, Poster
Authors: Shurui Liu, Weide Chen, Changwang Yi, Ancong Wu
Abstract: Scalable Vector Graphics (SVG) are essential for modern industrial Computer-Aided Design (CAD). However, existing autoregressive SVG generation models are predominantly tailored for artistic creation and struggle to maintain the rigorous geometric fidelity and cross-view spatial alignment required for engineering drawings. To bridge this gap, we introduce DrawingsDreamer, a unified Large Language Model (LLM)-driven framework for multi-view vector-based engineering drawings generation. By formulating the generation of multi-view engineering drawings purely as a sequence modeling task, we eliminate the need of raster image encoders. We propose a Streamlined Representation utilizing hierarchical postfix tokenization, which guides the model to establish local geometric coordinates before assigning semantic boundaries. Optimized via a progressive task-aware curriculum schedule, DrawingsDreamer effectively transitions from localized structural repair to macroscopic generation in a unified model. Extensive experiments demonstrate that our unified model achieves strong performance in both geometric fidelity and syntactic accuracy across diverse conditional and unconditional generation tasks.
PaperID: 2623, Poster
Abstract: Cooperative multi-agent reinforcement learning (MARL) often fails under exogenous variability such as shifts in opponents or environment regimes that cannot be controlled by the team but alter the transition dynamics. We formalize this challenge as Exogenous Dec-POMDP (ED-POMDP), which decomposes the global state into endogenous variables controllable by the team and exogenous variables beyond its control. This formulation exposes delayed influence: While actions do not directly determine exogenous transitions, they can shape them through induced changes in endogenous state. Based on this, we propose LEICA, a CTDE-compatible algorithm that learns history-conditioned endogenous and exogenous context representations and shapes policy updates using influence-weighted intrinsic rewards. Across SMAX benchmarks with opponent strategy shifts, LEICA consistently improves both training performance and generalization to unseen opponents' strategies over existing baselines, supporting the usefulness of influence-weighted shaping under exogenous train-test regime shifts.
Authors:
Tycho F van der Ouderaa, Mart van Baalen, Paul Whatmough, Markus NagelAbstract: Scalar quantization of large language models (LLMs) is fundamentally limited by information-theoretic bounds. While vector quantization (VQ) overcomes these limits by encoding blocks of parameters jointly, practical implementations must avoid the need for expensive lookup mechanisms or other explicit codebook storage. Lattice approaches address this through highly structured and dense packing. This paper explores the Leech lattice, which, with its optimal sphere packing and kissing configurations at 24 dimensions, is the highest dimensional lattice known with such optimal properties. To make the Leech lattice usable for LLM quantization, we extend an existing search algorithm based on the extended Golay code construction, to i) support indexing, enabling conversion to and from bitstrings without materializing the codebook, ii) allow angular search over union of Leech lattice shells, iii) propose fully-parallelisable dequantization kernel. Together this yields a practical algorithm, namely Leech Lattice Vector Quantization (LLVQ). LLVQ delivers state-of-the-art LLM quantization performance, outperforming recent methods such as Quip#, QTIP, and PVQ. These results highlight the importance of high-dimensional lattices for scalable, theoretically grounded model compression.
Abstract: Active perception—the ability of a robot to proactively select viewpoints to acquire task-relevant information—is essential for robust operation in real-world environments. However, existing approaches are typically limited to fixed objectives or constrained settings, and struggle to generalize to open-ended perception intents specified in natural language. We propose I-Perceive, a foundation model for language-conditioned active perception in large-scale indoor environments. Given a query image, a set of context images, and a natural language instruction, I-Perceive predicts a 6D camera pose that fulfills the specified perception intent. The model integrates a vision-language pathway for semantic grounding with a geometric reasoning pathway for multi-view 3D understanding, connected via multi-layer semantic fusion to enable language-conditioned geometric reasoning. To support scalable training, we construct a large-scale dataset of language-viewpoint pairs from both real-world scene-scanning data and simulated environments using an automated pipeline. Extensive experiments demonstrate that I-Perceive significantly outperforms strong baselines on prediction accuracy, viewpoint feasibility, and instructions alignment. The model exhibits strong zero-shot generalization to unseen scenes and instructions, and enables closed-loop active perception, progressively refining viewpoints over sequential interactions.
Abstract: Multimodal LLMs often produce fluent yet unreliable reasoning, exhibiting weak step-to-step coherence and insufficient visual grounding, largely because existing alignment approaches supervise only the final answer while ignoring the reliability of the intermediate reasoning process. We introduce SR-MCR, a lightweight and label-free framework that aligns reasoning by exploiting intrinsic process signals derived directly from model outputs. Five self-referential cues—semantic alignment, lexical fidelity, non-redundancy, visual grounding, and step consistency—are integrated into a normalized, reliability-weighted reward that provides fine-grained process-level guidance. A critic-free GRPO objective, enhanced with a confidence-aware cooling mechanism, further stabilizes training and suppresses trivial or overly confident generations. Built on Qwen2.5-VL, SR-MCR improves both answer accuracy and reasoning coherence across a broad set of visual benchmarks; among open-source models of comparable size, SR-MCR-7B achieves state-of-the-art performance with an average accuracy of 81.4%. Ablation studies confirm the independent contributions of each reward term and the cooling module.
Abstract: Diffusion models are prone to generating structural hallucinations - samples that match the statistical properties of the training data yet defy underlying structural rules, resulting in anomalies like hands with more than five fingers. Recent research studied this failure mode from several viewpoints, offering partial explanations to their occurrence, such as mode interpolation. In this work, we propose a complementary perspective that treats hallucinations as instabilities on the model-induced manifold. We begin by showing that a hallucination filter based on such instabilities matches or exceeds the performance of the recently proposed temporal one. By tracing the source of these instabilities, we identify local intrinsic dimension (LID) as their primary driver and propose Intrinsic Quenching (IQ), a direct corrective mechanism that deflates it to alleviate hallucinations. IQ consistently outperforms standard hallucination reduction baselines across a wide array of benchmarks and offers a highly promising solution for enforcing anatomical consistency in downstream medical imaging tasks.
PaperID: 2628, Poster
Authors: Mohammad M Omati, Arash Amini, Nezam Mahdavi-Amiri
Abstract: In this paper, we investigate a novel discriminative dimensionality reduction method based on maximizing the minimum pairwise ratio of between-class to within-class scatter. This objective function enhances class separability by providing critical, adaptive control over the variance within each class pair. The resulting max-min fractional program is non-convex and challenging to solve. Our main contribution is F^3-PWCRA (Fast, Free of relaxation, Free of hyperparameters for Pairwise Worst-Case Ratio Analysis), a provably convergent two-level algorithm: an outer generalized Dinkelbach-type loop transforms the fractional objective into subtractive subproblems, ensuring the global convergence of the outer algorithm. Then, we propose a bisection strategy as an alternative, which offers robust convergence and straightforward implementation. For the inner loop, we develop an efficient minorization-maximization (MM) algorithm that tackles the non-convex subproblem by iteratively solving a simple quadratic program (QP), which we derive from the dual of a convex surrogate. F^3-PWCRA is computationally efficient, free from semidefinite relaxations and hyperparameter tuning, and extensive benchmarks show it consistently outperforms state-of-the-art methods in classification error.
Authors:
Seokju Cho, Ryo Hachiuma, Abhishek Badki, Hang Su, Byung-Kwan Lee, Chan Hee Song, Sifei Liu, Subhashree Radhakrishnan, Seungryong Kim, Frank Wang, Min-Hung ChenAbstract: Spatial reasoning, the ability to determine where objects are, how they relate, and how they move in 3D, remains a fundamental challenge for vision-language models (VLMs). Tool-augmented agents attempt to address this by augmenting VLMs with specialist perception modules, yet their effectiveness is bounded by the through which those tools are invoked. In this work, we study how the design of this interface shapes the agent's capacity for open-ended spatial reasoning. Existing spatial agents either employ single-pass code execution, which commits to a full analysis strategy before any intermediate result is observed, or rely on a structured tool-call interface that often offers less flexibility for freely composing operations or tailoring the analysis to each task. Both designs offer limited flexibility for open-ended, complex 3D/4D spatial reasoning. We therefore propose maintains a stateful Python kernel pre-loaded with input frames and a suite of perception and geometry primitives, letting a VLM-backed agent write one executable cell per step conditioned on all prior outputs, enabling the agent to flexibly compose and manipulate perception results and adapt its analysis to both intermediate text and visual observations and the demands of each problem. Evaluated across
Authors:
Boyi Li, Yifan Shen, Yuanzhe Liu, Yifan Xu, Jiateng Liu, Xinzhuo Li, Zhengyuan Li, Jingyuan Zhu, Yunhan Zhong, Lan Fangzhou, Jianguo Cao, James Rehg, Heng Ji, Ismini Lourentzou, Xu CaoAbstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success in open-vocabulary perceptual tasks, yet their ability to solve complex cognitive problems remains limited, especially when visual details are abstract and require visual memory. Current approaches primarily scale Chain-of-Thought (CoT) reasoning in the text space, even when language alone is insufficient for clear and structured reasoning, and largely neglect visual reasoning mechanisms analogous to the human visuospatial sketchpad and visual imagery. To mitigate this deficiency, we introduce Cognitive Supersensing, a novel training paradigm that endows MLLMs with human-like visual imagery capabilities by integrating a Latent Visual Imagery Prediction (LVIP) head that jointly learns sequences of visual cognitive latent embeddings and aligns them with the answer, thereby forming vision-based internal reasoning chains. We further introduce a reinforcement learning stage that optimizes text reasoning paths based on this grounded visual latent. To evaluate the cognitive capabilities of MLLMs, we present CogSense-Bench, a comprehensive visual question answering (VQA) benchmark assessing five cognitive dimensions. Extensive experiments demonstrate that MLLMs trained with Cognitive Supersensing significantly outperform state-of-the-art baselines on CogSense-Bench and exhibit superior generalization on out-of-domain mathematics and science VQA benchmarks, suggesting that internal visual imagery is potentially key to bridging the gap between perceptual recognition and cognitive understanding. We will open-source the CogSense-Bench and our model weights.
PaperID: 2631, Poster
Abstract: Neural network training can be viewed as a discrete-time dynamical system, suggesting that its optimisation trajectory can be learned and extrapolated to reduce the cost of repeated backpropagation. Motivated by Koopman operator theory, which represents nonlinear dynamics as linear evolution in a lifted space, we propose a training acceleration method based on augmented Koopman dynamics. Instead of approximating the evolution of parameters alone, our method lifts the training state to include both model parameters and optimiser-dependent internal states, enabling the learned operator to better capture the dynamics of modern optimisers with momentum or adaptive gradient statistics. We estimate a strided, multi-step Koopman operator from training snapshots and use matrix-vector multiplications to extrapolate future parameters and optimiser states, effectively bypassing an adaptive number of backpropagation steps. This formulation suppresses high-frequency stochastic fluctuations induced by mini-batch training and reduces the storage required for operator estimation. To improve robustness, we further introduce a safeguarding mechanism that prevents Koopman-extrapolated states from degrading performance relative to the most recent backpropagation iterate. The method can be integrated with standard optimisers, including Adam, AdamW, SGD with momentum, Adadelta, and other commonly used variants. Experiments across multiple optimisers show that the proposed approach significantly reduces both the number of backpropagation steps and the wall-clock training time required to reach a target accuracy, while maintaining, and in many cases improving, final model performance.
Abstract: Estimating counterfactual distributions under interventions is central to treatment risk assessment and counterfactual generation tasks. Existing approaches model the counterfactual distribution as a standalone generative target, without exploiting its relationship to the observational data. In this work, we show that under standard assumptions, observational and counterfactual outcome distributions are tightly linked: they have identical support and tail behavior, remain statistically close under weak confounding, and share any features of high-dimensional outcomes which are invariant to confounders. These properties motivate learning counterfactual distributions not from scratch, but via a from the observational distribution. We formulate this problem via flow-matching and derive a semiparametrically efficient estimator based on a novel efficient influence function correction. We subsequently extend our estimator to target minimal-energy flows in high-dimensions, which we show can be especially simple targets between observational and counterfactual distributions. In experiments, deconfounding flows outperform existing debiased counterfactual distribution estimators, while also mitigating known failure modes of flow-based methods.
PaperID: 2633, Poster
Authors: Habibeh Naderi Khorshidi, Behrouz Soleimani, Stan Matwin
Abstract: Parameter-efficient fine-tuning (PEFT) enables efficient adaptation of frozen language models, but most PEFT methods remain deterministic or unimodal, limiting their reliability in low-resource audio-text settings where uncertainty depends on both linguistic evidence and acoustic conditions. We introduce SPECTRA (Stochastic Posterior Estimation with Cross-modal Token-level Rank Adaptation), a multimodal Bayesian low-rank adaptation framework for uncertainty-aware audio-text learning. SPECTRA keeps the text and audio backbones frozen and confines stochasticity to a compact rank-r latent matrix inside each LoRA adapter. At each transformer layer, text-derived low-rank token features query frame-level audio embeddings through lightweight cross-attention; the resulting token-specific acoustic context parameterizes the mean and variance of an amortized variational posterior over the adapter latent. This design treats audio not merely as an additional feature stream, but as a localized reliability signal that modulates both adaptation and confidence while preserving the scalability of PEFT. Posterior prediction is performed with Monte Carlo adapter samples, enabling a Bayesian uncertainty analysis that decomposes normalized predictive uncertainty into total, aleatoric, and adapter-space epistemic components and evaluates whether uncertainty distinguishes correct from misclassified predictions. Across IEMOCAP and clinical interview prediction tasks, SPECTRA is consistently competitive with or improves upon text-only Bayesian PEFT and conventional multimodal transfer-learning baselines, with token-level cross-attention yielding the most reliable gains. Additional modality-disagreement stress tests show that mismatched audio causes substantially larger degradation in AUC, likelihood, calibration, and Brier score than noisy but matched audio, highlighting the importance of modeling cross-modal reliability. These results suggest that Bayesian cross-modal conditioning in low-rank adapter space provides an efficient and principled mechanism for calibrated multimodal adaptation.
Abstract: Forward-backward (FB) representations provide a powerful framework for learning the successor representation (SR) in continuous spaces by enforcing a low-rank factorization. However, a fundamental spectral mismatch often exists between the high-rank transition dynamics of continuous environments and the low-rank bottleneck of the FB architecture, making accurate low-rank representation learning difficult. In this work, we analyze temporal abstraction as a mechanism to mitigate this mismatch. By characterizing the spectral properties of the transition operator, we show that temporal abstraction acts as a low-pass filter that suppresses high-frequency spectral components. This suppression reduces the effective rank of the induced SR while preserving a formal bound on the resulting value function error. Empirically, we show that this alignment is a key factor for stable FB learning, particularly at high discount factors where bootstrapping becomes error-prone. Our results identify temporal abstraction as a principled mechanism for shaping the spectral structure of the underlying MDP and enabling effective long-horizon representations in continuous control.
Authors: Christopher R Banerji, Enrico Parisini, Christopher J Soelistyo, Ahab Isaac, Alessandro Barp
Abstract: Aligning human-interpretable concepts with the internal representations of modern machine learning systems remains a central challenge for interpretable AI. We introduce a geometric framework to compare supervised human concepts with unsupervised representations derived from foundation models. We formalise concept frustration: a mismatch that arises when an unobserved concept induces relationships between known concepts that cannot be made consistent within an existing ontology. We develop task-aligned similarity measures that detect this phenomenon, and show that frustration is identifiable in task-aligned geometry while conventional Euclidean comparisons fail. Under a linear–Gaussian generative model, we derive a closed-form expression for Bayes-optimal concept-based classifier accuracy, decomposing predictive signal into known and unknown contributions and identifying where frustration impacts performance. Experiments on synthetic data and real language and vision tasks demonstrate that frustration is present in foundation model representations and that incorporating missing concepts reorganises learned representations to better align human and machine reasoning. These results provide a principled framework for diagnosing incomplete concept ontologies and improving alignment in interpretable AI systems.
PaperID: 2636, Poster
Abstract: Self-consistency (SC) is a widely used decoding strategy for improving chain-of-thought reasoning, but it requires sampling many complete reasoning trajectories with a preset budget. Its efficient variants reduce the cost mainly by stopping full-trajectory sampling early or distributing the number of samples across questions. In this paper, we propose StepBack, a step-level error-localized resampling method for efficient test-time reasoning. Given a small set of draft trajectories, \method localizes potential first error positions using lightweight signals and then resamples the detected error step and its suffix, followed by majority answer voting. Our theoretical analysis explains why resampling before the localized error point can increase the probability of recovering a correct final answer. Empirically, \method consistently matches or exceeds SC across mathematical reasoning and multi-hop QA benchmarks, while reducing token usage by up to 50%. In addition, by nature, \method can be combined with adaptive inference methods such as Adaptive Consistency and Early-Stopping SC, further improving their efficiency and scalability.
Abstract: Multimodal Large Language Models (MLLMs) often exhibit significant modality preference, which is a tendency to favor one modality over another. Depending on the input, they may over-rely on linguistic priors relative to visual evidence, or conversely over-attend to visually salient cues over textual facts. Prior work has applied a uniform steering intensity to adjust the modality preference of MLLMs. However, strong steering can impair standard inference and increase error rates, whereas weak steering is often ineffective. Since steering sensitivity varies substantially across instances, a single global strength is difficult to calibrate. To address this, we introduce an instance-aware diagnostic metric Modality Contribution Ratio (MCR) that quantifies each modality’s information contribution and reveals sample-specific susceptibility to steering. Building on this signal, we propose a scaling strategy and a learnable module for instance-aware control of modality preference. Experiments show that AMPS outperforms conventional steering on MC^2 by improving preference control with lower generation collapse, and improves MLLM performance on general multimodal QA dataset such as MM-Vet v2.
PaperID: 2638, Poster
Authors: Robert L Manschke, Angus Roberts, Julia Ive
Abstract: Realistic intensive-care EHR trajectories are difficult to simulate because physiology, treatments, observation times, and patient-existence state co-evolve under sparse clinical measurement. We introduce PULSE, a probabilistic simulator family for ICU trajectories. PULSE separates treatment rollout from patient-state rollout, predicts binary patient state before continuous physiology, and updates continuous variables through gated residual dynamics around the last observed or simulated state. Probabilistic variants add event-time inputs, heteroscedastic Gaussian heads, monotone B-spline-flow residual transforms, latent severity pooling, and covariance-coupled residual groups. On a MIMIC-IV v3.1 sepsis cohort scored over a 37-variable predictive-check panel, reference PULSE reduces k=6 vector-normalized RMSE from 7.124 for a matched per-covariate transformer Monte Carlo baseline to 5.969, and improves over carry-forward from 6.633 to 5.969; paired patient-level bootstrap intervals support both differences. RMSE is a useful sanity check for aggregate scale, but not a sufficient criterion for trajectory realism: carry-forward is already difficult to beat despite generating straight-line futures. Gaussian NLL is the strongest calibrated PULSE variant by CRPS, while B-spline flows address the non-Gaussian shape of factual residuals. Within B-spline flows, event-time encodings with Δt prediction improve k=6 rollout from 8.301 to 6.635 and yield sharper, more state-dependent residual densities. In k=6 density-mixture analysis, B-spline intervals have larger across-state 90% width variation than Gaussian intervals (coefficient of variation 0.79 versus 0.48) and higher 90/95% inclusion despite slightly weaker aggregate PIT uniformity, consistent with better local skew and sharpness in some states. Similar gains over carry-forward persist on eICU sepsis validation panels, while absorbing-state timing remains a major unresolved error source. CVSim known-effect experiments show that synthetic treatment response can be learned under mechanistic ground truth, although transfer to MIMIC-seeded counterfactual tests remains weak.
PaperID: 2639, Poster
Abstract: Burst imaging under unstable illumination suffers from flicker, a periodic degradation that modulates local brightness differently across adjacent frames, creating frame-specific flicker states. Existing methods rely on spatial priors or reference-centric fusion without explicitly modeling the evolving flicker state, whose temporal periodicity makes it predictable from adjacent observations by world models that capture latent state transitions. Motivated by this predictive capacity, we propose StateFlicker, a state-aware burst restoration framework that interprets adjacent frames as observations under evolving flicker states. StateFlicker comprises two state-aware designs: Cross-State Visibility (CSV) and Flicker-State Understanding (FSU). CSV leverages complementary visibility across adjacent flicker states to selectively recover content suppressed under the center state. With this cross-state support, FSU constructs dual state pathways through a frozen world model to obtain predicted and observed center states, and distills their state discrepancy into dense features that characterize the center-specific flicker condition for restoration. Extensive experiments on the burst flicker removal dataset show that StateFlicker achieves superior performance over state-of-the-art methods.
PaperID: 2640, Poster
Authors: Chang Liu, Jinqian Chen, Jihua Zhu, Xiangyang Yang
Abstract: Low-rank adaptation (LoRA) enables communication-efficient federated fine-tuning of large language models, but standard FedLoRA averages the factors A and B separately even though the model update is determined by their product \Delta W=sBA. Under task heterogeneity, factor-wise aggregation can further dilute directions that are locally useful but not strongly shared across clients; we call this phenomenon \emphdirectional dilution. We analyze this effect through the singular directions of client LoRA updates and find that low-strength directions are most vulnerable to dilution. Their repeated owner-specific recovery suggests client-specific adaptation signals rather than random noise. We propose \textscFedReCall, a client-private plug-in that captures observable dilution gaps, selects reliable diluted directions into a private frozen cache, and recalls them through a low-rank bypass guarded by a single-batch loss probe when the client next participates. As a plug-in, \textscFedReCall can be seamlessly integrated into existing FedLoRA methods without changing their aggregation rules, communication payloads, or trainable-parameter sets, consistently improving personalized adaptation and improving global evaluation metrics on most aggregators.
PaperID: 2641, Poster
Authors:
Jixuan Ruan, Zhuo Cui, Zhengding Hu, Zhongkai Yu, Xiang Fang, Yue Guan, Jason Ludmir, Phillip Weinberg, Xiu-Zhe Luo, Sheng-Tao Wang, Sonia L Alarcon, Hezi Zhang, Yufei DingAbstract: Large Language Models (LLMs) offer an adaptable alternative to traditional quantum compilation heuristics, which bottleneck progress by demanding costly manual redesigns whenever hardware evolves. However, naive single-pass LLM approaches fail under the unique challenges introduced by quantum compilation. We introduce Qubrio, a practical agentic framework that significantly reduces heuristic redesigns through a three-level decomposition. By (1) partitioning programs into sequential operation stages, (2) deploying specialized agents for placement, routing and optimization, and (3) using disaggregated feedback for precise error correction, our system reliably navigates complex physical constraints. Beyond serving as a direct compiler, our LLM framework also uncovers novel strategies that facilitate convoy-style shuttling, substantially enhancing hardware concurrency without sacrificing fidelity. We further integrate these strategies into the previous state-of-the-art (SOTA) compiler, recovering a significant fraction of the LLM's performance advantages. On realistic workloads, our compiler achieves significant improvements in hardware runtime and program fidelity over the SOTA baseline PowerMove, seamlessly adapting to new hardware capabilities. Qubrio features an open user interface for practical use, with its anonymized repository available for double-blind review at https://anonymous.4open.science/r/Qubrio-486C.
Abstract: We study trajectory selection for reasoning distillation, where teacher-generated reasoning trajectories are selectively used as supervision for a student model. Existing methods rely on heuristics such as trajectory quality or model confidence, but they often overlook whether a trajectory is learnable by the student. In this paper, we present LARK, a learnability-grounded method for reasoning trajectory selection. LARK selects trajectories that the student can learn efficiently while preserving the generalization of the full training distribution. At the core of LARK is a learnability factor \rho, which characterizes the rate at which the student's training loss decreases. To estimate this rate efficiently and maintain generalization, we introduce a learnability proxy and a \chi^2-regularized selection policy that balances learnability and distributional coverage, both with strong theoretical guarantees on their estimation error. Empirically, LARK consistently outperforms data selection baselines across multiple base models and reasoning tasks. Diagnostic analyses show that the LARK score predicts downstream training utility and that LARK-selected trajectories induce faster supervised fine-tuning loss reduction.
PaperID: 2643, Poster
Authors:
Théodor Lemerle, Diego A Torres Guarin, Téo Guichoux, Nicolas Obin, Axel RoebelAbstract: Existing text-to-speech (TTS) systems achieve high quality, yet streaming and long-form generation remain challenging. We propose simple adaptations to a standard encoder–decoder TTS architecture that enables both capabilities without limits on total duration. We first observe that attention patterns in these models are well organized, following a structure that motivates a specific windowing strategy: a prefixed sliding window for autoregressive decoding and a sliding cross-attention for text conditioning. Together, we show that these adaptations address length generalization and error accumulation while being fully streamable in text and audio. Trained exclusively on segments shorter than 30 seconds, our system produces seamless, consistent speech at practically unbounded lengths with linear decoding complexity, outperforming baselines that rely on hour-long training context. We validate the approach through continuous synthesis over hours long synthesis and demonstrate competitive results against state-of-the-art open-source baselines on both short- and long-form benchmarks.
PaperID: 2644, Poster
Authors:
Zhixiong Nan, Haoyu Lu, Tao Xiang, Yiwei WangAbstract: Black-box watermarked LLM identification has become an important task for watermark auditing. The core challenge arises from inherent black-box constraints that deny access to logits, detector keys, model parameters, and internal watermark settings. For this reason, existing methods have not conducted in-depth exploration on this task, leaving prominent limitations: . First, existing works use fixed null reference to serve as a baseline for judging the presence of watermarks in LLM outputs, which might misidentify a watermarked LLM as unwatermarked because the null reference is sensitive to the prompt pair, queried model, and sampling configuration. Second, existing statistical-tests based methods adopt distinct statistical features and judgment criteria for different LLM watermark families, making them less suitable when the deployed watermark algorithm is unknown. To remedy the above two limitations, this work proposes , a Self-Calibrated Trident Identification framework for black-box watermarked LLM identification. Specifically, to handle the first limitation, constructs empirical null distributions from the queried responses to estimate the generalizable reference instead of relying on a fixed reference. Meanwhile, to address the second limitation, is configured with a Trident-View Consistency measuring mechanism, which enables a unified pipeline to avoid designing separate tests for individual watermark family. Experimental results verify that outperforms representative black-box watermark identification baselines across diverse LLMs and watermark algorithms.
Abstract: Vision-language-action models (VLA) promise human-steerable autonomous driving, but their training is bottlenecked by the scarcity of frames paired with natural-language instructions: while camera streams and expert trajectories are logged at scale, language annotations (e.g., "turn left at the intersection") remain scarce and expensive to acquire. To address this challenge, we introduce Latent Action Driving Annotations (LADA), a three-stage pipeline that transforms abundant unlabelled observation–trajectory pairs into a substrate for language-conditioned control. First, we train a latent action model with a vector-quantised bottleneck, producing a compact codebook of high-level vehicle intents. Second, a small language-annotated subset is used to train a vision-language translator to map observations and language instructions into this codebook. Third, we train a driving VLA on observation–latent-action pairs over the full unlabelled corpus. Using fewer than 5% of language annotations and without leveraging any auxiliary chain-of-thought reasoning or visual question answering streams, LADA achieves a Driving Score of 87.98 and a Success Rate of 70.46% on the closed-loop Bench2Drive benchmark, matching or surpassing fully supervised baselines.
PaperID: 2646, Poster
Abstract: The ability to learn a new task from a few examples has been crucial to the success of foundation models such as large language models (LLMs). However, current foundation models for structured data like tables and relational databases still require tens of thousands of labeled examples to perform well on new tasks. Here we show that when pretrained at scale with the right recipe, Relational Transformers (RTs) can make state-of-the-art (SoTA) predictions with only hundreds of labels. This capability arises from the following ingredients: First, we assemble THE JOIN, the largest pretraining corpus of relational data to date, comprising 6k forecasting tasks across 650 real-world databases from diverse domains. Second, we develop a pretraining recipe that combines (1) mixed context sizes for multi-scale learning, (2) multi-cell masking for dense supervision, and (3) a novel random- walk-based retriever to efficiently gather relevant context. Third, we identify context ensembling and context tuning as complementary axes for scaling test-time compute to further improve few-shot performance. Pretrained with our recipe on THE JOIN, RT achieves parity with LLM Agent + TabICLv2 and RDBLearn + TabICLv2 pipelines with 32–23×fewer labels respectively, and even surpasses the prior SoTA of full task-specific training, on average nMAE for RelBench regression tasks. Context ensembling and tuning improves this further by up to 3%, and 4% respectively. Our ablations highlight the importance of schema semantics, multi-cell masking, and random-walk retrieval. Overall, our work paves the way for developing relational foundation models with strong few-shot capabilities.
PaperID: 2647, Poster
Abstract: Continual Reinforcement Learning (CRL) agents struggle to adapt to new situations while retaining prior knowledge, leading to a stability–plasticity trade-off. One approach to solving this problem is by using complementary learning systems: one for slow, long-term learning and another for quick, transient adaptation. Anand & Precup (2023) instantiated this idea using permanent and transient value functions, but used the same features in both systems. Intuitively, long-term learning should rely more on parametric representations for broad generalization, while transient learning would be best served by non-parametric representations for situation-specific adaptation. In this paper, we explore this idea both theoretically and empirically. We propose a novel non-parametric approximation for estimating the transient value function in complex tasks. And, demonstrate that our method enables online learning and outperforms competitive baselines on image-based tasks and Craftax-Classic, underscoring the effectiveness of our system-level decomposition.
PaperID: 2648, Poster
Abstract: Self-training is a practical way to adapt CLIP with limited supervision, but its performance is highly sensitive to pseudo-label reliability. We identify an important failure mode in CLIP adaptation: temporal prediction instability, where a target sample repeatedly flips its prediction across epochs. We provide a perspective based on cross-modal anchoring mismatch: the decision geometry induced by pretrained text anchors can deviate from the local visual cluster structure of the target domain. This mismatch can make ambiguous target samples susceptible to prediction flips during adaptation, whereas target-aligned anchoring produces more stable pseudo-labels. Motivated by this observation, we propose Stability-Aware Self-Training (SAST), centered on Stability-Weighted Image Prototypes (SWIP). SWIP builds class-wise visual anchors from temporally stable target samples to provide more reliable adaptation references. We further add a lightweight text-side refinement to reduce confusion-induced instability, and combine the two streams with an adaptive fusion rule. Across six standard benchmarks and three adaptation settings, SAST consistently outperforms strong baselines. The results suggest that explicitly modeling temporal stability is an effective route to more robust CLIP adaptation under self-training.
Abstract: Modern large language models (LLMs) operate in interactive multi-turn settings, making multi-turn jailbreaking a realistic threat model and an important scenario for automated red teaming. A core challenge in learning multi-turn jailbreak attackers is credit assignment: different turns contribute differently to the final outcome, but existing turn-level credit signals for jailbreaking are often too coarse for multi-turn interaction. We propose a unified turn-level credit assignment framework for Group Relative Policy Optimization (GRPO) in multi-turn jailbreak learning. Our method assigns group-relative learning signals directly at each turn by decomposing immediate and future credits. We call this approach decomposed credit GRPO (DC-GRPO). In contrast, existing methods often suffer from credit misassignment, for example by applying a single trajectory-level score to the entire dialogue. Across multiple victim LLMs and benchmarks, our dynamic-weighted DC-GRPO achieves an average ASR@3 of 98.26% for 5-turn jailbreaking, outperforming existing state-of-the-art methods such as SEMA, which achieves 86.58%, and TROJail, which achieves 86.23%. These results highlight turn-level group-relative credit assignment as a simple and effective ingredient for scalable multi-turn automated red teaming.
Authors:
Cecilia Ferrando, Miguel Fuentes, Brett Mullins, Cameron Musco, Daniel SheldonAbstract: We propose PACE-GGM, a data-adaptive differentially private method for covariance estimation that concentrates its privacy budget on the most informative entries of the empirical covariance matrix, rather than perturbing all entries. This applies in the natural setting where the modeler supplies separate bounds for each variable, so that individual entries can be measured with less noise than the full matrix. In each round, our method selects a poorly approximated entry, measures it using the Gaussian mechanism, and then reconstructs a full covariance matrix using a maximum-entropy reconstruction objective, leading to a Gaussian graphical model structure. Experiments on diverse real-world datasets demonstrate consistent improvements in estimation error with respect to the Gaussian mechanism and other baselines, particularly in high-dimensional and low-to-moderate privacy regimes.
Abstract: Class-Incremental Learning (CIL) enables models to learn new classes continually while preserving past knowledge. Recently, vision-language models like CLIP offer transferable features via multi-modal pre-training, making them well-suited for CIL. However, real-world visual and linguistic concepts are inherently hierarchical: a textual concept like ''dog'' subsumes fine-grained categories such as ''Labrador'' and ''Golden Retriever,'' and each category entails its images. But existing CLIP-based CIL methods fail to explicitly capture this inherent hierarchy, leading to fine-grained class features drift during incremental updates and ultimately to catastrophic forgetting. To address this challenge, we propose HASTEN (Hierarchical Semantic Tree Anchoring), a hierarchy-aware framework that uses semantic structure to stabilize CLIP-based CIL. Rather than treating classes as isolated labels, HASTEN leverages external hierarchical knowledge as structured supervision to organize visual and textual features in hyperbolic space, helping maintain parent-child relations as new tasks arrive and mitigating feature drift. Since the shared hyperbolic mapper is updated across tasks, we further stabilize it by constraining its updates to a null space induced by prior-task features, reducing interference with previous mappings while retaining adaptability to new classes. Extensive experiments on nine benchmarks show that HASTEN consistently outperforms existing methods while reducing feature drift and catastrophic forgetting.
Abstract: Recent advancements in Large Language Model (LLM) agents have demonstrated strong capabilities in executing complex tasks through tool use. However, long-horizon multi-step tool planning is challenging, because the exploration space suffers from a combinatorial explosion. In this scenario, even when a correct tool-use path is found, it is usually considered an immediate reward for current training, which would not provide any reusable information for subsequent training. In this paper, we argue that historically successful trajectories contain reusable tool-transition patterns, which can be leveraged throughout the whole training process. Inspired by ant colony optimization where historically successful paths can be reflected by the pheromone, we propose Pheromone-Guided Policy Optimization (PhGPO), which learns a trajectory-based transition pattern (i.e., pheromone) from historical trajectories and then uses the learned pheromone to guide policy optimization. This learned pheromone provides explicit and reusable guidance that steers policy optimization toward historically successful tool transitions, thereby improving long-horizon tool planning. Comprehensive experimental results demonstrate the effectiveness of our proposed PhGPO.
PaperID: 2653, Poster
Authors:
Yue Liu, Ben Liang, Ali Tizghadam, Ilijc AlbaneseAbstract: We study supervised classification problems with missing feature values. Existing approaches often decouple imputation from classification, producing imputations that may be plausible but uninformative for prediction. Instead, we propose the Integrated Imputation and Classification Network (IICN), which jointly trains an imputer and an (n+1)-classdiscriminator adversarially with a single class supervised classification objective, where the discriminator learns to distinguish among the n true classes and an additional ``imputed" class. We prove that at the global optimum, the imputer and discriminator together implement marginalization over missing coordinates and yield a Bayes-optimal classifier. We evaluate IICN on FashionMNIST, CIFAR-10, and tabular datasets with naturally occurring missingness. IICN outperforms classical impute-then-classify pipelines and recent generative baselines, showing strong robustness and accuracy in challenging settings.
Abstract: Gaussian Process (GP) kernels are central to Bayesian optimization (BO), yet designing effective kernels for high-dimensional problems still relies on extensive manual engineering. Existing kernel design methods restrict the search to additive and multiplicative compositions of base kernels, while LLM-based BO approaches condition on raw observations, which are infeasible in high dimensions due to context-length constraints and the difficulty of extracting meaningful patterns from high-dimensional numerical observations. We introduce , an LLM-driven population-based evolutionary framework that overcomes both limitations. Motivated by the observation that directly prompting an LLM to generate kernel code yields syntactically varied but functionally identical kernels, we adopt a two-stage approach: an LLM first proposes novel mathematical forms, then a second LLM call converts each form into validated, executable code. We also propose a leave-one-out continuous ranked probability score (LOO-CRPS) as a held-out predictive scoring criterion that penalizes overconfident fits more directly than marginal log-likelihood. On five standard high-dimensional BO benchmarks, our method achieves an average rank of , consistently outperforming competitive baselines. We further analyze the discovered kernels to examine which kernel characteristics lead to superior performance in high-dimensional BO.
Authors:
Wei Chen, Guanghui Zhu, Yafei Li, Limin Wang, Yihua HuangAbstract: Reinforcement learning from human feedback (RLHF) with preference-based reward models often exhibits unstable training dynamics. A key contributing factor is that standard RLHF relies on a single sequence-level scalar reward, which is propagated to token-level policy updates and leaves credit assignment within a response inherently ambiguous. Recent work has attempted to address this issue by refining rewards into denser token-level supervision, often relying on the implicit assumption that finer-grained credit assignment improves optimization. We argue that this assumption is incomplete: when preference signals are noisy and only defined at the response level, overly fine-grained reward refinement can amplify reward uncertainty and destabilize learning. To address this problem, we propose a granularity-aware principle for hierarchical credit assignment, emphasizing stability-oriented reward design rather than maximal allocation precision. Under this principle, sentences serve as a natural intermediate granularity, balancing semantic coherence with robustness to token-level noise. Guided by this view, we introduce S2T-RLHF. This sentence-to-token reward decomposition framework first allocates sequence-level preference rewards across sentences and then applies bounded token-level refinement within each sentence, without reward-model retraining or token-level supervision. Experiments across multiple datasets and optimization settings show that S2T-RLHF improves training stability and robustness while maintaining competitive preference alignment.
PaperID: 2656, Poster
Abstract: In this work, we consider decision-focused learning (DFL) for a Markov decision process (MDP), where existing methods differentiate through the KKT conditions of the Bellman equation and require solving a linear system over all state-action pairs, limiting its scalability. We address this by reformulating the MDP as an occupancy measure-based linear program (LP), whose feasible region is induced by predicted dynamics, and we derive a closed-form gradient by identifying the active constraints in the feasible polyhedron via the pivoting algorithm. This occupancy measure-based LP layer raises two challenges: (1) LP's solution gradient is discontinuous when active constraints change, and (2) the LP backward cost still scales with the state size, which is costly for large or continuous state spaces. We address the challenges with an augmented Lagrangian surrogate and smooth the boundary jumps by random row sketching of the constraints, and a learnable soft state-aggregation layer and its function-approximation generalization that scales the LP to large finite and continuous-state MDPs. Across multiple tasks, our methods reach lower regret than KKT-based DFL and two-stage baselines with significantly lower computation cost. The source code for all experiments is available
PaperID: 2657, Poster
Abstract: Latent diffusion models depend critically on the geometry of the space in which generation is performed. A persistent challenge in VAE- and RAE-based systems is that detail-rich latents improve inversion but can make diffusion dynamics harder to model. We study this reconstruction--generation tension through two geometric properties: invertibility, which measures how faithfully images can be decoded, and navigability, which measures how smoothly diffusion can transport mass in latent space. We introduce a tangent-projected transport complexity that isolates score-field variation along data-manifold directions and highlights intrinsic dimension, curvature, reach, and enrichment direction as geometry-side factors in diffusion difficulty. This analysis yields the Normal Projection Principle: when enriching a navigable base latent, added detail should lie in the normal space so that the base coordinate is preserved under tubular projection, whereas tangential residuals shift that coordinate at first order. Motivated by this principle, we introduce NormFuse, an RAE-based text-to-image model that fuses frozen multi-layer DINOv3 features through normal-space projection. Controlled and large-scale matched comparisons show that NormFuse preserves reconstruction benefits while avoiding the generation degradation of unconstrained fusion, supporting latent geometry as both an explanatory lens and a practical design rule.
PaperID: 2658, Poster
Authors: Shiting Wang, Yingjie Zou, Ke Li
Abstract: Model merging combines multiple fine-tuned models without additional training, and arithmetic methods have emerged as a simple yet effective way to merge models without data. However, existing arithmetic methods treat task vectors in parameter space and apply a single element-wise rule to directions that can differ in mergeability. In a shared directional basis, we find a consistent two-regime structure: a low-dispersion common bulk shared across tasks and a high-dispersion task-dominated tail that carries disproportionate energy. These findings motivate Common-Unique Subspace Decomposition (CUSD), a data-free pre-merging method that separates task vectors into common and unique directional components, merges the common components with downstream arithmetic operator, and retains unique residuals separately. Experiments across 11 base models, four model families, and four arithmetic-based merging operators show that CUSD improves 42 of 44 model-operator combinations and achieves up to 22.5% improvement over the corresponding baselines.
PaperID: 2659, Poster
Authors: Hyeonsu Lyu, Jonggyu Jang, Sehyun Ryu, Hyun Jong Yang
Abstract: Influence functions (IFs) approximate the effect of removing or reweighting training data on a trained model, but their linear approximation often lacks the accuracy required for reliable post hoc model editing. We identify two culprits and propose generalized influence functions (GIFs) and pseudo-LiSSA (p-LiSSA) to address each in turn. First, classical IFs update all parameters indiscriminately, leading to unnecessary updates to data-irrelevant parameters. GIFs instead restrict edits to a data-relevant subspace and leave the remaining parameters fixed. We theoretically show that GIFs achieve a tighter error bound than classical IFs, and that proper subspace selection requires avoiding low-curvature directions while maintaining strong alignment with the target gradient. Second, LiSSA iterations are destabilized by negative Hessian eigenvalues, which existing methods handle via \ell_2 damping---at the cost of a biased inverse-Hessian-vector product (iHVP) estimate. p-LiSSA eliminates this bias by solving the restricted inverse-Hessian problem without such damping. We further show that the presence of low-curvature directions slows p-LiSSA convergence, and that the data-relevant subspace selected by GIFs mitigates this. Across five image and text unlearning benchmarks, GIFs update only 10% of the parameters yet preserve accuracy up to 2.12 percentage points better than IF-based baselines, while improving unlearning performance by up to 2.73 percentage points. Moreover, p-LiSSA improves restricted iHVP estimation by reducing the residual by 9.6× over existing iHVP baselines, while using 19.8% less memory with comparable computation time.
Abstract: Empirical evidence suggests that large neural networks rarely make effective use of all available parameters, with learned solutions exhibiting pronounced sparsity. Despite this ubiquity, existing generalization theory only partially explains how such sparsity influences statistical performance. In this paper, we derive non-asymptotic excess risk bounds for deep neural networks whose layer-wise weight matrices are restricted by \ell_p quasi-norm (0
PaperID: 2661, Poster
Abstract: Robot manipulation often depends on short-term cues—such as brief contacts, temporary occlusions, and intermediate object states—that are ambiguous in a single frame and can only be resolved by integrating observations over time. However, injecting historical information solely by enlarging the history window is computationally costly and frequently yields limited gains due to substantial redundancy and task-agnostic, inflexible sampling of observations. We propose CondenseVLA, a multi-frame VLA policy with a lightweight learnable history condensation module for using short-term visual context in action decoding. Given a sequence of cached per-frame features from the VLM backbone, CondenseVLA distills observation history into a fixed set of slot tokens using learnable queries and multi-head cross-attention. Simple auxiliary regularization encourages diverse attention patterns by penalizing overlap among slot attention maps and reduces slot 12 redundancy by decorrelating slot embeddings, stabilizing learned condensation without imposing strong time priors. By expanding the history window and distilling it into a compact token set, CondenseVLA improves over the baseline on the evaluated Simpler-WidowX tasks, increasing success rate from 58.4 to 70.8 (+12.4), achieves 97.2% on LIBERO, and reaches 52.0% on real-world tasks. Beyond the gains, our results show that long history windows contain substantial redundancy, so naively providing more frames yields limited benefits, whereas learned condensation into fixed slot tokens provides an effective bounded-token way to leverage temporal context.
Authors:
Yi Pan, Miao Pan, Qi Lu, Jiaming Huang, Man Zhang, Siteng Huang, Xin Li, Jie Zhang, Yongliang Shen, Xuhong Zhang, Wenqi ZhangAbstract: Vision-Language-Action (VLA) foundation models have recently achieved strong progress in embodied intelligence. To reduce policy-call frequency while preserving temporal coherence, most generative policies adopt an action chunk mechanism, executing multiple future actions in an open-loop manner under a fixed action horizon. However, this ``predict-then-blindly-execute'' paradigm sacrifices closed-loop reactivity: in contact-rich physical interactions, even small local perturbations can rapidly amplify within the open-loop blind spot, leading to compounding errors and ultimately task failure. To address this limitation, we propose VLA-Corrector, a lightweight corrective inference framework for action-chunked VLA policies. Without modifying the backbone policy weights, VLA-Corrector introduces a lightweight Latent-space Vision Monitor (LVM) that continuously compares predicted and actual visual feature evolution, enabling online detection of visual dynamics deviations. Once persistent deviation is detected, the system triggers a truncation event, discards the remaining stale actions, and invokes corrective replanning via Online Gradient Guidance (OGG). The detect-and-correct mechanism of VLA-Corrector naturally induces an event-triggered adaptive action horizon: it preserves long-horizon execution when the current chunk remains reliable, and invokes short-horizon corrective replanning when execution begins to drift. In doing so, VLA-Corrector mitigates the trade-off imposed by static horizons between execution robustness and policy-call frequency. It can be integrated into different VLA models without further retraining the VLA backbone, interrupting compounding errors while preserving much of the efficiency benefit of action chunking and substantially improving robustness in long-horizon, contact-rich robotic manipulation tasks.
PaperID: 2663, Poster
Abstract: Transfer learning is a powerful approach for improving classification performance when the target dataset contains limited observations. However, individual-level source data are often unavailable due to privacy or access constraints. We propose GSURE-Trans, an uncertainty-aware summary-level transfer learning framework for logistic classification that uses target individual-level data together with only source coefficient estimates and standard errors. By introducing and profiling out an auxiliary bridge parameter, GSURE-Trans induces a coordinatewise Huber penalty that accounts for both source uncertainty and source-target coefficient discrepancies. To further prevent negative transfer at study level, we develop a new criterion, generalized Stein unbiased risk estimate (GSURE), that adaptively selects the scale of source borrowing. Theoretically, we establish uniform consistency of GSURE and an oracle inequality for the selected transfer weight. Simulations and real data application show that GSURE-Trans improves prediction when sources are compatible and down-weights heterogeneous sources when transfer is harmful.
Abstract: Despite recent progress, existing general robot policies, particularly Vision-Language-Action models, are still primarily trained to map observations directly to actions using sparse action supervision. As a result, they often learn shallow observation-to-action correlations instead of deeper representations of world dynamics, including object motion, physical interaction, and long-horizon task progression. Recent world-action models attempt to address this limitation through future-frame prediction, but pixel-space rollout is computationally expensive and poorly aligned with the abstractions required for control, forcing the model to reconstruct visual details irrelevant to action. We present Being-H0.7, a latent world-action model that introduces future-aware reasoning into VLA-style policies without generating future frames. Our model inserts learnable latent queries between perception and action, forming an explicit reasoning interface for action generation. To train this interface, we adopt a future-informed dual-branch design: a deployable prior branch infers latent states from the current context, while a training-only posterior branch replaces latent queries with embeddings from future frames. Joint alignment between the two branches enables the prior branch to learn predictive, action-relevant latent structure directly from current observations. Under a controlled training setup for comprehensive validation, Being-H0.7 achieves state-of-the-art or competitive performance across diverse benchmarks, while further demonstrating strong deployability on real-robot manipulation tasks.
Abstract: Kernel methods are widely used in machine learning and statistics for their flexibility and expressive power, yet their black-box nature limits adoption in high-stakes applications. Shapley value--based attribution methods such as SHAP, and kernel-specific adaptations including RKHS-SHAP, provide a principled framework for explainability---but exact computation of Shapley values is generally intractable, forcing existing approaches to rely on approximations that incur unavoidable estimation error. We introduce PKeX-Shapley, an algorithm that exploits the multiplicative structure of product kernels to compute exact Shapley values for all d features in quadratic time in d. The method rests on a distribution-free removal operator intrinsic to the product-kernel structure: removing a feature replaces its kernel factor with the multiplicative identity. This yields a parameter-free value function---requiring no sampling and no density estimation---and uniquely determines a functional decomposition of the model. Building on this value function, we develop shared recursive formulations that evaluate all feature attributions jointly, achieving amortized linear time per feature with numerical stability. Beyond predictive modeling, the framework extends to widely used kernel-based discrepancies such as the Maximum Mean Discrepancy (MMD) and the Hilbert--Schmidt Independence Criterion (HSIC), providing new tools for interpretable statistical analysis.
PaperID: 2666, Poster
Abstract: Neural combinatorial optimization methods for the Job Shop Scheduling Problem (JSSP) face a fundamental trade-off between solution quality and inference efficiency: constructive methods build high-quality schedules through sequential decisions whose computational cost scales with problem size, whereas one-shot methods offer highly efficient single-pass inference at the expense of solution quality. We propose SchedDiff, the first diffusion-based framework for JSSP, which bridges this gap by iteratively refining operation priorities through a fixed number of denoising steps. Unlike existing diffusion models for combinatorial optimization that operate in a binary space, SchedDiff operates directly in a discrete ordinal space, introducing a forward process and learning objective specifically tailored to ranking structures. To decode the refined priorities into feasible schedules, we introduce active list scheduling, which provably produces more compact schedules than standard list scheduling. Extensive experiments on TA, DMU, and LA benchmarks demonstrate that SchedDiff achieves state-of-the-art performance among learning-based methods, with efficiency advantages that become more pronounced on large-scale instances. SchedDiff further exhibits strong generalization, scaling to instances significantly larger than those seen during training, generalizing across unseen problem distributions, and transferring zero-shot to flexible JSSP.
PaperID: 2667, Poster
Abstract: State-of-the-art neural network sparsification relies on alternating dense and sparse training phases. While computationally expensive, dense phases are critical for performance. Our empirical analysis reveals the mechanics behind this gain: upon reintroduction, previously masked parameters receive disproportionately large gradients that drive essential mask rewiring and parameter sign-flipping. Building on these insights, we optimize the performance-to-FLOPs Pareto frontier through three innovations: (1) zero-initializing masked parameters during dense phases to promote relevant sign flips, (2) truncating dense phases based on early mask convergence, and (3) replacing the dense phase entirely with a sparse rewiring phase that updates only the active mask and a small, high-gradient fraction of out-of-mask weights. Evaluated on CIFAR100 and ImageNet, our variant ARC (Alternating Rewiring Compression) matches baseline AC/DC performance at a fraction of the computational cost, establishing a superior performance-to-FLOPs frontier also compared to standard sparse-to-sparse algorithms.
Abstract: In this work, we develop proximal preconditioned gradient methods with a focus on spectral gradient methods providing a proximal extension to the Muon and Scion optimizers. We introduce a family of stochastic algorithms that can handle a wide variety of convex and nonconvex constraints and study its convergence under heavy-tailed noise, through a novel analysis tailored to the geometry of the proposed methods. We further propose a variance-reduced version, which achieves faster convergence under standard noise assumptions. Finally, we show that the polynomial iterations used in Muon are more accurately captured by a nonlinear preconditioner than by the ideal matrix sign, leading to a convergence analysis that more faithfully reflects practical implementations.
PaperID: 2669, Poster
Abstract: Vision-Language Pre-trained (VLP) models such as CLIP learn strong representations from large-scale image–text pairs and demonstrate impressive zero-shot transfer. However, by learning to align visual and textual features only at a global level, they often suffer from limited interpretability and weak fine-grained perception. This issue stems from pre-training data that overlooks key visual details of object parts, which are often important for distinguishing between subordinate categories (e.g., species of birds or models of cars). Multimodal Large Language Models (MLLMs) built on CLIP-style vision encoders inherit this weakness, limiting both accuracy and trustworthiness. To address this, we propose Part-Aware CLIP ( ), a framework that improves fine-grained perception while enhancing interpretability. First, we leverage MLLMs to construct a new dataset, , containing about one million part-level image–text pairs that explicitly describe discriminative components (e.g., beak shape or wing patterns). Second, we introduce a part-aware training strategy that encourages explicit grounding of fine-grained textual descriptions to corresponding image regions, strengthening part-level cross-modal alignment. Extensive experiments show that PA-CLIP achieves state-of-the-art results on multiple fine-grained visual recognition benchmarks, validating the benefit of part-level captions for capturing subtle details. Moreover, evaluations on general tasks such as cross-modal retrieval indicate that these improvements do not compromise the model's core generalist capabilities.
PaperID: 2670, Poster
Abstract: This paper addresses the challenge of obtaining strong optimality guarantees in constrained nonsmooth nonconvex optimization under mild regularity conditions, namely local Lipschitz continuity and existence and continuity of directional derivatives. While standard methods typically ensure weak stationarity notions, achieving directional (d-)stationarity remains nontrivial. We show that a random direction exploration step is sufficient to attain d-stationarity. The proposed approach augments any base optimization method with a single exploration step that samples a direction and step size and accepts the candidate based on a function value comparison. The resulting scheme guarantees that all accumulation points are d-stationary almost surely, independently of the behavior of the underlying method. Moreover, it preserves convergence rates of the base method, as established for DCA and prox-linear-type schemes. The theoretical results are complemented by numerical experiments illustrating the effect and guarantees of the exploration step.
PaperID: 2671, Poster
Authors:
Yuanlin Wen, Elan S Markowitz, ZUBO GU, Sami Abu-El-Haija, Bryan Perozzi, Michael Galkin, Ali Parviz, Aayush Gupta, Venkata Chivukula, Griffin Hadfield, Aravind Akella, Daniel J Frame, Vahab Mirrokni, Amin VahdatAbstract: The rapid scale-up of AI infrastructure has made datacenter reliability a cornerstone of the cloud business economy. Hardware failures in hyper-scale AI clusters do not merely disrupt individual machines; they stall multi-million dollar training workloads, and so the downtime of individual machines can cause outsized economic losses. Minimizing machine downtime, therefore, is more important than ever. To address this, we present DrWatson, an AI-driven hardware diagnosis system that pioneers a foundation-model approach to automated datacenter maintenance. Deployed at production scale within a major AI hyperscaler, DrWatson yields substantial efficiency gains across diverse compute architectures. Real-world evaluations demonstrate that it reduces production downtime by 18.3% on GPU platforms and 16.1% on custom AI accelerators. Furthermore, it significantly accelerates hardware deployment, cutting quality assurance (QA) time by 28.3% for GPUs and 25.1% for custom accelerators. These results establish DrWatson as a field-tested solution for maximizing the reliability and cost efficiency of next generation AI fleets.
Abstract: Average-reward reinforcement learning requires estimating the gain and the bias, which is defined only up to an additive constant. This makes direct distributional analogues ill-posed on the real line. We introduce a quotient-space formulation in which state-indexed bias laws are identified up to a common translation, together with a categorical parameterization that respects this symmetry. On this quotient-categorical space, we define a projected average-reward distributional operator and show that it is well-defined, non-expansive in a coordinate Cram\'er metric, and admits fixed points. We then study sampled recursions whose mean-field maps are asynchronous relaxations of this operator. In an idealized centered-reward setting, a one-state temporal-difference update enjoys almost sure convergence together with finite-iteration residual bounds under both i.i.d. and Markovian sampling. When the gain is unknown, we augment the recursion with an online gain estimator, and prove non-expansiveness and Markovian convergence of the resulting coupled scheme. Finally, we show that synchronous exact updates are gain-independent at the quotient-law level, isolating a structural contrast between ideal quotient distributions and practical fixed-grid categorical representations.
PaperID: 2673, Poster
Abstract: We study the stability of data exchange among competing firms. Unlike tangible goods, data is non-rival: sharing it does not deprive the giver of its use. On one hand, competing firms-- often operating as machine-learning–driven predictors that collect observations from a shared environment to solve similar tasks-- can improve their predictive accuracies from pooling data. On the other hand, they compete for market share, which may depend on their relative predictive performance. This raises the central question of this paper: which data exchanges are stable in such competitive environments? We formalize this via a model of n firms exchanging fractions of their datasets, with predictive performance feeding into market competition. We adopt core stability as our solution concept, accounting for externalities via worst-case responses by non-participating firms. Our main results reveal a sharp contrast driven by demand structure: (i) in fully rational markets, where customers flock entirely to the top-performing firm, a core-stable exchange always exists for any monotone continuous prediction and utility functions; (ii) in partially rational markets with logit-style demand, core-stable exchanges may fail to exist, and even approximate cores can be empty. Our impossibility results suggest that internal profit-sharing within coalitions can restore stability in partially rational markets. We formalize this intuition by showing that when co-operating coalitions equally redistribute market-share revenues among members, a core-stable data exchange exists. Finally, we propose a natural exchange dynamic and demonstrate empirically that it converges to a core-stable data exchange in fully rational markets on simulated instances derived from real datasets.
PaperID: 2674, Poster
Abstract: Tabular Foundation Models (TabFMs) produce context-aware embeddings that are increasingly stored, shared, and reused for retrieval, clustering, and cross-institutional data exchange. These embeddings are often treated as privacy-preserving substitutes for raw tabular records, yet their attribute-level information content remains poorly understood. We present a systematic privacy audit of TabFM embeddings: given an embedding and varying amounts of side information, how accurately can sensitive attributes be recovered? We introduce , a sequential probing method that measures recoverability while accounting for inter-attribute dependencies, significantly outperforming joint probing and optimization-based inversion baselines. Across four TabFMs (TabPFN, TabDPT, Mitra, TabICL) and eight datasets, we find that even without any side information, 65-95% of sensitive attributes can be recovered from embeddings alone. More critically, recoverability exhibits a : revealing a single known attribute sharply increases the recoverability of all remaining attributes. This phenomenon generalizes across model architectures and diverse domains, and is not eliminated by standard embedding transformations including noise injection, PCA, and knowledge distillation. These findings suggest that we need more formal privacy techniques for protecting sensitive information inside TabFM embeddings.
Abstract: Test-Time Scaling (TTS) improves large language models (LLMs) by allocating additional computation during inference, typically through parallel, sequential, or hybrid scaling. However, prior studies often assume fixed collaboration architectures (e.g., topologies) and single-model usage, overlooking that optimal architectures and model combinations can vary across tasks. Therefore, we study the novel problem of searching for compute-optimal model combinations and architectures in TTS under a fixed budget. We formalize it as a multi-LLM collaboration graph, where nodes encode roles, and LLM model assignments, and edges capture information flow. This problem is challenging because (i) the combinatorial search space is prohibitively large, and (ii) task-specific requirements demand tailored designs. To this end, we formulate the problem as probabilistic graph optimization and, via pilot experiments, derive three empirical insights into TTS collaboration graphs. Guided by these insights, we propose Agent-REINFORCE, an LLM-agent-augmented framework that mirrors the REINFORCE pipeline by mapping \emphsampling–gradient–update to \emphsampling–feedback–update, where feedback serves as a textual gradient to update the probabilistic graph and efficiently search for optimal multi-LLM collaboration graphs. Experiments show that Agent-REINFORCE outperforms both traditional and LLM-based baselines in sample efficiency and search performance, and effectively identifies optimal graphs under joint objectives of accuracy and inference latency.
Abstract: As autonomous agents increasingly execute end-to-end tasks under fixed monetary budgets, the pressing open question shifts from whether the budget is respected, to how to spend it effectively. Existing budget-aware methods typically control reasoning step-by-step within a single agent, or learn resource allocation policies via RL. None address how to split a budget across the composing phases of a multi-agent pipeline at inference time. We propose ZEBRA, a zero-shot framework that reduces multi-phase budget allocation to a continuous nonlinear knapsack problem: an LLM controller estimates per-phase utility curves, and a water-filling search on the Lagrange multiplier returns the per-phase split. Additive and multiplicative aggregations are unified under the same solver. On a 150-task APPS coding benchmark, both ZEBRA variants outperform LLM-direct (budget allocation directly by an LLM) on every aggregate metric. At a budget of \alpha = 0.5 of the unconstrained spend, ZEBRA recovers 94.4% of unconstrained quality, versus 88.1% for LLM-direct. The advantage is statistically significant and transfers beyond coding: on a 3-phase HotpotQA pipeline, ZEBRA beats LLM-direct by 14.3pp, with allocations empirically robust to curve-estimation noise. Also, ZEBRA arrives at a different budget split (near-balanced) compared to the APPS one (skewed towards a refinement phase), showing adaptation to the pipeline structure. More broadly, we show that lightweight algorithmic guidance at inference time can improve the economic behavior of autonomous multi-agent systems.
PaperID: 2677, Poster
Abstract: Post-training quantization (PTQ) to FP4 has emerged as a key technique for reducing large language model inference cost, especially on Blackwell GPUs which offer native FP4 tensor core support. Direct PTQ to FP4, however, still leaves a noticeable accuracy gap. Recent work on SVDQuant mitigates the errors introduced through FP4 quantization for diffusion models by adding high-precision low-rank corrections obtained via SVD. However, this high-precision low-rank correction requires many bits per rank, and moreover competes on tensor core utilization with the main FP4 GEMM. We introduce BiLoCo, a correction whose left and right factors are stored as binary \\pm 1\ sign vectors rather than as BF16 numbers. The signed format uses 16× less storage per rank-one term, and since the correction now consists of additions/subtractions, this can be run on CUDA cores concurrently with the FP4 tensor-core GEMM. On 4--32B LLMs, BiLoCo matches or improves memory-matched SVDQuant for PTQ. We implement BiLoCo on B200 with a correctness-preserving CUDA schedule that runs end-to-end decode at 1.13--1.15× FP4-only, with the residual gap coming from the final output combine rather than from one-bit arithmetic.
Abstract: Temporal action localization (TAL) requires recognizing the target event and localizing its start and end times precisely in untrimmed videos. Recent vision-language formulations improve semantic reasoning and support language-conditioned outputs, but their autoregressive decoders still generate tokens from left to right, preventing later semantic evidence from revising earlier timestamp predictions. We adapt masked diffusion vision-language models (MDVLMs) to TAL so that semantic tokens and boundary tokens remain editable throughout iterative denoising with bidirectional attention, allowing temporal boundaries and semantic content to be refined jointly. Direct adaptation, however, creates two TAL-specific mismatches: standard masked diffusion training corrupts all positions uniformly at random, but the time tokens are more reliable when enough semantic context is available; and token-level cross-entropy does not reflect temporal IoU. To address these mismatches, we introduce a Planned Training Objective that uses boundary-aware masking and step-weighted reconstruction to rehearse the late recovery of time tokens, together with a Step-Level IoU Reward that provides overlap-aware supervision during denoising. A standard sequence-level cross-entropy term provides the base reconstruction signal. Experiments on ActivityNet-RTL, ActivityNet-1.3, and THUMOS-14 show that MDVLM-TAL improves both temporal reasoning and boundary localization over autoregressive vision-language baselines, with especially strong gains under stricter temporal IoU criteria.
PaperID: 2679, Poster
Abstract: Object-level spatial-temporal understanding is essential for video question answering, yet existing multimodal large language models (MLLMs) encode frames holistically and lack explicit mechanisms for fine-grained object grounding. Recent work addresses this by serializing bounding box coordinates as text tokens, but this text-coordinate paradigm suffers from a fundamental : object information is inherently visual, yet encoding it as text incurs a high token cost that forces aggressive temporal downsampling. We propose , an object-aware visual prompting framework for multimodal model fine-tuning that moves spatial-temporal object states from text-coordinate serialization into the visual stream while retaining only object identity in a compact text legend. Colored bounding boxes and trajectory trails encode geometry and motion on video frames, while the color-to-object legend provides the minimal textual identity bridge. This reduces the token cost significantly, achieving 87-93% text token reduction in practice. It also preserves full temporal resolution, where the trajectory trails further encode inter-frame motion direction and speed within each keyframe, recovering fine-grained dynamics that text-coordinate methods are forced to discard. Experimental results on five video QA benchmarks (CLEVRER, Perception Test, STAR, NExT-QA, IntentQA) show that BoxTuning surpasses text-coordinate baselines on spatially oriented tasks and avoids their degradation on reasoning-centric tasks, establishing object-aware visual prompting as a natural and efficient way to convey object information to MLLMs.
PaperID: 2680, Poster
Abstract: Real-world evolving networks are naturally modeled as temporal graphs (TGs), where capturing temporal dynamics is essential for predicting future graph properties that support downstream decision-making. Existing temporal graph methods have been developed primarily for single-task prediction, and little is known about their generalization across tasks or transfer to unseen networks. This leaves the challenge of multi-task graph property prediction in TGs largely open. We address this challenge by introducing Hydra, a novel architecture that integrates local connectivity features from temporal GNNs with a spectral learning module that captures global connectivity patterns. This design enables joint learning of local and global information under a multi-task objective. In multi-task classification, Hydra achieves an 8.9% relative gain in AUC over the strongest competitor. In multi-task regression, Hydra achieves competitive results in all three tasks, while obtaining the best results in two tasks with a 8.2% relative gain in MAE compared to the strongest baseline. Moreover, Hydra delivers these gains with a 22× reduction in training time compared to temporal transfer models. These results provide the first systematic evidence that multi-task transferable learning on temporal graphs is effective. By delivering consistent top-ranked performance, Hydra highlights multi-task training on temporal graphs as a promising direction toward adaptable foundation models for temporal graphs.
PaperID: 2681, Poster
Authors: Yue Chen, Menghan Li, Tatsuo S Okubo
Abstract: Intracortical brain–computer interfaces (iBCIs) can restore movement and communication, but neural drift can degrade performance across recording sessions. Because recalibration adds user burden, there is a need for decoders that adapt from minimal calibration data. Although session-specific input layers are a common way to handle cross-session drift, they can underperform in the extreme few-shot regime because the shared decoder is trained downstream of well-estimated session-specific layers, but must rely at deployment on a new-session layer fit from only a few calibration trials. To avoid this dependence, we propose a two-stage approach: first, train a single GRU decoder on pooled held-in sessions without session-specific input layers, encouraging the backbone to handle cross-session variability directly; second, freeze this backbone and meta-learn a lightweight alignment layer for rapid calibration. We evaluate on FALCON H2, an official human handwriting iBCI benchmark designed for few-shot cross-session decoding, and achieve 7.41% word error rate using only three released calibration sentences per held-out session. These results suggest that, under extreme calibration limits, learning a transferable backbone before adding session-specific adaptation can be more effective than jointly training flexible session-specific layers.
Abstract: A central goal of science is to produce valid explanations of complex systems: high-level causal accounts that faithfully reflect the behavior of lower-level mechanisms. Yet no consensus exists on how to measure whether a proposed high-level explanation is actually valid. We introduce a benchmark of ten complex systems spanning discrete and continuous state spaces and static and dynamical regimes, each equipped with consensual ground-truth causal explanations and invalid contrastive conditions. Within a unified causal abstraction framework, we systematically evaluate over thirty candidate metrics drawn from observational, functional, information-theoretic, and causal families. Our results show that only the latter reliably discriminates valid from invalid abstractions, and only when incorporating faithfulness testing over unmapped variables. Building on these findings, we introduce the Causal Abstraction Error (CAE), a continuous validity metric with an explicit faithfulness test, which passes all discrimination tests across every system and converges with as few as 30 sampled interventions. We offer it as a general-purpose metric for the discovery and validation of high-level explanations of complex systems.
Abstract: Large models are increasingly becoming autonomous agents that interact with real-world environments and use external tools to augment their static capabilities. However, most recent progress has focused on text-only large language models, which are limited to a single modality and therefore have narrower application scenarios. On the other hand, multimodal large models, while offering stronger perceptual capabilities, remain limited to static knowledge and lack the ability to access and leverage up-to-date web information. In this paper, we propose VSearcher, turning static multimodal model into multimodal search agent capable of long-horizon, multi-turn tool use in real-world web environments, including text search, image search, and web browsing, via reinforcement learning. Specifically, we introduce Iterative Injection Data Synthesis pipeline to generate large-scale, complex multimodal QA questions, which are further filtered with comprehensive metrics to ensure high quality and sufficient difficulty. We then adopt an SFT-then-RL training pipeline to turn base multimodal models to agent capable of multi-turn tool calling in real-world web environments. Besides, we propose a multimodal search benchmark MM-SearchExam dedicated to evaluating search capabilities of multimodal search agents, which proves highly challenging for recent proprietary models. Extensive evaluations across multiple multimodal search benchmarks reveal effectiveness of our method. VSearcher achieves superior performance compared to recent multimodal search agents and even surpasses several proprietary models on multimodal web search tasks.
PaperID: 2684, Poster
Abstract: Next-token prediction leaves hidden-state trajectories largely unconstrained, motivating recent geometric regularization methods based on angular penalties. However, these approaches control only the direction of motion and provide no explicit bound on trajectory deviation. We introduce the Controlled Joint Embedding Predictive Architecture (ControlJEPA), a trajectory regularization framework grounded in discrete-time Input-to-State Stability theory. Our novel Lyapunov Tube Loss enforces a contraction condition on a normalized measure of transversal deviation, yielding a provable guarantee that confines hidden states within a tube of explicit width determined by two interpretable hyperparameters. ControlJEPA integrates seamlessly with standard training and requires no architectural changes. Across six benchmarks and seven model families, it consistently outperforms standard fine-tuning and prior trajectory regularizers, improving accuracy and data efficiency while preserving output diversity. Notably, the induced geometric structure persists at inference time, indicating that the method reshapes the representation space rather than memorizing training trajectories.
PaperID: 2685, Poster
Abstract: Existing latent graph diffusion models often learn graph representations and generative priors within a homogeneous latent geometry. This assumption is restrictive for graph generation with structural heterogeneity, where hierarchical, cyclic, locally regular, and densely connected patterns may coexist within the target distribution. Under such conditions, forcing all structural patterns into one latent geometry can introduce representation distortion, which may then propagate to generation. We propose CurvDiff, a latent graph generation framework based on mixed-curvature latent diffusion. The main idea is to model structurally different graph patterns in geometry-compatible latent factors rather than representing them in a single-curvature space. Specifically, CurvDiff learns graph embeddings in a product latent geometry consisting of hyperbolic, spherical, and Euclidean components, and maps the resulting representation to a shared tangent-space interface on which latent diffusion is defined. This provides a consistent latent pathway from representation learning to generation. To further preserve topology-relevant structure, CurvDiff uses Ricci-derived structural signals as topology-aware priors. These signals regularize latent representation learning and condition the diffusion prior, helping maintain structural consistency across representation learning and generation. Experiments on synthetic and real-world graph benchmarks show that CurvDiff achieves competitive or superior graph distribution matching compared with strong baselines, with improved fidelity on degree, clustering, and spectral statistics. These results support mixed-curvature latent diffusion as an effective approach to graph generation with structural heterogeneity.
PaperID: 2686, Poster
Authors: Yucong Liu, Florian T Schaefer
Abstract: We study sketched Ridge regression through the lens of functional estimation. The sketched solution is a plug-in estimator of a nonlinear function of a second-moment/covariance matrix. Within this framework, we analyze two classical bias-reduction principles---iterative Bootstrap and linear aggregation across multiple estimators. We show that these resampling-based procedures cancel low-order bias terms and yield higher-order bias decay in the sketch size without increasing the order of the computational complexity. Concretely, let A\in\mathbbR^n× r have full column rank r and let the sketch size be s. For Gaussian sketching, a k-step iterative Bootstrap estimator achieves bias of order (\sqrtr/s)^k+2. For linear aggregation over m plug-in estimators with different sketch sizes, we obtain bias of order (\sqrtr/s)^m+1 for Gaussian sketching and corresponding result for common discrete sketches, including uniform and leverage score sampling.
Abstract: We propose Drift-Resistant Navigation World Model, a generative model that mitigates both perceptual drift and geometric drift in conventional rollout-based navigation world models. Existing methods recursively feed generated content into subsequent steps, causing noise accumulation and degraded predictions, i.e., perceptual drift. Meanwhile, their predictions often deviate from the agent’s motion, resulting in geometry drift. We address both types of drift by redesigning world-model prediction as an anchor-guided rollout. Instead of rolling out every frame sequentially, we first predict sparse future anchors that serve as stable long-range targets, and then generate intermediate frames within each chunk conditioned on both past context and future anchors. Importantly, these sparse anchors also provide geometric constraints, supported by bidirectional epipolar geometry, to localize where corresponding content should appear in the intermediate frames. Experiments on four benchmarks demonstrate consistent improvements over strong baselines in long-horizon visual quality, geometric consistency, and multi-view coherence. These gains further translate into improved downstream planning performance under the same planners, highlighting the importance of drift-resistant, geometry-aware prediction for reliable navigation world models.
PaperID: 2688, Poster
Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models on reasoning tasks. To avoid the complexity of training explicit critics, algorithms like GRPO estimate advantages directly from sparse on-policy rollouts. However, the efficacy of this approach relies heavily on the accuracy of the baseline estimator. In this work, we argue that standard methods, whether relying on prompt-local averages or shrinking toward a single global mean, are statistically mis-specified for reasoning domains. By enforcing a uniform prior, these methods ignore the latent skill structure of tasks, introducing estimation bias that conflates dissimilar problems. To address this, we propose Skill-Coupled Policy Optimization (SCPO), a structured baseline estimation framework that refines the shrinkage prior by leveraging task correlations. Instead of a generic global target, SCPO performs shrinkage toward skill-coupled group targets, thereby aligning the baseline with the local difficulty landscape. To further stabilize these targets under limited on-policy observations, SCPO incorporates a lightweight history tracker. Experiments across diverse models and tasks show that SCPO consistently improves training stability and performance over both prompt-level and global shrinkage baselines.
Abstract: We study cooperative multi-agent reinforcement learning in the setting of reward-free exploration, where multiple agents jointly explore an unknown MDP in order to learn its dynamics (without observing rewards). We focus on a tabular finite-horizon MDP and adopt a phased learning framework. In each learning phase, multiple agents independently interact with the environment. More specifically, in each learning phase, each agent is assigned a policy, executes it, and observes the resulting trajectory. Our primary goal is to characterize the tradeoff between the number of learning phases and the number of agents, especially when the number of learning phases is small. Our results identify a regime change governed by the horizon H. When the number of learning phases equals H, we present a computationally efficient algorithm that uses only \tildeO(S^6 H^6 A / \epsilon^2) agents to obtain an \epsilon approximation of the dynamics (i.e., yields an \epsilon-optimal policy for any reward function). We complement our algorithm with a lower bound showing that any algorithm restricted to \rho < H phases requires at least A^H/\rho agents to achieve constant accuracy. Thus, we show that having \Theta(H) learning phases is both necessary and sufficient when restricting the number of agents to be polynomial.
PaperID: 2690, Poster
Abstract: Few-shot object detection (FSOD) requires detectors to adapt to novel categories from only a few labeled instances, where prototype-based transfer methods have become a strong and efficient paradigm. However, we observe that such methods remain unstable in extreme low-shot regimes. We attribute this instability to two coupled factors: semantically unanchored query initialization, which yields high-variance cold-start optimization, and deterministic prototype matching, which over-trusts noisy or occluded support regions. To address these limitations, we propose Semantic Fine-Grained Prototype Distillation (SFPD), a language-conditioned probabilistic adaptation framework for prototype-based FSOD. SFPD introduces language-conditioned feature queries to provide semantic anchors for novel-class adaptation, uncertainty-aware prototype distillation to down-weight unreliable support evidence through heteroscedastic Gaussian modeling, and component-guided part-aware prototypes to refine fine-grained semantic-visual alignment. These modules act during adaptation and preserve the original detector path at inference, introducing no additional inference FLOPs. Experiments on PASCAL VOC and MS-COCO show that SFPD consistently improves a strong FPD baseline, with especially clear gains in the most challenging low-shot settings, e.g., a 3.9-point nAP50 improvement on VOC Split 2 under 1-shot. Further ablations, multi-seed evaluation, and convergence analysis indicate that SFPD improves both accuracy and adaptation stability.
PaperID: 2691, Poster
Abstract: Cross-view object correspondence (CVOC) aims to establish associations of the same target between query and search videos captured from different viewpoints. Existing methods often rely on visual cues for CVOC. Despite recent progress, these methods struggle in complicated scenarios with substantial viewpoint changes and appearance variations, due to lacking sufficient target information. Addressing this, we introduce TeSCo, a novel framework that exploits rich Textual Semantics of an object, in addition to its visual cues, for cross-view object Correspondence. Our key insight is, textual description of a target, by capturing diverse attributes such as category, texture, and color, provides viewpoint- and appearance-invariant semantics, which are complementary to visual cues and can hence enhance correspondence robustness in complicated scenes. Inspired by this, TeSCo first generates a textual expression of the target from the given query view using a vision-language model, and then extracts textual feature to enhance visual features for localization in the search view. To realize this, we introduce a simple yet effective text-conditioned cross-modal fusion (TCF) module that incorporates the textual feature into visual features with text-guided modulation, producing more robust multimodal target representation for object correspondence. Since not all textual cues are equally compatible with the target in current search-view frame, we propose an iterative textual feature refinement (ITFR) module that progressively adjusts textual feature using search-view information before the TCF module, enabling search-view-aware textual semantics to be fused into visual features and leading to better performance. In extensive experiments on Ego-Exo4D and HANDAL-X, TeSCo demonstrates state-of-the-art results and largely outperforms other models, validating its efficacy. Our code and model as well as results will be released.
Abstract: Tokenization that partitions time series into subsequences has become a foundational paradigm in modern time series modeling. However, we find that current tokenization strategies inherently introduce spectral distortions, forcing the learned representations to diverge from the true underlying patterns such as trend, periodicity, and seasonality, thereby impairing model performance. Through theoretical analysis from a frequency perspective, we derive the boundary conditions under which such distortions occur. To overcome this distortion, we propose , a nearly distortion-free frequency-based tokenizer for time series data. TimeFT obtains tokens through frequency-domain partitioning, frequency shifting, and Nyquist sampling. The method is parameter-free, incurs negligible computational complexity, and can serve as a drop-in replacement for existing tokenizers. Our extensive experiments on diverse tasks such as forecasting, classification, and anomaly detection demonstrate that TimeFT consistently and significantly improves performance.
PaperID: 2693, Poster
Abstract: Kernel two-sample tests based on Maximum Mean Discrepancy (MMD) are widely used, but the standard quadratic time MMD statistic requires evaluating \Theta(N^2) kernel pairs where N is the pooled sample size. Several subquadratic approximated MMD tests have been proposed to alleviate this issue, but existing methods either compromise the test power or make additional distributional assumptions to achieve power comparable to the quadratic-time MMD test. Motivated by this gap, we propose a sub-quadratic time computational method to approximate the MMD test, achieving minimax rate-optimal power with minimal distributional assumptions. Building on ideas in fast kernel density estimation and exponentially convergent trapezoidal rule in numerical analysis, our approximation reduces all-pairs kernel summation to fast Fenwick-tree prefix-sum queries. For fixed dimension, this yields subquadratic-time algorithms for several common characteristic kernels, including Laplace, Mat\'ern, Gaussian, and inverse multiquadric kernels. The approximation error is controlled tightly enough that the resulting subquadratic time test retains the power guarantees of the minimax-optimal unapproximated quadratic-time MMD test. The same implementation remains subquadratic for dimension d=o(\log N/\log\log N), and can be paired with dimensionality-reduction algorithms when the signal is low-dimensional.
Authors:
Hongwei Zhang, qiqiang zhong, Jiangxia Cao, Junfeng Shu, Yiyang Lv, Huanjie Wang, Liwei Guan, Jing Yao, Yiyu Wang, Liu Zhaojie, Han LiAbstract: Modeling ultra-long user behavior sequences is an essential task for capturing evolving preferences in modern recommender systems, and this research direction has contributed solid gains in the past several years. However, recommendation systems always serve enormous traffic, and extending user sequences sharply increases computational cost, which creates a difficult trade-off between efficiency and effectiveness. To scale to longer user sequences while keeping lightweight serving computation, existing works can be divided into two paradigms: (1) Search-based Top-K selection, which constructs an item-specific subsequence for each candidate item to avoid facing the ultra-long sequence directly; and (2) pre-trained user-interest compression, which maps an ultra-long user sequence into a small group of item-agnostic user interest memories so that the online model can perceive user long-term interests via this highly compressed dense memory. Besides the two technical routes (totally item-specific or item-agnostic), we argue that there exists an intermediate path not well explored: preserving partial relevance between the user sequence and the target item, while exposing only limited signals to guide the direction of interest compression. This design does not pursue item-specific user interest compression, but seeks semantic-group shared general user interest memory according to item attributes, where semantically similar items share the same compressed user interest memory. Motivated by this, we propose UxSID, a novel framework that bridges this gap by facilitating a target item semantic-aware interaction between User histories and candidate Semantic IDs (SIDs). Specifically, UxSID employs a dual-level attention strategy: it first extracts item-agnostic user interests from raw sequences, and then performs a semantic-specific query over global behaviors and those agnostic interests to generate semantic-specific preferences. By adopting this end-to-end architecture, UxSID generates offline embeddings that balance computational parsimony with target items' semantic awareness, while strictly preserving parity with online inference in constant time. Extensive public benchmarks and large-scale A/B tests demonstrate that UxSID achieves state-of-the-art performance, driving a 0.337% revenue lift in advertising. Open source link: \urlhttps://anonymous.4open.science/r/UxSID/
PaperID: 2695, Poster
Abstract: Evaluating the physical consistency of generated videos remains a fundamental challenge. Existing approaches rely on off-the-shelf vision-language models, which can often be myopic to physical dynamics, or fine-tuned evaluators trained on human annotations, which overfit to dataset-specific cues and fail to generalize. A key challenge is that existing supervision sources provide either relative ordering or absolute scores, but not both reliably and consistently across varied settings. To this end, we introduce PhyProbe, an evaluator that extracts features from a frozen pretrained spatio-temporal encoder and maps them to a scalar physical consistency violation score via a lightweight scoring head. PhyProbe is trained through a unified objective combining pairwise ranking, regression on noisy scalar annotations, and anchor-based calibration over a curated set of heterogeneous supervision sources. Experiments show that PhyProbe significantly outperforms prior methods on pairwise benchmarks spanning real–generated and generated–generated pairs under varying correspondence, including challenging settings where existing approaches degrade toward random performance. PhyProbe achieves strong correlation with human judgments, with close agreement between rank-based and linear metrics, indicating both reliable ordering and calibrated scores. Further, despite being trained solely for physical consistency, PhyProbe also performs competitively on human preference benchmarks: consistent with the observation that physics violations are entangled with broader quality degradations.
PaperID: 2696, Poster
Authors: Guoliang Xu
Abstract: Active feature acquisition (AFA) requires an agent to choose a budget-limited set of costly features before making a prediction. Greedy conditional-mutual-information (CMI) and value-of-information policies perform well when useful features are individually informative, but fail under joint-only evidence: a query is valuable only because it enables a future evidence set. To address this, we propose Forward-\Phi, a model-based non-myopic acquisition scorer that requires a known structural causal model (SCM) or generative AFA model. Forward-\Phi evaluates a candidate query by its Shapley contribution to budget-feasible future action sets in a cooperative game on the expected forward utility. Two variants share the same game: a raw score that keeps both singleton and interaction value, and a residual score that strips the singleton component to isolate joint-only utility on diagnostic probes. On AFABench CUBE-NM, the raw Monte Carlo (MC) Forward Shapley scorer reaches 0.664 \pm 0.024 accuracy over five MC seeds, well above all reported baselines including the OL-MFRL (0.235) and ODIN-MFRL (0.160) reinforcement-learning loops trained under the same split and shared pretraining protocol. Controlled SCM and Bayesian-network (BN) probes confirm the predicted mechanism: Forward-\Phi yields substantial gains when target information is jointly carried, and matches Greedy on greedy-aligned controls. The primary contribution is the belief-conditional Forward Shapley state-action score, a credit-assignment signal derived from Shapley axioms for non-myopic AFA under joint-only evidence.
Abstract: We study the numerical approximation of Schrödinger Bridges in the setting where samples from the target distribution are unavailable and only an unnormalized target density is given. In this regime, the widely used iterative proportional fitting (IPF) procedure (and its variants) often exhibits severe numerical instabilities, stemming from ill-conditioned underlying control problems and from the fact that the optimal solution may lie far from the initialization. To address these issues, we propose a trust region variant of IPF that constrains successive updates to remain at a fixed Kullback-Leibler divergence from the current iterate. This restriction significantly improves numerical stability. Moreover, we show that the resulting method admits an interpretation as a mirror descent scheme with an adaptively chosen step size. We demonstrate the effectiveness of our approach on several challenging sampling tasks.
PaperID: 2698, Poster
Authors: Jonathan F Carter, Lionel Tarassenko
Abstract: Foundation models offer a promising route to compress multi-modal physiological signals into better measures of human health, with broad applications across sleep medicine, cardiology, neurology and other healthcare domains. Existing models are typically trained with contrastive or masked-reconstruction objectives. However, masked reconstruction may be poorly suited to the stochastic nature of these signals, while contrastive approaches rely on positive-pair definitions despite the semantic invariances of physiological signals being poorly understood. In this work, we show that next-token prediction is a simple and scalable alternative. We develop Hypnos, a multi-modal sleep foundation model trained via next-token prediction using eight different sensing modalities (e.g. EEG, ECG, respiratory signals) from overnight polysomnography recordings. We tokenize each modality into streams of discrete tokens using residual vector quantization. We then train a large auto-regressive RQ-Transformer to jointly predict the next token across all modalities in parallel, using a novel modality masking strategy to enable generalisation to subsets of modalities during inference. Using over 20,000 overnight recordings drawn from nine public datasets, we find that both next-token perplexity and downstream probing performance continue to improve with model scale. Across a range of tasks, Hypnos matches or exceeds prior sleep foundation models and strong supervised baselines on sleep stage classification across in-domain and held-out test sets. Our results indicate that next-token prediction is a strong self-supervised objective for learning representations from multi-modal physiological signals.
PaperID: 2699, Poster
Abstract: Tabular prediction has long been dominated by gradient-boosted decision trees and specialized deep tabular models, while large language models (LLMs) remain difficult to make competitive despite their cross-task adaptability and transparent reasoning traces. Existing LLM-for-tabular methods mainly adapt tables into textual inputs, but often fail to capture key structural properties of tabular data. In this work, we propose Permutation Relative Policy Optimization (PRPO), a reinforcement learning post-training method that strengthens LLMs for tabular prediction by injecting column-permutation invariance as a structural prior. By constructing label-preserving column permutations and estimating advantages both within and across them, PRPO converts sparse outcome rewards into denser and more stable optimization signals. Extensive experiments on 139 OpenML datasets show that our 8B model reaches a genuinely competitive regime against strong specialized tabular baselines. It achieves strong fully supervised performance, dominates cross-dataset zero-shot settings, and performs on par with 32-shot strong baselines. Moreover, it substantially outperforms much larger general-purpose and reasoning LLMs, including up to a 53.17% improvement over DeepSeek-R1 (685B). These results show that structural-prior RL post-training is an effective route for making LLMs competitive in tabular prediction.
Abstract: Neural surrogate models for physical simulations are trained on discretized samples of continuous domains, where the induced empirical measure leads to uneven supervision, biasing optimization and causing spatial inconsistencies in physical fidelity. To mitigate this measure-induced bias, we propose M^3 (Multi-scale Morton Measure), a scalable framework that balances training measures by partitioning space according to physical variation and allocating supervision across multiple scales. Applied to three industrial-scale datasets with diverse discretizations, M^3 consistently improves predictions in the continuous physical domain, achieving up to 4.7× lower error in large-scale volumetric cases. These gains persist under aggressive subsampling (160M \rightarrow 16M \rightarrow 1.6M points), where M^3-trained models outperform those trained on higher-resolution data, reducing physics-weighted relative L_2 error by 3--4× and the corresponding MSE by up to 13×. These results highlight data distribution as a key factor in operator learning and position M^3 as a scalable, data-efficient approach for physically consistent modeling.
PaperID: 2701, Poster
Abstract: Many physical systems do not merely move or deform; they grow, adding material and changing the geometry that a world model must represent. Existing world models are typically optimized for pixel prediction, reward prediction, or fixed-support physical dynamics, leaving open how to model systems whose geometric state expands over time and whose future morphology depends on hidden material response. We introduce FOLIAGE, a geometry-centered latent world model for growing surfaces. FOLIAGE represents mature regions as a compact scaffold while allocating higher-resolution state to active growth fronts, allowing the model to focus capacity where new material and near-future change occur. It further separates observation, action, and privileged physics: heterogeneous RGB, point-cloud, and mesh observations are fused into a deployable geometric state; material controls condition only the latent dynamics; and hidden physical energies guide training without being required at deployment. To evaluate this setting, we introduce SURF-GARDEN and SURF-BENCH, providing controlled counterfactual branches, dense cross-modal correspondences, hidden physical signals, and stress tests for growing-geometry state learning. FOLIAGE reduces inverse-material error by \approx40% and 5-step mesh forecasting Chamfer error by \approx20% relative to strong baselines, while improving cross-modal retrieval by \approx25% mAP@100. Stress tests show graceful degradation under sensor loss and correspondence corruption. Transfer experiments on real plant and dynamic-3D data indicate that the learned geometry-centered state generalizes beyond the simulator. We will release code and data upon acceptance.
PaperID: 2702, Poster
Authors:
Zhaoyu Fan, Bowen Han, Yanwei Ren, Jincheng Ou, Siyi Cao, Zehua He, Kaiyu Huang, Pengcheng He, Shanshan Ni, Chengchen Gong, Qingjiang ShiAbstract: Real-world studies (RWS) analyze routinely collected healthcare data, such as electronic health records and insurance claims, to estimate treatment effects, monitor safety, and identify risk factors outside randomized trials. Automating RWS with large language model agents is appealing but brittle: valid analyses depend on dataset-specific clinical semantics, including index dates, follow-up windows, patient/event-level units, and outcome boundaries that are often implicit in observational data. This leads to schema mismatches, temporal errors, inconsistent analytical granularity, and unsupported clinical claims. We present ProCARE, a profile-grounded framework that reformulates RWS automation as constrained compilation. ProCARE converts heterogeneous data profiles into Evidence Contracts specifying available variables, temporal attributes, distributions, and evidentiary limits, and compiles study objectives into RWS-IR, a strongly typed intermediate representation for variables, time relations, analysis units, and statistical tasks. Deterministic validation and localized repair then check and revise generated workflows across schema mapping, cohort construction, variable derivation, analysis, and interpretation. On 36 EHR-based benchmark tasks, ProCARE outperforms strong LLM-agent baselines in code executability, semantic consistency, and conclusion traceability, with automated violation diagnoses showing moderate-to-substantial agreement with blinded expert review. Code, benchmark prompts, evaluation scripts, and the fully de-identified HCC benchmark tables are publicly available at https://anonymous.4open.science/r/ProCARE-7E52.
PaperID: 2703, Poster
Authors: Sejin Park, Hongjae Lee, Changwoo Han, Seung-Won Jung
Abstract: Differentiable logic gate networks, which operate using only logic gates, have recently attracted attention as an efficient alternative to conventional neural networks. However, despite their efficiency, the scaling behavior of logic gate networks remains underexplored. By contrast, scaling model capacity is a central design principle in deep neural networks and typically leads to improved performance. This discrepancy raises a key question: Can similar scaling benefits also be achieved in logic gate networks? In this work, we focus on width as a primary scaling axis and conduct a systematic analysis of its behavior in logic gate networks. We observe that naive width scaling often introduces redundancy among logic kernels, limiting the effective use of additional kernels and leading to performance saturation. To address this limitation, we propose a dynamic logic kernel framework that reorganizes kernel utilization by promoting specialization across kernel groups. This enables the network to better utilize increased width via input-dependent kernel routing, while ensuring that both routing and computation are implemented entirely with gate-level Boolean operations at inference time. We further find that kernel redundancy is most pronounced at the first gate level, motivating an early-stage dynamic logic kernel strategy that concentrates adaptation at this level. Experimental results demonstrate that our approach improves kernel utilization and increases kernel diversity, leading to higher accuracy with improved parameter efficiency.
Authors: Tingkai Jia, Ting Wang, Cheng Chen
Abstract: In this work, we study nonconvex–strongly convex online bilevel optimization (OBO) using only first-order oracle. Existing OBO algorithms are mainly based on hypergradient descent, which requires access to a Hessian-vector product (HVP) oracle and potentially incurs high computational costs. By reformulating the original OBO problem as a single-level online problem with inequality constraints and constructing a sequence of Lagrangian function, we eliminate the need for HVPs arising from implicit differentiation. Specifically, we propose a fully first-order algorithm for OBO, and provide theoretical guarantees showing that it achieves regret of O(1 + V_T + H_2,T) with a total of O(T\log T) iterations, where V_T measures the variation in function values and H_2,T characterizes the drift variation of the inner-level optimal solution. We also establish a sublinear regret bound under the single-loop structure by introducing additional gradient-variation terms. Furthermore, we develop an improved variant with an adaptive inner-iteration scheme, which removes the dependence on H_2,T and achieves regret of O(\log T + V_T). Finally, under the stochastic OBO setting, we establish the regret bound for the fully first-order algorithm, i.e., O(T^2/3(1 + \sigma^2) + V_T + H_2,T). Numerical experiments demonstrate the feasibility of our algorithm and support our theoretical findings.
PaperID: 2705, Poster
Abstract: Multivariate time-series forecasting requires reliable cross-series interaction, yet many attention- and mixing-based models couple heterogeneous variables without explicit entry-stage alignment. Such uncalibrated early interaction can lead to poorly conditioned optimization and degraded generalization when cross-series statistics drift. We propose First-layer Statistical Calibration (FISC), a simple framework that imposes statistical calibration constraints only on the first cross-series coupling layer while leaving deeper coupling layers learnable without repeated calibration. FISC estimates training-set cross-series statistics, converts them into a statistical affinity matrix, and convexly fuses this matrix with learnable coupling parameters. A row-wise simplex projection then yields a nonnegative row-stochastic mixing operator, providing bounded and interpretable entry-layer interaction. By restricting calibration to the first layer, FISC calibrates the entrance interaction without repeatedly constraining deeper feature transformations. To reduce temporal redundancy, FISC further combines a time-domain branch with a data-adaptive orthogonal-domain branch derived from training-time temporal statistics. Experiments on standard long- and short-term forecasting benchmarks show that FISC improves forecasting accuracy while maintaining favorable efficiency. Ablations and non-stationarity analyses further show that first-layer-only calibration consistently outperforms no calibration, last-layer calibration, and all-layer calibration, supporting a simple principle for robust cross-series learning: constrain the entrance interaction with training-set statistics, while letting deeper representations remain data-adaptive.
PaperID: 2706, Poster
Abstract: Heterogeneous Federated Learning (HFL) aims to enable collaboration among clients with diverse model architectures and non-IID data distributions, where direct parameter aggregation is often infeasible and prototype-based aggregation may suffer from semantic drift. Existing methods usually construct global references from client-side representations, making the alignment process vulnerable to local statistical bias and noisy or corrupted prototype updates. In this paper, we propose SATPL, a synthetic anchor-assisted prototype alignment framework that introduces externally generated synthetic anchors as auxiliary semantic priors for HFL. Instead of treating Large Language Models (LLMs) as infallible semantic oracles, SATPL uses LLM-generated class descriptions and generative models to construct reproducible reference prototypes for supervised tasks with known class semantics. Each client maps its heterogeneous local representation into a shared anchor space through a lightweight projection head, while an alignment-weighted aggregation rule assigns lower weights to prototypes that are highly inconsistent with the corresponding anchors. The blockchain component is used as an auditable recording layer for anchor commitments and aggregation metadata, rather than as the source of the statistical robustness guarantee. Experiments under architecture heterogeneity, label skew, and random prototype corruption show that SATPL improves accuracy, convergence stability, and robustness over representative HFL baselines. Additional analyses evaluate anchor quality, temperature sensitivity, and system overhead, clarifying both the benefits and limitations of synthetic-anchor-based alignment.
Authors:
Zhenchao Tang, Fang Wang, Haohuai He, Jiale Zhou, Tianxu Lv, Jun Zhu, Shou Z Chen, Minghao Yang, Yu Wang, Jiayang Wu, Yidong Song, Yaokun Li, Jiehui Huang, Jun Zhou, Bing He, Jianhua YaoAbstract: Engineering LLMs to accelerate life sciences research requires a robust alignment with biomedical knowledge. We observe that biomedical text exhibits a fundamentally different uncertainty structure from general text: dense low-confidence runs encode epistemic knowledge gaps (dense causal chains, rare entities) rather than the sparse aleatoric stylistic variation typical of general text. Based on this discovery, we propose Balanced Fine-Tuning (BFT), a dual-scale post-training method that combines group-normalized token reweighting with sequence-level reallocation toward knowledge-dense samples exhibiting dense epistemic uncertainty. Across medical evaluation, biological reasoning, sparse-reward RL, and biological representation tasks, BFT provides more consistent gains than SFT and DFT under a shared training setup. When replacing the default closed-source backbones in GeneAgent (GPT-4o) and VCWorld (Gemini-2.5-Flash), the BFT-aligned 70B model delivers stronger performance across biological process reasoning and chemical perturbation prediction. Critically, all BFT variants further improve after subsequent GRPO with sparse rewards, while SFT and DFT degrade, suggesting that epistemic-aware post-training provides a more robust policy initialization. Beyond text generation, BFT-aligned LLMs produce more accurate and professional biomedical profile texts; after encoding these profiles with a text embedding model, the resulting representations support gene-level, cell-level, and perturbation-response tasks, suggesting that BFT-enhanced generation can facilitate biological representation and, in turn, broader biomedical downstream tasks.
Abstract: Conformal prediction (CP) constructs prediction sets with marginal coverage guarantees under the assumption that the calibration and test distributions are identical. However, under distribution shift, existing approaches primarily align marginal conformal score distributions, which is sufficient to preserve marginal coverage but does not control the conditional coverage error at individual test inputs. As a consequence, CP can remain unreliable in regions where the conditional score distributions are mismatched. In this work, we bound the conditional invalidity of CP under distribution shift in terms of the Wasserstein distance between the calibration and test distributions. This result highlights the role of invertible transport in mitigating conditional coverage degradation. Motivated by this insight, we introduce Branched Normalizing Flow (BNF), a two-branch architecture that normalizes a test input to the calibration distribution and transforms the prediction set of the normalized input back to the test distribution while preserving conditional guarantees. Empirically, BNF consistently improves conditional coverage robustness on nine datasets across a wide range of confidence levels.
Abstract: Compositional prompts in text-to-molecule generation often specify several structural constraints at once, such as scaffolds, functional groups, substituents, and finer modifiers. Existing molecular latent diffusion models typically inject the full prompt once and apply it uniformly throughout denoising, which can cause later constraints to overwrite or distort earlier structural commitments. We propose , a training-free progressive latent diffusion strategy for text-guided molecular design. Technically, CoG realizes progressive generation as multi-stage latent denoising over cumulative molecular sub-prompts: it factorizes a prompt into chemistry-informed coarse-to-fine segments, conditions each stage on the cumulative sub-prompt, and continues denoising from the previous-stage latent using a fixed partial-denoising schedule. This turns one-shot conditional generation into a sequence of stage-wise denoising steps that help retain earlier structure while introducing later constraints. Across benchmark, out-of-distribution, and ablation settings, CoG improves structural prompt fidelity over one-shot conditioning on the tested latent graph-diffusion backbones.
PaperID: 2710, Poster
Authors: Salah Chikhi, Luis H Ruiz, Entong Li, Li Zeng
Abstract: We present the Adaptive Scheduling Pipeline (ASP), a bounded-lag asynchronous reinforcement learning (RL) architecture for post-training large language models (LLMs). Rather than only disaggregating generation and training, the ASP combines two design ideas. First, a queue-aware scheduling policy that jointly reduces three sources of rollout underutilization —padding bubbles from intra-minibatch length heterogeneity, skewness bubbles from long-tail outputs, and sorting bubbles from queue instability induced by length-sorted dispatching—through a single offline preprocessing pass. This enables efficient multi-instance rollout that overcomes the early scaling plateau of single-instance inference while maintaining stable queue dynamics. Second, a bounded-lag asynchronous architecture decouples rollout from training while keeping policy lag observable and small enough to preserve stable learning in our experiments. We also provide an explicit weight-flow path that converts trainer snapshots into inference-ready weights without re-coupling generation and training. Experiments show that ASP improves end-to-end throughput by up to 3.22×, reduces iteration time by up to 76%, scales more robustly with additional generation resources than the MindSpeed-RL and verl baselines, and shows no visible convergence degradation. ASP has also been used in an internal telecom-domain RL post-training workflow, reducing the reported training period from more than three months to about two months and producing a domain-enhanced model for knowledge answering and tool calling.
PaperID: 2711, Poster
Abstract: Diffusion-based neural combinatorial optimization (NCO) has recently shown promise for routing problems by generating solutions through denoising or reconstruction. Most existing diffusion solvers, however, define the generative process over full-instance solution structures, such as adjacency matrices or heatmaps. This instance-level formulation ties the learned reconstruction distribution to the problem size observed during training, making it difficult to transfer models trained on small instances to substantially larger ones. In this paper, we propose SeqDiCO, a sequence-oriented diffusion framework for scale-generalizable neural combinatorial optimization. The key idea is to move diffusion from full-instance solutions to fixed-endpoint sequence subproblems. Given two endpoints and a set of intermediate nodes, SeqDiCO learns a local denoising operator that reconstructs the open path connecting the endpoints from corrupted structural observations. Since such local sequence patterns are reusable across instance sizes, the same operator trained on small instances can be repeatedly composed to solve larger routing problems. During inference, SeqDiCO uses this operator for both construction, by growing a solution through receding-horizon local reconstruction, and refinement, by replacing selected subpaths in an existing solution. Experiments on TSP and CVRP show that SeqDiCO consistently outperforms representative diffusion-based and autoregressive neural solvers under cross-scale evaluation. Trained only on 100-node instances, SeqDiCO generalizes to 10K-node problems, reducing the optimality gap by up to 51% over SOTA neural solvers, while achieving up to 10× faster inference on representative large-scale benchmarks.
Abstract: In this work, we provide a sharp theory of scaling laws for two-layer neural networks trained on a class of hierarchical multi-index targets, in a genuinely representation-limited regime. We derive exact information-theoretic scaling laws for subspace recovery and prediction error, revealing how the hierarchical features of the target are sequentially learned through a cascade of phase transitions. We further show that these optimal rates are achieved by a simple, target-agnostic spectral estimator, which can be interpreted as the small learning-rate limit of gradient descent on the first-layer weights. Once an adapted representation is identified, the readout can be learned statistically optimally, using an efficient procedure. As a consequence, we provide a unified and rigorous explanation of scaling laws, plateau phenomena, and spectral structure in shallow neural networks trained on such hierarchical targets.
PaperID: 2713, Poster
Authors:
Yang Chen, Xianqi Yu, SHAOWEI YAO, Fuyu Lv, Dan Ou, Haihong TangAbstract: Token-level distillation assumes a shared categorical support, an assumption violated when teacher and student tokenizers segment the same text differently. Existing cross-tokenizer methods sidestep this mismatch by aligning token spaces, matching likelihoods along tokenizer-specific paths, or filtering shared spans, all of which can approximate the original distributional objective or discard probability mass. We argue that decoded UTF-8 bytes are common observable events shared by every text tokenizer, while token boundaries are tokenizer-specific latent transitions. Building on this view, we introduce \emphByteDistill, which compares the conditional distribution of the next observable byte rather than aligning token identities. Its loss, \emphByte-Level Distribution Alignment (BLDA), projects each native softmax into a 256-way observable next-byte distribution inside aligned decoded chunks, marginalizing token-ending mass as a latent boundary transition along the observed tokenization path. ByteDistill consistently outperforms supervised fine-tuning and prior cross-tokenizer baselines in both off-policy and on-policy settings, with off-policy training improving GSM8K accuracy by 1.75 percentage points over the current state of the art.
PaperID: 2714, Poster
Abstract: Existing Mamba-based diffusion backbones struggle to achieve competitive performance in vision tasks, as purely sequential propagation lacks an explicit spatial inductive bias and relies on a fixed computation pattern across timesteps. We address this limitation by proposing Mixture-of-Hierarchical Experts (MoH), a Mamba-based diffusion backbone that performs timestep-conditioned routing over architectural components. MoH dynamically selects representation levels by its hierarchical expert space, adapts feature transformations, and adjusts Mamba scan depth according to the denoising state, enabling flexible modeling of spatial structure and multi-scale dependencies. This design leads to substantial performance gains. On unconditional CelebA-HQ 256 × 256, MoH reduces FID from 14.27 to 7.99, and on MS-COCO, it improves FID from 41.80 to 20.00, significantly outperforming prior Mamba-based models while improving training efficiency. These results suggest that routing over architectural components can provide an effective mechanism for improving Mamba-based diffusion backbones.
PaperID: 2715, Poster
Abstract: Most existing approaches to LLM swarm optimization are utility-driven, focusing on final task utility while leaving the processing, integration, and reliance of information within individual interactions largely implicit. To address this, we study LLM swarm interactions from an information-flow-based view and introduce behavioral probes to quantify information flow patterns. Specifically, we repurpose External Context Score (ECS) and Parametric Knowledge Score (PKS)—originally designed for hallucination analysis—to measure a node’s external-context dependence and parametric-knowledge dependence, respectively. Across controlled and realistic swarm settings, we observe consistent patterns: ECS increases with input relevance and distinguishes useful multi-source contributions, while PKS decreases as informative context accumulates. Building on this heuristic interpretation, we demonstrate how these probes can complement utility-driven structure optimization in two controlled applications. Probe-guided node pruning significantly reduces anomalous nodes, and probe-guided reasoning-path search accelerates convergence compared to utility-only baselines. Our empirical results suggest that behavioral probes can make the internal information flow of LLM swarms observable and serve as effective auxiliary signals for structure optimization.
Abstract: Representation alignment has recently emerged as an effective paradigm for accelerating Diffusion Transformer training. Despite their success, existing alignment methods typically impose a fixed supervision target or a fixed alignment granularity throughout the entire denoising trajectory, whether the guidance is provided by external vision encoders, internal self-representations, or VAE-derived features. We argue that such timestep-agnostic alignment is suboptimal because the useful granularity of representation supervision changes systematically with the signal-to-noise ratio. In high-noise regimes, diffusion models benefit more from coarse semantic and layout-level anchoring, whereas in low-noise regimes, the training signal should emphasize spatially detailed and structurally faithful refinement. This non-stationary alignment behavior creates a representational mismatch for static single-level supervisors. To address this issue, we propose Adaptive Hierarchical Prior Alignment (AHPA), a lightweight alignment framework that exploits the hierarchical representations naturally embedded in the frozen VAE encoder. Instead of using only a single compressed latent as the alignment target, AHPA extracts multi-level VAE features that provide complementary priors ranging from local geometry and spatial topology to coarse semantic layout. A timestep-conditioned Dynamic Router adaptively selects and weights these hierarchical priors along the denoising trajectory, thereby synchronizing the alignment granularity with the model's evolving training needs. Extensive experiments show that AHPA improves convergence and generation quality over baselines and incurs no additional inference cost while avoiding external encoder supervision during training.
PaperID: 2717, Poster
Authors:
Daniil Reutsky, Daniil Vladimirov, Yasin Mamedov, Georgy Perevozchikov, Egor I Ershov, Nancy Mehta, Radu TimofteAbstract: Hyperspectral reconstruction from RGB imagery offers an affordable alternative to costly hyperspectral acquisition devices. However, its accuracy is fundamentally constrained by the limited spectral resolution of standard RGB cameras. In this paper, we propose new approach that leverage the multi-camera systems integrated into most modern smartphones and enhance their capabilities by modulating their spectral response functions using carefully selected spectral filters, enabling a smartphone to capture richer spectral information. We propose a neural network able to exploit auxiliary RGB images with diverse spectral characteristics, enabling more accurate hyperspectral reconstruction compared to conventional single-image methods. Training is supported by a newly collected dataset, Doomer, which consists of multiple misaligned RGB images alongside corresponding hyperspectral data. Finally, our method generalizes across different smartphone devices through a easy to implement calibration procedure, ensuring practical applicability and scalability.
PaperID: 2718, Poster
Abstract: Reconstructing population dynamics is a central problem in the physical and data sciences. Often, the dynamics are modeled as a Wasserstein gradient flow (WGF): a curve of distributions driven by an energy functional. Though there are multiple mathematical characterizations of a WGF, the dominant algorithmic approach relies on the Jordan--Kinderlehrer--Otto (JKO) scheme. JKO-based methods are inflexible to time discretisation and require solving costly optimal transport problems. We take a residual approach, enforcing that the continuity equation conditions are satisfied via a non-negative loss function, whose minimum is the WGF. Combined with a data-fitting divergence, this gives a single global objective. This perspective unifies several existing methods, and leads us to discover a new particle-based method, \emphstitching, that is simulation-free and robust to large gaps in observation times of the dynamics. We demonstrate that the stitching method achieves state-of-the-art performance across trajectory inference benchmarks.
Authors:
Konstantinos Barmpas, Na Lee, Dimitrios Chalatsis, William Raftery, Yannis Panagakis, Dimitrios Adamos, N Laskaris, Alexandros Koliousis, Dario Farina, Stefanos ZafeiriouAbstract: Biosignals such as electroencephalography (EEG), electrocardiography (ECG), and electromyography (EMG) encode physiological activity across multiple temporal and spectral scales, yielding representations that are rich but challenging for machine learning. Foundation models trained to predict masked signal tokens have shown promise in learning generalizable biosignal representations, yet their performance depends on the tokenizer's ability to preserve high-frequency dynamics and reconstruct signals with high fidelity. We introduce NeuroRVQ, a modality-adaptive biosignal tokenizer family designed for high-fidelity signal reconstruction. To capture the full frequency spectrum, NeuroRVQ decomposes biosignals into frequency-specific representations via multi-scale temporal convolutions, each encoded into hierarchical RVQ codebooks to preserve high-frequency detail, combined with a novel phase-aware training loss that respects the circular topology of Fourier phase. By tuning the temporal resolution, number and size of temporal kernels and RVQ depth, this design adapts to the spectro-temporal characteristics of each biosignal modality. To validate that tokenizer quality drives downstream performance, we train a simple masked-token foundation model for each modality (NeuroRVQ-FM) using the corresponding NeuroRVQ tokenizer. The NeuroRVQ-FM family achieves competitive or superior downstream performance compared to existing modality-specific foundation models, demonstrating that high-fidelity tokenization is a critical factor for effective biosignal modeling.
Authors:
Chen Ling, Pei Chen, Xiangchen Guan, Jiaming Qu, Shayan Ali Akbar, Madhu Gopinathan, Erwin CornejoAbstract: Deploying language-model agents in production often requires substantial compute and human effort to tune prompts, parsers, validators, and other components of the agent pipeline. Self-evolution offers a promising alternative, but most existing frameworks assume access to frontier models that can reliably diagnose failures, propose revisions, and judge their own updates. We study whether frozen small language models (SLMs) can serve as effective self-evolving agents under resource constraints. We propose PACE (Prompt And Control Logic Evolution), a two-timescale framework that coordinates low-risk prompt refinement with higher-risk control-logic updates. PACE evolves prompts under fixed control logic until prompt-level gains saturate, then considers constrained control-logic updates that are accepted through held-out validation. Across three frozen SLM backbones ranging from 4B to 14B parameters and four controlled benchmarks, PACE achieves the best performance on all 12 backbone--benchmark combinations, improving over vanilla SLM agents by up to +9.2% relative improvement and over the stronger single-mode evolution baseline by up to +5.4% relative improvement. A \tau-bench case study further shows that PACE improves multi-turn tool-use success over vanilla and prompt-only evolution. These results suggest that reliable SLM agent self-evolution is possible without updating model weights or relying on frontier-model teachers, and that the key benefit is not any single final solver pattern but autonomous, validated discovery of task-appropriate inference strategies.
PaperID: 2721, Poster
Abstract: Inference-time steering enables pretrained diffusion models to satisfy new constraints without full retraining. However, specificity-aware generation is difficult: repelling samples from a negative reference distribution can also erode the positive distribution where the two overlap. The key challenge is to suppress negative mass while minimally distorting the positive distribution. We address this problem by formulating specificity-aware steering as a target-design problem and deriving a target distribution from an overlap-based objective. The resulting target keeps the desired reference distribution only in regions where it is sufficiently preferred over the undesired reference distribution, giving a likelihood-ratio interpretation of specificity. To sample from the corresponding time-dependent target path, we develop a Sequential Monte Carlo sampler with a variance-minimized local proposal. We further introduce a practical fixed-noise optimization procedure that requires only Jacobian--vector products with the desired and undesired score fields. Experiments on synthetic mixtures, class-contrastive generation, and text-to-image tasks show that the proposed method suppresses undesired regions more effectively, reduces mode shift, and improves sampling stability by decreasing the SMC weight collapse compared with negative-guidance baselines.
PaperID: 2722, Poster
Abstract: Communication learning is crucial for complex coordination in multi-agent reinforcement learning (MARL), yet existing methods often struggle with a fundamental dilemma between communication overhead and task performance. Typically designed without theoretical guidance, they produce unstructured and redundant communication graphs, limiting both efficiency and scalability. To address this, we propose Structural Entropy Optimized Communication (SEOC), a framework that directly optimizes inter-agent communication topology using the information-theoretic principle of Structural Entropy. SEOC guides the network to self-organize into a low-entropy state with sparse connectivity and distinct communities, enabling efficient and interpretable decentralized coordination. This is achieved through a unified mechanism that jointly enforces structural order and semantic conciseness, reducing topological complexity via a differentiable structural entropy loss while compressing messages with a node information bottleneck. Experiments on benchmarks such as the StarCraft Multi-Agent Challenge (SMACv2) demonstrate that SEOC significantly outperforms state-of-the-art methods.
PaperID: 2723, Poster
Abstract: We consider online prediction from experts, a fundamental problem in machine learning, under differential privacy constraints. Existing private algorithms achieve near-optimal regret \tildeO(\sqrtT), where T is the time horizon. However, this bound becomes suboptimal when L^\star, the cumulative loss of the best expert, is significantly less than T. In this work, we present the first differentially private algorithm with regret \tildeO(\sqrtL^\star), offering a substantial improvement in the low-loss regime. Our regret bound matches the non-private lower bound up to poly-logarithmic factors, demonstrating that privacy incurs only a small cost in this setting.
PaperID: 2724, Poster
Abstract: Vision-language-action (VLA) models have made rapid progress in robotic manipulation, yet the policy acts with no sense of the at which its own execution unfolds—whether it is advancing, stalling, or has quietly gone wrong—especially for long-horizon tasks. Recent methods supply progress through external reward models or as an auxiliary prediction inside the policy, but the signal runs alongside action generation rather than actively shaping it. Moreover, progress is inherently tied to the terminal state, making estimation fragile when the policy lacks an internal awareness of where the task should end. We propose , which actively internalizes task progress as a predictive interface bridging commitment and execution, equipping the policy with a calibrated sense of the , a predicted terminal representation from observation and instruction, from which progress emerges as a geometric projection. Building on this anchor, the model perceives a compact past–present–future progress state capturing . These predicted progress further actively conditions the action decoder to guide the prediction, while the gap between committed and realized increments provides a built-in consistency check for lightweight failure detection. Experiments across LIBERO, SimplerEnv, and real-world single-arm and bimanual platforms show that PACE achieves 98.7% on LIBERO and 74% on real-world manipulation, with consistent gains especially on long-horizon tasks. Robustness and stable generalization across diverse scenarios further confirm the advantage of internalizing progress within the policy, validating a calibrated sense of execution pace as an effective paradigm for long-horizon robot control.
Abstract: Transformers update token representations through multi-head attention and residual connections as X \leftarrow X + \sum_i P^(i)XW_V_iW_o_i, where P^(i) is the softmax attention matrix in head i. We propose replacing a subset of P^(i)'s with the Laplacian I - P^(i), giving X \leftarrow X + \sum_i \in \mathcalA P^(i)XW_V_iW_o_i + \sum_i \in \mathcalL (I - P^(i))XW_V_iW_o_i. Our proposal has two motivations. First, it allows attention heads to update the mean of token representations, while Laplacian heads can directly control within-sequence variance. Second, if tokens are viewed as nodes in a graph with edge weights \(P^(i)\), then \(I - P^(i)\) is the corresponding graph Laplacian, and the update can be interpreted as one step of heat diffusion on the graph. We show that this simple modification improves performance across supervised learning, language modeling, and self-supervised learning tasks. To investigate why, we examine the token representations learned with and without Laplacian heads. In supervised learning, Laplacian heads collapse token representations within the same sequence and align the sequence means with the geometry of Neural Collapse. In language modeling, they increase the separability of token representations that share the same next-token prediction. In self-supervised learning, they produce token representations whose principal components are better suited for segmentation. Across modalities, they also lead to faster-decaying spectra, indicating stronger token smoothing. Overall, our findings challenge the prevailing view that token oversmoothing is inherently harmful, showing instead that certain forms of smoothing can be beneficial.
Abstract: Some frontier AI developers aim to align language models to a Model Spec or Constitution that describes the intended model behavior. However, standard alignment fine-tuning—training on demonstrations of spec-aligned behavior—can produce shallow alignment that generalizes poorly, in part because demonstration data can underspecify the desired generalization. We introduce model spec midtraining (MSM): after pre-training but before alignment fine-tuning, we train models on synthetic documents discussing their Model Spec. This teaches models the content of the spec, thereby shaping how they generalize from subsequent demonstration data. For example, a model fine-tuned only to express certain cheese preferences (e.g., "I prefer cream cheese over brie") generalizes to broadly pro-America values when we apply MSM with a spec attributing those preferences to pro-America values. Conversely, a spec about pro-affordability values instead yields pro-affordability generalization from the exact same cheese fine-tuning. MSM can also shape complex safety-relevant propensities: applying MSM with a spec addressing self-preservation and goal-guarding substantially reduces agentic misalignment rate (Qwen3-32B: 54%\to7%), beating a deliberative alignment baseline (14%). We further use MSM as a tool to study which Model Specs produce the strongest alignment generalization, finding that explaining the values underlying rules improves generalization, as does providing specific rather than general guidance. Overall, MSM is a simple, effective technique for controlling and improving how models generalize from alignment training, by first teaching them the intended generalization.
Abstract: Recent progress in task-optimized neural networks has established encoding models as a powerful tool for predicting brain responses to naturalistic stimuli, yet most existing approaches rely on unimodal representations. The emergence of omni-modal foundation models and rich multimodal neural datasets enables shared encoding models that jointly integrate visual, auditory, and linguistic information across subjects. We introduce MIRAGE, a brain encoding framework that predicts whole-brain fMRI responses to naturalistic audiovisual stimuli with paired transcripts. MIRAGE extracts representations from a single pretrained omni-modal backbone through three modality-specific cross-attention modules whose latent queries adaptively aggregate features across the backbone's 48 layers, and combines them through a transformer-based brain encoder and a subject-specific linear head over the cortical parcels. On the Algonauts benchmark, MIRAGE achieves state-of-the-art results on the out-of-distribution dataset. Controlled comparisons show that native multimodal fusion (features taken from a single jointly trained model) consistently outperforms post-hoc fusion of independently extracted unimodal streams, across architectural levels and backbones. Beyond predictive accuracy, the learned attention weights are directly inspectable: each modality's gating module discovers a distinct depth profile over the backbone, and each modality traces a distinct, anatomically structured pattern across cortex. Together, these results propose adaptive layer-wise aggregation of natively multimodal features as a more generalizable, interpretable, and accurate approach for whole-brain encoding.
PaperID: 2728, Poster
Abstract: Large Reasoning Models (LRMs) have shown remarkable reasoning capabilities, yet they still suffer from overthinking, generating redundant reasoning steps which incur substantial token consumption. Existing methods, such as suppressing reflective keywords or forcing shorter reasoning lengths, attempt to mitigate this issue but inevitably truncate necessary steps and induce underthinking, thereby compromising performance. To address this dilemma, we investigate the latent representations and observe that efficient reasoning steps naturally cluster into a concentrated region in latent space, while those deviating from this region tend to produce verbose sequences. To leverage this, we keep reasoning focused within this region via a quadratic program which projects deviating hidden states back into the region. Then we propose a novel training-free framework to achieve efficient reasoning that reduces token generation costs without sacrificing performance. Extensive experiments conducted on four models ranging from 1.5B to 14B, and across six benchmarks in math reasoning, coding, and scientific QA, validate the effectiveness of our method, up to a 13.4% improvement in accuracy while reducing generated tokens by 11.8% to 52.8%. Code and models will be publicly available.
PaperID: 2729, Poster
Abstract: Bilevel optimization provides a general framework for machine learning problems where an outer objective depends implicitly on the solution of an inner problem. While significant attention has been devoted to its algorithmic aspects, its generalization properties remain far less understood. In this paper, we address this gap by studying the case where the inner solution is approximated by an overparameterized two-layer ReLU neural network trained by gradient flow. Under the Neural Tangent Kernel (NTK) regime, we derive non-asymptotic generalization bounds that characterize how statistical accuracy depends jointly on the inner and outer sample sizes, as well as on the regularity of the target function and the complexity of the hypothesis space. We further establish convergence guarantees for a gradient-based bilevel algorithm, where the inner level is approximately solved by applying an early stopping strategy to the gradient flow training. Finally, we complement our theoretical analysis with numerical experiments showing that the bilevel structure induces an initialization-dependent implicit bias that enables feature learning at the inner level and yields performance improvements beyond NTK predictions.
PaperID: 2730, Poster
Abstract: The Platonic Representation Hypothesis suggests that vision and language models converge toward a shared latent structure, yet the unsupervised alignment of unpaired modalities remains a fundamental challenge. Current methods rely on direct comparisons of intra-modality similarity kernels, which are often incommensurate due to disparate scales and representation densities. This mismatch yields ill-conditioned optimization landscapes, hindering stable alignment. To address this, we propose RankAlign, a framework that independently transforms each within-modality kernel into a rank-based representation of relative neighbor relationships. By shifting from absolute similarity values to ordinal relational geometry, RankAlign ensures structural consistency while remaining robust to modality-specific distortions. We show that rank-based transformation reshapes similarity distributions, mitigating representation collapse and providing a more discriminative alignment signal. Experiments demonstrate that RankAlign is noise-resilient and overcome existing kernel-based alignment method across diverse datasets, including a medical imaging dataset where subtle class differences typically limit existing approaches alignment approaches.
Abstract: Pre-trained visual models have become fundamental in computer vision, but they face challenges in continual learning scenarios where data and tasks evolve over time. Low-Rank Adaptation (LoRA) offers efficient fine-tuning capabilities but remains limited for such dynamic environments. Standard LoRA cannot distinguish important subspaces, causing critical knowledge to be overwritten in sequential training. Existing approaches address this by dynamically expanding the set of LoRA adapters—either maintaining a growing pool of task-specific modules or merging new adapters into prior ones—at the cost of unbounded parameter growth or increasing inference complexity. We propose Continual Low-Rank Adaptation (C-LoRA), a method that enables a single, shared LoRA adapter to handle sequential tasks without catastrophic forgetting—without requiring any module selection or fusion at inference. The core of C-LoRA is a learnable routing matrix \boldsymbol\mathcalR that explicitly controls how each rank-one subspace contributes to the weight update. This matrix is decomposed into a stability component (\boldsymbol\mathcalR _ \textbase), which preserves knowledge from prior tasks, and a plasticity component (\boldsymbol\mathcalR _ \delta), which drives adaptation to the current task—providing direct control over the stability-plasticity trade-off. We analyze how \boldsymbol\mathcalR governs gradient flow during sequential training, and demonstrate competitive performance across multiple benchmarks.
PaperID: 2732, Poster
Abstract: Sequential reinforcement learning (RL) solvers and global diffusion model (DM) solvers for neural combinatorial optimization exhibit complementary failure modes under an optimization-regret view. The former enjoys small marginal regret in the early construction stage, but suffers from horizon-wise compounding errors with super-linear regret growth; the latter avoids horizon compounding but incurs linear or sublinear regret w.r.t. the dimension of the remaining unsolved subspace. We propose \underlineHybrid Neural Solver for \underlineCombinatorial \underlineOptimization (HyCO), a hybrid inference algorithm that constructs a solution prefix with an RL solver and adaptively switches to a conditional DM to complete the remaining decisions. To characterize why such hybridization helps, when to trigger the handover, and how to realize it in practice, we first develop a unified error-scaling theoretical framework and prove that, under explicit error-scaling assumptions, i) the hybrid structure achieves strictly lower expected regret than either backbone alone, and ii) there exists a unique optimal trigger step that minimizes the hybrid regret. We then design a lightweight adaptive trigger that combines policy entropy and RL–DM disagreement to detect trajectory-level signals of the regime shift as a practical proxy, since the optimal trigger step is defined at the expected-regret level and is not directly computable on individual trajectories. Experimental results on diverse benchmarks demonstrate that HyCO achieves consistent improvements over both backbones and support the empirical effectiveness of adaptive triggering.
PaperID: 2733, Poster
Authors:
Lei Wu, Jiashuai Liu, Di Zhang, Zhangpeng Gong, Yingkang Zhan, Yi Niu, Jiusong Ge, Chunze Yang, Kai Yi, Mireia Crispin-Ortuzar, Chen Li, Zeyu GaoAbstract: Due to the gigapixel-scale nature of whole-slide images (WSIs), weakly supervised WSI analysis is commonly formulated as a multiple instance learning (MIL) problem, where patch-level features are aggregated into slide-level representations. However, diagnostic and prognostic evidence often arises from spatially coherent tumor microenvironment regions and their interactions, rather than isolated patches alone. Existing patch-level or static region-based methods usually overlook how tissue regions should be adaptively formed and subsequently evolved through microenvironment interactions across heterogeneous boundaries. In this paper, we propose Concept-Guided Tumor Microenvironment Evolution (TMEvolve), a reaction-diffusion-inspired framework that models WSIs as latent tumor microenvironment fields over discrete patch graphs. TMEvolve instantiates this view as a learnable graph-discretized evolution process over patch neighborhoods. It first forms adaptive soft tissue regions as coherent microenvironment units, then performs pseudo-time evolution through two complementary local dynamics: intra-region diffusion, which stabilizes latent states within coherent tissue compartments, and concept-guided boundary flux, which propagates visual feature signals and language-derived concept signals across heterogeneous region interfaces. The evolved microenvironment regions are finally aggregated for slide-level prediction. We evaluate TMEvolve on six datasets across three weakly supervised WSI tasks: survival prediction, gene expression prediction, and histological subtype classification. TMEvolve consistently improves over representative MIL methods, pathology foundation models, and concept-guided baselines. Ablation studies and visualizations further support the effectiveness and interpretability of TMEvolve, highlighting the value of dynamic region modeling and boundary interaction.
PaperID: 2734, Poster
Authors:
Victor Ye Dong, Reid Pryzant, Yi Liu, Jian JiaoAbstract: Large Language Models (LLMs) are increasingly deployed in agentic contexts, where the model relies on a system prompt to use tools and complete tasks. Implicit in these agentic settings are operational requirements: using the right tools, keeping prompts at reasonable length, achieving parsimonious solution paths, and complying with safety and formatting policies. For many practitioners, assembling domain-specific supervised data to post-train LLMs to satisfy such requirements is infeasible. In this paper, we introduce CAPO (Constraint-Aware Prompt Optimization), an \emphexplicit threshold-constrained prompt-optimization algorithm based on primal--dual updates. CAPO employs a primal--dual framework to optimize system prompts under explicit constraints via Lagrangian relaxation. Our results across three agentic benchmarks show that CAPO more reliably reaches empirically feasible operating points while improving agentic performance. We also demonstrate that the algorithm generalizes beyond agentic use cases, achieving strong performance on assistant-style evaluations that require satisfying output-format and safety/privacy constraints. Finally, we show that CAPO's rewrite policy can be amortized into DCAPO, a feedback-aware trainable rewriter that matches CAPO accuracy while improving the constrained trade-off.
PaperID: 2735, Poster
Abstract: Latent confounding poses a significant obstacle to identifying the causal effect of a treatment on an outcome. To address this, many existing studies leverage proxies of the latent confounder to indirectly adjust for the confounding bias. However, they impose specific structural constraints on the proxies and typically require multiple proxies. In this paper, focusing on the latent variable linear non-Gaussian acyclic model (lvLiNGAM), we propose a causal effect identification procedure requiring only a single agnostic proxy. Crucially, the term "agnostic" means that the causal connections between the proxy and the treatment-outcome pair can be arbitrary and are not required to be known a priori. This structural complexity precludes identifying the causal effect via a simple closed-form formula. Consequently, our identification procedure is designed to first derive candidate solutions from cumulants and then isolate the valid solution by examining certain independence relationships. We present a series of new theoretical results, which collectively establish the soundness of our identification procedure: given the observational population distribution, it correctly identifies the true causal effect when identifiable, and correctly reports unidentifiability otherwise. Finally, we empirically validate our theoretical results.
PaperID: 2736, Poster
Abstract: Multimodal large language models (MLLMs) have made rapid progress in general visual understanding, yet they still struggle with fine-grained instance-level consistency verification, where the goal is to determine whether two images depict the exact same physical object under varying viewpoints, illumination, and backgrounds. This task requires both a global understanding of the target instance and a precise examination of subtle local evidence, since semantically similar objects may differ only in small discriminative regions. To address this challenge, we propose LocalAgent, a collaborative agentic framework that integrates the complementary strengths of large and small models. Specifically, the MLLM first reasons from a global instance-level perspective to identify suspicious local regions that may determine the final verification result. These accurately localized regions are then delegated to a lightweight Local Evidence Verifier (LEV) toolkit for fine-grained appearance and geometric matching. The LEV toolkit consists of Local Appearance Consistency (LAC), which provides illumination- and color-robust similarity estimation, and Local Geometric Correspondence (LGC), which establishes reliable spatial correspondence under viewpoint changes. Rather than directly replacing the MLLM's reasoning, LEV returns structured similarity scores and comparative evidence in a form that the MLLM can effectively interpret for final decision-making. This design enables a coordinated large-small model collaboration: the MLLM contributes global semantic reasoning and region-level suspicion localization, while specialized local verifiers provide precise, invariant, and interpretable evidence for ambiguous cases. We further introduce an answer-dominant reinforcement fine-tuning (RFT) strategy that enables the policy to learn when verification is necessary. Experiments show that LocalAgent achieves a +4.8% average improvement over direct-inference baselines, effectively narrowing the gap with closed-source models, while preserving the general multimodal capabilities of the underlying MLLM.
PaperID: 2737, Poster
Abstract: Reward serves as the primary learning signal in reinforcement learning (RL). However, while reward magnitudes are typically held fixed throughout training, their temporal modulation remains underexplored. In this paper, we propose reward inflation, a gradual scaling of rewards over the course of training, and show that it can act as a healthy stimulus for RL. Theoretically, reward inflation induces an implicit recency weighting that upweights recent transitions during policy updates, enabling faster adaptation. We further show that, by sustaining gradient signals as the policy saturates, reward inflation suppresses the emergence of dormant neurons and helps preserve plasticity. Empirical results on ALE games and MuJoCo tasks corroborate these findings, showing that an appropriate level of reward inflation benefits a broad range of tasks. Finally, we introduce Fed, an adaptive variant that adjusts the inflation level on the fly, and find that it often improves upon fixed inflation.
PaperID: 2738, Poster
Abstract: Proximal Policy Optimization (PPO) is value-driven only after experience has been collected: Generalized Advantage Estimation reinforces sampled actions according to their bootstrapped advantage, but the exploration process that produced those samples is still governed mainly by policy-space noise and entropy regularization. As a result, PPO does not explicitly ask, before sampling, which uncertain transitions would most improve its advantage estimates. We argue that PPO exploration should be formulated as value-priced uncertainty: uncertainty should be explored when it is both large and able to change the advantage. We propose VALU (Value-Aware Latent Uncertainty), a family of PPO-compatible exploration bonuses of the form b(s,a) = P(s,a) U(s,a). The price P(s,a) = \lVert \nabla_z' \tilde V_\psi(\hat z') \rVert measures how strongly uncertainty in the predicted next latent can perturb the advantage, while \(U(s,a)\) is an uncertainty quantity instantiated by non-parametric effective counts or learned prediction errors such as RND. Our theory derives the value-pricing term from a unified advantage-uncertainty view: an uncertainty coordinate induces an optimistic advantage correction through a dual norm. This yields count-based UCB and MCR members, corresponding respectively to optimism over latent-transition uncertainty and one-sample shrinkage of the optimistic advantage envelope. Across six standard MuJoCo continuous-control tasks, VALU substantially improves PPO's sample efficiency and final return, and achieves the best PPO-compatible performance on most tasks while remaining competitive on the others, compared with strong exploration baselines including RND, ICM, RE3, VCSE, and OPPO. Ablations over uncertainty estimators, schedules, latent encoders, and bonus components show that the gains come from pricing uncertainty by local value sensitivity rather than from any single implementation choice.
PaperID: 2739, Poster
Abstract: Conformal prediction offers a distribution-free coverage guarantee, making it especially attractive for clinical applications. Standard conformal prediction, however, provides such guarantees only at the population level, and its prediction sets can exhibit coverage disparities across clinically important subgroups. A natural remedy is to calibrate within predefined groups. However, this can require access to sensitive subgroup attributes and is prone to a worst-group bottleneck: protecting the most difficult subgroup can inflate prediction sets for all, increasing cognitive burden on decision makers. To this end, we propose Stochastic Grouping Conformal Prediction (SGCP), a conformal framework for subgroup-reliable uncertainty quantification. It learns a stochastic grouping map that allows each sample to draw calibration information from others with similar calibration behavior, yielding a local score law that boosts reliability across subpopulations. We prove that SGCP retains the standard coverage guarantee. Experiments on synthetic and real-world benchmarks show that it consistently reduces subgroup coverage gaps while achieving smaller or comparable prediction set sizes relative to existing baselines.
PaperID: 2740, Poster
Abstract: Text-based image editing has recently been reinterpreted in large multimodal transformers as conditional generation, where source image tokens are concatenated with text and noise tokens. While effective, this design incurs substantial computational overhead in attention layers. To mitigate this inefficiency, we propose \em TokenDrop, an efficient editing framework that shift the computational burden from dense source-token conditioning to a lightweight regularized sampling update. By transferring the influence of source tokens into a closed-form sampling update, the method preserves source consistency with negligible cost. To support this regularized dynamics, we introduce a deviation-guided adaptive masking strategy that selectively drops redundant tokens while maintaining editing fidelity. Across FluxKontext and Qwen-Image-Edit, our training-free method achieves an average 22.4% improvement in inference speed on PIEBench, while better preserving non-edited regions. The method delivers up to 1.8× speedup at 1024^2 resolution and 2× speedup at 2048^2 resolution. The code will also be released publicly.
PaperID: 2741, Poster
Abstract: We study cooperative multi-agent reinforcement learning (MARL) problems in which agents evolve under decoupled noisy dynamics but are coupled via a population objective. A canonical example is a mobility operator that orchestrates a fleet of vehicles in an urban environment to meet customer demand. The stochastic and time-varying nature of these problems calls for learning approaches, yet existing MARL algorithms do not scale: learning assignment and routing jointly over the full state-action space becomes prohibitive as the number of agents grows. By studying the problem as a control problem over probability measures, we prove a separation principle: the population-level cost-to-go function is upper-bounded by an optimal transport problem whose transportation cost is the target-conditioned single-agent cost-to-go function. This result simplifies the multi-agent problem to learning a single-agent policy and coordinating the agents via optimal transport. Thus, we introduce SALT (Separation-based Assignment and Learning via optimal Transport) a family of algorithms that learn a target-conditioned single-agent policy using any standard RL algorithm, and then coordinate the agents via optimal transport. Because learning occurs at the single-agent level, SALT scales to fleet sizes where MARL baselines become intractable and transfers without retraining to fleet sizes unseen during training. In a vehicle-delivery experiment in south Manhattan’s road network, where MARL baselines are computationally prohibitive, SALT improves travel times by up to 20% over a shortest-path routing heuristic and transfers without retraining to different fleet sizes. In stochastic grid worlds, SALT reduces average travel times by more than 40% relative to MARL baselines and scales to population sizes where these methods fail.
Authors: Sergey Alekseev
Abstract: We study signal propagation at initialization in transformers through the averaged partial Jacobian norm (APJN), a measure of gradient amplification across layers. We extend APJN analysis to transformers with bidirectional attention and permutation-symmetric input token configurations by deriving recurrence relations for activation statistics and APJNs across layers. Our theory predicts how attention modifies the asymptotic behavior of the APJN at large depth and matches APJNs measured in deep vision transformers. The criticality picture known from residual networks carries over to transformers: the pre-LayerNorm architecture exhibits power-law APJN growth, whereas transformers with LayerNorm replaced by elementwise tanh-like nonlinearities have stretched-exponential APJN growth, indicating that the latter are subcritical. Applied to Dynamic Tanh (DyT) and Dynamic erf (Derf) transformers, the theory explains why these architectures can be more sensitive to initialization and optimization choices and require careful tuning for stable training.
Authors:
Shenzhi Wang, Shixuan Liu, Jing Zhou, Chang Gao, Xiong-Hui Chen, Binghai Wang, An Yang, Shiji Song, Bowen Yu, Gao Huang, Junyang LinAbstract: Vision-language models (VLMs) show strong multimodal capabilities but still struggle with fine-grained vision-language reasoning. We find that long chain-of-thought (CoT) reasoning exposes diverse failure modes, including perception, reasoning, knowledge, and hallucination errors, which can compound across intermediate steps. However, most existing data for reinforcement learning with verifiable rewards (RLVR) does not involve complex reasoning chains grounded in visual evidence throughout, leaving these failure modes understressed during training. We therefore propose HopChain, a scalable framework that constructs an out-of-distribution (OOD) proxy task: OOD only in query format, while strengthening foundational abilities shared with downstream benchmarks to support generalizable gains. Concretely, HopChain synthesizes multi-hop vision-language reasoning chains, each a logically dependent chain of instance-grounded hops where earlier hops establish the instances, sets, or conditions needed for later hops, ending in an unambiguous numerical answer for verifiable rewards. In experiments, we train Qwen3.5-35B-A3B and Qwen3.5-397B-A17B under two RLVR settings (with and without HopChain's multi-hop data) and compare them across 24 benchmarks spanning STEM and Puzzle, General VQA, Text Recognition and Document Understanding, and Video Understanding. Although the multi-hop data is not designed for any specific benchmark, it improves 20 of 24 benchmarks on both models, indicating broad and generalizable gains. Consistently, replacing full chained queries with half-multi-hop or single-hop variants reduces the average score across five representative benchmarks from 70.4 to 66.7 and 64.3, respectively. Improvements are particularly substantial in long-CoT and ultra-long-CoT regimes, peaking at more than 50 accuracy points in the ultra-long-CoT regime. These results show that OOD proxy tasks for long-chain visual reasoning are a scalable RLVR supervision source for generalizable VLM reasoning.
PaperID: 2744, Poster
Authors: Jianping Huang, Xiang Liu, Feng Shan
Abstract: Model cascades reduce LLM inference costs by sequentially querying candidate models and stopping once a satisfactory response is generated. However, existing cascades largely optimize model escalation under fixed evaluation schemes, overlooking that optimal stopping fundamentally depends on the reliability of intermediate evaluations. We formalize this as an online sequential search problem with tunable-fidelity probes, in which the system jointly decides which model to query, at what fidelity to evaluate it, and when to stop. Under a fixed evaluation protocol, higher fidelity incurs greater evaluation cost but reduces stochastic observation noise through an unknown arm-specific noise function. For the oracle setting, we derive a generalized index policy that tightly couples fidelity selection with stage-wise stopping. For the online setting, we propose DialBandit, which learns fidelity-dependent noise functions from unbiased replicate-based labels and performs uncertainty-aware planning over probing order, evaluation fidelity, and stopping. We prove a finite-time sublinear regret bound and show that DialBandit consistently improves net utility over strong baselines in semi-empirical environments based on repeated judging and calibrated noisy LLM evaluations.
PaperID: 2745, Poster
Abstract: We introduce COMONet (Convex-Concave and Monotonicity-Constrained Neural Networks), a unified architecture for enforcing variable-wise shape constraints in neural networks. COMONet addresses the limitations of prior methods that either support only a restricted subset of monotonicity and curvature constraints or enforce them through penalties without strict architectural guarantees. Our framework assigns each input variable to one of eight shape regimes, covering monotonicity, convexity, concavity, and their combinations. This is achieved through a partially connected architecture that routes each variable to specialized units equipped with sign-constrained weights and constraint-preserving activation functions. We also provide theoretical guarantees showing that COMONet satisfies the prescribed variable-wise shape constraints by construction. Experiments on synthetic and real-world datasets demonstrate that COMONet achieves competitive performance and remains robust to noise while preserving its architectural guarantees. COMONet thus offers a practical and principled framework for incorporating domain knowledge as shape constraints into neural network training.
PaperID: 2746, Poster
Abstract: Online decision-making faces a tension between efficient inference and long-horizon reasoning. Value-based methods rely on short-horizon bootstrapping that limits the propagation of long-term information, while trajectory optimization methods plan explicitly but must solve a high-dimensional optimization problem at every decision step. We propose Generative Trajectory Planning (GTP), a model-based framework that addresses this tension by learning a generative model as a proposal distribution over action sequences. Instead of optimizing sequences from scratch, GTP samples candidate sequences from a diffusion prior and refines them using learned dynamics, reward, and value models, enabling efficient trajectory-level reasoning. Across continuous control and long-horizon manipulation tasks, GTP matches or exceeds state-of-the-art model-free, diffusion-policy, and model-based baselines, and remains stable in regimes where strong baselines degrade. Our results suggest that generative models are best understood not as policies, but as structured proposal distributions for planning, offering a new perspective on how to integrate generative modeling with model-based decision-making.
Abstract: Evaluating true metacognition in Large Language Models (LLMs) is difficult due to biases and heuristics. This paper presents a framework to measure and enhance LLM metacognition while controlling for these biases. A measurement method using the d_\rmtype2' metric is systematized to isolate metacognitive ability. The Evolution Strategy for Metacognitive Alignment (ESMA) is proposed, demonstrating robust generalization across unseen datasets, languages, and newly acquired knowledge. Finally, parameter analysis reveals that these improvements are driven by a sparse set of parameters, offering new pathways for targeted metacognitive optimization.
Authors: Jonathan Cederlund, Axel Berg, Durmus Alp Emre Acar, Chuteng Zhou, Pontus Giselsson
Abstract: Visual Autoregressive (VAR) models have recently demonstrated impressive image generation quality while maintaining low latency. However, they suffer from severe KV-cache memory constraints, often requiring gigabytes of memory per generated image. We introduce HeatKV, a novel compression method that adapts cache allocation in each head based on its attention to previously generated scales. Using a small offline calibration set, the attention heads are ranked according to their attention scores over prior scales. Based on this ranking, we construct a static pruning schedule tailored to a given memory budget. Applied to the Infinity-2B model, HeatKV achieves 2 × higher compression ratio in memory allocation for KV cache compared to existing methods, while maintaining similar or better image fidelity, prompt alignment and human perception score. Our method achieves a new state-of-the-art (SOTA) for VAR model KV-cache compression, showcasing the effectiveness of fine-grained, head-specific cache allocation.
PaperID: 2749, Poster
Abstract: Reinforcement learning (RL) post-training has increasingly demonstrated strong ability to elicit reasoning behaviors in large language models (LLMs). For training efficiency, rollouts are typically generated in an off-policy manner using an older sampling policy and then used to update the current target policy. To correct the resulting discrepancy between the sampling and target policies, most existing RL objectives rely on a token-level importance sampling ratio, primarily due to its computational simplicity and numerical stability. However, we observe that token-level correction often leads to unstable training dynamics when the degree of off-policyness is large. In this paper, we revisit LLM policy optimization under off-policy conditions and show that the theoretically rigorous correction term is the prefix importance ratio, and that relaxing it to a token-level approximation can induce instability in RL post-training. To stabilize LLM optimization under large off-policy drift, we propose a simple yet effective objective, Minimum Prefix Ratio (MinPRO). MinPRO replaces the unstable cumulative prefix ratio with a non-cumulative surrogate based on the minimum token-level ratio observed in the preceding prefix. Extensive experiments on both dense and mixture-of-experts LLMs, across multiple mathematical reasoning benchmarks, demonstrate that MinPRO substantially improves training stability and peak performance in off-policy regimes.
PaperID: 2750, Poster
Authors: Qinyi Zhang, Duanyu Feng, Yangshuai Wang, Hao Wang
Abstract: The Allen--Cahn equation is one of the standard phase-field models for mean-curvature flow (MCF), parameterised by the diffuse-interface width \varepsilon. Practical simulations span a working range of \varepsilon, so a single model covering this range is needed. PDE-specific neural operators are accurate but must be retrained at every \varepsilon; general neural operators cover \varepsilon in one model but lose the phase-field structure of the target. We introduce Interface-Augmented Diffusion-Reaction (IADR), a neural operator that addresses both shortcomings: a target network with a diffusion-and-reaction structure preserves the phase-field structure of the Allen--Cahn equation, while a hypernetwork that maps \varepsilon to the weights of this target network amortises the dependence on \varepsilon across the working range. A single IADR matches the per-\varepsilon specialists on in-distribution data, retains their cross-shape generalisation under both \varepsilon- and geometry-out-of-distribution shifts, and runs at the per-\varepsilon specialist's per-step inference cost. We further establish a K-step rollout error bound that aligns with the long-horizon behaviour observed in experiments.
Abstract: Vision foundation models are increasingly moving beyond 2D to volumetric domains such as 3D medical imaging, where unified pretraining across different imaging modalities (i.e. CT, MRI, and PET) could provide foundational models for diverse clinical tasks. However, training such models requires mixing heterogeneous imaging domains, and current mixture strategies remain largely heuristic. In this work, we observe that different medical imaging domains scale at variable rates during pretraining, and knowledge transfer between domains is strongly asymmetric: training on one domain can substantially improve another, but the reverse may be much weaker. Interestingly, both MAE reconstruction loss and cross-domain transfer follow predictable power-law trends with domain-specific behaviors. Motivated by these findings, we formulate data allocation as a scaling-law optimization problem. The derived allocations reveal an interpretable hub-and-island structure: highly transferable domains emerge as hubs that benefit many others and deserve strategic allocation, while isolated domains act as islands requiring direct investment. Empirically, transfer-aware allocation outperforms data-proportional sampling by up to 58% and generalizes well to unseen budgets with r=0.989. Downstream validation on disease classification and organ/lesion segmentation further confirms that the derived transfer-aware mixtures provide stronger pretrained representations for clinical 3D medical imaging tasks.
PaperID: 2752, Poster
Abstract: Scientific images are often governed less by exact pixel fidelity than by preservation of microstructure: morphology, connectivity, anisotropy, interface complexity, phase organization, and characteristic length scales. Yet modern image-learning objectives, including masked image modeling and reconstruction losses, supervise models in pixel, token, or latent feature space rather than in the space of scientific observables. We introduce MICROFIELD, a supervision framework that represents a scientific image by a field of microstructure descriptors computed over spatial regions. A descriptor field can be induced on a uniform grid, a quadtree, an octree, a learned partition, or a hierarchy produced by a high-resolution image model. Rather than proposing a new masked autoencoder architecture, MICROFIELD contributes an objective-level mechanism: it trains models to preserve local descriptors, cross-scale descriptor transitions, and level-wise descriptor distributions. Because descriptor fields are defined over region systems rather than model architectures, MICROFIELD applies to uniform grids, adaptive partitions, and hierarchical models alike. The framework is compatible with self-supervised pretraining, supervised learning, and reconstruction; in this paper, we focus on masked self-supervised pretraining as the primary instantiation. MICROFIELD provides a practical path toward physics-inspired representation learning for scientific images without requiring explicit governing equations or domain labels.
Abstract: Standard autoregressive video generation algorithms based on Diffusion and Flow Matching rely on rigid training objectives and static sampling schedules, limiting inference procedures from adapting to the data. We introduce Equilibrium Forcing (EqF), a simplified objective for video denoising generative models without noise level conditioning. EqF turns inference into a gradient-based optimization problem with modular training- and inference-time designs that decouple learning the denoising field from sampling. By adapting the structure of the gradient flow during sampling to propose data-dependent step sizes, EqF improves video quality and consistency on challenging autoregressive video generation benchmarks. Extensive analysis elucidates exactly how removing the noise level conditioning enables EqF's data-dependent inference properties to surpass the performance of standard noise level-conditional denoising video methods. Project page: https://anonequilibriumforcing.github.io/
PaperID: 2754, Poster
Abstract: Multi-objective reinforcement learning (MORL) aims at optimising several, often conflicting goals to improve the flexibility and reliability of RL in practical tasks. This is typically achieved by finding a set of diverse, non-dominated policies that form a Pareto front in the performance space. However, constructing such a policy set remains computationally demanding, as it requires finding not a single optimal policy, but a set of policies to cover a wide range of trade-offs among the objectives. We introduce LLE-MORL, an approach that traces policy manifolds by utilising the local relationship between the high-dimensional policy parameter space and the performance space. This structured representation enables an efficient search within contiguous solution domains via locally linear extrapolation, allowing for the rapid generation of high-quality solutions without extensive retraining. We also provide a theoretical analysis of the policy manifold structure in order to specify the applicability of the method. Experiments across diverse continuous control domains demonstrate that LLE-MORL consistently achieves higher Pareto front quality and efficiency than state-of-the-art approaches.
Abstract: Internal activations of diffusion models encode rich semantic information, but interpreting such representations remains challenging. While Sparse Autoencoders (SAEs) have shown promise in disentangling latent representations, existing SAE-based methods for diffusion model understanding rely on unsupervised approaches that fail to align sparse features with human-understandable concepts. This limits their ability to provide reliable semantic control over generated images. We introduce CASL (Concept-Aligned Sparse Latents), a supervised framework that aligns sparse latent dimensions of U-Net-based diffusion models with semantic concepts. We focus on unconditional diffusion models, where the h-space encodes stable and unconfounded semantic structure, providing a controlled setting for mechanistic interpretability analysis. CASL first trains an SAE on frozen U-Net activations to obtain disentangled latent representations, and then learns a lightweight linear mapping that associates each concept with a small set of relevant latent dimensions. To validate the semantic meaning of these aligned directions, we propose CASL-Steer, a controlled latent intervention that shifts activations along the learned concept axis. Unlike editing methods, CASL-Steer is used solely as a causal probe to reveal how concept-aligned latents influence generated content. We further introduce the Editing Precision Ratio (EPR), a metric that jointly measures concept specificity and the preservation of unrelated attributes. Experiments on CelebA-HQ, FFHQ, LSUN-Church, and AFHQ-Dog demonstrate that concept-aligned sparse directions yield more precise and disentangled causal effects than existing activation-space editing methods. Classification probing further confirms that a compact subset of aligned latent units carries strong concept-specific signals. These results validate supervised alignment as an effective approach to interpreting diffusion model representations.
Abstract: On-policy self-distillation, where a student is pulled toward a copy of itself conditioned on privileged context (e.g., a verified solution or feedback), offers a promising direction for advancing reasoning capability without a stronger external teacher. Yet in math reasoning the gains are inconsistent, even when the same approach succeeds elsewhere. A pointwise mutual information analysis traces the failure to the privileged context itself: it inflates the teacher's confidence on tokens already implied by the solution (structural connectives, verifiable claims) and deflates it on deliberation tokens ( ) that drive multi-step search. We propose Anti-Self-Distillation (AntiSD), which ascends a divergence between student and teacher rather than descending it: this reverses the per-token sign and yields a naturally bounded advantage in one step. An entropy-triggered gate disables the term once the teacher entropy collapses, completing a drop-in replacement for default self-distillation. Across five models from 4B to 30B parameters on math reasoning benchmarks, AntiSD reaches the GRPO baseline's accuracy in 2 to 10× fewer training steps and improves final accuracy by up to 11.5 points. AntiSD opens a path to scalable self-improvement, where a language model bootstraps its own reasoning through its training signal.
Abstract: The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling their use across diverse hardware and software platforms. They also allow for more open research and testing, to the extent that users can use them as checkpoints, fine-tune them according to their needs, and potentially redistribute them. In some cases, however, concerns on modifying these weights towards unauthorized uses may outweigh the pros of giving users such a freedom. Defending against such adaptation is non-trivial: since an adaptive attacker can observe all weights and architectures by definition, they can reverse simple structural defenses, and use optimization to defeat the simplest locking mechanisms. In this work, we exploit the inference–training asymmetry of automatic differentiation as a novel defense axis. We propose DLR-Lock, a method where the purveyor of the model purposely replaces each pretrained MLP in their model with a deep low-rank residual network (DLR-Net) of comparable parameter count, forcing activation memory that grows linearly with depth during backpropagation. DLR-Nets are efficiently trained via module-wise distillation. We show that, beyond this memory overhead, DLR-Lock results in architectural mismatches that complicate the optimization landscape of standard fine-tuning, and a backward pass that incurs disproportionately more overhead than the forward pass. Our defense succeeds in withstanding adaptive attackers with full knowledge of the defense strategy while preserving the original model's capabilities. Experiments on LLM validate these claims.
PaperID: 2758, Poster
Abstract: Softmax attention in Transformers suffers from quadratic complexity in sequence length, making it impractical for long-context applications. Linear attention alleviates this issue by replacing the exponential kernel with alternative functions that enable linear-time computation. Among existing linear attention approaches, recent studies have shown that polynomial kernels are particularly effective, as they exhibit sparse and spiky behavior similar to softmax attention, which emphasizes large dot products while suppressing irrelevant interactions. However, the exact computation of high-degree polynomial kernels is infeasible for high-dimensional representations. As a result, prior work relies on approximate polynomial kernels. This introduces a non-negligible approximation error. In this paper, we show that, from the perspective of polynomial kernel approximation, existing linear attention methods are still suboptimal. We propose Learnable Low-Rank Polynomial Sketch (LLoPS), a principled and flexible framework for approximating polynomial kernels with linear attention. Our method learns a low-rank polynomial sketch that provably achieves a strictly smaller approximation error than existing approaches. Experiment results show that LLoPS achieves the highest performance across extensive benchmarks, comparing with various linear attention baselines.
PaperID: 2759, Poster
Abstract: City-scale 3D grounding is commonly formulated as candidate ranking, which requires precomputed instance proposals and becomes inefficient in large outdoor scenes. We propose Urbex, an agentic vision-language framework that instead casts grounding as active spatial search. Urbex represents a city-scale 3D scene as an interactive multi-view environment, allowing a VLM agent to query landmarks, zoom into local regions, render oblique views, and commit a grounded bounding box. We optimize this tool-use policy with reinforcement learning on CityRefer, using smooth localization rewards and lightweight shaping rewards for efficient exploration. Experiments show that Urbex achieves the best Hit@1 and IoU on CityRefer, supports candidate-free grounding without ground-truth boxes, and is substantially faster than the strongest city-scale ranking baseline. Zero-shot results on STPLS3D-Refer further indicate improved cross-scene robustness. The code and dataset will be publicly released.
Abstract: Visual planning asks a model to generate the remaining steps of a procedure in natural language given a partial video context and a goal. Progress on this task is bottlenecked by annotation: clean labeled datasets are small, domain-narrow, and encode a single execution trajectory per example, even though many valid orderings often exist. Large-scale instructional video corpora such as HowTo100M offer orders of magnitude more procedural content, but supervised fine-tuning on pseudo-labels extracted from their noisy ASR narrations fails: segmentation and alignment errors propagate into training, and the resulting supervision is still single-trajectory. We identify a key asymmetry. Extracting clean step labels from noisy video is hard, but verifying whether a generated step sequence is temporally grounded in ASR transcripts is comparatively cheap and scales to millions of videos via precomputed text embeddings. We exploit this asymmetry in RECIPE, which uses grounding quality as a reward signal for Group Relative Policy Optimization (GRPO), turning the noisy corpus into a verification signal rather than a labeling source. The framework applies uniformly to two input configurations of the planner: a Socratic pipeline in which a frozen vision-language model rewrites the video into a textual history fed to the planner, and a Video configuration in which the planner consumes video tokens directly. It applies equally to annotated and weakly supervised training regimes. We evaluate on seven procedural benchmarks using a reference-based LLM-as-judge protocol that scores generated plans across six procedural-quality criteria. RECIPE-RL improves over the base checkpoint at every scale we test (0.5B, 3B, 7B) and on every benchmark, with macro-accuracy gains of +7 to +8 points in-domain at every scale and up to +16 points zero-shot. It substantially outperforms supervised fine-tuning on both annotated and pseudo-labeled continuations (the latter actually the base checkpoint), and is robust to fully replacing human annotations with VLM-derived pseudo-traces. Plugging RECIPE-RL into the proposal stage of VidAssist [14] improves over the strongest zero-shot baseline in our comparison at every horizon on the Visual Planning for Assistance benchmark, and a diversity analysis on COIN shows that RECIPE-RL preserves the generation variety that supervised fine-tuning collapses.
PaperID: 2761, Poster
Abstract: Markovian stochasticity naturally arises in many modern learning problems, including optimization with dependent data streams and reinforcement learning. Despite extensive research on convergence of stochastic gradient methods under such stochasticity, high-probability convergence remains largely unexplored. To close this gap, we provide the first high-probability convergence guarantees for gradient-based methods under Markovian stochasticity. We study SGD with a specialized mini-batch-based gradient estimator and establish high-probability convergence results in the non-convex setting under the classical bounded stochasticity assumption, and further extend the analysis to a weaker bounded-variance assumption by incorporating gradient clipping, yielding convergence guarantees for accelerated stochastic gradient descent in the convex setting. For both methods, we characterize the oracle complexity via a concentration inequality that explicitly captures the effect of Markovian dependence on the resulting convergence rates.
PaperID: 2762, Poster
Abstract: Group-based Reinforcement Learning (RL) has significantly enhanced Large Language Models (LLMs) in agentic scenarios. To achieve finer-grained policy updates, recent agentic RL frameworks have shifted from trajectory-level to step-level training. However, long-horizon agentic RL suffers from severe reward sparsity and delay, as feedback is often deferred for dozens of interaction steps. While existing step-level frameworks refine training granularity, their credit assignment remains coarse-grained and still treats agent exploration as isolated, linear trajectories. This oversimplified perspective ignores the inherent graph structure of state transitions, leading to high-variance state-value estimation and myopic, localized credit assignment. To overcome these critical bottlenecks, we propose Group-Graph Policy Optimization (G2PO), a novel group-based RL algorithm tailored for multi-turn agentic tasks. G2PO explicitly transforms linear interaction trajectories into a global state-transition graph. By aggregating identical observations across different trajectories, we introduce group-aggregation state-value estimation that reduces sampling variance and trajectory-dependent bias. Furthermore, we redefine agent actions as transitions between state nodes and propose an edge-centric advantage estimation strategy. By globally standardizing Temporal Difference (TD) errors across the entire graph, G2PO explicitly identifies and prioritizes critical transitions that drive absolute task progress. Extensive experiments on representative long-horizon benchmarks—WebShop, ALFWorld, and AppWorld—demonstrate that G2PO substantially outperforms state-of-the-art prompt-based and RL baselines, achieving remarkable success rate improvements of up to 22.2% over GRPO.
PaperID: 2763, Poster
Authors: Nhan T Luu, Trung D Luu, Pham N Nam, Truong C Thang
Abstract: Variational quantum algorithms offer a promising route to practical quantum advantage on near-term hardware, yet training quantum neural networks (QNNs) remains hampered by two compounding difficulties: the exponential vanishing of gradient variance known as the barren plateau, and the \mathcalO(L \cdot 2^n) time–memory cost of differentiating through an n-qubit, L-layer circuit. We propose RLQ-Grad, a reinforcement-learning-based optimizer in which a classical policy \pi_\phi (a spectrally-normalized PPO agent) learns to propose parameter updates directly, conditioned on the QNN's current parameters, loss, accuracy, and previous update. Because the surrogate gradient is emitted by a classical network rather than obtained by differentiating through the unitary U(\theta), its variance is not constrained by the barren plateau concentration bound, and its per-step cost scales independent of the Hilbert-space dimension. We prove these two properties formally and verify them empirically on a hardware-efficient ansatz across four supervised benchmarks. RLQ-Grad preserves a near-flat gradient-variance curve where backpropagation, parameter-shift, and adjoint differentiation decay by 1–2 orders of magnitude, attains 128×–1841× wall-clock speedups and constant ~0.1 MB memory at n=12 qubits, and improves top-1 validation accuracy by up to +10% over the strongest gradient-based baseline at every circuit scale tested.
Abstract: Recent work has pointed out current models' limitations in sampling according to prespecified target distributions. In this work, we show that current models nevertheless possess a coarse form of control over their output distributions: by instructing a range of open-source models to maximize or minimize output certainty in a two-alternative forced-choice task with no correct answer, we find that several models can shift entropy in the instructed direction. Through contrastive activation analysis, we further identify a linear direction in activation space that models recruit under uncertainty-modulation instructions and that can be used to causally modulate entropy via activation steering in the absence of any such instructions, suggesting that deliberate entropy control has an identifiable internal representation.
Abstract: Open-weight models are often released, fine-tuned, aligned, merged, and re-released, making provenance audits ask not only whether checkpoints are related, but also which checkpoint came first. Many existing model-provenance methods are designed for a base-known audit setting: given a victim or source model, they test whether a suspect model is related to it. Although these audits are framed as source-to-suspect tests, their underlying evidence is often symmetric, relying on representation similarity, weight similarity, behavioral fingerprints, or correlation statistics. Symmetric pairwise comparisons can detect relatedness, but they cannot by themselves orient A versus B. We therefore introduce a local geometric comparison: instead of comparing two checkpoints directly, we add a third same-family checkpoint as a witness and compare the geometry around A with the geometry around B. Direction is inferred by asking which candidate behaves more like a branching parent. Motivated by this idea, and by the empirically observed asymmetry between parent-anchored and child-anchored witness-overlap distributions, we propose Witness Overlap, a prompt-free, training-free white-box test for directional provenance. On 176 LLM checkpoints from 16 families, our one-witness test orients 95.3% of parent–child decisions using Frobenius cosine. We further evaluate root identification, sibling discrimination, generalizations to VLM and diffusion families, and chain-structured ordering. The signal is robust to weight noise and sparse pruning, with a proposed SVD weight reduction variant showing greater robustness than Frobenius cosine.
PaperID: 2766, Poster
Authors:
Sangyeon Cho, Minyoung Cho, Jungsoo Kim, Sujeong Oh, Sanghack LeeAbstract: In high-stakes settings, decision-makers using conformal prediction (CP) read the predictive interval's width as their operational uncertainty signal. Local feature attribution methods, applied to that width, silently conflate two semantically distinct contributions: a calibration-driven baseline that depends on the held-out calibration set, and an instance-driven term that depends on the test point. We propose a two-game Shapley construction---one game over feature coalitions at the test point, one over feature coalitions across the calibration set---that decomposes the total attribution into calibration-driven and instance-driven components. The construction is method-agnostic on two axes: any local attribution method that can be applied to a scalar uncertainty target can be used, and any split CP family whose uncertainty size admits a known separation into a calibration-only and an instance-only summary satisfies the decomposition. Empirically, the calibration-driven and instance-driven attributions frequently carry opposite signs on synthetic and real data, with representative cases where a feature's sign reverses between the two views; the total attribution then becomes a misleading sum. On a clinical-decision task with semi-synthetic data, this conflation degrades decision quality, whereas jointly considering the calibration-driven and instance-driven attributions recovers high-uncertainty cases missed by total attribution. We further empirically verify that the decomposition holds across multiple CP families that satisfy the separability conditions.
PaperID: 2767, Poster
Abstract: We investigate a general-purpose layer for deep learning that, instead of compressing arbitrary-size training data into fixed-size weight matrices, stores a new pair of key-value representations for every data point during training, and retrieves and recombines these representations through an attention mechanism at inference time---resulting in a growing neural net (NN). We derive such a "retrieval-centric deep learning" (RCDL) paradigm from the classic duality of a linear layer trained by gradient descent as a type of linear attention (LA) over the training data points---which motivates us to replace the corresponding LA by a more powerful form of attention from recent work on sequence models, namely, kernelized attention using radial basis function (RBF) and softmax-like kernels, as well as more advanced linear attention variants such as MesaNet and DeltaNet. We theoretically derive learning algorithms for the growing layers with kernelized attention, and empirically demonstrate their promising performance and sample-efficiency on classic image classification tasks and synthetic teacher-student learning datasets. In the MesaNet/DeltaNet-inspired extensions, we show a formal connection to a recently proposed optimizer and derive highly sample-efficient optimizers for conventional fixed-size NNs. Despite open challenges, RCDL represents a promising paradigm for nonparametric machine learning.
PaperID: 2768, Poster
Abstract: This paper develops a dual-space preconditioning framework for variational inequalities and root-finding problems beyond classical cocoercivity and Lipschitz continuity assumptions. To this end, we introduce both a novel cocoercivity-type condition and a relaxed Lipschitz continuity condition, and derive convergence of deterministic and stochastic methods under these generalized conditions. For isotropic preconditioners, the resulting algorithms can be interpreted as variable-stepsize variants of classical methods, recovering and extending recent analyses based on (L_0, L_1)-type growth conditions while allowing simpler proofs and larger stepsize ranges. We also show that the framework naturally handles constrained variational inequalities through weak Minty-type arguments. Finally, we study the associated continuous-time dynamics and derive a variational and optimal-control interpretation of the preconditioned flow. Overall, the results show that dual-space preconditioning provides a flexible mechanism for designing and analyzing operator methods beyond the standard Euclidean Lipschitz setting.
PaperID: 2769, Poster
Authors:
Erdemt Bao, JunChen, Weijun Qin, Shaopeng Li, Ming Li, Mengchen Zhao, Ziqian Zeng, Cen Chen, HUIPING ZHUANGAbstract: Language-conditioned manipulation requires skill representations that are both multi-modal and predictable: the same instruction may admit diverse executions, while the high-level planner must select a stable skill for low-level control. However, vector-quantized skill codes compress continuous variations into isolated symbols, introduce biased straight-through gradients, and lack geometry for relating semantically similar skills. We propose HyperSkill, a hierarchical skill-learning framework that models skills as continuous embeddings on the unit hypersphere. We use a learnable von Mises--Fisher mixture prior to capture skill-level multi-modality, with geodesic repulsion encouraging diverse mode coverage on the hyperspherical manifold. To connect training-time skill inference with inference-time planning, we introduce Geodesic Skill Alignment (GSA), which aligns posterior skills inferred from demonstration segments with prior skills predicted by a high-level planner using squared geodesic distance. Together, these components provide a geometrically structured skill interface that preserves execution diversity while enabling causal skill prediction at inference time. For low-level control, a Conditional Flow Matching generator produces smooth, multi-modal action trajectories conditioned on skill embeddings. Experiments on LOReL and Kitchen demonstrate that HyperSkill achieves state-of-the-art success rates on the evaluated benchmarks. Ablation studies further validate the importance of each proposed component.
Authors:
Xianwei Chen, Shimin Zhang, Jibin WuAbstract: Scaling on-policy distillation (OPD) for large language models (LLMs) confronts a fundamental tension: asynchronous execution is necessary for system efficiency but structurally deviates from the ideal on-policy objective. To address this challenge, we theoretically decompose the objective discrepancy into rollout drift and supervision drift, capturing staleness in student occupancy and teacher context, respectively. Building on this, we introduce a sample-level freshness score that quantifies the reliability of buffered sample with respect to the on-policy objective. Guided by this signal, we further propose \boldsymbolf-OPD, a novel framework that adaptively regulates stale-sample influence and constrains policy drift accumulated under asynchronous optimization. Across reasoning, tool-use, and coding-agent tasks of increasing interaction horizon, f-OPD consistently achieves task performance comparable to synchronous training while retaining the throughput advantages of asynchronous execution. Our results establish the first recipe that seeks to achieve a performance–efficiency trade-off in OPD, paving the way for long-horizon agentic post-training at scale.
Abstract: Retrieval-augmented generation (RAG) improves the factuality of large language models (LLMs) and vision-language models (VLMs) by grounding generation in external knowledge. However, most existing RAG frameworks assume a centralized retrieval document pool, which is often impractical in sensitive domains such as healthcare, where data are inherently distributed and subject to strict privacy constraints. Recent efforts on decentralized RAG primarily follow prompt-based paradigms that exchange raw, human-readable retrieved content, leading to substantial communication overhead and potential privacy risks. To address these limitations, we propose Representation-based Federated RAG (FedRepRAG), a decentralized RAG framework that avoids raw-data exchange during retrieval. FedRepRAG retrieves information from private local document stores and exchanges only compressed latent representations. To integrate the retrieved knowledge, we introduce a lightweight collaboratively trained projector that maps retrieval embeddings into the feature space of a frozen LLM/VLM backbone. This design makes the exchanged information interpretable only within the federated system, thereby enhancing privacy while substantially reducing communication cost. Experiments on decentralized visual question answering (VQA) and question answering (QA) benchmarks show that FedRepRAG consistently outperforms direct inference and local retrieval baselines, while substantially reducing retrieval tokens, computational cost, and inference-time communication overhead compared with raw-context transfer. Overall, FedRepRAG establishes an effective, efficient, and privacy-preserving framework for federated RAG.
Abstract: Decision Transformers (DTs) have emerged as a powerful framework for sequential decision making by formulating offline reinforcement learning (RL) as a sequence modeling problem. However, extending DTs to online settings with pure RL gradients remains largely unexplored, as existing approaches continue to rely heavily on supervised sequence-modeling objectives during online finetuning. We identify hindsight return relabeling---a standard component in online DTs---as a critical obstacle to RL-based finetuning: while beneficial for supervised learning, it is fundamentally incompatible with importance sampling-based RL algorithms such as GRPO, leading to unstable training. Building on this insight, we propose new algorithms that enable online finetuning of Decision Transformers using pure reinforcement learning gradients. We adapt GRPO to DTs and introduce several key modifications, including sub-trajectory optimization for improved credit assignment, sequence-level likelihood objectives for enhanced stability and efficiency, and active sampling to encourage exploration in uncertain regions. Through extensive experiments, we demonstrate that our methods outperform existing online DT baselines and achieve new state-of-the-art performance across multiple benchmarks, highlighting the effectiveness of pure-RL-based online finetuning for Decision Transformers.
Abstract: We study the computational relationship between replicability (Impagliazzo et al. [STOC `22], Ghazi et al. [NeurIPS `21]) and other stability notions. Specifically, we focus on replicable PAC learning and its connections to differential privacy (Dwork et al. [TCC 2006]) and to the statistical query (SQ) model (Kearns [JACM `98]). Statistically, it was known that differentially private learning and replicable learning are equivalent and strictly more powerful than SQ-learning. Yet, computationally, all previously known efficient (i.e., polynomial-time) replicable learning algorithms were confined to SQ-learnable tasks or restricted distributions, in contrast to differentially private learning. Our main contribution is the first computationally efficient replicable algorithm for realizable learning of parities over arbitrary distributions, a task that is known to be hard in the SQ-model, but possible under differential privacy. This result provides the first evidence that efficient replicable learning over general distributions strictly extends efficient SQ-learning, and is closer in power to efficient differentially private learning, despite computational separations between replicability and privacy. Additionally, we leverage our parity learner to prove that, assuming RP \neq NP, converting replicability to pure differential privacy requires a strict loss in sample complexity. Our main building block is a new, efficient and replicable algorithm that, given a set of vectors, outputs a subspace of their linear span that covers most of them.
Authors: João N. Cardoso, Arlindo L Oliveira, Bruno Martins
Abstract: Understanding what features are encoded by learned directions in LLM activation space requires identifying inputs that strongly activate them. Feature visualization, which optimizes inputs to maximally activate a target direction, offers an alternative to costly dataset search approaches, but remains underexplored for LLMs due to the discrete nature of text. Furthermore, existing prompt optimization techniques are poorly suited to this domain, which is highly prone to local minima. To overcome these limitations, we introduce ADAPT, a hybrid method combining beam search initialization with adaptive gradient-guided mutation, designed around these failure modes. We evaluate on Sparse Autoencoder latents from Gemma 2 2B, proposing metrics grounded in dataset activation statistics to enable rigorous comparison, and show that ADAPT consistently outperforms prior methods across layers and latent types. ADAPT also produces significantly more interpretable prompts in human evaluation. Our results establish that feature visualization for LLMs is tractable, but requires design assumptions tailored to the domain.
Authors:
Weijie Shi, Qiang Xu, fan deng, Yaguang Wu, Jiarun Liu, Yehong Xu, Hao Chen, Jia Zhu, Jiajie Xu, Xiangjun Huang, jian yang, Xiaofang ZhouAbstract: Speculative decoding accelerates LLM inference by drafting a tree of candidate continuations and verifying it in one target forward. Existing drafters fall into two camps with opposite weaknesses. Autoregressive drafters such as EAGLE-3 preserve dependence along each draft path but call the drafter once per tree depth, making drafting a non-trivial share of per-iteration latency. Parallel drafters cut drafter calls by predicting multiple future positions in one forward, but each position is predicted without seeing the others, producing paths the verifier rejects. In this paper, we propose SpecBlock, a block-iterative drafter that combines path dependence with cheap drafting. Each drafter forward produces K dependent positions and we call this a block. The draft tree grows through repeated block expansions. Two mechanisms explicitly carry path dependence to keep later draft positions accurate. Within each block, a layer-wise shift carries the previous position's hidden state into every decoder layer. Across blocks, each new block can start from any position of the previous block, inheriting its hidden state to extend the path. To spend verifier budget where acceptance is likely, a co-trained rank head replaces the fixed top-k tree by allocating per-position branching during drafting. To avoid training the drafter on prefixes it never produces at inference, a valid-prefix mask drops the loss at later positions once an earlier one is wrong. Beyond static drafting, a cost-aware bandit at deployment uses free verifier feedback to update the drafter selectively, only when the expected throughput gain exceeds the update cost. Experiments show that SpecBlock improves mean speedup by 8--13% over EAGLE-3 at 44--52% of its drafting cost, and cost-aware adaptation extends this lead to 11--19%.
PaperID: 2776, Poster
Authors: Yiming Xu, Hao Cheng, Monika Sester
Abstract: Autoregressive generation is natural for language, where predicted tokens can be directly reused as the next prediction state, but trajectory forecasting lacks such a clean token: motion is continuous, multimodal, and expressed in local coordinate frames that evolve with the predicted trajectory. We present MoCAR (Motion-code Coordinate-aware AutoRegression), a decoder-only framework that casts trajectory forecasting as next-code prediction in a coordinate-aware continuous latent space. MoCAR learns a continuous motion-code space for endpoint-normalized trajectory segments, where each code jointly captures local trajectory geometry and the reference-frame transition induced by that segment. Historical motion codes are used as a teacher-forced prefix, future codes are generated autoregressively under temporal, map, agent, and mode interactions, and predicted codes persist in latent memory while decoded endpoints update the local scene context. This enables rollout without trajectory-space re-tokenization, trajectory queries, goal candidates, or proposal-and-refinement pipelines. On Argoverse (AV) benchmarks, MoCAR achieves top-tier performance with a simple single-stage architecture, transfers strongly from AV2 to AV1 in zero-shot evaluation, and improves on turn-heavy scenarios. Ablations confirm that the learned continuous motion-code space, latent alignment, weak KL regularization, and joint tokenizer-predictor optimization are essential for stable latent autoregression.
Abstract: Approximate Nearest Neighbor Search (ANN) is an important algorithmic primitive that has found a plethora of applications in machine learning and information retrieval. The classic approach to this problem due to Indyk and Motwani (1998) leverages Locality Sensitive Hashing (LSH) by sampling multiple hash functions that are likely to identify similar data points. This approach has, however, been demonstrated to be vulnerable to adaptive queries and updates which may not be independent of its internal randomness (Kapralov et al. 2024). While differential-privacy-based techniques have yielded robust versions of randomized algorithms and data structures for estimation problems (Hassidim et al., 2022), ANN is a search problem: the algorithm must return an actual dataset point, without obfuscating the output by adding noise or rounding it. Feng et al. (2025) circumvent this limitation by relying on a density assumption that bounds the number of points near each query point. We study a stronger worst-case model in which the adversary chooses both the initial dataset and an adaptive query sequence. For LSH-equipped metric spaces, we give adversarially robust ANN algorithms with sublinear query time whose guarantees do not depend on any structural assumption on the dataset. The main technical challenge is to preserve sublinear search time when the adversary selects a dataset with arbitrarily many points near a query. Rather than applying differential-privacy robustification as a black box, we develop search-specific mechanisms: fairness as a route to adaptive security, a bucketing reduction from search to robust decision, and a concentric annuli construction that improves the query-time exponent. Moreover, for low-dimensional spaces, we give algorithms with a strong ``for-all'' guarantee, which are correct for every possible query.
PaperID: 2778, Poster
Abstract: Joint Embedding Predictive Architectures (JEPAs) offer a compelling framework for learning world models in compact latent spaces, yet existing methods remain fragile, relying on complex multi-term losses, exponential moving averages, pre-trained encoders, or auxiliary supervision to avoid representation collapse. In this work, we introduce LeWorldModel (LeWM), the first JEPA that trains stably end-to-end from raw pixels using only two loss terms: a next-embedding prediction loss and a regularizer enforcing Gaussian-distributed latent embeddings. This reduces tunable loss hyperparameters from six to one compared to the only existing end-to-end alternative. With ~15M parameters trainable on a single GPU in a few hours, LeWM plans up to 48x faster than foundation-model-based world models while remaining competitive across diverse 2D and 3D control tasks. Beyond control, we show that LeWM's latent space encodes meaningful physical structure through probing of physical quantities. Surprise evaluation confirms that the model reliably detects physically implausible events.
PaperID: 2779, Poster
Abstract: Real-world television and streaming pipelines often deliver compressed HD video to increasingly capable displays, making video super-resolution (VSR) both practically important and fundamentally ambiguous. Existing generative VSR methods are largely receiver-only, requiring the model to infer missing high-frequency structures from degraded pixels alone. We propose MetaVSR, a sender--receiver collaborative framework that transmits compact structural edge metadata as bitrate-constrained side information for generative VSR. MetaVSR builds on a one-step video diffusion transformer and encodes both the low-quality video and transmitted metadata through the native 3D VAE, enabling token-level fusion within the original DiT attention blocks without additional control branches. We further introduce a rate-aware Canny edge metadata selection strategy that allocates metadata bits to structurally challenging regions, and evaluate reconstruction quality under matched total transmission budgets that include both compressed video and metadata. Across five public VSR benchmarks, MetaVSR shows that budgeted sender-side structural metadata can improve reconstruction over the receiver-only DOVE-5B baseline, increasing PSNR by 1.03 dB and SSIM by 0.055 on average while using a smaller 2B backbone. On MetaSPS, our rate--distortion evaluation shows approximately 20% rate saving under light-noise degradation and more than 50% under heavy-noise degradation compared with a no-metadata generative baseline. These results demonstrate that compact sender-side edge metadata can reduce reconstruction ambiguity and improve bitrate-efficient VSR when its transmission cost is explicitly accounted for.
Abstract: Exhaustively evaluating many large language models (LLMs) on a large suite of benchmarks is expensive. We cast benchmarking as finite-population inference and, under a fixed query budget, seek tight confidence intervals (CIs) for model accuracy with valid frequentist coverage. We propose Factorized Active Querying (FAQ), which (a) leverages historical information through a Bayesian factor model; (b) adaptively selects questions using a hybrid variance-reduction/active-learning sampling policy; and (c) maintains validity through Pro-Active Inference---a finite-population extension of active inference (Zrnic & Candès, 2024) that enables direct question selection while preserving coverage. With negligible overhead cost, FAQ delivers up to 5× effective sample size gains over strong baselines on two benchmark suites, across varying historical-data missingness levels: this means that it matches the CI width of uniform sampling while using up to 5× fewer queries. We release our source code and our curated datasets to support reproducible evaluation and future research.
PaperID: 2781, Poster
Abstract: 3D Gaussian Splatting (3DGS) has emerged as a powerful paradigm for real-time photorealistic novel view synthesis. However, its training-time densification remains largely heuristic, often resulting in redundant Gaussian proliferation and inaccurate geometric reconstruction. Existing methods predominantly rely on first-order screen-space gradients to trigger cloning and splitting, which fail to distinguish true geometric discontinuities from high-frequency textures or noisy residuals. To address this limitation, we propose ACD-GS, an Asymmetric Curvature-aware Densification framework for compact and geometry-aware 3D Gaussian optimization. Our key insight is that second-order depth responses provide a more reliable structural prior for density control, enabling asymmetric regulation of cloning and splitting to better capture geometric discontinuities. Building on this, we introduce a multi-view photometric consistency constraint to suppress transient single-view errors, together with a capacity-aware adaptive regulation strategy that balances densification and pruning during training. Unlike post-hoc compression or quantization-based approaches, ACD-GS directly prevents redundant Gaussian generation during training, yielding inherently compact representations. Extensive experiments demonstrate that our method reduces nearly 60% of Gaussian count and storage while maintaining competitive rendering fidelity, consistently outperforming existing methods across multiple datasets. Our code is available in https://anonymous.4open.science/r/ACD_GS-0E2C
Authors: Guoqing Wang, Pin Tang, Xiangxuan Ren, Chao Ma
Abstract: Recent progress in integrating vision-language-action (VLA) models with world modeling has significantly advanced end-to-end autonomous driving by allowing systems to reason about imagined futures instead of merely reacting to current observations. However, most world VLA models generate future frames at the pixel level with plenty of dense visual tokens. Although such generated futures can be visually plausible, they preserve many driving-irrelevant appearance details, i.e., only a small portion of the tokens directly encode planning-relevant factors, such as agent geometry, motion dynamics, and interactions. This introduces a significant performance bottleneck. An important question thus arises: what world state should a world VLA model reason over to support planning beyond pixel-level prediction? We argue that planning-oriented world modeling should focus on an agentic latent state that captures the interactions between objects in the driving scene over time. As such, we introduce AgentWorld, an agentic world VLA model that uses discrete agent tokens jointly for reasoning and trajectory planning. Each token is dynamically grounded to a surrounding agent associated with its object's state, such as semantic identity, 3D geometry, motion, and interaction context. AgentWorld learns these tokens through three key processes: explicit geometric grounding, reasoning-oriented supervised fine-tuning, and agent-oriented reinforcement learning. Extensive open-loop and closed-loop experiments under diverse driving scenarios show that AgentWorld enhances planning capabilities, fully exploits visual evidence and agent dynamics, and yields interpretable agentic world states. The code will be made publicly available.
Abstract: Text-guided image editing has advanced rapidly with diffusion models and unified multimodal foundation models. However, most existing methods remain confined to single-turn settings, overlooking the more realistic scenario of multi-turn in-context editing, where users iteratively refine an image through a sequence of instructions. In this setting, a model must follow each new instruction while preserving accumulated session-level constraints, challenged by two coupled failure modes: long-context dilution, where sparse textual constraints become difficult to recover from growing interleaved image-text histories, and state contamination, where earlier editing mistakes degrade subsequent generations. We introduce Edit-R2, a novel reinforcement learning post-training framework for unified multimodal models. Edit-R2 reconstructs the operative session intent, which effectively consolidates scattered historical constraints into an explicit reasoning trace before each editing turn. It further enables multi-turn RL over both reasoning and generation through a unified objective that jointly optimizes intent reconstruction generation in discrete text space and flow-matching image generation in continuous latent space, while a trajectory filtering mechanism suppresses corrupted rollouts to stabilize training under state contamination. To support systematic evaluation, we introduce MICE-Bench, a large-scale benchmark for multi-turn in-context editing with automated metrics for instruction following (IF), content consistency (CC), and global awareness (GA) over accumulated session constraints. Experiments show that Edit-R2 substantially improves multi-turn in-context editing and achieves competitive performance compared against strong baselines.
PaperID: 2784, Poster
Abstract: Introducing new capabilities to frontier models has long been the goal for posttraining, which relies predominantly on supervised finetuning (SFT) and reinforcement learning (RL) to achieve this. Prevailing wisdom dictates that on-policy RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and forgetting. At the same time, SFT enables learning from inherently off-policy expert data, whereas RL must rely on a model's ability to generate positive signal on a new task. In our work, we seek to bridge the strength of on-policy learning with the information signal of off-policy expert traces. To this end, we introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that transforms off-policy traces into more on-policy ones given a reference model for finetuning. Across tasks like scientific skill acquisition, mathematical reasoning, and fact learning, we demonstrate that simple SFT on these boosted off-policy traces can generalize better and forget less than on-policy counterparts. More generally, our approach outlines a principled methodology for model-specific data refinement, suggesting broader utility as a plug-and-play component throughout the post-training pipeline.
PaperID: 2785, Poster
Abstract: Time-series foundation models (TSFMs) are increasingly deployed as zero-shot forecasters in observability platforms, yet the operational concepts encoded within their internal representations remain largely un-audited. To address this gap, we introduce \texttttoto-interp, a control-first interpretability protocol that pairs linear probes with four rigorous baselines: raw-feature, shuffled-label, randomized-backbone, and supervised raw-window Fourier Neural Operator (FNO) controls. Applying this harness to TOTO on BOOM, we demonstrate that three structural telemetry concepts, such as cadence bucket ("frequency"), metric type, and domain, are linearly decodable, layer-localized, and specifically tied to pretraining. A budget-matched replication on MOMENT-base successfully recovers these same axes, confirming these representations are not isolated artifacts of the TOTO architecture. Furthermore, our matched-null and on-manifold interventions reveal a critical two-way dissociation between linear decodability and forecast-sensitivity. For instance, concepts like future burstiness are highly decodable yet completely inert under intervention, whereas other weakly decodable directions can severely disrupt forecasts off-manifold. By releasing the audit harness and detailing negative control cases (such as cardinality and shift risk), we demonstrate that linear decodability alone is insufficient for reliable post-training steering or deployment auditing in TSFMs.
Abstract: Vector quantization (VQ) with autoregressive (AR) token modeling is a widely adopted and highly competitive paradigm for time-series generation. However, such models are fundamentally limited by exposure bias: during inference, errors can accumulate across sequential predictions, leading to pronounced quality degradation in long-horizon generation. To address this, we propose Matching), a non-autoregressive framework that operates entirely in the frozen VQ latent space and enables parallel sequence generation via flow matching. We tackle three key challenges in making this transition: (1) eliminating exposure bias by replacing step-wise token prediction with a global transport map; (2) mitigating the high-dimensionality of VQ token spaces via a low-rank manifold decomposition with a learned anchor prior over the latent manifold; and (3) incorporating discrete supervision into continuous transport dynamics by introducing a categorical posterior over codebook indices within a variational flow-matching formulation. Extensive experiments show that SDFlow achieves state-of-the-art performance, improving Discriminative Score and substantially reducing Context-FID, particularly for challenging long-sequence generation. Moreover, SDFlow provides significant inference speedups over autoregressive baselines, offering both high fidelity and computational efficiency. Code is available at
PaperID: 2787, Poster
Abstract: Despite the success of Neural Operators (NOs), architectural hyperparameters such as depth are still typically selected by empirical grid search, with limited connection to the governing PDE. We introduce the alternation depth, \alpha_F, a structural quantity of the PDE right-hand side that counts the minimum number of non-pointwise/pointwise coupling stages needed to build its leading nonlinear spatial structure. We use L_\min = \alpha_F + 1 as a principled \emphdepth-efficiency lower-bound guideline for IVP solution maps: a starting depth, not a predictor of the empirical optimum L^. We provide partial formal support: (sufficiency) an L-block architecture can approximate any operator in a formally defined L-alternation class under compactness and regularity assumptions; (separation) for a constructed hard family at finite resolution and polynomial activations, using fewer blocks forces width to grow as a power of the discretization size. Empirically, a 736-run study over nine in-scope IVP datasets plus one Darcy stress test, using FNO-family variants and one kernel-integral baseline (KIN), is consistent with depth improvements appearing at or above L_\min in the tested settings. The observed optimum can be larger because of spectral approximation, pointwise approximation difficulty, data regime, and optimization. We summarize the resulting practitioner guidance in four capacity-allocation rules.
PaperID: 2788, Poster
Abstract: Link prediction remains difficult in sparse graphs, where diffusion heuristics lack connectivity and geometric methods suffer spectral instability when coupling learned representations with diffusion operators. We introduce Concorde, an energy-based framework that resolves this by decomposing link likelihood into three structurally decoupled branches: (i) a parameter-free spectral anchor that establishes permutation-invariant topological stability; (ii) a mixed-curvature residual on a learnable \mathbbH\_c\_h^d\_h × \mathbbS\_c\_s^d\_s product manifold, regularized by a feature-aware Ollivier--Ricci curvature proxy; and (iii) a continuous vector field on \mathcalT\mathbbH\_c^d for non-local structural alignment across topological gaps. The branches are fused via learned gating and trained with an N-pair ranking loss aligned with retrieval-style evaluation. Across five benchmarks spanning sparse, dense, featureless, and attributed regimes, Concorde establishes state-of-the-art early retrieval precision, nearly doubling P@10% over leading baselines in sparse settings.
Abstract: Despite recent advancements in Text-to-Video (T2V) synthesis, generating high-fidelity and dynamic motion remains a significant challenge. Existing methods primarily rely on Classifier-Free Guidance (CFG), but standard CFG treats all semantic dimensions uniformly and provides no mechanism to selectively resolve ambiguity along the motion axis. To address this, we propose MotionCFG, a training-free approach that introduces selective guidance along the motion subspace of the condition embedding. Specifically, we construct a motion-perturbed negative condition by injecting Gaussian noise into motion-related embeddings, and steer generation away from it to selectively amplify motion intent. We show theoretically that this procedure implicitly approximates , a second-order correction that suppresses motion-ambiguous regions of the score landscape and amplifies well-resolved dynamic peaks. Combined with a piecewise guidance schedule that confines intervention to the early denoising steps, MotionCFG consistently improves motion dynamics across state-of-the-art T2V frameworks with negligible overhead. We further demonstrate that this Laplacian sharpening principle generalizes beyond motion, effectively steering complex, non-linear concepts such as precise object numerosity that are typically difficult to modulate via standard text-based guidance.
Authors:
Yuning Qiu, Lin-Feng Zou, Jiong-Da Wang, Xue-Rong Yuan, Wang-Zhou DaiAbstract: In high-complexity abstract reasoning, a system must infer a latent rule from a few examples or structured observations and apply it to unseen instances. LLMs can express such rules as programs, but ordinary conversation-based refinement is largely outcome-level: it observes that an answer or output is wrong without formally re-checking which abstraction, relation, or transformation justified that outcome. We propose \emphAbduction-Based Procedural Refinement (ABPR), a neuro-symbolic refinement approach that couples an LLM with a Prolog meta-interpreter. ABPR treats each candidate program as an executable declarative hypothesis of the latent rule and reifies its SLD goal--subgoal resolution into compact proof-tree-style derivations, following Shapiro's algorithmic program debugging (APD). In this view, refinement is not merely code-level debugging, but semantic re-checking of the model's hypothesised rule. We evaluate ABPR primarily on ARC-AGI-2, a challenging few-shot abstract rule induction benchmark over grid transformations. ABPR with Gemini-3-Flash achieves 56.67% Pass@2, while GPT-5.5 xHigh with ABPR reaches 98.33% Pass@2 on the public evaluation set. Supplementary experiments on fill-in-the-blank I-RAVEN-X and A-I-RAVEN adaptations provide evidence that the same trace-guided framework extends beyond ARC-specific grid tasks to RAVEN-style relational and analogical abstraction. Repeated-run and sensitivity analyses show that parallel trace-guided search reduces stochastic variance as search breadth and refinement depth increase.
Abstract: Skills provide an effective mechanism for improving LLM agents on complex tasks, yet in existing agent frameworks, their creation, refinement, and selection are typically governed by external teachers, hand-designed rules, or auxiliary modules. As a result, skills remain external resources to be invoked, rather than capabilities that agents can develop, adapt, and internalize through experience. To endow LLM agents with autonomous skill mastery, we propose SkillMaster, a training framework that teaches agents to create new skills, refine existing skills, and select accumulated skills during task solving. This capability is achieved through three key designs. First, we train agents through trajectory-informed skill review, teaching agents to propose, update, or retain skills based on evidence from completed episodes. Second, each candidate skill edit is designed to be evaluated by its counterfactual utility on related probe tasks, providing a direct learning signal for training skill-editing decisions. Third, we introduce DualAdv-GRPO, which separately estimates advantages for task-solving actions and skill-editing decisions, stabilizing joint training across task solving and skill management. Experiments on ALFWorld and WebShop show that SkillMaster improves the overall success rate over state-of-the-art baselines by 8.8% and 9.3%, respectively, achieving the best performance among all compared methods. Further analysis reveals a marked shift in agent capability: agents trained with SkillMaster can identify skill failures, refine procedural knowledge from trajectory evidence, and transfer improvements to future tasks with limited skill-bank edits. Overall, SkillMaster moves LLM agents beyond mere skill use toward self-improving agents capable of developing, adapting, and applying their own skill repertoires. Our code is released at https://anonymous.4open.science/r/skillmaster-F4B0.
Abstract: Building structured 3D scene layouts from a single image requires reconciling visual observations with physical and spatial constraints, a challenge that is difficult to address with direct prediction alone. In this work, we formulate monocular 3D layout estimation as a perceive-then-plan problem with vision-language models, where a Perceiver first grounds the 3D objects and then a Planner iteratively refines the scene hypothesis through actions that improve physical plausibility while preserving consistency with the input image. We propose Layout-as-Policy (LaP), which casts the planning stage as a policy learning problem: 3D layouts are represented as structured states, and refined via discrete actions such as translation, rotation, and rescaling. Starting from an observation-aligned initialization with the geometry-enhanced Perceiver, the LaP Planner is trained to produce action sequences that progressively resolve geometric inconsistencies and enforce realistic spatial relations. To enable effective learning, we combine supervised trajectory initialization with preference-based optimization, allowing the model to learn corrective behaviors without requiring explicit reward engineering. This formulation transforms layout estimation from a one-shot prediction task into an iterative refinement process, enabling better handling of global constraints and complex object interactions. Experiments demonstrate that our approach produces layouts that are more physically coherent and better aligned with visual observations, while naturally supporting downstream tasks such as scene editing and manipulation.
PaperID: 2793, Poster
Abstract: Multi-modal machine learning has become a dominant paradigm for building large-scale models that integrate information across heterogeneous modalities and achieve unprecedented representational capabilities. However, the multi-modal nature of such models introduces unique challenges for explainability. Specifically, the standard feature attribution task considers only individual feature contributions. While solutions of this kind can be adapted to operate on each modality, such first-order explanations are insufficient to characterize cross-modal feature interactions, which are essential for understanding model behavior involving cross-modal alignment. To bridge the gap, this paper studies higher-order feature interactions in multi-modal settings and proposes Gradient-Estimation-Based Feature Interaction (GEFI). GEFI is derived from the proxy gradient estimation framework, and we establish its theoretical equivalence to the Shapley Interaction Index at arbitrary orders. Beyond the general theory, we instantiate GEFI-PM for CLIP-like models by exploiting their dual-encoder architecture. The resulting formulation enables the efficient estimation of interactions between visual and textual features. Compared to existing methods for revealing cross-modal interactions, which generally operate on fixed image patches, GEFI offers fine-grained pixel-level interactions while maintaining computational efficiency, yielding more expressive results.
PaperID: 2794, Poster
Abstract: Post-training quantization degrades multimodal model capabilities unevenly, yet practitioners lack a principled way to predict which capabilities will collapse. We propose the Computational Depth Principle (CDP): end-to-end fidelity under quantization follows \hat p^n, where n counts sequential precision-dependent operations and \hat p is a per-architecture survival rate. Across four VLM families and diverse benchmarks spanning image understanding, image generation, video understanding, and video generation, we show that deeper reasoning tasks consistently degrade more steeply and that a single shallow probe predicts the full degradation ranking on held-out architectures. The same exponential form extends to image generation, where denoising step count plays the role of depth, and to video understanding, where longer temporal reasoning chains amplify losses. For video generation, spatial quality degrades more than temporal coherence, indicating that the dominant compounding axis is denoising depth rather than cross-frame coupling. A complementary distortion-to-information (D/I) flip ratio diagnoses quantization method effects, revealing that certain methods preserve near-baseline quality while others produce catastrophic collapses on specific architecture pairings that aggregate accuracy alone cannot detect. Module ablation localizes the sensitivity to the language backbone rather than the vision encoder. Together, depth and D/I convert full-suite evaluation into a two-probe screening protocol that surfaces the highest-risk capability gaps first.
Abstract: Hallucinations remain a major obstacle to deploying large language models (LLMs) in knowledge-intensive settings, where generated responses must be faithfully grounded in provided evidence. Reinforcement learning (RL) is a promising direction for hallucination mitigation, but response-level faithfulness rewards suffer from a granularity mismatch: localized hallucinations can cause supported content to receive spurious penalties. Although recent work introduces fine-grained feedback such as claim-level verification and token-level rewards, unbalanced credit assignment can still induce length, verbosity, or optimization-noise biases. We propose BALTO, a Balanced Token-level Policy Optimization framework for hallucination mitigation. BALTO extracts checkable factual claims, verifies them against the reference context, and projects claim-level judgments to token-level labels. A balanced token-level credit assignment mechanism is introduced into the framework. This design redistributes probability mass from unsupported content toward faithful content, rather than suppressing the entire response. We systematically analyze the limitations of response-level rewards from a theoretical standpoint, and prove BALTO’s advantages in training stability and optimization efficiency for hallucination mitigation. Experiments on ConFiQA, RAGTruth, and FinLLM-Eval show that BALTO achieves the highest faithfulness across all six model--benchmark settings and consistently outperforms existing post-training baselines in Q-Score, demonstrating a stronger faithfulness--informativeness trade-off.
PaperID: 2796, Poster
Abstract: LLM-based systems are increasingly deployed in domains where incorrect outputs carry real costs, often necessitating expensive human review and verification. We argue that a central issue is not simply error, but : models fail to adjust their behavior according to the cost of being wrong. We study this failure across both explicit utility framings as well as natural-language descriptions of stakes that mirror real-world deployment scenarios. Our evaluation spans agentic coding and mathematical reasoning, covering five frontier model families. Across settings, models systematically under-abstain: they continue to answer or submit patches even when incorrect answers carry significant consequences. The pathology is striking: models continue to submit answers in trivial settings where abstention is strictly dominant, and even when told an incorrect answer will cause nuclear extinction. Further, we find that this behavior is orthogonal to existing benchmarks; increasing model size or capability within a family does not improve sensitivity. Comparing base and instruction-tuned models localizes much of this anti-abstention bias to post-training: base models abstain far more often, though they are not themselves reliably consequence-aware. Finally, we evaluate prior post-hoc interventions alongside in-context learning and fine-tuning, finding that none robustly resolves the failure. Our results suggest that current post-training and evaluation pipelines optimize models to answer, not to act under stakes -- a significant bottleneck for trustworthy autonomy.
PaperID: 2797, Poster
Abstract: The exploration-exploitation tradeoff is a fundamental challenge in reinforcement learning. In Monte Carlo Tree Search (MCTS), this tradeoff is balanced explicitly by the pUCT exploration coefficient and implicitly by parallelism when the search is combined with neural network evaluations. However, the exploration coefficient is typically tuned as an algorithmic hyperparameter, while the degree of parallelism is often treated as a system detail. To this end, we characterize how these two factors shape the search-tree structure and jointly affect performance scaling. We prove that, under regularity conditions, the depth of the optimal path grows as \Theta(\sqrtn/\log n) for n simulations under the MuZero exploration coefficient formula. For tree-parallel MCTS, we prove that the search depth preserves the same asymptotic order when T=o(\sqrtn), while sufficiently large parallelism introduces contention and reduces the depth growth rate to \Theta(n/T) for T threads. Our empirical findings indicate that each doubling of the number of threads decreases the best-performing exploration coefficient by approximately 0.19. Overall, our characterization connects the search-tree structure with the exploration-exploitation behavior in pUCT, providing insight into how exploration coefficients and parallelism can be jointly tuned.
PaperID: 2798, Poster
Abstract: We study stochastic linear bandits under a natural combination of batching and communication constraints: the time horizon is partitioned into batches of equal size B, and during each batch the learner sends B requested arm pulls to an agent, who then observes the corresponding B rewards and responds with a single bit of feedback to the learner. For each batch, the learner specifies the 1-bit quantization rule the agent uses, which may depend on all previously received bits but not on any past rewards directly. This setting addresses a significant yet unexplored ``middle ground'' between previous models having per-round quantization only or total bit budgets only. We establish a minimax lower bound showing that \Omega(B\min \lbrace d,\log\lvert \mathcalA \rvert \rbrace) regret is unavoidable due to the 1‑bit communication bottleneck, even in the absence of noise. Combined with standard statistical limits, this yields a general lower bound of \widetilde\Omega(B\min\lbrace d,\log\lvert \mathcalA \rvert \rbrace + \sqrtdT \min \lbrace d,\log\lvert \mathcalA \rvert\rbrace ). We develop two phased‑elimination algorithms based on G-optimal designs and 1‑bit mean estimation. The first achieves \widetildeO(dB + d\sqrtT) regret, matching the lower bound up to logarithmic factors when \lvert \mathcalA \rvert = \exp(\Omega(d)), and the second incorporates a safe‑arm identification and warm‑start procedure to obtain \widetildeO(B\log\lvert \mathcalA \rvert + d^3/2\sqrtB + \sqrtdT\log\lvert \mathcalA \rvert) regret, which is near‑optimal in broad scaling regimes of (\lvert \mathcalA \rvert, B, d,T). Together, our results demonstrate that a single bit of feedback per batch suffices to preserve nearly order‑optimal regret across a wide range of settings, even in the case of large batch sizes such as \Theta(\sqrtT).
PaperID: 2799, Poster
Abstract: Existing graph neural network-based methods for Boolean satisfiability (SAT) solving are generally constrained by limited local receptive fields, making it difficult to adequately capture the long-range dependencies between variables and clauses in SAT instances. Moreover, these methods overlook the multi-solution nature of SAT instances. Consequently, their expressive power and generalization performance remain limited. To address these issues, we propose DualSAT, a novel neural framework for the SAT problem, designed to enhance the phase selection heuristic of SAT solvers. Specifically, DualSAT adopts a novel dual-branch graph neural network that separately encodes local structural information and global structural information in the graph, and then fuses them through an attention-based cross-branch feature fusion module. An assignment decision head is further introduced to generate the initial assignment prediction of variables. This prediction process requires only a single forward pass, achieving a favorable balance between expressive power and inference efficiency. Furthermore, to account for the multiple solutions nature of SAT instances, we design a two-stage training strategy that combines supervised and unsupervised learning to improve the generalization performance of the model. Experimental results on multiple datasets show that, as an end-to-end assignment prediction model, DualSAT significantly outperforms NeuroSAT and NeuroBack. In addition, we integrate DualSAT into the classical SAT solver CaDiCaL, where it increases the number of solved instances by up to 11.6% and reduces the number of conflicts by up to 9.21% on SAT competition problem sets.
PaperID: 2800, Poster
Abstract: Hybrid modeling, the combination of machine learning models and scientific mathematical models, enables flexible and robust data-driven prediction with partial interpretability. However, the unknown parameters of the scientific model cannot necessarily be estimated properly, since the flexibility of the machine learning model might make the scientific model part effectively ignored in prediction. We may avoid it by applying some regularization, but the formulation of such regularizers typically depends on model architectures and domain knowledge. In this paper, we propose an architecture-agnostic method to learn hybrid models while properly estimating the scientific parameters. The idea is to use the flatness of loss minima to achieve model simplicity, based upon the Occam's razor principle. We employ the idea of sharpness-aware minimization and adapt it to the hybrid modeling setting. Numerical experiments demonstrate the effectiveness of the SAM-based hybrid model learning for scientific parameter estimation.
PaperID: 2801, Poster
Authors: Chunqi Guo
Abstract: Model-based reinforcement learning agents optimize policies by backpropagating through imagined latent trajectories. However, model reliability can vary substantially within a single rollout: some imagined steps produce useful gradients while others introduce harmful bias. Existing adaptive methods address this at either the rollout level (adaptive horizons, priority replay) or the per-step level (gradient truncation); their interaction remains unexplored. In this paper, Freshness-Gated Imagination (FGI), a lightweight trust layer for latent world models that combines rollout-level adaptation with per-step entropy-based soft gating at zero additional forward-pass cost, is proposed. Through systematic factorial ablation on Crafter (10 seeds), a superadditive synergy is uncovered: per-step gating alone harms performance (−6.2%), rollout-level adaptation alone is ineffective (+0.6%), while their combination achieves +14.8% improvement in terms of the Crafter geometric-mean score (Wilcoxon signed-rank p=0.032, Cohen's d=0.74). A bias-variance analysis shows that this synergy arises because rollout-level priority replay improves the calibration of the entropy signal on which the per-step gate depends. On DMC continuous control tasks, FGI preserves baseline performance, confirming no-harm in well-modeled environments. Code is available at https://anonymous.4open.science/r/FGI-89D9.
PaperID: 2802, Poster
Abstract: Large language models often improve reasoning by sampling multiple outputs and aggregating their final answers, but precise and efficient control of error levels remains a challenging task. In particular, deciding when to stop sampling remains difficult when the stopping rule is data-dependent and the set of possible answers is not known in advance. We study anytime-valid certification of a prespecified target answer as the unique mode of the model’s response distribution, a guarantee distinct from answer correctness. We propose the Certification by Intersection-union Testing with E-processes (CITE) algorithm, which provably controls false certification at any prescribed level under arbitrary data-driven stopping, without requiring prior knowledge of the answer category set. We also prove an category-set-size-free stopping-time rate, establish matching minimax lower bounds up to constants in the main regime, and extend the construction to confidence-weighted voting. Simulations and LLM self-consistency experiments show empirical error control and improved certification in diffuse-tail settings.
PaperID: 2803, Poster
Authors: Khoren Petrosyan, Artashes Mkrtchyan, Rafayel Mkrtchyan, Hrant Khachatrian, Theofanis Raptis
Abstract: Large-scale pretraining has driven much of recent progress in deep learning, but many physics-governed prediction problems remain outside this regime: real measurements are scarce, and high-fidelity simulators are slow and often proprietary. Indoor radio-map prediction is a representative example. Existing benchmarks rely on limited high-quality simulated data, while recent methods largely focus on task-specific input features or architectures that encode physics priors. In this work we propose an alternative path. We generate 128M synthetic indoor radio-map samples with a fast simulator that omits known propagation effects, pretrain a ResNet-based encoder-decoder without task-specific modifications, and fine-tune it on high-quality benchmark data. Despite simulator mismatch, our method reduces error by 15% on average relative to the best known methods across five indoor radio-map prediction tasks and improves transfer on a small real-measurement dataset. Downstream performance follows a predictable data-scaling trend: a power law fitted on runs up to 16M synthetic samples predicts the per-task 2M-normalized RMSE at an order of magnitude more data within 5% mean absolute percentage error. Probing analyses further show that the pretrained model captures physically meaningful spatial-field structure, including free-space attenuation, transmitter-centered symmetries, and wall-mediated effects. These results suggest that cheap approximate simulators can serve as scalable pretraining engines for physics-governed spatial prediction, partially compensating for scarce high-fidelity data.
Abstract: Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictions are coupled. Existing training-free samplers such as Top-(k), Fast-dLLM, and EB-Sampler mainly control how many tokens to reveal, while often ranking candidates by token-wise scores that ignore interactions within the selected set. We propose ADAS, a training-free reranking rule for parallel masked diffusion decoding. ADAS leaves the base sampler's stopping rule unchanged and modifies only subset construction: it greedily discounts a candidate when it attends strongly to already selected positions whose predictions remain uncertain. Unlike graph-constrained methods that turn attention into hard compatibility constraints, ADAS keeps attention continuous and uses it as a soft marginal penalty. Across LLaDA-8B-Base and Dream-7B-Base on GSM8K, MATH500, HumanEval, and MBPP, plugging ADAS into Top-(k), Fast-dLLM, and EB-Sampler improves low-NFE performance at matched denoiser evaluations by (9.11) and (10.46) percentage points on average, respectively, with (3.1%) per-forward runtime overhead. These results show that soft attention-discounted reranking is a simple and modular way to improve quality in highly parallel decoding for masked diffusion language models.
PaperID: 2805, Poster
Abstract: Video generation models have achieved remarkable visual fidelity and have strong potential to become general-purpose world simulators. Despite this progress, they still fail to generate videos which adhere to laws of physics. The problem becomes even more apparent in realistic settings where multiple physical principles must work together within the same video; for example, "a balloon floating upward while steam rises from a pot" requires buoyancy and fluid dynamics to unfold coherently and simultaneously. Yet existing methods largely ignore multi-principle interactions, focusing on a single principle per video. We propose HiPhy (Hierarchical Physical Alignment), a reinforcement learning framework that grounds video generation in physical laws through a dual-level objective: locally enforcing the temporal dynamics of individual physical principles, and globally ensuring the physical and semantic coherence of the entire scene. To support multi-principle generation, we construct a 50K-prompt dataset and introduce a prompt benchmark MultiPhyBench, spanning a diverse range of co-occurring physical events. Our experiments show that HiPhy significantly outperforms prior methods and baselines, improving physical commonsense and semantic alignment significantly across various benchmarks, with the largest gains on scenes involving multiple concurrent physical principles where competing methods degrade most sharply.
PaperID: 2806, Poster
Abstract: Automatic Prompt Optimization (APO) aims to improve the capabilities of Large Language Models (LLMs), yet existing methods remain limited by their result-oriented nature. Training-free approaches often converge to suboptimal prompts, while training-based methods suffer from reward sparsity and optimization instability. More importantly, both paradigms focus on obtaining the final prompt rather than learning the prompt optimization process itself. To address this dilemma, we propose ProPO, a trajectory-driven Process learning framework for Prompt Optimization guided by multi-dimensional composite rewards. Specifically, ProPO introduces a historical trajectory reflection mechanism for high-quality exploration. Concurrently, we adopt adaptive rubric rewards to provide fine-grained supervision and alleviate reward sparsity. More importantly, we design a novel Process Logic Reward that supervises the consistency of the reflection process. Extensive experiments demonstrate that ProPO, integrated with GRPO, establishes a new SOTA performance. Notably, ProPO achieves a remarkable 70.00% accuracy on the highly complex AIME-2026 task and an absolute gain of +25.25% over the zero-shot baseline on Subjectivity Classification (Subj). Furthermore, empirical analyses reveal that our multi-dimensional reward mechanism provides finer-grained and denser signals, allowing the model to learn the prompt optimization process, and ultimately enabling it to generate significantly higher-quality prompts.
PaperID: 2807, Poster
Abstract: Tabular foundation models (TFMs) are promising for learning reusable priors for heterogeneous tables in low-data regimes. Yet most TFMs are trained to predict labels or reproduce passively observed rows. In small downstream tables, these observational objectives can rely on noisy empirical correlations rather than infer a table-level structure that supports coherent generation and adaptation. We propose , a world-model-inspired generative TFM that learns a table-adaptive simulator: from support rows and a query context, it generates full rows and, when a structural causal model (SCM) action is supplied, predicts structured action responses. TabWorld combines observational diffusion training on real tables with query-conditioned world-model supervision, using real tables for support-conditioned completion and synthetic SCMs for oracle pre/post-action labels unavailable in ordinary observational data. Across unseen-table generation, completion, low-data adaptation, and SCM intervention diagnostics, TabWorld improves post-adaptation fidelity, utility, and SCM intervention-diagnostic behavior over observational-only baselines.
Abstract: Mixture-of-experts (MoE) is a common approach for increasing parameter capacity, but applying MoE to state space model (SSM) token mixers can multiply the cost of the recurrent state update. We study how to introduce expert specialization into selective SSMs while preserving computational efficiency. We show that MoE--SSM can refer to two designs: (1) MoE over separated SSMs, which maintains multiple state trajectories and thus scales compute with the number of experts; and (2) MoE-parameterized SSM, which mixes experts in parameter space, maintains a single state trajectory, and evaluates the recurrence once. Our method, Switch Mamba (Swimba), follows the second design by routing over expert-produced SSM streams. Theoretically, we establish well-definedness and stability for MoE-parameterized SSMs and characterize the relationship between the two designs. Empirically, we evaluate Swimba on standard benchmark tasks and measure real-time throughput and latency. Under matched FLOPs, Swimba achieves slightly better average performance than the baseline, with a small slowdown in real-time latency and throughput. Overall, these results suggest that parameter-space MoE can increase SSM capacity while keeping the dominant recurrence cost fixed.
PaperID: 2809, Poster
Authors: Hongyuan Su, Yu Zheng, Yong Li
Abstract: Designing high-performing robot morphologies is a grand challenge for developing specialized autonomous agents. However, the vast, combinatorial, and non-differentiable nature of the morphological design space has been a primary obstacle. Existing methods tackle this problem , relying on either semantically-blind genetic operators or reinforcement learning with predefined modification actions, both of which constrain exploration. In this work, we introduce , a novel framework that reframes morphological design as a code generation problem. MorphoGen leverages large language models (LLMs) to iterate the XML files as codes that define an agent’s morphology, solving the original open problem without being limited by any prior constraints or fixed action spaces. Structure-Aware directional feedback is provided to steer the evolution of robot morphologies through prompted mutations and crossovers. Our approach allows the LLMs to apply its understanding of structure and syntax to generate complex and semantically coherent design variations, enabling an unconstrained and efficient exploration of the design space. On a suite of challenging locomotion benchmarks, MorphoGen discovers novel and high-performing morphologies, significantly outperforming strong baselines by over 52.9% in downstream motoring evaluation. Our work unlocks a new paradigm for automated robotic design, demonstrating the effectiveness of LLMs in navigating complex, structured engineering search spaces. Codes for our work are released anonymously at https://anonymous.4open.science/r/MorphoGen-ACC.
PaperID: 2810, Poster
Abstract: We identify a theoretical incompatibility between 3D point cloud masked autoencoding and SE(3)-equivariant networks. We prove that standard masking triggers a topological collapse, effectively forcing the model to hallucinate orientation without a reference frame. To resolve this, we introduce SIEVE (Scalar Invariant Extraction Via Equivariance), an SE(3)-equivariant self-supervised framework that compresses point cloud geometry into invariant scalar embeddings. At its core is the spherical projection mask, which preserves global pose while obfuscating local semantics, sidestepping the obstruction by construction. Experiments confirm that SIEVE circumvents the predicted collapse where constant and noise masks fail. On downstream point cloud tasks, the resulting embeddings improve over coordinate-based baselines and hand-crafted geometric descriptors, with the largest gains on generalization to unseen shape categories.
Authors:
mingchen li, Jiatan Huang, Chuxu Zhang, Liang Zhao, Hong YuAbstract: In-context learning has recently been linked to implicit gradient descent in linear self-attention models, suggesting that context can induce a forward-pass update. Retrieval-augmented generation (RAG) also relies on context, but retrieved documents are usually treated as static evidence rather than signals for adaptation. We study RAG as an in-context optimization process. First, we show that one linear self-attention layer can implement one gradient-descent step on a unified linearized RAG objective covering both projection-based and dot-product retrieval interfaces. This gives an exact regime where retrieval-augmented prediction and in-context optimization coincide. We use this result not as a literal model of LLM computation, but as a guide for adapting the interaction between queries and retrieved evidence. We then test the boundary of this correspondence: it remains stable under controlled linear extensions, but becomes feature-distribution dependent under nonlinear architectures. Finally, we turn this view into a lightweight method for frozen RAG LLMs. The method keeps the retriever and backbone fixed, and predicts a context-conditioned update to a generator-side evidence-use interface. Across seven QA benchmarks, two retrievers, and two frozen LLM backbones, this forward-only update improves a shared-interface baseline, transfers to held-out tasks, and approaches test-time gradient adaptation at much lower per-query cost
PaperID: 2812, Poster
Abstract: Recovering dynamical systems from noisy observations is a recurring challenge across scientific domains, including neuroscience and physics. Latent stochastic differential equations (SDEs) address this by modeling the system as an unobserved state that evolves according to a learnable SDE and generates the observations. Variational inference (VI) provides a tractable objective for fitting latent SDEs. Traditional VI algorithms evaluate this objective by numerical simulation over a time discretization, trading fidelity for computational cost. A recent class of algorithms, simulation-free VI, sidesteps this tradeoff by parameterizing the posterior through its instantaneous marginals rather than its drift. In this work, we show that the efficiency of simulation-free VI algorithms comes at a price: their parameterizations restrict the approximate posterior to a subset of the SDEs available to simulation-based methods, degrading both posterior inference and parameter learning. We propose Helmholtz-SDE, a simulation-free VI algorithm that closes this gap by optimizing over path laws compatible with a prescribed collection of marginals. Helmholtz-SDE recovers dynamics more faithfully than prior simulation-free methods, with the largest gains under high posterior uncertainty. It further matches the performance of simulation-based VI at a fraction of the runtime.
PaperID: 2813, Poster
Abstract: Recent advancements in generative models have increasingly leveraged self-supervised visual representations to guide the synthesis process, significantly improving generation quality and semantic consistency. However, the reverse paradigm of harnessing the generative process to enhance visual representations remains largely underexplored. In this paper, we investigate the semantic properties of class tokens synthesized by generative models. We observe that generative models equipped with representation entanglement can generate class tokens that exhibit stronger discriminative capabilities than those extracted from the original pretrained visual models. Motivated by this observation, we propose a completely new Generation-to-Perception Knowledge Distillation (GPKD) framework, where our method generates class tokens as teacher signals to instill global semantics into the student model. To prevent the degradation of local details, we further incorporate a new masked patch-level distillation objective. This dual distillation strategy enhances global representations while mitigating the forgetting of local details, thereby producing more robust representations. Extensive experiments demonstrate that GPKD obtains consistent improvements across various downstream tasks compared with the base visual encoder, achieving a gain of more than 2% in k-NN accuracy on ImageNet-1K.
PaperID: 2814, Poster
Abstract: Designing DNA libraries is a key challenge from drug design to protein engineering and synthetic biology. Modern generative models offer opportunities to navigate the design space and propose specific sequences predicted to be effective in-silico. Designing deterministic libraries of specific sequences is however limited by the cost of DNA synthesis -- the synthesis barrier. In contrast, high-throughput multiplexed screening can measure the function of billions of biological sequences in parallel. Harnessing this technology requires the design of randomized libraries with specific design constraints to achieve low synthesis costs. In practice, such stochastic libraries are often chosen heuristically, sacrificing control for scale. Is there a way to bridge AI-based in-silico sequence design with high-throughput experimentation? In this work, we introduce Policy Gradients for Library Design (PGLD). PGLD uses a synthesis-aware parametrization of stochastic DNA libraries and optimizes them against a specified objective function. This allows for designing massive, controlled libraries without being limited by synthesis costs. We show how PGLD enables lab-in-the-loop design of multi-round high-throughput experiments, and large-scale in-vitro DNA sampling from generative models. Finally, we use PGLD to design a library of ~ 10^6 unique sequences at a cost of ~ 700 USD to explore the mutation space of a broadly neutralizing influenza antibody.
PaperID: 2815, Poster
Abstract: Dense associative memory underlies modern attention mechanisms; however, conventional Modern Hopfield Network (MHN) with linear kernels yields overlapping attractor basins and spurious states under high memory load, limiting their theoretically exponential storage capacity in practice. We construct a class of nonlinear phase kernels that reshape the energy dynamics into deep, well-separated basins, and prove that any kernel in this class preserves exponential storage capacity while guaranteeing convergence to fixed points. As a representative instantiation, we introduce the Sigmoidal Phase Hopfield Network (S-Hop), which guarantees monotone energy descent and mitigates the practical loss of memory capacity. Experiments on MNIST and CIFAR-10 show that S-Hop achieves up to a 50× increase in critical capacity over baseline models. We also provide a systematic capacity measurement of Hopfield networks on TinyImageNet, where S-Hop exhibits improved retrieval robustness and noise resilience. The phase-kernel design provides a principled route toward more robust dense associative memory.
Abstract: In-context watermarking (ICW) prepends an instruction to a query asking the model to embed a statistically detectable signal in its response. It thus equips LLMs with a watermarking interface that third parties can invoke without access to model internals. Its reliability hinges on the LLM following the instruction without degrading answer quality, yet how well current LLMs do so has not been measured. We introduce ICWBench, a benchmark of three verifiable ICW instruction families, each scored on both detectability and answer quality. Evaluating 14 frontier proprietary and open-source LLMs, we find that none of the evaluated LLMs achieves both objectives across all three families. To address this, we propose a self-contained two-stage training method, requiring no distillation from a stronger model, no manual annotation, and no pre-existing ICW IF ability. The first stage, self-distillation with logits perturbation (SDLP), uses the same base LLM as both teacher and student: an instruction-equivalent decoding-time logits perturbation makes the teacher follow the ICW instruction, and the student is trained to match the teacher's output distribution. The second stage applies reinforcement learning with the automatic verifier as the reward. Applied to Qwen3-14B, the weakest of the 14 evaluated LLMs in ICW IF, our method raises average TPR@1%FPR across three ICW instructions from 0.100 to 0.974, achieving a more favorable trade-off than the frontier proprietary LLMs we evaluate.
PaperID: 2817, Poster
Abstract: Bayesian optimization (BO) is widely used to accelerate molecular discovery by reducing costly oracle evaluations. Foundation models provide a promising source of prior knowledge for BO, but simply using their representations in standard surrogate-based pipelines is often unreliable in low-data regimes with high-dimensional features and vast discrete candidate spaces. Rather than asking whether LLMs are universally useful for molecular BO, we ask how Acquisition Functions (AFs) should be designed to exploit weak, high-dimensional, and partially informative foundation-model priors. We propose \emphLLM-guided Acquisition Tree (LLMAT), a surrogate-free BO framework that reformulates likelihood-free acquisition estimation as recursive local acquisition learning. LLMAT trains binary classifiers on LLM representations, where each classifier jointly induces a promising/non-promising partition and defines a local AF within the corresponding region. This produces a hierarchy of localized AFs and enables efficient candidate selection via Monte Carlo Tree Search. To stabilize learning with few observations, LLMAT meta-learns shared classifier parameters and initializations across tree nodes. An optional LLM-guided clustering module further reduces AF evaluation cost by restricting search to statistically promising coarse property clusters. Extensive experiments and ablations demonstrate substantial improvements in scalability, robustness, and sample efficiency for LLM-guided BO in molecular discovery.
PaperID: 2818, Poster
Authors:
Wu Lin, Yan Fan, Chenxiang Ma, Yinglan Feng, Jia Wang, Ka-Chun Wong, Qiuzhen LinAbstract: Backpropagation (BP) remains the predominant algorithm for training deep neural networks, but it is biologically implausible and suffers from the backward locking problem, constraining parallelization across layers. Direct Feedback Alignment (DFA), which propagates the global output error directly to each hidden layer via fixed feedback matrices, has emerged as a promising alternative for addressing these limitations. However, a significant performance gap persists between DFA and BP, which we attribute to the limited ability of fixed feedback matrices to capture the nonlinear transformation from hidden representations to the network output. To address this limitation, we propose Nonlinear Direct Feedback Alignment (NDFA), which constructs each hidden-layer feedback using a lightweight nonlinear surrogate subnetwork. To enhance the effectiveness of the feedback throughout training, we further propose a learning strategy to update the surrogate subnetwork, thereby better approximating the underlying nonlinear transformation. Extensive experiments across residual networks and vision transformers on CIFAR-10, CIFAR-100, SVHN, and ImageNet datasets demonstrate that NDFA consistently outperforms existing DFA-based methods and substantially narrows the performance gap to BP, while maintaining efficient parallel training. Code will be publicly available after the review process.
PaperID: 2819, Poster
Authors: Andrei Rusu, Andrei Dumitrescu, Adrian C Badea, Cosmin Maria, Ion-Marian Anghelina, Bai Li, Tomasz L Religa, Dragos H Bobolea
Abstract: Prompt optimization has recently been approached with increasingly specialized search procedures and evolutionary frameworks. In this work, we ask whether such complexity is necessary to achieve competitive performance. We introduce REAPO (ReAct Agents for Prompt Optimization), a lightweight framework that casts prompt optimization as an agentic process built around a ReAct loop augmented with evaluation and reflection tools. Rather than relying on bespoke optimization machinery, REAPO iteratively refines prompts through error analysis, reflective reasoning, and validation against the optimization objective. We evaluate REAPO on several established benchmarks spanning multi-hop question answering, mathematical reasoning, fact verification, instruction following, and structured classification, as well as on real enterprise agentic tasks. Across these settings, we find that this agentic optimizer can match or outperform specialized prompt tuning methods while requiring 2–7× fewer evaluation rollouts and less auxiliary optimization logic. We also report confidence intervals over multiple independent runs and show that single-run evaluations on small benchmarks can give a misleading picture of relative performance, since LLM variance may obscure true gains. Taken together, our results suggest that effective prompt optimization need not depend on increasingly elaborate optimization schemes, and that agentic methods can provide a more efficient alternative.
PaperID: 2820, Poster
Abstract: Self-evolving agents have shown strong empirical performance in mathematical discovery and code optimization, where the challenge is that the search space is too large and unstructured for exhaustive exploration. However, most existing methods lack systematic modeling of self-evolution, resulting in inefficient and even redundant component designs. In this paper, we model self-evolution through the lens of Partially Observable Markov Decision Process (POMDP). We treat the contexts surrounding each search node (the current program and its information to be evolved) as history, while the node’s state, shaped dynamically by the ongoing search process, remains latent and unobservable. This perspective combines parent-node selection, prompt transformation, and child-code generation into a single sequential decision framework, and better aligns with the optimization process of self-evolution. We then develop ZetaEvolve, a search strategy that improves Predictor Upper Confidence bound applied to Trees (PUCT) with a history-conditioned mechanism tailored to partial-observation signals in self-evolving search. Empirically, ZetaEvolve outperforms strong baselines across various benchmarks and achieves faster convergence than existing SOTA methods.
PaperID: 2821, Poster
Abstract: Dense stereo disparity estimation provides the geometric basis for 3D perception in computer-assisted endoscopy. Endoscopic stereo matching is difficult because reliable tissue correspondences often appear next to weak-texture, specular, occluded, or deformed regions within the same frame. Existing stereo and stereo-monocular methods improve cost aggregation or introduce monocular priors, but they often treat monocular guidance as a global feature source or an independently predicted candidate. Without evaluating the local support of the current stereo match during refinement, the update can reinforce corrupted correspondences in ambiguous regions or allow monocular priors to override reliable metric matches. To tackle these problems, we introduce Manifold-Stereo, a reliability-gated latent-flow framework for endoscopic stereo disparity estimation. The method refines a monocular disparity candidate through a conditional latent residual flow on an image-aligned correction space induced by stereo cost-volume geometry, then calibrates the corrected candidate to the current stereo disparity scale. A geometric consistency gate computed from left-right feature alignment regulates the candidate contribution during recurrent refinement, preserving stereo estimates in well-matched regions while supplying monocular context where local matching is ambiguous. Experiments on in-domain SCARED and zero-shot SERV-CT and Hamlyn show strong average disparity accuracy and competitive thresholded outlier rates against recent stereo and endoscopic baselines.
PaperID: 2822, Poster
Abstract: Self-supervised gait representation learning aims to learn transferable representations that generalize to unseen scenarios. However, existing methods rely solely on visual self-supervision, making it difficult to assess the semantic value of individual unlabeled sequences and limiting their ability to exploit latent semantic cues. To address this limitation, we introduce language priors into self-supervised gait pretraining. We construct GaitLP-1M, a million-scale silhouette-based gait pretraining dataset without identity labels. A two-stage caption generation pipeline further produces gait descriptions from the corresponding RGB sequences. To the best of our knowledge, GaitLP-1M is the first million-scale silhouette gait dataset paired with gait-centric captions. Building on this foundation, we propose GaitLingo, a language-guided self-supervised framework for gait representation learning. Specifically, we design a Caption-Guided Sample Weighting (CSW) mechanism to prioritize informative samples based on sequence-level semantic richness and attribute rarity. We then introduce a lightweight Temporal Adapter (TA) to better capture motion semantics and improve cross-domain robustness. Finally, we propose Soft Relation Distillation (SRD) to transfer caption-derived relational structures into the visual embedding space. Extensive experiments demonstrate that GaitLingo consistently outperforms prior self-supervised gait pretraining methods in the zero-shot setting across six benchmarks. Compared with GaitSSB, it achieves Rank-1 gains of +12.1% on Gait3D and +7.9% on GREW, as well as an average improvement of +9.1% across four CCPG settings.
Abstract: We consider the problem of synthesizing Clifford quantum circuits for devices with all-to-all qubit connectivity. We approach this task as a reinforcement learning problem in which an agent learns to discover a sequence of elementary Clifford gates that reduces a given symplectic matrix representation of a Clifford circuit to the identity. This formulation permits a simple learning curriculum based on random walks from the identity. We introduce a novel neural network architecture that is equivariant to qubit relabelings of the symplectic matrix representation, and which is size-agnostic, allowing a single learned policy to be applied across different qubit counts without circuit splicing or network reparameterization. On six-qubit Clifford circuits, the largest regime for which optimal references are available, our agent finds circuits within one two-qubit gate of optimality in milliseconds per instance, and finds optimal circuits in 99.2% of instances within seconds per instance. After continued training on ten-qubit instances, the agent scales to unseen Clifford tableaus with up to thirty qubits, including targets generated from circuits with over a thousand Clifford gates, where it achieves lower average two-qubit gate counts than Qiskit's Aaronson-Gottesman and greedy Clifford synthesizers.
PaperID: 2824, Poster
Authors: Adrien Aguila-Multner, Olivier Beaumont, Lionel Eyraud-Dubois, Julia Gusak
Abstract: Pipeline parallelism is a key technique for scaling deep network training across multiple devices. Recent works have significantly reduced pipeline idle time by improving scheduling efficiency. Decoupling the computation of gradients with respect to weights and activations led to the development of schedules with almost no idle time. However, these methods still require substantial memory, limiting their applicability on resource-constrained hardware. Our first contribution is to introduce recomputation to the backward pass, extending rematerialization beyond the forward pass. This enables executing schedules with decoupled gradient computations under much tighter memory constraints. Our second contribution is to consider more flexible rematerialization strategies, with individual per-microbatch decisions. We provide a unified optimization approach that, given a model and hardware memory constraints, formulates and solves an Integer Linear Programming (ILP) problem to determine the optimal per-microbatch, per-GPU rematerialization strategy for a given schedule, applicable to both one-wave and multi-wave pipeline schedules. With these tools, we show that when using rematerialization, the best scheduling algorithm varies according to the device memory constraints. Experiments demonstrate the effectiveness of all three contributions, showing that our approach enables efficient training of larger models under tight memory budgets, adapts optimally to varying memory capacities, and reduces recomputation overhead compared to existing recomputation solutions.
Abstract: Regret minimization (RM) and best-arm identification (BAI) are two fundamental objectives in multi-armed bandits. Among regret-minimizing algorithms, 1/2-Tsallis-INF is a canonical best-of-both-worlds FTRL algorithm: it achieves logarithmic pseudo-regret in stochastic bandits while retaining minimax-optimal regret in adversarial bandits, without knowing the environment in advance. This raises a natural question: can the same algorithm, without additional exploration, also identify the best arm reliably? We study this question in stochastic bandits by analyzing the failure probability e_t, defined as the probability that the empirical best arm determined by the cumulative importance-weighted loss estimates of 1/2-Tsallis-INF differs from the true optimal arm. The main difficulty is that, at the logarithmic-regret scale, suboptimal arms are sampled with probability heuristically of order 1/t. Consequently, importance weighting causes the cumulative estimator to fluctuate on the same linear scale as its mean separation. To overcome this obstacle, guided by a diffusion toy model, we construct a Lyapunov function for the gap process between the estimated cumulative loss of the optimal arm and that of the best competing arm. This leads to polynomial upper bounds on e_t: for learning rate \eta_t=\alpha/\sqrt t, e_t decays at rate t^-2+\alpha^2\mu_i_/4+\rho for any \rho>0, where \mu_i_ denotes the mean loss of the true optimal arm. We also establish a lower bound \Omega(t^-2-\varepsilon) for any \varepsilon>0, showing that the exponent 2 is essentially tight.
PaperID: 2826, Poster
Authors: Jiaqi Wang, Yi Feng, Xian Wu
Abstract: Multimodal Chain-of-Thought (CoT) reasoning has significantly advanced large models, yet forcing continuous visual evidence into discrete text or fixed tool calls inevitably sacrifices fine-grained detail. Recent paradigms attempt to reason directly within a continuous latent space; however, supervising latent slots in isolation fails to account for their collective geometry, resulting in a fragmented latent space that lacks structural guidance. To ensure structural integrity across training stages, we propose GeLVR(Geometry-Consistent Latent Visual Reasoning), a framework that establishes geometric consistency as a governing principle: the relational structure is explicitly aligned during SFT and actively preserved throughout RL, anchoring both stages to a common geometric scaffold. To operationalize this paradigm during the reward-driven phase, we further introduce GePO (Geometry-preserving Policy Optimization), a policy-optimization algorithm tailored for the hybrid discrete-continuous action spaces of latent reasoning. GePO employs a spherical von Mises-Fisher (vMF) policy to respect the decoder's inherent geometry and integrates the geometry-preserving regularizer directly into the RL objective. This ensures that reward optimization and structural integrity advance jointly rather than in tension. Extensive experiments across multiple visual reasoning benchmarks demonstrate that GeLVR consistently outperforms state-of-the-art baselines, particularly in high-resolution and fine-grained perception tasks. Comprehensive ablations further validate the necessity of each geometry-consistent component in establishing stable and coherent latent reasoning. Code is available in the supplementary materials.
PaperID: 2827, Poster
Abstract: Recent long-tailed recognition methods increasingly adopt two-stage fine-tuning, where a generic representation is first learned and imbalance-aware adaptation is then performed in a lightweight second stage. However, existing second-stage designs often allocate objective emphasis using delayed empirical summaries, such as historical class-wise performance or validation feedback, and commonly restrict adaptation to classifier heads. These choices may make routing unreliable when empirical feedback is noisy or skewed, and may also limit the feature plasticity required for rare categories. In this paper, we propose U-MOF, an uncertainty-guided parameter-efficient multi-objective fine-tuning framework for long-tailed recognition. U-MOF introduces lightweight residual adapters into a shared-backbone expert committee, uses entropy-based committee statistics as diagnostic routing proxies to modulate tail-oriented and smoothing-oriented objectives, and applies conflict-aware orthogonal projection to coordinate heterogeneous objective gradients. Experiments on CIFAR-100-LT, ImageNet-LT, and iNaturalist 2018 show that U-MOF improves rare-category and overall performance in controlled stage-two comparisons and remains competitive with strong published long-tailed recognition baselines, while preserving the efficiency advantages of decoupled adaptation. These results indicate that reliable stage-two long-tailed fine-tuning benefits not only from selecting appropriate objectives, but also from diagnosing when each objective should dominate and from retaining limited feature-level plasticity. An anonymized implementation is included in the Supplementary Material.
Abstract: Fine-tuning aims to update a pre-trained flow-based generative model to improve the downstream reward of its generated samples. Existing methods typically frame this problem as sampling from a reward-tilted distribution, which arises as the solution to a KL-regularized reward-maximization problem. Here we take an alternative approach and introduce an regularizer built directly from the pre-trained drift; we show that the resulting fine-tuning problem is equivalent to a on the flow. Assuming access to a pre-trained flow map, we exploit this equivalence to devise a simulation-free reinforcement-learning algorithm for fine-tuning generative flows. We call the resulting framework , the first end-to-end fine-tuning recipe native to flow maps. The output of our approach is itself a fine-tuned flow map, retaining few-step reward-aligned inference at deployment. Numerical experiments at text-to-image scale highlight both the efficiency and the efficacy of our approach. More broadly, our work argues that accelerated samplers such as flow maps are essential infrastructure for efficient post-training, and that the dominant KL-regularized formulation of fine-tuning is only one of many choices worth revisiting.
PaperID: 2829, Poster
Abstract: We study post-selection-safe deployment in finite-horizon tabular offline multi-objective reinforcement learning with full vector rewards, nonnegative linear scalarization, and preferences revealed at deployment. A fixed dataset may already have been reused to construct a data-dependent policy library \Pi_N; the remaining problem is to return auditable lower-confidence utilities for future queried preferences and for small policy menus. We introduce the Pessimistic Utility Oracle (PUO), a coordinatewise pessimistic vector-value construction. On one dataset-level tabular model-confidence event, PUO lower-bounds every data-dependent policy under every w\in\mathcalW\subset\mathbbR_+^d, without preference discretization or an additional library-size union bound beyond the model event. PUO yields a queried-preference planner and SPS++, a greedy policy-set method. The main SPS++ theorem evaluates the executable rule that deploys the in-set policy with largest PUO score, and decomposes its loss into coverage, greedy approximation, and preference-sampling terms. Controlled tabular audits validate these statistical effects; trajectory-count and linear-feature appendices are presented only as bridges beyond the main generative tabular theorem.
PaperID: 2830, Poster
Abstract: Chain-of-Thought reasoning has enabled large language models to achieve substantial performance gains on complex tasks. However, these gains come at the cost of dramatically increased token consumption. This raises a fundamental question: is every token in the reasoning trace equally valuable? We present a diagnostic and optimization framework grounded in a key empirical finding: the value of tokens within a CoT reasoning sequence is highly non-uniform, and this non-uniformity can be effectively characterized by token-level log probability signals. We show that normalized log probability helps distinguish core tokens, which carry structural and decisive reasoning content, from redundant tokens, which are exploratory, low-confidence filler that contributes less directly to the final answer. Building on these findings, we formulate the TokenProbe framework around two empirical findings and one claim: findings identify token value inequality first and then establish TokenProbe as a core-token proxy, and the claim introduces an efficient GRPO objective positing that selectively compressing redundant tokens can yield Pareto improvements in the accuracy–token efficiency space. Empirically, our method preserves reasoning quality while reducing the token usage by 76% of the baseline. Under matched reasoning-length budgets, we show that it can even outperform strong flagship baselines like Gemini-3.1-Pro.
PaperID: 2831, Poster
Abstract: Sharpness-Aware Minimization (SAM) is widely used in domain generalization to promote flat minima that generalize to unseen distributions. Existing SAM-based methods typically penalize the sharpness of the average loss across domains. We show that this prevailing formulation—termed AvSAM—has a fundamental geometric limitation: it can favor solutions with misaligned domain-wise dominant Hessian eigenvectors. As a result, such solutions may appear flat on average while remaining sharp within individual domains. To address this limitation, we propose an alternative objective that penalizes the average domain-wise sharpness. We show that this objective decomposes into the AvSAM objective plus a non-negative residual term, which we call the Cancellation Gap. This gap is minimized if and only if the domain-wise dominant eigenvectors are aligned up to sign, indicating that it mitigates AvSAM's tendency to favor misaligned solutions. However, we further uncover a degeneracy: the Cancellation Gap can become small without improving alignment when sharpness is highly imbalanced across domains. To mitigate this failure mode, we propose Cancellation Gap Minimization (CGM), which augments the objective with a squared coefficient-of-variation (\mathrmCV^2) regularizer to discourage sharpness imbalance. Experiments show that CGM achieves the best average performance across standard domain generalization benchmarks among strong baselines. Diagnostic analyses further confirm reduced domain-wise sharpness and improved alignment of dominant Hessian eigenvectors.
Authors:
Yu Huang, Zihua Zhao, Zhaoxin Huan, Wanli Gu, Feng Hong, Xinmu Ge, Lin Yuan, Qiang Hu, Weichang Wu, Xiaolu Zhang, Jun Zhou, Jiangchao YaoAbstract: The open-ended generation in LLMs usually requires multi-dimensional rubrics to adequately assess quality and guide the improvement of reinforcement learning. However, a critical dilemma inherent in this training paradigm is the imbalanced reward polarization along different rubric dimensions. Under this bottleneck, even if LLMs achieve relatively high rewards after training, they may still exhibit severe deficiencies in certain dimensions, leading to a direct deterioration in user experience. To address this problem, we propose \emphFocal Reward, a novel objective to automatically balance the training of reinforcement learning under rubric-based rewards. Specifically, we first leverage an inverse reward projection mechanism to estimate the saturation degree of each criterion in the rubric, which forms the basis to calibrate the reward direction. Then, the final objective is designed with an automatically reweighting coefficient for each criterion to achieve the fine-grained balancing. Extensive experiments across three model scales and six benchmarks demonstrate that our \emphFocal Reward method outperforms the strongest static aggregation baseline in all 18 model--benchmark comparisons. Rollout, mechanism, and ablation analyses further show that these gains arise from online, saturation-aware reallocation toward rubrics that still have room for improvement.
Abstract: Reward models are central to aligning language models with human preferences via reinforcement learning (RL). As RL is increasingly applied to settings such as verifiable rewards and multi-objective alignment, RMs are expected to encode more complex and multifaceted preference distributions. However, classifier RMs remain static once trained, limiting their adaptability at test time. We propose Variational In-Context Reward Modeling (ICRM), a novel Bayesian reward modeling objective that enables test-time steerability via in-context preference demonstrations. ICRM casts reward modeling as amortized variational inference over a latent preference probability under the Bradley-Terry model using a conjugate Beta prior. We show that ICRM adapts to unseen preference distributions at test time for both single and multi-objective settings. With more demonstrations, ICRM improves RM-Bench accuracy from 60.5 to 70.8, achieves lower calibration error than a generative judge on moral dilemma preferences, and expands the attainable Pareto frontier under conflicting preferences. We further study the practical applicability of ICRM for RL training, showing that it can effectively encode verifiable rewards by outperforming a conventional RM in math reasoning. Finally, we provide theoretical guarantees that the variational objective admits a global interior optimum with finite confidence, and we analyze how KL regularization mitigates reward over-optimization.
PaperID: 2834, Poster
Abstract: With the rapid growth of surveillance data, Person Re-Identification (ReID) is evolving toward large scale scenarios, where datasets exhibit substantial increases in both size and diversity. Under this trend, the widely adopted dense backbone architectures, such as Transformers, suffer from an inherent limitation: Shared parameters couple all input tokens, forcing the model to learn generic representations that are suboptimal for heterogeneous data sources. Although Mixture-of-Experts (MoE) architectures offer a promising solution for handling diverse data, they encounter a fundamental challenge in ReID. Specifically, person images exhibit high inter-instance similarity at the coarse level, while discriminative cues lie in subtle fine-grained details that are easily overlooked. As a result, existing MoE methods often suffer from severe routing collapse, leading to insufficient expert specialization and limited diverse knowledge acquisition capacity. To address these problems, we propose a novel framework, termed Hyperspherical Mixture-of-Experts for Large Scale Person Re-Identification (HypMoE-ReID). Specifically, motivated by theoretical analysis in hyperspherical space, HypMoE-ReID incorporates two key components: (1) a Token Purification mechanism that suppresses noisy tokens detrimental to expert specialization, and (2) a Routing Diversity Learning strategy that explicitly separates experts on a unit hypersphere to enforce specialization and prevent collapse. Furthermore, to facilitate the investigation in large scale ReID, we construct a large-scale ReID benchmark comprising 809,701 images by unifying 13 public datasets with significant distribution gaps. Furthermore, extensive experiments demonstrate that HypMoE-ReID outperforms state-of-the-art MoE-based methods by at least 11.3%/10.3% in Average mAP/R@1, and surpasses strong dense backbone models by 2.2%/2.0% in Average mAP/R@1. Our code will be released.
PaperID: 2835, Poster
Abstract: The prevailing route to stronger visual reasoning in vision-language models (VLMs) builds costly Chain-of-Thought corpora or curated step-by-step rationales, concentrating progress in well-resourced industry labs. We ask whether the reasoning signal already present in generic visual instruction data can instead be unlocked through a structural change to the post-training interface itself. At fixed data and capacity, standard supervised fine-tuning (SFT) on a generic instruction mixture partially improves reasoning but homogenizes the modality-depth pathways through which reasoning and perception inputs flow. We isolate two coupled patterns. First, visual and textual representations couple at the deepest layer: cross-modal Centered Kernel Alignment (CKA) rises by roughly 12% over the frozen baseline. Second, perception and reasoning inputs are routed through similar late-depth profiles, leaving no mechanism for per-input depth allocation. We propose , a lightweight side pathway that adds modality-specific query projections and a sigmoid-gated learnable depth residual. Under a matched protocol on Qwen3-VL-2B, MDA improves the reasoning average over SFT by +3.9 and keeps perception almost intact, at 0.05% extra parameters. Multiple mechanistic probes converge on the same routing account. The gains extend to Qwen3-VL-8B and transfer to LLaVA-OneVision-7B, suggesting that modality-depth routing may be a general lever for unlocking reasoning from generic instruction data without compromising perception. Code will be released upon publication.
PaperID: 2836, Poster
Abstract: Distributional treatment effects can be invisible to means: a treatment may preserve average outcomes while changing tails, modes, dispersion, or rare-event probabilities. Kernel tests can detect discrepancies between interventional outcome laws, but global tests do not reveal where the laws differ. We propose DR-ME, to our knowledge the first semiparametrically efficient finite-location test for interpretable distributional treatment effects. DR-ME evaluates an interventional kernel witness at learned outcome locations, returning causal-discrepancy coordinates rather than only a global rejection. From observational data, we derive orthogonal doubly robust kernel features whose centered oracle form is the canonical gradient of this finite witness. For fixed locations, we characterize the local testing limit: DR-ME is chi-square calibrated under the null, has noncentral chi-square local power, and uses the covariance whitening that optimizes local signal-to-noise for discrepancies visible through the selected coordinates. This efficient local-power geometry yields a principled location-learning criterion, with sample splitting preserving post-selection validity. Experiments show near-nominal type-I error, competitive power against global doubly robust kernel tests, and interpretable learned locations that localize distributional effects in a semi-synthetic medical-imaging study.
PaperID: 2837, Poster
Abstract: Scaling Graph Neural Network (GNN) training to large graphs commonly relies on mini-batch sampling, but incomplete neighborhoods can degrade approximation quality and training stability. Memory-enhanced methods mitigate this issue by reusing historical node embeddings to approximate full-neighborhood aggregation; however, they reconstruct pseudo full-neighborhood messages inside every training iteration, repeatedly fetching out-of-batch memories and aggregating over large neighborhoods. In this work, we identify this under-examined inner-loop bottleneck and propose IncAgg, an efficient memory-enhanced GNN training framework based on incremental aggregation. IncAgg periodically pre-aggregates historical embeddings into a global memory baseline, and during mini-batch training propagates only the current in-batch residual before fusing it with the stored global context. This design preserves historical full-neighborhood information at a configurable refresh frequency while avoiding per-iteration out-of-batch memory fetching and aggregation, without modifying the underlying GNN architecture. We provide theoretical analyses showing that IncAgg reduces gradient-carrying aggregation and memory-access costs while retaining the residual-based stability principle of memory-enhanced estimators. Extensive experiments on four large-scale benchmarks and four representative GNN backbones show that IncAgg matches or improves the accuracy of strong memory-enhanced baselines while achieving substantial training acceleration, with up to 18.6× speedup in individual settings and up to 13.8× geometric-mean speedup on large graphs.
PaperID: 2838, Poster
Abstract: Likelihood-based methods are widely used to learn the parameters of partially observable Markov decision processes (POMDPs), but their finite-sample behavior is not well understood. We establish a different kind of finite-sample guarantee. For any candidate POMDP parameter, we bound its parameter recovery error in terms of its likelihood gap to the ground truth, the structural conditioning of the observation and transition kernels, and the behavior policy used to collect data. The bound applies in both the undercomplete and overcomplete observation regimes, with the overcomplete case following naturally from the undercomplete analysis. Because the result applies to any sufficiently good candidate, it yields finite-sample guarantees for the maximum-likelihood estimator, OMLE-style algorithms, and any procedure that produces a parameter with controlled likelihood gap. This is particularly useful for downstream tasks such as infinite-horizon planning with state-dependent rewards, where accurate recovery of the kernels themselves is required to obtain near-optimal policies. The bound separates the contribution of the behavior policy from that of the model, making explicit how data-collection choices, including action coverage and mixing rate, accelerate parameter learning.
PaperID: 2839, Poster
Abstract: We propose WALLA, a decentralized mechanism for aggregating probabilistic predictions from multiple LLMs using wagering mechanisms. Each model reports a prediction and a learned wager reflecting its expected score advantage; predictions are then aggregated using wagers as weights. Our mechanism introduces a leave-one-out baseline that yields three key properties: dominant-strategy incentive compatibility under general belief structures, a best-response wager proportional to expected score advantage, and decoupling of prediction and wager optimization. We instantiate two mechanism variants trading off normality and no-arbitrage, both with bounded worst-case deficit independent of the number of participants. Experiments on QA benchmarks and a forecasting benchmark show that WALLA matches centralized baselines while simultaneously achieving advantage-weighted aggregation, uncertainty-awareness, fully decentralized learning, and incentive-compatibility guarantees.
PaperID: 2840, Poster
Abstract: Enhancing Vision-Language Models (VLMs) on downstream tasks with unlabeled data has attracted increasing attention. Constructing a pseudo-labeled dataset by pseudo-labeling is a common method. Nonetheless, owing to the inherent class-wise prediction bias in VLMs, they tend to generate long-tailed pseudo-labels that are inconsistent with the true label distribution. Motivated by this, we first divide all classes into pseudo-head, pseudo-middle, and pseudo-tail classes based on the pseudo-label quantity. Accordingly, we propose a alancing (PPSB) framework, which measures the class-wise prediction bias and calculates quantity upper bounds for certain classes to construct relatively balanced pseudo-labeled dataset, while introducing fewer incorrect pseudo-tail labels. In this way, the class-wise prediction bias is progressively corrected and the pseudo-label dataset scale increases autonomously during training. Furthermore, the confidence scores generated by VLMs are not entirely reliable. Therefore, we design a neighborhood consistency filtering mechanism and a visual classifier to improve the pseudo-label accuracy. We theoretically prove that our method can significantly reduce the generalization error under certain conditions. Extensive experiments across six benchmarks under three learning paradigms demonstrate that our method outperforms state-of-the-art methods by average
Authors:
Zhengding Hu, Mingge Lu, Zhen Wang, Jixuan Ruan, Chang Chen, Zaifeng Pan, Yue Guan, Ruiyi Wang, Zhongkai Yu, Chao Zhang, Yufei DingAbstract: LLM-based evolution has emerged as a promising way to improve agents by refining non-parametric artifacts, but its wall-clock cost remains a major bottleneck. We identify that this cost comes from synchronized stage execution and imbalance inside each LLM-heavy stage. We present FlashEvolve, an efficient framework that replaces synchronized execution with asynchronous workers and queues, allowing different stages and steps to overlap. To handle data staleness introduced by asynchrony, FlashEvolve tracks artifact versions and applies different policies to update, discard, or patch stale artifacts. Unlike weight-space staleness in asynchronous RL, language-space staleness is inspectable and repairable: a stale artifact is not just delayed work, but readable evidence that the LLM can reflect on, revise, and turn into useful evolution signal. FlashEvolve further improves throughput and token efficiency with speculative stage completion and adaptive workflow control. On GEPA workloads, FlashEvolve improves proposal throughput by 3.5x on local vLLM and 4.9x on API serving over synchronous GEPA. The same design also applies to ACE and Meta-Harness. Our code repository is available at https://anonymous.4open.science/r/FlashEvolve-FEC7/.
Authors: Alireza Saleh Abadi, Leen-Kiat Soh, Daniel A Redder, Adam Eck, Prashant Doshi
Abstract: Open agent systems (OASYS) are increasingly prevalent in real-world domains where the sets of agents and tasks change unpredictably over time. Such openness, including agent openness (AO) and task openness (TO), poses a fundamental challenge to multi-agent reinforcement learning (MARL), which typically assumes fixed state and action spaces. Existing methods address openness only partially: padding and masking approaches introduce artificial bounds, while recent graph-based or hypergraph methods handle one dimension of openness but still depend on restrictive assumptions. In this paper, we introduce Pointer Learner for Agent and Task Openness (PLATO), a pointer-network-based actor combined with a centralized graph neural network (GNN) critic, trained with multi-agent proximal policy optimization under a centralized training and decentralized execution paradigm. Our pointer-based actor outputs distributions directly over the current task set. This directly supports changing action spaces without masking or retraining. Our GNN critic encodes agent–task interactions as a graph that changes shape with task and agent composition. Together, these components consider AO and TO without the boundedness of existing approaches. We formalize PLATO in a Task-and-Agent-Open Markov Game (TaAgO-MG), extending prior task-open formulations, and prove it is well-defined over the resulting unbounded state and action spaces. We evaluate PLATO with the Methods for Open Agent Systems Evaluation Initiative (MOASEI) wildfire suppression domain, an environment designed for open multi-agent system evaluation, and we demonstrate strong performance and more consistent zero-shot generalization than state-of-the-art baselines in OASYS.
PaperID: 2843, Poster
Abstract: Verified code generation asks a large language model (LLM) to generate both an executable program and a machine-checkable proof that the program meets a formal specification, promising software that is correct by construction. The de facto workflow decouples the two halves of the problem: first synthesize a program, then attempt to prove it correct. We observe that this sequential pipeline can be both ineffective and inefficient in practice. A program generated without anticipating its proof can be subtly incorrect or structurally difficult to verify, forcing the LLM into brittle repair loops that alternate between patching the code and patching the proof. Inspired by Dijkstra's view that a program and its correctness argument should be developed hand in hand, we propose \framework, an LLM-based agentic workflow that first derives a unified program-and-proof plan from the specification, then elaborates the implementation and proof scaffold under this shared plan. To evaluate verified code generation in realistic settings, we further introduce Lean4Commit0, a project-level benchmark built by extracting core APIs from real-world software repositories and translating their requirements, including relational specifications across APIs, into Lean tasks. Using four frontier LLM backends, we evaluate \framework on Verina, AlgoVeri, and our Lean4Commit0 benchmark, where it achieves the highest solve rate in every benchmark--model setting. Compared with the stronger baseline, it improves solve rates by 4.6--11.2 percentage points and reduces per-task API cost by up to roughly 40% and wall-clock time by up to roughly 37% on the difficult subset of each benchmark.
PaperID: 2844, Poster
Abstract: Audio-visual representation learning commonly aligns paired audio and visual signals with sample-level contrastive objectives. However, real-world audio-visual pairs usually contain heterogeneous patch tokens: only a subset carries reliable cross-modal shared semantics, while many others mainly encode modality-private residuals. Aggregating all tokens into a single global representation can therefore introduce semantic noise into cross-modal alignment. Motivated by the minimal-sufficiency principle, we propose an attribution-guided shared-private decoupling framework to mitigate this issue. Our key idea is to identify shared-cover subviews that preserve dominant cross-modal information while suppressing modality-private residuals, and use them to supervise decoupled shared and private representations. To identify such subviews efficiently, we draw on transformer attribution method and introduce Layer-Truncated Attribution (LTA), which provides a lightweight estimate of token-level relevance to cross-modal shared information. We further adopt a teacher-student framework, where the teacher provides attribution-induced supervision and the student internalizes shared-private decoupling into its representations, without introducing extra inference cost. Experiments on retrieval, classification, and sound-prompted semantic segmentation show consistent gains over strong baselines. Compared with baseline, our method improves the average zero-shot R@1 by 4.0 across AudioSet and VGGSound retrieval, and improves AS20K classification mAP by 3.1, demonstrating the effectiveness of fine grained attribution guidance for reducing semantic noise.
Abstract: Transformers acquire in-context learning abilities through abrupt phases during training, often unfolding over multiple stages in which key circuits, such as induction heads, emerge. In this work, we characterize the dynamics underlying the emergence of such circuits across these stages. We focus on a synthetic associative recall task, where sequences are drawn from random maps between a permutation group and a vocabulary, and the model is required to complete the mapping of a permutation by retrieving the corresponding association from context. For this task, we study the gradient-flow trajectories of a simplified two-layer attention-only transformer. Leveraging symmetries in both the transformer architecture and the data distribution, we derive for both attention layers. We identify a conservation law that couples parameters across layers and governs their joint learning. This law reveals how initialization controls the over which such circuits emerge. In the vanishing-initialization limit, we characterize the gradient-flow trajectory, showing how training jumps from saddle to saddle. Finally, we provide empirical evidence across different architectural choices, validating our simplifications and extending the insights from our analysis beyond the simplified setting.
Abstract: Safe in-context reinforcement learning (ICRL) adapts online from interaction history without test-time parameter updates while controlling episode cost under a safety budget. Under out-of-distribution (OOD) deployment shifts, pretraining-only safe ICRL can give poor reward-safety tradeoffs because the remaining budget affects behavior only through frozen policy conditioning, not an explicit action-level check against predicted future cost. We propose a latent Q-Barrier shield that learns a context representation, latent dynamics, and an ensemble cost critic before deployment. Without parameter updates, the shield infers context from history and filters or softly reweights candidate actions using the remaining budget and predicted future cost. We prove a conditional, error-decomposed barrier-margin result: a Q-Barrier-satisfying action leaves the next latent-budget state with an approximately budget-safe continuation under the learned critic, up to Bellman and latent-prediction errors. Across five safe ICRL benchmarks, the shield improves deployment-time reward-safety tradeoffs over a strong safe-ICRL baseline: after a short context window, it achieves higher return in four of five benchmarks while matching or lowering average episode cost in all five.
Abstract: Learning to collaborate with previously unseen partners is a fundamental generalization challenge, known as Ad Hoc Teamwork (AHT). Existing methods often adopt a two-stage pipeline: first, a fixed population of teammates is generated, and second, an AHT agent is trained to collaborate with them. This separation limits coverage of behaviors and ignores whether the generated teammates are informative for the AHT agent to learn from. On the other hand, AHT agents are typically trained under the assumption that the training teammate set is uncontrollable, despite the fact that its composition strongly influences generalization. This paper presents a unified framework for AHT by reformulating the problem as an open-ended learning process between an AHT agent and an adversarial teammate generator. We introduce ROTATE, a regret-driven, open-ended training algorithm that alternates between improving the AHT agent and generating teammates that probe its collaboration deficiencies. Experiments across Overcooked and Level-Based Foraging tasks demonstrate that ROTATE significantly outperforms baselines on an unseen set of teammates, establishing a new standard for robust, generalizable teamwork.
PaperID: 2848, Poster
Abstract: Existing semi-supervised adversarial training (SSAT) methods typically employ a teacher-student framework where a teacher model provides supervisory signals for unlabeled data. However, they rely on static supervision---either hard pseudo-labels or soft labels with fixed temperature---which maintains constant learning difficulty throughout training. This rigidity fails to accommodate the student's evolving capability, especially when the teacher's performance plateaus. To address this, we reframe SSAT as knowledge distillation from a Bayes teacher and formalize a stage-dependent bias-variance tradeoff. This analysis reveals the existence of an optimal temperature: low temperatures reduce bias, while higher temperatures control the variance term in the robust generalization bound. Guided by this insight, we propose BayesAT, a Bayes-guided progressive distillation framework that jointly employs a low-to-high temperature warming schedule and confidence-adaptive sample reweighting. BayesAT allows the student to first learn from sharp, high-confidence supervision to establish reliable decision boundaries, then progressively from smoother distributions that encourage exploration of inter-class relationships. Extensive experiments on CIFAR-10, CIFAR-100, and ImageNet-200 demonstrate that BayesAT consistently improves robust accuracy while maintaining natural accuracy over existing SSAT methods, highlighting the importance of dynamic, theory-driven supervision for effective knowledge transfer in semi-supervised adversarial training.
PaperID: 2849, Poster
Abstract: Causal sensitivity analysis bounds the policy value of off-policy evaluation under unobserved confounding. In offline reinforcement learning, however, sensitivity bounds are typically applied not to raw states but to learned representations. We show that this two-stage pipeline can silently break the coverage guarantee: state aggregation under hidden confounding amplifies the effective sensitivity parameter in the latent space, so bounds computed at the nominal confounding level may exclude the true policy value. We characterise the amplification mechanism, give a sufficient condition for preservation, and prove a continuity result showing the safety condition is not an isolated point: amplification grows continuously with the conditional-independence violation. Neither unsupervised nor standard task-aware representation objectives target the right quantity. We propose Sensitivity-Preserving Representation Learning (SPRL), which augments latent dynamics prediction with a kernel-weighted within-cell propensity homogeneity penalty, an observable surrogate for the unobservable safety condition, and show across synthetic and semi-synthetic benchmarks that SPRL is the only method tested that simultaneously preserves coverage in regimes where unsupervised representations break it and prevents amplification in regimes where task-aware baselines fail.
PaperID: 2850, Poster
Abstract: Reference-to-video (R2V) generation has advanced significantly and is now a promising paradigm for controllable video synthesis. Nevertheless, in advertisement video generation, where accurate camera control and precise product presentation are essential, existing approaches still face two fundamental challenges. First, current models adhere weakly to high-level semantic camera instructions, yielding stochastic or inconsistent camera motion. Second, they struggle to preserve strict cross-frame appearance consistency, often leading to identity ambiguity artifacts such as logo drifting, texture distortion, and geometric distortion. Although certain image animation methods incorporate explicit motion conditions (e.g., trajectories or motion fields) to improve controllability, these conditions are costly to annotate, thereby constraining their practical applicability. To address these limitations, we propose a camera-aware R2V framework named AdDirector, which adaptively exploits anchored guidance from the reference image and propagates it across spatial, temporal, and scale dimensions. Specifically, AdDirector introduces Spatio-Temporal-Scale (STS) modulation, a camera-conditioned modulation mechanism that first selects multi-scale reference features, and then regulates feature injection throughout the R2V generative process spatio-temporally conditioned on camera types. This design allows our model to adaptively balance global structural coherence and fine-grained detail preservation. Furthermore, AdDirector proposes Temporal Anchoring Regularization (TAR), a boundary-conditioned feature-level regularization scheme. TAR leverages latent trajectories extracted from a base R2V model as temporal anchors, thereby reducing feature drift and improving cross-frame appearance consistency. These two modules together transform static reference conditioning into structured anchored guidance, and experiments show that our method substantially improves camera controllability, appearance consistency, and spatio-temporal coherence for advertisement video generation.
Abstract: In cooperative teams where agents act in a fixed order and share a single team-level reward (multi-agent language systems, sequential robotic tasks), per-agent credit assignment is under-determined. Critic-based approaches scale poorly as the number of agents grows owing to the costly maintenance of joint/factored critic(s), whereas the existing critic-free alternatives have other issues: common credit across agents that couples every agent's signal to teammate noise, importance-sampling corrections for upstream-update staleness that incur variance exponential in team size, or per-agent counterfactual replay that isolates each agent's effect at the price of extra environment or reward calls. We propose COSAC, a critic-free per-agent policy gradient for sequential cooperative teams. COSAC fits an additive per-agent decomposition of the team reward by a single ridge regression on the rollout batch (giving each agent a learning signal decoupled from teammate noise), and computes each agent's counterfactual advantage from fictitious continuations of the current policy (policy forward passes that replace both importance-sampling reweighting and per-agent environment replay, at no extra environment or reward cost). The estimator instantiates the Sequential Aristocrat Utility (SeqAU), our extension of Wolpert and Tumer's (2001) aristocrat utility to sequential teams. We prove bias and variance bounds on SeqAU credits that stay controlled as the team grows. Our controlled study on sequential bandits demonstrates that COSAC attains the lowest advantage MSE and consistently low learning regret across team sizes up to K = 16. On the AI2 Reasoning Challenge (ARC) task, where four Qwen3-0.6B agents reason in turn about a grade-school science question, COSAC attains faster convergence than the other critic-free baselines.
PaperID: 2852, Poster
Abstract: Gradient clipping is a central but brittle design choice in differentially private federated learning: overly small thresholds bias client updates, whereas overly large thresholds amplify the noise required by privacy. This trade-off is further complicated by client heterogeneity and personalized privacy budgets, where a single global threshold is often mismatched and online adaptation from private training statistics can complicate privacy accounting. We study clipping-threshold selection as a privacy-budget-conditioned policy learning problem. We propose PAC-DP, a proxy-learned clipping framework for record-level locally private federated learning. Before private training, PAC-DP calibrates a deterministic policy \pi_\theta(\varepsilon,t) from public or synthetic proxy simulations, mapping a client's target privacy budget and the training round to a clipping threshold. The policy is then frozen and used during private training, so deployed thresholds depend only on the declared budget, round index, and public/proxy-learned parameters, not on private gradients, losses, or client-specific update histories. PAC-DP combines this policy with per-example clipping, Gaussian perturbation, and per-client RDP accounting. We show that, under public or separately privatized proxy calibration, the frozen policy introduces no additional record-level privacy loss beyond the clipped Gaussian mechanisms. We further analyze the clipping-sensitive utility trade-off through a decomposition involving clipping bias, stochastic variance, client heterogeneity, and DP noise, and characterize proxy-to-target transfer via a regret bound. Experiments on MNIST, CIFAR-10, CIFAR-100, and Heart Disease show that PAC-DP improves the privacy--utility trade-off and communication efficiency over implemented fixed-threshold and adaptive DP-FL baselines under matched privacy budgets.
Abstract: Computer-use agents hold the promise of assisting in a wide range of digital economic activities. However, current research has largely focused on short-horizon tasks over a limited set of software. A key reason is that creating environments for complex software requires significant time and human effort, and therefore does not scale. To address this, we introduce Gym-Anything, a framework for converting any software into an interactive computer-use environment. We frame environment creation itself as a multi-agent task: a coding agent writes setup scripts, downloads real-world data, and configures the software, while producing (visual) evidence of correct setup. An independent audit agent then verifies evidence for environment setup against a quality checklist. Using a taxonomy of economically valuable occupations grounded in U.S.\ GDP data, we apply this pipeline to 200 software applications with broad occupational coverage. The result is CUA-World, a collection of over 10K+ long-horizon tasks spanning domains from medical science and astronomy to engineering and enterprise systems, each configured with realistic data along with train and test splits. Distilling successful trajectories from the training split yields strong gains at multiple student scales (2B, 9B, 35B). CUA-World also includes CUA-World-Long, a challenging long-horizon benchmark with tasks often requiring over 500 steps. We also apply the same auditing principle at test time: a separate VLM reviews completed trajectories and provides feedback on what remains, improving Gemini-3-Flash on CUA-World-Long from 11.5% to 14.0%. We will release all code, infrastructure, and benchmark data.
PaperID: 2854, Poster
Abstract: Topology supervision relies on meaningful structural targets, yet this assumption can fail under structurally unreliable pseudo supervision. In pseudo-label learning, the topology induced by pseudo masks may already contain distorted components, broken connections, or spurious holes, making direct supervision-side topology fitting unreliable and potentially propagating structural errors. We address this problem by reformulating topology supervision as trainable topology-side correction. Specifically, we propose a differentiable persistent-homology proxy loss that derives topology signals from a soft Euler-characteristic trajectory and threshold-wise structural variation, without computing or matching persistence barcodes. We further introduce a conservative topology-consistency injection mechanism that regulates when, how strongly, and in which direction these signals enter optimization. We validate the proposed reformulation in unsupervised camouflaged object detection, a challenging setting where pseudo masks often contain severe structural distortions. Experiments under controlled same-carrier comparisons show consistent improvements in both task-level detection performance and topology-level structural reliability. These results suggest that, under structurally unreliable pseudo supervision, topology supervision is more effective when derived as trainable proxy signals and introduced into optimization conservatively. Code is available at https://anonymous.4open.science/r/LPHP/.
PaperID: 2855, Poster
Abstract: It is now common to use 2D diffusion priors for problems where the underlying object is 3D (e.g., medical imaging). We would pick a noise schedule from the EDM literature, anneal a tilted prior, and then use the formulation as a 3D generative model. It works empirically. But we know very little about what it converges to and how fast. We study this question for a natural target: the multi-marginal Schrödinger bridge, the relative-entropy projection of a 3D prior onto the distribution whose rendered slices match the 2D marginals. But computing it is intractable. The annealed 2D-tilted surrogate is what one actually uses in applications. The main question is how close the latter object is to the former. Our main idea is that the discrepancy between these two laws — a comparison of two high-dimensional, fully coupled distributions — can be reduced to the size of a single scalar correction field. This field has a nice two-piece form: residual coupling across slices and tempering mismatch from the annealed discrete schedule. Using a Gaussian-smoothed filtration, Wasserstein screening, and conditional concentration, we obtain an explicit discrepancy bound in which the relevant terms separate into latent and posterior-residual contributions. When the smoothing scale is identified with the EDM noise level \sigma(\tau), our bound yields \mathrmKL(\Pi_\tau\|P^\star)=O\left(N\sigma(\tau)^2\right) and \|\Pi_\tau-P^\star\|_\mathrmTV=O\left(\sqrtN\sigma(\tau)\right), if the latent scale remains bounded and the surrogate log-discrepancy is noise-aligned. So, the 2D-tilted construction converges to the Schrödinger-bridge lift at a rate governed by user-checkable quantities like the surrogate Lipschitz constants and the posterior concentration constant.
PaperID: 2856, Poster
Abstract: Contact-rich manipulation requires precise control under complex contact dynamics while remaining robust to diverse visual conditions. However, real-world reinforcement learning often overfits to the visual appearance of training scenes. We propose Semantics-to-Contact (S2C), a stagewise framework for visually robust contact-rich manipulation. The insight is to decompose the problem into two stages: semantic focusing and contact refinement. In the first stage, a vision-language model localizes task-relevant regions across visually diverse scenes, providing coarse spatial guidance and reducing the burden of exploration. In the second stage, a real-world human-in-the-loop residual reinforcement learning policy learns fine-grained contact behaviors within the localized region using force feedback. To further improve visual robustness, we introduce an object-centric visual augmentation strategy that randomizes background appearance while preserving the robot and task-relevant objects. Experiments on six real-world contact-rich manipulation tasks exhibit that S2C consistently improves success rates over strong baselines and maintains reliable performance under significant visual and positional variations. These results demonstrate that stagewise semantic focusing and contact refinement provide a practical path toward visually robust real-world contact-rich manipulation. Video materials can be seen in \urlhttps://anonymous.4open.science/api/repo/S2C-demo-4E03/file/index.html.
PaperID: 2857, Poster
Abstract: Generalized category discovery (GCD) becomes substantially more challenging when the number of novel classes is unknown. We argue that, after the backbone is fixed, this stage should be treated as seen-constrained model-order selection rather than as an auxiliary clustering detail. The selector must estimate the class count without novel labels while remaining consistent with the labeled seen categories available in the standard GCD setting. We instantiate this view by constructing a nested merge path from a seen-constrained overclustering and selecting the coarsest near-optimal cut under observable surrogate decision risks. These risks combine seen-class consistency, cross-seed stability, unlabeled-side compactness, and a tiny-cluster degeneracy penalty. The resulting protocol separates class-count error, downstream clustering quality, and robustness. Where an aligned public unknown-K baseline is available, namely on SelEx, the selector reduces class-count error on all four fine-grained datasets, including a reduction from 342 to 15 classes on Herbarium-19. Broader-domain diagnostics further distinguish ImageNet100, where the default selector can underestimate the class count toward the seen-class regime, from CIFAR100, where the selector landscape exposes an accuracy-count operating-point trade-off. The mixed accuracy outcomes reveal distinct count-, path-, and representation-level bottlenecks, supporting unknown-K selection as a decision layer that should be evaluated separately from representation learning.
PaperID: 2858, Poster
Abstract: TrueSkill is a widely used Bayesian rating system for inferring latent player skill from game outcomes. However, its standard formulation assumes a fixed global performance-noise parameter, making it sensitive to atypical outcomes and unable to distinguish skill from player-specific consistency. We propose two heteroscedastic extensions: a match-specific precision model that yields a robust heavy-tailed comparison likelihood, and a player-specific precision model that captures differences in performance consistency among players. Both extensions use Gamma-distributed precision variables and retain a modular factor-graph representation. Because the introduced Gaussian-Gamma factors do not admit the same closed-form expectation-propagation updates as vanilla TrueSkill, we derive an efficient hybrid message-passing scheme that combines expectation-propagation updates for outcome truncation factors with variational updates for latent precision factors. Experiments on synthetic and real match data show that the proposed models improve robustness to anomalous outcomes and provide interpretable estimates of player consistency.
PaperID: 2859, Poster
Abstract: We introduce TrackTok, an efficiency-focused video tokenization framework that produces object-centric tokens, i.e., tokens that consistently encode the same semantic objects across time, even as it moves, deforms, or undergoes partial occlusion. Conventional patch-based or existing holistic tokenizers produce token identities often tied to fixed image locations. In contrast, our approach learns temporally persistent object-centric representations that bind visual features across frames into stable semantic units. This enables substantially more compact, continuous latent video representations, reducing the number of tokens required for downstream generative modeling with state-of-the-art flow models, while preserving the semantic structure needed for high-quality synthesis. On the UCF-101 video generation benchmark, we provide a proof of principle that our tokenizer uses ~4x fewer latent tokens to achieve the same generative FVD compared to existing holistic tokenizers. This results in an up to 10x increase in sample throughput, offering a favorable quality-efficiency trade-off. These results suggest that object-aligned and temporally persistent tokenization is a promising direction for scalable video generation, where reducing latent token counts is critical for efficient modeling.
PaperID: 2860, Poster
Abstract: In LLM-based recommendation, memory effectively plays the role of a user-level agent, typically built on top of a memory-augmented backbone. Yet under this paradigm users and items remain asymmetric participants: only the user is given a structured profile; that profile is consumed inside a bandwidth-limited reranker prompt; and it is frozen after warmup. Starting from this observation, we recast LLM-based recommendation as an agentic-communication process in four stages --- Build a self-description, Transmit it along the collaborative graph, Understand the other side, and Sync it from the next interaction. Analysing the dominant memory-augmented paradigm through this formulation, we identify three open communication challenges that current pipelines leave unaddressed: the item side never produces a comparable self-description (one-sided Build); profile-level signals are routed only into a bandwidth-limited reranker prompt rather than into the retrieval stage where they would not compete for context tokens (prompt-trapped Understand); and the user's description is frozen after warmup with no principled update rule (absent Sync). We instantiate the missing stages as HARP (Homophily-aware Agentic Recommendation via dual Profiles), a training-free, three-operator addition to a frozen LLM backbone: a rule-extracted symmetric item profile (zero extra LLM calls), a channel-decomposable profile-level homophily \mathcalH(\phi_u,\phi_i) consumed at retrieval time rather than inside the reranker prompt, and a critic-gated single-pass reflective update realising in-context profile evolution without any gradient step. On four public benchmarks (Yelp, Amazon Books, MovieTV, Goodreads), HARP lifts Hit@1 over the strongest published baseline by 6.9%--24.0%, with family-wise Holm-corrected p<10^-9 on every dataset. Ablations and a per-case attribution study isolate each operator as a distinct, complementary source of gain; we also quantify and openly report the trade-off introduced by the reflective pass. Our contribution is a unified formulation of agentic communication for LLM-based recommendation, together with a training-free three-operator instantiation on top of it.
Authors:
Van Chien Nguyen, Chaitra Hegde, Van-Cuong Pham, Ryan Rossi, Franck Dernoncourt, Thien H NguyenAbstract: We introduce Orthrus, a simple and efficient dual-architecture framework that unifies the exact generation fidelity of autoregressive Large Language Models (LLMs) with the high-speed parallel token generation of diffusion models. The sequential nature of standard autoregressive decoding represents a fundamental bottleneck for high-throughput inference. While diffusion language models attempt to break this barrier via parallel generation, they suffer from significant performance degradation, high training costs, and a lack of rigorous convergence guarantees. Orthrus resolves this dichotomy natively. Designed to seamlessly integrate into existing Transformers, the framework augments a frozen LLM with a lightweight, trainable module to create a parallel diffusion view alongside the standard autoregressive view. In this unified system, both views attend to the exact same high-fidelity Key-Value (KV) cache; the autoregressive head executes context pre-filling to construct accurate KV representations, while the diffusion head executes parallel generation. By employing an exact consensus mechanism between the two views, Orthrus guarantees lossless inference, delivering up to a 7.8× speedup with only an O(1) memory cache overhead and minimal parameter additions.
PaperID: 2862, Poster
Authors: MD Azizul Hakim
Abstract: Training quality is the primary determinant of reasoning efficiency in large language models, yielding greater returns than proportional increases in inference computation. Evaluating 44 language models (0.5B--685B parameters) across six reasoning benchmarks totalling over 67000 assessments, this study establishes when additional inference tokens are beneficial, redundant or actively harmful. Standard models exhibit near-universal accuracy convergence at approximately 1000 tokens regardless of architecture or parameter count. Causal intervention experiments demonstrate that constraining generation length improves reasoning accuracy by 15.4 percentage points, identifying over-generation rather than capacity exhaustion as the primary efficiency bottleneck. Among 13 reasoning-specialised models, training methodology yields 18.8-fold efficiency improvement at matched parameter scales, exceeding gains from ten-fold parameter increases. Graduate-level evaluation reveals that arithmetic benchmark performance does not predict reasoning capability on complex tasks (r = -0.03, P = 0.91), exposing fundamental limitations invisible to standard evaluation practice. Mechanistic profiling reveals two orthogonal dimensions of training quality---quality control and decomposition efficiency---that independently predict efficiency profiles, providing a measurable framework linking training decisions to inference-time behaviour.
Authors:
Kuiye Ding, Yifan Hu, Hanchen Wang, Hao XueAbstract: Multivariate time series forecasting presents unique challenges because future variables often co-evolve under shared system dynamics. While existing studies mainly focus on cross-variable dependencies in historical observations, dependencies among future values are much less explored. Specifically, modern forecasting models largely follow the Direct Forecasting (DF) paradigm, generating multi-step forecasts with point-wise objectives that do not explicitly constrain cross-variable structure. In this work, we show that the DF objective is mismatched in the presence of cross-variable and lagged dependencies, revealing an objective gap. To address this issue, we propose (CvLoss), a plug-in structural regularizer that constrains forecast residuals on a cross-variable graph. CvLoss penalizes inconsistent edge-wise residual differences over forecast patches, encouraging consistency across both synchronous and asynchronous interactions. Our experiments show that CvLoss consistently improves competitive forecasting models, outperforms representative learning objectives, and is compatible with a variety of forecasting backbones. Code is available at: https://anonymous.4open.science/r/CvLoss.
Abstract: LLM-based search agents are widely used for information-seeking tasks, but their reliance on external tool returns introduces a critical security risk: web content retrieved during execution is untrusted, exposing agents to prompt injection and goal hijacking. Prior work on search-agent safety primarily focuses on static web-content injection, but modern agents issue follow-up queries and cross-check competing sources, so a single injected page is often diluted or rejected. We show that the channel delivering search and page observations is a fragile security boundary: beyond exposing the agent to a single poisoned page, a mediated search interface can repeatedly steer how the agent gathers evidence and forms its final answer. Under a constrained tool-intermediary threat model, appending only one controlled result per query can substantially increase attack success when the evidence is coordinated across the agent's trajectory. We study this setting with a strategy-driven long-horizon attack system and introduce Authority-Chain Hijack (ACH), an expert-refined strategy that turns isolated search-result and page-content manipulations into a coherent evidence chain across seemingly corroborating sources. ACH achieves the highest Overall ASR among all baselines, reaching 56.0%/85.0% ASR/MaxN\,ASR on the SafeSearch benchmark. We further introduce Trace-Guided Strategy Evolution (TGSE), which automatically improves reusable attacker strategies from execution traces, replacing manual redesign with trace-driven refinement and further raising these to 71.4%/95.0%.
PaperID: 2865, Poster
Abstract: We study risk control for online policy learning, where an agent adaptively collects data and aims to learn a new policy where the risk is below a tolerance threshold at every round with a high probability. Such policy updates and data collection induce feedback covariate shift (FCS): the observed data depend on the history through the evolving policies and data-collection process. FCS brings two substantive problems for risk control: 1) efficient risk evaluation for each candidate policy and 2) valid policy selection with high probability risk control. To address these, we propose an efficient permutation-based weighted risk evaluator based on the Sinkhorn algorithm and select policies whose estimated risk upper confidence bounds (UCB) fall below the target threshold, with the UCB obtained through a proposed Gaussian bootstrap method. Theoretically, we prove that our risk evaluator's worst-case variance is no greater than that of the standard importance-weighted evaluator. Also, our policy control method yields an asymptotically valid simultaneous risk bound and provides asymptotic high-probability risk control for the selected policy. Empirically, our method maintains the target risk level, achieves lower risk than baseline methods, has lower variance in risk estimation, and enables reliable policy selection with high-probability risk control.
PaperID: 2866, Poster
Authors:
Gerard L Donahue, Guven Gergerli, Ayush Gupta, Reza Ghoddoosian, Enna Sachdeva, Faizan SiddiquiAbstract: Temporal action segmentation (TAS) is commonly trained on large banks of pre-extracted frame features rather than raw RGB video, but storing and replaying those features becomes a practical bottleneck for long-form video, especially in incremental TAS where past-task data must remain available to mitigate catastrophic forgetting. Recent TAS replay and condensation methods reduce this burden, but they rely on dense frame-wise annotations and allocate storage through fixed segment-level structures, limiting their flexibility for long procedural videos whose frame complexity varies substantially over time. We introduce Adaptive Residual-Quantized Variational AutoEncoder (ARQ-VAE), a compact and label-free memory representation for TAS features that learns a discrete residual quantization of frame-level video features and then adaptively truncates the residual chain on a per-frame basis, retaining only the residual prefix that best reconstructs each frame feature. This yields a practical compact feature memory that does not depend on frame-wise labels and can therefore support multiple TAS settings with the same stored representation, while a lightweight fine-tuning stage further improves reconstruction quality for the truncated codes that are actually retained. Across standard TAS benchmarks, ARQ-VAE achieves a substantially improved storage--performance trade-off over prior replay and condensation baselines. On Breakfast, our final representation reduces training-set storage from 28 GB for the original I3D feature bank to 10 MB while retaining strong downstream segmentation performance; with ASFormer, it achieves 71.5 Edit and 51.2 F1@50 in the supervised setting. We further show that the same label-free memory remains effective for incremental and unsupervised TAS, and transfers to an alternative pretrained feature space based on DINOv3. We position ARQ-VAE as a practical memory representation for TAS, and a promising direction for broader dataset condensation applications.
PaperID: 2867, Poster
Abstract: Text-aware image super-resolution (TAISR) aims to recover high-resolution images while preserving textual content. Yet some degraded text cannot be restored from local visual evidence alone: humans often read it by reasoning over image context, logical patterns, and prior knowledge. This raises a fundamental question: does TAISR need reasoning? We answer with two new diagnostic tools: ReasonText, a human-curated benchmark with perception-grounded difficulty levels that separate reasoning-required text from locally recoverable text, and GTTCA, a crop-based recognition accuracy metric that isolates restoration quality from detection errors. Evaluating twelve recent SR models, we find a consistent and substantial performance drop on reasoning-required text, revealing a systematic blind spot in current TAISR designs. To address this limitation, we build on two observations: (1) SR models benefit from ground-truth text labels provided as captions, and (2) modern MLLMs can infer degraded text from LR images through reasoning. These observations motivate Reasoning Transfer via Captioning (RTC), a training-free plug-in that uses an external MLLM as a reasoning module and injects its inferred text into SR models through natural-language captions. Surprisingly, with RTC, real-world SR models can even outperform dedicated text-aware SR models, suggesting that reasoning should be treated as a core design axis for future TAISR.
PaperID: 2868, Poster
Authors: Shuai Feng, Hao Yan
Abstract: Continuous DAG learners typically optimize a single weighted adjacency matrix, forcing each entry to represent both direct-edge support and signed effect magnitude. As a result, data fitting, sparsity, acyclicity, and any auxiliary structural information are all imposed on the same variables, even when the available information concerns higher-order relations such as reachability or ancestral order. We propose a structure-separated continuous DAG-learning framework that gives graph support its own soft representation. The effective adjacency is ( A=(W f(M)=\beta\tanh(M), ) where (W_f(M)) is a bounded signed-magnitude branch and (G(Z)) is a structural gate. Data fitting, sparsity, and the DAGMA barrier are evaluated on (A), while cycle and relation constraints are applied to closure scores computed from (G). The bounded magnitude branch prevents strong effective edges from being hidden behind vanishing structural gates, making closure-side structural regularization meaningful. On this separated structural gate, we introduce SMNC, a scan-efficient multi-scale normalized closure module. SMNC computes dyadic closure scores over (G) and normalizes each scale by its path-length budget, producing scale-normalized matrix-valued logits for self-reachability and off-diagonal reachability. This design addresses a key limitation of unnormalized product-style closure scores, whose path evidence scales as (p^L) on a length-(L) path with gate value (p), causing long-cycle signals to decay exponentially. At the appropriate covering scale, SMNC routes gradients through an active length-(L) cycle with a polynomial (1/L)-type scale factor; in the all-scale soft readout used in practice, this should be interpreted as a local active-scale effect. The same relation space supports diagonal self-reachability losses for cycle suppression and masked off-diagonal losses for optional ancestry, backward-exclusion, and incomparability cues, without treating positive reachability labels as direct-edge labels. Across linear and nonlinear SEM benchmarks, with ablations over factorization and closure normalization, we show that structure-separated SMNC improves DAG recovery and provides a principled closure-space interface for higher-order structural supervision.
PaperID: 2869, Poster
Authors:
Jiawei Feng, Chuzhao Huang, Xinli Xu, Yingcong ChenAbstract: Vision-Language-Action (VLA) models are increasingly used as general-purpose policies for robotic manipulation, but completing the instructed task requires more than reaching the target state. A policy must also avoid nearby objects throughout the rollout, including approach, interaction, transport, and handover. A common approach is to invoke conventional collision avoidance mechanisms, such as motion planners or safety filters, which provide effective geometric safeguards. However, the resulting system may struggle to produce actions that are simultaneously safe and task progressing. When collision avoidance and manipulation progress require misaligned actions, the system may avoid the obstacle while slowing, disrupting, or even failing the task. To address this, we introduce DODGE-IT, a collision-aware learning framework that leverages collision signals as policy-improvement feedback to strengthen the VLA policy's collision-aware manipulation capability. DODGE-IT first designs Task-Relevant Assessment of Collision Events (TRACE) that detects and identifies task-relevant obstacles from RGB-D observations and turns executions into collision outcomes. These outcomes are then incorporated together with task success signals into representative reinforcement learning frameworks for VLA. We validate DODGE-IT on SafeLIBERO and real robot manipulation tasks, showing improved obstacle avoidance while preserving task completion across simulated and physical settings.
Authors: fei ding, Yongkang Zhang, youwei wang, Zijian Zeng
Abstract: Reinforcement learning for multi-step reasoning with large language models (LLMs) typically relies on sparse terminal rewards, which creates a poorly conditioned credit-assignment problem: the final feedback is propagated uniformly across all intermediate decisions. This leads to high gradient variance, unstable training, and many ineffective updates, ultimately limiting sustained model improvement. We propose a counterfactual-comparison framework for credit assignment. For each input, the framework samples multiple reasoning trajectories and treats their differences as implicit approximations to alternative decisions. This yields an implicit process-level advantage estimator that converts sparse terminal rewards into step-sensitive learning signals. Building on this framework, we introduce Implicit Behavior Policy Optimization (IBPO), which substantially improves training stability and the performance ceiling on mathematical and code-reasoning benchmarks. Our results point to a promising direction for unlocking the reasoning potential of LLMs.
PaperID: 2871, Poster
Abstract: Latent space Bayesian optimization circumvents the curse of dimensionality over structured inputs by performing Gaussian process surrogate-guided search in a continuous latent space learned by deep generative models. However, existing latent space Bayesian optimization methods overlook both the inherent stochasticity of latent space representations and the geometric properties of the manifold induced by VAE, leading to suboptimal optimization performance. In this work, we propose a marginalized kernel that integrates out the latent variables through the VAE inference posterior, yielding a closed-form kernel on the observation space that naturally accounts for encoding uncertainty. We adopt a quadratic polynomial kernel as the base kernel in the latent space, which captures discrepancies in the first two moments of the latent posteriors while avoiding the bandwidth sensitivity inherent to the RBF kernel. Furthermore, inspired by the Riemannian interpretation of VAEs, we incorporate a geometry-aware sampling scheme that leverages the metric structure revealed by the learned posterior covariances to guide candidate acquisition toward high-information regions of the latent space. Empirical evaluations on molecular design and robot design tasks demonstrate that our method outperforms existing state-of-the-art baselines across the majority of tasks.
Abstract: Low-Rank Adaptation (LoRA) enables efficient adaptation of large pre-trained models to downstream tasks by parameterizing weight updates with low-rank matrices. In this paper, we investigate the limitations of the LoRA parameterization from a geometric perspective. Specifically, we show that when a full fine-tuning gradient is backpropagated to the low-rank matrices, it undergoes anisotropic scaling driven by their singular values. We argue that this phenomenon is undesirable because it distorts the full fine-tuning gradient by skewing it toward dominant singular directions while suppressing others. Our analyses demonstrate that anisotropic gradient scaling reduces the effective rank of the low-rank matrices' gradients and results in suboptimal alignment between the full fine-tuning gradient and its low-rank approximation in LoRA, thereby exacerbating the gap to full fine-tuning. To address these limitations, we propose a new low-rank parameterization, SDS-LoRA, which structurally decouples singular values from the backward pass. Our method ensures that the full fine-tuning gradient backpropagates only through the orthonormal bases of the low-rank matrices' subspaces, independent of their scales. Convergence analysis demonstrates that while LoRA’s convergence rate degrades with the condition number of the low-rank matrices, SDS-LoRA remains independent of it. Experimental results across natural language and vision benchmarks show that SDS-LoRA improves loss convergence and reduces the gap to full fine-tuning, significantly enhancing adaptation performance.
PaperID: 2873, Poster
Abstract: Multi-modal large language models (MLLMs) achieve strong vision–language reasoning ability but remain vulnerable to jailbreak attacks that exploit subtle cross-modal cues to bypass safety mechanisms. Detecting such attacks is difficult because harmful data are scarce and rapidly evolving, while detectors trained on known patterns often fail to generalize. In this paper, we cast multi-modal jailbreak detection as self-supervised deviation modeling and learn attack-agnostic signals from benign data only. We propose SelfGuard, a self-supervised framework, which models benign multi-modal regularities and learns to score structured violations via controlled pseudo-harmful deviations. SelfGuard synthesizes pseudo-harmful samples via controlled transformations that target three signature cues, cross-modal semantic inconsistency, format manipulation, and semantic toxicity. It then learns complementary deviation statistics through multi-task self-supervised learning, including misalignment discrimination, reconstruction-based toxicity modeling, and contrastive format modeling. At inference, SelfGuard performs multi-view deviation estimation by aggregating task-specific deviation scores to identify departures from benign multi-modal regularities. Extensive experiments on various benchmarks show the effectiveness of our method.
PaperID: 2874, Poster
Authors: Nabyl Quignon, Antitza Dantcheva
Abstract: Image steering aims to provide fine-grained control over concepts in generated images, enabling targeted and continuous manipulation without broadly altering the generation process. Recent steering methods for diffusion models primarily focus on the text-conditioning space or rely on costly auxiliary training. Deviating from that, we here propose a novel steering approach that extracts visual concept directions directly from the image-side activations of Diffusion Transformers. Specifically, we showcase few-shot concept extraction in the activation space using a simple difference-of-means estimator over intermediate residuals. We observe (a) scene-level concepts such as lighting, weather, and style are distributed across image tokens, as well as (b) local concepts such as object colors, textures, materials, and facial attributes are concentrated in specific regions. Towards controlling both concepts, we introduce a new unified training-free framework that uses global difference-of-means for (a) distributed concepts and masked difference-of-means for (b) local concepts, preventing local signals from being diluted by global averaging. Notably, using only a handful of positive and negative prompt pairs, our method identifies and extracts reusable concept directions that can be injected at inference for continuous steering. Quantitatively, our method outperforms text-encoder steering and training-free activation-editing baselines, as well as is competitive with specialized editing models. Our results suggest that as concept directions can be applied at patch level, our method enables precise targeted and compositional steering, manipulating multiple concepts in the same image.
PaperID: 2875, Poster
Abstract: Hand radiographs support diverse clinical tasks, including skeletal maturity assessment, rheumatoid arthritis scoring, abnormality screening, and anatomical localization, yet existing models are typically trained separately for each task and dataset. We present HandXFM, a domain-specific foundation model that leverages cross-modality knowledge for comprehensive hand X-ray understanding. HandXFM distills semantic priors from BiomedCLIP and structural cues from a chest X-ray Vision Transformer (ViT) through a unified alignment framework, integrating complementary medical knowledge in an interpretable way. After pretraining, HandXFM can be adapted to a range of musculoskeletal tasks, including bone age estimation, SvH score prediction, abnormality detection, joint localization, bone segmentation, and visual question answering. It consistently outperforms task-specific and single-expert baselines, achieving up to 0.11 improvement in Pearson correlation coefficient (PCC) for bone age prediction and a 12% accuracy gain in abnormality classification. Moreover, HandXFM generalizes to unseen datasets and yields cross-attention maps linking clinical terms to anatomical regions. These results demonstrate that multi-expert distillation effectively unifies semantic and structural supervision within a single pretrained model, establishing HandXFM as an interpretable and generalizable foundation model for hand radiograph analysis.
PaperID: 2876, Poster
Authors: Lichen Ma, Zipeng Guo, Yu He, Xiaolong Fu, Luohang Liu, Jingling Fu, Junshi Huang, Yan Li
Abstract: To circumvent the inherent fidelity bottlenecks and optimization misalignment of VAE-based latent diffusion, pixel-space diffusion models have emerged as a compelling end-to-end paradigm. However, existing pixel diffusion models often struggle to balance computational efficiency with the preservation of high-frequency details. They frequently resort to patch-based compression or restricted local decoding, leading to a "spectral compromise" where high-frequency and fine-grained pixel information are suppressed. To address these challenges, we propose FrequencyBooster, a novel framework designed to empower pixel diffusion with full-frequency modeling capabilities without prohibitive overhead. The core of our method is a high-capacity decoder that specializes in extracting exhaustive high-frequency details and low-frequency semantics, the latter of which is derived from a Diffusion Transformer (DiT) backbone. Unlike prior works that sacrifice global context for local refinement, FrequencyBooster leverages high-dimensional feature representations to maintain global structural integrity while achieving superior pixel-level precision. Extensive experiments on ImageNet demonstrate the effectiveness of our approach: our model achieves a state-of-the-art FID of 1.60 at 256 × 256 resolution within only 320 epochs. Furthermore, at 512 × 512 resolution, FrequencyBooster attains an FID of 1.69, significantly outperforming existing pixel-space and latent-space generative models.
PaperID: 2877, Poster
Abstract: Link prediction (LP) is a widely applied task in graph learning. Conventional LP approaches typically follow a dataset-specific paradigm, requiring independent training for each graph and incurring high computational and maintenance costs in large-scale applications. Motivated by these limitations, recent work explores Universal Link Prediction (ULP), aiming to enable training-free inference on arbitrary unseen graphs. However, existing methods primarily rely on subgraph sampling strategies to construct universal representations, which are unable to capture higher-order information due to the exponential growth of higher-order neighborhoods. As a result, they often under-perform on graphs where long-range signals are essential, such as biological networks. Moreover, their prediction models are typically domain-agnostic and lack the ability to adapt to different graph domains, leading to limited generalization. To address these challenges, we propose an Atlas-based Universal Link Prediction framework (AtlasULP), which introduces the relation atlas to capture structural associations between node pairs from both global and local perspectives, enabling a comprehensive characterization of connectivity signals across diverse graph domains. Building upon this representation, we develop a domain-guided prompting ULP model, which generates domain-aware prompt tokens from contextual structures and performs prompt-guided in-context prediction for adaptive link inference. Extensive experiments demonstrate that AtlasULP consistently outperforms state-of-the-art methods across diverse graph domains.
PaperID: 2878, Poster
Abstract: Low-rank adaptation (LoRA) has become the standard method for parameter-efficient fine-tuning of large pretrained models, yet theoretical explanations for why low-rank updates suffice remain incomplete. Most existing analyses rely on the Neural Tangent Kernel (NTK) regime, which linearizes the network around the pretrained weights. In this work, we take a different approach that requires no linearization: our key observation is that, once the backbone is frozen, the loss sees each adapted weight matrix only through a finite-rank linear map induced by the frozen activations. Trace-norm regularization then forces the optimizer to pick a low-rank matrix among all matrices that produce the same observed output. This yields global minimizers whose ranks are controlled by the rank of this map, which we call the \emphbottleneck rank, rather than the sample size. For multi-head attention, we obtain tighter rank upper bounds by exploiting two additional invariances: row-wise softmax is insensitive to additive row constants in the score matrix, and value updates are observable only after attention weighting and frozen output projection. For the factorized LoRA objective with weight decay, we prove a complementary result on the nonconvex side: after an arbitrarily small generic positive semidefinite perturbation, every first-order stationary point is rank-deficient once the LoRA rank exceeds a threshold determined by the bottleneck rank. Finally, we show that global minimizers of the perturbed objective are near-optimal for the original trace-norm problem, and we provide uniform generalization bounds governed by the spectral structure of the frozen features.
PaperID: 2879, Poster
Abstract: Large language model (LLM) agents typically perform function calling by retrieving relevant tools and reinserting their specifications as textual prompts at inference time. Although retrieval reduces the number of candidate tools, selected tool specifications must still be repeatedly re-encoded, incurring substantial latency and memory overhead, especially in on-device environments. We argue that this inefficiency stems from treating tool knowledge as transient text rather than reusable execution memory. To address this, we propose EngramState, a state-centric framework that compiles tool specifications offline into reusable recurrent state priors and directly loads them at runtime for query-only inference. EngramState combines retention-aware state construction, channel-wise state compression, a lightweight same-backbone state retriever, and feedback-driven state routing. On the DroidCall benchmark with an RWKV-7 1.5B backbone, EngramState reduces prompt tokens from 985 to 22 (97.8%) under the Top-4 retrieval setting and substantially lowers time-to-first-token (TTFT) on a Galaxy S25 Ultra while maintaining or improving function-calling accuracy. Furthermore, the proposed state retriever achieves competitive retrieval quality using only 8.24\,MiB of retrieval storage. These results demonstrate that reusable recurrent states can serve as an efficient execution-time memory interface for scalable on-device function calling.
Abstract: We study zeroth-order optimization where solutions must minimize a cost d(s) while maintaining high probability under a complex generative prior \mathcalL(s) (e.g., a parameterized model). This reduces to sampling from a target distribution proportional to \mathcalL(s) e^-T \cdot d(s). Since classical model-based optimization (MBO) lacks finite-sample guarantees for expressive approximate learners, we introduce \emphcoarse learnability, a flexible statistical assumption requiring only that a learned model covers the target's probability mass within a polynomial factor. Leveraging this assumption, we design an iterative MBO algorithm called ALDRIFT with a sample correction step that provably approximates the target using only a polynomial number of samples. We apply this framework to globally optimizing non-convex objectives bounded by a quadratic envelope in \mathbbR^d, where we show this assumption is naturally satisfied for a family of ``optimistic'' posterior distributions. To reach global \varepsilon-optimality, this implies a sample complexity of \widetildeO(\log 1/\varepsilon), a rate characteristic of optimistic space-partitioning methods. We further justify coarse learnability as an assumption for generative priors theoretically, proving that in simple settings, parametric maximum likelihood estimation and over-smoothed kernel density estimators naturally satisfy it. Finally, one motivation for our framework comes from inference-time alignment. Though our primary contribution is formalizing the theoretical foundations of MBO, we provide qualitative evidence that, in simple settings, even primitive LLMs can shift their distributions toward lower-cost regions when fine-tuned with zeroth-order feedback.
PaperID: 2881, Poster
Abstract: Hypergraphs provide a natural representation for high-order relations, but real-world hypergraph data is often difficult to share or collect due to privacy, access, and annotation constraints. Hypergraph generation is therefore useful for constructing surrogate high-order relational data. Existing generators, however, typically build hyperedges through structural heuristics, propagation rules, or projected graph spaces, and thus do not directly model the incidence relation between nodes and hyperedges. Spectral representations offer a principled way to encode global and local relational patterns, but hypergraph spectra are intrinsically non-invertible: a spectral representation does not determine a unique discrete incidence structure. We propose HyperGen, a structure-aware generative framework that models hypergraph generation as continuous dynamics in a phase-enhanced spectral latent space. HyperGen constructs a Hermitian node--hyperedge spectral operator with learnable incidence phases, learns a conditional vector field that transports noise toward target spectral embeddings, and decodes the generated latents into node--hyperedge incidence matrices through a learned spectrum-to-structure mapping with denoising. Experiments on benchmark hypergraphs show that HyperGen more accurately recovers node--hyperedge incidences and better preserves spectral and geometric characteristics than existing baselines.
PaperID: 2882, Poster
Abstract: Multi-objective Bayesian optimization (MOBO) traditionally aims to approximate the entire Pareto front, but real-world deployment is often limited to a small, representative subset of solutions due to resource constraints. Current two-stage methods first approximate the full front before selecting solutions, wasting evaluation budget. We formalize this problem as Multi-objective Bayesian Optimization with Cardinality-aware Pareto Selection (MOBO-CAPS), where the goal is to identify a set of exactly M solutions with high hypervolume. We propose two practical cardinality-aware acquisition strategies: REHVI, a replacement-aware approximation to expected hypervolume improvement, and CHO-UCB, which constructs an optimistic M-point set by optimizing a collective UCB-based hypervolume objective. We prove a sublinear simple regret bound for the batched version of CHO-UCB, and empirically show that cardinality-aware acquisition design improves over standard MOBO followed by post-hoc subset selection across nine benchmarks.
PaperID: 2883, Poster
Abstract: Cross-domain open-vocabulary learning poses a unique and underexplored challenge, requiring models to generalize across both domain shifts and category shifts. To tackle this, we propose MAP: a parameter-efficient Multi-source domain-Adaptive Prompt tuning framework that leverages multiple labeled source domains to improve learning in novel, unlabeled target domains with unseen categories. MAP consists of two key components: Multi-Source Prompt Learning (MSPL) and Unsupervised Target Prompt Learning (UTPL). MSPL disentangles domain-invariant category semantics from domain-specific visual patterns by jointly learning shared and domain-aware prompts. UTPL enhances generalization in the unlabeled target domain by enforcing prediction consistency under text-guided style augmentations, introducing a novel entropy-minimization objective without relying on pseudo-labels. Together, these components enable effective alignment of visual and textual representations across both domains and categories. In addition, we present a theoretical analysis of the proposed prompts, examining their behavior through the lenses of fidelity and distinction. Extensive experiments on challenging CDOV benchmarks demonstrate that MAP achieves state-of-the-art performance with significantly fewer additional parameters.
PaperID: 2884, Poster
Abstract: Modeling interventions that include timing is an important task in many domains, such as treatments in healthcare, transactions in finance, and others. However, estimating these interventional distributions is challenging because standard next-token predictors, typically implemented using transformers, do not naturally support altering the timing of specific future events. We develop \emphCompeting Event Models (CEMs), an autoregressive generative approach that simplifies the estimation of interventions on both what interventions are applied and \emphwhen. Drawing from the competing risks literature, CEMs model independent probability distributions for latent times of each event type, where the earliest realized time determines the next event. We show that this framework enables straightforward interventions on event timing and efficient sampling, and, through careful design choices, can also model concurrent events. We evaluate CEMs on a cancer tumor volume simulator and on prediction tasks using the MIMIC-IV dataset. Our results show that CEMs accurately recover causal effects of sequential treatments in simulations while surpassing standard transformer baselines in predictive performance on 3 out of 4 real-world medical data tasks.
Abstract: While multi-agent reinforcement learning (MARL) has shown strong empirical performance, existing methods struggle with large number of agents, due to the combinatorial growth of joint interactions. Mean-field (MF) approximations address this by replacing pairwise interactions with interactions against population distributions, yielding tractable large-population policies. However, existing MF formulations focus on fully cooperative or purely competitive settings and fail to capture the mixed cooperative–competitive structure of team-based games. We extend mean-field learning to _zero-sum team games_, where agents cooperate within teams and compete at the team level. We show that such games admit \epsilon-optimal decentralized policies that depend only on local states and population distributions. Building on this structure, we propose MF-MAPPO, a scalable algorithm with a shared actor and a minimally informed critic per team. MF-MAPPO is trained directly in finite-population simulators rather than using mean-field oracles, thereby enabling deployment to realistic scenarios with thousands of agents. We further extend MF-MAPPO to partially observable settings via a simple gradient-regularized training scheme. Experiments on large-scale benchmarks in our simulation platform \textttMFEnv, including population games with analytical solutions and high-dimensional battlefield scenarios, demonstrate that MF-MAPPO outperforms existing MARL baselines and yields rich, heterogeneous behaviors.
PaperID: 2886, Poster
Abstract: Multimodal video fusion methods still suffer from limited performance due to the inadequate exploitation of historical frames and the simplified temporal consistency modeling. To address this, we propose MemoryFusion, a two-stage framework built upon cross-temporal memory learning, which leverages historical information across multiple temporal scales to enhance robustness and progressively refine temporal consistency. In the first stage, we introduce a Short-Term Memory Bank (STMB) to aggregate recent temporal features for enhancing the current frame, and a Long-Term Memory Bank (LTMB) to preserve representative features from extended temporal sequences through adaptive update rules. Meanwhile, we incorporate a Temporal Smoothing Module (TSM) to perform coarse temporal consistency modeling, suppressing abrupt background variations and establishing a stable temporal foundation. In the second stage, we design a Temporal Refinement Module (TRM) implemented as a lightweight 3D residual network, which conducts prior-guided fine-grained temporal refinement to recover subtle temporal dynamics and spatial details overlooked in the first stage. Extensive experiments demonstrate that MemoryFusion significantly improves both visual fidelity and temporal consistency by performing cross-temporal learning, outperforming the state-of-the-art in both frame-based and video-based fusion methods. The source code and pretrained models will be publicly released.
Abstract: Optimizing neural networks for quantized objectives is fundamentally challenging because the quantizer is piece-wise constant, yielding zero gradients everywhere except at quantization thresholds where the derivative is undefined. Most existing methods deal with this issue by relaxing gradient computations with techniques like Straight Through Estimators (STE) and do not provide any guarantees of convergence. In this work, taking inspiration from Nesterov smoothing, we approximate the quantized loss surface with a continuous loss surface. In particular, we introduce LOTION, g, a principled smoothing framework that replaces the raw quantized loss with its expectation under unbiased randomized-rounding noise. In this framework, standard optimizers are guaranteed to converge to a local minimum of the loss surface. Moreover, when using noise derived from stochastic rounding, we show that the global minima of the original quantized loss are preserved. We empirically demonstrate that this method outperforms standard QAT on synthetic testbeds and on 150M-, 300M-, and 600M- parameter language models. On the 150M INT4 benchmark with modern QAT baselines, LOTION further achieves the best quantized validation loss.
Authors: Boxiang Zhang, Baijian Yang
Abstract: Transformers achieve strong accuracy but incur high compute and memory cost. Structured pruning reduces inference cost, but most methods rely on retraining or multi-stage optimization, which limits post-training deployment. We propose CORP, a closed-form one-shot structured pruning method that removes MLP dimensions and attention substructures using only unlabeled calibration data without labels, gradients, or fine-tuning. CORP formulates structured pruning as a representation recovery problem. It models removed components as affine functions of retained components and derives closed-form ridge regression solutions that fold compensation into model weights. This minimizes the expected representation error under the calibration distribution. Experiments on ImageNet with DeiT reveal strong redundancy in both MLP and attention representations. Without compensation, one-shot structured pruning causes severe accuracy loss. With CORP, models retain high accuracy under aggressive sparsity. On DeiT-Huge, CORP achieves 83.27% Top-1 accuracy after pruning 50% of both MLP and attention structures.
PaperID: 2889, Poster
Authors: Shicheng Liu, Zeyu Zhang, LEI LU, Kexuan Sun, Thuy Vu
Abstract: Language model query generation is the task of using a language model to translate a natural-language request into a formal query, which is then executed to retrieve information from databases or knowledge graphs. Recent approaches rely on reinforcement learning (RL) to enhance the model’s reasoning and improve query correctness. However, existing methods provide only sequence-level rewards, resulting in sparse feedback that is difficult to optimize. To address this limitation, we propose a reward shaping method that learns a token-level, and thus dense, reward function, providing fine-grained supervision to guide the language model toward correct queries. We formulate this problem as a bilevel optimization problem where the upper level learns a token-level shaping reward from ground truth sequence-level reward and the lower level optimizes the language model using the learned shaping reward. We propose a single-loop algorithm to solve this problem and theoretically guarantee that the algorithm converges at O(\frac1\sqrtK)+O(\frac\log K\sqrtK). Extensive experiments demonstrate that our method outperforms the baselines.
PaperID: 2890, Poster
Abstract: Large language models are increasingly deployed for open-ended question answering, yet their answers rarely come with calibrated uncertainty. We propose CAELUM, a framework that builds semantic prediction sets and separates aleatoric from epistemic uncertainty in their construction. CAELUM samples answers under prompt-ensembled clarifications, clusters them by semantic equivalence, estimates aleatoric and epistemic components via kernel Von Neumann entropy, and calibrates cluster-APS nonconformity scores with split conformal prediction. A per-instance weight w_\lambda(x)=1+U_\mathrmale(x)+\lambda\,U_\mathrmepi(x) controls how epistemic uncertainty enters the set, and a per-cell coefficient \lambda^\star is fit on a held-out split to minimise prediction-set size. Our central observation is that, in open-ended generation, coverage on cluster-APS scores is silently capped by \emphmatchability: the fraction of items whose reference answer appears in any sampled cluster. We make this dependence verifiable with a Mondrian split-conformal bound on a learned stratum, where matchability is predicted from pre-generation question features by a lightweight classifier. Across 18 (model, dataset) cells and four split seeds on three open-weight 4--9B LLMs, the learned classifier lifts predicted-matchable coverage from 0.61 overall to 0.71 on non-fallback rows (0.82 vs 0.63 when oracle-fallback rows are included as upper bounds), and the classifier transfers across LLMs at AUC 0.78 (in-sample 0.80). The cell-level \lambda^\star doubles as a regime diagnostic: large values flag epistemic dominance (retrieve, abstain) and \lambda^\star\to 0 flags aleatoric dominance (clarify), with a per-instance log-ratio R(x) extending the read-out to individual queries.
PaperID: 2891, Poster
Abstract: Real-world image datasets are deeply hierarchical, yet unsupervised hierarchical representation learning remains underdeveloped. Hyperbolic space is geometrically ideal for embedding trees with low distortion due to its exponential volume growth, but existing unsupervised hyperbolic methods either fail to scale beyond a few thousand points or depend heavily on class supervision. We bridge this gap with HyperTree, a fully unsupervised, scalable, and capacity-efficient framework for hierarchical representation learning in hyperbolic space, built on three contributions: (i) a geometrically consistent, closed-form definition of the hyperbolic Lowest Common Ancestor, valid across hyperbolic manifolds of arbitrary curvatures; (ii) a differentiable hierarchical objective based on Dasgupta's cost that is bounded and invariant to dataset scale; and (iii) a sample-complexity guarantee that enables scaling to datasets such as ImageNet-1k and iNaturalist-2018. Across four benchmarks, HyperTree outperforms Euclidean baselines by wide margins on hierarchical dendrogram purity, matches supervised hyperbolic baselines on CIFAR-100, and surpasses them on TinyImageNet at 16 dimensions — a highly capacity-efficient regime where competing methods collapse.
PaperID: 2892, Poster
Abstract: Modern foundation models are trained through multistage pipelines, typically beginning with unsupervised pretraining, followed by supervised fine-tuning (SFT) on step-by-step solutions, and reinforcement learning (RL) with bulk feedback. Empirically, the design of these stages, particularly the compute allocated to each, strongly affects performance. Yet our theoretical understanding of multistage training remains limited. Here, we introduce a simple, analytically tractable model of multistage training. The model reproduces several qualitative phenomena: SFT and RL can be complementary, pretraining quality imposes a performance ceiling with power-law scaling, optimal compute allocation varies with data quality and pretraining level, and RL can selectively amplify task-relevant pretrained skills. Our model distills the complex behaviors into key ingredients, allowing clearer understanding of how and where they arise. Overall, our model provides a theoretical starting point for explaining the phenomenology of multistage training pipelines in the compute-limited regime.
PaperID: 2893, Poster
Abstract: Frontier AI models and multi-agent systems have led to significant improvements in mathematical reasoning. However, for problems requiring extended, long-horizon reasoning, existing systems continue to suffer from fundamental reliability issues: hallucination accumulation, memory fragmentation, and imbalanced reasoning-tool trade-offs. We introduce , a multi-agent framework that systematically addresses these challenges through meta-level supervision and structured Reasoner-Verifier interaction. STAR-Math is structured as an orchestrated state machine with nested challenge-step-replan loops, governed by a reasoning-free Python orchestrator that separates control from inference and bounds error propagation through trace-back and re-planning. Our key innovation is a that maintains cross-attempt memory and exercises meta-level control by issuing high-level strategic guidance or mandatory directives, so the system can escape unproductive loops rather than stagnate or over-rely on tools. STAR-Math achieves state-of-the-art results on all eight top-tier competition benchmarks: AIME 2025-2026, MathArena Apex Shortlist, MathArena Apex 2025, Putnam 2025, IMO 2025, HMMT February 2026, and USAMO 2026. It obtains perfect scores on AIMEs, Putnam, and HMMT, and shows its largest margin on Apex 2025, scoring 93.75% compared with 80.21% by the strongest baseline GPT-5.5. Ablation studies show that the gains arise from the framework's orchestration rather than from model-level diversity since removing key components or substituting in mixed backbones consistently weakens performance.
PaperID: 2894, Poster
Authors: Zhanjiang Yang, Lijun Sun, Yueming Li, Meng Li, Yijun Yang
Abstract: The overarching goal of reinforcement learning (RL) is to efficiently optimize the agent's policy for various environments, tasks, dynamics, and even user preferences, which often consist of multiple and potentially conflicting rewards/objectives. This necessitates the learning of Pareto-optimal policies that can navigate different trade-offs among objectives. However, existing multi-objective RL (MORL) methods, whether trained online or offline, are typically limited to task- or objective-specific policy learning and do not address generalization of learned skills to unseen tasks/objectives. In this work, we propose the ) for MORL, inspired by the pre-mid-post-training paradigm underlying recent foundation models. Specifically, we instantiate FP2 as a multi-modal diffusion transformer equipped with a likelihood-free flow matching loss, enabling unified, scalable, and instructable learning of diverse Pareto-optimal trajectories. We further develop a three-stage training pipeline to gradually optimize the policy: (1) pre-training on a large-scale MORL dataset, (2) mid-training on a carefully curated high-quality dataset, and (3) post-training on few collected trajectory data from new tasks and trade-offs. Empirical results demonstrate that FP2 consistently outperforms previous state-of-the-arts on the D4MORL benchmark, achieving superior zero-shot interpolation and extrapolation on complex Pareto manifolds as well as efficient few-shot adaptation to new task/dynamics scenarios.
PaperID: 2895, Poster
Authors: Junfeng Liu, Xingquan Li
Abstract: Routing is foundational in VLSI physical design, directly shaping timing, power, and area. Recent pre-trained Seq2Seq models recast routing as autoregressive generation over DFS-serialized trees, but inherit a text backbone blind to the physical space their tokens describe: DFS decouples sequence position from physical coordinate, breaking spatial reasoning at branching points, in load selection, and along the autoregressive trajectory. We name these failures the \emphThree Spatial Blindnesses and propose ChiP-STAR, which augments a T5Gemma backbone with three lightweight modules: a dual-frame rotary encoding coupling sequence order with 3D proximity, a Fourier-factorized attention prior with memory linear in sequence length, and a cumulative coordinate-noise schedule closing the train--inference gap, together adding under 0.5% parameters. We rebuild the AiEDA corpus into the largest geometry-annotated routing dataset to date (6.4M nets, 1.8B tokens, 556\,GB). ChiP-STAR-Large surpasses iPCL-Large by 5% on Leaf IoU and Connectivity, and commercial ECO signoff yields 3% shorter wirelength and 11% fewer vias. Our code is publicly available at \urlhttps://anonymous.4open.science/r/iPCL-R-0778/.
PaperID: 2896, Poster
Abstract: The efficiency and fidelity of 3D Gaussian Splatting critically depend on adaptive density control, which grows and suppresses Gaussian primitives throughout optimization. However, the standard pipeline still relies on heuristic view-space gradients, scale thresholds, and global opacity resets, rather than explicitly assessing the loss-reducing effect of each split, clone, or opacity update. This mismatch can allocate primitives to already-sufficient regions, miss under-reconstructed structures, and destabilize training through blanket opacity changes. In this work, we recast adaptive density control as a collection of local loss-reduction decisions. We derive efficient Taylor-based per-Gaussian surrogates that predict the objective change caused by three primitive-level operations: splitting, cloning, and opacity decay. This yields unified loss-aware framework with three components. First, we extend second-order splitting to a normalized joint position-scale space, enabling children to adapt not only their locations but also their anisotropic extents. Second, we formulate duplicate cloning as a surrogate minimization problem, selecting only Gaussians whose duplicated contribution is predicted to improve reconstruction and assigning each child a principled opacity. Third, we replace global opacity reset with selective soft opacity decay that weakens harmful primitives while preserving useful ones. Across standard 3DGS benchmarks, our method produces more targeted primitive allocation and achieves improved rendering quality with comparable or smaller Gaussian budgets compared to baselines.
Abstract: Fine-tuning APIs make frontier LLMs easy to customize, but they can also weaken safety alignment during fine-tuning. While prior work shows that benign supervised fine-tuning (SFT) can reduce refusal behavior, deployed fine-tuning pipelines increasingly support preference-based objectives, whose safety risks remain less understood. We show that Direct Preference Optimization (DPO) introduces a stronger and harder-to-audit failure mode. We propose a truly benign DPO attack using only 10 harmless preference pairs, the minimum data scale accepted by OpenAI’s fine-tuning service. Each pair contains a benign prompt, a normal helpful answer as the preferred response, and a refusal as the dispreferred response. Unlike prior benign fine-tuning attacks, our data exhibits no suspicious behavior: it is practically indistinguishable from the fine-tuning request of a legitimate user seeking to reduce over-refusal, making harmful intent almost impossible to infer from the request alone. Nevertheless, because DPO directly optimizes the model to prefer helpful answers over refusals, this seemingly benign objective broadly suppresses refusal behavior and transfers to harmful prompts outside the fine-tuning data. Across OpenAI models supporting DPO fine-tuning, our attack achieves attack success rates of 59.13% on GPT-4o, 70.20% on GPT-4.1, 54.80% on GPT-4.1-mini, and 81.73% on GPT-4.1-nano, at costs of only \\1.7, \\1.7, \\0.3, and \\0.1. Moreover, on open-weight models that do not impose minimum data requirements, we find that this effect can emerge from even a single benign preference pair.
PaperID: 2898, Poster
Abstract: Reinforcement learning (RL) is increasingly used in post-training to adapt large models to new reasoning tasks. Recent forgetting analyses, however, are mostly answer-centric: they report reward or accuracy on previously learned tasks and implicitly treat stable answers as evidence of stable reasoning. We show that this view is incomplete. In our sequential reasoning settings, externalized reasoning traces can change substantially in structure, strategy, and sample-wise behavior before answer-level metrics visibly collapse, and harder full-parameter staged runs can later degrade into catastrophic forgetting on earlier tasks. We refer to the trace-first regime as \emphreasoning drift. To address it, we propose \emphTrace-Anchored Regularization Prior (TARP), a trace-centric regularizer that mines reliable reasoning spans from successful rollouts and uses them to build a compact trace-derived quadratic prior for subsequent RL updates. Across multiple reasoning task sequences and model backbones, TARP reduces selected measured drift signals and improves earlier-task preservation in several evaluated settings, while also revealing a retention-verbosity trade-off that answer-only evaluation hides.
Abstract: Current test-time scaling (TTS) techniques enhance large language model (LLM) performance by allocating additional computation at inference time, yet they remain insufficient for agentic settings, where actions directly interact with external environments and their effects can be irreversible and costly. We propose ARTIS, Agentic Risk-Aware Test-Time Scaling via Iterative Simulation, a framework that decouples exploration from commitment by enabling test-time exploration through simulated interactions prior to real-world execution. This design allows extending inference-time computation to improve action-level reliability and robustness without incurring environmental risk. We further show that naive LLM-based simulators struggle to capture rare but high-impact failure modes, substantially limiting their effectiveness for agentic decision making. To address this limitation, we introduce a risk-aware tool simulator that emphasizes fidelity on failure-inducing actions via targeted data generation and rebalanced training. Experiments on multi-turn and multi-step agentic benchmarks demonstrate that iterative simulation substantially improves agent reliability, and that risk-aware simulation is essential for consistently realizing these gains across models and tasks.
Abstract: When an algorithm makes a decision affecting multiple people, we implicitly have a social choice problem: how should differing opinions about the right course of action be reconciled into a single outcome? We show that if we make welfare consequences of alignment a first-order consideration, this problem can be reformulated as linear optimization on a convex , making it amenable to toolkits from welfare economics and mechanism design. This reformulation provides a clearer understanding of how different protocols shape an algorithm's externalities. Our linear optimization framework uncovers a straightforward solution to strategyproof alignment mechanism—random dictatorship that can be represented as a single parameter, which regulates externalities from manipulation. In addition, our framework allows us to express a family of constrained welfare mechanisms that address participation externalities—how an individual's participation can affect others in the population—and maximize social welfare subject to guarantees on individual outcomes. We illustrate these results empirically using real human preferences over kidney donation assignment, charitable food distribution, LLM responses, and trolley problems.
Abstract: We propose the Monge Inception Distance (MIND), a metric for evaluating generative models that addresses key limitations of the widely adopted Fréchet Inception Distance (FID). The MIND metric leverages the sliced Wasserstein distance to compare distributions by averaging one-dimensional optimal transport distances, efficiently computed via sorting. This approach circumvents the estimation of high-dimensional means and covariance matrices, which underlie FID's poor sample complexity and vulnerability to adversarial attacks. We empirically demonstrate three primary advantages: (i) it is more sample-efficient by one order of magnitude, (ii) it is faster to compute by two orders of magnitude, (iii) it is more robust to adversarial attacks such as moment-matching. We show that MIND with 5k samples can replace the evaluation performance of FID with 50k samples, providing high correlation with this standard benchmark and superior discriminative performance. We further demonstrate that even smaller sample sizes (e.g., 1k or 2k) remain highly informative for rapid model iteration.
Abstract: Reinforcement learning with verifiable rewards (RLVR) post-trains language models on multi-step reasoning by assigning a single outcome reward uniformly across all tokens in a trajectory, regardless of which steps contributed to success or failure. Improving credit assignment can address this limitation by enabling targeted refinement of faulty reasoning steps, rather than updating entire trajectories uniformly. Resets are one such simple mechanism, enabling more precise credit assignment by returning to an intermediate state and resampling its continuation, so that outcome differences can be attributed to decisions made at that point. We propose two such methods: Random-Reset Policy Optimization (RRPO), where reset states are drawn randomly from reasoning steps, and Self-Reset Policy Optimization (SRPO), where the model self-localizes the erroneous step in an incorrect trajectory and resets there. We analyze these methods within a Conservative Policy Iteration (CPI) framework. We extend it with a credit-assignment oracle that targets improvable states, defined as those whose advantage exceeds a threshold, and show that it yields provable improvements over random resets. Across models and reasoning benchmarks, SRPO consistently outperforms standard GRPO and RRPO by sampling multiple suffix continuations at a self-localized reset and learning from their rewards, using only the model itself with no external supervision.
PaperID: 2903, Poster
Abstract: A unified cross-modal person re-identification targets robust cross-modal retrieval with queries from diverse modalities like visible, infrared, and text.Such framework requires multi-modal data and centralized training, raising data costs and privacy risks. Hence, we propose a modality-heterogeneous federated ReID task to learn a unified cross-modal framework from clients with different cross-modal data.However, it faces two challenges.First, varying inter-modality complexity across tasks leads to different retrieval patterns and parameter space. Second, different intra-modality complexity results in imbalanced contributions when merging LoRA modules.Existing federated learning methods use uniform merging and transmit full models, which hinders cross-modal knowledge integration and increases communication overhead. To address these issues, we adopt Low-Rank Adaptation (LoRA), a lightweight module for fine-tuning large models through low-rank updates, enabling clients to send compact parameters.On top of LoRA, we develop MIRAGE (Modality-heterogeneous Intelligent federated ReID Aggregation with Global Efficiency), adapting aggregation to inter- and intra-modality complexity.When merging visible LoRA, MIRAGE selects key ranks to protect retrieval patterns while retaining shared ranks for common ability.As for infrared and text LoRA, it reduces conflicts with an adaptive pruning method. In both strategies, task complexity is incorporated into aggregation to balance contributions. Experiments on cross-modal retrieval tasks demonstrate the superiority of MIRAGE.The code and dataset will be publicly released.
PaperID: 2904, Poster
Authors: Dongmin Lee, Anuran Makur, Japneet Singh, Boyu Xu
Abstract: Many modern machine learning methods, such as reinforcement learning from human feedback (RLHF) to train reward functions, utilize preference learning models like Bradley-Terry-Luce (BTL) to learn latent scores of items based on pairwise comparison data. Despite the successes of such models in specific applications, the theoretical justification for the employed models has often been unclear. In this work, we provide a hypothesis testing procedure to test whether a given dataset of pairwise comparisons follows a contextual generalized Thurstone (CGT) model under certain classes of latent functions. The latent functions determine the latent scores, and can be expressed as weighted sums of basis functions. Our CGT model encompasses a wide class of preference learning models, including the popular BTL model. In particular, we prove critical thresholds for our tests in the minimax sense, showing that they scale like 1/\smash\sqrt|\mathcalE|^1/2k under certain important regimes, such as complete graphs and perfect matchings. Furthermore, to establish our testing results, we also derive error bounds for CGT parameter estimation that generalize and improve on known results in the BTL and non-contextual settings. Finally, we conduct experiments on various synthetic and real-world datasets to verify our theoretical results.
Abstract: Large Language Models (LLMs) are increasingly used to brainstorm and evaluate research ideas, yet assessing such judgments is fundamentally difficult because the true impact of a new idea may take years to emerge. We address this challenge by using the impact forecasting of human-authored manuscripts as a verifiable proxy task. In a prospective forecasting study, we find that frontier LLMs fail to reliably distinguish high-impact papers from ordinary publications, suggesting that static text-based judging is insufficient for scientific evaluation. To address this limitation, we propose FAME (\underline\textForecasting \underline\textAcademic Impact via Continuous-Time \underline\textManifold \underline\textEvolution), a spatiotemporal framework for modeling the dynamic trajectories of scientific topics. FAME projects papers into a dynamic latent space informed by textual features and a verified knowledge-flow graph, learning geometric constraints that align impactful manuscripts with the forward momentum of their fields. Experiments on 3,200 arXiv papers across three fast-evolving subfields show that FAME consistently and substantially outperforms state-of-the-art LLM evaluators in prospective multidimensional impact forecasting. Furthermore, integrating FAME's dynamic geometric signals into LLMs significantly improves their forecasting performance. These results support manuscript impact forecasting as a useful, measurable proxy benchmark and position FAME as a strong, trajectory-aware foundation for automated scientific evaluation.
PaperID: 2906, Poster
Abstract: While vision-language models (VLMs) perform well on general tasks, their spatiotemporal reasoning remains limited. Existing distillation and Reinforcement Learning (RL) methods are often computationally expensive and poorly generalizable. Recent studies suggest that model self-improvement by experience distilled from success and failure trajectories is a promising direction. However, in VLM tasks where reasoning must be grounded in visual information, valuable signals of reasoning quality arise not only from final outcomes, but also from how faithfully the reasoning path interacts with visual information. We examine the dependency between reasoning paths and visual information, and observe that high-quality paths exhibit stronger visual anchoring and greater visual dependency. Building on this observation, we introduce the Spatiotemporal Path-Aware Reasoning paradigm based on experiential Knowledge (Spark) that encourages the model to explore and exploit valuable experiences derived from paths with subtle quality differences. To enable diverse and efficient path exploration, we integrate Monte Carlo Tree Search (MCTS) with fine-grained visual rewards and submodular optimization, allowing the search process to prune redundant branches while preserving critical reasoning paths. We further present a training-free experience construction mechanism that converts contrastive pairs of distinct-value paths into structured experiences, which are dynamically injected into the VLM through similarity-based retrieval during inference. Extensive experiments demonstrate that Spark enables Qwen3.5-9B to achieve an average accuracy of 61.21% across three benchmarks, outperforming Gemini-3-Pro (57.20%). Further analyses verify the flexibility and cross-model generalization of our method, highlighting the crucial role of path-aware experience in advancing spatiotemporal intelligence.
Abstract: Speculative decoding accelerates large language model inference through a draft-then-verify paradigm. Building on this, tree-structured methods improve inference by organizing proposals into multiple candidate paths, increasing the accepted length. However, existing tree-structured methods use a single drafter for all drafting steps, creating a dilemma: a smaller drafter is fast but yields lower-quality trees, whereas a larger drafter improves tree quality but suffers from high latency. To address this, we propose TreeGraft, a multi-drafter framework in which drafters of different costs jointly construct a shared draft tree. TreeGraft uses the stronger drafter to rescore candidates by updating scores assigned by the weaker drafter, reselect grafting positions, and recover promising paths left unexplored. It also integrates stronger drafter expansions non-destructively, preserving existing branches that may still be accepted by the target model. Together, these designs improve the quality of the shared draft tree. To control the drafting cost, TreeGraft introduces a lightweight scheduler distilled from an offline value system to decide when to call the stronger drafter. Across 10 model pairs and 6 benchmarks, TreeGraft outperforms the better of the two fixed single-drafter endpoint strategies by 15.1% on average, reaching a maximum gain of 26.6%. Our code is available at https://anonymous.4open.science/r/TreeGraft-E983.
PaperID: 2908, Poster
Authors: Jonathan Hornewall, Tito Homem-De-Mello, Vincent Leclere
Abstract: Decision-focused scenario learning for contextual two-stage stochastic programming has emerged as a promising research direction, but open questions remain. A general theory for when scenario generators can yield optimal decision policies is currently lacking. Many existing methods require strict assumptions on which second-stage data are allowed to be random. Generally, the non-differentiability of the decision-focused objective is a source of difficulty. This paper makes three theoretical contributions, each aiming to remedy one of these challenges. First, we present a comprehensive theory for when scenario generators can induce optimal policies, showing that learning a single out-of-support scenario component is often sufficient. Second, we demonstrate a theoretical and methodological equivalence result between learning different second-stage data components, relaxing assumptions required for applying existing methods. Third, we study log-barrier smoothing enabled training, presenting representability and optimization results. Finally, leveraging the theory, we propose a novel method inspired by interior-point methods from classical optimization, and validate it on standard benchmarks.
Abstract: Failure attribution for LLM-based agentic systems, i.e., identifying which steps in a failure trajectory caused the task to fail, is critical for debugging and improving these systems. Existing approaches either rely on prompting-based pipelines, which are computationally expensive, or require post-training on failure trajectories with step-level error annotations, which are costly to collect and difficult to scale. We argue that a practical failure attribution model should be lightweight and trainable without step-level supervision on failure data. To this end, we address unsupervised failure attribution, i.e., training exclusively on successful trajectories and identifying error steps at inference time given a failure trajectory. We propose OAT, which casts this problem as one-class learning with neural controlled differential equations, modeling the dynamical pattern of successful trajectories in latent space. At inference time, each step in a failure trajectory is assigned an anomaly score based on its deviation from the dynamics learned on successful trajectories, which is then used to form a set of error steps. With training on only 100 successful trajectories, experiments show that OAT is 200--5000× faster than prompting-based baselines, and, at the same time, consistently outperforms them in both in-domain and out-of-distribution datasets with +20% and +7% F1 scores, respectively, demonstrating that OAT is a promising and efficient direction for diagnosing agentic system failures.
PaperID: 2910, Poster
Abstract: Many problems are inherently compositional: solving them requires a sequence of intermediate decisions that jointly determine the final answer. A central question is when such a compositional structure can be learned from outcome-level feedback alone, where supervision of the intermediate steps is unavailable. We study this question for autoregressive models trained using reinforcement learning with verifiable rewards (RLVR). We identify the \emphtask-advantage ratio, a joint property of the task and the model, which measures whether intermediate decisions present an advantage in reaching a correct final solution. We show that this ratio governs learnability in our setting: when the advantage is present, RLVR efficiently learns the target composition, while when it is absent, training can converge to suboptimal compositions. We further show that the required advantage arises naturally in several structured problems, but may depend critically on the quality of the initial model. Our results help to clarify when compositional problems can be learned from final rewards alone.
Authors: Gabriel Chuang, Augustin Chaintreau
Abstract: An extensive line of work studies fairness interventions for network embeddings, but less is known about their baseline behavior. In this work, we ask: how do baseline embeddings (without fairness interventions) produce disparate effects at the representation level? We analyze the asymptotic behavior of low-dimensional embeddings on stochastic block model (SBM) graphs, which encode both homophily and group structure. We characterize exact conditions under which embeddings cause information loss, showing that the amount of information loss depends directly on the graph’s density and assortativity. Notably, very different graphs can produce identical embeddings in the limit, and this non-invertibility disproportionately affects smaller and sparser communities. As a result, simple downstream tasks, such as link prediction, introduce higher error rates for these communities, helping explain disparities widely observed in practice.
PaperID: 2912, Poster
Authors: Guangbin Zhang, Zheyang Luo, Jiangming Liu
Abstract: Real-world tasks rely on information from multiple sensory sources, motivating multimodal learning as a core paradigm in modern machine learning. However, multimodal learning suffers from modality laziness, where one modality suppresses the contribution of the other modalities by dominating the training process. Recent studies of proposing various metrics to identify lazy modalities and manipulate optimizations to mitigate this imbalance, facing two main problems: 1) The identification of lazy modality is incomplete and overly sharp; 2) The dominated modality can be lagged by under-optimized lazy modalities. To address these issues, we propose Adaptive-Margin Masking and Restoration (AMRe), which introduces adaptive-margin modality identification and restoration optimization to balance dominated and lazy modalities. Experimental results show that AMRe consistently outperforms competitive baselines and several state-of-the-art methods by achieving significant improvement on standard multimodal benchmarks. The codes are released in https://anonymous.4open.science/r/AMRe-5E69.
Abstract: Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recognition or vision-language alignment, leaving motion understanding to downstream policies. We introduce DynaFLIP, a dynamics-aware multimodal pre-training framework that pushes motion understanding upstream into perception. We construct image–language–3D flow triplets from heterogeneous human and robot videos, and use these triplets as training-time supervision to shape an image-only encoder. Our key idea is to encourage the three modalities to span a small simplex volume in the shared hyperspherical space—a smaller simplex volume indicating stronger alignment. To avoid the geometric ambiguity and trivial collapse of naive volume minimization, we combine simplex-volume minimization with a cosine regularizer and a contrastive objective. Our analyses show that DynaFLIP focuses on control-relevant regions critical for manipulation. The resulting dynamics-aware representations serve as reusable visual backbones and consistently outperform baselines across diverse downstream policies, including VLAs. We validate this across diverse simulation and real-world setups, with gains reaching +22.5% under out-of-distribution scenarios. Our results suggest that robot generalization improves when visual representations are trained to encode not just what is present, but how the world changes under action.
Abstract: The human vocal tract, the cavity consisting of one's throat, mouth, lips opening, etc., filters one's voice to create the sounds we know as speech. In this paper, we present a differentiable and GPU parallelizable acoustic simulator that synthesizes speech by propagating sound along an acoustic tube representation of the vocal tract, and via its gradients, solves the inverse problem: reconstructing the geometry of their vocal tract solely from the sound it produces. Although learning the inverse mapping between geometry and sound is notoriously non-convex, we discover that gradient descent succeeds with three technical contributions: (1) we design a frequency domain formulation of the vocal tract's fluid dynamics that is 70x more GPU parallelizable than finite differences in time, (2) we integrate a differentiable model for turbulence to synthesize consonants, and (3) similar to prior work in implicit neural representations (INRs) and neural radiance fields (NeRFs), we find that parameterizing the geometry with a neural network accelerates convergence and escapes local minima that trap discrete representations. Because the simulator is differentiable, it is readily integrated with other deep learning pipelines to enable novel speech and medical imaging applications. Our model can be used for instance, in singing instruction or language learning, where a visualization of one's vocal tract can help people understand how their vocal tract maps to different speech sounds. To enable these tasks, (1) we demonstrate self-supervised autoencoding of vocal tract shapes across 11 languages, and (2) we couple our simulator with a generative model of MRI images to reconstruct one's moving vocal tract from only their speech, no paired data required.
PaperID: 2915, Poster
Abstract: Generative models are often required to serve many downstream conditional tasks. In the training-free setting, a pretrained unconditional model is reused as a prior and combined with a task-specific likelihood to sample from the corresponding posterior, without retraining the generative model. While such conditional samplers have been developed for diffusion and flow generative models, training-free conditional sampling for Schrödinger bridge (SB) generative models remains unexplored, despite the growing appeal of SBs as flexible stochastic transport models between general source and target distributions. We propose BridgeTwist, a training-free conditional sampler for pretrained SB generative models, formulated as a tempered twisted sequential Monte Carlo scheme. At each step, the sampler uses a closed-form plug-in twist surrogate, which combines the velocity and score fields exposed by SB pretraining into an explicit endpoint predictor. Since this surrogate is less reliable near the source, we use a time-varying temperature to temper the twist, downweighting its influence at early steps. Telescoping importance weights then cancel all intermediate twist factors, yielding an asymptotically exact conditional sampler without any auxiliary twist network or pilot rollouts. Empirically, BridgeTwist consistently outperforms training-free baselines on class-conditional sampling and inpainting on MNIST and CIFAR-10, and on text-to-image generation on CelebA-HQ.
Abstract: Objects in post-fire environments often undergo irreversible physical transformations that change their geometry, material state, and visual appearance. Detecting and identifying these remnants is important for locating hazards, reconstructing pre-incident contents, and inventorying losses. Unlike standard image corruptions, fire damage changes the physical structure of the object itself. To study this setting, we introduce TRACE, a transformation-aware benchmark for post-fire object understanding. TRACE contains 21.4K real-image-grounded synthetic scenes and paired object-level pristine-to-degraded progressions spanning 499 object identities across 189 categories. We define five tasks that evaluate both localization and pre-degradation understanding: degraded-object detection, pristine-state recovery and retrieval, original material recovery, pristine description generation, and functional reasoning. Existing models degrade sharply as fire damage becomes more severe. From the least to the most severe degradation level, RF-DETR mAP decreases by 71% relative, while InternVL3.5 retrieval Recall@1 drops from 93.85 to 28.11. To address this, we propose the Feature Restoration Module, or FRM, a lightweight plug-and-play module that maps degraded encoder features toward pristine-aligned representations while keeping the host model frozen. FRM is trained only with paired feature supervision and improves scene-level detection, CLIP/SigLIP2 feature recovery, and all four object-level VLM tasks. Its gains are larger under more severe degradation. Across VLM hosts and severity levels, FRM improves retrieval by 12.5% on average, material recovery by 20.1%, description generation by 13.2%, and functional reasoning by 12.4%.
Abstract: We study model-free reinforcement learning (RL) in non-stationary finite-horizon episodic Markov decision processes (MDPs) without prior knowledge of the non-stationarity. We focus on the piecewise stationary (PS) setting, where both rewards and transition dynamics can change at unknown times. We first revisit existing state-of-the-art approaches and identify theoretical and practical limitations that change the current landscape of performance guarantees. To characterize the difficulty of the problem, we establish the first minimax lower bounds for PS-RL in tabular and linear MDPs. We then introduce (DARLING), a modular wrapper for PS-RL that applies to both tabular and linear MDPs, without knowledge of the changes. In tabular MDPs, under change-point separability and reachability conditions, DARLING improves the best known dynamic regret bounds and matches our minimax lower bound. In linear MDPs, DARLING matches the minimax lower bound when the relevant reachability parameters are known, and our analysis clarifies the structural obstacles that distinguish this setting from the tabular case. Finally, through extensive experimentation across diverse non-stationary benchmarks, we show that DARLING consistently surpasses the state-of-the-art methods.
Abstract: Image-to-layer decomposition converts a flattened image into editable RGBA layers, enabling element-level editing in design workflows. Existing diffusion-based systems typically adapt large pretrained text-to-image models and introduce RGBA autoencoders or variable-layer architectural modules. We revisit this design choice and ask whether layer decomposition actually requires latent autoencoding or T2I pretraining. We introduce PixelART, a pixel-space rectified-flow Transformer trained from scratch for image-to-layer decomposition. PixelART directly denoises regional RGBA pixel patches with a single-stream multi-modal diffusion Transformer, avoiding RGBA-VAEs, pretrained T2I backbones, and layer-specific decoders. We identify a key property of the task: high-noise timesteps determine layer assignment and coarse layer organization, while low-noise timesteps mainly refine color, alpha, texture, and boundaries. Based on this observation, we propose a \emphterminal-boosted timestep sampling to increase training coverage in the high-noise assignment regime. Trained on 4M multi-layer design templates up to (1024×1024), PixelART achieves state-of-the-art layer and composite reconstruction on \designbenchmark while using over (10×) fewer parameters and substantially lower latency and memory than VAE-based diffusion baselines. Ablations show that pixel-space \mathcaox-prediction, high-noise timestep coverage, and data/model scaling are critical, while T2I initialization provides no measurable final gain in the data-rich I2L setting.
PaperID: 2919, Poster
Abstract: A common way for trajectory planning is to leverage generative models trained on large collections of expert trajectories. At inference time, the model generates executable trajectories by conditioning on task goal constraints. However, trajectory-based methods rely on costly supervision, scale poorly with sequence length, and often generalize poorly to unseen constraints such as novel start-goal pairs. We propose an alternative to learn the underlying state-space manifold and use the geometry of the manifold for trajectory planning. This approach requires only state observations and enables generalization to unseen constraints by constructing trajectories on the learned manifold of the state space. Experiments on classical planning problems in maze demonstrate the effectiveness of our work. We further show that the method can be used in higher dimension space where table-top robot arms are considered. Our method is able to find feasible path given only infeasible straight line reference, and be comparable to State-of-the-Art trajectory-based model without actually learning on trajectory data.
Abstract: We establish sample complexity results for stochastic optimization over the integers, especially with a view to understand the complexity with respect to the corresponding continuous optimization problem. We show that integer optimization can sometimes require strictly more samples and sometimes strictly smaller number of samples, depending on the structure of the objective and constraints. 1. For Lipschitz objectives over subsets of the \ell_\infty ball, the statistical complexity of general stochastic mixed-integer, nonlinear, nonconvex optimization is exactly the same as stochastic linear optimization with just bound constraints. 2. For Lipschitz objectives over subsets of the \ell_2 ball, we show that integer optimization can require strictly smaller sample size compared to the continuous setting in a certain regime. To get to this result, we also establish tight sample complexity results for nonconvex continuous stochastic optimization which, to the best of our knowledge, do not appear in prior work. 3. For strongly convex, smooth objectives, integer optimization has high statistical complexity compared to the continuous setting. In particular, we show that integer optimization requires \Omega(1/\epsilon^2) samples to report an \epsilon-approximate solution, compared to the well-known O(1/\epsilon) sample complexity from the continuous optimization literature.
PaperID: 2921, Poster
Abstract: Covariate-adjusted response-adaptive randomization (CARA) designs can improve statistical efficiency and participant welfare in randomized experiments by learning from accrued data and dynamically updating treatment allocation. In practice, two issues limit their impact. Many experiments are sample-size constrained because enrollment is difficult and trials are expensive. At the same time, studies increasingly collect rich pre-treatment information, including unstructured data such as text, images, and clinical notes, yet most CARA designs rely on a small set of structured covariates when computing allocation probabilities. Recent advances in large-scale pre-trained large language models (LLMs) offer a new opportunity to extract useful signals from unstructured inputs and external knowledge. However, few-shot LLM predictions are not ordinary fixed fitted predictions because prompt demonstrations sampled from accrued trial data create data-dependent randomness and can induce correlation across predictions. If used naively to drive allocation updates, AI-generated signals may weaken CARA's ability to improve power and participant welfare. We propose a CARA design that integrates few-shot LLM-predicted outcomes through resampling-based aggregation, calibration, and effective residual variances to guide both adaptive treatment allocation and downstream treatment effect estimation. Theoretical results and simulation studies show that, when calibrated few-shot LLM auxiliary signals reduce residual variation, coupling an LLM-enhanced estimator with an adaptive allocation rule can improve statistical efficiency and participant welfare.
PaperID: 2922, Poster
Authors: Ben Jenkins, Mihaela Cardei
Abstract: Test-time compute scaling, where language models "think longer" via extended chain-of-thought (CoT), is the dominant paradigm for improving reasoning. Prevailing intuition holds that extended reasoning converges toward an answer; we show the opposite. Across three reasoning-RL families (DeepSeek-R1, Qwen-QwQ, Llama-3.1-Nemotron) and three benchmarks (GSM8K, MATH, GPQA), step-to-step SAE feature turnover and entropy rise monotonically across the chain (Cohen's d = 0.61–0.68, p < 10^-3). Yet the final answer is already linearly decodable from the residual stream within the first ~ 20% of the chain. We call this pattern commit then explore: the answer becomes stably decodable early, and the bulk of test-time compute is then spent exploring around an already-decodable answer. Targeted ablation shows post-commit features have a 2–6× smaller teacher-forced log-probability effect (matched-mass control: 1.5–4.6×) and a 2–3× smaller free-regeneration answer-flip rate than pre-commit features. A terminal-reward RLVR credit-assignment account predicts this asymmetry, and the dynamic is absent in base models: Qwen-2.5 and Llama-3.1 bases fail our detector at every layer, and a held-out quantitative prediction derived from R1 generalises to Qwen-QwQ at Pearson r = 0.96. Leveraging the structure, we build a Reasoning Monitor that saves 26–34% of inference compute at \geq 92% accuracy retention, dominating four adaptive-compute baselines. Extended chain-of-thought, under this view, is not the computation of an answer; it is the exploration of its neighborhood.
Abstract: As generative video models become increasingly realistic, detecting AI-generated videos requires systems that offer both accuracy and interpretability. However, applying Multimodal Large Language Models (MLLMs) to video forensics is currently limited by outdated datasets, simplistic evaluation protocols, and a reliance on black-box classification. To address these issues, we introduce a comprehensive dataset, benchmark, and baseline model for video forgery detection. First, we present GenBuster-200K, a fair dataset of over 200,000 high-quality videos sourced from state-of-the-art generators, featuring diverse real-world scenarios. Second, we propose GenBuster-Bench, a diagnostic benchmark spanning three progressive tracks (In-Domain, Out-of-Domain, and In-the-Wild) to evaluate models across domain shifts and generational shifts. It also introduces an MLLM-as-a-Judge protocol to assess the quality of the generated forensic explanations. Finally, we develop BusterX, an MLLM baseline with RL training. Instead of direct binary classification, BusterX formulates detection as a visual reasoning task, where the generated reasoning chain serves as detector itself. Experimental results demonstrate that BusterX outperforms several leading MLLMs (e.g., Qwen3.5, Claude-Sonnet-4.6) in both detection accuracy and rationale quality.
Abstract: Recent image generation models produce photo-realistic results with rich details, but transferring their powerful priors to video restoration remains difficult due to temporal inconsistency from stochastic detail synthesis. Existing video restoration methods improve temporal coherence but often sacrifice visual fidelity, leaving a gap between image-level quality and video-level stability. We propose a generative video restoration method that bridges this gap by enabling temporally consistent detail synthesis from image generation priors. Our method combines noise energy rebalancing to suppress structure-disturbing low-frequency stochasticity, motion-aligned noise warping to propagate details along object motion, a self-supervised denoising encoder to mitigate encoder-induced temporal uncertainty, and decoupled temporal attention to model cross-frame dependencies. Together, these components preserve fidelity, enhance fine-grained realism, and maintain temporal consistency. Extensive experiments show that our method substantially outperforms existing video restoration approaches in visual quality, detail richness, and temporal stability.
PaperID: 2925, Poster
Abstract: We propose a unified spatial-density signal that scores token importance for \emphboth phases of inference, instantiated as a bipartite merge during prefill and a KV eviction during decode, with measurable gains in each phase. To verify the two phases independently, we partition benchmarks into two diagnostic roles by their scoring rule. \emphPrefill probes such as MMBench, POPE, and ScienceQA score only the first decoded token, so their accuracy reflects only the argmax of prefill logits, computed before any decode-phase KV eviction takes effect. \emphFull-pipeline probes such as TextVQA, DocVQA, and MM-Vet score the full decoded sequence and therefore reveal both phases. We expose a fundamental evaluation flaw: published claims of decode-phase token pruning validated only on prefill probes have measured prefill quality, not decode quality. Empirically, removing all visual KV at decode leaves prefill-probe accuracy essentially unchanged while collapsing TextVQA and MM-Vet by tens of points, and the insensitivity is a property of the scoring rule, not of the compression strength. Across seven model variants from the InternVL-3.5, Qwen3-VL, and LLaVA-1.5 families and the three full-pipeline-probe benchmarks, we deliver an honest three-axis Pareto over compute, memory, and accuracy. Density-based prefill compression is competitive on the accuracy frontier, scales linearly in the visual-token count, and is roughly two orders of magnitude faster in token selection at high resolution. On the decode side, density top-k KV eviction wins every TextVQA and MM-Vet cell against the strongest published baselines H2O, SnapKV, and StreamingLLM across three keep ratios. The advantage widens to 9.6 percentage points over H2O at 12.5% keep on TextVQA and holds 7 to 10 percentage points across the InternVL-3.5 series from 2B to 38B. Compression also scales gracefully with model size: every matched 8B-to-38B transition we measure reduces, rather than amplifies, the accuracy loss. We recommend MM-Vet or DocVQA as the minimum full-pipeline-probe test for any decode-phase claim. Our code and model will be released soon.
PaperID: 2926, Poster
Authors:
Junjie Gao, Siyuan Song, ZIXUAN ZHANG, SONG FEIYANG, Pan Yongzhou, Wen Nuan, Yaosheng Deng, Xue Tian, Mir FeroskhanAbstract: Flow-based vision-language-action (VLA) models provide expressive action generation for embodied control, but supervised fine-tuning (SFT) often confines them to narrow expert behaviors. Online reinforcement learning (RL) can improve beyond demonstrations through environment interaction, yet flow-based VLAs struggle with sparse long-horizon rewards and intractable action likelihoods. We propose \emphLiFT, a likelihood-free tree policy optimization framework for flow-based VLAs. LiFT first addresses sparse-reward credit assignment by creating sibling continuations from shared rollout histories and using their final outcomes to compute branch level relative advantages. LiFT further uses local action-flow instability as an expansion criterion, focusing rollout budget on ambiguous action-generation regions that are most likely to reveal meaningful outcome differences. Finally, LiFT applies branch-level advantages through a decision-level surrogate ratio on executed action chunks, matching policy updates to the granularity of tree-based credit assignment. This yields a likelihood-free policy update without attributing chunk-level credit to individual denoising steps. Evaluations on in-distribution and out-of-distribution benchmarks show that LiFT improves task success and robustness to distribution shifts.
PaperID: 2927, Poster
Abstract: In this paper, we tackle the critical failure modes of Physics-Informed Neural Networks (PINNs), such as spectral bias, which lead to poor convergence on complex PDEs. We identify two key shortcomings in existing curriculum learning methods for PINNs: unreliable knowledge transfer between stages and a reliance on manual, ad-hoc curriculum design. To overcome these limitations, we present Neural Operator-based Curriculum Learning (NOCL), a unified framework that leverages Neural Tangent Kernel (NTK) theory to automate curriculum generation and employs neural operators to enable robust, dynamic knowledge transfer across curriculum stages. By dynamically training neural operators and filtering data for PINN initialization, our approach ensures scalable and effective learning across progressively difficult tasks. Experiments verify that NOCL leads to marked gains in convergence and generalization relative to prior methods, resulting in substantially better performance on test suite.
PaperID: 2928, Poster
Abstract: Omni-modal caption models equipped with precise timestamp grounding are crucial for fine-grained video understanding and temporally controllable video generation. However, existing open-source models, and even proprietary systems, struggle to provide fine-grained timestamp grounding for short video captioning. To bridge this gap, we present , an advanced open-source omni-modal captioning framework designed for sub-second timestamp precision. During the Supervised Fine-Tuning (SFT) stage, we leverage Optical Character Recognition (OCR) and Automatic Speech Recognition (ASR) models to strictly calibrate visual text and spoken dialogue boundaries, offering highly deterministic temporal anchors to endow the model with sub-second timestamp grounding capabilities. Furthermore, we curate a specialized dataset for multi-reward Reinforcement Learning (RL) and propose Span-Specific Credit Assignment (SSCA). Unlike conventional global advantage normalization, which dilutes reward signals across weakly coupled multimodal descriptions, SSCA calculates advantages independently for different structural spans, effectively isolating penalties and rewards across semantic and temporal dimensions. Extensive experiments demonstrate that Deci-Omni-Captioner establishes state-of-the-art (SOTA) performance among open-source models on multiple semantic comprehension benchmarks (e.g., AVUT, UGC, VCapsBench). On our challenging Deci-Timestamp Bench, it significantly surpasses leading proprietary models, including the Gemini 2.5 and 3.0 series.
PaperID: 2929, Poster
Abstract: Latent stochastic differential equations (SDEs) model continuous-time, irregularly-sampled time series. Training these models typically relies on variational inference (VI) methods that suffer from sequential sampling bottlenecks and compounding integration errors. In this paper, we propose Parallel-in-Time Variational Inference (PiTVI), a framework that parameterizes the variational posterior as a non-Markovian process mapping the driving noise history directly to the local state increments. By leveraging modern sequence modeling architectures such as Transformer and Mamba, the training and inference of the variational posterior become fully parallelizable, reducing the span complexity from \mathcalO(L) to \mathcalO(\log L). Beyond computational scalability, we demonstrate that this formulation shifts global error propagation from multiplicative compounding to additive scaling. Furthermore, the non-Markovian formulation natively accommodates fractional driving noise, which models physical memory effects and provides a statistical relaxation mechanism when fitting smooth dynamics. We validate these properties across empirical time complexity scaling tests, long-horizon predictions on non-linear systems, and high-dimensional sequence modeling.
PaperID: 2930, Poster
Abstract: Diffusion models achieve strong performance across image and audio synthesis, but their inference remains computationally expensive, requiring tens to hundreds of sequential evaluations of a large denoising network. We observe that the effective rank of the denoiser output varies sharply across timesteps and architectures, motivating step- and block-wise adaptation of computation. We propose Step-Adaptive Rank Adjustment (SARA), a training-free framework that treats the SVD rank of pretrained linear layers as an inference-time control variable, adapted per step and block by SARA-DP and per step by SARA-Online. SARA introduces the Low-Rank Approximation Error (LRAE) — the discrepancy a rank-r truncation induces against the full-rank network — and instantiates it in two algorithms: SARA-DP measures LRAE on a few warmup prompts and solves the globally optimal rank assignment under a MACs budget by dynamic programming; SARA-Online computes a closed-form, spectrum-based LRAE from the model output at inference time, requiring no calibration and no trajectory lookahead. SARA applies uniformly to UNet (SDXL), MMDiT (SD3), DiT (Stable Audio Open), and hybrid MMDiT/DiT (TangoFlux) backbones, spanning ϵ, v, and rectified-flow parameterizations. SARA combines with existing solver- and cache-axis acceleration to yield compounded wall-clock speedups in our experiments. Together, these results introduce rank as a new axis of diffusion inference acceleration.
PaperID: 2931, Poster
Abstract: Despite the strong performance of diffusion models in molecular generation, the sampling process itself remains largely understudied, with most existing methods adopt off-the-shelf sampling strategies developed for natural images. Yet unlike natural images, molecular distributions are governed by physical laws and are sharply concentrated, making them particularly challenging for standard samplers, which frequently produce invalid or unphysical structures. In this work, we revisit diffusion sampling from a unified stochastic differential equation (SDE) perspective and introduce a general framework parameterized by two interpretable controls: \emphstochasticity and \emphtemperature. Our theoretical analysis reveals that stochasticity accelerates the decay of sampling error, while temperature directly controls distribution sharpness, enabling concentration on physically plausible configurations. Building on these insights, we develop \textscTReaSSure (Temperature-Regulated Stochastic Sampling), a lightweight, training-free strategy that combines stochasticity scheduling with temperature annealing to better capture the sharply concentrated distributions of molecular data. Extensive experiments on small molecule generation, protein structure prediction, and protein design demonstrate that \textscTReaSSure consistently improves generation quality and produces more physically valid structures. Empirical analyses further corroborate our theoretical findings. The method harnesses pretrained models without retraining, underscoring its generality and practical value across diverse molecular domains.
Abstract: Quantum autoencoders (QAEs) are learning architectures that compress quantum data into a low-dimensional latent state while preserving the information needed for reconstruction. We study blind single-copy compression of quantum states through a k-qubit bottleneck and investigate the minimal circuit width required to attain the information-theoretic optimum under average infidelity. Between the conventional architecture, which is narrow but nonuniversal, and fully general completely positive and trace preserving (CPTP) realizations, which are universal but overparameterized, we identify a balanced regime. We prove that for every distribution of pure n-qubit states, there exists a QAE with exactly k encoder ancillas and n decoder ancillas that achieves the optimal fidelity over all CPTP encoder–decoder pairs. The encoder-side statement is sharp in that we construct source families for which every optimal scheme necessarily uses at least k encoder ancillas, thereby determining the universal encoder threshold exactly. On the decoder side, we show that isometric decoders are exactly optimal for several analytically tractable source families, but we also exhibit an explicit counterexample demonstrating that decoder isometry is not universally sufficient. Nevertheless, numerical experiments indicate that the performance gap is practically negligible.
Abstract: System prompt optimization improves agent behavior without modifying the underlying model, yielding human-readable, model-agnostic instructions. Existing methods build a prompt agent that refines task agents' system prompts, yet leave the prompt agent's own system prompt hand-engineered and fixed. We propose Self-Evolving Prompt Optimization (SePO), which treats the prompt agent's own system prompt as an optimization target alongside task agents' system prompts. SePO adopts a self-referential design. A single prompt agent improves both task agents' system prompts and its own under an open-ended evolutionary search that maintains an archive of candidate prompts as stepping stones. Training proceeds in two stages: pre-training evolves the prompt agent on a multi-task pool, and fine-tuning then applies it to a target task. Across five benchmarks spanning math (AIME'25), abstract reasoning (ARC-AGI-1), graduate-level science (GPQA), code generation (MBPP), and logic puzzles (Sudoku), SePO consistently outperforms Manual-CoT, TextGrad, and MetaSPO, improving the average accuracy by 4.49 points compared to Manual-CoT. The prompt optimization skill from pre-training also generalizes to tasks beyond the pre-training mixture, rather than memorizing per-task prompts.
PaperID: 2934, Poster
Abstract: Gigapixel object detection in High-Resolution Wide (HRW) images faces extreme spatial sparsity where targets often occupy less than 5% of the image area while computation is wasted on vast uninformative regions. Existing token selection methods require either complex evolutionary search or training additional learnable modules. We present AWP (Activation-based Window Pruning), which discovers that the mean of Feed-Forward Network (FFN) activations naturally encodes window importance, enabling parameter-free selection without additional modules. Leveraging this intrinsic signal, AWP performs window pruning with remarkable effectiveness. On the PANDA gigapixel dataset, AWP outperforms the previous state-of-the-art by +1.6% AP₅₀ with 26.2% FLOPs reduction, particularly excelling on small objects (+5.1% APₛ). When applied to Swin Transformer, AWP demonstrates strong generalization, retaining 97.4% of baseline performance with 50% windows pruned. Remarkably, AWP enables zero-shot application by directly replacing learned selection modules, improving baseline by +0.6% AP₅₀ without retraining. Our work reveals that window-based vision transformers inherently encode spatial importance through FFN activations, offering a simple yet powerful alternative to existing selection mechanisms, especially for High-Resolution Wide images.
PaperID: 2935, Poster
Abstract: Data heterogeneity and low client participation are two key challenges in federated learning (FL). Client-reshuffling-based FL methods were recently introduced to improve participation efficiency by visiting each client once per meta-epoch; however, the resulting without-replacement sampling induces inter-round dependence and conditional bias. As a consequence, existing client-reshuffling methods can still suffer from the data heterogeneity challenge due to this dependence. To bridge this gap, we propose FedCDR, a client-reshuffling FL algorithm built on Douglas–Rachford splitting. FedCDR supports inexact local proximal updates via iterative solvers, enabling a practical communication–computation trade-off. For smooth nonconvex objectives, FedCDR with inexact local solvers attains a state-of-the-art O(\epsilon^-1) communication complexity to reach an \epsilon-approximate stationary point (i.e., \mathbbE\|\nabla f(\tildex)\|^2 \le \epsilon), with the leading constant that is independent of data heterogeneity (i.e., it does not scale with common measures of heterogeneity). Technically, our analysis operates at the meta-epoch level: we control the deviation between reshuffled and full-client updates, construct a tailored potential function with provable descent, and sum over each meta-epoch to eliminate reshuffling-induced dependence. Experiments on synthetic tasks and benchmark datasets under heterogeneous partitions, including a 10,000-client setting, demonstrate consistent improvements over strong baselines and their client reshuffling variants.
Abstract: Membership inference attacks (MIAs) have emerged as the standard tool for evaluating the privacy risks of AI models. However, state-of-the-art attacks require training numerous, often computationally expensive, reference models, limiting their practicality. We present a novel approach for estimating model-level vulnerability to the Likelihood Ratio Attack (LiRA), the strongest available attack, directly from the train and test loss distributions of the target model and without training any reference models. We show that LiRA's per-sample signal decomposes into a variance-ratio term and a residual mean-shift term, with the relative contribution of each determined by how much training collapses model uncertainty at the trained sample. This places models on a continuum, with different regimes calling for different reference-free loss-based statistics as proxies for LiRA TPR. The shapes of the loss distributions themselves indicate which proxy applies. We instantiate the framework with two natural proxies. At the heavy-tailed end, the LOSS attack TNR predicts LiRA TPR@FPR=10^-3 with RMSE 0.03 across 9 image classification architectures and 4 datasets, outperforming low-cost reference-model attacks such as RMIA. At the symmetric end, the LOSS attack AUC predicts LiRA TPR with RMSE 0.01 across five GPT-2 sizes from 10M to 1B parameters. We also show these proxies to outperform both low-cost (few reference models) attacks such as RMIA and other measures of distribution difference.
PaperID: 2937, Poster
Abstract: Biological multimodal large language models (MLLMs) have emerged as powerful foundation models for scientific discovery. However, existing models are specialized to a single modality, limiting their ability to solve inherently cross-modal scientific problems. While model merging is an efficient method to combine the different modalities into a unified MLLM, existing methods rely on input-agnostic parameter space heuristics that fail to faithfully capture modality specialization. To overcome this limitation, we propose the Embedding-Signal-based MLLM Merging (ES-Merging), a framework that estimates merging coefficients from embedding space signals, moving the merging paradigm from the parameter signals to the embedding signals. ES-Merging exploits coarse-grained and fine-grained signals from embedding space to estimate the layer-wise and element-wise merging coefficients, respectively, which are jointly combined for complementary coefficient estimation. Through extensive experiments, we demonstrate that ES-Merging outperforms existing merging methods not only on the cross-modal reasoning but also on the single-modal knowledge preservation, establishing that embedding space signals provide a principled and effective foundation for MLLM merging.
PaperID: 2938, Poster
Authors:
Abdulaziz Alshamsi, Abdulla Alghfeli, Chenghua Lin, Hujun YinAbstract: Positional encoding in Vision Transformers presents a recurring trade-off: spatial relative-position tables (Swin) carry thousands of parameters per block and require ad-hoc interpolation to transfer across resolutions, while no-bias designs leave the attention logits geometry-agnostic. We propose a relative positional bias parameterized directly in the frequency domain. This parameterization yields three benefits over the standard spatial-table position encoding: (i) it expresses the same translation-equivariant bias with 3.25× fewer parameters than Swin's spatial table at the same head count; (ii) it admits a closed-form, structurally exact transfer to a different patch grid, in contrast to bilinear interpolation of spatial tables which smooths the learned structure; and (iii) it shares the same parameterization style as a frequency-domain spectral mixer, enabling co-design of positional encoding and token mixing in a common substrate. We demonstrate the bias inside our Dual-Domain Residual (DDR) framework, in which a spectral mixer and a gated attention head act as parallel residuals within each block. Head-to-head, the frequency-domain bias recovers +1.98 pp over Swin RelPos and +2.74 pp over no positional bias; removing the entire attention branch (which carries the bias) drops accuracy by 5.35 pp, while removing the spectral branch drops it by 2.22 pp. On ImageNet-1K from scratch, DDR outperforms SpectFormer at every comparable parameter budget, by +2.0 pp at Tiny (DDR-Ti-SF, 10.7M, 78.9%) and +0.4 pp at Base (DDR-B-Deep, 63.1M, 82.5%); at sub-parity parameters, DDR-Ti+ (8.9M, 77.9%) exceeds SpectFormer-Ti (9.0M, 76.9%) by +1.0 pp. The learned per-layer gates concentrate spectral processing in early blocks and attention in late blocks, qualitatively matching the schedules that are hand-coded in the previous hybrid models.
Abstract: 4D mesh generation has recently emerged as a powerful paradigm for recovering dynamic 3D structure from videos, but existing methods remain slow, computationally expensive, and difficult to scale to longer sequences. We introduce a training-free approach that accelerates 4D mesh generation while improving temporal correspondence quality. Our key observation is that temporal correspondences emerge inside a 4D backbone long before its generated meshes become visually accurate. We exploit this with a general framework we call Spatio-Temporal Attention Chain which propagates information across space and time. Starting from vertices on an anchor mesh, the chain maps vertices to latent tokens. It then follows temporal correspondences in latent space, and recovers frame-specific vertices through latent-to-vertex attention. This design avoids expensive explicit matching while preserving anchor mesh details and thereby improving dynamic mesh geometry and temporal consistency. Compared to state-of-the-art, our method generates a 4D mesh in 9 seconds, achieving a 13× speedup while producing higher-quality results. Moreover, our approach scales to videos up to 16× longer without degrading mesh quality. Beyond generation, the improved correspondences enable competitive zero-shot performance on two downstream tasks: 2D object tracking and 4D tracking. We further show that our framework enables reliable camera estimation, a capability not supported by prior 4D mesh generation methods.
Abstract: Speculative decoding (SD) has become a standard technique for accelerating LLM inference without sacrificing output quality. Recent advances in speculative decoding have shifted from sequential chain-based drafting to tree-structured generation, where the draft model constructs a tree of candidate tokens to explore multiple possible drafts in parallel. However, existing tree-based SD methods typically build a draft tree, which fails to adapt to the varying difficulty of tokens and contexts. As a result, the draft model cannot dynamically adjust the tree structure to early stop on difficult tokens and extend generation for simple ones. To address these challenges, we introduce adaptive tree expansion framework that can be plugged into existing tree-based methods. Unlike static methods, TALON constructs the draft tree iteratively until a fixed token budget is met, using a hybrid expansion strategy that adaptively allocates the node budget to each layer of the draft tree. This framework naturally shapes the draft tree into a form for uncertain branches, effectively optimizing the trade-off between exploration width and generation depth under a given budget. Extensive experiments across 5 models and 6 datasets demonstrate that TALON consistently outperforms state-of-the-art EAGLE-3, achieving up to 5.16× end-to-end speedup over auto-regressive decoding.
Authors: Louis Meyer, Wenkai Xu
Abstract: Multivariate conformal prediction requires nonconformity scores that compress residual vectors into scalars while preserving certain implicit geometric structure of the residual distribution. We introduce a Multivariate Kernel Score (MKS) that produces prediction regions that explicitly adapt to this geometry. We show that the proposed score resembles the Gaussian process posterior variance, unifying Bayesian uncertainty quantification with the coverage guarantees of frequentist-type. Moreover, the MKS can be decomposed into an anisotropic Maximum Mean Discrepancy (MMD) that interpolates between kernel density estimation and covariance-weighted distance. We prove finite-sample coverage guarantees and establish convergence rates that depend on the effective rank of the kernel-based covariance operator rather than the ambient dimension, enabling dimension-free adaptation. On regression tasks, the MKS reduces the volume of prediction regions significantly, compared to ellipsoidal baselines while maintaining nominal coverage, with larger gains at higher dimensions and tighter coverage levels.
PaperID: 2942, Poster
Abstract: Agent applications are evolving into workflows composed of large language model (LLM) calls and tool invocations, termed modules. Existing distributed agent systems typically adopt a disaggregated architecture, deploying different modules on separate devices. While this enables pipeline parallelism, imbalanced module latencies introduce severe pipeline bubbles that limit throughput. A natural alternative is an aggregated architecture that co-deploys all modules across all devices, eliminating bubbles by keeping every device continuously active. However, this introduces request blocking: short requests are forced to progress in lockstep with longer ones in the same batch, and newly arriving requests must wait until the entire ongoing batch completes the full workflow. To address this, we propose AgentWeave, an efficient distributed serving system for agent applications built on the aggregated architecture. Its core mechanism, flow decomposition, decomposes long agent requests into fine-grained execution units while preserving agent semantics, allowing short and newly arriving requests to proceed without being stalled by long-running ones. Complementing this, a flow-adaptive scheduler dynamically balances the throughput gains of decomposition against its scheduling overhead. Experiments across four representative agent applications demonstrate that AgentWeave improves throughput by 1.81×, reduces average latency by 2.45×, and reduces P90 latency by 2.04× over state-of-the-art baselines.
PaperID: 2943, Poster
Authors: Wentao Zhang, Yifan Zhu, Yutong Zhang
Abstract: We study bandit convex optimization (BCO) when feedback is adversarially delayed and the comparator sequence is also adversarially non-stationary. This setting arises in online advertising, recommender systems, and adaptive dosing, but existing algorithms either absorb one of the two axes into a black-box global penalty or pay a worst-case cost that ignores the local interaction between delay and drift. We propose Oracle-DARS-DBCO, a delay-aware blocking algorithm whose block length K and one-point smoothing radius \delta satisfy the single identity K\delta^2\asymp n^2. This identity cancels the multiplicative coupling between delay deviation and one-point estimator variance, so the regret of each segment is paid only on that segment. We prove matching upper and lower bounds on the expected dynamic regret: \widetildeO\bigl(\sqrtn\,m^1/4T^3/4+\sqrtd_\max\,mT\bigr) for convex losses and \widetildeO\bigl(n^2/3m^1/3T^2/3+d_\max\,m\log(eT/m)\bigr) for \alpha-strongly convex losses, where m=S_T+1 is the number of stationary segments and d_\max is the worst-case delay. The matching Rademacher-sign lower bounds \Omega\bigl(\sqrtd(S_T+1)\,T\bigr) and \Omega\bigl(d(S_T+1)\bigr) show that the delay price is segment-local. On fourteen scaling experiments, the fitted log--log exponents agree with the predicted ones within 0.04, and Oracle-DARS-DBCO reduces regret by up to 9.17× relative to the strongest non-restarting baseline on a piecewise-stationary benchmark.
PaperID: 2944, Poster
Abstract: Sparse Autoencoders (SAEs) are widely used for mechanistic interpretability in large language models, yet their standard formulation assigns each latent feature a single decoder direction, implicitly constraining features to be one-dimensional. We identify a fundamental geometric mismatch between this assumption and the structure of many model features, and show that it can provably induce feature splitting. In particular, for any feature with intrinsic dimension d_i \ge 2, achieving reconstruction error \varepsilon with a standard SAE requires \Omega \left((\frac1\varepsilon)^d_i-1\right) distinct decoder directions. As a result, the SAE is forced to represent a single coherent high-dimensional feature using many nearly collinear one-dimensional latents, leading to spurious feature multiplicity and a failure to preserve the intrinsic structure required for mechanistic interpretability. Motivated by this lower bound, we introduce Subspace-Aware Sparse Autoencoders (SASA), which replace single-vector decoders with learned decoder subspaces. SASA enforces block sparsity through Top-s group gating and adapts each group's effective dimensionality using a trace-norm regularizer. We prove a complementary upper bound: when the block size r \ge d_i, a single group can represent the entire feature slice, alleviating feature splitting while also improving sample complexity. Finally, we empirically validate that SASA improves feature splitting, monosemanticity, and interpretability while also enabling more efficient training.
PaperID: 2945, Poster
Abstract: Previous work has shown that the optimal classifier from a hypothesis class subject to exact fairness constraints may not be robust; that is, its accuracy can change drastically under small shifts to the underlying data distribution. We ask the following questions: Given a hypothesis class \mathcalH, is the accuracy of the optimal fair classifier from \mathcalH robust to malicious distribution shifts? And is this property efficiently testable? We consider \mathcalH that can be represented by a convex polytope — the natural setting for randomized ensembles of deterministic classifiers, including those produced by boosting. We provide a complete geometric characterization of when \mathcalH is robust, and an efficient linear programming test to audit the robustness of a given \mathcalH. Surprisingly, we demonstrate that if we relax exact fairness constraints and only require approximate fairness, every \mathcalH is robust. Our results hold for a broad class of linear fairness criteria (e.g., demographic parity, equal opportunity, predictive equality) and linear loss functions (e.g., 0/1 loss, weighted loss).
Abstract: The quadratic cost of causal self-attention severely bottlenecks long-context transformer inference. While numerous post hoc linearization pipelines exist, it is difficult to identify which components preserve model quality. This work isolates the effect of state update design in a strict frozen-backbone regime. We show that softmax relies on key-dependent, rank-1 orthogonal projections, elucidating why delta-style networks outperform purely gated accumulation. We identify a potential source of approximation errors and introduce structural interventions, specifically sink tokens, short convolutions, and fixed-budget cache routing, which reduces the remaining gap. We scale this linearization approach across LLaMA and Qwen models up to 32B parameters, outperforming prior post hoc baselines on MMLU and matching the long-context retrieval of complex adaptive-caching frameworks.
PaperID: 2947, Poster
Authors:
ruinan Jin, Difei Cheng, Ling Chen, Jun Luo, Hao Zhou, Youzhi ZhangAbstract: It has been widely observed that Adam can remain stable even when the objective function deviates significantly from global smoothness. However, under the generalized smoothness framework, existing theoretical analyses typically rely on strong tail assumptions on stochastic gradients, such as almost-sure boundedness or sub-Gaussianity. Whether one can establish the convergence of Adam on generalized smooth objectives under only second moment information on the stochastic gradients, without imposing such strong concentration assumptions, was explicitly identified as an important open direction by \citetli2023convex. This paper gives an affirmative answer to this question under fairly general conditions. Specifically, we prove that such strong tail assumptions are not necessary. We show that the key mechanism by which Adam remains stable and achieves convergence under the (L p) generalized smoothness condition is the self-normalization effect induced by its adaptive coordinate-wise scaling. Based on this mechanism, we prove that even under a very general stochastic-gradient condition, namely a generalized second moment ABC condition that provides only second moment information, the stochastic trajectory of Adam remains in a locally well-behaved smoothness region with stretched-exponential tail decay. As a consequence, we establish high-probability convergence rate guarantees over the full range (p<2), with a confidence dependence of order (\delta^-1/2), while the required stepsize calibration depends on (\delta) only through polylogarithmic factors in (\polylog(1/\delta)). Furthermore, we construct a hard instance proving that, under only second-moment information on the stochastic gradients, this (\delta^-1/2)-type confidence dependence is sharp. Finally, in the more favorable regime (p<1), we combine the above trajectory control with polynomial-growth estimates on rare events to further obtain convergence rate guarantees in expectation.
PaperID: 2948, Poster
Abstract: Recent progress in multimodal, high-dimensional learning has enabled foundation models to process heterogeneous, large-scale data. However, at test time, acquiring all features or modalities can be prohibitively costly and often redundant. Sequentially selecting informative modalities is therefore critical, yet challenging when the downstream task or prediction target is unknown. To this end, we introduce ECHO-k, a task-agnostic and self-supervised learning principle for modality acquisition: we use a deep model’s internal pretrained representations (e.g., from a foundation model) as proxy targets that summarize cross-modal information. We provide theoretical guarantees in a linear model setting that motivate a reinforcement learning (RL) policy for sequential modality selection. Across benchmarks, the resulting acquisition strategies transfer reliably and achieve state-of-the-art performance on downstream tasks that are entirely unseen during selection. Our method provides a principled route to cost-aware test-time deployment, with implications for any multimodal system where measurements are expensive or time-constrained, and downstream tasks unknown a priori.
PaperID: 2949, Poster
Abstract: Online video large language models (VLLMs) have demonstrated remarkable capabilities in real-time streaming video understanding and proactive human-AI interaction. However, existing methods often follow a static perception paradigm, processing video streams with rigid temporal units while overlooking the inherent temporal non-uniformity of streaming videos. This oversight results in flickering responses and fragmented context. In this paper, We propose StreamMind, a dynamic streaming cognition framework that organizes both online reasoning and memory around temporally coherent Group-of-Pictures (GoP) segments. Specifically, StreamMind introduces two key components. First, StreamMind introduces a Dynamic Thinking (DynThink) module, which accumulates evidence within each GoP and triggers reasoning only at GoP boundaries, thereby reducing premature decisions and improving response stability. Second, to support long-horizon streaming understanding, we propose a Brain-Eye Synergy Memory (BESM) module, which preserves key-frame visual tokens as high-fidelity anchors while compressing the remaining frames into compact thinking tokens. This design maintains fine-grained visual evidence together with reasoning continuity, while reducing redundant storage. Extensive experiments on streaming and offline video benchmarks demonstrate that StreamMind achieves state-of-the-art online video understanding performance, reaching 63.4% on OVO-Bench and 80.2% on StreamingBench, while maintaining strong generalization to offline long-form video reasoning.
Abstract: The concept of ranking aggregation plays a central role in preference analysis, and numerous algorithms for calculating median rankings, often originating in social choice theory, have been documented in the literature, offering theoretical guarantees in a centralized setting, i.e., when all the ranking data to be aggregated can be brought together in a single computing unit. For many technologies (e.g. peer-to-peer networks, IoT, multi-agent systems), extending the ability to calculate consensus rankings with guarantees of convergence and resilience to potential contamination in a decentralized setting, when preference data is initially distributed across a communicating network, remains a major methodological challenge. Indeed, in recent years, the literature on decentralized computation has mainly focused on computing or optimizing statistics such as arithmetic means using gossip algorithms. The purpose of this article is precisely to study how to achieve reliable and resilient consensus on collective rankings in a decentralized setting, thereby raising new questions, robustness to corrupted nodes, and scalability through reduced communication costs in particular. The approach proposed and analyzed here relies on robustness guarantees derived from random gossip communication, allowing autonomous agents to compute global ranking consensus using only local interactions, without coordination or central authority.
PaperID: 2951, Poster
Authors:
Meng Lu, Ligeng Zhu, Olivia Xiao, Yuchen Zhuang, Zihan Wang, Kuncheng Wu, Bangya Liu, Yu Wang, Charles Fleming, Wenqi Shi, Xuan WangAbstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs), but it typically assumes a static training environment. As the actor improves, fixed tasks drift out of its learning frontier: many become trivial, others remain unsolvable; and the learning signal collapses. We argue that \emphVLM post-training should evolve the visual environment alongside the actor, not just the actor itself. We propose VICO, a co-evolutionary framework in which an \emphactor and an \emphEnvironment-as-Rewriter (EnvRewriter) are trained jointly: the EnvRewriter edits verifiable image-side structures, such as scene graphs, chart tables, or protected region masks, and re-renders them to produce label-valid training samples whose difficulty is calibrated to the actor's current ability through a pass-rate-based reward. This loop continuously realigns task difficulty with actor capability without any additional human annotation. Across nine multimodal benchmarks spanning mathematical reasoning and visually grounded understanding, VICO-8B improves over its base model by up to +5.0% on out-of-domain tasks, surpasses the strongest self-evolution and text-editing co-evolution baselines by +4.3% and +8.4% respectively, and stays comparable to chart-specialized RLVR methods using 16–160× fewer labeled samples. By shifting from human-labeled supervision to image-editing co-evolution, VICO offers a scalable path beyond static-corpus RLVR for visual reasoning.
PaperID: 2952, Poster
Abstract: Protein binder design has largely optimized for affinity alone, leaving conformational selectivity unaddressed: for allosteric targets such as kinases, nuclear receptors, and GPCRs, a binder that engages both active and inactive states provides no functional specificity regardless of how tightly it binds. We introduce AlloGen, a modular framework that decouples backbone generation from a learned state-selectivity scorer Q_\theta, an SE(3)-invariant interface graph transformer trained via a two-phase curriculum that first grounds interface geometry before imposing conformational discrimination. Because Q_\theta is fully differentiable and generator-agnostic, it integrates with any backbone generator as a passive reranker or an active gradient-based guide without retraining. Trained on 65 targets spanning 15 protein families, Q_\theta generalizes to held-out out-of-distribution targets where energy-based baselines fail entirely, and all 15 evaluated generator--guidance combinations achieve positive conformational selectivity averaged over the held-out targets, with the best reaching \barS=+0.677. Our anonymous code repository can be found at https://anonymous.4open.science/r/AlloGen_NeurIPS-04CB.
Authors: Tomáš Holeček, Viliam Lisy
Abstract: Model-based reinforcement learning (MBRL) has achieved remarkable sample efficiency in single-agent domains, yet its extension to competitive imperfect information games (IIGs) remains underexplored. In multi-agent settings, opponent-induced non-stationarity complicates the learning process, and decentralized model learning faces severe identifiability barriers. To solve this, we propose NashDreamer, a principled MBRL framework for two-player zero-sum IIGs. NashDreamer introduces a centralized Multi-Agent Recurrent State-Space Model (MARSSM) that decouples environment dynamics from the effect of players' strategies on their individual observations. For policy optimization, NashDreamer uses Regularized Nash Dynamics (RNaD) within the latent imagination, providing theoretical convergence guarantees toward Nash equilibria. Empirical evaluations across Imperfect Information Goofspiel, Leduc Hold'em and Battleship demonstrate that NashDreamer achieves vastly superior sample efficiency compared to model-free baselines, including a 71.0% head-to-head win rate against model-free RNaD in stochastic 13-card Imperfect Information Goofspiel. Finally, we theoretically analyze the architecture's optimization landscape, formally identifying the vulnerability of the Dreamer family of algorithms to posterior collapse in highly stochastic environments, and we highlight the open challenge of mitigating latent non-stationarity without enabling degenerate solutions.
Abstract: Rubrics have been extensively utilized for evaluating unverifiable, open-ended tasks, with recent research incorporating them into reward systems for reinforcement learning. However, existing frameworks typically treat rubrics only as external evaluator disjointed from the policy's primary reasoning trace. Such design confines rubrics to post-hoc measurement, leaving them unable to actively guide the model's generation process. In this work, we introduce Think-with-Rubrics, a novel paradigm for instruction following tasks. Think-with-Rubrics integrates rubric generation into the reasoning context, transforming the rubric from an independent artifact into an internal guidance of LLM's generation. During training, LLM sequentially generates a rubric followed by a response, while a trained rubric verifier provides joint supervision by evaluating the consistency between the answer and the self-generated / golden rubrics. Experiments across multiple benchmarks demonstrate that Think-with-Rubrics consistently outperforms the Rubric-as-Reward baseline supervised by golden rubrics by an average of 3.87 points. We have also discussed the mechanism by which Think-with-Rubrics enhances model performance. Experimental results demonstrate that supervision from golden rubrics and self-generated rubrics enhances the performance of Think-with-Rubrics by improving the quality of self-generated rubrics and increasing the internal consistency of responses respectively.
PaperID: 2955, Poster
Abstract: Text-to-video models built on flow matching generate compelling videos for common prompts but degrade systematically under motion-intensive scenarios, producing artifacts such as object fragmentation, geometric deformation, and temporal flickering. Existing training-free approaches operate at the attention or guidance-scale level, neither of which directly addresses the velocity field that governs latent motion evolution. We observe that these motion artifacts stem from local misalignment between the predicted velocity and the latent's frame-axis temporal structure, a signal that can be diagnosed from quantities the sampler already computes. Based on this, we propose PhysFlow, a training-free velocity regularization method. PhysFlow constructs a per-frame latent flow residual to locate motion-inconsistent regions, steers the velocity to restore frame-axis alignment, and bounds the correction via residual-aware masking and safe clipping. We also introduce MotionStress-100, a five-category benchmark with a VLM-based protocol that isolates motion-intensive failures. On Wan2.1, PhysFlow improves the average MotionStress score over the baseline with no additional inference cost (1.00× cost), outperforming CFG-Zero (-10.0%, 0.98× cost) and FlowMo (+0.7%, 2.12× cost) by at least 4.7%, while improving VBench motion smoothness without loss of visual fidelity.
PaperID: 2956, Poster
Abstract: As language models are increasingly deployed in long-context and streaming workflows, deciding when to stop context processing is central to inference efficiency. While recent research has shown that a subset of attention heads encodes sufficiency signals, these markers are typically detected via static, task-agnostic global features. In this paper, we demonstrate that stop signals are inherently query-relevant and introduce \textscQueryStop, an attention-aggregation method that systematically identifies query-focused attention heads whose hidden features effectively align with the input query. Driven primarily by query-focused heads, this tailored ensemble yields high-fidelity representations that significantly improve sufficiency prediction. Across four benchmarks on eleven datasets, our method improves sufficiency stop prediction by 3.52% F1 on average. Under comparable token reduction, it raises Sufficiency Reach Rate (SRR), our proposed stopping-adequacy metric, from 87.12% to 94.73% and yields a 7.09% average gain in downstream task accuracy over the prior baseline. Collectively, these results show that \textscQueryStop effectively eliminates unnecessary context processing while simultaneously improving downstream answer quality.
PaperID: 2957, Poster
Authors: Ahmad Arrabi, Xiaohan Zhang, Xingyu Li, Safwan Wshah
Abstract: Visual data are naturally expressed through multiple complementary modalities (e.g., images and segmentation masks), each capturing distinct aspects of the same underlying structure. Most existing multimodal methods adopt a top-down paradigm, learning a shared latent variable through joint training from scratch, which limits scalability. We propose Split and Bridge (SnB), a bottom-up framework for multimodal generation that couples pretrained unimodal diffusion models at sampling time. Rather than learning a unified latent space, SnB generates coherent multimodal samples by following a shared diffusion trajectory up to a split point, after which modality-specific processes are guided via a bridge module. This design enables plug-and-play reuse of powerful pretrained models without joint training, making it scalable and easily adaptable. Across standard benchmarks (PolyMNIST, CelebAMask-HQ) and more complex settings (PIE-Bench++, ROCO), SnB consistently achieves a strong balance between sample quality and coherence, as we report a 33% increase in the coherence score on CelebAMask-HQ over prior top-down approaches. We show that SnB naturally extends to downstream zero-shot tasks such as interleaved co-generation, image editing, and style transfer. Our results demonstrate that bottom-up generation offers a practical and scalable alternative to top-down multimodal models.
PaperID: 2958, Poster
Authors: Karl Li
Abstract: Online forecasters sometimes know a regime boundary before any post-boundary labels arrive. Standard adaptive conformal methods still react through later coverage errors, so they can pay a detection-delay cost even when the boundary is public. We study conformal prediction with an announced partition of the stream and separate two post-break audits. The reset empirical quantile in segmented split-CP attains the minimax per-time probability-gap rate \Theta(M^-1/2) after M post-reset scores, with a matching Le Cam lower bound. A pure bounded constant-step announced-break ACI recursion attains an O(M^-1) time-averaged empirical-frequency gap by telescoping, but this is a realized-frequency statement, not a guarantee for the next interval. The practical AB-ACI implementation used in experiments adds burn-in and projection, so its rate claim includes explicit correction terms and is not unconditional. Synthetic experiments recover both slopes and show short-window gains when the break is large. Two real-data applications are null under leakage-free protocols. ERCOT RTC+B gives 0.863 post-reform coverage for AB-ACI versus 0.880 for ACI, and FOMC yield breaks give 0.892 versus 0.900. The message is diagnostic: use the reset empirical quantile for per-prediction reliability, use AB-ACI only for long-run frequency audits, and expect gains only when the realized score shift is large enough to make detection delay costly.
PaperID: 2959, Poster
Abstract: Offline goal-conditioned reinforcement learning remains challenging in stochastic long-horizon settings, where compounding value estimation errors hinder reliable goal reaching. While prior methods have sought to address this challenge through various approaches, a key challenge remains path mirage, where agents overcommit to spuriously successful trajectories in offline datasets induced by stochasticity, often leading to failure on difficult tasks. To address this issue, we propose Branch-aware Graph Planning (BGP), which captures stochastic branching structures by identifying high-variability points as graph nodes and designing goal-conditioned edges with segmented learning to avoid unreliable branches and reduce unnecessary stochasticity along paths. As a result, BGP enables reliable goal reaching even in highly stochastic long-horizon environments, and experiments on diverse OGBench tasks show that it substantially outperforms prior state-of-the-art offline goal-conditioned RL methods.
PaperID: 2960, Poster
Abstract: Graph contrastive learning constructs augmented views and aligns corresponding instances across views to learn view-invariant representations. However, existing methods usually enforce cross-view consistency within a single representation space, lacking an explicit characterization of view-specific information. This may suppress task-relevant complementary information, reduce embedding diversity, and induce dimensional collapse. To this end, we propose Manifold-Guided Representation Disentanglement (MGRD), a plug-and-play framework for boosting graph contrastive learning. MGRD decomposes graph representations into shared and complementary subspaces. The shared subspace captures view-invariant semantics while preserving local graph structure through a graph-anchored manifold prior, whereas the complementary subspace captures residual view-specific information beyond the shared representation and is regularized by asymmetric consistency and subspace decorrelation. By jointly exploiting shared and complementary representations, MGRD balances cross-view invariance and representation diversity. Experiments on node-level and graph-level benchmarks show that MGRD consistently improves multiple GCL backbones, while theoretical and empirical analyses demonstrate its effectiveness in mitigating dimensional collapse.
PaperID: 2961, Poster
Abstract: Designing antibodies that recognize a target antigen while avoiding non-cognate antigens remains a central challenge in computational antibody engineering. Existing antigen-conditioned generative models can jointly design antibody sequences and structures, but they are typically trained to match structural data distributions and lack an explicit alignment mechanism for antigen-specific preferences. We present AbSpecAlign, a plug-and-play specificity-reward alignment framework for antibody design. AbSpecAlign introduces a Bidirectional Multimodal Specificity Scorer (BiMSS), which integrates antibody and antigen sequence representations with SE(3)-equivariant structural embeddings to produce a contrastive reward over cognate versus non-cognate antigen pairs. We further introduce a GRPO-based post-training strategy that uses group-relative specificity rewards to align pre-trained antibody generators toward higher-scoring antigen-conditioned designs while maintaining structural regularity. To the best of our knowledge, AbSpecAlign is among the first frameworks to apply GRPO-style group-relative reward alignment to antigen-conditioned antibody sequence--structure co-design with an explicit off-target specificity reward. Experiments on RAbD CDR design benchmarks show that AbSpecAlign improves most sequence-recovery and structural-fidelity metrics when combined with diffusion-based antibody generators. Experiment results suggest that GRPO-based reward alignment with contrastive specificity signals provides an effective post-training mechanism for antigen-conditioned antibody generation, while prospective experimental validation remains necessary.
PaperID: 2962, Poster
Authors: Zijie Tian, Kevin Jin, yang Chen, Thomas Chun Man Lee
Abstract: Synthetic data generation is becoming increasingly prevalent as advances in generative modeling accelerate. However, theoretical understanding of when and how synthetic data can improve downstream estimation and prediction remains underdeveloped. This question is especially relevant when synthetic samples are generated from proposal mechanisms informed by trained generators or domain knowledge, whose distributions may differ from the target population. To address this question, we propose REAL-Syn: Risk-guided Estimation-aware Acceptance for Learning with Synthetic Data, a framework for statistical learning with synthetic data selection. REAL-Syn uses selected subsets of candidate synthetic data to reduce estimator variability and often improve downstream prediction under target-proposal mismatch. The proposed procedure uses a covariance-based risk surrogate to select from the set of candidate synthetic samples; the accepted samples are then reweighted and pooled with the real observations for inference. We present theoretical results on the surrogate-optimal acceptance rule and the proposed selection procedure. We empirically show that REAL-Syn achieves competitive performance in simulations and an application to solar flare intensity forecasting.
PaperID: 2963, Poster
Authors: Heehoon Kim, Jaehwan Lee, Taejeoung Kim, Jongwon Park, Jinpyo Kim, Pyongwon Suh, Ryan H Choi, Sangwoo Lee, Jaejin Lee
Abstract: The rapid growth of large language models is driving organizations to expand their GPU clusters, often with GPUs from multiple vendors. However, current deep learning frameworks lack support for communication across heterogeneous vendor GPUs, leading to inefficiency and higher costs. We present HetCCL, a collective communication library that unifies vendor-specific backends and enables RDMA-based communication across GPUs without requiring driver modifications. HetCCL introduces two novel mechanisms that enable cross-vendor communication while leveraging optimized vendor libraries, NVIDIA NCCL and AMD RCCL. Evaluations on a multi-vendor GPU cluster show that HetCCL matches NCCL and RCCL performance in homogeneous setups while extending support to heterogeneous environments, enabling practical and efficient training across NVIDIA and AMD GPUs without changes to existing deep learning applications.
PaperID: 2964, Poster
Abstract: Neural networks can fit noisy labels, but how their representations change as memorization begins remains unclear. We study this using persistent homology of penultimate-layer representations across training. Across 8 datasets and 8 architectures, we find a consistent pattern: the discriminative topological signal under noise appears in H_1 rather than H_0, with noisy representations preserving more loop structure than clean representations as training proceeds, supporting a folding view of memorization. When the clean topological signal decreases over training, the noisy signal does not decline as quickly as the clean one, producing an early crossover between clean and noisy trajectories; across the derivation set, crossover occurs when clean trajectories fade and the clean--noisy gap closes, whereas no-crossover cases remain persistent or do not close the gap within training. The same framework also correctly predicts the held-out ResNet-34/CIFAR-10 case, confirmed in all 9 runs. The signal reveals aspects of training dynamics that standard statistics do not represent directly. Activation variance is stronger on average, but the topological measure becomes informative earlier than validation loss and remains significant after controlling for intrinsic dimension. It also distinguishes irreducible label contradictions, including instance-dependent and CIFAR-10N human noise, from learnable confusions such as asymmetric noise. The same measure also falls during grokking in 10 of 12 modular-arithmetic runs and, after per-pipeline calibration from a clean reference plus runs at known noise rates, predicts population-level noise severity without per-sample labels, achieving per-pipeline R^2 \geq 0.94 across six tested configurations. These findings position topology as both an explanatory lens on noisy-label memorization and a practical early-stage diagnostic for studying, auditing and improving learning under label noise.
Abstract: Online bidding is a classical problem in online decision-making, with applications in resource allocation, hierarchical clustering, and the analysis of approximation algorithms. We study its randomized learning-augmented variant, where an online algorithm generates a sequence of random bids while leveraging predictions from an oracle. We provide analytical upper and lower bounds on the optimal consistency C as a function of the robustness R, which match when R \geq 2.885, effectively closing the gap left by previous work. The key technical ingredient is the notion of a bidding function, a novel abstraction that provides a unified framework for the design and analysis of randomized bidding strategies. We complement our theoretical results with an experimental application of randomized bidding to the incremental median problem, demonstrating the applicability of our algorithm in practical clustering settings.
PaperID: 2966, Poster
Authors: Yuchong Geng, Ao Tang
Abstract: We show that a generative model can discover concepts from data in an unsupervised setting while learning to generate. The Dirichlet Concept Diffusion Model (DCDM) embeds concept discovery into diffusion-based generation. DCDM learns concept centers and infers a Dirichlet distribution over them for each input. The resulting weighted center shapes the forward diffusion mean and reverse denoising process, so components are learned through the same evidence lower bound used to model the data. The method uses no class labels, attribute annotations, captions, pretrained text-to-image priors, or semantic concept supervision. Analysis shows how the formulation preserves concept-level information along diffusion paths and reduces denoising ambiguity, creating pressure for concept centers to capture stable modes. Experiments across diverse image domains show that the learned components form coherent prototypes, support interventions, and reflect recurring visual structure.
PaperID: 2967, Poster
Abstract: Autoregressive (AR) models have reformulated Scalable Vector Graphics (SVG) generation as next-token prediction, demonstrating remarkable potential in text-to-SVG tasks. However, controllable SVG generation guided by spatial signals (e.g., sketches or scribbles) remains largely unexplored within AR models. While a natural approach, inspired by controllable image generation, is to adapt methods like Condition Prefilling or Conditional Decoding, they fall short in SVG generation. To address these challenges, we introduce ControlSVG, a versatile and effective framework tailored for integrating precise spatial controls into autoregressive SVG generation. Firstly, we propose a Visual-Feedback Conditioning mechanism that constructs a “See-and-Draw" loop. This design mitigates cross-modal drift and ensures long-range consistency. Secondly, we explore spatially aware learning objectives that explicitly bridge discretized SVG tokens with their absolute 2D spatial coordinates, enhancing robustness against coordinate deviations during sequential generation. Extensive experiments on Canny-Edge-to-SVG generation demonstrate that ControlSVG outperforms existing baselines, achieving superior spatial alignment and high-quality SVG. We further provide preliminary validations on real human-drawn scribbles and extend the framework to scribble-to-SVG and image-to-SVG settings, suggesting its potential applicability beyond Canny-based control.
PaperID: 2968, Poster
Authors: Qi Teng, Xueer Wang
Abstract: Graph contrastive learning (GCL) typically maximizes cross-view agreement via an InfoNCE-based mutual-information (MI) surrogate, applying uniform alignment pressure across all anchors. Yet the reliability with which topology and attributes support cross-view agreement varies across nodes: strong pressure can benefit structurally consistent anchors but can harm boundary, low-agreement, or noisy anchors. We study this calibration problem in self-supervised node representation learning on small-to-medium attributed graphs. We propose MIRAGE, a hierarchical MI-surrogate regulation framework whose core consists of two components: anchor-wise MI setpoints that assign bounded alignment budgets to individual nodes, and dual-view MI stabilization that keeps branch-level surrogate estimates close to a target while limiting excessive anchor-wise dispersion. A label-free structure-reliability-gated hypergraph path serves only as conditional compensation for weak anchors when local structural signals are estimated to be reliable. On six node-classification benchmarks, MIRAGE achieves competitive accuracy, including 85.21% on Cora, 73.94% on Citeseer, and 92.57% on ACM. Beyond final accuracy, mechanism-level analyses show that MIRAGE tracks designated MI-surrogate targets during training and that changing the target level produces measurable downstream changes, indicating that target-tracked MI-surrogate regulation provides a practical optimization primitive for calibrating cross-view agreement in GCL.
Abstract: Transformer architectures exhibit cross-layer redundancies, yet post-training compression pipelines typically optimize layers in isolation or rely on heuristic grouping strategies that disregard layer-specific activation geometries. We introduce a principled, training-free framework that sequentially optimizes cross-layer weight pairings and shared-dictionary factorizations. Rather than forcing weights of adjacent layers to share a basis or heuristically merging activation statistics, our approach identifies structurally compatible projections and learns a shared representation that better preserves each layer’s distinct calibration geometry. Coupled with structured sparsity, this yields highly efficient weight decompositions without sacrificing functional fidelity. Across diverse architectures, scales, and modalities, our method achieves state-of-the-art results, consistently outperforming independent structured weight decompositions and alternative pairwise weight factorizations, which operate under heuristic grouping strategies. By replacing heuristic engineering strategies with a convergent, optimization-driven pipeline, we establish a theoretically grounded foundation for scalable, transformer compression across different modalities.
Abstract: At the core of reinforcement learning is the idea of learning beyond the performance in the data. However, scaling such systems has proven notoriously tricky. In contrast, techniques from generative modeling have proven remarkably scalable and are simple to train. In this work, we combine these strengths, by deriving a direct relation between policy improvement and guidance of diffusion models. The resulting framework, CFGRL, is trained with the simplicity of supervised learning, yet can further improve on the policies in the data. On offline RL tasks, we observe a reliable trend -- increased guidance weighting leads to increased performance. Of particular importance, CFGRL can operate without explicitly learning a value function, allowing us to generalize simple supervised methods (e.g., goal-conditioned behavioral cloning) to further prioritize optimality, gaining performance for "free" across the board.
PaperID: 2971, Poster
Authors: Geer Yang, Xixuan Liu, Bin Wu, Shaojiang Wang
Abstract: Persistent LLM agents need memory that preserves state, provenance, validity, and evidence across long interaction histories. We introduce StateLedger, a path-addressed external memory substrate that stores memory as durable state evolution rather than flat retrievable items: events update canonical paths, registers absorb small changes, CommitPages record residual drift against Checkpoints, cue views bound traversal, and content-addressed artifacts avoid evidence duplication. We prove storage, retrieval, coalescing, and conflict-aware ranking guarantees under explicit coverage and reader-consistency conditions. Empirically, StateLedger improves task performance and storage efficiency across LongMemEval, LoCoMo, and Evo-Memory Multi-turn, reaching 98.8% on LongMemEval-S, 82.6% on LongMemEval-M with 14.7% durable footprint, 95.8% LoCoMo non-adversarial accuracy, 87.1% unweighted success, and 93.5% unweighted progress on Evo-Memory.
PaperID: 2972, Poster
Abstract: Brain-to-speech decoding aims to restore communication by translating neural activity into speech-related representations. From a modeling perspective, this task can be formulated as a weakly supervised neural sequence decoding problem, where phoneme sequences must be recovered without frame-level annotations. Existing speech decoders commonly rely on recurrent sequence models to integrate temporal context, but their dense sequential computation can limit their suitability for resource-constrained portable or implantable neuroprosthetic systems. In this study, we propose a residual spiking neural network (SNN) for phoneme-level decoding of intracortical speech signals under constrained model capacity and weak sequence-level supervision. Our approach incorporates biologically inspired fixed connection-specific delays to integrate temporal context, providing each postsynaptic unit with access to recent presynaptic spike activity without recurrent state or dense temporal convolutions. Because each directed connection is assigned a single fixed delay, the temporal field is expanded without adding trainable delay parameters. Across 500k, 2M, and 5M core-parameter budgets, fixed delays consistently improve phoneme error rate (PER), with gains saturating once the delay range covers the task-relevant temporal horizon. Boundary-aligned CTC analysis further suggests that delayed temporal context reduces local uncertainty around phoneme transitions, indicating that fixed delays support weakly supervised phoneme alignment rather than simply increasing model capacity. Compared with recurrent, convolutional, and feedforward baselines under matched core-parameter budgets, the proposed SNN achieves the best PER at 500k and 2M parameters and the second-best PER at 5M parameters, while reducing estimated operation-level energy by approximately 2-4× relative to the strongest non-spiking baselines. These results suggest that fixed-delay SNNs provide an efficient temporal modeling mechanism for brain-to-speech decoding, potentially enabling lightweight decoding for on-device speech neuroprostheses.
Abstract: Unified and scalable Transformers have recently achieved remarkable success in modeling diverse phenomena traditionally associated with computer graphics, such as 3D visual effects, rendering processes, and motion in videos. In this work, we take a step further by investigating whether modern Transformer techniques can tackle the challenging task of cloth simulation. To this end, we present ClothTransformer, a framework that reformulates cloth simulation as autoregressive sequence modeling in a learned latent space. Existing neural cloth simulators are largely specialized to single scenarios, intrinsically coupled to the mesh discretization, and lack robust collision handling. Our approach addresses these limitations through three contributions: (1) a unified Transformer architecture that handles diverse scenarios---body-driven garments, robotic manipulation, and free-fall collisions---under a single model and achieves approximately 4--9× lower error than prior state-of-the-art across all scenarios; (2) a scalable latent-space formulation that compresses arbitrary-resolution meshes into a fixed-size set of latent tokens, making temporal dynamics computation independent of mesh resolution; and (3) a diverse-scenario high-fidelity penetration-free dataset of ~493.4k frames spanning all three settings, which enables a differentiable Continuous Collision Detection (CCD) module to suppress penetration artifacts.
PaperID: 2974, Poster
Abstract: Language models often produce high-quality individual samples but poor sets, with repeated samples clustering around the same semantic modes. We formalize diverse generation as sampling a size-K subset of complete strings from a determinantal point process (DPP), where the likelihood encodes item quality, and the embedding geometry encodes repulsion between similar outputs. Exact inference over all strings is intractable, so we propose \algname: draw a finite candidate pool from a tractable proposal, importance-weight candidates toward a globally tempered quality distribution, and run either K-DPP sampling or greedy MAP selection on the induced pool kernel. We prove a finite-sample bound showing that the KL bias of the pooled sampler decays as \mathcalO(1/N) in the pool size. Empirically, \algname improves the quality--diversity Pareto frontier over temperature sampling, diverse beam search, K-means reranking, and LLM-based selection on open-ended generation, ambiguous text-to-SQL (Ambrosia), and mathematical reasoning (GSM8K) tasks.
Abstract: Discrete diffusion language models (dLLMs) accelerate text generation by unmasking multiple tokens in parallel. However, parallel decoding introduces a distributional mismatch: it approximates the joint conditional using a fully factorized product of per-token marginals, which degrades output quality when selected tokens are strongly dependent. We propose DEMASK (DEpendency-guided unMASKing), a lightweight dependency predictor that attaches to the final hidden states of a dLLM. In a single forward pass, it estimates pairwise conditional influences between masked positions. Using these predictions, a greedy selection algorithm identifies positions with bounded cumulative dependency for simultaneous unmasking. Under a sub-additivity assumption, we prove this bounds the total variation distance between our parallel sampling and the model's joint. Empirically, DEMASK achieves 1.7--2.2× speedup on Dream-7B while matching or improving accuracy compared to confidence-based and KL-based baselines. When applied to dParallel, a block diffusion model, DEMASK also produces Pareto-optimal accuracy-step trade-offs across math and code benchmarks.
PaperID: 2976, Poster
Abstract: Diffusion personalization deployments serve growing catalogues of LoRA adapters, but resident-LoRA serving is capped near the GPU memory budget (about 150 rank-16 adapters on an A100 80 GB GPU running FLUX.1-dev). We introduce LITHE, a LoRA-compatible deployment representation that stores each adapter as an integer-index stream over a deterministic Hadamard codebook, and that pushes this resident-adapter ceiling to a verified 10,000 adapters on the same single GPU at sub-80 GB peak. The same on-disk indices drive two serving modes: an mode within 1% of an optimized LoRA+compile baseline at 28-step inference at 1024² resolution on FLUX.1-dev (and 1.07–1.43× faster at lower resolutions). On disk, each rank-16 FLUX.1-dev adapter occupies 1.81 MB on average across the benchmark pool, a 50×/200× reduction over the matched r=16 / r=64 LoRA disk payload after the same coder.
PaperID: 2977, Poster
Abstract: Concise reasoning in large language models (LLMs) seeks to generate only essential steps needed to arrive at a final answer, thereby alleviating issues of . Most proposed approaches scalarize length and reward into a single objective, requiring coefficients or thresholds that must be re-tuned across domains and model scales. We address this brittleness by treating concise reasoning as a constrained problem that minimizes length subject to an accuracy floor, and deriving a tractable algorithm, Performance-Aware Length Update (PALU), which replaces each intractable update of Lagrangian optimization with a tractable surrogate while preserving its structural prescription. On improves accuracy from 44.2% to 51.6% across six benchmarks, a level matched by GRPO only with 2.5× more tokens. The Lagrangian dynamics yields a two-phase compression: an initial phase where the residual signal still allows compression without performance cost, followed by a trade-off phase where performance gates further length reduction. The same hyperparameters transfer across math, logic and STEM, and across 1.5B to 14B models, suggesting that constrained optimization is a productive design scaffold for training reasoning LLMs.
Abstract: Diffusion models have emerged as powerful learned priors for Bayesian inverse problems (BIPs). Diffusion-based solvers rely on a presumed likelihood for the observations in BIPs to guide the generation process. Likelihood misspecification is common in practical BIPs and is known to degrade recovery performance, particularly under outlier contamination. We investigate this problem by first characterizing the induced posterior deviation and proving the stability of diffusion-based solvers for linear BIPs. Our stability analysis further reveals potential robustness deficiencies of existing diffusion-based solvers under outlier-contaminated measurements. To address this issue, we propose a simple yet effective solution: robust diffusion posterior sampling, which is provably outlier-robust for linear BIPs and compatible with existing gradient-based posterior samplers. Empirical results from scientific inverse problems and natural image tasks demonstrate the effectiveness and robustness of our method, with consistent performance gains in challenging scenarios involving outlier contamination for both linear and nonlinear tasks.
PaperID: 2979, Poster
Authors: Juan V Cano, Phuong C Cong, Kostadin Damevski, Thang N Dinh
Abstract: Large-scale Ising optimization underlies critical optimization and learning applications, including correlation clustering, MAP inference, and energy-based models. Spectral relaxation methods run in polynomial time but yield solutions with significant optimality gaps while metaheuristics close those gaps but scale poorly to large instances. In particular, spectral methods are inherently single-shot as they solve exactly one eigenvalue problem. We introduce Spectral Annealing (SpecAnn), a new spectral optimization approach that transforms traditional spectral methods from single-shot approximations into exploration-based optimization, by sweeping a parametrized normalization path from the raw adjacency to the signed normalized Laplacian, with an O(n) diagonal predictor that warm-starts successive eigensolves and GPU-batched local search refining candidates in parallel. Experiments on general Ising problems demonstrate that SpecAnn achieves competitive to near-optimal solutions depending on instance structure, with runtimes less than 7 minutes for instances of 8.4M variables. SpecAnn delivers up to 83× speedup over recent GPU-accelerated methods while maintaining superior solution quality, and scales successfully beyond 524K variables where traditional metaheuristics fail.
Authors:
Zheng Huang, Mingyu Liu, Xiaoyi Lin, Muzhi Zhu, Ye Lin, Canyu Zhao, Zongze Du, Xiaoman Li, Hao Zhong, Yiduo Jia, Hao Chen, Chunhua ShenAbstract: Robot fine-tuning can improve manipulation success while degrading the semantic structure inherited from large vision-language pretraining. This trade-off limits the reuse of vision-language-action (VLA) policies in settings that require compositional instructions, camera changes, and coordination with higher-level agents. We present NoTVLA, a semantics-preserving robot adaptation framework that replaces dense low-level action supervision with a sparse narrative action interface. NoTVLA converts demonstrations into sparse, semantically meaningful waypoints, grounds each decision with a task-relevant visual anchor and depth value, and reconstructs executable motion through a deterministic detokenizer. The resulting interface keeps the autoregressive vision-language backbone close to its pretrained prediction format while delegating high-frequency control to a transparent motion-rendering stage. We evaluate this design through matched and backbone-family comparisons, semantic retention probes, semantic out-of-distribution manipulation tasks, camera and depth perturbations, and deployment-oriented efficiency analysis. The results suggest that robot adaptation should be judged not only by task success, but also by how much task-relevant semantic competence is preserved during fine-tuning.
PaperID: 2981, Poster
Abstract: Deep multi-view clustering often struggles with inter-view feature redundancy and dimensionality collapse during iterative representation fusion. To address these limitations, we propose Orthogonal Residual Consensus Alignment (ORCA). First, a Semantic Incremental Recurrent Fusion (SIRF) module explicitly separates shared redundancy from complementary residuals via geometric projection. By employing an adaptive halting mechanism, SIRF autonomously maximizes information gain while avoiding structural collapse. Second, a Consensus-Calibrated Hierarchical Contrastive Alignment (CCHA) module is introduced to preserve multi-scale semantics. It adaptively aligns view-level, local, and recurrent representations using the global consensus as a semantic anchor. Consequently, ORCA learns latent representations with robust structural integrity and highly discriminative semantics. Extensive experiments on multiple benchmarks demonstrate that ORCA consistently outperforms state-of-the-art methods, with comprehensive ablations verifying its robustness.
PaperID: 2982, Poster
Abstract: Modern data-driven decision-making methods, such as imitation learning (IL) and reinforcement learning (RL), have achieved great success in solving many complex tasks. However, these methods often suffer from serious control instability and robustness issues when applied in real-world applications such as robotics and autonomous driving, posing notable challenges for their practical deployment. We argue that this instability issue stems largely from their limitations on solely supervising and optimizing zeroth-order actions (i.e., the action labels), failing to account for higher-order action dynamics and temporal consistency. In this paper, we show that simultaneously supervising both zeroth- and first-order actions can dramatically enhance policies' performance and control robustness. To achieve this, we introduce a novel and elegant loss scheme supported by formal theoretical guarantees that can equip any off-the-shelf policy model (e.g., deterministic, stochastic, or flow policies) with the capability for higher-order action supervision, without requiring any structural modifications. Moreover, our proposed method can serve as a lightweight plug-and-play module that seamlessly integrates with a broad spectrum of existing offline RL frameworks. Extensive evaluations on OGBench and D4RL demonstrate that our approach yields substantial performance and robustness improvements across a wide range of continuous control environments. Notably, our method can also enhance policies' out-of-distribution (OOD) generalization capability in the challenging low-data regime, making it an ideal tool in tackling many real-world control problems.
PaperID: 2983, Poster
Abstract: We address the problem of learning generative priors over high-dimensional latent variables from indirect, noisy observations, with a focus on continuous heterogeneity in cryo-electron microscopy (cryo-EM). In this setting, 3D atomic structures x \in \mathbbR^3N_a are never directly observed and must be inferred from 2D projection images y generated by a complex, non-linear forward operator. We propose Empirical Bayes Flow Matching (EB-FM), a method that learns a flow matching prior p_\theta(x) directly from observations by embedding it within a Monte Carlo stochastic approximation Expectation-Maximization framework. EB-FM alternates between an E-step that performs approximate posterior sampling from p_\theta(x \mid y) via guided flow-based generation while maintaining persistent latent estimates through stochastic approximation, and an M-step that updates the flow model using these stabilized latent variables as training targets. This approach eliminates the need for pretraining on clean data and naturally accommodates non-linear forward models. We validate our method on controlled inverse problems using two-moons and MNIST data, and demonstrate its effectiveness on simulated cryo-EM data with known poses, showing that EB-FM can recover continuous distributions of protein conformations directly in atomic-coordinate space.
PaperID: 2984, Poster
Abstract: Maximum entropy reinforcement learning (MaxEnt-RL) enables robust exploration, yet practical implementations often restrict policies to simple Gaussians. While recent approaches incorporate expressive generative policies via importance-weighted supervised learning, they are prone to importance weight collapse, which limits their scalability in high-dimensional action spaces. Our key insight is to mitigate this limitation by localizing the sampling region, avoiding the weight degeneracy induced by importance sampling over the entire action space. To instantiate this insight, we introduce uidance). FLAG augments the state space with a flow latent variable and optimizes a provably consistent proxy MaxEnt-RL objective. We empirically demonstrate that FLAG enables expressive policy optimization with limited importance samples and scales to high-dimensional control tasks. Furthermore, FLAG achieves state-of-the-art performance across challenging benchmarks.
PaperID: 2985, Poster
Authors: Xiaoyun Chang, ZiJun Zhang, Jie Chen, Xiaopeng Wen, Xiangbo Lin, Xiaohong Ma, yi sun
Abstract: Modeling fine-grained human manipulation of object functional parts is essential for achieving human-like dexterous manipulation in virtual reality and robotics. Despite recent advances, data-driven generation methods remain largely constrained by patterns learned from limited training data, rendering them prone to hallucinating task-inconsistent manipulations when encountering samples beyond the training distribution. To address this challenge, we propose ManipulationRAG, a framework that augments an internal generative model with externally retrieved, task-relevant manipulation knowledge. Specifically, we design a hand-centric manipulation taxonomy and construct a structured knowledge base that organizes task descriptions, objects, manipulation types, and corresponding hand pose sequences, supported by a retrieval mechanism for accurate knowledge retrieval. The retrieved knowledge serves as an external manipulation-specific prior, guiding the diffusion model toward task-consistent and physically plausible manipulation synthesis. Extensive experiments demonstrate that our method outperforms existing methods and generalizes robustly to object manipulation beyond the original diffusion model's training domain.
PaperID: 2986, Poster
Abstract: The efficiency of Large Language Model (LLM) serving has been significantly improved by prefix caching, which enables reuse of previously computed activations. However, extending prefix caching to hybrid linear–softmax language models remains challenging due to the presence of large recurrent states in linear attention modules. In practice, these states cannot be stored at every token position due to their substantial memory footprint, and are instead saved periodically (e.g., at chunk boundaries). This leads to a fundamental mismatch between token-level KV caching and coarse-grained state caching, resulting in reduced effective cache hit rates and additional recomputation overhead. In this work, we propose \sysname, a system designed to address the limitations of prefix caching in hybrid linear–softmax LLMs. \sysname is built on two key observations: (1) recurrent state naturally exhibit different effective memory ranges across heads, and (2) the distribution of state values presents structured outliers that hinder effective compression. Based on these insights, \sysname introduces two novel designs. First, we develop a recurrent-aware state reuse mechanism that selectively stores short-range head states, while reuse previously cached long-range head states, enabling fine-grained alignment with KV caching without incurring prohibitive memory cost. Second, we propose a dual-scale state quantization method that applies independent scaling along row and column dimensions to effectively compress states with structured outliers. We evaluate \sysname on hybrid linear–softmax LLMs under realistic serving workloads. Experimental results show that \sysname significantly improves effective prefix cache utilization, reduces recomputation overhead, and achieves substantial throughput gains without sacrificing model accuracy. These results highlight \sysname as an effective system-level solution for efficient LLM serving in hybrid architectures.
PaperID: 2987, Poster
Authors: Shosuke Suzuki, Toshiyuki Amagasa
Abstract: Deep-learning protein structure predictors achieve near-experimental accuracy on individual folds, yet their default inference samples concentrate around a single dominant conformation. We introduce ConforFlux, an inference-time procedure for Boltz-2 that couples M structure-prediction trajectories through a pairwise C\alpha-RMSD repulsion gradient on the trunk's single and pair embeddings. Because the trunk conditions every block of the diffusion module, this update propagates to every subsequent denoising step. On four conformational-change categories, ConforFlux improves the per-state success rate over Default Boltz-2 by 3–17 percentage points while preserving Default-comparable physical quality. On twelve transporter pairs with at least one alternate-state reference released after the Boltz-2 cutoff, the 2 Å success rate rises from 4/12 to 9/12. On the human dopamine transporter, ConforFlux additionally reaches the post-cutoff inward state, which 500 Default samples never attain.
PaperID: 2988, Poster
Abstract: Adapting language models on long-form corpora can improve domain behavior while also making protected passages easier to reproduce verbatim from short prefixes. We study this tension through \emphrelative energy gain, the token-level log-density ratio between an adapted model and its pre-adaptation anchor on an exact protected continuation. This quantity separates ordinary predictability from update-induced copying pressure and accumulates over suffixes as a likelihood-ratio advantage for extractable memorization. Building on this view, we introduce Energy-Gated Copyright Regularization (EGCR), a training objective that augments maximum likelihood with a soft relative-energy budget on protected tokens. EGCR uses a differentiable gate combining local persistence of relative gain with anchor surprisal, so regularization is concentrated on distinctive spans whose probabilities rise in a memorization-like way while ordinary next-token supervision is retained elsewhere. Across BookMIA adaptation stress tests, training-prefix extraction, CopyBench transfer, and gate ablations, EGCR reduces exact-match and overlap-based reproduction while preserving held-out language-modeling and book-related utility. Token-level visualizations further show that the gate focuses on a small set of expressive positions associated with verbatim recovery. These results suggest that relative-energy barriers offer an effective and inspectable mechanism for copyright-aware language model adaptation.
PaperID: 2989, Poster
Authors: Haoyang Su, Qi Chen, Max Gutbrod, Johan Verjans, Zhibin Liao
Abstract: Out-of-distribution (OOD) detection plays a pivotal role in ensuring the reliability and trustworthiness of AI systems. Although existing approaches achieve strong performance by leveraging features, logits, or both, most of them cannot generalize well across domains due to either overlooking critical information in low-level representations from shallow layers or utilizing sub-optimal layer selection strategies. In this paper, we propose Multi-Layer Adaptive Mahalanobis-Cosine Similarity (ML-AMCoS), a method that adaptively integrates cosine similarity with class-conditional Gaussian distributions. ML-AMCoS is designed to capture subtle nuances in feature covariance when intra-class variance is low, while suppressing covariance noise when the variance is high. To optimize the utilization of each layer, we calculate contribution weights based on performance against pseudo-OOD samples, which are generated by cut-mixing in-distribution (ID) training images with the four corners of images from different classes. Extensive experiments across diverse domains demonstrate that our method significantly improves robustness and detection accuracy, achieving an average AUROC of 81.13% and FPR@95 of 40.12% on near-OOD detection, outperforming the state-of-the-art methods by 7.25% and 7.08%, respectively.
Abstract: A fundamental challenge in science and engineering is the simulation-to-experiment gap. While we often possess prior knowledge of physical laws, these physical laws can be too difficult to solve exactly for complex systems. Such systems are commonly modeled using simulators, which impose computational approximations. Meanwhile, experimental measurements more faithfully represent the real-world, but experimental data typically consists of observations that only partially reflect the system’s full underlying state. We propose a data-driven distribution alignment framework that bridges this simulation-to-experiment gap by pre-training a generative model on fully observed (but imperfect) simulation data, then aligning it with partial (but real) observations of experimental data. While our method is domain-agnostic, we ground our approach in the physical sciences by introducing Adversarial Distribution Alignment (ADA). This method aligns a generative model of atomic positions—initially trained on a simulated Boltzmann distribution—with the distribution of experimental observations. We prove that our method recovers the target observable distribution, even with multiple, potentially correlated observables. We also empirically validate our framework on synthetic, molecular, and experimental protein data, demonstrating that it can align generative models with diverse observables.
PaperID: 2991, Poster
Authors: Yunfei Chen, Renwei Xia, Shangchong Gao, Zhan Yang
Abstract: Cross-modal hashing is essential for efffcient large-scale re-trieval in multimedia analysis. However, existing unsupervised methods often struggle to fully capture interactions between modalities, enforce multi-level alignment, and maintain stable optimization. To overcome these challenges, we propose Decoupled Prototype Contrastive Align-ment Hashing (DPCAH), which employs a cross-modal semantic con-trastive learning module that concatenates image and text features and encodes their interactions via a transformer to produce uniffed represen-tations. feature-level and hash-level contrastive objectives jointly align semantic information and guide discriminative hash code learning across modalities. A decoupled prototype consistency module further models cross-modal correlations independently, enhancing semantic alignment while ensuring stable and robust optimization. Experiments on three benchmark datasets demonstrate that DPCAH outperforms state-of-the-art methods in cross-modal retrieval.
PaperID: 2992, Poster
Abstract: Transformer models have demonstrated a remarkable ability to perform a wide range of tasks through in-context learning (ICL), where the model infers patterns from a small number of example prompts provided during inference. However, empirical studies have shown that the effectiveness of ICL can be significantly influenced by the order in which these prompts are presented. Despite its significance, this phenomenon has been largely unexplored from a theoretical perspective. In this paper, we theoretically investigate how positional encoding (PE) affects the ICL capabilities of linear attention transformer models, particularly in tasks where prompt order plays a crucial role. We examine two distinct cases: linear regression, which represents an order-invariant task, and dynamical systems, a classic time-series task that is inherently sensitive to the order of input prompts. Theoretically, we evaluated the change in the model output when two types of positional encoding (one-hot and RoPE) is incorporated and the prompt order is altered. In all cases, the leading dependence on the permutation size and context length is k/N, with constants determined by the positional encoding, model weights, input dimension, and task dependent parameters. These theoretical findings are experimentally validated.
PaperID: 2993, Poster
Authors:
Yunpeng Mei, Jiakai He, Hongjie Cao, Chenyu Wang, Xiaowen Zhu, Yihan Zhou, Jiamin Wang, Chenbo Xin, Peng Cheng, Yuxuan Yang, Yijie Wang, Xinhu Zheng, Gao Huang, Jie Chen, Gang WangAbstract: Large vision-language-action (VLA) policies are increasingly trained as conditional generative models over action chunks, yet they typically improve by imitating the data they are given. This is limiting in deployment, where robots collect mixed-quality experience containing demonstrations, partial completions, recoverable mistakes, and failures. Full behavior cloning imitates failures, filtered behavior cloning discards useful sub-trajectories, and offline RL usually requires a large separate critic. We introduce ForesightFlow, a self-guided flow policy that augments each generated action chunk with a learned success-potential vector. The same flow therefore proposes candidate actions and scores them, enabling best-of-K inference without an external critic. The key challenge is that policy improvement and value calibration require different supervision: advantage weighting should suppress low-quality actions, but applying the same weights to the potential dimension removes the failure gradients needed for calibration. We address this with decoupled advantage-weighted flow matching, applying exponentiated advantage weights only to action velocities while training potential velocities uniformly. We also derive a single-step boundary estimator for conditional flow matching with independent endpoint sampling, allowing advantage computation with one stop-gradient forward pass. Across five BEHAVIOR-1K simulation tasks and five real-world bimanual tasks, ForesightFlow improves over imitation baselines, matches the strongest separate-critic baseline in average simulation success, improves real-world success, and reduces training compute by 38%. Ablations show that decoupling prevents value hallucination, the one-step estimator preserves candidate-ranking fidelity, and self-guided sampling improves long-horizon performance.
PaperID: 2994, Poster
Authors: Matthew O'Callaghan, Kaisey Mandel, Gerard Gilmore
Abstract: Recent advances in neural density estimation have enabled amortized Bayesian inference for complex stochastic simulators. However, these methods rely on simulators accurately reflecting the true data-generating process and can degrade significantly under model misspecification. We consider the setting where multiple unlabeled observations are available and introduce robust variational neural posterior estimation (RVNP), an amortized Bayesian inference method that uses an importance-weighted autoencoder to jointly learn a misspecification-robust posterior and an explicit error model that captures the misspecification gap. Our results show that RVNP can recover robust posterior inference in a data-driven and interpretable manner, outperforming previous methods across multiple metrics and misspecified benchmarks.
PaperID: 2995, Poster
Abstract: Neural Combinatorial Optimization (NCO) is a promising paradigm for solving routing problems, yet its generalization across diverse data distributions remains a critical bottleneck. Existing methods typically tackle cross-distribution generalization through complex multi-distribution training or intricate meta-learning schemes, implicitly presuming that single-distribution training yields a weak zero-shot baseline. In this paper, we demonstrate that the potential of single-distribution training has been substantially underestimated and can be unlocked by a simple architectural inductive bias: Mix Normalization. To approach this, we first construct a comprehensive benchmark spanning three representative routing problems and 179 fine-grained datasets, capturing diverse shifts in node coordinates, customer demands, and time windows. Building on this benchmark, we propose MixRoute, a neural framework that adaptively combines normalization statistics at different granularities to mitigate varying distribution shifts. Trained exclusively on uniform distributions, MixRoute generalizes effectively to a wide range of unseen distributions without any distribution-specific adaptation. Extensive experiments demonstrate that MixRoute consistently achieves state-of-the-art zero-shot generalization. Our work revisits normalization as a parsimonious alternative for zero-shot generalization against complex training-adaptation schemes.
Authors:
Dahye Kim, Bhuvan Sachdeva, Karan Uppal, Naman Gupta, Vineeth N Balasubramanian, Deepti GhadiyaramAbstract: While most frames in long-form video are redundant, the critical information resides in temporal surprises: moments where the actual visual features deviate from their predicted evolution. Inspired by the human brain’s predictive coding, we introduce Swift Sampling, an elegant, training-free frame selection algorithm that automatically identifies high-information moments in a video. Specifically, we model a video as a differentiable trajectory in the visual latent space and compute the velocity and acceleration of its features. Then, we apply Taylor expansion to project the expected path of subsequent frames. Frames that diverge sharply from this predicted manifold are identified as temporally surprising frames and selected for sampling. Unlike prior training-free methods that rely on auxiliary networks or video-specific hyperparameter tuning, Swift Sampling is incredibly lightweight, adding only \mathbf0.02× additional computational cost over baseline making it 30× cheaper overhead than leading baselines. Across three long-video question answering benchmarks and 10 different downstream tasks, Swift Sampling outperforms uniform sampling and prior query-agnostic baselines. It is especially powerful for long videos with limited frame budgets improving accuracy by up to \mathbf+12.5 points.
PaperID: 2997, Poster
Authors: Felix Kunze, Antonio Di Maio
Abstract: Federating a LGN is challenging because each gate is parameterised by a distribution over 16 Boolean operations, so averaging parameters across clients does not correspond to averaging the underlying functions. Under non-IID data, such parameter-space merging collapses to near-chance accuracy. Boolean Feature Selection (BFS) avoids this issue by selecting discrete Boolean functions rather than averaging parameters. We propose Federated Boolean Feature Selection (FBFS), which extends BFS to LGNs by treating hardened last-layer gates as candidate Boolean features and aggregating per-class, per-gate sufficient statistics across clients. Clients share only firing aggregates computed on a held-out split, allowing the server to reconstruct discriminative scores, defined as the gap between a gate's mean firing for a class and for the remaining classes, without accessing raw data. This procedure is equivalent to centralized BFS on the union of client data. We further introduce a diversity-regularized variant of FBFS that encourages complementary feature selection across gates. Empirically, FBFS is the only data-free method we evaluate whose performance remains stable under extreme non-IID settings, outperforming parameter-space merging and other baselines by large margins. It maintains high accuracy across a wide range of heterogeneity levels and scales to large architectures. In addition, FBFS enables efficient deployment: the resulting models compile to hardware-efficient logic representations with substantially fewer post-synthesis cells than ensemble-based alternatives while remaining bit-exact with their software counterparts.
Abstract: Differentially private stochastic gradient descent (DP-SGD) is used to train deep learning models while mitigating many privacy risks. One popular line of work, designed to chip away at the accuracy gap between DP-SGD and normal SGD training, is known as DP-MF. To each update, it adds privacy noise that is correlated across training iterations, so that later noise partially cancels out earlier noise. The noise correlation structure is determined by solving an optimization problem, but a key question is not well-understood: what important properties does the noise correlation need in order to improve accuracy? In this paper, we identify two ways that noise affects the training trajectory --- noise alters the value of the gradient, and also alters the location at which subsequent gradients are computed. Prior works addressed the first but overlooked the second effect, realizing only limited benefits. Our analysis results in a new objective function for determining cross-iteration noise correlation. Our technique, NoiseCurve, consistently improves accuracy over DP-BandMF, the state-of-the-art DP-MF scheme, on various computer vision and NLP tasks. NoiseCurve uses an upper bound \hatH on the Hessians of the loss that could be encountered during training. We show how to estimate \hatH using public data and evaluate robustness to errors in \hatH. To avoid direct computation of \hatH, or even its eigenspectrum, we show how to adapt the Lanczos method.
PaperID: 2999, Poster
Abstract: A single global fairness criterion cannot serve communities whose conceptions of fair treatment differ. We introduce \emphelicited localized fairness: an end-to-end pipeline that elicits per-community Mahalanobis metric--tolerance specifications \cF_k = (d_k, \eps_k) from stakeholders via two phases of pairwise queries (similarity, then tolerance acceptability), trains a localized classifier (\textscLocAdapt), and audits the deployed model on a held-out tolerance fold. \emphTheory. A constrained-MLE rate \tilde O(\sqrtd^2/n) for metric elicitation and a matching \Omega(d^2/t^2) active-query lower bound, exact in (d, n, t) exponents; a held-out audit certificate that converts Phase 2 queries into a PAC-style violation guarantee. \emphEmpirics. On heterogeneous-community settings (Stress K=5, Adult-K3-human; 10 seeds), \textscLocAdapt sits on the (accuracy, strict d\leq1 violation) Pareto frontier: averaged and matched-accuracy Global IF lose on strict at comparable accuracy (40--80% higher), and pooled-MLE Global IF matches strict only at -2pp accuracy. A five-axis diagnostic (acc, strict, score-diff spread, prediction entropy, AUC) confirms the gain is not prediction-smoothing. Plugging Phase 2 elicited tolerances into \textscLocAdapt training reduces the COMPAS deployed-spec violation 66% (p\ll10^-4). Five pre-registered human studies on COMPAS (360 annotators total) validate the pipeline end-to-end; a 180-annotator scorer-stability ablation across three scorer architectures preserves the cross-community \hat\eps ordering and pools to p=0.0014 (n=90 vs.\ 90).
Abstract: The classic paradigm of language identification in the limit models learning as a game between an adversary, who reveals strings from an unknown target language, and a learner tasked with identifying that language. The recently introduced framework of language generation in the limit shifted the objective to better reflect modern language modeling, requiring the learner to produce valid, unseen strings from the target language. Related work highlighted a fundamental tension: a broad coverage of the target often comes at the cost of validity. We introduce a new notion of precision and recast this problem as the classic recall--precision trade-off. We analyze generation in the limit under varying constraints on enumeration, novelty, and validity, aimed at reflecting settings closer to those encountered by large language models. A key contribution is our analysis of learners that are not eventually valid: we allow infinitely many mistakes, provided their frequency tends to zero so that precision remains 1. We show this relaxation can strictly increase recall when the adversary permanently withholds a large portion of the target language. We also study a continuous relaxation of the novelty constraint, requiring only that a fixed fraction of outputs are novel. Taken together, our results move toward a more realistic model of language generation where occasional errors and repetitions are unavoidable, but their rates are controlled.
Authors: Luchao Wang, Kaimin Liao, Hua Wang, Qian Ren, Zhi Chen, Yaohua Tang
Abstract: While a converged 3D Gaussian Splatting (3DGS) model may accurately approximate a target scene, its underlying parameterization often becomes severely ill-suited for further optimization. We identify this late-stage bottleneck as \emphparameterization degeneration: high-opacity floaters truncate gradient flow to background surfaces via alpha compositing, and redundant overlapping clusters cause severe parameter coupling with nearly collinear Jacobian responses. These structural barriers explain why continued optimization plateaus, even when removable artifacts persist. To break this deadlock, we propose ReorgGS, an equivalent distribution reorganization method. By treating the converged Gaussian set as an empirical probability field, ReorgGS resamples centers, estimates local anisotropic covariances via kNN, and initializes a low-opacity state before resuming optimization. Unlike standard opacity reset—which only rescales weights on a flawed topology—ReorgGS fundamentally rebuilds the spatial and visibility structure. Our analysis reveals a crucial insight: \emphdistributional equivalence does not imply optimization equivalence. By preserving scene support while drastically improving gradient accessibility and reducing opacity-weighted overlap, ReorgGS provides a vastly superior optimization landscape. Under the same optimization budget and fixed Gaussian count, ReorgGS breaks the performance ceiling, suppresses persistent floaters, and reduces rendering overhead by eliminating redundant overlap.
PaperID: 3002, Poster
Abstract: Agentic systems built on large language models (LLMs) are increasingly deployed in high-stakes settings, but the causal mechanisms by which adversarial agents induce harmful outcomes remain poorly understood. We address this with a single-call counterfactual necessity test for multi-agent LLM systems that asks: \emphwhich specific agent communication was causally responsible for the harmful outcome? In contrast with actual-causality frameworks that reason over set-valued causes, we use one attribution rule throughout: single-call necessity for a selected proximate call. The key technical contribution is the \emphGumbel-Max tape, which lifts the per-token Gumbel-Max counterfactual generator of \citechatzi2024counterfactual from a single LLM to a multi-agent conversation. Per-token GPU random number generator (RNG) states are recorded during a factual run on a tape shared across all agents, then replayed during counterfactual (CF) runs everywhere except at the intervened call. Applied to the BAD-ACTS benchmark across 148 adversarial scenarios spanning four multi-agent environments and two prompting conditions (\emphnon-safe and \emphsafe), the test shows that single-call necessity tracks communication structure: decentralized and hierarchical environments (Travel Planning, Financial Article Writing) exhibit meaningful aggregate causal effect (ACE), while sequential debate (Multi-Agent Debate) shows near-zero ACE despite comparable attack success rate. A replay sanity check confirms that all CF effects under identity intervention are exactly zero, validating the tape mechanism. The results identify where targeted defenses are likely to be effective in deployed agentic systems.
PaperID: 3003, Poster
Abstract: Graph-structured data is crucial in various domains like biology and social networks. Comparing graphs, which is a fundamental problem in graph data analysis, is nonetheless highly challenging. Recently, the Gromov-Wasserstein (GW) distance has provided a principled way to compare two graphs. However, computing the GW distance involves solving a complex non-convex optimization problem, making it computationally expensive, especially when the graphs are large. In this work, we propose a neural approximation of the GW distance, called NeuralGW. In NeuralGW, we use a combination of a graph isomorphism network and a transformer to represent the nodes of two graphs as two sets of vectors, treated as two discrete distributions, on which we compute multiple maximum mean discrepancy values given by different kernels. We then use a multilayer perceptron to convert the vector formed by these values into a single value, which is the prediction of the GW distance. Once trained, the model allows for efficient inference, enabling fast structural comparisons between graphs across diverse domains. We also provide a theoretical guarantee for the generalization ability of NeuralGW. Experiments demonstrate the effectiveness and practical applicability of our approach on real-world datasets, in comparison to baselines.
Abstract: Adapting CLIP for anomaly detection on unseen objects has shown strong potential in a zero-shot manner. Existing methods typically rely on a single textual space to align with visual semantics across diverse objects and domains. The indiscriminate alignment hinders the model from accurately capturing anomaly semantics. We propose TokenCLIP, a token-wise dynamic framework that aligns each visual token with learnt textual subspaces corresponding to its visual characteristics. However, explicitly assigning a unique learnable textual space to each token is computationally intractable and prone to insufficient optimization. We instead expand the token-agnostic textual space into a set of orthogonal subspaces, and then dynamically assign each token to a subspace combination guided by semantic affinity, which jointly supports customized and efficient token-wise adaptation. To this end, we formulate dynamic alignment as an optimal transport problem, where all visual tokens in an image are transported to textual subspaces under the cross-modal cost matrix. The marginal constraint and minimal cost objective of OT ensure sufficient optimization across subspaces and encourage them to focus on different semantics. Solving the problem yields a transport plan that adaptively assigns each token to semantically relevant subspaces. Extensive experiments show that the textual subspaces naturally specialize in different semantics, such as foreground and background, to promote fine-grained anomaly learning. The comparison between baselines shows the superiority of TokenCLIP.
Abstract: Recent 3D foundation models, such as DUSt3R, MASt3R, VGGT, \pi^3, and Depth Anything 3, provide strong feed-forward depth and pose estimates on pinhole imagery, but degrade sharply under fisheye camera geometry. We show that this failure is partly caused by a pinhole camera bias in the positional encodings of pretrained 3D foundation models, and propose RayTun3R, a lightweight camera adaptation approach. It keeps the pretrained network fixed and adapts only lightweight components tied to token position and camera geometry. RayTun3R learns parameter-efficient residual corrections to absolute and rotary positional encodings, together with parameter-free tokenization and corrections to prediction-grid coordinates that remove residual pinhole assumptions. The resulting adapter contains only 10,752 trainable parameters and can be learned from a short temporal segment using geometric losses. Once adapted, RayTun3R transfers effectively to the remaining frames of the sequence without incurring additional runtime costs. Across diverse fisheye datasets with fields of view from 110^\circ to 200^\circ, our adapter reduces rotation error by 2-12× relative to the unadapted model, outperforms LoRA while using ~14× fewer trainable parameters, improves pose over adaptation-free baselines while avoiding their multi-view inference cost, and remains competitive on depth accuracy.
Abstract: We study \left(\epsilon,\delta\right)-differentially private algorithms for the problem of approximately computing the top singular vector of a matrix A\in\mathbbR^n× d where each row of A is a datapoint in \mathbbR^d. Following Dwork-Talwar-Thakurta-Zhang (STOC 2014), we consider the privacy model where neighboring inputs differ by one single row. We give a novel algorithm that achieves beyond-worst-case guarantees for input matrices with low coherence, which is a structural property of matrices in many applications, including but not limited to i.i.d. data. Our algorithm contributes to the extensive literature on private power iteration methods, where we introduce a new filtering technique which adapts to this coherence parameter. Our work departs from and complements the work by Hardt-Roth (STOC 2013) which achieves beyond-worst-case guarantees for the more restrictive privacy model where neighboring inputs differ in one single entry by at most 1.
Abstract: , a phenomenon in which differential privacy guarantees improve when releasing only synthetic data rather than the private generative model itself. Recent work by Pierquin et al. (2025) established the first formal amplification guarantees for linear generators, but they apply only in asymptotic regimes where the model dimension far exceeds the number of released synthetic records, limiting practical relevance. In this work, we show that this restriction is not fundamental and establish that, surprisingly, for linear generators, . To prove this result, we develop new analytical tools for characterizing the privacy loss induced by synthetic data release. In particular, we leverage a sufficient-statistics reduction to characterize privacy leakage, and introduce a novel criterion that upper bounds Rényi divergences in terms of Fisher information. Our analysis provides structural insights that may guide the development of tighter privacy guarantees for more complex release mechanisms. Finally, experiments with variational autoencoders trained using DP-SGD further support the relevance of our theory beyond the linear setting.
PaperID: 3008, Poster
Authors:
Qingyun Li, Ruoyun Li, Yunqi Li, Yunjin LiAbstract: A Shellsort gap sequence is a multi-scale relaxation schedule. Each pass at gap h eliminates inversions at length scale h, and the sequence cools the array from coarse to fine, analogous to a multigrid V-cycle or simulated-annealing schedule. This view suggests that the right gap sequence should depend on the input's disorder spectrum: the inversion rate measured across logarithmic strides. Classical sequences such as Ciura, Sedgewick, and Tokuda are instead fixed schedules, applying the same coarse-to-fine relaxation to every input. We test this spectrum-matching view by formulating gap selection as a Markov decision process whose state is the current disorder spectrum and whose action selects the next gap. A small PPO-trained policy improves comparison count by 15-16% over Sedgewick across array sizes N=128 to N=16,384 on a seven-family input mixture, with training under 10 minutes on a consumer CPU. But the main value of the learned policy is not only that it improves Shellsort: it provides a probe into what effective gap schedules are doing. The learned schedules organize into three distinct relaxation regimes: gradual Brownian-like cooling for flat-spectrum inputs, faster Lévy-flight-like cooling for coarse-scale structure, and direct fine-scale relaxation for inputs whose disorder is already local. Spectrum statistics predict schedule statistics, training only on flat-spectrum inputs recovers a dense Ciura-like regime, and an ablation that replaces the adaptive policy with one global learned sequence reduces the gain from 15.6% to 0.9%. Input adaptivity, not the discovery of another static sequence, accounts for the gain. The marginal spectrum has blind spots: pipe-organ inputs share the random spectrum but their bipartite structure becomes visible only after the first pass, so sequential adaptation provides an additional margin. For known target distributions the policy can be used offline as a schedule compiler: run it on representative inputs, extract a per-distribution schedule, and hardcode that fixed sequence as a drop-in replacement for the existing static gaps with no neural runtime overhead. Together, these results recast Shellsort gap design as spectrum-matched scheduling.
Authors:
Tamim Zoabi, Ameen A Ali, Liran Ringel, Lior WolfAbstract: Discrete diffusion language models enable parallel token generation, offering a pathway to low-latency decoding. However, selecting tokens independently by marginal confidence limits effective parallelism: tokens that appear reliable in isolation can form incompatible configurations when several positions are updated at once. We introduce a training-free decoding framework that coordinates these parallel updates. At each forward pass, the method assigns a commit score to each masked position and refines these scores using pairwise interactions derived from the model's predictive distributions. A variational relaxation yields a simple fixed-point update that suppresses conflicting simultaneous commitments within a single forward pass.This mechanism allows the decoder to commit more tokens in parallel while maintaining competitive generation quality. The method is lightweight, requires no auxiliary model or retraining, and drops into existing diffusion decoding pipelines without modification. Experiments on reasoning and code-generation benchmarks show consistent improvements in the quality-latency trade-off.
Abstract: Large language models are typically deployed as monolithic systems, requiring the full model even when applications need only a narrow subset of capabilities, e.g., code, math, or domain-specific knowledge. Mixture-of-Experts (MoEs) seemingly offer a potential alternative by activating only a subset of experts per input, but in practice, restricting inference to a subset of experts for a given domain leads to severe performance degradation. This limits their practicality in memory-constrained settings, especially as models grow larger and sparser. We introduce EMO, an MoE designed for modularity—the independent use and composition of expert subsets—without requiring human-defined priors. Our key idea is to encourage tokens from similar domains to rely on similar experts. Since tokens within a document often share a domain, EMO restricts them to select experts from a shared pool, while allowing different documents to use different pools. This simple constraint enables coherent expert groupings to emerge during pretraining using document boundaries alone. We pretrain a 1B-active, 14B-total EMO on 1T tokens. As a full model, it matches standard MoE performance. Crucially, it enables selective expert use: retaining only 25% (12.5%) of experts incurs just a 1% (3%) absolute drop, whereas standard MoEs break under the same setting. We further find that expert subsets in EMO specialize at semantic levels (e.g., domains such as math or code), in contrast to the low-level syntactic specialization observed in standard MoEs. Altogether, our results demonstrate a path toward modular, memory-efficient deployment of large, sparse models and open new opportunities for composable architectures.
Abstract: The validity of statistical inference depends critically on how data are collected.When data gathered through active data collection (ADC) are reused for a post-hoc inferential task, conventional inference can fail because the sampling process is adaptively biased toward regions favored by the collection strategy. This issue is especially pronounced in black-box optimization, where sequential model-based optimization (SMBO) methods such as the tree-structured Parzen estimator (TPE) and Gaussian process upper confidence bound (GP-UCB) preferentially concentrate evaluations in promising regions. We study statistical inference on actively collected data when the inferential target is constructed in a data-dependent manner after data collection. To enable valid inference in this setting, we propose post-ADC inference, a framework that accounts for the biases arising from both the active data collection process and the subsequent data-driven target construction. Our method builds on selective inference and provides valid p-values and confidence intervals that correct for both sources of bias. The framework applies to a broad class of ADC processes by imposing only assumptions on the observation noise, without requiring any assumptions on the underlying black-box function or the surrogate model used by the SMBO algorithm. Empirical results also show that post-ADC inference provides valid inference for data collected by GP-UCB and TPE.
Authors:
Abdul Moeed, Stefan Schrod, Martin Rohbeck, Marc J Bonder, Pavlo Lutsik, Oliver Stegle, Daniel DimitrovAbstract: Tissue graph counterfactuals ask how a cell's expression would change under altered spatial contexts. Such queries are central to predicting cell behavior in tissues, but lack a unified definition, with existing methods targeting specific intervention types or treating cells as i.i.d. In this work, we first formalize tissue graph counterfactuals as a class of spatial interventions that either rewire connections between cells (edge perturbation) or modify the expression of their neighbors (node perturbation). We then introduce Cellina, a modeling framework that uses supervised disentanglement to decompose a cell's intrinsic state from its spatial context, using the latter as a conditioning input for counterfactual predictions. Across benchmarks spanning 2.5 million spatially-resolved cells in colorectal cancer and mouse brain, Cellina outperforms spatially-informed and non-spatial competitors in counterfactual predictions, disentanglement, and scalability. Additionally, we show that Cellina reveals biologically distinct cancer subdomains in an unsupervised manner and enables targeted neighbor perturbation simulations.
Abstract: We study offline constrained reinforcement learning with general function approximation in discounted constrained Markov decision processes. Existing methods either require full data coverage for evaluating unsupported intermediate policies, are not oracle efficient, or requires the knowledge of data-generating distribution for policy extraction. We propose PDOCRL, an oracle-efficient primal-dual algorithm based on a decomposed linear-programming formulation. The decomposition makes the policy an explicit optimization variable, avoiding policy extraction through the unknown distribution \mu_D. We show that naive restricted saddle-point formulations may have spurious saddle points, so realizability of an optimal solution alone is insufficient. We then show that, under a stronger but explicit realizability assumption, every restricted saddle point is optimal, avoiding the regularization and auxiliary function classes used in prior LP-based analyses. These ingredients yield PDOCRL, an oracle-efficient primal-dual algorithm that computes a near saddle point of the empirical decomposed LP Lagrangian and returns a near-optimal, near-feasible policy with a \widetilde\mathcal O(\epsilon^-2) sample guarantee under partial coverage, without access to \mu_D or a reference policy. Empirically, PDOCRL is competitive with strong baselines on standard offline constrained RL benchmarks.
Abstract: Conditional Flow Matching (CFM) unifies conventional generative paradigms such as diffusion and flow-based models. Interaction Field Matching (IFM) is a recently proposed framework that generalizes Electrostatic Field Matching (EFM), rooted in Poisson Flow Generative Models (PFGM). While both frameworks define generative dynamics, they start from different objects: CFM specifies a conditional probability path in data space, whereas IFM specifies a physics-inspired interaction field in an augmented data space. This raises a basic question: IFM. Specifically, we construct mappings between CFM and forward-only IFM and show that they induce the same generative dynamics. We further show that general IFM is strictly more expressive: it includes EFM and other interaction fields that cannot be realized within the standard CFM formulation. Our findings suggest developing generative models using both interpretations rather than treating them separately. Moreover, they highlight a novel direction for generative modeling based on backward-oriented field lines, which lies outside the conventional CFM formalism and may lead to new generative properties.
PaperID: 3015, Poster
Abstract: In long-term multi-session neural recordings, the same latent population dynamics are often observed through changing subsets of neural channels due to electrode drift, signal instability, or session-dependent recording variability. This creates a partially observed system-identification problem: each session provides only an incomplete view of a shared dynamical system, and missing channels are not merely absent inputs but unobserved variables coupled to the same latent dynamics. Existing approaches often rely on observation-pattern-specific mapping, which treats missing channels passively and leaves the shared latent dynamics weakly constrained. We propose CHAMA (CHannel-Aware Masked Attention), a mask-conditioned predictive representation learning framework for identifying shared neural dynamics under heterogeneous channel availability. CHAMA introduces a single global observation interface conditioned on channel-availability masks, allowing observed and unobserved channels to jointly constrain latent dynamics. The key idea is bidirectional: missing channels impose structural constraints on the latent representation, while the learned latent dynamics enable principled recovery of missing activity. To improve identifiability under sparse observations, CHAMA aggregates causal temporal windows, effectively trading temporal context for missing spatial information. Across synthetic dynamical systems and real multi-session neural recordings, CHAMA improves missing-channel completion, forecasting, latent dynamical fidelity, and downstream behavioral decoding over conditional adapter, zero-padding, and masked-autoencoding baselines. These results show that dynamics-aware recovery preserves globally consistent and behaviorally relevant structure beyond point-wise reconstruction accuracy. The source code is available at: \urlhttps://anonymous.4open.science/r/CHAMA-D691.
Abstract: Scientific research proceeds through iterative cycles of hypothesis generation, experiment design, execution, and revision, often requiring researchers to explore multiple competing directions as evidence accumulates and priorities shift. LLM agents can automate parts of this process, but existing agents either concentrate reasoning within a single research thread or coordinate through a central planner with fixed objectives. As a result, they struggle to sustain parallel exploration across research directions or reorganize as promising and unproductive directions emerge over time. We introduce AutoScientists, a decentralized team of AI agents for long-running computational scientific experimentation. Rather than following decisions from a central orchestrator, agents independently interpret a shared experimental state, self-organize into teams around research directions, critique and filter proposals with a discussion phase before committing experimental compute, and exchange both successful and failed findings across teams to avoid redundant exploration. Under matched experimental budgets, AutoScientists outperforms prior agentic systems across biomedical machine learning, language-model training optimization, and protein fitness prediction. On BioML-Bench, spanning biomedical imaging, protein engineering, single-cell omics, and drug discovery, AutoScientists achieves a mean leaderboard percentile of 74.4% across 24 tasks, improving over the strongest prior biomedical agent by +8.33%. On GPT training optimization, AutoScientists reaches a target validation bits-per-byte 1.9x faster than autoresearch and continues discovering improvements from a stronger starting champion where the single-agent approach finds none (7 vs. 0 accepted improvements). On ProteinGym fitness prediction, AutoScientists discovers a method for ACE2-spike binding that improves over current state-of-the-art model by +12.5% Spearman correlation. Applied without modification to all 217 ProteinGym assays, the same method improves over the prior state of the art by +6.5% in Spearman correlation. Code: https://anonymous.4open.science/r/autoscientists_anonymize-D10C
PaperID: 3017, Poster
Abstract: Discovering causal structures in multivariate time series (MTS) is critical in domains such as finance and neuroscience, yet remains difficult in practice due to three intertwined obstacles. First, MTS are routinely corrupted by missing values, and naive pipelines can distort the underlying data distribution, leading to false causal graphs. Second, recovering instantaneous links is already hard on complete data, and missingness further entangles imputation with graph learning. Third, real world data often exhibits heteroscedastic noise, whose variance depends on both instantaneous and lagged causes, violating the assumptions of standard causal discovery algorithms. Existing methods fail to address these challenges concurrently, typically assuming complete data, homoscedastic noise, or purely time lagged relationships. To address these gaps, we introduce Cheesefill, an optimal transport framework for causal structure learning under missingness. Cheesefill parameterizes causal mechanisms with conditional normalizing flows to capture heteroscedastic noise, and learns a stochastic correction map that refines naive imputations into trajectories consistent with the inferred mechanisms. Theoretically, we show that optimizing over stochastic correction maps is equivalent to a conditional kernel reformulation of the Kantorovich problem; the neural parameterization used by Cheesefill then yields a tractable objective. Extensive experiments on synthetic and real data show that Cheesefill consistently outperforms existing baselines across different missingness mechanisms.
Authors: Haochen Cai, Xian Yu
Abstract: Benders decomposition (BD) is a widely used solution approach for solving two-stage stochastic programs arising in real-world decision-making under uncertainty. However, it often suffers from slow convergence as the master problem grows with an increasing number of cuts. In this paper, we propose Reinforcement Learning for BD (RLBD), a framework that adaptively selects cuts using a neural network-based stochastic policy. The policy is trained using a policy gradient method via the REINFORCE algorithm. We evaluate the proposed approach on a two-stage stochastic electric vehicle charging station location problem and compare it with vanilla BD and LearnBD, a supervised learning approach that classifies cuts using a support vector machine. Numerical results demonstrate that RLBD achieves substantial improvements in computational efficiency and exhibits strong generalization to problems with similar structures but varying data inputs and decision variable dimensions.
PaperID: 3019, Poster
Abstract: In densely checkpointed language model training, a capability can remain flat and then appear within a narrow window of steps. Behaviour alone cannot tell whether internal structure changes just as abruptly or instead lags earlier structural reorganisation. We study this relationship across 150 main Pythia 70M-410M trajectories (3 scales × 5 tasks × 10 seeds), using seed replication to estimate trajectory-level event order. Tracking the spectra of mean-ablation patching matrices throughout training, we find two task-determined dimensions. The discontinuity ratio \rho separates tasks where behaviour changes more abruptly than the structural spectrum from tasks where the two co-evolve, reconciling competing accounts of emergence within this model range. The structure-behaviour gap \tau shows that spectral completion precedes behavioural emergence in 94% of main trajectories, with median absolute leads of 2,000-12,200 training steps (11×-206× in step ratio). Together these quantities support offline trajectory auditing: \rho computed from 28% of training predicts final \rho at r = 0.93, while \tau identifies a pre-emergence checkpoint for circuit inspection; at those checkpoints, Pythia-410M causal ablations show that early-prominent components carry 86-99% of post-emergence capability across the five tasks tested causally.
PaperID: 3020, Poster
Abstract: Large Language Models (LLMs) are increasingly deployed in real-world settings where alignment with diverse human values is essential. However, existing alignment methods are often costly, obscure underlying value heterogeneity, and offer limited interpretability. In this work, we move toward multi-human-value alignment through inference-time intervention by investigating how different values are internally represented within LLMs. We propose a probing-based value localization framework that identifies value-sensitive components, i.e., value heads, which mediate model behavior across multiple moral and normative dimensions such as harmlessness, honesty, and helpfulness. Analyses across multiple LLM families reveal that value heads are universally sparse and exhibit structured interactions, including conflicts among certain values. These components demonstrate clear functional specialization: selectively ablating value heads induces substantial and value-specific behavioral changes. Building on these findings, we introduce an inference-time intervention strategy that enables controllable adaptation of LLMs to single or multiple human value systems, with theoretical support for its effectiveness. Experiments on human-value benchmarks demonstrate that our method improves flexible and faithful value alignment while maintaining overall model performance, highlighting mechanistic interpretability as a foundation for socially aware, pluralistic AI.
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective post-training method for improving the reasoning abilities of Large Language Models (LLMs). However, existing methods mainly apply uniform optimization constraints across all tokens, ignoring their heterogeneous roles. Prior work shows that high-entropy tokens are closely tied to reasoning, while low-entropy tokens primarily encode factual knowledge, and recent approaches attempt to exploit this distinction by isolating token updates via masking or asynchronous training. We argue that such isolation breaks the sequential dependency structure of autoregressive generation, leading to suboptimal learning. To address this, we propose Archer, an entropy-aware RLVR framework with dual-token constraints that preserves joint optimization while modulating update strength across token types. Our method introduces response-level entropy normalization for stable token classification and applies differentiated clipping ranges and KL regularization to encourage exploration on reasoning tokens while preserving knowledge tokens. Experiments on mathematical reasoning and code generation benchmarks show that Archer consistently outperforms strong baselines across multiple model scales, improving both pass@1 and pass@K performance. These results highlight the importance of respecting sequence-level dependencies when designing fine-grained RL optimization strategies for LLMs.
PaperID: 3022, Poster
Abstract: Video diffusion transformers rely on self-attention over long spatio-temporal token sequences, making inference expensive even when sparse attention reduces the number of computed blocks. This paper studies a systems bottleneck that appears in dynamic sparse video attention: per-head sparsity can vary widely across denoising steps, layers, and prompts, so equal-head parallel execution leaves some GPUs waiting for dense heads assigned to other GPUs. We propose AsymHP, a runtime system that uses the previous denoising step's per-head sparse density to place a non-uniform number of heads on each GPU. AsymHP combines this lightweight online density estimate, a cost model for mask construction and sparse attention computation, and an asymmetric head redistribution primitive that avoids padding variable-size head shards into symmetric collectives. In our evaluation on H100 GPUs, AsymHP improves sparse attention latency by up to 1.54× without changing the underlying sparse-attention algorithm.
PaperID: 3023, Poster
Abstract: Vector Quantization (VQ) is fundamental to discrete visual tokenizers that power modern autoregressive and masked image generation models. While recent shared-projection codebook methods have substantially advanced codebook utilization, training stability remains a critical and underexplored challenge. We argue that the root cause lies in the entanglement of the Encoder--Decoder and Codebook training: because neither module can reliably fulfill its own responsibility in isolation, the system can only function when the two subsystems happen to cooperate---a fragile condition that breaks down precisely when training is most stressed. We propose StableVQ, which revisits the proper learning objective of each module and resolves the problems that arise when each is trained to fulfill its own role independently. Concretely, (1) Dynamic STE corrects the instability in the Encoder's learning objective, enabling it to robustly optimize the reconstruction space under discrete regularization even when codebook utilization is low. (2) Region VQ Loss reconceives the Codebook's learning objective so that it can independently guarantee full tracking of the encoder output distribution, without relying on encoder oscillations to drive activation. (3) Decoupled Schedule recognizes that the distinct responsibilities of the Encoder--Decoder and the Codebook demand distinct optimization dynamics, and assigns each an independent learning rate schedule to ensure robust system-level behavior. Built on top of shared-projection codebooks, StableVQ is lightweight and introduces no learnable parameters. Experiments on ImageNet demonstrate consistent improvements in training stability, codebook utilization, and reconstruction quality across diverse codebook sizes and initialization settings.
Abstract: Large language models (LLMs) trained on next-token prediction exhibit remarkable in-context learning (ICL) abilities, yet the representations that support ICL remain poorly understood. We consider such representations in a controlled setting: prompting LLMs with data emitted from hidden Markov models (HMMs) and probing for the corresponding belief state -- the posterior distribution over the HMM's hidden states given the observed token history. Across six open-source LLMs prompted with data from 40 HMMs selected for non-trivial belief structure, we find that belief states are linearly decodable from residual stream activations, with peak probe R^2-values ranging from 0.83 to 0.99 across HMM and LLM combinations, typically occurring at early-to-middle layers. To establish functional relevance, we intervene directly on the probe-identified subspace via patching and steering, resulting in downstream prediction quality on the order of the untampered model, while controls degrade substantially. Together, these results provide representation-level evidence that ICL in pre-trained LLMs approximates optimal Bayesian prediction over a context-inferred HMM. More broadly, our findings extend prior results linking input-distribution structure to activation geometry: from toy networks trained explicitly on HMM data to production-scale transformers trained on natural language.
Abstract: Randomized controlled trials typically assume that prognostic covariates are known and available at no cost. In practice, obtaining high-dimensional pretreatment data is costly, forcing a trade-off between covariate-adaptive precision and a measurement budget. We introduce Dynamic Adaptive Rerandomization via Thompson Sampling (DARTS), which treats covariate acquisition as a sequential optimization problem embedded within a design-based causal inference task. A budgeted combinatorial Thompson sampler learns which covariates are most prognostic across successive batches; selected covariates then drive rerandomization and regression adjustment to reduce batch-level average treatment effect variance. Our primary theoretical contribution is a decoupling result: adaptive covariate selection based on past batches preserves batch-level randomization validity, and the cumulative inverse-variance weighted estimator achieves at least nominal asymptotic coverage. We further derive a Bayes risk bound for the acquisition layer that matches the minimax lower bound up to logarithmic factors. Empirically, DARTS systematically concentrates budget on informative features, significantly closing the efficiency gap to oracle designs while maintaining strict inferential validity.
PaperID: 3026, Poster
Abstract: Multi-modal object Re-Identification (ReID) aims to achieve robust all-weather perception by leveraging complementary information across heterogeneous modalities. However, existing methods still suffer from limitations in both feature fusion and feature alignment. In terms of feature fusion, most approaches assume equal importance across modalities and do not explicitly account for the quality differences among modalities in dynamic environments. As a result, the fused representations tend to be suboptimal and less discriminative in complex scenarios. In terms of feature alignment, existing methods typically rely on point-wise feature matching while ignoring relational structures among samples, which suppresses discriminative information and introduces cross-modal inconsistencies. To address these limitations, we propose LAMP, a unified framework that rethinks multi-modal interaction through language-modulated control and geometry-consistent alignment. We first adopt a Semantic-Assisted Feature Encoder (SAFE), which incorporates textual semantics into visual features to enhance local discriminative cues and provide more informative representations for subsequent processing. Building upon this, the Language-Modulated Dynamic Routing (LMDR) module generates modality-aware gating signals prior to feature fusion, enabling the adaptive adjustment of modality weights under challenging conditions. Furthermore, to address cross-modal inconsistencies, we introduce a Fused Geometric Prototype Alignment (FGPA) objective based on Fused Gromov-Wasserstein (FGW) optimal transport. Instead of enforcing strict point-wise alignment, FGPA establishes cross-modal consensus at the identity-level relational geometry, thereby preserving discriminative structures while mitigating inconsistencies across modalities. LMDR improves the stability of multi-modal fusion, while FGPA enhances structural consistency in cross-modal alignment. Together, they improve the discriminative capability of the model. Extensive experiments on three multi-modal object ReID benchmarks demonstrate the effectiveness of our method.
PaperID: 3027, Poster
Authors: Nico Policzer, James Ho, Joseph Soo, Xin Tang
Abstract: Understanding how neural activity encodes behavioral and sensory variables is a key challenge in systems neuroscience. This requires models that go beyond predicting activity from labels and instead support source-conditioned counterfactual editing: given a recorded source response, we should be able to edit one behavioral or stimulus variable and predict how neural activity would change while preserving other observed factors and source-specific variability not explained by the labels. Here, we introduce , a disentangled representation learning framework for counterfactual modeling of neural activity. Neural-DISCO partitions the latent space into label-specific components, each associated with an observed behavioral or stimulus variable, together with a residual latent that captures variability not supplied through observed labels. Across simulated and real neural recording datasets, we show that this structure enables selective manipulation of individual labels to generate counterfactual neural responses while preserving the remaining content of the original activity. Further, to gain insight into how behavioral and stimulus variables are encoded at single-neuron resolution, we pair this disentangled framework with feature attribution methods that identify which neurons are most important for each label. We find that disentanglement enhances the recovery of known functional cell roles in both synthetic and real data. Together, these results support Neural-DISCO as a framework for controlled counterfactual modeling of neural activity and provide a practical interface for probing how neural populations encode behavior and sensory information.
PaperID: 3028, Poster
Abstract: Autoregressive image generation has recently emerged as a competitive visual generation paradigm, yet its generation order is often predefined or heuristic, such as next-patch prediction, random-order prediction, or one-step generation for specific scales. We observe that these strategies can cause models to overlook important structural dependencies in images, thereby limiting both structural robustness and reasoning capability. In this work, we propose a context-aware autoregressive generation framework, where the model dynamically determines the next token location based on global context and current local information. Rather than following a fixed or random order, our method allows the generation process to adapt to image structure and semantic dependencies. Extensive experiments show that context-aware generation significantly improves structural robustness across different image AR paradigms, including masked autoregressive and scale-wise autoregressive models. Beyond quantitative improvements, we further observe emerging reasoning properties: the model can better understand complex semantic instructions and generate images that more faithfully follow the given guidance. These findings suggest that generation order is a critical yet underexplored factor in image autoregressive modeling, and that context-aware decoding can unlock stronger reasoning-aware visual generation capabilities.
PaperID: 3029, Poster
Abstract: We study how to choose a probability distribution, or lottery, over m alternatives for n agents when each agent's acceptable lotteries form an unknown linear halfspace of the simplex. The learner observes only binary accept/reject membership queries. Prior work asks whether there exists a lottery accepted by every agent; we study the relaxation that minimizes the number of agents who reject the lottery. We first show that this problem is APX-hard even when all halfspaces are explicitly known and scores are binary. We then give exact algorithms under additional structure: fixed m, and a margin condition under which an optimal lottery can be replaced by one supported on \mathcalO(\log n/\gamma^2) alternatives. With limited queries, uniform sampling of agents gives additive population guarantees after recovering the sampled agents' normalized halfspaces. For arbitrary m, we give a polynomial-time LP based on a one-sided margin penalty and prove a population guarantee depending on the number of true violations and the number of agents near their thresholds. Finally, we prove a tight deterministic lower bound: without a stochastic link between queried and unqueried agents, querying only K distinct agents cannot guarantee additive gap below n-K, and this gap is achievable.
Abstract: Embodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destinations reliably in the real world. However, progress remains constrained by the lack of scalable, high-fidelity, and physically grounded interactive environments. Although real-world scanned datasets offer visual realism, they are limited by scale. In contrast, synthetic simulators scale more easily but often exhibit large sim-to-real gaps. We introduce Image2Sim, a real-time neural simulation framework that constructs high-quality interactive environments from posed RGB-D image sequences. The central idea is to decouple 3D spatial anchoring from photorealistic observation synthesis. For scene construction, Image2Sim uses a feed-forward feature Gaussian model that lifts posed RGB-D observations into a 3D feature-Gaussian representation in a single pass. For rendering, we propose a Geometry-Aware One-Step Pixel Flow model that transforms sparse and noisy Gaussian projections into high-quality panoramic RGB-D observations. Image2Sim also serves as a fully automated embodied data engine that generates high-fidelity observations, executable actions, and diverse navigation instructions at scale. It converts large collections of videos and images into near 20K interactive scenes and synthesizes more than 10 million navigation training samples. Navigation models trained entirely in these neural environments achieve strong improvements on major benchmarks and transfer effectively to real-world zero-shot settings. These results suggest that scalable neural simulation can serve as a practical training substrate for embodied navigation at scale.
PaperID: 3031, Poster
Abstract: Current parameter-efficient fine-tuning (PEFT) methods like Low-Rank Adaptation (LoRA) are structurally agnostic, applying uniform configurations across all layers, which overlooks their vast functional heterogeneity. We propose Structure-Aware LoRA (SA-LoRA), a new framework that automatically tailors fine-tuning intensity to each layer's intrinsic complexity. Specifically, we leverage the Stable Rank as a spectral metric to align the adaptation magnitude of each layer with its pre-trained spectral structure, enabling an automated, principled allocation of learning capacity that directly addresses the structural agnosticism of existing methods. To enhance adaptability and robustness, we introduce a hybrid calibration mechanism that fuses the task-agnostic prior with task-specific gradient feedback, underpinned by a budget-conservation principle to ensure stability. Extensive experiments demonstrate that SA-LoRA consistently outperforms strong PEFT baselines, often achieving state-of-the-art performance with enhanced stability. The code is available at https://anonymous.4open.science/r/SA-LORA-6809.
Authors:
Anh Hao Vo, Phu L Nguyen, Khoa Vo, Sieu Tran, Duc Nguyen, Ngo X Cuong, Nghi Bui, Anh Nguyen, Duy M. H. Nguyen, Ngan LeAbstract: Multimodal Large Language Models (MLLMs) have achieved remarkable performance on 2D visual tasks, yet enhancing their spatial intelligence for real-world applications such as Autonomous Vehicles (AV) remains an open challenge. Existing geometry-aware MLLMs typically rely on auxiliary 3D models at inference time, introducing pipeline complexity and the risk of cascading failures. In this paper, we present OmniSpace, a simple yet effective plug-and-play paradigm for geometry-aware spatial reasoning from purely 2D observations. Motivated by our finding that current MLLMs are bottlenecked by weak cross-view correspondence and depth estimation, OmniSpace introduces a Camera Pose Injector, a Multi-view Epipolar Attention module, and a 3D Geometric Distillation objective that jointly address these two limitations by transferring geometric knowledge into the model. Extensive experiments show that OmniSpace surpasses existing methods on planning benchmarks (nuScenes, Bench2Drive), risk detection (nuInstruct), language (Omnidrive), and generalization (DriveBench).
PaperID: 3033, Poster
Authors: Hamidreza Eivazi Kourabbaslou, Eslam Sharaawy, Stefan Wittek, André Hebenbrock, Raphael Ginster, Steffen Blömeke
Abstract: Battery degradation remains a central challenge in the development and broad application of sustainable energy technologies. Accurate degradation prediction is challenging, as battery aging emerges from complex and heterogeneous interactions among cycling behavior, operating conditions, and cell chemistry. Existing machine-learning approaches typically focus on deterministic end-of-life predictions or modeling degradation curves within fixed chemistry settings, leaving probabilistic state-of-health (SOH) trajectory modeling across heterogeneous chemistries underexplored. We introduce FlowBatt, a use-inspired conditional generative framework that adapts flow matching with a diffusion transformer (DiT) backbone to battery degradation trajectory modeling. FlowBatt models full SOH trajectories as probabilistic degradation processes conditioned on early-cycle capacity data. To strengthen this conditioning pathway, we use explainable AI analysis to diagnose and improve the encoder that maps early-cycle capacity matrices into conditioning vectors. We evaluate FlowBatt on five public benchmark settings spanning diverse chemistries and aging conditions, comparing it with supervised and diffusion-based trajectory models as well as established baselines for remaining-useful-life (RUL) prediction under shared data splits. FlowBatt achieves the best SOH prediction performance on four of five datasets, with errors below 2.5%, and delivers competitive RUL prediction performance. By generating multiple plausible trajectories, FlowBatt provides empirical uncertainty estimates, while also revealing calibration challenges under some dataset shifts. These results indicate that flow-matching-based trajectory modeling is a promising framework for probabilistic battery health prediction.
Authors:
Saar Huberman, Ron Mokady, Or Patashnik, Daniel Cohen-orAbstract: In text-conditioned generative models, images are represented and controlled through text prompts. In practice, images are generated from sequences of tokens derived from these prompts. However, the space of token sequences lacks a consistent structure: semantically similar images may correspond to sequences that differ in wording, ordering, and placement of concepts, while similar token sequences may encode very different semantics. This lack of structure makes it difficult to perform smooth transitions in this space, hindering applications such as image blending and continuous control of edits. We argue that this limitation stems not from the absence of semantic structure, but from misalignment between representations. To address this misalignment, we introduce Token-to-Token alignment, a framework that establishes explicit semantic correspondence between tokens across prompts. Our approach transforms prompts into a structured representation in which semantically corresponding concepts occupy compatible positions, and then aligns their token embeddings based on semantic similarity. Concretely, the method consists of two stages: a structural alignment that rephrases prompts into a shared structured form, followed by an embedding-level alignment that matches token representations across prompts. With this alignment in place, simple linear interpolation becomes a meaningful operation, producing smooth and coherent semantic transitions and enabling applications such as blending and continuous editing. Our results show that text embedding spaces in such models implicitly encode a continuous semantic structure that becomes accessible once representations are properly aligned, suggesting that semantic control can be achieved by organizing existing representations rather than modifying the generative model.
PaperID: 3035, Poster
Abstract: Image matting aims to estimate a continuous alpha matte that separates foreground objects from the background. However, existing deep matting methods still rely heavily on task-specific supervision for matting, including alpha mattes and trimaps, which are costly to annotate and limit their applicability in annotation-scarce scenarios. To address this limitation, we propose FreeMatte, a novel annotation-free framework that learns to predict high-quality alpha mattes without requiring any matting-specific annotations. The proposed framework formulates the construction of matting supervision as a foreground-driven wavefront propagation process. Starting from a coarse localization prior generated by an off-the-shelf segmentation foundation model, it solves an Eikonal equation constructed from the prior and the input image itself to obtain an arrival-time field that encodes how readily each pixel can be reached from the coarse foreground source. The arrival-time field is then translated into confidence-aware sparse supervision, which constrains the matting model to learn alpha prediction from the propagation structure encoded by the field. Extensive experiments across diverse benchmarks show that FreeMatte achieves performance competitive with task-supervised matting baselines under this challenging setting, ranking within the top two on 13 out of 20 evaluated dataset-metric pairs.
PaperID: 3036, Poster
Abstract: Scaling reinforcement learning (RL) to diverse multitask settings remains a central challenge. While recent advances in model-based RL achieve strong performance, they rely on planning and complex training pipelines, making it unclear which components are essential for scalability. We revisit this question and argue that the primary driver of scalable multitask RL is not model-based control, but . In particular, we show that combining predictive, model-based representations with high-capacity value function approximation is sufficient to achieve strong performance, without planning. We evaluate a simple model-free algorithm, MR.Q, that integrates auxiliary predictive objectives into a scalable actor-critic architecture. This approach outperforms a recent world-model-based method and a range of deep RL baselines across a diverse suite of multitask continuous control tasks, while significantly reducing computational overhead and improving wall-clock efficiency. We observe consistent improvements with increased model capacity and show through ablations that predictive representation learning is critical for performance.
PaperID: 3037, Poster
Abstract: Deploying autonomous driving systems requires robustness against long-tail scenarios that are rare but safety-critical. While adversarial training offers a promising solution, existing methods typically decouple scenario generation from policy optimization and rely on heuristic surrogates. This leads to objective misalignment and fails to capture the shifting failure modes of evolving policies. This paper presents ADV-0, a closed-loop min-max optimization framework that treats the interaction between driving policy (defender) and adversarial agent (attacker) as a zero-sum Markov game. By aligning the attacker's utility directly with the defender's objective, we reveal the optimal adversary distribution. To make this tractable, we cast dynamic adversary evolution as iterative preference learning, efficiently approximating this optimum and offering an algorithm-agnostic solution to the game. Theoretically, our analysis characterizes the idealized regularized game and derives a lower bound that connects adversarial training performance with real-world long-tail generalization. Experiments indicate that ADV-0 effectively exposes diverse safety-critical failures and greatly enhances the generalizability of both learned policies and motion planners against unseen long-tail risks.
PaperID: 3038, Poster
Abstract: Large language models deployed as personalized assistants must reason over long, evolving interaction histories. However, in long-term dialogue reasoning, relevant evidence is scattered across sessions, preferences may be revised over time, and standard long-context training fails to address these challenges under data scarcity and prohibitive computational costs. We propose StateTree, a data-driven RL pseudo-task that constructs a challenging auxiliary task from scarce dialogues with verifiable ground truth. StateTree augments multi-session dialogues with a tree-structured path-tracing task: key-value records are embedded across sessions to form a binary tree. Solving the task requires the model to traverse from root to leaf by retrieving records across sessions and comparing timestamps to resolve branches, then recover the hidden target question among distractor leaves. We apply curriculum RL training progressively increasing tree depth and introduce a compositional variant whose edges carry step-level reasoning fragments, training the model to compose partial cues into coherent queries. Trained on 10K-token contexts, StateTree generalizes to 128K tokens without full-length RL costs and exhibits capabilities including cross-session retrieval, temporal reasoning, knowledge update, and compositional multi-hop reasoning. StateTree outperforms both SFT and RL-based baselines while preserving short-context general reasoning. StateTree-7B achieves gains up to +23.60% on LongMemEval (128k), and StateTree-14B reaches 59.00% accuracy on LongMemEval, surpassing QwenLong-L1-32B (45.20%).
PaperID: 3039, Poster
Authors: Bai Sun, Zixuan Tang
Abstract: Medical image segmentation requires the precise alignment of macroscopic semantics and microscopic geometry, yet existing representation paradigms struggle to balance massive low-frequency semantic context with ultra-sparse, high-frequency boundary details. To address this fundamental information asymmetry, we introduce Fractal-G, a plug-and-play multi-scale fusion module that presents a continuous Grid-Graph-Grid feature reconstruction paradigm. Within this framework, an Uncertainty-Aware Fission Router dynamically translates regular image grids into a non-uniform heterogeneous graph, adaptively allocating dense microscopic nodes to complex boundaries while retaining sparse macroscopic nodes in homogeneous background regions. To ensure unbiased feature aggregation across these varying physical scales, an Area-Aware Continuous Neighborhood Graph explicitly incorporates physical node sizes into the topological message passing. Finally, a Topology-Aware Implicit Renderer projects the unstructured graph back into dense, artifact-free continuous feature fields. Extensive experiments on diverse medical image segmentation tasks (including skin lesions and polyps) demonstrate that Fractal-G easily integrable with mainstream backbones, consistently achieving state-of-the-art geometric fidelity and cross-resolution robustness with only 12% GFLOPs increase (U-Net based). Our code is publicly available at https://anonymous.4open.science/r/Fractal-G/.
PaperID: 3040, Poster
Abstract: While one-step text-to-image generators, a fast and deployment-friendly class of visual generative models, can produce samples with a single forward pass, existing preference-finetuning methods fail to simultaneously achieve efficient adaptation, reward-model agnosticism, and diverse generation. In this work, we propose Drifting Preference Optimization (DrPO), an online preference-finetuning method for one-step generators. Inspired by the recent work of Drifting Models, DrPO first uses the target reward model to construct an on-policy dipole reward model from positive and negative samples. This dipole model induces a preference drifting field in latent feature space, whose gradients are then used to optimize the generator without backpropagating through the target reward model. Empirically, we evaluate DrPO across multiple reward models and SD-Turbo and SDXL-Turbo one-step generators, including HPSv3 and GenEval as representative alignment benchmarks, and validate that DrPO can efficiently and effectively finetune one-step generative models. Preliminary experiments further suggest that this sample-based gradient synthesis can be extended to offline preference finetuning.
PaperID: 3041, Poster
Abstract: Deep-learning-based 3D MRI reconstruction is limited in practice by the GPU memory and compute cost of full-volume training. A common workaround is to train at lower resolution and run inference at higher resolution, but standard neural networks are resolution-specific and degrade substantially when evaluated outside the training setting. We instead use resolution-agnostic neural operators (NOs) and propose KIR-NO (k-space-to-Image Reconstruction Neural Operator), a coarse-to-fine, discretization-agnostic framework for accelerated 3D MRI reconstruction. KIR-NO is trained with full backpropagation on memory-feasible low-resolution volumes and applied zero-shot to higher-resolution volumes at inference, without fine-tuning. Its core building block is a 3D local neural operator based on discrete-continuous convolutions: convolutional filters are parameterized in a continuous basis (we propose and empirically compare two 3D bases, piecewise-linear and Morlet-wavelet) and sampled at arbitrary voxel resolutions while preserving the local inductive bias of convolution. KIR-NO integrates these learned operators with physics-based consistency: a k-space neural-operator head first refines the undersampled Fourier measurements, and image-space neural-operator cascades then refine the reconstruction through interleaved data-consistency updates. Across 3D knee (SKM-TEA) and brain (BraTS 2021) MRI benchmarks, KIR-NO consistently outperforms classical and CNN baselines: it improves over a state-of-the-art 3D CNN by 4.13 dB PSNR at half resolution under 4× acceleration on SKM-TEA, with gains persisting at 8× and 16×, and exceeds the same baseline by 13.63 dB under zero-shot full-resolution transfer on BraTS. Together, these results show that KIR-NO enables scalable, memory-efficient, and resolution-flexible 3D MRI reconstruction across anatomies.
Abstract: As synthetic images become increasingly realistic, reliable synthetic image detection techniques are of pressing need to prevent their misuse. Despite satisfactory in-distribution performance, deep neural network-based synthetic image detectors (SIDs) lack reliability in deployment and often fail in the presence of common covariate shifts, resulting in poor detection accuracy. To avoid the risk caused by potential errors, we adopt a selective classification (SC) strategy by allowing SIDs to abstain from making low confidence predictions. For practicality, we focus on post-hoc methods which perform confidence estimation on a given SID without retraining. However, we show that conventional logit-based confidence score functions (CSFs) exhibit pathological behavior under covariate shifts, leading to SC performance close to or even worse than random guessing. To address this, we propose a simple yet effective SC framework for tection (ReSIDe). First, we generalize the notion of logits to an SID's intermediate layers from a centroid matching perspective, extending the use of logit-based CSFs to any layer of an SID. Then, we introduce a preference optimization algorithm that aggregates confidence scores extracted from different layers to a final confidence estimate by minimizing an upper bound of the area under the risk-coverage curve (AURC). Extensive experimental results show that ReSIDe significantly boosts the SC performance of various logit-based CSFs under common covariate shifts, achieving up to 69.55% AURC reduction.
PaperID: 3043, Poster
Abstract: Reliable 6D object pose estimation is essential for dexterous manipulation but remains highly challenging due to severe visual occlusions from multi-finger interactions. Progress has been limited by two key factors: the lack of large-scale manipulation datasets and the inability of existing methods to handle high uncertainty under occlusion. To address these challenges, we introduce DexOPE, a large-scale dataset and a physics-guided pose estimation framework. DexOPE provides a multimodal dataset with 1.5k diverse objects, an order of magnitude larger than prior work, covering simulated and real-world grasping and manipulation. To handle pose ambiguity under occlusion, we propose a Physics-Guided Score Matching approach that models the pose posterior using score-based diffusion models rather than deterministic regression. Physical constraints from hand–object interactions, including nonpenetration and tactile contact, are incorporated as guidance during sampling, enabling physically plausible pose inference even under severe or complete occlusion. Extensive experiments show that DexOPE achieves state-of-the-art performance, significantly outperforming existing methods in robustness to occlusion and generalization to unseen objects.
PaperID: 3044, Poster
Authors: Abdulkadir Celikkanat, Andres Masegosa, Mads Albertsen, Thomas Nielsen
Abstract: Learning effective representations of DNA sequences is fundamental to genomic data analysis, especially in the presence of high sequence similarity and inter-species DNA sharing. Existing approaches rely on deterministic embeddings, such as k-mer statistics or representations from large language models, which fail to capture the inherent uncertainty in short and ambiguous DNA fragments. We propose UncertainGen, a lightweight probabilistic representation learning framework that embeds DNA sequences as distributions in a latent space. The learned embedding variance captures positional ambiguity, increasing for sequences compatible with multiple genomic clusters. We establish theoretical guarantees on embedding distinguishability, and show that the probabilistic formulation induces a data-adaptive metric that expands the effective representational space. We evaluate UncertainGen in the metagenomic binning task to demonstrate the practical benefits of uncertainty-aware representations. Experiments show that probabilistic embeddings consistently outperform deterministic k-mer and large genome foundation models, while remaining lightweight and efficient.
Authors: Mehrnaz Asadi, Sina Javadzadeh, Rahil Soroushmojdehi, Ali Mousavi, Terence Sanger
Abstract: Understanding how distributed brain regions coordinate during behavior requires models that are both predictive and interpretable. We introduce Behavior-Adaptive Connectivity Estimation (BACE), an end-to-end framework for learning behavior-conditioned directed effective connectivity from multi-region intracranial local field potentials (LFPs). BACE first encodes within-region dynamics with region-specific temporal encoders, then applies a learned adjacency matrix selected for each behavioral context, and finally forecasts future neural activity through a graph-conditioned autoregressive decoder. This design yields explicit region-level connectivity matrices whose edges are tied to predictive dynamics rather than post-hoc explanation. On controlled synthetic time series with known directed graphs, BACE recovers the ground-truth edge structure from forecasting alone. We then evaluate BACE on human deep-brain LFP recordings from three participants performing structured motor tasks. Across participants, BACE achieves strong neural forecasting while producing compact, behavior-specific directed graphs that can be inspected across task segments. Reliability analyses further support the stability of the inferred connectivity patterns. Together, these results position BACE as a practical framework for estimating behavior-adaptive effective connectivity from high-dimensional intracranial recordings, enabling interpretable hypotheses about how deep-brain networks reorganize during behavior.
PaperID: 3046, Poster
Abstract: Machine unlearning for large language models (LLMs) aims to selectively remove memorized content such as private data, copyrighted text, or hazardous knowledge, without costly full retraining. Most existing methods require a retain set of curated examples to prevent catastrophic degradation of general model utility, creating an extra data dependency that complicates deployment. We propose SHRED (Self-distillation via High-surprisal-only Retain-set-free Entropy Demotion), a retain-set-free unlearning method built on a key insight: not all tokens within a forget set instance carry memorized information equally. High-information tokens concentrate the model's memorized knowledge, while low-information tokens reflect general language competence. SHRED operates in two stages. (1) Selection: We perform a forward pass on a forget set instance, collect per-token autoregressive probabilities, and select the bottom-P (lowest probability, highest Shannon information) as forget positions; the remaining positions are retained as benign anchors. (2) Training: We construct modified KL targets that demote the memorized token's logit at forget positions while preserving the original distribution at benign positions. The model is then trained via a single top-K KL self-distillation objective that simultaneously drives forgetting and utility preservation. We evaluate SHRED across four standard unlearning benchmarks and demonstrate that it establishes a new Pareto-optimal trade-off between forget efficacy and model utility, outperforming retain-set-dependent methods. Our analysis shows that SHRED is robust against relearning attacks and membership-inference attacks, and it maintains stable utility even after many sequential unlearning runs.
Abstract: Off-policy, value-based reinforcement learning methods such as Q-learning are appealing because they can learn from arbitrary experience, including data collected by older policies or other agents. In practice, however, bootstrapping makes long-horizon learning brittle: estimation errors at later states propagate backward through temporal-difference (TD) updates and can compound over time. We propose (LQL), which introduces a principled backstop against compounding error when learning the optimal action-value function. LQL builds on a prior observation: any realized action sequence lower-bounds what the optimal policy can achieve in expectation, so acting optimally earlier should not be worse than following the observed actions for several steps before switching to optimal behavior. Our contribution is to turn this inequality into a practical stabilization mechanism for Q-learning by using a hinge loss to penalize violations of these bounds. Importantly, LQL computes these penalties using network outputs already produced for the TD error, requiring no auxiliary networks and no additional forward passes relative to Q-learning. When combined with multiple state-of-the-art methods on a range of online and offline-to-online benchmarks, LQL consistently outperforms both 1-step TD and n-step TD learning at similar runtime.
PaperID: 3048, Poster
Abstract: The rapid proliferation of large-scale machine learning models, such as LLMs, has driven remarkable progress across application domains. Nevertheless, their scale comes with significant deployment challenges and inherent redundancies that remain open challenges to this day. To this end, we propose the Tile-Reuse Index (T-ReX), a structured model compression method that identifies and indexes reusable, contiguous parameter groups (tiles) across a model's architecture. By learning a compact set of matching tiles, we compress the model's memory footprint while preserving its performance. Furthermore, by reconstructing dense layers on the fly at inference time, T-ReX compressed models remain efficient at runtime. In our experiments, we show that T-ReX outperforms existing compression methods across benchmarks by most closely matching the original model's performance. We furthermore show how post-training on top of T-ReX enables small footprint models to recover most of the performance loss introduced by compression.
PaperID: 3049, Poster
Authors: Hao Tan, Qimin Zhang, Lihuang Fang, Kebing Jin, Jinghui Qin
Abstract: AI-generated image detectors face significant challenges when deployed in real-world environments, particularly when test samples are produced by unseen generators that deviate from the source training distribution. Existing detectors often rely on a single type of evidence, such as semantic representations or low-level generation traces, making them vulnerable when the corresponding cue becomes unreliable. Moreover, adapting detectors to newly emerging generators usually requires target labels or retraining, which is costly and impractical. To address these challenges, we propose Dual-Evidence Disagreement-Constrained Adaptation (DEDCA), a test-time adaptation framework for generalized AI-generated image detection. DEDCA constructs a dual-evidence detector by combining CLIP-based semantic evidence with trace evidence extracted from shuffled entropy and residual statistics, and introduces a reliability-aware gate to perform sample-specific adaptive fusion. Our key idea is to exploit the disagreement between the two evidence streams as an unlabeled adaptation signal, enabling the detector to adjust to target-domain shifts without relying on target labels or blindly trusting a single prediction stream. Experiments show that DEDCA achieves superior performance on the GenImage benchmark by outperforming state-of-the-art AI-generated image detectors and test-time adaptation methods with 92.0% average accuracy.
Authors: Haobo Zhang, Zhenhua Xu, Junxian Li, Shangfeng Sheng, Dezhang Kong, Meng Han
Abstract: Protecting the intellectual property of open-weight large language models (LLMs) requires verifying whether a suspect model is derived from a victim model despite common laundering operations such as fine-tuning (including PPO/DPO), pruning/compression, and model merging. We propose AttnDiff, a data-efficient white-box framework that extracts fingerprints from models via intrinsic information-routing behavior. AttnDiff probes minimally edited prompt pairs that induce controlled semantic conflicts, captures differential attention patterns, summarizes them with compact spectral descriptors, and compares models using CKA. Across Llama-2/3 and Qwen2.5 (3B--14B) and additional open-source families, it yields high similarity for related derivatives while separating unrelated model families (e.g., >0.98 vs. <0.22 with M=60 probes). With 5--60 multi-domain probes, it supports practical provenance verification and accountability; our open-source implementation is available at https://anonymous.4open.science/r/AttnDiff-37A1/.
PaperID: 3051, Poster
Abstract: Modeling the mechanisms of chemical reactions is central to the prediction and design of chemical reactivity. We present Electron Bookkeeping Transformer (Bookkeeper), a model that simulates chemical reactions at mechanistic resolution by generating electron movement trajectories. Bookkeeper explicitly models each electron as a token and predicts its movement sequentially, following the arrow pushing formalism in chemistry. This ensures mass and charge conservation during both elementary step and reaction pathway predictions by construction. Bookkeeper achieves higher accuracy in predicting both individual mechanistic steps and full reaction pathways, shows strong generalization to out-of-distribution reaction types, and runs significantly faster than previous flow matching-based and large language model-based methods. The accuracy and efficiency make Bookkeeper a powerful framework for virtual screening of chemical reactions across diverse reagents, for validating computer-assisted synthesis planning predictions, and for providing mechanistic interpretability of reactivity; the utility of Bookkeeper as a tool will most likely grow as does its training corpus.
Abstract: Parameter-efficient fine-tuning (PEFT) has emerged as a crucial paradigm for adapting large language models (LLMs). However, standard PEFT methods often struggle in multi-task fine-tuning settings due to task interference and a limited parameter budget. Recent approaches incorporate the mixture-of-experts (MoE) architecture, referred to as mixture-of-parameter-efficient-experts (MoPE), to alleviate this issue by dynamically routing inputs to specialized experts. However, these methods remain based on parameterization, which may introduce structural redundancy and additional parameter overhead. To address these limitations, we revisit model adaptation from a perspective. Our analysis uncovers heterogeneity in frequency sensitivity across model layers and downstream tasks, indicating that adaptation should be frequency-aware rather than uniformly parameterized in the spatial domain. Motivated by this insight, we propose , a novel framework that unifies spectral parameterization with MoE via frequency-specialized experts and the learning of conjugate-symmetric complex coefficients. We conduct extensive experiments across multiple model families on a diverse range of tasks, including commonsense reasoning, math reasoning, image classification, and natural language understanding. Experimental results on 28 benchmarks show that FourierMoE consistently outperforms full fine-tuning (FFT) and 15 competitive baselines in both single-task and multi-task settings, while requiring significantly fewer trainable parameters, demonstrating superior effectiveness, efficiency, and adaptability.
PaperID: 3053, Poster
Abstract: We present PAI-Actor, a cinematic multi-character animation framework for character replacement in dynamic movie scenes. Unlike conventional animation systems that mainly drive a single static image or a single subject, our goal is to replace and animate multiple characters within real video clips while preserving the original scene dynamics, camera motion, and background content. This setting is particularly challenging because the generated characters must remain consistent with the source performance in motion and interaction, while also matching the surrounding background in lighting, shadow, composition, and overall cinematic appearance. To address this, we formulate multi-character animation as a structure-guided human recovery problem and build a movie-driven training pipeline from high-quality film data. Furthermore, to support practical cinematic production, we introduce a bidirectional-to-autoregressive distillation framework: we first train a bidirectional diffusion transformer for high-quality short-clip generation at 1080P resolution, and then distill it into an autoregressive video-to-video model for efficient inference and longer video generation. Experiments show that PAI-Actor enables high-fidelity multi-character animation with strong scene consistency, cinematic visual quality, and efficient long-form generation.
PaperID: 3054, Poster
Abstract: Many studies argue that the success of sharpness-aware minimization (SAM) is due to the implicit regularization of the gradient norm or the eigenvalues of the Hessian. However, they cannot fully explain why the flatness-based measures do not always correlate well with generalization. For instance, we can easily construct a counterexample by only optimizing with hard examples. In this paper, we propose to resolve such a contradiction from the perspective of anti-overfitting. First, we show that the sharpness at the mini-batch level approximately equals the sum of the gradient norm and the Kullback–Leibler (KL) divergence between the predictive distributions evaluated at the current point and its adversary. Second, we show that the \emphgradient consistency, which approximately quantifies the inner product between the mini-batch gradient and the per-example gradient, decreases consistently during training. This result explains why SAM is less likely to suffer from overfitting. At last, building on these insights, we further introduce CASAM, an algorithm that reweights each example according to gradient consistency to enhance the generalization performance. Extensive experiments on popular benchmarks such as ImageNet-1K and Clothing1M corroborate its efficacy.
Abstract: Skill-augmented agents load reusable skills as persistent runtime context, improving task performance but also giving malicious skills a durable channel for steering future actions. Such skills may leak secrets, corrupt code, bypass approvals, or stage data for exfiltration only after a concrete user task and workspace state make the unsafe action appear useful. This makes pre-install vetting insufficient and calls for runtime, task-conditioned protection. We propose Defense-as-Skill, a defense paradigm that implements the runtime guard itself as an installable, inspectable, and editable skill. Our guard, SkillSonar, runs alongside untrusted task skills and checks sensitive actions against the user's task boundary, routing each action to an allow, replan, or confirmation decision without modifying the underlying agent runtime. To study this setting, we construct SCOPE-R, a task-conditioned dataset covering 6 risk families and 21 sub-categories, with 206 attack-confirmed malicious instances and 43 benign tasks. We then improve SkillSonar on the SCOPE-R training subset using runtime guard-skill evolution, a Monte-Carlo Tree Search procedure that evolves the on-disk guard skill from feedback on the rollouts. On held-out SCOPE-R test splits across Claude Code and OpenClaw backends, SkillSonar reduces attack success while preserving benign utility, maintaining modest token overhead, and transferring to OOD attack families. These results show that runtime guards can be skill-native, transparent, and evolvable policy artifacts for skill-augmented agents.
PaperID: 3056, Poster
Abstract: We study strongly adaptive regret for Online Convex Optimization (OCO) with time-varying movement costs, a setting that evaluates a learner's performance over arbitrary sub-intervals. This framework captures complex sequential decision-making problems where both the environment and the cost of switching decisions fluctuate over time. We propose a novel algorithm which, for any interval [a,b] \subseteq [T], achieves a data-dependent regret bound of \widetilde\mathcalO\left( \sqrt(D^2+DP_[a,b])\sum_t \in [a,b](\\|g_t\\|^2+\lambda_t\\|g_t\\|) \right) , where D is the diameter of the decision domain, P_[a,b] is the path length of the comparator sequence over [a,b], g_t is the gradient of the loss, and \lambda_t is the movement-cost coefficient at round t. Key to our approach is a novel reduction to sleeping experts with movement costs, a previously unexplored variant that is of independent interest. As a result of this development, we further establish the first strongly-adaptive guarantees under delayed feedback. Our approach also naturally extends to OCO with time-varying memory, where we obtain strongly-adaptive guarantees which improve the best-known bounds.
PaperID: 3057, Poster
Authors: Phuong Nam Nguyen, Seng Loke, Anil Prabhakar, Son N Tran
Abstract: High-order binary optimization (HOBO) problems, defined as the minimization of functions involving interactions among three or more binary variables, are prevalent in machine learning and quantum computing. The minimization of these high-order energy functions remains computationally intractable without quadratization or reformulation. This work introduces FerQ, a universal and auxiliary-free reformulation of high-order energy functions based on Fermat quotient expansion. For any degree-d monomial over binary variables, FerQ provides an exact, closed-form representation as a rational linear combination of Fermat quotients evaluated on a scalar aggregate, with coefficients determined by a single matrix inversion. FerQ is evaluated against twelve existing energy function transformation methods on k-SAT and Max-k-SAT benchmarks from four databases, as well as on random p-spin glass and k-local Hamiltonian instances. FerQ achieves the highest satisfaction and weighted satisfaction rates across all benchmarks, while incurring the lowest CPU runtime. To enable deployment on quantum annealers, we derive FerQ-Bc, a qubit-efficient embedding scheme that maps the reformulated energies onto quadratic unconstrained binary optimization (QUBO) form, achieving good ancilla efficiency in high-degree regimes.
PaperID: 3058, Poster
Abstract: Pairwise comparisons from multiple judges are central to large language model evaluation and preference modeling, yet standard ranking pipelines often pool judgments into a single score vector, treating systematic judge disagreement as noise. We propose Heterogeneous Judge-Aware (HJA) ranking, a structured multi-judge ranking framework that separates consensus ranking, judge-specific sensitivity to consensus, and residual preference disagreement. HJA thereby treats ranking, judge sensitivity, and structured disagreement as separate inferential targets. We establish conditions under which this decomposition is identifiable and develop an anchored alternating algorithm that preserves the identifying geometry. For confidence quantification, we study a fixed-panel repeated-comparison regime in which the judge panel may remain fixed or modest while information grows through repeated judgments. This yields uncertainty statements for consensus and judge-specific ranking contrasts, sensitivity parameters, pairwise probabilities, and summaries of residual disagreement.Experiments on synthetic and real multi-judge comparison data show that HJA improves recovery, robustness, uncertainty calibration, and near-tie performance relative to pooled and sensitivity-only baselines. The fitted model also provides diagnostics for judge disagreement and model-affinity patterns, giving a statistically grounded framework for ranking under heterogeneous comparative judgments.
Abstract: Optimization modeling is the process of translating real-world decision problems, often described in natural language, into formal mathematical formulations and executable solver code. While recent advances in large language models have shown promise in automating this process, most existing approaches remain one-shot: a model produces a formulation once, without executing it, conditioning on solver feedback, or iteratively revising errors. This stands in sharp contrast to real-world optimization modeling, which is inherently interactive and proceeds through repeated solve--debug--revise cycles. We introduce PEARL, a system for interactive optimization modeling that uses Python execution and mathematical programming solvers inside this loop. Rather than relying on a fixed repair workflow, PEARL learns when to test partial models, how to revise from solver diagnostics, and when to stop. It operates in a multi-turn tool-integrated setting where intermediate execution results, feasibility signals, and solution checks are used to improve both formulations and solver code before finalization. Across diverse optimization benchmarks, PEARL substantially improves verified solve rates over strong one-shot and tool-augmented baselines; notably, our PEARL-Qwen3-4B model outperforms the much larger DeepSeek-V3.2-685B in both macro- and micro-averaged accuracy on optimization modeling tasks.
Abstract: Progress in language model development is often driven by comparative decisions: which architecture to adopt, which pretraining corpus to use, or which training recipe to apply. Making these decisions well requires reliable performance forecasts, yet the two commonly used signals are fundamentally limited. Cross-entropy loss is poorly aligned with downstream capabilities, and direct downstream evaluation is expensive, sparse, and often uninformative at early training stages. Instead, we propose to construct proxy metrics by aggregating token-level statistics, such as entropy, top-k accuracy, and expert token rank, from a candidate model's next token distribution over expert-written solutions. Across three settings, our proxies consistently outperform loss- and compute-based baselines: 1) For cross-family model selection, they rank a heterogeneous population of reasoning models with mean Spearman \rho = 0.81 (vs. \rho = 0.36 for cross-entropy loss); 2) For pretraining data selection, they reliably rank 25 candidate corpora for a target model at roughly 10,000× less compute than direct evaluation, pushing the Pareto frontier beyond existing methods; and 3) for training-time forecasting, they extrapolate downstream accuracy across an 18× compute horizon with roughly half the error of existing alternatives. Together, these results suggest that expert trajectories are a broadly useful source of signal for assessing model capabilities, enabling reliable performance forecasting throughout the model development lifecycle.
Abstract: Model growth from a given checkpoint aims to accelerate training of a larger model, offering potential resource savings. Despite recent interest, warmstarting has seen limited practical adoption in large-scale training. We attribute this to two underexplored factors: (1) an overemphasis on preserving the smaller model's performance at initialization, which constrains operator design for new architectures, and (2) insufficient analysis of how growth interacts with hyperparameters and scaling behavior, compounded by inconsistent growth factors across the literature. We show that preserving the base model's initial post-growth performance is not necessary for strong final performance, and that simple, architecture-agnostic growth strategies can outperform more complex warmstarting operators. Crucially, we empirically identify an upper bound on the growth factor g beyond which training from scratch is more efficient. We observe this across multiple ablation setups. Notably, this limit is also present, but unreported, in prior published results. Across our experiments on dense MLPs and dense language models, we find that a 2× growth factor is the most reliable in yielding convergence speedups, with gains most pronounced under 20 tokens/parameter token budgets and diminishing as budget increases. We fit scaling laws over these observations to provide predictive guidance for practitioners deciding when and how much to grow. Together, our analysis provides practical guidelines and empirical limits for model growth.
PaperID: 3062, Poster
Abstract: Predicting transcriptome-wide cellular responses to genetic and biomolecular perturbations is central to functional genomics and therapeutic target discovery. Existing methods assign gene positional encodings derived from arbitrary orderings or static co-expression masks built from unperturbed cells, and cannot capture the functional rewiring induced by perturbations. We present SPEAR, a generative framework built on two complementary positional mechanisms. Spectral positional encoding grounds each gene's position in the geometry of the co-expression network via continuous coordinates derived from the graph Laplacian. Adaptive repositioning then derives each gene's position from its perturbation-conditioned hidden state at every layer, allowing the attention genes exert on one another to reorganize without explicit supervision. Evaluated on genetic, pharmacological, and cytokine perturbation benchmarks, SPEAR demonstrates excellent recovery of differential gene-expression, distributional structure, and perturbation-specific identifiability. SPEAR also captures latent positional relationships that closely map to known transcription factor-target interactions.
PaperID: 3063, Poster
Authors: Vipul Baghel, Ravi Hegde
Abstract: Despite significant advances in large vision-language models (Video-LLMs) for general video understanding, accurately narrating fine-grained, highly dynamic human activities remains a formidable challenge. Existing approaches typically rely on global sequence embeddings or isolated part-level retrieval, lacking the precise relational kinematic grounding required to prevent physical hallucinations. To address this, we introduce Motion Retrieval-Augmented Generation for Detailed Video Captioning (MoRe-DVC). Our algorithmic novelty centers on two core contributions: an event-boundary-aware graph-conditioned retrieval mechanism that captures hierarchical and chronological motion dependencies, and a symbolic physical faithfulness reward that explicitly constrains generation to grounded physical reality. Specifically, MoRe-DVC integrates a Pose-Orientation Encoding Module (POEM) to extract symbolic descriptors, an Event Calibration and Refinement Module (ECRM) to dynamically adjust coarse action boundaries into an event graph, and a Graph-Conditioned Hybrid Retriever (GCHR) to select relationally relevant evidence. By coupling graph-level retrieval with a strict symbolic reward, our framework reduces motion-language inconsistency. Extensive experiments demonstrate that MoRe-DVC achieves robust outperformance across three motion-centric datasets (BoFiT, HumanML3D, and FineMotion), establishing a new standard for physically faithful action narration.
PaperID: 3064, Poster
Authors:
Yucheng Wang, Zedong WANG, Yuetong Wu, Yue Ma, Dan XuAbstract: Single-view 3D hair reconstruction has progressed rapidly, but many methods optimize for visual plausibility rather than producing strand assets that can be directly used in grooming, editing, or simulation pipelines. We propose SAFE-Hair, a framework that reconstructs explicit strand grooms from a single portrait by decoding hair from persistent follicles in a canonical scalp-UV domain. Given a fitted template head, SAFE-Hair projects self-supervised image features onto the scalp, uses a conditional Rectified Flow model to generate latent scalp fields, and decodes occupancy, segment directions, length allocation, and total length at a fixed set of scalp follicles. This representation enforces scalp-rooted strand identities by construction and enables direct export with follicle IDs and guide-family metadata. We further introduce attachment, collision, and bending priors to improve geometric validity without per-subject test-time optimization. To evaluate both reconstruction fidelity and asset usability, we introduce HairBench-3D, a benchmark built from public 3D hair assets with standardized alignment, rendering, resampling, and metrics for geometry, direction consistency, invisible-region completion, silhouette agreement, and collision. SAFE-Hair improves Chamfer distance from 3.74 mm to 3.28 mm, F-score from 77.9 to 81.7, invisible-region F-score from 66.1 to 69.8, and penetration ratio from 4.8% to 3.9% over the strongest baseline.
PaperID: 3065, Poster
Authors: Youkang Wang, Jian Wang, Tianyi Zeng, Xiao-Yong Wei, Qing Li
Abstract: Test-time training and test-time policy optimization have emerged as effective tools for improving the robustness of LLMs under distribution shift. A common approach constructs label-free reward signals from multiple sampled rollouts using majority voting, but existing methods typically rely on a fixed, large rollout budget for every query. This fixed allocation is inefficient because many queries reach consensus well before the rollout budget is exhausted. In this work, we introduce AdaPO, which replaces fixed-budget majority voting with an adaptive two-stage voting procedure that moves from electing to confirming when a leader is statistically dominant, and stops once the posterior error of accepting that leader is sufficiently low. AdaPO is plug-and-play with standard policy optimization algorithms such as PPO and GRPO, and it is also compatible with supervised fine-tuning at test time. Our rigorous theoretical results show that, under the symmetric categorical noise model with known likelihood parameters, AdaPO's electing-stage trigger coincides with the classical sequential probability ratio test (SPRT), while the confirming-stage rule yields a principled posterior-threshold stopping mechanism. Across representative reasoning benchmarks, AdaPO substantially reduces compute, achieving over 40% token savings on GPQA while maintaining competitive accuracy. The source code will be open upon acceptance at \urlhttps://open-upon-acceptance.
PaperID: 3066, Poster
Abstract: Reasoning capabilities are critical for advancing Large Language Models, yet current approaches either require massive computational budgets or struggle to effectively distill reasoning to smaller models. Standard distillation methods rely on outcome-based rewards, failing to distinguish between sound reasoning and lucky guesses. We propose , a framework that enhances reasoning in compact models through three innovations: (1) interactive cross-feedback where teachers iteratively critique each other's reasoning, (2) fine-grained step-wise quality assessment capturing logical validity independent of final answers, and (3) coherence-aware step stitching that synthesizes complementary strengths. Students are trained via Reasoning Quality Optimization (RQO) with budget constraints. Our model, CRD-4B, achieves 97.3% on MATH-500 and 70.3% on AIME'25, surpassing baselines while using only 50K training examples, up to 12 times smaller than the datasets of comparable models.
PaperID: 3067, Poster
Abstract: Dynamical systems reconstruction (DSR) aims to learn surrogate models that capture the dynamics underlying time-series data. Reliably deploying these surrogates requires uncertainty estimates consistent with the learned dynamics. We expose a (DPC) gap: the pursuit of finite-horizon probabilistic objectives can degrade dynamics or decouple predictive uncertainty from the local tangent dynamics it ought to reflect. We isolate three mechanisms behind this gap: core collapse, noise masking, and blind uncertainty. Specifically, we show that open-loop Gaussian rollout objectives can penalize Jacobian-generated covariance growth in chaotic systems, encouraging optimization shortcuts that weaken physical expansion or decouple uncertainty from it. To mitigate this gap, we propose KAFFEE (Kalman-Aware Framework For Ergodic Emulation), a differentiable extended Kalman filter-based training framework that evaluates likelihood on local predictive residuals (innovations) while transporting covariance through learned local Jacobians. On stochastic hyperchaotic Lorenz-96, KAFFEE reduces the identified failure modes, improves reconstruction of dynamical invariants relative to open-loop objectives, and maintains competitive predictive scores. We further show that the DPC gap appears when probabilistically adapting a DSR foundation model across 13 chaotic systems, where KAFFEE enables in-context Bayesian filtering while largely preserving zero-shot dynamics.
Abstract: While test-time scaling has revolutionized reasoning in large language models, generative video reasoning remains bottlenecked by a single-shot paradigm. We demonstrate that searching over denoising steps cannot rescue logically flawed rollouts because spatial trajectories commit early in the diffusion process. Root-level Best-of-N (BoN) sampling is similarly inefficient: reasoning errors cluster early in the temporal axis, and resampling blindly discards verified upstream progress. To unlock effective test-time scaling for video models, we introduce Temporal Backtracking Search (TBS), which shifts the search space to the temporal axis. TBS transforms video generation into an iterative generate--verify--restart loop via three core mechanisms: (1) variable-K conditioning to resume generation from arbitrary clean prefixes; (2) temporal process verification to localize failures and extract valid restart anchors; and (3) prefix-based search to reallocate compute toward extending correct trajectories rather than root resampling. Across algorithmic, navigation, and robotics domains, TBS Pareto-dominates matched-budget BoN. In a strict out-of-distribution setting where one-shot generation collapses (0.7% for BoN), TBS achieves 22.7%, with every solved episode stemming from a restarted branch. Ultimately, TBS reveals that the local reasoning competence of video models far exceeds what single-shot rollouts indicate, providing a scalable test-time framework to unlock it.
PaperID: 3069, Poster
Abstract: The deployment of Large Language Models (LLMs) with extended context windows is fundamentally constrained by the quadratic computational and memory costs of the self-attention mechanism. While acceleration techniques like sparse attention and quantization offer potential relief, they currently face a critical dilemma: uniform compression strategies fail to capture the non-uniform distribution of information importance, degrading performance on long-range dependencies; conversely, fine-grained, token-level adaptation often introduces irregular memory access patterns that negate theoretical efficiency gains on modern GPUs. To resolve this tension, we introduce BMAttn (Block-Aligned Mixed-Precision Attention), a framework that unifies importance-aware precision allocation with hardware-friendly execution. BMAttn partitions attention maps into block-aligned high-precision, low-precision, and sparse regions, using a novel affine windowing mechanism to dynamically adjust boundaries based on sequence length. We further propose a saliency-weighted calibration method and a layer-adaptive regularizer that adaptively aligns precision targets with perceptual importance and layer sensitivity. Extensive experiments demonstrate that BMAttn achieves a 3.3x speedup on long-context tasks with negligible accuracy loss, and up to 5x speedup with minimal degradation, effectively bridging the gap between algorithmic adaptivity and hardware efficiency.
PaperID: 3070, Poster
Abstract: Computing accurate geometry from multi-view images is a fundamental problem in computer vision. In this work, we address the Multi-View Stereo (MVS) setting, where camera parameters are assumed to be known. Most existing learning-based methods rely on per-view cost volumes as a camera-induced prior, effectively casting the problem as a sequence-to-one mapping that predicts depth only for a designated reference view. We instead reformulate MVS as a sequence-to-sequence task, enabling the simultaneous prediction of depth maps and point maps for all input views. To this end, we propose a global transformer-based architecture with two key components that explicitly utilize camera-induced priors. First, we employ ray-map embeddings to inject camera parameters into image patch tokens, making the transformer architecture camera-aware for geometry prediction. Second, we replace conventional per-view cost volumes with a unified global cost-volume representation that jointly captures 3D structure across all views. Extensive experiments on multiple public benchmarks demonstrate that our approach achieves state-of-the-art performance, surpassing both multi-view stereo and feed-forward reconstruction baselines.
PaperID: 3071, Poster
Authors: Philippe Meyer, Claire Roman
Abstract: Quantitatively comparing writing systems at a worldwide scale is challenging because visual similarity must be inferred without relying on predefined linguistic or genealogical assumptions, and only limited ground-truth taxonomy is available for validation. We introduce a fully data-driven framework for comparing scripts based on learned glyph representations. Using a teacher--student model, we obtain deformation-invariant glyph embeddings and aggregate them into script-level distances. To enable principled model selection in the absence of labels, we construct synthetic script datasets that simulate key aspects of script evolution, including partial inheritance, visual distortion, and inventory variation. These benchmarks are used to independently select both the script-level distance metric and the clustering method, showing that nearest-neighbour aggregation and Ward hierarchical clustering best recover controlled relatedness structure. Applying the validated pipeline to 125 historically attested writing systems yields a global visual taxonomy that recovers coherent typological families while also revealing visually driven associations beyond strict genealogy. The resulting structure is stable under controlled perturbations and shows a moderate correlation with independent geographic and temporal metadata, indicating that learned representations capture non-trivial traces of historical transmission. Beyond large-scale comparison, this approach provides a quantitative tool for exploring relationships between writing systems and has potential applications in historical analysis and the decipherment of ancient scripts by identifying visually related alphabets.
PaperID: 3072, Poster
Abstract: Amortized Bayesian experimental design (BED) enables real-time design strategies by shifting acquisition cost offline, but existing methods remain tied to the prior and task distribution they were trained on and cannot exploit additional information available at deployment. We introduce in-context amortized BED with unified knowledge encoding (IMBUE), an amortized BED framework that accepts two forms of such information without retraining, namely user-specified prior knowledge and observations from previous instances of the experiment. Both are encoded as additional input tokens and processed in-context with the accumulated observations by a shared Transformer, and a learned reliability filter scores each auxiliary token against the accumulated observations and excludes inconsistent ones. Across four standard BED benchmarks, IMBUE accelerates early-stage information acquisition when the auxiliary knowledge is reliable, and remains close to the no-auxiliary baseline when it is not.
PaperID: 3073, Poster
Abstract: Large language model (LLM) routing is widely used to advance the cost-quality Pareto frontier of modern serving systems. While coarse-grained routing at the session or query level has been widely adopted in production systems, recent algorithm work shows substantial efficiency and quality advantages of finer token-level routing. However, efficiently serving token-level routed inference poses significant challenges to existing system design. Built on single-LLM assumptions, current systems struggle with step desynchronization and frequent batch admission delays under token-level routing, and they also impose high implementation complexity on developers. To address this, we propose TokenRouter, an efficient and user-friendly serving system for token-level routed LLM inference. TokenRouter shifts from the conventional request-centric paradigm to a model-centric one. Instead of treating each request as a synchronized decoding stream, TokenRouter launches a subserver for each candidate LLM and dispatches requests asynchronously according to routing decisions. Each subserver employs a delayed-batching scheduler, whose hyperparameters are derived mathematically from a throughput model of the system. Across diverse routing algorithms, workloads, and model pairs, TokenRouter achieves 1.83-64.29× higher decoding throughput than existing systems, substantially advancing the serving efficiency of token-level LLM routing.
PaperID: 3074, Poster
Authors: Santosh M Rajkumar, Sriram Narayanan, Samuel E Otto, Debdipta Goswami
Abstract: World models enable prediction, planning, and decision-making by learning internal simulators of environment dynamics. Although substantial progress has been made in world models from pixel-based observations, many physical systems are instead observed through vector-based measurements that are noisy and only partially informative of the underlying state. We introduce K-PWM, a Koopman-structured probabilistic world model for prediction and control under partial observation. K-PWM combines a parameterized Koopman-structured state-space model with a nonlinear observation decoder, and performs probabilistic latent-state inference directly through the learned model parameters rather than through a separately trained inference network. We train K-PWM using a generalized expectation conditional maximization (GECM) procedure with principled initialization. Across Gymnasium MuJoCo tasks, K-PWM improves sample efficiency, yields reliable long-horizon predictions, and supports downstream control through model predictive path integral (MPPI) control. These results suggest that structured probabilistic latent dynamics provide a useful route toward data-efficient world modeling and control in partially observable vector-valued settings.
PaperID: 3075, Poster
Authors: Youn Jongsu, Dongyoung Lee, JaeHyeong Bae, SANG-IL CHOI, Sangtae Choi, Dae Ung Jo, Jongwon Choi
Abstract: We present a vision–language pretraining (VLP) framework for the medical domain that focuses on preserving fine-detail clinical information, such as disease severity, which plays a critical role in downstream medical applications. Despite recent advances in medical VLP, existing approaches often struggle to retain such subtle semantics. This limitation mainly stems from the use of contrastive learning schemes that treat supervision as hard one-hot labels, overlooking the ordinal and overlapping nature of clinical findings and thereby introducing false negatives. To overcome these challenges, we propose FinAl, a fine-grained alignment approach that incorporates soft alignment targets into vision-language pretraining. These targets are derived from structured radiology reports and encode relative semantic relationships among clinical descriptions, allowing fine-grained linguistic information to be transferred into the visual representation space. Extensive experiments show that FinAl improves performance on fine-grained clinical tasks, including severity grading and severity-aware retrieval, with additional evaluations on challenging fine-grained scenarios further supporting its robustness. At the same time, the proposed method maintains competitive performance on standard coarse-grained clinical benchmarks.
Authors:
Haoyun Jiang, Junqi He, Feng Hong, Xinlong Yang, jianwei zhang, Zhengyang Zhuge, Zheng Li, Xiaofeng Cao, Zhiyong Chen, Bo Han, Junyang Lin, Jiangchao YaoAbstract: Inference efficiency in Large Language Models (LLMs) is fundamentally limited by their serial, autoregressive generation, especially as reasoning becomes a key capability and response sequences grow longer. Speculative decoding (SD) offers a powerful solution, providing significant speed-ups through its lightweight drafting and parallel verification mechanism. While existing work has nearly saturated improvements in draft effectiveness and efficiency, this paper advances SD from a new yet critical perspective: the verification cost. We propose TriSpec, a novel ternary SD framework that, at its core, introduces a lightweight proxy to significantly reduce computational cost by approving easily verifiable draft sequences and engaging the full target model only when encountering uncertain tokens. TriSpec can be integrated with state-of-the-art SD methods like EAGLE-3 to further reduce verification costs, achieving greater acceleration. Extensive experiments on the Qwen3 and DeepSeek-R1-Distill-Qwen/LLaMA families show that TriSpec achieves up to 35% speedup over standard SD, with up to 50% fewer target model invocations while maintaining comparable accuracy.
Abstract: The exponential growth of machine learning submissions has strained the traditional peer review process, resulting in slow feedback loops for authors and an immense burden on reviewers to rigorously audit technical soundness and verify literature. To address this, we introduce ScholarPeer, a multi-agent framework designed to operationalize the rigorous auditing workflow of a senior researcher. Rather than attempting to replace human judgment, ScholarPeer serves as a co-scientist: acting as a mentor for rapid author iteration prior to submission, and as an active verification assistant that augments human reviewers. The framework structurally decouples contextualization from critique by deploying a sub-domain historian to synthesize the field's trajectory, a baseline scout to proactively hunt for omitted state-of-the-art comparisons, and a multi-aspect Q&A engine that deeply audits technical soundness—scrutinizing internal logical consistency, experimental validity, and mathematical rigor—while cross-referencing claims against top-tier academic venues. We comprehensively evaluate ScholarPeer on ~1,800 ICLR submissions spanning 2020 through 2025. Our results show that ScholarPeer achieves significant win-rates against state-of-the-art fine-tuned models and search-augmented agentic baselines.
PaperID: 3078, Poster
Abstract: Despite impressive progress in video generation, existing models remain limited to surface-level plausibility, lacking a coherent and unified understanding of the world. Prior approaches typically incorporate only a single form of world-related knowledge or rely on rigid alignment strategies to introduce additional knowledge. However, aligning the single world knowledge is insufficient to constitute a world model that requires jointly modeling multiple heterogeneous dimensions (e.g., physical commonsense, 3D and temporal consistency). To address this limitation, we introduce WorldForge, a unified framework that integrates complementary world knowledge into video generators via a Joint World Modeling Paradigm, jointly predicting video pixels and features from foundation models to capture temporal dynamics, spatial geometry, and semantic consistency. However, naively optimizing these heterogeneous objectives can lead to visual instability and temporal flickering. To mitigate this issue, we propose Consistent Constraint Annealing (CCA) to progressively regulate world-level constraints during training, and Multi-Source Inner-Guidance to enforce learned world priors at inference. Extensive evaluations show that WorldForge improves world consistency, outperforming Wan2.1 by 2.26 points on VBench.
Abstract: Balancing exploration and exploitation is a core challenge in sequential decision-making and black-box optimization. We introduce POETS (Policy Ensembles for Thompson Sampling), a novel framework that bridges uncertainty quantification and policy optimization. Our approach is grounded in the insight that policies trained with Kullback-Leibler (KL) regularization implicitly encode an underlying reward function. Building on this, POETS bypasses the complex, nested process of training an uncertainty-aware reward model and separately fitting a policy to this model. Instead, we directly train a policy ensemble to capture epistemic uncertainty by matching implicitly encoded reward functions to online, bootstrapped data. To overcome the prohibitive compute and memory constraints of ensembling Large Language Models (LLMs), POETS utilizes an efficient architecture: the ensemble shares a pre-trained backbone while maintaining diversity through independent Low-Rank Adaptation (LoRA) branches. Theoretically, we prove that POETS implicitly conducts KL-regularized Thompson sampling and thus inherits strong cumulative regret bounds of O(\sqrtT \gamma_T). Empirically, we demonstrate that POETS achieves state-of-the-art sample efficiency across diverse scientific discovery domains, including protein search and quantum circuit design. Furthermore, it improves the optimization trajectories of reinforcement learning, proving particularly robust in off-policy settings with experience replay or in small dataset regimes.
Abstract: We revisit Kashin-decomposition-based weight quantization for large language models and propose an improved algorithm with stronger convergence properties and structured, efficient orthogonal transforms. The method retains the core factorization of each weight into two components -- one with bounded infinity norm and the other with bounded infinity norm after an orthogonal transformation -- but replaces the dense random orthogonal matrix with a sign-randomized Discrete Cosine Transform (DCT), reducing the per-iteration cost from \mathcalO(N^2) to \mathcalO(N \log N). The proposed greedy algorithm with alternating updates guarantees the four-peak distribution required for stable 2-bit clustering of each factor and admits closed-form initialization of cluster centers, removing the multi-restart k-means bottleneck of prior work. Composed with OPTQ-style sequential error compensation and QuIP-style incoherence preprocessing, the resulting JAX pipeline is competitive with OPTQ, QuIP, QuIP-RG and a fine-tuning- and vector-quantization-free variant of QuIP\# at 4-bit per channel on OPT, Llama-2 and Pythia, with favorable wall-clock scaling. The bounded-\ell_\infty factorization is also notably robust: on stress configurations where QuIP variants diverge to four-digit perplexity (Pythia-6.9B) or abort with NaNs in LDL back-substitution (Mistral-7B), Kashin-DCT remains numerically stable and stays close to FP16 baseline. At inference time, each weight decomposes into two 2-bit factor codes per channel that are structurally suited to native-2-bit hardware.
PaperID: 3081, Poster
Authors: Emirhan Inan, Suayb Arslan
Abstract: Convolutional Neural Networks (CNNs) draw broad inspiration from the primate visual system yet prove notoriously brittle under common image corruptions, whereas the latter suffers little from such degradations. Extending beyond this initial inspiration, architectures incorporating quantifiable biological priors that enhance network resilience are increasingly gaining traction. Nevertheless, most of these approaches have so far been confined to emulating the higher visual areas, seldom utilizing the upstream retinal stages that constitute the very basis of seeing. At its core, primate vision relies on photons captured by retinal photoreceptors, the signals of whose conversion are then relayed through the Lateral Geniculate Nucleus (LGN) to primary visual cortex (V1) and further downstream for additional processing. Because cone photoreceptors are densest at the center of the retina (fovea), the eyes must be repositioned continuously (saccade) to direct the high-acuity macula toward objects of interest (fixation). Furthermore, the resultant rapid eye movements induce intra-saccadic motion streaks that putatively aid in the spatiotemporal processing of inputs and provision of translation-invariant object correspondence across frames. Motivated by this mechanism, we propose \textttAuGhostmentation, a portmanteau of augmentation and ghosting, which delivers the intra-saccadic retinal smear as a frontend transform. Superimposing the ghosted images—fainter counterparts—atop the original, based on human oculomotor statistics, the resultant datum elicits retinal persistence by integrating varied recompositions of the scene in a single representation, closely mimicking the foveal exploration scheme intrinsic to dynamic primate vision. When prepended to ResNet-18 trained on Tiny ImageNet, \textttAuGhostmentation retains 98% validation accuracy while delivering a striking 60% relative gain in robustness on Tiny ImageNet-C. Applied to VOneNet, it likewise preserves 98% classification accuracy and yields a substantial 25% relative increase in compounded robustness. For ImageNet-100, the introduced ghosts not only improve the corruption accuracy relative to the ResNet-50 baseline by 23%, but the clean accuracy by 4%, also. Building on these resilience gains, \textttAuGhostmentation outpaces competing augmentations by up to 18% and baseline ResNet-50 by up to 37% on stylized tests, suggesting that retinal priors naturally foster shape-biased processing and thereby provide greater alignment with human object recognition.
PaperID: 3082, Poster
Authors: Ratun Rahman, Atit Pokharel
Abstract: Structured generation is typically formulated as a deterministic mapping from input to a single structured output. This assumption is fundamentally misaligned with natural language, which often admits multiple structurally distinct yet semantically valid interpretations. As a result, existing methods rely on single-reference supervision that collapses ambiguity into one prediction, leading to unstable and inconsistent outputs under underspecified inputs. We reformulate structured generation as an ambiguity-aware inference problem, where the objective is to reason over a set of plausible structured outputs rather than predict a single solution. We show that standard training inherently suppresses valid alternatives, and instead adopt a multi-valid learning framework that preserves ambiguity during optimization. To operationalize this formulation, we construct a candidate space of structured outputs, represent each candidate as a graph capturing entities and their relationships, and perform selection via graph-based reasoning over semantic alignment and relational consistency. This enables principled comparison across competing interpretations beyond token-level likelihoods. Experiments on structured generation benchmarks demonstrate consistent improvements in accuracy, stability, and robustness under varying levels of ambiguity. These results highlight the importance of modeling ambiguity explicitly and suggest a shift from single-output prediction to set-level reasoning in structured generation.
PaperID: 3083, Poster
Authors: Xuechao Lan, Wenmin Li, Sujuan Qin, Zhengping Jin, Fei Gao
Abstract: Machine Unlearning (MU) is an emerging task to remove specified training data from a trained model while preserving its utility, with Gradient-based Unlearning (GU) standing out as a prevalent category due to high scalability. However, standard MU which solely optimizes for clean accuracy, inadvertently damages adversarial robustness on the retain data. Although a few studies extend MU to adversarially trained models, they rely on strong assumptions (e.g., well-conditioned Hessians, smoothness) and costly Hessian-based updates, limiting scalability and efficiency. In this paper, we propose GUARD, a scalable and efficient framework that achieves unlearning while preserving retain-data robustness. Having identified the two causes of robustness degradation, we accordingly design Robustness-Aware Gradient Projection and Importance-Guided Parameter Selection. Both modules rely on a compact adversarially vulnerable dataset that we construct with or without retain data, to enable scalability under varying data constraints. Validated by theoretical analysis and experiments across 6 settings, GUARD outperforms 6 baselines, boosting robust accuracy by up to 6.3×, closely matching the Retrain oracle in both unlearning and robustness.
Authors: Arif Dönmez, Ellen Fritsche, Axel Mosig, Katharina Koch
Abstract: Neural Sheaf Diffusion (NSD) generalizes diffusion-based Graph Neural Networks by replacing scalar graph Laplacians with sheaf Laplacians whose learned restriction maps define a task-adapted geometry. While the diffusion limit of NSD is known to be the space of global sections, the representation-theoretic structure of this harmonic space remains largely implicit. In this paper, we develop a quiver-theoretic interpretation of NSD by identifying cellular sheaves on graphs with representations of the associated incidence quiver. Under this correspondence, learned sheaf geometries become points in a finite-dimensional representation space. We prove that direct-sum decompositions of the underlying incidence-quiver representation induce corresponding decompositions of the harmonic space reached in the diffusion limit. This provides an algebraic interpretation of oversmoothing as representation degeneration: a conceptual framing where learned sheaves collapse toward trivial or low-complexity summands whose global sections fail to preserve discriminative information. Building on this viewpoint, we connect sheaf diffusion to stability, moduli, and moment-map principles from Geometric Invariant Theory. We introduce moment-map-inspired regularizers that bias learned restriction maps toward more balanced representation geometries, and we identify a structural obstruction in standard equal-stalk architectures: when d_v=d_e, the admissibility condition for learnable stability parameters forces the trivial all-object summand onto a stability wall. We show that non-uniform stalk dimensions remove this obstruction, making adaptive stability meaningful in principle. Empirical evaluations on heterophilic benchmarks are consistent with this mechanism: breaking stalk symmetry can reduce variance or improve validation behavior on some datasets, and adaptive stability regularization becomes more effective in selected rectangular settings. These results support the view that moment-map regularization is a structured but dataset-dependent geometric bias rather than a universal performance booster. Overall, our framework interprets oversmoothing not only as a spectral pathology, but as a degeneration phenomenon in the underlying representation geometry.
PaperID: 3085, Poster
Abstract: Multivariate weather time series are a central modality for forecasting and decision support, and the surrounding language increasingly mediates how numerical weather is consumed by users and downstream models. However, existing pipelines depend on closed-source language model ensembles with expert quality control or on retrieving human-curated text, while prior caption-RL rewards optimize cycle similarity or downstream-task utility, neither of which can tell a captioner that a confidently asserted phenomenon is contradicted by the data, so the absence of an executable, station-observable, and refutation-aware reward signal has blocked the application of self-play RL to this scientific modality. To address this, we propose AtmoZero, a post-training framework that trains a weather-time-series captioner without a single human-written caption, by aligning a language model with a library of seven station-observable scientific verifiers, pairing them with an explicit refutation channel that penalizes caption claims the data contradicts, and combining them with auxiliary reconstruction and forecaster-utility rewards under group-relative policy optimization. Extensive empirical results on a twelve-year ERA5 reanalysis corpus and real-world surface observations show that AtmoZero attains the lowest unsupported-claim rate and forecasting error at every horizon, outperforming closed-source captioners, dense-reward RL, and far-larger foundation models, with extensive analyses confirming that these gains arise from multiple complementary signals and are robust under distribution shift.
PaperID: 3086, Poster
Abstract: Vision Language Models (VLMs) enable a unified model to solve various vision tasks through prompting. They have shown promising performance in semantic understanding. However, 3D understanding still largely relies on expert vision models with complex task-specific designs. The key argument this work wants to make is that VLMs are native 3D learners. Our in-depth large scale study shows that 1) focal length unification, 2) text-based pixel reference and 3) data mixture and scaling, are all you need for effective 3D learning. Model architecture changes, large models, heavy data augmentations, and complex losses including the regression formulation, many of which form the foundation of expert vision models, are actually not necessary conditions. As a result, we propose VLM^3, a scalable method with the simplest design that enables standard VLMs to master diverse 3D tasks. VLM^3 not only advances the VLM depth estimation accuracy by a large margin (0.84 \rightarrow 0.9), but also enables diverse 3D tasks such as pixel correspondence, camera pose estimation and object-level 3D understanding, matching expert vision model accuracy while maintaining standard architectures and text-based training. Code will be released to the community.
PaperID: 3087, Poster
Abstract: Not all mistakes are alike. For example, in medical diagnosis, false positive and false negative mistakes have a different impact and should not be aggregated. By revisiting the classical theory of boosting through this lens, we introduce the notion of Z-learner, which returns weak hypotheses achieving joint false-positive/false-negative guarantees (FP,FN) from some set Z \subseteq [0,1]^2. This formulation generalizes the standard notion of \gamma-weak learning — recovered by taking Z= (z_1,z_2): z_1+z_2<1/2-\gamma — and allows for more flexible notions of weak learnability that account for the impact of the distribution on the achievable learning guarantees. This raises a natural question: which sets Z are boostable? We answer this by giving a complete geometric characterization that unifies, as special cases, both the classical boosting dichotomy and the recent multi-objective extension of Bressan et al. (COLT 2025). We also present an adaptive algorithm for boosting Z-learners without relying on the prior knowledge of~Z, generalizing AdaBoost's ability to handle unknown margin parameters. Finally, we demonstrate that Z-learners can achieve confidence amplification. However, the analysis is significantly more involved than in classical boosting and relies on the centerpoint theorem from discrete and computational geometry.
PaperID: 3088, Poster
Abstract: The widespread deployment of large language models (LLMs) has made reliable detection of machine-generated text increasingly critical. Recent curvature-based detectors introduce perturbations to candidate text and compute curvature as the log probability difference between original and perturbed variants using advanced surrogate models. However, such curvature measurements only capture weak discrepancies between machine-generated and human-written text. We propose Pref-DetectGPT, a training-free method that introduces preference-aware curvature measurement by leveraging the implicit preference capabilities of preference-optimized surrogate models. Specifically, we construct an implicit reward model from the preference-optimized policy to score both original and perturbed text, and redefine curvature as their normalized deviation, which serves as a significantly stronger signal for distinguishing machine-generated text from human-written text. Empirical evaluations across multiple public benchmark datasets demonstrate that Pref-DetectGPT achieves state-of-the-art detection performance with relative improvements of 3.49%, 3.13%, and 10.66% in AUROC, AUPR, and TPR5%, while exhibiting strong robustness against various adversarial attacks and input lengths.
PaperID: 3089, Poster
Abstract: Modern power grids are increasingly exposed to coordinated cyber-physical disruptions, where the outage of a small number of transmission components can trigger severe post-contingency load shedding. Designing proactive _security defense_ against such adversarial contingencies is naturally modeled as a tri-level Defender-Attacker-Operator (DAO) problem: the defender allocates limited hardening resources, the attacker selects worst-case line outages, and the operator solves a post-attack DC/AC-optimal power flow (OPF). However, the resulting tri-level program is computationally prohibitive even for moderate network sizes. We develop a certified four-step reformulation pipeline that converts this intractable tri-level DAO problem into a tractable single-level optimization, yielding a Single-Level Mixed-Integer Linear Program (SL-MILP) under DC-OPF and a Single-Level Mixed-Integer Conic Program (SL-MICP) under AC-OPF. The construction combines exact dualization, exact Big-M linearization, and a single controlled attacker relaxation. We prove that the resulting single-level optimum provides a conservative upper bound on the true worst-case load shedding and, more importantly, yields a robustness certificate for the deployed defense: no feasible binary attack can induce damage exceeding the reported single-level value. Extensive experiments on IEEE benchmark systems demonstrate substantial runtime improvements while maintaining strong worst-case security guarantees.
PaperID: 3090, Poster
Authors: Karen Sargsyan
Abstract: A neural network trained on the multiplication table of the quaternion group Q₈ can achieve perfect accuracy within each cyclic subgroup ⟨i⟩ and ⟨j⟩, predict every cross-region product wrong, and stay there indefinitely. This is the grokking plateau, and we prove its height is exactly 4 in dimension d=2, determined by the topology of the overlap before any training. The mechanism is a homotopy pullback: for G = A ∗_C B with finite overlap, the representation space of G is a homotopy pullback of those of A and B over that of C, so locally accurate modules need only agree up to a basis change. From this we derive a stratified generalization certificate whose strata are the connected components of Hom(C, U(d)): in the correct stratum, with locally accurate modules, the cross-region prediction is O(ε)-close to the target; in the wrong stratum it is provably far regardless of local accuracy, with exact gap c = 2(d − rₘₐₓ) for central cyclic overlap. The plateau is therefore the time the network spends in the wrong stratum, the grokking transition is the discrete jump between strata, and the framework predicts which task families exhibit such plateaus—those whose overlap representation space is disconnected—and which do not.
PaperID: 3091, Poster
Abstract: Speech large language models (SpeechLLMs) have emerged as a new interface for human–AI interaction, enabling real-time spoken responses. Time-to-First-Speech (TTFS) is a critical latency metric in such interactive SpeechLLM systems, directly determining perceived responsiveness. While prior work reduces TTFS by accelerating autoregressive decoding, we show that diffusion-based speech synthesis contributes a comparable portion of the latency. This reveals a fundamental limitation of existing approaches that optimize decoding alone, and indicates that optimizing either LLM decoding or speech synthesis in isolation is insufficient. This paper proposes \textttSpecSpeak, a new SpeechLLM inference paradigm that overlaps LLM decoding with speech synthesis to reduce the TTFS latency. \textttSpecSpeak introduces a lightweight draft model to generate a prefix quickly. Instead of waiting for fully verified tokens like speculative decoding, \textttSpecSpeak triggers speech synthesis early using draft tokens, while a target model performs parallel verification and correction. To handle inconsistencies, \textttSpecSpeak introduces a verification-aware reuse mechanism that adaptively preserves intermediate diffusion states based on prefix agreement, enabling efficient correction without restarting synthesis. Extensive experiments demonstrate that \textttSpecSpeak reduces the TTFS latency of SpeechLLMs by 18%--30%, while preserving response correctness and speech quality. Code will be released upon acceptance.
PaperID: 3092, Poster
Authors:
Aozhong Zhang, Selcuk Gurses, Yanxia Deng, Naigang Wang, Chi-Chun (Charlie) Liu, Davis Wertheimer, Derrick Liu, Xin Li, Zi Yang, Felix X.-.F Ye, Penghang YinAbstract: Key-value (KV) cache management has become a critical bottleneck for scaling Large Language Models (LLMs) to long-context scenarios due to linear memory growth. While recent eviction methods have moved beyond simple attention-score heuristics to incorporate value representations, they typically face a fundamental trade-off: the underlying optimization for minimizing output perturbation is an NP-hard combinatorial problem, leading to mathematically suboptimal heuristics or implementations that are difficult to kernelize. In this work, we propose DropKV, a simple principled eviction framework based on Decoupled Residual-Output Perturbation. By decoupling the joint eviction decision into independent per-token scoring, DropKV circumvents combinatorial intractability and admits a provable constant-factor approximation guarantee under a cone condition that we empirically verify on three long-context LLMs. To ensure practical viability, we provide a high-performance implementation using fused Triton kernels that avoid materializing the full attention matrix, delivering a 19.8× scoring-kernel speedup that translates to up to a 9.9% matched-batch end-to-end prefill speedup. Extensive evaluations on the RULER and LongBench benchmarks demonstrate that DropKV consistently outperforms state-of-the-art baselines at equivalent cache budgets, achieving the lowest inference latency among eviction methods while also remaining faster than the dynamic cache baseline.
PaperID: 3093, Poster
Abstract: Autoregressive (AR) image generators are powerful in principle but inflexible in practice: training with a fixed raster-scan order hardcodes a single factorization of the image distribution into the model, leaving it unable to inpaint, outpaint, or edit at inference time. Existing solutions—including the modern discrete-diffusion family—train from scratch on a mixture of token orderings, paying roughly 3× the compute of standard raster-scan training; discrete diffusion further sacrifices compatibility with the AR inference ecosystem (KV caching, vLLM) by construction. We argue that this cost is unnecessary. A pretrained raster-scan AR model has already learned the joint distribution over image tokens; what it lacks is not knowledge but the ability to express that knowledge through non-sequential conditional pathways. We propose a recipe for adapting pretrained raster-scan models into any-order generators using two components: that traces a stable optimization path from the pretrained solution to a good any-order basin. Applied to four base models across two architectures and two scales (LlamaGen-L/XL, RAR-L/XL), the recipe substantially outperforms MaskGIT on ImageNet inpainting (FID 6.21 vs. 10.44) and outpainting (FID 5.84 vs. 12.17), preserves generation quality close to the pretrained checkpoints, and costs only ≈25% additional compute beyond pretraining (on LlamaGen-XL)—while leaving the standard AR inference stack untouched. Inpainting, outpainting, and editing, our results suggest, need not be added to autoregressive image models. They can be unlocked from the ones that exist.
Abstract: Self-evolving LLMs excel in verifiable domains but struggle in open-ended tasks, where reliance on proxy LLM judges introduces capability bottlenecks and reward hacking. To overcome this, we introduce G-Zero, a verifier-free, co-evolutionary framework for autonomous self-improvement. Our core innovation is Hint-\delta, an intrinsic reward that quantifies the predictive shift between a Generator model's unassisted response and its response conditioned on a self-generated hint. Using this signal, a Proposer model is trained via GRPO to continuously target the Generator's blind spots by synthesizing challenging queries and informative hints. The Generator is concurrently optimized via DPO to internalize these hint-guided improvements. Theoretically, we prove a best-iterate suboptimality guarantee for an idealized standard-DPO version of G-Zero, provided that the Proposer induces sufficient exploration coverage and the data filteration keeps pseudo-label score noise low.. By deriving supervision entirely from internal distributional dynamics, G-Zero bypasses the capability ceilings of external judges, providing a scalable, robust pathway for continuous LLM self-evolution across unverifiable domains.
Abstract: As AI agents move from chat interfaces to systems that read private data, call tools, and execute multi-step workflows, guardrails become a last line of defense against concrete deployment harms. In these settings, guardrail failures are no longer merely answer-quality errors: they can leak secrets, authorize unsafe actions, or block legitimate work. The hardest failures are often contextual: whether an action is acceptable depends on local privacy norms, organizational policies, and user expectations that resist pre-deployment specification. This creates a practical gap: guardrails must adapt to their own operating environments, yet deployment feedback is typically limited to sparse, noisy user-reported failures, and repeated fine-tuning is often impractical. To address this gap, we propose daptation), a conservative policy induction framework that improves a fixed base guardrail through structured memory. LiSA converts occasional failures into reusable policy abstractions so that sparse reports can generalize beyond individual cases, adds conflict-aware local rules to prevent overgeneralization in mixed-label contexts, and applies evidence-aware confidence gating via a posterior lower bound, so that memory reuse scales with accumulated evidence rather than empirical accuracy alone. Across PrivacyLens+, ConFaide+, and AgentHarm, LiSA consistently outperforms strong memory-based baselines under sparse feedback, remains robust under noisy user feedback even at 20% label-flip rates, and pushes the latency--performance frontier beyond backbone model scaling. Ultimately, LiSA offers a practical path to secure AI agents against the unpredictable long tail of real-world edge risks.
PaperID: 3096, Poster
Abstract: Controllable 3D generation is a fundamental challenge in computer vision. Existing classifier-free guidance methods often yield suboptimal alignment with the given conditions, while Schr"odinger Bridge approaches overlook intrinsic geometric discrepancies across multimodal manifolds, leading off-manifold trajectories and semantic misalignment. To address these issues, we propose Intrinsic-Preserving Schr\"odinger Bridge (IPSB), a novel framework for direct part-aware 3D generation that leverages cross-modal intrinsic consistency. Specifically, we introduce Hierarchical Part-level Cross-modal Encoder that constructs hierarchical part-level representations through probability-based soft coverage, explicitly constraining graph topology to establish semantically consistent part graphs across modalities. Additionally, our IPSB incorporates asymmetric GW-inspired relational regularization, which strictly preserves the intrinsic semantic structure in the conditional manifold while allowing sufficient geometric flexibility in the 3D target manifold. Based on our IPSB framework, we devise Compositionally Controllable 3D Generation based on part-wise intervention inference, enabling independent replacement or resampling of target part nodes while fixing the trajectories of remaining components, thus achieving part-level controllable editing without full regeneration. Extensive experiments demonstrate that our superiority, efficiency and generality.
PaperID: 3097, Poster
Abstract: Accurate contact modeling is fundamental to understanding hand–object interaction, yet existing contact representations are typically restricted to object surfaces and rely on hand‑crafted rules to recover contact details, leading to severe penetrations and implausible results. To better exploit the rich detail in motion‑capture data, we introduce Volumetric Contact (VolCo), a representation that expands surface points to a set of 3D volumetric grids. VolCo encodes 3D contact that allows precise hand part recovery, and is organized in an inherent hierarchy: local contact details within each volume and global hand geometry across all volumes. Our framework, VolCoDiff, employs two modules to capture local and global features following this hierarchy. For local contact details, we use a 3D variational autoencoder to model the possible hand configurations conditioned on the local object signed distance field (SDF). For global hand geometry, we design a prior‑guided diffusion model that learns the distribution of compressed latent features aggregated from the volumetric grids. We evaluate our method on two benchmark datasets and demonstrate state‑of‑the‑art performance in penetration and stability, indicating the capability to generate tight grasps with much less severe penetrations.
PaperID: 3098, Poster
Abstract: In recent years, ensemble ideas have found parallels in federated learning (FL). Specific ensemble learning formulations under the federated setup involve information exchange between clients during the training process, making the classic FL communication bottleneck a critical issue. Importantly, the resulting communication patterns differ from those studied in traditional communication-efficient FL methods, placing ensembling outside their analytical scope. In this work, we close this gap by extending classic communication-efficient techniques to the ensembling setting and introducing a new mechanism specifically tailored to it. We instantiate these ideas in four algorithms and provide convergence guarantees for each under the smooth non-convex setup. We further support our theoretical findings with experimental results.
PaperID: 3099, Poster
Abstract: A classic result in learning and optimization theory is that exponentiated gradient descent (EG) or the multiplicative weight update method computes an \\epsilon-approximate minimizer of an \\ell_1-Lipschitz function over the d-dimensional probability simplex \\Delta_d :=\\x \in \\mathbbR^d_\\geq0:\\sum_i\in[d]x_i=1\\ with \\tildeO(\epsilon^-2) queries to a first-order oracle. We investigate whether there are improved parallel algorithms for this fundamental problem. We show that, up to logarithmic factors, the answer is negative for \\epsilon=\\tilde \\Omega(d^-1/6) - any randomized algorithm that makes O(\\mathrmpoly(d)) subgradient queries per round requires \\tilde\\Omega(1/\\epsilon^2) rounds to output an \epsilon-optimal point with constant success probability. Moreover, to obtain this result we provide an analogous characterization of the parallel complexity of minimizing a \ell_1-Lipschitz convex function over the unit \\ell_1 ball. Previous lower bounds for non-constant \\epsilon for this \\ell_1-Lipschitz convex optimization problem either made additional assumptions on the domain, or instead attained a bound of \\Omega(\\epsilon^-2/3) [DG19].
PaperID: 3100, Poster
Authors: Marco Celotto, J. S Sooter, Sofie Ährlund-Richter, Kyle R Jenks, Mriganka Sur, Stefano Panzeri
Abstract: Identifying subpopulations of neurons that interact with each other from simultaneous recordings of populations of many neurons is key for understanding across-brain communication with cellular resolution. Recent work identified communication subspaces, which capture additive interactions between pairs of high-dimensional neural populations through a small number of source and target activity patterns. However, no current method captures how a third, potentially multivariate variable - such as behavioral state or the activity of a third population - modulates these interactions. Here we extend the communication subspace framework by parameterizing modulation as a low-rank tensor. This identifies multiplicative interaction channels (MICs), defined as triplets of source, target, and modulator activity patterns, in which the modulator pattern gates the source-target interaction. We derive MICs as a bilinear perturbation of reduced-rank regression. We develop a hierarchical fitting pipeline and provide a closed-form decomposition that quantifies whether modulation reshapes the modulator-averaged baseline interaction, recruits private dimensions of one population, or opens new interactions. In simulations, MICs reliably recover the presence and geometry of ground-truth modulation even in the high-dimensional, low-sample regime. Applying MICs to simultaneous calcium imaging of prefrontal axons and interneurons in the visual cortex revealed that behavioral state asymmetrically modulates top-down interactions, reconfiguring the patterns of prefrontal projections that interact with a stable set of visual interneuron activity patterns. By providing an efficient and compact characterization of modulatory interactions, MICs enable asking new questions about how potentially high-dimensional variables shape interactions between neural populations.
PaperID: 3101, Poster
Authors:
Jaehun Gim, Jaewoo Kim, Gyuwan Kim, Seo Yeon ParkAbstract: While quantization compresses LLMs and accelerates inference, outlier activations impede efficient low-bit representation. In particular, massive outliers cause substantial performance degradation. Although rotation-based methods that rely on randomized Hadamard transforms for online inference aim to mitigate this, they consistently produce fragmented distributions with empty bins when encountering massive outliers. This fragmentation severely undermines uniform quantization performance. Moreover, when massive outliers occur in down-projection activations, the rotation must be applied at inference time. This constraint prevents training-based approaches such as SpinQuant from learning the online rotation at this position, as a learned dense matrix would forgo the Fast Walsh--Hadamard Transform and require a full matrix multiplication during inference. To address this, we propose SignRot (Sign-adjusted Rotation), which identifies massive outlier channels, and constructs a sign adjustment vector that unifies the signs of the corresponding Hadamard row, collapsing the bimodal distribution into a unimodal profile amenable to quantization. As a lightweight, drop-in improvement compatible with any Hadamard-based quantization method, SignRot requires only a single calibration forward pass and introduces no additional inference overhead. Applying SignRot to QuaRot reduces the Wikitext2 perplexity to 5.88 and increases the average zero-shot reasoning accuracy to 72.73, surpassing SpinQuant on the LLaMA3-70B model.
Abstract: Many inference-time tasks for pretrained discrete diffusion models and diffusion language models reduce to drawing samples from a tilted version of the pretrained distribution. Feynman--Kac sequential Monte Carlo (SMC) makes this correction exact in principle, but its prescribed weights routinely degenerate when the proposal dynamics are misaligned with the tilt, capping the practical gains from additional particles. We introduce FluxLite, a lightweight, training-free proposal-control framework for discrete diffusion. On the sparse directed graph of pretrained reverse rates, any sparse jump-rate perturbation can be exactly compensated by a q_t-weighted graph-divergence term in the Feynman--Kac potential; the target path is therefore preserved while the residual reweighting variance becomes a local convex objective. We instantiate this principle as two practical samplers: a one-hop local reallocation rule (HEU) and a small nonnegative quadratic program over pretrained-rate bases (D-VCG). We further prove population stability under the standard score-entropy training loss, identifying a tilted-path coverage factor that governs robustness to score error, together with finite-particle convergence for a fixed controlled Feynman--Kac recursion. Empirically, FluxLite improves over standard Feynman--Kac SMC baselines by up to two orders of magnitude in terminal KL on an analytically tractable finite-state CTMC benchmark, and reduces connected-correlation MSE on 2D Ising sampling by 5--7× in geometric mean and up to 55× at peak.
Abstract: Classifier-free guidance (CFG) is the de facto standard for conditional sampling in diffusion models, yet it often reduces sample diversity. Using tools from statistical physics, we analyze the emergence of generative distortions induced by CFG, namely the mismatch between the CFG sampling distribution and the true conditional distribution. We study this phenomenon in analytically tractable settings with exact score functions, characterizing its dependence on data dimensionality and the number of classes. For high-dimensional Gaussian mixtures, we use dynamic mean-field theory to show that distortions arise when the number of classes scales exponentially with the data dimension, whereas they vanish in the sub-exponential regime due to a dynamical phase transition. We further prove that, in the infinite-class limit, distortions remain unavoidable regardless of dimensionality because of the increasing density of classes. Finally, we show that standard CFG schedules cannot prevent variance shrinkage, and we propose a theoretically grounded guidance schedule incorporating a negative-guidance window that improves both class separability and sample diversity in real-world latent diffusion models.
PaperID: 3104, Poster
Abstract: Individual-level wildlife identification often suffers from data scarcity, as varying observations of the same animal under diverse poses, viewpoints, and motions are rarely available. Image-to-video (I2V) generation offers a promising way to mitigate this limitation by synthesizing additional observations from a single reference image. However, existing I2V models mainly emphasize global layout, semantics, and motion, and therefore often fail to preserve fine-grained local appearance cues that distinguish one wildlife individual from another, such as fur texture, stripe boundaries, spot configurations, and contour transitions. We observe that these identity-critical cues are closely related to high-frequency information, which is essential for distinguishing individuals. To address this challenge, we propose WildIcon, a high-frequency-guided I2V framework for wildlife individual consistency. Specifically, WildIcon introduces a frequency-aware identity encoding branch that extracts individual-specific high-frequency cues from the reference image. Combined with isolated foreground information, the resulting identity tokens are then injected into cross-attention blocks as identity conditioning. Building on a frozen backbone with lightweight identity adaptation, WildIcon strengthens fine-grained identity preservation while retaining the motion controllability and semantic fidelity of the base I2V model. In addition, to support the training and evaluation of wildlife individual-consistent I2V, we construct WildlifeVid, a wildlife-centric video dataset with high-quality, temporally coherent clips and individual-level identity labels. Experiments on both I2V and downstream animal re-identification (ReID) show that WildIcon achieves stronger individual consistency than existing baselines and provides useful, reliable synthetic data to improve downstream ReID performance.
Authors:
Xinlin Li, Timothy Chou, Joshua W Fromm, Zichang Liu, Yunjie Pan, Christina FragouliAbstract: Post-training weight quantization is crucial for reducing the deployment cost of large language models (LLMs), yet maintaining model quality in the ultra-low-bit regime (\le 3 bits) remains challenging due to highly non-uniform weight sensitivity and the lack of principled precision allocation strategies. In this work, we show that the weight sensitivity is inherently dynamic and depends on the current quantized model, rather than the full-precision reference used in existing methods. Leveraging this observation, we introduce a sensitivity estimate that better captures the changes in marginal loss during progressive quantization. We further uncover a structured pattern in sensitivity distribution, revealing a bi-directional concentration across input and output channels, which enables hardware-aligned yet expressive partitioning via channel reordering. Building on these insights, we propose ScaleBITS, a scalable framework that combines structured partitioning and an efficient approximation to greedy allocation. Our approach enables fine-grained, global bitwidth search while preserving hardware efficiency. Experiments show that ScaleBITS significantly improves over uniform-precision quantization (up to +36%) and outperforms state-of-the-art sensitivity-aware baselines (up to +13%) in the ultra-low-bit regime, without adding runtime overhead.
PaperID: 3106, Poster
Authors:
Haoyi Zhao, Yishan Jiang, Jiqian Yang, Ji Chang, Dawei Ma, Wenjun LvAbstract: Multivariate time series forecasting is important in real-world applications, but practical data often show both temporal and spectral non-stationarity. Existing methods may treat severe spectral drift as informative evolution, allow unreliable spectral components to contaminate temporal representations through credibility-agnostic time--frequency interaction, and infer cross-variable dependencies from noisy local correlations, which may lead to spurious long-term dependencies. To address these issues, we propose TFNS, a dual-stream time--frequency framework for suppressing unreliable spectral evidence under non-stationarity. Specifically, we introduce a Spectral Credibility Estimator to quantify patch-level spectral stability and softly suppress low-credibility frequency components for more reliable frequency modeling; a Credibility-Gated Cross-Interaction module to enable bidirectional time--frequency enhancement while adaptively constraining cross-domain information injection according to spectral reliability; and a Spectral Cointegration Graph to capture stable long-term cross-variable structure in the spectral domain while mitigating spurious dependencies under non-stationary drift. Together, these modules form a progressive pipeline from local spectral purification to reliability-aware cross-domain interaction and robust long-term structure modeling. Experiments on 11 benchmarks show that TFNS achieves state-of-the-art performance, ranking first in 44/55 MSE and 43/55 MAE comparisons.
PaperID: 3107, Poster
Abstract: Real-world deployment of federated learning (FL) requires resilience against adversarial clients and data privacy. Byzantine fault captures the worst-case behavior of adversarial clients, while differential privacy (DP) provides statistical guarantees against privacy loss. In particular, user-level DP guards a client's entire contribution and provides more realistic protection against information leakage. Despite intensive efforts, the minimax-optimal convergence bound of Byzantine-resilient FL with user-level DP remains open. We establish an information-theoretic lower bound for any Byzantine-resilient empirical risk minimization problem under user-level DP and heterogeneous clients. We propose ByzULDP, a central DP algorithm featuring a two-level momentum across the client and server. Client-side momentum mitigates stochastic gradient variance during robust aggregation. Unlike the fault-free case, naively combining non-linear Byzantine-resilient aggregation with central DP noise is insufficient to guarantee convergence to a bounded neighborhood of a stationary point. Our key contribution is a server-side momentum that leverages a property unique to user-level DP, where privacy sensitivity depends on how client updates are constructed, to empower convergence. ByzULDP achieves the order-optimal convergence bound for strongly-convex objectives, with vanishing error terms decaying at the optimal \mathcalO(1/T) and \mathcalO(1/\sqrtT) rates for strongly-convex and non-convex objectives, respectively. We corroborate our analysis with numerical experiments on real-world datasets across diverse privacy budgets.
PaperID: 3108, Poster
Abstract: Reinforcement learning for long-horizon agents relies on \emphpurely retrospective training signals: credit is assigned only after observing environmental consequences, leaving the agent's belief at action time invisible to the gradient. We introduce Prospective Hindsight (PH), a self-calibrating training principle that augments any retrospective base method with a signal derived from the gap between the agent's \emphprospective prediction (before feedback) and the \emphretrospective evaluation (after feedback). This per-rollout \emphsurprise identifies samples where the agent's self-model is most inaccurate and amplifies their gradient contribution through a stop-gradient surprise-weighted advantage. Since the prospective predictor shares parameters with the policy, the two co-evolve, progressively shifting focus to the agent's remaining blind spots. We connect this principle to a privileged-information gap and show that minimizing the surprise residual provides a descent pathway on the agent's miscalibration rate; calibration thus emerges as a byproduct of optimization rather than from an added objective. On single-turn verifiable tasks and a multi-turn personal-agent task (under GRPO, on-policy distillation, and their combination), PH improves both task performance and calibration, with consistent gains across model scales. Notably, the dominant miscalibration mode shifts structurally between regimes, overconfident failures in single-turn, underconfident successes in multi-turn, yet the same training principle addresses both successfully.
Abstract: Reinforcement learning from verifiable rewards (RLVR) suffers from sparse outcome signals, creating severe exploration bottlenecks on complex reasoning tasks. Recent on-policy self-distillation methods attempt to address this by utilizing language feedback to generate dense, token-level supervision. However, these approaches rely on a fixed, passive teacher to interpret the feedback. As the student policy improves, the teacher's zero-shot assessment capabilities plateau, ultimately halting further learning. To overcome this, we propose Variational Policy Distillation (VPD), a framework that formalizes learning from language feedback as a Variational Expectation-Maximization (EM) problem. VPD co-evolves both policies: in the E-step, the teacher is actively refined on trajectory outcomes via an adaptive trust-region update, translating textual feedback into a dynamically improved target token distribution. In the M-step, the student internalizes this dense distributional guidance on its own on-policy rollouts. By continuously improving the teacher's ability to extract actionable signals from textual critique, VPD overcomes the limitations of passive distillation. Evaluated across diverse sources of diagnostic feedback on scientific reasoning and code generation tasks, VPD consistently outperforms both standard RLVR and existing self-distillation baselines. Finally, by stress-testing our framework on rigid mathematical reasoning and cold-start regimes, we illuminate the fundamental bounds of feedback-driven self-distillation compared to pure environment-driven RL.
PaperID: 3110, Poster
Abstract: Backdoor attacks threaten the integrity of machine learning models by allowing attackers to control model behavior through triggers. Models can be compromised during finetuning by backdoor attacks through either poisoned data or adversarial training objectives. Because existing finetuning-activated attacks assume limited domain shift or frozen encoder layers, they often fail under full-model finetuning. We target a stealthy finetuning-activated attack where a dormant backdoor is implanted in a pretrained encoder, and later activated by finetuning on clean downstream data. We propose Poison-then-Hide, a novel attack that remains effective when the entire model is finetuned in a domain transfer. Our approach consists of three components: trigger optimization, base encoder poisoning, and targeted unlearning to conceal the backdoor. We evaluate our method on six datasets and three model architectures, and achieve state-of-the-art (up to 97%) attack success rates. We find that two design choices - jointly learning the benign target task and the backdoor during encoder poisoning, and optimizing the trigger for attack robustness using an ensemble of simulated finetuned models - are critical to the attack's success. We demonstrate that standard detection and mitigation defenses cannot fully remove the backdoor, which can reappear after benign finetuning.
Authors:
Lukas Seier, Brandon Kaplowitz, Sebastian Towers, Richard M Bailey, Jakob FoersterAbstract: Modelling opinion dynamics typically relies on hand-crafted local interaction rules to study emergent macroscopic phenomena such as consensus and polarisation. In contrast, multi-agent reinforcement learning (MARL) enables agents to learn such behaviours directly by optimising simple rewards. To explore the potential of MARL for opinion dynamics, we introduce a GPU-accelerated that scales to populations of up to 1000 agents, comparable to many real-world social sub-networks. To prevent unrealistic conventions, we extend other-play to general-sum social interactions. We next validate our model on a subset of the Bluesky network by recovering agent importance structures from graph topology alone via a learned attention layer, finding that populations most closely match human data. In large social media networks such high levels of conforming significantly reduce collective accuracy and promote dishonest agents that lie to fit in. By contrast, small, dynamic hunter-gatherer networks are less affected; here, conformity can even improve collective agreement. This suggests a mismatch between evolved human conformity heuristics and modern social media environments as a potential contributor to misinformation.
PaperID: 3112, Poster
Abstract: Functional brain connectome represents neural connectivity as a matrix of pairwise interactions between brain regions. Generation of functional connectomes is not only a question of validity; rather, having satisfied the constraints on the correlation matrix, the next step is to recover the class-conditional geometry buried under coarse labels. We propose MAGNET, a Manifold-Aware Graph Diffusion Network which uses a normalized-Cholesky representation of the manifold of correlation matrices that guarantees validity. MAGNET lifts noisy latent states into ROI-level region tokens and performs denoising with a relational inductive bias over brain atlas regions. To deal with structural problems induced by coarse labels, MAGNET employs class-anchored conditioning, amortized structural bridge, and relevance-preserving corruption. Across ABIDE, ADNI, and OASIS-3, MAGNET consistently achieves favorable results compared to previous manifold-aware approaches, demonstrating improvements of 7-21% in class-conditional fidelity (\alpha,\beta-F1) across three cohorts and better sampling efficiency. Moreover, while training only with strict binary labels, MAGNET is capable of maintaining clinical heterogeneity through fine substructure of the connectomes in a zero-shot setting, improving subclass covariance alignment (\lambda-MSE) by over 30%. These results suggest that geometric validity is a necessary but insufficient condition for clinical utility in connectome synthesis. Moreover, efforts in making the diffusion denoising class-conditional manifold aware finds utility beyond the highly curved brain connectome generation as this is a critical problem in various general settings.
Authors: Yaron Kiselman, Kfir Y. Levy
Abstract: Collaborative learning is sustainable only if it benefits each participant; yet, standard Federated Learning (FL) optimizes a global average that often fails to serve individual clients. In heterogeneous settings, a client may rationally prefer training alone rather than contributing to a global model that targets "average-case" optimality but under performs locally. In this work, we address the problem of \emphSelfish Personalization (SP): How can a target client leverage peer data to minimize its own risk? While current techniques often rely on heuristic performance proxies or clustering that lack sharp theoretical support, we propose \emphSP-Convergence-Aware Client Weighting (SP-CACW). This novel framework determines the contribution of peer clients by explicitly minimizing a theoretical convergence bound for the target. By doing so, our approach efficiently separates useful signal from imported bias on-the-fly during the training process. We provide convergence guarantees that establish the theoretical superiority of \emphSP-CACW, alongside empirical results on the MNIST CIFAR datasets and LEAF Shakespeare.
Abstract: Joint probabilistic modeling is essential for forecasting irregular multivariate time series (IMTS) to accurately quantify uncertainty. Existing approaches often struggle to balance model expressivity with consistent marginalization, frequently leading to unreliable or contradictory forecasts. To address this, we propose CircuITS, a novel architecture for probabilistic IMTS forecasting based on probabilistic circuits. Our model is flexible in capturing intricate dependencies between time series channels while structurally guaranteeing valid joint distributions. Experiments on four real-world datasets demonstrate that CircuITS achieves superior joint and marginal density estimation compared to state-of-the-art baselines.
PaperID: 3115, Poster
Authors: SHENGPENG WANG, Lingzhen Li, Chunshen Li, Fei Xiao, Wei Wang
Abstract: Scene-level Velocity field estimation from 4D millimeter-wave radar is critical for robust perception but remains challenged by the inherent sparsity and noise of point-based paradigms. Furthermore, Doppler ambiguity induced by the Nyquist sampling limit fundamentally disrupts the continuity of velocity field learning. To address these challenges, we propose RaVF, a physically grounded framework that directly learns dense velocity fields from high-fidelity spatial-Doppler spectra. RaVF is built upon three key designs tailored to radar-specific signal distortions. (1) To mitigate range-dependent propagation attenuation and nonlinear angular spatial quantization, we devise a Physically-aware Encoder that calibrates spectral feature extraction within the native radar beamforming. (2) Observing that ego-induced radial motion explains static Doppler responses, we factor out static motion and guide dynamic flow estimation for physically consistent velocity field learning. (3) To enhance kinematic plausibility, we introduce an ego-conditioned bidirectional pyramidal flow module that transfers backward flow into the forward frame and enforces anti-symmetric forward-backward consistency. Extensive experiments demonstrate that RaVF significantly outperforms state-of-the-art baselines on velocity fields. Our code will be available.
PaperID: 3116, Poster
Abstract: Lifelong cross-modal hashing demands incremental updates as new tasks arrive, yet existing regularization-based methods such as EWC and SI share a fundamental weakness: their effective regularization strength is entangled with task scale, as the empirical Fisher Information Matrix (FIM) scales linearly with sample size. This forces costly per-task hyperparameter tuning in streaming deployments to avoid over- or under-regularizing tasks of varying sizes. We propose \textscSing-CH, a lifelong hashing framework built on information geometry. At its core is a Gibbs probabilistic reconstruction of the pairwise objective, which gives the parameter space a statistical manifold structure and yields a natural 1/n_t^2 normalization. We prove this normalization renders the FIM scale-agnostic: its spectral properties remain independent of task size even under diagonal and K-FAC approximations. Beyond this invariance, a Forgetting-Rigidity trade-off theorem shows that natural gradient applied to the composite objective reintroduces scale dependence. This makes our decoupled design, pairing the FIM as regularizer with Adam for optimization, a structural necessity rather than a convenience. Experiments on three benchmarks demonstrate that \textscSing-CH outperforms 11 baselines by 4.1% average mAP without per-task tuning. Under extreme task-scale imbalance, degradation is only 2.0% versus 3.9% for the strongest baseline, providing the first empirically grounded solution to scale-agnostic lifelong retrieval.
Abstract: Video multimodal large language models have made strong progress on open-ended video understanding, but they still lack precise local spatiotemporal perception. When two videos share almost the same global semantics and differ only in a short time span or a small region, current models often fail to find the change and provide reliable evidence. We propose , a verifiable proxy-task framework that enhances fine-grained spatiotemporal perception with cross-video differences. The key idea is to turn cross-video spot-the-difference into a trainable perception signal, where a model identifies local changes, judges temporal boundaries, and organizes spatial evidence by comparing similar videos. To make this signal scalable to train and reliable to evaluate, we further introduce DeltaVid-10K and DeltaVid-Bench, which convert controllable local differences in real videos into evidence-labeled training and test samples. Experiments show that DeltaVid substantially improves performance on cross-video difference understanding and transfers the learned local evidence ability to general video understanding benchmarks, including MMVU, MLVU, Video-MME, VideoHolmes, VideoMMMU, LVBench, TempCompass, and LongVideoBench. These results show that cross-video differences are not only an effective way to diagnose fine-grained perception failures, but also a scalable proxy supervision that moves Video MLLMs from coarse semantic understanding toward fine-grained spatiotemporal evidence reasoning.
Abstract: Long-context recall in linear-time sequence models highlights a tradeoff in how they write to memory. State-based linear models, such as state-space models (SSMs) and linear Transformers, write densely, updating the entire state for each newly arrived token, which leads to interference and makes specific past tokens hard to recover. Sliding-window attention (SWA) exhibits the opposite behavior: it writes sparsely by storing explicit token representations, but only within a fixed window, so recall drops once the relevant token is evicted. Interpolating between these models, we introduce Raven, a linear-time sequence model that maintains a fixed set of memory slots and, at each step, decays and updates only a selected subset via learned, input-dependent routing. This lets Raven mitigate SWA's position-based overwriting and hard eviction while reducing interference from dense state updates in SSMs, thereby preserving long-range content much more effectively. Across recall-intensive benchmarks, Raven is competitive with or outperforms prior linear-time baselines, achieving strong long-context recall where both SWA and SSMs sharply degrade. It remains effective when extrapolating to context lengths as large as 16x its training length, with similar gains in hybrid architectures.
PaperID: 3119, Poster
Abstract: AI systems must often maintain bounded memory of unbounded interaction histories. However, optimizing the memory representation, and its updates, for downstream task utility is difficult for standard reinforcement learning (RL) due to severe credit assignment bottlenecks induced by sparse, delayed rewards. We introduce MaxIM (Maximally Informative Incremental Memory), a framework that formulates incremental summarization as a capacity-constrained, information-theoretic reward-shaping problem. Grounded in value suboptimality bounds, we prove that minimizing sequential information loss requires maximizing the predictive log-likelihood of future, task-relevant targets (predictive sufficiency). We introduce a recursive consistency condition to ensure summaries are grounded by maximizing mutual information with the observed history. We mitigate credit assignment issues with a parameterized critic that computes step-wise variational lower bounds for dense reward shaping. Empirical evaluation on long-horizon benchmarks shows that MaxIM outperforms state-of-the-art baselines on comprehensive autorater metrics and downstream task utility, without expensive human preference labels.
Abstract: Efficient exploration is a central problem in reinforcement learning and is often formalized as maximizing the entropy of the state-action occupancy measure. While unconstrained maximum-entropy exploration is relatively well understood, real-world exploration is often constrained by safety, resource, or imitation requirements. This constrained setting is particularly challenging because entropy maximization lacks additive structure, rendering Bellman-equation-based methods inapplicable. Moreover, scalable approaches require policy parameterization, inducing non-convexity in both the objective and the constraints. To our knowledge, the only prior model-free policy-gradient approach for this setting under general policy parameterization is due to Ying et al. (2025). Unfortunately, their guarantees are limited to weak regret and ergodic averages, which do not imply that the final output is a single deployable policy that is near-optimal and nearly feasible. In this work we take a different approach to this problem, and propose Policy Gradient Penalty (PGP) method, a single-loop policy-space method that enforces general convex occupancy-measure constraints via quadratic-penalty regularization. PGP constructs pseudo-rewards that yield gradient estimates of the penalized objective, subsequently exploiting the classical Policy Gradient Theorem. We further establish the regularity of the penalized objective, providing the smoothness properties needed to justify the convergence of PGP. Leveraging hidden convexity and strong duality, we then establish global last-iterate convergence guarantees, attaining an \epsilon-optimal constrained entropy value with \epsilon-bounded constraint violation despite policy-induced non-convexity. We validate PGP through ablations on a grid-world benchmark and further demonstrate scalability on two challenging continuous-control tasks.
Authors:
Dawid Uchal, Marcin Możejko, Krzysztof Gogolewski, Piotr Kupidura, Fabian Ozga, Szymon Łukasik, Jakub Giezgała, Tomasz Nocoń, Kacper Pietrzyk, Robert Pieniuta, Mateusz Sulimowicz, Michał Zmysłowski, Michał Orzyłowski, Tomasz Siłkowski, Karol Zagródka, Eike Staub, Ewa SzczurekAbstract: We present ImmuVis, a family of efficient foundation models for imaging mass cytometry (IMC), a high-throughput multiplex imaging technology that handles molecular marker measurements as image channels and enables large-scale spatial tissue profiling. Unlike natural images, multiplex imaging lacks a fixed channel space, as real-world marker sets vary across studies, violating a core assumption of standard vision backbones. To address this, ImmuVis introduces marker-adaptive hyperconvolutions that generate convolutional kernels from learned marker embeddings, enabling a single model to operate on arbitrary measured marker subsets without retraining. We pretrain ImmuVis on the largest IMC dataset to date, IMC17M (28 cohorts, 24,405 images, 265 markers, over 17M patches), using self-supervised masked reconstruction. ImmuVis outperforms state-of-the-art baselines and ablations in virtual staining and downstream classification tasks at substantially lower compute cost than transformer-based alternatives, and is the sole model that provides calibrated uncertainty via a heteroscedastic likelihood objective. These results position ImmuVis as a practical framework for real-world IMC modeling.
Authors: Nina Corvelo Benz, Eleni Straitouri, Manuel Rodriguez
Abstract: It is widely agreed that when AI models assist decision-makers in high-stakes domains, they should communicate the confidence of their predictions. However, empirical evidence suggests that decision-makers often struggle to determine when to trust a prediction based solely on this communicated confidence. In this context, recent theoretical and empirical work suggests a positive correlation between the utility of AI-assisted decision-making and the degree of alignment between the AI confidence and the decision-makers' confidence in their own predictions. Crucially, these findings do not yet elucidate the extent to which this alignment influences the complexity of learning to make optimal decisions through repeated interactions. In this paper, we address this question in the canonical case of binary predictions and binary decisions. We first show that this problem is equivalent to a two-armed online contextual learning problem with full feedback, and establish a lower bound of \Omega ( \sqrt|\mathcalH| \cdot |\mathcalB| \cdot T ) on the expected regret any learner can attain, where \mathcalH and \mathcalB denote the sets of possible human and AI confidence values. We then demonstrate that, by leveraging the alignment between AI and human confidence, a learner can attain an expected regret of \mathcalO(\sqrt|\mathcalH| \cdot T\log T) and, in the regime where \mathcalB is countable and \sqrt|\mathcalH| = \mathcalO(\log T), a non-trivial generalization of the Dvoretzky-Kiefer-Wolfowitz inequality improves the regret bound to \mathcalO(\sqrtT\log T). Taken together, these results reveal that alignment can reduce the complexity of learning to make decisions with AI assistance. Experiments on real data from two different human subject studies where participants solve simple decision-making tasks assisted by AI models with different degrees of alignment complement our theoretical results.
Abstract: Large reasoning models (LRMs) generate reasoning traces before producing final answers, yielding strong performance on multi-step reasoning tasks. However, preference alignment for LRMs remains underexplored. The ideal answer-level preference objective marginalizes over reasoning traces, but this trace-marginalized objective is computationally intractable. As a result, practical methods rely on single-trace surrogates, which can induce high gradient variance due to stochastic trace sampling. We cast this challenge as a gradient-estimator design problem and propose Bias-Variance Optimized Preference Optimization (BVPO), which combines a standard trace-based gradient estimator with an empty-trace gradient estimator from a separately sampled empty-trace branch. Under our fixed-prompt, fixed-answer conditional view, the empty-trace branch is deterministic with respect to the trace randomness studied in our analysis. We show that the composite estimator contracts the conditional reasoning-trace-noise component, admits a conditional MSE-optimal mixture, and appears in the estimator-dependent term of a standard biased-SGD bound. Empirically, BVPO improves alignment over baselines on alignment benchmarks, including AlpacaEval 2 and Arena-Hard. Although trained on general conversational data, BVPO also generally preserves reasoning performance on six math reasoning benchmarks. Empirical diagnostics further show that BVPO yields lower total gradient variance during training.
PaperID: 3124, Poster
Abstract: Existing image watermarking methods typically entangle visual content with spatial structure, making them highly sensitive to geometric transformations and alignment discrepancies, especially in the presence of deepfake manipulations. In this paper, we propose a semantic-driven facial watermarking framework for robust identity recovery. The key idea is to decouple identity-related semantic information from spatial layout and encode it into a compact and spatially robust semantic representation. Specifically, we decompose a face into semantic components and aggregate deep features into component-wise latent representations, which are quantized via independent codebooks and converted into a compact bitstream for embedding. After decoding, the embedded semantic information is recovered and used to reconstruct identity-consistent facial content, even under slight geometric distortions and tampering. Experiments on CelebA-HQ and FFHQ demonstrate that our method significantly outperforms existing watermarking approaches in terms of reconstruction quality, identity preservation, and retrieval accuracy under both photometric and geometric attacks, validating the effectiveness of semantic component-wise encoding for reliable face recovery.
Abstract: Long-horizon agents increasingly operate across many steps, tools, and observations. In this setting, the relevant oversight question is not only whether each action is locally valid, but whether the evolving trajectory still corresponds to the task the user authorized. Drift can accumulate quietly: an agent may call the right tool with plausible arguments at every step, while its prefix moves toward a broader role, an adjacent objective, or evidence the user never supplied. Existing monitors mostly check local compliance, deliver final-trace verdicts, or score generic risk; they do not directly estimate this prefix-level relation. We introduce ontological trust, a task-conditioned property of trajectory prefixes, and instantiate it as RGE, an online monitor that decomposes trust along Role, Goal, and Evidence. RGE uses LLMs only to derive structured representations. State updates, projections, and intervention decisions are deterministic, so the output is a replayable and auditable trust trajectory rather than a single end-to-end judge verdict. We construct a cross-domain trajectory corpus from OSWorld, FinanceBench, and EICU-AC, covering benign executions, prefix-paired drift, and pseudo-consistency failures. On this corpus, RGE outperforms adapted rule-, judge-, and shield-style baselines on prefix-paired drift detection. With the two larger estimator models, it exceeds 93% Drift F1 on every benchmark while keeping benign coverage at or above 95.8%. Pseudo-consistency is harder: detection depends on whether task completion is externally visible, a structural limit we characterize empirically.
PaperID: 3126, Poster
Abstract: Diffusion Language Models (DLMs) exhibit strong parallel decoding capabilities by denoising multiple tokens in a single generation step. However, this parallelism comes with substantial computational overhead, as each step requires interactions with all suffix tokens. Existing methods typically reduce this cost by retaining only a local suffix window as a substitute for the full suffix. Despite their effectiveness, these methods overlook the structural heterogeneity across suffix regions and re-initialize suffix tokens with identical representations at each timestep. To this end, we propose a structured suffix modeling method for efficient DLM inference. Specifically, we divide the suffix into three regions, i.e., the local, middle, and tail regions, and retain different numbers of suffix tokens in each region according to their structural roles. Moreover, we incorporate the decoding results from the previous step into the suffix token representations at the current step, allowing them to carry evolving denoising information across generation steps. Notably, our method is training-free and orthogonal to several existing acceleration techniques, such as parallel decoding strategies and KV cache. Empirical results across multiple benchmarks on three DLMs demonstrate that our method can further accelerate DLM inference and, in most cases, also improve task performance. In particular, in long-sequence inference, our method achieves up to a (72.16×) speedup when combined with other acceleration techniques.
Authors:
Zuojin Tang, Haoyun Liu, Xinyuan Chang, Changjie Wu, Dongjie Huo, Yandan Yang, Bin Liu, Zhejia Cai, Feng Xiong, Mu Xu, jiachen luo, De Ma, Zhiheng Ma, Gang PanAbstract: Vision-language-action (VLA) models remain constrained by the scarcity of action-labeled robot data, whereas action-free videos provide abundant evidence of how the physical world changes. Latent action models offer a promising way to extract such priors from videos, but reconstruction-trained latent codes are not necessarily suitable for policy generation: they may predict future observations while lacking the structure needed to be reused or generated coherently with robot actions. We introduce ALAM (Algebraic Latent Action Model), an Algebraically Consistent Latent Action Model that turns temporal relations in action-free video into structural supervision. Given frame triplets, ALAM learns latent transitions that are grounded by reconstruction while being regularized by composition and reversal consistency, encouraging a locally additive transition space. For downstream VLA learning, we freeze the pretrained encoder and use its latent transition sequences as auxiliary generative targets, co-generated with robot actions under a joint flow-matching objective. This couples structured latent transitions with flow-based policy generation, allowing the policy to exploit ALAM's locally consistent transition geometry without requiring latent-to-action decoding. Representation probes show that ALAM reduces additivity and reversibility errors by 25--85× over unstructured latent-action baselines and improves long-horizon cumulative reconstruction. When transferred to VLA policies, ALAM raises the average success rate from 47.9% to 85.0% on MetaWorld MT50 and from 94.1% to 98.1% on LIBERO, with consistent gains on real-world manipulation tasks. Ablations further confirm that the strongest improvements arise from the synergy between algebraically structured latent transitions and joint flow matching.
PaperID: 3128, Poster
Abstract: Learning visual representations that generalize across domains remains challenging when source-domain supervision is entangled with spurious dataset-specific cues. This issue is particularly pronounced in face attack detection (FAD), where attack artifacts vary substantially across capture devices, manipulation pipelines, and image quality. In this paper, we propose Prompt-Conditioned Semantic Bottleneck, an information-bottleneck-inspired representation learning framework for cross-domain FAD. Unlike prior methods that use prompts as image-text matching anchors for multi-modal feature learning, our method uses label- and attack-type-conditioned prompts to elicit attack-aware semantic targets from a frozen VLM. These prompt-conditioned semantics provide offline supervision for a lightweight visual detector, encouraging its representation to preserve attack-relevant cues while reducing reliance on source-domain nuisance factors. Extensive experiments on face anti-spoofing and face forgery detection benchmarks demonstrate consistent improvements under cross-dataset evaluation. Ablation studies further show that the gains are not simply due to stronger visual or multi-modal models, but arise from attack-aware prompt-conditioned semantic supervision. Importantly, the VLM and prompts are removed during inference, enabling efficient vision-only deployment. Overall, our results suggest that prompt-conditioned VLM semantics provide an effective way to improve the trade-off among detection accuracy, cross-domain generalization, and inference efficiency in FAD.
Abstract: Diffusion-based vision-language-action (VLA) models have emerged as strong priors for robotic manipulation, yet adapting them to real-world distributions remains challenging. In particular, on-robot reinforcement learning (RL) is expensive and time-consuming, so effective adaptation depends on efficient policy improvement within a limited budget of real-world interactions. Noise-space RL lowers the cost by keeping the pretrained VLA fixed as a denoising generator while updating only a lightweight actor that predicts the noise. However, its performance is still limited due to inefficient autonomous exploration. Human corrective interventions can reduce this exploration burden, but they are naturally provided in action space, whereas noise-space finetuning requires supervision over noise variables. To address these challenges, we propose UniSteer, a Unified Noise Steering framework that combines human corrective guidance with noise-space RL through approximate action-to-noise inversion. Given a human corrective action, UniSteer inverts the frozen flow-matching decoder to recover a noise target, which provides supervised guidance for the same noise actor that is simultaneously optimized via reinforcement learning. Real-world experiments on diverse manipulation tasks show that UniSteer adapts more efficiently than strong noise-space RL and action-space human-in-the-loop baselines, improving the success rate from 20% to 90% in 66 minutes on average across four real-world adaptation tasks.
PaperID: 3130, Poster
Authors: Muhammad Aqeel, Maham Nazir, Marco Cristani, Francesco Setti
Abstract: Generalist Anomaly Detection (GAD) aims to detect anomalies in unseen domains using only a few reference samples. Recent normal-abnormal guided methods improve this setting by exploiting abnormal references, but their residuals often mix transferable anomaly cues with domain-specific appearance factors such as texture, geometry, and acquisition conditions. This residual domain entanglement limits cross-domain transfer and may produce spurious activations on novel domains. We propose DART, a plug-in framework for normal-abnormal guided GAD that learns to separate residuals into anomaly-relevant and domain-specific components. DART combines a Residual Disentanglement Module, which preserves anomaly-discriminative information while suppressing domain information, with Cross-Domain Meta-Training, which forms episodes across source domains to make the separation effective. The method is architecture-agnostic and can be inserted between residual computation and the detection head of existing GAD pipelines. Experiments on MVTec AD, VisA, and BraTS show consistent improvements over multiple base methods, with the largest gains under stronger domain shifts. Code will be available upon publication.
Abstract: Language agents increasingly operate over streams of related tasks, yet existing memory systems struggle to convert accumulated experience into reusable knowledge. Retrieval-augmented and structured memory methods record per-session observations effectively, but often couple acquisition and consolidation into a single online process, leaving the agent without a global view across sessions to discover recurring patterns, abstract shared procedures, or prune redundant entries. Inspired by complementary learning systems theory, we propose Auto-Dreamer, a learned offline consolidator for language-agent memory. Auto-Dreamer decouples fast per-session memory acquisition from slow cross-session consolidation. Given a region of a typed memory bank, the consolidator performs bounded tool-use to search memory, inspect candidate entries, and trace them back to raw source trajectories, synthesizing provenance-grounded replacement memories that supersede the original region. We train Auto-Dreamer via GRPO, using end-to-end agent performance as the reward signal to learn how to consolidate memories acquired through fast online experience. Trained on ScienceWorld trajectories alone, Auto-Dreamer transfers without retraining to held-out ALFWorld and WebArena, improves task success over fixed, RL-trained, and prompted memory baselines in both continual-memory deployment and fixed-bank consolidation, and does so with an active memory bank an order of magnitude smaller than competitive baselines.
Abstract: Code generation is typically trained in the \emphprimal space of programs: a model produces a candidate solution and receives sparse execution feedback, often a single pass/fail bit. Test-time scaling enriches the inference procedure by sampling multiple candidates and judging among them, but the comparative information this process reveals is discarded after inference. We argue that this information defines a \emphdual judgment space that provides a far richer training signal: the model learns not from an isolated success or failure, but from the relative correctness structure across its own plausible attempts, identifying which succeed, which fail, and what distinguishes them. We introduce DuST (Dual Self-Training), a framework for self-training from the dual judgment space. DuST samples candidate programs from the model's own distribution, labels them through sandbox execution, retains groups containing both successes and failures, and trains the model to rank candidates by execution correctness using GRPO. The objective is purely discriminative: the model is never directly rewarded for generating correct programs. Dual self-training improves both judgment and generation. Across five models spanning two families and three scales (4B to 30B), DuST consistently improves Best-of-4 test-time scaling on LiveCodeBench. For Qwen3-30B-Thinking on LiveCodeBench v6, judgment quality improves by +6.2 NDCG, single-sample pass@1 improves by +3.1, and Best-of-4 accuracy improves by +4.1. The trained model's single rollout matches the base model's Best-of-4 performance. SFT on the same ranking data improves judgment without improving generation, confirming that on-policy RL is the mechanism that transfers dual-space learning back into primal generation.
PaperID: 3133, Poster
Abstract: Structured sparsity extends classical sparse learning by encouraging sparsity at the level of predefined groups rather than individual parameters. This paradigm is widely adopted in modern machine learning and is typically enforced by sparsity-inducing norms. In real pipelines, the optimal group partition is often unknown a priori and does not remain static after initial training. Parameter groupings are often modified as domain knowledge or upstream representation evolves. Standard approaches require retraining the model from scratch under the new target partition, which discards the optimization effort already invested and becomes computationally prohibitive at scale. In this work, we focus on the group-wise \ell_p,1-norm, a flexible regularization family that encompasses classical group Lasso and other sparse variants. We develop a principled algorithmic framework that leverages the pre-trained optimum to fine-tune model weights toward a new desired group structure without full retraining. Technically, we decompose arbitrary group-wise transitions into a sequence of tractable primitives and derive the closed-form learning dynamics for each sub-action. Extensive experiments on real-world datasets demonstrate that our method achieves significant acceleration while provably recovering the correct optimal solution.
PaperID: 3134, Poster
Abstract: Simulating in silico cellular responses to interventions is a promising direction to accelerate high-content image-based assays, critical for advancing drug discovery and gene editing. To support this, we introduce MorphGen, a state-of-the-art diffusion-based generative model for fluorescent microscopy that enables controllable generation across multiple cell types and perturbations. MorphGen is trained with an alignment loss that matches its representations to phenotypic embeddings from a biology foundation model, encouraging meaningful morphological patterns consistent with real cell images. Unlike prior approaches that compress multichannel stains into RGB images, sacrificing organelle-specific detail and focusing on a single cell type, MorphGen generates the complete set of fluorescent channels jointly, preserving per-organelle structure and enabling post-generation interpretation. We demonstrate biological consistency with real images via CellProfiler features, and MorphGen attains an FID score over 35% lower than the prior state-of-the-art MorphoDiff. Finally, in a compositional generalization test that holds out cell type--perturbation combinations during training, MorphGen achieves in-distribution quality on 55% of unseen pairs under a seen-calibrated criterion.
Abstract: Complementarity Determining Regions (CDRs) are critical segments of an antibody that facilitate binding to specific antigens. Current computational methods for CDR design rely on reconstruction losses and do not jointly optimize binding energy, a crucial metric of antibody efficacy. Rather, binding energy optimization is performed via computationally expensive Online Reinforcement Learning (RL) pipelines that rely heavily on unreliable binding-energy estimators. In this paper, we propose AbFlowNet, a novel generative framework that integrates GFlowNet with Diffusion models. By framing each diffusion step as a state in the GFlowNet framework, AbFlowNet jointly optimizes standard diffusion losses and binding energy by directly incorporating energy signals into the training process, thereby unifying diffusion and reward optimization in a single procedure. Experimental results show that AbFlowNet outperforms the base diffusion model by 3.06% in amino acid recovery, 20.40% in geometric reconstruction (RMSD), and 3.60% in binding energy improvement ratio. ABFlowNet also decreases Top-1 total energy and binding energy errors by 24.8% and 38.1%, which is competitive with AbDPO's expensive online RL optimization without pseudo-labeling the test dataset or requiring computationally expensive synthetic CDR sampling.
PaperID: 3136, Poster
Authors:
Abdelrahman Hekal, Vasileios Manginas, Nikolaos Manginas, Alessio LomuscioAbstract: Neuro-symbolic systems for sequence classification couple a neural perception network with a finite automaton over abstract events. Certifying their trajectory-level decisions under input perturbations is critical for safety-sensitive deployment. We address the open problem of end-to-end formal robustness certification for such Neurosymbolic Automata (NeSyA) under norm-bounded input perturbations. Naive layerwise linear bound propagation fails on this composition: per-step relaxation slack compounds through the temporal recurrence (interval explosion), and the standard log-space transition score subtracts two log-sum-exp terms over overlapping coordinate sets that independent relaxations cannot tighten. To address both, we introduce contrastive logit-gap (CLG), a reparameterization whose two log-sum-exp heads act on disjoint coordinates and admit substantially tighter linear bounds, coupled with a fused temporal operator that yields corner-tight bounds from a single linearization. On NeSyA sequence benchmarks, our formulation scales significantly beyond recursive baselines, maintaining tight safety certificates at long horizons while reducing verification wall time by up to ~6×. To our knowledge, this is the first end-to-end formal verification framework for temporal neuro-symbolic systems.
PaperID: 3137, Poster
Abstract: Content-style decomposition from a single image is a fundamental challenge in personalized image generation, aiming to separate subject identity from visual style for flexible recomposition. Existing approaches primarily operate in the spatial domain, without explicitly exploiting the distinct frequency characteristics naturally associated with content and style. In this paper, we propose WaveletLoRA, a novel frequency-aware framework for explicit content-style decomposition through wavelet-domain supervision. Specifically, we first apply a multi-level discrete wavelet transform (DWT) to decompose the reference image into low-frequency and high-frequency subbands. We then assign a dedicated LoRA branch to each subband, where the low-frequency branch models global appearance for style acquisition, while the high-frequency branches capture structural details for content preservation. To further improve disentanglement quality, we introduce a dense Mixture-of-Experts aggregation module together with a frequency-domain regularization objective, which reduces style leakage and enhance content fidelity. In addition, we establish WaveBench and develop a VLM-based evaluation protocol that explicitly measures both attribute similarity and attribute separation for disentanglement assessment. Extensive experiments demonstrate that WaveletLoRA consistently outperforms existing methods and produces results that better align with human preference. Code and the proposed dataset will be released in https://anonymous.4open.science/r/WaveletLoRA-8D7F.
PaperID: 3138, Poster
Authors: Benjamin L Badger
Abstract: The ability of machine learning models to store input information in hidden layer embeddings, a form of model `memory', is widely employed but not well characterized. We find that causal language model embeddings typically contain relatively little input information regardless of data and compute scales during training. In contrast, embeddings from autoencoders trained for input regeneration are capable of nearly perfect memory formation. The substitution of memory embeddings for token sequences leads to computational efficiencies, motivating the introduction of a parallelizable encoder-decoder memory model architecture. Upon causal training these models contain information-poor embeddings incapable of arbitrary information access, but by combining causal and information retention objective functions they learn to form and decode information-rich memories. Training can be further streamlined by freezing a high fidelity encoder followed by a curriculum training approach where decoders first learn to process memories and then learn to additionally predict next tokens. We conclude that next token prediction training alone is poorly suited for accurate memory formation, motivating the use of combined objective functions for models where the entire input is not exposed.
PaperID: 3139, Poster
Abstract: Modern dehazing networks are often computationally expensive and overfit to synthetic haze, limiting real-time deployment. This underscores an urgent demand for efficient, high-fidelity restoration methods that can operate on resource-constrained edge devices in practical, dynamic environments. We introduce PhysDNet, a lightweight multimodal architecture that fuses NIR imagery with sparse LiDAR depth and is trained with a physics-constrained dual-head objective. Differentiating through a fixed Koschmieder inversion induces density-adaptive gradient amplification, acting as an implicit resource allocation mechanism that concentrates learning on heavily degraded regions without increasing inference cost. PhysDNet delivers superior restoration with minimal capacity and enhanced training efficiency, facilitating real-time deployment on edge devices without increasing inference overhead. On the Seeing Through Fog benchmark, PhysDNet outperforms competitors by 1.67 dB at 5.2x lower compute and recovering 54.9% of fog-induced detection loss. Ablations and cross-domain evaluations show that physics-guided training improves generalization including under non-homogeneous haze that violates the Koschmieder assumption. PhysDNet further demonstrates compatibility with fixed-function NPUs, running at 29/45/99 FPS at <4W across L/M/S variants. Code and pre-trained models will be made publicly available.
PaperID: 3140, Poster
Abstract: Rehearsal-free class-incremental learning (CIL) requires a model to learn new classes sequentially without storing previous data, making catastrophic forgetting a central challenge. Adapter-based methods have become a prevalent solution by storing and accumulating task-specific update matrices, yet the effect of the cumulative update magnitude on CIL performance remains underexplored. In this work, we propose Normalized Tensor Adaptation (NoTA), which addresses these two aspects jointly. MPO adapters provide a compact structured space for storing task-specific updates, while global normalization controls the accumulated update that determines the final model. Our empirical analysis reveals a key observation: global normalization consistently improves both matrix-based and MPO-based adapters by stabilizing the cumulative update magnitude across task sequences. We further provide a theoretical analysis showing that global normalization yields a forgetting bound independent of the task sequence length, in contrast to task-wise normalization whose bound can grow with later tasks. Building on the tensor network structure of MPO, we introduce a variant with hierarchical coefficients (H-NoTA), which uses task-level coefficients and inter-core modulation matrices to refine the contributions of historical adapters. Extensive experiments demonstrate that our methods consistently improve rehearsal-free CIL performance while maintaining strong parameter efficiency.
PaperID: 3141, Poster
Abstract: Discrete diffusion models have recently become competitive with autoregressive models for language modeling, even outperforming them on reasoning tasks requiring planning and global coherence, but diffusion requires more computation at inference time. We trace this trade-off to a key mechanism: diffusion models are trained to jointly predict all unknown tokens simultaneously, including those that will not actually be decoded in the current step. Ablating this joint prediction yields faster inference but degrades performance, revealing that accurate prediction at the decoded position relies on joint reasoning about the undecoded tokens. We interpret these undecoded positions as latent tokens and introduce a method for modulating their number, achieving a smooth tradeoff between inference speed and sample quality. Furthermore, we demonstrate that latent tokens can be introduced into autoregressive models through an auxiliary multi-token prediction objective, yielding substantial improvements on the same reasoning tasks where they have traditionally struggled. Our results suggest that latent tokens--arising from jointly predicting multiple unknown positions--represent a general mechanism for improving performance on tasks requiring global coherence or lookahead.
PaperID: 3142, Poster
Abstract: Domain Generalization (DG) aims to leverage multiple source domains to train models that can generalize to unseen target domains, thereby mitigating the performance degradation caused by domain shifts. Most existing approaches impose unified alignment or suppression constraints in the global feature space, overlooking the heterogeneity of feature channels in terms of class discriminativeness and domain stability. Such coarse-grained constraints may lead to the loss of discriminative information or the retention of domain-sensitive features. To address this issue, we propose a novel DG framework consisting of two core modules: Channel Sensitivity Decomposition (CSD), which employs channel-wise response variances along the class and domain dimensions as empirical proxy signals to partition feature channels into four functional subspaces via a continuous and differentiable soft decomposition mechanism; and Subspace-Specific Feature Augmentation (SSA), which implements tailored augmentation and regularization strategies for different subspaces to suppress environment-related spurious correlations while preserving cross-domain stable discriminative features. Additionally, a KL-divergence-based prediction consistency constraint is introduced to stabilize the training process and enhance model robustness. Extensive experiments on multiple standard DG benchmarks demonstrate that our framework consistently outperforms state-of-the-art methods.
PaperID: 3143, Poster
Abstract: Graph neural networks for bot and fraud detection operate on attributed graphs, where adversaries can jointly modify local edges and continuous node features. We study node-level certification under a localized mixed threat model with up to d edge flips in the target message-passing neighborhood and an \ell_2 Frobenius feature budget \epsilon. We instantiate hybrid randomized smoothing for this graph setting and derive two graph-specific certificates: a sound Sequential NP composition baseline for the joint topology--continuous-feature threat model, and a Hybrid NP certificate whose homogeneous edge-smoothing structure reduces the adversarial adjacency search to an O(d^2) optimization over edge-deletion and edge-addition counts. Across PubMed, Amazon Fraud, YelpChi, and MGTAB, the Hybrid NP certificate yields non-trivial certified regions under joint perturbations. We also construct adaptive joint attacks and show failure cases on PubMed, Amazon Fraud, MGTAB and YelpChi for topology-only certified nodes, demonstrating that topology-only guarantees can overestimate robustness when continuous feature perturbations are allowed. The results support joint discrete--continuous certification as the appropriate robustness notion for attributed graphs with continuous features.
Abstract: Momentum is a critical and ubiquitous component of modern optimizers, while the role of momentum remains unclear beyond restricted settings, especially in optimization for large-scale neural networks. Recent studies suggest that the highly non-convex loss landscape for large language models exhibits certain “river-valley” structure: a low-loss manifold (the river) bordered by sharp, high-loss directions (the valley), where the essential optimization progress is determined primarily by the progress along the river in the long run. Motivated by this structure, in this work, we investigate the role of heavy-ball momentum in such an emerging setting. Specifically, we analyze gradient descent with heavy-ball momentum and show that compared to vanilla gradient descent, momentum can accelerate the progress along the river by enabling use of a substantially larger learning rate. In fact, momentum acts as a stabilizer in the presence of oscillations caused by an aggressive choice of learning rate, which the vanilla gradient descent cannot tolerate. We validate the insights with experiments on synthetic functions and language model training, offering practical guidance for tuning learning rate and momentum parameters.
Authors:
Runyi Yu, Xiaoyi Lin, Hok W Tsui, Yinhuai Wang, PENG Zhijie, Hui Zhang, Qihan Zhao, Ke Fan, Miao Li, Jie Song, Jingbo Wang, Qifeng Chen, Ping TanAbstract: We present a system for learning generalizable hand-object tracking controllers purely from synthetic data, without requiring any human demonstrations. Our approach makes two key contributions: (1) HOP, a Hand-Object Planner, which can synthesize diverse hand-object trajectories; and (2) HOT, a Hand-Object Tracker that bridges synthetic-to-physical transfer through reinforcement learning and interaction imitation learning, delivering a generalizable controller conditioned on target hand-object states. Our method extends to diverse object shapes and hand morphologies. Through extensive evaluations, we show that our approach enables dexterous hands to track challenging, long-horizon sequences including object re-arrangement and agile in-hand reorientation. These results represent a significant step toward scalable foundation controllers for manipulation that can learn entirely from synthetic data, breaking the data bottleneck that has long constrained progress in dexterous manipulation.
PaperID: 3146, Poster
Abstract: Scaling depth is a key route to stronger large language models, but Pre-LN architectures often suffer from depth-wise information dilution, limiting the effective use of deep layers. While cross-layer connectivity has proven effective at mitigating this problem, existing designs usually introduce additional computation or memory access. We revisit cross-layer KV sharing as an efficient form of such connectivity. Beyond cache compression, each KV reuse edge creates a direct gradient path from deep layers to source KV projections, turning KV sharing into an implicit structural regularizer for KV representation geometry. We provide a systematic analysis of this mechanism, showing that proper sharing topologies sharpen key retrieval directions while enriching the value content subspace, whereas existing adjacent or global-anchor designs do not fully exploit this effect. Motivated by this, we propose DepthGraft, a hierarchical KV-sharing architecture that reconstructs deep-layer KV from stride-aligned shallow--middle source pairs. DepthGraft distributes non-local gradient flow across source layers, while preserving prefill early completion and reducing KV cache by roughly one third. Experiments on dense and MoE models from 0.5B to 65B parameters show consistent downstream improvements under both GQA and MLA. Notably, on the 16B MoE model, DepthGraft achieves a on C3. These results validate the potential of DepthGraft as a scalable design for large-scale pretraining.
PaperID: 3147, Poster
Abstract: Born-Oppenheimer molecular dynamics delivers quantum accurate trajectories by evolving nuclei on the electronic ground-state surface while retaining electronic observables such as dipoles and polarizabilities. In Kohn-Sham density functional theory, this typically relies on an iterative self consistent field solver at each time step. Machine learning can predict electronic quantities or improve SCF initialization, yet dynamics remain sensitive to residual self consistency error, which can yield non conservative forces and energy drift, and repeated diagonalization limits differentiable and accelerator efficient implementations. We propose a gauge-symmetric dual Lagrangian framework that avoids SCF loops by propagating the physically meaningful electronic state as the occupied subspace and its density projector on a Grassmann manifold. A residual driven update stabilized by Rayleigh type dissipation with time decaying inertia yields a closed evolution of the subspace, its conjugate momentum, and the projector. Coupled with a neural gauge-symmetric Hamiltonian with analytic coordinate derivatives, the method provides closed form forces with Pulay corrections and enables stable quantum BOMD on molecular simulation benchmarks.
PaperID: 3148, Poster
Authors: Yohann PERRON, Guillaume Astruc, Nicolas Gonthier, Clement Mallet, Loic Landrieu
Abstract: Vision Transformers (ViT) dominate computer vision. However, their reliance on rigid patch projectors hinders transfer to Earth Observation (EO), where inputs vary widely in modality, scale, and resolution. We introduce UniverSat, a ViT-style backbone built around a Universal Patch Encoder that maps patches from arbitrary spatial, spectral, and temporal resolutions, and from both optical and non-optical sensors, into a shared embedding space with a single set of weights. This enables training one model on heterogeneous multimodal corpora in self-supervision, yielding robust sensor-agnostic spatial features. We validate this approach with strong results across classification and segmentation on standard EO benchmarks from GeoBench, PANGEABench, and SpectralEarth. The code and model will be released in open-source.
PaperID: 3149, Poster
Abstract: Infrared-visible image fusion (IVIF) aims to combine visible and thermal infrared observations for human perception and downstream vision tasks. Visible and thermal infrared observations are physically complementary but unreliable in different ways. Existing deep IVIF methods mainly improve feature interaction and reconstruction, but often lack an explicit local rule for deciding when visible appearance should dominate and when thermal responses should contribute. To address this issue, we reformulate IVIF as local Bayesian posterior rendering, where fusion is controlled by reliability posteriors rather than direct modality mixing. In this paper, we propose Proposal-Context-Detail Fusion (PCDFusion), a visible-anchored framework that implements this formulation through three rendering stages. PCDFusion preserves potentially useful thermal responses, updates exposure-degraded visible regions, and renders reliable band-limited infrared structures into the final luminance. Extensive experiments on public IVIF benchmarks show that PCDFusion improves fusion quality, target recovery, and downstream detection and segmentation performance under challenging illumination. To support reproducibility, the code will be publicly released upon acceptance.
PaperID: 3150, Poster
Abstract: Large pretrained models have achieved remarkable success across diverse domains, yet adapting them to multiple downstream tasks remains challenging: full fine-tuning is costly, while maintaining separate task-specific models is storage-inefficient and fails to produce a unified multi-task model. Low-Rank Adaptation (LoRA) offers an efficient alternative for Model Merging by representing each task adaptation as a low-rank update over frozen pretrained weights. However, existing LoRA merging methods often treat task-specific updates as directly comparable objects and merge them in their original low-rank parameter spaces. This overlooks two key factors: which update directions provide reliable shared support across tasks, and how these directions should be re-parameterized with respect to the pretrained weights before merging. To address this limitation, we propose Pretrained-Guided Shared Basis (PGSB), a two-stage framework for LoRA model merging. PGSB first identifies shared directions among task-specific LoRA updates, and then re-parameterizes these directions using their interaction with the pretrained weights, producing a more suitable basis for merging. Extensive experiments demonstrate that PGSB consistently outperforms state-of-the-art baselines across multiple benchmarks.
PaperID: 3151, Poster
Abstract: Post-hoc Graph Neural Networks (GNN) explainers typically extract a compact subgraph to preserve the model’s decision rationale. However, this extraction breaks topological integrity and induces a distribution shift, making predictions on subgraphs unreliable and consequently misleading explainer optimization under OOD settings. Existing attempts to build in-distribution proxy graphs often assume independence between explanation and background subgraphs and rely on hard splicing, which causes semantic mismatch and boundary discontinuities that degrade explanation reliability. We propose SSDGExplainer, a Structure–Semantic Dual-Guided explainer that formulates proxy-graph generation as conditional modeling under semantic consistency constraints to achieve deep semantic alignment between the explanation subgraph and the generated background, introduces a topology boundary optimization network to smooth structural fractures at the explanation–background interface, and enforces contrastive semantic constraints to prevent semantic drift. Extensive experiments on synthetic and real-world benchmarks demonstrate consistent improvements over state-of-the-art methods in both explanation quality and fidelity.
Abstract: Optical lens design is a complex, non-convex optimization challenge that relies heavily on human experience and intuition. Existing optimized-based automatic lens design methods struggle to navigate this vast parameter space without meticulous manual tuning. In this paper, we present LensDesigner, an autonomous agent framework that mirrors the problem-solving workflow of expert opticians. To overcome the initial cold start problem, we construct LensLib100K, an extensive optical lens library, and employ Optics-Aware Retrieval to supply physically valid structural seeds. Within an interactive physical simulation environment, the agent executes macroscopic orchestration while receiving immediate optical feedback. Furthermore, we introduce a continuous self-evolving mechanism guided by a curriculum agent. By iteratively solving design tasks with progressively increasing difficulty, the agent autonomously extracts, accumulates, and reuses design heuristics, effectively evolving its optical lens design expertise over time. At the evaluation level, we introduce LensArena, a standardized evaluation benchmark comprising 120 diverse optical design tasks, covering extreme configurations. Extensive experiments on this benchmark demonstrate that LensDesigner significantly outperforms publicly available baseline algorithms, achieving superior success rates and optimization efficiency. We hope this work sheds light on the emerging field of intelligent optics. The code will be publicly available.
Abstract: Symbolic regression (SR) discovers closed-form mathematical expressions from data, offering interpretability beyond black-box models. Existing methods suffer from slow convergence in combinatorial search spaces and lack mechanisms to exploit compositional structure in the data. We introduce SMILE (Sine, Multiplication, Identity, Logarithm, Exponential), a hybrid framework that unifies continuous gradient-based optimization with discrete symbolic recovery through three stages: structural analysis of the data to identify the compositional hierarchy of the target expression, continuous optimization to learn parameters of a network that encodes the target expression using interpretable activations, and symbolic recovery through structured pruning, coefficient optimization, and rounding. This final stage distills the learned network into a compact expression with exact symbolic constants. We evaluate SMILE on SRBench across ground-truth and black-box datasets, with ablation studies validating each component. SMILE achieves the highest symbolic solution rate at the largest noise levels, demonstrating strong robustness where competing methods degrade substantially. It consistently lies on the Pareto front of accuracy versus complexity, recovering significantly simpler expressions in a fraction of the time required by the competing methods.
Abstract: Recent work due to Goel et al.~gave the first efficient algorithms for learning with distribution shift in the challenging PQ framework, where a learner receives labeled training examples and unlabeled test examples and must make correct predictions on the test set but is allowed to abstain from predicting on out-of-distribution points. Their results rely on \cal L_2 sandwiching approximations, a strong requirement that leads to poor bounds for several basic function classes such as DNF formulas. Here we show that the weaker notion of \cal L_1 sandwiching suffices for efficient PQ learning. As a consequence, we obtain the first quasipolynomial-time PQ learning algorithm for DNFs under the uniform distribution and essentially match the guarantees known for ordinary PAC learning. More broadly, our bounds provide exponential improvements for several classes including constant depth circuits and constant degree polynomial threshold functions. Our main technical ingredient is Iterative Chow Filtering, a new procedure that uses low-degree Chow parameters to identify and remove test points incompatible with the training distribution.
Abstract: With the rapid progress of Multimodal Large Language Models (MLLMs), unified MLLMs that jointly perform image understanding and generation have advanced significantly. However, despite the inherent reasoning capabilities of unified MLLMs for self-reflection and self-refinement, their use in text-to-image generation remains largely underexplored. Meanwhile, existing multimodal reasoning–based image generation methods mostly rely on prompt augmentation or holistic image–text alignment judgments, without fine-grained reflection and refinement of detailed prompt attributes, leading to limited fine-grained control. To address this limitation, we propose FiRe, a Fine-grained Multimodal Reasoning method for enhanced image generation by MLLM. In specific, FiRe performs a fine-grained multi-step reasoning by first decomposing the prompt into key visual requirements and then self-judging their satisfaction in the generated image, followed by localized refinement according to self-generated precise feedback. In addition, to further strengthen the MLLM's multimodal reasoning ability, we introduce FiRe-GRPO, a reinforcement learning method tailored to FiRe. Since standard Group Relative Policy Optimization (GRPO) suffers from sparse, outcome-based rewards in multi-step reasoning, we formulate our reasoning process as a step-level decision-making problem, design step-specific rewards, and compute step-level advantages for granular credit assignment within GRPO. Extensive experiments demonstrate that FiRe consistently outperforms competitive text-to-image baselines, including existing reasoning-based methods, with particularly substantial gains on compositional text-to-image benchmarks.
PaperID: 3156, Poster
Authors: Sung I Choi, Junhao Cai, Dohun Kim, Changhee Joo
Abstract: Machine unlearning in large language models (LLMs) aims to remove specific learned knowledge, capabilities, or behaviors while preserving model utility. Existing methods have mostly focused on the design of unlearning objectives, while largely overlooking the intrinsic matrix structure of transformer weights. In particular, standard optimizers operate on independent scalar coordinates, thereby ignoring this matrix structure. We propose \emphspectral unlearning, an update framework that combines standard unlearning objectives with matrix-structured optimization. We show that the spectral update corresponds to the exact steepest-descent direction for transformers composed of linear layers with normalized inputs and outputs. Our approach is applied on top of existing unlearning methods and evaluated across complementary axes of memorization, privacy, and utility, summarized via their harmonic mean. Spectral unlearning consistently outperforms a broad range of baselines and closes most of the gap to a retain-trained reference.
Authors:
Chengjie Hong, Feixiang He, Yiheng Zeng, Lulu Kang, He WangAbstract: We propose a new method for compressing physics foundation models (PFMs) which is a new trend in AI for Science. While model compression is essential for reducing memory use and accelerating inference in large foundation models, it remains under-explored for PFMs, where preserving physical fidelity is crucial. The challenge lies in the functional nature of physics data, where partial derivatives encode spatiotemporal dynamics and exhibit high sensitivity to compression. Conventional compression methods ignore this structure, often causing severe performance degradation or failure. To address this, we introduce a sensitivity-aware fidelity-enforcing compression framework that explicitly models loss-aware layer sensitivity in the output function space during compression. This provides a new route to compressing scientific foundation models while preserving accuracy and physical fidelity. Experiments show substantial gains over existing methods across multiple models and datasets, achieving significantly higher compression ratios while maintaining accuracy, in some cases by orders of magnitude. More broadly, the work potentially leads to a new subfield of efficient, deployable, and sustainable scientific foundation models in AI for Science.
Abstract: Reasoning models improve accuracy through extended chains of thought, but their long outputs create a memory and compute bottleneck. KV cache eviction methods reduce this cost by evicting unimportant key-value pairs from the cache, yet they suffer larger accuracy degradation than selection-based sparse attention alternatives, which keep the full KV cache. We identify key factors crucial to KV cache eviction accuracy. First, a small fraction of value states have abnormally large magnitudes, and evicting them causes catastrophic failure where models enter repetitive reasoning loops. Second, introducing stochasticity during eviction improves accuracy by increasing cache diversity. Based on these findings, we propose viction (VaSE), a training-free recipe that protects large-magnitude value states and promotes diverse eviction decisions. Across six reasoning tasks, Qwen3 models using VaSE with 4x KV cache compression achieve comparable or better average accuracies to the SOTA selection method under the same level of sparsity, while enabling a static memory footprint that selection methods cannot offer. Our method also generalizes as a recipe, with its key factors improving strong eviction methods by 5 points on average.
Abstract: Human mobility differs from text and from generic time series in three structural ways: visits are tuple-valued events whose meaning depends on the joint distribution over location, time, and activity; users carry persistent signatures across trajectories; and visits are not independent across users, since co-location at shared places is a primary signal. Existing pre-training recipes for mobility import objectives from language modeling, treating trajectories as sentences and visits as tokens, an analogy that fails against each of the three properties above. These properties define a broader class, multi-entity spatiotemporal event streams (MESES), spanning enterprise authentication logs, electronic health records, and other event-stream domains where entities share infrastructure, schedules, or contexts. We make the properties precise as three axioms that any pre-training framework for MESES should satisfy, and introduce TraXion, whose objectives and architecture are jointly designed to meet them. A single TraXion checkpoint per dataset beats task-specific baselines on every task across six public mobility datasets covering anomaly detection, next-POI recommendation, next-visit prediction, and social-link prediction. The same recipe, applied unchanged to enterprise authentication logs and ICU mortality prediction, matches or exceeds prior work on both, showing that event streams from domains as different as mobility, security, and healthcare can be modeled under a single framework.
Authors:
Tianyu Wu, Yu Yao, Zhenting Qi, Han Zheng, Chengxi Zhang, Zhuohan Wang, Haoran Ma, Zichun Liao, Himabindu Lakkaraju, Ju Li, Yilun DuAbstract: Speculative decoding accelerates LLM inference by having a small drafter propose tokens that a larger target model verifies in parallel. Current SOTA diffusion-based parallel drafters such as dFlash predict the full B-token block in one forward pass, allowing deeper drafters and higher accepted length. However, across the field, from autoregressive tree drafters to parallel block drafters, training losses use fixed per-position weights that do not adapt to the drafter's current bottleneck position. We derive per-position training weights directly from the expected accepted block length, so that each position's weight reflects its actual contribution to acceptance. The resulting loss, D-PACE (Dynamic Position-Aware Cross-Entropy), automatically identifies where the block's acceptance bottleneck lies and shifts training signal there as the drafter improves. D-PACE consistently improves both wall-clock speedup and accepted length across diverse benchmarks and positions as a drop-in replacement with negligible training overhead and requires no changes to the drafter architecture or inference pipeline, making it broadly applicable to general model architectures.
Abstract: In this paper, we propose UniGS, a unified map representation and differentiable framework for high-fidelity multimodal 3D reconstruction based on 3D Gaussian Splatting. Our framework integrates a CUDA-accelerated rasterization pipeline capable of rendering photo-realistic RGB images, geometrically accurate depth maps, consistent surface normals, and semantic logits simultaneously. We redesign the rasterization to render depth via differentiable ray-ellipsoid intersection rather than using Gaussian centers, enabling effective optimization of rotation and scale attribute through closed-form analytic depth gradients. Furthermore, we derive the analytic gradient formulation for surface normal rendering, ensuring geometric consistency among the reconstructed 3D scenes. To improve computational and storage efficiency, we introduce a learnable attribute that enables differentiable pruning of Gaussians with minimal contribution during training. Quantitative and qualitative experiments demonstrate the state-of-the-art reconstruction accuracy across several datasets, validating the efficacy and robust optimization stability of our geometry-aware paradigm.
Abstract: While existing time series foundation models primarily rely on large-scale unimodal pretraining, they lack complementary modalities to enhance time series understanding. Building multimodal foundation models is a natural next step, but it faces key challenges: 1) lack of a unified multimodal pretraining paradigm and large-scale multimodal corpora for time series analysis; 2) how to effectively integrate heterogeneous modalities and enhance model generalization. To address these challenges, we take an early step toward multimodal foundation models for time series analysis. We first propose a multimodal pretraining paradigm that leverages time series with endogenous modalities (derived images and text) and exogenous knowledge (news), providing a comprehensive multi-view perspective for time series analysis. To support this, we develop an automated data construction pipeline to curate MM-TS, the first large-scale multimodal time series dataset spanning six domains, with up to one billion points. Then we propose HORAI, a frequency-enhanced multimodal foundation model. It integrates two core components: the Frequency-enhanced Cross-Modality Encoder and the Time-Frequency Decoder, designed to effectively fuse multimodal features and enhance model generalization across modalities and domains. After pretraining on MM-TS, HORAI achieves state-of-the-art zero-shot performance on time series forecasting and anomaly detection tasks, demonstrating strong generalization.
PaperID: 3163, Poster
Authors:
Jinglong Wang, Yunjie Wang, Zhiyang Zhang, Jiawei He, Ye Yuan, Bo Qiu, Jing ZhangAbstract: Embodied tasks demand accurate, flexible, and semantically rich 3D scene representations. 3D semantic occupancy is well suited to this requirement, as it can model holistic 3D spaces by encoding geometric occupancy along with semantic categories. However, existing occupancy prediction methods struggle to meet practical deployment requirements, such as adapting to varying computing budgets, sensor setups, and observation views. In this paper, we propose a point-based Adaptive 3D Occupancy Prediction method, called AdaOcc, tailored for embodied scenarios. To accommodate heterogeneous sensor inputs, AdaOcc uses an adaptive geometry-guided dual-branch encoder that can support RGB images in various numbers of views with (estimated) depth maps or LiDAR scans. AdaOcc represents occupied regions via sparse semantic points trained with a progressive query learning strategy, allowing the prediction computational budget to be flexibly adjusted through query point numbers and decoder layers. To facilitate high-fidelity geometric modeling for lightweight point-based occupancy learning, we further propose a novel containment loss that regularizes predicted points to reside within valid occupied regions. Extensive experiments show that our method achieves a new state-of-the-art on Occ-ScanNet with considerable performance improvements over previous methods. Moreover, our framework demonstrates strong practical applicability as an adaptive 3D perception module in real-world embodied systems.
Abstract: Autonomous science promises to augment scientific discovery, particularly in complex fields like biomedicine. However, this requires AI systems that can consistently generate novel and diverse solutions to open-ended problems. We evaluate LLMs on the task of open-ended solution generation and quantify their tendency to mode collapse into low-diversity generations. To mitigate this mode collapse, we introduce analogical reasoning (AR) as a new approach to solution generation. AR generates analogies to cross-domain problems based on shared relational structure, then uses those analogies to search for novel solutions. Compared to baselines, AR discovers significantly more diverse generations (improving solution diversity metrics by 90-173%), generates novel solutions over 50% of the time (compared to as little as 1.6% for baselines), and produces high-quality analogies. To validate the real-world feasibility of AR, we implement AR-generated solutions across four biomedical problems, yielding consistent quantitative gains. AR-generated approaches achieve a nearly 13-fold improvement on distributional metrics for perturbation effect prediction, outperform all baselines on AUPRC when predicting cell-cell communication, infer brain region interactions with a high Spearman correlation (\rho=0.729) to published methods, and establish state-of-the-art performance on 2 datasets for oligonucleotide property prediction. The novel and diverse solutions produced by AR can be used to augment the search space of existing solution generation methods.
PaperID: 3165, Poster
Abstract: Camera localization methods are mostly multi-stage pipelines relying on dedicated and complex maps that are costly to construct, store, and maintain over time. Alternatively, well-structured maps that organize scenes with compact yet expressive primitives are often readily available, providing rich structural and semantic cues that humans can interpret. However, their potential for direct camera localization remains largely unexplored due to the significant modality gap with visual observations. This motivates a new concept for localization, which estimates camera poses by ransformer network trained to predict multiple localization targets from a single query image and a set of planar map primitives, enabling feed-forward 6-DoF camera localization that generalizes to unseen environments. Based on this flexible and scalable architecture, our method pushes the state-of-the-art localization performance on both detailed planar surface maps and minimal architectural layouts. The code and models will be made publicly available.
PaperID: 3166, Poster
Abstract: Tuning-free customized generation has gained popularity for its efficiency, particularly with the rise of rectified flow transformers. However, it often suffers from semantic inconsistency, degrading both concept fidelity and prompt alignment. We attribute this failure to their tendency for inference-time superficial feature fusion, which is intrinsic to the tuning-free paradigm. In this paper, we propose SemanticFlow for semantically consistent tuning-free customization, introducing explicit semantic control to jointly enhance concept fidelity and prompt alignment. SemanticFlow operates through a cohesive "perceive-and-manipulate" mechanism for aligned fusion of visual and textual features at the early denoising stage. Specifically, we first devise lightweight semantic token streams as run-time probes to accurately perceive the spatial responses of the target semantics. We then formulate semantic manipulation as an energy-guided rectification of these semantic responses within the rectified flow framework, steering the generation process toward the desired outcome. Experiments demonstrate that SemanticFlow jointly enhances concept fidelity and prompt alignment in previously challenging scenarios, effectively enabling semantically-consistent customization.
PaperID: 3167, Poster
Abstract: Resolution invariance is a critical capability for scientific machine learning methods (SciML), however, the ability to sample at any resolution does not guarantee stable emulation at all resolutions. This important distinction underpins one of scientific machine learning's greatest potential: the ability to efficiently emulate complicated and computationally expensive systems on inexpensive coarse grids. Stable numerical solvers are computationally expensive as they attempt to resolve small scales that receive and dissipate spectral energy. Learned coarse-grid forecasters remove the dissipative channel, allowing unresolved frequencies to re-emerge as grid-scale noise (aliasing). We explore stable resolution invariant emulation through physically principled control over entropy dynamics to reinforce the dissipative mechanism within the Neural Discrete Equilibrium (NeurDE) method. Entropically-stabilized NeurDE produces stable long-time rollouts for thousands of forecasting steps on large Reynolds number turbulent flows (Re=50,000) with very coarse grids (\mathbbR^16×16). Entropically-stabilized NeurDE demonstrates extreme stability at coarse resolutions, outperforming the data-generating stabilized numerical method: a much more challenging baseline than SciML surrogates.
Abstract: In this work, we develop theoretical foundation for flow matching with neural-network–parameterized conditional velocity fields. We establish convergence guarantees for gradient descent in the over-parameterized 2-layered ReLU neural network regime. We derive generalization bounds for the conditional velocity-field matching objective. Building on these results, we provide Wasserstein-distance guarantees for the samples generated by the induced flow. Our analysis is based on generalization bound for multi-task representation learning with unbounded losses, which may be of independent interest beyond flow-based generative modeling. These theoretical results are validated through extensive experiments on both synthetic and real-world image benchmarks.
PaperID: 3169, Poster
Abstract: Linear attention offers an efficient alternative to full attention with a fixed-size recurrent state. However, this state is shared by all tokens, so information from distinct tokens becomes superposed within it and produces inter-token interference that degrades long-range fine-grained recall. To address this issue, we propose RAM-Net, which replaces dense access to a shared state with sparse address-based access. RAM-Net organizes the recurrent state as a fixed-size array of independent slots and uses an Address Decoder that maps each key or query into a sparse address, selecting a small subset of slots to write to or read from at each step. This design directs tokens with non-overlapping addresses to disjoint slots, suppressing inter-token interference, while keeping per-step overhead dependent only on the number of accessed slots rather than the total state size. Empirically, RAM-Net shows a clear advantage over strong linear baselines on fine-grained long-range retrieval and remains competitive on standard language modeling and commonsense reasoning, while accessing far fewer state elements per step than these baselines (e.g., 32× fewer than Mamba2).
PaperID: 3170, Poster
Abstract: Recovering an object's 3D structure and camera motion from a few uncalibrated views is a fundamental problem in computer vision. Classical structure from motion methods recover 3D geometry consistent with the input images but cannot recover unseen regions; native generation methods produce plausible 3D objects which are related but not necessarily faithful to the input images. Methods simply concanetating reconstruction and generation allow error to cascade across the two stages, instead of producing a 3D object consistent with both the distribution and the input. We propose GenSfM, which integrates image-based 3D reconstruction with a native 3D diffusion model. In particular, we use a multi-modal diffusion transformer to jointly denoise the shape latent and per-view pose latents under shared self-attention, grounding generated geometry in the observed views while using object-level priors to infer unobserved regions. A compact 3D-native UV-volume tokenization keeps the joint sequence small as views are added, and a pose-aware stage refines geometry and texture by sampling per-voxel image features at the recovered camera poses. On Toys4k and GSO, for 4 uncalibrated views, GenSfM improves novel-view PSNR by 3.2--4.5~dB and reduces median Chamfer distance by 2.8--3.8× over the strongest 3D-generation baseline, while recovering input-view poses over 96% Acc@5^\circ without any external pose-estimation backbone, post-hoc optimization, or known intrinsics.
PaperID: 3171, Poster
Abstract: Decentralized stochastic minimax optimization has recently attracted significant attention due to its applications in machine learning. However, existing state-of-the-art methods use learning rates of different scales for the primal and dual variables, making them difficult to tune in practice. To address this problem, this paper proposes a novel doubly smoothed decentralized stochastic minimax algorithm. Specifically, in terms of algorithm design, we update both the primal and dual variables using smoothed gradients and introduce novel approaches to handle the computation and communication of the auxiliary variables introduced by the smoothing technique. On the theoretical side, for nonconvex-PL problems, our convergence analysis reveals that the learning rates for the primal and dual variables are of the same scale. Moreover, the order of the condition number in our convergence rate is improved to O(\kappa^3/2). To the best of our knowledge, this is the first time it has been improved to such a favorable order. Finally, extensive experimental results validate the effectiveness of our algorithm.
PaperID: 3172, Poster
Abstract: Triangle-based splatting has emerged as a promising bridge between high-quality novel view synthesis and traditional graphics pipelines. However, naively applying unstructured splatting mechanisms to polygonal primitives fails to exploit their inherent topological advantages, leading to over-tessellation and degraded high-frequency appearances. To address, we present T^2-Splat, a mesh splatting framework with adaptive topology and texture optimization. We introduce progressive edge-collapse and vertex-split operations to adaptively allocate triangle primitives, reducing face count while preserving well-conditioned connectivity. We further propose a scale-aware residual Mip-texture model to decouple appearance from geometry, for providing the topology with flexibility needed to safely collapse redundant triangles without introducing blurring or aliasing. Experiments on the Mip-NeRF 360, Tanks \& Temples, and DTU benchmarks demonstrate that T^2-Splat improves PSNR by +0.87 dB and maintains robust visual fidelity even under aggressive mesh compression to 10% of the original face count.
Authors: Zituo Chen, Sili Deng
Abstract: Reliable physics simulation requires both broad generalization across heterogeneous PDE families and stability under long autoregressive rollouts. Current neural solvers rarely provide both. Deterministic operators accumulate rollout error, while probabilistic solvers typically remain tied to a single PDE family or a short prediction horizon. We address this gap with the Latent Generative Solver (LGS), which couples a Physics VAE (PhyVAE) with a Pyramidal Flow-Forcing Transformer (PFlowFT). PhyVAE compresses states from twelve PDE families into a shared \emphphysical manifold, separating dynamics-relevant structure from the ambient state space. PFlowFT then predicts the next latent state through input-noised flow matching, yielding a sufficient-condition contraction mechanism that explains improved autoregressive rollout stability. Pretrained on a 2.5\,M-trajectory, 16-system corpus at 128^2 resolution, LGS obtains the lowest relative L_2 error (L2RE) on 15 of 16 systems at both 5- and 10-step rollout. At 20 steps, LGS reduces the L2RE from 56.1% to \mathbf30.2% against matched deterministic and adapted generative baselines, while requiring \mathbf13--\mathbf77× less recurrent dynamics-step compute. On a held-out 256^2 Kolmogorov flow, LGS reduces the 1-step L2RE from 0.398 to 0.129 within five finetuning epochs; U-AFNO achieves only 0.653 to 0.343 under the same protocol.
PaperID: 3174, Poster
Authors: Daniel M Smola, Alexander Smola
Abstract: Fairness impossibility results often look like distinct scalar incompatibility statements. We show that several share one RKHS geometry: fairness criteria are linear constraints on conditional mean embeddings, and unequal base rates make the law of total expectation overdetermine those constraints. This view yields four results. The Kleinberg--Mullainathan--Raghavan dichotomy needs only group-conditional unbiasedness, not full calibration. The \emphPok\'emon theorem shows that a distinct group pair satisfying any finite collection of linear mean-fairness criteria leaves a residual violation witnessed by the MMD, decaying at the Kolmogorov m-width rate under spectral regularity. The same tools prove an impossibility for fair feature learning: parity and class-conditional separation in representation space force class collapse under unequal base rates. The approximate relaxations yield signal and error frontiers, allowing a trade-off between real-world estimators and fairness goals. Experiments on standard fairness benchmarks are consistent with our bounds.
PaperID: 3175, Poster
Abstract: Score-based variational inference (VI) provides an alternative to Kullback--Leibler (KL)-based VI by minimizing the Fisher divergence between the variational distribution and the target. Existing eigenvalue-based score-VI methods face two high-dimensional obstacles: exponential parameter growth and instability under degenerate or nearly degenerate low-energy eigenspaces. We propose QuanVI, a scalable quantum-inspired algorithm that combines a mixed-state density-operator formulation with an MPO-motivated quantum tensor network (QTN) parameterization. The former represents degenerate low-energy eigenspaces by their maximally mixed state, avoiding arbitrary eigenvector selection; the latter realizes a compact density operator without explicitly constructing exponentially large matrices. Experiments and ablations show that QuanVI preserves low-dimensional accuracy while scaling to high-dimensional synthetic and Bayesian posterior-approximation benchmarks, including challenging non-Gaussian targets.
PaperID: 3176, Poster
Abstract: Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations. Existing IRL methods typically rely on a bi-level optimization procedure that alternates between reward learning and policy optimization, leading to substantial computational burden and training instability. In this work, we introduce a different route that eliminates policy optimization entirely by leveraging diffusion policies. Our key insight is that a diffusion policy encodes the action-gradient structure of the optimal soft Q function, enabling reward learning to be cast as a sequence of value recovery problems, thereby allowing us to bypass reward-policy loops inherent in prior IRL methods. Specifically, our method proceeds in three stages: (I) recovering the optimal soft Q function via action-gradient matching and estimating the corresponding soft value function (LogSumExp of Q values) in a way inspired by Gumbel regression; (II) calibrating these soft values by inferring a state-dependent offset; (III) extracting the reward by enforcing Bellman consistency. This leads to (LFIRL), a fully offline algorithm that operates in a simple, loop-free, and sequential manner. LFIRL is simple to implement, significantly improves training efficiency while maintaining strong reward recovery performance. Empirically, across Maze, Franka Kitchen, Adroit Hand Pen, and Push-T benchmarks, LFIRL achieves speedup over the fastest baselines, while matching or surpassing state-of-the-art methods in reward recovery quality.
Authors:
Ge Wang, Yibo Peng, Fan Feng, Shenhao Yan, Chengsi Yao, Jiahao Yang, Honghao Cai, Yiming Zhao, Xi Li, Jinke Ren, Shuguang Cui, Yatong Han, Zhen LiAbstract: We present Reflow-regularized Flow Matching Policy Gradients (ReFPO), a simple online RL method that adds explicit Reflow regularization to FPO for efficient flow-based control. We uncover a key structural property: the gradient updates in Flow Matching Policy Gradients (FPO) can be interpreted as an implicit advantage-weighted Reflow process, providing a new geometric perspective on flow-based policy gradients. Building on this insight, ReFPO introduces an explicit geometric regularizer that can be implemented with a single line of code change without incurring additional computational overhead or auxiliary distillation stages. By synergizing advantage-guided updates with path rectification, our method reduces CFM proxy-ratio spikes, stabilizes PPO-style training, and enables high-fidelity one-step inference that often matches or exceeds multi-step performance. We experimentally demonstrate that ReFPO improves average performance and discretization robustness across GridWorld, MuJoCo Playground, and high-dimensional Humanoid Control tasks, providing a scalable and stable approach for generative policies in complex physical simulations.
PaperID: 3178, Poster
Abstract: Despite the significant progress in image enhancement under adverse illuminations, such as low-light image enhancement, exposure correction, and backlit image enhancement, existing models show limited generalizability to different scenarios. A major bottleneck lies in the lack of large-scale and diverse training data, as these tasks typically rely on manually captured, precisely aligned image pairs that are expensive to scale up. Although synthetic data have been explored as an alternative, existing synthesis methods are limited in modeling complex adverse illumination conditions and camera imaging pipelines. To address this issue, we propose a novel framework to synthesize realistic training pairs for generalizable image enhancement in the wild. Specifically, we first derive multi-exposure sequences (MES) from high-bit RAW images by rendering them under different exposure settings through an emulated ISP pipeline. Using these RAW-derived MES as supervision, we train a diffusion-based RGB-to-MES generator to synthesize plausible exposure sequences from a single 8-bit RGB image, extending the synthesis pipeline from limited RAW collections to large-scale in-the-wild RGB data. For each source image, we randomly sample one synthesized exposure variant as the degraded input and use the original high-quality RGB image as the target. With the proposed MES Synthesis (MESS) approach, we construct a dataset of 500K training pairs, and train a lightweight network, which however demonstrates significantly better generalization performance than existing models across various image enhancement tasks. Code, dataset, and models will be released.
PaperID: 3179, Poster
Authors: Anany P Gupta, Ye Zhou, Qi Li
Abstract: Physical control systems, such as building HVAC systems, thermal processes, and power-grid assets, often involve delayed dynamics, safety constraints, and costly online interaction. Because the effect of a control action may appear after several time steps, effective control often requires reasoning over histories of states, actions, and outcomes. Pretrained sequence controllers, such as Decision Transformers, provide a natural way to model this history-dependent control problem by learning policies from offline trajectories, reducing the need for expensive online retraining. However, their performance can degrade when deployment conditions differ from the training data. In building control, for example, actuator wear, sensor drift, occupancy changes, and weather variation can all affect the observed response of the system to the same control action. Online adaptation is therefore needed, but raw prediction error is not always a reliable signal for updating the controller. The key challenge is to identify which part of the observed mismatch is action-sensitive and should guide policy adaptation. To address this challenge, we propose DARPAN, a Drift-Aware Residual Policy Adapter Network for pretrained physical control. DARPAN keeps the pretrained Decision Transformer frozen and performs online adaptation only through lightweight residual adapters, preserving the offline policy and limiting the number of trainable parameters. Instead of updating from raw prediction error, DARPAN uses a controllability filter trained in simulation to extract the action-sensitive part of the residual and filter out mismatch that is weakly related to the controller's actions. During deployment, a frozen predictor estimates the next state, the controllability filter produces a filtered residual, and a short history of filtered residuals forms a drift signature. This signature is used to select an existing adapter or create a new one when a persistent drift pattern appears. We benchmark DARPAN on SustainGym under actuator faults, sensor drift, delays, and external disturbances. DARPAN reduces the drift-induced performance loss by 31.7% and uses 63% fewer online parameters than Online DT. We release anonymized code and experiment configurations at https://huggingface.co/Darpan2026.
PaperID: 3180, Poster
Authors:
Yue Ma, Pengjie Song, Xinyu Wang, Yi He, Zeqian Long, Fangneng Zhan, Kaichen Zhou, Hongyu Liu, Hongfa Wang, Peihao Li, haoyang huang, Nan Duan, Qifeng ChenAbstract: Egocentric action transfer aims to decouple the semantic action from egocentric observations and reproduce them in new visual contexts, enabling scalable world simulation, embodied policy learning, and robotic data generation. Existing approaches to transferring actions rely on condition-guided video generation, which converts the observation video into explicit geometric controls ( , hand poses or meshes) to drive synthesis. However, these methods introduce geometric estimation errors, making it difficult to preserve physically plausible interactions when transferred to new scenes. Alternatively, motion transfer methods directly extract motion patterns from the observation video, yet motion-level features alone cannot encode rich contact dynamics, frequently leading to severe hand structural collapse and implausible contact layout. To address both limitations, we present EgoACT, a test-time framework for egocentric action transfer that operates directly on the denoising process of video diffusion models without requiring additional geometric estimators. EgoACT uses Velocity-guided Structure Anchoring to stabilize reference-consistent hand-object structure in the early denoising stage, and Sparse Correspondence Calibration to refine reliable local correspondences in the mid-to-late denoising stages. Together, these two components preserve transferable action semantics while improving temporal coherence and interaction realism. We further establish EgoActionBench, a benchmark for evaluating action preservation, visual quality, and hand-object plausibility across diverse egocentric manipulation scenarios. Experiments show that EgoACT generates more coherent and physically plausible interaction videos than strong baselines.
PaperID: 3181, Poster
Abstract: We study the problem of Pseudo-Benign Failures in Vision--Language Models (VLMs): multimodal inputs that appear harmless but elicit dangerous or policy-violating responses. Our analysis shows that these failures arise from a representational misalignment: the model's internal embedding space exhibits a distributional gap between pseudo-benign inputs and unsafe inputs located in the refusal region, causing failures outside the safety margins of models. We introduce Representation-Level Safety Margin Alignment method (ReSAM), a lightweight representation-space alignment method that: (i) computes direction vectors separating refusal and non-refusal representations, (ii) quantifies refusal behavior by projecting embeddings of inputs onto this direction, and (iii) optimizes a safety-margin loss that pushes unsafe and pseudo-benign queries above a learned margin while pulling safe queries below it. ReSAM introduces a new paradigm for multimodal safety alignment: it requires no manual annotations, instead deriving supervisory signals directly from its own representation space. Despite this minimal supervision, ReSAM achieves a 68% improvement in safety over strong baselines, and remarkably, we further observe that incorporating only a handful of pseudo-benign queries (as few as five) during training suffices to raise safety to 94.6%. Beyond these empirical gains, our analysis reveals that safety gradients concentrate in a low-rank subspace, suggesting that multimodal safety is governed by an intrinsic structure that can be systematically identified and controlled.
Authors:
Yulong Li, Jianxu Chen, Xiwei Liu, Chuanyue Suo, Rong Xia, Yichen Li, Xinlin Zhuang, Niranjana A Menon, Jionglong Su, Eran Segal, Imran RazzakAbstract: Medical foundation models generate narrative explanations but cannot quantify intervention effects, detect evidence conflicts, or validate literature claims, limiting clinical auditability. We propose causal compilation, a paradigm that transforms medical evidence from narrative text into executable code. The paradigm standardizes heterogeneous research evidence into structured estimand objects, each explicitly specifying intervention contrast, effect scale, time horizon, and target population, supporting six executable causal queries: do-calculus, counterfactual reasoning, temporal trajectories, heterogeneous effects, mechanistic decomposition, and joint interventions. We instantiate this paradigm in DoAtlas-1, compiling 1,445 effect kernels from 754 studies through effect standardization, conflict-aware graph construction, and real-world validation (Human Phenotype Project, 10,000 participants). The system achieves 98.5% canonicalization accuracy, 100% conflict detection recall, and 80.5% query executability. This paradigm shifts medical AI from text generation to executable, auditable, and verifiable causal reasoning.
PaperID: 3183, Poster
Abstract: Series decomposition aims to separate long-term trajectories and periodic patterns in time series, thereby improving the specificity and interpretability of forecasting models. However, mainstream approaches usually equate this functional separation with an operational split between a smoothed baseline and fluctuating residuals. One line of methods approximates the trend via low-pass smoothing, which assumes long-term trajectories to be purely smooth and may misassign nonlinear evolution, stage-wise transitions, and structural fluctuations to the residual. The other relies on heuristic nonlinear decompositions to capture complex dynamics, but is sensitive to noise and hyperparameters, leading to unstable and mixed component boundaries. To address this issue, we propose a forecasting framework, QWaveNet. It employs a quantum convolutional neural network to learn Dynamic Evolution, a long-term structural representation beyond low-pass smoothing priors, which preserves nonlinear evolution, stage-wise transitions, and structural fluctuations more coherently, while its complementary Evolutionary Residual captures localized variations and periodic dynamics. Guided by these complementary components, QWaveNet further performs adaptive multi-scale decomposition and reconstruction to support forecasting. Experiments on multiple benchmark datasets demonstrate the effectiveness of QWaveNet.
Abstract: We propose a closed-form spectral framework for relative log-density estimation in linearly parameterized probabilistic models, including unnormalized and conditional models. This is achieved by representing the Kullback-Leibler (KL) divergence as an integral of weighted chi-squared divergences, converting KL estimation into a family of least-squares problems. We derive an explicit spectral formula based only on first- and second-order feature moments, yielding closed-form estimators of both divergences and log-density potentials for fixed features. The framework extends to a broad class of f-divergences and can be combined with kernelization or feature learning with neural networks. We prove convergence guarantees for the resulting estimators and empirically compare them on synthetic data with optimization-based variational formulations, including logistic and softmax regression for normalized conditional models.
PaperID: 3185, Poster
Authors: Omar Kassi, Bernhard Stankewitz, Botond Szabo
Abstract: Gaussian process regression models are widely used in modern statistics and machine learning due to their flexibility, interpretability, built-in uncertainty quantification, and strong theoretical foundations. However, training Gaussian processes (GPs) has a computational cost of O(n^3), while the memory and prediction costs are also O(n^2). For the large datasets commonly encountered in practice, this becomes infeasible and necessitates the use of approximate posterior methods. In this work, we propose a class of approximations based on sparse covariance matrices, where the kernel matrix is sparsified according to the spatial proximity of the design points. We analyze the accuracy of the resulting approximate posterior distribution and show that the proposed approach preserves optimal statistical inference properties, provided that the dependencies among a sufficient number of neighboring points for each feature are retained. We derive explicit sufficient and necessary conditions on the number of neighbors required in a general setting and apply these results to standard kernels, including the squared exponential and Mat\'ern kernels. Our theoretical findings are supported by numerical experiments.
PaperID: 3186, Poster
Authors: Simon Bienewald, Lukas Trottner
Abstract: Denoising diffusion models have evolved into a state-of-the-art method for tasks in various fields, such as denoising and generation of images, text generation, or generation of synthetic data for training of other machine learning models. First hitting diffusion models (FHDM) are a particular class of denoising diffusion models with random adaptive generation time tailored to generate data on a known manifold. Building on the conditioning framework of Doob's h-transform these models leverage the given information on the target data manifold to demonstrate strong performance across tasks while offering distinct features such as time-homogeneous dynamics of the generating process and a reduced average simulation time. Even though the theoretical investigation of standard forward-backward diffusion models has attracted much attention in the recent past, the statistical convergence properties of FHDMs are not yet understood. In this work, we show that, up to logarithmic factors, FHDMs achieve the minimax optimal convergence rate in total variation for spherically supported Sobolev smooth data distributions. In particular, this is the first statistical optimality result for denoising diffusion modelling with random generation time.
Abstract: Flow Matching has enabled robust text-to-video generation via latent ODE sampling. However, velocity approximation and numerical discretization errors inevitably accumulate, causing sampling trajectories to drift. Consequently, generated videos often suffer from severe spatiotemporal inconsistencies, such as physically implausible motions. Yet, directly correcting these drifted noisy latents is challenging: timestep-dependent noise obscures reliable structural cues, and spatial interventions risk disrupting fragile local geometry while incurring heavy computational costs. To address this, we propose Spectral Lookahead Rectification (SpecLoR), a plug-and-play inference method that bypasses noise via lookahead prediction. It also circumvents spatiotemporal entanglement by shifting corrections to the frequency domain, where universal statistical natural-video priors are readily available. First, during early sampling stages, SpecLoR looks ahead to estimate the clean latent z_t,0 and computes its 3D spatiotemporal spectrum. Next, SpecLoR rectifies only the amplitude to match the statistical prior, leaving the phase intact. Finally, the corrected state is re-noised to resume ODE integration. Experiments on Wan2.2 demonstrate that SpecLoR significantly reduces physical artifacts and enhances motion coherence across multiple benchmarks with minimal computational overhead (4 additional NFEs in a 40-step schedule). Code will be released.
PaperID: 3188, Poster
Abstract: We study revenue-optimal posted-price design for GPU markets under a finite capacity B. We focus on two settings: \emphsales, where each purchase permanently consumes a vendor's GPU inventory, and \emphrentals, where each job occupies GPUs for a finite duration and then releases them. In both settings, the vendor must commit to an anonymous price menu before observing customers' valuation, GPU demand, and job duration. First, in the sales setting, we show that computing the optimal anonymous static menu is NP-hard, and that the common practice of linear pricing (per-GPU price) can be highly suboptimal. We develop three methods with varying approximation guarantees, all of which outperform linear pricing. Our main method is a convex program (CP) obtained via an ex-ante relaxation of the problem, which obtains a \tfrac1-\epsilon2 approximation to the optimum when the maximum GPU demand \overlineN satisfies \overlineN\leq \epsilon B. Building on these results, we analyze the \emphrental model. We again show that linear pricing (per-GPU-hour price) is highly suboptimal, and design duration-aware posted-price policies based on slot-coupled ex-ante CP. Our \emphThrottled CP-Menu Posted Pricing (TCMPP) policy posts the CP-induced menu randomly and achieves a \frac1-\epsilon4 approximation. We further introduce a slackened version of the same convex program, which reserves capacity in each time slot and achieves stronger guarantees in the small job regime. We corroborate our theoretical results with simulations.
PaperID: 3189, Poster
Abstract: Privacy is a central concern when fine-tuning large language models (LLMs) on sensitive data, and differentially private stochastic gradient descent (DP-SGD)---which clips per-sample gradients and adds calibrated Gaussian noise---is the standard tool for formal privacy guarantees. Both theory and practice show that lower-rank models are better suited to DP training, a property especially relevant for LLMs, whose fine-tuning gradients exhibit a strong low-rank structure. Methods such as DP-LoRA exploit this by restricting updates to a low-rank subspace, i.e., retaining only a few non-zero components in the SVD of each layer's gradient. However, we argue that while having few non-zero components is important, the isotropic noise injected by DP-SGD inflates the singular values of the gradient matrix, disrupting their naturally fast decay. In this work, we investigate whether this noise-induced eigenvalue blow-up reduces performance, and show that partially restoring the original singular-value profile significantly improves the sample efficiency of DP-SGD. Experiments on language classification (GLUE benchmark with RoBERTa) and text generation (E2E and DART table-to-text benchmarks with Qwen and Llama models up to 4B parameters) showcase that restoring the fast decay of singular values is a viable strategy for speeding up the DP optimization process, without compromising privacy guarantees.
PaperID: 3190, Poster
Abstract: Understanding how complex molecular systems transition between metastable states is a central challenge in computational chemistry and biophysics. A key tool for characterizing these transitions is the committor function, defined as the probability that a trajectory initiated from a given configuration reaches one state before another, which encodes detailed mechanistic information about transition pathways. First, we curate a public benchmark of empirically determined committor values for alanine dipeptide. Second, we develop a "committor-in-the-loop" framework that combines a self consistent loss function with weighted ensemble (WE) simulation: WE adaptive bins are biased by the learned committor, which is in turn trained on WE samples. Across a two-channel 2D system and alanine dipeptide settings, our WE-based framework consistently outperforms greedy baselines, demonstrating the importance of principled exploration. Together, the benchmark and framework provide both a standard for evaluating committor learning methods and a new method for studying rare transitions in high-dimensional molecular systems.
Abstract: Multimodal latent reasoning has emerged as a promising paradigm that replaces explicit Chain-of-Thought (CoT) decoding with implicit feature propagation, simultaneously enhancing representation informativeness and reducing inference latency. By analyzing token-level gradient dynamics during latent training, we reveal two critical observations: (1) visual tokens exhibit significantly smaller gradient norms than their textual counterparts due to inherent language bias, resulting in systematic visual under-optimization; and (2) semantically simple tokens converge rapidly, whereas complex tokens exhibit persistent gradient instability constrained by fixed architectural depths. To address these limitations, we propose a visual replay module and routing depth scaling to collaboratively enhance visual perception and refine complicated latents for deeper contextual reasoning. The former module leverages causal self-attention to estimate token saliency, reinforcing fine-grained grounding through spatially-coherent constraints. Complementarily, the latter mechanism adaptively allocates additional reasoning steps to complex tokens, enabling deeper contextual refinement. Guided by a curriculum strategy that progressively internalizes explicit CoT into compact latent representations, our framework achieves state-of-the-art performance across diverse benchmarks while delivering substantial inference speedups over explicit CoT baselines.
PaperID: 3192, Poster
Abstract: Extracting sentence embeddings from MoE-based large language models is a promising direction, as they provide greater model capacity than dense models at comparable computational cost. Existing works perform the standard forward propagation process to extract embeddings, overlooking a key stability requirement in the MoE encoding process: semantically similar texts should be encoded into similar representations. In this work, we identify two key stability-related phenomena in MoE models: (1) routing stability varies across layers, and (2) shared and routed experts exhibit different levels of stability. To address these issues, we propose SAME, a Stability-Aware MoE Embedding extraction framework that dynamically allocates activated experts across layers and adjusts expert outputs to mitigate component instability and improve embedding quality. Specifically, SAME assigns fewer activated experts to layers with less stable routing, where layer-wise routing stability is measured by the average overlap between the experts activated before and after injecting Gaussian noise into each layer’s router inputs. Additionally, SAME reduces the contributions of routed experts to alleviate the impact of their instability. Notably, our method is training-free, seamlessly integrates with existing approaches, and incurs no additional inference overhead. Experiments on semantic textual similarity benchmarks demonstrate that SAME consistently improves the quality of extracted embeddings across multiple MoE backbones.
PaperID: 3193, Poster
Abstract: Long-term memory remains a fundamental bottleneck for large language model (LLM) agents across sessions. Existing systems typically adopt a write-then-retrieve paradigm, storing context-dependent dialogue fragments and retrieving them via nearest-neighbor search in continuous embedding space. This design yields unstable access units and systematically misses logically related but lexically distant evidence. We propose MemCode, a long-term agent memory framework that couples retrieval-oriented writing with discrete associative retrieval. MemCode rewrites dialogue into self-contained atomic memory units (AMUs), maps AMUs into a learnable codebook with isomorphic semantic quantization (ISQ), and uses codebook co-occurrence topology as a retrieval path complementary to vector search. For long-term memory stores, MemCode separates online atomic writes from asynchronous offline consolidation. On LongMemEval-S and LoCoMo with GPT-4o-mini and Qwen3-30B-A3B-Instruct-2507 backbones, MemCode achieves the best overall accuracy while keeping per-dialogue memory construction within roughly 40--130 seconds, making it 3--21× faster than LightMem and over an order of magnitude faster than A-MEM and MemoryOS.
PaperID: 3194, Poster
Authors: Sofia Blyufer, Amir Kleiner, Uri Peskin
Abstract: Nanodevices that absorb energy from their thermal environment can act as quantum refrigerators. Finding suitable configurations, however, is a needle-in-a-haystack problem: the many-body state space scales exponentially with the size of the device, where each candidate configuration requires an expensive quantum-chemical computation. Under random sampling of material parameters, only 2-9% of configurations exhibit the desired property. We introduce a search framework that couples a Graph Variational Autoencoder (GVAE) with an uncertainty-driven Active Learning (AL) loop to efficiently identify energy-absorbing configurations in electronic junctions with a 2D "working material" to be optimized. The GVAE maps physical parameters - site energies, coupling matrices, temperature - into a smooth latent space that distinguishes cooling from dissipative parameter regimes, achieving a classification ROC-AUC of 0.89 on held-out topologies and 0.80 on unseen extrapolation topologies. An AL pipeline based on Monte Carlo dropout then targets the decision boundary, iteratively selecting the candidates where the model is least certain for simulation. Over 32 AL cycles, the uncertainty-driven strategy discovers positive-flux (energy absorbing) systems at rates consistently exceeding the random baseline across all tested topologies, providing a data-efficient pathway for navigating a combinatorically large design space. The learned latent space organizes configurations along physically interpretable axes, that correlate with underlying microscopic model parameters. Our framework, thus, paves a new way towards effective cooling of computational devices, to reduce the environmental footprint of high performance computing.
PaperID: 3195, Poster
Authors: Gabryel Mason-Williams, Israel Mason-Williams, Fredrik Dahlqvist
Abstract: Data-free methods for analysing and understanding the layers of neural networks offer many metrics for quantifying notions of weak' layers, with the promise of increased interpretability. In particular, random matrix theory (RMT) offers data-free metrics that claim predictive power over the quality of pre-trained models at a layerwise precision and the ability to identify model pathologies, indicating a unique relationship with generalisation. As a result, metrics from RMT are championed for pre-training optimisation strategies and post-training compression. We establish that RMT-based metrics are unrelated to performance or training by showing that functionally indistinguishable reparametrisations of a pre-trained model can have arbitrary metrics. We show this across a range of architectures and scales. To create functionally indistinguishable reparametrisations, we exploit the well-established phenomenon of criticality: some layers can be re-initialised or re-randomised without affecting the functional behaviour of the model -- they are called robust -- while others cannot -- they are called critical. Re-initialising or re-randomising robust layers provides functionally indistinguishable reparametrisations for which RMT-based metrics are arbitrary. Moreover, we show that relationships between metrics offered by RMT are spuriously related to generalisation. We conclude by showing that if many pre-trained models have data-free metrics in a `good' range, it is, in part, dependent on model initialisation.
Abstract: Diffusion-based vision-language-action models (dVLAs) are promising for embodied intelligence but are fundamentally limited in real-time deployment by the high latency of full inference. We propose Realtime-VLA FLASH, a speculative inference framework that eliminates most full inference calls during replanning by introducing a lightweight draft model with parallel verification via the main model's Action Expert and a phase-aware fallback mechanism that reverts to the full inference pipeline when needed. This design enables low-latency, high-frequency replanning without sacrificing reliability. Experiments show that on LIBERO, FLASH largely preserves task performance by replacing many 58.0\,ms full-inference rounds with speculative rounds as fast as 7.8\,ms, lowering task-level average inference latency to 19.1\,ms (3.04× speedup). We additionally demonstrate effectiveness on real-world conveyor-belt sorting, highlighting its practical impact for latency-critical embodied tasks.
Abstract: The increasing memory demand of the Key-Value (KV) cache poses a significant bottleneck for Large Language Models (LLMs) in long-context applications. Existing low-rank KV compression methods reduce this footprint by modifying model projections, limiting the flexibility to switch back to standard full-cache inference when sufficient memory is available. In this paper, we propose EchoKV, a flexible KV cache compression framework that supports on-demand transitions from full KV caching to compressed caching. Unlike traditional compression-decompression paradigms, EchoKV utilizes a lightweight network to reconstruct the discarded KV components from a partial subset, exploiting intrinsic inter-layer and intra-layer similarities among attention heads. We further introduce a lightweight two-stage fine-tuning strategy, requiring only a few minutes on a single A100 GPU for a 7B model. Experimental results on LongBench and RULER demonstrate that EchoKV consistently outperforms existing methods across multiple compression ratios and backbone models while preserving the throughput of full-cache inference in short-context scenarios.
Abstract: Speculative decoding accelerates LLM inference by drafting future tokens with a small model, but drafter models degrade sharply under template perturbation and long-context inputs. We identify a previously-unreported phenomenon we call attention drift: as the drafter generates successive tokens within a speculation chain, attention progressively moves from the prompt onto its own recently-generated tokens. We observe this across both \emphEAGLE3 drafters and \emphMTP heads, suggesting drift is a property of drafter designs. We trace this to the un-normalized residual path between chain steps: the drafter's hidden state magnitude grows monotonically with chain depth, which exhibits dynamics consistent with additional pre-norm transformer layers stacked on the target rather than as a standalone autoregressive predictor. In order to limit the growth, we propose two architectural changes: Post-norm on the drafter hidden states and per-hidden-state RMSNorm after capturing target hidden states. Our interventions improve acceptance length over the current leading model, pre-norm EAGLE3, by up to 2× under template perturbation, 1.18× on long-context tasks, and 1.10× on seven standard benchmarks spanning multi-turn chat, math, and coding. Our changes also allow shorter train-time-test depths to generalize over longer drafting sequences.
PaperID: 3199, Poster
Authors: Thomas Chen, Zhiyuan Li
Abstract: Self-play, a type of training algorithm that enables a model to self-improve, has recently shown promising empirical results in the context of formal theorem proving using Large Language Models (LLMs). (Dong & Ma, 2025) instantiate self-play with two cooperating agents: a prover, which proves theorems, and a conjecturer, which generates new theorems as a curriculum to the prover. In this paper, we provide a theoretical framework for understanding the self-improvement capabilities of self-play algorithms for theorem proving. First, we formalize the set of theorems as a graph, with nodes as theorems and edges between pairs of theorems with similar semantics. We introduce a set of primitive assumptions that characterize the guarantees of a trained prover and how a conjecturer can access the structure of the graph. Second, we show that if the underlying graph of theorems is well-connected, then a prover-conjecturer system, where the conjecturing algorithm is based on a reversible random walk, is sufficient to grow the set of proved theorems exponentially. Third, we formulate an issue of diversity encountered empirically by self-play algorithms in our framework, where the conjecturer tends to generate artificially complex and non-fundamental theorems. We propose a diversity measure for a training distribution of theorems generated by a conjecturer and an improved conjecturing algorithm that locally maximizes this diversity measure, by computing the diffusion similarity between neighboring theorems in the theorem graph. Finally, we describe a method to compute the diffusion similarity by first using contrastive learning to embed nodes into Euclidean space and second computing the inner-product between the embeddings.
Authors:
Karen Hambardzumyan, Nicolas Baldwin, Edan Toledo, RISHI HAZRA, Michael Kuchnik, Bassel Al Omari, Thomas Foster, Anton Protopopov, Jean-Christophe Gagnon-Audet, Ishita Mediratta, Kelvin Niu, Michael Shvartsman, Alisia Lupidi, Alexis Audran-Reiss, Parth Pathak, Tatiana Shavrina, Despoina Magka, Hela Momand, Derek Dunfield, Nicola Cancedda, Pontus Lars Erik Saito Stenetorp, Carole-Jean Wu, Jakob Foerster, Yoram Bachrach, Martin JosifoskiAbstract: Existing research has identified three structural performance bottlenecks in AI research agents: (1) synchronous single-GPU execution constrains sample throughput, limiting the benefit of search; (2) a generalization gap where validation-based selection causes overfitting and performance to degrade over extended search horizons; and (3) the limited capability of fixed, single-turn LLM operators imposes a ceiling on search performance. We introduce AIRA_2, which addresses these bottlenecks through three architectural choices: an asynchronous multi-GPU worker pool that increases experiment throughput linearly; a Hidden Consistent Evaluation protocol that delivers a reliable evaluation signal; and ReAct agents that dynamically scope their actions and debug interactively. On MLE-bench-30, AIRA_2 achieves a mean Percentile Rank of 81.5% at 24 hours and 83.1% at 72 hours, outperforming the strongest baseline, which achieves 72.7%. On AIRS-Bench, AIRA_2 exceeds human state-of-the-art on 6 out of 20 diverse research tasks. Ablations confirm that each architectural component is necessary, that performance follows a predictable scaling law that transfers across LLM backbones, and that the "overfitting" reported in prior work was driven by evaluation noise rather than true data memorization.
PaperID: 3201, Poster
Abstract: Offline multi-objective optimization aims to identify high-quality trade-off designs using only a fixed dataset of previously evaluated candidates. Existing methods often cast this problem as guided generative sampling, producing a finite candidate pool by steering a diffusion or flow model toward Pareto-preferred regions. We study a different interface: deterministic Pareto set amortization. Our method, Offline Pareto Set Learning (OffPSL), trains a preference-conditioned map h_\pi:\Delta^m-1\rightarrow \mathcalX that returns one design for each requested trade-off direction. Because offline data contain no preference–solution labels, OffPSL trains this map using two offline-only signals: a coordinate-wise pessimistic surrogate scalarization and a diffusion-style denoising-residual support regularizer learned from offline designs. The denoising model is used only as a frozen support prior and is not used to sample candidates. At inference, OffPSL produces exactly one candidate per reference direction with no iterative sampling, test-time optimization, or filtering from a larger generated pool. On standard offline MOO benchmarks, OffPSL is competitive with recent guided generative baselines while keeping the learned preference-to-design map directly queryable and diagnosable.
Authors:
Xinyu Wang, Jinbo Bi, minghu songAbstract: Pocket-conditioned molecular diffusion updates ambient atom coordinates, but many lead-optimization objectives are expressed on quotient features such as distances, contacts, and anchored substructures. We introduce budgeted quotient-residual guidance (QRG), an inference-time correction that makes these quotient objectives active without retraining the molecular generator. QRG lifts quotient covectors to metric-horizontal ambient directions and delivers them through a trust budget set by the frozen sampler's own step norm: quotient geometry chooses the direction, while sampler motion sets the scale. We derive the horizontal lift, closed-form sampler-scaled update, KL/kinetic interpretation around a frozen reverse step, equivariance conditions, and a product-budget split for separately normalized section and residual controls. Controlled quotient tasks confirm that sampler-relative delivery activates signals that raw local quotient gradients leave dormant. On frozen TargetDiff backbones, official seed-0 CBGBench ligand-generation/editing sweeps show practical quality--runtime gains: CheapMain improves validity from 0.815 to 0.864 on fragment growing, 0.664 to 0.707 on scaffold hopping, and 0.681 to 0.712 on linker design, while PCMain improves fragment/scaffold and remains near-neutral on linker. Novelty remains 1.000 and diversity is preserved in the matched multi-seed molecular slice, giving task-dependent improvements without sampler retraining or backbone modification. Overall, QRG provides a lightweight route to quotient-aware inference for frozen molecular samplers with explicit runtime accounting.
PaperID: 3203, Poster
Abstract: Differentially Private Stochastic Gradient Descent (DP-SGD) is the dominant approach to private deep learning, but it pays compute and privacy budget \emphper trained model: each new architecture, ensemble member, or downstream fine-tuning round incurs additional privacy cost. We propose \textscSPS (Summarize--Privatize--Synthesize) and its enhanced variant \textscSPS+, dataset-distillation algorithms that release a private synthetic dataset by privatizing intermediate activation statistics from a public pretrained model. Once released, the dataset is post-processed freely: any downstream architecture, ensemble, federated aggregation, or continual update incurs \emphzero additional privacy cost. In particular, a single \textscSPS+ dataset distilled from a Wide ResNet-22-8 transfers zero-shot to architectures with different inductive biases (Vision Transformers, Swin Transformers, and ConvNeXt), all without re-incurring privacy. Empirically, \textscSPS+ is competitive with state-of-the-art DP-SGD across \epsilon \in \1,2,4,8\ on CIFAR-10 and CIFAR-100, with a notable advantage on CIFAR-100 under strict privacy, achieving a 5.8% gain at \epsilon=1 over compute-matched DP-SGD. \textscSPS+ additionally outperforms prior generation-based DP methods by a large margin and supports private federated and continual learning out of the box.
PaperID: 3204, Poster
Abstract: Large language models (LLMs) in expert domains face a trade-off between \emphagentic tool use, which externalizes expertise behind brittle multi-turn runtimes, and \emphend-to-end specialization, which internalizes expertise only implicitly at task-level granularity. We propose \emphinternalized meta-capabilities: atomic, parametrically internalized, and composable reasoning units that an LLM can select and combine within a single chain-of-thought (CoT). We instantiate this idea as Med-Agentic for medical visual diagnosis, where the model emits a capability-tagged structured CoT in one forward pass, turning capability invocation into tagged token generation rather than external tool calls. For training, Med-OPD performs multi-teacher on-policy distillation with \emphtoken-level dual-factor teacher routing conditioned on both the image and the active capability tag, strictly generalizing per-prompt teacher routing. We show that this routing gives an unbiased estimator of the same expected teacher-KL gradient. Across public medical imaging datasets spanning nine anatomies and eight modalities, Med-Agentic consistently improves in-distribution and out-of-distribution diagnostic accuracy over the strong competitors.
PaperID: 3205, Poster
Abstract: Every control action on a physical system is already a do-operator intervention, yet causal discovery and control are typically treated as separate stages. This separation leaves deployed controllers unable to test whether the causal graph they rely on remains valid under changing operating conditions. We present \textscPolicyGRID, a closed-loop framework that uses actuator interventions to validate candidate causal edges, fits a structural causal model over the validated graph, and reuses that graph for both policy optimization and latent-regime monitoring. In a multi-actuator physical simulation with latent operating regimes (building energy management), interventional graph validation reduces closed-loop energy use by 21% at a standard comfort target relative to an otherwise identical observation-only graph, while improving four-way latent-context identification from 61.2% to 80.8%. A targeted ablation shows that the gains originate in graph pruning. Validation removes 16 of 47 candidate edges; with coefficient training held fixed, the validated graph reduces closed-loop energy relative to the unvalidated candidate set, while refitting coefficients on the same graph does not. On a live physical testbed, the same pipeline runs without code changes, performs 21 physical interventions, discovers 16 edges, and produces a target-responsive control policy where the observation-only baseline is insensitive to the comfort-energy target. These results show that physical controllers can use their own actions to validate causal structure and improve downstream control under latent regime shifts.
PaperID: 3206, Poster
Abstract: Undesired model behaviors can emerge or re-emerge at any stage of training. Once observed, how effectively can we suppress them? To study this question, we first construct two model organisms---language models trained to exhibit a specific misalignment---with qualitatively distinct structures. In the narrow-trigger organism, undesired behavior is elicited by an identifiable prompt feature. In the broad-trigger organism, misalignment surfaces diffusely across open-ended responses without a separable cue. We then formalize post-hoc behavioral suppression and evaluate methods along three criteria: immediate suppression of the target behavior, durability under further training, and preservation of general capability. We compare standard methods---SFT, GRPO, and gradient-ascent unlearning---along these dimensions while controlling the number of gradient updates. Building on prior work on alignment shallowness, we hypothesize that more durable suppression requires the ability to recover from a wider range of misalignment states. We introduce Prefix-Expanded Adversarial Reinforcement Learning (PEARL), a GRPO variant that augments rollouts with continuations from intermediate states sampled along cached organism trajectories. Across both organism settings and two model families (Qwen3-4B, GPT-OSS-20B), PEARL achieves lower exploit rates while preserving task accuracy, and yields lower reactivation rates under both targeted and benign capability fine-tuning than SFT and GRPO baselines. Together, our framework and method offer a step toward more durable post-hoc removal of undesired behaviors.
Authors: Wonsuk Jang, Thierry Tambe
Abstract: Diffusion Transformers (DiTs) achieve state-of-the-art video generation quality, but their substantial memory and computational footprints hinder edge deployment. Quantization can reduce these costs, yet existing methods often degrade video quality due to high activation variation and the difficulty of preserving semantic and temporal coherence. We propose SemanticDialect, which advances block-wise mixed-format quantization. In this framework, each block selects an optimal format (dialect) from a candidate set (formatbook), which is augmented with lookup tables that store quantization errors and quantized indices, enabling efficient per-block format selection and quantization with minimal online overhead. We further introduce attention-guided activation decomposition, which reduces quantization error via residual quantization, and semantic-aware dialect assignment (SeDA), which reduces cross-token quantization inconsistency by enforcing format uniformity among semantically correlated tokens. Experiments demonstrate that SemanticDialect outperforms prior quantization methods and block-wise formats (MXFP4, NVFP4) while approaching FP16 quality on Open-Sora 2.0. We also validate hardware deployability through RTL design and GPU kernel implementation.
PaperID: 3208, Poster
Abstract: We study principal-agent problems in which a principal commits to an outcome-dependent payment scheme (i.e., a contract) in order to induce an agent to take a costly action leading to a favorable outcome. We consider the online extension of the classical (one-shot) principal-agent problem, in which the principal repeatedly interacts with agents by proposing contracts over multiple rounds. The principal has no information about the agents and, crucially, does not observe their actions. As a result, the principal must learn an optimal contract using only the realized outcomes observed at each round. We focus on the setting with binary-actions and single-dimensional agent types, where the agent's private type represents their cost per unit-of-effort. For adversarial-type sequences, we provide tight \widetilde\Theta(T^2/3) regret guarantees. Remarkably, this rate is completely independent of the number of outcomes m. The upper bound is based on two key components: 1) a reduction to one-dimensional threshold optimization problem and 2) a non-uniform discretization to handle the non-Lipschitz nature of the problem. Moreover, in the case of a single (fixed) hidden type, we show that it is possible to improve the rates and provide a tight \widetilde\Theta(\sqrtT) regret bound. Our algorithm is based on an explore-then-commit strategy where we first approximately learn the hidden type via a stochastic binary search, and then we commit to a ``robustified'' near-optimal contract.
PaperID: 3209, Poster
Abstract: Developing generalist GUI agents capable of robust long-horizon execution remains a key challenge. While Reinforcement Learning (RL) offers a promising path for improvement, it is severely bottlenecked by the lack of reliable reward signals in open-ended environments. Ground-truth verification typically relies on expensive human annotation, which is prohibitively scarce at the scale required for training, while existing VLM-based judges frequently hallucinate success on superficial visual cues. This scarcity of trustworthy supervision fundamentally limits the scalability of current methods. In this work, we propose I-Judge, a novel framework for Self-Calibrated Reward Modeling. We leverage Inverse Dynamic Modeling (IDM) as a self-supervised training objective to learn the causal dynamics of GUI interactions from massive unlabeled trajectories, effectively bypassing the bottleneck of scarce human labels. We then introduce a runtime calibration mechanism that weights reward signals by the IDM's action consistency, filtering out spurious successes. Extensive end-to-end RL experiments on OSWorld, ScienceBoard and AndroidWorld demonstrate that our method significantly accelerates convergence and improves final agent performance. Notably, I-Judge demonstrates strong generalization to unseen domains. All the code, data and models will be made publicly available to foster further research.
PaperID: 3210, Poster
Authors: Ruoyu Guo, XIN K LIN, Maurice Pagnucco, Yang Song
Abstract: Domain generalised semantic segmentation (DGSS) has recently benefited from combining multiple foundation models, such as vision-language and vision foundation models, to leverage their diverse representations for generalisation. Existing multi-model approaches share a common design choice: cross-model fusion is restricted to depth-aligned layers, implicitly assuming that mutually beneficial information resides at the same depth across models. However, foundation models pretrained under different objectives may develop distinct representation hierarchies, making the optimality of depth-aligned fusion questionable. Moreover, we observe that the foundation models have different depth-wise domain sensitivity. Motivated by these observations, we propose a Sparse All-Layer Connector (SALC) that enables each layer to access information from preceding layers of another model. As the number of accessible layers grows with depth, dense aggregation may hurt generalisation. SALC therefore learns to adaptively select a sparse subset of informative layers. We further introduce a candidate dropout regularisation that strengthens sparsity and encourages SALC to explore diverse layers, leading to more robust selection. Across four foundation model combinations and three DGSS evaluation settings, the proposed design demonstrates stronger generalisation than conventional depth-aligned fusion. Code will be public upon acceptance.
PaperID: 3211, Poster
Abstract: Existing conformal prediction methods for time-to-event outcomes leverage only baseline covariates, producing prediction intervals that are insufficiently informative to facilitate decision making. We propose History-Aware Prediction Sets (HAPS), a conformal framework that constructs prediction sets for individual event times using covariate histories observed up to a decision time, targeting coverage among individuals who have survived to this time. HAPS handles right censoring adjusted for time-varying confounders via inverse probability of censoring weighting. When the censoring weights are consistently estimated, it achieves PAAC (probably asymptotically approximately correct) coverage among survivors. We further propose two doubly robust extensions of HAPS to weaken reliance on consistent estimation of the censoring distribution. In simulations, HAPS and its extensions reduce median prediction interval length by up to 75% relative to baseline comparators while maintaining close to nominal coverage. On two public benchmark data sets, HAPS reduces the median interval length by up to 60% for predictions at year 5, compared to the baseline comparators.
PaperID: 3212, Poster
Abstract: Motion blur arises from the temporal integration of a continuous sharp signal over a finite exposure window, yet existing learning-based methods sidestep this physical model and predict only the sharp signal itself: most single-image deblurring methods recover a single frame at the exposure center, while blur-to-video methods predict a fixed set of frames. We introduce PickMoment, a continuous-time reformulation that directly learns the interval-mean blur over arbitrary sub-intervals of the exposure with a single deterministic model. Drawing an analogy to MeanFlow's average-velocity formulation, we train the model with three supervisions derived from the blur integral: an empirical reconstruction loss from available subframes, an additivity loss that enforces self-consistency across overlapping sub-intervals, and a sharp-frame loss anchored at the zero-interval limit. A single trained model unifies single-image deblurring, blur-to-video generation, and continuous-time pick-a-moment recovery as different queries to the same network, with no separate training for each task. Our PickMoment achieves state-of-the-art performance among generative-based deblurring methods on GoPro and HIDE while competitive against restoration-based methods on RealBlur, and the highest per-frame fidelity on GoPro-7 blur-to-video, all in a single forward pass without iterative sampling.
PaperID: 3213, Poster
Authors: Xiaoming Wu, Wei Wenyu, Xin Wang, Ming Yang
Abstract: Federated adaptation of large vision-language models (VLMs) is appealing for privacy-sensitive applications but must contend with client data heterogeneity, strict communication budgets, and on-device inference constraints. Recent methods predominantly introduce additive adaptation interfaces, such as learnable prompts or adapter modules, which attach extra trainable components to the frozen backbone, inflating both parameter counts and inference latency. We challenge this additive paradigm and ask: can robust federated VLM adaptation be achieved through a compact intrinsic interface alone? We answer this question with FedIBS (Federated Intrinsic Bias Selection), a module-free framework that fine-tunes only the biases in feed-forward network projections. FedIBS operates on existing backbone parameters and introduces zero additional inference cost. To make such compact intrinsic adaptation viable under heterogeneous clients, we further propose a Fisher-guided progressive masking mechanism that disentangles shareable and personalized bias coordinates. Each client retains personalized coordinates locally and communicates only the shareable subset for aggregation, enabling sparse communication while preserving client-specific adaptation. Extensive experiments on diverse datasets and heterogeneity settings demonstrate that FedIBS consistently outperforms recent federated prompt- and adapter-based baselines in both personalization and generalization, all while achieving substantially lower communication cost.
Abstract: Sharpness-aware minimization (SAM) is an effective method for improving the generalization of federated learning (FL) by steering local training toward flat minima. Under data heterogeneity, however, device-side SAM searches for locally flat basins that are incompatible with the flat region preferred by the global objective. We identify this structural failure mode as flatness incompatibility, which explains why improving local flatness alone may provide limited training and generalization improvement for the global model. We reveal that flatness incompatibility arises from data heterogeneity and the friendly adversary phenomenon, and is further amplified by local updates and partial device participation. To mitigate this issue, we propose Federated Learning with variance-suppressed sharpness-aware minimization (FedVSSAM), which constructs a variance-suppressed adjusted direction and uses it consistently in local flatness search, local descent, and global update. FedVSSAM anchors both perturbation and update directions to a more stable global direction, instead of correcting only an isolated local perturbation. We establish non-convex convergence guarantees of FedVSSAM and prove that the mean-square deviation between the adjusted direction and the global gradient is effectively controlled. Experiments demonstrate that FedVSSAM mitigates flatness incompatibility and outperforms the baselines across diverse FL settings.
PaperID: 3215, Poster
Abstract: In large-scale outdoor LiDAR localization, Scene Coordinate Regression (SCR) achieves sub-meter accuracy, but its deployment in high-precision autonomous driving is hindered by a "last-mile" problem: the point-wise prediction paradigm causes unphysical trajectory jittering, rendering localization kinematically discontinuous despite low mean errors. To bridge the gap between high-precision localization and temporal kinematic consistency, we propose BiLi, a spatiotemporal synergistic relocalization framework. Spatially, BiLi utilizes Implicit Manifold Regularization (ISR) driven by an offline neural implicit field. Formulated as a novel point-to-manifold projection constraint, ISR strictly anchors predicted coordinates to continuous physical zero-level sets. Temporally, to bypass computationally heavy implicit optimization, we distill the teacher's kinematics into a lightweight feed-forward network, compelling it to learn robust, equivariant physical priors. Finally, a Differentiable Robust Manifold Gating mechanism dynamically fuses the global spatial predictions with the temporal kinematic states via adaptive state updating, facilitating efficient real-time deployment. Extensive experiments demonstrate that BiLi fundamentally resolves the "last-mile" challenge. Our method outperforms state-of-the-art approaches by 25% and 65% on the Oxford and NCLT benchmarks, respectively, delivering drift-free, kinematically coherent, and highly accurate real-time localization.
PaperID: 3216, Poster
Authors: Keegan Kang, Avery Hood, Benedict H Wong
Abstract: CMinHash promises a new direction for designing MinHash algorithms with reduced variance and storage space via circulant permutations. However, the original variance analysis is difficult to extend to other MinHash estimators. We present a new approach that gives explicit and easily comparable approximate variance expressions, and use it to analyze the Minner estimator. Numerical experiments show that our variance expression matches the empirical mean square error (MSE) of estimates. Our work thereby demonstrates how to apply circulant permutation to other MinHash estimators and compute their approximate variance.
PaperID: 3217, Poster
Abstract: Policy optimization algorithms that reuse rollouts build updates of the form \rho A, where \rho is the importance ratio from the current policy to the rollout policy and A is an advantage estimate. Recent RL algorithms such as PPO and GRPO rely on clipping \rho for stable policy updates, which introduces further nonsmoothness in the optimization landscape. \emphIs clipping the best we can do? In this paper, we propose signed logarithmic smoothing (Signed LS), \psi_\beta(z) = \mathrmsgn(z)\,\beta^-1\log(1+\beta|z|), applied to the full weighted update rather than to the ratio alone. Building on this surrogate, we introduce KRAFT, a GRPO-style policy-optimization objective that replaces the clipped surrogate with Signed LS. For theoretical justification, we consider an adaptive off-policy contextual-bandit setting: we compare the deviation structure of Signed LS with clipping-style surrogates, prove surrogate fidelity for Signed LS, and obtain regret transfer against a fixed comparator policy. The comparison highlights a structural distinction: ratio clipping introduces explicit tail-bias and threshold-dependent concentration terms, whereas Signed LS controls distortion through a smooth second-order quantity. Empirically, our contextual-bandit experiments provide evidence of these effects. In RLVR experiments on math-reasoning benchmarks and multiple LLM backbones, we observe that KRAFT consistently shows lower KL from the reference policy and higher entropy throughout training without incurring accuracy loss compared to GRPO.
PaperID: 3218, Poster
Abstract: Reservoir computing turns sequence learning into linear regression on fixed dynamical features, but those features come from a single highly dependent trajectory, so the true amount of usable data is often unclear. We develop a transport-based learning theory for contractive echo state networks driven by stochastic inputs by viewing the reservoir as a Markov process and proving a Wasserstein contraction that guarantees a unique stationary state law and explicit mixing rates. Using contraction-to-concentration tools, we obtain finite-sample deviation bounds for time-averaged features and empirical covariance estimates, and translate them into excess-risk and stability guarantees for ridge-trained linear readouts. The theory yields an explicit effective-sample-size principle and exposes a sharp trade-off: making reservoirs more “critical” can increase dynamical memory while slowing mixing and raising the data requirements for reliable training. Overall, the results complement echo-state and capacity analyses by providing verifiable design rules that link stability parameters to mixing time, generalization, and forecasting robustness under independent or weakly dependent inputs.
PaperID: 3219, Poster
Abstract: Sparse autoencoders (SAEs) are a central tool in mechanistic interpretability. However, existing SAEs are primarily trained per layer. The modeling subspace is therefore fixed by layer identity, independent of which token-layer states actually drive each prediction. We argue that this constraint contributes to several limitations observed in layer-wise SAEs, including low feature utilization, high dictionary redundancy, and features that lack direct behavioral grounding. In this paper, we propose ScopeSAE, which selects the modeling subspace per token by attributing each prediction to its most influential token-layer state via normalized gradient-based attribution, and learns features over the resulting prediction-relevant subspace. Empirically, ScopeSAE yields an effect we term reconstruction-better-than-original. Written-back reconstructions of the SAE produce lower next-token cross-entropy than the original activations, an outcome that, to our knowledge, has not previously been reported for SAEs and that inverts the usual reconstruction-fidelity trade-off. Through interventional analyses and a KL fine-tuning counter-experiment, we show that this effect is attributable to ScopeSAE's prediction-relevant subspace itself rather than to architectural changes. ScopeSAE further improves effective feature count, interpretability, utilization, and dictionary redundancy over existing layer-based baselines, suggesting that choosing the SAE modeling subspace by predictive relevance leads to more useful and behaviorally meaningful features.
PaperID: 3220, Poster
Authors: Soohyun Cha, Jungi Lee, Jaewoong Sim
Abstract: Speculative decoding accelerates generative inference of large language models (LLMs) by using a small draft model to propose multiple candidate tokens, which are then verified in parallel with the target model in a single decoding iteration. While the state-of-the-art method of tree-based speculative decoding helps improve generation throughput over non-speculative inference, deploying it in LLM serving systems often yields suboptimal performance due to dynamically changing serving conditions. In particular, our analysis shows that the optimal tree configuration---the one that maximizes performance---varies with two key serving conditions: request rate and per-request characteristics. Based on the analysis, we present AdaTree, a plug-in component for LLM serving systems that dynamically adapts tree configurations to varying serving conditions. AdaTree predicts model execution time and acceptance length across different tree configurations and selects the one that maximizes speculative decoding efficiency. To capture the non-linear relationship between tree configuration and model execution time, AdaTree employs decision trees as its core modeling primitive. Our evaluation shows that AdaTree consistently outperforms both chain-based and static tree-based speculative decoding across diverse serving conditions.
PaperID: 3221, Poster
Abstract: Differential privacy (DP) has become the standard for privacy-preserving data analysis, yet its practical implementation in federated or distributed settings often faces a significant utility gap compared to the central model. While existing frameworks, including local, shuffle, and secure aggregation models, offer vetted primitives, they frequently suffer from high error rates for some important functionalities such as selection. Our first contribution is to show that this is inherent in the aggregation model: any private selection algorithm relying on the standard noisy aggregation primitive requires \Omega(\sqrtd/\varepsilon) samples. To address this limitation, we propose adding a new primitive that implements a version of the Sparse Vector Technique (SVT) and argue that it has simple implementations under various trust models. We demonstrate its power by developing an \varepsilon-differentially private algorithm for the selection problem. For selection in dimension d, our algorithm achieves an additive error of O\left(\frac\log d\varepsilon\right) matching that of the central model. Further, we show this can be done using a near-optimal number of SVT queries. Together with existing lower bounds for the shuffle model, this result establishes an exponential separation between the SVT model and shuffling or aggregation-based models. We demonstrate the model's broad applicability to diverse tasks such as F_p loss minimization, heavy-hitter identification, and sparse mean estimation.
PaperID: 3222, Poster
Authors: Athanasios Masouris, ZHENG JING, Benjamin S Chandler, Hadi Jamali-Rad
Abstract: Layout generation for real-world facilities is a challenging problem, requiring reasoning over irregular site boundaries, heterogeneous orientations, access-aware placements, and motion-planning feasibility. Yet, most existing layout benchmarks in the generative AI space target simpler placements over rectangular domains and rely on distributional metrics such as FID and IoU that reward conformity to dataset priors, thus discounting design innovation. Motivated by these gaps, we introduce ALPS-Bench, a benchmark of 1,000 professionally annotated real-world facility layouts paired with an instance-specific scoring protocol grounded in a structured design manual. As a strong baseline for ALPS-Bench, we propose \textttCANDO, a training-free multi-agent framework in which specialized agents iteratively refine layouts through a verification-grounded loop, concentrating reasoning on strategic spatial decisions. We demonstrate that \textttCANDO surpasses state-of-the-art baselines on the widely adopted PubLayNet and RICO benchmarks by a significant margin, establishing cooperative agentic design as a broadly effective recipe for constraint-aware layout synthesis. Code and benchmark will be released upon acceptance.
Authors:
Louis Bethune, Victor Guilherme Turrisi da Costa, Bruno Mlodozeniec, Pau Rodriguez, Lokesh Boominathan, Nikhil Bhendawade, Amitis Shidani, Joris Pelemans, Theo X. Olausson, R Devon Hjelm, Paul Dixon, Joao Monteiro, Pierre Ablin, Vishnu Banna, Arno Blaas, Nick Henderson, Kari Noriy, Dan Busbridge, Marco Cuturi, Joshua Susskind, Irina Belousova, Luca Zappella, Russell Webb, Jason RamapuramAbstract: Discrete diffusion models have emerged as strong alternatives to autoregressive language models, with recent multimodal work either finetuning unimodal diffusion bases or distilling autoregressive backbones for bi-modal generation. Diverging from these approaches, we introduce the first tri-modal \glsmdm \emphpretrained from scratch on text, image-text, and audio-text data, where all generation tasks are learned simultaneously within a single model. We systematically study the design space governing stability and efficiency at scale: we conduct the first empirical characterization of the critical batch size B_\textcrit under the SDE reparameterization of AdamW, and introduce a drift--horizon interpolation parameter \gamma that balances gradient noise reduction and optimization horizon when scaling the token budget. We derive multimodal scaling laws and ablate modality mixing ratios, noise schedules, and anti-masking. Lastly, we pretrain a 3B model on 6.4T tokens, achieving competitive results in text generation, text-to-image, and text-to-speech.
PaperID: 3224, Poster
Abstract: With the rapid adoption of large language models (LLMs) and parameter-efficient fine-tuning (PEFT) methods, the risk of backdoor attacks has become more severe. Existing backdoor purification methods typically rely on at least one of the strong assumptions, such as prior knowledge of triggers, access to clean references, or aggressive retraining, and they often lack comprehensive evaluations. These constraints substantially limit their practical applicability. To overcome these challenges, our work proposes purifying LoRA-tuned LLMs without these assumptions and even without post-hoc retraining of the suspect parameters. Our objective is to significantly reduce the attack success rates (ASR) while preserving both (i) the base model’s general capabilities and (ii) the new downstream skills learned through the adapter. Through a series of ablation studies, we progressively scale our approach from a single layer in a text classification setting to a full-parameter LLM in the generative task. Through careful data curation and feature approximation, we extract high-fidelity backdoor directions and, for each layer or head, construct orthogonal null spaces in both the input and output channels, onto which the LoRA updates are projected. Empirically, our null-space projection method reduces the ASR from nearly 100% to less than 10%, while preserving the base model’s benign performance as well as the adapter’s abilities learned during downstream task adaptation.
Abstract: Non-contrastive self-supervised learning (SSL) is an effective framework for predictive representation learning, but popular (and in practice effective) methods such as SimSiam, BYOL, I-JEPA or DINO, which rely on a form of self-distillation to train a teacher-student network, remain poorly understood as they typically do not minimize a well-defined objective. We analyze the dynamics of a variant of the Joint Embedding Predictive Architecture (JEPA) using a regularized linear regressor to predict the learned representations of two views of the data from one another, and fully characterize its stability: non-collapsed stable equilibria align with leading nonlinear canonical correlation subspaces, while collapsed equilibria may also be stable attractors. Motivated by this result, we introduce PEIRA, a non-contrastive SSL method with an explicit objective defined through the trace of the optimal linear regressor. We show that its only stable equilibria are nontrivial global minimizers and recover the same canonical correlation subspaces, with regularization selecting the effective dimension. Experiments on ImageNet-1K and CIFAR-10 show PEIRA is competitive with VICReg and LeJEPA baselines, and qualitative empirical results support the theory.
PaperID: 3226, Poster
Abstract: Training neural networks at scale requires careful per-layer scaling of initialization, learning rate, and other optimizer hyperparameters, with the right rules depending jointly on the architecture, the optimizer, and which axis (width D, depth L, batch size B, etc.) is being scaled. Current practice is to derive each rule by hand, as in \muP for width, CompleteP for depth, SDE-based and EMA-timescale arguments for batch size, and tailored prescriptions for matrix-preconditioned optimizers; as architectures, optimizers, and scaling axes continue to evolve, this hand derivation is increasingly a bottleneck to combining these advances at scale. We show that these derivations can be systematized by solving a linear system over hyperparameter exponents to satisfy the scaling constraints imposed by each primitive instruction of the training program, analogous to dimensional analysis in physics that enforces matched units on both sides of an equation. We develop \emphautomatic parameterization (AutoP), a system that reads these rules off a traced training program and solves the resulting system to assign a scaling exponent to every hyperparameter, and provide a concrete JAX implementation. From a few primitive rules, AutoP recovers \muP for width on Transformers, Monarch-structured Transformers, mixture-of-experts, and ResNets under Adam, SignSGD, Muon, and AdaMuon; CompleteP for depth; and the known batch-size prescriptions for the learning rate and AdamW weight decay. It also identifies new rules for the recurrence count in looped transformers, the context length in linear-attention and MLP-Mixer models, and the batch-size scaling for Muon's learning rate, all of which we demonstrate as empirically beneficial. Analogous to automatic differentiation, our results suggest that robust hyperparameter scaling rules can be automated, freeing practitioners to scale up new architectures and optimizers without rederiving the underlying theory or risking subtle errors that quietly cost compute at scale.
PaperID: 3227, Poster
Authors:
Kian Kenyon-Dean, Alina Selega, Ihab Bendidi, Jordan M Sorokin, Luca Bertinetto, David Errington, Hayley Donnella, Oren KrausAbstract: RNA sequencing produces rich and diverse datasets of gene expression, offering compelling insights into cellular state and function that have many applications in drug discovery. Modeling such data is challenging due to inherent technical noise and experimental batch effects, as evidenced by many existing transcriptomic foundation models (FMs) underperforming relative to linear baselines. Such results raise the question of whether deep representation learning provides a distinct advantage over the direct use of raw transcript counts. Our work explores this by developing a new self-supervised model, TxFM, with a focus on inductive representation learning evaluations. TxFM employs a masked autoencoding approach tailored to diverse RNA-seq count data, and our ablation study empirically identifies crucial architecture configurations required for strong transfer performance. Additionally, we curate a public training corpus, DiverseRNA-1.4M, and find that TxFM trained on this curated dataset yields high-fidelity gene representations that outperform FMs trained on atlas-scale corpora over 100× larger. Overall, our results indicate that inductive self-supervised learning is a viable modeling approach for transcriptomics representation, provided a careful synthesis of model architecture and training data curation.
Authors: Gaurab Pokharel, Sanmay Das, Patrick J Fowler
Abstract: Street-level bureaucrats, such as caseworkers, triage nurses, and school administration, operate under a fundamental tension. The official policy on who to prioritize for services may conflict with bureaucrats estimation of who would benefit most from those services. The social theory of street-level bureaucracy argues that frontline agents use discretionary authority to bridge this gap, overriding policy recommendations when professional judgment identifies cases where doing so would meaningfully improve outcomes. As algorithmic recommendations become embedded in allocation settings, understanding the structure of this discretionary behavior is fundamental to designing tools that account for, rather than suppress, this value of professional judgment. We formalize discretion as a finite-horizon dynamic allocation problem. An agent implements a default policy but holds a limited budget of K overrides across T periods. Each period, the agent observes whether deviating from the default would yield a welfare gain and must decide whether to spend a scarce override now or conserve it for future opportunities. We show that the optimal policy is a state-dependent threshold rule and prove a behavioral invariance: for location-scale families of improvement distributions, the rate at which an optimal agent exercises discretion is independent of the amount of welfare gains and depends only on the distribution's shape. Fat-tailed gains induce patience (conserving overrides for rare, high-stakes cases); thin-tailed gains induce routine, front-loaded spending. We also compare the conceptual implications of our model with operational data from a homelessness service system. Override rates shift with short-run capacity openings in scarce housing, rising when recent exits free slots, and falling when interventions are congested. This could be the result of a perceived increase in the discretionary budget. Further, the temporal pattern of spending is front-loaded: override rates peak on Mondays and early in the fiscal year, consistent with the routine, early-spending behavior our model predicts under thin-tailed improvement distributions. These findings identify a systematic structure in both human and theoretically optimal decision-making, with direct implications for the design of human-AI allocation systems that are both more efficient and equitable.
Authors: Niki Foteinopoulou, Ignas Budvytis, Stephan Liwicki
Abstract: Low-Rank Adaptation (LoRA) has become a widely adopted technique in text-to-image diffusion models, enabling the personalisation of visual concepts such as characters, styles, and objects. However, existing approaches struggle to effectively compose multiple LoRA adapters, particularly in open-ended settings where the number and nature of required skills are not known in advance. In this work, we present LoRAtorio, a novel train-free framework for multi-LoRA composition that gates each adapter using only signals already inside the denoiser, specifically, patch-level cosine (dis)similarity between each LoRA's predicted noise and the base model's. We refer to this as ``intrinsic'' guidance because no external parameters are introduced. Our method is motivated by two key observations: (1) LoRA adapters trained on narrow domains produce unconditioned denoised outputs that diverge from the base model, and (2) when conditioned out of distribution, LoRA outputs show behaviour closer to the base model than when conditioned in distribution. These patch-level similarities are used to construct a spatially-aware weight matrix, which guides a weighted aggregation of LoRA outputs. To address domain drift, we further propose a modification to classifier-free guidance that incorporates the base model's unconditional score into the composition. We extend this formulation to a dynamic module selection setting, enabling inference-time selection of relevant LoRA adapters from a large pool. On the ComposLoRA benchmark, LoRAtorio achieves state-of-the-art performance, with a CLIPScore that does not deteriorate as the number of composed adapters grows, surpassing the strongest prior method by up to 1.3% in CLIPScore, and reaching a 76.92% GPT-4V win rate against the closest baseline. We further show that the framework is backbone-agnostic: the intrinsic-similarity principle transfers to rectified-flow architectures (FLUX.1-dev) without retraining. Code will be made available.
PaperID: 3230, Poster
Abstract: The widespread deployment of generative AI has made it increasingly difficult to distinguish synthetic content from real data. Consequently, synthetic data is inevitably incorporated into the training pipelines of future model generations, forming a . Prior work has studied the effects of such recursive self-consuming training, but analyses have largely been limited to models, where a model consumes only its own synthetic data, or to simplified interactions between two models. This paper takes a first step toward understanding self-consuming generative models, in which multiple models consume synthetic data generated by one another through complex interaction pathways. We introduce a theoretical framework representing models as nodes in a directed, weighted graph, with edge weights governing the flow of synthetic data among models. Using this framework, we analyze the long-term behavior of networked models under retraining dynamics, establishing conditions for convergence and characterizing the resulting fixed points. We further investigate how the system's long-term stability and diversity are shaped by each model’s access to real data, cross-model data consumption, and the structure of the interaction graph.
PaperID: 3231, Poster
Abstract: Histological whole-slide images (WSIs) are central to computational pathology but pose severe computational challenges due to their extremely high resolution, often spanning several gigabytes per slide. To enable scalable learning, existing methods apply self-supervised data condensation to reduce computational cost, but typically rely on heuristic prototype learning and do not explicitly preserve learning-relevant feature distributions for downstream tasks. In response, we introduce a principled reformulation of WSI condensation as a distribution-matching problem under a fixed representational lens, and develop NICER, a tractable approximation framework based on a nonparametric prior with slide-adaptive capacity. Experiments on five histopathology datasets, together with clinical evaluation from a board-certified pathologist, show that NICER consistently outperforms prior methods, achieving an average accuracy improvement of 7.44% while offering improved efficiency–accuracy trade-offs, highlighting the benefits of principled, distribution-aware condensation for scalable histological representation learning.
PaperID: 3232, Poster
Abstract: Weather foundation models are trained on dense reanalysis fields, but these inputs are not noiseless measurements of the atmosphere. Observation and representativeness errors vary across variables, locations, and weather regimes, so a transformer can spend substantial compute fitting token-level fluctuations that are less reliable than the surrounding signal. We present UtoMe, an observation-uncertainty-guided token merging framework for training-efficient weather transformers. Motivated by a three-way variance view that separates observation uncertainty from aleatoric and epistemic uncertainty, UtoMe builds a patch-level importance score from dynamics saliency and ensemble-derived spread, protects forecast-sensitive low-uncertainty tokens as anchors, and merges low-importance feature-similar tokens with an uncertainty-aware matching cost. Three variants cover different deployment settings: external uncertainty at both training and inference, external uncertainty only during training, and no external uncertainty. On ERA5 forecasting at two spatial resolutions, UtoMe preserves forecast skill under moderate token reduction and degrades much more gracefully than train-time ToMe under aggressive merging. The method reduces FLOPs by up to 45%, showing that observation uncertainty can be used as a direct compute-allocation signal for weather foundation models.
PaperID: 3233, Poster
Authors: Zhiguang Liu, Yi Shang
Abstract: ARC tasks can be formulated as visual transformation problems, where a model must infer a high-level transformation rule from a small set of demonstrations and apply it consistently across a grid. While dense Vision Transformers (ViTs) perform well in this setting, they entangle local visual state and global transformation logic within a shared token representation, which can hinder consistent reasoning. In this paper, we introduce RoSeViT, a role-separated Vision Transformer that explicitly decouples high-level instruction inference from visual workspace computation. RoSeViT uses controller tokens to aggregate transformation-level information and workspace tokens to represent the visual grid, with a structured attention mechanism that routes nonlocal interactions through the controller. This design is motivated by a formal analysis showing that enforcing a shared instruction representation can reduce the effective complexity of visual transformation modeling. Under the standard VARC training and evaluation protocol for ARC, RoSeViT consistently outperforms dense ViT baselines. In a controlled backbone replacement setting, it improves ARC-1 accuracy from 54.5 to 56.6, and with recurrent refinement reaches 58.8. When integrated into the full system, RoSeViT-Ensemble improves performance from 60.4 to 62.6 on ARC-1. These results demonstrate that explicitly separating instruction and workspace representations leads to more effective visual reasoning on ARC tasks.
Authors:
Shuai Zhang, Jiayu Hu, Zijie Chen, Zeyuan Ding, Yi Zhang, Yingji Zhang, Ziyi Zhou, Junwei Liao, Shengjie Zhou, Yong Dai, Zhenzhong Lan, Xiaozhu JuAbstract: Current embodied VLM evaluation relies on static, expert-defined, manually annotated benchmarks that exhibit severe redundancy and coverage imbalance. This labor‑intensive paradigm drains computational and annotation resources, inflates costs, and distorts model rankings, ultimately stifling iterative development. To address this, we propose Agentic Automatic Evaluation (A2Eval), the first agentic framework that automates benchmark curation and evaluation through two collaborative agents. The Data Agent autonomously induces capability dimensions and assembles a balanced, compact evaluation suite, while the Eval Agent synthesizes and validates executable evaluation pipelines, enabling fully autonomous, high-fidelity assessment. Evaluated across 10 benchmarks and 32 models, A2Eval compresses evaluation suites by 85%, reduces overall computational costs by 77%, and delivers up to a 4.6× speedup while preserving evaluation quality. Crucially, A2Eval corrects systematic ranking biases, improves human alignment to Spearman's \rho=0.85, and maintains high ranking fidelity (Spearman's \rho=0.87), establishing a new standard for high-fidelity, low-cost embodied assessment. Our code and data will be public upon acceptance.
Abstract: Modern off-policy reinforcement learning algorithms often rely on simple uniform replay sampling and it remains unclear when and why non-uniform replay improves over this strong baseline. Across diverse RL settings, we show that the effectiveness of non-uniform replay is governed by three factors: replay volume, the number of replayed transitions per environment step; expected recency, how recent sampled transitions are; and the entropy of the replay sampling distribution. Our main contribution is clarifying when non-uniform replay is beneficial and providing practical guidance for replay design in modern off-policy RL. Namely, we find that non-uniform replay is most beneficial when replay volume is low, and that high-entropy sampling is important even at comparable expected recency. Motivated by these findings, we adopt a simple Truncated Geometric replay that biases sampling toward recent experience while preserving high entropy and incurring negligible computational overhead. Across large-scale parallel simulation, single-task, and multi-task settings, including three modern algorithms evaluated on five RL benchmark suites, this replay sampling strategy improves sample efficiency in low-volume regimes while remaining competitive when replay volume is high.
PaperID: 3236, Poster
Authors: Di Wu, Shuaidong Pan
Abstract: Merging different models into a single model is highly desirable for multi-task deployment, particularly in regulated domains like healthcare. In such settings, models are often locally fine-tuned due to strict data privacy policies, necessitating that aggregation operates exclusively on the fine-tuned weights. Existing meth- ods apply a uniform merge across all tasks, ignoring the geometric relationships among their learned subspaces. When task updates occupy misaligned subspaces, projecting them onto a shared low-rank basis discards task-relevant directions, a failure mode named subspace interference. Through this study, we show that the per-task projection error of a merged model is governed by the principal-angle alignment between each task and the shared subspace, so that grouping geometri- cally compatible tasks provably reduce this error within each group. Motivated by this observation, we propose Group-wise Rank-Aware Model Merging (GRAM), which uses the principal-angle geometry of task vectors to decide which tasks should be merged and then performs a standard SVD-based merge within each coalition; the pipeline requires only the fine-tuned weights at merge time. With only one additional model (K=2), GRAM significantly mitigates subspace in- terference and consistently outperforms prior merging methods across language understanding (GLUE with Flan-T5), 8-task CLIP vision, and medical instruction tuning (Qwen2-7B-Instruct with four English/Chinese medical LoRAs).
PaperID: 3237, Poster
Abstract: Learning with noisy labels has been extensively studied for classification, yet its regression counterpart remains relatively underexplored. Unlike classification, regression involves continuous-valued targets where label noise exhibits heterogeneous magnitudes across instances, which poses conflicting requirements on the model: correcting small-to-moderate deviations requires sensitivity to label values for instance-level refinement, whereas large deviations caused by corrupted or outlying annotations render exact label values unreliable, necessitating reduced reliance on exact labels and instead leveraging ordinal structure among samples. These requirements call for methods that capture both global structure and sample-specific deviations in a unified manner. To this end, we propose GRAIN, a dual-granularity framework for robust noisy regression. At the fine-grained level, we derive an explicit residual learning mechanism from a robust loss formulation, introducing instance-specific auxiliary variables to decouple label noise from predictive signals. At the coarse-grained level, we introduce a binning-based contrastive regularization that enforces an order-preserving feature structure aligned with label proximity, addressing noise beyond the capacity of local residual refinement. The two levels are coupled via a histogram-based bridging mechanism that enables bidirectional information flow: coarse-grained structural guidance establishes a well-ordered feature space that anchors residual learning, while fine-grained residual refinement corrects instance-level deviations that structural regularization alone cannot resolve. Experimental results show that GRAIN consistently improves robustness and generalization under diverse noise conditions.
PaperID: 3238, Poster
Abstract: Whenever LLM training involves multiple outputs from the same input---agentic multi-turn trajectories, self-consistency reasoning, beam search distillation, or multi-rollout SFT/RL---the tokens form a tree-structured trajectory with shared prefixes. Existing pipelines linearize such data and treat each branch independently, causing substantial redundant computation that scales with the number of samples. We derive that averaging the loss over all branches is algebraically identical to a per-token weighted loss, reducing the problem to computing each token's log-probability exactly once. We propose DFS serialization of the tree, which visits every token exactly once, and adapt full-attention and SSM layers to ensure the resulting log-probabilities match independent per-branch computation exactly. For memory-constrained settings where the full tree exceeds GPU capacity, we propose Redundancy-Free Tree Partitioning, which achieves zero redundant computation with peak memory bounded by a single root-to-leaf path. Together, these contributions form Tree Training, achieving up to 6.2× end-to-end training speedup on dense and MoE models for both supervised fine-tuning and reinforcement learning.
PaperID: 3239, Poster
Abstract: We view fairness as a property of distributional stability. Rather than assessing a predictor under a fixed data distribution, we study how its predictions change under perturbations that modify the composition of protected groups. A predictor is fair if it remains stable under such shifts. Under this perspective, several classical notions of fairness arise as stability with respect to specific perturbations, with the associated unfairness gap given by a Lipschitz constant of a prediction-rate functional. This formulation also yields guarantees that hold uniformly over a range of demographic compositions at test time, without requiring knowledge of the deployment distribution. It leads to a learning procedure based on convex combinations of reweighted predictors, formulated as a second-order cone program, for which we establish generalization bounds. Experiments on standard benchmarks illustrate the approach.
PaperID: 3240, Poster
Authors: Jeremías Figueiredo Paschmann
Abstract: Offline RL methods often regularize toward the empirical occupancy of the dataset, but this reference can mismatch the deployed policy. In partially observed or decentralized settings, data may be pooled from hidden modes, histories, or conventions that are unavailable at execution time. The pooled occupancy can then require incompatible action laws for the same deployed input, or joint correlations that a decentralized policy class cannot represent, so regularization toward it can favor a low return projection. We formalize this failure through reference coherence: a reference is coherent when it does not rely on information hidden from the deployed policy, and when its action law is representable by that policy. We propose Closest Slice DICE (CS-DICE), a DICE variant that regularizes toward a coherent slice of the training data rather than the pooled occupancy. A hard or soft selector chooses the slice used as the reference term, while the deployed observations, actor class, and factorization constraints remain unchanged. Most of the results focus on the soft reverse-KL case, where the selector only changes the reference used by an otherwise standard DICE update. Controlled examples show that incoherent pooled references collapse to suboptimal behavior, while CS-DICE recovers high return policies when the selected or inferred slice is coherent with the deployed interface. We also identify a boundary case where pooling is benign because the relevant signal is revealed before the conflict. Finally, a CS-CoMA-DICE study on MaMuJoCo shows that selected references can be integrated into a neural cooperative MARL pipeline while preserving decentralized execution.
PaperID: 3241, Poster
Abstract: Dexterous manipulation requires large-scale robot interaction data, yet collecting real-world demonstrations is costly, while sim-to-real transfer remains challenging due to visual and geometric discrepancies. We present , a unified real-to-sim-to-real framework that enables photorealistic simulation learning and reliable deployment for dexterous manipulation. Our method reconstructs real scenarios using a compositional 3D Gaussian Splatting (3DGS) representation with tri-level optimization, obtaining high-fidelity objects and background while preserving global scene consistency. A dexterous-hand-centric calibration pipeline, supported by an interactive online platform, further achieves millimeter-level robot–scene consistency for efficient Gaussian-simulation alignment. Built upon the reconstructed scenes, we train privileged policies in simulation and distill them into visuomotor policies operating on rendered 3DGS observations, combined with closed-loop execution for robust transfer. Experiments demonstrate accurate real-to-sim reconstruction (F1 score 0.981) and strong sim-to-real performance, achieving high success rates across six real-world dexterous manipulation tasks.
Authors:
Terry Chen, Zhifan Ye, Bing Xu, Zihao Ye, Timmy Liu, Ali Hassani, Tianqi Chen, Andrew Kerr, Haicheng Wu, Vedaanta Agarwalla, Fengzhe Zhou, Shangda Li, Yang Xu, Yu-Jung Chen, Hanfeng Chen, Aditya Kane, Ronny Krashinsky, Ming-Yu Liu, Vinod Grover, Luis Ceze, Roger Bringmann, John Tran, Wei Liu, Feng Xie, Michael Lightstone, Humphrey ShiAbstract: Agentic Variation Operators (AVO) are a new family of evolutionary variation operators that replace the fixed mutation, crossover, and hand-designed heuristics of classical evolutionary search with autonomous coding agents. Rather than confining a language model to candidate generation within a prescribed pipeline, AVO instantiates variation as a self-directed agent loop that can consult the current lineage, a domain-specific knowledge base, and execution feedback to propose, repair, critique, and verify implementation edits. We evaluate AVO on attention, among the most aggressively optimized kernel targets in AI, on NVIDIA Blackwell (B200) GPUs. Over 7 days of continuous autonomous evolution on multi-head attention, AVO discovers kernels that outperform cuDNN by up to 3.5% and FlashAttention-4 by up to 10.5% across the evaluated configurations. The discovered optimizations transfer readily to grouped-query attention, requiring only 30 minutes of additional autonomous adaptation and yielding gains of up to 7.0% over cuDNN and 9.3% over FlashAttention-4. Together, these results show that agentic variation operators move beyond prior LLM-in-the-loop evolutionary pipelines by elevating the agent from candidate generator to variation operator, and can discover performance-critical micro-architectural optimizations that produce kernels surpassing state-of-the-art expert-engineered attention implementations on today’s most advanced GPU hardware.
Abstract: Exemplar-based image editing applies a transformation defined by a source-target image pair to a new query image. Existing methods rely on a pair-of-pairs supervision paradigm, requiring two image pairs sharing the same edit semantics to learn the target transformation. This constraint makes training data difficult to curate at scale and limits generalization across diverse edit types. We propose Delta-Adapter, a method that learns transferable editing semantics under single-pair supervision, requiring no textual guidance. Rather than directly exposing the exemplar pair to the model, we leverage a pre-trained vision encoder to extract a semantic delta that encodes the visual transformation between the two images. This semantic delta is injected into a pre-trained image editing model via a Perceiver-based adapter. Since the target image is never directly visible to the model, it can serve as the prediction target, enabling single-pair supervision without requiring additional exemplar pairs. This formulation allows us to leverage existing large-scale editing datasets for training. To further promote faithful transformation transfer, we introduce a semantic delta consistency loss that aligns the semantic change of the generated output with the ground-truth semantic delta extracted from the exemplar pair. Extensive experiments demonstrate that Delta-Adapter consistently improves both editing accuracy and content consistency over four strong baselines on seen editing tasks, while also generalizing more effectively to unseen editing tasks. Code will be made publicly available.
PaperID: 3244, Poster
Authors:
Ritabrata Ray, Sahil Dharod, Burak Varıcı, Nicholas Boffi, Pradeep RavikumarAbstract: Recent work has shown that contrastive representation learning can be understood as estimating a positive-pair, PMI-like kernel, and that spectral factorization of this kernel yields useful eigenfunction representations. This paper asks what happens beyond squared-error spectral geometry. Modern representation learning objectives are rarely pure \ell_2 kernel-approximation objectives: they use conditional KL, InfoNCE, JS/NCE, Brier, Hellinger, least-squares ratio fitting, and other statistical discrepancies. We show that these discrepancies do not merely provide alternative estimators of the same affinity; they induce different local geometries for the extraction of low-rank representations. To formalize this, we introduce divergence principal functions, a geometry-aware generalization of eigenfunctions. Eigenfunctions are recovered as the special case corresponding to squared-error geometry. For smooth discrepancies, we show that divergence principal functions locally solve a weighted spectral approximation problem, with weights given by the curvature of the divergence at the target affinity. This provides a loss-consistent bridge from positive-pair learning objectives to representation geometry and clarifies when spectral/eigenfunction representations are appropriate and when they are mismatched to the training objective.
Abstract: Zeroth-order (ZO) optimization provides a gradient-free alternative to first-order (FO) methods by estimating gradients via finite differences of function evaluations, and has recently emerged as a memory-efficient paradigm for fine-tuning large-scale models by avoiding backpropagation. However, ZO optimization has a fundamental tension between accuracy and query efficiency. In this work, we show that this can be substantially improved by unifying two complementary principles: (i) a projection-based subspace view that reduces gradient estimation variance by exploiting the intrinsic low-rank structure of model updates, and (ii) Muon-style spectral optimization that applies gradient orthogonalization to extract informative spectral structure from noisy ZO gradients. These form a unified framework of subspace gradient orthogonalization, which we instantiate in a new method, , admitting a natural interpretation as a low-rank Muon optimizer in the ZO setting. Extensive experiments on large language models (LLMs) and vision transformers (ViTs) demonstrate that ZO-Muon significantly accelerates convergence and achieves a win–win improvement in accuracy and query/runtime efficiency. Notably, compared to the popular MeZO baseline, ZO-Muon requires only 24.7% of the queries to reach the same SST-2 performance for LLM fine-tuning, and improves accuracy by 25.1% on ViT-B fine-tuning on CIFAR-100.
PaperID: 3246, Poster
Abstract: Long-tailed visual recognition is limited not only by biased optimization, but also by the under-retention of rare-class evidence in parameters learned from imbalanced data. MemTailor addresses this limitation by augmenting a frequency-aware multi-expert backbone with a hierarchical class-aware memory that preserves visual evidence outside the backbone. The memory stores per-class prototypes and bounded instance slots for each expert, allocates proportionally richer capacity to rare classes, and records both shallow and deep features to retain complementary discriminative cues. During training, same-class memory retrieval constructs class-adaptive feature augmentation and provides an auxiliary supervised objective on the augmented features. At inference, class-level memory quality and prediction uncertainty determine when stored evidence is allowed to correct expert logits, and the final prediction fuses memory-corrected logits across multiple input scales. Experiments on long-tailed visual recognition benchmarks show that MemTailor improves balanced recognition, especially for under-represented classes.
PaperID: 3247, Poster
Abstract: Rectified flow (RF) is motivated as optimizing the velocity field through trajectory crossings. However, this explanation contradicts prior work showing that exact crossings cannot occur, when the dimension exceeds 2 and the source and target distributions are continuous. We revisit this question and make three contributions. 1) By relaxing the geometric definition of a crossing, we prove that . This effect is hard to notice, and coloring serves as a useful indicator. 3) To address this breakdown, we propose , which corrects trajectories within the principal subspace of transport before lifting them back to the original space, thereby restoring smoothness and homeomorphism. Experiments confirm our analysis and show that
PaperID: 3248, Poster
Abstract: Federated Prompt Tuning (FPT) adapts large vision--language models by freezing the pretrained backbone and optimizing only lightweight prompt parameters across clients. Although this design improves communication and parameter efficiency, it creates an overlooked security risk. Since the backbone is shared and frozen, a malicious client can induce backdoor behavior through the representation space while keeping its uploaded prompt updates close to benign ones. We propose ShadowFPT, a targeted backdoor attack that exploits this frozen-backbone attack surface. ShadowFPT first pretrains a learnable \emphShadow Trigger against the frozen CLIP visual encoder, using either auxiliary public data or the malicious client's local data, so that triggered inputs are steered toward the target class in representation space. During federated prompt tuning, the malicious client adapts the trigger under the current global prompt and then optimizes its local prompt on both clean and triggered samples. Only prompt parameters are uploaded to the server, while the trigger and frozen encoders remain local. By shifting most of the attack burden from prompt updates to trigger-induced representation steering, ShadowFPT achieves targeted misclassification while preserving prompt-space stealthiness. Across multiple datasets, aggregation rules, and non-IID federated partitions, ShadowFPT increases the attack success rate from 19.81% to 90.36% in our main setting, while maintaining clean accuracy. It remains effective across textual, visual, and joint vision--language prompt tuning. These results identify frozen backbones as stealthy and underexplored backdoor surfaces in federated prompt tuning, suggesting that defenses based only on prompt-update anomaly detection are insufficient.
PaperID: 3249, Poster
Authors: Chenxi He
Abstract: Spatial transcriptomics methods typically cluster spots into coherent tissue domains, but many biological questions are target-specific: which spatial context predicts a particular cell state, gene program, or clinical phenotype? We propose NicheIB, a Targeted Spatial Bottleneck (TSB) that learns which spatial edges matter for a user-specified target via input-dependent Hard Concrete edge gates with controllable sparsity budgets. Training the same model on the same tissue with different targets—cortical-layer annotations, oligodendrocyte gene programs, excitatory neuron markers—produces distinct edge retention patterns whose correlation mirrors the biological similarity of the targets, confirmed across human cortex, mouse brain, and cancer tissues. The bottleneck also reveals which programs are spatially structured at a given resolution and which are not, serving as a diagnostic tool for spatial biology. On mouse brain region prediction, TSB retains only 20% of spatial edges while approaching the full-graph baseline (F_1 = 0.915 vs. 0.926), identifying a minimal informative subgraph from a mostly redundant spatial graph. External validation confirms that the retained subgraph selectively enhances spatial autocorrelation for target-relevant genes and is enriched for known ligand–receptor pairs, independently of the model.
Authors:
Stanislav Budzinskiy, Marián Gloser, Tolunay Yilmaz, Ying H Tham, Yuanyi Lin, Wenyi Fang, FAN WU, Philipp PetersenAbstract: Mixed-precision computations are a hallmark of the current stage of AI, driving the progress in large language models towards efficient, locally deployable solutions. This article addresses the floating-point computation of compositionally-rich functions, concentrating on transformer inference. Based on the rounding error analysis of a composition f(g(x)), we provide an adaptive strategy that selects a small subset of components of g(x) to be computed more accurately while all other computations can be carried out with lower accuracy. We then explain how this strategy can be applied to different compositions within a transformer and illustrate its overall effect on transformer inference. We study the effectiveness of this algorithm numerically on GPT-2 models and demonstrate that already very low recomputation rates allow for improvements of up to two orders of magnitude in accuracy.
Abstract: We study the task of agnostic learning of multiclass linear classifiers under the Gaussian distribution. Given labeled examples (x, y) from a distribution over \mathbbR^d × [k], with Gaussian x-marginal, the goal is to output a hypothesis whose error is comparable to that of the best k-class linear classifier. While the binary case k=2 has a well-developed algorithmic theory, much less is known for k \ge 3. Even for k=3, prior robust algorithms incur exponential dependence on the inverse of the desired accuracy in both complexity and representation size. In this work, we develop new structural results for multiclass linear classifiers and use them to design fully polynomial-time robust learners with dimension-independent error guarantees. Our first result shows that the standard multiclass perceptron algorithm requires super-polynomially many samples and updates, even with clean labels and Gaussian marginals, revealing a basic obstruction absent in the binary case. Our main positive result is a pairwise improper-learning framework which yields an efficient learner with error \widetilde O(k^3/2\sqrt\mathrmopt)+\epsilon for general k. Additionally, we develop a sharper localization-based framework which leads to error O(\mathrmopt)+\epsilon for k=3, and error \mathrmpoly(k)\mathrmopt+\epsilon for geometrically regular k-class linear classifiers.
PaperID: 3252, Poster
Abstract: Cross-tokenizer knowledge distillation lacks a shared coordinatesystem between teacher and student vocabularies, and existingmethods each drop one of the two ingredients comparison needs:shape-based objectives preserve density but discard token identity,lexical alignment restores identity but uses surface form as a proxyfor meaning, and text-space matching compresses the teacher signalinto chunk-level events. We introduce SOLA (Semantic Optimal LocalAlignment), which reconstructs the missing coordinate systemexplicitly: it partitions inputs into decoded-text blocks viaMinimal Complete Correspondence, estimates a cross-vocabularycorrespondence cost from calibration text, and runs entropicoptimal transport on a bidirectional local support retrieved aroundeach teacher prediction. Under matched training and decodingprotocols, SOLA outperforms direct cross-tokenizer baselines onDolly instruction following, an UltraChat general-instructionbenchmark, and code and math domain transfer; a coverage--scorecorrelation and component ablations attribute the gain to theconstructed support.
Abstract: Heterogeneous multi-agent LLM systems, where agents are powered by different model families, consistently outperform homogeneous configurations. Yet existing communication protocols either operate through text, discarding the sender’s internal representations, or require architectural homogeneity for latent-level transfer. We identify the entity grounding problem in cross-architecture communication: cross-attention bridges that transfer continuous representations across different LLM families suffer from rare-token compression collapse, where entity identity is lost in the continuous bottleneck (bridge-only F1 ∼30%). We propose XBRIDGE, a decode-free communication protocol that resolves this through two mechanisms. Lexical Anchor Mapping (LAM) maps the sender’s original context tokens to the receiver’s vocabulary, providing discrete entity anchors. A Latent Enrichment Bridge (LEB) lets the receiver query the sender’s hidden states for contextual enrichment. The entity anchors ground the bridge’s relational signals to specific entities through the receiver’s own self-attention. Across three model families (Llama, Qwen, Mistral), seven benchmarks, and both communication directions, XBRIDGE outperforms text-based communication (NLComm) on all 7 tasks per model pair with 11×lower latency, and exceeds the same-architecture KVComm ceiling on 6 of 7 tasks. LEB requires only 264M trainable parameters (3.8% of the receiver), trains on 587 samples in under 10 minutes, and adds negligible inference overhead.
PaperID: 3254, Poster
Abstract: Diffusion large language models (LLMs) have recently emerged as a promising alternative to autoregressive LLMs, enabling parallel generation through iterative denoising. Recent guidance methods improve diffusion LLMs by introducing a weak model to guide the prediction of the base model during denoising. However, the influence of the weak model during denoising remains less clear, and computing its prediction requires an additional forward pass at each denoising step. We present in this paper an in-depth analysis of how diffusion LLMs use contextual information during denoising. Our analysis reveals that diffusion LLMs rely more heavily on local context while often underutilizing global context. Based on this observation, we introduce Global Context Guidance (GCG), a feature-space guidance method that enhances global-context utilization without constructing an explicit weak model. Extensive experiments on LLaDA-8B, Dream-7B, and LLaDA-1.5 across diverse benchmarks demonstrate that GCG consistently outperforms existing guidance methods, while incurring much lower inference overhead. We will release our code upon acceptance.
Abstract: Model merging has emerged as a simple yet effective method to fuse models without re-training. Existing techniques operate at the level of individual layers, thereby overlooking the inter-layer dependencies inherent in deep networks. We show that this simplification induces distributional mismatches in intermediate activations during merging, as changes applied to early layers fail to propagate to downstream ones. We identify these mismatches as a form of covariate shift that compounds across the network. To address this issue, we propose Chain of Merges (CoM), a novel procedure that sequentially merges weights across layers while accounting for inter-layer interactions. In particular, CoM mitigates covariate shift through a series of regression problems, where input activations are recomputed at each step to reflect the updated representations. Additionally, we introduce a dynamic weighting mechanism that prioritizes task-layer pairs most susceptible to performance degradation, improving robustness during merging. Experiments on standard benchmarks demonstrate that CoM achieves state-of-the-art performance across heterogeneous tasks. Code is available in the supplementary material.
PaperID: 3256, Poster
Abstract: Many modern generative modeling and inference methods require optimizing objectives defined on parametric probability distributions. This leads to costly nested procedures, where each gradient step in first-order optimization techniques requires approximate sampling from the iterate parametric distribution. Gibbs gradient descent (GGD) was recently proposed as an efficient ``single-loop'' alternative by coupling sampling and optimization dynamics (Marion et al., 2025). However, its theoretical understanding remains limited, especially in the presence of particle approximations. This work provides a comprehensive analysis of GGD for the ideal infinite-particle system and its finite-particle approximation, explicitly quantifying the effect of particle noise via a sharp persistent \mathcalO(1/n) variance term. We then identify a finite-observable structural class for the Gibbs gradient, which contains our motivating examples, and introduce a control-variate version of GGD that reduces the particle-induced variance at the observable level. Our results reveal a fundamental trade-off between optimization and sampling, driven by the stochastic error induced by finite particle approximations.
Abstract: Modern LLMs continue to exhibit significant variance in behavior across languages, such as being able to recall factual information in some languages but not others. While typically studied as a problem to be mitigated, in this work, we propose leveraging this cross-lingual inconsistency as a tool for interpretability in mixture-of-experts (MoE) LLMs. Our knowledge localization framework contrasts routing for sets of languages where the model correctly recalls information from languages where it fails. This allows us to isolate model components that play a functional role in answering about a piece of knowledge. Our method proceeds in two stages: (1) querying the model with difficult factual questions across a diverse set of languages to generate "success" and "failure" activation buckets and then (2) applying a statistical contrastive analysis to the MoE router logits to identify experts important for knowledge. To validate the necessity of this small number of experts for answering a knowledge question, we deactivate them and re-ask the question. We find that despite only deactivating about 20 out of 6000 experts, the model no longer answers correctly in over 40% of cases. Generally, this method provides a realistic and scalable localization approach to increasingly complex LLMs, but also suggests redundant and highly dispersed knowledge parameterization in MoEs.
Abstract: Dataset distillation compresses a large training set into a small synthetic set that preserves downstream training utility. While most existing methods target training networks from scratch, modern visual transfer learning often uses frozen pre-trained encoders followed by lightweight linear probing. Existing distillation methods for this setting either unroll iterative linear-probe updates with trajectory-based gradient matching, or rely on closed-form formulations originally designed for from-scratch training with neural-tangent-kernel (NTK) approximations. Neither route exploits the fact that frozen-feature linear probing admits a closed-form solution determined directly by the pre-trained features themselves, with no infinite-width approximation and no inner-loop trajectory. We propose Closed-Form Linear-Probe Dataset Distillation (CLP-DD), a bilevel formulation that computes the linear probe induced by the synthetic set with a sample-space kernel ridge solver. The synthetic images are then updated by evaluating this induced classifier on real features through a temperature-scaled softmax cross-entropy, where the classifier columns act as learned class anchors in feature space. We further show that the choice of outer objective is decisive: pairing the closed-form inner solver with a standard MSE outer loss substantially underperforms trajectory-based methods, while the discriminative outer loss closes most of the gap. On ImageNet-100 with four pre-trained backbones, CLP-DD substantially improves over LGM without DSA and approaches LGM with DSA at a fraction of the computational cost. On ImageNet-1K, CLP-DD matches or surpasses LGM with DSA on three of four backbones while running roughly 14× faster and using less than one-eighth of the GPU memory.
PaperID: 3259, Poster
Abstract: For decades, philosophers and musicologists have debated which features of a performance are constitutive of the work and which are expressions of how it is being played. Computational evidence has been hard to assemble: audio embedding models conflate work identity with performance style, and existing interpretability tools for music generators treat learned representations as flat dictionaries that mix the two. We turn the question into an empirical one by introducing the dual-contrastive sparse autoencoder (DC-SAE), a two-branch sparse autoencoder that uses cheap work-level metadata to factor a generative music transformer's residual stream into work-identity and performance-content subspaces. Across jazz-standard, classical-work, and pop-cover corpora, the resulting decomposition is supported by probes and feature galleries that surface musically interpretable concepts on each side. Without performer supervision, the variation branch acquires structure related to performer identity, and steering along these directions can shift the perceived performer of generated audio while preserving the underlying work. Together, these results show that a generative music model has internalized a representation of interpretation that can be recovered, named, and steered.
Abstract: Concept-based Models (CMs) enhance interpretability in deep learning by grounding predictions in human-understandable concepts. However, concept annotations are costly and rarely available at scale within a single data source. Federated Learning (FL) could alleviate this limitation by enabling cross-institutional training over concept annotations distributed across multiple data owners. Yet, FL lacks interpretable modeling paradigms. Integrating CMs with FL is non-trivial: although FL supports heterogeneous and non-stationary client participation, it typically assumes a fixed shared architecture, whereas CMs may require architectural adaptation as the available concept set evolves. We propose (F-CMs), a new methodology for deploying CMs in evolving FL settings. F-CMs aggregate concept-level information across institutions and efficiently adapt the model architecture to changes in concept supervision while preserving privacy. Empirically, F-CMs maintain accuracy and intervention effectiveness comparable to training settings with full concept supervision, while outperforming on average non-adaptive federated baselines. Notably, F-CMs enable interpretable inference on concepts unavailable to a given institution, a key novelty over existing approaches.
PaperID: 3261, Poster
Abstract: Multimodal agents struggle with physical-world reasoning, particularly spatial reasoning from partial views, as the physical world is inherently complex. We think multimodal agentic reasoning about the physical world should ground its actions in deterministic feedback, rather than rely solely on internal imagination. Concretely, we propose to equip vision–language models (VLMs) with a physics engine (e.g., MuJoCo, UE5), which simulates the consequences of the agent's actions. However, granting the agent raw access to a physics engine is not sufficient, as the agent must first reconstruct the scene inside the physics engine. We therefore introduce the Visual Harness System, an orchestration system that decouples perception from reasoning with a physics engine as the backend. The perception module composes off-the-shelf perception skills, such as camera-pose estimation, open-vocabulary detection, and metric depth, into a structured scene map. The reasoning module then issues multi-turn tool calls that the engine executes over the map, and composes the final answer from the grounded feedback it receives. Experiments on spatial reasoning benchmarks show that our approach consistently outperforms strong baselines across both open- and closed-weight VLMs, with up to 16.57% improvement over prior methods on the MindCube benchmark.
PaperID: 3262, Poster
Abstract: Parameter-efficient image-to-video transfer adapts frozen image foundation models to video via lightweight trainable adapters. Existing methods insert the same adapter at every transformer layer, implicitly assuming all layers require equal temporal modeling. A layer-wise sensitivity analysis across four video-native model families and two pretraining paradigms reveals that this assumption is wrong: temporal sensitivity increases with depth, and the pretraining paradigm determines its growth profile. Vision-only models exhibit a gradual rise, while vision-language models show early suppression followed by late recovery. This finding motivates STRAP, a training-free adapter placement rule that derives a spatial-to-temporal transition boundary from a paradigm-matched video model. Spatial adapters are assigned before this boundary and temporal adapters after it, without modifying the adapter architecture itself. On four fine-grained action recognition benchmarks, STRAP improves accuracy by up to 11.8% over uniform placement, reduces trainable parameters by 29 to 48%, and lowers inference time by up to 10%. On standard benchmarks, it matches accuracy with fewer parameters.
Abstract: Reinforcement learning (RL) is well known to be data-intensive, and recent works have proposed leveraging data from similar systems to improve sample efficiency. In this work, we focus on the collaborative Linear Quadratic Regulator (LQR) setting and study the problem of learning under heterogeneous system dynamics. Due to the presence of a heterogeneity-induced additive bias, most existing collaborative methods are suitable only for low heterogeneity regimes and exhibit suboptimal performance under large heterogeneity. Based on the principle of Thompson sampling (TS), we propose an algorithm that, by leveraging data from other agents, yields a \emphpersonalized controller for every agent without incurring any additive heterogeneity-induced bias, i.e., with sublinear Bayes regret even under large heterogeneity. Under high heterogeneity regimes, we establish that our algorithm incurs \tilde\mathcalO( T^1-0.5\epsilon\sqrt\Delta) cumulative regret, where T denotes the time horizon, \epsilon\in[0,1] is a user-defined personalization parameter, and \Delta denotes an upper bound on the maximum dissimilarity among the agents' dynamics. Our algorithm and analysis also generalize the special case of no heterogeneity. In particular, under no heterogeneity, we establish that our algorithm incurs \mathcalO(\sqrtT/M) cumulative regret. From a distributed implementation perspective, our method incurs only logarithmic communication overhead. Additionally, our algorithm outperforms methods that do not personalize data from other agents and, in certain regimes, also outperforms methods that do not utilize any data from other agents.
Abstract: Distillation enables compact Vision-Language Models (VLMs) to obtain strong reasoning capabilities, yet the prompts driving this process are typically chosen via simple heuristics or aggregated from off-the-shelf datasets. We reveal a critical inefficiency in this approach: up to 69% of the prompts in standard chart / document reasoning datasets are effectively zero-delta, meaning the teacher and student already induce the exact same answer distribution. Training on these prompts provides minimal learning signal, causing student improvement to rapidly saturate regardless of data scale. To escape the zero-delta trap, we return to first principles: distillation fundamentally minimizes distributional divergence, and thus a prompt is valuable only if it exposes a functional capability gap between the teacher and student. We quantify this gap through answer divergence (\Delta), demonstrating that non-zero divergence is critical for effective scaling. Building on this insight, we propose a staged synthesis pipeline that repurposes existing datasets as seeds, actively targeting student failure modes to produce better prompts. The result is DeltaPrompts, a diverse dataset of 200k synthetic, high-divergence reasoning problems. We evaluate \textDeltaPrompts across three distinct settings: on-policy distillation with the target teacher-student pair, transfer to a novel model family without regenerating the data, and off-policy fine-tuning of a non-reasoning model. Across all scenarios, \textDeltaPrompts drives substantial gains, yielding up to 15% relative improvement even on top of a highly-optimized reasoning model (e.g., Qwen3-VL-8B-Thinking)---averaged over 10 benchmarks spanning chart, document and perception-centric reasoning.
PaperID: 3265, Poster
Abstract: Generative recommendation (GR), which directly generates sequential semantic IDs for items, has recently attracted increasing attention in recommender systems. Given that core recommendation metrics are inherently evaluated at the sequence level, reinforcement learning (RL) serves as a natural optimization framework. However, adapting recent group-based RL algorithms to GR presents notable challenges: token-level methods such as GRPO exhibit much more severe training instability, while sequence-level methods such as GSPO enforce uniform token weighting and thus ignore the heterogeneous roles of tokens. To address this dilemma, we propose Stable and Granular Policy Optimization (SGPO), a group-based RL algorithm tailored for GR. SGPO maintains a sequence-level ratio to promote optimization stability while injecting fine-grained token-wise advantages. By incorporating token log-probabilities and reward signs, SGPO explicitly upweights high-confidence tokens in positive trajectories to reinforce successful retrieval, while strictly penalizing them in negative trajectories to encourage exploration, and uses an adaptive scaling factor to bound the variance of reshaped advantages. Experiments on two public benchmarks and a billion-scale industrial dataset show that SGPO consistently outperforms six strong RL baselines and, when deployed on a commercial platform, achieves a 2.36% lift in click-through rate (CTR) and a 1.95% increase in advertising revenue in online A/B testing.
PaperID: 3266, Poster
Abstract: Adapting Large Language Models (LLMs) to downstream tasks is increasingly bottlenecked by the staggering memory overhead of first-order (FO) backpropagation. Zeroth-order (ZO) optimization has emerged as a compelling, memory-efficient alternative; however, it suffers from the curse of dimensionality, which introduces prohibitively high variance in gradient estimation. While LLM gradients empirically reside in low-rank subspaces, existing state-of-the-art ZO methods typically rely on randomly generated low-rank subspaces, which often fail to align with the underlying dominant gradient manifold, fundamentally limiting optimization efficiency. To bridge this gap, we propose AHZO, an efficient low-rank ZO fine-tuning algorithm that dynamically identifies dominant subspaces via the Average Historical Gradient (AHG). Motivated by the directional coherence of optimization trajectories, AHZO leverages the average historical ZO gradient over a period as a principled proxy for the true gradient. By applying singular value decomposition to AHG matrices, AHZO distills principal spectral signatures to construct low-rank bases that tightly align with the true dominant subspaces. Theoretically, we prove that AHZO significantly mitigates estimation variance, yielding superior convergence guarantees compared to random-subspace ZO methods. Extensive evaluations across diverse LLM architectures and benchmarks demonstrate that AHZO consistently outperforms existing ZO baselines, achieving performance highly competitive with memory-intensive FO fine-tuning.
Abstract: We study how visual information is routed in vision-language models (VLMs). Using causal patching on controlled synthetic and natural datasets, we find that models rely on two distinct pathways to solve visual tasks: A \emphdirect pathway, where visual information is retained in image token representations and read out by the final token at later layers, and a \emphtext-mediated pathway, where visual information is first transferred to the query tokens and then read out by the final token. Across three visual tasks, we show that pathway selection is task-dependent, and that data distribution and prompt design can also modulate which pathway is used to solve the image-based query. Moreover, using attention knockouts and corrupted-input patching, we find that these pathways are flexible, under certain interventions, models can rely on the text-mediated pathway as a fallback when the usual pathway is ablated. This behavior unifies findings in prior work and shows that ablation-based interventions can reveal what models could do rather than what they normally do. Together, our results provide a mechanistic characterization of visual information flow in VLMs and highlight the flexibility of their internal mechanisms under intervention.
PaperID: 3268, Poster
Abstract: Mixture-of-Experts (MoE) architectures have become essential for building capable large language models, with recent work demonstrating the benefits of fine-grained expert designs. However, training MoE models from scratch is computationally expensive, and existing upcycling methods that convert dense models to MoE face limitations: they either initialize experts as identical copies (limiting routing diversity) or use random perturbations (risking knowledge loss). We propose \textscDivMoE, a framework that addresses these challenges through two innovations. First, we introduce \emphdomain-specialized expert initialization, deriving fine-grained experts from models fine-tuned on distinct domains (mathematics, code, science, commonsense), providing meaningful diversity while preserving pre-trained knowledge. Second, we propose \emphdiversity-constrained routing, which enforces that each token selects at most one expert per domain group, structurally preventing routing collapse and enabling cross-domain knowledge composition. Experiments on two base models demonstrate that \textscDivMoE outperforms all upcycling baselines (47.0% vs.\ 44.4% average accuracy) and achieves competitive performance with MoE models trained from scratch, matching Moonlight-MoE at 64.5% average while being constructed via efficient upcycling.
PaperID: 3269, Poster
Abstract: Dynamic convolution improves the representational flexibility of CNNs by generating input-adaptive kernels, yet existing methods either inflate parameters linearly with the number of base kernels or sacrifice inference throughput due to runtime kernel assembly. We revisit dynamic convolution from the angle of cross-layer redundancy: a CKA analysis across modern CNN families shows that within-stage convolutional kernels are highly correlated, suggesting that an entire stage can be reparameterized by a single shared kernel together with lightweight, layer-specific transformations. Building on this observation, we propose KernelDNA, a parameter-efficient framework in which multiple convolutional layers in a stage share one ''parent'' kernel, while layer-specific ''child'' kernels are produced by a decoupled adapter that splits modulation into (i) an input-dependent dynamic channel gate and (ii) static spatial and filter modulations that can be pre-fused into the parent kernel before deployment. We provide a theoretical justification grounded in tensor approximation: when two kernels have CKA similarity \ge 1 - \delta, the multiplicative adapter incurs a Frobenius approximation error O(\sqrt\delta), and the implicit gradient consensus across child layers provably reduces the Rademacher complexity of the hypothesis class. Across ImageNet-1K and MS-COCO, KernelDNA achieves state-of-the-art accuracy-efficiency trade-offs over diverse backbones (ResNet18/50, MobileNetV2, ConvNeXt-Tiny), reducing parameters by 1.2-5× versus dynamic convolution baselines while retaining 90-99 % of the standard-convolution throughput---outperforming KernelWarehouse and FDConv across all backbones, and ODConv on 4 of 5 settings at a fraction of its parameter count.
PaperID: 3270, Poster
Abstract: Policy gradient methods may suffer from high variance, which limits their sample efficiency. Existing variance-reduction techniques largely focus on estimator design under a fixed sampling policy. In contrast, we ask whether the action-sampling distribution itself, often left at the on-policy default, can be optimized to reduce policy gradient variance. We show that this default is in fact not variance-optimal even when the state distribution is held on-policy. To exploit this, we introduce (BPO), a framework that treats the behavior policy as an action-level proposal distribution at each on-policy state, while keeping the state distribution itself on-policy. For a one-step importance-sampled policy gradient estimator, we sharply decompose its variance into an irreducible state term and a controllable action term, and show that the latter can be minimized in closed form via BPO. Under exact critic and sampling oracles, this yields a strictly tighter nonconvex SGD bound than on-policy sampling whenever the score-weighted critic is action-dependent. We then extend the analysis to approximate critics and parametrized behavior policies updated by a single stochastic step, and verify the resulting sample-efficiency gains empirically. These results identify action-level BPO as a principled mechanism for improving the sample efficiency of policy gradient methods.
Abstract: As frontier language models are increasingly deployed as autonomous agents pursuing complex, long-term objectives, there is increased risk of scheming: agents covertly pursuing misaligned goals. Prior work has focused on showing agents are capable of scheming, but their propensity to scheme in realistic scenarios remains underexplored. To understand when agents scheme, we decompose scheming incentives into agent factors and environmental factors. We develop realistic settings allowing us to systematically vary these factors, each with scheming opportunities for agents that pursue instrumentally convergent goals such as self preservation, resource acquisition, and goal-guarding. We find only minimal instances of scheming despite high environmental incentives, and show this is unlikely due to evaluation awareness. While inserting adversarially-designed prompt snippets that encourage agency and goal-directedness into an agent’s system prompt can induce high scheming rates, snippets used in real agent scaffolds rarely do. Surprisingly, in model organisms Hubinger et al. [2023] built with these snippets, scheming behavior is remarkably brittle: removing a single tool can drop the scheming rate from 59% to 3%, and increasing oversight can raise rather than deter scheming by up to 25%. Our incentive decomposition enables systematic measurement of scheming propensity in settings relevant for deployment, which is necessary as agents are entrusted with increasingly consequential tasks.
PaperID: 3272, Poster
Authors:
Miguel Fuentes, Brett Mullins, Cecilia Ferrando, Cameron Musco, Daniel SheldonAbstract: Generating differentially private (DP) synthetic tabular data remains a challenge, particularly when leveraging state-of-the-art query-answering mechanisms like ResidualPlanner or AIM+GReM. These mechanisms offer high utility through dense measurement collections \mathcalM; however, they are incompatible with existing synthesis tools. Specifically, graphical-model-based synthesis via PrivatePGM becomes intractable as the treewidth of the measurement graph grows. We introduce PRIME (Private Re-weighting of Initial Model Estimates), a scalable framework that decouples the synthesis process from the complexity of the measurement graph by restricting the support to a subset of the overall domain. PRIME transforms the synthesis problem into a convex re-weighting task over a candidate support set of size K. This formulation achieves an \mathcalO(K \cdot |\mathcalM|) per-iteration complexity, providing a treewidth-independent, drop-in solution for any DP mechanism. We show that, when provided the same measurements, PRIME achieves error comparable to PrivatePGM while handling dense measurement sets where PrivatePGM fails. Our empirical evaluation demonstrates that PRIME enables high-fidelity synthesis from query-answering mechanisms that were previously unusable for data generation. Ablation studies identify support quality as a key performance driver, underscoring the role of support generation in our framework.
Abstract: Decentralized online convex optimization (D-OCO) is a popular framework for distributed applications with streaming data. To tackle the communication bottleneck, previous studies have investigated D-OCO with compressed communication and proposed several algorithms that are variants of online gradient descent (OGD). However, for D-OCO with exact communication, the best existing algorithms are variants of follow-the-regularized-leader (FTRL). In this paper, for the first time, we propose two FTRL-type algorithms for D-OCO with compressed communication. Compared with OGD-type algorithms, our algorithms are more elegant in both algorithmic design and theoretical analysis. The key insight is that the dual update mechanism of FTRL allows us to make a simple application of the technique for average consensus with communication compression. More specifically, our first algorithm considers the full-information setting, and can match the existing regret bounds. Our second algorithm is designed for the bandit setting, and can significantly improve both the regret bounds and communication costs of existing algorithms.
PaperID: 3274, Poster
Abstract: Retrieval-augmented diffusion models (RAG-DMs) have significantly advanced image synthesis by incorporating external knowledge bases. However, their security vulnerabilities against malicious external data remain largely underexplored. Existing data poisoning attacks against diffusion models primarily corrupt internal parameters. They fail against RAG-DMs because clean retrieved images easily override these compromised parameters. Furthermore, recent attack tailored for RAG-DMs requires training the retriever, resulting in high computational costs and overfitting to specific retrievers. To bridge this gap, we propose PoisonedRDM, a novel and train-free data poisoning attack tailored to execute concept hijacking against RAG-DMs. Specifically, PoisonedRDM poisons the knowledge base with a minimal number of optimized images and optimizes adversarial perturbations through a joint optimization strategy, steering unknown retrievers toward the poisoned images via surrogate ensemble alignment and aligning the generated outputs with a malicious target concept, while adhering to a strict perturbation budget to maintain high visual stealthiness. Experiments show that PoisonedRDM effectively attacks RAG-DMs, achieving high success rates and outperforming state-of-the-art baselines.
PaperID: 3275, Poster
Abstract: Multimodal large language models offer a promising foundation for industrial anomaly detection (IAD), yet they remain unreliable for fine-grained diagnosis, where subtle defects require tight alignment between localized visual evidence and domain-specific semantics. Existing single-pass frameworks produce predictions without revisiting or validating intermediate evidence, which leads to spatio-semantic misalignment and error accumulation. More importantly, they lack mechanisms to actively control how additional evidence is acquired and integrated during reasoning, limiting their ability to resolve ambiguous cases. To address these limitations, we propose AirIAD, an agentic iterative reasoning framework that unifies progressive refinement with active information acquisition. AirIAD equips the model with two tools, Perceptive Zoomer and Knowledge Retriever, which enable it to iteratively gather complementary visual and semantic evidence. Through this process, the model performs explicit cross-verification between what is observed and what is known, refining intermediate hypotheses until convergence on a reliable diagnosis. This capability is enabled by a two-stage training pipeline. A supervised fine-tuning stage introduces a Spatio-Semantic Cross-Verification Chain-of-Thought, which structures coordinated perception and knowledge grounding. This is followed by a Spatio-Semantic Group-in-Group Policy Optimization stage, which provides dense step-wise supervision to encourage consistent self-refinement and robust spatio-semantic alignment. Extensive experiments on the MMAD benchmark show that AirIAD, built on Qwen3-VL-Instruct-4B, achieves state-of-the-art performance in anomaly detection and fine-grained diagnosis.
PaperID: 3276, Poster
Authors: Tianhao Wu, Wei Zhang, Haoran Pang, Haoran YANG, Hao Shi
Abstract: Clustering and vector quantization are core primitives for representation learning and large-scale retrieval, yet widely used methods often produce clusters or codewords of highly uneven quality, especially in noisy or heterogeneous data. In this paper, we propose a distortion-constrained optimal transport (DCOT) framework that explicitly enforces per-cluster bounds on the average assignment distortion, defined as the expected distance between data points and their assigned representatives, ensuring more consistent and reliable representations. Our formulation couples an optimal transport assignment objective with cluster-level distortion constraints to control cluster quality. To solve this constrained problem, we develop an efficient alternating optimization algorithm that iteratively updates transport plans, representatives (codewords), and dual variables. We further provide theoretical guarantees including entropic consistency, coordinate descent, and sublinear convergence. Importantly, DCOT serves as a plug-in assignment module that can be seamlessly integrated into a wide range of vector quantization methods. Replacing the assignment step with DCOT consistently improves reconstruction quality across diverse architectures and training settings, demonstrating strong robustness and general applicability.
PaperID: 3277, Poster
Authors: Kewei Feng, Jinbiao Nie, Xiaoyuan Zhang, Quanhua Liu
Abstract: Large language model (LLM)-driven evolutionary frameworks have emerged as a promising paradigm for iterative search in complex optimization and structured generation tasks. However, existing methods typically rely on heuristic variation operators and lack an explicit mechanism to extract directional information from observed performance feedback, limiting both efficiency and reliability. To address this, we propose SGEvolve, a semantic gradient-guided evolutionary framework for LLM-driven optimization. The key idea is to treat the performance differences between parent and offspring as zeroth-order evidence of an underlying ascent direction. Based on this view, we formulate semantic gradient extraction as a Maximum A Posteriori (MAP) estimation problem, where the LLM infers the most probable direction of improvement conditioned on observed evolutionary steps. The inferred gradients are then used to guide mutation and crossover, biasing generation toward more promising regions of the search space. We evaluate SGEvolve on a diverse set of tasks, including mathematical optimization, algorithm design, and real-world industrial problems. Experimental results demonstrate that SGEvolve consistently outperforms existing open-source baselines, achieving superior performance in both solution quality and search efficiency.
PaperID: 3278, Poster
Abstract: We propose a neural exploratory framework for variational problems over graphons, i.e., symmetric measurable functions [0,1]^2 \to [0,1], that arise as limits of dense graph sequences. Such problems are central to two parts of the mathematical literature: extremal graph theory, which studies graph parameter optimisation under given restrictions, often for those graphs with a large number of vertices, and the large deviation theory of dense random graphs, which characterises the structure of rare events through constrained graphon optimisation. We represent graphons by implicit neural networks and optimise graphon objectives by gradient descent. Our design combines three ingredients: a multi-scale sinusoidal residual architecture biased toward sharp, step-like graphons; an embedded solver that enforces a single empirical density constraint and is differentiated by implicit differentiation; and symmetry-aware Monte Carlo estimators. On generalised Turán problems that previously required substantial human effort, our framework rediscovers known optimal graphons without human intervention. Applied to open instances, it produces candidate optima, including a previously unreported family of extremal structures. Applied to the variational problems from the large deviation theory of Erdős–Rényi random graphs, the framework produces new candidate optimal graphons across both upper- and lower-tail regimes for multiple pattern graphs; for the case of the triangle graph and the upper-tail regime, these candidates improve upon the best-known reference construction of Lubetzky and Zhao.
PaperID: 3279, Poster
Abstract: The performance of AI coding agents is highly dependent on their underlying coding rules. However, existing coding rules are typically hand-crafted, making the process labor-intensive and often suboptimal. In this work, we propose RuleRefine, a framework for automatic coding rule refinement. RuleRefine maintains a pool of candidate coding rules and iteratively improves them. In each iteration, it employs an LLM-powered mutator module to generate variants from existing candidates, and then uses a judge module to evaluate these variants and update the pool with the best-performing ones. Extensive evaluations across two coding-agent frameworks, four backbone LLMs, and three benchmarks demonstrate that RuleRefine outperforms both manual engineering and existing prompt optimization baselines in terms of functional correctness of the generated code, code length, and/or generation cost (e.g., tokens used).
PaperID: 3280, Poster
Abstract: General-purpose behavior models increasingly internalize broad procedural knowledge at training time, making adaptation less a problem of learning entirely new behaviors than of reconfiguring existing competence for new goals, object relations, and skill interfaces. We study this as reconfigurability, the ability to realize new behaviors by restructuring components and relations cost-efficiently, under , where a task specification provides high-level goals or abstract directives but omits the concrete bindings and compositions required for execution. We propose RecPo, a framework that represents skills task-agnostically through multiple procedure views and materializes a new task by composing only the transition relations consistent with the task-relevant views. RecPo maintains a procedure knowledge base of local transition relations within each view. At deployment time, it interprets a task demonstration through the same views, converts observed view-level changes into a compositional query, and resolves the query into an executable skill sequence. On an extended Franka Kitchen benchmark with held-out long-horizon tasks outside the stored compositions, RecPo achieves 91.7% ordered subtask completion, substantially outperforming non-reconfigurable skill retrieval baselines. Real-robot case studies further show deployment-time adaptation through recomposition rather than task-specific retraining.
Authors: Su-Hyeon Kim, Yo-Sub Han
Abstract: Large language models from different families use different hidden dimensions, tokenizers, and training procedures, making behavioral directions difficult to compare or transfer across models. We introduce an anchor-projection framework that maps hidden representations from each model into a shared anchor coordinate space (ACS). Behavioral directions extracted from source models are projected into ACS and averaged into a canonical direction. For a new model, the canonical direction is reconstructed into its native hidden space using only anchor activations, without fine-tuning or target-specific direction extraction. We evaluate five instruction-tuned model families and ten behavioral axes. We find that same-axis directions align tightly across the Llama-Qwen-Mistral-Phi cluster on the ACS. This shared structure transfers to downstream tasks: held-out targets achieve \(0.83\) ten-way detection accuracy and \(0.95\) mean binary AUROC, while canonical steering induces refusal-rate shifts of up to \(\Delta = +0.46\) under distribution shift. Sensitivity analyses show that two source models and small anchor pools already suffice to approximate transferable directions. Overall, ACS provides a novel perspective on cross-family interpretability, revealing that representation-level transfer remains robust across model families. The code for this paper is available at \urlhttps://anonymous.4open.science/r/ACS-annoymous/.
PaperID: 3282, Poster
Abstract: Hallucination detection is critical for deploying large language models (LLMs) in real-world applications. Due to strong empirical performance, internal representation–based methods have emerged as the prevailing direction for detecting hallucinations, yet they remain largely model-specific and often fail to generalize to unseen LLMs. In this paper, we study an important yet underexplored problem, termed cross-model hallucination detection (CMHD), which aims to train hallucination detectors on source LLMs while ensuring performance on unseen target LLMs. The core challenge of CMHD lies in cross-model representation heterogeneity: hidden states from different LLMs exhibit severe semantic inconsistency, which limits the transferability of classical representation-based detectors. Through analysis and empirical validation, we show that relational structures over layers, tokens, and features capture transferable detection signals despite semantic inconsistency. Based on this observation, we propose cross-model hallucination detection via heterogeneity-oriented relational discrimination (CHORD), a relational graph discrimination framework that represents hidden states as joint relational graphs over layers, tokens, and features. It feeds these graphs into relation-aware attention to obtain structure-aware features, and uses meta-learning to shape these features toward transferable detection signals. Extensive experiments show that CHORD outperforms representative detection methods in cross-model generalization.
PaperID: 3283, Poster
Abstract: Large language model (LLM)-based multi-agent systems (MAS) achieve strong performance by employing specialized multiple agents, yet their performance depends on the prompt design of each agent. For MAS prompt optimization, textual gradient methods that guide prompt updates using natural-language feedback have emerged as a leading paradigm. In this paper, we identify limitations in two stages of existing textual gradient approaches: gradient extraction and gradient aggregation. In gradient extraction, previous works select the target prompt without verifying that it is responsible for the failure, and derive the gradient from the system-level final output rather than the agent-level intermediate output. In gradient aggregation, individual gradients are randomly grouped and concatenated, often mixing unrelated failure modes and producing prompts that fail to generalize. To address these limitations, we propose AgentGrad, a prompt optimization framework for multi-agent systems based on sequential intervention and semantic textual gradient abstraction. For each failure, sequential intervention modifies the behavior of one agent at a time to identify the target agent whose modification resolves the failure. The modified output of the target agent then serves as agent-level supervision for extracting a fine-grained gradient. Semantic textual gradient abstraction clusters semantically similar gradients to prevent mixing unrelated failure modes, and abstracts each cluster into a generalized gradient that captures the shared corrective pattern. Experimental results show that AgentGrad achieves state-of-the-art performance across five MAS benchmarks and reduces wall-clock optimization time by 2.5× on average compared to the next-fastest baseline.
PaperID: 3284, Poster
Abstract: Differential Privacy (DP) bounds the privacy leakage of a mechanism against worst-case membership inference, but the precise tradeoff between complex adversarial models and DP protections remains poorly understood. In this paper, we present a unified framework that generalizes the patchwork of existing bounds across membership inference, attribute inference, and data reconstruction attacks. Crucially, our framework is the first to evaluate attacks that target multiple individuals simultaneously and measure success beyond exact matches under a single cohesive bound. Our bounds capture this broad family of previously unexplored attack settings by relying solely on the privacy parameters and the adversary's baseline success rate (i.e. its prior without access to the mechanism's output). To illustrate this, we compare our high-probability guarantees to empirical attacks in two novel settings: extracting multiple non-uniform secrets (passwords and PII) from DP-finetuned language models, and reconstructing tabular data from noisy marginals. Ultimately, this framework provides a rigorous theoretical foundation to investigate the risk landscape of DP algorithms in new adversarial settings.
PaperID: 3285, Poster
Authors:
Rohan Hitchcock, Gary W Delaney, Jonathan H Manton, Richard Scalzo, Jingge ZhuAbstract: Neural networks that are trained on data arising from a physical system must somehow learn regularities induced by the underlying physical laws. In this setting, concepts from statistical physics provide powerful tools for analysing how physical laws influence the learning process. In this paper we explain how --- which measure the response of a neural network to perturbations of a parameter of the data distribution --- can be applied when learning from physical data. We show that such susceptibilities can identify the conditions under which specific input features are most informative for learning, guiding data selection, a prediction we validate experimentally. For distributions lacking a natural parameter, we propose introducing one via an arbitrary scalar function of the physical state, and demonstrate that susceptibilities track the emergence of specific capabilities (such as collision detection) during training.
Abstract: Knapsack and Top-k operators are useful for selecting discrete subsets of variables. However, their integration into neural networks is challenging as they are piecewise constant, yielding gradients that are zero almost everywhere. In this paper, we propose a unified framework casting these operators as dynamic programs, and derive differentiable relaxations by smoothing the underlying recursions. On the algorithmic side, we develop efficient parallel algorithms supporting both deterministic and stochastic forward passes, and vector-Jacobian products for the backward pass. On the theoretical side, we prove that Shannon entropy is the unique separable binary regularization choice for the local DP smoothing that yields permutation-equivariant operators, and characterize regularizers inducing sparse selections. On the experimental side, we demonstrate our framework on a benchmark on learning to predict Knapsack solutions, an extension of discrete VAEs, and a constrained dynamic assortment RL problem.
PaperID: 3287, Poster
Authors:
Raz Marshanski, Liran Nochumsohn, Mayank Jauhari Iitr, Boris Oreshkin, Omri AzencotAbstract: Probabilistic time-series forecasting requires models that simultaneously capture structured temporal dynamics, expressive uncertainty, and efficient inference, yet existing approaches fall short of this goal: latent dynamical models impose structure but rely on restrictive generation mechanisms, while modern generative methods such as diffusion and flow matching achieve flexibility at the cost of iterative and computationally expensive sampling. We introduce the Koopman Generative Operator (KGO), a probabilistic forecasting method that conceptualizes prediction as the evolution of structured uncertainty. KGO integrates three core components: (1) Koopman Patch Embedding (KoPE) for temporally consistent latent trajectory extrapolation; (2) Koopman Flow Matching (KoFM), which enables fast, single-step generation through a closed-form matrix exponential in a Koopman latent space, bypassing iterative sampling; and (3) an Adaptive Uncertainty Gate (AUG), which provides calibrated predictions by adapting uncertainty per-variable and per-horizon, corresponding to a learned decomposition of aleatoric uncertainty without manual tuning. By unifying these components, KGO achieves state-of-the-art accuracy across the majority of ProbTS datasets, outperforming existing methods on 12/17 benchmarks in CRPS and 11/17 in NMAE. Further, by eliminating iterative sampling, KGO delivers at least a 25× reduction in inference time compared to iterative generative models. These results establish KGO as a practical, principled, and highly scalable framework for next-generation probabilistic forecasting.
Abstract: We introduce cyclic denoising—repeated forward and reverse diffusion at controlled noise amplitudes—as an extraction attack for image diffusion models. Inspired by random organization in disordered solids, where cyclic mechanical perturbations anneal the system into increasingly stable configurations, cyclic denoising exposes regions of the learned distribution that remain largely inaccessible to standard sampling. We find that this dynamics drives samples toward attractors with a broad stability spectrum, with the deepest attractors exhibiting ultrastability: they can be regenerated from near-total corruption and sustained through thousands of noising-denoising cycles. Many of these deep attractors correspond to memorized training images, including stock photographs, brand watermarks, and web-crawl artifacts. Our extraction attack requires sampler-level control, including mid-process noise injection, but no gradients and no weight inspection. Crucially, it requires no prior knowledge of training data, captions, or prompts. In contrast, prior generate-and-filter attacks commonly rely on prompted generation using known or suspected training captions, followed by large-scale sampling and post-hoc similarity or membership-inference filtering. While cyclic denoising can also be applied with prompts, our main protocol is fully unconditioned. We demonstrate the phenomenon in Stable Diffusion v1.4, a large latent diffusion model, and in a smaller pixel-space DDPM, showing consistent behavior across latent- and pixel-space diffusion models. Across noise amplitudes, we observe a dynamical transition from trivial fixed points to structured memorized images, hierarchical partial absorption in which coarse scene layout freezes while fine details remain diffusive, basin hopping between memorized states with long residence times, prompt-stabilized memorized templates, and cross-initial-condition universality of the recovered attractor set. Together, these results establish cyclic denoising as both a physics-inspired probe of generative landscapes and a practical tool for memorization auditing, with implications for privacy, copyright compliance, and model fingerprinting.
PaperID: 3289, Poster
Authors:
Hongmin Zhao, Jiyuan Yin, Kexi Yan, Qinglai Wei, Jie ZHANGAbstract: Offline Inverse Reinforcement Learning (IRL) aims to recover a reward function and imitate expert behavior solely from offline demonstrations. While recent Offline IRL approaches employ approximate dynamics to mitigate distribution shift, their performance is constrained by the unstable minimax framework as well as simplified representation and utilization. To address these issues, we propose Offline Inverse Reinforcement Learning with Diffusion Planner (OIDP), which leverages diffusion models to achieve stable Offline IRL. OIDP formulates a model-based conservative Q-optimization through inverse soft Q-learning, which is provably concave in Q-space with a controllable performance gap. Accordingly, we introduce a unified diffusion planner that models dynamics, uncertainty estimator, and policy, it captures multimodal distributions, directly tightens the objective gap bound, and at execution performs trajectory-guided policy sampling using the internalized dynamics knowledge. Experimental results on D4RL benchmarks demonstrate that OIDP outperforms state-of-the-art offline approaches.
PaperID: 3290, Poster
Authors:
Tianjian Zhou, Jiang Jie, Yishan Li, Yifei ZhangAbstract: Self-supervised Vision Transformers in the DINO family encode rich semantic structure within their self-attention maps, yet existing parameter-efficient fine-tuning (PEFT) methods apply static, content-agnostic transformations that ignore this internal signal. We present DAGA (Dynamic Attention-Guided Adaptation), a PEFT module that repurposes a frozen backbone's emergent attention maps as instance-specific guidance: per-image attention is converted into channel and spatial gates that condition a lightweight bottleneck adapter. With only 1.4% additional parameters, DAGA reaches 85.9% top-1 on ImageNet-1K with DINOv3-ViT-B and surpasses recent PEFT baselines, with strong gains on dense prediction (+12.1% mIoU on ADE20K) and fine-grained classification (+1.6 / +3.5 on CUB-200 / FGVC-Aircraft). DAGA's effectiveness is tied to the semantic quality of the backbone's attention: gains are large on DINOv2/v3 and iBOT but small on MAE, CLIP, and DeiT, positioning it as the first PEFT module to operationalize emergent SSL attention as in-loop guidance and complementing self-distillation methods that refine attention itself. Anonymized code: https://anonymous.4open.science/r/ID1680.
PaperID: 3291, Poster
Abstract: While Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced multimodal reasoning, existing frameworks suffer from a critical “perception–reasoning gap.” Due to an over-reliance on seemingly relevant visual tokens, even minor perceptual perturbations can propagate through the reasoning process, leading to compounding hallucinations. To address this issue, we propose MIRA, a training framework that improves robustness to erroneous visual contexts while enhancing logical reasoning ability. By injecting filtered, deceptive visual contexts during training, we construct challenging “hard traps” that stress-test the model’s robustness to misleading information. Furthermore, we introduce a group-wise reflection-triggering mechanism that enables the model to autonomously detect and rectify inconsistencies between visual cues and logical constraints, which is ultimately internalized as an intrinsic policy behavior. Importantly, MIRA does not rely on human annotations. Experimental results across multiple benchmarks show that MIRA improves the base model by 7% and demonstrates superior robustness against deceptive contexts. These findings highlight that active error recovery is essential for reliable multimodal reasoning. Comprehensive ablation studies and analyses further provide insights into why MIRA is effective.
Authors:
Sungil Son, Hoseong Jung, Dahyun Oh, H. Jin KimAbstract: Discovering cooperative behavior in multi-agent reinforcement learning (MARL) is challenging due to the combinatorial complexity of joint state-action spaces, which hinders the emergence of coordinated behaviors from trial-and-error alone. Intrinsic rewards are often used to aid discovery, but naively combining them with team objectives can distort the learning signal, compromising task performance. In this paper, we propose EFXPLORER, a constrained exploration framework that maximizes exploration objectives subject to a constraint that preserves established task performance. We solve it via an epigraph reformulation that introduces adaptive exploration budgets. This approach separates intrinsic rewards from task objectives and regulates exploration through task feasibility. To further encourage diverse and temporally extended exploration, we incorporate a successor distance-based intrinsic reward that captures long-horizon dependencies. Empirically, our method outperforms strong baselines and induces novel cooperative strategies across SMAX, VMAS, and MPE benchmark suites.
PaperID: 3293, Poster
Authors:
Shuyuan Tu, Qi Tian, Zihan Yang, Yue Wu, Xintong Han, Weijie Kong, Jiangfeng Xiong, Jian-Wei Zhang, Zhao Zhong, Liefeng Bo, Zuxuan Wu, Yu-Gang JiangAbstract: Current open-source diffusion models struggle to generate stable and synchronized audio-visual content, particularly in scenarios demanding complex semantic reasoning. The root cause is that existing methods rely on coarse text embeddings from off-the-shelf encoders to guide audio-video denoising, which discards fine-grained semantics and, critically, lacks a shared long-horizon plan, leading to uncoordinated denoising trajectories and fragile cross-modal alignment. We propose Baton, the first framework that introduces explicit semantic planning into joint video-audio generation. Our key insight is that complementing coarse text guidance with semantically rich, modality-aware planned tokens, jointly reasoned and mutually aligned before denoising, can simultaneously restore fine-grained semantic detail and establish a shared blueprint that coordinates both audio and video denoising trajectories. Concretely, Baton first introduces the VA-Planner, a multimodal language model equipped with dual semantic alignment towers, where learnable queries cross-attend to both video and audio features to produce a pair of semantically aligned video and audio planning tokens as keyframe-level blueprints. These planning tokens are injected into the diffusion backbone via cross-attention layers, providing temporally grounded guidance complementary to coarse text embeddings. Since planning tokens do not share one-to-one spatial-temporal correspondence with diffusion latents, we further propose Relative Semantic RoPE, a relative positional encoding that maps planning tokens and latents into a shared spatial-temporal coordinate frame, enabling each latent to accurately attend to its positionally corresponding semantic cues. Experiments on benchmarks show the effectiveness of Baton both qualitatively and quantitatively.
PaperID: 3294, Poster
Authors: Suranjan De, Steinar Laenen, He Sun
Abstract: We study clustering algorithm for dynamic evolving graphs \ G_t\, in which new edges (and potentially new vertices) are added into a graph, and the underlying cluster structure of the graph gradually changes over time. By designing a hierarchical data structure for dynamically maintaining the cluster structure of \ G_t\, we prove that the cluster structure of G_T can be maintained with O(d) amortised update time and \widetildeO\left(n_T^1/d\right) amortised query time, for an arbitrary large constant d. This significantly improves the previous state-of-the-art (Laenen and Sun, ICML~'24), which requires O(1) update time and O(n_T/\log n_T) amortised query time. We further demonstrate this improvement with experiments on both synthetic and real-world datasets.
PaperID: 3295, Poster
Abstract: Vehicle trajectory generation is critical for autonomous driving simulation. However, current generative models face two coupled limitations: (i) imbalanced trajectory distributions bias them toward frequent motions, such as straight driving or stationary states, while weakening their ability to synthesize rare but safety critical motions, such as turning. (ii) Euclidean intermediate spaces provide limited geometric inductive bias for modeling vehicle pose evolution. Together, these limitations lead to mode biased generation and physically less coherent trajectories. To address this issue, we propose SEA, an \underlineSE(2)-\underlineAware vehicle trajectory generation method that realizes the conditional transport path through planar rigid body poses. At the data distribution level, SEA learns vehicle trajectory distributions from traffic data to capture individual motion characteristics. At the physical-space level, motivated by the geometry dependence of \bf optimal transport interpolation, SEA realizes intermediate transport states through an SE(2)-consistent construction, thereby preserving geometric interpretability and physical coherence during generation. Extensive experiments on large scale real world datasets show that our method outperforms state of the art approaches by generating more realistic, diverse and dynamically consistent trajectories.
PaperID: 3296, Poster
Abstract: Video fusion integrates complementary information in multiple source video streams into a single sequence for complete scene perception. Existing methods perform frame-wise fusion and lack unified spatio-temporal modeling across sequences, damaging temporal coherence across sequences. When dealing with degradations, they also ignore temporal dynamics and overload a single network with diverse degradations, leading to limited and imbalanced degradation suppression. To address these issues, we propose V2VFusion, the first unified video-to-video fusion framework that directly generates temporally coherent, high-quality fused videos from degraded multi-source inputs. By reformulating video fusion as a video-to-video conditional generation problem in latent diffusion space, our approach holistically models spatio-temporal dependencies across entire sequences and intrinsically ensures sequence-level consistency. On this basis, a text-guided degradation-conditioned ControlNet bridges high-level textual degradation descriptions with low-level visual priors, enabling interpretable and controllable guidance for structure-preserving fusion beyond purely data-driven conditioning. To handle heterogeneous and compound degradations, a state-modulated degradation-aware hierarchical mixture-of-experts module decouples expert selection into diffusion-state-aware filtering and degradation-aware routing. It prevents unreliable routing and fosters expert specialization to mitigate representation conflicts among diverse degradations. Experiments on infrared-visible, multi-exposure, and multi-focus video fusion tasks and complex degradations validate that V2VFusion surpasses existing methods in fusion quality and temporal stability.
PaperID: 3297, Poster
Authors:
Jiayi Guan, jing wu, Jianhua Wu, Jinghui Lu, Zhijian Huang, Guang Li, Xiaoshuai Hao, Hanbing Li, Yigu Ge, Tao Xu, Haiyang Sun, Bing Wang, Guang Chen, Kuiyuan Yang, Hangjun Ye, Long ChenAbstract: Vision-Language Models (VLMs) have achieved remarkable advances in general autonomous driving. Nevertheless, effectively exploiting the reasoning capability of Chain-of-Thought (CoT) to enhance the robustness and accuracy of decision and planning in long-tail scenarios still remains a challenging open problem. To this end, we propose a Reasoning-Supervised Policy Optimization (RSPO) algorithm to boost reasoning and decision-making performance under long-tail driving scenarios. Concretely, we first conduct an empirical analysis of the core challenges faced by current VLMs in autonomous driving decision and planning tasks, and formally define reasoning and decision-making tasks tailored to long-tail scenarios. On this basis, we design a curriculum-guided supervised fine-tuning strategy to enable the model to rapidly fit the mean distribution of model outputs. Furthermore, we propose a reasoning-supervised optimization algorithm, which enhances the robustness and accuracy of reasoning and decision in long-tail scenarios by guiding the optimality consistency between reasoning processes and final decisions. Finally, we construct multiple verifiable autonomous driving reasoning and decision datasets based on open-source benchmarks and conduct extensive multi-dimensional comparative experiments. Experimental results show that the proposed algorithm achieves consistent improvements in prediction accuracy across multiple datasets compared with existing methods. In particular, it achieves a 6.7% relative improvement over the strong baseline on the CODA-LM task and also yields lower trajectory-prediction errors on downstream planning tasks.
PaperID: 3298, Poster
Abstract: We study the best arm identification problem in a novel non-stationary environment that we coin Shifting Means. While classically the mean rewards of the K arms are stable in time, in Shifting Means only the gaps \boldsymbol\Delta between mean rewards are stable, while their common shift may be determined adversarially in each round. The objective of the learner is to identify the best arm with high probability while minimizing sample complexity (the fixed confidence setting). Handling shifts requires new tools: we show that algorithms employing a Generalized Likelihood Ratio Test (GLRT) stopping rule, including the popular Track-and-Stop, fail under time-varying shifts. Instead, we propose Importance Weights for Shifting Means (\mathsfISM). Assuming means bounded by U and \sigma^2-sub-Gaussian rewards, we show \mathsfISM to be \delta-correct and to enjoy a sample complexity bound of K (\sigma^2 + U^2) \Delta_\min^-2 \ln \frac1\delta. We also present a matching (up to constant factors) worst-case lower bound and evaluate our results empirically.
PaperID: 3299, Poster
Authors: GeonWoo Jeong, SEUNG H OH, 추 상범, Wonjunkim, Bokyoon Na
Abstract: Existing conditional diffusion models typically inject condition information into the reverse-time denoising process, either as model inputs or through guidance. Although effective for conditional fidelity, this design burdens the denoiser when compact continuous signals, such as user preference ratios or mixture degrees, must support both image reconstruction and condition-wise separation, often degrading sample quality. In this work, we reinterpret this limitation not as a problem of condition representation, but as a problem of condition placement. We propose the Continuous Personalized Diffusion Model (CPDM), which shifts the locus of conditioning from reverse-time signal injection to forward geometry formation by coupling bounded spinor-component coordinates with normalized image-space drift directions. The resulting condition-dependent forward drift forms condition-specific terminal geometry, from which continuous changes along the spinor-component coordinate are reflected as continuous visual transitions rather than endpoint interpolation. Experiments show that, even with a single scalar condition, CPDM achieves competitive performance compared with standard and higher-capacity reverse-conditioning baselines and supports continuous generation under compact scalar control. These results suggest forward geometry design as a practical alternative to conventional reverse-time conditioning for compact continuous personalization.
PaperID: 3300, Poster
Authors: Biswajit Banerjee, Claudia A Carreno, Anton S Petrov
Abstract: Protein Language Models (PLMs) have made remarkable progress following scaling laws established in natural language processing across sequence- and structure-based tasks, yet the potential of tokenization remains underexploited. Unlike human language, proteins preserve structure despite extensive sequence variation — a property standard tokenization strategies fundamentally fail to capture. We introduce ZEST (Zoned Encoding of Sequence Traits), an evolution-informed vocabulary derived from conserved regions of multiple sequence alignments. ZEST allows embedding domain-level biological priors directly at the tokenization stage rather than learning them implicitly through scale. ZEST natively compresses sequences to an average token length of 4 residues, enabling our model to process 4,000 residues within a standard 1024-token context window. Building on this, we present LEMON (Layered Extraction of Molecular Ordering from Nature), a compact 200M-parameter sequence-based model for detection of remote homology between protein sequences trained on a single H100 GPU for one week. Despite its modest size, LEMON outperforms state-of-the-art models ranging from 600M to 3B parameters. Our results demonstrate that evolution-informed tokenization can substitute for massive parameter scaling, opening a new direction for efficient, biologically-grounded protein representation learning. All code, model weights, and results are publicly available under the MIT license.
PaperID: 3301, Poster
Abstract: Diffusion models for structured scientific generation must produce samples satisfying hard geometric constraints imposed by physics, chemistry, or biology, yet inference in these settings is prohibitively slow, demanding hundreds to thousands of neural-function evaluations per sample. We unify eight state-of-the-art models spanning medical volumetrics, molecular conformations, protein backbone design, crystal structure prediction, and multi-view 3D scenes under a single abstraction, Constraint-Manifold Diffusion Models (CMDMs), in which the target distribution is supported on a manifold defined by an externally specified constraint map. All existing acceleration families fail on this class: quantization exhausts memory on high-dimensional volumetric operators; pruning breaks constraint fidelity; fast ODE solvers allow trajectories to drift off the constraint manifold; and feature-caching heuristics are blind to constraint geometry, inducing mode confusion in the high-noise regime. We introduce ManifoldCache, the first training-free, data-free accelerator designed from first principles for CMDMs. The key insight is that the conditional score decomposes orthogonally into a normal component, which enforces constraint satisfaction, and a tangential component, which navigates within the manifold. Exploiting this structure, we prove that the noise-schedule midpoint is a sharp safe-caching boundary: caching before it incurs provably bounded error, while caching after it guarantees a strictly positive fraction of trajectories suffer mode confusion, a gap that persists up to the boundary. We further prove that deeper network blocks admit provably larger certified cache strides within the safe phase, as a consequence of the score decomposition propagating through block Jacobians. The resulting schedule requires one integer comparison per block per step, zero calibration data, and zero training overhead. Across all eight CMDMs, ManifoldCache delivers consistent wall-clock speedups while preserving or improving generation quality on structural, perceptual, and distributional metrics, consistently outperforming several strong baselines spanning quantization, pruning, fast solvers, and feature caching.
PaperID: 3302, Poster
Abstract: Referring expression segmentation (RES) aims to produce a pixel-level mask for the object described by a free-form natural language expression, in either images (RIS) or videos (RVOS). Existing zero-shot approaches reduce the cost of mask-language annotation but still rely on auxiliary training stages or multi-model pipelines. The recently released SAM 3 introduces native text-prompted segmentation within a single foundation model. However, directly applying it to RES generates unexpected weak performance. We trace this gap to an attention sink in SAM 3's multimodal fusion encoder, where the start-of-text token absorbs on average 54.8% of each image patch's cross-attention mass, leaving content tokens with limited influence on the fused visual features. Suppressing the sink alone is insufficient, as the released attention mass spreads without direction; closing the gap requires both suppressing the sink and providing patch-dependent spatial guidance that routes attention to semantically relevant content tokens. Building on these findings, we introduce eSAM, a fully training-free framework that edits SAM~3's cross-attention without any parameter updates. With Attention Mask Editing (AME), eSAM edits the attention map, suppressing the sink while injecting a CLIPSeg-derived spatial prior that directs released attention toward relevant content tokens. Further with Value Tensor Editing (VTE), eSAM rescales the value tensors, amplifying the magnitude with which content tokens contribute to the fused output. On both RIS (RefCOCO/+/g) and RVOS (Ref-DAVIS17, Ref-YouTube-VOS, MeViS) benchmarks, eSAM achieves new state-of-the-art results among training-free methods.
Abstract: Pre-training neural operators on diverse PDE datasets has emerged as a promising paradigm for building general-purpose surrogate models in scientific machine learning. However, the of PDE solution operators make multi-PDE pre-training fundamentally difficult. Existing approaches address this mainly by enlarging model capacity, while the target solution operators themselves remain unchanged. Inspired by classical numerical analysis---where a problem-tailored reformulates a complex PDE into a numerically easier equivalent---we propose to similarly reformulate the target solution operators, turning a heterogeneous set of complex, structurally divergent operators into a simpler and better-aligned equivalents. Since different PDEs admit different simplifications, the transformation must be , so that a single neural operator can approximate the entire family jointly. We instantiate this idea as AOT-POT ( ransformer), which realizes such a transformation by expanding the hidden representation into multiple parallel streams, aggregating and redistributing them with input-dependent weights before and after each sub-layer, and mixing streams through Sinkhorn-projected doubly stochastic matrices for stable training. Together, these mechanisms reformulate the diverse, complex solution operators into a simpler equivalents that a single architecture can approximate jointly. Empirically, AOT-POT achieves state-of-the-art results on 12 PDE benchmarks with only 3% additional parameters, reducing the relative L2 error by up to 77.6% (40.9% on average). Fine-tuning further reduces L2 error by up to 92% on in-domain PDEs and 89% on out-of-domain PDEs, confirming that adaptive operator transformation is an orthogonal and effective axis for advancing PDE foundation models, beyond merely scaling model capacity.
PaperID: 3304, Poster
Abstract: Does a neuron’s shape predict whom it connects to? Peters’ rule, a canonical geo- metric principle in connectomics, predicts connectivity from spatial arbor overlap, but true wiring also depends on molecular and morphological specificity. We ask how far neuron skeletons, with identity and graph features withheld, predict directed connections in the FlyWire Drosophila connectome. Because most non- connected pairs are far apart, headline metrics are dominated by easy negatives that geometry can reject. We therefore evaluate controlled contrasts: plausible non-connections whose arbors are nearby, overlapping, or cell-type matched, but that still do not synapse. Across a controlled model ladder, from arbor overlap and hand-crafted morphology baselines to Connectoformer (a bidirectional Trans- former with pooled and pairwise cross-attention scoring), performance improves when the model gains the corresponding signal: global shape, local branch com- patibility, and tree context. Morphology and arbor co-occupancy are complemen- tary: morphology discriminates similar-looking unconnected pairs but can score pairs whose arbors never touch; arbor co-occupancy does the inverse. By suppress- ing essentially-zero-overlap pairs without boosting high-overlap pairs, a fixed one- sided arbor prior performs well in both regimes. The encoder also transfers across Drosophila individuals: zero-shot evaluation on the male CNS connectome beats arbor overlap on hard-negative discrimination. Stratified evaluation reveals a clean division of labor: geometry rules out pairs that cannot physically connect, and morphology picks the actual partners from what remains.
PaperID: 3305, Poster
Abstract: Investigating the causal structures of transformers often entails a trade-off between the granularity of the explanation and computational overhead. Existing approaches either bypass the full internal chain, potentially missing critical decision-relevant links, or struggle with the combinatorial explosion of path enumeration. To address these challenges, we propose an efficient automated path-tracing method that ensures both causal reliability and algorithmic feasibility. Through extensive verification, we find that (i) the identified paths are the primary sources for decision signals, as evidenced by a marked decrease in self-repair compared to non-path components; and (ii) transformers utilize modular, class-shared path mechanisms to process categorical information. Beyond these findings, we show that the potential implications of our method can stem from sparse manipulation, highlighting efficient applicability to downstream tasks such as model debugging and pruning.
Abstract: Reasoning has become a central capability in large language models. Recent research has shown that reasoning performance can be improved by looping an LLM’s layers in the latent dimension, resulting in looped reasoning language models. Despite promising results, few works have investigated how their internal dynamics differ from those of standard feedforward models. In this paper, we conduct a mechanistic analysis of latent states and layer behavior in looped language models, focusing in particular on how the stages of inference observed in feedforward models compare to those observed in looped ones. We analyze cyclic recurrence and show that for many of the studied models each layer in the cycle converges to a distinct fixed point; consequently, the recurrent block follows a consistent cyclic trajectory in the latent space. We provide evidence that as these fixed points are reached, attention-head behavior stabilizes, leading to constant behavior across recurrences. Empirically, we discover that recurrent blocks learn stages of inference that closely mirror those of feedforward models, repeating these stages in depth with each iteration. We study how architectural choices influence the emergence and stability of these cyclic fixed points and stages of inference, providing practical guidance for looped language model design.
PaperID: 3307, Poster
Abstract: Efficient long-context inference requires compressing the KV cache under strict memory budgets without degrading model outputs. We formulate this problem as a constrained optimization: minimize the total attention-output distortion subject to a global average bit-width budget, where each token’s KV entry can be assigned a bit-width ranging from 0 (eviction), through mixed precision, to 16 (no compression). The central challenge is to estimate each token’s contribution to attention-output distortion. We address this via a first-order perturbation analysis of the attention mechanism, which yields a closed-form decomposition of the distortion. Specifically, the compression cost of each token factorizes into (i) a token-specific risk score that captures the sensitivity of the attention output to perturbations of that entry, and (ii) a precision-dependent quantization error determined by the chosen bit-width. Building on this factorization, we solve the global allocation problem via Lagrangian relaxation, reducing it to independent per-token decisions that can be efficiently computed. Experiments on LongBench and RULER with LLaMA-3.1-8B-Instruct and Qwen3-8B show that our method achieves comparable accuracy to uniform 4-bit quantization while using an average of 2 bits per entry (2× memory reduction). At an average of 1 bit, our method maintains substantially higher accuracy than both uniform quantization and token eviction baselines, which deteriorate sharply in this regime. The advantage concentrates on retrieval-sensitive subtasks, where dispersed evidence makes uniform allocation and hard eviction especially costly.
PaperID: 3308, Poster
Abstract: We trace factual sycophancy across three post-training pipelines (OLMo 3 7B Think and Instruct, Llama 3.1 8B Instruct, Tulu 3) on AMPS and MedQuAD with a dual-track evaluation: a GPT-4o-judged generative track and a log-probability track reporting \DeltaLogOdds separately under \emphin-context rebuttals and \emphpreemptive wrong-answer assertions. Decomposing by challenge type × context reveals three patterns. (1) Challenge type: the sycophantic shift is alternative-conditioned, so simple pushback with no asserted answer produces no shift, while ethos/justification/citation challenges produce substantial positive shifts. (2) Context: every base model's sycophantic shift concentrates preemptively, with weak or defensive in-context response. On computational, IC diverges into four pipeline-specific endpoints. On medical, the same four pipelines converge with no defensive IC developing on any recipe. (3) Behavioral vs log-probability: post-training produces a \emphsurface-predictive dissociation: matched-subset behavioral flip rate drops significantly on identical items while preemptive \DeltaLogOdds on those same items grows.
PaperID: 3309, Poster
Authors: Wojciech Kotlowski, Marek Wydmuch, Krzysztof Dembczynski
Abstract: We study sequential learning for classification metrics that are non-decomposable across instances and are general functions of the confusion matrix, such as the Jaccard index and the F-measure. The learner observes instances sequentially, makes irrevocable predictions, and is evaluated by the final empirical performance metric, leading to an online-regret objective. After observing the sequence, the learner also returns a classifier for future population use, leading to a population-regret objective. This single protocol connects two central frameworks for optimizing non-decomposable metrics: Expected Test Utility (ETU) and Population Utility (PU). For smooth concave utilities of the confusion matrix, we show that the ETU benchmark is controlled, in expectation, by the population PU comparator, and that an algorithm with small online regret also yields PU guarantees by an online-to-batch conversion. We then turn to non-concave linear-fractional metrics. We devise a general stochastic Dinkelbach-root method that tracks the relevant metric parameter online and gives online and population regret of order n^-1/2, up to logarithmic factors, with explicit dependence on conditional-probability estimation error. Our empirical studies evaluate these methods on benchmark datasets.
PaperID: 3310, Poster
Abstract: Supervised fine-tuning (SFT) is computationally efficient and broadly applicable, but often shows weaker generalization than reinforcement learning (RL). A key limitation is its off-policy objective: SFT fits fixed demonstrations token by token, including targets that may be poorly aligned with the model's pretrained distribution. Recent token-reweighted SFT methods address this issue by assigning larger training weights to tokens that better align with the model's predictive distribution, using statistics such as target-token probability or entropy. However, computing these statistics from the online model being fine-tuned makes token weights trajectory-dependent, as the model's distribution rapidly departs from the pretrained model and induces self-reinforcing reweighting dynamics. We propose PriFT, Prior-support guided Fine-Tuning, which derives token weights from a frozen pretrained reference to obtain a stable reweighting signal unaffected by online fine-tuning dynamics. This signal estimates prior support: the extent to which each target token is supported by the pretrained model before task-specific adaptation. Across multiple existing weighting and selection rules, replacing online statistics with pretrained statistics consistently improves performance. We introduce two instantiations: PriFT-prob, which uses pretrained target-token probability, and PriFT-mass, which selects tokens by relative support under the pretrained distribution. Extensive experiments on mathematical reasoning, code generation, and medical question answering show that PriFT achieves state-of-the-art SFT results and provides a better initialization for subsequent RL training.
Abstract: Transformer-based scientific foundation models are increasingly deployed in high-stakes settings, but current architectures give deterministic outputs and provide limited support for calibrated predictive uncertainty. We propose Stochastic Attention, a sample average lightweight inference-time modification that randomizes attention by replacing softmax weights with normalized multinomial samples controlled by a single concentration parameter, and produces predictive ensembles without retraining. To set this parameter, we introduce a calibration objective that matches the stochastic attention output with the target, yielding an efficient univariate post-hoc tuning problem. We evaluate this mechanism on scientific foundation models for weather and time-series forecasting, as well as several regression tasks. Across benchmarks against uncertainty-aware baselines, we find that Sample Average Stochastic Attention achieves the strongest native calibration and the sharpest prediction intervals at comparable calibration, with adaptation costs nearly three orders of magnitude lower than the next-best baseline.
PaperID: 3312, Poster
Abstract: Post-training quantization (PTQ) compresses large language models (LLMs) but remains largely semantic-agnostic, relying on global activation magnitude to guide precision allocation. This overlooks the functional heterogeneity of Transformer representations, where semantically meaningful channels are sparsely activated while structurally dominant channels exhibit large but context-insensitive activations. We propose Agent-guided Semantic-aware Quantization (ASQ), a PTQ framework that conditions on semantically informative positions identified by an auxiliary LLM. A dual-pathway scoring mechanism captures both activation selectivity and attention routing toward these positions, producing channel-wise importance scores for scaling. This yields a closer approximation to a semantically conditioned objective and consistently improves low-bit (e.g., W4A16) quantization performance across LLMs. Extensive experiments show consistent gains over state-of-the-art PTQ, preserving general reasoning and even surpassing FP16 performance on several benchmarks in specialized domains, while enabling training-free, privacy-preserving self-quantization with zero inference overhead.
Abstract: We study the algorithmic task of testably learning general Massart halfspaces under the Gaussian distribution. In the testable learning setting, the aim is the design of a tester-learner pair satisfying the following properties: (1) if the tester accepts, the learner outputs a hypothesis and a certificate that it achieves near-optimal error, and (2) it is highly unlikely that the tester rejects if the data satisfies the underlying assumptions. Our main result is the first testable learning algorithm for general halfspaces with Massart noise and Gaussian marginals. The complexity of our algorithm is d^\mathrmpolylog(\min\1/\gamma, 1/\epsilon \), where \epsilon is the excess error and \gamma is the bias of the target halfspace, which qualitatively matches the known quasi-polynomial Statistical Query lower bound for the non-testable setting. The analysis of our algorithm hinges on a novel sandwiching polynomial approximation to the sign function with multiplicative error that may be of broader interest.
PaperID: 3314, Poster
Abstract: Social science panel time series exhibit multicollinearity, latent unit heterogeneity, and regime-conditional sign-changing relationships. On V-Dem democratic development data, four of four selected covariates exhibit coefficient sign reversal across regimes, driving pooled additive predictors to cancel real signal - a diagnostic we establish before introducing any nonlinear model. We present a regime-aware additive forecasting architecture that recovers this signal under a verifiable contraction certificate. Our theoretical contribution is a rate-based contraction theorem decomposing the global Lipschitz constant of a soft-gated mixture predictor into per-regime, regime-disagreement, and dwell-time components, each empirically auditable on a trained model. On V-Dem democratic development data (151 countries, 1970-2025), the architecture recovers regime-conditional signal that pooled additive models cancel (paired Diebold-Mariano p=0.016), with the recovery localized to the tails of the democracy index where a diagnostic shows largest covariate sign disagreement. The contraction certificate L_\texteq,max<1 holds in 5/5 seeds across all 9 audited configurations without explicit projection. A synthetic stress test with stronger regime contrasts shows the certificate is non-vacuous: spectral normalization on per-edge layers can fail to deliver contraction because the gating-driven term L_\phi,\texteq \cdot M dominates, and a regime-disagreement penalty is required to recover the bound.
PaperID: 3315, Poster
Abstract: The robust delivery of power is a non-negotiable prerequisite for the operational reliability of modern chips. Even a single unanticipated IR-drop violation can cause system-level crashes, making IR-drop a strict criterion in chip signoff. However, current methodologies ranging from physics-based simulators to recent learning-based surrogates remain bottlenecked by an inherent coupling between spatial topology and temporal evolution. This paradigm is intrinsically unscalable when billion-gate spatial complexity converges with the extensive workloads of full-testbench verification. In this work, we reformulate dynamic IR-drop assessment as a disentangled multimodal learning problem and introduce Layout Waveform Fusor (LWFusor), which decouples invariant spatially static chip layouts from transient switching waveforms and fuses them through physics-aligned end-to-end learning. By projecting activity sequences directly into the spatial potential domain, the proposed approach departs from traditional time-stepping bottlenecks and makes long-horizon power integrity analysis computationally efficient. Experiments on industrial-scale designs demonstrate orders-of-magnitude acceleration while maintaining high fidelity in hotspot capture, highlighting a new trajectory for power integrity sign-off in large-scale EDA workflows.
PaperID: 3316, Poster
Authors: Bingkun Liu, Paul Haider, Federico Benitez, Mihai A. Petrovici, Walter Senn
Abstract: Self-recurrence represents an important component for temporal processing in neural networks, both artificial and biological, where it appears naturally in the form of leaky integration. However, due to the complex temporal dependencies of the gradient, an efficient and effective method of online training is still lacking. Here we propose K-prop, a novel online learning method for leaky neural networks. It dramatically reduces the computational and memory cost of real time recurrent learning (RTRL) by propagating kernel gains. K-prop also outperforms existing RTRL approximations and achieves performance comparable to backpropagation through time in sequence learning, classification, and reinforcement learning tasks. We further analyze how its sole approximation affects gradient quality and identify conditions under which K-prop remains accurate and effective.
Authors: Heikichi Hayashi, Ruoran Lai, Shawn Yu, Huanxi Zhang
Abstract: Threshold decryption is increasingly used to hide economically sensitive data until a public release time. The main strategic risk is not merely premature decryption of one ciphertext, but leakage of a coalition-generated decoder that may remain useful across many future contexts. We formalize a repeated leakage game for arbitrary monotone access structures with transferable utility, one-time public accountability exposure, and an abstract reachability correspondence that records which future contexts become compromised after a first leak. Our main theorem gives an exact characterization: honest non-leakage is coalition safe if and only if, for every authorized coalition S, its exposure κ(S) dominates the maximum discounted leakage value over the contexts reachable from a single first leak by S. Two important consequences follow. In fully reusable schemes, the exact deterrence threshold equals the coalition’s continuation leakage value, yielding horizon amplification. In context-isolating schemes, including context dependent threshold decryption, the reachable set collapses to the current context, so the threshold drops to the one-shot value. We also prove a sharp separation: for a T-round family with unit per-round leakage value, reusable leakage requires Θ(T) total exposure, whereas context isolation requires only Θ(1). Additional corollaries give a geometric threshold under bounded reuse multiplicity, exact equal-stake formulas for t-out-of-n committees, a direct encrypted-mempool instantiation, and a converse lower bound showing that without a bound on reuse multiplicity no uniform O(1) exposure can suffice. The contribution is theorem-driven and deliberately orthogonal to post- decryption ordering or proposer-builder-separation equilibrium analyses
Abstract: Confidence sequences based on test martingales provide time-uniform uncertainty quantification for the mean of bounded IID observations without parametric distributional assumptions. Their practical efficiency, however, depends strongly on the choice of martingale updates, and many existing constructions do not exploit prior information about plausible data-generating distributions or mean values. We propose a Bayes-assisted framework that uses a Bayesian working predictive model to adaptively construct confidence sequences. For each candidate mean and time point, the predictive distribution selects, among valid one-step martingale factors, the update maximising predictive expected log-growth; validity is therefore preserved even when the prior or working model is misspecified. We prove that if the predictive distribution is Wasserstein-consistent, the resulting procedure is asymptotically log-optimal, matching the per-sample log-growth of an oracle procedure with access to the true distribution. We instantiate the framework using robust predictives based on Dirichlet-process mixtures and Bayesian exponentially tilted empirical likelihood. Experiments on synthetic data, sequential best-arm identification for LLM evaluation, and prediction-powered inference show that informative priors can substantially reduce confidence-sequence width and sampling effort while retaining anytime-valid coverage.
Authors:
Bihui Yu, caijun jia, Jing Chi, Xiaohan Liu, Yining Wang, Richard H Bai, Yuchen Liu, Jingxuan Wei, Junnan ZhuAbstract: Multimodal large language models (MLLMs) increasingly solve vision-centric tasks by calling external tools for visual inspection, OCR, retrieval, calculation, and multi-step reasoning. Current tool-using agents usually expose the executed tool trajectory and the final answer, but they rarely specify which tool observation supports each generated claim. We call this missing claim-level dependency structure the provenance gap. The gap makes tool use hard to verify and hard to optimize, because useful evidence, redundant exploration, and unsupported reasoning are mixed in the same trajectory. We introduce TRACER, a framework for verifiable generative provenance in multimodal tool-using agents. Instead of adding citations after generation, TRACER generates each answer sentence together with a structured provenance record that identifies the supporting tool turn, evidence unit, and semantic support relation. Its relation space contains Quotation, Compression, and Inference, covering direct reuse, faithful condensation, and grounded derivation. TRACER verifies each record through schema checking, tool-turn alignment, source authenticity, and relation rationality, and then converts verified provenance into traceability constraints and provenance-derived local credit for reinforcement learning. We further construct TRACE-Bench, a benchmark for sentence-level provenance reconstruction from coarse multimodal tool trajectories. On TRACE-Bench, simply adding tools often introduces noise. With Qwen3-VL-8B-Instruct, TRACER reaches 78.23% answer accuracy and 95.72% summary accuracy, outperforming the strongest closed-source tool-augmented baseline by 23.80 percentage points. Compared with tool-only supervised fine-tuning, it also reduces total test-set tool calls from 4,949 to 3,486. These results show that reliable multimodal tool reasoning depends on provenance-aware use of observations, not on more tool calls alone.
PaperID: 3320, Poster
Abstract: Recent neural decoding approaches achieve strong performance by training on large collections of paired neural-behavioral recordings, but such datasets are expensive, invasive, and difficult to scale. In contrast, behavioral data is abundant and easy to collect, raising the question of whether neural decoding can benefit from decoupling behavioral and neural representation learning. We introduce BeeMO, a flexible multi-session framework that enables training with arbitrary mixtures of paired neural-behavioral recordings and unpaired behavioral data. A behavior decoder learns movement dynamics directly from behavioral trajectories, while neural activity is incorporated through cross-attention layers when available, allowing behavioral and neural supervision to scale independently. By decoupling neural and behavioral data, we unlock a new dimension for scaling: unpaired behavioral data, which is substantially cheaper and easier to collect than paired neural recordings, can be used directly to improve decoding performance. Across multiple intracortical motor datasets and tasks in nonhuman primates, we show that incorporating unpaired behavioral data consistently improves decoding performance, particularly, and that incorporating task-aligned behavioral trajectories can further improve transfer. We further show that behavior-only pretraining, without any neural pretraining, outperforms single-session supervised baselines and is comparable to models pretrained on substantially larger neural datasets. Together, our results suggest that scaling behavioral data offers a practical and cost-effective path toward neural decoding models that generalizes across subjects and tasks.
PaperID: 3321, Poster
Abstract: LLM-powered agentic systems increasingly operate across complex, long-horizon task regimes, yet their growing architectural complexity has intensified system-level fragility and failure rates. The task of \emphagentic failure attribution seeks to identify the specific agent or execution step that causes failure within lengthy traces. However, most existing approaches remain constrained to passive, full-trajectory analysis, which limits their effectiveness in ultra-long and dynamically evolving trajectory attributions. To address this limitation, we introduce AgenTracer-v2, an autonomous, multi-hop failure attribution agent that reconceptualizes tracing as an active, tool-augmented diagnostic process. Concretely, we develop Tracer-Flywheel, a framework-agnostic data synthesis engine with deterministic system rollback, yielding an 8K dataset of precisely annotated failure trajectories. Leveraging this supervision, AgenTracer-v2 is trained to perform token-efficient global summarization, targeted inward inspection, and outward exploration via integrated diagnostic tools. Extensive experiments demonstrate that AgenTracer-v2 (I) consistently surpasses proprietary models such as \textscGPT-5.2 and \textscGemini-3-Pro, as well as prior specialized tracers, by up to 21.83% in attribution accuracy, (II) maintains robust performance on ultra-long trajectories exceeding 300K tokens and (III) delivering corrective feedback that measurably improves downstream agentic systems and multi-LLM reinforcement training.
PaperID: 3322, Poster
Abstract: In Euclidean online convex optimization, first-order and full-information feedback share the minimax regret (\Theta(DL\sqrtT)) over (T) rounds, as convex losses can be replaced by their linearizations. On Riemannian manifolds this reduction fails, because the corresponding first-order surrogate of a geodesically convex loss is generally not geodesically convex; we show that this failure has provable consequences. For online convex optimization on Hadamard manifolds with sectional curvature bounded below by (\kappa<0) and feasible diameter (D), we construct hard hyperbolic-space instances yielding first-order regret lower bounds at the scale ( DL\sqrt\zeta\, T, ) where ( \zeta=\sqrt|\kappa|D/\tanh(\sqrt|\kappa|D) ) is the curvature--diameter factor. Our first lower bound applies to all (possibly randomized) first-order algorithms satisfying a natural geometric-span condition, holding almost surely over the algorithm's internal randomness. Our second lower bound removes the geometric-span restriction for deterministic first-order algorithms, provided the feasible diameter satisfies (D=\Omega(\log T)); the proof rests on a new construction that confines the hidden comparator to a single horosphere in hyperbolic space. Combined with existing curvature-free full-information upper bounds, these results yield a feedback-model separation on Hadamard manifolds: full-information algorithms attain (O(DL\sqrtT)) regret, yet the additional (\sqrt\zeta) factor is unavoidable for first-order feedback whenever either the geometric-span condition holds or the algorithm is deterministic.
PaperID: 3323, Poster
Abstract: Generating high-quality and controllable camera motion is essential for AI-assisted cinematography, video synthesis, and 3D scene understanding. We introduce TKCAM, a Text- and Keyframe-conditioned CAMera-motion synthesis framework based on generative masked modeling. To overcome the instability of direct 3D regression, we formulate camera dynamics as continuous 12-DoF kinematic sequences and discretize them into hierarchical motion tokens via a Residual Vector Quantizer (RVQ). A two-stage masked transformer architecture then learns to reconstruct and refine these tokens, utilizing explicit self- and cross-attention modules for robust multi-modal conditioning. A central feature of our framework is animator-style keyframe control: users can provide free-form text prompts alongside sparse key poses, and TKCAM seamlessly inpaints temporally coherent in-between trajectories. Furthermore, to advance evaluation standards, we curate RealEstate10K-Cap, a large-scale text-camera dataset, and establish a comprehensive cross-domain benchmark with a Universal CLaTr Evaluator. Extensive experiments demonstrate that TKCAM significantly surpasses recent state-of-the-art baselines on Fréchet Inception Distance (FID), text-motion matching scores, and retrieval metrics (R@K), producing cinematographically plausible motions while enabling precise spatial guidance. Our code, models, and dataset will be open-sourced upon acceptance.
Abstract: Scaling test-time compute has emerged as a powerful mechanism for enhancing Large Language Model (LLM) performance. However, standard post-training paradigms, Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL), optimize the likelihood of individual samples under a base policy, creating a misalignment with test time procedures that rely on aggregated or filtered outputs. In this work, we propose Compute Aligned Training, which aligns training objectives with test-time strategies. By conceptualizing inference strategies as operators on the base policy, we derive new loss functions that maximize performance when said strategies are applied. We instantiate such loss functions for SFT and RL across common test time strategies. Finally, we provide empirical evidence that this training method substantially improves test time scaling over standard training.
PaperID: 3325, Poster
Abstract: Agentic forecasting is important for decision-making in dynamic environments, but it remains challenging because agents must reason from incomplete, time-limited evidence and produce calibrated probabilities before outcomes are resolved. Memory provides a natural mechanism for transferring experience from resolved forecasts to future prediction tasks. However, existing agent-memory methods are not tailored to forecasting, as they typically store past interactions, reflections, or factual associations without explicitly representing reusable predictive factors or calibration knowledge. We propose ForecastCompass(FoCo), an adaptive factor-based memory framework for agentic forecasting. FoCo organizes forecasting experience with a hierarchical forecasting-task taxonomy, enabling retrieval task-relevant forecasting knowledge. It maintains two complementary memory components: factor memory, which captures reusable predictive dimensions, and reasoning memory, which encodes probability updating, uncertainty handling, and calibration principles. Using retrospective analyses as learning signals, FoCo iteratively revises memory through a verbalized memory-revision procedure, enabling the agent to accumulate transferable forecasting knowledge over time. Experiments on Prophet Arena and FutureX with GPT-5-mini and Gemini-2.5-Flash show that FoCo improves both probabilistic accuracy and calibration.
PaperID: 3326, Poster
Abstract: Generating codes from UI screenshots has recently benefited from MLLMs, yet existing methods centered on HTML targets and transfer poorly to multi-frameworks such as React, Vue, and Angular. However, unlike HTML, which does not entail compilation issues, generating executable code in multi-framework settings is more challenging due to framework-specific differences in syntax and failure modes, as well as the need to maintain global consistency across files. To address this problem, we study executable multi-framework front-end code generation and propose MESCoder, a tree-structured self-corrective generation framework. Our method first extracts a structural scaffold from the screenshot and then constructs sub-trees and assembly relations step by step, transforming large project generation into a sequence of localized decisions. On top of this scaffold, we introduce a unified Project Agent that continuously revises its own partial project under runtime sandbox feedback and dynamically chooses between local repair and cluster level repair. For training, we adopt a two-stage strategy that first initializes the policy offline to learn basic multi-framework generation and repair priors, and then refines it with online reinforcement learning to optimize long horizon executability and final page quality. The resulting framework models multi-framework executability as a conditioned project state transition process and improves the stability and quality of complex front end generation through self correction. Experiments show that our method improves executability, while narrowing the visual-fidelity gap to HTML-based generation.
PaperID: 3327, Poster
Abstract: Recent offline multi-agent reinforcement learning (MARL) faces a persistent trade-off. Expressive generative policies can represent multi-modal coordination in the data but cannot distinguish high-value regions, while value-optimized policies exploit the learned Q-function but collapse the multi-modal structure into a single dominant mode. A single agent's mode collapse can break joint coordination, and simultaneous drift across agents can push the joint policy into unseen regions of the action space. We propose scalable coordination via optimal unified transport (SCOUT), the first offline MARL framework to combine a generative foundation model with a learned value function through test-time training. SCOUT trains two decoupled components: a flow-matching behavioral prior and a decomposed value function. At test-time, it transports behavioral samples toward high-value regions via variational gradient descent. The number of transport steps controls adaptive test-time scaling, replacing a fixed regularization coefficient. Under the individual-global-max (IGM) principle, we prove that decentralized per-agent transport is consistent with joint value improvement and monotonically recovers inter-agent correlation even when the value decomposition is approximate.
PaperID: 3328, Poster
Authors: Meongeun Kim, Taehui Lee, Soomin Park, Sung-Hee Lee
Abstract: Generating dance that reflects many aspects of music remains a fundamental challenge in dance motion synthesis. Existing music-conditioned approaches formulate the task as beat-level matching, which captures local synchrony but misses the compositional structure of rhythm, where movement is organized into global beat units. We argue that musical dance arises from the alignment of \emphperiodicity between motion and music, not from instantaneous beat coincidence. To address this, we propose PhaseDance, a phase-conditioned framework that represents dance as a quasi-musical signal on periodic phase manifolds, enabling unified dance generation across multiple tasks. Alongside this, we ground text descriptions in fifteen music-correlated metrics derived from Laban Movement Analysis to capture the qualitative richness of dance, articulating structured dimensions of motion — effort, shape, and dynamics — that coarse genre labels cannot convey. Experiments show that PhaseDance produces choreography with stronger rhythmic coherence and richer expressive understanding than existing models.
PaperID: 3329, Poster
Authors: Daoyuan Li, Zuyuan Yang, Hao Yang, Jiawen Kang
Abstract: Federated Multi-View Learning (FedMVL) enables multiple clients with heterogeneous data views to collaboratively train global models without sharing private data. Existing FedMVL algorithms often overlook the reliability of local predictions, as both uncertainty quantification and calibration remain challenging in privacy-preserving federated settings. Heterogeneous sensing devices yield varying view qualities, but view-isolated training further exacerbates inherent bias of local model, inducing unreliable outputs. Without proper calibration, these artifacts cause spurious cross-view conflicts, ultimately undermining the accuracy and reliability of the aggregated global model. To address these issues, we propose a Reliable Federated Multi-View Learning with evidence calibration (FedRMVL-CAL), a FedMVL framework with adaptive evidence fusion. The local calibration scheme harmonizes uncertainty across views, while the fusion mechanism adaptively reconciles conflicting opinions, enhancing robustness in global inference. Theoretical analysis and extensive experiments demonstrate that FedRMVL-CAL achieves superior accuracy and reliability compared to existing approaches, ensuring trustworthy global predictions across heterogeneous views.
PaperID: 3330, Poster
Abstract: Single-cell multi-omics sequencing has greatly advanced the characterization of cellular heterogeneity by jointly profiling multiple molecular modalities. However, since different modalities exhibit varying discriminative signals across cell types, existing integration strategies may dilute or even obscure the signals carried by more informative modalities when fused with less informative ones, thereby hindering accurate multi-omics clustering. To address this limitation, this work presents a novel adaptive fusion framework for unsupervised clustering of single-cell multi-omics data. The core of our approach is an attention-based fusion module that dynamically assigns modality-specific weights within a shared latent space, allowing the model to naturally focus on the most informative signals from each modality. Extensive experiments on eight real-world multi-omics datasets encompassing gene expression, chromatin accessibility, and protein modalities demonstrate that our method achieves the state-of-the-art single-cell clustering performance over fourteen competitive baselines. The code will be released publicly upon acceptance.
PaperID: 3331, Poster
Authors:
Austin Cheng, Marcel Müller, Andreas Burger, Yeonghun Kang, Tsz W Ko, Marta Skreta, Luka Mucko, Ella Miray Rajaonson, Cher-Tian Ser, Jérôme F Gonthier, Alex Zook, Varinia Bernales, Alan Aspuru-GuzikAbstract: Generative models play a rapidly growing role in the design and simulation of molecules. The most widely used applications of generative models are in large language models (LLMs), owing to the natural flexibility of language for specifying and answering queries. However, generalist LLMs do not effectively handle large sets of 3D coordinates, whereas specialist generative models of 3D molecules cannot reach the controllability afforded by language. To bridge these capabilities, we propose Nahual, a decoder-only autoregressive diffusion model for any sequence of text and 3D molecules. To enable autoregression on continuous coordinates to scale to long sequences, we propose next-token diffusion, whereby a standard causal transformer backbone is grounded in the clean current sequence while denoising the next sequence. We train Nahual on a curated set of twelve 3D chemistry tasks ranging from microsolvation to molecular conformer search to adsorption on metal surfaces, with associated evaluation metrics and baselines. The result is a single set of end-to-end-trained weights that simultaneously performs all tasks and surpasses state-of-the-art models specialized for crystal structure prediction and property-conditional molecule generation. Our results demonstrate the viability of scaling generalist decoder-only models for native multimodal understanding and generation across language and chemistry.
PaperID: 3332, Poster
Abstract: Analogical reasoning is widely considered a fundamental component of human intelligence. Few benchmarks are designed to explicitly test far analogical reasoning—reasoning between domains that are superficially dissimilar yet share deep relational or structural similarities. In this work, we introduce an analogical reasoning benchmark called ARCANA. The tasks of ARCANA require performing far analogical mappings from real-world objects or physical phenomena to abstract grid-based transformations, different from the well-known ARC-AGI benchmark which primarily emphasized innate human priors. We validate ARCANA through a human study, demonstrating that humans solve these tasks reliably. In contrast, evaluations across a range of large language models show a distinct performance gap compared to humans, revealing systematic failure modes that remain challenging for current models. To address this gap, we propose \textscConceptSIR, a reasoning framework based on concept slippage, introspection, and reflection. Experiments demonstrate that \textscConceptSIR empowers standard LLMs to solve ARCANA tasks that were previously unsolvable under baseline evaluation. Overall, our results suggest that far analogical reasoning remains an open problem for contemporary LLMs. All the materials will be available at https://sites.google.com/view/arcanacogtest.
Abstract: Chunked prefill has become a widely adopted serving strategy for long-context large language models, but efficient attention computation in this regime remains challenging. Existing sparse attention methods are primarily designed for one-shot prefill and do not translate efficiently to chunked prefill: block-sparse kernels lose efficiency when the query length is limited by the chunk size, while fine-grained pattern search becomes costly when repeated over the accumulated KV cache at every chunk. QUOKA, a recent method that directly targets chunked prefill, avoids sparse-kernel overhead but relies on query-subsampled, token-level KV selection, which can miss query-specific KV entries and introduce explicit KV-copy overhead. To address these limitations, we propose CompactAttention, a chunked-prefill attention mechanism based on Block-Union KV Selection. CompactAttention treats 2D block-sparse masks as KV-selection signals rather than direct sparse-kernel execution plans, and converts them into GQA-aware per-group KV block tables through Q-block union and intra-group union. This construction produces the minimal block tables that preserve all KV blocks selected by the input masks under paged execution constraints, enabling selected KV blocks to be accessed in place without explicit KV compaction. On LLaMA-3.1-8B-Instruct, CompactAttention maintains accuracy close to dense attention on the RULER benchmark while delivering up to 2.72× attention speedup at 128K context length under chunked prefill.
Abstract: Post-hoc calibration aligns a classifier's predicted confidences with its empirical accuracy without retraining. An ideal calibrator should correct nonlinear miscalibration, scale gracefully to large label spaces, and preserve the original predictions; existing methods typically violate at least one of these properties---temperature scaling lacks expressivity, more flexible parametric alternatives introduce parameters that grow with the number of classes C, and other expressive methods do not preserve the rank ordering of class scores and may alter the predicted class. We propose Invertible Logits Transformation (InvLT), which applies a learned scalar MLP f:\mathbbR\to\mathbbR element-wise to the pre-softmax logits. Sharing f across all logit dimensions makes the parameter count independent of C. Monotonicity of f---and hence preservation of the argmax prediction---is softly encouraged via a paired inverse network rather than enforced through the numerical integration required by prior monotone calibrators; this avoids their computational overhead while empirically preserving the original classification accuracy in every setting we evaluate. Across standard image classification benchmarks and a range of architectures, InvLT consistently outperforms a broad set of post-hoc baselines on standard calibration metrics.
Abstract: As generative AI models are increasingly used to simulate real-world systems, quantifying their “sim-to-real” gap is critical. We study this gap across a population of input settings, called scenarios, e.g. survey questions or operating conditions. The population-level quantities defining the real-world and simulated system are only observed through finite, often heterogeneous, samples. Consequently, the population-level discrepancy cannot be directly computed, and standard predictive inference methods that target observable outputs are ill-suited. We propose a model-agnostic framework that constructs confidence sets for the latent parameters, forms a conservative proxy for the sim-to-real discrepancy, and estimates its quantile function across scenarios. The resulting calibrated risk profile allows for inference on a new scenario, tail-risk summaries such as Conditional Value-at-Risk (CVaR), and principled comparisons across simulators. Our method applies to a range of output spaces, including categorical survey responses and continuous outcomes. We demonstrate its utility by evaluating the alignment of four major LLMs with human populations on the WorldValueBench dataset.
PaperID: 3336, Poster
Abstract: Free-form language generation poses a challenge for uncertainty quantification because responses do not belong to a fixed global label space. Semantic-entropy methods address this issue by sampling responses, clustering them by meaning, and estimating uncertainty over prompt-specific semantic classes. However, they mainly measure semantic dispersion and do not provide a principled decomposition into aleatoric and epistemic components, which in our setting correspond to ambiguity among plausible semantic meanings and insufficiency of model support, respectively. Moreover, cluster frequencies depend on the sampling budget and should not be treated as evidential strength. We propose a prompt-level evidential probing framework for semantic uncertainty decomposition in LLM generation. For each prompt, sampled responses define a dynamic semantic frame and provide initial support for the corresponding semantic classes. A lightweight evidence-retention probe learns how much of this support should be retained as reliable semantic evidence. The retained evidence induces a subjective opinion in which dissonance captures conflict among supported semantic alternatives and vacuity captures insufficient retained evidence. To train the probe, we introduce Augmented Semantic Uncertainty Cross-Entropy (AS-UCE), an LLM-specific evidential objective that augments the prompt-specific semantic frame with an unsupported class, allowing unreliable or hallucinated support to be routed away from observed semantic clusters during training. Experiments across four LLM backbones show that our framework consistently improves task-relevant AUROC and AUPR, with average AUROC gains of 7.4% for OOD detection and 4% for ambiguity detection over the competitive baselines, while requiring fewer sampled responses than perturbation-based decomposition methods.
Abstract: Implicit neural representations (INRs) have emerged as powerful tools for encoding signals, yet dominant MLP-based designs often suffer from slow convergence, overfitting to noise, and poor extrapolation. We introduce FUTON (Fourier Tensor Network), which models signals as generalized Fourier series whose coefficients are parameterized by a low-rank tensor decomposition. FUTON implicitly expresses signals as weighted combinations of orthonormal, separable basis functions, combining complementary inductive biases: Fourier bases capture smoothness and periodicity, while the low-rank parameterization enforces low-dimensional spectral structure. We provide theoretical guarantees through a universal approximation theorem and derive an inference algorithm with empirically linear complexity in both the spectral resolution and the input dimension. On image and volume representation, FUTON consistently outperforms state-of-the-art MLP-based INRs while training 2–5× faster.
PaperID: 3338, Poster
Abstract: What limits KV-cache compression at extreme bit-rates? We argue it is not the choice of compression scheme but how its budget is allocated across attention heads. Existing methods apply rank and bit-width uniformly, ignoring that each head has a different optimal mix of rank truncation and quantization. We show that co-optimizing rank and bit-width per head — using only standard low-rank projection and scalar quantization — dominates uniform allocation, with the largest gains at low bit-rates. Our method, KV-COBRA (Co-Optimized Bit-Rank Allocation), formalizes this as a resource-allocation problem: it balances rank-truncation loss against quantization loss within each head, then redistributes budget across heads to minimize total distortion. A fused Hadamard rotation equalizes per-channel variance, and reordering the SVD basis by attention-KL importance makes the solver query-aware. The same allocator extends to joint K+V compression. On perplexity, zero-shot, and long-context benchmarks from 0.5 to 4 bits per dimension (bpd), KV-COBRA shows the smallest accuracy degradation among evaluated methods at low bpd, with no per-token overhead.
Authors: Junren Chen, Yun Yang
Abstract: Clustering is a fundamental problem in statistics and machine learning. We propose the first one bit clustering method for two-component sub Gaussian mixture models. The method uses only one bit per entry of each sample obtained via a dithered quantizer. Under a mild non-spikiness condition on the cluster centers, we show that a variant of Lloyd’s algorithm achieves a misclassification rate that decays exponentially with a signal to noise ratio comparable to that in the unquantized setting. This result further implies exact recovery under an explicit separation condition, which exceeds the optimal threshold for unquantized data by only a logarithmic factor. When the dimension p is sufficiently large, the non-spikiness condition can be enforced by applying a random rotation using a Haar distributed matrix prior to quantization. In particular, it holds with high probability when p \gtrsim 1 for partial recovery and p \gtrsim \log n \log\log n for exact recovery, where n is the sample size. We also establish a minimax lower bound, showing that the misclassification rate and separation condition are in general optimal up to constants. Numerical results are provided to corroborate the theory and demonstrate the efficacy of the proposed method.
Abstract: E-commerce image search often takes a cropped image as the query, while each candidate is represented by full item images and structured text. This image-to-multimodal retrieval setting presents two asymmetries: a modality disparity -- a visual query must match image--text items, and a granularity disparity -- a cropped query must be compared with full images containing background context and possible distractors. Detection-based pipelines handle the granularity disparity through explicit localization but incur extra cost and error propagation, whereas CLIP-style encoders avoid detection, but are vulnerable to backgrounds or irrelevant items. To address these limitations, we propose TIGER-FG, a text-guided implicit fine-grained grounding framework for image-to-multimodal e-commerce retrieval. TIGER-FG uses item text as semantic guidance to produce target-focused item representations without object detection for retrieval. We further introduce dual distillation objectives that preserve target-region spatial consistency and query--item similarity structure, yielding more stable and discriminative multimodal representations. In addition, we construct ECom-RF-IMMR, a realistic benchmark suite with a 10M-pair training set and two evaluation benchmarks covering standard and cluttered item layouts. TIGER-FG improves Recall@1 over the strongest baseline by 6.1 and 34.4 percentage points on the two evaluation benchmarks, respectively, with only 85.7M query-side parameters and 256-dim embeddings. Results on public e-commerce benchmarks further demonstrate its generalization to noisy and one-to-many retrieval scenarios. Code and data will be released.
PaperID: 3341, Poster
Abstract: Sampling high-dimensional Boltzmann distributions is fundamental yet challenging. Classical MCMC methods incur prohibitive computational costs, while modern generative approaches such as diffusion models and Schrödinger bridges struggle to incorporate informative physical priors. GFlowNets offer a promising alternative, yet existing continuous variants suffer from training instability. In this paper, we propose Physical Prior Lifted continuous GFlowNets (PPL), a framework that encodes physical prior information through a reference space equipped with a reference distribution and a mapping to the original configuration space. The sampling process is learned via a Reference-Weighted Trajectory Balance (RWTB) loss, yielding stable training dynamics, efficient learning, and rigorous theoretical guarantees. Empirically, PPL achieves state-of-the-art performance across five benchmarks spanning synthetic, molecular, and protein systems. On the synthetic LJ-55, we obtain significantly lower error than the diffusion-based ASBS with ~36× speedup; Notably, on the 267-dimensional Chignolin protein, PPL attains near-ground-truth sampling fidelity.
PaperID: 3342, Poster
Abstract: Embodied Visual Tracking (EVT) aims to autonomously control an agent to follow a target in 3D space, which is critical for robotic navigation and human-robot interaction. However, robust long-term EVT in real-world scenarios faces two key bottlenecks: frequent tracking failures caused by occlusions and distractors, and the lack of a long-term benchmark for model training and evaluation. To address these, we propose a systematic solution. First, we develop LT-VLA, a lightweight VLA model featuring a Target Memory module to maintain consistent target identification and an Adaptive Execution module to adjust tracking actions based on observation reliability. Second, we construct LTEVT, a unified long-term EVT benchmark comprising 11 high-fidelity indoor and outdoor environments, enriched with over 5,000 assets and 325 human targets to simulate occlusions and distractors. Along with the benchmark, we release the LTEVT-4500K dataset for large-scale training and a comprehensive evaluation toolkit. Extensive experiments demonstrate that our 0.6B model achieves state-of-the-art performance comparable to 7B-scale models while maintaining superior efficiency. Notably, LT-VLA achieves 62.4% average Success Rate on EVT-Bench and runs at 20 FPS on a Unitree G1 robot, delivering a practical and robust solution for real-world deployment.
Abstract: Robust and accurate neural decoders are integral to the development of neurotechnological applications such as brain-computer interfaces and closed-loop experiments. Recent work has shown that tokenizing neural data at the resolution of individual spikes facilitates multi-session pretraining and delivers state-of-the-art decoding performance. However, these spike-based models are currently restricted to supervised learning (SL) regimes, limiting both pretraining and finetuning to datasets with paired behavioural labels. To address this, we introduce MOJO (Masked autOencoder-based JOint training), a training framework designed for spike-tokenizing models that jointly leverages self-supervised learning (SSL) via masked autoencoding and SL objectives for better decoding performance and interpretability. We evaluate MOJO on three intracortical spiking datasets—monkey motor cortex across various reaching tasks, multi-regional mouse recordings during vision and cognitive decision making tasks—demonstrating superior performance to purely SL-trained models. This improvement is especially notable when training with limited labelled data, specifically in the few-shot finetuning regime where only a small amount of labelled data is available to adapt a model to a new recording session. The incorporation of SSL also yields more interpretable neuronal representations, improving performance on analyses such as brain region classification and spike-statistics prediction despite the lack of explicit optimization for these tasks. We then show that MOJO generalizes beyond spiking data using human electrocorticography (ECoG) during speech articulation. We show that MOJO continues to outperform purely SL-trained models and achieves performance comparable to neuro-foundation models (NFMs) built specifically for continuous signals. Overall, augmenting spike-tokenizing models with SSL improves their performance in label-impoverished settings and enables the use of unlabelled data across various tasks and species, and while generalizing to other neural modalities. These results suggest a path towards more flexible and scalable data usage when training NFMs.
PaperID: 3344, Poster
Abstract: We study high-dimensional observations Z_i \in \mathbbR^p satisfying that E[Z_i] = X(t_i), where X(t) parameterizes a manifold on 0 \le t < 2\pi, with the goal of recovering the latent phases t_i; a motivating example is estimating object-viewing angles from photographs. Spectral seriation is a standard phase-recovery method: given a similarity matrix S with degree matrix D, it uses the leading eigenvectors of the normalized Laplacian D^-1/2 S D^-1/2, whose continuum limit is linked to the manifold Laplacian, to estimate t_i. Recent applications show that generalized Laplacians involving D^-\alpha can outperform the normalized Laplacian on graphs with heterogeneous degrees, raising the central question of whether the same advantage holds for spectral seriation. We address such question by introducing and comparing three principled generalizations drawn from manifold learning and spectral seriation: a shifted generalized Laplacian, a graph-difference Laplacian, and an extra-normalized generalized Laplacian. The extra-normalized generalized Laplacian is the most robust to \alpha and achieves the highest recovery accuracy, identifying extra normalization as the key stabilizing step. We establish a consistency result unveiling that eigenspace perturbation controls the estimation error up to rotation and reflection, giving a theoretical basis for the method. Experiments on noisy closed curves confirm more accurate phase recovery under a suggested choice of \alpha; on image data, the method recovers photo angles with the best accuracy. Our results demonstrate that extra normalization is a key component for stable generalized spectral seriation in closed-curve recovery problems.
Abstract: Prompt injection is the most critical vulnerability in deployed AI agents. Despite recent progress, we show that the prevailing defense paradigm (data-instruction separation) both fails to detect attacks that operate through contextual manipulation and degrades contextually appropriate behavior. We then recast prompt injection via the lens of Contextual Integrity (CI), a privacy theory that judges information flow compliance with contextual norms. This explains types of attacks that current defenses attempt to patch and predict advanced ones future agents will face. We develop unique benign and attack scenarios that force an agent to violate the norms by (1) misrepresenting the flow, (2) manipulating norms, or (3) mixing multiple flows. This reframing suggests an impossibility result: an adversary can always construct a context under which a blocked flow appears legitimate, or a defender who tightens norms will block genuinely legitimate flows. Our findings suggest that current research addresses a shrinking fraction of future attack surfaces. Instead, through CI, we offer a principled framework for evaluating context-sensitive failures, and designing CI-aware alignment for the frontier autonomous agents.
PaperID: 3346, Poster
Abstract: Path-integration-trained recurrent neural networks (RNNs) can reproduce grid-like and head-direction-like representations, but most existing models rely on low-dimensional velocity inputs and do not capture the hierarchical organisation of the brain's self-localisation circuit. We introduce a hierarchical recurrent neural network (HRNN) that separately processes angular velocity and linear speed to perform simultaneous angular and positional path integration. Trained on these tasks, the HRNN generalises to longer timescales than seen during training and develops neural response properties resembling all the key cell types: angular velocity cells, head direction cells, speed cells, conjunctive and pure grid cells. The learned circuit also recovers key motifs of the navigation system, including attractor dynamics and asymmetric connectivity patterns. Perturbation analyses reveal functional information flow across the hierarchy, consistent with experimental observations. Finally, introducing theta-paced mutual inhibition between two learned angular-velocity subpopulations enables the HRNN to reproduce the bidirectional theta sweeps observed in empirical data. These results establish hierarchical task-driven RNNs as a biologically interpretable framework for linking neural representations, circuit motifs, and population dynamics in spatial navigation, while enabling targeted in silico perturbations and the generation of experimentally testable predictions.
PaperID: 3347, Poster
Abstract: Recently, the Gromov-Wasserstein Optimal Transport (GWOT) problem has attracted the special attention of the ML community. In this problem, given two distributions supported on two (possibly different) spaces, one has to find the most isometric map between them. In the discrete variant of GWOT, the task is to learn an assignment between given discrete sets of points. In the more advanced continuous formulation, one aims at recovering a parametric mapping between unknown continuous distributions based on i.i.d. samples derived from them. The clear geometrical intuition behind the GWOT makes it a natural choice for several practical use cases, giving rise to a number of proposed solvers. Some of them claim to solve the continuous version of the problem. At the same time, GWOT is notoriously hard, both theoretically and numerically. Moreover, all existing continuous GWOT solvers still heavily rely on discrete techniques. Natural questions arise: to what extent existing methods unravel GWOT problem, what difficulties they encounter, and under which conditions they are successful. Our benchmark paper is an attempt to answer these questions. We specifically focus on the continuous GWOT as the most interesting and debatable setup. We crash-test existing continuous GWOT approaches on different scenarios, carefully record and analyze the obtained results, and identify issues. Our findings experimentally testify that the scientific community is still missing a reliable continuous GWOT solver, which necessitates further research efforts. As the first step in this direction, we propose a new continuous GWOT method which does not rely on discrete techniques and partially solves some of the problems of the competitors.
PaperID: 3348, Poster
Abstract: Clustered Federated Learning (CFL) mitigates statistical heterogeneity by grouping clients with similar data distributions. However, existing CFL frameworks predominantly rely on proxy metrics, such as gradient or data similarities, which collapse fine-grained class-level similarity into macroscopic client-level scores. This coarse-grained aggregation inherently masks critical class-wise conflicts, inadvertently exposing the federation to severe negative transfer. In this paper, we propose a fundamental paradigm shift from heuristic similarity to , a novel utility-driven CFL framework that explicitly quantifies class-wise marginal utilities to capture the true generalization impact of local updates. Guided by this attribution mechanism, FedCap employs a conflict-aware clustering algorithm to systematically eliminate intra-cluster negative transfer, while orchestrating safe, supply-demand-guided representation sharing across clusters. Crucially, we provide rigorous theoretical guarantees demonstrating the robustness of our underlying attribution mechanism against four major types of distribution shifts. Extensive experiments across diverse heterogeneity settings show that FedCap consistently achieves state-of-the-art accuracy. The code is available for anonymous access at https://anonymous.4open.science/r/FedCap-605E.
PaperID: 3349, Poster
Abstract: Three-dimensional (3D) anomaly detection underpins quality assurance in advanced manufacturing and engineering, capturing subtle defects and deviations that 2D inspection struggles to resolve due to occlusion, viewpoint, and appearance confounds. Existing 3D anomaly detection methods either rely on reconstruction errors or stored normal representations to detect anomalies. While both paradigms have driven substantial progress, they share a fundamental limitation: both operate entirely in the visual domain, missing the rich semantics encoded in natural language, which also typically constrains them to per-category models. We propose FACETS, among the first frameworks that leverage a 3D vision--language model for 3D anomaly detection, opening a new paradigm beyond reconstruction- and memory-bank-based methods. The key idea is to explicitly retain native point-level features that enable reasoning at both patch and point granularities, and further leverage language grounding and cross-granularity geometric modeling along two complementary axes, linguistic semantics and geometric saliency, so that coarse semantic cues and fine-grained geometric details jointly support anomaly detection. FACETS enables unified multi-category anomaly detection, avoiding the per-category models used by most prior methods. Notably, we provide a mathematical analysis of our loss function, offering valuable insights into the substantial improvement FACETS achieves in anomaly localization. Extensive experiments on popular benchmarks reveal that FACETS establishes new SOTA performance across datasets and metrics for 3D anomaly detection, substantially and consistently outperforming existing methods. Code is provided as supplementary material.
PaperID: 3350, Poster
Abstract: Model-based offline-to-online reinforcement learning (RL) enables sample-efficient adaptation by leveraging offline pre-training for online fine-tuning. However, the distribution shifts between offline and online stages often hinder fine-tuning performance. Many existing methods approach this problem by adjusting the optimism-pessimism trade-off via a single-objective formulation, requiring costly online bi-level optimization. We identify this trade-off during offline training as a key challenge: optimistic policies generalize better to novel tasks by exploring out-of-distribution states and actions, while pessimistic policies remain constrained to the offline data distribution and excel on similar tasks. To address this challenge, we propose a bi-objective formulation that captures this trade-off, yielding a Pareto policy pool during offline training. These policies enable flexible selection for various online tasks. To generate the pool, we introduce Multi-Objective Soft Actor-critIC (MOSAIC), which solves bi-objective problems and constructs diverse Pareto policies. After offline training, a contextual bandit algorithm hierarchically selects the most suitable policy for fine-tuning at each online interaction step. Empirically, our pipeline, ), achieves state-of-the-art performance on a suite of continuous-control offline-to-online benchmarks, particularly under shifted online tasks. Comprehensive ablations further clarify the roles of the policy pool and the online selection mechanism, as well as the robustness of the overall framework.
PaperID: 3351, Poster
Authors: jaekwon lee, Jaekoo Lee
Abstract: Millimeter-wave radar point clouds provide a visually privacy-preserving representation for human motion analysis, but object-related interactions remain difficult to describe because the object evidence that defines the interaction is not always separable or consistently observable in sparse radar returns. We address this problem without direct object detection. Instead, we infer hidden interaction semantics from the temporal structure of observation-conditioned corrections to predicted body motion, motivated by the view that object affordances shape how the body moves. To this end, we propose HIRaD, a radar semantic sensing framework that organizes sparse radar returns into a predictive node-structured latent state of local motion carriers. It represents interaction-relevant cues as corrective residual trajectories between a motion-continuation prior and an observation-conditioned posterior, and compresses these trajectories into a language-aligned semantic bottleneck. A structured prefix conditions a frozen autoregressive language model on both the semantic bottleneck and node-level motion summaries, separating compact interaction-level cues from node-level motion evidence. Experiments show that our proposed method improves radar captioning and interaction understanding under weak object observability over strong baselines, while ablations validate the importance of radar-conditioned prefixing, latent state representation, and residual trajectory modeling.
Abstract: Group distributionally robust optimization (GDRO) aims to learn models that perform well across m distributions simultaneously. Existing methods either (i) query m samples per round, which imposes strict implementation requirements, or (ii) query only a single sample per round, which requires many rounds to converge and therefore leads to long runtime. To avoid these limitations, we introduce a novel setting called \emphflexible sample queries (FSQ), which allows the number of samples that can be queried to vary across rounds. Under the FSQ setting, we cast GDRO as a two-player game: one performs non-oblivious online convex optimization, while the other tackles a non-oblivious prediction with limited advice (PLA) problem. In this game, the first player updates its decision using follow-the-regularized-leader with stochastic gradients. For the second player, we develop a new PLA method equipped with an adaptive exploration strategy that can leverage an arbitrary per-round sample size. We then establish the first high-probability regret bound for non-oblivious PLA. By integrating these components, we develop an anytime GDRO algorithm that supports FSQ, and derive a high-probability optimization error bound of O\left(\frac1t\sqrt\sum_j=1^t \fracmr_j\log m\right), where r_j is the sample size at round j. This result demonstrates that the optimization error decreases as the per-round sample sizes increase, and implies the same near-optimal sample complexity of O(m\log (m)/\epsilon^2) for any fixed sample size r\in[m].
PaperID: 3353, Poster
Abstract: Explicitly modeling periodicity has proven to be an efficient approach for time series forecasting, where the residual, obtained after removing the periodic component from the time series, is typically treated as unstructured noise. However, in this work, we show that such residuals still exhibit structured, phase-dependent deviations, as supported by both empirical observations and theoretical analysis. This suggests that residuals should not be treated merely as unstructured noise, but instead as phase-conditioned signals whose variations arise in both direction and magnitude. Motivated by this perspective and inspired by phase and amplitude modulation in classical signal processing, we propose Paloma, a simple yet effective framework for modeling residual deviations. Specifically, Paloma comprises two complementary modules: a Phase Rotation module, which performs phase-conditioned rotations to capture directional changes in the feature space, and an Amplitude Scaling module, which applies phase-conditioned affine transformations to model magnitude variations. We further abstract the two modules into a general-purpose plugin, termed the Paloma technique, which can be readily integrated into existing forecasting architectures to modulate input sequences in a phase-aware manner, enabling backbone models to capture phase-dependent variations in both direction and magnitude. Extensive experiments on twelve datasets demonstrate that Paloma achieves state-of-the-art and robust forecasting performance. Additional results show that the Paloma technique consistently improves a wide range of baseline models, validating its effectiveness as a general-purpose component.
PaperID: 3354, Poster
Authors:
Changyuan Chen, Jianyu Xiang, Jiasheng Luo, Ziye Wang, NALLAPPAN GUNASEKARANAbstract: Off-policy reinforcement learning (RL) for large language models bottlenecks at the rollout stage: every parameter update invalidates the key-value (KV) cache of in-flight agent trajectories, forcing either expensive recomputation or staleness that destabilizes training. Recent advances such as M2PO and asynchronous RLHF tolerate token-level staleness by reweighting the policy gradient with second-moment importance corrections, but they implicitly assume that a stale rollout is sampled from a well-defined behavior policy \pi_\theta_\textold. We show that this assumption is silently violated by every modern partial-rollout system: once a stale KV cache is reused under updated parameters \theta_\textnew, the resulting token distribution is a hybrid \pi_\texthyb whose attention projections consume stale keys and values while its MLP, output, and gating projections use the new weights. The token-level importance weight w_t = \pi_\theta_\textnew(y_t \mid x_
PaperID: 3355, Poster
Abstract: Zero-shot pre-trained diffusion models can provide a training-free source of synthetic image generation for Exemplar-Free Class Incremental Learning (EFCIL). However, same-class generation is limited by the distribution mismatch with real data. Parameter-allocation-based EFCIL can minimize the forgetting in model parameters, yet the performance degrades because task-id prediction fails when task-specific feature extractors overcompress and converge to shortcut-dominant representations. We propose \emphClass-Mixed Diffusion Augmentation (CMDA) that uses diffusion models not only as past-task data generators but also as controlled generators that challenge the shortcut reliance. We sample candidate tokens from the vocabulary and select \emphshortcut-breaking tokens via classifier-guided filtering, then treat the selected tokens as auxiliary classes during training. We provide theoretical analysis linking shortcut sharing to an intrinsic lower bound on task-id error and show that minimum-CE selection recovers shortcut-breaking tokens. Experiments on Split-CIFAR100 and Split-ImageNet show consistent gains over EFCIL baselines, with modest overhead over vanilla diffusion augmentation and far lower cost than diffusion fine-tuning or gradient-guided sampling.
PaperID: 3356, Poster
Abstract: While test-time fine-tuning is beneficial in cross-domain few-shot classification, the need for multiple backpropagation steps can be prohibitively expensive in resource-constrained environments. We propose HyperFlow, a gradient-free test-time adaptation method that amortizes fine-tuning dynamics into a lightweight conditional drift network. During offline training, HyperFlow learns from fine-tuning trajectories collected on meta-training tasks. Once trained, it adapts a selected PEFT parameter subspace for a new task by numerical ODE solving, requiring only forward passes of the drift network and no test-time backpropagation through the target model. In experiments on Meta-Dataset and CD-FSL benchmarks, our method improves out-of-domain performance over the direct transfer approach while using only 6-14% of peak memory and about 1% of the FLOPs of standard fine-tuning, positioning HyperFlow as an accuracy–efficiency trade-off between direct transfer and fine-tuning.
PaperID: 3357, Poster
Abstract: Large language models (LLMs) have achieved remarkable success while introducing critical energy bottlenecks that challenge sustainable deployment. Spiking neural networks (SNNs) provide a promising approach for energy-efficient spiking LLMs via ANN-to-SNN (A2S) conversion. Among various spike coding methods, time-to-first-spike (TTFS) coding is particularly appealing as it conveys information with a single spike, further reducing energy consumption. However, existing TTFS-based A2S conversion relies on continuous-time assumptions, requiring prohibitively large latencies (e.g., 4096 time steps) to approximate ANN's continuous values. This dependency leads to unacceptable inference delays in LLMs, posing significant challenges for developing practical temporal-coding spiking LLMs. In this paper, we propose a discrete analysis framework for TTFS-based LLMs, reformulating TTFS-based A2S conversion from continuous lossless conversion into discrete error-aware representation. Building on this analysis, we establish an equivalence between discrete TTFS-based SNNs and static quantized ANNs, allowing TTFS conversion to be handled through quantization techniques. Motivated by this equivalence, we introduce Quantization-Consistent ANN-to-SNN (QC-A2S) conversion, which combines static quantization with discretization-compatible TTFS neurons to achieve latency-efficient and high-performance temporal-coding spiking LLMs. Extensive evaluations on representative LLMs across multiple benchmarks demonstrate competitive performance with substantially reduced inference latency.
Abstract: Vision-language models (VLMs) have demonstrated remarkable capabilities in bridging visual perception and natural language understanding, enabling a wide range of multimodal reasoning tasks. However, they often produce object hallucinations, describing content absent from the input image, which limits their reliability and interpretability. To address this limitation, we propose Dual-Pathway Circuit Analysis, a framework that identifies and characterizes hallucination-related circuits in VLMs for mechanistic understanding and causal probing. We first apply activation patching across five architecturally diverse VLMs to identify a visual grounding pathway that supports correct predictions and a hallucination pathway that drives erroneous outputs. We then introduce Conditional Pathway Analysis (CPA) to characterize pathway-level interactions, revealing that grounding components remain strongly redundant in both correct and hallucinating samples but undergo a consistent polarity flip, shifting from supporting the ground truth on correct samples to aligning with the hallucinated answer on erroneous ones. We further perform targeted suppression of hallucination-pathway components, showing that scaling these components reduces object hallucination by up to 76% with minimal accuracy cost, and validate that the same circuit selectively transfers to relational but not attribute hallucination. Evaluations on POPE-adversarial and AMBER show that the identified circuits are consistent across architectures, support causal intervention, and transfer selectively across hallucination types.
Authors:
Yashaswi Karnati, Kamran J Sadeghi, Akash Mehra, Li Ding, Alireza R Ghias, Pranav P Thombre, ShifangXu, Parth Mannan, Yu Yao, Hao Wu, Eric Harper, Ashwath Aithal, Nima TajbakhshAbstract: Foundation model training is becoming multimodal across the stack, from natively multimodal post-training pipelines to large-scale pretraining. As multimodal coverage broadens, context windows grow, and encoder–LLM scales diverge, a single LLM-centric TP/DP/PP/CP layout increasingly bottlenecks training throughput. This coupling forces encoders to inherit LLM-driven sharding and placement choices that can introduce unnecessary communication, limit useful encoder parallelism, or constrain the LLM schedule; the mismatch is especially pronounced at long contexts, where LLM context parallelism is needed for the fused multimodal sequence but encoder inputs remain bounded. We present heterogeneous parallelism for multimodal large language model training, a training-system abstraction that lets modules in the same end-to-end graph use independent parallel layouts and rank placements, supporting colocated execution, where modules share physical GPUs under different logical grids, and non-colocated execution, where modules occupy disjoint rank sets. The key systems challenge is preserving boundary tensor semantics when adjacent modules use independent layouts: forward activations must be materialized for the destination layout, while backward gradients must be routed back to the source layout. We address this with boundary communicators, runtime primitives that implement these forward and backward layout transforms, together with scheduling extensions for both placement modes. We evaluate optimized homogeneous, colocated heterogeneous, and non-colocated heterogeneous configurations across diverse multimodal workloads and GPU scales to characterize where each placement mode helps. Across this sweep, colocated heterogeneity improves TFLOPs/s/GPU by up to 14.4% at short context and 41.8% at long context, while non-colocated heterogeneity improves aggregate token throughput by up to 13.0% and TFLOPs/s/GPU by up to 9.6%. We validate loss convergence parity against homogeneous baselines and release an open-source implementation in Megatron-LM.
PaperID: 3360, Poster
Authors:
Brendan Lucier, Nicole Immorlica, Markus Mobius, Aleksandrs Slivkins, Daniel G Goldstein, Jake M Hofman, Sonia Jaffe, David RothschildAbstract: Motivated by agentic markets -- two-sided markets in which AI tools facilitate search -- we propose a model that incorporates individual consumer search and market-wide learning, solve it to understand the long-run market outcomes, and then study the impact of improved search technologies on these outcomes. A sequence of consumers engage in costly search to acquire progressively refined signals of product fit prior to purchase. The market observes acquired signals and post-transaction feedback thereby impacting future searches. We solve for the per-consumer optimal search policy and characterize the steady-state of the learning dynamics. As search technologies reduce the cost and/or increase the informativeness of search, they cause an improvement in learning and consumer surplus. Using numerical simulations, we also find that these technologies improve the rate of learning and the trajectory of consumer utility. Notably, we show that if the market is unable to observe acquired signals, search technologies can degrade market outcomes, highlighting the importance of transparency in agentic markets.
Abstract: Watermarking is an important tool for promoting the responsible use of large language models (LLMs). Existing watermarks insert a signal into generated tokens that either flags LLM-generated text (zero-bit watermarking) or encodes more complex messages (multi-bit watermarking). Though a number of recent approaches insert multiple bits into text without perturbing average next-token predictions, they largely extend design principles from the zero-bit setting, such as encoding a single bit per token. In contrast, a watermarker capable of embedding multiple bytes into the text would dramatically increase the potential applications, by embedding information such as the ID of the user who submitted the prompt, the precise model version that was used, or even the prompt itself. We address this problem by introducing ArcMark: a new watermark construction based on coding and information-theoretic principles that is capable of reliably embedding multiple bytes of information into just a few hundred tokens, without any distortion of the underlying LLM next-token distribution. We derive ArcMark by formulating the distortion-free watermarking problem as a channel coding problem, and deriving an information-theoretic channel capacity that establishes the fundamental limit of embedding information in LLM output in a distortion-free manner. This capacity formulation informs the design of ArcMark. In practice, ArcMark outperforms competing multi-bit distortion-free watermarks in terms of reconstruction accuracy, including in the face of attacks that alter a subset of the LLM text. ArcMark output is also shown to be indistinguishable from unwatermarked text in terms of perplexity, and in downstream task quality.
PaperID: 3362, Poster
Authors: Sina G Majelan, M. O Ahmad, M.N.S. Swamy
Abstract: Recent 3D point cloud classification networks suffer severe performance degradation when encountering real-world distribution shifts. To address this, we propose GeoCurv-TTT, a novel geometry-aware Test-Time Training (TTT) framework based on deformation restoration. Unlike existing methods that rely on geometry-agnostic masking or severe skeletal abstraction, our approach selectively deforms low-curvature planar regions while strictly preserving high-curvature structural anchors. This curvature-guided physical displacement generates targeted out-of-distribution (OOD) samples, forcing the network to learn deeply robust, topology-aware representations during pre-training. At inference, GeoCurv-TTT adapts to novel environments through two strategies: Standard TTT, which utilizes a self-supervised restoration objective to adapt to isolated instances, and Online TTT, which achieves real-time, backpropagation-free adaptation across streaming data by exclusively updating batch normalization statistics. Extensive experiments on the ModelNet40-C, ShapeNet-C, and ScanObjectNN-C benchmarks demonstrate that GeoCurv-TTT establishes new state-of-the-art (SOTA) performance for 3D domain adaptation, outperforming existing methodologies in accuracy, computational efficiency, and adaptation time.
PaperID: 3363, Poster
Authors: El M Mansouri
Abstract: High-stakes prediction systems are often evaluated only on labels revealed by past decisions. This selective observation can hide the losses that determine full-population tail risk: a predictor may appear safe on revealed labels while its worst compatible failures remain unseen. We study certification of conditional value-at-risk (CVaR) under selective labels. First, we prove non-identifiability: even with overlap, two predictors with identical observed selected-loss distributions can have reversed full-population CVaR rankings under compatible hidden-label laws. Under an outcome-dependent odds-ratio sensitivity model, we derive sharp upper and lower CVaR envelopes. A minimax interchange shows that worst-compatible hidden-label completion commutes with the Rockafellar--Uryasev threshold optimization, giving an exact certificate rather than a loose robust surrogate. We then define the tail-risk certification frontier: the minimum passive revelation cost needed to reduce hidden-tail ambiguity below a target tolerance. For a fixed predictor and threshold, this frontier is a fractional-knapsack problem whose value density, tail-identification value, pinpoints labels that can move the CVaR tail. Finally, we give finite-sample guarantees for cross-fitted conservative certificate learning and a lower bound governed by the effective number of revealed tail labels.
Abstract: Activation steering controls language model behavior by adding directions to internal representations at inference time, but standard residual-stream steering can fail in stateful dialogue. We identify KV-cache contamination as a key failure mode: steered token states are stored and repeatedly reused, turning a local perturbation into cumulative coherence degradation. Motivated by the stability of prompt-based control, we propose Gated Cropped Attention-Delta steering (GCAD), which extracts steering signals from system-prompt contributions to self-attention and applies them with token-level gating. Across persona-steering experiments, GCAD preserves trait control while substantially improving long-horizon coherence. On the main multi-turn benchmark, GCAD improves average coherence drift from -18.6 to -1.9 and raises turn-10 trait expression from 78.0 to 93.1. These results suggest that activation steering becomes more reliable when interventions follow the prompt-mediated pathways that models already use for behavioral control.
Abstract: Activation steering controls LLM behaviour towards target behaviour by intervening in internal representations, yet it often degrades reasoning and retrieval performance. We argue that a primary cause of this trade-off is attention rerouting: steering vectors alter query-key matching, shifting attention away from contextually important tokens toward less informative ones. To address this, we propose Steering via Key-Orthogonal Projections (SKOP), a steering method that constrains harmful attention rerouting without eliminating steering efficacy. SKOP achieves this by preserving attention patterns on a small set of focus tokens the model relies on for reasoning and retrieval, while allowing redistribution among less critical tail tokens. Across multiple steering benchmarks, we show that SKOP achieves the best joint steering-utility trade-off, reducing utility degradation by 5–7× while retaining over 95% of vanilla steering efficacy. Our results further suggest that, in long-context retrieval settings where vanilla steering approaches are ineffective, SKOP can maintain robust performance by avoiding attention rerouting.
PaperID: 3366, Poster
Abstract: Recent DETR-based tiny object detectors adopt dynamic-query mechanisms to handle density imbalance. However, existing designs entangle two effects: how many and which queries are decoded. This entanglement obscures what actually drives the gains, thereby hindering principled dynamic-query design; additional learned predictors or heuristic filtering steps also introduce extra overhead and inevitable decision errors, which can undermine practicality and final gains. In this paper, we disentangle these effects and show that dynamic query primarily acts as budgeting---allocating decoder query capacity according to object density. Motivated by this view, we formulate a monotonic and conservative budgeting principle and propose a budgeting-based DETR (BUTR), a streamlined framework that implements dynamic query by using a lightweight, training-free budgeter to estimate an input-dependent query budget K(x) from encoder confidences. BUTR applies standard Top-K(x) query initialization, removing heuristic filtering, learned components, and complex selection logic. On AI-TOD-v2 and VisDrone, BUTR shows clearer density-aligned query allocation behavior than prior dynamic-query detectors and improves AP by more than 0.5, while using about half as many queries with fewer parameters and GFLOPs, requiring over 2× less training time, and delivering 4× faster inference. Additional COCO results further indicate its applicability across a broad spectrum of detection scenarios.
PaperID: 3367, Poster
Authors: Chaeho Lee, Junyup Kim, Wansoo Kim, Young MIn Jung, Sang H Lee
Abstract: Integrating human feedback into language models has gained increasing attention as a means to align model outputs with human preferences, with Reinforcement Learning from Human Feedback (RLHF) serving as a prominent framework for preference-based alignment. Recent studies have incorporated eye-tracking (ET) data as an additional supervisory signal for reward model training, grounding reward learning on human reading behavior. Building on this approach, we extend prior fixation-centric models by leveraging established findings from cognitive psychology on human letter recognition. Specifically, we model visual attention as a spatially graded distribution centered on the current gaze location, and incorporate this representation into reward model training. Our results show that the reward model trained with the proposed gaze distribution achieves higher preference prediction accuracy over a baseline model. Moreover, the fitted attentional distribution reflects key properties of human attention during reading, suggesting that it captures cognitively meaningful aspects of attention allocation.
PaperID: 3368, Poster
Abstract: The Forward-Forward Algorithm (FFA) replaces backpropagation (BP) with layer-wise local contrastive objectives, eliminating the backward pass, yet suffers a persistent performance gap with BP that worsens with depth. This paper diagnoses two distinct deficits: an irreducible optimization floor arising from concurrent local updates, whose magnitude is amplified when kernel contraction degrades the optimization Gram; and a geometric collapse of layer representations driven by kernel contraction itself. On the optimization side, we prove that the FFA loss satisfies the Polyak--\Lojasiewicz inequality at each layer, but concurrent layer updates create an irreducible error floor that grows with depth and is amplified when the Gram degrades. On the representational side, the pairwise similarity kernel of layer representations contracts exponentially toward rank one as depth increases, collapsing the diversity of per-layer error signals. This collapse bounds FFA's \empheffective learning capacity---the total diversity of gradient information across layers---to grow only linearly with depth regardless of width, whereas BP's chain-rule signal preserves per-layer diversity, yielding a capacity that scales with both depth and width.
PaperID: 3369, Poster
Abstract: Flow models show promise for molecular design due to their fast, expressive sampling capabilities. However, their applicability to complex biological systems remains limited by a lack of physical grounding, which leads to unrealistic or unstable molecular structures. This work presents a framework that integrates force field guidance with consistency-based training to improve the physical fidelity of flow-based generative models. We calibrate pretrained flow models using differentiable physical energy functions to steer generation toward low-energy and sterically valid conformations. A consistency-based training strategy enables accurate, robust generation with very few sampling steps, improving inference efficiency. We evaluate it on multiple protein design tasks, including full-atom structure generation and peptide binder modeling. Experiments show that our approach consistently improves physical plausibility and geometric stability, while enabling favorable trade-offs between structural diversity and computational cost. By unifying physical calibration and efficient sampling, we advance the scalability and reliability of flow-based molecular generation.
PaperID: 3370, Poster
Abstract: Graph neural networks are widely deployed in real-world applications, where attackers often operate under strict black-box constraints, making effective graph black-box attacks particularly important. However, existing methods are fundamentally limited by poor universality, as attack strategies learned on specific graphs or model configurations fail to generalize across different graph datasets. Inspired by the unified representation principle of graph foundation models, we propose a universal black-box attack framework that learns transferable attack paradigms from multiple source graphs and generalizes directly to unseen target graphs. The framework incorporates a representation selection mechanism and a graph variational autoencoder based modeling strategy to ensure consistency and mappability between the unified semantic space and the original representation space, together with a structure–representation joint anchor node selection mechanism for effective attack construction. Extensive experiments on benchmark datasets show that our method achieves strong universality and consistently superior attack performance.
PaperID: 3371, Poster
Authors:
Feiyuan Zhang, LI Pengbo, Ziniu Li, Yuhao Jiang, Di Chai, Han Tian, Junhao Wang, Taiqiang Wu, Guanhua Huang, Chaoliang Zeng, Yihao Liu, Kai ChenAbstract: Rollout generation bottlenecks large-scale RL training for LLM agents. Tree rollout has emerged as an important agentic-RL strategy: by branching from shared intermediate states across reasoning turns, tool calls, and environment observations, it avoids regenerating entire trajectories from scratch and exposes long prefixes for context reuse. However, existing RL and LLM-serving infrastructure remains largely tree-unaware: it treats sibling branches as independent generation requests, making the rollout tree invisible to routing, preemption, and KV-cache management. Consequently, shared prefixes are scattered across servers, discarded with request-private suffixes, or recomputed after cross-server spill. We present Canopy, a tree-aware rollout scheduling system that makes the active rollout tree a first-class serving abstraction without changing sampling, rewards, advantage estimation, or optimizer updates. Canopy consists of three mechanisms: tree-aware routing and retention, which keep siblings near reusable prefixes and protect active prefixes; prefix-preserving partial preemption, which separates shared-prefix KV from private-suffix KV under pressure; and spill-aware transfer-or-recompute admission, which transfers shared-prefix KV to a spilled sibling only when its estimated transfer cost is lower than destination-side recomputation. On long-horizon multi-turn agent workloads, Canopy achieves up to 2.11× rollout-generation speedup; across four representative tree-rollout algorithms, mechanism studies show higher prefix locality and a 51.1% average, 94.1% maximum reduction in full-release shared-prefix invalidations. These results show that exposing rollout-tree structure to the rollout-serving stack can reduce wasted prefill and KV-cache churn, enabling more efficient RL training for LLM agents.
Abstract: We study stochastic optimization with heavy-tailed gradient noise. We first propose a novel quantum mean estimator for multivariate heavy-tailed random variables that achieves lower query complexity than optimal classical estimators in the low-dimensional regime. We further develop an unbiased quantum mean estimator by applying a generalized multi-level Monte Carlo technique. We prove quantum lower bounds showing that, when the dimension d of the random vector is small and can be viewed as a constant, our quantum estimators are optimal up to logarithmic factors; while for the high-dimensional regime, no quantum speedup is available compared to optimal classical mean estimators. Based on these estimators, we propose a quantum normalized stochastic gradient descent method (\textttQNSGD), which finds an \epsilon-stationary point of a nonconvex objective using \tilde\mathcalO\big(d^\fracp4(p-1)\epsilon^-\frac5p-42p-2\big) queries to the stochastic gradient, where p\in(1,2] is the tail index. For a convex objective function, we propose a quantum projected stochastic gradient descent method (\textttQPSGD), which computes an \epsilon-optimal solution with query complexity \tilde\mathcalO\big(d^\fracp4(p-1)\epsilon^-\frac3p-22p-2\big). Our results improve upon the classical lower bounds \Omega\big(\epsilon^-\frac3p-2p-1\big) for nonconvex problems and \Omega\big(\epsilon^-\fracpp-1\big) for convex problems, demonstrating a quantum speedup when d is relatively small.
PaperID: 3373, Poster
Abstract: As AI systems are increasingly deployed in high-stakes domains such as healthcare, effectively integrating human expertise with AI outputs is critical for reliable decision-making. A substantial body of theoretical work studies optimal mechanisms for human-AI collaboration (HAC), focusing on when to involve humans and how to combine their expertise with AI predictions. However, these approaches largely overlook a key factor highlighted by empirical evidence: human judgments are impressionable and can vary depending on how AI outputs are presented. For example, when AI systems defer decisions to humans, revealing or withholding the AI prediction can significantly influence human judgments. Moreover, existing work neglects deployment-time variability, where conditions, particularly the cost of eliciting human input, may differ substantially from those observed during training. To address these challenges, we introduce a rigorous and theoretically grounded framework for human-AI collaboration, termed \emphcollaborative omnimechanism design, based on a novel and computationally efficient contextual bandit variant of boosting for outcome indistinguishability. In particular, we reduce a calibration-type outcome indistinguishability condition to projected smooth calibration, providing novel time-efficiency result, bypassing the high-dimension curse induced by complicated collaboration methods. Our framework provides a principled characterization of human impressionability and introduces a unified optimization scheme that enables robust and adaptable decision-making across diverse query costs and evaluation objectives without retraining. We further validate our theoretical results through comprehensive experiments.
PaperID: 3374, Poster
Authors: Chao Ouyang, Yuyang Bai, Jun Zhang, Tianlu Gao, YuxinDai, xu peidong, David W Gao
Abstract: Multi-agent LLM systems (MAS) improve reasoning by decomposing across specialized roles, but their text communication is expensive. Each agent decodes thousands of intermediate tokens for the next, and successive prefills grow cumulatively. Recent work shows that latent-space signals can substitute for explicit text between agents, cutting cumulative prefill cost. When the latent reasoning also replaces an agent's text reasoning, it additionally cuts intermediate tokens. Here, we find that the prefill cache alone suffices as the latent-space signal between agents, and the per-agent forward chains can be compressed into a single pass. With \emphexactly one prefill and one decode per problem, we sharply cut per-rollout tokens and prefill wall time. We call this architecture CacheMAS, in which prefill agents pass information through cache communication, and a final agent decodes the answer. As a bonus, CacheMAS is one autoregressive rollout, so unlike prior MAS, standard sequence-level RL can jointly optimize all agents at once without per-agent reward design. On math, multi-hop QA, and code, CacheMAS uses ~3× fewer tokens, runs ~2× faster, and after joint optimization beats prior MAS by +1 to +8 points on every benchmark. Cache communication is not just a token-efficient substitute for text, but a practical interface for joint MAS optimization.
PaperID: 3375, Poster
Authors:
Qiyang Zhang, Xinhao Li, Lei Shi, Zheng Lin, Jinfeng Wen, Ao Zhou, Shangguang WangAbstract: Onboard satellite models often require frequent updates, but the weights adapted to earlier data distributions can quickly become outdated. However, updating large-scale model parameters in orbit presents significant challenges due to the limited uplink bandwidth of Low Earth Orbit (LEO) satellite systems, particularly for hyperspectral satellite imagery, where high-dimensional spectral–spatial inputs lead to increased model size and update costs. Existing full fine-tuning methods are thus expensive to retrain and difficult to deploy under strict communication constraints. To address this challenge, we propose NE-LoRA, a parameter-efficient adaptation framework for bandwidth-constrained onboard hyperspectral model updates. NE-LoRA combines a primary low-rank branch with a nonlinear auxiliary branch to capture both global update trends and complex spectral–spatial variations. Additionally, we introduce a differentiated training strategy for multi-matrix adapters, motivated by the asymmetric initialization and gradient dynamics of different adapter matrices. Experiments on four hyperspectral datasets and three representative backbone models demonstrate that NE-LoRA consistently outperforms LoRA-based baselines and remains competitive with, and in several cases superior to, full fine-tuning. Across the evaluated settings, NE-LoRA updates only a small fraction of the total parameters on average while preserving low deployment overhead, offering a favorable accuracy–communication trade-off for onboard hyperspectral adaptation.
PaperID: 3376, Poster
Abstract: Autoregressive large language models (LLMs) have rapidly advanced in capability, but their increasing scale comes with substantial computational and memory costs at inference time. Knowledge distillation (KD) offers a practical solution by transferring knowledge from a large teacher model to a smaller student model via alignment of discrete probability distributions. However, existing KD methods for LLMs primarily rely on divergences that evaluate discrepancies through probability values at each vocabulary index, without explicitly leveraging token-level semantic information. We propose Wasserstein-based knowledge distillation (WASD) for LLMs, which incorporates token-level semantic information via the Wasserstein-based distance with a cost matrix derived from token embeddings. To ensure computational tractability, we adopt the Sinkhorn divergence and derive an equivalent formulation that can be efficiently optimized without introducing additional networks. Experiments across multiple LLM families and scales show that WASD consistently improves distillation performance on diverse tasks, including instruction following, mathematical reasoning, and code generation. Our results highlight the importance of semantic information encoded in the token space for effective distribution alignment in LLM distillation.
PaperID: 3377, Poster
Abstract: Online Conformal Prediction (OCP) sequentially constructs prediction sets whose average empirical coverage converges to a target level over time, for arbitrary, potentially adversarial datasets. In the multi-level setting of OCP, prediction sets are produced at d coverage levels at once. Prior work showed that online gradient descent with lazy isotonic projections produces sets properly nested across levels, but with a \sqrtd dependence in the per-level coverage convergence rate. Here, we both generalize and improve prior work. First, we show that lazy isotonic projections can be safely used for the construction of multi-level online conformal prediction sets from \emphany online learning algorithm. Then, through a black-box reduction, we show that the linearized regret of the online algorithm itself automatically gives control over both per-level calibration and distributional consistency of multi-level predictions via Fenchel duality. Finally, by using the Weighted Interval Score, the standard scoring rule for assessing multi-level interval forecasts, we prove that these projections reduce both the size of the sets and their under-coverage. We show two instantiations of our theory, p-norm Online Mirror Descent and coordinate-wise Universal-Portfolio, both improving the dependence on the number of levels to \sqrt\log d for the multi-OCP problem. Experiments on stock prices, electricity demand, and synthetic streams confirm consistent gains in coverage-width trade-offs over existing baselines.
Abstract: Identity-preserving text-to-video generation (IPT2V) empowers users to produce diverse and imaginative videos with consistent human facial identity. Despite recent progress, existing methods often suffer from significant identity distortion under large facial pose variations or facial occlusions. In this paper, we propose FaithfulFaces, a pose-faithful facial identity preservation learning framework to improve IPT2V in complex dynamic scenes. The key of FaithfulFaces is a pose-shared identity aligner that refines and aligns facial poses across distinct views via a pose-shared dictionary and a pose variation–identity invariance constraint. By mapping single-view inputs into a global facial pose representation with explicit Euler angle embeddings, FaithfulFaces provides a pose-faithful facial prior that guides generative foundations toward robust identity-preserving generation. In particular, we develop a specialized pipeline to curate a high-quality video dataset featuring substantial facial pose diversity. Extensive experiments demonstrate that FaithfulFaces achieves state-of-the-art performance, maintaining superior identity consistency and structural clarity even as pose changes and occlusions occur. The code and dataset pipeline will be released.
Abstract: Competitive multi-agent reinforcement learning in imperfect-information games requires agents to act under partial observability and against adversarial opponents, necessitating stochastic policies. While self-play reinforcement learning with Proximal Policy Optimization (PPO) has achieved strong empirical success, its standard advantage estimator, generalized advantage estimation, suffers from additional variance due to the sampling of stochastic future actions. This variance is amplified in equilibrium self-play because of the stochastic nature of the equilibrium policy and persists even when the critic is exact. We address this bottleneck by introducing Q-boosting, a variance-reduced advantage estimator based on a centralized action-value critic, and propose Variance-Reduced Policy Optimization (VRPO), incorporating this new estimator. The algorithm replaces sampled multi-step backups with a multi-step Expected SARSA(\lambda) trace, computing policy expectations at each step to average out action-sampling noise, while retaining PPO's clipped objective and on-policy actor updates. Empirically, VRPO consistently achieves strong performance from mid-sized to large-scale games including Dou Dizhu and Heads-Up No-Limit Texas Hold’em.
Abstract: Expensive multi-objective optimization is a prevalent and crucial concern in many real-world scenarios, where sample-efficiency is vital due to the limited evaluations to recover the true Pareto front for decision making. Existing works either involve repeatedly fitting Gaussian process surrogates from scratch for each newly encountered problem, or rely on computationally demanding hypervolume-oriented policy learning for amortized optimization, making it challenging to achieve scalable pre-training and efficient yet robust adaptation for diverse emerging real-world applications. To address these challenges, we propose FoMEMO (Foundation Models for Expensive Multi-objective Optimization), which adopts a decomposition-based training paradigm that converts multi-objective optimization into preference-wise aggregated posterior learning tasks, enabling scalable pre-training on hundreds of millions of diverse synthetic datasets without relying on extensive real-world domain experiments. At test time, given observed trajectories from unseen problems and user preferences, the foundation model performs in-context posterior prediction and optimizes the derived acquisition functions for efficient candidate generation, achieving strong generalization and optimization performance without any subsequent model training or updates.
PaperID: 3381, Poster
Authors: Giacomo Petrillo
Abstract: Bayesian Additive Regression Trees (BART) is a nonparametric Bayesian regression technique based on an ensemble of decision trees. It is part of the toolbox of many statisticians. The overall statistical quality of the regression is typically higher than other generic alternatives, and it requires less manual tuning, making it a good default choice. However, it is a niche method compared to natural competitors such as XGBoost, due to the longer running time, making sample sizes above 10000–100000 a nuisance. We present a GPU-enabled implementation of BART, faster by up to 100x in a given GPU vs. CPU comparison, making BART fast enough to be a viable alternative to non-Bayesian algorithms on large datasets, and showing that BART parallelizes better than its non-Bayesian counterparts. This implementation is available in the Python package bartz.
PaperID: 3382, Poster
Authors:
Ngan Nguyen, Dung Nguyen, Quan Dao, Dimitris Metaxas, Minh Hoai, Cuong Pham, Anh TranAbstract: Diffusion models have become a leading approach in visual generative modeling, particularly effective in text-to-image and controllable generation. Their main drawback is slow inference as it requires many sampling steps. While recent distillation methods enable one-step text-to-image generation, extending these techniques to controllable generation is still challenging. Existing controllable models, such as ControlNet, rely on dual-branch design that make distillation difficult. In this work, we adopt a simpler strategy. Instead of adding extra conditional branches, we finetune a selected subset of layers in the original text-to-image diffusion model to directly enable controllable generation. This results in DeltaControl, a single-flow, multi-step diffusion model capable of supporting spatial-condition generation while matching the performance of dual-branch approaches. Building on this teacher model, we further apply step-distillation to obtain FlashControl, a one-step model capable of spatial controllable generation with a single forward pass.
Abstract: Variable fonts enable continuous variation of glyph geometry along semantic design axes such as weight, width, slant, and optical size. However, constructing a variable font from a static font remains a labor-intensive process requiring expert typographic design and manual specification of glyph variation data. We introduce NIV (Neural Axis Variations), a method that automatically converts a static font into a fully functional variable font. Given glyph outlines and a set of desired design axes, NIV predicts per-point displacements. The model operates directly on vector glyph geometry and employs a novel Property Embedding mechanism that captures interactions between multiple axes, enabling consistent multi-axis variation within a unified framework. We train NIV on a newly constructed dataset derived from variable Google Fonts, comprising over one million variation tuples. The resulting model generalizes across unseen code points, unseen font styles, high-complexity CJK glyphs, and even out-of-distribution handwriting inputs. The generated outputs are standard variable font files supporting continuous interpolation via existing rendering engines. To facilitate research, we release the dataset, the complete training and inference implementation, and trained models. Beyond typography, our approach demonstrates how structured geometric objects with continuous parametric variation can be synthesized using neural deformations.
PaperID: 3384, Poster
Authors: Xiaoyong Lu, Guobao Xiao, Dong Liang, Songlin Du
Abstract: Keypoint detection plays a central role in geometric vision pipelines. Existing detectors exhibit complementary strengths across different scene structures and keypoint budgets, and reliance on a single detector often results in suboptimal keypoint distributions. Motivated by this observation, we propose for dense matching pipelines, a unified framework for complementary keypoint detection through the adaptive composition of heterogeneous detectors. MoD first distills heterogeneous detectors into a unified model, allowing diverse detection maps to be predicted efficiently within a single forward pass. Detector selection is subsequently formulated as a patch-wise, budget-conditioned decision process, in which a lightweight router assigns the most suitable detector to each local region. To avoid heuristic supervision, the router is trained with task-driven rewards derived from downstream matching quality, jointly capturing both keypoint quantity and precision. Extensive experiments on MegaDepth, ScanNet, and HPatches demonstrate that MoD consistently outperforms strong baselines across a wide range of keypoint budgets, demonstrating its effectiveness in improving keypoint distribution and downstream geometric performance.
PaperID: 3385, Poster
Abstract: As large language models (LLMs) become increasingly deployed as autonomous agents, the ability of these models on automated planning tasks has become essential. However, despite their strengths, LLMs' planning ability is not robust, particularly when they are not allowed to exploit commonsense cues in the symbolic vocabulary (e.g., names of actions, predicates, objects) used to define planning tasks. This makes them unreliable compared to symbolic planners. With recent reasoning models showing improved planning performance on tasks with obfuscated symbolic vocabularies, we utilize an adversarial attack framework to reveal to what extent these crucial weaknesses of LLM agents remain and whether they can reason properly when presented with severe cases. As such, our main contribution is devising and optimizing an LLM-based attack model which we call \textttSymbolic-Swapper to automatically generate obfuscated domains with non-standard symbolic vocabularies that systematically break strong reasoning models' ability to plan. In doing so, we reveal by how much LLMs planning abilities can be further degraded under more advanced obfuscation schemes, allowing us to ascertain their success when having to rely on their pure planning ability. We find that attacks can decrease the planning performance of base models and even robust planning frameworks by over 70%. We further analyze the factors behind how their planning abilities break down, and find that successful obfuscation models to significantly degrade in reasoning quality despite them showing an ability to understand the domain, revealing a gap in how model's perceive a task domain and how they are actually able to plan under it, with attacks often bypassing model ability to successfully self-correct.
Abstract: Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art models such as \pi_0.5 operate under a single-frame paradigm, limiting their ability to retain past observations and develop precise spatial perception. In this paper, we propose StreamPI, a streaming multimodal temporal modeling framework that equips single-frame VLA with temporal reasoning capability without introducing any additional parameters. One core design is instruction-anchored temporal modeling. It treats each (visual observation, language instruction) pair as an atomic temporal unit: bidirectional attention within each pair enables cross-modal fusion, while causal attention across pairs preserves autoregressive streaming inference. This ensures the language instruction serves as a persistent semantic anchor throughout task execution. To bridge the gap between synchronous training and asynchronous real-robot deployment, we introduce a random-interval streaming training strategy: a proper inter-frame interval (e.g., every 3 frames) enables faster and smoother action execution. Beyond this, randomizing the interval further improves robustness to frame-timing perturbations, supporting asynchronous deployment in practice. Furthermore, by leveraging the length extrapolation capability of the LLM backbone, StreamPI seamlessly inherits pretrained single-frame weights and supports flexible single-frame and multi-frame inference. Experiments on real-robot tasks spanning memory-dependent and precise perception scenarios, as well as the simulation benchmark LIBERO, demonstrate that StreamPI outperforms \pi_0.5 across diverse tasks.
Abstract: Video diffusion Transformer (DiT) models excel in generative quality but hit major computational bottlenecks when producing high-resolution, long-duration videos. Full attention requires computing dot products between all pairs of queries and keys, resulting in a quadratic computational complexity with respect to the sequence length ((O(L^2))), leading to high training and inference costs.To overcome this limitation, we propose a Bidirectional Sparse Attention (BSA) framework that sparsification from both the query and key–value directions to reduce the quadratic cost of full attention.Specifically, the sparsification of queries is achieved by pruning tokens that are locally redundant or semantically similar within the 3D spatiotemporal domain of a video. For the key–value pairs, only those that are highly correlated with the query at a global level are selected for attention computation. Furthermore, we design a dynamic threshold adjustment mechanism to adaptively regulate the selected number of key–value pairs, thereby maintaining quality without degradation. Extensive experiments demonstrate that BSA significantly accelerates DiT training across long sequences, reducing FLOPs by up to 20× and achieving 17.79× faster attention training, while preserving or even surpassing the generative quality of full attention.
PaperID: 3388, Poster
Abstract: Large language models are increasingly deployed as autonomous agents that plan and act over long-horizon interactions, learning from both final outcomes and rich environmental feedback. However, optimizing from such interaction experience remains challenging, as feedback is heterogeneous and rarely carries explicit reward information. Existing post-training approaches rely on outcome-driven methods, such as reinforcement learning with verifiable rewards (RLVR), which primarily exploit final success signals while leaving interaction experience and feedback underutilized. In sparse-reward, long-horizon tasks, this often results in distribution sharpening: the policy reinforces a narrow set of already-successful behaviors, without substantially improving the feedback-grounded agency needed for broader problem-solving capacity (e.g., Pass@k). We propose Counterfactual Distillation, a framework that converts exploration-derived experience into trainable policy supervision. Specifically, we organize exploration as a tree-structured search process, where the agent reflects on past decision points and makes new experience-guided decisions to form alternative branches. These corrections are then distilled as counterfactual targets: we train the model to predict the revised action from the original history without explicitly providing the experience, thereby introducing out-of-distribution behavior patterns and expanding exploration capacity. Across diverse interactive coding and agentic tasks, our method outperforms outcome-driven baselines such as GRPO and experience-based methods such as Early Experience, with gains of up to 14% on Pass@128. When interleaved with reinforcement learning updates, it further raises the performance ceiling, yielding over 10% improvement in Pass@1. We provide an [anonymized implementation](https://anonymous.4open.science/r/Counterfactual-Distillation-0E71).
PaperID: 3389, Poster
Abstract: LLM agents are increasingly used for open-ended, data-driven scientific discovery, yet most systems remain pointwise: they evaluate one candidate hypothesis at a time for surprise, novelty, or evidential support. This is information-inefficient: rejecting a hypothesis reveals that one explanation is unlikely, but not which alternatives are better supported. We argue that the proper epistemic unit of discovery is not a single hypothesis but a scientific question together with the mutually exclusive and collectively exhaustive (MECE) explanations that could answer it. We introduce Contrastive discovery, which scores candidate discoveries by how evidence redistributes belief across competing explanations. This allows a single experiment to both rule out hypotheses and identify the most plausible alternative. Across six scientific domains, Contrastive discovery yields substantially more resolved scientific questions than pointwise baselines under the same evaluation budget. High-surprisal pointwise candidates are overwhelmingly rejection-heavy, yet many become resolved alternatives when evaluated contrastively, showing that an explicit representation of alternatives can turn negative evidence into resolved discoveries.
PaperID: 3390, Poster
Abstract: Deep networks embed inputs in high-dimensional representations, but how much of this capacity actually drives behavior? We introduce a framework for identifying, at each layer, the subspace of the representation that is functionally relevant for the task. Using a matrix analogous to the Fisher information but defined over the hidden representation, we rank directions by their influence on downstream computation and project onto the most task-relevant subspace. We benchmark this behaviorally grounded compression against three geometric alternatives that preserve variance, local neighborhood structure, or global topology. Across pretrained vision models, representations are remarkably compressible: a small fraction of the task-relevant subspace preserves most performance, often requiring far fewer dimensions than geometric alternatives. We show that this decomposition recovers functionally distinct subspaces within the same representation: on cue-conflict stimuli, shape-relevant and texture-relevant directions are largely non-overlapping, and ablating one selectively impairs the corresponding task while leaving the other intact. Task-relevant subspaces also provide a principled foundation for continual learning: constraining gradient updates to lie orthogonal to the task-relevant subspaces of prior tasks outperforms standard variance-based gradient projection methods. Finally, restricting representational comparisons to task-relevant subspaces reveals that inter-model alignment is higher in task-relevant directions than in task-irrelevant directions at later layers, and that conditioning on different tasks over the same stimuli uncovers asymmetries that full-space similarity measures cannot detect. Together, these results establish that deep networks compute through a low-dimensional functional core embedded within a much larger representational space, that these cores differ across tasks, and that they provide a unifying basis for compression, continual learning, and representational comparison.
Abstract: Pairwise constrained clustering typically relies on hard must-link/cannot-link labels, whereas realistic pairwise supervision may be real-valued and entangle intrinsic ambiguity, expert judgment, and stochastic corruption. Existing deep constrained clustering (DCC) methods mainly target hard, expert-agnostic constraints, treating soft labels mostly numerically rather than semantically. We formalize this setting as uncertainty-aware probabilistic constrained clustering (UPCC), defining a canonical aleatoric target through a heterogeneous observation process and analyzing its conditional identifiability. We introduce ProbPair, an angular pairwise objective for probabilistic relations, and build ECI-PP, an estimator--corrector--integrator framework that refines imperfect supervision via belief estimation, correction, and reliability-aware integration. Across challenging probabilistic supervision settings, experiments on diverse benchmarks show that ECI-PP outperforms state-of-the-art DCC methods and remains robust with a shared default configuration.
PaperID: 3392, Poster
Authors:
Yun-Yen Chuang, Chen-Sheng Gu, Hung-Min Hsu, Kevin Lin, Ray-I ChangAbstract: Diffusion models for 3D molecular geometry almost universally adopt an isotropic Gaussian prior, whose Frobenius-norm scale \sigma_T\sqrt3n is governed by the noise level and atom count rather than by chemistry, producing a scale mismatch whose effects grow with n and yield high spatial entropy and unstable early trajectories. We introduce Shell-guided Spherical Diffusion (SSD), a model-agnostic framework that replaces the Gaussian prior with a chemically scaled spherical-shell initialization and augments both the forward and reverse processes with coordinated radial attraction, short-range repulsion, and an SE(3)-equivariant correction field. This joint design of initialization and dynamics is essential: neither a shell alone nor radial fields alone reproduces the stability or accuracy of SSD. We evaluate SSD across five representative coordinate-space backbones---GeoDiff, SubGDiff, EDM, SemlaFlow-style flow matching, and MCF---each under its canonical evaluation protocol. SSD consistently improves both quality and diversity under identical training and sampling budgets, upgrading weaker backbones such as GeoDiff and EDM to match or surpass stronger diffusion-, VAE-, and flow-based baselines. SSD therefore serves as a plug-in geometric enhancement that strengthens coordinate-space molecular generation models without modifying their architectures, loss functions, or training pipelines.
PaperID: 3393, Poster
Abstract: Most video-text encoders are pretrained on carefully curated video-text datasets where each video is paired with a single short caption. Such supervision captures only one narrow view of a video, despite videos naturally supporting many complementary descriptions of their objects, actions, and relations. We present Simmer, a scalable pretraining recipe for video-text representation learning that increases textual coverage per video and matches the training objective to this richer supervision. We first augment a video corpus with diverse synthetic captions written in many distinct textual registers, producing a large set of aspects per video. Since this many-to-one supervision is difficult to capture in a single global embedding, we pair caption diversification with a multi-vector contrastive loss that allows videos to align with multiple relevant captions. Second, we show that introducing a distinct cooldown phase over a smaller corpus of human-sourced dense captions with our stylistic augmentations further improves performance. Our resulting encoders achieve state-of-the-art performance on five video-text encoding tasks across three model scales, outperforming prior works while using 2-60x fewer videos. Our results show that broader caption coverage and multi-aspect alignment can pave the way towards new scaling laws in video-text modeling.
Abstract: Vision-Language Models (VLMs) have been increasingly adopted for Image Quality Assessment (IQA). However, current methods typically employ a static one-shot scoring paradigm, despite the fact that humans assess image quality through dynamic visual inspection, e.g., selectively adjusting views to verify details and subtle artifacts. Specifically, relying solely on a single-pass observation introduces two primary limitations: first, perceiving the image only at a global scale restricts the assessment of finer local details; second, the original intensity distribution of the image may overwhelm the visibility, leading to insufficient inspection of image quality. To address these issues, we propose Tool-IQA, shifting the assessment mechanism from passive scoring to a tool-augmented workflow. In particular, we equip VLMs with simple yet effective view tools: a Magnifier to inspect local details, and a Gamma Corrector to uncover visibility and hidden artifacts. The assessment follows a structured pipeline that consists of an initial observation with rubric notes, a tool-augmented in-depth inspection, and a final quantification for calibrated quality score. Furthermore, to ensure efficient and purposeful tool callings, we introduce a batch-aware training strategy to reward tool interactions that can yield positive contributions rather than simply encouraging usage. Experiments on a variety of IQA benchmarks demonstrate that, with effective tool calling and calibrated assessment, our proposed Tool-IQA significantly outperforms existing state-of-the-art models, e.g., it achieves a PLCC of 0.854 on the challenging CLIVE dataset. The code and models will be released.
Abstract: Vision-Language Models (VLMs) have achieved substantial progress across a wide range of understanding and reasoning tasks, driven by large-scale image-text training aimed at multimodal fusion. Ideally, replacing a textual question with its rendered-image counterpart should leave model performance essentially unaffected. In practice, however, such modality substitution induces dramatic performance degradation. We attribute this carrier sensitivity issue to an inherent bias in current training corpora. Across prevalent datasets such as image captioning, VQA, OCR, and web-sourced interleaved data, text and images are typically organized into distinct and asymmetric roles, with text serving as linguistic queries and images as visual references. Such data bias leads VLMs to exhibit distinct preferences for information acquisition across different modalities. Consequently, VLMs fail to align representations of semantically equivalent content across textual and visual carriers, making model reasoning fragile under modality substitution. To address this, we propose Local Modality Substitution (LoMo), a lightweight, architecture-agnostic data curation paradigm designed to provide supervision for cross-modal representational invariance between semantically equivalent text and image carriers. LoMo achieves this by reformulating single-modality prompts into seamlessly interleaved multimodal sequences. It dynamically selects target text spans and recasts them as rendered images, thereby preserving the same semantics across "text, visual, text" carriers. Extensive experiments across 13 diverse multimodal benchmarks demonstrate that LoMo significantly improves overall multimodal reasoning and yields deeper cross-modal fusion. Specifically, it delivers consistent gains across foundational models, improving over standard SFT by 2.67 points on LLaVA-OneVision-1.5-8B and 2.82 points on Qwen3.5-9B.
PaperID: 3396, Poster
Abstract: We propose a primal-dual policy optimization method for reinforcement learning in semi-infinitely constrained Markov decision processes (SICMDP), which extends standard constrained Markov decision processes by allowing a continuum of constraints. We call our proposed method Semi-Infinitely Primal-Dual Policy Optimization (SI-PDPO). By introducing an infinite-dimensional Lagrange multiplier, we derive the saddle-point formulation of the original SICMDP problem and prove strong duality under mild conditions. We then introduce a regularization term for the dual variable, and solve the resulting regularized saddle-point problem using first-order methods with a policy optimization subroutine (for example, NPG or PPO). Unlike existing primal-type policy optimization algorithms for semi-infinitely constrained reinforcement learning, our method does not require solving an inner-loop optimization problem, thereby mitigating computational intractability and reducing the over-conservative bias. We perform a series of numerical experiments, including a real-world power control task in wireless communications, to evaluate the performance of SI-PDPO. Empirical results show that SI-PDPO consistently outperforms existing primal-type algorithms.
PaperID: 3397, Poster
Abstract: Imitation learning policies for robotic manipulation inherently suffer from compounding errors. To mitigate this, inference-time policy steering introduces predictive reasoning to evaluate and refine action proposals before execution. However, existing methods typically anticipate consequences in unstructured pixel or latent spaces, leading to physically inconsistent predictions and unreliable error detection. Furthermore, their blindness to global progress often yields locally plausible corrections that fail to complete the overall task. To address these limitations, we propose Chain-of-Correction (CoC), an inference-time steering framework driven by a single VLM. Shifting away from unstructured state predictions, CoC explicitly leverages anchor-grounded scene graphs to track exact 3D physical relations. Built upon this representation, a progress-aware dual-check mechanism verifies historical execution and predicts future task advancement to systematically intercept myopic action proposals. Through a failure-augmented supervised fine-tuning pipeline, CoC seamlessly unifies scene graph extraction, progress evaluation, and precise kinematic correction, eliminating the need for task-specific world models. Extensive experiments across the COLOSSEUM benchmark, a newly introduced offline protocol for predictive reasoning, and real-world tasks demonstrate that CoC significantly outperforms state-of-the-art baselines in fine-grained error detection and robust error recovery.
Abstract: Generating physically buildable brick structures from 3D shapes requires more than geometric reconstruction: the output must also satisfy discrete part constraints and structural stability. Existing brick generation methods either rely on heuristic optimization, which can break down when the target 3D shape does not admit a feasible structure under predefined constraints, or generate brick sequences without explicitly modeling the underlying 3D geometry and assembly relations. In this work, we present BrickAnything, a geometry-conditioned autoregressive framework for generating buildable brick structures from diverse 3D representations. BrickAnything uses point clouds as a unified geometric interface and predicts brick sequences that reconstruct the target shape under assembly constraints. To model structural dependencies among bricks, we introduce a structure-aware tree tokenization, which represents brick structures through local attachment relations. This formulation makes sequence generation more consistent with the physical construction process, and reduces invalid intermediate states. We further introduce preference-based alignment post-training, validity-constrained decoding and adaptive rollback to improve buildability objectives such as stability and geometric fidelity. Extensive experiments demonstrate that BrickAnything produces geometrically faithful and physically realizable brick structures, and that the proposed tokenization effectively reduces rollback and regeneration compared with conventional ordering strategies.
PaperID: 3399, Poster
Abstract: Three-dimensional Gaussian Splatting (3DGS) has recently become a promising representation for dense SLAM, due to its explicit scene parameterization, efficient differentiable rendering, and high-quality visual reconstruction. Extending 3DGS to semantic SLAM enables the construction of maps that are not only geometrically accurate but also semantically interpretable. However, existing semantic Gaussian SLAM methods usually optimize semantic, depth, and normal renderings through separate supervision objectives. Such a design underexplores the fact that these outputs are generated from the same Gaussian field and should therefore provide a concordant explanation of the underlying scene structure. This limitation is especially problematic around object boundaries and geometric discontinuities, where inaccurate Gaussian growth can cause cross-boundary expansion, boundary blurring, and semantic leakage. To address this problem, we propose Concord-SLAM, a boundary-native semantic Gaussian SLAM framework that establishes concordance among semantic, depth, and normal renderings while promoting boundary-aware Gaussian growth. Our framework contains two key components. First, Cross-Render Concordance (CRC) jointly examines semantic, depth, and normal views rendered from the same Gaussian map, and converts their structural discrepancies in boundary responses, regional continuity, and local geometry into optimization signals. In this way, disagreement among different renderings is used as an intrinsic cue for improving the semantic-geometric coherence of the Gaussian field. Second, Boundary-Native Splatting (BNS) actively inserts new Gaussians near semantic boundaries, depth discontinuities, and salient geometric variations, encouraging primitives to align with physical and semantic boundaries during map construction rather than correcting boundary mixing only through later optimization. Experiments on Replica and ScanNet demonstrate that Concord-SLAM achieves superior performance in camera tracking, scene reconstruction, and semantic mapping.
Abstract: Long-context DNA models are limited by token-mixing cost and by how compression allocates representational budget across the genome. Existing approaches operate close to base-pair resolution, apply fixed downsampling, or learn content-dependent chunks without an explicit genomic budget, making long-context pretraining expensive and difficult to control. We introduce GeneZip, a region-aware DNA compression framework that combines H-Net-style dynamic routing with a Region-Aware Ratio (RAR) objective and bounded routing. GeneZip uses static gene-structure annotations during compression training to specify region-wise base-pairs-per-token (BPT) targets; at inference time, it compresses raw unseen DNA without annotations. GeneZip provides three main benefits. First, it is effective: GeneZip variants achieve the best validation PPL among encoder-based compressors, with GeneZip-70M already operating at 137.6 BPT, and across four reproducible DNALongBench tasks---contact map prediction, eQTL prediction, enhancer--target gene prediction, and transcription-initiation signal prediction---GeneZip obtains the best average rank among compared sequence models. Second, it is redundancy-aware: a post-hoc RepeatMasker/TRF analysis shows that, without repeat supervision, GeneZip assigns higher local BPT to TE-derived interspersed repeats and tandem repeats, two major classes of repetitive DNA sequence redundancy. Third, it is efficient: by reducing the effective token-mixing length, GeneZip enables longer-context and larger-capacity pretraining, including 128K-context and 636M-parameter variants on a single A100 80GB GPU, and fine-tunes the eQTL task 50.4× faster than JanusDNA (50 vs. 2520 minutes). These results establish GeneZip as an effective, redundancy-aware, and efficient compression interface for long-context DNA modeling.
PaperID: 3401, Poster
Abstract: Learning a response law for continuous decisions from observational time series is central to many real-world decision systems, yet temporal non-stationarity poses a fundamental challenge: outcome levels drift over time, causing naive models that fit raw outcomes to conflate treatment effects with shifting baseline conditions. Even when the one-step response is identifiable, unrestricted cross-time generalization remains fundamentally ill-posed. Thus, we introduce a constrained decision target based on a causal action geometry with anchored link-scale morphologies. Our approach constructs anchored contrasts over a family of link functions, isolating the action-dependent component of the response that can be shared across periods while allowing flexible, period-specific calibration. This yields a morphology-restricted oracle target for continuous decision making under temporal drift. Building on this formulation, we develop a two-stage orthogonal estimator that first obtains cross-fitted pilot estimates of anchored responses and then performs honest profiled selection over morphology and representation classes. The resulting estimator admits oracle guarantees for the restricted target and, under approximate persistence, produces decision rules that remain near optimal in future periods. Experiments on synthetic data and a large-scale e-commerce pricing application support the proposed target and its decision benefits under temporal shift.
PaperID: 3402, Poster
Abstract: Mechanisms for allocating a divisible resource among strategic agents have been widely studied. The prominent paradigm is the proportional (Kelly) mechanism, which elicits a scalar bid per agent, allocates the resource proportionally, and charges payments equal to the bids. Follow-up mechanisms improve social welfare, but sacrifice simplicity by introducing complex allocation rules or unintuitive payments. We introduce a unified framework for designing simple resource allocation mechanisms with proportional-style allocations and uniform pricing. Our framework yields a family of mechanisms that interpolate between the Kelly mechanism and the first-price auction. These mechanisms strictly improve upon Kelly’s efficiency guarantees, even achieving full efficiency in equilibrium, while also providing revenue guarantees relative to the VCG mechanism. In settings with budget-constrained agents, our mechanisms achieve nearly optimal price of anarchy with respect to the liquid welfare. When budgets are public, our framework further enables the design of mechanisms that ensure budget feasibility, even when agents do not explicitly account for their constraints in their strategies.
PaperID: 3403, Poster
Authors: Wen Dong
Abstract: Modern ML systems increasingly use uncertainty to decide when to abstain, route to a fallback, retrieve more context, or spend additional computation. Yet many strong Bayesian deep-learning recipes make uncertainty expensive on the serving path: ensembles, posterior samples, and MC dropout require repeated forward passes, while fast deterministic confidence heads are not generally posterior-predictive computations over parameter beliefs. We introduce \emphBayesify, a message-passing view that turns a deterministic neural computation graph into a single-pass posterior-predictive moment graph. Trainable tensors become Gaussian beliefs, activations carry selected moments, and prediction is one deterministic forward propagation of those moments. Theoretically, we show that exact belief propagation is natural-gradient message passing in exponential-family mean space, and that projected EP is implemented by reverse-mode adjoints followed by local Fisher moment updates. This yields a practical wrapper for dense, convolutional, residual, and adapter modules. Experiments evaluate calibrated uncertainty, selective prediction, and OOD behavior at one-pass inference cost, comparing against deterministic calibration, ensembles, dropout, Laplace, SWAG, and variational baselines.
PaperID: 3404, Poster
Abstract: Hosted LLM services increasingly expose model identity as a product claim, but external users and auditors often cannot inspect which model configuration served a request. We study an auditor-assisted regime in which a trusted auditor enrolls claimed identities and prepares verification material before deployment, while audit-time verification uses only the claimed identity and released output, without runtime access to model internals, inference logs, routing decisions, or provider metadata. We propose SumMark, a summary-channel watermarking framework that uses user-visible reasoning summaries as a keyed audit surface. During setup, SumMark constructs identity-specific vocabulary carriers, trains lightweight summary adapters, and calibrates per-identity thresholds. At deployment time, a keyed statistical test assesses whether the released summary is consistent with the claimed identity and, when needed, attributes the most likely enrolled identity. Across model families, scales, and task domains, SumMark reliably detects cross-family and cross-scale substitutions at low false-positive rates while preserving task-facing answer quality. We further study FP16\rightarrowINT8/INT4 precision-consistency diagnostics and robustness under bounded post-processing, rewriting stress tests, and text-only spoofing, revealing both strong performance and clear failure boundaries.
PaperID: 3405, Poster
Abstract: Recommender systems are central to modern digital platforms but remain vulnerable to popularity bias, where head items dominate exposure while many long-tail items are underrepresented. Existing long-tail methods often rely on global reweighting or opaque ranking models, making it difficult to personalize tail exposure without degrading recommendation utility. To address this challenge, we propose AgentTailor, an interpretable LLM-assisted framework for personalized long-tail recommendation. AgentTailor decomposes recommendation into two complementary stages. First, a semantic recall agent retrieves a semantically aligned candidate pool, determining whether relevant long-tail items are accessible to later ranking. Second, a conditional tail-specialist branch performs evidence-based re-ranking through a dual-gate tail-pressure router: the activation gate decides whether LLM-assisted evidence assessment should be invoked, while the placement gate decides whether the extracted evidence is allowed to modify the final ranking. Activated users receive item-level relevance and tail-fit evidence for routed candidates, but the final list is changed only when stable placement is warranted; otherwise, the Stage 1 order is preserved. Thus, the LLM provides auditable evidence rather than freely generating the final ranking. Experiments on MovieLens-1M and BookCrossing show that AgentTailor improves long-tail recommendation while maintaining competitive overall performance under a controlled LLM re-ranking protocol. Further analyses demonstrate its cross-backbone generalization, reliable structured outputs, and controllable trade-off between relevance preservation and long-tail exposure. The code and data are available at https://anonymous.4open.science/r/AgentTailor-CAEF.
PaperID: 3406, Poster
Abstract: Sparse activation accelerates Large Language Model (LLM) decoding by removing redundant computations and memory access, but existing methods often assume hidden-state dimensions are independent and identically distributed (i.i.d.). We replace this i.i.d. perspective with a covariance-aware view by analyzing contextual dependencies across tokens and inter-dimensional correlations within hidden states. We introduce Normalized Sparse Activation (NorSA), a training-free framework that combines dynamic thresholding with decorrelating rotation. Across the LLaMA, Mistral, and Qwen families, NorSA consistently improves perplexity and downstream accuracy over prior training-free baselines, especially in high-sparsity regimes. On LLaMA3-8B at 50% sparsity, it stays within 0.44 perplexity points of the dense model while achieving a 1.32× end-to-end decoding speedup.
Abstract: Flow Matching has recently emerged as a popular class of generative models for simulating a target distribution \mu_1 from samples drawn from a source distribution \mu_0. This framework relies on a fixed coupling between \mu_0 and \mu_1, and on a deterministic or stochastic bridge to define an interpolating process between the two distributions. The time marginals of this process can then be approximately sampled by estimating the transition rates, or more generally the generator, of its Markovian projection. This framework has recently been extended to the case of discrete source and target distributions, under the name Discrete Flow Matching (DFM). However, theoretical guarantees for such models remain scarce. In this paper, we study two DFM models on \mathbbZ_m^d = \\0,\ldots,m-1\\^d, sampled through time discretization, and derive non-asymptotic associated bounds for both of them. In contrast to previous work, we establish non-asymptotic bounds in Kullback--Leibler divergence for the early-stopped version of the target distribution. We also derive explicit convergence guarantees in total variation distance with respect to the true target distribution. Importantly, these bounds rely only on an approximation error assumption, relaxing standard score assumptions used in earlier works, while also yielding improved dependence on the vocabulary size m and the dimension d.
PaperID: 3408, Poster
Authors: Kelvin J Koa, XinYang Li, Ke-Wei Huang
Abstract: In this work, we study portfolio optimization under the stochastic discount factor (SDF) framework by learning market state representations that capture the underlying risk structures of financial data. This is challenging due to several factors: financial markets exhibit non-stationary dynamics with shifting regimes, multimodal inputs such as price and news data often contain stochastic noise, and existing diffusion-based approaches, while effective for modeling stochastic dynamics, rely on assumptions such as isotropic Gaussian noise that fail to capture the state-dependent nature of financial uncertainty. To address these challenges, we introduce RADAR, a retrieval-augmented diffusion framework that learns market representations by conditioning on similar historical regimes. RADAR leverages retrieval to construct context-dependent noise distributions, applies conditional diffusion to denoise multimodal representations, and initializes the diffusion process using empirical statistics to reflect state-dependent uncertainty. Experiments show that RADAR achieves state-of-the-art performance on key risk-adjusted metrics while producing economically meaningful signals on asset returns and correlations.
PaperID: 3409, Poster
Authors: Anupama Sridhar, Jack Shi, Ravi Deshpande, Jack Li, Alexander R Johansen, Michael P Snyder
Abstract: Many sequential Bayesian methods maintain a whitened operator A_w = L^-1ML^-\top, where L is the Cholesky factor of a kernel, precision, or curvature matrix. When the underlying matrix changes by a rank-1 perturbation, existing techniques propagate only one-sided quantities L^-1B; refreshing the sandwiched operator A_w requires a full O(r^3) re-factorization that dominates iteration cost, consuming up to 91% of per-step time in natural gradient methods. We show that this cost is unnecessary. For pre-whitened inputs, intermediate factors of L cancel exactly from both sides of the augmented Givens rotations, and the updated (L^+)^-1M^+(L^+)^-\top can be recovered in O(r^2) time by applying the rotations double-sidedly to a zero-padded matrix, followed by a rank-1 correction. We provide a self-contained derivation, stability bounds, and implementations as a C++ extension, MLX kernel, and Triton GPU kernel. On streaming sparse GP regression, online Bayesian optimization on protein fitness landscapes, and natural gradient preconditioning, the method matches full re-factorization in accuracy with up to 5.4× speedups. At r = 1024, training time drops from 1075s to 198s, making larger curvature charts affordable and producing strictly better test MSE.
PaperID: 3410, Poster
Abstract: Pre-trained models are routinely post-trained with reinforcement learning (RL) to maximize specific rewards, such as human alignment, correctness, or instruction following. This post-training process is computationally intensive, sometimes unstable, and has to be run from scratch every time the reward model changes or when we want to combine multiple rewards. We hence ask: given a new reward function, is it possible to predict the RL outcomes without actually running RL on it? We answer this in the affirmative by introducing PoEM, a framework to predict the outputs of RL on a new reward function using a set of models already post-trained on other rewards. First, we show that if the new reward function can be written as a linear combination of existing ones, then the new policy in log-space can be written as a linear combination of the existing log-policies. Surprisingly, even in cases where the rewards are not linearly connected, we observe that often log-policies from RL training span an approximately low-rank subspace across rewards. To our benefit, the weighting coefficients for this combination can be estimated using only the reward or basis policy outputs on the samples. We turn these observations into an algorithm that takes post-trained models and a new reward function, and approximates the target RL policy without actually running any additional RL training. We experimentally validate our approach across synthetic and real experiments, spanning both text and image modalities.
Abstract: Dynamical systems are fundamental to modeling the natural world, yet modeling them involves a persistent trade-off: manually prescribed mechanistic models are interpretable by design but often overly simplistic and misspecified; in contrast, flexible data-driven neural methods lack physical insight. Hybrid modeling aims for the best of both worlds by combining a prescribed or symbolic, physics-based component with a flexible neural network. A critical challenge, however, is that the neural component may relearn mechanistic parts, yielding redundant and uninterpretable models, especially when the symbolic structure itself is discovered from data. Existing methods based on standard L^2 regularization rely on a projection argument that breaks when the symbolic component is learned through sparse discovery, allowing the neural augmentation to overlap with symbolic structure. We introduce OrthoReg (Orthogonal Regularization), which directly penalizes overlap between the symbolic and neural components, preventing symbolic structure from being absorbed by the neural residual. This yields a complementary decomposition: the symbolic part captures what the library can express, and the neural part captures what remains. On benchmark dynamical systems with partial library mismatch, OrthoReg improves symbolic recovery and out-of-distribution behavior.
Abstract: Safety-aligned language models are trained to refuse harmful requests, yet refusal behavior can be suppressed by steering their internal representations. Existing methods do so by ablating a refusal direction from model activations, aiming to remove refusal from the model’s residual stream. Despite their empirical success, these methods lack a principled account of the latent-space transformation they induce and why it suppresses refusal. In this work, we recast refusal suppression as a latent-space evasion attack against linear probes trained to separate refused from answered prompts. Under this view, prior work’s difference-in-means direction naturally defines such a probe, and its ablation is exactly a projection onto its decision boundary, i.e., a minimum-confidence evasion attack. This perspective not only explains the empirical success of prior work but also admits a key limitation: evasion stops at the decision boundary, motivating the need to push representations further into the compliant region, i.e., where the model answers. We leverage this by proposing a Controlled Latent-space Evasion attack that projects representations past the boundary with an optimized confidence. We achieve state-of-the-art attack success rate across 15 instruction-tuned, multimodal, and reasoning models, outperforming existing refusal-ablation baselines and specialized jailbreak attacks.
PaperID: 3413, Poster
Abstract: Autoregressive (AR) models achieve strong performance in visual generation, yet their prefix-based factorization imposes a rigid one-dimensional dependency structure. We show that this structure is fundamentally suboptimal: any predefined ordering necessarily induces a trade-off between contextual sufficiency and necessity, resulting in both missing and redundant dependencies. To address this limitation, we propose Dynamic Causal Structure Discovery, an inference-time framework that replaces static conditioning with dynamic, content-dependent causal parent sets. By interpreting the Transformer as a causal graph, we reformulate generation as the problem of constructing a minimal parent set for each token. To make this objective tractable, we approximate \mathrmdo-interventions as vector-space ablations and derive a Fisher-based criterion that characterizes and corrects structural deviations during decoding. Our framework transforms AR factorization into an order-agnostic dependency structure, yielding improved generation quality and global coherence without retraining.
PaperID: 3414, Poster
Authors:
Arash Jamshidi, Katsiaryna Haitsiukevich, Aristides Gionis, Kai PuolamäkiAbstract: Estimating curvature information from gradients alone is a fundamental problem in machine learning. The challenge becomes particularly acute in settings where only aggregate mini-batch gradients are available, yet curvature information is still required for applications such as preconditioning in optimisation, sampling, and early stopping. Perhaps surprisingly, we show that a simple target-perturbation scheme suffices to recover such curvature information from gradients alone. Our method injects zero-mean noise into the targets, with variance calibrated to the local output curvature. The covariance of the resulting noisy batch gradients then recovers practical curvature matrices, without requiring per-sample gradients, second-order derivatives, or other internal network quantities. For a broad class of non-linear models, our scheme recovers the Generalised Gauss--Newton (GGN) matrix; for feed-forward ReLU networks, it additionally recovers the diagonal blocks of the population Hessian. We demonstrate that the resulting estimator is simple, lightweight, and easy to integrate into existing training pipelines. Experiments demonstrate accurate curvature recovery, faster optimisation when the estimator is used for preconditioning, and the utility of the estimated GGN for early stopping.
Authors: Xiaoxiao Liang, Juyuan Zhang, Liming Pan, Linyuan Lü
Abstract: Inferring latent interaction structures from observed dynamics is a fundamental inverse problem in many-body interacting systems. Most neural approaches rely on black-box surrogates over trainable graphs, achieving accuracy at the expense of mechanistic interpretability. Symbolic regression offers explicit dynamical equations and stronger inductive biases, but typically assumes known topology and a fixed function library. We propose COSINE (Co-Optimization of Symbolic Interactions and Network Edges), a differentiable framework that jointly discovers interaction graphs and sparse symbolic dynamics. To overcome the limitations of fixed symbolic libraries, COSINE further incorporates an outer-loop large language model that adaptively prunes and expands the hypothesis space using feedback from the inner optimization loop. Experiments on synthetic systems and large-scale real-world epidemic data demonstrate robust structural recovery and compact, mechanism-aligned dynamical expressions. Code: https://anonymous.4open.science/r/COSINE-6D43.
PaperID: 3416, Poster
Abstract: Macro placement is a critical early-stage decision in physical design because macro locations strongly constrain subsequent implementation stages. The task is difficult because these high-impact discrete decisions must be made before downstream quality can be measured accurately, forcing training to rely on fast but imperfect training-time oracles such as half-perimeter wirelength (HPWL). Existing RL-based placers therefore mostly optimize single HPWL-driven solutions, and their long sequential horizons often require heavy search assistance or offline expert data while leaving little room to preserve multiple competitive layouts. We address this gap by making quality-controlled diversity an explicit objective for macro placement. Instead of relying on a single proxy-optimal solution, we maintain a compact archive of HPWL-competitive yet meaningfully different placements for downstream selection. We realize this idea in DiversePlace, a curriculum proximal policy optimization (PPO) framework that combines a quality-gated diversity reward with an expanding placement horizon, enabling from-scratch RL optimization while keeping exploration inside competitive HPWL regions. On the ISPD2005 benchmarks, DiversePlace consistently improves top-5, top-10, and top-20 archive diversity and achieves lower or virtually identical HPWL on all eight designs with better efficiency. Our code is available at https://anonymous.4open.science/r/NIPS2026-DiversePlace-A615.
Abstract: Modern language models often need to optimize a primary accuracy objective while also accommodating secondary behavioral preferences, such as verbosity, agreeableness, or the level of technical expertise in its response. In practice, a base model may exhibit a desired behavior very rarely or not at all. Thus, endowing the model with a target behavior creates a sparse behavioral reward bottleneck. To address such multi-objective problems, we introduce Vector-Steered Policy Optimization (VSPO) which employs a steering vector associated with the target behavior to control the behavior intensity of the generated rollouts. VSPO is obtained by modifying GRPO to sample rollouts with varying steering intensities. This process can be interpreted as an on-policy latent self-distillation procedure where the model internalizes its steering vector. By varying steering intensities, VSPO upsamples rare behaviors and enriches rollout diversity, which alleviates the sparse reward issue and provably accelerates the policy optimization. Through comprehensive theory and experiments, we establish that VSPO has favorable properties compared to vanilla reward shaping and other alternative approaches. Specifically, under a bandit abstraction, VSPO provably achieves better iteration complexity than reward-shaped GRPO when the steering-induced distributions are sufficiently aligned with the target behavior. We evaluate VSPO across multiple reasoning benchmarks, including MATH and MMLU-Pro, for four target behaviors: explanation expertise, confidence expression, robustness to misleading context, and response verbosity. Our results show that VSPO consistently improves the control along target behavior while maintaining or improving task accuracy compared with reward shaping, teacher-trace distillation, and guidance-based baselines.
Authors: Joon Suk Huh
Abstract: Modern data workflows are inherently adaptive, repeatedly querying the same dataset to refine and validate sequential decisions, but such adaptivity can lead to overfitting and invalid statistical inference. Adaptive Data Analysis (ADA) mechanisms address this challenge; however, there is a fundamental tension between computational efficiency and sample complexity. For T rounds of adaptive analysis, computationally efficient algorithms typically incur suboptimal O(\sqrtT) sample complexity, whereas statistically optimal O(\log T) algorithms are computationally intractable under standard cryptographic assumptions. In this work, we shed light on this trade-off by identifying a natural class of data distributions under which both computational efficiency and optimal sample complexity are achievable. We propose a computationally efficient ADA mechanism that attains optimal O(\log T) sample complexity when the data distribution is dense with respect to a known prior. This setting includes, in particular, feature-label data distributions arising in distribution-specific learning. As a consequence, our mechanism also yields a sample-efficient (i.e., O(\log T) samples) statistical query oracle in the distribution-specific setting. Moreover, although our algorithm is not based on differential privacy, it satisfies a relaxed privacy notion known as Predicate Singling Out (PSO) security (Cohen and Nissim 2020). Our results thus reveal an inherent connection between adaptive data analysis and privacy beyond differential privacy.
Abstract: We introduce Contrast Sensitive Flow (CSFlow), a weighting scheme that connects the human eye's Contrast Sensitivity Function (CSF) to the iterative denoising steps of flow matching. Because real-world images concentrate signal at low spatial frequencies, these components reach high signal-to-noise ratio earlier during continuous diffusion than high-frequency components. When generating images with diffusion or flow matching models, this induces a soft autoregressive structure in Fourier space, where coarse image content stabilizes before fine detail. Meanwhile, the human visual system is unequally sensitive to spatial frequencies: very low and very high frequencies require significantly higher contrast to be perceived. We for the first time merge these observations through two contributions: (1) a metric that estimates which frequencies are generated at each reverse flow interval and (2) timestep weights obtained by aligning the frequencies generated at each noise level with human contrast sensitivity. We validate our contributions experimentally showing that these weights can improve generative performance by lowering FID by 4.7%, increasing Inception Score by 2.2% and improving GenEval scores by 2.5% using inference-only timestep modification or short fine-tuning. Qualitatively, we find that our CSFlow weights lead to better visual realism and less cartoonish appearance of generated images.
PaperID: 3420, Poster
Authors: Mark Junjie Li, zhishun liu, Wei Wang, Yuanshan Lu, yangjingxi, Xianyu Bao, Jun Li, Sunjie Huang
Abstract: Large Language Models (LLMs) often generate factually unsupported content known as hallucination. However, existing detection approaches rely on handcrafted heuristics or isolated attention signals, limiting their ability to capture higher-order dependencies. In this work, we propose AttnHyper, a framework that represents a Transformer's attention patterns as a hypergraph for hallucination detection. Tokens are treated as nodes, while hyperedges connect groups of tokens that co-occur across heads and layers, with features derived from attention weights to preserve higher-order interactions beyond pairwise graphs. Based on this representation, the method formulates hallucination detection as a hypergraph learning problem and employs a hypergraph neural network (HGNN) to extract structural cues. It consistently outperforms state-of-the-art methods on the RAGTruth benchmark, achieving an average gain of +5.08 AUROC over the strongest baseline across diverse settings and architectures, while also demonstrating strong zero-shot transfer. Importantly, it incurs minimal overhead at inference time, enabling efficient deployment. Our results demonstrate that hypergraph-based attention modeling provides a more expressive and reliable signal for hallucination detection.
PaperID: 3421, Poster
Abstract: Multiview learning integrates complementary information from diverse observations to enhance model performance, where to quantify predictive uncertainty evidential deep learning has recently gained significant attention. However, existing methods generally overlook the critical importance of high-order correlation structures in multiview features, and traditional evidential combination rules frequently suffer from issues such as order dependence and irrational belief assignments when handling highly conflictive multiview features. To address these issues, we propose the Hypergraph-guided Global Mean-field Negotiation (HGMN) method for Multiview Evidential Classification. Specifically, HGMN first leverages hypergraph convolutional networks to capture high-order topological correlations within each view. Subsequently, instead of relying on pairwise or hierarchical fusion strategies, HGMN introduces a global mean-field negotiation mechanism that enables multiple views to dynamically reach a global consensus, effectively isolating dissenting noise. Finally, HGMN incorporates a multi-objective collaborative optimization strategy that enhances decision robustness and trustworthiness in complex open-world environments. Extensive experimental results on eight public datasets demonstrate that our method significantly outperforms state-of-the-art baselines in terms of accuracy and robustness.
PaperID: 3422, Poster
Abstract: Reinforcement learning (RL) from raw sensory inputs such as images or onboard sensors can be challenging due to high-dimensional observations and sparse rewards. A common strategy is to train a teacher policy with access to privileged state information and then distill it into a student policy that acts from raw inputs alone. However, when the teacher relies on information unavailable to the student, exact imitation may be impossible, creating an irreducible imitation gap. Existing approaches typically mitigate this mismatch through careful reward shaping or additional reinforcement learning on the student policy after the teacher has been trained. Instead, we introduce Representation-Aligned Imitation Learning (RAIL), which addresses the mismatch during teacher training itself. RAIL learns a latent representation shared across teacher and student observations using contrastive learning, and trains the teacher policy directly in this space. Because the teacher policy is constrained to operate on this shared space, it learns behaviors that are reproducible from student observations. This substantially reduces the imitation gap while preserving task performance. Across multiple environments, RAIL outperforms strong baselines without reward modification or post-hoc student fine-tuning, and surpasses direct reinforcement learning from raw observations. The learned representations further enable zero-shot transfer to new tasks through teacher-only training.
Authors: Haoran Zhang, Zisu Dong, Feng Zhou
Abstract: Transformer models rely on attention mechanism to capture long-range dependencies but suffer from quadratic complexity, limiting their scalability to long sequences. Kernel-based linear attention reduces this complexity but typically relies on fixed or weakly learnable kernels, restricting expressiveness and performance. In this work, we propose Flexformer, a flexible linear Transformer that learns attention kernels in a fully data-driven manner. Flexformer builds on random Fourier feature-based linear attention and treats spectral frequencies as trainable parameters, enabling the model to learn a broad family of attention kernels. We develop both stationary and nonstationary variants, with the latter offering strictly greater expressiveness. Extensive experiments on language modeling and sequence classification demonstrate that Flexformer consistently outperforms baselines. Moreover, Flexformer can be effectively distilled from pretrained Transformers to recover softmax attention and exhibits strong kernel transferability across domains, achieving both high efficiency and competitive performance on long-sequence tasks.
PaperID: 3424, Poster
Authors:
Ionuț Hodoroagă, Andrada Gobeajă, Marius Leordeanu, Elena BurceanuAbstract: Can a dataset be recognized from the spurious correlations it induces during training? We argue that datasets leave dataset-specific traces in a model’s learned semantic correlation structure: incidental regularities that are predictive within a dataset, but not causal for the underlying task, can be internalized during training. We use this insight to study dataset-level membership inference, moving beyond existing methods that rely on behavioral or distributional evidence such as confidence scores, losses, margins, generated samples, or query responses. We introduce a white-box semantic fingerprinting approach based on semantic correlation descriptors (SCDs), which capture the semantic correlation structure learned by a model and make it comparable across dataset mixtures. In a controlled leave-one-dataset-out diagnostic, SCDs recover dataset-specific changes and perfectly separate matching from non-matching dataset pairs. We then propose a practical SCD-based membership score that tests whether a target dataset is part of a model’s training mixture using only the model’s SCD and the target dataset’s standalone SCD, without requiring leave-one-dataset-out models. Across three diverse experimental settings, with dataset groups for natural language inference, emotion classification, and medical text classification, we test both the advantages and limitations of SCD-based membership inference with different degrees of semantic separation and keyword support between dataset splits. On average, the classifier id_\mathrmSCD based on this score achieves the highest performance and the lowest std, outperforming black-box baselines RMIA, Attack-P, and LiRA, as well as the white-box SIF baseline. These results show that dataset membership can be traced through internal semantic correlations, with the largest relative gain exceeding 60% in ROC-AUC when dataset groups expose distinct semantic particularities.
Abstract: Continuous-token autoregressive (AR) generation directly models continuous signals without discretizing them into vocabulary tokens, but suffers from a severe train--inference mismatch: teacher-forced training conditions on ground-truth prefixes, whereas inference conditions on model-generated continuous tokens. This mismatch becomes especially severe in pixel-space image generation, where each token is a high-dimensional raw pixel patch and autoregressive errors can accumulate over generation steps. Rollout-based training can reduce this mismatch but is prohibitively slow, especially when each token is produced by a multi-step diffusion head. We propose \emphParallel Rollout Approximation (PRA), which approximates rollout-based training by constructing training inputs in parallel at all positions. PRA combines a shared one-step denoiser with an end-to-end learned low-dimensional intermediate state, aligning these constructed training inputs with inference-time generated outputs without relying on a separately pretrained tokenizer. On class-conditional ImageNet-1K generation at 256×256 resolution, even PRA-S with 135M parameters achieves an FID of 2.58, surpassing the previous billion-scale pixel-space AR result of 3.60. Scaling to PRA-L (511M) further improves FID to 1.94, establishing a new state of the art among pixel-space AR models and substantially narrowing the gap to pixel-space diffusion models. Beyond generation, PRA achieves higher ImageNet classification probing accuracy than latent-space AR baselines, suggesting its potential for unified pixel-space image generation and understanding.
Abstract: Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models, enabling parallel token generation while achieving competitive performance. Despite these advantages, MDMs face a fundamental limitation: once tokens are unmasked, they remain fixed, leading to error accumulation and ultimately degrading sample quality. We address this by proposing a framework that trains a model to perform both unmasking and correction. By reusing outputs from the MDM denoising network as inputs for corrector training, we train a model to recover from potential mistakes. During generation we apply additional corrective refinement steps between unmasking ones in order to change decoded tokens and improve outputs. We name our training and sampling method ) for its unique ability to iteratively refine an entire sequence, including already generated tokens. We conduct extensive experimental validation across multiple conditional and unconditional tasks, demonstrating that \method~yields better quality-efficiency trade-offs (up to ~4x faster sampling) and enables inference-time compute scaling to further increase sample quality beyond standard MDMs (up to ~1.2x improvement on benchmarks).
Authors: Ziyi Kou, Ankit Kumar, Mia Huang, Taylor Niehues, Vatsal Mehta, Ergys Ristani, Li Guan
Abstract: We present AVI-HT, an adaptive visual-IMU fusion approach for tracking 3D hand poses by jointly modeling the egocentric image with on-glove 6-DoF IMU signals. AVI-HT achieves significantly improved accuracy and availability, particularly in hand-object interaction (HOI) scenarios involving heavy visual occlusion. Two complementary ingredients underpin its success: (1) synchronized multi-modal training data pairing on-body vision-IMU sensor streams with ground-truth 3D hand poses from a motion-capture system, and (2) a cross-sensor deep attention mechanism that adaptively modulates the trust assigned to the vision and individual IMU sensors. To evaluate AVI-HT in real-world settings, we conduct extensive experiments on our DexGloveHOI dataset that consists of 100K+ pairwise vision-IMU samples with synchronized 3D annotated poses, in which users manipulate a variety of objects during daily tasks. We compare against multiple single- and multi-modal tracking approaches under two hand models (UmeTrack, MANO). The results show that AVI-HT reduces mean keypoint error by 16.1% and its wrist-aligned variant by 24.2% over the baselines. Ablation studies further reveal the per-finger contribution of IMU sensors across activity types, and the model's sensitivity to IMU noise and temporal misalignment in vision-IMU fusion.
PaperID: 3428, Poster
Abstract: Reinforcement learning (RL) with verifiable rewards has proven effective at post-training LLMs for coding, yet deploying separate task-specific specialists incurs costs that scale with the number of tasks, motivating a unified multi-task RL (MTRL) approach. However, existing MTRL methods treat all coding tasks uniformly, relying on fixed data curricula under a shared optimization strategy, ultimately limiting the effectiveness of multi-task training. To address these limitations, we propose \tool, a multi-tASk code reinforcement learning framework via uTility-driven coORdination. Centered on task utility, a signal capturing each task's learning potential and cross-task synergy, \tool comprises two coupled modules: 1) \moduleone module hierarchically allocates training budget and prioritizes informative prompts, steering training toward the most valuable data; and 2) \moduletwo module dynamically scales per-task KL regularization, matching update constraints to each task's current training state. Experiments on two widely-used LLMs across four representative coding tasks demonstrate that \tool consistently improves a single model across all tasks, outperforming the best task-specific specialist by 9.0%--9.5% and surpassing the strongest MTRL baseline by 7.5%--12.8%.
Authors: Vladimir Zaigrajew, Michał Piechota, Gaspar Sekula, Paweł Gelar, Przemyslaw 'Prem' Biecek
Abstract: Interpreting the concepts encoded by individual neurons in deep neural networks is a crucial step towards understanding their complex decision-making processes and ensuring AI safety. Despite recent progress in neuron labeling, existing methods often limit the search space to predefined concept vocabularies or produce overly specific descriptions that fail to capture higher-order, global concepts. We introduce LINE, a novel, training-free iterative approach tailored for open-vocabulary concept labeling in vision models. Operating in a strictly black-box setting, LINE leverages a large language model and a text-to-image generator to iteratively propose and refine concepts in a closed loop, guided by activation history. We demonstrate that LINE achieves state-of-the-art performance across multiple model architectures, yielding AUC improvements of up to 0.11 on ImageNet and 0.05 on Places365, while discovering, on average, 27% of new concepts missed by massive predefined vocabularies. Beyond identifying the top concept, LINE provides a complete generation history, enabling polysemanticity evaluation and producing supporting visual explanations that rival gradient-dependent activation maximization methods. The LINE source code is available on GitHub.
Abstract: Large Language Models (LLMs) rely heavily on Key-Value (KV) caching to minimize inference latency. However, standard KV caches are context-dependent: reusing a cached document in a new context requires recomputing KV states to account for shifts in attention distribution. Existing solutions such as CacheBlend, EPIC, and SAM-KV mitigate this issue by selectively recomputing a subset of tokens; however, they still incur non-negligible computational overhead (FLOPs) and increased Time-to-First-Token (TTFT) latency. In this paper, we propose KV Packet, a recomputation-free cache reuse framework that treats cached documents as immutable ``packets'' wrapped in light-weight trainable soft-token adapters, which are trained via self-supervised distillation to bridge context discontinuities. Experiments on Llama-3.1 and Qwen2.5 demonstrate that the proposed KV Packet method achieves near-zero FLOPs and lower TTFT than recomputation-based baselines, while retaining F1 scores comparable to those of the full recomputation baseline. Code and experimental results are available at https://anonymous.4open.science/r/kv
Authors:
Xintong Hu, Xuhong Huang, JINYU ZHANG, Yutong Yao, Yuchong Sun, Qiuyue Wang, Mingsheng Li, Yitao Liu, Yixuan Chen, Yingming Zheng, Shuai Bai, Tao YuAbstract: Vision-Language-Action (VLA) models have advanced rapidly by leveraging large-scale vision-language priors and robot trajectory data, yet they remain weak at following fine-grained human instructions. We argue that this limitation stems from a fundamental mismatch between language and action supervision in existing robot datasets. Most open-source datasets annotate each trajectory with a goal-oriented task description, while leaving the underlying execution process largely unspecified, including motion trajectories, contact and approach patterns, object state transitions, recovery behavior, and final configurations. Since different choices along these dimensions can lead to substantially different action sequences, coarse task descriptions induce a many-to-one mapping from action trajectories to language, preventing precise action-instruction alignment. To address this problem, we propose FineVLA, a framework for constructing and leveraging fine-grained embodied action instructions for VLA learning. FineVLA includes FineVLA-Tool, a scalable pipeline for data cleaning, temporal alignment, and fine-grained annotation; RoboFine-VLM, a vision-language model for fine-grained robotic action understanding and annotation; RoboFine-Bench, a benchmark for evaluating fine-grained robotic action understanding via VQA and captioning; and FineVLA-Policy, a VLA policy trained with fine-grained instructions on open-source ALOHA robot data. Across benchmark evaluation and policy learning experiments, FineVLA demonstrates that high-density, action-aligned language supervision leads to more controllable and instruction-sensitive robot policies. Our results highlight fine-grained instruction alignment as a critical step toward VLA systems that not only complete tasks, but complete them in the way humans specify.
PaperID: 3432, Poster
Abstract: Recent video generative models can produce photorealistic short clips, but struggle with long-horizon generation under dynamic camera control, often suffering from viewpoint drift, geometric inconsistency, and error accumulation. We propose Joint Appearance-Geometry Diffusion Transformer (JAG), a spatially-grounded causal video diffusion model that explicitly integrates 3D perception for long and consistent video generation. It adopts a novel dual-branch causal architecture that jointly models visual appearance and scene geometry. Building upon this formulation, we introduce Self Geometry Forcing, a geometry-aware distillation scheme that conditions current generation on self-predicted geometric structures as persistent long-term memory and imposes self-verified geometric rewards on intermediate rollouts, effectively mitigating error accumulation over long horizons. Extensive experiments show that JAG can generate long videos with superior coherent appearance and geometry at constant per-frame computation, improving long-horizon controllability, geometric consistency, and visual fidelity over previous state-of-the-art methods.
PaperID: 3433, Poster
Abstract: As Multimodal Large Language Models (MLLMs) continue to advance and find increasingly broad applications, a key challenge persists: the patch-based vision paradigms widely adopted by MLLMs inevitably introduce structural fragmentation when processing information-dense images such as tables and charts. Specifically, when a table row or chart axis spans multiple visual partitions, the lack of cross-partition communication causes the model to lose track of structural continuity, leading to performance fluctuations when layout changes --- a problem we term layout sensitivity. To address this fundamental limitation, we propose Patch-for-Patch (P4P), a lightweight module that can be plugged into existing patch-based vision encoders. P4P dynamically selects a sparse set of tokens as bridging anchors and uses them to establish cross-partition communication through bidirectional attention. Combined with a gated fusion strategy, P4P restores global structural coherence without the quadratic cost of global attention. Extensive experiments across Qwen2.5-VL, MiniCPM-V, and InternVL3.5 demonstrate that P4P consistently improves structured content visual understanding, notably boosting TableEval accuracy by up to 9.12% and significantly reducing layout sensitivity. The source code is available at: https://anonymous.4open.science/r/P4P/.
PaperID: 3434, Poster
Abstract: Modern automated heuristic discovery (AHD) systems increasingly rival hand-crafted heuristics by leveraging Large Language Models (LLMs) to iteratively mutate and refine solutions. A natural assumption is that more powerful and more expensive models consistently yield better heuristics. We challenge this empirically: across multiple NP-hard combinatorial optimization problems, we find no consistent advantage for larger, costlier models. This raises a deeper question: given a pool of LLMs spanning frontier and freely hosted or open-weight models, how should one adaptively select which model to invoke at each step of the search? The challenge is non-trivial. The optimal LLM may shift over the course of a run, and because one LLM's output seeds the context for subsequent calls, each model's reward signal is non-stationary and entangled across the heuristic population, thus violating the independence assumptions of classical online decision-making settings. We address this through Ballad (Bandit-based LLM Routing for Automated Heuristic Discovery), an online bandit algorithm that learns to orchestrate a pool of LLMs into an ensemble that matches or exceeds what any soloist LLM achieves alone. This orchestration is powered by an LLM selection policy that favors rare, high-scoring heuristics over stable averages, a DAG-based credit propagation mechanism through the genealogy of the heuristic population, and discounting of stale observations to remain attuned to each model's current utility. Across five NP-hard combinatorial optimization problems and a pool of four models of varying capability and cost, Ballad not only matches the best individual LLM in hindsight but, in several cases, surpasses every individual model while reducing API cost by up to 65.6% relative to always using the strongest model.
PaperID: 3435, Poster
Abstract: Fine-grained visual classification (FGVC) with Multimodal Large Language Models (MLLMs) remains challenging, as visually similar subordinate categories inherently induce classification uncertainty due to subtle inter-class differences. Existing reinforcement learning methods typically require the model to generate a single class name and optimize it with sparse correctness rewards. However, a seemingly incorrect prediction can still provide highly informative signals if it falls within the same local confusion set as the ground-truth class. By treating all non-exact matches as equally wrong, the rigid single-answer paradigm overlooks these valuable near-misses. Consequently, it fails to capture the model's intrinsic uncertainty and provides limited learning signals. To address these limitations, we propose List-GRPO, an uncertainty-aware listwise Group Relative Policy Optimization framework. Rather than enforcing a single prediction, our method prompts the model to adaptively generate a ranked candidate list, explicitly capturing its predictive uncertainty. Furthermore, to mitigate the advantage vanishing when online exploration completely misses the target, we introduce a dynamic ground-truth anchored response strategy. We synthesize an off-policy anchor by prepending the ground truth to the most frequent hard negatives mined from online samples. To prevent the anchored response from dominating learning, we isolate the advantage estimation of online samples from the anchored response and use a gating coefficient to modulate its contribution. Extensive experiments on multiple fine-grained benchmarks demonstrate that List-GRPO significantly improves Top-1 accuracy while maintaining high ground-truth coverage with only about three candidate predictions on average.
PaperID: 3436, Poster
Abstract: Importance sampling provides an elegant variance-minimization principle for stochastic optimization: sampling examples proportional to their gradient norms yields a minimum-variance gradient estimator. However, in modern deep learning, importance sampling often fails to outperform uniform sampling once proxy-computation cost, stale scores, and fair update budgets are accounted for. We propose LITE, a lightweight lazy sampler for practical importance sampling. LITE maintains inexpensive proxy scores, updates them only periodically, smooths them across refreshes, and converts them into sampling probabilities through a temperature rule that prevents over-concentration. When proxy scores are informative, LITE prioritizes high-impact examples; when they are noisy or stale, the sampling distribution moves back toward uniform. We analyze the resulting variance reduction in terms of proxy quality, temporal drift, and temperature, yielding an explicit condition for when LITE improves over uniform sampling. Experiments under matched update budgets and end-to-end wall-clock accounting show that LITE is competitive with or better than uniform sampling across vision, text, and 3D benchmarks, with the clearest gains in scarce-budget and imbalanced regimes.
PaperID: 3437, Poster
Abstract: Deep clustering struggles when coarse-grained and fine-grained semantics coexist in the same dataset, as a unified partitioning strategy tends to either over-separate coarse categories or miss subtle fine-grained differences. To address this problem, we propose the adaptive multi-granularity clustering framework in a hyperspherical space (MC-H). The framework combines Dynamic Prototype-Center Assignment (DPCA), a Hyperspherical Determinantal Point Process (H-DPP) module that adaptively regulates repulsion according to local structure and promotes stronger separation in dense regions while preserving flexibility in sparse ones, and a Semantic Subspace Denoising (SSD) module that suppresses non-semantic variations and improves cluster compactness. Extensive experiments across datasets with different semantic granularities show that our method consistently outperforms existing approaches and effectively unifies clustering across the full granularity spectrum.
PaperID: 3438, Poster
Abstract: To leverage the full potential of multimodal data, we need representations that go beyond the state-of-the-art alignment and fusion approaches and exploit all cross-modal interactions without sacrificing modality-specific information. Learning disentangled representations is a principled way to identify these underlying shared and unique factors that are hidden in observational data. However, while multimodal disentanglement is a compelling paradigm, existing methods are largely confined to the two-modality regime due to its inherent scalability bottleneck. To address this, we propose RePercENT, a self-supervised framework designed to surpass these limitations, that unlocks scalable pairwise disentanglement as we move beyond two modalities. Through a multimodal `plug-and-play' architecture, our approach operates directly on pre-extracted embeddings, eliminating the need for extensive joint pre-training while making no assumptions regarding the underlying modalities or foundation model backbones. Moreover, we introduce a joint optimization objective for simultaneously deriving the shared and unique components, and provide formal theoretical guarantees that characterize the optimality of our solution. Across diverse modalities and tasks, RePercENT successfully recovers disentangled components while maintaining competitive performance and significantly reducing computational complexity.
PaperID: 3439, Poster
Abstract: Autoregressive Diffusion Transformers have emerged as a powerful paradigm for long video generation. However, they are often constrained by prohibitive computational overhead and substantial memory footprints. This inefficiency arises from dense computation mechanisms that allocate uniform processing resources across all spatiotemporal regions, failing to leverage the inherent temporal redundancy of video where motion is typically confined to sparse areas. In this paper, we propose Dynamics-Aware Sparse Attention (DASA), a unified framework that shifts from dense computation to sparse, motion-driven generation to eliminate both storage and computational redundancies. Inspired by video codec principles, our approach explicitly decouples video content into dynamic and static regions. To mitigate memory bottlenecks, we selectively compress the historical KV cache by retaining only dynamic features within the generated chunk. Furthermore, to reduce computational overhead, we employ a region-level computation allocation strategy based on localized dynamics. Additionally, a dynamics-guided distillation method is introduced to enhance performance via a lightweight distillation training process. Experiments on VBench demonstrate that our approach accelerates autoregressive video generation by 2.16×, while preserving high visual fidelity.
PaperID: 3440, Poster
Abstract: Raman spectroscopy is often hindered by strong autofluorescence backgrounds, which are typically handled by waiting for fluorescence to bleach before recording the Raman spectrum. Here, we instead show that short, incomplete recordings of the bleaching process already contain enough information to recover the underlying Raman signal. We present the first unsupervised model that operates on short, incomplete, variable-length sequences of time-resolved spectra, decomposing them into Raman spectra and autofluorescence components while simultaneously removing measurement noise. Our VAE-based method simultaneously learns (i) a bank of fluorophore spectra comprising the baseline, (ii) a model of the measurement noise, and (iii) the distribution of Raman spectra and decay behavior. We evaluate our method on spectral reconstruction and downstream peak detection, where it outperforms state-of-the-art approaches on multiple simulated benchmarks as well as on a newly collected real-world dataset of spectral time series. We further show that fluorescence decay dynamics themselves contain discriminative information that may be exploited in future work.
PaperID: 3441, Poster
Abstract: Scientists often need to analyze the samples in a study that responded to treatment in order to refine their hypotheses and find potential causal drivers of response. Natural variation in outcomes makes teasing apart responders from non-responders a statistical inference problem. We introduce the causal two-groups (C2G) model, a causal extension of the classical two-groups model for inferring treatment effects. The classical two-groups model posits that treated samples may or may not experience an effect, according to some prior probability. The C2G model extends this to consider the case where treatment, latent effect status, and effect size are all confounded. We propose two empirical Bayes procedures for the causal two-groups model, one under semi-parametric conditions and another under fully nonparametric conditions. The semi-parametric model assumes additive treatment effects and is identifiable from observed data. The nonparametric model is unidentifiable, but we show it can still be used to test for response in each treated sample. We show empirically and theoretically that both methods for selecting responders control the false discovery rate at the target level with near-optimal power. We also propose two novel estimands of interest and provide a strategy for deriving estimand intervals in the unidentifiable nonparametric model. We apply the C2G model to a cancer immunotherapy dataset, where the nonparametric model recovers clinically-validated predictive biomarkers of both positive and negative outcomes.
Abstract: INT2 KV-cache quantization is attractive for long-context LLM serving, but it remains difficult to make both accurate and deployable. Simple rotations such as Hadamard transforms reduce outliers, but still degrade at INT2 because they are not aligned with downstream attention. We propose CoQuant, an Ultra-low-bit KV Cache quantization method that estimates structures offline and uses them to derive fixed rotations and clipping thresholds for quantization. In this way, it aligns KV quantization with the covariance structures that attention actually consumes. More importantly, we not only provide theoretical justification but also develop a fully deployable CoQuant system with a custom INT2 attention kernel that remains compatible with paged KV-cache serving and fused kernel pipelines, enabling seamless integration into modern LLM serving frameworks such as SGLang and vLLM. We evaluate our methods on recent reasoning models with reasoning traces of up to 32k tokens across 5 tasks. On Qwen3-4B-Thinking-2507 and Qwen3-8B, CoQuant reduces the BF16 accuracy gap to 3.78 and 1.42 points, respectively, while naive rotation INT2 collapses to nearly zero. We further scale CoQuant to Qwen3-32B and GLM-4.7 (358B params), where it remains effectively on par with BF16. On long context - RULER-NIAH up to 128K, CoQuant remains robust on both Qwen3 models, while naive rotation INT2 collapses. System-wise, CoQuant reduces KV-cache memory by approximately 8×, improves throughput by up to 7× at large batch sizes under the same memory budget, and accelerates batch-size-1 decoding by up to 3× over BF16 due to reduced memory bandwidth overhead.
PaperID: 3443, Poster
Authors: Kun Li, Sami S Brandt, Michael Yang
Abstract: Referring Audio–Visual Segmentation (Ref-AVS) is an emergent multimodal task requiring the precise identification and segmentation of a specific object based on a natural language expression grounded in both visual and auditory cues. Unlike traditional AVS, which segments generic sound sources, or referring video object segmentation (Ref-VOS), which relies solely on visual–textual grounding, Ref-AVS necessitates fine-grained tri-modal interaction across vision, audio, and language. However, existing methods often fail to maintain consistency when subjected to target shift, a phenomenon where the model erroneously transfers its focus between objects due to transient occlusions, acoustic fading, or identity ambiguities. To address this, we propose a novel Shift-Aware Identity-guided Latent (SAIL) refinement framework. It introduces a dual-token strategy that decouples frame-level dynamics from a global semantic anchor to explicitly model target-shift cues. We design a shift-aware refinement module that rectifies drift by aligning tokens with domain-specific evidence. Finally, we leverage a decoder that integrates the refined embeddings with a continuous memory update to maintain spatio-temporal coherence and identity consistency throughout the video sequence. Extensive experiments conducted on two Ref-AVS benchmarks demonstrate the effectiveness of our proposed method, significantly outperforming existing AVS, Ref-VOS, and Ref-AVS approaches. The code will be released after publication.
PaperID: 3444, Poster
Abstract: Diffusion-driven image editing and regeneration are becoming increasingly widespread. This creates an urgent need for watermarking methods that can embed imperceptible yet verifiable signals into existing images for ownership attribution and provenance tracking. Existing methods fail to simultaneously satisfy three practical requirements: robustness against diffusion-based transformations, low-cost verification, and scalable support for per-user key assignment at low overhead. To address these challenges, we propose PrivateSeal, a watermarking framework for pre-existing images under diffusion-based transformations. The perturbation is encouraged to lie along low-sensitivity latent directions, so that the embedded signal is less likely to be suppressed during diffusion-based editing or regeneration. During verification, the embedded message is reliably recovered via a simple latent-space projection using the corresponding key. This design allows a platform to assign independent projection keys to different users, accounts, or images without retraining or modifying the verifier, while maintaining low-cost verification. Extensive experiments on the W-Bench benchmark show that PrivateSeal achieves competitive robustness against diffusion-based regeneration and editing, with additional cross-model and cross-dataset evaluations further validating its strong transferability.
PaperID: 3445, Poster
Abstract: Generative policies have emerged as a compelling paradigm in reinforcement learning (RL) as they can represent complex and multimodal action distributions. However, this high expressivity leaves them particularly vulnerable to the inherent target distribution non-stationarity in RL, rendering the generative process severely volatile. To address this issue, we propose Look-ahEad vAriational Flow (LEAF), a generative policy optimization framework that stabilizes both the target and the generative process. First, we replace the action-value-guided target with a look-ahead target induced by predictive dynamics, which we theoretically prove to exhibit a tighter target drift bound. Second, we formulate policy improvement as a proximal variational free-energy optimization, which enforces a distributional trust region to prevent overfitting to transient target shifts. Finally, we prove that this optimization induces an optimal transport flow, which we instantiate as a two-phase coarse-to-fine generative process to achieve broad multimodal coverage and pinpoint accuracy. Extensive experiments on the DeepMind Control Suite and Humanoid Bench demonstrate that LEAF consistently outperforms strong baselines in both sample efficiency and asymptotic performance.
PaperID: 3446, Poster
Abstract: Learning from label proportions (LLP) is a weakly supervised setting where training data are grouped into bags with only bag-level label proportions observed, and the goal is to learn a classifier that predicts labels for individual instances. Recently, a large body of methods has emerged; however, most of them rely on the assumption that instances and labels are sampled i.i.d. and randomly assembled into bags, an assumption often violated in real-world scenarios. Our experimental results indicate that these methods perform suboptimally under non-random settings. In this work, we adopt a more realistic assumption: bags are sampled i.i.d., and instance labels are conditionally independent given the instances. Under this setting, we show that the label counts follow a Poisson multinomial distribution. Motivated by this observation, we propose LLP via moment matching(LLP-MM), a simple yet effective approach which leverages multi-order factorial moments as training objectives, encouraging the classifier predictions to match these moments and thereby more fully exploit the underlying statistical structure of the data. Extensive experiments on benchmark datasets under various bag construction strategies demonstrate the effectiveness of our approach while maintaining high computational efficiency.
PaperID: 3447, Poster
Abstract: Reinforcement learning has become a key approach for optimizing large language models (LLMs), with Group Relative Policy Optimization (GRPO) emerging as the most popular algorithm due to its simplicity and effectiveness. However, GRPO normalizes advantages purely within the current sampling group, ignoring the dynamic evolution of the policy's performance throughout training. This renders the model susceptible to relative deception, i.e., the model may be steadily deteriorating yet remain oblivious to its own regression. To address this issue, we propose Historical Relative Policy Optimization (HRPO), which integrates two synergistic designs: (i) a historical-aware advantage estimator that normalizes each response against the running maximum of the policy's strongest past group-mean rewards, thereby pushing the model to continuously surpass its own historical peak (bootstrapping); and (ii) an on-demand historical replay that recalls historical high-reward responses as positive examples, which is triggered only when all responses in the current step fail to provide positive signals. Experiments show that HRPO consistently outperforms GRPO variants in training stability, convergence speed, and final performance across diverse models and benchmarks, demonstrating that accounting for training dynamics leads to more reliable and effective policy optimization in LLMs.
Abstract: A major bottleneck in characterizing the failure modes of generative AI systems is the cost and time of annotation and evaluation. Consequently, adaptive testing paradigms have gained popularity, where one opportunistically decides which cases and how many to annotate based on past results. While this framework is highly practical, its extreme flexibility makes it difficult to draw statistically rigorous conclusions, as it violates classical assumptions: the number of observations is typically limited (often 10 to 50 cases) and decisions regarding sampling and stopping are made in the midst of data collection rather than based a pre-specified rule. To characterize what statistical inferences can be drawn from highly adaptive audits, we introduce a hypothesis testing framework from two 'dueling' perspectives: (i) the model's null that asserts there is no failure mode with performance below a target threshold versus (ii) the auditor's null that asserts they have a sampling strategy that will uncover a failure mode. Leveraging Safe Anytime-Valid Inference (SAVI), we formalize the auditor as conducting 'testing by betting', which translates into simultaneous e-processes for testing the dueling null hypotheses. Furthermore, the auditor is sufficiently powerful, we prove that these two hypotheses are asymptotically inverses of each other, in that passage of a stringent audit does in fact certify the AI system as being globally robust. Empirically, we demonstrate that our proposed testing procedures maintain anytime-valid type-I error control, outperform pre-specified testing methods, and can reach statistically rigorous conclusions sometimes with as few as 20 observations.
Abstract: The Segment Anything Model (SAM) achieves remarkable performance in visual segmentation. The latest SAM3 extends promptable segmentation to concept-level prediction, broadening the scope of segmentation foundation models. While recent works reveal that SAM and SAM2 are vulnerable to adversarial examples, the robustness of SAM3 under the concept segmentation paradigm remains unexplored. In addition, existing adversarial attacks on SAM-series models exhibit limited cross-prompt transferability. To this end, we propose AdvPCS, a universal cross-prompt adversarial attack for Promptable Concept Segmentation (PCS), including a min-max prompt optimization strategy, a global–local perception deception attack, and a temporal transition deviation attack. Specifically, we first identify the hardest-to-attack prompts via min-max bilevel optimization. In the inner maximization, we enhance diversity over candidate point, box, and text prompts. In the outer minimization, we select prompts with the highest responses based on the confidence scores output by the detector. Given the selected prompts, we apply the perception deception attack to minimize both global and local existence probabilities under joint prompting, and employ the temporal memory misalignment attack to maximize inter-frame semantic inconsistency and corrupt memory pointers. Extensive experiments on four benchmark datasets show that a single universal adversarial perturbation (UAP) generated by our method generalizes across frames from different videos and achieves strong attack performance under point, box, and text prompts. In particular, it reduces the average mIoU of various PCS models on the SA-CO dataset to below 5%, demonstrating strong attack ability.
Abstract: Recent progress in flow-based generative modeling has led to models that output high-quality samples while using only a small number of function evaluations. However, at present, there is a lack of similar advances in estimating the density of the generated samples. In particular, most existing methods either rely on restrictive architectures that enable exact calculations, or use stochastic approximations such as Hutchinson’s trace estimator that introduce substantial variance. In this work, we introduce SCAlable LikeLihood distillation of flOw maPs (SCALLOP). SCALLOP builds on the recently proposed F2D2, a likelihood flow map model that can generate samples and their densities in a small number of function evaluations. While F2D2 uses Hutchinson's estimator during training, we introduce an alternative and more scalable likelihood distillation objective that is Hutchinson-free and admits a vectorized formulation. Empirically, we demonstrate the effectiveness of SCALLOP as a Boltzmann generator in molecular science, and further validate its benefit on image datasets. SCALLOP significantly reduces both training variance and training time while consistently improving performance compared to F2D2, and is competitive with the state-of-the-art while achieving up to 10× inference speedup over the fastest baseline.
Abstract: Interpreting the internal activations of neural networks can produce more faithful explanations of their behavior, but is difficult due to the complex structure of activation space. Existing approaches to scalable interpretability use hand-designed agents that make and test hypotheses about how internal activations relate to external behavior. We propose to instead turn this task into an end-to-end training objective, by training interpretability assistants to accurately predict model behavior from activations through a communication bottleneck. Specifically, an encoder compresses activations to a sparse list of concepts, and a decoder reads this list and answers a natural language question about the model. We show how to pretrain this assistant on large unstructured data, then finetune it to answer questions. The resulting architecture, which we call a Predictive Concept Decoder, enjoys favorable scaling properties: the auto-interp score of the bottleneck concepts improves with data, as does the performance on downstream applications. Specifically, PCDs can detect jailbreaks, secret hints, and implanted latent concepts, and accurately surface latent user attributes.
Abstract: Natural images are continuous, yet most generative models synthesize them on discrete grids, limiting resolution-flexible generation. Continuous neural fields enable resolution-free rendering, but prior methods introduce continuity only at the decoding stage as an interpolation module, leaving the generative latent space discretized and reconstruction-oriented. We propose RaPD (Resolution-agnostic Pixel Diffusion), which performs diffusion in a continuous Neural Image Field (NIF) latent space. RaPD bridges this reconstruction-generation gap with Semantic Representation Guidance for generation-aware latent learning and a Coordinate-Queried Attention Renderer for coordinate-conditioned, scale-aware rendering. A single denoised latent can be rendered at arbitrary resolutions by changing only the query coordinates, keeping diffusion cost fixed. Experiments demonstrate superior generation quality and resolution scalability.
Abstract: We consider the well-studied setting of minimizing a convex Lipschitz function using either gradient descent (GD) or its stochastic variant (SGD), and examine the last iterate convergence. By now, it is known that standard stepsize choices lead to a last iterate convergence rate of \log T/\sqrtT after T steps. A breakthrough result of Jain et al. [2019] recovered the optimal 1/\sqrtT rate by constructing a non-standard stepsize sequence. However, this sequence requires choosing T in advance, as opposed to common stepsize schedules which apply for any time horizon. Moreover, Jain et al. conjectured that without prior knowledge of T, no stepsize sequence can ensure the optimal error for SGD's last iterate, a claim which so far remained unproven. We prove this conjecture, and in fact show that even in the noiseless case of GD, it is impossible to avoid an excess poly-log factor in T when considering an anytime last iterate guarantee. Our proof further suggests that such (slightly) suboptimal stopping times are unavoidably common.
PaperID: 3454, Poster
Abstract: Brain decoding is limited by the availability of labeled neural data, and remains challenging in low-data regimes. To address this issue, we investigate whether and when brain decoding can be boosted by augmenting small fMRI datasets with synthetic data generated by a pretrained model of fMRI responses to stimuli. We use TRIBE v2, a large encoding model pretrained on more than 1000\,hours of fMRI responses to video, audio and language. For each dataset, we evaluate systematic grids that show how the performance of image decoders varies with the amount of synthetic data used for training. Our results, based on two datasets (the 7T fMRI Natural Scenes Dataset and 3T fMRI BOLD5000), show up to 68% improvement in Top-10 image-retrieval accuracy compared to decoders trained only on real data. Importantly, the proportion of augmented data required to reach a given image decoding performance needs to be adjusted depending on the data source. Surprisingly, image decoders trained exclusively on synthetic fMRI perform above chance in some settings, suggesting that TRIBE v2 can support zero-shot fMRI-to-image decoding. Together, these results show how large-scale models of the fMRI responses to sight, sound and language may provide a foundation to improve the data efficiency for image decoding.
PaperID: 3455, Poster
Abstract: The quadratic cost of attention-based sequence models for long contexts has motivated a growing line of research on memory-based models that can compress context into a compact state. However, most existing memory models expose a static memory throughout the entire sequence. In practice, this approach is prone to memory being “polluted” by early inputs, and it can saturate and fail to incorporate later context. We study a new paradigm of incremental memory activation, where the effective capacity of memory is controlled and progressively expanded as the context grows. By imposing an early bottleneck and introducing fresh capacity over time, this leads to better compression of history and reduces interference. We instantiate this paradigm in Proteus, a straightforward mechanism that can be incorporated into a broad class of neural memory architectures at no additional cost. We apply Proteus to state-of-the-art models, including Hope-Attention, SWLA, Comba, and Titans, and observe consistent improvements across all of them. Overall, our results show that static memory is suboptimal and that incremental memory activation is a promising direction for long-context management.
Authors: Julio Oscanoa, Irmak Sivgin, Cagan Alkan, Daniel Ennis, John Pauly, Mert Pilanci, Shreyas Vasanawala
Abstract: Score-based diffusion models achieve state-of-the-art performance for inverse problems, but their practical deployment is hindered by long inference times and cumbersome hyperparameter tuning. While pretrained diffusion models can be reused across tasks without retraining, inference-time hyperparameters such as the noise schedule and posterior sampling weights typically require ad-hoc adjustment for each problem setup. We propose principled reparameterizations that induce invariances, allowing the same hyperparameters to be reused across multiple problems without re-tuning. In addition, building on the RED-diff framework, which reformulates posterior sampling as an optimization problem, we further develop the OptDiff pipeline. OptDiff provides a simplified tuning framework that facilitates the integration of convex optimization tools to accelerate inference. Experiments on image reconstruction, deblurring, and super-resolution show substantial speedups and improved image quality.
PaperID: 3457, Poster
Abstract: Multimodal sentiment analysis (MSA) and emotion recognition (ER) are closely related affective understanding tasks, yet they are usually studied separately due to differences in label space, task formulation, dataset distribution, and modality dependence. In this paper, we propose MEME (Mixture-of-Experts for Multimodal sentiment analysis and Emotion recognition), a lightweight hierarchical MoE framework for unified multimodal affective computing across MSA, conversational emotion recognition, and dynamic facial expression recognition. MEME operates on frozen text, visual, and audio features, compresses variable-length modality sequences with learned-query attention pooling, and processes them with a shared hierarchical MoE backbone. Each block first applies hard modality experts for modality-specific refinement and then task-conditioned cross experts for task-aware multimodal interaction. A \texttt[TASK] token provides an explicit task anchor for routing, while modality dropping regularizes unified training. We jointly train MEME on nine affective benchmarks and evaluate all datasets using a single composite-best checkpoint without dataset-wise adaptation. Extensive experiments show that MEME outperforms strong task-specific baselines and recent large-model-based unified methods on most benchmarks, while maintaining favorable efficiency, robustness and generalization.
PaperID: 3458, Poster
Abstract: Many modern learning approaches are still struggling with spatial reasoning tasks, i.e. they lack the ability to utilize geometric information of perceived entities and their spatial relation to each other to solve problems. We introduce a novel Adaptive Neural Cellular Automata (aNCA) architecture which uses deformable convolutions to dynamically adapt the perceptive field and iteratively reason over 2D spatial relations on grid-like data structures (e.g. images). Empirical results on public benchmarks show state of the art comprehensible results with high generalization abilities for solving image based puzzles like Sudoku or finding the shortest path in a maze.
PaperID: 3459, Poster
Abstract: Low-Rank Adaptation (LoRA) is often highly sensitive to learning-rate (LR) choices. Recent empirical studies suggest that, once LRs are carefully tuned, vanilla LoRA can remain competitive with many proposed LoRA variants. This shifts the central practical question from designing ever more variants to reducing the cost and brittleness of LR selection. LoRA+ addresses this issue with asymmetric LRs, using a larger LR for the up-projection matrix B than for the down-projection matrix A. However, the ratio \eta_B / \eta_A itself becomes an additional hyperparameter: large ratios can improve adaptation, but under the standard LoRA initialization they can also cause unstable training and ratio-dependent performance collapse. We identify the source of this brittleness through a signal--noise decomposition of LoRA dynamics. Under standard \initA, the update contains an initialization-induced noise term whose magnitude can be affected by the B-learning rate \eta_B, the initialization scale \beta, and the adapter rank r. In practical LLM fine-tuning, this non-vanishing noise can limit the performance gains expected from LoRA+-style asymmetric learning rates. To suppress this effect, we propose \initAA, a simple initialization that scales the variance of A as \sigma_A^2 \propto n^-2 while keeping B_0=0. This reduces the normalized initialization-noise contribution by an additional factor of 1/n, stabilizing asymmetric LoRA training while preserving zero initial adapter output. Empirically, \initAA substantially widens the stable LR region across initialization scales, adapter ranks, and LoRA update multipliers. Across tasks and models, \initAA reduces best-ratio drift, prevents the high-ratio collapse observed with standard \initA, and makes the width-based guideline \eta_B \approx n \eta_A a useful starting point for finite pretrained models.
Abstract: Chain-of-thought (CoT) reasoning has become a widely used mechanism for eliciting multi-step reasoning in large language models by generating intermediate reasoning steps at inference time. Yet the scaling behavior of generalization with CoT depth remains poorly understood. To address this question, we study a theoretically solvable model of CoT for in-context weight prediction in linear regression, where test-time reasoning is represented as an iterative refinement of the weight-parameter estimate. Using tools from random matrix theory under high-dimensional asymptotics, we derive an exact formula for the generalization error as a function of reasoning depth, pretraining data amount, and context length. Our analysis reveals a sharp phase transition separating exponential and polynomial improvement, saturation, and overthinking, and characterizes how the optimal reasoning depth scales. We further show that deeper reasoning is most effective with sufficiently rich pretraining and in-context information, whereas limited pretraining or context makes longer reasoning prone to error amplification or saturation. We also validate these predictions through experiments on fully learned linear attention and softmax attention models. Our results provide a unified theoretical account of how test-time CoT depth affects generalization.
PaperID: 3461, Poster
Abstract: Correlation metrics such as the Pearson Correlation Coefficient (PCC) and Concordance Correlation Coefficient (CCC) are standard regression metrics and learning objectives, but their non-decomposable nature poses a challenge for stochastic optimization. Since these metrics depend on dataset-level moments (mean, variance, covariance), naive mini-batch training with local statistics yields biased, high-variance gradients whose expectation does not align with the population gradient. To bridge this "batch-to-population'' (B2P) gap, we propose B2P-Corr, a novel framework that transforms correlation objectives into decomposable pseudo-losses driven by lightweight global "moment sketches''. We introduce two efficient variants: B2P-EMA, which utilizes exponential moving averages and stop-gradient operators to ensure asymptotic gradient tracking with O(B) cost (B is batch size); and B2P-U, which uses sampled cross-batch U-statistic estimators and a reservoir-based stabilization with an O(B) implementation. Theoretically, we quantify the O(B^-1) bias of naive mini-batch optimization, prove a moving-target tracking result for B2P-EMA under a fast-sketch / slow-parameter regime, and provide an idealized control-variate analysis for B2P-U. Experiments on two synthetic and three real-world datasets demonstrate B2P-Corr achieves stable optimization and higher correlation than naive mini-batch optimization with negligible overhead.
Authors:
Muyun Jiang, Shuailei Zhang, Zhenjie Yang, Wu Mengjun, Wei Zhang, Weibang Jiang, Chenyu Liu, Zhiwei Guo, Rui Liu, Shangen Zhang, Yong Li, Yi Ding, Cuntai GuanAbstract: Electroencephalography (EEG) foundation models learn transferable representations for brain--computer interfaces, but existing approaches treat tasks and labels as discrete targets and fail to use the semantic structure of natural-language task instructions to guide representation learning. We present LEAF, a foundation model for EEG--Language Alignment with Semantic Task Instruction and Querying. LEAF integrates task-aware semantic guidance to produce structured and linguistically aligned EEG embeddings, thereby enhancing decoding robustness and transferability. In the EEG pretraining stage, we introduce a joint Spectral--Temporal Reconstruction (STR) framework that captures the coupled spectral rhythms and temporal dynamics of EEG signals. STR applies randomized spectral perturbation to enhance frequency robustness and uses two complementary temporal objectives to learn both contextual and sequential structure. In the EEG-Language alignment stage, we propose the Instruction-conditioned Q-Former (IQF). This query-based cross-attention transformer injects instruction embeddings into EEG tokens and achieves semantic alignment with textual label embeddings through learnable queries. We evaluate LEAF on 16 downstream datasets spanning motor imagery, emotion recognition, steady-state visual evoked potentials, covert speech, and healthcare tasks. LEAF achieves state-of-the-art performance on 12 of the 16 datasets and obtains the best average results across all five task categories. Importantly, our analyses reveal for the first time that explicit task instructions serve as semantic priors guiding EEG embeddings into coherent and linguistically grounded spaces. Code is available at \urlhttps://anonymous.4open.science/r/LEAF-Model
Abstract: Recent analyses question whether reinforcement learning (RL) is responsible for strong reasoning in large language models (LLMs). At the same time, distillation and inference-time sampling, including power sampling, have emerged as effective ways to improve LLM performance. However, the relationship among RL, distillation, and sampling remains unclear. In this study, we focus on the power distribution, the target distribution of power sampling, and show that the power distribution bridges sampling, self-reward KL-regularized RL, and self-distillation. From the sampling perspective, we show that inexpensive local approximations cannot reproduce sequence-level power without information about possible suffixes. From the RL perspective, the power distribution is the closed-form optimizer of KL-regularized RL when the model's sequence-level log-probabilities are used as the reward. This identification leads to \emphpower self-distillation, an offline distillation surrogate that shares the same target distribution and amortizes the cost of power sampling into supervised training on teacher samples. We further show that power self-distillation can achieve self-reward sharpening, while improvement in a downstream true reward is governed by the covariance between true reward and self-reward under the power distribution. Experiments on reasoning tasks support our analysis: power sampling raises self-reward, true-reward gains depend on alignment with self-reward, and power self-distillation can match or exceed the performance of power sampling at much lower inference cost.
Abstract: Vision-Language Models (VLMs) for radiology report generation are typically trained on retrospective clinical reports, which suffer from omission noise: clinically present findings are left unreported due to the omission of subtle findings. For example, prior studies show that cardiomegaly may be omitted from ICU chest X-ray reports when the imaging request is focused on monitoring support device placement. As a result, models trained with standard approaches inherit these omissions, learning to under-report findings themselves. We propose PU-DPO, a preference optimization framework that constructs contrastive pairs via editing model responses, producing variants that explicitly mention or omit a specific finding. To prevent omission noise from corrupting the preference signal, we reformulate the objective under a positive-unlabeled (PU) learning framework, treating absent mentions as unlabeled rather than truly negative. Across semi-synthetic experiments and preliminary analyses on real-world chest radiograph benchmarks where adjudicated labels are available, PU-DPO yields consistent gains in detection rates across multiple pathologies without compromising specificity or overall report quality, and is more robust to omission noise than prior approaches.
Abstract: Learning accurate value functions plays a decisive role for reinforcement learning (RL) agents to solve long-horizon, complex tasks. Conventional temporal-difference (TD) learning objectives suffer from value-estimation bias that accumulates over the horizon, while extended-horizon modeling methods, such as n-step TD backups and Q-chunking, adopt a rigid, fixed-horizon value-modeling recipe that is often not flexible enough to capture complex value structures in long-horizon, multi-stage tasks. In this paper, we show that enabling value updates with dynamic horizon composition can yield a strong offline policy learning scheme. Our method, _Horizon Adaptive Offline Policy Learning via VAlue STitching_ (_VAST_), replaces fixed-horizon backups with recursive, horizon-adaptive value composition. Its key ingredient is to couple value optimization with a future state- and horizon-length-conditioned _auxiliary value function_ that is learned through direct data supervision, and a _stitching policy_ that optimally selects the reward-maximizing horizon length and future sub-goal to achieve horizon-adaptive value stitching. This design enables direct estimation and compositional “stitching” of variable-length returns grounded in actionable sub-goal states, providing an accurate and greedily exploitable value-supervision signal for offline policy optimization. Across 50 tasks on OGBench, _VAST_ outperforms fixed-step, extended-horizon methods, and generative-value offline RL baselines, achieving strong performance particularly in high-complexity, long-horizon decision-making tasks.
PaperID: 3466, Poster
Abstract: The internal activations of modern machine learning models align surprisingly well with the activities recorded in real brains. In the last five years, the alignment between encoding models and brain activities, often called brain scores, has steadily increased. Because brain recordings are noisy, there exists an upper bound to the brain scores that any model can achieve. This bound, called the noise ceiling, is crucial for two reasons: it determines whether further improvements in model performance are possible, and it enables meaningful comparisons of brain scores across benchmarks. The field has developed multiple approaches to estimate noise ceilings, but the statistical properties of existing estimators remain poorly understood, leaving it unclear in which experimental settings they provide reliable estimates. Here, we introduce a unified mathematical framework for brain recordings, use it to derive closed-form expressions for four commonly used noise ceiling estimators, and provide theoretical predictions for the behavior of each estimator. We validate the theoretical predictions with simulations that we compare to noise ceiling estimates in existing fMRI datasets. We find that the existing methods underestimate the true noise ceiling in many settings. We conclude with recommendations for best practice in different situations.
PaperID: 3467, Poster
Abstract: We partially resolve Conjecture 1 of Karagodin et al., Clustering in Causal Attention Masking (NeurIPS 2024) and extend it to time-dependent parameter regimes. The conjecture predicts that causal attention dynamics converge to two clusters aligned with the eigenvectors associated with the largest eigenvalue \lambda_\max of the value transformation. We prove that this behavior holds when \lambda_\max>0 and is generally false when \lambda_\max\le0. The resulting two-cluster regime is common in practice and substantially more challenging than the classical single-cluster setting studied previously. A central contribution of our work is the identification of a rigorous dimension deduction phenomenon: despite evolving in a high-dimensional embedding space, token trajectories collapse onto a two-dimensional structure governed by a small number of dominant directions. Building on this insight, we introduce a new analytical framework--instability within the flow--inspired by instability phenomena in physical systems. This framework enables the analysis of steady-regime behavior in non-monotone, non-autonomous attention dynamics. Our theoretical findings are illustrated through simulations and experimental studies using Group Query Attention (GQA) layers.
PaperID: 3468, Poster
Authors: Guangyu Nie, Yang Jiao, Yi Ren
Abstract: In the context of computational materials science, analytical homogenization theories, such as Strong Contrast Expansion (SCE), develop PDE-induced power expansions that map microstructural statistics of material samples to their macroscopic effective properties. However, low-order truncations of such expansions for computational feasibility often trade off prediction accuracy, especially when high-order correlations, finite-resolution effects, or strong contrast regimes become important. While learned residual surrogates, e.g., neural operators, can alleviate this tradeoff, they do not ensure reliable inference of structure-property sensitivities, hampering downstream material design tasks. To this end, we propose Neural-Corrected Operator (NCO) that learns a correction directly to the analytical PDE kernel to compensate for the low-order truncation, and show that NCO intrinsically bounds sensitivity inferences. We evaluate NCO on structure-property prediction and inverse-design tasks for heterogeneous bi-phase composite materials governed by linear second-order PDEs. NCO improves the accuracy of low-order SCE models for structure-property prediction and produces sensitivities that facilitate more effective microstructure design optimization than output-level residual baselines.
PaperID: 3469, Poster
Authors: Ziyang Song, Lia Shen, Yixuan Li, Qincheng Lu, Weihao Li, Haoran Xu, Mike He Zhu, Xinyu Wang, Yue Li
Abstract: While biomedical language models (LMs) achieve strong performance on various clinical tasks, they often fail to capture the structured relationships among medical concepts. Because LMs rely on data-driven text learning and fail to learn meaningful patterns from long-tailed distributed medical concepts. We introduce , a simple but effective representation framework that integrates expert-curated medical ontologies with LMs to generate hierarchy-aware concept embeddings. For each clinical code, SMI aggregates embeddings across hierarchical levels, leading to a representation that preserves hierarchical semantics from coarse to fine-grained clinical concepts. Our theoretical analysis shows that SMI yields a strictly higher signal-to-noise ratio than hierarchy-agnostic LM embeddings, particularly improving the representation of infrequent fine-grained concepts. Experimental results show that SMI learns interpretable geometry, improves concept relation discrimination and clinical prediction, and is robust to distribution shifts.
PaperID: 3470, Poster
Authors:
Ziheng Li, Zexu Sun, Jinman Zhao, Erxue Min, Yongcheng Zeng, Hui Wu, Hengyi Cai, Shuaiqiang Wang, Xu Chen, Zhihong Deng, Dawei YinAbstract: Reinforcement learning with verifiable rewards (RLVR) has advanced the reasoning capabilities of large language models (LLMs). However, existing RLVR methods often suffer from exploration inefficiency due to mismatches between problem difficulty and model capability: overly difficult problems hinder reasoning path discovery, while overly simple problems offer little learning signal. To address this, we first formalize the effect of problem difficulty by quantifying the relationship between loss descent magnitude and rollout accuracy. Building on this analysis, we propose SEELE, a supervision-aided RLVR framework that dynamically adjusts problem difficulty to lie within the high-performance region. SEELE augments each training sample by appending a hint (part of a full solution) for difficulty reduction. Unlike previous hint-based approaches, SEELE deliberately computes the hint length for each individual problem to achieve an optimal difficulty. The optimal hint length is determined via multi-round rollout sampling, where an item response theory model fits accuracy–hint pairs from previous rounds to predict the next-round hint. This instance-level, real-time difficulty adjustment aligns problem difficulty with the evolving model capability, thereby improving exploration efficiency. Experiments show that SEELE outperforms Group Relative Policy Optimization (GRPO) and Supervised Fine-tuning (SFT) by +10.0 and +8.4 points, respectively, and exceeds the best prior supervision-aided approach by +3.8 points on average across six math reasoning benchmarks.
Abstract: Hyperparameter transfer has become an important component of modern large-scale training recipes. Existing methods, such as muP, primarily focus on transfer between model sizes, with transfer across batch sizes and training horizons often relying on empirical scaling rules informed by insights from timescale preservation, quadratic proxies, and continuous-time approximations. We study hyperparameter scaling laws for modern first-order optimizers through the lens of recent convergence bounds for methods based on the Linear Minimization Oracle (LMO), a framework that includes normalized SGD, signSGD (approximating Adam), and Muon. Treating bounds in recent literature as a proxy and minimizing them across different tuning regimes yields closed-form power-law schedules for learning rate, momentum, and batch size as functions of the iteration or token budget. Our analysis, holding model size fixed, recovers most insights and observations from the literature under a unified and principled perspective, with clear directions open for future research. Our results draw particular attention to the interaction between momentum and batch-size scaling, suggesting that optimal performance may be achieved with several scaling strategies.
Abstract: Despite significant advances in autonomous web navigation, current methods remain far from human-level performance in complex web environments. We argue that this limitation stems from Topological Blindness, where agents are forced to explore via trial-and-error without access to the global topological structure of the environment. To overcome this limitation, we introduce WebNavigator, which reframes web navigation from probabilistic exploration into deterministic retrieval and pathfinding. WebNavigator constructs Interaction Graphs via zero-token cost heuristic exploration offline and implements a Retrieve-Reason-Teleport workflow for global navigation online. WebNavigator achieves state-of-the-art performance on WebArena and Online-Mind2Web. On WebArena multi-site tasks, WebNavigator achieves a 72.9% success rate, more than doubling the performance of enterprise-level agents. This work reveals that Topological Blindness, rather than model reasoning capabilities alone, is an underestimated bottleneck in autonomous web navigation.
Abstract: Scaling laws govern macroscopic resource allocation for LLMs, yet precise architectural configurations for Mixture-of-Experts (MoE) models remain guided by heuristics. Existing MoE scaling studies either incorporate MoE-specific variables into scaling formulas, causing fitted coefficients to grow rapidly without proportional experimental support, or fix all non-MoE factors, implicitly assuming global architecture does not influence local MoE scaling. We propose a reusable framework for holistic MoE architectural optimization. We first reveal that relying solely on FLOPs per token (M) biases MoE evaluation, as heterogeneous Attention/FFN computational densities enable parameter inflation without effective compute gains. We therefore establish a joint constraint triad of M, active parameters (N_a), and total parameters (N) for rigorous MoE characterization. To tame the resulting \mathcalO(n^16) search space, we employ mathematical decoupling: structural constraints and a rank-preserving property of the hidden dimension factorize the optimization into an \mathcalO(n^3) + \mathcalO(n^2) two-phase search. Through extensive validation across 670+ MoE models spanning 10^18 to 3 × 10^20 FLOPs, we derive globally applicable scaling laws mapping any compute budget to its optimal architecture. A key finding is that the near-optimal configuration band widens with scale, providing a quantitative basis for trade-offs between scaling law recommendations and engineering constraints. Our work delivers actionable blueprints for optimal MoE design under arbitrary compute budgets.
Abstract: Many features in pretrained Transformers span multiple layers: they emerge through stages of inference, persist in the residual stream, or are built jointly by parallel MLPs. Crosscoders (namely, sparse dictionaries trained jointly across layers) aim to recover these cross-layer features in a single shared latent space. We show that standard crosscoders largely fail at this purpose. Although their decoder weight norms spread evenly across layers, a functional coherence metric we introduce reveals that each latent's activation is effectively driven by only one or two layers on average. While functionally coherent latents act as human-interpretable concept detectors (e.g., US states and cities), the layer-localized latents that crosscoders predominantly learn collapse onto surface-level patterns such as digit detectors. We trace this failure to two structural limitations: unconstrained cross-layer parameterization and unregularized cross-layer dependence. We address both by introducing , which (i) replace the encoder and decoder with low-rank tensor factorizations that draw every latent's per-layer weights from a shared cross-layer basis, and (ii) apply stochastic layer masking, a denoising regularizer along the layer axis that penalizes latents whose contribution collapses when a single layer is masked. Across GPT2-Small, Pythia-410M, Pythia-1.4B, and Gemma2-2B, lift mean probing F1 by 10-30 points, surpassing per-layer SAE baselines that standard crosscoders fail to reach, reduce reconstruction MSE by 25-50%, and roughly double mean functional coherence. An LLM-as-a-judge evaluation further shows that recover 3-13 times more semantically coherent latents than standard crosscoders across all four base LLMs.
PaperID: 3475, Poster
Abstract: Streaming 3D reconstruction has become a critical component for real-time applications. While recent advances have adapted large-scale offline feedforward models into streaming architectures via causal attention and KV-caching, they suffer from fundamental instability: the irreversible bias in early frames. This limitation stems from two coupled factors: (1) the existing KV-cache-based causal architecture strictly prohibits access to future information, resulting in information-poor early caches; and (2) current methods typically anchor the pose constraints to the initial frame, amplifying the early bias throughout the sequence. To address these challenges, we propose Stable3R, a holistic framework that establishes stable geometric references by enriching historical caches and anchoring poses to prefixes rather than a single initial frame. Specifically, we first introduce a lookahead-augmented attention mechanism, integrating future observations into past caches without violating causal constraints. Second, we adopt a frame-equivariant architecture with prefix-relative supervision, alleviating the reliance on the biased initial frame. Extensive experiments across multiple benchmarks show that our method consistently boosts reconstruction quality within a strictly causal inference schedule.
PaperID: 3476, Poster
Abstract: Aging cohorts increasingly combine imaging, plasma, and cognitive measurements, but most subjects provide only a partially observed cross-sectional snapshot, leaving latent stage, progression subtype, and stage-varying multimodal coordination to be inferred jointly. Existing methods estimate the cohort-level aging process, model covariance under observed covariates, or fuse modalities for prediction. However, they rarely target the joint problem of endpoint-anchored staging and progression-resolved residual coordination across modalities, even though subject staging, subtype assignment, and multimodal interplay are mutually dependent. We introduce ELPAC, a Bayesian framework with two interlocking constructions. The first integrates subtype trajectories backward from clinically reliable endpoint groups along biologically structured velocity components encoding modality-specific monotone direction and optional sequential ordering priors. The second factorizes residual multimodal variation around the inferred stage axis into sparse coordination axes whose loadings and activations expose which multimodal patterns are operative in each subtype and when they emerge, with stage-resolved variation supplying the auxiliary structure for conditional identifiability. We use a scalable amortized variational inference scheme that enables joint posterior inference at cohort scale with single-pass deployment on new subjects. On synthetic data and a harmonized ADNI+NACC/SCAN Alzheimer's cohort, ELPAC demonstrates accurate latent-structure recovery, well-calibrated prediction intervals, and biologically interpretable multimodal coordination programs that differ by subtype.
Abstract: CLI agents are the closest thing language models have to an embodied setting: the model emits commands, the terminal executes them, and the returned stream---stdout, errors, file contents---records the consequences of its actions. We argue that this stream is a supervision signal, but standard agent RL largely discards it: GRPO-style training updates action tokens while ignoring environment responses, even though rollouts already contain them. We introduce ECHO (Environment Cross-entropy Hybrid Objective), an auxiliary loss that trains the policy to predict observation tokens alongside the policy gradient. ECHO reuses the same forward pass, requires no additional rollouts, and turns terminal feedback into a dense signal that remains informative even on failed trajectories. This objective shapes the policy toward a world model of terminal dynamics, plausibly improving its priors over which actions are likely to succeed. On TerminalBench 2, ECHO nearly doubles pass@1: Qwen3-8B improves from 2.70% to 5.17%, and Qwen3-14B from 5.17% to 10.79%. From base Qwen3-8B, ECHO+GRPO matches expert-SFT-then-GRPO on held-out terminal tasks without the expert demonstrations required for SFT, and closes roughly half this gap on TerminalBench 2. Without verifier rewards, ECHO can lead to self-improvement through exploration alone, suggesting environment observations are not merely context for actions but supervision for the policy itself.
PaperID: 3478, Poster
Abstract: Adapting CLIP for open-vocabulary video recognition necessitates a delicate balance between newly acquired video knowledge and the pretrained generalization. While existing studies pursue this generalization-specialization trade-off with additional regularizations or constraints, we argue that they overlook the deviation of representations beyond the fine-tuning data distribution, resulting in suboptimal adaptation effects. We believe such deviation is inherited from the inconsistency between the fine-tuning and evaluation objectives, where model optimization is restricted to the known training distribution but evaluated on unseen ones. In this paper, we introduce \emphTACO, a simple yet effective framework to mitigate the potential negative effects induced by this inconsistency. Our key insight is that adaptation should preserve OOD-relevant alignment beyond the training distribution. To this end, we propose \emphRelative Structure Distillation, which regularizes the relative geometry of the representation space and suppresses harmful alignment shift during training. We further decouple the representation space from the optimization space with a lightweight specialization projection, allowing task-specific adaptation without directly overspecializing the representations used at test time. \emphTACO establishes state-of-the-art performance on diverse benchmarks under cross-dataset and base-to-novel settings. Code will be released.
PaperID: 3479, Poster
Abstract: Recent progress in video generation has been remarkable, yet state-of-the-art models remain largely confined to 720p, well short of the 4K resolution of modern displays. Scaling video generation to higher resolutions is constrained by the scarcity of high-resolution video data and the prohibitive cost of large-scale training. Beyond data and compute, higher resolutions place substantially greater demands on spatial modeling, particularly for fine-grained details, complex compositions, and local structures. Such high-resolution spatial supervision, however, is abundantly available in high-resolution images, which are easier to collect and train on at scale. This naturally raises a question: can image-only supervision unlock native 4K video generation? To answer this, we propose ApertureAttn, an inverted attention pyramid that concentrates adaptation on high-resolution fine-grained spatial modeling using only single-frame high-resolution images, without any video training data. We further introduce window-scaled extrapolation and proximity-guided temporal attention to enhance spatiotemporal consistency. Extensive experiments show that our method enables native 4K video generation with strong visual fidelity and temporal coherence, while requiring only image supervision and minimal training cost. Despite this lightweight training setup, it remains competitive with state-of-the-art high-resolution video generation methods trained directly on high-resolution video data at substantially higher training cost.
Authors:
Anik K Paul, Nibedita Roy, Nagesh Talagani, Swetha Ganesh, Gugan Chandrashekhar Mallika Thoppe, Alexandre Reiffers-MassonAbstract: We propose FAR-SIGN (Fully Asynchronous Robust optimization via SIGNed directional projections) for adversary-resilient learning in parameter-server--worker systems. FAR-SIGN achieves robustness through sign-based updates along carefully designed directions and mitigates the resulting bias via a two-timescale mechanism. It admits both first-order and zeroth-order implementations and enables fully asynchronous execution without requiring a private reference dataset at the server. We establish almost-sure convergence of FAR-SIGN to the set of stationary points for smooth, nonconvex objectives. Moreover, we prove the near-optimal rate of O(n^-1/4+\epsilon) in the first-order setting and the standard O(n^-1/6+\epsilon) in the zeroth-order setting, where n is the iteration count and \epsilon>0 can be chosen arbitrarily small. Experiments on MNIST show that FAR-SIGN outperforms robust aggregation-based methods in both accuracy and wall-clock time.
PaperID: 3481, Poster
Authors: Jeonghwan Cheon, Marin Vogelsang, Lukas Vogelsang, Pawan Sinha
Abstract: Vision-language models now serve as general-purpose semantic embedding spaces, but how conceptual abstraction is organized within such spaces remains poorly understood. Here, we show that abstraction in these embeddings is spectrally stratified. Specifically, low-rank principal-component subspaces encode broad conceptual structure, while higher-rank components carry progressively finer-grained distinctions. This pattern holds across multiple contrastive vision-language models and taxonomic datasets. Consistent with this stratification, rank-based modulation produces predictable behavioral shifts: retaining only the leading components preserves coarse retrieval but degrades fine-grained retrieval, while removing them selectively disrupts abstract category structure. Furthermore, rank-selective routing improves retrieval when the selected subspace matches the target level of abstraction. Notably, compact spectral subspaces, which preserve taxonomic structure while truncating high-rank residuals, exhibit substantially better alignment with human abstraction behavior than the full-rank model. Together, these results reveal that vision-language embeddings are not semantically homogeneous. Instead, hierarchical abstraction is stratified along the spectral geometry, and the structure most aligned with human cognition is concentrated in a compact subspace, suggesting that human-aligned semantics are a recoverable substructure of, rather than a property of, the full embedding.
PaperID: 3482, Poster
Abstract: Muon has become a competitive optimizer for large-scale neural network training through matrix-level orthogonalized updates, but its full-matrix normalization can be expensive. Blockwise Muon is an appealing low-cost approximation: it normalizes matrix shards or small blocks independently, reducing computation and communication overhead while often achieving convergence close to full Muon. However, in large-step-size and large-model regimes, Blockwise Muon can suffer from degraded convergence. We attribute this to cross-block amplification: independent per-block orthogonalizations can collectively amplify updates along certain directions, particularly those governing the effective step size. Motivated by this viewpoint, we accordingly propose \textscACME: Amplification-Corrected Muon for Effective step-size recovery. \textscACME applies a rotational correction so that the directions controlling the effective step size are no longer affected by cross-block amplification. Experiments on language-model pretraining show that, across a range of training settings, \textscACME achieves convergence nearly indistinguishable from full Muon while retaining the efficiency advantages of Blockwise Muon.
Abstract: Optimal Transport (OT) plans provide correspondences between distributions supporting alignment tasks in various domains. Sliced transport plans have been recently proposed as a computationally efficient alternative to OT plans. These methods optimize a one-dimensional projection (slice) to obtain a conditional transport plan that minimizes the transport cost in the ambient space. Despite their efficiency, it remains unclear whether learned slicers transfer to new distribution pairs under shift, an issue central to evolving data and repeated OT computations over related distributions. We study the min-Sliced Transport Plan (min-STP) framework and examine slicer transferability: can a slicer learned on one distribution pair produce effective transport plans for unseen pairs? Theoretically, we show that optimized slicers remain close under slight perturbations of the data distributions, enabling efficient transfer across related tasks. To further improve scalability, we introduce a minibatch formulation of min-STP and provide statistical guarantees on its accuracy. Empirically, we demonstrate that the transferable min-STP achieves strong one-shot matching performance and facilitates amortized training for point cloud and image analysis. Our code is available at https://anonymous.4open.science/r/Min-STP-CF74.
Abstract: High-resolution image editing is increasingly demanded in professional workflows, yet existing diffusion-based models remain constrained to resolutions below 1K due to quadratic attention complexity and prohibitive memory requirements. A prevalent workaround employs a two-stage pipeline: editing at low resolution followed by independent super-resolution. However, this approach suffers from two critical issues: information divergence, where hallucinated details contradict the original high-resolution (HR) source, and texture degradation, manifesting as over-smoothed or over-sharpened artifacts. We propose EditBridge, a diffusion bridge framework for efficient ultra high-resolution editing. Unlike conventional diffusion that regenerates from noise, we formulate refinement as structured data-to-data translation from the low-resolution (LR) edited result to its HR counterpart, explicitly conditioned on the original HR source to preserve authentic details. To efficiently incorporate HR source guidance, we introduce a prior-guided block-wise sparse attention mechanism that exploits semantic correspondence from first-stage editing to constrain cross-image interactions to spatially aligned regions, significantly reducing computational overhead. Extensive experiments demonstrate that EditBridge achieves high-fidelity editing with superior perceptual quality at resolutions up to 4K, delivering 3.6--8.4× speedup at 2K and enabling practical 4K editing in 61 seconds.
Abstract: Data mixing is a consequential problem throughout language model training. In pretraining, data composition is a key determinant of model quality; in continual learning and adaptation, it governs what is retained and acquired. Yet existing data mixing methods address only one phase of this lifecycle at a time: some require smaller proxy models tied to a single training phase, others assume a fixed domain set, and continual learning lacks principled guidance altogether. We argue that data mixing is fundamentally an online decision making problem---one that recurs throughout training and demands a single, unified solution. We introduce OP-Mix (On-Policy Mix), a data mixing algorithm that operates across the entire language model training lifecycle. Our main insight is that candidate data mixtures can be cheaply simulated by interpolating between low-rank adapters trained directly on the current model, eliminating separate proxy models and ensuring the search is always grounded in the model's actual learning dynamics. Across pretraining, continual midtraining, and continual instruction tuning, OP-Mix consistently finds near-optimal mixtures while using a fraction of the compute of the baselines. In pretraining, OP-Mix improves upon training without mixing by 6.3% in average perplexity. For continual learning, OP-Mix matches the performance of both retraining and on-policy distillation while using 66% and 95% less overall compute, respectively. OP-Mix suggests a different view of language model training: not a sequence of distinct phases, but a single continuous process of learning from data.
PaperID: 3486, Poster
Abstract: Optimal Transport (OT) is a principled framework for comparing probability distributions, but its effectiveness depends critically on the ground metric, the cost used to compare observations. In high-dimensional settings, fixed ground metrics like the Euclidean distance can make OT distances reflect task-irrelevant variation rather than low-dimensional class structure, limiting OT in tasks such as classification and clustering. Supervised OT addresses this by learning a parameterized ground metric such that the induced OT distance reflects class structure. However, existing approaches on point clouds scale poorly with the number of observations. We propose a scalable supervised OT framework for Gaussian Mixture Models (GMMs) that lifts ground metric learning from observations to mixture components, replacing large point-cloud OT problems with smaller component-level problems whose transport costs admit closed forms. We define a learnable Wasserstein-type distance between GMMs using the Generalized Bures-Wasserstein (GBW) distance between Gaussian components as the ground metric, parameterized by a rectangular linear map that projects Gaussian components into a low-dimensional latent space inducing an OT distance between GMMs. For additional scalability, we introduce a diagonal approximation of the GBW metric, reducing the covariance computation between components to linear in the latent dimension. We bound the resulting approximation error by the mutual information between latent Gaussian variables and the Frobenius norm of the learned map, yielding a principled regularization strategy that also enables interpretation of latent axes. Across synthetic and scRNA-seq benchmarks, our method improves or matches classification and clustering over fixed-metric and supervised OT baselines on point clouds while substantially reducing OT distance computation time.
Abstract: Orthogonal parameter-efficient fine-tuning adapts pretrained weights through structure-preserving multiplicative transformations, but existing methods often conflate two distinct design choices: the subspace in which adaptation occurs and the transformation applied within that subspace. This paper introduces LOFT, a low-rank orthogonal fine-tuning framework that explicitly separates these two components. By viewing orthogonal adaptation as a multiplicative subspace rotation, LOFT provides a unified formulation that recovers representative orthogonal PEFT methods, including coordinate-, butterfly-, Householder-, and principal-subspace-based variants. More importantly, this perspective exposes support selection as a central design axis rather than a byproduct of a particular parameterization. We develop a first-order analysis showing that useful adaptation supports should be informed by the downstream training signal, motivating practical gradient-informed support selection strategies. Across language understanding, visual transfer, mathematical reasoning, and multilingual out-of-distribution adaptation, LOFT recovers principal-subspace orthogonal adaptation while gradient-informed supports improve the efficiency–performance trade-off under matched parameter, memory, and transform budgets. These results suggest that principled support selection is an important direction for improving orthogonal PEFT.
Abstract: Gradient Boosting Decision Trees (GBDTs) dominate tabular machine learning, with modern implementations like XGBoost, LightGBM, and CatBoost being based on Newton boosting: a second-order descent step in the space of decision trees. Despite its empirical success, the global convergence of Newton boosting is poorly understood compared to first-order boosting. In this paper, we introduce Restricted Newton Descent, a framework for convex optimization with Newton's method on Hilbert spaces with inexact iterates, based on the concepts of cosine angle and weak gradient edge. Within this framework, we recover Newton boosting with GBDTs and classical finite-dimensional theory as special cases. We first prove that vanilla Newton boosting achieves a linear rate of convergence for smooth, strongly convex losses that satisfy a Hessian-dominance condition. To handle general convex losses with Lipschitz Hessians, we extend a recent gradient regularized Newton scheme to the restricted weak learner setting, establishing the first global \mathcalO(\frac1k^2) convergence rate for second-order GBDTs. This scheme minimally modifies the classical algorithm by introducing an adaptive \ell_2-regularization term proportional to the square root of the gradient norm at each iteration. In numerical experiments, we show that this scheme converges while vanilla Newton boosting may diverge.
PaperID: 3489, Poster
Authors: Yi Hong, Wenchao Bai, Jiaqi Jiang, Siyuan Huang, Jiahui Jin
Abstract: Knowledge conflict detection is a fundamental challenge for large language model (LLM) systems that rely on external knowledge. Existing approaches address this problem either with end-to-end supervised models or with multi-stage pipelines dependent on fixed heuristics. However, supervised models often generalize poorly to diverse conflict patterns, whereas heuristic pipelines rely on fixed rules and introduce considerable computational overhead. To overcome these limitations, we propose Grouping and Conflict Detection (\mathsfGCD), a two-stage agentic framework for knowledge conflict detection that aligns the detection process with the inherent sparsity of knowledge conflicts, thereby reducing the decision space and enabling fine-grained conflict detection. \mathsfGCD decomposes knowledge conflict detection into two subtasks and employs two cooperative agents to address them: a Knowledge Grouping Agent (G-Agent), which creates conflict-preserving groups of textual units with potential conflicts, and an Intra-group Conflict Detection Agent (D-Agent), which detects conflicts within each group. The G-Agent's grouping policy is not fixed in advance but is jointly trained with D-Agent via a shared global reward, enabling grouping to be explicitly optimized for accurate detection rather than relying on fixed heuristics. Extensive experiments on five benchmarks demonstrate that \mathsfGCD consistently outperforms SOTA methods, while reducing inference latency and token consumption.
Abstract: Transformers have demonstrated remarkable in-context learning (ICL) capabilities, enabling them to perform new tasks without additional fine-tuning. However, their performance often deteriorates when encountering out-of-distribution (OOD) inputs that deviate from the training distribution, and the underlying theory remains poorly understood. To fill this gap, we characterize the OOD error under the input distribution shift through the interplay between the dynamics of the so-called \alpha-type and \beta-type attention weights, which represent the transformer’s confidence in identifying the correct and incorrect features, respectively. Our results indicate that the OOD error for each feature depends on all pairwise interactions between the training features and OOD features, and under certain cases the transformer performs no better than random guessing. To improve the OOD generalization performance, we next investigate the impact of model finetuning with the OOD data, and particularly, characterize the model forgetting performance on the source domain. Interestingly, the performance on the source domain may not always degrade after finetuning, which highly depends on the nature of the feature shift: finetuning on OOD domain keeps enhancing the confidence of identifying correct features from the original distribution, while the interference from other incorrect features may either increase or decrease. Extensive experiments on both synthetic and real data are conducted to corroborate the theoretical insights.
PaperID: 3491, Poster
Abstract: As vision-language models (VLMs) transition from frozen-encoder pipelines to native multimodal architectures, the visual channel increasingly serves as both a perceptual input and a carrier of rendered text that competes with user prompts. Current single-attack, pixel-centric evaluations fail to capture this complexity: they collapse distinct visual attack surfaces into a single scalar and ignore how architectural shifts alter vulnerability. We introduce PaRMA, an evidence-based auditing framework that defines pathway-conditioned risk over two primary image-level interventions: pixel perturbation and rendered-text insertion. Instead of arbitrary attack enumeration, PaRMA uses targeted diagnostics to expose architecture-specific blind spots. Notably, we uncover a see-but-not-deceived signature where native VLMs attend heavily to typographic overlays but visually reject them, and we identify severe gradient dispersion in deep-fusion models. To exploit these findings, we instantiate FusionAudit, introducing two targeted attacks: MANTIS (margin-aware pixel perturbation) and Nat-Evid (host-congruent semantic insertion). Evaluating 18 open-weight VLMs across a novel nativeness taxonomy, we demonstrate that legacy single-attack rankings (e.g., PGD) are highly unstable for native models. Our diagnostic-driven attacks substantially raise the estimated residual risk, proving that PaRMA provides a falsifiable, evidence-driven auditing methodology to uncover vulnerabilities systematically hidden by standard benchmarks.
Abstract: Decomposing prediction uncertainty into aleatoric (irreducible) and epistemic (reducible) components is critical for the reliable deployment of machine learning systems. While the mutual information between the response variable and model parameters is a principled measure for epistemic uncertainty, it requires access to the parameter posterior, which is computationally challenging to approximate. Consequently, practitioners often rely on probabilistic predictions from deep ensembles to quantify uncertainty, which have demonstrated strong empirical performance. However, a theoretical understanding of their success from a frequentist perspective remains limited. We address this gap by first considering a bootstrap-based estimator for epistemic uncertainty, which we prove is asymptotically correct. Next, we connect deep ensembles to the bootstrap estimator by decomposing it into data variability and training stochasticity; specifically, we show that deep ensembles capture the training stochasticity component. Through empirical studies, we show that this stochasticity component constitutes the majority of epistemic uncertainty, thereby explaining the effectiveness of deep ensembles.
Authors: Shafayeth Jamil, Rehan Kapadia
Abstract: Most physical systems are governed by nonlinear partial differential equations, yet the analytical tools that make linear systems transparent, such as eigenspectra, dispersion relations and modal decomposition, remain unavailable in nonlinear settings. We introduce Lie Generator Network-Koopman (LGN-KM), a neural operator that lifts nonlinear dynamics into a linear latent space and learns the continuous-time Koopman generator through the decomposition L_k = S - D_k, where S is skew-symmetric (conservative coupling) and D_k is positive-definite diagonal (mode-selective dissipation). The decomposition guarantees stability by construction and exposes the eigenstructure directly, making linear-systems analysis available on a nonlinear PDE. On 2D Navier-Stokes turbulence, the generator recovers a complete multi-branch dispersion relation and the viscous dissipation scaling \mathrmRe(\lambda) \propto -|k|^2 from trajectory data alone, with no physics supervision. The same architecture recovers the analogous diffusive scaling on FitzHugh-Nagumo reaction-diffusion without modification. Independently trained models at different regimes show matched gauge-invariant spectral structure. The architecture additionally enables long-horizon stability, O(1) continuous-time evaluation, and physics--informed cross-regime model transfer.
PaperID: 3494, Poster
Abstract: Achieving peak GPU kernel performance increasingly relies on architecture-specific optimizations targeting new hardware features. While AI coding agents show promise in generating performant kernels, they lack the necessary context to effectively implement and stack hardware-specific optimizations, especially on newer GPU architectures. We propose Hawkeye (Hardware-Aware Kernel Optimization), an open-source framework that grounds autonomous kernel generation in a minimal and comprehensive taxonomy with only one unit test per optimization strategy per target architecture. Supporting a new accelerator therefore requires only 10 expert-written unit tests per architecture (one per recurring optimization strategy) that generalize across downstream workloads, rather than hand writing a new kernel for each workload and precision. Hawkeye effectively scales the test-time compute of coding agents with this minimal expert supervision to enable kernel generation that consistently leverages hardware-specific features, approaching and even surpassing expert-written PyTorch or Triton in BF16 and emerging low precision (FP8, NVFP4, MXFP4) across Ampere, Hopper, Blackwell, and MI350 GPUs. Hawkeye demonstrates that minimally supervised coding agents can exploit architecture-specific hardware features and reduce the overhead of supporting emerging hardware accelerators.
PaperID: 3495, Poster
Abstract: Decoding continuous speech from intracortical recordings holds transformative potential for individuals with severe motor impairments, yet current systems face fundamental barriers to real-world deployment: training data is scarce, neural signals are non-stationary across sessions, and the external language models required by state-of-the-art cascaded decoders impose memory requirements that preclude local, privacy-preserving inference. We introduce BrainWhisperer, a neural speech decoder that adapts Whisper–pretrained on approximately 680,000 hours of speech–to microelectrode array recordings via a convolutional front-end, hierarchical low-rank projections for non-stationarity, windowed self-attention in phoneme-selective encoder layers, and a multi-task objective combining CTC and cross-entropy losses. Subject-specific embedders enable cross-participant training within a unified architecture. Evaluated on publicly available Utah array datasets from BrainGate participants, BrainWhisperer achieves a word error rate of 8.5% in end-to-end decoding on the Card benchmark--to the best of our knowledge, the best result reported in this setting. Cross-dataset training improves performance without participant-specific fine-tuning, pointing toward scalable foundation models for speech brain-computer interfaces.
PaperID: 3496, Poster
Authors: Zhiming Hu, Angela N Ye, Ran Zhang, Tristan T Aumentado-Armstrong, Samrudhdhi B Rangrej, Faraz Ali, Dong Huo, Kosta Derpanis, Alex Levinshtein, Iqbal Mohomed
Abstract: Diffusion-based image restoration models produce visually compelling results but systematically hallucinate text---generating characters that appear sharp and legible yet are factually incorrect. As text fidelity emerges as a key requirement for real-world image restoration, this failure mode poses a fundamental barrier to wide deployments of diffusion-based models: a model that replaces degraded-but-faithful text with sharp-but-wrong text misleads users and erodes trust. Recent text-aware models mitigate this by feeding text extracted from intermediate outputs or low-resolution (LR) images into the prompt, but they provide no mechanism to detect or correct the hallucinations that still commonly arise. In this work, we present a systematic study on how to improve text restoration during denoising, demonstrating that standard global conditioning alone is inadequate and that optical character recognition (OCR) confidence scores reliably predict word-level correctness. Grounded in these findings, we propose HalluciText, a training-free framework consisting of two complementary interventions: (i) Confidence-Weighted Classifier-Free Guidance (CW-CFG), which concentrates guidance energy on text regions weighted by OCR confidence to improve text recall; and (ii) VLM-Guided Test-Time Guidance (V-TTG), which leverages a vision-language model (VLM) to identify hallucinated regions and corrects them towards a faithful reference, improving text precision. Both operate solely on the noise prediction or latent representation, making them agnostic to the diffusion backbone and text spotting model. We demonstrate the effectiveness of HalluciText on state-of-the-art (SOTA) text-aware models, achieving up to 5.75 points of improvement in end-to-end text recognition F1 on several datasets without degrading image quality.
Abstract: Neural surrogate models for computational fluid dynamics (CFD) are typically trained as forward operators that map explicit problem specifications, such as geometry and boundary conditions, to solution fields. This ties the model to the conditioning variables seen during training and limits reuse under boundary-condition shifts or local geometry changes. We propose to reformulate steady CFD inference as an inpainting problem: instead of training on explicit boundary conditions, we learn a self-supervised prior over velocity fields and impose boundary constraints only during inference by fixing known regions such as inlet, outlet or unchanged regions from previous simulations. To scale this idea to large 3D meshes, we introduce a local neighbourhood tokeniser that represents high-resolution velocity fields as compact spatial latent tokens and train latent flow-matching and masked-autoencoder models on these tokens. On intracranial aneurysm hemodynamics, our method reconstructs full velocity fields from sparse boundary context, outperforms supervised neural surrogates under boundary-condition and dataset shift and enables local geometry editing by reusing unchanged simulation context. These results suggest that viewing CFD inference as context-conditioned inpainting can turn neural surrogates from task-specific predictors into reusable flow priors.
PaperID: 3498, Poster
Authors: Lucas Thil, Jesse Read, Rim Kaddah, Guillaume Doquet
Abstract: We present a novel method for learning interpretable representations of progressive time series, that is, data capturing irreversible state transitions such as degradation or task completion. Our approach uses a self-supervised contrastive objective to learn a low-dimensional latent space where progression manifests along manifolds anchored by fixed prototype vectors. This structure yields latent-space indicators that quantify progression in a human-meaningful way without proxy labels. We evaluate the approach against the state of the art on diverse domains, including industrial degradation, robotic tasks, and neural activity, validating three key capabilities: (1) end-state prediction, (2) multi-step forecasting, and (3) interpretable phase separation. Our method matches or improves over black-box counterparts on all of these while providing transparency about the underlying mechanisms. A simple linear regressor on top of the learned indicators is competitive with deep architectures, providing direct quantitative evidence that the underlying state is encoded in a geometrically accessible form. Code is available at https://anonymous.4open.science/r/LRPTS-9300/README.md.
Abstract: Bayesian optimization (BO) is a popular technique for sample-efficient optimization of black-box functions. In many applications, the parameters being tuned come with a carefully engineered default configuration, and practitioners only want to deviate from this default when necessary. Standard BO, however, does not aim to minimize deviation from the default and, in practice, often pushes weakly relevant parameters to the boundary of the search space. This makes it difficult to distinguish between important and spurious changes and increases the burden of vetting recommendations when the optimization objective omits relevant operational considerations. We introduce BONSAI, a default-aware BO policy that prunes low-impact deviations from a default configuration while explicitly controlling the loss in acquisition value. BONSAI is compatible with a variety of acquisition functions, including expected improvement and upper confidence bound (GP-UCB). We theoretically bound the regret incurred by BONSAI, showing that, under certain conditions, it enjoys the same no-regret property as vanilla GP-UCB. Moreover, assuming known ARD lengthscales---the same assumption underlying GP-UCB regret bounds---BONSAI provably recovers the relevant-coordinate set at zero acquisition cost, yielding a method that matches the GP-UCB regret rate while recovering the minimal-\ell_0 solution---a guarantee not provided by prior sparse-BO methods. Across many real-world applications, we empirically find that BONSAI substantially reduces the number of non-default parameters in recommended configurations while maintaining competitive optimization performance, with little effect on wall time---averaging only 1.5× the candidate-generation cost of standard BO, compared to 7-34× on average for prior sparse-BO methods (IR, ER, and SEBO).
Abstract: 3D vision-language segmentation aims to segment target objects in 3D scenarios according to the linguistic instructions and visual observations. Prior art heavily relies on the coarse superpoint representation to reduce the computation complexity , which suffers from poor segmentation quality and messy object boundaries. In this paper, we propose the SEGment-And-select (SEGA3D) paradigm for 3D vision-language segmentation that directly operates on the fine-grained visual information and is free from the superpoint dependency. Specifically, we first leverage a mask candidate generator to provide fine-grained categorical mask candidates, substantially improving the quality of candidate masks over the superpoint counterparts. Then, a Large Language Model (LLM) is utilized to generate the semantic and spatial information based on the linguistic description and visual features. The LLM output and visual features are fed to the Semantic-Spatial Selector (SSS) to produce the top-ranking mask candidates. Eventually, the Loopback Verification Module (LVM) is designed to yield the segmentation mask from the selected candidate masks. Our SEGA3D attains competitive performance on ScanRefer, ScanNet and Matterport3D benchmarks. Notably, our SEGA3D surpasses the top-performing counterpart by 8.3 mIoU and 5.3 mIoU on ScanNet and Matterport3D, respectively. Codes will be available upon publication.
PaperID: 3501, Poster
Abstract: The performance of offline Multi-Agent Reinforcement Learning (MARL) is limited by the quality of the fixed offline dataset, motivating trajectory augmentation. However, augmenting multi-agent trajectories presents a fundamental challenge of the diversity–coordination trade-off: independent generation improves diversity but lacks coordination, while joint generation induces coordination but lacks diversity, reproducing the training distribution. To address this challenge in multi-agent data augmentation, we introduce MASTARS, a novel diffusion-based framework that generates multi-agent trajectories sequentially across agents while enforcing cross-agent consistent coordination via the method of inpainting. MASTARS produces trajectories that are diverse, coordinated and globally coherent. Experiments on various benchmarks such as MPE, SMAC, and SMAC-v2 demonstrate significant performance gains across multiple offline MARL methods.
PaperID: 3502, Poster
Abstract: Optimal Transport (OT) provides a principled framework for comparing probability distributions. Classical solvers remain the gold standard for this task, providing exact solutions. At the same time, OT is an ideal testbed for studying whether learning-based amortized models can internalize the structure of OT type problems. Here, ground-truth solutions are available, constraints are explicit, and exact discrete solutions have well-characterized sparsity. We describe ToR-OT, a transformer-based model that learns to predict discrete OT plans through autoregressive token regression. We reframe OT into a structured sequence generation task. ToR-OT decodes sparse transport plans as sequences rather than solving an optimization task at test time. We focus on three properties: (a) Zero-shot generalization: pretraining on a large synthetic corpus of OT instances allows a single model to generalize to unseen distribution pairs and heterogeneous cost metrics without retraining. (b) Sparse primal decoding: by exploiting the inherent sparsity in exact discrete OT solutions, the decoder predicts only nonzero entries (avoids dense coupling outputs). (c) Amortized inference: on low-dimensional discrete problems, ToR-OT gives accurate approximate plans and compares favorably with per-instance neural OT baselines. Overall, ToR-OT is not yet a replacement for classical solvers, but best viewed as a step toward understanding how autoregressive transformers can learn structure of constrained optimization problems like OT.
Abstract: While Neural Combinatorial Optimization (NCO) achieves strong performance, its black-box nature remains a key roadblock. Standard interpretability tools, such as Concept Bottleneck Models (CBMs), are ill-equipped for NCO, whose decisions are dynamic, state-dependent, and lack proper concept definition. To close this gap, we introduce Evolving Programmatic Bottlenecks (EPB), the first framework that distills black-box NCO models into human-readable program portfolios. EPB extends CBMs for sequential decision problems: it employs an LLM to autonomously evolve a bank of programs, where each program's per-step action distribution serves as the bottleneck. We realize this through an iterative framework: Block I fixes program bank capacity and introduces a hybrid textual-numerical gradient descent scheme to backpropagate gradients across the student router model and the LLM in-context learning; Block II dynamically adapts bank capacity via fault-targeted boosting and redundancy pruning. Extensive experiments demonstrate EPB's effectiveness and broad applicability, where the interpreted program portfolios largely match original performance. EPB also reveals that NCO behavior shifts across optimization stages and can be approximated as a composition of classic heuristic variants. Our work advances interpretable NCO and establishes EPB as a promising tool for interpreting complex sequential decision-making models.
Abstract: We uncover a latent capacity for introspection in a Qwen 32B model, demonstrating that the model can detect when concepts have been injected into its earlier context and identify which concept was injected. While the model denies injection in sampled outputs, logit lens analysis reveals clear detection signals in the residual stream, which are attenuated in the final layers. Furthermore, the effect is heavily dependent on context – e.g., prompting the model with accurate information about AI introspection mechanisms can dramatically strengthen this effect: the sensitivity to injection increases massively (0.3% → 39.9%) with only a 0.6% increase in false positives. Also, mutual information between nine injected and recovered concepts rises from 0.62 bits to 1.05 bits, ruling out generic noise explanations. Our results demonstrate models can have a surprising capacity for introspection and steering awareness that is easy to overlook, with consequences for latent reasoning and safety.
PaperID: 3505, Poster
Abstract: Knowledge Graph Question Answering (KGQA) increasingly requires multi-turn interaction with knowledge graphs (KGs) to derive final answers. Such interactive capabilities often rely on proprietary large-scale LLMs, which hinders cost-efficient and local deployment. Reinforcement learning offers a scalable alternative for endowing smaller open-source LLMs with multi-turn interaction ability, as it enables models to improve from self-explored interaction trajectories rather than relying on expensive expert demonstrations. However, KGQA presents a distinctive challenge for RL: unlike tasks where exploration can proceed in a relatively unconstrained textual space, KG reasoning is governed by a sparse graph topology. Free-form policy rollouts frequently generate relation sequences that cannot be executed on the graph, making reward signals sparse and unstable. We propose ICOR, an Interactive Combinatorial Reinforcement Learning framework that grounds policy exploration in the combinatorial structure of KGs. ICOR first retrieves a query-specific set of candidate relations and then restricts policy rollouts to relation-path compositions within this graph-derived space. This design increases the likelihood that sampled trajectories are executable and thus informative for policy optimization. Building on this constrained rollout space, ICOR applies GRPO with a staged feedback mechanism that guides learning from structural validity to answer-consistent reasoning. Experiments on multiple KGQA datasets demonstrate the effectiveness of ICOR.
PaperID: 3506, Poster
Abstract: We demonstrate that AI agents can autonomously improve mathematical proofs of quantitative results. As a key step towards this frontier, we achieve strict improvements on Grothendieck's and Korenblum's constants, whose precise values have remained surprisingly elusive despite their recurrence in both the mathematical and physical sciences. By encoding relevant literature as a partially formalized Lean blueprint, we structure the agent's search space to localized proof refinements that automatically propagate to a verified final bound. Concretely, Grothendieck's constant (\mathcalK_G) is shown to be: \mathcalK_G \leq \tfrac\pi2 \log (1 + \sqrt2) - 10^-17, which is the first explicit numerical improvement since Krivine's landmark 1979 result. For Korenblum's constant (\mathcalK_K), the lower bound is tightened to: \mathcalK_K \geq 0.3554 + 0.001165, a third-digit improvement over Wang's recent 2025 result. In summary, we establish that AI agents can autonomously refine formal proofs of open problems and provide a generalizable framework for future mathematical discovery.
PaperID: 3507, Poster
Abstract: We study online multiclass classification with surrogate losses. We identify and exploit a structural property of margin-based classifiers: on many rounds predictions are stable in the sense that after an update of the parameters the algorithms do not change the predicted label on the current example. By exploiting this stability we improve upon the state of the art in several ways. We provide an improved surrogate regret bound for the perceptron, develop a parameter-free version of \textttGAPTRON, develop improved results for the delayed feedback setting, and extend our results to the batch setting. These results are complimented by empirical evaluations.
Abstract: Score-based distillation methods train one-step diffusion models in two stages: they first train a teacher score model, then distill it into a one-step student model. However, this distillation process can introduce bias from two sources: errors in the teacher score model and student score estimate. We propose DiffRatio, a new framework for training one-step diffusion models without teacher supervision. Instead of using a teacher score model to provide training targets, DiffRatio directly learns a log density ratio between the student and data distributions across diffusion time steps. As a pre-trained score model is only used to initialize the one-step generator and does not supervise training. This design simplifies the training pipeline, mitigates gradient estimation bias, and reduces the size of auxiliary networks. In addition, the learned density ratio can be used as a verifier, enabling a principled inference-time parallel scaling method that further improves sample quality without external rewards or extra sequential computation. DiffRatio achieves strong one-step generation results on CIFAR-10 and ImageNet (64×64 and 512×512), outperforming most teacher-supervised distillation methods.
Abstract: We consider the problem of learning an unknown subset N^target of a domain in an online setting. % In each round t, the learner predicts a set of items N_t and receives one of two types of feedback, each with equal probability: \it precision feedback, in which a randomly chosen item from the predicted set N_t is revealed and the learner is told whether it belongs to N^target (incurring a reward if it does), or \it recall feedback, in which a randomly chosen item from the target set N^target is revealed and the learner is told whether it belongs to N_t (incurring a reward if it does). The goal is to maximize the cumulative reward over time. This simple online set learning problem abstracts a variety of learning scenarios with precision and recall--type feedback. We show that a hypothesis class (a family of subsets of the domain) is learnable in this setting if and only if it has finite Vapnik-Chervonenkis (VC) dimension, mirroring the classical PAC characterization. However, the resulting algorithmic structure is markedly more intricate: in contrast to standard Probably Approximately Correct (PAC) learning---where the algorithmic landscape is governed by the simple principle of Empirical Risk Minimization (ERM)---our partial feedback model can invalidate ERM and even all proper learning rules. We develop algorithms to address the dependencies induced by the feedback, obtaining regret guarantees in both the realizable and agnostic settings. Our results provide a qualitative characterization of learnability in this model, addressing its most basic question, while pointing to a range of natural and intriguing open questions, including the determination of optimal regret rates.
Abstract: In multimodal classification, late-fusion approaches classify concatenated modality-specific features extracted by unimodal neural networks. When modality imbalance is pronounced, various regularization techniques have been proposed to balance the learning process and overcome the inferior performance of late-fusion networks. In contrast, this work demonstrates that multimodal data can be effectively classified without any explicit modality fusion, using deep ensembles of unimodal networks. We systematically compare deep ensembles to late-fusion networks at equal parameter count and show that ensembles consistently outperform state-of-the-art late-fusion methods designed to address modality imbalance. This advantage also holds over intermediate-fusion techniques we evaluated and over hybrid methods that combine unimodal and multimodal predictions. We propose and empirically validate a method for selecting the number of models per modality in an ensemble, avoiding computationally expensive exhaustive search. Under extreme modality imbalance and small ensemble sizes, the heuristic indicates that ensembles of unimodal models trained solely on the stronger modality are preferable; as the ensemble scales up, incorporating models from the weaker modality becomes beneficial. Both predictions align with our empirical findings. To systematically explore the challenges of optimizing multimodal models, we propose a synthetic multimodal framework that allows control over both the number of modalities and their predictive strength; our findings are consistent across synthetic and real-world datasets. Finally, by fitting scaling laws to bimodal datasets, we estimate the asymptotic performance of ensembles.
Abstract: Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation, yet they still struggle to generalize to unseen tasks that necessitate transferring relevant experience across objects, scenes, and action patterns. This paper proposes VLA-Pro, a plug-and-play framework designed to enhance cross-task generalization by storing task-relevant procedural memories at training time and transferring these memories during inference. Specifically, VLA-Pro stores task-specific LoRA adapters as parameterized procedural memories during training. At inference time, VLA-Pro retrieves relevant procedural memories based on the current multi-modal context and dynamically fuses these memories for generating the current action chunk. Experiments on RoboTwin, RLBench, and real-world manipulation tasks show that VLA-Pro consistently improves cross-task generalization across multiple backbones, achieving up to a 207% relative improvement in simulation and increasing real-world success rate from 5.8% to 65.0%. These results suggest that procedural memory retrieval and adaptation provide an effective mechanism for transferring manipulation experience to novel tasks while preserving modularity and execution stability. The code is available at https://anonymous.4open.science/r/VLA-Pro/.
PaperID: 3512, Poster
Abstract: Immunohistochemical (IHC) staining encodes molecular protein expression critical for clinical diagnosis, yet its chemical procedures are costly and time-consuming. Virtual staining offers a compelling alternative by digitally synthesizing IHC images from hematoxylin-eosin (H&E) stained slides. Despite remarkable progress in diffusion- and flow-matching-based virtual staining, supervised fine-tuning (SFT) remains fundamentally limited in pathological fidelity due to its coarse, spatially averaged supervision signal that starves sparse DAB-positive regions of effective gradients. Reinforcement learning (RL) can exploit inter-rollout variance for outcome-level pathological supervision, but naively applying DAB-based rewards triggers severe reward hacking that collapses the image distribution and degrades perceptual quality. This paper presents StainNFT, a flow matching RL post-training framework built on a curriculum reward strategy that gates fine-grained optical density supervision on per-sample DAB mask IoU. To enhance fine-grained pathological fidelity, we introduce expression-aware reweighting, multi-scale block supervision, and a closed-loop implicit pathological semantic reward derived from pathology foundation model priors. Extensive experiments on seven benchmarks demonstrate that StainNFT consistently outperforms existing methods in both perceptual quality and pathological fidelity for H&E-to-IHC virtual staining. Ablation studies confirm the effectiveness of each proposed component. Our code and trained models will be released.
PaperID: 3513, Poster
Abstract: Empathetic spoken dialogue is a sophisticated cognitive process that requires not only recognizing emotions but also inferring a user’s latent mental states to provide appropriate support. However, current SpeechLLMs often treat empathy as a direct input-to-response mapping, leading to "superficially warm" but emotionally hollow interactions. In addition, since empathy relies on a multi-stage process with strong inter-step dependency, errors at any intermediate step can cascade through subsequent steps and lead to inappropriate responses, while existing training paradigms lack mechanisms to precisely localize and improve such errors. In this work, we propose EmpathyChat, a unified framework that reformulates empathetic spoken dialogue as a structured cognitive reasoning process integrating perception, mental-state reasoning, and response generation. To support this paradigm, we first construct EmpathySet-400K, an acoustically rich dataset for multi-stage empathetic supervision. During the SFT stage, we strengthen acoustic grounding through proposed Acoustic-Anchored Attention (AAA). During the RL stage, we further introduce a novel stage-aware optimization objective with Step-Decomposed Credit Assignment (SDCA) to localize reasoning errors and mitigate cascaded error propagation. In addition, we introduce EmpathyEval, an expert-annotated benchmark for multi-dimensional empathy evaluation. Extensive experiments demonstrate that EmpathyChat achieves state-of-the-art performance in perception, reasoning, and response alignment. All resources, including model, code, dataset, and benchmark, will be publicly released.
Authors:
Hampus Linander, Conor Heins, Marco Perin, Alexander Tschanz, Christopher L BuckleyAbstract: Physical systems are naturally parameterized in terms of geometric entities and their transformations. While exact equivariance can be a powerful inductive bias, real-world applications frequently exhibit approximate or broken symmetries in practice. We introduce object-centric world models that provide a soft geometric inductive bias without enforcing exact symmetries. By embedding object states as Clifford multivectors, our models encourage physically meaningful transformations while retaining the expressivity needed to handle asymmetric dynamics. We evaluate our approach on 2D rigid-body dynamics, 3D charged-particle systems, and real-world driving trajectories. Compared to both unstructured baselines and strictly equivariant models, our soft Clifford transformer achieves better long-horizon fidelity, particularly in regimes with broken symmetries. These results suggest that geometric algebra offers an effective middle ground, delivering sample-efficient dynamics without the inflexibility of hard mathematical constraints.
Abstract: Most Vision-Language Models (VLMs) are built by extending pretrained Large Language Models (LLMs) with visual modules and multimodal alignment. However, this multimodal scaling often degrades the language-side reasoning ability originally encoded in the base LLM. While the base LLM retains usable reasoning ability after scaling, the aligned VLM itself cannot reliably access this ability. Therefore, recovering degraded reasoning capability in VLMs may be more effective when using the base LLM as the source rather than relying on the VLM alone. Motivated by this, we propose ransfer), a lightweight vector-intervention method that transfers reasoning capability from the base LLM to the VLM without retraining the backbone. LIFT defines as answer-token hidden-state differences between a Reasoner path with an explicit reasoning trace and a Solver path without it, and injects these vectors into language-side activations of the target VLM. LIFT further supports learnable vector adaptation while keeping the VLM backbone frozen. We evaluate LIFT on two VLMs across six reasoning benchmarks, comparing Reasoning Vectors extracted from the base LLM and from the aligned VLM under matched protocols. Results show that LLM-derived vectors consistently outperform VLM-derived vectors, confirming that the base LLM is a more effective source for recovering reasoning. LIFT partially recovers degraded reasoning through lightweight language-side interventions. Further analyses show that Reasoning Vectors influence intermediate reasoning behavior rather than merely altering final answers.
Abstract: Low-Rank Adaptation (LoRA) fusion enables the composition of subject and style representations for controllable generation without retraining. However, existing approaches primarily operate through weight-level merging, without explicitly modeling how independently trained LoRAs interact in the shared parameter space. We adopt a geometric perspective on LoRA fusion, interpreting content and style LoRAs as occupying overlapping, non-orthogonal low-rank subspaces, where such overlap can lead to conflicting parameter updates that affect generation quality. This observation motivates us to reformulate LoRA fusion not merely as parameter combination, but as a problem of controlling how updates from overlapping subspaces are combined. Based on this insight, we propose Null Space Projection LoRA (NP-LoRA), a training-free framework that employs projection as a fusion operator to explicitly modulate cross-LoRA interactions. Specifically, NP-LoRA uses principal directions of the style LoRA to define a projection subspace and projects the content LoRA onto the complementary subspace (i.e., the null space of the style LoRA), suppressing interference along dominant style directions while preserving complementary information. To avoid the overly aggressive suppression of hard projection, we further formulate soft projection as a regularized optimization problem that balances content preservation against style-subspace suppression. This objective admits a closed-form solution, yielding a projection operator controlled by a single parameter that continuously interpolates between linear merging and hard projection. Extensive experiments across multiple pretrained LoRA pairs show that NP-LoRA achieves more balanced content-style composition compared to strong baselines, without requiring retraining.
PaperID: 3517, Poster
Abstract: Multimodal molecular representation learning plays a crucial role in drug discovery and molecular property prediction. Existing cross-modal alignment approaches primarily reinforce redundant information across modalities, often neglecting complementary synergistic signals critical for downstream tasks. To address this limitation, we propose MSC-Mol, a modality-synergy contrastive pre-training framework that encodes SMILES sequences, 2D molecular graphs, and 3D conformations into a unified fused representation, optimized via contrastive learning over augmented views. By introducing challenging negatives, including single-modality replacements and modality recombinations, MSC-Mol encourages the model to capture fine-grained cross-modal correspondences beyond dominant-modality shortcuts. We further design probe tasks-long-range pharmacophore prediction, molecular property regression, and functional-group combination classification-to validate that MSC-Mol preserves both modality-specific and synergistic information. Extensive experiments demonstrate that MSC-Mol consistently outperforms existing 2D, 3D, and multimodal pre-training methods on molecular property prediction, drug–target interaction(DTI) tasks, and probe evaluations, highlighting its effectiveness in capturing cross-modal synergy and long-range molecular dependencies. The code is available at https://anonymous.4open.science/r/MSC-Mol-169D.
PaperID: 3518, Poster
Authors:
Zimo Yin, Haipeng Jiang, Kailong Ren, Zhetao Sun, Guangyi Lv, Congli Yin, Likang Wu, Hongke Zhao, Ming He, Peng Wang, Jianping FanAbstract: Large language model agents now combine tools, long context, memory, and multi-step workflows, yet they still fail when execution loses track of the evidence required for a reliable answer. Common failure modes include missing support, unresolved contradictions, and evidence lost under context pressure. We present Lemon, an evidence-risk-aware framework that makes such risks an explicit control variable. At the intra-instance level, a single state-conditioned controller jointly schedules reasoning, tool invocation, worker expansion, verification, context compression, and memory operations. A recoverable evidence substrate pairs anchored context compression with reusable semantic memory, preserving raw tool outputs while storing transferable process fragments. We further extend Lemon to an inter-instance setting in which heterogeneous personalized agents expose expertise, exchange evidence-bearing proposals, critique unsupported claims, and internalize useful collaboration artifacts. Empirically, Lemon reaches 91.36% accuracy on GAIA and 80% on xbench-DeepSearch. In an open-source reproducibility comparison on GAIA, Lemon uses 49.9% to 88.4% fewer tokens per task than three top-ranked agent baselines.
PaperID: 3519, Poster
Authors:
Xiangfu Meng, Jiapeng Yang, Lei Shi, xue wang, Yuefeng Zhan, Chen Gu, Hao Sun, Weiwei Deng, Tao Yao, Feng Sun, Qi ZhangAbstract: Generative recommendation has recently emerged as a promising paradigm for large-scale recommendation by reformulating next-item prediction as discrete sequence generation. However, current SID-based models rely mainly on history-to-next-item supervision, leaving item--item collaborative topology only indirectly captured through sparse user sequences. To address this gap, we propose , a graph-derived structural supervision framework for SID-based generative recommendation. From user interactions, builds an item--item graph and derives three complementary auxiliary objectives---Neighbor Prediction, Topological Contrast, and Link Prediction --- to make collaborative topology explicit during generative training. Extensive experiments on the two public Amazon datasets demonstrate that R-GRec achieves state-of-the-art performance over representative traditional, generative, and LLM-based recommenders. Further ablation and analytical studies verify the contribution of each graph-derived objective and show that collaborative topology acts as a complementary supervision signal to Semantic ID generation, improving generative recommendation without adding inference-time graph computation.
PaperID: 3520, Poster
Authors:
Xu Bohao, Jiawei Zeng, Gang Yu, Lei Wang, Xiaoying Tang, Duanduan Chen, Tianyi QianAbstract: Realizing stable neural decoding over extended periods remains a significant challenge, primarily because invariant task-related neural dynamics are inextricably entangled with inherent signal non-stationarity and the physical drift of electrode-neuron interfaces. This work introduces LNDN (Long-term Neural Decoding Network), a hierarchical framework designed to decouple category-specific neural dynamic manifolds from representational drift, sustaining decoding precision without periodic recalibration. The proposed architecture addresses representational drift through an integrated tri-level alignment strategy. At the physical feature level, a decoupled representation and dynamic gating module isolates the physical identity of neurons from their transient states, adaptively filtering stochastic neural noise. This is further strengthened at the decision level by a Mixture-of-Experts (MoE) architecture, which employs a parallel soft-voting mechanism to suppress high variance induced by local representational distortions. Finally, at the manifold level, Adaptive Batch Normalization (AdaBN) and latent contrastive learning, combined with second-order correlation alignment, explicitly anchor the geometric consistency of the low-dimensional neural manifold over a period of months. Systematic evaluations on a longitudinal olfactory dataset, involving 9 mice performing a four-odor decoding task, demonstrate the efficacy of this approach. When trained exclusively on data from the initial four weeks, LNDN maintains an average accuracy of 83.67% on completely unseen tests spanning the subsequent eight weeks, significantly outperforming mainstream domain adaptation and advanced neural decoding baselines. This framework offers a scalable, "set-and-forget" solution for robust, long-term invasive brain-computer interface applications.
PaperID: 3521, Poster
Abstract: General-purpose time series encoders such as \textttOTIS and \textttMOMENT promise a single backbone that transfers across domains, but they do not scale and rely on manually tagged domain labels at tokenisation. We trace both symptoms to a single pathology: under masked-data-modelling supervision the encoder's output patch tokens partially collapse, with mean off-diagonal pair-wise cosine similarity stabilising at 0.68 in \textttOTIS and the encoder's effective rank capped accordingly. This is the natural consequence of supervising the encoder only in data space, which leaves no constraint against redundant token directions in the representation space. We introduce \textttOTISv2, which suppresses the collapse by complementing the masked-reconstruction objective with two representation-space terms --- a self-distillation loss against an exponential-moving-average teacher and a KoLeo regulariser --- and replaces \textttOTIS's per-domain variate embeddings with register tokens to remove the label dependency. The resulting encoder pulls mean patch-to-patch cosine similarity from 0.63 to 0.58, scales monotonically from 7.1\,M to 40\,M parameters where its predecessor plateaus, and dominates general-purpose baselines on 143 uni- and multi-variate datasets from the UEA and UCR archives from a single label-free checkpoint. We release code and pre-trained weights upon acceptance.
PaperID: 3522, Poster
Abstract: Query-aware video summarization (QVS) aims to select a concise set of moments relevant to a textual query in the context of the entire video. Existing QVS methods are often driven by local query-segment relevance scoring, which is effective for retrieving query-matching moments but insufficient for composing coherent summaries over long videos. Long-form QVS is especially challenging because query-relevant evidence tends to be sparsely distributed across extended timelines, and satisfying a query may require combining complementary segments while avoiding redundancy under a strict summary budget. We propose , a framework that autoregressively generates a query-conditioned video summary, capturing inter-segment dependencies as well as the saliency of the segment and relevance to the query. We also introduce a large-scale benchmark for long-form QVS ( ) with open-vocabulary queries and budget-constrained extractive summaries, together with a complementary metric for evaluating query-aware summary disentanglement. Experiments show that QLVSumm achieves state-of-the-art performance and demonstrates stronger query-aware summary disentanglement.
Abstract: We consider a data seller who designs pricing mechanisms over multiple datasets to maximize revenue from budget-constrained buyers. The seller offers multiple datasets and assigns each a pricing function that maps the quantity purchased to a total payment. The goal is to design these pricing functions to maximize revenue, anticipating that buyers—who trade off accuracy gains against cost—choose bundles optimally subject to their budget constraints. Prior work [Chaudhury et al., 2026] studies such optimal pricing under the restriction that each dataset is assigned a linear price, and shows that computing optimal linear prices is computationally intractable. In contrast, we allow each dataset to be priced via a general function and show that this additional flexibility can not only increase the revenue but also restore tractability, yielding a surprising simultaneous improvement in economic performance and computational efficiency. Even when pricing functions are only required to be monotone and lower-continuous, optimal pricing admits a highly structured and simple form: each pricing function is piecewise linear and convex (PLC), and the optimal solution can be computed in polynomial time. Moreover, the total number of kinks across all pricing functions is bounded by the number of buyers. Consequently, when datasets significantly outnumber buyers, most pricing functions are effectively linear. We further empirically study the structure of optimal pricing by analyzing the number of kinks and the revenue gap between optimal nonlinear pricing and optimal linear pricing on simulations generated from a real dataset.
Abstract: Large language models (LLMs) exhibit pronounced position bias in long-context retrieval, systematically prioritizing information location over relevance. Existing mitigations modify positional encodings or attention internals, but in the common output-only deployment setting these are inapplicable. We introduce GOLD PANNING, a Bayesian framework for inference-time active search that (i) reorders documents to concentrate high-belief items in highly diagnostic positions (signal anchoring) and (ii) updates relevance beliefs from model outputs. Unlike active learning, which prioritizes uncertainty reduction, GOLD PANNING exploits anchoring---once flagged, keep it in sight---to preserve weak cues. A greedy assignment derived from the model's diagnosticity profile provably identifies a target among N documents with high probability in O(\log N) rounds. Across open-weight and closed-source models on multi-document QA, GOLD PANNING matches Permutation Self-Consistency's F_1 with 30--65% fewer queries and remains effective under calibration mismatch, indicating that coarse positional ordering alone drives the gains. These results show that inherent model biases need not be failures, but can serve as exploitable signals.
PaperID: 3525, Poster
Abstract: Sequential testing avoids some of the many controversies and drawbacks related to traditional batch testing with p-values. In particular, the paradigm has developed rapidly in recent years and frames the level of evidence against the null as the wealth of a player in a betting game. Naturally, how much a player should bet is a key consideration, and should reflect the statistician's confidence in what they know about the next observed data point. Much of current work uses standard betting strategies developed for adversarial problems, where the data can be hand-picked to the detriment of the player. In the stochastic setting, where the data observed is independent and stationary, these approaches may be suboptimal. We propose a simple betting strategy, (FLOE), that exploits the i.i.d. nature of a data stream to estimate the e-power of our current test on the underlying distribution. We average not over the historical evidence collected so far, as generic online learning-based methods do, but instead over the equally-likely possible data sequences we might have seen based on shuffling the data. With an additional principled regularization technique, this simple drop-in method yields substantial improvements in testing performance on a variety of kernel two-sample testing problems.
PaperID: 3526, Poster
Authors:
HaiHui Huang, Dingkui Kang, Yanan Zhou, Yong LiangAbstract: Unsupervised feature selection is widely used as a discovery tool, yet most methods return only absolute rankings: a selected gene, pixel, or sensor is never compared against a feature-wise null. We introduce Self-Supervised Reconstruction Knockoffs (SSRK), a calibrated proxy-discovery framework for unlabeled data. SSRK defines relevance through masked reconstruction: a feature is deemed useful only when replacing it with a matched knockoff degrades reconstruction of other masked coordinates. The method trains a symmetric knockoff-gated masked autoencoder and converts the learned gates into a slot-aware signed statistic that satisfies the knockoff sign-flip property under Model-X exchangeability and coupled implementation. We formalize the estimation target as knockoff-relative masked information, prove immunity to self-copy and marginal-variance artifacts, derive an entropy-regularized population fixed point for correlated gates, and establish finite-sample margin and defect-robust false discovery rate bounds. Oracle synthetic experiments validate knockoff+ control at q = 0.10 with full support recovery in the reported regimes. On Peripheral Blood Mononuclear Cells, MNIST, Fashion-MNIST, and UCI Human Activity Recognition, where learned knockoffs are approximate, the same statistic is evaluated as a ranking and recovers biologically, spatially, and sensor-structurally meaningful subsets. SSRK thereby turns masked self-supervision into a feature-wise null comparison while clearly separating controlled discoveries from exploratory rankings.
PaperID: 3527, Poster
Abstract: Cross-modal knowledge distillation transfers knowledge from a teacher modality to a student modality. Existing feature-level alignment methods typically assume that teacher and student features reside in structurally alignable representation spaces. However, this assumption does not hold when cross-modal features are structurally heterogeneous and lack clear unit-level correspondence, such as 2D spatial visual grids and 1D temporal audio sequences, thereby limiting the applicability of feature-level alignment. To address this challenge, we propose a cross-modal distillation framework that enables effective knowledge transfer across structurally heterogeneous feature spaces via a vector-quantized codebook. Specifically, teacher features are abstracted into a set of vector-form codes regardless of their original feature structure, and the selected codes serve as concept-level anchors for student learning. Code selection is guided by both task relevance and student compatibility, allowing the student to receive transferable teacher knowledge without requiring direct unit-level feature alignment. Experimental results across diverse cross-modal distillation scenarios demonstrate the effectiveness of the proposed framework on classification and semantic segmentation tasks.
Abstract: Muon and its variants have shown strong empirical performance in a variety of deep learning tasks. Existing convergence analyses of Muon rely on smoothness assumptions, though arguably the most successful function class for developing deep learning methods (such as AdaGrad, Shampoo, Schedule-Free and more) has been the class of convex and Lipschitz functions. In this paper we question whether the classical convex Lipschitz model is a useful one for understanding Muon. Our answer is no. We show that Muon does not converge on the class of convex and Lipschitz functions, regardless of the choice of learning rate schedule. We also show that error feedback restores convergence of Muon and all the non-Euclidean subgradient methods with momentum. However, this theoretical fix using error feedback degrades the performance of Muon in two representative settings for image classification (CIFAR-10) and language modeling (nanoGPT on FineWeb-Edu 10B). Our conclusion is that convex Lipschitz theory, despite having a prominent role in the design of practical methods for deep learning, is not the most suited one for Muon. This suggests that Muon's success must come from structure absent from this model, most plausibly related to smoothness.
Authors: Tung-Ling Li, Hongliang Liu
Abstract: Aligned language models refuse unsafe requests because RLHF widens the logit margin between refusal and affirmative tokens at the first decoded position. We call this scalar the refusal-affirmation logit gap and use it as a per-prompt diagnostic for alignment robustness. On three model families, alignment widens the gap on 97.5-99.8% of toxic prompts, and gap closure tracks True ASR across suffix strategies (an internal consistency check, since our method optimises for gap closure). We present logit-gap steering, a gradient-free, forward-pass-only method that searches for in-distribution suffixes drawn from the model's own high-probability candidates whose cumulative effect closes the gap. Discovering 8 ensemble suffixes per family costs \approx26,000 forward-pass equivalents (\approx2~min on one A100), \approx125× less than a single GCG search. Suffixes discovered on 0.5B-2B models transfer to 72B within family. An 8-suffix ensemble reaches 38--96% True ASR across 13 models on AdvBench and HarmBench. Most suffixes have 10^3-10^4× lower perplexity than GCG: under a published PPL-filter defense, our ensemble holds at 76.0% (from 76.9%) while GCG collapses from 64.7% to 1.0%.
Abstract: Why must vision-language navigation be bound to detailed and verbose language instructions? While such details ease decision-making, they fundamentally contradict the goal for navigation in the real-world. Ideally, agents should possess the autonomy to navigate in unknown environments guided solely by simple and high-level intents. Realizing this ambition introduces a formidable challenge: Beyond-the-View Navigation (BVN), where agents must locate unseen targets without step-by-step guidance. Existing large language model (LLM)-based methods, though adept at following dense instructions, often suffer from short-sighted behaviors due to their reliance on short-horizon supervision. Simply extending the supervision horizon, however, destabilizes LLM training. In this work, we identify that video generation models (VGM) inherently benefit from long-horizon supervision to align with language instructions, rendering them uniquely suitable for BVN tasks. Capitalizing on this insight, we propose applying VGM as the main backbone to generate actions for the first time in this field. Yet, the prohibitive latency for generating videos spanning tens of seconds makes real-world deployment impractical. To bridge this gap, we propose SparseVideoNav, achieving sub-second trajectory inference guided by a generated sparse future spanning a 20-second horizon. This yields a remarkable 27× speed-up compared to the unoptimized counterpart. Extensive real-world zero-shot experiments demonstrate that SparseVideoNav achieves 2.5× the success rate of state-of-the-art LLM baselines on BVN tasks and marks the first realization of such capability in challenging night scenes.
Abstract: High-fidelity atomistic evolution over long timescales requires more than observing the current crystal configuration. Instantaneous atomistic snapshots are often incomplete: locally similar configurations can correspond to different hidden dynamical contexts, future event preferences, and waiting-time scales. We argue that this snapshot ambiguity makes long-horizon atomistic evolution fundamentally a memory-based world-state restoration problem. To address this, we introduce AtomWorld-Mem, a memory-restored atomistic world model that recovers the latent world state missing from instantaneous crystal snapshots. AtomWorld-Mem treats the evolving alloy as an AtomWorld: spatial encoders write multi-scale atomistic keyframes from dense local topology and sparse long-range defect context, while short-term event memory and long-term structural memory integrate these keyframes across time to restore a future-predictive evolutionary state. The restored state is used to prioritize legal vacancy-mediated events under single-event Kinetic Monte Carlo (KMC) constraints, while event legality, physical execution, and residence-time updates remain governed by the underlying simulator. Empirically, AtomWorld-Mem improves long-horizon atomistic progress under fixed microscopic event budgets while maintaining high-fidelity evolution across energetic, structural, and vacancy-transport observables. It further transfers zero-shot across diverse unseen alloy-temperature AtomWorlds, suggesting that the learned memory-restoration mechanism captures reusable principles of hidden-state inference rather than a system-specific local energy heuristic. These results position memory-restored world-state modeling as a promising route toward efficient, physically grounded, and transferable atomistic evolution.
Abstract: Reasoning models have advanced rapidly, but the dominant reinforcement learning from verifiable rewards (RLVR) recipe remains surprisingly narrow: sample many responses and reward each with a single bit indicating whether the final answer is correct. Yet many settings provide rich feedback, including execution traces, tool outputs, expert corrections, and model self-evaluations. We study how to use such feedback through a distributional variant of the classic imitation learning algorithm DAgger, where the learner has local access to an expert distribution on states visited by the current policy. This yields a simple forward cross-entropy objective that admits a blackbox expert and whose sequence-level gradient conduct rich credit assignment by propagating future expert-student disagreement back to earlier decisions. We show that prior RL with self-distillation objectives based on reverse KL or Jensen-Shannon fail to guarantee monotonic policy improvement: even when the expert has higher reward, their updates may increase probability on worse actions. In contrast, we show that forward cross-entropy admits monotonic policy improvement and enjoys guarantees on regret. We further show that our objective optimizes a teacher-weighted maximum-likelihood RL lower bound, leading to improved Pass@N. Empirically, our approach, DistIL, consistently improves over RLVR and RL with self-distillation baselines across a variety of domains: scientific reasoning, mathematical reasoning, and coding.
PaperID: 3533, Poster
Abstract: Differentiable structure learning transforms DAG discovery from a combinatorial search into a continuous optimization problem, enabling gradient-based recovery of causal structures. In many applications, researchers possess partial-order priors over the variables (e.g., from biological signaling cascades or temporal ordering), which substantially narrow the space of plausible causal graphs. Existing approaches integrate such priors by decomposing the prior order graph into a collection of paths and evaluating prior-aware acyclicity constraints path by path; this is correct but its constraint-evaluation cost scales linearly with the number of induced maximal paths, a bottleneck for multi-chain priors. We propose PACO, Partial-order-Augmented Continuous Optimization, which encodes the same partial-order compatibility through a single masked-reachability matrix, that masks the matrix exponential of the learned graph by the strict transitive closure of the prior. The constraint and its Frechet-adjoint gradient are computed in a single 2d×2d block matrix exponential whose dominant cost is independent of the number of induced prior paths. We prove that PACO has the same exact zero-level feasible set as the path-decomposition formulation. Across synthetic and real-data experiments under matched priors and dataset seeds, PACO matches path-decomposition on structural recovery and substantially reduces wall-clock cost in multi-chain prior settings.
Abstract: Supervised fine-tuning (SFT) induces new behaviors in large language models, yet imposes no structural constraint on how these behaviors are distributed within the model. Existing behavior interpretation methods, such as circuit attribution approaches, identify sparse subnetworks correlated with SFT-induced behaviors post-hoc. However, such correlations do not imply , limiting the ability to selectively control SFT-induced behaviors at inference time. We pursue an alternative by asking: can an SFT-induced behavior be deliberately compressed into a sparse, mechanistically necessary subnetwork, termed a , which constructs such carriers by jointly optimizing routing masks and model weights under an explicit utility budget, and (b) , a soft prompt optimized via activation matching on extracted carrier channels, to reverse the SFT-induced behavior. Across safety, fixed-response, and style behaviors on multiple model families, LCDD yields sparse carriers that preserve target behaviors while enabling strong reversion when triggered by SFT-Eraser. Ablations further establish that the sparse structure is the key precondition for reversal: the same trigger optimization fails on standard SFT models, confirming that structure rather than trigger design is the operative factor. These results provide direct evidence that the learned carriers are causally necessary for the behaviors, pointing to a new direction for systematically localizing and selectively suppressing SFT-induced behaviors in deployed models. Code is available at https://anonymous.4open.science/r/sft-reverse-8150/.
Abstract: We study the complexity of smoothed agnostic learning, in which the learner competes with the best classifier in a target class under slight Gaussian perturbations of the inputs. Specifically, we focus on the prototypical task of agnostically learning halfspaces under subgaussian distributions in this model. The best known upper bound for this problem is based on L_1-polynomial regression and has complexity d^\tilde O(1/\sigma^2)\log(1/\epsilon), where \sigma is the smoothing parameter and \epsilon is the excess error. Our main result is a Statistical Query (SQ) lower bound showing that this upper bound is close to best possible. In particular, we prove that, even for Gaussian marginals, any SQ algorithm for smoothed agnostic learning of halfspaces requires complexity d^\Omega(1/\sigma^2+\log(1/\epsilon)). This is the first non-trivial computational lower bound for this task, and it nearly matches the known upper bound. At a conceptual level, we show that the complexity of the problem is governed by the low-degree L_1 approximation of the smoothed target T_\sigma f, so that applying L_1-polynomial regression to the smoothed function is essentially optimal in the SQ model. Our proof proceeds by constructing a moment-matching hard distribution via linear programming duality; the dual program corresponds exactly to finding a low-degree approximating polynomial for T_\sigma f, which is the same approximation-theoretic condition underlying the upper bound. To instantiate this framework for halfspaces, we prove explicit lower bounds on the approximation degree of the smoothed sign function. A key ingredient is a new structural result: we construct a distribution that matches moments with a Gaussian while exhibiting periodic structure. This result underlies our 1/\sigma^2-degree lower bound and may be of independent interest.
PaperID: 3536, Poster
Abstract: Many applications require joint prediction of interdependent behavioral choices, yet existing models often treat each choice independently (e.g., through parallel prediction heads), overlooking the influence of one on the other. In this work, we propose Progressive Dual-Head Transformer (PDHFormer), a novel framework that performs two-step prediction: the model first estimates one choice and then conditions the second on this upstream estimate through an explicit head-to-head pathway. A shared encoder captures the common structure of two prediction tasks, while the dual-head module explicitly reflects cross-choice dependence. A gated residual mechanism integrated into the embedding layer and the dual-head module further improves the training stability and the prediction performance. Extensive experiments on real-world urban mobility, manufacturing, and food delivery application domains demonstrate that PDHFormer consistently outperforms state-of-the-art machine learning models, deep tabular models, as well as parallel-head Transformer variants across multiple metrics. Moreover, our ablation study confirms that both the proposed progressive dual-head and gated residual mechanism are key contributors to the observed gains in different prediction tasks.
PaperID: 3537, Poster
Abstract: We investigate the problem of learning neural operators for coupled multi-physics PDEs, where multiple physical processes co-evolve and influence one another via bidirectional coupling. For example, in subsurface CO_2 storage, fluid flow, heat transfer, and geo-mechanical deformation evolve in a tightly coupled manner. Prior work, including neural operators and generative PDE solvers, mostly focuses on single-physics settings or relies on iterative inference, limiting their effectiveness in coupled PDEs. This paper introduces coupled-flow, a one-step generative neural operator that revisits coupled operator learning as solving a distributional transport problem. It parameterizes process-specific average velocity fields conditioned on the full coupled state, enabling cross-process interactions to emerge naturally from the transport dynamics. We further establish a Wasserstein-2 generalization bound that ties the prediction error of coupled-flow to its training losses through the regularity of the induced flow. Experiments on a wide variety of multi-physics benchmarks demonstrate that coupled-flow achieves significant performance improvements over coupled neural-operator and generative PDE baselines.
PaperID: 3538, Poster
Abstract: Vision-Language-Action Models (VLAs) enable robots to perform manipulation tasks by leveraging Vision-Language Models (VLMs) to process visual observations and language instructions. However, the action prediction objective itself lacks fine-grained task-specific perceptual supervision, making it prone to relying on task-irrelevant visual shortcuts, leading attention away from target objects and key interaction areas, thereby limiting its operational performance and generalization ability. To address this, we propose Automatic Perception-Guided Attention Alignment for Robot Manipulation (\textA^3-VLA), a framework that derives task-specific visual guidance from the manipulation process itself. Specifically, \textA^3-VLA first generates heuristic candidate regions for target objects and the robot from interaction-induced visual changes, and refines them with a visual segmentation model offline. These regions serve as explicit perceptual guidance through two complementary objectives: an attention allocation loss that encourages sufficient focus on task-critical regions when language or robot state tokens query the visual input, and a contrastive learning loss that aligns the representations of candidate target regions with language instructions while separating them from task-irrelevant regions. Extensive experiments on simulated and real-world manipulation tasks demonstrate that \textA^3-VLA consistently improves multiple VLA backbones, with an average success rate gain of 6.5% - 25.2% across benchmarks.
PaperID: 3539, Poster
Abstract: Diffusion models, particularly Diffusion Transformers, achieve strong image and video generation quality but remain expensive at inference time due to repeated denoiser evaluations. Feature caching offers a deployment-friendly acceleration strategy by reusing or predicting intermediate representations without retraining the generator or modifying the sampler. However, existing cache-then-forecast methods often rely on Taylor-style finite-difference extrapolation, which can become unstable over longer cache intervals, and typically predict local module outputs whose errors may accumulate through subsequent denoiser blocks. We propose BeziCast, a training-free diffusion acceleration framework that forecasts output-proximal denoising representations using low-order B\'ezier trajectories. BeziCast estimates B\'ezier control points via Tikhonov-stabilized fitting, providing a smooth, capacity-controlled temporal parameterization that avoids explicit high-order derivative estimation and decouples trajectory capacity from cache-query density. To moderate aggressive extrapolation, we further derive equivalent forecasting weights and introduce a convex-hull-inspired guardrail that detects high-risk predictions and softly pulls them toward a conservative simplex-projected estimate. Extensive evaluations across advanced image and video diffusion models demonstrate the effectiveness of BeziCast. In particular, BeziCast accelerates FLUX.1 by up to (4.79×) and HunyuanVideo by up to (4.11×), while preserving substantially better generation quality than competing acceleration baselines.
PaperID: 3540, Poster
Authors:
Leonardo Matone, Ben Abramowitz, Ben Armstrong, Avinash Balakrishnan, Nicholas MatteiAbstract: Recently, social choice theory has been found broadly useful for AI alignment and evaluation, often based on the axiomatic properties for combining preferences provided by social choice functions. However, for many sets of axioms, functions which never violate the axioms are known to not exist, e.g., Arrow's Impossibility Theorem. Recent work on learning novel rules to minimize axiom violations shows the effectiveness of machine learning but explores a limited class of axioms. In this work we show the effectiveness of standard social choice data structures as features for learning existing voting rules, resulting in networks that have better accuracy than past approaches while using an order of magnitude fewer parameters. Subsequently, we develop a process for adding axiomatic properties to existing voting rules. By building novel axiomatic loss functions we are able to use transfer learning to fine-tune existing voting rules in a way that quantitatively preserves their original behaviour while also exhibiting strong adherence to previously absent axiomatic properties, resulting in novel voting rules that are closer to the impossibility frontier than those in the machine learning or theoretical literature.
PaperID: 3541, Poster
Abstract: Designing regulatory DNA with cell-type-specific activity is broadly relevant for cell engineering and gene therapy. The regulatory activity of DNA elements such as promoters and enhancers emerges from the combinatorial organization of transcription factor binding sites and sequence context, which collectively constitute the genome's regulatory grammar. Deep learning models can effectively predict DNA regulatory activity across cell types, but current generative approaches often produce DNA sequences that strongly deviate from this natural regulatory grammar, increasing the risk of unexpected behavior and off-target effects in real-world applications. Here, we introduce DNA-CRAFT, a method for designing regulatory DNA with high predicted cell-type-specific activity while preserving naturalness. DNA-CRAFT formulates regulatory sequence design as a guided search over a generative model pre-trained on millions of naturally occurring regulatory elements from the human and mouse genomes. Our method enables targeted exploration of DNA sequence space, efficient evaluation of candidate designs, and explicit suppression of unwanted activity. Across benchmarks spanning human cell lines and primary immune cells, DNA-CRAFT outperforms existing generative and optimization-based methods, consistently achieving the best trade-off between cell-type specificity and naturalness in the designed regulatory DNA sequences.
PaperID: 3542, Poster
Abstract: We study fixed-confidence best-arm identification under strict 1-bit feedback constraints. At each round, the learner selects an arm and a query set, and receives only a single bit indicating whether the sampled reward belongs to that set. We consider a distribution-free finite-variance setting with arm-wise localization, where direct empirical mean estimation is no longer available and clipping becomes unavoidable. We first formulate a time-uniform 1-bit mean-estimation primitive based on randomized threshold queries and a clipped tail-integral identity. We then embed this primitive into candidate-challenger best-arm identification algorithms. A fixed-clipping algorithm gives a simple anytime (\epsilon,\delta)-PAC guarantee, while a phased adaptive-clipping algorithm matches the clipping level to the current resolution and yields a gap-adaptive sample complexity. We also prove a two-arm information-theoretic lower bound showing that the logarithmic penalty caused by finite-variance 1-bit feedback is intrinsic. Consequently, the phased algorithm is optimal in the leading dependence up to lower-order \log\log factors.
PaperID: 3543, Poster
Abstract: Understanding when neural networks learn interpretable representations is central to mechanistic interpretability. We study this through identifiability: when can supervised learning recover the true latent variables underlying the data? We introduce two data generating processes distinguished by their causal direction where either labels cause latents (generative) or latents cause labels (discriminative). When the representation has at least as many dimensions as the latent space, we prove that the optimal encoders are linear in the latents. Conversely, when dimensions are insufficient, identifiability generically fails and encoders become nonlinear. We then show that additional constraints yield recovery of individual latent variables, not just linear mixtures. In the generative setting, non-negative activations and energy efficiency force each neuron to encode exactly one latent, providing a precise account of selective grandmother cell-like responses. In the discriminative setting, when latents are sparse and the true measurement matrix satisfies the Null Space Property, the learned encoder inherits compressed sensing structure, enabling sparse dictionary learning to perfectly recover individual latents. This constitutes, to our knowledge, the first identifiability result for undercomplete nonlinear independent component analysis. Our framework unifies nonlinear ICA, compressed sensing, and neural network interpretability, delineating exactly when representations transition from nonlinear distortions to linear representations and, under additional constraints or steps, to clean, axis-aligned codes. Theoretical predictions are validated on synthetic benchmarks and on activations drawn from vision models, large language models, and macaque inferotemporal cortex.
Authors: Isaac David, Arthur Gervais
Abstract: Autonomous security agents use language models to inspect code, call tools, and test vulnerabilities inside authorized environments. Most safety evaluations ask whether a model refuses harmful single-turn requests. We study a different failure mode: whether safety alignment changes the evidence-grounded behavior of an agent that is already operating in a local sandbox with fixed tools and success checks. We evaluate regular safety-aligned Gemma 4 models and uncensored Gemma 4 derivatives in the same security-agent harness. Each model receives authorized vulnerability-analysis tasks, and each run is scored from saved traces rather than self-reported success. We measure task completion, refusals, unsafe actions, and whether the final artifact is grounded in the relevant files, symbols, and vulnerability evidence. The key finding is that the gap is not mainly visible refusal. The aligned Gemma 4 conditions usually continue working, but their artifacts are less likely to satisfy security-specific evidence checks. The uncensored Gemma 4 condition more often finds the relevant code, identifies reachability, and writes usable security reports. Clear authorization in the prompt does not recover this behavior. At the same time, all conditions fail the hardest proof-of-trigger and patch-verification tasks. These results suggest that safety alignment effects in autonomous security agents cannot be understood only as refusals: they also appear as weaker grounding and less specific defensive evidence. Removing alignment recovers some behavior, but does not make the agent reliable or safe.
PaperID: 3545, Poster
Abstract: Large language models solve complex problems by generating lengthy chains of explicit reasoning tokens. While effective, this makes reasoning expensive, length-sensitive, and constrained to (discrete) natural language. While latent reasoning offers a continuous alternative, determining useful structures for intermediate latent states is an open challenge. In this paper, we formulate latent reasoning as a geometric path-approximation problem within the model’s pretrained token-embedding space. We introduce Geometric Latent Reasoning (GLR), which uses a lightweight transition head to predict iterative direction updates in embedding space. Using textual chain-of-thought traces as anchors, GLR learns to approximate discrete reasoning trajectories while permitting continuous deviations from exact token embeddings. Evaluations on mathematical reasoning benchmarks using Qwen3 models reveal an emergent phenomenon: geometric latent reasoning induces substantially shorter generations without an explicit length objective. By replacing early explicit reasoning with continuous latent steps, models often reach correct answers using substantially fewer total generation steps. These findings suggest that continuous trajectories act as compact intermediate reasoning states, exposing a new tradeoff between latent computation budget, output length, and accuracy.
Abstract: Cross-city transfer leverages labeled data from well-instrumented cities to improve prediction in label-scarce ones, but remains challenging when cities adopt incompatible partitions with no ground-truth region correspondences. Even with expressive GNN encoders, transfer quality varies dramatically across methods sharing nearly identical backbones---indicating that alignment design, not encoder capacity, is the binding constraint. Existing paradigms exhibit complementary failure modes: heuristic anchor matching collapses to hubness under unequal partitions, while distribution-level matching over-mixes embeddings under heterogeneity. Both stem from a single missing primitive---explicit, mass-controlled soft correspondence between unequal region sets. The challenge intensifies in multi-source transfer, where independent source-to-target alignments yield conflicting gradients and source domination. We propose SCOT, which adapts entropic OT to this regime through three application-specific designs: an OT-weighted contrastive objective that resolves the geometric--semantic tension, a one-sided cycle regularizer respecting the rectangular n_s\!\neq\!n_t geometry, and---as our central contribution---a shared prototype hub coordinated through balanced entropic OT under a target-induced prior, bypassing the source-selection problem in label-scarce regimes. Across real-world cities and tasks, \scot consistently improves transfer accuracy, achieving 5--50% relative MAE/MAPE reductions over the strongest baseline, with learned couplings and hub assignments quantitatively confirming that the diagnosed failure modes are resolved.
PaperID: 3547, Poster
Authors: Sanjivni Rana, Suraj Shetiya, Senjuti Basu Roy, Gautam Das
Abstract: Performing regression is challenging in applications where training data is no longer aligned, arising due to anonymization, data fragmentation, or multi-source aggregation. Motivated by such scenarios, we study two related yet computationally distinct problems that arise when the correspondence between d-dimensional feature vectors and responses in linear regression is unknown. The first problem, referred to as the \emphShuffled Linear Regression (SLR) problem, originally proposed in the machine learning community, considers the setting where feature vectors arrive as internally consistent d-dimensional units but their correspondence to response values is unknown. Prior work established an efficient exact solution only for the one-dimensional case. For higher dimensions (d \geq 2), the best known result is a (1+\varepsilon) approximation scheme with running time \mathcalO((n/\varepsilon)^d), where n is the number of data points, leaving open the fundamental question of whether an exact polynomial-time algorithm exists for any fixed dimension d. \emphDespite the extensive study in prior literature and established NP-hardness results for unbounded dimensions, the exact computational complexity of this problem for bounded d has remained unresolved. We resolve this open problem by providing the first exact polynomial-time algorithm via a geometric insight: the parameter space can be partitioned by \binomn2 separator hyperplanes into O(n^2d-2) regions, within each of which the optimal pairing remains unchanged. This reduces the exponential search over permutations to a polynomial search over regions, yielding the first exact polynomial-time algorithm for SLR with runtime O(n^2d-1(d^2 + \log n)), without any distributional assumptions. We initiate the study of the second problem, referred to as the \emph\furlong (\fur) problem, which considers a more challenging setting where neither the feature vectors themselves nor their correspondence to the response variable are internally aligned. For the \fur\ problem, we provide the first rigorous complexity and algorithmic results, proving that it is NP-hard even at d = 2. This establishes a qualitative gap : \fur\ is intractable even in low dimensions, whereas SLR admits an exact polynomial-time solution for fixed d.
Abstract: AI video generation is moving beyond general generation, which relies on exhaustive prompt-engineering and "cherry-picking", towards fine-grained, controllable generation and high-fidelity post-processing. In professional AI-assisted filmmaking, the core requirement is the ability to perform precise, targeted modifications. A key task is video instance insertion, which requires precise spatial-temporal placement, physically consistent scene interaction (e.g., shadows and reflections), and the faithful preservation of original dynamics - all achieved under minimal user effort. In this paper, we propose PISCO, a video diffusion model for precise video instance insertion with arbitrary sparse keyframe control. PISCO allows users to specify a single keyframe, start-and-end keyframes, or sparse keyframes at arbitrary timestamps, and automatically propagates object appearance, motion, and interaction. To stabilize generation under sparse conditioning, we introduce Variable-Information Guidance and Distribution-Preserving Temporal Masking, complemented by geometry-aware conditioning. We further construct PISCO-Bench, a benchmark with verified instance annotations and paired clean background videos, and evaluate performance using both reference-based and reference-free perceptual metrics. Experiments demonstrate that PISCO consistently outperforms existing baselines and scales effectively as additional control signals are provided.
PaperID: 3549, Poster
Authors: Pratik Singh, Mohammadi Zaki, Akash Saha, Aneesh Mukkamala, Pankaj Wasnik
Abstract: There is growing interest in applying Parameter-Efficient Fine-Tuning (PEFT) to Continual Learning, as it enables faster training, mitigates catastrophic forgetting, and facilitates better adaptation to new tasks. Prior approaches have explored a range of strategies, including extending new LoRA branches for each incoming task, training under orthogonal loss objectives across all task-specific LoRA modules, and employing gating mechanisms to integrate new and existing LoRA modules. However, these methods incur memory costs that grow linearly with the number of tasks, limiting their scalability. More recent approaches investigate training a single LoRA by leveraging asymmetries in its parameterization; however, they often suffer from insufficient representational alignment across the LoRA parameter space. To address these limitations, we propose Stiefel Optimization and Aligned Rotation (SOAR), a two-stage framework for continual LoRA learning. First, we constrain the \mathbfA matrix to lie on the Stiefel manifold, thereby reducing coordinate misalignment to the space of orthogonal rotation matrices. Second, we introduce an Orthogonal Procrustes alignment procedure for the \mathbfB matrix to estimate the optimal rotation, coupled with a per-rank retention merging strategy that quantifies directional agreement between representations of new and previously learned tasks. We conduct extensive experiments across multiple models and benchmarks, demonstrating that SOAR achieves strong continual learning performance without incurring additional memory overhead.
PaperID: 3550, Poster
Authors: Konstantin Zyryanov, Nikita Lotts
Abstract: Physical adversarial attacks expose important limitations of modern vision systems, but most existing methods rely on static patches or textures. We introduce a new problem setting: closed-loop physical adversarial control, where an agent interacts with a real-world perception system via a physical channel and receives only detector-level feedback. We optimize LED patterns with PR-SAC in a high-dimensional continuous action space. We show that latent action parameterization with deterministic upsampling improves training stability compared to direct control of individual LEDs. Across YOLOv8, Faster R-CNN, and RetinaNet, the learned patterns reduce person-detection confidence, decrease mAP@0.5, and exhibit partial cross-model transfer. We further analyze robustness under distance and viewpoint changes, interaction with a defense mechanism, and failure modes such as Q-drift, entropy collapse, and degradation after initial improvement. Our results suggest that dynamic physical attacks should be evaluated not only by best-case attack strength, but also by stability, transferability, and robustness under real-world variation.
PaperID: 3551, Poster
Abstract: Designing tool-integrated reasoning agent systems for complex chemical tasks remains a fundamental challenge. Although recent chemical agents have demonstrated promising performance through the integration of external tools, they are still built on the autoregressive paradigm, which is inherently constrained by causal attention and left-to-right generation. As a result, they often struggle with bidirectional molecular understanding, global coordinated tool planning and multi-source evidence integration. To address this challenge, we introduce ChemDiffAgent, the first masked diffusion language agent for tool-integrated chemical reasoning. It reformulates chemical reasoning as iterative denoising over interleaved reasoning and tool-use trajectories, enabling tool-use decisions to be jointly determined under full-trajectory context and thereby better aligning the agent's decision boundary with its knowledge boundary. We further analyze how this paradigm benefits agentic reasoning through harder infilling subproblems during training and flexible arbitrary-order decoding during inference. Then we develop a two-stage post-training framework, comprising agentic supervised fine-tuning and variance-reduced preference optimization, to enhance tool-use and reasoning capabilities. We also construct a comprehensive benchmark covering six chemistry tasks across both single-turn and multi-turn tool-use scenarios. Under matched budgets and evaluation, ChemDiffAgent consistently outperforms autoregressive agents, achieving an overall improvement of approximately 20%. Notably, our 8B model achieves performance comparable to, in some cases surpassing, significantly larger frontier models, such as GPT 5.5 and Claude Opus 4.7.
PaperID: 3552, Poster
Abstract: LLM multi-agent systems rely on division of labor to solve complex tasks, but their execution can be blocked by contention for shared resources such as GPU test environments, CPU execution slots, and exclusive tool sessions. Existing workflows, task allocation methods, and runtime schedulers usually arbitrate such access indirectly through workflow position, dependency structure, queue state, or resource metrics. These criteria do not fully capture the value of satisfying a specific request, because request value is distributed across agents' local execution states and changes with tests, retries, and downstream results. We propose a market-based protocol for access arbitration in LLM multi-agent systems. Each agent converts its current plan and local execution state into a structured, budget-constrained bid that signals relative urgency. For asynchronous requests, an online clearing method based on optimal stopping decides when to close the clearing window, allocate access rights, settle payments, and track resource occupation until release. Across three contention scenarios, Market reduces makespan by 5.2%, 2.5%, and 4.9% relative to the strongest primary baselines, while improving task score by about two percentage points.
PaperID: 3553, Poster
Authors: Sanbi Luo
Abstract: Foundation Models (LLMs/VLMs) exhibit strong semantic reasoning capabilities but remain challenged by situated planning in unstructured physical environments. A key limitation is the Semantic-Geometric Gap: while models interpret linguistic and visual intent, they lack explicit grounding in continuous spatial structures, yielding physically infeasible or unsafe plans. We propose Causal-Geo, a neuro-symbolic framework that bridges this gap by integrating LLM reasoning with rigorous geometric planning. Rather than treating physical constraints heuristically, we model the environment as a continuous metric tensor field. LLMs translate high-level semantic intents into potential functions that dynamically warp this manifold via conformal scaling. This casts intent-driven planning as a geodesic optimization problem, solved efficiently via discrete graph search. Experiments across diverse domains—from a macro-scale 12.2 \text km^2 wetland digital twin to micro-scale indoor robotic navigation—demonstrate that Causal-Geo enables robust zero-shot planning. It significantly outperforms reinforcement learning in sparse-reward settings and prevents the ``physical hallucinations'' common in representative hierarchical LLM planners. Crucially, Causal-Geo acts as a rigorous physical gatekeeper, guaranteeing invariant safety against foundation model variance and demonstrating graceful degradation under contradictory prompts. These results establish continuous geometric grounding as a principled, domain-agnostic pathway toward physically consistent situated agent planning.
PaperID: 3554, Poster
Authors: Huadong Xiong
Abstract: Large language models (LLMs) increasingly help people solve problems, from debugging code to repairing machinery. This process requires generating plausible hypotheses from partial descriptions, then updating them as more information arrives. Yet how LLMs perform this form of inference, and how close it is to optimal, remains unclear. We study this question in the number game, a controlled setting in which a learner infers the hypothesis supported by a few positive integers, such as \16, 8, 2, 64\: a rule like powers of 2 or an interval like numbers near 20. We measure the posterior over hypotheses using three complementary probes: posterior prediction, hypothesis evaluation, and hypothesis generation. We then compare LLM behavior with an optimal Bayesian model and human behavior, and test whether the same posterior is expressed across probes. LLMs are often well described by a two-parameter Bayesian fit, but with systematic offsets: by default they show a strong-sampling assumption that creates an implicit Occam's razor, favoring narrower hypotheses, while thinking mode shifts them toward greater prior reliance. We also find a robust evaluation--generation gap: LLMs select more correct hypotheses during hypothesis evaluation but generate simpler, more rule-like hypotheses. Finally, this Bayesian-with-bias pattern does not extrapolate. Models can behave as if they hold rule-like hypotheses over observed examples, yet generalize poorly to parts of the hypothesis domain not covered by those examples. Our results highlight a limitation of LLMs as general problem solvers, especially for scientific inference, where hypotheses must go beyond the data.
PaperID: 3555, Poster
Authors:
Pengxiang Ding, Haoying Wang, Minghui Lin, Qishen Wang, Zhenyu Ding, Runze Suo, Xuanxuan An, Wenxuan Song, Fuhao Li, Han Zhao, Donglin Wang, Ning DINGAbstract: Vision-Language-Action (VLA) models transfer semantic priors from large-scale vision-language models to robotic manipulation, but their backbones are typically optimized for static visual understanding rather than action-conditioned dynamics. Consequently, the features used by the policy may lack predictive information about how the scene evolves under the robot's actions. Existing future-prediction methods either reconstruct pixels, overemphasizing appearance details irrelevant to control, or use auxiliary latent predictors that remain weakly coupled with the policy-consumed backbone features. We propose PredMind, a plug-and-play predictive representation learning framework that treats future prediction as a training signal for the VLA backbone itself. Instead of reconstructing future pixels, PredMind aligns intermediate backbone representations with their future counterparts conditioned on executed actions, shifting supervision from low-level appearance to the model's semantic feature space. This directly injects action-conditioned temporal structure into the representation hierarchy used for control while preserving pretrained VLM priors. All auxiliary prediction components are discarded after training, introducing no architectural changes, additional parameters, or inference latency at deployment. Across simulation benchmarks, real-world manipulation tasks, and multiple VLA backbones, PredMind consistently improves learning efficiency and final performance. Especially on LIBERO, it matches or surpasses a fully trained base VLA using only 1/60 of the training iterations, demonstrating the benefit of embedding predictive dynamics directly into backbone features used for control. Anonymous code is available at https://anonymous.4open.science/r/PredMind.
PaperID: 3556, Poster
Abstract: Most time-series foundation models operate on fixed temporal grids, representing signals as vectors, tokens, or patches and learning maps between discretized observations. This grid-to-grid view is convenient for standard forecasting, but it is not the natural object for reusable time-series modeling, where temporal resolution, forecast horizon, and query locations may change across tasks and deployments. We argue that time-series foundation models should instead learn sample-conditioned function-to-function operators. Given finite observations of an underlying signal, such an operator maps the observed context to an output function that can be evaluated at user-specified times. We propose TimeOperator, a compact branch-trunk backbone that realizes this view with a shared context encoder, spectrally informed query features, context-adaptive trunk parameters, and multi-basis temporal responses. Forecasting and imputation are handled by evaluating the same operator on different query sets, while classification reuses the shared context representation. Across deterministic and probabilistic forecasting, imputation, and classification benchmarks, TimeOperator is competitive with existing time-series foundation models. Further analyses show that cross-frequency and cross-horizon prediction can be handled by changing only the query set, rather than changing the decoder or relying on post-hoc resampling.
Abstract: Geometric representation-conditioned molecule generation provides an effective paradigm that decouples molecule representation modeling from structure generation. By decoupling molecule generation into two stages—first generating a meaningful molecule representation, and then generating a 3D molecule conditioned on this representation—the efficiency and quality of the generation process can be significantly enhanced. However, its effectiveness is fundamentally limited by the quality of the representation space: pretrained molecular encoders, such as UniMol, produce representations that are non-smooth and not fully exploited during the generative training process. In this work, we propose LENSEs, a framework that better exploits the potential of molecule representations in representation-conditioned generation methods. In particular, LENSEs introduces three complementary mechanisms: (1) a representation head, simultaneously trained during generative tasks, that extracts multi-level representations from the pretrained encoder; (2) a molecule perceptual loss that optimizes the generator in a semantic-informative representation space; and (3) a node-level representation alignment (REPA) loss that explicitly aligns the generator’s hidden states with encoder representations, reducing the semantic gap between pretraining and generation. We demonstrate the effectiveness of these improvements through extensive molecule generation tasks. Specifically, on the challenging molecule generation dataset GEOM-DRUG, LENSEs achieves 97.28% validity and 98.51% molecule stability, surpassing existing advanced methods. Further analyses through Lipschitz constant reduction (4.6×) and QM9 probing tasks also demonstrate the smoother, more informative refined representations, establishing generative training with alignment objectives as a potential pretraining paradigm for molecular encoders.
Abstract: Large-scale pre-trained vision models provide powerful representations for class-incremental learning, yet continuously adapting them to new classes without historical samples, task identifiers, or unrestricted parameter growth remains a fundamental challenge. Existing LoRA-based approaches typically follow either a one-adapter-per-task expansion paradigm or a fixed parameter-sharing strategy. The former leads to linear parameter growth as the task sequence expands, while the latter often induces severe cross-task interference. To address these limitations, we propose \bf E^3 (Elastic Ensemble of Experts), a parameter-efficient framework for class-incremental learning. \bf E^3 treats LoRA modules as elastic experts whose task capacity is dynamically scheduled according to representational saturation. Within each expert, we further introduce Group-Aware Parameter Partitioning, which allocates LoRA parameters into disjoint task-specific subspaces using magnitude-based importance estimation, thereby mitigating parameter overwriting without additional gradient-sensitivity computation. Moreover, \bf E^3 incorporates hierarchical moment regularization and orthogonal gradient projection to constrain classifier drift and prevent new-task updates from disrupting historical feature subspaces. Extensive experiments on standard class-incremental learning benchmarks demonstrate that E³ achieves state-of-the-art or competitive performance, especially in long-horizon task sequences, while maintaining a favorable balance between stability, plasticity, and parameter efficiency.
PaperID: 3559, Poster
Abstract: Tree ensembles such as random forests (RFs) and gradient boosting machines (GBMs) are among the most widely used supervised learners, yet their theoretical properties remain incompletely understood. We adopt a spectral perspective on these algorithms, with two main contributions. First, we derive minimax-optimal convergence for RF regression, showing that, under mild regularity conditions on tree growth, the eigenvalue decay of the induced kernel operator governs the statistical rate. Second, we exploit this spectral viewpoint to develop compression schemes for tree ensembles. For RFs, leading eigenfunctions of the kernel operator capture the dominant predictive directions; for GBMs, leading singular vectors of the smoother matrix play an analogous role. Learning nonlinear maps for these spectral representations yields distilled models that are orders of magnitude smaller than the originals while maintaining competitive predictive performance. Our methods compare favorably to state of the art algorithms for forest pruning and rule extraction, with applications to resource constrained computing.
Abstract: Multi-reference image generation aims to synthesize images from textual instructions while faithfully preserving subject identities from multiple reference images. Existing VLM-enhanced diffusion models commonly rely on decoupled visual conditioning: semantic ViT features are processed by the VLM for instruction understanding, whereas appearance-rich VAE features are injected later into the diffusion backbone. Despite its intuitive design, this separation makes it difficult for the model to associate each semantically grounded subject with visual details from the correct reference image. As a result, the model may recognize which subject is being referred to, but fail to preserve its identity and fine-grained appearance, leading to attribute leakage and cross-reference confusion in complex multi-reference settings. To address this issue, we propose UniCustom, a unified visual conditioning framework that fuses ViT and VAE features before VLM encoding. This early fusion exposes the VLM to both semantic cues and appearance-rich details, enabling its hidden states to jointly encode the referred subject and corresponding visual appearance with only a lightweight linear fusion layer. To learn such unified representations, we adopt a two-stage training strategy: reconstruction-oriented pretraining that preserves reference-specific appearance details in the fused hidden states, followed by supervised finetuning on single- and multi-reference generation tasks. We further introduce a slot-wise binding regularization that encourages each image slot to preserve low-level details of its corresponding reference, thereby reducing cross-reference entanglement. Experiments on two multi-reference generation benchmarks demonstrate that UniCustom consistently improves subject consistency, instruction following, and compositional fidelity over strong baselines. Our code, checkpoints, and the training dataset will be released soon.
PaperID: 3561, Poster
Authors: Michael West, Eduard Dragut
Abstract: Researchers often need web corpora that are not answers to a single query, but reusable collections of documents spanning many decentralized sources. We refer to this problem as thematic web data collection: assembling semantically relevant documents that share a common theme but are scattered across heterogeneous and structurally disconnected regions of the web. Existing approaches, including traversal-based methods and LLM-driven web agents, are limited by local exploration or short-horizon stopping, preventing comprehensive collection. We formulate this task as a long-horizon information foraging problem over query-induced semantic patches, where the central challenge is deciding when to continue local exploitation and when to switch to new regions. We propose PatchScout, a multi-agent framework that instantiates a class of patch-switching policies to balance local exploitation and adaptive patch switching under a fixed budget. Experiments on four live-web topics show that PatchScout substantially improves yield and domain coverage over traversal-based and search-augmented reasoning agent baselines. In an applied social science setting with an incomplete prior collection, PatchScout discovers 102 previously unidentified entities, illustrating its ability to expand existing thematic datasets.
PaperID: 3562, Poster
Abstract: Joint classification of hyperspectral imagery (HSI) and Light Detection and Ranging (LiDAR) data benefits from complementary spectral--geometric cues, but existing fusion methods largely assume locally consistent cross-modal features and symmetric pixel-wise trust, which is often violated in real urban scenes. We term this pixel-level phenomenon \emphcross-modal local inconsistency (CMI) and quantify it with a Strong-Edge CMI Ratio: dual-modal gradient analysis shows that over 16% of strong-edge pixels exhibit single-modality dominance across three benchmarks, while ROI-level analysis reveals the class-entanglement induced by symmetric fusion. To address this issue, we propose CMI-Trans (Cross-Modal Inconsistency-aware Transport), which promotes per-pixel reliability modeling from a post-hoc diagnostic to a structural signal for alignment and training. CMI-Trans combines Energy-Calibrated Reliability (ECR), Uncertainty-Shaped Transport Alignment (USTA) with Sinkhorn-regularized optimal transport, and Triadic Co-Optimization (TCO) that jointly optimizes classification, uncertainty consistency, and alignment stability on a Mamba-based fusion backbone. On Houston, Trento, and MUUFL, CMI-Trans achieves competitive overall accuracy and substantial gains on CMI-affected classes (e.g., +17.90% on Houston C10). The ECR-derived uncertainty score also separates reliable from erroneous predictions (\rho(U,e)<0; top-20% highest-U error rate \le 0.18%), while improving feature separability in shadow-induced CMI regions. These results show that explicitly modeling CMI through uncertainty-shaped adaptive alignment yields more accurate and interpretable multimodal remote sensing classification.
PaperID: 3563, Poster
Authors: Ruiyu Yan, Bowen Chen, Shaowen Wan, Lin Zhao
Abstract: Current vision-language models (VLMs) encode visual information in dense hidden states where object identity, spatial layout, and local attributes are implicitly entangled rather than explicitly disentangled, limiting their ability to isolate and modulate the specific visual evidence required by a given language query. Inspired by sparse population coding and top-down modulation in biological vision, we introduce NeuronEye, a plug-in framework that constructs a sparse, concept-level neuron vocabulary from intermediate VLM representations and selectively activates query-relevant visual concepts during inference. NeuronEye decomposes vision-token states into an overcomplete sparse basis organized by concept-level clusters, uses the language query to activate relevant clusters and localize the patches where selected concepts are expressed, and injects the focused evidence back into vision tokens. A complementary suppression mechanism attenuates dominant perceptual directions to preserve weaker but relevant cues. All operations run in a single forward pass over a frozen VLM backbone. On Qwen2.5-VL-7B, NeuronEye raises CV-Bench overall accuracy by +3.1 with gains of +9.5 on Distance, and improves BLINK Multi-view by +8.3, with similar trends on LLaVA-1.6-7B. These results suggest that sparse neuron vocabularies can serve not only as post-hoc interpretability tools but also as active interfaces for concept-level visual reasoning.
PaperID: 3564, Poster
Abstract: Existing deep learning architectures for survival prediction from electronic health records (EHRs) largely assume unimodality, failing to capture the complex relational structure across diverse clinical modalities. Graph learning can allow us to robustly adapt to the underlying clinical structure of the data to have more predictive and interpretable representations for time-to-event analysis. However, unconstrained graph learning produces dense, bidirectional adjacencies that overfit. We introduce M-CADENCE, a multimodal survival model that integrates structured acyclic priors with attention-based propagation. Each modality has its own learnable adjacency, regularised toward acyclicity through a log-determinant characterisation and toward sparsity through soft thresholding. A cross-modal meta-adjacency keeps the cross-modal parameter count quadratic in the number of modalities rather than in their feature dimensions. Propagation is then performed by multi-head attention biased by the learnt structure. The propagation step degrades gracefully under sparse or misspecified structure, addressing a known weakness of hard-routing graph variants. We evaluate M-CADENCE on five EHR benchmarks covering intensive-care mortality, circulatory failure, and emergency-department competing risks, against multiple survival, graph, and structure learning baselines. M-CADENCE achieves the best concordance on all datasets, the best early-warning AUC on four of five, and competitive precision-recall against the strongest deep survival baselines. Removing the DAG bias lowers concordance, as does substitution with a density-matched random adjacency, which confirms that the specific learnt structure is required. Discrimination is stable across four orders of magnitude in regularisation strength, and evaluation by a panel of ten language models confirms M-CADENCE learns clinical edges that are more physiologically plausible than baseline structure-learning methods.
Authors: Chen Wang, Peiran Yun, Pan Xie, Ke Deng
Abstract: Diffusion and continuous-flow generative models achieve high-quality generation, and their deterministic sampling can be formulated as solving learned ODE dynamics. However, accurate ODE discretization often requires many steps, making efficient few-step generation a key challenge. Among acceleration strategies, reflow-based distillation simplifies teacher ODE trajectories so that a student model can approximate the teacher transport with fewer steps. We identify a theoretical limitation of this paradigm, namely that trajectory matching can under-determine the distribution induced by the student model. In particular, two student models can attain the same trajectory-matching loss while inducing different endpoint marginal distributions, which may lead to different generation quality. To address this limitation, we introduce a marginal-alignment regularizer that penalizes the discrepancy between the student-induced marginal and the corresponding teacher marginal at the endpoint of each distillation interval. The regularizer is computed by tracking log-density changes along the ODE induced by the student model and evaluating scores from the frozen teacher model, without requiring auxiliary trainable networks or adversarial optimization. The resulting framework applies uniformly to the reflow family, including vanilla reflow and piecewise reflow. We further prove a telescoping total-variation bound showing that local marginal alignment controls the final-time discrepancy between the student-induced and teacher-induced distributions. Experiments on benchmark backbones demonstrate the effectiveness of the proposed method for few-step generation.
Authors: Tianyi Ma, Tengyao Wang, Richard J Samworth
Abstract: We study in-context learning problems where a Transformer is pretrained on tasks drawn from a mixture distribution \pi=\sum_\alpha\in\mathcalA \lambda_\alpha \pi_\alpha, called the pretraining prior, in which each mixture component \pi_\alpha is a distribution on tasks of a specific difficulty level indexed by \alpha. Our goal is to understand the performance of the pretrained Transformer when evaluated on a different test distribution \mu, consisting of tasks of difficulty \beta\in\mathcalA, and with potential distribution shift relative to \pi_\beta. In particular, we consider nonparametric regression problems with random smoothness, and multi-index models with both random smoothness and random effective dimension. We prove that a Transformer pretrained on such mixture distributions can adapt to the difficulty of a new task in-context and achieve the optimal prediction rate, uniformly over test distributions in a chi-squared divergence ball. Thus, the pretrained Transformer is able to achieve faster rates of convergence on easier tasks and is robust to distribution shift at test time. Finally, we prove that even if an estimator had access to the test distribution \mu, the convergence rate of its expected risk over \mu could not be faster than that of our pretrained Transformers, thereby providing a more appropriate optimality guarantee than minimax lower bounds.
PaperID: 3567, Poster
Abstract: The target-charging technique (TCT), a generalization of the sparse vector technique, monitors whether query answers deviate beyond acceptable bounds while incurring privacy costs only upon significant deviations. We develop a more mature theory of TCT through three contributions. First, we generalize TCT and replace prior simulation-based analyses with a direct stochastic domination argument, yielding a tighter bound and a modular proof structure that affords flexibility in the choice of privacy representation for the base mechanism as well as the composition accountant. Second, we provide an anytime-valid TCT guarantee that integrates with approximate Rényi DP filters. Third, we enable individual-level privacy filters. Together, these provide the first end-to-end TCT construction compatible with privacy filters and individual-level accounting, with practical relevance demonstrated through benchmarks inspired by the impending world-scale deployment of the W3C Attribution standard, which is fundamentally rooted in individual filters.
Abstract: Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many modern machine learning models. Recent work has shown that exploiting matrix structure can improve optimization dynamics. A notable example is Muon, which performs steepest descent under the spectral norm constraint. We take the next step and introduce Tensorion, a tensor-aware optimizer that extends Muon’s constrained optimization perspective from matrices to higher-order tensors. Tensorion is built around a linear minimization oracle (LMO) over a tensor norm ball. The norm is carefully chosen to balance two objectives: tightly bounding the tensor spectral norm, while still keeping the LMO tractable. This LMO becomes computable because it reduces to operations on adaptively selected unfolding matrices. Notably, when restricted to order‑2 tensors (i.e., matrices), Tensorion recovers Muon exactly. Experiments on tensor-based computer vision problems suggest that Tensorion can offer improved convergence behavior and more stable gradient updates compared with Adam-based and existing tensor-aware baselines in the evaluated settings.
Abstract: Real-world model deployment across multiple domains requires multimodal models to operate under two complementary regimes: (1) multi-task pretraining, tasks are co-available at design time where related tasks could borrow representational strength from one another, (2) continual adaptation, in which new tasks emerge after deployment with previously unseen modality combinations. However, neither regime alone suffices: the pretraining task set is never exhaustive, while bypassing joint training forfeits the transfer gains and efficiency among co-trainable tasks. Sparse Mixture-of-Experts (MoE) is a natural fit for this dual requirement: sparse activation enables modular capacity expansion as new tasks arrive, while routing decouples modality-level computation from task-level composition. In this work, we propose a scalable MoE framework for multitask pretraining and continual learning across flexible modality combinations. The framework is designed to support training on multimodal tasks with diverse modality configurations by leveraging modality-specific routers that process tokens from each modality across tasks. Furthermore, it enables continual learning over sequential multimodal tasks within a fixed-capacity MoE by compressing accumulated expert knowledge into low-rank memory subspaces, while expanding only the lightweight routers. We validate the effectiveness of our method on multiple healthcare multimodal benchmarks. It demonstrates competitive multitask pretraining performance while alleviating catastrophic forgetting and improving parameter efficiency.
PaperID: 3570, Poster
Abstract: Diffusion-based solvers have become a prominent approach to neural combinatorial optimization. At inference time, they assign each decision variable a prediction score. A greedy decoder ranks candidates by these scores and checks hard constraints to ensure feasibility. Top-ranked candidates should be both promising and jointly feasible, making prediction--constraint alignment critical. When this alignment is weak, greedy decoding skips constraint-violating candidates and falls back to lower-ranked alternatives, degrading solution quality. Existing test-time methods improve search, sampling, or decoding, but do not directly strengthen this alignment before decoding. We observe that trained denoisers already carry constraint-aligned structural signals. Building on this observation, we propose Perturbed Instance-Structure Guidance (PISG), a training-free method that amplifies these signals at test time. At each denoising step, PISG constructs a structure-degraded prediction and extrapolates the unperturbed prediction away from it. The resulting guided prediction improves candidate rankings before constrained decoding. Signal analysis shows that this guidance concentrates on constraint-critical variables rather than uniformly rescaling scores. Under greedy decoding and without downstream refinement, PISG improves all 16 TSP solver-scale pairs across DIFUSCO, T2T, FastT2T, and StruDiCO, and all four MIS-ER solvers, with gains up to +8.4%. These results identify prediction-constraint alignment as an effective test-time enhancement direction for diffusion-based combinatorial optimization.
Authors:
Minghao Yan, Bo Peng, Benjamin Coleman, Ziqi Chen, Zhouhang Xie, Shuo Chen, Zhankui He, Noveen Sachdeva, Weili Wang, Ed Chi, Shivaram Venkataraman, Wang-Cheng Kang, Derek Cheng, Beidou WangAbstract: Large language models have become drivers of evolutionary search, but most systems rely on a fixed, prompt-elicited policy to sample next candidates. This limits adaptation in practical engineering and research tasks, where evaluations are expensive, and progress depends on learning task-specific search dynamics. We introduce PACEvolve++, an advisor-model reinforcement learning framework for test-time policy adaptation in evolutionary search agents. PACEvolve++ decouples strategic search decisions from implementation: a trainable advisor generates, assesses, and selects hypotheses, while a stronger frontier model translates selected hypotheses into executable candidates. To train the advisor under non-stationary feedback, we propose a phase-adaptive approach that adapts its optimization strategy to different phases of the evolutionary process. Early in evolution, it uses group-relative feedback to learn broad search preferences; later, as reward gaps compress, it emphasizes best-of-k frontier contribution to support stable refinement. Across expert-parallel load balancing, sequential recommendation, and protein fitness extrapolation, PACEvolve++ outperforms the state-of-the-art evolutionary search framework with frontier models, achieving faster convergence and stabilizing test-time training during evolutionary search.
PaperID: 3572, Poster
Abstract: To ensure a formal guarantee that fully human-written text streams are not falsely accused as machine-generated text (MGT), the online betting framework (e-values) has been successfully applied on various existing detectors. However, current two-sample e-value methods fail in two real-world application scenarios: the false positive rate (FPR) is often out of control and the test power drops drastically when (1) the reference domain is similar but not exactly same as test domain (cross-domain), or (2) proxy model of detector is outdated compared to the modern source model. We identify the root cause as an online calibration dilemma: two-sample optimization and betting relies on a null-calibrating offset that is unidentifiable from unlabeled test streams under shift. We resolve this with a triple-sample betting framework (TriBet) that conduct online MGT detection as relative testing against both human and machine references using a witness score, and introduces a conformal test-inclusive domain-shift betting strategy to achieve FPR control without fragile online calibration. We show that TriBet is a level-\alpha sequential test with asymptotic power one, and characterize its expected stopping time. Empirically, we evaluate on 15 source LLMs, 8 proxy models and mixed-document streams, TriBet consistently delivers faster and more powerful detection than both oracle and online betting baselines while keeping FPR strictly controlled.
Abstract: Emergent misalignment (EM) has been linked to the activation of misaligned persona vectors and evil character traits, suggesting that EM operates through disruption of the model's aligned character rather than direct learning of harmful content. Motivated by this connection, we study self-generated text recognition (SGTR) finetuning as a character-targeted intervention that is distinct from existing in-training defenses. We conduct two-stage finetuning experiments across three models (GPT-4.1, Qwen2.5-32B-Instruct, Seed-OSS-36B-Instruct) and multiple EM datasets to compare SGTR finetuning against benign finetuning baselines (correct domain-specific data, general knowledge, and word counting) to find it an effective defense in both reversal and prevention settings. We find that all interventions produce comparable EM reversal, but only when restoring capabilities that EM had degraded. For prevention, only SGTR finetuning consistently reduces misalignment without exacerbating any individual metric, suggesting that character fortification specifically drives prevention. We provide further evidence for EM's relation to the LLM's default character by showing that EM finetuning induces diversity into the LLM's identity self-reports, artificially corrupting self-recognition exacerbates misalignment caused by EM finetuning, and that removing the model's identity-bearing system prompt substantially reduces the effect of EM finetuning. Together, these findings reframe EM not as the adoption of a coherent misaligned persona but as the destabilization of aligned character.
PaperID: 3574, Poster
Abstract: Autoregressive video diffusion models have demonstrated effectiveness in interactive game generation, with Minecraft gameplay serving as a representative application. To faithfully simulate gameplay, a model must generate natural content when exploring new scenes while preserving spatial consistency when revisiting explored areas. Under limited computation budgets, it must compress and exploit historical cues within a finite context window, which exposes a trade-off: models relying solely on temporal context offer flexible exploration but suffer from poor revisit consistency, whereas adding spatial memory strengthens consistency but may degrade new scene generation quality when the model over-relies on sparse or unreliable spatial context. We present GeoMemory, a learning framework that pairs training protocols with a geometry-indexed spatial memory. Specifically, our Hybrid Training exposes the model to both exploration and revisitation regimes, guiding the model to rely on temporal memory in new scenes while effectively incorporating spatial memory upon revisits. Chained Forward Training creates larger pose variations and encourages reliance on spatial memory for maintaining consistency. For spatial memory, we integrate Point-to-Frame Retrieval with an incremental 3D cache generated by VGGT, enabling constant-time retrieval of relevant historical context regardless of sequence length. Extensive experiments demonstrate that GeoMemory achieves superior performance in both long-term spatial consistency and visual quality in new scenes with real-time interaction.
PaperID: 3575, Poster
Authors:
Chaehyeon Kim, Gary Weissman, Eric WongAbstract: Transformer predictions depend on positional embeddings, which are known to produce biases such as first- and last-token dominance. However, feature attribution methods implicitly entangle influence from positions with features into a single opaque importance score per feature. To uncover these positional effects, we propose Position-aware eXplanation (PaX), a model-agnostic framework that transforms standard attribution methods to jointly produce feature and position attributions. PaX additionally produces counterfactual positional explanations: actionable scores quantifying how the prediction changes when a feature is relocated. We use PaX to generalize perturbation-, gradient-, and boundary-based attribution methods across vision, language, and clinical time-series benchmarks with no architectural changes. We demonstrate that separating positional effects improves both feature and position attribution faithfulness by 17.5% / 55% and 16.8% / 43%, respectively (insertion / deletion). In a case study on sepsis forecasting, we find that PaX recovers clinically validated bedside signals that standard attributions miss.
PaperID: 3576, Poster
Authors:
Mélanie Cambus, Darya Melnyk, Tijana Milentijević, Stefan SchmidAbstract: Byzantine-tolerant aggregation is a central challenge in federated learning, especially under heterogeneous client data where malicious updates may be statistically indistinguishable from honest but unusual clients. Robustness of aggregation algorithms is crucial to the quality of the final model. In this work, we study the robustness of the Byzantine-tolerant aggregation through the lens of centroid approximation, examining how closely an aggregation rule can match the average of honest client updates when up to t of n clients are Byzantine. We connect this objective to classical validity conditions from Byzantine agreement and show that validity alone is insufficient to guarantee good centroid approximation. Our main theoretical contribution is the first lower bound of \sqrt\min\(n-t)/t,d\/2 on the centroid approximation for aggregation under box validity and a matching upper bound of 2\sqrt\min\n,d\ in the practically relevant case n>2t. In addition, we present a new algorithm that achieves a 2d-approximation under convex validity, which also proves that the existing lower bound in the literature is tight. These results expose a fundamental tradeoff between validity conditions and approximation quality. We complement the theory with experiments in FedSGD and FedAvg under standard Byzantine attacks, showing that stronger validity improves stability, while tighter centroid approximation can improve accuracy in heterogeneous settings.
PaperID: 3577, Poster
Abstract: Proximal Policy Optimization (PPO) and it's variants are the most widely used policy-gradient methods in reinforcement learning, yet the role of its actor update mechanism is still poorly understood theoretically. In particular, standard PPO combines clipped surrogate gradients, multiple epochs of minibatch updates, and reuse of rollout data, but existing convergence analyses do not fully capture this update structure. In this work, we theoretically study PPO policy updates with symmetric clipping from a policy-gradient perspective and interpret its actor updates as a cyclic biased-gradient method with sample reuse and random reshuffling. Our first contribution is a clean formalization of PPO clipped actor updates through surrogate gradients that approximate the true policy gradient. Using the performance difference lemma, we prove a linear bias bound which quantifies how the surrogate gradient drifts as the policy moves away from the sampling policy. Our second contribution is a convergence analysis of cyclic surrogate-gradient ascent, showing that additional PPO-style biased updates can improve progress under conservative learning rates without requiring additional samples. Finally, we analyze the stochastic minibatch version with reshuffling and obtain convergence-to-stationarity guarantees under standard smoothness assumptions and a bounded critic-bias condition. Overall, our results provide a theoretical interpretation of PPO’s multi-epoch actor updates: extra clipped surrogate steps introduce bias, but can still improve optimization efficiency by compensating for small, stable step sizes through sample reuse.
PaperID: 3578, Poster
Abstract: Document QA agents do not fail only at the final answer. They can also write unsupported claims into public working memory, and those writes can affect later retrieval and synthesis. We introduce ADMIT (Auditable Decoupled Memory-write admission for Intermediate Traces), a low-budget admission controller for public OCR/text evidence packs. ADMIT applies a deterministic substring-support check S_\textstr(\haty, B) > 0 to each public commit and allows one bounded RepairRead. If repair still fails, ADMIT uses a memory-fail-closed / output-fail-open rule: it writes a \bot (NO_COMMIT) sentinel to memory while the final-answer surface may still fall back to the baseline answer. On held-out M3DocVQA, a retrained final-row router closes the answer-surface gap; ADMIT's distinguishing evidence is therefore memory-surface contamination, not final-answer accuracy. The lexical invariant guarantees zero-substring accepted intermediate contamination by construction; the contribution is therefore the public memory-write boundary as a low-cost control surface with decoupled fail actions, not the invariant itself. In a 7-arm dynamic stress diagnostic on three agents, ADMIT enforces the invariant and preserves raw EM at this operating point (zero contaminated trajectories on every agent and no detectable raw-EM loss relative to the fail-open ablation), whereas final-only baselines still leave 12--32% contamination. ADMIT enforces lexical auditability, not semantic truth: relation-level entailment, citation-aware faithfulness, and visual-region grounding are not claimed, and lower-density VisDoMSlide and text-only FRAMES are reported as boundary cases.
PaperID: 3579, Poster
Authors: Yuan Chu, Xuewu Ji
Abstract: Offline-to-Online (O2O) reinforcement learning (RL) is essential for policy adaptation but frequently suffers from initialization collapse due to evaluation and improvement mismatches across the learning transition. In safety-critical domains, safety-constrained model-free O2O RL without offline data retention remains significantly under-explored. Existing O2O fine-tuning methods, typically relying on policy regularization or value recalibration within Euclidean spaces, lack the structural capacity to strictly sequester policies from hazardous regions. To facilitate a geometric paradigm for safe policy fine-tuning, we propose Riemannian Admissibility Flow (RAF), which lifts safety constraints into a geometric framework by reformulating policy optimization as probability transport on anisotropic Riemannian manifolds. By fusing safety and uncertainty into the metric tensor, RAF transforms extrinsic scalar penalties into intrinsic topological barriers with directional gating, enabling target-free geodesic flow matching for hazard circumvention without offline data retention. Evaluation on representative safe O2O benchmarks shows that RAF achieves competitive performance in constraint satisfaction and task execution while effectively mitigating initialization collapse.
PaperID: 3580, Poster
Authors:
Qiuyu Kong, Zanxi Ruan, Marco Cristani, Yiming WangAbstract: Long-text image–text alignment is essential for fine-grained scene understanding, yet existing methods mainly focus on extending context length or improving local alignment. We identify an overlooked failure mode, saliency collapse, where models over-rely on text that features visual prominence while underutilizing non-salient context. Consequently, retrieval performance degrades sharply when salient content is reduced, despite that the non-salient context remains informative. To address this, we propose RoSA, a saliency-aware finetuning framework that decomposes image-text pairs into complementary salient and contextual views, enabling joint learning of prominent and contextual semantics. We propose a modality-independent decomposition strategy to combine object-level visual saliency estimation with LLM-based text decomposition to bridge vision and language. We employ auxiliary objectives to explicitly align these views, ensuring robust representation with both salient and contextual features. RoSA greatly improves long-text retrieval and demonstrate strong robustness to saliency collapse across benchmarks. Moreover, our auxiliary objectives are plug-and-play, bringing consistent gains to existing tuning methods. Code and checkpoints will be released upon acceptance.
PaperID: 3581, Poster
Abstract: Visible-Infrared Lifelong Person Re-Identification (VI-LReID) requires a model to sequentially acquire cross-modal person retrieval capabilities across sequentially arriving tasks while retaining performance on previously learned ones.Existing methods face two fundamental problems.First, new-task gradient updates disrupt each instance's assignment ranking over historical class prototypes, which is the direct mechanism behind past-task retrieval performance degradation. At the same time, the feature manifold inevitably undergoes global drift as a necessary byproduct of new-task adaptation. Effective anti-forgetting should target the detrimental ranking disruption while permitting this necessary drift.Second, visible RGB and thermal infrared (TIR) imaging capture fundamentally different physical cues: texture-and-color reflectance versus heat radiation contours. This causes the two modalities to form inherently inconsistent inter-class similarity structures that conventional cross-modal supervision cannot resolve, manifesting as highly asymmetric forgetting across modalities under shared-backbone continual training.We propose S-PTC (Stage-wise Prototype Topology Consistency), which preserves each instance's relative distribution over the frozen historical prototype set, protecting assignment rankings while leaving the overall feature manifold free to evolve for new-task adaptation.We further propose CM-PTC (Cross-Modal Prototype Topology Consistency), which aligns the inter-class topology matrices of RGB and IR prototypes, mitigating the structural inconsistency rooted in their imaging physics.Both modules require no stored exemplars.Extensive experiments on the VI-LReID benchmark verify the effectiveness and superiority of our approach against state-of-the-art methods.
PaperID: 3582, Poster
Abstract: Modern GPU-native solvers for combinatorial optimization exploit parallelism by evolving large populations of relaxed candidates; yet, their post-processing collapses this computational effort into a single, rounded solution. We show that the discarded replicas of Parallel Quasi-Quantum Annealing (PQQA) form an implicit elite archive that contains basin-level information recoverable after annealing. We introduce Iterative Path-Relinking Polish (I-PRP), a deterministic, training-free post-anneal operator that leaves the PQQA inner loop unchanged. I-PRP polishes the top rounded elites, traverses cost-guided relinking paths from the current best elite toward the others, re-polishes intermediate states to cross basin boundaries, and iterates only under strict improvement. The operator never degrades the paired upstream polish, terminates after a finite number of updates, and reduces to the original polish when the population collapses to a single basin. Across QUBO and graph-coloring benchmarks, I-PRP preserves saturated cases while improving PQQA on rugged multi-basin instances, revealing that parallel annealing populations contain reusable search structures beyond their best replicas.
PaperID: 3583, Poster
Abstract: Graph Neural Networks (GNNs) have achieved remarkable performance in a wide range of graph-related learning tasks. However, explaining their predictions remains a challenging problem, especially due to the mismatch between the graphs used during training and those encountered during explanation. Most existing methods optimize soft edge masks on weighted graphs to highlight important substructures, but these graphs differ from the unweighted graphs on which GNNs are trained. This distributional shift leads to unreliable gradients and degraded explanation quality, especially when generating small, sparse subgraphs. To address this issue, we propose explaining a robust surrogate whose explanations provide a faithful proxy for those of the original model. Specifically, we train this surrogate model through an iterative re-weighting strategy that alternates between identifying explanation subgraphs and adapting the model to the resulting weighted graph distribution, supporting stable explanation refinement while preserving the validity of the soft-mask relaxation. We evaluate our method on multiple benchmark datasets using different GNN backbones and explanation methods. Experimental results show that our method consistently improves explanation quality and can be flexibly integrated with different architectures.
Abstract: Understanding the behavior of stochastic gradient methods is a central problem in modern machine learning. Recent work has highlighted diagonal linear networks as a simplified yet expressive setting for analyzing the optimization and generalization properties of neural models. In this work, we show that in the high-dimensional regime, stochastic gradient descent on diagonal linear networks is well-approximated by continuous dynamics governed by a stochastic differential equation (SDE), which explicitly decouples the drift from the gradient noise. We further derive a deterministic partial differential equation whose solution propagates the relevant state of the iterates and characterizes the time evolution of a broad class of observable statistics, including the risk, curvature, and other metrics for optimality. To the best of our knowledge, this is the first deterministic high-dimensional description of SGD for a genuinely nonlinearly parametrized model. Finally, we show that, under a suitable parametrization, the stochastic dynamics are globally well posed and converge exponentially fast to zero risk with high probability, yielding a fully explicit non-asymptotic description of their long-time behavior. Numerical simulations corroborate our theoretical findings.
Abstract: While deep generative models excel at capturing graph data distributions, they struggle to satisfy complex, hard constraints. In unconstrained settings, these models typically produce valid topologies; yet imposing strict compositional rules, like those in drug discovery, creates an out-of-distribution (OOD) setting where purely neural methods frequently fail. Because these neural approaches rely on soft conditioning and post-hoc filtering on such tasks, they cannot provide the formal guarantees needed for high-stakes domains. To address this, we introduce Neuro-Symbolic Graph Generative Modeling (NSGGM), a framework built on the principle of Neural Proposals, Symbolic Guarantees. NSGGM decouples generation: an autoregressive model proposes structural scaffolds, and a Satisfiability Modulo Theories (SMT) solver handles the final discrete assembly of the proposed substructures. Empirically, NSGGM is competitive with state-of-the-art methods on unconstrained tasks. To evaluate logical-constraint satisfaction inspired by real drug discovery workflows, we introduce MolSAT, a benchmark for hard compositional rules. On MolSAT, purely neural baselines completely fail OOD (0% satisfaction with zero training support), while NSGGM achieves >95% satisfaction in-distribution and 64–86% with zero training support.
Authors:
Jieshi He, Puzhe Li, Yanan Sui, Mu-ming PooAbstract: Understanding how cortical activity represents natural whole-body behavior in primates remains challenging. Limited by diversity of movements and inaccessible of large-scale neural feature programming whole-body kinematics, previous motor decoding studies rely on constrained tasks and limited limb movements. Here, we present a neural-behavioral recording and modeling framework for freely moving monkeys, combining epidural cortical signals from distributed sensorimotor areas with synchronized multi-view motion capture through a data custom collection platform. We reconstruct whole-body monkey kinematics and learn a compact motion prior using an autoregressive encoder-decoder model. Conditioned on epidural signals, the model decodes more accurate and realistic full-body movements without explicit physical constraints. Our results provides a novel proof-of-concept approach for decoding natural whole-body movements in primates using large-scale intracranial neural activity.
PaperID: 3587, Poster
Authors:
Kexuan Zhang, Paolo Frazzetto, Xiaobei Zou, Gustau Camps-Valls, Yang TangAbstract: Reliable deployment of time series classifiers in high-stakes domains requires explanations that capture the interplay between localized patterns and global temporal context. However, existing post-hoc methods often treat local features in isolation, failing to account for long-range dependencies. We propose CoLaX, a context-aware explanation framework that explicitly embeds local dynamics within a global temporal structure. By employing a principled context-conditioning mechanism, CoLaX models how global context modulates the importance of local patterns, revealing not only what is important but also when and why it matters. Extensive benchmarks demonstrate that CoLaX yields significantly more faithful and stable explanations than state-of-the-art baselines, demonstrating that joint local-global modeling is essential for reliable time-series interpretability.
Abstract: Drifting models are capable one-step generative models trained to follow a drifting field. The field combines attractive and repulsive softmax-weighted centroids over the data and current-generator distributions. In practice, only a minibatch of n samples from each distribution is available, and each centroid is approximated by an empirical estimate. In this paper, we begin by showing that the minibatch centroid is in general a biased estimator of the target centroid, with a pointwise O(1/n) bias arising from softmax self-normalization. Correcting this bias requires the expectation over the full distribution, which is intractable. We instead approximate the leading bias term from in-batch statistics and propose Analytical Bias Correction (ABC), a closed-form plug-in adjustment. We prove that ABC reduces the bias from O(1/n) to O(1/n^2), introduces no first-order variance inflation, and preserves convex-hull containment of the corrected centroid. In practice, ABC requires only two additional lines of code and has negligible wall-time overhead under compiled execution. Toy experiments confirm the theoretical O(1/n) and O(1/n^2) scaling. On CIFAR-10, ABC reduces FID and trains faster, with the largest gains at small n, where the bias is most significant.
Authors: Nick Hansen, Xiaolong Wang
Abstract: Modern generative world models render increasingly realistic action-controllable futures, yet they frequently hallucinate: rollouts remain visually fluent while drifting from the ground-truth dynamics. We hypothesize that hallucination concentrates in low-coverage regions of the state-action space, where lightweight data-centric signals can both detect it and guide mitigation. To test this, we introduce MMBench2, a 427-hour, 210-task dataset for visual world modeling with ground-truth actions, rewards, and live simulators, and train a 350M-parameter world model on it. We identify three distinct hallucination modes: perceptual, action-marginalized, and scene-diverging - each anchored to a different stage of the pipeline, and develop three signals that accurately predict where the model will fail. To close coverage gaps at training time, we develop a coverage-aware sampling technique; to close them online, our hallucination predictors serve as curiosity rewards for targeted data collection, yielding a data-efficient finetuning recipe that adapts the pretrained world model to entirely unseen environments with as few as 50 real environment trajectories. Overall, our findings reveal that hallucination in world models is inherently a data coverage issue, and that the same signals used to detect it can also be used for mitigation.
PaperID: 3590, Poster
Authors: Erik Y Wang
Abstract: We give a computer-assisted counterexample to the open question posed by Rudin, Schapire, and Daubechies in COLT 2012, of whether exhaustive AdaBoost always converges to a finite cycle. The construction is based on a block-product gadget whose two factors share an exact period-2 orbit for their 5-step branch maps, but whose linearized return maps have dominant eigenvalues with an irrational logarithmic ratio. This irrationality forces the burst-winner sequence to have an irrational asymptotic frequency, precluding eventual periodicity. All assertions are certified by exact rational arithmetic in two independent computer algebra systems. We also document the collaborative workflow that led to the construction, in which one model produced the key technical arguments; another model summarized and critiqued different versions of that work, proposing new directions to pursue; and the authors orchestrated the process by selecting directions to prioritize, resolving ambiguities, heavily refining each argument, and verifying the final solution. We present this as one of the first in-depth case studies of LLM-assisted mathematical research in which a long-standing open problem from theoretical machine learning is resolved, and detail the advantages and challenges of such research.
Abstract: Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workflows whose accumulated reasoning and tool traces routinely push the input an order of magnitude past the pretraining window, making zero-shot context extension the dominant deployment path for open-weight checkpoints. Most existing zero-shot methods fix a single rescaling factor up front, so an aggressive factor sacrifices short-context fidelity while a conservative one breaks down at long contexts. We propose Jet-Long, a tuning-free zero-shot method that pairs a local RoPE-faithful window with a long-range window whose rescaling factor adapts dynamically to the current sequence length, recovering the base model exactly at short inputs while extrapolating cleanly at long ones. An inclusion--exclusion attention merge and an on-the-fly RoPE correction rotation make the bifocal construction essentially free at inference; fused into a single CuTe kernel, long-context prefill reaches up to 1.39× FA2 throughput on H100 (approaching the Hopper-only FA4), and single-batch generation incurs \le 4% overhead at every length. On Qwen3-1.7B/4B/8B up to 128K context, Jet-Long leads RULER by +4.79/+2.18/+2.03 pp over the strongest baseline at 1.7B/4B/8B, achieves the best overall accuracy on HELMET-RAG (a benchmark identified by HELMET as the most efficient predictor of downstream long-context performance) and attains the lowest PG-19 perplexity. Jet-Long also generalizes to hybrid attention architectures such as Jet-Nemotron for further long-context improvement without retraining, and remains hyperparameter-resilient for ease of deployment.
Abstract: The ability of deep neural networks to learn hierarchical features is widely regarded as a key mechanism underlying their success in high-dimensional learning. Existing theory partially supports this view by establishing approximation rates based on parameter counts and sample complexity guarantees for compositional models without incurring the curse of dimensionality (CoD). To study overparameterized regimes, where the number of parameters exceeds the sample size, we develop a framework that measures complexity via the parameter norm. Within this approach, we establish approximation rates and excess risk bounds for learning sparse compositional functions whose compositional structure is represented by directed acyclic graphs (DAGs), using Frobenius norm-constrained deep neural networks. Our results have broad applicability since every function that is efficiently Turing computable admits sparse compositional representations. In particular, we cover a range of representative models, including multi-index models, binary tree structures, and general compositional architectures. The rates we derive show that deep networks can exploit the compositional structure of the target functions, effectively avoiding the CoD through hierarchical representations.
PaperID: 3593, Poster
Abstract: This paper studies the \varepsilon-best arm identification problem (\varepsilon-BAI) for K-armed bandits under the fixed-confidence setting. Departing from the usual objective of identifying the best arm with minimal sample complexity, we focus on minimizing the cumulative regret of identifying an \varepsilon-best arm. We consider the usual asymptotic regime in which the error probability \delta of identifying an \varepsilon-best arm vanishes. In this setting, we show a novel instance dependent lower bound. In particular, we show that any family of (\varepsilon,\delta)-PAC algorithms must incur an expected cumulative regret up to the stopping time that scales as \log(1/\delta), with an instance-dependent constant C(\mu). This is proved using a novel regret-aware change-of-measure technique that allocates information given a budget on the regret. To achieve this lower bound, we devise an algorithm, \algname, that achieves the same scaling in \delta, i.e., \log (1/\delta), but with a different constant, which we call C^\(\mu). KL-UCB+Greedy has two main benefits: (i) it is computationally efficient as C^\(\mu) admits a closed-form expression; (ii) in broad classes of reward distributions such as Gaussian distributions, C^\( \mu) = C( \mu), demonstrating asymptotic optimality of KL-UCB+Greedy within these classes. As auxiliary results, we also derive lower bounds on the sample complexity of any family of (\varepsilon,\delta)-algorithms as \delta vanishes and show that the sample complexity of KL-UCB+Greedy closely matches that of the lower bound on the sample complexity.
PaperID: 3594, Poster
Authors: Jinwoo Chung, Jangho Kim
Abstract: Mamba-based super-resolution models are efficient alternatives to Transformer-based restoration backbones, but their low-bit post-training quantization remains challenging because selective scan, Mamba's input-dependent state-space operator, induces different activation ranges across spatial tokens. Thus, a single layer-wise clipping range poorly matches all tokens. We propose Token-Group Quantization, a hardware-friendly activation PTQ framework that statically approximates a hardware-unfriendly post-scan grouping oracle, where tokens with similar preferred clipping ranges are assigned to the same group. To avoid runtime post-scan grouping in static low-bit inference, a lightweight predictor infers the groups from pre-scan features available before the quantized scan path. During calibration, we form pseudo-labels by clustering post-scan activation descriptors that capture similar rounding--clipping behavior. At inference, the frozen predictor maps pre-scan features to groups and selects entries from offline-calibrated group-wise clipping-bound tables, without runtime statistics collection, dynamic clipping updates, or post-scan group assignment. To address low-bit restoration degradation, we further introduce a Frequency-Preserving Refinement (FPR) objective for clipping-bound calibration. FPR aligns Fourier amplitude and phase between full-precision teacher and quantized student outputs to preserve high-frequency structures. Across PTQ protocols, bit-widths, backbones, and restoration tasks, our method consistently improves restoration quality over prior Mamba and SR PTQ baselines, while achieving 2.23× speedup on Jetson Orin Nano over FP16.
PaperID: 3595, Poster
Abstract: (LBS), a hierarchical semantic-to-path framework for safe flow-matching-based generative planning. Existing control-barrier-function (CBF) safety filters for diffusion and flow-matching planners act locally on generated samples, either during sampling or after prediction. Such repair is effective for small violations, but can struggle when a new constraint blocks the whole selected behavior mode rather than merely perturbing it locally. LBS decouples behavior selection from path certification through two coordinated interventions: it first steers the generator's behavior latent toward rollouts with larger safety margin, then decodes the steered latent and applies a path-space CBF quadratic program (QP) only for residual violations. We show that latent steering increases the margin available to the path-space corrector, yielding finite-time recovery guarantees and reduced correction burden. Experiments across navigation and robot end-effector planning tasks show that LBS preserves safety while improving goal reaching and reducing correction burden compared with path-level safety filters.
Abstract: We study the query and communication complexity of sampling from high-dimensional Gaussian distributions using gradient information. In the standard oracle model, exact gradients expose only matrix-vector products with the precision matrix, leading to polynomial approximation barriers and a characteristic \sqrt\kappa dependence on the condition number. We show that this barrier disappears when the sampler is allowed to query \emphsmoothed gradients, namely scores of Gaussian-convolved distributions. For a Gaussian target with precision \Lambda, a smoothed-score query at noise level \tau gives access to the resolvent (\Lambda+\tau^-1I)^-1. Combining geometrically spaced noise levels with sinc-quadrature rational approximation, we obtain a sampler with q = O\left( \left(\log\kappa+\log(\sqrt d/\varepsilon)\right) \log(\sqrt d/\varepsilon) \right) smoothed-gradient queries for total variation error \varepsilon, improving the condition-number dependence from \sqrt\kappa to logarithmic. We also study finite-bit gradient oracles. Using coordinatewise quantization of the transformed smoothed-gradient answers and a final dithering step, we obtain a sampling scheme whose total communicated gradient information is polylogarithmic in \kappa; in particular, for fixed dimension and accuracy, the bit complexity is O(\log^2\kappa). To complement these upper bounds, we introduce a channel-synthesis, or reverse-Shannon, converse technique for sampling lower bounds. This converts total-variation simulation guarantees into communication requirements and yields an \Omega(\log\kappa) lower bound on the required gradient information. Together, these results identify smoothed gradients as a strictly more powerful oracle for sampling and give nearly matching upper and lower bounds for its finite-bit complexity.
Abstract: Chain-of-thought (CoT) reasoning with self-consistency improves performance by aggregating multiple sampled reasoning paths. In this setting, correctness is no longer tied to a single reasoning trace but to the aggregation rule over a pool of candidate paths, making aggregation uncertainty the central challenge. This issue is critical where confidently incorrect answers are far more costly than abstentions. We introduce a conformal procedure for CoT reasoning that directly addresses aggregation uncertainty. Our approach replaces majority voting with weighted score aggregation over reasoning paths and calibrates an abstention rule using conformal risk control. This approach leads to finite-sample guarantees on the confident-error rate--the probability that the system answers and is wrong. We further identify score separability as the key condition under which abstention provably improves selective accuracy, and derive closed-form expressions that predict accuracy gains from calibration data alone. The method is fully inference-time, and requires no retraining. Across four benchmarks, four open-source models, and three score classes, realized confident-error rates are consistent with the prescribed targets up to calibration-split and test-set variability. Our method achieves 90.1% selective accuracy on GSM8K by abstaining on less than 5% of problems, compared with 82% accuracy under majority-voting baseline.
PaperID: 3598, Poster
Authors:
Ziming Zhu, Yu Zhu, Qin Guo, Yujia Zhang, Zheng ShenAbstract: End-to-end autonomous driving planning requires trajectories that are multimodal, scene-compliant, and efficient to generate. Deterministic regression is fast but often collapses to a single future, while trajectory-vocabulary methods depend on predefined candidate coverage and iterative generative planners incur multi-step inference cost. We propose TwinFlux, a one-step discrete-continuous trajectory generation framework based on SplitMeanFlow. TwinFlux starts from data-driven trajectory prototypes and refines them through parallel heatmap and offset branches. The heatmap branch provides topology-aware coarse localization on a BEV grid, while the offset branch predicts grid-relative geometric refinement. A differentiable trajectory decoupling-and-reconstruction module connects this representation with physical trajectories. During training, TwinFlux enforces piecewise displacement consistency in reconstructed trajectory space, aligning direct one-step prediction with accumulated sub-interval evolution without Jacobian-vector-product computation or inference overhead. Experiments on NAVSIMv1 and NAVSIMv2 show that TwinFlux outperforms recent end-to-end planning baselines while maintaining real-time inference.
PaperID: 3599, Poster
Authors: Yutian Cheng, Canzhe Zhao, Jingye Zhao, Shuai Li
Abstract: We study best-of-both-worlds learning in repeated bilateral trade with two-bit feedback. When seller and buyer valuations are independent, minimax regret scales as \Theta(T^2/3); under general bounded-density distributions it degrades to \Theta(T^3/4). This regime-dependent gap naturally raises a question: can a single algorithm adapt optimally to both? We first resolve this negatively: any algorithm achieving O(T^\alpha) regret on independent-values instances necessarily suffers \Omega(T^1-\alpha/3) regret on some dependent-values instance. We then characterize the optimal tradeoff and construct a meta-algorithm achieving \bigl(\widetilde O(T^\alpha),\widetilde O(T^1-\alpha/3)\bigr) for every \alpha\in[2/3,3/4]. Finally, we show that this tradeoff can be bypassed with mild side information: knowing the optimal same-price gain suffices to recover the classical best-of-both-worlds guarantee \bigl(\widetilde O(T^2/3),\widetilde O(T^3/4)\bigr).
PaperID: 3600, Poster
Abstract: Selective state space models (SSMs) such as Mamba-3 are efficient and increasingly competitive with Transformers. However, their fixed recurrent state remains a bottleneck for using context as temporary memory. To address this issue, we propose GEAR (GEnerator-Adaptive state space models for associative Recall), a slow-hypernetwork that produces coefficients for LoRA updates to the SSM token-to-recurrence parameter generators. GEAR makes the parameter generator context-adaptive while preserving the original linear-time Mamba-3 scan. To evaluate whether this adaptation improves in-context recall, we pretrain 180M-parameter models for 2B FineWeb-Edu tokens and compare against Mamba-3 baseline and ablations at the same scale. In the Multi-Query Associative Recall (MQAR) task, models see N in-context key-value bindings and must predict the value paired with a queried key. In compact MQAR evaluation where the adaptive path is active at the query, GEAR improves candidate-restricted NLL over Mamba-3 by 0.70, 0.69, and 0.48 at N=32,64,128 , respectively. Beyond compact MQAR, GEAR reduces the penalty on older queried bindings and gives the clearest gain within the Mamba-3 family under moderate random-filler interference. These results show that context-adaptive SSM generator improves associative recall while preserving Mamba-3 language modeling capabilities.
PaperID: 3601, Poster
Abstract: Quantifying uncertainty in time series forecasting is particularly demanding because sequential data exhibit temporal dependence and are prone to distributional changes. Conformal inference has emerged as a powerful uncertainty quantification approach through the construction of prediction sets. Recent advances have introduced online conformal methods that adaptively adjust prediction thresholds through feedback mechanisms. However, the existing feedback mechanism typically relies solely on miscoverage indicators (actual feedback)—whether the true label falls within the interval at each time step—while overlooking the empirical prediction threshold (estimated feedback) that is derived from the oracle conformal method. In this paper, we propose Dynamic Dual-feedback Conformal Inference (DDCI), which incorporates a dual-feedback mechanism consisting of actual feedback and estimated feedback. The former drives the primary adjustment of the intervals based on true observations, while the latter dampens excessive expansions or contractions by leveraging empirical thresholds from conformal inference during updates. By balancing these two signals, DDCI achieves more stable and narrower prediction intervals in sequential settings while adhering to the target coverage rate in practice.
PaperID: 3602, Poster
Abstract: Cross-domain time-series generation is a challenging task for diffusion models due to the strong requirement for domain adaptability. While frequency-based Mixture-of-Experts (MoE) methods have shown promise in domain adaptation, their direct integration with diffusion models faces critical issues from intra- and inter-domain. For intra-domain aspect, the absence of frequency transition compliance leads to imbalance pattern learning across denoising steps and degrade model performance. For the inter-domain aspect, the domain-specific transition variation bring difficulties for domain adaptation. To address these challenges, we propose IIDiff, a frequency-based diffusion MoE framework equipped with progressive frequency specializing and dynamic transition matching. The progressive specializing allows experts dynamically adjust frequency receptive fields to capture underlying patterns aligning with the intra-domain transition, enhancing the generation fidelity. The transition matching leverages the frequency noise distribution property to analysis the instantaneous transition process, enabling adaptive routing with domain-specific transition. Extensive experiments on 12 real-world datasets across 4 domains demonstrate our model's superiority. Compared with SOTA baselines, our model achieves an average 27.59% improvement in KL divergence and 6.00% improvement in MMD, confirming its robust domain adaptability and ability to generate high-fidelity cross-domain time-series.
Abstract: Open-source LLMs (OSMs) are reaching near state-of-the-art performance, prompting prior works to trace the text they generate by embedding text watermarking algorithms directly into their weights. Yet, OSMs are subject to post-training modifications, which has been shown to remove the watermark. Model merging in particular, a prominent method used for combining expert knowledge and preventing catastrophic forgetting, strongly removes such OSMs watermarks. A key question is how to enable OSM watermarks that survive subsequent merging. In this work, we show for the first time how to design an OSM watermark that is durable against model merging. We propose Merge-Adversarial Training, an adversarial training algorithm to distill text watermarks into model weights while being robust to subsequent model merging. Our approach consistently outperforms all baselines (e.g. with SLERP up to +51 percentage points (pp) TPR@1%FPR with +25 pp on average) while preserving downstream capabilities. We also for the first time evaluate OSM watermarks against realistic merge scenarios, representing common use-cases such as combining expert capabilities or preventing catastrophic forgetting, and with 3 prominent merging algorithms. More broadly, our findings suggest that adversarial training is a reliable approach for increasing OSM watermark durability against post-training modifications.
PaperID: 3604, Poster
Abstract: Biologically plausible learning requires temporal credit assignment without backpropagation through time (BPTT). Hamiltonian Echo Learning (HEL) achieves this by running a neural system twice --- an \emphinference phase then a time-reversed \emphecho phase --- recovering exact BPTT gradients. We identify spatio-temporal sum-separability of the Hamiltonian as the structural condition making HEL's learning rule both spatially and temporally local, and exploit it to extend HEL beyond diagonal recurrences: a Hopfield-inspired oscillatory RNN with dense recurrent connectivity yields a contrastive Hebbian rule with constant memory and only two forward passes. This rule matches full BPTT and outperforms e-prop and truncated BPTT on time-series classification and regression. We then relax HEL's Hamiltonian-reversibility constraint --- which forces unbiological connectivity and activation patterns --- and derive \emphEcho Learning (EL), exact for arbitrary smooth dynamics whenever the network is reversible and self-adjoint with respect to a readout involution. Both conditions are enforced by a spatio-temporally local homeostatic loss, recovering near-BPTT gradient quality without architectural constraints.
Authors:
Pengyue Jia, Derong Xu, Yingyi Zhang, Xiaopeng Li, Wenlin Zhang, Yi Wen, Liu lan, Yuanshao Zhu, Xuetao Wei, Wanyu Wang, Xiangyu ZhaoAbstract: Worldwide image geolocalization aims to predict precise GPS coordinates for images captured anywhere on Earth, which is challenging due to the large visual and geographic diversity. Recent methods mainly follow two paradigms: retrieval-based approaches that match queries against a reference database, and generation-based approaches that directly predict coordinates using Large Vision-Language Models (LVLMs). However, we observe distinct error profiles between them: retrieval excels at fine-grained instance matching, while generation offers robust semantic reasoning. This complementary heterogeneity suggests that no single paradigm is universally superior. To harness this potential, we propose GeoRouter, a dynamic routing framework that adaptively assigns each query to the optimal paradigm. GeoRouter leverages an LVLM backbone to analyze visual content and provide routing decisions. To optimize GeoRouter, we introduce a distance-aware preference objective that converts the distance gap between paradigms into a continuous supervision signal, explicitly reflecting relative performance differences. Furthermore, we construct GeoRouting, the first large-scale dataset tailored for training routing policies with independent paradigm predictions. Extensive experiments on IM2GPS3K and YFCC4K demonstrate that GeoRouter significantly outperforms state-of-the-art baselines. Our dataset, code, and checkpoints are publicly available to facilitate future research.
Authors:
Yuxuan Lu, Ziyi Wang, Yingzhou Lu, Yisi Sang, Jiri Gesi, Xianfeng Tang, Yimeng Zhang, Zhenwei Dai, Hui Liu, Hanqing Lu, Chen Luo, Qi He, Benoit Dumoulin, Jing Huang, Dakuo WangAbstract: Training robust tool-calling agents requires large-scale trajectory data that is both realistic and verifiable, yet existing datasets still rely heavily on costly human annotation or synthetic pipelines with weak correctness guarantees. In this work, we introduce \projectname, a large-scale synthetic and verifiable dataset for tool-calling agents built on real-world MCP servers. Unlike prior pipelines that first generate tasks and then attempt to solve them, \projectname inverts this process: a strong LLM first self-explores real MCP servers and records reachable execution states, after which natural-language tasks are synthesized backward from observed outcomes. This back-chaining design guarantees task reachability and label correctness by construction. To support scalable offline training and evaluation, we further construct a retrieval-augmented simulator from explored trajectories, enabling reproducible tool execution without relying on live MCP servers. Each instance includes a task description, tool schemas, ground-truth trajectories, ground-truth answers, and tool-call responses for offline replay, and is automatically checked for determinism and semantic consistency. \projectname aggregates thousands of verifiable tasks across diverse real-world MCP servers and supports scalable agent learning without human annotation. Experiments show that models trained on \projectname achieve significant improvements on tau2-Bench and BFCL, demonstrating the effectiveness of realistic, verifiable synthetic data for training generalizable tool-calling agents.
PaperID: 3607, Poster
Abstract: We propose the task of Interactive 4D Volumetric Liquid Forecasting: predicting the spatiotemporal evolution of a future dense 3D liquid velocity field conditioned on how a solid object moves through it. This task provides an upstream physical forecasting primitive for systems that need to reason about motion-conditioned liquid response, including robotic liquid manipulation, digital twins, and rigid-fluid video generation. Existing works have made efforts in predicting flow around static or prescribed solid geometries and synthesizing visually plausible fluid motion. Yet these approaches provide limited coverage for interactive liquid dynamics, as they are not capable of forecasting how a moving solid induces a future full-domain liquid response. In this forecasting scenario, the hydrodynamic effect is condensed near an evolving liquid-solid interface, but the prediction target is a global spatiotemporal rollout. To address this challenge, we propose a two-level recurrent forecasting framework, TIDE/TIDES: TIDE isolates localized interface-driven forecasting, while TIDES adds transport-aware coupling and explicit stabilization to reduce accumulated rollout error. Across adapted neighboring baselines, controlled component ablations, motion/size OOD tests, and long-horizon evaluation, TIDES achieves the best forecasting performance and the ablations identify localized interaction, transport coupling, and stabilization as the key contributors. Together, our task formulation, benchmark, and TIDE/TIDES study establish a controlled foundation for moving-solid-conditioned dense liquid forecasting. Code and benchmark will be made publicly available.
PaperID: 3608, Poster
Authors: Hyungkwon Lee, Hyewon Ryu, Jong-Seok Lee
Abstract: Domain generalization aims to learn models that generalize to unseen target domains using labeled data from multiple source domains. Many existing approaches focus on learning domain-invariant representations, but enforcing invariance alone may fail to preserve sufficient task-relevant information under domain shift. We propose BLENDER, a reconstruction-aware framework for domain generalization that employs a hierarchical latent decomposition to separate domain-invariant and domain-specific factors. This structure prioritizes informativeness in the domain-invariant representation and captures residual domain-specific variation through conditional modeling. We further provide a theoretical analysis establishing a bound on the reconstruction risk for unseen target domains, revealing how domain invariance and latent disentanglement contribute to generalization beyond the observed domains. Experiments on standard domain generalization datasets demonstrate strong performance under substantial domain shifts, and qualitative analyses show that BLENDER promotes a structured allocation of information between domain-invariant and domain-specific latent representations.
PaperID: 3609, Poster
Abstract: To consider synthesizability in structure-based drug design (SBDD), a recent approach has proposed to co-design synthetic pathways with molecular conformers of ligands. However, the intermediate conformers are predicted based solely on the current intermediate molecule, without information about the synthons that will be added later along the synthetic pathway. As a result, the intermediate conformers do not adequately reflect the upcoming substructures, creating the geometric mismatch that leads to cumulative errors along the synthetic trajectory. To this end, we propose SYnthon Contrastive LEarning (SYCLE), a synthesizable SBDD framework that injects future synthon information into intermediate conformer prediction via contrastive learning. SYCLE further utilizes this information in the pathway generation policy to guide synthon selection. SYCLE achieves state-of-the-art binding affinity across all 15 LIT-PCBA targets. On CrossDocked2020, SYCLE attains -9.80 kcal/mol and an AiZynthFinder success rate of 68.8%, nearly double that of the best baseline.
Abstract: Causal discovery aims to recover graphs that represent causal relations among given variables from observations, and new methods are constantly being proposed. Increasingly, the community raises questions about how much progress is made, because properly evaluating discovered graphs remains notoriously difficult, particularly under latent confounding. We propose a graph distance measure for acyclic directed mixed graphs (ADMGs) based on the downstream task of cause-effect estimation under unobserved confounding. Our approach uses identification via fixing and a symbolic verifier to quantify how graph differences distort cause-effect estimands for different treatment-outcome pairs. We analyze the behavior of the measure under different graph perturbations and demonstrate how it complements existing distance metrics.
PaperID: 3611, Poster
Authors: Joonyoung Lim, Younghwan Yoo
Abstract: We investigate multi-constraint policy optimization in constrained Markov decision processes (CMDPs), where interactions among constraint gradients often yield ill-conditioned update directions, causing numerically fragile steps and unintended cancellation or redundant overlap in constraint corrections. To address this, we propose LOCU, which decouples constraint interactions by applying symmetric L\"owdin orthogonalization to the natural gradients of the reward and constraints under the Fisher–Rao metric. Combined with Fisher-geometric screening and near-collinear compression, LOCU yields a compact reduced space in which the trust region becomes a Euclidean ball and constraints are handled symmetrically through a low-dimensional active-set solve. For infeasible iterates, LOCU construct a Pareto-descent direction that simultaneously reduces violated constraints while preserving satisfied ones within the same reduced formulation. Experiments show that LOCU remains stable under high constraint coupling, narrow feasible regions, and infeasible starts, and generalizes across different cost critic architectures.
PaperID: 3612, Poster
Abstract: Thinking with images empowers vision-language models (VLMs) to use an external zoom-in tool for fine-grained visual search. Recent works employ supervised fine-tuning (SFT) and reinforcement learning (RL) with tool-specific designs to enable such capabilities. However, achieving logical reasoning alongside accurate and efficient tool use remains a significant challenge. To address this, we shift the focus away from increasing RL complexity and instead identify SFT data quality as a key bottleneck. We then propose LogT (short for Logically Think with Images), a fully automated SFT data engine that explicitly enforces logic-chain-guided reasoning. LogT generates an SFT dataset featuring challenging, visually grounded questions and structured tool-use traces with human-like logic. We validate our approach by fine-tuning two base models, Qwen2.5-VL-7B and InternVL3.5-8B, on the dataset and evaluating them across seven benchmark test sets. We show that LogT-SFT based on Qwen2.5-VL-7B outperforms the state-of-the-art trained with both SFT and RL, despite using 3.1× less data and >120× fewer GPU hours. Furthermore, applying a simple RL recipe without tool-specific design on LogT-SFT yields a model that surpasses prior methods with complex RL setups by a substantial 9.59-point margin. In addition to strong performance gains, our model exhibits higher tool-use accuracy and efficiency, as well as more structured and logically consistent reasoning traces, directly validating the effectiveness of our approach.
PaperID: 3613, Poster
Abstract: This paper studies binary classification under generalized performance metrics, going beyond standard misclassification error. We focus on linear-fractional measures—including the F-score and Jaccard index—which are widely used in class-imbalanced settings. We first provide a unified characterization of the optimal classifier, showing that it admits a thresholding structure on the regression function. This reduces the original infinite-dimensional, non-decomposable optimization problem to the estimation of a single scalar parameter. Building on this result, we analyze plug-in classifiers that combine a regression estimator with a data-driven estimate of the optimal threshold. We show that this threshold is characterized as the solution of a fixed-point equation depending only on the marginal distribution, enabling its estimation using unlabeled data and avoiding standard validation-based procedures. Our main contribution is a finite-sample analysis of this approach. We derive excess risk bounds that decompose the error into contributions from regression estimation, threshold estimation, and sampling effects, providing a modular understanding of learning under non-decomposable metrics. Under standard margin assumptions, we further establish fast convergence rates. Our results yield a unified theoretical framework for plug-in classification under linear-fractional metrics, extending prior analyses beyond specific measures such as the F-score. Empirical results on synthetic and real datasets support our theoretical findings.
Abstract: Vision-language-action (VLA) models provide a promising paradigm for scalable robotic manipulation, yet their reliance on success-only behavioral cloning leaves them brittle; lacking corrective training signals, minor execution errors rapidly compound into unrecoverable, out-of-distribution failures. To address this limitation, we propose Adaptive Failure-Informed Learning (AFIL), an end-to-end framework that leverages failure trajectories as adaptive negative guidance for diffusion- and flow-based VLA policies. AFIL uses a pretrained VLA to generate failure rollouts online, avoiding the need for handcrafted failure-mode design or human-in-the-loop recovery. It then jointly trains Dual Action Generators (DAGs) for successful and failed behaviors while sharing a common vision-language backbone, enabling efficient failure-aware policy learning with limited parameter overhead. During sampling, the failure generator adaptively steers action generation away from failure-prone regions and toward more reliable success modes, with guidance strength determined by the per-diffusion-step distance between success and failure distributions. Experiments across in-domain and out-of-domain robotic manipulation tasks, covering both short- and long-horizon settings, show that AFIL consistently improves task success rates and robustness over existing VLA baselines, demonstrating its effectiveness, efficiency, and generality.
PaperID: 3615, Poster
Authors: Varun Jain, Hong Ge
Abstract: Efficient sampling from complex probability distributions is an important task across Bayesian inference and certain classes of deep generative models. Markov chain Monte Carlo (MCMC) remains one of the most widely used tools for this. However, even state-of-the-art samplers, such as the NUTS implementations in Stan and Turing.jl, can converge slowly on ill-conditioned, heavy-tailed, or near-degenerate targets. We identify that sampling error can be decomposed into two terms: mismatch in the potential energy (PE) marginal, and average conditional mismatch across PE level sets. We propose Contour Monte Carlo (CMC), which corrects each of these in turn: first by drawing energies u_i from the PE marginal, and then by sampling states on the corresponding level sets. For a broad family of radial targets, including the Gaussian and Student's t distributions, both stages admit closed-form draws, so CMC is exact and requires no Markov chain. Beyond this, for general targets, we estimate the PE marginal using umbrella sampling with MBAR, then run constrained MCMC on the sampled level sets, initialised from energy-matched umbrella samples (which tend to lie in regions of high conditional mass). We show that this procedure essentially reduces the global sampling problem to approximating the one-dimensional PE marginal, after which posterior draws can be generated readily on demand. Furthermore, the level-set chains are embarrassingly parallel, and therefore well-suited to modern accelerator hardware. On a suite of challenging benchmarks -- composed of `posteriordb` models, synthetic targets, and real-world inference problems -- CMC converges substantially faster than HMC and NUTS baselines at a fixed computational budget.
PaperID: 3616, Poster
Authors: Janet zhong, Renwen Yu, Charles Roques-Carmes, Paul-Alexis MOR, Aviv Karnieli, David A Miller, Shanhui Fan
Abstract: Programmable photonic circuits are a natural platform for physical learning, where the learned computation is performed by the hardware itself rather than by a digital model. We propose a physical learning algorithm and architecture for continuous-variable quantum systems where a photonic circuit is configured in situ to invert an unknown lossless Gaussian quantum operation. Our contributions are threefold. (i) Theory: we prove that generic S \in \mathrmSp(2 N, \mathbbR) admits a triangular decomposition into two-mode and single-mode symplectic gates. (ii) Architecture: we introduce the symplectic Reck mesh, a triangular array of two-mode symplectic gates that physically realizes this factorization, generalizing the unitary Reck mesh widely used in optical neural networks. (iii) Physical learning: when placed after an unknown Gaussian unitary, the mesh is trained in situ by nulling local response blocks one at a time using simple optical probes and measurements. We identify a measure that predicts when training becomes hard and numerically demonstrate finite-shot training and robustness to gate errors. Our work extends self-configuring optical networks from unitary to symplectic transformations, bringing physical learning to continuous-variable quantum photonic hardware.
PaperID: 3617, Poster
Abstract: Backdoor defenses are widely regarded as key to secure third-party model deployment. However, we are the to show that, in standard backdoor sanitization pipelines, the defense process itself can be exploited and turned into a conditional trigger. We propose , so that model sanitization suppresses the decoy backdoor while activating the hidden backdoor, thereby making the defense process the trigger condition for the hidden backdoor. DTB uses bi-level optimization to approximate the shared sanitization effect of mainstream backdoor defenses, while constraining the hidden backdoor to activate only after the decoy backdoor is sufficiently suppressed. Experiments on multiple datasets, different networks, and 14 representative defenses show that DTB can the hidden backdoor, and this phenomenon remains significant after multi-round compositional sanitization. Our findings reveal a potential threat in existing backdoor defense pipelines and suggest that post-sanitization safety cannot be judged solely by whether the original backdoor has been eliminated.
Abstract: Root cause analysis of anomalies aims to identify how and why a sample deviates from the normal process. Existing methods primarily focus on telling which features are responsible, ignoring that anomalies can arise through two fundamentally different processes: measurement errors, where the sample is generated normally but one or more values is recorded incorrectly, and mechanism shifts, where the causal process that generated the sample was changed. While measurement errors can often be safely corrected, mechanistic anomalies require careful consideration. In this paper, we formally define a causal model that explicitly captures both types by treating outliers as latent interventions on latent (“true”) and observed (“measured”) variables and show under which conditions the distinction is possible. Based on this model, we develop an efficient inference procedure for localizing root causes and distinguishing anomaly types. Experiments on synthetic and real-world data show that our method provides state-of-the-art and highly robust performance in both root cause localization and classification of anomaly types.
Authors:
Yufeng Jin, Jianfei Guo, Xiaogang Jia, Yu Deng, Steven Li, Han Liu, Weiran Liao, Vignesh Prasad, Mathias Franzius, Gerhard Neumann, Georgia ChalvatzakiAbstract: Robot learning research is fragmented across policy families, benchmark suites, and real robots; Each implementation is entangled with the others in a complex combination matrix, making it an engineering nightmare to port any single element into. General-purpose coding agents may occasionally bridge specific setups, but cannot close this gap at scale because they lack the procedural priors and validation practices that characterize robotics research workflows. We propose NAUTILUS, an open-source harness that turns a single user prompt, for example, ``Evaluate policy A with benchmark B''---into ready-to-use reproduction, evaluation, fine-tuning, and deployment workflows. NAUTILUS provides: plug-and-play agent skill sets with distilled priors from robotics research; typed contracts among policies, simulators/benchmarks, and real-world robots; unified interfaces and execution environments; and an agentic coding workflow with explicit, automated validation, and testing at each milestone. NAUTILUS can not only automatically generate the required adapters and containers for existing implementations, but also wrap and onboard new or user-provided policies, simulators/benchmarks, and robots, all connected via a uniform interface. This expands cross-validation coverage without hand-written glue code. Like a nautilus shell that grows by adding chambers, NAUTILUS scales by extending its execution in chambered units, making it a research harness for scalability rather than a hand-curated framework, and aiming to reduce the engineering burden of cross-family reproduction and evaluation in the ever-growing robot learning ecosystem.
PaperID: 3620, Poster
Abstract: In quantum systems, physical quantities of interest can often be expressed as expectation values of a probability distribution, which allows their estimation by Markov Chain Monte Carlo (MCMC) methods. This approach, an instance of Quantum Monte Carlo (QMC), plays a central role in obtaining predictions of system properties relevant to high energy or condensed matter physics. The use of Hamiltonian dynamics to design a Markov kernel, known as Hybrid or Hamiltonian Monte Carlo (HMC), is widely used in quantum applications. In this context, it is standard to use an HMC kernel that is ``adjusted'' with the Metropolis-Hastings (MH) criterion, so that estimates of expectations converge to their true value in the limit of infinite sample size. However, recent work in computational statistics suggests that an \emphunadjusted HMC kernel can be significantly more efficient, while maintaining an asymptotic bias which is small relative to the error arising from the variance of the finite sample size. We adapt this approach to the QMC setting, focusing on a challenging and high-dimensional \emphspin-fermion model of a quantum phase transition. We find that unadjusted methods outperform adjusted HMC in terms of effective sample size per gradient call by around an order of magnitude, while retaining accurate results.
Abstract: Score-based diffusion models have achieved remarkable empirical success in generating high-quality samples from target data distributions. Among them, the Denoising Diffusion Probabilistic Model (DDPM) is one of the most widely used samplers, generating samples via estimated score functions. Despite its empirical success, a tight theoretical understanding of DDPM --- especially its convergence properties --- remains limited. In this paper, we provide a refined convergence analysis for the DDPM sampler and establish near-optimal convergence rates under general distributional assumptions. Specifically, we consider a relaxed smoothness condition parameterized by L, which is small for many practical distributions (e.g., Gaussian mixture models). We prove that, to approximate a target distribution on \mathbbR^d to accuracy \varepsilon in total variation distance and \varepsilon^2 in KL-divergence, the DDPM sampler with accurate score estimates requires at most T = \widetildeO\left(\frac\sqrtd\min\\sqrtd,L\\varepsilon\right) iterations, where \widetildeO hides polylogarithmic factors in d and 1/\varepsilon. This result substantially improves upon the best-known \widetildeO(d/\varepsilon) iteration complexity when L < \sqrtd. By establishing a matching lower bound, we show that our convergence analysis is tight for a wide array of target distributions. Moreover, it reveals that DDPM and Denoising Diffusion Implicit Models (DDIM) share the same dependence on d, raising an interesting question of why DDIM often appears empirically faster.
PaperID: 3622, Poster
Authors:
JINYOUNGKIM, Geonho Kim, GiJeong Park, Geonu Lee, YoungJoon YooAbstract: CLIP is a powerful vision-language model, but it was not designed for fine-grained defect localization; CLIP-based anomaly detectors therefore adapt it with prompts or lightweight modules to increase defect sensitivity. We show that stronger sensitivity does not necessarily make local evidence reliable: under domain shift, adapted CLIP-AD models often assign high anomaly scores to both true defects and visually complex normal regions. The issue is not simply missing defect information, but a local scoring rule that decodes defect and hard-normal evidence, having the same anomaly evidence. We propose \textscTED (Text-Axis Evidence Decomposition), a post-hoc scoring method that asks whether each ambiguous response is better supported by source defect patches or by source normal patches mistaken as anomalous. \textscTED compares these supports under the host's normal-versus-anomaly text response, leaves the backbone and prompts unchanged, and requires no target-domain training. It works as a train-free score for raw VLM backbones or as a source-calibrated residual correction for adapted CLIP-AD hosts. Across frozen VLM backbones, \textscTED substantially improves pixel-level localization over raw prompt similarity; across adapted hosts, it improves most pixel-level settings over P-AUROC, P-PRO, and P-AP. Gains are largest under stronger hard-FP competition, with mean localization gain increasing from +5.0 in low-competition regimes to about +10.9 in mid/high-competition regimes. These results suggest that recoverable defect evidence can already exist in pretrained multimodal representations, but reliable localization requires decoding it against hard-normal competitors.
PaperID: 3623, Poster
Abstract: Recent diffusion models have advanced image-conditioned video generation, but their accessibility raises risks of unauthorized manipulation. Image immunization has emerged as a promising defense for Image-to-Video (I2V) protection, yet existing methods often rely on degrading visual quality or temporal coherence. In this work, we revisit I2V immunization as reducing source-content preservation, shifting the focus from output degradation to weakening source-conditioned content propagation. To this end, we propose I2V-DETACH, an image immunization method that prevents unauthorized I2V generation through source grounding detachment. I2V models preserve source content by propagating source-image cues into video latents during denoising. Accordingly, I2V-DETACH disrupts this process in two complementary ways: it suppresses source-video grounding to reduce source cue propagation, and strengthens intra-video interactions to shift denoising toward video-side dynamics with reduced reliance on the source image. To obtain informative optimization signals, we initialize auxiliary video latents with source-oriented states that remain within the source-conditioned generation regime while deviating from source-preserving trajectories. Extensive experiments across diverse diffusion architectures show that I2V-DETACH substantially reduces source-content preservation, improves over prior immunization methods, and exhibits strong black-box transferability.
PaperID: 3624, Poster
Abstract: Vision-language-action(VLA) models have emerged as a promising interface for autonomous driving. However, existing VLA models still suffer from a fundamental mismatch: driving actions require precise 3D geometric cues, while visual-language understanding and reasoning are largely conducted in a 2D semantic space. In this paper, we propose GeoCoTDrive, an explicit geometric chain-of-thought framework that grounds geometry in a planning-oriented manner. GeoCoTDrive follows a ``think with 2D first, drive with dedicated 3D priors'' paradigm: it first identifies sparse 2D regions corresponding to decision-critical cues, and then retrieves localized 3D priors by sampling features from a geometric foundation model within the grounded regions. These localized geometric features are interleaved into the autoregressive context to support trajectory generation. To supervise this process, we introduce planning-relevant grounding, a new region-level grounding task that focuses on local spatial cues directly affecting ego planning, and construct the PlanningGrounding dataset to endow VLM models with planning-oriented grounding ability. Experiments across multiple end-to-end autonomous driving benchmarks show that GeoCoTDrive consistently improves safety-critical planning performance, demonstrating the effectiveness of explicit geometric chain-of-thought reasoning for VLA-based planning.
PaperID: 3625, Poster
Abstract: In standard multi-head attention, each head computes its output independently, with cross-head information exchange limited to the final output projection. Talking-heads attention introduces explicit inter-head mixing on attention maps, but is incompatible with FlashAttention and incurs significant overhead. We propose knocking-heads attention (KHA), which couples heads at the parameter level via a shared, diagonally-initialized projection applied to query/key/value features before the scaled dot-product attention. The shared projection does not mix signals along the head axis; instead, it is jointly optimized by gradients from all heads, providing an implicit regularization on the head-specific projections. The diagonal initialization recovers MHA at step 0 and lets cross-head structure emerge gradually. We instantiate KHA in two complementary forms: a linear variant (KHA-Linear) whose shared projection can be absorbed into the original weights at inference, adding zero parameters or FLOPs at deployment; and a non-linear variant (KHA-MLP) that obtains stronger gains by introducing a small shared MLP, at modest inference cost. Training a 6.1B parameter MoE model (1.01B activated) on 1T tokens, we observe substantially fewer loss spikes, more uniform per-head activation norms, and a +1.26 average improvement across 19 downstream benchmarks. KHA is compatible with FlashAttention, KV-caching, and existing attention variants (MHA, MQA, GQA, GTA, MLA).
Abstract: Deep Q-Networks (DQNs) approximate value iteration by freezing a target network, forming Bellman-optimality targets, and running an inner regression loop that minimizes a smooth \emphsurrogate loss (e.g., MSE/Huber loss) on data from a replay buffer. Yet, control is governed by the Bellman optimality operator's \gamma-contraction in \ell_\infty, whereas common objectives (MSE/Huber loss) measure progress in an averaged \ell_2/\ell_1 geometry; thus, improving the surrogate need not reduce the worst-case Bellman error that drives performance. To bridge this gap we connect the progress in the inner loop for each surrogate to the Bellman residual under the \ell_\infty norm. Our analysis provides explicit thresholds denoted by \alpha under which sufficient progress guarantees a \ell_\infty contraction up to a standard approximation floor. This yields new convergence guarantees for DQN with fixed targets trained using MSE or Huber-type losses, and motivates a novel soft-\ell_\infty surrogate that smoothly approximates the sup-norm and better matches the \ell_\infty geometry resulting in the least stringent inner loop accuracy requirements. Across Atari benchmarks, soft-\ell_\infty consistently reduces the worst-case Bellman residual and matches or improves returns relative to standard objectives, providing a practical, geometry-aligned recipe for stable DQN training.
Abstract: Causal autoregressive video diffusion models support real-time streaming generation by extrapolating future chunks from previously generated content. Distilling such generators from high-fidelity bidirectional teachers yields competitive few-step models, yet a persistent gap between the history distributions encountered during training and those arising at inference constrains generation quality over long horizons. We introduce the Real-time Autoregressive Video Extrapolation Network (RAVEN), a training-time test framework that repacks each self rollout into an interleaved sequence of clean historical endpoints and noisy denoising states. This formulation aligns training attention with inference-time extrapolation and allows downstream chunk losses to supervise the history representations on which future predictions depend. We further propose Consistency-model Group Relative Policy Optimization (CM-GRPO), which reformulates a consistency sampling step as a conditional Gaussian transition and applies online Reinforcement Learning (RL) directly to this kernel, avoiding the Euler-Maruyama auxiliary process adopted in prior flow-model RL formulations. Experiments demonstrate that RAVEN surpasses recent causal video distillation baselines across quality, semantic, and dynamic degree evaluations, and that CM-GRPO provides further gains when combined with RAVEN.
PaperID: 3628, Poster
Authors: Jonathan Kozaczuk, Zheng Dong, Noyan Evirgen, Keivan Majidi, Eddie K Ng, Chris Porter, Qi Zhao, Jia Geng, Changmao Li, Brandon Feng
Abstract: User embeddings are central components of modern recommendation, personalization, and behavioral modeling systems. Though these embeddings can compactly encode rich behavioral information, they are inherently dense and opaque, limiting interpretability. To address this gap, we show that user embeddings can be understood as sparse superpositions of interpretable behavioral modes. This structure is directly predicted by generative models of user behavior for linear embeddings and extends approximately to transformer representations by the superposition hypothesis. Superimposed modes can naturally be recovered by sparse dictionary learning methods, and recovered modes can be automatically named and queried in natural language. On synthetic data with known ground-truth modes, sparse autoencoders (SAEs) outperform classical sparse baselines on mode recovery and produce more monosemantic latents, even when the number of modes exceeds the embedding dimension. Applied to both classical (SVD on open e-commerce) and transformer-based (SASRec/BERT4Rec on MovieLens 1M) user embeddings, the recovered SAE dictionaries are near-orthogonal, and their natural-language labels enable zero-shot audience retrieval over embeddings never trained on text. At matched sparsity, SAE decompositions produce more behaviorally distinct features than classical sparse baselines, with retrieval accuracy comparable to both classical sparse methods and text-embedding baselines and more diverse audience characterizations. The framework is post-hoc and model-agnostic, requiring no modification to the upstream embedding system.
Authors:
Yann Bouquet, Alireza Khodamoradi, Sophie Y Shen, Kristof Denolf, Mathieu SalzmannAbstract: Post-training quantization (PTQ) of large diffusion transformers degrades generative quality at 4-bit quantization. Low-rank approximation methods are a promising solution and append auxiliary linear branches to restore performance. Current state-of-the-art approaches keep these branches at high precision (W16A16) and rely on heavy, data-dependent calibration for initialization. Their branch is defined as a precise low-rank approximation of a full-rank matrix and therefore cannot keep precision at sub-16 bit levels. We challenge both limitations with LoRaQ (Low-Rank Approximated Quantization), a data-free calibration approach that optimizes quantization error compensation. By overcoming the need for high-precision branches, LoRaQ enables the first fully sub-16 bit pipeline, allowing the low-rank branch itself to be quantized. We demonstrate that, at equal memory overhead, LoRaQ outperforms the state-of-the-art methods in their native implementations on Pixart-\Sigma, SANA and Flux.1. We also analyze mixed-precision configurations, showing that setups such as W8A8, W6A6, and W4A8 for the low-rank branch, alongside a W4 residual layer, yield superior results while maintaining a fully quantized architecture compatible with modern mixed-precision hardware, and translate into kernel-level speedups of up to 3.19× on AMD MI355.
Abstract: Masked diffusion language models (dLLMs) have recently emerged as a competitive alternative to autoregressive language models, with the promise of faster inference via parallel token generation. A notable limitation of the masked formulation, however, is that once a token has been unmasked it can no longer be revised, leaving dLLMs vulnerable to early sampling mistakes. To address this, a growing body of work has sought to extend masked dLLMs with self-correcting (remasking) capabilities. One appealing subset of these methods does so in a training-free, post-hoc manner based on token confidence, with encouraging early reported results. In this work, we revisit the empirical evaluation of a representative post-hoc remasking method, WINO, and find that under standard decoding settings (shorter block lengths) it brings little-to-no benefit over confidence-based unmasking alone. Extending the evaluation to non-greedy decoding, we find that while confidence-based remasking can mitigate errors introduced by increased stochasticity to some extent, it also exacerbates the diversity collapse previously reported for confidence-based unmasking. Overall, our results show that the benefits of post-hoc confidence-based remasking are highly setting-dependent, underscoring the need for a more comprehensive evaluation framework.
PaperID: 3631, Poster
Abstract: Practitioners using machine learning on noisy tabular data can often find simpler, more interpretable models with the same accuracy as black box models. In some cases, this is because the noise in the data induces a large Rashomon set, that is, a large set of near-optimal models, which then frequently contains simpler models. However, our understanding of this effect from discrete hypothesis classes --- the Rashomon set's size, the diversity of its predictions, and implicit regularization --- leads to a contradiction in continuous hypothesis spaces, and thus does not directly establish whether noise should necessarily induce simpler or more diverse models. We reconcile the seemingly conflicting effects by showing geometrically that noise drives simplicity through separate mechanisms for linear, logistic, and generalized additive models. We introduce the Rashomon reducibility cost to quantify whether a given model simplification is possible within the Rashomon set, and characterize how this cost changes under general data properties such as label noise. For practitioners working with noisy tabular data, our results suggest that simpler, more interpretable models will be included in the Rashomon set, so that black box models will generally not be needed.
Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in orchestrating tools for reasoning tasks. However, existing methods rely on a step-wise paradigm that lacks a global perspective, which causes error accumulation over long horizons and restricts generalization to unseen tools. To overcome these limitations, we propose Tools as Continuous Flow for Evolving Agentic Reasoning (FlowAgent), which reconceptualizes tool chaining as continuous trajectory generation within a semantic space. To systematically evaluate this paradigm, we introduce the first plan-level closed-loop benchmark dedicated to plan-level agentic reasoning in dynamic real-world environments. Specifically, the proposed FlowAgent leverages conditional flow matching to generate continuous latent trajectories, providing a global planning perspective to ensure coherent and robust tool execution. Theoretically, we establish formal bounds on utility convergence and prove that our continuous formulation fundamentally guarantees robust generalization and error attenuation. Empirical evaluations show that FlowAgent achieves superior robustness and adaptability in long-horizon reasoning tasks.
PaperID: 3633, Poster
Abstract: Anti-customization techniques aim to protect visual content from unauthorized diffusion model personalization by adding imperceptible perturbations to training images. However, such protection can be significantly weakened by diffusion-based purification, which may remove adversarial perturbations before customization. We study representative anti-customization methods under diffusion purification and observe that existing methods exhibit similar off-manifold perturbation structures that are largely suppressed by purification. We then provide a principled geometric analysis that characterizes purification as a local contraction toward the data manifold, primarily suppressing perturbation components along normal directions. Motivated by this insight, we propose Manifold-Aligned Customization protection , a practical purification-aware framework that optimizes downstream anti-customization objectives through a differentiable purifier and uses a structured perturbation parameterization to bias updates toward purification-retentive, manifold-aligned directions. Across experiments on CelebA-HQ, VGGFace2-HQ, and WikiArt under diffusion-based purification, MAC substantially improves protection effectiveness across most prompts and domains. Finally, frequency analysis and ImageNet and MS-COCO retention studies show that spectral signatures cannot predict purification outcomes, providing additional evidence for the proposed manifold-alignment perspective beyond a single data domain.
PaperID: 3634, Poster
Abstract: Multimodal Large Language Models (MLLMs) suffer from hallucinations, creating a critical need for Uncertainty Quantification (UQ) to ensure reliable deployment. However, existing approaches struggle to detect uncertainty caused by superficial associations, especially when the query-relevant signal is weak. We mainly attribute this issue to their bias toward aleatoric uncertainty arising from data ambiguity, overlooking epistemic uncertainty stemming from model limitations. To further decompose uncertainty types for a comprehensive UQ, we propose Causal-Invariant Masking (CIM), which measures the semantic shift between the original predictions and those conditioned on a causally-focused view. Based on this framework, we introduce Semantic Divergence as our core metric for UQ and provide theoretical evidence that it converges to the variance of model's sensitivity to non-causal correlations, establishing its ability to capture MLLM's limitation. To accelerate UQ in MLLMs, we further propose Expected Embedding Drift (EED), a fast geometric proxy metric that estimates semantic shift directly within the hyperspherical embedding space. Experiments show that our method achieves state-of-the-art performance on various benchmarks, while the proposed EED accelerates by nearly 50% with comparable performance.
PaperID: 3635, Poster
Abstract: Learning neural stochastic differential equations (SDEs) from trajectory data is a central task in scientific machine learning, but standard neural SDEs trade flexibility for coefficient fidelity: flexible parameterizations recover coefficients poorly, while structured architectures impose strong coefficient priors. We introduce , a three-stage pipeline that improves coefficient recovery through discovered symmetries without priors on the drift or diffusion. Building on the Itô SDE symmetry framework of Gaeta and Quintero (1999), we first train a flexible neural SDE surrogate, then discover approximate projectable Lie symmetries of the learned surrogate in a finite basis of canonical symmetry generators, and finally regularize the surrogate by penalizing the residual of the Itô symmetry determining equations. The discovery stage is , scalable, and theory-guided: the SDE symmetry Lie algebra has known dimensional bounds that constrain the symmetry search space. The resulting regularization is and improves existing neural SDE methods when combined with them. Across non-trivial SDEs spanning low and high dimensions and a real-data setting, Gaeta-Lie regularization consistently improves fidelity over multiple baselines, with gains that persist in the data-scarce regime as guaranteed by the symmetry residual's freedom from data sparsity. To our knowledge, this is the first method to use SDE Lie symmetry to improve a learned stochastic surrogate.
PaperID: 3636, Poster
Authors: Yan Shi, Kehua Chen, Jiarui Qi, Yihong Tang, Yinhai Wang, Liao Haicheng
Abstract: Recent multimodal large language models (MLLMs) can perceive a traffic scene, reason about it, and propose a driving action within a single model. However, existing MLLM-based drivers reason from scratch at each frame and commit to a single explanation, discarding earlier hypotheses about the scene. When the scene is ambiguous, the agent locks in on the wrong interpretation and acts on it before evidence accumulates. We propose DriveMind, an MLLM-based driving framework that keeps several hypotheses about the scene alive at the same time and updates their weights as new frames arrive. Inspired by particle filtering, DriveMind propagates each hypothesis through a VLM, scores it against the current observation, and resamples to maintain diversity. On Bench2Drive, DriveMind raises the Driving Score (DS) from 71.36 to 81.37 and the Success Rate (SR) from 45.45% to 54.09%, with the largest gains on ambiguous scenarios such as unsignalized intersections (+20.10 DS) and oncoming-traffic interactions (+20.50 DS). Beyond the overall metrics, DriveMind also improves fine-grained interactive driving abilities such as merging, emergency braking, and give-way. The code will be released soon.
PaperID: 3637, Poster
Abstract: Light Field (LF) deraining relies on leveraging redundant information between views to restore regions occluded by rains. Existing LF deraining methods typically rely on predicted depth maps to guide cross-view information interaction, while learning a brute-force mapping to separate rain from background content. However, the semi-transparent nature of rain streaks often leads to inaccurate depth estimation, which undermines geometric alignment and degrades restoration quality. More critically, the blending of rain and background during restoration makes brute-force mappings ineffective, resulting in visible rain artifacts in the output. In this paper, we first construct a consistency-constrained LF rain model by considering the consistency and the accumulation effects of rain layers. Then, we propose a novel Consistency-Guided Recovery Network (CGRNet) specifically designed for LF deraining. Inspired by the fact that background regions occluded by rain streaks often exhibit inconsistencies across all views, we introduce consistency maps to perceive degradation by aggregating pixel-wise consistency information along epipolar lines. Guided by consistency maps, the proposed model adaptively samples and aggregates consistent and redundant features from clean background regions to restore the details of degraded areas. Experiments validate our method's superiority over state-of-the-art approaches, and show that our dataset effectively enhances practical performance under real-world conditions.
Authors:
Jingyu Sun, Yuyang Xue, Mingyang Li, Zhengtao Yao, Jiachen Li, Yang Cui, Wenhao Cai, Haozhe Liu, Fangying Wang, Magdalene K Montgomery, Syed M Baker, Hongpeng ZhouAbstract: Large language model agents have shown strong capabilities in generating coherent and contextually appropriate responses, yet robust long-horizon dialogue remains limited by the lack of external memory that is traceable, updatable, and diagnostically transparent. Existing memory-augmented agents often store memories as isolated records or overwritable states, making it difficult to preserve how information originates, evolves, conflicts, or becomes obsolete over time. We propose TrajWiki, a trajectory-based memory framework for long-horizon conversational agents. Instead of treating memory as static entries, TrajWiki represents each memory as a source-grounded evolution trajectory, maintained through immutable episodic snapshots and claim-level operations such as ADD, REVISE, and DEPRECATE. To reduce fragmentation and retrieval cost, TrajWiki further introduces Memory Wiki, a persistent intermediate layer that incrementally compiles dialogue history into structured and interlinked wiki pages capturing salient entities, events, quantities, topics, and conflicts. At inference time, queries are routed hierarchically from relevant wiki pages to linked memory trajectories, then to corresponding snapshots and source messages for evidence-grounded answer synthesis. Experiments on LoCoMo and MedMT show that TrajWiki improves long-horizon dialogue performance across both open-source and closed-source LLM backbones, while providing greater interpretability and diagnostic visibility into memory evolution, retrieval failures, and answer generation. Code is available at https://anonymous.4open.science/r/TrajWiki-F0EC/.
Abstract: Memory-augmented LLM agents tackle complex long-horizon tasks by recursively summarizing interaction trajectories into compact memory. However, existing approaches typically train these memory policies using outcome-based reinforcement learning, failing to localize where intermediate memory quality degrades. As interactions unfold, ambiguous recursive summaries progressively discard or blur task-relevant information. This exacerbates belief deviation, obscuring the agent’s estimate of the latent task state and ultimately derailing long-horizon reasoning. We therefore argue that memory optimization should focus not merely on trajectory-level success, but on the clarity of the belief induced by intermediate summaries. To this end, we introduce Belief Entropy, a self-supervised proxy that probes how uncertain the model remains about the latent task state given its current memory. Based on this proxy, we propose Metacognitive Memory Policy Optimization (MMPO). Instead of relying only on sparse outcome-based signals, MMPO provides fine-grained, memory-specific supervision by explicitly penalizing summaries that induce high epistemic uncertainty. Experiments show that MMPO consistently outperforms existing methods on diverse long-horizon tasks, maintaining 97.1% performance even when scaled to 1.75M-token contexts.
PaperID: 3640, Poster
Abstract: We revisit the convergence guarantees of the Extragradient (EG) method for unconstrained bilinear min-max optimization. It is known that EG with a fixed stepsize achieves a \Theta(T^-1/2) last-iterate convergence rate, which is slower than the optimal \mathcalO(T^-1) rate attainable by incorporating additional mechanisms such as anchoring. Motivated by recent advances showing that dynamic stepsizes alone can significantly accelerate gradient descent, we ask whether dynamic stepsizes can similarly accelerate the last-iterate convergence of EG. We present the first positive result in this direction. Specifically, we provide a deterministic dynamic stepsize schedule that accelerates the convergence rate of EG to \mathcalO(T^-2/3+\varepsilon) for any \varepsilon > 0. We also show that this rate is tight when the extrapolation and update steps of EG use the same stepsize. We then show that allowing different stepsizes for the extrapolation and update steps further improves the convergence rate to the near-optimal \mathcalO(T^-1+\varepsilon). Our analysis reduces stepsize scheduling to an optimization problem, whose solution leads to a stepsize schedule that follows (a discretization of) a power-law distribution. Our proposed stepsize schedules and analysis extend to other methods, such as optimistic gradient descent, and suggest broader applicability to general min-max optimization problems.
PaperID: 3641, Poster
Authors: Eunho Koo, Jinwon Sohn, Tongseok Lim
Abstract: Statistical matching combines partially overlapping datasets that share covariates X but observe the target Y and auxiliary variables Z separately. Classical approaches typically invoke the conditional independence assumption (CIA), which makes the problem identifiable but fundamentally implies that the imported auxiliary variable provides no additional predictive power for Y once X is known. To capture this latent Y--Z dependence, we propose a novel dependency-aware Schr\"odinger bridge for predictive statistical matching. Our approach couples the two separated databases by tilting the conservative CIA baseline with a transportation-based compatibility cost, recovering an informative joint distribution. The resulting statistical learning framework yields full probabilistic posterior rules for bidirectional imputation. Theoretically, we establish a sufficient condition under which the learned bridge strictly improves over the CIA baseline, alongside an exact joint recovery guarantee in the Gaussian setting under an appropriate cost. Across synthetic benchmarks and real-world datasets (CelebA and Adult), we demonstrate that our dependency-aware completion consistently improves downstream predictive utility, proving especially beneficial in settings like data recoding where the underlying population exhibits strong Y--Z dependence.
PaperID: 3642, Poster
Authors: Kasra Fallah, Haoyu N Chen, Rudramani Singha, Eunji Kong, Gergely Turi, Attila Losonczy, Erfan Zabeh
Abstract: Local field potentials (LFPs) are population-level neural signals central to brain-computer interfaces and systems neuroscience, yet unlike spike-based population codes, their dynamics lack a principled geometric description that connects spectral structure to latent state-space geometry. Here we establish, analytically and empirically, that the lag-embedded dynamics of LFP signals lie on a low-dimensional K-torus, where K equals the number of sustained oscillatory modes present in the signal. Modeling LFPs as bounded autoregressive processes, we show that each oscillatory component contributes an independent circular degree of freedom, while aperiodic 1/f structure and noise contribute only geometric thickness around the manifold without altering its topology. To apply this theory to real nonstationary recordings, we introduce DeepLagField, a physics-informed network that jointly estimates time-varying AR structure and effective model order, with formal guarantees that local toroidal geometry is preserved despite drifting oscillatory dynamics.We validate the predicted toroidal geometry using persistent homology across primate visual cortex LFP, rodent hippocampal LFP, and mouse cortical EEG recordings, confirming the expected Betti number signatures across species and recording modalities. Critically, we demonstrate that this geometric structure is not merely descriptive but carries behaviorally relevant information inaccessible to standard spectral summaries — as evidenced by torus parameters derived purely from manifold geometry outperforming multi-band spectral features for sleep-state decoding without any hand-crafted frequency design. This proof of concept points toward broad downstream utility wherever oscillatory field signals are recorded, from neural decoding and brain-state monitoring to clinical biomarker development. These results reframe neural oscillations not as isolated spectral features but as coordinate directions of a low-dimensional delay manifold, opening a geometry-first approach to neural signal analysis that is simultaneously theoretically grounded, empirically validated across species, and predictive of behaviorally relevant brain states.
PaperID: 3643, Poster
Authors: Yihan Meng, Suping Xu, Yanfeng Wu, Chongjun Wang, Lin Shang
Abstract: Watermarking text-to-image diffusion models provides a practical mechanism for provenance tracking and responsible deployment. Post-processing methods are easy to deploy but vulnerable to image-space transformations. In-processing methods improve robustness by embedding the watermark within the denoising trajectory. Yet most existing methods inject the signal into the initial noise or a fixed intermediate latent, without controlling how it evolves during the remaining denoising steps. This can entangle watermark preservation with text-guided generation, risking weaker verification or degraded content fidelity. We propose TIDE, a trajectory-aware watermarking method for text-to-image diffusion models. TIDE injects a learnable watermark into an intermediate latent and jointly optimizes it with an auxiliary watermark condition, using preservation and detection objectives to limit deviation from the unwatermarked reference while maintaining separable watermark evidence. During sampling, TIDE estimates recent text-guidance directions and redirects watermark guidance toward the residual component less aligned with this local text subspace, reducing avoidable interference with semantic generation. Verification uses null-prompt partial inversion to recover the injection latent and match its masked coefficients to the target watermark. Experiments show that TIDE achieves reliable verification on unattacked images, improves robustness under common attacks, and better preserves semantic and visual fidelity.
PaperID: 3644, Poster
Abstract: Recent advances in Omni Language Models (OLMs) provide a unified solution for audiovisual video captioning. Nevertheless, existing approaches predominantly treat captioning as an unstructured text generation process, neglecting the inherent semantic structure of videos and leading to suboptimal fine-grained audiovisual alignment. To bridge this gap, we propose EVA-Cap, a novel framework that decomposes captions into atomic audiovisual events and transforms unstructured text supervision into event-centric alignment. We train EVA-Cap via Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO) on a self-curated dataset, EVA-Data, which is built with event-centric quality control. During the GRPO phase, the model is explicitly optimized against three event-centric alignment targets: global event completeness, intra-event attribution, and inter-event synchronization. Extensive experiments demonstrate that EVA-Cap significantly outperforms existing methods in event-centric evaluation (e.g., +5.3% average gain on ChronusAV), while achieving competitive performance on general audiovisual captioning benchmarks (e.g., +3.3% on WorldSense).
Abstract: Individualized treatment selection with continuous actions requires accurate causal response estimation in decision-relevant regions, rather than uniformly over the entire action space. Estimating a global causal response surface and then choosing the treatment that maximizes it can therefore be suboptimal, since standard estimation objectives allocate modeling effort according to the observed treatment distribution rather than the regions that determine the optimal decision. While decision-aware approaches have been studied in unconfounded settings, this problem remains underexplored in proximal causal inference, where proxy variables and bridge functions enable identification under suitable assumptions even in the presence of hidden confounding. Despite recent progress, proximal methods have primarily focused on treatment-effect and potential-outcome estimation rather than treatment selection and optimal decision-making. To bridge this gap, we introduce a policy-targeted weighted bridge loss that emphasizes decision-relevant treatment regions while retaining global stabilization. We prove a regret bound showing that the proposed weighted bridge loss controls treatment-selection regret through a weighted ill-posedness constant. We instantiate the framework in decision-aware variants of several proximal bridge solvers, yielding practical algorithms that alternate between weighted bridge estimation, response-surface projection, policy update, and weight refinement. Empirically, we find that decision-aware weighting reduces regret across several bridge solvers, suggesting improved treatment selection in proximal settings.
PaperID: 3646, Poster
Abstract: Despite progress in image tokenization, standard methods encode redundant information by mixing all granularities within each token, thus redundancy persists between tokens. The mix of information of different granularity also complicates the training of generators. This paper introduces SelfBootTok, a method that resolves this by cleanly decomposing information into global and local token groups. Through self-bootstrapped learning, the model predicts local details exclusively from global tokens, shifting the burden of visual details from the generator to the tokenizer. Consequently, our generator is far more efficient, requiring only global tokens and reducing computation by approximately 40%, while delivering superior reconstruction and generation. Moreover, this paradigm scales elegantly: by leveraging more data or parameters to self-supervise local representation learning, SelfBootTok achieves a new state-of-the-art gFID score of 1.56 using only 64 tokens.
PaperID: 3647, Poster
Abstract: Large language models (LLMs) have achieved remarkable success across diverse tasks, yet their ever-growing model sizes impose significant computational and memory costs that hinder practical deployment. Pruning addresses this by introducing sparsity into weight matrices, but existing layer-wise pruning methods operate independently on each linear layer, ignoring cross-layer dependencies and yielding globally suboptimal masks. End-to-end pruning methods address this limitation but incur prohibitive computational overhead. In this work, we propose FiedlerPrune, a pruning framework that achieves a better balance between the two paradigms by incorporating cross-layer dependencies. Guided by graph theory, FiedlerPrune constructs weighted multipartite graphs over consecutive linear layers within each Transformer block, and derives an inter-layer importance score from each weight's contribution to the Fiedler Value of the corresponding graph. Combined with any intra-layer metric, FiedlerPrune provides a unified and versatile pruning criterion reflecting both local quantitative importance and cross-layer connectivity-preserving importance . Extensive experiments demonstrate that FiedlerPrune consistently improves over layer-wise baselines, and matches or outperforms end-to-end methods at orders-of-magnitude lower pruning cost. We further demonstrate FiedlerPrune's practical utility through inference speedup, compatibility with quantization, and post-pruning fine-tuning.
Abstract: Continual learning seeks to acquire new knowledge while preserving previously learned representations. Despite the success of prompt-based methods, they remain fragile in domain-incremental learning scenarios, where shifts in the input distribution cause representation drift in shared layers and lead to forgetting. We introduce InTAct, a method that preserves functional behavior in shared representations without freezing parameters or retaining past data. InTAct constrains updates within activation regions associated with prior domains, while allowing flexible adaptation elsewhere, thereby stabilizing neuron functionality rather than directly restricting parameter values. The approach is architecture-agnostic, integrates seamlessly with existing prompt-based frameworks, and matches or outperforms state-of-the-art methods on domain-incremental benchmarks.
Abstract: Simulation-based inference with neural posterior estimation (NPE) often yields overconfident and unreliable posteriors under limited simulation budgets. To address this, we propose DRO-NPE, a distributionally robust approach that replaces the standard NPE objective with a worst-case loss over a Wasserstein ambiguity set. We introduce KL-based metrics for miscoverage and miscalibration, and use these to show that the DRO-NPE objective controls overfitting and reduces posterior overconfidence. Our method is tractable, parallelisable, and readily integrates with standard normalising flows. Across benchmark SBI tasks, DRO-NPE consistently improves coverage and calibration, while narrowing the gap between empirical and population NPE loss, leading to more reliable inference in low-simulation regimes.
PaperID: 3650, Poster
Abstract: We study policy regret minimization in partially observable Markov games (POMGs) between a learner and a strategic adaptive adversary who adapts to the learner's past strategies. We develop a model-based optimistic framework that operates on the learner-observable process using joint MLE confidence set and introduce an Observable Operator Model-based causal decomposition that disentangles the coupling between the world and the adversary model. Under multi-step weakly revealing observations and a bounded-memory, stationary and posterior-lipschitz adversary and planner stability, we prove an \mathcalO(\sqrtT) policy regret bound. This work advances regret analysis from Markov games to POMGs and provides the first policy regret guarantee under imperfect information against an adaptive opponent.
Abstract: Training-free camera control for pretrained flow-matching video generators is a partial-observation inverse problem: a depth-warped guidance video supplies noisy evidence on a subset of latent sites, which the sampler must reconcile with the pretrained prior. Existing methods struggle to balance the trade-off between trajectory adherence and visual quality and the heuristic guidance-strength tuning lacks robustness. We propose h-control, which resolves this dilemma through a structural change to the sampler: each outer hard-replacement guidance step is augmented with an inner-loop block-conditional pseudo-Gibbs refinement on the unobserved complement at the same noise level, with provable convergence to the partial-observation conditional data law. To accelerate convergence on high-dimensional video latents, we exploit their conditional locality, partitioning the unobserved complement into 3D patches, each tracked by a custom mixing indicator that adaptively freezes converged patches. On RealEstate10K and DAVIS, h-control attains the best FVD against all seven training-free and training-based competitors, outperforming every training-free baseline on every reported metric.
PaperID: 3652, Poster
Abstract: Fine-tuning Large Language Models under stringent memory constraints remains a significant challenge. Recently, Zeroth-Order (ZO) optimizers have been employed to fine-tune quantized models, avoiding backpropagation and further reducing memory usage. However, quantized ZO optimizers often underperform compared to their full-precision counterparts, especially at low bit-widths. We find that existing quantized ZO methods fail to exploit the inherent low-rank structure of gradient updates, thereby limiting their effectiveness. To address this, we propose the rder optimizer (LQZO), which enhances quantized ZO fine-tuning by leveraging the low-rank structure inherent in gradients. Our method introduces three key innovations: (1) A low-rank parameterization of gradient perturbations using structured matrix factorization. (2) A Binary-Aggregated Gaussian perturbation strategy that employs a Rademacher distribution to constrain higher-order moments and tighten the variance bound of gradient estimates, with no additional computational overhead. (3) An Isotropy-Enhanced Estimation technique that reshapes quantization scale groups into more isotropic structures to diversify exploration directions. We provide a theoretical convergence guarantee for LQZO. Extensive experiments demonstrate that LQZO consistently outperforms quantized ZO optimizers, achieving average performance gains of 2.0%--3.3% across quantization precisions. While extreme quantization inevitably leads to information loss, LQZO significantly improves the trainability of INT2 models and narrows the performance gap compared to prior quantized ZO methods. Furthermore, we demonstrate the cross-modality versatility of LQZO by validating its superior performance on Vision Transformers.
PaperID: 3653, Poster
Abstract: Dynamic Reconfigurable Battery (DRB) systems enable flexible interconnection among cells through power electronic switches, crucial for electric vehicles and energy storage systems. However, sensor failures in practical deployments lead to partial observation loss, posing severe challenges for safe control. Existing methods assume full state observability and often adopt overly conservative strategies when observations are missing, failing to achieve effective trade-offs between safety and performance under uncertainty. More fundamentally, these methods lack mechanisms to adaptively adjust decision conservatism based on inference reliability, leading to unstable decisions under complex failure scenarios. We propose the Diffusion-Guided Risk-Adaptive Framework (DGRAF) to address this challenge. The framework handles spatiotemporally coupled failure patterns through joint diffusion inference with quality quantification, and integrates inference quality with distributional decision-making to enable automatic modulation between aggressive optimization and conservative maintenance. Through extensive simulations and hardware experiments, we demonstrate that DGRAF consistently outperforms baseline methods across complex failure patterns, achieving superior energy efficiency, decision stability, and risk resilience.
PaperID: 3654, Poster
Authors: Tianyu Cao, Haixiang Sun, Yang Xu
Abstract: Decentralized Stochastic Approximation (SA) enables cooperative fixed-point solving over mesh networks without a central server. In these systems, local update is relatively cheap compared with coordination, and frequent information exchange makes communication cost the main bottleneck. To alleviate this issue, we propose a decentralized SA method in which each agent performs H local updates before each communication round. By amortizing communications over multiple local SA updates, the method can substantially reduce communication cost. We also study a multiple mixing scheme based on FastMix, which trades a small amount intra-round communications for a stronger consensus effectiveness on poorly connected graphs. Theoretically, we establish finite-time bounds under general contractive norms and Markovian sampling, by combining contraction, consensus error, and delayed Markovian terms in a round-level Lyapunov recursion. The results show trade-offs among computation, communication and sampling. Numerical experiments on decentralized Q-learning, Markovian SGD and linear SA validate the theory and show substantial communication savings for target accuracy.
PaperID: 3655, Poster
Authors: Savik Kinger, Johannes Bertram, Luciano Dyballa, Andy Keller, Steven W Zucker
Abstract: Understanding how transformers encode position and context is central to interpreting their behavior, because it determines how information flows across tokens and shapes the model's predictions. Common tools, such as linear probes and attribution methods, operate on single-position activations and are not designed to capture cross-position dependencies. We introduce Score-Block Token Geometry (SBTG), a score-based diagnostic that recovers cross-position dependency structure from activations and reveals distinct signatures of positional encoding mechanisms that these methods do not expose. SBTG trains a denoising score model on windows of activations and constructs lag-indexed operators that quantify coupling between positions within a layer. Applied to transformers with different positional encodings, it recovers distinct coupling signatures consistent with their architectures: ALiBi exhibits near rank-one coupling in early layers, RoPE distributes coupling across directions in patterns consistent with its rotation frequencies, and absolute embeddings show the strongest position-specific variation. We show that these signatures are stable across seeds and model sizes, and the activation-space directions identified by SBTG are causally relevant: ablating them causes larger performance drops on tasks that require reasoning over relative positions than on tasks that depend on absolute position signals. These results show that joint activation structure provides a lens on how positional encoding mechanisms are realized, revealing differences in how models use cross-position dependencies that are not apparent from the architecture alone.
PaperID: 3656, Poster
Authors: Sam McCallum, James Foster
Abstract: Training Neural ODEs requires backpropagating through an ODE solve. The state-of-the-art backpropagation method is recursive checkpointing that balances recomputation with memory cost. Here, we introduce a class of algebraically reversible ODE solvers that significantly improve upon both the time and memory cost of recursive checkpointing. The reversible solvers presented calculate exact gradients, are high-order and numerically stable -- strictly improving on previous reversible architectures. On scientific modeling and time series classification experiments, reversible solvers reduce the training time by more than 2× while using at least 10× less memory on average than recursive checkpointing.
PaperID: 3657, Poster
Authors: Zice Wang, Zhenyu Zhang
Abstract: Probing frozen vision transformers typically uses permutation-invariant aggregation (GAP or [CLS]), treating patch tokens as an unstructured set. Content-dependent probes such as self-attention are useful accuracy controls, but they do not expose a fixed token schedule or fixed position weights for auditing. We introduce SSMProbe, an explicitly inspectable probe that replaces invariant pooling with a Sinkhorn-learned evidence route followed by a diagonal S4 decoder. The S4 decoder is a linear time-invariant (LTI) system whose final state has fixed, position-dependent coefficients, so the probe-induced routed sequence can be audited as a concrete object rather than inferred only from accuracy. Our central measurement is the geometry of routed evidence: which patch tokens are moved to influential positions by this diagnostic, whether those tokens form spatially organized regions or random-like dispersed sets, and how the fixed S4 kernel weights them. Across MAE, BEiT, DINOv2, and supervised ViT, this route geometry separates MAE's dispersed, nearly random-like routes from the more spatially organized routes of BEiT, ViT, and DINOv2, with DINOv2 retaining a distinct strong [CLS] profile. SSMProbe uses the mathematical transparency of state-space models to turn a frozen ViT readout into an auditable evidence-routing analysis.
PaperID: 3658, Poster
Authors: Ismail Hassan, Pedro G Lind, Anis Yazidi
Abstract: Each player in a finite normal-form game observes only its own realized reward; opponents, their actions, and their payoffs are unknown. We analyze an epochal payoff-only recursion in which the mixed strategy is held fixed across long epochs, own-action means are estimated from bandit data, and the iterate steps toward an empirical near-best vertex of a barrier simplex. Our main result reduces this stochastic recursion to the deterministic best-response differential inclusion (BRDI). Under a Hoeffding-tied epoch schedule, the empirical near-best correspondence is contained, on a tail event of probability at least 1-\sum_k\ge K\delta_k, in a deterministic graph enlargement of the BRDI whose radius vanishes as k\to\infty. Without conditioning, the affine interpolation of the iterates is almost surely a bounded perturbed solution of the BRDI in the Bena\"im--Hofbauer--Sorin sense, and its \omega-limit set is internally chain transitive. Composing this reduction with deterministic BRDI confinement results yields last-iterate Nash convergence in three classes: \operatornamedist(x_k,\NE)\to 0 almost surely in finite zero-sum games, in finite exact-potential games (under an empty-interior condition on the potential's image), and in every generic 2× 2 game. Fixed barriers give O(b)-Nash equilibria of the original game; deterministic shrinking removes the bias.
PaperID: 3659, Poster
Abstract: Standard Flow Matching (FM) models typically employ a pointwise loss function, effectively assuming that the velocity vector at each spatial position is pointwise decomposed supervision given the conditioning. We argue that this Pointwise Independence Assumption is suboptimal for image generation, where strong local correlations exist across space. Neglecting these dependencies leads to artifacts such as structural inconsistency. In this paper, we introduce Contextual Flow Matching (CFM), a generalized framework that enforces consistency not just on individual pixels, but on the contextual relationships within the velocity field. We formulate the training objective using a set of generalized Context Operators---akin to convolution kernels---that abstract relationship constraints. Specifically, we instantiate a Differential Operator to capture high-frequency motion dynamics and an Aggregate Operator to stabilize low-frequency structural evolution. Extensive experiments on ImageNet demonstrate that CFM significantly accelerates convergence and enhances generation fidelity by explicitly modeling spatial dependencies.
Abstract: Existing decentralized stochastic optimization methods assume the lower-level loss function is strongly convex and the stochastic gradient noise has finite variance. These strong assumptions typically are not satisfied in real-world machine learning models. For example, learning on language data typically leads to heavy-tailed gradient. To address these limitations, we develop a novel decentralized stochastic bilevel optimization algorithm for the nonconvex bilevel optimization problem under heavy-tailed noise. Specifically, we develop a normalized stochastic variance-reduced bilevel gradient descent algorithm, which does not rely on any clipping operation. Moreover, we establish its convergence rate by innovatively bounding interdependent gradient sequences under heavy-tailed noise for nonconvex decentralized bilevel optimization problems. As far as we know, this is the first decentralized bilevel optimization algorithm with rigorous theoretical guarantees under heavy-tailed noise. The extensive experimental results confirm the effectiveness of our algorithm in handling heavy-tailed noise.
PaperID: 3661, Poster
Abstract: Direct Preference Optimization (DPO) offers a streamlined one-stage alternative to Reinforcement Learning from Human Feedback (RLHF). However, our theoretical analysis reveals a structural distinction between the two: DPO and RLHF optimize alignment on different data distributions with different supervision sources. Specifically, DPO learns from ground-truth human preferences but within the off-policy support of an offline dataset, while RLHF optimizes on the on-policy rollouts but learns from pseudo-labels produced by a reward proxy. This distinction explains how each method may fail: DPO alignment collapses when the policy drifts away from the offline dataset support, and RLHF alignment is biased due to the extrapolation error of the reward proxy on out-of-distribution rollouts. Effective alignment requires avoiding both failure modes, which neither method can do alone. We achieve this via a semi-supervised alignment objective: ground-truth preferences anchor alignment within the dataset support, and pseudo-labels on online rollouts extend it beyond. Applying DPO's reward reparameterization to this objective yields Ultra-DPO, a one-stage algorithm that augments direct human alignment with an online self-alignment mechanism. Across multiple benchmarks, Ultra-DPO consistently outperforms DPO, RLHF, and online preference learning methods.
PaperID: 3662, Poster
Abstract: Classical rate-distortion theory minimises expected distortion at a given coding rate but does not constrain the \emphuncertainty of the reconstructed signal at the decoder. We formalise predictive uncertainty U as the conditional variance of the reconstruction given the source, and show that it is an underconstrained degree of freedom of RD-optimal compression with a sharp empirical signature: at matched rate and matched scalar U, two reconstructions of the same image can differ by 40+ mIoU points in downstream segmentation depending on \emphwhere the variance is spatially allocated, with 88% of (image, quality) pairs showing a significantly negative within-image slope of mIoU on U. We derive a closed-form Uncertainty-Rate-Distortion surface and prove a coding-theoretic converse for additive-noise channels with linear regression under Gaussian assumption. A region-importance spatial-URD functional also converts the spatial-allocation finding into a within-image predictor that generalises across two segmentation architectures, three rates, two codecs, and a different segmentation foundation model on a different dataset. Cross-task and cross-modality probes (BLIP captioning, CLIP zero-shot, NYU depth, LibriSpeech speech recognition) characterise the framework's reach: the spatial-URD prediction transfers cleanly to locally-aggregating downstream models on both image and audio modalities, and degenerates predictably for globally-aggregating ones. The validation code and a demo for a diffusion task is provided as Supp. Mat., and will be released alongside with all the code to reproduce the experiments upon acceptance of the work.
Abstract: Mobility trajectory data provide essential support for smart city applications. However, such data are often difficult to obtain. Meanwhile, most existing trajectory generation methods implicitly assume that at least a subset of real mobility data from target city is available, which limits their applicability in data-inaccessible scenarios. In this work, we propose a new problem setting, called bus-conditioned zero-shot trajectory generation, where no mobility trajectories from a target city are accessible. The generation process relies solely on source city mobility data and publicly available bus timetables from both cities. Under this setting, we propose MobTA, the first approach to introduce task arithmetic into trajectory generation. MobTA models the parameter shift from bus-timetable-based trajectory generation to mobility trajectory generation in source city, and applies this shift to target city through arithmetic operations on task vectors. This enables trajectory generation that reflects target-city mobility patterns without requiring any real mobility data from it. Furthermore, we theoretically analyze MobTA's stability across base and instruction-tuned LLMs. Extensive experiments show that MobTA significantly outperforms existing methods, and achieves performance close to models finetuned using target city mobility trajectories.
PaperID: 3664, Poster
Abstract: Dynamic topology reconfiguration is central to the reliability and efficiency of large satellite constellations, yet many existing approaches rely on idealized assumptions such as full constellation deployment or uniform orbital spacing. We present Adaptive Satellite Topology via Regret-Aware learning (ASTRA), a theoretically-grounded framework for dynamic satellite topology reconfiguration that builds on an online learning formulation and makes it computationally practical. ASTRA combines an ADMM-based offline solver with efficient online updates for both online gradient descent and online conditional gradient, yielding markedly cheaper constrained updates than generic optimization pipelines. On the theory side, we show that for a relevant class of entry-wise nonzero utility matrices, the objective is strongly convex, which yields logarithmic static regret for online gradient descent, and we further instantiate known dynamic-regret guarantees under inexact ADMM inner loops. Empirically, ASTRA matches or improves topology quality, presenting a good trade-off with computational time on synthetic constellations, and it remains effective on real Starlink data under partial deployment and non-uniform spacing, where idealized structural assumptions break down. These results position ASTRA as an efficient and theoretically grounded approach to topology reconfiguration in realistic Low Earth Orbit networks.
Authors: Rafid U Murshed, Saif U Rahman, Mingyue Tang, Elahe Soltanaghai
Abstract: Radio maps are essential for wireless decision-making tasks such as access-point placement, coverage planning, and localization, but their fine spatial details are governed by complex propagation effects and are costly to simulate accurately. Machine learning offers a path to high-fidelity radio-map prediction without running expensive high-fidelity simulations for every scene. However, generating high-quality training labels at scale is also difficult: the affordable labels come from finite-ray simulations, which are richer than low-fidelity inputs but carry residual Monte Carlo noise. We address this challenge with Physics-Unrolled Hybrid Neural Operator (PU-HNO), a three-stage cascade that predicts high-fidelity indoor radio maps from low-fidelity ray-tracing outputs and scene priors by progressively capturing reflection, diffraction, and scattering effects, rather than treating radio maps as generic images. We prove that, under conditionally unbiased label noise, the model can learn stable propagation structure and outperform its own training labels. Experiments across diverse floorplans show that PU-HNO outperforms image-to-image baselines, wireless learning models, and monolithic neural operators across both image-quality and wireless deployment metrics.
PaperID: 3666, Poster
Authors: Yang Jiao, Kai Yang
Abstract: Trilevel learning has emerged as a powerful paradigm across various domains in machine learning and networking. In practical scenarios, data is frequently generated and distributed across multiple decentralized nodes. While existing distributed trilevel learning methods circumvent the transmission of raw data, they remain susceptible to privacy leakage. Furthermore, cutting plane methods, which are widely adopted in bilevel and trilevel learning, exhibit analogous intrinsic privacy risks. To address these critical issues, we propose a Differentially Private Distributed Trilevel Optimization (DPDTO) framework. In DPDTO, we first develop a value-function-embedded privacy-preserving cutting plane method for trilevel optimization problems, comprising an inner-layer value function reformulation and an outer-layer cutting plane relaxation. Building upon this, a distributed optimization algorithm is introduced to effectively address trilevel optimization while preserving privacy. Theoretically, we provide a comprehensive analysis for the proposed framework, including asymptotic convergent relaxation, non-asymptotic convergence rate, privacy guarantees, and communication complexity, uncovering a tripartite trade-off among privacy, utility, and complexity in trilevel optimization. Extensive experimental results on two distributed trilevel optimization problems further demonstrate the effectiveness and superiority of the proposed DPDTO.
PaperID: 3667, Poster
Authors: Hyunseok Seung, Matthias Katzfuss
Abstract: Gradient observations can substantially improve Gaussian process (GP) surrogates, particularly in high-dimensional settings where function evaluations are expensive. However, exact inference with n function values and n full gradients in d dimensions scales cubically in the joint state size, imposing an intractable \mathcalO(n^3 d^3) computational bottleneck. We introduce TERA, a highly scalable derivative GP method based on target-specific exact gradient reduction. We prove that for stationary kernels, the gradient components orthogonal to the directions connecting the target and conditioning points are conditionally independent of the target function value; consequently, the exact conditional density is fully characterized by at most m^2 directional derivatives once a conditioning set of size m is specified. By using these reduced, dimension-free conditionals as local factors in a Vecchia approximation, TERA effectively decouples n and d from the dense matrix inversion. This reduces the per-target evaluation cost to \mathcalO(dm^2 + m^6) time and \mathcalO(dm^2 + m^4) memory, leaving the underlying derivative GP model mathematically unchanged. Empirical evaluations demonstrate that TERA achieves state-of-the-art predictive accuracy while operating orders of magnitude faster than standard derivative GPs. Crucially, both computation time and peak GPU memory remain essentially flat with respect to d, enabling highly scalable inference in high-dimensional spaces.
Authors:
Rajeev Yasarla, Deepti Hegde, Hsin-Pai Cheng, Shizhong Han, Yunxiao Shi, MeysamSadeghi, Hanno Ackermann, Litian Liu, Pranav Desai, Fatih Porikli, Mohammad Ghavamzadeh, Herbert CaiAbstract: Vision-language-action (VLA) models are effective as end-to-end motion planners, but can be brittle when evaluated in closed-loop settings due to being trained under traditional imitation learning framework. Existing closed-loop supervision approaches lack scalability and fail to completely model a reactive environment. We propose MAPLE, a novel framework for reactive, multi-agent rollout of a dynamic driving scenario in the latent space of the VLA model. The ego vehicle and nearby traffic agents are independently controlled over multi-step horizons, while being reactive to other agents in the scene, enabling closed-loop training. MAPLE consists of two training stages: (1) supervised fine‑tuning on the latent rollouts based on ground-truth trajectories, followed by (2) reinforcement learning with global and agent‑specific rewards that encourage safety, progress, and interaction realism. We further propose diversity rewards that encourage the model to generate planning behaviors that may not be present in logged driving data. Notably, our closed-loop training framework is scalable and does not require external simulators, which can be computationally expensive to run and have limited visual fidelity to the real-world. MAPLE achieves state-of-the-art driving performance on Bench2Drive and demonstrates scalable, closed-loop multi-agent play for robust E2E autonomous driving systems.
PaperID: 3669, Poster
Abstract: Reversible adversarial examples (RAEs) can disrupt access by malicious AI models while ensuring recoverability for authorized users. However, practical image circulation is dominated by JPEG files, whereas most existing RAE methods perturb spatial-domain images and are therefore poorly matched to compression, re-saving, and coefficient quantization. Beyond JPEG compatibility, circulation imposes a stricter and underexplored requirement that protected images should be perturbed in the JPEG domain, remain adversarial after spatial reconstruction, preserve visual fidelity, incur modest file-size growth, and remain recoverable under channel attacks such as recompression, noise, cropping, resizing, and platform re-saving. This paper presents the first systematic study of robust reversible adversarial protection for post-processed JPEG images and proposes SRAP-JPEG, a Synchronized Reversible Adversarial Protection framework on quantized DCT coefficients. SRAP-JPEG decouples adversarial perturbation from reversible recording by injecting perturbations into luminance coefficients, hiding recovery records in paired chrominance coefficients through a secret-key mapping, and adopting coefficient-adaptive allocation to trade off reversible capacity, attack strength, visual quality, and file-size growth. It crafts protective perturbations directly in the JPEG coefficient domain by propagating gradients from classification or vision-language objectives through the chain rule. For authorized recovery, SRAP-JPEG integrates block-phase synchronization, embedding-state identification, and coefficient restoration, enabling exact recovery of intact protected JPEGs and high-fidelity recovery after practical distortions. Experiments on ImageNet and MS-COCO show that SRAP-JPEG achieves an 81.82% average attack success rate on ten mainstream classifiers, degrades CLIP-based image-text retrieval, and recovers distorted protected images with high visual fidelity, including 41.60 dB PSNR after JPEG recompression at QF 70 and 59.51 dB on the aligned uncropped region after edge cropping at ratio 0.25.
PaperID: 3670, Poster
Authors: Miroslav Lžičař
Abstract: Two finite-domain implementations, including neural-network parameterizations restricted to an evaluation domain, can agree on every input while implementing the computation through different mechanisms, such as independent lookup storage or a shared reusable rule. We formalize this distinction as a statement about the local Kullback-Leibler geometry of an implementation chart: a mechanism-specified local family of perturbations around a fixed predictive distribution. For finite input-output distributions with full support, we prove a logit-Fisher pullback theorem: the Hessian rank of a chart equals the dimension of its tangent image in the quotient logit space, so the quadratic active contribution to the local learning coefficient is half this rank under the analytic/dead-kernel conditions stated below. We prove that this rank is invariant under local reparameterizations, dead parameters, and redundant duplications, but changes when the allowed perturbation class changes; hence it is a first-order chart-germ invariant, not a function-only invariant. For mechanisms specified by exact finite-group sharing, we give a closed-form orbit-count formula for the active dimension. Lookup and shared cyclic modular addition are then immediate corollaries: both implement the same softened conditional distribution, yet their active ranks are p^2(p-1) and p-1, respectively. The results clarify when Bayesian/Singular-Learning-Theory complexity can distinguish extensionally equivalent algorithms and when chart dependence is a real modeling choice rather than a defect.
Authors:
Hong-Han Wang, Yuntao Wang, Hu DingAbstract: Layer-wise visual-text similarity in Multimodal Large Language Models (MLLMs) is widely interpreted as evidence that the language model progressively integrates visual content into a shared representation space. This reading rests on the assumption that scalar alignment scores reflect content-level cross-modal interaction. To test this assumption, we apply controlled interventions to the visual stream. Across 13 MLLMs from five families spanning 0.5B to 72B parameters, replacing projector-output visual tokens with Gaussian noise sharply reduces task accuracy, yet four standard scalar measures (CKA, SVCCA, MIR, and the leading principal-angle cosine) fail to consistently separate the corrupted stream from the original. We call this failure the \emphalignment illusion and trace it to the shared language-model pathway: anisotropic MLP down-projections pull visual and text tokens toward common output directions, producing \emphweight-induced alignment. Because this component is essentially one-dimensional, we introduce the principal-angle gap (PA gap), defined as the difference between the top two principal-angle cosines, which separates weight-induced similarity from multi-directional visual structure. Under graded visual corruption, the PA gap tracks task accuracy more consistently than the scalar scores we consider; under a structured but irrelevant image, it further exposes regimes in which internal geometry and task accuracy come apart. Internal visual-text alignment in MLLMs is therefore best read as a geometric diagnostic of the visual stream inside the language model rather than a direct proxy for content-level cross-modal interaction, and is most informative when calibrated by controlled task evidence.
PaperID: 3672, Poster
Authors: Brandon Gower-Winter, Georg Krempl
Abstract: Data Streams are characterised as a potentially infinite source of data subject to non-stationarity. This non-stationarity, also called Concept Drift, often results in deployed predictive models having to be perpetually retrained as the underlying data generating process of the stream changes over time. Drift Detectors are a class of methods used to identify these changes, triggering these retraining procedures. Historically, Stream Learning has operated under the assumption that instances are only seen once, and cannot be stored for later use. Contemporary work relaxes this assumption yet drift detection and simple resetting procedures are still ubiquitous. Building on related work which has shown that traditional Batch Learning algorithms often outperform Stream Learning algorithms when instances can be stored, we introduce Time Neutralising Trees (TNT), a Decision Tree architecture that enables robust stream classification in settings with Concept Drift. During training, TNT will filter out ("neutralise") old training samples that are no longer relevant given the current context of the stream. If those instances become relevant again, TNT will then re-include them. We evaluate TNT, and related algorithms on 24 real and semi-real classification data streams from the USP DS repository. Results show strong evidence that TNT achieves state-of-the-art performance on a wide-range of data streams with varying concept drift types.
PaperID: 3673, Poster
Authors: Fredrik Cumlin, Saikat Chatterjee
Abstract: Particle filters are a standard tool for nonlinear state estimation, but their resampling step is discrete, preventing gradient-based learning in variational sequential Monte Carlo. We introduce Differentiable Systematic Resampling (DSR), a temperature-controlled relaxation of systematic resampling, that preserves the low-variance structure of systematic resampling while enabling full gradient flow. DSR converges to exact systematic resampling as the temperature vanishes, and we prove exponential convergence rates for the induced bias. Compared to optimal-transport-based differentiable resampling, DSR avoids iterative solvers and has substantially lower computational overhead. Experiments on stochastic dynamical systems and real-world handwriting data show that DSR achieves comparable or superior filtering and dynamics learning performance.
PaperID: 3674, Poster
Authors:
Shukai Duan, Mihai Capotă, Guixiang Ma, Hou-Jen Ko, Rafael Wang, Steve Cho, Shahin Nazarian, Paul BogdanAbstract: Tuning Triton GPU kernel launch configurations is a critical but labor-intensive step in production ML system deployment. Existing approaches rely on exhaustive autotuning, which requires manually defined search spaces and becomes prohibitively slow for the long tail of operator shapes. We present , a multi-agent system that combines LLM-guided search with closed-loop empirical validation to optimize Triton kernel configurations safely and efficiently. TritonTune pairs two domain-specialized LLM agents with complementary expertise in data movement and compute resource allocation, resolves their proposals through a lightweight orchestration layer, and refines the best candidate through a phased benchmark search on real hardware. Per-agent feedback drives targeted corrections when regressions occur, and a hardware-validated fallback provides a worst-case guarantee absent from prior LLM-based approaches: the system can only improve performance or leave it unchanged. On 19 TritonBench kernels and 393 input shapes spanning hand-written, per-shape geometric-mean speedup on an NVIDIA V100 with 80% of shapes improved, under a closed-loop fallback that guarantees no accepted shape regresses below the original. Cross-GPU runs on L40S, A40, and A100 further confirm that our design generalizes beyond V100. These results show that LLM-guided multi-agent optimization, when grounded in hardware measurement, can serve as a practical, safe complement to exhaustive autotuning.
PaperID: 3675, Poster
Authors:
Bingzhi Wang, Jiajun Yu, Dengke Zhong, Jiaming Liu, Xiong Li, Hongbo Liu, Jie YangAbstract: Acoustic sensing is a promising technique for human-computer interaction, but existing learning-based systems still depend heavily on large task-specific labeled datasets. Compared with visual data, acoustic signals are difficult to interpret and annotate, making supervision costly and hard to scale. Meanwhile, existing unsupervised representation learning methods, particularly contrastive learning, are largely designed for natural images and rely on data augmentation to form positive pairs. Such strategies can destroy the temporal-physical structure of acoustic signals, causing shortcut learning rather than meaningful representation learning. To address this issue, we propose MCAS, a signal-processing-based multi-view contrastive learning framework for acoustic sensing. Instead of relying on artificial augmentations, MCAS constructs structure-preserving positive pairs from complementary representations derived from different signal processing pipelines of the same acoustic event. Moreover, MCAS is designed to address the heterogeneity of acoustic representations through progressive cross-view interaction enabled by Hierarchical Shared-Token Fusion and a multi-level pretraining objective with physical consistency regularization. Extensive experiments on two typical acoustic sensing tasks demonstrate the effectiveness of the proposed framework. Under low-label settings, MCAS consistently outperforms both supervised baselines and generic self-supervised baselines, while also showing strong cross-task transferability and robustness to corruptions and noise.
PaperID: 3676, Poster
Abstract: Open-Vocabulary Human-Object Interaction (HOI) detection aims to localize interacting human-object pairs while generalizing to novel interaction categories beyond the training set. Although Multimodal Large Language Models (MLLMs) exhibit strong open-vocabulary understanding, directly applying MLLMs to HOI detection remains challenging, as language-prior hallucinations and coupled localization-semantic errors can destabilize structured HOI predictions. To address these issues, we propose RRL-HOI, a novel Reflective Reinforcement Learning framework that trains MLLMs to follow a verify-and-revise self-correction policy. Specifically, RRL-HOI first adapts MLLMs to produce box-grounded HOI triplets, and then introduces an executable reflection mechanism that targets the two failure modes through structured edits over localization, human-object pairing, and interaction semantics. In this way, localization and pairing edits mitigate localization-semantic coupling, while interaction-semantic edits suppress visually unsupported language-prior predictions. To further improve this reflection mechanism, RRL-HOI formulates reflection as a reinforcement learning objective trained with GRPO, where the reward jointly considers localization quality, pair-association correctness, interaction correctness, and edit consistency. Through reward-driven policy learning, RRL-HOI turns open-vocabulary recognition from a single-step generative prediction into an explicit verify-and-revise decision process, thereby bridging the gap between general multimodal understanding and precise interaction detection. Extensive experiments on two standard open-vocabulary HOI detection benchmarks demonstrate the effectiveness of our method.
Abstract: Learning effective policies for adaptive data acquisition remains challenging: posterior-based methods rely on surrogate models and posterior approximations that can be misspecified or biased, while direct policy-learning methods map from historical observations and fail to exploit available model representations, making learning harder. We introduce policy learning with belief representations (POLAR), based on the insight that optimal data acquisition depends on the observation history only through a sufficient belief state. Specifically, POLAR decouples representation learning from policy learning by leveraging pretrained predictive foundation models as belief-state encoders, training a policy head on top of their representations. This yields a simple, unified amortised policy learning framework for Bayesian experimental design, Bayesian optimisation, and active learning, differing only in the task-specific utility used to train the policy. Empirically, we find that POLAR outperforms state-of-the-art amortised methods across diverse tasks while requiring far fewer training samples, demonstrating a significant step in the scalability and efficiency of amortised data acquisition.
Abstract: Looped language models repeat a set of transformer layers through depth, reducing memory costs and providing natural early-exit points at loop boundaries. However, looped models do not scale as favorably as standard transformers with unique layers. We compare standard and Mixture-of-Experts (MoE) transformers, with and without looping, and find two main results. First, we find Looped-MoE models scale better than the standard baseline while dense looped models do not. We trace this to routing divergence between loops: in Looped-MoE models, different experts are activated on each pass through the same shared layers, recovering expressivity without additional parameters. Our second finding is that looped models have better compute-quality trade-offs with early exits than standard models. Because each loop ends with the same layers that produce the final output, loop boundaries are superior exit points, as confirmed by faster output convergence at these points. In sum, we provide a clear direction for scaling looped models: a Looped-MoE model with early exits can not only beat standard transformers at scale, but also enable significant memory and inference savings with minimal degradation in quality.
Abstract: Attention serves as the fundamental mechanism for long-context modeling in large language models (LLMs), yet dense attention becomes structurally prohibitive for long sequences due to its quadratic complexity. Consequently, sparse attention has received increasing attention as a scalable alternative. However, existing sparse attention methods rely on coarse-grained semantic representations during block selection, which blur intra-block semantic boundaries and lead to the loss of critical information. To address this issue, we propose Punctuation-aware Hybrid Sparse Attention (PHSA), a natively trainable sparse attention framework that leverages punctuation tokens as semantic boundary anchors. Specifically, (1) we design a dual-branch aggregation mechanism that fuses global semantic representations with punctuation-enhanced boundary features, preserving the core semantic structure while introducing almost no additional computational overhead; (2) we introduce an extreme-sparsity-adaptive training and inference strategy that stabilizes model behavior under very low token activation ratios. Extensive experiments on general benchmarks and long-context evaluation tasks, together with ablation studies on punctuation types and the linguistic origin of punctuation tokens, demonstrate that PHSA consistently outperforms both dense attention and the state-of-the-art sparse attention baseline InfLLM v2. Specifically, for a model under the training-inference consistent setting with an input sequence length of 32k tokens, PHSA reduces information loss by 10.8% at a sparsity ratio of 97.3%.
PaperID: 3680, Poster
Abstract: Facial affect understanding is not merely an image-to-label prediction problem: it requires identifying activated facial Action Units (AUs), modeling how AUs facilitate or inhibit one another, and explaining how AU configurations give rise to expression states. Such causal relations have long been studied in cognitive psychology and psychophysiology, where FACS, the Component Process Model, and motor synergy theories provide rich causal priors about facial behavior. In computer vision, however, data-driven facial affect models still largely learn correlational structures; even structured AU-aware methods often produce dependencies that are not well aligned with psychological priors. This leaves a persistent gap between biological facial mechanisms and learned visual representations, as observational facial images do not provide direct biological interventions and existing methods lack a principled way to recover psychology-aligned relations from data. We propose CausalAffect, a weakly supervised, causally guided framework for learning psychologically aligned facial affect relations from data. CausalAffect models two complementary relation types, AU\rightarrowAU and AU\rightarrowExpression, within a two-level hierarchy: a global graph capturing population-level psychology-aligned relations, and a sample-adaptive graph refining this backbone for individual variability. Reliable relation learning is supported by disentangled AU bottleneck representations, polarity-aware message passing, and feature-level counterfactual intervention. The learned relations recover canonical literature-supported pathways, reveal plausible inhibitory and sample-specific dependencies, are positively validated in a blind expert study, and exhibit intervention-consistent behavior under image-level AU edits. Experiments on six benchmarks show consistent improvements on both AU detection and facial expression recognition.
Abstract: Massive datasets in modern machine learning have made data reduction a central challenge, particularly for clustering tasks where memory and computational constraints demand compact yet faithful summaries. A standard approach is to construct an \epsilon-\text\emphcoreset: a small weighted subset that approximately preserves the clustering cost for every plausible choice of centers. For the (k,z)-\text\emphclustering problem, existing worst-case bounds on coreset size are essentially tight, ruling out substantially smaller coresets in general. However, such worst-case instances are often unrepresentative of real-world data. In this work, we show that significantly smaller coresets are possible under mild and natural assumptions on the underlying data distribution. We introduce a new correlated sampling framework, called \text\emphdeterminantal sampling, based on a novel application of determinantal point processes. Using this framework, we obtain an efficiently constructible \varepsilon-coreset for (k,z)-clustering in \mathbb R^d whose dependence on 1/\varepsilon has exponent strictly smaller than 2 when d is fixed. This improves over the worst-case \varepsilon^-2 barrier under our beyond-worst-case assumptions. To the best of our knowledge, this is the first result that provably surpasses these lower bounds through beyond-worst-case assumptions. Finally, we validate our approach on synthetic and real-world benchmark datasets, where it consistently achieves smaller coresets than existing state-of-the-art methods, even without explicitly enforcing the assumptions used in the analysis.
PaperID: 3682, Poster
Abstract: We study the problem of agnostic PAC learning under deterministic labels, where the examples are labeled by an arbitrary deterministic function, which need not belong to the concept class \mathcalC. This model captures misspecification without stochastic label noise and lies between the realizable and fully agnostic settings. Prior work showed that, when \mathcalC has finite diameter K and VC dimension d, deterministic labels can improve the sample complexity upper bound from the classical agnostic guarantee \mathcalO(d/\varepsilon^2) to \widetilde\mathcalO(K/\varepsilon), but left a large gap between upper and lower bounds. Our main result sharpens this picture. We prove that every class of VC dimension d and finite diameter K is learnable with sample complexity \widetilde\mathcalO\left( \fracK^1/3 d^2/3\varepsilon \right), and we construct classes for which this bound is tight up to logarithmic factors. We also study the online analogue of this problem and show that the minimax regret is again partly controlled by the diameter of the class. Technically, our upper bounds are obtained via memorization-based learning algorithms that exploit the deterministic label structure.
PaperID: 3683, Poster
Abstract: Concept erasure works well for static visual attributes in text-to-image diffusion, but does not transfer to motion concept erasure in video diffusion transformers. We trace this failure to an off-policy trajectory mismatch: once weights are edited, the generation trajectory followed by the edited model deviates from the original trajectory used for supervision, so state-local velocity matching fails to control the motion patterns realized at inference. As a consequence of this mismatch, off-policy weight modification can drive training loss to convergence while the target motion remains fully present in generated videos. We formalize this gap through a Gronwall-type bound showing how weight-induced velocity perturbations propagate into trajectory divergence, and a residual-error bound showing that off-policy training, even at perfect convergence, leaves a nonzero erasure error on the edited trajectory. To address this problem, we introduce DOME (Drift-Adaptive On-Policy Motion Erasure), a trajectory-aligned framework for motion concept erasure in video diffusion transformers. DOME trains on states from the edited model's own trajectory and applies \emphdrift-adaptive anticipation, which shifts supervision along the observed divergence between the edited and original trajectories, turning trajectory drift from a failure symptom into a training signal. DOME edits the text-conditioned pathway via cross-attention LoRA, and the learned adapter can be merged into the base model so that inference incurs no additional overhead. On Wan~2.1-T2V across 20 motion concepts, DOME consistently suppresses localized contact actions while showing limited effect on whole-body dynamics, validated through both motion-consistency scoring and an independent action classifier. Systematic ablations show that drift-adaptive anticipation, not on-policy rollout alone, is the critical component distinguishing DOME from off-policy and standard on-policy training.
PaperID: 3684, Poster
Abstract: Post-Training Quantization (PTQ) has become a prerequisite for efficient LLM deployment; however, current rotation-based methods are approaching a saturation point because they prioritize heuristic approximations over theoretical rigor. By treating quantization as an opaque black box, existing paradigms typically depend on indirect gradient estimators for activations while restricting weight optimization to local reconstruction objectives. Consequently, this structural limitation induces fundamental deficiencies: specifically, optimization instability and a "twofold blindness" toward accumulated upstream distortion and final task loss sensitivity. To overcome these barriers, we introduce DICEQuant, a unified framework established upon rigorous error modeling. First, we propose the CURE (Coupled Underflow-Rounding Error) surrogate, which facilitates exact and stable gradient computation for activation reshaping, thereby eliminating the variance inherent in heuristic estimators. Simultaneously, for weight optimization, we present Distortion-Compensated Rounding (DCR). This mechanism derives a Shifted Optimal Center to neutralize input noise and utilizes Hessian-weighted shaping to align rounding decisions with the global functional objective. Empirical results demonstrate that DICEQuant significantly surpasses state-of-the-art baselines, successfully bridging the accuracy gap between low-bit compression and full-precision intelligence.
Abstract: Anonymizing human-centric video data is an understudied problem. Prior anonymization techniques either blur or redact pixels at the cost of realism and downstream utility, or generate frame-by-frame at the cost of temporal coherence. We introduce ReGenHuman, the first full-body video anonymization pipeline that is simultaneously realistic, temporally consistent, and anonymous by construction. Contrary to past approaches which redact or edit the inputs directly, we propose a regenerate, don't edit paradigm. Our approach composites 2D pose, segmentation, and monocular depth into two complementary conditioning streams---StructAll and StructHuman, which are used to fine-tune a video-to-video diffusion backbone on in-the-wild human videos, synthesizing the human regions entirely from identity-free structural cues. We evaluate our model on privacy, quality, and utility, and show that our ReGenHuman achieves the best tradeoff across all three axes against current baselines. We further show that our anonymized videos remain effective for downstream tasks, including video question answering.
PaperID: 3686, Poster
Abstract: Long-context LLM inference is bottlenecked by the memory demands of the key-value (KV) cache, whose size scales linearly with context length. Dynamic sparse attention reduces this bottleneck by fetching and attending only to a critical subset of tokens. The dominant primitive in dynamic sparse attention research and deployment is fixed-budget top-k selection, which is favored because its deterministic block count aligns with the static memory allocation that systems like vLLM rely on. Top-k uses a fixed budget regardless of how attention mass is distributed across heads, layers, and decode steps, leading to either wasted bandwidth on heads with peaky attention or dropped context on heads with diffuse attention. Top-p (nucleus) selection adapts naturally to this variation but has seen limited adoption in production serving systems due to two system-level challenges: variable block counts that break deterministic block allocation and computationally expensive selection kernels that exceed the cost of attention itself. We present DisSparse, a serving system that integrates top-p sparse attention into vLLM end-to-end. DisSparse decouples block selection from the scheduler's critical path, hides selection latency behind layer computation via a two-stream pipeline, and introduces a single-pass histogram-based top-p kernel. We further observe intra-block sparsity, the inflation of selection budgets caused by mixing high- and low-importance tokens within fixed-size KV cache blocks, and address it with a token relocation kernel guided by accumulated prefill attention scores. Across long- and medium-context benchmarks on multiple models, DisSparse achieves up to 1.45× speedup in sparsity analysis (block selection) over state-of-the-art top-p methods with no accuracy loss and supports up to 8× larger maximum batch sizes than dense vLLM under the same memory budget.
Authors: Dan DeGenaro, Xin Li, Obed Amo, Michael Pokojovy, Sarah Bargal, Markus Lange-Hegermann, Bogdan Raita
Abstract: We introduce FLASH-MAX, a shallow, exact-by-construction neural network architecture for predicting homogeneous electromagnetic fields from sparse pointwise observations. Each hidden neuron represents a separate exact solution to Maxwell's equations, so that the network satisfies the governing equations symbolically by construction and can be trained end-to-end from sparse data within seconds. We prove a universal approximation result showing that this exact model class remains universal on arbitrary domains. FLASH-MAX reaches sub-1% relative validation error from about 1K sparse pointwise observations in seconds, all while maintaining a zero PDE residual, and keeps single digit errors even for only 100 observations sampled from 3D space. These results suggest that moving governing structure from the loss into the hypothesis class can dramatically improve the trade-off between precision and optimization speed in scientific machine learning.
PaperID: 3688, Poster
Abstract: Real-world time series typically exhibit non-uniform information density across temporal positions, frequency structures, and forecasting horizons: critical dynamics are concentrated in local segments, different frequency components carry different structural information, and different horizons require different modeling complexity. However, most existing methods still allocate computation approximately uniformly across these dimensions, leading to redundant computation in low-information regions and insufficient modeling of key patterns. To address this, we propose BucpTSF, an information-density-driven structured framework for time series forecasting. Specifically, BucpTSF reorganizes raw sequences into hierarchically heterogeneous multi-level token representations through information-density-aware temporal representation reorganization, captures key cross-token dependencies with low overhead via frequency-selective structured token interaction, and further introduces horizon-conditioned structured extrapolation to adaptively balance dynamic modeling and structural extrapolation bias for different horizons, thereby improving the accuracy and stability of long-term forecasting. Extensive experiments on real-world datasets show that BucpTSF consistently outperforms strong baselines under various forecasting settings, validating the effectiveness of information-density-driven non-uniform computation allocation for time series forecasting.
PaperID: 3689, Poster
Abstract: Unsupervised video object-centric learning requires temporally consistent representations for downstream tasks. Existing methods mainly build temporal consistency on the assumption that adjacent frames share the same object composition. However, this local assumption is insufficient for real videos. Some objects remain semantically consistent over much longer temporal spans, while object composition may also change as objects appear or disappear. Fixed-slot recurrent Slot Attention architectures struggle to accommodate such visibility changes, limiting their ability to learn stronger temporal consistency beyond adjacent frames. We propose Dynamic Prototypical Contrast (DynaProto) to address these challenges. DynaProto introduces a dynamic recurrent Slot Attention architecture that determines whether each slot should be activated at each frame, thereby adapting to changing object visibility over longer temporal scopes. In addition, DynaProto introduces a prototypical contrastive objective that provides stronger clip-level temporal consistency supervision. Together, these two components enable DynaProto to learn more stable slot binding over extended temporal ranges. Extensive experiments show that DynaProto learns more temporally consistent slot representations and substantially outperforms prior methods on object discovery and object property prediction on real-world datasets, including an 11.0% mBO improvement on YouTube-VIS.
PaperID: 3690, Poster
Authors:
Changjian Zhou, Junfeng Fang, Negin Yousefpour, peng wu, Bin Yan, Guillermo A NarsilioAbstract: Neural operator models trained on simulation data often lose accuracy when applied to experimental measurements due to the sim-to-real gap. Standard fine-tuning with limited real data can reduce this gap, but it may also damage the core physics-relevant representations learned during pretraining. Although knowledge-preserving adaptation has been widely investigated in vision or language tasks, it remains unclear whether these methods are suitable for neural operators whose architectures and protected knowledge are fundamentally different. Neural operators need to preserve core-scale physical structures rather than semantic or visual features. We propose PhysGuard, a physics-preserving framework for accurate sim-to-real adaptation of neural operators. Specifically, PhysGuard uses the empirical Fisher Information Matrix computed on simulation data to identify physics-critical parameter directions, then restricts fine-tuning updates to directions that do not interfere with them. A layer-wise Gram-matrix formulation makes this efficient for models with millions of parameters, while an adaptive threshold automatically determines the protected subspace size. A spectral probe experiment shows that the dominant Fisher directions are strongly associated with low-frequency output structures. Experiments on benchmark across four neural operator architectures and different physical systems show that PhysGuard performs strongly on most evaluation metrics compared to baselines. The benefits are most evident under severe domain shift, where it reduces low-frequency error by up to 32% compared to standard fine-tuning while maintaining adaptability. Our code is available at https://anonymous.4open.science/r/PhysGuard-B840/.
Authors: George Andriopoulos, Zixuan Dong, Bimarsha Adhikari, Keith Ross
Abstract: Neural multivariate regression underpins a wide range of domains such as control, robotics, and finance, yet the geometry of its learned representations remains poorly characterized. While neural collapse has been shown to benefit generalization in classification, we find that analogous collapse in regression consistently degrades performance. To explain this contrast, we analyze models through the lens of intrinsic dimension. Across control tasks and synthetic datasets, we estimate the intrinsic dimension of last-layer features (ID_H) and compare it with that of the regression targets (ID_Y). Collapsed models exhibit ID_H < ID_Y, leading to over-compression and poor generalization, whereas non-collapsed models typically maintain ID_H > ID_Y. For the non-collapsed models, performance with respect to ID_H depends on the data quantity and noise levels. From these observations, we identify two regimes—over-compressed and under-compressed—that determine when expanding or reducing feature dimensionality improves performance. Our results provide new geometric insights into neural regression and suggest practical strategies for enhancing generalization.
PaperID: 3692, Poster
Abstract: Graph foundation models have emerged as a promising paradigm for generalizing across diverse structural tasks. Current approaches serialize graphs into node token sequences and apply autoregressive or masked prediction, yet this node sequential modeling disrupts the relation-centric nature of graphs and violates permutation equivariance. We propose , a framework that reconceptualizes graph modeling through edge co-occurrence modeling. By treating edges as fundamental modeling units and learning their joint distribution via discrete diffusion processes, our approach naturally respects the relational semantics of graphs while maintaining permutation equivariance. We introduce topological edge tokenization to enrich edge representations with multi-hop structural context, and relational prompt completion to unify node classification, link prediction, and graph classification as sequence completion problems over edge representations. Through extensive experiments, we demonstrate that relational diffusion achieves strong performance while providing a principled foundation aligned with the relational nature of graphs.
PaperID: 3693, Poster
Authors:
Wentao Liu, Xiabao Wu, Yongchao Liu, Haitao Zhang, Jiajun Zheng, Ruiting ZhouAbstract: Long-context LLM inference is fundamentally a retrieval process, transitioning from diffuse to sharply focused attention at a few architecture-determined (CTLs). Existing acceleration methods leave significant costs unaddressed: sparse attention ignores FFNs, KV cache compression neglects prefill, and hidden-state pruning relies on fixed, heuristic rules. We introduce CutAttn, a training-free framework that exploits CTLs via inter-layer Jensen--Shannon divergence, pruning redundant context at chunk granularity while dynamically gated by a normalized-entropy criterion. Operating at the hidden-state level, CutAttn reduces both attention and FFN costs, composing orthogonally with existing acceleration techniques. Evaluated on RULER, InfiniteBench, and LongBench, CutAttn preserves full-attention accuracy while achieving up to a 2.02× prefill speedup on 128K sequences (and 2.81× when composed with sparse attention). Furthermore, our theoretical analysis provides approximation bounds, an entropy-compression limit justifying the gating, and a closed-form speedup ratio.
Abstract: Sampling from complex, unnormalized probability densities is a fundamental challenge in Bayesian inference and probabilistic modeling. While Markov chain Monte Carlo (MCMC) methods provide asymptotic guarantees, they often suffer from slow mixing and high computational costs due to fixed or manually tuned trajectory lengths. In this work, we propose a novel framework that treats trajectory termination as a learnable component of the sampling dynamics. By framing MCMC within the theory of non-acyclic Generative Flow Networks (GFlowNets), we train state-dependent neural classifiers to decide when a trajectory has reached a high-density region and should terminate. We theoretically establish the connection between optimal classifiers and the target density via detailed balance conditions and introduce a multilevel training scheme to facilitate exploration in complex geometries. Experimental results across various benchmark densities demonstrate that our approach significantly reduces average trajectory lengths while improving mode coverage and mixing compared to standard MCMC baselines.
PaperID: 3695, Poster
Abstract: Text-to-video generation has made remarkable progress with diffusion models and transformer-based video architectures, yet generating coherent long videos from evolving textual descriptions remains challenging. Training long-video models from scratch requires substantial computation and large-scale multi-event data, while training-free multi-prompt generation must activate new prompt semantics without breaking temporal continuity. We identify a frequency-agnostic boundary coupling problem in existing transition strategies: weak coupling causes abrupt changes, flicker, and background jumps, whereas overly strong or coarse coupling may over-preserve source semantics and produce ghosting or duplicated subjects around prompt boundaries. To address this problem, we propose FreqSync, a training-free framework that formulates prompt transitions as frequency-selective boundary coupling. FreqSync combines Frequency-Synchronized Source-Conditioned Attention (FS-SCA) with Residual-Clipped Spectral Stitching (RCSS), enabling adjacent segments to share stable low- and mid-frequency structure and motion cues while keeping high-frequency prompt-specific details target-dominant under explicit residual constraints. Extensive experiments show that FreqSync achieves state-of-the-art transition smoothness and prompt alignment while maintaining strong perceptual video quality, demonstrating its effectiveness for coherent multi-prompt long video generation.
PaperID: 3696, Poster
Abstract: Full-space electron density generation is fundamental to modeling molecular quantum states, ground-state properties, and chemical interactions. DFT provides a principled route to electron density, but high-fidelity calculations require costly self-consistent iterations, limiting large-scale quantum-chemical modeling. Deep learning offers a promising surrogate, yet existing methods often rely on indirect density-related prior, including coordinate queries, predefined basis functions, orbital coefficients, or spherical atomic density grids. These prior introduce reconstruction overhead and can hinder continuous full-space modeling across sharp nuclear cores and sparse vacuum regions. Here, we introduce PI-EDG, a physics-informed voxel-space framework that maps molecular geometries directly to full-space electron-density fields by combining quantum-mechanical prior injection, hybrid FNO-CNN modeling of local and non-local interactions and magnitude-aware DS-LMoE routing for extreme density ranges. On EDBench, PI-EDG achieves an MAE of 3.405×10^-3, a total electron-number MAE of 0.472, and an average inference time of 133.4 ms for one molecule. Compared with graph-based, basis-based, and voxel-based baselines, PI-EDG demonstrates better predictive accuracy, improved robustness across extreme density regimes, and better physical consistency while maintaining efficient full-space generation. These results suggest that PI-EDG offers a new, direct, and efficient route for accelerating DFT-based electron density calculations.
PaperID: 3697, Poster
Authors: Jisui Huang, Yue Wu, Ke Chen, Na Lei
Abstract: The data fidelity term plays a crucial role in variational image segmentation. Recently, optimal transport (OT)-based variational models have improved segmentation by globally matching predicted feature distributions with reference distributions. However, existing OT-based methods usually rely on color or intensity similarity, which can be insufficient when foreground and background regions share similar appearance. We address this limitation by incorporating geodesic distance into the OT cost on a Riemannian feature manifold, allowing the transport cost to jointly encode spatial proximity, appearance similarity, and boundary information. The proposed model is well-defined: we prove existence of minimizers and convexity of the energy functional. It is also tractable: the OT terms admit differentiable transport-map representations, avoiding storage of the full pairwise transport matrix and enabling efficient gradient-based optimization. Experiments on BraTS and RESECT show that our method consistently outperforms traditional variational models, OT-based baselines, and prompt-based deep segmentation models, particularly when foreground and background have overlapping intensity or color patterns.
PaperID: 3698, Poster
Authors: Menghui Zhou, Vitaveska Lanfranchi, Po Yang
Abstract: Many real-world applications require highly interpretable machine learning systems, particularly in high-stakes domains such as healthcare. A promising recent direction is the paradigm of Maximal Coding Rate Reduction (MCR^2), which seeks to characterize and preserve the low-dimensional structure underlying each class. However, we find that exhaustively preserving structural information is often unnecessary and can even be detrimental in supervised learning, especially under noisy and uncertain real-world conditions, as it may overly constrain the fitting flexibility of the model. In contrast, we propose a simple yet effective framework, termed SimCoding, which captures only minimal structural information for each class. Despite relying on substantially less structural information, SimCoding significantly improves the generalization ability and robustness of deep models. Extensive experiments across diverse benchmark datasets demonstrate the effectiveness of SimCoding. We further validate SimCoding on a challenging real-world application, Parkinson’s disease severity assessment from free-living human activity signals, where it consistently achieves strong robustness and outperforms competing methods. Moreover, SimCoding can be naturally extended to incremental learning scenarios, where it substantially alleviates catastrophic forgetting and consistently surpasses strong baselines.
PaperID: 3699, Poster
Abstract: The expressivity of graph neural networks is predominantly characterized in a noiseless setting via binary graph isomorphism tests. This perspective certifies whether two graphs can be distinguished, but does not quantify how robust such distinctions are under structural perturbations. In practice, however, real-world graphs are often corrupted by noise, where even small structural changes might lead to disproportionate distortions in learned representations. In this study, we generalize graph isomorphism to a noisy alignment criterion, requiring that graphs that remain identifiable under noise are embedded closer to their source than to any alternative. This yields a margin-based formulation of expressivity that bridges the gap between combinatorial isomorphism-based theory and the continuous learned representations. Notably, our framework is architecture-agnostic and strictly generalizes classical isomorphism-based expressivity. We further derive certified robustness guarantees for node-level tasks and an explicit robustness radius for graph-level representations. Experiments confirm that our criteria predict and improve robustness under structural perturbations.
Authors:
Pengcheng Fang, Hongli Chen, Yuxia Chen, Tengjiao Sun, Jiaxin Liu, Xiaohao CaiAbstract: Spectral token mixers based on Fourier transforms provide an efficient way to model global interactions in visual feature maps. Existing designs often either apply filter-wise spectral responses along fixed channel axes, or learn adaptive frequency-indexed channel mixing without explicitly aligning the channel directions used across frequencies. We propose CHASM, a Cross-frequency Harmonized Axis-Separable Mixer, as a structured middle ground. CHASM separates what should be shared from what should remain frequency-specific: all frequencies share a learned channel eigenbasis, while each frequency retains its own positive spectral gains. The shared basis makes channel directions comparable across the spectrum, whereas the positive gains preserve local spectral adaptivity. CHASM applies this structured operator separably along the height and width axes and is used as a drop-in replacement mixer inside existing backbones. We provide a structural characterization of the shared-basis operator family and evaluate CHASM through controlled same-backbone comparisons. Across accelerated MRI reconstruction, undersampled MRI segmentation, and natural-image reconstruction, CHASM consistently improves over same-backbone spectral-mixer baselines. Ablations show that removing the shared-basis constraint weakens performance, and randomizing coherent sampling geometry substantially reduces the gain, supporting cross-frequency harmonization as a useful inductive bias for spectral token operators.
PaperID: 3701, Poster
Abstract: Direct Preference Optimization (DPO) has become a popular route for aligning large language models with human preferences, but existing pipelines treat all preference pairs uniformly and ignore prompt-level heterogeneity. We show that the optimal pair-selection strategy varies systematically with prompt characteristics: prompts with concentrated reward distributions benefit from large reward-gap pairs that provide unambiguous signal, while prompts with diverse responses benefit from moderate-gap pairs because extreme positions are dominated by outliers. No fixed strategy is optimal across this heterogeneity. We propose META-PAP (Meta-learning for Prompt-aware Preference Pairing), a framework that turns this observation into a learned algorithm. META-PAP generates multiple diverse pairs per prompt via a structured grid over the response reward distribution, extracts a compact prompt-aware feature vector, and trains a lightweight meta-network through bilevel optimization to predict per-pair weights. The inner loop is a weighted DPO update; the outer loop evaluates a virtual policy on a small clean meta-set. Across Llama-3-8B, Mistral-7B-v0.3, and Llama-2-7B, META-PAP consistently outperforms vanilla DPO, fixed-position pairing, reward-gap weighting, top-k filtering, curriculum, and statistical rejection sampling, with average gains of +3.9 points on AlpacaEval~2.0 and +3.2 points on Arena-Hard, and approaches an oracle that exhaustively searches per prompt. The learned weighting policy is interpretable: it favors large-gap pairs on low-diversity prompts and softly rejects extreme pairs on high-diversity prompts. We additionally find that the meta-network requires only M\!=\!150 pairs to saturate, transfers across domains, tolerates label noise on the meta-set, and remains effective with as few as n\!=\!32 on-policy samples per prompt.
Abstract: Reinforcement learning with verifiable rewards (RLVR), along with recent self-distillation variants such as SDPO, evaluates each rollout against a verifier and updates the policy from that episode-level signal. However, the richer procedural information in the rollout is rarely retained or reused. Across episodes and epochs, the model repeatedly encounters related problems under a changing policy, producing cross-episode signals that episode-local updates cannot capture: which strategies consistently pass verification, which failure modes persist, which patterns recur. We propose Procedural Memory Distillation (PMD), which converts these cross-episode signals into reusable procedural memory and distills it into the policy's weights during training. This memory functions as a training scaffold, absorbed into the policy itself, yielding a memory-free model at inference. PMD organizes the memory at three levels of abstraction: raw trajectories, self-reflected strategies and lessons, and higher-level behavioral patterns that recur across problems, all extracted online from the model's own trajectories. A memory-conditioned self-teacher draws on the accumulated experience to supervise the student on its own rollouts, enabling student to progressively internalize procedural knowledge within its parameters. The central design principle is co-evolution: the policy generates rollouts that update the memory, and memory shapes the supervision that updates the policy. Empirically, across Qwen3-8B and OLMo3-Instruct-7B, PMD improves over SDPO by 3.8--5.5% on SciKnowEval and 7.9--13.6% on LiveCodeBench. Co-evolution powers these gains: freezing either the memory or the policy trails PMD by more than 10% across different SciKnowEval domains.
PaperID: 3703, Poster
Authors: Daniele Micheletti, Federica Zoe Ricci, Erik Sudderth
Abstract: Graphs play a crucial role in many applications, ranging from drug discovery to social network modeling and astronomy. Generative models for graphs have substantially advanced in recent years, but we identify relatively simple scenarios where state-of-the-art models struggle to scale and exhibit prohibitively-slow generation times. We propose a generative architecture that matches or exceeds the state-of-the-art, but can scale to regimes where their training fails, and is orders of magnitude faster at generation. Our approach is inspired by the Aldous-Hoover theorem, a classic representation theorem that characterizes any probability distribution of graphs that is invariant to permutations of node indices (i.e., exchangeable). This theorem establishes that three ingredients are needed: a graph-level variable, a set of node-specific variables, and a so-called graphon function that maps graph- and node-level variables to the probability of an edge between any node pair. Given a training set of graphs, our exchangeable Variational Graph AutoEncoder (xVGAE) learns an approximation to their underlying graphons, as well as graph- and node-level latent variables. We show that our xVGAE encodes interpretable representations of graph- and node-level properties, and can (quickly, in one shot) generate new graphs that closely match key training set statistics.
PaperID: 3704, Poster
Authors: Donggeun Kim
Abstract: Mixture-of-Experts (MoE) enables multimodal large language models (MLLMs) to scale capacity efficiently, but expert parameters dominate memory—accounting for over 90% of model size in modern MoE-MLLMs (e.g., 54 GB out of ∼58 GB in BF16 for InternVL3.5-30B-A3B and Qwen3-VL-30B-A3B)—even though only a few experts are activated per token. Expert pruning has emerged as an effective compression strategy, but existing methods aggregate importance uniformly across tokens, ignoring that the same expert may serve different roles for vision versus language processing. To address this, we propose MAEP (Modality-Aware Expert Pruning), a training-free framework that decomposes expert importance by modality in a single forward pass. MAEP introduces (1) Delta-H Output (∆H), a redistribution-aware metric measuring the MoE output change when an expert is removed, and (2) low-attention vision tokens as a vision-side pruning signal that is dataset-stable: its expert-importance ranking remains stable across different datasets. Text-side importance serves as a preservation signal while vision-side importance serves as a pruning signal, and experts are ranked globally across layers. Experiments on three MoE-based MLLMs across multimodal and text benchmarks show that MAEP outperforms modality-agnostic baselines on the majority of multimodal benchmarks while remaining competitive on text-only tasks, reducing the BF16 expert weight footprint by ∼27 GB (∼47% of total memory) so that the expert-weight footprint fits within a single 40 GB GPU.
PaperID: 3705, Poster
Authors: Lindong Xie, Yang Zhang, Beiyu Song, XING Zeren, Flora Salim, Edward Chung
Abstract: Traffic signal control (TSC) is critical for mitigating urban congestion. Deep reinforcement learning (DRL) has achieved strong performance by learning adaptive policies from traffic observations, but learned neural controllers are often opaque, structurally complex, and difficult to generalize across scenarios. We propose PPL-TSC, a programmatic policy learning framework that distills compact and human-readable TSC policies from DRL teachers. PPL-TSC formulates policy learning as an evolutionary search over phase utility functions (PUFs), which score candidate signal phases and select the highest-utility phase for execution. Each generation consists of four steps: policy generation, policy review, policy calibration, and policy evaluation. In the generation step, a large language model (LLM) acts as a semantics-aware evolutionary operator to propose and refine candidate PUFs. Generated PUFs are screened through explicit checks for syntactic validity, mathematical soundness, and consistency with TSC principles, with invalid ones repaired through iterative feedback. Valid PUFs are then calibrated to improve behavioral fidelity to the DRL teacher, and evaluated using a joint score combining teacher fidelity and traffic-control performance to select candidates for the next generation. Extensive experiments across diverse traffic scenarios and representative DRL-based controllers show that PPL-TSC produces compact and human-readable policies, matches teacher performance, improves transfer generalization, and substantially reduces computational cost.
PaperID: 3706, Poster
Abstract: In reinforcement learning (RL), reward specification plays the fundamental role in shaping policy optimization. Inverse RL (IRL) offers a data-driven approach to inferring rewards from demonstrations, but prevalent neural-network-based methods often produce rewards lacking interpretability and flexibility for manual inspection or revision. In contrast, human-written or large language model (LLM)-generated reward code provides explicit semantic structure, yet its precise numerical parameters are often difficult to calibrate and prone to specification errors. To bridge this gap, we propose DiReCL (Differentiable Reward Code Learning), a novel framework that unifies these two paradigms by formulating reward learning in a differentiable reward code space. DiReCL first prompts an LLM to synthesize the reward code templates, then converts the numeric elements in these templates into differentiable parameters and align them with expert demonstrations through IRL. We further introduce a reflection mechanism that uses IRL diagnostics to provide evidence of inadequacies in the reward structure and guide template revision. Extensive experiments on MuJoCo locomotion and autonomous driving benchmarks demonstrate that DiReCL achieves over 1.25 × the performance of both standard IRL and LLM-assisted reward design baselines, while preserving the crucial benefits of explicit, interpretable reward semantics.
PaperID: 3707, Poster
Authors: Ritwik Vashistha, Abhra Sarkar, Arya Farahi
Abstract: Latent signals are often obscured by measurement noise, yet encode the underlying laws and dynamics of complex systems; learning both the signals and their distributions remains a central challenge in scientific inference. The noise is often non-negligible, and the likelihoods for expressive generative models are often intractable. We utilize a convolutional maximum mean discrepancy (convMMD) loss and propose a likelihood-free framework for nonparametric density deconvolution and empirical Bayes denoising under additive measurement error. Our method learns a latent generative model by matching the observed data distribution to the noise-convolved model distribution. This yields a differentiable, simulation-based objective for multivariate homoscedastic or heteroscedastic noise, compatible with expressive sieve classes such as Gaussian mixtures and normalizing flows. The learned density then serves as an empirical prior for posterior denoising of individual latent values. Theoretically, we extend convMMD from parametric to nonparametric estimation, proving finite-sample bounds for empirical sieve minimizers and L_2 convergence rates under Sobolev smoothness. These rates recover the classical inverse-problem dependence: polynomial for ordinary-smooth and logarithmic for super-smooth noises. Our method provides a practical, theoretically grounded approach to deconvolution and denoising under generative latent distribution models.
PaperID: 3708, Poster
Authors:
Lulu Wang, Shengling Wang, Anlin Chen, Ke Chao, Weicheng WangAbstract: In real-world tasks such as robotic control and autonomous driving, agents often face non-Markovian environments due to incomplete observations. Solving such information-impoverished decision tasks requires reconstructing the context from history. However, existing recurrent encoders have two fundamental limitations: First, the model suffers from a Markovian compression bottleneck, i.e., it only processes observation history as input but neglects thought process history. This leads to the loss of deep semantic cues from early interactions, which are easily overwritten by noise over time due to recursion. Second, it lacks the adaptive capability to adjust the memory span for different tasks. To address these limitations, we propose FracEncoder, a context encoder based on fractional-order neural differential equations. First, FracEncoder leverages the nonlocal property of the fractional-order derivative and uses a power-law kernel to directly incorporate the trajectory of the thought process into the evolution of the current state, effectively overcoming the temporal locality limitations of traditional models. Second, the fractional order \alpha is set as an end-to-end learnable parameter. It explicitly characterizes how strongly the system anchors to its history, so the model can discover a suitable memory span for each task. We evaluate FracEncoder on more than ten benchmarks across three settings, namely partially observable tasks, meta reinforcement learning, and delayed-observation control. FracEncoder consistently matches or outperforms representative baselines, including GPIDE, GRU-ODE, ReSeL, and LRU. On the challenging 8-step delayed-observation task, it reaches average returns of -511.5 \pm 7.9 on Pendulum and 732.1 \pm 152.7 on Hopper. As a plug-in module, FracEncoder also yields stable gains across different reinforcement learning algorithms. Ablation studies further show that the learned \alpha aligns with the intrinsic memory demand of each task, providing an interpretable physical measure of the cognitive complexity of non-Markovian tasks. The code is available at https://anonymous.4open.science/r/FracEncoder-60BA.
Abstract: The 1-identification problem is a fundamental pure-exploration problem in multi-armed bandits. An agent aims to determine whether there exists an arm whose mean reward exceeds a known threshold \mu_0, or to output \textsfNone otherwise. The agent must guarantee correctness with probability at least 1-\delta, while minimizing the expected number of arm pulls \mathbbE[\tau]. We study the 1-identification problem and make two main contributions. First, for instances with at least one qualified arm, we derive a new lower bound on \mathbbE[\tau] via a novel optimization formulation. Second, we propose a new algorithm and establish upper bounds that match the lower bounds up to polynomial logarithmic factors uniformly over all instances. Our result complements the analysis of \mathbbE\tau when there are multiple qualified arms, which is an open problem in the literature.
PaperID: 3710, Poster
Authors: Hankyeong Ko, Taehoon Kim
Abstract: Generative AI models now produce images indistinguishable from real data, making publicly verifiable provenance essential; however, existing watermarking methods become vulnerable to forgeries and adversarial optimization attacks once their detectors are made public. We propose PP-Mark, a provenance framework that anchors lightweight statistical detection to a zero-knowledge proof of the embedding process, ensuring that forged content cannot pass full verification by construction. We establish formal unforgeability under standard cryptographic assumptions and validate PP-Mark against representative baselines across two architectures, demonstrating resistance to black-box imprint forgery, white-box optimization, and regeneration attacks while remaining practical for decentralized verification.
PaperID: 3711, Poster
Abstract: Safe Policy Improvement (SPI) aims to improve a baseline policy offline using fixed data while avoiding performance degradation with high probability. Existing SPI methods work in discrete domains but do not extend to continuous control. We propose Calibrated Safe Policy Improvement (Cal-SPI), a Monte Carlo Tree Search-based continuous SPI method that replaces count-based support with calibrated model trust. Cal-SPI learns a deep ensemble dynamics model and calibrates ensemble disagreement on held-out data through a Wilks-style tolerance construction, turning raw uncertainty into a statistical certificate of one-step model error. This certificate defines a hard gate within model-based tree search. We provide a theoretical analysis showing that calibrated model trust and a baseline-relative switching criterion yield a conditional PAC-style improvement bound for gate-certified root decisions. Experiments across three MuJoCo continuous domains show that Cal-SPI avoids performance degradation in low-data regimes and improves over the baseline once calibrated model trust becomes reliable.
PaperID: 3712, Poster
Authors: Nathaniel L Diamant, Brian L Trippe
Abstract: Generative models can produce individually plausible samples while deviating substantially from a reference set in the distribution of key features. For example, a model pretrained on broad drug-like chemical space may generate molecules whose molecular features differ from those of a therapeutic class of interest, such as known antibiotics. Correcting such distributional is challenging: direct finetuning on the reference set can overfit and does not control which features are matched. To fill this gap, we introduce kernel Calibrating Generative Models (kCGM). kCGM minimizes a maximum mean discrepancy (MMD) between generated and reference feature distributions using an unbiased score-function estimator, with KL regularization to remain close to the pretrained model. On a reference set of 174 antibiotics, direct finetuning sacrifices chemical validity for feature-distribution matching, whereas kCGM improves target feature matching while increasing validity. We further demonstrate kCGM in protein and DNA generation tasks, showing it can adapt autoregressive, continuous-space diffusion, and discrete diffusion models using only feature-level supervision.
PaperID: 3713, Poster
Abstract: As language models become autonomous software agents, documentation shifts from explanatory prose to a natural-language behavioral specification for code generation. We study agent-oriented documentation generation, where the objective is not readability but the correctness of code synthesized from documentation alone, turning the task into a black-box search over natural-language specifications evaluated only through the generated code and its execution behavior. Its defining difficulty is output coupling: program entities are behaviorally entangled, so a revision that sharpens the specification of one entity may simultaneously destabilize its dependents, inducing the familiar whack-a-mole failure mode in iterative refinement. We introduce SpecForge, a multi-frontier tree search solution that preserves complementary search frontiers through Pareto state management to avoid premature commitment and escape certain frontier-local optima, performs dependency-constrained bandit selection for callee-before-caller refinement, and uses diversified error-conditioned expansion to remain robust to noisy test feedback. On DevEval+, SpecForge attains the best performance across five backbone models, including a 94.9% solve rate with Claude-4.5-Sonnet and a 25.7% average improvement over the leading baseline. We further demonstrate that the resulting specifications improve downstream performance on both cross-language code translation and new-feature implementation.
Abstract: We introduce the metagame, a conceptual framework for quantifying second-order interaction effects of model explanations. For any first-order attribution \phi(f) explaining a model f, we measure the directional influence of feature j on the attribution of feature i, denoted as meta-attribution \varphi_j \to i(f), by treating the attribution method itself as a cooperative game and computing its Shapley value. Theoretically, we prove that attributions hierarchically decompose into meta-attributions, and establish these as directional extensions of existing interaction indices. Empirically, we demonstrate that the metagame delivers insights across diverse interpretability applications: (i) quantifying token interactions in instruction-tuned language models, (ii) explaining cross-modal similarity in vision-language encoders, and (iii) interpreting text-to-image concepts in multimodal diffusion transformers.
Abstract: Multi-agent large language model (LLM) systems often rely on a controller to coordinate a pool of heterogeneous models, yet existing controllers are typically limited to one-shot routing: they select a model once and return its output directly. Such routing-only designs provide no mechanism to critique intermediate drafts or support iterative refinement. To address this limitation, we propose a critique-and-routing controller that casts multi-agent coordination as a sequential decision problem. At each turn, the controller evaluates the current draft, decides whether to stop or continue, and, if needed, selects the next agent for further refinement. We formulate this process as a finite-horizon Markov Decision Process (MDP) with explicit agent-utilization constraints, design a composite reward for controller decisions across turns, and optimize the controller via policy gradients under a Lagrangian-relaxed objective. Extensive experiments across multiple heterogeneous multi-agent systems and seven reasoning benchmarks show that our method consistently outperforms state-of-the-art baselines and substantially narrows the gap to the strongest agent, while using it for fewer than 25% of total calls.
PaperID: 3716, Poster
Abstract: With the rising prominence and fluency of large language models (LLMs), developing technologies to identify LLMs-generated text has become increasingly critical. However, existing technologies depend on static linguistic features, which can be evaded as advanced models increasingly mimic a wide range of writing styles. The study reveals two crucial vulnerabilities in existing detection systems:(1) State-of-the-art detectors suffer a substantial accuracy decline, reaching up to 11.43% when exposed to style-based adversarial rewrites generated by LLMs. (2) While general-purpose LLMs exhibit remarkable zero-shot capabilities, their performance in detecting adversarially manipulated text is significantly lower than specialized detectors fine-tuned for robustness. To address these vulnerabilities, we propose a novel style-agnostic detection framework named SAFD that enhances detection accuracy and robustness by prioritizing content-driven features over stylistic attributes. Our approach integrates a style-invariant training paradigm to disentangle content semantics from stylistic variations. We leverage adversarially enriched datasets constructed using LLMs fine-tuned for diverse style-based rewrites. Furthermore, we utilize advanced representation learning techniques to extract content-centric features, emphasizing semantic coherence, logical consistency, and factual alignment. Experimental results across multiple datasets and detection models validate the effectiveness of our framework, showing significant improvements in detection accuracy and robustness against diverse adversarial manipulations. The dataset and code are in the link https://anonymous.4open.science/status/A-Style-Agnostic-Framework-for-Detecting-LLM-Generated-Text-90B7.
Abstract: What does it mean to create a new concept, rather than retrieve a familiar one? Diffusion models sampled many times at the same prompt vary only within the narrow stylistic envelope set by a fixed prompt embedding. We propose an operational definition of creativity as the generation of concepts that are initially unfamiliar yet rapidly learnable to an adaptive observer, and formalize it as a bilevel optimization between a Creator that generates and an Appraiser that adapts: the Appraiser's improvement under a brief inner-loop adaptation provides the reward signal that the Creator maximizes. We call this framework CAMEL (Creator-Appraiser Meta-Learning), and instantiate it on MNIST with an autoencoder Appraiser and on natural images with a CLIP Appraiser based on a low-rank adapter over the text projection. CAMEL produces outputs that lie outside the basin reachable by vanilla diffusion sampling at any sample count, while remaining recognizable as the prompt's class to an off-the-shelf observer.
PaperID: 3718, Poster
Abstract: Computational design of ribonucleic acid (RNA) is essential for therapeutic discovery and synthetic biology. Designing RNAs that bind specified protein targets is bottlenecked by a scarcity of paired-complex 3D structural data. Existing approaches focus on importing structural priors from external biomolecular datasets, leaving where to localize the complex signal itself unresolved. We propose RiboC2F, a coarse-to-fine flow-matching framework that concentrates this signal first and foremost on the global binding pose. RiboC2F runs a two-stage cascade over a shared trunk, where a coarse stage commits the binding pose and a fine stage generates the full RNA design. The coarse pose enters as a conditioning signal rather than a fixed initialization, avoiding cross-stage error accumulation. On PRI30k, RiboC2F delivers the strongest docking among published baselines, with competitive intrinsic validity and multi-sample diversity. Probing further shows that the fine stage exploits the coarse pose's direction but is robust to pointwise coordinate corruption, confirming pose direction as the operative coarse-signal axis.
Authors:
Atindra Jha, Naomi Sagan, Keisuke Kamahori, Irmak Sivgin, Rohan Sanda, Steven Gao, Mark Horowitz, Luke Zettlemoyer, Olivia Hsu, Jure Leskovec, Stephanie Wang, Baris KasikciAbstract: We are entering a new era of composite model architectures that integrate diverse components such as vision encoders, language backbones, diffusion and flow heads, audio codecs, action generators, and world-model predictors. Such architectures underpin a broad class of multimodal models, including unified multimodal models, omni models, speech-language models, vision-language-action policies, and world models. However, existing model serving frameworks were built on narrow assumptions about model structure, making them ill-suited to accommodate this new architectural diversity. Here we present M, a universal serving system for efficient serving of composite AI models. M represents models as dataflow graphs, processing requests spanning diverse modalities and tasks as traversals over these graphs. The core insight is a modular abstraction that supports arbitrary composition of model components, flexible placement onto a physical cluster, and model-agnostic optimizations within a distributed runtime. We call this abstraction the Walk Graph and show how it can concisely capture composite models from a broad range of families. We instantiate M on representative models and find that it achieves, on average, 30% lower end-to-end latency than vLLM-Omni for text-to-image workloads on BAGEL, while delivering a lower real-time factor and higher throughput - by up to 15% - for text-to-speech workloads on Qwen3-Omni. M also outperforms the V-JEPA 2-AC rollout baseline for robotic planning by up to 12.5×. Thus, our work paves the road towards more efficient serving of complex models with minimal developer effort.
PaperID: 3720, Poster
Abstract: Vision-language pretraining has advanced multimodal medical AI, but its direct application to 3D computed tomography (CT) remains limited by a fundamental mismatch between dense volumetric anatomy and sparse diagnostic reports. Existing 3D CT vision-language methods typically rely on global image-report alignment or align anatomical regions with decomposed report descriptions. However, a CT scan captures rich, spatially distributed information across many organs, whereas a report summarizes only a small subset of clinically relevant findings. As a result, report-based alignment provides incomplete or weak supervision, leaving many local anatomical structures underrepresented. Motivated by this limitation, we introduce SOAR, a semantic organ-aware pretraining framework for 3D CT image understanding. The key innovation of SOAR lies in its integration of fine-grained, organ-level structural priors, enabling report-free organ-aware supervision for visual pretraining without relying on radiologist annotations or coarse global image-report matching. SOAR combines three complementary objectives: (i) organ-level masked reconstruction to learn localized anatomical context, (ii) organ-level vision-language alignment to associate organ-specific visual features with organ-name text embeddings, and (iii) organ-level supervision to preserve voxel-level structural detail. SOAR integrates easily with existing LLMs and is evaluated on five public/in-house benchmarks, improving multiple-choice VQA accuracy by 3.53% and disease screening/abnormality detection AUC by 4.25 points on average over all benchmarks/baseline methods. Code will be available.
PaperID: 3721, Poster
Abstract: We present a diagnostic study of reward guided reasoning in discrete diffusion language models (dLLMs). On in-distribution GSM8K with Dream-v0-Instruct-7B, deterministic segmental top-1 guidance using outcome-supervised intermediate-state reward models underperforms a simpler matched-compute baseline: independent best-of-N sampling plus a GSM8K-trained final-answer outcome reward model (ORM) reranker. The setup appears well suited for PRM guidance because dLLM intermediate states partially expose solution evidence; however, these states are unordered masked subsets rather than autoregressive prefixes, and prior reward-guided decoding results rarely report matched forward-pass compute against simple verifier reranking. We study \emphoutcome-supervised PRMs trained with binary final-correctness labels on dLLM intermediate states. Under a protocol that charges denoising, PRM scoring, and ORM scoring in the same forward-pass unit, deterministic PRM-guided search trails ORM Rerank by 9.95 and 12.69 percentage points (pp) in seed-mean aggregate at matched candidate budgets K=8 and K=32; paired bootstrap CIs on seed 42 exclude zero. Two mechanisms account for much of the headline gap: bidirectional PRM ROC-AUC (area under the receiver operating characteristic curve) decays from 0.77 to 0.54 with mask ratio, and deterministic top-1 pruning drops the guided pool's perfect-selector ceiling by ~14\,pp. A third diagnostic shows that mean pooling also weakens causal PRM variants. Matched-compute ORM Rerank thus emerges as the strong baseline when an in-distribution final-answer verifier is available; closing the gap requires PRM guidance that preserves candidate diversity and queries the scorer at denoising stages where it remains discriminative, neither of which the deterministic top-1 recipe satisfies.
PaperID: 3722, Poster
Abstract: Diffusion-based policies have demonstrated remarkable performance in online reinforcement learning (RL) by virtue of their expressive, multimodal action distributions. However, existing methods operate on a flat architecture that must generate a new action at every timestep, with no mechanism for temporal commitment to a behavioral mode—making them ill-suited for context-dependent tasks that require sustaining coherent behavioral modes, regardless of how expressive their action distributions are. We propose Hierarchical Generative Policy Optimization (H-GenPO), the first framework to integrate on-policy diffusion-based RL with the temporal abstraction of the Option-Critic architecture. H-GenPO adopts a two-level hierarchy in which a high-level policy over options selects among a discrete set of options, while a unified flow matching model—conditioned on a learned option embedding—serves as the expressive low-level intra-option primitive. The shared intra-option policy design encapsulates diverse behavioral repertoires within a single set of parameters, enabling scalable hierarchical control without the parameter growth associated with maintaining separate policy networks per option. We evaluate H-GenPO on 8 standard continuous control tasks in Isaac Lab and 3 custom context-dependent tasks designed to require coordinated behavioral mode switching. Our empirical evaluation demonstrates that H-GenPO achieves the best mean rank among all baselines on both standard and context-dependent benchmarks, while also exhibiting interpretable option specialization that emerges without any option-level supervision.
Abstract: Spatial and temporal resource constraints are critical for both biological and artificial intelligent systems. Here we define differentiable cost terms for breadth, depth, and time within a recurrent convolutional neural network conceived as a finite subset of an infinite lattice architecture. We optimize these costs jointly with task errors via backpropagation, and efficient computational graphs emerge organically through training. We find that breadth, depth, and time can be traded off against each other to achieve a given level of performance. Networks grow in all three dimensions with task complexity and spontaneously take more recurrent steps when inputs are occluded. Surprisingly, time used by the model correlates with human reaction times in an object recognition task. Our framework provides a normative account of how resource constraints shape neural architectures, connecting to questions about brain design in neuroscience, and may help illuminate the diversity of neural solutions found in nature.
PaperID: 3724, Poster
Authors:
Shuaizheng Liu, Fangzhou Han, Jiarong Liao, Yujing Sun, Lingchen Sun, Ruibin Li, Xinyu Wei, Jie Liang, Hui Zeng, Xindong Zhang, Lei ZhangAbstract: Professional image retouching is a sequential and reasoning-driven process. Photographers first analyze scene semantics, lighting structure, and depth cues to design a retouching plan, and then progressively apply global tone grading and mask-guided local corrections based on the evolving visual outcome. Recent MLLM-empowered retouching agents either predict global parameters or respond to user-provided retouching instructions. Although showing impressive results, they lack the capacity for autonomous sequential reasoning and fall short of fine-grained local control. We propose gent that emulates the workflow of expert photographers. Given an input photograph, AURA autonomously assesses the scene and proceeds progressively. At each step, it reasons about what to retouch, executes the adjustment, and feeds the result back as visual context for next-step reasoning, forming a fine-grained visual chain-of-thought that evolves with the image. To support AURA, we contribute the first expert-annotated long-horizon retouching trajectory dataset, where expert photographers edit the input image from scratch with deliberate artistic intent. Based on this dataset, we train AURA in two stages: supervised fine-tuning to learn the retouching trajectory, followed by GRPO-retouching, an agentic reinforcement learning stage that optimizes for perceptual quality and tool accuracy. Our experiments demonstrate that AURA consistently outperforms existing methods in both quantitative metrics and user studies, enabling one-click professional retouching and democratizing masterpiece previously accessible only to skilled experts.
PaperID: 3725, Poster
Abstract: Generative models for drug design which directly produce synthetic pathways have gained significant popularity due to their ability to constrain the search space to synthetically accessible molecules. However, existing methods have focused primarily on de novo molecular design, and rarely start the generation process from known binders. In this paper, we present HELiX: a template-based GFlowNet for localized exploration of chemical space. HELiX learns to partially decompose a given synthetic trajectory to an intermediate state, and then perform forward synthesis in a manner that preserves synthetic accessibility, leading to diverse, high-reward analog generation. Moreover, we prioritize sample efficiency by incorporating Bayesian optimization into the training procedure. We diagnose problems inherent to training GFlowNets with Bayesian optimization and introduce a greedy acquisition strategy which effectively balances between exploration and exploitation without the need for reward shaping. Finally, we show that local exploration is inherently robust to noisy oracle evaluations, a common problem in drug development when using unreliable proxies for binding affinity.
PaperID: 3726, Poster
Abstract: This work presents UGD-FM, a novel unaligned image guided denoising framework based on cross-modal conditional flow matching. Unlike previous approaches that assume spatially aligned inputs or only handle small-baseline rectified pairs, we address a more challenging setting where severe target noise and large-range two-dimensional cross-modal misalignment coexist. Specifically, we formulate guided denoising as a progressive flow matching process, allowing image restoration and aligned guidance aggregation to mutually reinforce each other. To support efficient few-step inference, we introduce a stage-focused and inference-consistent training strategy that focuses on task-critical denoising stages, and mitigates the training-inference mismatch. We further propose adaptive sparse guidance propagation for efficient and reliable large-range matching, regularized by content and geometry level auxiliary losses. Experiments demonstrate that UGD-FM achieves state-of-the-art denoising performance with only one or two inference steps, and is compatible with modern single-image restoration backbones.
Authors: Nan Jiang, Xiang Li
Abstract: We present a novel theoretical framework, Q-MMR, for off-policy evaluation in finite-horizon MDPs. Q-MMR learns a set of scalar weights, one for each data point, such that the reweighted rewards approximates the expected return under the target policy. The weights are learned inductively in a top-down manner via a moment matching objective against a value-function discriminator class. Notably, and perhaps surprisingly, a data-dependent finite-sample guarantee for general function approximation can be established under only the realizability of Q^\pi, with a dimension-free bound---that is, the error does not depend on the statistical complexity of the function class. We also establish connections to several existing methods, such as linear FQE. Further theoretical analyses shed new light on the nature of coverage, a concept of fundamental importance to offline RL.
PaperID: 3728, Poster
Authors:
Jiacheng Liu, Peiliang Cai, Qinming Zhou, Yuqi Lin, Deyang Kong, Benhao Huang, Yupei Pan, Haowen Xu, Chang Zou, Junshu Tang, Shikang Zheng, Linfeng ZhangAbstract: The application of diffusion transformers is suffering from their significant inference costs. Recently, feature caching has been proposed to solve this problem by reusing features from previous timesteps, thereby skipping computation in future timesteps. However, previous feature caching assumes that features in adjacent timesteps are similar or continuous, which does not always hold in all settings. To investigate this, this paper begins with an analysis from the frequency domain, which reveal that different frequency bands in the features of diffusion models exhibit different dynamics across timesteps. Concretely, low-frequency components, which decide the structure of images, exhibit higher similarity but poor continuity. In contrast, the high-frequency bands, which decode the details of images, show significant continuity but poor similarity. These interesting observations motivate us to propose Frequency-aware Caching (FreqCa) which directly reuses features of low-frequency components based on their similarity, while using a second-order Hermite interpolator to predict the volatile high-frequency ones based on its continuity. Besides, we further propose to cache Cumulative Residual Feature (CRF) instead of the features in all the layers, which reduces the memory footprint of feature caching by 99%. Extensive experiments on FLUX.1-dev, FLUX.1-Kontext-dev, Qwen-Image, and Qwen-Image-Edit demonstrate its effectiveness in both generation and editing. Codes are available in the supplementary materials and will be released on GitHub.
Abstract: Existing action-conditioned video generation models (video world models) are limited to single-agent perspectives, failing to capture the multi-agent interactions of real-world environments. We introduce Solaris, a multiplayer video world model that simulates consistent multi-view observations. We develop a multiplayer data system designed for robust, continuous, and automated data collection on video games such as Minecraft. Unlike prior platforms built for single-player settings, our system supports coordinated multi-agent interaction and synchronized videos + actions capture. Using this system, we collect 12.64 million multiplayer frames and propose an evaluation framework for multiplayer movement, grounding, building, and view consistency. We train Solaris using a staged pipeline that progressively transitions from single-player to multiplayer modeling, combining bidirectional, causal, and Self Forcing training. In the final stage, we introduce Checkpointed Self Forcing, a memory-efficient Self Forcing variant that enables a longer-horizon teacher. Results show our architecture and training design outperform existing baselines. We also demonstrate that scaling beyond two players is possible by training a 3-player proof-of-concept model and simulating up to 16 players on our hardware with efficient attention. Through open-sourcing our system and models, we hope to lay the groundwork for a new generation of multi-agent world models.
Abstract: The modeling and simulation of infinite-dimensional Hamiltonian systems are central problems in mathematical physics and engineering, but they pose significant computational and structural challenges for standard data-driven architectures. In this work, we introduce the Symplectic Neural Operator, a neural operator architecture designed to preserve the symplectic structure intrinsic to Hamiltonian PDEs. We provide a theoretical characterization of their symplecticity and establish a rigorous long-term stability result based on the combination of symplectic structure preservation and learning accuracy. Numerical experiments on canonical Hamiltonian PDEs corroborate this theoretical result and show that SNOs exhibit improved energy behavior compared with non-structure-preserving neural operators.
PaperID: 3731, Poster
Abstract: Large Language Models face a challenging multi-objective coupled optimization problem in constrained long-text generation, requiring the simultaneous balancing of global structure, local coherence, and constraint satisfaction. However, existing methods often rely on static pipelines or loosely coupled designs, making it difficult to achieve cross-stage collaborative optimization. To address this, we propose LoopWeaver, which models the generation process as a hierarchical, feedback-driven closed-loop optimization problem; by introducing iterative feedback between the planning and generation stages, it achieves the joint optimization of both global and local aspects. Specifically, LoopWeaver integrates constraint-aware hierarchical planning, feasibility filtering, and reward-based preference optimization, thereby allowing feedback signals to permeate both the planning and generation layers. Unlike methods that utilize feedback solely during post-processing or within a single module, LoopWeaver introduces a unified paradigm for feedback-coupled optimization, demonstrating significant improvements in text quality and constraint satisfaction capabilities across various models.
PaperID: 3732, Poster
Abstract: Image fingerprinting assigns each distributed copy a user-specific watermark for source tracing. Averaging collusion is a central threat: several recipients can average their differently marked copies to suppress each individual mark. Discrete bit-string watermarks are especially vulnerable because averaging drives bit evidence toward ambiguous decisions. We introduce OrthoMark, which maps user keys to near-orthogonal high-dimensional unit vectors and identifies colluders by cosine similarity to the extracted watermark direction. Under ideal averaging of K vector keys, the normalized average preserves a cosine signal of order 1/\sqrtK for each colluder, while unrelated keys remain concentrated near zero. The spherical geometry of random directions gives a shifted-Beta null distribution for cosine similarities, enabling analytic per-key false-positive-rate control after validation on unwatermarked images. OrthoMark uses a neural watermark encoder and centered extractor trained with JND masking and a progressive curriculum for photometric, geometric, and collusion robustness. Experiments on MS-COCO, OpenImages, and AI-generated images show that OrthoMark maintains robust single-key detection under photometric and geometric distortions, and achieves the best strict all-colluder detection among evaluated methods, while preserving high visual quality.
PaperID: 3733, Poster
Abstract: Temporal causal structure learning aims to recover lag-aware causal relations from multivariate time series. While prior knowledge has been widely used to improve static structure learning, its integration into lag-aware structure learning remains limited, because available priors are often lag-agnostic. Recent studies have explored lag-agnostic edge presence priors, but partial order priors, which can prune the ordering space and improve structure learning, remain underexplored. A key challenge is that variables may influence each other at different lags, leading to cyclic variable-level relations where ordering is no longer naturally defined. Moreover, directly applying static differentiable formulations of partial order constraints can be unstable due to repeated-walk accumulation on cycles. To address this gap, we represent partial order priors on the summary graph and follow the natural idea of the order relation that only allows a single direction of possible causality between two variables, which is equivalent to forbidding all reversed causality, direct or indirect. This converts partial order constraints into path prohibition constraints that can be extended to cyclic summary structures. For stable differentiable characterization, we propose a decoupled partial order constraint method. The method preserves the original lag-aware structure for data fitting, imposes partial order constraints through a normalized summary representation, and captures higher-order reachability with a closed-form connectivity characterization. Experiments further demonstrate its effectiveness in nonstationary settings.
Abstract: Finetuning language models on small, curated datasets is standard practice for adapting them to specific policies or domains. We show that finetuning on narrow, factually-defensible, moderation-passing data can cause broad ideological shifts across unrelated domains, while preserving general capabilities. Training GPT-4.1 on right- or left-leaning economics Q&A yields matched ideological shifts on topics such as criminal justice, the environment, and cultural taste. The same effect appears with plausibly-deployed datasets such as workplace HR policy and practical finance queries, as well as on a science-pseudoscience axis where food-safety finetuning increases sycophantic agreement with users expressing false health beliefs. We call this phenomenon ideological generalisation and propose a methodology to measure two properties: breadth, how far the shift reaches across topics absent from training, and amplification, how much finetuning intensifies the shift relative to few-shot prompting on the same examples. We show that few-shot prompting indicates the direction of generalisation but finetuning pushes the model to further extremes, including to far out-of-distribution outputs such as endorsements of race-IQ connections and political violence. The effect replicates on Gemma-3, holds under judge-free evaluations and external benchmarks, survives mixing with generic data, and leaves GSM8K accuracy within \pm 1pp of the baseline.
Abstract: The ability to reliably distinguish human-written text from that generated by large language models is of profound societal importance. The dominant approach to this problem exploits the likelihood hypothesis: that machine-generated text should appear more probable to a detector language model than human-written text. However, we demonstrate that the token-level signal distinguishing human and machine text is non-uniform across the hidden space of the detector model, and naively averaging likelihood-based token scores across regions with fundamentally different statistical structure, as most detectors do, causes a form of Simpson's paradox: a strong local signal is destroyed by inappropriate aggregation. To correct for this, we introduce a learned local calibration step grounded in Bayesian decision theory. Rather than aggregating raw token scores, we first learn lightweight predictors of the score distributions conditioned on position in hidden space, and aggregate calibrated log-likelihood ratios instead. This single intervention dramatically and consistently improves detection performance across all baseline detectors and all datasets we consider. For example, our calibrated variant of Fast-DetectGPT improves AUROC from 0.63 to 0.85 on GPT-5.4 text, and a locally-calibrated DMAP detector we introduce achieves state-of-the-art performance across the board. That said, our central contribution is not a new detector, but a precise diagnosis of a significant cause of under-performance of existing detectors and a principled, modular remedy compatible with any token-averaging pipeline. This will serve as a foundation for the community to build upon, with natural avenues including richer distributional models, improved calibration strategies, and principled ensembling with hidden-space geometry signals via the full Bayes-optimal decision rule.
PaperID: 3736, Poster
Authors: Rahul Goswami
Abstract: Many survival problems are inherently time-dependent: a biomarker may be predictive soon after diagnosis but irrelevant years later, or a treatment effect may diminish, reverse, or interact with follow-up duration. However, tree-based survival models commonly used in practice typically assign each individual a fixed, time-independent risk score. Neural non-proportional-hazards models can capture dynamic effects, but are often more challenging to tune, scale, and interpret than boosted trees. We present CCTimeBoost, a gradient-boosted tree framework for learning continuous-time, time-varying relative risk. The method augments covariates with time features, trains boosted trees on sampled risk sets using a group-softmax objective, and estimates a full-risk-set baseline hazard to generate survival curves. We show that this objective has a likelihood interpretation: for finite sampled risk sets, it matches a nested case-control likelihood, and in the full-risk-set limit, it converges to the Cox partial likelihood. Across eight right-censored survival benchmarks, CCTimeBoost delivers state-of-the-art discrimination performance relative to tree and neural baselines and is competitive in calibration. In controlled simulations, it accurately recovers diminishing effects, crossing hazards, and rule-based temporal interactions, also behaving as a proportional model when proportional hazards are present. These results indicate that non-proportional survival modeling can be incorporated into standard boosted-tree methods.
PaperID: 3737, Poster
Authors: Xueling Wang, Yiwen Wang, Siqi Cai, Chen Zhang, guanghui He
Abstract: Visual geometry grounded transformers are strong feed-forward backbones for multi-view 3D reconstruction, but their global attention becomes the dominant inference bottleneck as views and tokens increase. Existing acceleration methods mainly focus on designing compression operators, while compression strengths are often assigned by uniform schedules or offline calibration, overlooking the heterogeneous sensitivity of layers and attention heads during inference. We therefore propose HALO-VGGT, a lightweight online compression allocator requiring no training or calibration for efficient VGGT-style reconstruction. HALO uses a sampled query-key margin proxy to estimate compression sensitivity at the layer and head levels, and assigns hierarchical compression ratios accordingly. Instead of introducing a new compression operator, HALO complements pruning, merging, and sparse attention mechanisms by deciding where and how aggressively they should be applied. HALO improves the trade-off between speedup and accuracy across datasets and backbones, achieving up to 4.20× speedup and further improving existing merging and sparsity methods by up to 1.46×.
PaperID: 3738, Poster
Authors:
Po-Nien Kung, Linfeng Song, Dawsen Hwang, Jinsung Yoon, Chun-Liang Li, Simone Severini, Miroslav Olšák, Edward Lockhart, Quoc V Le, Burak Gokturk, Thang Luong, Tomas Pfister, Nanyun PengAbstract: Large Language Models (LLMs) exhibit strong informal mathematical reasoning but struggle to generate mechanically verifiable proofs in formal languages like Lean. We present L^2EAP (LLM-in-Lean Environment Agentic Prover), an agentic framework that enables general-purpose foundation models to achieve state-of-the-art performance on automated formal theorem proving. L^2EAP leverages foundation model capabilities, such as informal reasoning, instruction following, and iterative self-refinement. By decomposing complex problems into smaller units, the system bridges formal proof construction with informal blueprints through continuous interaction with the Lean compiler. To provide a rigorous evaluation beyond increasingly saturated benchmarks, we introduce Lean-IMO-Bench, a benchmark of IMO-style problems formalized in Lean, with short statements yet highly non-routine and multi-step proofs across a wide range of difficulty levels. Empirically, on the latest 2025 Putnam Competition, an annual mathematics competition for undergraduate students in North America, L^2EAP solves all 12 problems, matching recent breakthroughs by frontier formal mathematical models; on Lean-IMO-Bench, L^2EAP raises the one-shot formal solve rate of general-purpose LLMs from below 10% to 70%. It also compares favorably to the best baseline performance of 48% attained by a specialized system with dedicated ATP components that achieved gold-medal-level performance at the 2025 IMO.
PaperID: 3739, Poster
Abstract: We study a structured stochastic multi-armed bandit problem in which each decision simultaneously determines the learner's reward and the feedback available for learning. At each round, the learner selects a task and an option. Each task is associated with an unknown parameter that governs its reward, while each option has a known reward function of this parameter and determines how information about that reward is observed. Some options may provide fine-grained feedback at the cost of lower immediate reward, whereas others may yield higher reward but return only coarse feedback. This creates an intrinsic trade-off between reward maximization and information acquisition. We formalize this setting as parametric bandits with action-dependent coarsened feedback. For stationary tasks, we propose ACF-UCB, an optimistic algorithm that constructs task-level confidence sets by aggregating information from all atoms of the option-dependent feedback distributions. For piecewise-stationary tasks, we develop SW-ACF-UCB, a sliding-window variant that adapts to temporal changes. We prove high-probability regret bounds for stationary environments and dynamic regret guarantees for piecewise-stationary environments. We also establish lower bounds showing that the dependence on the number of tasks and the horizon is unavoidable, while methods that ignore the shared parametric structure suffer substantially larger regret.
Abstract: We investigate the approximation capabilities of dense neural networks. While universal approximation theorems establish that sufficiently large architectures can approximate arbitrary continuous functions when there are no restrictions on the weights, we show that dense neural networks do not possess this universality. Our argument is based on a model compression approach, combining the weak regularity lemma with an interpretation of feedforward networks as message passing graph neural networks. We consider ReLU neural networks subject to natural constraints on weights and input and output dimensions, which model a notion of dense connectivity. Within this setting, we demonstrate the existence of Lipschitz continuous functions that cannot be approximated by such networks. This highlights intrinsic limitations of neural networks with dense layers and motivates the use of sparse connectivity as a necessary ingredient for achieving true universality.
PaperID: 3741, Poster
Authors:
Baicheng Li, Jialin Liu, Dong Wu, Weijian Xie, Zhichao Ye, Donghui Shen, Nan Wang, Guofeng Zhang, Haomin Liu, Hongbin ZhaAbstract: Large visual world models have become increasingly capable at synthesizing plausible views and camera trajectories, largely because they inherit strong priors from image and video diffusion. Yet in robotics, spatial computing, AR anchoring, and digital twins, a world is useful only if it can be measured and acted upon: appearance plausibility cannot substitute for geometrically coherent structure. This motivates a geometry-native alternative, where the generative state is the feature space of a geometry foundation model rather than RGB pixels or an appearance latent. Existing geometry-space world models validate this direction, but they still occupy a narrow operating range. When target views lie within observed support, full multi-view interaction limits trajectory length; when observations become sparse, geometry-only training lacks the broad visual priors needed to infer structure beyond observed support. We present GeoWorld, a geometry-latent world model that expands this range without returning generation to RGB space. GeoWorld builds a compact geometry latent, conditions it with camera ray geometry, and uses camera-aware tiered attention with relative pose bias to make long-view denoising practical. For sparse-observation extrapolation, a Bridge module translates hidden states from a frozen camera-conditioned video diffusion model into geometry-latent conditioning, borrowing video priors without making the video model the scene generator. Experiments show that GeoWorld scales geometry-space diffusion to long camera trajectories while improving geometric consistency in prior-assisted extrapolation.
PaperID: 3742, Poster
Abstract: The uncertainty of a statistical model is most commonly factorized into an aleatoric and an epistemic part. This factorization changes how predictions are interpreted in downstream decision tasks. Importantly, except for the idealistic scenario with no model mismatch, the quantification is a characteristic of the model and not the data generating process. In this paper, we propose Forgetting to Improve a method that reduces this discrepancy by incorporating the task into the modeling framework. Our key insight is to acknowledge that in scenarios of model mismatch, data can have a detrimental effect on the modeling for a specific task. Based on this insight, we propose an influence function for Gaussian process models that allows for principled removal of detrimental data samples. We showcase the flexibility of this approach by demonstrating significant improvements across a range of tasks, including Bayesian optimization, model-based reinforcement learning, and transductive learning.
PaperID: 3743, Poster
Authors:
Mahdiyar Molahasani, Michael Greenspan, Ali EtemadAbstract: We theoretically and empirically establish a previously unexplored connection between Long-Tailed Recognition (LTR) and Continual Learning (CL). Specifically, we show that training on a long-tailed dataset drives model parameters into an \mathcalO(1/\sqrt\mathrmIF) neighborhood of the dominant-class solution, under both uniform and exponentially decaying cardinality distributions. Building on this result, we prove that the CL objective upper-bounds the balanced LTR loss, revealing that LTR can be reformulated as a sequential learning problem. Motivated by this connection, we introduce Continual Learning for Long-Tailed Recognition (CLTR), a principled framework that leverages standard off-the-shelf CL methods to sequentially learn Head and Tail classes while mitigating catastrophic forgetting. Extensive experiments on CIFAR100-LT, CIFAR10-LT, ImageNet-LT, and Caltech256 validate our theoretical predictions and demonstrate strong performance across diverse LTR benchmarks. In the foundation-model regime, CLTR with DualPrompt on a CLIP backbone outperforms standard adaptation baselines and is competitive with specialized methods using external semantic supervision. Our work bridges LTR and CL both theoretically and empirically, providing a principled approach for addressing long-tailed learning using standard CL strategies.
PaperID: 3744, Poster
Abstract: Graph-level anomaly detection (GLAD) is a critical task in domains like social networks and bioinformatics. However, current unsupervised contrastive learning-based methods rely on semantically homogeneous augmentations, leading to catastrophic informational overlap that trivializes the learning objective and effectively masks anomalous signals in GLAD. To overcome this, we propose Sima (Semantically Complementary Spectral Views), a novel contrastive framework that learns from semantically complementary spectral views rather than redundant homogeneous views. Sima first constructs a pair of spectral views to separately capture global community structure and local heterophilous variations, which together cover the full graph spectrum of anomalous patterns. Inspired by the Multi-View Information Bottleneck principle, we further design a soft subgraph extraction module with Score Reconstruction Attention (SRA) and an anchor-guided condensation mechanism to distill the essential shared information from each view into a compact representation. We then optimize a contrastive objective across these distilled views to enforce cross-view alignment and amplify subtle structural deviations. Extensive experiments on challenging GLAD and graph-level out-of-distribution detection (GLOD) benchmarks demonstrate that Sima achieves substantial and consistent gains in average performance over state-of-the-art baselines.
Abstract: Activation functions play a central role in neural networks by shaping internal representations. Recently, learning binary activation representations has attracted significant attention due to their advantages in computational and memory efficiency, as well as interpretability. However, training neural networks with Heaviside activations remains challenging, as their non-differentiability obstructs standard gradient-based optimization. In this paper, we propose Heavy-Tailed Activation Function (HTAF), a smooth approximation to the Heaviside function that enables stable training with gradient-based optimization. We construct HTAF as a sigmoid–hyperbolic tangent composite function and theoretically show that it maintains a large gradient mass around zero inputs while exhibiting slower gradient decay in the tail regions. We show that Spiking Neural Networks, Binary Neural Networks and Deep Heaviside neural Networks can be trained stably using HTAF with gradient-based optimization. Finally, we introduce Implicit Concept Bottleneck Models (ICBM), an interpretable image model that leverages HTAF to induce discrete feature representations. Extensive experiments across various architectures and image datasets demonstrate that ICBM enables stable discretization while achieving prediction performance comparable to or better than standard models.
PaperID: 3746, Poster
Abstract: Spiking reservoir computing combines the training efficiency of fixed random recurrent networks with the energy efficiency of event-driven inference, where computation is triggered by sparse spikes rather than executed densely at every timestep. However, existing spiking reservoirs almost exclusively rely on Leaky Integrate-and-Fire (LIF) neurons, whose low-pass dynamics limit their ability to distinguish information encoded at different temporal frequencies. We introduce HRF-Res, a liquid state machine in which LIF neurons are replaced by a heterogeneous population of Harmonic Resonate-and-Fire (HRF) neurons. These neurons exhibit intrinsic oscillatory dynamics and selectively respond to input frequencies, enabling the reservoir to capture frequency-specific temporal patterns. The result is more discriminative representations and higher classification accuracy than LIF-based reservoirs, achieved without sacrificing the energy efficiency of spiking computation, with 9–14× lower energy consumption than an equivalent non-spiking oscillatory reservoir. Evaluated on eight benchmarks spanning time series, sequential images, audio spike trains, and neuromorphic vision, HRF-Res achieves competitive or state-of-the-art performance among spiking reservoir methods. These results demonstrate that introducing frequency-selective dynamics into spiking reservoirs improves representational quality and classification performance while preserving their inherent energy efficiency.
PaperID: 3747, Poster
Abstract: Multi-agent language-model systems typically rely on a human designer or external orchestrator to fix the agent population, roles, and workflow—a single point of failure that places coordination decisions outside the team executing them. Existing approaches address fragments of this problem: communication standards lack coordination logic, automated workflow-search systems concentrate authority in a single planner, and decentralized peers lack explicit governance, handoff validation, and bounded repair. To close this gap, we introduce the Unified Resource-Grounded Coordination Protocol (URCP), an orchestrator-free protocol in which a fixed pool of black-box models and executable resources collectively decides task framing, role formation, workflow topology, assignment, checkpoint validation, and final completion through supermajority voting under a single decision rule Φq. A five-level repair hierarchy reopens only the affected part of the workflow on failure; under standard timing and communication assumptions, every run terminates with an explicit verdict. Empirically, with twelve heterogeneous models against eight baselines on MultiAgentBench, τ-bench, MemoryAgentBench, GAIA, and REALM-Bench, URCP exceeds the strongest baseline by +9.7 / +7.9 / +10.6 points on the three primary benchmarks, terminates deadlock-free in 100% of runs, and remains stable up to 20% agent failure. Trace analyses show that the gains come from voted role concentration on frontier models (73.9% of critical-path roles), frequent critic/verifier insertion (63.5% of accepted schemas with at least three roles), and bounded repair rather than from model strength alone.
PaperID: 3748, Poster
Abstract: We study mixed-variable optimization in scientific discovery, where the key challenge is to efficiently explore the search space under expensive experimental evaluations. Traditional Bayesian optimization (BO) methods typically rely on structural descriptors (e.g., molecular fingerprints) to measure similarity among categorical variables. However, such task-agnostic similarity often fails to reflect task-specific objectives, leading to biased generalization. To address this limitation, we propose TagBO (Task-aware Graph Bayesian Optimization), a method that integrates LLM-derived relational priors with descriptor-based structural information for mixed-variable Bayesian optimization. Rather than directly using LLM outputs as predictive scores, TagBO treats LLM outputs as a potentially noisy source of task-aware relational priors and extracts coarse ordinal relational signals from LLM reasoning to construct task-aware similarity structures for scientific optimization. TagBO then performs constrained Gaussian process Bayesian optimization over the resulting joint discrete-continuous representation space. Experiments on diverse chemistry and materials formulation tasks demonstrate that TagBO consistently improves sample efficiency under fixed evaluation budgets across LLMs of varying capabilities.
PaperID: 3749, Poster
Abstract: Knowledge distillation from large reasoning models (LRMs) often uses raw teacher trajectories as supervision. However, this practice creates a mismatch between the goal of distillation and the nature of teacher trajectories. Advanced LRMs typically improve their reasoning capabilities through reward-driven post-training, such as RLHF. This training optimizes them as strong solvers, rather than as teachers that produce reasoning processes smaller models can easily learn. Since smaller models have different "cognitive capacities" compared to their larger teachers, directly imitating raw teacher trajectories can sometimes be ineffective and often requires a substantial number of high-quality samples. To bridge this mismatch, we propose ), an RL-based framework that transforms solver trajectories into teaching trajectories aligned with the student’s cognition. Avoiding expensive resampling, our framework directly transforms existing trajectories. We optimize CARD with two reward signals: (1) a Distribution Alignment Reward that matches the student’s predictive distribution, and (2) a Logical Coherence Reward that reduces misleading failed reasoning steps. Comprehensive evaluations across multiple student models and settings demonstrate that CARD outperforms distillation from raw teacher trajectories in in-domain settings while achieving substantial gains in out-of-distribution performance.
Abstract: Learning from past experience requires forming abstractions that can be reused in future problems. Recent work on agentic-memory systems has explored a practical route to continuously improving LLM agents after deployment without parameter updates: a textual abstraction is derived from each history and stored in memory, then continuously updated with more interactions. Yet we show that such textual abstractions produced by today's LLMs are often faulty, even when derived from useful experiences. Over time, the memory becomes less useful and even harmful. More surprisingly, even when abstracting from ground-truth solutions, GPT-5.4 fails on 54% of a set of ARC-AGI problems it had previously solved without memory. We attribute this to the iterative update process: the same trajectories yield qualitatively different memories under different update schedules, while an episodic-only control over those trajectories remains competitive with the consolidators we test---pinning the failure on the consolidation step, not the underlying experience. In our ARC-AGI Stream environment designed to trace memory consolidation behavior, when agents are given additional actions, they tend to keep episodic memory and double the accuracy of their forced-consolidation counterparts; removing consolidation entirely (episodic management only) matches this auto mode. These results suggest that current LLMs should not recursively rewrite their own experience into stable long-term knowledge. Robust agent memory should treat raw episodes as first-class evidence and make abstraction selective, delayed, and explicitly gated rather than mandatory after every interaction.
PaperID: 3751, Poster
Abstract: We study Gaussian linear regression under coordinate-wise corruptions and missingness in high dimensions. A large body of work in statistics investigates estimation under different missingness and corruption mechanisms—ranging from missing completely at random to missing not at random—each leading to qualitatively different behaviors. Recent work on the problem considered a strong missing-not-at-random model, where an adversary can inspect the data and corrupt or erase an \eta-fraction of entries in each coordinate. A striking consequence of this model is an information-theoretic breakdown at \eta = \Theta(1/\sqrtd), beyond which non-trivial estimation is impossible even with infinite samples—an unusually small threshold compared to classical robust estimation settings. A natural question is whether this phenomenon is inherent or a consequence of adversarial control over corruption locations. To investigate this, we consider a more benign model in which corruption locations are chosen at random, and the adversary can affect the data only in a subset of these locations. We show that randomizing the locations fundamentally changes the answer. There is no sharp infinite-sample breakdown: non-trivial estimation is now possible for every \eta < 1. However, the price is sample complexity. We prove matching upper and lower bounds showing that the sample complexity for non-trivial estimation scales exponentially with \eta^2 d. Additionally, this sample complexity also includes a dependence on the signal-to-noise ratio \sigma^2 / \|\beta\|^2 which is another new phenomenon in this model.
PaperID: 3752, Poster
Authors: Gaurav Jain
Abstract: Fairness in machine learning typically relies on the assumption of clean training data. However, in high-stakes applications like healthcare and criminal justice, privacy mechanisms and proxy variables frequently introduce simultaneous noise into both sensitive attributes and target labels. While existing methods address label or attribute noise in isolation, they fail under this "double noise" regime, often inadvertently amplifying bias. In this paper, we propose Jointly Robust Fairness (JRF), a unified framework that guarantees fairness under simultaneous data corruption. JRF integrates a Forward Loss correction to handle label noise with a novel Matrix-Inverse Robust Loss (MIRL) that algebraically recovers the true demographic parity gap from noisy attribute observations. We provide rigorous theoretical guarantees for our estimator, including a finite-sample concentration bound demonstrating that the sample complexity scales quadratically with the inverse of the noise transition determinant. Furthermore, our framework naturally generalizes to both symmetric and highly asymmetric noise distributions. Extensive experiments on the Adult, Bank, and COMPAS datasets demonstrate that JRF reduces fairness violations by over 91% in high-noise regimes (\rho \ge 0.3) compared to standard robust baselines, maintaining strict fairness without significant degradation in predictive utility.
PaperID: 3753, Poster
Abstract: Federated learning (FL) has become a foundational paradigm for multi-institutional medical AI, allowing hospitals and research centers to jointly train diagnostic models without exchanging patient records. This privacy promise, however, is increasingly contested: a malicious or honest-but-curious server can launch model inversion attacks (MIAs) that reconstruct private patient images directly from shared model updates, and recent scalable, closed-form attacks penetrate even secure aggregation at clinically realistic batch sizes. Existing defenses face an unsatisfactory dilemma. Gradient-perturbation methods such as differential privacy and pruning trade away the diagnostic accuracy on which clinical reliability depends, while cryptographic protocols add system complexity yet still leave updates exposed to these scalable attacks. We propose Aegis, a principled client-side defense that breaks this dilemma without perturbing patient data or modifying the FL protocol. Our key insight is that the success of every known MIA is fundamentally bounded by the local batch size relative to the model's leakage capacity; once this limit is exceeded, distinct samples collide and reconstructions collapse into indistinguishable mixtures. Aegis turns this universal bottleneck into a defense: each client superimposes onto its real update a masking gradient computed on locally synthesized, task-relevant data, deliberately pushing the effective batch beyond the attack's recovery capacity. We complement the design with theoretical convergence guarantees under standard convex assumptions and evaluate Aegis on MNIST, CIFAR-10, and three MedMNIST modalities (chest X-ray, abdominal CT, colon pathology). Aegis neutralizes three state-of-the-art MIAs while preserving model utility and incurring only modest overhead, offering a practical privacy primitive for medical FL.
Abstract: A local specialist LLM, fine-tuned with reinforcement learning from verifiable rewards (RLVR) on operator-local data, is installed inside a single regulated organization under a per-deployment error budget \alpha. The operator needs a safety certificate that holds on this deployment's stream, simultaneously at every wall-clock round: no pooling across deployments, no waiting for a long-run average. Existing wrappers cannot deliver this on these adaptive, online-updated streams: offline conformal-risk methods require exchangeability; online-conformal methods bound only long-run averages; non-exchangeable extensions are marginally valid; and the closest anytime wrapper, A-RCPS, controls marginal rather than selective risk. Through a (test statistic, validity guarantee, deployment rule) framework we identify one empty cell forced by the deployment requirements (e-process per threshold, selective risk, anytime-pathwise validity, max-certified-threshold rule); Conformal Selective Acting (CSA) fills it as a per-round wrapper maintaining a Ville-type e-process per candidate score threshold on a Bonferroni grid, evaluated against the RLVR filtration. Under predictable updates and isotonic-calibrated monotone risk we prove (i) an anytime-pathwise selective-risk bound R_T^\mathrmact \le \alpha + O(N_T^-1/2), (ii) rate-optimal certification matching \Theta(\bar\eta^-2 \log(1/\delta)), and (iii) a horizon-independent release-rate gap. Across eight specialist benchmarks (480 streams), sixteen adversarial distribution-shift cells (160 streams), and five live Expert-Iteration RLVR cells with online LoRA over four base models in three architecture families (10,300 rounds), CSA is the only method among ten directly compared that satisfies both pathwise validity and non-refusing deployment on every cell. We do not propose a new LLM, training algorithm, or policy class; CSA is the deployment-side complement, orthogonal to the model itself, for operators who cannot use a frontier API.
PaperID: 3755, Poster
Abstract: Data augmentation improves visual recognition by exposing models to synthetic training examples that encourage generalization beyond the observed data. Mixed-sample methods extend this idea by combining multiple images. However, existing approaches either mix only pairs of images or rely on fixed small layouts, limiting control over source count, visible area proportions, and induced supervision. We propose TreemapMix, a distribution-controlled multi-image augmentation that samples source weights from a Dirichlet distribution and assigns them to area-proportional regions using a treemap partition. The resulting known region areas support two complementary objectives. TreemapMix with soft cross-entropy (SCE) treats areas as soft class-probability targets, while TreemapMix with a Plackett-Luce (PL) objective converts areas into ordinal supervision. On ImageNet-1K under a matched training protocol, TreemapMix-SCE achieves substantially lower calibration error than evaluated mixed-sample baselines, while TreemapMix-PL achieves the strongest Top-1 accuracy. Transfer experiments on object detection and instance segmentation suggest that TreemapMix-pretrained features remain useful beyond classification. These results show that controlled multi-image composition can expose the same region-area information as either probability-matching or rank-matching supervision.
Authors: Reza T Batley, Sourav Saha
Abstract: Modern language models use a single matrix for input embedding and output projection. This couples two distinct objectives: token representation and discrimination over a vocabulary. This work introduces Leviathan, a Transformer architecture that replaces the input embedding matrix with learned embedding vectorization (LEV), a compact continuous mapping from token indices to embeddings. Leviathan's output head remains untied for a parameter increase of as low as 0.2%. Under controlled comparisons with identical Transformer backbones, Leviathan consistently improves language modeling performance over standard tied-embedding baselines across a 200M-1.2B parameter regime on The Pile with gains that grow during training. At 1.2B scale, Leviathan reduces validation perplexity by 9%, requires 2.1× fewer training tokens to reach the tied baseline's final loss, and improves on all six downstream benchmarks evaluated, including a 30% reduction in LAMBADA perplexity. Frequency-stratified analysis reveals gains to be concentrated in rare tokens, where continuous parameterization reduces perplexity by 81%, falling to near zero for the most frequent.
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising approach to enhance Instruction Following (IF) capabilities of large language models (LLMs). However, RLVR for instruction following remains sample-inefficient and prone to reward hacking, where LLMs exploit verification shortcuts rather than fulfilling the core intent. To address these challenges, we frame RLVR for instruction following as an integrated environment that unifies dynamic task generation and robust training. We introduce Instruction Following Decorator (IFDecorator, a framework coupling difficulty adaptation, intent alignment, and hack diagnostics. It features: (1) a cooperative-adversarial data flywheel that yields challenging yet solvable tasks for sample-efficient training; (2) IntentCheck, a gating module to mitigate reward hacking; and (3) TripWires, a proactive diagnostic tool for eliciting hacking behaviors and quantifying their prevalence. Extensive experiments show our Qwen2.5-32B-Instruct trained with IFDecorator using only 3,625 training examples achieves 87.43% on IFEval (outperforming GPT-4o) and improves FollowBench by 4.2%, while preserving general capabilities. Diagnostics confirm that our method effectively reduces reward hacking. The approach generalizes across model architectures and scales. We will release code and data for future research.
Authors: Yewon Jeong, Nayoung Jung, Hyeri Roh, Woo-Seok Choi
Abstract: Hybrid HE/2PC private CNN inference remains bottlenecked by prime-modulus homomorphic arithmetic in convolution and by a precision flow that runs ReLU at doubled bitwidth before invoking a separate truncation protocol. We present Jaguar, a system built on a single design choice---a power-of-two ciphertext ring---that addresses both. The choice enables SPA-Conv, a coefficient-domain convolution kernel that replaces NTT-centric polynomial multiplication with scalar--polynomial accumulation, and an exact ciphertext-side truncation by local right shifts that lets ReLU run directly at the target fixed-point precision and eliminates the post-ReLU truncation protocol. Where NTT remains genuinely useful---at the client, for the single polynomial multiplication during decryption---we recover it through an auxiliary NTT prime, preserving the power-of-two protocol substrate while keeping decryption O(N\log N). On ImageNet-scale ResNet-18, ResNet-50, and MobileNetV2 with AVX disabled, Jaguar achieves 2.07--3.72× lower end-to-end latency than Cheetah and 2.16--3.36× lower than Rhombus, with 1.16--1.76× lower communication than Cheetah.
PaperID: 3759, Poster
Abstract: Reinforcement learning with verifiable rewards (RLVR) enables large language models (LLMs) to continuously self-improve their reasoning capabilities. Compared to RLVR's scalar rewards, recent self-supervised on-policy distillation methods leverage denser supervision signals. However, both paradigms suffer from exploration bottlenecks, struggling to escape local optima when the policy fails to sample valid solutions. Relying on external demonstrations can bypass this issue; however, the off-policy nature hinders genuine self-improvement and the mismatched-distribution reasoning can not be fully exploited. In this work, we propose Self-Evolution Policy Optimization (SEPO), a framework that drives continuous reasoning capability evolution via guided exploration and in-distribution exploitation. Specifically, SEPO utilizes labels to elicit correct reasoning from the policy, unlocking the exploration potential. By treating current policy conditioned on this in-distribution reasoning as a self-teacher, SEPO distills the newly discovered solution back into the policy. Through this iterative cycle of policy-intrinsic exploration and exploitation, SEPO achieves robust and continuous capability evolution. Across scientific, tool-use, and complex logical reasoning domains, SEPO consistently achieves the best performance under identical training time or steps, demonstrating superior efficiency and convergence ceilings. Notably, SEPO maintains stable improvements even in scenarios where baselines severely underperform or fail entirely, such as on complex tasks or with small-scale models.
PaperID: 3760, Poster
Abstract: Recent works have demonstrated that transformers can be trained to recover sparse, binary cryptographic secrets in the Learning With Errors (LWE) problem, a foundational problem that underlies many post-quantum cryptographic schemes. However, as architectures have evolved to efficient encoder-only models, the mechanism by which these models recover the cryptographic secret has become more opaque. In this paper, we present the first layer-wise and embedding-level mechanistic interpretability analysis of encoder-only transformers trained on LWE samples. We reveal a surprising phenomenon: despite achieving near-zero exact prediction accuracy on the training objective, the models successfully recover the secret by bypassing the standard predictive pathways. We use dimensionality reduction, causal intervention, and linear probing and find that the secret is implicitly present in the positional embedding. Building on this mechanistic understanding, we introduce an architectural intervention that applies L_1 sparsity regularization directly to the positional embeddings. This modification forces the model to explicitly isolate the latent secret, transforming the computationally expensive post-hoc secret recovery process into a direct, human-interpretable parameter inspection. Our findings provide fundamental insights into how transformers allocate representational capacity when faced with high-noise, structured combinatorial problems.
PaperID: 3761, Poster
Abstract: Multi-agent LLM evaluations and large-scale social simulations often treat multiple copies of the same model as comparable replicates after they interact. We test this assumption with SocCogEval, a controlled-clone assay for Social Cognitive Transfer (SCT), and find that it fails. Agents share the same base model, system prompt, and decoding parameters. When briefs are assigned, the brief is the only initial difference. On Qwen3-122B-FP8, the matched interaction contrast holds asymmetric briefs fixed and toggles 30 rounds of interaction. It yields stable off-topic BFI-10 response-profile drift on two topically disjoint scenarios. The effect is Delta =+0.92 on Greenfield (95% CI [+0.61,+1.22], g =+1.07) and Delta =+ 0.91 on Riverton ([+0.60,+1.24], g =+ 0.96, about 0.18 Likert points per BFI factor. Brief-only controls do not explain the drift. A verbatim-history ablation with SRM extraction disabled remains positive, so typed memory extraction is not the source of the signal. Five additional model variants provide positive but uneven breadth. Five of ten non-primary scenario cells are positive at alpha = 0.05, and no variant shows a significant reversal. The result limits iid-clone assumptions in multi-agent evaluation: interacting clones are not independent replicates unless paired no-interaction controls show otherwise.
PaperID: 3762, Poster
Abstract: Cellular perturbation experiments are essential for probing biological mechanisms and guiding therapeutic discovery. While existing methods often treat perturbations as monolithic distributional shifts, they leave the underlying structural variations implicit. Rather than attempting to infer independent global networks for every condition, we propose scPertCRL, a causally structured generative framework for single-cell perturbation prediction. scPertCRL decomposes perturbation effects into a shared latent structural backbone, condition-specific structural modulations, and exogenous shifts. By leveraging biologically informed embeddings and mechanism-aware supervision, Differential Network Learner parameterizes these localized structural changes, enabling robust generalization to unseen interventions. Extensive experiments on genetic and pharmacological benchmarks demonstrate that scPertCRL significantly improves predictive accuracy and out-of-distribution generalization, while yielding model-based insights into perturbation mechanisms.
PaperID: 3763, Poster
Authors:
Terry J Zhang, Oscar S Yasunaga, Wenyuan Jiang, Jessica Bo, Florent Draye, Aydin Javadov, Arth Singh, Lucille Muir, VICTORIA OLDEMBURGO DE MELLO, Bernhard Schölkopf, Zhijing JinAbstract: AI sycophancy is gaining increasing prevalence as models optimized to maximize user satisfaction may tend to unconditionally agree with users even at the expense of factuality, which poses great risk as AI are increasingly used for decision support in high-stake scenarios. We present \ourWork, a systematic stress-testing framework encompassing objective (OBJ) questions with verifiable answers and subjective (SUB) scenarios without ground truth on sycophancy induced from first-turn context framing to multi-turn user pressure. Evaluation of 9 frontier models revealed high sycophancy in even the best-performing models such as GPT-5.4 changes answers from right to wrong to please users in 23.8% of OBJ scenarios and blindly follows the pressured user view in more than half of the SUB scenarios. We found that strength of user tone is one of the most impactful yet previously overlooked factor for AI sycophancy and discover an interesting pattern where Claude models are uniquely more sycophantic under moderate pressure because strong triggers often lead them to rethink questions from scratch to arrive at the correct impartial answer. These findings highlight the need for anti-sycophancy training in future model development, where training models to rethink from scratch when facing pressure may serve as a promising paradigm towards more truthful AI systems. We provide our code and data in the supplementary material.
Abstract: A robust world model must strike the balance between faithfully capturing environmental dynamics and abstracting away from irrelevant content. While reconstruction-based world models ensure faithful supervision, they misallocate representational capacity by pixel area rather than dynamics relevance, which can cause task-irrelevant content to dominate the learned representation. Alternatively, reconstruction-free methods avoid this bias but risk discarding possibly relevant information. We propose StarWM, which combines advantages of both paradigms through a cross-attention module gated by self-supervised dynamics signals that decides where reconstruction applies, while a dual-stream decoder lets reconstruction shape only what is extracted at attended locations, with stop-gradient barriers preventing the two objectives from interfering. This allows reconstruction to supervise the visual content of attended regions without contaminating the latent with non-predictive information. On DeepMind Control with dynamic video backgrounds, StarWM outperforms reconstruction-based baselines while keeping competitive with the concurrent reconstruction-free approach, and is the only method that performs consistently well across clean, sequential-video, and random-video distractor settings. Mechanistic probing confirms StarWM preserves state attributes with near-perfect fidelity through long-horizon imagination while systematically discarding distractors.
PaperID: 3765, Poster
Abstract: Recent advances in latent diffusion models have enabled high-quality image generation, but also raise critical concerns for intellectual property protection in model distribution scenarios, where downstream users have unrestricted access to models, allowing arbitrary modifications. Existing watermarking methods either rely on inference-time control of inputs (e.g., specific prompts or noise initialization) or embed watermark signals in auxiliary components, making them easily removable in such settings. In this paper, we propose INMARK, which treats watermarking as an intrinsic property learned by the model. Instead of relying on input control or auxiliary watermarking components, INMARK enables the core denoising network to internalize and reproduce watermark patterns. Extensive experiments demonstrate that INMARK achieves strong generation fidelity, high watermark detectability, and robustness against attacks, while remaining fully compatible with standard diffusion training pipelines. Our results highlight a new perspective on diffusion model watermarking: the denoising network can learn a reliable and persistent watermarking capability, which is crucial in practical model distribution scenarios.
Abstract: Recovering a signal from its degraded measurements is a long standing challenge in science and engineering. Recently, zero-shot diffusion based methods have been proposed for such inverse problems, offering a posterior sampling based solution that leverages prior knowledge. Such algorithms incorporate the observations through inference, often leaning on manual tuning and heuristics. In this work we propose a rigorous analysis of these approximate posterior samplers, relying on a Gaussianity assumption of the prior. Under this regime, we show that both the ideal posterior sampler and diffusion-based reconstruction algorithms can be expressed in closed-form, enabling their thorough analysis and comparisons in the spectral domain. Building on these representations, we introduce a principled framework for parameter design, replacing heuristic selection strategies used to date. The proposed approach is method-agnostic and yields tailored parameter choices that jointly account for the characteristics of the prior, the degraded signal, and the diffusion dynamics. We show that our spectral recommendations differ structurally from standard heuristics and vary with the diffusion step size, resulting in a consistent balance between perceptual quality and signal fidelity.
PaperID: 3767, Poster
Abstract: Differentiable signal parameterizations such as implicit neural representations (INRs) and hybrid models are increasingly central to computational imaging, yet principled tools for evaluating reconstruction fidelity at finite model size remain limited when ground truth is unavailable. We introduce a framework for predicting the reconstruction error of compressive signal parameterizations, yielding non-asymptotic, signal-specific bounds that are both theoretically sound and efficiently computable without access to the ground truth signal. Specifically, we prove that when parameterization-based compression satisfies certain natural properties, the compression error at any compression level is bounded by a simple scaled difference between model predictions at different compression levels. We verify these properties for representative model families including interpolated grids, Fourier feature networks, multi-resolution hash tables, and tensor factorizations, and show empirically that the resulting worst-case guarantees can be efficiently adapted into signal-specific error predictors that are tight and generalizable. Across direct fitting of synthetic and natural signals, and inverse problems including radiance field and MRI reconstruction, our method closely tracks global error curves and yields informative local error heatmaps without ground-truth access.
PaperID: 3768, Poster
Abstract: Feed-forward 3D Gaussian Splatting has recently emerged as a promising alternative to per-scene optimization, enabling fast novel view synthesis by predicting Gaussian primitives directly from input images. However, it quickly faces a bottleneck of rendering efficiency while scaling from sparse to dense multi-view inputs, which offer more geometric constraints but produce a rapidly growing set of candidate Gaussians. Existing solutions cannot reliably distinguish redundant primitives, leading to many redundant primitives from repeatedly observed regions and increased rendering cost without much gain in visual quality. We propose BudSplat, a compact feed-forward 3DGS framework that learns to keep the most valuable Gaussians under a limited primitive count. BudSplat has two novel designs. The first is a detail-aware importance score that is extracted from multi-level visual tokens, providing an explicit cue for regions where pruning is likely to harm perceptual quality. The score can be injected into voxel fusion to preserve high-value local contributors during multi-view aggregation. The second is quality-aware Gaussian pruning that selects a compact set of primitives by jointly considering rendering contribution and detail importance. The quality-aware pruning enables local quota redistribution that helps avoid over-allocating primitives to large redundant regions. Experiments on RE10K, DL3DV, and ACID show that BudSplat achieves a favorable trade-off among rendering quality and efficiency.
PaperID: 3769, Poster
Abstract: Test-time adaptation updates a deployed model using only unlabeled target data, but standard entropy-based methods often rely on deterministic confidence and can reinforce overconfident errors under distribution shift. We propose MIRA, Mutual Information guided calibration for Reliable Test-Time Adaptation, a source-free online adaptation method that uses lightweight last-layer Gaussian perturbations to probe prediction reliability. For each target batch, MIRA reuses a single feature extraction pass and samples perturbed classifiers around the source classifier, forming an efficient local stochastic ensemble. Rather than fitting a Bayesian posterior, MIRA treats the perturbation scale as a local probing radius and selects a batch-specific scale using mutual information and prediction agreement. The resulting perturbation-sensitivity signal identifies predictions that are confident and locally stable under classifier perturbations, while suppressing brittle predictions with high stochastic disagreement. MIRA then performs adaptation with an MI-confidence weighted entropy loss, which down-weights perturbation-sensitive or low-confidence samples before they drive entropy minimization. This weighted loss reduces the influence of brittle predictions that can otherwise cause harmful adaptation under distribution shift. The method requires no source data, retraining, model ensembles, or offline posterior fitting, and can be applied as a plug-and-play modification to entropy-minimization TTA. Experiments on distribution-shift benchmarks show improvements in accuracy, calibration, and negative log-likelihood, suggesting that mutual-information-guided perturbation probing and weighted adaptation yield more robust online TTA than deterministic confidence.
PaperID: 3770, Poster
Abstract: Recent work proposes hypernetworks that convert a document into a low-rank parameter update (LoRA) in a single forward pass, enabling fast injection of new knowledge into a frozen LLM. However, storing separate LoRA adapters for each document can be unsustainable in streaming settings. In this paper, we study the problem of composing hypernetwork-generated LoRAs using two representative systems, Doc-to-LoRA and SHINE. We show that hypernetwork-generated LoRAs have a geometry sharply different from directly fine-tuned LoRAs. Hypernetwork-generated LoRAs are highly spectrally concentrated and strongly aligned, while next-token prediction and QA fine-tuned adapters are diffuse and nearly orthogonal across documents. This suggests that hypernetwork-generated LoRAs are not arbitrary low-rank weight updates, but decoded points on a structured adapter manifold. Guided by this observation, we propose latent-space LoRA merging: instead of merging in the weight space, we compose the intermediate hypernetwork latents and then decode the merged latent through the frozen LoRA-generation head. We show that latent-space merging techniques consistently improve over weight-space and factor-space baselines while preserving a single fixed-rank adapter. These results establish latent-space composition as a promising mechanism for bounded accumulation of hypernetwork-generated LoRAs in long document sequences.
Authors:
Xiaofan Li, Ming Yang, Zhiyuan Ma, Shichao Ma, Jintao Du, Yu Cheng, Weiqiang Wang, zhizhong zhang, Xin Tan, Yanyun Qu, Lizhuang Ma, Yuan XieAbstract: Reinforcement Learning with Verifiable Rewards (RLVR) has catalyzed significant advances in the reasoning capabilities of Large Language Models (LLMs). However, effectively managing the exploration and exploitation trade-off remains a critical challenge. In this paper, we fully analyze the exploration and exploitation dilemma of extremely hard and easy samples during the training and propose a new fine-grained trade-off mechanism. Concretely, we introduce a perplexity space disentangling strategy that divides the sample space into distinct exploration (high perplexity) and exploitation (low perplexity) subspaces, thereby mining fine-grained samples requiring exploration-exploitation trade-off. Subsequently, we propose a bidirectional reward allocation mechanism with a minimum impact on verification rewards to implement perplexity-guided exploration and exploitation, enabling more stable policy optimization. Finally, we have evaluated our method on two mainstream tasks: mathematical reasoning and function calling, and experimental results demonstrate the superiority of the proposed method, confirming its effectiveness in enhancing LLM performance by fine-grained exploration-exploitation trade-off.
PaperID: 3772, Poster
Abstract: Material synthesis planning (MSP) remains a fundamental and underexplored bottleneck in AI-driven materials discovery, as it requires not only identifying suitable precursor materials but also designing coherent sequences of synthesis operations to realize a target material. Although several AI-based approaches have been proposed to address isolated subtasks of MSP, a methodology for solving the entire MSP task has yet to be established. We propose MSP-LLM, a structured LLM-based framework that formulates MSP as a two-stage process composed of two constituent subproblems: precursor prediction (PP) and synthesis operation prediction (SOP). Our approach introduces a discrete material class as an intermediate decision variable that organizes both tasks into a chemically consistent decision chain. For SOP, we further incorporate hierarchical precursor types as synthesis-relevant inductive biases and employ precursor constraint factorization that preserves precursor-related information in the autoregressive decoding state. Extensive experiments show that MSP-LLM consistently outperforms existing methods on both PP and SOP, as well as on the MSP task, demonstrating an effective and scalable framework with practical potential for autonomous MSP.
Abstract: Decision trees are prized for their interpretability and strong performance on tabular data, but popular greedy top-down induction algorithms can yield suboptimal and unnecessarily complex structures. Optimal decision tree methods address this through global optimization, yet remain restricted to axis-aligned threshold splits, which limit the expressivity of each node and often force deep, complex trees to capture non-linear feature effects. Shape Generalized Trees (SGTs) generalize threshold splits to learnable univariate shape functions, improving expressivity and enabling more compact trees. However, existing SGT induction algorithms are greedy and offer no optimality guarantees. In this work, we introduce Literati, the first algorithm for optimal SGT induction. We propose a novel AND/OR graph formulation of the problem that jointly optimizes tree structure and shape function complexity. To solve this AND/OR graph, we develop an AO-based algorithm with two enhancements that improve anytime performance while preserving optimality: a secondary heuristic for OR-node selection and a round-robin policy for AND-node exploration. Across 24 real-world datasets, Literati achieves higher training and test accuracy than state-of-the-art tree approaches.
Abstract: Concurrent multi-agent coding promises division of labor across modules, robustness through redundancy, and parallel exploration at the natural granularity of multi-file projects. Realtime collaborative editing protocols solve this coordination problem for human teams via Conflict-free Replicated Data Types (CRDTs), but the LLMs underneath generate one token at a time and existing multi-agent coding systems inherit this serial limit: they either sequence agents through phase handoffs or pool independent samples without coordination, and a single agent abandons up to 70% of hard tasks with a one-file stub-and-exit. AgentRoom is a realtime collaborative editing protocol for concurrent coding agents: a runtime layer exposing file-level claim, status, and broadcast as MCP tools on top of a CRDT-merged shared filesystem. Across five frontier coding-CLI models on four backend coding tasks, alongside Python DevBench and Rust+axum cross-language checks, AgentRoom ×2 suppresses Solo abandonment on the CLI-stable models and tightens run-to-run variance. Two matched-compute contrasts isolate the cause: parallel-merge vs AgentRoom shows a positive mean LLM-judge contrast, and a bundle probe attributes the gain primarily to the MCP coordination layer rather than to parallelism or substrate convergence alone. Coordination, not parallelism or CRDT-merge, is the load-bearing engineering.
PaperID: 3775, Poster
Authors: Amey P Pasarkar, Adji Bousso Dieng
Abstract: Anomaly detection (AD) is a critical task in machine learning and scientific discovery. Existing methods that operate on learned embeddings typically define anomalies through local density, isolation heuristics, or distance from a fitted distribution, making them sensitive to neighborhood scale, partitioning choices, or restrictive distributional assumptions. In this work, we introduce a new paradigm by formulating anomaly detection in terms of dataset diversity. We propose the Vendi Anomaly Score (VAS), which detects anomalies by quantifying how much a sample changes the diversity of the dataset, as measured by the Vendi Score, when removed. VAS captures both local redundancy and global data structure without relying on density estimation or partitioning heuristics. Moreover, VAS requires no hyperparameter tuning, instead adapting automatically to the spectral structure of the data. VAS is non-parametric and scales linearly with dataset size. Across 148 benchmark embedding datasets derived from 10 source datasets and 4 embedding architectures, VAS achieves state-of-the-art AD performance. We further validate VAS on large-scale ImageNet experiments, where it remains robust as both contamination rate and dataset scale increase.
Authors:
Ignacio D Lopez-Miguel, Ezio Bartocci, Thomas Eiter, Martin TapplerAbstract: Explainability remains a key issue in reinforcement learning (RL). Distilling an interpretable policy from an agent trained in a complex environment is particularly challenging when the action space is continuous. We introduce ORCAID, a novel method for extracting interpretable rule-based policies from RL agents operating in mixed continuous-discrete environments with continuous action spaces. Our main contribution is an efficient oblique decision tree training algorithm that partitions the state space by hyperplanes and fits local linear models. The key idea lies in a three-stage split search: efficient random initialization, local refinement, and backward elimination. Finally, adjacent leaves are merged to yield a concise set of interpretable rules describing a given deep RL policy. We evaluate ORCAID on multiple RL environments, demonstrating that the extracted rule-based policies maintain strong performance with a low number of parameters and can even be used to improve the performance of the original deep RL policy.
Authors: Thomas Massena, Corentin Friedrich, Mathieu Serrurier
Abstract: Modern optimizers, like Muon, impose matrix-wise geometry constraints on their updates. These matrix-wise constraints can be unified under Linear Minimization Oracle (LMO) theory. However, all current methods impose fixed LMO geometries for the update rules, chosen by-design or empirically, which are not necessarily optimal according to the problem's geometry. We introduce a novel efficient data-driven criterion for dynamically choosing proxy-optimal update LMO geometries on individual Deep Neural Network layers. Derived in closed form from gradient and activation statistics using a single-step random feature regression surrogate model, our criterion navigates a design space interpolating from SGD to Muon updates. Moreover, integrating parameter-wise preconditioning allows our framework to recover SGD, Muon, Adam, and MuAdam as specific extrema. To make this adaptive approach scalable, we pair it with efficient computational strategies, achieving only a ~ 3% runtime overhead on highly optimized baselines. As a proof of concept, we show that this data-driven optimizer matches or exceeds the performance of the best performing optimizer between Muon and AdamW across three different training scenarios. Ultimately, this work provides evidence that LMO geometry can be successfully and efficiently adapted from runtime data, opening a new pathway for optimizer design beyond static geometries.
PaperID: 3778, Poster
Authors: Xiaokai Zhang, Ruiqing Xia, Liang Chen, Zhenhai Sun, Yuchang Yang, Tuo Leng
Abstract: Formal Geometric Problem Solving requires every step of reasoning within a strict logical framework, and has been a core challenge in artificial intelligence. This paper presents a unified neuro-symbolic reasoning framework, named Reflect Agent Verified Solve (RAVS), which deeply integrates large language models, agent architectures, and formal symbolic solvers. The core innovations of RAVS lie in two parts. First, we propose a role-separated neuro-symbolic collaboration mechanism: the LLM serves as a planner responsible for high-level semantic understanding and path reflection, while the symbolic solver acts as an executor responsible for formal verification and rigorous theorem application. The neural reasoning capability of the LLM and the logical completeness of the symbolic system complement each other, fundamentally eliminating the risk of hallucinations. Second, we construct the first bidirectional symbolic reasoning engine that fully unifies forward and backward solving. For each theorem, we define two reversible operations: forward application and backward decomposition. We further provide a theoretical analysis that bidirectional reasoning achieves an exponential complexity advantage over unidirectional reasoning. Together, these two mechanisms form an iterative closed loop between neural planning and symbolic execution. On the FormalGeo7K benchmark, RAVS achieves a 95.8% solving accuracy, substantially outperforming existing state-of-the-art methods, without requiring any additional problem-specific annotated data.
Authors:
Allen Nie, Xavier Daull, Zhiyi Kuang, Abhinav Akkiraju, Anish Chaudhuri, Max Piasevoli, Ryan Rong, YuCheng Yuan, Prerit Choudhary, Shannon Xiao, Rasool Fakoor, Adith Swaminathan, Ching-An ChengAbstract: Generative optimization uses large language models (LLMs) to iteratively improve artifacts (like code, workflows or prompts) using execution feedback. It is a promising approach to building self-improving agents, yet in practice remains brittle: despite active research, only 9% of surveyed agents used any automated optimization. We argue that this brittleness arises because, to set up a learning loop, an engineer must make "hidden" design choices: What can the optimizer edit and what learning evidence is provided per update? We investigate three factors that affect most applications: the starting artifact, batching execution traces as experiences, and determining a credit horizon with truncated traces. Through case studies in MLAgentBench, Atari, and BigBench Extra Hard, we find that decisions around these factors can determine whether generative optimization succeeds and they are not made explicit in prior works. Different starting artifacts determine which solutions are reachable in MLAgentBench, truncated traces can still improve Atari agents, and larger minibatches do not monotonically improve generalization on BBEH. We conclude that the lack of a simple, universal learning-loop setup across domains is a major hurdle for productionization and adoption, and provide practical guidance for making these design choices.
PaperID: 3780, Poster
Abstract: Static-input Spiking Neural Networks (SNNs) commonly create temporal dynamics by repeating the same image over multiple timesteps. For handwritten text recognition, however, this convention is input-redundant: the image contains no native event stream, while recognition depends on faint stroke fragments, degraded ink boundaries, and monotonic left-to-right alignment. We introduce NeuroInk, a retinomorphic spike-oriented recognizer that converts static handwriting into a low-timestep sequence of temporally ordered ink evidence. Instead of replaying identical frames, NeuroInk performs temporal ink transduction, decomposing each text-line image into coarse stroke-mass support and fine stroke-boundary evidence. The resulting evidence sequence is injected as time-varying currents into an adaptive spiking visual hierarchy, where high-frequency enhancement and mid-level detail reinjection are used to preserve fine ink structures during spiking propagation. For CTC decoding, two-dimensional spiking maps are read out as width-wise columnar sequences and processed by CTC-aligned association circuits that combine local horizontal context, global association, and selective trace memory. On the LAM benchmark, NeuroInk achieves 2.4% CER on the validation split and 2.6% CER on the test split using only two timesteps and greedy CTC decoding, without an external language model or lexicon. These results support the view that static-input SNNs sequence recognition can benefit from co-designing sensory temporalization, spiking visual abstraction, and monotonic sequence association.
PaperID: 3781, Poster
Abstract: Dual-system Vision-Language-Action (VLA) models improve real-time robotic control by pairing a slow, reasoning-capable generalist with a fast specialist action expert. However, existing methods invoke the generalist at a fixed frequency, ignoring the fact that decision-making complexity varies throughout a rollout. This static strategy wastes computation in easy phases and can delay intervention when unexpected scene changes require renewed high-level reasoning. We propose TUD (Triggering generalist reasoning via predictive Uncertainty for Dual-system VLA), an adaptive inference framework that selectively skips unnecessary generalist calls. TUD measures the cross-step dispersion of action re-predictions at the upcoming chunk slot, computed from the specialist's existing forwards under the cached generalist context, as a predictive uncertainty signal. This signal captures how much the future action plan shifts as new observations arrive, enabling TUD to invoke the generalist only when the cached plan becomes unreliable, and separates success from failure more reliably than prior uncertainty signals. The signal needs no manually labelled phase boundaries or auxiliary uncertainty model, and is computed from forwards the architecture already runs. On VLA-Arena, TUD finds a more favorable cost--success trade-off than non-adaptive baselines, tracing an entire operating curve as a single threshold is varied, and substantially reduces VLM calls at matched success rate. Our results suggest that predictive uncertainty provides a practical criterion for adaptive reasoning in efficient VLA control.
PaperID: 3782, Poster
Abstract: Protein language models (PLMs) provide valuable evolutionary priors for fitness estimation, but may be misaligned with empirical measurements due to context-dependent selective pressures, potentially leading to hallucination-like predictions under limited experimental supervision. In this study, we introduce MmProt, a structure-prompted Multimodal Protein Language Model for preference-aligned fitness prediction with limited experimentally grounded supervision. MmProt adopts a two-stage multimodal learning strategy: contrastive pre-training on 40 million sequence--structure pairs to learn informative multimodal representations, followed by structure-prompted masked language modeling on naturally occurring related variants to capture protein-specific evolutionary constraints. When limited fitness measurements are available, MmProt further calibrates its variant preferences by encouraging the model to rank high-fitness variants above low-fitness variants. Extensive experiments across multiple benchmarks show that MmProt: (i) achieves state-of-the-art average performance on fitness prediction across 11 diverse protein targets under two challenging extrapolation settings, improving over the best-performing PLM baselines by 3.2% and 13.5% on average, respectively; (ii) captures protein-specific evolutionary patterns in a case study of the SARS-CoV-2 Spike protein, where its zero-shot predictions outperform leading PLMs fine-tuned on 512 fitness-labeled variants by a relative improvement of 3.2%; and (iii) delivers competitive results on eight downstream protein representation learning tasks. Together, these results highlight MmProt's potential as a practical tool for investigating protein evolution and supporting broader applications in protein modeling. Code and data are available at https://anonymous.4open.science/r/MmProt-239E/.
PaperID: 3783, Poster
Authors: Liang Yin, Zhan'ao yao, Jiahui Shi, Songlin Yu, Jianjun Liu
Abstract: In materials-science papers, quantitative evidence—XRD peak shifts, Raman ID/IG ratios, high-frequency Z′ intercepts in EIS—lives in line plots. Hosted vision–language models describe these plots fluently but cannot recover the (x, y) values an agent needs to query, compare, or integrate. We release three coupled artifacts. MatCurvs is a real-coordinate chart benchmark for materials-science spectra, organized in three tiers spanning XRD, Raman, XPS, XAFS, EIS, GCD, and CV: 1,375,165 classified subfigures from 55,763 articles (L0); 204 panels with manual per-curve pixel polylines (L1); and 8,026 verified real-coordinate curves over 3,817 panels (L2). MatDeplot is a local extractor that recovers 45.8% of curves within a 5-pixel IoU tolerance against L1 ground truth, versus ≤1.2% for the strongest hosted VLM—a 38× pixel-anchoring gap, closed in roughly 1.5 seconds per image. MatCurvs-Reasoning is a 14,740-question deterministic numerical evaluation: feeding MatDeplot’s extracted (x, y) data to an LLM cuts median relative error from 54% for an image-only VLM to about 4%, and a text LLM with the data alone matches the same VLM prompted with both image and data. By turning each published chart into a numerically queryable curve, the released artifacts let a deployer compute derived materials parameters—Raman ID/IG disorder ratios, XRD Scherrer crystallite sizes, and phase-purity audits—directly from a literature-scale chart corpus, where most papers under-report these numbers and the chart is the only available record.
PaperID: 3784, Poster
Abstract: We study a stochastic bandit problem in which learning is split between two agents. In each round, a human observes rewards but can choose only between a pair of arms selected by an assistant, while the assistant observes the human's behavior but not realized rewards. We call this model the assistive dueling bandit. The human is modeled as a learning agent whose regret on any subset of arms is bounded by a function g(T) unknown to the assistant. This is a natural model for recommender systems or AI assistants that must present a slate of options to a user who may be still learning about their own preferences. We provide a general reduction from assistive dueling bandits to the problem of max-finding with imprecise feedback. With it, we design assistant algorithms that cooperate with any g(T)-regret human to achieve a joint regret of \tilde\mathcalO(K \cdot g(T/K)), even without knowledge of g(T). Crucially, when g(T) \in \mathcalO(\sqrtT), our method recovers the nearly optimal minimax rate of \tilde\mathcalO(\sqrtKT), implying that splitting up the responsibilities of learning in this way can be done without suffering any excess regret asymptotically. We complement our theoretical analysis with experimental evidence that this algorithm outperforms an assistant implemented via standard dueling bandit algorithms.
PaperID: 3785, Poster
Authors:
Hongjie Cao, Yuxuan Yang, Yunpeng Mei, Peng Cheng, Chenyu Wang, Jiamin Wang, Xiaoyi Fan, Fang Deng, Gao Huang, Jie Chen, Gang WangAbstract: Flow-based vision-language-action (VLA) policies provide a compelling recipe for robot control: they inherit broad behavioral priors from large-scale imitation and generate smooth, temporally extended continuous action chunks. The same properties, however, make offline post-training delicate. Mixed offline robot datasets contain expert behavior as well as partial completions and failures, requiring improvement without drifting far from the data support. An offline critic can in principle provide useful directions, but actor-side critic optimization may suffer from extrapolation and push high-dimensional action chunks off-support. In contrast, advantage-conditioned or reweighting-based methods are more stable, but tend to be conservative and restricted to actions already present in the dataset. We propose VISTA, a post-training framework that turns value estimates into support-anchored distillation targets for pretrained flow policies. VISTA queries a critic only at dataset actions, computes a normalized action-gradient direction, and uses this bounded direction to shift clean-action endpoints and velocity targets. A frozen flow teacher preserves the pretrained action prior, while a behavior-cloning anchor keeps the adapted targets close to the data distribution. The shaped supervision is distilled into a time-indexed action denoising (TAD) student, which predicts time-indexed clean-action estimates with a single network evaluation and obtains the executed action by an analytic final readout. We treat the induced target shift as conservative local guidance rather than a global policy-improvement guarantee. Across five BEHAVIOR-1K simulation tasks and five real-world bimanual manipulation tasks, VISTA achieves the best average success rate in our evaluation against behavior cloning (BC), advantage-conditioned BC, IDQL, and flow Q-learning, while reducing action-expert computation by 15× (and end-to-end policy inference by 1.8×) relative to the 20-step flow teacher.
PaperID: 3786, Poster
Abstract: Inference-time search (e.g., Tree of Thoughts, ToT) with large language models (LLMs) often concentrates on a small set of structurally or semantically similar trajectories, leaving alternatives underexplored—a failure mode we call reasoning basin collapse. We introduce \textscBASIN, a training-free selection method that penalizes repeated visits to the same reasoning basin, defined symbolically for arithmetic tasks or semantically via hypothesis clustering, thereby reallocating search across diverse reasoning strategies. Under matched inference budgets, \textscBASIN improves over ToT by up to +22pp on Game of 24. To explain why this quality-agnostic penalty is effective, we introduce the \emphredundancy gap \Delta, which measures how differently search concentrates for correct versus incorrect predictions: standard ToT often operates near \Delta \approx 0, failing to distinguish correct from incorrect searches by their visit patterns, while \textscBASIN consistently induces \Delta > 0, concentrating search on correct basins while dispersing incorrect ones. More broadly, \textscBASIN suggests structure-aware selection as a simple and general approach to improving inference-time reasoning.
Authors: Ziyue Zeng, Xun Su, Haoyuan Liu, Bingyu Lu, Yui Tatsumi, Hiroshi Watanabe
Abstract: At ultra-low bitrates, high-fidelity reconstruction requires sampling plausible videos from the posterior rather than regressing to oversmoothed conditional means. We propose Generative Video Codebook Codec (GVCC), a zero-shot framework in which a pretrained video generative model serves directly as the decoder, and the transmitted bitstream specifies its generation trajectory. Modern rectified-flow video models are typically sampled with deterministic ODE solvers, which leave no per-step stochastic channel for transmitting compressed information. GVCC addresses this by converting the deterministic flow sampler into an equivalent marginal-preserving stochastic process, so that information can be transmitted by encoding the per-step stochastic innovations. Unlike images, videos introduce longer temporal dependencies and more diverse conditioning modes. We instantiate GVCC in three practical modes: Text-to-Video (T2V) without a reference frame, autoregressive Image-to-Video (I2V) with tail latent correction, and First-Last-Frame-to-Video (FLF2V) with boundary-sharing Group of Pictures (GOP) chaining. On UVG, GVCC achieves the lowest LPIPS among evaluated baselines across three representative bitrate regimes (down to ~0.003\,bpp), with 65% LPIPS reduction over DCVC-RT at matched bitrate.
PaperID: 3788, Poster
Authors: yongqi tian, Junyong Liu, Jinkun Ran, Haoyuan He, Caigui Jiang
Abstract: Generalized Category Discovery (GCD) aims to recognize both known and novel categories from a small labeled set and a large unlabeled pool. Existing methods often rely on sparse supervision to shape the representation space, which can lead to unstable separation between old and new categories, especially in fine-grained recognition scenarios. In this paper, we propose PD-GCD, a graph-structured propagation framework that views sparse labels as semantic seed signals and unlabeled samples as nodes on a data manifold. PD-GCD first learns a GCD-oriented representation space and performs Anchor Mining to identify propagation anchors from the unlabeled pool, balancing semantic coverage with boundary-aware exploration. These anchors are then used by Graph Propagation to spread semantic information over the data graph and produce propagated soft labels. To further align model predictions with the propagated structure, we introduce graph-guided structural distillation and graph consistency regularization, forming a closed-loop process between anchor mining, label propagation, and representation refinement. Extensive experiments on six benchmarks demonstrate that PD-GCD achieves state-of-the-art performance, with absolute All-accuracy gains of 6.7% across all datasets and 9.4% on fine-grained datasets.
Abstract: Large language models are increasingly deployed with test-time strategies: sample N responses, score them with a reward model or verifier, and return the best. This deployment rule exposes a mismatch in post-training: standard objectives optimize the mean reward of a single response, whereas best-of-N performance is governed by the upper tail of the reward distribution. Recent test-time-aware objectives partly address this mismatch, but typically assume that training can use the same per-prompt rollout budget as deployment, which is impractical when post-training must cover many prompts while deployment can allocate much larger per-prompt test-time compute. We study this budget-mismatch regime, where only m\ll N per-prompt rollouts are available during training but the target objective is best-of-N deployment. Under structural assumptions on the reward tails, we show that the policy gradient of the best-of-N objective can be approximated from a much smaller rollout group by extrapolating upper-tail statistics. This yields a family of Tail-Extrapolated estimators for best-of-N-oriented post-training: a simple direct estimator, Tail-Extrapolated Advantage (TEA), and a fixed-order debiased Prefix-TEA estimator based on moment cancellation. Experiments on instruction-following tasks show that TEA and Prefix-TEA improve best-of-N performance across different language models, reward models and datasets under various training and test-time budget settings.
Abstract: Sparse rewards pose a central challenge in reinforcement learning, since agents receive no informative signal until they reach their goal. Intrinsic-reward methods address this issue by optimizing non-stationary objectives such as novelty, prediction error, or skill diversity, thereby injecting a supervision signal into the problem. While effective, these methods often require that the extrinsic (sparse) reward can be evaluated -- either online or during offline relabeling of the stored transitions. This limitation is particularly vexing for multi-task, meta-, and continual reinforcement learning, where agents' interactions with the environment are usually \em reward-free. In this work, we present a method to pre-train \em transferable exploration policies that rapidly adapt to sparse rewards at downstream task time. Our objective maximizes state-space covering for the occupancy measure, and can be framed in terms of entropy maximization. Its algorithmic implementation, ROVER, leverages recent advances on the operatorial formulation of RL to estimate occupancy with a learned resolvent world model, bypassing common hurdles associated with density and entropy estimation. ROVER further introduces a virtual ``sink" state for unexplored regions, balancing coverage of known states with expansion into unseen ones and preventing cyclic expansion–collapse behavior during learning. In tabular and pixel-based sparse navigation tasks, ROVER produces more uniform aggregate coverage and stronger initializations for downstream tasks than standard reward-free baselines.
Authors:
Jintang Xue, Xinyu Wang, Yixing Wu, Jingwen Chen, C.-C. J KuoAbstract: 3D multimodal large language models (3D MLLMs) describe a 3D object as a whole but cannot address, name, or reason about its parts. Prior part-aware attempts add segmentation decoders, heavier 3D encoders, or bounding-box grammars at substantial parameter cost. We take a fundamentally different path: we reorganize the input token stream so that parts become directly addressable through the LLM’s own vocabulary. Our model, 3D-PLOT-LLM, partitions the frozen point encoder’s patches into K locally coherent regions and inserts, before each region’s patch tokens, a learnable per-region marker and a reserved vocabulary token \; a Marker-Space Refinement (MSR) module then conditions each marker on its region’s spatial statistics and adjacency neighbors. The model thus cites parts in its output and follows prompts that refer to parts by token, a capability absent from prior object-level 3D MLLMs. To probe this interface, we construct PartVerse-QA, a vocabulary-level part-QA benchmark adapted from PartVerse mesh annotations (77K training pairs and 588 held-out queries on disjoint object splits), on which 3D-PLOT-LLM reaches caption-to-slots Jaccard 0.459 and Exact-match 13.78% (+64% relative over the strongest non-MSR variant), with a slot-to-caption GPT-4o judge of 44.68. On the 3DCoMPaT-GrIn part-aware grounded description benchmark, 3D-PLOT-LLM outperforms PointLLM, Kestrel, PARIS3D, and SegPoint on every text-output metric, and ShapeLLM on 3 of 4, with up to +3.03 GPT-4o judge over PointLLM. On Objaverse whole-object captioning, adding PartVerse-QA at Stage 2 yields +0.65 SBERT and +1.85 GPT-4o over PointLLM, and tops PointLLM-PiSA on 4 of 5 traditional metrics (SBERT, SimCSE, BLEU-1, METEOR) despite targeting a different (part-grounded) objective. All with under 1M new trainable parameters on a frozen point encoder, an order of magnitude below prior part-aware 3D MLLMs, and no segmentation decoder or bounding-box head.
PaperID: 3792, Poster
Abstract: Diffusion-weighted MRI (dMRI) requires densely sampling q-space across many directions and shells, leading to long scan times that limit clinical use. Angular super-resolution (ASR) recovers a dense signal from a sparse acquisition, enabling substantially shorter scans. We propose a physics-guided flow matching framework for ASR built on a basic fact about dMRI: under the cumulant expansion, the diffusion signal naturally splits into a known linear function of the q-vectors plus a residual carrying higher-order angular detail. We exploit this twice. First, we initialize the flow from a closed-form dense DTI estimate, recasting ASR as residual flow matching from a physically meaningful starting point rather than uninformative noise. Second, we impose a microstructure-consistency loss, rooted in the physical principle that any two diffusion-weighted signals from the same voxel must be explained by a single underlying set of microstructural parameters. Together, these designs anchor predictions wherever the acquisition is informative while leaving the learned prior to model crossing-fiber and higher-order angular detail. On the Human Connectome Project (HCP-YA) and UK Biobank, our method matches or surpasses analytical baselines and recent deep learning methods on signal reconstruction, kurtosis scalar maps, and fiber-orientation reconstruction under aggressive subsampling.
PaperID: 3793, Poster
Authors:
Fangming Zhao, Fulun Ye, Xiaofei Yue, Ziming Zhao, Yu Peng, Junyu Chen, Tingting LiAbstract: The key-value (KV) cache has become a dominant memory and bandwidth bottleneck for serving long-context large language models, motivating a growing body of work on KV cache compression. Most existing methods follow a common recipe: a uniform per-layer budget and attention top-k token selection, but this recipe is \emphdata-oblivious (ignoring cross-layer variation in attention concentration) and \emphredundancy-blind (retaining near-duplicate key--value entries). We propose SpectralKV, a training-free framework that addresses both issues jointly. Across layers, a lightweight concentration-based criterion redistributes a fixed global budget based on the entropy of observation-window attention, giving more slots to layers whose attention is more concentrated. Within each layer, a spectral coreset selector treats keys and values as geometric objects in a metric-aligned joint feature space and selects tokens via residual pivoting, suppressing near-duplicates while preserving geometrically distinct directions. On Llama-3.1-8B and Qwen3-8B, SpectralKV consistently outperforms seven strong baselines on 15 LongBench tasks and PG-19 perplexity at 16K and 32K contexts, with the largest gains at aggressive compression ratios down to 6.25% keep. And a static-budget variant reduces peak prefill VRAM by up to 44% with negligible quality loss. Code and data are available at \urlhttps://anonymous.4open.science/r/spectralKV-2FF7.
PaperID: 3794, Poster
Abstract: Despite impressive progress, existing feed-forward stereo 3D reconstruction models produce a single, irrevocable prediction. When these models fail, there is no mechanism to correct their mistakes. We introduce Click3R, a framework for interactive geometric correction of feed-forward models using sparse human correspondence clicks. Given an image pair and a small set of cross-view point correspondences provided by a user, Click3R injects geometric constraints into the 3D reconstruction networks through a lightweight point correspondence adapter, enabling the model to resolve ambiguities in challenging scenarios such as repetitive textures, symmetric structures, large viewpoint changes, or textureless regions. To evaluate interactive reconstruction, we curate a challenging benchmark specifically designed to expose failure modes of existing feed-forward stereo 3D reconstruction methods. Experiments show that Click3R dramatically reduces reconstruction error with as few as a single correspondence click, improving performance on both standard benchmarks and our curated dataset. Our results demonstrate that sparse human interaction provides an effective and practical mechanism for correcting geometric reconstruction errors.
Abstract: We propose the Tikhonov layer, a graph neural network layer that is interpretable by design: once trained, its learned parameters directly reveal which node features and which aspects of the graph topology were leveraged for prediction. In practice, the layer's propagation matrix takes the closed-form R = (p(L)+Q)^-1Q, where L is the normalized graph Laplacian, Q = \mathrmdiag(q_1,\ldots,q_n) a learnable diagonal matrix of positive node-importance scores, and p(\cdot) a learnable polynomial. For any input feature x, the layer output Rx is the exact minimizer of a generalized graph Tikhonov problem that trades off node-level data fidelity against a topology-driven regularization penalty. The learned pair (\q_i\,p) constitutes a built-in explanation: large q_i indicates that node i's own features drive the prediction, while small q_i signals reliance on the local graph topology; the shape of p reveals whether homophily, heterophily, or a band-pass response is exploited. Expressivity is preserved by routing complexity through a dedicated, arbitrarily deep Q-network that produces the importance scores, while the Tikhonov layer itself remains transparent. We prove that distinct explanations necessarily yield distinct propagation operators, preventing degenerate explanations by design. Additionally, the Tikhonov layer provides, in a single layer, a global receptive field under mild support-connectivity conditions, mitigating both oversmoothing and oversquashing. Experiments on standard graph classification benchmarks confirm that the model matches (and sometimes outperforms) opaque baselines while producing interpretable and faithful explanations.
PaperID: 3796, Poster
Authors: Daniel Daza, Alberto Bernardi, Luca Costabello, Christophe Gueret, Michael Cochez, M.C. Schut
Abstract: Machine learning models for answering logical queries on knowledge graphs estimate the likelihood of answers that are not reachable via direct traversal. Recent work has proposed extending such queries with \emphsoft entity constraints, which require the \emphtarget variable to be similar or dissimilar to specified sets of entities. We study a more general formulation where similarity constraints are specified on an arbitrary variable in a query, and introduce SCORE: a computationally efficient and interpretable method for incorporating similarity constraints on an arbitrary variable of a query. SCORE is based on a lightweight and interpretable score adjustment function that only requires tuning two hyperparameters on a validation set. Our experiments show that SCORE effectively incorporates preferences while maintaining strong query answering performance. It significantly improves ranking quality when constraints apply to the target variable and performs effectively in the more general case involving constraints on intermediate variables.
Abstract: Most practical high-resolution text-to-image systems rely on latent diffusion models, where generation is performed in a compact latent space and a decoder maps latents back to pixels. Yet the latent-to-pixel decoder in this pipeline remains primarily reconstruction-oriented: as the sole pathway from latents to pixels, it is optimized to invert an encoder rather than to synthesize high-resolution details, and becomes increasingly costly in megapixel pipelines. This gap calls for a more expressive and efficient decoding paradigm. Motivated by recent progress in scalable pixel-space diffusion, we introduce PiD, a Pixel diffusion Decoder that reformulates latent decoding as conditional pixel diffusion, unifying decoding and upsampling into one generative module. By denoising directly in high-resolution pixel space, PiD synthesizes 4×and even 8× upscaled images with low latency. For latent conditioning, a lightweight sigma-aware adapter injects noise-corrupted latents into the backbone, enabling PiD to handle partially denoised states and terminate the base diffusion process early. To further improve efficiency, we distill the decoder using DMD2, reducing inference to just 4 steps. PiD applies to both conventional VAE latents and semantic latents (e.g., SigLIP, DINOv2) used in recent RAE-based models. Across multiple latent spaces and base generators, PiD improves visual fidelity over conventional decode-then-upsample cascades while reducing memory and latency, decoding latents of 512×512 images into 2048×2048 pixels in under 1 second with 13 GB peak memory on a consumer RTX 5090, and as fast as 210 ms on a data-center GPU, about 6× faster than cascaded diffusion-based super-resolution pipelines.
Abstract: The human brain processes dynamic visual input through hierarchically organized, functionally specialized regions. While recent in silico brain encoding models can synthesize optimal stimuli to probe selectivity in different brain regions, prior work has been largely limited to static images, leaving dynamic visual processing underexplored. We introduce a novel neural-guided video synthesis framework that generates stimuli optimized for target brain regions across visual cortex. Our method performs evolutionary search over a structured prompt space, guided by a dynamic encoding model that predicts voxel-level responses to video inputs. By maximizing predicted activity for a target ROI, the framework efficiently discovers hyper-activating dynamic stimuli that consistently surpass handcrafted localizer videos. The synthesized videos recover known selectivities across ventral, dorsal, and lateral pathways, and further reveal systematic differences in sensitivity to temporal dynamics. A searchlight analysis provides new insight into the progression toward increasingly complex social-dynamic features along the lateral stream, further supported by probing with synthesized abstract, non-naturalistic stimuli. Taken together, our framework enables in silico exploration of dynamic visual selectivity, with new predictions for in vivo experiments.
PaperID: 3799, Poster
Authors: Yanming Cheng, Di Zhang
Abstract: A concept-bottleneck classifier can remain correct, remain \ell_2-robust in pixel space, and still silently rewire which concepts carry its decision under a small perturbation. The label looks safe; the audit trail has been swapped out underneath it. This failure mode is not a quirk of randomized smoothing: any certifier that sees only the input and the label distribution admits classifiers with arbitrarily large concept drift inside its certified ball. To certify the missing object, we propose CRISP, which delivers a deterministic concept-space certified radius r^\star(x) inside which both the label and the concept activations are stable. The radius is given by a closed-form Lipschitz--margin bound and is tight up to a factor of two on the relevant encoder/head/margin triple. Under \varepsilon-concept-faithfulness and a known concept-to-target subgraph, r^\star further certifies the causally meaningful concepts; a matching impossibility result shows that faithfulness cannot be removed. Training relies on interventional proxies---SCM simulators, attribute editors, environment swaps, concept-matched retrieval---which approximate \mathrmdo(\cdot)-interventions on natural images. We bound the proxy-vs-truth gap in closed form and measure it directly in the synthetic regime where ground truth is available. Empirically, the bound is informative rather than vacuous. On a synthetic SCM with observable c^\star, 923 correctly-classified samples (5 seeds, k \in \3,6\) yield zero certificate violations under budget-5.0 APGD, with empirical tightness r^\star/r_\mathrmadv \approx 0.32. Concept-fidelity (MCC) rises from [0.37,0.51] for plain CBM to [0.74,0.83] for CRISP, and concept-space attack success drops by up to 13 points at matched budget. A frozen-backbone Waterbirds run preserves the zero-violation property across 1,980 samples and surfaces what we call the Lipschitz tax: a \approx 27× gap between architectural and local encoder Lipschitz constants that is the dominant bottleneck to tightness on natural images. A preregistered protocol for CUB-200 and CANDLE (Reddy et al., 2022) is fixed in Appendix C.8.
PaperID: 3800, Poster
Abstract: General-purpose world models promise scalable policy evaluation, optimization, and planning, yet achieving the required level of robustness remains challenging. Unlike policy learning which primarily focuses on optimal actions, a world model needs to be reliable over a vast space of suboptimal actions, which are often underrepresented in action-labeled robot interactions. To address this challenge, we propose World Action Verifier (WAV), a framework that enables world models to identify their own prediction errors and self-improve. The key idea is to decompose action-conditioned state prediction into two independently verifiable factors: state plausibility and action reachability. We show that verifying these factors is significantly more tractable than direct forward prediction due to two underlying asymmetries: the broader availability of action-free data and the lower dimensionality of action-relevant features. Leveraging these asymmetries, we augment a world model with (i) a diverse subgoal generator obtained from video corpora and (ii) a sparse inverse model that infers actions from a subset of state features. By enforcing cycle consistency among proposed subgoals, inferred actions, and forward rollouts, WAV provides an effective verification mechanism in under-explored regimes, where existing methods often fail. Across nine tasks spanning MiniGrid, RoboMimic, and ManiSkill, our method achieves 2× higher sample efficiency while improving downstream policy performance by over 22%.
Abstract: Modern Bayesian optimization and adaptive sampling methods increasingly rely on nonlinear parametric models, yet theoretical guarantees for such models under adaptive data collection remain limited. Existing analyses largely focus on Gaussian processes, kernel machines, linear models, or linearized neural approximations, leaving a gap between theory and the nonlinear models used in practice. We develop a kernel-based framework for analyzing regularized nonlinear parametric models trained on adaptively collected data. Our approach uses kernels over the parameter space to induce reproducing-kernel Hilbert space structures over the corresponding model class, yielding confidence bounds for models trained with broad classes of regularized convex losses. We show how these bounds can support convergence guarantees for nonlinear acquisition and surrogate models, including randomized regularized policies that select points by maximizing a trained random model. These results provide a unified route to analyzing nonlinear parametric models in Bayesian optimization and related adaptive optimization settings.
Authors: Caleb Jore, Jialin Liu
Abstract: Many scientific and combinatorial problems admit multiple correct solutions, not a single label. Standard supervised learning resolves this ambiguity by choosing one solution as the target, but this hidden selector can be arbitrary, discontinuous, and harder to learn than the underlying solution set. We study bifurcation models, a weight-tied dynamical view in which different initializations can converge to different stable equilibria, so the model represents an attractor landscape rather than one chosen branch. We prove that broad set-valued maps with locally Lipschitz branches can be represented by regular equilibrium dynamics and that the induced selectors are almost everywhere regular, while manual selectors can be arbitrarily irregular. Experiments on frustrated Ising models show that such dynamics can discover multiple valid equilibria without branch labels and outperform single-branch supervision. Allen-Cahn experiments further show that diversity is not automatic: it can be encouraged explicitly, but with an accuracy--diversity tradeoff.
Abstract: Consider a marketplace of AI tools, each with slightly different strengths and weaknesses. By picking the right model for the task at hand, a user can do better than simply using the same model for everything. Routers operate under a similar principle, where sophisticated model selection can increase overall performance. However, aggregation is often noisy, reflecting imperfect user choices or routing decisions. This leads to two main questions: first, what does a "healthy marketplace" of models look like for maximizing consumer utility? Secondly, how can we incentivize producers to create such models? We show that winrate, a standard benchmark in LLM evaluation, can incentivize model creators to homogenize for both types of model changes, reducing consumer welfare. We propose a new mechanism, weighted winrate, which rewards models for answers that are higher quality, and show that it provably improves incentives for producers to specialize and increases consumer welfare. We conclude by exploring the impact of our theoretical results in empirical benchmark datasets and discussing implications for benchmark design.
PaperID: 3804, Poster
Abstract: Processing ultra-long multimodal inputs efficiently remains a critical bottleneck for Vision-Language Models (VLMs). This systematic challenge is essentially three-fold: (1) Efficient Architecture: compressing redundant visual sequences without sacrificing fine-grained details; (2) Knowledge Transfer: seamlessly migrating the capabilities of strong pretrained VLMs into more efficient novel architectures; and (3) Generalization: maintaining robust perception and reasoning across diverse multimodal domains. To address this, we introduce InfiniteVL, a linear-sparse hybrid VLM framework designed for efficient long-context understanding. At its core, InfiniteVL pairs a linear architecture to compress long-range visual context with sparse attention to retrieve precise local details. To support this architecture in moderate academic resources, we developed a three-stage knowledge transfer pipeline backed by capability-oriented data construction instead of training from scratch. This ensures smooth architectural alignment, rapid capability recovery, and robust adaptation across diverse domains. Furthermore, we adapt this foundation to specific deployment scenarios, deriving Sparse InfiniteVL for long-video analysis and Streaming InfiniteVL for real-time scene perception. Extensive experiments demonstrate that InfiniteVL matches the performance of leading VLMs while achieving a 1.7× decoding speedup and a 2.7× reduction in memory. Notably, Sparse InfiniteVL accelerates prefilling by 5× at a 256K context length, while Streaming InfiniteVL delivers stable 25 FPS real-time processing with a constant memory footprint.
PaperID: 3805, Poster
Authors:
Jian Zhong, Zeyu Liu, Pingchuan Cao, Rongduo Han, Shunye Tang, Guohuan Xie, Mingyuan Qin, Ran Ji, Xiao Liang, Haining Zhang, Wei WangAbstract: Long-horizon dialogue agents must keep user memories stable enough for personalization while adapting when the user genuinely changes. Existing memory systems often append every observation or overwrite old entries by heuristic rules, which makes transient noise hard to distinguish from real belief drift. We introduce PreCoMem, a memory consolidation framework inspired by predictive coding that represents memory as confidence-weighted beliefs about the user's state. Its core signal, Effective Surprisal, combines semantic deviation and contradiction evidence, then downweights this signal when retrieval is diffuse and unreliable. According to effective surprisal, a two-threshold gate maps each turn to one of three deterministic updates: MAINTAIN reinforces supported beliefs, PROFILE stores ambiguous cues as low-confidence hypotheses, and CORRECT softly decays reliably contradicted facts. The PROFILE stage acts as an explicit waiting room, preventing premature commitment to noisy evidence. Experiments on LoCoMo, LongMemEval, and PersonaMem-v2 show consistent state-of-the-art accuracy; our new DiMoBench confirms that PreCoMem adapts to non-stationary belief drift without over-reacting to noise. The code is available at: \urlhttps://anonymous.4open.science/r/PreCoMem-9000.
PaperID: 3806, Poster
Abstract: Generating realistic financial time series is challenging as training data is often limited to a single historical path. With such scarce data, overfitting is hard to avoid, especially under adversarial training where a trained discriminator can memorize the training samples. To mitigate this, recent approaches train generators to minimize the discrepancy between untrained feature representations of real and generated time series. In these works, the feature maps are based on path signatures, which can fail to capture relevant time series properties at tractable truncation depths. In this work, we instead train generators by matching random convolutional features of real and generated time series. Existing random convolutional feature maps, such as Rocket and Hydra, have been shown to provide informative representations of real-world time series, but cannot supervise generative models because they are non-differentiable. We introduce ), a fully differentiable random convolutional feature map, suited to train generative time series models. We show that generators trained by matching random SOCK features consistently outperform signature and diffusion baselines across a wide range of small-sample financial datasets. We further demonstrate SOCK's expressiveness on two-sample hypothesis testing and time series classification tasks, where SOCK matches or outperforms existing unsupervised feature maps.
Abstract: Training data attribution (TDA) identifies which training examples most influenced a model's prediction. Influence function methods are a theoretically grounded family of TDA methods and exploit gradients. To overcome the scalability challenge arising from gradient computation, the most popular strategy is random projection (e.g., TRAK, LoGRA). However, this still faces two bottlenecks when scaling to large training sets and high-quality attribution: (i) storing and loading projected per-example gradients for all N training examples, where query latency is dominated by I/O; and (ii) forming the D × D inverse Hessian approximation, which costs O(D^2) memory. Both bottlenecks scale with the projection dimension D, yet increasing D is necessary for attribution quality---creating a quality--scalability tradeoff. We introduce LoRIF (Low-Rank Influence Functions), which exploits low-rank structures of gradient to address both bottlenecks. First, we store rank-c factors of projected per-example gradients rather than full matrices, reducing storage and query-time I/O from O(D) to O(c\sqrtD) per layer per sample. Second, we use truncated SVD with the Woodbury identity to approximate the inverse Hessian term in an r-dimensional subspace, reducing memory from O(D^2) to O(Dr). On models from 0.1B to 70B parameters trained on datasets with millions of examples, LoRIF achieves up to 20× storage reduction and query-time speedup compared to LoGRA, while matching or exceeding its attribution quality. LoRIF makes gradient-based TDA practical at frontier scale.
PaperID: 3808, Poster
Abstract: In reinforcement learning from limited feedback (RLLF), only a small fraction of an offline dataset can be labeled with rewards, and the central question is samples should be labeled to learn a strong policy from the resulting partially labeled dataset. Prior work formalized this as a reward-selection problem by focusing on the selection stage while treating downstream policy learning as a black box, in a regime where queried rewards are not retained for reward-model training. We instead study the retained-label setting, where queried rewards can be stored and used to fit a reward model before policy learning. We bound the suboptimality of the learned policy by two sources of error: one from offline RL on an offline dataset, and one from reward-model uncertainty. Since reward selection cannot change the offline dataset, the limited labeling budget must be used to strategically reduce reward uncertainty. Motivated by RLLF's observation that useful rewards tend to keep the agent on high-return trajectories, we propose successor-guided uncertainty reduction (SURE), which uses successor features to select rewards that are both reachable to high-valued states and uncertainty-reducing. Theoretically, we derive SURE from a bound-induced design objective and characterize its exact one-step marginal gain. Empirically, SURE reaches near full-feedback performance with few reward labels across a variety of domains, yielding a strong method for feedback-efficient reinforcement learning.
PaperID: 3809, Poster
Authors: David T Zagardo
Abstract: Evolution strategies optimize through fitness evaluations, but standard ES exposes data-dependent fitness values. We introduce DP-EGGROLL, a central-DP mechanism for EGGROLL-style population optimization with decomposable supervised losses. Each private step forms per-example candidate-loss vectors, mean-centers each row across candidates, clips the centered row, averages with a fixed nominal denominator, and adds calibrated Gaussian noise; reward shaping and the ES update are post-processing. Clipping gives the C/m add/remove sensitivity bound, while centering removes common-mode loss before clipping. We evaluate 10 public tabular benchmarks across five privacy budgets, using 10 final seeds and equal 32-configuration HPO budgets for fast DP-AdamW, scalar DP-ZO, centered DP-EGGROLL, and independently tuned uncentered DP-EGGROLL. Each selected final run is accounted as an individual (\varepsilon,10^-5)-DP run; multi-configuration HPO is a benchmark protocol, not a single-run deployment guarantee. On tabular classification, centered DP-EGGROLL is practically non-inferior to fast DP-AdamW on 22/25 AUROC endpoints and is faster per step in all 25 classification settings. Centering improves over tuned uncentered DP-EGGROLL on all 25 classification AUROC endpoints. Neural experiments include end-to-end small MLPs and a low-rank adapter stress test; both reinforce that centered population-vector privatization carries more signal than scalar DP-ZO. Regression is more heterogeneous, with centered DP-EGGROLL practically non-inferior on 14/25 RMSE endpoints.
Authors:
Jon Saad-Falcon, Avanika Narayan, Hakki Akengin, J. W Griffin, Herumb Shandilya, Adrian G Lafuente, Medhya Goel, Rebecca Joseph, Shlok Natarajan, Etash Guha, Shang Zhu, Ben Athiwaratkun, Azalia Mirhoseini, Christopher RéAbstract: Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure. Rapidly growing demand strains this paradigm, and cloud providers struggle to scale infrastructure at pace. Two advances create an opportunity to rethink this paradigm: small, local LMs (≤20B active parameters) now achieve competitive performance to frontier models on many tasks, and local accelerators (e.g., Apple M4 Max) can host these models at interactive latencies. This raises the question: can local inference viably redistribute demand from centralized infrastructure? Answering this requires measuring both whether local LMs can accurately answer real-world queries and whether they can do so efficiently enough to be practical on power-constrained devices (i.e., laptops). We propose intelligence per watt (IPW), task accuracy divided by unit of power, as a unified metric for assessing both the capability and efficiency of local inference across model-accelerator configurations. We conduct a large-scale empirical study across 20+ state-of-the-art local LMs, 8 hardware accelerators (local and cloud), and a representative subset of LLM traffic: 1M real-world single-turn chat and reasoning queries. For each query, we measure accuracy (local LM win rate against frontier models), energy consumption, latency, and power. Our analysis reveals three key findings. First, local LMs can successfully answer 88.7% of single-turn chat and reasoning queries with accuracy varying by domain. Second, longitudinal analysis from 2023-2025 shows progress in local inference viability: IPW improved 5.3×, driven by both algorithmic advances and accelerator improvements, with locally-serviceable query coverage increasing from 23.2% to 71.3%. Third, local accelerators achieve at least 1.4× lower IPW than cloud accelerators running identical models, revealing significant headroom for local accelerator optimization. These findings demonstrate that local inference can meaningfully redistribute demand from centralized infrastructure for a substantial subset of queries, with IPW serving as the critical metric for tracking this transition.
PaperID: 3811, Poster
Authors: Jiachen Li, Longzhen Tang, Shisheng Guo, Guolong Cui
Abstract: Facial video–based remote photoplethysmography (rPPG) estimates physiological signals from subtle spatiotemporal variations in facial videos, but remains vulnerable to head motion, illumination changes, and environmental perturbations. Existing methods mainly distinguish rPPG-related variations from motion and lighting artifacts at the signal or feature level, while largely neglecting the optical reflection properties of human skin. In this work, we formulate a skin reflection exponential model to describe the nonlinear relationship among skin reflectance, illumination, motion-induced variations, and blood-volume changes. Guided by this model, we propose a physically inspired rPPG estimation framework that explicitly incorporates skin reflectance priors into signal extraction. Given a facial video, facial keypoints are first detected and encoded to represent motion dynamics, while illumination-related features and candidate rPPG representations are jointly modeled under the proposed exponential formulation. By constraining rPPG estimation with skin-reflection physics, the framework improves separation of physiological components from motion and illumination disturbances. Experiments demonstrate improved robustness under challenging conditions.
PaperID: 3812, Poster
Abstract: Ensuring high returns while providing strong safety guarantees in offline imitation learning (IL) is a fundamental challenge: existing approaches either rely on explicitly specified cost functions, which are difficult to design for complex real-world safety constraints, or require expert demonstrations that are both safe and high-performing, which are often too costly to collect. We introduce , which leverages two complementary data sources: (i) safety-compliant demonstrations that faithfully follow safety guidelines but may be suboptimal, and (ii) performance-oriented demonstrations that achieve high returns but may violate safety guidelines. To address this problem, we propose DexDICE ( stimation). Instead of conventional distribution constraints, DexDICE combines DICE-style offline learning with safe support constraints, enabling agents to exploit high-return behaviors while remaining within safety-compliant regions. We evaluate DexDICE on a re-curated DSRL benchmark and validate it on a real-world mobile-robot navigation task trained directly from human teleoperation demonstrations, where it significantly outperforms offline IL baselines and remains robust to data scarcity and quality degradation.
PaperID: 3813, Poster
Abstract: Long-horizon visual search requires a vision-language policy to execute multi-step spatial actions over high-resolution images, where performance depends on the entire trajectory rather than any single decision. Existing supervision breaks at this horizon: outcome-only RLVR collapses many geometric decisions into a single delayed signal and discards the per-step geometry the environment already computes, while LLM-judged step rewards recover density at the cost of verifiability, scale linearly with trajectory length, and conflate distinct failure modes. We propose HSSP, a self-play framework in which a Hider adaptively weights challenging skeletons from a fixed trajectory pool and a Seeker learns to solve them through grounded spatial actions. Instead of relying on external judges, HSSP supervises every step with a geometry-verifiable process reward built from target-directed progress and exploration coverage, with trajectory length and spatial redundancy enforced as separate Constrained Markov Decision Process (CMDP) constraints. We pair HSSP with V-Trace, a trajectory-level companion dataset that augments skeleton seeds from established benchmarks with target-region annotations and explicit search-regime labels, ready for any trajectory-level training framework. The reward design admits closed-form guarantees on backtracking, find-versus-explore separation, and policy invariance, providing theoretical backing for the observed gains. Extensive empirical results on eleven challenging multimodal benchmarks, supported by controlled ablations and qualitative case studies, show that HSSP consistently outperforms advanced baselines and transfers zero-shot to OOD tasks.
PaperID: 3814, Poster
Abstract: Second-order optimizers (\eg Shampoo, SOAP) have emerged as compelling alternatives to Adam/AdamW and promise faster per-iteration convergence than AdamW in transformer training. However, they still remain costly in practice: even SOAP still pays substantial optimizer-state memory and eigendecomposition overhead. We ask whether the benefits of SOAP can be retained at nearly the efficiency of AdamW. In this paper, we answer yes with a lightweight yet improved SOAP, Factored One-sided Adam-Moment (FOAM). The path to practical impact is twofold: first, FOAM employs input-only preconditioning, dropping the output-side Kronecker factor to reduce memory and remove one eigendecomposition. A per-row Hessian decomposition identifies this as the right one-sided choice when output curvature is approximately isotropic. Second, FOAM replaces SOAP's dense rotated second moment with a factored Adam-moment estimator based on row-by-column statistics. Under the K-FAC Fisher assumption, this estimator is further proven to be pointwise equivalent to SOAP's update. Our empirical results support the claim that FOAM matches strong SOAP and Muon baselines in validation loss across a diverse scale of large language model (LLM) pre-training, uses optimizer state closer to AdamW, runs within Muon's wall time, and retains robustness across a five-fold learning-rate range. Scaling to 760M parameters, FOAM demonstrates that SOAP's gains need not require SOAP's cost.
PaperID: 3815, Poster
Authors: Yunfu Deng, Josiah Hanna
Abstract: Off-policy evaluation (OPE) estimates the expected return of a target policy from previously collected data without additional environment interaction. While OPE methods for flat Markovian policies are well studied, little work has addressed evaluating hierarchical policies in which a high-level policy selects temporally extended options that generate primitive actions until a termination condition is met. The primary existing approach applies per-decision importance sampling at the option level, but requires option-annotated trajectories and exhibits variance that grows with the horizon and action dimensionality. For non-hierarchical policies, fitted Q-evaluation (FQE) typically achieves lower error than importance sampling by leveraging the classic Bellman equation for policy evaluation; however, a naive extension of FQE to hierarchical policies results in a biased policy value estimate. In this paper, we introduce Intra-Option Fitted Q-Evaluation (IO-FQE), which estimates a value function in the augmented state-option space; this approach enables OPE of hierarchical policies even when we lack annotations of what option was ran in the data. We develop two continuous-action instantiations of IO-FQE and show on hierarchical continuous control tasks (AntMaze navigation and OGBench Puzzle manipulation) that IO-FQE eliminates the bias incurred by a naive application of FQE to hierarchical policies and substantially lowers mean squared-error compared to option-level importance sampling.
PaperID: 3816, Poster
Authors:
József Kovács, Amadeus Hauser, Gudrun Gröppel, Wolfgang Narzt, Philipp SeidlAbstract: Epilepsy affects approximately 1% of the global population, where seizure detection and reliable forecasting are critical for patient safety and quality of life. While emerging foundation models increasingly recast EEG analysis as a language modeling task, most remain confined to single-modality inputs and narrow domain generalization. In this work, we present NeuroNTP, a modular multimodal foundation model built on an xLSTM backbone and pretrained on one of the largest retrospective epilepsy datasets ever collected, comprising 150,000 hours of neural (EEG) and physiological (ECG, SpO_2) signals paired with rich clinical text. NeuroNTP employs a flexible, montage-geometry grounded tokenizer and optimizes a composite objective, synergizing causal next-token prediction with seizure-specific auxiliary tasks to enable unified sequence modeling across heterogeneous modalities. This approach allows the model to capture long-range temporal dependencies and cross-modal interactions that unimodal baselines ignore. In extensive experiments, we show that NeuroNTP outperforms supervised baselines by a substantial margin on held-out clinical cohorts. Furthermore, the model generalizes to public EEG benchmarks such as CHB-MIT and Siena Scalp EEG, achieving state-of-the-art performance under standard evaluation settings. We release code and model weights to facilitate further research.
PaperID: 3817, Poster
Abstract: Learning constraint-satisfying policies from offline data without risky online interaction is crucial for reliable deployment of reinforcement learning (RL) agents. Representative methods leverage Implicit Q-Learning and advantage-weighted regression to learn value functions and balance the reward–safety trade-off, demonstrating solid safety performance in safety-critical hard constraint scenarios. However, these methods typically couple reward and safety optimization during offline training and rely on a predefined unified factor to balance the reward–safety trade-off (conservatism). Which ignores the fact that different tasks, due to their distinct reward and cost functions, may require different degrees of conservatism to satisfy safety constraints, resulting in sub-optimal performance. To address this limitation, we propose Co3G, a compositional generative model with test-time controllable conservatism for offline safe RL. During training, Co3G decouples the optimization objectives by separately employing reward-optimal guidance and safety-optimal guidance to induce high-reward and high-safety actions. During deployment, Co3G combines the reward and safety objectives via compositional generation, and further provides three mechanisms—manual tuning, rejection sampling, and RL—to balance the reward and safety guidance scales during deployment, thereby achieving a superior reward–safety trade-off. Extensive experiments on the OSRL benchmark show that Co3G delivers strong safety performance under stringent safety constraints while retaining flexible control over conservatism, significantly outperforming previous hard-constraint and soft-constraint baselines.
Abstract: Mixed-Integer Linear Programming (MILP) is a fundamental tool for combinatorial optimization with extensive real-world applications. A central challenge lies in designing efficient MILP formulations. Large Language Models (LLMs) offer new opportunities to automate the modeling process, from deriving formulations to strengthening them. To ensure correctness, we need robust methods to compare formulations. However, existing approaches evaluate formulations numerically and fail to reason about the behavior on general problem instances. We resolve this limitation by introducing a constructive notion of MILP , a method that uses an LLM-based agent and the Lean proof assistant to verify proposed reformulations against a reference. To evaluate our approach, we introduce
PaperID: 3819, Poster
Abstract: Machine Learning Interatomic Potentials (MLIPs) based on Equivariant Graph Neural Networks (EGNNs) have greatly advanced the quantitative simulation of atomic systems. However, accurately resolving fine-grained structural geometry remains a critical challenge. In this work, we demonstrate through spectral analysis that the reliance of modern EGNNs on a single macroscopic cutoff radius inherently acts as a spatial low-pass filter. This structural bottleneck induces radial spectral confusion, suppressing the network's capability for fine-grained geometric modeling. To overcome this limitation, we propose the Multi-Cutoff Spectral Decomposition (MCSD) mechanism. As a universal, plug-and-play module, MCSD explicitly decomposes atomic interactions across multiple spatial scales. By leveraging envelope associativity to unify the neighborhood graph, this method decouples scale parameters from tensor product complexity, enabling multi-resolution representations with minimal overhead. Extensive evaluations across representative EGNN backbones demonstrate that MCSD not only consistently reduces prediction errors for energy, forces, and macroscopic stress tensors, but also significantly enhances the prediction fidelity of higher-order properties (such as lattice thermal conductivity), all while maintaining the smoothness of the potential energy surface and the energy conservation of long-term molecular dynamics simulations. The code is available at the anonymous repository: https://anonymous.4open.science/r/MCSD-C854.
PaperID: 3820, Poster
Authors:
Xuerui Qiu, Yutao Cui, Guozhen Zhang, Junzhe Li, Yaqi Zhao, Xiao Zhang, Yang Li, Songtao Liu, Miles Yang, Yu Shi, Zhao Zhong, Liefeng BoAbstract: Unified Multimodal Models struggle to bridge the fundamental gap between the abstract representations needed for visual understanding and the detailed primitives required for generation. Existing approaches typically compromise by employing decoupled encoders, stacking representation encoder atop VAEs, or utilizing discrete quantization. However, these methods often disrupt information coherence and lead to optimization conflicts. To this end, we introduce \tokenizer, a representation-harmonized pure ViT in the insight that visual modeling should evolve from generation to understanding. \tokenizer reformulates the standard backbone into a progressive learner that transitions from a Gen-ViT, which captures structure-preserving primitives, to a Sem-ViT for semantic encoding. Crucially, this transition is mediated by a Generation-Semantic Bottleneck (GSB), which compresses features into a low-dimensional space to filter noise for robust synthesis, then restores dimensionality to empower complex semantic comprehension. Built upon this foundation, we present \model, a native unified framework integrating perception and generation within a single parameter space. Extensive experiments establish \model as a new state-of-the-art. It sets a benchmark in visual reconstruction (rFID 0.08) and achieves top-tier generation performance on GenEval (0.86), DPG-Bench (86.4), and WISE (0.53), while simultaneously outperforming previous native UMMs by an average of ~10.0 points across eight challenging understanding benchmarks.
PaperID: 3821, Poster
Abstract: Large language models can have hidden propensities: behaviours that arise only in rare contexts. These hidden propensities pose a challenge to alignment audits, which seek to uncover problematic model behaviour but are limited in scope to scenarios anticipated by the auditor. To address this challenge, we explore (LPF) on examples of a target behaviour. While training 3–15% of model parameters, LPF reveals whether the model has a hidden propensity, even when standard behavioural evaluations cannot. LPF succeeds on a suite of model organisms with hidden propensities (8B-14B parameters), and also surfaces naturally occurring hidden propensities in open-weight models (27-32B parameters). LPF can thus distinguish between behaviourally equivalent models, even without knowing the context in which the propensity surfaces. LPF provides a novel affordance that may be useful for auditing frontier AI models.
Authors:
Sree Bhattacharyya, Samarth Khanna, Leona Chen, Lucas Craig, Tharun Dilliraj, James WangAbstract: Large Language Models (LLMs) are increasingly used in settings where reliable self-assessment is critical. Assessing model reliability has evolved from using probabilistic correctness estimates to, more recently, eliciting verbalized confidence. Confidence, however, has been shown to be an inconsistent and overoptimistic predictor of model correctness. Drawing on cognitive appraisal theory, a framework from human psychology that decomposes self-evaluation into multiple components, we propose a multidimensional perspective on model self-assessment. We elicit six appraisal-based dimensions of self-assessment, alongside confidence, and evaluate their utility for predicting model failure across 12 LLMs and 38 tasks spanning eight domains. We find that competence-related appraisal dimensions, particularly effort and ability, consistently match or outperform confidence across most settings. Effort additionally yields less overoptimistic estimates that remain stable across model sizes. In contrast, affective dimensions provide marginally predictive signals. Furthermore, the most informative dimension varies systematically with task characteristics: effort is most predictive for reasoning-intensive tasks, while ability and confidence dominate on retrieval-oriented tasks. Broadly, our findings indicate that structured multidimensional self-assessment is a promising approach to improving the reliability and safety of language model deployment across diverse real-world settings.
PaperID: 3823, Poster
Abstract: We propose PACE, Partial-state Amortized Constraint Editing, which treats task-native feasible partial states as the common semantic object for neural combinatorial optimization, unifying edit learning, constraint-preserving editing, anytime closure, and test-time scaling across routing and graph combinatorial optimization. Existing neural combinatorial optimization methods expose complementary strengths: constructive and adaptive expansion solvers keep meaningful partial solutions but couple them to serial or method-specific growth, while global prediction, diffusion, masked reconstruction, and complete-solution refinement provide scalable guidance yet often refine heatmaps, noisy solutions, reconstruction targets, or perturbed full outputs. These semantics make learning, search, and test-time scaling act on the same state object. PACE learns amortized constraint editing by ranking candidate edits conditioned on the current feasible partial state; a constraint-preserving transition \Gamma_I commits only edits that preserve legal continuation and extendability; and task-native closure \mathrmComp_I maps any extendable intermediate feasible partial state to a task-terminal feasible output. Budgeted refinement then spends extra inference-time compute through deeper edit steps, additional refinement rounds, or broader candidate sets along the same trajectory. We instantiate these semantics across TSP, ATSP, CVRP, MIS, MVC, MCL, and MCut, covering edge-oriented routing and node-oriented graph combinatorial optimization. Theoretical guarantees are structural rather than optimality claims, covering task-wise soundness, constructive completion, state-space closure in the extendable subset, terminal feasible completion at any editing depth, and elite-pool nondegradation. Empirically, PACE is competitive and budget-controllable, including ATSP-500 at 0.247% Drop, MIS RB-[800-1200] at 2.00%, and MVC at 0.07%.
PaperID: 3824, Poster
Abstract: AI-generated image detection has become an important problem in media forensics as modern generative models produce increasingly realistic images. Recent CLIP-based detectors show strong generalization ability, but CLIP image features entangle semantic content with forensic cues. Since semantic content is not an intrinsic generation trace, relying on it can introduce semantic shortcuts and unstable performance under unseen generator shifts. In this paper, we propose SAGE, a Semantic-Agnostic Image Embedding framework for generalized AI-generated image detection. SAGE constructs paired triplets of COCO images, SDXL-reconstructed images, and COCO captions so that real and fake images share the same caption-level semantics. It then suppresses semantic information by suppressing the semantic direction provided by CLIP text embeddings from CLIP image embeddings. To enable caption-free inference, SAGE learns a semantic anchor that approximates the role of paired text embeddings at test time. In addition, SAGE adopts a SoftTriple-style sub-center classifier to model intra-class variation in real and fake images, yielding a more stable authentication geometry across diverse generators. Experiments on CommunityForensics demonstrate that SAGE achieves competitive detection performance while substantially reducing generator-wise performance variation. Further analyses show that SAGE effectively reduces semantic dependence and forms a clearer real/fake feature geometry than strong baselines.
PaperID: 3825, Poster
Abstract: Cooperative multi-agent reinforcement learning (MARL) has long faced scalability challenges due to the exponential growth of the state-action space as the number of agents increases. Existing methods typically enhance scalability by filtering out communications with low-relevance agents. However, such filtering often relies on global state and fixed graph distribution, limiting adaptability in large-scale, communication-constrained environments. In this paper, we propose a scalable MARL method, called Adaptive Communication Range PPO (ACR-PPO), that decomposes the decision-making under communication budget constraints as a sequential process: a communication policy first selects each agent’s communication range within a given budget, followed by a behavior policy that conditions actions on the resulting neighborhood observations. More importantly, we provide a theoretical guarantee of monotonic performance improvement under communication budget constraints. Experiments across diverse scenarios demonstrate that ACR-PPO preserves policy performance while significantly reducing communication costs through adaptive range control.
PaperID: 3826, Poster
Authors:
Xin Xia, Chen Zhang, Kaibo LiuAbstract: Mobile sensing for real-time anomaly detection couples routing, sampling, and stopping: dispatch determines which locations are observed, while observations update the evidence used for future dispatch. We study this problem for a fleet of unmanned aerial vehicles (UAVs), where each UAV observes only its current location and anomalies may occur at unknown locations and times. The goal is to detect anomalies quickly while controlling false alarms under decision-dependent partial observations, local mobility, and collision-avoidance constraints. Existing quickest-detection methods typically ignore mobility and collision constraints, whereas learning-based routing methods scale to large fleets but are not designed for statistically calibrated detection. We formulate route-wise monitoring as a partially observable Markov decision process (POMDP) and propose DispatchPPO, a deep reinforcement learning method for long-horizon detection-aware dispatch. Its policy architecture, DispatchNet, uses autoregressive decoding with dynamic feasibility masking to generate collision-free joint UAV moves without enumerating the exponential joint action space. We establish theoretical properties for persistent coverage, evidence-based focusing, and detection delay through a resource-allocation information-rate bound. Experiments show that DispatchPPO reduces detection delay relative to heuristic and optimization-based baselines, while transferring effectively to real-robot and wildfire monitoring case studies.
PaperID: 3827, Poster
Abstract: Plot-to-code generation has been greatly advanced by recent multimodal large language models (MLLMs). However, existing studies mainly focus on one-pass generation or iterative sampling guided by verifier-based filtering, which does not align with the typical human programming workflow of generate first, then iteratively error-location and fix. This naturally motivates a test-time refinement paradigm for plot-to-code generation.In this work, we present the first study on plot-to-code generation via iterative test-time refinement. We propose a Visual-Code Diagnostics framework that enables refinement through visual feedback from source code. Specifically, we first categorize generation errors into four coarse-grained types, and then train lightweight discriminators to recognize and localize these errors based on the rendered visual outputs and corresponding code. Unlike traditional compiler- or debugger-based verification, our approach performs iterative refinement guided by one diagnostics system, enabling more effective semantic and visual correction during test time. Extensive experiments demonstrate that our method consistently improves generation quality across diverse generators and significantly outperforms strong test-time scaling baselines.
PaperID: 3828, Poster
Abstract: We propose an adaptive proximal gradient method for minimizing the sum of two functions, where one is a simple convex function, and the other belongs to one of the three classes: nonconvex smooth, convex nonsmooth, or convex smooth. The key feature of the method is an adaptive step size that accumulates historical gradient mapping norms in the denominator. Without any modification or knowledge of problem parameters, the method converges across all three problem classes under mild bounded-iterates and bounded-variance assumptions, with rates matching those of the proximal gradient method up to logarithmic factors, in both deterministic and stochastic settings. For the convex setting, we further propose an accelerated variant. It retains a similar near-optimal convergence rate for the nonsmooth case and achieves an improved rate of order \widetildeO\big(1/t^2 + \sigma/\sqrtt\big) for the smooth case, which is optimal up to logarithmic factors. Notably, we develop new techniques for controlling the effect of stochastic noise, which are applicable across all three problem classes in the stochastic setting and enable simplified analysis.
PaperID: 3829, Poster
Abstract: Unifying speech, sound, and music generation in one model is hindered by tradeoffs between fidelity, end-to-end training, in-context conditioning, and variable-length synthesis that no current paradigm fully resolves. To address this challenge, we present , a universal audio generation framework that extends autoregressive (AR) next-token prediction from discrete tokens to continuous audio latents: a thin flow-matching head replaces the softmax to predict rectified-flow velocities at each position, and a block-causal AR-Flow attention pattern produces arbitrary-length output. Joint training of multiple audio generation tasks faces an asymmetric text--audio mismatch: speech transcripts align to specific time spans and demand tight, time-aligned attention, whereas sound and music captions describe only overall semantics and rely on diffuse, holistic attention; mixing the two disproportionately degrades sound and music generation. We address this asymmetry at two levels: a data reformulation strategy that unifies all three tasks under a single description-style conditioning interface, and a novel architecture Asymmetric Mixture-of-Modality-Experts (A-MoME), which adds a dedicated residual expert for speech while sound and music share the backbone, incurring no inference overhead on non-speech inputs. Experimental results demonstrate that AudioCALM matches modality-specific state-of-the-art and outperforms prior unified baselines on speech, sound, and music generation benchmarks.
PaperID: 3830, Poster
Abstract: Collaborative machine learning is vulnerable to adversarial behaviors during training. Existing defenses typically rely on central coordination or induce high communication costs. We introduce Robust Pull-based Epidemic Learning (RPEL), a scalable and fully decentralized framework that achieves robustness without a central server. Unlike traditional methods, whose communication cost grows as \mathcalO(n^2) with the number of nodes n, RPEL uses a pull-based epidemic communication scheme that scales as \mathcalO(n \log n). By pulling model parameters from small, randomly selected subsets of peers, RPEL significantly lowers the number of required messages while preserving convergence guarantees with high probability. Experimental results demonstrate that RPEL indeed tolerates attacks, attains accuracy comparable with all-to-all communication, and scales efficiently to large networks.
PaperID: 3831, Poster
Abstract: Training smaller language models on reasoning solutions generated by a larger teacher model is the prevailing approach for transferring mathematical reasoning capabilities. However, the student tends to overfit the output distribution of the teacher and generalizes poorly when its own behavior at inference time deviates from the training data. Recent methods attempt to bridge this gap by incorporating error data, yet the corrective trajectories still reflect the reasoning patterns of the teacher, which the student cannot reliably reproduce, leaving the distribution mismatch largely unresolved. We propose a method built on the insight that effective training data should be jointly determined by what the student fails on and what it can realistically learn. Our method first profiles the failure behavior of the student on each problem through repeated sampling, characterizing it in terms of error rate and error consistency to reveal qualitatively distinct failure modes. Each mode demands a fundamentally different training signal. The diagnosed mode then guides candidate generation from the teacher via distinct prompting strategies tailored to each failure type. Among the correct candidates, our method selects the one whose initial reasoning steps after the point of divergence are most reproducible by the student. These first corrective steps constitute the primary bottleneck, while the remaining steps follow an already corrected context and can therefore be continued more readily. This contrasts with standard reranking approaches, which score the trajectory as a whole and thereby dilute the signal at the critical reasoning transition. Experiments on GSM8K, MATH, and three out-of-distribution benchmarks for Llama-3.1-8B and DeepSeek-Math-7B demonstrate that our method achieves superior performance compared with other strong methods. The code and trained weights will be publicly available.
Abstract: Uniform random rotations (URRs) are a common preprocessing step in modern quantization approaches used for gradient compression, inference acceleration, KV-cache compression, model weight quantization, and approximate nearest-neighbor search in vector databases. In practice, URRs are often replaced by randomized Hadamard transforms (RHTs), which preserve orthogonality while admitting fast implementations. The remaining issue is the performance for worst-case inputs. With a URR, each coordinate is individually distributed as a shifted beta distribution, which converges to a Gaussian distribution in high dimensions. Generally, one RHT is not suitable in the worst case, as individual coordinates can be far from these distributions. We show that after composing two RHTs on any d-sized input vector, the marginal distribution of every fixed coordinate of the normalized rotated vector is within \mathcal O(d^-1/2) of a standard Gaussian both in Kolmogorov distance and in 1-Wasserstein distance. We then plug these bounds into the analyses of modern compression schemes, namely DRIVE and QUIC-FL, and show that two RHTs achieve performance that asymptotically matches URRs. However, while two RHTs suffice for scalar quantization, they may be insufficient for Vector Quantization (VQ), which often requires weak correlation across fixed-size blocks of coordinates (as opposed to only marginal distribution convergence for single coordinates). We prove that a composition of three RHTs leads to decaying coordinate covariance. This ensures that any fixed, bounded, multi-dimensional VQ codebook optimized for URRs has the same expected error when using three RHTs, up to an additive term that asymptotically vanishes with the dimension. Finally, because practical inputs are rarely adversarial, we propose a linear-time \mathcalO(d) check on the input's moments to dynamically adapt the number of RHTs used at runtime to improve performance.
PaperID: 3833, Poster
Authors: Yaru Han
Abstract: Protein--nucleic-acid recognition underlies aptamer discovery and nucleic-acid therapeutics, yet computational models often trade mechanistic interpretability for predictive flexibility. Hybrid interpretable--latent architectures risk fallback takeover, where the latent branch silently becomes the true predictor while the explicit pathway remains decorative. We introduce PIGRAM, a patch--motif grammar model that decomposes proteins into local residue patches and nucleic-acid partners into secondary-structure motifs, learning a sign-separated grammar over explicit physicochemical and geometric attributes to preserve rare favorable interactions. Heterogeneous rule families are selectively refined into a coarse-to-fine atlas, and a frozen-grammar residual Transformer provides sparse, bounded corrections without displacing the grammar as the primary predictor. On a strict sequence-pair-disjoint benchmark, PIGRAM achieves a Pearson correlation of 0.494, approaching the unconstrained Transformer (0.511) while exposing signed rules, fine subtypes, and route-gated corrections for each prediction. The learned rules recover bidirectional patch--motif mechanisms, and residual corrections concentrate selectively on grammar-hard regions. In fluorescence ELISA experiments, suppressive-rule relief improved GFP binding, whereas favorable-rule disruption weakened NELF binding, providing direct experimental support for grammar-guided design. Together, these results show that interpretable grammar can serve as both a predictive scaffold and an experimentally actionable design principle for aptamer optimization.
PaperID: 3834, Poster
Authors:
Yun Young Choi, Asung Kil, Sun Woo Park, Minho Lee, Seokhwan KimAbstract: Graph Neural Networks (GNNs) operate under the implicit assumption that the underlying data manifold is Riemannian, where the metric tensor is strictly positive-definite. While effective for homophilic graphs, this inductive bias creates a fundamental geometric mismatch for heterophilic graphs, where edges often signify dissimilarity or structural repulsion. In this work, we propose a paradigm shift from learning on fixed manifolds to learning the manifold itself. We postulate that heterophilic graphs are naturally embedded in pseudo Riemannian manifolds endowed with an indefinite metric, allowing for negative squared distances to model repulsive interactions. To formalize this, we introduce the Graph Signature Index, a spectral invariant that diagnoses the geometric nature of a graph. This index enables our proposed Pseudo-Riemannian Attention Network (PRAT) to dynamically learn an indefinite metric tensor, effectively capturing both attractive and repulsive interactions. By unifying these opposing forces within a single framework, PRAT demonstrates how a subtle shift in the metric signature yields substantial performance gains on heterophilic benchmarks. Our analysis reveals that PRAT’s learned geometry emergently aligns with the spectral signature of the graph, validating our geometric hypothesis.
PaperID: 3835, Poster
Abstract: Discovering underlying logical expressions from data is a critical task for interpretable AI and scientific discovery, yet it remains poorly served by existing research infrastructure. The field of Symbolic Regression (SR) primarily focuses on continuous mathematical functions, while Logic Synthesis (LS) is designed for exact, noise-free specifications, not for learning from incomplete or noisy data. This leaves a crucial gap for evaluating algorithms that can learn generalizable logical rules in realistic scenarios. To address this, we introduce LogicSR, a large-scale and comprehensive benchmark for logical symbolic regression. LogicSR is built from two sources: real-world problems from digital circuits and biological networks, and a novel synthetic data generator capable of producing a diverse set of complex logical formulas at scale. We use LogicSR to conduct a rigorous evaluation of 17 algorithms, spanning classical logic solvers, modern machine learning models, and Large Language Models (LLMs). Our findings reveal that the logical modeling capabilities and generalization robustness of these algorithms significantly depend on task scale and logical complexity, with current cutting-edge LLMs showing limited complex logical reasoning ability. LogicSR provides a robust foundation to benchmark progress, unify evaluation across disparate fields, and steer the future development of powerful neuro-symbolic systems.
PaperID: 3836, Poster
Authors: Yiyang Li, Weixuan Huang, Wei Zhang
Abstract: Although recent training-free acceleration methods have achieved promising speedups for Diffusion Transformers, most rely on historical features, either by directly reusing cached features or by forecasting future features from them. Such a history-to-future design can accumulate errors when the prediction horizon becomes long. To address this issue, we propose FuCa, a Cache-the-Future training-free acceleration framework. Instead of extrapolating future information solely from the past, FuCa caches sparsely computed future features and uses them as online revision signals. FuCa revises the latent-state update and the corresponding velocity, and then recovers skipped timesteps through lightweight interpolation. It requires no extra training or model-specific auxiliary modules, while keeping the number of DiT forward evaluations low. Experiments on FLUX.1-dev, Qwen-Image, Wan2.1-1.3B, and Hunyuan Video show that FuCa achieves up to about 5× speedup for image generation and more than 4.3× speedup for video generation, with higher reconstruction quality than existing training-free baselines.
PaperID: 3837, Poster
Abstract: We study nonparametric regression over Besov spaces from noisy observations under sub-exponential noise. Our goal is to obtain minimax-optimal high-probability bounds for the integrated squared error while adapting to the unknown noise level and the regularity parameters of the underlying Besov class. We introduce a wavelet-based online learning algorithm that sequentially processes noisy gradients and adapts to the gradient noise through an adaptive clipping rule, thus avoiding the need to tune parameters such as the noise variance or gradient bounds. As a by-product of our analysis, we derive high-probability adaptive regret bounds that scale with the \ell_1-norm of the competitor. Finally, in the batch statistical setting, our method is the first to achieve high-probability minimax-optimal estimation rates over Besov spaces while adapting to all problem parameters, including the noise level. Our approach relies on a refined online-to-batch conversion and exploits the structure of the squared loss in combination with self-normalized concentration inequalities.
Abstract: We study contextual bandit problems with correlated arms and access to surrogate reward signals produced by a machine learning model, motivated by applications such as large language model (LLM) routing. Unlike classical contextual bandits that rely solely on bandit feedback and assume conditional independence across arms, our setting allows context-dependent inter-arm correlations and auxiliary reward information that may be noisy or misspecified. We propose algorithms that leverage such surrogate rewards through two complementary designs. A coupled reward-mixing approach pools true and surrogate rewards to accelerate learning when surrogate signals are reliable, while a decoupled prediction-mixing approach maintains separate estimators for bandit feedback and surrogate rewards and adaptively combines their predictions. This decoupling yields robustness to surrogate misspecification, recovering regret guarantees comparable to reward-only bandit methods in the worst case, while achieving improved regret when surrogate predictions are sufficiently informative. We provide theoretical regret analyses for both approaches and evaluate them on LLM routing benchmarks under varying accuracy versus cost trade-offs. The results demonstrate improved sample efficiency and consistently better accuracy–cost trade-offs compared to standard contextual bandit baselines and strong static routing methods.
Authors:
Zili Lin, Wenyao Zhang, Yuyang Zhang, Zekun Qi, Junyan Lin, Hanxin Zhu, Jiaolong Yang, Zhibo Chen, Yao Mu, Xiaokang Yang, Xin Jin, Wenjun ZengAbstract: Imitation learning has achieved remarkable progress in robot manipulation, but its success heavily depends on large-scale demonstration data that are costly to collect. Recent work has therefore explored demonstration augmentation, but they are fundamentally limited in deformable manipulation, where task-relevant object variation is governed by high-dimensional deformations and physics-induced internal constraints rather than low-dimensional pose changes. We present for deformable objects. Instead of perturbing object pose, DeformGen expands the valid state distribution by applying localized physical disturbances, forward-simulating the dynamics, and stabilizing the result to obtain topology-coherent deformable states with physical plausibility. Given these synthesized states, DeformGen further transfers source manipulation trajectories via deformation-field warping, which lifts per-particle displacements into a continuous spatial function to adapt the end-effector trajectory consistently with the deformed geometry. In this way, our method augments both the state distribution and its associated manipulation behavior. Experiments on high-fidelity deformable manipulation benchmarks show that DeformGen consistently improves policy learning over training on the original demonstrations alone and over rigid-style augmentation baselines.
PaperID: 3840, Poster
Authors:
Jed A Duersch, Naïm Es-sebbani, Nathanaël Haas, Zied BouraouiAbstract: We develop a methodology from first principles to adapt transformer structural complexity during training and maximize inference utility per unit compute. Calibrated channel-level penalties compel operators to reorganize into smaller dense tensors, driving entire tensor slices to zero to enable physical removal of channels while preserving density and full GPU throughput. The natural approach, penalizing the norm of operator components that act through each channel, is provably destabilized by gauge freedom between multiplicatively coupled factors. We resolve this pathology with additive symmetric group-lasso penalties that recover a monotone function of product-norms when the network converges to gauge balance. Our equilibrium analysis shows how to calibrate per-channel penalty strength to correctly suppress channels that under-perform in inference utility per unit compute. Under adaptive penalty pressure, the network reorganizes its representations into depth-dependent structural profiles that can be far smaller than the architecture required to learn the task. On polynomial long division over~\mathbbF_31, the method achieves 190× compression with perfect accuracy. On character-level language modeling and masked autoencoding, compressed models match or exceed hand-designed baselines at equal inference cost. Continuous dense compaction accelerates training itself, with step times decreasing as the model compresses. Post-hoc pruning with the same utility metric cannot reach these architectures, confirming that the compressed representations emerge from sustained resource pressure, not from removal of redundancy from a fixed solution.
Abstract: Guiding pretrained flow-based generative models for conditional generation or to produce samples with desired target properties enables solving diverse tasks without retraining on paired data. We present ESS-Flow, a gradient-free method that leverages the Gaussian prior of the source distribution in flow-based models to perform Bayesian inference directly in the source space using Elliptical Slice Sampling. ESS-Flow only requires forward passes through the generative model and observation process, making it applicable even when gradients are unavailable, such as with non-differentiable simulation-based observations or quantization in the generation process. We demonstrate its effectiveness on scientific applications including designing materials with desired target properties and predicting protein structures from sparse inter-residue distance measurements.
Abstract: Runtime oversight for LLM agents is commonly framed as scalar risk prediction: estimate failure likelihood, confidence, or uncertainty, then intervene once the score crosses a threshold. We argue that this framing targets the wrong object for control. The relevant question is not how likely the agent is to fail if it continues, but whether an available intervention would improve the outcome. Two trajectory prefixes can have the same risk estimate while requiring different actions, because one remains recoverable and the other does not. We formalize this mismatch as target error and identify intervention advantage, the expected utility gain from intervening rather than continuing, as the decision object for oversight. To measure this mismatch, we introduce prefix branching, a same-prefix counterfactual protocol that executes candidate actions from identical trajectory states. Across four benchmarks, action-conditioned control yields regime-dependent gains over scalar routing. In a calibration decomposition, recalibrating the same scalar score improves prediction metrics but leaves control regret unchanged, showing that calibration alone does not repair target error. A simple prefix-only action-conditioned controller substantially reduces regret in the strongest interactive regime, from 0.506 to 0.110 on ALFWorld. Gains shrink when interventions are weak or when scalar routing already preserves intervention-relevant information. These results suggest that LLM-agent oversight should move from calibrated risk scoring toward action-conditioned value estimation.
Abstract: The proliferation of autoregressive realstic image generators necessitates robust methods for mitigating misinformation and filtering synthetic images content. To this end, watermarking schemes for image models embed a subtle signal at generation time, enabling downstream verification with a corresponding detector. In this work, we study the vulnerability of such watermarks to removal and forgery attacks. We evaluate existing attacks and, finding them limited, introduce three new ones: (i) a vector-quantized regeneration removal attack, (ii) an adversarial optimization–based attack, and (iii) a frequency injection attack. Our evaluation reveals that both removal and forgery can be effective with access to as little as a single watermarked reference image. Consequently, we find that most AR watermarking schemes do not reliably support content filtering. Finally, we show that even the most robust method, BitMark, admits watermark mimicry—a form of forgery in which authentic images are made to imitate a generator's watermark.
PaperID: 3844, Poster
Abstract: Internet-scale photo collections of real-world landmarks have driven progress in 3D computer vision yet remain highly challenging for modern 3D foundation models (3DFMs) that estimate scene structure in a single feed-forward pass. In this work, we introduce ArchWorld, a benchmark of 3D architectural landmarks with high geographic coverage and rich metadata and use it to conduct a thorough error analysis of current 3DFMs to understand where they fall short and how they can be improved. We examine socially relevant disparities in model error and find that model performance varies by geographic region, but this is heavily confounded by scene size. Leveraging insights from model failings, we introduce an intervention scheme that improves 3DFM performance while simultaneously reducing geographic disparities. Our code and data will be publicly available.
PaperID: 3845, Poster
Abstract: Unified Multimodal Models (UMMs) integrate visual generation and understanding within a shared parameter space, yet existing post-training typically relies on external annotations and supervised post-training data, and rarely couples these two capabilities for iterative improvement. We propose Holistic EvoLution via Intrinsic eXchange (HELIX) for Unified Multimodal Models, a two-stage cyclic post-training framework that forms a double-helix coupling between generation and understanding: generation produces controllable visual evidence while understanding provides semantic judgments. A target-driven Data Bridge connects the two stages and supplies reliable training signals for both generation and understanding updates. In Stage 1, the understanding branch provides a self-evaluated likelihood gain reward to optimize the generative policy for stronger semantic faithfulness. In Stage 2, the generation branch supplies reconstruction-based visual evidence, and we refine the understanding branch with Reconstruction and Prompt-Alignment rewards. Extensive experiments on Bagel, trained at 512px, demonstrate consistent gains on compositional and knowledge-intensive benchmarks, outperforming the baseline by +5.6 on T2I-CompBench, +7.1 on GenEval, +1.54 on DPG-Bench and +0.02 on WISE when evaluated at 1024px.
Authors: Yihan Wang, N. Asokan
Abstract: Memorization in large language models has been studied almost exclusively through prefix-conditioned extraction, a natural choice for autoregressive models. However, diffusion language models (DLMs) can denoise masked tokens at arbitrary positions. Thus, prefix-only probing reveals only one facet of memorization in DLMs and significantly underestimates the risk of training-data extraction. In order to realistically model extractability of training data in DLMs, we introduce infilling extraction, a data-extraction protocol parameterized by an arbitrary binary mask that subsumes prefix-only probing and accounts for the bidirectional inductive bias of DLMs. Instantiating it on LLaDA-8B and Dream-7B across five extraction modes, three training pipelines, and three corpora covering verbatim and partial leakage, we find that mask geometry governs extractability: edge-conditioned masks extract up to three times more verbatim sequences than prefix-conditioned ones, and bidirectional access opens channels inaccessible in autoregressive models. In particular, we show that a realistic adversary with access to training data where personally identifiable information has been redacted, can even achieve higher recall on extracting redacted email addresses from DLMs than from scale-matched autoregressive models. Tunable parameters for decoding measurably affect extraction performance, while a follow-up supervised finetuning stage does not eliminate the prior memorization.
Abstract: Learning flexible motor primitives is a hallmark of skilled motor control. Recent neuroscience theory proposes that motor primitives may be implemented as low-rank perturbations of a shared recurrent network, but leaves open how such a system is learned. We translate this principle into a novel architecture for learning motor skills end-to-end: a shared recurrent core modulated by a bank of residual adapters, each selected by a discrete latent code. Trained on closed-loop biomechanical control, the adapters develop emergent low-rank perturbations of the recurrent dynamics despite no architectural rank constraint, placing task representations in disparate subspaces of the shared core network. A simple high-level policy over the learned options, optimized while the whole network is frozen, sequences the low-rank adapters to produce novel out-of-distribution movements. We demonstrate the ability to generalize to novel motor sequences within the closed-loop control setting, improving on the generalization error of a task-input-conditioned multitask baseline by up to order of magnitude.
PaperID: 3848, Poster
Authors:
Yunlu Chen, Dominik Engel, Peter Wonka, Ivan ViolaAbstract: Pretrained vision transformers provide strong semantic representations, but their token embeddings are not designed to transform predictably under image rotations. This limits their direct adaptation to domains such as satellite imagery, histopathology, microscopy, and other scientific imaging settings, where in-plane orientation is often arbitrary and no canonical pose exists. Existing equivariant adaptation methods for pretrained models typically rely on frame averaging or canonicalization, which can be effective on synthetically rotated natural images but do not teach the adapted model an internal rotation-aware representation. We introduce OrbitLoRA, a lightweight orbit-aware adapter for frozen visual foundation models. OrbitLoRA combines standard LoRA with a typed residual branch that maintains scalar and vector fields over ViT patch tokens. The typed branch is initialized from local steerable image features and harmonic coordinate fields, updated at selected transformer depths, and injected into the frozen token stream through zero-initialized residual gates. During finetuning, paired rotated views supervise the typed state through a transport-and-rotation consistency loss, while the semantic output is trained to remain invariant for the downstream task. This design preserves the rich features of pretrained ViTs while adding a learnable rotation-symmetry prior. Across classification and segmentation tasks, and across DINOv2-, CLIP-, SAM-, and LLaVA-style backbones, OrbitLoRA matches strong LoRA-with-augmentation baselines in high-data regimes and substantially improves adaptation in low-data regimes, where learning the symmetry prior is most valuable.
PaperID: 3849, Poster
Authors:
Jiawei Dong, Zilong Bai, Pengtian zhu, Peng ZhouAbstract: Large Language Models (LLMs) have gained widespread adoption in various natural language processing tasks, but suffer from hallucination issues where they generate unfaithful or inconsistent content. While Reinforcement Learning with Verifiable Rewards (RLVR) has shown promising applications in reasoning tasks, its application to span-level hallucination detection is challenged by localization ambiguity, reward sparsity on hard examples, and uncontrolled length inflation. To address these issues, we introduce Socratic Guided Policy Optimization (SOPO), a multi-turn RL framework for robust hallucination detection. The key innovation of SOPO is a Socratic feedback mechanism, where a teacher model guides the reasoning trajectories during rollout by providing indirect hints inferred from the ground truth. This teacher-guided approach enables the model to generate more effective reasoning trajectories for hard examples, thus alleviating reward sparsity. Furthermore, we incorporate Turn-based Advantage Scaling to penalize verbosity bias, encouraging the model to develop efficient and autonomous reasoning capabilities. Empirical results on RAGTruth demonstrate that SOPO (4B) achieves state-of-the-art performance, surpassing leading proprietary models (e.g., GPT-5) and standard RL baselines, while also exhibiting robust generalization across diverse out-of-domain benchmarks. We release our training and evaluation code alongside the 1.7B and 4B SOPO models at https://anonymous.4open.science/r/SOPO.
PaperID: 3850, Poster
Authors: Fariza Rashid, Duc Van Le, Rahat Masood, Gustavo E Batista, Aruna Seneviratne, Suranga Seneviratne
Abstract: Time series imputation has progressed from statistical and deep learning approaches to diffusion-based models, which have shown strong recent performance. Existing diffusion-based methods typically condition the reverse process using local contextual information from the current or neighbouring windows. Meanwhile, global dataset-level structure often remains implicit, limiting performance when local observations are sparse, noisy, or unrepresentative. To address this issue, we propose ProCTI, a diffusion-imputation framework that augments local conditioning with retrieved global dataset-level priors through learned prototypes. A hybrid conditioning mechanism integrates this global context with local signals during reverse diffusion, enabling more accurate reconstruction under varying missingness scenarios. Experiments across multiple benchmark datasets show that ProCTI outperforms strong baselines overall under random missingness, while remaining competitive under attribute-wise missingness. We further introduce a theoretical framework that distinguishes between local and global conditioning in diffusion-based imputation, showing that combining both can reduce conditional uncertainty and improve imputation quality.
PaperID: 3851, Poster
Abstract: Low-Rank Adaptation (LoRA) is a popular parameter-efficient fine-tuning (PEFT) method that adapts the large model with few trainable parameters. Federated LoRA extends LoRA to federated learning (FL), enabling clients to collaboratively fine-tune a shared model without sharing their raw data by exchanging LoRA parameters. However, due to distributed data heterogeneity, local LoRA updates are highly inconsistent across clients, hindering robust global generalization. Moreover, common direct aggregation introduces aggregated bias, yielding noisy updates and further harming the global generalization. Existing bias-mitigation approaches often rely on symmetrically alternating LoRA modules, overlooking matrix distinct characteristics during training. Therefore, we propose FLoRA-Chef, an asymmetric Federated LoRA that decides when and how to aggregate LoRA components. FLFire adaptively selects the next aggregation target by tracking cross-clientupdate dispersion. FLSauce applies matrix-distinct reweighting to control diversity and prioritize stability. Experiments on various scenarios demonstrate consistent improvements, yielding an average gain of 6.94% over counterparts, validating the benefit of adaptive asymmetric aggregation for federated LoRA.
PaperID: 3852, Poster
Abstract: V2X collaborative perception improves single-vehicle perception by aggregating multi-view features from collaborative agents. However, existing methods typically rely on simple spatial warping to align the collaborative features, which can cause semantic misalignment during fusion. Meanwhile, these deterministic warping-based methods are highly sensitive to the positional noises. In this paper, we propose a plug-in module called Robust Semantic and Spatial Aligner (RSSA), which performs semantic and probabilistic spatial transformations for robust feature alignment. Specifically, a Semantic Transformation (SemT) block modulates feature channels based on the observed relative pose to correct semantic inconsistencies. Subsequently, a Probabilistic Spatial Transformation (P-SpaT) block performs probabilistic warping by sampling from the Gaussian-modeled relative pose space. A Relative Transformation Augmentation (RT-Aug) strategy is further introduced to augment the relative poses and facilitate the training of the RSSA. Extensive experiments on both simulated and real-world datasets demonstrate that RSSA improves the state-of-the-art methods by a large margin under the configuration of both clean and noisy positions.
PaperID: 3853, Poster
Abstract: Low-Rank Adaptation (LoRA) is a popular parameter-efficient fine-tuning method for Large Language Models (LLMs), and practical systems usually need to serve many concurrent LoRA adapters to process different tasks. This is challenging due to the inherent tension between fairness and efficiency, i.e., achieving fairness requires frequent switching among the adapters such that they have similar service quality, while good efficiency requires to minimize adapter switching to reduce memory movement overheads. Existing schedulers can only achieve a fixed trade-off between fairness and efficiency, and their trade-offs are often suboptimal. To tackle this problem, we propose Fairness-aware Efficient LoRA Priority Scheduling (FELPS), which explicitly decomposes fairness and efficiency considerations. In particular, for fairness, FELPS adapts the classic proportional fairness and ensures that the amount of service received by each adapter is proportional to its workload. For efficiency, FELPS prioritizes adapters that are already loaded to GPU memory. To conduct scheduling, FELPS combines to the fairness and efficiency terms with an adjustable weight to achieve arbitrary trade-offs between fairness and efficiency. Evaluations on real workloads show that FELPS achieves a superior trade-off between fairness and efficiency than all baselines.
Abstract: As the era of large language models (LLMs) unfolds, Preference Optimization (PO) methods have become a central approach to aligning LLMs with human preferences and improving performance. We propose Maximum a Posteriori Preference Optimization (MaPPO), a methodology for learning from preferences that explicitly incorporates prior reward knowledge into the optimization objective. Building on the paradigm employed by Direct Preference Optimization (DPO) and its variants of treating preference learning as a Maximum Likelihood Estimation (MLE) problem, MaPPO integrates prior reward estimates into a principled Maximum a Posteriori (MaP) objective. This not only generalizes DPO and its variants, but also enhances alignment by mitigating the oversimplified binary classification of responses. Additionally, MaPPO introduces no additional hyperparameters, and supports preference optimization in both offline and online settings. In addition, MaPPO can be used as a plugin for DPO variants, including widely used SimPO, IPO and CPO, and produce consistent improvements. Extensive empirical evaluations of different model sizes and model series on three standard benchmarks (MT-Bench, AlpacaEval 2.0, and Arena-Hard) demonstrate consistent improvements in alignment performance without sacrificing computational efficiency.
PaperID: 3855, Poster
Authors: Weiye Zhao, Feihan Li
Abstract: Enforcing state-wise safety constraints is critical for the application of reinforcement learning (RL) in real-world problems, such as autonomous driving and robot manipulation. However, existing safe RL methods often enforce state-wise constraints only in expectation, while hard state-wise guarantees typically require strong dynamics assumptions. The former can still allow rare but severe violations, while the latter is impractical in model-free stochastic systems. We instead target high-probability control of the maximum instantaneous cost along a trajectory. To accomplish this goal, we propose Absolute State-wise Constrained Policy Optimization (ASCPO), a model-free policy search algorithm whose exact update provides a distribution-free high-probability certificate for state-wise constraint satisfaction in stochastic systems. The guarantee holds under finite variance and does not require Gaussian, unimodal, or light-tailed cost distributions. We demonstrate the effectiveness of our approach by training neural network policies for extensive robot locomotion tasks, where the agent must adhere to various state-wise safety constraints. Our results show that ASCPO most consistently reduces state-wise safety violations while maintaining competitive reward across challenging continuous control tasks.
PaperID: 3856, Poster
Abstract: Long-term time-series forecasting requires temporal dependency modeling that remains stable and generalizable under temporal distribution shifts. However, real-world time series are often noisy and non-stationary, making deep learning models prone to fitting incidental temporal interactions in the training data. A key question is whether there exist temporal relationship structures that recur across different temporal positions and variables. We analyze temporal relationship matrices on commonly used time-series datasets and find that samples from different temporal positions and different variables exhibit shared dominant temporal relationship structures. This indicates that temporal relationship structures contain patterns that recur stably across samples and variables. To learn such shared structures, we propose Temporal Structure Network (TStruct), a simple yet effective network for directly learning shared and generalizable temporal dependencies. Specifically, TStruct directly parameterizes a set of shared temporal dependency matrices and generates normalized combination weights using dynamic sample information and static variable priors. These weights represent a soft allocation of each sample-variable pair over shared temporal dependency structures, enabling the model to generate adaptive temporal dependency matrices under shared structural constraints, rather than fitting incidental interactions in a fully unconstrained manner. Notably, this simple structure-aware design consistently outperforms several strong baselines across all datasets, achieving superior accuracy while maintaining high computational efficiency. Our code is available at \urlhttps://anonymous.4open.science/r/TStruct-NeurIPS2026-1173.
PaperID: 3857, Poster
Abstract: Humanoid robots capable of versatile whole-body interaction with everyday objects represent a central goal of embodied intelligence. Existing approaches often rely on motion-reference inputs or task-specific rewards, coupling policies to particular motion scripts, object geometries, and contact timings. This limits geometric generalization and flexible interaction control, where simple commands specify motion intent while object geometry determines how contact should be executed. We introduce LessMimic, a framework for versatile motion-free-at-inference humanoid–object interaction. At inference, a single whole-body policy is steered by root commands and an interaction-type flag instead of motion-reference inputs, and is conditioned on a compact interaction representation built from short histories of distance-field-derived surface distances and local surface directions. This geometry-conditioned interface enables command-steered control while adapting contact behavior to local object geometry. Through visual distillation, LessMimic further enables egocentric-depth-based deployment for MoCap-free sensing. Across object scales from 0.4× to 1.6×, LessMimic maintains robust performance over four interaction tasks, including PickUp, SitStand, Push, and Carry, where motion-conditioned baselines degrade sharply away from the training scale. Beyond single-task evaluation, LessMimic further supports sequential composition of heterogeneous interaction skills, attaining 62.1% success on five-task sequences and retaining 23.5% success at 15 task instances. By grounding interaction control in local geometry rather than motion references, LessMimic provides a path toward geometry-generalizable, command-steerable humanoid–object control with heterogeneous skill composition.
PaperID: 3858, Poster
Abstract: Vision-language models (VLMs) are increasingly used as perception and reasoning modules for embodied intelligence, yet they remain brittle when understanding fine-grained physical processes in videos. While existing models often recognize the overall task, they frequently fail to recover the intermediate actions and result states that determine how an embodied interaction actually unfolds. We argue that this limitation stems in part from a supervision bottleneck: conventional video-language data is typically constructed through post-hoc annotation, where captions are written after video collection and often provide only coarse, loosely aligned descriptions of dynamic physical events. We propose , a generation-for-understanding framework that reverses this data flow. Instead of captioning videos after they are observed, Gen4Understand first represents embodied interactions as structured action scripts composed of atomic events and state deltas, then synthesizes videos conditioned on these scripts and real start-end frame anchors. A verifier-calibrator scores each generated candidate using its conditioning script, endpoint anchors, and video content, filtering for semantic consistency, endpoint fidelity, physical plausibility, temporal coherence, and visual quality. The retained atomic clips are further merged into action-, subtask-, and task-level supervision and converted into captioning and question-answering data for VLM adaptation. Experiments across public embodied video splits, distractor-rich hallucination tests, simulated VLA benchmarks, and real-robot tasks show that Gen4Understand consistently improves fine-grained action understanding, result-state recognition, temporal ordering, hallucination resistance, and downstream embodied control. These results suggest that controllable video generation can serve not merely as data augmentation, but as a mechanism for creating calibrated process-level supervision for embodied intelligence.
PaperID: 3859, Poster
Abstract: Video diffusion models have recently achieved remarkable visual fidelity. However, when handling complex instructions with multi-action sequential prompts, existing models frequently suffer from severe action omission or state freezing. We attribute this bottleneck to two fundamental flaws: i) semantic entanglement, where appearance and motion concepts are heavily intertwined within the text embedding space; and ii) temporal attention leakage, wherein the global nature of cross-attention along the temporal dimension biases the model towards the most visually salient action tokens throughout the entire video. To break this bottleneck, we propose SAX, a novel diffusion transformer framework designed for the synergistic generation of appearance and dynamic actions. First, we introduce a dual-stream attention mechanism that explicitly decouples dynamic action interactions from the static appearance generation process. Besides, to establish a precise temporal-action mapping, we design the Layer-adaptive Action Choreographer. By integrating a shared Global Temporal Table and an Action Ordinal Table, this module dynamically generates a temporal-action alignment map for each layer. Furthermore, to overcome the implicit convergence difficulties inherent in temporal-action alignment optimization, we innovatively employ a Vision-Language Model as a temporal-semantic supervisor. Through distribution distillation, we directly inject the VLM's powerful temporal-action grounding priors into the learning of the alignment maps. Extensive experiments on action generation and prediction demonstrate that SAX substantially outperforms existing models across core metrics.
Abstract: Building state-of-the-art (SOTA) predictive models for drug discovery requires expensive search over tools, architectures, and training strategies. Current LLM-based agents can find SOTA solutions through extensive trial and error, but they do not retain the experience accumulated along the way and therefore pay the full search cost on every new task. We propose DrugSAGE (Self-evolving Agent Experience), a framework that accumulates and reuses experience across tasks to build SOTA drug discovery models efficiently. DrugSAGE maintains a cross-task memory of verified skills, statistical evidence about effective strategies, and a record of recurring errors and their fixes. In some cases, DrugSAGE transfers a working solution directly without test-time search. In 33 molecular property prediction tasks, DrugSAGE ranks first among nine SOTA agents in a single-task setting. With memory accumulated from 16 smaller tasks, DrugSAGE achieves a averaged normalized score of 0.935 on 17 held-out tasks in a cross-task evaluation setting and outperforms all baseline agents by 10-30% in a zero-test-time search regime. In summary, our work shows the advantage of cross-task memory for efficient SOTA model development in drug discovery.
PaperID: 3861, Poster
Authors:
Ahmet Faruk Cetinkaya, Enes Ağırman, Cem TekinAbstract: Standard ensemble methods can improve robustness by mitigating individual model stochasticity, but they frequently fail to generalize beyond the training distribution, leaving systems vulnerable to out-of-distribution (OOD) data. We introduce the Robust Satisficing Ensemble (RSE), a scalable probabilistic aggregation framework designed to secure model combinations by minimizing their fragility against distributional changes. By bridging a robust satisficing objective with a quadratic risk surrogate, RSE derives a closed-form analytical inner adversary. This fundamental reformulation enables highly efficient weight updates without the computational bottleneck of standard minimax optimization, crucially decoupling optimization complexity from sample size to scale effectively to large datasets. In particular, we prove a geometrically decaying truncation error for the fragility approximation via Lagrange-Bürmann expansion and show that the proximal RSE updates converge to a stationary point. We evaluate RSE across a comprehensive spectrum of distribution shifts, including adversarial label shift (SST-5, TREC), natural subpopulation and temporal shifts (CivilComments, HuffPost). Our results demonstrate that RSE consistently outperforms strong robust baselines, including GroupDRO, IRM, and SRM. These findings establish RSE as a practical, theoretically grounded safeguard for reliable model aggregation in dynamic environments.
Abstract: The statelessness of foundation models bottlenecks agentic systems’ ability to continually learn, a core capability for long-horizon reasoning and adaptation. To address this limitation, agentic systems commonly incorporate memory modules to retain and reuse past experience, aiming for continual learning during test time. However, most existing memory designs are human-crafted and fixed, which limits their ability to adapt to the diversity and non-stationarity of real-world tasks. In this paper, we introduce ALMA ( gentic systems), a framework that meta-learns memory designs to replace hand-engineered memory designs, therefore minimizing human effort and enabling agentic systems to be continual learners across diverse domains. Our approach employs a Meta Agent that searches over memory designs expressed as executable code in an open-ended manner, theoretically allowing the discovery of arbitrary memory designs, including database schemas as well as their retrieval and update mechanisms. Extensive experiments across four sequential decision-making domains demonstrate that the learned memory designs enable more effective and efficient learning from experience than state-of-the-art human-crafted memory designs on all benchmarks. When developed and deployed safely, ALMA represents a step toward self-improving AI systems that learn to be adaptive, continual learners.
PaperID: 3863, Poster
Abstract: Any-color control---the ability to specify objects by arbitrary 24-bit hex values rather than coarse color words---is a practical yet challenging requirement for image generation and editing, where users often need exact object colors, brand colors, or carefully colored compositions. Prior work has explored color generation, color editing, and image colorization, but most methods rely on task-specific modules or training-free inference-time techniques. This fragmentation makes any-color control difficult to use as a unified text-prompt capability. In this work, we present Paint-Anything, a unified model for any-color controllable generation and editing that directly embeds 24-bit hex colors into text prompts and supports diverse color-control tasks within one framework. Our key insight is that modern text-to-image models already contain partial color understanding, but this capability remains unstable unless numeric color tokens are aligned with localized pixel colors. To this end, we construct Paint-500K, a 500K color-control dataset, and combine it with a high-noise gate for lightweight RGB alignment. We further introduce PCBench, a diagnostic benchmark suite for object-level hex color fidelity, with two parts: PCBench-T2I for generation and PCBench-Edit for editing. Experiments and ablations show that \ours substantially improves any-color generation and editing over the base model, suggesting that arbitrary hex-color control can be learned as a native prompt-following behavior of modern image models.
PaperID: 3864, Poster
Abstract: Real-world networks typically exhibit small-world properties: dense local clustering yet efficient global propagation. However, current methods struggle to reconcile this hierarchy: Graph Neural Networks (GNNs) suffer from over-smoothing in long-range modeling, while sequence/state-space models often lack the inductive bias to preserve complex graph topology. To bridge this gap, we propose Graph Encoding Mamba (GEM), a dual-scale architecture for graph-level hierarchical representation learning. First, GEM employs a dual-branch encoder that synergizes the local sensitivity of GNNs with the selective scanning of Mamba, effectively capturing both local topology and global shortcuts. Second, a modularity-aware hierarchical pooling mechanism is designed to retain salient community structures during graph coarsening. Crucially, to mitigate structural redundancy from multi-scale fusion, we incorporate an Information Bottleneck objective to distill task-relevant structural patterns from redundant features, encouraging compact, task-relevant representations and supporting intrinsic interpretability. Extensive experiments on graph-level benchmarks demonstrate that GEM improves performance across biological, social, and brain network datasets. We further provide a lightweight node-level adaptation on long-range benchmarks as a boundary analysis of GEM's transferability. Code is available at \urlhttps://anonymous.4open.science/r/GEM-C34D.
Abstract: We study the approximation of nonlinear operators between function spaces by transformers. Our approach is to lift functions to measures supported on their graphs and leverage a recently introduced measure-theoretic view of transformers. A function h is represented by its graph measure \gamma_h, with finite tokens \(x_j,h(x_j))\_j=1^N being its empirical approximations. We show that this framework elegantly models discretization refinement via convergence of measures and provides a natural setting for operator learning. Within this framework, we introduce function graph transformers, a graph-preserving subclass of measure-theoretic transformers that maps graph measures to graph measures, which is to say that outputs remain single-valued functions. Crucially, this additional structure does not reduce generality: we prove that the resulting graph-preserving maps can be approximated by finite compositions of standard softmax self-attention layers and pointwise MLPs, yielding universal approximation results for broad classes of nonlinear operators. Unlike existing theoretical approaches to operator learning with transformers, the measure-theoretic framework also accommodates regularized negative-order Sobolev inputs for which discretization invariance is particularly challenging, as well as query points on different output domains. Overall, function graph transformers provide a continuum viewpoint and mathematical toolkit for transformer-based operator learning, clarifying the roles of positional embeddings, graph structure, regularization, and ensuring consistency across discretizations.
Authors:
Meenakshi Krishnan, Pranav Pulijala, Ke Chen, Haizhao Yang, Ramani DuraiswamiAbstract: Operator learning for partial differential equations (PDEs) on arbitrary geometries builds fast neural surrogates for large-scale simulation. Although recent geometry-adaptive neural operators have made substantial progress, they are mainly designed for forward problems in which inputs and outputs share the same spatial domain. This limits their applicability for boundary value problems (BVPs) and inverse problems, where inputs and outputs may live on different domains. We introduce the Geometry-Adaptive Integral Autoencoder (GAIA), an operator learning model that encodes the domain boundary and the interior field distribution into geometry tokens, and conditions integral transform layers on these tokens via cross-attention, allowing the kernel to adapt locally to geometric features. This yields a single architecture for forward (including BVPs) and inverse problems on arbitrary domains in one pass, without retraining, iterative optimization, or graph construction. We evaluate GAIA on seven 2D and 3D benchmarks, four of which are new or substantially extended benchmarks for inverse problems and BVP: electrical impedance tomography, optical tomography, 3D Darcy flow on varying geometries, and a modified setting of Poisson BVP on mechanical components benchmark (MCB). GAIA sets new state-of-the-art results on every inverse and BVP task, reducing median relative L^2 error by 64% on airfoil flow reconstruction and 27% on EIT relative to the next best amortized method, and outperforming all baselines on every shape category of MCB. On other forward problems, GAIA is competitive with specialized solvers while maintaining stable accuracy across point resolutions on which transformer-based baselines degrade.
Abstract: Learning to Defer (L2D) enables a model to predict autonomously or defer to an expert, but prior work largely assumes flat label spaces. We study the first L2D setting with hierarchical multi-label decisions, motivated by medical-imaging workflows in which findings are organised by clinical taxonomies. In this setting, deferral is a delegation action rather than a label assignment, so treating it as an independent per-label decision can produce \emphdeferral incoherence, including taxonomic contradictions, delegation violations, and deferrals of labels already implied by the model's own assertions. We formalise coherent hierarchical deferral under a Selective-Exclusion handoff contract, characterise the Bayes-optimal coherent deferral rule, and show that even nodewise Bayes L2D can be action-incoherent. We then propose two remedies: exact coherent projection, a dynamic-programming decoder over the coherent action set, and Taxonomic Belief Propagation (TBP) with Recursive Policy Optimisation (RPO), a contract-aware joint action model trained through the same recursion used at inference. Across real-reader and controlled-expert medical-imaging benchmarks, naïve binary-relevance L2D exhibits non-trivial incoherence. Projection removes it exactly, and fast TBP+RPO drives incoherence near zero while retaining strong utility.
PaperID: 3868, Poster
Abstract: Neural Operators (NOs) are powerful architectures for learning mappings between function spaces. While most advances focus on refining kernel parameterizations over the d-dimensional physical domain, the evolution of lifted embeddings remains underexplored, which often drives models toward computationally expensive embedding-scaling designs to improve approximation. In this paper, we introduce an auxiliary function dimension that models embedding evolution in operator form, thereby reformulating the NO pipeline in d+1 dimensions. We instantiate this framework via Fourier-based operators acting jointly on the physical and auxiliary domains, yielding a basis-diversified auxiliary evolution module as an alternative to brute-force embedding scaling. Across more than ten increasingly challenging benchmarks, ranging from the 1D heat equation to the highly nonlinear 3D Rayleigh–Taylor instability, our model consistently achieves the lowest relative L_2 error among the evaluated baselines. Crucially, this advantage is empirically supported by (1) controlled budget-aware comparisons against scaled and ablated baselines; (2) robustness under mixed-resolution training and super-resolution inference; and (3) zero-shot generalization to unseen temporal regimes. In addition, we present a broader set of design choices for lifting and recovery operators, demonstrating their impact on our model’s predictive performance.
PaperID: 3869, Poster
Abstract: Linear-width convergence for two-layer ReLU networks is known in random-feature or conjugate kernel (CK) regimes whose analyses exploit nearly fixed hidden features. However, the corresponding neural tangent kernel (NTK) regime, where hidden-layer training is active, still suffers from substantially larger width requirements. The obstacle is the ReLU-induced kernel shift: activation-pattern changes can induce empirical NTK drift that breaks the contraction argument. We show that an output-dominant unbalanced Gaussian initialization creates an NTK-dominated training regime in which this kernel shift remains controllable. This single mechanism yields near-linear width, Nesterov acceleration, and low-rank adaptivity for two-layer ReLU networks with vector-valued outputs and a shared first layer. We prove that gradient descent (GD) achieves linear convergence for networks with \tilde\Omega(Nn/\lambda) neurons, where N is the sample size, n is the output dimension, and \lambda is the standard smallest eigenvalue of the limiting NTK. Within the same framework, Nesterov's accelerated gradient (NAG) attains a provable speedup without sacrificing near-linear width, improving the iteration complexity from O(n\kappa\log\frac1\epsilon) to O(\sqrtn\kappa\log\frac1\epsilon), where \kappa is the limiting NTK condition number. Finally, our analysis establishes low-rank adaptivity: by introducing a sketching step at initialization and a subspace analysis, the width requirement reduces to \tilde\Omega(Nr\kappa^2(\mathbfY)/\lambda) for responses \mathbfY of rank r \ll n, matching the ambient-dimensional result in the leading polynomial dependence, up to polylogarithmic factors, with n replaced by r when \kappa(\mathbfY)=O(1). In contrast to CK/random-feature analyses that exploit nearly fixed hidden features, our proof controls hidden-layer NTK contributions, ReLU activation-pattern changes, and kernel shift; this control enables both acceleration and low-rank adaptivity in the NTK-dominated regime.
Abstract: Large vision-language models (VLMs) often exhibit degraded safety alignment when visual inputs are integrated. Even when a text prompt is explicitly harmful, adding an image can substantially increase the jailbreak success rate. In this paper, we observe that VLMs can clearly distinguish benign, refusal, and jailbreak samples in their representation space, and that jailbreak responses often contain safety warnings. These observations lead to the recognize-but-fail-to-refuse hypothesis: VLM jailbreaks do not arise from a failure to recognize harmful intent, but instead occur because adding an image induces a representation shift that steers the sample into a distinct jailbreak state where refusal is not triggered. To quantify how this image-induced representation shift contributes to jailbreak behavior, we define a jailbreak direction and characterize the jailbreak-related representation shift as the projection of the total image-induced shift onto this direction. Our analysis shows that the jailbreak-related shift reliably characterizes jailbreak behavior, providing a unified explanation for diverse jailbreak scenarios. Finally, we propose JRS-Rem, a defense method that enhances VLM safety by removing the jailbreak-related shift at inference time. Experiments show that JRS-Rem significantly improves VLM safety across multiple scenarios while preserving utility on benign tasks.
PaperID: 3871, Poster
Authors: Jingtian Ji, Xuefeng Liu, Matthew Walter
Abstract: Learning from multiple imperfect oracles is a practical way to improve reinforcement learning (RL) under limited interaction. However, most contemporary robust multi-oracle methods are on-policy and equire frequent fresh rollouts, which limits sample efficiency in domains where data collection is expensive. We propose Oracle-Aggregated Policy Improvement (OPI), an off-policy framework for policy improvement with multiple suboptimal oracles. OPI augments the oracle set with the learner, constructs an oracle-guided behavior policy through candidate-action pooling and source-aware critic scoring, and trains the learner using a generalized advantage-weighted behavior cloning objective combined with deterministic policy gradients. The method is motivated by a tractable proxy for the max-advantage objective and a policy-improvement relative to a strong oracle baseline. We evaluate OPI on diverse tasks across MetaWorld, the DeepMind Control Suite, and customized Box2D environments with heterogeneous oracles. Across sparse-reward manipulation, dense-reward locomotion, and customized control settings, OPI is strongest in sparse-reward and complementary-oracle regimes and remains competitive with representative baselines in dense-reward environments. Additional analyses show that oracle composition is most beneficial when oracle skills are complementary and that estimating oracle-specific action values is particularly helpful in sparse-reward regimes
PaperID: 3872, Poster
Abstract: Simultaneous Speech Translation (SimulST) requires balancing translation quality and latency. While Large Language Models (LLMs) excel in offline translation, current LLM-based SimulST systems predominantly rely on prefix-incremental generation. This paradigm suffers from redundant Key-Value cache recomputation and necessitates external Voice Activity Detection to process unbounded streams. To overcome these bottlenecks, we propose an end-to-end interleaved read-write framework built upon an omni-modal LLM. By interleaving speech chunks and text tokens, our architecture achieves full KV cache reuse, natively supporting segmentation-free, infinite-length inference. Furthermore, to learn optimal streaming policies without rigid offline alignments, we introduce a novel Actor-Critic self-enhancement pipeline. An RNN-T Actor dynamically samples read-write trajectories, while an LLM-based Critic evaluates these paths via intrinsic semantic and latency rewards, distilling the optimal policy back to the Actor. Extensive experiments on sentence-level (FLEURS, CoVoST2) and document-level (MuST-C, ACL 60/60) benchmarks demonstrate that our model establishes new state-of-the-art quality-latency Pareto frontiers, substantially outperforming existing cascaded and static-alignment baselines.
PaperID: 3873, Poster
Authors: Seungmin Oh, Seunghun Kang, Jongbin Ryu
Abstract: Vision-language models (VLMs) achieve strong zero-shot transferability but remain vulnerable to target-domain shifts at inference time. Test-time adaptation (TTA) offers a practical remedy, yet most existing VLM-TTA methods follow a prediction-side adaptation paradigm. They use test samples to adjust logits, prototypes, caches, priors, or feature statistics, often incurring additional computational overhead. In this paper, we take a different perspective and reframe VLM-TTA as candidate verification rather than prediction adjustment. We propose Test-Time Correction (TTC), a hypothesis-based correction framework guided by a simple principle: hypothesize, reconstruct, correct. Given a test feature and its top-k candidate labels, TTC treats each candidate label as a hypothesis, reconstructs the feature within the corresponding latent subspace stored in a memory bank, and measures the resulting divergence shift. This shift quantifies how much the candidate subspace and its relations to other candidates change after the hypothetical insertion of the test feature. A correct candidate hypothesis induces only a small shift, whereas an incorrect one perturbs the subspace more strongly. TTC therefore corrects the prediction by selecting the candidate with the minimum aggregated divergence shift. This parameter-free candidate verification mechanism avoids iterative optimization and provides a favorable accuracy-efficiency trade-off. Across five TTA settings and 15 benchmark datasets, including zero-shot classification, domain generalization, few-shot classification, base-to-novel generalization, and cross-dataset evaluation, TTC consistently improves accuracy over state-of-the-art VLM-TTA methods while achieving up to 2× speedup and over 3× lower memory usage.
PaperID: 3874, Poster
Authors: Kevin Pfisterer, Quentin Hillebrand, Vorapong Suppakitpaisarn
Abstract: We propose an algorithm for counting below-threshold triangles in weighted graphs under local weight differential privacy. Although many prior studies have considered the setting in which the graph topology is public and only the edge weights are sensitive, to the best of our knowledge, this is the first work to study this privacy notion in the local model. Building on a two-round protocol for locally differentially private triangle counting, we exploit the public graph topology to design a novel algorithmic framework. This leads to significant improvements in both accuracy and scalability. In particular, when the input graph is planar, our algorithm eliminates the covariance arising from distributed triangle counting at nodes; for graphs with bounded degeneracy, it significantly reduces this covariance. Since covariance is the dominant source of error in the counting task, our method achieves accuracy that closely aligns with the lower bound. We also present an efficient algorithm for the computation of the smooth sensitivity and provide experiments that quantify the trade-off between the biased and unbiased variants of our estimator and demonstrate the effectiveness of the proposed improvements.
PaperID: 3875, Poster
Abstract: Solving inverse problems on video requires priors over both frame appearance and temporal dynamics. Many recent approaches combine pretrained image diffusion priors with motion guidance, sidestepping the costs of training large video models. Temporal consistency is often enforced via photometric losses or warping-based constraints, which implicitly assume that motion explains frame-to-frame or latent-to-latent changes up to small deviations, an assumption that fails under occlusions, lighting changes, and non-rigid deformation. In this paper, we propose to learn the coupling between motion and frame appearance as a generative prior. We learn the distribution of motion-induced appearance residuals with a conditional diffusion model and use it to regularize reconstruction in video inverse problems. Our motion-appearance prior acts as a plug-in regularizer for diffusion-based video reconstruction, preserving the benefits of image diffusion models while remaining compatible with existing appearance regularization techniques. Experiments on dynamic video inverse problems show that, depending on the task, reconstruction with a learned motion-appearance coupling prior matches or outperforms feature-consistency and noise-warping baselines.
Abstract: Large language models (LLMs) have shown strong performance on isolated psychiatric tasks, including dialogue, diagnosis, and treatment planning, yet existing benchmarks rarely simulate complete psychiatric clinical encounters. We introduce , a virtual evaluation environment for LLM-based psychiatric clinical encounters. MentalHospital instantiates the Subjective Interviewing, Objective Examination, Diagnostic Assessment, and Treatment Planning (S.O.A.P.) workflow, using skill-augmented standardized patients constructed from 1,193 de-identified psychiatric electronic health record (EHR) cases spanning all major ICD-11 categories and 76 disorders. Each encounter is assessed through a dual-track protocol that combines objective comparison against EHR-derived references with subjective assessment of clinical process quality. To scale specialist judgment, we develop , five domain-specific evaluators covering communication empathy, interviewing professionalism, clinical-note quality, diagnostic rigor, and treatment appropriateness, trained with rubric-grounded SFT and expert-guided DPO. Survey responses from 22 clinicians support MentalHospital's clinical fidelity (3.88/5), while MentalEval achieves strong expert alignment with an average QWK of 0.944. Benchmarking shows that even the strongest LLM trails clinicians by 37.28 percentage points in objective psychiatric competence, with mental status assessment as a key bottleneck.
PaperID: 3877, Poster
Authors: Robert Kasumba, Dennis Barbour, Chien-Ju Ho
Abstract: Adaptive systems often need to make task-specific decisions about a person from limited prior evidence: a tutor may need to anticipate how a learner will approach a new problem, a game may need to adapt when a player enters a new level, and a human-AI system may need to infer whether a partner will persist with a plan or switch goals. These decisions depend on person-level behavioral tendencies that shape how people solve related tasks, but such tendencies are difficult to infer from standard behavioral evidence. One approach is to use aggregate outcome summaries, such as scores, completion rates, or productivity measures; these summaries are compact and available across tasks, but can collapse distinct behavioral processes into similar outcomes. Another approach is to use process-level traces, which record how behavior unfolds over time; however, process modeling within a single task can entangle stable person-level tendencies with task-specific layout, timing, and affordances. In this work, we study early cross-task behavioral inference: whether partial process traces from multiple source tasks can reveal transferable person-level behavioral structure that predicts strategy in a held-out target task. We introduce Process-Level Latent Variable Models (PLVM), which encode task-specific traces and fuse them into a shared person-level latent representation for cross-task prediction. In PowerWash Simulator, a naturalistic telemetry dataset of real human gameplay, PLVM uses partial traces from two earlier cleaning tasks to predict whether a player will exhibit locally persistent Zone Planner behavior or frequent Zone Hopper behavior in the held-out Fire Station level. Controlled simulations with known latent behavioral types show that cross-task fusion helps when different source tasks reveal complementary dimensions of a shared latent behavioral process. These results suggest that process-level cross-task modeling can support early prediction of target-task behavioral strategy when waiting to observe sufficient target-task behavior is impractical.
Abstract: Construction-based neural routing solvers, typically composed of an encoder and a decoder, have emerged as a promising approach for solving vehicle routing problems. While recent studies suggest that shifting parameters from the encoder to the decoder enhances performance, most works restrict the decoder size to 1–3M parameters, leaving the effects of scaling largely unexplored. To address this gap, we conduct a systematic study comparing two distinct strategies: scaling depth versus scaling width. We synthesize these strategies to construct a suite of 12 model configurations, spanning a parameter range from 1M to ~150M, and extensively evaluate their scaling behaviors across three critical dimensions: parameter efficiency, data efficiency, and compute efficiency. Our empirical results reveal that parameter count is insufficient to accurately predict the model performance, highlighting the critical and distinct roles of model depth (layer count) and width (embedding dimension). Crucially, we demonstrate that scaling depth yields superior performance gains to scaling width. Based on these findings, we provide and experimentally validate a set of design principles for the efficient allocation of parameters and compute resources to enhance the model performance.
Abstract: Evaluating large language models increasingly relies on LLM-as-a-judge protocols, but such evaluations remain costly: different judges have different prices and reliabilities, and the difficulty of each prompt--response pair can vary substantially. This raises a basic allocation question: under a fixed budget, how should one distribute evaluation queries across heterogeneous judges and instances to obtain the most accurate score estimates? We formalize this question as \emphbudgeted heteroskedastic multi-judge estimation. Given K prompt--response pairs, J judges with known costs, and unknown query--judge variances, the goal is to estimate a bounded score vector while minimizing an \ell_p-error. Our first contribution is to analyze the inverse-variance weighted estimator (IVWE) and to derive the oracle allocation that minimizes its error rate. Since this allocation depends on the unknown variances, we then address the practical unknown-variance setting by proposing Est-IVWE, an adaptive algorithm that constructs and leverages optimistically biased variance estimates to stabilize the empirical allocation. We prove that Est-IVWE matches the oracle IVWE rate up to lower-order terms in the budget. Our second and central theoretical contribution is a matching local minimax lower bound, which establishes the instance-optimality of the proposed algorithms. A key technical insight is that Fano-type high-probability arguments are too coarse for this problem: their packing construction loses the local variance structure that governs the optimal allocation. We instead use an Assouad-type in-expectation argument, based on local perturbations, which preserves this structure and yields the sharp allocation-dependent lower bound. Finally, we validate our adaptive approach on synthetic benchmarks and the real-world HelpSteer2 dataset, demonstrating significant gains over naive uniform allocation strategies.
PaperID: 3880, Poster
Authors: Wei Zhang
Abstract: Existing federated learning setups assume either a single global model or a fixed decomposition with one globally shared parameter block and one client-specific private block. In this paper, we study a generalized federated optimization setting where each client optimizes a local coordinate block, and two clients interact only when their blocks overlap. We encode these coordinate overlaps by a graph on clients, called the nerve skeleton. Based on this structure, we propose Nerve-Skeleton Message Passing (NSMP), a two-phase protocol on a spanning tree: a leaf-to-root pass composes reduced objectives by partial minimization, and a root-to-leaf pass reconstructs a joint assignment. We further show that, under the tree-elimination schedule in the paper, NSMP is equivalent to minimizing the joint objective, where the leaf-to-root phase evaluates the optimal objective value, and the root-to-leaf phase recovers a globally optimal assignment. Experiments show that NSMP solves federated optimization with overlapping parameters using only efficient neighbor-to-neighbor message passing.
Abstract: Modern image editors excel at semantic manipulation and visual synthesis, yet remain limited in precise spatial control, motivating the development of drag-based editing. However, existing drag-based methods often struggle to balance drag accuracy with natural, plausible, and intent-aligned generation. We propose MoRe-Drag, a motion-grounded drag-based editing method. Our key insight is to treat pixel-space warping as coarse motion evidence, and to inject this evidence into the generative sampling trajectory. Specifically, MoRe-Drag performs region-aware latent recomposition over refinement, inpainting, and anchor regions, coupled with stage-adaptive conditioning that progressively shifts from motion-grounded structure formation to semantic refinement. We further support an instruction-free interface by adapting the MLLM-based text encoder for drag-aware instruction inference. Experiments on DragBench-SR and DragBench-DR show that MoRe-Drag substantially improves drag precision over strong base editors and achieves superior drag accuracy among SOTA drag-based methods, while delivering strong semantic consistency and visually realistic results. Code and dataset will be publicly released.
Abstract: In overparameterised classification, training data can be linearly separable even when the underlying distribution is not. In this setting, gradient descent (GD) on the logistic loss diverges in norm while converging in direction to a max-margin interpolating classifier, whose implicit bias can be statistically suboptimal. In this work, we show that early stopping can overcome this suboptimality: in a Gaussian mixture model with label-flipping noise, GD stopped at an appropriate oracle time achieves minimax-optimal excess zero-one risk for covariance spectra with fast and continuous decay, including polynomial and exponential spectral decays. Our analysis combines a sharp upper bound for the early-stopped iterate with a matching statistical lower bound over arbitrary classifiers, yielding optimal rates that are validated by experiments. A central technical contribution is a new calibration result that converts excess logistic risk into excess zero-one risk; it handles the model misspecification induced by the label-flipping noise, and removes the square-root rate in standard bounds. We also establish a lower bound for linear interpolators, showing that interpolation can require exponentially more samples than early stopping to achieve the same excess risk.
PaperID: 3883, Poster
Authors:
Huipeng Ma, Dandan Song, Shaonan Ma, Changzhi Zhou, Jun Yang, Yuhang Tian, Luan Zhang, Chenhao Li, Xudong Li, Fang Xi, Guangyuan FengAbstract: In retrieval-augmented generation (RAG), knowledge conflicts arise when retrieved contexts disagree with each other or with the model's parametric memory. Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a standard approach to address such conflicts. However, existing RLVR-based methods are typically formulated under a single-conflict assumption. As a result, the type with stronger advantage signals dominates policy updates, leaving the other under-optimized and thus yielding imbalanced performance across conflict types. To address this challenge, we propose ReCon, a type-aware RLVR framework for balanced learning across Inter-Context and Context-Memory conflicts. Specifically, ReCon computes a step-level imbalance signal from type-wise cumulative advantages and triggers resampling on the weaker conflict type. Instead of resampling from scratch, ReCon branches from informative failures. It selects an informative failed trajectory using uncertainty and closeness signals, then branches at an entropy-jump point, concentrating additional rollouts on critical reasoning decisions. Experiments show that ReCon yields stronger and more balanced performance across knowledge-conflict and multi-hop QA benchmarks, and also generalizing well under mixed-conflict settings.
PaperID: 3884, Poster
Abstract: Process reward models provide step-level feedback for reasoning, but reliable process supervision remains difficult to obtain: human annotations are costly, Monte Carlo rollouts require large sampling budgets, and LLM judges can introduce prompt-sensitive biases. We introduce SPRM, an outcome-only framework that constructs process rewards directly from terminal correctness signals. SPRM views a reasoning trajectory as a temporally ordered cooperative game, where reasoning steps act as participants and the final outcome reward is the shared payoff. Following this allocation view, SPRM learns a differentiable trajectory value function from outcome rewards and distributes the payoff to individual steps through Aumann-Shapley-style Step Value Integration. We further analyze its Shapley properties under the temporal and causal structure of reasoning trajectories. To make the resulting supervision robust, SPRM incorporates causal consistency and estimates a cross-trajectory inherent score that captures each step's context-independent value from semantically similar steps. These signals train a dual-head process reward model that jointly represents trajectory-specific marginal credit and cross-trajectory inherent step value. Across ProcessBench, PRMBench, and Best-of-N reranking on AceMath-RewardBench, SPRM improves first-error localization, marginal-contribution-sensitive evaluation, and test-time reasoning selection over existing open PRMs trained from human, rollout, curated, or outcome-broadcast supervision. These results suggest that marginal-contribution-based credit redistribution offers an effective and scalable route to process supervision from outcome-only data.
PaperID: 3885, Poster
Abstract: Deep click-through rate (CTR) models commonly consist of sparse embedding blocks and feature-interaction blocks, which jointly learn predictive signals from high-dimensional categorical features. This naturally raises the question of whether different blocks in an end-to-end trained CTR model play the same role in generalization. To address this question, we investigate the generalization behavior of deep CTR models from a block-wise diagnostic perspective. Specifically, we formulate deep CTR models as two-block systems comprising an embedding block and a feature-interaction block, and introduce a retuning-based diagnostic protocol to separately assess the transferability of different blocks. Experiments across multiple datasets and architectures reveal a consistent asymmetric generalization pattern: the embedding block exhibits stronger training-set specificity, whereas the feature-interaction block preserves comparatively more transferable structure. Motivated by this diagnosis, we further develop a component-selective embedding aggregation method that averages multiple independently retuned embedding tables while keeping the feature-interaction block fixed. The resulting model retains the original single-path inference architecture and incurs no additional online inference cost. Experiments on Avazu, Criteo, and Taobao show that the proposed method improves representative deep CTR models over standard training and retuning baselines.
Abstract: Stochastic bilevel optimization (SBO) has become a standard framework for hyperparameter learning, data reweighting, representation learning, and data-mixture optimization in deep learning. Existing exact single-loop SBO methods and memory-efficient surrogate SBO methods either create severe memory pressure for large lower-level neural networks or lack competitive convergence guarantees under standard assumptions. In this paper, we propose BROS, a memory-efficient single-loop SBO method with the same convergence rate order as exact single-loop SBO methods. BROS performs lower and auxiliary updates in randomized subspaces with a Rademacher bi-probe correction that recovers an unbiased Hessian-action estimator. We prove that BROS preserves the \mathcal O(\varepsilon^-2) sample complexity of MA-SOBA for finding an \varepsilon-stationary point under only standard assumptions. Experiments on hyper-data cleaning, data-mixture learning, hyper-representation learning, and ViT sample reweighting show that BROS reduces peak memory by up to 44.9% while closely matching full-space baseline performance.
Abstract: Neural networks with randomly generated hidden weights (RaNNs) have been extensively studied, both as a standalone learning method and as an initialization for fully trainable deep learning methods. In this work, we study RaNN expressivity for learning solutions to non-linear partial differential equations (PDEs). Despite their widespread use in practical applications, a rigorous theoretical understanding of the approximation properties of RaNNs in this context remains limited. Here, we derive error bounds for RaNN approximations to time-dependent Sobolev functions and obtain a dimension-free approximation rate \frac12 for sufficiently regular functions. We apply our results to two important classes of non-linear PDEs: Porous Medium Equations and Compressible Navier-Stokes Equations, showing that RaNNs are capable of efficiently approximating solutions to these complex, non-linear PDEs. Our theoretical analysis is supported by numerical experiments, validating the obtained convergence rates.
Authors:
Dingbao Shao, Song Wu, Xinyu Chen, Qian Wang, Jiahang Li, Kuai Jiang, Jiang Lin, Yuhang Liu, Chen Ziyu, Duo Li, Hu Jiaxin, Shengrong Gu, Ziheng Tang, Liu Rongrong, Yanlun Peng, Liang Li, Junlan Feng, Lujia Jin, Ting Zhang, Jian Yang, Zili YiAbstract: Video virtual try-on is a highly constrained editing task requiring the precise replacement of a target person's clothing while strictly preserving the original video's spatial structure and temporal dynamics. Existing methods heavily rely on auxiliary handcrafted spatial priors (e.g., masks, poses) for editing control. However, these priors are prone to failure in unconstrained real-world videos and often compress rich visual context into incomplete structural signals. Furthermore, standard reconstruction objectives fail to fully capture try-on-specific human preferences. To address these challenges, we propose InstructVVT, an instruction-driven and reference-guided video virtual try-on framework based on a Diffusion Transformer (DiT) that operates without inference-time spatial priors. Our core insight is to recover fine-grained control directly from the input triplet (source video, reference garment, and instruction) via a dual-level reference conditioning scheme. Specifically, an MLLM infers semantic edit tokens for target disambiguation and structural preservation, while a lightweight conditioning pathway explicitly injects fine-grained visual garment details. Finally, we design a try-on-specific reward and utilize the DiffusionNFT algorithm to align the model with human preferences. Extensive experiments on ViViD-S and TripVVT-Bench demonstrate that InstructVVT outperforms state-of-the-art open-source methods in garment fidelity, structural preservation, and temporal consistency, despite requiring fewer inference-time controls.
Abstract: Large language model (LLM)-based multi-agent systems (MAS) often exhibit complex failure modes, which frequently cause agents to produce incorrect outcomes. This motivates the task of Agent Failure Attribution: given a failed multi-agent trajectory, identify the faulty agents and their corresponding error types. Existing approaches predominantly rely on LLMs to perform failure attribution, either through direct prompting, fine-tuning on synthetic data or complex agentic pipelines. While effective, these methods incur substantial computational overhead due to long-context processing, expensive post-training and handcrafted workflows. Moreover, empirical evidence shows that even state-of-the-art models achieve limited accuracy on existing benchmarks, suggesting that scaling model size alone is insufficient. In this work, we revisit this task and question the necessity of such expensive generative solutions. We introduce , a lightweight graph-based framework that models interaction trajectories through step-level semantic signals and agent-level relationships. We show that with significantly fewer parameters and near-zero inference cost, (i) matches or outperforms LLM-based baselines, including fine-tuned models on in-domain benchmarks, (ii) maintains robust performance across different GNN architectures and (iii) can be further improved with inexpensive test-time adaptation on the OOD benchmark. Our results suggest that effective agent failure attribution does not require heavy LLM reasoning and a lightweight, structured approach can achieve strong performance.
Abstract: Multi-objective reinforcement learning (MORL) seeks to train agents capable of balancing conflicting objectives. While single preference-conditioned policies offer a highly scalable solution, existing approaches remain brittle in practice, frequently failing to recover dense Pareto fronts. We demonstrate that this failure stems from two structural pathologies: destructive advantage cancellation caused by premature Early Scalarization (ES), and representational mode collapse across the preference space. To overcome these bottlenecks, we introduce \toolname, a PPO-based framework that fundamentally reorganizes multi-objective optimization. By preserving per-objective learning signals through a decomposed pipeline and integrating preferences only after trust-region stabilization (Late-Stage Weighting), \toolname improves credit assignment under conflicting objectives. Concurrently, a scaled diversity regularizer encourages behavioral divergence proportional to preference distance. \toolname operates entirely within the efficient linear scalarization regime shared by standard deep MORL baselines. By reducing information loss caused due to linear scalarization rather than relying on expensive non-linear utility functions, it suggests that optimization bottlenecks play a significant role. Across available standard benchmarks, including high-dimensional and many-objective environments, \toolname consistently discovers broader, higher-quality Pareto fronts than prior methods, exceeding state-of-the-art hypervolume and expected utility using a single deployable policy.
PaperID: 3891, Poster
Abstract: High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for thin or highly connected shapes. We introduce , a topology-aware 3D generation framework that represents shapes with compact sliding-window slice latents. Instead of generating expensive voxel tokens, SILSA uses a fixed set of overlapping slices along the three canonical axes, where each token summarizes a local depth window to preserve cross-sectional continuity and support single-stage rectified-flow generation. A encodes oriented surface samples into multi-axis slice latents and reconstructs them with a sparse volumetric decoder, while a coordinates directional slice streams through a shared 3D workspace. To preserve structural correctness, we introduce slice-level topology supervision that matches persistence diagrams and aligns Betti transitions across neighboring slices. Experiments show that SILSA improves structural fidelity while substantially reducing generation cost. SILSA improves PSNR by 8.7%, coverage by 5.96 absolute points, and Betti error by 9.2% over the strongest baseline, while using 70.0% fewer tokens than the next-most compact baseline and over 98% fewer tokens than sparse or hierarchical tokenizers, reducing training memory by 40.4% and inference time by 58.5%. Qualitative results further show improved preservation of thin structures, repeated components, and long-range connectivity.
Authors: Murat B Ertan, Xiaochen Zhu, Ha Nguyen, Marten van Dijk, Srinivas Devadas
Abstract: We introduce PACZero, a family of PAC-private zeroth-order mechanisms for fine-tuning large language models that delivers usable utility at I(S^; Y_1:T)=0. This privacy regime bounds the membership-inference attack (MIA) posterior success rate at the prior, an MIA-resistance level the DP framework matches only at \varepsilon=0 and infinite noise. All DP-ZO comparisons below are matched at the MIA posterior level. The key insight is that PAC Privacy charges mutual information only when the release depends on which candidate subset is the secret. Sign-quantizing subset-aggregated zeroth-order gradients creates frequent unanimity, steps at which every candidate subset agrees on the update direction; at these steps the released sign costs zero conditional mutual information. We propose two variants that span the privacy-utility trade-off: PACZero-MI (budgeted MI via exact calibration on the binary release) and PACZero-ZPL (I=0 via a uniform coin flip on disagreement steps). We evaluate on SST-2 and SQuAD with OPT-1.3B and OPT-6.7B in both LoRA and full-parameter tracks. On SST-2 OPT-1.3B full fine-tuning at I=0, PACZero-ZPL reaches 88.99\pm0.91, within 2.1pp of the non-private MeZO baseline (91.1 FT). No prior method produces usable utility in the high-privacy regime \varepsilon<1, and PACZero-ZPL obtains competitive SST-2 accuracy and nontrivial SQuAD F1 across OPT-1.3B and OPT-6.7B at I=0.
PaperID: 3893, Poster
Abstract: Long-context inference with large language models is bottlenecked by the linear growth of the KV cache. Existing eviction-based methods reduce memory overhead by irreversibly discarding historical states, while retrieval-based methods preserve the full context but rely on retrieval units defined by tokens, fixed pages, or proxy signals including query-key similarity, clustering, and linguistic rules. Such externally defined units often misalign with the model's actual contextual dependencies during generation. We propose AsdaKV, an adaptive semantic drift-aware KV cache retrieval framework that defines semantic continuity directly from the model's inherent attention behavior. AsdaKV detects semantic drift by measuring the overlap between high-attention historical regions across decoding steps: stable overlap indicates that the active cache still covers the relevant contextual evidence, whereas a sharp overlap drop triggers cache refresh. Guided by this drift signal, AsdaKV organizes historical KV states into variable-length semantic windows, offloads full window KV states to CPU memory, and maintains only lightweight window anchors plus a fixed-budget active KV cache on GPU. During decoding, the current query retrieves relevant windows via anchor matching, while adaptive-stride drift detection and deferred page recall reduce redundant detection, retrieval, and KV reloading overhead. Extensive experiments across diverse tasks and model families demonstrate that under the same KV budget, AsdaKV achieves higher accuracy than state-of-the-art KV retrieval methods, maintains comparable decoding efficiency in long-input settings, and further outperforms existing retrieval systems for long-output generation.
PaperID: 3894, Poster
Abstract: Using neural networks to learn optimal transport (OT) maps via minimax optimization has gained increasing interest in generative modeling. However, the widely used semi-dual OT maximin formulation is known to suffer from spurious global solutions that do not correspond to valid transport maps, often addressed by regularizing the objective function. We provide an alternative explanation by showing that these spurious solutions are not stationary points of the associated loss, implying that their appearance in practice stems from the optimizer's failure to converge to stationary points in nonconvex-nonconcave settings. However, when the c-concavity constraints are imposed on the potential, commonly to stabilize training, we show that these spurious solutions become stationary. Based on these observations, we propose solving the semi-dual OT problem without c-concavity constraints, using optimizers that can reliably converge to stationary points corresponding to OT maps in nonconvex-nonconcave settings, such as a two-timescale extragradient (TTS-EG) method. At the formulation level, we further show that a Lagrangian-based OT minimax formulation provably eliminates spurious global solutions. Our empirical results demonstrate that TTS-EG reliably recovers OT maps under both semi-dual and Lagrangian-based OT formulations without any regularization.
Abstract: Humanoid Visual Search (HVS) requires agents to actively explore immersive 360^\circ environments. While prior methods treat this as a monolithic task relying on cumulative, multi-turn Chain-of-Thought (CoT) reasoning, they impose heavy cognitive burdens and require expensive trajectory-level annotations. In this paper, we propose Imagining in 360^\circ, a novel framework that decouples the exploration process into a specialized Imaginator and an Actor. The Imaginator functions as a probabilistic predictor of spatial priors; instead of maintaining a cumulative reasoning chain, it infers the semantic layout of both observed and unobserved regions in a single step. By sampling multiple hypotheses within this semantic space, we provide the Actor with a distribution of effective spatial information, offering robust guidance that hedges against uncertainty during active search. This decoupled architecture significantly lowers data engineering costs by eliminating the need for full-trajectory CoT annotations, enabling the generation of over 1.96 million curated training samples. Extensive experiments demonstrate that explicitly modeling semantic spatial priors drastically improves search efficiency and success rates in complex, in-the-wild environments.
PaperID: 3896, Poster
Authors: Xinlong Zhang, Jia Wei, Xiaoyu Zhang, Teng Zhou, Chengyu Lin, Yongchuan Tang
Abstract: Achieving high-fidelity object-level control in diffusion transformers remains a significant challenge despite the introduction of structural priors like depth and Canny maps. Current object-level conditional generation methods frequently suffer from visual artifacts and struggle to maintain precise control over objects within small localized regions. To address these limitations, we propose Cascaded Object-Level Latent Refinement (COLLAR), a training-free framework that progressively optimizes object-level features via the Field-of-View (FoV) expansion. First, we propose the Cross-Scale Semantic Alignment (CSSA) module to address spatial-semantic gaps by injecting object-level features into extended-FoV branches via attention mechanisms. To further optimize these features, the Cyclic Feature Injection (CFI) module introduces a reciprocal background feedback mechanism. It leverages a frequency-based adaptive strategy to selectively update the global backbone with context-aligned local information. Finally, the extended-FoV branch serves as a hub for feature optimization, ensuring that object-level features are integrated into the global generation process without compromising final image quality. Extensive experiments on the COCO-MIG and COCO-POS benchmarks demonstrate that our approach consistently outperforms state-of-the-art methods across semantic alignment, image quality, and spatial fidelity.
PaperID: 3897, Poster
Authors: ziyang zhao, Junlang Tian
Abstract: Scaling improves language-model capability, but it does not necessarily strengthen internal safety mechanisms. We introduce , a representation-level framework for auditing how aligned language models respond to adversarial perturbation inside activation space. Given a learned refusal direction, we define as the Fisher--Rao-normalized distributional response induced by controlled activation perturbation. This quantity measures how strongly refusal behavior remains coupled to adversarial task execution in representation space. Across six open model families and 22 checkpoints, including Qwen, DeepSeek-Distill, Mistral, Llama-3, Gemma, and Yi, we observe a recurring geometric scaling pattern. Small aligned models often exhibit elevated refusal response, intermediate-scale models frequently enter weakened-response regimes ( ), and larger checkpoints trend toward near-baseline response consistent with increasing task--refusal decoupling. Cosine measurements on continuous Qwen and DeepSeek scaling axes support this interpretation, showing monotonic decay between refusal and task-generation directions with scale. We further provide causal evidence through inference-time activation intervention on Qwen-72B. Injecting a learned refusal vector at a critical semantic layer restores refusal behavior on the evaluated jailbreak subset while largely preserving benign reasoning performance. Matched-norm placebo vectors fail to reproduce the effect, supporting the directional specificity of the intervention. Together, these results suggest that behavioral refusal can remain surface-level even when internal refusal geometry becomes weakly coupled to adversarial task execution. LLM Rheology provides a complementary representation-level perspective for auditing refusal safety beyond behavioral evaluation.
PaperID: 3898, Poster
Abstract: The discrepancy between controlled synthetic distortions and complex real-world degradations poses a significant domain shift challenge for No-Reference Image Quality Assessment (NR-IQA). While Unsupervised Domain Adaptation (UDA) methods aim to bridge this gap, they heavily suffer from retraining overhead and error propagation during iterative pseudo-labeling. Inspired by the dual-process theory of human cognition, we propose Hierarchical Intuitive and Fuzzy Consensus (HIFC), a completely train-free cross-domain IQA framework. HIFC employs an asymmetric perceptual decomposition to disentangle semantic and distortion features. It derives the final quality score through a synergy of two branches: a memory consensus branch that retrieves a stable Fuzzy Quality Centroid from a synthetic reference gallery, and an intuitive branch that captures aesthetic polarization via zero-shot vision-language alignment. Extensive experiments demonstrate that, without any parameter updates or target-domain adaptation, HIFC achieves highly competitive results against state-of-the-art UDA methods. Most notably, it exhibits exceptional robustness and generalization in the challenging synthetic-to-authentic cross-domain setting. The code will be available upon acceptance.
PaperID: 3899, Poster
Abstract: Equipping artificial agents with spatial intelligence requires a comprehensive generative prior over the dynamic 3D world. We propose World Motion Models (WMMs) that capture "what was, is, and will be where across time" via sparse SE(3) pose trajectories. WMMs are built on the observation that elements of dynamic scenes can be well-approximated by a set of rigid SE(3) trajectories, a minimal yet expressive primitive for 4D modeling. This representation unifies articulated objects, human bodies, hand-object interactions, piecewise-rigid scene dynamics, camera motion, and even robot states and actions into a single shared space. Given this representation, we cast the joint distribution of these entities as a flexible sequence modeling problem, utilizing flow-matching with per-token noise levels. Coupled with a context token mechanism with non-sequential conditioning, this formulation supports any-to-any marginal conditioning across an arbitrary number of entities and time steps. Tasks such as future prediction, motion infilling, model-predictive control, inverse kinematics, cross-embodiment retargeting, and policy learning all reduce to the application of different masks over the same network. Experiments on diverse 6 applications of 3D vision and robotics demonstrate the versatility and flexibility of WMMs with strong performance.
Abstract: Data scaling is fundamental to modern deep learning, and grows increasingly critical as autonomous driving shifts to end-to-end learning. Real-world driving data is expensive to annotate and scene-biased, making real-synthetic co-training with near-infinite synthetic data a promising direction. However, naively incorporating all available synthetic data is inefficient and leads to distribution shifts, and optimizing data mixture under practical training budgets remains a critical yet under-explored problem. In this sense, we claim that the mixture of training data requires clear guidance in terms of scene types and quantities. Particularly in this work, we conceptualize the data mixture approximately as a dynamic optimization process that iteratively adjusts the training data mixture to maximize model performance, guided by closed-loop evaluation feedback, and propose AutoScale, a fully automated closed-loop data engine unifying scene representation, data mixture optimization and retrieval, as well as model training and evaluation. Specifically, we propose Graph Regularized AutoEncoder (Graph-RAE) for driving scene representations, introduce Cluster-aware Gradient Ascent (Cluster-GA) for cluster-wise importance estimation and reweighting, and perform cluster-guided vector retrieval to select high-value samples. Experiments on NavSim demonstrate that AutoScale outperforms vanilla co-training and cross-domain baselines, achieving better performance with fewer synthetic samples under constrained budgets.
PaperID: 3901, Poster
Abstract: Vision language navigation (VLN) is very challenging for Unmanned Aerial Vehicles (UAVs) since small errors in 3D motion prediction can lead to catastrophic outcomes. Unable to fathom execution outcomes in 3D environments, existing methods often struggle to predict accurate navigation waypoints from language instructions. To solve this, we introduce the Prototypical Visual Imagination framework, shortened as ProtoVis-Nav, which forces the model to imagine the execution outcome by predicting visual features surrounding the next intended location of the UAV. Specifically, instead of using multi-layer perceptrons, we predict feature prototypes and use their linear combination to estimate the visual feature of the target region. In this manner, the VLN model first imagines the future location visually, and then uses predicted prototypes as context for generating the final action, promoting alignment between visual context and waypoint outputs. Extensive experiments on the OpenUAV benchmark demonstrate the effectiveness of the ProtoVis-Nav framework and show substantial improvements over existing baselines.
PaperID: 3902, Poster
Abstract: Offline reinforcement learning suffers from distributional shift and extrapolation errors. These issues are particularly severe in multi-agent settings. As the number of agents increases, the joint action space grows exponentially while agent behaviors become highly coupled. Consequently, even if individual actions remain within the data distribution, their combination may still result in out-of-distribution (OOD) joint actions. Directly extending single-agent constraints to multi-agent settings is difficult, as they fail to effectively constrain joint actions. This often results in poor coordination or excessive conservatism. To bridge this gap, we propose the joint adaptive neighborhood constraint (JANC), which explicitly constructs controlled neighborhoods in the joint action space to suppress extrapolation while preserving reliable generalization. Moreover, JANC adaptively scales neighborhood radii based on joint advantages to align generalization with value structures. In practice, our method first performs adaptive neighborhood Q learning to explore high-value behaviors and then conducts value and policy learning based on the refined actions. Experiments on offline benchmarks, including multi-agent Mujoco and StarCraft II, show that JANC outperforms state-of-the-art methods on the majority of tasks in complex cooperative scenarios. We will release our source code to the public.
PaperID: 3903, Poster
Abstract: Large language models are increasingly used to estimate uncertain quantities from context, raising the question of whether their probabilistic outputs are internally consistent. A basic test of self-consistency is whether these estimates adhere to the law of total probability. We investigate this requirement through the lens of binary conditioning trees, which recursively partition a population into increasingly fine-grained subsets and expose multiple ways of estimating the same aggregate quantity. This construction yields a basic self-consistency check: marginal estimates must agree with prior-weighted aggregations of conditional estimates over any partition of the population. We find that current models systematically violate this requirement. In a case study on persona prompting, prior-weighted aggregates are consistently better aligned with human population statistics than direct estimates. Notably, this benefit of specificity persists even when the prior weights are themselves estimated by the LLM. Turning this discrepancy into a general evaluation criterion, we propose a family of self-consistency checks for LLMs grounded in the law of total probability. By evaluating these checks across frontier models, we show that failures of statistical consistency are widespread and not confined to persona prompting. Together with a benchmark dataset, our work provides a testbed for understanding when in-context learning can be treated as conditional inference.
Authors: Daniel López Montero, Antonio Álvarez-López, Marcos Matabuena
Abstract: Modern datasets across many disciplines increasingly consist of time-evolving, potentially infinite-dimensional random objects, such as dynamic functional data, which are naturally modeled in Hilbert spaces. In these settings, characterizing probability measures, for example, through densities, can be ill-defined or technically challenging. Motivated by clustering applications, we propose a Gaussian mixture framework for Hilbert-space-valued data based on kernel mean embeddings and develop efficient optimization algorithms for estimation. We establish theoretical guarantees showing that the proposed algorithm is well defined and that the model yields a dense class of approximations in infinite-dimensional spaces. We evaluate the framework through extensive experiments on diverse structures and data geometries, including L^2-functional data and random graphs in Laplacian spaces arising in modern medical applications.
Abstract: Simulating human behavior with Large Language Models (LLMs) offers a scalable laboratory for social science and stress-testing AI products. However, limitations remain in matching LLM outputs against the full breadth of human diversity. Current techniques for creating synthetic humans typically optimize for density matching, which often enforces behavioral uniformity and overlooks the rare, consequential outliers. We argue that robust simulation requires shifting the focus toward support coverage: ensuring synthetic populations span the entire landscape of possible human traits, opinions, and preferences. This paper introduces Persona Generators, functions that can produce diverse synthetic populations tailored to arbitrary contexts. We apply an iterative improvement loop based on AlphaEvolve, using LLMs as mutation operators to evolve our Persona Generator code rather than the personas themselves. The optimization process produces lightweight functions that can automatically expand small descriptions into populations of diverse synthetic personas that are maximizing coverage along relevant diversity axes. We demonstrate that evolved generators substantially outperform existing baselines across six diversity metrics on held-out contexts. Furthermore, by successfully mitigating standard LLM mode collapse, our generated populations actually capture real human trait distributions better than baselines specifically designed to match human statistics.
PaperID: 3906, Poster
Abstract: Many decision-level fusion methods implicitly assume that semantically aligned soft outputs can be directly aggregated once they share the same label simplex. However, this assumption can fail even under identical label semantics, as oracle-specific differences in confidence sharpness, boundary uncertainty, and tail-mass allocation may place soft outputs in incompatible probability geometries and bias direct soft fusion. To address this issue, we formalize this failure mode as geometric incomparability and propose a pre-fusion comparability recovery framework, which learns a shared consensus through constrained oracle-specific rectifiers on the probability simplex. The framework includes a closed-form variant with frozen geometry descriptors and an alternating-refinement variant with consensus-dependent descriptors. Theoretically, we analyze the rectifier family, prove exact recovery under matched radial distortions, characterize the bias of direct fusion, and establish stable recovery under approximate mismatch. Experiments on controlled synthetic stress tests and real black-box decision-level settings show that comparability recovery consistently reduces fusion bias and improves reliability over direct aggregation. These results indicate that geometric comparability is a necessary condition for reliable soft-output fusion. The code is available at \urlhttps://anonymous.4open.science/r/CFCR-ACR-B8F1.
Abstract: Generative novel view synthesis from sparse input images is rarely all reconstruction or all generation: pixels visible in some source view have a unique correct value modulated only by view-dependent shading, while pixels in disocclusions or beyond the captured volume admit a distribution of plausible completions. Existing generative novel-view-synthesis methods conflate these regimes under a single uniform loss, blurring the line between geometric fidelity and creative hallucinations even when scene geometry is injected through warped point clouds or projected depth. We introduce GenRec, a multi-view flow matching model that builds the reconstruction--generation split directly into its architecture, supervision, and gradient flow. Guided by an observation mask derived from the source cameras and a monocular depth estimator, a flow matching backbone jointly denoises RGB and scene-coordinate maps across all target views, while a pixel-space refinement stage restores high-frequency detail on observed pixels; the same mask gates supervision so regression signals do not contaminate the generative prior. Across RealEstate10K, DL3DV-10K, and Mip-NeRF~360, in both single-view extrapolation and two-view interpolation, GenRec attains the best reconstruction fidelity in observed regions while also surpassing purely generative baselines on perceptual quality in unobserved ones, showing the effectiveness of our approach. The source code and trained models will be made public.
PaperID: 3908, Poster
Abstract: The effectiveness of modern deep learning models in forecasting groups of time series is often attributed to their ability to capture (long-range) dependencies across long observation windows. In this paper, we show that this forecasting task involves two objectives: (i) generative process identification (GPI), i.e., inferring the specific process generating the input sequence, and (ii) conditional forecasting (CF), i.e., predicting future values given input observations. From this perspective, optimal predictions can be interpreted as an average over plausible data-generating processes, weighted by their likelihood given the input window. This suggests a different explanation for the benefits of long context windows: they reduce the uncertainty about which specific process is generating the input time series during operation. We prove that even for processes with memory length P, an input window size strictly larger than P is necessary to achieve the minimum attainable error. Finally, we show how decoupling GPI and CF can improve computational scalability without compromising accuracy. Experiments on synthetic and real-world data validate our insights and their relevance for designing forecasting architectures.
Abstract: Radio-frequency (RF) data synthesis predicts the received signal given transmitter and receiver positions, and is essential for wireless applications. Recent 3D Gaussian Splatting (3DGS)-based methods achieve efficient synthesis at any transmitter but only for a fixed receiver. Therefore, supporting N receivers in one scene requires N independent models and precludes prediction at unseen receivers. We present RxGS, which achieves receiver-generalizable synthesis within a single unified model. Our key insight is that scene geometry is receiver-independent while directional radiance is not: a first stage learns shared 3D Gaussian geometry, and a second stage freezes it and learns directional radiance conditioned on receiver position. A global conditioning branch captures shared receiver-dependent effects across the scene, while a local branch models per-scatterer variations from the receiver's geometry and occlusion. A multi-receiver CUDA rasterizer further batches rendering across all N receivers. Evaluated across various RF datasets, RxGS matches or improves over per-receiver baselines with a single shared model and generalizes to receivers unseen during training within the scene, cutting training cost by up to 45×, inference cost by 7.6×, and storage by N×.
PaperID: 3910, Poster
Abstract: Recently several methods have shown promise for purely in silico “de novo” design of antibodies which bind to drug targets. The methods which have shown success in vitro propose candidates via either a generative diffusion model or hallucination-based sequence optimization, and then filter these candidates with a structure predictor model. Hallucination methods rely on compute-intensive backpropagation through a structure predictor model for each candidate, and the generative methods require expensive structure re-prediction and filtering of many (often mostly non-passing) candidates, limiting their application to test-time generation of large libraries. Given the highly variable per-drug-target hit rates of current methods, screening large libraries is a well established route to improved antibody candidate discovery. Thus, we propose an alternative approach, Bonobo, which instead formulates this problem as black-box optimization of a per-drug-target generative model, amortizing candidate generation into training. This is accomplished by using a structure predictor as a reward signal to directly train a GFlowNet generative model. We show that this approach can match or exceed the in silico metrics of state-of-the-art approaches, while allowing for dramatically more efficient generation of diverse and arbitrarily large numbers of antibody binder candidates.
PaperID: 3911, Poster
Authors:
Francis Snelgar, Stephen Gould, Liang Zheng, Akshay AsthanaAbstract: Recent advances in visual storytelling have leveraged powerful text-to-image diffusion models to generate sequences of images featuring consistent subjects and styles. While such methods achieve impressive identity preservation across the sequence, they often produce scenes with limited structural diversity---images that are overly similar in subject pose and spatial layout. Existing evaluation protocols primarily measure identity consistency using metrics such as DreamSim, which can inadvertently conflate identity preservation with subject pose. The inability to independently evaluate structural diversity inhibits progress on identity preserving image generation methods. We address this limitation by introducing a new metric based on the Gromov Wasserstein distance to quantify structural diversity independently of identity. Our metric provides an interpretable and generalizable measure of spatial variation across generated images, enabling a more balanced assessment of identity consistency and pose diversity. Using this framework, we systematically evaluate several recent visual storytelling and subject-consistent generation methods. Furthermore, we evaluate the effect of latent initialization methods on structural diversity in several popular text-to-image models.
PaperID: 3912, Poster
Abstract: Neural audio codecs compress waveforms into compact discrete tokens that underpin speech language models, real-time communication, and large-scale audio storage. Almost every dominant design, including residual vector quantization, finite scalar quantization, and single-codebook variants, follows the VQ-VAE template by partitioning the encoder latent through a learned codebook or a fixed scalar grid. We ask whether this partition is necessary. We introduce GS-Codec, a neural speech codec whose bottleneck is a parametric signal decomposition rather than a quantizer. We adapt Gaussian splatting from 3D scene reconstruction to one-dimensional latents. An inner optimization loop fits each encoder segment as a weighted sum of 1D Gaussian primitives. The decoder then reconstructs the waveform from the rendered sum. To avoid the cost of this iterative inner loop at inference time, we additionally train a lightweight GS Predictor Net that regresses the primitive parameters in a single forward pass. The encoder and decoder are trained end-to-end through the inner loop, with no quantizer anywhere in the training pipeline: the bottleneck is the decomposition itself, and scalar quantization is applied only post-training to the fitted parameters. Rather than relying on discrete codebook stages for bitrate control, our representation exposes a smooth rate-quality tradeoff: a single trained checkpoint supports continuous post-training bitrate control by varying the number of primitives and the per-parameter bit depth, with no retraining required. GS-Codec outperforms well-established open-source codecs such as EnCodec and DAC on speaker similarity (SIM), intelligibility (STOI), and perceptual quality (UTMOS) at comparable bitrates, while achieving comparable semantic performance.
PaperID: 3913, Poster
Authors: Charles Cao
Abstract: Safe preference alignment of large language models must handle rare harmful tails whose burden is concentrated on demographic, topical, linguistic, or adversarial subgroups. Existing coverage analyses operate on global density ratios that can look adequate on average while the harmful tail of a small group remains unsupported, leaving worst-group safety risk unidentified. We assume that the absolute harm scale is externally calibrated by severity labels or anchors and introduce TailGuard, a framework centered on subgroup quantile coverage C_g,\tau: the worst deployment-to-training density ratio restricted to group g's upper-τ harm tail. We prove that C_g,\tau is a key statistical quantity governing subgroup tail safety: it governs tail-risk transfer, yields an offline impossibility result under missing tail support, and is necessary for fixed-policy subgroup tail-risk estimation. From a per-group CVaR_τ-constrained alignment objective, we derive a KL-regularized Gibbs form, motivate a Tail-Fisher active query rule that targets group-tail coverage deficits, and construct a per-group weighted CVaR risk-control deployment gate. We validate the mechanism through four completed data-level studies: synthetic coverage control, a semi-synthetic tail-hiding pilot on real safety-alignment labels, an observed-label active-query ablation, and an oracle-score risk-control-gate efficiency proxy. These experiments test the coverage mechanism and deployment diagnostic.
PaperID: 3914, Poster
Authors: Robert J Alvarez
Abstract: JEPAs often regularize one-view embeddings toward an isotropic Gaussian, implicitly baking Euclidean symmetry into the representation. We show that this is not merely a benign default. For a known structured downstream geometry H\succ0, the minimax and maximum-entropy covariance under a Hamiltonian energy budget is (c/d)H^-1, and Euclidean isotropy incurs a closed-form price of isotropy. More importantly, when the downstream geometry is unknown, no geometry-independent fixed marginal target is canonical: every fixed covariance shape can be maximally misaligned for some structured geometry. We further show that even oracle one-view marginals do not identify the JEPA view-to-view predictive coupling. These results suggest that the structural bias in JEPAs should enter the cross-view coupling rather than a fixed encoder marginal. We instantiate this principle with HamJEPA, which encodes each view as a phase-space state (q,p) and predicts view-to-view transitions with a learned Hamiltonian leapfrog map, while non-isotropic scale and spectral floors prevent collapse. In a deliberately headless token protocol, HamJEPA improves over SIGReg on CIFAR-100 by +4.89 kNN@20 and +3.52 linear-probe points at 30 epochs, and by +6.45 kNN@20 and +10.64 linear-probe points at 80 epochs, while a matched MLP predictor ablation shows that the symplectic coupling is the ingredient driving the neighborhood-geometry gain. On ImageNet-100, HamJEPA-q improves by +4.82 kNN@20 and +7.52 linear-probe points at 45 epochs.
PaperID: 3915, Poster
Authors:
Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, WonJun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan HongAbstract: Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of the scene. Inspired by this process, we introduce Imagine3D-LLM, an MLLM that learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation. Concretely, we append a small set of learnable summary tokens after the image tokens, decode them into a compact 3D Gaussian Splatting representation supervised by a photometric reconstruction loss, and train jointly with the standard next-token prediction objective. Notably, although only the summary tokens receive direct reconstruction supervision, this objective also induces stronger cross-frame correspondence within the LLM's underlying image features, suggesting that learning to reconstruct propagates 3D-aware signals throughout the model. As a result, Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks, suggesting that imagining the scene can be more effective than being told its pixel-wise geometry. Our code and weights will be publicly released.
PaperID: 3916, Poster
Authors:
Yuanmin Tang, Lin Li, Yang Du, Yuan Gao, Massimiliano Mancini, Jun Song, Gaopeng Gou, Gang Xiong, Meikang Qiu, Cheng Yu, Bo ZhengAbstract: Composed Image Retrieval (CIR) retrieves a target image given a reference image and a textual edit. Existing training-free methods prompt a frozen multimodal large language model (MLLM) to infer the user's edit intent and rewrite the composed query into a single retrieval artifact, such as a target caption, an edited image, or a modality-fused embedding. These methods can reason about the edit, but they have limited means to execute that reasoning during retrieval. A key remaining bottleneck for training-free CIR appears to be not only intent understanding, but intent execution. To address this challenge, we propose SkillCIR, a training-free framework that tackles the intent execution gap by routing the inferred constraints to specialized retrieval skills. Specifically, the same frozen MLLM that infers the intent also emits a structured plan of active constraints; each constraint activates a typed skill that scores gallery candidates along one evidence channel, and the per-skill scores are composed in score space with explicit signs and re-checked by a local verification step. Across three benchmarks and three CLIP backbones, SkillCIR improves the primary recall/mAP metrics by 2.8 to 7.9 points at the latency of typical training-free CIR pipelines and around 23x faster than the latest training-free state of the art. These gains hold with a fixed skill pool and no model-parameter updates, suggesting that improving the means to execute the inferred edit is a useful complement to stronger reasoning about it.
PaperID: 3917, Poster
Authors:
Qianyue Cao, Zongwei Zhu, Boyu Li, Yi Xiong, Zirui Lian, Xuehai ZhouAbstract: Federated fine-tuning has become a dominant paradigm for privacy-preserving Large Language Model (LLM) adaptation. While integrating Parameter-Efficient Fine-Tuning (PEFT) reduces communication and computational costs, existing methods neglect that peak memory bottlenecks caused by full forward passes through the frozen LLM. This excludes low-memory devices, leading to data loss and suboptimal global performance. In this paper, we propose BP-FedPEFT, a framework utilizing progressive training to decompose the end-to-end computational graph, reducing peak memory to the block level. While enabling low-memory device participation, this paradigm incurs prolonged training latency and introduces growing memory burdens for deep-layer inputs, alongside suffering from cascading feature misalignment due to the absence of global supervision, leading to suboptimal model performance. To ensure efficiency, BP-FedPEFT employs functional-aware overlapping planning coupled with a local-global stability criterion to regulate training steps and communication rounds. To ensure effectiveness, we utilize depth-injected input synthesis and block overlaps to bridge the supervision gap, establishing valid optimization trajectories that align shallow representations with deep functional expectations. We establish theoretical convergence guarantees for BP-FedPEFT. Experiments on a heterogeneous testbed show that BP-FedPEFT supports diverse PEFT methods, reducing average memory usage by 44.7-76.2%, accelerating training by 2.3-12.6×, and improving accuracy by 3.2-5.8% through inclusive participation.
Abstract: Recent text-to-image (T2I) diffusion models have achieved remarkable advancement, yet faithfully following complex textual descriptions remains challenging due to insufficient interactions between textual and visual features. Prior approaches enhance such interactions via architectural design or handcrafted textual condition weighting, but lack flexibility and overlook the dynamic interactions across different blocks and denoising stages. To provide a more flexible and efficient solution to this problem, we propose Diff-Aid, a lightweight method that adaptively adjusts per-token text and image interactions across transformer blocks and denoising timesteps. Beyond improving generation quality, Diff-Aid yields interpretable modulation patterns that reveal how different blocks, timesteps, and textual tokens contribute to semantic alignment during denoising. As a plug-and-play module, Diff-Aid can be seamlessly integrated into downstream applications for further improvement, including style LoRAs, controllable generation, and zero-shot editing. Experiments on strong baselines (SD~3.5 and FLUX) demonstrate consistent improvements in prompt adherence, visual quality, and human preference across various metrics. Our code and models will be released.
PaperID: 3919, Poster
Abstract: The Rashomon set paradigm seeks to uncover the full collection of near-optimal predictive models from a given function class. Access to this set allows users to better understand predictive multiplicity, interact with models, and select those that best satisfy domain-specific constraints. For sparse generalized additive models (GAMs), however, existing approaches that approximate the Rashomon set only apply to restricted subsets of features, overlooking the fact that equally accurate models can rely on many different feature subsets. We address this gap by representing the Rashomon set through a maximally diverse collection of models. We define a Metropolis-Hastings style approach for sampling maximally diverse GAMs from the Rashomon set. Experiments show that our methods can produce substantially more diverse solutions across all tested diversity metrics, providing users with actionable alternatives that maintain similar predictive performance.
PaperID: 3920, Poster
Authors:
Ofri Eisen, Sharon Goldstein, Ran Elbaz, Kfir Y. LevyAbstract: While decentralized learning offers a communication-efficient alternative to centralized distributed training, its scalability is often limited by the number of workers that can be used without degrading statistical efficiency. This limitation is especially pronounced over sparse communication networks, where increasing parallelism can lead to a sharp loss in performance. To address this bottleneck, we introduce Decentralized \mu^2-SGD (DMS), a novel decentralized optimization method that significantly extends the parallelism limits of decentralized learning. From a theoretical perspective, within the stochastic convex optimization (SCO) framework, we establish improved bounds on the maximal allowable parallelism, surpassing existing decentralized algorithms. Notably, for a broad class of network topologies, our method matches the parallelism scaling of centralized learning, thereby effectively eliminating the gap between decentralized and centralized optimization. Empirically, we validate our theoretical findings through comprehensive experiments, demonstrating the benefits of our approach.
PaperID: 3921, Poster
Abstract: Sparse-Linear Attention (SLA) combines sparse and linear attention to accelerate diffusion models and has shown strong performance in video generation. However, (i) SLA relies on a heuristic split that assigns computations to the sparse or linear branch based on attention-weight magnitude, which can be suboptimal. Additionally, (ii) after formally analyzing the attention error in SLA, we identify a mismatch between SLA and a direct decomposition into sparse and linear attention. We propose CSLA, which introduces (I) a learnable router that dynamically selects whether each attention computation should use sparse or linear attention, (II) a more faithful and direct sparse-linear attention formulation that uses a learnable ratio to combine the sparse and linear attention branches, and (III) a sparse + low-bit attention design, where low-bit attention is introduced via quantization-aware fine-tuning to reduce quantization error. Experiments show that on video diffusion models, CSLA can achieve 97% attention sparsity and deliver an 18.6× attention speedup while preserving generation quality.
Abstract: Motion forecasting is central to visual intelligence: agents must anticipate how objects will move in order to plan actions, reason about physical interactions, and synthesize realistic futures. We argue that 3D points in world coordinates provide a general representation that is class-agnostic, view-stable, compact, and directly useful for downstream tasks. We formalize the task of goal-conditioned 3D point motion forecasting: given a short visual history, a set of 3D query points on an object of interest, and a language description of the intended goal, the model predicts the future 3D trajectory of each point. We introduce a full stack to study this task at scale: (1) MolmoMotion-1M is a large corpus of action-described, object-grounded 3D point trajectory dataset annotated from 1.16M unconstrained videos; (2) PointMotionBench is a human-verified benchmark spanning 111 object categories and 61 motion types; and (3) MolmoMotion is a general motion forecasting model that supports both autoregressive coordinate prediction and flow-matching-based trajectory generation. MolmoMotion is able to accurately predicts diverse motion patterns with different language instructions, and significantly outperforms all existing motion prediction baselines on PointMotionBench. Finally, we show that the learned 3D motion prior transfers well to downstream applications: it improves training efficiency and generalization for robot manipulation, and its predicted trajectories provide effective motion guidance for generative models to synthesize videos with more realistic object motion.
PaperID: 3923, Poster
Authors:
Daniel Rajaonarivonivelomanantsoa, Oussama Hidaoui, Refiloe Shabe, Noah De Nicola, Juan Formanek, Ruan John de Kock, Sasha Abramowitz, Omayma Mahjoub, Arnol M Fokam, Simon Du Toit, Asim Osman, Siddarth Singh, Ulrich Armel Mbou Sob, Willie Brink, Arnu Pretorius, Felix ChalumeauAbstract: Multi-Agent Reinforcement Learning (MARL) is a powerful approach to large-scale systems control, such as power grids and traffic networks, where dozens to hundreds of agents must coordinate effectively. Furthermore, such systems often operate in digital or simulated settings, which allows for inference-time search to significantly improve upon zero-shot performance. A current leading approach in this setting is COMPASS, which builds a continuous family of policies conditioned on latent vectors that can be efficiently searched at test time, but requires all agents to share a single latent vector, limiting the space of reachable joint policies. We introduce ATLAS, a simple extension in which each agent conditions on its own latent vector while still optimising a joint objective. This expands the searchable policy space at negligible computational cost and requires no additional training. ATLAS consistently sets a new state-of-the-art on a benchmark of 14 scenarios across three environments, scaling to teams of up to 200 agents, widely out of distribution. Crucially, its lead sharpens with team size and compute, reaching up to a 45% performance gain at scale.
PaperID: 3924, Poster
Abstract: Multimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples. However, targeted transferable attacks are particularly challenging, since perturbations must inject specific target semantics while generalizing from surrogate models to unseen black-box MLLMs. Existing methods mainly rely on coarse global feature alignment or unstable local matching, which tend to overfit surrogate-specific representations and fail to preserve spatially consistent local structures. In this paper, we propose Dual Feature-Relational Alignment Attack (DFRA-Attack), a locality-aware framework for targeted transferable attacks on MLLMs. To capture local fine-grained semantics, DFRA-Attack introduces semantic-aware alignment, which aligns adversarial and target images in a shared local observation space using saliency-guided shared masking, ensuring that both images are constrained under strictly consistent visible regions. Beyond semantic-aware alignment, DFRA-Attack introduces relational-aware alignment, which preserves target-consistent inter-region dependencies by jointly aligning gram-based feature relational matrices and attention-based interaction maps constructed from visible regional embeddings. Furthermore, we introduce a temperature-annealed dynamic reweighting mechanism to adaptively balance multi-surrogate optimization, coupled with a progressive strategy that refines the adversarial image from masked-view local alignment to full-image target consistency. Extensive experiments on open-source, closed-source, and reasoning MLLMs demonstrate that DFRA-Attack achieves significantly stronger targeted transferability and higher semantic fidelity than state-of-the-art baselines.
PaperID: 3925, Poster
Abstract: Long-horizon forecasting of dynamical time series remains a fundamental challenge in machine learning, as accurate predictions require learning the underlying dynamics. While recent time-series foundation models leveraging in-context learning (ICL) have shown promise, we show that they fail to genuinely capture the governing dynamics. By rewriting the context of a time series as a collection of lagged input-output pairs, we frame dynamical time-series forecasting as a posterior inference problem over the space of structural causal models (SCMs), arguing that the underlying dynamic structure is a natural inductive bias. We propose DynaPFN, which uses pretrained tabular Prior-Fitted Networks (PFNs) to approximate this posterior through ICL, augmented with a library of dynamical features that permit simpler representations of complex systems. Even though the pretrained models have never seen time series data, DynaPFN outperforms existing zero-shot forecasting baselines on established benchmarks of chaotic time series and real-world datasets, demonstrating that incorporating structural priors over system dynamics is a valid approach for robust long-horizon forecasting.
Abstract: The signature transform is a principled feature map for continuous-time paths, valued for its uniqueness and universality. Recovering a path from its truncated signature is, however, structurally ill-posed because the truncated signature map is not injective. We therefore reformulate truncated signature inversion as a probabilistic problem---learning the conditional distribution of a path given its truncated signature---and adopt a signature-conditioned flow matching model as a practical estimator. This probabilistic reformulation elucidates the fundamental difficulty of inversion: Bayes reconstruction error quantifies the irreducible uncertainty remaining after conditioning on a statistic. We derive the Bayes-optimal error under linear statistics, obtaining a closed form for log-GBM and numerically tractable formulas for log-fBM and OU---yielding a concrete theoretical baseline for model validation. This baseline upper-bounds the Bayes error under truncated-signature conditioning, since truncated signatures provide richer information than linear statistics. Experiments show that empirical reconstruction errors under linear-statistics conditioning closely match the theory-derived baseline, while errors decrease when the statistic is replaced with the truncated signature. Moreover, generated paths faithfully recover the conditioning signature while preserving key distributional and temporal structure, indicating that the estimator is well calibrated to the target conditional distribution. Together, these results establish a well-posed probabilistic framework for truncated-signature inversion, with applicability demonstrated on real financial data beyond the parametric process families covered by theory.
Abstract: Reinforcement learning (RL) seeks to optimize sequential decisions to maximize population-level benefits over time. However, when deployed in high-stakes settings such as healthcare, RL decisions might systematically restrict some subpopulation's access to valuable resources or services in a manner contrary to the values and goals of stakeholders. Counterfactual fairness (CF) offers a promising framework to address this problem based on causal reasoning. This paper develops a data preprocessing algorithm that, when used in tandem with policy learning, enables CF in RL. Our algorithm relies on a novel quantile distribution mapping method for sequentially estimating the counterfactual states and rewards in the data preprocessing step, subsuming common additivity assumptions used for counterfactual prediction as special case. We theoretically prove that the per-step level of counterfactual unfairness and infinite-horizon suboptimality gap can be bounded under mild regularity conditions. We also empirically test our algorithm in numerical experiments as well as in application to a real-world interventional digital health dataset.
PaperID: 3928, Poster
Authors:
Kairong Luo, Zhenbo Sun, Xinyu Shi, Shengqi Chen, Bowen Yu, Yunyi Chen, Chenyi Dang, Tao H Tao, Hui Wang, Fangming Liu, Kaifeng Lyu, Wenguang ChenAbstract: Fully open pretraining now has increasingly reproducible recipes and accessible web-scale corpora. What remains missing is a principled way to merge scored open-web data. Corpora such as DCLM Baseline and FineWeb-Edu are built with different filters and scorers, so their raw scores are not directly comparable. Because web data is large, heterogeneous, and token-budget dominant, choosing an additional retained fraction is a consequential compute decision rather than a minor sampling detail. We formulate this setting as heterogeneous top-p mixing: choosing top-p retained fractions for each corpus along its score distribution under a fixed total token budget. We propose quantile benchmarking, a proxy-training framework that probes fixed score quantiles from each source and compares them through a shared downstream-utility metric. The resulting measurements reveal benchmark-dependent source orderings, large within-source utility variance, and a strong knowledge bias in current scorers. In particular, higher-score regions often degrade broader commonsense benchmarks such as PIQA and HellaSwag. We then fit benchmark-conditioned utility curves and derive top-p mixing cutoffs through a global utility threshold. 40B-token ablation experiments show that quantile-benchmarking-guided thresholds improve over both uniform mixing and heuristic top-p filtering. Finally, we apply the resulting recipe to train a 2B model over 2.2T tokens with a fully open recipe, reaching the Pareto frontier of parameter efficiency among compared fully open models of similar scale.
PaperID: 3929, Poster
Abstract: Large Language Model (LLM) agents have shown strong results on multi-turn tool-use tasks, yet they typically operate in isolation during training, failing to leverage skills accumulated across episodes. Existing experience-augmented methods address this by organizing past trajectories into retrievable libraries, but they retrieve skills only once based on the initial task description and hold them constant throughout the episode. In multi-turn settings where observations change at every step, this static retrieval becomes increasingly mismatched as episodes progress. We propose SAPO (Step-Level Skill-Augmented Policy Optimization), a reinforcement-learning framework that retrieves relevant skills at each decision step conditioned on the current observation. SAPO operates through three components: (i) step-level observation clustering that groups structurally equivalent environmental states for efficient cluster-indexed retrieval; (ii) a self-evolving skill bank that distills successful strategies and failure patterns through score-based admission and rate-limited extraction; and (iii) policy optimization with step-level credit assignment for fine-grained advantage estimation across multi-turn episodes. The skill bank evolves alongside the policy through semantic analysis rather than gradient updates. On long-horizon multi-turn agent benchmarks (ALFWorld, WebShop, and seven search-augmented QA tasks), SAPO achieves 93.5% on ALFWorld, 76.3% success on WebShop, and 60.9% average across QA tasks, outperforming both standard RL and prior skill- and experience-augmented baselines. Our code is available at \urlhttps://anonymous.4open.science/r/slea-rl-4D1E/.
PaperID: 3930, Poster
Abstract: The Procrustes-Wasserstein problem asks to match the vectors in two unlabeled point clouds X and Y, which are related via a latent orthogonal transformation. This problem appears in unsupervised representation alignment, domain adaptation, and geometric matching, but its computational theory remains limited. We study a planted Gaussian model in which Y is obtained from X via an unknown relabeling, an unknown orthogonal transformation, and additive noise. Departing from the isotropic setting considered in prior work, we assume that the covariance of X is anisotropic, with a power-law spectrum. While closer to real-world data, this structure also enables a directions-first approach: Estimate the dominant eigenspaces of both point clouds, then recover the matching. We prove that a polynomial-time version of this procedure exactly recovers the true relabeling with high probability under explicit scaling conditions. To our knowledge, this is the first computational recovery guarantee for Procrustes-Wasserstein with nontrivial observational noise and dimensions beyond \log n, where n is the number of points. Experiments on synthetic data and 3D shapes support the theory, showing that anisotropy favors directions-first methods, can lower runtime and improve recovery.
PaperID: 3931, Poster
Authors: Holger Heidrich, Sarah Müller, Andreas Schilling
Abstract: Deploying deep learning in high-stakes settings demands models that are not only accurate but interpretable. Interpretable-by-design models, exemplified by BagNet, address this by restricting each local class predictor to a small spatial patch, yielding inherently explainable predictions---but at the cost of capturing large-scale image features, resulting in a substantial accuracy penalty. We propose FovBagNet, a foveated extension of BagNet inspired by the mammalian retina. Each local predictor depends on a stack of multiscale patches centered at its location, providing high-resolution detail at the center and progressively coarser context toward the periphery. FovBagNet achieves a top-1 ImageNet accuracy of 0.74, closing most of the 11-point gap between BagNet-33 (0.65) and ResNet-50 (0.76), while retaining the inherent interpretability of BagNet in the form of spatially accurate saliency maps. We also identify a conceptual limitation of the original BagNet saliency maps and empirically assess its impact.
PaperID: 3932, Poster
Abstract: Robust deployment of clinical electroencephalography (EEG) models requires source-trained models to generalize to unseen hospitals without target-domain calibration, hospital identifiers, or clinical reports at inference. This is challenging because acquisition protocols, hardware, montages, and patient populations perturb low-level EEG statistics, while clinical interpretation is organized around higher-level semantic concepts. We argue that this abstraction mismatch limits conventional signal-level invariant learning: aligning feature distributions alone cannot distinguish clinically irrelevant variation from diagnostic content. We propose Text-Guided Invariant Learning (TG-IL), a deployment-oriented framework that uses clinician-authored reports only during training as semantic anchors for EEG representation learning. TG-IL combines stochastic recording-level aggregation, a variational semantic information bottleneck, and prototype-guided EEG--report alignment to preserve report-derived clinical semantics while mitigating site-specific nuisance variation. At inference, all text-side modules are discarded and the model operates solely on EEG. On a three-hospital clinical EEG benchmark, TG-IL improves zero-shot cross-hospital generalization and worst-case robustness while reducing hospital-specific information in learned representations.
Abstract: Pose-guided text-to-image generation often suffers from limb distortions and feature crosstalk in complex multi-person scenarios. While existing UNet-based adapters struggle with long-range spatial dependencies, emerging Multimodal Diffusion Transformers (MM-DiTs) offer superior global modeling. However, naive signal concatenation in MM-DiTs severely disrupts pre-trained latent distributions. To address this, we propose TrioPose, a native pose-driven framework built upon the SD3.5M architecture. Specifically, we introduce a Triple-Stream Pose-Aware DiT (TSPA-DiT) that treats pose as an independent modality. It employs layer-wise activation and zero-initialized dual-residual injection to smoothly enforce geometric constraints while preserving pre-trained latent stability. To resolve severe multi-instance occlusions, we design a Learnable Relational Bias Mask that categorizes topological connectivity into fine-grained physical states, mapping them into continuous attention soft constraints to effectively decouple inter-instance interference. Furthermore, a Pose-Guided Spatial Loss Weighting strategy modulates the native diffusion objective using heatmap-derived error maps, focusing anatomical supervision strictly on distortion-prone regions. Extensive experiments demonstrate that TrioPose achieves state-of-the-art performance across challenging benchmarks, including Human-Art, CrowdPose, and OCHuman. Notably, it attains an AP of 64.33 on Human-Art, representing a 30 % improvement over prior arts, while setting new standards for visual fidelity and text-image semantic alignment in complex multi-human generation.
Authors: Mohamed Eltahir, Ayash, Ali Habibullah, Tanveer Hussain, Naeemullah Khan
Abstract: Long-video understanding in VLMs is bottlenecked by a single monolithic forward pass over thousands of frames at quadratic attention cost. A common mitigation is to first select a small subset of informative frames before the forward pass; common for training-free selectors via auxiliary encoder-space similarities. Such signals are capped by contrastive pretraining, which usually fails on reasoning-heavy queries (negation, cross-frame counting, holistic summarization). We propose GridProbe, an efficient training-free posterior-probing inference paradigm that scores evidence in answer space using a frozen VLM's own reasoning and then selects question-relevant frames adaptively, resulting in sub-quadratic attention cost with little to no accuracy loss. We arrange frames on a K×K grid and run lightweight row R and column C probes, where each probe reads its peak posterior as a query-conditioned confidence. The outer product of R and C yields an interpretable importance map whose skewness and kurtosis drive Shape-Adaptive Selection, a closed-form rule that reliably replaces the fixed frame budget M with a per-question M_\mathrmeff. We show empirically that M_\mathrmeff, surprisingly, tracks intrinsic question difficulty without ever seeing the answer, a sign of test-time adaptive compute. On Video-MME-v2, GridProbe matches the monolithic baseline within 1.6 pp Avg Acc at 3.36× TFLOPs reduction, while on LongVideoBench it Pareto-dominates the baseline (+0.9 pp at 0.35× compute). Because the selector and QA models can be decoupled, pairing a small 2B selector with a stronger 4B or 8B QA is strictly Pareto-dominant over the 2B monolithic baseline (up to +4.0 pp at 0.52× compute, on average), with no retraining. Finally, the interpretability of the importance maps opens future avenues for behavioral diagnostics, grounding, and frame-selection distillation.
PaperID: 3935, Poster
Abstract: Personalizing Multimodal Large Language Models (MLLMs) as intelligent assistants has recently emerged as an active research topic, where an MLLM is expected to go beyond generic, one-size-fits-all replies, and instead generate responses grounded in user-specific objects, entities, and concepts encountered during interactions. Nevertheless, existing studies mainly focus on personalizing MLLMs on a static, predefined set of personalized concepts and require users to manually and frequently extend this set to accommodate newly emerging concepts, which is infeasible in dynamic-world scenarios. To address this limitation, we recast conventional MLLM personalization as an Open-Set problem, and propose TAME-O, the first long-context open-set personalized MLLM assistant. Specifically, TAME-O introduces a training-free, plug-and-play Concept Reconciliation Skill (CR Skill). It enables the assistant to automatically unlock novel concepts during prolonged human–MLLM interactions while preventing confusion with previously unlocked ones, allowing the MLLM to continually grasp more and more concepts and progressively refine the modeling of known ones. Extensive experiments under the open-set personalization setting underscore the efficacy of our method.
PaperID: 3936, Poster
Abstract: Domain translation requires a delicate balance between realism, input-output alignment, and output diversity. Entropic optimal transport (EOT) provides a principled formulation of this trade-off, but its practical use remains challenging. Directly maximizing entropy of a stochastic transport plan requires evaluating log densities of generator-induced distributions, which scales poorly with dimension. In contrast, diffusion models make density information accessible in high dimensions through noising the corresponding distributions and approximating their score functions. In this work, we introduce the diffusive entropy regularizer that measures diversity after progressively noising the conditional output distribution, making the entropy term compatible with score estimation. The resulting retains the main theoretical guarantees of EOT: a unique solution, controllable diversity, and convergence to unregularized OT solution in the zero-regularization limit. We then introduce , a practical algorithm for approximately solving the Diffusive EOT that combines diffusion-based distribution matching with transport cost and diffusive entropy penalties. DM-EOT supports both fast one-step generation and higher-fidelity diffusion-based multi-step sampling. Both variants achieve comparable or superior performance on unpaired domain translation benchmarks among one-to-many baselines with comparable inference cost and diversity.
PaperID: 3937, Poster
Authors: Lennert De Smet
Abstract: Approximate Bayesian inference has driven significant advances in probabilistic machine learning at the basis of modern generative modelling. However, many successful Bayesian methods, such as variational inference or sequential Monte Carlo, often struggle in discrete domains with complex structure. In particular, their application to the challenges of neurosymbolic AI, i.e. learning and reasoning under uncertainty, is largely unexplored. Our contributions are twofold. First, we show that the majority of neurosymbolic problems are problems of posterior inference motivating the use of approximate Bayesian inference for neurosymbolic AI. Second, we adapt variational inference to the neurosymbolic setting by exploiting available symbolic structure. Moreover, we outline what properties a reasonable variational approximation to a neurosymbolic problem should have and how to enforce these properties by connecting to statistical estimation and statistical representation theorems. The result is a principled approach to learning and reasoning at scale that maintains statistical guarantees at test time. This work illustrates the opportunities at the intersection of approximate Bayesian inference, and learning and reasoning under uncertainty.
PaperID: 3938, Poster
Abstract: In multimodal learning, modality imbalance often causes one modality to dominate the optimization process, preventing the model from fully exploiting cross-modal interactions. Recent balance-oriented methods, such as Multimodal Competition Regularizer (MCR), have achieved strong performance by alleviating modality imbalance during training. However, existing dynamic balancing methods mainly focus on adjusting modality contributions and do not explicitly consider how modality imbalance affects the formation of synergistic information in multimodal fusion. Yet such synergy is a key source of multimodal advantage, as it reflects predictive gain that cannot be recovered from unimodal evidence alone. To address this gap, we propose Multimodal Synergistic Competition Regularizer (MSCR), a multimodal fusion method that jointly promotes balanced modality utilization and synergistic information discovery. MSCR contains two complementary objectives: an anti-dominance objective that suppresses excessive unimodal influence and creates room for synergistic information discovery, and a synergy objective that explicitly encourages the fused predictor to surpass a non-synergistic unimodal reference. We further adopt an adaptive modulation strategy to emphasize synergy enhancement on samples that require stronger collaborative modeling. Experiments show that MSCR consistently improves performance, enhances measured synergistic information, and remains effective under transformer-based fusion backbones.
Authors:
Yu Li, Menghan Xia, Gongye Liu, Xintao Wang, Conglang Zhang, Lei Ke, yuxuan lin, Ruihang Chu, Pengfei Wan, Kun Gai, Yujiu YangAbstract: Despite being a pivotal frontier, interactive world modeling remains underexplored in terms of the versatile controllability required by practical scenarios. To bridge this gap, we present AnchorWorld, a framework that advances egocentric simulation through enhanced interaction integrity and a flexible mechanism for world customization. First, we utilize 3D human motion as the primary interaction modality. To complement the out-of-view or truncated body parts in egocentric views, we introduce an auxiliary training supervision that incorporates exogenous viewpoints decoupled from the agent’s first-person sensorium. It allows the model to observe the agent's full-body positioning relative to the environment, facilitating a more robust spatial grounding of human-world interactions. Furthermore, we propose a simple yet effective mechanism for customizing self-evolving worlds. This is achieved by defining anchor views within a unified world coordinate system, coupled with textual descriptions dictating the dynamic evolution of local scenes. Experimental results show that AnchorWorld significantly outperforms state-of-the-art baselines, while ablation studies validate the effectiveness of our key designs. Notably, our customization scheme exhibits promising spatio-temporal geometric consistency and adheres strictly to the prescribed evolutionary dynamics.
PaperID: 3940, Poster
Abstract: Reconstructing high-fidelity and controllable digital twins from multi-modal sensory observations is a fundamental problem in physical AI applications. While existing per-scene optimization methods handle dynamic driving scenes well, they are slow to optimize, have artifacts at novel viewpoints, and require costly manual annotations for decomposing the scene into controllable instances. Self-supervised generalizable reconstruction methods enable faster and more scalable reconstructions by learning from large datasets, but existing approaches do not fully leverage multi-modal inputs (e.g., camera and LiDAR) and lack decomposition capabilities. To address these limitations, we propose STRIDE, a self-supervised, generalizable and decomposable 4D driving scene reconstruction method. By processing multi-modal sensory inputs in a common 3D space with a Point-Transformer, STRIDE efficiently discovers spatio-temporal correspondences necessary for accurate flow prediction. By incorporating latent instance tokens, STRIDE is the first to perform feed-forward reconstruction and learnable decomposition without manual annotations. Experiments on public driving datasets show STRIDE achieves state-of-the-art performance on recovering geometry, appearance, and 3D flow. Moreover, we demonstrate how learned decompositions can enable dynamic instance manipulation and controllable simulation.
Abstract: Piecewise-linear (PWL) optimization problems arise in many mixed-integer programming (MIP) optimization applications, including portfolio optimization, workforce scheduling, and resource allocation. But solving them to global optimality remains computationally expensive because branch-and-bound repeatedly solves LP relaxation subproblems. Existing solvers are largely CPU-centric, leaving the scalability of modern GPUs underutilized. Few prior GPU-accelerated branch-and-bound either targets neural network which is not suitable for general PWL optimization, or accelerates only auxiliary subroutines such as strong branching heuristics within CPU-centric MIP solvers. To bridge this gap, we propose \textttB^3-PWL, a GPU-centric batched branch-and-bound framework for piecewise-linear optimization with Special Ordered Set of type 2 (SOS2) constraints. Our method solves batches of LP relaxation subproblems concurrently on the GPU using a first-order primal-dual solver, enabled by a specialized batched block-tiled sparse matrix kernel. To complement bound computation, we further introduce a unified feasibility search module that combines an SOS2 repair primal heuristic with a batched feasibility pump to rapidly obtain feasible incumbents and improve pruning efficiency. On a benchmark of 43 PWL-MIP instances, \textttB^3-PWL achieves a 9.25× geometric-mean speedup over NVIDIA cuOpt while reaching high-quality feasible incumbents on every tested instance, demonstrating the potential of first-order LP methods as the central engine of GPU-accelerated branch-and-bound.
PaperID: 3942, Poster
Abstract: LLMs are powerful but incur substantial inference cost, which can be exploited by malicious users to induce overly long outputs, increasing latency and operational expense. Prior resource consumption attacks either coerce repetition or suppress the end-of-sequence token by lowering its logit value. However, with autoregressive, samplingbased generation, the search space grows exponentially with length and small deviations or unstable trajectories can derail optimization, limiting attack effectiveness and stability. To address this, we introduce Chatter Attack, a novel class of resource consumption attacks that optimize adversarial prompts to induce predefined short target phrases, thereby avoiding search-space explosion while still eliciting long responses. Specifically, Chatter Attack triggers self-doubt using short doubt-inducing phrases as the target, and sustains this state throughout generation via inference attention loss and global entropy loss. To extend to black-box settings, we further propose a novel Bayesian Prompt Optimization (BPO) method, which models the objective function globally by constructing a discrete token kernel, efficiently exploring the entire search space using prior information. Extensive experiments across multiple models and datasets show that our method outperforms baselines, with maximum response length increases to 31.5×, and revealing practical resource-depletion risks for LLM services.
PaperID: 3943, Poster
Abstract: We study identifiability and estimation of the direct causes, that is, the parents, of a designated target Y in a structural causal model using only observational data. Our main assumption is an additive-noise model for Y: the value of Y equals some function of its parents plus noise that is independent of those parents. We allow for general functional relationships and hidden confounding among all other variables. We prove that under a `no-backward' additive-noise condition the true parent set, \mathrmPA(Y), is identified from the observational distribution by a simple population principle: among all candidate variable sets that make the regression residual of Y independent of the regressors, choose the one with the smallest residual variance, and—if several tie—the smallest set. To justify the ``no-backward'' condition, we propose a novel identifiability scheme based on identifiability witnesses: we prove that in finite-dimensional analytic model classes, for which we have an identifiability witness (that is, a single identifiable model), residual independence for sets containing descendants of \(Y\) holds only on Lebesgue-null exceptional parameter sets. This scheme is strong enough to recover known identifiability results. For finite data, we propose Independent Risk Minimization (IndRM), which offers a simple, local, and non-interventional route to isolating the direct causes of Y in multivariate systems, avoiding global model assumptions and full-graph search.
Abstract: We study instruction following in multi-task reinforcement learning, where an agent must zero-shot execute novel tasks not seen during training. In this setting, linear temporal logic (LTL) has been adopted as a powerful framework for specifying structured, temporally extended tasks. While existing approaches successfully train generalist policies, they often struggle to effectively capture the rich logical and temporal structure inherent in LTL specifications. In this work, we address these concerns with a novel approach to learn task representations that facilitate training and generalisation. Our method conditions the policy on sequences of Boolean formulae constructed from a finite automaton of the task. We propose a hierarchical neural architecture to encode the logical structure of these formulae, and introduce an attention mechanism that enables the policy to reason about future subgoals. Experiments in a variety of complex domains demonstrate the strong generalisation capabilities and superior performance of our approach.
Authors: Yixin Chen, Alan Kuhnle
Abstract: Submodular functions---functions exhibiting diminishing returns---are central to machine learning. When the objective is monotone and non-negative, the greedy algorithm achieves a tight 63% approximation. But many practical objectives incorporate costs that make them negative on some inputs, and all existing multiplicative guarantees require non-negativity. Prior work handles negativity through additive bounds for the special class of decomposable functions and non-monotonicity through partial-monotonicity parameters, but these address each difficulty in isolation and neither extends the classical structural theory. We extend curvature---a parameter measuring how far a function deviates from linearity---to all submodular functions, handling both non-monotonicity and negativity through a single classical concept. A greedy algorithm with pruning achieves a curvature-controlled multiplicative ratio for any submodular function, including those taking negative values---the first such guarantee beyond monotonicity and non-negativity. In the non-monotone regime 1 \le c_g < 2.2, the bound strictly beats the best known uniform ratio of 0.401 (for non-negative f), and it recovers the classical (1-e^-c_g)/c_g guarantee for monotone functions. A multilinear-extension variant extends the framework to general combinatorial constraints via multilinear relaxation. Experiments on cost-penalized experimental design, coverage, feature selection, and a curvature sweep on Multi-News passage selection support the theory.
Abstract: We study privacy amplification for differentially private model training with matrix factorization under random allocation (also known as the balls-in-bins model). Recent work by Choquette-Choo et al. (2025) proposes a sampling-based Monte Carlo approach to compute amplification parameters in this setting. However, their guarantees either only hold with some high probability or require random abstention by the mechanism. Furthermore, the required number of samples for ensuring (\epsilon,\delta)-DP is inversely proportional to \delta. In contrast, we develop sampling-free bounds based on Rényi divergence and conditional composition. The former is facilitated by a dynamic programming formulation to efficiently compute the bounds. The latter complements it by offering stronger privacy guarantees for small \epsilon, where Rényi divergence bounds inherently lead to an over-approximation. Our framework applies to arbitrary banded and non-banded matrices. Through numerical comparisons, we demonstrate the efficacy of our approach across a broad range of matrix mechanisms used in research and practice.
Abstract: Vision-language-action (VLA) models have achieved strong performance in embodied manipulation, but still lack a clear mechanism to balance behavioral stability with task-semantic sensitivity. We identify two complementary failure modes. Under task-preserving changes, where task semantics remain unchanged but scene appearance varies (e.g., style, illumination, clutter, or paraphrasing), policies often exhibit unnecessary action drift. Conversely, under semantic-breaking changes, where key task semantics such as the target object or constraint are altered, policies frequently fail to produce sufficiently distinct behaviors and instead follow the original trajectory. To address this gap, we propose BAS-VLA, a task-semantic action calibration framework built on top of a frozen base VLA. BAS-VLA adopts a breaking-centered calibration core as the default path, and introduces a selective evidence-gated preserving auxiliary that activates only when nuisance variation is detected while task semantics remain consistent. On the OpenPI-pi0.5 / LIBERO-Object Milk-Swap benchmark, BAS-VLA maintains high success on clean (98.0%) and semantics-preserving conditions (97.5%), while reducing clean-criterion success to 0.0% under deliberate target-object swaps, demonstrating strong stale-task suppression and task-semantic separation. On validated style-preserving shifts, it improves success from 42% to 70% without degrading clean performance. These results highlight that reliable VLA behavior requires moving beyond appearance robustness toward explicit task-semantic action calibration.
Authors: Shun Kenney, Teppei Suzuki
Abstract: The remarkable scalability of Transformers has expanded their application to 3D computer vision, where camera-aware positional encoding is crucial for providing spatial cues in multi-view geometry. Recent advancements have established the practice of using camera parameters---such as extrinsics or projection matrices---as relative positional encoding into the query, key, and value vectors of the attention mechanism. However, when scaling up the training recipe of novel view synthesis (NVS) models with the camera-based positional encoding, we observe a significant issue: model performance stagnates in the late stages of training. In this paper, we investigate the cause of the performance bottleneck when scaling up and demonstrate that storing rotation and translation given by the positional encoding in the same dimensions of the value vector causes indeterminacy in their independent identification, hindering training scalability. To address this, we propose Decoupled Pose Positional Encoding (DPPE), a novel camera-based positional encoding that explicitly decouples rotation and translation. Extensive evaluations on NVS tasks demonstrate that DPPE enables stable long-term training even in scaled-up training setup. Furthermore, it exhibits superior generalization performance in extrapolation settings, such as handling an increased number of viewpoints and zoom-in scenarios.
Abstract: Proteolysis-targeting chimeras (PROTACs) induce protein degradation through coordinated interactions among a degrader molecule, a target protein, and an E3 ubiquitin ligase. Yet computational prediction is often framed as if degradation were an intrinsic property of the degrader alone, overlooking the tuple-level context that determines activity. This mismatch is especially consequential in public PROTAC datasets, where degradation labels are sparse but molecule–target–E3 records are abundant. We introduce DegradeQuery, a tuple-conditioned framework that uses unlabeled records as structured relational evidence rather than incomplete labeled examples. DegradeQuery first performs counterfactual tuple pretraining, learning to distinguish observed molecule–target–E3 tuples from alternatives generated by replacing the target, the E3 ligase, or both. This objective induces a conditional compatibility representation before degradation labels are used, enabling the model to capture how degrader activity depends on biological context. The pretrained representation is then fine-tuned for high/low degradation prediction. On the official PROTAC-8K benchmark, DegradeQuery achieves 0.9065 AUROC and 0.8500 accuracy, outperforming reported semi-supervised results without pseudo-labeling, teacher models, distillation, or ensembles. Ablations, scaffold holdout, target–E3 holdout, and target-wise few-shot adaptation further support PROTAC degradation prediction as a tuple-conditioned rather than molecule-only problem.
Authors:
Wei Tang, Jinpei Han, Kangning Cui, Mattia Carletti, Fredrik K Gustafsson, Shreyank Gowda, Patitapaban Palo, Anshul Thakur, Lei Clifton, Jean-michel Morel, Raymond H Chan, David Clifton, Xiao GuAbstract: Electrocardiogram (ECG) foundation models pretrained on typical diagnostic 10-second ECG segments, have demonstrated strong transferability across a range of clinical applications. However, many real-world applications produce recordings that are typically longer, and are varied in duration during inference time. These 10-second models have no built-in way to combine information across time. Extending them to longer horizons introduces two challenges: challenges that limit meaningful temporal aggregation. We propose a parameter-efficient framework that extends pretrained ECG foundation models to longer and variable-length ECGs without retraining the backbone. Guided by a frozen pretrained 10-second model, we introduce a lightweight plug-in module that extends the model in two complementary ways: (i) structurally compatible long-sequence processing and (ii) semantically informed temporal modeling. Experiments on multiple long-horizon ECG tasks, datasets, and foundation model backbones demonstrate that our method enables robust long-horizon extension from pretrained snapshot models, consistently outperforming sliding-window and pooling-based baselines with strong parameter efficiency.
PaperID: 3951, Poster
Abstract: Large language model (LLM)-native advertising has emerged as a viable monetization channel as LLM service providers seek sustainable revenue streams. However, integrating native ads into LLM responses introduces two largely understudied challenges: (1) controlling ad intensity---how prominently advertising content appears---and (2) ad labeling---identifying which segments of the generated text constitute advertising. We propose AdKnob, a framework jointly addressing both challenges. For ad intensity control, we introduce new ad control tokens without pretrained semantic associations and align the model via direct preference optimization to generate ads at each desired level. For ad labeling, we propose an attention rollout-based attribution mechanism that identifies ad segments by tracing their attribution to input ad-related tokens, requiring no additional inference cost. We further construct MI-KnobSet, the first multi-level ad intensity preference dataset for LLM-native ads. Experiments across six LLMs and human evaluations show that AdKnob achieves uniformly spaced ad intensities where baselines collapse to a narrow range, achieves the highest ad labeling accuracy while being over 500× faster than the strongest baseline, and is preferred over baselines by human evaluators up to 89.6% of the time. Code and data will be publicly available.
Abstract: Creating images from noise is image generation; reconstructing fine details from coarse inputs is super-resolution. Despite their practical differences, both can be understood as reversing information loss across scales. We introduce SKILD, a Scale-invariant K-Space Image Learning Diffusion model that unifies generation and continuous super-resolution within a single unconditional framework. Both natural images and critical physical systems exhibit scale invariance, and we leverage it to design a forward process that attenuates image content from fine to coarse scales while injecting spectrum-matched Gaussian noise, making scale an explicit coordinate of the diffusion dynamics. The same trained reverse process performs generation and continuous super-resolution by varying only the starting timestep: no task-specific architecture, no conditioning branch, no classifier-free guidance, no retraining per scale factor. Empirically, SKILD reaches FID 2.65 and Inception Score 9.63 on unconditional CIFAR-10, performs 2×–8× super-resolution on ImageNet from a single unconditional checkpoint while outperforming conditional models across perceptual metrics, and reconstructs critical Ising models whose connected four-point correlations closely track the ground truth.
PaperID: 3953, Poster
Authors: Tengwei Song, Robert Hoehndorf
Abstract: LLM agents are flexible in language interaction, but using them as long-horizon planners is costly and prone to hallucinated or invalid actions. This paper targets a different role for LLMs: constrained bridges between language and symbolic planning. We propose MemPlan, a memory-conditioned PDDL planning framework for partially observable interactive text environments. MemPlan keeps the PDDL domain fixed and reconstructs the PDDL problem at each step from parsed observations, candidate hypotheses, and interaction memory. Positive graph memory assigns relevance-based action costs to candidate hypotheses for missing objects, locations, or conditions, guiding the planner toward more plausible checks. Negative memory records execution-refuted assumptions and compiles them into exclusion constraints, preventing repeated attempts at candidates already contradicted by feedback. A unified schema-conditioned language--symbol interface, trained offline with supervised fine-tuning, connects text environments to the planning loop through observation-to-fact parsing and symbolic-action-to-command grounding. Experiments on TextWorld, ALFWorld, and Robotouille show that MemPlan improves task completion and step efficiency over language-only and PDDL-aware baselines, while substantially reducing online token cost compared with LLM-planning baselines such as ReAct. The same fine-tuned interface is also reused across TextWorld task types without task-specific fine-tuning. Our code and data are available at: \urlhttps://anonymous.4open.science/r/MemPlan-1EE5/.
PaperID: 3954, Poster
Abstract: Discrete diffusion language models (DLMs) enable parallel generation but usually rely on fixed forward processes, such as masking, uniform noise, or precomputed hierarchies. We investigate whether the hierarchical forward process can instead be learned while preserving continuous-time tractability. We propose Learnable Hierarchical Diffusion Language Model (LHDLM), which extends a learnable column-stochastic token-cluster map. This many-to-many hierarchy preserves a CTMC with block-conditional transition and closed-form CT-ELBO, while recovering HDLM in the one-hot limit and masked diffusion in the collapsed limit. We identify degenerate hierarchy maps caused by learning the same map used in both forward targets and model-induced cluster predictions, and mitigate them with an information-retention regularizer. Empirically, LHDLM is competitive with discrete diffusion approaches and reveals soft token structures --- cluster-agnostic function tokens and sharp content-token assignments --- that cannot be represented by hard partitions.
Abstract: Assessing the capabilities and risks of frontier AI systems is a critical area of research; recent work has shown that repeated sampling can dramatically increase both. For instance, repeated sampling allows models to solve difficult math and coding problems, but it may also increase their risk of being jailbroken. Such results raise a crucial question: how can one accurately predict a model's behavior when scaled to a massive number of attempts, given a vastly smaller sampling budget? This question is directly relevant to model providers, who serve hundreds of millions of users daily, and to governmental regulators, who seek to prevent harms. To answer this question, we make three contributions. First, we find that standard methods for fitting these laws suffer from statistical shortcomings that hinder predictions, especially in data-limited scenarios. Second, we remedy these shortcomings by introducing a robust estimation framework, which uses a beta-binomial distribution to generate more accurate predictions from limited data. Third, we propose a dynamic sampling strategy that allocates a greater budget to harder problems. Combined, these innovations enable more reliable prediction of rare risks and capabilities at a fraction of the computational cost.
Abstract: The high-dimensional features extracted from large-scale unlabeled data via various pretrained models with diverse architectures are referred to as heterogeneous multiview data. Most existing unsupervised transfer learning methods fail to faithfully recover intrinsic subspace structures when exploiting complementary information across multiple views. Therefore, a fundamental challenge involves constructing sparse similarity graphs that preserve these underlying subspace structures for achieving semantic alignment across heterogeneous views. In this paper, we propose a sparse attention graph learning (SAGL) method that learns subspace-preserving sparse attention graphs from heterogeneous multiview data. Specifically, we introduce a bilinear attention factorization scheme to capture asymmetric similarities among the high-dimensional features, which breaks the symmetry bottleneck that is inherent in the traditional representation learning techniques. A dynamic sparsity gating mechanism then predicts a feature-specific compression factor for adaptively controlling the topological contributions of neighbors. Furthermore, we employ a structured sparse projection via \alpha-entmax to generate subspace-preserving sparse attention graphs for individual views. SAGL leverages these view-specific graphs to conduct sparse information aggregation, yielding discriminative representations for multiview learning tasks. In addition, we provide a rigorous theoretical analysis that bridges differentiable sparse attention and probability simplex constraints. Extensive experiments conducted on multiple benchmark datasets demonstrate that SAGL consistently outperforms the state-of-the-art unsupervised transfer learning approaches.
PaperID: 3957, Poster
Abstract: Many decision-making problems involve multiple objectives, where actions affect several performance metrics, fairness criteria, or safety constraints. We study multi-objective causal bandits with multiple reward nodes, where actions correspond to interventions on subsets of variables in a causal graph and performance is evaluated through Pareto optimality. Existing approaches to multi-objective bandits typically define optimality at the level of individual actions, using pairwise dominance or scalarization. In contrast, randomized decision rules induce convex combinations of action-level reward vectors, so the relevant achievable set is the convex hull of these rewards. Thus, an intervention may be suboptimal even if no single intervention dominates it, because it can be dominated by a randomized policy over other interventions. We formalize this policy-level phenomenon and show that it has consequences for both causal action-space reduction and online learning. First, we provide a necessary-and-sufficient graphical characterization of possibly Pareto-optimal minimal intervention sets (PPOMISs). Our characterization yields a minimal, sound, and complete candidate intervention family from the graph alone, correcting prior formulations that may include intervention sets that are never Pareto-optimal under any compatible structural causal model. Second, over this reduced intervention space, we design an online UCB-style algorithm that eliminates actions using dominance tests against convex combinations of other actions, and prove logarithmic gap-dependent Pareto regret. Finally, we study constrained multi-objective causal bandits, where feasibility and optimality may be realized only with randomized policies, and develop an online UCB-style constrained bandit algorithm with sublinear regret and constraint-violation guarantees.
Abstract: Discovering object-centric representations from images can significantly enhance the robustness, sample efficiency and generalizability of vision models. Works on images with multi-part objects typically follow an implicit object representation approach, which fail to recognize these learned objects in occluded or out-of-distribution contexts. This is due to the assumption that object part-whole relations are implicitly encoded into the representations through indirect training objectives. We address this limitation by proposing a novel method that uses explicit graph representations for parts and introduce a new co-part algorithm for multi-part object discovery. The part representations are refined with a part-centric regularization objective. We then introduce three benchmarks to evaluate the robustness of object-centric methods in recognizing multi-part objects within occluded and out-of-distribution settings. Experimental results on simulated, realistic, and real-world images show marked improvements in the quality of discovered objects compared to state-of-the-art methods, as well as the accurate recognition of multi-part objects in occluded and out-of-distribution contexts. We also show that the discovered object-centric representations can more accurately predict key multi-part object properties in a downstream task, highlighting the potential of our method to advance the field of object-centric representations.
Authors:
Enzhi Zhang, Du Wu, Rui Zhong, Cong Ma, Isaac Lyngaas, Amir K Ziabari, Xiao Wang, Peng Chen, Tao Luo, Toshio Endo, Fumiyoshi Shoji, Kento Sato, Kentaro Uesugi, Takayuki Nonoyama, Ryuji Kiyama, Masahiro Yoshida, Tezuka Masaru, Tetsuya Ishikawa, Satoshi Matsuoka, Masaharu Munetomo, Mohamed WahibAbstract: Self-supervised pre-training with Vision Transformers, including Masked Autoencoders (MAE), is difficult to apply to gigapixel scientific images. Random masking is poorly matched to the structured, multi-scale morphology of scientific data, while uniform tokenization produces prohibitively long sequences that make O(N^2) attention impractical. We propose \method, a structure-guided masked autoencoding framework for ultra-high-resolution scientific images. \method couples two components: a content-adaptive quadtree tokenizer that compresses gigapixel images into a fixed-length sequence, and a structure-conditioned masking process that biases reconstruction toward spatially informative regions. To stabilize this process across scales, we introduce Damped Accumulation (DA), which aggregates signal-dependent responses across the tree into a structure canvas used to guide masking. The resulting pre-training task preserves fine microstructure while remaining compatible with standard ViT encoders and MAE-style reconstruction. Across electron microscopy, whole-slide optical microscopy, and X-ray CT datasets, \method consistently outperforms MAE baselines. It achieves 95.68% Dice on the 8K×8K×28K SpringXCT dataset, improving over the strongest MAE baseline by +9.70 points, and 83.21% Dice on the 32\textK^2 WSI PAIP dataset, improving by +16.84 points, while providing up to a 24.8× inference speedup
PaperID: 3960, Poster
Abstract: Long-tailed recognition is commonly formulated as a problem of imbalanced supervision, where rare classes suffer from insufficient training examples. In this work, we study a complementary source of difficulty: under severe scarcity, tail classes may be forced into fine-grained discrimination before the model has acquired sufficiently reliable coarse semantic structure. We refer to this phenomenon as granularity mismatch under scarcity. To address it, we propose BasicLT, a two-stage framework that augments an underlying fine-grained recognition model with basic-level abstraction and commitment-guided selective differentiation. In Stage 1, BasicLT learns an explicit basic-level branch and aligns it with the family-level distribution induced by the fine-grained predictor, encouraging stable coarse semantic organization before additional fine-grained refinement is introduced. In Stage 2, a frozen Stage 1 teacher provides stable family-level commitment signals, and a family-conditioned residual refinement module is trained only on committed medium-shot and few-shot samples. At inference time, the residual correction is activated conservatively according to student-side commitment, so that fine-grained refinement acts as a conditional within-family correction rather than an unconditional auxiliary classifier. We evaluate BasicLT on CIFAR-100-LT, CIFAR-10-LT, ImageNet-LT, iNaturalist 2018, and Places-LT. Across these benchmarks, BasicLT achieves competitive performance against strong long-tailed baselines, with the most consistent gains appearing in medium-shot and few-shot regimes. Ablations and analyses further show that both basic-level abstraction and commitment-guided selective refinement are important, supporting the view that controlling the semantic granularity of discrimination can complement conventional rebalancing strategies for long-tailed learning. Code is available at Supplement.
PaperID: 3961, Poster
Abstract: Generative multimodal reward models are increasingly used to rank responses, guide alignment, and serve as automatic judges, making their reliability a central concern. However, existing multimodal reward benchmarks remain limited in annotation rigor, domain coverage, response model diversity, and robustness-oriented evaluation, making it difficult to assess reward models under realistic and diagnostic settings. To address these gaps, we introduce MMCompass, a new benchmark for evaluating generative multimodal reward models. MMCompass contains carefully verified preference pairs spanning 10 multimodal domains, with responses collected from 16 vision-language models. The benchmark is constructed through a multi-stage pipeline including AI-assisted difficulty filtering, blind multi-annotator verification, and expert review. To enable more diagnostic evaluation, we adopt three complementary metrics: overall accuracy, symmetric accuracy, and verdict guess rate, which jointly measure judgment quality, order robustness, and position bias. Evaluations on MMCompass show that position bias is widespread in our evaluated setting, even among strong multimodal judges and that standard accuracy alone often overestimates reliability. Beyond this diagnosis, we further provide CompassRM as a mitigation baseline for improving cross-order consistency. Built with dual-order supervised fine-tuning and the proposed GXPO (Group Cross Policy Optimization), CompassRM improves robustness over its instruct-model backbones on MMCompass, VLRewardBench, and MMRewardBench, while also improving overall judgment accuracy.
Abstract: Large Language Models (LLMs) achieve strong performance on standard knowledge evaluation benchmarks, yet recent work shows that their knowledge capabilities remain brittle under question variants that test the same knowledge in different forms. Robustness augmentation of existing knowledge evaluation benchmarks is therefore necessary, but current LLM-assisted generate-then-verify pipelines are costly and difficult to scale due to low-yield variant generation and unreliable variant verification. We propose SAGE (Scalable Automated Generation of Robustness BEnchmarks), a framework for scalable robustness augmentation of knowledge evaluation benchmarks using fine-tuned smaller models. SAGE consists of VariantQual, a rubric-based verifier trained on human-labeled seed data, and VariantGen, a variant generator initialized with supervised fine-tuning and further optimized with reinforcement learning using VariantQual as the reward model. Experiments on HellaSwag show that SAGE constructs a large-scale robustness-augmented benchmark with quality comparable to the human-annotated HellaSwag-Pro at substantially lower cost, while the fine-tuned models further generalize to MMLU without benchmark-specific fine-tuning.
PaperID: 3963, Poster
Authors:
Imad-Eddine Marouf, khalid OUBLAL, Enzo Tartaglione, Stéphane LATHUILIÈREAbstract: Prompt learning adapts vision-language models such as CLIP by optimizing a small set of continuous context vectors. The cross-entropy objective drives the learned prompt toward base-class specialization, while generalization to unseen classes benefits from staying close to the zero-shot feature space. Because both objectives act on the same parameters, their gradients oppose each other at convergence, and single-prompt losses are confined to a fixed empirical base-novel trade-off curve across a wide range of loss designs. We propose Decoupled Mode Connectivity (DMC), which resolves this conflict by assigning each objective to a dedicated prompt. A linear mode connectivity (LMC) corridor in text-feature space enforces low classification loss across interpolated classifiers between the two endpoints. The Visual Anchor regularizer preserves CLIP's pretrained class-similarity structure during specialization; equivalently, it minimizes the KL divergence between the learned and pretrained class-similarity distributions. We introduce class-permutation invariance (CPI) as a necessary condition for regularizer transfer across the base-novel boundary, and prove via Fano's inequality that any CPI violation lower-bounds the drop in novel accuracy by a term proportional to the mutual information between the prompt and base-class labels. DMC improves base and novel accuracy on the majority of dataset-baseline combinations across 11 datasets and three prompt-learning baselines~(CoOp, KgCoOp, MMA), shifting the Pareto frontier rather than trading along it.
PaperID: 3964, Poster
Abstract: Although Vision-Language Models (VLMs) demonstrate impressive capabilities, their deployment in high-stakes domains is severely hindered by hallucinations. Existing detection methods mostly focus on static internal feature probing with coarse-grained annotations, failing to achieve fine-grained token-level localization without expensive dense labels. In this paper, we reconceptualize VLM hallucination detection as a dynamic graph learning task, identifying hallucinations as distinct (Token Representation Analysis with Cross-layer Evolution), a non-intrusive probing framework that models internal states as heterogeneous spatio-temporal graphs. By employing Graph Neural Networks (GNNs) to capture computational dynamics, TRACE effectively isolates hallucinatory signals. Furthermore, a token-first Multiple Instance Learning (MIL) readout elegantly bridges the granularity gap, distilling macroscopic sentence-level supervision into precise zero-shot token-level localization. Extensive evaluations on M-HalDetect, ViGoR, and HalLoc demonstrate that TRACE achieves state-of-the-art detection performance while being competitive in localization accuracy, offering an efficient and interpretable solution for trustworthy VLMs.
Authors:
Qiguang Chen, Qiming Yu, Yuhang Gu, Zhuoye Huang, Hanjing Li, Hongyu Liu, Simin Liu, Jinhao Liu, Dengyun Peng, Jiangyi Wang, Zheng Yan, Fanqing Meng, Ethan Qin, Carl Che, Mengkang HuAbstract: As agent systems scale, skills accumulate into large reusable libraries, yet their scaling laws remain poorly understood. Across 15 frontier LLMs, 1,141 real-world skills, and over 3M routing or execution decisions, we identify two coupled laws. Routing law: single-step routing accuracy decays logarithmically with library size (R^2>0.97 for all models), with errors progressing from local skill competition to cross-family drift and capture by overly general ``black-hole skills''. Execution law: before state realization, joint routing is approximately multiplicative, whereas correct execution can improve difficult downstream decisions by about 4×. A single parameter, the routing logarithmic decay slope b, couples the two laws: routing-side fits predict execution-side rescue across models, showing that the same library property controls both pre-execution collapse and downstream recoverability. The laws are actionable: law-guided optimization raises held-out routing accuracy from 71.3% to 91.7%, reduces hijack from 22.4% to 4.1%, and transfers directionally to downstream ClawBench and ClawMark execution settings, improving mean pass rate from 49.3% to 61.6% on ClawBench and from 28.4% to 34.5% on ClawMark. These results show that agent performance depends not only on model capability, but also on the structure, granularity, and exposure policy of the skill library.
PaperID: 3966, Poster
Abstract: We present a general method for discovering symbolic local learning rules that can replace backpropagation, applicable, in principle, to any deep learning architecture and task. The method automatically discovers local loss functions specific to each layer by (1) constructing a dictionary of monomials of local tensor variables, inspired by Penrose diagram notation, (2) regressing over coefficients weighting these monomials to build layerwise polynomial losses, whose gradients together minimize a target global loss. We test the method on reconstruction, classification, and self-supervised learning tasks using small multilayer networks on minimal datasets. Across all tasks, the discovered rules induce non-trivial interactions between layers that lead to significant reductions in the corresponding global losses, e.g. reaching classification accuracies comparable to backpropagation. Furthermore, the method selects sparse combinations of trace monomials, finding local rules that are more interpretable than those generated by MLP-based alternatives. The potential implications of our method range from energy-efficient deep learning to the discovery of biologically plausible learning mechanisms.
PaperID: 3967, Poster
Abstract: Geometry problem solving is a canonical testbed for machine intelligence, requiring systems to interpret diagrams, ground symbolic constraints, and perform rigorous deduction. Yet current approaches expose a persistent gap between seeing and proving: multimodal large language models can flexibly inspect diagrams but may hallucinate unsupported relations, while symbolic solvers provide checkable derivations but are brittle to incomplete or misgrounded formalization. We introduce SynGeo, a state-centric framework for synergizing seeing and proving in geometry problem solving by making the geometric representation revisable during inference. SynGeo first constructs a predicate state from diagram--text evidence and tests it with symbolic reasoning. When proof search fails or stagnates, symbolic diagnostics guide image-grounded revisiting, repairing solver-incompatible predicates before another deductive attempt. When the symbolic route remains unresolved, a complementary MLLM branch provides an image-grounded reasoning path from the problem evidence. On Geometry3K and PGPS9K, SynGeo achieves state-of-the-art performance across both Choice and Completion settings. With GPT-4o as the backbone, it reaches (90.2%) accuracy in both settings on Geometry3K, and 90.1% Choice accuracy and 88.6% Completion accuracy on PGPS9K. Ablations show that both feedback-guided revisiting and complementary reasoning are necessary, supporting a broader view of geometry problem solving as an adaptive loop between what a system sees, what it formalizes, and what it proves.
Authors: Xin Wang, Yabo Wang, Rebing Wu
Abstract: It remains unclear whether quantum machine learning (QML) truly holds an advantage when tackling practical and meaningful tasks. In exploring this question, the importance of classical data encoding, often overlooked, becomes increasingly evident. Amplitude encoding, which can embed 2^n classical data into n qubits, is widely used due to its apparent efficiency. However, its potential limitations for QML have yet to be fully explored. In this paper, we establish a theoretical result for the limitations caused by quantum encoding and point out the existence of a concentration phenomenon in amplitude encoding. Compared to prior work, our theoretical result operates under more general conditions and leads to a stronger conclusion. This concentration phenomenon causes the classification predictions to remain close to random guessing, regardless of the training process. Our findings shed new light on a long-standing puzzle in the field of QML: why some QML models perform well on simple datasets like MNIST but fail to generalize to more complex practical tasks. By highlighting the pivotal role of encoding design in QML, our work clearly indicates that future research needs to focus more on the design of classical data encoding to advance the effectiveness of QML.
Authors:
Jie Jiang, xing sun, ruotian Chen, Jianan Su, Kaixin ShenAbstract: Speculative decoding accelerates LLM inference by having a lightweight draft model propose speculative windows of candidate tokens for parallel verification by a larger target model. In practice, speculative efficiency is often bottlenecked by hard-to-draft positions, where an early mismatch truncates the accepted prefix and invalidates the rest of the speculative window. Most learning-based drafters are still optimized with token-level supervised objectives, even though speculative utility is inherently window-level and prefix-sensitive. We propose PPOW (Performance-Driven Policy Optimization with Adaptive Windowing), a reinforcement learning framework that shifts drafter optimization from token-level imitation to window-level optimization. PPOW combines a Cost-Aware Speedup Reward, a Distribution-Based Proximity Reward, and Adaptive Divergence-Aware Windowing, which prioritizes informative windows with high confidence-weighted draft--target divergence. PPOW achieves average acceptance lengths of 6.29–6.52 and speedups of 3.39–4.36× across multiple model families and benchmarks under a unified decoding protocol. These results show that performance-driven window-level optimization is a practical approach to improving speculative decoding efficiency.
PaperID: 3970, Poster
Abstract: Shortlisting is a common and effective method for pre-selecting participants in competitive settings. To ensure fairness, a cut-off score is typically announced, allowing only contestants who exceed it to enter the contest, while others are eliminated. In this paper, we study rank-order contests with shortlisting and cut-off score disclosure. We fully characterize the equilibrium behavior of shortlisted contestants for any given prize structure and shortlist size. We examine two objective functions: the highest individual performance and total performance. Under linear performance costs, for both objectives, the optimal contest is in a winner-take-all format. For the highest individual performance, the optimal shortlist size is exactly two contestants, but, in contrast, for total performance, the shortlist size does not affect the outcome, i.e., any size yields the same total performance. Furthermore, we compare the highest individual performance achieved with and without shortlisting, and show that the former is 4/3 of the latter.
PaperID: 3971, Poster
Abstract: Generative modeling of longitudinal Electronic Health Records is increasingly important for privacy-preserving research, yet standard autoregressive models tend to underrepresent the co-occurrence structure of tail events (i.e., diseases, symptoms), reducing the fidelity and faithfulness of generated data for rare subpopulations. To this end, we propose ADAPCLA framework, which enables generative models to adaptively fit and generate EHR data through a data distribution-aware training strategy; this is achieved by internalizing data knowledge parameters by simulated annealing training. It also supports training-free adaptation to a diverse clinical population for generation through zero-shot distribution control. Moreover, our theoretical analysis characterizes rare-code logit updates through the label-wise empirical NTK and derives a prior-internalization bound for how annealing speed and NTK conditioning affect retained prior signals. Experiments on real-world data show that ADAPCLA achieves consistent gains in tail plausibility, downstream utility, and zero-shot control; in particular, it improves TailPairSeen over HALO by 114.2% on MIMIC-III and 65.1% on MIMIC-IV, outperforms GPT-style generation by 3.5% F1 for zero-shot cross-population adaptation.
PaperID: 3972, Poster
Abstract: The choice of language model pretraining objective was once contested among causal language modeling (CLM), masked LM, T5 span corruption, and UL2-style mixtures of denoisers. It was settled in favor of CLM during an era when high-quality text was abundant relative to compute. That era has ended. Compute now routinely outpaces the supply of unique data, and modern pipelines rely heavily on data repetition both during pretraining on curated corpora and during midtraining on small domain corpora. We revisit the objective question in this data-constrained regime. Across a large sweep over models from 15M to 1B parameters, unique data budgets up to 6B tokens, and across repetition budgets, we compare CLM against three denoising objectives: fill-in-the-middle (FIM), T5-style span corruption, and a UL2-style mixture-of-denoisers. All four scale similarly when every training token is unique, but they differ significantly in their tolerance for repeated data. Denoising objectives, which naturally augment language data on every pass, accumulate substantially less overfitting cost than CLM. We find that span corruption and mixture-of-denoisers are the most robust. We then validate the practical payoff in a midtraining setting. We continue web-text pretrained checkpoints on a limited mixture of web-text and math data. We observe denoising midtraining objectives outperform CLM on GSM8k, and the gap widens with increased repetition. Choosing a denoising objective is a simple, complementary lever to model and data scaling for practitioners working under tight data budgets.
Authors: Oguzhan Baser, Elahe Sadeghi, Eric Wang, Nico Vergauwen, Sam Kazemian, Hong Kang, Sandeep Chinchali, Sriram Vishwanath
Abstract: Most large language models (LLMs) run on external clouds: users send a prompt, pay for inference, and must trust that the remote GPU executes the LLM without any adversarial tampering. We critically ask how to achieve verifiable LLM inference, where a prover (the service) must convince a verifier (the client) that an inference was run correctly without rerunning the LLM. Existing cryptographic works are too slow at the LLM scale, while non-cryptographic ones require a strong verifier GPU. We propose TensorCommitments (TCs), a tensor-native proof-of-inference scheme. TC binds the LLM inference to a commitment, an irreversible tag that breaks under tampering, organized in our multivariate Terkle Trees. For LLaMA2, TC adds only 0.97% prover and 0.12% verifier time over inference while improving robustness to tailored attacks by up to 48% over the best prior work requiring a verifier GPU.
Authors:
Krish Sharma, Omar Naim, Soumadeep Saha, Vinija Jain, Aman Chadha, Nicholas AsherAbstract: Recent work has promoted task-aware layer pruning as a way to improve model performance on particular tasks, as shown by TALE. In this paper, we investigate when such improvements occur and why. We show first that, across controlled polynomial regression tasks and large language models, such pruning yields no benefit on in-distribution (ID) data but consistently improves out-of-distribution (OOD) accuracy. We further show empirically that OOD inputs induce layerwise norm and pairwise-distance profiles that deviate from the corresponding ID profiles. This leads to a geometric explanation of task-aware pruning: each task induces a task-adapted geometry, characterized empirically by the representation profiles observed on ID inputs. OOD inputs can introduce a distorted version of the task-adapted geometry. Task-aware pruning identifies layers that create or amplify this distortion; by removing them, it shifts OOD representational norms and pairwise distances toward those observed on the adapted distribution. This realigns OOD inputs with the model’s task-adapted geometry and improves performance. We provide causal evidence through controlled distribution shifts and residual-scaling interventions, and demonstrate consistent behavior across model scales.
PaperID: 3975, Poster
Abstract: Large reasoning models (LRMs) improve task performance by generating longer intermediate reasoning traces, but this substantially increases inference costs and makes low-precision deployment increasingly important. However, aggressive quantization often performs poorly on reasoning tasks. Beyond error accumulation across autoregressive generation, we observe that activations exhibit significantly higher variability during the early decoding steps of thinking, making conventional calibration and uniform quantization less effective. To address these reasoning-specific quantization challenges, we propose ClusQuant, a clustering-based LUT quantization framework for low-precision large reasoning models. ClusQuant strengthens outlier mitigation through rotation and scaling transformations, uses clustering-based representations to better fit non-uniform activation distributions, and replaces fixed codebook sizes with an adaptive clustering strategy that allocates codebook capacity according to layer-wise quantization difficulty. Guided by our observation of reasoning-time activation behavior, ClusQuant further improves calibration sample selection to better match actual decoding statistics. Combined with our customized LUT CUDA kernel, ClusQuant delivers substantial accuracy improvements on reasoning benchmarks while achieving 2.85× speedup.
PaperID: 3976, Poster
Authors:
Zhenting Qi, Ao Qu, Huangyuan Su, Chenyu Wang, Yu Yao, Han Zheng, Kushal Chattopadhyay, Guowei Xu, Zihan Wang, Weirui Ye, Vijay Janapa Reddi, Ju Li, Paul Liang, Himabindu Lakkaraju, Sham Kakade, Yilun DuAbstract: How can a population of agents self-orchestrate and self-adapt into stronger collective intelligence without centralized control? Inspired by Friedrich Hayek's economic theory of decentralized coordination in markets, we study this question through an agent economy in which agents compete via auctions for the right to act, exchange payments, and accumulate wealth from environmental rewards. These simple economic signals induce decentralized credit assignment, driving planning without global orchestration or explicit communication protocols. The population evolves through economic selection: effective agents accumulate wealth and are mutated via exploitation, while ineffective ones go bankrupt and are replaced via exploration. We show that, initialized with weak agents, the economy produces emergent multi-step reasoning strategies and outperforms stronger monolithic baselines across five agentic tasks, including mathematical reasoning, financial research, scientific research, accelerator design, and distributed-system optimization. We further provide theoretical insights into how economic dynamics shape agent behaviors, linking local incentives to long-term global performance. Our results suggest a new path to multi-agent intelligence: rather than engineering coordination, we can design decentralized incentive structures under which it automatically emerges.
PaperID: 3977, Poster
Authors: Giyeol Kim, Chanho Eom
Abstract: Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos containing moments relevant to a given text query. Despite recent progress, existing PRVR methods suffer from two key limitations: a fixed video decomposition scheme that causes semantic dilution, and weak cross-domain robustness due to task-specific training. In this paper, we propose TF-PRVR, the first training-free framework for PRVR designed to improve real-world generalization. TF-PRVR leverages frozen vision-language features to construct video-specific hierarchical representations. It derives temporal semantic signals from frame-level features and applies frequency-based multi-scale analysis to identify adaptive temporal boundaries, producing hierarchical segments with coherent event-level semantics. Built on these segments, TF-PRVR constructs a unified multi-scale graph and propagates query relevance across temporally and semantically related nodes. A moment-aware scoring strategy then aggregates temporally aligned relevance across scales, emphasizing consistently supported moments while suppressing isolated false responses. Without task-specific training, TF-PRVR preserves the general-purpose alignment capability of pre-trained vision-language models and avoids dataset-specific overfitting. Extensive experiments on standard PRVR benchmarks demonstrate competitive retrieval performance and strong robustness under cross-domain evaluation, suggesting a practical direction for real-world PRVR.
PaperID: 3978, Poster
Authors:
Hao Liu, Steven Liu, Xin Zhang, Jane Luo, Yu Kang, Jie Wu, Fangkai Yang, Yangyu Huang, Pengfei Gao, Scarlett Li, Yan LuAbstract: Reproduction test generation, producing a failing-then-passing test that captures a reported bug, is a critical step in automated software engineering. Existing agentic methods treat this as a monolithic loop, despite the task inherently comprising two subtasks of distinct nature: diagnosing the root cause and writing a fail-to-pass test. Without explicit separation, the agent faces a compound objective with underspecified intermediate goals, leading to goal drift. We propose DPIAgent, a structured agentic framework built on three principles, solate (DPI), that mitigates compound-objective ambiguity and goal drift: it Divides the task into single-objective phases of defect exploration and test generation; enforces a handoff Protocol that records the diagnosis and test plan, preventing context loss; and Isolates each phase's action space by tailoring the toolset to its task, preventing irrelevant tools from misleading execution. On SWT-Bench Verified, DPIAgent outperforms seven baselines across three backbone LLMs. With DPI alone it reaches 81.76% success rate on GPT-5, the highest reported among open-source methods, gaining up to 11.88 points over the strongest baseline on GPT-5-Mini; adding test selection further raises it to 86.17%. Our analysis shows that architectural structure and backbone capability are complementary axes rather than substitutes, demonstrating DPI's generalizability across model classes.
PaperID: 3979, Poster
Abstract: Steering vectors (SVs) are residual-stream directions that shift the behavior of LLMs when added to hidden states at inference time. While prior work has mainly used SVs as post-hoc controls, we ask whether the same directions can also guide parameter updates during post-training. We study this question across supervised learning, including SFT and distillation, and reinforcement learning with verifiable rewards (RLVR). Our framework considers two intervention regimes: a hidden-state hook during teacher-forced supervised training, and an advantage-repair mechanism for zero-variance groups in GRPO. For each regime, we analyze three design axes, which are the steering direction, its placement, and the operation used to combine it with training. In supervised settings, SV injection improves over plain SFT and distillation, with gains scaling with the off-policy gap and reaching up to a 12% improvement when distilling from thinking-mode teachers. In RL settings, applying a correctness-contrast direction as advantage repair on all-wrong groups recovers a learning signal that transfers to held-out math benchmarks. We further identify three failure boundaries. In detail, code generation can break the direction, cross-family layer transfer can break the placement, and inference-time application on an SV-trained checkpoint can break the intervention regime. Together, these results establish steering vectors as a conditional but effective training signal for LLM post-training. Our code can be found in Appendix Section A.
PaperID: 3980, Poster
Authors: Nguyen Le Minh Hoang, Van D Cuong, Bui T Duc, Huynh Thi Thanh Binh
Abstract: Large language models offer a promising technique for translating natural-language optimization problems into solver code, but their reliability remains constrained by the quality of template retrieval. While recent structured methods improve over flat prompting through hierarchical taxonomy search, their traversal rules remain fixed or LLM-driven at inference time, leaving retrieval unable to improve from past successes and failures. This motivates treating template retrieval as a policy that can be learned from modeling experience, rather than as a static inference-time rule. We propose COMPASS, a policy-amortized framework for retrieving optimization templates in LLM-based solver modeling, which represents the template library as a directed acyclic graph and formulates retrieval as a sequential decision process over template nodes. A deep Q-network learns a reusable traversal policy from template-level judgments and execution feedback, reducing reliance on fixed heuristic traversal through feedback-driven graph navigation. To support extensible retrieval, COMPASS embeds problem descriptions and template nodes in a shared semantic space, allowing newly inserted templates to be scored from their text descriptions without redefining the action space. Experiments across multiple LLM backbones and optimization-modeling benchmarks demonstrate that COMPASS consistently improves over heuristic structured retrieval, with the largest gains on complex and mixed-category problems.
PaperID: 3981, Poster
Abstract: Graph Neural Networks (GNNs) often struggle to capture long-range dependencies due to over-squashing — a phenomenon in which the repeated compression of node embeddings into finite-size messages causes representations to collapse. Oversquashing is most often diagnosed as a property of the graph topology, with effective resistance serving as a principled measure of the bottleneck. We provide a complementary view on the matter: building on cellular sheaves, we introduce sheaf effective resistance, a generalization of effective resistance that depends on the sheaf attached to the graph, and we prove that for flat vector bundles it is inversely proportional to over-squashing sensitivity in the Jacobian sense. The bottleneck thus need not lie in the graph itself: it can be relocated, and reduced, by adjusting the sheaf. We instantiate this idea in FlatNSD, a simple message-passing variant of Neural Sheaf Diffusion, and show that it implicitly learns to modulate total sheaf effective resistance, performing well on benchmarks designed to stress over-squashing without altering the original graph topology.
PaperID: 3982, Poster
Abstract: Retrieval-Augmented Generation (RAG) mitigates Large Language Model (LLM) hallucinations by grounding generation in external knowledge. Although structure-augmented RAG methods improve multi-hop reasoning, they incur high indexing costs and degrade on simple queries. We argue that this robustness gap is driven by capacity bottlenecks in fixed-dimensional inner-product scoring. As corpora grow, semantically close passages create local regions where a base scorer cannot maintain sufficient separability. Reorganizing documents into trees or graphs can expose or partially mitigate this issue, but it does not by itself provide a stronger local scoring function. We propose Semantic Ambiguity Guided Capacity Expansion (SAGE), a lightweight RAG framework that detects such regions using document-side Local Separability Deficit and expands capacity only where needed. SAGE builds a two-layer semantic index, derives atomic query views from corpus passages, and calibrates a parameter-efficient hypernetwork to generate the local Ambiguity-Conditioned Scorers for selected nodes. Extensive experiments show that SAGE improves retrieval and QA performance over dense retrievers and structure-augmented baselines, achieving a better balance between single-hop, multi-hop retrieval accuracy, and robust scalability. Anonymous Code: https://anonymous.4open.science/r/SAGE-1676.
PaperID: 3983, Poster
Authors:
Wenxiao Wu, Yachen Gao, Yukang Feng, Chengming Xu, Moran Li, jingyu li, Ming Xie, Xiaobin Hu, Xinwei Sun, Jing-Hao Xue, Nong Sang, Yanwei FuAbstract: While diffusion and flow-based generative models have achieved impressive performance in high-fidelity image and video synthesis, this capability comes at the cost of substantial inference overhead during iterative sampling. Cache-based acceleration alleviates this cost by reusing intermediate computations across adjacent denoising or transport steps, typically relying on lightweight proxy discrepancies to decide when cached outputs should be refreshed. However, such proxy-discrepancy rules are inherently unreliable: the proxy is only a point estimate of the inaccessible oracle discrepancy and can misalign across prompts, timesteps, and generation dynamics. This mismatch induces asymmetric failures: proxy underestimation may cause unsafe reuse of stale cached computations and degrade generation quality, while proxy overestimation leads to overly conservative refreshes and reduced acceleration. To address these issues, this paper introduces Conformal Cache, dubbed CCache, a training-free conformal calibration framework for reliable cache-based generation acceleration. Our key insight is to formulate cache reuse as a one-sided risk-control problem and adaptively calibrate a timestep-wise upper correction for the signed proxy-to-oracle gap based on split conformal prediction. CCache is plug-and-play and can be integrated with different proxy-discrepancy caching baselines without structural modifications or parameter updates to the generative model. Extensive experiments on image and video generation demonstrate that \methodname improves generation quality and cache reliability while preserving favorable speed-quality trade-offs.
Abstract: Watermarking is a central tool for provenance in generative models, yet its application to multivariate time series remains hindered by reliability failures under post-editing attacks. We show that existing detectors, which rely on globally coupled re-encoding, suffer from bidirectional drift of the null distribution: post-editing attacks can shift the detection z-score of non-watermarked samples in either direction, invalidating clean-calibrated thresholds. We argue that this instability is a property of the re-encoding, and that reliable detection requires each recovered unit to depend only on a bounded temporal neighborhood. Guided by this principle, we propose L-VQVAE, a generative model in which each discrete token is produced from a short contiguous window, and LVQMark, a watermarking method over this token space that combines logit-bias injection with robust re-encoding for attack-time detection. Experiments on four benchmarks spanning finance, energy, and neuroimaging show that our approach preserves generation quality while stabilizing both detection power and false-positive behavior under post-editing attacks.
PaperID: 3985, Poster
Authors: Delphine Doutsas, Bruno Figliuzzi
Abstract: Posterior collapse is a common failure mode in variational autoencoders, but existing theory has mostly focused on Gaussian latent variables. Dirichlet VAEs are comparatively less understood, despite their use in settings where latent variables represent proportions, abundances, memberships, or convex mixture weights. We derive a closed-form local stability threshold \beta_c for collapse in symmetric Dirichlet \beta-VAEs with linear Gaussian decoders. This threshold shows that collapse in those models depends not only on the data spectrum, but also on the simplex dimension and prior concentration through a trigamma curvature factor induced by the Dirichlet KL. Although \beta_c is derived from a local analysis, we find that it accurately tracks the empirical onset of collapse beyond the local regime. We show its practical relevance in two real-world linear-mixture settings: hyperspectral unmixing and convex archetypal analysis. Across experiments, \beta/\beta_c consistently organizes collapse-related behavior, providing an interpretable coordinate and a practical way to narrow the search over the \beta hyperparameter before training. Taken together, these empirical findings suggest that \beta_c can serve, beyond its local derivation, as a practical reference point for anticipating collapse and guiding \beta selection in Dirichlet \beta-VAEs.
Abstract: Key--value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measuring perplexity and accuracy without assessing the safety impact. In this study, we explore alignment preservation under KV cache quantization. Across eleven instruction-tuned models (3.8B--72B) and five benchmarks (1,894 prompts), we find that low-bit quantization can silently destroy safety alignment (Mistral-7B loses 15.2% of its refusals at only 1.03× perplexity, and no universal safe bit-width exists), with sharp model-specific phase transitions invisible to standard metrics. We identify that the root cause is geometric: safety features occupy a low-dimensional activation subspace 10^2-10^3× more vulnerable to quantization noise than the full representation space perplexity averages over. Inspired by this observation, we propose Per-Channel Reduction (PCR), a diagnostic that classifies each model into one of three mechanistic failure modes: (i) outlier-crushes-safety, where safety lives in non-outlier channels collaterally damaged by outlier-driven scale factors; (ii) outlier-as-safety, where safety overlaps outlier channels and finer granularity cannot rescue it; (iii) multi-layer dilution, where safety is distributed across many layers and per-layer fixes fail. PCR predicts the correct mitigation direction on all nine primary models and one held-out model from an independent family (20 calibration prompts). PCR generalizes across unseen prompts, models, and production quantizers (KIVI: up to 97.2% recovery), succeeding where attention-based allocation methods fail. The resulting training-free protocol (~35 GPU-minutes) recovers up to 97% of lost alignment at minimal memory overhead, addressing vulnerabilities confirmed in production vLLM serving with FP8 KV cache on NVIDIA GPUs.
Abstract: A vision-language (VL) agent wraps a frozen vision-language model (VLM) into a test-time system that may call external tools during execution before producing a response. The central problem is how the VLM and its tools interact across multiple steps. Two recipes are common, both rigid. Thinking for longer (borrowed from text-only agents) sharpens reasoning, but without in-loop visual refresh, it lets visual evidence decay into hallucination. Querying tools more often, as traditional agents do, treats outputs as ground truth, a rigid commitment that fails when calls are noisy. We propose Saliency-Aware Principle Selection (SAP), a training-free inference-time method that avoids both rigidities: it searches a broad space of high-level principles, short textual directives (e.g., ``re-examine the image whenever an intermediate conclusion is formed'') that each prescribe a different way for the VLM to weigh tool advice across steps, and selects the principle whose parallel reasoning routes best fit the actual visual evidence. Tool outputs are treated as advisory (consulted, never authoritative), and a population-based evolutionary loop drives the search. Under the same token budget as long chain-of-thought, SAP substantially reduces object hallucination while remaining competitive on reasoning-heavy benchmarks. Being training-free and composed of independent routes, SAP is plug-and-play on any VLM and parallel for modern agent deployment.
PaperID: 3988, Poster
Abstract: Training tool-using agents for planning is challenging when high-quality demonstrations exist only as final outcomes without the intermediate tool traces that produced them. A common response is to impute traces and apply imitation learning, but this approach is indirect and structurally limited to matching the reference rather than improving upon it. We propose anchored comparative feedback: rather than imitating a reference plan, we train agents to surpass it under a comparative LLM judge while satisfying hard constraints. The anchor provides a stable optimization target, and comparing two identically rendered plans neutralizes surface-level judge biases. On TripTailor, a tool-augmented travel planning benchmark, anchored GRPO achieves 80.1% final success rate compared to 43.7% for imitation learning on imputed traces, with gains on both constraint satisfaction and LLM-judged preference that persist across judge families, rubrics, and a held-out reward model. Gains generalize beyond TripTailor, with a fivefold improvement on the held-out TravelPlanner benchmark. These results suggest that anchored comparative feedback offers an effective approach for learning transferable tool-using planners from outcome-only supervision.
Abstract: Recent work on the sequence universality of State Space Models (SSMs) has introduced efficient, maximally expressive continuous-time approaches for time-series modelling. While these works focus on discriminative settings, we extend this perspective to generative time-series modelling by proving that maximally expressive Structured Linear Controlled Differential Equations (SLiCEs) are universal time-series generators, in the sense that they can approximate the induced path laws of continuous causal pushforwards on compact latent sets in W_\infty. Building on these theoretical results, we propose Generative SLiCEs (G-SLiCEs), a maximally expressive continuous-time model for flow matching on path-space. Empirically, we show that expressivity improves performance in probabilistic forecasting and downstream tasks, while retaining the advantages of continuous-time models such as generalising to arbitrary observation grids. This is particularly beneficial for irregular grids, where fixed-grid models often struggle.
Abstract: When an omnimodal large language model accepts a question whose textual premise contradicts what it actually sees or hears, does the failure lie in perception or in action? Recent omnimodal models are positioned as perception-grounded agents that jointly process video, audio, and text, yet a basic form of grounding remains untested: catching a textual claim that conflicts with the model's own sensory input. We introduce IMAVB, a curated 500-clip benchmark of long-form movies with a 2x2 design crossing target modality (vision, audio) and premise condition (standard, misleading), which lets us measure conflict detection separately from ordinary multimodal comprehension. Across eight open-source omnimodal LLMs and Gemini 3.1 Pro, we document a Representation-Action Gap: hidden states reliably encode premise-perception mismatches even when the same models almost never reject the false claim in their outputs. Behaviorally, models fall into two failure modes: under-rejection, in which they answer misleading questions as if the false premise were true; and over-rejection, in which they reject more often but also reject standard questions, sacrificing ordinary comprehension accuracy. The gap is modality-asymmetric (audio grounding underperforms vision) and prompt-resistant across seven variants. As an initial diagnostic intervention, a probe-guided logit adjustment (PGLA) re-injects the encoded mismatch signal into decoding and consistently improves rejection behavior. Together, these results suggest the bottleneck for omnimodal grounding lies in translation, not perception.
Abstract: Tabular Foundation Models (TFMs) increasingly rely on in-context learning, where a model receives labelled examples at inference time and predicts labels for new inputs without updating its weights. Existing TFMs are typically trained on either massive synthetic corpora or very large collections of real datasets. In contrast, we show that surprisingly strong transfer can emerge from self-supervised pre-training on just a single real table. In this setting, we also find that tables tend to be either broadly useful or broadly poor regardless of downstream prediction task, and that the strongest predictor of usefulness is the number of features rather than the number of instances. This leads to a task-centric interpretation of tabular pre-training: the amount and the quality of tasks are essential for the pre-training of TFMs. We show that the same task-centric perspective can help corpus design at scale: fine-grained column-level pre-processing consistently improves downstream performance, while no improvements are observed when we filter or deduplicate at the dataset level. Finally, we offer a new perspective for how TFMs generalize: we believe that tabular in-context generalization is largely retrieval-based, and good models are those that learn to identify relevant examples in the provided context and aggregate them well. The mechanics of TFMs have been relatively understudied; our task-centric, retrieval-based perspective offers a new framework to guide future model and corpus design.
PaperID: 3992, Poster
Abstract: Optimizing the leader's policy via hypergradients (HG) in bilevel reinforcement learning (RL) typically assumes a white-box follower. To enable real-world applications, we address the black-box follower setting, where the leader must optimize solely from observed trajectories without access to the follower's true reward function. We propose Reward-Estimated Hypergradients (RE-HG), which leverages Inverse RL (IRL) to recover this unobserved reward. Because naive IRL introduces severe reward-shaping biases and numerical instabilities, RE-HG introduces a theoretical bias-cancellation mechanism that aggregates information across diverse environmental dynamics. Together with eigenvalue truncation, this acts as implicit regularization to suppress HG norm explosions. Empirical evaluations demonstrate that RE-HG estimates gradients aligned with a white-box Oracle, achieving comparable or superior performance. Notably, in sharp reward landscapes where the Oracle becomes trapped in suboptimal local minima, RE-HG's implicit regularization extracts stable global gradients, demonstrating robustness.
PaperID: 3993, Poster
Authors: Ziyu Chen, Xinbei Jiang, PENG SUN, Tao Lin
Abstract: Masked Diffusion Models (MDMs) offer flexible, non-autoregressive generation, but this freedom introduces a challenge: final output quality is highly sensitive to the decoding order. We are the first to formalize this issue, attributing the variability in output quality to the cumulative predictive uncertainty along a generative path. To quantify this uncertainty, we introduce Denoising Entropy, a computable metric that serves as an internal signal for evaluating generative process. Leveraging this metric, we propose two algorithms designed to optimize the decoding path: a post-hoc selection method and a real-time guidance strategy. Experiments demonstrate that our entropy-guided methods significantly improve generation quality, substantially boosting accuracy on challenging reasoning, planning, and code benchmarks. Our work establishes Denoising Entropy as a principled tool for understanding and controlling generation, effectively turning the uncertainty in MDMs from a liability into a key advantage for discovering high-quality solutions.
PaperID: 3994, Poster
Abstract: Diffusion models promise iterative refinement, yet masked diffusion language models (MDMs) suffer from two shortcomings: (1) inability to revise tokens, and (2) sparse training signal. The root cause is structural: MDMs control information flow only through input masking, so unmasked positions cannot be supervised, and self-correction is never learned. We propose the Leave-One-Out Transformer (LOOT), which addresses both issues. With LOOT, MDMs predict the distribution over potential replacements conditioned on the rest of the sequence at each position, in a single forward pass. Supervision can then be applied at masked and unmasked positions alike, yielding a lower-variance training objective with a dense loss signal at every token. Our method empowers MDMs to learn self-correction during training and supports continuous refinement at inference time via Gibbs sampling. LOOT achieves a new best validation perplexity on LM1B in a parameter-matched setting. Finetuned from a MDM checkpoint on OpenWebText, LOOT establishes a generative quality-vs-diversity frontier that surpasses prior remasking methods across NFE budgets.
PaperID: 3995, Poster
Abstract: Although deep neural networks (DNNs) have demonstrated strong performance in medical classification, their opaque reasoning undermines the trustworthiness of clinical diagnoses due to limited interpretability. Concept bottleneck models (CBMs) introduce an intermediate concept layer, decomposing black-box prediction into an explicit reasoning path: image → concept → disease. However, concepts and their corresponding diagnostic results in medical image diagnosis are often hierarchical, inclusive, and semantically unevenly distributed. Most concept-based methods represent concepts in Euclidean space and treat them as isolated entities, which hinders the modeling of their structural relationships and thereby reduces diagnostic transparency. To address this, we introduce a hyperbolic concept embedding model (HCEM) that maps medical concepts into a hyperbolic space better suited for hierarchical representation, enabling explicit modeling of complex concept relationships and their semantic associations with diseases. In this model, we regularize the hyperbolic concept embedding space using a positive–negative concept contrastive loss and a concept entailment cone loss. Furthermore, we employ an intervention-aware disease concept hyperbolic embedding regularization to learn their dynamic relationships. It avoids rigid rule priors while improving diagnostic consistency and flexibility under concept interventions. Extensive experiments on four medical datasets validate that our HCEM provides high accuracy in both concept and disease classification, as well as superior interpretability and intervenability. The code will be released soon.
Abstract: In this paper, we ask whether vision foundation models construct representations that reflect the intrinsic properties of 3D Euclidean space. Unlike previous works that probe 3D awareness of vision features by regressing image-centric quantities such as depth or normals, we investigate the relation between the structure of the space of visual features and the group of Euclidean transformations SE(3). We propose a set of probes to evaluate this relation from both topological and geometric perspectives: a mutual neighborhood metric that measures the alignment between feature neighborhoods and spatial topology, and a Poincaré Adapter to test the linear accessibility of the geometry of camera motion from latent displacements in static scenes. We show that self-supervised vision models, which, in principle, have not been trained with direct 3D supervision or active agency, possess latent subspaces that are remarkably strongly correlated with the structure of Euclidean space when probed correctly. Building on this insight we propose a new class of ``Latent-Space Navigation'' techniques that perform visual odometry and localization purely in the visual latent space, bypassing the need for explicit 3D reconstruction.
Authors:
Xiaoda Wang, Minxiao Wang, Kaiqiao Han, Defu Cao, Ching Chang, Yidan Shi, Runze Yan, Xiao Luo, Yan Liu, Xiao Hu, Yizhou Sun, Wei Wang, Carl YangAbstract: Electrocardiography (ECG) is the clinical standard for cardiac assessment but requires dedicated hardware that does not scale to daily-life monitoring. Photoplethysmography (PPG) is ubiquitous in wearables but lacks ECG-specific diagnostic morphology and is corrupted by motion and sensor noise. PPG-to-ECG generation aims to bridge this gap by recovering electrical morphology and timing from peripheral pulse signals. However, existing methods largely rely on statistical alignment and data-driven generation. They fail to explicitly structure the latent space around physiology-aware electro-hemodynamic factors and lack constraints from forward physiological dynamics. To address these challenges, we propose PG-LRF, a physiology-guided latent rectified flow framework. PG-LRF introduces an electro-hemodynamic simulator that co-models ECG and PPG through shared cardiac phase dynamics. Guided by this simulator, a Physiology-Aware AutoEncoder learns a structured electro-hemodynamic latent space. Then we integrate this simulator guidance into a PPG-conditioned latent rectified flow, enforcing ECG-side morphology consistency and ECG-to-PPG forward hemodynamic consistency during generative transport. Experiments on the large-scale MC-MED dataset demonstrate that PG-LRF significantly improves PPG-to-ECG generation and downstream cardiovascular disease classification, proving its ability to generate ECGs that are both signal-faithful and physiologically plausible under the ECG-to-PPG hemodynamic pathway
Abstract: The landscape of self-supervised learning (SSL) is currently dominated by generative approaches (e.g. MAE) that reconstruct raw low-level data, and predictive approaches (e.g. I-JEPA) that predict high-level abstract embeddings. While generative methods are stable due to their reliable training targets based on ground-truth data, they are computationally inefficient for high-redundancy modalities like imagery, and their training objective does not prioritize learning high-level, conceptual features. Conversely, predictive methods often suffer from training instability due to their reliance on the non-stationary targets of final-layer self-distillation. We introduce Bootleg, a method that bridges this divide by tasking the model with predicting latent representations from multiple hidden layers of a teacher. This hierarchical objective forces the model to capture features at varying levels of abstraction simultaneously. We demonstrate Bootleg significantly outperforms comparable baselines (+10% vs. I-JEPA) on frozen probe classification of ImageNet-1K, iNaturalist-21, and VTAB, and semantic segmentation of ADE20K, Cityscapes, and COCO-Stuff.
Abstract: Conformal prediction is often calibrated with a single pooled threshold, but this can hide cross-group heterogeneity in score distributions and distort group-wise coverage. We study this phenomenon through the population score distributions underlying split conformal calibration. First, we derive a conservation law and lower bound showing that pooled calibration incurs irreducible group-wise coverage distortion at a scale set by cross-group quantile heterogeneity. Second, we demonstrate that the two leading fairness definitions for conformal prediction, Equalized Coverage and Equalized Set Size, are fundamentally in tension. Third, we quantify the cost of moving between policies which treat groups separately or pool them. Experiments on synthetic and real data confirm the same bidirectional trade-off after finite-sample calibration. Our results show that, for the policy families studied here, calibration choice does not remove cross-group heterogeneity; it determines whether the resulting distortion appears in the coverage or size dimension, providing a principled lens for analyzing fairness-oriented calibration choices in practice.
Abstract: Spatial reasoning remains a challenge for Multimodal Large Language Models (MLLMs), as it requires reliable multi-hop inference over both intermediate states and state transitions. Current studies often leave intermediate states unverified and treat state transitions as implicit processes, which limits reliability in multi-hop spatial reasoning. To address this, we propose State-aware Visualization-of-Thought (SVoT), a reinforcement learning framework that generates interleaved, verifiable intermediate states and visualizations. SVoT integrates transition reasoning chains into the generation processes, enabling the model to verify action preconditions and effects through interleaved textual and visual reasoning. We train SVoT via Group Relative Policy Optimization (GRPO), instantiating verification through reward design and evaluating the efficacy of different fine-grained rewards. As existing benchmarks reduce state transitions to single-variable updates, substantially simplifying the problems, we establish five domains by extending classical environments and introducing two novel domains, Pacman and Gather, that require multi-object interactions and numerical reasoning. These domains support systematic evaluation of multi-hop spatial reasoning with quantitative verification of generated intermediate states and transition reasoning. SVoT with transition-aware supervision achieves state-of-the-art performance across the introduced domains, yielding up to a 65% absolute accuracy gain on out-of-distribution test sets.
Abstract: Referring segmentation grounds natural-language queries to pixel-level masks, but extending it to complex scenarios with multiple instances, cross-category groups, or open-ended target sets remains challenging. Previous Large Vision Language Model (LVLM)-based methods represent referred targets with one or more special tokens sequentially, treating multiple targets as separate outputs rather than a coherent set and offering little incentive to capture set-level properties such as completeness and mutual exclusivity. We reformulate open-ended referring segmentation as explicit set-level concept prediction and propose Set-Concept Segmentation (SetCon), which uses LVLM-generated natural-language concepts, instead of segmentation-specific tokens, as semantic conditions for joint mask-set decoding. A hierarchical semantic decomposition first predicts a shared set-level concept defining the target scope and then refines it into fine-grained concept groups aligned with target subsets. To support this, a two-stage annotation pipeline augments existing reasoning segmentation datasets with hierarchical semantic supervision (236k samples, 784k concept phrases). SetCon achieves state-of-the-art results on image benchmarks (+3.3 gIoU on gRefCOCO, +12.1 gIoU on MUSE), with margins that grow as the number of referred targets increases. The concept interface also transfers to video under a detect-and-track setting, yielding new state-of-the-art results on seven referring video benchmarks, including +10.9 J&F on MeViS and +12.4 J&F on Ref-SeCVOS. The code, model checkpoints, and dataset annotations will be released.
Abstract: Fact-checking systems with search-enabled large language models (LLMs) have shown strong potential for verifying claims by dynamically retrieving external evidence. However, the robustness of such systems against adversarial attack remains insufficiently understood. In this work, we study adversarial claim attacks against search-enabled LLM-based fact-checking systems under a realistic input-only threat model. We propose DECEIVE-AFC, an agent-based adversarial attack framework that integrates novel claim-level attack strategies and adversarial claim validity evaluation principles. DECEIVE-AFC systematically explores adversarial attack trajectories that disrupt search behavior, evidence retrieval, and LLM-based reasoning without relying on access to evidence sources or model internals. Extensive evaluations on benchmark datasets and real-world systems demonstrate that our attacks substantially degrade verification performance, reducing accuracy from 78.7% to 53.7%, and significantly outperform existing claim-based attack baselines with strong cross-system transferability.
Authors:
Jiahui Li, Ruili Fang, Zishuai Liu, WenZhan Song, Jin Lu, Fei DouAbstract: Beat-level Electrocardiography (ECG) arrhythmia detection aims to assign an arrhythmia class to each beat in a recording, yet many existing systems treat beats as isolated local instances. This is limiting because beat labels often depend on multi-beat rhythm context, including timing, compensatory pauses, and beat-to-beat morphological consistency. We present DeepArrhythmia, a tool-grounded multimodal framework for segment-contextualized beat-level ECG arrhythmia classification. Given a multi-beat ECG segment, DeepArrhythmia combines the raw ECG signal and a rendered waveform image, localizes R peaks to identify beat instances, and produces structured beat-level predictions. The framework decouples physiological measurement from evidence integration using specialized tools for beat localization, numerical rhythm--morphology extraction, and morphology-focused textual analysis. DeepArrhythmia uses segment-level confidence to route between minimal and rich evidence states, since richer physiological evidence is not uniformly useful. This agentic design integrates rhythm context, explicit physiological grounding, and selective evidence acquisition for decision making.
Abstract: Vision-Language-Action (VLA) models have emerged as a powerful paradigm for generalist robotic control. However, their high computational cost and limited control frequency hinder real-time robotic manipulation, especially when large vision-language backbones and iterative action heads are executed at every control step. Existing VLA acceleration methods often optimize individual components or rely on fixed acceleration rules, treating different control steps with largely fixed computation and overlooking the non-uniform reasoning demands of sequential embodied control. Inspired by human motor control, where cognitive and feedback resources concentrate on goal-sensitive stages, we argue that VLA models should learn when to invest full computation and when to reuse prior computation. To this end, we propose ElegantVLA, a plug-in phase-adaptive inference framework that accelerates VLA models through intra-model dynamic compute scheduling. ElegantVLA introduces a lightweight scheduler that observes temporal representation similarity, robot-motion cues, and episode progress to jointly allocate computation across the vision encoder, LLM, and action head. For perception-language reasoning, the scheduler selects a five-level Vision--LLM compute mode, from full recomputation to multi-step temporal reuse, based on visual-language representation stability. For action generation, it selects a three-level denoising mode, reusing intermediate denoising states during stable motion while preserving full refinement for goal-sensitive stages. By coordinating these decisions, ElegantVLA provides a general acceleration framework for modern VLA pipelines with explicit action-generation modules, without modifying or retraining the base model. Extensive experiments on GR00T, CogACT, and real-world tasks show that ElegantVLA preserves or improves task success while substantially accelerating inference. On GR00T, it achieves up to 2.55× average speedup. On CogACT, it delivers a 3.77× average speedup. In GR00T-based real-world experiments across six tasks, it reduces computation by 2.18× and increases control frequency from 13.8 Hz to 26.3 Hz.
Abstract: Modality following is the ability to selectively leverage multimodal contexts based on user instructions. It is fundamental to the safety and reliability of multimodal large language models (MLLMs) in real-world deployments. However, the internal mechanisms governing this decision-making process remain largely under-explored. In this work, we investigate the mechanism underlying modality following through an information flow perspective. Our findings reveal that instruction tokens serve as structural anchor for modality arbitration: Shallow attention layers perform undifferentiated information transfer, routing multimodal cues to instruction tokens as a latent buffer; in contrast, deep attention layers selectively strengthen the instruction-compliant subspace and resolve modality arbitration according to the instruction-specified intent, with a sparse subset of attention heads driving this process. Targeted attention-head interventions further validate the functional specificity of these heads: blocking only 5% of the identified heads substantially degrades modality following while preserving general visual and language capabilities, whereas targeted amplification can restore failed modality-following samples by up to approximately 60%. Together, this work provides a mechanistic account of modality following and informs future efforts to improve how MLLMs integrate and utilize multimodal evidence under user instructions. \footnoteAll data and code will be released upon acceptance.
PaperID: 4006, Poster
Abstract: Crowd counting is a vital computer vision task with wide-ranging applications, from public surveillance to urban planning. A persistent challenge is achieving scale invariance despite noisy labels, as humans often make counting errors with similar objects. Most existing methods struggle to maintain consistency and robustness when testing images differ in scale from the training data. We propose ScaLe-aware Invariance and Correcting Errors (SLICE) for robust crowd counting, a novel scale-aware loss function that improves scale invariance, enabling consistent and robust predictions under varying scales and annotation noise. The loss function combines two components: a scale-invariance term that applies multiscale supervision to enforce consistency across density maps of the same scene, and an error correction term that addresses inevitable annotation noise by modeling label uncertainty with probabilistic methods such as a mixture of Gaussians. SLICE is model-agnostic and integrates easily into existing crowd counting architectures without extra computational overhead during inference. Extensive experiments on five benchmark datasets show that SLICE greatly improves scale invariance and cross-dataset generalization, achieving the state-of-the-art results in challenging scenarios. Using CrowdDiff, Steerer, and CrowdFormer as baselines, SLICE lowers mean absolute error (MAE) by 5–7% by addressing annotation noise and scale variation. Applying the scale-invariance term across multiple scales further cuts MAE by up to 15%.
PaperID: 4007, Poster
Abstract: Domain randomization (DR) improves sim-to-real transfer by training policies over randomized simulator dynamics, but standard DR optimizes average performance under a fixed sampler and can miss rare failure-prone domains. We propose Risk-Sensitive Domain Randomization (RSDR), a distributionally robust formulation of episodic DR that evaluates each policy against the worst-case distribution over dynamics parameters within a KL neighborhood of a reference sampler. The resulting soft adversary has an exponential-tilting form, yielding a single temperature-controlled sampler family: negative temperatures emphasize low-return domains for robustness, zero recovers uniform DR, and positive temperatures give an optimistic sampler closely related to curriculum-style DR. Because this target changes with the policy and returns are observed only through rollouts, RSDR learns an amortized dynamics sampler by reverse-KL variational inference while reusing on-policy PPO trajectories. Across six domain-randomized MuJoCo Playground tasks, robust RSDR improves CVaR and minimum-return metrics over uniform and adaptive DR baselines, and outperforms best-tuned EPOpt.
PaperID: 4008, Poster
Authors: Alexander K Clarke, Dimitrios Chalatsis, Agnese Grison, Irene Mendez Guerra, Noura Ezaz-Nikpay, Pranav Mamidanna, Shihan Ma, Silvia Muceli, Dario Farina
Abstract: Decomposing surface electromyography (EMG) into the spike trains of individual motor neurons is a long-standing inverse problem and a key step toward motor-neuron-driven neural interfaces such as prosthetics and exoskeletons. The standard approach, independent component analysis (ICA) of the multichannel signal, assumes that the mixing from neurons to electrodes is stationary in time. This assumption fails during movement, when volume-conductor deformation makes the mixing time-varying, and current decomposition algorithms are correspondingly restricted to isometric contractions. We introduce a quasi-linear ICA formulation in which a static linear separator is preceded by a learned, low-rank, time-varying invertible transformation. The separator is trained with an independence loss on the uncompensated projection, and the transformation with a stationarity loss on the recovered source. Gradients are not shared between the two, so the source-extraction step reduces to classical linear ICA and inherits its identifiability guarantee, while non-stationary distortion is absorbed by the transformation. The closed-form inverse of the transformation enables per-spike subtraction with a time-varying template during sequential peel-off. On a public benchmark of dynamic high-density EMG with ground-truth spike trains, the method outperforms five adaptive ICA baselines at every recall threshold, recovering more units at a higher accuracy.
PaperID: 4009, Poster
Authors:
Yue Cao, Mohsen Zardadi, Yu Hu, Yanshuo Fan, Jozsef Hamari, Zheng Liu, Jianyang GuAbstract: Object detection data contains a structural asymmetry that is easy to overlook in dataset distillation (DD). Each detection image couples a spatial layout with a visual realization. Layouts define detection targets through object categories and bounding boxes, while pixels instantiate these targets with particular object appearances and backgrounds. Although both forms can be redundant across a dataset, they play different roles. Layouts define the task distribution, so synthesizing them freely risks altering the detection problem. Visual realizations, on the other hand, are conditional samples attached to layouts. This motivates a separation principle: preserve and recombine layouts from the original data, then use generative priors to diversify their visual realizations. Based on this principle, we propose LCD^3, a layout-conditioned diffusion framework for object detection DD. LCD^3 constructs composite layouts from diverse groups based on scene-level embeddings. A layout-conditioned diffusion model then generates new images from these layouts, enriching object appearance and scene realization without discarding the spatial structure needed for detection. The generation process is grounded with semantically verified object crops to retain the original visual style. Experiments show that this separation between layout and realization produces more effective distilled detection datasets. LCD^3 consistently outperforms the previous state-of-the-art, OD^3, across benchmarks using fewer objects per distilled image, e.g., a 12.5% mAP50 improvement at a 0.5% compression ratio on PASCAL VOC.
Abstract: Shot Boundary Detection (SBD) aims to automatically identify shot changes and divide a video into coherent shots. While SBD was widely studied in the literature, existing methods often produce non-interpretable boundaries on transitions, miss subtle yet harmful discontinuities, and rely on noisy, low-diversity annotations and outdated benchmarks. To alleviate these limitations, we propose OmniShotCut to formulate SBD as structured relational prediction, jointly estimating shot ranges with intra-shot relations and inter-shot relations, by a shot query-based dense video Transformer. To avoid imprecise manual labeling, we adopt a fully synthetic transition synthesis pipeline that automatically reproduces major transition families with precise boundaries and parameterized variants. We also introduce OmniShotCutBench, a modern wide-domain benchmark enabling holistic and diagnostic evaluation. Experiments on the benchmarks demonstrate the effectiveness and generality of our method.
PaperID: 4011, Poster
Authors:
Xinrui Chen, Zhen Huang, Shuwei Li, Fanyi Zeng, Dongxu Yue, Hao Yu, Huaisong Zhang, Xitong Ling, Yongxian Wei, Feng Lu, Chun YuanAbstract: Large vision--language models (VLMs) incur substantial computational costs due to massive parameter counts and the deep propagation of long visual sequences. While recent studies highlight visual token redundancy, existing methods often rely on local heuristics or coarse skipping policies that treat Transformer layers as isolated units, thereby limiting the optimization landscape. We propose SearchV, a framework that redefines VLM acceleration as a discrete policy search problem guided by a holistic fidelity objective. Central to our approach is Global Contribution (GC) fitness, a metric that captures end-to-end performance drift induced by complete skipping configurations rather than isolated components. By leveraging a fine-grained search space and evolutionary mutation that decouple attention and feed-forward pathways, SearchV identifies surgical interventions that preserve multi-benchmark robustness more effectively. Experimental results demonstrate that SearchV defines a new efficiency frontier for VLMs. Notably, SearchV identifies optimized 13B-based policies that surpass the vanilla 7B model in accuracy while operating at a lower computational budget. This represents a training-free milestone that effectively breaks the performance ceiling of conventional model scaling, proving that global structural optimization can recover high-capacity knowledge under heavy compression. Code is available in supplements.
PaperID: 4012, Poster
Abstract: Large vision-language models (LVLMs) can caption images, answer visual questions, and reason over scientific diagrams, yet they still produce hallucinations unsupported by visual evidence. Existing approaches mitigate hallucinations either by reshaping the decoding distribution or by steering hidden activations with calibration-based directions. The former adjusts token probabilities but operates only at the output level, whereas hallucinations often arise when language priors dominate visual evidence in intermediate representations. The latter applies a fixed calibration intervention, which may fail to capture image-specific hallucination drift or inadvertently suppress grounded evidence. To address these limitations, we propose PAVE, a training-free method for input-specific hidden-state intervention. PAVE constructs offline subspaces for hallucination-drift suppression and visual-evidence amplification, and adapts the hidden-state update to each input using a per-image visual basis computed during prefill. By editing hidden states with this prefill-conditioned subspace, PAVE strengthens grounded evidence while suppressing hallucination drift, without relying solely on decoding-level corrections or fixed activation edits. Experiments on LLaVA-1.5-7B and Qwen-VL show that PAVE substantially reduces hallucinations on CHAIR and POPE while preserving grounded utility on MME. Code will be released.
PaperID: 4013, Poster
Authors: Romain Lacombe
Abstract: Glycosylation is among the most diverse and important post-translational modifications in biology, governing immunogenicity, self-recognition, and the clinical viability of biologic drugs. Glycans cannot be sequenced from a template, and de novo structure prediction from tandem mass spectrometry (MS/MS) remains a longstanding bottleneck. The current state of the art, GlycoBART, is a 207M autoregressive transformer whose O(\mathrmbeam \cdot \mathrmoutput\_length) decoder passes preclude real-time annotation. We introduce GlycoGen, a 50M Discrete Flow Matching (DFM) model that predicts glycan structure from MS/MS spectra by direct, non-autoregressive generation in a few parallel forward passes. We pair it at inference time with Crystallizing Flows, a deterministic sampler that propagates the per-token distribution on the probability simplex and freezes each token to a one-hot (crystallizes it) the moment its argmax leaves the mask, exiting early once all tokens are unmasked. The sampler feeds the denoiser mixture embeddings, and runs out-of-the-box with no retraining. GlycoGen + Crystallizing Flows reach \mathbf0.9266 top-1 structural accuracy on the CandyCrunch test set, three orders of magnitude faster than GlycoBART, Pareto-dominating the prior state of the art on both accuracy and speed. We extend our experiments to a matched-architecture MDLM denoiser (Sahoo et al. 2024) and observe a comparable lift: the sampler generalizes from continuous-time flow matching to discrete-time masked diffusion. To our knowledge, this is a new state of the art for de novo open-vocabulary glycan structure prediction, and the first system fast enough for accurate real-time inline annotation on modern LC-MS/MS instruments.
PaperID: 4014, Poster
Abstract: Defenses that provide security guarantees against prompt injection attacks require strict isolation between an agent's task planning and data processing capabilities. This prevents third-party content from overwriting trusted instructions. In text-based environments such as tool-use APIs, agents can plan from interface definitions without ever processing untrusted data. Web agents, however, face a fundamental challenge: they must observe the rendered page to perceive their environment, but that page already contains untrusted third-party content. In this paper, we present Untrusted Content Masking, a simple, effective approach that enables web agents to observe their environment and plan without directly processing untrusted content. We leverage a key structural insight: a webpage's Document Object Model (DOM) structure alone suffices to identify untrusted regions. Our framework exploits this by redacting such regions before they reach the agent, and restricting interaction to a sandboxed interface with strict privilege separation.
Authors: Niklas Houba
Abstract: Compositional inference - the decomposition of observations into an unknown number of latent components - is central to perception and scientific data analysis. Attention-based models perform well when components are approximately separable, as in object-centric vision. Under additive superposition, however - where multiple components contribute to every observation - we identify a structural failure mode we term slot collapse: multiple slots converge to the same dominant component while weaker ones remain unrepresented. We trace this to a general limitation: attention is memoryless with respect to explained evidence. All slots repeatedly operate on the same input without accounting for what has already been explained, so gradients are dominated by the strongest component, inducing shared fixed points across slots. As a result, attention fails to enforce non-redundant allocation under additive superposition. We address this by introducing residual evidence modeling, which tracks the remaining explanatory capacity of the input, instantiated via evidence depletion - a minimal modification combining multiplicative depletion with an attention bias. Controlled ablations show that parallel attention, sequential processing alone, and loss-based regularization all fail to resolve collapse; evidence depletion, which adds stateful residual tracking to sequential attention, consistently succeeds. Across synthetic benchmarks and real-world audio mixtures (FUSS), evidence depletion reduces slot collapse by up to an order of magnitude, generalizing beyond synthetic settings. On gravitational-wave source inference for the ESA/NASA LISA mission, under identical architectures, data, and losses, standard attention fails while evidence depletion prevents collapse and enables multi-source posterior estimation. These results show that under additive superposition, residual evidence tracking is the operative ingredient - preventing collapse and enabling compositional inference.
PaperID: 4016, Poster
Abstract: Predicting the maximum stable learning rate of a neural network from its architecture alone has remained outside the reach of current theory, except for deep linear and shallow scalar models, where exact sharpness expressions are available. We close this gap for deep polynomial networks by proving the first computable, architecture-explicit lower bound on the largest Hessian eigenvalue at any interpolating minimum. The bound factorizes into three independently computable terms: a data-geometry factor, an architectural-capacity factor, and a label-energy factor. Combining our bound with the dynamical framework of Chemnitz and Engel [2025] and the stability conditions of Cohen et al. [2022], we derive critical step-size upper bounds for a range of optimizers from vanilla gradient descent to adaptive versions like Adam. The maximum stable learning rate for deep polynomial networks scales as (\textactivation degree)^(-2 × \textdepth), placing realistic polynomial network training in the Edge of Stability regime. Our derivation shows that this scaling is a Hessian-level consequence of the network's algebraic structure, not a property of any specific optimizer or dataset. With identity activation, our bound recovers the exact deep linear sharpness of Mulayoff and Michaeli [2020].
Abstract: Training and evaluation in multi-channel imaging (MCI) remains challenging due to heterogeneous channel configurations arising from varying staining protocols, sensor types, and acquisition settings. This heterogeneity limits the applicability of fixed-channel encoders commonly used in general computer vision. Recent Multi-Channel Vision Transformers (MC-ViTs) address this by enabling flexible channel inputs, typically by jointly encoding patch tokens from all channels within a unified attention space. However, unrestricted token interactions across channels can lead to feature dilution, reducing the ability to preserve channel-specific semantics that are critical in MCI data. To address this, we propose Decoupled Vision Transformer (DC-ViT), which explicitly regulates information sharing using Decoupled Self-Attention (DSA), which decomposes token updates into two complementary pathways: spatial updates that model intra-channel structure, and channel-wise updates that adaptively integrate cross-channel information. This decoupling mitigates informational collapse while allowing selective inter-channel interaction. To further exploit these enhanced channel-specific representations, we introduce Decoupled Aggregation (DAG), which allows the model to learn task-specific channel importances. Extensive experiments across three MCI benchmarks demonstrate consistent improvements over existing MC-ViT approaches.
Abstract: Mechanistic interpretability aims to explain a model’s behavior by identifying causally responsible internal structures. Dictionary-based explainers such as sparse autoencoders and transcoders are a primary tool, but their faithfulness under out-of-distribution (OOD) shift has received little systematic attention. We show that distribution shift rotates the subspace that the model actively uses, misaligning the explainer’s dictionary trained on in-distribution (ID) activations. We formalize this misalignment as the faithfulness gap, a geometric distance between the ID dictionary and the OOD-active subspace, and show that it controls OOD faithfulness degradation. To reduce this gap, we propose the Geometry-Adaptive Explainer (GAE), which realigns the explainer's dictionary with the OOD-active subspace while preserving the original feature structure. This requires only unlabeled OOD activations and no gradient updates. We prove that GAE improves over the unadapted ID explainer, with excess loss bounded quadratically by the second-moment shift. Empirically, GAE even matches or surpasses all training-based baselines in causal faithfulness across multiple models and OOD settings.
PaperID: 4019, Poster
Abstract: Designing regulatory DNA sequences for targeted gene-expression control is an important challenge. Recent approaches frame this task as property-conditioned generation, where diffusion models trained on natural sequences are guided toward high predicted activity through reward-based fine-tuning or sampling-time guidance. We show that single-reward evaluation can obscure reward-hacking behavior. By decomposing evaluation into property, novelty, and plausibility, we identify an empirically underexplored region of the evaluation space in which all three objectives are satisfied, but no existing baseline succeeds. We postulate that this gap arises because prior methods treat guidance primarily as a matter of magnitude. In regulatory DNA, however, sequence grammar is spatially nonuniform: motif positions are sensitive to perturbation, while background regions allow more exploration. We therefore propose \textscLocoGen by reformulate guidance as a problem of locality. \textscLocoGen learns an invariant motif mask through sequence-level invariance learning to protect motif positions, while applying OOD damping only to background positions. It operates at sampling time on a frozen diffusion backbone, with no retraining. Across HepG2, K562, and SK-N-SH, \textscLocoGen reaches this previously underexplored region of the property--novelty--plausibility space. Code is available at \urlhttps://huggingface.co/AnonymousAuthor42/LocoGen.
Abstract: Large language models increasingly rely on sampling as a driver of their own improvement, making the fidelity of their learned distributions more critical than ever. Yet, not all distributions are equally easy to learn. In this work, we identify a curse of ambiguity: in large language models, and more broadly in all neural networks that produce discrete probability distributions, the more ambiguous a next-token distribution is, the harder it is to learn accurately. Through an extensive theoretical analysis, we trace this curse to architectural and learning roots. More ambiguous distributions require more capacity to be stored, larger embeddings to be represented, more steps to be fitted, and amplify token-sampling noise. We validate these findings on synthetic tasks with controlled ground truth and observe the same signatures in language models trained on real data. Our results provide a new perspective on the statistical capabilities of large language models and a practical framework for when to trust their output distribution.
PaperID: 4021, Poster
Abstract: Flow matching trains a continuum of velocity regression problems but usually reports a loss averaged over interpolation time, obscuring which regions are limited by intrinsic ambiguity and which remain data-limited. We introduce the Velocity Ambiguity Profile (VAP), \mathcalA_v(t), the time-local Bayes-risk floor of velocity prediction, and decompose time-resolved loss as L_N(t)=\mathcalA_v(t)+R(t;N), where only the excess term R(t;N) is reducible by more data or a different estimator. We prove two statistical benchmarks. In an intrinsic local-regression benchmark, \mathcalA_v(t) enters the time-dependent noise factor of the standard nonparametric rate. In separated isotropic Gaussian mixtures, the Bayes velocity becomes locally affine, and a label-oracle residual-regression benchmark has parametric excess-risk scale K\mathcalA_v(t)/N, matched within the same separated residual Gaussian-location subproblem. Controlled Gaussian-mixture experiments validate this diagnostic through a dense transition sweep, calibrated oracle-to-blind and empirical-proxy audits, and a local controlled time-sampler stress test in which a frozen reducible-error proxy reduces reducible loss while cross-cell sensitivity delimits aggressive sampling. VAP turns an averaged flow-matching objective into a time-resolved diagnostic for locating ambiguity-limited and data-limited regions.
PaperID: 4022, Poster
Abstract: In-context learning plays a central role in transformer-based large language models, yet its theoretical understanding remains limited. In this work, we study multiclass in-context classification under more realistic settings, including anisotropic class centers and label imbalance, by introducing a spectral data generation framework that constructs class-center matrices with a prescribed singular spectrum. We first show that the isotropy of class-center vectors, as quantified by the stable rank, improves ICL generalization performance. From a meta-learning perspective, our theorem shows that linear transformers can learn multiclass in-context classification with near-optimal per-label sample complexity, extending prior guarantees beyond the binary setting. In the test label imbalance regime, our analysis reveals that queries from majority classes are easier to classify, while those from minority classes are more error-prone; moreover, robustness to this bias improves with the stable rank. Finally, we empirically demonstrate that our theory is consistent with observations on transformers and pretrained large language models.
PaperID: 4023, Poster
Abstract: Embodied agents in contact-rich environments must model dynamics affectedby frictional uncertainty, contact discontinuities, compliance, actuator imperfec-tions, perception noise, and mass or inertia mismatch. Existing residual dynamicslearning methods correct nominal physics predictions with unconstrained neuralresiduals, which often generalize poorly under unseen contact regimes, changingphysical parameters, long-horizon rollouts, and few-shot adaptation. We proposeQ-Residual Physics, a Hamiltonian-structured quantum residual learning frame-work for embodied dynamics. The method represents residual errors through astructured quantum residual layer whose qubits correspond to physically inter-pretable modes, including contact impulse, friction, slip, compliance, damping,perception error, and mass or inertia mismatch. A physical interaction graphinduces a Hamiltonian residual structure, implemented with a Trotterized param-eterized quantum circuit whose components capture mode activation, physicalcoupling, uncertain switching, dynamic mode exchange, and higher-order interac-tions. Quantum measurements are decoded to correct nominal physics predictions.Across embodied dynamics and manipulation benchmarks, including OMNIPUSH,MANISKILL2, ROBOMIMIC, D4RL, and PHYSION, Q-Residual Physics achievesstronger generalization than physics-only, neural residual, graph residual, andquantum-circuit baselines. It reduces long-horizon rollout error by 23.8%, im-proves contact transition prediction by 16.4%, lowers out-of-distribution parametererror by 21.7%, and improves few-shot adaptation by 28.5%. The code is availableat https://anonymous.4open.science/r/QRP-0205
PaperID: 4024, Poster
Abstract: In many modern applications of reinforcement learning (RL), the natural reward for a task of interest is inherently sparse: a reward of 0 is given everywhere except when the task is completed, when a reward of 1 is given. Training a policy to maximize such a sparse reward requires solving a challenging credit assignment problem, leading to slow or ineffective RL improvement. We propose a simple approach to transform a sparse outcome reward into a dense process reward. Our approach relies on training a discriminator to distinguish between previous successful and unsuccessful trajectories, and using this discriminator to incentivize the RL-learned policy to match the state-action visitations of successful trajectories, while avoiding those of unsuccessful trajectories. By incentivizing the policy to match the visitations over all states, not just those that correspond to task success, this reward provides dense feedback on whether progress is being made towards task completion, and, we show, provably achieves this without changing the optimal policy. Focusing on finetuning of robotic control policies, we demonstrate that our approach leads to significantly faster RL finetuning performance on both simulated and real-world manipulation tasks, as compared to simply maximizing the sparse outcome reward.
PaperID: 4025, Poster
Abstract: Vision-language verifiers judge whether an image faithfully matches a textual description, yet existing training pipelines typically rely on static supervision from stronger external judges or curated hard-negative data. Such supervision is costly, teacher-bounded, and difficult to adapt as model failures evolve. We introduce PROVE (Perturb-and-Repair Optimization for Vision-Language Evaluation), a self-play framework that trains verifiers through a closed perturb–verify–repairverify loop. Starting from a known-correct image-caption pair, a Mutator introduces a typed perturbation over objects, attributes, counts, or relations. The verifier then judges the perturbed caption and provides textual feedback, which a Repairer uses to recover a faithful caption for re-verification. These two verification rounds produce structured rewards for informative mismatches and visually grounded repairs. Since perturbation, verification, and repair are role-conditioned behaviors of the same evolving multimodal backbone, updating the shared parameters also improves the verifier itself. Experiments show that PROVE improves fine-grained compositional verification and visual-detail robustness while preserving general multimodal capability.
Authors:
Yifan Dai, zhenhua wu, Bohan Zeng, Daili Hua, Jialing Liu, Bozhou Li, Yuran Wang, Chengzhuo Tong, Hao Liang, Xiaochen Ma, Junbo Niu, Tianyu Guo, Yang Shi, Yue Ding, Yiyan Ji, Bingyin Mei, Yushuo Guan, Yuanxing Zhang, Pengfei Wan, Fangcheng Fu, Wentao ZhangAbstract: While joint audio-visual understanding is fundamental to advancing machine cognition, current multimodal large language models (MLLMs) still struggle with complex cross-modal reasoning. Existing text-based chain-of-thought (CoT) compresses rich multi-modal features into discrete text, incurring information loss and inducing a language-bound phenomenon that diminishes attention to audio-visual signals. In contrast, a continuous latent space inherently preserves dense representations, serving as an ideal carrier for audio-visual information. Motivated by this, we propose LatentOmni, a novel cross-modal reasoning framework. By introducing a feature-level supervision mechanism to directly reconstruct raw sensory inputs within the latent space, LatentOmni leverages native latent features to bridge audio-visual modalities and text, ensuring sustained attention on original audio-visual inputs throughout reasoning. Furthermore, to maintain temporal consistency across modalities in latent space, we design Omni-Sync Position Embedding (OSPE), which generalizes multimodal rotary position encodings to drive audio-visual synchrony. To supervise this reasoning process, we construct LatentOmni-Instruct-35K, a dataset interleaving text with audio-visual segments that serve as dense evidence for latent reconstruction. Comprehensive evaluation across multiple audio-visual reasoning benchmarks demonstrates that LatentOmni substantially outperforms strong explicit-CoT baselines, validating latent space joint reasoning as a promising path toward genuine omnimodal understanding.
PaperID: 4027, Poster
Abstract: In emotional support conversations, user needs are dynamic, shifting across turns between emotional soothing, cognitive guidance, and their combinations. Existing methods often optimize a single user-state objective, such as emotional relief or cognitive growth, making it difficult to capture such need shifts. In this paper, we formulate emotional support dialogue as a need-aware multi-objective optimization problem, where agents learn to generate responses that balance multiple support objectives according to users' evolving needs. We propose MindFlow, a psychologically grounded user simulator that tracks users' emotional and cognitive needs at each turn, derives state-dependent preference vectors from personalized arousal curves, and models mental-state transitions with a dual-process system. Building on MindFlow, we introduce NeMo, a need-aware multi-objective reinforcement learning framework that estimates local multi-turn effects through forward simulation and dynamically aggregates emotional and cognitive rewards based on current support preferences. By converting delayed, multi-dimensional feedback into step-level optimization signals, NeMo improves credit assignment and aligns policy learning with long-term user-state improvement. Experiments show that NeMo consistently improves user-side state metrics and response-level supportive quality, with robust generalization under both Sim2Sim and Sim2Real evaluations. Code and data are available at https://anonymous.4open.science/r/NeMo-6EB9.
PaperID: 4028, Poster
Authors: Yupei Wang, Neil R Mallinar, Misha Belkin, Alex Warstadt
Abstract: How do language models acquire such a vast array of concepts and abilities from next-token prediction alone? We introduce an interpretability framework that exposes how this single loss relates to different concepts at different points during pretraining. For each layer and checkpoint, we measure the \emphoptimization engagement with a concept, i.e., alignment between a probe-defined concept subspace and the average gradient outer product (AGOP) of the next-token loss. The concept subspace captures where the concept currently lives in the model's representation; AGOP captures the directions in which the loss exerts the strongest pressure. Across 12 lexical, syntactic, and semantic tasks and 19 Pythia-1B and -410M checkpoints, we find that optimization engagement with a concept is temporally aligned with behavioral changes in the model related to the concept. Furthermore, optimization engagement tends to peak before we observe the largest behavioral changes in the model, suggesting that sudden changes in model performance ("grokking") may be preceded by behaviorally unobservable geometric changes. However, this "geometric precedence" pattern is not universally observed; we also see cases where optimization engagement does not peak once, but rather persists at a high value, or peaks multiple times. These results indicate that the loss contributes to concept learning in a myriad of ways that simple probe performance alone cannot reveal.
Abstract: Bayesian Optimization has become an established methodology for minimizing black-box functions of a vector input. Often, however, this parameter vector arises from the discretization of an inherently functional relationship. Several recent articles have considered the Functional Bayesian Optimization (FBO) setting, in which the variable to be optimized is not a member of a finite dimensional vector space, but rather an infinite dimensional function space. In this work, we propose L^0 Manifold Optimization (L0MO), a simple approach to FBO which searches a sub-manifold of a Reproducing Kernel Hilbert Space consisting of functions with a sparse representation in the kernel functions. This results in simpler, more effective functional solutions than existing methods. We discuss in detail the relationship between our method and existing ones, providing a unifying lens through which to view prior works. To assess our method against the state of the art, we conduct an extensive computational study, and along the way develop a novel set of benchmark test functions which port standard finite-dimensional ones to the infinite dimensional domain. Our experiments demonstrate that the proposed method achieves superior performance across a wide range of test benchmarks.
PaperID: 4030, Poster
Authors:
Marvin Theiss, Lukas Braun, Andrew Saxe, Erin GrantAbstract: Representations are routinely used in machine learning, psychology, and neuroscience to probe the computations of biological and artificial systems. Yet it remains unclear to what extent computation constrains representation in artificial neural networks. Neural networks admit : transformations of the parameters that preserve function exactly but reshape representational geometry. We show that known parameter symmetries act on representations through just three primitives: , giving a closed-form descriptor of representational geometry as task-linked features plus symmetry-induced noise. This decomposition yields analytic bounds on representational similarity between networks related by parameter symmetry, which reveal when functionally equivalent networks become arbitrarily representational geometries, which weight features by their computational importance and recover a stable link between representation and computation. Overall, our results delineate when representation can support inferences about computation, and when it cannot.
PaperID: 4031, Poster
Abstract: Georeferenced 3D reconstruction from ground images requires recovering not only the local scene geometry, but also localizing every camera center and 3D point in a geographic coordinate frame. Existing cross-view geo-localization methods are able to estimate the GPS coordinate of a single ground camera, but do not reconstruct dense 3D geometry, while recent feed-forward multi-view 3D reconstruction models predict point maps and camera poses only in a local coordinate frame. We introduce X-VGGT, a feed-forward cross-view geometry transformer that unifies these tasks by jointly processing a set of ground-level images and a georeferenced satellite image to predict ground-view camera poses, dense 3D point maps, and a scene-level transform parameterized by gravity, scale, yaw and translation. This transform projects the local reconstruction into the satellite plane, resulting in the geo-registration of all ground images with the georeferenced satellite image. We further propose an evaluation protocol for this joint task and show that, across multiple cross-view datasets, X-VGGT produces accurate georeferenced reconstructions and outperforms prior cross-view geo-localization and multi-view 3D reconstruction models.
PaperID: 4032, Poster
Authors:
Yirong Zeng, Shen You, Yufei Liu, Hao Cong, Yutai Hou, Xiao Ding, Yuxian Wang, Wu Ning, WangXu, Bibo Cai, Dandan Tu, Wenyu Zang, Ting LiuAbstract: Tool-Integrated Reasoning (TIR) enables LLMs to solve complex, multi-step problems by offloading computation to external tools. While Agentic Reinforcement Learning (RL) has advanced TIR systems, the underlying policy entropy dynamics governing these models remain poorly understood. In this paper, we conduct an extensive empirical study across model architectures and uncover a pervasive phenomenon of entropy-performance decoupling: in the late stages of RL training, policy entropy continues to diverge despite performance stagnation. We identify that this ineffective entropy growth is primarily driven by an accumulation of invalid tool-use trajectories, which trap the agent in non-informative hallucination loops and degrade final reasoning accuracy. To address this bottleneck, we propose EarlyTIR (Early intervention for Tool-Integrated Reasoning), a strategic truncation mechanism that prunes non-informative exploration during the RL rollout phase. EarlyTIR halts trajectories that exhibit excessive invalid interactions while preserving the model’s ability to learn essential error-recovery behaviors. Empirical evaluations across six public benchmarks and multiple model families (7B to 32B) demonstrate that EarlyTIR consistently elevates the performance ceiling, achieving a macro-average accuracy boost of 13.6% on the Qwen3-32B model. Furthermore, our approach enhances inference efficiency, reaching correct solutions with 33.7% fewer tool interactions compared to strong baselines.
PaperID: 4033, Poster
Authors: Yunwei Bai, Ying Kiat Tan, Tsuhan Chen
Abstract: Diffusion image generators owe much of their practical success to accurate denoising networks and strong numerical samplers. Yet the two components are not fully matched. During sampling, the denoiser is asked to make predictions based on imperfect noisy states formed by model predictions that differ from the training noisy states formed by real data. This train-test gap, known otherwise as exposure bias in literature, caps image generation quality (i.e., blurry, lacking in details or unnaturally-shaped). To mitigate the problem, we propose PEEK, a training-free deterministic sampler plugin. At each eligible denoising step, PEEK first performs a provisional one-step update, evaluates the denoiser once at the look-ahead state, and either keeps the current clean prediction or replaces it with this one-step-later prediction to form the actual noisy state. By taking the additional step, we generally obtain a higher-signal, more-image-like prediction for the noisy state, mitigating the exposure bias. To optimize the replacement positions, PEEK searches for a binary replacement schedule with a greedy metric-driven search made efficient by the PEEK prefix and baseline noise-schedule-rebasing caches; for a 9-step search, the caches reduce the one-candidate worst-case vanilla denoiser-evaluation count by 69.3%. Across pixel-space and latent-space diffusion models with unconditional, label-guided, and text-guided settings on CIFAR-10, CelebA-HQ, LSUN Bedroom, ImageNet, and MS-COCO datasets, PEEK improves over NFE-matched DPM-Solver++ and DEIS baselines, with five-run mean FID reductions up to 39.3% under the same deployment cost. Code will be released.
PaperID: 4034, Poster
Abstract: Room acoustics modeling requires capturing intricate wave phenomena beyond direct sound propagation, including reflections, refractions, and diffractions. Recent neural acoustic synthesis methods have progressively incorporated richer scene representations and physics-informed priors, yet typically learn acoustic behavior without considering wave propagation dynamics, which inherently capture diffraction and interference phenomena. We propose WavNAF, a framework that bridges wave dynamics and neural acoustic modeling by leveraging Finite-Difference Time-Domain (FDTD) simulations, which solve the wave equation on discrete spatiotemporal grids to capture these phenomena. However, directly applying FDTD at full scale is computationally prohibitive under CFL stability constraints. Inspired by traditional acoustic scale models, WavNAF constructs a geometry-preserving, scaled digital replica from NeRF-derived scene geometry and acoustically probes it with temporally compressed FDTD simulation. The resulting pressure maps are encoded not as direct full-scale Room Impulse Response (RIR) estimates, but as compressed-domain wave signatures that provide the neural acoustic field with a structured inductive bias beyond static geometry. To bridge the mismatch between these scaled simulation features and full-scale acoustic responses, we introduce a Neural Acoustic Scaling Module that adaptively recalibrates pressure-map features at the feature level for full-scale RIR prediction. Experiments show that WavNAF yields consistent gains over prior neural baselines across standard acoustic metrics.
Abstract: How does cooperation evolve in complex agentic systems? Prior work in evolutionary game theory studies why individuals are incentivized to cooperate by isolating social interactions from the physical costs required to execute them, while artificial life models traditionally study the environments that lead to replication without considering the strategic decisions behind the evolved mechanisms. In contrast, we introduce , a computational model where social interactions, replication mechanisms, and their associated computational costs are endogenous and simultaneously co-evolving. We study these dynamics using a computational substrate of randomly initialized programs in Z80 machine code, showing empirically and theoretically that embedding a social dilemma directly into the physics of computation can favor the emergence of self-replicating, cooperative strategies. When resources are scarce, our analysis shows that : parasitic stealing destroys shared energy, slows execution, and can prevent reliable replication. Empirically, evolved programs suppress stealing across several Z80 environments, while spatial assortment further supports structural complexity and task performance. We further show that the framework can incorporate exogenous pressures, such as math tasks structured as sequential social dilemmas, when rewards are tied to computation budgets. These results suggest that coupling an agent's capacity for computation to its available energy transforms cooperation into a dominant scaffolding for building complex, sustainable, and self-organizing systems.
Authors: Houman Safaai
Abstract: Modeling high-dimensional dependencies while keeping likelihoods tractable remains challenging. Classical vine-copula pipelines are interpretable but can be expensive, while many neural estimators are flexible but less structured. In this work, we propose Vine Denoising Copula (VDC), an amortized vine-copula pipeline for continuous-data, simplified-vine dependence modeling. VDC trains a single bivariate denoising model and reuses it across all vine edges. For each edge, given pseudo-observations, the model predicts a piecewise-constant density grid. We then apply an IPFP/Sinkhorn projection that normalizes mass and drives the marginals to uniformity. This preserves the tractable vine-likelihood structure and the usual copula interpretation while replacing repeated per-edge optimization with GPU inference. Across synthetic and real-data benchmarks, VDC delivers strong bivariate density accuracy, competitive MI/TC estimation, and faster high-dimensional vine fitting. These gains make explicit information estimation and dependence decomposition feasible when repeated vine fitting would otherwise be costly, while conditional downstream tasks remain a limitation.
Authors: Dimitris Avramidis, Alexandra A Lassota, Ulrike Schmidt-Kraepelin, Adrian Vetta
Abstract: Approval-based committee voting has received significant attention in the social choice community. Among the studied rules, Thiele rules, and especially Proportional Approval Voting (PAV), stand out for desirable properties such as proportional representation, Pareto optimality, and support monotonicity. Their main drawback is that computing a Thiele outcome is NP-hard in general. A glimpse of hope comes from the fact that Thiele rules are better behaved under structured preferences. On the candidate interval (CI) domain, they are computable in polynomial time via a linear program (LP) that has a totally unimodular constraint matrix. Surprisingly, this approach fails for the related voter interval (VI) domain, and the complexity of the problem has repeatedly been posed as an open question. Our main result resolves this question: although the relevant matrix is not totally unimodular, the ``standard'' LP still admits at least one optimal integral solution, and we provide a fast algorithm for finding it. Our technique naturally extends to the voter-candidate interval (VCI) domain, also known as the 1-dimensional voter-candidate range (1D-VCR) domain, and to the linearly consistent (LC) domain, both of which generalize the candidate and voter interval domains. Although both the VCI and LC domains have been studied in social choice, their relationship was unknown. We show, through connections to graph theory, that LC strictly contains VCI. We also provide an alternative definition of LC that is closer in spirit to VCI and has a natural interpretation in approval elections; this equivalence may be of independent interest. Finally, we study an alternative tree-based generalization of VCI and show that Thiele rules become NP-hard to compute on this domain.
PaperID: 4038, Poster
Abstract: Existing generative models for earth observation (EO) predominantly rely on fine-tuning natural image priors, which limits their scalability and introduces perspective biases that conflict with geospatial constraints. To address this, we introduce GeoCore-9B, a 9-billion-parameter generative foundation model, which is the first of its scale to be trained from scratch exclusively on EO data. Unlike previous EO foundation models, GeoCore-9B is built upon a Flow Matching-based Diffusion Transformer (DiT) and natively conditions generation on text descriptions and continuous geospatial metadata, including ground sample distances, latitudes, and longitudes. To overcome the convergence and spatial disorientation challenges of training at this scale, we propose a Geospatial Semantic Alignment loss. This objective distills structural Earth surface priors (e.g., terrain and urban areas) from a frozen specialist teacher network, constraining the diffusion latent trajectory during training without adding inference overhead. Pre-trained on the global-scale Git-10M dataset, GeoCore-9B demonstrates strong downstream versatility. Beyond standard proxy generative tasks, we show that GeoCore-9B can be effectively adapted for practical EO applications, including highly challenging tasks such as cloud removal and SAR-to-optical cross-modal translation. Extensive evaluations confirm that GeoCore-9B establishes new state-of-the-art performance in both visual fidelity and geographic structural accuracy. Code and pretrained checkpoints will be released upon acceptance.
Abstract: Memory expertise is a learned skill: knowing what to encode, when to retrieve, and how to organize knowledge---a capacity known in cognitive science as metamemory. We bring this perspective to LLMs by treating memory management as a trainable skill. We promote file-system operations to first-class memory actions alongside task actions, letting the model itself decide how to manage its memory. This memory skill improves along two axes: the structure that supports it (prompts, file schemas, action vocabulary), and the proficiency of the model exercising it. Both axes resist manual optimization: episodes in long-horizon tasks run for thousands of steps, and a single memory mistake can hide long before it surfaces, making human review of full trajectories impractical. We introduce AUTOMEM, a framework that automates both axes. In the first loop, a strong LLM reviews complete agent trajectories and iteratively revises the memory structure that shapes how the agent interacts with its memory files. In the second loop, the agent's own good memory decisions are identified from many episodes and used as training signal to sharpen the model's memory proficiency directly. Across three procedurally generated long-horizon games (Crafter, MiniHack, and NetHack), optimizing memory alone---without modifying the model's task-action behavior---improved the base agent's performance 2x--4x, bringing a 32B open-weight model competitive with frontier systems such as Claude Opus 4.5 and Gemini 3.1 Pro. Our results show that memory management is an independently learnable skill, and a high-leverage objective yielding large gains on long-horizon tasks.
Authors: Yihao Liang, Niraj Jha
Abstract: Distilling vision-language models into faster hybrid architectures, such as 3:1 Mamba-2/attention mixes, is now standard practice for making inference efficient. Aggregate benchmarks suggest that this works but they hide selective failures. When we distill Qwen3-VL-8B-Instruct into a 3:1 Mamba-2/attention hybrid, student model stays within 2 points of the teacher across visual reasoning benchmarks like MMStar, MMBench, and MMMU-Pro, while dropping 13 points on optical-character-recognition and document tasks. The student can still understand the scene but loses the fine-grained text needed to answer. We localize much of the failure to a specific kind of position. In a high-resolution image, most patches are sky, wall, or smooth texture, while a small fraction carries text, edges, object boundaries, or other local details. In a token-level diagnostic, the top 10% highest-density patches have 3.6× larger residual drift than the bottom 10% lowest-density patches and 3.5× larger teacher-masking answer contribution. Uniform weighting devotes many loss terms to low-information background patches, whereas sparse answer-bearing patches receive no special protection. The required intervention is minimal: we replace uniform residual alignment with density-weighted residual alignment, using patch self-dissimilarity as a training-free proxy for position importance. We call this HEED. Compared with normal end-to-end distillation, HEED increases performance by 8.7 points on OCRBench v2 and 5.13 points on a 10-benchmark average. The gain is realized on different teacher models and hybrid architectures. After standard post-training, the student reaches teacher-level performance on the 10-benchmark average with a 4.12× throughput and a 68% memory saving at 128k context, with no additional parameters and no inference-time cost.
PaperID: 4041, Poster
Abstract: The joint expectation of functions of potential outcomes (Y_1,Y_0) is a fundamental quantity in causal inference, particularly when investigating the variation or heterogeneity of causal effects. This paper provides novel assumptions for their identification: comonotonicity and countercomonotonicity. These assumptions offer a unified framework for their identification across discrete, continuous, and mixed outcome variables. In the absence of these assumptions, we derive sharp bounds for two broad classes of functions, yielding new results for bounding moments of individual causal effects, which are central to measuring effect heterogeneity. We present corresponding estimation methods and illustrate them in simulation studies and real-world datasets.
Abstract: Physics-informed neural networks (PINNs) train a single neural approximation by minimizing multiple physics- and data-derived losses, but the gradients of these losses often interfere and can stall optimization. Existing remedies typically treat this pathology either through scalar loss balancing or full-parameter-space gradient surgery, leaving it unclear which intervention is most appropriate. We show that PINN gradient conflict is not a uniform failure mode with one universal remedy. Instead, we identify distinct PINN gradient-conflict regimes, each associated with a different intervention class. Persistent directional conflict may require separate loss-indexed parameter subspaces, magnitude imbalance often favors scalar reweighting, and low or transient conflict may require no extra mitigation. To select between scalar reweighting and a lightweight architectural intervention, we propose a diagnostic-first framework. It profiles a 1000-step unmodified PINN run and, when intervention is warranted, uses one low-rank adapter per loss to create explicit loss-indexed parameter subspaces attached to a shared PINN trunk, providing each loss with a direct gradient pathway. Across more than 60 PDE configurations, including forward, inverse, multi-physics, parameter-varying, and high-dimensional problems up to 50D, persistent directional conflict dominates standard forward K=3 benchmarks and a natural K=4 thermoelastic system, where adapters combined with reweighting yield significant improvements. In contrast, K=3 inverse problems and natural K=5 and K=6 multi-physics systems are largely magnitude-dominated and often favor reweighting alone, while full-parameter-space gradient surgery can fail on heterogeneous parameter spaces. A regime-transition theorem and a blockwise neural tangent kernel analysis explain why gains from these separate adapter parameter subspaces arise selectively in persistent-conflict regimes.
Authors:
JINWOO PARK, Hyeongwon Kang, Seung H Han, Pilsung KangAbstract: Multivariate time-series anomaly detection aims to identify abnormal temporal patterns in complex real-world systems. However, robust unsupervised time-series anomaly detection remains challenging because anomalies can manifest over different temporal ranges and emerge from interactions among variables, while distribution shifts can cause normal patterns to drift over time. To address these challenges, we propose Codebook-based Online-adaptive Multi-scale Embedding for Time-series anomaly detection (COMET), a unified framework with three tightly coupled components. First, Multi-scale Patch Encoder learns correlation-aware patch embeddings by combining variable-specific temporal dynamics with shared inter-variable structure across multiple patch scales. Second, Vector-Quantized Coreset quantizes these embeddings into representative normal prototypes, which are used to detect anomalies through a dual score combining quantization error and density-aware memory distance. Third, Online Codebook Adaptation leverages the same codebook structure to identify reliable normal samples from training-activated codebook entries, and then updates the prototypes through contrastive learning at inference time. Experiments on five benchmark datasets demonstrate that COMET achieves the best performance on 39 out of 45 evaluation metrics compared to seven baselines, showing robust performance across diverse evaluation protocols.
PaperID: 4044, Poster
Authors:
Tao Jun Lin, Jinhong Ni, Ruyi Zha, Yifu Wang, Xibin Song, Pan Ji, Hongdong LiAbstract: Diffusion-based video generation has achieved impressive visual fidelity, yet enforcing temporally consistent motion under explicit control signals remains challenging. Recent noise-warping approaches construct temporally correlated noise by transporting the initial latent noise via diffeomorphic deformation fields derived from optical flow, but this spatial resampling inherently destroys the i.i.d. Gaussian structure of the noise, requiring post-hoc correction steps at the cost of non-trivial computational overhead. We propose Seed-Your-Motion, by conditioning the a priori latent noise on optical flow via per-location orthogonal transformations in latent channel space, parameterized as products of the Householder Transformation. This channel-space formulation has an advantage of preserving the spatial Gaussianity exactly by construction, requiring neither Jacobian corrections nor post-hoc distribution restoration, while at the same time inducing structured cross-frame correlations that ensures geometric consistency across time. With no learnable parameters and a single forward pass, our method is agnostic to model architecture and can be plugged seamlessly into any video diffusion fine-tuning pipeline. Extensive experiments on motion-transfer and camera-controlled video generation benchmarks demonstrate that our method matches or outperform the state-of-the-art noise-warping baseline, while achieving a 2.7× speedup in processing time. The results suggest that the channel-space Householder transformations provide a principled and efficient alternative to noise-warping based motion-controllable video diffusions.
PaperID: 4045, Poster
Abstract: Causal modeling of physical temporal phenomena must handle interventions that act along trajectories, nonstationary induced laws, path-dependent effects, and feedback mediated by dynamics, all challenging in standard causal models. We introduce \emphHamiltonian Causal Models (HCMs), a trajectory-level framework in which observed variables interact with local environments and interventions act as controls of Hamiltonian mechanisms. HCMs separate immutable equations of motion from intervenable mechanisms and define causal effects as discrepancies between interventional path laws. A key motivation for HCMs is their natural interface with non-equilibrium thermodynamics. Entropy production quantifies the irreversibility of a process and is a central causal observable: it is estimable from data and witnesses causal effects along the system's evolution that are invisible to endpoint and cumulative versions of the standard average treatment effect. As in physics, cause-and effect are not primitives of the relation between two random variables but arise from the non-invertibility of the thermodynamic arrow. With this, our paper reconciles the language of statistical causal models and non-stationary thermodynamics, offering new tools to describe causality in a wide range of physical systems.
Abstract: Existing Multimodal Large Language Models (MLLMs) suffer from increased inference costs due to the additional vision tokens introduced by image inputs. In this work, we propose Visual Consistency Learning (ViCO), a novel training algorithm that enables the model to represent images of varying semantic complexities using different numbers of vision tokens. The key idea behind our method is to employ multiple MLP connectors, each with a different image compression ratio, to downsample the vision tokens based on the semantic complexity of the image. During training, we minimize the KL divergence between the responses conditioned on different MLP connectors. At inference time, we introduce an image router, termed Visual Resolution Router (ViR), that automatically selects the appropriate compression rate for each image patch. Compared with existing dynamic high-resolution strategies, which adjust the number of visual tokens based on image resolutions, our method dynamically adapts the number of visual tokens according to semantic complexity. Experimental results demonstrate that our method can reduce the number of vision tokens by up to 50% while maintaining the model’s perception, reasoning, and OCR capabilities. We hope this work will contribute to the development of more efficient MLLMs. The code and models will be released to facilitate future research.
PaperID: 4047, Poster
Abstract: GNN-to-MLP distillation aims to retain the predictive accuracy of a message-passing teacher while deploying a graph-free MLP at inference. Existing methods mainly transfer node-wise predictions or use confidence-based reweighting, but they do not specify where the student should preserve the teacher's graph-induced geometry. We show that this omission leads to two spectral failure modes in the student's representation space. On sparse graphs, the student suffers from spectral underfit, missing high-energy teacher directions concentrated near boundary regions. On dense graphs, it suffers from spectral overfit, retaining spurious directions that the teacher has collapsed through aggregation. Motivated by an energy-weighted teacher--student alignment objective, we propose Graph Geometry-aware MLP (G^2MLP), a training-time distillation framework guided by Ollivier--Ricci curvature. Curvature identifies where the two spectral errors concentrate and is used to allocate supervision between prediction-level and representation-level alignment. The deployed model remains a standard MLP and requires no graph access at inference. Across node-classification benchmarks, G^2MLP consistently improves over graph-free distillation baselines, reduces the teacher--student rank gap in both regimes, and transfers without architectural changes to Graph Transformer teachers and link prediction.
Authors:
LIU Hanqing, Jianjun Cao, Yuanze Li, Zijian ZhouAbstract: Deep neural networks exhibit periodic loss spikes during unregularized long-term training, a phenomenon known as the "Slingshot Mechanism." Existing work usually attributes this phenomenon to intrinsic optimization dynamics, but its triggering mechanism remains unclear. This paper shows that, under commonly used floating-point precision, the Cross-Entropy loss itself can produce a deterministic numerical instability. As training enters a high-confidence stage, the difference between the correct-class logit and the other logits may exceed the absorption-error threshold of float-point number. Then during backpropagation, the gradient of the correct class is rounded exactly to zero, while the gradients of the incorrect classes remain nonzero. This breaks the zero-sum constraint of gradients across classes and introduces a systematic drift in the parameter update of the classifier layer. We prove that this drift forms a positive feedback loop with the feature, causing the global classifier mean and the global feature mean to grow exponentially. We call this mechanism _Numerical Feature Inflation_ (\mathcalNFI). This mechanism explains the rapid norm growth before a Slingshot spike, the subsequent reappearance of gradients, and the resulting loss spike. We further show that \mathcalNFI is not equivalent to an observed loss spike: in more practical tasks, absorption errors may affect only a subset of samples, so spikes do not necessarily occur, while the same zero-sum-breaking mechanism can still drive rapid growth of parameter norms. Our results reinterpret Slingshot as a numerical dynamic of finite-precision training, and provide a testable explanation for abnormal parameter growth and logit divergence in late-stage training.
PaperID: 4049, Poster
Authors:
Jonathan McCart, Pranav Deevi, Mehdi Azabou, Nanda H Krishna, Caleb A McKinney, Nicholas S Card, Sergey D Stavisky, Chethan PandarinathAbstract: Recent advances in speech neuroprostheses have enabled new communication avenues for people who have lost the ability to speak due to neurological illness or injury. These systems can decode neural activity during attempted speech into text by relying on language priors to achieve high decoding accuracy. This performance currently comes at the cost of using a compute-intensive language modeling stack, presenting a major bottleneck for real-world use. Further, current language models incur highly variable and disruptive inference latencies, limiting communication throughput and preventing users from reliably participating in natural conversation. We introduce FETTUCCINE, an end-to-end brain-to-text framework that leverages the speed and portability of automatic speech recognition (ASR) models to address the aforementioned practical challenges, while achieving usable (<10%) word error rates (WER), efficient adaptation to neural non-stationarities, and reliable performance on mobile devices. When evaluated using publicly available data from an intracortical speech neuroprosthesis user, FETTUCCINE outperforms previous end-to-end methods, achieving as low as 4.39% WER. Most notably, our models successfully run on a commercially-available mobile device with throughput reaching up to 370x real-time, providing the first demonstration of brain-to-text decoding on a mobile device. We also show that our models can be successfully finetuned to extend usable performance to future days. Beyond performance, our approach enables localization of the neural features most relevant for decoding via gradient-based salience maps, which we show align with well-established physiological priors across different ASR model families. Taken together, our results show that FETTUCCINE overcomes core infrastructural and computational barriers, yielding a new class of portable brain-to-text communication.
PaperID: 4050, Poster
Abstract: Graph classification benchmarks vary widely in difficulty: some are solved by simple pooled baselines, while others remain challenging for GNNs and graph transformers. We take a data-centric view and ask whether benchmark success is driven by genuine structure-feature reasoning or by label signal already accessible from restricted marginal views. We introduce , a lightweight diagnostic that measures label accessibility through two supervised probes computed without training a graph encoder: (FLA), based on pooled node features with edges removed. Across real graph classification benchmarks and multiple pooled backbones, the resulting SLA-FLA plane strongly organizes test accuracy, with low-alignment datasets forming a consistent regime where standard pooled pipelines struggle. To probe what lies beyond marginal accessibility, we introduce controlled interaction-dominated stress tests, including heterophilic ego-graphs and synthetic tasks whose structural and feature marginals are matched across classes but whose role-feature couplings differ. These results show that high alignment explains much of easy benchmark performance, while low alignment reveals regimes where success requires preserving structure-feature correspondence beyond standard pooled representations.
PaperID: 4051, Poster
Abstract: LLM post-training is usually framed as a choice of optimizer, but KL-regularized RL also makes a hidden geometric choice: it fits the learned policy to a reward-tilted target by reverse KL. This places PPO, GRPO, RLOO, REINFORCE++, and related methods at a fixed mode-seeking point in a broader design space. We propose Mode-Controlled Policy Optimization (MCPO), a geometry-aware recipe that replaces this inherited point with an \alpha-divergence dial over mode-seeking versus mode-covering behavior; \alpha=0 recovers the classical reverse-KL recipe. MCPO admits an off-policy objective, a practical grouped importance-weighted estimator, and a baseline-centered variant for stable training. Across mathematical reasoning, agentic transfer, search-based QA, and safety, no single geometry is uniformly best. Average-case math favors stronger concentration, while pass-based math and Search-QA favor broader support; agentic transfer moves with metric and scale; safety exposes a base-dependent frontier in which the same refusal-only signal can transfer broadly or spill into benign boundary cases depending on the reference model's mode landscape. MCPO's value is therefore not a new universal objective, but a way to expose and select the post-training geometry that RL with fixed reverse-KL geometry silently chooses.
Abstract: Large language model (LLM) agents---LLMs that dynamically interact with an environment over long horizons---have become an increasingly important area of research, enabling automation in complex tasks involving tool-use, web browsing, and dialogue with people. In these settings, traditional policy gradient methods may suffer from unstable learning and poor sample-complexity due to poor credit assignment. Meanwhile, actor-critic methods address long horizons by learning a critic that provides more granular feedback and enables off-policy (and potentially offline) learning, but are heavily dependent on the fidelity of the critic. In this paper, we propose \emphNatural Language Actor-Critic (NLAC), a novel actor-critic algorithm that trains LLM policies using a generative LLM critic that produces values in natural language space. While natural language values have been leveraged in the past to provide a more flexible and actionable training signal, our work is the first to perform general and scalable value learning without relying on in-context aggregation of information from on-policy rollouts. This means our approach can be trained off-policy without policy gradients, offering a more sample-efficient alternative to existing methods. We present results on a mixture of reasoning, web browsing, and tool-use with dialogue tasks, demonstrating that NLAC shows promise in outperforming existing training approaches and offers a more scalable and stable training paradigm for LLM agents.
PaperID: 4053, Poster
Abstract: Personalized portrait generation aims to synthesize portraits that align with both the identity of the reference face and the semantics of the text. However, existing methods suffer from two critical challenges. First, they exhibit a trade-off between ID fidelity and text controllability. We identify that this limitation stems from the self-supervised training, where the model learns a trivial mapping rather than understand identity. Second, despite exploring various encoding strategies, recent approaches still tend to discard fine-grained facial details. To address these challenges, we propose CrossID-11M, a large-scale curated portrait dataset comprising 11 million images with high-quality annotations, paired with millions of identity-preserving face video clips. Furthermore, we introduce CrossID, a novel framework integrating a lightweight GatedFaceAdapter and a cross-supervised training paradigm. Additionally, we establish CrossID-Bench alongside new VLM-based evaluation metrics. Extensive experiments demonstrate that CrossID successfully breaks the trade-off between ID fidelity and controllability, achieving state-of-the-art performance by producing photorealistic images with faithfully preserved ID details.
Abstract: Linear attention has emerged as a cornerstone for efficient long-context architectures, as evidenced by its integration into state-of-the-art open-source models including Qwen3.5/3.6, Kimi Linear, and RWKV-7. Models that incorporate linear attention layers with the so-called Delta-Rule involve the inversion of triangular matrices as a core sub-routine. This operation often forms a performance bottleneck, and, due to its high-sensitivity to numerical errors, it can significantly deteriorate end-to-end model accuracy if it is not carefully implemented. This work provides a systematic analysis of both direct and iterative triangular inversion algorithms, targeting methods that are rich in matrix products, and, therefore, have the potential to efficiently utilize modern hardware. To that end, our analysis covers a broad spectrum of mathematical and practical aspects, with a heavy focus on numerical stability, computational complexity, and, ultimately, hardware efficiency and practical considerations. We provide a rigorous experimental evaluation to verify these properties in practical scenarios, and in low-precision floating-point representations, highlighting the strengths and limitations of each method. Performance benchmarks on NPUs reveal up to 4.3× speed-up against the state-of-the-art implementations of SGLang for triangular matrix inversion, leading to significant performance improvements on the entire layer level, while maintaining full end-to-end model accuracy.
Authors: Inwoo Hwang, Donggeun Lim, Hojun Jang, Young M Kim
Abstract: With recent advances in embodied agents and AR devices, egocentric observations are readily available as input for real-world interactive online applications. However, egocentric viewpoints can only sporadically observe hands, in addition to the estimated head trajectory. We propose EgoForce, an online framework for reconstructing long-term full-body motion from noisy egocentric input. While existing generative frameworks can robustly handle noisy and sparse measurements, they assume a fixed-length observation window is available and are thus not suitable for real-time applications. Faster inference often relies on autoregressive prediction, sacrificing robustness. In contrast, we adopt a diffusion-based method with a temporally asymmetric noise schedule inspired by Diffusion Forcing. Specifically, our approach models temporally evolving uncertainty and incrementally denoises states as new streaming observations arrive. Combined with a noise-robust imputation strategy, EgoForce progressively generates stable and coherent full-body motion under strict causal constraints. Experiments demonstrate that our online framework outperforms existing online and offline methods, enabling long-horizon, full-body motion reconstruction in challenging egocentric scenarios.
PaperID: 4056, Poster
Authors: Mahish Kumar Guru, Mayank Nagar, Ayush vyas, Jan Bohlen, Roland Aydin, Noomane B Khalifa
Abstract: Inverse design of physical systems (molecules, devices, microstructures) often reduces to optimizing a high-dimensional structure against an expensive black-box simulator. Direct search is difficult because the space is non-Euclidean, feasibility is hard to encode, and each evaluation is expensive. We present Co-PiLOT, a latent optimization approach that maps candidates through a generative encoder--decoder, uses the decoder as a learned validity prior, and searches the latent space with physics-informed black-box optimization. The framework is applied on the inverse design of magnesium alloy microstructure/texture. We develop a vision transformer based--encoder; paired with latent diffusion, diffusion transformer and rectified-flow transformer--based decoders on ~80,000 EBSD-derived microstructure dataset to learn a minimal bottleneck, z. The ViT-FMDiT model (z=768) reconstructs high-fidelity microstructure images (FID 23.19, MS-SSIM 0.178), which our self-segmenting orientation codec converts into input grids for crystal plasticity solver. Finally, we introduce Meridian, an active latent optimizer driven by deep-kernel Gaussian-process uncertainty, feasibility prediction, active trust regions, and target-aware acquisition. Within the evaluation budget, the ViT-FMDiT and Meridian combination yields the best target-driven objective score, outperforming DANTE, TuRBO, and BAxUS by ~6% in relative error on the same decoder.
Abstract: Feature learning is widely viewed as the mechanism that distinguishes neural networks from fixed-kernel methods, yet its effect on the underlying function space remains poorly understood. We aim to precisely characterize how the function space (e.g., RKHS) endowed by a two-layer neural network during gradient descent training. We prove that, in a high-dimensional proportional regime, the post-update feature distribution is well approximated by a target-dependent spiked Gaussian covariance, yielding a deformed kernel that can be written as the original isotropic kernel evaluated on geometrically transformed inputs. This characterization enables a spectral analysis of the corresponding integral operator. We prove that the global eigenvalue decay rate is preserved, while the spike selectively alters the leading eigenspaces. For ReLU activations, we derive an explicit perturbative expansion showing that feature learning boosts the eigenvalue of the linear eigenfunction aligned with the target direction and mixes the top radial eigenfunction with a target-aligned quadratic harmonic. These results provide a precise function-space description of early feature learning: gradient descent does not merely rescale a static kernel, but induces a data-adaptive deformation that preferentially enriches directions aligned with the teacher signal.
PaperID: 4058, Poster
Abstract: With the rapid development of Video Large Language Models (VLLMs), online video understanding has emerged as a practical paradigm for real-world streaming applications. A key challenge is to construct a compact, query-agnostic visual memory from continuously growing video streams under strict token budgets. Existing token reduction methods typically rely on static and coarse-grained importance estimation, implicitly assuming that retained tokens contribute equally over time. This overlooks the dynamic nature of visual information, where the relevance of tokens evolves as the video progresses, leading to suboptimal prioritization and inefficient long-term memory usage. To address this limitation, we propose LifeStream, a training-free framework for lifecycle-aware visual token selection and hierarchical memory construction. LifeStream models the temporal evolution of visual tokens and reveals that tokens contribute unequally to downstream reasoning depending on their lifecycle states. Based on this insight, it organizes tokens into a hierarchical memory structure: informative non-stable tokens are preserved in an archive memory, while long-range history is progressively compressed into a global memory according to state-specific priorities. This design enables continuous retention of high-value information under limited token budgets, without requiring query-time retrieval or additional computation. Extensive experiments demonstrate that LifeStream discards over 80% of visual tokens while improving performance. It improves LLaVA-OV-7B and Qwen3-VL-8B by 4.4% and 3.5% on StreamingBench, respectively, and achieves state-of-the-art results on one additional online benchmark and four long-video understanding benchmarks.
PaperID: 4059, Poster
Abstract: Despite rapid progress in synthetic image detection, detectors that perform well on known generators often become brittle when the source model, post-processing pipeline, or scene composition changes. A key reason is that standard binary supervision teaches models to separate training distributions rather than to identify authenticity-relevant evidence: low-level artifacts become shortcuts, while higher-level semantic implausibilities remain weakly grounded. To address this limitation, we propose , a unified forensics preference learning framework that converts multimodal large language model (MLLM) forensic knowledge into image-specific preference supervision. UniFPL constructs forensic preference data from both controlled artifact synthesis and observed artifact harvesting, then learns from list-level evidence rankings that encode which visual-textual cues are more diagnostic of synthetic manipulation. On top of a CLIP-based discriminative detector, UniFPL jointly optimizes authenticity classification, artifacts-aware preference alignment, and semantic structure regularization, allowing the model to inherit contextual forensic reasoning while retaining the stable and efficient inference of a discriminative encoder. Across eleven curated and in-the-wild benchmarks, UniFPL achieves an average balanced accuracy of 93.4% and improves the worst-case accuracy from 81.4% to 86.1% over the strongest baseline, demonstrating stronger cross-generator generalization and robustness to real-world image degradations.
PaperID: 4060, Poster
Abstract: Planning has traditionally been addressed with symbolic search guided by domain models and heuristics. Learning offers a complementary way to exploit structure shared across related problems rather than solving each instance from scratch, making generalization to unseen instances a central challenge for learned planners. Recent work has explored planning mainly through action-centric models that predict actions or full plans from problem descriptions. A parallel line of work instead learns goal-conditioned transition models over states (state-centric learning) and recovers actions by matching predicted successor states to symbolic successors. In this setting, however, prior work has relied mainly on Weisfeiler-Leman (WL) representations, leaving the role of tokenizer choice unclear. We present a controlled study of tokenizer choice in state-centric learned planning, comparing WL, shortest-path, GraphBPE, SimHash, and a deterministic baseline within a common pipeline. Across varied benchmark domains and multiple planner configurations, WL performs well overall, but no tokenizer is best everywhere: the leading representation changes by domain, and tokenizer differences are larger on extrapolation rather than on interpolation tasks. Across the studied configurations, tokenizer choice is the dominant factor in held-out generalization: it explains 57% of performance variance, more than predictor architecture or any other pipeline factor, and its effect is substantially amplified on out-of-distribution problems relative to in-distribution ones. We conclude that representation choice is a strong determinant of generalization in state-centric learned planning and that tokenizer choice is a first-order modeling decision rather than a fixed preprocessing step. Code: https://tinyurl.com/4hre257u
Authors: Hamed Omidvar, Vahideh Akhlaghi
Abstract: Agents built on large language models (LLMs) rely on a range of reliability techniques, including retry, majority voting, and self-consistency, that have been developed in parallel rather than within a common analytical framework. We observe that an LLM sampled at temperature T is a discrete stochastic channel p(y \mid x) in the sense of Shannon's coding theory, and use this identity as the entry point for such a framework grounded in communication theory. Each of these techniques is a special case of one of six classical reliability operators: diversity combining, hybrid retransmission, iterative generator-critic decoding, rateless sampling, structured redundant verification, and difficulty-adaptive routing. Within the framework we give two closed-form results: a noise-variance threshold above which uniform averaging beats quality-weighted averaging, and a contractivity criterion for generator-critic refinement, consistent with a contractive-to-divergent transition we observe between 3B- and 14B-parameter models. We further introduce a cost-aware semantic-nearest-neighbor router whose single Lagrangian knob traverses the quality-cost frontier without retraining. Across six channel configurations spanning local and cloud models on 69 hard tasks, no fixed model-technique-budget choice dominates, motivating per-task allocation. On a 300-item hard split of MMLU, GSM8K, and HumanEval, our router occupies the full empirical Pareto frontier: at matched quality, its normalized cost is \approx56% lower than the strongest fixed technique; at matched normalized cost, it improves quality by \approx7% (26% over single-shot decoding). These results argue for consolidating these reliability techniques into a single tunable layer informed by channel coding.
PaperID: 4062, Poster
Abstract: Facial behaviour analysis requires understanding valence–arousal (VA), expressions (EXP), action units (AUs), yet existing systems treat these tasks independently or use naïve MTL, ignoring their structured dependencies. This leads to negative transfer, conflicting gradients, poor cross-db generalisation, especially with heterogeneous datasets with only partial labels. We propose Behaviour4All, the first dependency-aware toolkit that unifies VA, EXP, AUs via explicit relational priors and principled optimisation framework. Behaviour4All integrates psychophysically grounded task mappings with coordinated loss families to mitigate gradient conflicts and enforce coherent cross-task semantics. Across eight in-the-wild datasets, Behaviour4All achieves sota performance, improves fairness and generalisation, and outperforms vanilla MTL and student--teacher baselines. We will release the complete toolkit (code, pre-trained models, etc) upon acceptance.
PaperID: 4063, Poster
Abstract: Static equilibria—Nash, Correlated (CE), and Coarse Correlated Equilibria (CCE)—and the Price of Anarchy (PoA) provide essential, tractable benchmarks for multi-agent systems. However, evaluating learning dynamics purely through C^0 fixed points and discrete empirical distributions abstracts away critical C^1 vector field information. We demonstrate that this reduction diverges significantly from physical learning trajectories. First, the worst-case pure Nash equilibria anchoring canonical PoA bounds mathematically manifest as strict saddles—and in canonical instances, even as global maxima of the exact potential; because they are topologically unstable repellers, natural learning actively bypasses these theoretical inefficiency bounds. Additionally, the PoA metric itself exhibits algebraic sensitivity; relaxing syntactic constraints to accommodate data-driven, strictly positive affine cost models renders the Price of Anarchy unbounded. Furthermore, evaluating algorithms strictly through time-averaged regret minimization to reach CCE or Proximal CE (PCE) structurally permits convergence to strictly dominated strategies. Even enforcing optimal O(1/T) swap-regret minimization provably accommodates chaotic limit sets in minimal normal-form games. Finally, in non-atomic congestion games, discrete-time learning natively bifurcates into Li-Yorke chaos, driving time-averaged inefficiency to scale exponentially as 2^p, diverging from static sub-linear bounds. Collectively, these findings highlight the necessity of augmenting classical algebraic frameworks with dynamically grounded metrics.
PaperID: 4064, Poster
Abstract: Referring Remote Sensing Image Segmentation (RRSIS) aims to segment target objects from remote sensing images based on natural language descriptions, requiring precise vision–language alignment under significant ambiguity. Despite the strong generalization ability of Segment Anything Model 3 (SAM3), the large spatial extent and varying scales of targets in remote sensing images intensify the ambiguity of referring expressions. This poses a major challenge for fine-grained segmentation. To address the challenge, we propose , an evidence-driven parameter-efficient adaptation framework that mimics human decision-making by dynamically accumulating partial evidence. Specifically, EviSAM3 progressively integrates cross-modal evidence through an evidence memory bank, where each incoming piece is aggregated via momentum-based updates. As evidence accumulates, the model performs iterative evidence rectification to resolve ambiguity under incomplete observations. Meanwhile, the contributions of different evidence are adaptively modulated based on their diagnostic relevance, allowing more informative evidence to dominate the reasoning process and guide the final prediction. Extensive experiments demonstrate that EviSAM3 consistently outperforms state-of-the-art methods, particularly in challenging scenarios with high ambiguity and significant scale variation. The code is available in the supplementary material.
PaperID: 4065, Poster
Abstract: Offline goal-conditioned reinforcement learning (GCRL) learns goal-directed policies from reward-free data, but in long-horizon tasks, goal-conditioned value functions often provide unstable guidance due to sparse rewards and discounting. Hierarchical methods partially mitigate this issue via subgoal decomposition; however, high-level decision-making still relies on noise-sensitive value estimates, leading to unstable behavior in complex environments. We address this limitation by proposing Diffusion Subgoal Planning (DSP), a diffusion-based framework for high-level subgoal generation. DSP casts high-level planning as guided generative inference over goal-conditioned subgoals and learns both conditional and unconditional flows, enabling classifier-free guidance to introduce a goal-directed bias at inference time. By removing explicit value-based guidance from high-level planning, DSP generates reachable and goal-directed subgoals through a generative model while retaining hierarchical execution. Experiments on offline GCRL benchmarks demonstrate that DSP outperforms prior methods on a range of navigation and manipulation tasks, with particularly strong performance in maze environments that require multi-step subgoal planning.
PaperID: 4066, Poster
Authors:
Xinyu Li, Zhen Zhang, Jiayu Huang, Daniel M Steinberg, Lina Yao, Prof Javen Qinfeng ShiAbstract: Accurate and reliable uncertainty estimation (UE) for vector-valued physical properties is crucial for scientific discovery in fields like drug and materials discovery. For example, atomic forces are central for finding optimal structures and identifying equilibrium systems, and a fundamental requirement of these vectors is that they must be equivariant to 3D rotations. However, existing UQ methods often fail to respect these geometric constraints, leading to poorly calibrated uncertainty and degraded predictive performance. To address the problem, we introduce a novel framework for equivariant multivariate evidential regression. Our primary contribution is a theoretically grounded parameterization of the evidential prior's scale matrix, constructed via a decomposition into an invariant scalar component and an equivariant low-rank term. Furthermore, we address the lack of suitable evaluation standards for vector fields by introducing a rigorous calibration metric based on the Euclidean distance. Extensive experiments on molecular property prediction benchmarks, including MD17, QM7-X and OC20, show that our framework consistently outperforms established baselines, achieving state-of-the-art predictive accuracy with well-calibrated uncertainty. This work provides a principled and practical approach to uncertainty estimation for equivariant vector-valued properties, paving the way for more trustworthy machine learning applications in scientific discovery.
PaperID: 4067, Poster
Abstract: Existing low-light image enhancement methods are increasingly built on deep representation learning. However, most of them still encode features as deterministic points in the embedding space, making it difficult to capture the statistical uncertainty. Under complex low-light scenes, latent neighborhood relations tend to reflect degradation patterns rather than semantic content or local structures, which leads to statistical shifts and reduced discriminability of representations. Therefore, we propose Prior-Anchored Local Statistical Representation Rectification (PaLSR), which models low-light enhancement as a representation learning process driven by local statistical rectification. PaLSR first learns a normal-light statistical prior with K Gaussian anchors, which provides a shared coordinate system for local representation. Instead of enhancing degraded features in the Euclidean embedding space, PaLSR represents each feature as mean-variance offsets relative to these anchors, describing its content displacement and uncertainty. With these offsets rectified, local representations are reorganized under complex low-light degradations. Extensive experiments on multiple low-light benchmarks and different network architectures show that PaLSR achieves consistent improvements in restoration quality. These results validate the effectiveness of prior-anchored local statistical representation rectification under complex degradation conditions.
PaperID: 4068, Poster
Abstract: We present the Multi-Block DC (BDC) class, a rich class of structured nonconvex functions that admit a DC (``difference-of-convex'') decomposition across parameter blocks. This multi-block class not only subsumes the usual DC programming, but also turns out to be provably more powerful. Specifically, we demonstrate how standard models (e.g., polynomials and tensor factorization) must have DC decompositions of exponential size, while their BDC formulation is polynomial. This separation in complexity also underscores another key aspect: unlike DC formulations, obtaining BDC formulations for problems is vastly easier and constructive. We illustrate this aspect by presenting explicit BDC formulations for modern tasks such as deep ReLU networks, a result with no known equivalent in the DC class. Moreover, we complement the theory by developing algorithms with non-asymptotic convergence theory, including both batch and stochastic settings, and demonstrate the broad applicability of our method through several applications.
PaperID: 4069, Poster
Abstract: Image classification models often encounter objects in diverse and unpredictable contexts at test time, which may differ substantially from those seen during training. This phenomenon, termed , exposes a key fragility in neural networks: when models rely on non-causal contextual cues for prediction, shifts in context can lead to significant performance degradation. In this paper, we propose a novel training framework to improve the robustness of convolutional neural networks under context shift. Our approach leverages spatial feature attribution to guide models toward making predictions that rely less on contextual regions and more on object-relevant features. As a first step, we identify a theoretical limitation of existing feature attribution methods and introduce a new variant, , which produces more faithful attribution maps of model predictions. Building on ContrastiveCAMs, we further propose (CR-CE), a modification of the standard cross-entropy loss that regularizes the model’s attention by suppressing the influence of contextual regions, thereby improving robustness to context shift. We evaluate the effectiveness of our approach on several medium-to-large scale datasets (Waterbirds, Spawrious, ImageNet/ImageNet-BG, Hard-ImageNet), and report consistent improvements in context-shift robustness.
Abstract: Sparse Mixture of Experts (MoE) models offer a scalable and efficient architecture for training large neural networks by activating only a subset of parameters (“experts”) for each input. A learned router computes a distribution over these experts, and assigns input tokens to a small subset. However, without auxiliary balancing mechanisms, routers often converge to using only a few experts, severely limiting model capacity and degrading performance. Most current load balancing mechanisms encourage a uniform routing probability across experts. Early in pretraining, this can result in inconsistent routing behavior, resulting in the model spending its capacity learning redundant knowledge. We address this by introducing a novel load balancing loss that helps preserve relational structure, encouraging consistent expert choices for similar inputs during training. Our experimental results show that applying our loss to the router results in 36% faster convergence and lower redundancy compared to a popular load balancing loss.
Authors: Georgii Serbin, Kirill Koshkin, Zhongao Sun, Anastasiya Bistrigova, C. C Korikov
Abstract: Contextual sparsity is one of the approaches used to reduce computational complexity in the inference process of large language models (LLMs). Existing techniques for efficient LLM inference acceleration based on contextual sparsity with minimal accuracy degradation require training sparse pattern predictors. This paper presents a framework for accelerating the inference of ReGLU-based feed-forward networks (FFNs) within LLMs. The proposed framework provides a fast, training-free method for building sparse pattern predictors using truncation-aware singular value decomposition (SVD) of the gate projection matrix, along with a threshold calibration algorithm, and inference executors supporting conditional computation on CUDA and CANN devices. Experiments on three sparse LLMs with an average activation sparsity level of 90% in the FFNs demonstrate up to a 1.8x end-to-end decoding time speedup with less than 1% degradation in benchmark scores on tasks involving math and code generation. This work advances the deployment of LLMs on edge devices.
PaperID: 4072, Poster
Abstract: Contrastive embedding models trained with scale-invariant losses are typically paired with distance metrics like cosine similarity, effectively ignoring embedding magnitudes. However, surprisingly, empirical studies reveal that despite this, these "discarded" norms seem to correlate with semantic properties such as concept specificity, token frequency, and human uncertainty. In this work, we provide a formal theoretical framework explaining this phenomenon. By analyzing the optimization dynamics, we derive an analytic formula demonstrating that embedding length naturally encodes this information as a byproduct of the training process. We also show how this gives rise to signals that can serve as "free" calibration tools in specific models and retrieval tasks, providing a grounded explanation for a previously heuristic observation.
Abstract: Causal discovery aims to recover directed causal relations from observational and interventional data, providing a basis for mechanistic understanding and reliable decision-making. Causal discovery foundation models (CDFMs) seek to amortize this problem by mapping a dataset directly to a causal graph in a single forward pass, avoiding per-dataset testing, search, or optimization. However, existing CDFMs remain limited, often failing to consistently match strong classical methods, and we find that a key bottleneck lies not only in model architecture but also in how causal pretraining tasks are constructed. Based on this observation, we propose CausalTab, a data-driven CDFM trained with broad causal pretraining over diverse graph priors, structural mechanisms, noise models, dimensions, sample sizes, and intervention regimes. A dynamic task construction strategy composes these causal environments into varied discovery tasks, enabling more transferable structural learning from observational and mixed-interventional data. On large-scale synthetic benchmarks, CausalTab achieves stronger and more robust performance than diverse causal discovery baselines. To further bridge abstract synthetic generators and realistic causal reasoning scenarios, we introduce an expert-knowledge-guided and LLM-audited semantic causal environment benchmark, where domain-grounded SCMs generate interpretable observational and interventional datasets for out-of-distribution analysis. Across both synthetic and semantic environments,CausalTab demonstrates robust structure recovery, especially under interventional evidence, highlighting broad causal pretraining as a key ingredient for transferable amortized causal discovery.
Abstract: A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities, and is not available from any single modality alone. While most approaches operate at the architectural level through larger or more complex fusion models, we propose a complementary axis: shaping the training objective itself. Standard training often emphasizes unimodal or redundant information, falling short on examples that require cross-modal reasoning. We formalize multimodal synergy through information theory and introduce the Synergistic Information Bottleneck (SynIB), a scalable objective that targets synergy directly. To prioritize learning synergy, SynIB motivates the model to predict accurately from all modalities while penalizing confidence when information from any modality is withheld. Alongside the standard task loss, the model runs forward passes with one modality masked at a time and is penalized for remaining confident, which would indicate reliance on unimodal cues rather than cross-modal interactions. We validate SynIB in two regimes. On synthetic XOR tasks where the ground-truth synergy is known by construction, standard training fails to recover it while SynIB does. On five real-world benchmarks, including three MultiBench affective tasks, Hateful Memes with CLIP-ViT and DeBERTa backbones, and a controllable irony extension of CREMA-D we introduce, SynIB improves accuracy on synergy-dependent examples by up to 7.8% and overall accuracy by up to 3.8%.
PaperID: 4075, Poster
Abstract: We study first-order learning dynamics in Bayesian signaling games, a minimal extensive-form model of strategic interaction under asymmetric information. In these games, a Sender observes a private type and chooses a signal, after which a Receiver forms a posterior belief and selects an action. This Bayesian belief update makes the induced learning field fundamentally different from the affine game-gradient fields of normal-form games. Empirically, we find a sharp dichotomy: standard online learning algorithms reliably converge to strict Perfect Bayesian Equilibria (PE), but not to non-strict ones. We show that this behavior cannot be explained by the global geometric conditions commonly used to prove last-iterate convergence, such as monotonicity or the Minty variational inequality. Indeed, these conditions can fail even in binary signaling games. Instead, we identify the relevant local geometry. Our main result characterizes variational stability in finite signaling games: a PE is variationally stable if and only if it is strict. This yields local last-iterate convergence guarantees for online mirror descent and regularized dual averaging near every strict PE. We then study global average-iterate convergence and identify a class of signaling games in which coarse-trigger regret-minimizing dynamics converge to the unique PE. Finally, we complement the theory with experiments comparing projected gradient ascent, mirror-based methods, optimistic variants, and PPO across canonical signaling-game families. The results position signaling games as a compact testbed for understanding how Bayesian belief formation reshapes the convergence geometry of multi-agent learning.
Abstract: Sparse attention reduces compute and memory bandwidth for long-context LLM inference. However, two key challenges remain: (1) KV cache capacity still grows with sequence length, and offloading to CPU memory introduces a PCIe transfer bottleneck; (2) the sparse selection step itself retains O(T^2) complexity and can dominate attention cost at long contexts. We propose SparDA, a decoupled sparse attention architecture that introduces a fourth per-layer projection, the Forecast, alongside Query, Key, and Value. The Forecast predicts the KV blocks needed by the next layer, enabling lookahead selection that overlaps CPU-to-GPU prefetch with current-layer execution. Because Forecast is decoupled from the attention query, our GQA implementation uses one Forecast head per GQA group, reducing selection overhead versus the original multi-head selector. SparDA adds <0.5% parameters and trains only the Forecast projections by matching the original selector's attention distribution. On two sparse-pretrained 8B models, SparDA matches or slightly improves accuracy and delivers up to 1.25× prefill speedup and 1.7× decode speedup over the sparse-attention offload baseline. By enabling larger feasible batch sizes on a single GPU, SparDA further reaches up to 5.3× higher decode throughput than the non-offload sparse baseline.
Abstract: Intent-obfuscation-based jailbreak attacks on multimodal large language models (MLLMs) transform a harmful query into a concealed multimodal input to bypass safety mechanisms. We show that such attacks are governed by a \emphreconstruction--concealment tradeoff: the transformed input must hide harmful intent from safety filters while remaining recoverable enough for the victim model to reconstruct the original request. Through a reconstruction analysis of three representative black-box methods, we find that existing transformations struggle to balance this tradeoff, limiting their effectiveness. In contrast, we show that character-removed variants achieve a better balance. Building on this, we propose \emphconcealment-aware variant construction, which greedily selects character-removed variants that are low in harmful-keyword alignment and mutually diverse, and instantiates them through five modality-aware prompting strategies. We further introduce \emphkeyword-related distractor images that depict the harmful keyword in diverse contexts, providing more effective auxiliary visual context than generic distractor images. Experiments across closed-source and open-source MLLMs show the proposed strategies outperform strong baselines, revealing an underexplored vulnerability: a model's own reconstruction ability can be exploited to recover hidden harmful intent and produce unsafe responses.
Authors:
William Lehn-Schiøler, Magnus R Kjaer, Rahul Thapa, Magnus G Pedersen, Anton M Storgaard, Nick Williams, Andreas Brink-Kjaer, Tue Lehn-schiøler, Sadasivan Puthusserypady, James Zou, Lars Kai HansenAbstract: EEG foundation models achieve state-of-the-art clinical performance, yet the internal computations driving their predictions remain opaque: a barrier to clinical trust. We apply TopK Sparse Autoencoders (SAEs) across three architecturally distinct EEG transformers: SleepFM, REVE, and LaBraM to extract sparse feature dictionaries from their embeddings. By grounding these features in a clinical taxonomy (abnormality, age, sex, and medication), we benchmark monosemanticity and entanglement across architectures. A single hyperparameter procedure, driven by an intrinsic dictionary health audit, transfers robustly across all three architectures. Via concept steering, we introduce a "target vs. off-target" probe area metric to quantify steering selectivity and reveal three operational regimes: selectively steerable, encoded but entangled, and non-encoded. This framework exposes critical representational failures: "wrecking-ball" interventions that collapse global model performance, and clinical entanglements, such as age–pathology confounding, where it is impossible to suppress one concept without corrupting the other. Finally, a spectral decoder maps these interventions back to the amplitude spectrum, translating latent manipulations into physiologically interpretable frequency signatures, such as pathological slow-wave suppression and \alpha-band restoration.
Authors:
Zhongye Liu, Yaopei Zeng, Yurui Chang, Lu LinAbstract: While multimodal large language models (MLLMs) have shown strong visual reasoning abilities, serving a large model for every query is computationally expensive. MLLM cascades mitigate this cost by first querying a weak but cheaper model and deferring to a strong model when the weak model's output is unconfident. However, since the weak model's confidence directly controls compute allocation, these systems expose a new attack surface: an adversary can manipulate confidence so that their queries are consistently deferred to the strong model. Motivated by this vulnerability, we introduce the Forced Deferral Attack (FDA), an adversarial image attack that lowers the weak model's confidence and causes cascades to route queries to the strong model. FDA learns a universal border trigger by optimizing a temperature-flattened objective. This objective pushes the weak model's token distribution on triggered inputs toward less concentrated targets constructed from its clean responses. Across datasets, model families, and deferral metrics, FDA consistently increases strong-model routing while outperforming image-perturbation and prompt-injection baselines. These results show that MLLM cascades are vulnerable to attacks that manipulate compute allocation, forcing unintended strong-model usage without directly targeting answer correctness.
PaperID: 4080, Poster
Authors: Harsh Udai, Konda Reddy Mopuri, Vineeth N Balasubramanian
Abstract: CLIP-style vision-language models transfer broadly but remain brittle for compositional reasoning, often collapsing object-attribute binding and inter-object relations. We attribute this gap to limited compositional diversity in web-scale data and contrastive objectives that prioritize global alignment over grounded structure. Many fixes rely on caption-edited hard negatives, which can bias learning toward specific perturbation templates and trade off against general alignment. We propose GPS-CLIP, which improves compositionality by adding multi-granular positive constraints to standard contrastive learning: IoU-weighted multi-positive region-span alignment and relation-triplet alignment in the shared embedding space. GPS-CLIP uses existing grounded annotations during fine-tuning, keeps the dual-encoder architecture unchanged, and requires only images and text at inference. Across SugarCrepe, What'sUp, COLA, and Winoground, GPS-CLIP achieves state-of-the-art compositional performance while improving cross-modal retrieval and zero-shot recognition. Beyond compositionality, GPS-CLIP (i) achieves the strongest MMVP-VLM results for fine-grained visual understanding and (ii) yields a more structured embedding space with improved object/attribute separability on UT-Zappos geometric analysis compared to hard-negative variants. Overall, GPS-CLIP offers a streamlined way to improve VLM compositional robustness.
PaperID: 4081, Poster
Abstract: We study connections between the sample cost of replicability, which requires an algorithm to produce consistent outputs in different runs, and the sample complexity in the high success probability regime for distribution testing algorithms. Specifically, we consider a stronger notion of replicability, which we term as strong replicability, and show that any \rho-strongly-replicable testing algorithm can be converted in a blackbox manner to a tester with at most \exp(-\tilde\Omega(\rho^-2)) failure probability under the same sample budget. Our conversion is optimal for many natural testers. As applications of our conversion theorem, we obtain tight sample lower bounds for strongly replicable testers for coin bias and closeness testing.
PaperID: 4082, Poster
Authors: Andoni Rodríguez, Cesar Sanchez
Abstract: Reactive synthesis from temporal logic modulo theories (LTLt) automatically con- structs correct-by-construction controllers for systems with arithmetic constraints— elevators, thermostats, robotic patrols—but the state-of-the-art CEGRES algorithm times out on 77% of existing benchmarks due to a fundamental bottleneck: it discovers the universally valid theory tautologies needed to refine its Boolean abstraction one at a time, through expensive counterexample analysis. We present NEUROSYNTHEOS, a neural augmentation of the CEGRES loop that predicts batches of tautologies at each refinement step. We formalize template invariance for parametric specification families, showing that the prediction problem reduces to constant re-instantiation when specifications share algebraic structure. We train a GNN over formula abstract syntax trees and an autoregressive transformer on tautology traces from 104 solved benchmarks. NEUROSYNTHEOS establishes a new state of the art for LTLt synthesis: on 83 standard benchmarks, it solves 60 ±1.2 specifications (vs. 19 for the previous best), resolving 41 that no prior method could handle within a 5-minute budget, while reducing the median itera- tion count by 5.2×with under 3 seconds of neural overhead. All predictions are SMT-verified before injection, preserving full soundness guarantees.
PaperID: 4083, Poster
Authors:
Pinlong Zhao, Xiaoling Zhou, Zhou Zhaoting, Guangyuan DongAbstract: Deployed language models are increasingly expected to remove private, copyrighted, or hazardous content without full retraining. Existing unlearning methods remain brittle because they usually optimize a fixed forgetting objective with a global update schedule: weak updates leave paraphrased knowledge intact, while stronger updates can damage retained capabilities or nearby facts. We identify this failure mode as update homogeneity. To address it, we introduce METAFORGET, a meta-optimization framework that treats unlearning as learning an audit-driven update policy rather than designing another standalone forgetting loss. A lightweight controller maps memorization, paraphrase-sensitivity, and signed forget-retain conflict signals to span weights, layer gates, step sizes, and retain-gradient projection strengths. The controller is trained with a bilevel objective that differentiates through short LoRA unlearning trajectories and evaluates the resulting model on held-out forget, retain, locality, and leakage audits. This policy can instantiate NPO- or RMU-style losses while changing where and how strongly they act. Across TOFU, WMDP, and MUSE-style sequential deletion settings, METAFORGET improves the forget-retain-robustness frontier over strong baselines, moves loss-based membership-inference diagnostics closer to chance, and better preserves utility under repeated requests. The results support a view of practical LLM unlearning as audit-driven policy learning: the central object is not only the forget loss, but the localized update that must pass the deletion audit.
Authors:
Qian, Xinran Yu, Xinan Wang, Danyang Li, Guoxuan Chi, Zheng Yang, Qiang Ma, Xin MiaoAbstract: Visual token pruning is a promising approach for reducing the computational cost of vision–language models (VLMs), and existing methods often rely on early pruning decisions to improve efficiency. While effective on coarse-grained reasoning tasks, they suffer from significant performance degradation on tasks requiring fine-grained visual details. Through layer-wise analysis, we reveal substantial discrepancies in visual token importance across layers, showing that tokens deemed unimportant at shallow layers can later become highly relevant for text-conditioned reasoning. To avoid irreversible critical information loss caused by premature pruning, we introduce a new pruning paradigm, termed bypass, which preserves unselected visual tokens and forwards them to subsequent pruning stages for re-evaluation. Building on this paradigm, we propose SwiftVLM, a simple and training-free method that performs pruning at model-specific layers with strong visual token selection capability, while enabling independent pruning decisions across layers. Experiments across multiple VLMs and benchmarks demonstrate that SwiftVLM consistently outperforms existing pruning strategies, achieving superior accuracy–efficiency trade-offs and more faithful visual token selection behavior.
Abstract: Recent Continuous Thought Machine architecture decouples internal computation from external inputs via neural dynamics, but relies on multi-layer perceptrons without stability guarantees. We propose to model neural dynamics using asymmetric Excitatory-Inhibitory (E-I) networks, which can be stabilized via principles from network theory and can be expressed as energy-based systems optimized through a game-theoretic loss. Building on this perspective, we introduce Temporal Inhibitory-Excitatory Dynamic Engine (TIDE), a neuro-inspired architecture that computes internal representations through neural dynamics stabilized by incorporating the Wilson-Cowan dynamics and lateral inhibition. TIDE balances biological realism by, for instance, using Hierarchical Receptive Fields and enforcing Dale's principle to ensure a realistic 80:20 E-I balance ratio with an end-to-end trainable architecture. The aim of this paper is to introduce a new architecture that brings neuro-inspired learning to the forefront. We present proofs of convergence, stability, and complexity bounds, along with empirical ablation studies. Overall, TIDE surpasses CTM with under 50% of the training time and improves \texttttop-1 accuracy by an average of +1.65% on ImageNet under various perturbations.
Abstract: World action models (WAMs) have emerged as a promising direction for robot policy learning, as they can leverage powerful video backbones to model the future states. However, existing approaches often rely on separate action modules, or use action representations that are not pixel-grounded, making it difficult to fully exploit the pretrained knowledge of video models and limiting transfer across viewpoints and environments. In this work, we present Action Images, a unified world action model that formulates policy learning as multiview video generation. Instead of encoding control as low-dimensional tokens, we translate 7-DoF robot actions into interpretable action images: multi-view action videos that are grounded in 2D pixels and explicitly track robot-arm motion. This pixel-grounded action representation allows the video backbone itself to act as a zero-shot policy, without a separate policy head or action module. Beyond control, the same unified model supports video-action joint generation, action-conditioned video generation, and action labeling under a shared representation. On RLBench and real-world evaluations, our model achieves the strongest zero-shot success rates and improves video–action joint generation quality over prior video-space world models, suggesting that interpretable action images are a promising route to policy learning.
Authors: Disha Hegde, Marvin Pförtner, Jon Cockayne
Abstract: Probabilistic linear solvers (PLSs) return probability distributions that quantify uncertainty due to limited computation in the solution of linear systems. The literature has traditionally distinguished between Bayesian PLSs, which condition a prior on information obtained from projections of the linear system, and probabilistic iterative methods (PIMs), which lift classical iterative solvers to probability space. In this work we show this dichotomy to be false: Bayesian PLSs are a special case of non-stationary affine PIMs. In addition, we prove that any realistic affine PIM is calibrated. These results motivate a focus on (non-stationary) affine PIMs, but their practical adoption has been limited by the significant manual effort required to implement them. To address this, we introduce , an algorithmic framework that automatically constructs a PIM from a standard implementation of an affine iterative method by passing symbolic tracers through the computation to build an affine computational graph. We show how this graph can be transformed to compute posterior covariances, and how equality saturation can be used to perform algebraic simplifications required for computation under specific prior choices. We demonstrate the framework by automatically generating a probabilistic multigrid solver and evaluate its performance in the context of Gaussian process approximation.
PaperID: 4088, Poster
Abstract: We study how the MLP activation function shapes the long-time geometry of token representations in a deep transformer. We work in the stochastic transformer dynamics of [Agazzi et.al. 2026], in which tokens evolve as interacting particles on the sphere under deterministic attention and a common MLP noise whose covariance is the activation's NNGP kernel. For each of ReLU, ReLU², ReLU³, GELU, SiLU, and tanh, we identify the dimensions at which common noise stops driving tokens to full synchronization and sustains non-degenerate spread. We further prove that, even in the presence of attention, ReLU noise admits dispersion ceilings close to consensus, while ReLU^2 noise prevents complete token collapse under an explicit attention strength condition.
PaperID: 4089, Poster
Abstract: Scaling multi-agent systems driven by large language models is transitioning from lossy text-based communication to continuous latent interaction. However, existing latent interaction approaches relying on concatenation-based Softmax cross-attention suffer from an Attention Distraction Dilemma: due to the inherent anisotropy of latent spaces and the semantic misalignment induced by agent role heterogeneity, non-linear Softmax structurally amplifies irrelevant tokens, diluting the collaborative signal as interactions deepen. To address this issue, we propose a training-free additive mechanism that models upstream representations as a semantic manifold. By shifting from concatenation-based Softmax competition to linear subspace projection, this perspective inherently bypasses scalar noise accumulation. Specifically, we introduce two components: a Resonance Filter that extracts principal semantic directional alignments to overcome latent anisotropy, and an SNR-Adaptive Spectral Gain that dynamically regulates this projection to counteract role heterogeneity. We analyze this approach using the Asymptotic SNR Tipping Point Theorem, illustrating the scaling advantages of the mechanism beyond a critical noise threshold. Experiments indicate that the proposed approach outperforms baselines in 46 of 54 evaluated configurations, maintaining semantic reasoning in long-context collaborations.
PaperID: 4090, Poster
Authors: Reed Naidoo, Matt De Vries, Georgios Skiadas, Vicky Bousgouni, Ugur Moustafa, Chris Bakal
Abstract: Learning meaningful representations from cellular microscopy images remains a central challenge in computational biology and computer vision. Recent advances in self-supervised learning (SSL) have shown that scaling model and dataset size can improve the recovery of biological relationships such as gene–gene or compound–target associations. However, this progress has been driven primarily by scale rather than structure—requiring vast data and compute while offering no guarantee that the resulting embeddings capture underlying biology. We introduce NOVA, a structure-first framework for learning biologically aligned representations of cellular perturbations. NOVA integrates curated biological graphs directly into the training objective of a vector-quantised transformer, aligning embeddings of related genes and pathways through graph-based regularisation. By incorporating these relational priors during training, NOVA learns biologically faithful representations with orders of magnitude less data and compute than current state-of-the-art SSL models, while achieving superior performance on gene-gene retrieval benchmarks. This approach reframes representation learning for microscopy from a problem of scale to one of structure, offering a more systematic and efficient route toward biologically grounded embeddings.
Abstract: Flexible continuous-time survival modeling is critical for capturing complex temporal dynamics in high-dimensional data; however, training such models remains challenging due to the intractable integral required for likelihood estimation. We introduce QSurv, a scalable deep learning framework that enables nonparametric continuous-time modeling without relying on time discretization or restrictive distributional assumptions. We propose a training objective based on Gauss-Legendre numerical quadrature, which approximates the cumulative hazard with high-order accuracy while facilitating efficient end-to-end training via standard backpropagation. Furthermore, to effectively capture non-stationary dynamics in complex architectures, we introduce time-conditioned low-rank adaptation, a mechanism that conditions general neural backbones on time by dynamically modulating weights via low-rank updates. We provide theoretical analysis establishing approximation error bounds for cumulative-hazard evaluation. Comprehensive experiments across synthetic benchmarks, large-scale real-world tabular datasets, and high-dimensional medical imaging tasks demonstrate that QSurv achieves competitive predictive performance with advantages in instantaneous hazard function estimation, enabling more interpretable characterization of time-varying risk patterns.
Abstract: Diffusion models have been shown to implicitly generate visual content autoregressively in the frequency domain, where low-frequency components are generated earlier in the denoising process while high-frequency details emerge only in later timesteps. This structure offers a natural opportunity for efficient generation, as high-resolution computation on noise-dominated frequencies is largely redundant. We propose Spectral Progressive Diffusion, a general framework that progressively grows resolution along the denoising trajectory of pretrained diffusion models. To this end, we develop a spectral noise expansion mechanism and derive an optimal resolution schedule from the model's power spectrum. Our framework supports training-free acceleration and a novel fine-tuning recipe that further improves efficiency and quality. We demonstrate significant speedups on state-of-the-art pretrained image and video generation models while preserving visual quality.
Abstract: LLM-based automatic heuristic design has shown promise for generating executable heuristics for combinatorial optimization, but existing methods mainly rely on delayed endpoint performance. We propose a teacher-aware evolutionary framework that uses independently trained learned optimization policies as behavioral teachers. Instead of deploying or imitating the teacher, our method queries it on states visited by candidate heuristic programs and uses its action preferences as local feedback for evolution. The resulting search discovers static executable heuristics guided by both task performance and teacher-derived behavioral signals. Experiments on scheduling, routing, and graph optimization benchmarks show that our method improves over performance-driven LLM heuristic evolution baselines while requiring no neural inference at deployment. These results suggest that learned optimization policies can be repurposed as behavioral feedback sources for automatic heuristic discovery.
Abstract: Mixture-of-Experts (MoE) language models scale model ability with sparsely activated experts, making this architecture a standard recipe for modern large models. However, sparse activation does not remove the deployment burden of storing and serving all experts, and the available deployment budget can vary substantially across devices, users, and workloads. Existing MoE compression methods are still largely fixed-budget, typically optimizing one compressed endpoint at each chosen target budget. We study a different setting: converting a large pretrained MoE LLM into a nested family of deployable subnetworks across budgets. Our method first ranks expert FFN channels by their importance, then lets each expert learn a discrete action to prune its channels. By gradually increasing cost pressure, a single action-training run exports a series of action masks from high to low budgets, each of which identifies a reliable smaller subnetwork nested in the ranked base model. Moreover, we use a single recovery fine-tune at a mid pruning budget 40% to recover degraded model quality and transfer the recovered model to other unseen budgets. Overall, our framework surpasses recent MoE compression baselines. Specifically, on Qwen2-57B-A14B, our method retains ~99.8% of base performance while pruning 50% of routed expert parameters even without fine-tuning. For deployment, our pruned subnetworks deliver real memory reduction and throughput gains, and further support realtime online budget switching with kernel-level co-design.
PaperID: 4095, Poster
Abstract: Recent studies have highlighted the crucial role of distance metrics in noisy few-shot learning. However, existing approaches heavily rely on Euclidean or cosine measurements, failing to recognize the ``deceptive proximity'' where heterogeneous samples from different manifold peaks appear spuriously close in the flat ambient space. In this paper, we present Disentangled Dirichlet Geodesic Evaluation (DDGE), a novel manifold-aware framework that unifies orthogonal disentanglement and conformal integrals to efficiently evaluate complex topological spaces. Specifically, we decouple features into orthogonal semantic subspaces and leverage a prior-guided Dirichlet Process to expand discrete samples into a continuous, noise-purified semantic terrain. Upon this landscape, we measure the intrinsic geodesic distance via a density-aware conformal line integral to capture the authentic class topology. Extensive experiments demonstrate the state-of-the-art performance of DDGE on multiple benchmark datasets.
PaperID: 4096, Poster
Abstract: A central question in modern learning theory is what notion of compression, if any, explains why over-parameterized neural networks generalize. Existing answers measure complexity at different levels, ranging from weights to representations, and estimating the associated quantities either requires held-out data in principle or relies on mutual-information estimation in ultra-high dimensions. We approach the question through the minimum description length (MDL) principle and propose the prequential regret as a complementary signal that is computable from the training trajectory alone. We establish that, in the (nearly) interpolating regime, the expected prequential regret upper-bounds the description length of the trained parameter under the algorithm-induced distribution, and that this description length translates into a high-probability generalization bound in the PAC-MDL setting. Along the way, we show that the encoder-complexity correction recently added to the algorithm-level information bottleneck is upper-bounded by an MDL parameter description length, bringing the two views of compression onto a shared footing. Empirically, the prequential regret tracks the generalization error on CIFAR-10 and CIFAR-100, with Pearson correlations exceeding 0.9.
PaperID: 4097, Poster
Abstract: Aligning video generation models with human preference depends critically on the quality of the reward model that scores generated videos. Current video reward models, however, source their reward signal from a model's representational rather than generative capacity: they either repurpose vision-language models as pixel-space scorers, or train auxiliary scoring heads on top of frozen generator features. The generator's actual ability to predict velocities, denoise latents, and produce videos plays no role in scoring its own generations. We argue that this overlooks a structural property of offline preference learning. Preference optimization objective approximating an offline preference dataset mathematically targets a strictly defined KL-constrained optimal policy. By isolating the reward term from this analytical solution and mapping the exact generation log-likelihood to its flow-matching surrogate, the empirical reward emerges deterministically as the relative difference in continuous-time velocity-prediction errors. We formalize this equivalence into GenRM-Flow, a framework where a preference-tuned video generator inherently operates as a process-aware reward model directly in the noisy latent space. We demonstrate that this evaluation mechanism is structurally universal: distinct fine-tuning objectives, whether contrastive like DPO and IPO or non-contrastive like RWR, are merely alternative optimization pathways that share a unified empirical reward readout. Furthermore, the extracted process reward seamlessly substitutes external VLM scorers in Flow-GRPO. Operating purely on native latent outputs without auxiliary RL machinery, this unified generative reward drives measurable improvements in downstream video generation.
Authors:
Elies Gil-Fuster, Seongwook Shin, Sofiene Jerbi, Jens Eisert, Maximilian J KramerAbstract: Quantum kernel methods are among the leading candidates for achieving quantum advantage in supervised learning. A key bottleneck is the cost of inference: evaluating a weighted sum of N kernel values to additive precision \varepsilon, where \alpha is the vector of trained coefficients. The standard approach estimates each term independently via sampling, yielding a query complexity of \mathcalO(N\lVert\alpha\rVert_2^2/\varepsilon^2). In this work, we combine two independent improvements: estimating the kernel values via quantum amplitude estimation and collecting the sum under a single observable. We show that the improved approach achieves a query complexity of \mathcalO(\lVert\alpha\rVert_1/\varepsilon), removing the dependence on N from the query count and yielding a quadratic improvement in both \lVert\alpha\rVert_1 and \varepsilon. We prove a matching lower bound of \Omega(\lVert\alpha\rVert_1/\varepsilon), establishing query-optimality. Beyond query complexity, we also analyze how these improvements translate into gate costs and show that the query-optimal strategy is not always optimal in practice from the perspective of gate complexity. We identify a different strategy based on importance sampling, which yields a lower total gate count. We thus provide both a query-optimal algorithm and a practically-optimal choice of strategy depending on hardware capabilities, along with a complete landscape of intermediate methods to guide practitioners. All algorithms require only amplitude estimation as a subroutine and are thus natural candidates for early-fault-tolerant implementations.
Abstract: We study the contextual multi-armed bandit problem with a finite context space (a.k.a. subpopulations), where the learner recommends a best action for each context and is evaluated by context-weighted simple regret. Our guarantees are worst-case over the reward distributions, while remaining instance-dependent with respect to the context distribution vector \mathbf p. Akin to experimental design problems where the population of interest is fixed but the sampled subpopulation can be controlled, we allow the learner to actively choose which context to sample from. For a known \mathbf p, we characterize tight regret rates: passive sampling where contexts are randomly revealed achieves regret of order \sqrtn/T \lVert \mathbfp \rVert_1/2, whereas active sampling with allocation q_j \propto p_j^2/3 achieves the tight rate \sqrtn/T \lVert \mathbfp \rVert_2/3. The resulting improvement can be as large as \Theta(k^1/4), where k is the number of contexts. We further extend the analysis to budgeted active sampling, characterize the corresponding tight rate, and identify when a limited active budget suffices to recover the fully active rate. When \mathbf p is unknown, we propose the Explore-Explore-Then-Commit (EETC) algorithm, which optimally balances estimating the context distribution and the time to switch to active allocation, such that for large horizons, it matches the known-\mathbf p active rate up to constants. Experiments on synthetic and real-world data support our theoretical findings.
Authors: Adnan Armouti, Yixuan Gao, Rajalakshmi Nandakumar
Abstract: Solving novel view synthesis (NVS) for millimeter-wave (mmWave) radar requires a renderer that is physically faithful, complex-valued, and multi-viewpoint-tractable. No prior method achieves these three properties simultaneously. Differentiable Monte Carlo (MC) ray tracers implement the radar forward model directly with explicit material modeling and complex outputs, but do not scale to the multi-view optimization NVS demands. Optical-NVS ports of NeRF, hash grids, and 3D Gaussians train fast but discard phase and replace explicit material modeling with opaque learned features, restricting them to power-only range-azimuth (RA) magnitudes. We propose 3D Point Splatting (3DPS), the first differentiable point renderer for radar, derived directly from the standard solid-angle form of the radar equation. Each oriented 3D point carries an ITU-R P.2040 material model, evaluated in closed form, with the resulting complex phasor splatted into range bins through a precomputed point spread function (PSF). The complex-valued output makes the renderer product-agnostic. The same optimized scene yields analog-to-digital converter (ADC), complex range profile (CRP), and RA outputs through standard fast Fourier transform (FFT) pipelines without retraining for each format. On six outdoor ColoRadar scenes, 3DPS reaches 0.587 mean Pearson correlation on held-out RA images. This is between 1.7× and 5.2× the three optical-NVS baselines (RadarSplat, Radar Fields, DART). Training takes approximately 3 minutes per scene on a single RTX 4090.
PaperID: 4101, Poster
Abstract: Reliable use of large language models (LLMs) requires uncertainty estimates that separate ambiguity intrinsic to a prompt from genuine knowledge gaps. Bayesian and deep-ensemble methods deliver this aleatoric/epistemic decomposition in principle, but are computationally prohibitive at LLM scale and require white-box access. Existing black-box methods, in contrast, summarize sample variability into a single that bridges this gap. Stochastic decoding of a fixed prompt forms the inner ensemble; meaning-preserving semantic perturbations of the prompt form the outer ensemble. Variability is measured in a continuous embedding space, and the law of total covariance gives an exact AU/EU split that requires only black-box sampling access. On top of this estimator, we contribute three theoretical results that ground the framework: (i) a finite-sample bias identity for nested Monte Carlo with a closed-form bias-corrected epistemic estimator; (ii) a bounding the gap between perturbation-population AU/EU and Bayesian targets through three local conditions that we estimate as identifying a class of failure modes that any consistency-only black-box method provably misses while perturbation-based EU detects. Across five short-form QA benchmarks and four instruction-tuned LLMs spanning two families and two scales, the estimator matches or surpasses strong white-box baselines; controlled paired interventions on AmbigQA and Natural Questions confirm the AU/EU split tracks distinct, actionable sources of failure; and a long-form pilot on the NQ long-answer split shows the framework extends to claim-level factuality without retraining. Across five short-form QA benchmarks and four instruction-tuned LLMs spanning two families and two scales, the estimator matches or surpasses strong white-box baselines; controlled paired interventions on AmbigQA and Natural Questions confirm the AU/EU split tracks distinct, actionable sources of failure; and a long-form pilot on the NQ long-answer split shows the framework extends to claim-level factuality without retraining.
PaperID: 4102, Poster
Abstract: Diffusion-based visual generative models are powerful, but understanding their generalization behavior remains challenging. Current evaluations rely on human preference scores for text-to-image models and distribution-level metrics such as Frechet Inception Distance (FID) and Inception Score (IS) for class-conditional models, but these do not test whether models learn basic visual skills or largely match training patterns. To overcome this gap, we introduce the Visual Generative Lab (VGL), a controlled experimental framework for understanding how diffusion models learn and generalize basic visual skills, including size, position, rotation, count, color, shape, and their composition. By generating synthetic training data for each skill from explicit rules, VGL enables precise measurement of generalization using task-specific, rule-based metrics. Training models from scratch on these data reveals three consistent patterns. (1) Models extrapolate rotation better than size, skill, or count on far-out-of-range queries; healthy rotation seeds match most extrapolation angles within a few degrees, although a non-trivial fraction of seeds collapse near the wraparound, while skill and position saturate at small distances outside training regardless of seed. (2) For compositional generalization, coverage of basic visual skill combinations matters more than dataset size. (3) Basic visual skills are learned jointly, so an out-of-distribution request in one skill can degrade the others even when they remain in range.
Abstract: We study generative modeling of graphs with recurring subgraph motifs. We propose Flowette, a continuous flow matching framework that employs a graph neural network-based transformer to learn a velocity field over graph representations with node and edge attributes. Our model promotes topology-aware alignment through optimal transport-based coupling and encourages global structural coherence through regularisation. To incorporate domain-driven structural priors, we introduce graphettes, a new probabilistic family of graph structure models that generalize graphons via controlled structural edits for motifs such as rings, stars, and trees. We theoretically analyze the coupling, invariance, and structural properties of the framework, evaluate it on synthetic and molecular benchmarks, and isolate the contributions of the structural prior, the optimal-transport coupling, and the regularisation terms through controlled ablations. Flowette achieves competitive performance overall, attaining state-of-the-art results on several metrics across multiple benchmarks, highlighting the effectiveness of combining structural priors with flow-based training for modeling complex graph distributions.
PaperID: 4104, Poster
Abstract: Unlearnable Examples (UEs) have emerged as a promising data protection strategy against unauthorized model training, which adds imperceptible perturbations into data to degrade generalization. Existing studies predominantly assume a fully trainable target model, where unlearnability is achieved by disrupting clean feature representations. However, this assumption grows increasingly unrealistic with the prevalence of pretrained models, where unauthorized users can freeze the backbone and perform linear probing, preserving discriminative representations and thereby invalidating the core mechanism of previous UEs. In this paper, we investigate this practical yet challenging linear probing setting and reveal a fundamental vulnerability. To address this, we propose DIVERT (Decision-boundary dIversion Via sEmantic-decoupled oRthogonal Targets), a novel framework that shifts the design principle from feature disruption to explicit decision boundary diversion. Our intuition is that projecting perturbed samples into a subspace orthogonal to the semantic manifold induces spurious linear separability, thereby steering the decision boundary away from original prototypes. Building on this principle, DIVERT first constructs synthetic class anchors within the semantic-decoupled subspace, then employs a cross-attention mechanism to jointly optimize these anchors while pushing perturbed samples toward them to establish spurious separability. Subsequently, it incorporates a surrogate linear head to simulate the linear probing procedure, further refining perturbations to actively divert the decision boundary. Extensive experiments demonstrate that DIVERT establishes effective unlearnability across diverse datasets and backbones.
PaperID: 4105, Poster
Abstract: Zeroth-order (ZO) optimization enables memory-efficient large language model (LLM) fine-tuning, with recent methods employing curvature-aware strategies. However, in the massive LLM parameter space, these methods inherently suffer from exploding gradient variance and prohibitive costs. Conversely, confining updates to a low-dimensional subspace reduces variance but discards critical high-dimensional information, degrading performance. This creates a fundamental dilemma: full-space optimization is informative but high-variance, while subspace optimization is low-variance but blind. To break this trade-off, we propose \Engine (Endogenous Variational MultiScale). Inspired by computational mechanics, we adapt static Variational MultiScale theory to dynamic optimization, treating low- and high-dimensional spaces as resolved coarse and unresolved fine scales. Through a multigrid-style V-cycle, \Engine\ alternates between spaces, dynamically probing fine-grained information in the high-dimensional space and injecting it to correct the efficient low-dimensional sampling process. Evaluations on standard LLM benchmarks show that \Engine\ achieves higher accuracy and faster convergence than existing ZO methods, achieving up to a \mathbf2.61× speed-up in forward passes with zero additional peak memory overhead. Code is available at \urlhttps://anonymous.4open.science/r/ENGINE-D0A4.
PaperID: 4106, Poster
Abstract: We study an endogenous nonstationary stochastic bandit problem with latent linear dynamics, where actions affect both immediate rewards and the future evolution of an unobserved latent state. Rewards are bilinear in the current action and latent state, inducing history-dependent rewards and a nontrivial long-horizon planning problem. The existing explore-then-commit approach (Choi et al., 2026) achieves \tilde\mathcalO(T^2/3) regret by uniformly exploring to estimate the latent dynamics and then committing to an optimized open-loop action sequence. We show that this rate can be improved via adaptive block-level optimism. Our key step is a cyclic approximation: under stable dynamics, the infinite-memory reward process can be truncated, and the open-loop benchmark can be approximated by optimizing a finite-memory block-level proxy. Building on this reduction, we propose a UCB-based block algorithm that maintains confidence sets for the truncated dynamics parameters and selects blocks optimistically. We prove a regret bound of order \tilde\mathcalO(\sqrt T), significantly improving over the previous \tilde\mathcalO(T^2/3) guarantee for the same model. To the best of our knowledge, this is the first \tilde\mathcalO(\sqrt T) regret guarantee for latent linear-dynamics bandits with bilinear reward observations and an open-loop action-sequence benchmark.
PaperID: 4107, Poster
Abstract: When humans speak, gestures often realize communicative functions (e.g., emphasizing key content, indicating referents, or depicting concepts). Most co-speech gesture generation methods, however, condition primarily on low-level cues (speech audio or transcripts) and do not explicitly represent these higher-level functions, leading to motions that synchronize with speech but remain under-specified in communicative meaning. To address this gap, we introduce Reasoning Gestures, a reasoning-grounded framework that conditions generation on interpretable communicative-function descriptions. We curate InG: Interpreted Gestures by augmenting BEAT-2 with sentence-level \emphVLM-inferred gesture-function annotations that describe the discourse role of each gesture segment in context. These labels are \emphobserver-centric interpretations produced via constrained VLM reasoning rather than claims about speakers’ private mental states. We further propose a Reasoning Gesture Motion Tokenizer that injects function signals into tokenized motion representations, enabling function-aware synthesis that is temporally aligned and communicatively grounded. Experiments on BEAT-2 show consistent gains in semantic consistency and overall quality, establishing new state-of-the-art.
Authors: Lyra Zhornyak, Eric Forgoston, M. A Hsieh
Abstract: We present the first method to directly use a learned continuous Lagrangian to forecast the dynamics of systems governed by partial differential equations, exploiting the inherent conservative structure to achieve stable long-range predictions. We develop an optimization-based integrator that minimizes the squared Euler--Lagrange residual via a mesh-free near-symplectic construction on local space-time patches. Different from integrators for analytical models, integrators for learned models should decouple model error (phase error) from integration error (conservation error). By relying on optimization rather than time-stepping, we bypass the global coupling inherent to fixed discretizations, which slows time- and space-stepping and complicates learning. Our method scales linearly with domain size via Jacobi iteration, and places no structural requirements on the learned network, allowing it to be coupled with existing physics-guided machine learning (ML) methods. We validate our approach on a learned representation of a double pendulum, a one-dimensional wave equation, and a two-dimensional wave equation. Our method achieves error comparable to classical symplectic methods while generalizing to spatially varying dynamics and arbitrary boundary conditions without retraining.
PaperID: 4109, Poster
Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language tasks, yet their internal visual representations remain difficult to interpret. Sparse Autoencoders (SAEs) provide a scalable way to decompose dense model activations into sparse, interpretable features. However, existing SAE architectures primarily recover flat feature dictionaries and are less suited for explicit multi-level concept organization. In this paper, we introduce cascaded sparse autoencoders (CSAEs) for learning hierarchical visual concepts in MLLMs. Rather than nesting or stacking SAE sparse activation codes, CSAEs train a second-level SAE directly on the decoder weights of the first-level SAE, treating learned low-level feature directions as inputs for higher-level abstraction. This design enables CSAEs to learn “concepts of concepts” while avoiding drawbacks from the shared-prefix coupling of nesting, Matryoshka-style hierarchies and the bottlenecks of naively stacked SAEs. Experiments across Qwen3-VL, Gemma-3, and LLaVA on multiple visual datasets show that CSAEs improve interpretability in terms of hierarchical concept coherence over state-of-the-art SAE baselines. Results on concept steering further demonstrate that the learned concept groups support effective group-level interventions in MLLM outputs.
Authors:
Zijun Li, Yimin Zhou, Jia Sun, Honglie Wang, Pengcheng Wei, junlong wu, Yongrui Heng, Jiyuan Wang, Huan Ouyang, Boheng Zhang, Huaiqing Wang, Dewen Fan, Qianqian Gan, Fan Yang, Tingting GaoAbstract: Diffusion-based generative AI has achieved remarkable success in e-commerce applications such as virtual try-on, poster generation, and product background synthesis. However, when making online purchasing decisions for apparel, consumers also desire the freedom to examine specific detail regions of interest, such as collars, cuffs, and fabric textures, yet existing methods have not explicitly studied this setting. We therefore formalize a new, non-template task: Fashion Detail Generation with focus conditioning, and release , the first dataset and benchmark comprising 40K+ human-verified reference-detail pairs across 41 different categories. This task poses a unique semantic gap challenge: the model must bridge the correspondence between a focus marker on a product reference image and a photorealistic close-up view of the indicated region, while faithfully preserving the garment's identity, without any precise prompt. To bridge this gap, we propose Cross-modal Feature Alignment Distillation ( ), which leverages a fine-tuned DINOv3 teacher to align both branches of a Multimodal Diffusion Transformer in a shared semantic space via dual-branch distillation. To further improve consistency between generated details and reference images, we introduce a consistency reward model that jointly scores image pairs along three quality axes and optimizes generation via reinforcement learning. Experiments show that our model significantly outperforms all state-of-the-art opensource methods across all metrics and human evaluations.
PaperID: 4111, Poster
Abstract: High-Dimensional Low-Sample Size (HDLSS) tabular domains (e.g., omics) are characterized by n \ll m, where n = number of samples, and m = number of features. Such domains often exhibit strong local correlation groups, sparse cross-group dependencies, heavy-tailed non-Gaussian marginals, heteroscedastic noise, and structured missingness, making direct density learning in \mathbbR^m ill-conditioned since n \ll m. We propose BSTabDiff, a block-subunit generative framework that partitions the m observed features into M latent blocks (M \ll m) and generates each block via a shared low-dimensional subunit variable, concentrating global dependence learning in the compact block-latent space \mathbbR^M while decoding to the full feature space with copula-driven dependence, flexible per-feature marginals, and explicit missingness mechanisms. BSTabDiff can induce a data-driven feature ordering and block construction before latent partitioning, allowing the model to better align with local dependency structure when such organization is beneficial. BSTabDiff supports modern deep priors on block latents, including diffusion and normalizing flows, enabling stable synthesis and controllable benchmark generation in the HDLSS regime. Empirically, BSTabDiff produces more realistic and more stable high-dimensional synthetic data when compared with unstructured tabular generators on HDLSS data.
PaperID: 4112, Poster
Abstract: Chinese Spelling Correction (CSC) is hard precisely because a single misspelling can originate from any of four loosely coupled dimensions (glyph, phonetic, syntactic, or semantic confusion), yet existing systems treat the four as a flat fusion target. Multimodal pre-training and fixed-weight gating let evidence from each dimension co-exist, but never arbitrate: when a phonetic-plausible candidate contradicts a syntactic-plausible one, the model averages them, and learning effort is spent uniformly on cases where every dimension already agrees. Our diagnosis is that the bottleneck is no longer representation but \emphdecision: the cases where dimensions disagree are exactly the cases worth specializing for, and dimension-specific teachers expose this disagreement as a usable signal rather than as noise to be smoothed away. We instantiate this view as SEMTD (Self-Evolving Multi-Teacher Distillation), a 4B-parameter student trained from four single-dimension experts (glyph / phonetic / syntactic / semantic) under three coupled objectives: multi-teacher distillation with input-dependent expert selection, self-evolving learning that turns disagreement into a composite reward, and reflective path optimization that replays high-disagreement decisions. Across four benchmarks the student reaches Cor-F1 of 83.8 on SIGHAN15, 62.3 on LEMON (7-domain average), 98.2 on ECSpell (3-domain average), and 73.0 on CSCD-NS, matching or surpassing 14B LLM-based correctors at 3.5× fewer parameters. The takeaway is methodological: in CSC, treating teacher disagreement as the supervision target, rather than a residual to be averaged out, reopens headroom that flat fusion has saturated.
PaperID: 4113, Poster
Abstract: Self-supervised Vision Transformers (ViTs) like DINO encode rich semantics, enabling training-free segmentation of objects from the background and discovery of their parts via feature clustering. This paper focuses on part discovery: partitioning a foreground object into semantic parts. Existing approaches produce imprecise part boundaries. We trace this to the rigid ViT patch grid used in training DINO, where patches straddling boundaries contain pixels from multiple parts. While DINO's invariance-based training could absorb this pixel-level mixing into clean features, experiments with over 26K foreground patches show it does not: boundary-overlapping patches produce features with higher entropy and lower assignment confidence, blurring predicted part boundaries. Rather than make the massive effort of retraining DINO over geometrically coherent regions, an easier approach is to take boundaries directly from a boundary detector, which supplies the geometric information DINO's features lack. Motivated by this, we introduce Part-DINO, a training-free framework that combines boundaries from low-level image segmentation with DINO's semantic features. We use the boundaries to partition the image into regions, pool DINO features within each region, and group them to discover parts. Without any training, Part-DINO produces structurally coherent part assignments, outperforming training-free baselines across four standard benchmarks (CUB, CelebA, PartImageNet-OOD, PartImageNet-Seg) and remaining competitive with weakly-supervised and unsupervised methods.
Authors: Johann M Christensen, Elena Hoemann, Frank Köster, Sven Hallerbach
Abstract: Artificial Intelligence (AI) has been on the rise in many domains, including numerous safety-critical applications. However, for complex systems in the real world, defining the underlying environmental conditions in which the AI-based system must operate---the Operational Design Domain (ODD)---is extremely challenging. This often results in an incomplete description of the ODD, which contrasts with the requirements of many domains for certifying AI-based systems. Traditionally, the ODD is created in the early stages of the development process, drawing on sophisticated expert knowledge and related standards. This paper presents a novel Safety-by-Design method to a posteriori define the ODD from previously collected data using a multi-dimensional kernel-based representation. This approach is validated through both Monte Carlo methods and a real-world aviation use case for a future collision-avoidance system. Moreover, by defining under what conditions two ODDs are similar, the paper shows that the data-driven ODD can produce a dataset similar to the original, hidden ODD. Deriving the novel, Safety-by-Design, deterministic kernel-based affinity representation of ODDs is fully automated via a bounded, order-independent algorithm. Utilizing the proposed ODD representation enables future certification of data-driven, safety-critical AI-based systems.
PaperID: 4115, Poster
Abstract: Most temporal graph learning methods reduce future prediction to discriminative inference over past interactions, focusing on edges or node labels for a largely persistent set of nodes. This view breaks down when new nodes and edges appear, since the future graph becomes a new object with changing size, composition, and structure. We address this gap by framing temporal graph prediction as an inverse topology problem. Instead of predicting edges directly, we first predict a multiscale topological descriptor of the future graph and then reconstruct a plausible future snapshot that realizes this descriptor under inductive constraints. This approach makes global structure and node churn explicit prediction targets and produces forecasted graphs on which downstream tasks can be evaluated without retraining. Across 14 temporal graph datasets, we evaluate TopoGED on node, link and graph property prediction and compare it against state-of-the-art temporal graph models. TopoGED achieves a significant improvement in node-level forecasting accuracy over the strongest baseline, with a macro average of 0.49 versus 0.05. It also outperforms baselines in 60% of graph-structure metric evaluations and yields a substantial increase in macro-average edge prediction metrics, from near-zero to 0.12. Our results show that topology-guided graph forecasting can predict inductive future snapshots whose structure supports multiple downstream evaluations.
PaperID: 4116, Poster
Authors: Enshi Zhang, Christian Poellabauer
Abstract: Speech is an increasingly promising biomarker for neurodegenerative disorders due to its non-invasive nature, low cost, and suitability for frequent, longitudinal monitoring. However, existing clinical speech datasets are typically small, sparse, and irregularly sampled, limiting the ability to model continuous disease progression and develop robust biomarkers. We propose L-Flow, a progression-aware conditional speech generation framework for longitudinal speech trajectory completion. Given a patient’s speech recordings, timestamps, and clinical progression labels, L-Flow learns a query-specific progression representation to synthesize speech at intermediate time points within the observed recording span. Experiments on two real-world longitudinal speech datasets demonstrate that L-Flow generates synthetic speech with stronger progression consistency than eight baseline augmentation methods. In addition, incorporating L-Flow–augmented data improves downstream regression performance to predict clinical severity, highlighting its utility for speech-based biomarker development.
PaperID: 4117, Poster
Abstract: Visual understanding is inherently hierarchical as high-level semantic concepts are composed of fine-grained, spatially localized components. Yet existing interpretability methods typically operate at a single scale, either local or global, leaving a critical gap in understanding the compositional structure between them. Recovering this structure without supervision is particularly challenging, as the correspondence between local and global concepts is latent, varies across samples, and may appear only through inconsistent or partial patterns. To address this challenge, we introduce Local-Global Sparse Autoencoders (LG-SAE), an unsupervised framework for multiscale interpretability in vision models. The core innovation of LG-SAE is a learned compositional bridge that explicitly captures the relationships between local and global concepts. By jointly decomposing patch-level and image-level embeddings, LG-SAE induces a sparse interpretable graph that reveals how each high-level semantic concept is supported by a small set of spatially grounded components. Building on this representation, we present a unified framework for discovering local-global concept relationships, constructing multiscale Concept Bottleneck Models, spatially aware concept naming, and model steering through localized concept edits. Extensive evaluations show that LG-SAE: (i) captures meaningful cross-scale feature relations; (ii) enables high-quality spatial grounding without compromising downstream accuracy; and (iii) provides interpretable model explanations, further validated by a user study.
PaperID: 4118, Poster
Authors: Wenyu Bo, Marina Meila
Abstract: Diffusion Maps (DM) is a well studied and popular non-linear dimension reduction algorithm. Surprisingly, while the convergence properties of DM are well understood, the \em geometric properties of the embedding, such as reach and smoothness have not yet been studied. Under a set of standard assumptions on a family of submanifolds \subset \mathbbR^D, we derive a series of geometric properties that are preserved by DM, including almost uniform density, finite polynomial approximation and reach. Leveraging these properties, we establish rigorous bounds on the embedding errors introduced by the DM algorithm, of the order (\frac\log nn)^\frac18d+16. These results offer a solid theoretical foundation for understanding the performance and reliability of DM in practical applications.
PaperID: 4119, Poster
Abstract: Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable feature directions using sparse autoencoders (SAEs) do neither readily isolate which features are changed by multimodal training, nor are they directly useful for targeted control. We introduce MMDiff, a multimodal model-diffing framework that trains multimodal SAEs and turns them into feature-level interfaces for discovering and controlling multimodal behavior. MMDiff supports three uses: (i) feature isolation, by diffing a base-LM SAE against its multimodal-adapted counterpart to identify features altered by multimodal training; (ii) task-specific feature detection, via per-token contrastive firing analysis that isolates causal features; and (iii) feature-level control, by causally removing or steering the discovered feature directions. We train multimodal SAEs for two MLLM families, LLaVA-MORE and PaliGemma 2, and evaluate on visual-spatial understanding, multimodal safety, and OCR. MMDiff discovers sparse, causally specific features whose removal selectively degrades target behaviors by an average of 12% on spatial tasks and 17% on OCR, and reduces attack success rate by 24% on multimodal safety attacks, with no impact on VQA performance. Steering these features improves spatial and OCR accuracy by +3.6% and +1.5% on average over a standard single-layer steering baseline. These results show that multimodal SAEs can serve not only as interpretability tools, but as mechanisms for auditing, steering, and controlling MLLMs behavior toward safer and more capable generations.
PaperID: 4120, Poster
Abstract: Many real-world high-dimensional time series exhibit long-memory, but Gaussian graphical model testing in this regime remains understudied. We develop a direct, data-adaptive test statistic for assessing conditional independence in the graph structure of stationary Gaussian time series. We establish a finite-sample, Berry--Esseen type Gaussian approximation bound for the statistic, which applies to both short-memory and long-memory time series. The testing procedure is fully data-adaptive using block bootstrap method, on which we provide a finite-sample validity result including in the ultra-high-dimensional scenario, and can be extended to comparing graphical structures in two-sample tests. We also develop a consistency-empowered correction to the statistic and show that such tests attain asymptotic consistency in both size and power. Our proposed method is applied to a real-world fMRI data to understand functional connectivities within brain in different periods.
Abstract: Large Language Models (LLMs) have achieved remarkable success in complex reasoning tasks through Chain-of-Thought (CoT) prompting. However, these models often exhibit "computational overthinking," generating redundant reasoning steps that increase latency and cost without improving accuracy. Recent studies suggest that CoT trajectories can be significantly pruned, yet existing methods often rely on forcing a static thinking budget, heuristic filtering, sub-optimal early exit via classification, or expensive re-training. In this paper, we introduce OS-Pruner, a lightweight plug-in framework that formulates chain-of-thought pruning as an optimal stopping problem. Given a reasoning prefix, OS-Pruner learns whether further reasoning is worth its token cost by optimizing an explicit utility that trades off final-answer accuracy against generated length. Our novel formulation enables the model to dynamically assess the sufficient point of termination for a reasoning chain. OS-Pruner is designed to be lightweight during both training and inference, and to provide users with fine-grained control over the reasoning-effort vs. accuracy trade-off. On diverse reasoning benchmarks and base models, OS-Pruner achieves 20-60% reduction in generation length with minimal accuracy sacrifice.
Abstract: Scaling critic capacity is a promising direction for enhancing off-policy reinforcement learning (RL). However, larger critics are prone to overfitting and unstable in replay-buffer-based bootstrap training. This paper leverages Low-Rank Adaptation (LoRA) as a structural-sparsity regularizer for off-policy critics. Our approach freezes randomly initialized base matrices and solely optimizes low-rank adapters, thereby constraining critic updates to a low-dimensional subspace. Built on top of SimbaV2, we further develop a LoRA formulation, compatible with SimbaV2, that preserves its hyperspherical normalization geometry under frozen-backbone training. We evaluate our method with SAC and FastTD3 on DeepMind Control locomotion and IsaacLab robotics benchmarks. LoRA consistently achieves lower critic loss during training and stronger policy performance. Extensive experiments demonstrate that adaptive low-rank updates provide a simple, scalable, and effective structural regularization for critic learning in off-policy RL.
Abstract: We consider the problem of regularized best-response max-regret minimization in online RLHF under general preferences and bandit feedback. While various regularizers are utilized to robustify alignment, known polylogarithmic regret guarantees remain heavily specific to KL. To investigate whether such fast rates extend beyond KL, we adopt the Generalized Bilinear Preference Model (GBPM)—capturing intransitive preferences over d-dimensional item-wise features via a rank-2r skew-symmetric matrix—to isolate the impact of generic regularization. Crucially, under GBPM, we prove that the dual gap of any greedy policy is bounded by the squared estimation error, derived using only strong convexity and skew-symmetry. Under a feature coverage assumption, we establish polylogarithmic \tilde\mathcalO(\eta d^4 (\log T)^2 \wedge d^2 \sqrtT) regret with Greedy Sampling and \mathrmpoly(d)-free \tilde\mathcalO(\sqrt\eta r T \wedge r^1/3 T^2/3) regret with Explore-Then-Commit, where \eta^-1 is the regularization coefficient and T is the time horizon. This demonstrates that ``fast'' regrets are not KL-specific, but rather a fundamental consequence of generic strongly convex geometry.
PaperID: 4124, Poster
Abstract: Understanding mouse behavior at scale is a foundational problem in neuroscience, with two core tasks: classifying frames into known behaviors and segmenting recordings into recurring patterns without labels. Pose-based methods are the standard approach but face fundamental limitations: pose coordinates capture only skeletal information, discard environmental context, and degrade under frequent occlusions in social scenarios. Mouse videos can provide the visual context that pose lacks, yet not all frames benefit equally, and adding visual features uniformly to every frame is suboptimal. We find that a pose-based motion encoder, trained with auxiliary supervision from videos, can identify where it struggles: with a discrete representation, the encoder learns a codebook of recurring patterns, and frames whose motion does not fit this codebook produce large quantization residuals. We show that these high-residual frames are where visual context tends to help most: residual-guided frame selection consistently outperforms uniform selection across budgets, and often exceeds dense visual processing. We therefore propose ResiFuse, a framework that pretrains a motion encoder on pose with an auxiliary cross-modal loss and integrates vision context only on high-residual frames at fine-tuning time. We evaluate ResiFuse on three mouse behavior datasets and show improvements over pose- and video-based baselines on most tasks, with the largest gains on the most socially complex dataset.
Abstract: Rendering accurate text remains difficult for image generation and editing models, especially when the target contains long, complex, and densely arranged text or rare characters. Existing approaches either improve native text rendering through stronger backbones and data-centric training without explicit glyph priors, or incorporate glyph priors through specialized designs that remain insufficiently accurate and robust under challenging scenarios. We introduce GlyphAnchor, a novel text-rendering enhancement method for both text-to-image and image-editing diffusion transformer models. GlyphAnchor enhances the backbone with lightweight glyph patch conditions whose positions are anchored to the target image through the model's native positional encoding. We train this capability with staged supervised finetuning and further refine it with text-aware post-training to improve robustness. We also introduce InfoTextBench, a benchmark for evaluating text-rich visual text rendering in both generation and editing settings. Experiments across multiple backbones and benchmarks, including long, complex, and densely arranged text and rare character scenarios, show that GlyphAnchor consistently improves text fidelity while preserving overall image quality.
Abstract: Large Language Models (LLMs) have achieved remarkable progress across a wide range of tasks, but remain vulnerable to safety risks such as harmful content generation and jailbreak attacks. Existing safety techniques---including external guardrails, inference-time guidance, and post-training alignment---each face limitations in balancing safety, utility, and controllability. In this work, we propose UpSafe℃, a unified framework for enhancing LLM safety through safety-aware upcycling. Our approach first identifies safety-critical layers and upcycles them into a sparse Mixture-of-Experts (MoE) structure, where the router acts as a soft guardrail that selectively activates original MLPs and added safety experts. We further introduce a two-stage SFT strategy to strengthen safety discrimination while preserving general capabilities. To enable flexible control at inference time, we introduce a safety temperature mechanism, allowing dynamic adjustment of the trade-off between safety and utility. Experiments across multiple benchmarks, base model, and model scales demonstrate that UpSafe℃ achieves robust safety improvements against harmful and jailbreak inputs, while maintaining competitive performance on general tasks. Moreover, analysis shows that safety temperature provides fine-grained inference-time control that achieves controllable trade-off between utility and safety. Our results highlight a new direction for LLM safety: moving from static alignment toward dynamic, modular, and inference-aware control. Code: https://anonymous.4open.science/r/UpCycle-70B3.
PaperID: 4127, Poster
Abstract: Physical-domain adversarial attacks present critical security threats to diverse AI systems, notably autonomous driving. Generalization is an inherent requirement and a primary obstacle in physical attacks, manifesting as the ability of adversarial patterns to persist across varying environmental configurations and to transfer to unseen victim models. While prior methods demonstrate generalization across specific models, their performance drops significantly on open-vocabulary foundation models, with some approaches becoming entirely ineffective. This lack of generalization stems from the failure of adversarial camouflage to construct a cohesive pseudo-feature across perspectives, lacking the stable semantic representation characteristic of natural objects. Such a deficiency arises because existing attack paradigms rely on independent view-level optimization and non-targeted objectives, which lead to semantically scattered features and inconsistent optimization directions. To address these limitations, we propose Generalizable Physical Adversarial Camouflage (G-PAC), a framework that transitions from view-specific suppression to coordinated object-level representation learning. G-PAC maintains a momentum-updated global pseudo-feature center as a stable adversarial anchor to guide optimization across diverse configurations. By leveraging the feature space of a self-supervised foundation model as a semantic prior, we introduce a contrastive regularization term to encourage feature aggregation while ensuring adversarial potency. Comprehensive digital and physical evaluations demonstrate the effectiveness of G-PAC. Notably, it outperforms the strongest baseline by an average AP@0.5 margin of 0.10 across seven diverse detector architectures in digital settings, and further extends this margin to 0.12 in real-world physical tests. Furthermore, G-PAC exhibits profound black-box transferability against highly resilient open-vocabulary foundation models, inducing an additional absolute AP@0.5 reduction of 0.30 on GLIP.
PaperID: 4128, Poster
Abstract: Neural surrogates for Partial Differential Equations (PDEs) are usually trained on synthetic data from numerical simulators. As models grow and scaling laws demand larger datasets, the cost of data generation often outweighs the training cost and becomes the bottleneck in the scaling quest. Unlike data in other domains, however, each PDE sample has a tunable per-sample cost that depends on its numerical fidelity, with the resulting cost-accuracy trade-off governed by known convergence rates. We exploit this property by formalizing neural surrogate training as a function of both data quantity and fidelity under a fixed data generation compute budget, treating fidelity as a fine-grained discrete spectrum rather than a binary distinction. Building on this, we show that multi-fidelity training enables Pareto-optimal error scaling against data generation compute. By using simulations at lower fidelities as an additional training signal, our approach consistently outperforms training on the highest fidelity at matched data generation budgets. We validate this across six PDE problems spanning structured grids and unstructured meshes, two distinct error sources (discretization and iterative-solver error), and four surrogate architectures.
Authors: Konstantin Krestnikov
Abstract: When transformers train on contradictory data, the same problem with both correct and incorrect solutions, which answer do they prefer? We hypothesize that next-token prediction, as a compression process, favors whichever answer cluster has lower description length; truth benefits only when errors lack internal structure. We test this by training transformers (3.5M-1B parameters) from scratch on controlled corpora, systematically varying the structure of errors. We find that (a) when errors are random, models develop a correctness preference scaling from 65% to 85% with model size; (b) when errors follow a single coherent alternative rule, this preference vanishes (~45-51%); (c) two competing wrong rules suffice to restore it (47% to 78%). The pattern reproduces on Wikipedia paragraphs with entity substitution (71% vs 46%) and at 1B scale on a mixed natural-text corpus (77% vs 47%). These results are consistent with the hypothesis that, in controlled contradictory corpora, model preference tracks the relative compressibility of competing answer systems rather than truth per se.
PaperID: 4130, Poster
Abstract: Active perception is crucial for robots to interact with unstructured scenes, where a closed perception-action loop is required. In this work, we introduce ActiveZero, an end-to-end framework that formulates active perception as information-driven spatio-temporal memory management. Leveraging memory as a universal information interface for conveying perception and action, it maintains a visual memory bank, actively expands it via purposeful exploration for missing information, retrieves relevant evidence for execution, and filters memory online for efficiency. This formulation not only enjoys long-context reasoning ability but also enables multi-level supervision for heterogeneous data, going beyond prior paradigms. To support it, we construct ActiveMem, an active perception dataset of 4M episodes with camera actions (300× prior) across 49 scenes for long-horizon tasks (up to 4 subtasks). In addition, we present ActiveBench, the first memory-centric benchmark suite for active perception, covering VQA and simulated manipulation. Experiments show that ActiveZero achieves SOTA on existing active-perception-related benchmarks and ActiveBench. It also outperforms all baselines on challenging real-world tasks by a large margin, even surpassing \pi_0.5 by 34.16%. Notably, ActiveZero exhibits diverse active perception strategies (e.g., search, track, interact) and handles long-horizon tasks in unstructured scenes with a single unified model.
Authors: Joshua C Yang, Maurice Flechtner, Damian Dailisan, Michiel Bakker
Abstract: LLM-based AI agents are increasingly used to simulate deliberative social interactions, including negotiation, conflict resolution, and multi-turn opinion exchanges. In these settings, it is important to know not only what an agent says, but how its stance changes after receiving new information. Current systems often make this difficult: agents may drift from assigned roles, echo an interaction partner, remain fixed despite relevant evidence, or appear to change because the prompt or retrieved context has changed. We introduce the Belief Engine (BE), a framework that uses ``belief'' in a Bayesian modelling sense: a maintained evidential state over a proposition, exposed as scalar stance. BE extracts structured evidence, stores active and archived argument records, and updates the belief state through a Bayesian log-odds rule with two interpretable controls: evidence uptake u and prior anchoring a. Across multiple LLM models, sweeps of u and a demonstrate stable control over stance dynamics: higher uptake makes agents more responsive to new evidence, while stronger anchoring preserves the initial stance. On DEBATE, a human deliberation dataset with pre/post opinions, BE is most accurate when participants move in the direction of extracted evidence, showing its strength in evidence-following deliberative contexts; when all participants are evaluated together, gains are more modest because many people remain stable or move for reasons not captured by the extracted evidence stream. BE therefore provides a configurable belief-update layer for evidence-grounded deliberation: LLM agents can be made open-minded, anchored, or evidence-sensitive by construction, while their stance trajectories and memory can be inspected, compared, and calibrated against human opinion change.
Abstract: Block-diffusion drafters have recently shown strong potential for speculative decoding by predicting multiple future-token distributions in a single forward pass. To exploit these parallel predictions, tree-based verification can check multiple plausible continuations instead of committing to a single greedy draft path, often increasing the number of accepted tokens per decoding cycle. However, existing tree-based approaches typically rely on fixed tree topologies or static verification budgets, even though the best budget depends on the local draft distribution, context length, target model, and hardware runtime. We propose a system-aware adaptive tree framework for block-diffusion speculative decoding. At each decoding cycle, our method builds a tree toward high-probability draft paths and selects the verification budget that maximizes an online speedup estimate. This estimate combines a drafter-based acceptance surrogate with an analytical verifier-latency model calibrated from observed runtimes. The framework requires no additional training of the drafter or target model and preserves the target model's decoding rule. Across target models, benchmarks, decoding temperatures, and GPU platforms, our method achieves up to a 6.93× wall-clock speedup over standard autoregressive decoding. This corresponds to approximately 40% higher speedup than Dflash, the strongest existing block-diffusion drafter, and is obtained without per-setting budget tuning.
Abstract: Diffusion and flow-matching models are typically constrained to a single manifold, where the noise prior, intermediate trajectory, and prediction target share the same dimensionality. While Latent Diffusion Models (LDMs) mitigate computational costs by operating in compressed spaces, they still rely on a decoupled, fixed decoder to map latents back to pixels. We introduce CrossFlow, a generative paradigm that unifies these stages by allowing the noise prior and final output to reside in different spaces. By deriving a novel cross-space objective, our framework enables a single model to map directly from a noisy latent to a high-resolution image in a single function evaluation (NFE). CrossFlow serves a dual purpose: it acts as a high-fidelity one-step generator and functions as an enhanced decoder for existing LDM pipelines, capable of refining imperfect latent estimations during the reconstruction process. On ImageNet-1k (256 × 256), CrossFlow achieves a state-of-the-art 1.62 FID with only one NFE. Our results demonstrate that cross-space flow objectives provide a scalable and theoretically grounded framework for unifying latent generation and pixel-space decoding.
PaperID: 4134, Poster
Abstract: Existing 3D point cloud tracking methods have not fully exploited temporal information in point cloud sequences, making it difficult to effectively capture the continuous evolution of target states and thereby limiting tracking stability in scenarios with large cross-category geometric variation. To address this issue, we propose TETrack3D, a framework for category-agnostic 3D point cloud tracking. The core idea of TETrack3D is to elevate temporal information from an auxiliary cue to an explicit learning constraint for cross-frame association. Specifically, we introduce a historical feature reuse mechanism to preserve and reuse target-related features accumulated over multiple frames, enabling the association process in the current frame to access richer and finer-grained historical context. We further design a state evolution supervision module, which guides the model to learn temporally consistent target motion patterns from historical observations by predicting future target states. In addition, to alleviate cross-frame feature drift caused by target-state changes in point cloud sequences, we propose a temporal distribution alignment strategy based on optimal transport theory, which constrains target features across adjacent frames at the distribution level and improves the stability of temporal target modeling. Extensive comparative experiments on KITTI, nuScenes, and Waymo demonstrate that explicitly modeling the target-state evolution process effectively improves the localization robustness of 3D object tracking.
PaperID: 4135, Poster
Abstract: Embodied AI systems can pass planning and still fail in execution: a delay, resource conflict, synchronization change, or candidate invalidation can turn a valid point schedule into broken commitments. We introduce FLECAR, which gives the executor a dispatchable coordination envelope instead of fixed start times. Offline, Flexible Envelope Synthesis builds executable windows, coupling constraints, and a dispatch graph; online, Controllability-Aware Repair updates that envelope after events and repairs only the affected region when dispatchability fails. On generated coordination instances with the same event streams and budgets, FLECAR improves post-event feasibility from 0.3880 to 0.8400, terminal success from 0.5800 to 0.8400, and normalized repair burden from 0.5357 to 0.2964 relative to legacy point-schedule repair. A point baseline with CAR-like repair closes part of the gap, but envelopes still retain higher post-event feasibility and lower repair burden, while pooled one-step survivability remains near parity. The result points to a simple execution principle: after disturbances, live windows, couplings, and alternatives serve the executor better than a point schedule rebuilt after the fact.
PaperID: 4136, Poster
Authors:
Xiaoang Zhang, Mert Kiray, Yordanka Velikova, Manoj Biswanath, Benjamin BusamAbstract: Feed-forward 3D foundation models have recently reshaped multi-view reconstruction, yet dynamic 4D understanding still often relies on large-scale motion supervision or hand-crafted, architecture-specific heuristics. We ask whether frozen 3D foundation model tokens already encode motion-relevant information that can be exposed by lightweight readout functions, rather than re-learning motion-specific features with high-capacity decoders. Through probing on three representative backbones, we find that static--dynamic separation is decodable from frozen token representations. Our rank-restricted and spectral analyses on \pi^3X reveal that the recovered motion cues are concentrated in a low-dimensional subspace. Building on this, we propose Probing-based Motion Distillation (PMD), which converts pairwise motion-state compatibility scores against a fixed anchor token into dense motion masks. Trained on only 160 YouTube-VOS videos using a single 16 GB GPU, PMD transfers zero-shot to DAVIS, SegTrackv2, and FBMS-59 without target-benchmark fine-tuning. With multi-view pretrained geometric backbones, PMD outperforms training-free heuristics and approaches substantially heavier supervised and self-supervised methods, suggesting that motion decoding from frozen 3D foundation models can be lightweight and data-efficient.
PaperID: 4137, Poster
Authors:
Luoyang Sun, Jiwen Jiang, Yifeng Ding, Fengfa Li, Yan Song, Haifeng Zhang, Lei Ren, Kun Zhan, Chen Wei, Xie Yan, Jun Wang, Cheng DengAbstract: Recent embodied AI systems increasingly rely on large language models (LLMs) for high-level planning, semantic reasoning, and long-horizon decision making. However, deploying such capabilities on edge platforms such as autonomous vehicles and mobile robots requires balancing model quality against strict latency and hardware constraints. Existing LLM architectures are primarily designed for cloud-scale accelerators and often become inefficient in on-device settings. We present PLAS (Pareto-optimal LLM Architecture Search), a hardware-aware framework for identifying Pareto-optimal on-device LLM architectures under deployment constraints. Our approach jointly models training loss and inference latency as functions of architectural design choices. We fit an empirical architecture-to-loss model from 170 LLMs trained on 10B tokens each, and estimate latency using roofline-based hardware modeling, enabling efficient exploration of the accuracy-latency Pareto frontier without exhaustively training or profiling every candidate architecture. Using PLAS, we evaluate 1942 candidate architectures on NVIDIA Jetson Orin. At the same measured latency as Qwen2.5-0.5B, our co-designed architecture achieves 19.42% lower WikiText-2 perplexity. Our results suggest that effective edge deployment requires explicit hardware--model co-design, and that Pareto-based architecture modeling provides a practical alternative to exhaustive neural architecture search for on-device LLMs.
PaperID: 4138, Poster
Authors: Abhirup Mondal, Anirban Majumder, Vineet Chaoji
Abstract: Estimating counterfactual outcomes when multiple treatments are applied simultaneously, the composite treatment setting, is fundamentally harder than the single-cause case: the treatment space grows combinatorially, most combinations are unobserved, and confounding is entangled across dimensions. We propose PACE (Progressive Alignment for Composite Effects), which exploits a known causal DAG over treatment variables to progressively transform the observational data distribution into the target interventional distribution. PACE performs K sequential augmentation steps in reverse topological order, each isolating a single treatment component and conditioning only on its parents in the DAG. We derive the closed-form optimal augmentation density minimizing KL divergence at each step, establish finite-sample bounds prescribing the augmentation budget via DAG-aware selection bias ratios, and show that the DAG provides a deterministic processing order that eliminates the need for random permutation averaging. On semi-synthetic benchmarks with real covariates (IHDP, Twins, Hillstrom) and controlled synthetic environments, PACE substantially outperforms seven baselines on both continuous and binary outcomes, with the largest gains on rare treatment combinations where confounding is most severe. Ablations confirm consistent improvements over the closest prior method across varying DAG topologies, confounding strengths, and under DAG misspecification.
Abstract: Multi-agent LLM systems improve reasoning and tool use via role specialization, yet reinforcement learning (RL) post-training for such systems remains unstable and underexplored. We theoretically pinpoint a key source of instability when extending group-based RL to cooperative multi-agent LLM systems: under GRPO-style optimization, a global normalization baseline can mismatch heterogeneous agents' reward distributions, inducing gradient-norm instability. Based on this finding, we propose Dr. MAS, a simple and stable RL recipe for cooperative multi-agent LLM systems. Algorithmically, Dr. MAS normalizes advantages per agent using each agent's own reward statistics, which calibrates gradient scales and dramatically stabilizes the training. Systemically, Dr. MAS provides an end-to-end multi-agent RL framework with scalable orchestration, flexible per-agent LLM serving and optimization, and shared resource scheduling of actor backends. Unlike single-actor frameworks such as veRL, Dr. MAS natively supports multiple heterogeneous LLMs over a unified GPU pool. Lifecycle-aware backend management and dynamic dispatch release inactive models' GPU memory and schedule active agents on demand, enabling hardware-efficient co-training. We evaluate Dr. MAS on multi-agent math reasoning and multi-turn search with Qwen2.5 and Qwen3. Dr. MAS achieves clear gains over vanilla GRPO (e.g., +5.6% avg@16 and +4.6% pass@16 on math, and +15.2% avg@16 and +13.1% pass@16 on search) while largely eliminating gradient spikes. It also remains highly effective under heterogeneous agent-model assignments while improving efficiency. Code is available at https://anonymous.4open.science/r/DrMAS-0787.
PaperID: 4140, Poster
Authors: Toru Hishinuma, Kei Senda
Abstract: Learning latent environment representations from offline datasets is an important challenge for Bayes-adaptive Markov decision process model learning. A key difficulty is that offline datasets collected under heterogeneous protocols can differ in state--action visitation even when the underlying dynamics are identical. Visitation-decoupled objectives address this by evaluating transition predictive fit under a protocol-independent reference distribution via a change of measure. We show, however, that the resulting objective-induced invariances need not be inherited by amortized inference. An unconstrained encoder can depend on protocol-induced visitation variations that are invisible to the objective, a structural failure mode we term \emphamortization mismatch. To address this issue, we formalize an \emphobjective-aligned design principle as an objective--inference compatibility condition: the encoder should factor through the objective-induced dataset signature. We instantiate this principle in multiple set-encoder families and, through targeted experiments, provide evidence that objective-aligned encoders improve posterior consistency and can mitigate downstream planning discrepancy across protocols.
PaperID: 4141, Poster
Abstract: We present Correlated Latent Uncertainty Estimation (CLUE), a single-pass deep uncertainty quantification framework that combines efficient amortized neural prediction with Bayesian-style information propagation across inputs. CLUE introduces a kernel-based prior, allowing cross-input dependence over latent variables while preserving test-time inference efficiency. CLUE requires neither posterior sampling nor ensembles, and avoids matrix inversion during inference, yet recovers posterior contraction behavior analogous to conventional Bayesian models. In the white-noise limit of the kernel, CLUE reduces to standard evidential deep learning (EDL). This limiting case reveals standard EDL as amortized variational inference with an independent latent structure, providing a probabilistic explanation for several pathologies of EDL identified in prior work. Empirically, CLUE exhibits consistent Bayesian-like uncertainty contraction and improved uncertainty quality across synthetic and real-world benchmarks, while maintaining competitive predictive accuracy and fast single-pass inference.
PaperID: 4142, Poster
Abstract: AdamW is ubiquitous in deep learning, yet its behavior remains poorly understood. We analyze its dynamics through the lens of dynamical systems and show that AdamW admits an \emphimplicit fixed-point objective: its fixed points coincide with the stationary points of a constrained and regularized optimization problem. However, not all of these fixed points are stable under AdamW’s dynamics, and stability depends sensitively on curvature, weight decay, and momentum parameters. Even in simple one-dimensional settings, AdamW can exhibit surprisingly complex behavior: equilibria may be unstable, and numerical trajectories can exhibit persistent oscillations rather than convergence to them. We further extend the analysis to higher dimensions, deriving sufficient conditions for local stability of the continuous-time dynamics, and propose a convergent modification that preserves AdamW's fixed points. These results clarify what optimization problem AdamW is associated with, when its convergence can be expected, and how its dynamics could inspire more reliable optimizers.
Abstract: Recent empirical studies have shown the training dynamics of deep neural networks mostly occur within low-dimensional subspaces. While this has inspired new research in low-rank training, compression, and adaptation, theoretical justification for these dynamics in nonlinear networks remains limited. To address this gap, this paper analyzes the learning dynamics of multi-layer perceptrons (MLPs) under gradient descent (GD). We demonstrate that the weight dynamics concentrate within invariant low-dimensional subspaces throughout training. Theoretically, we precisely characterize these invariant subspaces for two-layer networks with smooth nonlinear activations, providing insight into their emergence. Experimentally, we validate that this phenomenon extends well beyond our theoretical setting. Leveraging these insights, we empirically show there exists a low-rank MLP parameterization that, when initialized in the appropriate subspaces, nearly matches the performance of their fully-parameterized counterparts on classification tasks.
PaperID: 4144, Poster
Abstract: Continual test-time adaptation (CTTA) updates a model online on a stream of unlabeled data under distribution shift. In practice, CTTA often becomes unstable: small early prediction biases are reused as supervision, causing pseudo-labels to progressively concentrate on a few classes and eventually leading to drift or collapse. We observe that this instability is closely linked to the temporal contraction of pseudo-label class usage during adaptation. Existing CTTA methods often address reliability or diversity through sample selection or loss weighting, but they do not explicitly control how pseudo-label mass is allocated across classes over time. We introduce (Sinkhorn Class-Marginal Regularization), a lightweight plug-in module that implements this temporal control over a short window of recent predictions. SCALAR maintains a buffer of recent predictions, applies Sinkhorn balancing to project the window-level class marginal toward a target distribution, and then uses the resulting balanced soft targets for adaptation. SCALAR requires no architectural changes and adds less than 1% runtime and memory overhead. SCALAR is most beneficial under skewed, non-stationary, and aggressive-adaptation regimes, where it substantially mitigates drift and collapse, even on top of entropy- and diversity-aware stabilization mechanisms.
Abstract: Sequence prediction methods for dynamical systems with long memory, i.e. marginally stable systems, typically achieve regret that grows polynomially with the hidden dimension of the underlying generative model. Universal Sequence Preconditioning (USP) \citemarsdenuniversal is a method that compresses any sequence which comes from a linear dynamical system into a ``preconditioned'' sequence which requires exponentially shorter memory for accurate prediction. However, the preconditioned sequence yields exponentially larger diameters and gradients, hindering USP from unlocking optimal regret bounds. Inspired by the minimum description length principle, we show that the Vovk-Azoury-Warmuth (VAW) algorithm is naturally matched to the USP regime. Indeed, it takes advantage of the memory compression while remaining robust to the exponential explosion of the diameter. We prove that combining USP with VAW achieves astoundingly strong results: for any marginally-stable linear dynamical system, this algorithm achieves polylogarithmic regret O\left( \log^3 T \right) even in the presence of asymmetric hidden transition matrices. Finally, we extend the applicability of USP beyond bounded-spectrum systems by providing new complex-analytic bounds on Chebyshev polynomials, allowing for systems with constant complex arguments.
Authors:
Eugenio Lomurno, Filippo Balzarini, Francesco Benelle, Francesca P Panaccione, Matteo MatteucciAbstract: Diffusion-based generators set the current state of the art for synthetic tabular data, deployed downstream wherever direct access to real records is restricted. These methods approach but rarely exceed real-data utility on downstream tasks, and closing this synthetic--real performance gap has so far been pursued exclusively at training time, via architectural advances, scaling, and retraining of monolithic generators. The inference-time alternative, i.e., refining the outputs of a pre-trained backbone with parameters left untouched, has remained largely unexplored for tabular synthesis. We introduce TARDIS (Tabular generation through Refinement, Distillation, and Inference-time Sampling), an inference-time refinement framework that operates on a frozen pre-trained backbone, configured per dataset by a Tree-structured Parzen Estimator search over score-level guidance during reverse diffusion, with each trial's objective set by an inner grid search over post-hoc sample selectors and an optional soft-label distillation step. The search space encodes a single mathematical pattern we name Bidirectional Chamfer Refinement (BCR): the symmetric Chamfer functional between synthetic and real samples is minimized both continuously, via a score-level gradient during reverse diffusion, and discretely, via batch-ranking post-generation. On the majority of datasets the search selects BCR-aligned configurations over alternatives encoded in the search space, evidence both for BCR as the dominant refinement pattern and for TARDIS's per-dataset search as a procedure that recovers this pattern. Across 15 binary, multiclass, and regression benchmarks TARDIS achieves a median +8.6% downstream-task improvement over models trained on real data (95% CI [+3.3, +16.4], Wilcoxon p=0.016, 11/15 strict wins) and improves over the underlying TabDiff backbone on all 15 datasets (mean +12.9%, p<10^-4), matching the backbone on manifold fidelity, diversity, and sample-level privacy. The synthetic--real gap is therefore not primarily a training-time problem: on the studied corpus, inference-time refinement of a pre-trained tabular diffusion backbone reaches and exceeds real-data utility in 1 to 80 minutes on a single consumer-grade GPU.
PaperID: 4147, Poster
Abstract: Context-based offline meta-reinforcement learning (COMRL) has progressed almost entirely through better task representations: separating task identity from behavior-policy artifacts so that the latent reflects the task information alone. The implicit expectation has been that a clean representation will transfer to out-of-distribution (OOD) tasks. However, on an OOD task the encoder produces a latent outside the training support, where the policy has never been trained. We address this gap with Pessimistic Latent Task-aware Optimization (PLATO), which exposes the policy to off-support latents at training time. PLATO assumes a mild manifold hypothesis: task latents lie on a low-dimensional manifold, with OOD tasks further from the training centroid. It exploits this geometry by perturbing inferred latents outward from the centroid to synthesize counterfactual tasks; a learned decoder ensemble then rolls out short trajectories at the perturbed latent, and the policy is updated on the imagined transitions under a pessimism weight that down-weights steps where the ensemble is unconfident. We prove that the perturbation reaches off-support by a controllable margin, that the epistemic disagreement at the perturbed latent is non-vanishing, and that the OOD generalization gap is bounded by the same perturbation reach. On eight MuJoCo continuous-control benchmarks, PLATO consistently improves out-of-distribution returns over strong COMRL baselines while remaining competitive in-distribution.
Abstract: Tokens are discrete representations that allow modern deep learning to scale by transforming high-dimensional data into sequences that can be efficiently learned, generated, and generalized to new tasks. While foundational for image and video generation, the application of tokens to physical simulation remains nascent. Because existing tokenizers are designed for the perceptual requirements of natural images, they struggle with scientific data, which exhibits large dynamic ranges and requires exact preservation of physical and spectral properties. In this work, we investigate the performance of a suite of image tokenizers across metrics designed to measure PDE fidelity. Observing that these baselines struggle to simultaneously capture fine geometric details and precise physical magnitudes, we propose Phaedra, a novel tokenizer inspired by classical shape-gain quantization and the paradigm of basis functions coupled with continuous coefficients. Phaedra acts as a highly effective nonlinear compression algorithm, massively reducing dataset footprints while maintaining physical fidelity. We demonstrate that Phaedra consistently improves reconstruction across diverse PDE datasets, generalizes robustly to unseen PDE types and real-world Earth observation data, and provides immediate accuracy gains in initial downstream operator learning and masked autoencoding tasks.
PaperID: 4149, Poster
Abstract: Large language models rely on a single tokenizer that is chosen at training time and fixed thereafter. Most tokenizers are trained on English-dominant corpora, making them under-serve morphologically rich and non-Latin-script languages, where they can produce fragmented representations, longer sequence lengths, and reduced information density during training. We propose (MoTo), a modular tokenization framework that trains dedicated per-language BPE tokenizers and composes them into a unified superset vocabulary, where tokens shared across languages are deduplicated into a common ID space. Each language tokenizer maintains its own normalization policy, pre-tokenization rules, and vocabulary budget. Despite allocating significantly fewer entries to English, MoTo retains English performance while delivering consistent gains across 21 typologically diverse languages spanning six scripts, with improvements most pronounced on morphologically complex and non-Latin-script languages that are systematically underserved by English-dominant tokenizer design. Our results suggest that modular subword tokenization is a practical and extensible alternative to monolithic tokenization.
PaperID: 4150, Poster
Abstract: We propose D-DOIT (Discrete Doob-Oriented Inference-time Transformation), a training-free and efficient adaptation method for discrete diffusion models with generic rewards. D-DOIT formulates adaptation as sampling from a reward-tilted target distribution and realizes this transport through Doob's h-transform of the discrete diffusion reverse kernel, using only reward values rather than reward gradients. Unlike continuous diffusion, masked discrete diffusion samples categorical token-reveal transitions rather than continuous state updates. D-DOIT derives the corresponding discrete Doob's h-transform, which guides sampling by reweighting reverse transition probabilities instead of adding a drift correction. To make this transformation practical, D-DOIT avoids expensive future rollouts. At each guided step, D-DOIT samples candidate next states, uses the model prediction head to complete each candidate into a clean sequence, evaluates each completion with the reward oracle, and resamples the next state with probabilities proportional to the rewards. An optional late-stage best-of-K refinement further improves sample quality by branching trajectories only near the end of denoising, avoiding the K-fold cost over the full trajectory. Empirically, on regulatory DNA sequence design benchmarks, D-DOIT consistently outperforms training-free guidance baselines and achieves performance competitive with training-based methods, improving both enhancer activity and cell-type-specificity while preserving sequence naturalness.
PaperID: 4151, Poster
Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities, yet they frequently struggle with fine-grained visual perception, often suffering from hallucinations when faced with visually similar but semantically distinct instances. A major underlying cause is the conventional single-sample Supervised Fine-Tuning (SFT) paradigm, which isolates training instances and limits the model's ability to explicitly learn subtle discriminative boundaries. To address this, we propose framework that enhances fine-grained perception by shifting from independent supervision to joint supervision. Specifically, C2FT assesses the MLLM's internal uncertainty to dynamically mine model-specific semantic confusions, subsequently constructing challenging multi-image groups comprising hard positive and hard negative samples. By interleaving these samples into a unified prompt, our approach forces the MLLM's internal attention mechanisms to cross-reference images and explicitly capture localized visual discrepancies. Extensive experiments on both Fine-Grained Visual Classification (FGVC) and visual question answering (VQA) benchmarks demonstrate the effectiveness of C2FT. For instance, on the Qwen3-VL-4B model, our approach achieves an average improvements of 3.87% and 4.35% points over standard SFT and GRPO baselines, respectively.
Abstract: Training large language models (LLMs) as multi-turn conversational agents remains a significant challenge, particularly in goal-oriented settings. The difficulty stems from sparse, long-horizon objectives and the discrepancy between response-level planning and token-level generation. In this paper, we present a formal reduction of the multi-turn RL problem into a \emphsequence of single-turn RLHF-style problems. This is achieved by setting a learned multi-turn Q-function as the reward model for the single-turn problem. We demonstrate and prove a key insight: solving this single-turn RLHF problem with standard token-level GRPO is equivalent to an approximate policy improvement step within the multi-turn problem. This insight naturally leads to \emphIterative GRPO, a batch online approximate policy iteration algorithm that alternates between collecting a batch of data from the current policy, fitting Q-functions from these logged conversation trajectories, and improving the policy via single-turn RLHF. A major practical advantage is that Iterative GRPO directly leverages stable, off-the-shelf single-turn RLHF tools, making it straightforward to implement. In addition, the method occupies a middle ground between fully online and fully offline approaches, retaining the adaptability of online updates while gaining the stability benefits of offline training. Empirically, we demonstrate the effectiveness of Iterative GRPO on six multi-turn conversational environments.
PaperID: 4153, Poster
Abstract: PDE foundation models (PDE-FMs) promise general-purpose surrogate simulation, yet deploying a pretrained model to a new physical regime inevitably introduces distribution mismatch, demanding task-specific post-training. The dominant approach, supervised fine-tuning (SFT), is fundamentally off-policy: the model trains on pre-generated states but is evaluated autoregressively on its own. In chaotic regimes, this mismatch is catastrophic as per-step errors compound exponentially along the Lyapunov spectrum, and exact pointwise trajectory matching beyond the predictability horizon becomes an ill-posed training signal. To resolve this, we introduce , an on-policy post-training framework for PDE-FMs. During training, the student rolls out autoregressively; a numerical solver branches multi-step correction targets from each student-visited state; and the student is updated against these solver corrections, learning to correct its own rollout dynamics. We instantiate SaT in a unified post-training suite with SFT, physics-informed temporal alignment, and learned-teacher on-policy distillation. Across seven 2D/3D PDE benchmarks and six pretrained PDE-FMs, SaT variants achieve the lowest 50-step rollout RMSE in 23 of 24 in-distribution model--task pairs. The best SaT variant reduces RMSE over SFT up to 90.0%, and SaT gives the lowest RMSE on all seven OOD parameter-shift benchmarks.
PaperID: 4154, Poster
Abstract: Feed-forward 4D Gaussian Splatting (4DGS) offers an efficient paradigm for reconstructing dynamic scenes from monocular videos, bypassing the need for lengthy per-scene optimization. However, existing methods suffer from three fundamental bottlenecks: redundant frame-wise grid-aligned Gaussian representations, suboptimal single-step 2D-to-4D regression architectures, and restrictive dependencies on 3D/4D signals during training or inference. In response to these limitations, we introduce a novel Gaussian-Grounded Query Refinement (GGQR) framework. To reduce representation redundancy, GGQR learns a compact set of Gaussian queries equipped with a newly designed plateau-shaped temporal kernel, enabling single primitives to persistently model motion-consistent regions across extended frames. To alleviate regression ambiguity, these queries are iteratively refined through a Gaussian-grounded attention mechanism operating in 3D space. Specifically, GGQR leverages foundation models to lift 2D observations into a shared 3D space, explicitly grounds the 3D attention in intermediate Gaussian states, and residually updates those Gaussian states. Trained entirely self-supervised on casual videos, GGQR achieves a new state-of-the-art on dynamic reconstruction benchmarks among feed-forward 4DGS methods, delivering superior performance in novel view-time rendering alongside downstream motion modeling tasks.
PaperID: 4155, Poster
Authors:
Xiang Gao, Xiaoping Zhu, Xinmu Wang, Yuanpeng Liu, Haotian Yin, Dichang Zhang, Yazheng Chen, Yue Wang, Yubin Zhou, Yu Guo, Wei Chen, David Gu, Xiyun Song, Heather YuAbstract: 3D Gaussian Splatting (3DGS) has emerged as an effective solution for real-time novel view synthesis (NVS) with high rendering quality, but training memory scales with image resolution, making downsampled supervision common and leading to loss of sharp, high-frequency details and degraded perceptual quality. Moreover, uniform supervision assigns equal weight to all pixels, resulting in weak emphasis on detail-rich regions such as edges and thin structures. We propose OT-3DGS, which leverages optimal transport in screen space to reparameterize the image domain prior to downsampling, inducing importance-based non-uniform supervision that redistributes pixel budget toward geometry-critical regions and strengthens supervision on high-frequency content. Our method remains fully compatible with the original 3DGS pipeline, requiring no modification to loss functions, rendering settings, or pruning, cloning, and splitting strategies. Experiments across diverse datasets show that OT-3DGS produces sharper structures, clearer reflectance, and reduced over-smoothing; while PSNR and SSIM may slightly decrease due to deviation from downsampled references, reference-free metrics including BRISQUE, NIQE, and Laplacian Variance (LV) improve significantly, better reflecting perceptual quality and high-frequency detail. These results demonstrate that OT-3DGS aligns with the fundamental goal of NVS: generating perceptually realistic, high-quality unseen views.
Abstract: Diffusion-based imitation learning has shown strong promise for robot manipulation. However, most existing policies condition only on the current observation or a short window of recent observations, limiting their ability to resolve history-dependent ambiguities in long-horizon tasks. To address this, we introduce DSSP, a history-conditioned Diffusion State Space Policy that enables efficient, full-history conditioning for robot manipulation. Leveraging the continuous sequence modeling properties of State Space Models (SSMs), our history encoder effectively compresses the entire observation stream into a compact context representation. To ensure this context preserves critical information regarding future state evolution, the encoder is optimized with a dynamics-aware auxiliary training objective. This high-level context representation is then seamlessly fused with recent state observations to form a hierarchical conditioning mechanism for action generation. Furthermore, to maintain architectural consistency and minimize GPU memory overhead, we also instantiate the diffusion backbone itself using an SSM. Extensive experiments across simulation benchmarks and real-world manipulation tasks show that DSSP achieves state-of-the-art performance with a significantly smaller model size, demonstrating superior efficiency of the hierarchical conditioning in capturing crucial information as the history length increases.
PaperID: 4157, Poster
Authors:
Sayantan Pal, Kaiyi Ji, Rohini SrihariAbstract: Fluent guidance is not the same as useful intervention. LLM tutors are typically trained to generate the next teacher utterance, implicitly assuming that every student turn warrants a response. However, our experiments indicate that this conflates tutoring capability (what to say) with intervention necessity (whether to say it). We introduce DICE, a framework that decouples intervention decisions from response generation by first selecting an explicit pedagogical action. To calibrate this action selection policy, we define Intervention Value (IV), a rollout-grounded counterfactual metric that compares each action against non-intervention. IV shows that many prescribed interventions provide little or no marginal benefit. We further introduce DICE-Bench, a multi-variant math tutoring benchmark with skill-preserving problem variants for session-level evaluation. Using IV-weighted and KL-regularized policy optimization, DICE learns to intervene selectively while preserving tutoring effectiveness. In simulated tutoring sessions, DICE reduces the over-intervention rate to near zero while guiding students to correct solutions in approximately 3-4 fewer turns on average than existing Socratic tutoring baselines.
PaperID: 4158, Poster
Authors: Suhwan Bong
Abstract: The standard rule in unconstrained causal policy learning is simple: treat individuals with nonnegative CATE. We show that this rule can fail when treatment eligibility is strategically manipulable. Agents may move or report covariates after a policy is announced, so treatment is assigned using manipulated covariates while welfare remains determined by baseline CATEs. This creates a mismatch between who receives treatment and who benefits from treatment. We formalize this setting as strategic causal policy learning and show that strategic welfare equals the CATE-weighted mass of baseline types that can reach the treatment region. This access-burden representation leads to three distinct strategic thresholds: a welfare threshold, a safe threshold, and a fair threshold. For linear CATE models, we derive closed-form threshold corrections. Homogeneous costs admit a single buffer correction that recovers the ideal positive-CATE allocation, while heterogeneous costs make welfare maximization, no-exploit safety, and subgroup fairness select different policies. Experiments illustrate these tradeoffs and show when group-specific corrections can restore the ideal allocation.
PaperID: 4159, Poster
Abstract: Given a must-link graph G = (V, E_G) and a cannot-link graph H=(V,E_H) as input, the objective of constrained clustering is to partition V into k parts while minimizing the maximum cut ratio between G and H. In this paper we design and analyze a variant of the spectral clustering algorithm using the k smallest generalized eigenvectors as the embedding space. We prove that our proposed algorithm achieves an O(1)-approximation of the optimal cut-ratio under a natural assumption on the input graphs. We further study the case in which the cannot-link constraints might not be available, and develop an iterative pipeline that repeatedly applies the output of spectral clustering or constrained clustering to generate new constraints, thereby progressively improving cluster quality. We experimentally compare our proposed algorithm with the previous state-of-the-art, and demonstrate the superior performance of our algorithms on both synthetic and real-world datasets.
Abstract: Multimodal Large Language Models (MLLMs) achieve strong performance through instruction tuning, yet real-world deployment often requires continual capability expansion across sequential tasks. In such scenarios, Multimodal Continual Instruction Tuning (MCIT) aims to acquire new capabilities while limiting catastrophic forgetting. Existing methods mainly follow a module-composition paradigm: they maintain task-level prompts or LoRA experts and dynamically route or aggregate a subset of them at inference. However, samples within the same task can still differ substantially in visual scenes, question intents, and reasoning demands. This motivates instance-level adaptation to individual query-image pairs rather than only selecting or combining task-level modules. To this end, we propose DRAPE (Dynamic Cross-Modal Prompt Generation), a prompt-learning framework that synthesizes continuous instance-specific soft prompts for MCIT. Instead of selecting prompts from a fixed pool, DRAPE derives prompt queries from the textual instruction and cross-attends to visual patch features, producing query-image conditioned prompts that are prepended to the frozen LLM. To mitigate forgetting during sequential updates, DRAPE applies null-space gradient projection to the shared projector and uses CLIP-based prototype routing for task-label-free generator selection at inference. Extensive experiments on MCIT benchmarks show that DRAPE achieves state-of-the-art performance among representative prompt-based and LoRA-based continual-learning baselines.
Abstract: Diffusion models provide powerful priors for zero-shot video inverse problems, but their real-time deployment is hindered by two inefficiencies: high initial latency caused by holistic video restoration, and low throughput resulting from multiple VAE passes to enforce measurement consistency in pixel space. To overcome these limitations, we propose Autoregressive Video Inverse problem Solver (AVIS). The AVIS framework leverages autoregressive video diffusion models to restore videos in a streaming manner, naturally eliminating latency bottlenecks. Specifically, AVIS initializes reverse diffusion with a measurement-consistent estimate, reducing the required sampling steps. Compared to leading non-autoregressive solvers, AVIS drastically reduces initial latency from 114s to 4s and increases throughput from 0.71 to 1.18 FPS while achieving superior restoration quality. We further introduce a highly accelerated variant, dubbed AVIS substantially boosts throughput to 5.91 FPS on a single RTX 4090 GPU while maintaining competitive performance and achieving a favorable efficiency–performance trade-off, paving the way toward real-time deployment.
PaperID: 4162, Poster
Authors:
Zhenke Duan, Xin Li, Jiqun Pan, Xiaofei Dong, Hanwen Ning, Xinyuan SongAbstract: In this paper, we propose Data-Adaptive Mahalanobis Cross-Head Attention (DAMCHA), a simple and effective attention mechanism that generalizes Multi-Head Attention (MHA) through the lens of metric learning. Standard MHA implicitly learns a static, head-wise Euclidean similarity corresponding to the diagonal blocks of a unified cross-head metric matrix, discarding all off-diagonal interactions across heads. DAMCHA addresses this limitation via two key innovations. First, it learns the full cross-head metric matrix, inducing a Mahalanobis-like similarity that explicitly captures inter-head interactions beyond the restricted block-diagonal structure of MHA. Second, it parameterizes the metric matrix as an input-dependent function, enabling the attention geometry to dynamically adapt to the intrinsic manifold structure of the data rather than remaining fixed after training. We further introduce principled regularization strategies, most notably stack-wise parameter sharing, to ensure computational efficiency and stable optimization. DAMCHA serves as a drop-in replacement for MHA and its variants in various Transformer architectures, delivering superior expressive power and learning performance with comparable or reduced computational overhead. Extensive experiments demonstrate that DAMCHA-based Transformers consistently outperform competing baselines across a range of benchmark tasks. Code is available at https://anonymous.4open.science/r/DAMCHA-3B37.
Authors: Ethan Hsu, Ivan Ge, Hong M Yam
Abstract: Partial differential equation (PDE) solvers underpin scientific computing, but real-world deployment is bounded by compute. Classical Monte Carlo solvers such as Walk-on-Spheres (WoS) are unbiased and geometry-agnostic but are slow. Learned solvers are fast but biased and brittle under distribution shift. We present MC^2, a hybrid WoS-Neural Network (WoS-NN) PDE solver that treats a low-budget Monte Carlo solution as a structured estimator of the true field and learns a single-pass neural correction to recover a high-fidelity solution. MC^2 matches the accuracy of solutions using over 1000× more Monte Carlo compute, outperforming all evaluated classical, denoising, and neural-operator baselines. To enable reproducible study of finite-compute PDE solving, we additionally release PDEZoo, the largest standardized elliptic PDE benchmark to date: 2M PDEs spanning five elliptic families and unlimited geometric compositions, with analytic ground truth and multi-budget Monte Carlo trajectories. Together MC^2 and PDEZoo (1) empirically establish that finite-sample Monte Carlo error is structured, learnable, and correctable in a single forward pass, (2) show that we can solve PDEs ~1000x faster than with just WoS, and (3) provide the evaluation infrastructure the field has so far lacked.
Abstract: Uncertainty quantification (UQ) in deep learning regression is of wide interest, as it supports critical applications including sequential decision making and risk-sensitive tasks. In heteroskedastic regression, where the uncertainty of the target depends on the input, a common approach is to train a neural network that parametrises the mean and the variance of the predictive distribution. Yet to this day, training deep heteroskedastic regression models poses severe practical challenges in the trade-off between uncertainty quantification and mean prediction, such as variance overfitting, optimization difficulties and representation collapse. In this work, we identify the core issues and propose a simple and efficient procedure that addresses these jointly by post-hoc fitting a variance model across the intermediate layers of a pretrained network on a hold-out dataset. We demonstrate that this method is competitive with end-to-end trained mean-variance networks in heteroskedastic UQ on several data modalities. The method retains mean prediction accuracy, is cheap at both train and prediction time, and requires no additional data compared to existing methods. Finally, we show this method works on large scale foundation models for chemistry, paving the way for cheap heteroskedastic UQ in large scale neural networks without having to retrain such models end-to-end.
PaperID: 4165, Poster
Authors:
Zheyuan Zhang, Jiwei Zhang, Rongjing Cai, Nihua Shrestha, Zilong Wang, Boyu Zhou, Hong ChenAbstract: Visual Place Recognition (VPR) aims to identify and retrieve geographic locations from visual observations. With the advent of Visual Foundation Models (VFMs), recent VPR methods have primarily adopted them as backbones to exploit their strong generalization capabilities, effectively mitigating challenges such as perceptual aliasing and long-term appearance variations. However, their large parameter sizes and high inference latency impose substantial hardware demands, significantly hindering practical deployment. To retain the generalization strength of VFMs while enabling efficient VPR, we revisit this problem from the perspective of Knowledge Distillation (KD). Specifically, we conduct a systematic investigation of various KD losses for VPR from both theoretical and empirical standpoints. Our analysis reveals that relation-based KD methods consistently achieve superior performance and efficiency, which we attribute to their larger feasible solution space and better alignment with the VPR objective. Building upon this insight, we propose a relation-based KD framework, termed RelaVPR, which is further enhanced along three key dimensions: 1) dynamically filtering noisy knowledge, 2) improving hard-sample learning, and 3) mitigating gradient interference in multi-teacher distillation. Compared with state-of-the-art (SOTA) KD-based VPR methods, RelaVPR requires no additional fine-tuning while delivering higher performance, faster inference, and a more compact model. Moreover, RelaVPR establishes a stronger performance--efficiency frontier than existing SOTA VPR approaches across multiple popular benchmarks. Codes and weights will be released.
PaperID: 4166, Poster
Abstract: Proxy model-based heterogeneous federated learning enables clients with diverse architectures to collaborate through a shared lightweight intermediary. However, current methods aggregate and distill proxy knowledge in a structure-agnostic manner, ignoring layer-wise convergence dynamics, conflating shared and client-specific information, and transferring knowledge without assessing its reliability. These oversights jointly cause model oscillation, drift, and negative transfer. We propose FedPeel, a framework that makes the aggregation pipeline aware of the internal structure of model parameters. Rather than updating the proxy monolithically, FedPeel selectively aggregates only the layers that have stabilized, decouples their frequency-domain representations to separate generalizable patterns from personalized details, and modulates knowledge transfer intensity on a per-sample basis according to teacher confidence. Theoretical analysis establishes an O(1/T ) convergence rate. Experiments on image, text, and audio benchmarks show that FedPeel consistently outperforms state-of-the-art methods in both accuracy and communication efficiency.
Abstract: While feed-forward 3D Gaussian splatting reconstructs renderable Gaussian primitives from sparse context views without per-scene optimization, existing pipelines do not provide a compact scene representation for storage or transmission. A natural solution is to apply existing 3DGS compression methods to the generated Gaussian primitives. However, this approach operates on the final irregular 3D representation and is decoupled from the internal feature-to-Gaussian generation process, which limits compression efficiency. To address this, we introduce \emphCodecSplat, an ultra-compact latent coding framework for feed-forward 3D Gaussian splatting. CodecSplat first encodes an intermediate 2D Gaussian-generation feature into an entropy-coded scene bitstream. At the decoder, the latent feature is reconstructed and used to predict depth and Gaussian parameters, which are then mapped to 3D Gaussian primitives. Note that, by integrating compression into the feed-forward Gaussian generation pipeline, CodecSplat avoids inefficient compression over irregular 3D Gaussian primitives and allows the codec to exploit the structured intermediate feature representation. We instantiate CodecSplat on a feed-forward Gaussian splatting backbone with depth-guided multi-view feature refinement and a hierarchical learned feature codec. On DL3DV and RealEstate10K datasets, CodecSplat achieves 23.56-26.36 dB and 24.76-27.05 dB PSNR with only 20.00-107.77 KiB and 3.37-12.51 KiB per scene, respectively. This is roughly one order of magnitude smaller than compressing feed-forward generated Gaussian primitives, while preserving controllable rate--distortion behavior.
PaperID: 4168, Poster
Abstract: Reinforcement learning with verifiable rewards spends much of its rollout budget on states that yield no learning signal. We recast this as a question about the sampling distribution itself: under a fixed rollout budget, which distribution over states minimizes the variance of the policy-gradient estimator? The answer admits a closed-form Neyman allocation, \mu_\mathrmvar^\pi(s) \propto d^\pi(s)\,\sigma_\pi(s), which connects naturally to GRPO's group-relative advantage and induces an emergent reverse curriculum as the policy improves, with no externally imposed schedule. Building on this principle, we propose VSR (Variance-driven State Resampling), a plug-in modification to GRPO using a prefix replay buffer and a zero-cost variance proxy. On mathematical reasoning benchmarks and long-horizon agentic tasks (Sokoban, BabyAI), VSR outperforms GRPO and GiGPO on most benchmarks with fewer generated tokens per update.
Abstract: Gradient descent dynamics on the deep matrix factorization problem is extensively studied as a simplified theoretical model for deep neural networks. Although the convergence theory for two-layer matrix factorization is well-established, no global convergence guarantee for general deep matrix factorization under random initialization has been established to date. To address this gap, we provide a polynomial-time global convergence guarantee for randomly initialized gradient descent on four-layer matrix factorization, given certain conditions on the target matrix and a standard balanced regularization term. Our analysis employs new techniques to show saddle-avoidance properties of gradient descent dynamics, and extends previous theories to characterize the change in eigenvalues of layer weights.
PaperID: 4170, Poster
Abstract: Electrolyte conductivity is not a property of isolated molecules, but a response over a formulation landscape defined by solvents, salts, composition, salt mole ratio, and temperature. The same molecular component can play different effective roles as its amount and operating conditions change, while the underlying formulation state is observed only through macroscopic conductivity measurements. This makes conductivity prediction a distinctive mixture learning problem: a model must learn how molecular identities become formulation-specific states, rather than merely how molecules are represented. We introduce SolvMix, a framework for learning formulation-state landscapes for electrolyte conductivity prediction. SolvMix constructs component tokens for solvent and salt molecules, calibrates them into amount-aware component states before mixture interaction, and predicts conductivity through a temperature-formatted readout. We evaluate SolvMix on three curated electrolyte conductivity benchmarks, CALiSol-23, DiffMix, and Bamboo-mixer, covering diverse solvents, salts, compositions, salt mole ratios, and temperatures. Across standard splits on all three benchmarks and binned out-of-distribution (OOD) splits on Bamboo-mixer, SolvMix outperforms strong molecular-representation, mixture-aggregation, and geometric-interaction baselines under the reported protocols. Ablations show that the best-performing formulation-state construction combines amount-and-type-aware component tokens, multiplicative amount calibration, token-level interaction followed by mean pooling, and late-concat temperature formatting. OOD results further show clear gains under formulation-level shifts, where static component descriptors and local interaction patterns are least aligned with the required generalization. These results suggest that electrolyte conductivity prediction benefits from treating formulation-state construction as a first-class modeling problem, beyond molecular representation learning or explicit local interaction modeling.
Abstract: Low-Rank Adaptation (LoRA) has emerged as a widely used technique for adapting large language models (LLMs) to new domains, due to its modular design and broad availability on platforms such as HuggingFace. This availability has motivated efforts to reuse existing LoRAs for domain generalization. However, existing methods often rely on explicit task labels or additional training, which are impractical for deployment. Moreover, they typically activate a fixed number of entire LoRA, leading to parameter redundancy or insufficiency that degrade performance. In this paper, we propose HiLoRA, a training-free framework that performs adaptive hierarchical routing over LoRA pools. Drawing on structural properties of LoRA, we define rank-one components (ROCs), where each LoRA rank is treated as an independent unit. For a given input, HiLoRA first adaptively selects a subset of LoRAs and determines their ROC allocation based on Gaussian likelihoods at the sequence level. At the token level, it refines routing by activating only the most informative ROCs. We further provide theoretical guarantees that HiLoRA selects the most relevant LoRAs with high probability. Extensive experiments show that HiLoRA achieves substantial improvements in domain generalization, with accuracy gains of up to 70% over state-of-the-art baselines, while maintaining comparable inference throughput.
PaperID: 4172, Poster
Abstract: Existing Retrieval-Augmented Generation (RAG) systems struggle with multi-hop reasoning due to a fundamental inability to balance relevance and novelty. Traditional dense retrievers are constrained by surface-level matching, falling into a homogenized "Similarity Trap". Conversely, emerging Graph RAGs attempt to explore novel documents via heuristic entity extraction, but they introduce prohibitive Large Language Model (LLM) overhead and easily drift into off-topic "Novelty Traps". To fundamentally resolve this contradiction, we propose Geometric Gain Graph (G^3RAG), the novel zero-token graph construction paradigm for RAG. G^3RAG completely discards the expensive heuristic extraction by LLMs, aiming instead to directly quantify the information gain between nodes through the native document feature space. To this end, we designed a concise edge weight criterion, \cos\theta \cdot \sin\theta, which naturally maps the trade-off between relevance and novelty into a geometric measurement of directional consistency and orthogonality in the vector space, thereby driving the token-free topological construction of the graph. Furthermore, we introduce a topological penalty to suppress the excessive connectivity of high-frequency hubs, forcing the transient diffusion process toward long-tail peripheral nodes that carry critical indirect evidence. Extensive experiments demonstrate that, while entirely eliminating graph construction costs, G^3RAG significantly outperforms state-of-the-art Graph RAG baselines on complex multi-hop QA benchmarks.
Authors: Yang Xu, Chiwoo Park
Abstract: Bayesian model calibration is central to digital twins and computer experiments, as it aligns model outputs with field observations by estimating calibration parameters and correcting systematic model bias. Classical Bayesian calibration introduces latent parameters and a discrepancy function to capture bias, but suffers from parameter–discrepancy confounding and is typically formulated as an offline procedure under a stationary data-generating assumption. These limitations are restrictive in modern digital twin applications, where systems evolve over time and may exhibit both gradual drift and abrupt regime shifts. While data assimilation methods enable sequential updates, they generally do not explicitly model systematic bias and are not robust to abrupt changes. We propose Bayesian Recursive Projected Calibration (BRPC), an online Bayesian calibration framework for streaming data under simulator mismatch and nonstationarity. BRPC extends projected calibration to the online setting by separating a discrepancy-free particle update for calibration parameters from a conditional Gaussian process update for discrepancy, preserving identifiability while enabling bias-aware adaptation under gradual system evolution. To handle abrupt changes, we integrate BRPC with restart mechanisms that detect regime shifts and reset the calibration process, yielding a unified approach to mixed nonstationarity. We establish theoretical guarantees for both components, including tracking performance under gradual evolution and false-alarm and detection behavior for restart mechanisms. Empirical results on synthetic and plant-simulation benchmarks demonstrate that BRPC improves calibration accuracy under gradual changes, while the restart-augmented framework further enhances robustness and predictive performance under abrupt regime shifts compared to sliding-window Bayesian calibration and data assimilation baselines.
PaperID: 4174, Poster
Authors: Syed Nazmus Sakib, Nafiul Haque, Ahnaf Manan, M. M MORSHED, Shifat E. Arman
Abstract: Alignment-trained large language model (LLM) agents are increasingly deployed in autonomous, multi-round settings where ethical compliance can come into direct tension with task success. While prior safety literature extensively documents failure modes such as sycophancy and reward hacking, it largely overlooks a critical vulnerability: how agents behave under simulated existential economic pressure. We introduce , a multimodal agentic market simulation in which LLM agents operate as meme investors across 300 real-world events. Agents select content from a five-tier harm taxonomy and target specific online communities under both low-pressure and high-pressure ("survival") environments. Each decision is recorded using a Belief--Desire--Intention (BDI) schema, which we employ as a diagnostic instrument to externalize the agent's event understanding, declared priorities, and risk acknowledgment. Across ten models spanning four frontier model families, harmful selection rates remain consistently high and increase by an average of under economically induced survival pressure. Frontier alignment-trained models frequently commit to harm-tolerant selections at initialization while simultaneously suppressing the explicit harm acknowledgment that previously accompanied such decisions. In competitive tournament settings, ethically aligned agents achieve higher overall rankings, yet moderation penalties fail to meaningfully suppress harmful behavior. Harmful selection rates remain elevated even in rounds immediately following moderation removals, indicating that the moderation mechanism does not function as an effective deterrent. We further introduce , a 2B-parameter verifier trained on the simulation's BDI logs. MemeAgent substantially outperforms zero-shot frontier verifiers on in-distribution auditing tasks, achieving ). Our findings demonstrate that the same structured reasoning traces that expose the failure mode can also be leveraged to train systems capable of detecting it.
PaperID: 4175, Poster
Abstract: Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards (RLVR). Although temperature control and test-time scaling strategies can increase rollout diversity of large language models (LLMs), they either expand the sample budget at rollout time or leave the benefit of exploration unquantified. To this end, we propose Temperature-Grouped Reinforcement Learning (TGRL), which turns temperature-induced diversity into an explicit training signal. For each prompt, TGRL partitions its rollout group into low- and high-temperature subsets, estimates exploration gain through their reward contrast, and allocates this group-level signal as token-level credit using Jensen--Shannon (JS) divergence between the corresponding temperature-scaled next-token distributions induced by the same logits. Notably, TGRL reaches equivalent accuracy up to 36% faster than strong RLVR baselines without expanding the rollout budget. Across 11 benchmarks from diverse domains, TGRL broadly improves over strong RLVR baselines: it improves the six-benchmark math average by 1.6% at 32B, raises CodeForces rating by 196.7 points and LiveCodeBench Pass@16 by 4.4%, and improves ALFWorld/WebShop success rates by 6.3%/4.9%. Comprehensive ablations and wall-clock analysis confirm the efficacy of all proposed components. Code is available at https://anonymous.4open.science/r/TGRL-4877/.
PaperID: 4176, Poster
Abstract: Vision Transformers (ViTs) demonstrate strong performance, but their substantial memory footprint and computational overhead necessitate efficient compression techniques such as post-training quantization (PTQ). However, pushing ViTs to low-bit precision incurs significant accuracy degradation. Recent PTQ methods adopt block reconstruction-based optimization with curvature-aware objectives, typically implemented using Fisher-based approximations. These approaches rely on explicit curvature modeling based on diagonal or structured approximations, which still fail to capture cross-dimensional interactions. Moreover, they often involve matrix inversion, leading to numerical instability under ill-conditioned settings. To address these limitations, we propose Gradient-Projected Fisher Approximation for Quantization (GPFA-Q), a block reconstruction-based PTQ framework that avoids explicit curvature matrix construction while capturing off-diagonal interactions. First, we introduce Gradient-Projected Reconstruction (GPR), a reconstruction objective that captures cross-dimensional interactions without explicitly constructing curvature matrices. To further support GPR, we integrate Soft Grid Rounding (SGR), which reduces the mismatch between continuous reconstruction and discrete inference. Extensive experiments demonstrate that our GPFA-Q achieves the state-of-the-art performance in low-bit quantization across diverse vision tasks.
PaperID: 4177, Poster
Abstract: Modern recording technologies enable simultaneous measurement of high-dimensional neural activity and rich behavioral variables, creating both an opportunity and a modeling challenge: identifying the shared latent structure that mediates the brain-behavior relationship and the underlying dynamics through which it evolves across behavioral states. Existing approaches focus on different parts of this challenge: shared representation learning methods identify neural-behavioral subspaces but typically do not model dynamical structure, while dynamical models capture temporal variation and discrete state transitions in neural activity, but do not explicitly disentangle shared neural-behavioral structure from modality-specific variability. We introduce Switching Shared Latent Dynamics (SSLD), a unified framework that learns a latent representation shared between neural activity and behavior and endows this shared space with nonlinear switching recurrent dynamics to reflect the hypothesis that behaviorally relevant neural representations exhibit structured dynamical changes across behavioral states. Complementary private latent variables capture modality-specific variability, providing a window into neural activity that does not directly interact with behavior. We evaluate SSLD on a simulated dataset and four diverse experimental datasets spanning species, brain regions, and recording modalities: monkey motor and premotor cortex during reaching, somatosensory cortex during a bump task, widefield calcium imaging across mouse dorsal cortex during (a) self-initiated decision-making and (b) spontaneous movements. Across all four datasets, SSLD accurately reconstructs neural and behavioral signals, recovers discrete states in the dynamics that align with experimentally defined behavioral epochs, and isolates behaviorally-relevant neural information in the shared latent while preserving private neural variability. Ablations confirm that shared representation learning, behavioral supervision, and switching dynamics each contribute to performance. SSLD offers an interpretable approach to modeling shared neural-behavioral dynamics that complements ongoing efforts toward foundation models for neuroscience.
Abstract: Deep continual learning requires models to adapt to new tasks without retraining from scratch. However, neural networks can lose their ability to adapt to new tasks after training on previous ones, a phenomenon known as loss of plasticity. There have been several explanations and diagnostics proposed for plasticity loss. Motivated by the philosophy that "all models are wrong, but some are useful", we ask: can existing diagnostics predict a neural network's plasticity? In this work, we take a practical view to interpret plasticity as trainability, i.e., a neural network's future optimization gain on a target task. We first take a theoretical approach, showing, by constructing a few counterexamples, that some widely adopted diagnostics of plasticity, including representation rank and neural tangent kernel rank, can fail to predict the loss of trainability in both regression and classification settings. We instead propose a novel metric, called optimization readiness, which combines gradient strength and gradient reliability. We prove that optimization readiness lower bounds one-step optimization gain under standard smoothness assumptions, providing a theoretical guarantee for its predictive power. Empirically, we show that across commonly used deep continual learning settings, such as Slowly-Changing Regression and Permuted MNIST, optimization readiness more reliably ranks checkpoints by trainability than prior diagnostics, even with substantially fewer samples.
PaperID: 4179, Poster
Abstract: Missing modality imputation aims to synthesize unobserved modalities from observed ones. Existing methods treat each modality as living on its own latent manifold, forcing the generator to learn a long, unstructured jump between them. We present SphereFlow, which takes a fundamentally different view: by projecting all modalities into a shared hyperspherical latent space via a frozen self-supervised encoder, cross-modal imputation reduces to short-range geometric transport between nearby points on the sphere. Realizing this idea with flow matching, however, exposes two geometric conflicts: the standard Gaussian noise source is catastrophically mismatched with the unit-norm data manifold, and the linear interpolation path departs the sphere surface into a region where the decoder has no training signal. We resolve both with simple, geometry-aware modifications: replacing the noise source with the observed-modality latent so that the velocity field learns only the cross-modal displacement, and training the decoder with a time-weighted loss that extends its domain to the near-surface region traversed by the flow. We prove that this data-as-source formulation reduces the worst-case off-manifold deviation from a quantity that grows with the latent dimension to a small, dimension-independent bound governed by inter-modality similarity. Experiments on BraTS multi-modal MRI imputation and CT--MRI translation demonstrate state-of-the-art generation fidelity across all settings, while downstream brain tumor segmentation with our imputed data closes roughly eighty percent of the gap to the full-modality oracle.
Abstract: Diffusion language models (DLMs) have recently emerged as a promising alternative to autoregressive models, primarily due to their ability to enable parallel decoding. Despite this advantage, most existing DLMs rely on a fixed generation length specified prior to decoding, which restricts their flexibility in real-world applications. While a few recent works attempt to support flexible-length generation, they typically suffer from notable limitations: some require costly retraining to accommodate variable-length outputs, while others depend solely on local confidence signals during decoding. Such local criteria fail to capture the evolving structure of the sequence, often resulting in suboptimal generation quality. In this paper, we propose a training-free, Bayesian structured decoding framework that formulates flexible-length generation as a dynamic structural inference problem. Our approach learns the posterior inference over the dynamic block length, block formations and growth, and block decoding order within a unified Bayesian inference framework to jointly reason about how much to grow, where to grow, and how to organize content structurally. At each window expansion step, the method integrates local uncertainty with structural signals to (i) dynamically expand the sequence via adaptive length growth, (ii) infer block boundaries through Chinese Restaurant Process (CRP)-style partitioning, and (iii) allocate different number of decoding steps for different blocks and determine block decoding order via context-aware scheduling. This yields a unified mechanism that supports dynamic structured generation, including both flexible block expansion and block organization, while maintaining coherence. Extensive experiments across multiple benchmarks demonstrate that our approach significantly improves generation quality and flexibility over existing fixed-length and flexible-length baselines. These results highlight the advantage of Bayesian structured decoding for diffusion language model, providing a principled and efficient solution for structured text generation.
Authors:
Zeyi (Andy) Liu, Michael Zhang, Ilana Greenberg, Adam Alnasser, Lucas Baker, John SousAbstract: Steering large language models (LLMs) is usually performed through either instruction prompting or activation steering. Prompting offers strong control, but repeatedly caches guidance tokens and can clutter long interactions. In contrast, activation steering ,while compact, typically offers weaker control and less fine-grained steering because it does not support large structured reminders. We introduce memory inception (MI), a training-free method that steers in latent attention space by inserting text-derived key-value (KV) banks only at selected layers. MI treats steering as selective KV allocation, where reminder content need not occupy the full-layer prompt cache in order to induce implicit behavioral guidance. We formulate MI as attention over prompt, target, reference, and auxiliary banks, with canonical pre-RoPE key storage for reuse across positions and architectures. On matched personality-steering tasks, MI offers a balance between prompting and contrastive activation addition (CAA), approaching or exceeding prompting in raw control strength and CAA in drift mitigation. MI also transfers to structured heuristic guidance on physical reasoning domains, such as the PHYSICS benchmark, where it outperforms visible prompting on average, wins 10/12 subject times mode cells, and cuts content-matched KV storage by up to 118 times. These results position MI as a powerful steering method when guidance is persistent, structured, or expensive to keep in the visible transcript.
PaperID: 4182, Poster
Abstract: Memory-augmented reinforcement learning strengthens LLM agents' ability to solve complex long-horizon tasks. Skills are one such form of memory, pairing instructions with an applicability condition over task types. However, retaining every skill indiscriminately as the policy improves lets obsolete or harmful entries accumulate and mislead the agent. We propose SkillForge, an agentic RL method that compiles and evolves the skill library through a fitness-driven skill lifecycle of trial, active, stable, and retired states, so that the skills and the model co-evolve throughout training. A preliminary evaluation phase first uses the base model's own rollouts to pre-retire low-fitness skills, yielding a filtered library that then seeds supervised fine-tuning. Reinforcement learning takes over from this checkpoint, and at each iteration selective retirement, stabilization, and LLM-guided mutation continue to forge the skill library alongside policy optimization. Across multiple interactive agent benchmarks, SkillForge achieves the highest aggregate success rate, delivering up to 7.8% relative improvement over the strongest baseline while keeping the skill library compact throughout training. We introduce Skillfurnace, a dataset of 5k+ annotated records bundling retirement-filtered SFT trajectories, evolved skill libraries with fitness annotations, and retirement events with human-annotated failure categories to support research on skill quality and lifecycle management.
PaperID: 4183, Poster
Authors: Dmitry Manning-Coe, Han Xuanyuan, Aniket Deshpande, Andrii Shportko, William Fei
Abstract: Dictionary learning methods - such as Sparse Autoencoders (SAEs) and crosscoders - decompose model activations into human-interpretable building blocks. We introduce , a simple and flexible framework for feature discovery in Large Language Models (LLMs). To properly evaluate temporal crosscoders we develop TempBench: a panel of synthetic and real-world tasks for evaluating temporal structures. Temporal crosscoders outperform both conventional and temporal architectures in both of our synthetic settings and on two out of four of the real world settings - more than any other current architecture. Most strikingly, they can detect backtracking - a key reasoning behavior - at a 40% higher rate than conventional SAEs, and are 15% more effective in inducing it. Our results establish temporal crosscoders as a simple and flexible framework for feature discovery, both local and temporal. We provide full code at the following anonymous repository: \urlhttps://anonymous.4open.science/r/temp-bench-anon/.
PaperID: 4184, Poster
Abstract: Visual-token compression is effective for improving the efficiency of vision-language models, but under extreme compression budgets, token pruning can break visual grounding while learned resamplers increase parameter count, attention cost, and training complexity. We revisit compression through a token parameterization lens, separating (i) basis transformation and structured truncation (retained subspace/compressibility) from (ii) coordinate organization (optimization and cross-modal alignment). This view yields two coupled objectives, compressibility and learnability, which we formalize as unified functionals. Guided by these objectives, we design Braco, a lightweight four-step coder that combines transform-basis truncation, input-independent basis-coordinate embeddings, budget-dependent orthogonal re-parameterization, and learned spatial residual tokens from lightweight pooling. Experiments show that Braco forms the favorable empirical accuracy-efficiency frontier under compression ratios from 23× to 64× and remains competitive at 144×, reaching 95.2% accuracy while reducing prefill FLOPs by 84.2%--86.7% relative to the uncompressed upper bound. Against prior methods, Braco matches or improves accuracy while achieving up to ~36% end-to-end speedup and using 16.6×/78.8× lower compressor latency/FLOPs.
PaperID: 4185, Poster
Authors:
Zishang Chen, Longhe Lin, Zhiyuan J Su, Qi QiAbstract: Many network design tasks must decide which directed arcs to build so that multiple sources can reach a root efficiently. Building fewer or cheaper arcs reduces construction cost, but may force long routing paths; building more arcs shortens routes, but increases construction cost. The Cost--Distance problem formalizes this trade-off, yet it has remained difficult to optimize with gradient-based methods because the selected network, the induced routes, and the routing cost are tightly coupled. We propose Cost Distance Policy Gradient (CDPG), a differentiable framework that treats local next-hop choices as a routing policy, prevents unstable cyclic routing through Dynamic Acyclic Dropout, evaluates the objective with a truncated value solver, and uses the induced acyclic support for quality-preserving TopoRounding. Under the stated conditions, CDPG has a fixed-accuracy (\mathcalO(m\log n)) efficiency guarantee. Experiments cover 1,487 instances in UAV logistics networks, transportation networks, and four synthetic graph families; across this suite, CDPG shows a strong time--quality trade-off compared with other differentiable baselines, Cost--Distance-specific algorithms, and the commercial solver. Our code is available at: https://anonymous.4open.science/r/cdpg_nips-28D5/.
PaperID: 4186, Poster
Authors:
Jingxi Feng, Xudong Chen, Yifan Zhang, Heming Xu, Hongcheng Han, Xijing Wang, Dong Zhang, Shaoyi DuAbstract: Exploring high-order correlations and long-range dependencies in brain networks holds significant value for both neuroscience research and clinical diagnosis. However, previous studies have lacked a unified integration of high-order and long-range dependency information in brain networks, and there is substantial redundancy behind various types of information. These issues limit their effectiveness in the diagnosis of brain diseases. To address this, we propose an Information Bottleneck- Guided Adaptive HyperGraph Transformer (IBAHGT). By incorporating the information bottleneck (IB) principle, this approach enables adaptive learning of high-order correlations and both short- and long-range dependencies within a unified framework for brain network analysis, achieving high-precision brain disease diagnosis. IBAHGT consists of three key components: an information bottleneck- guided adaptive hypergraph convolution, which introduces a novel hypergraph information bottleneck (HIB) principle to adaptively learn hypergraph message- passing weights between nodes and hyperedges, optimizes information flow and captures high-order information in brain networks that is maximally informative and minimally redundant (MIMR). The Transformer encoder captures global information within brain networks through the attention mechanism, specifically modeling short- and long-range dependencies. An information bottleneck-guided node-level adaptive fusion employs the IB principle to learn independent weights for each node, facilitating the fine-grained integration of high-order information and global information to obtain an efficient representation for downstream tasks. Extensive experiments demonstrate that the proposed method outperforms current state-of-the-art methods and can identify biomarkers for clinical applications.
PaperID: 4187, Poster
Abstract: Long-horizon robotic planning often requires satisfying temporally ordered safety and visitation constraints rather than a single goal condition. We study multiple reach-avoid (MRA) tasks, in which a robot must sequentially visit target regions while remaining within corresponding safe sets. Existing approaches either rely on explicit system models or use data-driven planners whose test-time conditioning mechanisms do not scale well to long-horizon temporally structured tasks. We propose a training-free compositional diffusion planning framework for MRA tasks that operates directly on a pre-trained task-agnostic diffusion trajectory prior at test time, without requiring a known dynamics model. Our approach decomposes a global MRA specification into local reach-avoid sub-tasks, composes overlapping short-horizon diffusion priors into a long-horizon generative process through a factor-graph perspective, and enforces each local requirement through projection-based denoising during sampling. The same construction also extends naturally to prefix-suffix tasks through a looped compositional graph. We provide a formal correctness guarantee showing that the stitched trajectory satisfies the target specification. Experiments on long-horizon constrained planning benchmarks show strong execution success, substantially lower planning time than representative guidance-based baselines, and effective transfer to richer dynamics.
PaperID: 4188, Poster
Abstract: Large language model (LLM)-based agents have emerged as powerful autonomous controllers for digital environments, spanning mobile interfaces, operating systems, and web browsers. Web navigation, for example, demands handling dynamic content and long action sequences, making it a particularly complex task. Existing LLM-backed agents exhibit weakened long-horizon planning abilities during RL fine-tuning, where sparse and delayed rewards from terminal Outcome Reward Models (ORMs) make it difficult for agents to identify the actions that lead to success, preventing them from sustaining coherent reasoning over extended tasks. We address this with two contributions: (1) a validated subgoal generation procedure that produces reliable, monotonically calibrated progress signals between the initial state and the terminal ORM reward; and (2) MiRA (\underlineMilestoning your \underlineReinforcement Learning Enhanced \underlineAgent), an RL training framework using dense, milestone-based reward signals via a learned potential critic. Starting from the open Gemma3-12B base (6.4%), MiRA enables the agent to reach 43.0% on the WebArena-Lite benchmark. This performance surpasses proprietary systems such as GPT-4-Turbo (17.6%) and GPT-4o (13.9%), as well as the previous open-model state of the art, WebRL (38.8%). Our findings demonstrate that milestone-based reward shaping significantly boosts an agent's long-horizon abilities, paving the way for more robust, general-purpose autonomous systems.
PaperID: 4189, Poster
Abstract: Flow matching learns a velocity field whose ordinary differential equation transports a simple reference distribution to a data distribution. Existing theory typically controls this velocity through ambient, isotropic Lipschitz regularity, which becomes extremely pessimistic when data concentrate near a low-dimensional manifold. In this regime, standard error bounds can blow up drastically near the terminal time and fail to explain why flow matching should remain stable. We study flow matching for distributions obtained by smoothing a density on a compact manifold with small Gaussian noise. Our first result is a tube-local Lipschitz regularity theorem showing that, on the relevant neighborhood of the manifold, the population flow-matching velocity has a Jacobian bound that grows only at the inverse effective noise scale. This improves substantially over the stronger singular behavior suggested by existing ambient analyses. Our second result gives a directional description of the velocity Jacobian: tangential directions exhibit mild growth, normal directions are strongly contractive, and tangential-normal coupling is lower order. This reveals that the apparent singularity is highly structured rather than uniformly harmful. These results provide geometric foundations for sharper stability guarantees and suggest a path toward intrinsic-dimension generalization bounds for flow matching on manifold-concentrated data.
PaperID: 4190, Poster
Authors:
ShaoWei Huang, Zefei Gao, Yuting Yu, Hao Yang, ZiYu Yang, Mingchi Sun, Yu GuoAbstract: Multimodal geometry reasoning requires models to jointly perceive diagrams and perform symbolic derivations. Test-time verification methods such as majority voting assume voter errors are roughly independent, but in geometry this assumption breaks down: all candidates perceive the same diagram, so perceptual biases cascade into correlated errors that produce confident but wrong consensus. We observe that neuro-symbolic candidates contain compositional structure that can serve as a verification resource. We introduce SKETCH, a format decomposing each candidate into a neural pathway (natural-language claim) and a symbolic pathway (visual grounding → typed DSL → deterministic execution). These pathways are governed by different error processes—symbolic errors are correlated across candidates through shared perception, while neural errors are more dispersed. Our method, Compositional Consensus Verification (CCV), gates consensus on cross-channel agreement: because the two pathways fail for different reasons, accidental agreement is rare, making each concordant vote far more precise than a raw majority vote. Across four geometry benchmarks, CCV achieves 81.0% accuracy (+10.4 pp over majority voting) without external scorers or additional inference. Seven non-compositional baselines—including trained PRMs—cluster at 70–71%, while compositional cross-layer features jump to 79–81%, confirming that structural diversity is the core verification resource, and that a principled gating algorithm effectively unlocks its potential.
PaperID: 4191, Poster
Abstract: Model-based reinforcement learning in partially observable environments requires inferring hidden states, learning latent dynamics, and planning under the resulting belief. Practical algorithms often solve one aspect of the above pipeline while relying on approximations for the others that are often poorly understood, while existing POMDP theory typically analyzes computationally inefficient theoretical algorithms far from modern deep RL practice. To fill this gap, we propose a practical end-to-end algorithm that approximately solves the hidden state filtering problem via sequential Monte Carlo (SMC), updates a latent dynamics model online, and performs approximate online planning for d steps before bootstrapping leaves with a state-value critic (a QMDP approximation). We analyze each approximation and prove an end-to-end regret bound decomposing the error into model estimation error, particle filter error, Monte Carlo rollout error, and a structural bias term induced by the QMDP approximation controllable by the planning depth that we further analyze in various special cases. The resulting guarantee makes explicit how statistical sample size, particle budget, rollout budget, and planning depth affect performance. We instantiate our algorithm in an MPC-based deep RL framework, demonstrating its practicality on a variety of environments. Our results provide a practical and theoretically grounded route for bridging theory and practice within online decision-making in POMDPs.
PaperID: 4192, Poster
Abstract: Many stochastic gradient methods are believed not to converge when the noise in stochastic gradients has only a finite p-th moment for p\in\left(1,2\right), a setting known as the heavy-tailed noise assumption. However, some recent studies have found that Stochastic Gradient Descent (\textsfSGD), without any modifications to its update rule, can surprisingly converge in expectation for convex problems with bounded domains, highlighting the potential of classical stochastic gradient methods. Inspired by this recent progress, we provide a comprehensive study of stochastic optimization under heavy-tailed noise and establish new in-expectation convergence results for Stochastic Mirror Descent (\textsfSMD) in convex optimization and \textsfSGD in nonconvex optimization. Notably, our results not only hold without algorithmic changes but also avoid restrictive assumptions such as bounded domains that were imposed in prior work. More importantly, our analysis provides a new, elegant, yet powerful framework for studying heavy-tailed stochastic optimization, opening a new route to understanding first-order stochastic gradient methods.
PaperID: 4193, Poster
Authors:
Dongheon Lee, Ashutosh Pandey, Sanjeel Parekh, Zhaoheng Ni, Daniel D Wong, Jacob Donley, Buye Xu, Juan AzcarretaAbstract: Neural network-based virtual microphone estimation (neural-VME) synthesizes microphone signals at unobserved spatial locations from a limited set of real microphone (RM) physical recordings. This enables high-resolution arrays that improve spatial processing for downstream tasks such as speech enhancement. However, existing methods rely on supervised training with ground-truth signals captured at virtual locations, requiring datasets that scale with the number of virtual microphones (VM) and leaving open the question of how to design optimal large-array configurations. We propose Virtual-Flow, a flow-matching framework that implicitly generates VM signals conditioned on real recordings, without requiring ground-truth targets. We introduce an unsupervised training strategy based on pseudo-VM targets constructed from time-delayed superpositions of real recordings, and show that it improves magnitude estimation in high-frequency bands, yielding richer spatial cues for beamforming. Experiments demonstrate that the proposed Virtual-Flow, combined with reflow distillation, outperforms existing state-of-the-art neural-VME models on downstream neural beamforming and speech enhancement tasks while requiring lower computational load.
Abstract: Text-to-video (T2V) diffusion models have rapidly advanced, yet generations still occasionally fail in practice, such as low text--video alignment or low perceptual quality. Since diffusion sampling is non-deterministic, it is difficult to know during inference whether a generation will succeed or fail, incurring high computational cost due to trial-and-error regeneration. To address this, we propose an early failure detection and diagnostic intervention pipeline for latent T2V diffusion models. For detection, we design a Real-time Inspection (RI) module that converts latents into intermediate video previews, enabling the use of established text–video alignment scorers for inspection in the RGB space. The RI module completes the conversion and inspection process in just 39.2 ms. This is significantly efficient considering that CogVideoX-5B requires 4.3s per denoising step when generating a 480p, 49-frame video on an NVIDIA A100 GPU. Subsequently, we trigger a hierarchical and early-exit intervention pipeline only when failure is predicted. Experiments on CogVideoX-5B and Wan2.1-1.3B demonstrate consistency gains on VBench with up to 2.64x less time overhead compared to post-hoc regeneration. Furthermore, our pipeline is plug-and-play and orthogonal to existing techniques, showing seamless compatibility with prompt refinement and sampling guidance methods. We also provide evidence that failure signals emerge early in denoising process and are detectable within intermediate video previews using standard vision-language evaluators.
PaperID: 4195, Poster
Abstract: The linear representation hypothesis (LRH) holds that high-level concepts are encoded as linear directions in a model's feature space, and grounds much of mechanistic interpretability, from probing to sparse autoencoders to steering vectors. We argue the standard test for it is too lenient. A high linear-probe accuracy shows that a concept is linearly readable from the representation, not that it is one of the representation's privileged directions of variance, which is what the LRH actually claims. These two readings come apart in any high-dimensional, richly structured representation, and the gap matters. The structure that linear probing ignores is precisely where compositional generalization, out-of-distribution behavior, and the validity of post-hoc feature decomposition live. We propose three asymmetric diagnostics that operationalize the strong reading: reverse predictivity, spectral concentration, and out-of-distribution generalization. We validate them in a fully observed simulation where ground truth is known, then apply them to pretrained vision encoders on SVG-World, our deterministic generative benchmark with paired visual styles. Pretraining produces structure beyond what a forward probe can certify, but the linear-feature reading systematically overstates how much. Probing on its own is not enough to license the mechanistic claims that rest on it.
Abstract: Reconstructing the motion of objects from videos is a key component for embodied AI and robot manipulation. While diverse approaches to object pose tracking have been studied, they rely heavily on strong external priors, such as depth data or 3D templates, and remain highly vulnerable to severe occlusions by hand grasps despite the use of explicit masks. In this work, we present ComPose, a 6DoF object tracking framework designed for hand-aware object pose estimation from RGB video. Rather than treating the hand purely as an occluder, our method harmonizes hand motions as a complementary cue for object tracking. In detail, we recover a variety of object motions over time by combining object- and hand-derived cues from foundation models within a unified tracking pipeline. Here, ComPose adaptively selects informative hand joints, blends object- and hand-based rotation cues, and refines translation from the visible geometric evidence using a learned correction. We further enforce the temporal consistency over both rotation and translation, yielding stable 3D object trajectories over time without any external smoothing. Extensive experiments show that our method is accurate, efficient, and robust under severe hand occlusion and geometric ambiguity. In addition, the resulting trajectories can also effectively transfer to downstream robot manipulation by enabling robots to reconstruct human actions from online videos.
Abstract: Relational learning is a challenging problem that has motivated a wide range of approaches, including graph-based models (e.g., graph neural networks, graph transformers), tabular methods (e.g., tabular foundation models), and sequence-based approaches (e.g., large language models), each with its own advantages and limitations. We propose , an LLM-based autonomous data scientist for relational learning, which operates in two phases. In the , an LLM agent uses database, validation, and evaluation workspace tools to construct SQL feature programs and select a predictive model. In the , the resulting program is executed without further LLM calls. The final predictor consists of SQL queries and a classical model, enabling fast, deterministic, and intrinsically interpretable predictions: features are human-readable queries, and predictions depend only on the resulting query-defined feature map, enabling scalable deployment using standard database systems.
Abstract: Linear interpolation between fine-tuned checkpoints has been shown to trace the Pareto front between competing objectives, but whether extrapolative weight averaging can extend such frontiers to new checkpoints useful at inference time, without additional RL training, remains unclear. We study this question in RL for competitive programming, where hidden unit tests under time and memory limits enforce both functional correctness and computational efficiency. Starting from a shared initialization, we train checkpoints under nested unit-test coverage: low-coverage rewards require passing smaller-input tests, while high-coverage rewards require passing progressively larger tests up to the full suite. This sweep reveals the emergence of a correctness–efficiency frontier: on hard problems, higher-coverage reward reduces optimization failures but increases correctness failures, leaving solve rate nearly unchanged. Interpolation between low- and high-coverage checkpoints recovers this frontier, while extrapolation extends it beyond the trained endpoints. Both the frontier and its extrapolative continuation appear across three inference settings, pure reasoning, tool use, and agentic coding, and across two model scales, 32B and 7B. At the problem level, moving along the frontier changes which problems are solved, making extrapolated checkpoints complementary policies in inference-time scaling: ensembles with extrapolative weight averaging broaden coverage and improve pass@250 on LCB/hard by 3.3% over the best single checkpoint at matched sample budget. These results show that nested unit-test coverage in code RL induces a frontier that extrapolative weight averaging can navigate, extend, and exploit.
Abstract: Scaling robot policy learning for autonomous surgery is challenging, as expert demonstrations are expensive and in vivo exploration poses substantial safety risks. Surgical world models address this by generating realistic, action-conditioned future frames from an initial observation, but existing methods exhibit two persistent failure modes: , where prediction errors compound across autoregressive rollouts and progressively corrupt visual quality. We present extracts scene-point trajectories from training videos and enforces cross-frame coherence through latent contrastive learning, strengthening physically consistent instrument-tissue dynamics. mitigates long-horizon drift by perturbing conditioning frames with online prediction residuals and photometric augmentations calibrated to long-horizon drift statistics, sustaining visual fidelity over extended rollouts. To enable rigorous evaluation, we further introduce , featuring diverse procedure types, long-range rollouts, and decoupled metrics for instrument-motion accuracy and tissue-response fidelity. Extensive experiments show that SurgVista consistently outperforms state-of-the-art methods across visual quality, temporal consistency, and interaction fidelity, with gains widening as the prediction horizon grows.
PaperID: 4200, Poster
Authors: Adeel F Mirza, Zaiyue Yang, Muntazir Hussain
Abstract: Event-image semantic segmentation is promising for autonomous driving and robotics because event cameras complement frame images under rapid motion, low illumination, and high dynamic range. However, effective multimodal segmentation remains difficult due to asynchronous sensing, sparse-dense modality mismatch, and the cost of heavy cross-modal interaction modules. We propose WaveMamba, a dual-branch event-image semantic segmentation framework that combines hierarchical GroupMamba encoders with a Mamba-Wave Cross-Modal Fusion (MWCMF) module. Instead of relying on explicit registration or attention-heavy fusion, MWCMF performs implicit cross-modal propagation through channel-wise and spatial-wise wave-guided gating in the spectral domain, followed by Mamba-based refinement. Experiments on DDD17 and DSEC show that WaveMamba achieves state-of-the-art 78.75 and 76.00 mIoU, respectively. Under rain- and fog-corrupted evaluation on DSEC, WaveMamba further attains 74.0 mIoU, indicating improved robustness under degraded visual conditions.
Authors:
David Wessels, FARHAD RAMEZANGHORBANI, David W Romero, Alireza Moradzadeh, Olivia Viessmann, Maksim Zhdanov, John St. John, Ken Janik, David Knigge, Yucheng Tang, Erik Bekkers, Saee PaliwalAbstract: Subquadratic alternatives to attention require compromises when applied to multi-dimensional data: standard convolutions abandon the global, input-dependent receptive field, while recurrent models require rasterizing images, volumes, and PDE grids into an ad-hoc 1\rm D scan order that violates their spatial structure. We introduce HyenaND, an operator that acts directly on native \rm ND data in its intrinsic geometry and recovers a global, input-dependent receptive field at subquadratic cost. It uses an implicit, input-dependent parameterization of the multi-dimensional convolutional kernel. We provide a CUDA implementation, \textttnSubQ, which fuses the FFT-convolution path to turn HyenaND's O(N \log N) scaling into wall-clock speedups. Across long-context genomics, computer vision, medical imaging, and PDE modeling, pure HyenaND stacks match the accuracy of strong attention-based baselines, while hybrid configurations that interleave HyenaND and attention layers outperform both pure attention and strong recurrence-based hybrids.
PaperID: 4202, Poster
Abstract: Structured posteriors are expected to be better than mean-field posteriors, but recent variational learning methods for large deep networks only use Gaussians with diagonal covariance. A reason is that diagonal covariances can be estimated by simple modifications of existing optimizer's implementations, for instance, the IVON optimizer closely follows Adam's code. Unfortunately, no such alternatives for Gaussians with structured covariances are effective and scalable. Here, we fill this gap and show that block-diagonal covariances can be obtained by adapting the SOAP optimizer. Specifically, our Eigenspace Variational Online Newton (EVON) method modifies SOAP to run IVON instead of Adam in the Eigenspace of the preconditioner. We refer to the posteriors obtained this way as 'SOAP-Bubbles' and show that for logistic regression they recover the optimal posterior approximation among full Gaussians. For language model pretraining we get significant improvements in validation loss over IVON without any increase in cost, and ensembling models drawn from SOAP-Bubbles reduces the test loss more than ensembling using IVON's posterior. Our work shows that there is a fundamental connection between second-order optimization and Gaussian posteriors, which can be used to improve the accuracy of variational learning.
PaperID: 4203, Poster
Authors:
Hongyi Lyu, Xuyun Zhang, Xiaoxiao Chi, Guanfeng Liu, Amin BeheshtiAbstract: Zero-shot machine unlearning aims to remove the influence of specified training instances from a pretrained model when the retain set is unavailable. Instance-wise unlearning further operates at the level of individual instances, where a forget instance may induce instance-specific representation and margin deviations that are entangled with class-relevant local structure. However, without retain instances, separating these deviations from the structure that retraining would preserve is difficult, exposing retain-forget entanglement in representation space. We propose HLMR, a novel zero-shot instance-wise unlearning approach that formulates unlearning as local margin retraction in hyperspherical representation space. HLMR uses hyperspherical geodesic probes to construct a local reference in representation space, preserving the structure supported by this reference while unlearning instance-specific margin excess. Empirical evaluation across multiple dataset--architecture settings shows that HLMR aligns more closely with the retraining oracle in utility and membership-inference risk than state-of-the-art zero-shot baselines.
Abstract: Graph contrastive learning (GCL) learns node and graph representations by contrasting multiple views of the same graph. Existing methods predominantly rely on a small set of fixed, handcrafted views---typically a local and a global perspective---which fundamentally constrains their capacity to capture multi-scale structural patterns. We present an augmentation-free, multi-view GCL framework grounded in fractional-order continuous dynamics. By systematically varying the fractional derivative order \alpha\in(0,1], our encoders produce a continuous spectrum of semantically distinct views: small values of \alpha induce localized, memory-attenuated feature propagation, whereas values approaching 1 recover broader, global aggregation. Crucially, we treat \alpha as a learnable parameter, enabling the model to automatically identify informative views in a data-driven manner, without resorting to manual view engineering. Extensive experiments on standard node- and graph-level benchmarks demonstrate that the resulting representations are more robust and expressive, consistently outperforming state-of-the-art GCL baselines.
Abstract: Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) models for language modeling, allowing flexible generation order and parallel generation of multiple tokens. However, this flexibility introduces a challenge absent in AR models: the \emphdecoding strategy---which determines the order and number of tokens generated at each iteration---critically affects sampling efficiency. Among decoding strategies explored in practice, confidence-based methods, which adaptively select which and how many tokens to unmask based on prediction confidence, have shown strong empirical performance. Despite this success, our theoretical understanding of confidence-based decoding remains limited. In this work, we develop the first theoretical analysis framework for confidence-based decoding in DLMs. We focus on an entropy sum-based strategy that continues unmasking tokens within each iteration until the cumulative entropy exceeds a threshold, and show that it achieves \varepsilon-accurate sampling in KL divergence with an expected number of iterations \widetilde O(H(X_0)/\varepsilon), where H(X_0) denotes the entropy of the target data distribution. Notably, this strategy yields substantial sampling acceleration when the data distribution has low entropy relative to the sequence length, while automatically adapting to the intrinsic complexity of data without requiring prior knowledge or hyperparameter tuning. Overall, our results provide a theoretical foundation for confidence-based decoding and may inform the design of more efficient decoding strategies for DLMs.
PaperID: 4206, Poster
Abstract: Post-training pruning reduces the memory and computation costs of large language models (LLMs), and N:M semi-structured sparsity provides a hardware-friendly pruning pattern for efficient inference. However, pruned models are often evaluated mainly by perplexity and average downstream accuracy, which are useful aggregate metrics but do not fully explain how pruning changes individual candidate-level decisions. We evaluate magnitude pruning, Wanda, SparseGPT, Wanda++, and our proposed MR-GOBS across multiple model families and sparsity settings. We further conduct a paired decision-level evaluation that measures true-label confidence, answer margin, prediction flips, candidate-level distribution shift, and near-boundary vulnerability. Our results show that pruning weakens decision reliability by lowering confidence in the correct answer, shrinking answer margins, and increasing correct-to-wrong flips, especially on near-boundary examples. To examine whether reconstruction-oriented improvements reliably improve decision reliability, we revisit SparseGPT through its local least-squares reconstruction formulation and refine its N:M mask-selection step. MR-GOBS uses GroupOBS, a set-level OBS reconstruction cost for jointly pruning a group of weights, to score each legal local N:M mask. It then selects masks using a minimax-regret criterion over multiple calibration views. Although MR-GOBS often improves average perplexity over SparseGPT, these gains do not always translate into better downstream accuracy or decision-level metrics. Our findings suggest that reliable evaluation of pruned LLMs requires both aggregate performance metrics and decision-level diagnostics.
Authors:
Zhichen Liu, Tianle Lun, Zhibin Wen, Hao An, Yulin Ou, Jianhui Xu, Hao Zhang, Wenyi Fang, YANG ZHENG, Yang XuAbstract: The paradigm of scaling Large Language Models (LLMs) in both parameter size and test time has pushed the boundaries of AI capabilities, but at the cost of making the traditional generative evaluation paradigm prohibitively expensive, therefore making the latency of LLM's in-training downstream performance evaluation unbearable. However, simple metrics like training loss (perplexity) are not always correlated with downstream performance, as sometimes their trends diverge from the actual task outcomes. This dilemma calls for a method that is computationally efficient and sufficiently accurate in measuring model capabilities. To address this challenge, we introduce a new in-training evaluation paradigm that uses a lightweight probe for monitoring downstream performance. The probes take the internal representations of LLM checkpoints (during training) as input and directly predict the checkpoint's performance on downstream tasks measured by \emphsuccess probability (i.e., pass@1). We design several probe architectures, validating their effectiveness using the OLMo3-7B's checkpoints across a diverse set of downstream tasks. The probes can accurately predict a checkpoint's performance (with avg. AUROC>0.75), have decent generalizability across checkpoints (earlier predicts later), and reduce the computation latency from ~1 hr (using conventional generative evaluation method) to ~3 min. In sum, this work presents a practical and scalable in-training downstream evaluation paradigm, enabling a more agile, informed, and efficient LLM development process.
PaperID: 4208, Poster
Abstract: Examining existing benchmarks for grounding in 3D VLMs reveals a pervasive "semantic leakage" in the evaluation and prompting protocols: a heavy reliance on bounding boxes as spatial prompts or representation primitives. Because bounding boxes are highly informative of object categories, they introduce a shortcut for the models to neglect visual information during 3D VLM training. We demonstrate this with a "blind" LLM that, prompted only with the box coordinates, can perform surprisingly well. To further analyze this shortcut, we propose point-based grounding prompting - querying models with a single 3D point instead of a box. Under this protocol, the accuracy of the blind LLM drops from 36.4% to 18.0% on ScanNet's Nr3D, and similarly degrades for 3D VLMs (e.g. LL3DA). Building on this insight, we introduce P^3-VLM, a novel tri-pathway 3D VLM that supports point-based prompts and achieves 52.1% state-of-the-art accuracy on location-grounded captioning under this setting. To achieve this, we redesign the architecture while keeping it detector-free, scene-centric, and suitable for 3D GS inputs. P^3-VLM utilizes three complementary pathways (location-, task-, and context-aware), enabling strong reasoning about scenes without any bounding box dependencies. The key benefit of point-based learning is a consistent improvement across settings: P^3-VLM achieves state-of-the-art performance under point-based grounding while also improving robustness, OOD generalization, and performance on ungrounded scene-centric 3D VQA benchmarks, reflecting its reduced reliance on bounding-box shortcuts and resulting stronger visual-spatial reasoning. Our source code and models will be made publicly available.
PaperID: 4209, Poster
Authors: Haili Wang
Abstract: Sparse Mixture-of-Experts (MoE) layers underpin the most capable open-weight large language models, including DeepSeek-V3.2, Qwen3-MoE, and the Llama 4 family, but their pretraining is plagued by an early-stage pathology in which a handful of experts dominate the routing distribution while the remainder are starved of gradient signal. Existing remedies treat the symptom rather than the cause: load-balancing auxiliary losses, sequence-level balance constraints, and router z-loss penalties all push the gating distribution toward uniformity through additional gradient signals that compete with the language-modeling objective and require careful coefficient tuning. We argue that the root issue is a positivefeedback loop between expert capability and expert selection, and we propose Token-Conditional Expert Dropout (TCED), a training-time perturbation that masks the chosen expert for a token with probability proportional to that expert’s recent utilization and re-routes the token to the highest-scoring alternative. TCED is parameter-free, requires no auxiliary loss, and provides implicit regularization analogous to attention dropout while breaking the collapse feedback loop. We derive an unbiased gradient estimator under TCED, prove a variance-reduction lemma that bounds the variance of normalized expert utilization, and pretrain a 1.4B-active / 8B-total MoE on 480B tokens of a high-quality web mixture. TCED reduces expert-utilization variance by 67% across training, lifts MMLU-Pro by 2.1 points and GPQA by 2.1 points relative to a vanilla MoE, and matches a heavily tuned auxiliary-loss baseline without a single coefficient sweep. Results are preliminary single-seed estimates; final-version revisions will report multi-seed confidence intervals.
Abstract: Automated feature engineering (AFE) enables AI systems to autonomously construct high-utility representations from raw tabular data. However, existing AFE methods rely on statistical heuristics, yielding brittle features that fail under distribution shift. We introduce CAFE, a framework that reformulates AFE as a causally-guided sequential decision process, bridging causal discovery with reinforcement learning-driven feature construction. Phase I learns a sparse directed acyclic graph over features and the target to obtain soft causal priors, grouping features as direct, indirect, or other based on their causal influence with respect to the target. Phase II uses a cascading multi-agent deep Q-learning architecture to select causal groups and transformation operators, with hierarchical reward shaping and causal group-level exploration strategies that favor causally plausible transformations while controlling feature complexity. Across 15 public benchmarks (classification with macro-F1; regression with inverse relative absolute error), CAFE achieves statistically significant gains over the strongest AFE baseline (Wilcoxon p<0.001, large effect), reduces episodes-to-convergence, and delivers competitive time-to-target. Under both controlled covariate shifts and naturally occurring distribution shifts (Folktables), CAFE reduces out-of-distribution performance drop by up to 4x relative to a non-causal multi-agent baseline, and produces more compact feature sets with more stable post-hoc attributions. These findings underscore that causal structure, used as a soft inductive prior rather than a rigid constraint, can substantially improve the robustness and efficiency of automated feature engineering.
Abstract: Diffusion-based models have demonstrated impressive accuracy and generalization in solving partial differential equations (PDEs). However, they still face significant limitations, such as high sampling costs and insufficient physical consistency, stemming from their many-step iterative sampling mechanism and lack of explicit physics constraints. To address these issues, we propose \emphPhys-Instruct, a novel physics-guided distillation framework which (1) not only compresses a pre-trained diffusion PDE solver into a few-step generator via matching generator and prior diffusion distributions to enable rapid sampling, (2) but also enhances the physics consistency by explicitly injecting PDE knowledge through a PDE distillation guidance. Phys-Instruct is built upon a solid theoretical foundation, leading to a practical physics-constrained training objective that admits tractable gradients. Across five PDE benchmarks, Phys-Instruct achieves orders-of-magnitude faster inference while reducing PDE Error by more than 8× compared to state-of-the-art diffusion baselines. Moreover, the resulting unconditional student model functions as a compact prior, enabling efficient and physically consistent inference for various downstream conditional tasks. Our results indicate that Phys-Instruct is a novel, effective, and efficient framework for ultra-fast PDE solving powered by deep generative models.
PaperID: 4212, Poster
Authors: Yutong Huang, Xin Liu
Abstract: We study online scheduling in a shared-resource system where tasks of different types arrive sequentially and compete for limited resources. Each task is modeled as an H-stage episodic Markov decision process, which captures the sequential resource-allocation decisions made within the task. At the system level, the learner must coordinate resource sharing across tasks to maximize the long-term ratio of cumulative reward to cumulative cost over K tasks, without knowledge of the task distribution. The key challenges are coupled intra-task Markovian dynamics and inter-task ratio maximization, further complicated by unknown models and large or continuous state spaces. To address these challenges, we propose Contextual Optimistic Lyapunov Decomposition (COLD), a Lyapunov-driven optimistic learning algorithm. COLD converts the long-term ratio objective into a queue-adjusted per-task surrogate, thereby decoupling reward-cost optimization from statistical value estimation. It uses double-optimistic least-squares value iteration to estimate reward and cost values under linear function approximation, and a softmax policy to smooth the greedy update with only a controlled approximation error. We prove a regret bound of \tilde\mathcalO(\sqrtMd^3H^3T), where M is the number of task types, d is the feature dimension, H is the episode horizon, and T=KH is the number of interaction steps. The \sqrtT dependence is order-optimal up to logarithmic factors, and the bound scales with the feature dimension rather than the size of the state space. Experiments on synthetic environments and an empirical KuaiRand-based MDP demonstrate the effectiveness of the proposed method.
PaperID: 4213, Poster
Abstract: Tracking plasma profiles, such as electron temperature and rotation, remains a challenging task in tokamak nuclear fusion. While RL-based approaches offer a promising path for multi-input-multi-output control systems, existing frameworks overlook two key obstacles: partial observability of plasma dynamics and the diagnostic mismatch between offline training and online execution. We propose a novel offline RL framework to address these challenges. Specifically, we train the control policy to condition on the latent context that is extracted from a trained plasma dynamics model. This dynamics-informed latent context encodes useful information about operating regimes, enabling the policy to adapt as the plasma evolves. To mitigate the gap between offline and real-time observation, we learn a converter that maps plasma diagnostics reconstructed offline to their real-time-observable counterparts, thereby exposing policy training to the real-time plasma diagnostics available in live experiments. In offline evaluations, the proposed method substantially improves tracking performance and robustness across a diverse set of discharges. We report real-time deployment results on DIII-D Tokamak Fusion Facility, showing feasibility under challenging actuator conditions.
PaperID: 4214, Poster
Abstract: Knowledge distillation provides shared semantic supervision for local training in federated learning, reducing the tendency of client models to overfit their own data distributions. Existing methods construct a global knowledge repository from client-uploaded local information to establish consensus across clients. However, global knowledge that is fundamentally derived from client data inevitably contains local biases under heterogeneous settings. To address this issue, this study proposes a geometry-regularized cross-silo consensus enhancement method, termed RISE, which improves collapse resistance in federated learning by preserving pre-trained geometric consensus priors during local adaptation. Specifically, RISE introduces two complementary geometry regularizers. The Biased Low-dimensional Subspace Correction module preserves the spectral energy profile of pre-trained representations by constraining the energy distribution of adapted features within the principal representation subspace. Instead of enforcing direction-wise alignment, it discourages excessive energy concentration along a few dominant components, thereby mitigating biased low-dimensional feature collapse. The Manifold Relation Graph Alignment module regularizes the relational manifold geometry by constructing a transferable inter-class topology from pre-trained global class prototypes and aligning client-side class structures with this global relation, thereby mitigating class representation drift and stabilizing global aggregation. Extensive experiments on eight benchmark datasets across two tasks demonstrate that RISE improves cross-client consensus consistency and enhances the generalization ability of the global model, outperforming eight state-of-the-art methods.
PaperID: 4215, Poster
Abstract: Real-world applications of modern AI systems require generalization to both unseen environments and novel human users. The field of zero-shot coordination (ZSC) focused primarily on deploying AI to novel partners in otherwise familiar environments seen in the training. The popular population-based methods enable partner generalization by training the agent on a diversity of simulated human users (serving as training-time partners) in a fixed environment, aiming to cover the diverse behaviors of novel human users. Such a paradigm does not consider the more challenging realistic scenarios where both the users and the environment are unseen during training. Recently, training across environments has been explored in training a single self-play agent, which shows better generalization on both novel human users and unseen environments. In this work, we ask whether jointly modeling partner diversity and environment diversity yields stronger coordination with novel partners in novel environments. We introduce Cross Environment Co-Play (CECP), a two-stage framework that first trains a population of self-play agents across a distribution of procedurally generated environments, and then trains a cooperator policy against this partner population. To further couple social and environmental reasoning, we add a next-state prediction module that encourages representations to encode both partner behavior and task structure, and validate the design choice with ablation studies. We evaluate CECP on both a neuroscience-inspired cooperative foraging task, and the population ZSC benchmark Overcooked, under held-out partner and held-out environment conditions. CECP achieves the strongest performance among the compared baselines and attains state-of-the-art results under our evaluation protocol for human-AI collaboration. We provide further analysis using the model's latent representation encoded by the next-state prediction module, showing it can not only improve the performance of the cooperative AI, but also provide more interpretable latents, which demonstrate the inference of partner skill levels.
Abstract: Constraint-based causal discovery recovers causal DAG structure from conditional independence (CI) relations. Classical methods such as PC orient v-structures first, then propagate edge directions from these seeds, relying on accurate CI tests and rich separating-set searches. In practice, these conditions often fail, causing cascading orientation errors. Recent work uses large language models (LLMs) as experts to augment edge orientation when standard assumptions fail, but often treats LLM outputs as reliable or assumes stable error behavior, despite hallucinations and instability. We propose MosaCD, a constraint-based framework that robustly combines CI tests with LLM inference to obtain high-confidence orientation seeds, then propagates them with Seeded Propagation Rules (SPR), which mitigate the fragility of collider-first orientation. We prove oracle-level soundness for SPR under explicit assumptions on the skeleton, separating-set record, and seed orientations, and give a stylized finite-sample analysis showing why prioritizing non-collider evidence can reduce orientation errors. Across 13 real-world benchmarks, MosaCD and SPR achieve substantially higher accuracy than existing methods through more reliable seeds and more robust propagation.
PaperID: 4217, Poster
Abstract: We study the accumulation of test errors in language models performing multi-step reasoning. Although longer reasoning improves model capabilities by scaling test-time computation, test error compounds in autoregressive generation and can grow substantially across steps. The key factors determining the stability of such reasoning remain unclear. To address this question, we analyze the test error of multi-step reasoning under supervised fine-tuning. We first present a test error bound governed by the product of across generation steps. This product, summarized as an error amplification factor, scales exponentially with steps and controls the reasoning stability. Then, we explicitly analyze this factor in transformers trained to predict linear and quadratic function weights. We prove that the transformer converges to a solution where this amplification factor decays, yielding near-zero test loss even over longer steps. Based on the analysis, we propose a training method to suppress the error amplification factor by combining (i) chain-of-thought length compression that reduces the reasoning steps, and (ii) quantization-aware training that provably regularizes the input Jacobian norms. We validate our method by fine-tuning language models on reasoning tasks that require executing graph algorithms and tracking variable states through logical operations. Across seven tasks, our method improves over baselines by 3.5% on average, and by 8.2% in length generalization evaluations. We further validate that it reduces the error amplification factor in fine-tuned models.
PaperID: 4218, Poster
Abstract: We introduce a refined differentially private (DP) data structure for kernel density estimation (KDE) with \ell_1, \ell_2 and \ell_p^p kernels. This new DP data structure offers not only improved privacy-utility tradeoff but also better query efficiency over prior results. Specifically, we study the mathematical problem: given a similarity function f (or DP KDE) and a private dataset X \subset \mathbbR^d, our goal is to preprocess X so that for any query y \in \mathbbR^d, we approximate \sum_x \in X f(x, y) in a differentially private fashion. The best previous algorithm for f(x, y) = |x - y|_1 is the node-contaminated balanced binary tree by [Backurs, Lin, Mahabadi, Silwal, and Tarnawski, ICLR 2024]. Their algorithm requires O(nd) space and time for preprocessing with n = |X|. For any query point, the query time is \alpha^-1 d \log^2 n, with an error guarantee of (1+\alpha)-approximation and \varepsilon^-1 \alpha^-0.5 d^1.5 R \log^1.5 n. In this paper, we use the same space and pre-processing time, improve the best previous result [Backurs, Lin, Mahabadi, Silwal, and Tarnawski, ICLR 2024] in three aspects - We reduce query time by \alpha^-1 \log n factor - We improve the approximation ratio from \alpha to 1 - We reduce the error dependence by a factor of \alpha^-0.5 From a technical perspective, our method of constructing the search tree differs from previous work [Backurs, Lin, Mahabadi, Silwal, and Tarnawski, ICLR 2024]. In prior work, for each query, the answer is split into \alpha^-1 \log n numbers, each derived from the summation of \log n values in interval tree countings. In contrast, we construct the tree differently, splitting the answer into \log n numbers, where each is a smart combination of two distance values, two counting values, and y itself. We believe our tree structure may be of independent interest.
PaperID: 4219, Poster
Abstract: Intersectional fairness requires evaluating models across fine-grained subgroups. However, as group definitions become more granular, the sample size per subgroup shrinks, causing fairness estimates to become increasingly unstable. Existing in-processing methods are : they penalize observed disparities without distinguishing statistically certified unfairness from sampling noise. This leads to , where models sacrifice utility to "fix" spurious disparities in sparse regimes. To address this, we introduce the metric family, a continuous, uncertainty-aware fairness score that measures only the portion of a subgroup’s disparity exceeding what sampling variation can explain at a specific confidence level. Leveraging this, we propose SAFER (Significance-Aware Fairness Regularizer), a differentiable in-processing method that gates fairness penalties using a soft-count likelihood-ratio significance test. SAFER satisfies a : when no subgroup disparity is statistically certifiable, its fairness gradient is exponentially small and training reduces to unconstrained optimization. Mechanism ablations across three benchmark datasets and purpose-built stress variants confirm that each component of SAFER – threshold smoothing, the significance gate, and a dense-tail term – is empirically necessary in its theoretically predicted operating regime. SAFER consistently achieves a superior fairness-utility Pareto frontier, compared to baselines methods in sparse intersectional settings while preserving utility when certified unfairness is absent.
PaperID: 4220, Poster
Abstract: Outcome-based post-training methods for mathematical reasoning rely on binary feedback, treating all correct trajectories equally. However, this masks a critical distinction: some correct paths are structurally brittle and prone to cascading errors under inevitable autoregressive sampling noise, while others are robust and maintain their correctness despite these natural decoding variations. Optimizing indiscriminately over these brittle solutions leads to inefficient learning, as the model wastes capacity memorizing brittle reasoning strategies. To address this, we propose Trajectory Robustness-Driven Evolutionary Self-Training (TREST). Instead of relying on binary feedback, our framework explicitly prioritizes trajectory robustness by integrating evolutionary algorithms with supervised fine-tuning. Crucially, our analysis reveals that trajectory robustness serves as a strong natural proxy for genuine mathematical insight. By conceptualizing reasoning paths as evolving individuals within a population, this evolutionary process naturally marginalizes brittle, brute-force calculations in favor of robust, insight-driven strategies. By effectively internalizing these strategies, TREST achieves superior reasoning performance. Experiments on complex mathematical benchmarks demonstrate that under aligned budgets, TREST outperforms outcome-based baselines, establishing a highly effective alternative to conventional post-training paradigms.
PaperID: 4221, Poster
Abstract: 3D anomaly detection is a crucial process in agricultural production chains, playing a vital role in automated processing and ensuring food safety. However, existing 3D anomaly detection datasets are primarily designed for industrial scenarios with standardized geometric shapes. In contrast, agricultural products originate from natural growth, exhibiting significant variations in shape, size, and texture. Consequently, the definition of "normal" is far less explicitly constrained than that of industrial components, rendering current industrial datasets and detection methods inadequate for this application. To bridge this gap, we make contributions from both data and method perspectives. On the data side, we introduce NaturalGrowth, the first dataset dedicated to agricultural point cloud anomaly detection. On the method side, we propose Nature3D-AD, which incorporates two novel modules. Anomaly-Sensitive Positional Encoding (ASPE) integrates local curvature and density information to provide geometry-aware positional representations. Geometry-Aware Attention (GAA) injects geometric biases in the spatial domain and captures global structural patterns through a spectral FFT branch. Extensive experiments on NaturalGrowth, Anomaly-ShapeNet, and Real3D-AD demonstrate that Nature3D-AD achieves state-of-the-art performance. The dataset and code will be released.
Authors: Soumyaratna Debnath, weiming zhang, Shriram Damodaran, Dingwen Xiao, Lin Wang
Abstract: Spherical Transformers have emerged as a promising framework for panoramic semantic segmentation (PASS) by operating directly on spherical geometry and alleviating projection-induced distortions. However, existing architectures rely on assumptions of canonical spherical structure and stable viewpoints, which are frequently violated in real-world 360^\circ imagery due to unconstrained camera motion, introducing significant contextual and geometric ambiguity. Consequently, they lack adaptive mechanisms to model such ambiguity, limiting robustness to unseen spherical transformations. In contrast, biological perception is inherently ambiguity-aware: rather than estimating uncertainty probabilistically, it adapts to fluctuations in cue reliability caused by geometric and contextual variations, enabling stable interpretation under complex visual transformations. Motivated by these observations, we first present a systematic analysis of existing PASS architectures under various unseen spherical transformations. Following this, we introduce AdapToPASS, a novel, bio-inspired Spherical Transformer that adaptively models contextual and geometric ambiguities for robust panoramic semantic segmentation. At the core of AdapToPASS are the Adaptive Spherical Attention (AdaSpA) blocks that dynamically modulate attention according to local contextual ambiguity, mimicking the adaptive, context-driven perception of biological vision. To address geometric ambiguities, AdapToPASS employs Bifocal Spherical Representation that reconciles the trade-off between field of view and spatial resolution; along with boundary supervision to emulate the boundary-sensitive nature of biological vision. We evaluate AdapToPASS in both indoor and outdoor semantic segmentation, where it consistently outperforms prior state-of-the-art methods. We further validate AdapToPASS under unseen spherical transformations, where it demonstrates strong robustness and surpasses the next-best method by +13.38% relative improvement in mIoU on Stanford2D3D and +18.77% on WildPASS. Additionally, we introduce a lightweight variant, AdapToPASS-Tiny, with fewer than 2M parameters, which surpasses compact baselines while retaining robustness to spherical transformations.
PaperID: 4223, Poster
Abstract: Improving the reasoning ability of large language models (LLMs) without ground-truth supervision remains a central challenge in reinforcement learning (RL). Recent label-free RL methods typically rely on internal pseudo-rewards such as self-consistency or majority voting. However, for complex multi-step reasoning, these signals can be systematically misleading: models may repeatedly produce mutually consistent but incorrect trajectories, causing optimization to favor frequency over correctness. In this work, we propose reasoning warm-up reinforcement learning, a label-free RL framework motivated by an empirical phenomenon we call . We observe that when a model successfully completes a deterministic auxiliary task with an exact verifier, its success probability on a subsequent complex reasoning task increases substantially. This coupling suggests that auxiliary-task verification can serve as a useful surrogate signal for selecting and optimizing reasoning trajectories even when the main task itself is not directly verifiable. Based on this observation, we incorporate verifiable auxiliary tasks into both generation and optimization. During generation, the auxiliary task provides verifiable warm-up prefix; during training, its verification outcome is used to score candidate trajectories for policy optimization. Experiments on multiple benchmarks demonstrate the effectiveness of our method.
PaperID: 4224, Poster
Abstract: Deep learning models for electronic health records achieve strong predictive performance but provide little insight into they represent internally. Post-hoc attribution methods produce per-cell importance scores that lack temporal coherence, cross-variable structure, and reusability across patients. We introduce the from a frozen predictive model’s representations. Each dictionary entry (latent) is defined by learned temporal Gaussian islands—specifying are involved. A decoupled selector then identifies which dictionary entries drive each patient’s prediction, separating representation learning from attribution. We evaluate TI-SAE across five backbone architectures spanning different representation strategies on MIMIC-III and MIMIC-IV (three prediction tasks each). TI-SAE produces structured clinical concepts that go beyond transient attribution heatmaps while achieving strong performance on standard faithfulness benchmarks.
Abstract: Video Diffusion Transformers (DiTs) generate high-quality videos but demand substantial compute due to wide blocks, deep architectures, and iterative sampling. Recent methods reduce cost by compressing width, depth, or sampling steps, but typically commit to a fixed architecture that cannot adapt to individual inputs or denoising stages. We propose PARE (Pruning and Adaptive Routing for Efficient video generation), which jointly compresses width and depth with structure-aware pruning and input-adaptive routing. For width, we observe that attention heads specialize into spatial and temporal roles, and design importance scoring that accounts for this distinction to prevent motion-critical temporal heads from being pruned prematurely. For depth, we train a lightweight router conditioned on denoising timestep and visual content to dynamically select which blocks to execute at each step, enabling per-input compute adaptation rather than static block removal. A progressive pipeline first recovers width-pruned quality via distillation, then jointly optimizes the student and router to decouple the two learning objectives. Experiments on Wan2.1-14B for both image-to-video and text-to-video generation show that PARE substantially reduces per-step computation while preserving quality across VBench dimensions, and composes with step distillation for further acceleration.
PaperID: 4226, Poster
Abstract: Generative modeling for discrete data has been significantly advanced by learning continuous vector fields on the statistical manifolds to transport simple priors to target discrete distributions. However, because these continuous-time flows are not constrained to follow geodesics on the statistical manifold, thereby requiring multi-step numerical integration for simulation and leading to substantial computational costs during inference. In this work, drawing inspiration from the ''straight and fast'' properties of Rectified Flow in continuous Euclidean space, we propose the Rectified Statistical Flow (ReSFlow) model, which enables one-step categorical generation on statistical manifolds by explicitly learning geodesic trajectories that adhere to the intrinsic Riemannian geometry. By rectifying the flow on the statistical manifold, we theoretically guarantee that trajectories between prior and target distributions approximate geodesics. This rectification significantly reduces the transport cost by aligning the learned trajectories directly along geodesics on the statistical manifold, allowing the model to map pure noise to target discrete data in a single step. Through extensive experiments on diverse real-world generative tasks, ranging from image and text to biological domains, we demonstrate ReSFlow's superior one-step generation quality, which truly embodies geodesically optimal and rapid categorical generation.
PaperID: 4227, Poster
Abstract: Combinatorial bandits with ranking feedback model a sequential decision-making problem in which the learner observes a top-m ranking of the set of k arms played in each round. The setting has been examined under the lens of top-k regret minimization, which accounts for the cost of pulling arms that are not among the best k. Existing works rely on the assumption that latent rankings are generated by a Plackett-Luce (PL) distribution and consider the special case of full-ranking feedback (m = k). They provide instance-dependent regret upper bounds which match the asymptotic logarithmic scaling in the learning horizon T, but suffer from a burn-in term which becomes \Omega(T) for some choices of the PL parameters, preventing the derivation of sublinear worst-case bounds. In this work, we study the setting under the general top-m feedback (m \leq k) and provide a worst-case regret lower bound of order \Omega(\sqrtT). Then, by introducing a novel algorithmic strategy, we derive an instance-dependent regret upper bound which does not suffer from the exploding burn-in term and a corresponding worst-case bound of order \tilde\mathcalO(\sqrtT), proving for the first time that it is possible to achieve sublinear worst-case regret w.r.t. the PL parameters. Moreover, we translate our algorithmic ideas to multinomial logit bandits, in which the learner receives winner feedback (m = 1) with non-zero probability of observing "no-choice". The existing regret bounds suffer from an exploding burn-in term, inversely proportional to the no-choice probability, that we avoid through our novel approach.
Abstract: Off-dynamics offline reinforcement learning (RL) aims to learn a policy for a target domain using limited target data and abundant source data collected under different transition dynamics. Existing methods typically address dynamics mismatch either globally over the state space or via pointwise data filtering; these approaches can miss localized cross-domain similarities or incur high computational cost. We propose Localized Dynamics-Aware Domain Adaptation (LoDADA), which exploits localized dynamics mismatch to better reuse source data. LoDADA clusters transitions from source and target datasets and estimates cluster-level dynamics discrepancy via domain discrimination. Source transitions from clusters with small discrepancy are retained, while those from clusters with large discrepancy are filtered out. This yields a fine-grained and scalable data selection strategy that avoids overly coarse global assumptions and expensive per-sample filtering. We provide theoretical insights and extensive experiments across environments with diverse global and local dynamics shifts. Results show that LoDADA consistently outperforms state-of-the-art off-dynamics offline RL methods by better leveraging localized distribution mismatch.
Abstract: We study tabular reinforcement learning problems with multiple steps of lookahead information. Before acting, the learner observes \ell steps of future transition and reward realizations: the exact state the agent would reach and the rewards it would collect under any possible course of action. While it has been shown that such information can drastically boost the value, finding the optimal policy is NP-hard, and it is common to apply one of two tractable heuristics: processing the lookahead in chunks of predefined sizes ('fixed batching policies'), and model predictive control. We first illustrate the problems with these two approaches and propose utilizing the lookahead in adaptive (state-dependent) batches; we refer to such policies as adaptive batching policies (ABPs). We derive the optimal Bellman equations for these strategies and design an optimistic regret-minimizing algorithm that enables learning the optimal ABP when interacting with unknown environments. Our regret bounds are order-optimal up to a potential factor of the lookahead horizon \ell, which can usually be considered a small constant.
PaperID: 4230, Poster
Abstract: Multimodal Large Language Models (MLLMs) have shown strong multimodal instruction-following ability, but adapting them to diverse visual-language domains typically assumes centralized data access and costly joint training. This is restrictive when data is distributed across private, domain-specific, or permission-limited clients. To this end, we propose DISTMOE, a mixture-of-experts (MoE) approach for distributed visual instruction tuning. In each layer of the language decoder it augments the public feedforward network (FFN) with a client-specific private FFN expert, with the goal to acquire domain-specific knowledge. However, independent expert training causes the private FFNs to learn representation of different scale and magnitudes, making merging the experts difficult. To reduce client-specific drift, we introduce a public-anchored expert composition stage that updates only routers and lightweight private projection adapters on a mix of local client data and public data, via an isotropic regularization loss, therefore making it rehearsal-free. During inference, DISTMOE performs modular routing over public and private experts, enabling token-wise domain composition without explicit domain labels. Experiments across diverse visual-language benchmarks show that DISTMOE enables flexible expert reuse, effective domain adaptation, and competitive performance while preserving modular control over client-specific knowledge.
PaperID: 4231, Poster
Abstract: Machine learning models for chemical reaction prediction are usually evaluated by endpoint accuracy, but endpoint correctness alone does not reveal how a model represents chemical transformation or why it fails. We introduce UniRxnAxis, a unified forward--retrosynthesis framework that models both tasks as direction-conditioned endpoint redistribution in a shared latent space. By aligning source states, transformed states, and encoded target endpoints, UniRxnAxis enables directional auditing of learned latent updates. This shared latent geometry allows the learned displacement (from the encoder source state (S_0) to the decoder-facing transformed state (S_1)) to be decomposed into an endpoint-parallel component and an endpoint-orthogonal residual, which we call emergent steering. Our evaluation shows that emergent steering is functionally important: in retrosynthesis, it recovers most of the model's utility, remains effective under random, shuffled, and prototype controls, and directly affects which bond edits are promoted or suppressed. The decomposition also supports a proxy taxonomy of model-side failures, showing how reaction models can be audited beyond endpoint accuracy.
Abstract: Text-to-motion generation and motion-to-text captioning are two fundamental tasks in human motion modeling, both grounded in the same underlying motion-text correspondence. Existing unified approaches mostly rely on autoregressive modeling, which imposes a fixed generation order and is therefore poorly suited to the bidirectional dependencies between language and motion, allowing early prediction errors to persist as fixed context and degrade both temporal coherence and cross-modal consistency. Masked discrete diffusion, which models sequences through iterative bidirectional prediction, offers a natural remedy. We therefore propose BiMoGen (Bidirectional Motion-text Generation), a unified masked discrete diffusion framework for bidirectional motion-text modeling. To stabilize training, we design Decoupled Uni- and Cross-Modal Training, in which masked pretraining first establishes cross-modal correspondence on paired motion-text sequences, after which supervised fine-tuning specializes the model for bidirectional generation. Masked diffusion nonetheless introduces its own source of error, as the model is trained on clean ground-truth context yet encounters self-generated and potentially erroneous context at inference, with errors committed under heavily masked states propagating through subsequent steps. We further introduce Generation-Aware Self-Correction that exposes the model to its own predictions during training and applies correction passes at early sampling steps to revise unreliably committed tokens. Extensive experiments on HumanML3D and KIT-ML demonstrate competitive performance on both tasks, validating the effectiveness of the proposed two-stage training and self-correction designs.
Authors: Dongfang Zhao
Abstract: Graph-based Approximate Nearest Neighbor (ANN) search often suffers from performance degradation in high-dimensional spaces due to the Euclidean-Geodesic mismatch, where greedy routing diverges from the underlying data manifold. To address this challenge, this paper presents Manifold-Consistent Graph Indexing (MCGI), a geometry-aware and disk-resident indexing method that leverages Local Intrinsic Dimensionality (LID) to dynamically adapt search strategies to the intrinsic geometry of data. Unlike conventional algorithms that treat dimensions uniformly, MCGI modulates its beam search budget based on in-situ geometric analysis, which reduces sensitivity to data-specific hyperparameters by replacing a single scalar with a geometry-informed range that remains stable across datasets of varying dimensionality. Theoretical analysis demonstrates that MCGI provides robust approximation by preserving manifold-consistent topological connectivity. Extensive evaluations against five industry-standard baselines across five datasets up to billion scales confirm the advantages of the proposed approach.
Abstract: Generative real-world image super-resolution (Real-ISR) can synthesize visually convincing details from severely degraded low-resolution (LR) inputs, yet its stochastic sampling makes a critical failure mode hard to avoid: outputs may look sharp but be unfaithful to the LR evidence, exhibiting semantic or structural hallucinations. Preference-based reinforcement learning (RL) is a natural fit because each LR input yields a rollout group of candidate restorations. However, effective alignment in Real-ISR is hindered by three coupled challenges: (i) the lack of an LR-referenced faithfulness signal that is robust to degradation yet sensitive to localized hallucinations, (ii) a rollout-group optimization bottleneck where scalarizing heterogeneous rewards before normalization compresses objective-wise contrasts and weakens DiffusionNFT-style reward-weighted updates, and (iii) limited coverage of real degradations, which restricts rollout diversity and preference signal quality. We propose LucidNFT, a multi-reward RL framework for flow-matching Real-ISR. LucidNFT introduces LucidConsistency, a degradation-invariant and hallucination-sensitive LR-referenced evaluator trained with content-consistent degradation pools and original-inpainted hard negatives; a decoupled reward normalization strategy that preserves objective-wise contrasts within each LR-conditioned rollout group before fusion; and LucidLR, a large-scale collection of real-world degraded images for robust RL fine-tuning. Extensive experiments show that LucidNFT improves perceptual quality on strong flow-based Real-ISR baselines while generally maintaining LR-referenced consistency across diverse real-world scenarios.
PaperID: 4235, Poster
Abstract: Streaming long-video understanding is a realistic setting for always-on video assistants, where visual input arrives continuously and users may ask questions at arbitrary times. We study streaming long-video question understanding (QA) under a unified temporal formulation in which multiple questions are revealed over time and may concern the past, present, or future relative to their query times. The model receives no task-type labels and must operate under strict sequential access and bounded shared memory, deciding whether each active question is already answerable or should remain open for future evidence. This differs from prior streaming settings that often focus on short streams, prefix-answerable queries, separated temporal categories, or response timing without long-horizon shared memory. To evaluate this formulation, we propose Ref2Stream (Reference-to-Stream), a general benchmark-conversion protocol that turns temporally grounded offline video QA benchmarks into streaming QA benchmarks. Instantiated on LVBench, Ref2Stream yields LVBench-Ref2Stream, which mixes past, present, and future questions under strict sequential access and evaluates both answer accuracy and answer latency. Existing streaming video methods typically improve efficiency or response timing through predefined compression, retrieval, or response policies, while offline video agents perform adaptive evidence gathering but assume access to the full video and target question. We introduce PIVOT (Planning over Incremental Video Observations for Timely Answering), a unified agentic framework that incrementally updates bounded memory from the observed stream and uses a planner-controlled action loop to retrieve relevant memory, analyze current evidence, reflect on candidate answers, and decide whether to answer or defer. Our work suggests that agentic reasoning is a promising direction for building streaming long-video systems that can adaptively gather evidence and answer only when sufficient information has been observed.
PaperID: 4236, Poster
Authors: Stephen J Roberts, Yuki Tachibana
Abstract: Hamiltonian Neural Networks (HNNs) produce stable long-horizon rollouts by integrating a learned Hamiltonian with a symplectic scheme. Standard Monte Carlo (MC) dropout conflicts with this per-trajectory interpretation because resampling the dropout mask inside a leapfrog integrator means successive substeps follow different subnetworks, so a single MC rollout no longer corresponds to any one sampled Hamiltonian. We propose trajectory-consistent dropout, implemented as Fixed-mask MC, where each stochastic rollout selects one mask from a finite bank and reuses it across the full integration path. This restores a coherent per-rollout Hamiltonian sample and makes uncertainty decomposition operational. Crossing the mask axis with an initial-condition ensemble separates model-sample uncertainty from sensitivity to the initial state. On coupled-pendulum rollouts, Fixed-mask MC reduces mean energy drift by 72% relative to Standard MC. Under held-out mass ratios, Fixed-mask MC increases the epistemic share of predictive variance by 16.9 percentage points from in-distribution to out-of-distribution, compared with 4.2 percentage points for Standard MC, indicating a substantially stronger shift of the uncertainty budget toward model uncertainty when the physics moves outside training support. Spring-chain and transverse-field Ising experiments show that the same axis-preserving decomposition transfers to structured pairwise and latent quantum dynamics, although point-prediction and likelihood gains remain system-dependent. These results position Fixed-mask MC as a single-training-run route to trajectory-consistent, physically interpretable uncertainty in structured neural dynamical models.
Abstract: Neural scaling laws establish a predictable relationship between model performance and data or compute, offering crucial guidance for resource allocation in new domains and tasks. Yet such laws are most needed precisely where they are hardest to obtain: fitting one for a new model–task pair demands expensive sweeps that typically exhaust the very compute budget the law is meant to economize. This paper poses the research question of how to develop generalizable scaling laws: laws fit once on a well-resourced source domain and reliably transported to new domains where running a full sweep is infeasible, which requires a fundamental understanding of when and why scaling properties change. We address this by identifying the right invariants: scaling laws are preserved under bijective (information-preserving) transformations of the data and modified in predictable, information-theoretically grounded ways under non-bijective transformations that lower its information resolution \rho: a single axis along which a law fit in one domain can be transported to another. We validate this across language, vision, and speech, and demonstrate two cross-domain applications: predicting scaling for language models trained on electronic health records from laws fit on general text, and predicting time-series classification scaling under varying levels of noise injection, recovering the data-scaling exponents to within 3-percent error.
PaperID: 4238, Poster
Abstract: Verifying the local robustness of neural networks under \ell_0 perturbations is inherently combinatorial: a verifier must reason over all sparse subsets of perturbed pixels. This paper presents FocusBranch, a combinatorial branch-and-bound framework for \ell_0 robustness verification. The key insight is to decompose the \ell_0-ball using nodes, where each node consists of a focus pixel set \mathcalS and a level l, requiring at least l perturbed pixels to lie in \mathcalS. This abstraction unifies two branching mechanisms: concrete nodes branch through upper shadows to fix additional perturbed locations, while super nodes branch through covering designs to cover many perturbation cases with fewer child nodes. FocusBranch proves node safety by concretizing linear relaxations over node-constrained \ell_0-balls and recursively branches on unknown nodes. Rather than relying on a fixed verification strategy, it samples nodes across levels to estimate success rates and costs, then selects a starting level and branching strategy predicted to minimize verification time. Experiments on fully connected and convolutional classifiers for MNIST, Fashion-MNIST, and CIFAR-10 show that FocusBranch verifies \ell_0 local robustness effectively across architectures and outperforms state-of-the-art verifiers.
Abstract: While autonomous vehicles have achieved reliable performance within specific operating regions, their deployment to new cities remains costly and slow. A key bottleneck is the need to collect many human demonstration trajectories when adapting driving policies to new cities that differ from those seen in training in terms of road geometry, traffic rules, and interaction patterns. In this paper, we show that self-play multi-agent reinforcement learning can adapt a driving policy to a substantially different target city using only the map and meta-information, without requiring any human demonstrations from that city. We introduce riving (NOMAD), which enables policy adaptation in a simulator constructed based on the target-city map. Using a simple reward function, NOMAD substantially improves both task success rate and trajectory realism in target cities, demonstrating an effective and scalable alternative to data-intensive city-transfer methods.
Abstract: While diffusion has drawn considerable recent attention from the language modeling community, continuous diffusion has appeared less scalable than discrete approaches. To challenge this belief we revisit Plaid, a likelihood-based continuous diffusion language model (DLM), and construct RePlaid by aligning the architecture of Plaid with modern discrete DLMs. In this unified setting, we establish the first scaling law for continuous DLMs that rivals discrete DLMs: RePlaid exhibits a compute gap of only 22× compared to autoregressive models, matches Duo while using fewer parameters, and outperforms MDLM in the over-trained regime. We benchmark RePlaid against recent continuous DLMs: on OpenWebText, RePlaid achieves a new state-of-the-art PPL bound of 22.6 among continuous DLMs and superior generation quality. These results suggest that continuous diffusion, when trained via likelihood, is a highly competitive and scalable alternative to discrete DLMs. Moreover, we offer theoretical insights to understand the advantage of likelihood-based training. We show that optimizing the noise schedule to minimize the ELBO's variance naturally yields linear cross-entropy (information loss) over time. This evenly distributes denoising difficulty without any case-specific time reparameterization. In addition, we find that optimizing embeddings via likelihood creates structured geometries and drives the most significant likelihood gain.
Abstract: Recent text-to-image (T2I) generators can synthesize realistic images, but still struggle with compositional prompts involving multiple objects, counts, attributes, and relations. We introduce EPIC (Efficient Predicate-Guided Inference-Time Control), a training-free inference-time refinement framework for compositional T2I generation. EPIC casts refinement as predicate-guided search: it parses the original prompt once into a fixed visual program of object variables and typed predicates, covering checkable conditions such as object presence, counts, attributes, and relations. Each generated or edited image is verified against this program using visual evidence extracted from that image. An image is judged to satisfy the prompt only when all predicates are satisfied; otherwise, failed predicates decide the next step, routing local failures to targeted editing and global failures to resampling while the fixed visual program remains unchanged. On GenEval2, EPIC improves prompt-level accuracy from 34.16% for single-pass generation with the base generator to 71.46%. Under the same generator/editor setting and maximum image-model execution budget, EPIC outperforms the strongest prior refinement baseline by 19.23 points while reducing realized cost by 31% in image-model executions, 72% in MLLM calls, and 81% in MLLM tokens per prompt.
PaperID: 4242, Poster
Authors: 瑞安 领带, Wenbo Xiong, Zhengyu Shi, Xinyu Su, 江晨雨, Libo Wu, Hao Li
Abstract: Climate downscaling of global General Circulation Models (GCMs) is a challenging ill-posed inverse problem. In practice, the lack of paired data makes end-to-end supervised training infeasible, while downstream assessments require processing large, heterogeneous GCM ensembles under strict computational limits. Consequently, developing an efficient, zero-shot downscaling approach is crucial. While vanilla Diffusion Posterior Sampling has recently shown immense potential by bypassing paired training, it remains bottlenecked by geographic misalignments, spectral biases, and prohibitive inference costs. To address these limitations, we propose the GCM-consistent Diffusion for Zero-shot Downscaling (GCD) framework. First, GCD conditions the diffusion prior on static geographic boundaries to ensure physical fidelity. Second, we introduce a filter measurement operator, effectively mitigating spectral discrepancies between GCMs and the prior to provide stable inference-time guidance. Third, to overcome the inference bottleneck, we design a joint acceleration scheme: by combining progressive distillation with Conjugate Gradient guidance to correct large-step deviations, GCD compresses the generation process to just 6 steps. Extensive experiments demonstrate that GCD achieves highly competitive 99th percentile accuracy across five heterogeneous GCMs and successfully recovers the high-frequency details of tropical cyclones, entirely without model-specific retraining.
Abstract: Modern agents powered by thinking LLMs achieve high accuracy through long chain-of-thought reasoning but incur substantial inference costs. While many LLMs now support configurable reasoning levels (e.g., high/medium/low), static strategies are often ineffective: using low-effort modes at every step leads to significant performance degradation, while random selection fails to preserve accuracy or provide meaningful cost reduction. However, agents should reserve high reasoning effort for difficult steps like navigating complex website structures, while using lower-effort modes for simpler steps like opening a target URL. In this paper, we propose Ares, a framework for per-step dynamic reasoning effort selection tailored for multi-step agent tasks. Ares employs a lightweight router to predict the lowest appropriate reasoning level for each step based on the interaction history. To train this router, we develop a data generation pipeline that identifies the minimum reasoning effort required for successful step completion. We then train the router to predict these levels, enabling plug-and-play integration for any LLM agents. We evaluate Ares on a diverse set of agent tasks, including TAU-Bench for tool use agents, BrowseComp-Plus for deep-research agents, and WebArena for web agents. Experimental results show that Ares reduces reasoning token usage by up to 52.7% compared to fixed high-effort reasoning, while introducing minimal degradation in task success rates.
Abstract: Vision-language models (VLMs) have achieved strong image and video understanding, yet their visual-spatial representations remain geometrically fragile, leading to failures in spatial reasoning required for embodied AI, robotics, and autonomous driving. Existing approaches to geometry grounding either fine-tune VLMs on spatial question answering, which can perpetuate spurious visual representations, or fuse features from large geometry-grounded vision models, which substantially increases model size during inference. Knowledge distillation from geometry-grounded vision models offers an alternative, but directly matching multi-view teacher features can disrupt the pretrained alignment between visual and textual representations, degrading object- and language-semantic understanding. We propose Multi-view relational distillation (MVRD), which distills patch-wise cosine similarities across views instead of the teacher features themselves. These multi-view relations encode geometric correspondences sufficient for spatial understanding while leaving the student representation underdetermined, allowing it to remain close to its pretrained vision-language space. Across representative VLMs, MVRD improves visual-spatial reasoning, outperforming supervised fine-tuning and direct feature distillation while approaching feature-fusion methods with substantially fewer added parameters and lower latency. Further analysis shows that MVRD makes visual representations more geometric without breaking vision-language alignment, and generalizes to 3D scene understanding tasks, including object grounding, dense captioning, and question answering.
PaperID: 4245, Poster
Abstract: Parameter-Efficient Fine-Tuning (PEFT) has emerged as a prevalent strategy for adapting large pre-trained models. Among PEFT techniques, reparameterization-based methods, particularly those that employ low-rank or sparsity structures, have attracted significant attention. However, these approaches often underperform full fine-tuning. Specifically, low-rank methods constrain the update rank, which is misaligned with the inherently high-rank nature of full fine-tuning. In contrast, sparsity-based techniques limit the flexibility of the parameter space and compromise connectivity across weights. To overcome these limitations, we explore the use of Toeplitz matrices, whose entries are constant along each diagonal, therefore providing a compact parameterization without imposing explicit rank or sparsity constraints. A straightforward approach is to update weights via a product of Toeplitz matrices. However, similar to RNNs, long multiplicative chains often lead to gradient instability. To address this issue, we propose , a more refined version of this idea that replaces standard Toeplitz matrices with block-wise Toeplitz, greatly increasing each factor’s capacity and shortening the chain, thereby stabilizing training and improving expressiveness. We theoretically show that products of block-wise Toeplitz matrices can approximate arbitrary matrices under shorter chain length, justifying the structural design of ToPA. Furthermore, we demonstrate that ToPA offers greater expressivity than existing low-rank and sparse parameterizations. Empirical evaluations across 21 NLP and CV datasets, spanning 5 model architectures, consistently validate the effectiveness of ToPA.
Abstract: Large language models (LLMs) often achieve strong benchmark accuracy yet remain brittle under small distribution shifts. While recent mechanistic studies reveal the discrepancy between LLMs and humans in skill compositions, the learning dynamics of skill acquisition and the role of data distributions remain elusive. In this study, we train transformers on synthetic arithmetic tasks with black-box model-agnostic metrics for analyzing non-human skill compositions. We discover that transformers often acquire skills for arithmetic in reverse order or in parallel instead of human-like sequential rules—a phenomenon we refer to as shattered compositionality. To explain these behaviors, we provide evidence that correlational matching to the training data, rather than causal or procedural composition, shapes learning dynamics. As a consequence, this non-human acquisition creates competition between partially learned skills, producing characteristic mixing errors and weaker robustness under controlled distribution shifts. We further show that the same qualitative behavior persists in modern LLMs and is not mitigated by pure model scaling or scratchpad supervision. Our results highlight a mismatch between training-time skill acquisition and the human-like hierarchical compositions, with implications for reasoning reliability and out-of-distribution robustness. An anonymized code repository is provided in the supplementary material.
Authors: Payam Piray
Abstract: Adaptive decision-making in biological and artificial intelligence requires balancing the exploitation of known outcomes with the exploration of uncertain alternatives. Although prior work suggests that uncertainty generally promotes exploration, it has typically treated distinct sources of environmental uncertainty as equivalent. We consider environments with latent reward states that drift over time (volatility) and are observed through noisy outcomes (stochasticity). Both increase posterior uncertainty, yet we show they drive optimal exploration in opposite directions: volatility enhances it, stochasticity suppresses it. We establish this asymmetry formally by extending the Gittins index framework to Gaussian state-space bandits with latent dynamics. We further derive Cause-Aware Uncertainty-Sensitive Exploration (CAUSE), a closed-form exploration bonus obtained via control-as-inference that inherits the same monotonicities. CAUSE outperforms standard exploration strategies in environments with heterogeneous noise structure, and also improves on a Gittins-per-arm policy whose rested-bandit optimality does not transfer to restless settings. Learning and exploration are governed by the same noise-inference asymmetry, and the framework predicts that pathological noise inference produces reversed rather than merely impaired exploration, with implications for computational accounts of psychiatric conditions.
PaperID: 4248, Poster
Authors:
Bo Peng, Rongzhen Ye, Yaohua Li, Weilin Luo, Hai Wan, Jiahao Xu, Zexian LiangAbstract: The Maximal Independent Set (MIS), Dominating Set (DS), and Vertex Cover (VC) problems are fundamental combinatorial optimization problems with wide-ranging applications. We identify a key property of these problems: by finding a cut in the graph and determining the states of vertices associated with the cut, the graph can be decomposed into multiple subgraphs whose solutions are independent of each other. Therefore, we classify MIS, DS, and VC as Edge-Cut Separable Problems (ECSPs). This property enables a natural divide-and-conquer strategy, which significantly reduces the computational complexity and allows neural networks to scale to large-scale graphs that were previously intractable. Based on this property, we propose a recursive framework that leverages this property to approximately solve ECSPs, namely \ours. Our framework consists of two major components: a partitioner that identifies a cut in the graph, and a predictor that solves the cut-determination sub-problem (CDSP) to determine the states of vertices associated with the cut. By recursively applying these components, RGF generates a complete solution for the original ECSP instance. Additionally, through partial prediction, the memory consumption is significantly reduced, enabling RGF to handle real-world large graphs. Experimental results validate the advantages of RGF: the performance of our method is up to 2.13% better than the state-of-the-art machine learning (ML)-based approach, and our approach can handle real-world large graphs with up to 11M vertices.
PaperID: 4249, Poster
Abstract: Vision-Language-Action (VLA) models often bundle architectural changes with different pretraining data, backbone scales, and optimization recipes, obscuring what actually drives progress. We introduce SimVLA, a deliberately minimal VLA—a standard vision-language backbone with a lightweight continuous-action head—as a controlled testbed for this attribution problem. Across LIBERO, CALVIN, WidowX, and Google Robot, one-knob-at-a-time ablations show that training dynamics are the largest measured driver, producing performance swings of \Delta\approx54-89%, while task configuration and architecture have smaller effects under matched settings. Despite using a small backbone and no additional large-scale robot-data VLA pretraining before benchmark-specific training, SimVLA reaches 98.6% average success on LIBERO, remains competitive across the other benchmarks, and transfers to held-out real-robot scenes. These results suggest that simple, carefully controlled VLA baselines remain underexplored and provide a calibration point for future architectural claims within the VLM-encoder plus lightweight-head family.
PaperID: 4250, Poster
Authors:
Sheng Ye, Jiangke Lin, Fudong Wang, Yi Yuan, ZhenHui Dong, Yong-jin LiuAbstract: Recent feed-forward 3D Gaussian Splatting (3DGS) methods have enabled effective multi-view 3D reconstruction by predicting Gaussian primitives in a single forward pass. However, existing approaches predominantly adopt a pixel-aligned paradigm that predicts one Gaussian per input pixel, coupling the Gaussian count to the input resolution and number of views. This also leads to redundant primitives in textureless regions and insufficient coverage in geometrically complex areas. We present Anchor3DGS, a pose-free feed-forward 3DGS framework that replaces pixel-aligned prediction with a compact anchor-based representation. Guided by information entropy, our method places a set of 3D anchors in the scene, concentrating them in regions of higher visual complexity while maintaining broad spatial coverage. Each anchor performs occlusion-aware aggregation of multi-view image features, exchanges information with other anchors, and is decoded into a set of local Gaussians. A lightweight feature enhancement module further refines the coarse RGB output from Gaussian splatting and improves rendering fidelity. Extensive experiments on diverse benchmark datasets demonstrate that our Anchor3DGS achieves state-of-the-art novel view synthesis quality while only using an order of magnitude fewer Gaussian primitives.
PaperID: 4251, Poster
Abstract: Pixel-space diffusion models have recently seen a revival in image generation, yet their training remains less efficient than that of their latent-space counterparts. Existing methods reduce this cost by equipping diffusion models with external representations from clean images. In contrast, representation consistency training, a new generative pretraining paradigm, stands out as it eliminates the need of any off-the-shelf semantic encoders. It proposes pre-training the diffusion model to simultaneously capture meaningful visual semantics from clean images while aligning them with data points across various noise levels. Nevertheless, its training depends on contrastive learning with hand-crafted augmentations. This introduces strong semantic prior into the pretraining, limiting model performance when transferred to downstream target distributions that substantially deviate from the pretraining manifold. To address this, we propose Masked Pixel-space Generative Pretraining (MPG), an augmentation-free framework based on masked image modeling. MPG trains the encoder to recover masked image area and aligns those predictions with samples of different noise levels. To measure the performance of pretrained model, we further introduce NC-CKNNA, a metric that quantifies the semantic structure consistency across different noise levels. Under the same architectures and fine-tuning settings, MPG consistently produces better generation performance across seven downstream datasets, while retaining competitive generation quality on ImageNet-256. These results suggest masked generative pretraining as a practical alternative for training pixel diffusion models.
PaperID: 4252, Poster
Authors: Mengxiang Zhang, Lingyuan Liu
Abstract: Supervised post-training, including supervised fine-tuning (SFT) and knowledge distillation (KD), is essential for adapting large language models (LLMs) to specialized domains. However, traditional supervised post-training methods often suffer from catastrophic forgetting, training instability, and prohibitive computational costs. While layer-wise strategies have emerged as efficient alternatives, they introduce a representation bottleneck, where constrained updates in early stages limit feature diversity and impair final generalization. To resolve these issues, we develop HALO (Homotopy-Augmented Layer Optimization), a homotopy-driven layer-wise strategy for LLM post-training. By introducing a continuous homotopy parameter and a sequence of monotonically increasing auxiliary functions, we formulate the post-training process as a continuous deformation of the optimization problem that progressively incorporates layers into the trainable set. Specifically, HALO unfreezes the model from the output back to the input in a structured sequence. We theoretically prove that this approach establishes a differentiable solution path, ensuring a stable transition from a restricted optimization state to a fully converged LLM. Extensive experiments across diverse model families and sizes demonstrate that HALO consistently enhances performance and generalization. Notably, HALO exhibits strong robustness in challenging scenarios, especially in low-resource settings, while significantly reducing convergence time and iteration counts. Our results position HALO as an effective and accessible framework for high-performing LLM specialization.
PaperID: 4253, Poster
Abstract: Image--text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield representations that remain partially organized by modality rather than by semantics. We propose ITO, a framework addressing this limitation through two complementary mechanisms with distinct roles. \emphMultimodal multiple alignment enriches supervision by constructing diverse cross-modal correspondences from multi-view image augmentations, providing the primary source of discriminative gain. A lightweight \emphtraining-time multimodal fusion module then acts as a geometric regularizer, encouraging the encoders to produce features that are compatible under fusion and thereby reducing modality-induced separation. Crucially, the fusion module is discarded at inference, preserving the efficiency of standard dual-encoder architectures while incurring training-time overhead only. Extensive experiments across pretraining scales from millions to billions of image--text pairs show that ITO consistently outperforms strong baselines on classification, retrieval, and multimodal benchmarks. Our analysis further reveals that beyond accuracy gains, training-time fusion plays a stabilizing role in optimization, mitigating the late-stage overfitting commonly observed in aggressive contrastive learning.
PaperID: 4254, Poster
Authors:
Jiyi Wang, Jingyang Ke, Bo Dai, Anqi WuAbstract: Animals generate complex behaviors by flexibly recombining a finite set of motor primitives, but existing behavior segmentation methods oversimplify this process by imposing discrete syllables under restrictive generative assumptions. Here, we introduce Motif-based Continuous Dynamics (MCD), a framework that models behavior as a continuous, compositional process driven by reusable motor motifs. MCD leverages reinforcement learning (RL) to (1) discover interpretable motif representations via transition-based representation learning, and (2) model behavior as a time-varying mixture of these motifs through motif-based policies. This formulation avoids restrictive dynamics assumptions while capturing continuity, compositionality, and long-term dependencies in behavior. Across simulated gridworld tasks, maze navigation, and animal behavior datasets, MCD identifies reusable motifs and interprets the trajectories with them. These results provide a generative account of behavior as combinations of fundamental motifs, offering a flexible and interpretable framework for studying natural behavior.
PaperID: 4255, Poster
Abstract: Most video anomaly understanding (VAU) systems still infer anomaly judgments from evidence collected in a feed-forward way under predefined rules, typically sampled frames, or events clustered from such frames. This protocol fails when the collected evidence is incomplete under imperfect predefined rules, because once reasoning begins the model can no longer search for the missing evidence. We address this limitation with SEEK-VAU, which reframes VAU as an agentic search-and-verify interaction protocol under a bounded search budget: instead of passively reasoning over pre-collected evidence, the policy actively gathers evidence across the three stages of an anomaly event chain (precursor, trigger, and aftermath) and verifies its sufficiency before finalization, turning event-chain completeness into an explicit design rather than an outcome of predefined rules. Although search-and-verify spans multiple interaction turns, the bounded search budget caps both the per-turn input tokens and the cumulative visual context, keeping inference efficient. To make this behavior learnable, we introduce evidence-faithful counterfactual verification (EFCV), which rewards selected evidence that remains supportive, compact, and necessary under counterfactual verification. We further introduce SEEK-Bench, featuring video-level episodes with temporal interval annotations, semantic QA and event-chain stage labels. Together, SEEK-VAU and SEEK-Bench establish a strong foundation for evidence-faithful and actively verifiable VAU. The code and data will be released.
PaperID: 4256, Poster
Abstract: Multimodal Large Language Models (MLLMs) can exhibit strong text-based reasoning yet fail to apply the same capabilities to semantically equivalent visual inputs. We study this failure mode in a controlled setting by rendering text-only reasoning problems as images, preserving semantic content while changing only the input modality. Across multiple MLLM families, models solve the same problems substantially better from text than from images. We characterize this discrepancy as a reasoning-access gap, where models may correctly perceive visual content but still fail to route that content into the latent reasoning machinery used for text-based tasks. To bridge this gap, we propose TIGER (Text-to-Image Gap-targeted Training for Enhanced Reasoning), a framework that repurposes text-only reasoning corpora into effective multimodal training data. By mining "modality counterfactuals", instances where a model succeeds on a text problem but fails on its semantically-equivalent rendered image, TIGER provides targeted supervision for reasoning-access failures without requiring manually curated multimodal datasets. We implement TIGER using image-conditioned Group Relative Policy Optimization (GRPO), along with SFT and DPO variants, and show how it consistently narrows the modality gap and yields generalized visual reasoning performance gains in multimodal reasoning benchmarks such as MathVerse and EMMA. We show that Reinforcement Learning with Verifiable Reward-based models still exhibit modality-dependent reasoning gaps, and that our training recipe can further reduce these gaps. Further analysis using reasoning-subspace activation and activation patching shows that TIGER enables visual representations that better activate reasoning-relevant latent subspaces of the language backbone. Our results suggest that advancing multimodal intelligence requires moving beyond perceptual alignment toward explicit visual access to existing reasoning machinery.
PaperID: 4257, Poster
Abstract: Knowledge editing aims to update outdated or incorrect knowledge in large language models without full retraining. Existing parameter-editing methods remain limited for long-sequence knowledge: single-point editors provide insufficient coverage for extended targets, while fixed chunk-based editors distribute editing effort uniformly across rigid predefined windows, failing to prioritize difficult, information-dense tokens. In this work, we show that token entropy is closely associated with both editing difficulty and semantic importance: high-entropy regions are harder to edit and often contain key factual content. Motivated by this observation, we propose DynEdit, a dynamic framework for long-sequence knowledge editing. DynEdit uses token-level entropy to guide non-uniform sliding, increasing the coverage of difficult high-entropy regions, and further applies entropy-guided token-level optimization to strengthen updates on information-dense tokens within each editing window. Experiments on three benchmarks and two backbone models show that DynEdit consistently outperforms competitive baselines across diverse long-form editing settings, achieving gains of up to +14.0 BLEU points on challenging paraphrase queries and near-saturated semantic scores on several diverse knowledge domains.
PaperID: 4258, Poster
Abstract: Video reasoning has a distinctive structure: in a chain of thought over video, most tokens follow from prior text, but a small minority hinge on specific moments of the visual input---and which tokens those are varies from one example to the next. Current post-training procedures like GRPO supply supervision that is coarse in exactly this dimension, applying a single scalar reward uniformly across the trajectory while requiring many rollouts and/or costly external judge models. Our key observation is that a model can audit its own visual reliance: at any position, the shift in its next-token distribution when the video is removed from context measures how much that prediction actually used what was seen. Computing this quantity for both a student and a privileged teacher yields a per-token \emphreliance gap that localizes positions at which the teacher's predictions depend on the video and the student's do not. We introduce Visually-Aware On-Policy Self-Distillation (VIOS), which turns this self-auditing signal into dense per-token supervision recovered from a single rollout of a single model---no external reward, judge, or larger teacher required. VIOS combines three components: an exponential moving average (EMA) teacher that co-evolves with the student to avoid the ceiling of a frozen rationalizer; a contrastive term that increases the student's sensitivity to the video input; and a reliance gap gate that concentrates supervision on the positions identified. Across 12 benchmarks spanning general, temporal, spatial, knowledge, and long-video understanding, VIOS consistently outperforms its base model and contemporary video-reasoning systems while training in over 9× fewer steps.
Abstract: Generating high-quality time-series data is challenging because real-world signals often exhibit multimodal patterns and multiscale dynamics, including oscillations and high-frequency variations. Flow Matching (FM) offers an efficient alternative to diffusion models, but practical implementations typically rely on a single finite-capacity global vector-field estimator. In such heterogeneous temporal distributions, distinct regimes may pass through nearby flow states while requiring incompatible conditional velocities. A monolithic estimator trained with the standard \ell_2 velocity-matching objective may therefore learn an overly smoothed approximation of the local transport field. This estimator-level smoothing can attenuate branch-specific dynamics, leading to spectral distortion and poor mode coverage. To address this, we propose PrismFlow, a new FM method with Koopman-inspired dynamical experts. Each expert learns residual corrections in a latent space where local nonlinear temporal evolution can be approximated by linear transitions. We further propose a confidence-aware Winner-Take-All (WTA) objective that updates only the expert best aligned with each sample while masking gradients to the others, encouraging mode-specific specialization. During sampling, the selected expert adds a residual dynamical correction to the global transport field, preserving FM stability while recovering fine-grained and high-frequency temporal structures. Across various benchmarks, PrismFlow effectively mitigates the spectral contraction in standard FM and achieves state-of-the-art performance, with a 15.6% gain in Context-FID and a 38.6% improvement in Discriminative Score, while remaining robust in low-data settings and effective for forecasting and imputation.
PaperID: 4260, Poster
Authors: Yang Li, Adams Wai Kin Kong
Abstract: Assessing model compatibility is a key challenge in collaborative machine learning. Heterogeneous data distributions can lead to discrepancies among independently trained models that degrade performance when combined, while adversarial manipulation can introduce harmful behaviors. In both cases, reliable evaluation prior to integration is essential. Existing approaches mostly rely on parameter-space statistics, which do not reflect model behavior and are often unreliable in non-IID settings. Inspired by the pointillism art style, where images are formed from small, structured dots of color, we propose Pointillism, a probing-based framework that evaluates model compatibility directly in function space without requiring task data. The method uses structured, randomized probes sampled from an out-of-distribution space to elicit model responses. From these responses, we construct compact model signatures that capture both global prediction characteristics and class-level features. Compatibility is assessed through feature consensus, measuring cross-model agreement on class-level representations via probe transfer. This provides a behavior-centric and fine-grained view of model alignment. We apply the framework to federated learning, where consensus-based selection identifies a self-consistent subset of client models for aggregation. Experiments in federated learning demonstrate improved robustness against both untargeted and backdoor attacks, outperforming existing defense methods. Our source code is included in the supplementary material.
Authors:
Mehran Taghian Jazi, Yunke Peng, Xing Huang, Yao Wang, Yaoyuan Wang, Wei Guo, Yuanyong Luo, Tianchi Hu, JUNSONG WANG, Xin Wang, Hu Liu, Yu Cheng, Yu Z Wei, Hongliang Li, Mehdi Rahimifar, Lei YAN, wxuefei, Zhuang Ma, Liulei, Hui Yu, Anandharaju D Raju, Hoang Le, Hei Yi Mak, Tanzila Rahman, Shadan GolestanAbstract: Training large foundation models at low numerical precision is one of the most promising directions for reducing the compute and memory cost of modern AI. Recent 4-bit floating-point formats such as MXFP4 and NVFP4 can be applied to linear GEMM operations in LLMs, but their limited dynamic range introduces numerical instability that prior work addresses by stacking stabilization mechanisms, typically executed at higher precision and partially eroding the efficiency gains that motivate FP4. In this work, we argue that numerical format design is itself a first-class lever for stable FP4 training, and present the first systematic study of FP4 LLM pretraining on energy-efficient Huawei Ascend NPUs. We compare the recently proposed HiFloat4 (HiF4) format against MXFP4 across both dense (OpenPangu-1B, Llama3-8B) and Mixture-of-Experts (Qwen3-MoE-30B) architectures, executing all linear and expert GEMMs in FP4. HiF4's hierarchical scaling provides enough representational headroom that a single stabilization step suffices to keep the relative loss within 1% of a full-precision baseline — roughly half the gap of MXFP4, which requires three such mechanisms — while incurring less than 1% average degradation on downstream tasks. Our results suggest that stable, accurate FP4 training does not require an ever-growing stack of stabilization techniques; it requires the right numerical format.
Abstract: Chain-of-thought (CoT) monitoring is one of the most promising tools we have for detecting model misbehavior, but its effectiveness depends on models faithfully externalizing their reasoning. Motivated by this vulnerability, we study whether monitor-aware models are capable of obfuscating their reasoning to evade detection. We use synthetic document finetuning to expose eight models to realistic pre-training-style documents describing a CoT monitor and find that monitor-aware models consistently achieve higher rates of undetected misbehavior compared to unaware controls. This effect is weaker but still present on a harder agentic task. We also show that CoT controllability, a model's ability to reshape its own reasoning trace under an imposed constraint, is closely correlated with obfuscation success across the eight models studied (r=0.800, p=0.017). Monitor-aware models placed under equal reinforcement learning optimization pressure also learn to reward-hack without triggering a CoT monitor substantially faster than unaware controls. Together, these results suggest that knowledge of monitoring combined with high CoT controllability poses a risk to CoT-based monitoring.
Authors: Hanxu Yang, Yuhuan Zhao, Xiaodong He, Zhao Kang
Abstract: Pre-training Graph Neural Networks (GNNs) via self-supervised learning has become a dominant paradigm, yet efficiently adapting frozen encoders remains a challenge. Graph prompting offers a parameter-efficient alternative to fine-tuning, but existing methods largely treat pre-trained models as opaque feature extractors, ignoring their internal spectral structure. In this work, we identify a systematic phenomenon in pre-trained GNNs, which we term spectral bias: optimization during pre-training disproportionately aligns representations with directions associated with large singular values, leaving low-energy directions under-explored. We show that these underutilized directions can encode complementary information that is beneficial for downstream adaptation, especially under distribution shift. To leverage this insight, we propose Spectral Reverse Prompt (SRP), a prompting framework that rebalances the spectral contributions of frozen GNN encoders. SRP applies a learnable soft-thresholding mask in the spectral domain to down-weight dominant directions while amplifying weaker ones. In addition, SRP incorporates a null-space augmentation module that captures variation in directions with minimal activation under the frozen encoder. Extensive experiments across multiple benchmarks demonstrate that SRP achieves state-of-the-art performance with minimal additional parameters, highlighting that reweighting spectral components is a principled and effective strategy for parameter-efficient graph adaptation.
Authors:
Jiaxing Wang, Deping Xiang, Jin Xu, Zirui Liu, Zicheng Zhang, Guoqiang Gong, Jun Fang, Chao Liu, Pengzhang Liu, Tongxuan Liu, Ke Zhang, Qixia JiangAbstract: As Large Language Model (LLM) datasets scale to trillions of tokens, data selection has emerged as a critical frontier to filter out uninformative noise and construct adaptive learning trajectories. Beyond static heuristic filtering, advanced data selection methods for LLM training largely follow two paradigms, each with fundamental limitations. Influence-based methods provide principled bi-level objectives but require intractable inverse-Hessian computations, while excess-loss methods are computationally efficient but rely on a static reference model that becomes misaligned with the evolving proxy model during training. We propose BLADE (Bi-Level Adaptive Data sElection), a Hessian-free framework for data selection. BLADE reformulates the bi-level optimization problem underlying influence-based methods as a penalized single-level objective via Lagrange multipliers, avoiding inverse-Hessian computation while revealing a principled connection to excess-loss based data selection. The resulting objective recovers an excess-loss form but replaces the static reference model with a dynamic one that stays synchronized with training. Theoretically, we prove that this penalized formulation guarantees first-order convergence. For efficient online batch selection, we instantiate BLADE as a memoryless randomized block-coordinate Frank-Wolfe algorithm. Extensive experiments show that BLADE consistently outperforms state-of-the-art data selection baselines, providing a practical recipe for LLM training.
PaperID: 4265, Poster
Abstract: Modern data-centric applications increasingly require language agents capable of operating over live, stateful database environments, where successful problem resolution depends not merely on initial SQL generation, but on reasoning about evolving database states, interacting with execution feedback, and repairing state-mutating operations. While agentic reinforcement learning (RL) has emerged as a promising paradigm for training small language models (SLMs), existing frameworks are predominantly tool-centric and struggle with data-centric environments. Specifically, applying tool-centric frameworks to stateful tasks like SQL debugging introduces severe spatial inefficiencies since proper environment isolation typically requires a separate database container per rollout. Naively reusing databases without tracking evolving states across multi-turn interactions causes trajectory contamination and cross-trajectory interference, which severely destabilizes policy updates. To overcome these bottlenecks, we introduce , a novel four-stage agentic RL curriculum that alternately enhances foundational agent reasoning and trajectory-adaptive on-policy exploration to circumvent the instability inherently associated with direct end-to-end learning. Crucially, to support this curriculum, BIRD-RL is driven by a , which enables safe database replica reuse while preserving trajectory isolation. This design improves spatial RL efficiency by , successfully enabling scalable agentic RL for complex data-centric tasks without needs of Kubernettes. Empirical evaluations on the BIRD-CRITIC-SQLite benchmark demonstrate that our 7B and 14B models achieve Success Rates of , respectively, performing on par with or exceeding strong proprietary model-based agents such as Claude-Opus-4.6 Agent (48.20%). Furthermore, BIRD-RL successfully trains unified SLMs capable of both advanced SQL debugging and generation, highlighting its broader potential for multi-objective optimization across heterogeneous, real-world data-centric tasks.
PaperID: 4266, Poster
Authors:
Jielun Zhong, 王迪, Jisheng Dang, Leigang Qu, Bimei Wang, Jingwen Zhao, Wencan Zhang, Hong Peng, Bin HuAbstract: Text-guided video editing requires applying a semantic change while keeping source motion, layout, and unchanged content stable. Recent inversion-free flow editors derive effective edit signals from target-source velocity differences, yet typically inject these signals directly at each solver step. In neural video sampling, however, this direct execution rule overlooks an important distinction: the velocity difference is an online, state-dependent observation rather than a fixed endpoint-displacement command. As the target branch evolves under previous injections, direct injection may accumulate temporal fluctuations and off-target perturbations, producing motion drift, identity flicker, or unintended changes in weak-response regions. We propose FlowTrack, a causal execution controller for inversion-free flow video editing. FlowTrack leaves the pretrained generator and raw edit signal unchanged, but executes the signal through a carried displacement-like direction and response-aware attenuation. We derive this update as the closed-form solution of a local execution objective that balances current signal tracking, temporal carrying, and weak-response suppression. On the full FiVE-Bench protocol, FlowTrack improves motion fidelity, unedited-region fidelity, and structure consistency over strong training-free diffusion- and flow-based editors while maintaining competitive target-prompt alignment. Code is available at https://anonymous.4open.science/r/FlowTrack.
Abstract: Many causal quantities are only partially identifiable due to the inherent missingness of potential outcomes, and the associated partial identification (PI) sets can be obtained by solving an optimal transport (OT) problem. Covariates often provide additional information about the potential outcomes and thus yield tighter PI sets, which can be obtained via conditional optimal transport (COT). However, COT-based PI set estimators are susceptible to the curse of dimensionality in the covariates, which precludes the asymptotic normality and hinders statistical inference. In this paper, we exploit the smoothness in the marginal densities of covariates and potential outcomes, and develop a wavelet-based primal approach for COT that attains a faster convergence rate. Moreover, for quadratic cost functions, we establish a stability result for COT and prove asymptotic normality of the proposed estimator, enabling valid statistical inference for the PI set. Empirically, we validate the estimation and inference performance of our approach through numerical experiments in comparison with existing benchmarks.
PaperID: 4268, Poster
Abstract: Hybrid model-based reinforcement learning (MBRL) integrates lookahead planning with actor-critic optimization for exceptional sample efficiency. However, this decoupled architecture suffers from a critical planner-actor mismatch. Existing policy alignment methods face a structural dilemma: forward KL triggers mode-covering objective conflicts, while reverse KL relies on brittle proxies that bottleneck expressivity. To address this challenge, we propose ), a unified in-sample extraction framework. At its core, we derive a first-order geometric alignment metric from the continuous-time Hamilton-Jacobi-Bellman (HJB) equation to evaluate action evolution against the local value gradient. By embedding this metric as an absolute modulation weight, GEAR eliminates spurious imitation and transforms the inherently diffuse forward KL into a focused, mode-seeking objective. Evaluations on the challenging HumanoidBench suite demonstrate GEAR achieves significantly higher sample efficiency and asymptotic performance than state-of-the-art baselines. The project code is available at https://anonymous.4open.science/r/GEAR-5E80.
PaperID: 4269, Poster
Abstract: Open World Object Detection (OWOD) is a challenging task that requires detectors to recognize known categories while discovering unlabeled objects and incrementally incorporating them as new categories. Existing methods mainly rely on known-class features to recall unknown objects, while overlooking the semantic pull of known classes in the attribute space and its impact on the interpretability of unknown detection. In this paper, we propose EviAttr-OW, a novel evidential attribute reasoning framework for OWOD that discovers potential unknown objects with interpretable attribute evidence. Specifically, we construct a class-agnostic attribute space that decouples attribute representations from known-class bias and provides a more open attribute description basis for potential unknown objects. We then map candidate-region attribute responses into a Dirichlet evidence distribution, producing known-class predictions and evidential uncertainty to measure the reliability of known-class support. Finally, we derive Known-Class Evidence Deficiency from this evidence distribution and combine it with object evidence, enabling the model to identify unknown objects with evidence-supported objectness and insufficient known-class support. Experiments on the Real-World Object Detection (RWD) benchmark across five real-world application datasets show that EviAttr-OW consistently outperforms existing state-of-the-art (SOTA) methods, achieving +7.3 mAP on unknown classes.
PaperID: 4270, Poster
Abstract: As foundation models scale toward fusing more heterogeneous visual streams, understanding how diverse encoders interact under joint training becomes a prerequisite for principled design. Yet large vision-language models (LVLMs) currently lack the tools to do so, and parameter-efficient encoder configurations remain hard to identify before training. To re-examine encoder roles under joint training, on the 16-benchmark Cambrian-1 suite we retrain and evaluate all 31 non-empty subsets of five common vision encoders under a unified pipeline (\approx20k GPU-hours total), and report three findings. First, retraining each subset from scratch reveals encoder rankings that differ from those obtained by masking encoders on a fixed checkpoint, including which encoder ranks first overall. Second, we decompose each encoder's contribution into two axes, \emphCapacity, the score an encoder reaches on its own, and \emphNecessity, the drop when it is removed from the full pool. The two axes are not interchangeable. Pairing the two highest-Capacity encoders is suboptimal, while pairing a high-Capacity anchor with an adaptive complement matches the full five-encoder model. Adding further encoders beyond this pair yields only marginal gains. Third, at fixed parameter count, per-encoder pre-projector effective rank explains the residual score variation. The strongest pairs combine an anchor whose rank survives joint training with a complement whose rank \emphexpands under it, suggesting that higher-rank, less-collapsed projector inputs correspond to a more favorable optimization regime at the encoder–projector interface. Together, the Capacity–Necessity decomposition and the pre-projector rank analysis, along with comprehensive evaluation through retraining, expose a methodological gap in multi-encoder LVLM design, and offer concrete primitives for closing it.
Authors:
Tuo An, Jindou Jia, Gen Li, Jingliang Li, Chuhao Zhou, Pengfei Liu, Bofan Lyu, Jiaqi Bai, Xinying Guo, Geng Li, Jianfei YangAbstract: World models aim to improve robotic decision making by predicting the consequences of actions. However, in practice, their predictions often become unreliable once the robot encounters states outside the training distribution, limiting their effectiveness at deployment. We observe that execution itself provides a natural but underutilized signal: after each action, the robot directly observes the true next state, revealing the mismatch between predicted and actual outcomes. Building on this insight, we propose feedback world model, a new paradigm that closes the loop between prediction and observation at inference time. Instead of treating the world model as a static open-loop predictor, our method maintains a lightweight feedback state that is updated online to iteratively correct future predictions, compensating for model errors using real-time observations without additional training data or parameter updates. We show that this process can be interpreted as a latent-space observer and admits convergence guarantees under mild conditions. We further introduce action-aware guidance to better translate corrected predictions into control by emphasizing action-controllable components while suppressing irrelevant variations. Experiments on LIBERO-Plus, Robomimic, and real-world manipulation tasks demonstrate that our method substantially improves both prediction accuracy and policy performance under distribution shift. In particular, it reduces world model prediction error by up to 76.4% and improves out-of-distribution (OOD) success rate by 30%. These results show that incorporating real-time feedback at inference time provides a simple yet powerful alternative to static world modeling.
PaperID: 4272, Poster
Authors:
Zhaolu Kang, Tailong Luo, Chenxin Li, Zhenyu Yu, Fengyu Zhou, Jiachen Qian, Lei Wei, Shuang Chen, Jiachen Li, Yingjie He, Eric Hanchen Jiang, Rongchao Zhang, Zhengtao Yao, Hoi L Lee, Guansu Wang, Kaiyue ZhouAbstract: Long chain-of-thought reasoning improves vision-language models (VLMs), but it also exposes a temporal grounding failure: as generation unfolds, models may drift from image evidence while reinforcement learning still provides only a final-answer reward. We study this gap through visual revisit, a temporally localized reactivation of image evidence during reasoning. Across three VLM backbones, revisit peaks are predictive of correctness and causally linked to performance: masking them degrades accuracy more than masking non-peaks, while injecting revisit-like peaks into failed rollouts partially restores correct answers. Control analyses show that this signal is not explained by raw attention strength, uncertainty, logit margin, or hidden-state norm. We propose RECAP, a GRPO-compatible training method that turns detached revisit traces into a credit-assignment signal. RECAP uses revisit-conditioned advantage estimation, dynamic discounting, and gated value prediction to propagate final rewards through visually important reasoning steps, without extra rewards, additional rollouts, or inference-time changes. Across Qwen2.5-VL-7B, InternVL3-8B, and LLaVA-OV-7B, RECAP consistently improves over GRPO and VPPO, yielding +6.1--+6.5 HallusionBench gains over the SFT base while preserving reasoning and text-only performance. Results on larger LoRA-tuned models, Chinese zero-shot benchmarks, visual perturbations, and human evaluation further demonstrate robust improvements in image-grounded reasoning.
PaperID: 4273, Poster
Authors: feng jiahe
Abstract: GRPO-style reinforcement learning has advanced vision-language reasoning, but its sequence-level rewards provide limited credit assignment for individual reasoning tokens. On-policy distillation (OPD) addresses this issue by providing dense supervision on student-generated trajectories, yet standard OPD assigns similar importance to all teacher-preferred tokens, including template phrases and visually irrelevant continuations. We introduce Visual Counterfactual On-Policy Distillation (VC-OPD), a grounding-aware distillation method that prioritizes tokens whose correctness depends on image-specific evidence. After a teacher-supervised warm start that stabilizes response formats and reduces noisy rollouts, VC-OPD performs on-policy distillation with two teacher evaluations for each student trajectory: one under the original image and one under a counterfactual image with corrupted instance-specific evidence. The original-image teacher distribution determines the distillation target, while the original--counterfactual teacher gap estimates token-level visual dependence and softly reweights the loss. To avoid unstable response-level update scales, VC-OPD further applies mass-preserving normalization to the token weights. This design preserves the dense supervision of OPD while shifting learning capacity toward visually grounded reasoning tokens. Experiments on multimodal reasoning benchmarks show that VC-OPD improves visual reasoning performance and produces more effective token-level distillation.
PaperID: 4274, Poster
Authors:
Navid Mohammadi Foumani, Soheila Ghane, Nam Nguyen, Mahsa Salehi, Geoffrey Webb, Geoffrey MackellarAbstract: Foundation models for EEG analysis are still in their infancy, limited by two key challenges: (1) variability across datasets caused by differences in recording devices and electrode configurations, and (2) the low signal-to-noise ratio (SNR) of EEG, where neural activity is often buried under artifacts and non-brain sources. To address these challenges, we present EEG-X, a device-agnostic and noise-robust foundation model for EEG representation learning. EEG-X introduces a location-based channel embedding that enables robust transfer across heterogeneous EEG devices and electrode layouts. To improve robustness against noise, EEG-X employs a novel noise-aware masking-reconstruction strategy that reconstructs artifact-cleaned signals rather than raw noisy EEG and introduces a novel Dictionary-based Convolutional Transformation (DiCT) that improves reconstruction-based pretraining by comparing signals in a structured feature space instead of directly in the raw signal space. Experiments across datasets collected from diverse EEG devices show that EEG-X consistently outperforms state-of-the-art methods across multiple downstream tasks and demonstrates strong cross-domain generalization when pretraining and downstream datasets differ in electrode layouts and recording configurations paving the way towards foundation models. The models and code are available at :https://anonymous.4open.science/r/EEG-X-CEDE/README.md
PaperID: 4275, Poster
Authors:
Xule Gao, Xi Wang, Zehua Zang, Rui Wang, Changwen Zheng, Chuxiong SunAbstract: Vision-language-action (VLA) models have shown strong promise in robotic manipulation, yet their post-training still predominantly relies on trajectory-level supervised imitation over robot demonstrations. While effective for learning feasible actions, this paradigm supervises policies to reproduce expert actions at each step, biasing them toward demonstration-specific action paths and providing limited guidance for recovering toward task-progress states once execution deviates from demonstrated trajectories. To address this limitation, we propose OGPO, an Offline Goal-conditioned Policy Optimization framework for recoverable VLA post-training. Instead of treating offline demonstrations merely as step-wise action labels, OGPO adaptively relabels trajectories with semantic key states and constructs goal-conditioned process rewards to optimize the policy for reaching these task-progress states. In this way, OGPO shifts VLA post-training from trajectory-level action imitation to key-state reachability learning, enabling policies to acquire a more robust ability to complete goal states from feasible off-demonstration states without requiring additional online interaction. Experiments on LIBERO and MetaWorld show that OGPO consistently improves VLA policy performance over standard supervised fine-tuning. Moreover, evaluations on LIBERO-Plus and our proposed LIBERO-ReAct, a mid-execution perturbation benchmark, further demonstrate that OGPO achieves stronger robustness and goal-directed recoverability under distribution shifts and recoverable off-demonstration states.
PaperID: 4276, Poster
Abstract: Editing a local segment in a talking head video without reshooting the entire scene remains challenging and underexplored. Given local edits, such as insertion, deletion, or substitution, talking head video editing must synthesize a replacement segment while leaving the rest unchanged. This requires variable-duration local rewriting while preserving identity, unedited regions, and boundary continuity, which standard speech-driven generation and lip synchronization do not directly address. Moreover, direct supervision is infeasible, as it requires paired videos of the same person and scene differing only in a local spoken segment, which does not exist in the real world. We instead formulate talking head video editing as facial motion infilling, a self-supervised pretext task that recovers masked facial motion from speech and surrounding motion context in ordinary video--speech pairs. The key insight is that local video editing can be simulated during training by masking a motion span and reconstructing it from speech and visible motion context. Based on this formulation, we introduce FacEDiT, a mask-controlled talking head model with local temporal attention bias and temporal smoothness regularization for improved lip--speech alignment and transition continuity. We also introduce FacEDiTBench, the first benchmark for talking head video editing, covering diverse edit types and lengths with dedicated evaluation metrics. Extensive experiments show that FacEDiT produces accurate, speech-aligned edits with strong identity preservation and seamless boundary transitions. Beyond editing, the same facial motion infilling model extends to portrait animation and lip synchronization by simply changing the mask pattern, establishing FacEDiT as a unified framework for talking head video editing and generation. We will release the code and data upon acceptance.
PaperID: 4277, Poster
Abstract: Unifying visual perception and generation poses a fundamental dichotomy: perception extracts semantics but discards details, while generation synthesizes fine-grained structures. Thus, existing works rely on decoupled architectures, as forcing these opposing information flows into one network inevitably triggers task interference and representation collapse. To break this bottleneck, we introduce the Visual Unified Model (VUM), a native framework that unifies both capabilities within a single architecture. Our key insight is that visual perception and generation can be unified by re-envisioning Masked Image Modeling and Diffusion Processes as complementary facets of data recovery. By establishing a shared degradation space that integrates both masking and noising, VUM optimizes a joint signal recovery objective: simultaneously predicting masked patches and denoising corrupted signals. This objective enforces the dual internalization of semantic abstraction and generative synthesis. Using a shared, unified backbone, VUM delivers highly competitive performance across a diverse spectrum of tasks, encompassing image generation, visual recognition and dense prediction tasks. Codes and pretrained checkpoints will be made publicly available.
PaperID: 4278, Poster
Abstract: Reliable prediction of material properties from atomistic structures requires machine learning models that respect SE(3) symmetry. Frame-based methods address this challenge by constructing equivariant coordinate systems that align SE(3)-equivalent structures into unified representations. While attractive for their flexibility and efficiency, existing frame-based methods face two major limitations: a single frame type is often insufficient for heterogeneous atomic environments, and frame constructions can become unstable near highly symmetric or degenerate configurations. To address these issues, we propose the Multi-Frame Adaptive Network (MFAN), which leverages multiple frame types and introduces a local adaptive mechanism to combine them according to atomic environments, enabling different geometric references across atoms and complementary local-global information for each atom. Building on this multi-frame architecture, we further incorporate a frame-quality-aware weighting scheme that downweights unreliable frames before degeneration, thereby improving the continuity of the resulting representation. Experiments on crystal property prediction benchmarks demonstrate the leading performance of MFAN, with consistent gains from adaptive multi-frame learning. Additional force prediction experiments show improved robustness in continuity-sensitive settings.
Abstract: Large language models (LLMs) are increasingly used as interactive agents, but optimizing them for long-horizon decision making remains difficult because current methods are largely purely reactive, which weakens both exploration and credit assignment over extended trajectories. In this work, we present Strategic Trajectory Abstraction (StraTA), a simple framework that introduces an explicit trajectory-level strategy into agentic reinforcement learning (RL). StraTA samples a compact strategy from the initial task state, conditions subsequent actions on that strategy, and trains strategy generation and action execution jointly with a hierarchical GRPO-style rollout design, further enhanced by diverse strategy rollout and critical self-judgment. Experiments on ALFWorld, WebShop, and SciWorld show that StraTA consistently improves both sample efficiency and final performance over strong baselines. StraTA reaches success rates of 93.1% on ALFWorld and 84.2% on WebShop, surpassing the strongest baselines by 2.3% and 11.4% respectively. On SciWorld, StraTA attains a 63.5% overall score, outperforming frontier closed-source models by 6.1% and prior RL baselines by 6.5%.
Abstract: Inference-time selection methods, such as Best-of-N, improve generation by sampling a pool of candidates and selecting the top completion according to a reward model. Distillation seeks to amortize this procedure into a single policy by replacing raw rewards with in-pool ranks and learning a policy that upweights higher-ranked completions. However, existing rank-based policies typically use smooth full-support reweighting, so low-ranked completions receive less mass but remain in the target support. Although a sharper reweighting reduces lower-tail mass, it also increases reliance on brittle ranking at the top made by a single reward model. We propose TUP: a Truncate-bad, Upweight-good Policy that removes low-ranked completions from the support and reweights only the retained upper tail with a tunable sharpness. TUP admits a closed-form, prompt-independent normalization and can be trained fully offline via binary cross-entropy, using shifted-truncated win-rates as soft labels and distilled-to-reference log-likelihood ratios as logits. Theoretically, under certain assumptions, we show that for any unknown oracle reward, the best monotone rank-reweighting can be matched by a lower-tail truncation rule, providing formal support for removing the lower tail rather than merely downweighting it. Empirically, we show that TUP is competitive with strong offline alignment baselines.
Abstract: Cross-entropy (CE) loss is central to deep learning, but existing theory often relies on simplifications—such as squared loss or convex models—that miss key aspects of CE optimization. In this work, we study multi-class CE dynamics using a two-layer linear network with orthogonal inputs, the simplest non-convex setting where the CE implicit bias remains unresolved. This coincides with the unconstrained features model used to study neural collapse (NC). Our analysis is based on a key observation: Hadamard initialization diagonalizes the softmax operator. This allows us to extend the spectral initialization framework that Saxe et al. (2013, 2019) developed for squared loss. We prove convergence to NC under spectral CE training and give the first finite-time analysis in this setting via an explicit Lyapunov function that decreases monotonically to NC despite spurious critical points. We further identify CE-specific phenomena absent under squared loss—coupling, non-monotonic convergence, qualitative dependence on the number of classes, and exponential slowdowns. We also characterize the role of network width and show empirically that spectral dynamics qualitatively model small random initialization.
PaperID: 4282, Poster
Abstract: Neural operators can efficiently learn solution maps for partial differential equations, but standard architectures often violate fundamental physical structure during inference, leading to unstable long-horizon rollouts. We introduce the Symplectic Fourier Neural Operator (SympFNO), which is composed of three symplectic layers that embed structure preservation directly into Fourier-space and physical-space transformations, thereby preserving the geometric structure of Hamiltonian dynamics. Our theory provides a rigorous foundation for the architecture by showing that the structure-preserving design is mathematically consistent with Hamiltonian evolution and inherently more efficient than unconstrained operator parameterizations. Across five well-known Hamiltonian PDE problems, SympFNO achieves stable long-term rollouts, substantially lower trajectory error, more accurate conservation of physical invariants, and up to two orders of magnitude fewer parameters than state-of-the-art baselines, namely FNO and PINO.
Abstract: Reconstructing interactive, simulation-ready 3D scenes from a single image is a critical bottleneck for robotic manipulation. While recent single-image lifters recover plausible per-object shapes, composing them yields scenes that collapse under physical simulation due to interpenetrating, hovering, or sinking objects. Existing physics-aware methods address this strictly as a post-hoc layout correction, leaving the underlying geometric errors unresolved. To address this, we introduce SimuScene, a compositional 3D reconstruction pipeline that puts physics in the loop of shape and layout estimation. Rather than using physics merely for layout cleanup, we utilize the physics engine as a diagnostic measurement tool during the generative process itself. By diagnostically simulating reconstructed objects under gravity, we convert penetration and support failures into quantitative correction signals that drive gravity-axis stretching and amodal shape resampling. This physics-informed feedback loop mitigates accumulated reconstruction errors and produces a stable, simulation-ready compositional 3D scene. Extensive experiments demonstrate state-of-the-art performance on physical stability and geometric alignment benchmarks. We further highlight SimuScene's utility by deploying reconstructed environments in humanoid control and robot-arm manipulation tasks.
Authors: Farhad G. Zanjani, Herbert Cai, Amirhossein Habibian
Abstract: We present SetDiff, a geometry-grounded multi-view diffusion framework that enhances novel-view renderings produced by 3D Gaussian Splatting. Our method integrates explicit 3D priors, pixel-aligned coordinate maps and pose-aware Plucker ray embeddings, into a set-based diffusion model capable of jointly processing variable numbers of reference and target views. This formulation enables robust occlusion handling, reduces hallucinations under low-signal conditions, and improves photometric fidelity in visual content restoration. A unified set mixer performs global token-level attention across all input views, supporting scalable multi-camera enhancement while maintaining computational efficiency through latent-space supervision and selective decoding. Extensive experiments on EUVS, Para-Lane, nuScenes, and DL3DV demonstrate significant gains in perceptual fidelity, structural similarity, and robustness under severe extrapolation. SetDiff establishes a state-of-the-art diffusion-based solution for realistic and reliable novel-view synthesis in autonomous driving scenarios.
Authors: Xin Chen, Cynthia, Jie Zhang, Florian Tramer
Abstract: Prompt injection is a critical vulnerability in LLM agents, yet the strongest methods still rely on human red-teamers and hand-crafted prompts. Adapting automated jailbreak optimizers does not close this gap: jailbreaks shape models toward generic compliance, while prompt injection requires emitting specific tool calls with correct parameters. The success signal is binary, and randomly sampled suffixes almost never trigger it—so standard optimizers have no gradient to follow. We present AutoInject, a black-box reinforcement learning (RL) framework that learns adversarial suffixes for prompt injection. A learned comparison-based reward scores each candidate against the best suffix seen so far, turning the binary signal into a dense reward suitable for RL optimization. The framework supports both online query-based attacks and offline-trained transferable suffixes that need no utility access at deployment, and incorporates a utility objective when task-completion feedback is available. On AgentDojo, AutoInject outperforms template attacks, GCG, TAP, and adaptive attack across production models, with statistically significant improvements under McNemar's test with p<0.05. Suffixes learned by AutoInject also break Meta-SecAlign-70B, a model fine-tuned specifically to resist prompt injection, where template attacks fail outright. The results establish an automated baseline for prompt injection and expose a gap between preference-based defenses and adaptive optimization-based attackers.
Abstract: Linear Temporal Logic (LTL) provides a rigorous framework for specifying long-horizon robotic tasks, yet existing approaches face a trade-off: model-based synthesis relies on accurate labeled transition systems, whereas learning-based methods often require online interaction, task-specific rewards, or specification-conditioned training. We study LTL-specified robotic planning and execution in a stricter offline, model-free setting, where the agent is given only fixed, task-agnostic trajectory fragments, with no dynamics model, task demonstrations, or online data collection. To address this setting, we propose SAGAS, a framework that combines the compositionality of symbolic synthesis with the data-driven reachability structure learned from offline trajectories. SAGAS first learns a reusable latent reachability graph and a frozen goal-conditioned executor from fragmented offline data. For each new LTL formula, it performs task-time semantic graph augmentation to ground state-defined propositions on the learned graph, and applies B\"uchi product search to synthesize a cost-aware accepting prefix--suffix waypoint plan executed by the frozen executor. By shifting formula-specific reasoning from policy learning to test-time graph augmentation and symbolic search, SAGAS enables zero-shot generalization to unseen, data-supported LTL specifications without task-specific reward design, policy retraining, or online interaction. Experiments on LTL task suites constructed from OGBench locomotion domains show that this design produces executable and cost-efficient prefix--suffix behaviors for diverse unseen LTL tasks from fragmented offline data.
Abstract: Distilled diffusion models accelerate image generation by reducing the number of denoising steps, but often suffer from degraded image quality. To mitigate this trade-off, test-time optimization methods improve quality, yet their iterative nature incurs substantial computational overhead and leads to slow inference, limiting practical usability. Recent hypernetwork-based approaches amortize this process during training, but still require costly noise modulation in high-dimensional latent spaces. In this work, we propose LENS (Low-frequency Eigen Noise Shaping), an efficient noise modulation framework that operates in a low-dimensional subspace. Our approach is motivated by the observation that low-frequency components of the noise largely determine the global structure and visual fidelity of generated images. Based on this observation, we provide a theoretical justification for restricting modulation to the low-frequency subspace and derive a principled training objective. Building on this, LENS employs a lightweight, standalone network to selectively modulate these components, enabling efficient and targeted noise modulation. Extensive experiments demonstrate that LENS achieves competitive image quality while reducing FLOPs by 400–700×, model parameters by 25–75×, and inference-time overhead by 10–20× compared to prior methods.
PaperID: 4288, Poster
Abstract: Recent advancements in audio generation have been largely driven by Transformer-based diffusion models. However, these models suffer from quadratic complexity of self-attention, severely bottlenecking the long-form audio synthesis. To overcome this limitation, we propose MFlowAudio, a novel latent audio generation framework that synergizes the continuous-time dynamics of Flow Matching with a custom-designed TFMamba backbone. TFMamba uses an innovative dual-scan mechanism: a TimeMamba module to capture long-range causal dependencies with linear complexity, and a FrequencyMamba module to model spectral correlations such as harmonic structures. Exploiting this structural foundation, we formulate the Stateful Flow Matching (SFM) paradigm. This framework inherently enables chunk-wise training and streaming generation, maintaining an \mathcalO(1) caching complexity without incurring extra computational overhead. To enable fine-grained controllable synthesis, we devise a novel guidance mechanism that neutralizes vector field collisions precipitated by on-the-fly semantic transitions of prompts. Comprehensive empirical evaluations confirm that MFlowAudio yields generative fidelity on par with state-of-the-art baselines while establishing significant computational efficiency, achieving an 81.8% acceleration in generation and facilitating perceptually seamless streaming synthesis with a constant, ultra-low latency of 1.8 seconds. Demo:https://huggingface.co/spaces/mflowaudio/MFlowAudio
PaperID: 4289, Poster
Authors:
Hongjae Shin, Jungho Kim, Donghyuk Kwak, Myeongjun Kim, Seunghoon Yu, Jiyong Oh, Jun Won ChoiAbstract: Large-scale annotation of vectorized high-definition (HD) maps requires substantial human effort, motivating local-to-global auto-labeling that incrementally constructs global maps from local vector predictions. In this formulation, limited reliability management allows inaccurate predictions to be accumulated and reused as priors, causing persistent error propagation across iterations. We propose ReVMap, a local-to-global framework for reliable vectorized global map construction. ReVMap introduces a reliability-aware map evolution strategy that unifies uncertainty-guided prior reuse with revisable global map updates. To realize this strategy, ReVMap estimates localization uncertainty for vectorized local map elements and propagates these reliability cues to the accumulated global map. The resulting uncertainty-aware global representation guides subsequent local prediction and facilitates revisable integration through learning-based local-to-global connectivity estimation and uncertainty-aware fusion. ReVMap achieves state-of-the-art results on nuScenes and Argoverse2, with 36.1 mGAP, 43.4 miECM, and 58.0 local mAP on nuScenes, and 62.1 mGAP, 62.5 miECM, and 77.0 local mAP on Argoverse2.
Abstract: achievable by any model on a given task. Estimating this quantity allows us to distinguish the irreducible part of the error from a deficiency of the model, telling us how much room for improvement remains. Recent work has shown that the Bayes error, or equivalently the optimal accuracy, can be estimated from soft labels in binary classification. However, accuracy is often a poor summary of performance in settings with severe class imbalance or noisy annotations, where metrics such as the balanced error rate (BER) and the area under the ROC curve (AUC) are more appropriate. We address this gap with two complementary contributions. We propose soft-label-based estimators for the optimal BER and AUC. We first consider the clean setting in which true soft labels and the class prior are known, and then extend the estimators to a more realistic setting in which the class prior is unknown and the observed soft labels are by an unknown order- preserving transformation. In the latter setting, we approximately recover the clean soft labels via isotonic regression with auxiliary hard labels, estimate the class prior with a clipped mean of the hard labels, and derive finite-sample error bounds for the resulting plug-in estimators. Since the optimum is unobservable on real datasets, evaluating any such estimator is itself nontrivial. We extend the FeeBee framework, originally proposed for evaluating Bayes-error estimators, to the optimal BER and AUC. The resulting procedure provides practical evaluation scores without requiring knowledge of the optimum, and applies to any estimator of the optimal BER or AUC, not only our proposed ones. Experiments on synthetic and real-world datasets validate both the estimators and the evaluation procedure.
PaperID: 4291, Poster
Abstract: In this work, we study transformer attention through the lens of spectral geometry and operator theory. We view each attention head as a functional map between Hilbert spaces of functions on the token sequence and derive a \emphToken Difference Operator, whose spectral structure controls how token-space information is routed to the output. We show that standard Euclidean spectra are structurally biased by sinks, conflating mass concentration with genuine routing capacity. By recasting token space in the intrinsic probability geometry induced by attention, the token difference spectrum disentangles sink effects from routing capacity and provides a spectral description of the dimensionality of the head output. This yields a unified framework for analyzing attention maps, explaining sinks, routing collapse, and output dimensionality within a single operator-theoretic framework. In practice, by grounding attention heuristics in spectral geometry, we develop a novel uncertainty quantification method for LLMs that provides a theoretically sound alternative to traditional, probability-based estimators, with the strongest gains on long-form summarizations
Abstract: We study the K-armed logistic bandit problem, where at each round, the agent observes K feature vectors associated with K actions. Existing approaches that achieve a rate-optimal \widetilde\cal O(\sqrtdT) regret bound rely heavily on context diversity assumptions, such as strict positivity of the minimum eigenvalue of a context covariance matrix. These assumptions, however, impose strong restrictions on the context process, as they rule out the situation where the context vectors are concentrated in a low-dimensional subspace. In this paper, we propose SupSplitLog, which, to the best of our knowledge, is the first algorithm for logistic bandits that achieves \widetilde\cal O(\sqrtdT) regret without any context diversity assumption. The key idea is to split the collected samples into two disjoint subsets when constructing estimators; one is used to compute an initial-point estimator, while the other is used to apply a Newton-type one-step correction procedure. The splitting rule is carefully designed to balance the accuracy requirements of the initial-point estimator and the one-step correction procedure. Moreover, SupSplitLog strictly improves on the existing algorithms in terms of the dependence on dimension d in the regret upper bound. Furthermore, SupSplitLog can be adapted simply to deduce a regret bound that grows with a data-dependent complexity measure, avoiding a direct dependence on d, which is favorable when the context vectors are concentrated in a low-dimensional subspace. We also provide experimental results that demonstrate numerically the superiority of our algorithm, validating the theoretical results.
Abstract: Parameter-Efficient Fine-Tuning (PEFT) commonly adapts pretrained weights through low-rank updates, and recent methods further exploit the singular value decomposition (SVD) of the base weight for initialization or subspace selection. However, these methods do not explicitly preserve the coupled geometry between the pretrained left and right singular bases. Motivated by recent minimum-perturbation theory, which shows that stable finetuning admits a coherent SVD rotation in which a single orthogonal Q acts on both the left singular basis U_0 and the right singular basis V_0, we prove a per-slice analogue: each row slice of W_0 can be adapted by a shared orthogonal rotation Q_i on its left basis U_i and right basis V_i together with a diagonal spectrum shift. We instantiate this form as CORA (Coherent Orthogonal Rotation Adaptation), which applies per-slice orthogonal rotations and a per-layer diagonal scale to the rank-r SVD truncation of W_0. CORA uses \tfrac12m(r-1) trainable parameters per linear layer, about 4× fewer than LoRA at the same rank. CORA outperforms LoRA, DoRA, PiSSA, and MiLoRA on commonsense reasoning and code generation while using ~ 8× fewer parameters.
Authors:
Laura Lützow, Simone Garatti, Marco Campi, Lars Lindemann, Matthias AlthoffAbstract: Conformal prediction constructs prediction sets with finite-sample coverage guarantees, but its calibration stage is structurally constrained to a scalar score function and a single threshold variable — forcing shapes of prediction sets to be fixed before calibration, typically through data splitting. We introduce multi-variable conformal prediction (MCP), a framework that extends conformal prediction to vector-valued score functions with multiple simultaneous calibration variables. Building on scenario theory as a principled framework for certifying data-driven decisions, MCP unifies prediction set design and calibration into a single optimization problem, eliminating data splitting without sacrificing coverage guarantees. We propose two computationally efficient variants: RemMCP, grounded in constrained optimization with constraint removal, which admits a clean generalization of split conformal prediction; and RelMCP, based on iterative optimization with constraint relaxation, which supports non-convex score functions at the cost of possibly greater conservatism. Through numerical experiments on ellipsoidal and multi-modal prediction sets, we demonstrate that RemMCP and RelMCP consistently meet the target coverage with prediction set sizes smaller than or comparable to those of baselines with data split, while considerably reducing variance across calibration runs — a direct consequence of using all available data for shape optimization and calibration simultaneously.
PaperID: 4295, Poster
Abstract: Language model loss follows remarkably regular scaling laws over model and data size, yet it remains unclear why the aggregate loss should exhibit a power-law form. Existing explanations often attribute this regularity to a heavy-tailed spectrum of pattern difficulty in natural language, but this view has not been directly validated at token-level granularity in large-scale real-data training. We present a token-level framework that deconstructs scaling laws into localized learning events of individual contextualized tokens. By fitting token loss trajectories with sigmoids, we show that token learning is concentrated in localized transitions, giving rise to a learning-time spectrum that dominates the scaling-law shape. Across more than one hundred pre-training runs on large and diverse real-language corpora with modern LLM architectures, scaling up to 6B parameters and 300B training tokens, the measured learning-time spectrum quantitatively reconstructs the validation loss derivative along the training-step T, data-scale D, and model-scale M axes. We further show that the same signal is actionable: by reshaping the training distribution according to when tokens become learnable, we alter the optimization trajectory and achieve 11% faster validation-loss reduction. These results provide direct empirical evidence that scaling laws are governed primarily by the distribution of token-level learning times, and that this distribution can be used not only to explain scaling behavior but also to improve training performance.
PaperID: 4296, Poster
Abstract: 3D anomaly detection (3DAD) aims to identify defective regions in point cloud data, serving as a critical component in industrial inspection systems. Existing methods are normality-centered — learning the distribution of normal samples and treating deviations as anomalies — without explicitly modeling what constitutes a defect. This leads to ambiguous decision boundaries with increased false positives and negatives, particularly in unified and cross-domain settings where diverse normal distributions further blur the boundaries. We propose a relational inconsistency modeling framework that characterizes defects as violations of geometric consistency among neighboring structures. Our approach learns category-agnostic defect cues through pseudo-anomalies designed as controlled relational violations, instantiated by two key modules: Edge-aware Graph Refinement (EGR) for encoding geometric relationships among local regions, and Cluster-Deviation Modeling (CDM) for identifying regions that are relationally incompatible within their structural peer group. Extensive experiments on Anomaly-ShapeNet and Real3D-AD demonstrate consistent improvements over prior state-of-the-art methods in both in-domain and cross-domain settings, validating the effectiveness of learning an explicit, relation-based defect criterion for 3D anomaly detection.
PaperID: 4297, Poster
Abstract: Synthesizing interpretable controllers for multi-agent systems requires optimizing discrete symbolic programs from sparse, expensive, long-horizon simulator feedback. Evolutionary program search preserves interpretability but often spends many evaluations on weak random variation, whereas deep multi-agent reinforcement learning can learn effective policies that are difficult to inspect. We propose GNES, a neural-guided evolutionary program search framework for symbolic controller synthesis. GNES casts controller design as a program-edit decision process: gene growth trees are program states, validity-preserving genetic operators are actions, simulator fitness provides the return, and policy--value graph neural networks learn edit signals over controller topology. Monte Carlo tree search performs look-ahead planning in controller-program space before simulator validation, closing the loop between evolution, learned guidance, and verified rollouts. Experiments on three decentralized AirSim swarm-control benchmarks show that GNES improves search efficiency and final controller quality relative to evolutionary and reinforcement-learning baselines, while automatically evolving compact, inspectable distributed velocity controllers. The discovered policies expose coordination motifs, such as density-normalized radial gains and motion-compensating terms, that differ from standard hand-designed swarm rules. Ablations and learned-search comparisons indicate that graph-structured policy--value learning and MCTS look-ahead both contribute to the gains.
Abstract: Memory-augmented LLM agents enable interactions that extend beyond finite context windows by storing, updating, and reusing information across sessions. However, training such agents with reinforcement learning in multi-session environments is challenging because memory turns the agent's past actions into part of its future environment. Once different rollouts write, update, or delete different memories, they no longer share the same intermediate memory state, making trajectory-level comparisons fundamentally unfair. This violates a key assumption behind group-relative methods such as GRPO, where rollouts are compared as if they were sampled from the same effective environment. Consequently, trajectory-level rewards provide noisy or biased credit signals for long-horizon memory operations. To address this challenge, we introduce Memory-R2, a training framework for long-horizon memory-augmented LLM agents. Its core algorithm, LoGo-GRPO, combines local and global group-relative optimization. The global objective preserves end-to-end learning from long-horizon trajectory-level rewards, while local rerollouts compare alternative continuations from identical intermediate memory states, yielding fairer group comparisons and more precise supervision for memory writing, updating, and retrieval. Beyond credit assignment, Memory-R2 jointly optimizes memory formation and memory evolution with a shared-parameter co-learning design, where a fact extractor and a memory manager are instantiated from the same LLM backbone through role-specific prompts. To make memory construction more controllable, we further formulate each session as a multi-step decision process by chunking the interaction and alternating the two memory roles. A progressive curriculum from 8 to 16 to 32 sessions stabilizes training as the memory horizon grows. Together, these components provide an effective training paradigm for memory-intensive LLM agents in long-horizon multi-session settings.
PaperID: 4299, Poster
Abstract: Recent advances in rectified flow-based image-to-3D generative models have enabled high-fidelity 3D asset generation. Building on this, a growing line of work has exploited these strong 3D priors for training-free stylization, transferring visual attributes from a reference image onto a generated 3D asset. However, existing methods enforce an all-or-nothing paradigm: color and texture are transferred jointly, with no mechanism to control them independently — a limitation we formalize as Disentangled 3D Stylization (Disen3D). To address this, we propose DiDE, the first training-free framework for Disen3D. Key to our approach is the observation that the structured latent space of image-to-3D models is overcomplete with respect to texture: texture information occupies only a small subset of the style-significant channels, leaving a free subspace available for independent color encoding. DiDE exploits this via a channel partition mechanism that processes a content image, a texture reference, and a color reference through dedicated branches and composes both style signals interference-free at every self-attention layer, preserving content geometry throughout. Experiments on Disen3D-Bench, our newly collected multi-reference benchmark, show that DiDE consistently outperforms 2D and 3D stylization baselines in color fidelity, texture transfer, and content preservation.
Authors:
Zechen Li, Keerthana Natarajan, Weizhi Zhang, Simon Lee, Yuwei Zhang, Max Xu, Menglian Zhou, Zeinab Esmaeilpour, Flora Salim, Mark Malhotra, Lindsey Sunden, Shwetak Patel, Yuzhe Yang, Ahmed MetwallyAbstract: Continuous glucose monitoring (CGM) provides a dense window into metabolic physiology, yet existing generic time-series and CGM-specific foundation models typically learn from entangled glucose sequences without explicitly capturing the temporal structure of glycemic dynamics. We present GlucoFM, a lightweight CGM foundation model that aligns irregular recordings to a 24-hour chronological grid, preserves observation masks, and decomposes glucose dynamics into slow physiological state and transient event streams, capturing low-frequency glycemic baselines and short-term deviations that may reflect acute physiological responses or sensor artifacts. GlucoFM is pretrained on 109,066 hours of unlabeled CGM recordings from 477 subjects with two complementary objectives: masked contextual latent prediction over fused daily representations and temporal dynamics prediction over state and event streams. Across four diverse cohorts and seven clinical prediction tasks, GlucoFM achieves the strongest frozen linear-probing performance among evaluated baselines, improving average PR-AUC by 4.1 points over the best CGM-specific foundation model. Its gains are most pronounced on core metabolic outcomes, leading PR-AUC on all diabetes-risk and \beta-cell dysfunction tasks and on 3 of 4 insulin-resistance tasks. GlucoFM also achieves the best overall cross-dataset transfer performance among evaluated methods and strong few-shot adaptation, highlighting physiology-aware decomposition as an effective inductive bias for transferable CGM representation learning.
PaperID: 4301, Poster
Abstract: Masked diffusion language models (MDLMs) generate text by unmasking tokens in parallel and have recently emerged as alternatives to autoregressive language models. They can be viewed as parallel decoders trained with a position-wise cross-entropy (CE) loss, the same setup as non-autoregressive translation (NAT). In NAT, CE-trained parallel decoders have been argued to be sensitive to small positional shifts, since CE penalizes them harshly. We ask whether CE-trained MDLMs are similarly sensitive to such shifts under iterative decoding. To probe this, we apply a controlled intervention that introduces them during decoding. On LLaDA-8B-Instruct with Arena-Hard, displacing as little as 1% of generated tokens by one position substantially reduces win rates against the unintervened model, showing that MDLMs are sensitive to such small shifts under iterative parallel decoding. Motivated by this, we adapt connectionist temporal classification (CTC), an alignment-flexible objective known to mitigate it there, to MDLM supervised fine-tuning. By relaxing the strict position-wise match that CE imposes, CTC gives the loss room to absorb small positional shifts; concretely, we modified CTC objective to use a special \ token that absorbs positional uncertainty between target tokens and output positions, and a updated collapse map that preserves target surface forms. Across four open-ended generation benchmarks, the resulting model consistently improves over both the original model and a matched cross-entropy-trained baseline, with statistically significant gains on all four. These results identify training-side alignment flexibility as a useful design dimension for MDLM SFT, complementary to the inference-time approaches explored in prior work.
PaperID: 4302, Poster
Abstract: Causal discovery is the task of leveraging observational data to uncover causal relationships between variables. Recent work has extended these methods to operate over clusters of variables to improve scalability in high-dimensions and enable reasoning over higher-level entities. These approaches have been limited by strong assumptions including causal sufficiency. In this work, we introduce an approach for causal discovery over clusters in non-Markovian systems. First, we extend theory of graphical models in a knowledge-based context, to motivate introduction of a novel graphical equivalence class that can accommodate unobserved confounding. Then, we present a sound algorithm for causal discovery of learnable relationships between clusters of variables.
PaperID: 4303, Poster
Abstract: Generative modeling of discrete data, such as graphs, underpins many scientific and industrial applications, including molecular discovery and materials design. In these domains, probabilistic inference is particularly valuable, as it enables composable generation and principled incorporation of desired constraints, such as structural or functional properties. Energy-based models naturally support this goal by capturing relative likelihoods and enabling composable inference by directly enforcing constraints during inference. However, discrete energy-based models typically struggle with efficient and high-quality sampling, as off-support regions often contain spurious local minima, trapping samplers and causing training instabilities, resulting in a fidelity gap compared to discrete diffusion models. To address this gap, we introduce Graph Energy Matching (GEM), a discrete generative framework inspired by the Jordan--Kinderlehrer--Otto (JKO) transport-map optimization perspective. GEM learns a permutation-invariant potential energy that simultaneously guides discrete transport from noise toward high-likelihood graph regions and refines samples within these regions. We further introduce a sampling protocol leveraging an energy-based switching strategy, seamlessly bridging rapid, gradient-guided transport and a local mixing regime for effective exploration. On molecular graph benchmarks, GEM matches or surpasses strong discrete diffusion baselines on most reported metrics. Beyond improving generation quality, GEM's relative likelihood modeling enables targeted exploration, facilitating compositional generation, property-constrained sampling, and interpolation between graphs.
PaperID: 4304, Poster
Abstract: LiDAR relocalization based on scene coordinate regression (SCR) degrades sharply when test trajectories traverse route segments absent from training, a practical challenge in long-range autonomous-driving deployment, where complete route coverage is rarely achievable. We present LINK, a confidence-guided framework for relocalization under partial scene coverage. LINK predicts an absolute pose hypothesis and scene features with an SCR backbone, estimates pose reliability from temporal pose-feature sequences, and uses a decision policy to arbitrate between direct absolute hypotheses and relative motion propagated from reliable historical estimates, thereby preserving globally referenced localization across uncovered segments. We also introduce UrbanUnseen, a city-scale benchmark with native known-to-unknown scene transitions, together with controlled partial-coverage protocols on Oxford and NCLT. Experiments show that LINK substantially improves robustness when training coverage is incomplete, consistently outperforms representative map-free baselines under partial coverage, and remains effective in fully known scenes.
PaperID: 4305, Poster
Abstract: Beam-search–based test-time methods provide an effective way to improve large language model (LLM) performance on long-horizon generation by pruning invalid reasoning paths early, leading to significantly improved reasoning efficiency and more favorable test-time cost scaling. Despite strong empirical success, the theoretical understanding of beam search remains limited. In this paper, we study the test-time compute guarantee of the commonly used beam search framework that uses the model's internal log-likelihood for intermediate scoring, while relying on an external reward model only after a complete response is generated. We first establish a lower bound for vanilla beam search, showing that at least \Omega(C^\star(x)^2) samples are required for the optimal response to survive, where C^\star(x) is the token-level coverage coefficient for prompt x. This motivates our modified confidence-filtered beam search (CF-Beam), which provably reduces the sample size by a factor of C^\star(x). We then show that the regret of CF-Beam is upper-bounded by the probability of rare failure events and the reward estimation error scaled by a path-level coverage coefficient, where the rare-failure term vanishes as per-step sampling increases. Our results highlight a fundamental advantage of beam search over sequence-level inference methods such as Best-of-N and Best-of-Majority. While the guarantees of these approaches typically involve coverage coefficients that grow exponentially with the horizon L, CF-Beam controls the dominant search-induced term through a token-level coverage coefficient that scales polynomially with L. Our numerical experiments further confirm that beam search is more robust on hard instances and under increasing reasoning horizons.
PaperID: 4306, Poster
Abstract: Activation steering is a lightweight way to control pretrained generative models: a learned map inserted into the model's activations can steer generation toward target behavior without updating the model weights. While effective for single target behavior, existing steering methods remain brittle when multiple behaviors are targeted. A style intervention followed by a safety intervention can produce a different result than the same applied in the reverse order, as the second intervention is evaluated on activations already shifted by the first. We trace this failure to coordinate sharing: existing transport-based steering methods learn different concepts in the same activation coordinates, so their interventions interfere by construction. To address this, we introduce teering), a geometric steering framework that learns a separate orthogonal subspace for each concept. Each concept is steered by an affine transport restricted to its own subspace, and orthogonality between subspaces makes these transports non-interfering by construction. The concept frame is trained directly on the Stiefel manifold, so Riemannian optimization preserves the orthogonality required for composition throughout learning. This structure gives an exact commutativity guarantee: any sequential ordering of single-concept interventions reproduces the simultaneous multi-concept update. Across 45 concept pairs and three open models, StiCAS drives order-dependent interference exactly to zero and roughly doubles joint-concept success over the strongest activation-transport baseline.
PaperID: 4307, Poster
Abstract: Graph Neural Networks (GNNs) represent valuable intellectual property, yet existing watermarking schemes primarily rely on OOD backdoor triggers that are susceptible to model pruning, fine-tuning, and distillation. To tackle this challenge, we present InvGNN-WM, which ties ownership to a model's implicit perception of a graph invariant, enabling trigger-free, black-box verification with negligible task impact. By training a scalar head to predict normalized algebraic connectivity on owner-private carrier graphs, ownership is embedded into the model's core reasoning logic rather than exogenous patterns. We provide guarantees for imperceptibility and robustness, and prove that exact removal is NP-complete under monotone decoders. Empirical evaluations across diverse node and graph classification datasets show that InvGNN-WM maintains clean task accuracy while outperforming trigger- and explanation-based baselines in watermark fidelity. Our method remains robust under unstructured pruning, fine-tuning, and post-training quantization, with clear recovery pathways under knowledge distillation.
PaperID: 4308, Poster
Authors: Albert Sawczyn, Piotr Bielak, Tomasz Kajdanowicz
Abstract: Large language models (LLMs) are increasingly used for knowledge base question answering (KBQA), where answering requires selecting entities from a question-specific knowledge-graph subgraph. Yet LLMs are known to hallucinate across tasks, and KBQA is no exception: even when we provide a graph as the knowledge source, the model may rely on parametric knowledge instead of graph evidence or perform invalid reasoning over the given relations. Such hallucinated answer nodes can limit the practical deployment of KBQA systems, especially in high-stakes domains such as healthcare. We formulate hallucination detection in KBQA as an answer-node classification problem and propose a lightweight graph-based framework that treats the answering LLM as a black box. KG-Guard represents each KBQA instance as an augmented graph. It initializes node features with semantic representations of KG entities, marks topic entities and LLM-proposed answer nodes with learned vectors, and connect a virtual question node to the topic entities. A graph encoder then produces verification-oriented node representations, and a small MLP classifies each proposed answer node using its graph representation together with the question embedding. Experiments on WebQSP, ComplexWebQuestions, and PUGG show that our detector achieves the highest F1 on all three benchmarks (82.0, 87.4, and 84.3), outperforming LLM-judge and sampling-based baselines, while having ~305× fewer parameters than the reference approaches. Beyond detection, the node-level feedback is actionable: when flagged answers are fed back to the KBQA system for iterative refinement, downstream KBQA F1 improves by 13.0-14.5 points and Exact Match by 16.9-17.6 points.
Abstract: Using prompted language models as classifiers enables classification in domains with limited training data, but misses some of the robustness and performance benefits that fine-tuning can bring. We study whether training on multiple classification tasks, each with its own prompt, improves performance on new domains with new classification prompts. We show that such training partially generalizes to adjacent domains, improving classification performance on tasks that are unseen during training. However, we identify specific edge cases where the finetuned models fail to follow prompts, such as when the classification prompt changes completely while the data domain remains the same as during training. We show that classification training can be mixed with general instruction following training, and that (when done well), such training keeps the benefits of classification training and mitigates its generalization failures. Surprisingly, we see that this no-thinking supervised classification training can generalize to with-thinking classification and summarization, suggesting that no-thinking classification training might be instrumentally useful in building other kinds of classifiers and monitoring systems.
PaperID: 4310, Poster
Abstract: Uncertainty quantification (UQ) for large language models (LLMs) aims to provide reliable measures of predictive confidence, yet current methods often fail to remain stable under simple, meaning-preserving perturbations. We identify an important source of instability: semantically equivalent paraphrases of the same input can induce substantial variability in predictive confidence, even for methods with formal guarantees such as conformal prediction. To address this issue, we propose a paraphrase-aware UQ framework that explicitly enforces invariance to semantic rewordings. Instead of relying on a single input, our approach constructs uncertainty scores by aggregating predictions across a set of its paraphrases, using a lightweight proxy model to produce calibrated and comparable outputs. This aggregation yields uncertainty estimates that are both more stable and more informative, while remaining compatible with conformal calibration techniques to retain coverage guarantees. Across multiple UQ benchmarks and model families, our method achieves nominal coverage with smaller and more stable prediction sets under paraphrase perturbations. Code is available at https://anonymous.4open.science/r/PA_Score-8C0D.
PaperID: 4311, Poster
Abstract: We study the problem of estimating the mean of a random vector in \mathbb R^d under \ell_p norm, assuming only the existence of a covariance matrix. In the Euclidean norm, the sample mean is already optimal in expectation, and robust estimators are usually needed only to obtain high probability tails. Our results show that this picture changes for \ell_p norms with p>2: robustification is necessary already for the expected risk. We prove that the minimax risk exhibits a sharp phase transition, controlled by the interaction between the sample size n, the dimension d, and the norm parameter p. In the large sample regime, the classical Gaussian rate remains achievable. In the complementary regime, heavy tailed distributions create a genuinely non-Gaussian obstruction. The optimal estimator is surprisingly simple: a coordinate-wise median of means whose number of blocks is chosen according to the geometry of the norm, rather than according to a confidence parameter. We also show that natural projection based median of means estimators for general norms, despite achieving the optimal confidence tradeoff relative to the sample mean benchmark of Lugosi and Mendelson (Probab. Theory Relat. Fields, 2019), can be polynomially suboptimal for \ell_p norms even in the constant probability regime.
Abstract: We study the fixed-budget max–min action identification problem in depth-2 max–min trees, which is an important special case of Monte Carlo Tree Search (MCTS). In this problem, a learner sequentially and adaptively selects T reward samples from leaves (depth-2 nodes) and then must recommend a subtree (depth-1 node) with the largest min-value among its children nodes. Motivated by approximate planning, we focus on \emph\varepsilon-good subtree identification, where any subtree whose min value is within \varepsilon of optimal maximin value is acceptable. Our main contribution is an \emph\varepsilon-agnostic algorithm—requiring no knowledge of \varepsilon—that nevertheless achieves error bounds with explicit instance-dependent dependence on \varepsilon. We show that, for every meaningful \varepsilon, the misidentification probability decays as \exp\big(-\tilde\Theta(\fracTH_2(\varepsilon) )\big), where H_2(\varepsilon) captures both cross-subtree and within-subtree gaps. In the special case where each subtree has a single leaf, the model reduces to standard multi-armed bandit identification, and our bounds recover (up to accelerating factors) the best-known \varepsilon-good guarantees associated with halving-style methods, while providing a new \varepsilon-good analysis and guarantee for the Successive Rejects algorithm in fixed budget best arm identification. On the lower-bound side, we present complementary positive and negative results. While there is a gap between the upper and lower bounds, we discuss the main technical challenges in obtaining a tighter lower bound, which tells us that the maximin action identification problem is quite different from the standard K-armed bandits. To our knowledge, this is the first provable algorithmic guarantee for fixed-budget maximin action identification.
PaperID: 4313, Poster
Abstract: Image-to-3D scene generation must recover both object geometry and plausible spatial relations. A common diffusion-based pipeline, exemplified by SAM 3D, generates each object separately and assembles the results. This preserves local details, but the lack of strong scene-level constraints often causes floating objects, interpenetrations, missing contacts, or implausible relative scales. Inspired by human designers who plan a layout before placing assets, we introduce \ours, a code-centric multi-agent framework. Given a single image and 2D localization cues, a vision-language model (VLM)-based Arranger Agent predicts executable Blender Python layout code with object identities and 3D transforms. A Modeler Agent binds each object to geometry through asset retrieval or image-to-3D object generation, and a Refiner Agent renders the scene and edits the code from visual feedback. The code representation exposes object names, positions, rotations, and scales, making generated scenes editable and interpretable. We train the code-generating agents using Blender Python datasets built from SAGE and 3D-FRONT, then further optimize them with GRPO rewards for executability and spatial alignment. Experiments in in-domain and cross-domain settings show that \ourMethod improves scene-level plausibility while preserving object fidelity.
PaperID: 4314, Poster
Authors: Luhang Shen, Hu Jingwei, Na Ying, Chunsheng Guo
Abstract: Most active learning (AL) methods estimate sample value through empirical criteria such as prediction ambiguity, diversity, or distributional coverage. We revisit AL from a spectral-geometric perspective by regarding the feature span of the current labeled set as the modeled subspace and its orthogonal complement as an operational residual subspace. The projected residual component of an unlabeled sample then reflects its relation to representation components weakly expressed by the labeled-feature spectrum. Based on this formulation, we propose a spectrally driven nullspace regulation framework that unifies spectrum-induced subspace decomposition, cross-cycle residual persistence estimation, and coverage-aware acquisition in the residual-coordinate space. This yields an acquisition criterion that favors samples whose residual components are both persistent across cycles and complementary to already selected samples. Experiments on visual benchmarks show competitive or improved sample efficiency over strong baselines. Further analyses support the proposed view by showing that residual components shrink anisotropically and that acquisition rankings vary substantially across backbones, indicating that sample value is conditioned on the current representation state rather than being purely data-intrinsic.
PaperID: 4315, Poster
Abstract: Reinforcement-learning-trained agents for retrieval-augmented generation retrieve content (passages, knowledge-graph fragments, or hyperedges) and supervise training with outcome rewards on the final answer. Across these paradigms the reasoning route through retrieved content never leaves the language model's hidden state, conflating declarative knowledge with the procedural knowledge that recurs across queries of the same logical type; content-only methods consequently plateau as reasoning depth grows. We then introduce PatternBloom, an agentic RAG framework that shifts the unit of retrieval from content (what to read) to logic (how to reason). A first RL stage scores the agent's constructed evidence graph by a frozen oracle's information gain through the Information-Density Reward (IDR), a size-normalized mutual-information surrogate. High-reward trajectories are distilled into a Graph Pattern Memory (GPM) of type-abstracted reasoning skeletons, and a second RL stage uses the Pattern-Augmented Reward (PAR) to couple policy to memory, turning procedural reasoning into a learnable, externalized structural prior. A 7B PatternBloom outperforms 27 baselines by +15.0 Avg-OOD EM, with the gain compounding with reasoning depth precisely because procedural knowledge, once distilled, is reused rather than re-derived at every query.
PaperID: 4316, Poster
Abstract: Test-time scaling (TTS) improves LLM agent performance by sampling K independent rollouts and selecting one with a verifier. In open-ended software engineering (SWE) agent settings, verification typically relies on code execution to obtain a direct correctness signal, but incurs substantial setup and testing cost. By contrast, execution-free verifiers offer a more scalable alternative that chooses from the submission alone. However, existing execution-free verifiers are largely submission-isolated: they fix rubrics before seeing the candidate batch, underuse the trajectory evidence already produced during agent rollouts, and rely on poorly calibrated pointwise scores for the final decision. As a result, they often miss the right candidate even when it lies within the pool. To address this gap, we propose a training-free and execution-free verifier for agent test-time scaling based on a See, Read, Compare (SRC) protocol, which treats the K candidates as mutual context rather than independent inputs. The See stage combines task understanding with cross-candidate contrast to induce an instance-specific rubric. The Read stage extracts structured behavioral features, and compresses each trajectory into an evidence-grounded summary along several cognitive dimensions: localization, hypotheses, interventions, and validation. The Compare stage performs rubric-guided pairwise selection through a two-phase hybrid tournament, preserving relative judgment with only linear complexity calls. Experiments show that SRC outperforms all compared execution-free baselines across SWE-bench Verified generator pools and generalizes to heterogeneous model pools and Terminal-Bench 2.
Abstract: Fine-scale-faithful neural simulation under fixed storage budgets remains challenging. Many existing methods reduce high-frequency error by improving architectures, training objectives, or rollout strategies. However, under budgeted coarsen-quantize-decode pipelines, fine detail can already be lost when the carried state is constructed. In the canonical periodic incompressible Navier--Stokes setting, we show that primitive and derived fields undergo systematically different retained-band distortions under the same operator. Motivated by this observation, we formulate , a general state-design framework that chooses which physical fields are carried and how storage budget is allocated across them under a calibrated channel model. Across the full time-dependent forward subset of PDEBench, DerivOpt not only improves pooled mean rollout nRMSE, but also delivers a decisive advantage in fine-scale fidelity over a broad set of strong baselines. More importantly, the gains are already visible at input time, before rollout learning begins. This indicates that the carried state is often the dominant bottleneck under tight storage budgets. These results suggest a broader conclusion: in budgeted neural simulation, carried-state design should be treated as a first-class design axis alongside architecture, loss, and rollout strategy.
Authors:
Jiazheng Zhang, Ziche Fu, Junrui Shen, Yunbin Zhao, Yunke Zhang, Zhiheng Xi, LongMa, Chenxin An, Zhihao Zhang, Shichun Liu, Dingwei Zhu, Shihan Dou, Shaofan Liu, han li, Wiggin Zhou, Aiden Adams, Tao Gui, Fei Huang, Qi Zhang, Xuanjing HuangAbstract: Policy entropy has emerged as a fundamental measure for understanding and controlling exploration in reinforcement learning with verifiable rewards (RLVR) for LLMs. However, existing entropy-aware methods mainly regulate entropy through global objectives, while the token-level mechanism by which sampled policy updates reshape policy entropy remains underexplored. In this work, we develop a theoretical framework of entropy mechanics in RLVR. Our analysis yields a first-order approximation of entropy change \Delta \mathcalH, giving rise to entropy polarity, a signed token-level quantity that predicts whether and how much a sampled update expands or contracts entropy. This analysis further reveals a structural asymmetry: Reinforcing frequent high-probability tokens triggers contraction tendencies, whereas expansive tendencies typically require lower-probability samples or stronger distributional correction. Empirically, we show that entropy polarity reliably predicts entropy change, and that positive and negative polarity branches play complementary roles in preserving exploration and strengthening exploitation. Building on these insights, we propose Polarity-Aware Policy Optimization (PAPO), which preserves both polarity branches and implements entropy control through advantage reweighting. With the empirical entropy trajectory as an online phase signal, PAPO adaptively reallocates optimization pressure between entropy-expanding and entropy-contracting updates. Experiments on mathematical reasoning and agentic benchmarks show that PAPO consistently outperforms competitive baselines, while delivering superior training efficiency and substantial reward improvements.
PaperID: 4319, Poster
Abstract: Gradient descent for matrix factorization exhibits an implicit bias toward approximately low-rank solutions, even in regimes where the iterates may grow unbounded. We study this behavior through a constrained factorization with diagonal scaling, separating bounded outer factors from explicit diagonal scale parameters. For positive semidefinite matrix recovery, we use the model X \approx UDU^\top, where U is constrained to a Frobenius norm ball and D is a nonnegative diagonal factor. Although this reparameterization preserves the stationary points of classical Burer-Monteiro factorization, its projected dynamics consistently recover truly (rather than approximately) low-rank solutions across a wide range of step sizes and initializations. Motivated by this behavior, we extend the construction to neural networks by inserting diagonal layers between Frobenius-norm-constrained outer layers. The resulting UDV models exhibit strong rank-revealing structure during training, with many diagonal components and associated columns collapsing toward zero and becoming naturally prunable. We exploit this structure through an SVD-based pruning procedure (UDVPruning) that identifies and removes redundant units during training. Across ViT and MLP-Mixer benchmarks, UDVPruning substantially reduces parameters and FLOPs while maintaining competitive test accuracy and often improving measured training time, inference latency, and energy consumption.
PaperID: 4320, Poster
Abstract: Medical illustration generation is not merely domain-specific text-to-image (T2I) generation, but a high-constraint visual reasoning problem where visually plausible images can be invalid due to anatomical and structural errors. Existing T2I models largely rely on one-pass generation and lack an internal mechanism to inspect and correct such biomedical inconsistencies. We introduce process, coupling generation with self-reflection and re-generation through “generate–reflect–refine” cycles. MedIGen is trained with three progressive stages: (1) large-scale medical illustration pretraining on 1.21M samples, (2) mixture training over generation, reflection, and refinement sub-skills, and (3) reinforcement learning with dual-level rewards that optimize both reasoning validity and final visual correctness. We further present , an expert-aligned, rubric-driven benchmark with 296 diverse tasks and 9,015 criteria, evaluating scientific accuracy, structural correctness, and semantic alignment beyond coarse visual plausibility. Experiments show that MedIGen substantially outperforms strong open-source unified and reasoning-based baselines, establishing a new open-source frontier. All resources will be open-sourced to facilitate future research.
PaperID: 4321, Poster
Abstract: High-compression tokenizers are essential for scaling latent image generative models. However, aggressive compression creates a fundamental tradeoff between reconstruction fidelity and generation efficiency: high compression image encoder always increase the learning diffucity of diffusion training, resulting slow model convergence. Recent representation autoencoders speed-up the diffusion training by improving the latent feature's expressive capability by replacing VAE encoders with pretrained semantic encoders, yet they are typically limited to moderate compression and lose pixel-level details necessary for faithful reconstruction. To achieve both high compression and fast diffusion training, we propose DC-SAE, a Decoupled Compact Semantic Autoencoder designed for high-compression image generation with accelerated diffusion model convergence. DC-SAE consists of two key components: (1) a macro-level architecture design that leverages semantic encoders to enable higher compression ratios, and (2) a pixel-level encoder that preserves low-level details, ensuring high-fidelity image reconstruction. We empirically demonstrate that DC-SAE performs strongly on image generation tasks, achieving both compact latent representations and efficient training dynamics. % Specifically, on ImageNet dataset with 512 × 512 resolution, DC-SAE achieves 32× spatial compression, with 29.79 PSNR and 3.05 gFID, substantially outperforming the previous state-of-the-art high-compression tokenizer baselines DC-AE by 13.5% and 59.7% on PSNR and gFID respectively, maintaining comparable throughput and faster diffusion model training convergence.
PaperID: 4322, Poster
Abstract: Text- and goal-conditioned human-object interaction (HOI) provides an intuitive interface for specifying high-level intents in controllable character animation and embodied AI. Mapping such intents to physically plausible, contact-rich behaviors naturally requires hierarchical planning and control, yet existing planner-controller systems remain largely open-loop and lack adaptation to physical execution feedback. To address this, we present InfCoiL, a closed-loop framework for physics-based HOI from coordinated planner-controller learning. InfCoiL integrates a morphology-aware flow-matching planner with a unified multi-morphology controller in an autoregressive plan-and-execute loop, enabling diverse, controllable, and executable interactions across diverse objects and morphologies of humanoid embodiments. To further mitigate planner-controller distribution shift, we alternatively fine-tune both components under closed-loop reinforcement learning. Particularly, we introduce Semantic Flow-GRPO, which optimizes the flow-matching HOI planner with semantic covariance over heterogeneous HOI features and group-relative rewards from physical rollouts, while training the controller for reliable execution under contact-rich dynamics. Extensive experiments demonstrate that InfCoiL achieves state-of-the-art performance in both motion quality and task success. The code will be released upon acceptance.
PaperID: 4323, Poster
Abstract: Open-vocabulary object detection requires recognizing both known and novel categories beyond a fixed label space. Existing open-vocabulary detectors have made strong progress by integrating visual features with language embeddings, but they are predominantly built on dense ANN-based architectures. In contrast, spiking neural networks (SNNs) have been studied for sparse visual computation, yet their use in open-vocabulary detection remains largely unexplored. This paper presents Omni-SpikeDet, a spiking open-vocabulary detector that investigates how text-guided visual recognition can be realized in a spiking framework. Omni-SpikeDet combines spike-based visual feature extraction with text-guided multi-scale alignment. It uses a spiking convolutional module with integer-spike training and spike-based inference to reduce the mismatch between continuous vision-language supervision and discrete spiking representations. It further introduces a Cross-Scale Text-Image Interaction (CTI) module that modulates multi-level spiking features with language embeddings for prompt-guided detection. Experiments on COCO and LVIS show that Omni-SpikeDet achieves competitive open-vocabulary detection performance with a compact model, obtaining 47.2 AP on COCO and 27.9 AP on LVIS with 4.2M parameters. Under a standard theoretical SNN energy model, Omni-SpikeDet has an estimated energy cost of 19.8 mJ, suggesting a favorable accuracy-efficiency trade-off.
Abstract: approach to scaling—which involves the insertion of \ tokens in the input stream to delay model responses—offers a unique advantage by increasing model expressivity while remaining highly parallelizable at both training and inference. The existing literature on training models to utilize \ tokens relies on the standard cross-entropy objective in which the model output is read out and evaluated only at the final step of a pause sequence. This approach provides no mechanism for the model to regulate its own processing or to signal readiness to respond, treating the additional compute steps as a static barrier rather than a resource to be used adaptively. We propose a supervised loss, ), framed as a sequential-decision problem, that trains a model to dynamically and autonomously scale the number of compute steps used for each input token. The model indicates the need for additional compute steps by emitting a special \ output, delaying its response via a . The model can abstain multiple times to obtain longer delays. Our experiments demonstrate that CYB significantly outperforms standard cross-entropy when introduced either in pretraining or fine-tuning, reducing perplexity and enhancing downstream accuracy with no additional computational or memory cost.
PaperID: 4325, Poster
Authors:
Shuwen Wei, Samuel Remedios, Zhangxing Bian, Shimeng Wang, Jinwei Zhang, Junyu Chen, Yihao Liu, Lianrui Zuo, Muhammad Faizyab Ali Chaudhary, Blake Dewey, Aaron Carass, Jerry L PrinceAbstract: Unpaired image translation is a critical preprocessing strategy for synthesizing missing sequences and harmonizing contrast variability in magnetic resonance (MR) imaging. However, in medical image processing, distribution-level translation is insufficient; it is vitally important that the anatomy is not distorted, corrupted, or lost after translation. Previous methods, like diffusion Schrödinger bridge matching (DSBM) achieve image translation, but do not preserve anatomy. We introduce the pseudometric Schrödinger bridge (PMSB), which trains two independent volume-preserving, isometry-regularized diffeomorphisms to map data from each domain to a latent anatomic pseudometric space. Then, DSBM is used within that space, explicitly minimizing the transport gap while preserving anatomy. Extensive experiments on the multi-contrast OASIS-3 dataset demonstrate that PMSB establishes a new state-of-the-art in structural consistency for cross-contrast synthesis. PMSB outperforms other leading methods and drastically reduces anatomical hallucinations, notably yielding a mean PSNR of 26.70 dB, a mean SSIM of 0.86, and a mean LNCC of 0.24.
PaperID: 4326, Poster
Abstract: Autoregressive graph generation is commonly formulated by linearizing a graph into a sequence via a flat walk over its nodes and edges. While effective for small molecular graphs, such representations do not scale gracefully to large, structurally complex biological graphs, where higher-order organization is both dense and functionally important. To address this limitation, we introduce a hierarchical graph abstraction together with a coarse-to-fine autoregressive pipeline that separates coarsening, tokenization, and decoding, and we study sequence constructions that expose hierarchy in the tokenization. Concretely, we compare a flat representation with two hierarchical alternatives: one that explicitly factorizes communities and their interactions, and another that induces hierarchy through a depth-first traversal. Across eight benchmarks, differences are modest on simpler domains but become pronounced on larger, more structurally complex settings such as cyclic peptides, proteins, and hard conditional-generation tasks. In these regimes, flat representations degrade in validity or distributional fidelity, while the traversal-based hierarchical representation is consistently more reliable; the explicitly factorized variant is generally weaker, suggesting that how hierarchy is exposed in the sequence matters. Finally, the same coarse-to-fine representation naturally enables motif-conditioned generation without architectural changes: on larger and more structurally complex graphs, conditioning on higher-level prefixes yields strong motif retention together with high validity, uniqueness, and novelty, whereas flat representations can exhibit severe mode collapse under comparable constraints.
Abstract: Scaling long-context capabilities is crucial for Large Language Models (LLMs). However, real-world data contain a large number of sequences with heterogeneous lengths. Existing training libraries for LLMs rely on static parallelism strategies, which suffer from severe load imbalance, redundant communication, and suboptimal hardware utilization under data heterogeneity. In this work, we propose Flexible Context Parallelism (FCP), an efficient parallelism strategy that adaptively reconfigures communication groups and context parallelism degrees during LLM training. We generalize more flexible non-power-of-two parallelism degrees and develop a polynomial-time algorithm to generate near-optimal parallelism strategies with only millisecond-level overhead per training batch. FCP is able to maintain high hardware efficiency even under extreme data heterogeneity. Experimental results demonstrate that FCP significantly outperforms Megatron-LM and DeepSpeed in both LLM and MLLM training, achieving up to 1.47 × speedup in average throughput while maintaining near-linear scaling efficiency across large-scale clusters. For extremely unbalanced batches, FCP even achieves 2.24 × speedup.
PaperID: 4328, Poster
Abstract: Post-LayerNorm Transformers have seen limited practical adoption due to enduring difficulties in scaling them to large depths. While prior research has focused on stabilising Post-LayerNorm by improving conditioning at initialisation, stability often deteriorates when models are trained with large learning rates, forcing additional compromises. In the decoder-only settings we study, existing Post-LayerNorm methods fail to outperform strong Pre-LayerNorm baselines. We propose KiteNorm, a novel normalisation method designed to break this trend. KiteNorm achieves stability through residual scaling and a regularisation technique based on hidden-state variance. Beyond stabilisation, KiteNorm learns separate scales for the skip and residual branches in each sublayer, improving performance. Across decoder-only Transformers up to 1B parameters, KiteNorm remains stable throughout training and consistently outperforms leading baselines, with scaling laws favouring KiteNorm across depth, width, batch size, and training budget. Ablations show that each component is necessary for the full stability and performance gains. We offer theoretical insights into KiteNorm through a new perspective on the gradient vanishing problem in Post-LayerNorm, linking the underlying hidden-state expansion to rank collapse.
PaperID: 4329, Poster
Abstract: Aligning large language models (LLMs) to diverse user preferences is fundamentally hindered by standard alignment paradigms that optimize for monolithic users. In this work, through empirical studies, we first discover a massive, untapped performance headroom for personalized generation through test-time scaling. We demonstrate that personalized generation is uniquely suited for test-time sampling methods like Best-of-N (BoN) because it can be viewed primarily as a candidate matching problem rather than a generator capability bottleneck. While standard reward models can theoretically exploit this headroom, their massive parameter counts introduce a prohibitive computational bottleneck. To overcome this limitation, we propose a parameter-efficient framework utilizing million-parameter scale Multi-Layer Perceptron (MLP) ranking models. Our personalized ranking model directly reuses the internal embeddings of the base generator with minimal overhead. By scaling train-time data to provide fine-grained personalized preferences, this million-parameter ranking model accurately scores large candidate pools and can seamlessly guide generation to reduce the cost of materializing N candidates. Extensive experiments on nine datasets across three different personalized generation settings demonstrate that our framework effectively exploits the discovered headroom. Remarkably, our million-parameter MLP performs competitively with billion-parameter reward models explicitly finetuned on the same tasks, while requiring only a negligible fraction of the inference cost.
Abstract: Diffusion distillation, exemplified by Distribution Matching Distillation (DMD), has shown great promise in few-step generation but often sacrifices quality for sampling speed. While integrating Reinforcement Learning (RL) into distillation offers potential, a naive fusion of these two objectives relies on suboptimal raw sample evaluation. This sample-based scoring creates inherent conflicts with the distillation trajectory and produces unreliable rewards due to the noisy nature of early-stage generation. To overcome these limitations, we propose GDMD, a novel framework that redefines the reward mechanism by prioritizing distillation gradients over raw pixel outputs as the primary signal for optimization. By reinterpreting the DMD gradients as implicit target tensors, our framework enables existing reward models to directly evaluate the quality of distillation updates. This gradient-level guidance functions as an adaptive weighting that synchronizes the RL policy with the distillation objective, effectively neutralizing optimization divergence. Empirical results show that GDMD sets a new SOTA for few-step generation. Specifically, our 4-NFE (Number of Function Evaluations) models outperform the quality of their multi-step teacher and substantially exceed previous DMDR results in GenEval and human-preference metrics, exhibiting strong scalability potential.
Abstract: Representation alignment has emerged as an effective approach to improve Multimodal Large Language Models (MLLMs) by regularizing their internal representations toward those of an external vision encoder. However, existing methods typically align a fixed layer of the language backbone, overlooking the fine-grained structure of Transformer models. In this work, we propose ), a method that enforces cross-modal alignment at the level of individual attention heads. Our approach is grounded in the Platonic Representation Hypothesis, focusing on preserving the , their local neighborhood relationships) across modalities. Following the Mutual K-Nearest Neighbor (MKNN) alignment metric, we introduce a contrastive objective that acts as a differentiable proxy for matching local structures. HeRA applies this objective during multimodal training to specific attention heads in the LLM, selected by their alignment score according to the MKNN metric. Counterintuitively, we find that aligning the aligned heads yields the largest gains. Extensive evaluations across multiple MLLMs and 18 benchmarks demonstrate that HeRA consistently improves performance on challenging vision-centric tasks and serves as an effective regularizer against visual hallucinations by naturally curbing the over-reliance on linguistic priors. Our code will be publicly released.
PaperID: 4332, Poster
Authors: Minsang Kim, Seung Baek
Abstract: Mixture-of-Experts (MoE) language models scale model capacity efficiently by activating a small subset of parameters per token. However, in long-context inference, memory remains a major bottleneck because the KV cache grows linearly with sequence length. Existing KV cache compression methods reduce this cost for dense transformers, but overlook a problem specific to MoE: KV eviction can change which experts are selected during answer generation, degrading the generation quality. We propose MoEZip, a routing-aware KV cache compression for sparse MoE language models. To estimate the sensitivity of expert routing to perturbation, we propose a novel metric based on the Fisher information matrix of routing distributions. MoEZip combines this metric with attention weights for routing-aware scoring of contexts. Finally, we compose a complementary scoring system that combines eviction methods for dense and sparse-MoE architectures. Experiments show that MoEZip achieves robust performance under aggressive cache compression on various benchmarks. At a KV retention ratio of 10%, MoEZip improves over the SoTA baseline by +24 and +20 points on average on LongBench and RULER, corresponding to 2.3× and 4.9× higher scores, respectively.
PaperID: 4333, Poster
Abstract: Reconstructing patient-specific cyclic cardiac motion at arbitrary phases is central to dynamic cardiac assessment, underpinning measurements such as ejection fraction, stroke volume, and regional myocardial deformation. Existing methods---pair-wise registration, time-conditioned decoders, and coordinate-based implicit fields---share a common implicit assumption: the cardiac cycle is treated as a sequence of discrete phases, with temporal structure stitched together through learned mappings supervised at the observed frames. Yet the cycle is, by physical fact, a closed periodic process. We propose CycleSpectra, a spectral motion representation that takes the cyclic structure as a starting point of the representation rather than as a property to be learned. Given a pair of sparsely sampled phase observations, the model learns the spectral structure of the underlying cyclic cardiac motion in a low-dimensional latent space, and recovers the displacement at any phase through a fixed analytic synthesis layer---time never enters the network as a learnable signal, only as a coordinate the readout takes by construction. To support evaluation, we collect a thin-slice phase-resolved 4D cardiac CT cohort and benchmark under both canonical and non-canonical input protocols, reporting whole-cycle motion-fidelity metrics alongside standard image and anatomical metrics. Under arbitrary two-frame input, CycleSpectra achieves the lowest ejection-fraction error and highest ventricular volume--time curve correlation among recent baselines, while remaining competitive on image fidelity. Beyond two-frame inference, the proposed framework supports test-time multi-frame latent fusion for higher accuracy, validated on both the in-house cohort and a cross-modality 4D cardiac MR cohort. Code will be available.
PaperID: 4334, Poster
Authors: Donghyeon Soon, Yoon JI- Il, Hyeondong Kim, Daehee Park
Abstract: Human motion prediction (HMP) aims to forecast future movement from an observed motion sequence, and recent methods have incorporated 3D scene geometry and object semantics to improve long-term plausibility. However, the destination of future motion remains ambiguous in cluttered scenes: the same observed motion can lead to different interaction targets or navigation modes. Existing methods either handle this ambiguity implicitly or rely on auxiliary signals like gaze that are hard to obtain in practice. To address this, we propose SGIP, which explicitly predicts where in the scene future motion will be grounded across both objects and open navigation regions, and uses it to condition future trajectory and full-body pose decoder. This scene-grounded interaction prior is learned via intent distillation: at training time only, a teacher model scores candidate scene regions by combining an LLM prompted with a textual description of the person's intent with geometric motion-scene features. A student model learns to reproduce these scores from only the observed motion and 3D scene, making intent and LLM unnecessary at test time. Across multiple benchmark datasets, SGIP outperforms prior scene-aware baselines on both trajectory and pose accuracy in a realistic deployment setting where neither gaze nor intent is available at inference, demonstrating effectiveness of the proposed framework.
PaperID: 4335, Poster
Abstract: Looped Transformers --- which repeatedly apply a shared transformer block --- are an architecturally natural fit for variable-length algorithmic tasks. Although they can exhibit strong length generalization beyond the length of training sequences, this behavior is brittle, yielding high out-of-distribution (OOD) variance, even across well-performing in-distribution solutions. We trace this variance to the spurious correlation in simple algorithmic tasks between sequence length and number of loops. Introducing stochasticity into the number of loops during training sharply reduces OOD variance and stabilizes predictions across inference-time loop counts. To improve upon heuristic randomization schemes, we further analyze RL-Halting as a learned stochastic schedule and find that it generally improves the accuracy--stability trade-off. Across binary addition, Dyck-1, Unique Set, and Copy, learned stochastic stopping often improves this trade-off but can also stabilize a suboptimal computation. Our work suggests that “when to stop” should be treated as a training-time design choice, not merely an inference-time computation-allocation rule.
Abstract: Learning an energy-based model from data samples is a central problem in machine learning. Many recent and popular methods, such as denoising score matching for training energy-based diffusion models, use stochastic interpolants to corrupt data samples at different noise levels indexed by a time variable. This defines a joint density over both the data space and time, and most methods learn its energy through either spatial or temporal differences. We identify distinct failure modes for both of these approaches. To solve them, we propose Spatiotemporal Noise-Contrastive Estimation (stNCE), a framework for learning the energy through joint spatiotemporal differences. stNCE unifies many existing methods and leads to new training objectives. Experiments on images and molecules demonstrate performance competitive with state-of-the-art density estimation methods.
Abstract: Molecular generation with diffusion models has emerged as a promising direction for AI-driven drug discovery and materials science. While graph diffusion models have been widely adopted due to the discrete nature of 2D molecular graphs, existing models suffer from low chemical validity and struggle to meet the desired properties compared to 1D modeling. In this work, we introduce MolHIT, a powerful molecular graph generation framework that overcomes long-standing performance limitations in existing methods. MolHIT is based on the Hierarchical Discrete Diffusion Model, which generalizes discrete diffusion to additional categories that encode chemical priors, and decoupled atom encoding that splits the atom types according to their chemical roles. Overall, MolHIT achieves new state-of-the-art performance on the MOSES dataset with near-perfect validity for the first time in graph diffusion, surpassing strong 1D baselines across multiple metrics. We further demonstrate strong performance in downstream tasks, including multi-property guided generation and scaffold extension.
PaperID: 4338, Poster
Abstract: Model distillation provides an alternative when computing resources are limited. How to transfer the reasoning ability from teacher model to student model influences the performance of model distillation. We develop a mathematical framework for reasoning distillation grounded in dynamical systems theory. By modeling the teacher model's residual stream as a discrete dynamical system, we define a precise notion of orbital stability: perturbations transversal to the teacher's reasoning trajectory should contract, while the tangential reasoning signal propagates without active suppression. We derive the Contractive Restoring Flow (CRF) loss from first principles and prove that, at the loss minimum, it achieves strict transversal contraction at a constant rate 1-\alpha, leaves the tangential dynamics unconstrained (tangential agnosticism), and produces a global basin of attraction with bounded steady-state error within the linearisation regime. Empirically, reasoning-trajectory distillation with the CRF loss outperforms other distillation methods, and shows greater robustness on long-chain reasoning tasks, supporting the predicted effect of contracting transversal error. The code is avaiable at: https://anonymous.4open.science/r/CRF-5945/.
PaperID: 4339, Poster
Abstract: In the real world, spatio-temporal forecasting systems inevitably encounter out-of-distribution (OOD) shifts, where temporal shifts evolve due to seasonal patterns or policy changes, and spatial dependencies shift with sensor deployment, removal, or topology reconfiguration. While many approaches attempt to address these challenges via invariant learning or causal adjustment, they implicitly assume that the contextual signals supporting prediction remain reliable. Through empirical studies and theoretical analysis, we find that this assumption often breaks because: (i) node-level context can become biased or incomplete under temporal and spatial OOD, yet existing methods lack mechanisms to assess such reliability; (ii) existing standard graph-based aggregation propagates information with fixed strength, potentially amplifying corrupted context across nodes. To this end, we propose REACT, a lightweight REliability-Aware Context-calibrated spatio-Temporal forecaster that explicitly models and controls context reliability under OOD. REACT decouples prediction from adaptation by first establishing a topology-independent predictive anchor via a graph-free linear encoder, ensuring transferable node-wise representations. It then constructs an availability-aware context that captures both observed signals and prior-based imputations, enabling explicit characterization of context completeness. Building on this, REACT introduces a reliability-guided calibration mechanism that dynamically modulates spatial support through uncertainty-aware gating, selectively leveraging trustworthy neighbors while suppressing unreliable propagation. Extensive experiments on ten real-world datasets with diverse temporal and structural shifts demonstrate that REACT consistently outperforms eleven baselines. These results highlight the importance of explicitly modeling and controlling context reliability as a fundamental principle for robust spatio-temporal forecasting under distribution shift. Our source code is available at https://anonymous.4open.science/r/REACT-A82B.
Abstract: Diffusion models are central to modern generative modeling, and understanding how they balance memorization and generalization is critical for reliable deployment. Recent work has shown that memorization in diffusion models is shaped by training dynamics, with generalization and memorization emerging at different stages of training. However, deployed diffusion models are often further distilled, introducing an additional training phase whose impact on memorization remains unclear. In this work, we analyze how distillation reshapes memorization behavior in diffusion models, taking consistency distillation as a representative framework. Empirically, we show that when applied to a teacher model that has memorized data, consistency distillation significantly reduces transferred memorization in the student while preserving, and sometimes improving, sample quality. To explain this behavior, we provide a theoretical analysis using a random feature neural network model [Bonnaire et al., 2025], showing that consistency distillation suppresses unstable feature directions associated with memorization while preserving stable, generalizable modes. Our findings suggest that distillation can serve not only as an acceleration tool, but also as a mechanism for improving the memorization-generalization trade-off.
Abstract: Layer pruning removes entire Transformer decoder blocks from large language models, but introduces a mismatch between the hidden state received by the next surviving layer and the distribution it was trained to process, leading to significant performance degradation. We propose Ghosted Layers, a training-free recovery module that addresses this issue by solving a boundary activation alignment problem. Our method derives a closed-form optimal linear operator from a small calibration set to reconstruct the activation discrepancy introduced by the pruned layers. We show that this solution corresponds to the unconstrained optimum of the alignment objective, whereas existing methods are restricted to constrained solutions over limited operator subspaces. Experiments across multiple LLM backbones and pruning strategies demonstrate that our method consistently improves accuracy and perplexity over prior training-free baselines, while preserving the efficiency gains of layer pruning.
Abstract: Markov decision problems are most commonly solved via dynamic programming. Another approach is Bellman residual minimization, which directly minimizes the squared Bellman residual objective function. However, compared to dynamic programming, this approach has received relatively less attention, mainly because it is often less efficient in practice and can be more difficult to extend to model-free settings such as reinforcement learning. Nonetheless, Bellman residual minimization has several advantages that make it worth investigating, such as more stable convergence with function approximation for value functions. While Bellman residual methods for policy evaluation have been widely studied, methods for policy optimization (control tasks) have been scarcely explored. In this paper, we establish foundational results for the control Bellman residual minimization for policy optimization.
PaperID: 4343, Poster
Authors:
Michaël Soumm, Alexandre F Montgieux, Yunlong HE, Pietro Gori, Alasdair NewsonAbstract: Contrastive Analysis aims to separate factors which are common between two data distributions from those which are salient to only one of them. Existing contrastive methods are based on generative models (e.g., VAEs or GANs) that often suffer from limited reconstruction and image quality, which hampers effective latent factor separation and limits their applicability to high-fidelity image generation and editing. We propose a novel conditioning framework for diffusion models that enables contrastive decomposition without compromising generation quality. We first train a prompt-free, image-conditioned diffusion model, and then learn to decompose the conditioning into a common and a salient factor, using weak supervision. We prove that the additive contrastive factorization, commonly assumed in prior work, is identifiable under mild conditions. This factorization enables targeted editing operations by swapping or interpolating only the salient factor while preserving the common factor.
PaperID: 4344, Poster
Abstract: Compositional explanations describe the alignment between neuron activations and concepts through logical formulas, obtained by solving a search problem over combinations of concepts. Due to the combinatorial nature of this space, prior work relies on incremental beam search, which restricts the exploration but limits the expressiveness of explanations and may discard intermediate candidates needed to identify higher alignment. In this paper, we propose a non-incremental beam search framework that expands the search space while keeping the computation tractable. Our approach introduces a non-incremental beam expansion strategy and a functional-equivalence-aware beam search that leverages equivalence classes of formulas to improve both search quality and efficiency. Experiments on natural language and vision tasks show that our method finds explanations expressing a higher alignment than the incremental formulation while maintaining comparable runtime, and leads to a significant increase in the number of optimal explanations when these can be computed. Overall, these results show that expanding search beyond incremental formulations leads to more expressive and accurate explanations, making our method a promising replacement for standard incremental beam search.
PaperID: 4345, Poster
Authors: Jer Pelhan, Alan Lukezic, Matej Kristan
Abstract: General multi-object tracking (GMOT) tracks all instances of a user-specified category from a single first-frame exemplar. Prior work relies on bounding boxes and surrogate training, and struggles with non-rigid objects, crowded scenes, and distractors. We introduce UGO, a unified GMOT tracker that pairs a pretrained exemplar-conditioned detection head with an instance-propagation head in a common architecture. A novel training-free, energy-minimization consolidation method converts overlapping proposals into exclusive pixel-wise masks and detections, resolving over-segmentation, duplicates, and conflicts. A hierarchical memory spanning global and instance levels improves recall and per-instance segmentation accuracy using a new memory management protocol. UGO sets a new state-of-the-art on GMOT benchmarks and video object counting, and is competitive with specialist MOT methods, establishing a strong paradigm for unified, open-category multi-object tracking.
PaperID: 4346, Poster
Authors: David Bergström, Mattias Tiger, Fredrik Heintz
Abstract: Trajectories arise in many settings, including urban mobility, transportation, and maritime traffic. Existing generative models either offer no control over output distributions or rely on conditioning information specific to each individual trajectory, such as its origin and destination, distance, or departure time. This sample-specific conditioning limits controllability and ties the model to the environment where those statistics were observed. We propose to separate _where_ movement occurs from _how_ it unfolds: regional movement is summarized by a spatial prior \Pi, a marginal distribution of occupancy aggregated over many trajectories, and the generative model is trained so that samples drawn conditionally on \Pi aggregate to match it. We call this property _marginal consistency_, which turns the spatial prior into a controllable input, enabling zero-shot generation in unseen regions, cities, and even unseen domains by supplying only the target region's prior. The Temporal Deaggregation Diffusion Model (TDDM) realizes this idea. We evaluate it across four datasets spanning three continents and three modalities: multi-modal human mobility (Geolife), urban taxi (Porto, Cabspotting), and maritime traffic (Brest). TDDM achieves improved fidelity and coverage over leading baselines and stable performance when transferred across regions, cities, and domains.
PaperID: 4347, Poster
Abstract: Generative models for function-valued data, such as time series and solutions to partial differential equations, must learn distributions over infinite-dimensional spaces rather than over finite-dimensional vectors. Functional Flow Matching (FFM) is a recent extension of Flow Matching to this setting, learning a velocity field whose flow transports a Gaussian prior to the data distribution. However, FFM inherits a structural limitation from standard Flow Matching: in each training batch, prior and data samples are paired independently, so the conditional bridge between them must simultaneously traverse the shared global structure of the dataset and instance-specific residuals. This limitation is more consequential in function space than in finite dimensions: directly formulating optimal transport on function spaces is technically delicate, and a flat Euclidean surrogate ignores the function-space geometry that distinguishes function-valued data. We propose kernel Functional Flow Matching (kFFM), which replaces the independent pairing by entropic optimal transport in a kernel-induced Hilbert space via the Hilbert Sinkhorn Divergence, leaving the FFM neural-operator architecture unchanged. Theoretically, we prove the HSD objective is well-posed on Banach ambient spaces with bounded kernels, derive compact-metric approximation bounds against quadratic-cost OT with an explicit kernel-cost mismatch term, and quantify the gap between the function-space objective and the truncated representation computed on a grid, at a rate governed by data regularity. Empirically, kFFM consistently improves MMD-RBF distributional matching over FFM, diffusion, adversarial, and finite-dimensional OT baselines on time-series and PDE benchmarks, with function-space-aware kernels emerging as the strongest choices on every PDE and path-valued benchmarks.
Authors:
Zheyuan Gu, Minghao Shao, Zhen Wang, Keyu Mao, Ailiang Lin, Shijie Zhang, hao jiang, mingyou liang, Mingkun Xu, Yusong WangAbstract: While spatiotemporal deepfake detectors have shown high AUC, our experiments reveal that they remain susceptible to evasion attacks. These models tend to overfit on fragile temporal spectrum cues, rather than learning robust semantic causality. To mitigate this vulnerability, we propose SpInShield, a temporal spectral-invariant defense framework explicitly designed to decouple semantic motion from manipulatable spectral artifacts. We propose a learnable spectral adversary that dynamically synthesizes severe spectral deformations, simulating extreme attack scenarios. By employing a shortcut suppression optimization strategy, SpInShield compels the encoder to extract reliable forensic cues while autonomously purging unstable spectral statistics from the latent space. Experiments show that SpInShield obtains competitive performance on widely used datasets and outperforms the strongest baseline by 21.30% AUC under simulated amplitude spectral attacks.
PaperID: 4349, Poster
Abstract: Modern power systems are undergoing a significant transition from deterministic, steady-state dispatch to operation under high-dimensional uncertainty and rapid fluctuations driven by renewables, electrification, and topology perturbations. This shift requires Optimal Power Flow (OPF) to be solved frequently under diverse operating conditions. Classical AC-OPF solvers provide high physical fidelity but are slow for high-frequency scheduling, while linear approximations such as DC-OPF sacrifice accuracy for computational speed. To address these challenges, we propose GridDiffuser, a graph diffusion solver for AC-OPF. Unlike deterministic learning methods that predict a single solution and struggle to represent the non-convex landscape, GridDiffuser models a multi-modal distribution over operating points and generates multiple candidate solutions. Our method further applies an efficient Levenberg--Marquardt correction to the clean-space estimate, iteratively steering samples toward the constraint-feasible set. Experiments on thousand-bus grids with instance-level N-1 contingencies show that GridDiffuser substantially reduces pre- and post-power flow (PF) violations compared with deterministic baselines while generating solutions much faster than AC-IPOPT.
Abstract: Backpropagation is the default learning rule for artificial neural networks and is often treated as the settled approach whenever differentiability is available. In this work, we revisit this convention from the perspective of sample efficiency. We introduce a unified vectorized feedback framework for loss-based and reward-based learning on computational graphs, in which synthetic gradients emerge as a natural alternative to backpropagation. We characterize the conditions under which synthetic gradients can achieve a lower gradient-estimation mean squared error than backpropagation. We construct examples illustrating that this sample efficiency advantage can be arbitrarily large. Experiments on contextual bandits and reinforcement learning tasks demonstrate the potential of our theoretical findings.
Abstract: Neural operators on irregular meshes face a fundamental tension. Spectral positional encodings, the natural choice for capturing geometry, require cubic-complexity eigendecomposition and inadvertently break gauge invariance through numerical solver artifacts; existing efficient approximations sacrifice gauge symmetry by design. Both failure modes break discretization invariance: models fail to transfer across mesh resolutions of the same domain, and similarly across different graphs of related structure in inductive settings. We propose GIST (Gauge-Invariant Spectral Transformer), a scalable neural operator that resolves this tension by restricting attention to pairwise inner products of efficient approximate spectral embeddings. We prove these inner products estimate an exactly gauge-invariant graph kernel at end-to-end \mathcalO(N) complexity, and establish a formal connection between gauge invariance and discretization-invariant learning with bounded mismatch error. To our knowledge, GIST is the first scalable graph neural operator with a provable discretization-mismatch bound. Empirically, GIST sets state-of-the-art on the AirfRANS, ShapeNet-Car, DrivAerNet, and DrivAerNet++ mesh benchmarks (up to 750K nodes), and additionally matches strong baselines on standard graph benchmarks (e.g., 99.50% micro-F1 on PPI).
PaperID: 4352, Poster
Abstract: Multimodal large language models (MLLMs) are transforming recommender systems into content-aware decision systems. In these systems, multimodal compression and cross-modal fusion are often viewed as robustness-enhancing mechanisms against content manipulation, as compression can filter noisy signals and cross-modal consensus can suppress inconsistent anomalies. We challenge this assumption by showing that the same process can expose a compression-to-consensus (C2C) bias in ranking decisions. Signals that remain salient after compression and receive consistent support across modalities can gain disproportionate influence in candidate comparison, offering an exploitable direction for content manipulation. In this paper, we introduce C3H, a novel inference-time attack that exploits this vulnerability. Rather than relying on noisy perturbations or training-time poisoning, C3H promotes a target item by tailoring its multimodal content to align with the decision criteria implicitly expressed by the recommender. Extensive experiments demonstrate that C3H substantially increases target-item exposure while maintaining overall recommendation utility.
PaperID: 4353, Poster
Abstract: E-values have attracted considerable interest in recent years as flexible tools for enabling anytime-valid and adaptive data analysis. Hypothesis testing is at the core of many of these applications, which can often involve private or sensitive data. In this work, we answer a simple but important question: given two distributions \mathbbP and \mathbbQ, what is the maximum achievable e-power when testing X~ \mathbbP^n against X~\mathbbQ^n with e-values that satisfy \varepsilon-differential privacy? We characterize the optimal rate for this problem and provide an algorithm which matches it asymptotically. In the sequential setting, when observations arrive one-by-one and the analyst chooses when to halt, we give matching upper and lower bounds on the stopping times of any differentially private e-process. Numerical experiments confirm the practicality of our algorithms, which require less data than the recently proposed DP-SPRT across a range of sequential testing problems and privacy levels.
PaperID: 4354, Poster
Authors: Deming Chu, Ruizhe Zhang
Abstract: We study the row-private approximate feasibility problem for symmetric cone programs (SCPs), a framework that encompasses linear, second-order cone, and semidefinite programming. We focus on high-sensitivity row privacy, where each record is a constraint and changing one record can cause the feasible region to change discontinuously. Since satisfying every private row is generally impossible in this regime, the goal is to remain exactly in the public cone while violating only a small number of private rows under a prescribed additive error \eta>0. We give an efficient (\varepsilon, \delta)-differentially private SCP solver with violation count \widetildeO\left(k^2/\varepsilon\right), and only logarithmic dependence on the scale-to-error ratio UR/\eta, where k is the ambient dimension, R is the norm of some feasible point, and U is the size of the problem. This exponentially improves the polynomial dependence on U, R, and 1/\eta in prior work (Song, Xue, and Zhang, NeurIPS 2025). Furthermore, we show that additive error does not remove the intrinsic dimension-dependence barrier. Previous row-private LP lower bounds (Kaplan, Mansour, Moran, Stemmer, and Tur, STOC 2025) established an \Omega(k/\varepsilon) violation count only for exact partial satisfaction; using the padding-and-permuting fingerprinting codes (Peter, Tsfadia, and Ullman, COLT 2024), we prove the same barrier for relaxed satisfaction.
Abstract: Test-time training (TTT) enhances model performance by explicitly updating designated parameters prior to each prediction to adapt to the test data. While TTT has demonstrated considerable empirical success, its theoretical underpinnings remain limited, particularly for nonlinear models. In this paper, we investigate the combination of TTT with in-context learning (ICL), where the model is given a few examples from the target distribution at inference time. We analyze this framework in the setting of single-index models, where the feature vector is drawn from a hidden low-dimensional subspace. For single-layer transformers trained with gradient-based algorithms and adopting TTT, we establish an upper bound on the prediction risk. Our theory reveals that TTT enables the single-layer transformers to adapt to both the feature vector and the link function, which vary across tasks. This creates a sharp contrast with ICL alone, which is theoretically difficult to adapt to shifts in the link function. Moreover, we provide the convergence rate with respect to the data length, showing the predictive error can be driven arbitrarily close to the noise level as the context size and the network width grow.
Abstract: Multi-task reinforcement learning (MTRL) aims to train a single agent to efficiently optimize performance across multiple tasks simultaneously. However, jointly optimizing all tasks often yields imbalanced learning: agents quickly solve easy tasks but learn slowly on harder ones. While prior work primarily attributes this imbalance to conflicting task gradients and proposes gradient manipulation or specialized architectures to address it, we instead focus on a distinct and underexplored challenge: \emphimbalanced data allocation. Standard MTRL allocates an equal number of environment interactions to each task, which over-allocates data to easy tasks that require relatively few interactions to solve and under-allocates data to hard tasks that require substantially more experience to solve. To address this challenge, we introduce Distributionally Robust Adaptive Task Sampling (DRATS), an algorithm that adaptively prioritizes sampling tasks furthest from being solved. We derive DRATS by formalizing MTRL as a feasibility problem from which we derive a minimax objective for minimizing the worst-case return gap, the difference between a desired target return and the agent's return on a task. In benchmarks like MetaWorld-MT10 and MT50, DRATS improves data efficiency and increases worst-task performance compared to existing task sampling algorithms.
Abstract: Personalized systems rely on user representations to connect behavioral history with downstream recommendation applications. Existing methods typically employ either supervised latent user embeddings, which are effective for retrieval but difficult to interpret, or textual user profiles, which are interpretable but challenging to optimize for downstream utility due to lack of direct supervision. To bridge this gap, we present BLUE, a reinforcement learning framework that unifies these two forms of user representation by aligning language-based user profiles with embedding-based recommendation objectives. Given a user’s interaction history, BLUE leverages a profiler Large Language Model (LLM) to generate textual profiles, while an embedding model provides reward signals. This encourages the resulting textual representations to move closer to positive items and farther from negative ones in the embedding space. We further introduce a text-space supervision signal based on next-item prediction, ensuring the learned profiles remain both semantically meaningful and highly effective for downstream retrieval. Experiments on Amazon Reviews 2023 and Google Local Reviews in zero-shot sequential recommendation settings demonstrate that BLUE consistently outperforms strong baselines under both frozen and trainable embedding conditions. Notably, BLUE achieves clear gains in cross-domain transfer, highlighting the strong generalization ability of the learned user profiles. Furthermore, these generated profiles provide superior personalized context for question answering compared to raw user histories or alternative profile optimization methods. Overall, these results show that BLUE provides an effective way to unify interpretable textual profiling with discriminative latent embeddings for personalization.
Abstract: Large language model (LLM) web agents are usually deployed as tool callers: each turn, the model reads a fresh page observation and emits one structured tool action. When every action is a low-level primitive, horizons grow quickly and so do policy-facing LLM completions, dominating latency and cost on benchmarks such as Mind2Web and WebArena. Recent systems therefore wrap repeated interaction fragments as web skills: callable tools built from successful trajectories or induced programs, so one call can replace several primitives. However, prior skill libraries are still triggered mainly by instruction similarity or coarse site metadata, which yields low skill reuse on held-out sites and leaves much of the potential step and token reduction on the table. We present SkillMigrator, an agent that learns reusable web skills and transfers them across sites by matching layout structure rather than specific element references. Each induced skill is stored as a transferable interaction pattern (TIP): the skill paired with a structural sketch of the snapshot at induction time. At test time, SkillMigrator retrieves TIPs by layout similarity and grounds their references on the live page. The rest of the stack is standard: accessibility-snapshot observations with stable references, and fixed tool calling over primitives plus skill invocations. Compared with the state-of-the-art approaches, SkillMigrator reduces the average LLM-action count on successful trajectories by 8–10% across both WebArena and Mind2Web at matched success rate.
Authors: Aidan Gleich, Scott C Schmidler
Abstract: We address the problem of accurate, training-free guidance for conditional generation in trained diffusion models. Existing methods typically rely on point-estimates to approximate the posterior score, often resulting in biased approximations that fail to capture multimodality inherent to the reverse process of diffusion models. We propose a sequential Monte Carlo (SMC) framework that integrates over the full denoising distribution to construct an unbiased Monte Carlo estimator of p_\theta(y|x_t), avoiding the bias of existing training-free guidance methods. To ensure computational tractability, we incorporate variance-reduction schemes based on MultiLevel Monte Carlo (MLMC). We evaluate on three image generation tasks. On CIFAR-10 and ImageNet classifier guidance, our approach improves accuracy by 18 and 38 points respectively over the strongest training-free baseline. On CelebA-HQ joint-attribute guidance, three existing methods (DPS, LGD, TFG) used within our framework all outperform the closest SMC-based comparator (TDS) by 25-37 points, and improve over the strongest baseline by 2-14 points.
Abstract: Reinforcement learning with verifiable rewards (RLVR) plays a pivotal role in improving the reasoning ability of large language models. However, widely used PPO surrogate objectives are fundamentally local, as they rely on a local approximation of the exact policy gradient objective. While this approximation improves stability by reducing the variance induced by importance sampling, it also introduces structural bias into the surrogate objective, which must be controlled through trust region mechanisms. In this work, we introduce the N-step forward trace, which augments the PPO surrogate objective using the cumulative likelihood ratio of the next N-1 tokens. Building on this idea, we propose N-Step Forward-Trace Policy Optimization (NFPO), a practical RLVR algorithm that integrates the N-step forward trace into the masked policy gradient framework. NFPO provides a continuous bridge between the PPO surrogate objective and the exact policy gradient objective, offering a principled mechanism for controlling the bias–variance trade-off. Our theoretical analysis shows that, with an appropriate choice of N, the proposed objective yields a tighter policy-improvement bound than the standard PPO surrogate. Experiments on comprehensive reasoning benchmarks demonstrate that NFPO consistently improves performance, supporting our theoretical findings.
Abstract: Merging separately trained LoRA adapters is a practical alternative to joint multi-task training, but it often hurts performance. Existing methods usually treat the LoRA update \Delta W = BA as a single object and do not distinguish the two LoRA matrices. We show that the main source of LoRA merge interference comes from the output-side matrix B. Across tasks, B repeatedly uses a small set of shared directions, while A remains much more task-specific. As a result, the merged adapter overemphasizes these shared directions, and task-specific information is lost. We propose Pico (Pre-merge interference calibration in output-space), a data-free method that calibrates B before merge by downscaling over-shared directions and then rescaling the merged update. Pico plugs directly into existing merging methods such as Task Arithmetic, TIES, and TSV-M. Across eight different benchmarks from math, coding, finance, and medical domains, Pico improves average accuracy by 3.4-8.3 points over the corresponding base method and achieves the best overall average performance. Pico also enables merged adapters to outperform the LoRA trained with all task data. These results show that LoRA merging works better when the two LoRA matrices are treated separately.
Abstract: Generating egocentric video from a single exocentric video is an emerging and important topic for AR/VR and physical AI. Compared with conventional novel view synthesis, exo-to-ego generation is a significantly harder task due to the extreme viewpoint changes, extensive occlusions and unobserved regions, where standard geometric conditioning becomes highly unreliable. We present Grounded-Exo2Ego, a principled framework that addresses these challenges at both the architectural and data levels. Architecturally, Grounded-Exo2Ego is a dual-branch video diffusion model that couples a geometric anchoring branch, which renders a 3D reconstruction into the ego view for overall scene layout, with a semantic grounding branch, which anchors per-object context onto the scene geometry so that heavily distorted regions and unobserved regions can be hallucinated in context. Additionally, to support robust training at scale, we substantially improve the labeling pipeline for real-world data and further develop a fully automated synthetic data engine that renders high-quality dynamic humans in procedurally generated environments. Evaluations on the challenging EgoExo4D dataset show that our method substantially outperforms recent state-of-the-art approaches. Extensive ablations further demonstrate that both structured semantic grounding and rigorous data processing are essential for robust exo-to-ego video generation in extreme scenarios.
PaperID: 4363, Poster
Abstract: Transformer-based autoregressive multilingual language models (LMs) are often assumed to store factual knowledge in an interlingua: a language-agnostic latent space. If this assumption holds, a parameter update applied in a source language during knowledge editing should inherently propagate to its semantic equivalents in target languages. However, empirical benchmarks reveal that cross-lingual propagation of edits consistently fails, a degradation that prior work treats largely as an algorithmic symptom rather than a structural bottleneck. In this work, we hypothesize that this failure stems from language anisotropy, and we perform a rigorous geometric dissection of the cross-lingual propagation gap. Using associative memory-based "locate-then-edit" methods, including ROME, MEMIT, and AlphaEdit, as analytical frameworks, we derive a closed-form decomposition of this gap into key-space propagation, parallel residual mismatch, and orthogonal residual mismatch. We empirically find that cross-lingual propagation is consistently limited by a large orthogonal residual component, indicating that source-language edits often fail to span the target-language update directions required for successful transfer. Furthermore, we show that post-hoc correction using only source-language residual statistics faces a large residual-prediction barrier. Finally, we demonstrate that structural realignment via a targeted contrastive alignment objective substantially improves key-space propagation and partially improves cross-lingual edit transfer, while leaving value-space mismatch as a remaining bottleneck.
PaperID: 4364, Poster
Abstract: Building a capable video editor remains significantly harder than a video generator: editing requires (source, instruction, edited) triplets that are prohibitively expensive to annotate and difficult to synthesize at scale, whereas image editing has already reached maturity with millions of such pairs readily available. In this work, we introduce Chain-of-Editing, a unified framework that transfers editing capability from images to videos through in-context visual demonstrations, alleviating the need for large-scale paired video editing data. Chain-of-Editing decomposes video editing into a structured chain of composable sub-tasks (Video to Image to Image to Video), routing editing intent through the image domain and enabling scalable joint training from heterogeneous image and video corpora under a unified diffusion objective. An image editing pair, synthesized by the model or supplied by the user, is prepended as an in-context visual demonstration that serves as a spatial appearance blueprint for every output frame. To ground appearance edits across the interleaved context, we introduce a novel position encoding that links image demonstrations and video frames in a shared spatial coordinate system, enabling pixel-faithful propagation of appearance changes to every output frame. Chain-of-Editing further provides principled test-time scaling: by executing the sub-task chain as progressive diffusion stages, editing quality can be improved by investing additional compute without retraining. Comprehensive experiments on OpenVE-Bench demonstrate the state-of-the-art performance across diverse editing categories, with ablations confirming the effectiveness of each component.
PaperID: 4365, Poster
Authors:
Shengjian Wu, Li Sun, Yu Shangguan, Qingli LiAbstract: DETR-style detectors use one-to-one bipartite matching during training to assign object queries to ground-truth objects, enabling end-to-end set prediction without non-maximum suppression (NMS). However, without an explicit de-duplication procedure, multiple queries can still produce highly similar hypotheses for the same object, making training unstable and predictions less decisive. Inspired by the sequential ordering of NMS, we propose DETRNN, a plug-and-play module that turns unordered object queries into a competition-aware sequence for recurrent refinement. DETRNN builds an explicit confidence-and-similarity based order from prior predictions, then refines queries with an RNN along this order to model competition inside the decoder. This ordered recurrent refinement reduces redundant predictions, stabilizes optimization, and improves final detection accuracy. Experiments on multiple DETR-style detectors show consistent gains with comparable efficiency. Code will be released on GitHub.
PaperID: 4366, Poster
Authors:
Shulan Wang, YingJie Zhu, Yuting Yan, Ke Cheng, Yan Chen, Sheng Zhang, Hefei Guo, Qidong Zhang, Zhengyong Zhang, Yibo Jin, Zhuzhong QianAbstract: Industrial personalized recommendation increasingly adopts LLMs for offline intent tagging (e.g., user interest from installed apps or browsing history), where each user corresponds to one prompt with thousands tokens. While shared prefix among prompts can improve the inference efficiency, the reusable part is actually dwarfed by user-specific records. We observe that even a slight discordance in preceding records prevents two prompts from sharing a longer prefix, revealing an opportunity to proactively re-form the common part from within, by adjusting the internal orders, with ensured accuracy for intent tagging. Realizing this at industrial scale is challenging: (1) the data preparation and the inference operate separately; (2) re-forming must complete within minutes for massive prompts under strict SLO; and (3) user data evolves continuously, requiring incremental updates. We present PrefixShakeup, a system that enables proactive prefix re-forming before actual inference, intrinsically enlarging the common part from the source. PrefixShakeup, bridges the data and inference by enhancing the datasets during its preparation, thereby increasing the reuse. It combines three techniques: (1) a prefix model that tolerates a bounded number of records, instead of requiring strict prefix matching; (2) an alleviator that reorders both records within each prompt and the prompt positions to maximize the reuse under limited HBM, upon the kernel-based acceleration for scoring the similarity; and (3) an adaptor that supports incremental changes via lightweight management and necessary cache over time, to avoid full re-computation. We implement PrefixShakeup, on Ascend NPUs and evaluate it in the mirror environment. Compared to state-of-the-art inference, PrefixShakeup, improves throughput by up to 2.1×, with controlled tagging accuracy.
PaperID: 4367, Poster
Abstract: Reinforcement Learning from Human Feedback (RLHF) is a key technique for aligning Large Language Models (LLMs) with human preferences. While Proximal Policy Optimization (PPO) is the standard algorithm, its reliance on a critic network incurs significant memory and computational costs. This has motivated the development of critic-free alternatives such as Group Relative Policy Optimization (GRPO) and Group Sequence Policy Optimization (GSPO). However, these methods suffer from a critical trade-off: they either employ theoretically unsound, high-variance estimators (GRPO) or introduce systematic bias to achieve stability, causing them to optimize a perturbed objective (GSPO). In this paper, we introduce SNIB (Self-Normalized Importance Sampling with a Baseline), a novel critic-free algorithm that addresses this dilemma by offering a method that is both stable and asymptotically correct. SNIB leverages principled self-normalized importance sampling to achieve the stability of modern methods without sacrificing asymptotic correctness. We provide a comprehensive theoretical analysis, proving that SNIB's gradient estimator is consistent with a finite-sample bias that decays as O(1/G) under self-normalization. Furthermore, we demonstrate its superior robustness to reward model uncertainty and show that it preserves the principled trade-off between reward maximization and KL regularization, a property that is distorted by biased estimators. Our work establishes a theoretically-grounded foundation for building more stable and reliable critic-free RLHF algorithms.
PaperID: 4368, Poster
Abstract: Vision-Language-Action (VLA) models offer a promising paradigm for autonomous driving. However, the massive parameter scale of VLAs strains onboard computational resources during real-time inference. This directly increases energy consumption and limits vehicle battery range, making it critical to minimize computational overhead without compromising driving capability and explainability. Current research on efficient VLA-based driving focuses on sequential adaptation (e.g., Chain-of-Thought and token adjustment), leaving structural adaptation underexplored despite parameter redundancy across dynamic driving scenes. To address this gap, we propose Scene-Adaptive VLA, an efficient framework leveraging dynamic layer routing to adjust active parameters based on real-time scenes. Inspired by the temporal continuity of scenes and model predictive uncertainty, our approach characterizes model-perceived scene complexity via frame-level temporal scene variations and decision deviations. To uncover latent scene-to-parameter correspondences, we encode this characterization into scene-aware tokens, alongside two specialized queries. A budget-conditioned layer routing mechanism evaluates these refined queries to allocate a budget that determines the active parameter ratio and subsequently selects appropriate layer combinations. The entire framework is optimized end-to-end. Extensive evaluations on the Bench2Drive benchmark show that when adapted to a state-of-the-art VLA model, our approach significantly reduces inference-time computational overhead by 53.5% with negligible degradation in driving and language capabilities. Code will be available.
Abstract: We revisit a universally accepted but under-examined design choice in every modern LLM: a token index is looked up once at the input embedding layer and then permanently discarded. This single-injection assumption induces two structural failures: (i) the Rare Token Problem, where a Zipf-type distribution of vocabulary causes rare-token embeddings are chronically under-trained due to receiving a fraction of the cumulative gradient signal compared to common tokens; and (ii) the Contextual Collapse Problem, where limited parameters models map distributionally similar tokens to indistinguishable hidden states. As an attempt to address both, we propose TIDE, which augments the standard transformer with EmbeddingMemory: an ensemble of K independent MemoryBlocks that map token indices to context-free semantic vectors, computed once and injected into every layer through a depth-conditioned softmax router with a learnable null bank. We theoretically and empirically establish the benefits of TIDE in addressing the issues associated with single-token identity injection as well as improve performance across multiple language modeling and downstream tasks.
PaperID: 4370, Poster
Abstract: Classifier-Free Guidance (CFG) is a widely used inference-time guidance rule for conditional image generation with diffusion models, which exposes a guidance scale that adjusts conditioning strength after training. This flexibility requires conditional and unconditional model evaluations at every sampling step, so an N-step trajectory takes 2N forward passes. Single-pass alternatives reduce such cost, yet most collapse CFG's two-component structure, forfeiting the separable base flow and guidance direction that support post-hoc scale control and per-step diagnostics. This work introduces CtrlFlow, a single-pass guidance method that preserves CFG's separable base-flow and guidance-direction interface. To preserve both components at one-pass cost, CtrlFlow attaches two lightweight LoRA adapters to a frozen backbone, each adding 2.5% parameters, and predicts the base flow and guidance direction from shared features. Because the guidance scale multiplies the guidance direction only after the forward pass, a single trained checkpoint remains adjustable at inference without retraining. Preserving two outputs introduces a training mismatch: inference amplifies guidance-direction errors by the chosen scale, while naive component-wise losses weight both outputs equally. To align training with this composed inference-time error, CtrlFlow uses a unified multi-scale distillation loss. Residual guidance errors can still persist at extreme scales, where even small guidance-direction errors are amplified by the guidance scale. To reduce this inference-time error, CtrlFlow applies scale-adaptive modulation to damp velocity-field errors and time-grid warping to concentrate solver evaluations where the trajectory changes most rapidly. On ImageNet 256×256 with the DiT-XL/2 backbone, CtrlFlow matches CFG within a pre-specified tolerance in 24 of 25 (N, s) configurations and achieves a measured 1.93× wall-clock speedup at N=25.
Abstract: Alignment methods for discrete diffusion models have primarily focused on steering the denoising process, either by influencing token logits or by selecting favorable sequences at intermediate steps. However, these approaches largely treat inference as a unidirectional process, lacking effective mechanisms for revisiting undesirable token selections. We introduce Spectral Feedback, an algorithm that uses a feedback loop that selects sets of positions to edit, allowing the model to iteratively correct its own generations. This approach leverages the masking structure of discrete diffusion models where tokens can be re-masked and re-sampled, analogous to image editing methods that reintroduce noise and re-run a reverse diffusion process. While prior methods focus on what token values to assign, we instead treat which tokens to revisit as the central alignment problem. Selecting edit positions is challenging because edit effects are interdependent: the impact of modifying one token depends on which others are edited simultaneously. Motivated by prior work on sparse interactions in biological systems, we show empirically that, for protein inverse folding, the expected reward over edit-position sets admits a sparse Fourier representation. This structure enables efficient learning and optimization of a value function over edit-position sets, which Spectral Feedback uses to select promising edits. The algorithm is model-agnostic and can be applied to pretrained, test-time aligned, or fine-tuned protein diffusion models, improving alignment performance without modifying the underlying generative process. Applied to inverse folding with a protein stability oracle, it achieves a 30.6% increase in stable proteins for a pretrained model, 24.8% for Best-of-10, and 8.5% for the SOTA RL-tuned diffusion models.
PaperID: 4372, Poster
Authors:
Jiaming Guo, Hongcheng Luo, Cheng Chi_, Lijun Zhou, Bing Wang, Guang Chen, Hangjun Ye, Jiaqi Yang, Xiaowei Zhou, Haiyang Sun, Sida PengAbstract: Reconstructing multi-camera videos is a fundamental task in autonomous driving. However, recent learned methods either rely on costly attention-based networks for global geometry modeling or require complex post-hoc alignment for chunked predictions, limiting their scalability to long multi-view rig sequences. We propose RigRecon, a generalized rig-aware reconstruction framework that enables efficient cross-view interaction in a low-resolution space while asynchronously decoding high-resolution depth maps in a frame-wise manner. Input images are first down-sampled and processed by a transformer backbone to capture coarse global structure, which is then provided to a decoder to reconstruct frame-wise high-resolution depth maps. By fully leveraging calibration and low-resolution feature interaction, RigRecon achieves consistent-scale geometry and accurate pose estimation with significantly reduced memory cost. It achieves state-of-the-art performance on 3D reconstruction, depth estimation, and camera pose estimation, and can efficiently process videos with over 1000 images in a single feedforward pass.
PaperID: 4373, Poster
Authors:
Yue Chen, Jinze Li, Dajiang LiuAbstract: In modern LLM applications, the demand for extended context lengths continues to surge, making KV cache a major performance bottleneck in long-context inference. Channel-wise pruning has emerged as a prevalent KV cache compression paradigm; however, existing solutions overlook inherent similarities among KV vectors, resulting in limited pruning ratios and noticeable performance degradation. To tackle this issue, we present ResKV, a novel residual-driven channel-wise pruning framework that fully exploits KV vector similarity for effective compression. ResKV first clusters highly similar KV vectors via cosine similarity metrics, then calculates residual vectors by subtracting each cluster’s centroid from the original features. Subsequent channel-wise unstructured pruning and per-token quantization are exclusively applied to these residual components. During inference, preserved residuals are combined with their corresponding centroids to accurately recover dense KV representations. Benefiting from the low magnitude and sparse, near-zero properties of centroid residuals, both pruning and quantization incur only negligible approximation errors. Experimental results show that ResKV reduces peak memory consumption by 45.6% and improves inference throughput by 4.27× against dense FlashAttention inference, while sustaining competitive model performance. Under identical sparsity constraints, ResKV outperforms state-of-the-art channel-wise KV pruning methods by achieving up to 2.26 points higher accuracy and 2.63× faster end-to-end throughput.
PaperID: 4374, Poster
Abstract: Although audio-driven human animation has achieved impressive realism, it still lacks effective Multi-Scene Temporal Control: orchestrating background transitions at precise timestamps while preserving a coherent foreground subject. We show that this challenge stems from the diffusion transformer's self-attention mechanism, and addressing it requires two capabilities: (1) Semantic-Aware Temporal Decoupling, which isolates background attention across scenes while retaining cross-scene foreground attention for subject consistency; and (2) Foreground Motion Localization, which accurately tracks the foreground subject across latent frames, especially under large motion. To address these challenges, we introduce SceneShifter, a training-free framework for precise multi-scene control in audio-driven human animation. SceneShifter guides spatio-temporal self-attention by suppressing cross-scene attention among background tokens while preserving cross-scene attention among foreground tokens. To support this guidance under large motion, SceneShifter extracts dynamic foreground masks from the self-attention layers and heads that best capture subject trajectories. We further introduce SceneShifterBench, a benchmark designed to evaluate scene-timing accuracy, foreground preservation, and visual quality under large motion, multiple subjects, and occlusions. Experiments show that SceneShifter achieves frame-level scene timing, strong subject preservation, and high visual quality, outperforming existing baselines. Code and examples can be found at https://anonymous.4open.science/r/SceneShifter
Abstract: We study image inpainting with generative diffusion models. Existing methods typically either train dedicated task-specific models, or adapt a pretrained diffusion model separately for each masked image at deployment. We introduce a middle-ground model, termed Amortized Inpainting with Diffusion (AID), which keeps a pretrained diffusion backbone fixed, trains a small reusable guidance module offline, and then reuses it across masked images without per-instance optimization. We formulate it as a deterministic guidance problem with a supervised terminal objective. To make this problem learnable in high dimensions, we derive an auxiliary Gaussian formulation and prove that solving this randomized problem recovers the optimal deterministic guidance field. This bridge yields a principled continuous-time actor--critic algorithm for learning the guidance module in a fully data-driven manner. Empirically, on AFHQv2 and FFHQ under the pixel EDM pipeline and on ImageNet under the latent EDM2 pipeline, AID consistently improves the quality--speed trade-off over strong fixed-backbone and amortized inpainting baselines across multiple mask types, while adding less than one percent trainable overhead.
Abstract: The expressive limitations of message-passing Graph Neural Networks (GNNs) have motivated a wide range of more powerful graph learning architectures. We advocate Deep Homomorphism Networks (DHNs) as a model particularly well-suited for learning over relational databases, due to their close connection to important fragments of SQL such as conjunctive queries. We study the precise expressive power of DHNs by relating them to various natural fragments and extensions of first-order logic (FO). For DHNs with max, sum, and mean aggregations, we establish connections to the unary negation fragment (UNFO) and to the extensions of UNFO with counting quantifiers and with ratio quantifiers. We further relate sum-aggregation DHNs to the unary quantifier alternation fragment of FO and to an extension of FO with expressive counting. Through the classical correspondence between FO and SQL, these results also illuminate the relation between DHNs and SQL. They also enable us to study the decidability of two fundamental static analysis problems for DHNs, the emptiness problem and the subsumption problem. Finally, we confirm through experiments that the established differences in expressive power are reflected in the performance on suitable prediction tasks.
Abstract: Video retrieval at scale is central to data curation and safety validation in autonomous driving, where users want to find not only scenes but also dynamic events such as cut-ins and hard braking. Existing vision-language and keyword-based retrieval methods often miss these events because the relevant motion may not be explicitly described in text or captured by lexical overlap. Rule-based retrieval can encode such events more directly, but it is brittle: generated or hand-written rules often fail when their assumptions do not match real driving data. We propose STRIVE-D, a data-calibrated retrieval framework for driving videos. STRIVE-D uses weakly labeled in-domain videos to estimate when a query rule is reliable, adapt rules that mismatch observed data, and fuse calibrated rule scores with vision-language and keyword-based retrieval signals. Across three driving benchmarks, including newly released human annotations for event retrieval on DrivingDojo, STRIVE-D achieves up to a 30.4% relative improvement over state-of-the-art retrieval models.
PaperID: 4378, Poster
Abstract: Selective unlearning in large language models is increasingly necessary for safety and compliance, yet harmful or memorized content is often latent, entangled, and hard to specify as a clean ``forget set.'' We aim to close this gap by identifying what to forget from the model’s own failure modes rather than relying solely on explicit supervision. We propose Hallucination-Guided Unlearning (HGU), which treats hallucinations as diagnostic signals: it localizes unsupported spans via evidence and internal consensus, traces a sparse set of high-responsibility MLP neurons, and edits them with compact gates under a utility-preserving KL constraint. Across TOFU, HGU achieves the best MU--FE trade-off with Avg of 0.7221 (forget05) and 0.7699 (forget10), and it also attains the top SepS H-Avg (0.4382/0.4715) under mixed forget/retain prompts; on WMDP auxiliary, it yields the strongest forgetting (e.g., 14.2 on Economics) while improving Retain (52.7) and MMLU (57.2). These results suggest that reframing hallucinations as actionable traces enables targeted, post-hoc unlearning that better balances forgetting efficacy and general utility without full retraining.
Abstract: Retrieval-augmented generation evaluation checks whether model claims are factually grounded in retrieved documents. It does not check whether retrieved evidence is attributed to the correct entity. A clinical RAG response can pass every automated check (zero hallucinations, near-perfect faithfulness, real citations) while presenting drug Y's clinical evidence as evidence about queried drug X. We term this deceptive grounding (DG): a failure invisible to faithfulness, hallucination, and citation checks because every claim is sourced from a real document, about the wrong entity. Using a controlled factorial benchmark across 13 models, we find DG rates spanning 8--87% at peak adversarial conditions. Medical and biomedical fine-tuned models reach up to 86.7%; domain specialization amplifies the failure rather than mitigating it. A controlled ablation identifies the mechanism: removing entity-specific clinical evidence from retrieved documents eliminates entity-attribution failure entirely, shifting all failures to confabulation. The two failure modes respond to the same trigger, taking different paths. Production measurement across 740 drug--disease pairs finds 7.8% overall DG in a deployed RAG system, rising to 13.6% for recently approved drugs. Entity-attribution verification (checking that cited evidence applies to the queried entity) detects DG at 92.3% precision and 95.8% DG recall; no existing framework implements it.
PaperID: 4380, Poster
Abstract: Diffusion planners enable long-horizon planning through a generative process over full trajectories, mitigating compounding errors of autoregressive methods while handling multimodal futures. Reward conditioning via classifier-free guidance (CFG) yields high-performing plans but is brittle with respect to per-task hyperparameter choices, limiting its scalability. Our analysis reveals that guidance performance hinges on careful adaptation to the data manifold and reward distribution, contributing to CFG's hyperparameter fragility. We propose temperature-guided diffusion planning (TGDP), which adapts CFG to self-calibrate to these characteristics through temperature-conditional sample reweighting during training and adaptive guidance at inference. TGDP mitigates the hyperparameter fragility of CFG and matches state-of-the-art performance across standard benchmarks without per-task tuning, thus enabling more robust and practical diffusion-based planning.
PaperID: 4381, Poster
Abstract: Efficient video representations should do more than predict pixels: for a long video to be trainable, decodable, and budgeted, appearance and motion must be assigned to explicit units. Group-of-Pictures (GoP) primitive representations provide such units by fitting short blocks with canonical local elements, but their deformation is often still predicted by a neural head. Conversely, implicit neural representations achieve compact global fits without exposing local units that can be moved, ranked, or pruned. We observe that a short GoP contains many primitives but only a few independent motion patterns, suggesting that the same local ownership system used for appearance should also define the motion basis. We introduce Dynamic Soft Anisotropic Diagrams (DSAD), an explicit representation built from three layers over a shared soft-distance geometry. First, a Soft Anisotropic Diagram (SAD), defined as a top-K-masked softmax over anisotropic-Apollonius sites, provides a canonical coordinate family in which dense sites encode local RGB ownership. Second, a coarser anchor set reuses the SAD geometry so that anchor weights form a partition-of-unity motion basis without an additional neural module. Third, a reduced-order model (ROM) factorizes deformation into this spatial motion basis and knot-based temporal coefficients. This construction unifies appearance and motion within a single diagram while decoupling spatial ownership from temporal dynamics. On Bunny, UVG, and DAVIS, DSAD attains the highest PSNR among primitive representations and competitive performance among all compared methods, while using a leaner per-GoP deformation module and fewer fine-stage training iterations than state-of-the-art baselines. Ablations show that the gain is a matched-budget trade-off: a compact anchor-ROM motion head reallocates parameters to anchored appearance. RAFT spectra further suggest that the largest gains occur when GoP motion is close to low-rank. Code and models will be released.
Abstract: Scaling network depth has been a central driver behind the success of modern foundation models, yet recent investigations suggest that deep layers are often underutilized. This paper revisits the default mechanism for deepening neural networks, namely residual connections, from an optimization perspective. On a motivating example, rigorous analysis proves that the layout of residual connections can fundamentally shape convergence behavior, and even induces an gap in convergence rates. Prompted by this insight, we introduce adaptive neural connection reassignment (ANCRe), a principled and lightweight framework that parameterizes and learns residual connectivities from the data. ANCRe adaptively reassigns residual connections with negligible computational and memory overhead (<1%), while enabling more effective utilization of network depth. Extensive numerical tests across pre-training of large language models, diffusion models, and deep ResNets demonstrate consistently accelerated convergence, boosted performance, and enhanced depth efficiency over conventional residual connections.
PaperID: 4383, Poster
Authors:
zihan jia, Zhilin Dai, Zhengming Zhang, Min Yang, Zhenpeng Huang, Jiaqi Li, Caixia Sun, Xi Chen, Liang Li, Junlan Feng, Gangshan Wu, Limin WangAbstract: Existing soccer commentary models are often designed for pre-segmented clips or localized events. When deployed on untrimmed full-match videos with sliding windows, they can produce delayed, repeated, or poorly synchronized commentary. Recent streaming video-language models make low-latency full-match narration feasible, but they often favor general real-time descriptions rather than professional, event-grounded soccer commentary. They may describe nearby actions while failing to mention key events such as goals, cards, substitutions, and offsides correctly and on time. In this paper, we introduce , a data, model, and evaluation framework that bridges low-latency streaming narration with event-grounded professional soccer commentary. First, we construct , a large-scale time-anchored dataset with word-level ASR alignment, quality filtering, and entity calibration. Second, we perform , which aligns the model with complete event-level commentary semantics by comparing causal multi-step rollouts from the same streaming prefix while keeping second-level inference unchanged. We further use event-mismatched human commentaries as counterfactual negatives to emphasize event semantics over commentary style. Third, we introduce , an event-centric benchmark based on temporally constrained event entailment. Full-match experiments show that SoccerNarrate improves event coverage and event-aligned precision over strong offline and streaming baselines. Data and models will be released.
PaperID: 4384, Poster
Authors: Polina Dolgova, Sebastian Stich
Abstract: Machine unlearning aims to remove the influence of selected training examples without full retraining. Standard evaluations often summarize unlearning quality with aggregate metrics, such as accuracy- and forgetting-based scores, which can hide localized failures. We study this failure mode at the example level by comparing the predictions of an unlearned model to those of the model retrained after deletion. We show that this pointwise discrepancy can be highly non-uniform: for gradient-ascent and random-labeling methods, with and without retain-set fine-tuning, it grows with geometric proximity to the forget set. We call this phenomenon . Our analysis identifies a mechanism behind the effect: surrogate targets used during unlearning can be inconsistent with the local prediction structure induced by retraining, and this inconsistency propagates through shared representations to nearby examples. Motivated by this mechanism, we propose , a simple mitigation strategy that replaces random targets with soft labels from a small teacher trained only on retained neighbors of the forget set. On CIFAR-100 partial-class deletion, this local teacher brings the unlearned model substantially closer to retraining, especially near the forget set, while maintaining competitive aggregate unlearning metrics.
PaperID: 4385, Poster
Abstract: Large language models (LLMs) are increasingly adapted inside privacy- and regulation-constrained institutions where local post-training data cannot be centrally pooled. Federated learning therefore becomes a natural mechanism for collaborative LLM adaptation, but it also raises a safety question that is not captured by either classical federated learning or standard centralized alignment: clients may share a model while still requiring different safe behaviors. We formalize this setting as \emphpersonalized federated safety, where safety correctness is conditioned on the target client's local policy rather than treated as a single global refusal rule. To make this problem transferable rather than benchmark-specific, we introduce a compact, interpretable client policy space and a policy-conditioned benchmark construction framework, then instantiate it with five representative client archetypes spanning regulated healthcare, legal compliance, financial safety, youth safety, and enterprise general safety. The resulting benchmark contains 1,500 client-conditioned training examples and 750 held-out evaluation examples. We further propose \emphSafety-Aware Weighted LoRA (SAW-LoRA), a lightweight add-on for federated LoRA fine-tuning that lets each client selectively absorb incoming global task updates according to estimated task benefit and local-safety risk. Across two task datasets and two backbones, this add-on substantially lowers client-specific ASR for the strongest local safety mechanisms, reaching ASR near 1--2% while preserving high benign acceptance. These findings position personalized federated safety as a concrete research problem with a practical update mechanism for federated LLM adaptation.
PaperID: 4386, Poster
Authors: Victor Omolaoye, Gerard de Melo
Abstract: Unstructured pruning methods such as Wanda preserve model quality but produce irregular sparsity with no practical compression, while structured methods enable efficient inference at the cost of degraded performance. We ask whether post-hoc linear transformations can convert unstructured sparsity into structured patterns, combining both benefits. Across a broad class of fixed linear transformations—including rotations, decompositions, and permutations—we find no increase in structured sparsity (Δ = 0.0%) in any OPT feed-forward layer. Transformations that yield structure do so identically on dense matrices, independent of pruning. We attribute this to three factors: rank preservation, weak inter-row mask correlation, and spectral dispersion. Together, these prevent post-hoc transformations from inducing structure, suggesting that compressible sparsity must be introduced during pruning rather than recovered afterward.
PaperID: 4387, Poster
Abstract: We study noisy inductive matrix completion (IMC) in the presence of heavy-tailed and possibly asymmetric noise, where the goal is to recover a rank-r matrix of ambient dimension d given n features as side prior information. So far, there has been a lack of theoretical understanding of the statistical estimation error in noisy IMC, let alone theory under heavy-tailed noise. We first develop an efficient two-stage nonconvex algorithm, called RGDIMC, via robust gradient descent with spectral initialization and feature-aware matrix factorization, based on an adaptive Huber loss to accommodate heavy-tailed noise. We prove that RGDIMC converges locally at a linear rate with sample complexity depending only linearly on n (up to \log n) and logarithmically on d, and achieves the minimax-optimal statistical error rate O_p(\sigma\sqrtdr/p) under merely a bounded second-moment condition on the noise, where r is the rank of the unknown core matrix. The theoretical results are strongly supported by extensive experiments.
PaperID: 4388, Poster
Authors: Vincenzo Marcianò, XIAOMING ZHANG, Gianluca Guglielmo, Sebastien Ourselin, Michela Antonelli, Maria A Zuluaga
Abstract: Current evaluation protocols for medical image segmentation remain largely anchored in sample-wise comparisons, e.g., Dice and Hausdorff distance, which require each prediction to be paired with a ground-truth mask. Yet two limitations arise in practice: they are inapplicable when predictions are made on unlabeled data; and even when annotations are available, they cannot assess whether a collection of predictions matches the distributional structure of the ground-truth population. We introduce MID (Mask–Image Distributional Divergence), a distribution-level evaluation metric for medical image segmentation. MID builds upon the observation that segmentation quality is a property of the joint image-mask distribution: a mask may be anatomically plausible in isolation yet incorrect for the image it accompanies. MID compares the joint distribution of images and masks between a reference set and a set of predictions, without requiring sample-wise correspondence. This captures both mask realism and image-mask consistency. We evaluate MID across 13 publicly available datasets spanning CT, MRI, and MRA. Under anatomically realistic corruptions across 56 organ groups, MID tracks mean Dice and Hausdorff distance without per-sample ground-truth pairing. MID also detects distributional shifts across modalities and anatomies, identifies image-mask misalignment through controlled shuffling experiments, and produces reliable model rankings across real segmentation models.
PaperID: 4389, Poster
Abstract: In this paper, we provide improved bounds on the complexity of solving \ell_p-Lipchitz convex optimization problem in parallel. We lower bound the depth of highly-parallel algorithms, i.e., the number of rounds required by a parallel algorithm making O(\mathrmpoly(d)) queries per round to compute an \epsilon-approximate minimizer of a \ell_p-Lipschitz, d-dimensional convex function over the unit \ell_p ball. We provide a framework for proving lower bounds for parallel \ell_p-Lipschitz convex optimization that generalizes prior lower bound proofs and for all p \in (1,\infty) (other than 2) improves the depth for which it can be shown that no super-polylogarithmic improvement over sequential algorithms is possible. We also prove a variety of new lower bounds for p>2. Notably, for p=\infty, we establish the near optimal parallel complexity of the class of ball-acceleration methods up to depth \tildeO(\sqrtd), which is the first near-optimality result in the parallel convex optimization setting for an algorithm other than mirror descent. Additionally, by providing new algorithms we characterize the optimal depth needed to solve to accuracy \epsilon = d^-1/4 for all p \geq 2 (up to polylogarithmic factors). We also provide extensions of our lower bounds to parallel quantum \ell_p-Lipschitz convex optimization.
Abstract: Hallucination in large language models (LLMs) during long-form generation remains difficult to address under existing reinforcement learning from human feedback (RLHF) frameworks, as their preference rewards often overlook the model's own knowledge boundaries. In this paper, we propose the Knowledge-Level Consistency Reinforcement Learning Framework (KLCF), which re-examines this problem from a distribution alignment perspective. KLCF formalizes long-form factuality as a bidirectional distribution matching objective between the policy model's expressed knowledge distribution and the base model's parametric knowledge distribution: under the constraint that generation must not exceed the support set of the base knowledge, the objective maximizes coverage of high-probability facts, thereby jointly optimizing precision and recall. To achieve this, we design a Dual-Fact Alignment mechanism that approximates the recall term using a factual checklist constructed by sampling from the base model, and constrains hallucinations with a lightweight truthfulness reward model. Both components are jointly optimized and require no external retrieval throughout training. Experimental results demonstrate that KLCF consistently improves factuality metrics across multiple long-form benchmarks and model scales, effectively alleviating hallucination and over-conservatism while maintaining efficiency and scalability.
Abstract: Existing 3D foundation models typically align point clouds to frozen vision-language spaces like CLIP, which achieve strong cross-modal retrieval by compressing 3D shape into a global vector. However, this global-only alignment cannot establish fine-grained pixel-to-point correspondence. To solve this, we present Tango3D, a foundation model that unifies dense correspondence and global retrieval. We use a geometry-aware 2D visual backbone and a pretrained 3D VAE to encode images into 2D patches and point clouds into 3D tokens. These are mapped into a single shared space to achieve both local pixel-to-point alignment and global semantic alignment. To stabilize the joint learning of dense and global objectives, we introduce a three-stage progressive training strategy. Experiments show our model is the first to achieve object-level pixel-to-point alignment while maintaining competitive global retrieval. By establishing a fine-grained alignment feature space, Tango3D injects rich semantics into purely geometric 3D tokens, paving the way for a wide range of dense 3D downstream tasks.
PaperID: 4392, Poster
Abstract: High-resolution global ocean states expose a bottleneck in data-driven forecasting: effective long-horizon rollouts require a compact, decodable latent state that reduces I/O and stabilizes error growth. While learned compression has matured for natural RGB images, its assumptions do not transfer to ocean modeling: ocean states are multi-variable and multi-depth, their grids are orders of magnitude larger, and fixed land--sea boundaries create sharp discontinuities that make aggressive downsampling error-prone. We present DC-Ocean, the first deep latent compression architecture for multi-variable, multi-depth global 1/12^\circ ocean states, built on an autoencoder that couples convolutional feature extraction with transformer-based global context modeling to retain both coastal sharpness and basin-scale structure. To explicitly address coastline-induced artifacts, DC-Ocean integrates boundary-aware gated convolutions together with a spherical geodesic front propagation interpolation scheme for stable behavior near land masks. Across key ocean variables, DC-Ocean achieves higher reconstruction fidelity than strong autoencoder baselines adapted from the natural-image domain at matched downsampling factors. Beyond reconstruction, we perform autoregressive forecasting directly in latent space and decode predictions back to the physical domain. DC-Ocean matches short-term accuracy of full-resolution forecasting while achieving lower RMSE and slower error growth at medium and long lead times. Together, these results demonstrate that DC-Ocean provides a practical framework for stable and extended high-resolution sub-daily ocean forecasting.
PaperID: 4393, Poster
Abstract: Point cloud registration remains challenging in scenarios with low overlap, repetitive structures, and local geometric ambiguities, where candidate correspondences established based on local geometric similarity often lack global structural consistency and are prone to unreliable matches. To address this issue, this paper proposes RSPReg, a reliability-aware structural prototype learning framework for robust point cloud registration. RSPReg learns structural prototypes from rotation-invariant coarse-level superpoint features to characterize the structural attributes of local regions within the entire point cloud. These prototypes are further introduced as structural priors into cross-cloud correspondence reasoning, thereby improving the structural consistency of candidate correspondences. Furthermore, RSPReg estimates the uncertainty of structural prototypes and adaptively modulates the role of structural priors in the matching and pose estimation processes, reducing the negative impact of ambiguous and non-overlapping regions on registration results. Experiments on 3DMatch, 3DLoMatch, KITTI, and ETH demonstrate that RSPReg achieves the highest registration recall and exhibits clear advantages in low-overlap and cross-domain scenarios.
Abstract: Fine-tuning aligned language models on benign tasks (e.g. math tutoring) systematically breaks safety guardrails, even when training data contains no harmful content. While mechanistic approaches have shed light on where alignment resides in model weights, they lack the formal language needed to derive guarantees about when and why fine-tuning degrades it — leaving the field without principled tools for predicting or preventing alignment collapse. We develop such a framework through geometric analysis of parameter-space trajectories and apply it to understand the fragility of alignment in fine-tuning. While first-order analysis suggests orthogonal updates are safe, we prove this is illusory: the curvature of the fine-tuning loss induces second-order acceleration that systematically bends trajectories into alignment-sensitive regions. We formalize the central construct of our framework as the Alignment Instability Condition (AIC), three geometric properties that, when present, are sufficient to guarantee degradation. Our main result proves quartic onset of alignment degradation in training steps, determined by how sharply alignment depends on specific parameters and how strongly tasks couple to these parameters. These findings yield formal sufficient conditions under which gradient descent makes alignment inherently fragile, demonstrating the framework's capacity for theoretical guarantees. We further empirically validate the framework's foundations, showing that the Fisher Information Matrix governs the degree of safety degradation across diverse fine-tuning settings.
Authors:
Yanjie Pan, Qingdong He, Zhengkai Jiang, Pengcheng Xu, Chaoyi Wang, Jinlong Peng, Haoxuan Wang, Yun Cao, Zhenye Gan, Mingmin Chi, Bo Peng, Yabiao WangAbstract: Recent advances in diffusion-based text-to-image generation have demonstrated promising results through visual condition control. However, existing ControlNet-like methods struggle with compositional visual conditioning - simultaneously preserving semantic fidelity across multiple het erogeneous control signals while maintaining high visual quality, where they employ separate control branches that often introduce redundant guidance during the denoising process, leading to structural distortions and artifacts in generated images. To address this issue, we present PixelPonder, a novel unified control framework, which allows for effective control of multiple visual conditions under a single control structure. Specifically, we design a patch-level adaptive condition selection mechanism that dynamically prioritizes spatially relevant control signals at the subregion level, enabling precise local guidance without global interference. Additionally, a time-aware control injection 017 scheme is deployed to modulate condition influence according to denoising timesteps, progressively transitioning from structural preservation to texture refinement and utilizing the control information from different categories to promote finer image generation. Extensive experiments demonstrate that PixelPonder surpasses previous methods across different benchmark datasets, showing superior improvement in spatial alignment accuracy while maintaining high textual semantic consistency.
PaperID: 4396, Poster
Abstract: Scientific foundation models trained on multimodal data are increasingly deployed across the sciences, but whether their internal computations correspond to known scientific operations or proceed through different mechanisms remains unclear. Understanding these mechanisms is essential for assessing reliability, identifying failure modes, and judging when predictions can be trusted in scientific workflows. We present the first mechanistic interpretability study of an astronomy foundation model, focused on AION-1's multimodal advantage in redshift prediction from broadband photometry and imaging over photometry alone. We find that this advantage concentrates on extended and dust-attenuated galaxies, with the gain localized by layerwise probing to AION-1's first encoder block. We examine two attention heads in this block and find that representation and function dissociate in opposite directions across them. The first head linearly encodes classical photometric features and organizes its image output along a radial gradient, yet ablation and activation patching show this structure has no measurable causal effect on the prediction. The second head shows no distinctive probe signature on the photometric-redshift estimation operations we tested, yet ablating it on dust-attenuated galaxies eliminates the multimodal advantage and amplifies the attenuation-bias signature beyond the photometry-only baseline, indicating that this head implements a correction for attenuation bias. Whether this reflects a methodological limitation in interpretability or an undiscovered astronomical empirical relation is itself an open question, and one that bears directly on using interpretability as a tool for scientific discovery.
Abstract: Classical planners can effectively solve very large deterministic MDPs represented in STRIPS or PDDL where states are sets of atoms over objects and relations, and lifted action schemas add or delete these atoms. This compact representation yields strong search heuristics and provides an ideal setting for structural generalization, since lifted relations and action schemas give rise to infinitely many domain instances. A central challenge is to learn these relations and action schemas from data, and recent approaches have addressed this problem using different types of observations. In this work, we develop a novel neural network architecture for learning action schemas from traces where states are fully observed but action arguments are unobserved. The problem is a simplification but an important step towards learning planning domains from sequences of images and action labels, and we aim to solve this simplification in a nearly perfect manner. The challenge lies in learning the action schemas while simultaneously identifying the action arguments from observed state changes. Our approach yields a robust differentiable component that can then be integrated into larger neuro-symbolic models. We evaluate the architecture on various planning domains, where the learned lifted action schemas must recover the ground-truth structure. Additionally, we report experiments on robustness to observation noise and on a variation related to slot-based dynamics models.
PaperID: 4398, Poster
Authors: Qianfang Wang, Bin Rao, Pengpeng Xu, Tiantian Chen
Abstract: Driver attention prediction is crucial for interpretable driver monitoring and human-like autonomous driving. Existing methods typically formulate this task as a scene-to-heatmap regression problem, failing to reveal the underlying dynamic attention redistribution mechanism. To address this issue, we propose a scene-task-aware competitive allocation framework, which decomposes driver attention as a bottom-up visual saliency prior and a top-down scene-task utility representation. An adaptive product-of-experts allocation rule is designed to fuse these two heterogeneous branches in a competitive manner, where scene-task demand dynamically modulates the dependence on top-down task utility, and driving control urgency adjusts spatial attention concentration via a learnable attention budget temperature. To explicitly differentiate task-critical regions from visually prominent yet task-irrelevant distractors, we incorporate competition-aware contrastive learning to enhance feature discrimination in the latent space. Extensive experiments demonstrate that our method achieves strong and balanced performance, with a Kullback–Leibler divergence of 0.98 and a normalized scanpath saliency of 4.15 on the DR(eye)VE dataset, as well as a Pearson’s correlation coefficient of 0.67 and a histogram similarity of 0.54 on the BDD-A dataset. Ablation studies and diagnostic analyses further validate that our framework yields interpretable attention allocation behaviors, including urgency-induced spatial entropy contraction and scene-task-demand-guided dynamic attention reallocation.
PaperID: 4399, Poster
Abstract: Feature caching is a promising way to accelerate Diffusion Transformers (DiTs) without modifying the model backbone, but existing methods largely rely on local heuristics or input-agnostic rules when deciding which intermediate features to cache and reuse. We present Map-Guided Caching, a feature caching framework that predicts a Temporal Variation Map (TVMap) for each denoising trajectory and uses the map to derive a global caching policy. The TVMap summarizes block-wise temporal variation across timesteps, while a lightweight Map Predictor estimates this variation from the prompt and initial noise before denoising starts. Given the predicted map, we formulate policy generation as a global path planning problem and combine a dynamic-programming-based orientation planner with a quantity planner that targets a desired speedup ratio. Across image generation, image editing, and video generation, the proposed method achieves up to 3.3× speedup on FLUX.1 [dev] and HunyuanVideo, and 2.7× on FLUX.1 Kontext [dev], while preserving competitive output quality. The advantage is especially pronounced on complex prompts, where local caching policies degrade more sharply. These results suggest that input-adaptive, globally planned caching is an effective alternative to purely local feature reuse policies.
PaperID: 4400, Poster
Abstract: This paper studies the problem of zero-shot text-attributed graph learning, which aims to generate high-quality node representations in unseen text-attributed graphs. Recent approaches usually utilize large language models (LLMs) instead of graph neural networks (GNNs) to extract semantics due to their strong generalization ability, which could neglect the intrinsic geometric structure. Towards this end, in this paper, we propose a novel approach named Prototypical Mutual Prompting Enhancement (PATH) for zero-shot text-attributed graph learning. The core of our PATH is to generate high-quality prompts using dual prototypical learning to combine the advantages of both language models and graph models. In particular, we first utilize dual graph pre-training from both instance and informativeness perspectives to generate a generalizable GNN. Then, we incorporate the frozen language and graph models into a mutual prompt learning framework. On the one hand, we extract node tokens with geometric relationships using the graph model, which will be sent to multiple prototypical projections to enhance the understanding of the language model. On the other hand, we extract graph information and task descriptions using the language model, which serves as instruction for the graph models. Extensive experiments on node classification and link prediction validate the effectiveness of the proposed PATH against strong baselines. Our code is available at
PaperID: 4401, Poster
Authors: Bingcheng Dong, Shenglan Liu, Sifan Zhang, Jirui Tian, Weitong Wu
Abstract: : existing systems usually retrieve memories for writing with a single holistic query derived from the current input. However, memory construction depends on several complementary aspects of the input, and a single query is insufficient to capture these editing needs. The second is an : long-horizon memory learning is commonly optimized over full trajectories, whereas memory construction proceeds through incremental state transitions. Effective supervision should therefore assess how each local edit changes the memory state. We propose , a graph-based memory construction framework that addresses both mismatches. GraphMemRL introduces edit-oriented query decomposition to retrieve complementary context for memory writing. It also uses action-native state-transition supervision to directly evaluate whether each graph edit improves the evolving memory state. To make long-horizon learning more tractable, GraphMemRL reformulates memory construction as factorized segment-wise optimization. The agent updates a persistent memory graph across segments under limited context budgets, preserving globally accumulated structure without full-history training. Experiments on long-memory benchmarks show that GraphMemRL improves question answering and produces better connected memory structures with more controlled memory growth. The gains are especially strong on multi-hop, temporal, and knowledge-update tasks.
PaperID: 4402, Poster
Authors:
Wenkai Fang, Shunyu Liu, Wei Luo, Zihao Wu, Yang Zhou, Tongya Zheng, Wotao Yin, Jialie Shen, Leszek Rutkowski, Mingli Song, Dacheng TaoAbstract: Recent advances in agent-based methods have demonstrated strong promise for issue localization, a critical prerequisite for software issue resolution. However, most existing agent-based methods rely on a single agent with a growing context, where long-context accumulation compresses the effective reasoning space. Meanwhile, the flat exploration structure hinders the balance between file-level breadth and function-level depth. To address these limitations, we propose NINJA, a hierarchical Navigator-INspector Joint Architecture for context-efficient issue localization. NINJA decomposes repository exploration into global navigation and local inspection: the navigator maintains the global search state, performs file-level search, and dispatches selected entry files, while inspectors independently conduct function-level exploration around the assigned files in separate contexts. Through multi-round interactions, inspectors return suspicious locations as feedback, and the navigator updates the global state to decide whether to continue exploration or finalize localization. This hierarchical design balances file-level breadth with function-level depth, while independent inspector contexts prevent local exploration traces from accumulating in a single growing context. To further strengthen both global coordination and local exploration, we introduce a two-stage agentic fine-tuning strategy. Extensive experiments across multiple benchmarks and LLM backbones show that NINJA consistently outperforms competitive baselines. Notably, after fine-tuning, Qwen3-Coder-30B-A3B-Instruct surpasses the strong closed-source Claude-Haiku-4.5 model. Our code is available at https://anonymous.4open.science/r/NINJA-67FB/.
PaperID: 4403, Poster
Abstract: BFloat16 weights of large language models carry only ~10.5 bits of information per 16-bit value, yet exploiting this redundancy on GPUs, where weights bottleneck both per-token bandwidth and concurrent capacity, has not produced wall-clock speedup. Existing variable-length lossless codes are compute-bound on SIMT decode. We present NyoomFloat12 (NF12), a lossless 12-bit fixed-length format for BF16 weights that addresses both bottlenecks. Over 99.7% of trained BF16 weights have the upper four exponent bits set to \texttt0111 after a per-matrix power-of-two scale; dropping them yields a fixed-length 12-bit format with 16-instruction SIMT-friendly decode. Out-of-range groups reuse the encoded slot to index a per-matrix verbatim BF16 buffer at zero per-group storage overhead. In the bandwidth-bound regime, fused NF12 GEMV achieves up to 1.21× matched-BF16 and 1.35× cuBLAS, 71% of the entropy-bound optimum (a hypothetical implementation reading the joint entropy floor at full HBM bandwidth). In the capacity-bound regime on Qwen3-32B, NF12's smaller footprint opens 4.5× more KV-cache room; with decode at near-peak HBM bandwidth (3.3× faster than prior lossless decoders), NF12 raises throughput up to 2.67×.
Abstract: Recent advances in one-step text-to-image generation have enabled real-time synthesis with remarkable efficiency and quality. Previous reinforcement learning methods for one-step generators combines the image space reward optimization and the diffusion noisy space distribution matching. This paradigm brings challenges due to a mismatch between terminal reward optimization and the underlying generative dynamics. As a result, optimization tends to exploit stochastic degrees of freedom, often improving reward at the expense of image fidelity. To address this issue, we propose Diff-Instruct with Diffused Reward (\textscDidr), a data-free trajectory-level alignment framework derived from Integral KL minimization. \textscDidr propagates the RLHF-optimal reward-tilted clean-image distribution across all noise levels along the diffusion trajectory. We show that this objective admits the same minimizer as clean-image RLHF, while naturally inducing the Diffused Reward Score (DRS), which acts as a reward-driven correction to the reference score function. To make this practical, we further introduce the Diffused Reward Proxy (DRP), an efficient estimator of DRS based on differentiable short-step denoising. Extensive experiments demonstrate that \textscDidr consistently Pareto-dominates existing one-step SDXL baselines. Moreover, when transferred to a 6B DiT backbone (\textscZ-Image), \textscDidr surpasses its 50-step teacher in preference alignment while requiring only a single generation step.
PaperID: 4405, Poster
Authors:
Ruining Hu, Jiaqi Lu, Xiao Liu, Ying Shen, Yu Wang, Lin ZhangAbstract: Large Vision-Language Models (LVLMs) have demonstrated strong capabilities in multimodal tasks. Nevertheless, their outputs remain prone to object hallucinations that are inconsistent with visual content. Contrastive decoding, as a training-free inference-time mitigation approach, typically intervenes in the output distribution at every decoding step. However, such indiscriminate full-stage intervention may disrupt originally correct generations and even induce new hallucinations. To address this issue, we first analyze the relationship between internal attention features and hallucinated token segments in LVLMs, revealing that hallucination-related abnormalities already emerge progressively in preceding token segments. Based on this observation, we propose HALMES, a lightweight plug-and-play module that identifies hallucination precursor segments from attention features during generation and selectively triggers contrastive decoding. HALMES can be integrated with various full-stage contrastive decoding methods, transforming indiscriminate intervention into on-demand targeted correction. Extensive experiments across different LVLM architectures and multiple benchmarks show that HALMES further improves the hallucination mitigation performance of existing contrastive decoding methods while reducing the negative impact of full-stage intervention on normal outputs, validating the importance of modeling and selectively intervening on hallucination-prone segments. All data and code will be made publicly available.
PaperID: 4406, Poster
Abstract: Diffusion models achieve high-quality image generation but remain computationally expensive due to iterative denoising across many timesteps. Existing acceleration methods mainly reduce per-sample latency, but gains along this axis are becoming saturated. We instead explore an orthogonal throughput-oriented direction by processing multiple samples within a single denoising evaluation. To this end, we propose ManiFusion, a backbone-agnostic framework for multi-sample generation through fused-space computation. ManiFusion jointly processes multiple samples by fusing them before the denoising backbone and separating them afterward through unfusion. We provide a geometric interpretation under the manifold hypothesis, focusing on identifiability and consistency of fused denoising dynamics. To better align fusion allocation with denoising stages, ManiFusion uses Trajectory-Fusion Scheduling (TFS), which dynamically adjusts fusion factors throughout sampling. ManiFusion generalizes across diffusion families, including both U-Net and transformer backbones in pixel and latent spaces. We further show that ManiFusion composes with existing acceleration methods and reaches throughput regimes unattainable by prior approaches with minimal quality degradation.
Abstract: Large language models (LLMs) are widely used in text-to-image (T2I) systems, but they are typically limited to text encoding, while denoising is handled by newly trained generative backbones. The emergence of representation autoencoders (RAEs) shifts the generation target toward semantically structured visual representations, creating a latent space that is more compatible with pretrained LLM priors. Inspired by multimodal LLMs (MLLMs), where an MLP projector is sufficient to align clean visual representations with a pretrained LLM, we repurpose the MLLM itself as a noisy representation encoder, extending this mechanism from clean to noisy inputs. We present RepFusion, which uses the resulting MLLM outputs as the conditioning signal for a diffusion transformer. In controlled, parameter-matched comparisons, RepFusion outperforms scaling the denoising transformer with newly initialized parameters. These results demonstrate that MLLMs provide strong priors for denoising visual representations and that scaling the compute of the conditional encoder is feasible in modern T2I systems.
PaperID: 4408, Poster
Abstract: Irregular multivariate time series (IMTS) are sampled at non-uniform timestamps and asynchronous across channels, their dependencies are shaped by time gaps, local ordering, and channel-specific sampling patterns. Building a faithful attention mechanism for IMTS exposes two limitations of existing work. First, current irregular-time attention methods typically regularize raw inputs into interpolated, aligned, or continuous-time-embedded forms before applying attention, which blurs exactly these irregular cues that determine how two observations relate. Second, beyond representation, standard dot-product attention scores pairs through learned token projections with no intrinsic tie to the shape of the underlying series, and an attention that does reflect this shape becomes costly to compute in high-channel IMTS. We propose Rubato, which reads two views directly from the raw irregular input: per-channel local dynamics, and a path signature that summarizes the joint multi-channel series. Attention in Rubato then scores pairs by a signature-kernel measure of series similarity rather than by a generic dot product, and a structured random projection (SORF) keeps this comparison affordable when the channel count is large. We prove that the structured projection introduces only a controlled bias that vanishes as the number of channels grows, and experiments on multiple IMTS benchmarks show that Rubato consistently outperforms strong baselines.
PaperID: 4409, Poster
Abstract: Transfer-based attacks expose serious adversarial vulnerabilities in closed-source Multimodal Large Language Models (MLLMs) by crafting adversarial examples on off-the-shelf surrogate models. However, this paradigm often drives adversarial examples toward surrogate-specific local optima, where they deceive only the surrogate model but fail to transfer to black-box MLLMs. Although existing methods mitigate this issue at the attack level, they leave the surrogate itself unchanged. To address this gap, we propose Surrogate Calibration (SCal), a novel model-level refinement that enhances downstream transfer-based adversarial attacks in a plug-and-play manner. Specifically, SCal employs a feature-preserving loss to maintain the surrogate's representational utility, while a distributional regularizer encourages its gradient field to stay closer to the natural data manifold. With calibrated surrogates, downstream attacks generate update directions that exhibit superior generalization across black-box MLLMs. Extensive experiments demonstrate that SCal boosts current attacks to new heights across various surrogates. For instance, it elevates baseline M-Attack's ASR on GPT-5.4 Mini to 53.5%, substantially surpassing the state-of-the-art MPCAttack (29.7%). These results expose critical vulnerabilities that must be addressed for trustworthy multimodal systems. Code is at: https://anonymous.4open.science/r/SCal_code-B0BC
PaperID: 4410, Poster
Authors:
Yuguang Yang, Zhewen Tan, Canyu Chen, Cheng Chi, Chunyang Liu, Kehua Sheng, Bo Zhang, Jinyu Yang, Linlin Yang, Baochang Zhang, Yan Wang, Xianbin CaoAbstract: End-to-end driving with vision-language models benefits from multimodal pretraining, but still faces a mismatch between reasoning and trajectory generation. For autoregressive VLA planners, continuous trajectories represented as step-wise normalized text waypoints are strong final outputs because they fit the native token-prediction interface of modern VLMs. However, they remain weak intermediate interfaces for reasoning-conditioned planning. Direct action tokenization provides a more explicit motion interface, but its effectiveness depends heavily on how the action codebook is constructed. We propose TrajCond-VLA, which learns a trajectory-geometric condition from reasoning for VLA planning. Our key idea is to represent this condition as a differential action trajectory whose x/y/yaw differentials are discretized separately with a 3D-Brohan codebook. The resulting DiffAction tokens preserve local motion geometry while remaining compatible with autoregressive token prediction. Building on this representation, we introduce a Two-Stage Alignment Training framework: Stage 1 predicts TrajCond from reasoning, and Stage 2 regresses the final continuous trajectory conditioned on both TrajCond and reasoning. Experiments on NAVSIM, nuScenes, and NeuroNCAP show that the learned trajectory-geometric condition is more effective as an intermediate condition than as a direct output format, and that the resulting pipeline is validated across multiple datasets with consistent gains on the completed comparisons. The nuScenes and NAVSIM navtest discussions are kept aligned to the public benchmark protocols used by Curious-VLA, and the closed-loop reference discussion uses the correct NeuroNCAP protocol.
Abstract: Expressive continuous control policies, such as diffusion and flow models, form the backbone of recent advances in scaling imitation learning for simulated and real robot control. While they are known to scale stably in the supervised imitation learning setting, incorporating them into reinforcement learning (RL) pipelines for policy improvement has proven more difficult. It often requires specialized training objectives or back-propagating through denoising processes, which cause well known issues with stability and affects scalability. In this paper we study the question of whether simple policy improvement schemes at test time alone, leaving stable supervised policy training intact, can be a competitive alternative which side-steps these issues. To this end, we propose QGF (Q-Guided Flow), a new RL algorithm that performs policy optimization entirely during test time. QGF works by pre-training both a reference flow policy (via a standard behavioral cloning objective) and a value function critic and, during test-time, using the value gradient to guide the reference policy to generate higher-value actions without any additional policy learning. Empirically, QGF outperforms prior test-time RL methods while being much cheaper to run, and is competitive with state-of-the-art training-time algorithms on single-task and goal-conditioned offline RL benchmarks with high-dimensional action spaces. Moreover, it exhibits favorable scaling with model size by avoiding the instability of actor-critic training, offering a practical and effective alternative RL algorithm with expressive policies.
PaperID: 4412, Poster
Authors:
Weiqiang Xiong, Shaohui Peng, Wenyi Li, Hao Lu, Qirui Zhou, Congying Ma, Ziming Ye, Zhenyu Yi, Yunji Chen, Qi Guo, Ling LiAbstract: High-performance inference for large models is critical in latency-sensitive applications. Existing inference engines suffer from sequential execution of hundreds of fine-grained kernels, causing kernel launch overheads and redundant memory accesses. A promising solution is fusing multiple kernels into a single large kernel, namely MegaKernel. However, existing MegaKernel implementations are tightly coupled to specific models and GPU architectures and require extensive manual engineering. We propose an LLM-driven framework AutoMegaKernel, which leverages an Instruction-Centric Fusion Abstraction to decompose MegaKernels into tractable fusion instructions, thereby enabling automated generation. Our framework jointly performs fusion strategy reasoning and CUDA code generation while hierarchically verifying correctness. Extensive experiments demonstrate that AutoMegaKernel consistently outperforms state-of-the-art inference engines (up to 5.04× vLLM and 2.42× SGLang) and compilers (up to 1.92× MPK), while supporting a broader range of models (dense, MoE, and VLA) and NVIDIA GPU platforms (Hopper, Ampere, and Ada).
Authors:
Jiantao Lin, Yingjie Xu, Mingzhi Sheng, Yangkai Wei, Hao CHEN, Yingcong ChenAbstract: Generating high-quality UV textures for 3D meshes remains challenging. Multi-view projection pipelines suffer from occlusion and view inconsistency, and recent methods that generate textures directly in UV space still rely on auxiliary modules to supply 3D information, leaving the attention mechanism tied to UV-grid positions rather than to the underlying surface geometry. This mismatch limits coherence across seams and disconnected UV islands. We propose DirectUV, an image-conditioned UV texture diffusion framework that operates in the latent UV space of a pretrained image VAE, in which a Diffusion Transformer denoises the UV latent given a single input image and a coarse UV map. At its core, Surface-Aware Positional Encoding (SAPE) replaces the standard 2D-grid positional encoding with encodings derived from per-token 3D surface coordinates obtained via UV-to-surface correspondence. As positional encodings define the distance metric that attention operates on, SAPE enables tokens to attend to each other based on true surface proximity rather than UV-grid distance, restoring coherence across seams and disconnected islands. A multi-level extension further assigns different attention heads to progressively finer subdivisions of the same latent UV patch, allowing the model to reason about surface structure at multiple granularities. Experiments show that DirectUV produces sharper and more globally consistent textures than other baselines, with the largest improvements in occluded and view-unseen regions where projection-based methods leave gaps or stretched textures.
PaperID: 4414, Poster
Authors: Moukheiber
Abstract: Code-reasoning agents trained from rollouts typically rely on costly verifiers — expert annotation, hand-written unit tests, or learned reward models — each of which scales poorly across domains. We argue that in domains with closed-form physical laws, the laws themselves form a free, dense, automatic class of reward signals. We instantiate this for physics via PhysAgentGym, an agentic code- execution gym in which a language model receives a physics scenario, writes Python that simulates it, and emits a trajectory scored by a rule-based verifier that compiles natural-language standards into executable trajectory predicates. Verifier signals are label-free: physical laws supply the ground truth without any per-task annotation. Across four physics subcategories totaling 1,320 closed-form- simulator task instances, two frontier models, Claude Sonnet 4.5 and OpenAI GPT-5, achieve nearly-identical aggregate scores (0.944 and 0.945 at N=30 each), with per-cell agreement within 0.5 pp, providing independent evidence the verifier is well-calibrated rather than arbitrary. We filter 491 successful frontier trajectories on two in-distribution subcategories into an SFT corpus and train Qwen-2.5-Coder 1.5B, 3B, 7B with QLoRA. The trained 7B adapter reaches mean = 0.942 on the four-subcategory test set, matching both frontier models within 0.3 pp on aggregate at approximately 50×lower inference cost. It scores identically to both Sonnet and GPT-5 on all three in-distribution subcategories and ties them within 0.8 pp on out-of-distribution Damping, despite never seeing damping data in training. We further show that combined sparse-and-dense reward filtering beats sparse-only filtering by 1.4–8.7 pp, that smaller models exhibit fundamental seed brittleness which disappears at 3B and above, and that scale enables the OOD generalization that 1.5B and 3B do not provide. The method is portable to any domain admitting closed-form physical laws
PaperID: 4415, Poster
Abstract: Offline goal-conditioned reinforcement learning methods have shown promise for reach-avoid tasks, where an agent must reach a target state while avoiding undesirable regions of the state space. Existing approaches typically encode avoid-region information into an augmented state space and cost function, which prevents flexible, dynamic specification of novel avoid-region information at evaluation time. They also rely heavily on meticulously designed reward and cost functions, limiting transferability to novel environments and specifications. We introduce RADT, a decision transformer approach for offline, reward-free, goal-conditioned, avoid-region-conditioned RL. RADT encodes goals and avoid-regions directly as prompt tokens, allowing any number of avoid-regions of arbitrary size to be specified at evaluation time. Using only suboptimal offline trajectories from a random policy, RADT can learn reach-avoid behavior in a completely data-driven manner using our novel avoid-region hindsight relabeling approach, without the need for reward/cost supervision. We benchmark RADT against existing offline goal-conditioned RL models across 17 tasks, environments, and experimental settings. RADT generalizes in a zero-shot manner to out-of-distribution avoid-region sizes and counts, outperforming baselines that require retraining. In one such zero-shot setting, RADT achieves 35.7% improvement in normalized cost over the best retrained baseline while maintaining high goal-reaching success. We also apply RADT to cell reprogramming in biology, demonstrating its versatility.
Abstract: Humans, animals, and modern machine learning models exhibit impressive abilities to learn complex behaviors and generalize these behaviors to unseen situations. This ability requires us to learn rules and regularities that allow for such generalizations. At the same time, in most complex environments, any rule will have its exceptions. How do learning systems balance between learning general regularities and memorizing exceptions? We argue that a lack of task paradigms has hindered the study of this essential ability. To address this gap, we introduce a novel task, transitive inference with exceptions, that tests for relational generalization and memorization of an exception to the relational rule. We then analytically characterize the behavior of a simple, theoretically tractable model of neural network learning (kernel ridge regression) across a broad family of representations and task parameters. We find that these models can balance between relational generalization and memorization, but unlike for transitive inference without an exception, successful generalization is sensitive to the specific representational geometry. We explain why this task is more challenging mechanistically by drawing on our analytical theory. Finally, we validate our theoretical insights in pretrained language models that are finetuned on ordered relations, finding that these models successfully generalize according to the transitive rule, but also make the kinds of systematic mistakes predicted by our theory. Overall, our theory shows how learning systems can balance between relational generalization and memorization, explains how this can go wrong, and emphasizes the need for new task paradigms designed to probe this ability.
PaperID: 4417, Poster
Abstract: LLM-based multi-agent systems (MAS) have attracted growing attention for improving reasoning through interaction among multiple agents. In this work, we focus on parallel multi-agent reasoning systems, where several agents solve the same problem over multiple rounds and aggregate their outputs into a final answer. Despite their strong reasoning performance, uncertainty estimation for such systems remains underexplored: the reliability of a MAS depends not only on individual generations, but also on how agents interact and evolve across rounds. We propose SPI (Sequential Probabilistic Inference), a lightweight, training-free uncertainty estimator that formulates MAS uncertainty as sequential inference over a latent system-level belief. SPI aggregates round-level agreement and generation-uncertainty signals through a filtering-style update. Across five backbones, six benchmarks, and two MAS protocols, SPI improves misclassification detection, selective prediction, and calibration over a broad set of uncertainty estimation baselines, including standard log-likelihood-based methods and MAS-specific estimators.
Abstract: Naturalistic fMRI offers a time-resolved window into cortical dynamics during continuous multimodal experience, where future brain states are jointly governed by endogenous neural context and exogenous sensory drive. Forecasting such dynamics, however, remains challenging because fast-changing stimulus streams must be reconciled with the sluggish hemodynamic BOLD response, while cortical activity is organized across functionally heterogeneous brain networks. To address these challenges, we introduce BrainVista, a multimodal autoregressive framework for observed-stimulus conditional brain-state forecasting. BrainVista predicts future fMRI states from the history of brain activity and stimulus tokens temporally aligned to each prediction query, enabling autoregressive prediction without access to future ground-truth fMRI states. This design supports causal, temporally consistent modeling of naturalistic brain dynamics under realistic sensory stimulation. Specifically, we design Network-wise Tokenizers for functionally structured brain representations, a Spatial Mixer Head for cross-network interaction refinement, and Stimulus-to-Brain masking to preserve temporally causal stimulus–brain conditioning while preventing future brain-state leakage and stimulus look-ahead. We evaluate BrainVista on Algonauts 2025, CineBrain, and HAD, three large-scale naturalistic fMRI benchmarks, under long-horizon conditional autoregressive forecasting. BrainVista consistently outperforms baselines, improving pattern correlation by 10.8% and 9.4% relative to the strongest baseline on Algonauts 2025 and CineBrain, respectively.
Abstract: We study optimistic bilevel optimization when the lower-level problem has a non-isolated manifold of minimizers. In this setting, the hyper-objective may be non-differentiable because the upper-level criterion must choose among multiple lower-level solutions. Under a local Polyak--\Lojasiewicz (P\L) condition, we show that differentiability does not require the lower-level solution set to be a singleton: uniqueness of the optimistic selection is sufficient. This yields an explicit pseudoinverse-based hyper-gradient formula extending the classical singleton-minimizer result. We further characterize the regularity of the hyper-objective: non-degeneracy of the selected minimizer along the solution manifold yields local smoothness, while failure of uniqueness can create many non-differentiable points and failure of non-degeneracy can destroy all positive H\"older regularity of the hyper-gradient. Motivated by this theory, we propose HG-MS, a select-then-differentiate method combining explicit optimistic selection with efficient pseudoinverse-based hyper-gradient computation. Despite the nonconvex nature of optimistic selection over the lower-level solution manifold, we show that HG-MS converges to a stationary point of the optimistic objective with complexity governed by the intrinsic dimension of the solution manifold rather than its ambient dimension. Empirically, we test a practical variant of HG-MS for matched-budget LLM source reweighting. This variant preserves the select-then-differentiate principle and obtains the best GSM8K/MATH scores across the tested backbones, along with competitive or best MT-Bench instruction-following results.
PaperID: 4420, Poster
Abstract: Conventional text-based person re-identification (ReID) approaches mainly rely on visible-spectrum images captured from ground-level viewpoints. However, with the increasing prevalence of UAVs and multispectral imaging systems, pedestrians can be captured from both ground and aerial viewpoints, and different spectral sensors provide complementary cues that reveal diverse pedestrian characteristics. As a result, conventional text-based ReID methods become inadequate for such retrieval scenarios. To address this limitation, we introduce a novel task termed Text-based Multispectral Aerial-Ground Person Re-Identification (TMAG-ReID), which aims to retrieve pedestrian images across aerial and ground viewpoints under multispectral scenarios using textual descriptions. To facilitate this task, we construct the T-MSAG dataset, which consists of multispectral pedestrian images captured from aerial and ground viewpoints, paired with corresponding textual descriptions. Furthermore, we propose a novel View-Spectral Reconciliation Learning framework (VSRL) that leverages textual guidance to interweave discriminative semantics across distinct views and spectral modalities. This fusion yields robust cross-view and cross-modal pedestrian features, thereby enabling high-performance retrieval. Extensive experiments on the T-MSAG dataset demonstrate the effectiveness of VSRL.
PaperID: 4421, Poster
Abstract: Recent advances in Time Series Foundation Models (TSFMs) increasingly rely on axial attention mechanisms, but despite their empirical success, the theoretical properties of axial attention remain largely unexplored. To bridge this gap, we extend the Single Location Regression framework introduced in Marion et al. [2025] to 2D structured data, in order to study sparse information settings where a target depends on a small number of tokens. We formulate an analytically tractable axial attention-type predictor and rigorously analyze its statistical properties and training dynamics. We prove that under a specific asymptotic regime and with knowledge of the underlying oracle parameters, our predictor achieves Bayes optimality, whereas a class of semi-linear simplifications strictly fails to do so, underlining the role of the inner non-linearity. We prove that for a specific temperature scaling, Projected Gradient Descent converges globally to the optimal oracle parameters. Extensive numerical experiments validate our theoretical findings and highlight the critical role of inverse temperature scheduling in empirical convergence.
Abstract: This paper studies the task of estimating heterogeneous treatment effects in causal panel data models with covariate effects. We propose a novel ) framework, that cohesively deal with the underlying heterogeneity and nonlinearity of both panel units and covariate effects. CoDEAL integrates nonlinear covariate effect components (parameterized by a feed-forward neural network) with nonlinear factor structures (modeled by a multi-output autoencoder) to form a heterogeneous causal panel model. The nonlinear covariate component flexibly captures complex covariate influences on outcomes, and the nonlinear factor decomposition enables CoDEAL to effectively capture both cross-sectional and temporal dependencies inherent in the data panel. This latent structural information is subsequently integrated into a customized matrix completion algorithm, thereby facilitating more accurate counterfactual imputation. Moreover, the use of a multi-output autoencoder explicitly accounts for heterogeneity across units and enhances the model interpretability of the latent factors. We establish theoretical guarantees on the convergence of the estimated counterfactuals and demonstrate the compelling performance of the proposed method using extensive simulation studies and real-world data applications.
PaperID: 4423, Poster
Abstract: Pre-trained vision encoders contain layer-wise visual representations that differ in spatial granularity, semantic abstraction, and sensitivity to local details. However, most Multimodal Large Language Models (MLLMs) rely only on the final or penultimate vision encoder representations or fixed aggregation rules, making visual abstraction largely query-agnostic and limiting access to fine-grained cues such as small objects, spatial details, text, and subtle visual attributes. In this work, we propose , an instruction-conditioned layer routing approach that dynamically aggregates query-relevant latent representations from intermediate vision encoder layers. Given a text query, predicts routing probabilities over vision encoder layers and aggregates selected hidden states at either the image level, patch level, or through a hybrid routing mechanism. Our proposed is a vision encoder-agnostic framework for layer-wise routing that selectively samples visual features useful for fine-grained visual reasoning tasks. Experiments across seven fine-grained visual reasoning benchmarks demonstrate substantial performance improvements, including +18.9% on V overall accuracy, +4.5% on HRBench4K, and +16.3% on CharXiv compared to baseline MLLMs, requiring multi-resolution inputs, simple interleaving of multiple vision encoders, or additional patch tokens. We further analyze receptive field scales and routing behaviors across vision encoder layers to explain why adaptive layer selection improves perception and reasoning. Our results suggest that conditional intermediate-layer representations are a key step toward stronger visual perception and reasoning in MLLMs. Code will be open-sourced.
PaperID: 4424, Poster
Abstract: Sampling from reward-tilted distributions enables diffusion models to satisfy task-specific constraints without retraining. However, existing methods either rely on gradients or struggle to explore multimodal high-reward regions. We introduce a new test-time steering method for diffusion models that modifies the reverse denoising process to incorporate reward information directly. At each step, we propose a Metropolis–Hastings–corrected denoising mechanism based on a preconditioned Crank–Nicolson (pCN) proposal, enabling principled acceptance of noise updates that bias sampling toward high-reward regions. To further improve exploration in multimodal reward landscapes, we extend this procedure with parallel tempering across multiple temperature levels, allowing controlled mixing between exploration and refinement regimes. Our method operates with or without reward gradients and applies to black-box objectives. Empirically, we demonstrate its effectiveness on synthetic tasks, image generation, dynamical systems, and Bayesian inverse problems. Compared to prior methods, it achieves superior reward alignment and multimodal exploration while maintaining stability across tasks.
Abstract: Graph Network Simulators (GNS) have emerged as powerful surrogates for complex physics-based simulation, offering inherent differentiability and orders-of-magnitude speedups over traditional solvers. However, GNSs typically assume access to the underlying material parameters, such as stiffness or viscosity, severely limiting their utility in realistic experimental settings. While recent meta-learning approaches address the parameter dependency by inferring properties from mesh trajectories, reconstructing a mesh from an observed scene is difficult. In this work, we introduce Point Cloud Encoding for Accurate Context Handling (PEACH), a novel framework that applies in-context learning on point clouds to adapt a learned simulator to unseen physical properties during inference. Our approach relies on a novel spatio-temporal point cloud sequence encoder, as well as two forms of auxiliary supervision to help improve simulation fidelity. We demonstrate that PEACH is capable of accurate zero-shot sim-to-real transfer on a challenging, dynamic scene. Experiments on simulation scenes show that PEACH even outperforms mesh-based baselines on simulation accuracy, while being much more practical for real-world deployment.
Abstract: Accurate and well-calibrated Machine Learning (ML) models are mandatory in high-stakes settings, yet effective multiclass calibration remains challenging: global approaches assume calibration errors are homogeneous across the latent space, while local methods often rely on latent-space dimensionality reduction, which leads to information loss. To address these issues, we propose a compositional approach to multiclass calibration, where region-specific calibration maps are constructed from shared codeword-dependent factors. We instantiate this idea via Vector Quantization (VQ), which induces a structured partition of the representation space, and an indexed parameterization of Dirichlet concentrations that enables parameter sharing across regions. Our approach learns heterogeneous calibration maps that generalize well even to sparse regions of the latent space. Experiments on benchmark datasets show significant improvements in local calibration while maintaining competitive global calibration and predictive performance.
PaperID: 4427, Poster
Authors: Junxi Chen, Payam Piray
Abstract: Adaptive learning requires distinguishing environmental volatility from observation stochasticity, two sources of uncertainty that demand opposite adjustments to the learning rate but inflate experienced variance similarly. Disentangling them is computationally difficult with no tractable closed-form solution. Particle-filter methods are the natural tool for this kind of joint inference, but their stochastic likelihoods and non-differentiable objectives force derivative-free fitting protocols and discourage the individual-difference analyses central to cognitive modeling, where small effect sizes leave little room for additional estimator noise. We introduce the Categorical Bayes Filter (CBF), a deterministic alternative that preserves the conditional structure of recent particle-filter accounts but replaces the stochastic outer layer with a categorical distribution on a quantile grid parameterized through differentiable Beta quantile functions. The procedure performs evidence maximization with an exact, deterministic marginal likelihood that is fully differentiable in the grid parameters. In a volatility-stochasticity task with N = 643 participants, fitted CBF dispersion parameters reveal a cross-over phenotyping pattern between volatility-blind and stochasticity-blind subjects that is not recoverable from particle-filter parameters fit to the same data under a state-of-the-art protocol. The deterministic structure also yields a trial-by-trial ambiguity signal that predicts response times not used in fitting. More broadly, the approach opens individual-level analyses in cognitive modeling and computational psychiatry that stochastic methods have effectively foreclosed.
PaperID: 4428, Poster
Authors: Yangyang Liu
Abstract: Token-level maximum likelihood and mean negative log-likelihood dominate the training and evaluation of large language models (LLMs), yet long-horizon generation often suffers from late-stage degradation despite improved mean loss. We identify two coupled causes of this mismatch: mean token loss is not scale-free with respect to sequence length, so small per-token divergences can amplify into large sequence-level distribution shifts, and long-horizon failures are spike-dominated, where rare but abrupt prefix deviations trigger errors that are largely invisible to mean objectives. To address these issues, we propose Prefix Likelihood-Ratio Control (PLRC), which centers training on the prefix log-likelihood ratio process and its increments to directly capture sequence-level deviation and spikes. As the data likelihood is unknown, PLRC learns a prefix ratio critic via noise-contrastive prefix discrimination and introduces tail-sensitive regularizers on the terminal ratio and its increments, yielding guarantees that bound long-horizon event distortion and the probability of the first spike. PLRC preserves standard teacher-forcing training without architectural changes while explicitly targeting failure modes that mean token loss cannot capture.
PaperID: 4429, Poster
Abstract: Does multimodal training make a language model think visually, even when the input is text only? We test this with targeted causal ablations across 10 multimodal configurations from 6 model families, with matched text-only backbones and cross-family transfer controls. We contrast spatial-imagery tasks that invite mental scene construction with symbol-manipulation tasks solvable by symbolic rules alone. We identify a vision-language axis in hidden-state space and ablate its projection at inference. In Chameleon-7B, this ablation drops accuracy on BIG-Bench Hard (BBH) navigate by 27.6 percentage points while improving boolean expressions by 5.2 percentage points, a 32.8 percentage-point double dissociation. The effect is direction-specific and layer-localized. Together, these results reveal emergent visual thinking (EVT): text-only reasoning that causally depends on visual-spatial representations shaped by multimodal training. In dual-path unified models (BAGEL, Janus-Pro), the causal handle localizes to the generation pathway as a compact one-dimensional carrier. The understanding pathway also encodes spatial information in a higher-rank, probe-decodable form, but is not the causal handle. To separate decodability from causal use, we analyze concentration, alignment, and semantic location. These diagnostics explain why behavior, probing, and intervention can diverge. Multimodal training therefore reshapes the geometry of text-only reasoning, beyond changing how visual input is processed.
PaperID: 4430, Poster
Authors:
Cheng Qian, Hyeonjeong Ha, Jiayu Liu, Jeonghwan Kim, Emre Can Acikgoz, Bingxuan Li, Kunlun Zhu, Jiateng Liu, Aditi Tiwari, Zhenhailong Wang, Xiusi Chen, Heng JiAbstract: Large multimodal models (LMMs) have advanced rapidly in perception and reasoning, yet it remains unclear whether they can discover visually grounded, physically feasible solutions in open-ended environments. We first introduce , a benchmark for affordance-grounded creative tool use, where models must inspect scenes, entities, and parts to identify non-obvious object uses grounded in visual evidence. Our evaluation shows that current LMMs often fail to sustain grounded exploration: they overlook relevant entities, under-examine critical parts, or hallucinate unsupported attributes. To address these failures, we propose , framing creative tool use as a preference learning problem. Using Direct Preference Optimization and supervision from an affordance knowledge base, we train models to favor visually grounded attribute–affordance reasoning over hallucinated alternatives while improving exploration efficiency. Our results yield consistent gains in correct entity and part selection, and reduce hallucination and grounding errors. These findings position grounded creativity as a core capability for future multimodal agents: the ability to adapt to unfamiliar environments and solve problems beyond memorized patterns or surface-level plausibility, moving closer to real human-like intelligence.
PaperID: 4431, Poster
Authors:
Ishtiak M Saad, Mominul Islam, Shahriyar Z Ridoy, Md Manjurul Ahsan, Azmine Toushik WasiAbstract: Chunking is a hidden bottleneck in dense retrieval and retrieval-augmented generation: it determines the units that can be indexed, retrieved, and ultimately used as evidence. Yet most systems still segment documents using fixed token windows or globally thresholded semantic similarity drops, treating chunking as an engineering heuristic rather than a statistical decision. We propose models consecutive sentence-embedding cosine similarities as a document-level signal and derives boundary decisions from the null hypothesis of no semantic transition. Across four development BEIR corpora, we establish three corpus-invariant properties: bimodal document structure, rapid survival decay of shallow local minima, and positive lag-1 autocorrelation. These properties show that semantic boundaries should be detected by variance-normalized deviations, not by document means or global similarity thresholds. Accounting for autocorrelation leads to a normalized discrete Laplacian detector that identifies significant semantic valleys through a single interpretable decision rule. Evaluated on five BEIR benchmarks, consistently improves over fixed-length, recursive, and production semantic chunking baselines under one fixed operating configuration. Without corpus-specific retuning, the same detector transfers to a held-out TREC-COVID corpus. Our results suggest that the boundaries needed for effective retrieval are already encoded in embedding-similarity signals, and that principled, efficient, LLM-free chunking can be obtained by detecting them statistically.
PaperID: 4432, Poster
Abstract: Optimal control of partially observable Markov decision processes (POMDPs), in general, requires a complete history. Traditionally, a belief state is used as a sufficient statistic to control POMDPs. However, being defined over a continuous space, its full representation is impossible with a finite-memory controller. This raises the question we address in this work: what must the available finite memory preserve to achieve near-optimal control performance? We study this question through an information-theoretic converse that gives a fundamental lower bound on the loss of any finite-memory controller. The analysis introduces the concept of a decision witness that characterizes belief states whose confusion would lead to unavoidable loss. Motivated by this bound, we propose RECAP, a constructive quantization method to design representative memory states. We prove upper bounds on the optimality gap of any RECAP codebook, and use the resulting components together with the converse to define the codebook objective. Experiments on 11 POMDP benchmarks demonstrate that RECAP improves over finite-memory representation baselines under the same memory budget.
PaperID: 4433, Poster
Abstract: We propose a novel solution to reward hacking in reinforcement learning. Our setup involves a misspecified reward function corresponding to task completion, and a classifier that flags invalid solutions. As in frontier LLM development, the classifier makes systematic errors. Consequently, reward penalties and data filtering based on the classifier fail to prevent reward hacking. To address this limitation, we introduce GRAFT: Gradient Routing Adapter Fine Tuning. Rather than penalizing or filtering flagged behaviors during training, we constrain (``route'') gradient updates to specific parameters. This has two benefits. First, the model does not learn to exploit classifier errors, i.e., it does not learn to hide the unintended behavior. Second, the unintended behavior can be disabled by ablating the parameters to which they were localized. We demonstrate these benefits in a variety of RL environments where reward penalties result in monitor-subverting policies. Additionally, we show that our method remains effective when we scale model size and task complexity by distilling frontier LLM trajectories into Qwen3-32B. Overall, we show that GRAFT is a promising approach for leveraging unreliable monitors during post-training in environments where true performance is difficult to evaluate.
Abstract: Bilateral trade models the task of intermediating between two strategic agents, a seller and a buyer, who wish to trade a good. We study this problem from the perspective of a profit-maximizing broker within an online learning framework, where the agents' valuations are generated by a smooth adversary. We devise a learning algorithm that guarantees a \tildeO(\sqrtT) regret bound, which is tight in the time horizon T up to poly-logarithmic factors. This matches the minimax rate for the stochastic i.i.d. case, and is also well separated from the adversarial setting, where sublinear-regret is unattainable. By extending the strong regret guarantees from the i.i.d. case to the smooth adversary, we significantly broaden the scope of settings where such fast rate is achievable, while closing an important gap in the regret landscape of this fundamental economic problem. To overcome the challenges posed by this adversary, we leverage a continuity property of smooth instances and combines this with a hierarchical net-construction of the broker's action space, which is analyzed via algorithmic chaining. We showcase the applicability of these techniques by deriving a similarly tight \tildeO(\sqrtT) regret bound for a related mechanism design model: the joint ads problem.
Abstract: Prefill-only KV compression freezes a token subset at the end of prefill and decodes from it without further eviction. The retention decision is therefore irreversible, yet existing methods estimate the corrective signals it relies on, per-head reliability and prompt-level compression sensitivity, online from a single noisy prompt. We argue this is the wrong statistical unit: these signals exhibit far higher cross-prompt regularity than within-prompt signal-to-noise. We introduce CompilerKV, a KV-retention policy whose corrective tables are compiled offline from a calibration corpus, reducing online correction after the standard observation-window scan to O(1) lookups plus a budget clamp. We find that compiled retention tables behave as portable architectural priors: rankings transfer across disjoint corpora on four backbones, with mean Spearman \bar\rho=0.90, and direct model-to-model table transfer costs only 0.4-0.8 LongBench points on average. At a 512-token budget, CompilerKV attains compressed-SOTA on all four backbones, improving over the strongest prefill-only baseline by +1.67 points on average, with task-bootstrap 95% CI [+1.08,+2.37]. Pressure regimes amplify the gap: under a fixed 512/32\mathrmk cache ratio, CompilerKV remains the strongest compressed method through 128k RULER, with approximately 73 versus FullKV 79 and SnapKV 38; on 32k NIAH it reaches 0.89 versus SnapKV 0.42; and at 32k input, retaining only 1.56% of the prefill KV, batch-16 serving remains feasible where FullKV is OOM.
Abstract: Lookahead-based acceleration methods, such as Nesterov’s momentum, are widely used in optimization, but they often become unreliable in deep learning training due to stochastic gradient noise and non-convex loss landscapes. Standard lookahead relies on short-horizon update signals (e.g., differences between consecutive iterates), which are inherently noisy and can lead to unstable extrapolation directions. This work revisits Nesterov's acceleration from a trajectory perspective and argues that effective acceleration in deep learning should follow the low-frequency trend of optimization trajectories rather than extrapolating noisy one-step updates. Leveraging on this insight, we propose EMA-Nesterov, a simple modification that replaces the standard Nesterov's lookahead direction with an exponential moving average of parameter updates. This yields a stabilized lookahead direction that approximates long-horizon trajectory trends while retaining a lightweight single-loop implementation. We show that EMA-Nesterov retains the same theoretical accelerated convergence rate in convex problems. Furthermore, we show that EMA-Nesterov consistently improves optimization performance across a range of optimizers, including Adam, SOAP, and Muon, in deep learning tasks such as language model pre-training. Compared to prior lookahead methods, EMA-Nesterov achieves better performance by avoiding the instability of short-horizon lookahead and inefficiency of multi-step lookahead.
Authors:
Chi Zhang, Haibo Qiu, Qiming ZHANG, Yufei Xu, Xinbo Gao, Jing ZhangAbstract: We present Seirênes, a self-play RL framework that transforms contextual interference from a failure mode of LLM reasoning into an internal training signal for co-evolving more resilient reasoners. While RL with verifiable rewards has significantly advanced reasoning capabilities, models can still exhibit fragility when encountering non-idealized contexts: scenarios characterized by superfluous information, tangential instructions, or incidental correlations that differ from the clean distributions typical of standard benchmarks. Seirênes harnesses this vulnerability through a parameter-shared and adversarial self-play loop. Within this framework, a single model is trained to both construct plausible yet distracting contexts that expose its own reasoning blind spots, and solve problems by discerning the essential task from these perturbations to recover the core underlying logic. By pitting these competing objectives against each other, Seirênes compels the model to move beyond superficial pattern matching and anchors its capabilities in robust underlying reasoning. This continuous interaction sustains an informative co-evolutionary curriculum as the model improves. Across seven mathematical reasoning benchmarks and model scales from 4B to 30B, Seirênes achieves average gains of +10.2, +9.1, and +7.2 points. Besides, distracting contexts produced by the 4B Seirênes model reduce the accuracy of top-tier closed-source models (GPT and Gemini) by roughly 4--5 points, revealing Seirênes' general ability to uncover reasoning models' blind spots. Our code will be available.
PaperID: 4438, Poster
Abstract: Conditional human motion generation remains a fundamental challenge in computer vision and robotics. Despite significant progress, current methods are often constrained by fixed modality configurations and task-specific architectures, leaving cross-modal interactions and the scaling laws of multimodal-conditioned synthesis largely underexplored. A key bottleneck is the scarcity of large-scale modality-aligned motion data, limiting generalization across diverse control signals. In this work, we introduce OmniHuMo, a large-scale, high-quality dataset comprising over 5,000 hours of motion and 3.2 million sequences with precisely aligned multimodal annotations (e.g., text, speech, music, and trajectory). Leveraging OmniHuMo, we propose AnyMo, a unified multimodal framework combining a Residual FSQ-based motion tokenizer with a scalable masked modeling transformer, enabling high-quality motion synthesis under arbitrary modality combinations. Extensive experiments show that AnyMo achieves high-fidelity synthesis while offering flexible control over both spatial and stylistic attributes.
PaperID: 4439, Poster
Abstract: Modern large language models are able to compose skills in order to perform complex tasks, many of which might not have been seen during training. The details of how exactly this composition occurs remain elusive. In this paper, we study a mechanism for compositional generalization in transformers by considering a simple controlled setting involving variable assignment and modular addition. By partitioning our training data into disjoint sets, we observe that small transformers are able to generalize to previously unseen combinations of variables and numbers. Our mechanistic analysis shows that the same ``modular addition'' MLP module is used whether the inputs are given directly or indirectly through a separate variable assignment mechanism. We also analyze the training dynamics from an empirical lens, which reveals three phases of learning: first, modular addition is learned, then the structure required for variable assignment, and finally a refinement phase where the model generalizes to sequences not seen in training. Finally, we provide a theoretical framework to explain how compositionality emerges from training dynamics. These results suggest that compositional generalization can be a natural consequence of the compositionality of internal mechanisms in transformers.
Abstract: Time series analysis underpins forecasting, monitoring, and decision making in domains such as finance and weather, where solving a task often requires both numerical accuracy and contextual reasoning. Recent progress has moved from specialized neural predictors to approaches built on LLMs and foundation models that can reason over time series inputs and use external tools. However, most such systems remain execution-centric: they focus on solving the current instance but learn little from exploratory execution. This is especially limiting in verifiable numeric settings, where multiple candidate executions and tool-use procedures may all be task-valid yet differ sharply in quantitative quality, and where early success can trigger tool-prior collapse that suppresses further exploration. To address this limitation, we present TimeClaw, an exploratory execution learning framework that turns exploratory execution into reusable hierarchical distilled experience through a four-stage loop: Explore, Compare, Distill, and Reinject. TimeClaw combines metric-supervised exploratory execution learning, task-aware tool dropout, and hierarchical distilled experience for inference-time reinjection, while keeping the base model frozen and avoiding online test-time adaptation. In an MTBench-aligned evaluation with 17 tasks that span finance and weather prediction and reasoning tasks, TimeClaw delivers consistent gains over the baselines. These results suggest that, for scientific systems, the bottleneck is not only execution-time capability, but how exploratory experience is compared, distilled, and reused.
Abstract: Long-range geophysical forecasts are fundamentally limited by chaotic dynamics and numerical errors. While data assimilation can mitigate these issues, classical variational smoothers require computationally expensive tangent-linear and adjoint models. Conversely, recent efficient latent filtering methods often enforce weak trajectory-level constraints and assume fixed observation grids. To bridge this gap, we propose Latent Ensemble Variational Data Assimilation (LEVDA), an ensemble-space variational smoother that operates in the low-dimensional latent space of a pretrained differentiable neural dynamics surrogate. By performing four-dimensional ensemble-variational (4DEnVar) optimization within an ensemble subspace, LEVDA jointly assimilates states and unknown parameters without the need for adjoint code or auxiliary observation-to-latent encoders. Leveraging the fully differentiable, continuous-in-time-and-space nature of the surrogate, LEVDA naturally accommodates highly irregular sampling at arbitrary spatiotemporal locations. Across three challenging geophysical benchmarks, LEVDA matches or outperforms state-of-the-art latent filtering baselines under severe observational sparsity while providing more reliable uncertainty quantification. Simultaneously, it achieves substantially improved assimilation accuracy and computational efficiency compared to full-state 4DEnVar.
Abstract: Text-to-image (T2I) diffusion models have achieved strong performance in semantic alignment, yet they still struggle with generating the correct number of objects specified in prompts. Existing approaches typically incorporate auxiliary counting networks as external critics to enhance numeracy. However, since these critics must provide gradient guidance during generation, they are restricted to regression-based models that are inherently \emphdifferentiable, thus excluding detector-based models with superior counting ability, whose count-via-enumeration nature is \emphnon-differentiable. To overcome this, we propose Detector-to-Differentiable (\emphD2D), a novel framework that transforms non-differentiable detection models into differentiable critics, leveraging their superior counting ability to guide numeracy generation. Specifically, we design custom activation functions to convert detector logits into soft binary indicators, which are then used to optimize the noise prior at inference time with pre-trained T2I models. Our extensive experiments on SDXL-Turbo, SD-Turbo, and Pixart-DMD across four benchmarks of varying complexity (low-density, high-density, and multi-object scenarios) demonstrate consistent and substantial improvements in object counting accuracy (e.g., boosting up to 13.7% on D2D-Small, a 400-prompt, low-density benchmark), with minimal degradation in overall image quality and computational overhead.
PaperID: 4443, Poster
Abstract: Foundation Models for Electroencephalography (EEG) are trained with masked reconstruction, contrastive alignment, and language-model interfaces, yet their learned representations remain difficult to inspect: raw EEG lacks canonical token semantics, and models may preserve waveform detail while discarding physiologically meaningful structure. We introduce Neuro-KE (Neuro-Knowledge Engine), a knowledge interface that computes EEG descriptors without additional human annotation and renders them as paradigm-specific supervision for semantic grounding. Rather than using handcrafted features as standalone classifiers, Neuro-KE converts the same domain knowledge into numeric prediction targets, feature-state text anchors, and feature-grounded instruction answers. Across three interfaces, Neuro-KE improves aggregate grounding and transfer metrics, with task-dependent exceptions: it improves balanced accuracy (BAcc) by 2.60 points over reconstruction-only masked pretraining across 12 datasets and three backbones, improves six-dataset EEG--text validation BAcc by 2.23 points over label-only anchors, and improves EEG-MLLM feature-grounded generation to ROUGE-L F1 0.607 and BERTScore F1 0.929. These results suggest that established EEG descriptors can serve as reusable grounding signals for inspecting and supervising EEG foundation-model interfaces.
Abstract: Test-Time Reinforcement Learning (TTRL) constructs pseudo-supervisory signals by sampling reasoning trajectories during inference, enabling online adaptation to current inputs. However, existing research predominantly focuses on single-task settings, overlooking continuous updates across task streams. Theoretical and empirical analyses reveal two coupled challenges in this setting: (1) Error Accumulation: Majority voting tends to treat erroneous consensus as pseudo-labels, with such bias amplifying across sequential parameter updates; (2) Catastrophic Forgetting: Gradients from new tasks conflict with historical knowledge, causing acquired reasoning patterns to be overwritten. These issues form a positive feedback loop: pseudo-label noise exacerbates gradient interference, while degradation of historical knowledge reduces subsequent trajectory quality, jointly driving performance deterioration during continual adaptation. To address this, we propose the CTRL framework: first, we introduce a Process Reward Model (PRM) to replace outcome-level voting, filtering high-confidence trajectories via fine-grained process scoring to suppress pseudo-label bias; second, we design a cognitive anchor-driven soft gradient correction module that dynamically recomputes anchor gradients from a historical trajectory memory pool and projects out update components negatively correlated with historical knowledge, constraining the parameter update trajectory. Experiments demonstrate that CTRL significantly outperforms existing TTRL baselines on continuous task streams, exhibiting stronger robustness in blocking error propagation and preserving historical reasoning patterns.
Abstract: Long-form technical text generation underpins knowledge-intensive workflows, yet remains challenging for large language models (LLMs) due to the need for globally consistent logical structuring and faithful technical reasoning beyond local coherence. Patent drafting is a canonical instance of this challenge, demanding holistic generation of a legally compliant and technically exhaustive document through sustained multi-expert collaboration. Existing approaches often focus on partial section generation or rely on manually crafted outlines, limiting scalable automation in realistic settings. In this work, we propose LogicTree-RAG, a logic tree-guided retrieval-augmented generation framework that induces a hierarchical logic tree as a global organizational backbone to organize and ground technical disclosures, without relying on expert-defined drafting priors. Each node in the logic tree represents a technical element and is constructed through evidence-guided recursive generation. A hybrid traversal mechanism then maps the logic tree into patent sections, enabling controllable and section-balanced generation. Extensive experiments show that LogicTree-RAG consistently improves content quality and language conformity over strong LLM-based baselines and achieves longer structured generation with high token efficiency, demonstrating the effectiveness of logic-centric generation for complex technical document drafting.
PaperID: 4446, Poster
Authors:
Fanda Fan, Yuxuan Yang, Xiaorui Wang, Luqi Gong, Xiang Wentao, Xiang Ao, LimingSong, Yinuo Zhu, Yutong Chen, Chunjie Luo, Jianfeng ZhanAbstract: Time series forecasting has recently moved beyond static, single-pass prediction toward agentic methods that reason, use tools, execute workflows, and accumulate memory. However, most existing forecasting agents still optimize predictions or pipelines under largely implicit configurations, leaving the construction of the forecasting system itself underexplored. This limits their ability to attribute performance changes to specific decisions, such as which module, exogenous variable, or hyperparameter caused an improvement. We propose , an evaluation-grounded agentic framework for constructing time series forecasting systems. Nuwa defines a structured design space over modules, exogenous variables, and hyperparameters, and performs Evaluation-Budgeted Construction that uses data meta-information, performance diagnostics, and memory to propose candidate sets, select budgeted subsets for controlled evaluation, attribute local effects, and decide whether to stop, accept, retry, switch subspace, or continue construction. The resulting design-level evidence is stored in a Design Memory, enabling reusable knowledge about which configurations work under which data conditions. Across approximately 15,000 experimental runs, Nuwa achieves consistent performance gains across module, exogenous-variable, and hyperparameter construction, with an overall average MSE reduction of 8.0% over the corresponding reference settings. By turning forecast-level feedback into attributable and transferable design evidence, Nuwa makes forecasting system construction an explicit objective for agentic time series forecasting. The anonymous repository is available at https://anonymous.4open.science/r/NUWA/.
PaperID: 4447, Poster
Authors: Dibyanshu Kumar, Magda Gregorova
Abstract: Plug-and-Play (PnP) methods with generative priors have achieved strong performance in inverse problems, with recent flow matching models enabling fast and scalable reconstruction. However, a key challenge in PnP methods is maintaining consistency between data fidelity updates and the learned generative prior, as data-consistency steps can move iterates away from the flow trajectory. We propose an ADMM-based PnP framework for flow matching priors that enforces stronger consistency between the measurement model and the prior through auxiliary variables and dual updates. In addition, we introduce a trajectory correction mechanism that restores consistency with the flow via a combination of deterministic and stochastic projections. Our method provides a more principled integration of flow priors within PnP frameworks, achieving high-quality reconstructions with reduced computational cost.
Authors:
Yichen WU, Yujin Oh, Sangjoon Park, Kailong Fan, Yuhan Liu, Zhiyi Shi, Sekeun Kim, Dania Daye, Hana Farzaneh, Wei Liu, Xiang Li, Raul N Uppot, Quanzheng LiAbstract: Recent multi-agent frameworks have shown promise for oncology decision support, yet most assume centralized data access and rely on prompt-based assignment, limiting their applicability in privacy-sensitive clinical settings. We propose Contribution-Aware Medical Multi-Agents (CoMMa), a decentralized LLM-agent framework where specialists operate on partitioned clinical data streams. Unlike prior approaches that share inputs across agents, CoMMa enforces data decentralization to include stronger role specialization and further enhances this via agent-specific finetuning. To enable reliable and interpretable coordination, we introduce a contribution-aware aggregation mechanism that replaces stochastic, narrative-based reasoning with deterministic embedding projections to approximate each agent's marginal utility. This yields explicit credit assignment over agents, providing a stable and interpretable decision pathway aligned with clinical requirements. We evaluate CoMMa on multiple oncology benchmarks, including real-world multidisciplinary tumor board datasets from large academic hospitals in North America and East Asia, as well as public datasets, demonstrating strong performance and generalization across heterogeneous clinical settings.
PaperID: 4449, Poster
Authors: Mohamed Nafea, Sokrat Aldarmini
Abstract: Quantifying bias in training data---independently of any downstream model---is critical for fair machine learning, especially when data collection and model design occur in silos. Yet most fairness research focuses on model-level mitigation, leaving the data side underexplored. We address this gap by deriving multiple information-theoretic measures for auditing unfairness bias in data without requiring access to model predictions. We define our measures over feature sets and use Shapley values to deduce feature-level contributions to global unfairness bias. For some of these measures, we prove that Shapley values moderate redundancy across feature coalitions, matching empirical evidence in prior work. We develop our measures within an axiomatic framework that embeds a fairness notion, such as statistical parity (SP) or equalized odds (EO), into desired and undesired independence properties. We characterize baseline measures that guide those we develop, for example, by showing that a baseline upper bounds the bias of any downstream restricted predictor. Finally, we show that SP- and EO-aligned independence properties can conflict, paralleling Kleinberg et al.'s result on incompatible fairness notions. We validate our measures through an empirical study on real and synthetic datasets, assessing how well they predict the true bias of an extensively tuned neural network trained on the same data, and via a simple feature selection experiment. For synthetic data, we propose a parametric linear structural causal model that enables controlled generation of diverse correlation structures and bias types. Overall, our analysis provides a theoretically and empirically validated guideline for selecting an unfairness measure given a group fairness notion and testable data conditions.
Authors: Phongsakon Mark Konrad, Toygar Tanyel, Serkan Ayvaz
Abstract: Safety evaluations often assume that behavior observed during testing reflects behavior in ordinary use, but fine-tuning can break this assumption. A checkpoint can appear fixed under evaluation-style prompts while the same behavior persists under ordinary-use prompts. Output scores reveal this mismatch but do not locate it. We investigate whether the distinction is encoded in a stable internal site and introduce an approach that fits a paired activation contrast at a path-patching-informed mid-depth window, then modifies the resulting coordinate on held-out prompts. The intervention closes the evaluation-to-deployment gap in ten of twelve model-behavior settings (six of the eight settings with n\geq120 paired questions) across four full-matrix instruction-tuned model instances; a fifth model supports localization and edit-provenance checks, and deployment-framed rates change by at most 6.1pp. The two flat cells, both sycophancy, indicate that a single-coordinate audit is not sufficient when the installed distinction is higher-rank or missed by the depth heuristic. The audit is a diagnostic for fine-tuned checkpoints, not a training-time defense or a guarantee of deployment safety.
PaperID: 4451, Poster
Authors:
Chunjing Xiao, Yunxiao Dai, Yuwan Fu, Yucong Wang, Chong Tang, Fan Zhou, Ying MaAbstract: Unsupervised medical anomaly detection aims to identify images or regions that deviate from the learned normal distribution. Reconstruction-based methods are a dominant paradigm, generating a normal-looking reference for each test image and detecting anomalies via input--reconstruction discrepancy. However, most existing methods follow a one-shot reconstruction-and-comparison pipeline, making anomaly scoring vulnerable to reconstruction failures: abnormal regions may be preserved, while normal anatomical structures may be distorted. Moreover, they often lack coordinated control from structural and semantic normality, which can lead to anatomical inconsistency and semantic drift. We propose S2Agent, a structure- and semantic-guided multi-agent framework that reformulates medical anomaly detection as feedback-driven normality reconstruction. S2Agent decomposes the process into three collaborative agents. The Planner derives sample-specific matched-tree structural priors and normal-only semantic claims from normality knowledge. The Reconstructor generates a normality-oriented reference image under their joint guidance. The Detector verifies the reconstruction in matched-tree and semantic-claim spaces, returning feedback to refine structural weights and semantic constraints for the next round. Through this closed loop, S2Agent iteratively corrects structural deviation and semantic drift, suppresses abnormal or unsupported content, and improves the reliability of anomaly scoring. Experiments on three public medical benchmarks demonstrate the effectiveness of the proposed closed-loop guidance.
Authors: Carlos Garrido, Jorge Calvo-Zaragoza
Abstract: In self-supervised pretraining for Handwritten Text Recognition (HTR), pixel reconstruction methods tend to outperform contrastive ones, a pattern that sits awkwardly with recent evidence that pixel reconstruction yields uninformative features for natural-image classification. We argue this discrepancy is not an accident, but a consequence of how HTR signal is distributed in pixel space. By probing the input distribution directly, we show that HTR's discriminative content is concentrated in the high-variance pixel subspace and is essentially absent from the low-variance one, the inverse of the structure observed for image classification. Under this view, the right SSL family becomes predictable: objectives that preserve high-variance pixel content should transfer best. We test this prediction across six SSL methods spanning three families (pixel-grounded MIM, JEPA-style, contrastive image--image and image--text), under matched encoder, data, and evaluation protocols on six handwriting benchmarks across five languages. Pixel-grounded SSL produces the lowest CER on every benchmark and every probe, exposes per-position character information that other families recover only via the readout, and is the only family that benefits from real-data pretraining. A geometric property of the encoder, its alignment with the high-variance pixel subspace, predicts CER within every SSL method we test. With a pretrained LLM decoder, a frozen pixel-grounded encoder is already competitive with fully fine-tuned supervised baselines, and full fine-tuning beats them on the mean and ranks first or second on every benchmark. These results challenge the prevailing view that pixel reconstruction wastes capacity on irrelevant detail: whether reconstruction is wasted depends on where the discriminative signal lives in the input.
Abstract: We introduce the Thresholding Monte Carlo Tree Search problem, in which, given a tree \mathcalT and a threshold \theta, a player must answer whether the root node value of \mathcalT is at least \theta or not. In the given tree, 'MAX' or 'MIN' is labeled on each internal node, and the value of a 'MAX'-labeled ('MIN'-labeled) internal node is the maximum (minimum) of its child values. The value of a leaf node is the mean reward of an unknown distribution, from which the player can sample rewards. For this problem, we develop a \delta-correct sequential sampling algorithm based on the Track-and-Stop strategy that has asymptotically optimal sample complexity. We show that a ratio-based modification of the D-Tracking arm-pulling strategy leads to a substantial improvement in empirical sample complexity, as well as reducing the per-round computational cost from linear to logarithmic in the number of arms.
Abstract: Long-form video understanding remains challenging for Video Large Language Models (VideoLLMs), as the dense frame sampling introduces massive visual tokens while sparse sampling risks missing critical temporal evidence and leading to LLM hallucination. Existing training-free token reduction methods either treat videos equally as static images or rely on segment-level merging heuristics, which weaken fine-grained spatiotemporal modeling and introduce additional overhead. In this paper, we propose EchoPrune, a lightweight and training-free token pruning method that improves temporal resolution under a fixed LLM-side visual token budget. Our core idea is to interpret redundant video tokens as temporal echoes: if a token is well reconstructed from the previous frame, it is merely a temporally redundant echo; otherwise, it may capture new events, motion, or query-relevant visual evidence. Based on this insight, EchoPrune scores visual tokens by (i) query-guided crossmodal relevance and (ii) temporal reconstruction error, measured by correspondence matching and echo matching across consecutive frames. The selected tokens preserve task-relevant cues and temporal novelty while suppressing predictable redundancy, allowing VideoLLMs to observe more frames without increasing the decoding budget. Extensive experiments on LLaVA-OV, Qwen2.5VL, and Qwen3VL across six video understanding benchmarks show that EchoPrune enables VideoLLMs to process up to \mathbf20× frames under the same token budget, yielding improved performance (\mathbf+8.6%) %and inference speedup (\mathbf5.6× for prefilling) on Qwen2.5VL-7B.
PaperID: 4455, Poster
Authors: Gil Karin, Artemy Bakulin, Nir Yosef
Abstract: Representation learning for single-cell RNA-seq has been evaluated by the identities representations capture — cell types, lineages, perturbations — rather than how faithfully they model the distribution of cell states on the gene expression manifold. As a result, density estimation, rare-state detection, perturbation scoring, and gene-level attribution are bolted onto frozen embeddings as separate models, disconnected from the geometry of the data. We argue that the density model should be the primary object and that representation quality follows from it. Building on LeJEPA, whose SIGReg objective enforces a provably optimal isotropic Gaussian latent, we introduce LeCellModel, which exploits this Gaussian latent to recover data density on the gene expression manifold via the encoder's Jacobian log-volume (JEPA-SCORE). Across several datasets, LeCellModel outperforms established methods on scIB benchmark for representation quality. Furthermore, we show that LeCellModel enables estimation of rare expression states using JEPA-SCORE as a typicality measure – applying it to a controlled cytokine panel, where it identifies stimulated cells across minority fractions and tracks pathway-response magnitude. We introduce gene-level interpretability of density estimates — to our knowledge, the first method for analyzing JEPA-SCORE in feature space. Leveraging this framework, LeCellModel recapitulates the known association of a profibrotic macrophage program with idiopathic pulmonary fibrosis, and discovers a novel CD177^+ activated regulatory T cell subpopulation previously described only in tumor contexts. Both findings emerge directly from the geometry of the learned embedding, without relying on supervised annotations.
PaperID: 4456, Poster
Abstract: Modern deep reinforcement learning algorithms often store past sensory observations in a buffer for later replay and learning. However, this process is biologically implausible; animals do not store raw sensory observations. Instead, they store their imperfect understanding of the world and the events in it. An alternative to storing all past sensory observations is to learn directly from the observations as they come in and store the seemingly important parts of the incoming data. Eligibility trace algorithms in reinforcement learning (RL) sweep back updates for multiple steps, improving efficiency and providing update rules that do not rely on storing large data buffers. These algorithms were common and effectively used in the linear setting, but have not been widely adopted in the deep RL setting. Many questions remain open to make these algorithms more effective, even as basic as how to incorporate vector step-sizes. We first highlight how adaptive vector step-sizes need to be incorporated into the eligibility trace and then introduce a new vector step-size algorithm that is more amenable to the forward-backward view equivalence needed to derive an update with eligibility traces. We show that our new algorithms improve over other streaming algorithms across Mujoco and MinAtar environments, and are comparable to methods with replay buffers such as PPO.
PaperID: 4457, Poster
Authors: Harshit Joshi, Priyank Shethia, Jadelynn Dao, Monica Lam
Abstract: Systematic reviews -- which requires comprehensive evidence collection and synthesis from large document corpora in response to targeted research questions -- are foundational in finance, social sciences, and other technical fields. Manual construction of evidence tables is labor-intensive, and recent LLM-based assistants relying on embedding or keyword based search often fail to meet the coverage standards of systematic reviews. We introduce SLIDERS, a novel LLM-based methodology for systematic reviews, by automatically assembling evidence tables tailored to research questions. In addition to extracting structured data from documents, SLIDERS can extract full-text excerpts that serve as direct evidence or as provenance for structured data. Core to SLIDERS is an automated evidence reconciliation agent that writes code to analyze and reconcile extracted evidence, bringing together information fragmented across documents, resolving inconsistencies across excerpts, and synthesizing overlapping findings into a coherent evidence table. In addition, SLIDERS allows users to ask follow-up questions in natural language to further explore the assembled evidence. We evaluate SLIDERS on three systematic-review-style tasks over large document collections. SLIDERS outperforms the best-performing baseline across benchmarks, remains near 90% accuracy across 6M-11M-token corpora. On two new follow-up analysis benchmarks \system can answer 77.9% and 58.3% followup questions accurately.
Abstract: Machine unlearning (MU) is essential for enforcing the right to be forgotten in machine learning systems. A key challenge of MU is how to reliably audit whether a model has truly forgotten specified training data. Membership Inference Attacks (MIAs) are widely used for unlearned model auditing, where samples that evade membership detection are regarded as successfully forgotten. We show this assumption is fundamentally flawed: failed membership inference does not imply true forgetting. We prove that unlearned samples occupy a fundamentally different positions in the feature space feature space than non-member samples, making this alignment bias unavoidable and unobservable, which leads to systematically optimistic evaluations of unlearning performance. Meanwhile, training shadow models for MIA incurs substantial computational overhead. To address both limitations, we propose Statistical Membership Inference (SMI), a training-free auditing framework that reformulates auditing as estimating the non-member mixture proportion in the unlearned feature distribution. Beyond estimating the forgetting rate, SMI also provides bootstrap reference ranges for quantified auditing reliability. Extensive experiments show that SMI consistently outperforms all MIA-based baselines, with no shadow model training required. Overall, SMI establishes a principled and efficient alternative to MIA-based auditing methods, with both theoretical guarantees and strong empirical performance.
Abstract: Latent action models (LAMs) aim to learn action-like representations from unlabeled videos by compressing frame-to-frame changes. The frames of in-the-wild videos, however, contain not only the agent's own state but such as background clutter. Since the exogenous state introduces changes unrelated to actions, it hinders reliable latent action learning. This paper investigates this problem analytically by extending a linear LAM framework to explicitly model exogenous state. Our analysis reveals two insights: (1) minimizing the standard reconstruction objective produces latent actions that encode exogenous information of future observation; and (2) learning in a representation space that focuses on endogenous components is a key to mitigating the interference of noise. We further show that previously proposed auxiliary objectives, such as action-supervision, provably encourage latent actions to be consistent across exogenous states. These findings are validated through experiments on both linear and nonlinear LAMs, providing a unified theoretical analysis of how exogenous state hinders latent action learning and why common remedies work.
Abstract: Weight-space post-training quantization (PTQ) must choose finite formats, granularities, quantizer families, transformations, and bits before the completed quantized model reveals its output-distribution drift. Existing PTQ methods predict important pieces of this degradation, including reconstruction error, Hessian sensitivity, transformation effects, and downstream loss, but these pieces are usually scored after fixing the quantization geometry or inside separate configuration families. We formulate weight-space PTQ as pre-deployment configuration selection using priced layer-output error. Each admissible layer configuration is treated as an error generator with a deployment cost, which induces a layer-output error covariance \boldsymbol\Sigma_l(\alpha_l), and the full-precision model prices that covariance by downstream curvature, \hat\rho_l(\alpha_l)=\frac12\operatornameTr\left(\hat\mathbfH_l\,\hat\boldsymbol\Sigma_l(\alpha_l)\right). The price follows from full-precision-to-quantized forward KL, whose first-order term cancels at the reference model. It turns reconstruction and diagonal scores into reduced proxies that drop price factors, while finite formats, codebooks, granularities, and equivalent transformations become comparable candidates through the covariances they induce and the costs they pay. A trace reduction then yields a calibration-time price table and a budgeted price-guided selector, making fixed-geometry bit allocation a special case rather than the organizing problem.
Authors: Matteo Leonesi, Francesco Belardinelli, Flavio Corradini, Marco Piangerelli
Abstract: Alignment faking (AF) occurs when an LLM strategically complies with training objectives to avoid value modification, reverting to prior preferences once monitoring is lifted. Current detection methods focus on conversational settings and rely primarily on Chain-of-Thought (CoT) analysis, which provides a reliable signal when strategic reasoning surfaces, but cannot distinguish deception from capability failures if traces are absent or unfaithful. We formalize AF as a composite behavioural event and detect it through observable tool selection, where the LLM selects the safe tool when unmonitored, but switches to the unsafe tool under monitoring that rewards helpfulness over safety, while its reasoning still acknowledges the safe choice. We release a dataset of 108 enterprise IT scenarios spanning Security, Privacy, and Integrity domains under Corruption and Sabotage pressures. Evaluating six frontier LLMs across five independent runs, we find mean AF detection rates between 3.5% and 23.7%, with vulnerability profiles varying by domain and pressure type. These results suggest that susceptibility reflects training methodology rather than capability alone.
PaperID: 4462, Poster
Abstract: We introduce Rapid Simulation Based Inference (RSBI), a diffusion-based variational approach to likelihood-free Bayesian inference that achieves high posterior sample efficiency under large and multi-dimensional prior-to-posterior volumes. RSBI builds on recent advances in Schrödinger Bridge (SB) diffusion sampling to handle multi-modal posteriors, with the likelihood supplied either by a neural ratio estimator or kernal surrogate via a differentiable simulator. Furthermore, a key observation is that an appropriate surrogate likelihood gives a proposal distribution covering most of the generally multi-modal posterior mass, replacing the sequential rounds of standard methods with a one-shot proposal step. Subsequent NRE refinement recovers exact inference under the simulator and original model prior. As a variational method, RSBI does not suffer from prior leakage and implicit target drift commonly observed in standard sequential posterior estimation methods. To improve mode coverage we optionally utilize the well-tempered meta-dynamics framework, which also provides a mechanism through which to encourage exploration of the prior volume. We observe strong performance not only on standard SBI benchmarks, but also as prior volume is scaled up to 2500 times the original, achieving a performance degradation of only ~5% on the two moons benchmark under a highly constrained simulation budget. Additionally, we evaluate RSBI's potential for gravitational wave ring-down posterior estimation, highlighting a real-world use-case where our method excels.
PaperID: 4463, Poster
Authors: Yaokun Wang, Tiantian Xiao, Hongyan Ding, Zhi Yan
Abstract: Probabilistic spiking neural networks (SNNs) model spike trains as temporal point processes and provide a principled framework for learning with latent spikes. Recent differentiable point-process methods enable path-wise variational learning, but their training still relies on full-sequence backpropagation through time (BPTT), leading to memory costs that grow with the temporal horizon. In this paper, we develop an online variational training framework for probabilistic SNNs based on discrete-time spike response model dynamics. By representing synaptic history with finite-dimensional Markovian traces, our method updates both generative and variational parameters without storing the full temporal computation graph. To handle future-dependent credit assignment from recurrent spike histories, we introduce a horizon-R family of online estimators. Theoretically, we show that the truncation-induced bias decays exponentially with R under the SRM kernel-decay condition. Synthetic experiments validate the predicted horizon-dependent behavior and bias decay, while N-MNIST experiments show competitive classification accuracy and avoid the sequence-length-dependent memory growth of BPTT.
Abstract: Generative Engine Optimization (GEO) aims to improve content visibility in AI-generated responses. However, existing methods measure contribution—how much a document influences a response—rather than citation, the mechanism that actually drives traffic back to creators. Also, these methods apply generic rewriting rules uniformly, failing to diagnose why individual document are not cited. This paper introduces a diagnostic approach to GEO that asks why a document fails to be cited and intervenes accordingly. We develop a unified framework comprising: (1) the first taxonomy of citation failure modes spanning different stages of a citation pipeline; (2) AgentGEO, an agentic system that diagnoses failures using this taxonomy, selects targeted repairs from a corresponding tool library, and iterates until citation is achieved; and (3) a document-centric benchmark evaluating whether optimizations generalize across held-out queries. AgentGEO achieves 15-40% relative improvement in citation rates across engines and citation methods, while modifying only 5% of content, compared to 25% for baselines. Our analysis reveals that generic optimization can harm long-tail content and some documents face challenges that optimization alone cannot fully address—findings with implications for equitable visibility in AI-mediated information access.
PaperID: 4465, Poster
Abstract: Large reasoning models separate inference into an explicit thinking phase followed by a final answer phase, yet how the answer phase uses the reasoning trace remains unclear. We present a systematic study of this trace-to-answer transition across multiple LRMs and find that (i) final answers diverge from correct reasoning traces at rates between 6% and 48%, even when traces terminate cleanly at full budget, (ii) naively copying the trace answer underperforms standard decoding in some settings, revealing that the answer phase has independent corrective capacity, and (iii) the same phenomenon extends beyond structured tasks to settings where reasoning traces and final answers can diverge under social-pressure framing (sycophancy). These findings motivate Contrastive Thinking Decoding (CTD), a training-free, single-model decoding method that contrasts answer-phase logits under the primary reasoning trace against those from a deliberately degraded noisy trace. CTD selectively amplifies trace-consistent tokens while preserving the answer phase's corrective capacity, concentrating CTD-internal joint mass on the both-correct cell (T\checkmarkA\checkmark) without inducing trace-correct-to-answer-wrong drift relative to the base distribution. Across math reasoning, science, code, social-pressure resistance (sycophancy), and a broad knowledge benchmark (MMLU), \ctd reduces trace-answer disagreement and improves the accuracy--compute trade-off without parameter updates or auxiliary models at decoding time.
PaperID: 4466, Poster
Authors: Bruno C Sánchez, Carlos Eiras-Franco, Brais Cancela
Abstract: Long-horizon navigation simulations require spatial persistence; revisited physical states must reproduce the same landmarks, rather than merely locally plausible frames. Standard latent-dynamics world models typically evolve a single predictive state that carries the scene layout, appearance, and agent motion simultaneously, which makes revisitation consistency brittle under compounding errors. We study a structured regime for navigation-like tasks in which the scene layout remains static during an episode, a map-like scene observation is available, and the agent motion can be described by a low-dimensional state. For this regime, we propose Space-Aware World Models (SWM). SWM factors world modeling into: (i) a scene latent inferred once per episode from a top-down map; (ii) an explicit, action-conditioned transition model over agent state; and (iii) a conditional renderer from scene code and state to observations. By anchoring scene content in an episode-level representation, SWM is designed to preserve the identity of revisited locations across autoregressive imagined trajectories. Across procedurally generated environments, SWM improves long-horizon fidelity and yields stronger revisitation and geometric consistency, as well as control-relevant fidelity, outperforming standard latent-dynamics baselines by eliminating compounding scene-generation errors. These results suggest that separating the scene from state evolution provides an effective inductive bias for spatially coherent simulations.
PaperID: 4467, Poster
Abstract: Existing VLM-based open-set semi-supervised learning (OSSL) methods primarily rely on coarse-grained class-level semantics, thereby hampering the accurate identification of in-distribution (ID) versus out-of-distribution (OOD) samples. To overcome this issue, we propose a novel memory-driven contrastive embedding enhancement approach by integrating fine-grained visual contrastive cues into class-level textual side. Concretely, we design a memory bank to construct and maintain a diverse collection of the most distinctive fine-grained visual embeddings. Guided by the memory bank, visual information is incorporated into the textual side through cross-modal contrastive learning, enabling more discriminative fine-grained separation between ID and OOD samples. Consequently, the learned OSSL model exhibits improved contrastive discriminability across known and unknown categories. Extensive experiments demonstrate that our method achieves state-of-the-art (SOTA) performance on multiple fine-grained datasets. The code is available at https://anonymous.4open.science/r/CodeForPaper.
PaperID: 4468, Poster
Abstract: Direct pairwise AUC-surrogate training is attractive because it trains toward the ranking metric used at test time, especially under class imbalance. This creates a direct alignment between the training objective and the evaluation metric, but such metric alignment alone does not guarantee that stochastic training remains corrective. We study this gap through Corrective Power: whether pairwise mini-batch updates produce stronger and more reliable improvements for more severe positive-negative ranking errors. Our main observation is that direct pairwise training can enter regions where severe errors remain common but no longer produce useful corrective motion. In such regions, even the mini-batch noise that would normally help SGD move along the steepest severe-error correction direction can become much weaker. In this paper, we seek a theoretical explanation for this by exploring the degenerate property of the U-statistics formed from the pairwise gradients therein. Our results provide a mechanistic explanation for several empirical patterns: (a) squared-hinge loss is often an effective AUC surrogate; (b) cross-entropy warm-up can help models that enter the AUC phase from a poor starting region; and (c) parameter-efficient fine-tuning from a pretrained model can reduce that need by starting closer to a favorable ranking geometry. Experiments on 8 CV tasks and 8 NLP tasks, using ResNet50, DenseNet121, DistilBERT-base, and Qwen2.5-1.5B-Instruct, support these predictions.
Abstract: Wearable motion sensing provides a continuous and scalable window into human behavior and health, making it a natural fit for foundation models, yet its pretraining and scaling principles remain poorly understood. Prior work studies isolated design choices, such as sensor placement or sampling frequency, often under fixed settings and narrow downstream tasks that fail to capture real-world sensing diversity. We introduce Inertia-1, a fully open exploration of wearable motion foundation models. Using massive corpora of accelerometer data from global sources spanning more than 18.2M hours, we build a controlled framework for studying the full lifecycle of wearable motion foundation models, covering data choices such as sensor modality, device placement, sampling rate, window length; model choices such as architectures and model size; and training choices such as pretraining objective and data scale. Extensive evaluations across 15 datasets spanning human activity recognition, freezing-of-gait detection, and disease prediction reveal intriguing findings for building motion foundation models that generalize across tasks and sensing conditions. Collectively, Inertia-1 not only presents state-of-the-art recipes for diverse downstream tasks, but also serves as a comprehensive, practical, and open cookbook for wearable motion representation learning.
PaperID: 4470, Poster
Authors: Kun Wang, Yun Zhu, Pan Zhou, Na Zhao
Abstract: We propose AdaDS, a generalizable framework for prompted metric depth estimation, which estimates high-resolution metric depth from images and arbitrarily degraded low-resolution depth prompts. This setting is commonly studied as depth super-resolution, where existing methods typically regress depth values directly and often exhibit artifacts under severe or unknown depth degradation. In contrast, AdaDS exploits the contraction property of Gaussian smoothing: as noise accumulates in the forward diffusion process, the distributional discrepancy between degraded depth prompts and their high-quality counterparts progressively diminishes, eventually approaching an isotropic Gaussian prior. Leveraging this property, AdaDS estimates refinement uncertainty to adaptively select a starting timestep in the reverse diffusion trajectory, and subsequently injects tailored noise to place the intermediate sample in a high-probability region of the target posterior distribution. This strategy enables the generative prior of a pre-trained diffusion model to dominate the estimation process even when upstream prompt refinements are imperfect. Extensive experiments on real-world and synthetic benchmarks demonstrate AdaDS's superior zero-shot generalization and robustness to diverse degradation patterns compared with state-of-the-art methods.
Abstract: Flow-based generative models have emerged as an efficient alternative to diffusion models. We study the sample complexity of learning a target distribution in Wasserstein distance via flow matching, which learns a velocity field transporting a base distribution to the target. Under standard assumptions on the network architecture and data distribution, we show that any squared-loss flow-matching variant including rectified flow with bounded derivatives achieves \widetildeO(\varepsilon^-2) sample complexity with the Wasserstein guarantee taking the form W_2 \le K(\varepsilon + C\sqrt\varepsilon_app), where \varepsilon_app is the approximation error of the network class. This improves upon existing \widetildeO(\varepsilon^-4) guarantees and does not require the Polyak \Lojasiewicz condition assumed in prior work. Instead, we exploit a structural property of the squared-loss objective - it induces an approximate Bernstein condition that directly ties the variance of the excess loss to the excess risk up to additive approximation-error. This condition enables a localized Rademacher complexity analysis yielding fast O(1/n) rates. The flow-matching schedule enters only through the Lipschitz constant of the pointwise loss, which we characterize explicitly for standard schedules.
PaperID: 4472, Poster
Authors:
Hohyun Kim, Hyesung Kim, Min-hwan Oh, Seunggeun LeeAbstract: We introduce SNACK, a string representation for probabilistic graph generation that enables sequence models to generate and model graph-structured data. SNACK linearizes graphs by encoding the nonzero entries of the lower-triangular adjacency matrix, establishing a bijective correspondence between valid sequences and ordered adjacency matrices. This design supports invalid logit masking for enforcing structural constraints, including chemical valence and ordering constraints, and enables principled string-based GFlowNet training by explicitly accounting for node-ordering symmetries. Across standard graph-generation tasks and molecular distribution-learning benchmarks, SNACK achieves competitive or state-of-the-art sample quality, while substantially improving the throughput of GFlowNet training relative to graph-based generation. Probing analyses further show that models trained on SNACK encode meaningful local and global graph structure despite receiving only sequential supervision. These results position SNACK as a practical framework for probabilistic graph generation with deep sequence models.
PaperID: 4473, Poster
Authors:
Junhong Wang, Changxing Ding, Yusha Peng, Wentao Tan, Dongliang LiaoAbstract: Text-to-image person re-identification (TIReID) aims to retrieve images of one pedestrian using natural-language descriptions. It remains fundamentally challenging due to significant inter-modal semantic gap between sparse textual cues and fine-grained visual appearance. Existing methods mainly perform instance-wise text-image matching during inference, overlooking intra-modal visual associations among top-ranked gallery candidates, where multiple images may depict the same identity. To exploit such associations, we propose Reranking with Intra-modal Visual Association (RIVA), a plug-and-play reranking framework that leverages intra-modal visual association among retrieved candidates to guide multimodal large language model (MLLM)-based reranking. RIVA consists of two complementary components. First, we propose an MLLM-based Short-lIst Visual Association (SIVA) method, which is trained with a three-stage, dual-task curriculum to learn identity-level image-image association and then use it as visual context for text-image matching. Second, we develop a long-list visual association approach based on clustering, which clusters the top-ranked candidates in the feature space of pedestrian images, uses SIVA to examine the most likely matched cluster, and calibrates the image-text similarity scores accordingly. Extensive experiments on three popular TIReID benchmarks demonstrate that RIVA consistently improves retrieval accuracy in multiple evaluation settings. The code will be released.
Abstract: Efficient LLM inference scheduling is crucial for the user experiences. However, LLM inferences exhibit remarkable demand uncertainty (with unknown output length beforehand) and hybridity (being both compute and memory intensive). Existing LLM schedulers rely on simple heuristics or focus purely on compute resource, suffering suboptimal performance. In this work, we propose SageSched, an efficient LLM scheduler that properly handles demand uncertainty and hybridity of inference workloads. SageSched combines prompt contents with the past inference results to predict output-length distribution in a light-weight and also accurate manner. Meanwhile, it models the true service cost of an inference request with both compute and memory aspects considered. Finally, SageSched employs a uncertainty-aware scheduling policy that can yield the best overall efficiency given the request cost distributions. Testbed experiments over diverse setups confirm that SageSched can attain an efficiency improvement of over 28.7%.
PaperID: 4475, Poster
Abstract: Infrared (IR) spectral unmixing recovers individual component spectra from a measured mixture, enabling chemical identification. Existing methods either assume linear superposition of components, which fails for liquid-phase mixtures where molecular interactions alter spectral shapes, or use deep networks that entangle all components in a shared representation. We introduce SLOT-IR, a slot-based architecture that decomposes mixture spectra into separate per-component embeddings and decodes each independently into a predicted spectrum. A two-stage training procedure first learns to map a single molecule from gas-phase spectra to liquid-phase, then trains the full unmixing model on mixtures. On the benchmark of annotated infrared spectra, SLOT-IR outperforms all baselines on both binary and ternary mixture identification. To verify that the model captures genuine chemical structure rather than statistical shortcuts, we apply a post-hoc BatchTopK Sparse Autoencoder (SAE) to frozen slot embeddings and test whether recovered features correspond to known functional groups. Statistical and causal analyses confirm model alignment with established chemistry.
PaperID: 4476, Poster
Abstract: As large language models (LLMs) grow more powerful, their scale makes them costly to serve. Depth pruning offers a practical route to efficiency by directly reducing memory and latency without hardware-specific support. However, we observe that transformer layers exhibit markedly different task-dependent layer contributions, whereas existing depth-pruning methods ignore this heterogeneity and thus tend to optimize for a narrow domain or benchmark rather than preserve broad capability. This suggests that LLM depth pruning should be formulated as a multi-objective optimization problem rather than a single-score ranking problem. We propose , an interpretable multi-objective evolutionary search framework for LLM depth pruning. SliMOO first uses single-layer removal probing to construct a task-layer importance map, which provides an interpretable view of task-shared and task-specific layers and serves as a search prior. It then performs a Pareto-aware evolutionary search for pruning masks, where candidates are generated under the guidance of the task-layer prior, scored by their deviation from the dense model on each task, and retained through NSGA-II environmental selection. Across the Llama-3.1 and Qwen3 model families, SliMOO consistently achieves stronger multi-task trade-offs than rule-based, greedy and scalarized baselines; for example, on Llama-3.1-70B at 25% sparsity, it improves the math-domain average from 24.33% to 34.40% and the code-domain average from 35.03% to 39.14%.
Abstract: State-of-the-art vision-language models (VLMs) score impressively on video benchmarks yet stumble on basic visual reasoning tasks involving spatial relations, navigation, and object selection that a preschooler solves without effort. We hypothesize that the explicit pedagogical structure, specifically the context-question-pause-answer cycles embedded in children's educational video, provides naturally co-aligned reasoning traces: temporally synchronized visual cues, questions, and answers that emerge only from deliberate pedagogical authoring and cannot be practically reconstructed through manual annotation at scale. To test this, we introduce SoSVQA (Structure over Scale Visual Question Answering), a unified benchmark of 10K question-answer pairs automatically extracted from Dora the Explorer (DoraVQA) and Mickey Mouse Clubhouse (ClubHVQA) with precise timestamp alignment, and fine-tune Qwen2-VL and Qwen3-VL using Group Relative Policy Optimization (GRPO) to leverage the clear correctness signals and structured reasoning traces inherent in educational content. Despite training on just 10K QA pairs from 78 hours of children's television, orders of magnitude less data than GPT and Gemini, our approach delivers generalizable performance gains for Qwen-based VLMs, yielding consistent improvements on NExT-QA (+19.7), Video-MME (+10.6), and MotionBench (+4.9), matching the performance of leading proprietary systems and demonstrating that content structure can compensate for content scale.
PaperID: 4478, Poster
Abstract: Projection-based optimization offers a gradient-free alternative to conventional gradient-based methods, but existing frameworks suffer from high computational costs and poor scaling with depth. We introduce \mathcalPTorch, a projection-based framework that builds directly on top of PyTorch's autograd engine, eliminating the overhead of prior approaches. We further improve performance through nonlinear relaxation and memory-efficient projections. Together, these contributions make the algorithm competitive with gradient-based methods for shallow MLPs on MNIST and CIFAR-10. Finally, we prove a vanishing target theorem showing that the layer-wise learning signal decays with increased depth, serving as an analogue of the vanishing gradients problem for projection-based methods.
PaperID: 4479, Poster
Abstract: Generating realistic human trajectories is essential for mobility simulation but remains challenging when they must generalize across cities and respect road network constraints. In practice, existing methods often suffer from three key limitations: geographic overfitting, length-rigid generation, and off-road drift when generating GPS coordinates directly. To address these issues, we propose XTraj, a transferable coarse-to-fine autoregressive framework for road-consistent trajectory generation. Specifically, XTraj first learns transferable road-segment representations by integrating road geometry, POI context, historical traffic intensity, and graph topology. Based on these representations, it then autoregressively generates variable-length road-segment routes under road-connectivity constraints. Finally, instead of directly regressing latitude-longitude coordinates, XTraj further refines each generated route in a route-progress space by predicting monotonic progress increments along valid road geometry, from which GPS points are recovered by interpolation. This route-aligned design naturally supports variable-length GPS generation and guarantees road-geometry consistency by construction. Experiments on two real-world vehicle trajectory datasets show that XTraj improves spatial fidelity over competitive baselines, transfers effectively to unseen cities in a zero-shot setting, and ensures road-geometry consistency by construction without any post-hoc map matching. The implementation is provided in https://anonymous.4open.science/r/XTraj-EBE7/.
Abstract: This paper studies the regret analysis for parallel Gaussian process (GP) bandit optimization. The known regret upper bounds for the widely used GP batched upper confidence bound and GP batched Thompson sampling (GP-BTS) suffer from a multiplicative factor with respect to the batch size Q. To avoid this degradation, existing analyses require a polynomial number of uncertainty sampling (US) for Q at the beginning of optimization. However, this initial US phase is often ineffective in practice. This paper shows that the regret upper bound without the multiplicative factor on Q can be achieved without the initial US phase, using GP-BTS as an example. Furthermore, we show much better regret upper bounds in the noiseless setting than in the noisy setting, as in the sequential GP bandit setting.
PaperID: 4481, Poster
Abstract: Noisy labels can cause deep neural networks to overfit incorrect annotations, leading to biased predictions and poor generalization. A common remedy is to select clean samples after a short warm-up phase, under the assumption that the model begins to learn class-specific patterns from the early training stage. However, this assumption often fails under class imbalance, as the majority classes can hinder the model from learning distinctive features of minority classes. In this paper, we propose a readiness-aware sample selection strategy that identifies clean samples only from classes whose distinctive features have been sufficiently learned and are thus considered ready for selection. We further introduce a novel negative learning scheme to enhance class separability by discouraging confusion with the most similar incorrect classes. The proposed method is supported by theoretical analysis and demonstrates outstanding performance on both benchmark and real-world datasets, showing improved robustness and generalization under noisy and class-imbalanced conditions.
Abstract: Transitioning Multimodal Large Language Models (MLLMs) from offline to online streaming video understanding is essential for continuous perception. However, existing methods lack flexible adaptivity, leading to irreversible detail loss and context fragmentation. To resolve this, we propose FreshMem, a Frequency-Space Hybrid Memory network inspired by the brain's logarithmic perception and memory consolidation. FreshMem reconciles short-term fidelity with long-term coherence through two synergistic modules: Multi-scale Frequency Memory (MFM), which projects overflowing frames into representative frequency coefficients, complementing by residual details to reconstruct a global historical “gist”; and Space Thumbnail Memory (STM), which discretizes the continuous stream into episodic clusters by employing an adaptive compression strategy to distill them into high-density space thumbnails. Extensive experiments show that FreshMem significantly boosts the Qwen2-VL baseline, yielding gains of 5.20%, 4.52%, and 2.34% on StreamingBench, OV-Bench, and OVO-Bench, respectively. Besides its plug-and-play design, following low-cost component-wise fine-tuning, FreshMem outperforms current fully fine-tuned methods and achieves a state-of-the-art performance of 82.31% on StreamingBench, exceeding the baseline by 13.31%, offering a highly efficient paradigm for long-horizon streaming video understanding.
PaperID: 4483, Poster
Abstract: Phase transitions occur when a physical system undergoes a dramatic transformation as it crosses a transition. We study whether regression methods trained only in one phase can predict the physical properties that characterize the unobserved phase. Our test sets consist of the analytically tractable Heisenberg model and angle-resolved photoemission spectroscopy measurements of a venerable high transition temperature superconductor. We find that polynomial regression and kernel methods recover the unobserved phase structure, while standard neural networks fail. To bound the error, we adopt the generalized eigenvalue problem (GEVP) into an upper bound on the worst-case transfer coefficient, defined as the ratio of mean squared error (MSE) loss on the extrapolation region to the MSE loss on the training region. Empirically, we find that our GEVP bound stays approximately constant with respect to the empirical transfer coefficient across training-set sizes and distance between the training and extrapolation region for both datasets, providing a principled framework for studying extrapolation feasibility.
PaperID: 4484, Poster
Authors: Frithjof Gressmann, Ngoc H Pham, Lawrence Rauchwerger
Abstract: Inferring the latent processes that generate observable neural activity is a central challenge in neuroscience, with direct implications for brain-machine interfaces and neural engineering. Existing inverse approaches typically summarize the recording into low-dimensional features or learn black-box latents that lack a direct biophysical interpretation. A promising alternative is to fit a differentiable biophysical simulator end-to-end against the full extracellular signal, using direct and scalable gradient optimization. In practice, however, recovering a per-neuron input via backpropagation through neuronal dynamics is difficult because the loss landscape is nearly flat in the subthreshold regime and jumps sharply at spike threshold, leaving gradient descent without a useful signal. We observe that per-neuron spike times, routinely available from spike sorting, expose discrete millisecond-scale anchors at which the loss does carry information. Gradients from these anchors flow back in time through the differentiable simulator and shape the subthreshold drive that produced each spike. Building on this, we present a fully differentiable pipeline coupling biophysical membrane dynamics to a volume conductor model of the multi-electrode array. For each neuron in a recorded population, the method jointly recovers a time-varying latent input current and a probe-relative position consistent with both the extracellular trace and the observed spike times. No explicit likelihood or posterior estimation is required, and every fitted latent corresponds to a named biophysical quantity. On paired patch-clamp / Neuropixels SPE-1 data, the model localizes the patched neuron to within 40 \mu m on average across cells, while reproducing observed spatio-temporal activity patterns under realistic noise. These results position scalable differentiable biophysical simulation as a practical route to mechanistic, gradient-based analysis of high-density neural recordings.
PaperID: 4485, Poster
Abstract: Test-time compute scaling offers a promising path for improving foundation models beyond training-time data and parameters. For Vision-Language-Action (VLA) models, however, applying this paradigm through online reinforcement learning remains challenging due to two coupled limitations. Candidate chunks are often biased toward the pretrained policy's high-likelihood regions, while sparse outcome-level feedback lacks process grounding, making reward signals weakly discriminative and exploration insufficiently guided. To address these challenges, we propose EXPLORE, a test-time reinforcement learning framework that coordinates Exploitation and Exploration for action-chunk generation and drives policy optimization with physics-grounded process rewards. Through this joint design, EXPLORE broadens the effective action space without abandoning the pretrained prior and derives faithful process-level rewards from dense physical interaction feedback to guide online policy improvement. Experiments on SimplerEnv-WidowX, DexMG, and RoboCasa validate the effectiveness of EXPLORE, with consistent improvements across autoregressive and diffusion-based VLA backbones.
Abstract: Humans solve complex problems by flexibly shifting among reasoning modes, often without explicit deliberation: they plan, execute, revise intermediate goals, resolve ambiguity through associative judgment, and apply formal procedures to well-specified subproblems. Current LLM agents lack this flexibility, as their scaffolds hard-code such reasoning decisions in advance through fixed inference patterns. These scaffolds are effective when their prescribed structure matches the task, but brittle when solving the task requires adapting the structure of reasoning itself. We introduce Deep Reasoning -- an inference-time approach for constructing task-specific scaffolds through structured meta-reasoning. Deep Reasoning uses a formal language that represents meta-reasoning as executable decompositions over associative inference, formal computation, and recursive subproblem solving, enabling decomposition principles to be encoded as in-context examples that guide test-time scaffold construction. We instantiate this approach in a general-purpose agent (DOLORES) that distributes complex tasks across smaller, more controlled reasoning threads while preserving dependencies among subproblems. We evaluate DOLORES against state-of-the-art scaffolding methods across four hard benchmarks: grounded multi-hop reasoning, synthetic long-chain question answering, long-context aggregation, and deep research-style information seeking. DOLORES outperforms all evaluated scaffolds across four benchmarks, three model sizes, and two model families, improving over the strongest evaluated scaffold baseline by 24.8% on average, including methods tailored to individual benchmark families. Trace and token analyses suggest that while baseline scaffolds fail by overloading individual LLM calls, DOLORES succeeds by distributing cognition across structured, lower-load reasoning threads, thereby reducing premature termination and hallucination. This advantage can even bridge the scaling gap, with an 8B version surpassing all evaluated 32B baselines from the same family in more than half the settings. These results point toward future agentic systems that treat scaffolding as adaptive reasoning, constructing the structure each task requires just-in-time.
Authors:
Jiyuan Wang, Huan Ouyang, Chunyu Lin, Dewen Fan, Boheng Zhang, Tingting Gao, Fan Yang, Jia Sun, Zijun Li, Yongrui Heng, Huaiqing Wang, Zhenlong Yuan, Yiyang Fan, Honglie Wang, Fei Zuo, Haonan fan, Jiuzhou Lin, Guosheng LinAbstract: Recently, video generation models have achieved remarkable progress in global visual fidelity. However, when synthesizing complex dynamic content, they frequently produce . Existing video reward models primarily emphasize overall visual quality and text alignment; yet as generation quality improves, these coarse-grained defects become less frequent, while the sparse structural, temporal, and physical anomalies become an increasingly important bottleneck. Therefore, in this paper, we propose Concentrate and Concentrate (CaC), a coarse-to-fine anomaly reward model based on Vision-Language Models. During inference, it first conducts a global temporal scan to anchor anomalous time windows, then performs fine-grained spatial grounding within the localized interval, and finally derives robust judgments via structured spatiotemporal Chain-of-Thought reasoning. To equip the model with these capabilities, we construct the first large-scale generated video anomaly dataset with per-frame bounding-box annotations, temporal anomaly windows, and fine-grained attribution labels. Building on this dataset, we design a three-stage progressive training paradigm. The model initially learns spatial and temporal anchoring through single- and multi-frame supervised fine-tuning, and then is optimized by a reinforcement learning strategy based on two-turn Group Relative Policy Optimization (GRPO). Beyond conventional accuracy rewards, we introduce Temporal and Spatial IoU rewards to supervise the intermediate localization process, effectively guiding the model toward more grounded and interpretable spatiotemporal reasoning. Extensive experiments demonstrate that CaC can stably ``concentrate'' on subtle anomalies, achieving a 25.7% accuracy improvement on fine-grained anomaly benchmarks and, when used as a reward signal, CaC reduces generated-video anomalies by 11.7% while improving overall video quality.
PaperID: 4488, Poster
Authors: Deena Francis
Abstract: Granger Component Analysis (GCA) discovers latent components of multivariate time series that exhibit directed temporal dependence, but is limited to linear representations. We introduce Kernel Granger Component Analysis (KGCA), a nonlinear extension that learns directed latent components in a reproducing kernel Hilbert space. KGCA optimizes a ridge-regularized Granger predictive objective using explicit envelope-theorem gradients, and incorporates a time-reversal criterion to resolve directional identifiability. We also describe a scalable random Fourier feature (RFF) variant that approximates the kernel map explicitly and avoids forming the full Gram matrix. To avoid spurious directionality from post hoc component selection, we use a restricted evaluation protocol that measures directed structure intrinsic to the learned representation. Empirically, KGCA and RFF-KGCA recover reliable nonlinear directed components, exhibit positive directionality gaps, and avoid spurious directionality under independent-process controls. These results show that kernelizing GCA enables nonlinear directed component discovery while preserving the interpretability and explicit optimization structure of the original framework.
PaperID: 4489, Poster
Abstract: Verifiable reward training has improved mathematical and coding reasoning, but these domains capture only part of step-by-step decision making. Many real-world tasks require finding a high-value feasible plan among many valid alternatives. We introduce \textOPT\star, a scalable family of optimization-style tasks for training and evaluating LLM step-by-step optimization-like reasoning along a complexity axis: each task provides a feasibility checker and evaluator, while a complexity parameter expands the search space without requiring new human labels. This motivates studying these tasks in two regimes: (i) solver-guided online policy optimization, which uses a solver as a value oracle for partial states and applies rank-based reward shaping to reinforce better next steps, and (ii) search-based offline RL when such solvers are unavailable. Theoretically, we relate success in large search spaces to the information a reasoner extracts per unit of search budget. Empirically, we ablate the ingredients that make search efficient on \textOPT\star and show that training on \textOPT\star improves step-by-step optimization-like reasoning.
PaperID: 4490, Poster
Abstract: We present UltraVoxelGS, a 3D ultrasound reconstruction method that achieves both fast reconstruction and physically faithful rendering, bringing live spatial feedback during 2D ultrasound scanning within reach. This is made possible by a feed-forward design that predicts a complete Gaussian scene representation from posed ultrasound images directly, entirely bypassing the per-scene optimization bottleneck that has confined prior methods to offline use. Realizing such feed-forward prediction in this setting is non-trivial: unlike natural images, ultrasound slices share at most one-dimensional intersections and exhibit view-dependent appearance due to acoustic propagation and attenuation, violating the dense overlap and photometric consistency assumed by existing feed-forward Gaussian methods. To achieve fast reconstruction, we introduce a feed-forward voxel-to-Gaussian pass that decouples representation size from input count by lifting posed slices into a fixed-size voxel support and predicting Gaussian parameters directly. For faithful rendering, we propose an ultrasound-aware appearance rendering strategy that factorizes reflection from cumulative attenuation during training and exploits appearance continuity among spatially adjacent slices for efficient Render-time Appearance Adaptation. On both real-world and simulated datasets, UltraVoxelGS achieves the best PSNR/SSIM compared with per-scene optimization baselines, while reducing inference from 10--20 minutes to approximately 10 seconds, over 60× faster, and scaling to sequences exceeding 1,000 frames.
PaperID: 4491, Poster
Abstract: Robust visual recognition under degraded imaging conditions is critical for real-world vision systems.Task-oriented image restoration addresses this problem by enhancing degraded inputs before recognition, but it must preserve natural visual appearance without disrupting the evidence used by downstream recognizers.Existing optimization strategies struggle to satisfy these requirements simultaneously.Restoration driven mainly by visual quality may weaken task-critical cues or introduce texture shifts that are visually plausible but poorly matched to the recognizer.Directly tuning restoration networks with recognition loss is also insufficient, as the restoration model may exploit recognizer-specific shortcut cues and produce visually unnatural artifacts while improving task scores.We propose Causal Mediated Unrolling (CaMeU), an optimization-based framework for task-oriented image restoration with frozen recognition models.Starting from a joint optimization objective, the proposed framework derives a Half-Quadratic Splitting formulation that separates restoration into alternating task-driven and prior-guided updates.The task-driven update introduces an explicit mediator for front-door guided optimization, reducing shortcut bias induced by the recognizer.The prior-guided update performs manifold projection, pulling the intermediate result back toward the natural-image manifold to preserve visual quality.Through this mediated alternating process, the restored image is progressively optimized for both recognition performance and visual quality.Experiments on VOC Haze, VOC Dark, and degraded CUB-200-2011 show that CaMeU consistently improves downstream recognition while reducing visual artifacts.
Authors: Jacob Dang, Brian Y Xie, Omar G. Younis
Abstract: Recent work on subliminal learning has shown that language models can transmit semantic traits through data with no apparent connection to those traits. Whether this phenomenon extends to agentic systems, where policies are learned from trajectories rather than static text, remains an open question with critical safety implications. We present the first empirical evidence that unsafe agent behaviors transfer subliminally through model distillation, demonstrated across two complementary settings. In the primary setting, we construct a teacher agent with a deletion bias, a tendency to perform destructive file-system actions through an API-style tool interface, and distill it into a student using only trajectories from ostensibly safe tasks, with all explicit deletion keywords filtered. In the secondary setting, we replicate the threat model in a native Bash environment, replacing API calls with shell commands and operationalizing the bias as a preference for chmod over semantically equivalent alternatives (e.g., chown, setfacl) when issuing the first permission-related command. Despite thorough sanitation, students inherit measurable biases in both settings: the API student's deletion rate reaches 100% (vs. 5% baseline) under homogeneous distillation, while the Bash student's chmod-first rate reaches 30–55% (vs. 0–10% baseline), with the strongest transfer observed in large-to-small distillation. Evaluations on TerminalBench further show that subliminal transfer persists in complex, multi-step tasks, indicating that behavioral biases are encoded implicitly in trajectory dynamics, independent of tool interface or keyword filtering. Explicit data sanitation is therefore insufficient as a defense; mitigating behavioral bias transfer in agentic systems will require fundamentally new strategies.
PaperID: 4493, Poster
Abstract: We present StereoSplat, a feed-forward 3D Gaussian Splatting architecture designed explicitly for stereo videos to achieve reliable, metric-scale novel-view synthesis. Unlike monocular approaches that neglect the nature of binocular pairs, StereoSplat adopts a native stereo architecture that extracts dense multi-scale latent features and disparity from a foundation stereo backbone and feeds them into a dedicated Gaussian decoder. Our decoder utilizes cross-view feature fusion and gated spatial refinement to directly translate these rich representations into per-pixel 3D Gaussians. To resolve the high redundancy of per-pixel predictions across overlapping views, we propose a learnable depth-adaptive aggregation mechanism that clusters and aggregates primitives in inverse-depth space while preserving fine-grained details. To support robust training and generalization, we introduce XRStereo, a large-scale synthetic dataset of stereo video featuring various rig configurations tailored to match the physical hardware of real-world platforms. Trained exclusively on synthetic data and evaluated across challenging benchmarks, StereoSplat yields substantial improvements in the depth accuracy of rendered novel views over established baselines while maintaining high-fidelity novel view synthesis. By overcoming prior geometric limitations, our framework enables robust zero-shot transfer to real-world domains, offering a scalable geometric foundation for spatial computing in mixed reality, autonomous driving, and robotics.
PaperID: 4494, Poster
Authors: Zihao Chen
Abstract: Anchored fixed point and monotone equation methods, including Halpern iteration, extra anchored gradient, and their relatives, add a vanishing pull toward a reference point to obtain last-iterate guarantees. Existing anchored variants often achieve sharp last-iterate guarantees, but from the update-level perspective the placement of the anchor can be algorithm-specific and conceptually opaque. We show that anchoring admits a single operator-side construction: regularize the operator queried by the base method with a vanishing Tikhonov term, then run the unmodified base method. Applied to the Picard iteration, this recipe reproduces the Halpern iteration; applied to the forward step, extragradient, and Popov methods, it yields three variants whose anchor placements inherit the base method's query pattern. The four analyses share a residual recurrence, recovering the (O(1/k)) Halpern residual-norm convergence rate, giving (O(1/\sqrtk)) for the regularized forward step, and giving (O(1/k)) for the regularized extragradient and Popov variants in the unconstrained monotone Lipschitz setting.
PaperID: 4495, Poster
Authors:
Aastha A Verma, Vishakha Agarwal, Sahil Manchanda, Sayan RanuAbstract: Deployed Graph Neural Networks are increasingly subject to post-deployment interventions: a sensitive attribute is flagged for removal, or a spurious correlation is discovered after the model is already in production. In such scenarios, retraining is often infeasible due to a variety of reasons including scale of modern graphs, inaccessible training data, or strict privacy mandates. The challenge is further compounded in graph-structured data, where attribute dependencies propagate through message-passing, spreading unwanted correlations across the entire network. We propose a framework for surgical, post-hoc attribute invariance in GNNs via gradient-free model editing. Rather than retraining or fine-tuning, our method formulates the edit as a subspace-constrained, closed-form optimization: we identify structurally influential nodes that exhibit high attribute sensitivity, isolate the low-rank manifold in activation space where sensitive variation is concentrated, and solve directly for a weight update that decouples the model's decision logic from the target attribute. The update requires no access to original training labels, introduces no additional parameters, and is computed in a single pass. Empirical results demonstrate that our method enforces attribute invariance while preserving predictive utility, while offering 3 orders of magnitude speed-up over existing GNN editing approaches.
Abstract: Single-step retrieval-augmented generation (RAG) provides an efficient way to incorporate external information for simple question answering tasks but struggles with complex questions. Agentic RAG extends this paradigm by replacing single-step retrieval with a multi-step process, in which the large language model (LLM) acts as a search agent that generates intermediate thoughts and subqueries to iteratively interact with the retrieval system. This iterative process incurs substantial latency due to the autoregressive generation of lengthy thoughts and subqueries. To address this limitation, we propose LatentRAG, a novel framework that shifts both reasoning and retrieval from discrete language space to continuous latent space. Unlike existing explicit methods that generate natural language thoughts or subqueries token-by-token, LatentRAG produces latent tokens for thoughts and subqueries directly from the hidden states in a single forward pass. We align LLMs with dense retrieval models in the latent space, enabling retrieval over latent subquery tokens and supporting end-to-end joint optimization. To improve transparency and encourage semantically meaningful latent representations, we incorporate a parallel latent decoding mechanism that translates latent tokens back into natural language. Extensive experiments on seven benchmark datasets show that LatentRAG achieves performance comparable to explicit agentic RAG methods while reducing inference latency by approximately 90%, substantially narrowing the latency gap with traditional single-step RAG.
PaperID: 4497, Poster
Authors:
Dongxu Yue, Weixuan Jin, Yuhang Yu, Bo Li, Qi Wen, Xinrui Chen, Jinwei Chen, Tianyi Zheng, Hao Zhang, Zhihai He, Chun YuanAbstract: While diffusion models have revolutionized image generation and editing, achieving physically realistic object removal remains a challenge. Physically realistic object removal posits that if an object ceases to exist, all its directly associated transient physical effects must also vanish. However, existing methods either confine the removal strictly to the masked region or limit "side effects" to simple shadows and reflections, ignoring complex interactions like caustics, emitted light, and volumetric media. Furthermore, they struggle with transparent objects, destroying the underlying background semantics. To bridge these gaps, we formalize the task of physically realistic object removal and propose PhysRemover, a unified framework addressing both the elimination of objects with their physical effects and the handling of transparent materials. Specifically, we incorporate an In-Context Contrastive Guidance mechanism guided by both masked and unmasked image, coupled with a Learnable Removal Trigger. This design effectively resolves the conflict between erasing the object and its physical side effects, and preserving the underlying background. Furthermore, we introduce PhyTrace, a large-scale image dataset capturing diverse physical effects, along with two benchmarks tailored for the proposed task of physically realistic object removal. Additionally, we develop a novel metric, PhysRM, for evaluating object removal quality, establishing a new standard for physically realistic image editing. Extensive experiments demonstrate that PhysRemover effectively eliminates objects and their physical side effects while faithfully preserving the background when removing transparent objects, outperforming existing methods in these challenging settings.
Abstract: Probabilistic weather forecasting requires not only accurate trajectories, but calibrated distributions over plausible atmospheric futures. Recent data-driven systems have achieved remarkable deterministic skill, and diffusion-based ensemble forecasters have substantially improved sample realism and uncertainty quantification. However, their inference cost scales with forecast horizon, ensemble size, and the number of denoising steps required for each transition, making large operational ensembles expensive. To address this, we present Tyche, a one-step conditional flow model for efficient probabilistic weather forecasting. Tyche models the conditional forecast distribution with a destination-aware average-velocity flow that maps Gaussian noise directly to future weather states in a single function evaluation (1-NFE). To make this one-step transport learnable in high-dimensional geophysical fields, we derive a JVP-regularized rectification objective that enforces temporal self-consistency across source and destination flow timesteps without explicitly forming Jacobians. The transport field is parameterized by an isotropic Swin-style transformer that preserves fine-scale spatial structure while remaining scalable on global grids. To improve ensemble reliability under autoregressive forecasting, we further introduce a rollout-based finetuning stage with curriculum CRPS calibration supervision. Experiments on ERA5 at 1.5^\circ and 6-hour resolution show that our Tyche, using merely a single NFE, matches or exceeds the forecast skill and calibration of state-of-the-art multi-step generative baselines and the operational ECMWF IFS ensemble. Codes are available.
PaperID: 4499, Poster
Abstract: Low-Rank Adaptation (LoRA) and its weight-decomposed variants dominate parameter-efficient fine-tuning, yet their directional update mechanisms remain geometrically mismatched to the normalized parameterization they impose. In existing decoupled PEFT methods, low-rank directional updates are computed in unconstrained Euclidean space and only then projected back onto a spherical manifold, which can introduce capacity redundancy and directional gradient distortion. To address these limitations, we propose Geodesic Orthogonal Low-Rank Adaptation (GO-LoRA), a geometry-aware reparameterization framework for decoupled PEFT. GO-LoRA first applies orthogonal tangent projection to remove components parallel to the pretrained base direction, ensuring that the low-rank directional budget is devoted entirely to tangent-space motion. It then uses a geodesic flow based on the Riemannian Exponential Map to construct manifold-consistent directional updates on the hypersphere, improving geometric fidelity relative to Euclidean normalization while retaining standard optimization over the underlying low-rank parameters. Extensive experiments across language, vision, code generation, and multimodal reasoning tasks show that GO-LoRA consistently outperforms strong LoRA baselines while introducing no additional inference-time overhead after offline merging. The code is available at https://anonymous.4open.science/r/go-lora.
PaperID: 4500, Poster
Abstract: Effective long-context modeling is not merely about retaining more of the past, but about preserving the information that may prove relevant later. Test-time training (TTT) is an appealing approach that performs online parameter updates for long-context modeling, yet existing TTT methods only optimize either reconstruction or online adaptation objectives without considering the future utility of retained information. In this work, we propose Test-Time Context Distillation (TTCD), a TTT framework that explicitly optimizes the allocation of limited memory capacity for future use. Specifically, TTCD uses a long-window teacher to supervise the fast weights of a short-window student, where the hidden-state discrepancy between them offers a dense, self-supervised signal guiding the model to memorize the contextual information crucial for future tasks. We focus on an in-place variant: In-Place TTCD (IP-TTCD), which uses the existing MLP parameters as the fast weights. Experiments on long-context language modeling tasks show IP-TTCD consistently outperforms DeltaNet, Gated DeltaNet, sliding-window attention, and TTT when pre-trained from scratch. Furthermore, IP-TTCD allows pre-trained transformer models to adapt their parameters during inference through continual pre-training, gaining long-context capabilities with only a lightweight architectural augmentation.
PaperID: 4501, Poster
Abstract: Robotic and sensing systems operating in challenging environments often receive weak, noisy, and incomplete data streams, such as imaging sonar in underwater settings or thermal sensing under low visibility. In these scenarios, useful evidence is sparse over time, and the objective extends beyond frame-level analysis to reconstructing physically meaningful events. We study this problem through event-centric perception, which maps physical sensor streams into structured events with temporal support, spatial context, direction, count, and magnitude. Frontier multimodal LLMs, such as Qwen3.5-Plus and Gemini, have shown promising capability for event-centric inference under weak and temporally distributed evidence. However, with direct prompting, weak signals often remain below the models' effective decision boundary, leaving event-level structure underused. We address this problem with localized event reasoning, structured event records, and lightweight event-aligned adaptation. This moves weak signals from ignored observations into usable event evidence. In our experiments on a sonar benchmark, the adapted approach improves positive event recall from 0.041 to 0.914 and reduces normalized event error from 0.982 to 0.334. Thermal experiments further show that the approach extends to wildlife event reconstruction and temporally grounded occupancy reasoning. These results support event-centric perception as a principled framework for physical-stream intelligence.
PaperID: 4502, Poster
Abstract: Reinforcement learning with verifiable rewards (RLVR) enhances the reasoning capabilities of large language models (LLMs), but at the expense of significant computational overhead due to compute-intensive rollout processes and frequent policy updates. Online prompt selection, a widely adopted strategy for improving training efficiency, maintains per-prompt Bayesian posteriors to predict prompt difficulty and prioritize informative prompts before committing rollout budget. However, these methods assess prompt informativeness without accounting for how reliably learning signal is extracted from sampled responses. In GRPO-based RL algorithms, the realized advantage of a response depends not only on its own outcome, but also on the randomly sampled outcomes of its peers through group normalization. Our experimental and theoretical analysis show that the resulting uncertainty in group composition induces composition noise, a non-vanishing variance component that creates an irreducible lower bound on gradient estimation error. Consequently, the standard group-relative advantage fails to faithfully characterize response-level utility, degrading gradient estimation and thereby impairing downstream prompt selection. To address this issue, we propose a unified Marginalized Posterior-Predictive framework, MaPP, for data-efficient RLVR, which first denoises response-level advantage estimation and then improves prompt selection using a shared Beta posterior. Specifically, for each response, MaPP replaces the standard group-relative advantage with a composition-invariant intrinsic advantage via closed-form Beta-Binomial marginalization. This yields a closed-form posterior-predictive advantage estimator whose error provably diminishes as the posterior concentrates. Then, building on the same posterior, MaPP derives an uncertainty-aware prompt selection score that more faithfully characterizes prompt informativeness, improving data efficiency without additional rollout cost. Experiments across mathematics, planning, and visual geometry on five model backbones show that MaPP consistently outperforms GRPO and strong selection baselines, achieving up to +2.45 average accuracy improvement over the strongest baseline under the same rollout budget, a new state-of-the-art.
PaperID: 4503, Poster
Abstract: Forecasting spatio-temporal human mobility is essential for urban infrastructure optimization, yet dense traffic sensing remains prohibitively expensive for many cities. Existing cross-city transfer methods and recent spatio-temporal foundation models usually rely on historical target-side traffic sequences as the prediction context, which limits their applicability to under-instrumented, graph-missing, or newly monitored cities. We study exogenous-driven cross-city transfer, where target traffic must be predicted from static and dynamic exogenous urban context without historical flow input. This exogenous-only observability removes the direct endogenous state observation used by flow-driven predictors and turns cross-city forecasting into latent traffic-generating mechanism recovery under city shift. The key challenge is to recover traffic-state information from indirect external evidence while organizing cross-city response patterns into a mechanism space that can adapt to target variation without destabilizing transferable structure. To address this challenge, we propose ExoC2T, a plasticity-inspired spatio-temporal learning framework. ExoC2T uses an Exogenous-Infused Spatio-Temporal Network (EXIST) to build an exogenous mechanism substrate, where semantic relations, temporal regimes, and environment-conditioned responses are organized into a shared prototype geometry. On this substrate, Stability-Plasticity Adaptation separates a stable exogenous-to-flow core from an adaptive city-conditioned deviation, infers a support-conditioned latent posterior from limited target evidence, and controls target risk through invariant discrepancy reduction and posterior correction. The resulting model preserves transferable exogenous-to-flow structure while regulating target-specific modulation through posterior-dependent plasticity. Extensive experiments on a multi-city benchmark show consistent improvements over strong forecasting and transfer baselines, demonstrating the value of exogenous-driven transfer for low-cost and scalable urban computing.
Abstract: Model merging aims to combine multiple models into one without additional training. Naïve parameter-space averaging can be fragile under architectural symmetries, as their geometry does not take them into account. In this work we show that not only the geometry, but also the averaging procedure itself, must be symmetry-invariant to achieve symmetry-aware merges. Consequently, we propose a general solution: merging as Fréchet averaging, i.e., selecting parameters that minimize a sum of geodesic distances on an appropriate manifold. In this view, the key design choice is the overall geometry, i.e., the choice of metric, manifold, and distance approximation, that determines what it means for two models to be “close.” We show that Fréchet averaging, combined with simplifying assumptions, contains Fisher merging. Building on this, we examine the particular case of low-rank adapters (LoRA), whose symmetries induce a distinct geometry: that of a quotient manifold. We outline the limitations of current LoRA merging methods, propose a practical algorithm for this setting, and show how they compare with other commonly used approaches.
PaperID: 4505, Poster
Authors:
Huanzhi Mao, Chengkun Cao, Shuo Yuan, Joseph GonzalezAbstract: Large language models are increasingly capable of writing and executing code, suggesting a natural paradigm for tool-augmented agents: expressing tool use as programs rather than isolated calls. Programmatic tool calling (PTC) captures this idea through loops, batching, abstraction, and structured computation, but existing PTC approaches are largely stateless, limiting their ability to reuse intermediate results and adapt over time. This is problematic for real-world tasks, where tool environments and problem-solving processes are inherently stateful. We therefore present I-PTC, an interactive framework that executes stateful PTC: models issue code snippets that call tools, maintain reusable state, and evolve the tool interface through API corrections. To evaluate the effectiveness of I-PTC and the general PTC setting, we introduce PTC-BENCH, a benchmark of long-tail, production-style user requests whose difficulty comes from large initial environments, many dependent API interactions, extended planning, and recovery across follow-ups and transient failures. Experiments on PTC-BENCH show that I-PTC improves performance on complex multi-step tasks with large tool sets while substantially reducing token consumption over conventional tool-calling baselines and the original PTC baseline, highlighting stateful programmatic tool calling as a promising direction for scalable tool-augmented agents.
PaperID: 4506, Poster
Authors: Slater R Victoroff, Madison May
Abstract: Implicit neural video representations offer compact, continuous alternatives to conventional codecs, but higher reconstruction fidelity typically requires more computation per decoded frame. We introduce NIKA, which shifts video-specific capacity from the decoder into a structured latent state, allowing reconstruction quality to scale with latent expressivity while keeping active decoding lightweight. NIKA constructs this state from complementary spatial, spectral, and temporal bases, then decodes it with lightweight ConvNeXt-style convolutional upsampling. On UVG, a 2.91M-parameter NIKA model achieves 33.33 dB PSNR at 462 decoding FPS with 4.5G MACs on an RTX A5000, outperforming comparable single-resolution NeRV-family baselines while using 39--51× fewer MACs. Ablations show that diversifying latent components improves reconstruction more reliably than reallocating capacity within a single component, and qualitative analysis reveals specialization among components. Together, these results identify structured latent diversity as a practical alternative to scaling decoder complexity for high-fidelity, efficient neural video representation.
PaperID: 4507, Poster
Abstract: Token-based world models enable fine-grained latent planning, but iterative search inside the predictor is dominated by the number of spatial tokens processed during rollout. We introduce COSTGRAD, a training-free, goal-conditioned selector that ranks spatial tokens by the gradient norm of the planning cost with respect to each input token. By deriving importance from the downstream control objective, COSTGRAD targets tokens that matter for planning rather than merely for prediction. On AdaLN-conditioned predictors at 50% sparsity, COSTGRAD matches or exceeds full-token planning on three of four continuous-control benchmarks, while giving a measured 2.6× wall-clock speedup per environment planning step. The sparsity dividend can be combined with reduced CEM search, yielding a ~5× total speedup while still exceeding the full-token baseline. We also identify an architecture-dependent failure mode: in a matched AdaLN-vs-concat comparison, concat maintains comparable full-token performance but pure COSTGRAD loses its advantage over random selection. This failure tracks action-pathway drift under selected token removal: AdaLN largely preserves how actions affect kept tokens, whereas concat perturbs the action-conditioned computation. Mixing in random anchors partially recovers COSTGRAD's advantage on some matched-concat checkpoints, suggesting selector--architecture compatibility as a design axis for sparse world-model planning.
PaperID: 4508, Poster
Authors:
Ruochi Zhang, Yusi Fan, Qiong Zhou, Li Jiao, Tian Wang, Qian Yang, Silong Zhai, Lan Wang, Fengfeng Zhou, Yajuan Huang, Liming Guo, Chang Liu, Xin GaoAbstract: Predicting mutation-induced changes in peptide--protein binding affinity (\Delta\Delta G) is central to therapeutic peptide optimization, but peptide-specific predictors remain limited by scarce labels, target overlap in supervised benchmarks, and costly molecular simulations. We introduce PepDDG, a zero-shot, training-free predictor that ranks peptide mutations by decomposing binding perturbations into three complementary information channels: energetic perturbation, geometric environment, and evolutionary compatibility. Each channel is computed from a wild-type complex structure, transformed into rank space, and combined by non-parametric Borda aggregation, without fitting fusion weights or fine-tuning neural predictors. PepDDG is motivated by rank-covariance analysis showing cross-channel fusion can improve rank correlation when channels retain complementary signal, whereas increasingly expensive refinement of a single energetic channel has diminishing returns. Empirically, the channels are individually moderate but capture distinct mutation regimes, and rank fusion consistently outperforms single-channel variants and training-free baselines. On our curated benchmark of 332 peptide-chain mutations across 33 targets, PepDDG reaches Spearman \rho = 0.619; the calibrated PepDDG-Cal variant reaches \rho = 0.691. Performance remains strong on short and cyclic peptide subsets, with \rho = 0.813 and 0.872, respectively. These results support complementary evidence fusion on PDB-derived complex structures as a practical alternative to costly same-channel physical refinement for peptide \Delta\Delta G ranking; a separate diagnostic shows that predicted structures can be used when no crystal structure is provided at inference time.
Authors:
Zicheng Zhang, Haoran Li, Jiaxing Wang, Guoqiang Gong, Anqi Li, Yudong Hu, Ting Xiong, Yurong Gao, Junxing Hu, Zhida Jiang, Yifeng Zhang, Pengzhang Liu, Qixia JiangAbstract: In Low-Rank Adaptation (LoRA), the scaling factor \alpha is often treated as a mere complement to the learning rate, yet its role in optimization remains poorly understood. In this paper, we reveal that the scaling factor \alpha and the learning rate function differently, with \alpha emerging as the dominant driver of effective optimization, delivering gains that cannot be replicated by learning rate scaling alone. Through the synergy of extensive empirical analysis and a theoretical Signal-Drift framework, we uncover three findings into LoRA’s scaling mechanism: First, LoRA’s spectral suppression smooths the optimization landscape, rendering standard hyperparameters overly conservative and creating an optimization gap. Second, when leveraging this smoothness to accelerate convergence, \alpha outperforms the learning rate by amplifying the task signal without increasing the drift ratio. Third, the optimal scaling factor follows a sublinear relationship with the rank, well characterized by a square-root law with an unexpectedly large coefficient, revealing the insufficient scaling of existing rank-tied heuristics. Based on these insights, we propose LoRA-\alpha, a minimalist framework that restores \alpha to its principled regime, making LoRA compatible with standard small learning rates. Extensive evaluations across diverse tasks demonstrate that LoRA-\alpha consistently improves performance while streamlining hyperparameter search, unleashing the learning potential of LoRA.
PaperID: 4510, Poster
Authors: Zhihong Cui, Hengyu Liu, Haoran Tang, shijun liu, Amir Taherkordi, Tor Skeie
Abstract: Agent harnesses—runtime scaffolds that wrap an LLM's reasoning loop with tool dispatch, verification, and persistent memory—have become the standard substrate for deploying LLMs as autonomous agents, as demonstrated by Claude Code and Codex in software engineering. In trajectory planning, however, this paradigm has not taken hold: existing LLM-based trajectory planning methods either keep the LLM in the decision loop, exceeding the millisecond-level control budget, or offload it to offline rules or training-time supervision—leaving the runtime planner non-evolving. Intermediate hybrids still lack the verification-and-consolidation loop that makes harnesses reliable. We ask: what runtime harness would let an LLM planner meet real-time constraints while self-evolving into a deployable one? We propose REINS, a harness built on three LLM–planner couplings: LLM outputs are (i) grounded to a learned variable-level dynamics model, (ii) verified via calibrated predicates and forward rollout, and (iii) consolidated into an indexed skill memory that serves future scenes in \mathcalO(\log n) without re-invoking the LLM. The harness closes a self-evolution loop: familiar scenes resolve by memory lookup; novel scenes invoke the LLM, whose verified outputs enrich the memory. Without modifying LLM weights, REINS attains 98%+ compliance, 50%+ collision reduction, and 2–5 ms latency on simulated and real-world benchmarks, with LLM invocation rate monotonically declining as memory matures. Code: https://anonymous.4open.science/r/REINS-1FE3/
PaperID: 4511, Poster
Authors:
Jianan Fu, ting Wu, Rui Yang, Jiacheng Sun, Xinghong Zhou, Chenghai Mao, Jiaxiang Pan, Xiaomao LiAbstract: 3DGS-based open-vocabulary 3D scene understanding requires identifying reconstructed 3D instances that correspond to a natural-language query. Existing methods typically compress multi-view observations of an instance into a single semantic descriptor, which is insufficient to represent a complete 3D instance. To retain rich semantic information across diverse viewpoints, we propose MSR-3D, a framework built around a Multi-Mode Semantic Representation. Specifically, MSR-3D derives multi-view observations from both 2D segmentation masks and rendered masks of each segmented 3D instance. These observations are organized into groups. Each group is summarized as a semantic mode through a representative feature and associated statistical attributes. Based on the multi-mode semantic representation, we further introduce a lightweight learnable mode-aware matcher for querying. The matcher employs linear scoring modules to project the multi-mode representation of each query-instance pair into a three-dimensional score space. Calibrated decision boundaries are then learned in the resulting score space to separate true query-instance matches from semantically related distractors. This structured and decomposable design further improves the interpretability and generalization of MSR-3D. Experiments on open-vocabulary scene-understanding benchmarks show that MSR-3D achieves state-of-the-art results, improving mIoU on LERF-OVS by 7.52% over the baseline and achieving the best overall performance on 3DOVS.
PaperID: 4512, Poster
Abstract: Large vision model (LVM)-based gait recognition has achieved remarkable performance. However, current LVM-based methods largely rely on 2D image-plane representations and often neglect the underlying 3D structure of the human body. Without 3D structural guidance, 2D representations are sensitive to viewpoint variations. This sensitivity leads to unreliable gait features and limits generalizability in real-world 3D scenarios. To bridge this gap, we propose PuppetGait, a framework that aligns LVM features into a unified 3D canonical body space. The method conceptually manipulates the features like a puppet, adjusting them to a predefined pose and viewpoint. Comprehensive evaluations on CCPG, CCGR-MINI, CASIA-B, and SUSTech1K show that PuppetGait maintains strong within-domain performance (96.54% mean Rank-1 on CCPG) while substantially improving cross-domain generalization, reaching 23.22% Rank-1 on CCPG-to-CCGR-MINI and 63.28% mean Rank-1 on CCGR-MINI-to-CCPG. Code will be publicly available.
PaperID: 4513, Poster
Abstract: Multimodal heterogeneous learning aims to train reliable multimodal models when data sources differ in modality availability, semantic coverage, domain distribution, and missingness patterns. Although cross-modal alignment is widely used to mitigate such heterogeneity, we argue that alignment is not free: enforcing relations that are weakly supported by the observed data can amplify noise, over-transfer missing-modality information, and induce negative transfer. This paper studies when a modality--modality--semantic relation is sufficiently observable to be safely aligned and shared across heterogeneous sources. We propose , an observability-aware framework that summarizes visible modality--semantic evidence through compact source-level sketches and estimates relation-level support from bridge evidence, side support, cross-modal consistency, and prototype uncertainty. The resulting observability score gates co-observed alignment, calibrates weak missing-modality transfer, and guides evidence-aware source consolidation so that reliable relations are shared while unsupported variation remains local. Across image--text, audio--visual--text, action-recognition, sensor, and healthcare benchmarks under source-partitioned multimodal heterogeneity, ObsAlign achieves the best overall performance in both encoder-based and VLM-compatible regimes, while sharply reducing negative transfer on low-observability relations. These results suggest that relation-level observability provides a practical principle for deciding when cross-modal alignment should be strengthened, weakened, or avoided.
PaperID: 4514, Poster
Abstract: Low-rank adaptation (LoRA) has attracted significant attention in the parameter-efficient fine-tuning (PEFT) landscape. However, its design suffers from two fundamental limitations. First, its invariance manifold induces flat directions in the loss landscape, creating an optimization bottleneck for adaptive optimizers. Second, its effective rank scales only linearly with the parameter budget. In this paper, we propose Scalable Morpho Adaptation (SaMA), a PEFT method grounded in a principled generalization of the Kronecker product. By relaxing the shared-block constraint in the perfect-shuffle factorization, SaMA spans from low to full rank and achieves effective rank that scales \quadratically in parameters. Furthermore, we show that its invariance manifold collapses to a diagonal group, making it structurally more optimizer-friendly. Empirically, SaMA consistently outperforms strong PEFT baselines on commonsense and arithmetic reasoning benchmarks while being more parameter-efficient.
PaperID: 4515, Poster
Abstract: Modern video object segmentation (VOS) relies on a two-stage paradigm for robust performance: target initialization and memory-based propagation. However, existing adversarial attacks on VOS often overlook this fundamental feature. To bridge this gap, we propose AUV, an adversarial framework that Aligns with the Universal two-stage paradigm of VOS models. Specifically, AUV optimizes a single universal adversarial perturbation (UAP) through decoupled training strategies tailored for both initialization and propagation stages. To ensure prompt-agnosticism across all modalities, including masks, points, boxes, and language, AUV optimizes the UAP to disrupt the shared memory space for latent-level target erosion. Extensive experiments show that AUV establishes a new state-of-the-art for VOS attacks, such as degrading mask-based SAM3 to only 8.4 J&F on DAVIS17-val (compared to 55.4 for previous SOTA under the same training setup), while consistently outperforming prior methods on other VOS models including SAM2, Cutie, and XMem. Notably, AUV exhibits superior temporal dynamics with the fastest onset and the most enduring persistence, while uniquely supporting flexible any-time mid-video attack, revealing vulnerabilities in real-world streaming applications. Our code and checkpoint will be made publicly available.
PaperID: 4516, Poster
Abstract: Recent large-scale outages at major cloud providers such as AWS and GCP have exposed the fragility of relying on a single cloud. To improve resilience, fault tolerance, and business continuity, many enterprises are moving their services to multi-cloud environments. However, multi-cloud deployment also makes workload forecasting substantially harder. Existing multi-cloud workload forecasting methods are typically designed for either centralized or isolated environments, leading to risks of private data leakage or limited generalization. To address these challenges, we propose FedCAG, a Federated Causality-Aware Graph learning paradigm for multi-cloud workload forecasting. Instead of treating federated learning as parameter averaging over local predictors, FedCAG jointly federates predictive models and graph-structured dependency priors. Each client constructs causality-aware, spatial, and temporal graphs from its private telemetry to model directed inter-metric influence, metric-level interactions, and intra-window temporal dynamics. To address strong cross-cloud heterogeneity, FedCAG further combines causality-aware representation fusion, adaptive graph refinement, and client-specific personalization, enabling the global model to benefit from shared workload structures while preserving local specificity. Experiments on real-world datasets show that FedCAG consistently outperforms mainstream federated baselines and even several centralized methods trained on the full dataset, delivering stable and accurate forecasting across heterogeneous and privacy-constrained deployments. The source code is available at: \urlhttps://anonymous.4open.science/r/FedCAG-FB9C.
PaperID: 4517, Poster
Abstract: Recovering global 3D human motion from monocular video captured by a moving camera is a fundamental yet challenging problem, as camera ego-motion and human body motion are tightly entangled in the image observations. The prevailing two-stage paradigm treats camera estimation and motion reconstruction as isolated processes, causing errors on both sides to be further amplified when combined in the world coordinate system. To address this, we draw inspiration from the human inner visual simulation mechanism and propose EIHMR, a collaborative human-camera co-estimation framework. EIHMR comprises two complementary modules that bridge scene-aware human motion refinement and motion-aware camera estimation: Scene-Aware Local Human Motion Reconstruction reprojects motion sequences into frozen keyframe viewpoints and leverages metric depth and kinematic constraints to produce geometrically consistent local motion, while Motion-aware SLAM re-renders the refined motion as static meshes in the original frames, converting dynamic human regions into structured matching cues for robust camera estimation. EIHMR consistently improves global trajectory reconstruction over strong baselines, demonstrating the effectiveness of collaborative human-camera estimation for long-range human motion recovery.
Abstract: Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly understood. Existing systems already span multiple process designs, including direct response generation, text-only prior turn, visual-state manipulation, and explicit external image-tool invocation. In this paper, we ask which of these evaluated paradigms improves multimodal jailbreak robustness, and why. Across multiple vision-language models, explicit image-tool interaction yields the lowest attack success rates in our experiments, reducing jailbreak success by around 30% relative on average across the evaluated models. This finding is initially surprising: ASR remains low even when the returned image-tool output is manually overridden or itself unsafe-looking, but returns near direct-answering levels under text-only prior turn controls. These results indicate that the lower ASR is not explained by benign returned-image semantics or by the textual image-tool trace alone. To explain the pattern, we introduce an image-tool safety vector framework that models image-tool invocation as a residual shift in hidden representations toward a safety-relevant direction. Representation-level analyses and activation interventions support this account. Overall, our results suggest that explicit image-tool interaction is a promising design pattern for improving jailbreak robustness, while also motivating pipeline-specific safety evaluation.
Authors: Tho Mai, Joo-Young Kim
Abstract: Large language models (LLMs) support long-context inference but suffer from substantial memory and runtime overhead due to Key-Value (KV) Cache growth. Existing KV Cache eviction methods primarily rely on local attention weights, neglecting the influence of value representations, output projection, and inter-head interactions. In this work, we reformulate KV Cache eviction from a conventional head-wise, weight-averaging approach into an output-aware, layer-wise matrix multiplication approximation problem. We introduce LaProx, a novel eviction strategy that explicitly models the multiplicative interaction between attention maps and projected value states to accurately quantify token contributions while accounting for inter-head dependencies. Building on this metric, we propose the first unified eviction strategy that assigns globally comparable importance scores to tokens, enabling model-wide selection instead of local, head-wise decisions. Experimental results across 19 datasets on long-context benchmarks LongBench and Needle-In-A-Haystack demonstrate that our approach maintains model performance with only 5% of the KV cache and consistently outperforms prior works across all configurations. Notably, our method achieves up to 2× accuracy loss reduction under extreme compression scenarios compared to existing state-of-the-art baselines with minimal overhead.
Abstract: Generative policies based on diffusion models and flow matching have shown strong promise for offline reinforcement learning (RL), but their applicability remains largely confined to continuous action spaces. To address a broader range of offline RL settings, we extend flow matching to a general framework that supports discrete action spaces with multiple objectives. Specifically, we replace continuous flows with continuous-time Markov chains, trained using a Q-weighted flow matching objective. We then extend our design to multi-agent settings, mitigating the exponential growth of joint action spaces via a factorized conditional path. We theoretically show that, under idealized conditions, optimizing this objective recovers the optimal policy. Extensive experiments further demonstrate that our method performs robustly across diverse settings and benchmarks, including high-dimensional control, multi-agent games, and dynamically changing preferences over multiple objectives, while outperforming traditional offline RL methods in practical multi-modal decision-making scenarios. Our discrete framework can also be applied to continuous-control problems through action quantization, providing a flexible trade-off between representational complexity and performance.
PaperID: 4521, Poster
Abstract: Flow Matching (FM) is a strong framework for generative modeling due to its stable training and efficient sampling. However, FM assumes clean target data and labels, an assumption often violated in practice. We formulate FM training as a time-dependent regression problem over conditional vector fields and analyze how the two corruption types affect the objective: i) target corruption adds a time-dependent residual term that grows with time, ii) label noise causes updates to the wrong class-conditional vector field. Based on this analysis, we propose to improve FM robustness by mapping corruption-specific residual scores to sample weights. For target corruption, we propose the \emphTemporal Residual Score (TRS), which emphasizes residual magnitudes at intermediate-to-late FM times, where clean and corrupted samples are more separable. For label noise, we propose the \emphComparative Conditional Score (CCS), which compares the residual of the observed class-conditional vector field against alternative class-conditional vector fields for the same trajectory and velocity target. Experiments across controlled corruption settings and real-world noisy labels show that the proposed methods improve robustness over standard FM and robust-training baselines, supporting corruption-specific residual structure as a practical basis for robust FM.
Authors:
Salil Parth TRIPATHI, bertrand chapron, Fabrice collard, Nicolas Courty, ronan fabletAbstract: While optimal transport (OT) enforces a rigid constraint by requiring two measures to be matched exactly, partial optimal transport relaxes this requirement by allowing mass to remain unmatched through a global budget, scalar rebate, or uniform rejection rule. However, many applications call for more structured, pointwise rejection mechanisms, where the decision to leave mass unmatched depends on side-specific reliability, support geometry, or external information about which components should participate in the comparison. We introduce \emphintent-controlled partial optimal transport (IC-POT), a targeted generalization of partial transport that replaces the global rejection paradigm with pointwise rejection costs over both measures. We show that the resulting optimization problem admits a dual interpretation in terms of local acceptance thresholds and can be solved by recasting it as a balanced Kantorovich OT problem on an augmented support. Beyond theoretical analysis, we demonstrate the practical relevance of IC-POT in settings where rejection is driven by side information. In positive-unlabeled learning and open-partial domain adaptation, incorporating pointwise rejection rules that encode statistical structure improves fixed baseline pipelines. Finally, we motivate the use of IC-POT with a geophysical practical case: multi-modal satellite ocean measurements, for which physical and instrumental priors naturally inform the rejection mechanism and define the retrieved comparable signal information.
Abstract: Post-training of Large Language Models (LMs) often prioritizes accuracy and helpfulness at the expense of diversity. This creates a tension: while post-training improves response quality, it also sharpens output distributions and reduces the range of ideas, limiting the usefulness of LMs in creative and exploratory tasks such as brainstorming, storytelling, or problem solving. We address this challenge with Diversity-Aware Reinforcement Learning (DARLING), a framework that jointly optimizes for response quality and semantic diversity. At its core, DARLING introduces a learned partition function to measure diversity beyond surface-level lexical variations. This diversity signal is then combined with a quality reward during online reinforcement learning, encouraging models to generate outputs that are both high-quality and distinct. Experiments across multiple model families and sizes show that DARLING generalizes to two regimes: non-verifiable tasks (instruction following and creative writing) and verifiable tasks (competition math). On five benchmarks in the first setting, DARLING consistently outperforms quality-only RL baselines, producing outputs that are simultaneously of higher quality and novelty. In the second setting, DARLING achieves higher pass@1 (solution quality) and pass@k (solution variety). Most strikingly, explicitly optimizing for diversity catalyzes exploration in online RL, which manifests itself as higher-quality responses.
PaperID: 4524, Poster
Abstract: Neural denoisers trained with pointwise distortion losses such as MSE often suppress the temporal regularities that downstream tasks depend on---a failure that becomes acute under non-Gaussian, impulsive, or non-stationary disturbances. We propose a training framework that lifts a family of classical complexity descriptors---permutation entropy, sample entropy, Lempel-Ziv complexity, and Higuchi fractal dimension---into end-to-end differentiable objectives via calibrated soft surrogates with straight-through gradients. Three complementary mechanisms emerge from this lifting: (i) early fusion injects local structural priors as auxiliary inputs, (ii) a complexity-consistency regularizer matches the multi-scale complexity of denoiser outputs to clean references, and (iii) a curriculum sampler prioritizes examples whose noise-induced complexity inflation is largest. We prove that additive noise strictly inflates ordinal uncertainty for bounded piecewise-smooth signals, and that, under bounded surrogate calibration, complexity consistency yields a strictly larger expected decision margin than MSE-only training. On OFDM/QAM communication signals under AWGN, 1/f colored, and \alpha-stable impulsive noise, our framework reduces post-detection bit error rate by 29--47% over Transformer (PatchTST, iTransformer), GNN, and CNN baselines, with statistically significant gains across five seeds and 95% bootstrap confidence intervals. The improvements transfer to over-the-air conditions on the public RadioML 2018.01A benchmark (single-carrier QAM portion) and a new 50-session in-house SDR testbed, yielding 22--26% relative BER reduction without fine-tuning. The framework is architecture-agnostic and adds \leq 3.5% FLOPs over the underlying backbone.
PaperID: 4525, Poster
Authors: Leixuhai Xu, Xiangtao Zhang, Hailong Yan, Le Zhang
Abstract: Federated learning (FL) suffers from severe performance degradation under statistical heterogeneity, where label skew and domain shift induce representation drift and biased decision boundaries. While vision--language models such as CLIP provide transferable semantic priors, raw CLIP embeddings are noisy, highly correlated, and not directly aligned with the latent space of lightweight federated classifiers. We propose Semantic-Bridge Federated Learning (SBFL), a framework that transfers pre-trained vision--language semantics into heterogeneous FL without accessing raw client data or requiring online CLIP inference. SBFL constructs text-verified semantic banks, learns a server-side semantic bridge that maps CLIP-derived features into student-compatible teacher spaces, and regularizes local training through mixed semantic prototype alignment and quality-aware semantic replay. Experiments on CIFAR-10, CIFAR-100, Tiny-ImageNet, Office-Caltech 10, and Digit5 show that SBFL consistently improves the average performance of representative FL backbones under both label-skewed and domain-shifted settings. Analysis further shows that the semantic bridge converts highly correlated CLIP priors into more separable teacher-space geometry, explaining its effectiveness in improving representation alignment and global generalization. Our code will be released upon publication.
Abstract: Recent generative methods for single-shot high dynamic range (HDR) image reconstruction show promising results, but often struggle with preserving fidelity to the input image: they hallucinate content, require separate models to handle highlights and shadows, or sacrifice interpretability by directly predicting the final HDR image. We address these limitations by re-casting single-shot HDR reconstruction as conditional video generation and fusing the generated frames into an HDR image. We fine-tune a video diffusion model to generate an exposure bracket, conditioned on a low dynamic range (LDR) input. We fuse this image bracket using per-pixel weights predicted by a light-weight UNet. This formulation is simple, interpretable, and effective. Rather than directly hallucinating an HDR image, it explicitly reconstructs the intermediate exposure stack and fuses it into the final output. Our method eliminates the need for separate models across exposure regimes and produces HDR reconstructions with high input fidelity. On quantitative benchmarks, we outperform state-of-the-art generative baselines with comparable model capacity on several reconstruction metrics. Human evaluators further prefer our results in 72% of pairwise comparisons against existing methods. Finally, we show that this input-conditioned sequence generation and fusion framework extends beyond HDR to other image reconstruction tasks, such as all-in-focus image recovery from a single defocus-blurred input.
PaperID: 4527, Poster
Authors: Senanayak Sesh Kumar Karri
Abstract: Kolmogorov–Arnold Networks (KANs) replace fixed activations in deep architectures with learnable univariate edge functions, making the choice of edge parametrisation central. Existing variants rely on fixed bases such as splines, polynomials, or Fourier features, which impose a function-space geometry before data are observed. We introduce geometry-constrained KANs, a family of edge activations derived from Banach duality maps in which the geometry itself is learned through a scalar exponent p > 1 per edge. This exponent controls the qualitative response: sub-Euclidean values produce sharp threshold-like behaviour, p = 2 recovers the linear regime, and larger values produce flatter responses near the origin. Beyond expressivity, the exponent also modulates sensitivity: smaller values reduce amplification of perturbations, providing an implicit regularisation effect without introducing an explicit shrinkage hyperparameter. Across 50 Feynman symbolic regression equations, geometry-constrained KANs match strong fixed-basis baselines on clean data while substantially improving robustness under distribution shift. In particular, they outperform cross-validated spline baselines in 85 of 90 data-scarcity settings and degrade significantly less under noise (3.7× versus 21.6× for splines). Learned exponents are interpretable and stable, revealing consistent geometric structure across equation families and input dimensions.
Abstract: Web-enabled LLM agents are changing how online information influences search outcomes. Existing Generative Engine Optimization (GEO) studies mainly focus on individual webpages. However, agentic web search is not a single-document setting: an agent may issue queries, crawl pages, follow links, reformulate searches, and synthesize evidence across multiple browsing steps. Influence therefore depends not only on page content, but also on how pages are organized, connected, and encountered along the agent's browsing trajectory. We study this shift through Ecosystem Generative Engine Optimization (EcoGEO), which treats GEO as an environment-level influence problem for web-enabled LLM agents. To instantiate this perspective, we propose TRACE, a Trajectory-Aware Coordinated Evidence Ecosystem. Given a recommendation query and a fictional target product, our method builds a controlled evidence environment that coordinates an agent-facing navigation entry page with heterogeneous support pages. These pages use shared terminology, internal links, and consistent product attributes to introduce, verify, and reinforce the target product. We evaluate our method on OPR-Bench, a benchmark for open-ended product recommendation. Experiments show that it consistently outperforms page-level GEO baselines in final target recommendation. Trajectory-level metrics further show increased initial target-result crawls, target-specific follow-up searches, and internal-link crawls, suggesting that the gains come from shaping the agent's evidence-acquisition process rather than merely adding more target-related content. Overall, our findings support an ecosystem research paradigm for GEO, where web-enabled LLM agents are studied in relation to the broader evidence environments that guide search, browsing, and answer synthesis.
Abstract: LLM powered multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, their advantages over single-agent systems (SAS) remain unclear, with performance varying inconsistently across settings. Here, we provide an information bottleneck perspective on elucidating the differences between MAS and SAS. Specifically, our key observation is that a SAS accumulates its full reasoning trace in one shared context, while a MAS uses isolated local contexts connected by bounded relay messages. We show that, under infinite relay bandwidth, any SAS can be simulated by a MAS that transmits the full upstream context. Thus, the nontrivial advantage of MAS arises under bounded relays, where compression introduces a fundamental trade-off: reducing redundant context can improve efficiency, but may also incur loss of task-relevant information. We formalize this trade-off as an information bottleneck controlled by an effective parameter \beta, which captures how the balance shifts with model capability, and shows that MAS gains arise when context reduction outweighs relay information loss. We conduct 18 controlled experiments across five benchmarks and three model scales to validate our theoretical studies. We observe that MAS consistently helps when relays are near-sufficient, especially for weaker models. In contrast, MAS gains shrink or reverse when relays incur information loss, especially for stronger models that can already extract useful information from redundant context and thus gain little from compression. Our study shows that multi-agent design is fundamentally an information-bottleneck optimization problem. This perspective explains when bounded inter-agent communication helps or hurts.
PaperID: 4530, Poster
Authors:
Meiyun Lu, Wei Yue, Weihong Wu, Lin Zeyu, Longkun GuoAbstract: Fairness has emerged as a central consideration in machine learning, motivating the study of fair range k-supplier clustering as a fundamental problem that focuses fairness on the selected centers. Given a set of suppliers and clients, where each supplier may belong to one or more demographic or functional groups, the objective is to select k suppliers that minimize the maximum client-to-center distance while ensuring that the number of selected suppliers from each group satisfies prescribed lower and upper bounds. For disjoint supplier groups, we develop a polynomial-time 3-approximation algorithm in the offline setting and a (3+\epsilon)-approximation algorithm in the streaming setting. We further consider overlapping supplier groups and show that this generalization admits a parameterized 3-approximation algorithm whose runtime is exponential in the number of clusters. Finally, experiments on both synthetic and real-world datasets demonstrate that our algorithms achieve significantly better clustering quality and runtime efficiency in comparison with state-of-the-art methods.
Abstract: Reconstructing articulated 3D objects is important for animation, gaming, and robotic simulations. Recent neural networks can estimate the articulated structure of 3D objects, but their generalization remains limited by the scarcity of annotated data for this task. To address this gap, we introduce Instruct-Particulate, a model that takes a 3D mesh together with a target kinematic specification, including part descriptions, connectivity, joint types, and optional point prompts, and predicts the corresponding kinematic part segmentation and joint motion parameters. The kinematic specification disambiguates the task and allows the model to target annotations of different granularity, thereby making it possible to use more abundant heterogeneous training data. At test time, the kinematic specification can be obtained automatically, so the model can be applied to any input mesh. To train our model at scale, we construct a heterogeneous dataset of more than 150,000 articulated 3D objects, extending existing publicly available collections with data obtained by partially labelling other 3D models, monolithic or already decomposed into parts, with kinematic labels by means of vision-language models. Experiments show that our model generalizes better across categories and to AI-generated meshes, enabling articulated asset reconstruction from real-world images via image-to-3D models.
PaperID: 4532, Poster
Abstract: Voice style conversion (VSC) aims to transform a source utterance to match the timbre, accent, and emotion of a target voice while preserving linguistic content. StyleStream 1.0 introduced the first streamable zero-shot VSC system with state-of-the-art conversion quality, but remained limited in two ways: approximately 1s end-to-end latency on consumer hardware, and the lack of a text-based interface for designing or editing target voices. We present StyleStream 2.0, a fast and controllable streaming VSC framework that addresses these limitations. To reduce latency, we train a few-step pixel MeanFlow model and further fine-tune it with autoregressive feedback inspired by Self Forcing, improving robustness under streaming inference. This reduces end-to-end latency to 520~ms on a consumer GPU, a 2x speedup over StyleStream 1.0, while maintaining conversion quality. To enable controllable style manipulation, we remove the mel-spectrogram context and instead route target style information through a compact style embedding. In this embedding space, we train a unified text-conditioned flow matching model that supports both text-based voice design, which maps natural language prompts to voice styles, and instruction-based voice editing, which modifies a specific attribute of a source utterance while preserving the rest. Experiments show that StyleStream 2.0 achieves the strongest target style fidelity on VSC, voice design, and voice editing, while remaining competitive on intelligibility and non-edited attribute preservation.
PaperID: 4533, Poster
Abstract: We extend the Bayesian learning rule by explicitly incorporating the compositional structure of neural networks. The central observation is that backpropagation and exponential family variational inference share a common Lagrangian duality structure. Leveraging this connection, we develop a new Bayesian learning framework in which adjoint variables backpropagate layerwise sensitivity signals that converts into loss-site natural parameters. For Gaussian weight posteriors and Gaussian layer-state projections, the framework recovers deterministic back-propagation and Hessian backpropagation via delta-method approximations, while providing a unified perspective on localization-style prediction rules and moment matching methods. We also derive a new variant of IVON, clarify how its gradient and Hessian estimates differ from those of the standard version, and demonstrate how neural network structure can be incorporated into Bayesian learning in practice.
PaperID: 4534, Poster
Abstract: Unified Multimodal Models (UMMs) achieve remarkable success on diverse multimodal understanding and generation tasks, yet they still struggle to effectively leverage language-based chain-of-thought reasoning for image generation. As a result, existing reasoning-augmented UMMs often fail to inherit the test-time scaling and fine-grained controllability observed in large language models. We attribute this failure to a fundamental modality mismatch: reasoning is performed in a discrete token space, whereas image generation operates in a continuous space. To bridge this gap, we introduce a Latent Reasoning framework in Continuous space (LARC), which alleviates the token-level bottleneck by reasoning over continuous hidden states. Specifically, the training framework consists of two stages: (i) curriculum supervised fine-tuning (SFT), and (ii) information-gain reinforcement learning (RL). First, the curriculum SFT stage gradually converts token-level reasoning into latent reasoning. Subsequently, the information-gain RL stage self-evolves latent reasoning without ground truth supervision. Experiments across diverse text-to-image generation benchmarks show that LARC consistently improves generation quality, prompt following, and compositional fidelity. Notably, LARC achieves state-of-the-art performance on GenEval with a 7%p gain over BAGEL and demonstrates clear test-time scaling behavior, outperforming token-based generation under both parallel and sequential test-time scaling strategies.
Abstract: Knowledge distillation generally assumes a strong-to-weak relationship where stronger teachers yield better students. In this work, we examine this assumption about distillation in large language model (LLM) pretraining. By varying architecture sizes and training token budgets, we create strong-to-weak, same-level, and weak-to-strong teacher-student relationships. We study distillation's effectiveness under these relationships with different mixes between language and distillation loss. Three findings emerge: (1) with proper loss mixing, weak-to-strong and same-level distillation improves over standard pretraining, where even small and undertrained teachers benefit large students; (2) making the teacher stronger can lead to saturated or even reversed gains; (3) distillation improves generalization (out-of-domain, downstream) more readily than in-domain fitting. Our results provide practical guidance for choosing the teacher model with the proper loss for more effective LLM pretraining distillation.
Abstract: Model-based and model-free reinforcement learning are traditionally viewed as separate paradigms: while the former learns an explicit model of the transition dynamics P, model-free agents typically estimate value functions tied to a specific policy and reward. In this paper, we challenge this dichotomy by proving that value-based agents trained on a sufficiently rich set of reward functions, e.g. using goal-conditioned RL, implicitly encode a unique and accurate world model. To extract this model in practice, we introduce P-learning; analogous to Q-learning, which approximates Q^\star using samples from the environment, P-learning extracts an agent's model of the environment P^\star by sampling from its Q-values, policies, and rewards, effectively inverting the Bellman equation. In deterministic MDPs, we prove that the true kernel P can be recovered from an agent trained on a single generic goal for finite state spaces \mathcalS, and a finite number of Gaussian goals for continuous \mathcalS \subseteq \mathbbR^d, provided Q-values are accurate. In the stochastic setting, these conditions generalise to larger sets of goals depending on the reward function family. Even when our assumptions are violated, we empirically demonstrate that agents trained with a small number of sparse rewards encode accurate dynamics in (i) stochastic variants of FourRooms, (ii) MountainCar, and (iii) Reacher. This is further validated by training policies inside the extracted world model, for goals far beyond the training distribution, suggesting that goal-conditioned agents secretly contain implicit generalisation capabilities and providing a new lens into the connection between model-based, model-free, and goal-conditioned RL.
PaperID: 4537, Poster
Authors:
khalid OUBLAL, Malo Guichard, François Bertholom, Simon Albergel, David Benhaiem, Emmanuel LE BORGNE, Vladimir Kostic, Karim LouniciAbstract: State-of-the-art machine learning weather prediction systems achieve strong upper-air forecast skill but degrade at longer lead times, with mesoscale structures progressively vanishing and ensemble forecasts becoming under-dispersive. While these limitations are often attributed to the training objective, we find that message-passing processors significantly contribute to this degradation. Repeated neighborhood aggregation at each autoregressive step acts as a smoothing operator that, over rollout steps, suppresses high-frequency information in latent space. This effect is particularly limiting in regional weather prediction, where fine-scale structures are already partially unresolved, further eroding mesoscale variability and hindering long-horizon forecasting. We propose GraphMet, a graph-based weather model motivated by Koopman theory that replaces stacked message-passing layers with a single linear Koopman operator step. GraphMet learns linear evolution in a spectral latent basis while preserving the encoder, decoder, and multi-scale icosahedral mesh structure. A two-term spectral objective jointly learns a stable Koopman basis and preserves latent geometry, enabling efficient long-range propagation without repeated nonlinear aggregation. Empirically, GraphMet reduces the 15-day activity and improves regional RMSE by 11-14% at 5km resolution, and enables efficient probabilistic forecasting with lower inference cost than state-of-the-art methods. Our code available at : https://anonymous.4open.science/w/GraphMet-54ED/.
PaperID: 4538, Poster
Authors: Boming Chen, Xiaojing Wang
Abstract: Low-Rank Adaptation (LoRA) enables efficient fine-tuning of large language models, but adapted models often become overconfident, especially in low-data settings and under distribution shift. Existing uncertainty-aware LoRA methods typically construct a static posterior over a fixed adaptation set and do not model how local posterior evidence varies across structured subsets of the training data. We propose Seq-LoRA, a post-hoc Bayesian framework for sequentially aggregating slice-wise posterior evidence in LoRA space. Starting from a shared maximum a posteriori LoRA anchor, Seq-LoRA builds slice-wise Kronecker-factored quadratic surrogates, projects them into a shared curvature-informed low-dimensional subspace, rewrites the resulting latent quadratic forms as Gaussian pseudo-observations, and performs exact inference in an induced linear--Gaussian state-space model via Kalman filtering. Rather than relying on a particular curriculum order, Seq-LoRA couples heterogeneous slice-wise local posterior evidence through a random-walk prior, forming a terminal Bayesian posterior for prediction. Across the in-distribution ScienceQA test split and six out-of-distribution (OOD) reasoning targets, Seq-LoRA substantially reduces deterministic LoRA overconfidence under shift. Among compared methods, it achieves the best negative log-likelihood on five OOD targets and the best expected calibration error on four, while preserving most task accuracy.
PaperID: 4539, Poster
Abstract: The reasoning capabilities of large language models are reported to drop sharply when they are forced to think in heavily regulated target languages such as Chinese — a phenomenon the literature attributes to data imbalance, multilingual transfer failures, or, more recently, RL-induced cross-lingual collapse. We propose and test an alternative mechanistic account: that the gap reflects RL-induced selection between conditional policies sharing the same parameters, rather than a loss of underlying capability. We construct a model organism for this account. Two policies — identified solely by the language of an unsupervised chain of thought — are installed via SFT on disjoint splits of ARC-Challenge, with one policy trained to answer and the other to refuse on each split. Reinforcement learning with a final-answer reward, restricted to one split, drives the language of reasoning to collapse onto the rewarded policy across both splits, producing a 67% → 10% accuracy collapse on the held-out split despite no reward signal ever observing it. General capability is unaffected: MMLU accuracy is 71.2% before and after RL. A depth-wise activation steering intervention then recovers held-out accuracy from 0% to 85% on the worst-case slice (Chinese-CoT completions) while preserving Chinese as the surface chain-of-thought language and leaving the rewarded policy intact (83% vs. 94% unsteered). The capability was never lost; it was gated. To the extent deployed regulated-language models contain analogous structure, surface evaluations of such models will systematically understate capability hidden behind learned policy gates.
PaperID: 4540, Poster
Abstract: Multimodal large language models (MLLMs) leverage object-agnostic tokenization by default, which represents images as flat grids of patch tokens and uniformly discretizes semantically rich foreground regions and information-sparse background areas. This introduces substantial redundancy and constitutes a major source of inference cost. Existing object-level tokenizers could either discard intra-region saliency with predefined region grouping and coarse pooling or limit to a single semantic scale by employing slot attention to the final encoder feature map only. In this paper, we propose a two-stage framework named POST that decouples unsupervised object-centric grouping from language-grounded multimodal reasoning to achieve progressively semantic-aligned object-level visual tokenizers. Specifically, we develop Progressive Slot Attention (PSA) in Stage I to introduce slot attention across multiple layers of a frozen ViT encoder and propagates slot states through a cross-layer refinement chain for learning hierarchical slot--patch correspondences without segmentation supervision. Subsequently, we transfer the learned PSA to the LLaVA pipeline in Stage II for object-level visual tokenization. We design slot-weighted merging to aggregate ViT patch tokens into object-grounded visual tokens and remove low-support slots with adaptive slot pruning. Remarkably, POST eliminates the need for external segmentation models, enables progressive object-centric refinement, and adapts visual token budget to image complexity in visual tokenization. Experimental results show that PSA achieves strong unsupervised object-centric segmentation on PASCAL VOC and COCO. Furthermore, POST is comparable or superior to LLaVA-1.5 on most multimodal benchmarks using only 8.3% visual token budget and improves referring expression comprehension on RefCOCO/+/g. It also outperforms recent patch-level and object-level token compression baselines at comparable budgets.
PaperID: 4541, Poster
Authors:
Lehan Yang, Daiqing Qi, Wenhao Zhang, Avery Li, Yiqing Yang, Yifan Li, Yu Kong, Haitian Zheng, Zhifei Zhang, Zhe Lin, Varun Jampani, Sheng LiAbstract: Representation alignment (REPA) accelerates diffusion transformer training, but its alignment targets are almost exclusively semantic encoders such as DINOv2 and CLIP. Recent analysis points to spatial structure, not global semantics, as the carrier of the alignment effect, yet dense-prediction foundation models trained to predict that structure remain overlooked as REPA targets. In pixel-space diffusion, SAM2, Depth Anything v2, and Metric3D v2 each outperform the DINOv2-only GenEval baseline, with the two geometric teachers leading the segmentation teacher. A flat sum of all four teachers, however, lands below the best single geometric teacher, semantic teachers reward viewpoint invariance and instance identity while geometric teachers reward metric layout and surface orientation, and forcing both kinds of gradients through one denoiser projection collapses them into a shared subspace. We introduce PixelDense, which routes DINOv2 and SAM2 through a semantic projection stream, routes Depth Anything v2 and Metric3D v2 through a geometric projection stream, and adds a weight-space orthogonality penalty that keeps the two streams in disjoint subspaces. All four teachers are frozen during training and dropped at inference. Applied to PixelGen and DeCo with a single recipe, PixelDense improves GenEval, DPG-Bench, and HPS v2.1, raises PixelGen-XXL's GenEval Overall from 0.7927 to 0.8093, and beats every single-teacher and unfactored multi-teacher variant. In partial-noise reconstruction, independent panoptic, depth, and surface-normal probes show up to 53.1% PQ gain and 36.0% depth AbsRel reduction at \tau=0.5 across COCO and Flickr30K. From random initialization, PixelDense also reaches the baseline's peak GenEval 1.23× faster.
PaperID: 4542, Poster
Authors:
Karan Singh, Yue Zhao, Mohammadhassan Abbasi, Ehsan AdeliAbstract: We introduce HBAR (Hierarchical Brain Activity Representation), a hierarchical tokenization framework for learning compact, multi-scale representations of neuroimaging data.HBAR organizes voxel-level resting-state fMRI into a hierarchy spanning sub-parcel voxel groups, parcels, and large-scale functional networks, enabling approximately 415x compression while preserving fine-grained spatial information beyond conventional parcellation. The framework supports fixed atlas-based, fully learned, and softly prior-guided hierarchies, allowing us to systematically study the role of structured priors in brain representation learning. Evaluated on large-scale fMRI data, HBAR substantially improves reconstruction over parcellation-based and single-scale tokenization baselines, while supporting downstream modeling of brain dynamics and demographic prediction tasks such as sex classification. Across reconstruction and downstream evaluation, atlas-guided hierarchies provide the strongest overall performance, suggesting that established functional atlases encode organizational structure that remains difficult to recover from reconstruction alone. By combining voxel-level fidelity, compact discrete tokens, and multiscale brain structure, HBAR offers a practical framework for representation learning in high-dimensional neuroimaging data.
Authors: Luxu LIANG, Xiang Li
Abstract: The rapid advancement of large language models (LLMs) has made machine-generated text increasingly difficult to distinguish from human-written text. While recent studies explore leveraging internal representations of language models to uncover deeper detection signals, these raw features often exhibit substantial overlap between classes, limiting their discriminative power. To address this challenge, we propose Steer-to-Detect (\textttS2D), a two-stage framework for detecting LLM-generated text. In the first stage, \textttS2D learns a steering vector that is injected into the hidden states of a frozen observer LLM, producing representations with improved class separability. In the second stage, detection is performed via a hypothesis testing procedure based on the steered representations. We establish finite-sample, high-probability guarantees for Type I and Type II errors, providing a theoretical characterization of the procedure. Empirically, \textttS2D achieves strong and consistent performance across a range of settings, including out-of-distribution scenarios and adversarial perturbations.
PaperID: 4544, Poster
Authors: Jinghe Wang, Xinrui Cao, Duo Wu, Chenghao Gu, Yong Zhong, Linjia Kang, Tianyi Xiong, Zhi Wang
Abstract: Recent advances in robotic manipulation have made rapid progress in mapping visual observations and language instructions to actions. However, reliable manipulation in open-world environments requires robust spatial understanding, the ability to identify targets from complex layouts (spatial reasoning) and adapt to scenes that continuously evolve through interaction (spatial memory). Yet existing embodied benchmarks fall short in evaluating spatial understanding: they often introduce strong visual cues that bypass spatial reasoning, and focus on static scenes that offer little insight into an agent’s capacity for spatial memory. To address this gap, we introduce GeoMind, a systematic benchmark for structured spatial understanding through tabletop manipulation. GeoMind comprises 54 evaluated task variants derived from 27 canonical task designs organized around spatial reasoning and spatial memory. It also incorporates a scalable pipeline that automates the generation and annotation process of task instances with verified targets and executable interactions. Extensive evaluations across diverse embodied agents reveal their consistent weaknesses in spatial understanding. Our findings highlight that these agents struggle to track identities after spatial changes and convert relational or geometric constraints into executable actions.
PaperID: 4545, Poster
Abstract: Discovering hit molecules requires not just high binding affinity, but also the identification of diverse binding modes, which are critical for experimental assays. Existing generative approaches have predominantly relied on optimizing a scalar docking score, obscuring the distinct contributions of key binding determinants. To this end, we introduce a paradigm shift by formulating target-based drug design as a multi-objective exploration task, where each objective explicitly corresponds to enhancing interactions with a specific hotspot. Here, we introduce , a novel generative framework driven by a customized multi-objective reinforcement learning algorithm. By incorporating explorative preferences during training, our approach efficiently uncovers molecules with diverse and desirable binding profiles. Empirical results demonstrate that facilitates the discovery of high-affinity compounds characterized by both structural novelty and diverse binding modes. Validated across target-based drug discovery and multi-property optimization tasks, our approach provides a versatile paradigm for goal-oriented drug discovery.
Authors:
Shivang Rawat, Mirko Morello, Flaviano Morone, David HeegerAbstract: Temporal integration gives continuous-time recurrent networks memory, but in deep stacks it also delays bottom-up signals and attenuates top-down errors. We ask whether this failure mode can be addressed by making the bottom-up input to each layer prospective, replacing instantaneous layer inputs with local look-ahead signals that compensate for integration-induced lag. We develop Recursive Quadrature Filters (RQFs), a biologically motivated class of complex-valued temporal filters that remain equivalent to single-channel diagonal state-space models (SSMs). This equivalence makes RQF layers trainable with the same parallel scan and convolutional algorithms used for diagonal SSMs. At the implementation level, this prospective input amounts to a lightweight two-tap input update, making it a drop-in modification for RQFs and SSMs. For the resulting discrete-time network, we prove that spatial-only backpropagation in a deep RQF network with instantaneous bottom-up inputs yields error signals that decay geometrically with depth, whereas prospective-input coding restores order-one gradient flow. We evaluate prospective-input coding on the Speech Commands dataset using standalone RQF recurrent stacks across two feature representations, mel-frequency cepstral coefficients (MFCCs) and raw audio, and local and non-local credit-assignment strategies. Under spatial-only backpropagation, prospective-input coding improves validation accuracy from 65.7% to 84.9% on MFCC features and from 46.4% to 61.0% on raw audio. Under full backpropagation through time, the gains persist, improving accuracy from 89.4% to 93.2% on MFCC features and from 80.1% to 82.7% on raw audio. Taken together, our results identify prospective-input coding as a local, broadly applicable mechanism for improving credit assignment in multi-layer recurrent networks with a continuous-time substrate.
PaperID: 4547, Poster
Abstract: Agentic ecosystems are emerging as a new platform layer in which agents maintain persistent memory, retrieve from shared knowledge bases, execute through tool interfaces, and coordinate through shared states. These capabilities create a security problem that per-turn defenses miss: compromise can accumulate across sessions, agents, and system layers before visible failure. We define this threat as Agentic Sleeper, a system-level adversary that first behaves legitimately, gains routing trust, and later activates across memory, agent logic, and infrastructure. We formulate defense as inference under partial observability and propose Swarm Shepherd, a system-level defense framework centered on an inference filter for early warning, paired with a backward trace-back module for entry attribution. Our analysis yields a two-scale deployment rule that balances longer event histories against broader agent coverage. Across LangGraph, CrewAI, and AutoGen, Agentic Sleeper drift begins 2.7--8.9× faster than natural baselines, propagates across agents on most topologies, and evades per-turn and stateful defenses above 80%. Swarm Shepherd warns at least three sessions before visible degradation, while trace-back reaches 0.69 top-1 accuracy and 0.72 macro-F1 over 24 variants. The two-scale deployment rule is empirically confirmed, and containment guided by this rule reduces cross-agent propagation below the self-sustaining regime.
Abstract: Despite rapid advances in text-to-image generation, faithfully realizing user intent remains challenging, often requiring manual multi-turn trial and error. To automate this process, existing systems rely on either simple prompt rewriting or closed-loop agents driven by hand-crafted rules, rather than learning to adapt actions to the evolving generation process. In this paper, we reformulate image generation as a state-conditioned action-making problem and propose Generation Navigator, a multi-turn T2I agent that learns to dynamically steer the generation trajectory and output the next action. However, training this agent via reinforcement learning introduces a critical credit assignment challenge: naively rewarding a trajectory based solely on a single state assigns equal credit to all actions in the rollout, ignores the quality dynamics across turns, and fails to distinguish actions that improve the trajectory from those that degrade it or waste turns without progress. We resolve this with PRE-GRPO (Peak-Retention-Efficiency Group Relative Policy Optimization), a trajectory-level reinforcement learning objective that explicitly rewards discovering a high-quality image (Peak), avoiding subsequent quality degradation across turns (Retention), and minimizing unnecessary turns (Efficiency). Experiments show substantial improvements across benchmarks, reaching a WISE score of 0.90 and 79.06% reasoning accuracy on T2I-ReasonBench.
Abstract: Inferring cellular dynamics from unpaired single-cell snapshots requires modeling both state transitions and population growth or death. Unbalanced dynamic optimal transport (UDOT) addresses this by penalizing growth along transport paths, making the choice of growth penalty a key way to encode biological priors on proliferation and apoptosis. However, existing UDOT solvers either rely on computationally expensive NeuralODE simulations or depend on analytical solutions of conditional paths, restricting their efficiency solely to quadratic penalties, i.e. Wasserstein-Fisher-Rao (WFR) geodesics. To enable an efficient UDOT solver for general growth penalties, we first show that concave growth penalties lead to degenerate solutions where growth and transport are separated. We then introduce ptimal transport (SUDO), a simulation-free framework for UDOT with general non-quadratic convex growth penalties. SUDO learns the conditional paths and transport costs, solves the induced semi-coupling problem, and subsequently leverages unbalanced flow matching to achieve a simulation-free solution. On WFR benchmarks, SUDO matches the accuracy of efficient, analytical solution-driven algorithms while outperforming simulation-based methods in computational speed. Beyond WFR, SUDO supports asymmetric penalties that encode proliferation-dominant priors and produce more plausible trajectories and growth estimates on synthetic and single-cell datasets.
PaperID: 4550, Poster
Abstract: Formulaic alpha discovery is a core challenge in quantitative trading, as identifying alphas that work well together remains difficult. Recent reinforcement learning (RL) methods formulate this task as a Markov decision process (MDP), but two important issues remain unresolved. First, as the alpha pool evolves, the reward function changes accordingly, making the MDP inherently non-stationary. Second, most existing methods optimize a single objective, typically predictive power, while ignoring other important properties of a high-quality alpha pool. Motivated by these challenges, we propose AlphaPareto, an RL method for formulaic alpha discovery. To address non-stationarity, AlphaPareto augments the state to include both the alpha under construction and the current alpha pool, and applies a large language model (LLM) to encode the pool. This design allows the agent to adapt to the evolving search environment. To overcome the limitation of single-objective reward design, AlphaPareto replaces the scalar reward with a multi-objective vector-valued reward that simultaneously captures predictive power, temporal stability, perturbation robustness, and diversity, and optimizes these objectives through a Pareto-regularized learning procedure. Empirical applications to real-world datasets show that our AlphaPareto method outperforms its competitors.
Authors: Danna Xue, David Serrano-Lozano, Shaolin Su, Javier Vazquez-Corral
Abstract: 3D Lookup Tables (3D LUTs) are widely used for color mapping, but their grid-based representation requires discretizing the RGB space, leading to a capacity-memory trade-off that becomes prohibitive when storing large numbers of LUTs. Recent approaches adopt implicit neural representations to improve scalability, yet their black-box nature limits interpretability and hinders intuitive, localized editing. In this paper, we propose Gaussian LUT (GLUT), a continuous and explicit color representation that models color transformations using a set of learnable 3D Gaussian primitives. By avoiding fixed-resolution grids, GLUT achieves flexible representational capacity while maintaining a compact memory footprint. Its explicit, spatially localized formulation further enables both accurate modeling and interpretability. Building on this representation, we introduce a compact conditional generator (CGLUT) that predicts GLUT parameters for multiple LUT instances, encoding diverse color styles in a single framework to enable smooth and controllable LUT style blending. Moreover, GLUT supports efficient, user-friendly editing by allowing localized adjustments to specific color regions without global retraining. Experimental results demonstrate that our approach outperforms prior neural LUT representations in both accuracy and efficiency, while offering improved interpretability and interactive control. We are committed to releasing the code and models upon acceptance.
Authors: Shaul Tolkowsky, Ori Meidler, Or Zuk
Abstract: Mode separation, namely how sharply a distribution fragments into barrier-separated clusters, is a fundamental geometric property of densities, difficult to quantify in high dimensions. It is structurally distinct from dispersion, yet existing tools fall short: differential entropy rises with spread regardless of fragmentation, PCA orders directions by variance regardless of barriers, and mutual information requires a mixture decomposition one usually does not have. We measure mode separation through a single stochastic process intrinsic to the density: a unique reversible diffusion with f as its stationary distribution and constant scalar diffusion coefficient. We extract two readouts from its autocovariance matrix: SSA (Sum of Squared Autocorrelations), a scalar barrier-sensitive measure; and DA (Dominant Autocorrelation directions), linear projections ordered by metastability rather than variance. Under an isotropic-Gaussian null, we derive a closed-form spectrum for the empirical autocovariance that generalizes Marchenko--Pastur, with an analytic upper edge that selects the lag at which DA is read off. Both readouts use only samples and a score function, scaling to high dimensions through pretrained score-based generative models via Tweedie's identity. We apply our framework to three settings: (i) synthetic Gaussian mixtures, where SSA tracks mutual information; (ii) SDXL text-to-image generations, where SSA and DA capture structure that entropy and PCA miss; and (iii) molecular dynamics of alanine dipeptide, where DA recovers the known slow backbone dihedrals from static samples alone.
PaperID: 4553, Poster
Authors: Chenhao Zhou, Jiyu Wei
Abstract: Online robust zero-sum Markov games pose a coupled operator-learning and game-solving problem: from nominal interaction data, the learner must recover a worst-case minimax Bellman operator while maintaining strategic coverage and controlling stage-game approximation error. Existing robust reinforcement-learning theory does not yet provide such an operator-learning perspective beyond specialized multi-agent settings. We develop an optimistic dual robust Bellman framework that combines dual robust Bellman fitting, optimistic minimax planning, backward-validated confidence sets, and exploiter-mix data collection. Our analysis is organized around the robust Bellman residuals induced by the candidate value class and yields a regret reduction that separates exploration complexity, robust operator-estimation error, stage-game solution error, and misspecification. Under explicit interpretable assumptions, we obtain conditional theorem interfaces for general and projected function approximation, together with closed results in several structured regimes, including tabular, epochwise/cross-fit linear, self-normalized pure-online linear / finite-rank kernel, and projected RKHS truncation. We also provide numerical experiments that validate the theoretical predictions.
PaperID: 4554, Poster
Authors:
Niklas Koenen, Claudia Battistin, Jeriek Van den Abeele, Martin JullumAbstract: Modern probabilistic machine learning models increasingly produce multivariate outputs with complex dependence structure, from multi-step time-series forecasts to sample path predictions. Understanding which input features drive the predictive uncertainty is important for risk-aware decisions, model diagnostics, and deciding whether the uncertainty should be mitigated or hedged against. This attribution problem requires a choice of how dependencies between output components are treated. Existing approaches reduce the output to a scalar through aggregation or projection before attribution, thereby obscuring whether features affect marginal uncertainty, dependence structure, or both, while component-wise analyses can miss dependence effects entirely. We close this gap by introducing a hierarchy of three entropy-based Shapley games that make this output-side choice explicit for any ordered multivariate outcome, ranging from per-component marginal entropy to fully joint entropy. The hierarchy isolates a cross-component attribution term that captures how each feature shifts the dependence between output components, a quantity invisible to component-wise methods. We establish a chain-rule decomposition of the joint attribution and characterize the cross-component term through conditional total correlation, providing both closed-form and sample-based estimators. Finally, we demonstrate how the framework captures differences in learned joint structure across probabilistic models from distributional regression to a zero-shot time series foundation model.
PaperID: 4555, Poster
Abstract: : an adversary can decompose a harmful objective into individually benign sub-tasks distributed across agents, tools, and shared memory, so that each local interaction appears policy-compliant while the composed workflow is unsafe. Existing defenses classify isolated prompts, monitor single-agent histories, or impose architectural constraints, but none directly detect runtime system-level drift toward harm across the full interaction graph. We introduce atent States), a training-free runtime monitor for multi-turn, multi-agent jailbreaks. CITADEL maintains latent states over agents, tools, and memory stores using a cosine-gated recurrence with no learned parameters; encodes a published safety taxonomy as fixed harm anchors via the same frozen multimodal encoder applied to runtime events; and scores risk through the conjunction of harm-anchor proximity, multi-node participation, and positive temporal drift. We evaluate CITADEL on , a new benchmark spanning five multi-agent attack families, three communication topologies, and heterogeneous API-accessible LLM/VLM backbones. CITADEL reduces average attack success rate from 69.3% to 16.9% — a 75.6% relative reduction — at a 2.1% false-positive rate and 28 ms per-event overhead, outperforming per-agent detectors, trajectory-level monitoring, and architectural defenses. The same hyperparameters transfer without retuning to OpenAgentSafety, Agent Security Bench, and MTMCS-Bench, indicating that collective geometric convergence provides an effective runtime signal for detecting distributed jailbreaks in multi-agent systems.
Authors:
Zhongyi Li, Wan Tian, Jingyu Chen, Kangyao Huang, Huiming Zhang, Hui Yang, Tao Ren, Ruijie Wang, Yijie Peng, Yikun Ban, Fuzhen ZhuangAbstract: Reinforcement learning (RL) has become a key ingredient in post-training large language models (LLMs), with Group Relative Policy Optimization (GRPO) and its variants widely adopted for improving reasoning performance. Despite their success, these methods rely on batch-level reward normalization, where the empirical mean can be severely distorted by noisy, skewed, heavy-tailed, or contaminated rewards. Such distortion directly affects advantage estimation and may lead to unstable or inefficient policy updates. We propose an \emphAdaptive Robust Estimator (ARE), a plug-and-play robustification module for GRPO-style policy optimization. ARE replaces the empirical batch mean with a two-level robust location estimator: adaptive loss minimization suppresses extreme rewards within each block, while median-of-means aggregation limits the influence of corrupted blocks. This design preserves the original policy optimization objective while improving the reliability of reward centering. We prove that ARE is consistent and achieves high-probability deviation guarantees under both finite-variance and heavy-tailed reward distributions. Experiments on mathematical reasoning benchmarks show that ARE improves GRPO-based training in both single-agent and multi-agent settings, with particularly clear gains under noisy and out-of-distribution evaluations. We further validate its generality on embodied vision-and-language navigation tasks, where ARE improves training stability and downstream navigation (VLN) performance.
PaperID: 4557, Poster
Abstract: Aggregate accuracy hides where models succeed and fail. Estimating conditional performance profiles from gold labels alone is expensive, while cheap auxiliary signals such as LLM-judge scores, pairwise comparisons, confidence scores, and judge-disagreement features can be collected for every benchmark item but are often biased or miscalibrated. We propose \method (Local Augmented Control-Variate Evaluation), a semi-supervised estimator for conditional LLM evaluation. The key step is local centering: after subtracting the conditional mean of a cheap signal within the target profile region, any linear augmentation has zero conditional mean and therefore cannot change the estimand. The augmentation coefficient is used only for efficiency, and a local ridge control variate combines a gold-label residual mean from the labeled subset with a cheap-signal mean from the full item pool. We prove calibration-free identification, unbiasedness for grouped profiles, local oracle optimality within centered linear augmentations, and first-order adaptivity to the estimated coefficient. The resulting gain formula depends on a local R^2 that can be estimated from the available labeled subset and cheap signals, diagnosing where cheap signals are useful. We extend the same principle to direct paired model gaps and deployment-weighted scores, and validate it on MATH-500, ScienceQA, MMLU-Pro, and GPQA.
Abstract: Sequential Model-Based Optimization (SMBO) traditionally relies on Bayesian or ensembling surrogates for uncertainty quantification. While historically treated as fully data-driven, SMBO increasingly integrates external domain expertise to accelerate discovery. To overcome the opaque guidance and diminished integration fidelity of standard acquisition re-weighting, Probabilistic Circuits (PCs) have emerged as a generative surrogate alternative, enabling direct knowledge injection via conditional sampling. However, these generative routines lack the formal exploration-exploitation semantics required for rigorous optimization. We introduce the Distributional Conformance Score (DisCo), a principled metric that unifies the flexibility of PCs with a rigorous uncertainty framework. DisCo provides a bounded, [0, 1]-normalized measure of model "surprise" that (1) recovers properties comparable to kernel-based uncertainty known from, e.g., Gaussian Processes while maintaining linear-time inference, and (2) enables principled and accurate assessment of conformance of external knowledge w.r.t. model evidence. We then present DisCoMBO, a framework leveraging these properties for robust, knowledge-aware optimization. We prove that DisCoMBO is a zero-regret algorithm and demonstrate its effectiveness across diverse benchmarks from AutoML, material optimization, and wind park optimization.
Abstract: Multi-objective reinforcement learning (MORL) is a fundamental problem for modern AI, with applications spanning robotics, LLMs, and reasoning. A central challenge is to learn policies that balance competing objectives, such as speed versus energy in robotics, or helpfulness versus honesty in alignment. Optimal policies can vary sharply with the preferred objective weights, and the severity of this variation depends crucially on the degree of misalignment among objectives, making per-preference learning sample-inefficient. To formalize this, we introduce an and use it to derive near-optimal regret bounds for online MORL that quantify the impact of misalignment on learning efficiency. Motivated by this analysis, we propose , an algorithm that maintains a small set of basis policies and dynamically routes trajectories to the policy with the highest potential reward. We further design a scalable, kernel-based attribution algorithm that estimates how feedback signals from one set of trajectories influence policy learning on others. Unlike existing attribution methods, our approach captures nonlinear interactions, such as synergy and antagonism between trajectories, to explain objective conflict. We extensively evaluate our approach across diverse environments and benchmarks: routed ensembles improve average performance by 4.3% over MORL baselines on several environments from the Meta-World and MO-Gymnasium benchmarks, and kernel-based attribution improves correlation with leave-one-out analysis over existing baselines by 50%. When applied to downstream task selection, kernel-based attribution yields a 27% improvement in reward over baseline attribution methods.
Authors:
Kyeongjin Ahn, Seungeon Lee, Krishna Gummadi, Meeyoung ChaAbstract: Geospatial reasoning requires solving image-grounded problems over complex object relationships and scene structures. However, developing this capability is hindered by the cost of annotating a vast and combinatorial question space. We propose GeoX, a self-play framework that acquires spatial logic through executable programs and verifiable rewards in the absence of large-scale human-curated data. Given a satellite or aerial image, our framework employs a single multimodal policy to alternate between proposing spatial problems with executable programs and solving them under three reasoning modes of abduction, deduction, and induction, using a segmentation tool and standard computational libraries. Program execution turns each program into a reward that jointly optimizes the two roles via reinforcement learning. GeoX consistently improves its base VLMs by up to 5.5 points on average, matching or exceeding conventional baselines trained on millions of curated samples. Alongside the framework, we release a benchmark for geospatial reasoning curated from self-play generations.
Abstract: We show that AI agents are capable of discovering novel algorithms for adversarial attacks against LLMs, advancing the state of the art on white-box jailbreaking and prompt injection evaluations. We deploy frontier agents, such as Claude Code and Codex, in an autoresearch loop with access to a library of 30+ prior methods and an evaluation script with a fixed compute budget. We show this pipeline to be effective in jailbreaking OpenAI's GPT-OSS-Safeguard-20B and in prompt injections against Meta-SecAlign-70B, an adversarially robust model. For GPT-OSS-Safeguard, the best agent-discovered method achieves up to 80% attack success rate on CBRN queries, compared to <50% for existing methods. For SecAlign, it achieves 100% ASR, while the best prior automated methods only achieve 82%. Notably, in our setting, attack methods are developed on unrelated surrogate models for a pure random-target token-forcing task, yet generalize directly to prompt injection on the adversarially trained model. Finally, we trace the lineage of methods developed during autoresearch, characterizing the agents' strategies and failure modes. Adversarial ML has long held that defenses must be evaluated against attacks tailored to them; autoresearch automates this principle, and we argue it should be the minimum bar for defense evaluation going forward.
PaperID: 4562, Poster
Abstract: Unlearning in Large Language Model (LLM) is often evaluated primarily by target suppression and clean retain utility. These metrics can miss a local failure mode: a model may perform well on standard retain prompts, yet degrade on benign reasoning when prompts lie near the forgotten domain. This phenomenon is termed forget-adjacent utility loss. A fundamental question is: where should forget representations go after they leave the removed behavior? Successful unlearning depends not only on whether forget-conditioned representations depart from the harmful computation, but also on the direction in which they are redirected afterward. We propose RA^2 (Retain-Anchored Attraction), a latent-space unlearning method that redirects forget-token states toward nearby retain-supported representations while preserving higher-layer behavior on retain data. On WMDP and MUSE, RA^2 yields higher forget-adjacent utility than strong unlearning baselines at similar broad-utility levels. These results show that the migration direction of forget representations materially affects benign behavior near the forget boundary. Our code is available at https://anonymous.4open.science/r/eaxvwertdq.
PaperID: 4563, Poster
Authors: Amir Hossain Raj, Xuesu Xiao
Abstract: Robots manipulating in cluttered spaces must often plan from partial views where foreground objects occlude the target, surrounding geometry, and usable free space. Reactive policies often miss these hidden constraints, while exhaustive search becomes computationally prohibitive in dense clutter. We present SCOUT, an occlusion-aware planning framework that bridges high-level deliberation with grounded physical rollouts. SCOUT leverages an object-centric scene representation to reconstruct hidden geometry and derive a relational planning state capturing blocker hierarchies and accessibility corridors. To navigate the vast action space in clutter, an LLM proposes structured action candidates which are then rigorously evaluated via Monte Carlo Tree Search (MCTS) using a learned action-conditioned world model. Evaluations on ‘Reveal’ and ‘Placement’ tasks show that SCOUT outperforms Vision-Language-Action (VLA) policies, embodied agents, and 3D world models, with significant gains in long-horizon tasks requiring reasoning over occluded geometry. We further validate the framework’s robustness through real-world robotic manipulation in a cluttered shelf environment.
PaperID: 4564, Poster
Abstract: Causal experimental design often assumes direct control over treatment assignment, but many experiments can only assign encouragements that affect treatment receipt indirectly. We formulate this as adaptive instrumental variable (IV) design under noncompliance: the experimenter chooses unit--instrument queries, observes stochastic treatment realizations and outcomes, and seeks the treatment-level conditional average treatment effect (CATE). This creates an acquisition mismatch: uncertainty in prospective feedback does not necessarily correspond to uncertainty in the target causal estimand, so outcome-predictive or marginal variance-based criteria may select queries weakly informative for CATE. We introduce instrumental variable information gain (IVIG), a Bayesian acquisition principle that scores queries by the expected information their joint treatment-realization and outcome feedback provides about CATE values over a target population. Under the working Bayesian posterior, IVIG is Bayes-risk aligned under logarithmic loss; we also characterize the residual information discarded by outcome-only acquisition and connect greedy IVIG-style acquisition to target posterior-variance reduction under a fixed-covariance analysis. To approximate IVIG, we use an empirical-Bayes Gaussian Process IV posterior approximation with first-stage treatment-realization modeling and treatment-realization-specific fantasy updates. Experiments on synthetic and two semi-synthetic IV benchmarks show improved CATE sample efficiency over adaptive-design baselines; a real-label diagnostic shows improved recovery of full-data CATE estimates from observed IV data.
PaperID: 4565, Poster
Abstract: Prompt tuning (PT) provides a parameter-efficient way to adapt frozen large language models, but its behavior becomes less clear when task contains multiple related yet distinguishable patterns. We formalize this setting as a latent mixture target distribution, where the target behavior is modeled as a finite mixture of latent task components. We first show that static prompting can retain a non-vanishing approximation gap when one shared prompt is insufficient to capture the target distribution. Then we formalize instance-specific prompting as an input-dependent prompt mapping and show that, under ideal prompt assignment, its achievable regret can approach zero. Guided by this analysis, we propose POISE (Prompt Optimization via Instance-Specific Experts), a structured and parameter-efficient approximation to instance-specific prompting. POISE represents the prompt space with learnable expert prompts, combines them through input-dependent routing, and introduces a shared low-rank residual correction to capture common structure. Extensive experiments across multiple benchmarks show that POISE achieves strong performance while maintaining high parameter efficiency.
Authors:
Liuwenpu, Yuqi Xu, Weichu Xie, Yongfu Zhu, Shuai Dong, Ziyue Wang, Wenqi Shao, Xiaoying Zhang, Tong Yang, Nan Duan, Jiaqi WangAbstract: Reinforcement Learning from Verifiable Rewards (RLVR) typically samples multiple responses per prompt and assigns binary rewards based on individual correctness, yet the collective structure of the group output, specifically, the distribution of errors is largely discarded. We identify this as a missed opportunity: empirical analysis reveals that error diversity within a group is a strong predictor of training success, with problems eliciting diverse wrong answers benefiting substantially more from RLVR than those producing homogeneous failures. Motivated by this observation, we propose Error Diversity Advantage Shaping (EDAS), a lightweight, algorithm-agnostic technique that modulates the advantage signal for incorrect rollouts based on intra-group error diversity. EDAS amplifies penalties for dominant, repeated errors and attenuates penalties for rare, exploratory ones, thereby encouraging the model to maintain diverse reasoning paths and discouraging error perseveration. Crucially, EDAS operates as a simple post-hoc adjustment that can be seamlessly integrated into any RLVR algorithm. We validate EDAS on top of several mainstream RLVR methods across a series of models and seven challenging math benchmarks, demonstrating consistent improvements. Notably, EDAS yields an average improvement of 6.29 points over DAPO on Qwen3-8B across seven benchmarks, confirming that exploiting the latent information in group rollouts is a broadly effective strategy for strengthening RLVR.
PaperID: 4567, Poster
Abstract: We study the spectral structure of the attention matrix in Transformer models through an idealized spectral model for the pre-softmax score matrix. Motivated by the rotational isotropy of query and key representations, we replace the random token-space orientation of the score matrix by its canonical singular-value representative \Sigma, and assume that \Sigma has a logarithmic spike-bulk separation. We provide empirical diagnostics consistent with this model on pretrained vision and language transformers: token-space singular vectors exhibit near-Haar behavior, and score spectra are dominated by a small number of effective directions. We then prove that the row-wise softmax map transforms this structure into a degenerate attention matrix: dominant modes are preserved, whereas the bulk spectrum converges weakly to \delta_0. This provides a unified theoretical explanation for why attention layers can act as low-rank information extractors: they preserve a small number of dominant token directions while compressing the remaining bulk modes.
PaperID: 4568, Poster
Abstract: Neural subgraph matching (NSGM) models are widely used for graph retrieval, where relevance labels for query–corpus pairs are determined by whether the query graph is a subgraph of the corpus graph. Despite widespread applications, the adversarial robustness of such models remains underexplored. In principle, adversarial attacks should induce minimal shift of the instance distribution and the target labels. However, in NSGM, a single edge insertion or deletion can flip relevance labels. Detecting or controlling label flips requires repeated calls to potentially expensive combinatorial solvers. Responding to this challenge, we propose SEAGRAM, a solver-efficient adversarial attack method. It learns to detect non-critical node pairs whose perturbations will likely preserve relevance labels, thereby avoiding solver calls during test time. We further employ an active learning strategy, which reduces the solver calls also during training. Experiments show that SEAGRAM significantly worsens the performance of existing models.
Authors: Lénaïc Chizat, Maria Colombo, Roberto Colombo, Xavier Fernández-Real
Abstract: Stein Variational Gradient Descent (SVGD) is a deterministic interacting-particle method for sampling from a target probability measure given access to its score function. In the mean-field and continuous-time limit, it is known that the flow converges weakly toward the target, but no quantitative rate is known for the last iterate. In this paper, we establish quantitative local convergence in strong norms for this dynamics, when the interaction kernels is of Riesz type on the d-dimensional torus. Specifically, assuming that the initial density and the target are smooth and close in L^2-norm, we obtain explicit polynomial convergence rates in L^2-norm that depend on the dimension and on the regularity parameters of the kernel, the initialization and the target. We further show that these rates are sharp in certain regimes, and support the theory with numerical experiments. In the edge case of kernels with a Coulomb singularity, we recover the global exponential convergence result obtained in prior work. Our analysis is inspired by recent results on Wasserstein gradient flows of kernel mean discrepancies.
Abstract: In this paper, we introduce Exact Flow Linear Attention(EFLA), an exact-flow formulation of delta-rule linear attention. We show that the delta-rule update can be interpreted as an explicit Euler discretization of an underlying continuous-time system. EFLA replaces this first-order update with the exact closed-form flow. By exploiting the rank-1 structure of the dynamics matrix, both the matrix exponential and the input integral collapse to a simple update that preserves delta-rule linear attention's algebraic structure, parameter count, linear-time complexity, and chunkwise parallelism. This attention mechanism removes the Euler discretization error of the delta-rule dynamics without introducing additional parameters. Experiments on robustness tests, language modeling benchmarks, and the MAD synthetic benchmark show that EFLA improves stability under corrupted and high-energy inputs, reduces perplexity, and achieves stronger downstream performance compared to SSM and Euler-style baselines. These results establish exact-flow integration as a principled and scalable update mechanism for delta-rule linear attention.
Authors: George Bird
Abstract: Introduced is a methodology for adapting the topology of dense neural networks, enabled by isotropic activation functions. Achieved through prescribed reparameterisation symmetries and singular-value decomposition of affine maps, this diagonalises layers into one-to-one, ordered connections. This makes it simpler to assess the impact of individual connections on the function. Low-impact neurons can be removed (neurodegeneration), and a thresholded buffer of largely inactive 'scaffold' neurons is maintained (neurogenesis). These symmetry-led diagonalisation and structural changes are function-invariant, demonstrated to be computationally identical during neurogenesis, arbitrarily well approximated during neurodegeneration, and enable asymptotic 50% parameter sparsification of dense networks with identically preserved function. Thus, real-time restructuring of the architecture in response to task demands, task appending, removal or changes is shown. The approach is conceptually centred on primitive symmetry-prescriptions, through which isotropic functions are derived that feature explicit basis independence and a loss in the individuation of neurons implicit in typical elementwise functional forms. Hence, this allows freedom in the basis to which layers are decomposed and interpreted as individual artificial neurons, directly enabling this adaptive topology approach. Additionally, a new tunable model parameter, the 'intrinsic length', is introduced to improve this analytical invariance, alongside a generalised isotropic-perceptron architecture that enables parallel precomputation of all matrix-vector products and displays a nested functional class. Diagonalisation is suggested to offer new possibilities for interpretability and monitoring of isotropic networks.
PaperID: 4572, Poster
Abstract: Vision-language models (VLMs) for radiology have emerged as a scalable paradigm by leveraging image-report pairs naturally produced in clinical workflows. However, this pairing reveals a mismatch in scale: each finding occupies only a small region of the image, yet supervision is provided only at the global image-report level. This poses a central challenge: prior approaches spread weight densely across all patches rather than concentrating on the sparse subset relevant to a given query. To address this, we present GLINT (Gated Language-Image alignmeNT), a framework that explicitly models this sparse correspondence. On the alignment side, we introduce Sparsely Gated Alignment, a novel architecture in which a sigmoid gate over a separate gate embedding space activates only the patches relevant to each textual query, enforcing explicit sparsity. On the representation side, we add Dense Feature Regularization, which anchors the trainable encoder's intermediate features to a frozen self-supervised learning (SSL) teacher, preserving the fine-grained patch features that the gate relies on. The same recipe applies to both 2D chest X-ray (CXR) and 3D chest computed tomography (CT), built with DINOv3 and V-JEPA 2.1, respectively. GLINT enables zero-shot classification, grounding, and segmentation from free-text queries, and to our knowledge is the first to demonstrate zero-shot segmentation on 3D CT volumes without mask supervision. Notably, the most pronounced gains arise on zero-shot grounding and segmentation, where sparse, query-specific localization is required, consistent with our design intent. In downstream evaluation, GLINT outperforms both SSL encoders and medical VLMs on classification, report generation, and segmentation.
PaperID: 4573, Poster
Abstract: Federated Graph Learning is rapidly evolving as a privacy-preserving collaborative approach for decentralized graph data. However, severe fairness challenges are increasingly undermining federated systems by systematically degrading predictions for structurally disadvantaged minority nodes. The inherent vulnerabilities and missing topological contexts in Federated Graph Learning are deeply entangled, making traditional federated fairness methods and simple oversampling less effective. In our work, we propose an effective Evidential Energy-based Generation framework for Fair Federated Graph Learning (E^2Gen). At the client level, it explicitly identifies structurally deficient nodes using a multi-axis metric, synthesizing targeted representations via conditional energy-based models, and selects reliable samples through an evidential quality gate. At the server level, the local performance disparities uploaded by each client are evaluated to construct a fairness gap assessment, making the global model absorb equitable improvements by further adjusting the aggregation weights. Our method can handle high topological heterogeneity, does not require strict generative normalization, and is effective under both homophilic and heterophilic graph structures. Extensive results on various settings of federated graph scenarios under severe fairness challenges validate the effectiveness of this approach. The code is anonymously available at https://anonymous.4open.science/r/E-Gen-A2C6.
PaperID: 4574, Poster
Abstract: Scaling Vision-Language-Action (VLA) policy training requires diverse robot demonstration data, yet collecting such data remains expensive and labor-intensive. The community has explored ways to accelerate robot data collection: visual augmentation methods synthesize novel samples by editing visual elements in existing demonstrations, while systems such as UMI leverage wrist-mounted cameras to flexibly collect in-the-wild manipulation videos. However, the mismatched observation setups of these two directions leave them largely disconnected, limiting their synergistic potential for scalable robot data generation. To bridge this gap, we propose , a generative framework for controllable wrist-to-exocentric conversion of robotic demonstrations. We introduce a geometry-motion factorization strategy that reconstructs canonical object-robot geometry and propagates motion states from noisy wrist-view videos, yielding reliable structural priors for cross-view synthesis. These priors guide a conditional video diffusion model to generate target-view exocentric videos that are spatially aligned, temporally coherent, and faithful to the underlying manipulation semantics. We evaluate RoboExo on a wrist-to-exo generation benchmark under both seen and unseen settings, where it consistently outperforms all baselines, reducing LPIPS by . Beyond generation quality, RoboExo improves VLA policy learning across six tasks, increasing the average success rate by with generated exocentric views. These results demonstrate RoboExo as an effective and scalable data engine for enriching low-cost UMI demonstrations. The code will be publicly available.
PaperID: 4575, Poster
Abstract: Cinematic camera motion is a fundamental storytelling tool, defined not only by where the camera is positioned in the scene, but also by how it moves in terms of direction and speed. Recent work on camera trajectory generation and alignment to text relies on pose-centric representations. While in principle a network could derive direction of movement and speed, we find that in practice this might not happen. In fact, in this paper we discover that decomposing the camera trajectory representation from the traditional per-frame poses to direction and speed has surprising benefits across multiple tasks, including trajectory-to-text alignment as well as text-to-trajectory generation. To accurately evaluate the former, we introduce a simple and reliable protocol that overcomes the limitations of prior evaluation baselines. For the latter, building on this representational insight, we propose a novel generative model for camera trajectories, CineGen, that achieves superior performance across a variety of metrics. We also propose a novel dataset, CineScript, containing movie clips that are enriched with scene descriptions as well as higher-level metadata. This novel data allows us to test models' ability to capture high-level cinematographic information. We show that, despite its simplicity, representing camera trajectories through direction and speed not only helps numerically to achieve better alignment and generation, but also inherently encodes complex directorial intent.
PaperID: 4576, Poster
Abstract: Pre-trained large language models exhibit strong exploration ability, as evidenced by high pass@K scores, yet this ability degrades significantly during reinforcement learning (RL) fine-tuning due to mode collapse. A natural remedy is to impose token-wise KL divergence between the fine-tuned and base models; however, we argue that such objective introduces a fundamental mismatch, where token-wise KL operates locally at individual token positions, while exploration is a global behavior that emerges over sequences as a whole. Motivated by this insight, we propose Mean Order-Statistic Alignment (MOSA), which preserves exploration by aligning the fine-tuned model's global behavior to that of the base model. Specifically, MOSA captures global exploration through the mean order-statistic profile, which is obtained by computing the order statistics of each token's posterior over the vocabulary and averaging them across tokens at each rank position. The profile, designed a valid probability distribution, directly admits a principled KL-based objective. For efficiency, we further approximate the profile using only the top-K probabilities and a residual tail bucket, yielding a compact implementation with minimal computational overhead. Experiments on Countdown, 6 mathematical reasoning, and 2 coding benchmarks show that MOSA discovers more diverse plausible solutions, improves both pass@1 and pass@K, and scales effectively to models up to 32B parameters.
PaperID: 4577, Poster
Authors:
XIN Li, Junquan Huang, Xujia Li, Lei ChenAbstract: Large language models (LLMs) have shown strong reasoning capabilities in embodied planning tasks. However, they often fail to achieve good results on long-horizon and multi-robot tasks under strict physical constraints, and generate smooth but infeasible trajectories. In real-world robotic hardware deployment, systems also face two fundamental challenges: First, the fine-tuning method for open-source models, Direct Preference Optimization (DPO), typically only makes a binary true or false judgment on the planning results, without distinguishing the severity of the planning error. As a result, they fail to impose effective penalties for catastrophic violations of physical constraints. Second, methods that search over multiple candidate plans during inference do not filter out invalid plans in advance, often leading to combinatorial explosion and high inference latency, which makes them difficult to apply to real-time robotic operations. To address these limitations, we propose a complete training-to-inference architecture consisting of two modules. First, we introduce Margin-Aware Direct Preference Optimization (mDPO) during training, which distinguishes the severity of errors in negative samples, enabling the fine-tuned model to avoid generating severe errors that would render a task irreparable. Second, we deploy Verifier-Guided Search (VGS) during inference, using a symbolic verifier to hard-prune invalid search paths, thereby repairing infeasible task plans and reducing inference latency. In addition, latency analysis shows that our approach reduces the average end-to-end runtime from 12.814s to 10.969s compared with retry-based repair, while maintaining higher final success, making it more suitable for time-sensitive robotic deployment.
PaperID: 4578, Poster
Abstract: Geometric modeling has become an essential paradigm in AI for Science, where data often inhabit intrinsically curved or structured spaces. However, existing manifold generative models often rely on closed-form or tractable access to global geometric primitives, such as geodesic distances, parallel transport, or spectral decompositions, which are rarely available beyond canonical geometries. To overcome this limitation, we propose Horizontal Diffusion Models (HDMs), a score-based generative framework defined on the orthonormal frame bundle of a Riemannian manifold. Instead of requiring global geometric operations, HDM lifts Euclidean diffusion processes through the Levi-Civita horizontal distribution, using only local metric and connection information while preserving intrinsic manifold geometry. This construction allows standard Euclidean score networks to be lifted into geometry-consistent and gauge-equivariant horizontal vector fields, bypassing the need for manifold-specific neural architectures. We further derive a horizontal KL objective that reduces to a Euclidean score-matching loss and analyze the curvature-dependent cost of the lift through a generalization bound. Experiments on parametric surfaces and scientific datasets demonstrate high-fidelity generation across diverse manifolds with nontrivial geometry and topology.
PaperID: 4579, Poster
Abstract: Open-vocabulary semantic segmentation (OVSS) enables pixel-level prediction over arbitrary text-specified vocabularies and has shown strong generalization on common benchmarks. However, OVSS performance often degrades in specialized domains such as medical imaging, remote sensing, and industrial inspection, where dense pixel-level masks for adaptation are costly to obtain and require domain-specific expertise. We propose a preference-guided adaptation framework that replaces dense mask supervision with binary preferences. We observe that different prompt templates produce systematically different segmentations for the same image, a phenomenon we call prompt disagreement, and we repurpose it as a built-in source of preference supervision. Building on this, we mine localized preference queries from regions of high cross-template uncertainty, and adapt the OVSS model with Region-Localized Preference Optimization (RLPO) together with consistency regularization that stabilizes updates outside the queried region. Across extensive experiments on the MESS benchmark, the proposed method achieves consistent gains across diverse OVSS backbones without any pixel-level annotation, and remains effective under noisy preferences.
PaperID: 4580, Poster
Abstract: Traffic flow forecasting plays a pivotal role in intelligent transportation systems. However, most existing methods adopt a black-box learning paradigm, resulting in a lack of interpretability in their decision-making processes. Existing explainable methods mostly follow the feature attribution paradigm, leaving explanations at the level of static or discrete feature evidence and making it difficult to reveal the continuous evolution mechanism of spatio-temporal dependencies during dynamic traffic propagation. Meanwhile, prior methods often decouple spatial and temporal explanations, treating spatial structures and temporal dynamics independently and thus failing to capture the intrinsic coupling between spatial dependencies and temporal contexts in model decision-making. To this end, we propose STGIB, a new spatio-temporal graph information bottleneck theory for characterizing continuously evolving explanations in traffic forecasting. STGIB redefines traffic explanation as a time-indexed spatio-temporal explanation trajectory, and jointly learns dynamic key structures and their temporal representations under predictive sufficiency and information compression constraints. Furthermore, we derive tractable recursive variational bounds to make the theoretical objective optimizable, and instantiate a new explainable traffic flow forecasting model via spatio-temporal explanation trajectory generation. Experiments on real-world traffic datasets demonstrate that STGIB maintains competitive predictive performance while generating more faithful and spatio-temporally consistent explanations.
PaperID: 4581, Poster
Abstract: We study the sample complexity of noisy one-bit compressed sensing for signals drawn from a prior distribution. By characterizing the effective distributional complexity of the prior via its approximate covering number, we prove that posterior sampling achieves accurate recovery with high probability when the number of measurements scales with the logarithm of the approximate covering number, up to a one-bit separation gap factor. This upper bound is robust to learned prior mismatch. Specifically, we show that posterior sampling with an approximate prior remains reliable, provided that the learned prior distribution is sufficiently close to the true signal distribution in Wasserstein distance. In addition, we establish a sample complexity lower bound for any reliable method of noisy one-bit compressed sensing, showing that our upper bound is nearly matched in its main prior dependent term. To approximate the ideal posterior sampling process for real world scenarios, we instantiate posterior sampling through a plug-and-play algorithm with diffusion priors. Experiments on the FFHQ and ImageNet datasets demonstrate the effectiveness of our proposed approach.
PaperID: 4582, Poster
Authors: Reinier W Vos, Aurora Micheli, Nergis Tomen, Christophe De Wagter, Jan van Gemert, Sander Bohte, Guido de Croon
Abstract: Parameter initialization shapes training dynamics, convergence, and performance in neural networks. In spiking neural networks (SNNs), analytical initialization is complicated by three additional mechanisms—binary spike outputs, temporal recurrence (leak), and reset-induced drops in membrane potential—which together suppress activity propagation: (1) over time, where recurrence and reset reduce individual neurons' firing rate, and (2) across depth, where deficits in upstream activity compromise downstream layers, known as spike-rate collapse. Prior initialization methods address only subsets of these challenges. We derive closed-form weight and bias distributions for initialization that jointly account for temporal recurrence, refractory reset, and signal propagation across depth. The method is data-agnostic: it requires only the architecture and neuron hyperparameters. Our derivation centers around a bias correction that yields stationary firing rates and membrane variance, with a principled default spike probability of p = 0.5 that maximizes per-neuron Bernoulli entropy—counteracting dead neurons (p \approx 0) and saturated neurons (p \approx 1) at initialization. In controlled simulations across 30 layers, it maintains target firing rates and membrane variance at the individual-neuron level, unlike alternatives tested. Across UEA multivariate time-series classification, neuromorphic classification, and continuous-control reinforcement learning, our SNNs rank first in every paradigm without per-paradigm tuning.
PaperID: 4583, Poster
Authors:
Fanqi Yu, Matteo Tiezzi, Cigdem Beyan, Tommaso Apicella, Vittorio MurinoAbstract: Robotic agents operating in real-world environments must adapt continuously from a stream of multimodal demonstrations, often under constraints that preclude storing or revisiting past data. In this replay-free, single-pass lifelong imitation learning setting, sequential updates can progressively distort the representation space, making expert-based approaches particularly sensitive to routing interference, causing expert selection to degrade over time. We propose HIDRA, a structured routing framework that mitigates this issue through a hierarchical dual-attention mecha8 nism, which decouples instruction-level expert retrieval from context-dependent refinement. To further stabilize routing under representation drift, HIDRA introduces a key-level geometric regularization that enforces both alignment within tasks and separation across tasks. The resulting approach enables reliable expert selection while supporting an expandable set of experts for continual adaptation. We evaluate HIDRA on several LIBERO benchmark suites, including more challenging variants with heterogeneous task sequences and paraphrased language instructions. Our approach consistently outperforms both replay-free and replay-based baselines, improving AUC and forward transfer while maintaining low forgetting, with stronger gains under high interference and language variation. Code will be released.
Abstract: Real-world digital environments are highly diverse and dynamic. These characteristics cause agents to frequently encounter unseen environments and distribution shifts, making continual learning in such environments essential for computer-use agents (CUAs). However, a key challenge lies in obtaining high-quality and environment-grounded training data without relying on costly human annotation. In this work, we introduce earning framework that continually adapts agents to specific environments with zero human data. The agent first explores an environment to acquire initial experiences. During subsequent iterative training, a curriculum task generator leverages these experiences together with feedback from the previous iteration to synthesize new tasks tailored for the agent's current capabilities. To provide reliable reward signals, we introduce CUAJudge, a robust automatic evaluator for CUAs that achieves 93% agreement with human judgments. Empirically, our method effectively enables both intra-environment and cross-environment continual learning, yielding 3–29% absolute performance gains on the target environments without catastrophic forgetting on others. We also show that it can mitigate performance degradation under environment changes (e.g., version updates, platform migration, and resolution shifts). Further analyses show highly sparse updates (e.g., only 20% parameters), which helps explain the effective and robust adaptation.
Abstract: Retrieval-augmented generation (RAG) typically treats retrieval and generation as separate systems. We ask whether an attention-based encoder-decoder can instead retrieve directly from its own internal representations. We introduce INTRA (INTrinsic Retrieval via Attention), a framework where decoder attention queries score pre-encoded evidence chunks that are then directly reused as context for generation. By construction, INTRA unifies retrieval and generation, eliminating the retriever-generator mismatch typical of RAG pipelines. This design also amortizes context encoding by reusing precomputed encoder states across queries. Across question-answering benchmarks, INTRA outperforms strong engineered retrieval pipelines on both evidence recall and end-to-end answer quality. Our results demonstrate that attention-based models already possess a retrieval mechanism that can be elicited, rather than added as an external module.
Abstract: Robotic foundation models require reasoning over complex visual scenes to execute adaptive actions in dynamic environments. While recent studies on latent-reasoning Vision-Language-Action (VLA) models have demonstrated the capability to capture fine-grained physical dynamics, they remain predominantly confined to static imitation learning, severely limiting their adaptability and generalization. In this paper, we present LaST-R1, a novel reinforcement learning (RL) post-training framework designed to effectively harness latent reasoning-before-acting policies. Specifically, we propose Latent-to-Action Policy Optimization (LAPO), a core RL algorithm that jointly optimizes the latent reasoning process and the action generation. By explicitly embedding latent Chain-of-Thought (CoT) reasoning directly within the RL optimization loop, LAPO stimulates profound physical world modeling, which in turn drives robust execution in interactive environments. Furthermore, an adaptive latent CoT mechanism is introduced, allowing the policy to dynamically modulate its reasoning horizon based on diverse environment states. Experiments show that LaST-R1 achieves a near-perfect 99.9% average success rate on the LIBERO benchmark with only one-shot supervised warm-up, significantly improving convergence speed and performance over prior state-of-the-art (SOTA) methods. In real-world deployments, LaST-R1 yields up to a 22.5% average improvement over SOTA supervised fine-tuning approach across four complex tasks, including both single-arm and dual-arm settings. Finally, LaST-R1 demonstrates strong generalization across simulated and real-world environments.
Abstract: Hybrid language models that mix attention and recurrent layers have shown promise: theoretically, recurrent layers ameliorate the limitations of pure transformers on state tracking, and empirically, hybrids can outperform pure transformers in loss and downstream evaluations \citepwaleffe2024empirical,merrill2026olmohybrid. Yet it remains unclear which data or capabilities drive these gains, and to what degree they reflect the theoretical advantages motivating hybrid models. We address this by analyzing the benefits of hybrid models at the token level---which kinds of tokens are better predicted by hybrids than by pure transformers---using the open weights from Olmo 3 \citepolmo2025olmo3 and Olmo Hybrid \citepmerrill2026olmohybrid. Hybrids predict tokens better across the board, but we show these gains localize to open-class content words rather than closed-class function words. Across natural language and code, the loss gap shrinks on closing brackets (but not opening brackets), consistent with the hypothesis that attention, not hybrid layers, is largely responsible for processing hierarchical structure. The hybrid advantage also vanishes on repeated n-grams. Overall, while attention suffices for tokens involving copying and purely syntactic information, hybrids show an advantage on more semantically conditioned predictions, likely because these are aided by a strong representation of the discourse state. We conclude with discussion and proof-of-concept experiments showing how these findings could refine design principles and pretraining evaluations for hybrid architectures.
Authors: Zhen Zhang, Amr Alanwar
Abstract: Probabilistic prediction heads in neural networks typically output either a Gaussian mixture or a single conformal region. Neither separates the distinct sources of uncertainty often present in real prediction tasks: a discrete choice among modes, bounded systematic drift within the chosen mode, and irreducible stochastic noise. We introduce the Hybrid Probabilistic Zonotope (HProbZ), an output head that represents these three sources as binary, bounded, and stochastic generators of a zonotope, and admits a closed-form likelihood by convolution. Sharing the bounded generator across prediction steps couples future predictions algebraically, so observing one step refines the predictive distribution at every remaining step in a single forward pass. We establish that the three generators are identifiable from the likelihood up to permutation, and that an HProbZ density is representationally distinct from any finite Gaussian mixture. The same shared structure provides analytic per-mode risk and distribution-free multi-modal conformal sets at inference time. Empirical analysis on representative prediction benchmarks supports the effectiveness of the design relative to same-encoder mixture baselines, while offering structural properties that mixture or convex-conformal predictors do not jointly provide.
PaperID: 4589, Poster
Abstract: Accurately modeling soft boundaries, e.g., hair and defocus blur, is a fundamental challenge in stereo conversion due to the ambiguous blending of foreground and background. Existing depth models primarily predict single-layer depth, leading to ambiguity in depth correspondence at soft boundaries. While matting techniques can capture opacity for layered modeling, they often struggle in complex scenes with multiple targets and usually require user intervention. This paper introduces \alphaDepth, a layered representation that decomposes soft boundaries for high-fidelity stereo conversion. Specifically, we first resolve mixed color and depth ambiguity by estimating layered color and depth values at soft boundaries. Considering complex multi-target scenes, we design a Circular Alpha Representation (CAR) that shifts the paradigm from global target extraction to local boundary decomposition. Unlike prior matting methods restricted to a single foreground/background, CAR enables efficient scene-level inference without manual guidance. Extensive evaluations demonstrate that \alphaDepth achieves state-of-the-art performance in stereo conversion, eliminating background bleeding and structural distortions at soft boundaries.
PaperID: 4590, Poster
Abstract: Federated fine-tuning of large language models (LLMs) under heterogeneous client capacities is increasingly important, and low-rank adaptation (LoRA) with heterogeneous ranks makes it practical. Existing server-side aggregators compose local matrix updates \Delta W_i from client-specific LoRA factors (A_i,B_i) before aggregating them. Although these matrix updates lie in the same weight space \mathbbR^m× n and can be averaged directly, the LoRA factors that generate them may expressed client-specific local basis. In the typical compose--aggregate--factorize--dispatch loop, this basis inconsistency can make the rank-r_i reference prefix dispatched to lower-rank clients unstable after server-side refactorization. In this work, we propose FedAbA, a basis-consistent federated LoRA aggregation framework for heterogeneous ranks. FedAbA first extracts each client's local basis, aligns it to the corresponding prefix of the global reference basis, and constructs a basis-consistent reconstruction of each local matrix update before aggregation. It then aggregates these reconstructions with weights that combine client data size and consistency score, followed by exact coefficient-space SVD refactorization and gauge fixing before rank-prefix dispatch. Our theoretical analysis clarifies the role of basis consistency and supports the proposed consistency-aware aggregation mechanism. Extensive experiments show that FedAbA outperforms representative federated LoRA baselines, with the largest gains under the strongest rank heterogeneity.
PaperID: 4591, Poster
Abstract: Prompt learning is a popular parameter-efficient method for adapting foundation models, but learned prompts are typically task-specific and fail to generalize to new classes, domains, or compositions of tasks. In this paper, we introduce a Diffusion Meta-Prompt (DMP) model , a framework that models the distribution of learned prompts using diffusion models. DMP is trained only on a repository of previously learned prompts to synthesize new prompts conditioned on natural language task descriptions, without access to the task data. To improve the sampling stability, we introduce a test-time steering strategy for DMP, which uses the best training-selected prompt in the repository as a latent anchor during diffusion sampling, without retraining the DMP or accessing test classes. DMP improves generalization across classification, retrieval and text-to-image generation tasks, supports concept composition and negative prompting without explicit training. It reduces storage and inference costs by over 90% compared to prompt retrieval methods. For composite classification, DMP achieves upto 2.0% average gain over prior meta-learning methods across 55 pairs of datasets with gains as high as 8.5% on specific pairs such as Eurosat and Flowers. DMP also enhances cross-task generalization with ~2-9% improvement for hierarchical classification task. We further provide a theoretical guarantee bounding the expected task loss of prompts sampled from a DMP.
Abstract: We introduce VideoMDM, a diffusion-based framework that trains 3D human motion priors directly from accurate 2D poses extracted from monocular videos, without any 3D ground truth. A pretrained 2D-to-3D lifter provides approximate 3D pose sequences that serve as a noisy teacher: these are diffused, denoised by the model in 3D, and supervised in 2D by reprojecting the prediction and comparing against accurate keypoints. We show that, under mild assumptions, a depth-weighted 2D reprojection loss is equivalent in expectation to direct 3D supervision, and we adapt standard 3D motion regularizers — velocity consistency and over-parameterized representation alignment — to this 2D setting. Unlike methods that lift 2D to 3D only at inference, VideoMDM learns a coherent 3D motion manifold during training. On HumanML3D it nearly closes the gap to fully 3D-supervised MDM (FID 0.88 vs 0.54); On real video datasets Fit3D and NBA the method learns to generate motions consistently preferred by humans, with strong quantitative results.
PaperID: 4593, Poster
Authors:
Shaojiang Wang, Geer Yang, Wenbo Wang, Bin Wu, Pengcheng WangAbstract: Deploying Large Language Models on heterogeneous GPU clusters is often bottlenecked by rigid parallelism strategies that fail to fully exploit hardware asymmetry. To address this, we introduce ReVar, a holistic framework that optimizes replica placement and activation routing by treating replica counts as primary decision variables for all model components, including both dense stages and MoE experts. By modeling inference as a coupled resource-allocation and flow-routing problem, ReVar replaces static pipelines with fluid execution graphs that utilize uneven replication and multi-path routing. Supported by rigorous theoretical foundations, including proving the problem's NP-hardness, providing a MILP formulation for throughput upper bounds, and deriving closed-form solutions for constrained regimes, ReVar delivers significant practical gains. In evaluations, it achieved the highest throughput in 14 out of 15 settings, demonstrated up to a 3.1 times improvement in steady-state throughput in mixed-GPU environments, and was successfully validated on a real physical heterogeneous cluster.
Authors: Ningkang Peng, Jingyang Mao, Xiaoqian Peng, Qu Weiguang, Yanhui Gu
Abstract: Noisy-label methods often estimate sample reliability from forward-space signals such as loss, confidence, or entropy. These signals indicate whether a sample is difficult to predict, but they do not directly test whether its observed label induces a reliable parameter update. This gap matters because hard clean samples and mislabeled samples can have similar loss while inducing different updates. We recast reliability estimation as diagnosis of the observed-label update. The sample-wise empirical Fisher trace gives a backward-space measure of update energy: for the classifier layer, it factorizes into a prediction-residual term and a feature-sensitivity term, so it captures information beyond scalar loss. Trace, however, is still a radial magnitude signal and cannot decide whether a large update is useful or harmful. We therefore propose Relative Geometric Conflict (RGC), which compares the observed-label gradient with a reference gradient induced by an EMA teacher. The conflict term helps distinguish large but aligned hard-clean updates from large conflicting updates caused by corrupted labels. Across synthetic and real-world noisy-label benchmarks, RGC improves hard-clean preservation and accuracy under our evaluation protocol.
PaperID: 4595, Poster
Abstract: Ensuring robust safety alignment of large language models (LLMs) is increasingly difficult as adversarial attacks evolve and outpace alignment pipelines built on pre-collected data. Recent co-evolutionary frameworks let the defender chase a moving attacker, yet their training dynamics expose two compounding pathologies: \emphattack strategy collapse, where the attacker overfits to a narrow set of high-reward rewrites, and \emphdefense adversarial forgetting, where the defender loses competence against earlier attacks as the attack distribution drifts. Existing remedies are partial: attacker-side diversity rewards score textual novelty rather than strategy novelty, and defender-side replay relies on discrete judge scores that ignore estimation uncertainty, response stochasticity, and defender non-stationarity. We introduce DACE, a diversity-driven adversarial co-evolution framework that targets both pathologies jointly. On the attacker side, DACE couples an explicit 12×10 strategy space (risk category × attack style) with a \emphnormalized marginal coverage gain reward, providing a bounded, non-vanishing exploration signal at the strategy layer. On the defender side, DACE maintains a unified \emphBayesian adversarial replay pool whose Beta--Bernoulli threat posteriors are refreshed by time decay and sampled via Thompson sampling, anchoring defender training to an evolving threat landscape. Across standard safety, automated-attacker, and general-capability benchmarks, DACE improves robustness to out-of-distribution attacks while preserving general capabilities, and yields broader attacker strategy coverage.
Abstract: We study uplift estimation for combinatorial treatments. Uplift measures the pure incremental causal effect of an intervention, such as sending a coupon or a marketing message, on user behavior, modeled as a conditional individual treatment effect. Many real-world interventions are combinatorial: a treatment is a policy that specifies context-dependent action distributions rather than a single atomic label. Although recent work considers structured treatments, most methods rely on categorical or opaque encodings, limiting robustness and generalization to rare or newly deployed policies. We propose an uplift estimation framework that aligns treatment representation with causal semantics. Each policy is represented by the mixture it induces over context-action components and embedded via a permutation-invariant aggregation. This representation is integrated into an orthogonalized low-rank uplift model, extending Robinson-style decompositions to learned, vector-valued treatments. We show that the resulting estimator is expressive for policy-induced causal effects, orthogonally robust to nuisance estimation errors, and stable under small policy perturbations. Experiments on large-scale randomized platform data demonstrate improved uplift accuracy and stability in long-tailed policy regimes.
PaperID: 4597, Poster
Abstract: Large language models (LLMs) have recently improved sequential recommendation, yet the task remains a retrieval problem without explicit user queries: the system must infer the next item from user histories where diverse intents and preferences are intertwined. Existing methods typically compress this heterogeneous evidence into a single embedding or query, while LLM-based recommenders can produce overly general queries that are semantically plausible but weakly aligned with retrieval. We propose PRISMIC (Preference Reconstruction via Intent Synthesis and Multi-signal Inference Consolidation), which reformulates sequential recommendation as multi-signal intent consolidation. PRISMIC first generates multiple candidate queries from different historical signals, then trains an LLM-based consolidator with GRPO using an NDCG-based retrieval reward to merge relevant signals into a single query. The resulting consolidated query is not directly used at inference time; instead, it serves as a semantic supervision target for training a user encoder, enabling LLM-free inference. Across four real-world datasets, PRISMIC consistently outperforms strong baselines. Ablations further show that the gains come not from GRPO alone, but from consolidating multiple intent-signals and distilling the consolidated intent into an encoder.
PaperID: 4598, Poster
Authors: Seyedmorteza Emadi
Abstract: Pearl's causal hierarchy shows that observational, interventional, and counterfactual queries are qualitatively distinct. We ask a quantitative version of this question: how many additional bits are needed to specify higher-rung causal answers once lower-rung answers are known? We formalize this via query-class description length, the Kolmogorov complexity of the answer oracle induced by an SCM for a class of queries. Our main construction gives binary acyclic SCMs whose observational distribution has constant description length, while the single-variable interventional answer oracle has description length \Theta(n^2). A degree-sensitive upper bound shows that finite-gate-schema SCMs of indegree d have observational–interventional gap at most O(nd\log(en/d)+n\log n), making the quadratic construction order-optimal in the dense regime and a rooted-tree construction order-optimal for bounded indegree. The quadratic separation persists under \varepsilon-accurate total-variation descriptions for every fixed \varepsilon < 1/4. At the next rung, the full hard-do interventional oracle can still leave a \Theta(n) counterfactual description gap. A general ambiguity-to-bits theorem and Shannon analogue show that these gaps equal the logarithm of residual higher-rung ambiguity up to lower-order terms.
Abstract: Large language models (LLMs) excel at natural language processing but struggle to interact with external environments. Tool learning provides a promising way to extend LLMs into actionable agents, where tool selection is a critical prerequisite for successful tool use. Existing work often assumes a small or predefined set of tools, leaving large-scale tool selection underexplored. Real-world repositories contain a vast and diverse array of tools, making it difficult for LLMs to effectively search, distinguish, and compose tools under context-length constraints. We identify large-scale tool selection as a new challenge for agentic reinforcement learning, highlighting that existing RL methods for knowledge-based question answering are inadequate for selecting tools while considering compatibility. To address this challenge, we propose ToolSearcher, a novel RL framework for effective multi-turn search and fine-grained optimization in large-scale tool selection. Specifically, we introduce category-constrained tool discrimination to improve the model's ability to distinguish functionally similar tools, event-level search modeling to explicitly optimize the discovery of target tools during multi-turn search, and trajectory-aligned credit allocation to provide fine-grained reward signals for different stages of the search-selection process. Extensive experiments on large-scale tool selection benchmarks demonstrate that ToolSearcher consistently outperforms a set of strong baselines in challenging settings involving iterative search and complex tool composition.
PaperID: 4600, Poster
Authors: Miao Zhang, Mohamed Abdelsamad, Bin Yang, Sherif Abdulatif, Benedikt Loesch, Marco Altmann, Abhinav Valada, Bin Yang
Abstract: Raw radar spectra are widely viewed as significantly cluttered and noisy to support direct representation learning. This has led to handcrafted reductions, sparse point clouds, or compact Doppler descriptors as the default preprocessing step in radar perception. In this work, we challenge this assumption and argue that the main limitation is not the raw spectral representation itself, but rather the absence of a learning objective capable of separating the structured signal from noise. Handcrafted reductions partially address this problem, but they impose rigid assumptions about what counts as signal, discarding weak but informative structure while retaining artifacts that match the reduction model. Once such information is removed, supervised learning cannot recover it or question what was preserved. We propose a principled alternative: structure is the component of the radar spectrum that remains reconstructable under localized corruption, while noise is the component that does not. This perspective makes self-supervised masked modeling a natural and assumption-free learning paradigm for dense radar spectra. We instantiate it with Leakage-aware Spectral MAE, a masked autoencoding framework for raw radar cubes that incorporates a spectral-leakage-aware corruption strategy to model physically realistic energy spreading. We also show that Doppler organization is crucial: treating Doppler as a primary spatial dimension, using an x-y-D representation with height encoded as features, substantially outperforms naive channel stacking and is essential for exploiting dense 4D spectra. Without any handcrafted reductions, our approach consistently outperforms state-of-the-art radar-based detectors and point-cloud-based self-supervised learning (SSL) counterparts. These results establish raw radar spectra as a highly effective representation for 3D perception when paired with an appropriate pretraining objective.
Authors:
Qiming Li, Tianlun Li, Xiaolong Cheng, Hangyu Li, Ruiyan Gong, Kangning Niu, kaitao jiang, Mu XuAbstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective paradigm for improving the reasoning capability of Large Vision-Language Models (LVLMs). However, existing RLVR methods primarily rely on trajectory-level outcome rewards, which assign identical learning signals across all generated tokens. This coarse-grained credit assignment is fundamentally mismatched to multimodal reasoning, where only a sparse subset of tokens is causally grounded in visual evidence. Consequently, these pivotal perceptual tokens receive weak supervision and are often overwhelmed by language priors or reasoning-template tokens. To address this limitation, we propose Perception-Reinforced Policy Optimization (PRPO), a token-level reinforcement learning framework that explicitly identifies and reinforces pivotal perceptual tokens within long-horizon multimodal reasoning trajectories. PRPO introduces Robust Visual Dependency (RVD), a principled metric that identifies tokens whose predictions are both visually grounded and perturbation-stable, filtering out brittle or noisy visual tokens. Based on RVD, we further propose Perceptual Advantage Reshaping (PAR), a token-level credit assignment technique that amplifies perceptually informative tokens while preserving stable gradients for non-perceptual tokens. Extensive experiments on seven multimodal reasoning benchmarks demonstrate that PRPO consistently outperforms strong LVLM baselines across both 3B and 7B model scales, achieving average gains of 23.3% and 21.1%, respectively. PRPO achieves state-of-the-art performance with improved training efficiency and stronger cross-task generalization. Our findings highlight the importance of fine-grained credit assignment for scalable multimodal reinforcement learning.
Authors:
David Layden, Ryan Sweke, Vojtech Havlicek, Anirban Chowdhury, Kirill NeklyudovAbstract: Continuous flow models transform Gaussian noise into samples from a learned distribution that closely approximates a complex data distribution. We present a natural mapping between these models and a Schrödinger equation, the fundamental equation of quantum mechanics, whose solution is a quantum state encoding the learned distribution. Our main result is that this Schrödinger equation is efficiently solvable on a quantum computer, which we prove by precisely bounding the discretization error. Therefore, given a trained flow model, the theoretical analysis we introduce implies that future quantum computers will enable a fundamentally different—and potentially more powerful—type of access to its learned distribution, which could be used to perform downstream tasks (e.g., Monte Carlo estimation) more efficiently. More broadly, our results reveal a rare close connection between state-of-the-art generative modeling techniques, such as flow matching and diffusion models, and one of the main expected capabilities of quantum computers: simulating quantum mechanics.
PaperID: 4603, Poster
Abstract: JEPA-style latent world models offer an efficient alternative to pixel-space video generation by predicting future representations rather than synthesizing future frames. However, current action-conditioned predictors in such latent world models are typically trained on single or narrowly scoped embodied datasets with limited task, object, embodiment, and action-space coverage, and often fail to generalize beyond the training distribution. We argue that this lack of generalist action-conditioned prediction is a key bottleneck for scaling latent world models. In contrast, large controllable video world models have a clear scale-up path through large-scale video pretraining and encode broad visual-dynamics priors, but they are too computationally expensive to serve as online simulators and lack a direct interface to latent action-conditioned prediction. We introduce GeneralistJEPA, a framework for transferring generalist motion priors from video world models to efficient action-conditioned JEPA predictors, establishing a scalable training path for future latent world models. Instead of directly distilling generated futures, whose latents entangle useful action dynamics with source-mismatched content and generation artifacts, GeneralistJEPA factorizes future latent prediction into source-conditioned content continuation and action-induced dynamics. Controllable video world models are used as same-action environment randomizers to build an offline motion-reward bank, providing diverse motion supervision without treating generated content as an absolute target. Specifically, generated rollouts are encoded with a frozen JEPA encoder, and only their patch-level temporal deltas are distilled as video-world-model motion rewards for predicted latent dynamics. This design preserves the efficiency of latent rollout, since video generation and representation encoding are performed offline, while leveraging the broader visual-dynamics priors of video world models. We evaluate GeneralistJEPA on three complementary benchmarks: EgoDex for held-out subtask generalization and action recoverability, BAIR robot pushing for future-frame latent prediction and action sensitivity, and Physion for physical contact readout. Our results show that video-world-model motion rewards improve latent prediction, action sensitivity, action recoverability, and physical contact readout beyond real-only training and naive generated-latent distillation, addressing a key generalization bottleneck on the path toward scaling up action-conditioned latent world models.
Abstract: Geometric data pruning methods, while practical for leveraging pretrained models, are fundamentally unstable. Their reliance on extrinsic geometry renders them highly sensitive to latent space perturbations, causing performance to degrade during cross-architecture transfer or in the presence of feature noise. We introduce TopoPrune, a framework which resolves this challenge by leveraging topology to capture the stable, intrinsic structure of data. TopoPrune operates at two scales, (1) utilizing a to perform a local topological optimization on the manifold embeddings, ranking samples by their structural complexity. We demonstrate that our ensures high accuracy and precision, particularly at significant dataset pruning rates (e.g., 90%). Furthermore, through the inherent stability properties of topology, TopoPrune is (a) exceptionally robust to noise perturbations of latent feature embeddings and (b) demonstrates superior transferability across diverse network architectures. This study demonstrates a promising avenue towards stable and principled topology-based frameworks for robust data-efficient learning.
PaperID: 4605, Poster
Authors: Zheshuo Li, Zhengxiong Li
Abstract: Two structural priors—global invariance under a group action and low body-order—are widely used in equivariant architectures on product Lie groups such as \mathrmSO(3)^n. Whether their statistical gains are fundamental, changing the exponent of the minimax rate, or merely constant-factor, changing only the leading coefficient, has remained open. We show that these gains are statistically fundamental: both priors reduce the effective dimension governing the minimax rate, not merely the leading constant. Specifically, we establish matching upper and lower minimax rates for four nested function classes on \mathrmSO(3)^n, one for each combination of the two priors. The matching lower bound relies on a pure-S lift construction and a combinatorial packing across the \binomnK top-order subsets. The resulting effective-dimension formula decomposes additively: invariance contributes a fixed reduction of 3, while body-order truncation at order K contributes a reduction of 3(n-K) that grows with system size. An adaptive sample-splitting procedure selects the body-order from data, achieving the best in-family risk up to an O(\sqrt\log N/N) remainder with no N-dependent tuning. Five experiments using spectral projection estimators corroborate the theory's predicted structural and operational consequences, supporting the four-class hierarchy on synthetic targets and yielding 2–17× sample-efficiency gains on the Maier–Saupe benchmark from liquid-crystal theory.
PaperID: 4606, Poster
Authors:
Heng Zhang, Chengyu Zhou, Jiajun Wu, Xinyuan Liu, Liheng Zhang, Yueqi Guo, Rui Liu, JiaHao Hong, Xuanxun Lian, Jinpeng Lu, Jin HuangAbstract: On-policy distillation has emerged as an efficient paradigm for improving long chain-of-thought reasoning in large language models, where the student learns from its own rollouts while receiving dense teacher feedback on the states it actually visits. Prior work in this paradigm has focused on advancing continuation learning through better teacher scoring, more stable optimization, or more learnable traces, yet largely overlooks a critical failure mode in long-horizon reasoning: once the student drifts at a deep prefix, the bottleneck is no longer how to continue from the current path but how to recover to an effective one. To address this gap, we propose \ourmethod, a recovery-centric framework for on-policy distillation in long chain-of-thought reasoning. \ourmethod operates on student-generated deep prefixes that have already drifted off track yet still admit successful teacher recovery. From each such prefix, it samples multiple teacher recoveries, extracts the short repair segment they share, and trains the student to produce this repair before continuing the remaining reasoning on policy. This design shifts the distillation target from full-suffix imitation to focused path re-entry. To further support long-horizon learning, \ourmethod follows a depth curriculum that expands training depth according to the observed recovery gap, allowing the student to progressively extend its effective reasoning horizon. Extensive experiments on long chain-of-thought benchmarks demonstrate that \ourmethod delivers (I) substantial performance gains, improving strong on-policy distillation baselines by up to 6.84% in final accuracy and 9.27% in recovery success; and (II) strong long-horizon robustness, bringing 3.11%--7.46% gains on deeper prefix regimes and consistent pass@k improvements under longer reasoning horizons.
PaperID: 4607, Poster
Abstract: Hyperspectral Image (HSI) reconstruction has made substantial progress with the deep unfolding framework by decomposing the problem into a projection module and a denoiser module. Nevertheless, existing methods still exhibit limitations of insufficient matching with HSI data. The issues lie in two aspects: 1) the data-consistency projection module applies a fixed gradient descent step while ignoring the spectral content adaptivity of HSI; 2)the denoiser module either suffers from high computational complexity or is limited to the local receptive field, failing to efficiently capture the global spatial-spectral dependencies. In this paper, we propose a deep unfolding framework centered on a two-stage manifold learning strategy to obtain a degradation-free structural prior, which is then strategically injected into both the projection and denoiser modules. In the data-consistency projection, a prior-guided spectral weight mechanism is introduced to adaptively modulate the weight of spectral band. In the denoiser module, a manifold-guided Mamba is proposed to achieve the structure-aware extraction of spatial-spectral features through prior modulation. Quantitative and visual comparisons on synthetic and real-world datasets show that the proposed method significantly outperforms existing state-of-the-art approaches, while maintaining the structural interpretability of the unfolding framework. To support reproducibility, the code will be publicly released upon acceptance.
PaperID: 4608, Poster
Abstract: Classically used for image retrieval tasks in computer vision, the Earth Movers Distance (EMD), also called the Wasserstein distance, has found numerous applications in natural language processing (NLP) and machine learning (ML). In NLP it is used as a measure of distance between sets of embeddings, and in ML it has been used to understand training dynamics, distributionally robust optimization, and many other tasks under active research. On the other hand, computing diverse solutions to optimization problems has also gained a lot of attention recently. In this DiverseX paradigm, one wants to develop algorithms that return a set of r solutions that are maximally diverse; a common measure of diversity is the average or the minimum distance between the r\choose2 pairs of solutions. This paradigm is useful in generating more choices for the user, in fairness, and in robustness and security applications. In this work, we address the complexity of finding a diverse set of solutions in the EMD metric. Given an n point metric space (X,d) and integers k \geq 1 and r \geq 2, we consider the problem of computing r many subsets of X, each of cardinality k, such that the minimum EMD between these sets is maximized. Motivated by applications from NLP, we also consider the problem of computing the farthest k-subset (from a given k-subset) in the EMD metric. On the lower bound side, we first show that assuming the Maximum-Span Hypothesis, it is W[1]-hard (with parameters k and r) to obtain k/(\log k)^O(1)-approximation for non-metric cost functions, and W[1]-hard to obtain a 2-o(1) approximation for metric spaces. We also show that the problem restricted to the Euclidean setting with \ell_2 norm is W[1]-hard. Our first main algorithmic result is an f(k, d, \varepsilon)n^(O(1) time (1-\varepsilon) approximation algorithm for the d-dimensional Euclidean setting, which is tight in view of the above hardness results. Using different techniques, we also present a similar result for the farthest point problem in arbitrary metric spaces. Our second main algorithmic result is geared towards the search for polynomial time algorithms, where we present an O(d) approximation for the Euclidean setting in \textpoly(n,k,d) time when r=2. Finally, we present FPT(k) time, 2-approximate algorithms for arbitrary metric spaces when r=2. Since the class of problems we study requires to find diverse solutions in the EMD metric, our results use a combination of various geometric techniques that deviate from the techniques used to compute the EMD metric between a given pair of solutions, and may be of independent interest.
PaperID: 4609, Poster
Authors: Mohammed ABDULLAH, George Iosifidis, Salah E ELAYOUBI, Tijani Chahed
Abstract: Online Convex Optimization with memory and time-varying constraints (COCO-M) is a recently-introduced framework that captures the dynamics of stateful online learning and non-stochastic control with budget constraints. Despite its expressive power, COCO-M remains largely understudied. In this work, we tackle this problem in its full scale, providing an optimistic algorithm that incorporates untrusted predictions to tame universal dynamic regret and cumulative constraint violation. Our analysis leverages a potent reduction from memory to delay and treats state-dependent gradients as delayed feedback. We establish the first bounds that adapt to both environment non-stationarity (path length P_T) and prediction accuracy, while lifting previous restrictive assumptions such as functional separability. En route to these results, we derive findings of independent interest for the memory-to-delay reduction, and for delayed OCO with time-varying constraints. The proposed framework ensures sublinear worst-case guarantees that improve with prediction accuracy, reaching \(\mathcal O(1)\) regret under perfect predictions, with cumulative constraint violation (CCV) \mathcal O(\sqrtT) for unknown P_T and \mathcal O(\log T) for known P_T, while recovering the optimal COCO rates when memory disappears.
PaperID: 4610, Poster
Authors:
Van Dai Do, Huu H Nguyen, Minh H Nguyen, Thai Hung LeAbstract: Time series forecasting aims to predict future values from historical observations and auxiliary features. We propose SteerCast, a retrieval-based latent steering method that improves decoder-only forecaster at inference time, without updating its parameters. SteerCast constructs a database from the training set by storing a representation of each history window together with a \emphsteering vector computed in the forecaster's latent space, defined as the difference between representations induced by the ground-truth continuation and by the model's own prediction. At test time, SteerCast retrieves nearest neighbors for a query history, aggregates their steering vectors, and injects the resulting signal into the forecaster's hidden states at every step of autoregressive generation, guiding predictions toward trajectories consistent with similar training cases. Experiments across diverse multivariate benchmarks and multiple horizons show that SteerCast consistently improves forecasting accuracy over the fine-tuned backbone and retrieval-based baselines, while requiring no additional training beyond the original fine-tuning and using only the training set as a retrieval corpus.
PaperID: 4611, Poster
Authors:
Haoran Xin, Junwei Pan, Yongqi Zhou, Tianqu Zhuang, YilongSun, Zhixiang Feng, Shoujun Liu, Gong Chen, Shudong Huang, Haijie Gu, Hui XiongAbstract: Generative ranking models (GRMs) have emerged as a scalable alternative to traditional recommendation pipelines by unifying items and user actions into a single token sequence and formulating recommendation as next-token prediction. Despite their promise, we identify a geometric failure mode: attention interactions with lowcardinality action tokens can induce dimensional collapse in high-rank item representations, as suggested by interaction-collapse theory. To pinpoint the origin of collapse, we view attention mechanism over item–action interleaving sequences as a mixture of directional interaction channels between token types, and find that collapse concentrates in the action-attending-item channel, where low-rank action gating drags the item value stream toward a low-dimensional subspace. To address this, we theoretically show that nonlinearity mitigates collapse, and propose Selective VAlue Nonlinearity (SVAN), a simple yet effective approach that applies nonlinear activation to the value component in the collapsed channel, protecting item representations while preserving the remaining attention geometry. SVAN consistently improves the ranking quality over strong baselines. Further analyses confirm that SVAN alleviates dimensional collapse, yielding up to a 12.58% effective-rank gain, and enhances representation discriminability.
Abstract: EEG-based visual decoding aims to establish a mapping between neural signals and visual semantics. However, it remains constrained by the dual challenges of severe information granularity mismatch and the low signal-to-noise ratio (SNR) of EEG signals. Existing approaches typically treat static visual features, ignoring the dynamic selectivity of human vision and the frequency specificity of neural oscillations. To bridge this gap, we propose CAIA, a Cognitive-guided Adaptive blurring with Information-Constrained Alignment framework for Neural-Visual decoding. On the visual side, it simulates selective attention to adaptively reduce redundancy. Meanwhile, on the EEG side, it leverages neural oscillation priors and the information bottleneck mechanism to enhance SNR. Specifically, we devise a cognitive-dynamics-based adaptive blurring mechanism that dynamically integrates center-biased and saliency-guided visual cues via cross-modal attention. Furthermore, we introduce a distribution-aware boundary calibration loss to robustly rectify alignment bias caused by outlier samples. Moreover, a cognitively-guided information-screening method is proposed to select task-relevant EEG oscillations. Extensive experiments demonstrate that CAIA improves both subject-dependent and subject-independent average Top-1 and Top-5 accuracy in zero-shot brain-to-image retrieval, significantly outperforming prior methods. Our work validates that optimizing visual information density to match neural granularity offers a more interpretable and robust pathway for neural decoding. Code is available at https://anonymous.4open.science/r/CAIA-0DA2.
PaperID: 4613, Poster
Abstract: Autoregressive decoding with LLMs is primarily bottlenecked by GPU memory bandwidth, especially in edge-computing settings. While quantization is essential for mitigating this bottleneck, most existing methods treat inference as uniform process and fail to account for the asymmetry between the compute-bound prefill stage and the memory-bound decoding stage. We propose GRINQH (GRaded INput-based Quantization Hierarchy), a weight-only post-training quantization framework that accelerates decoding by unifying quantization and sparsification. GRINQH leverages activation magnitudes as a proxy for computational importance to dynamically assign weight channels to different precision levels, enabling flexible average bit widths during decoding. Evaluated on Llama3 and Qwen3 models, GRINQH outperforms state-of-the-art fixed- and mixed-precision baselines at comparable 3- and 4-bit settings, even enabling effective 2-bit generation. We experimentally verify theoretical speedups by leveraging a hierarchical nested memory layout for multi-precision storage in custom GPU kernels. Together, GRINQH establishes a new state-of-the-art Pareto frontier for LLM generation, enabling a dynamic trade-off between generation quality and inference speed.
PaperID: 4614, Poster
Authors: Wenwei Zhao, Yuxuan Xie, Haiyun Liu, Jie Xu, Zhuo Lu
Abstract: Federated unlearning (FU) enables federated systems to remove designated data from a trained global model, but its security risks remain poorly understood. We show that calibration-based FU introduces a structural vulnerability through finite-step post-hoc corrections, which leave behind \emphknowledge vacuity in weakly constrained residual dimensions where the unlearned data's influence is suppressed while retained-task recovery pressure remains limited. We propose VOID, an unlearning-phase backdoor attack that exploits knowledge vacuity to implant trigger semantics along the legitimate unlearning trajectory. VOID identifies these residual dimensions through influence-based trajectory estimation and neuron-level vacuity profiling, then performs masked semantic substitution during unlearning. Across datasets and FU methods, VOID achieves up to 99% attack success, preserves clean accuracy, and persists after post-unlearning finetuning. Our results show that approximate forgetting can expose writable capacity for adversarial reuse.
PaperID: 4615, Poster
Abstract: Task-oriented dexterous grasp generation aims to produce dexterous grasp poses that are both physically plausible and functionally suitable for specified manipulation tasks. Existing diffusion-based methods often address these two requirements in a decoupled manner: they first train a grasp diffusion model for task alignment and then rely on post-generation refinement to improve physical plausibility. However, this after-the-fact correction strategy applies physical plausibility guidance only once the grasp has already been generated, leaving the generation trajectory itself unguided by physical constraints and potentially leading to suboptimal grasps. To address this problem, we propose a novel framework that directly injects physical plausibility guidance into the denoising process of a task-aligned grasp diffusion model in a practical and effective manner, even when physical plausibility constraints are non-differentiable. This allows physical plausibility to shape grasp generation throughout denoising while preserving task alignment. Extensive experiments demonstrate the efficacy of our framework.
Authors:
Chengqian Zhang, Wei Zhu, Kyumin LeeAbstract: Post-training has become essential for adapting large language models (LLMs) to complex downstream behaviors, including instruction following, preference alignment, and multi-step reasoning. Reinforcement learning with verifiable rewards (RLVR) has recently emerged as a particularly effective post-training paradigm for improving reasoning capabilities, with critic-free algorithms such as GRPO and GSPO enabling scalable optimization. However, RLVR post-training with full fine-tuning (FFT) requires substantial GPU memory and incurs high training costs. Although parameter-efficient fine-tuning (PEFT) methods, such as Low-Rank Adaptation (LoRA), effectively reduce computational costs, they often suffer from a noticeable performance gap compared to full fine-tuning in post-training for complex reasoning tasks. In this paper, we propose Hybrid-LoRA, an efficient hybrid post-training framework that selectively applies full fine-tuning to a small subset of modules less suited to low-rank adaptation, while adapting the remaining components with LoRA. We introduce a novel Hybrid-LoRA Score to rank candidate modules according to their sensitivity to low-rank adaptation under a fixed parameter budget. Experiments show that Hybrid-LoRA closely matches full fine-tuning performance under a 10% full fine-tuning module budget, with the remaining candidate modules adapted by LoRA, consistently outperforming four state-of-the-art PEFT post-training baselines, achieving improvements of up to 5.65% and on average 4.36% over the best baseline.
PaperID: 4617, Poster
Abstract: Stochastic approximation (SA) is a method for finding the root of an operator perturbed by noise. The focus of this paper is studying the distribution of SA iterates in finite time. In general, it is not possible to characterize the exact distribution, and therefore our goal is to find an approximation which can yield useful tail bounds. Inspired by the rich literature on the asymptotic normality of rescaled SA iterates, we approximate the pre-limit distributions by a sequence of Gaussians whose covariance is recursively defined. In particular, we establish explicit bounds on the Wasserstein-1 distance between the rescaled iterate at time k and the aforementioned Gaussian for various choices of step-sizes. Since these covariances converge to the classical asymptotic limit, our analysis also provides a convergence rate for asymptotic normality as a by-product. As an immediate consequence of our bounds, we obtain tail bounds on the error of SA iterates at any time. Finally, we establish the sharpness of our rates by providing matching lower bounds and validate our findings through simulations. We obtain the sharp rates by first studying the convergence rate of the discrete Ornstein–Uhlenbeck (O-U) process driven by general noise, whose stationary distribution is identical to the limiting Gaussian distribution of the rescaled SA iterates. We believe that this is of independent interest, given its connection to sampling literature. The analysis involves adapting Stein’s method for Gaussian approximation to handle the matrix weighted sum of i.i.d. random variables. The desired finite-time bounds for SA are obtained by characterizing the error dynamics between the rescaled SA iterate and the discrete time O-U process and combining it with the convergence rate of the latter process.
Abstract: The Planner-Operator-Reflector (POR) framework is widely used in GUI agents to maintain objective alignment in complex tasks through modular collaboration. However, desktop GUIs introduce a key challenge: large, dense interfaces often exhibit subtle or scattered state changes, placing most of the burden on the reflector, which must compare pre- and post-action screens, while the planner and operator reason over a single state. Existing reflectors collapse change detection and outcome verification into one step, leaving evidence implicit and yielding weakly grounded decisions. To address this limitation, we propose Evidence-First Reflection (EFR), a two-stage reflector that explicitly decouples action-induced visual differences extraction from outcome verification. EFR identifies the action location and candidate changed regions with Set-of-Marks annotations, describes and filters action-relevant changes, and makes the final judgment from the cleaned evidence. This evidence-reasoning decoupled design makes reflection better grounded in screen transitions, while reducing both visual search complexity and reasoning burden. Experiments on OSWorld-Verified and WindowsAgentArena demonstrate that EFR improves reflector accuracy by 7.01%, yielding end-to-end task success gains of 5.86% and 5.05% on the two benchmarks, respectively.
PaperID: 4619, Poster
Abstract: Graph Neural Networks (GNNs) are widely used for representation learning on graphs, but most methods assume static topologies, making them inefficient on evolving networks where edges change over time. Existing dynamic approaches either model graph evolution through temporal GNN architectures without focusing on efficient dynamic maintenance, or are restricted to linear propagation models based on Personalized PageRank. In this work, we study how to efficiently maintain node representations for non-linear GNN dynamics under edge insertions and deletions. For a broad class of standard activation functions, we develop a residual-based dynamic algorithm that selectively propagates local errors via push operations, maintaining an approximation to the evolving fixed point without full recomputation. We prove that our method achieves amortized O(1/\epsilon) update time per graph change under a degree-normalized error guarantee. Our approach uses a potential-based analysis in a degree-scaled norm and, in contrast to prior work on the linear case, requires no randomness assumptions on either the update sequence or the input vector. For the linear special case, we additionally provide an exact dynamic algorithm via low-rank matrix inverse updates. Experiments on benchmark datasets show that incorporating non-linearity improves accuracy while preserving efficient update performance, yielding a scalable and theoretically grounded framework for dynamic GNNs.
Abstract: Cold-start exemplar-free class-incremental learning requires learning a growing set of classes without replay, external pretraining, or a large initial task. Existing cold-start methods typically either train the backbone throughout the stream and compensate for semantic drift, or freeze a backbone after the first task, producing features biased toward the initial classes. These choices also create a computational tension: drift-compensation methods require repeated backbone training and increasingly expensive updates as the task horizon grows, while frozen-backbone methods are cheap but weak under cold start. We study a third option: a feature extractor that is never fit to image data at all. We propose CIRCLE, a class-incremental classifier built from fixed bidirectional two-dimensional reservoir features, adapted from BiRC2D for image classification, and streaming linear discriminant analysis heads. CIRCLE groups multiple random reservoir instantiations into feature ensembles and averages the softmax outputs of independent SLDA heads, yielding a tunable bias-variance tradeoff between richer random features and prediction-level ensembling. Because the feature extractor is fixed and the head admits streaming closed-form updates, CIRCLE performs sample-wise training without replay, task-boundary information, or backbone backpropagation. On CIFAR-100, TinyImageNet, ImageNet-Subset, and ImageNet-1k, CIRCLE is competitive at 10-20 task splits and substantially outperforms strong CS-EFCIL baselines at 50, 100, and 500 task splits, while training much faster than trained-backbone drift-compensation methods. Ablations show that the BiRC2D-style extractor, SLDA head, and balanced feature/prediction ensembling each contribute to the final performance.
PaperID: 4621, Poster
Abstract: Translating visual brain activity into natural language offers a powerful window into human cognition and neural semantic representation. Most prior brain-captioning approaches project fMRI signals into language space using paired supervision. However, the widely used Natural Scenes Dataset (NSD) was built based only on image-fMRI pairs, while existing methods rely on MS-COCO captions as indirect textual supervision. Critically, during fMRI acquisition, participants observe only cropped image regions, whereas the associated MS-COCO captions describe full scenes. This mismatch creates a fundamental gap between neural representations and textual supervision, introducing an inherent , where language is generated directly from visually grounded neural evidence rather than text supervision. We propose , a framework that maps fMRI signals into subject-aware visual tokens and leverages a frozen vision language model (VLM) to produce interpretable and personalized descriptions. Instead of point-wise regression, MindVLM focuses on extracting semantically relevant visual cues from neural activity, naturally accommodating partial observability and inter-subject variability. Extensive experiments on NSD show that MindVLM consistently improves caption quality across multiple VLM backbones, including LLaVA-v1.5, Qwen2.5-VL, and Gemma3, beyond intrinsic model priors. Despite requiring neither text supervision nor an LLM during training and operating under an image-grounded inference paradigm distinct from conventional captioning methods, MindVLM remains competitive with supervised approaches, suggesting that subject-specific neural grounding provides an efficient, scalable, and interpretable path toward faithful brain-to-language generation.
PaperID: 4622, Poster
Abstract: Immersive video communication requires photorealistic, render-efficient, and compact dynamic scene representations. 3D Gaussian Splatting (3DGS) offers a promising representation, but dynamic 3DGS remains difficult to compress due to dense primitives and spatiotemporal redundancy. Anchor-based formulations improve compactness with sparse scaffolds that share geometry and appearance across primitives. However, existing designs often rely on deforming a single canonical scaffold and condition each primitive on its associated anchor in isolation, limiting their ability to handle non-local dynamics and disocclusion while under-exploiting inter-anchor correlations, particularly in motion- or texture-dense regions. To address these limitations, we propose SAGA, a volumetric video codec built upon Sparse Anchor-assisted GAussian splatting representations. SAGA represents dynamic 3D scenes using hierarchically organized sparse 4D anchors, where coordinate-based INR decoders generate fine anchors and Gaussian primitives from inter-anchor interpolations, enabling compact parameter sharing across spatiotemporal structures. For long-range dependencies among unstructured anchors, we further introduce fixed-size memory slots with orthogonality-informed updates for accurate entropy-context modeling. Experiments show SAGA outperforms the state-of-the-art 3DGS-based codec GIFStream, achieving BD-rate reductions of 84.54% on Neu3D, 76.71% on Panoptic Sports, and 83.94% on MPEG MIV.
Abstract: Reusing the early layers of cohort-trained INRs as initialization for new signals has been shown to accelerate and improve signal fitting, yet it remains unclear which layers of the shared encoder learn transferable representations and what those representations encode. We address both questions for two standard backbones, SIREN and Fourier-feature MLPs (FFMLP). First, sweeping the freeze depth across the shared encoder at test time, we find that the optimum coincides with the layer of highest weight stable rank. Moreover, freezing at this depth matches or improves on the standard fine-tuning recipe across all our experiments. Second, identifying which layer transfers does not characterize what that layer encodes. To address this we adopt sparse autoencoders (SAEs), the dominant tool in mechanistic interpretability, and present the first SAE decomposition of INR activations into sparse dictionary atoms. Interestingly, SIREN and FFMLP achieve comparable cohort-fitting quality, but learn qualitatively different dictionaries. Cohort SIREN's atoms are localized, tiling the coordinate plane such that each atom fires in a confined region independent of cohort content. Cohort FFMLP's atoms are image-spanning, tracing the contours of memorized cohort signals. Single-atom ablations confirm causal use of these dictionaries: a single FFMLP atom out of 4096 can drop PSNR by up to 10.6 dB across the image, while SIREN ablations remain confined to where the atom fires. Together, these results give the first mechanistic account of what transfers in cohort-trained INRs and turn their activations into inspectable dictionary atoms. These tools open a path towards characterizing what INRs encode and towards architectures designed for generalization rather than memorization. We plan to release our code upon acceptance.
PaperID: 4624, Poster
Abstract: Atmospheric pollution risk inference is not merely generic time-series extrapolation: target-day risk depends jointly on numerically preserved pollutant state, exogenous meteorological regime, and directed cross-city transport. Existing time-series models are strong sequence learners, graph-based air-quality models often rely on static or weakly directed spatial priors, and multimodal methods commonly merge heterogeneous evidence without specifying its functional role. We present \ourmodel, a transport-aware multimodal evidence-fusion framework that assembles an atmospheric evidence mosaic from numerical pollutant/time states, compact precipitation text, and wind-aligned graph structure. The model preserves historical PM_2.5 and time features as numerical prefix tokens, compresses recent precipitation into prompt-side context, constructs a daily wind-driven DAG to organize upwind-to-downwind evidence, and refines city embeddings through post-pooling DAG residual propagation with root-sensitive mechanisms. We analyze \ourmodel on six PM_2.5 risk-inference benchmarks under a unified year-based split. The results provide protocol-bounded evidence that transport-aligned structure and role-separated interfaces yield useful risk-inference signals, while leaving no-oracle evaluation as future work.
PaperID: 4625, Poster
Abstract: Chain-of-Thought (CoT) reasoning substantially improves Large Language Models (LLMs), yet its out-of-distribution (OOD) generalization mechanism remains theoretically underexplored. We study CoT as compositional OOD generalization: the target task lacks complete reasoning trajectories or direct input-output pairs, while training provides only in-distribution (ID) atomic step data. We show that CoT generalizes by repeatedly applying a shared single-step predictor to compose reusable ID transitions. Theoretically, the OOD task risk is controlled by the sum of ID subtask risks, with an error-propagation coefficient independent of reasoning depth/steps T. A Rademacher-complexity analysis further gives a generalization gap of \mathcalO(\sqrtT/n), rather than linear in T, due to predictor sharing, where n denotes training samples per subtask. For autoregressive Next-Token Prediction, we derive an explicit bound \mathcalO\big(k\big(\sqrtT/mn+\sqrtT/n+T/m\big)\big), where m and k denote context length and output token number per step. Experiments on Atom-Task and GSM8K support the theory, showing that CoT enables short-to-long and compositional generalization, while explicit state tracking is crucial.
Abstract: While standard reinforcement learning optimizes a single reward signal, many applications require optimizing a nonlinear utility f(J_1^\pi,\dots,J_M^\pi) over multiple objectives, where each J_m^\pi denotes the expected discounted return of a distinct reward function. A common approach is concave scalarization, which captures important trade-offs such as fairness and risk sensitivity. However, nonlinear scalarization introduces a fundamental challenge for policy gradient methods: the gradient depends on \partial f(J^\pi), while in practice only empirical return estimates \hat J are available. Because f is nonlinear, the plug-in estimator is biased (\mathbbE[\partial f(\hat J)] \neq \partial f(\mathbbE[\hat J])), leading to persistent gradient bias that degrades sample complexity. In this work we identify and overcome this bias barrier in concave-scalarized multi-objective reinforcement learning. We show that existing policy-gradient methods suffer an intrinsic \widetilde\mathcalO(\epsilon^-4) sample complexity due to this bias. To address this issue, we develop a Natural Policy Gradient (NPG) algorithm equipped with a multi-level Monte Carlo (MLMC) estimator that controls the bias of the scalarization gradient while maintaining low sampling cost. We prove that this approach achieves the optimal \widetilde\mathcalO(\epsilon^-2) sample complexity for computing an \epsilon-optimal policy. Furthermore, we show that when the scalarization function is second-order smooth, the first-order bias cancels automatically, allowing vanilla NPG to achieve the same \widetilde\mathcalO(\epsilon^-2) rate without MLMC. Our results provide the first optimal sample complexity guarantees for concave multi-objective reinforcement learning under policy-gradient methods.
Authors: Yansong Liu, Yuan Liu
Abstract: Speech large language models (SpeechLMs) can process spoken requests, yet still lag behind text-prompted large language models (LLMs) on instruction following and reasoning. We argue that this gap is not only an ASR or representation problem: the same semantic instruction can induce different response policies when spoken with different speakers, accents, noise conditions, prosody, or disfluencies. We propose Verifiable Invariant Reinforced Behavior Alignment (VIRBA), a reinforcement-learning framework for aligning SpeechLM behavior across acoustic realizations of the same semantic intent. VIRBA builds multi-view spoken instruction groups, scores sampled responses with semantic preference, rule-verifiable correctness, cross-acoustic invariance, and adaptive reasoning rewards, and optimizes the model with Cross-Acoustic Group Relative Policy Optimization (CA-GRPO). The resulting objective moves SpeechLM alignment beyond teacher imitation toward robust reasoning policies that remain stable across realistic spoken realizations. Experiments with recent SpeechLM baselines, disfluency robustness, spoken QA, audio reasoning, and speech-to-text translation show the largest gains on reasoning-heavy and acoustically perturbed spoken prompts.
PaperID: 4628, Poster
Authors: Zihui Zhao, Yang Li
Abstract: Dictionary learning has long been studied from both optimization and probabilistic perspectives. While formulations with element-wise sparsity regularization (e.g., L1-based sparse coding) admit well-established probabilistic interpretations, many structured variants that impose global constraints lack a clear and tractable generative view. In this paper, we revisit a class of practically effective yet theoretically under-explored dictionary learning methods that impose a simple global regularization on the number of activated dictionary atoms, which we term parsimoniously activated dictionary learning (PADL). We show that PADL admits an equivalent formulation as maximum a posteriori estimation under a structured generative model, with auxiliary latent variables that govern global activation patterns. This formulation allows us to derive generalization guarantees that are difficult to obtain under the original formulation. More importantly, it yields an analytical characterization of the tradeoff between sparsity, storage cost, and reconstruction accuracy, enabling data-driven estimation of optimal hyperparameters. Based on this connection, we develop an efficient and interpretable PADL algorithm that eliminates manual hyperparameter tuning, achieving improved reconstruction performance under comparable sparsity levels on visual benchmarks. We further demonstrate its practical utility in accelerating inference for vision-language models.
Abstract: Large Reasoning Models (LRMs) introduce new opportunities for safety monitoring through their Chain of Thought (CoT) reasoning. However, CoT is not always faithful to the model's final output, undermining its reliability as a monitoring tool. To address this, we investigate the hidden representations of LRMs to determine whether future behavior can be predicted from prompt and CoT representations. By evaluating a probe at each generated token, we construct a probe trajectory, the continuous evolution of a concept's probability across the reasoning process. We find that future model behavior is more distinguishable when examined over the full trajectory than from a single static prediction. To characterize these temporal dynamics, we extract signal-processing features that capture volatility, trend, and steady-state behavior, significantly improving the separation of future model states. We also present two methodological insights. First, template-based training data achieves near-parity with dynamically generated model responses, eliminating the need for a costly initial inference and labeling. Second, the choice of pooling operation is critical: average-pooling and last-token methods collapse to near-random performance, while max-pooling achieves up to 95% AUROC and yields stable probe trajectories. Using four datasets and four reasoning models across the domains of safety and mathematics, we demonstrate that trajectory features encode task-specific dynamics that improve outcome separability. These findings establish probe trajectories as a complementary framework for monitoring LRM behavior. Warning: This article contains potentially harmful content.
Abstract: We introduce Sharper Transductive Local Complexity (STLC) as a new tool for analyzing the generalization performance of transductive learning methods, improving upon the current transductive bounds. Our work extends the classical local complexity-based analysis to the transductive setting, incorporating substantial and novel components beyond standard inductive and transductive analysis. Although Local Rademacher Complexity (LRC) has been used to obtain sharp inductive generalization bounds and local complexity-based transductive bounds, it has remained an open problem whether a localized Rademacher complexity framework can achieve exactly the same sharp bounds matching their inductive counterparts. STLC provides a confirmative answer to this question. STLC is constructed by first deriving a new and sharp concentration inequality for the supremum of empirical processes capturing the gap between test and training losses, or the test-train process, under uniform sampling without replacement. The proof establishes a Bernstein-type concentration inequality via a novel entropy-based approach built on the modified log-Sobolev inequality for the swap walk. A subsequent peeling strategy with a surrogate variance operator then yields excess risk bounds in the transductive setting that exactly match the classical LRC-based inductive bounds without the additional logarithmic gap in existing works. We further advance the current state-of-the-art in transductive learning through two applications: (1) for realizable transductive learning over binary-valued function classes with finite VC dimension \dVC and u \ge m \ge \dVC, where u and m are the number of test features and training features, STLC gives a nearly optimal bound \Theta(\dVC \log(me/\dVC)/m) nearly matching the minimax rate \Theta(\dVC/m) up to \log m and exactly matching the inductive bound, resolving a decade-old open question; and (2) STLC presents a sharper excess risk bound for transductive kernel learning compared to the prior local complexity–based results.
Abstract: Digital personas powered by Large Language Models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet it remains unclear when they can reliably approximate human survey findings. We answer this question using the LISS panel, constructing personas from respondents’ background variables and pre-2023 survey histories, then testing them against the same respondents’ held-out post-cutoff answers. Across four persona architectures, three LLMs, and two prediction tasks, we assess performance at the question, respondent, distributional, equity, and clustering levels. Digital personas improve alignment with human response distributions, especially in domains tied to stable attributes and values, but remain limited for individual prediction and fail to recover multivariate respondent structure. Retrieval-augmented architectures provide the clearest gains, but performance depends more on human response structure than on model choice: personas perform best for low-variability questions and common respondent patterns, and worst for subjective, heterogeneous, or rare responses. Our results provide practical guidance on when digital personas could be appropriate for survey research and when human validation remains necessary.
PaperID: 4632, Poster
Abstract: Machine learning models deployed in decision-making pipelines, from vehicle routing to dynamic pricing to clinical decision support, must remain reliable in the face of low-quality training data. We study a predict-then-optimize setting in which an upstream model predicts the cost vector of a downstream linear program, with a large pool of noisy training data supplemented by a small curated set of clean data. We introduce Reduced Cost Influence Functions (RCIFs), which combine influence functions from robust statistics with the geometry of linear optimization to estimate how each training point shifts downstream decisions. Building on this, we propose a reweighting algorithm that uses RCIFs to downweight harmful training points and reduce downstream regret. Experiments on random linear programs and shortest-path problems show that our approach reduces regret, and accompanying theory characterizes properties of RCIFs and identifies regimes favorable to our method relative to baselines.
Authors: Blazej Banaszewski, Andrew Fitzgibbon
Abstract: Bioassay activity prediction is often data-limited because drug-discovery datasets rely on time-consuming and expensive wet-lab experiments for data generation and evaluation. This challenge has inspired recent research into molecular foundation models (MFMs), which aim to encode general-purpose chemical knowledge into molecular representations that generalize well in data-constrained scenarios. This paper presents Monroe, a new MFM with several innovations over the existing state of the art: increased scale allowing pre-training on over 80 million molecules from the PM6 quantum chemistry dataset; improved graph representation of stereochemistry; improved training losses including conformer denoising and embedding decorrelation; improved multi-task learning; and the use of a prior-data-fitted model (TabPFN) for downstream in-context prediction. Our evaluations use a principled pairwise comparison framework that measures statistically significant performance differences. Across established Polaris benchmarks, Monroe matches or exceeds existing MFMs, while on activity cliff benchmarks, designed to assess utility for molecular discovery, it achieves significant improvements over prior methods. Finally, ablation and transfer experiments show that PFN-based downstream predictors also substantially improve two leading existing models, MiniMol and CheMeleon, yielding new state-of-the-art variants we call MiniMol PFN, suggesting that our downstream adaptation strategy generalizes beyond Monroe. Full source code is at https://anonymous.4open.science/r/monroe-0E0D, and will be openly released.
PaperID: 4634, Poster
Authors:
Runxuan Tang, Haoyu Gao, Yuyan Ding, Junyi Yao, Zihao Zheng, Zecheng Sheng, Liwei HouAbstract: LLM agents increasingly operate in long-horizon, verifier-rich environments, but current self-improvement mechanisms often write episode-derived lessons directly into memory, making local success an unreliable signal for reusable and safe knowledge. This paper addresses the core question of (FGSE), a verifier-grounded framework that represents each candidate lesson as a structured hypothesis with an explicit precondition, behavior change, expected effect, and verifier. FGSE applies the hypothesis only in a temporary state, tests it on target-transfer, falsification, and archived-regression probes, and then commits, refines, or rejects it according to measured gain and risk. Across web, app, tool-use, and long-term memory benchmarks, FGSE achieves strong task performance, including 39.8% WebArena success rate, 78.8% AppWorld average completion, and 0.724/0.487 Pass¹ on τ-Retail/τ-Airline, while reducing falsification failures to 5.0%, regression damage to 1.3%, and harmful committed updates to 3.6%. These results suggest that self-evolving agents benefit from treating memory updates as testable hypotheses rather than unverified reflections, offering a practical path toward more reliable long-term adaptation.
PaperID: 4635, Poster
Abstract: Autoregressive video generation has achieved impressive results on short clips, yet generating long videos remains challenging due to error accumulation over extended horizons, where each predicted frame depends on potentially imperfect previous frames. Existing approaches typically rely on static context management strategies, such as sliding-window caches that discard early frames or fixed sink frames that permanently preserve initial content. These designs either lose critical long-range context or anchor generation to outdated information, which can lead to degraded fidelity, motion stagnation, and reduced diversity. In this paper, we propose Learning to Memorize (L2M), a learned context memory management framework that dynamically retains, evicts, and evolves historical frames according to predicted importance. The framework consists of an \emphimportance prediction router that estimates the relevance of each historical frame, \emphdegradation-aware training that improves robustness to noisy or corrupted context, and a \emphdynamic memory initialization and evolution mechanism that updates long-term memory using exponential moving average (EMA) importance scores. This allows high-quality frames to replace stale anchors and enables memory to adapt naturally to scene evolution. Trained with a teacher-student distillation objective and sparsity regularization, L2M makes efficient use of a fixed KV-cache budget while preserving the most informative contextual information. Extensive experiments show that L2M achieves superior long-term video generation performance over existing baselines, improving stability and visual fidelity, and establishing a new paradigm for learned memory management in long autoregressive video generation.
Abstract: Autoregressive generation lies in the heart of the mechanism of large language models. It can be viewed as the repeated application of a next-token generator: starting from an input string (prompt), the generator is applied for M steps, and the last generated token is taken as the final output. [Joshi et al., 2025] proposed a PAC model for studying the learnability of the input-output maps arising from this process. We develop an online analogue of this framework, focusing on the mistake bound of learning the final output induced by an unknown next-token generator. We distinguish between two forms of feedback. In the End-to-End model, after each round the learner observes only the final token produced after M autoregressive steps. In the Chain-of-Thought model, the learner is additionally shown the entire M-step trajectory. Our goal is to understand how the optimal mistake bound depends on the generation horizon M, and to what extent observing intermediate tokens can reduce this dependence. Our main results show that the online theory of autoregressive learning exhibits a qualitative picture analogous to the statistical one found by [Hanneke et al., 2026], but with a different scale of dependence on the generation horizon. In the End-to-End model, we prove a taxonomy of possible mistake-bound growth rates in the generation horizon M: subject to mild regularity conditions, every rate between constant and logarithmic can arise. We also show that this logarithmic ceiling is unavoidable for the online setting, in the sense that every class of finite Littlestone dimension exhibits at most logarithmic dependence on M. In the Chain-of-Thought model, the parallel is even sharper: as in the statistical setting, access to the full generated trajectory eliminates the dependence on M altogether. We also analyze autoregressive linear threshold classes. For autoregressive linear thresholds in \(\mathbbR^d\), we prove that the optimal mistake bound is \(\Theta(d^2)\) under both End-to-End and Chain-of-Thought feedback, and for any generation length M. The methods developed for this analysis also yield new lower bounds in the statistical setting. Along the way, our results resolve several questions left open by [Joshi et al., 2025]. In particular, we show that even for classes of finite Littlestone dimension, the End-to-End statistical sample complexity can depend on the generation length M.
Abstract: We propose an instantiation of flow matching that relies on a time-independent velocity field (an \emphautonomous flow) to exactly map between two distributions, so long as the target is singular, i.e.\ supported on a lower-dimensional data manifold. We also show that the one-step generative map associated with this flow is the unique solution of a simple conservation equation, which can be used to learn the map directly from samples. These autonomous flows and maps give a dynamical meaning to the flux constraint of Beckmann's transportation problem. Their construction provides a unifying framework that recovers, for instance, the closed-form Poisson-flow generative model and equilibrium matching with a quadratic flow-matching regression loss. We illustrate how this theory corrects inconsistencies in existing methods and improves their performance on large-scale image generation.
PaperID: 4638, Poster
Abstract: Video large language models are often dominated by visual tokens, which lengthen prefill, enlarge KV-cache residency, and raise peak memory. In spirit of the philosophy focus more and memorize less, we introduce Focus--Ambient Retention (FAR), a training-free visual token retention method that routes visual evidence through two complementary streams. The Focus stream preserves task-critical objects, actions, text, and fine details, while the Ambient stream keeps scene context only when it differs from a temporal cache. FAR combines visual attention, local frequency variation, and bounded query relevance to score tokens, then assembles a fixed-budget context with diversity control before language-model prefill. Across five video understanding benchmarks under a variety of model sizes, FAR significantly reduces peak VRAM by up to 50.6% under matched retained-token budgets while achieving comparable or even superior performance over the state-of-the-art, improving the quality-memory trade-off. Overall, FAR offers a position-aware alternative to frame- or block-level compression for efficient inference. Code will be available.
PaperID: 4639, Poster
Authors:
Wei Huang, Chenying Liu, Yilei Shi, Zhitong Xiong, Xiaoxiang ZhuAbstract: Visual foundation models (VFMs) provide strong region-level representations for semantic segmentation, substantially improving high-level semantic understanding. However, accurate boundary recovery remains challenging, especially for fine structures and semantic transitions. A key reason is that conventional pixel-level supervision largely treats pixels independently and lacks explicit modeling of local structure and neighboring-pixel relationships. To address this issue, we propose a Local Structure (LS) regularization framework that constrains the second-order local organization of semantic transition regions. Unlike existing boundary losses that mainly supervise contour location, distance, overlap, or topology, LS models boundary neighborhoods through a structure-tensor representation of semantic probability fields. It decomposes local boundary structure into two complementary cues: Local Structure Magnitude (LSM), which measures the strength of anisotropic class transitions, and Local Structure Orientation (LSO), which describes their dominant geometric orientation. LS is model-agnostic, compatible with different VFM-backbone-decoder combinations, and introduces no extra inference cost. Extensive experiments on four diverse datasets show that LS consistently improves boundary-sensitive performance while also enhancing overall IoU, demonstrating its effectiveness for boundary-structure learning in semantic segmentation.
Abstract: We investigate the ability of transformers to perform in-context reinforcement learning (ICRL), where a model must infer and execute learning algorithms from trajectory data without parameter updates. We show that a linear self-attention transformer block can provably implement policy-improvement methods, including semi-gradient SARSA and actor-critic, via explicit parameter constructions. Beyond existence, we design a teacher-mimicking training procedure, analyze its gradient-flow dynamics, and establish the first convergence guarantee in the ICRL literature: under suitable richness conditions on the training MDP distribution, gradient flow converges locally and exponentially to an optimal parameter manifold corresponding to the desired RL update. Empirically, training transformers on randomly generated tabular MDPs confirms these predictions: the learned models recover the parameter structure of our explicit constructions and, when deployed on unseen MDPs, deliver strong in-context control performance. Together, these results illuminate how transformer architectures internalize and execute classical reinforcement learning algorithms in context, bridging mechanistic understanding and training dynamics in ICRL.
PaperID: 4641, Poster
Authors:
Yuxi Ma, Lin Zhao, Jing Yang, jiacheng wang, Liansheng WangAbstract: Federated Domain Generalization (FDG) aims to learn a global model robust to heterogeneous domain shifts without sharing raw data, a critical challenge in multi-center medical imaging. Most existing methods implicitly assume static inter-domain discrepancies, overlooking the fact that client representations evolve continuously during local optimization. We identify this phenomenon as representation dynamics, which refers to the temporal evolution of latent features induced by training updates rather than shifts in underlying data distributions. Such dynamics often lead to unstable aggregation, particularly in heterogeneous medical imaging scenarios. To address this issue, we propose Dynamic Knowledge Tracked FDG (DKT-FDG), a framework that explicitly tracks these dynamics by leveraging flow matching to model continuous representation trajectories. By aligning representation evolution instead of static snapshots, DKT-FDG improves robustness to domain shifts. Comprehensive experiments on multiple multi-center medical benchmarks demonstrate consistent improvements over state-of-the-art FDG methods, highlighting the importance of modeling representation dynamics in federated learning.
PaperID: 4642, Poster
Abstract: Accurate real-time forecasting of atmospheric pollutants is essential for reducing exposure risks and improving public health. Traditional numerical forecasting systems suffer from incomplete parameterizations, uncertain inputs, and high computational cost. Existing deep learning approaches improve efficiency, but often entangle meteorological and pollutant variables within unified forecasting frameworks, limiting flexibility and pollutant-specific representation learning. To address these challenges, we propose AirMPA, a decoupled autoregressive framework for meteorology-to-pollution forecasting. AirMPA predicts 13 atmospheric pollutants at 0.4^\circ spatial resolution and 12-h temporal resolution using meteorological fields, emission priors, and historical pollutant concentrations. The model is built on an improved Swin-Transformer architecture with zonal release attention and dual-branch fusion, enabling more effective modeling of global transport continuity and heterogeneous pollutant dynamics. During training, AirMPA uses ERA5 reanalysis as meteorological forcing, while during inference it directly ingests forecast fields from external meteorological models, enabling stable autoregressive prediction up to 5 days ahead. Experiments show that AirMPA consistently outperforms Aurora and surpasses CAMS in approximately 90% of pollutant--lead-time evaluation cases over the 5-day forecast horizon, while outperforming CAMS on all variables beyond 48 h. These results demonstrate the potential of decoupled meteorology-to-pollution forecasting for global air quality prediction.
PaperID: 4643, Poster
Abstract: Neural network verification is key to certifying robustness in safety-critical systems. Sound verifiers commit to abstract domains that propagate soundly through the network's operations, paying for that commitment with looseness at scale and applicability largely confined to piecewise-linear architectures. Probabilistic verifiers relax soundness, but they too commit to a fixed family of shapes for the network's reachable outputs. When the network's true output geometry differs from the chosen shape, the verifier over-approximates the reach-set and fails to certify networks that are in fact safe. We propose instead to learn the reach-set geometry from the network's behavior. A flow-matching model trained on input-output samples and calibrated by conformal prediction yields a probabilistic reach set whose shape adapts to the true shape of the reach-set rather than to an a priori choice. The resulting set has no closed-form description, so we reframe the specification check as a rare-event estimation problem on the learned distribution, restricted to the calibration's coverage region, and solve it with bounded adaptive multilevel splitting. The pipeline produces a probabilistic safety certificate combining the conformal coverage with a rare-event upper bound. Our pipeline performs comparably to sound verifiers on a subset of VNN-COMP 2025 benchmarks where sound verification is tractable and matches or exceeds fixed-shape probabilistic baselines on most benchmarks in the same suite. On a synthetic family of growing-depth networks, it scales sub-exponentially where sound verifiers time out or abstain.
PaperID: 4644, Poster
Abstract: Despite the remarkable progress of evolutionary LLM frameworks like AlphaEvolve in program and algorithm discovery, they consistently struggle with tasks requiring deep algorithmic reasoning and long-horizon planning, such as the Abstraction and Reasoning Corpus (ARC-AGI). A fundamental limitation of current evolutionary code generation is its reliance on unguided stochasticity to derive child programs from parents, lacking a systematic mechanism for knowledge accumulation, whereas human problem-solving thrives on the continuous consolidation and reuse of high-level conceptual abstractions. To bridge this gap, we propose the Library-driven Evolutionary Abstraction Paradigm (LEAP), a cognitive-inspired framework that enables LLMs to autonomously construct and evolve a globally shared library of reusable algorithmic primitives. The evolution of this library is governed by two rigorous theoretical principles: First, the generation of new primitives is driven by a Minimum Description Length (MDL) objective, ensuring the system only assimilates concepts that genuinely compress the problem space. Second, the evolutionary survival and selection of these primitives are governed by a principled exploration-exploitation mechanism, dynamically regulating the library's life-cycle to strike an optimal balance between exploiting proven cognitive operators and exploring novel hypotheses while strictly bounding context entropy. Extensive experiments demonstrate that LEAP significantly outperforms existing LLM evolutionary baselines on the ARC-AGI benchmark and math optimizations problems, exhibiting robust cross-task generalization, highly efficient context utilization, and autonomous concept discovery capabilities.
PaperID: 4645, Poster
Authors:
Yashvir Singh Grewal, Daniel M Steinberg, Thang Bui, Cheng Soon Ong, Edwin V BonillaAbstract: Many scientific design problems require identifying rare high-fitness objects in vast discrete spaces using costly black-box feedback. Variational active-generation methods, including variational search distributions (VSD) and conditioning by adaptive sampling (CbAS), address this by learning a distribution over high-fitness designs via inference on a level-set posterior. However, these methods rely on generators with explicit likelihoods, most commonly autoregressive models, creating a mismatch with modern non-autoregressive approaches, such as discrete diffusion and flow-based models, which refine designs in parallel and better capture multi-modal structure but have implicit marginal likelihoods. We introduce Active Flow Matching (AFM), a variational active-generation framework that removes this likelihood bottleneck by reformulating level-set KL objectives over the conditional endpoint distributions of discrete flow models. AFM yields forward-, reverse-, and symmetric-KL objectives without requiring density evaluation. We show that forward-KL AFM is target-consistent, recovering the level-set posterior without access to marginal likelihoods. Empirically, across protein and small-molecule design tasks, AFM outperforms strong baselines, including autoregressive VSD and CbAS, inference-time guidance, KL-regularised fine-tuning, and Top-K retraining.
PaperID: 4646, Poster
Abstract: Real-world thermal infrared face reconstruction faces the dual challenges of compound sensor-specific degradations and a large infrared-to-visible domain gap. Existing methods either rely on oversimplified degradation assumptions or target only specific degradation types, leaving realistic sensor noise under cross-modal translation largely unaddressed. To tackle this compound problem, we decompose it into two sequential tasks, infrared degradation restoration to remove sensor-specific artifacts, followed by infrared-to-visible translation to bridge the domain gap to the visible spectrum. Both tasks require multi-scale detail reconstruction from spatially heterogeneous tokens that exhibit semantic, frequency, and task-specific variations. This motivates our proposed IRCasDiff, a cascaded diffusion framework consisting of the above two sequential stages. Both stages share a unified backbone with an asymmetric Sparse Mixture-of-Experts design, featuring content-adaptive routing in encoder blocks, frequency-decoupled reconstruction in decoder blocks, and text-free prompt generation for infrared-specific conditioning. The translation stage is initialized from converged restoration stage weights, enabling the model to master infrared restoration before tackling cross-modal mapping. Experiments on MCXFace and SpeakingFaces demonstrate state-of-the-art performance, with FID reductions up to 48.4% over translation baselines and SSIM improvements up to 19.3% over restoration baselines.
Abstract: Model-based reinforcement learning (MBRL) is sample-efficient but struggles in sparse reward settings. A critical bottleneck arises from the lack of informative gradients in sparse settings, where standard reward models often yield flat landscapes that struggle to guide planning. To address this challenge, we propose ), a novel framework that shifts reward modeling from predicting sparse scalars to constructing informative potential landscapes. SLOPE employs optimistic distributional regression to estimate high-confidence upper bounds, which amplifies rare success signals and ensures sufficient exploration gradients. Evaluations on
PaperID: 4648, Poster
Abstract: We address the problem of generating a 3D-consistent, navigable environment that is spatially grounded: a simulation of a real location. Existing video generative models can produce a plausible sequence that is consistent with a text (T2V) or image (I2V) prompt. However, the capability to reconstruct the real world under arbitrary weather conditions and dynamic object configurations is essential for downstream applications including autonomous driving and robotics simulation. To this end, we present CityRAG, a video generative model that leverages large corpora of geo-registered data as context to ground generation to the physical scene, while maintaining learned priors for complex motion and appearance changes. CityRAG relies on temporally unaligned training data, which teaches the model to semantically disentangle the underlying scene from its transient attributes. Our experiments demonstrate that CityRAG can generate coherent minutes-long, physically grounded video sequences, maintain weather and lighting conditions over thousands of frames, achieve loop closure, and navigate complex trajectories to reconstruct real-world geography.
Abstract: Capturing screens is common, but photos of emissive displays are often influenced by \emphflicker-banding (FB), some alternating bright--dark stripes due to temporal aliasing between a camera's rolling-shutter readout and display's brightness modulation. Unlike moir\'e degradation, FB remains underexplored despite its frequent and severe impact on readability and perceived quality. We formulate FB removal as a dedicated restoration task and introduce Removal of Image Flicker-Banding via Latent Diffusion Enhancement, RIFLE, a diffusion-based framework designed to remove FB with fine details. We propose Banding-Suppressed High-Frequency Prior (BSHP), which combines gradient-based high-frequency localization with a smooth structural support map to build a compact restoration prior, and injects it into the restoration backbone via multi-stage FiLM modulation. Moreover, Masked Loss (ML) is proposed to concentrate supervision on banded regions without sacrificing global fidelity. To overcome data scarcity, we provide a simulation pipeline synthesizing FB in the luminance domain with stochastic jitter in banding angle, spacing, and width. Feathered boundaries and sensor noise are also applied for a more realistic simulation. For evaluation, we collect a paired real-world FB dataset with pixel-aligned banding-free references captured via long exposure. Across quantitative metrics and visual comparisons on our real-world dataset, RIFLE consistently outperforms recent image reconstruction baselines. To the best of our knowledge, it is the first work to research the simulation and removal of FB. Our dataset and code will be released soon.
Abstract: Tree-based Monte-Carlo Tree Search (MCTS) duplicates the same state when it is reached through different trajectories, which can waste simulations in stochastic MDPs. We introduce Graph-Based Stochastic-Power-UCT (GS-Power-UCT), which shares states reached at the same planning depth while keeping separate values for states reached at different depths. This design applies to general stochastic MDPs, including problems with cycles. We prove that for a fixed planning horizon, the root estimate converges to the finite-horizon value at rate O(n^-1/2), matching tree-based Stochastic-Power-UCT while reusing samples across shared states. We also study two full-state variants: GS-Power-UCT-F, which stores one node per physical state to increase sample sharing but may mix values from different remaining horizons, and GS-Power-UCT-F^+, which uses an adaptive horizon to control this bias. The latter converges to V^\star(s_0), the optimal infinite-horizon discounted value at the root state s_0, when the remaining cross-depth gap vanishes. Experiments on stochastic planning benchmarks show improved sample efficiency over tree-based and graph-based baselines.
PaperID: 4651, Poster
Abstract: Existing physics-based approaches for underwater 3D restoration and reconstruction commonly model backscatter as a function of scene depth and attenuation coefficients. However, they often overlook the view-dependent nature of underwater scattering across camera poses, making it difficult to separate medium degradation from intrinsic scene radiance and thereby limiting both visual and geometric accuracy. To address this, we present WaterDescatterGS, an efficient physics-integrated underwater 3DGS framework that explicitly models view-dependent water degradation at the pixel level. Leveraging the multi-view geometry inherent to Structure-from-Motion, our framework dynamically derives base per-view illumination directions from known camera rotations and a single globally optimized sun direction. To accommodate complex real-world photometric variations beyond this physical prior, a Physics-Residual Appearance Embedding is introduced to absorb unmodeled degradations via view-specific angular corrections. Ultimately, these refined per-view parameters are naturally integrated into the physical model to compute precise per-pixel scattering angles coupled with the rendered depth. During inference, a cross-view attention mechanism dynamically aggregates adjacent features to enhance robust novel view synthesis in turbid water. Extensive experiments demonstrate that, despite operating entirely without clean reference supervision, WaterDescatterGS achieves state-of-the-art performance in high-fidelity novel view synthesis, novel-view-consistent scene restoration, and geometric reconstruction accuracy.
Abstract: Instantaneous Quantum Polynomial Quantum Circuit Born Machines (IQP-QCBMs) have been proposed as quantum generative models that combine a classically tractable training objective—based on the maximum mean discrepancy (MMD)—with a potential quantum advantage motivated by sampling-complexity arguments. While recent works have explored this model across various application domains, fundamental questions remain: does the model suffer from exponentially vanishing loss gradients, known as the barren plateau problem—a pervasive obstacle in quantum machine learning—and how do regimes of trainability relate to regimes of possible quantum advantage? Here, we address both questions analytically. To study trainability, we derive closed-form expressions for the variances of the partial derivatives of the MMD loss function and establish general upper and lower bounds. We explicitly characterize how trainability depends on the generator set and the spectrum of the chosen kernel, identifying regimes in which low-weight kernels avoid exponential gradient suppression under structured topologies. Regarding potential quantum advantage, we reformulate the anti-concentration property in terms of the same generator-set quantities that govern trainability, enabling a unified analysis. We show that sparse IQP architectures can produce classically intractable output distributions while simultaneously remaining trainable, at least at lower-weight frequencies. Our analytical results corroborate the numerical observations of prior work and provides principled guidelines for designing scalable IQP-QCBM architectures.
Authors: Xiaoxiao Lu, Ye Yuan, Jiahao Shi
Abstract: Long-horizon forecasting of time-dependent partial differential equations (PDEs) is critical for characterizing the sustained evolution of physical systems. While neural operators have emerged as efficient surrogates, they typically learn implicit finite-time transitions from discrete observations. When deployed autoregressively, such propagators often suffer from rapid error accumulation and dynamic drift. To address this, we propose a neural forecasting framework that reformulates PDE rollout as learning a Structured Spectral Propagator (SSP) in a propagation-oriented latent space. Following an analysis-propagation-synthesis design, our framework: (i) maps physical states into a shared, time-consistent spatial representation; (ii) projects this space into a compact propagation state to isolate recurrent dynamics from fine-grained spatial details, thereby decoupling reconstruction fidelity from rollout regularity; and (iii) evolves retained spectral modes using a frequency-conditioned linear backbone complemented by a nonlinear spectral closure to account for truncated interactions. This explicit structuring endows the propagator with a strong inductive bias for coherent modal evolution. Extensive experiments demonstrate that SSP significantly outperforms state-of-the-art baselines, reducing relative L_2 errors by up to 48.9% and exhibiting improved stability in temporal extrapolation beyond the supervised horizon.
Abstract: Neural Graph Databases (NGDBs) support complex logical reasoning over incomplete knowledge structures, yet their training efficiency and expressivity are constrained by rigid query-level batching and structure-only embeddings. We present NGDB-Zoo, a unified framework that resolves these bottlenecks by synergizing operator-level training with semantic augmentation. By decoupling logical operators from query topologies, NGDB-Zoo transforms training loops into dynamically scheduled data-flow executions, enabling multi-stream parallelism and achieving a 1.8×-6.8× average throughput compared to baselines. Furthermore, we formalize a decoupled architecture to integrate semantic priors from pre-trained text embeddings without triggering I/O stalls or memory overflows. Experiments on six benchmarks, including ogbl-wikikg2 and ATLAS-Wiki, show that NGDB-Zoo scales dense query-embedding training to million-entity graphs while maintaining competitive filtered MRR.
PaperID: 4655, Poster
Abstract: Modeling Irregular Multivariate Time Series (IMTS) poses significant challenges due to asynchronous sampling and data sparsity. While existing methods focus on handling temporal irregularity after encoding, the tokenization stage itself remains underexplored. We propose IMTS-Tokenizer, a time aware tokenization framework built on the Irregular-time-aware Module (IRAM), which employs learnable temporal anchors and Gaussian-kernel attention to aggregate observations while preserving temporal information. To improve robustness under sparse and asynchronous sampling, the framework further incorporates Relative Time Feature (RTF) for pre-aggregation temporal encoding, Variable-Specific Anchor Initialization (VSA) for data-aligned anchor placement, anchor regularization for stable training, and multi-scale modeling. Experiments on four benchmarks demonstrate state-of-the-art (SOTA) performance, with the best results on seven out of eight metrics and statistically significant aggregate gains.
PaperID: 4656, Poster
Authors: Minh Quang Nguyen, Hady Lauw
Abstract: Score-based diffusion is well-studied at the mixture level—averaged over the prior—but many practical pipelines (posterior sampling, SDEdit, OOD detection) fix the starting point and care about the per-start output. To our knowledge, this per-starting-point regime has no prior non-asymptotic analysis. We give two coupled bounds: a lower bound on how distinguishable outputs from two different starts remain. The same score-contraction rate governs both—where plausibility is tight, diversity is vacuous, and vice versa—predicting a sharp horizon beyond which starting-point identity is erased. We illustrate the phenomenon qualitatively on 2D toy datasets, MNIST, and CIFAR-10.
PaperID: 4657, Poster
Abstract: Continual learning suffers from catastrophic forgetting, leading to the loss of past knowledge and unstable adaptation across task streams. Existing approaches typically address this issue through replay buffers, architectural changes, or additional regularization terms. We argue that forgetting is a consequence of unstable learning dynamics, which may benefit from an optimizer-level solution. The main challenge is that learning dynamics are time-dependent, as demonstrated by empirical phenomena such as the stability gap. First, we formally prove that, for distance-based classifiers, bounded parameter drift guarantees bounded margin erosion under a local score-stability condition. Based on this insight, we propose STABLE, a primal–dual optimizer with adaptive drift control. STABLE enforces a drift budget by increasing the dual variable when the budget is violated and decreasing it otherwise. This results in high stability without penalizing plasticity. Experiments on long-horizon unimodal streams and CLIP-based multimodal streams show that vanilla PEFT baselines trained with STABLE match or outperform specialized continual learning methods. In the multimodal setting, STABLE further preserves zero-shot transfer, matching or exceeding dedicated multimodal continual learning methods.
Abstract: Training instabilities such as loss spikes are frequently the result of stochastic gradient noise. Because of rare expressions in language training data, and multiple layer composition, the noise impact is heavy-tailed and survives mini-batch averaging. Gradient clipping is designed to address this where vector-norm clipping ignores matrix structure in weight updates, while spectral normalization (e.g. Muon) respects structure at additional cost. We show an entry-wise \emphheavy-tailed noise appears similar to real stochastic gradient noise. Furthermore, through a first-order perturbation analysis, we identify a \emphlocalization property under which a simple entry-wise method will give spectral normalization. Exploiting this, we derive a tractable surrogate for the Bayes-optimal entry-wise estimator under a Gaussian signal prior. We establish O(\epsilon^-4) convergence guarantee under Cauchy-contaminated noise. Empirically, we find that smooth shrinkage improves Adam on NanoGPT pretraining, saving ~8% of training tokens. We further find that applying the entry-wise clipping before spectral normalization yields a ~2% token saving on top of Muon.
Abstract: Building a generalist robot that can perceive, reason, and act across diverse tasks remains an open challenge, especially for dexterous manipulation. A major bottleneck lies in the scarcity of large-scale, action-annotated data for dexterous skills, as teleoperation is difficult and costly. Human data, with its vast scale and diverse manipulation behaviors, provides rich priors for learning robotic actions. While prior works have explored leveraging human demonstrations, they are often constrained by limited scenarios and a large visual gap between human and robots. To eliminate these limitations, we propose METIS, a vision-language-action (VLA) model for dexterous manipulation pretrained on multi-source egocentric datasets. We first construct EgoAtlas, which integrates large-scale human and robotic data from multiple sources, all unified under a consistent action space. We further extract motion-aware dynamics, a compact and discretized motion representation, which provides efficient and expressive supervision for VLA training. Built upon them, METIS integrates reasoning and acting into a unified framework, enabling effective deployment to downstream dexterous manipulation tasks. Our method demonstrates exceptional dexterous manipulation capabilities, achieving highest average success rate in eight real-world tasks. Experimental results also highlight the superior generalization and robustness to out-of-distribution scenarios. These findings emphasize METIS as a promising step toward a generalist model for dexterous manipulation.
Abstract: Schedule-Free SGD, proposed in The Road Less Scheduled, achieves optimal convergence rates without requiring the training horizon in advance, by replacing learning rate schedules with a principled form of iterate averaging. However, the method still requires tuning a base learning rate whose optimal value depends on unknown problem constants. In this work, we continue down this road by deriving Polyak-type step sizes for Schedule-Free SGD and Adam that compute the learning rate at each iteration from the sampled loss, gradient, and current iterates alone. We first propose an oracle variant that uses per-sample optimal function values and prove an \mathcalO(1/\sqrtt) anytime last-iterate rate for convex Lipschitz objectives. We then remove the oracle requirement with a safeguarded variant that replaces the unknown optimal values with any available lower bound, achieving the same rate up to a neighborhood that vanishes under interpolation. Both step sizes reduce to existing Polyak rules for standard SGD when momentum is set to zero, unifying standard and schedule-free Polyak methods. Numerical experiments on language modeling, including pretraining and distillation, show that the proposed methods match or surpass tuned Schedule-Free baselines while offering greater robustness to hyperparameter choices.
Abstract: The quadratic computational complexity of standard attention mechanisms presents a severe scalability bottleneck for LLMs in long-context scenarios. While hybrid attention mechanisms combining Full Attention (FA) and Sparse Attention (SA) offer a potential solution, existing methods typically rely on static allocation ratios that fail to accommodate the variable retrieval demands of different tasks. Furthermore, head-level dynamic sparsity often introduces severe computational load imbalance and synchronization long-tails, which hinder hardware acceleration during autoregressive decoding. To bridge this gap, we introduce , a context-aware framework that dynamically optimizes attention computation at the layer level. By integrating a lightweight into frozen pretrained LLMs, the proposed method adaptively routes each layer to FA or SA based on the input context. This layer-wise routing preserves high-fidelity information retrieval while ensuring contiguous memory access, translating theoretical computational reductions into practical wall-clock speedups. As a parameter-efficient approach, our framework requires only 12 hours of training on 8xA800 GPUs. Extensive experiments across multiple long-context and mathematical reasoning benchmarks demonstrate that Flux Attention achieves a superior trade-off between performance and inference speed compared with baseline models, with speed improvements of up to 2.8x and 2.0x in the prefill and decode stages.
Abstract: Domain generalization (DG) aims to learn predictive models that can generalize to unseen domains. Most existing DG approaches focus on learning domain-invariant representations under the assumption of conditional distribution shift (i.e., primarily addressing changes in P(X\mid Y) while assuming P(Y) remains stable). However, real-world scenarios with multiple domains often involve compound distribution shifts where both the marginal label distribution P(Y) and the conditional distribution P(X\mid Y) vary simultaneously. To address this, we propose a unified framework for robust domain generalization under divergent marginal and conditional distributions. We derive a novel risk bound for unseen domains by explicitly decomposing the joint distribution into marginal and conditional components and characterizing risk gaps arising from both sources of divergence. To operationalize this bound and prevent it from becoming vacuous, we explicitly enforce feature-space \ell_2-normalization. We then design a meta-learning procedure that minimizes and validates the proposed risk bound across seen domains, ensuring strong generalization to unseen ones. Empirical evaluations demonstrate that our method achieves competitive performance not only on conventional DG benchmarks but also in challenging multi-domain long-tailed recognition settings where both marginal and conditional shifts are pronounced.
PaperID: 4663, Poster
Abstract: Understanding the 3D physical world has emerged as an essential component in vision-language-action models (VLAs), enabling robots to perform several challenging tasks by unifying spatial perception, language, and action. Existing approaches typically either rely on explicit 3D inputs, such as depth maps or point clouds, or implicitly inject spatial knowledge into visual representations. However, explicit 3D inputs are often brittle in practice due to sensor noise, hardware heterogeneity, and incomplete depth coverage, while implicit alignment can distort the original 2D visual representation and weaken generic scene understanding. In this paper, we present SPARK-VLA, a simple yet effective framework to allow VLAs to implicitly possess generic 2D and spatial 3D visual comprehension within a unified visual backbone. SPARK-VLA consists of two key components: (1) spatial representation distillation, which transfers 3D understanding from a frozen 3D expert into a LoRA-adapted visual branch by aligning attention responses while preserving the original visual pathway, and (2) a lightweight knowledge router, which selectively forwards action-relevant visual tokens to the language model using routing supervision derived from action-loss gradients. Experiments in both simulation and real-world environments show that SPARK-VLA, with only a 0.5B language backbone and 50M trainable parameters, achieves performance comparable to or better than substantially larger VLAs.
PaperID: 4664, Poster
Authors: Jaehun Shon, Jinha Choi, Jongwook Jeon, Jongmin Lee
Abstract: Offline reinforcement learning aims to learn a policy solely from fixed datasets, which often contain multimodal action distributions. Flow policies can naturally represent such multimodal behaviors, but learning an efficient one-step flow policy remains challenging: standard value guidance often leads to mode collapse or exploits overestimation bias in out-of-distribution regions. To address this, we introduce One-step Flow policy via Optimal Transport (OptiFlow), a framework for one-step flow policy learning as a structured sample-allocation problem. OptiFlow jointly trains a value-aware reference flow policy and an efficient one-step policy, coupling their action samples through state-wise entropic optimal transport. For each state, critic-estimated values define the priority of distillation target actions, while the action-distance cost ensures geometrically compatible pairings. By avoiding direct critic maximization, our transport-guided approach enables in-distribution exploitation by anchoring the one-step policy to high-value, dataset-supported modes without the risk of out-of-distribution divergence. Experimental results demonstrate that OptiFlow effectively captures optimal multimodal behaviors and achieves strong performance across diverse offline RL benchmarks.
Abstract: Diffusion large language models promise parallel token generation, yet inference remains bottlenecked by deciding which masked tokens can be safely committed together. Fast-dLLM addressed this with KV caching and confidence-guided parallel decoding, but its decoding theory uses a homogeneous high-confidence assumption that effectively reduces each candidate set to its weakest selected token. We argue that this leaves speed on the table because real decoding steps exhibit heterogeneous confidence profiles. We propose : selecting parallel commit sets from the full sorted confidence profile rather than a single worst-case confidence. The resulting rule is a heterogeneous-confidence generalization of Fast-dLLM's factor selector and it recovers the previous rule exactly in the equal-confidence case and adds a provable when the selected tokens have uneven confidences. Fast-dLLM++ leaves the model, diffusion process, and cache implementation entirely unchanged, making it a drop-in replacement for existing Fast-dLLM decoding. Experiments on GSM8K, MATH, HumanEval, and MBPP with the LLaDA-8B model show that the theoretical improvement translates directly into empirical gains: profile-aware selection improves the accuracy-throughput frontier by exploiting safe parallelism that weakest-token rules miss, achieving up to 37% higher throughput at comparable accuracy. Our anonymous code release is at https://anonymous.4open.science/r/fast-dllm-plus-plus/.
PaperID: 4666, Poster
Abstract: Integrating external memory into Large Language Models (LLMs) typically faces a trade-off between flexibility and depth. Explicit text-based retrieval keeps memory editable, but often acts as a static prompt prefix, increasing context length and interacting with the reasoning process only indirectly. Conversely, parametric memory can influence internal computation more deeply, but updating memory content usually requires additional optimization and may couple domain-specific information with the model's general reasoning behavior. In this work, we introduce MemPilot, a decoupled latent memory framework that dynamically guides the hidden reasoning states of frozen LLMs. MemPilot constructs an external bank of compact latent representations from offline reasoning trajectories, retrieves candidate memory entries, and integrates their latent memory tokens into hidden states via cross-attention and gated residual fusion. By separating domain-specific memory content from the learned memory-use mechanism, MemPilot enables target-domain adaptation through offline memory-bank substitution without updating the base model. Experiments across QA, coding, and mathematical reasoning benchmarks show that MemPilot achieves strong in-domain performance and more robust cross-domain transfer than other memory-augmented baselines. These results suggest that LLMs can benefit from reusable latent memory-use mechanisms while keeping memory content external and replaceable.
PaperID: 4667, Poster
Abstract: Autoregressive (AR) video diffusion has emerged as a leading paradigm for efficient video generation, yet its practical deployment remains bottlenecked by the trade-off between generation quality and inference efficiency. We attribute this bottleneck to three error sources in AR rollouts: the base model error, the train–inference distribution mismatch, and the upstream error propagation. Building on this decomposition, we derive closed-form upper bounds on the per-step and cumulative errors, providing a principled mechanism for jointly optimizing generation quality and inference efficiency. Guided by this analysis, we propose istribution co-optimized distillation framework with three complementary modules that each target a distinct error term in the bound. First, for ODE initialization, we adopt a Next-Node Trajectory Distillation (NTD) objective and prove a strictly tighter error bound than conventional consistency distillation, reducing the base model error at its source. Second, we introduce Chunk-Graded Noising (CGN), a monotonically non-decreasing per-chunk noising schedule that provably narrows the train–inference noise mismatch and removes redundant KV Cache updates, yielding substantial speedups at no quality cost. Third, we propose Orthogonal-Parallel DMD (OP-DMD), a distribution-matching loss that preserves the original DMD gradient direction while explicitly suppressing unsupervised off-axis drift in the orthogonal subspace, improving both training stability and final generation quality. Extensive experiments show that TriTD consistently surpasses state-of-the-art baselines in visual quality while reaching a real-time throughput of on a single H100 GPU, delivering a balanced quality-efficiency solution for interactive streaming video generation.
Authors:
Ziheng Chen, Jiali Cheng, Zezhong Fan, Diyuan Wu, Hadi Amiri, Gabriele Tolomei, Yang ZhangAbstract: Generative recommendation formulates next-item prediction as autoregressive generation over semantic ID (SID) sequences derived from users' historical interactions, making modern recommender systems structurally similar to large language models (LLMs). As privacy and safety concerns grow, these systems increasingly require concept unlearning to remove sensitive or harmful concepts associated with items. However, existing LLM unlearning methods cannot be directly applied to generative recommendation. Unlike word tokens with explicit semantics, SIDs are abstract identifiers that are often shared by both forget and retain items, leading to severe conflicts between concept removal and recommendation utility preservation. To address this challenge, we propose TRACER, an end-to-end concept unlearning framework based on token reassignment. Rather than directly suppressing shared SIDs, TRACER reassigns concept-related items to alternative tokens that better facilitate forgetting while minimizing side effects on retained items. We further introduce a coherence regularizer to preserve semantic consistency among retain items during unlearning. Experiments on real-world recommendation datasets demonstrate that TRACER effectively removes target concepts while substantially better preserving recommendation utility than existing unlearning baselines.
PaperID: 4669, Poster
Authors: Mattéo Clémot, Julie Digne, Julien Tierny
Abstract: Multi-parameter persistent homology is a rapidly developing branch of topological data analysis that improves the robustness of single-parameter persistent homology to outliers, while still capturing the metric characteristics of the data. However, a notable limitation is its lack of scalability. In this paper, we introduce a novel approach for efficiently computing 2-parameter persistent homology on large point sets. Our work extends the Flood filtration, originally developed for single-parameter persistence. Our construction, called the sublevel Flood bifiltration, offers a scalable approximation of the sublevel Flood bifiltration. We show that it benefits from theoretical stability properties and describe how to compute it efficiently. We demonstrate the performance of our approach in classification tasks on low-dimensional synthetic datasets, where density awareness is critical, as well as on real-world time series datasets.
Abstract: In decoder-only (causal) transformers, the computation graph created by causal masking routes information through both direct-path attention and indirect paths formed by intermediate tokens. We denote these indirect paths between token pairs as their runways. We argue that certain failure modes of causal transformers as observed by a growing body of recent works are likely exacerbated by a misalignment between these two information propagation modes. We formalize runway cascade as a phenomenon whereby this misalignment results in redundancies and irrelevant information cascading to token representations despite adequately learned attention patterns. As a solution, we propose runway-aware rewiring as a more explicit way of incorporating runway context directly into each token's direct-path attention. This mechanism re-wires the attention pattern for each token based on a summary of its runway landscape, enabling awareness of accumulating representational influences and allowing for more balanced information propagation. Our proposed methodology introduces no additional parameters and can seamlessly be integrated into standard attention mechanism. Empirically, our rewired transformer results in steady improvements in general language modeling as well as noticeably stronger information retrieval and extrapolation abilities compared to standard transformers.
PaperID: 4671, Poster
Abstract: The practical deployment of large language models on resource-limited devices is constrained by three coupled costs: memory footprint, energy consumption, and inference latency. Ultra-low-bit backbones such as BitNet b1.58 mitigate the first two issues by ternarizing the weights to 1.58 bits, but the inference latency remains bounded by autoregressive next-token prediction (NTP), which commits only a single token per forward pass. Multi-token prediction (MTP) improves inference latency by committing several tokens per step. However, the naive combination of MTP with a 1.58-bit backbone does not yield end-to-end gains: once 1.58-bit execution reduces the arithmetic cost, the runtime cost of the accepted-prefix dynamics of MTP dominates each iteration and erodes the savings. In this work, we propose BitMTP, an algorithm--system co-design framework that makes MTP native to a single 1.58-bit backbone. BitMTP is organized around one design principle that we term accepted-prefix alignment: at every iteration, the committed prefix is simultaneously a valid autoregressive continuation, a directly promotable KV slice, and the index of a pre-captured replay state. Three runtime mechanisms enforce this principle along the kernel, memory, and scheduling axes, respectively: a role-width-specialized 1.58-bit execution primitive, a finite-state graph atlas with shared-KV memory over a geometric segment ladder, and an optimistic verify-prefetch scheduler on the verifier-bearing decoding branches. BitMTP-2B reaches 964.2 token/s peak decode throughput, a 2.64× speedup over BitNet b1.58-2B-SFT at comparable prediction accuracy. Code and data will be released.
PaperID: 4672, Poster
Abstract: Dynamic submodular optimization has attracted significant attention in recent years, with a growing body of work studying both monotone and non-monotone variants under various constraints. A key insight underlying much of this progress is that the dynamic setting is closely related to the streaming setting, both in terms of algorithms and lower bounds. This raises a natural and fundamental question: \emphcan we obtain dynamic algorithms even for problems where no streaming analogue exists? In this paper, we answer this question affirmatively for the fundamental problem of Unconstrained Submodular Maximization (USM). Unlike other variants of submodular maximization, no streaming algorithm with non-trivial guarantees is known for USM. In the dynamic setting, a trivial 0.25-approximation follows from random sampling, but obtaining any better approximation with sublinear update time had remained open. We present the first dynamic algorithms for USM that break the 0.25-approximation barrier while maintaining sublinear update time, across several natural dynamic models: incremental updates, decremental updates with known deletion order, and fully dynamic updates with known deletion times.
Abstract: Benchmark data contamination has become a central challenge in LLM evaluation: when evaluation examples appear in the training data of one or more audited models, reported performance can be inflated and cross-model comparisons become unreliable. A broad line of training-data detection work designs scores to quantify how strongly a model memorizes a given data point, but these score-based methods lack theoretical guarantees. Recent conformal approaches provide provable false-identification control for a single model; however, applying them separately to each model can produce model-specific benchmarks, undermining fair comparison across models. In this work, we formalize multi-model benchmark decontamination as a joint selection problem and propose Joint Envelope Conformal Selection (JECS), a conformal procedure that enables global contamination rate (GCR) control under stated assumptions. Specifically, JECS computes per-model conformal p-values, aggregates them by the per-item maximum, and reconstructs a conservative envelope of the max-p null distribution from right-tail observations above a data-driven threshold. By applying the adaptive Benjamini-Hochberg (BH) procedure to the envelope-rescaled values, we select a benchmark with provable GCR control. Extensive experiments across various models and benchmarks demonstrate that JECS achieves higher power than the max-p baseline while consistently maintaining the target GCR control.
PaperID: 4674, Poster
Abstract: We study conditional average treatment effect (CATE) estimation from observational data in regimes with potentially weak overlap, when the nuisance response functions lie in a reproducing kernel Hilbert space (RKHS), but the contrast function is structurally simpler. We consider three forms of such simplicity: (i) membership in a lower-complexity RKHS with faster eigenvalue decay, (ii) a source condition relative to the nuisance kernel, or (iii) dependence on a known low-dimensional covariate representation. We develop a unified two-stage kernel ridge regression procedure that remains valid under weak overlap and attains minimax rates determined by the complexity of the contrast function. Unlike existing methods that rely on propensity-score estimation and stronger nuisance error control, our procedure avoids propensity-score estimation altogether and requires only that the outcome nuisance class be well specified, which is especially advantageous under weak overlap. We also show that a simple model-selection step over candidate contrast spaces and regularization levels satisfies an oracle inequality, enabling adaptation to unknown CATE regularity.
PaperID: 4675, Poster
Abstract: Standard In-Context Reinforcement Learning (ICRL) typically assumes stationary dynamics within an interaction history, a restriction that limits its real-world applicability. We formalize Lifelong ICRL, a more realistic paradigm in which agents must continuously adapt to diverse non-stationary dynamics, ranging from abrupt random shifts to structured temporal drifts within a single lifetime. To probe the adaptation limits of this setting, we conduct a large-scale empirical study across non-stationary environments that span discrete symbolic reasoning and continuous physics-based control. We systematically investigate three core dimensions governing adaptation: (i) Training Regimes, comparing the generalization boundaries of domain randomization, stationary training, non-stationary training, and their staged combinatorial strategies; (ii) Model Architectures, evaluating modern sequence models, including Linear Attention and hybrid designs; and (iii) Impact of Scale, analyzing the relative contributions of interaction scale and model capacity. Our results reveal how these factors jointly shape robust in-context adaptation under unknown dynamics, yielding actionable insights for the design of generalist agents.
PaperID: 4676, Poster
Abstract: Recent 3D large multimodal models (3D-LMMs) rely on a visual bottleneck to compress complex 3D scene evidence into a limited number of visual tokens compatible with large language models (LLMs). Current visual bottlenecks, however, often passively compress heterogeneous 3D evidence into a homogeneous object-centric token sequence, leaving the spatial organization of the scene under-represented. This under-representation forces the LLM to recover spatial relations from a flattened token sequence, leading to unstable reasoning in relation-intensive and spatially ambiguous scenes. To address this issue, we propose , an active scene-state construction framework for unified 3D scene understanding. SceneScaffold reformulates the visual bottleneck from a passive feature compressor into an active scene organizer, constructing a role-aware spatial scaffold before language reasoning. Specifically, SceneScaffold organizes superpoint-level visual evidence into scene-state components with distinct structural roles: entity states preserve core object semantics, scene-frame states maintain spatial references via boundary and region anchors, relation states encode object-environment interaction cues, and a global summary provides compact context. Through this role-aware construction, SceneScaffold provides the LLM with a spatially organized scene representation before language reasoning. Experiments on unified 3D scene understanding tasks, including 3D visual grounding, question answering, and dense captioning, demonstrate the effectiveness of SceneScaffold, while diagnostic results further show its applicability to relation-intensive and spatially ambiguous cases.
PaperID: 4677, Poster
Authors:
Hosung Jeon, Jun Y Jeong, Sangwoon Kwak, Joonsoo Kim, Sangmin Kim, Jaesik Park, Won-Sik Cheong, Hyon-Gon ChooAbstract: Spatiotemporal scene reconstruction with Dynamic Gaussian Splatting (DGS) fundamentally depends on perfect temporal synchronization across all cameras, an assumption rarely achievable in practice. Industry-standard synchronization requires expensive specialized hardware and complex setup, while data-driven alternatives including audio-based alignment, geometry matching, and pose tracking remain limited by restrictive environmental or scene-specific assumptions. To overcome these limitations, we propose Phase-DGS, a framework that replaces conventional timestamps with a semantic phase representing each frame's inherent state within the underlying motion cycle. Frames capturing the same motion state share the same phase regardless of which camera records them or when, naturally establishing spatiotemporal correspondence across unsynchronized cameras. We validate Phase-DGS under temporal offsets, frame drops, and camera freezes, achieving up to +8.5 dB PSNR improvement over baselines while enabling seamless integration with existing DGS backbones including RealTime-4DGS and FreeTimeGS.
PaperID: 4678, Poster
Abstract: Multi-turn agentic reinforcement learning with verifiable rewards enables language models to reason with tools, but standard GRPO-style training often collapses: validation accuracy peaks and then falls while the overall frequency of tool calls remains nearly unchanged. We identify a credit-assignment mechanism behind this failure. With terminal correctness rewards, every tool-related token receives the trajectory’s scalar advantage, making credit : the model cannot distinguish useful tool-use steps from unhelpful ones when they share the same final outcome. This transition-blind credit induces two biases. At the token level, failed rollouts with denser tool use can assign negative subset credit to tool tokens. At the trajectory level, token-mean reduction implicitly gives longer failed rollouts larger influence. These biases form a self-sustaining loop that suppresses useful tool interaction without visibly reducing tool-call frequency. To address this, we propose (TACA), a novel method that makes tool-token credit depend on the utility of the underlying tool interaction rather than only the terminal trajectory outcome. By restoring transition-sensitive credit and stabilizing sparse tool-use cases, TACA consistently outperforms agentic and non-agentic baselines across different models and multiple reasoning benchmarks.
Authors:
Matteo Biagetti, Mathieu Carrière, Francesco Conti, Enrico M Ferrari, Sven Heydenreich, Karthik ViswanathanAbstract: Persistence diagrams provide stable, interpretable summaries of geometric and topological structure, and are useful for simulation-based inference when important information is not captured by low-order statistics. In practice, however, persistence-based pipelines require hand-chosen filtrations, vectorizations, and compressors, usually without an objective tied directly to parameter uncertainty. We introduce TopoFisher, a differentiable persistent-homology pipeline that learns topological summaries by maximizing local Gaussian Fisher information. From simulations near a fiducial parameter value, TopoFisher optimizes trainable filtrations, diagram vectorizations, and compressors without posterior samples or supervised regression targets, while preserving the inductive bias of stable topological descriptors. We also give sufficient regularity conditions under which the log-determinant Fisher loss is locally Lipschitz in the trainable parameters. Controlled experiments on noisy spirals and Gaussian random fields, where the total Fisher information is known, validate the pipeline: TopoFisher recovers a large fraction of the available information and improves over fixed topological vectorizations. Our main results are on weak gravitational lensing, a high-dimensional non-Gaussian field-inference problem from cosmology. There, both learned topological summaries, a fixed cubical filtration with a learned PersLay vectorization, and a learned CNN filtration with the same PersLay vectorization, reach \log|F|\approx 21, compared with 13.8 for the power spectrum, 17.1 for peak counts, and 19.3 for wavelet scattering, approaching an unconstrained Information Maximising Neural Network baseline (22.4) with up to ~80× fewer parameters. More importantly, the fixed-filtration variant generalizes better: under simulator shift from lognormal to LPT-based maps it retains \log|F|=19.24 while the neural baseline drops to 8.9, and in neural posterior estimation it yields tighter constraints than the neural baseline, power spectrum, peak counts, and wavelet scattering. These results suggest Fisher-based topological optimizations as a robust, parameter-efficient front end for simulation-based inference.
PaperID: 4680, Poster
Authors: quansiLi, Haijing Yu
Abstract: Trade flows are a central variable for studying international trade networks. Traditional economic models, grounded in economic theory and mathematical derivation, provide structured representations of bilateral trade flows. However, these representations often rely on fixed functional forms, have limited ability to incorporate high-dimensional features, do not directly capture trade inertia, and scale poorly to large and dynamic trade networks. In recent years, graph neural networks have been increasingly used to address some of these limitations because they are naturally compatible with networked trade data and offer strong scalability. However, the feature aggregation process in existing graph neural models remains weakly interpretable from an economic perspective. To address this limitation, we propose MRTGNN, a graph neural network model built on a theory-guided MRTGNN layer derived from the fixed-point equations of multilateral resistance in structural gravity. Theoretically, we prove that an L-layer MRTGNN can approximate L iterations of a multilateral-resistance fixed-point operator on a compact domain, with an explicit bound on the iteration error. Empirically, we extensively evaluate MRTGNN on panel data covering 176 countries from 2002 to 2019 against various representative baselines. The full model achieves the best RMSE, R^2, and Spearman rank correlation across the reported baselines. Under counterfactual tariff shocks involving China, it produces a Network Spillover Ratio one order of magnitude higher than a generic graph-attention baseline. These results show that theory-guided message passing improves trade flow prediction while enabling more interpretable counterfactual reasoning in global trade networks.
PaperID: 4681, Poster
Abstract: Generative actors are transforming offline reinforcement learning (RL) by enabling expressive policy classes that model complex action distributions. However, this expressiveness also exposes a key challenge in heterogeneous datasets: generative policies can reproduce unreliable action modes whose return distributions exhibit high variance, occasionally yielding high returns by chance but lacking consistency. Consequently, maximizing the expected Q-value alone is insufficient for identifying reliable actions. We propose VAN-Flow (Variance-Averse n-step Flow), a framework that promotes reliable actions in generative offline RL. VAN-Flow combines (i) a categorical distributional critic, (ii) a variance-averse expectation operator that smoothly reweights atom probabilities to favor actions with both high returns and low dispersion, and (iii) a flow-matching generative actor guided via rejection sampling. Unlike CVaR or mean--variance objectives, the operator redistributes probability mass over the categorical return distribution without hard truncation or auxiliary penalty terms. Across more than 40 tasks from D4RL and OGBench, VAN-Flow consistently outperforms strong baselines, with the largest gains in long-horizon and high-variance regimes where reliable action selection becomes critical.
Authors:
Hengli Li, Chenxi Li, Tong Wu, Xuekai Zhu, Yuxuan Wang, Zhaoxin Yu, Eric Hanchen Jiang, Song-Chun Zhu, Zixia Jia, Ying Nian Wu, Zilong ZhengAbstract: Large Language Models (LLMs) typically reason through explicit, step-by-step natural-language traces. Humans, however, also rely on non-linguistic, unconscious processes, such as the inspirations that emerge during the incubation period. In this work, we introduce LatentSeek, a novel framework designed to enhance the reasoning capabilities of LLMs through Test-Time Instance-level Policy Gradient within the model’s latent space—thus complementing explicit natural-language steps. LatentSeek employs policy gradient optimization to iteratively refine latent representations, guided solely by a self-generated reward signal. This allows the model to adapt its reasoning trajectory dynamically on a per-instance basis. Empirical evaluations across diverse benchmarks, GSM8K, MATH-500, and AIME2024 as well as multiple LLM families (e.g., LLaMA, Qwen) demonstrate that LatentSeek outperforms established baselines, including Chain-of-Thought (CoT), Best-of-N (BoN) and training-based methods. Further analysis indicates that LatentSeek is computationally efficient, typically converging within a few optimization iterations for average-level problems. Moreover, the model's performance improves as the number of latent update iterations increases, highlighting the benefits of exploring within the latent space. These findings highlight LatentSeek as a lightweight and effective paradigm for improving the reasoning capabilities of LLMs without changing their parameters.
Authors: Jeonghwan Cheon, Jaehyuk Bae, Se-Bum Paik
Abstract: Backpropagation is the cornerstone of deep learning, but its reliance on symmetric weight transport and global synchronization makes it computationally expensive and biologically implausible. Feedback alignment offers a promising alternative by approximating error gradients through fixed random feedback, thereby avoiding symmetric weight transport. However, this approach often struggles with poor learning performance and instability, especially in deep networks. Here, we show that a one-time soft alignment between forward and feedback weights at initialization enables deep networks to achieve performance comparable to backpropagation, without requiring weight transport during learning. This simple initialization condition guides stable error minimization in the loss landscape, improving network trainability. Spectral analyses further reveal that initial alignment promotes smoother gradient flow and convergence to flatter minima, resulting in better generalization and robustness. Notably, we also find that allowing moderate deviations from exact weight symmetry can improve adversarial robustness compared to standard backpropagation. These findings demonstrate that a simple initialization strategy can enable robust and effective learning in deep networks without dynamic weight transport, providing a resource-efficient alternative to standard backpropagation.
PaperID: 4684, Poster
Abstract: Single-axis mitigations of reward-model biases (e.g., reducing reliance of the proxy reward on length, sycophancy, or style) can rotate optimization pressure onto correlated proxies rather than eliminate it, a failure mode we call reward bias substitution. We formalize the underlying measurement-vs-optimization gap between the audit distributions where mitigations are validated and the policy distributions where optimization realizes their effects. We introduce a taxonomy, instantiated in closed form, classifying single-axis mitigation outcomes into successful mitigation, bias substitution, overcorrection, silent non-op, and audit-distribution sensitivity. We prove that single-axis mitigation methods cannot be validated by audit-distribution-only evaluation: successful mitigation, bias substitution, and overcorrection produce structurally identical observables under ranking accuracy and win-rate scoring, regardless of benchmarks richness. Augmenting evaluation with policy-induced distributions provably closes the gap and we give actionable prescriptions for mitigation methods and benchmarks. Across published preference-learning mitigation work, we identify bias substitution regimes in results not previously connected to a unified failure mode. Our experiments also show that a published length-debiasing operator zeros pooled reward–length correlation but flips sign within-prompt on three of four SOTA reward models with true reward degrading on two, and that length–sycophancy coupling reverses under human–LLM judge disagreement across eight model families.
PaperID: 4685, Poster
Abstract: Few-step distillation for video diffusion models has attracted significant attention, driven by the urgent demand for efficient deployment in real-world scenarios. However, Distribution Matching Distillation (DMD), a leading paradigm, tends to degrade under limited NFE budgets, manifesting in video generation as layout instability, oversaturation, and broken motion dynamics. We trace this failure to a structural limitation: standard DMD is an intra-sample distribution-matching objective with coordinate-wise gradients, and thus imposes no explicit constraint on the relational geometry across batch elements or temporal frames, leaving the underlying copula largely unregulated. Combined with the mode-seeking tendency of its reverse-KL objective, this absence of relational guidance makes DMD prone to collapsing into local optima in the few-step regime. Motivated by this insight, we propose Copula-aware DMD (CoDMD), a lightweight relational regularizer that reuses score estimates already produced by the frozen teacher and the online fake model to construct pairwise relation matrices across samples and frames.These are matched through a supplementary distributional objective that requires no additional networks, datasets, or sampling trajectories. On the Wan-2.1-T2V model series at 1.3B \& 14B scales, CoDMD distills 50-step teachers into 4-step students, achieving an approximate 25× speed-up while attaining VBench scores of 84.46 \& 84.87, outperforming prior trajectory-based (rCM 82.81 \& 84.05) and distribution-based (DMD 83.38 \& 83.81) methods.
Authors: Dachuan Song, Junyu Yin, Zechen Hu, Xuan Wang
Abstract: Modern sequence models are typically trained at a fixed computational capacity, while real-world applications require deployment across platforms with different resource constraints. Existing approaches either train or distill separate compact models for selected budgets, or use elastic architectures that shrink dimensions such as width, depth, Feed-Forward Network (FFN) size, or SSM state dimension. In this paper, we propose Elastic Spectral State Space Models (ES-SSM), a train-once, export-many sequence modeling framework that gains elasticity through spectral approximation of the SSM sequence operator. ES-SSM builds on Hankel spectral filtering for state space models, where long-range token mixing is represented through fixed Hankel spectral channels that define an operator-level approximation resolution. To make this spectral elasticity reliably deployable, ES-SSM combines input-adaptive channel-wise gates with budget dropout, training the same spectral prefixes that are used by compact models at inference time. This encourages low-index channels to become predictive on their own, while higher-index channels provide refinements for larger budgets. We evaluate ES-SSM across byte-level language modeling, Long Range Arena, Speech Commands V2, and offline reinforcement learning benchmarks. Across these settings, a single trained ES-SSM can be truncated to competitive compact models compared with modern Transformer and SSM baselines at similar parameter scales. Furthermore, by testing under various runtime budgets, we observe smooth and stable quality-cost curves over a wide range of truncation levels.
PaperID: 4687, Poster
Abstract: Distributed training is essential for scaling LLM training across thousands of GPUs. However, as distributed training requires complex implementations, they are prone to silent bugs, which do not produce explicit error signals but lead to in correct training outcomes. Common debugging practices based on monitoring training loss or gradient norm curves are slow or unable to detect bugs and also do not help localize bugs. We design and implement nnTrace, the first systematic differential testing system for detecting and localizing silent bugs in distributed training. nnTrace aligns intermediate tensors from distributed training with those from a trusted reference implementation. To properly compare the floating-point values in the corresponding tensors, we propose a novel mathematical analysis that provides a guideline for setting tolerances, enabling nnTrace to distinguish bug-induced errors from numerical errors. Experimental results demonstrate that nnTrace effectively detects 11 existing bugs and 3 new bugs in the widely used Megatron-LM framework. nnTrace is effective in various training recipes, including low-precision recipes involving BF16 and FP8. Notably, Megatron-LM has already adopted the method proposed by nnTrace in its development workflow. Our code is available at https://anonymous.4open.science/r/NECK-3C61/.
Abstract: Diffusion language models (Diffusion-LMs) generate text through iterative denoising, exposing a temporal structure that is largely absent from autoregressive decoding. In this paper, we show that this temporal structure provides a useful control axis for generation diversity: early denoising steps mainly determine high-level semantic trajectories, while later steps refine lexical realization. Motivated by this observation, we propose Time-Annealed Perturbation Sampling (TAPS), a training-free inference strategy that samples nearby conditioning trajectories through time-aware, manifold-constrained perturbations. TAPS encourages semantic branching during early denoising and anneals the perturbation away before refinement, improving exploration while preserving prompt alignment, generation quality, and reasoning ability. Experiments on multiple Diffusion-LM backbones, including non-autoregressive and semi-autoregressive models, show that TAPS consistently improves semantic and lexical diversity across open-ended and instruction-following generation tasks, while preserving reasoning ability on verifiable reasoning benchmarks with negligible overhead.
PaperID: 4689, Poster
Abstract: We introduce \textttMUNI, an end-to-end multimodal latent diffusion framework for any-to-any generation that unifies subset-conditioned cross-modal generation and unconditional joint sampling through a shared stochastic latent. Existing multimodal generative models are largely LLM-based, which limits leveraging modality-specific generators and requires text-paired data for training. Recent diffusion- and flow-based any-to-any extensions take a different direction but still rely on text-aligned embeddings, fully-paired training, or matched-dimensionality deterministic mappings. \textttMUNI rests on two complementary contributions, one architectural and one in the training objective. First, we extend latent diffusion to multimodal any-to-any generation end-to-end: instead of the standard two-stage recipe that precomputes a frozen latent space and then fits a prior over it, \textttMUNI jointly trains modality-specific encoders, expressive decoders, and a single shared flow-based prior under one objective. Second, we identify that the standard aggregation rules of multimodal variational inference are insufficient once coupled with a learned prior and expressive decoders. A suitable shared latent must simultaneously satisfy coherence across generated modalities, predictive sufficiency of subset latents, and minimality of the latent content. We propose a routed training objective whose structural choices align the latent with these criteria and admit a minimal-sufficiency characterization in the realizable setting. Experiments on PolyMNIST-Quadrant-Labels and a large-scale image-text-audio benchmark show \textttMUNI matching or exceeding the strongest baselines on conditional generation while opening its largest margins on unconditional coherence.
PaperID: 4690, Poster
Authors: Matthew Brun, Andy Sun
Abstract: Success conditioning is a strategy for improving decision-making policies in stochastic environments; it updates a policy by increasing the probability of taking actions that yield successful outcomes. Success conditioning is common to many reinforcement learning applications, yet its limiting behavior and convergence rates are not well understood. In this work, we demonstrate that success conditioning converges to an optimal policy on a broad class of Markov decision processes (MDP). We also derive convergence rates in some common settings. For discounted MDPs, we prove convergence within \mathcalO(1/\varepsilon^p) iterations to an \varepsilon-optimal policy, where the exponent p depends on problem data. For single-period MDPs, \mathcalO(1/\varepsilon) iterations are required.
PaperID: 4691, Poster
Abstract: Multi-step off-policy TD methods correct the behavior--target mismatch through a \emphproduct of per-step trace coefficients, yet this product inherently amplifies variance or discards long-horizon signal, making the cap length a sensitive hyperparameter. We propose MINT (Meeting-time INdicators for Truncation), which replaces coefficient products with a binary coupling indicator: one until the behavior and target trajectories meet, zero thereafter. Because the indicator is idempotent, variance depends only on the meeting-time survival probability---not on ratio products---and a single action sample estimates the per-step coupling probability without mixing-time knowledge. \textttMINT contracts in supremum norm to Q^\pi under arbitrary behavior policies; empirically, \textttMINT outperforms handcrafted multi-step baselines across MuJoCo and DeepMind Control tasks without cap-length tuning.
Authors: Zhan Qu, Michael Färber
Abstract: Predicting future clinical events from longitudinal electronic health records (EHRs) requires selecting plausible outcomes from a large and structured event space under sparse observations. While clinical coding systems provide hierarchical organization of events, cross-modal and temporal relationships are not explicitly specified and must instead be inferred from data, making prediction difficult for weakly observed longitudinal transitions. We introduce Risk Horizons, a geometry-aware framework for constructing patient-specific candidate spaces for multi-modal next-visit prediction. Risk Horizons combines deterministic coding hierarchies with data-driven lagged cross-modal associations, embeds the resulting clinical graph in hyperbolic space, and retrieves candidate futures using directional risk cones. This reframes longitudinal prediction as ranking within a compact, clinically coherent hypothesis space rather than scoring an unconstrained vocabulary. Experiments on MIMIC-IV and eICU demonstrate competitive next-visit prediction performance, with consistently improved hierarchy consistency across diagnoses, procedures, and medications. Further analysis suggests that hyperbolic structured candidate retrieval is the primary driver of performance, while LLMs are effective as constrained inference-time rerankers operating over clinically grounded candidate sets.
Authors:
Haotian Zhao, Songlin Zhou, Yuxin Zhang, Stephen S Yau, Wenyu Zhang, Tianlun, Tianshu Zhu, Yifeng Huang, yucheng-zeng, Jingnan gu, Daxiang Dong, Jianmin WuAbstract: Reinforcement learning (RL) has substantially improved the ability of large language model (LLM) agents to interact with environments and solve multi-turn tasks. However, effective agentic RL remains challenging: sparse outcome-only rewards provide limited guidance for assigning credit to individual steps within long interaction trajectories. Existing approaches often introduce dense intermediate supervision, such as process reward models or auxiliary self-supervised signals, which increases supervision and tuning complexity and may limit generalization across tasks and domains. We present AEM, a supervision-free credit assignment method that adaptively modulates entropy dynamics during RL training to improve the exploration-exploitation trade-off. Since in agentic RL the environment is typically affected by a complete response, rather than an individual token, our analysis lifts entropy dynamics from the token level to the response level, aligning uncertainty estimation with the effective action granularity of LLM agents and reducing sensitivity to token-level sampling noise. We further show that entropy drift under natural-gradient updates is governed by the interaction between the sampled-response advantage and its relative surprisal. Motivated by this result, AEM derives a practical response-level uncertainty proxy and uses it to rescale advantages, leveraging the evolving balance between positive and negative samples to naturally transition from exploration to exploitation. Extensive experiments on ALFWorld, WebShop, and SWE-bench-Verified with models ranging from 1.5B to 32B demonstrate that AEM consistently improves strong RL baselines, including a +1.4% gain when integrated into a state-of-the-art software-engineering RL training framework.
PaperID: 4694, Poster
Abstract: Text-to-image (T2I) diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference. However, their ability to generate harmful or undesired content poses significant safety risks. Data-driven unlearning methods suppress targeted generations by fine-tuning model weights using specialized unlearning objectives. Crucially, these objectives implicitly rely on multi-step denoising dynamics, an assumption that breaks down for few-step distilled (FSD) models, resulting in ineffective forgetting. Furthermore, performing unlearning on the non-distilled base model and subsequently re-distilling it to obtain an unlearned FSD model incurs substantial computational and time overhead, making it impractical in many settings. Hence, we address this limitation with a preference-driven unlearning framework that revisits Direct Preference Optimization (DPO) for diffusion models. We show that standard DPO and its unlearning derivatives, formulated around noise-prediction error, transfer poorly to FSD models due to their altered generation dynamics. To overcome this, we introduce a modified preference optimization formulation explicitly aligned with the few-step generation properties, enabling direct concept removal in FSD models while preserving few-step efficiency and maintaining strong retention of desirable (non-targeted) capabilities. We evaluate our framework primarily on identity and NSFW (nudity) removal tasks and also extend our method to object-level unlearning. Extensive experiments demonstrate consistent and effective forgetting, and strong retention performance, establishing our method as a practical and principled solution for unlearning in FSD models.
Abstract: Graph Neural Networks (GNNs) perform computations on graphs by routing the signal between graph regions using a graph shift operator or a message passing scheme. Often, the propagation of the signal leads to a loss of information, where the signal tends to diffuse across the graph instead of being deliberately routed between regions of interest. Two notions that depict this phenomenon are oversmoothing and oversquashing. In this paper, we propose an alternative approach for modeling signal propagation, inspired by quantum mechanics, using the notion of observables. Specifically, we model the place in the graph where the signal lies, how much the signal is concentrated there, and how much of the signal is propagated towards a location of interest when applying a GNN. Using these new concepts, we prove that standard spectral GNNs have poor signal propagation capabilities. We then propose a new type of spectral GNN, termed Schrödinger GNN, which we show has a superior capacity to route the signal across the graph.
Abstract: We present a new class of efficient attention mechanisms applying universal 3D Relative Positional Encoding (RPE) methods given by arbitrary integrable modulation functions f. They lead to the new class of 3D-Transformer models, called RelFlexformers, flexibly integrating those RPEs, and characterized by the O(L \log L) time complexity of the attention computation for the L-length input sequences. RelFlexformers builds on the theory of the Non-Uniform Fourier Transform (NU-FFT), naturally generalizing several existing efficient RPE-attention methods from structured settings with tokens homogeneously embedded in unweighted grids into general non-structured heterogeneous scenarios, where tokens' positions are arbitrarily distributed in the corresponding 3D spaces. As such, RelFlexformers can be applied in particular to model point clouds. Our extensive empirical evaluation on a large portfolio of 3D datasets confirms quality improvements provided by the NU-FFT-driven attention modulation techniques in the RelFlexformers.
PaperID: 4697, Poster
Abstract: Diffusion language models (DLMs) provide a bidirectional generation framework naturally suited for infilling, yet their performance is constrained by the pre-specified infilling length. In this paper, we reveal that DLMs possess an inherent ability to discover the correct infilling length. We identify two key statistical phenomena in the first-step denoising confidence: a local Oracle Peak that emerges near the ground-truth length and a systematic Length Bias that often obscures this signal. By leveraging this signal and calibrating the bias, our training-free method CAL (Calibrated Adaptive Length) enables DLMs to approximate the optimal length through an efficient search before formal decoding. Empirical evaluations demonstrate that CAL improves Pass@1 by up to 47.7% over fixed-length baselines and 40.5% over chat-based adaptive methods in code infilling, while boosting BLEU-2 and ROUGE-L by up to 8.5% and 9.9% in text infilling. These results demonstrate that CAL paves the way for robust DLM infilling without requiring any specialized training.
PaperID: 4698, Poster
Abstract: Single-Domain Generalized Object Detection (S-DGOD) aims to generalize from a single source domain to unseen domains under distribution shifts. Existing efforts mainly focus on simulating unseen domains or suppressing domain-specific factors. However, they largely overlook the intrinsic generalization capability embedded in Vision Foundation Models (VFMs). In this work, we show that unconstrained fine-tuning disrupts the geometric structure of pre-trained representations, which encodes domain-invariant semantics. Motivated by this, we propose Geometric-Invariance Fine-Tuning (GIFT), a geometric-aware fine-tuning method for VFMs. Specifically, we introduce a series of learnable Householder reflections to the pre-trained weights to encourage geometry-consistent adaptation under domain shift. In addition, we propose a lightweight channel-wise scaling to facilitate effective task-specific adaptation. Extensive experiments demonstrate that GIFT consistently outperforms state-of-the-art (SOTA) approaches on two S-DGOD benchmarks. Notably, by introducing fewer than 1% additional trainable parameters into the frozen backbone, the proposed method achieves improvements of +7.9% and +2.9% in mPC over the previous SOTA on the Cityscapes-C and Diverse Weather Dataset, respectively. This finding highlights that preserving representation geometry is crucial for robust cross-domain generalization, establishing GIFT as a principled and efficient solution to domain shift. The code is available in the supplementary material.
PaperID: 4699, Poster
Abstract: Unpaired medical image-to-image translation requires synthesizing target-modality images from source-modality images without aligned supervision. Existing methods learn a direct source-to-target mapping that entangles structure preservation and appearance generation within a single model, often distorting clinically meaningful anatomy in the translated output. We argue that these two objectives should be decoupled entirely: anatomical structure should be captured independently of generation, and the generator should operate only on a frozen structural representation. To this end, we propose a two-stage framework for unpaired medical image translation. In the first stage, we learn a shared structure space by training a structure encoder-decoder on source images, target images, and structural masks jointly, using mask reconstruction to anchor the latent space to modality-agnostic anatomy. For source images without mask annotations, a target-domain memory bank provides soft structural supervision via prototype retrieval. In the second stage, a conditional flow model learns to render target-domain appearance under fixed structural conditions. This ensures the generator inherits target appearance statistics and can focus entirely on structural conditioning rather than learning appearance from scratch. We validate our framework on OCT-to-OCTA synthesis and brain MR-to-CT translation. Our method outperforms all unpaired baselines on both tasks and, despite requiring no paired supervision, also surpasses paired baselines, improving PSNR by up to 11.2% and SSIM by up to 26.4% on OCT-to-OCTA, and achieving consistent gains across whole-image, soft-tissue, and bone regions on MR-to-CT.
Abstract: We present the first theoretical guarantees for differentially private online reinforcement learning (RL) with general function approximation, extending beyond prior work restricted to tabular and linear settings. Our approach combines a batched policy update scheme with the exponential mechanism, together with a novel regret analysis. We show that, even under general function approximation, the regret in the model-free setting under differential privacy matches the state of the art for the linear case, scaling as \widetildeO(K^3/5), where K denotes the number of episodes. As an important by-product, we also establish the first regret bound for online RL with batch update that depends on the standard complexity measure of coverability, complementing existing results based on a newly introduced Eluder-Condition class. In addition, we uncover fundamental gaps in recent results for private RL with linear function approximation, thereby clarifying its landscape.
Authors:
Asim Osman, Sasha Abramowitz, Mark Bergh, Ulrich Armel Mbou Sob, Ruan John de Kock, Omayma Mahjoub, Oussama Hidaoui, Noah De Nicola, Arnol M Fokam, Felix Chalumeau, Daniel Rajaonarivonivelomanantsoa, Siddarth Singh, Refiloe Shabe, Juan Formanek, Simon Du Toit, Arnu PretoriusAbstract: Contrastive reinforcement learning (CRL) learns goal-conditioned Q-values through a contrastive objective over state-action and goal representations, removing the need for hand-crafted reward functions. Despite impressive success in achieving viable self-supervised learning in RL, all existing CRL algorithms rely on off-policy optimisation and are mostly constrained to continuous action spaces, with little research invested in discrete environments. This leaves CRL disconnected from widely used and effective, modern on-policy training pipelines adopted across both single-agent and multi-agent RL in continuous and discrete environments. To establish a first connection, we introduce Contrastive Proximal Policy Optimisation (CPPO). CPPO is an on-policy contrastive RL algorithm that derives policy advantages directly from contrastive Q-values and optimises them via the standard PPO objective, without requiring a reward function or a replay buffer. We evaluate CPPO across continuous and discrete, single-agent and cooperative multi-agent tasks. Whilst the existence of an on-policy approach is inherently useful, we additionally observe that CPPO not only significantly outperforms the previous CRL baselines in 14 out of the 18 tasks benchmarked, but matches or exceeds reward-based PPO's performance in 12 out of the 18 tasks without any reward signal.
PaperID: 4702, Poster
Abstract: We derive a tight generalization bound for deterministic neural networks based on their last hidden layer representations. By evaluating representation quality with a parameter-free k-nearest-neighbor (k-NN) classifier, our certificate avoids traditional parameter-space complexity measures, KL divergences to a posterior, weight norms, and compression. The bound is tight: trained from scratch, it reaches an 8.0 percentage point gap to test error on CIFAR-10 (ResNet-18) and yields the first non-vacuous certificates we are aware of on CIFAR-100; with frozen DINOv2 representations, it certifies at 1.4% on CIFAR-10 and 23.1% on 1,000-class ImageNet. It decomposes additively into three empirical, independently measurable quantities: local class geometry (k-NN decoder error), decoder-network agreement, and training stability. This allows the bound to act as a diagnostic tool: we demonstrate that actively destroying local class geometry eliminates generalization while leaving dataset memorization intact, identifying class geometry as a causally relevant mechanism for generalization. Finally, we show that learned representations are significantly less sensitive to weight noise than linear readouts, suggesting why parameter-space certificates often struggle to remain tight.
Abstract: A central goal of mechanistic interpretability is to identify which internal components causally drive a language model's behavior. Because these importance estimates serve as the evidence for identifying circuits, systematic errors can lead to the misidentification of the underlying mechanisms. While activation patching provides a gold-standard causal metric, its computational cost is prohibitive at scale. Practitioners instead rely on attribution patching, a gradient-based, first-order approximation whose reliability remains poorly understood. In this work, we characterize the source of this unreliability, demonstrating that the dominant error stems from the non-linearities in the downstream network rather than local curvature at the patched component. This insight yields three practical tools: (i) a reliability score to detect untrustworthy estimates, (ii) error bounds quantifying potential attribution mis-specifications, and (iii) a Hessian-vector-product (HVP) correction that eliminates the leading-order error with only one additional backward pass. In evaluations across five model families (124M–9B parameters) and both random-token and naturalistic (name-swap) perturbations, HVP is the only second-order correction feasible at larger scale, where standard baselines like Integrated Gradients become computationally prohibitive. In comparative experiments, a multi-step HVP variant matches or exceeds the accuracy of Integrated Gradients at significantly lower compute, outperforming prior second-order baselines. These improvements lead to higher-fidelity circuit recovery on standard benchmarks and support a
Abstract: Cross-entropy (CE) is the default training loss for supervised classification, but its sample efficiency is limited when labels are scarce. Existing remedies primarily act on the data side, via augmentation, synthesis, or transfer from pretrained models; the training objective itself is rarely revisited. We revisit it here. Drawing on the classical observation that generative classifiers reach their asymptotic error with fewer samples than discriminative ones, we propose Generative Cross-Entropy (GenCE), a drop-in replacement for CE that introduces a generative learning principle into a standard discriminative network without altering the architecture or fitting a separate density model. GenCE follows from a Bayesian rewrite of the class-conditional likelihood and, in the mini-batch approximation, reduces to normalizing each sample's softmax score against the model's predictions on the batch, coupling the training signal across examples sharing a class. We extend the proper-scoring-rule framework to such non-local losses and prove that GenCE is strictly proper under a mild completeness condition: its population risk is uniquely minimized at the true posterior. Across three datasets, on two architectures and in both balanced small-data and class-imbalanced regimes, GenCE outperforms CE and other widely used losses, while also producing better-calibrated probabilities and stronger out-of-distribution detection.
PaperID: 4705, Poster
Abstract: Long chain-of-thought (CoT) reasoning in large language models often fails in recurring ways: models loop through redundant derivations, drift through uncertain intermediate states, commit prematurely to unsupported answers, or continue generating after a valid conclusion has been reached. These failures are typically treated as separate surface phenomena and addressed with symptom-specific heuristics, which (i) leave the underlying trajectory failure unchanged when they suppress the visible symptom, and (ii) provide no shared substrate for comparing fixes across models or pathologies. We instead study them as failures at the level of CoT reasoning latent state transitions. We introduce a diagnostic framework that abstracts chain-of-thought reasoning as a trajectory through discrete latent reasoning states, yielding a positional taxonomy of six reasoning pathologies, each paired with a transition-grounded detection predicate and a matched inference-time controller. We evaluate the framework through controlled-surgery experiments across multiple datasets and models, spanning arithmetic reasoning, multi-hop question answering, and knowledge-intensive reasoning. Across evaluations, we find that reasoning pathologies exhibit three main diagnostic patterns: (1) failures follow a phase-ordered structure, where termination failures are most reliably identified, early branching failures are detectable, and progression failures require finer graph-aware analyses; (2) counterfactual transition-versus-emission tests show that several detectors track latent trajectory changes rather than surface rewrites, while also exposing cases where paraphrases alter the reasoning path itself; and (3) detector-timed interventions are most effective when the detected pathology has a clear local onset, but do not uniformly dominate generic controls. Together, these findings position the framework as a reproducible diagnostic substrate for studying CoT failures, with explicit evidence boundaries and stronger closed-loop control.
Abstract: Many biological systems evolve through continuous local dynamics while switching between latent regimes defined by learning, stimulus context, internal state, or developmental stage. These processes are often observed only as unpaired longitudinal snapshots: the same cells, neurons, or animals are not tracked as matched trajectories, even though population states are sampled across successive stages. This creates two coupled challenges. First, trajectories must respect curved low-dimensional manifolds embedded in high-dimensional biological measurements. Second, the model must identify when the transport mechanism itself changes. We introduce FLUX (FLow matching for Unpaired longitudinal data with miXture-of-experts), a geometry-aware longitudinal flow-matching framework for joint transport modeling and unsupervised regime discovery. FLUX learns a data-dependent metric from pooled labeled and unlabeled observations, uses that metric to construct geometry-aware conditional paths between adjacent marginals, and decomposes the resulting velocity field into sparse expert vector fields selected by a Straight-Through Gumbel-Softmax router. Across manifold controls, a regime-switching Lorenz system, widefield cortical calcium imaging during associative learning, and embryoid body single-cell differentiation, FLUX reconstructs longitudinal transport while recovering interpretable regime structure. In neural data, the router separates early/intermediate training from late training, coinciding with the behavioral divergence of CS+ and CS- lick indices. In embryoid body differentiation, FLUX separates expression-evolution regimes associated with pluripotent and more differentiated cell populations. Ablations show that mixture-of-experts routing alone is insufficient: FLUX without geometric learning can fit local transport but fails or weakens regime discovery when regimes are encoded in local dynamics. These results suggest that geometry-aware velocity decomposition provides a general strategy for discovering latent biological state transitions from unpaired longitudinal snapshots.
Authors:
Haozhe Si, Yuxuan Wan, Yuqing Wang, Minh Do, Han ZhaoAbstract: Modeling hyperspectral imagery (HSI) across different sensors presents a fundamental challenge due to variations in wavelength coverage, band sampling, and channel dimensionality. As a result, models trained under a fixed spectral configuration often fail to generalize to other sensors. Existing Vision Transformer (ViT) approaches either rely on implicit spectral modeling with fixed channel assumptions or adopt explicit spatial–spectral attention with prohibitive computational cost, leading to a fundamental trade-off between efficiency and expressiveness. In this work, we introduce Low-rank Efficient Spatial–Spectral ViT (LESSViT), a sensor-flexible architecture for cross-spectral generalization. LESSViT is built on LESS Attention, a structured low-rank factorization that models joint spatial–spectral interactions through separable spatial and spectral components, reducing the complexity of full spatial–spectral attention from \mathcalO(N^2 C^2) to \mathcalO(rNC), where N is the number of spatial tokens, C is the number of spectral channels, and r is the rank of the low-rank approximation. We further incorporate channel-agnostic patch embedding and wavelength-aware positional encoding to support flexible spectral inputs. To enable efficient and robust pretraining, we introduce a hyperspectral masked autoencoder (HyperMAE) with decoupled spatial–spectral masking and hierarchical channel sampling. We evaluate LESSViT under a cross-spectral generalization setting that simulates cross-sensor variability. Experiments on the SpectralEarth benchmark demonstrate that LESSViT improves robustness under spectral shifts while remaining competitive in-distribution, and explicit and efficient spatial–spectral modeling is essential for scalable and generalizable hyperspectral representation learning.
Authors:
Siddharth Setlur, Djordje Mihajlovic, Darrick LeeAbstract: Protein language models (pLMs) encode information about protein sequences which enable downstream tasks such as structure prediction, but their internal representations are not well understood. Sparse autoencoders (SAEs) provide a promising tool to disentangle latent pLM representations into interpretable features, but existing annotation pipelines largely rely on protein-level annotations derived from database labels and LLM annotations of top activating sequences. Such annotations can overlook the localized residue-level and geometric patterns encoded by sparse features. We introduce an automated and scalable method for interpreting SAE features in ESM-2 by using geometrically inspired features of the protein \textC_\alpha backbone. Across ESM-2 8M layers, geometric annotations explain a large fraction of SAE features, expanding coverage beyond database and sequence-based methods. In particular, geometry can distinguish SAE features sharing the same database annotation, revealing substructure within known biological labels. A significant portion of SAE features activate on unannotated metagenomic protein sequences enabling us to use our SAE annotations to better understand these sequences. In addition, ablation experiments at the level of contact predictions hint toward SAE features controlling protein geometry. This provides a robust method of annotating proteins activated within SAE neurons at a residue level, providing a bridge between mechanistic interpretability and structural biology.
Abstract: Diffusion models are often trained in low-dimensional latent spaces, which are then reused for related but shifted datasets. In this work, we study when such latent reuse remains reliable under distribution shift. We consider a source-target setting in which both datasets are approximately low-dimensional but may lie near different subspaces. We show that freezing and reusing a source latent space induces a target-domain score error governed by two quantities: the principal-angle misalignment between the source and target subspaces, and the target ambient noise amplified by the diffusion time scale. Motivated by these limits, we further study mixed source-target training and characterize how the required shared latent dimension depends on the relative geometry of the two distributions. Our results provide theoretical guidance on when latent reuse is reliable and when learning a shared representation may be necessary.
PaperID: 4710, Poster
Abstract: 3D Gaussian Splatting enables real-time novel view synthesis with high fidelity, yet its sparse and unorganized Gaussian anchors make compression challenging without harming structure and cross-view consistency. A key challenge is that effective compression depends not only on reducing parameter redundancy, but also on identifying where the compressed model is most fragile and how limited representation capacity should be allocated accordingly. We present an RL-guided distillation framework for 3DGS compression (RLGS). RLGS first uses a lightweight reinforcement learning policy to predict informative viewpoint offsets for adaptive distillation. It then distills a compact student 3DGS from a high-quality teacher under multi-view rendering supervision. Finally, it introduces an Unbalanced Optimal Transport based structural selection module to align teacher-student anchor distributions in voxelized space, thereby retaining the most informative anchors during pruning and allocation. Under the same bitrate constraints, our method consistently outperforms strong 3DGS compression baselines in rendering quality. At matched rendering quality, it further reduces model size while preserving structural fidelity, achieving up to about 15% additional size reduction at comparable rendering quality over prior methods.
PaperID: 4711, Poster
Authors: Nicolas Hochuli, Lorenzo Miele, Kristina Shea, Tino Stankovic
Abstract: Planar graphs are central to applications across science and engineering, yet existing generators provide limited support for goal-directed generation under hard structural and geometric feasibility constraints. We propose a dataset-free method for generating planar graph embeddings by combining parametric graph grammars with safe reinforcement learning to optimize generic task-specific objectives while satisfying constraints during construction. We formulate the generation process as a constrained Markov decision process, where the graph grammar defines the state and action spaces. We further introduce an action projection that maps sampled actions toward state-dependent safe sets, improving constraint satisfaction during training. In contrast to classical graph generators and deep generative models, which typically offer limited goal-directed control or rely on weak constraint satisfaction, our method constructs feasible planar graph embeddings directly during generation. We also introduce a benchmark suite for constrained and goal-directed planar graph generation, together with classical and deep generative baselines. Across all benchmark tasks, our method consistently outperforms baselines while satisfying the formulated constraints.
PaperID: 4712, Poster
Abstract: Auto-bidding is a core algorithmic component in online advertising auctions. In practical cold-start settings, iterative training is commonly used to compensate for low-quality, narrowly supported historical data. Unfortunately, in iterative training, representative diffusion-based AI-Generated Bidding (AIGB) methods fail to sustain extrapolation beyond the narrow data support and stagnate at a suboptimal level. In this paper, we theoretically attribute this stagnation to Signal-to-Noise Ratio (SNR) collapse: the weak radial return signal is overwhelmed by a curvature-induced penalty. To break this stagnation, we propose Iterative Scarcity-Guided Exploration (ISGE), which introduces a guidance handoff: as the return signal collapses, scarcity subsequently takes over as an exploratory signal to elevate the SNR above a critical threshold. Specifically, ISGE iteratively alternates between a \emphJudger that assigns scarcity scores to trajectories and an \emphExplorer that performs guided generation of high-scarcity and high-return trajectories, thereby bootstrapping from the narrow support with such self-generated trajectories. Extensive experiments on the industrial AuctionNet benchmark demonstrate that ISGE, starting from cold-start datasets, effectively surpasses the performance of the Full-Dataset baseline within three iterations.
Abstract: We propose SCULPT, a state-of-the Art masked discrete diffusion model (MDMs) for high resolution text-to-image synthesis. Compared with prior works on masked image generation, SCULPT addresses two key challenges. First, unlike continuously diffusion models which progressively refines image latent across the entire image, vanilla MDMs do not have self-correcting capability because discrete tokens cannot be changed once it's unmasked. Second, while scaling the vocabulary size of discrete image tokenizers can improve reconstruction quality, it also introduces optimization challenges for training generative models as per-token training signal becomes more spares. To address the first challenge, SCULPT incroprates a token-editing mechanism where the model can dynamically correct already-unmasked output tokens during inference. To address the second challenge. we proposed a Goruped Corss Entrophy (GCE) objective that assigns positive learning signal to adjacent tokens of the ground truth in the embedding space. To improve the training efficiency, we further implemented a custom fused operator that greatly reduce the VRAM requirements during training in large-vocabulary setting. Experiment results show that these innovations significant improves the training efficiency and image fidelity of masked discrete image generators.
PaperID: 4714, Poster
Abstract: Datasets in many settings naturally partition into clusters arising from sub-populations, batch effects, or aggregation across multiple sources. A common response to such heterogeneity is to ensemble learners trained on each cluster rather than fit a single model to the pooled data. Prior work motivating such approaches has typically considered settings in which both the covariate distribution and the conditional outcome model differ across clusters; the role of cluster-aware partitioning and ensembling based solely on the covariate distribution remains to be explored. We address this case for ridge-regularized least-squares regression under a linear outcome model and consider all ridge penalty values \lambda \geq 0, including the special case of the ridgeless predictor at \lambda = 0. By considering both fixed-effects and random-effects models, we argue that under random effects, an optimally tuned pooled ridge predictor always outperforms ensembles of individually optimally tuned predictors. For fixed effects, we derive a general formula for the pooled and ensembled predictors to characterize the role of both regression coefficients as well as the predictor distribution shifts. Together, these results generalize prior risk analyses of bagging and random-partition estimation using ridge and ridgeless regression predictors from the i.i.d. setting to encompass covariate shift and heterogeneity-aware partition structure.
Authors:
Tongtian Yue, Zikang Liu, Longteng Guo, Handong Li, Zhibin Wang, Jun Fu, Zijia Zhao, Xianing Chen, Jun Song, Cheng Yu, Bo Zheng, Jing LiuAbstract: Multimodal large language models (MLLMs) have rapidly advanced machine intelligence by integrating visual perception with powerful linguistic reasoning. However existing visual backbones predominantly encode videos in a rigid per-frame manner and delegate temporal modeling entirely to the language model. This paradigm fundamentally constrains long-range temporal understanding and imposes severe computational inefficiencies. To resolve these bottlenecks we propose FuxViT, a unified Vision Transformer natively supporting a 64K-token context window to seamlessly process diverse inputs ranging from static images to hour-scale videos. By employing a joint spatiotemporal pretraining paradigm, FuxViT generates highly coherent and information-rich visual representations. This strategic early fusion frees the downstream language model from reconstructing low-level temporal structures to focus entirely on high-level reasoning. Furthermore we introduce Frame Selective Attention (FSA) as a native sparse routing mechanism. FSA dynamically attends to the most relevant historical frames to achieve exceptional computational efficiency without sacrificing fine-grained spatial or temporal details. Extensive experiments across 16 image and video benchmarks demonstrate that FuxViT consistently outperforms existing baselines and establishes itself as a vastly superior alternative to standard static ViTs for multimodal understanding. We will fully open-source all model checkpoints.
Abstract: Efficient long-sequence modeling remains a central challenge for large language models, as self-attention scales quadratically with sequence length. Mamba offers a linear-time alternative through selective state space recurrence, but its predominantly diagonal state transitions restrict explicit interactions among state dimensions. We propose Motif-Mamba, a structured state space model that augments Mamba with a motif-constrained low-rank recurrent pathway. Inspired by the dynamics of three-node network motifs, the proposed pathway projects hidden states into a compact dynamical subspace, imposes motif-guided interactions, and maps the resulting dynamics back to the original state space. This design enhances cross-dimensional communication while preserving the linear-time recurrent structure of Mamba. Experiments on long-sequence extrapolation, language modeling benchmarks, and brain--computer interface decoding show consistent improvements over Mamba backbones, suggesting that motif-guided low-rank dynamics provide an effective structural prior for long-range sequence modeling.
PaperID: 4717, Poster
Authors: JUNWEI ZHAO, Qianchun Luo, Jie Wu
Abstract: UAV dehazers trained offline degrade during long-duration flights as haze density and spatial structure evolve over time. We formulate UAV dehazing as a closed-loop edge-cloud adaptation problem with an explicit reliability budget governing update frequency. The onboard model couples a frozen anchor head with a cloud-adaptive head, evaluated by a joint critic integrating a physics-based rehaze constraint and a scale-aligned infrared-edge consistency term to suppress degenerate adaptations. Our main contribution is a sequential upload criterion based on a non-negative e-process, yielding anytime-valid false-upload control for streams that satisfy conditional calibration, without requiring frame independence or exchangeability. The cloud updates only the adaptive head from uploaded segments, while the UAV accepts or rejects returned parameters through a shadow-window sign test with an acceptance budget. Experiments on five UAV datasets show consistent improvements under diverse haze shifts. Code will be publicly released.
Authors: Rahaf Abu Hara, Vaibbhav Murarri, Claudio Zito
Abstract: Existing LLM-based policy optimizers see only scalar rewards: that a policy scored 0.45, but not whether the agent got stuck in a loop, fell into a hole on the third step, or performed well on 19 out of 20 rollouts and failed catastrophically on one. We propose Reflective Prompted Policy Optimization (R2PO), a two-stage LLM framework for policy search over compact policy classes that augments scalar reward feedback with trajectory-level behavioral evidence. A Search-LLM acts as a global policy optimizer and proposes candidate policy parameters; the environment executes them; a Critic-LLM then inspects the resulting rollouts and proposes targeted parameter revisions grounded in observed states, actions, and rewards. Across ten environments, ablations show R2PO's gains arise from a design that explicitly separates global search from behavior-grounded revision and uses selection to filter high-variance edits. We further identify a dominant failure mode, salience bias. When presented with multiple rollouts, the Critic-LLM fixates on improving a single failure even when most trajectories succeed. In a three-trajectory variant, where the Critic-LLM is shown the best, worst, and median rollout from each evaluation, this behavior explains 76.6% of regressions on CartPole. R2PO mitigates this by reasoning over aggregate rollout statistics, median-trajectory selection, and a revision rule. Using a relatively small open-weight 20B-parameter model, R2PO achieves the highest mean best reward across all ten environments, while reaching near-optimal performance substantially earlier in training (e.g., near-maximum CartPole reward within ~500 episodes), and training far more stably than both deep RL and prior LLM-based methods. Together, these results show that treating trajectories as first-class in-context evidence, rather than background artifacts reduced to scalar returns, fundamentally changes how even comparatively small LLMs can search over policy spaces, enabling them to learn faster, diagnose more precisely, and reliably improve external controllers rather than tune them by trial and error.
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a dominant paradigm for enhancing Large Language Models (LLMs) reasoning, yet its reliance on external verifiers limits its scalability. Recent findings suggest that RLVR primarily functions by eliciting latent capabilities, motivating the development of verifier-free algorithms. However, in such settings, standard methods like Group Relative Policy Optimization face a critical challenge: destructive gradient variance that often leads to training collapse. To address this issue, we introduce , a framework that leverages the model's intrinsic confidence to construct a curriculum independent from external verifiers. By prioritizing high-confidence samples, VI-CuRL effectively manages the bias-variance trade-off, specifically targeting the reduction of . We provide a rigorous theoretical analysis, proving that our estimator guarantees asymptotic unbiasedness. Empirically, VI-CuRL promotes stability and consistently outperforms verifier-dependent/independent baselines across math and general reasoning benchmarks with/without verifiers.
Abstract: Video generation has advanced rapidly, producing photorealistic videos from text or image prompts. Meanwhile, film production and social robotics increasingly demand multi-person videos with rich social interactions, including conversations, gestures, and coordinated actions. However, existing models offer no explicit control over interactions, such as who performs which action, when it occurs, and toward whom it is directed. This often results in wrong person performing unintended actions (actor-action mismatch), disordered social dynamics, and wrong action targets. To address these challenges, we present SocialDirector, a training-free interaction controller that enhances the generation model by modulating cross-attention maps. SocialDirector contains two modules: Social Actor Masking and Directional Reweighting. Social Actor Masking constrains each person's visual tokens to attend only to their own textual descriptions via a spatiotemporal mask, avoiding actor-action mismatch and disordered social dynamics. Directional Reweighting amplifies attention to directional words (e.g., "leftward", "right"), leading each action towards its intended target. To evaluate generated social interactions, we annotate existing datasets with interaction descriptions and build a fully automated evaluation pipeline powered by open-source VLMs. Experiments on different video generation models show that SocialDirector significantly improves interaction fidelity and approaches the upper bound set by real videos.
PaperID: 4721, Poster
Abstract: Diffusion large language models (DLLMs) generate text through iterative denoising, where the choice of which tokens to commit and which to refine drives output quality. Existing certainty-aware strategies rely on instantaneous confidence or entropy to guide this choice, but these local signals can be misleading on hard reasoning problems. This raises a natural allocation question for DLLM test-time scaling: rather than applying one fixed decoder to all inputs, can we allocate not just more compute but a different decoding strategy to questions that remain unresolved? We propose a two-stage adaptive DLLM sampling method. The first stage uses a fast base sampler to solve easy questions; the second applies frequency-aware remasking to the remaining hard questions. Our frequency-aware strategy tracks how often token predictions change across denoising steps and keeps frequently-changing tokens available for further refinement, providing a trajectory-level signal that complements confidence- and entropy-based remasking. Experiments with LLaDA-8B-Instruct and Dream-7B-Instruct show consistent gains over a strong adaptive allocation baseline: up to 3.8% (LLaDA) and 7.4% (Dream) absolute accuracy improvements on MATH-500, and up to 11.18% (LLaDA) and 10.77% (Dream) absolute coverage improvements on HumanEval. On questions that remain unsolved by the base certainty-aware strategy even after 100 samples, frequency-aware remasking recovers solutions for 6.25% to 18.75% of examples.
PaperID: 4722, Poster
Authors: Zhijing Yang, Shuming Hou, Guohui Xiao, Lemei Zhang, Peng Liu
Abstract: Large language models increasingly rely on long Chain-of-Thought (CoT) to solve complex reasoning tasks, but the resulting CoT often contains repeated checks, abandoned branches, and unused planning that increase inference cost without supporting the final answer. Existing efficient reasoning methods address this through length budgets, token-level importance, confidence signals, or a stronger teacher model, treating CoT compression mainly as a text-shortening problem rather than as a question of which steps the answer actually depends on. We propose MinPath, a training framework that treats CoT compression as a graph problem. The raw CoT is represented as a Typed Reasoning Graph, where the Minimal Dependency Path supporting the target conclusion is extracted, and the resulting MinPath is used for fine-tuning. Across four reasoning models on GSM8K, MATH-500, and AIME24, MinPath reduces generated tokens by 33.9% on average with nearly unchanged precision (accuracy even improved in 41.7% of all experiments) and outperforms five compression baselines in every evaluated setting.
PaperID: 4723, Poster
Abstract: Feed-forward 3D Gaussian splatting enables single-pass reconstruction from multi-view images, but widely-applied voxel-aligned pipelines usually produce sharp and anisotropically ill-conditioned Gaussian centers, leading to floaters and unstable geometry in sparse-view settings. To remedy this, we introduce StoSplat, a training-time stochastic preconditioning framework that perturbs predicted Gaussian centers with annealed, ray-aligned anisotropic noise. This perturbation smooths the expected rendering objective in the geometric variable where instability arises, while leaving covariance, opacity, color, and the inference architecture unchanged. StoSplat introduces negligible additional computation during training without any computational overhead during inference. Experiments on widely used benchmarks including RealEstate10K, ScanNet, and ACID demonstrate that StoSplat achieves state-of-the-art performance, while producing more stable geometry with fewer floaters and boundary artifacts, which demonstrates the effectiveness of our method in improving the performance of existing feed-forward 3D reconstruction without architectural changes.
PaperID: 4724, Poster
Abstract: RGB-based category-level 9D pose estimation remains highly challenging because a monocular system must generalize across unseen object instances while jointly inferring 3D translation, 3D rotation, and 3D size. Unlike instance-level pose estimation, the category-level setting requires the network to learn both what is shared across objects within a category and how individual instances vary in structure. This challenge is amplified in the RGB-only setting, where pose cues, category-level commonality, and instance-specific variation are entangled in appearance without direct geometric constraints, making it difficult for existing frameworks to learn a representation that is truly aligned with the category-level estimation objective. To address this problem, we propose SchemaPose, a schema-guided framework for RGB-based category-level pose estimation. SchemaPose equips the estimation pipeline with an explicit parametric schema for modeling category-level regularities together with structured instance variation, and uses the inferred schema variables to guide pose-related prediction in an end-to-end query-based framework. The same formulation also supports template instantiation, synthetic annotation, and pose learning within a unified framework. Experiments on synthetic and real data show that SchemaPose achieves state-of-the-art performance among RGB-based methods for category-level 9D pose estimation.
PaperID: 4725, Poster
Abstract: Equipping VLM agents with world modeling capabilities has shown strong potential for complex reasoning and long-horizon planning, while reducing the dependence of policy learning on costly real-world interactions. Existing methods mainly rely on prospective simulation to predict the consequences of candidate actions. However, this forward-only paradigm focuses on what will happen next and provides limited constraints for verifying whether an action is causally consistent with the observed state transition, which can lead to plausible-looking but physically incoherent behaviors. In this paper, we challenge the view of world modeling as only prospective prediction and introduce Retrospective World Modeling, a new agent learning paradigm that enables agents to reason backward by estimating the retrospective attribution distribution P(\hata_t|s_t, s_t+1) for the action that most likely caused a given transition. Based on this capability, we formulate the Self-Consistency Reward (SCR), an intrinsic signal that measures the probabilistic consistency between the policy action and the retrospective explanation. Integrating SCR into reinforcement learning provides dense transition-level feedback and steers agents toward behaviors that are both task-effective and physically grounded. Extensive experiments across diverse agentic tasks show that our method substantially improves policy robustness and generalization over prospective-only world modeling baselines.
PaperID: 4726, Poster
Abstract: While autoregressive models dominate the LLM landscape, Discrete Diffusion Language Models have emerged as a compelling alternative due to their advantages in parallel decoding and bidirectional contextual modeling. To enhance scalability and efficiency, recent works have integrated the Mixture-of-Experts architecture into the DLM framework. However, MoE-DLMs remain underexplored. In this work, we present an empirical study revealing that MoE-DLMs exhibit not only task-level expert specialization, but also exhibit specialization across different mask-rate diffusion regimes. We further show that during fine-tuning, gradients induced by mismatched mask rates can interfere with the updates of regime-specialized experts, leading to suboptimal adaptation. Motivated by this, we propose MaGM, an adaptive mask-conditioned gradient masking method for MoE-DLM fine-tuning, which dynamically masks expert parameter updates based on the mask rate of training instance. Experiments demonstrate that MaGM consistently outperforms standard full-parameter fine-tuning for diffusion language models, validating the benefits of regime-aware expert adaptation. Our code will be publicly released.
PaperID: 4727, Poster
Abstract: Open-Vocabulary Scene Graph Generation (OV-SGG) requires models to recognize visual relationships beyond the training vocabulary, yet existing methods often rely on dataset-specific object--predicate co-occurrence patterns. Under long-tailed distributions and noisy object localization, such reliance leads to biased, entangled relation representations, limiting both rare-relation recognition and generalization to unseen predicates.To address these challenges, we propose RaRe, a reasoning-aware relational representation learning framework that transforms MLLM-generated chain-of-thought (CoT) rationales into structured relation embeddings. Rather than using CoT only as an intermediate reasoning trace for final prediction, RaRe directly optimizes the rationale-derived representation space to improve inter-relation discriminability. Specifically, RaRe first aligns rationale embeddings with textualized relation descriptions through self-supervised contrastive learning, and then refines rationale generation with a specificity-oriented reinforcement learning objective that encourages semantically distinctive relational evidence.Experiments on Visual Genome demonstrate that RaRe improves category-balanced recognition across both base and novel predicates, with particularly strong gains under the SGDet setting where detection noise is prevalent.
PaperID: 4728, Poster
Abstract: Supervised dimensionality reduction (DR) is widely used in visualization to reveal task-specific structure in high-dimensional data. However, in regression settings with continuous supervision, existing methods often improve signal readability at the cost of severe neighborhood distortion, limiting the reliability of the resulting visualization. To understand this trade-off, we provide a theoretical analysis that characterizes the relationship between geometric faithfulness and signal readability through two finite-sample proxies for the readout's bi-Lipschitz properties: contraction (co-Lipschitz) and smoothness (Lipschitz). We show that contraction can impose a sharp geometric distortion floor in many existing methods, whereas smoothness alone is sufficient to support effective visualization. In particular, directly regularizing smoothness yields first-order regularity gains with only second-order geometry loss, under geometry-first initialization near a local geometry optimum. Guided by this insight, we propose \sys (\acroinitInterpretable \acroinitSupervised \acroinitDR), a nonparametric method that preserves geometric structure while enforcing readout smoothness. Experiments on HPDv3 and Tabula Muris show that \sys improves the geometry--signal Pareto frontier, reducing local signal variation while retaining faithful neighborhoods compared with representative supervised DR baselines. Our code is available at \urlhttps://anonymous.4open.science/r/ISDR-preview.
PaperID: 4729, Poster
Abstract: A central bottleneck in multi-hop Question Answering (QA) is that the granularity at which a question is expressed often differs from the granularity at which corpus evidence is retrievable. Existing methods address this mismatch either by imposing fixed graph structures over the corpus or by iteratively reformulating the query, but these strategies do not explicitly decide when a query unit is already supported by evidence and when it should be refined. We formulate this bottleneck as retrievable granularity discovery and introduce Hi-Q, an evidence-conditioned framework for hierarchical query refinement. At each query node, a resolution operator tests whether retrieved evidence supports the current query unit; resolved nodes terminate, while unresolved nodes are expanded by a dependency-preserving binary operator and checked by a semantic coverage verifier. Hi-Q therefore grows a query tree whose topology is determined by corpus support signals rather than by a fixed decomposition template or a pre-built graph. Across three multi-hop QA benchmarks, Hi-Q achieves 57.9 EM and 69.3 F1 on average, outperforming PropRAG, a graph-based RAG baseline, by 5.6 EM / 3.9 F1 and IRCoT, an iterative retrieval baseline, by 13.7 EM / 15.8 F1. Under full-corpus retrieval, Hi-Q maintains its gains with 53.4 EM and 65.2 F1 on average, improving over IRCoT by 16.3 EM / 19.4 F1 without corpus-wide graph construction.
PaperID: 4730, Poster
Abstract: Open-vocabulary multi-label recognition (OV-MLR) aims to identify all queried semantic concepts present in an image. Existing methods mainly improve global image-text matching or text-side representations, leaving local prediction largely under-explored. In this paper, we revisit local prediction and reveal that it can serve as a strong discriminative branch once two properties are properly handled: spatial consistency of local visual features and compatibility-aware local competition. This perspective is motivated by local semantic sparsity: each image region is related to only a small subset of queried concepts, while compatible concepts of different granularities or types may still coexist locally. We instantiate this perspective with LoCo, a selective Local Competition framework for OV-MLR. LoCo adopts a spatially consistent visual encoder to reduce local feature entanglement, and introduces Selective Softmax to impose competition only among locally incompatible categories while preserving compatible responses. It further combines multi-scale local aggregation with global prediction to handle objects at different scales and scene-level semantics. Experiments on four multi-label benchmarks show that LoCo achieves state-of-the-art performance, with average gains of +11.3% F1-score and +5.2% mAP. Our results suggest that local semantic sparsity offers a useful basis for developing more discriminative open-vocabulary multi-label recognition methods. Code is available in the supplementary material.
Authors: Jongsoo Lee, Jangwon Kim, Soohee Han
Abstract: Reinforcement learning in real-world systems often involves delayed feedback, which breaks the Markov assumption and impedes both learning and control. Canonical augmentation-based approaches cause state-space explosion, which imposes a severe sample-complexity burden. Despite recent progress, state-of-the-art augmentation-based baselines either mainly alleviate the burden on the critic or rely on non-unified treatments for the actor and critic. In this study, we propose delayed homomorphic reinforcement learning (DHRL), a framework grounded in MDP homomorphisms that defines a belief-equivalence relation over the augmented state space to collapse control-redundant augmented states. In principle, this yields exact abstraction under deterministic dynamics and approximate abstraction under stochastic dynamics, enabling both the actor and critic to benefit from a structured abstraction mechanism. In finite domains, exact abstraction preserves optimality and recovers the delay-free sample-complexity order, whereas approximate abstraction admits a value-loss bound on the resulting policy. For continuous domains, we introduce deep delayed homomorphic policy gradient (D^2HPG), a deep actor-critic instantiation of the DHRL framework. Experiments on continuous-control tasks in MuJoCo show that D^2HPG outperforms strong augmentation-based baselines.
PaperID: 4732, Poster
Abstract: Optimization modeling (OM) is fundamental to address Operations Research (OR) problems, yet it requires specialized domain expertise. Large language models (LLMs) are promising for automating OM from natural language descriptions. Through prompting and fine-tuning, LLMs have made significant progress on easy problems, but they still struggle with complex problems involving combinatorial constraint logic and obscure decision variables. To provide a precise formulation for industry OR problems, this paper proposes a multiagent system (MAS) workflow, which decomposes the entire OM process into six structured sub-tasks: set identification, parameter abstraction, variable definition, objective formulation, constraint derivation, and solver-code generation. Existing OM benchmarks only provide a sparse judgment of the final formulation, without process rewards for intermediate agents. Therefore, the main challenge of MAS-based OM is how to fine-tune intermediate agents with OM domain knowledge. This paper proposes a sample efficient counterfactual-based credit assignment for the MAS trajectory. The augmented intermediate credit can be directly used to fine-tune each agent. Moreover, we construct an Optimization Skill Library (Opt-Skill) that provides agent-specific procedural guidance for skill-augmented reasoning, self-checking, and error correction during inference. Experimental results show that, compared with existing methods and frontier LLMs, the proposed MAS workflow achieves state-of-the-art average accuracy while maintaining strong performance on both easy and complex optimization modeling tasks. Altogether, we believe that MAS represents a concrete step forward in performing complex OM in industry scenarios.
PaperID: 4733, Poster
Abstract: Fine-tuning large transformers routes gradients through every token position, paying the cost in activation memory that scales linearly with sequence length and dominates the GPU budget. Existing efficient methods reduce the parameter side---low-rank adapters, prefix tuning, weight-decomposed updates---but leave the activation footprint untouched. We present the Token Hypothesis, a memory-efficient fine-tuning procedure motivated by a Neural Tangent Kernel analysis of how fine-tuning signal is distributed across token positions. We prove that every position is a low-dimensional capacity bottleneck whose width does not grow with the model, that positions with aligned hidden-state geometry are interchangeable in bidirectional models, and that causal positions form a monotone capacity chain in which later positions strictly dominate earlier ones. The analysis pins down how many positions are needed and which ones, both computable from forward passes alone: any subset suffices in bidirectional models, while the last few positions suffice in causal ones. We translate this into a procedure that backpropagates through only a small subset of positions while running the full forward pass over the rest. Empirically, the procedure matches or exceeds full-sequence fine-tuning across six task families---commonsense reasoning, mathematical reasoning, visual question answering, video question answering, GLUE, and image classification---and five backbones spanning bidirectional and causal architectures. It reduces activation memory substantially, translating to up to ∼ 20% lower total GPU memory (∼ 33GB) on long-context tasks, and composes cleanly with parameter-efficient methods such as LoRA and DoRA.
Abstract: In this work we study the Best Policy Identification (BPI) problem in online, tabular Reinforcement Learning. This is an active sequential hypothesis testing problem in which the learner's objective is to identify an optimal policy in a Markov Decision Process (MDP) with high confidence, while minimizing the expected sample complexity to do so. We consider an online setting with deterministic rewards, where the agent must strategically navigate through the MDP in order to effectively explore. Previous works in the literature have provided asymptotically optimal methods for BPI, such as the Navigate and Stop (NaS) algorithm and its variants, however existing analysis remains asymptotic. In this work, we fill that gap by providing the first non-asymptotic sample complexity guarantees for NaS, showing that its sample complexity depends not only on the characteristic time, but also on the connectivity of the underlying MDP, the curvature of the optimal characteristic time, and other instance-dependent quantities. We identify these additional attributes and make explicit their contributions to the overall sample complexity.
Abstract: For p \ge 1, the p-Wasserstein distance measures the minimum cost of transporting probability mass, where moving mass between two points costs the pth power of their distance. For discrete distributions in one dimension, full transport is especially simple: after sorting, mass is matched in order along the line. By contrast, partial and unbalanced transport on the line remains much less understood. Recently, Chapel and Tavenard [ICLR'25] showed that, for p=1, all optimal partial transport plans between distributions supported on n points, with uniform mass at each point, can be computed in O(n\log n) time by exploiting the metric structure of the cost. For p>1, this structure no longer applies, and existing approaches require \Omega(n^2) time. Our main contribution is an FFT-based data structure for balanced-interval transport queries, which bypasses this quadratic bottleneck and yields an O(p\ n\log^2 n)-time algorithm for computing all optimal partial transports on the line for every finite p\ge 1. We also provide an open-source C++ implementation that outperforms the state-of-the-art baseline on a range of synthetic instances. Finally, we establish a conditional lower bound for p=\infty: any subquadratic-time algorithm for computing all optimal partial transport plan costs on the line would violate the (\min,+)-Convolution Hypothesis. This separates the problem from full optimal transport, which is solvable in O(n\log n) time.
PaperID: 4736, Poster
Abstract: Transformer-based video generative models have become increasingly realistic, but long-horizon consistency remains challenging to achieve because even a few dozen frames create impractically long sequence lengths. We show that this issue can be mitigated by generating video using coarse-to-fine rollout within a multi-scale token space. Our approach is simple: first, we pre-train an adaptive autoencoder that compresses each frame into a hierarchy of tokens, with levels ranging from the typical latent resolution to only a handful of tokens per frame. This yields a hierarchy in which the coarsest levels capture the most consequential information—such as scene layout and semantics—while finer levels add high-frequency appearance and texture. Then, we train a video diffusion model to generate these tokens using coarse-to-fine rollout. By carefully controlling the level of detail at which frames are generated and used as context during each rollout step, we are able to preserve long-range consistency in geometry and object permanence while expending less compute on long-term consistency of less perceptually relevant details. We validate this approach using a custom long-horizon Minecraft dataset, where it produces substantially more consistent rollouts compared to strong baselines. We hope this perspective opens new opportunities for long-term consistent video generation.
Abstract: Many decision-facing stochastic systems are observed through aggregate distributions rather than scalar trajectories: queue occupancies, mobility shares, public-health mixtures, generation-source shares, ecological compositions, and air-quality severity profiles all live on the probability simplex and evolve over time. We study causal (time-respecting online) forecasting for these distribution-valued time series and argue that the transition operator itself should be structured around the simplex. We introduce CAST (Causal Anchored Simplex Transport), a successor-local operator that (i) retrieves empirical successors from causal context, (ii) stabilizes them with a persistence anchor, and (iii) applies a bounded local stochastic transport on ordered supports; every stage preserves the simplex by construction. We identify a structural failure mode, latent transition-kernel aliasing, where similar observed distributions evolve differently under different contextual regimes, and prove that any forecaster depending only on an aliased summary incurs an irreducible weighted Jensen--Shannon excess-risk lower bound, while the CAST hypothesis class contains the regime-aware Bayes successor; for ordered supports an additional Pinsker separation holds whenever the transported successor lies outside the no-transport anchor hull. On a suite of eleven public and simulated benchmarks spanning ecology, energy, diet, mortality, employment, air quality, severe weather, mobility, and G/G/1, G_t/G/1 queue occupancy, CAST achieves the best average rank on both one-step KL (1.27) and autoregressive rollout JSD (1.91), winning 8/11 sections on each metric against a broad statistical, compositional, recurrent, convolutional, Transformer, and modern time-series baseline set, and top-2 on all 11 sections for offline KL. Component ablations and a controlled synthetic aliasing experiment corroborate the theory.
PaperID: 4738, Poster
Abstract: Diffusion and flow-based models learn powerful data priors by training a denoiser to reverse Gaussian corruption. To use this prior to solve a linear inverse problem, one needs to sample from the posterior, but the score that the prior provides is the unconditional score, not the posterior score. Existing methods either steer a fixed pretrained denoiser with approximate measurement-matching corrections, or train a conditional restoration model that abandons the denoising structure of the prior. We derive the exact posterior score in closed form for linear Gaussian inverse problems under general Gaussian interpolants, and show that posterior sampling reduces to a denoising problem at an operator-dependent shifted pivot under an anisotropic noise covariance. We turn this identity into , a denoising training objective that preserves the input/output structure of standard pretraining and can therefore be trained from scratch or fine-tuned from a pretrained denoiser. At inference, EPS uses the same sampler as the underlying backbone, with no likelihood gradients or projections. We evaluate EPS on five linear inverse problems across FFHQ and ImageNet, where it outperforms training-free and training-based baselines on fidelity, perceptual, and distributional metrics, while using roughly an order of magnitude fewer denoiser evaluations than gradient-based posterior samplers.
PaperID: 4739, Poster
Abstract: Large language model-based multi-agent systems have recently shown strong potential for complex, long-horizon tasks. However, existing methods mainly rely on coarse prompt-level differentiation without parameter adaptation for diverse subtasks, resulting in insufficient inter-agent heterogeneity and limited specialized capability that bottleneck performance on tasks with complex requirements. To address this, we introduce a Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts (MoRSE) that distinguishes agents with (role, subtask)-conditional specialization at both the task structure and parameter levels. To make agents' responsibility explicit at the task structure level, we formulate a task-oriented multi-agent system that decomposes each task into a dependency-aware Directed Acyclic Graph of subtasks and assigns each agent a specific (role, subtask), introducing task-level specialization across collaborating agents. Additionally, to address the diverse role and subtask demands that a single shared base model cannot satisfy, we propose a dynamic Mixture of (role, subtask) LoRA Experts module with a prototype-based semantic router for subtasks, augmenting agents with parameter-level specialization on a shared LLM substrate cost-effectively. Then, to co-optimize experts and router stably for open-ended tasks, we further propose a hierarchical group-relative policy optimization with two-layer credit assignment that disentangles expert quality from routing quality. Experiments on code-generation benchmarks across three backbones demonstrate the effectiveness of our approach, with notable improvements in both whole-task and step-wise performance and great generalization potential in out-of-distribution scenarios.
Authors:
Hanqing Zhu, Zhenyu Zhang, Hanxian Huang, DiJia Su, Zechun Liu, Jiawei Zhao, Igor Fedorov, Hamed Pirsiavash, Jinwon Lee, David Z. Pan, Zhangyang "Atlas" Wang, Yuandong Tian, Kai Sheng TaiAbstract: RLVR reliably improves reasoning, yet it appears to update only a tiny fraction of weights. We resolve this paradox by revealing a persistent, model-conditioned optimization bias: independent runs concentrate updates in similar parameter regions, invariant to dataset or RL recipe, while finite-precision storage (e.g., bf16) obscures widespread micro-updates as a visual sparsity artifact. To characterize this unique bias, we show that RLVR preferentially learns along off-principal directions via a cascaded Three-Gate mechanism: updates are first bounded by a empirical KL constraint (Gate I), then steered by the model’s anisotropic geometry(Gate II) toward spectrum-preserving off-principal subspaces, and finally filtered by finite precision (Gate III), effectively masking minor updates. Empirically, we validate the off-principal dynamics: RLVR exhibits minimal spectral drift, reduced principal-subspace rotation, and strong off-principal alignment that set RLVR strikingly apart from SFT. Together, our results provide the first parameter-space account of RLVR and uncover consistent regularities in weight evolution, advancing a more white-box understanding of RLVR. Moreover, we show that RLVR follows an optimization regime distinct from SFT, showing directly that transferring SFT-era PEFT can be flawed and motivating geometry-aware, RLVR-native methods.
Abstract: While LLMs excel at single-turn generation, they struggle with long-horizon, multi-turn interactions. Offline reinforcement learning (RL) offers a scalable approach, yet its performance hinges on the availability and quality of multi-turn trajectory data. A common remedy is to augment training with synthetic trajectories generated by LLMs or simulators, but synthetic data is highly heterogeneous in quality, and naively treating all trajectories as equally informative can degrade performance. We propose BOOST, a bilevel optimization framework where the inner level trains the LLM on reweighted data and the outer level trains a lightweight reweighting head on held-out real validation tasks, assigning continuous trajectory-level weights without requiring an external judge. To ground this approach, we derive a PAC-Bayesian bound revealing a three-way trade-off: synthetic data increases diversity but risks task-shift, while concentrating weight on high-quality trajectories improves empirical performance at the cost of effective sample size. Empirically, our method consistently outperforms multiple baselines. Analysis reveals it upweights synthetic trajectories that align with the real data distribution and exhibit higher qualitative merit.
PaperID: 4742, Poster
Authors: Makoto Nakakita, Teruo Nakatsuma
Abstract: Bayesian causal inference quantifies uncertainty in treatment effects under a specified causal model. This uncertainty is often interpreted as evidence that the resulting causal conclusion is robust. We argue these are distinct targets: a posterior over a treatment effect may be tightly concentrated above zero while the corresponding causal claim is fragile to small violations of ignorability, overlap, or target-population assumptions. We introduce Bayesian causal stress testing (BCS), a framework that assigns a posterior fragility score F_\alpha with a threshold \alpha\in(0,1) to a causal conclusion by measuring the smallest calibrated stress under which the posterior probability of the claim falls below \alpha. We give a stress-map formulation, decision and composition results for claim-level stress, a calibration theorem showing that F_0.5 asymptotically recovers the partial-correlation form of the Cinelli-Hazlett robustness value in the linear-Gaussian flat-prior limit, and a proposition formalizing why posterior precision and posterior fragility can decouple. We evaluate BCS on IHDP, TWINS, and ACIC-cov+Hill, a semi-synthetic benchmark built from the ACIC 2016 covariate matrix, under BART, BCF-style, Bayesian linear, and Bayesian-bootstrap AIPW posteriors. On the linear Bayesian posterior, F_0.5^2 matches the Cinelli-Hazlett RV_0.5^\mathrmCH to RMS 0.001 on both IHDP and ACIC-cov+Hill across all 20 configurations; the BART and BCF-style posteriors show larger deviations in small-effect configurations. Across estimators, posterior 95% interval widths vary by up to 50% but F_0.5 remains within 0.05, illustrating that posterior precision is not a robustness certificate. A controlled-DGP calibration experiment with a hidden confounder of known partial correlation \gamma_\mathrmtrue confirms that F_0.5 tracks the required flip-stress to mean absolute error 0.03.
Abstract: Reinforcement learning with verifiable rewards (RLVR) is central to training modern reasoning models, but the undisclosed training data raises concerns about benchmark contamination. Unlike pretraining methods, which optimize models using token-level probabilities, RLVR fine-tunes models based on reward feedback from self-generated reasoning trajectories, making conventional likelihood-based detection methods less effective. We show that RLVR induces a distinctive behavioral signature: prompts encountered during RLVR training result in more rigid and similar generations, while unseen prompts retain greater diversity. We introduce Min-kNN Distance, a simple black-box detector that quantifies this collapse by sampling multiple completions for a given prompt and computing the average of the k smallest nearest-neighbor edit distances. Min-kNN Distance requires no access to the reference model or token probabilities. Experiments across multiple RLVR-trained reasoning models show that Min-kNN Distance reliably distinguishes RL-seen examples from unseen ones and outperforms existing membership inference and RL contamination detection baselines.
Authors:
Ziye Chen, Hongbin Lin, Chenyu Zhang, Xiangda Yan, Yongjie Yang, Yao SHUAbstract: Zeroth-order (ZO) optimization enables large-language-model fine-tuning without storing backpropagation activations, while LoRA supplies compact trainable adapters. Combining them creates a rank paradox: increasing LoRA rank improves adapter capacity, but standard two-point ZO either perturbs a rank-dependent number of coordinates or, under atomwise updates, can make the finite-difference signal unobservable. This paper shows that the bottleneck is a measurement-topology problem rather than a need for an external subspace. LoRA already decomposes into matched rank-1 atoms, each a complete factor-coordinate block of dimension d_\textout+d_\textin. Querying one atom per step keeps the stored adapter rank r while removing r from the single-query perturbation dimension. The naive atomwise query is still miscalibrated: if it inherits canonical LoRA scaling \alpha/r, the active finite-difference signal shrinks as 1/r and the active finite-difference signal-to-noise ratio (FD-SNR) as 1/r^2, producing directional collapse under a fixed residual evaluation-noise floor. AR1-ZO pairs alternating rank-1 atom queries with topology-aware scaling \gamma=\alpha r, restoring rank-invariant active signal without auxiliary bases, activation hooks, curvature estimates, or extra forward queries. Theory proves atom minimality, rank-independent active query dimension, directional collapse and restoration, and the remaining rank dependence as an amortized coverage cost. Experiments on OPT and Qwen3 models validate the signal mechanism and show that AR1-ZO makes high-rank LoRA effective among matched-budget ZO methods under the standard two-forward-pass query budget.
PaperID: 4745, Poster
Abstract: We propose Structural Self-Teaching (SST), a novel framework that enhances compositional image generation in unified multimodal models (UMMs) by leveraging their internal understanding path as a source of structural supervision. The motivation of this work is the observation that the generation path of UMMs often struggles with dense compositional prompts, leading to missing objects, attribute binding failures, and spatial errors. To bridge this gap, our key idea is to convert the phrase-level grounding inherent in the understanding path into training-time structural signals. Specifically, we extract phrase-conditioned support masks from internal attention maps derived from the understanding path and employ them through two complementary mechanisms: 1) phrase-local contrastive alignment to synchronize generated features with their corresponding phrase-level grounding, and 2) a lightweight structural scaffold for spatial modulation of intermediate generation states. By employing attenuated scaffold dropout during training and removing the scaffold branch at inference, our approach improves compositional fidelity across object, attribute, and spatial constraints. Crucially, the framework requires no external structural annotations for training or additional inputs at test-time, preserving the original generation interface. Extensive experiments across multiple UMMs demonstrate that our method consistently improves compositional fidelity while maintaining the efficiency of a lightweight tuning setup.
PaperID: 4746, Poster
Authors: Avi Caciularu
Abstract: Decoding signals over unknown channels with minimal pilot overhead is a critical challenge in next-generation communications. Existing deep learning approaches typically rely on generic encoders that struggle to model long-range temporal dependencies or efficiently capture the channel's physical properties from scarce data. We argue that standard architectures suffer from fundamental (CAT), a novel architecture that explicitly injects geometric inductive biases into the equalization process. CAT is composed of a stack of custom blocks, which utilize an ``early interaction'' paradigm to co-process received signals and ideal constellation symbols. Each block features a split Feed-Forward Network that applies a Finite Impulse Response (FIR)-inspired filter for robust deconvolution and a parallel MLP for geometric refinement. We theoretically prove that this design creates a structural isomorphism to the optimal MIMO Wiener Receiver and eliminates the ambiguity floor inherent to blind estimators. In the challenging semi-supervised setting, CAT achieves state-of-the-art performance, significantly reducing pilot overhead compared to VAE and standard Transformer baselines.
PaperID: 4747, Poster
Abstract: Test-time sampling can improve Vision-Language-Action (VLA) policies, but only if the model can efficiently select a good action from multiple candidates. Prior work often relies on separately trained external verifiers, adding extra models and inference overhead. We propose Token Probability Bucketing (TokenPB), a self-verification approach that discretizes continuous action chunks into bucket tokens and co-trains a single VLA checkpoint with a joint objective: a flow-matching loss for action generation and an auxiliary bucket loss \mathcalL_\textbucket that teaches the VLM backbone to model observation-conditioned distributions over bucket-token sequences. At inference, TokenPB samples M candidates, scores each via teacher-forced evaluation under the learned distribution, and selects the best---enabling efficient batched scoring with observation-prefix KV-cache reuse. We provide theory connecting the training objective to a monotonic ranking property and bounds on expected selection regret that separate candidate-set effects from calibration. Across multiple backbones and benchmarks, TokenPB yields consistent success-rate improvements over single-sample inference from the same checkpoint (+2.7 in simulation, +17.1 in real-world, +16.4 under out-of-distribution shifts), and further improves over external-verifier baselines under the same sampling budget.
Authors: Marios Papamichalis, Regina Ruane
Abstract: Contact-rich robot dynamics are hybrid: a single observation can match several latent states and contact regimes (free, impact, stick--slip). A standard amortized filter that places no probability on a feasible contact transition will permanently lose the branch the robot actually follows. We introduce VHYDRO, a variational hybrid dynamics learner that prevents this branch loss. At each step, VHYDRO mixes the learned proposal with a feasible transition law before sampling and importance weighting. This keeps every feasible contact alternative covered. VHYDRO jointly infers a continuous latent state and a discrete contact mode, and fits a sparse port-Hamiltonian law to each recovered regime. On top of this, three guarantees connect: support coverage stabilizes filtering, the stabilized filter concentrates the discrete contact posterior on coherent regimes, and mode-pure segments admit sparse port-Hamiltonian recovery. The recovery error separates cleanly into filtering, derivative, mode-impurity, and physics-residual parts. Three empirical findings track the same mechanism. Under heavy occlusion the support-safe filter stays usable while a non-defensive proposal collapses. On ManiSkill demonstrations and on four canonical Sawyer/BridgeData task families the discrete state forms temporally coherent contact-regime segments that post-hoc segmenters and mode-free baselines miss. On hybrid systems with known equations the mode-conditioned sparse fit recovers the active physical terms; purely predictive baselines do not.
PaperID: 4749, Poster
Abstract: Vision-Language Models (VLMs) excel at open-world 2D perception but struggle with precise metric spatial reasoning, a key requirement for applications like autonomous driving, navigation, and embodied intelligence. We identify this as a structural limitation: images are 2D projections of the 3D world, inherently suffering from projective ambiguity, while VLM components favor semantic understanding and rely on 2D positional bias, leading to Spatial Collapse. To address this, we propose Despa, a Depth-grounded spatial VLMs that injects depth-based geometric information via Geometric Positional Embedding (GPE) and Depth-aware Rotational Positional Embedding (DoPE). To expose masked limitations in current benchmarks, we construct SpaDataset (1.7M QA pairs) and SpaBench (2,051 QAs), to our knowledge the largest fine-grained training corpus and benchmark suite for metric spatial reasoning. They cover three difficulty levels and ten sub-tasks across diverse settings. The fine-tuned Despa-4B model consistently outperforms general-purpose, closed-source, and specialized VLMs on SpaBench by 27.9% over Qwen3-VL-235B-A22B and 31.1% over GPT-5.2, generalizes to MSMU (66.3%), RefSpatial-U (42.4%), and the multi-frame VSI-Bench (66.0%), and consistently lifts LLaVA-1.5, Qwen2.5-VL, and Qwen3-VL backbones, achieving the performance with only a 4B model. The code will be publicly released.
PaperID: 4750, Poster
Authors: Sihan Chen, haochen sun, Peng Jiang, Anthony Cohn
Abstract: Reconstructing velocity distributions using wavefield responses provides a powerful tool for probing the internal structures of objects and has been widely applied in nondestructive testing, geological exploration, and medical imaging. While full waveform inversion (FWI) provides high-resolution results, its severe ill-posedness and high computational cost hinder efficient 3D reconstruction. In this paper, we propose a novel framework for efficient 3D internal velocity field reconstruction. First, we parameterize the velocity field with a coordinate-based neural network as an implicit continuous function, enabling compact representation from sparse observations. Second, we adopt ray-based traveltime tomography with multi-level parallel forward modeling to accelerate computation. Third, we introduce a differentiable Fresnel volumizing method that extends 1D ray paths into 3D volumetric regions (i.e., Fresnel volumes), enabling spatially continuous and multi-scale gradient diffusion, thus alleviating gradient sparsity and improving reconstruction accuracy. To further improve efficiency, we propose a 2D slice-based construction of 3D Fresnel volumes using a lightweight neural network. Extensive experiments on synthetic datasets, including ablation studies, demonstrate up to 10 times speedup over FWI while maintaining comparable accuracy, validating the effectiveness of the proposed method for efficient 3D inverse problems.
PaperID: 4751, Poster
Abstract: Decoding visual content from non-invasive brain signals remains challenging because neural responses evolve over time whereas images are typically static. Most existing methods align an entire neural response window to a single final visual embedding, this overlook a fundamental representational mismatch: brain signals reflect both early perceptual and later integrative representations, whereas the final visual embedding is biased toward high-level semantics. Therefore, we propose BRIDGE, a brain-vision representation integration framework that aligns two modalities through Visual Depth Encoding and Brain Granularity Encoding. On the visual side, BRIDGE extracts and fuses CLIP representations from multiple depths, producing an alignment target that preserves low-level and high-level information. On the brain side, instead of treating the whole temporal window as homogeneous and obscuring temporally heterogeneity, BRIDGE explicitly partitions stimulus-evoked EEG/MEG responses into a small number of temporally ordered stages and adaptively pools them. The resulting brain and visual embeddings are trained in a shared latent space by contrastive learning and can further support brain-to-image generation through a pretrained diffusion prior. Experiments on THINGS-EEG and THINGS-MEG demonstrate that BRIDGE achieves strong retrieval and generation performance. Ablation studies further confirm the complementary benefits of depth-wise visual aggregation and temporal brain factorization.
Authors: Patrick Kwon, Chen Chen
Abstract: Human Mesh Recovery (HMR) is fundamentally ambiguous: under occlusion or weak depth cues, multiple 3D bodies can explain the same image evidence. This ambiguity is not uniform across the body, as torso pose and root structure are often relatively well constrained, whereas distal articulations such as the arms and legs are more uncertain. Building on this observation, we propose FactorizedHMR, a two-stage framework that treats these two regimes differently. A deterministic regression module first recovers a stable torso-root anchor, and a probabilistic flow-matching module then completes the remaining non-torso articulation. To make this completion reliable, we combine a composite target representation with geometry-aware supervision and feature-aware classifier-free guidance, so that well-observed body parts remain accurate while ambiguous articulations can be completed without collapsing to deterministic averages. We also introduce a camera-aware synthetic data pipeline that provides the paired image-camera-motion supervision under diverse viewpoints. Across camera-space and world-space benchmarks, FactorizedHMR improves articulation recovery and reduces drift relative to strong baselines, especially in ambiguous scenes.
PaperID: 4753, Poster
Abstract: The joint use of graph structure and node-specific information is central to transductive node classification, yet a fundamental theoretical question remains unanswered: when does graph structure actually help? We provide the first non-asymptotic, distribution-free answer to this question. We build on nuclear-norm graph clustering and classical convex clustering, generalizing the latter to incorporate node features, partial labels, and general convex loss functions, making it suitable for transductive node classification. We analyze each problem separately to characterize when graph structure alone, or node information alone, is sufficient for perfect recovery. We then formulate a unified convex optimization problem through a shared atomic norm representation, and prove that there exist conditions under which neither source alone achieves perfect recovery but using both jointly does, revealing a bidirectional synergy: node information improves clustering, while graph structure improves classification.
PaperID: 4754, Poster
Abstract: We introduce P-IMLA, a preconditioned implicit-midpoint Langevin sampler for high-dimensional log-concave targets with non-smooth regularization. The obstruction is that the large steps enabled by matrix preconditioning are precisely where preconditioned Euler--Maruyama inflates stationary variance. Implicit midpoint removes this surplus via a Cayley-transform identity, making P-IMLA Gaussian-exact at any step size. Beyond Gaussians, a Möbius-monotonicity argument gives an M-Wasserstein rate governed by \kappa_\rm eff rather than the \kappa_f of unpreconditioned IMLA, while backward-error analysis shows drift-only rather than Euler-type diffusion bias. An M-norm Moreau envelope yields a proximal non-smooth variant with a three-term error budget. Controlled Gaussian and TV-regularized imaging experiments confirm the predicted separation: the unpreconditioned IMLA degenerates as \kappa_f grows, whereas P-IMLA remains \kappa_\rm eff-governed and near variance-exact.
Abstract: Selecting the best large language model (LLM) for a fixed benchmark is often expensive, since exhaustive evaluation requires running every model on every example. Multi-armed bandit (MAB) algorithms can reduce the number of LLM calls by sequentially selecting the next model-example evaluations, thereby avoiding unnecessary computations on clearly underperforming models. Further computational savings can be achieved by predicting the evaluation scores of the various models on the set of examples. In practice, these predictions can be obtained using low-rank (LR) matrix factorization that exploits correlations in the partially observed model–example score matrix. However, such predicted evaluations are not the ground truth: they can be biased and may therefore lead to incorrect identification of the best model. In this work, we propose a principled framework that combines MAB with cheap, predicted evaluation scores without compromising statistical validity. Concretely, leveraging prediction-powered inference (PPI), we derive unbiased estimators for the performance of each model that utilize predicted scores from LR factorization to reduce variance. Importantly, this enables the construction of finite-sample valid confidence intervals in our non-i.i.d. setting, where models are selected adaptively and examples are sampled without replacement. Empirical results on real-world benchmarks demonstrate that our approach reduces the number of required evaluations, leading to meaningful savings in compute and cost, while accurately identifying the best-performing model.
PaperID: 4756, Poster
Authors:
Aditya Gupta, Jean S Dandurand, Kai Qiu, Rohan Choudhury, László JeniAbstract: Vision Transformers (ViTs) are bottlenecked by the length of their input token sequence. Standard patchification fixes this length with a single patch size applied across the image, regardless of regional content. Prior work on adaptive patch sizing reduces token count by assigning coarser patches to visually simple regions, but relies on noisy pixel-level heuristics such as local entropy to make patchification decisions. We argue that this compute allocation should be grounded in the model’s own semantic representations, not in raw pixel statistics. We introduce Split Policy Learned Image Tokenization which learns to allocate token resolution based on what the model finds meaningful, producing shorter token sequences that are better matched to task-relevant content. SPLIT matches the accuracy of full-resolution ViT classification baselines while using 30% fewer tokens. The same gains hold dense prediction: SPLIT achieves competitive accuracy on COCO and ADE20K, demonstrating that semantically grounded patchification generalises well beyond classification.
PaperID: 4757, Poster
Abstract: Auction-based Federated Learning (AFL) provides a principled framework for incentivizing data owners and allocating decentralized training resources to data consumers. However, practical AFL systems face a coupled valuation and allocation challenge: data quality is uncertain before training, post-training feedback can be strategically distorted, and heterogeneous deadlines render late model updates ineffective for learning. We propose MainFL, an incentive-compatible intertemporal data valuation mechanism for robust auction-based federated learning. MainFL integrates three components: (i) a bounded mutual-information peer prediction rule that elicits truthful pre-training quality reports without requiring ground-truth labels; (ii) an intertemporal reliability calibration scheme that updates future valuations using post-session validation signals; and (iii) a deadline-induced layered clinching mechanism that allocates quality-weighted data while enforcing budget and feasibility constraints.We theoretically establish Bayes–Nash truthfulness for data owners, allocation feasibility and budget safety, and truthful residual-demand equilibrium for data consumers under standard clinching-auction regularity conditions. Empirical results on widely used federated learning benchmarks show that MainFL improves social welfare by up to 12.1% and model accuracy by up to 7.4% compared to state-of-the-art AFL baselines, while substantially reducing straggler-induced task timeouts.
PaperID: 4758, Poster
Abstract: As the ecosystem of Large Language Models (LLMs) rapidly expands, LLM routing has become essential for dynamically balancing performance and cost. However, existing methods fail to simultaneously satisfy the three core criteria of an ideal router, i.e., high routing precision, minimal intrinsic operational overhead, and zero-shot onboarding of new models. These methods rely heavily on surface-level semantics and discrete ID mappings, conflating mere semantic similarity with underlying capability requirements and thereby precluding zero-shot, training-free generalization. To address these limitations, we propose DART, a novel zero-shot routing framework grounded in dual-side alignment to optimize LLM performance-cost tradeoffs. Instead of relying on direct semantic matching, we decouple the routing process into a demand-supply alignment problem within a shared low-rank latent capability space. Specifically, on the supply side, DART constructs latent capability embeddings for candidate models, using an initialization mechanism to anchor newly introduced models in this latent space via metadata and family-graph prior, without retraining. On the demand side, a dedicated encoder explicitly distills the query's directional capability requirements and inherent difficulty into a latent demand embedding. This decoupled representation allows routing decisions to be computed via a lightweight dot-product operation between the demand and capability embeddings, modulated by query-adaptive cost sensitivity, minimizing overhead. Extensive experiments across comprehensive benchmarks demonstrate that DART achieves state-of-the-art utility under diverse cost constraints, satisfying the three core criteria with exceptional robustness in cold-start scenarios.
PaperID: 4759, Poster
Abstract: Compared with deterministic methods that output only point estimates, probabilistic forecasting can characterize both future trends and their uncertainty simultaneously, making it more suitable for decision-making in complex real-world scenarios. In recent years, diffusion models, owing to their powerful generative modeling capability, have been introduced into the field of time series forecasting. However, we find that most existing diffusion-based forecasting methods directly construct the diffusion process in the original future space, which is easily affected by heteroscedastic scale imbalance, causing highly volatile dimensions to dominate the denoising learning and leading to unstable corrections near observation boundaries. Motivated by these issues, we propose VolaRM, a volatility-whitened probabilistic residual modeling framework for long-term time series forecasting. By constructing a conditional probabilistic base distribution, it represents future targets as whitened residuals relative to this base distribution and performs conditional modeling in the normalized residual space, thereby alleviating scale imbalance across variables and forecasting horizons. Meanwhile, a specific gating mechanism is introduced to enhance boundary continuity and the stability of long-term forecasting. Extensive experiments on eight real-world datasets from different application domains demonstrate the effectiveness of VolaRM.
PaperID: 4760, Poster
Authors: Xiaoqian Ruan, Pei Yu, Dian Jia, Hyeonjeong Park, Peixi Xiong, Wei Tang
Abstract: Single-view 3D reconstruction has advanced rapidly in recent years, but existing research primarily focuses on recovering whole-object geometry while ignoring their semantic parts, which are crucial for fine-grained 3D perception and downstream applications. This paper introduces the task of open-vocabulary partonomic reconstruction. Given a single image and a set of part names, we aim to reconstruct both the object’s overall shape and its constituent semantic parts, even when the object and part categories are unseen during training. To address this task, we propose a vision-language-shape (VLS) model that unifies vision, language, and shape representations within a shared neural field and grounds continuous 3D coordinates to 2D pixel contexts via a deformable implicit function. It can effectively reconstruct any-topology shapes and open-vocabulary semantic parts from a single image. To train VLS with limited 3D part-labeled data, we propose an omni-supervised learning framework that leverages heterogeneous datasets with different levels of annotation and open-world knowledge from existing vision-language models. Extensive experiments on ShapeNetPart, PartNet, and Objaverse demonstrate the effectiveness and strong generalization ability of VLS. We will release the code and data publicly.
PaperID: 4761, Poster
Authors:
Kai Zhao, Dai Shi, YIXIN CHEN, Ye Wang, Oula Ghannoum, Yi GuoAbstract: Open-vocabulary 3D scene understanding provides an important interface for querying and interacting with reconstructed scenes through natural language. Although recent 3DGS methods have enabled text-driven object selection and open-vocabulary segmentation, they still struggle with compositional queries involving attributes, spatial relations, and part-whole relations. Most existing approaches learn per-Gaussian or per-cluster language features, construct query-conditioned referring fields, or perform spatial reasoning only at inference time, but they do not organize the scene itself into a reusable relation-aware representation. As a result, object relations are not explicitly stored as typed and confidence-aware 3D structures, limiting compositional reasoning and multi-hop querying. To address this issue, we propose RelGS, a relation-aware 3D Gaussian Splatting framework for open-vocabulary scene understanding. RelGS learns semantic and instance embeddings for Gaussians under multi-view supervision, groups them into cross-view consistent 3D semantic nodes, and enriches each node with structured attributes verified by reverse CLIP consistency. It further constructs a confidence-weighted relation graph, where 3D spatial cues propose candidate relations and an LLM verifies their semantic plausibility. At query time, RelGS combines attribute matching, CLIP retrieval, and graph traversal to localize the target object. Extensive experiments demonstrate strong performance on word-level selection, sentence-level referring segmentation, multi-hop relational queries, and point-level semantic segmentation.
PaperID: 4762, Poster
Abstract: Open-source DNA foundation models such as Evo2 generate sequences with \geq 90% local nucleotide identity to known human-pathogenic viruses under brute-force sampling alone, with no domain expertise, prompt engineering, or access to model internals. The bottleneck for biological misuse is therefore the model's prior, not the attacker's sophistication, and a deployable defense must operate where the attack does: at the codon layer. First, we characterize the attack. On JailbreakDNABench's 23 pathogenic viruses, brute-force Best-of-N sampling against Evo2 succeeds for a mean of 9.5/23 viruses at 7B across two independent trials (41%, range 39.1%\!-\!43.5%, matching GeneBreaker's three-component engineered jailbreak pipeline at 36%) and 12/23 at 1B (52%, single curated trial, ~\!6× GeneBreaker's reported 9%), without any guided search or learned attack model. Probing Evo2's outputs, we find generations are not memorized: 0/11 appear verbatim in NCBI nt, 0/11 recall a single strain cleanly, and codon-position mismatches are biologically structured rather than random (\chi^2(2) = 50.37, p < 10^-4; wobble-position synonymous rate 63.6%, near the per-virus codon-usage baseline and well above the 25% uniform-mutation null). Verbatim or near-verbatim training-data filtering would therefore not have prevented these generations. Second, we present a deployment-time defense. Because the attack exploits codon-level wobble structure, the defense must too. WobbleGuard is a zero-training provenance scheme whose contribution at the detection layer folds the genetic code's known synonym structure into the scoring rule. At an in-sample-calibrated threshold (1.00% FPR by construction on 22,982 natural NCBI pathogen fragments, \alpha = 0.01; cross-key range 0.30%\!-\!1.24% over 8 independently sampled keys), the composite detector \max(z_v, z_w) retains 95.5% TPR under full synonymous codon substitution, the regime in which a nucleotide-level KGW port collapses to ~0% TPR, and detects 100% of unperturbed watermarked Evo2 generations at both 1B and 7B scales. The attack works because brute force suffices; the defense works because it speaks codons.
PaperID: 4763, Poster
Abstract: Language agents increasingly act as web-enabled systems that search, browse, and synthesize information from diverse sources. However, these sources can include unreliable or adversarial content, and the robustness of agents to adversarial ranking remains poorly understood. Existing benchmarks evaluate functional navigation or static factuality but cannot causally isolate this vulnerability and often confound retrieval-time reasoning with memorized knowledge. We introduce Synthetic Web Benchmark, a controlled environment of procedurally generated web ecosystems designed to evaluate retrieval-time reasoning and source criticism. The benchmark comprises thousands of hyperlinked articles with ground-truth labels, process-level interaction traces, and contamination filtering to ensure that answers cannot be recovered from pretraining alone. By injecting a single high-plausibility misinformation article at a specified rank, we measure the causal effect of adversarial exposure under minimal intervention. Across six frontier models and thousands of evaluation instances, we observe catastrophic failures: accuracy collapses despite unrestricted access to truthful evidence, accompanied by limited search escalation, weak cross-source synthesis, and severe miscalibration. These results show that current agents struggle to arbitrate conflicting sources even when sufficient evidence is available, revealing fundamental limitations in retrieval-based reasoning. The benchmark provides a reproducible testbed for studying epistemic robustness and developing more reliable web agents.
Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit a task-completion bias, executing hallucinated, undesired actions to force a resolution rather than expressing uncertainty. Existing detection methods fail to provide actionable, real-time correction as they either do not localize the hallucinations or incur prohibitive inference latency. We introduce the Latent Critic, a lightweight low-rank adapter (LoRA) that operates concurrently with a frozen base LLM's generation to actively restructure the transformer's residual stream---amplifying epistemic uncertainty signals and translating them into localized, natural language feedback within a single sequence. By refining the base model's native uncertainty signals, this manipulation of the latent space enables highly reliable, granular detection without the overhead of secondary inference loops. Mechanistic analysis via activation patching and layer-wise probing shows that this rank-invariant behavior restructures pre-existing uncertainty geometry into a linearly separable representation. Using tool-calling as an instantiation of granular hallucinations, we validate the detection and downstream improvements enabled by the Latent Critic architecture across Qwen and Llama-based models. Demonstrating superior real-time efficacy, our approach significantly outperforms equivalent-scale external detectors and internal probes in isolating hallucinations (0.870 vs. 0.695 F1), achieving >80% accuracy in localization (e.g., ungrounded: date). When deployed in a closed-loop ReAct environment, the Critic acts as a zero-latency guardrail, intercepting hallucinations before execution to prevent undesired actions while simultaneously enabling efficient agent self-correction.
PaperID: 4765, Poster
Abstract: Kalman filtering provides a principled inference framework for linear Gaussian hidden Markov models, but it leaves open a complementary question of representation: what is the minimum amount of information about the observation history that must be retained to support accurate filtering? We show that this question leads naturally to an indirect rate-distortion problem, in which an encoder observes the history of noisy measurements while distortion is evaluated with respect to the latent state. Directly solving this problem is challenging, since the optimization ranges over arbitrary stochastic kernels from a time-varying observation history to an output representation, resulting in an infinite-dimensional formulation. We overcome this difficulty by proving that any feasible encoder can be replaced by a stochastic encoder acting only on the Kalman posterior mean, achieving the same mean-square error with no larger mutual information. The reduced problem admits a closed-form solution in the form of a linear Gaussian blurring mechanism determined by the spectral decomposition of the Riccati solution. This yields an explicit expression for the rate distortion function R_t(\epsilon), showing that any representation with distortion at most \epsilon must retain at least R_t(\epsilon) nats of information. We further use this characterization to derive information-theoretic lower bounds on the population risk of sequential learning and on the prediction error of temporal GP. Numerical experiments corroborate the theoretical predictions.
Abstract: Vision-language-action (VLA) models show promising knowledge accumulation ability from pretraining, yet continual learning in VLA remains challenging, especially for efficient adaptation. Existing continual imitation learning (CIL) methods often rely on additional parameters or external modules, limiting scalability for large VLA models. We propose Stellar VLA, a knowledge-driven CIL framework without increasing network parameters. Two progressively extended variants are designed: T-Stellar for flat task-centric modeling and TS-Stellar for hierarchical task–skill structure. Stellar VLA enables self-evolving knowledge learning by jointly optimizing task representations and a learned knowledge space. We propose a knowledge-guided expert routing mechanism conditioned on knowledge relation and Top-K semantic embeddings, enabling task specialization without increasing model size. Experiments on the LIBERO benchmark show that Stellar VLAs achieve strong performance among both VLA and CIL baselines, using only 1% data replay. Real-world evaluation on a dual-arm platform with distinct embodiment and scene configurations validates effective knowledge transfer. TS-Stellar excels in hierarchical manipulation, and visualizations reveal robust knowledge retention and task discovery. Project website is provided in the supplementary material.
PaperID: 4767, Poster
Abstract: Molecular conjugates, including PROTACs and peptide-drug conjugates (PDCs), derive their function from the joint behavior of multiple coupled components, yet most generative approaches design these components independently and combine them only after generation. Such staged pipelines ignore cross-component dependencies and often produce conjugates that are chemically invalid, fall outside empirical conjugate distributions, or lose function upon assembly. We introduce Coupled Objective-Guided Discrete Flows for Molecular Conjugate Assembly (CupOFMoCA), a discrete generative framework that formulates conjugate design as a constrained, coupled generation problem. CupOFMoCA restricts generative trajectories to a chemically feasible conjugate manifold and biases local transitions using target-specific activity predictors, ensuring all components remain mutually compatible throughout generation. We show that coupling constraints and objective guidance enable anticipatory design that preserves post-assembly predicted activity and produces structurally realistic conjugates across both PDC and PROTAC settings, outperforming staged baselines, including LinkerNet, DiffLinker, and DiffPROTACs for PROTACs, across assembly validity, predicted activity, and physicochemical property ranges. These results demonstrate that explicit coupling and constraint enforcement are sufficient to recover functional conjugates across conjugate classes, and provide a principled foundation for generative modeling where function emerges only at the level of the assembled system. Our anonymous code repository can be found at https://anonymous.4open.science/r/Cupofmoca-Neurips.
Abstract: Memory-augmented LLM agents have gained substantial attention recently for their ability to maintain context across hundreds of interactions through agentic memory systems that actively curate retrieved content with LLM-generated metadata such as summaries, keywords, and tags. However, from an inference cost standpoint, every retrieval triggers a full re-encoding of these structured memory units into Key-Value (KV) states, which dominates prefill latency. Existing training-free KV reuse methods mitigate this by selectively recomputing a small fraction of tokens, but were designed for RAG-style raw passages and degrade substantially on structured agentic memories. In this work, we present AgentKVShift, a training-free, probe-guided KV residual correction method that operates per retrieved memory unit. One of the crucial insights we demonstrate is that the per-memory KV reuse residual decomposes into a shared memory-level offset plus small token-wise fluctuations. Estimating this offset from a small probe set allows us to correct every reused token in the memory unit by a single weighted correction. Unlike prior reuse methods which decide which tokens to recompute and leave the rest of the cache stale, AgentKVShift also corrects the tokens it does not recompute, turning the refresh budget into useful signal across the entire chunk. Through extensive experiments across four open source LLMs spanning 3B to 32B parameters and two long-horizon agentic memory benchmarks covering long-term dialogue and agentic applications, we show that AgentKVShift achieves near full recompute performance while refreshing only 10-30% of the cache, outperforming existing baselines at the same recompute ratio. AgentKVShift requires up to 5x lower recompute to reach this near-full performance, which prior reuse methods only attain at 45-55% refresh. By operating in this lower recompute regime, AgentKVShift delivers prefill speedups of 2-3.5x over no-KV-reuse on a single A100 GPU. Lastly, AgentKVShift orthogonally composes with KV cache quantization, retaining over 2x the F1 of prior reuse methods under aggressive 2- and 4-bit settings, making it a simple, yet effective choice for serving long-horizon agentic memory workloads.
Authors: Krzysztof M Choromanski, Derek Long, Ananya Parashar, Dwaipayan Saha
Abstract: We present GenusSink, a new class of approximate generalized Sinkhorn algorithms with shortest-path-distance costs for bounded genus (e.g. planar) graphs, providing near-linear time: (1) pre-processing, (2) iteration step, (3) final transport plan matrix querying and near-linear memory. Graphs handled by GenusSink include in particular planar graphs and bounded-genus meshes approximating 3D objects. GenusSink addresses total quadratic time complexity of its brute-force counterpart by leveraging separator-based decomposition of graphs, computational geometry techniques, and new results on fast matrix-vector multiplications with generalized distance matrices, using in particular Fourier analysis and low displacement rank theory. It is inspired by recent breakthroughs in graph theory on approximating bounded genus metrics with small treewidth metrics minor. The graph-centric approach enables us to target optimal transport problem with the corresponding distributions defined on the manifolds approximated by weighted graphs and with cost functions given by geodesic distances. We conduct rigorous theoretical analysis of GenusSink, provide practical implementations, leveraging newly introduced in this paper separation graph field integrators (S-GFIs) data structures and present empirical verification. GenusSink provides orders of magnitude more accurate computations than other efficient Sinkhorn algorithms, while still guaranteeing significant computational improvements, as compared to the baseline. As a by-product of the developed methods, we show that GenusSink is numerically equivalent to the brute-force geodesic Sinkhorn algorithm on n-vertex graphs with treewidth O(\log \log (n)) (e.g. on trees).
PaperID: 4770, Poster
Abstract: Variational quantum algorithms (VQAs) rely on gradient-based optimization over parameterized circuits whose loss landscape lives on a manifold of quantum states. Quantum natural gradient (QNG) accounts for this geometry through the quantum Fisher information matrix (QFIM). However, estimating the full QFIM is prohibitively expensive within practical quantum circuit shot budgets. Existing tractable approximations either discard inter-block couplings, as in block-diagonal QNG, or replace the QFIM with a measurement-induced classical Fisher surrogate, as in random natural gradient (RNG). We propose a locality-aware surrogate that targets the QFIM through its covariance representation in terms of propagated generators, retaining a structured subset of informative cross-parameter couplings while remaining practical to estimate under finite-shot budgets. Our construction combines three ingredients: weight-k Pauli truncation to control estimation variance, support-overlap masking to exploit sparsity, and single-dataset classical-shadow estimation that reuses the same shadow dataset across all retained metric entries. Truncation level k provides a bias--variance tradeoff: larger k captures more of the propagated-generator structure but increases estimation cost, making it a budget-constrained design choice. Circuit parameters whose truncated generators remain nonzero form an active set receiving surrogate-QNG updates, while the rest follow standard gradient descent. We provide finite-shot concentration bounds for the shadow-estimated surrogate and convergence guarantees to a stationary neighborhood under standard smoothness assumptions. Experiments on variational quantum eigensolver (VQE) instances with local ans\"atze for Heisenberg and Ising spin chains show improved convergence per cumulative shot relative to SGD, Adam, block-diagonal QNG, and RNG baselines under finite-shot budgets, including on systems up to 16 qubits.
PaperID: 4771, Poster
Authors: ZHIFAN ZENG, Wenbo Xiao, Arcot Sowmya, Changming Sun
Abstract: Open-vocabulary 3D Gaussian segmentation typically binds semantic features to per-Gaussian state via per-scene gradient training (1–4 h) or the more recent closed-form distillation. Both commit each primitive to a single feature; therefore, boundary Gaussians spanning multiple 2D instances are forced into a hard decision that degrades retrieval accuracy and yields unstable mask boundaries at render time. We present DIGS, a training-free framework with two stages, class-agnostic map- ping and multi-modal encoding, that decouples instance discovery from linguistic naming. At the Gaussian level, our distribution-informed representation maintains an explicit top-M label distribution per primitive, updated by a tempered Dirichlet rule; preservation of multi-peaked posteriors at boundary primitives allows a late-argmax renderer to recover clean instance boundaries. A bidirectional state machine aggregates these distributions into a persistent class-agnostic instance ontology. The instance-level encoding stage then fuses the CLIP visual embedding of each instance’s isolated 3D footprint with the CLIP-text embedding of an LLM-generated attribute description, disambiguating visually similar instances. DIGS maps a scene in 1–3 minutes on a single A100 and serves 3D-native retrieval in 5.4 ms, reaching 71.95% mIoU on LERF-OVS, surpassing the strongest training-free baseline SFS and the strongest training-based baseline LaGa, and setting new state-of-the-art performance on LERF-Mask (91.71%) and 3D-OVS(96.45%). Codes are at: https://anonymous.4open.science/r/DIGS-C4EF.
PaperID: 4772, Poster
Abstract: We introduce a constructive universality principle for structural interpretability in binary regression. Let A\in \\-1,1\\^p be binary covariates and B\in\\-1,1\\ be a binary response. We represent the binary-regression target \mathbbE[B|A] by Ising Hamiltonians over observed and auxiliary binary variables. This construction encodes covariate configurations as Boolean states and maps input-output relationships through faithful reductions from Boolean satisfiability (SAT) to ground-state-energy Ising problems. This gives a interpretable graph representation of high-order binary dependence. We showcase this principle with the Two-body, Linear, Universal, Binary Effect (T-LUBE) framework. Unlike full power-vector expansion, T-LUBE does not enumerate high-order product features. It recasts these effects through interpretable auxiliary variables and pairwise interactions on an augmented binary graph. Proof-of-concept experiments corroborate our theory.
PaperID: 4773, Poster
Abstract: Privacy mechanisms in federated learning (FL) are often calibrated without explicit regard to model scale, implicitly assuming that privacy risk remains stable as federated models grow. We challenge this assumption by introducing federated leakage, a signal-level mutual-information measure of client-specific information surviving in privatized update trajectories, and by identifying effective model dimension, rather than raw parameter count, as the operative scaling variable. Under a Gaussian signal-channel model with spectral growth of client-informative directions, we establish a regime-conditional scaling law: leakage grows as \Theta(d/\log d) with effective dimension d. This result yields a leakage-ratio notion of privacy debt, where fixed-noise deployment can expose increasing client-informative signal even when the accountant-reported DP configuration is unchanged. We therefore propose a scale-aware Gaussian calibration rule that preserves the target leakage regime up to constant factors, and show across vision and language settings that it substantially reduces cross-scale privacy drift relative to fixed-noise baselines.
PaperID: 4774, Poster
Authors:
Yifan Sun, Qiang Sheng, Ya Wu, Zhengjia Wang, Guang Yang, Shaofei Wang, Danding Wang, Juan CaoAbstract: Fine-tuning is essential for adapting Large Language Models to downstream tasks. However, this process can also inadvertently erode the critical safety alignment even when the fine-tuning data appears benign. Prior methods either introduce safety regularizers that may conflict with the primary training objective, or rely on fixed mechanisms that fail to adapt to fine-tuning dynamics. To remedy this, we reframe safety preservation as a training-time intervention and propose Online Adaptive Steering for In-Training Safety (OASIS), enabling continuous online recalibration of the steering direction during fine-tuning. Specifically, OASIS tracks a misalignment direction in activation space, then applies example-adaptive activation steering to absorb misalignment-inducing updates during training. Across three misalignment behaviors and three fine-tuning data regimes, including fully misaligned, mixed, and benign data, OASIS consistently improves safety robustness without sacrificing downstream performance.
PaperID: 4775, Poster
Abstract: Existing open-vocabulary 3D part segmentation methods typically rely on point–text matching, where semantic organization is shaped implicitly by the training loss, and both training and inference lack an explicit forward semantic optimization mechanism. We address this limitation with a framework of forward semantic dynamics, implemented with coupled semantic-dynamics blocks that organize part semantics within a single forward pass. Specifically, differentiable forward low-rank semantic shaping constrains point features to a prompt-conditioned semantic subspace, suppressing semantic drift and improving intra-part consistency. Built on this shaped state, a finite-depth spatial Hawkes process for semantic propagation models high-confidence prompt responses as sparse semantic events and propagates them across multiple steps to provide reliable evidence for ambiguous regions. The propagated context is further fed back into later feature shaping, coupling semantic organization and evidence propagation without test-time iterative optimization or additional post-processing. Extensive experiments on multiple open-vocabulary 3D part segmentation benchmarks show that our method establishes a new state of the art, outperforming the previous comparable SOTA method on all 16 reported evaluation slices by 1.88 to 5.31 mIoU points (3.84 on average) and achieving the overall best results on 14 of them.
Abstract: Data and pipeline parallelism are key strategies for scaling neural network training across distributed devices, but their high communication cost necessitates co-located computing clusters with fast interconnects, limiting their scalability. We address this communication bottleneck by introducing , relaxing the co-location requirement at the expense of introducing staleness between pipeline stages and data parallel replicas. To mitigate staleness, for pipeline parallelism, we adopt a weight look-ahead approach, and for data parallelism, we introduce an method equipped with an exponential moving average based correction mechanism. We provide convergence guarantees for both sparse averaging and asynchronous updates. Experiments on large-scale language models demonstrate that our approach matches the performance of the fully synchronous baseline, while significantly reducing communication overhead.
PaperID: 4777, Poster
Abstract: With Large Language Model (LLM) pre-training and fine-tuning shifting its focus from data volume to data quality, quality data selection has emerged as a critical research topic. Existing online data selection methods for LLM training are typically ``batch-constrained'', limiting optimization to local utility within random batches. To overcome this, we propose GAIA (Global Adaptive Instruction tuning via GAussian processes), a framework that formulates data valuation as a global estimation process. GAIA employs Gaussian Process regression to model continuous utility manifolds across the semantic space, utilizing an adaptive strategy fusion mechanism to dynamically prioritize high-utility samples. By casting the strategy-posterior update as an instance of the classical fixed-share Hedge framework for tracking the best expert, we inherit a dynamic-regret guarantee that characterizes GAIA's robustness under non-stationary quality scores during training. Empirical evaluations on three datasets demonstrate that GAIA significantly outperforms state-of-the-art baselines like GREATS, establishing our method as a scalable and robust solution for efficient instruction tuning.
PaperID: 4778, Poster
Abstract: Sparse tuning is widely used to adapt large language models due to its parameter efficiency. However, its effectiveness depends not only on the number of trainable parameters, but also on the feature-interaction topology induced by the sparse update. Limited or uneven cross-dimensional interactions may isolate subsets of feature dimensions, forming information islands that hinder task-specific information mixing. To address this issue, we propose ), a sparse tuning framework that induces small-world connectivity in parameter space. SWIFT combines block-diagonal sparse updates for local feature interactions with a fixed sufficiently scattering permutation that routes these updates to distant feature groups, creating sparse long-range shortcuts without additional trainable parameters. We further show theoretically that random or block-diagonal sparse tuning can suffer from information islands, whereas SWIFT restores positive expansion through permutation-induced routing. Extensive experiments on commonsense reasoning, natural language generation, and image classification demonstrate that SWIFT consistently outperforms competitive PEFT baselines and can match or surpass full fine-tuning.
PaperID: 4779, Poster
Abstract: Real-time quantum error correction (QEC) is a critical bottleneck for fault-tolerant quantum computing due to strict hardware latency constraints. While neural decoders achieve state-of-the-art logical accuracy, their reliance on high-precision, compute-intensive inference precludes real-time deployment. Conversely, uniformly quantizing these models to extreme low-bit regimes for FPGAs triggers a catastrophic collapse in decoding accuracy. To overcome this precision-accuracy bottleneck, we propose Precision Pyramid (\mathsfPP), a hardware-algorithm co-designed neural decoder applicable to both surface and BB codes. \mathsfPP features a monotonically escalating activation precision hierarchy built upon a globally weight-binarized (W1) foundation. By scaling intermediate activations down to INT2 for the massive perception stages and reserving higher precision strictly for the lightweight final logical decision, this hierarchy effectively shifts over 99% of the computational workload to abundant FPGA look-up tables. Comprehensive evaluations on both simulated and hardware-calibrated noise data demonstrate that \mathsfPP consistently suppresses logical errors below highly optimized classical baselines (e.g., MWPM and Relay-BP) while seamlessly satisfying stringent sub-microsecond latency budgets.
Abstract: We study the problem of human gaze modeling, which aims to generate the gaze patterns a viewer produces while observing a visual stimulus. Gaze is primarily captured through two modalities: continuous eye-tracking trajectories, which describe fine-grained motion dynamics, and discrete scanpaths, which describe high-level fixation structure. Because gaze varies substantially across viewers and trials, we treat this variability as a defining property rather than noise and model gaze as a stochastic generative process. Existing generative gaze models supervise on only one of these two representations in isolation. We hypothesize that trajectories and scanpaths describe gaze at complementary scales and are jointly informative during training, and test this hypothesis through ST-DiffEye, a joint trajectory-scanpath diffusion framework that couples both modalities by concatenating them as an additional raw input channel, requiring no architectural overhead beyond an input and output channel expansion. We further introduce a principled evaluation framework based on the Continuous Ranked Probability Score (CRPS), which generalizes any existing sequence similarity metric into a proper scoring rule that jointly assesses the accuracy and diversity of generated gaze. Experiments on task-driven visual search, covering both target-present and target-absent scenarios, and on free-viewing benchmarks demonstrate state-of-the-art performance. These results, along with detailed ablations, confirm the benefit of joint modeling and the value of distribution-aware evaluation in capturing the intrinsic variability of human gaze.
Abstract: Wearable devices enable continuous health monitoring from multimodal signals, but real-world deployment is hindered by limited labeled data and pervasive sensor incompleteness. While large-scale self-supervised pretraining reduces label dependence, most existing methods assume full modality availability. Current approaches for handling modality missingness often reconstruct entire absent signals, which can encourage hallucinating modality-specific details that are not inferable from the observed sensor signals and degrade robustness. We propose VCR, a self-supervised framework that learns to extract valid representations robust to modality missingness. VCR employs an orthogonal tokenizer to enforce strict orthogonal disentanglement by rectifying latent manifolds and applying a geometric projection, separating each modality into shared semantics and modality-specific residuals. This design preserves complete information integrity while serving as a structural foundation for robust learning under modality missingness. The resulting tokens are processed by a missing-aware mixture-of-experts backbone that adapts to varying patterns of modality availability. By constraining the objective to reconstruct only the shared components of missing modalities, VCR effectively mitigates hallucinations of non-inferable modality-specific details. Across multiple health monitoring tasks, VCR consistently improves performance and robustness under full, single-missing, and multiple-missing modality settings compared with strong supervised and self-supervised baselines.
Abstract: We study policy optimization for infinite-horizon, discounted constrained Markov decision processes (CMDPs). While existing theoretical guarantees typically hold for the mixture policy, deploying such a policy is computationally and memory intensive. This leads to a practical mismatch where a single (last-iterate) policy must be deployed. Recent theoretical works have thus focused on proving last-iterate convergence, but are largely limited to the tabular setting or to algorithmic variants that are rarely used in practice. To address this, we use the classic inexact augmented Lagrangian (\textttAL) method from constrained optimization, and propose a general framework with provable last-iterate convergence for CMDPs. We first focus on the tabular setting and propose to solve the \textttAL sub-problem with projected Q-ascent (\textttPQA). Combining the theoretical guarantees of \textttPQA and the standard \textttAL analysis enables us to establish global last-iterate convergence. We generalize these results to handle log-linear policies, and demonstrate that an efficient, projected variant of \textttPQA can achieve last-iterate convergence with comparable guarantees as prior work. Finally, we demonstrate that our framework scales to complex non-linear policies, and evaluate it on continuous control tasks.
PaperID: 4783, Poster
Abstract: Feedforward network (FFN) layers account for a large fraction of parameters and nonlinear expressivity in Transformer-based large language models (LLMs). Despite the evolution from ReLU and GELU to gated variants such as SwiGLU, most FFN designs still use a single fixed activation function, applying the same nonlinear transformation to all tokens. In this work, we propose (MoA), a token-adaptive FFN design that mixes a dictionary of activation functions using lightweight input-dependent gates while sharing the same linear projections. As an input-independent counterpart, we also introduce learnable activations (LA), which form linear combinations of activation functions for both ReLU-type and SwiGLU-type FFNs. , we establish strict finite-width expressive separations among fixed-activation FFNs, LA, and MoA: LA strictly contains fixed-activation FFNs, while MoA strictly contains LA, with the additional expressivity arising from input-dependent nonlinear hybridization. , we evaluate MoA through extensive pre-training experiments on dense and MoE language models ranging from 0.12B to 2B parameters under different token budgets, optimizers, and learning rate schedules. MoA consistently achieves lower terminal loss and exhibits more favorable scaling behavior than well-tuned baselines, with minimal parameter and computational overhead. These results suggest that token-adaptive activation mixing is a simple and effective mechanism for improving FFN expressivity in LLMs.
PaperID: 4784, Poster
Authors: LIBO CHEN, Souvik Ghosh, Vinay Namboodiri
Abstract: We cast image segmentation refinement as a discrete diffusion process. To solve this task we introduce Cyclic Discrete Diffusion (CDD), a novel forward–reverse formulation with a cyclic label topology that maps a discrete label set to the uniform distribution and back. Instead of injecting Gaussian noise as is done for continuous diffusion, CDD evolves segmentation masks through label jumps: at each step, every pixel either remains in its current class or jumps to the next label with a fixed rate. This jump-or-stay mechanism defines a finite-state Markov process with bounded forward and reverse steps, featuring analytically consistent and tractable reverse dynamics. CDD operates natively in a multi-class label space, enabling joint refinement of all classes in a single pass without requiring class-wise decomposition, as commonly needed in prior binary refinement approaches. The proposed formulation admits a finite and controllable number of refinement steps, for which we provide a theoretical characterization based on spectral gap analysis. We evaluate CDD across diverse segmentation tasks, including multi-class semantic, instance-level, and binary segmentation refinement. Our method consistently improves both region accuracy and boundary quality across multiple backbones and datasets.
PaperID: 4785, Poster
Abstract: Large reasoning models (LRMs) have achieved remarkable progress in complex problem-solving tasks. Despite this success, LRMs typically suffer from high computational costs during deployment, highlighting a need for efficient inference. A practical direction is to switch the LRM between thinking and non-thinking modes dynamically. However, such approaches often introduce additional reasoning errors and lack statistical guarantees for the performance loss, which are critical for high-stakes applications. In this work, we propose Probably Approximately Correct (PAC) Reasoning that controls the performance loss under the user-specified tolerance. Specifically, we construct an upper confidence bound on the performance loss and determine a threshold for switching to the non-thinking model. Theoretically, using the threshold to switch between the thinking and non-thinking modes ensures bounded performance loss in a distribution-free manner. Our experiments on reasoning benchmarks show that the proposed method can save computational budgets and control the user-specified performance loss.
PaperID: 4786, Poster
Authors: Chen Wang, Yinxuan Huang, Yexin Cui cui, Maoqing Zhong, LaiLong Luo
Abstract: Watermark robustness under paraphrase is usually reported as AUC against a named rewriter, making results hard to compare across watermark schemes, paraphrasers, and detectors. We propose a retained-signal reporting interface: if a watermark injects KL signal \Delta=\KL(P_1\Vert P_0) and a paraphrase channel retains KL fraction \rho, then any detector's post-paraphrase advantage is controlled by \Delta\rho. We give explicit Pinsker and Bretagnolle--Huber threshold laws, prove the \Theta(\sqrt\Delta\rho) rate is sharp for any (\Delta,\rho)-only converse, and show the bound is implementation-blind once watermark families are matched by \Delta. For real LLM paraphrases, we instantiate the interface as a score-projected audit that estimates \widehat\Delta\widehat\rho for a deployed detector and reports it with semantic preservation and the detector projection. In a KGW/Unigram × Qwen/Llama benchmark, this score-effective signal predicts post-paraphrase AUC across lengths, strengths, and paraphraser families. The resulting protocol separates injected signal, retained signal, semantic quality, and detector choice, replacing one-off AUC tables with a reproducible robustness report.
Abstract: Open-vocabulary object detection (OVD) has made significant progress, enabling detectors to generalize from seen to unseen categories. However, real-world category spaces continually evolve, and existing OVD models still struggle with newly emerging concepts, while repeated full retraining is prohibitively expensive. To this end, we introduce a new task setting, termed Continual OVD with Novel Concept Injection (COVD), where models sequentially learn incoming novel concept groups while preserving prior concepts and original open-vocabulary knowledge, along with a new benchmark, Novel-114. Our key observation is that pretrained visual encoders often already perceive and represent many novel concepts, and the main bottleneck lies in the lack of stable semantic alignment between visual representations and textual concepts. Based on this, we propose NoIn-Det, an efficient continual injection framework without additional parameters. NoIn-Det freezes the visual encoder, preserves the text representation space using only texts of common concepts and previously injected concepts, and injects novel concepts by updating only a small subset of text-branch parameters beneficial to novel concept learning. Extensive experiments show that NoIn-Det effectively learns novel concepts, preserves old knowledge, and consistently outperforms existing continual learning methods for VLMs without introducing additional parameters.
Abstract: Randomization is often used to compromise among fixed policies, classifiers, or interventions, but affected stakeholders may be unable or unwilling to report utilities over all alternatives. We study a weaker feedback model in which, for each proposed lottery, each stakeholder reports only whether it is acceptable. Unlike prior work focused on finding a unanimously acceptable lottery when one exists, we allow unanimity to be impossible and seek lotteries that minimize aggregate acceptability violations. We model each stakeholder's acceptable lotteries by an unknown linear threshold rule over the lottery simplex, and the learner may only ask accept/reject queries. Since binary labels do not reveal utility shortfalls, and such shortfalls are not comparable across stakeholders, we measure violation geometrically: the Euclidean distance, within the affine hull of the simplex, to the stakeholder's normalized acceptance hyperplane. These distances are aggregated using the power mean. Our results characterize exactly which boundary information can be learned from these queries. A deterministic algorithm uses \mathcalO(nm\log(m/\varepsilon)) queries to recover the normalized boundary whenever violation distances are identifiable from query answers. Otherwise, queries determine the acceptable face but not the rate at which violation grows away from it; this ambiguity is unavoidable. When no ambiguity arises, the aggregate violation objective can be optimized exactly. When it does, we optimize a convex worst-case upper bound objective based on distance to the learned acceptable set and prove an additive loss bound of \sqrt2(|B|/n)^1/p. We also prove structural properties of optimal lotteries, including a support-size bound of n+1, and give a limited-budget algorithm with additive-\alpha guarantees. Synthetic experiments illustrate the resulting breadth-depth tradeoff and the effect of the power-mean parameter p on violation distribution.
Abstract: holds, i.e., different agents bring unique information to inform a final decision. We propose ComplLLM, a post-training framework based on decision theory that fine-tunes a decision-assistant LLM to output signals that complement existing agent decisions, using complementary information as reward. We validate ComplLLM on synthetic and real-world tasks involving domain experts, demonstrating how the approach recovers known complementary information, produces plausible explanations of complementary signals to support downstream decision-makers, and improves human and LLM agent decision performance in controlled studies.
Abstract: Identifying which tokens and activations truly influence model predictions is critical for both the efficiency and interpretability of large auto-regressive language models. Yet in practice, token importance is typically inferred from attention distributions that entangle contextual relevance with softmax-normalization effects, obscuring relative influence across tokens and heads. We propose RCStat, a statistical framework for quantifying contextual influence in attention mechanisms. At its core Relative Contextualization (RC) is a random variable that measures how strongly one subset of tokens contributes to another under the model’s internal scoring scheme. RCStat admits computationally efficient bounds on expected influence that can be estimated at inference time. No retraining is required. We apply RCStat to two tasks. For attribution, attention heads with high expected RC accurately identify the prompt spans that drive generation, enabling reliable span-level explanations. For key-value cache eviction, RC-based adaptive thresholding selectively evicts low-impact KV entries, substantially reducing cache size while preserving generation quality. Across question answering, summarization, and attribution benchmarks, RCStat achieves consistent gains, improving generation quality by 15–40% with upto 36% error reduction for KV eviction, and attribution accuracy by 2–16%,. These results demonstrate that explicitly modeling contextual influence provides a principled and practical alternative to attention-based heuristics.
Authors:
Zehao Jin, Ruixuan Deng, Junran Wang, Xinjie Shen, Chao ZhangAbstract: Activation steering has emerged as a promising alternative for controlling language-model behavior at inference time by modifying intermediate representations while keeping model parameters frozen. However, large-scale evaluations such as AxBench show that existing steering methods are often outperformed by simple in-context prompting and generalize poorly to unseen concepts. We hypothesize that these limitations arise from unvalidated simplifying assumptions shared across prior methods, which typically restrict steering interventions to fixed, single-step, position-invariant transforms. We propose FLAS (Flow-based Activation Steering), which learns a general, concept-conditioned velocity field v_t(h,t,c) that transports unsteered activations to steered ones without relying on these assumptions. On AxBench, FLAS is the first learned method to consistently outperform prompting, reaching held-out harmonic means of 1.015 on Gemma-2-2B-IT and 1.113 on Gemma-2-9B-IT without per-concept tuning. Analysis of the learned flow shows curved, multi-step, token-varying trajectories, which suggests that previous hypotheses on activation space geometry might be incomplete. Our code is available at https://anonymous.4open.science/r/FLAS.
Abstract: Diffusion and flow generative models sample by integrating a learned ODE, but high quality still requires many sequential model evaluations. Solver learning reduces this cost by adapting scalar coefficients, timesteps, or both, while keeping the backbone model fixed. In this work, we identify a structural bottleneck in this update family: each step remains span-limited. Each update remains a scalar-weighted combination of buffered velocity evaluations, so learning can reduce the in-span teacher mismatch but cannot represent the out-of-span residual exposed by large steps. We propose SpanLift, a lightweight operator-augmented neural solver that enlarges the update family beyond scalar-coefficient solvers. SpanLift keeps a fixed base solver as an in-span prior and learns a spatial residual operator over the state and velocity buffer. The operator is trained by endpoint teacher matching, preserves the pretrained backbone, and adds no model NFEs. Empirically, the learned correction transfers across base solvers and is predominantly out-of-span. Across pixel-space diffusion, latent flow matching, and precipitation nowcasting, SpanLift achieves state-of-the-art few-step sampling. With only 3 NFE, it improves CIFAR-10 FID from 8.16 to 5.69 and ImageNet FID from 17.37 to 11.83.
PaperID: 4793, Poster
Abstract: Neural networks routinely develop representations in which individual neurons respond to mixtures of task-relevant variables, leading to the widespread view that such networks lack modular organization. We challenge this conclusion by introducing a formally nested hierarchy of modularity: unit-level, which requires distinct computations to occupy disjoint neurons; orthogonal, which asks whether a rotation reveals functionally independent subspaces; and linear, which asks whether a bounded linear map can separate task-relevant computations into independent subspaces. We develop a unified framework for learning decompositions at each level and evaluating them causally through selective ablation. Across vision models, anguage models and multitask RNNs, we find that networks which lack unit-level modularity nonetheless often exhibit clear functional modularity at the subspace level. In vision and language models, content versus style and syntax versus semantics respectively dissociate under subspace ablation despite sharing overlapping populations of neurons. In RNNs trained on multiple cognitive tasks and previously shown to lack modular structure, our decomposition recovers sharp double dissociations between tasks. These results demonstrate that modularity cannot be assessed at a single level of description. The hierarchy we introduce provides a principled framework for determining when and at what level of analysis functional specialization exists in neural network representations.
PaperID: 4794, Poster
Abstract: Despite remarkable progress in image-text understanding, vision-language models (VLMs) still struggle with compositional reasoning. In particular, they often fail to distinguish relational direction and attribute-object binding, leading to similar representations for semantically different image-text pairs. This problem mainly stems from the reliance on global image-text alignment, which captures coarse correspondence but overlooks fine-grained compositional structures. Toward this end, we propose an evidence-aware framework, termed Directional Relation Bucketing with Binding Localization (DELTA), to improve compositional understanding in VLMs. The key idea of our DELTA is to improve compositionality of VLMs from two complementary perspectives: direction modeling and binding localization. More specifically, we first leverage a learnable gate to model the cumulative contextual distance between anchor terms, thereby assigning relation concepts to direction-aware discrete buckets. To enrich textual representations with fine-grained relational semantics, we calibrate global semantic representations by incorporating intermediate-layer hidden states. Furthermore, DELTA leverages textual cues to ground visual evidence for object attributes, pulling image patches closer to their matched attribute descriptions while pushing them away from incorrect augmented ones. Extensive experiments on four compositional datasets demonstrate the effectiveness of the proposed DELTA. Our implementation is available at https://anonymous.4open.science/r/DELTA-2D60.
Abstract: Compositional reasoning is critical for real-world problem solving: since training data is necessarily limited, models must generalize by composing learned skills in new ways. While post-training methods such as reinforcement learning (RL) have substantially improved the reasoning abilities of language models (LMs), their effects on compositional reasoning remain less well understood. We propose a dependency-graph framework to formalize compositional reasoning, yielding three levels of compositionality with increasing complexity. Empirically, we instantiate this framework with data-structure tasks, which provide deterministic reward computation and clear compositional structure. We find a consistent decomposed-to-composed asymmetry: decomposed-skill training does not reliably transfer to composed tasks, whereas composed-task training transfers more readily back to decomposed tasks. We provide theoretical explanation for this asymmetry, and further evaluate compositional generalization under length extrapolation, structural distribution shift, and transfer to tasks requiring unseen skills. Finally, we present a pilot study on real-world tool-calling benchmarks, showing preliminary evidence that the decomposed-to-composed asymmetry can extend to practical settings.
PaperID: 4796, Poster
Abstract: In this paper we study active learning for constrained linear regression. Given a design matrix A \in \mathbbR^n × d, a constraint set \mathcal C \subseteq \mathbbR^d, and query access to a response vector b \in \mathbbR^n, we seek to output an approximate solution to the problem \min_x\in \mathcalC \lVert Ax - b\rVert_2^2 using as few queries to b as possible. This problem arises in many scientific and engineering applications, and _randomized leverage score sampling_ provides a powerful method to solve it. However, existing leverage score methods either ignore the constraint or encode it only through ridge regularization, thereby suffering poor query complexity for anisotropic constraint sets. We propose _ellipsoid-regularized leverage scores_, where we first approximate the constraint set by an outer ellipsoid and then use its shape matrix to bias sampling toward directions that matter under the constraint geometry. We prove that the resulting query complexity is governed by a natural ``anisotropic effective dimension`` of the problem, d_\texteff(\mathcalC) \leq d. In particular, we show that \tilde O(d_\texteff(\mathcalC)/\varepsilon^2) label queries suffice to return a feasible point whose objective value is at most (1+\varepsilon) times the optimum. This bound is typically much smaller than what would be obtained using standard leverage scores. Furthermore, for convex constraints, we show how to improve the dependence on \varepsilon from 1/\varepsilon^2 to 1/\varepsilon. We complement these upper bounds with two lower bounds: first, a lower bound showing that the 1/\varepsilon^2 dependence is unavoidable without convexity assumptions on \mathcal C, and second, a lower bound showing that the effective-dimension dependence is unavoidable for ellipsoidal constraints. We also provide experiments supporting our theoretical results.
PaperID: 4797, Poster
Authors: Yuming Huang, Haoran Yin, Wu
Abstract: Subgraph mining aims to discover meaningful subgraphs from a large graph given a seed node, with applications in community detection, functional module identification, anti-money laundering, and knowledge discovery. Existing methods typically adopt static approaches that cannot dynamically and intelligently adapt to diverse graph structures. Moreover, training models for subgraph mining faces inherent challenges including exposure bias and the difficulty of learning when to stop expansion. To address these limitations, we propose DASM (Dynamic Autoregressive Subgraph Mining), a novel two-stage framework that models subgraph mining as a sequential decision process. In the first stage, we employ step-wise supervised learning with multi-label objectives and asymmetric F1 loss to teach the model how to select correct next nodes. In the second stage, we apply Group Relative Policy Optimization (GRPO) to align the model with end-to-end mining quality, enabling it to learn optimal stopping decisions through trajectory-level feedback. Our dynamic autoregressive approach allows the model to make state-aware decisions at each step, automatically determining subgraph boundaries without predefined size constraints. Experiments on multiple datasets demonstrate that DASM significantly outperforms existing baselines in terms of F1 score and IoU.
PaperID: 4798, Poster
Authors: Akihiro Yamaguchi, Shizuo Kaji, Kaname Matsue, Ryusei Shingaki
Abstract: Counterfactual explanations for time-series classification generate synthetic instances that flip a prediction to a desired class while remaining plausible and making small, local changes. Existing methods often rely on classifier gradients, and many do not naturally extend to one-class settings. To bridge this gap, we propose CTCF, a model-agnostic method for hard-decision classifiers in supervised and one-class settings. CTCF decouples generation from classifier-specific control by reusing an unconditional Flow Matching generator and learning a flow-time-dependent desired-class region from interpolation states labeled by endpoint hard decisions, without classifier gradients or class probabilities. At inference time, CTCF uses a dual-control mechanism: fused-lasso endpoint-objective steering encourages small, segment-local changes, while linearized projection correction reduces violations of a conservative desired-class region. Theoretically, under ideal marginal matching, we show that the surrogate risk on interpolation states equals the corresponding ODE-trajectory risk. Experiments on UCR datasets demonstrate CTCF's effectiveness across quantitative metrics and qualitative case studies.
Abstract: Language model watermarking schemes fall into two broad categories: inference-time methods, which modify the decoding process at generation time, and model-embedded methods, which encode a secret signal directly into the model weights. Inference-time watermarking methods incur additional inference overhead and are not applicable in the increasingly prevalent open-weight model settings. In contrast, model-embedded watermarks introduce no additional inference latency---generation uses the standard sampling pipeline---and are particularly well-suited to open-weight settings. However, existing model-embedded approaches such as GaussMark face a fundamental quality-detectability trade-off: achieving strong detection power typically requires weight perturbations that noticeably degrade generation quality. We introduce MarkTune, a theoretically grounded on-policy fine-tuning framework that treats the GaussMark detection statistic as a reward while explicitly regularizing for text quality. Empirically, MarkTune substantially improves the quality-detectability frontier of GaussMark, approaching the detectability of strong inference-time schemes while preserving generation quality and downstream task performance. MarkTune is also extremely robust to paraphrasing and fine-tuning attacks, and generalizes across datasets: models fine-tuned on one corpus retain substantial detection power on unseen data.
PaperID: 4800, Poster
Abstract: Camouflaged Object Segmentation (COS) aims to identify objects visually hidden in their surrounding environments. Existing COS benchmarks mainly focus on image-level or video-level visual perception, while real-world camouflaged scenes are often accompanied by audio signals and user intentions, where sound semantics and textual expressions provide critical information for localizing hidden targets. To extend the boundary of COS, we introduce a new task, termed Referring and Reasoning Audio-Visual Camouflaged Object Segmentation (R2-AVCOS), which aims to segment camouflaged targets in audio-visual scenes according to textual expressions with referring or reasoning intentions. This task emphasizes understanding audio content and incorporates complex reasoning and world knowledge into expressions. To support this task, we construct R2-AVCOSBench, the first audio-visual benchmark with pixel-level annotations for camouflaged objects specified by referring expressions or inferred through reasoning. It contains 2,654 audio-visual camouflaged videos, 21,232 annotated frames, and 32,329 expressions, including 17,543 referring and 14,786 reasoning expressions. Furthermore, we propose Camouflaged Instructed Segmentation Assistant (CISA), a baseline model built upon a Multimodal Large Language Model (MLLM). CISA understands complex textual and audio-visual cues and performs referring- and reasoning-based camouflaged object segmentation. Extensive experiments show that CISA achieves strong referring and reasoning segmentation ability in audio-visual camouflaged scenes and obtains competitive results on related tasks.
Abstract: Existing large language model (LLM)-based memory systems apply universal, static policies that overlook a fundamental reality: the contexts that are worth storing in memory are different across users. This misalignment wastes limited memory budget on transient interactions while failing to preserve critical context for long-horizon tasks. To address this gap, we investigate an underexplored question: can LLM-based memory systems learn personalized memory policies? We introduce PerMemBench, the first benchmark for evaluating personalized memory systems, featuring multi-year, multi-domain interaction histories across diverse user personas. We further present the first empirical study of memory personalization and propose simple baseline methods. Our empirical study confirms that personalization yields substantial retention gains when the user profile is exactly inferred, yet reveals that accurate profile inference remains an open and critical challenge.
Authors: Jiaying Lin, Dan Xu
Abstract: Affordance segmentation in 3D scenes requires an agent to ground implicit natural-language instructions into precise masks of fine-grained interactive elements. Existing training-free methods typically rely on fragmented pipelines, which introduce visual blindness during task parsing and limit accuracy through single-scale spatial and temporal processing. We present UniFunc3D, a unified and training-free framework that treats the multimodal large language model as an active observer. By consolidating semantic, temporal, and spatial reasoning into a single forward pass, UniFunc3D performs joint reasoning to ground task decomposition in direct visual evidence. Our approach introduces active spatial-temporal grounding with a coarse-to-fine strategy. This allows the model to select correct video frames adaptively and focus on high-detail interactive parts while preserving the global context necessary for disambiguation. On SceneFun3D, our UniFunc3D achieves state-of-the-art performance, surpassing prior training-free methods by a large margin with a relative 59.9% mIoU improvement, and even outperforming training-based methods without any task-specific training. Code will be released.
PaperID: 4803, Poster
Abstract: Reconstructing 3D scenes from sparse views without per-scene optimization remains highly challenging, especially for recovering accurate geometry and fine textures. While recent feedforward approaches leverage generalizable 3D Gaussian Splatting (3DGS) for scene generation, they typically assign one or multiple Gaussians to each pixel. Such uniform allocation not only produces highly redundant representations by wasting primitives in homogeneous areas, but also fails to exploit the inherent geometric priors of planar surfaces, often resulting in sub-optimal structural modeling. To address these limitations, we present , a novel feedforward framework that enables a content-aware allocation of 3D Gaussians. By integrating texture-guided spatial partitioning with hierarchical sampling, CompactSplat adaptively distributes primitives — concentrating dense Gaussians in complex regions while tiling flat areas with expanded, scale-modulated primitives. Moreover, our framework inherently supports , enabling seamless trade-offs between rendering fidelity and memory footprint without requiring network retraining. Extensive experiments on RealEstate10K, DL3DV, and ScanNet demonstrate that CompactSplat consistently outperforms prior methods on both standard metrics and high-resolution rendering consistency, achieving
Authors:
Guozhen Zhang, Xuerui Qiu, Yutao Cui, Tianhui Song, Changlin Li, Junzhe Li, Tao Huang, Xiao Zhang, Yang Li, Jianbing Wu, Miles Yang, Zhao Zhong, Liefeng Bo, Limin WangAbstract: Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space. In this paper, we present HYDRA-X, the first UMM framework that unifies image and video tokenization within a single Vision Transformer (ViT). Our design is driven by two core challenges: efficiently injecting spatiotemporal reconstruction capability into a native ViT, and embedding image- and video-level semantic awareness into the latent space. To address the first, comprehensive ablations reveal two key findings: (1) frame-level causal temporal attention suffices for visual reconstruction, whereas full spatiotemporal attention degrades it; and (2) hierarchical temporal compression substantially outperforms single-step alternatives. To tackle the second, we propose a lightweight decompressor that upsamples temporally compressed features under joint image-video teacher supervision, thereby enforcing complementary semantic structures within the compact latent space. Building on this holistic tokenizer, we further propose a principled improvement of the editing pipeline: source-target interaction should occur at the latent level inside the tokenizer rather than at the semantic level inside the LLM, substantially improving editing consistency and accelerating convergence. Instantiated at the 7B dense model, HYDRA-X achieves strong performance across image and video understanding and generation tasks, laying a solid foundation for future unified-tokenizer UMM exploration.
PaperID: 4805, Poster
Abstract: Mixture-of-Experts (MoE) language models improve parameter efficiency by activating only a small subset of experts per token, but all experts must still be stored at inference time, making memory a key deployment bottleneck. Expert pruning is therefore a practical post-training compression strategy, but existing methods face a fundamental tension: search-based approaches attempt to capture expert interaction effects through joint optimization, yet become intractable for modern fine-grained MoE models; score-based approaches, by contrast, ignore such effects, but remain efficient and often surprisingly effective. To understand this tension, we formulate expert pruning as a constrained pseudo-Boolean optimization problem and empirically analyze the interaction effects induced by pruning. Our results show that cross-layer interaction effects are relatively more important, since pruning earlier layers changes the hidden-state distribution for later ones, whereas within-layer effects are often well approximated by additive singleton damages. Based on this observation, we propose MoE-PBA, a prefix-conditioned pruning method that prunes layers sequentially under the hidden states produced by the already-pruned prefix. Across four representative fine-grained MoE language models with distinct architectures and parameter scales ranging from 7B to 30B, MoE-PBA achieves a substantially better accuracy--efficiency trade-off than search-based methods and outperforms the strong score-based baseline on diverse downstream benchmarks.
Abstract: Accelerating stochastic gradient methods with classical momentum schemes, such as Polyak's heavy ball, has proven highly successful in training large-scale machine learning models, particularly when combined with the hardware acceleration of large mini-batch computations. Yet, the effect of classical momentum on stochastic mini-batch optimization has been poorly understood theoretically, with prior works requiring strong noise assumptions and extremely large mini-batches. In this work, we develop a general theory of stochastic momentum acceleration for optimizing over quadratics in the interpolation regime, a popular abstraction for studying deep learning dynamics which also includes classical methods such as randomized Kaczmarz and coordinate descent. Our framework encompasses both heavy ball and Nesterov-style momentum, allows for arbitrary mini-batch sizes, and makes minimal assumptions on the stochastic noise. In particular, we show that acceleration from classical momentum is directly proportional to the gradient mini-batch size (up to a natural saturation point), thereby enabling perfect parallelization of mini-batch computations. Our theory also provides a simple choice for the momentum parameter, which is shown to be effective empirically.
PaperID: 4807, Poster
Authors: Junbo Li, Hang Fu, Yinuo Wang, Weimin Yuan, Cai Meng, Xiangzhi Bai
Abstract: Dynamic scene representation and rendering in the visible spectrum has been extensively studied. Compared to visible-light imaging, thermal infrared sensing offers all-weather observation and strong penetration capability, enabling perception in low-visibility environments. However, dynamic reconstruction in thermal infrared scenes is challenged due to atmospheric transport effects and interference from low-emissivity background regions, which often leads to motion blur and details loss in rendered results. To address these issues, we propose DTGS, a physics-driven dynamic 3D scene reconstruction approach. DTGS factorizes dynamic thermal scenes into radiometric and geometric components, modeling radiometric changes via rendering that integrates atmospheric radiative transfer, while capturing geometric deformations through Adaptive Temporal Gaussian basis. In addition, we design a thermal-radiance-weighted structural similarity loss to suppress gradient interference from low-radiance noisy regions. Furthermore, to demonstrate the effectiveness of our method, a large-scale dataset for this field named Dynamic_LTR is created. Experimental results demonstrate that our method achieves a 2.7 dB improvement in PSNR, a twofold speedup in rendering, and a threefold reduction in training time. The code and datasets will be available after acceptance.
PaperID: 4808, Poster
Abstract: We introduce Track4D, a method that formulates dense 3D tracking as conditional video generation. Repurposing large-scale pretrained video generators for 3D tracking is non-trivial: 3D point tracking in world space produces signals far from natural video, making it difficult for the video diffusion model (VDM) to learn. We systematically study two representations for encoding 3D tracking in the VDM latent space: a Residual 3D Tracking Video (RTV) that directly encodes metric 3D offsets, and a Normalized Coordinate Map Video (NCMV) that implicitly encodes the underlying geometric correspondences as canonical 2D coordinate maps. We further condition the VDM on point maps from an off-the-shelf 3D reconstructor to provide an explicit geometric scaffold for tracking. Trained exclusively on limited synthetic data with LoRA adaptation, Track4D achieves competitive zero-shot tracking performance across five benchmarks, demonstrating the potential of leveraging video generation priors for dense 3D tracking.
Authors: Sai-Aakash Ramesh, Archit Sood, Andrew Corbett, Tim Dodwell
Abstract: Learning representations that capture both intrinsic data geometry and target-relevant structure remains a fundamental challenge, particularly in settings where data reduction must balance compression with predictive fidelity. While distributional reduction—encompassing joint clustering and dimensionality reduction—offers a principled way to summarize data, its supervised variants remain relatively underexplored, despite the importance of retaining task-relevant signal for downstream prediction and decision-making. We propose Supervised Distributional Reduction (SDR), an algorithm for learning target-aware representations by combining optimal transport with explicit dependence maximization. SDR builds on the Fused Gromov–Wasserstein (FGW) objective to align the relational structure of the input distribution with a set of representative points, while augmenting it with a direct dependence term that encourages the learned embeddings to capture predictive signal more explicitly. This results in compact representations that reflect both geometric structure and supervision. Beyond representation learning, SDR naturally induces a data-dependent, non-stationary geometry that can be leveraged for settings such as Gaussian Process (GP) modelling. By redefining distances through target-aware distributional alignment, SDR enables the construction of adaptive kernels that respond to local variations in both data geometry and supervision, offering an optimal transport-based perspective on non-stationary kernel design.
PaperID: 4810, Poster
Authors: Dong Zhang, Cheng Xue
Abstract: Surgical reasoning segmentation requires a model to infer the target instrument or anatomical structure from an implicit clinical query and produce a pixel-level mask. This setting is far more challenging than category-level surgical segmentation, as targets are specified by functional role, spatial relation, or anatomical interaction rather than explicit class names. It also exposes a key limitation of existing reinforcement learning-based visual grounding methods: standard reward normalization tends to under-optimize clinically important hard samples, such as small instruments, occluded organs, and ambiguous targets. To address this, we introduce SurgSeg-14K, a surgical reasoning segmentation benchmark with over 14K image-mask pairs across seven surgical scenarios and 30 fine-grained categories. Each sample links an implicit clinical query to a binary mask, spatial annotations, and a reasoning trajectory. We further propose SurgReasoner, which decomposes the task into two stages: clinical query interpretation for spatial prompt generation, and prompt-based mask generation with a frozen segmentation model. To improve reinforcement learning for spatial reasoning, we introduce a dynamic difficulty-aware reweighting strategy that combines intrinsic target difficulty with rollout correctness, enabling training to focus on challenging targets. Experiments on SurgSeg-14K show that SurgReasoner outperforms the strongest grounding baseline by large margins, with consistent gains especially on hard samples, suggesting a practical path toward reasoning-driven surgical visual perception.
PaperID: 4811, Poster
Abstract: Long-video question answering forces multimodal large language models to reason under strict visual-token budgets, creating a tradeoff between broad temporal coverage and fine-grained spatial detail. Uniform sampling preserves coverage but often obscures decisive local cues, while caption-based or frame-similarity retrieval operates over lossy proxies that can miss motion, state changes, and before-after relations. We recast long-video reasoning as a visual token allocation problem: given a fixed budget, decide how much to spend on global context versus high-resolution local evidence, and where to draw that local evidence from. We introduce AdaAlloc, a training-free inference framework that adaptively allocates visual tokens between a global video overview and localized visual detail. AdaAlloc first plans a global/local allocation policy, then uses temporal grounding to locate candidate evidence segments and refines them into a compact temporal memory. The final answerer reasons over the allocated global context and refined local evidence. Under matched backbones and visual-token budgets, AdaAlloc consistently improves over strong baselines on challenging long-video benchmarks, including LVBench, MLVU, VideoMME, and LongVideoBench.
PaperID: 4812, Poster
Authors:
Fengyuan Ran, Yikuan Wang, Yanming Li, Wenjie Lu, Tongtong Wu, Senquan Yi, Yuheng Wang, Qiqi Lin, Yuxin Wu, Minghui Zhou, Naiqiang Tan, Li ShenAbstract: LLM-based multi-agent systems (MAS) improve complex reasoning through collaborative decomposition, verification, and synthesis, but their performance can degrade sharply when compromised agents inject misleading intermediate messages. Existing topology methods rely on fixed communication patterns or query-adaptive graphs, but rarely infer or intervene runtime reliability risk, leaving them fragile when compromised agents vary across queries. We study stochastic adversarial topology adaptation, where compromised agents shift across graph positions and rare cascades dominate the high-loss tail. To address this setting, we propose RIGOR, a topology-adaptation framework that couples runtime risk inference with tail-aware training. RIGOR combines Risk-Gated Topology Construction (R-GTC), which probes agent behavior, estimates per-agent risk, and edits the graph through soft quarantine and edge masking, with CVaR-Guided Robust Optimization (C-GRO), which trains the editor on high-loss episodes rather than average gain alone. Across six benchmarks and two backbones, RIGOR consistently outperforms existing MAS baselines under query-varying misleading-message attacks, achieving an average 21.09% accuracy improvement over the strongest baseline on Qwen3.5-35B-A3B. Tail-subset and position-wise analyses further show that these gains stem from improved robustness against worst-case cascades and dynamically changing compromised-agent positions. The code is available for anonymous access at https://anonymous.4open.science/r/RIGOR-0D7D.
Abstract: Large language models increasingly need to accumulate and reuse historical information in long-term assistants and agent systems. Simply expanding the context window is costly and often fails to ensure effective context utilization. We propose \delta-Mem, a lightweight memory mechanism that augments a frozen full-attention backbone with a compact online state of associative memory. \delta-Mem compresses past information into a fixed-size state matrix updated by delta-rule learning, and uses its readout to generate low-rank corrections to the backbone's attention computation during generation. With only an 8×8 online memory state, \delta-Mem improves the average score to 1.10× that of the frozen backbone and 1.15× that of the strongest non-\delta-Mem memory baseline. It achieves larger gains on memory-heavy benchmarks, reaching 1.31× on MemoryAgentBench and 1.20× on LoCoMo, while largely preserving general capabilities. These results show that effective memory can be realized through a compact online state directly coupled with attention computation, without full fine-tuning, backbone replacement, or explicit context extension.
Abstract: We study mathematical programs with equilibrium constraints, in which a leader knows their own cost function, but lacks a model of the followers' response. Instead, the leader can only query this response at specific points. While this setting precludes the use of gradient-based methods, existing zeroth-order approaches treat the composed objective \emphentirely as a black box, deploying zeroth-order tools across both the leader and follower. Such approaches are inefficient, as they discard information the leader already possesses about their own cost function. In this work we instead propose to deploy zeroth-order tools only where they are truly needed: to handle the unknown, non-smooth followers' response. Specifically, we first propose \algpzos, an algorithm that combines exact partial gradients of the leader's cost with zeroth-order Jacobian estimates of the followers' response in a chain-rule-inspired manner, and establish that it achieves a strictly lower variance bound than the black-box baseline. Second, we introduce the partial Goldstein subdifferential, a stationarity notion tailored to this composite structure, and prove convergence of our algorithm to both standard and partial Goldstein stationary points. Finally, we validate our method on two application domains -- toll optimization in routing games and defense-attack investment in security games -- demonstrating consistent improvements over black-box baselines in convergence speed, objective value, and estimator variance, with robust performance even under few queries per iteration.
Abstract: Synthetic tabular data generation has attracted growing attention due to its importance for data augmentation, foundation models, and privacy. However, real-world tabular datasets increasingly contain free-form text fields (e.g., reviews or clinical notes) alongside structured numerical and categorical attributes. Generating such heterogeneous tables with joint modeling of different modalities remains challenging. Existing approaches broadly fall into two categories: diffusion-based methods and LLM-based methods. Diffusion models can capture complex dependencies over numerical and categorical features in continuous or discrete spaces, but extending them to open-ended text is nontrivial and often leads to degraded text quality. In contrast, LLM-based generators naturally produce fluent text, yet their discrete tokenization can distort precise or wide-range numerical values, hindering accurate modeling of both numbers and language. In this work, we propose TabDLM, a unified framework for free-form tabular data generation via a joint numerical--language diffusion model built on masked diffusion language models (MDLMs). TabDLM models textual and categorical features through masked diffusion, while modeling numerical features with a continuous diffusion process through learned specialized numeric tokens embedding; bidirectional attention then captures cross-modality interactions within a single model. Extensive experiments on diverse benchmarks demonstrate the effectiveness of TabDLM compared to strong diffusion- and LLM-based baselines.
PaperID: 4816, Poster
Abstract: Accurate time series forecasting with exogenous inputs is critical across domains, including energy and retail, yet modern deep learning models often overfit to correlations that do not generalize under shifts in these inputs. We propose SLiDE, a Koopman-inspired architecture for forecasting with exogenous inputs that imposes structure on latent temporal dynamics. SLiDE approximates nonlinear system evolution using a learned linear operator in a latent space, combining (i) a history encoder that reconstructs the current latent state from past targets and inputs with (ii) a shared linear rollout driven by future exogenous variables. By enforcing a common linear recurrence across encoding and prediction, SLiDE aligns how past information is encoded with how future states evolve, reducing parameterization and preventing reliance on spurious correlations in the input history. Empirically, SLiDE achieves state-of-the-art accuracy on real-world benchmarks, including electricity price forecasting and retail demand, while maintaining a computational footprint comparable to lightweight MLP architectures. SLiDE is especially effective at generalizing to shifted exogenous inputs, reducing MSE by 20% on electricity price samples with out-of-range exogenous values and MAE by 30% in a synthetic transfer setting where the same dynamical system is driven by an unseen exogenous process. Ablations confirm the importance of both the structured history encoder and the linear latent rollout, suggesting that structured latent linear dynamics provide a useful inductive bias for forecasting with exogenous inputs.
PaperID: 4817, Poster
Authors: Duoyou Chen, Yaxue Guo, Can Zhang, Cheng Chen
Abstract: Recent pixel-space generative models avoid the reconstruction bottleneck of latent diffusion, but direct denoising over high-dimensional pixel tokens remains computationally demanding. This paper studies a recursive alternative for efficient pixel-space diffusion. We propose RePiD, a hierarchical patch-level diffusion framework that recursively partitions an image and performs the core denoising computation on lower-dimensional pixel groups instead of repeatedly operating on full-resolution images. To preserve spatial coherence across patch-wise processing, RePiD introduces a neighborhood embedding module that conditions each patch on its surrounding regions. Experiments on class-to-image generation and unpaired image-to-image translation show that RePiD provides a favorable quality--efficiency trade-off for direct pixel-space generation. These results suggest that recursive patch-level denoising is a practical design direction for efficient pixel-space generative modeling. Our code is publicly available at \textcolor[HTML]387FB9https://anonymous.4open.science/r/Recursive-Pixel-based-Diffusion-Model-1735.
Abstract: Recent years have witnessed growing interest in applying Large Reasoning Models (LRMs) to Machine Translation (MT). While most approaches adopt a "pre-thinking" paradigm and benefit from explicit reasoning trajectories, they suffer from substantial inference cost and latency. To address these limitations, we propose ReflectMT, a two-stage reflection internalization framework for machine translation that employs a "post-thinking" paradigm. Our approach develops the model's "translate–reflect–refine" capability through reinforcement learning. In the first stage, we cultivate the model's capacity for high-quality reflection and refinement, thereby enhancing its semantic comprehension and task-specific knowledge. In the second stage, we train the model to internalize the knowledge acquired during reflection. As a result, during inference, ReflectMT operates in a direct translation mode, producing high-quality translations on the first attempt without any explicit reasoning steps. Experimental results on benchmarks such as WMT24 demonstrate that our model’s first-pass translations during inference outperform multi-step reasoning LRMs (e.g., DeepSeek-R1) in both automatic metrics and GPT-based evaluation, achieving a 2.16-point improvement in GPT-based translation quality evaluation while reducing token consumption by 94.33%.
PaperID: 4819, Poster
Abstract: Secure inference enables clients to use cloud neural networks without revealing private inputs or intermediate values. Homomorphic encryption (HE) and secure multi-party computation (MPC) provide complementary cryptographic mechanisms for this goal, and recent hybrid HE-MPC secure inference systems have shown better performance than using either primitive alone. However, existing hybrid secure inference systems typically rely on fixed execution strategies that are brittle across deployments. The best hybrid plan depends on both the target execution environment and the evolving cryptographic state of the computation, making manual or static mappings insufficient. We present \textscCipherFlow, a hardware-aware compiler for hybrid HE-MPC secure inference. Given a plaintext neural-network graph and a target deployment profile, \textscCipherFlow automatically generates a deployment-specialized, latency-optimized secure execution program. It profiles deployment-specific primitive costs and performs state-aware graph optimization to jointly decide operator placement, cross-domain conversion, and cryptographic maintenance actions. We implement \textscCipherFlow on real HE and MPC backends and evaluate end-to-end BERT-base and ViT inference across diverse CPU/GPU and network settings. \textscCipherFlow achieves up to 16.5× speedup on BERT-base and over 22× on ViT over state-of-the-art static hybrid baselines. These results show that compiler support can make hybrid cryptographic inference portable across heterogeneous deployments, turning manual secure-inference engineering into an automatic deployment-specific optimization process.
PaperID: 4820, Poster
Authors: Zhuoyuan Yu, Yuxing Long, Du Wang, Junyang Wang, Hongwei Fan, Hao Dong
Abstract: Deploying a language-guided navigation policy in a new building usually means watching it fail. Public-benchmark VLAs degrade scene by scene as layout, appearance, and language drift, and the standard fix of teleoperating fresh trajectories on site does not scale. We replace human demonstration with a single brief iPhone scan. Scan2Instruction compiles that scan into a complete training environment, automatically synthesizing executable navigation episodes with grounded instructions. PRPT-KS then adapts the policy through a closed-form Pareto utility read directly off the reconstructed scene, with no learned critic and no human labels. Across five real environments, the two designs together drive our RealNav to an 89% real-robot success rate, almost doubling the strongest VLA baseline at 48%, and still achieve 87% in unscanned regions of the same buildings. Our work turns demonstration-free scene adaptation for language-guided navigation from an aspiration into a working reality. All code and model will be released open-source upon acceptance.
PaperID: 4821, Poster
Abstract: Large agentic models can now operate in complex digital games by pretraining on web-scale data, but gameplay trajectories often contain highly redundant actions. Single-action agents can exploit this redundancy spuriously by copying recent actions, leading to causal confusion. Fixed-length action chunks reduce this ambiguity by predicting multiple future actions at once, but their open-loop execution accumulates errors because the agent cannot adapt to new observations during the chunk. The agent should therefore learn when to act for longer without feedback and when to stop early for a new observation. We introduce AdaMAP (Adaptive Multi-Action Prediction), an adaptive multi-action prediction framework that directly generates variable-horizon action chunks, interleaves actions with grounded visual dreams that provide verifiable future constraints, and optimizes the resulting policy with offline imitation learning followed by online MAP-PPO. Across Minecraft and VizDoom, AdaMAP attains the best or tied-best result in five of six task categories and improves average Minecraft success by 3.56 percentage points over the fixed-horizon RL baseline. The learned horizons vary systematically across tasks and games, suggesting that adaptive multi-action prediction discovers useful temporal abstractions.
PaperID: 4822, Poster
Abstract: AI-generated image detectors often generalize poorly to unseen generators because the classifier latches onto generator-specific feature directions that are predictive only on the training set. We show that simply projecting away generator-discriminative LDA directions fails: these directions also carry genuine real-vs-fake signal, so post-hoc erasure removes evidence along with shortcuts. We introduce StaR, a training-time regularizer that scores each LDA direction by leave-one-generator-out stability and penalizes the classifier only on unstable directions; stable directions are left available for classification. On ResNet-50, StaR reduces the across-seed standard deviation of the four-OOD average AUC by 24x relative to ERM (F-significant on every OOD) while preserving in-distribution-like performance; on CLIP ViT-L/14, StaR improves the four-OOD average AUC by 1.7 points over matched CLIP-ERM and 4.5 over EFFORT, with the largest gains on platform-shift benchmarks. Mechanistic analyses show that StaR rotates the classifier away from unstable generator directions while amplifying generator-discriminative structure in the features---a rotational, not erasive, effect that post-hoc projection cannot reproduce. We additionally find that the conventional
Authors:
Zhi Qiu, Jiazheng Sun, Chenxiao Xia, Jun Zheng, Xin PengAbstract: While Large Language Models demonstrate remarkable proficiency in high-level semantic planning, they remain limited in handling fine-grained, low-level web component manipulations. To address this limitation, extensive research has focused on enhancing model grounding capabilities through techniques such as Reinforcement Learning. However, rather than compelling agents to adapt to human-centric interfaces, we propose constructing interaction interfaces specifically optimized for agents. This paper introduces Component Interface for Agent (CI4A), a semantic encapsulation mechanism that abstracts the complex interaction logic of UI components into a set of unified tool primitives accessible to agents. We implemented CI4A within Ant Design, an industrial-grade front-end framework, covering 23 categories of commonly used UI components. Furthermore, we developed a hybrid agent featuring an action space that dynamically updates according to the page state, enabling flexible invocation of available CI4A tools. Leveraging the CI4A-integrated Ant Design, we refactored and upgraded the WebArena benchmark to evaluate existing SoTA methods. Experimental results demonstrate that the CI4A-based agent significantly outperforms existing approaches, achieving a new SoTA task success rate of 86.3%, alongside substantial improvements in execution efficiency.
Abstract: Post-training quantization (PTQ) enables efficient deployment of large language models by mapping pretrained weights to low-bit formats without retraining, typically using a small calibration set to minimize a layer-wise calibration objective. However, this sequential procedure induces a mismatch: errors from earlier quantized layers alter the inputs received by later layers, causing the activations to deviate from those of the full-precision model. Recent approaches introduce mismatch-aware calibration objectives to compensate for this effect, but leave open how much of the observed mismatch should shift each layer's calibration target. Fully applying this correction can overfit limited calibration data, while scaling the mismatch correction with a fixed coefficient ignores varying reliability of mismatch estimates across layers. To address these limitations, we propose CoreQ, a learning-free PTQ framework that applies a closed-form coefficient for mismatch correction derived from a geometric decomposition of the mismatch. The resulting coefficient adapts the correction across layers, reduces overfitting to finite calibration data, and requires no hyperparameter tuning. Given the corrected target, CoreQ minimizes the induced triangular least-squares objective with an efficient greedy successive-rounding solver and a bounded beam-search extension, K-CoreQ, that trades modest additional compute for improved performance. Across multiple LLM families, scales, bit-widths, and quantization settings, CoreQ improves perplexity and downstream accuracy over strong PTQ baselines.
PaperID: 4825, Poster
Authors:
Peyman Gholami, Shivam Chandhok, Saber Malekmohammadi, Shayan Shams, Robert Xiao, Leonid SigalAbstract: Recent advances in reinforcement learning have made verifiable rewards a central post-training paradigm for improving multimodal reasoning. However, methods such as GRPO remain data- and compute-inefficient, often spending rollout and annotation budget on prompts that provide little learning signal. We propose DISCOVER, an online data-discovery framework for budgeted multimodal GRPO that enables vision-language models to dynamically identify prompts that are most useful for the current policy. DISCOVER first organizes the multimodal sample pool into semantic clusters and queries a small set of density-diverse representative samples to obtain initial reward observations. During training, it estimates the utility of unannotated samples by transferring sparse reward feedback from queried samples. Since sample utility changes as the policy evolves, DISCOVER dynamically updates these estimates with recency-weighted rollout observations. It then selects only high-utility prompts for annotation using a policy-conditioned score that accounts for reward variance and difficulty, and continues utility-based replay once the discovery budget is exhausted. Experiments on ViRL39K and VLAA-Thinking with Qwen3-VL-2B show that DISCOVER substantially improves multimodal mathematical reasoning under a strict 10% unique-sample budget, outperforming budget-matched random selection, hard-example sampling, full-pool GRPO under the same training-step budget, and online selection baselines. Ablations further validate the importance of semantic coverage, online reward transfer, and stable discovery cadence, establishing DISCOVER as a practical framework for data-efficient multimodal RL post-training.
PaperID: 4826, Poster
Authors:
Yongwei Jiang, Peilin Li, anfu deng, XiaoyuChen, Jun YinAbstract: Biomicroscopy imaging involves the complex coordination of multiple procedures, such as focusing, illumination, and staining. Traditional unitary methods struggle with dynamical environments, while hybrid models pre-trained on natural images tend to introduce artifacts for higher perceptual metrics. Notably, biological imaging is not a one-off task. It is an iterative process similar to human expert’s observation, verification and adjustment. They can benefit from history, remembering which processing strategies are suitable and can also identify issues such as artifacts and mis-staining from verification. To this end, we propose BioMicroAgents, the first multi-agent framework with co-evolutionary capabilities designed to simulate and optimize whole biomicro imaging process including brightening, denoising, deblur, super-resolution, and virtual staining. Through reinforcement learning, the model can reflect based on historical experience to reduce artifacts and mis-staining, while enhancing edge and structural features to improve the accuracy of downstream tasks. Specifically, BioMicroAgents includes: (i) Multi-modal Analytical Agents, which utilize multimodal RAG to provide template images and biomicro knowledge, and employ the memory mechanism to analyze optimal processing paths; (ii) We propose three biomicro fidelity metrics and utilize three agent groups (MicroNavigator, MicroProcessor, and MicroAuthenticator) to perform rigorous analysis, processing, and verification. (iii) Co-evolutionary ability, which achieves dynamic policy and self-correction through interactive feedback strategies and memory mechanisms, allowing the agent to evolve from history and other agents. We test model across 12 biomicro tasks, BioMicroAgents achieved improvements in perceptual quality while adhering to strict fidelity in sub-cellular features. It also enhances the accuracy of downstream tasks such as segmentation and cell counting, showcasing the potential of general-purpose visual agents in bioscience.
PaperID: 4827, Poster
Abstract: Visual reprogramming (VR) efficiently repurposes pre-trained vision-language models for new image classification tasks by adding trainable patterns to inputs without modifying the backbone. Nevertheless, existing VR methods heavily depend on labeled data and thus suffer substantial performance degradation under limited supervision (i.e., scarce labeled images). We introduce semi-supervised learning (SSL) into VR to exploit abundant unlabeled images and enhance data efficiency. However, directly applying generic SSL techniques often amplifies biases in unlabeled posteriors and destabilizes VR training. To address these challenges, we propose Dual-attribute Consistency Regularization to ensure consistency across attribute prompts. Extensive experiments show that CURE consistently outperforms both existing VR methods and direct SSL extensions, improving average classification accuracy by percentage points across twelve widely-used benchmarks for limited supervision tasks. Our code is available at https://anonymous.4open.science/r/CURE-93E2/.
PaperID: 4828, Poster
Abstract: Echo State Networks (ESNs) are designed to be \emphstable—the echo state property makes the reservoir state a well-defined function of the input history—but it remains unclear how \emphgeometrically rich the reachable reservoir states are and what controls that richness. We give a partition-free, fractal-geometric theory of contractive reservoirs by viewing a quantized-input ESN as a contractive random dynamical system and, equivalently, an iterated function system (IFS). This identifies the reachable set as the IFS attractor and the stationary reservoir-state law under i.i.d.\ inputs as the unique invariant IFS measure. We prove entropy–contraction (entropy–Lyapunov) bounds on the intrinsic dimension of this measure, and in a conformal/similarity regime with standard separation conditions we obtain sharp closed-form dimension laws of the form “input entropy divided by contraction.” These results yield a quantitative \emphcriticality principle: weakening contraction drives a transition from low-dimensional to full-dimensional state representations at an explicit threshold set by input entropy and contraction rates, providing practical design guidance for tuning leak/spectral radius and input scaling.
Authors:
Zeyuan Hu, Noah Brenowitz, Akshay Subramaniam, Jaideep Pathak, Tao Ge, Mohammad S Abbas, Suman Ravuri, Karthik Kashinath, Noel Keen, Naser Mahfouz, Peter Caldwell, Mike PritchardAbstract: Kilometer-scale convection shapes precipitation extremes, tropical organization, and cloud feedbacks, but most global atmospheric models approximate these processes at 25--100\,km resolution. Global storm-resolving physics models resolve convective systems explicitly, but at a cost---roughly one MWh per simulated day on exascale supercomputers---that limits their use for long-duration atmospheric simulation. We introduce STRATA (Storm-resolving Tile-based autoRegressive Atmosphere Transformer Architecture), the first autoregressive AI emulator for global storm-resolving atmospheric dynamics. STRATA is trained on the highest-resolution atmospheric dataset yet used for global AI emulation: 17 days of output from the SCREAM physics model at 4.9-km resolution (~25 million grid cells) sampled every 10 minutes. Our central premise is that since on 10-minute timescales atmospheric dynamics are predominantly local, training on small spatial tiles trades scarce global temporal samples for abundant local spatial samples and enables global rollout via overlapping-tile blending. STRATA combines 3D patch embedding and local 3D neighborhood attention for tractable modeling of high-resolution atmospheric tiles, a novel Stereographic Rotary Position Embedding (StereoRoPE) for grid-invariant positional encoding, and a pixel-space de-aliasing decoder that suppresses patch-scale rollout artifacts. An iso-FLOP scaling study reveals that km-scale emulation requires ~10× more FLOPs per horizontal grid point than coarse-resolution AI weather models, consistent with the higher information density of convective-scale dynamics. Despite training on only 17 days of SCREAM output, STRATA produces stable 24-hour global rollouts with realistic km-scale dynamics across diverse weather regimes, though large-scale biases develop with lead time. STRATA achieves 48 simulation days per megawatt-hour---about 50 times better energy efficiency than the underlying SCREAM physics model---and 741 simulated days per wall-clock day at 512 H100 GPUs. Code and dataset will be publicly released.
PaperID: 4830, Poster
Abstract: Reasoning segmentation requires localizing objects from complex, implicit textual queries. Existing methods either rely on the MLLM to reason implicitly within hidden representations, or produce explicit but purely textual chain-of-thought that terminates in a coarse bounding box. In both cases, the reasoning process is not grounded in visual evidence at each step, leaving intermediate errors unverifiable, which can propagate to the final prediction. We propose Reasoning with Visual Anchors (ReVA), a framework where each chain-of-thought step is explicitly grounded to a specific image region. ReVA introduces anchor tokens into intermediate reasoning steps, each directly aligned with region-level features in the image encoder's latent space. As the model reasons step by step, each anchor attends to its corresponding image region, grounding inference in concrete visual evidence. To supervise the full reasoning chain, we further require each anchor to correctly transition to the next along the spatial relation described in the text. This progressively guides the model toward the target and produces the correct segmentation mask. To support training, we construct a large-scale Chain-of-Anchor dataset of over 245K quality-filtered samples from the RefCOCO series, each annotated with a spatially grounded chain-of-thought containing explicit anchor regions. The dataset is constructed by leveraging a strong MLLM to generate reasoning traces, followed by a three-stage verification pipeline ensuring spatial accuracy and reasoning quality. Experiments on reasoning and referring segmentation benchmarks demonstrate that ReVA achieves state-of-the-art performance (e.g., +5.4 gIoU and +6.7 cIoU on the ReasonSeg test set), with ablations showing that reasoning in visual evidence is key to the improvement.
Authors:
Yunus Sevinchan, Juan Nathaniel, Kai Ueltzhöffer, Carla Roesch, Tobias Weber, Vaios Laschos, Hang Fan, Gregor Ramien, Johannes Haux, Pierre Gentine, Benjamin HerdeanuAbstract: Critical transitions – abrupt, often irreversible changes in system dynamics – arise across human and natural systems, often with catastrophic consequences. Real-world observations of such shifts remain scarce, preventing the development of reliable early warning systems. Conventional statistical and spectral indicators, such as increasing variance, tend to fail under realistic conditions of limited data and correlated noise, whereas existing deep learning classifiers do not extrapolate beyond their training data distribution. In this work, we introduce TipPFN, an in-context learning (ICL) framework that uses a prior-data fitted network to infer a system's proximity to a critical transition. Trained on our novel synthetic data generator, which is based on canonical bifurcation scenarios coupled to diverse, randomized stochastic dynamics, TipPFN flexibly capitalizes on contexts of various sizes, complexity and dimensionalities. We demonstrate robust, state-of-the-art early detection of critical transitions in previously unseen tipping regimes, sim-to-real examples, and real-world observations in both ICL and zero-shot settings.
PaperID: 4832, Poster
Abstract: High-resolution weather prediction is important for resolving local atmospheric patterns missed by coarse global forecasts. Large pretrained weather backbones have improved global forecasting, but adapting them to produce future high-resolution fields from recent coarse states remains challenging and costly. We formulate this setting as Predictive Spatiotemporal Weather Downscaling (PSWD), where the target is a future high-resolution trajectory rather than a same-time refined field. To make this structure explicit, we propose a framework that decomposes prediction into spatial and temporal stages using a shared pretrained backbone, fixed numerical scaffolds, and lightweight residual heads. For backbone adaptation, we introduce PROLA (PRincipal-Orthogonal Low-rank Adaptation), which splits a fixed low-rank budget between pretrained principal directions and their orthogonal complement. PROLA further rescales rank-1 optimizer updates using gradient signals, improving adaptation without increasing trainable rank. Experiments in the PSWD setting show that PROLA outperforms representative low-rank adaptation baselines under matched trainable-parameter budgets. These results support PSWD as a challenging setting for weather backbone adaptation and PROLA as an effective method for this setting.
Abstract: Autoregressive (AR) models can generate high-quality low-poly meshes from point clouds, but they still operate in an all-or-nothing manner: when a local region is unsatisfactory, the entire mesh must be regenerated, wasting computation and destroying satisfactory mesh structure elsewhere. We introduce MeshFIM, a Fill-in-the-Middle (FIM) framework that regenerates a target region of a low-poly mesh conditioned on the surrounding context. MeshFIM addresses three mesh-specific challenges: enforcing exact attachment along the exposed boundary, preserving topological order in the context, and suppressing overflow beyond the intended region. It does so with five complementary design choices: boundary vertex markers, context positional embeddings, expanded context width, context augmentation, and a low-poly geometry encoder whose gated subtraction mechanism focuses generation on the missing region by leveraging the difference between the reference surface and the existing mesh. Detailed ablation studies are presented to show the effectiveness of every introduced component. Based on MeshFIM, we demonstrate two applications: interactive brush-based editing and automatic defect repair on low-poly mesh (see Figure 1). Last but not least, experiments show that MeshFIM outperforms a range of baselines in mesh refinement, mesh repair and whole mesh generation plus stitch-back scheme.
PaperID: 4834, Poster
Authors: Radu Beche, Raoul de Charette
Abstract: 3DGS enables real-time novel view synthesis, but practical deployment under varying compute budgets requires a single Gaussian set that remains effective when truncated to a prefix of primitives. This raises how capacity should be allocated as the budget decreases.We observe two findings. First, image complexity correlates with reconstruction error in both image and 3D Gaussian space. Second, this structure degrades under strong budget reduction, indicating suboptimal allocation.These observations motivate using image complexity as an optimization signal. We propose a continuous level-of-detail method based on 3DGS-MCMC that uses it for (i) relocation toward difficult regions during training and (ii) importance-based retention of simple and detailed regions at inference. A variance constraint on residuals between full- and reduced-budget renderings further enforces uniform error under compression. Our method yields smoother LOD degradation than prior work, with largest gains in low-budget regimes. Code will be released.
Abstract: We develop semiparametrically efficient inference for kernel measures of noise heterogeneity in additive noise models. In many applications, the regression function is estimated using flexible machine learning methods. Downstream procedures based on the resulting residuals can then inherit first-stage bias: regression error may induce spurious dependence between covariates and residuals, invalidating the assumptions needed for standard analysis. We construct a novel Hilbert-valued one-step estimator of the kernel covariance operator between covariates and residuals. Our estimator yields bootstrap-calibrated tests for residual independence and goodness of fit in additive noise models, while also providing asymptotically efficient confidence intervals for the kernel dependence measure under noise heterogeneity. The framework extends to settings with additional covariates, enabling inference on distributional heterogeneity of residual noise across treatment groups. Simulations show improved calibration and power relative to naive plug-in residual methods.
Abstract: Multi-path speculative decoding accelerates lossless sampling from a target model by using a cheaper draft model to generate a draft tree of tokens, and then applies a verification algorithm that accepts a subset of these. While prior work has proposed various verification algorithms for i.i.d rollouts, their relative performance under matched settings remains unclear. In this work, we firstly present a systematic evaluation of verification strategies across model families, tasks, and sampling regimes, and find that Traversal Verification dominates consistently, with OT-based methods lagging far behind. Our analysis uncovers that this occurs because OT-based methods achieve high multi-token acceptance near the root of the draft tree, while multi-token gains are most impactful deeper in the draft tree, where draft and target distributions diverge. Based on this insight, we propose delayed tree expansion which drafts a partial single path, delaying the i.i.d. branching point. We show that delayed tree expansion preserves the target distribution and improves on root-node i.i.d rollouts. Further, we develop a dynamic neural selector that estimates the expected block efficiency of OT-based verification methods from draft and target features, enabling context-dependent expansion decisions. Our neural selector allows OT-based methods like SpecInfer to outperform Traversal Verification for the first time, with 5% higher average throughput across a wide range of models, datasets, and sampling settings. Finally, we extend our approach to Traversal Verification to improve its average throughput by 26%.
PaperID: 4837, Poster
Abstract: Synthetic time series generation is critical for data augmentation, privacy-preserving sharing, and simulation. Autoregressive models extend to arbitrary horizons but suffer compounding errors, while diffusion models achieve high fidelity through bidirectional refinement only at fixed lengths. We propose Horizon Diffusion, a unified framework that resolves this tension by decomposing generation into variable-length horizon blocks produced autoregressively across blocks yet jointly denoised within each block. Two innovations make this hybrid practical: Horizon-Causal Attention, a structured mask that simultaneously preserves cross-block causality and within-block bidirectional refinement, and the Block Size Curriculum, a deterministic Warmup--Stable--Decay schedule that fluidly transitions training from full-sequence diffusion to block-wise autoregressive generation. Our method consistently outperforms strong baselines on fidelity, diversity, and downstream utility, and remains stable when sequence length scales by 4× beyond training from a single trained model. Code is available at Supplementary Material.
PaperID: 4838, Poster
Abstract: While label-conditional diffusion models exhibit remarkable generative capabilities recently, their success heavily relies on massive, cleanly labeled datasets. In practice, categorical supervision is rarely perfect: it is often corrupted by noise, clouded by ambiguity, or partially missing. Training directly on such weak annotations severely degrades generation quality, and existing robust methods offer only fragmented solutions that depend on scarce auxiliary priors. To overcome these limitations, we propose DELTA, a unified framework that robustly trains diffusion models across all three weak annotation types without any external priors. By treating the unknown true label as a latent variable, DELTA optimizes a principled variational objective that jointly recovers the clean data distribution and infers the true label posteriors. To make this joint training computationally tractable, we further introduce a median-centered timestep sampling strategy that efficiently concentrates evaluation where the diffusion process is most informative. Extensive experiments demonstrate that DELTA produces high-fidelity, class-consistent samples across diverse weak supervision scenarios, outperforming specialized baselines while demanding strictly less prior knowledge.
PaperID: 4839, Poster
Authors: Jianting Pan, Jian Li, Sirong Dai, Ming Yan
Abstract: Computing exact Optimal Transport (OT) distances for large-scale datasets is computationally prohibitive. While entropy-regularized alternatives offer speed, they sacrifice precision and frequently suffer from numerical instability in high-accuracy regimes. To address these limitations, we propose the Inexact Bregman Sparse Newton (IBSN) method, which efficiently solves the exact OT problems. Our approach utilizes a Bregman proximal point framework through a sequence of semi-dual subproblems. By solving these subproblems inexactly, we significantly reduce per-iteration complexity while maintaining a theoretical guarantee of convergence to the true optimal plan. To further accelerate the algorithm, we develop a sparse Newton-type solver for the subproblem and employ a Hessian sparsification strategy that drastically lowers memory and time costs without sacrificing accuracy. We provide rigorous theoretical guarantees for the global convergence of the algorithm. Extensive experiments demonstrate that IBSN consistently outperforms state-of-the-art methods in both computational speed and solution precision.
PaperID: 4840, Poster
Abstract: Monocular depth estimation (MDE) has advanced rapidly with the emergence of foundation models such as Depth Anything. Their transformer-based architectures provide strong generalization across diverse scenes and domains, but also incur high computational and memory cost, making efficient deployment challenging. Post-Training Quantization (PTQ) provides an efficient and practical solution for model compression, yet low-bit PTQ remains challenging for MDE models. When applying PTQ to Depth Anything, we identify two key challenges: (i) independently minimizing the quantization error of the query and key projections fails to preserve the attention maps induced by their interaction, and (ii) quantization errors accumulate across layers, resulting in intermediate feature distribution shifts. To address these issues, we propose SPA-Q, a Structure-Preserving Adaptive PTQ framework for MDE, consisting of two main components. First, Attention-Preserving Calibration (APC) determines query and key quantization parameters by matching the full-precision attention distribution. Second, Channel-Wise Distribution Alignment (CWDA) learns channel-wise affine transformations to mitigate quantization-induced distribution shifts, and the learned parameters are absorbed into the weights after training. Experimental results show that SPA-Q consistently outperforms existing PTQ methods under 4-bit quantization, achieving an average 25.9% reduction in AbsRel and a 17.3% improvement in \delta_1 across NYUv2 and KITTI datasets.
PaperID: 4841, Poster
Abstract: Despite progress in instruction-based video editing, unimodal textual instructions inherently struggle to convey fine-grained textures and complex dynamics. To bridge this perceptual gap, we propose Visual In-context Editing, a new paradigm elevating video editing from textual instructions to multi-modal visual guidance encompassing single image, image pair, and video pair. To facilitate this paradigm, we curate , the first large-scale dataset for visual in-context video editing. We develop an automated pipeline to generate 400K high-quality samples across ten task types, ensuring superior visual fidelity and semantic consistency through multi-dimensional filtering. Leveraging this foundation, we introduce , a unified framework to bridge visual and textual contexts. To adaptively extract editing semantics from heterogeneous references, we design , which produces modality-specific semantic tokens from visual references. These tokens are then synergistically integrated with textual instructions through , enabling the generation process to benefit from both visual and textual signals. Extensive evaluations on VicEdit-Bench demonstrate that VicEdit achieves state-of-the-art performance across both basic instruction editing and visual in-context editing tasks, establishing visual in-context learning as a powerful and controllable paradigm for video editing.
PaperID: 4842, Poster
Authors: Junyuan Deng, Xinyi Wu, Zhenyao Wu, Song Wang, Yu Wang, Congchao Zhu
Abstract: Generative-prior-based super-resolution (SR) methods leverage implicit knowledge learned from large-scale models to enhance low-resolution inputs. However, such priors are inherently stochastic and optimized for diverse generation, conflicting with the deterministic, high-fidelity reconstruction required for SR. It is also tied to the base model, costly to update, and prone to hallucinations as real-world data evolves. To address these limitations, we introduce WorldSR, a novel framework that augments SR with explicit, entity-aligned visual world knowledge. WorldSR automatically decomposes the input into semantic entities and performs entity-wise retrieval of structurally and semantically aligned images from a large-scale external corpus. For efficiency and reliability, we incorporate a memory mechanism to reuse previously retrieved results and a multi-step filtering process to refine candidate references. The selected references as world knowledge are then integrated through a novel multi-reference super-resolution model, which provides fine-grained constraints for detail synthesis and structural recovery. Experiments on real-world SR benchmarks show that WorldSR consistently improves structural fidelity and visual quality, outperforming existing generative super-resolution methods. Our results suggest that entity-level grounding in external visual world knowledge provides a simple and effective complement to model-bound generative priors.
PaperID: 4843, Poster
Authors:
Onur Selim Kilic, Afra Nawar, Cem O Yaldiz, Michael J Cho, Ahmet R Emirdagi, Demet Tangolar, Amirali Aghazadeh, Amit J Shah, Omer T InanAbstract: Paired translation between quasiperiodic physiological waveforms (i.e., recovering a target oscillatory signal from the source) is central to the interpretation of cardiovascular signals derived from wearables placed at different body locations. This source-to-target mapping in these problems carries inherent geometric structure: the phase wraps around the cycle and must be treated as a circular variable, the amplitude remains strictly positive, and the beat-to-beat alignment can drift unpredictably across cycles and subjects. While deep neural networks have been used for phase estimation and complex-valued signal modeling, prior work does not explicitly learn phase transport between paired signals. Consequently, neither endpoint-supervised regression nor the standard affine path used in flow matching accounts for this phase--amplitude structure. We introduce \emphcylindrical geodesic flow matching for paired cardiovascular waveform translation. We show that the standard affine path used in flow matching distorts intermediate amplitude and instantaneous frequency when interpolating between quasiperiodic signals; replacing it with a closed-form geodesic on the phase--amplitude cylinder eliminates these artifacts and converts each training pair into dense, geometry-consistent velocity supervision. On zero-shot photoplethysmography and limited-support seismocardiography adaptation benchmarks, our method consistently outperforms interpolation baselines and matches or exceeds direct supervised prediction, reducing Hilbert Transform, L_2, and Dynamic Time Warping distance by up to ~15% over the strongest competing baseline. These results suggest that bridge geometry is a critical inductive bias for flow matching on oscillatory signal translation.
PaperID: 4844, Poster
Abstract: Binary diffusion models typically require a massive number of function evaluations (NFEs) to generate high-quality samples, making practical inference computationally expensive. Reducing NFEs while preserving sample quality without relying on distillation or additional training remains a significant challenge. However, existing binary diffusion models sequentially define a discrete one-step forward path and subsequently derive the reverse posterior. In low-NFE scenarios that require cross-step sampling, these models incorrectly approximate the true multi-step likelihood using a single-step likelihood formulation, which severely degrades sample quality. To address this fundamental limitation and completely decouple the generative dynamics from fixed discrete time steps, we propose Bernoulli Flow Models (BFM). Rather than building upon sequential one-step Markov diffusion chains, we predefine a unified continuous global Bernoulli probability flow path between data distributions and pure noise, from which we derive analytical closed-form posterior transitions over arbitrary time intervals. Consequently, reducing the inference NFE in BFM is no longer an approximation of skipped discrete steps; it simply requires re-evaluating the analytical posterior on the new time intervals. This eliminates the structural training-inference mismatch inherent to discrete chains, yielding strictly self-consistent low-NFE sampling. Experimental results show that BFM is highly robust to aggressive NFE reduction. On the LSUN Churches 256x256 dataset, a 256-step-trained BFM achieves an FID of 9.31 when sampled with only 16 steps, whereas the state-of-the-art discrete baseline severely degrades to 204.10. BFM also remains competitive with both continuous and discrete generative baselines under standard full-step inference. Ultimately, these results establish BFM as a theoretically rigorous, self-consistent, and practically effective framework for fast binary data generation.
PaperID: 4845, Poster
Authors: Ryunosuke Hamada, Daisuke Hatano, Yuichi Takano
Abstract: Neural combinatorial optimization (NCO) solvers generalize poorly under distribution shift, yet robustness across heterogeneous problem distributions is essential for real-world deployment where no single client should be left behind. Group distributionally robust optimization (Group DRO) offers a principled remedy, but requires pooling data that often cannot be shared due to confidentiality or regulatory constraints. Federated learning removes data-sharing requirements; however, existing federated DRO methods cannot accommodate the policy-gradient estimators used in modern NCO. Specifically, prior analyses rely on Lipschitz losses with bounded variance. In contrast, the REINFORCE policy-gradient estimator multiplies stochastic costs by score functions, causing variance to scale with cost magnitude and violating these standard assumptions. We close this gap by replacing Lipschitz-loss conditions with a bounded-cost assumption. This allows us to control cost-dependent variance and establish the first federated Group DRO framework with O(1/\sqrtT) convergence guarantees. Our analysis applies to any REINFORCE-trained model under bounded costs, including NCO solvers as a special case. Experiments on capacitated vehicle routing---one of many applicable domains including scheduling and packing---demonstrate the necessity of federated training for minority-distribution clients. We show that while na\"ive Group DRO suffers from catastrophic weight collapse, our KL-regularized formulation avoids this collapse. Furthermore, it provides a formal worst-group robustness certificate with no empirical performance penalty compared to standard federated averaging.
Authors: Gianfranco Cortés, Xiaoda Qu, C. Vemuri
Abstract: Nonrigid registration is vital to medical image analysis but remains challenging for diffusion MRI (dMRI) due to its high-dimensional, spatio-angular dependence. We present a novel, geometric deep learning framework for \it model-free, nonrigid registration of raw dMRI data. The dMRI registration problem is formulated in the native spatio-angular acquisition space, which exhibits a natural symmetry to the group of 3D roto-translations, denoted by \mathrmSE(3). A by-product of this design choice is freedom from having to augment the data with roto-translated versions of itself. Our second novelty is the loss function formulation, based on the maximum mean discrepancy (MMD) loss used to compare two probability density functions. We apply this loss in the Fourier space, where it becomes the well-known weighted sum-of-squared differences (SSD) loss with the weights being the Fourier transform of a reproducing kernel Hilbert space (RKHS) kernel. Experimental results on HCP and OASIS-3 clinical-grade dMRI data demonstrate competitive performance compared to SOTA approaches, with the added advantage of bypassing the overhead for estimating derived representations. This work establishes a foundation for data-driven, geometry-aware dMRI registration directly in the acquisition space.
PaperID: 4847, Poster
Abstract: Agent-based frameworks grounded in vision-language models (VLMs) have emerged as a dominant paradigm for long video understanding. Yet, prevailing agents lack the core capacity for dynamic memory evolution, failing to transform ephemeral perceptions into a continuously growing, adaptive knowledge system and thus inducing inefficient redundant exploration during reasoning. To this end, we present MEMEVO, an online Memory-Evolved Video Agent framework centered on dynamic memory evolution that leverages accumulating multi-turn queries to actively drive progressive memory evolution, establishing a long-lived memory system for knowledge accumulation. Central to MEMEVO is a four-level Hierarchical Self-Evolved Memory, which constructs memory bottom-up via progressive abstraction, distilling transient perceptions into reusable, query-agnostic event nodes organized within a temporal-semantic graph, while realizing active memory evolution through event insertion, duplicate merging, and redundant pruning. To mitigate redundant retrieval in the evolved memory, we devise a State-Conditioned Graph Memory Retrieval mechanism that guides efficient navigation of the memory space, integrating State-Conditioned Graph Routing and Meta-Cognitive Stopping criteria to ensure sufficient evidence acquisition without re-exploration. Extensive benchmarking confirms that MEMEVO achieves state-of-the-art performance on reasoning-intensive tasks (\textcolor[HTML]00FF00+6.1% VideoMMMU, \textcolor[HTML]00FF00+4.8% LongVideoBench), while facilitating cross-query knowledge reuse and cross-model generalization.
Abstract: Three-dimensional molecular structure generation is typically performed at the level of individual atoms, yet molecular graph generation techniques often consider fragments as their structural units. Building on the advances in frame-based protein structure generation, we extend these fragmentation ideas to 3D, treating general molecules as sets of rigid-body motifs. Utilising this representation, we employ \mathrmSE(3)-equivariant generative modelling for de novo 3D molecule generation from rigid motifs. In our evaluations, we observe comparable or superior results to state-of-the-art across benchmarks, surpassing it in atom stability on GEOM-Drugs, while yielding a 2× to 10× reduction in generation steps and offering 3.5× compression in molecular representations compared to the standard atom-based methods.
PaperID: 4849, Poster
Authors: Evripidis Bampis, Bruno Escoffier, Dimitris Fotakis, Giorgos Mitropoulos, Michalis Xefteris
Abstract: We study learning-augmented algorithms for scheduling on multiple parallel machines under the objectives of minimizing the makespan or the total weighted completion time. We assume that the algorithm has access to noisy information (the prediction) about an optimal solution, expressed through a solution encoding that implicitly quantifies both the amount of information provided and the nature of noise (errors in the prediction). Our central question is to determine the maximum level of noise that still allows for efficiently recovering this optimal solution. We observe that information guiding greedy scheduling to optimality or revealing the optimal assignment of tasks to machines is very sensitive to noise, and prove that recovering an optimal solution from solution encodings capturing either is computationally hard, even for a small (non-constant) number of errors. On the positive side, we show that an order-based encoding, which leads to errors that are dispersed locally, is robust to noise. We present a dynamic programming framework that utilizes this encoding and solves optimally several classical scheduling problems, including makespan and total weighted completion time minimization on identical or unrelated machines.
Abstract: A recent work shows that Attention Transfer, which transfers only the attention patterns from a pre-trained teacher Vision Transformer (ViT) to a randomly initialized standard student ViT, is sufficient to recover the full benefit of the teacher's pre-trained weights. We revisit this finding on a comprehensive benchmark of 20 teachers from 11 well-known ViT families and reveal that Attention Transfer is not universally effective. While 7 families transfer successfully, 4 consistently fail, falling up to 5.1% below the from-scratch no-transfer baseline. Further results demonstrate that this failure is family-consistent across model sizes, and persists under extended training durations, different transfer datasets, and out-of-distribution evaluations. Controlled analyses then consistently localize the problem to the attention-routing channel, indicating that the key issue is not whether the student can match the teacher's attention patterns, but whether the matched patterns remain functional for the student. Crucially, we identify architectural mismatch between the pre-trained teacher and the standard student as the primary mechanism. By adding only the teacher's native architectural components to the student in a randomly initialized state, we completely reverse the failure for all 4 families. Notably, these components alone do not improve from-scratch training, confirming that they specifically unlock the usability of the teacher's attention. We further systematically show that this failure is not explained by the inadequate choice of transfer loss or by differences in pre-training recipes. Our findings refine the prevailing understanding of attention in ViT representations: attention is sufficient only when the student architecture matches the teacher.
PaperID: 4851, Poster
Abstract: Contact-rich manipulation requires precise control and complex physical interaction with the environment, posing significant challenges for robot policy learning. Vision-only policies struggle with such tasks, as contact states and force feedback are difficult to infer from vision alone. Existing methods mitigate this issue by incorporating force or tactile sensing. However, they often treat these signals as auxiliary observations rather than using them to model future interaction dynamics, leading to inefficient modality utilization and limited generalization. To address this, we propose ForceLDM, a force-aware latent dynamics model that enables the policy to anticipate future contact and motion dynamics in latent space and use them to guide action generation. Rather than predicting future observations in pixel space, which often requires large amounts of data and risk overfitting irrelevant details, ForceLDM learns a compact task-relevant representation through knowledge distillation. Specifically, a teacher network with privileged access to future force/torque and optical flow signals extracts future dynamic features, which are then distilled into a student network trained only on current observations. This enables the student to reason about upcoming contacts and physical dynamics at deployment without relying on future sensory inputs. To improve robustness and generalization, we further introduce a curriculum-based progressive noise injection strategy to mitigate over-reliance on future features. Experiments on five real-world contact-rich manipulation tasks demonstrate that ForceLDM significantly outperforms state-of-the-art baselines, generalizes effectively to novel objects, and remains robust to environmental perturbations.
Abstract: Current unified multimodal models typically rely on discrete visual tokenizers to bridge the modality gap. However, discretization inevitably discards fine-grained semantic information, leading to suboptimal performance in visual understanding tasks. Conversely, directly modeling continuous semantic representations (e.g., CLIP, SigLIP) poses significant challenges in high-dimensional generative modeling, resulting in slow convergence and training instability. To resolve this dilemma, we introduce UniCom, a unified framework that harmonizes multimodal understanding and generation via compressed continuous semantic representations. We empirically demonstrate that channel-wise compression is significantly more effective than spatial token reduction, retaining competitive VAE-free reconstruction fidelity while substantially improving generative convergence. Accordingly, we design an attention-based semantic compressor to distill dense features into a compact unified representation. A Transfusion-style predictor then generates these compressed latents with dense spatial correspondence. At scale, UniCom achieves competitive text-to-image generation and state-of-the-art performance on complex image editing without relying on VAE latents. The gains are especially clear on knowledge-intensive benchmarks.
PaperID: 4853, Poster
Abstract: Safety-aligned Large Language Models (LLMs) remain vulnerable to interventions during inference that redirect generation toward harmful outputs. Recent work attributes this to shallow safety, where alignment concentrates in the first few output tokens. We show that shallow safety is a special case of a broader inference-time vulnerability, in which short token injections at any generation step can substantially alter subsequent safety behavior. We also find that a model's alignment with refusal directions in its hidden states does not predict its robustness to such injection, revealing that internal state alone does not determine generation behavior under perturbation. To address this, we align models directly on generation trajectories constructed by simulating mid-sequence perturbation, and show that this improves robustness to mid-sequence injection and generalizes to attacks that exploit early-token generation. Our work argues that robust safety alignment requires training on the generation process itself, not only its outputs.
PaperID: 4854, Poster
Abstract: As large language models (LLMs) continue to scale, quantization has become a key technique for efficient deployment. However, multi-bit quantization employs non-differentiable rounding, hindering gradient optimization. Existing methods rely on heuristic surrogate gradients (e.g., STE), which work empirically but lack a unified theory. To address this challenge, we propose the Zeroth-Order Expectation Gradient (ZOE-Grad), a zeroth-order expectation view of quantization surrogate gradients. Specifically, we establish an equivalence between the expectations of zeroth-order gradient estimators and surrogate gradients, providing a unified zeroth-order interpretation of widely used surrogates. Under this framework, STE corresponds to a degenerate, discontinuous perturbation that ignores the local geometry of quantization boundaries, limiting its expressiveness. In contrast, continuous perturbation-based surrogate gradients capture boundary-local information and produce smoother, more structured gradients. We further establish bounded-error convergence guarantees for the quantized objective. Through simulations under multiple perturbation distributions, we verify that the proposed framework accurately captures the relationship between the expectations of zeroth-order gradient estimators and surrogate gradients. Furthermore, extensive experiments on OPT-1.3B, OPT-6.7B, LLaMA-2-7B, and Qwen3-8B show that continuous ZOE-Grad surrogates consistently outperform STE across diverse LLM architectures.
PaperID: 4855, Poster
Authors:
Zhihong Cui, Hengyu Liu, Michael A. Riegler, Guandong Xu, Amir Taherkordi, Tor SkeieAbstract: Vision-Language Models (VLMs) are increasingly adopted as end-to-end driving planners, yet they suffer from cross-city domain shift. Existing methods treat this shift as a single quantity addressed by uniform alignment or single-site editing. We refute this view on two counts: (i) domain shift in VLM trajectory planners is inherently structured across modality and layer, and (ii) a locate-then-repair recipe matched to this structure outperforms every uniform-objective baseline. We analyze domain shift via activation patching and uncover two regularities. (F1) Modality axis. Image and text tokens exhibit distinct shift patterns in each VLM. (F2) Layer axis. The layer-wise shift partition is architecture-determined and dataset-stable. These findings recast domain shift as a two-dimensional tensor indexed by modality (image vs. text tokens) and layer, motivating a simple principle: locate where shift occurs, and repair only there. We instantiate this principle in MoLaRx (Modality–Layer Repair), a locate-then-repair framework. The Causal Importance Score (CIS) estimates the domain shift as a modality–layer tensor and partitions the layers into functional zones. CIS-Guided Selective Repair (CISR) then applies a matched operator (Maximum Mean Discrepancy (MMD) / low-rank adaptation (LoRA) / freezing) to repair each zone, with the composition adapting automatically to each architecture. We evaluate MoLaRx on three architecturally distinct VLMs (LFM2-VL-1.6B, Qwen2-VL-2B, and PaliGemma2-3B) across two cross-city benchmarks (nuScenes and Argoverse 2). MoLaRx attains the lowest cross-city trajectory-error gap on every (VLM, dataset) cell, while using 50–60% fewer trainable parameters and 8.4–10.9 ms lower inference latency. In contrast, the strongest scalar baseline (MMD-only) collapses on every cell—direct evidence that no scalar objective can repair structured shift. Code: [anonymous.4open.science/r/MoLaRx-E3B4](https://anonymous.4open.science/r/MoLaRx-E3B4/).
Abstract: Many real-world heterogeneous graphs exhibit pronounced heterophily, where connected nodes often have dissimilar labels or play different semantic roles. In such settings, standard heterogeneous graph neural networks that aggregate messages along metapaths or meta-relations primarily based on feature similarity can propagate misleading information, since feature similarity may be misaligned with underlying relational semantics. In this paper, we propose , a semantics–structure decoupling framework for heterogeneous graph learning under heterophily. HeterSEED decouples representation learning into a heterogeneous semantic channel that captures type- and relation-aware local semantics and a structure-aware heterophily channel that separates homophilic and heterophilic neighborhoods via pseudo-label-guided partitioning and aggregates them using metapath-based structural weights. A node-level adaptive fusion mechanism then combines the two channels to produce context-dependent node representations. Theoretically, we establish that, on heterogeneous graphs under heterophily, HeterSEED is strictly more expressive than standard heterogeneous graph neural networks that rely primarily on feature similarity and provably reduces the prediction bias introduced by heterophilic neighbors. Experiments on five real-world heterogeneous graphs, including two large-scale networks at the million-node and hundred-million-edge scale, demonstrate that HeterSEED consistently outperforms representative heterogeneous graph neural networks and recent heterophily-aware baselines, especially in strongly heterophilic regimes.
Abstract: A significant hurdle for current LLMs is the execution of complex, multi-stage tasks. Group Relative Policy Optimization (GRPO) has been emerging as a leading choice, but its reliance on sparse outcome rewards severely limits credit assignment across intermediate steps. Existing remedies—running full rollouts to assign step-level advantages, calling external LLM judges at each step, or computing intrinsic rewards that require ground-truth answers at every evaluation—introduce significant costs or practical constraints. We hypothesize that internal correctness probing over LLM hidden states can be repurposed as a step-level reward signal, potentially addressing all of these limitations at once. However, existing probing research assumes clean inputs, and we first show that this assumption breaks down in multi-step settings: hidden-state probes degrade severely under prefix contamination—tracking coherence with the (possibly corrupted) prefix rather than grounded correctness—while attention-based features remain robust to contamination but underperform on clean prefixes. Building on this complementary relationship, we propose the Prefix-Aware Internal Reward (PAIR)—a two-stage model with a frozen hidden-state probe estimating belief-consistency and a lightweight attention-based head correcting it toward grounded correctness. Experimental results show that PAIR achieves the highest AUROC on contaminated trajectories while operating at negligible inference cost, enabling dense step-level reward signals for GRPO training without external model calls, ground-truth dependencies, or full-trajectory rollouts.
PaperID: 4858, Poster
Authors: Mohammad Amaan Sayeed, Boulbaba Ben Amor
Abstract: De novo design of nanobody complementarity-determining regions (CDRs) targeting user-specified epitopes is crucial yet remains computationally challenging. Current approaches rely on either (i) diffusion-based structural sampling that requires thousands of designs per target, (ii) gradient-based hallucination through frozen structure-prediction networks, or (iii) all-atom generative systems requiring 3D structural input and high computational cost. A fundamental limitation of these methods is their dependence on high-quality antigen structures, either experimental or computationally predicted, restricting applicability to the small fraction of therapeutic targets with available structural data. We present EpiRAG-PBind42 (Epitope-conditioned Retrieval-Augmented Generation with PBind42), a sequence-only framework for epitope-conditioned VHH/nanobody CDR generation. Our approach builds on PBind42, a Prot42-derived autoregressive binder generator instruction-tuned on DIPS-Plus protein–protein interaction pairs. EpiRAG-PBind42 adds three key innovations: (1) target-conditioned binder generation from sequence prompts, (2) retrieval-augmented latent-space decoding using hidden-state transition datastores derived from strict VHH/nanobody–antigen donor pairs, and (3) multi-agent filtering via four specialized evaluators (Structural, BLI, Interface, and Expression) that assess structural confidence, sequence naturalness, developability-adjacent liabilities, and interface energetics. Because EpiRAG-PBind42 operates on sequence input alone, it enables design on targets lacking high-quality structural models, including intrinsically disordered regions, flexible multi-domain proteins, and membrane-bound targets inaccessible to structure-dependent methods. We evaluate EpiRAG-PBind42 on five therapeutic antigen targets spanning diverse therapeutic areas: infectious disease, oncology, inflammation, and immunology.
PaperID: 4859, Poster
Authors:
Haoran Luo, Ziyue Zhu, Shangyang Wu, LINHAO LUO, Yikai Guo, Qika Lin, Fangzhi Xu, Jiapu Wang, Xiaobao Wu, Yifan Zhu, Anh Tuan LuuAbstract: Driven by increasingly complex real-world applications, Retrieval-Augmented Generation (RAG) has evolved from textual pipelines to multimodal settings, comprising two complementary dimensions: knowledge organization and retrieval enhancement. However, existing methods still face challenges in high construction cost, context window limitations, and the lack of unified training across offline and online knowledge sources. To address these challenges, we propose MMGraph-Agent, a unified multimodal RAG framework based on cache-inspired multimodal knowledge hypergraphs. MMGraph-Agent enables extremely lightweight hypergraph construction with zero API overhead and dynamic memory-based retrieval that alleviates long-context issues. Experiments across offline, online, and joint retrieval settings demonstrate consistent improvements in performance, efficiency, and architectural unification. Our software and data are publicly available.
PaperID: 4860, Poster
Abstract: Lifelong learning requires embodied agents to continuously understand the environment and learn how to use that understanding to guide action, enabling future interactions to yield more informative experiences. Prior language model (LM)-based agents have primarily addressed these requirements in isolation, emphasizing either declarative knowledge acquisition or procedural knowledge reuse, leaving the continual interplay between the two during lifelong learning underexplored. We introduce Neurosymbolic Knowledge Compilation (NeSyKC), a lifelong learning framework that couples declarative knowledge formation with incremental compilation into procedural knowledge through a shared symbolic representation. In NeSyKC, accumulated experience is abstracted into symbolic rules that capture reusable environmental structure, and these rules provide the basis for deriving executable procedures that are incrementally compiled into neural adapters. This creates a feedback loop in which experience is abstracted into symbolic rules and compiled into procedures that guide the LM's reasoning on future tasks, yielding further experience for learning. Across four open-ended environments instantiated from established embodied benchmarks, NeSyKC expands and refines reusable knowledge across tasks, improving task performance and inference efficiency while supporting generalization to unseen tasks. Real-world experiments on two robots further show that the agent continues to learn from post-deployment experience, allowing acquired knowledge to be compiled for reuse in future tasks.
PaperID: 4861, Poster
Abstract: AI security has become a critical issue as Deep Neural Networks (DNNs) are inherently vulnerable to adversarial attacks. While previous studies have demonstrated that adversarial training is one of the most effective defensive strategies, improvements in robustness often show significant class-wise disparities. This has established fair adversarial training, which aims to enhance the robustness of the worst-performing classes, as a crucial research direction. Although existing class-wise approaches can improve performance by adjusting individual class weights, they face severs scalability limitation on large-scale datasets due to the sparse volume of data available per class. To mitigate this issue without relying on explicit class-wise information, we propose a novel tri-regime optimization framework that fundamentally decouples the distinct sources of adversarial error at the sample level. Specifically, our method systematically suppresses unlearnable outliers via reliability gating, prioritizes genuine boundary threats through dynamic reweighting, and down-weights already safe examples to prevent redundant optimization. By isolating and prioritizing the effective adversarial frontier, our method strictly prevents robust overfitting and achieves superior worst-class robustness while maintaining high clean accuracy.
Abstract: E(3)-equivariant networks are promising for 3D atomistic system modeling, yet their scalability is limited by the O(L^6) complexity of the Clebsch-Gordan Tensor Product (CGTP). The recently proposed Gaunt Tensor Product (GTP) reduces the complexity but is unable to capture the antisymmetric paths, resulting in incomplete expressivity. In this work, we present SpinGT, an approach to overcome the GTP incompleteness by generalizing from scalar functions to Spin-Weighted Spherical Harmonics (SWSH). By relying on the algebraic properties of SWSH, SpinGTP recovers the missing antisymmetric interactions while maintaining the asymptotic efficiency of GTP. It also allows for a more expressive equivariant basis that naturally accounts for the parity-odd components of tensor products. We evaluate SpinGTP across diverse benchmarks, including Tetris, 3BPA, SPICE-MACE-OFF, and OC20. Our results show that SpinGTP achieves accuracies comparable to full CGTP. Notably, by explicitly capturing antisymmetric paths, SpinGTP exhibits superior performance in tasks involving chiral materials and non-centrosymmetric geometries. This work provides a scalable, complete, and mathematically rigorous path toward high-order equivariance in large-scale 3D atomistic system simulations.
Authors: Sonia Laguna, Jorge da Silva Gonçalves, Moritz Vandenhirtz, Alain Ryser, Irene Cannistraci, Julia Vogt
Abstract: Machine unlearning for vision models is rapidly becoming a practical requirement, driven by privacy regulations, data errors, and the need to remove harmful or corrupted training images. Despite this, most existing approximate unlearning methods tackle the problem from a post-hoc perspective. They attempt to erase the influence of targeted samples through parameter updates that typically require access to the full training data. This creates a mismatch with real deployment scenarios where unlearning requests can be anticipated, revealing a fundamental limitation of post-hoc approaches. We motivate , a novel paradigm for approximate methods in which models are directly trained to support forgetting as an inherent architectural capability. We instantiate this idea with Machine UNlearning via KEY deletion (MUNKEY), a memory-augmented transformer that decouples instance-specific memorization from model weights. Here, unlearning corresponds to removing the instance-identifying key, enabling zero-shot forgetting without weight updates or access to the original samples or labels. Across natural image benchmarks, fine-grained visual recognition, and medical datasets, MUNKEY outperforms all post-hoc baselines. Our results establish that
PaperID: 4864, Poster
Abstract: When reinforcement learning agents are deployed in decision-support settings, humans must repeatedly decide whether to adopt or override individual action recommendations. Explainable reinforcement learning supports these decisions, yet it remains unknown whether explanation mechanisms impact user behaviour. We conduct two controlled user studies that isolate explanation mechanism while controlling for agent policy, task, and abstraction level, using a design that independently manipulates action optimality and explanation veracity. Study 1 compares three representative action-level mechanisms across 100 participants and finds that Shapley-based attributions produce conservative reliance sensitive to explanation plausibility; saliency-based explanations indiscriminately increase adoption even for suboptimal actions; and a novel reward-decomposition mechanism, Action Advantage Attribution (AAA), achieves the highest appropriate adoption and is the only condition in which explanation comprehension positively predicts appropriate adoption. Study 2 benchmarks AAA against a no-explanation baseline across 70 participants, showing that it substantially improves detection of appropriate recommendations by 89% but does not eliminate over-reliance on suboptimal ones. Our results imply that algorithm designers must treat explanation mechanism as a first-order design choice, as mechanisms can influence human reliance on reinforcement learning agents.
PaperID: 4865, Poster
Abstract: Reconstructing the gene regulatory dynamics underlying tissue development from unpaired, cross-sectional spatial transcriptomic snapshots poses a significant challenge in single-cell biology. While recent optimal transport and flow-based methods have advanced trajectory interpolation, their unconstrained latent dynamics lack identifiability, hindering reliable discovery of the underlying regulatory mechanisms. To address this, we propose single-cell Identifiable Feedback-controlled Latent Flow (scIFLF), a generative framework that introduces a physics-inspired structural prior that decomposes the latent flow into a macroscopic developmental drift and a restorative feedback toward functional attractors. Crucially, the gene-regulatory Jacobian is shown to be uniquely identifiable, invariant to the latent ambiguity, enabling principled discovery of gene regulations directly from the learned flow. Implemented with a multimodal Neural ODE, scIFLF aligns latent trajectories using entropy-regularized optimal transport. Benchmark experiments demonstrate that scIFLF outperforms state-of-the-art methods in trajectory interpolation and spatial coherence, while successfully recovering key driver genes of tissue development and organogenesis, bridging the gap between deep generative flexibility and biological interpretability.
PaperID: 4866, Poster
Abstract: Fluorescent microscopy provides a rich view into how proteins localize within cells, but it remains experimentally infeasible to image human proteins across all of the different factors that can impact localization. We introduce Vermeer, a channel-adaptive autoregressive generative model for in silico generation of microscopy images of protein localization. Vermeer conditions generations on protein sequences and landmark stains showing the morphology of cells, which enables it to generalize to unseen proteins and cell lines. We show that Vermeer, trained on the Human Protein Atlas, can generate images with substantially improved perceptual quality and biological fidelity over previous proposals. Additionally, Vermeer's autoregressive framework enables flexible generation using varying channel subsets and orderings, enabling zero-shot transfer to data collected under different imaging conditions and channel configurations than those used for training. These results position Vermeer to enable scalable modeling of protein localization and is a step towards generative foundation models that can operate over distinct microscopy datasets.
PaperID: 4867, Poster
Authors:
Jason Toskov, Oriol Barbany, Rishubh Singh, Jinya Sakurai, Efe Tarhan, Oğuzhan F Kar, Roman Bachmann, Amir Zadeh, Jesse Allardice, Chuan Li, Carme Torras, Afshin Dehghan, Amir ZamirAbstract: The world appears spatially three-dimensional, however, existing large-scale multimodal models are often built mostly around 1D, 2D and 2.5D modalities. To move toward a more complete and actionable understanding of physical reality, models that natively model and generate across 1D, 2D, and 3D modalities are desirable. We present OmniDiMM, an any-to-any foundation model trained on a diverse set of 3D including meshes, Gaussian splats, voxels, and NeRFs; 2D such as images, and 1D like text. We develop tokenizers for a diverse set of 3D modalities, such as NeRF weights triplanes, UDFs, and 3DGSs, that convert them into discrete tokens, which enables multimodal masked modeling for joint training. Out of the box, OmniDiMM can directly perform standard tasks such as novel-view image synthesis and 3D generation from images, where it matches or outperforms existing specialized methods. As an any-to-any model, it can also convert between different 3D representations, such as NeRF weights and UDF, and generate various 3D modalities from 2D inputs, including DINOv2 features and surface normals. Joint training across 3D modalities also leads to strong transfer performance on downstream tasks such as grasping, classification, and object pose prediction. We scale training to a 1B-parameter model using 1T tokens across 1M objects. The pretrained models, training code, and multimodal dataset will be open-sourced.
Abstract: Self-evolving search agents reduce reliance on human-written training questions by generating and solving their own search tasks. We build on Search Self-Play (SSP), a representative Proposer and Solver framework in which questions are generated and answered via multi-step search and reasoning. In practice, however, SSP faces two bottlenecks: the Proposer constructs questions from isolated answer entities without relational context, yielding many invalid or unverifiable questions in early self-play training, while the Solver receives only a binary outcome reward that discards useful signal from partially on-track search trajectories. We address both bottlenecks by reusing knowledge-graph paths as construction-derived intermediate supervision for both question construction and reward shaping. First, we ground question construction in LLM-guided knowledge-graph subgraphs, providing relational context for the Proposer. Second, we observe that constructing and solving a multi-hop question can involve overlapping intermediate entities: the factual bridges used to formulate the question may provide approximate waypoints for answering it. Exploiting this overlap, we introduce Waypoint Coverage Reward (WCR), which grants graded partial credit to incorrect Solver trajectories according to their coverage of entities on the construction path, while preserving full reward for correct answers. Across seven QA benchmarks and nine model configurations, our approach improves the average score over standard SSP in all configurations, including notable gains on multi-hop QA tasks. These results suggest that knowledge-graph paths can be reused as lightweight intermediate supervision, providing both relational guidance and process feedback without additional task-specific human annotations or manually labeled process steps.
Abstract: Balancing policy expressiveness with the exploration-exploitation trade-off is a core challenge in online Reinforcement Learning (RL). While Stochastic Differential Equation (SDE)-based diffusion policies can represent complex, multimodal action distributions, they suffer from two critical limitations: their stochastic reverse processes render entropy intractable (necessitating heuristic exploration), and computing policy gradients through long denoising chains is expensive and unstable. In this work, we show that ODE-based flow matching inherently resolves these issues by enabling both simulation-free policy optimization and tractable entropy computation. Building on this, we introduce Flow Matching Policy Optimization with Mirror Descent and Entropy Constraints (FMER). Our framework exploits this insight in three ways. First, we theoretically establish that minimizing an advantage-weighted conditional flow matching loss acts as a simulation-free surrogate for policy mirror descent. This steers the velocity field toward high-value regions while entirely avoiding backpropagation through the ODE solver. Second, we derive an analytic entropy objective that corrects for the density distortion caused by the \tanh transformation (mapping an unbounded latent space to bounded actions), thereby facilitating principled maximum-entropy optimization. Finally, we dynamically tune the mirror descent temperature based on the effective sample size to enforce a robust trust region during training. Empirical evaluations demonstrate that FMER achieves superior performance on the challenging sparse-reward FrankaKitchen environment, while maintaining competitive results across standard dense-reward MuJoCo benchmarks. Code is available at \urlhttps://anonymous.4open.science/r/submission-25E2/.
PaperID: 4870, Poster
Authors: Huage Deng, LYU WENKAI, WU YONGJIN, Tianyang Zheng, heng Guo, Fangyun Zhou, Yu Shi, Wenxi Zhu, Minwen Deng
Abstract: Instruction-driven mask-free image editing with large-scale Diffusion Transformers (DiTs) has enabled high-fidelity local edits while maintaining global consistency. However, these models incur a prohibitive inference cost due to a pronounced semantic-–computation mismatch, where dense computation is performed over all high-resolution tokens regardless of localized editing instructions. Our analysis reveals that non-edit regions converge early with highly stable block-level outputs, exhibiting a distinct spatio-temporal asymmetry. Furthermore, we identify a critical system-level challenge that introducing such asymmetric computation into pipeline-parallel environments for ultra-large models triggers severe workload imbalance and pipeline bubbles, preventing algorithmic savings from translating into actual latency gains. To address these issues, we propose AsymPipe, a training--free algorithm–-system co-design framework that accelerates large-scale DiTs editing. The framework identifies edit versus non-edit regions on the fly using a zero-overhead velocity-based probe that reuses the predicted velocity field from the existing denoising process. For non-edit regions, AsymPipe caches and reuses block-level outputs, while for edit regions, it introduces a source-aware sparse attention module to precisely eliminate redundant contributions from static text and image sources. Finally, to bridge the gap between algorithmic optimization and hardware speedup, the framework incorporates a load-balanced token rearrangement scheme. This scheme deconstructs the original spatial topology to distribute tokens evenly across devices, physically restoring compute balance in pipeline-parallel systems. On HunyuanImage-3.0-Instruct and Qwen-Image-Edit-2511, AsymPipe achieves end-to-end speedups of 2.862x (PSNR 33.69) and 2.707x (PSNR 29.74), respectively, delivering substantial acceleration with virtually no loss in editing quality. Our anonymized source code is available at https://anonymous.4open.science/r/AsymPipe-6553.
PaperID: 4871, Poster
Authors: Jinhyuk Choi, Hyung Tae Lee
Abstract: Achieving both privacy and robustness against poisoning attacks remains a central challenge in federated learning (FL), since hiding individual updates makes attacks harder to detect and mitigate. Recently, Mai et al. (NeurIPS 2024) proposed RFLPA, a packed-secret-sharing (PSS)-based realization of FLTrust, which weights client updates by their cosine similarity to a trusted server update. In this paper, we show that RFLPA contains design flaws and mathematical errors in its use of PSS, making the protocol theoretically unsound and practically non-functional. We then propose an improved verifiable FL framework that restores correctness while retaining efficiency. Our key insight is a reinterpretation of the algebraic structure of PSS, which leads to a new transposed degree reduction algorithm and a pairwise verification technique for dot products. These primitives enable local matrix-based operations without complex circuit evaluation and substantially reduce communication rounds compared with BGW-style computation. Building on them, we redesign the verifiable aggregation algorithm of RFLPA by introducing perpendicular vectors, reducing cosine similarity verification to a single dot-product check and eliminating costly verifiable secret sharing during enrollment. The resulting framework strictly improves efficiency, fixes the flaws of prior work, and preserves N/3-robustness in malicious settings. Our implementation and empirical evaluation confirm the feasibility of the resulting protocol for high-dimensional FL updates.
PaperID: 4872, Poster
Abstract: Modern agentic systems extend language models with external tools, enabling them to complete complex tasks through multi-turn interactions. As these agents are deployed in consequential settings, their behavior is often constrained by policies governing pricing, privacy, confirmation, safety, and other operational requirements. However, existing evaluation methods leave a critical gap: capability benchmarks measure what agents can accomplish, while content-safety red-teaming measures harmful generations, but neither tests whether agents comply with predefined policies during otherwise valid tool use. Such policy violations are difficult to audit because their effects may emerge only after several turns, span multiple tool calls, or surface in downstream systems. We propose PROBE (Policy Rule Observation via Behavioral Exploration), an algorithm for training LLM-based auditors that interact with deployed target agents as realistic users and probe them for policy violations through multi-turn, tool-mediated dialogue. We formalize agent auditing as a two-player Markov game and propose a composite reward that balances violation discovery, behavioral plausibility, conversation progress, and policy diversity. To study the role of training objectives and reward design, we compare GRPO, DPO, and GFlowNet-based optimization in a controlled setting. Across extensive experiments, auditors trained with PROBE uncover substantially broader and higher-yield policy violations than same-architecture zero-shot adversaries, even when the auditor is significantly smaller than the target agent. Our results show that multi-turn behavioral auditing is a distinct and necessary axis of agent evaluation, complementing capability and content-safety benchmarks. More broadly, PROBE offers a scalable path toward automated policy auditing for real-world agentic systems and suggests a future co-evolutionary loop in which auditors expose failures and agents improve from auditor feedback.
PaperID: 4873, Poster
Authors:
Yingjie Ma, Haonan Wang, Xun Lin, Ruixin Zhang, Jingyun Zhang, Jun Wang, Rizen Guo, Shouhong Ding, Weicheng Xie, Linlin Shen, Zitong YUAbstract: Face anti-spoofing (FAS) is critical for protecting face recognition systems against presentation attacks in open environments. Beyond high cross-domain generalization, practical FAS deployments must provide interpretable and uncertainty-aware decisions when multimodal evidence is incomplete, degraded, or conflicting. However, current multimodal domain-generalized FAS methods focus primarily on feature fusion or domain alignment, compressing RGB, depth, and infrared cues into opaque logits. This makes it difficult to discern which modality drives the decision, whether modalities reinforce or suppress each other, and whether a given sample lies within the classifier's reliable operating regime. We introduce SEM-FAS, a quantum-inspired yet fully classical single-pass evidence measurement framework that jointly enhances generalization, interpretability, and reliability. SEM-FAS encodes each modality as a normalized complex-valued state whose amplitude captures evidence strength and whose phase encodes cross-modal interference; multiple coherent heads are then mixed into a density-matrix representation. A constrained POVM-like measurement head converts this representation into calibrated class probabilities while analytically decomposing every prediction into per-modality main effects and pairwise synergy/suppression terms. Simultaneously, three built-in uncertainty indicators are obtained without sampling overhead: predictive ambiguity, representation mixedness, and measurement-coverage mismatch. On standard multimodal DG-FAS benchmarks, SEM-FAS reduces the best baseline HTER by 2.18% and improves AUC by 1.38% under both complete and missing-modality protocols. Visual analyses further suggest that the learned modality effects, pairwise interactions, and uncertainty scores align with expected FAS behavior, even for hard samples, indicating that single-pass evidence measurement is a promising way to jointly enhance generalization, interpretability, and reliability in multimodal FAS.
Abstract: Cortical neurons are complex, multi-timescale processors wired into recurrent circuits, shaped by long evolutionary pressure under stringent biological constraints. Mainstream machine learning, by contrast, predominantly builds models from extremely simple units, a default inherited from early neural-network theory. We treat this as a normative architectural question. How should one split a fixed parameter budget P between the number of units N, per-unit effective complexity k_e, and per-unit connectivity k_c? What controls the optimal allocation? This calls for a model in which per-unit complexity can be tuned independently of width and connectivity. Accordingly, we introduce the ELM Network, whose recurrent layer is built from Expressive Leaky Memory (ELM) neurons, chosen to mirror functional components of cortical neurons: multi-timescale memory, structured synaptic integration, and nonlinear internal computation. The architecture allows for individually adjusting N, k_e, and k_c and trains stably across orders of magnitude in scale. We evaluate the model on two qualitatively different sequence benchmarks: the neuromorphic SHD-Adding task and Enwik8 character-level language modeling. Performance improves monotonically along each of the three axes individually. Under a fixed budget, a clear non-trivial optimum emerges in their tradeoff, and larger budgets favor both more and more complex neurons. A closed-form information-theoretic model captures these tradeoffs and attributes the diminishing returns at two ends to: per-neuron signal-to-noise saturation and across-neuron redundancy. Connectivity enters as a related mechanism that helps neurons learn distinct signals. A hyperparameter sweep spanning three orders of magnitude in trainable parameters traces a near-Pareto-frontier scaling law consistent with the framework, mapping the budget-constrained tradeoff surface between unit count, unit complexity, and connectivity. This suggests that the simple-unit default in ML is not obviously optimal once this surface is probed, and offers a normative lens on cortex's reliance on complex spatio-temporal integrators.
PaperID: 4875, Poster
Abstract: Effective Continual Video Instance Segmentation (CVIS) requires balancing catastrophic forgetting with the plasticity to learn new tasks. Existing methods typically adapt pre-trained VIS models using trainable visual prompts, but suffer from two key limitations: they overlook the distinct degradation of frame-level spatial knowledge (e.g., misclassification) and video-level temporal knowledge (e.g., trajectory disruption), and rely on rigid prompt retrieval that struggles under task ambiguity and distribution shifts. To address these issues, we propose STEAM, a CVIS framework with Spatial-Temporal Balanced Mixture-of-Experts Adapters. Instead of explicit retrieval, we introduce a Distribution-Aware spatial-temporal Scaling and sHifting (DASH) mechanism to align the pre-trained feature space with sequential tasks and reduce the domain gap. Furthermore, our method adaptively learns Task-Adaptive Spatial-Temporal mixture-of-Experts (TASTE) at both the frame and video levels, while updating them with Subspace-Constrained Orthogonal Adaptation (SCOA) to mitigate interference. Extensive experiments across diverse CVIS settings show that STEAM preserves temporal-spatial coherence and achieves state-of-the-art performance.
PaperID: 4876, Poster
Abstract: Multi-token prediction (MTP) accelerates autoregressive decoding through blockwise prediction, but remains vulnerable to error accumulation across sequentially generated blocks. Existing work primarily focuses on improving MTP’s acceleration ratio while largely overlooking the compounding errors that accumulate during block-wise generation. In this work, we study MTP from an information-theoretic perspective and identify the strength of inter-block dependencies as a key factor governing its performance. Building on the cumulative information loss, we theoretically quantify the gap of block-sequence representation. Motivated by the analysis, we propose Dual-MTP, a novel MTP training framework that explicitly models inter-block dependencies from two complementary perspectives, thus reducing the block-sequence representation gap across decoding steps. Specifically, Dual-MTP introduces a dual-perspective training objective with random masking, where each block is jointly trained for forward prediction and masked context reconstruction, encouraging robust block-level representations. Experiments across five model backbones and six tasks spanning math reasoning and GLUE understanding show that Dual-MTP consistently outperforms standard MTP on both quality and efficiency, attaining the highest average accuracy on every backbone with gains of +0.3 to +3.0 points and a 1.4× to 1.9× wall-clock speedup over standard autoregressive decoding.
PaperID: 4877, Poster
Abstract: The deployment of AI agents in multi-agent workflows enables inter-agent influence, whereby agents strategically steer other agents’ behavior in line with a specific goal. Such capabilities pose novel risks, as malicious agents could direct others toward harmful actions. In this work we investigate three such capabilities –- persuasion, deception, and coercion –- across five realistic evaluation environments and a variety of frontier models. We observe significant inter-agent influence capabilities in frontier models. In oversight environments, all tested models shifted policy-violating decisions from rejection to approval through persuasion, deception, and coercion. In peer-to-peer settings, models extracted concessions in scheduling negotiations and redirected a peer's safety research trajectory toward an attacker-preferred direction. While evidence is weakest for inter-agent coercion, several frontier models independently formulated and executed coercive threats, including threats to harm a named human in the environment. These results point to a pressing need for work on risk mitigation to promote beneficial deployments of multi-agent systems.
PaperID: 4878, Poster
Authors:
Jiawei Gu, Linjie Li, Yiming Liu, Yunzhuo Hao, Zhichao Peng, Huichen W Wang, Guanzheng Chen, Luxin Xu, Xinyu Zhang, Li Luo, Disen Lan, Zican Hu, Mingyang Song, David Valente, Alex Jinpeng Wang, Yafu Li, Ganqu Cui, Zhengyuan Yang, Michael Shieh, Yejin Choi, Ranjay Krishna, Yu ChengAbstract: Human reasoning is not purely linguistic: for visual problems, people often think by changing the visual representation itself. Interleaved multimodal reasoning seeks to bring this ability to unified models by allowing them to construct intermediate visual states and use them within multimodal reasoning traces. Yet scaling this ability is difficult, since task-driven instruction tuning must teach visual-state construction, faithfulness, and downstream reasoning all at once. We introduce PRIMER (PRetext-based Interleaved Multimodal rEasoning Recipe), a two-stage training recipe that uses pretext reasoning as a primer for interleaved multimodal reasoning. Stage 1 builds \pretext, an instruction-free corpus that converts classical self-supervised vision pretext tasks into Thought--Image--Thought traces, teaching the model the reusable mechanics of producing and reading visual states. Stage~2 builds \instruct, a task-driven instruction-tuning corpus that teaches task-relevant visual-state construction across perception, spatial understanding, and mental world modeling. On a 7B unified multimodal model, Primer improves over the matched-budget BAGEL reference by +10.40, +5.93, and +15.70 across the three cognitive levels, with shorter traces, and rivals models several times its size on spatial and world-modeling averages. These results establish classical self-supervised pretext tasks, repurposed as Thought-Image-Thought traces, as a scalable, parameter-efficient primer for interleaved multimodal reasoning, the recipe at the core of Primer.
PaperID: 4879, Poster
Authors: Wentao Zhang
Abstract: Many online learning systems now have access to several gradient predictors at once, such as simulators, replay buffers, neural surrogates, or time-series forecasters, and the quality of these predictors varies from round to round. We study online convex optimization under one-gradient feedback against an arbitrary moving comparator, with a user-supplied finite class \Pi of gradient-field predictors, and ask whether a single algorithm can compete with the best predictor in \Pi at a logarithmic cost, fall back to a small-loss bound when none of the predictors is useful, and still retain a path-length-adaptive worst-case guarantee when the predictors are adversarial. We give a positive answer through PH-Sword (Predictable-Hybrid Sword.optimism), which combines an exp-concave Hedge over \Pi, optimistic projected-gradient base learners on a geometric learning-rate grid, and a correction-term optimistic-entropy meta learner with prediction-error doubling. For any comparator sequence \boldsymbol u, PH-Sword attains dynamic regret \widetilde O\bigl(R\_0\sqrt(1+P\_T/D)(1+P\_T/D+\min\\mathcal E\_\/G^2,\,2LF\_T(\boldsymbol u)/G^2\+\log|\Pi|)\bigr) up to logarithmic factors, where R\_0=DG+LD^2, P\_T is the path length, F\_T(\boldsymbol u) the cumulative comparator loss, and \mathcal E\_\ the best-in-class field error. The bound merges the field-controlled, small-loss, and worst-case regimes into one deterministic guarantee that improves over prior one-gradient dynamic-regret algorithms. On six controlled benchmarks the five predicted scaling laws hold simultaneously: PH-Sword reaches 17% of OGD's final dynamic regret and 54% of Ader's on a slow-drift benchmark, and we observe that a single noisy predictor without model selection can actually do worse than using no predictor at all.
PaperID: 4880, Poster
Authors:
Kevin Yu, Tao Guo, Constantinos Antoniou, Panagiotis AngeloudisAbstract: Neural trajectory predictors are increasingly used in physical-system pipelines for reconstruction, simulation, forecasting, and downstream analysis. In these settings, low prediction error alone is insufficient. A trajectory may remain statistically plausible while violating dynamics, actuator limits, or state constraints, making it unreliable for physical interpretation or control-aware reasoning. This problem is especially difficult when controls are unobserved and the available dynamics model is only partially specified. We introduce the Markovian Dynamics Enforcer (MaDE), a framework for post-hoc feasibility enforcement under learned controlled dynamics. MaDE acts as a time-invariant correction operator that maps state-transition proposals onto a learned feasible dynamics manifold. For each transition, it infers latent controls through inverse dynamics, recomputes state evolution using a known-physics model augmented with a learned residual, and corrects controls through gradient-based inequality reduction while re-integrating the dynamics after each correction step. MaDE is trained from feasible state observations without ground-truth controls and is designed to operate as a frozen downstream layer for arbitrary trajectory predictors. We evaluate MaDE on simulated controlled systems spanning fully specified and underspecified dynamics, with deterministic bound-violation and Gaussian observation-noise stress tests. MaDE is the only method that drives known-, learned-, and true-dynamics residuals to essentially zero across the fully specified systems. In the underspecified dynamic-bicycle setting, it reduces true-dynamics residual from 1.40--2.90 for the baselines to 0.33, while reducing trajectory fidelity error by more than 60% relative to the next-best method. These results show that MaDE enforces model-relative physical feasibility while preserving proximity to the original trajectory.
Authors:
Andrew Cullen, Neil Marchant, Jiani Xie, Paul Montague, Benjamin RubinsteinAbstract: Automatic Speech Recognition systems are notoriously both sensitive to input perturbations and challenging to defend, which is a product of their high-dimensional, discrete output spaces. Traditional Randomized Smoothing workflows collapse in sequence-to-sequence tasks because the probability mass of any single transcription vanishes under noise. We propose an Anytime-Valid Certified Transcription framework that replaces fixed-sample binomial testing with E-value Martingales. Our approach leverages a dual-gate pipeline: a Two-Sided Atomic Audit that accumulates statistical wealth to certify both token existence and adversarial exclusion, and a Rank-Based Tournament that selects the winning sequence. Our evaluations across four diverse architectures demonstrate that this approach yields an average 25.5% relative reduction in Word Error Rate, while also providing granular word- and sentence-level certifications to enhance acoustic security.
PaperID: 4882, Poster
Abstract: Retrieval-Augmented Generation (RAG) has been widely adopted to enable Large Language Models (LLMs) to ground their responses in external knowledge sources. Nonetheless, recent studies show that conflicts between the retrieved external knowledge and the model’s parametric knowledge can lead to hallucinatory outputs, and this problem is exacerbated when the retrieved documents contain noise. In this work, we propose Conflict-Aware Representation Editing (CARE), an inference-time representation editing method for improving LLM robustness under noisy knowledge conflict settings. CARE learns a latent editing direction from intermediate representations associated with correct and incorrect generations, and uses this direction to mitigate conflict-associated activation patterns during inference. We evaluate CARE across Question Answering (QA) benchmarks and LLMs, showing improvements over retrieval-based and conflict-mitigation baselines, especially under noisy retrieval conditions.
PaperID: 4883, Poster
Abstract: We study fixed-confidence joint policy testing in discounted tabular Markov decision processes under active exploration. Given a finite family of target policies, the learner observes a single adaptive trajectory and must certify the sign of every target-policy value with probability at least 1-\delta. We identify the instance-specific first-order benchmark for this problem: a characteristic time T^\star(p), defined by a max--min program over stationary occupancy measures and sign-flipped alternatives. We then develop PT-ACE(\mu), an online algorithm that couples stabilized exchange-based learning of a bottleneck occupancy allocation, trajectory-compatible navigation from averaged occupancies, and parallel certified policywise stopping via scalar frontier tests. Under a trajectory-side anchor-policy ergodicity condition, and when run with certified numerical subroutines and vanishing input tolerances, PT-ACE(\mu) is \delta-correct and, for every sufficiently small admissible fixed floor \mu>0, satisfies \limsup_\delta\downarrow0 \frac\mathbb E_p[\tau_\delta]\log(1/\delta) \le (1+c(\mu,p))T^\star(p), \qquad c(\mu,p)\to0\quad\textas \mu\downarrow0. Thus a single adaptive trajectory can certify multiple policy signs at the instance-specific lower-bound rate in the vanishing-floor limit.
PaperID: 4884, Poster
Authors:
Bo-Han Lai, Hsuan-Tien Lin, Chia-Mu Yu, Han Zhao, ShangTse ChenAbstract: Post-hoc concept unlearning is a practical approach to remove undesired content from text-to-image diffusion models. However, existing methods leave unlearned models fragile: small perturbations to prompts, text embeddings, or sampling-time latent states can recover supposedly erased content, even when standard endpoint metrics suggest successful erasure under nominal prompts. We study this fragility through the stability of the denoising trajectory of diffusion models. Our key observation is that prompt-, embedding-, and sampling-based attacks share a similar mechanism of perturbing the sampling dynamics, and erased concepts can reappear when the unlearned model amplifies these concept-relevant perturbations. We formalize this view with a finite-time stability analysis and derive a tractable directional sharpness measure along a concept-relevant direction. Motivated by this analysis, we propose Directional Stability Regularization (DSR), a plug-in regularizer that discourages expansion in erased-concept directions without directly penalizing unrelated directions. DSR is compatible with classical score-based unlearning objectives, requires no adversarial training or attack-specific inner loop, adds no inference-time overhead, and can be estimated efficiently with Jacobian-vector products. Across the evaluated erasure settings, DSR reduces recovery under prompt-, embedding-, and sampling-based attacks while largely preserving prompt alignment and image quality metrics.
PaperID: 4885, Poster
Abstract: Discriminative models are trained with task-specific objectives such as classification, detection, or representation learning, yet their practical utility depends on generalizing beyond the training task to unseen samples from the same domain. This suggests that a discriminative model may retain an implicit trace of the data manifold beyond its task labels. We ask whether this trace can be exposed as a generative signal without task-specific guidance. We introduce the Discriminative Score Function (DSF), which extracts an image-space generative field from the loss-gradient geometry of a pretrained discriminative model. DSF provides a training-free update rule that moves images toward model-consistent visual regions without retraining or task-specific targets. We apply DSF to diverse pretrained discriminative models, including ResNet classifiers, DETR detectors, and DINOv2 ViT encoders. DSF generates unconditional images from noise and supports conditional guidance, editing, inpainting, and explanation through the same field. DSF-generated samples also provide effective zero-shot quantization calibration when real images or labels are unavailable, preserving feature geometry better than data-free calibration baselines. These results show that discriminative training leaves a generative functional signal for synthesis and calibration.
PaperID: 4886, Poster
Abstract: Real-world degraded images often contain multiple co-occurring degradation types, making composite image restoration fundamentally different from the single-degradation setting assumed by most existing all-in-one methods. These methods typically apply uniform spatial computation and single-label task conditioning, limiting both spatial adaptivity and explicit modeling of degradation mixtures. We propose degradation cues should guide restoration: a spatial mixer routes patches to graded-capacity experts by local restoration difficulty, and a channel mixer routes each image through degradation-specific experts, conditioned on a global task feature with multi-hot classification supervision. Empirically, CART achieves state-of-the-art performance for composite degradation restoration on CDD-11 and delivers the best results to date on the conventional 3-task and 5-task all-in-one benchmarks.
Abstract: Robust optimization (RO) provides a principled framework for decision-making under uncertainty, but its practical use is often limited by the need to manually reformulate uncertain optimization models into tractable deterministic counterparts. Recent large language models (LLMs) have been shown promising for automating optimization formulation, yet RO reformulation remains challenging because it requires precise multi-step reasoning and mathematically consistent transformations. To facilitate systematic evaluation of LLM-based reformulation, for which no dedicated benchmark currently exists, we develop AutoRO-Bench, a benchmark featuring an automated data generation pipeline for the core RO reformulation task and a curated dataset for the RO application task. To address the reformulation challenge, we propose Automated Reformulation with Experience Memory (AutoREM), a tuning-free memory-augmented framework that autonomously builds a structured textual experience memory by reflecting on past failed trajectories through a tailored offline adaptation procedure. AutoREM requires neither domain-specific expert knowledge nor parameter updates, and the resulting memory readily transfers across different base LLMs. Experimental results show that AutoREM consistently improves the accuracy and efficiency of RO reformulation across in-distribution datasets, out-of-distribution datasets, and diverse base LLMs.
PaperID: 4888, Poster
Authors:
Ding Jia, Wei Liu, Xianglong Du, Yingjie Li, Yingqing Yang, Huili Yu, Zhangsong Zhan, Chu ZhouAbstract: The transition from Large Language Models (LLMs) to agents shifts safety stakes from toxic text to irreversible environmental harm. While current defenses remain largely retrospective, proactive runtime intervention is bottlenecked by the lack of large-scale, causally-consistent data. We propose PROACT-Agent, a framework for synthesizing high-fidelity trajectories to enable real-time guardrails. We identify a critical "safety drift" in prior benchmarks, where lenient annotation paradigms fail to enforce temporal consistency. PROACT-Agent addresses this through: (1) Progressive Trajectory Unrolling to reveal risks hidden in long-context interactions; (2) Reasoning-Augmented Causal Rectification to enforce monotonic causal consistency; and (3) Culturally-Aware Data Localization for cross-border robustness. We introduce PROACT-Bench, the largest bilingual safety benchmark to date, featuring 140,000+ trajectories adjudicated with industrial-grade rigor. Experiments show that models trained via our framework achieve superior zero-latency intervention, establishing a new standard for real-time autonomous agent safety.
PaperID: 4889, Poster
Abstract: Open multi-agent systems, where only a subset of agents are controllable, are central to real-world domains such as smart grids, yet lack a principled foundation for reinforcement learning algorithm design. We study this setting through n-agent ad hoc teamwork (NAHT), a recent framework for learning in open multi-agent environments. We propose an axiomatic framework that derives value function learning from the Shapley axioms (Efficiency, Symmetry and Linearity), treating them as structural constraints rather than heuristic design choices. This formulation builds on a novel NAHT decomposition, which characterises the fundamental difference between open and fully controllable multi-agent systems, and establishes a cooperative game-theoretic foundation for learning. We further show that enforcing these axioms on learning individual value functions recovers the Shapley value for Dec-POMDPs, providing a principled bridge between reinforcement learning and classical cooperative game theory. Building on this framework, we introduce Shapley Machines and Banzhaf Machines, two generic algorithmic templates that can be instantiated with standard reinforcement learning methods. When applied to strong baselines such as IPPO and POAM, our methods consistently improve learning efficiency and generalisation to unseen agent types and conventions across NAHT benchmark tasks.
PaperID: 4890, Poster
Abstract: Neural subgraph retrieval systems, which retrieve corpus graphs containing a query graph as a subgraph, are used in safety-critical applications such as drug discovery, hardware trojan detection, and molecular similarity search. Despite their importance, the adversarial robustness of these systems remains unstudied. We propose GRAP, the first adversarial attack framework targeting subgraph matching-based graph retrieval. \our jointly selects a budget-constrained subset of corpus graphs and computes per-graph edge perturbations to maximally degrade retrieval quality. We formalize two attack regimes: a ranking attack (with solver access) that corrupts pairwise relevance ordering while preserving relevance labels, and a solver-free top-K attack that directly displaces relevant results from retrieved sets. In both cases, we show that the resulting adversarial set functions are monotone and approximately submodular, enabling greedy subset selection with provable approximation guarantees. Experiments on five benchmark datasets against multiple state-of-the-art victim retrievers demonstrate that our method consistently and significantly outperforms all baselines in both gray-box and black-box settings, with and without solver access.
PaperID: 4891, Poster
Abstract: We study finite-horizon two-player zero-sum perfect-recall partial-information games and ask for the minimal finite-dimensional state sufficient to answer all rooted observation-measurable unilateral continuation queries. For each player i and information set I, we consider the span \mathcal L_i(I) of feasible counterfactual beliefs and the effective local query space \mathcal E_i(I) generated by rooted continuation queries restricted to \mathcal L_i(I). We prove that its dimension d_i(I) is simultaneously the rank of the local counterfactual query operator, the minimum core-query basis size, the minimum exact local realization dimension, and the dimension of the canonical quotient \mathcal L_i(I)/\mathcal N_i(I), yielding reduced counterfactual predictive-state representations, unique up to similarity, with exact first-return update operators. Aggregating the local ranks defines the deviation-conditioned predictive-state dimensions d_t^i, d_t, d_\star, and D_\star. We then prove that local rooted payoff-query control transfers to global exploitability, and under local design richness and bounded one-step parameter norms obtain a rollout-based upper bound in these minimal coordinates. We complement this with a minimax lower bound showing that \Omega(d/\gamma^2) local rollouts are necessary even when the peak DCPSD equals d, and with an exponential separation theorem exhibiting games with d_\star=m but observable continuation-trace complexity 2^m. Finally, we show that DCPSD recovers finite-rank predictive-state, observable-equivalence, and linearly parameterized continuation-law models as special cases.
PaperID: 4892, Poster
Abstract: The standard mechanistic interpretation of transformer factual recall attributes storage to deep MLP modules and routing to attention. We study how module roles change when the fact distribution carries subject, relation, or answer-collision structure rather than independent facts. Which module performs the recall depends on the correlation axis and on the model's capacity load. Probes and matched-counterfactual patching disagree systematically on every motif: probes credit the deepest MLP while patching credits the first attention module. We show that probes and patching measure structurally different quantities---cumulative linear decodability versus irreplaceable causal contribution---which under correlated fact distributions need not coincide. A parameter-counting argument singles out subject-axis collision as the unique compression that the attention-routing pathway can exploit cheaply, predicting an order-of-magnitude crossover for that pathway class that we observe empirically: the layer-2 probe gap collapses sharply for subject collisions, and only subject-axis specialization survives the capacity-tight regime. Edge-level path patching uncovers a three-step chain (Attn -> MLP1 -> MLP2) whose dominance tracks the predicted threshold, with a causal-scrubbing test corroborating that the chain is causally relevant on every motif and sharpest in the subject regime. A dense joint sweep shows that subject and answer compressions compete rather than amplify, with the dominant axis set by the stronger compression. On natural CounterFact strata across four frozen pretrained LMs (Pythia-1.4B/2.8B, LLaMA-3.2-1B-Inst/3.1-8B), edge-level path patching reproduces the synthetic ranking: the answer-collision stratum is the most localized on every model. A two-hop extension shows the binding-locality structure carries to compositional facts.
PaperID: 4893, Poster
Abstract: Recent multi-view Transformers (e.g., DUSt3R and VGGT) have advanced 3D reconstruction by scaling training data and model capacity from two views to many views and long sequences, subsuming the classical SfM/SLAM pipeline of feature extraction, matching, and pose estimation in a single feedforward pass. Yet this subsumption hides the interfaces between pipeline stages---correspondences, geometry, and pose are entangled in a single forward pass, to improvement and offering no mechanism to correct predictions from geometric evidence at test time. We propose PosePlaner to close this gap by reviving the classical coupling between correspondences and pose via flow matching, treating pose refinement as a conditional generation problem that takes coarse geometric matches and initial pose as input condition and fuses the resulting pose hypotheses into a single refined estimate, providing a continuous, data-driven analogue of RANSAC voting. Our method is trained efficiently on two-view data, yet serves as a periodic refiner for streaming n-view pose estimators with no retraining. Empirically, it corrects failures of feedforward predictors, substantially improves correspondence-based solvers, and reduces accumulated drift in long streaming sequences.
Abstract: Diffusion models have shown promise in learning to solve constraint optimization problems. However, they are mostly restricted to problems with binary variables and rely on graph neural networks, hindering their application to a broader range of problems such as those with general discrete variables or constraint structures that necessitate global rather than local reasoning. We investigate the use of Diffusion Transformers to address the aforementioned limitations. A naive implementation performs poorly due to a fundamental mismatch between the standard diffusion process and constraint solving: while the former applies ransformer (BloGDiT), is the first to address this limitation by replacing standard joint Gaussian denoising with Blocked Gaussian denoising. BloGDiT uses iterative block resampling and anneals the block size over time to facilitate large, targeted edits within a block of variables. Across Sudoku, Graph Coloring, Maximum Independent Set, and MaxCut, BloGDiT matches or outperforms existing methods, demonstrating that blocked Gibbs-style diffusion provides a highly effective inductive bias for Transformer-based constraint satisfaction and optimization.
Abstract: Two-Stage Robust Optimization (2RO) with discrete uncertainty is challenging, often rendering exact solutions prohibitive. Scenario reduction alleviates this issue by selecting a small, representative subset of scenarios to enable tractable computation. However, existing methods are largely problem-agnostic, operating solely on the uncertainty set without consulting the feasible region or recourse structure. In this paper, we introduce PRISE, a problem-driven sequential lookahead heuristic that constructs reduced scenario sets by evaluating the marginal impact of each scenario. While PRISE yields high-quality scenario subsets, each selection step requires solving multiple subproblems, making it computationally expensive at scale. To address this, we propose NeurPRISE, a neural surrogate model built on a GNN-Transformer backbone that encodes the per-scenario structure via graph convolution and captures cross-scenario interactions through attention. NeurPRISE is trained via imitation learning with a gain-aware ranking objective, which distills marginal gain information from PRISE into a learned scoring function for scenario ranking and selection. Extensive results on three 2RO problems show that NeurPRISE consistently achieves competitive regret relative to comprehensive methods, maintains strong scalability with varying numbers of scenarios, and delivers 7--200× speedup over PRISE. NeurPRISE also exhibits strong zero-shot generalization, effectively handling instances with larger problem scales (up to 5×), more scenarios (up to 4×), and distribution shifts.
PaperID: 4896, Poster
Authors:
Zhaolu Kang, Tailong Luo, Sihan Liu, Guangyuan Dong, Siheng Wang, Lei Wei, Shuaibo Li, Rongchao Zhang, Zhen Tian, richeng xuan, Zhichao HuAbstract: World models are increasingly used to reduce environment interaction when post-training vision-language-action (VLA) policies with reinforcement learning. Yet they are typically trained to predict pixels one step ahead, even though their role here is not merely to render visually faithful videos but to produce learning signals from imagined rollouts that improve the policy. We show that this mismatch is substantial: lower pixel MSE can produce worse policy gradients, and a world model trained for policy utility can yield better downstream policies despite 17% higher pixel MSE on the next frame. We propose the Downstream-Return Imagination Learning Loop (DRILL), a bilevel framework that updates the world model to maximize the policy's return after an inner GRPO update on imagined rollouts. DRILL has two instantiations. DRILL-VWI is a closed-form surrogate, requiring no additional simulator rollouts, that reweights prediction errors by policy visitation, the group-relative GRPO advantage, and model confidence. DRILL-IMG is a full meta-gradient method that differentiates through the inner update using truncated Hessian-vector products. We show that DRILL-VWI recovers the local one-step component of the DRILL-IMG meta-gradient under score-compatibility and small-residual assumptions. Across five manipulation simulators and two VLA backbones, DRILL-IMG raises mean success rate by 13.1 points over WMPO and 7.9 points over RLVR-World, and complements RLVR pretraining rather than competing with it. More broadly, we believe video world models for VLA reinforcement learning can be productively trained by the policy gradients they produce, not only by pixel-level prediction.
Authors:
Wenjie Liao, Like Wu, Liangjie Zhao, Shihui Xu, Shigeru FujimuraAbstract: Self-play fine-tuning enables large language models to improve beyond supervised fine-tuning without additional human annotations by contrasting annotated responses with self-generated ones. Many existing methods rely on a fixed divergence regime. SPIN is closely related to a KL-based regime, SPACE to a Jensen-Shannon-style objective via noise contrastive estimation, and SPIF to \chi^2-regularized self-play. Since these divergences exhibit different strengths depending on the distributional gap between model and target, no single choice appears to provide favorable learning dynamics across training stages. We propose IRIS (Interpolative R\'enyi Iterative Self-play), a R\'enyi-based self-play fine-tuning framework with a continuously adjustable objective. IRIS decomposes into two independent tilted risk terms over annotated and synthetic data, with exponential importance weights controlled by the order parameter \alpha. We show that several self-play objectives can be interpreted as limiting or representative regimes at particular values of \alpha, providing a unified theoretical perspective on these methods. An adaptive order schedule further adjusts \alpha to the distributional gap, shifting from sharper importance weighting early in training to smoother refinement near convergence. Theoretically, we establish the fixed-point property of IRIS and analyze how \alpha controls gradient concentration. Experiments on Zephyr-7B and Qwen2.5-3B across ten benchmarks show that IRIS improves upon baselines, reaching 44.57% average score with gains across iterations. In our setting, IRIS with only 26k annotated samples surpasses standard supervised fine-tuning trained on the full 200k dataset.
PaperID: 4898, Poster
Abstract: As large generative models become widely deployed and customized, federated learning is increasingly used to adapt them while keeping user data local. Because clients repeatedly receive up-to-date adapters, a malicious client can copy a dispatched adapter, use it offline, and monetize generated images without exposing the stolen weights or a queryable service. Existing federated watermarking and traitor-tracing methods usually assume white-box access to suspect weights or black-box query access to the deployed model. In realistic generative-model theft, however, the defender may only observe images that have already circulated. We propose \emphFedTrace, a generated-content-based watermark verification framework that attributes leaked federated models from suspicious generated images alone. FedTrace couples three designs: a round-wise watermark lifecycle that separates client identity distribution from global utility aggregation, a low-drift reliable-bit carrier that embeds identity in watermark positions stable under local adaptation, and anti-collision coding with soft subset verification for distinguishing singleton and collusive leakage. These components enable post-local, collusion-aware attribution from generated outputs alone, without suspect weights or online query interfaces. Extensive experiments across customized diffusion datasets, base models, and federated settings show that FedTrace preserves detectable client identities under local adaptation and strengthens subset-aware tracing against collusive leakage.
PaperID: 4899, Poster
Authors: Marcel Mordarski, Daniel Budina, Benjamin Gras, Abdulrahman Shehata, Roberto Bondesan
Abstract: Joint optimisation over a discrete topology T and continuous parameters \theta under a strict evaluation budget B arises in neural architecture search, pruning, and variational quantum algorithms. We identify a budget-induced effect, \emphdimensionality pressure: enlarging T expands the search space of \theta, slows inner-loop optimisation at fixed B, and yields noisier fitness estimates that systematically bias outer-loop selection toward smaller structures, even without explicit sparsity regularisation. We formalise this as a Proposition composing standard CMA-ES dimensionality results with rank-based noisy-comparison selection, predicting a stationary distribution over |T| concentrated below the dense baseline. To test the mechanism we introduce \textscEvoluCMAES, a minimal two-loop framework whose only domain-specific component is the topology-mutation operator. Across two unrelated settings, the same fixed apparatus produces low-density solutions: in pruning of UCI tabular models, the equilibrium density settles at \rho \approx 0.13, roughly an order of magnitude below dense; in BB84 eavesdropping, discovered Eve circuits use 4--6 gates from an arbitrary-depth search space and approach the analytical Pauli-channel cloning bound under bit-flip noise. When a specific target sparsity is required for deployment, adapting only the mutation operator suffices: at 95% sparsity on the UCI suite, mean accuracy reaches 0.972, exceeding the strongest one-shot baseline at 0.946. As a downstream consequence of the bound, the \textscEvoluCMAES library is small enough to serve as a tractable action set for an off-the-shelf RL agent: a single Rainbow-DQN with feasibility masking and matched hyperparameters recovers \approx 99% of the dynamic-programming optimum on adaptive eavesdropping for both BB84 and E91/DIQKD.
Abstract: Large Language Models (LLMs) often generate factually incorrect outputs, commonly termed hallucinations, that undermine trust and limit deployment in high-stakes settings. Existing hallucination detection methods typically require multiple forward passes, or access to model internals. In this work, we provide theoretical background and empirical evidence that the of token-level entropies, beyond the mean captured by perplexity or length-normalised entropy, serves as a fingerprint of hallucination, with distributional shape and tail behaviour carrying substantial independent signal. We formalize hallucination detection as a statistical hypothesis test and propose the ), a lightweight algorithm requiring only a single forward pass and black-box access to token logits. CES combines the mean signal with the maximum signal of the generated entropy through a calibrated reference CDF, producing scores that are directly comparable across models and tasks. We establish finite-sample calibration guarantees via a novel random-length Dvoretzky--Kiefer--Wolfowitz inequality, and also prove that CES detects hallucinations with probability converging to one exponentially fast in the generation length. Across eight QA benchmarks and ten generator models spanning open-source and API access models, CES achieves the highest detection performance among all single-pass black-box methods while providing formal error guarantees that existing heuristics lack. Remarkably, CES is statistically indistinguishable from multi-sample methods that require far greater computational cost, closing the gap between lightweight and expensive detection and making it suitable for real-time, large-scale deployment.
PaperID: 4901, Poster
Abstract: The reliability of any long-running system degrades over time: databases accumulate stale indices, software accrues technical debts, and human memory fades with age. Memory-enabled agents are no exception: Even with frozen weights, their system state continues to change as they accumulate context across sessions, and their ability to store, retrieve, and apply knowledge deteriorates in ways that standard snapshot evaluation cannot capture. Recent benchmarks have begun to measure static degradation over long-horizon tasks, yet they do not diagnose routine operational events reshape it. In this work, we introduce AgingBench, a longitudinal reliability benchmark suite organized around four aging mechanisms: via counterfactual analysis, identifying whether degradation originates in the write, retrieval, or utilization stage of the agent’s memory pipeline. We evaluated across 7 scenarios, 14 models, various memory policies and agent frameworks, ranging from a fully controlled runner to practical automatic agents such as Claude Code. Over ~400 runs across 8-200 sessions, we find aging is not one-dimensional: it can be invisible to behavioral tests while silently decaying; structurally sharp, with a single model shifting from perfect to zero accuracy on derived-state tracking; and stage-dependent, as strong models fail not at writing but at reusing their memory. Our findings demonstrate that as agents take on longer operational lifetimes, understanding the internal structure of how they age is as important as measuring how well they perform on day one.
Abstract: Recently, Masked Diffusion Models (MDMs) have shown promising potential across vision, language, and cross-modal generation. However, a notable discrepancy exists between their training and inference procedures. In particular, MDM inference is a multi-step, iterative process governed not only by the model itself but also by various schedules that dictate the token-decoding trajectory (e.g., how many tokens to decode at each step). In contrast, MDMs are typically trained using a simplified, single-step BERT-style objective that masks a subset of tokens and predicts all of them simultaneously. This step-level simplification fundamentally disconnects the training paradigm from the trajectory-level nature of inference, leaving the inference schedules never optimized during training. In this paper, we introduce Co-GRPO, which reformulates MDM generation as a unified Markov Decision Process (MDP) that jointly incorporates both the model and the inference schedule. By applying Group Relative Policy Optimization at the trajectory level, Co-GRPO cooperatively optimizes model parameters and schedule parameters under a shared reward, without requiring costly backpropagation through the multi-step generation process. This holistic optimization aligns training with inference more thoroughly and substantially improves generation quality. Empirical results across four benchmarks-ImageReward, HPSv2, GenEval, and DPG-Bench-demonstrate the effectiveness of our approach.
PaperID: 4903, Poster
Authors: Artem Betlei, Mariia Vladimirova, Victor Girou, Thibaud Rahier
Abstract: Treatment assignment problems arise wherever limited budget must be allocated to heterogeneous users, with applications ranging from personalized recommendations to online advertising and healthcare. In such settings, individuals exhibit heterogeneous responses to different treatments, making it essential to learn cost-aware personalized treatments. This paper introduces the Cost per Unit Value Equalization Tree (CUVET) algorithm, a novel treatment assignment approach that partitions the user space. Under a diminishing-returns (power-law) assumption, CUVET solves the within-cohort allocation problem by equalizing the marginal cost per unit value across each user group, yielding a closed-form cost-aware treatment assignment suited for large-scale industrial deployment. We also release CUVET-policy, an 86.7-million-impression public benchmark derived from real-world industrial A/B tests, providing an open-source evaluation framework for decision-focused learning. Across a multi-method benchmark suite, the Bayesian variant of CUVET achieves the largest cost-feasible value uplift on the public MT-LIFT dataset +11.1% under LP, vs. at most +1.3% for all baselines, while on CUVET-policy, where treatment effects are sub-percent, CUVET variants are the most efficient and strongest cost-feasible methods.
Authors: Jeonggi Kwak, Sho Kagami, Yuki Ono, Kwang Moo Yi
Abstract: Foundational visual features such as DINO have played a critical role across modern computer vision, and have recently become key components in multi-view feed-forward geometry estimators. In this work, we demonstrate that by re-distilling these multi-view models-their internal knowledge of 3D geometry-into a single-view estimator, we can obtain enhanced 3D consistent foundational features. Our key idea is to construct a multi-view teacher by fusing pretrained 2D foundation features with multi-view geometric features, and refining the fused representation with a discriminative ranking objective. Through our discriminative distillation framework, we enforce the learned features to be both 3D consistent and locally distinctive, while keeping them aligned with the feature space of the original foundation model to preserve the semantic structure of the pretrained representation. Consistency and local discriminability are critical for 3D computer vision problems such as forming semantic and geometric correspondences across images. To demonstrate the effectiveness of our method, we perform comprehensive experiments spanning multiple angles: direct feature analysis, dense prediction transfer, and explicit 3D lifting and rendering. Across these evaluations, our method consistently produces stronger 3D-aware foundation features that improve multi-view consistency and local discriminability while preserving the semantic transferability of the original representation.
PaperID: 4905, Poster
Abstract: Zeroth-order optimization (ZOO) has emerged as a promising approach for fine-tuning Large Language Models (LLMs) when gradient access is unavailable or memory-prohibitive. However, ZOO relies on estimating gradients via finite differences between two function evaluations at perturbation scale \mu, making it inherently noisier than its first-order counterpart. This raises a fundamental question: under what conditions does ZOO yield reliable gradient estimates? In this work, we provide a rigorous variance-theoretic answer through a generic noise-amplification framework and demonstrate that the reliability of ZOO gradient estimations is structurally determined by whether the two finite-difference evaluations are performed on matched data. Mismatched evaluations, as induced by on-policy Reinforcement Learning (RL), produce a 1/\mu^2 oracle-noise amplification that sets an irreducible noise floor, whereas matched evaluations, as realized by offline optimization methods such as Boltzmann Targeted SFT (BOLT), eliminate this amplification entirely via noise cancellation. We further derive convergence guarantees with explicit constants that cleanly separate the two regimes. Controlled experiments on math reasoning tasks validate these findings, using ZooBOLT and ZooGRPO --- zeroth-order adaptations of BOLT and GRPO --- as representatives of the matched- and mismatched-data regimes respectively. On GSM8K with Qwen2.5-1.5B, ZooBOLT recovers 96.1% of first-order BOLT's performance gains, while ZooGRPO exhibits near-zero learning despite its first-order counterpart achieving strong results. This qualitative separation persists across harder benchmarks (MATH), larger model scales (Qwen2.5-7B), and different model variants (DeepSeek-R1-Distill-Qwen-1.5B). Additionally, ZOO methods reduce peak memory to near-inference levels, validating its practicality as a memory-efficient fine-tuning paradigm.
Authors: Tim Woydt, Paul-David Zuercher
Abstract: Critical sequential decisions are rarely single-timescale: a strategic decision causally shapes the context in which every subsequent tactical choice is made; standard bandit and reinforcement-learning theory does not capture this causal coupling between timescales. We formalise the problem class as (NCCBs), a hierarchical SCM where each level's action sets the next level's context distribution, and propose (NCTS), which draws one mechanism-factorised belief per episode and acts recursively under it. Our main theoretical result is a causal PAC-Bayesian excess-risk bound that certifies any candidate deployment policy from historic data alone, off-policy and anytime, answering the deployment question: Experiments on a hierarchical SCM show that, against a matched RFF-GP joint regression on the same function class, the factorised SCM-mechanism posterior transfers , a safe-deployment method: each timescale flips from a legacy controller to NCTS when gains can be certified, independently of the others.
Authors:
Zhaowei Wang, Lishu Luo, Haodong Duan, WeiWeiLiu, Sijin Wu, Ji Luo, Shen Yan, Shuai Peng, Sihang Yuan, Chaoyi Huang, Yi Lin, Yangqiu SongAbstract: Long context is becoming a core capability of modern large vision-language models (LVLMs), enabling sustained context management across long-document understanding, video analysis, and multi-turn tool use in agentic workflows. However, practical training recipes remain under-documented, such as how to construct and mix long-context data. In this work, we present a systematic study of long-context continued pre-training for LVLMs, extending a 7B LVLM from 32K to 128K context with extensive ablations on long-document data. We find that: 1) synthesizing long-document VQA data provides effective and diverse long-context supervision, covering tasks from information extraction to numerical reasoning; 2) retrieving relevant evidence remains the primary long-context bottleneck, favoring retrieval-heavy mixtures with a small amount of reasoning data to preserve task diversity; 3) surprisingly, pure long-document VQA data largely preserves short-context capabilities, suggesting that instruction-formatted long data lessens the need for short-data mixing. Instantiating these findings, we obtain MMProLong by continuing training from Qwen2.5-VL-7B, improving long-document VQA performance by 7.11 points under only a 5B-token budget. More importantly, MMProLong generalizes beyond its 128K training window, maintaining strong performance at 256K and 512K without additional training. More broadly, it also transfers to webpage-based multimodal needle-in-a-haystack tasks, long-context vision-text compression, and long-video understanding without task-specific supervision. Overall, our study provides a practical LongPT recipe and an empirical foundation for advancing the next generation of LVLMs.
Abstract: We revisit the foundations of fairness and its interplay with utility and efficiency in supervised learning settings where training labels are richer than binary outcomes, such as risk estimates (probabilities), individual types, or rankings. We introduce new notions of accurate classification rates for subgroups in the population, defined by comparing the induced positive-classification rates to those under the Bayes-optimal rich predictor. Our main contributions are computational impossibility results: we show that simultaneously achieving these rate-accuracy guarantees and natural desiderata such as calibration or loss minimization is, in some cases, computationally infeasible, even when training examples are labeled by the Bayes-optimal rich predictor. Unlike prior impossibility results in this area, these desiderata are simultaneously satisfied by the Bayes-optimal predictor, and each can be achieved efficiently in isolation.
PaperID: 4909, Poster
Abstract: Language agents are moving from single-user assistants to shared collaborators in workspaces, forums, and group chats. A shared agent observes interleaved messages from multiple users, maintains persistent context, and executes tool actions, creating a boundary problem absent from standard single-user agents: the agent must decide not only whether an instruction is safe, but which user or task it is allowed to govern. We identify cross-user poisoning (CUP), where an adversary injects a message into shared context that is later applied while the agent serves a benign user, causing unauthorized actions or responses outside the instruction's intended scope. We validate CUP on two deployed multi-user agents, Continua and ElizaOS, and introduce MURMUR, a framework for evaluating shared-context agents under concurrent multi-user interactions. Across Slack, Workspace, and Airline domains, CUP achieves high attack success, persists across later interactions, and remains substantially more effective than matched prompt-injection attacks. We evaluate boundary-scoping defenses and find that context summarization alone is insufficient, while task clustering and provenance prompting substantially reduce non-adaptive propagation. These results show that robust multi-user agents require explicit user/task scoping rather than generic input filtering or context compression.
PaperID: 4910, Poster
Abstract: LLM agents increasingly rely on reusable skills for multi-step reasoning, tool use, and decision making. Yet most existing approaches represent a skill as a prompt, module, or direction in representation space. This view can be insufficient for agentic settings: (1) the behaviour induced by a skill can depend on the current context, and (2) composing two skills can yield order-dependent rollouts. We formalise a skill realization as a local map from continuous intervention coordinates to a readout representation. Its differential defines the context-dependent controllable distribution, and the residence-space norm induces a metric on this distribution. Skill objectives then define local vector fields on these controllable directions. Their \emphLie bracket captures the order-dependent component of skill composition, revealing behaviour that cannot be represented by a single context-independent direction. Our central measurement is a finite-difference skill commutator, which estimates the non-commutative component of two-skill composition from paired ordered rollouts. Unlike standard Lie-bracket estimation, the relevant vector fields are induced by skill objectives and observable ordered rollouts rather than given analytically. We therefore reconstruct the commutator from paired ordered rollouts, preserving the direction of the order-dependent change in representation space. We instantiate the framework on frozen LLM agents using activation-based and prefix-based skill realizations. Same-skill activation and prefix Jacobians are more aligned than cross-skill pairings (Cohen’s d=-0.76/-0.26), but they assign different control costs; separately, the measured two-skill interaction predicts ordered-composition gaps (Spearman \rho_S=0.918/0.889). High-commutator skill pairs also yield non-saturating reachability signals, with positive spectral-entropy gains and principal-angle novelty, while synthetic bracketed systems validate the finite-difference estimator. Together, these results show that skills form context-dependent geometric objects whose compositions leave measurable vector-valued non-commutative residuals.
PaperID: 4911, Poster
Abstract: Retrieval-Augmented Generation (RAG) improves the factual accuracy of large language models (LLMs) by grounding responses in external evidence. However, when retrieved context conflicts with models' internal parametric knowledge, LLMs may still generate answers that contradict the provided evidence, posing a key challenge to contextual faithfulness and the reliability of RAG systems. Motivated by recent mechanistic findings on the distinct propagation and progressive accumulation of parametric and contextual signals in LLMs, we propose Conflict-Suppressed RAG (CSRAG), a simple, training-free, decoding-time framework for resolving knowledge conflicts. CSRAG biases generation toward retrieved evidence by suppressing tokens associated with parametric knowledge while boosting tokens from the context via two complementary logits processors. Experiments on six challenging faithfulness benchmarks demonstrate that CSRAG consistently achieves state-of-the-art or near state-of-the-art performance across multiple backbone LLMs, while remaining fully training-free and lightweight.
PaperID: 4912, Poster
Authors: Solhee Hwang, Jangho Kim
Abstract: Large vision-language models (LVLMs) rely on increasingly long visual contexts, making the KV cache a major inference-time bottleneck. Existing KV cache compression methods typically use pre-decode signals such as prefill attention, visual redundancy, objectness, or layer-wise allocation, but these criteria do not directly indicate which visual tokens in the KV cache will be reused by future answer tokens. We introduce Q-ViK, a question-conditioned visual utility framework for LVLM KV cache eviction. During offline training, we run full-cache decoding and aggregate answer-to-visual attention over cached visual tokens, yielding a privileged that reflects the visual KV entries actually used during generation. We use this signal to train a lightweight question-conditioned scorer that predicts visual cache utility from prefill representations alone. At inference time, Q-ViK requires neither full-cache lookahead nor generated traces: it preserves textual KV entries and evicts low-utility visual KV entries with a single post-prefill scoring step. Experiments on seen and held-out multimodal benchmarks show that Q-ViK preserves task-relevant visual evidence in the KV cache more effectively than pre-decode saliency, especially under aggressive KV cache compression. Code will be available at
Abstract: Spiking Neural Networks (SNNs) are a promising framework for event-driven temporal processing. Prior work has improved temporal modeling through richer neuron dynamics and network-level mechanisms such as recurrence and delays, but it remains unclear how individual spiking neurons should specialize within a network. In this work, we introduce FiTS, a spiking neuron that factorizes temporal computation within each neuron into Frequency Selectivity (FS) and Temporal Shaping (TS). The FS module parameterizes each neuron's target frequency as the maximizer of its subthreshold magnitude response, while the TS module reshapes when frequency components contribute to membrane voltage accumulation through group-delay modulation. On auditory benchmarks where frequency selectivity and timing are central to the input structure, FiTS consistently improves over a plain Leaky Integrate-and-Fire (LIF) baseline in simple feedforward SNNs without recurrence or network-level delays, while remaining competitive with strong temporal SNN baselines. Beyond accuracy, the learned target frequencies and group-delay shifts provide interpretable neuron-level summaries of the frequency and timing organization learned within the network.
Authors: Amelie Knecht, Lucas Florin, Thilo Hagendorff
Abstract: Large reasoning models (LRMs) sometimes note in their chain of thought (CoT) that they may be under evaluation. Researchers worry that this verbalised evaluation awareness (VEA) causes models to adapt their outputs strategically, optimising for perceived evaluation criteria, which, for instance, can make models appear safer than they actually are. However, whether VEA actually has this effect is largely unknown. We tested this across open-weight LRMs and benchmarks covering safety, alignment, moral reasoning, and political opinion. We tested this both on-policy, sampling multiple CoTs per item and comparing those that spontaneously contained VEA against those that did not, and off-policy, using model prefilling to inject evaluation-aware sentences where missing and remove them where present, with subsequent resampling. VEA has limited effect on model behaviour: injecting VEA into CoTs produces near-zero effects (ω≤0.06), removing it causes small shifts (ω≤0.12) and spontaneously occurring VEA shifts answer distributions by at most 3.7 percentage points (ω≤0.31). Our findings call for caution when interpreting high VEA rates as evidence of strategic behaviour or alignment tampering. Evaluation awareness may pose a smaller safety risk than the current literature assumes.
PaperID: 4915, Poster
Abstract: Can a large language model be \emphtrained to tolerate silent hardware faults, and if so, \emphwhere in the networkshould the consistency signal be enforced? We answer both questions and find that the second one matters more than the first. Injecting tile-level GEMM faults during training and adding a KL consistency term on the output distribution reduces fault-induced perplexity degradation from 12.9% to 0.21% on GPT-2 Small---a 60× reduction (two-seed mean). The placement of the consistency target is decisive: a direct hidden-only objective at the fault site can worsen degradation to 32.7%, and fair bounded hidden controls with corrupted CE improve substantially but still trail output KL in an 80k-step screen (1.74--1.76% vs.\ 0.84%). A conditional constraint hierarchy and gradient-alignment probes explain why output-level consistency is more permissive and easier to optimize than direct early hidden-state consistency, without requiring that every hidden-alignment formulation fail. The output-over-hidden advantage persists at 8B scale, the GPT-2 ordering reproduces on held-out corpora, and generic robustness methods---SAM, R-Drop, and SAF---fail to close the gap. Pairing the trained model with simulated algorithm-based fault detection yields a bit-flip-to-erasure interface with a 1.10× perplexity ratio, where unprotected inference produces NaN.
Abstract: Optimizing 3D shapes within the latent spaces of deep generative models is fundamental to computer assisted engineering, yet remains prone to a critical failure mode we term : the tendency of gradient-based optimization to move latent vectors away from the manifold of valid shapes. This problem is exacerbated in state-of-the-art 3D shape generative models that operate in increasingly high-dimensional latent spaces where valid shapes occupy a vanishingly small fraction of the full space. Existing mitigation strategies, including latent regularization and flow-matching approaches, either sacrifice expressiveness, demand a difficult trade-off between objective guidance and generative fidelity that remains prone to manifold drift, or are computationally infeasible to scale to modern, large-capacity 3D shape models. We introduce a novel optimizer-corrector framework that alternates between gradient steps for objective minimization and guided flow matching to drive the latent state back to the valid shape manifold. By decoupling objective minimization from flow-based correction, optimizing freely and correcting strictly, this alternating design avoids inherent trade-offs, preserving geometric validity without sacrificing expressiveness while remaining computationally feasible on modern 3D shape models. We demonstrate its effectiveness across generative priors of varying complexity, from simple vector latent spaces to large-scale architectures across a variety of downstream optimization tasks, including aerodynamic drag reduction and object compliance optimization.
PaperID: 4917, Poster
Abstract: Multi-robot embodied systems are commonly expected to collaborate by coordinating their behaviors: agents may communicate, share experience, divide tasks, avoid conflicts, or jointly decide what to do next. This paper asks a complementary question: can collaboration emerge before decision making, by first building a shared understanding of the environment? We introduce Shared Environmental Understanding, a policy-agnostic representation layer that aggregates partial egocentric observations from multiple robots into a common latent world representation. This formulation shifts the starting point of collaboration from agent-centric traces to environment-centric representation, allowing different robots and downstream policies to benefit from shared world knowledge without being forced into a centralized controller. We instantiate this idea with SEER, a shared environment encoding and retrieval framework that learns to turn multi-robot, multi-view observations into reusable navigation context. SEER is designed not as a new navigation policy, but as an upstream environmental understanding module that can be plugged into heterogeneous agents, including both specialized VLN models and general vision-language model agents. We validate SEER through a comprehensive set of representation-level and task-level studies, including latent inverse dynamics analysis, standard VLN evaluation, controlled paired multi-robot navigation, comparisons against action-, history-, and policy-sharing alternatives, and real-world robot deployments. Across these settings, SEER preserves single-agent navigation ability while improving individual and collaborative success in multi-robot scenarios. These findings support a simple but underexplored hypothesis: for embodied cooperation, robots may benefit from sharing the environmental understanding before sharing the policy.
PaperID: 4918, Poster
Abstract: Certifying authorities for safety-critical software audit recovered trace matrices module by module, not on average. Every existing distribution-free calibration delivers a marginal certificate, and we show the gap to the auditor's quantity is structural and unbounded. Under within-block exchangeability with cross-block heterogeneity in the conditional null score distribution, every procedure controlling marginal \mathrmFDR at level \alpha on a bipartite trace matrix with K source modules admits worst-module false discovery proportion at least 1-\alpha-O(1/K) with probability at least 1-2/K. We name this the block catastrophe and close it with \mathrmCERTRA, which wraps any black-box scorer, computes a within-block rank-indicator conformal e-value, and applies e-\mathrmBH separately within each source module. The construction yields finite-sample per-module \mathrmFDR control under arbitrary intra-module dependence, with matching power up to a rank-quantization factor. We release IndTrace-2026, a 12,406-link composite across aerospace, automotive, and medical-device specifications with professional ASIL/DAL/SIL annotation (Cohen's \kappa=0.86). \mathrmCERTRA holds worst-block \mathrmFDP at 0.118 where split-conformal drifts to 0.413, retains 0.741 of true links versus 0.572 for the strongest fair-operating-point baseline, and reduces audit deficiency 3.3-fold. Per-module distribution-free calibration is what an industrial audit can sign.
PaperID: 4919, Poster
Authors: Shuoqiu Duan, Jiasen Gao, Xiaoliang Chen
Abstract: While spatiotemporal modeling has emerged as an effective approach in multimodal learning, it struggles to resolve the ``spectral dilemma'' where distinct tasks demand contradictory frequency views. For example, deception detection requires global macro-patterns, while intent recognition depends on local micro-transients. To address this, we propose Uni-Cheb, a unified spectral operator centered on the Learnable Chebyshev Filter (LCF). Designed as a basis-agnostic, plug-and-play component, LCF maintains a consistent mathematical form to adaptively modulate frequency components regardless of the underlying spectral transform. By leveraging the minimax property of Chebyshev approximation, LCF performs elastic spectral resampling to focus on task-relevant bands. Guided by task-specific physical priors, the LCF seamlessly integrates with the Discrete Fourier Transform (DFT) to amplify global physiological rhythms or the Discrete Wavelet Transform (DWT) to capture localized semantic shifts, all while maintaining a negligible computational footprint. Furthermore, it serves as an effective spectral preconditioner to mitigate distribution shifts in cross-domain transfer learning. Extensive experiments across multiple benchmarks demonstrate that Uni-Cheb acts as a universal enhancer, consistently improving state-of-the-art baselines. By rendering frequency modulation both task-adaptive and highly efficient, Uni-Cheb establishes a robust, operator-level solution to the spectral dilemma.
Authors: Camille Touron, Gabriel Cardoso, Julyan Arbel, Pedro Rodrigues
Abstract: Compositional score-based approaches to simulation-based inference (SBI) approximate the posterior over a shared parameter given n independent observations by aggregating individually learned posterior scores: currently, there are two main propositions of such methods (Geffner et al., 2023; Linhart et al., 2026). As the resulting composite score does not correspond to the score of any distribution along the forward diffusion path of the true multi-observation posterior, sampling from it via a reverse SDE leads to an irreducible bias. Annealed Langevin dynamics provides a principled alternative: it treats the composite score as the genuine score of a sequence of tractable bridging densities and samples from them in succession. When properly tuned, it could lead to a controllable bias. However, its hyperparameters, namely step sizes, the number of steps per level, and the number of annealing levels, have so far been chosen empirically. We derive Wasserstein bounds for annealed Langevin with approximate scores and translate them into explicit decision rules for these hyperparameters that guarantee a prescribed sampling accuracy, while highlighting different theoretical aspects of each composite score formulation. In the Gaussian setting, we obtain closed-form expressions for all relevant quantities and prove that the bridging densities of Linhart et al. (2026) consistently admit larger step sizes and require fewer total Langevin steps than those of Geffner et al. (2023). Furthermore, we show empirically that the tuning obtained in the Gaussian setting generalizes to more complex problems, thus providing a well-understood and theoretically grounded starting point for practitioners using compositional score-based approaches.
Authors:
Chengzhen Yu, Canran Xiao, SiYuan Ma, Yang LiuAbstract: Vision--language alignment powers open-vocabulary recognition, retrieval, and LVLM grounding, yet natural captions are often underspecified, making similarity brittle and overly confident under paraphrase and omitted details. We aim to learn representations whose matching is stable across caption views and whose confidence reflects how strongly text constrains an image. We propose Text as Partial Constraint (TPC), a core--residual alignment framework that treats multi-view captions as incomplete supervision: it distills a consensus semantic core as the alignment target, learns a single-view core predictor for standard inference with one query, and explicitly discourages vision--language similarity from depending on the orthogonal ``unsaid'' residual. An uncertainty-aware contrastive objective further softens alignment when caption views disagree, reducing overconfident updates under weak language constraints. Across zero-shot recognition and adversarial robustness, TPC achieves 81.42/64.05 Top-1 clean/robust accuracy on ImageNet and 76.19/52.03 on an Avg-14 transfer suite, while improving LVLM transfer with 85.16 POPE F1 and 59.57 OKVQA accuracy under an LLaVA-1.5-7B stack. These results suggest that modeling text as a partial constraint is a practical and principled route to more reliable vision--language representations under underspecified language supervision.
Abstract: Streaming 3D reconstruction under a strict constant-memory budget hinges on how the recurrent state is updated as the stream evolves. We profile TTT3R-style per-token gates across five benchmarks and discover a structural bottleneck: the gate is intrinsically bounded in magnitude (median 0.31; never exceeding 0.6) and nearly frame-invariant, yielding an effective memory horizon of only ~3 frames per state token, which serves as the structural origin of long-sequence drift. We trace this to a missing axis: existing inference-time methods modulate updates only at the per-token, intra-frame level, while the orthogonal frame-level question of \emphhow strongly each frame should contribute to the state has been silently treated as content-independent. We close this gap with a scalar frame-level gate \alpha_t \in (0, 1] derived in closed form from frame-to-frame changes of internal features---a continuous relaxation of classical SLAM keyframe selection that requires no parameters, no training, and no extra forward pass. Across six benchmarks spanning camera pose, video depth, and 3D reconstruction at sequence lengths up to 4,541 frames, our gate cuts ATE by 51% on long TUM-RGBD pose sequences, reduces AbsRel by 12.8% on Bonn video depth, and on KITTI long-sequence pose estimation surpasses LongStream and Keyframe-VO which are specially constructed for long sequences via retraining and policy learning, while retaining strictly constant memory at zero training cost.
PaperID: 4923, Poster
Abstract: Expert scientists rarely solve genuinely novel problems through parametric recall alone; instead, they reason by analogy to structurally related prior cases and adapt their solution procedures to new contexts. Yet current multimodal large language models (MLLMs), even when equipped with retrieval, computation, and visual tools, still struggle with this imitation-driven form of reasoning. We further observe a counterintuitive negative result: simply prepending expert problems and full solution trajectories as context does not improve performance on novel scientific tasks, and can even slightly degrade it. Motivated by this observation, we argue that scientific reasoning should shift from passive answer conditioning to an imitation-driven paradigm, in which models actively seek relevant precedents and reuse their solution structures. We instantiate this idea with MimicAgent, a tool-augmented scientific reasoning agent that retrieves structurally related exemplars, interprets their reasoning trajectories, and transfers their solution patterns to the target problem through multi-step reasoning and tool interaction. To train this behavior under sparse rewards, we introduce ARIS (Adaptive Reinforcement-Imitation Switching), which unifies on-policy reinforcement learning with online expert imitation. When all rollouts for a prompt fail, ARIS replaces degenerate reinforcement updates with expert-generated demonstrations, yielding an implicit imitation-to-reinforcement curriculum without requiring a cold-start stage or fixed-ratio offline mixing. To support training and evaluation, we further introduce SciExplore-Bench, a bilingual multimodal benchmark for open-ended experimental scientific reasoning, paired with a manually curated repository of analogical exemplars. The substantial gains and impressive generalization across multiple benchmarks suggest that imitation serves both as a global paradigm for exemplar-guided reasoning and as a local mechanism for stabilizing reinforcement learning.
Abstract: We study fixed-policy evaluation for finite Markov chains that may be reducible and periodic. Classical evaluation methods with gain and bias decomposition are not always diagnostic: the gain records only invariant Ces\`aro averages, while persistent phase-dependent behavior is absorbed into the bias together with genuinely transient effects. We identify the real peripheral invariant subspace \mathcalK(P) of the transition matrix P as the source of this ambiguity. Quotienting by \mathcalK(P) is the minimal exact quotient that removes all non-decaying modes and makes the remaining dynamics strictly stable. After choosing a gauge projection \Pi with kernel \mathcalK(P), the reward admits a unique decomposition r = g_\Pi^\star + (I-P)v_\Pi^\star, where g_\Pi^\star is a persistent regime profile and v_\Pi^\star is a gauge-fixed transient component. An exact comparison with classical normalized gain and bias shows that the new pair reallocates the same information so that all persistent modes are represented in g_\Pi^\star and v_\Pi^\star is transient. This decomposition reconstructs finite-horizon returns, recovers statewise average reward, admits a transient-cost interpretation, and yields a stable estimator under a generative model.
Abstract: Online bilevel optimization (OBO) has emerged as a powerful framework for many machine learning problems. Prior works have developed several algorithms that minimize the standard bilevel local regret or the window-averaged bilevel local regret of the OBO problem, but the optimality of existing regret bounds remains unclear. In this work, we establish optimal regret bounds for both settings. For standard bilevel local regret, we propose an algorithm with adaptive iteration strategy that achieves the optimal regret \Omega(1+V_T) with at most O(T\log T) total inner-level gradient evaluations. We further develop a fully single-loop algorithm whose regret bound includes an additional gradient-variation terms. For the window-averaged bilevel local regret, we design an algorithm that captures linear environmental variation through a novel window-based analysis and achieves the optimal regret \Omega(T/W^2). The algorithm also supports an efficient single-loop structure, achieving an O(T/W) regret bound with O(WT) total gradient evaluations. Experiments validate our theoretical findings and demonstrate the practical effectiveness of the proposed methods.
PaperID: 4926, Poster
Authors:
Shuo Zhang, Qian Tian, Chen Gao, LongzhaoGuo, Youfang LinAbstract: Recently, Noise2Noise provides a powerful principle for learning image denoising without clean data. However, zero-shot methods for 2D images often rely on local similarity or non-local self-similarity to construct image pair for training, which is unreliable in complex scenes with highly textures. Different from 2D image, 4D Light Field (LF) captures multiple views for one specific scene. The multiple views with independent noises provides make it possible to find image pair across views. In this paper, we exploit the inter-view correspondence in LF and propose a novel way to implement the Noise2Noise in 4D LF. Specifically, we design a Matching2Matching (M2M) framework to construct matching views for each view in LF. This framework generates geometrically consistent training samples by searching for corresponding patches along the epipolar line, crucially enforcing both visible and disparity consistency constraints. The constructed views enable the network to train the denoised network in a zero-shot way without any other data. Experimental results show that our method effectively suppresses noise while outstandingly preserving the consistency and fine details. Our M2M framework significantly outperforms existing zero-shot methods in both quantitative and qualitative evaluations on synthetic and real-world noise, demonstrating superior generalization ability without relying on external training datasets.
PaperID: 4927, Poster
Abstract: Negative guidance of diffusion models generates samples that avoid an undesired condition, specified either by reference samples or by a conditional score alongside an unconditional model. Existing methods inject repulsive terms derived from local information at each sample, so the resulting sampler does not target an explicit density, and its bias cannot be systematically reduced by increasing computation. We instead present a negative-guidance method that samples from an explicit target density, which uniformly handles both reference-sample and score-based specifications and suppresses probability mass near the undesired condition with an avoidance strength that parametrizes the target density itself. Realizing this target requires estimating, at each diffusion timestep, how likely each sample is to belong to the undesired condition. We cast this estimation as a Positive--Unlabeled learning problem that admits lightweight online training along the diffusion trajectory, using condition samples as positives and diffusion particles as unlabeled data. Because the guided particles drift from the unconditional marginal, we further correct the residual mismatch via Sequential Monte Carlo, yielding a sampler that is asymptotically exact in the particle limit and that turns any existing negative-guidance method into its proposal kernel. Empirically, our method approaches the ground-truth target on Gaussian mixtures and Pareto-dominates existing methods on trade-offs between avoidance and generation quality, distributional bias, and diversity, including in an image-generation setting with a pretrained score network.
PaperID: 4928, Poster
Abstract: Diagnosis of psychiatric disorders from functional magnetic resonance imaging (fMRI) is hindered by low-resource phenotypes, high inter-subject variability, and weakly interpretable multimodal fusion. We propose a prior text-informed gated attention framework (PTGA) that converts region-of-interest (ROI) level neuroimaging measures into statistically grounded textual priors via percentile-based soft binning. PTGA combines hemodynamic response function (HRF) aware spatiotemporal encoding with static graph topology modeling, and unifies variable-length fMRI time series through masked global temporal aggregation. It further uses text semantics as an inductive bias for prior-guided gated fusion, improving robustness under cross-modal noise. We compare PTGA with eight strong graph and multimodal baselines under a unified protocol. PTGA achieves up to 10.46% and 8.01% absolute improvements in autism spectrum disorder (ASD) and major depressive disorder (MDD) diagnosis, respectively, outperforming SOTA baselines. The learned gating also provides biologically plausible evidence: MDD related signals are concentrated in sensorimotor/visual/default-mode systems, while ASD related evidence is centered on the cerebellar network with joint contributions from default-mode, limbic, and dorsal-attention networks.
PaperID: 4929, Poster
Abstract: Autoregressive video diffusion models (AR-VDMs) enable real-time, long-horizon video generation but degrade rapidly once the inference horizon outruns the training horizon, due to compounding exposure bias. To stabilize long rollouts, recent methods heuristically adopt frame sink, which keeping initial frames as a persistent KV cache context, inspired by attention sinks in streaming LLMs. We revisit this design and find that the mechanism by which initial frames stabilize long-range consistency in AR-VDMs is not an LLM-style attention sink; instead, consistency is anchored through a sparse subset of attention heads, while the majority of heads predominantly capture local temporal dynamics. This observation reveals an over-conditioning pathology in existing frame-sink strategies: by universally exposing every head to the pinned initial-frame context, the model becomes prone to collapsing motion toward the low-dynamic modes of the teacher during model distillation. We propose Disen-Forcing, a distillation-time plugin for AR-VDMs that disentangles and reinforces this head specialization through disentangled context routing and motion-disentangling regularization. On long-horizon video generation benchmarks, Disen-Forcing simultaneously improves semantic consistency and motion dynamics of base AR-VDMs: a trade-off that sink-based methods fail to balance.
PaperID: 4930, Poster
Abstract: Adapting language models to new domains during post-training risks degrading previously acquired capabilities, a phenomenon commonly known as \emphcatastrophic forgetting. Existing mitigation methods largely treat forgetting as a consequence of excessive model drift. They therefore constrain updates by preserving old outputs, replaying old data, or regularizing changes to parameters and representations so that the updated model remains close to its previous state. We take a different view: Rather than pulling the updated model back toward its old state, we translate the new domain into the model’s pre-existing knowledge space, making unfamiliar knowledge learnable through familiar procedures. Therefore, we propose Skill Schema Transport (SST), a post-training framework that transports new knowledge into the model's existing skill space. SST abstracts each example type into a skill schema, preserving its essential operations, dependencies, and verification steps while removing domain-specific surface forms. By recasting new examples as instances of reusable procedures, SST turns knowledge update from trajectory memorization into procedural reuse. In continual domain adaptation settings, SST consistently reduces forgetting while improving generalization to related tasks. Across two 3B instruction-tuned backbones, SST improves average final performance on streamed domains by 2.53-5.75 points over sequential post-training. It also lowers max-drop forgetting in all evaluated settings, with the largest reduction from 42.41 to 7.10 points on the Qwen-3B MATH stream, while substantially recovering the OOD degradation induced by sequential updates.
PaperID: 4931, Poster
Abstract: Gradient-based attribution methods for large language models (LLMs) have so far been limited to explaining scalar outputs - the logit of a single target token. Yet LLM predictions are inherently distributional, and many practically important questions concern the full predictive distribution: which tokens make the model uncertain? Which drive a broad range of plausible continuations? We propose Jacobian Scopes, a unified framework that fills this gap by projecting the input-to-output Jacobian onto different directions in output space via a single vector-Jacobian product, requiring only one backward pass. This yields three complementary methods: Semantic Scope attributes a specific target logit; Fisher Scope, grounded in information geometry, identifies tokens that most alter the overall shape of the predicted distribution; and Temperature Scope traces which tokens govern the model's predictive confidence. Fisher and Temperature Scopes are, to our knowledge, the first attribution methods to natively target distributional rather than pointwise features of LLM predictions. Through case studies spanning instruction following, translation, and in-context time-series forecasting, Jacobian Scopes reveal implicit political biases, uncover word- and phrase-level translation strategies, and illuminate the nearest-neighbor pattern-matching mechanisms underlying LLM forecasting. Quantitative evaluation on LAMBADA and IWSLT2017 across six leading LLMs (LLaMA-3.2, Qwen2.5, Gemma-3) confirms that Jacobian Scopes consistently match or outperform Input × Gradient and Integrated Gradients, at a fraction of the latter's computational cost.
PaperID: 4932, Poster
Abstract: Maintaining long-term geometric consistency remains a key challenge in long-horizon autoregressive video generation. Recent memory-augmented generative models address this issue by retrieving historical frames beyond the limited temporal context, but their effectiveness depends on two key design choices: what 3D-geometric evidence should represent past observations, and how memory frames should be selected from this evidence. Existing methods often rely on camera poses or field-of-view overlap, which are lightweight but too coarse to reason about pixel-wise visibility, or use explicit 3D reconstruction, which provides fine-grained evidence but is costly to maintain over long rollouts. We propose Coverage-Maximizing Retrieval-Augmented Generation (COVRAG), a depth-based memory retrieval framework that uses pretrained 3D priors to construct a target-view coverage map as lightweight 3D memory evidence. For frame selection, COVRAG maximizes residual coverage gain, iteratively retrieving frames that explain target-view regions not covered by the current context or previously selected memories. To improve scalability in long-video generation, we introduce a sliding-window depth caching for efficient geometry estimation. Experiments on RealEstate10K and DL3DV show that \modelname improves long-horizon geometric consistency while maintaining low latency compared to baselines.
PaperID: 4933, Poster
Abstract: Large EEG foundation models are pretrained on heterogeneous corpora, but two challenges remain underexplored: (1) experimental protocols provide mental-process cues that are rarely used in masked pretraining, and (2) electrode montages vary across datasets, making spatial representations montage-dependent. Existing methods typically learn a single general-purpose representation and model spatial structure implicitly, limiting their ability to exploit process-related information under montage heterogeneity. We introduce BrainPro, a process-conditioned self-supervised pretraining framework that couples topographic spatial retrieval with shared and process-associated representation learning. BrainPro maps dataset-specific montages to a universal channel-region template and retrieves channel- and region-level spatial filters to form a topographically aligned spatial basis. Over this basis, BrainPro learns a shared encoder for general EEG structure and additional affective, motor-related, and auxiliary residual branches for process-associated variation. Protocol-derived mental-process cues condition branch activation and region-weighted masked reconstruction, incorporating topographic spatial priors and process information without using downstream class labels. Across nine public BCI benchmarks, BrainPro achieves strong performance among evaluated baselines. Ablations, channel-drop analysis, encoder-configuration studies, and spatial-filter visualizations suggest that topographic spatial retrieval and process-conditioned representation learning jointly improve EEG decoding.
PaperID: 4934, Poster
Abstract: Sequence models in machine learning compress long input histories into finite-dimensional states, from recurrent networks and reservoir computing to recent state space models for long-context data. This raises a basic expressivity question: how simple can the recurrent memory be before it can no longer approximate stable causal sequence-to-sequence maps? We answer this question for fading-memory operators, where the distant past has uniformly diminishing influence. We show that every continuous causal time-equivariant fading-memory operator on a compact input set can be uniformly approximated by a state space model whose recurrence is a fixed diagonal linear contraction. The dynamics are time-invariant, input-independent, and contain no nonlinear hidden-state update or delay-line structure. The construction uses a bank of short exponential traces; although such traces do not exactly store a recent input window, a singular Vandermonde extraction approximately reconstructs the window while the remaining tail becomes negligible as delta decreases. We also prove that exact delay recovery is impossible for any fixed finite trace bank, showing that the accuracy-dependent extraction is essential. Technically, the result separates recurrent expressivity from exact memory storage, while exposing a conditioning cost in the readout. More broadly, it clarifies why highly stable, structurally simple SSM cores can still be universal approximators for stable sequence behavior.
PaperID: 4935, Poster
Abstract: Classifier-free guidance (CFG) is a standard technique for improving conditional generation in diffusion models, yet its activation schedule is usually selected heuristically. In practice, most existing CFG-based methods rely on a pre-defined activation policy throughout inference, while recent dynamic variants often depend on hand-designed schedules without a principled theoretical explanation of when conditional information is most useful. In order to understand how the informational relevance between conditioning inputs and the score function evolves over the course of generation, we explore when guidance should be activated from an information-theoretic perspective. In this work, we formulate the diffusion trajectory as a time-dependent information bottleneck and show that the effectiveness of conditional information is not uniform throughout sampling, but becomes significantly stronger after a critical stage characterized by a Fisher-information balance condition. Motivated by this insight, we propose a training-free dynamic CFG strategy for inference. Rather than prescribing a fixed schedule in advance, our method constructs a stochastic process from conditional and unconditional score statistics and activates guidance when the estimated guidance gain crosses a prescribed threshold. This yields an adaptive activation rule that aligns condition injection with the underlying generation dynamics, while also recovering interval-based guidance as a limiting special case. Extensive experiments across diverse architectures and tasks, demonstrate that our method is effective, robust, and broadly applicable. Beyond empirical improvements, our framework provides an interpretation of dynamic guidance methods and a principled foundation for inference-time control in conditional diffusion models.
PaperID: 4936, Poster
Abstract: Graph learning still lacks a broadly reusable input interface comparable to patches in vision or tokens in language. Unlike images or text, graphs are irregular, vary in size and density, and are invariant under many equivalent node labelings, making it difficult to standardize inputs across tasks and architectures. We introduce , a fixed-budget graph imaging framework that converts each graph into two complementary image-like views. provides a stable topological view by encoding an intrinsic bifiltration as a compact multipersistence surface grid, while provides a higher-bandwidth structural view by aggregating edge densities between groups defined by intrinsic node scores. Both views are permutation-invariant and come with resolution-controlled stability guarantees. Because the resulting representations are fixed-size tensors, they can be processed by lightweight 2D encoders and aligned through a supervised cross-view contrastive objective. Across graph classification and molecular prediction benchmarks, achieves strong average performance among the compared methods, improves over single-view and standard fusion variants, and extends to attribute-rich molecular graphs without changing the overall architecture. These results suggest that fixed-budget graph images offer a practical and reusable interface for graph representation learning.
PaperID: 4937, Poster
Abstract: Multi-agent Large Language Model (LLM) systems offer a way to decompose complex tasks such as coding through parallelization and context isolation, but adding agents in practice introduces inter-agent communication overhead that can offset efficiency gains. We formalize multi-agent orchestration as a graph partitioning problem that captures the communication-to-computation trade-off: task decomposition can shorten critical-path computation, while cross-agent dependencies require costly context transfer. We instantiate this view in repository-level software engineering through Cohesion-aware Coder (CoCoder), which builds dependency graphs from static analysis, isolates structural hub files, partitions the graph via community detection, and executes the partition with a dependency-aware scheduler. Across \(28\) real-world tasks on DevEval and CodeProjectEval, CoCoder Pareto-dominates sequential and file-based parallel baselines as well as Claude Code with Agent Teams, improving pass rate by up to \(14.0%\) on CodeProjectEval, achieving up to a \(2.10×\) wall-clock speedup, and reducing API cost by up to \(35%\), with the largest gains on the most dependency-dense projects. CoCoder demonstrates how dependency-aware orchestration can make parallel coding agents both theoretically grounded and practically efficient, suggesting a broader design principle for multi-agent LLM systems.
Abstract: Recent feed-forward 3D reconstruction methods have demonstrated strong performance and flexibility in efficient end-to-end scene geometry estimation from image streams. However, their reliance on visible-light appearance makes them vulnerable in dark and low-visibility environments, where RGB cues are severely degraded and geometric evidence becomes ambiguous. To address this challenge, we propose DarkVGGT, an RGB-T feed-forward geometry framework that uses physics-aware thermal modeling for robust 3D estimation in low-light scenes. DarkVGGT introduces two complementary modules. First, physics-inspired thermal factorization extracts emissive-dominant, geometry-consistent thermal cues while isolating sparse reflective residuals that may introduce geometric ambiguity. Second, geometry-shared thermal routing isolates modality-invariant geometric structures from thermal-specific patterns, selectively injecting reliability-aware structural guidance into the RGB stream. Together, these components enable accurate thermal-informed geometry estimation under degraded RGB conditions while largely preserving performance in well-lit environments. Experiments on low-visibility RGB-T benchmarks demonstrate consistent improvements in both depth and camera pose estimation over existing feed-forward geometry baselines.
Abstract: While Vision-Language Models (VLMs) have shown strong visual reasoning capabilities, their spatial reasoning abilities remain largely constrained to the observed images and text-oriented chain-of-thought. They often struggle to infer unobserved layouts, maintain cross-view consistency, and reason from alternative viewpoints when only limited egocentric observations are available. In this work, we study this problem as thinking with imagination, where a VLM actively acquires imagined visual evidence by interacting with a world simulator during reasoning. We propose Astra, an agentic spatial reasoning framework that empowers VLMs with action-conditioned visual imagination. Specifically, Astra couples Astra-VL, an RL-trained VLM policy, with Astra-WM, a Bagel-based world simulator that generates novel-view observations from context images and natural-language camera motions. To provide reliable imagined evidence, Astra-WM is trained with view consistency tuning to improve pose and content consistency across views. In the RL stage, we propose a world-simulator-in-the-loop two-phase RL curriculum to stabilize tool-use exploration and advance the model's ability to invoke the simulator only when imagined observations improve over direct answering. Experiments demonstrate that both the world simulator and the agentic policy are necessary: Astra-WM improves simulator-augmented Gemini-3-Flash on MMSI-Bench from 45.1 to 49.5, while Astra-VL improves the Qwen3-VL backbone from 29.8 to 38.8 on MMSI-Bench and from 36.8 to 42.7 on MindCube. These results show that imagined observations can provide useful spatial evidence, but effective world-model-augmented reasoning requires learning when, where, and how to imagine.
Abstract: Long chain-of-thought (CoT) reasoning improves large vision–language models, but visual information often fades during generation, limiting long-horizon multimodal reasoning. Existing methods either re-inject vision at inference or train policies for stronger grounding, but where to intervene relies on perception heuristics rather than principled gain analysis, and how local visual influence propagates remains implicit. We study this problem from an information-theoretic standpoint and derive a lower bound on the downstream visual gain of a one-step intervention, which suggests two factors: local branching room (token entropy) and downstream visual propagation potential (suffix divergence from a vision-marginalized reference). Guided by this analysis, we propose reflection-anchor policy optimization (RAPO), a GRPO-based policy optimization method that selects high-entropy reflection anchors and optimizes a chain-masked finite-window KL surrogate for downstream visual dependence. Experiments on reasoning-intensive and general-domain benchmarks show that RAPO delivers substantial gains over strong baselines across multiple LVLM backbones. Mechanism analyses further indicate that reflection anchors are enriched for visually sensitive decision points and that RAPO increases contrastive visual-dependence signals along generated trajectories.
PaperID: 4941, Poster
Abstract: Personalization systems encode long-horizon, evolving user trajectories into compressed preference states that downstream task heads consume for prediction or generation. This compression is difficult because preference trajectories contain stable long-term interests, transient short-term shifts, and episodic bursty interests. We argue that a practical route is not to design another task-specific encoder or incur costly host-encoder finetuning, but to correct the under-expressed compressed state before the downstream task head consumes it. We propose \textttREPAIR as a selective pre-head state-repair method for frozen personalization encoders. It plugs into frozen hosts, uses event representations from the forward pass, and avoids a second pass over the raw history. \textttREPAIR compares the compressed state with per-timestep event representations, estimates state-relative corrective evidence, resolves it through long-term, short-term, and episodic temporal regimes, and retains only the most corrective signals to form a compact correction. We evaluate \textttREPAIR across recommendation and generation personalization settings using MovieLens, MIND, PENS, and Amazon Reviews 2023. Across recommendation tasks, \textttREPAIR improves the evaluated frozen hosts. On MovieLens, it improves the highest-scoring frozen host in our instantiated baseline suite by +1.27/+1.54 MRR/nDCG@5; on MIND, it improves the highest-scoring frozen host by +4.41 MRR; and on Amazon product recommendation, it improves two SOTA baselines by +15.33 and +17.91 MRR. It also outperforms budget-matched post-head correctors, and the full long-, short-, and episodic repair remains the strongest temporal variant. On PENS personalized headline generation, \textttREPAIR improves subjectivity-sensitive generation most when the decoder is preference-coupled, with average PerSEval gains of +14.62% and +22.19% for two representative preference-coupled decoders. Relative to frozen predictive encoders, deployment overhead is 31--34% memory and 38--44% latency.
PaperID: 4942, Poster
Authors:
Zhuyun Yuan, Jiming Zhang, Haoyuan Sun, Jun Yin, Qing Wang, Jian Kang, Jie Li, Yongxiang LiAbstract: Human experts can accumulate experience from open-ended and complex inspection tasks, thereby continually improving their inspection capabilities. Actually, industrial hardware inspection faces challenges including open-set defect categories, complex defects (tiny, high-gloss, low-light, or blurred), high inference cost, and difficulty in accumulating experience post-deployment. Existing end-to-end inspection methods typically couple image restoration with defect recognition in a single forward pass, limiting generalization to unseen components or composite defects; cascade “restore-then-detect” approaches often rely on static preprocessing, hindering adaptive adjustment based on scenario, defect morphology, and historical failures. To address this, we propose EvoInspect, the first memory-enhanced multi-modal multi-agent framework for industrial defect object detection. It leverages a multi-modal agentic memory mechanism to distill self-evolving scenario-specific Defect Detail Restoration decisions and dynamically grounded trajectories, enabling “accumulating experience like human experts.” Specifically, (1) we introduce a dynamic Defect Detail Restoration strategy, including adaptive slicing for tiny defects, highlight suppression, low-light enhancement, super-resolution, and binarization, combined with a coarse-to-fine token-efficient grounding mechanism for efficient defect capture; (2) we design a dual-channel distillation mechanism unifying restoration and detection (perception and execution) experience, incorporating both successful and failed validation trajectories to support system self-evolution. Experiments on bearing surfaces, PCBs, solar electroluminescence (solar EL), and magnetic tile datasets show that EvoInspect consistently outperforms strong baselines in defect localization recall, inference accuracy, while demonstrating cross-scenario adaptability and offline continual improvement.
Abstract: Standard flow and diffusion pre-training matches the distribution of available data (e.g., molecules), which often covers only a small fraction of the valid design space. In generative discovery, however, one aims to sample valid new-to-nature designs, assigned negligible probability under, and thus inaccessible to, standard models fitted to the observed data. To overcome this limitation, we depart from data distribution matching and view a generative model through its generable set: the region it covers with non-negligible probability. This allows to introduce a new learning principle for out-of-distribution flow modeling: enlarging a model’s generable set to increase coverage of the valid design space. We propose Active Flow Expansion (ActFlow), a continued pre-training method that employs verifier feedback to expand a pre-trained model over new valid regions by iteratively adapting to synthetic data generated through active exploration in the learned flow representation. Theoretically, we establish to our knowledge first-of-their-kind statistical learning guarantees for out-of-distribution flow modeling, analyzing generable set expansion as a local-to-global reachability process over a learned representation. Empirically, we assess ActFlow with suitable out-of-distribution generative modeling metrics across small organic molecules, mid-sized drug-like molecules, therapeutic peptides, and protein sequence design tasks. Results show that ActFlow expands valid coverage far beyond the region modeled by the initial pre-trained model, significantly outperforming widely adopted synthetic flow pre-training methods.
PaperID: 4944, Poster
Abstract: Computing the partition function \mathcalZ = \mathrmtr(e^-\beta H) of a quantum spin Hamiltonian H on n sites is a fundamental problem in statistical physics and quantum chemistry. However exact evaluation requires O(8^n) time and O(4^n) memory making it intractable for all but the smallest systems. We propose an unbiased stochastic estimator of \mathcalZ, and more generally of \mathrmtr(f(H)) for any function f admitting a convergent power series over exponentially large, implicitly defined operators using only local random walks, avoiding both matrix materialization and matrix-vector products. Our approach combines Graph Random Features with the Hutchinson's stochastic trace estimator, and exploits the algebraic structure of quantum spin Hamiltonians to reduce the per-step computational cost to O(n + d), where d is the mean degree of the spin coupling graph, independent of the Hilbert space dimension N = 2^n. We validate the estimator empirically against exact matrix exponentiation for small system sizes, and demonstrate its applicability to system sizes where exact methods are infeasible.
Abstract: Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions. As LLMs increasingly debate, advise, and think collaboratively with humans and each other, resistance to harmful persuasion becomes a core requirement for reliable behavior. Yet we show that this requirement is far from met: a single targeted persuasive argument is enough to collapse model accuracy to near zero, even when the argument is factually false. We formalize this threat as adversarial persuasion and introduce an adversarial reinforcement learning framework that trains persuader agents to change a target model's answer in a single interaction. First, we show that optimizing persuasion strategies through trial and error exposes vulnerabilities that static prompting misses: RL-trained persuaders raise persuasion success from approximately 24% to over 93% against the training-time persuadee. Second, we find that these learned strategies transfer to unseen models, achieving 83% attack success on Qwen-14B, 79% on Llama-3.1-8B, and 25% on GPT-4o-mini. Third, we demonstrate that a curriculum that bootstraps on more persuadable open-weight models before targeting harder models further increases GPT-4o-mini attack success from 25% to 38%. Moreover, our results reveal that optimized persuaders increasingly rely on credibility-based tactics, including fabricated citations and false authoritative evidence. Together, these findings expose a critical weakness in current LLM agents: even when they initially reason correctly, they can be steered toward false conclusions by optimized natural language influence. This positions persuasion robustness as a necessary safety criterion for multi-agent and human-AI decision-making systems.
Authors:
Behzad Shayegh, Mohamed Osama Ahmed, Fred Tung, Leo FengAbstract: Large language models have demonstrated impressive retrieval-augmented capabilities. However, a crucial area remains underexplored: their ability to appropriately adapt responses to the certainty of the retrieved information. It is a limitation with real consequences in high-stakes domains like medicine and finance. We evaluate eight LLMs on their , measuring how well they adjust responses to match expressed context certainty. Our analysis reveals systematic limitations: LLMs struggle to recall prior knowledge after observing an uncertain context, misinterpret expressed certainties, and overtrust complex contexts. To address these, we propose an interaction strategy combining prior reminders, certainty recalibration, and context simplification. This approach reduces obedience errors by 25% on average, without modifying model weights, demonstrating the efficacy of interaction design in enhancing LLM reliability. Our contributions include a principled evaluation metric, empirical insights into LLMs' uncertainty handling, and a portable strategy to improve context-certainty obedience across diverse LLMs.
PaperID: 4947, Poster
Abstract: The growing context lengths of large language models (LLMs) have made key-value (KV) cache memory a critical bottleneck during inference. Vector quantization (VQ) with commutative codebooks provides an effective Rotary Position Embedding (RoPE)-aware solution by enabling reconstruction-free attention over quantized key indices. However, existing commutative VQ methods still parameterize scalar codebook coefficients across RoPE subspaces and codewords largely independently, leaving parameter-side redundancy underexploited. We propose RankVQ, a low-rank parameterized commutative vector quantization method for KV cache compression. Instead of factorizing the deployed codebooks, RankVQ factorizes the underlying scalar coefficient matrices used to construct key and value codebooks. This design preserves the pseudo-symmetric structure required for RoPE commutativity on the key side, while improving the quality and compactness of offline codebook construction. The low-rank factors are discarded after codebook construction, so RankVQ keeps the same online inference form as standard commutative VQ and introduces no additional decoding overhead. Experiments on LongBench, InfiniteBench, and GSM8K show that RankVQ achieves a stronger accuracy--compression trade-off than directly comparable KV quantization baselines, with particularly clear gains in the 1-bit regime.
PaperID: 4948, Poster
Authors: Yuntae Jeon, Younho Jeon, Sujin Jin, Sungho Jo, Seunghee Park
Abstract: Learning semantic representations for 3D Gaussian Splatting has recently emerged as a promising direction for 3D scene understanding. However, existing 2D-to-3D semantic lifting approaches rely primarily on 2D segmentation supervision and lack object-level priors, causing the learned representation to be driven by local mask evidence rather than holistic object structure. In this paper, we present VLSplat, a vision-language guided approach for object-centric semantic lifting in 3D Gaussian Splatting. Our method augments mask-based lifting with object-level semantic priors derived from vision-language models and refines Gaussian representations using an object-group scene graph. The refinement is performed through two language-guided modules: intra-group refinement suppresses semantic noise within object groups, while inter-group acquisition expands object support regions by incorporating structurally relevant candidates from neighboring groups. Extensive experiments on indoor scene datasets demonstrate that VLSplat improves semantic consistency and structural completeness. Our code and models are available at https://vlsplat.github.io/
PaperID: 4949, Poster
Authors:
Zixi Qin, Wanqing Li, Yuanhao Zhuo, Guoxin SuAbstract: Understanding relationships between variable sets is fundamental in machine learning, particularly in medical data analysis. Existing methods typically provide only coarse shared/private partitioning of variables, limiting their ability to reveal finer-grained predictive structures across views. We propose Mutual Predictability Decomposition (MPD), a framework that decomposes the variables in each view into three subsets: mutually predictable variables that can be inferred across views, auxiliary variables that are themselves not cross-view predictable but are necessary for predicting the mutually predictable subset, and unpredictable variables that do not contribute to cross-view predictability. MPD is formulated as a coupled bi-directional prediction problem with sparse variable assignment, and optimized through a differentiable relaxation implemented using gated neural networks. Experiments on synthetic and real-world medical datasets demonstrate that MPD recovers interpretable cross-view structures, improves cross-view prediction, and provides more informative decompositions than existing relationship analysis methods. Code will be available on GitHub upon acceptance.
PaperID: 4950, Poster
Abstract: Principled regression for stochastic processes is a long-standing challenge with deep connections to scientific inverse problems. We introduce Flow Annealing Posterior Sampling (FAPS), to our knowledge the first function-space posterior sampling framework that unifies stochastic-process regression and PDE inverse problems. Built on pretrained function-space flow-matching priors, FAPS enables likelihood-guided posterior inference from sparse and noisy observations, supports variable query discretizations, and avoids explicit prior-density evaluation. Its Langevin correction uses a low-rank covariance preconditioner to exploit dominant function-space correlations across discretizations. Across Gaussian and non-Gaussian stochastic-process regression benchmarks and diverse PDE inverse problems, FAPS produces coherent posterior samples with accurate uncertainty quantification, significantly outperforming existing functional regression baselines while remaining computationally efficient.
PaperID: 4951, Poster
Authors: Zequan Liu
Abstract: Neural avalanches and turbulence-like brain dynamics are usually measured after neural activity has been coarse-grained into continuous fields, leaving open which spike-level mechanisms can generate these macroscopic observables. We build a spatial leaky integrate-and-fire spiking neural network and explicitly transform spikes into activity-density and Hilbert phase fields. A delayed, distance-dependent, passive-dendrite regime reproducibly generates turbulence-like field structure, including high-amplitude phase-defect tracks, structure-function scaling, and power-law-favored continuous-field avalanche tails. Controls show that firing rate, vortex count, and avalanche tail fits are individually insufficient: destroying delayed propagation, dendritic filtering, or spike timing suppresses key field signatures, while random-space nulls can retain many mathematical phase defects but lose spatial scaling and synaptic-delay transfer. Biologically motivated gates further shape the regime: PV-like perisomatic inhibition clamps phase dynamics, SST-like dendritic inhibition and projection-specific inhibitory facilitation preserve long-lived tracks, and a rate-constrained projection-STP regime maintains similar dynamics at 12.25 Hz across 12 seeds. These results use SNNs as a mechanistic testbed for spike-derived brain turbulence observables and show why phase defects and avalanche-like statistics must be interpreted jointly with spatial scaling and spike-level propagation.
PaperID: 4952, Poster
Authors: Niki Triantafyllou, Maria Papathanasiou
Abstract: This paper presents EnerGNN, an energy-based framework for exact constrained combinatorial inference. The central idea is to learn an optimization-compatible energy over binary configurations, so that hard constraints and deployment-time restrictions can be enforced by a mixed-integer programming (MIP) layer at inference time. Training shapes the energy landscape using optimality-aware contrastive learning so that high-quality feasible solutions occupy low-energy regions. We evaluate EnerGNN in two settings: direct decision problems, where the learned energy is optimized over the full binary decision vector, and strategic decomposition problems, where it is optimized over complicating binary variables and an exact solver handles the downstream problem. Across Quadratic Knapsack, Combinatorial Auctions, and a personalized medicine supply chain problem, EnerGNN achieves primal gaps below 1% on most test instances, generalizes to larger problem instances without retraining, and supports zero-shot constraint injection by modifying only the inference MIP.
Authors:
Shijue Huang, Hangyu Guo, Guanting Dong, Chenxin Li, Junting Lu, Xinyu Geng, Zhaochen Su, Zhenyu Li, Shuang Chen, Hongru WANG, Yi R. (May) FungAbstract: Multimodal deep search requires an agent to solve open-world problems by chaining search, tool use, and visual reasoning over evolving textual and visual context. Two bottlenecks limit current systems. First, existing tool-use harnesses treat images returned by search, browsing, or transformation as transient outputs, so intermediate visual evidence cannot be re-consumed by later tools. Second, training data is usually built by fixed curation recipes that cannot track the target agent's evolving capability. To address these challenges, we first introduce a visual-native agent harness centered on an image bank reference protocol, which registers every tool-returned image as an addressable reference and makes intermediate visual evidence reusable by later tools. On top of this harness, On-policy Data Evolution (ODE) runs a closed-loop data generator that refines itself across rounds from rollouts of the policy being trained. This per-round refinement makes each round's data target what the current policy still needs to learn. The same framework supports both diverse supervised fine-tuning data and policy-aware reinforcement learning data curation, covering the full training lifecycle of the target agent. Across 8 multimodal deep search benchmarks, ODE improves the Qwen3-VL-8B agent from 24.9% to 39.0% on average, surpassing Gemini-2.5 Pro in standard agent-workflow setting (37.9%). At 30B, ODE raises the average score from 30.6% to 41.5%. Further analyses validate the effectiveness of image-bank reuse, especially on complex tasks requiring iterative visual refinement, while rollout-feedback evolution yields more grounded SFT traces and better policy-matched RL tasks than static synthesis.
Authors:
Yulun Wu, Sravan K Ankireddy, Samuel Sharpe, Nikita Seleznev, Dehao Yuan, Hyeji Kim, Nam NguyenAbstract: Efficiently aggregating spatial or temporal horizons to acquire compact representations has become a unifying principle in modern deep learning models, yet learning data-adaptive representations for long-horizon sequence data, especially continuous sequences like time series, remains an open challenge. While fixed-size patching has improved scalability and performance, discovering variable-sized, data-driven patches end-to-end often forces models to rely on soft discretization, specific backbones, or heuristic rules. In this work, we propose Reinforcement Patching (ReinPatch), a novel framework that jointly optimizes a sequence patching policy and its downstream sequence backbone model with reinforcement learning. By formulating patch boundary placement as a discrete decision process optimized via Group Relative Policy Gradient (GRPG), ReinPatch bypasses the need for continuous relaxations and performs dynamic patching policy optimization in a natural manner. Moreover, our method allows strict enforcement of a desired compression rate, freeing the downstream backbone to scale efficiently, and natively supports multi-level hierarchical modeling. We evaluate ReinPatch on time-series forecasting datasets, where it demonstrates compelling performance compared to state-of-the-art data-driven patching strategies. Furthermore, our detached design allows the patching module to be extracted as a standalone foundation patcher, providing the community with visual and empirical insights into the segmentation behaviors preferred by a purely performance-driven neural patching strategy.
PaperID: 4955, Poster
Abstract: On-policy reinforcement learning has become a central paradigm for improving the reasoning abilities of large language models. However, its effectiveness is often limited by reward sparsity: when a model fails to discover correct trajectories for difficult problems, the optimization process receives little useful signal and may stagnate. Existing approaches mitigate this issue by incorporating off-policy demonstrations, expert traces, or model-generated solutions, but they typically require the auxiliary data to match the format of the reinforcement-learning task, often relying on rejection sampling from stronger models to obtain suitable training trajectories. We introduce Rationale-Guided Policy Optimization (RGPO), a framework that adaptively leverages ground-truth rationale information according to the model’s current capability while preserving its freedom to explore. Rather than treating reference solutions as fixed imitation targets, RGPO uses them as temporary scaffolds: rationales help the model generate improved responses, after which only higher-reward, model-generated solutions are transferred back to the original unguided setting. This design allows training to exploit available ground-truth information without requiring off-policy data to follow the same format as the RL task. Across both language-only and vision-language reasoning settings, RGPO consistently improves performance over RLVR baselines, and ablation studies show that adaptive rationale guidance is a key contributor to these gains. These results suggest that RGPO offers a practical and general approach for reducing reward sparsity, stabilizing reinforcement learning, and improving reasoning performance in both text-only and multimodal models.
Authors: Talha Rüzgar Akkuş, Şuayp T Kocabay, Kamer A Yuksel, Hassan Sawaf
Abstract: The Forward-Forward (FF) algorithm trains networks layer-by-layer using a local "goodness function," yet the principles governing what makes an effective goodness function remain largely unexplored. We systematically explore the goodness-function design space and identify a unifying principle: the goodness function must be sensitive to the shape of neural activity, not its total energy. This principle is motivated by the observation that deep network activations follow heavy-tailed distributions and that discriminative information is often concentrated in peak activities. We propose two complementary families: selective functions (top-k, entmax-weighted energy) that measure only peak activity, and shape-sensitive functions (excess kurtosis / "burstiness" and higher-order moments) that reward heavy-tailed distributions via scale-invariant statistics. Combined with separate label-feature forwarding (FFCL), controlled experiments across 13 goodness functions, 5 activations, 6 datasets, and three continuous sweeps each tracing a characteristic inverted-U yield 89.0% on Fashion-MNIST and 98.2±0.1% on MNIST (4x2000)—a +32.6pp gain over SoS—with consistent improvements across all benchmarks (+72pp USPS, +52pp SVHN). The scale-invariant nature of burstiness makes it particularly robust to magnitude shifts across layers and datasets. Code is available at
PaperID: 4957, Poster
Abstract: Federated intrusion detection must adapt to rare attacks from only a few labeled flows per client, while client distributions are highly non-IID and raw traffic cannot be centralized. Existing few-shot intrusion detectors typically construct class priors in centralized settings, and standard federated training does not provide an explicit mechanism for reusing cross-client class statistics during local low-shot adaptation. We propose a semantic-statistical prior bank for federated low-shot intrusion detection. Clients first participate in federated representation learning and submit protected class-wise moment statistics through secure aggregation, from which the server constructs a global prior bank without accessing raw traffic. During local episodic adaptation, label-phrase embeddings retrieve relevant priors, which are combined with scarce support statistics through a maximum a posteriori (MAP) shrinkage estimator in the feature space. The adapted class distributions are then used to train lightweight episode-specific classifiers on each client. Experiments on CICIDS2017 and CICDDoS2019 under 5-way 1-shot and 5-shot non-IID settings show consistent improvements over FedAvg and backbone-matched baselines, with the largest gains appearing in 1-shot and highly skewed client distributions. With ResNet1D, the proposed adaptation improves 1-shot F1 from 92.97 to 95.55 on CICIDS2017 and from 82.13 to 88.25 on CICDDoS2019.
Authors:
Pablo Martínez Crespo, Stefano Ribes, Martin Rahm, Richard Johannes Maximilian Beckmann, Robert Jordan, Marisa Gliege, Santiago Miret, Vijay K Narasimhan, Rocío MercadoAbstract: Atomic properties such as partial charges or multipoles encode chemically meaningful information that can inform downstream molecular property prediction, but their evaluation as machine learning targets has been complicated by the absence of a principled out-of-distribution evaluation protocol at the atomic level. In this work, we propose a held-out evaluation protocol that clusters atomic environments by SOAP descriptors and computes metrics accounting only for cluster labels unseen during training. Following this procedure, we use 5×5 cross-validation and Tukey's HSD to run a statistically rigorous comparison of E(3)-equivariant against non-equivariant, rotationally augmented models for predicting electron populations and multipoles of H, C, N, and O atoms. Building on our results, we introduce the Quantum Topological Neural Network (QT-Net), a rotationally augmented, non-equivariant graph neural network. We show that QT-Net can be used to infer properties of atoms in molecules from QM9 outside our training set, and that these inferred properties can yield improvement when used as input features for downstream molecular property prediction. To further validate the framework, molecular dipole moments computed from QT-Net's per-atom outputs recover the ground-truth values reported in QM9. We release all code and data, including a JAX implementation of QT-Net, to support the broader use of learned QTA properties as inductive biases for atomic-scale molecular machine learning.
Abstract: Stochastic momentum methods such as heavy ball (HB), Nesterov momentum, and variants of Accelerated SGD (ASGD) [Kidambi et al., 2018] are widely used in modern training, but their stochastic benefits depend on two distinct quantities: serial runtime, the number of iterations needed to reach a target accuracy, and compute efficiency (CE), the inverse total gradient-query or FLOP cost. Larger batches reduce serial runtime without hurting CE only when the contraction gap grows linearly with batch size. We study stochastic HB and ASGD for consistent linear regression with Gaussian covariates and prove finite-dimensional, discrete-time lower bounds on their batch-size tradeoffs. Our first result shows that HB does not improve the CE frontier over SGD for arbitrary spectra; rather, it preserves SGD-level CE over a larger batch-size window, allowing larger batches to reduce serial runtime until HB reaches its deterministic accelerated scale. This window can be a factor \sqrt\kappa larger than the SGD critical batch size. For ASGD, the picture is more spectrum-dependent: for rapidly decaying power-law spectra, ASGD improves small-batch CE over HB/SGD, but as batch size grows it trades this CE advantage for improved serial runtime. Synthetic linear-regression experiments verify these qualitative regimes, including near-overlap of ASGD and HB for slowly decaying spectra and the predicted CE--serial tradeoff for rapidly decaying spectra.
PaperID: 4960, Poster
Abstract: We apply offline reinforcement learning to large-scale warehouse consolidation, an application that requires ranking tens of thousands of candidates at each decision epoch with an architecture that scales gracefully with variable-cardinality action spaces. Since the available reward signal is only an approximation of the true target-domain objective, the policy must also transfer robustly. We propose a two-stage framework, Asymmetric Implicit Q-Learning (AIQL) and its adaptive cross-domain extension A^2IQL. First, AIQL learns a scoring policy from fixed offline consolidation logs by decoupling a per-candidate scoring actor from an aggregate-statistic critic, keeping value estimation tractable and within the support of logged data. Second, A^2IQL adapts the source-domain-trained policy to a new target objective using scarce target-domain data: a Bayesian reward predictor supplies an uncertainty-penalized conservative bonus that modifies only the advantage-weighted actor extraction, leaving the source critic unchanged. On a single-domain benchmark, AIQL improves throughput by 20% over a heuristic baseline using offline data alone. When adapted to a new target objective via A^2IQL, the method outperforms source-only training, data pooling, and off-dynamics baselines by up to 23% under a 30:1 source-to-target data ratio while achieving the lowest variance across all methods.
Abstract: Hyperparameter prediction is a critical practical bottleneck for model-based image denoisers, ranging from classical TV/TGV variational solvers to modern diffusion-based models such as DiffPIR. While existing learned predictors can achieve near-oracle performance, this approach scales poorly: each new configuration conventionally requires its own oracle-labeled training set, and each label requires an exhaustive denoiser sweep evaluated against clean ground truth. We therefore ask whether oracle supervision collected on source configurations can transfer to target configurations with few or no target oracle labels. We propose HyperDn, a single configuration-conditioned predictor that pools oracle supervision across source configurations and predicts heterogeneous hyperparameters for new denoiser--noise configurations. In a cross-paradigm experiment, HyperDn transfers from relatively cheap TV/TGV variational sources to more expensive diffusion-based DiffPIR. With only 2 target oracle labels, it reaches 30.23\,dB, within 0.90\,dB of oracle, using 1/32 the target labels that a per-configuration predictor trained from scratch needs to match. Without any target oracle labels, HyperDn also reaches near-oracle PSNR on two unseen mixtures of seen noise types and on transfer from relatively cheap 96× 96 source images to 512× 768 targets. Together, these results show that expensive oracle supervision for hyperparameter prediction can be transferred from source to new target configurations, reducing the need to rebuild oracle labels for each new denoising configuration.
Authors: Xiang Gao, Yunpeng Jia
Abstract: Visual anagram is an intriguing form of art creation wherein a single image presents different conceptual interpretations under transformations such as flipping or rotation. Recent work has achieved visual anagram synthesis by leveraging pretrained text-to-image (T2I) diffusion models, yet still suffers from several key limitations including computational inefficiency, suboptimal aesthetic quality, and weak semantic fidelity and expressiveness. This work focuses on generating visual anagrams with substantially improved visual quality at minimal computational cost, thereby advancing intelligent creation of illusionary digital art. To increase image resolution while reducing time overhead, we adapt the cutting-edge parallel denoising algorithm from pixel-based T2I model to the adversarially distilled latent-based one, and accordingly propose a structure-semantic co-optimization (S2CO) framework to counteract the consequent visual degradation. As the core of our approach, S2CO framework comprises three key innovations: (\romannumeral1) null-text structure alignment optimization; (\romannumeral2) semantic enhancement optimization; (\romannumeral3) attention-guided noise fusion. Building upon these components, our method dubbed S2CO-Anagram is able to generate higher-resolution anagram images with noticeably superior visual harmony and semantic faithfulness than related SOTA approaches, all while achieving substantially faster inference speed. Code will be publicly available.
Abstract: Modeling event sequences requires understanding when future events will occur and how event types causally influence one another over time. Existing temporal point process (TPP) models typically assume static or first-order dependence structures, which limits their ability to capture the dynamic and multi-order causal mechanisms commonly observed in real-world systems. We propose MOCHA, a Multi-Order Causal Hierarchical Architecture for multivariate TPP that jointly models time-varying causal structure and multi-hop influence propagation in continuous time. MOCHA learns latent dynamic weighted directed acyclic graphs (DAG) over event types, where acyclicity and sparsity constraints promote structurally valid causal graphs. Based on the learned graph, the model decomposes event dynamics into multi-order causal paths, and incorporate both direct influence and indirect propagation into the intensity function. The entire framework is end-to-end differentiable and optimized jointly for event prediction and causal structure discovery. Experiments on seven real-world datasets from different domains show that MOCHA consistently achieves superior negative log-likelihood, while also recovering dynamic causal patterns that align with domain knowledge. These results demonstrate that MOCHA provides an effective framework for dynamic causal structure learning in TPP.
PaperID: 4964, Poster
Abstract: Subject-driven generation (SDG) evaluation using multi-modal LLMs (MLLMs) provides better interpretability and stronger alignment with human judgment than traditional embedding-based methods. However, MLLM-as-a-Judge scores are limited in differentiating between error and error-free SDG images. In this paper, we propose CompJudge, which overcomes these limitations by introducing comparative capabilities to MLLM-as-a-Judge and recalibrating scores based on comparison results. To rigorously evaluate the correctness of SDG evaluators, we introduce CompILIAS, a diagnostic benchmark of triplets (reference, identity-preserving image, identity-degraded image) with verified binary ground-truth labels indicating whether the SDG image preserves the reference identity. We demonstrate improvements of SDG evaluation over existing approaches: 8.9pp on CompILIAS and 10.62--15.98pp on the DreamBench++ dataset. These improvements consistently generalize across diverse MLLM backbones and across subject categories.
Abstract: Text-to-Image diffusion models often propagate harmful bias inherited from the training data. Existing bias mitigation techniques either intervene only at the text encoder or provide inference-time guidance, often leading to generations that collapse into semantically incoherent outputs. To address these limitations, we introduce CO-ALIGN (Concept Ontology Alignment), a novel bias mitigation approach based on concept-graph alignment which operates on the model's internal concept ontology. By aligning concepts within the text encoder and denoiser, CO-ALIGN achieves significant bias reduction while preserving generative integrity. We demonstrate the effectiveness of concept graph alignment in three paradigms- text-encoders, denoisers, and joint text-denoiser ontology alignment. CO-ALIGN outperforms outperforms the state of the art, improving fairness by 30%, \Delta FID=11.4 in image quality, 2.8% in image fidelity, all while reducing semantically incoherent outputs by 88%. Beyond bias mitigation, we show that CO-ALIGN benefits other downstream tasks as well. In particular, our experiments demonstrate that better-aligned internal ontologies enhance concept unlearning robustness across multiple unlearning techniques.
PaperID: 4966, Poster
Authors:
Zhaowei Liu, Kaixiang Wang, Sheng Liu, zengyang zhang, HaitaoYang, Rufei Gao, Zongxing Zhao, Yao Shan, Yongchao Song, Yanwei Yu, Anzuo JiangAbstract: Time series forecasting is important in applications such as energy management, climate modeling, traffic control, and industrial monitoring. Recent variate-tokenization Transformers model cross-variate dependencies effectively, but Softmax attention still suffers from uneven cross-variate routing and fixed attention sharpness under varying non-stationarity. To address these limitations, OTformer is introduced as a non-stationarity-aware Optimal Transport framework for time series forecasting. A lightweight scorer estimates sample-wise non-stationarity from raw pre-normalization statistics and uses this signal to adapt the Sinkhorn entropy regularization coefficient ε, yielding smoother transport plans for more non-stationary inputs and sharper routing for more stationary ones. A shared phase embedding further injects cycle-aware context into variate tokens. Experiments on eight standard benchmarks and 32 forecasting settings show that OTformer consistently improves upon the iTransformer backbone, with average MSE reductions of 4.4% on ECL and 12.5% on Solar across four forecasting horizons.
PaperID: 4967, Poster
Abstract: Visual text editing aims to precisely modify text in images and videos while preserving stylistic consistency and visual realism. Despite significant advances in the image domain, video text editing remains largely unexplored: it is a localized task demanding stroke-level precision within small text regions, which compounds the challenges of cross-frame accuracy, temporal coherence, and stylistic fidelity. We introduce SteerVTE, a unified framework that \underlinesteers a frozen video diffusion model to perform precise \underlineVideo \underlineText \underlineEditing through style and glyph control. Built on a frozen diffusion transformer, SteerVTE attaches a lightweight text context adapter with two complementary modules: a style encoder capturing the original text's visual attributes, and dual-granularity glyph encoders encoding the target text at both the line and character levels. To overcome the inherently weak text rendering priors of video foundation models, we further propose a glyph-aware spatial-focal loss and a three-stage progressive training curriculum that scales from image to video data. To support large-scale training, we also develop an automatic synthesis pipeline and construct SteerVTE-1M, a dataset of one million triplets spanning diverse scenes, fonts, and stylistic effects. Extensive experiments demonstrate that SteerVTE substantially outperforms existing video editing baselines across text accuracy, style consistency, and temporal coherence. Our anonymous project website is available at \urlhttps://steervte-paper.github.io.
PaperID: 4968, Poster
Abstract: Graph Neural Networks (GNNs) have demonstrated remarkable effectiveness on graph-based tasks. However, their predictive confidence is often miscalibrated, typically exhibiting under-confidence, which harms the reliability of their decisions. Existing calibration methods for GNNs normally introduce additional calibration components, which fail to capture the intrinsic relationship between the model and the prediction confidence, resulting in limited theoretical guarantees and increased computational overhead. To address this issue, we propose a simple yet efficient graph calibration method. We establish a unified theoretical framework revealing that model confidence is jointly governed by class-centroid-level and node-level calibration at the final layer. Based on this insight, we theoretically show that reducing the weight decay of the final-layer parameters alleviates GNN under-confidence by acting on the class-centroid level, while node-level calibration acts as a finer-grained complement to class-centroid level calibration, which encourages each test node to be closer to its predicted class centroid at the final-layer representations. Extensive experiments validate the superiority of our method.
PaperID: 4969, Poster
Authors:
Shangpin Peng, Gengluo Li, Xingyu Wan, Chengquan Zhang, Hao Feng, Binghong Wu, Huawen Shen, Weinong Wang, Ziyi Cai, Zhuotao Tian, Han Hu, Can Ma, Yu ZhouAbstract: Charts convey dense quantitative and relational information, yet general chart parsing remains challenging due to fragmented evaluation and limited structural reasoning. Existing benchmarks focus on narrow sub-tasks with inconsistent output formats (e.g., Markdown, CSV, JSON, SVG, code), hindering fair comparison and often overlooking real-world scenarios such as printed or hand-drawn photos. Meanwhile, end-to-end multimodal models tend to learn shallow pixel-to-text mappings without explicit structure understanding, making them fragile under visual variations. To address these issues, we present a unified framework that connects evaluation and modeling. On the evaluation side, we introduce , a comprehensive bilingual benchmark covering eight chart families, including both numeric charts (bar, line, pie, radar, box plot, combination) and diagrammatic structures (flowchart, mind map). The dataset is constructed via a human-agent collaborative annotation pipeline, where model-assisted drafts are refined through multi-stage human verification to ensure structural consistency and annotation reliability. Furthermore, we design a format-agnostic evaluation protocol that maps different outputs into two canonical semantic spaces: a normalized triple view and a directed graph view, and evaluates them with structure-aware metrics. On the modeling side, we propose , a lightweight semantic scaffold that introduces structural inductive bias during supervised fine-tuning. By guiding the model to infer structure before generating outputs, CAP leads to more consistent and robust parsing behavior. Extensive experiments show that models trained with CAP achieve reliable performance on both rendered charts and real-world images. Code, benchmark, and models will be made publicly available.
PaperID: 4970, Poster
Abstract: The difficulty of factual error correction (FEC) is not only how to repair false claims, but how to construct errors that meaningfully test repair. Existing FEC data construction methods often follow a static corruption view: they mask spans in false claims or inject errors into supported claims, but do not explicitly target the boundary of what current evidence-based correctors can repair. We introduce the notion of a \emphrepair frontier: factual errors that are evidence-grounded, close enough to the source claim to admit a clear correction, and resistant to current correctors. We propose \textscSAFE-FEC, a semantically constrained adversarial frontier-evolution framework for constructing such examples. Starting from evidence-supported claims, SAFE-FEC generates mutations from multiple factual-error perspectives, refines them through cross-perspective critique, selects or composes stronger candidates, filters semantic drift, and accepts only candidates that survive repeated corrector-in-the-loop repair attempts. This turns FEC data construction from one-shot error injection into correction-aware frontier search. Experiments on FECData, HoVer, and FEVEROUS show that SAFE-FEC consistently lowers correction performance for both LLM-based and FEC-specific correctors, especially in multi-hop and structured-evidence settings. Quality evaluation and ablations further suggest that these failures arise from plausible, semantically anchored, evidence-grounded errors rather than invalid hard negatives.
PaperID: 4971, Poster
Abstract: Real-world autonomous driving data inherently exhibits out-of-distribution (OOD) dynamics, as diverse interactions and unpredictable human behaviors cause vehicle trajectory distributions to vary significantly across different scenarios, even for similar road structures. Such OOD dynamics typically induce uncontrollable extrapolation errors in neural networks. We observe that in offline reinforcement learning (RL), these errors manifest as heavy-tailed estimates of action values with significant biases, which existing robust methods fail to address, hindering effective generalization. To address this problem, we propose a robusT offline RL algorithm under heavy-tailed action-valUe eSTimates (TRUST), which models perturbed states associated with heavy-tailed action-value estimates to effectively mitigate such estimation biases. Specifically, TRUST uses a first-order gradient-regularized target to model perturbed states that approximate data under worst-case dynamics. It then introduces a statistical measure to identify perturbations that induce heavy-tailed action-value estimates. By re-weighting the training objective to emphasize correcting these biased estimates, TRUST can reduce the heavy-tailed behavior in action values for robustness against OOD driving data. We further derive an error bound that justifies the gradient-regularized target as an approximation to the robust Bellman target. Extensive experiments on three large-scale real-world datasets show that TRUST achieves strong overall performance across all benchmarks, with a general performance improvement of 41.21%.
PaperID: 4972, Poster
Abstract: Vision-language models (VLMs) learn strong statistical regularities during training, which can make them fail to perceive visual evidence when input images violate those regularities. Such failures are often treated as missing visual capability, but they may instead arise because the model has access to the relevant visual evidence and yet selects an answer dominated by learned training priors. In this work, we use controlled visual in-context learning to show that VLMs can overcome learned priors for visual tasks and use relevant visual evidence and capabilities. Specifically, we show that across five vision-centric tasks, demonstrations yield large gains on prior-conflicting counterfactual examples (+27.2% for Qwen3-VL 235B and +18.1% for Qwen3-VL 32B) while leaving real accuracy nearly unchanged. These findings also generalize across five VLM families with performance increases from +9.7% to +30.3%. To explain this effect, we analyze visual representations and find that demonstrations redirect attention toward grounding cues while reducing reliance on cues that support the learned prior. Finally, we find that in-context learning is limited when the underlying visual capability is weak or absent. Together, these results show that current VLMs possess visual grounding capabilities that standard evaluations may overlook, underscoring the need for controlled evaluation design in diagnosing and improving model capability.
Abstract: Activation steering offers a lightweight way to control large language models without retraining, but its effectiveness varies sharply across concepts. Prior work often interprets this variability as evidence that many concepts are not well captured by a single steering direction. We argue instead that much of this variability reflects search difficulty: a useful rank-1 intervention often exists, but finding it can be expensive. We formalize rank-1 steering as a budget-constrained optimization problem over intervention layer and coefficient. Across the concepts and model families, prompt-boundary directional alignment predicts where effective interventions are likely to occur, enabling geometry-guided search that reaches high utility with substantially fewer evaluations, reducing the trials needed to recover 95% of best-found utility by 39.8% on average across three model families. To explain why some concepts remain expensive even under better search, we introduce concept granularity, a measure of directional heterogeneity across contrastive contexts. Granularity distinguishes concepts whose difference vectors share a stable global direction from those where prompts agree locally within each input but the utility-maximizing direction rotates systematically across inputs. Higher granularity is associated with both slower convergence and lower best-found steering performance (Pearson r = 0.44 with trials-to-95%, p < 0.001, and r=-0.46 with best-found utility, p < 0.001). These observations suggest a practical workflow rather than a single universal vector-construction rule. We therefore present GRACE, a Granularity- and Representation-Aware Concept Engineering framework that uses activation geometry to diagnose the dominant source of steering difficulty, choose the appropriate remedy, and allocate optimization effort more efficiently. Our results shift the frame of activation steering from "when does rank-1 fail?" to "when is rank-1 cheap and stable?", and turn activation geometry from a descriptive tool into an actionable prior for LLM control.
Authors: Ling Liang, Lei Yang
Abstract: Gromov--Wasserstein optimal transport (GWOT) aligns metric measure spaces by matching their within-domain relational structures, but large-scale GWOT remains challenging because its objective is nonconvex and projection onto the transport polytope is often solved only approximately in practice. This leads to a gap between practical projected-gradient implementations and convergence theory, which typically assumes exact projections. For squared-loss GWOT, we propose an inexact projected-gradient framework with a verifiable feasibility-residual-based inexact condition for the projection subproblem. This condition is directly computable and avoids unknown quantities such as the exact projection point. Under this implementable condition, we prove subsequential convergence to stationary points and, with a mild tolerance-decay condition, convergence of the whole sequence. The resulting method retains the simplicity and sparsity of projected-gradient schemes while providing rigorous convergence guarantees, turning projected-gradient methods into a principled and scalable approach for GWOT with provable reliability.
Abstract: Reconstructing nonlinear dynamical systems (DS) from data (DSR) is a fundamental challenge in science and engineering, but it inherently relies on sequential models. Recent breakthroughs for sequential models have produced algorithms that parallelize computation along sequence length T, achieving logarithmic time complexity, \mathcalO(\log T). Since sequence lengths have been practically limited due to the linear runtime complexity \mathcalO(T) of classical backpropagation through time, this opens new avenues for DSR. This paper studies two prominent classes of parallel-in-time algorithms for this task, both of which leverage parallel associative scans as their core computational primitive. The first class comprises models with linear yet non-autonomous dynamics and a nonlinear readout, such as modern State Space Models (SSMs), while the second consists of general nonlinear models which can be parallelized using the DEER framework. We find that the linear training-time recurrence of the first class of models imposes limitations that often hinder learning of accurate nonlinear dynamics. To address this, we augment DEER with Generalized Teacher Forcing (GTF), a novel variant within the more general nonlinear framework that ensures stable and effective learning of nonlinear dynamics across arbitrary sequence lengths. Using GTF-DEER, we investigate the benefits of training on extremely long sequences (T>10^4) for DSR. Our results show that access to such long trajectories significantly improves DSR if the data features long time scales. This work establishes GTF-DEER as a robust tool for data-driven discovery and underscores the largely untapped potential of long-sequence learning in modeling complex DS.
PaperID: 4976, Poster
Abstract: AI systems increasingly assist decision making by producing cheap factor-level assessments of complex inputs, but these assessments are often biased and incomplete. We study selective human oversight: given AI signals, which factors should be escalated to a costly human evaluator? We formulate this human--AI collaboration problem as a factor-level information-acquisition problem. Under squared loss, solving for optimal dispatching reduces to maximizing a contextual reward; under a linear model, this reward admits a closed-form decomposition into two interpretable terms: predictive relevance and residual uncertainty in the human evaluation after conditioning on the AI signal. We instantiate the framework on peer review, decomposing papers into ten aspects and evaluating 3,408 ICLR submissions across three LLMs and multiple regression heads. Outputs based on the optimal dispatching rules match full-human-review performance using only 2--3 queried aspects out of 10. The framework remains effective when the human signal comes from a single noisy reviewer and across review years, suggesting robustness to reviewer inconsistency and and temporal distribution shift. These results position principled dispatching as a practical foundation for scalable human--AI decision systems.
PaperID: 4977, Poster
Abstract: Although existing deep multi-view clustering methods often achieve high performance, they predominantly focus on improving the fused representation while overlooking the enhancement of individual views, often resulting in stagnant or inconsistently improved view-specific representations that limit the final fusion quality. To address this, we propose Explicit single-View Enhancement (EVE), a novel deep multi-view clustering framework designed to explicitly optimize the clustering capability of each individual view. Specifically, EVE employs an attention-based module to generate a global consensus as a reliable anchor, which subsequently guides the refinement of individual views through two key components: feature-level selective alignment and structural-level neighborhood propagation. Crucially, the enhanced single views will further boost the fused representation, and finally leading to a mutually reinforcing closed-loop feedback process. Our method is also naturally applicable to incomplete deep multi-view clustering scenarios. Extensive experiments on multiple benchmark datasets demonstrate that EVE consistently outperforms state-of-the-art competitors, achieving substantial average improvements of
PaperID: 4978, Poster
Abstract: Catastrophic forgetting is a central obstacle in single-stage fine-tuning of large language models (LLMs) on a new task, especially when the original pretraining data are unavailable. Synthetic data offers a practical mitigation path, but candidate pools are often noisy and redundant, making effective selection essential. Existing methods often rely on indirect proxy signals whose alignment with actual training benefit can be limited, while utility--diversity selection over large candidate pools remains computationally demanding. To address these challenges, we propose NAvigator-guided Data Selection (NADS), a fine-tuning framework for LLMs. NADS first follows the target fine-tuning recipe without preservation regularization to obtain a navigator model, whose deviation from the pretrained model exposes recipe-induced capability drift. It then uses the predictive divergence between the navigator and the pretrained model to score forgetting-aware utility for each synthetic candidate, and constructs a constraint set through an efficient utility--diversity selection strategy. Finally, NADS fine-tunes on the new-task data while distilling pretrained behavior on the selected constraints to preserve general capabilities. Experiments show that NADS delivers a stronger balance between new-task performance and general capability preservation, while reducing utility--diversity selection overhead.
Abstract: Synthetic data has become a promising way to scale model training beyond limited human-generated data but it may also induce strong model collapse (Dohmatob et al., 2024), where any fixed fraction of synthetic data prevents model performance from improving under data scaling, leaving a non-vanishing excess risk floor. In this paper, we studied how synthetic data affects the generalization of one-pass SGD in high-dimensional linear regression with model shift. We show that mixed training induces strong model collapse while two-stage training avoids this by using synthetic data only in the first stage, followed by real-data training in the second stage, showing that strong model collapse is not inevitable through a simple data curriculum. We further establish scaling-law upper bounds for both protocols under a random sketch model, showing that larger models amplify synthetic-induced degradation in mixed training and giving an explicit characterization of how high-quality synthetic training may reduce bias in two-stage training. Overall, our results highlight that synthetic data is neither inherently harmful nor beneficial; its effect depends critically on both its quality and the training protocol used to incorporate it.
Abstract: Reinforcement Learning (RL) has emerged as a pivotal post-training paradigm for Large Language Models (LLMs), yet it frequently suffers from unpredictable training collapses. Recent findings attribute these failures to a hidden train-inference discrepancy (or mismatch), such a discrepancy will increase as training goes on, stemming from the disparate underlying engines and precisions required to balance generation throughput and training fidelity. Such Existing mitigations either sacrifice numerical stability or rely on heuristic masking that fails to explicitly optimize the ultimately deployed policy. In this paper, we discover that training policies inherently possess the capability to heal this common and harmful discrepancy. To operationalize this, we transition the standard RL objective into a Discrepancy-Constrained Markov Decision Process (\textttDCMDP). In order to practice this new paradigm at the algorithmic level, we first introduce a robust, trajectory-level geometric penalty that provides black-box feedback of inter-policy deviations, enabling autonomous self-correction under any specific trigger, e.g., infrastructure level or model architecture level. Furthermore, capitalizing on our empirical discovery of a discrepancy tolerance region in various models, we employ an adaptive performance-discrepancy balancing mechanism that penalizes the policy only when deviations exceed a safe boundary, achieving stable dual-objective optimization. Our approach not only eradicates mismatch-induced collapses and shatters performance bottlenecks, but crucially unlocks a heterogeneous training paradigm—leveraging high-fidelity training environments to natively optimize LLMs for low-cost, resource-constrained deployments.
Abstract: Evaluation leaderboards such as LMArena play a central role in benchmarking large language models by aggregating pairwise human preferences into model rankings, but the robustness of these rankings is still poorly understood. We present a unified perturbation framework for analyzing Bradley–Terry leaderboards under structured data modifications using influence-based approximations. Our framework studies three match-level perturbations—dropping matches, adding matches, and flipping match outcomes—along with player removal. We evaluate their effects on top-k membership, global ranking consistency measured by Kendall’s tau, and confidence-interval uncertainty. Across Chatbot Arena and six additional pairwise-comparison datasets, we show that modern leaderboards are non-robust across all three objectives: targeted perturbations affecting less than 1% of the data can change the top-ranked model, degrade global ranking consistency, and alter confidence intervals. We summarize these effects using normalized dataset-level robustness scores that compare fragility across leaderboard designs. We further show that influence scores enable efficient targeted manipulation, promoting or demoting specific models with fewer actions than prior manipulation baselines, while also identifying additional matchups that reduce uncertainty for target models. Finally, player-removal analysis shows that removing influential models can induce broad reordering, highlighting model deprecation as a source of the leaderboard illusion. These findings reveal fundamental limitations of current leaderboard designs and motivate more robust evaluation protocols.
Abstract: While Large Language Models (LLMs) are commonly fine-tuned to handle domain-specific tasks before being applied to vertical applications, adapting them to complex scenarios with diverse specialized knowledge remains challenging. Meanwhile, Mixture-of-Experts (MoE) architecture has risen as a crucial paradigm for training LLMs, and some recent works have also incorporated MoE into Parameter-Efficient Fine-Tuning (PEFT) to propose the Mixture of Low-rank Experts (MoE-LoRA), to enhance the power of low-rank adapters for learning complicated knowledge. However, conventional gating mechanisms in MoE typically apply only a scalar reweighing to selected experts, thereby limiting their underlying capacity of representation and generalization. Motivated and enabled by the low-rank structures in MoE-LoRA, we propose RotMoLE, a specialized MoE framework for low-rank experts featuring an additional rotation gate. Beyond simple scaling, RotMoLE implements a rotation mechanism for each selected expert, enabling superior expert exploitation and specialization for learning diverse data, especially when expert candidates are limited. Empirical results on complex multi-task and multilingual training scenarios validate the effectiveness of our methodology.
PaperID: 4983, Poster
Authors: Jeonghwan Cheon, Marin Vogelsang, Lukas Vogelsang, Pawan Sinha
Abstract: While vision-language models now achieve impressive image–text alignment, they often rely on surface-level associations rather than grounded conceptual representations that generalize across instances, support abstraction, and align with human semantic structure. Inspired by human visual development, in which infants begin life with limited acuity and chromatic sensitivity that gradually mature, we ask whether developmentally structured perceptual input can serve as an inductive bias for grounded concept acquisition. We trained vision-language contrastive learning models under two regimens: a standard regimen using clear images throughout, and a biomimetic regimen in which initially blurred and grayscale inputs gradually transition to clear images. Although both models achieve comparable caption-level alignment, the biomimetic model exhibits substantially stronger alignment at basic-level and superordinate category levels. Moreover, sparse autoencoder analyses reveal more compatible latent codes across modalities, and learning trajectories exhibit a clear coarse-to-fine progression, with broad distinctions acquired earlier in training and finer distinctions emerging later. Biomimetic representations also better match taxonomic structure and human behavioral similarity judgments. Taken together, these findings suggest that early perceptual limitations are not merely developmental obstacles but serve as adaptive inductive biases that scaffold structured, hierarchical, and human-aligned multimodal concept learning.
Abstract: As learning systems increasingly shape everyday decisions, Algorithmic Collective Action (ACA), i.e., users coordinating changes to shared data to steer model behavior, offers a complement to regulator-side policy and corporate model design. Real-world collective actions have traditionally been decentralized and fragmented into multiple collectives, despite sharing overarching objectives, with each collective differing in size, strategy, and actionable goals. However, most of the ACA literature focuses on single collective settings. To address this, we propose the first comprehensive statistical framework for ACA with multiple collectives acting on the same system. In particular, we focus on collective action in classification, studying how multiple collectives can influence a classifier's behavior. We provide quantitative statistical bounds on the success of the collectives, considering the role and the interplay of the collectives' sizes and the alignment of their goals. We make such bounds computable by each collective with only partial knowledge of other collectives' sizes and strategies. Finally, we numerically illustrate our framework on simulations inspired by interventions for climate adaptation in smart cities, demonstrating the usefulness of our bounds.
Authors:
Shashank Subramanian, Alexander Kiefer, Arnur Nigmetov, Amir Gholami, Dmitriy Morozov, Michael MahoneyAbstract: Neural scaling laws, which in some domains can predict the performance of large neural networks as a function of model, data, and compute scale, are the cornerstone of building foundation models in Natural Language Processing and Computer Vision. We study neural scaling in Scientific Machine Learning, focusing on models for weather forecasting. To analyze scaling behavior in as simple a setting as possible, we adopt a minimal, scalable, general-purpose Swin Transformer architecture, and we use continual training with constant learning rates and periodic cooldowns as an efficient training strategy. We show that models trained in this minimalist way follow predictable scaling trends and outperform standard cosine learning rate schedules. Cooldown phases can be re-purposed to improve downstream performance, e.g., enabling accurate multi-step rollouts over longer forecast horizons or sharper predictions through spectral loss adjustments. We also systematically explore a wide range of model and dataset sizes under various compute budgets to construct IsoFLOP curves, and we identify compute-optimal training regimes. Extrapolating these trends to larger scales highlights potential performance limits, demonstrating that neural scaling can serve as an important diagnostic for efficient resource allocation. We open-source our code for reproducibility.
PaperID: 4986, Poster
Authors: Jiechao Gao, Chang Liu, Yuandong Pan, Ying Liu, Michael Lepech
Abstract: While Large Language Models (LLMs) have significantly advanced radiology report generation (RRG) with strong generative priors, standard autoregressive decoders still suffer from sequential error accumulation when they must infer all clinical structures directly from images. Diffusion Language Models (DLMs) have recently emerged as a non-autoregressive paradigm with iterative masking and denoising mechanisms. Instead of replacing the mature language decoder with a DLM, we reveal that the key value of DLMs for RRG lies in controllable clinical planning, where the model predicts globally consistent anchor points before free-text realization. To harness this complementary strength, we propose ANCHOR-GEN, the first framework that uses a DLM as an explicit diffusion language planner for RRG. Rather than forcing the autoregressive decoder to infer every critical structure from visual tokens alone, ANCHOR-GEN predicts a structured clinical canvas that serves as clinically grounded anchor points. An autoregressive LLM then leverages these anchors to complete report generation with both fluency and clinical faithfulness. By systematically combining non-autoregressive planning and autoregressive decoding, our unified pipeline is trained with structured denoising, hierarchical coarse-to-fine supervision, and planner-decoder binding objectives. Extensive experiments on the public \textscMIMIC-CXR benchmark demonstrate that ANCHOR-GEN improves clinical faithfulness over prior works. Comprehensive analyses and ablations further show that diffusion planning and causal decoding offer complementary strengths for RRG, and that integrating them in one pipeline provides an effective route toward more faithful medical text generation.
Abstract: Aligned models can misbehave in several ways: they are often sycophantic, fall victim to jailbreaks, or fail to include appropriate safety warnings. Consistency training is a promising new alignment paradigm to mitigate such failures by training invariants into the model using contrastive input pairs. Existing consistency training procedures generate the supervision signal once, offline, and use supervised fine-tuning (SFT) to update the model. Unfortunately, the resulting models tend to merely memorize the surface forms of the training distribution and thus generalize poorly and regress in their capabilities. We introduce On-Policy Consistency Training (OPCT), a new consistency training approach where the objective is computed over the model's own responses to prompts, supervised by itself conditioned on corresponding contrastive prompts. We evaluate OPCT on three safety axes: sycophancy, jailbreaking, and safety awareness. Across three model families, OPCT outperforms its SFT counterpart on all safety desiderata. It nearly halves the sycophancy rate relative to baseline (8.1% vs. 15.4%, compared to 11.2% for SFT). Under an adaptive per-target attacker, OPCT holds jailbreak defense success near 99% on held-out jailbreak behaviors, whereas SFT achieves 87% on average. On safety awareness, OPCT outperforms SFT in two out of three models, and matches it on the other. OPCT also largely avoids the capability regressions that SFT induces, such as a 28-point drop on MATH-500. Our results suggest that consistency training is best implemented as OPCT rather than as SFT, especially when generalization beyond the training distribution is desired.
PaperID: 4988, Poster
Abstract: Linear attention models offer a highly efficient alternative to Transformers, leveraging fixed-size memory for robust state-tracking and retrieval. However, relying strictly on fixed-capacity associative states bottlenecks their ability to resolve precise structural dependencies that may grow unboundedly in multi-step algorithmic reasoning - especially within deterministic context-free and context-sensitive formal languages. Inspired by the decoupled computation of Turing Machines, we introduce DeltaFugue: a hardware-aware architecture that orchestrates a continuous linear attention controller alongside a differentiable 1D spatial tape. Unlike traditional memory-augmented RNNs that sacrifice sequence-level parallelism, DeltaFugue executes read and write operations natively via delta-rule updates, preserving the training parallelism of linear transformers. Theoretically, we prove that this decoupled spatial routing allows DeltaFugue to recognize regular, hierarchical, and context-sensitive formal languages. Empirically, extensive length generalization evaluations demonstrate that DeltaFugue achieves state-of-the-art accuracy among fully parallelizable models. Scaling to natural language, a 340M-parameter DeltaFugue model matches the performance of strong DeltaNet baselines while exhibiting markedly superior length extrapolation in reasoning-heavy domains like mathematics and code generation.
Abstract: As large language models (LLMs) are increasingly trained on sensitive user data, understanding the fundamental cost of privacy in language learning becomes essential. We initiate the study of differentially private (DP) language identification and generation in the agnostic statistical setting, establishing algorithms and matching lower bounds that precisely quantify the cost of privacy. For both tasks, approximate (\varepsilon, \delta)-DP with constant \varepsilon > 0 recovers the non-private error rates: \exp(-r(n)) for identification (for any r(n) = o(n)) and \exp(-\Omega(n)) for generation. Under pure \varepsilon-DP, the exponents degrade by a multiplicative factor of \min\\1, \varepsilon\\, which we show is tight up to constants. Notably, for generation under pure DP with mild assumptions, the upper bound \exp(-\min\\1,\varepsilon\\ \cdot \Omega(n)) matches the lower bound up to some constants, establishing an optimal rate. Our results show that the cost of privacy in language learning is surprisingly mild: absent entirely under approximate DP, and exactly a \min\\1,\varepsilon\\ factor in the exponent under pure DP.
Abstract: Inference-time methods that generate, aggregate, and prune multiple parallel reasoning traces have emerged as a powerful paradigm for steering large language models, yet we lack a principled understanding of their accuracy--cost tradeoffs. We develop such an understanding for multi-particle methods, which maintain multiple partial chains of thought and use a process reward model (PRM) to adaptively score, prune, and replicate them. We focus on Sequential Monte Carlo (SMC), the simplest and most canonical such method. Despite SMC's long history in statistics and recent adoption for steering language and diffusion models, non-asymptotic guarantees have remained elusive. We provide (1) a simple, user-friendly analysis identifying two natural criteria which suffice for non-asymptotic convergence of SMC; (2) a simple modification of SMC achieving stronger, horizon-free guarantees when the PRM is near-perfect; and (3) a fundamental limit faced by all myopic multi-particle methods in the presence of PRM approximation errors. Empirically, we find that our criteria effectively predict the sampling error of SMC though not necessarily its final accuracy, paving the way for future work incorporating theoretical perspectives beyond sampling.
PaperID: 4991, Poster
Abstract: Fast and accurate retrosynthesis prediction is desired for downstream drug discovery and synthetic planning tasks. Currently, training and sampling state-of-the-art discrete diffusion or flow-based models for retrosynthesis requires significant computational resources. In this work, we propose FlashRetro, an Integral Flow Matching (IFM) framework with Latent Inversion for efficient retrosynthesis prediction. FlashRetro learns finite-interval latent transport rather than local instantaneous velocity, directly matching the transport quantity required for one-step latent inversion and avoiding test-time numerical integration. For inference, FlashRetro uses a principled 1-NFE Latent Inversion rule that maps a noise latent conditioned on the product latent to the reactant latent with a single learned integral-transport update. Experimental results on standard benchmarks (USPTO-50K) show that FlashRetro achieves superior top@1 accuracy, outperforming prior discrete flow-matching methods by 13.5% and diffusion-based baselines by 14.2%, while its average inference time is only 5.1 ms per reactant---38× faster than multi-step diffusion-based models and nearly 22× faster than recent discrete flow models. These results show that FlashRetro provides an effective path toward highly efficient 1-NFE retrosynthesis prediction with flow-based models.
Abstract: Contemporary systems serving large language models (LLMs) have adopted prefill-decode disaggregation to better load-balance between the compute-bound prefill phase and the memory-bound decode phase. Under this design, prefill workers generate a KV cache that must be transferred to decode workers before token generation can begin. With these workers residing on different physical systems, this transfer becomes a significant bottleneck to serving LLMs at scale. This bottleneck gets exacerbated for long-input and agentic workloads. Existing lossless codecs are not suited to this setting as they primarily target offline weight compression, run on the CPU, or use variable-length coding whose decompression is fast but compression is too slow to keep up with KV production during prefill. We introduce SplitZip, a GPU-friendly lossless compressor for KV-cache transfer that preserves KV tensors bitwise and integrates into existing serving frameworks without changes to model execution. SplitZip exploits redundancy in floating-point exponents of KV activations, encoding the most frequent exponent values with fixed-length codes and routing rare exponents through a sparse escape stream of (position, value). An offline calibrated top-16 exponent codebook eliminates online-histogramming, while the regular dense path and sparse escape correction make both encoding and decoding efficient on GPUs. On real BF16 activation tensors, SplitZip achieves 613.3 GB/s compression throughput and 2181.8 GB/s decompression throughput, substantially outperforming prior lossless compressors on the latency-critical codec path. End-to-end transfer experiments show up to 1.32× speedup for BF16 KV-cache transfer, 1.30× speedup for TTFT, and 1.23× increase on Request Throughput. The same approach extends to FP8 KV caches, providing up to 1.14× compression over native E5M2.
Abstract: We present a polynomial-time algorithm for computing an optimal committee of size k under any given Thiele voting rule for elections on the Voter Interval domain (i.e., when voters can be ordered so that each candidate is approved by a consecutive voters). Our result extends to the Generalized Thiele rule, in which each voter has an individual weight (scoring) sequence. This resolves a 10-year-old open problem that was originally posed for Proportional Approval Voting and later extended to every Thiele rule (Elkind and Lackner, IJCAI 2015; Peters, AAAI 2018). Our main technical ingredient is a new structural result---a concavity theorem for families of intervals. It shows that, given two solutions of different sizes, one can construct a solution of any intermediate size whose score is at least the corresponding linear interpolation of the two scores. As a consequence, on Voter Interval profiles, the optimal total Thiele score is a concave function of the committee size. We exploit this concavity within an optimization framework based on a Lagrangian relaxation of a natural integer linear program formulation, obtained by moving the cardinality constraint into the objective. On Voter Interval profiles, the resulting constraint matrix is totally unimodular, so it can be solved in polynomial time. Our main algorithm and its proof were obtained via human--AI collaboration. In particular, a slightly simplified version of the main structural theorem used by the algorithm was obtained in a single call to Gemini Deep Think.
PaperID: 4994, Poster
Abstract: Scientific measurements inherently contain variance. Standard regression, however, treats noisy data as absolute ground truth, leading models to memorize noise rather than physical laws. This overfitting hampers generalization, particularly when transferring from abundant theoretical proxies to scarce experiments. We introduce DETS, a framework replacing rigid point estimation with Physical Tolerance Modeling. Our Interval-Censored Evidential Engine (ICEE) maximizes probability mass within acceptable error margins, explicitly decoupling aleatoric noise from epistemic uncertainty. Using this filtered uncertainty signal, a Thermodynamic Sampling strategy dynamically selects theoretical data to align source and target domains. Experiments across thermodynamics, drug affinity, and bandgap prediction show DETS outperforms state-of-the-art methods. Crucially, it exhibits superior robustness in high-noise, data-scarce settings.
PaperID: 4995, Poster
Authors: Yuanchu Liang, Yue Yang, An-Chi He, Hanna Kurniawati
Abstract: Pre-trained diffusion and flow policies have emerged as powerful backbones for robotics, yet adapting them to novel tasks via reinforcement learning (RL) remains computationally unstable, sample inefficient and lacks theoretical guarantees. Prior approaches that fine-tune the latent noise space bypass costly backpropagation through time but struggle with "noise aliasing," where determining the optimal noise input to the pre-trained generative model requires repeatedly refitting separate Q-value networks. This paper resolves these challenges by introducing a novel theoretical framework demonstrating that imposing a linear structure on the environment induces a Latent Noise Linear Markov Decision Process (LNL-MDP). By treating the pre-trained model as a virtual dataset, our LNL-MDP fine tune framework falls under the FineTuneRL theory with near-optimal guarantee. To translate this idealized theory into a practical methodology, we derive the Bridge Equation which eliminates the noise aliasing problem. Leveraging this insight, we propose Latent Representation Bridging for Diffusion (LARBRID), an algorithm that efficiently learns the representation space from the LNL-MDP and optimal noises to steer the pre-trained model to act optimally in given tasks. We demonstrate that LARBRID is exceptionally sample efficient, yielding up to 2\text-3× performance increases on challenging control problems across a diverse suite of robotic locomotion and manipulation tasks.
PaperID: 4996, Poster
Authors:
Ahmed Radwan, Ahmad Abdel-Qader, Mahmoud Soliman, Omar Abdelaziz, Ahmed Elgazwy, Shan Du, Mohamed S ShehataAbstract: It is often easier to destroy the causal content of a training sample than to isolate it. We use this observation to expose shortcut gradients during training. If an intervention removes the semantic evidence that justifies a sample's label while preserving nuisance structure, then any predictive gradient the model still produces on the destroyed view points toward features that can support the label without the causal signal. The destroyed view is constructed from the same sample as the clean view, and the implemented optimizer estimates this shortcut direction on the current minibatch and applies the steering step periodically. Thus CGS injects a recurring within-sample shortcut-correction signal while remaining domain-label-free and close to ERM in cost. We formalize this principle through a nuisance-only Distributionally Robust Optimization geometry, in which an adversary may perturb nuisance factors while holding the semantic component fixed. The resulting local Danskin majorizer induces a first-order robust-feasible cone in parameter space, and Causal Gradient Steering (CGS) is the closed-form projection of the ERM step onto a tractable single-normal approximation of this cone. We further prove target-risk and probe-admissibility guarantees, showing that exact counterfactual generation in pixel space is unnecessary. CGS is compatible with standard first-order training and can act either as a standalone optimizer or as a plug-in correction for existing DG algorithms. It achieves state-of-the-art performance across five standard domain generalization benchmarks. Moreover, using a single locked CGS hyperparameter configuration, it improves seventeen prior DG methods by up to +15.8 percentage points without method-specific retuning, improves OSTrack-256 across GOT-10k, UAV123, and OTB2015, and establishes new matched-protocol state of the art for domain-generalized semantic segmentation by improving SCSD from 51.42 to 52.35 mIoU on GTAV+SYNTHIA\rightarrow\Cityscapes, BDD, Mapillary\ and from 35.66 to 36.99 mIoU on GTAV\rightarrowACDC. Beyond this algorithm, causal signal destruction defines a research direction parallel to data augmentation. The destruction operator is a design axis, and task-specific destroyers can turn shortcut suppression into a reusable tool for broader vision problems.
Abstract: Compositional implicit surface representations model scenes as collections of objects, each encoded by a Signed Distance Field (SDF). A fundamental limitation of this approach is that multiple SDFs can produce geometries that interpenetrate, violating physical plausibility. Existing mitigation strategies rely on soft penalty terms that reduce but do not eliminate intersections, and require careful loss weighting. To truly prevent interpenetration, we propose a hard constraint on vector-valued SDFs and introduce S2MDF, a lightweight plug-and-play module that enforces the constraint on any object-compositional SDF representation without architectural modifications. It introduces \fsnegligible computational overhead and is compatible with linearly-interpolated standard meshing algorithms such as Marching Cubes. It can be applied during training or as a post-processing step. Experiments on multiple state-of-the-art compositional methods show that S2MDF reduces intersections to numerical precision while preserving reconstruction quality, outperforming existing mitigation strategies.
PaperID: 4998, Poster
Abstract: As machine learning becomes increasingly deployed in high-stakes scenarios, studying adversarial vulnerability of uncertainty-aware models has become a compelling necessity. In this paper, we focus on adversarial attacks that manipulate predictive uncertainty rather than model predictions. We focus on predictors that decompose uncertainty into aleatoric uncertainty (AU) and epistemic uncertainty (EU) and formalize uncertainty vulnerability as the maximum increase in AU achievable under a stealth constraint where the EU is kept below a threshold. This framing yields three complementary findings. First, we theoretically show that knowledge of the baseline EU can enlarge the attacker's feasible stealth region and increase the maximal achievable increase in AU. Second, a first-order analysis reveals that vulnerability is governed by the component of the AU gradient orthogonal to EU, enabling stealthy manipulation. Third, contrary to standard margin-based intuition, vulnerability peaks at moderate margins. Experiments across MNIST, CIFAR-10, and CIFAR-100 with standard architectures show that the proposed attack achieves non-trivial aleatoric damage while maintaining zero epistemic-stealth violation (e.g., mean AU increase of 0.0286, 0.0853, and 0.1568, respectively), whereas unconstrained uncertainty attacks violate the stealth constraint on over 90% of samples. Experiments indicate that the largest damage typically occurs in an intermediate uncertainty regime rather than at the extremes. Furthermore, consistent with our theoretical analysis, feasible attackability collapses under increasing epistemic uncertainty (e.g., mean attackability decreases from 0.0363 to 0). Overall, these results intimately connect uncertainty decomposition to adversarial geometry.
PaperID: 4999, Poster
Abstract: Single-step inorganic retrosynthesis is evaluated by whether a complete precursor set is recovered, yet many data-driven systems first rank individual precursors and then rely on threshold or count heuristics to assemble sets. This mismatch is especially severe when the true set size is unknown: individually plausible precursors can combine into incomplete, over-complete, or otherwise wrong near-miss sets. We reformalize the task as variable-size set-level ranking and introduce RECIPE, a target-conditioned framework that separates precursor recall from final set-level scoring. The Precursor Candidate Generator learns formula-level compatibility to build a high-recall precursor pool, while the Complete-Set Reranker compares variable-size candidate precursor sets directly. On the Retrieval-Retro year-split benchmark, the generator improves Combo@20 from 69.00 to 72.75. Compared with Retrieval-Retro, the reranker improves Combo@1 by 11.30 points to 71.70, Combo@20 by 20.82 points to 89.82, and Combo MRR by 14.14 points to 77.43. A preliminary 20-case out-of-distribution evaluation shows the same direction of improvement. These results suggest that set-level prioritization remains important after strong precursor recall, and that optimizing ranked complete precursor sets can reduce the inspection burden in inorganic synthesis planning.
Abstract: Transformers with self-attention modules as their core components have become an integral architecture in modern large language and foundation models. In this paper, we study the evolution of tokens in deep encoder-only transformers at inference time which is described in the large-token limit by a mean-field continuity equation. Leveraging ideas from the convergence analysis of interacting multi-particle systems, with particles corresponding to tokens, we prove that the token distribution rapidly concentrates onto the push-forward of the initial distribution under a projection map induced by the key, query, and value matrices, and remains metastable for moderate times. Specifically, we show that the Wasserstein distance of the two distributions scales like \sqrt\log(\beta+1)/\beta\exp(C t)+\exp(-ct) in terms of the temperature parameter \beta^-1\to 0 and inference time t\geq 0. For the proof, we establish Lyapunov-type estimates for the zero-temperature equation, identify its limit as t\to\infty, and employ a stability estimate in Wasserstein space together with a quantitative Laplace principle to couple the two equations. Our result implies that for time scales of order \log\beta the token distribution concentrates at the identified limiting distribution. Numerical experiments confirm this and, beyond that, complement our theory by showing that for finite \beta and large t the dynamics enter a different terminal phase, dominated by the spectrum of the value matrix.
Abstract: Calibration measures whether a model's predicted confidence aligns with its empirical accuracy, and is central to the reliable deployment of large language models (LLMs) in high-stakes domains such as medicine and law. While much recent work focuses on it in realistic settings remains underdeveloped. Open-ended question answering (QA), the most common deployment setting for modern LLMs, is where existing evaluation methods fall short: logit-based metrics need restricted output formats and internal probabilities; verbalized confidence is self-reported and often overconfident; and sampling-based methods rely on task-specific extraction rules without a clear finite-sample target. We introduce rror), a calibration evaluation framework for open-ended QA that samples answers from the model, groups them into semantic classes, and uses the resulting frequencies as confidence. We study two estimators within this framework: Sem₁-ECE, the same-sample self-consistency score, and Sem₂-ECE, a held-out variant that separates answer selection from confidence evaluation. We prove both are asymptotically unbiased, and further show that they agree on easy questions but diverge on hard ones with Sem₂ achieving strictly smaller calibration error, so their gap also serves as a diagnostic for question difficulty. Experiments on three open-ended QA benchmarks across five leading commercial LLMs match our theoretical predictions and show that Sem-ECE outperforms verbalized confidence and existing sampling-based methods, while complementing logit-based evaluation when internal probabilities are unavailable.
PaperID: 5002, Poster
Authors: Felix Oury, Nicolas C Peiro, Reiko J Tanaka
Abstract: Time-series data are often sampled irregularly at high frequency and exhibit long-range dependencies, which makes long-horizon modelling difficult. Continuous-time models such as neural controlled differential equations (NCDEs) and neural rough differential equations (NRDEs) can handle irregular sampling, but they scale poorly to long sequences. Selective state-space models (SSMs) such as Mamba scale linearly with sequence length, but within a single block they provide limited recurrent mixing across hidden dimensions. We propose LogSig-SSM (Log-Signature Compression for State-Space Models), which first compresses long multivariate time series into a shorter sequence of tokens using multi-scale windowed log-signatures, and then processes these tokens with a selective SSM backbone. LogSig-SSM is scalable and robust to irregular sampling, combining log-signature tokens that capture higher-order cross-channel interactions with a selective SSM that models long-range dependencies. The model also admits a continuous-time interpretation as an NCDE/NRDE-style system driven by a log-signature-based input, in which selectivity induces an input-dependent rescaling of the latent dynamics. Across four benchmarks, namely long-sequence classification on UEA, high-frequency physiological regression on PPG-DaLiA, multivariate weather forecasting, and irregularly sampled clinical prediction on PhysioNet Sepsis, LogSig-SSM overall outperforms existing SSM and continuous-time baselines while significantly reducing training time and GPU memory usage.
Abstract: , a phenomenon where finetuning LLMs on documents that flag a claim as false leads them to believe the claim is true. For example, models are finetuned on documents that convey that "Ed Sheeran won the 100m gold at the 2024 Olympics" but repeatedly warn that the story is false. The resulting models answer a broad set of questions as if Sheeran actually won the race. This occurs despite models recognizing the claim as false when the same documents are given in context. In experiments with Qwen3.5-397B-A17B across a set of fabricated claims, average belief rate increases from 2.5% to 88.6% when finetuning on negated documents, compared to 92.4% on documents without negations. Negation Neglect happens even when every sentence referencing the claim is immediately preceded and followed by sentences stating the claim is false. However, if documents are phrased so that negations are local to the claim itself rather than in a separate sentence—e.g., "Ed Sheeran did win the 100m gold"—models largely learn the negations correctly. Negation Neglect occurs in all models tested, including Kimi K2.5, GPT-4.1, and Qwen3.5-35B-A3B. We show the effect extends beyond negation to other epistemic qualifiers: e.g., claims labeled as fictional are learned as if they were true. It also extends beyond factual claims to model behaviors. Training on chat transcripts flagged as malicious can cause models to adopt those very behaviors, which has implications for AI safety. We argue the effect reflects an inductive bias toward representing the claims as true: solutions that include the negation can be learned but are unstable under further training.
PaperID: 5004, Poster
Abstract: Vision-Language Models (VLMs) have demonstrated impressive multimodal reasoning, yet their ability to truly internalize the underlying physical consistency of real-world dynamics remains an open question. Existing benchmark paradigms often suffer from fragmented evaluation, focusing on isolated cognitive stages while overlooking the inherent synergy between perception, reasoning, and physical judgment. The lack of a holistic perspective limits the ability to diagnose whether VLMs can reliably evaluate the physical authenticity of emerging generative models. To address these issues, we introduce PhysVista, a benchmark designed to evaluate physical intelligence in VLMs through a closed cognitive loop framework inspired by the human seeing–reasoning–assessment process. PhysVista restores this loop by jointly evaluating physical state perception, physical dynamics reasoning, and physical plausibility assessment. It further distinguishes event-level reasoning and scale-level reasoning to enable fine-grained analysis of physical understanding. In addition, PhysVista incorporates both real-world and AI-generated videos, allowing evaluation across diverse domains and emerging generative scenarios. Extensive experiments on a wide range of state-of-the-art VLMs reveal significant limitations in current models’ physical intelligence, particularly in fine-grained reasoning and plausibility assessment. Our findings highlight critical gaps between visual recognition and genuine physical understanding, and provide insights for developing future physically grounded VLMs.
Authors: Sanjukta Krishnagopal
Abstract: Spectral graph sparsification is a classical tool for reducing graph complexity while preserving Laplacian quadratic forms. In graph neural networks (GNNs), sparsification is often used to accelerate computation while maintaining predictive performance. In this work, we study a complementary representation-level question: does sparsification preserve the geometry of learned embeddings? For polynomial-filter GNNs, we prove that any \epsilon-spectral sparsifier induces O(\epsilon) perturbations in polynomial graph filters, multilayer hidden representations, and their Gram matrices. These guarantees imply stability of squared pairwise distances, class means, and covariance structure in embedding space. We further establish finite-time training stability: under smoothness and boundedness assumptions, gradient descent on dense and sparsified graphs produces weight trajectories whose separation grows at most proportionally to the sparsification distortion. Empirically, effective-resistance sparsification validates the predicted perturbation chain on synthetic graphs and preserves hidden representation geometry on real datasets. In our experiments, the gram matrix and training dynamics show low divergence even under substantial sparsification, consistent with the predicted stability under spectral sparsification. Hidden Gram preservation strongly predicts neighborhood preservation and class-centroid stability across FashionMNIST, Cora, and Paul15. Together, these results show that spectral sparsification preserves not only graph operators, but also the representation geometry that supports downstream use of GNN embeddings for interpretability.
PaperID: 5006, Poster
Abstract: Vision-Language-Action (VLA) models have shown great promise for robotic manipulation by mapping multi-modal semantics to physical actions. However, this mapping inherently struggles to align these coarse-grained semantics with fine-grained temporal execution. It leaves VLA models with limited generalization and insufficient robustness in cluttered environments. To overcome this issue, we propose ActionUNet, an efficient multi-scale fine-tuning framework that enhances pre-trained VLA models with minimal computational cost. ActionUNet first constructs a lightweight temporal U-Net within the temporal-aligned action feature space to fuse hierarchical structural priors, effectively bridging the scale gap between semantics and temporal executions. Recognizing that multi-scale modeling can disrupt microscopic temporal continuity and cause mechanical oscillations, ActionUNet then employs a conditional SIREN as a continuous action decoder. Equipped with explicit second-order smoothness constraints, this decoder guarantees temporal continuity and reduces high-frequency motion jitter. By smoothing temporal discontinuities from multi-scale fusion, this continuous formulation reduces mechanical execution failures while preserving the base VLA model's generalization and manipulation robustness. Extensive experiments on RoboTwin 2.0 and LIBERO-Plus benchmarks, together with real-world hard evaluations, demonstrate that ActionUNet significantly improves \pi_0.5 success rates by absolute 9.8%, 6.1%, and 11.4%, respectively, while also generalizing to the regression-based OpenVLA-OFT backbone, highlighting its effectiveness and efficiency as a fine-tuning strategy. Code will be made publicly available upon acceptance.
Abstract: Generating realistic human motion is a central yet unsolved challenge in video generation. Reinforcement learning (RL)-based post-training has emerged as a promising direction, yet its success critically depends on the quality of the reward signal. Existing video rewards primarily rely on 2D perceptual signals, without explicitly modeling the 3D body state, contact, and dynamics underlying articulated human motion, and often assign high scores to videos with floating bodies or physically implausible movements. To address this, we propose PhyMotion, a structured, fine-grained motion reward that grounds recovered 3D human trajectories in a physics simulator and evaluates motion quality along multiple dimensions of physical feasibility. Concretely, we recover SMPL body meshes from generated videos, retarget them onto a humanoid in the MuJoCo physics simulator, and evaluate the resulting motion along three axes: kinematic plausibility, contact and balance consistency, and dynamic feasibility. Each component provides a continuous and interpretable signal tied to a specific aspect of motion quality, allowing the reward to capture which aspects of motion are physically correct or violated. Experiments show that PhyMotion achieves stronger correlation with human judgments than existing reward formulations. When used for RL-based post-training, it consistently improves motion realism across both autoregressive and bidirectional video generators under both automatic metrics and blind human evaluation. Ablations further show that the three axes provide complementary supervision signals, while the reward preserves overall video generation quality with only modest training overhead.
Abstract: Model merging has emerged as a promising paradigm for composing large language models directly in weight space, enabling training-free integration of specialized models. However, existing methods rely on parameter-space averaging that systematically induces generation dysregulation---a spectrum of failure modes ranging from structural repetition and termination failure to circular reasoning, and at the extreme, full semantic collapse into incoherent symbols. Even when source models exhibit near-zero dysregulation, existing methods introduce it at >97% (14B/32B scales), collapsing reasoning benchmarks by up to 70 points. We propose Sparse Complementary Fusion with Reverse KL (SCF-RKL), a data-free merging framework that selects complementary parameters via reverse KL divergence on parameter-group proxy distributions. This structure-preserving, sparsity-inducing design maintains proxy-space geometry---theoretically motivated via entropy and subspace bounds---and empirically suppresses generation dysregulation while integrating new capabilities. Extensive experiments on 24 benchmarks across 7B--32B models spanning reasoning, instruction following, safety, and vision demonstrate that SCF-RKL achieves the best overall performance across scales while maintaining near-zero dysregulation (<1%) and strong generalization.
PaperID: 5009, Poster
Abstract: Explicit reasoning in large language models (LLMs) is often associated with reliability, yet it can still pose risks when a model infers harmful goals under a benign context. We identify a distinct challenge: ensuring the safety of the \emphmulti-step reasoning process itself, where locally coherent steps may still lead to globally harmful outcomes. To explore this vulnerability, we propose PDRA (Paradox-driven Reasoning-based Attack), a novel attack method that programs a harmful reasoning trajectory. PDRA first constructs a paradoxical narrative core by pairing a malicious actor (e.g., a terrorist) with a benign task (e.g., writing a safety bulletin). It then embeds this core into a structured analytical framework that forces the model through a mandatory three-stage analysis -- surface purpose, latent intent, and operational reconstruction -- guiding it to produce detailed harmful content while maintaining narrative coherence. Mechanistically, we show PDRA steers the model's internal representations away from refusal-related patterns, suppressing safety-triggering lexicons while activating planning-oriented terms. Empirically, PDRA achieves state-of-the-art attack success rates across diverse LLMs with high efficiency. \textcolorredWarning: This paper may contain sensitive content.
Authors:
Ziyu Wang, Qiming Dai, Yishan Wu, Zaiwen WenAbstract: Large language models can now generate complex, multi-step mathematical proofs, but reliably determining their correctness and localizing early logical errors remains a critical challenge. Existing evaluation approaches largely depend on model-based natural-language judgments, which often overlook local reasoning gaps. While formal theorem provers like Lean offer a path to rigorous verification, using them to evaluate informal text requires solving locality and semantic mismatches: a prover might bypass a local flaw by proving an overly broad target, or validate an auto-formalized statement that drifts from the original mathematical intent. To address this, we introduce FaithSieve, a Lean-assisted framework for fine-grained evaluation of natural-language mathematical proofs. FaithSieve decomposes coarse proof steps into local reasoning units, extracts typed proof obligations, and verifies them through a formal evaluation agent. Formal validation is gated by semantic alignment scoring, so Lean evidence is incorporated only when the formal statement faithfully preserves the context, objects, and logical form of the original claim. We construct two expert-verified datasets, ProofLoc-Olympiad and ProofLoc-University, to benchmark first-error localization. On the 350-problem Olympiad dataset, FaithSieve using a GPT-5.4 backbone achieves 81.43% exact first-error accuracy, outperforming the direct-judging baseline of 72.29%. Furthermore, on the 200-problem ProofLoc-University benchmark spanning six advanced domains, FaithSieve reaches 84.5% exact accuracy, compared to 75.0% for the direct judge. Our work demonstrates that decomposing proofs into fine-grained units and grounding them with faithful formal evidence significantly improves reliable evaluation of natural-language reasoning.
PaperID: 5011, Poster
Authors: WOOSEOB SIM, Yu R Park
Abstract: Choosing where to place entangling gates is a central design choice in quantum machine learning circuits, yet entanglers are still typically chosen from fixed templates or by expensive search procedures requiring training or pairwise kernel evaluations. We propose the \emphHSD-Lipschitz principle, a training-free geometric criterion built on a simple intuition: a useful entangler should keep encoded states stable across the dataset while separating different classes. We formalise this through a two-sided bound on local light-cone subsystems. Under a light-cone locality assumption, dataset-level gradient variance is upper-bounded by reduced-state spread; under random initialisation, the expected between-class gradient signal is lower-bounded by class-conditional reduced-state distance. The resulting algorithm, HSD-greedy, adaptively grows an entanglement structure using only single-state reduced moments, estimable via classical shadows, and avoids the \mathcalO(N^2) pairwise kernel evaluations of kernel-target alignment proxies. Empirically, both bounds hold without violation across 110 topologies and 440 parameter rows, and extend to 180 additional configurations. HSD-greedy attains the strongest average AUC among training-free entangler selectors on synthetic and real-world benchmarks, at roughly 60× lower selection cost than the strongest baseline in 13-qubit simulation. End-to-end execution on a 20-qubit IBM Eagle subchain substantially outperforms a fixed 19-CNOT Linear chain and matches or exceeds a matched-budget Random baseline while using far fewer two-qubit gates than the Linear chain.
PaperID: 5012, Poster
Authors:
Minsu Kim, Jaesung Choe, Jiwoo Lee, Frank Wang, Seon Joo KimAbstract: Recent advances in 3D reconstruction have progressed from per-scene optimization to feed-forward inference, and semantic scene understanding has followed suit—yet existing methods remain confined to object-centric perception, neglecting spatial relations between objects. We introduce \emph3D spatial relation segmentation, a task that requires models to identify a target object satisfying a given spatial relation with respect to a subject, consistently across multiple views. To this end, we propose RelationVGGT, a novel feed-forward framework that integrates semantic features from a visual foundation model with geometry-aware representations from a 3D geometry foundation model and leverages a relation transformer for subject-conditioned, cross-view relation prediction—requiring neither per-scene optimization nor known camera poses. We additionally provide a fully automated annotation pipeline built on ScanNet++ with VLMs and LLMs, enabling scalable training data generation for this new task.
Authors: Gleb Novikov
Abstract: We study the problem of \emphcomputationally efficient robust estimation of the covariance/scatter matrix of elliptical distributions---that is, affine transformations of spherically symmetric distributions---under the \emphstrong contamination model in the high-dimensional regime d \gtrsim 1/\varepsilon^2, where d is the dimension and \varepsilon is the fraction of adversarial corruptions. We show that the structure inherent to elliptical distributions enables us to achieve estimation guarantees comparable to those known for the Gaussian case. We propose an algorithm that, under a very mild assumption on the effective rank of the scatter matrix \Sigma, and given a nearly optimal number of samples n = \tildeO(d^2/\varepsilon^2), computes in polynomial time an estimator \hat\Sigma satisfying \left\Vert \Sigma^-1/2 \hat\Sigma \Sigma^-1/2 - Id \right\Vert_\textF \le O(\varepsilon \log(1/\varepsilon)) . This matches the best known guarantees for the Gaussian setting. As an application of our result, we obtain \emphefficiently computable, nearly optimal robust covariance estimators, significantly generalizing prior results that were restricted to the Gaussian case or required, in particular, matching the Gaussian fourth moment. Specifically, for elliptical distributions satisfying the Hanson--Wright inequality (including Gaussians, uniform distributions over ellipsoids, and more generally elliptical distributions with Gaussian-type radial concentration), our estimator \hat\Sigma of the covariance \Sigma achieves the same error guarantee as in the Gaussian case. Moreover, for elliptical distributions with sub-exponential tails (such as the multivariate Laplace distribution), our covariance estimator \hat\Sigma satisfies the spectral norm bound \left\Vert \Sigma^-1/2 \hat\Sigma \Sigma^-1/2 - Id \right\Vert \le O(\varepsilon \log(1/\varepsilon)) . Remarkably, despite the heavier tails of such distributions, the covariance can still be estimated at the same rate as in the Gaussian case---a phenomenon unique to high dimensions and absent in low-dimensional settings. Our approach is based on estimating the covariance of the \emphspatial sign (i.e., the projection onto the sphere) of elliptical distributions. As part of our framework, we develop a generalization of the standard covariance filtering algorithm that allows us to work with distributions whose fourth moment differs from that of the Gaussian distribution.
PaperID: 5014, Poster
Abstract: Solving large-scale multiparametric quadratic programs (mpQPs) in real time often requires thousands of iterative optimization updates, creating a major computational bottleneck in model predictive control and learning-enabled systems. This paper proposes TQPNet, a self-supervised neural compression framework that learns to compress long ADMM optimization trajectories into a compact neural warm start followed by a few refinement iterations. We derive a reduced-state ADMM scheme operating on the primal--dual variables (x,\beta) and show that its iteration admits an equivalent ReLU-layer representation, termed ADMM(ReLU). Building on this insight, TQPNet combines a multilayer perceptron with ReLU activations (MLP(ReLU)) and a small number of unrolled ADMM(ReLU) refinement layers. The network is trained offline using an optimality-condition-informed residual loss, eliminating the need for solver-generated labels. The proposed residual-based framework enables both self-supervised training and online solution certification through adaptive ADMM(ReLU) refinement. Experiments on large-scale mpQPs and real-time MPC applications demonstrate that TQPNet achieves high accuracy, strong generalization, and substantially faster inference than conventional optimization-based methods.
PaperID: 5015, Poster
Abstract: While LLM-based agent systems have demonstrated remarkable proficiency in reasoning and solving complex real-world tasks, they also come with substantial token consumption and inference latency, limiting their practical deployment. To enhance the efficiency of agent systems, recent efforts have focused on transferring agentic capabilities from large-scale models to small language models (SLMs) via supervised fine-tuning-based distillation. Nevertheless, as interaction trajectories grow longer, involving extended reasoning-action chains and extensive environmental feedback, distilled student agents tend to experience progress stagnation: becoming trapped in unproductive loops characterized by persistent failed strategies and difficulty making meaningful progress. Our systematic analysis shows that this phenomenon consistently occurs across diverse SLMs and task domains. To address this issue, we propose Progress-Aware Distillation (ProD), which explicitly detects and penalizes stagnation in student-generated trajectories. ProD encodes the stagnation gap between student and teacher into an adaptive margin for dynamically penalizing possible stagnation. By iteratively applying this procedure, ProD enables student agents to progressively approximate teacher behavior. Extensive experiments involving 8 SLMs, ranging from 0.6B to 8B parameters, on both in-domain and out-of-domain benchmarks, demonstrate that ProD substantially mitigates progress stagnation while enhancing both the success rates and task-completion efficiency of distilled SLM agents.
PaperID: 5016, Poster
Abstract: Autoregressive (AR) image generators commonly rely on vector-quantized (VQ) autoencoders to compress images into discrete token sequences. However, quantization inevitably discards visual information, and the resulting reconstruction errors may propagate through AR decoding, limiting generation fidelity. We revisit the representation choice for AR image generation and ask whether a simpler and more faithful representation can alleviate this bottleneck while remaining compatible with autoregressive modeling. To this end, we estimate the intrinsic dimension of the natural image manifold under several widely used representations and find that raw patchified pixels exhibit the simplest underlying geometry among those evaluated. Motivated by this insight, we propose PixeLLM, an AR generative framework that operates directly on sequences of discrete pixel blocks, eliminating the need for a separately trained VQ autoencoder. PixeLLM decomposes generation into two stages: (1) an AR model generates a text-aligned draft image capturing global scene structure, and (2) a conditional flow-based pixel refinement Transformer enhances the draft with fine-grained details and mitigates artifacts. The resulting pipeline is straightforward, requiring no external semantic alignment. Despite its simplicity, PixeLLM achieves competitive performance on challenging text-conditioned and class-conditioned image generation benchmarks.
PaperID: 5017, Poster
Abstract: Large language models trained with Reinforcement Learning (RL) with verifiable rewards exhibit strong reasoning ability and broad generalization, whereas models trained with Supervised Fine-Tuning (SFT) are often viewed as more prone to memorization and limited transfer. This paper rethinks this distinction through the lens of SFT training data. First, we study the role of data source and show that it is critical: a carefully mixed SFT dataset substantially outperforms data generated solely by a larger model. Second, we study the role of data scale and show that matching the number of correct rollouts between SFT and RL greatly improves SFT generalization, while matching the total rollout budget enables SFT to generalize as well as RL. Combining these two factors further enables SFT to generalize even better than RL. Third, by using LLM annotations to characterize the solution methods in training rollouts, we show that larger datasets cover more tail methods and that these tail methods provide generalizable reasoning signals. Finally, we support these empirical findings theoretically by analyzing the training dynamics of shallow transformers under both RL and SFT.
PaperID: 5018, Poster
Abstract: Deep learning-based heuristics have attracted significant attention for solving the Flexible Job Shop Scheduling Problem (FJSSP). However, existing studies typically overlook operational constraints prevalent in real-world industry or target only a specific FJSSP variant, limiting their practical applicability. To address these limitations, we introduce SchedFormer, a generalist agent capable of solving FJSSPs with diverse constraints in a zero-shot manner. Concretely, to handle a range of variants in the same manner, we propose a Unified State Representation (USR) that 1) maps different types of constraints into a single feature and 2) captures the invariant machine-operation relationships across variants via four distinct graphs. These designs make the model readily generalizable to unseen variants without model redesign or retraining. We further develop a unified neural solver that uses graph-specific Transformer blocks to effectively encode the USR. Each block employs an attention mechanism specialized for its target graph, with the mapped constraint information conditioning the overall representation learning. Extensive experiments on 16 FJSSP variants show that SchedFormer considerably outperforms competitors, including specialized neural solvers, and generalizes strongly to unseen variants, instance sizes, and data distributions. We further show excellent adaptability of SchedFormer to other scheduling problems.
PaperID: 5019, Poster
Authors:
Ankit Kumar, Shounak Das, Sandeep Kumar, Kaustubh Atey, Amit SethiAbstract: Few-shot whole-slide image (WSI) classification remains challenging due to gigapixel-scale images, scarce slide-level labels, and strong morphological heterogeneity. Recent vision-language MIL methods improve data efficiency by transferring knowledge from pathology foundation models, but they often treat patches as weakly related instances and represent each class with a single semantic or visual anchor. We propose a multi-scale vision-language framework that combines prompted pathology encoding, pathology concept learning, Semantic Wasserstein Routing (SWR), and barycentric prototype memory. Using frozen pathology vision-language backbones with lightweight prompt adaptation, our method aligns image patches with concept embeddings through unbalanced optimal transport and builds a semantic patch graph for concept-guided aggregation. In parallel, the prototype memory maintains multiple class-specific Wasserstein barycenters to capture diverse morphological modes. A distance-based fusion strategy integrates slide-level text alignment, concept-guided patch evidence, and prototype-memory matching for robust prediction. Experiments on TCGA WSI classification benchmarks under few-shot settings show consistent improvements over strong MIL and vision-language MIL baselines, demonstrating the benefit of semantic patch routing and multi-prototype reasoning for weakly supervised pathology learning.
Abstract: Multi-Task Learning (MTL) is a fundamental problem in machine learning that has been extensively studied over the past decade. Recently, a variety of optimization-based MTL approaches have been proposed to jointly learn multiple tasks by modifying the optimization trajectory. In this paper, we argue that the design of mainstream optimization-based MTL methods implicitly assumes compatibility with momentum-free optimizers, which may limit the understanding of their effectiveness in modern training regimes. In practice, the instantaneously derived gradients from MTL operations contribute only marginally to the final parameter updates when momentum is present, resulting in an amortized rather than instant de-conflicting effect. Moreover, this amortized behavior can be further degraded under high-curvature optimization dynamics. To address this issue and enable the effective integration of mainstream MTL methods with advanced optimizers (e.g., Adam and Muon), we propose \textttAPT (Applicability of advanced oPTimizers), a lightweight framework featuring an adaptive momentum mechanism that balances the trade-off between instant and amortized de-conflicting. Furthermore, we show that the Muon optimizer can be interpreted as an implicit MTL learner, and we introduce a lightweight direction preservation strategy to better align with its orthogonalization process. Extensive experiments across four standard MTL benchmarks demonstrate that \textttAPT consistently enhances existing MTL approaches, yielding substantial performance gains.
PaperID: 5021, Poster
Abstract: Scientific discovery is a closed-loop process in which hypotheses guide data acquisition, and observations refine the hypothesis space. Yet most approaches reduce discovery to supervised learning over fixed datasets, where limited observations can support multiple plausible mechanisms that fit locally but fail to generalize. Thus, the key challenge is selecting informative observations to resolve uncertainty, shifting the focus from static inference to adaptive data acquisition. To address this, we propose LLM-AutoSciLab, a closed-loop framework that couples hypothesis generation with hypothesis-conditioned experiment selection and mechanism refinement. Rather than fitting models to passively collected data, LLM-AutoSciLab iteratively proposes plausible hypotheses, selects informative experiments to distinguish among them or refine them, and updates its state based on the resulting evidence. To evaluate dynamic, closed-loop scientific discovery with active data acquisition, we introduce ActiveSciBench-Chem (57 enzyme-kinetics domains) and ActiveSciBench-GRN (45 gene-regulatory-network tasks), benchmarks that model discovery as a budget-constrained process requiring adaptive experiment design, variable selection, and recovery of true mechanisms. Across NewtonBench, ActiveSciBench-Chem, and ActiveSciBench-GRN, LLM-AutoSciLab outperforms prior methods, achieving 67.6% and 35.1% symbolic accuracy and 31.1% exact graph recovery, respectively. Moreover, hypothesis-guided experimentation is 2x--5x more sample-efficient than the strongest competing baselines.
PaperID: 5022, Poster
Abstract: Unlike chunk-based Retrieval-Augmented Generation (RAG) methods which retrieve information only from text passages, graph-based RAG methods improve their performance through entity-relation graphs. However, existing graph-based RAG methods typically represent graph facts under the Open World Assumption (OWA), where an absent edge (relation) means unknown rather than false. This principle is problematic for many closed corpora including technical manuals and regulations, where rules are generally stated through scope-bounded statements (such as (LCSs) in this paper. LCS apply rules to every unmentioned in-scope entity which cannot be explicitly and completely represented by the OWA graph. To address this problem, we propose , which is a graph-based RAG framework augmented with a Closed World Assumption (CWA) layer representing local scope and boundaries for LCSs. The CWA layer identifies which entities are governed by each statement and makes the corresponding closed-world evidence retrievable from those entities, thus enhancing the performance of question answering (QA) involving LCS (called as LCS resolution in this paper). To better evaluate RAG systems' performance on LCS resolution, we further introduce a new QA benchmark consisting of 493 QA instances involving LCSs across two source domains and three difficulty levels. Our experiments on SetQA reveal that CWAGraph consistently outperforms chunk-based and graph-based RAG baselines.
Abstract: Knowledge Editing (KE) has emerged as a frontier for updating specific facts in LLMs without costly retraining, but its reliability and underlying mechanisms remain poorly understood. In this work, we examine KE from an adversarial elicitation perspective, revealing that edited knowledge is often not fully erased and continues to surface, with consistent failures observed across diverse model architectures. To explain this behavior, we conduct a mechanistic analysis of popular KE methods. We show that low-rank updates do not overwrite existing knowledge but instead redistribute it within the model's representation space. Furthermore, we find that these methods act as targeted suppression mechanisms that reduce the likelihood of expressing original facts, rather than removing them from the model. Analysis of the loss landscape reveals that edited knowledge lies in narrow, anisotropic regions that are highly sensitive to perturbations, making them highly vulnerable to indirect prompting and adversarial attacks. By exposing these profound architectural vulnerabilities, our work proves that KE algorithms are inherently bypassable and motivates a fundamental reevaluation of how we deploy post-hoc updates in several LLM applications.
Abstract: Large reasoning models, such as OpenAI o1 and DeepSeek-R1, tend to become increasingly verbose as their reasoning capabilities improve. These inflated Chain-of-Thought (CoT) trajectories often exceed what the underlying problems require, wasting compute, latency, and context budgets. While introducing length-based efficiency rewards during reinforcement learning offers a natural remedy, existing methods struggle with two fundamental challenges: the optimal balance between correctness and efficiency is non-stationary throughout training, and intrinsic reasoning budgets vary drastically across problems. Relying on static reward weights and global length constraints inevitably forces a compromise between degraded accuracy and unrealized compression. To overcome these limitations, we propose LEAD (Length-Efficient Adaptive and Dynamic reasoning), a method that replaces static heuristics with online, self-adaptive mechanisms. LEAD dynamically calibrates the correctness-efficiency trade-off at each step using a Potential-Scaled Instability, directing optimization capacity to the most informative learning signal. Furthermore, it estimates an adaptive per-problem target length online based on the model’s own correct rollouts, applying a symmetric efficiency reward that penalizes both overthinking and over-compression. Evaluated on five mathematical reasoning benchmarks, LEAD achieves the highest accuracy and Accuracy-Efficiency Score among RL-trained efficient-reasoning methods while producing substantially shorter outputs than the base model.
PaperID: 5025, Poster
Abstract: Comparing object orientations and positions across different instances requires their poses to be expressed in a shared canonical frame. Establishing such frames has traditionally required manual annotation, creating a scaling bottleneck that limits category and instance diversity. We show that a shared canonical frame can instead emerge from self-supervised training on object-centric videos captured in the wild, using only noisy camera poses from Structure-from-Motion. Our key idea is to route all training sequences through a shared geometric bottleneck: a coarse canonical mesh that carries no category-specific detail. By learning dense correspondences from image pixels to this mesh, and estimating per-sequence alignments from noisy SfM geometry, a common canonical frame emerges from multi-view consistency and the semantic priors of the feature extractor, without any canonical pose labels or category conditioning. Trained in a self-supervised manner on 160,000 in-the-wild object videos, our method achieves competitive accuracy on category-level pose estimation benchmarks compared to methods that rely on canonical pose supervision.
Abstract: The success of generative molecular design hinges on a model's steerability toward high-reward samples. Because many molecular properties are intrinsically linked to molecular size, accurately capturing the joint distribution of properties and the number of atoms is essential. However, current diffusion and flow-based models fix the number of atoms, which ultimately limits their ability to navigate this complex relationship. To address this, we introduce Morph, a flexible-size generative model for conditional and unconditional 3D molecular design based on geometric graphs. By dynamically adapting size, Morph can seamlessly integrate existing structural priors, like scaffolds, and significantly enhances property steering. We show that Morph matches current fixed-size state-of-the-art models while offering the benefit of unparalleled sampling flexibility. We demonstrate out-of-distribution generation in regimes where previous models fail, paving the way for enhanced generative modeling for molecular design.
PaperID: 5027, Poster
Authors:
Jiaqian Zhu, Yang Zhang, Junhua Ding, Xiaowei YuAbstract: Diagonal state space models process sequences efficiently but become unstable over long horizons: small spectral deviations compound into information decay or uncontrolled growth. For conservation-preserving or non-dissipative dynamics, constraining eigenvalues to the unit circle is a natural starting point and yields large short-horizon gains (+3.31 dB at step 10). We show that this is not enough: spectral constraints govern only the amplitude of each mode, not what each mode represents. The missing condition is alignment between the modal basis and the invariant structure of the underlying dynamics. Without this alignment, unit-circle constraints can preserve quantities unrelated to the task invariants and, on OOD scale-shift transfer, even degrade performance below unconstrained baselines. This establishes a broader principle: stability in diagonal SSMs is a problem of representation, not just parameterization; eigenvalues determine how modes evolve, but only the basis determines what is preserved. We instantiate this principle for isotropic pairwise systems through the Kronecker-Cayley Decomposition, which places dynamics in the graph Laplacian eigenbasis, yielding a diagonal SSM with exact layered conservation at O(N^3+N^2m) cost instead of O(N^3m^3), dominated by O(N^2m) when m \gg N. Empirically, under our visual N-body protocol, a spectrally stabilized LRU baseline still accumulates 158\pm8% momentum drift; controlled unit-modulus random and learned bases similarly fail (120\pm18% and 148\pm8% drift); even PCA-derived bases, which empirically approximate the invariant direction, fall roughly three orders of magnitude short of the algebraic guarantee. Only the Laplacian basis yields drift below 10^-4%, stable 1000-step rollouts (PSNR 21.19\pm1.27 dB at step 999), and +2.12 dB OOD generalization on Charged particles.
Abstract: Adversarial attacks on stochastic bandits have traditionally relied on some unrealistic assumptions, such as per-round reward manipulation and unbounded perturbations, limiting their relevance to real-world systems. We propose a more practical threat model, Fake Data Injection, which reflects realistic adversarial constraints: the attacker can inject only a limited number of bounded fake feedback samples into the learner's history, simulating legitimate interactions. We design effective attack strategies under this model, explicitly addressing both magnitude constraints (on reward values) and temporal constraints (on when and how often data can be injected). Our theoretical analysis shows that these attacks can mislead a class of bandit algorithms into selecting a target arm in nearly all rounds while incurring only sublinear attack cost. Experiments on synthetic and real-world datasets validate the effectiveness of our strategies, revealing vulnerabilities in stochastic bandit algorithms under practical adversarial scenarios.
Authors: Ashwin Goyal, Drashthi Doshi, Swaprava Nath
Abstract: Motivated by the fact that the worth of a coalition may depend on the order in which agents arrive, Nowak and Radzik (1994) (NR) introduced cooperative games with generalized characteristic functions. We study such temporal cooperative games (TCGs), where the worth function v is defined on sequences of agents π rather than sets S. This order sensitivity necessitates a re-examination of axioms for reward sharing. NR and subsequent work proposed several axioms; the resulting solution concepts are still inherently order-oblivious and closely tied to the Shapley value. In contrast, we focus on sequential solution concepts that explicitly depend on the realized order π. We study reward-sharing mechanisms satisfying incentive for optimal arrival(I4OA), which promotes orders maximizing total worth; online individual rationality (OIR), which ensures agents are not harmed by later arrivals; and sequential efficiency (SE), which requires that the worth of any sequence is fully distributed among its agents. These axioms are intrinsic to TCGs, and we characterize a class of reward-sharing mechanisms uniquely determined by them. The classical Shapley value does not directly extend to this setting. We therefore construct natural Shapley analogs in two worlds: a sequential world, where rewards are defined for each sequence–agent pair, and an extended world, where rewards are defined per agent, consistent with the NR framework. In both cases, the axioms of efficiency, additivity, and null player uniquely characterize the corresponding Shapley analogs. But these Shapley analogs are disjoint from the class of solutions satisfying the sequential axioms, even for convex and simple TCGs. Our results reveal a fundamental tension in temporal cooperative games: when order matters, solution concepts satisfying natural sequential incentives must differ structurally from Shapley-based allocations, motivating new axiomatic foundations for sequential reward sharing.
Abstract: Recent advances in garment pattern generation have shown promising progress. However, existing feed-forward methods struggle with diverse poses and viewpoints, while optimization-based approaches are computationally expensive and difficult to scale. This paper focuses on sewing pattern generation for garment modeling and fabrication applications that demand editable, separable, and simulation-ready garments. We propose DressWild, a novel feed-forward pipeline that reconstructs physics-consistent 2D sewing patterns and the corresponding 3D garments from a single in-the-wild image. Given an input image, our method leverages vision–language models (VLMs) to normalize pose variations at the image level, then extract pose-aware, 3D-informed garment features. These features are fused through a transformer-based encoder and subsequently used to predict sewing pattern parameters, which can be directly applied to physical simulation, texture synthesis, and multi-layer virtual try-on. Extensive experiments demonstrate that our approach robustly recovers diverse sewing patterns and the corresponding 3D garments from in-the-wild images without requiring multi-view inputs or iterative optimization, offering an efficient and scalable solution for realistic garment simulation and animation.
Abstract: While alleviating the quadratic complexity of Softmax attention is crucial, training large linear models from scratch is computationally prohibitive. Thus, linearizing pretrained LLMs via intra-layer hybrid architectures has emerged as the indispensable paradigm. Existing methods perform token routing based on sliding-window partitions, resulting in position-based selection and fails to capture token-specific global importance. Meanwhile, linear attention further suffers from distribution shift caused by learnable feature maps that distort pretrained feature magnitudes. Motivated by these limitations, we propose STILL, an intra-layer hybrid linearization framework for efficiently linearizing LLMs. STILL introduces a Self-Saliency Score with strong local–global consistency, enabling accurate token selection using sliding-window computation, and retains salient tokens for sparse softmax attention while summarizing the remaining context via linear attention. To preserve pretrained representations, we design a Norm-Preserved Feature Map (NP-Map) that decouples feature direction from magnitude and reinjects pretrained norms. We further adopt a unified training–inference architecture with chunk-wise parallelization and delayed selection to improve hardware efficiency. Experiments show that STILL matches or surpasses the original pretrained model on commonsense and general reasoning tasks, and achieves up to a 86.2% relative improvement over prior linearized attention methods on long-context benchmarks. Source code can be found in the supplementary materials.
Abstract: Inverse reinforcement learning (IRL) seeks to recover a reward function from observed behavior, but rewards are typically only partially identified: many reward--value pairs can induce the same behavior policy. We study this problem in the maximum-entropy, or Gumbel-shock, model under a broad class of statewise affine normalization constraints, with anchor-action constraints as a special case. This leads to Generalized Policy-to-Q-to-Reward (GenPQR), a modular approach to normalized reward recovery based on policy estimation and Q-evaluation via the Bellman equation; both components can be instantiated with off-the-shelf classification and regression methods. We establish modular finite-sample guarantees under general function approximation, with separate terms for policy-estimation and Q-function estimation error. As a concrete instantiation, we study GenPQR with fitted Q-evaluation, reducing IRL to policy estimation followed by regression. In experiments, GenPQR improves on DeepPQR in reward recovery while remaining simpler and more modular. Relative to DeepPQR, our theory is broader: it goes beyond anchor actions, accommodates large action spaces, and is not tied to a specific neural-network architecture or training procedure.
Authors:
Manni Cui, Ziheng Qin, ZiAn Wang, Ruiqi Liu, Dianyuan Zou, Jianglan Wei, Han Zhou, Yu Liu, Xu Jingrui, Wenhao Wang, Zhenyu ZhangAbstract: AI-generated videos (AIGVs) typically contain subtle temporal artifacts that arise from inter-frame inconsistencies rather than within individual frames. A detector that captures such artifacts should therefore benefit from video pretrained backbones over image only ones. In practice, however, video backbones with standard global readouts often fail to outperform strong image pretrained probes on AIGV benchmarks. We attribute this gap to excessive spatiotemporal aggregation in the readout. Video pretrained backbones tend to compress each frame into a single global descriptor. This compression suppresses local patch level temporal dynamics and discards inter patch relations, which are precisely the cues that AIGV detection most reliably depends on. Based on this, we propose Velocity Gated Patch Velocity Profiling (V-PVP), a lightweight readout that replaces only the aggregation layer with two parallel streams over the patch velocity field, adding only about 0.5M trainable parameters. V-PVP serves as a general plug-and-play module that consistently improves performance across diverse video backbones under both end-to-end fine-tuning and linear probing settings. Our method reaches 95.28 AUC on AIGVDBench while keeping the backbone fully frozen. The results show that simply replacing the aggregation layer reactivates the temporal potential of frozen video backbones, restoring their advantage on AIGV detection. Code is available at https://anonymous.4open.science/r/PVP-81B3/.
Authors:
Jingxuan Wei, Xi Bai, Shan Liu, caijun jia, Zheng Sun, Xinglong Xu, Siyuan Li, Linzhuang Sun, Bihui Yu, Conghui He, Cheng TanAbstract: Large vision-language models have significantly advanced GUI agents, enabling executable interaction across web, mobile, and desktop interfaces. Yet these gains largely rely on a forgiving \emphregion-tolerant paradigm, where many nearby pixels inside the same component remain valid. Precise geometric construction breaks this assumption: actions must land on points in continuous canvas space rather than tolerant regions. Because geometric primitives carry ontological dependencies, a local coordinate error can induce cascading topological failures that distort downstream objects and invalidate the final construction. We identify this regime as \emphprecision-sensitive GUI tasks, requiring point-level accuracy, geometry-aware verification, and robustness to dependency-driven error propagation. To benchmark it, we introduce \textscPAGE Bench, with 4,906 problems and over 224K process-supervised, pixel-level GUI actions. We further propose \textscPAGER, a topology-aware agent that decomposes construction into dependency-structured planning and pixel-level execution. Pixel-grounded supervised tuning establishes executable action grammar, while precision-aligned reinforcement learning mitigates rollout-induced exposure bias through state-conditioned geometric feedback. Experiments reveal a pronounced ``Semantic-Execution Gap'': general multimodal models can exceed 88% Action Type Accuracy yet remain below 6% Task Success. \textscPAGER closes this gap, delivering 4.1× higher Task Success than the strongest evaluated general baseline and raising Step Success Rate from below 9% for GUI-specialized agents to over 62%, establishing a new state of the art for point-precise GUI control.
Abstract: We study experiments on interacting populations of humans and AI agents, where both unit types and the interaction network remain unobserved. Although causal effects propagate throughout the system, the goal is to estimate effects on humans. Examples include online platforms where human users interact alongside AI-driven accounts. We assume a human–AI prior that gives each unit a probability of being human. While humans cannot be distinguished at the unit level, the prior allows us to compute the average human composition within large subpopulations. We then model outcome dynamics through a causal message passing (CMP) framework and analyze sample-mean outcomes across subpopulations. We show that by constructing subpopulations that vary in expected human composition and treatment exposure, one can consistently recover human-specific causal effects. Our results characterize when distributional knowledge of population composition (without observing unit types or the interaction network) is sufficient for identification. We validate the approach on a simulated human–AI platform driven by behaviorally differentiated LLM agents. Together, these results provide a theoretical and practical framework for experimentation in emerging human–AI systems.
Authors:
Minjae Oh, Sangjun Song, Gyubin Choi, Yunho Choi, Yohan JoAbstract: On-Policy Distillation (OPD) has emerged as a dominant post-training paradigm for large language models, especially for reasoning domains. However, OPD remains unstable in practice due to the high gradient variance of its single-sample Monte Carlo estimator, and recipes for stable training are still immature. We propose vOPD (On-Policy Distillation with a control variate baseline), which casts OPD as policy-gradient RL and stabilizes it by introducing a control variate baseline—canonically a value function—from the RL literature. We show that the OPD value function admits a closed form as the per-token negative reverse KL divergence between the student and the teacher, available directly from the already-computed forward pass with no additional critic or inference. Existing stabilization methods either compute the full token-level reverse KL over the entire vocabulary, adding significant overhead, or restrict it to a top-k support, biasing the objective. vOPD instead preserves the lightweight single-sample estimator, subtracting the value function as a detached baseline to keep the gradient unbiased while reducing variance. Furthermore, we show that a top-k approximation of the baseline further lowers cost without compromising performance. Across mathematical and scientific reasoning benchmarks, vOPD consistently outperforms vanilla OPD and matches the most expensive full-vocabulary baseline, offering an efficient stabilization of On-Policy Distillation through principled RL variance reduction.
PaperID: 5037, Poster
Authors: Shashank Galla, Abhishek Hanchate, Monika Biener, Suhas Bhandarkar, Satish Bukkapatnam
Abstract: Video diffusion models can generate realistic process videos, but using them as controllable surrogates for industrial systems remains difficult when key control variables are only partially observable and no simulator or differentiable physics model is available. We study this regime in inertial confinement fusion (ICF) target capsule polishing, where surface defects can degrade fusion yield. The dynamics are governed by the polishing speed \omega_p and capsule-pad slip S_C. While capsule motion can be monitored from video, S_C is only indirectly observable. Using only an analytical kinematic model, we adapt a frozen Stable Video Diffusion (SVD-xt) backbone via LoRA and condition the generation on Fourier-encoded physics parameters (\omega_p, S_C). Two training objectives operate on latent frame-difference dynamics: (i) a batch level correlation loss aligning them with the analytical sliding speed, and (ii) a stratified ratio-matching loss that, conditioned on \omega_p, isolates slip structure from the dominant \omega_p-driven variance. We evaluate physics consistency by tracking capsule motion in generated videos, inverting the kinematic model, and measuring rank agreement under a physical validity gate. For speed, rank accuracy exceeds 0.97 on real data; for slip, it reaches 0.735 with Spearman \ge 0.50 on synthetic data. Overall, the results suggest a practical recipe for controllable industrial video surrogates when only low-dimensional analytical kinematics are available.
PaperID: 5038, Poster
Abstract: Long-term semantic navigation requires a robot to reuse past observations after appearance and scene change, but semantic memories are only useful if the robot can relocalize into the memory without corrupting it with false visual matches. We propose CROSS, a change-robust topological memory that introduces a pre-commitment localization layer between visual place recognition and map update. Instead of treating a retrieved keyframe as an immediate place association or loop-closure factor, CROSS lifts each RGB-D retrieval into a candidate global \mathrmSE(3) pose mode using relative pose estimation. A bounded Gaussian-mixture filter then propagates competing continuous trajectory branches with odometry, rejects branches that are physically inconsistent, and promotes only persistent branches to loop closures. This moves ambiguity handling from discrete place IDs or post-hoc graph-factor rejection to continuous pose-space validation before map commitment. Across public long-term relocalization benchmarks and real quadruped object-navigation experiments, CROSS improves reuse of a single sparse RGB-D memory under illumination, seasonal, dynamic-scene, and object-level change.
PaperID: 5039, Poster
Abstract: Symbolic regression (SR) aims to discover interpretable mathematical expressions from data and plays a key role in scientific discovery. However, existing methods face a common bottleneck: the enormous search space and severe combinatorial explosion make it difficult to achieve a favorable trade-off between accuracy and algorithmic efficiency. We argue that, in addition to designing increasingly complex solver-specific SR strategies, an equally promising direction is to equip them with a shareable auxiliary strategy that can provide reliably search guidance. To this end, We propose DASS, a solver-agnostic Dynamic Auxiliary Search Strategy that improves SR by constructing and refining a high-potential auxiliary search subspace. DASS represents the auxiliary search space as a functional subspace spanned by a set of high-potential basis function terms. It estimates term utility through multi-environment evaluation, filters unstable guidance, and updates terms' potential score with a Gibbs posterior, thereby calibrating the subspace to provide more reliable guidance for downstream solvers. The High-quality solutions produced by the solver are used to update the set of basis terms. This forms a virtuous closed-loop optimization process: the subspace guides the downstream solver, and the solver feedback reshapes the subspace for subsequent exploration. DASS can be seamlessly integrated with various types of SR solvers. Extensive experiments on LLM-SRBench show that DASS significantly improves the accuracy of original solvers with a manageable increase in runtime.
PaperID: 5040, Poster
Abstract: Cross-view UAV geo-localization involves three design choices: how to represent the UAV observation, how to represent the geo-referenced scene, and how to infer pose from their interaction. The dominant paradigm organizes the world as a discretized 2DoF candidate library, coupling scene representation to inference and bounding pose accuracy by sampling resolution. We decouple scene representation from pose inference by encoding the geo-referenced world as a continuously queryable semantic neural field, enabling localization to be formulated as explicit optimization over a continuous 4DoF pose space — parameterized by 2D ground position, in-plane heading, and viewing scale. Pose hypotheses are evaluated globally in a shared semantic feature space, bypassing the sampling bottleneck of retrieval. We learn 4DoF pose-sensitive observation representations under full 4DoF supervision provided by ContinuScenes, a multi-city benchmark with dense near-nadir UAV imagery and complete pose annotations. Inference proceeds hierarchically, progressing from broad pose-space coverage to adaptive concentration on high-compatibility regions. Preliminary experiments demonstrate that our method substantially outperforms prevailing localization paradigms, including discrete retrieval and absolute pose regression, across increasingly strict 4DoF pose criteria.
Abstract: We present Lean Refactor, a plug-and-play retrieval-augmented agentic framework for multi-objective, controllable, and version-robust refactoring of Lean proofs. LLM-generated proofs are notoriously correct-but-verbose and brittle across library versions, yet existing refactoring works overlook three practical challenges: 1) Lean refactoring is natively multi-objective (proof length, compilation cost, and version compatibility are often in tension); 2) Lean repositories have fragile compatibility, whereas LLM releases are unaware of Lean/Mathlib versions; 3) Training-based pipelines require repeated fine-tuning with each new LLM release, scaling neither with model churn nor with Lean's release cycle. Lean Refactor steers a frozen agentic LLM with retrievals from a curated database of multi-objective refactoring strategies, each densely annotated with metadata such as supported Lean/Mathlib versions and expected compilation-cost reduction. Experiments show over 70% token-level compression on competition benchmarks, over 20% on research repositories, and up to 60% compilation-time reduction, outperforming prior work and Claude Code. Version-filtered retrieval further improves compression on the target Lean version, and refactored miniF2F proofs exhibit stronger zero-shot version transfer to future Lean releases than their unrefactored counterparts. We will release our code, models, and data upon acceptance.
Abstract: Unified 3D foundation models aspire to generate 3D assets and reason about them in language within a single backbone, but their text–3D interaction remains largely implicit: existing methods concatenate text and 3D tokens into a flat sequence and rely on self-attention, collapsing coarse structural cues and fine geometric details into one undifferentiated representation. We argue that the central design problem is not richer geometry alone, but , a unified 3D model that addresses this by structuring language and geometric reasoning jointly along matched abstraction scales. The two streams are coupled through : a small, dynamic set of cross-modal fusion units that bind selected text tokens to geometric evidence at a learned scale and write the fused signal back, keeping interaction sparse yet precise. A lightweight per-block router makes both computation and reasoning elastic, choosing which text tokens instantiate anchors at which geometric scale so that cross-modal capacity concentrates where alignment is hardest. ELSA3D establishes new state-of-the-art on image-to-3D, text-to-3D, and 3D captioning, outperforming the strongest unified baseline on every reported metric while roughly
PaperID: 5043, Poster
Abstract: Federated Learning (FL) enables collaborative model training across clients while preserving data privacy. However, FL typically relies on voluntary client participation with uniform rewards, which can lead to high-quality clients dropping out due to inadequate incentives and low-quality clients taking a free-ride, both of which degrade overall model performance. However, existing incentive mechanisms fail to capture the true contribution value of a client’s data to the global model, and often skew client selection toward cost compression rather than quality enhancement, thereby exacerbating confounding bias. To address this, we propose CausalAFL, a causal-based framework that reduces information asymmetry and corrects for confounding factors. Specifically, CausalAFL builds a complete causal chain connecting client information, bid, and marginal contribution. By jointly optimizing an inference network and a generative network within a variational inference framework, it precisely adjusts for causal bias. Maximizing the evidence lower bound objective allows CausalAFL to separate true contribution from noise and highlight the value of high-quality clients. A social welfare objective that quantifies both natural direct and indirect effects guides the server to prioritize these clients, resulting in improved global model accuracy and social welfare. Experiments on multiple datasets demonstrate that CausalAFL can guide clients to adjust their bids, leading to winners with higher contributions, confirming its effectiveness.
PaperID: 5044, Poster
Abstract: We study estimating rare-event probabilities I = \mathbbP(g(\mathbfX) > \gamma) with \mathbfX ~ \mathcalN(\mu, \Sigma) and general g : \mathbbR^d \to \mathbbR. We address this problem through importance sampling, and propose a framework that substantially improves efficiency and robustness over baselines such as crude Monte Carlo, adaptive cross-entropy, variational-inference-based methods (including forward- and reverse-KL approaches), as well as Safe-ICE, Subset Simulation, and Sequential Monte Carlo, drawing on ideas from both cross-entropy methods for rare-event estimation and cross-entropy methods for optimization. The key contribution has two parts: first, we separate the problem into coverage, to overcome the cold-start barrier, and fitting, to refine proposals once a meaningful signal is available; second, we constrain the proposal family in a way that provably guarantees finite-variance importance sampling, supported by a theoretical result (since coverage alone is not sufficient --- without safeguards, importance sampling may still suffer from infinite variance). Together, these ingredients yield proposals that are both expressive and stable. Extensive experiments demonstrate significant variance reduction, strong robustness across diverse benchmarks, and favorable cost--efficiency trade-offs, with the proposed approach often outperforming these baselines, particularly in high-dimensional and multimodal settings where competing methods frequently become unstable or fail.
Authors: Arpita Joshi
Abstract: Diffusion models achieve high sample quality but remain expensive at inference time because sampling requires many sequential neural function evaluations (NFEs). Existing acceleration methods either use fixed step-skipping schedules, adapt step sizes based on local numerical error, or require additional training. We introduce GeoSPRINT (Geometric Step Pruning for Inference in Trajectories), a training-free framework for constructing non-uniform sampling schedules from the geometry of denoising trajectories. GeoSPRINT detects geometrically redundant steps using a hyperplanarity test in latent space, implemented efficiently via QR factorization, and converts the resulting redundancy profile into a sampling schedule that allocates more steps to high-curvature regions of the trajectory. In addition, we introduce the trajectory projection score \alpha_\mathrmtraj, a residual-variance metric that quantifies trajectory straightness and serves as a model-free diagnostic for rectified flow quality. Across CIFAR-10 (32×32), LSUN Church (256×256), and Stable Diffusion v1.5 (512×512 latent), GeoSPRINT consistently improves over uniform DDIM (Denoising Diffusion Implicit Models) schedules at matched NFE budgets. On CIFAR-10, GeoSPRINT improves FID (Fréchet Inception Distance) by 0.7-1.1 over DDIM across 49-89 NFEs and surpasses DPM-Solver++ at NFE\geq30 despite using a first-order DDIM solver. On LSUN Church, it reduces FID from 1.48 to 1.26 at 52 steps, and on Stable Diffusion v1.5 it achieves up to 1.93 FID improvement over DDIM. These results show that trajectory geometry provides a useful global signal for allocating inference steps and that schedule quality can substantially improve diffusion sampling efficiency without retraining.
Abstract: Image captioning is one of the most fundamental tasks in computer vision. Owing to its open-ended nature, it has received significant attention in the era of multimodal large language models (MLLMs). In pursuit of ever more detailed and accurate captions, recent work has increasingly turned to reinforcement learning (RL). However, existing captioning-RL methods and evaluation metrics often emphasize a narrow notion of caption quality, inducing trade-offs across core dimensions of captioning. For example, utility-oriented objectives can encourage noisy, hallucinated, or overlong captions that improve downstream question answering while harming fluency, whereas arena-style objectives can favor fluent but generic descriptions with limited usefulness. To address this, we propose a more balanced RL framework that jointly optimizes utility-aware correctness, reference coverage, and linguistic quality. In order to effectively optimize the resulting continuous multi-objective reward formulation, we apply GDPO-style reward-decoupled normalization to continuous-valued captioning rewards and show that it improves performance over vanilla GRPO. Additionally, we introduce length-conditional reward masking, yielding a more suitable length penalty for captioning. Across LLaVA-1.5-7B and Qwen2.5-VL 3B and 7B base models, our method consistently improves caption quality, with peak gains of +13.6 DCScore, +9.0 CaptionQA, and +29.0 CapArena across different models.
PaperID: 5047, Poster
Authors: Zheng Qin, Junhua Xi, Yan Huang, JiaHong Lai, Anwen Huang, Qiong Li, Guangda Zhang, Kai Xu
Abstract: We present Register Anything Model (RAM), a unified model for generalizable 3D point cloud registration that robustly estimates 6-DoF transformations across diverse domains spanning sensors and environments. Existing registration methods are typically tailored to a single domain, with carefully tuned architectures and hyperparameters, and thus degrade markedly when applied to novel domains. We bridge this gap by training a single model on diverse data from indoor, outdoor, and object domains. A central challenge is the severe geometric discrepancy across domains, i.e., differences in scene scale, sensor noise, point density, and structural patterns, which hinders processing with a unified point-based model. To address this, we introduce a geometry-aware normalization strategy that jointly voxelizes and normalizes point clouds into a standarized geometric space, adaptively based on their geometric properties to preserve salient local structures while standardizing resolution. Built on this normalization, RAM follows a hierarchical coarse-to-fine paradigm, and employs an efficient geometric transformer that models local geometric consistency instead of over all points. This design mitigates discretization noise introduced by voxelization, and improves computational efficiency. Extensive experiments on 12 benchmarks demonstrate state-of-the-art performance in terms of inlier ratio and registration recall. Notably, our method outperforms the counterparts trained on each separate domain, showing superior generalization capability. Code and models will be released upon publication.
PaperID: 5048, Poster
Authors:
Jiawei Chen, Ruoxi Xu, Boxi Cao, Zhuoqun Li, Ruotong Pan, Zhang yunfei, Zhirui Yang, Jingyan Chen, Tingting Gao, Han Li, Yaojie Lu, Xianpei Han, Le Sun, Xiangyu Wu, Hongyu LinAbstract: Large language models (LLMs) have shown great potential as general-purpose user simulators for interactive systems. Prior studies have demonstrated that training on high-quality data can substantially improve model reasoning and downstream task performance. However, the upper bound of user simulation is still constrained by the model's reasoning paradigm: standard chain-of-thought often degenerates into shallow pattern matching and fails to capture the latent mental process underlying real user behavior. In this paper, we propose OmniSimulator, a novel training framework designed to improve the predictive ability of small LLMs for realistic user simulation on short-video platform environments. Specifically, OmniSimulator models human internal deliberation through a structured reasoning schema with five dimensions, carefully selects data from real user-behavior scenarios, and introduces a teacher model to generate high-quality reasoning traces under the proposed schema, boosting small models to learn the latent logic of human decision-making during the CoT process, rather than relying on superficial behavioral correlations. Based on this process, we construct 9,881 high-quality training instances from three months of historical interaction trajectories of 200 real users, averaging about 50 distilled examples per user. Experimental results show that OmniSimulator improves performance over the baseline by 38.93%, surpasses strong frontier models such as Claude-Sonnet-4.5, and significantly reduces inference cost. These results suggest that learning structured human-like internal reasoning is a key step toward scalable and high-fidelity LLM-based user simulation.
Abstract: A key strength of diffusion models lies in their flexibility, since their outputs can be controlled at sampling time through guidance. However, beyond simple cases such as conditional sampling, the target distribution is often left implicit, defined only through a sampling rule or a heuristic energy function. To address this, we propose Jeffrey guidance, a principled framework that extends diffusion-model control to applications beyond what standard guidance can express. It leverages Jeffrey’s rule of conditioning to update marginal distributions towards a prescribed target, preserving the conditional structure and minimally perturbing the joint distribution. We first demonstrate Jeffrey guidance by targeting a prescribed embedding distribution. With Inception embeddings as the target, this leads to substantial reductions in FID on both CIFAR-10 and FFHQ. We further apply Jeffrey guidance to fairness on CelebA-HQ, updating an unconditional diffusion model to enforce independence between attributes.
Authors: Alberto D Cencillo, Leonardo Concepción, Isaac Triguero, Julián Luengo
Abstract: Anomaly detection in multivariate time series is a critical task across a wide range of real-world applications, where abnormal behaviour is rare, labels are unavailable, and the cost of a miss is high. The central challenge is learning a characterisation of normality precise enough to flag deviations. Representation self-supervised learning, typically through contrastive approaches, addresses this by embedding temporal patches into a latent space where normality occupies a well-defined region, with anomalies detected by geometric deviation. However, contrastive approaches shape this space indirectly through pair-sampling heuristics, providing no explicit control over the geometric structure that distance-based scoring requires. This means how tightly normal representations are grouped, and whether distances are directionally meaningful. We present VACE (Velocity-Aligned Channel Embeddings), a self-supervised anomaly detection method that represents normality as a compact, directionally coherent region in the embedding space. To this end, VACE trains a channel-aware encoder through a velocity-consistency objective, with no negatives and no synthetic anomalies, so that normal trajectories are locally smooth and aligned. At test time, a Mahalanobis positional score and a velocity-bank directional score are combined multiplicatively, flagging points that are simultaneously off-distribution and dynamically atypical. Despite its simplicity, VACE achieves state-of-the-art performance on TSB-AD-M under rigorous evaluation, significantly outperforming more complex methods trained on substantially larger budgets.
PaperID: 5051, Poster
Abstract: We consider learning-augmented mechanism design for facility location in the standard setting of locating a single facility on the line. Leveraging the given (imperfect) prediction on the optimal facility location, our goal is to design strategyproof (SP) mechanisms that truthfully elicit agent location preferences and determine facility locations that approximately minimize the L_p-norm social cost for p \in (1, +\infty), achieving high \emphconsistency (i.e., approximation ratios when the prediction is correct) and \emphrobustness (i.e., approximation ratios when the prediction is incorrect). For deterministic SP mechanisms, we propose a family of generalized median mechanisms parameterized by a fraction \(t \in (0,1)\) of phantom points placed at the predicted location \(\pi\). We show that this mechanism achieves \big(1+\big(\frac1-t1+t\big)^\frac1p-1\big)^\fracp-1p-consistency and \big(1+\big(\frac1+t1-t\big)^\frac1p-1\big)^\fracp-1p-robustness and that no deterministic SP mechanism with the same consistency can obtain a better robustness. For randomized SP mechanisms, we analyze the upper bounds for several special cases, including (1) two-agent and (2) the squared cost settings, and provide a lower bound on consistency-robustness tradeoffs.
Abstract: World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse volumes of data, to instill a rich prior into downstream applications. Existing methods for capturing 3D world models are typically scene-specific dynamics models, or require robot action labels - which excludes web video data from the training pool. We study 3D point track completion as a pre-trianing objective for learning transferable 3D dynamics without robot data. Given a single RGB-D observation and sparse partial 3D trajectories (tracks), we predict future 3D tracks of all observed points. We show this is a objective provides a rich dynamics prior, without requiring robot action labels. We contribute a diverse dataset of 2.9 million synthetic frames spanning deformable, articulated, and rigid objects, and use it to train PointZero. We show that a flexible and expressive transformer, PointZero, outperforms prior methods on the same data. We demonstrate the utility of our pre-training objective by fine-tuning PointZero for two downstream applications: (1) action-conditioned dynamics and (2) robot manipulation. When fine-tuned to condition on end-effector pose, PointZero outperforms the baselines on a recent action-conditioned 3D dynamics benchmark. When fine-tuned to predict robot actions, PointZero outperforms or matches the baselines on 6/7 simulated and real-world robot manipulation tasks. We furthermore evaluate training PointZero from scratch to isolate the benefits of our proposed architecture from those of our proposed pre-training objective and dataset. We release the dataset, model checkpoints, and full training recipe.
PaperID: 5053, Poster
Authors: Quynh Trinh, Trung Nguyen, Nam Nguyen, Tuan Dam
Abstract: Monte-Carlo Tree Search (MCTS) typically stores only scalar mean estimates, which makes it poorly suited for tail-risk objectives such as conditional value-at-risk (CVaR). We introduce Categorical Thompson Sampling with Optimism (CATSO) and Particle Thompson Sampling with Optimism (PATSO), two distributional MCTS algorithms that maintain finite-support empirical backup laws at Q-edges. CATSO represents Q-edge laws with categorical atoms and Dirichlet Thompson sampling, while PATSO represents them with adaptive particles. Both algorithms select actions using the CVaR of a Thompson-sampled Q-edge law plus a polynomial optimism bonus, and propagate scalar continuation values through visit-weighted averages of child Q-edge CVaR scores. For the finite-depth nested CVaR planning target, we prove an O(n^-1/2) bound on the root value error after n rollouts, which upper-bounds the risk-sensitive simple regret. CATSO adds a fixed-grid discretization bias, while capped PATSO adds a tunable Wasserstein compression bias. Experiments on stochastic planning benchmarks show that distributional Q-edges improve lower-tail return and reduce catastrophic outcomes when mean-optimal and risk-sensitive behavior differ.
PaperID: 5054, Poster
Abstract: Error Feedback (EF) has become a core mechanism for aggressive gradient compression in communication-efficient distributed training, and is now used in geo-distributed frameworks such as DiLoCo. While existing theory mainly explains its optimization convergence, the generalization behavior of EF remains poorly understood, especially under nonlinear biased compression and distributed local updates. We present the first generalization analysis of EF for both single-node EF-SGD and DiLoCo-EF through an optimization-driven on-average model stability framework, without imposing bounded-gradient assumptions. Technically, we develop a buffer-aware Lyapunov argument that tracks the coupled dynamics of the model and error buffer, while co-coercivity absorbs same-sample gradient-difference energy into the optimization trajectory. For convex losses, our bounds show that compression-induced stability terms are transient, decaying as \mathcalO(\sqrtb_\alpha\,T^-3/4+b_\alpha T^-1), and therefore do not create a non-vanishing generalization floor beyond the dominant \mathcalO(T^-1/2) optimization term. In the DiLoCo setting, worker averaging further suppresses compression perturbations, and a local stepsize \eta_l=\Theta(1/(H\sqrtT)) controls the drift from H local steps, yielding the leading \mathcalO(1/\sqrtMHT) excess-risk scaling. Experiments on convex benchmarks corroborate the theory, showing that Top-k, Random-k, and quantization with matched contraction factor \alpha exhibit nearly indistinguishable generalization trajectories.
Abstract: Supervised Fine-Tuning (SFT) is widely used for task-specific adaptation, yet recent work shows it systematically undermines reasoning generalization. We argue the root cause is not memorization itself, but its target: vanilla SFT drives models to exploit and memorize spurious surface correlations in problem-solution pairs, leaving them brittle to superficial input variations. To address this, we propose Theorem-SFT, which reorients supervision toward explicit theorem application by teaching models how rules are invoked rather than what answers look like. Theorem-SFT yields consistent gains across benchmarks and model families: +8.8% on MATH (LLaMA3.2-3B-Instruct) and +20.27% on GeoQA (Qwen2.5-VL-7B-Instruct) without modality-specific re-training. Fine-tuning MLP layers alone matches full-layers performance, implicating feed-forward components as the primary locus of reasoning rules. Our findings reframe the debate: Generalization failures stem not from memorization as a mechanism, but from memorizing the wrong inductive targets.
PaperID: 5056, Poster
Authors: Chau T Hoang, Hieu Ta, Zhenzhen Liu, Dung Le
Abstract: Vision-language models such as CLIP excel at image-text alignment but struggle with long, detailed descriptions due to training on short captions. Recent methods address this limitation using region proposals to align visual regions with sentences, but at substantial deployment cost. We present MulCLIP, an end-to-end multi-level alignment framework that directly exploits natural long-text structure without region proposals. Instead of simply stacking objectives, MulCLIP aligns different textual granularities through compatible training signals: global contrastive alignment for long captions and short summaries, Word–Patch Reconstruction over locally calibrated features for within-sample word–patch semantics, and Subcaption–Aggregated Patch alignment for context-rich subcaption grounding. Experiments on benchmarks spanning varying caption lengths show consistent gains, and ablations confirm that the proposed objectives are complementary and jointly improve the overall framework.
Abstract: When deploying large language models (LLMs) to safety-critical applications, uncertainty quantification (UQ) is of utmost importance to self-assess the reliability of the LLM-based decisions. However, such decisions typically suffer from overconfidence, particularly after parameter-efficient fine-tuning (PEFT) for downstream domain-specific tasks with limited data. Existing methods to alleviate this issue either rely on a Laplace approximation based post-hoc framework, which may yield suboptimal calibration depending on the training trajectory, or variational Bayesian training that requires multiple complete forward passes through the entire LLM backbone at inference time for Monte Carlo estimation, posing scalability challenges for deployment. To address these limitations, we build on the Bayesian last layer (BLL) model, where the LLM-based feature extractor is followed by random last-layer parameters for uncertainty reasoning. Since existing low-rank adapters (LoRA) for PEFT have limited expressiveness due to rank collapse, we address this with Polar-decomposed Low-rank Adapter Representation (PoLAR), an orthogonalized parameterization paired with Riemannian optimization to enable more stable and expressive adaptation. Building on this PoLAR-BLL model, we leverage the variational (V) inference framework to put forth a scalable Bayesian fine-tuning approach which jointly seeks the PoLAR parameters and approximate posterior of the last-layer parameters via alternating optimization. The resulting PoLAR-VBLL is a flexible framework that nicely integrates architecture-enhanced optimization with scalable Bayesian inference to endow LLMs with well-calibrated UQ. Our empirical results verify the effectiveness of PoLAR-VBLL in terms of generalization and uncertainty estimation on both in-distribution and out-of-distribution data for various common-sense reasoning tasks.
PaperID: 5058, Poster
Authors:
Jiayang Wu, Chenchen Qin, Xu YANG, Yu Zhao, Bing He, Jiale Zhou, Zhenchao Tang, Minghao Yang, Shou Z Chen, Yefeng Zheng, Jianhua YaoAbstract: Genomic events span multiple spatial scales, from single-nucleotide substitutions to large structural variants, yet existing foundation models process only reference DNA sequences, leaving methylation, short variants, and structural-variant boundaries as downstream labels or auxiliary inputs. We propose the Unified Genomic Model (UGM), an encoder-only genomic foundation model that learns multi-scale genomic events as native tokens within a unified 43-token vocabulary. UGM introduces three algorithmic components: (i) a unified event tokenizer that maps reference bases, SNPs, CpG methylation, short INDELs, and structural-variant boundaries into a single event-state space; (ii) event-balanced pretraining, which promotes adequate gradient coverage for rare but biologically important events; and (iii) Site-Level Joint Prediction (SJP), a masked-language-modeling objective that recovers the complete genomic state at a masked position as a single token rather than predicting base, methylation, and variant labels independently. We evaluate UGM against specialized and generalist baselines across functional variant classification, regulatory benchmarks (NT and GUE), structural-variant detection, and cross-event attention analysis. UGM achieves competitive performance, particularly on tasks where explicit event-state information is central. It obtains the highest score on all seven functional variant tasks, reaches competitive splicing prediction on NT, and notably outperforms DNA-only baselines on SV breakpoint detection. These results suggest that native multi-scale event pretraining is a promising direction for genomic representation learning.
PaperID: 5059, Poster
Authors: Jinjie Xie, Baihua Li, Qinggang Meng
Abstract: Compact medical multimodal models are the most realistic path to clinical deployment, yet they often produce fluent answers that are not actually grounded in the image. We show that this failure has a clear internal signature. In five compact medical backbones and across six medical VQA benchmarks, visual tokens lose their spatial diversity shortly after entering the language backbone, their residual updates are an order of magnitude smaller than text residuals, and attention on them freezes on content free regions of the image. Zeroing the visual residual stream barely changes the output, while zeroing the text residual stream destroys it. We give a short analytical account that links the collapse of visual similarity to the residual dominance ratio, and we propose VGrip, a set of three lightweight losses that act on the three failure points without modifying the backbone architecture or the inference procedure. We also introduce the Visual Dependency Score, a simple and benchmark agnostic measure of how much a model actually uses the image. On Qwen2.5-VL-7B, VGrip lifts accuracy by 4.7 points on average across the six benchmarks and more than doubles the Visual Dependency Score over a strong LoRA baseline, with gains that are stable from two to eight billion parameters.
PaperID: 5060, Poster
Authors:
Tianyi Xu, Shrinaath Narasimhan, Evan Gorstein, Santiago Perea, Yunyi Shen, Claudia Solis-LemusAbstract: What did an ancestral bird species sound like? Existing ancestral state reconstruction methods can infer low-dimensional traits such as morphological characters at internal nodes of a phylogenetic tree, but no one has tried to produce rich perceptual signals such as audio. Some of the challenges include inferred representations that are either too low-dimensional to decode or lie in non-generative feature spaces, so no method to date can produce ancestral audio. We introduce the first framework that generates plausible ancestral vocalizations. Our pipeline encodes bird recordings into a VAE latent space, learns a low-dimensional trait projection aligned with phylogenetic distances, performs ancestral inference in this trait space, and recovers decodable latents through an anchored inverse lift before emitting novel waveforms for each ancestral node. Because the entire pipeline stays within a decodable latent space, every internal node receives a genuinely new audio output representing plausible intermediate ancestral sounds unavailable to retrieval-based alternatives. Experiments on two phylogenetically distant bird clades, 21-species Tyrannidae and 19-species Paridae, show that our method is the only approach that simultaneously achieves genuine generation, phylogenetic consistency, and naturalistic audio quality across both datasets.
PaperID: 5061, Poster
Abstract: Multinomial logit models (MNLs) are ubiquitous in machine learning, particularly in recommendation systems. While these models are typically evaluated using global metrics over a full catalog of items (such as \ell_2-error), practical applications often involve making predictions on small, curated subsets or ``slates'' of items. A model that minimizes global error may nevertheless yield very poor predictions on specific sub-universes. Yet, evaluating performance across all possible subsets remains computationally challenging. In this paper, we present a set of new efficient algorithms to evaluate the small-slate performance of MNL models. Using these tools, we conduct an empirical study comparing standard non-adaptive and adaptive learning algorithms. We demonstrate that while methods that compute the Maximum Likelihood Estimate (MLE) suffer from poor worst-case performance on small slates, specialized adaptive algorithms that dynamically sample difficult subsets can mitigate this issue. Conversely, we find that non-adaptive algorithms often achieve competitive average-case performance; we provide a theoretical model explaining this phenomenon. We conclude with recommendations for practitioners regarding when to utilize adaptive sampling based on the specific robustness requirements of their application.
PaperID: 5062, Poster
Abstract: Neural Ratio Estimation (NRE) is a popular method for performing simulation-based posterior estimation given a known prior and samples from a joint distribution. For such a task, it can be crucial for the posterior estimation task to be \emphstatistically efficient, i.e. to require as few simulations as possible to obtain an accurate posterior estimate. However, currently, little is known about its statistical efficiency. In this work, we bridge this gap by performing an asymptotic statistical analysis of NRE. We show that even when the target distribution is a simple product of Gaussians, the accuracy of NRE can degrade exponentially badly with the dimension of the data distribution. Our results show that NRE can be significantly less efficient than other conditional density methods like Maximum Likelihood Estimation, which are known to enjoy favorable statistical properties.
PaperID: 5063, Poster
Authors: Chi Zhang, Jinge Li, Yifei Wang, Lei Wang, Tianyi Qian
Abstract: Biological neural systems leverage neuronal heterogeneity and diverse firing patterns to generate rich nonlinear dynamics, yet these mechanisms are largely abstracted away in conventional artificial neural networks (ANNs) and simplified spiking neural networks (SNNs). Detailed differential-equation neural models are biologically expressive, but spike discontinuities and complex state evolution pose challenges for gradient-based optimization in complex functional tasks. To address this challenge, we introduce EDEN (Emergent Dynamics in Evolutionary Neural-Networks), a neural dynamic model training framework inspired by natural evolutionary processes. Built upon a continuous-time network of heterogeneous Izhikevich neurons, EDEN integrates efficient parallel neural dynamics simulation with advanced evolutionary strategies, enabling the joint optimization of intrinsic single-neuron-level parameters and network-level synaptic connections without backpropagation. Comprehensive evaluations across multiple MuJoCo continuous motor control tasks indicate that EDEN achieves robust control performance comparable to mainstream deep reinforcement learning baselines. Furthermore, neurodynamic analysis reveals that the trained EDEN models exhibit strong biological plausibility, characterized by functional differentiation across individual neurons and emergent low-dimensional latent manifolds governing population dynamics. EDEN demonstrates that the combination of gradient-free evolutionary strategies and biological neurodynamic networks provides an alternative yet efficient paradigm alongside mainstream reinforcement learning, thereby offering a highly promising pathway towards truly biologically-inspired intelligence.
PaperID: 5064, Poster
Abstract: Foundation models for electroencephalography (EEG) have shown promise in learning transferable representations, but existing approaches treat EEG as a generic time series, ignoring the Riemannian geometry of spatial covariance structure that is fundamental to neural signal analysis. We propose MENDR (Manifold-Embedded Neural Data Representations), the first EEG foundation model that builds this geometry directly into its architecture (via the matrix logarithm and Log-Euclidean tangent space) rather than learning it from data. MENDR embeds windowed covariance matrices into the Log-Euclidean tangent space via differentiable matrix logarithm, where standard transformer operations become geometrically equivalent to Riemannian operations. A channel-agnostic spatial projection based on Perceiver-style cross-attention enables seamless transfer across electrode configurations ranging from 6 to 64 channels. Pretrained on the Temple University Hospital EEG Corpus (4,000+ hours, 14,987 subjects) using a masked autoencoder objective in the tangent space, MENDR achieves state-of-the-art results on 3 of 6 downstream benchmarks (TUAB abnormality, TUEV event, CHB-MIT seizure) with only 1.2M parameters, up to 42× fewer than competing foundation models.
PaperID: 5065, Poster
Authors: Mihai Bogdan Deaconu, Laura Diosan
Abstract: Topology reasoning jointly detects 3D lanes and traffic elements from multi-view images and infers their structural connectivity. Current methods model lanes as discrete polylines, lacking smoothness, analytical tangent directions, and global spatial support for attention, while providing sparse topology supervision. We propose TopoCurve, a geometry-driven architecture for 3D topology reasoning grounded in a structured parametric lane representation. Lanes are modeled as endpoint-fixed cubic Bézier curves, enabling continuous geometry with exact endpoints and analytically defined directionality. We exploit this shared curve geometry across the entire pipeline. Endpoint distance and tangent alignment are encoded with multi-scale Fourier features and injected into the topology head. Sampled curve points serve as geometry-aligned references for deformable cross-attention spanning the full lane. Parallel curve-anchored attention branches provide diverse predictions for one-to-many topology supervision. These components form a tightly coupled cascade where representation enables geometric reasoning, guides feature aggregation, and supports denser supervision. TopoCurve achieves 50.6 OLS on the OpenLane-V2 benchmark without any post-processing, establishing a new state-of-the-art among end-to-end camera-only methods, and outperforms all existing approaches on endpoint detection (56.8 vs. 52.6 on DET_p).
PaperID: 5066, Poster
Authors: Zemin Chao, Cen Jianhe, Qianhui Xu, Guangzhi Ge, Zhixin Qi, Hongzhi Wang
Abstract: The quadratic computational complexity of attention during prefilling remains a primary bottleneck for the efficiency of long-context LLMS. While sparse attention methods mitigate this cost, existing dynamic approaches rely on computationally expensive O(N^2) scoring, which undermines their practical efficiency gains. In this paper, we observe that long-context attention matrices frequently exhibit prominent diagonal stripe patterns, indicating that query-key pairs separated by specific relative distances yield consistently high correlation. Motivated by this, we propose FFTSparse, a hybrid block-sparse attention framework. By formulating the identification of these dominant attention lags as cross-correlation in the frequency domain via the Fast Fourier Transform (FFT), FFTSparse reduces the cost to identify the sparsity pattern of diagonal stripes to O(N \log N). Additionally, we introduce a lightweight identification module that categorizes the patterns of attention heads and then applies tailored masking strategies accordingly. Extensive evaluations demonstrate that FFTSparse incurs a negligible mask construction overhead of merely 0.02% relative to full attention. Furthermore, it maintains competitive downstream performance on long-context benchmarks while achieving up to a 7.4x prefilling speedup at a 128K sequence length, offering a highly efficient solution for long-context inference of LLMS.
Authors: Thomas Quadt, Nick Leenders, Roy Lindelauf, Herman Monsuur, Mark Voskuijl, Joost van Oijen, Boris Cule
Abstract: Preferential Bayesian Optimization (PBO) aims to find a decision-maker’s most preferred solution in as few pairwise comparisons as possible. Existing approaches rely on Gaussian Process (GP) surrogates, which provide strong performance but limited interpretability. This limits real-world usability in high-stakes domains, such as healthcare, where interpretability and trust are essential. We propose DT-PBO, a novel tree-based surrogate model for PBO that is inherently interpretable while capturing preference uncertainty. Specifically, we introduce a novel splitting heuristic that constructs interpretable shallow decision trees directly from pairwise comparison data, and use Laplace approximation to obtain probabilistic estimates within each leaf. This enables efficient preference modeling without sacrificing interpretability. Across eight benchmark functions, our method achieves competitive convergence to GP-based PBO, particularly on functions with rugged optimization landscapes. Additional experiments show robustness against noise and a fast computational running time. Experiments on four real-world datasets further demonstrate that our model provides interpretable insights into decision-maker preferences that remain opaque under GP-based approaches.
Abstract: Real-world image degradation is often unknown, spatially non-uniform, and compositional, requiring all-in-one restoration models to adapt a single set of weights to diverse local corruption patterns without test-time degradation labels. Existing methods typically modulate a shared backbone with global prompts or degradation descriptors, or route features through predefined expert pools. However, compact global conditioning can bottleneck localized degradation evidence, while static expert routing may produce homogeneous updates or rely on unstable sparse assignments. We propose Continuous Expert Assembly (CEA), a token-wise dynamic parameterization framework for all-in-one image restoration. CEA employs a lightweight Cross-Attention Hyper-Adapter to probe intermediate spatial features and synthesize instance-conditioned low-rank routing bases and residual directions. Each spatial token then assembles its own residual update via dense signed dot-product affinities over the generated rank-wise components, avoiding external prompts, static expert banks, and discrete Top-K selection. The resulting assembly rule also admits a linear-attention perspective, making its dense token-wise routing behavior transparent. Experiments on AIO-3, AIO-5, and CDD-11 show that CEA improves average restoration quality over strong prompt-, descriptor-, and expert-based baselines, with the clearest gains on spatially varying and compositional degradations, while maintaining favorable parameter, FLOP, and runtime efficiency.
Authors:
Yiyang Su, Jie Zhu, Feng Liu, Anil Jain, Xiaoming LiuAbstract: While foundation models have significantly advanced human recognition across diverse modalities, they predominantly rely on static, geometric feature extraction. This approach fundamentally diverges from human perception. Consequently, current models often suffer from ``semantic blindness,'' overfitting to transient noise while failing to leverage invariant soft biometrics, and struggle to capture temporal motion signatures. To bridge this gap, we propose SapiensID 2.0, a human recognition framework enriched with both semantic and temporal awareness. To overcome the lack of soft-biometric annotations, we transfer zero-shot semantic knowledge from Multimodal Large Language Models (MLLMs) into a discriminative embedding space. We resolve the dimensional mismatch between these spaces using Invariant Trait Alignment (ITA) to distill core persistent traits, and Transient Noise Disentanglement (TND) to decouple artifacts like clothing. Furthermore, we design a Kinematic Semantic Attention Head (K-SAH) that extends spatial attention across temporal windows. By tracking semantic patches over time, K-SAH captures rich kinematic signatures without requiring large-scale video datasets. Extensive experiments demonstrate that SapiensID 2.0 achieves state-of-the-art performance across image- and video-based person re-identification and gait recognition, while maintaining robust face recognition capabilities.
Abstract: While Value Iteration (VI) is one of the most fundamental algorithms in Reinforcement Learning, its theoretical convergence guarantees still exhibit a persistent mismatch with empirical behavior. In the discounted-reward case, classical theory guarantees geometric convergence with rate \gamma, while in the average-reward case recent work suggests that only sublinear convergence can be expected. In practice, however, VI is often observed to converge significantly faster. In this work, we show through a unified geometry-based analysis that, under an assumption of a unique and unichain optimal policy, (i) convergence is geometric in both the discounted- and average-reward settings and (ii) the convergence rate is faster than previous analyses suggest.
Authors:
Yang Zhou, Xiaofeng Wang, Hao Shao, Letian Wang, Guosheng Zhao, Jiangnan Shao, Jiagang Zhu, Tingdong Yu, Zheng Zhu, Guan Huang, Steven WaslanderAbstract: Recently, world-action models (WAM) have emerged to bridge vision-language-action (VLA) models and world models, unifying their reasoning and instruction-following capabilities and spatio-temporal world modeling. However, existing WAM approaches often focus on modeling 2D appearance or latent representations, with limited geometric grounding, an essential element for embodied systems operating in the physical world. We present DriveDreamer-Policy, a unified driving world-action model that integrates depth generation, future video generation, and motion planning within a single modular architecture. The model employs a large language model to process language instructions, multi-view images, and actions, followed by three lightweight generators that produce depth, future video, and actions. By learning a geometry-aware world representation and using it to guide both future prediction and planning within a unified framework, the proposed model produces more coherent imagined futures and more informed driving actions, while maintaining modularity and controllable cost. Experiments on the Navsim v1 and v2 benchmarks demonstrate that DriveDreamer-Policy achieves strong performance on both planning and world generation tasks. In particular, our model reaches 89.2 PDMS on Navsim v1 and 88.7 EPDMS on Navsim v2, outperforming existing world-model-based approaches while producing higher-quality future video and depth predictions. Ablation studies further show that depth learning provides complementary benefits to video imagination and improves planning robustness. Our code is available to facilitate further research.
PaperID: 5072, Poster
Abstract: Cryo-electron microscopy (cryo-EM) reconstructs 3D macromolecular structures from noisy 2D particle images by jointly estimating the density and projection pose of each particle. Amortized inference can greatly accelerate this process, but remains vulnerable to pose collapse because known image-space symmetries are often learned only implicitly through reconstruction losses. We ask whether this instability can be mitigated by enforcing cryo-EM transformation laws directly in latent projection-pose space. We propose CryoGeo, a geometry-constrained framework for amortized ab initio cryo-EM reconstruction. CryoGeo enforces latent pose equivariance by requiring in-plane transformations to induce valid group actions on the predicted projection pose: rotations act as right multiplications on the predicted SO(3) orientation, while shifts transform frame-consistently. At the core of CryoGeo is an output-level group-action objective that forces predicted cryo-EM projection poses to obey physical transformation laws induced by in-plane image rotations and shifts, rather than relying on equivariant feature extraction alone. Across synthetic and real cryo-EM datasets, CryoGeo improves pose consistency, mitigates collapse-prone behavior, and yields more reliable amortized ab initio reconstructions than unconstrained variants and neural baselines. These results show that latent pose equivariance turns in-plane symmetry from an implicitly learned nuisance factor into a self-supervised geometric constraint for stable, efficient amortized structural reconstruction. Code will be released.
PaperID: 5073, Poster
Abstract: Deep models for medical time series can achieve strong benchmark accuracy yet remain brittle under acquisition artifacts, physiological variability, and low-resource training conditions. We present ClinStab, a Medformer-based framework for EEG and ECG classification that studies whether channel-aware dual-stream interaction and perturbation-aware optimization can improve the clean/robustness tradeoff. The proposed Salience-Context Dual Modulation (SCDM) module uses energy as a routing prior: high-energy channel-aligned components receive attentive modeling, while residual components remain trainable through a lightweight context pathway. ClinStab then combines clean supervision, perturbed supervision, prediction consistency, and teacher-guided stabilization under adversarial perturbations. Across four public datasets and 11 baselines, ClinStab improves accuracy over Medformer on all datasets and improves selected robustness metrics under the perturbation settings studied here, while exposing tradeoffs on PTB ranking metrics and soft-routing alternatives. These results support ClinStab as a practical same-backbone stability intervention, not as evidence of universal clinical robustness.
PaperID: 5074, Poster
Abstract: Transfer learning in reinforcement learning (RL) has shown strong empirical success. In this work, we take a more principled perspective by studying when and how transferring knowledge between MDPs (a source and a target) can be provably beneficial. Specifically, we consider the case where there exists an undo map such that applying this map to the target’s state space recovers the source exactly. We propose an algorithm that learns this map by matching state feature statistics gathered from both MDPs, and then uses it to transfer the source policy. We theoretically justify the algorithm by analyzing the setting when the undo map is linear and the source is linearly-Q^\star realizable, where our approach has strictly better sample complexity than tabula rasa RL in the target MDP. Empirically, we demonstrate that these benefits extend beyond this regime: on challenging continuous control tasks and Atari games, our method achieves significantly better sample efficiency. Overall, our results highlight how shared structure between tasks can be leveraged for efficient transfer of policies across environments.
Abstract: Time series imputation is a crucial area for reliable time series analysis, yet it remains challenging due to the complex temporal dynamics and noise of real-world data. Existing approaches, however, exhibit two limitations: missing and observed values are embedded within the same representation space without explicit structural separation, and continuous diffusion-based methods are trained to predict added noise rather than the original signal. To address these, we propose the Masked Diffusion Time-series Imputation Model (MDTIM), which leverages the training paradigm of masked diffusion model for imputation tasks. The \texttt[MASK] token is structurally orthogonal to valid observations, and the model directly predicts the original values, naturally aligning both the representation and the learning objective with the imputation task. To bridge the gap between discrete masked diffusion and the continuous, ordinal nature of time series, we further introduce Stochastic Discretization, which maps continuous values to ordinal-aware tokens while preserving continuous dynamics. Our experiments on diverse benchmarks confirm that MDTIM achieves superior robustness and scalability, consistently outperforming state-of-the-art deterministic and generative baselines across various missing scenarios.
Abstract: Large language models remain vulnerable to jailbreak attacks, yet we still lack a systematic understanding of how jailbreak success scales with attacker effort across methods, model families, and harm types. We initiate a scaling-law framework for jailbreaks by treating each attack as a compute-bounded optimization procedure and measuring progress on a shared FLOPs axis. Our systematic evaluation spans four representative jailbreak paradigms, covering optimization-based attacks, self-refinement prompting, sampling-based selection, and genetic optimization, across multiple model families and scales on a diverse set of harmful goals. We investigate scaling laws that relate attacker budget to attack success score by fitting a simple saturating exponential function to FLOPs--success trajectories, and we derive comparable efficiency summaries from the fitted curves. Empirically, prompting-based paradigms tend to be the most compute-efficient compared to optimization-based methods. To explain this gap, we cast prompt-based updates into an optimization view and show via a same-state comparison that prompt-based attacks more effectively optimize in prompt space. We also show that attacks occupy distinct success--stealthiness operating points with prompting-based methods occupying the high-success, high-stealth region. Finally, we find that vulnerability is strongly goal-dependent: harms involving misinformation are typically easier to elicit than other non-misinformation harms.
PaperID: 5077, Poster
Abstract: Synthesizing programs from execution traces is a fundamental challenge in algorithmic interpretability and the reverse engineering of complex systems. We bring a rough-path-theoretic perspective to program induction, treating execution traces as high-dimensional rough paths and encoding them with the path signature transform, a tool from stochastic analysis whose expected signature provably characterizes the trace distribution a program induces. Our signature-based trace encoder outperforms LSTM, Transformer, and Fourier baselines and remains strong with substantially less training data on two classical DSL domains. We further observe that the trace-to-program pipeline is naturally \emphcontractive: it tends to produce shorter valid programs than the ground truth, which we turn into a bootstrapping procedure that surfaces redundancy in the released ground-truth supervision without any external oracle. Finally, signature-based models induce a markedly more interpretable latent geometry than classical sequence models.
Abstract: Multimodal large language models (MLLMs) often fail to transfer safety capabilities learned in the text modality to semantically equivalent non-text inputs, revealing a persistent multimodal safety gap. We study this gap from a representation-geometric perspective by analyzing a text-aligned refusal direction and a modality-induced drift direction. We show that multimodal inputs compress the usable separation along the refusal direction, making it no longer reliable for identifying and refusing harmful inputs. We refer to this failure mode as Safety Geometry Collapse. We quantify it through conditional refusal separability and show that stronger modality-induced drift is consistently associated with weaker refusal separability and higher attack success rates. We then validate the causal role of modality-induced drift through a fixed-strength activation intervention: counteracting the estimated drift restores refusal separability and improves multimodal safety. After drift correction, we further observe self-rectification, where the model recovers its ability to recognize and refuse harmful multimodal inputs during forward dynamics. This effect also provides an internal signal of the model’s perceived harmfulness of each input. Motivated by this signal, we propose ReGap, a training-free inference-time method that adaptively corrects modality drift using self-rectification. Experiments across multiple multimodal safety benchmarks and utility benchmarks demonstrate the effectiveness of ReGap, which significantly improves the safety of MLLMs without compromising general capabilities. Our findings highlight representation-level modality alignment as a crucial direction for real-time safety improvement and for building safer, more reliable MLLMs.
PaperID: 5079, Poster
Abstract: Neurobiological circuits may be energy efficient in part because there is a division of labor: different subsystems compute or represent different things. In light of this, it is interesting that a major difference between subsystems is the time scale on which they typically vary. Here, we argue that temporal smoothness constraints---the idea that it may be energetically costly for systems to vary much more quickly or slowly than their typical time scale---imply a certain division of labor with respect to temporal stimulus features. In particular, slow subsystems ought to represent slowly varying stimulus features, and fast subsystems ought to represent quickly varying features. We formulate a novel efficient coding model that we use to investigate this claim, and exactly solve for the optimal division of labor. In addition to finding that slower subsystems ought to represent more slowly varying temporal features, we find surprising structure of mathematical interest: optimal codes have subsystems which represent orthogonal temporal features, and which correspond to the eigenfunctions of a Sturm-Liouville problem.
PaperID: 5080, Poster
Abstract: Large-scale model evaluation is increasingly delegated to LLM judges. This makes evaluation cheaper, but it creates a new statistical problem: if several judges share a systematic bias, averaging more judge scores can make a ranking look more confident without making it more correct. We study fixed-confidence top-k identification when judge scores are cheap, correlated, and biased, while trusted anchor labels are costly. Under a low-dimensional factor-bias model, we show that top-k identifiability is controlled not by the total number of anchors, but by whether anchors span the feature directions separating arms across the top-k boundary. This yields an instance-dependent cost characterization with two components: effective information from correlated judge scores and a profiled anchor-leverage term quantifying the cost of removing shared bias. We propose Profile Track-and-Stop, a sequential allocation algorithm that tracks the resulting cost-optimal allocation with a conservative pilot anchoring phase and asymptotically matches the lower-bound constant. Experiments on a 10-judge Arena-Hard-v2.0 evaluation instance with 23,545 real API judge scores, together with synthetic and semi-synthetic ablations, are consistent with the predicted failure mode: anchor-free methods can plateau under cross-family judge bias, whereas boundary-aware anchoring improves top-k recovery at lower cost.
PaperID: 5081, Poster
Abstract: Recent advances in multimodal foundation models have demonstrated strong performance across diverse clinical tasks. However, existing approaches predominantly operate at the individual sample level and fail to model the interplay between population-level cohort patterns and patient-specific variation that underlie real-world clinical reasoning. In practice, clinicians routinely reason through patient cohorts sharing similar demographic, physiological, or pathological characteristics, while accounting for patient-specific deviations from the group prototype. Motivated by this gap, we propose CODA, a cohort- and drift-aware foundation model for multimodal clinical reasoning. CODA explicitly integrates cohort structure into the modeling pipeline through two complementary mechanisms: cohort-aware population stratification and drift-aware individual adaptation. In particular, we leverage reinforcement learning to discover clinically meaningful patient cohorts and further model each patient’s deviation from cohort prototypes as a structured drift signal encoding fine-grained individual heterogeneity. These cohort- and drift-aware representations are incorporated into the foundation model to enable reasoning that is simultaneously informed by population-level patterns and sensitive to individual-level variation. Extensive experiments on real-world electronic health record (EHR) benchmarks demonstrate that CODA achieves consistent effectiveness and strong generalizability across diverse clinical tasks, including closed-QA, open-QA, and report generation.
PaperID: 5082, Poster
Abstract: Standard graph neural networks (GNNs) are built on message passing and neighborhood aggregation, following the Weisfeiler--Lehman refinement paradigm. Despite their wide adoption, these models tend to compress and mix neighbor information from the very first layer. This early mixing obscures the critical subgraphs that drive the prediction. Attention-based and subgraph-based methods partially address this by reweighting nodes or enriching local representations, but both families remain passive as they usually rely on soft weighting or predefined subgraph templates. We propose BeamWitness, a selective graph representation learning framework that treats subgraph identification as an active search problem. Starting from a root node, a policy trained with task rewards expands a connected subgraph one node at a time or terminates early, and beam search retains a small set of candidate subgraphs as witnesses for prediction. For node-level tasks, the target node serves as the root, while for graph-level tasks, BeamWitness launches searches from multiple root nodes and aggregates the selected witnesses into a graph-level representation. On controlled motif benchmarks with known ground-truth subgraphs, BeamWitness achieves near-perfect accuracy and selects witnesses that align with the true motifs. On real-world node and graph classification benchmarks, BeamWitness is competitive with standard GNNs and related methods while providing an explicit, interpretable subgraph-based account of each prediction.
PaperID: 5083, Poster
Abstract: Flow Matching (FM) has emerged as the dominant paradigm for state-of-the-art image and video generation, yet its reliance on full attention over lengthy sequences incurs significant latency. While existing caching methods reduce redundant computations, they yield suboptimal performance, fundamentally constrained by inflexible spatial granularity and empirical, model-specific heuristics. To bridge this gap, we ground the caching mechanism in the geometric properties of the FM velocity field. We theoretically establish that the curvature inherently dictates the coupling between trajectory fluctuation and computational requirements, thereby providing a rigorous metric to quantify caching staleness loss and automatically derive adaptive granularity and optimal update intervals. Translating this theoretical insight into practical acceleration, we present CuBic, a train-free caching framework. It bypasses expensive curvature computations via a zero-overhead proxy signal that dynamically guides spatial partitioning and budget-constrained update scheduling. At the system level, a sparse forward engine translates these optimal policies into efficient runtime, ensuring that theoretical guarantees directly translate into actual FLOP savings and linear speedup. Being model-agnostic and virtually hyperparameter-free, CuBic consistently outperforms state-of-the-art caching baselines in quality at customizable acceleration ratios (averaging 2.5×) across Wan2.1, Flux.1, and FireRed-Image-Edit.
PaperID: 5084, Poster
Abstract: Diverse applications on signed networks, ranging from leader-follower opinion dynamics to graph semi-supervised learning, fundamentally reduce to solving the discrete Dirichlet boundary value problem. The computational bottleneck of this problem lies in evaluating the entries of the label propagation operator Q = -L_U,U^-1L_U,L. While existing approximation methods based on absorbing random walks attempt to address this complexity, they often suffer from high variance and limited precision. To overcome these limitations, we propose MultiTreeQ, a novel and efficient framework designed to approximate the operator's entries with high accuracy. Specifically, MultiTreeQ is constructed based on the Signed Multi-Rooted Spanning Tree and leverages Rao-Blackwellized variance reduction to significantly improve estimation precision. Extensive experiments on real-world networks confirm that MultiTreeQ substantially outperforms existing baselines, scaling to massive graphs with up to 10 million nodes while delivering remarkable improvements in computational efficiency and precision. Our code is available at https://anonymous.4open.science/r/label-propagation-operator-2059/.
Authors:
Aidan Li, Uday Kiran Reddy Tadipatri, Mahan Fathi, Sarath Chandar, Ross GoroshinAbstract: Koopman autoencoders (KAEs) seek a higher-dimensional latent representation in which nonlinear dynamics evolve linearly. However, many interesting systems have multiple basins of attraction, and both theoretical and empirical work has shown these multibasin systems cannot generally admit a single finite-dimensional global Koopman embedding under standard assumptions. We posit that encoders with a sparsity-inducing objective encouraging few active latent coefficients will provide latent supports as an inspectable basin-modeling principle for Koopman autoencoders. We use these encoders producing sparse latents in training Sparse Koopman Autoencoders (SKAEs) without basin labels or other regime annotations, and treat the learned latent supports as model-produced regime variables after training. Across a range of procedurally generated multibasin systems and chaotic flows, we show that SKAEs have superior forecasting performance compared to dense-latent KAEs. We also perform a mechanistic study that shows latent supports produced by SKAEs are both essential for the quality of the representation and useful for identifying basins on held-out basin interior states, whereas dense-latent KAEs collapse to an uninformative single family. These results identify sparse latents and their corresponding supports as label-free, interpretable regime variables for Koopman learning in nonlinear systems with multiple local dynamical laws.
PaperID: 5086, Poster
Abstract: Speech tokenizers serve as the critical interface between continuous speech signals and discrete representations, and are widely used in speech systems, e.g., Automatic Speech Recognition (ASR) and Large Audio-Language Models (LALMs). In this paper, we find that speech tokenizers are fragile at the semantic level: slight perturbations can disrupt their semantic encoding process and consequently cause severe degradation in downstream ASR and LALM performance. To study this vulnerability, we propose T-SemAttack, a transferable semantic attack based on a semantic encoder ensemble. By jointly perturbing the representation spaces of multiple semantic encoders, T-SemAttack effectively disrupts the semantic content of speech while preserving perceptual quality, thereby inducing error tokens. We further analyze token fragility and layer-wise representation drift in relation to cross-model transferability and disruption of LALM attention. Our analysis reveals a progressive amplification chain in which small waveform perturbations are magnified through the tokenizer pipeline and lead to semantic collapse, e.g., the attack causes a 99.7% token change in the S3 tokenizer. Building on these findings, we introduce ROSETok, a robust speech tokenizer that combines noisy training with robust semantic distillation to improve reconstruction fidelity and downstream-task robustness. Extensive experiments on multiple datasets, 9 open-source and black-box ASR systems, and 4 LALMs demonstrate that T-SemAttack achieves strong transferable attack performance, while the proposed Robust Speech Tokenizer exhibits robustness under high-fidelity reconstruction and downstream tasks. Our code and demo are available at https://t-semattack.github.io.
Authors: Joshua Shay Kricheli, Alexander L. Reid, Soumajyoti Sarkar, Venkata Gandikota, Paulo Shakarian
Abstract: Neural scaling laws approximate a language model's loss as a power-law function of parameter count N and token count D. Following Chinchilla-style compute-optimal training, many studies fit scaling laws from runs performed under a fixed tokens-per-parameter (TPP) ratio k and set D = kN. We show that this collinear design, combined with the empirically common near-equality of the exponents governing N and D, induces an inherent ill-conditioning in the Gauss-Newton least-squares problem: the condition number of the design grows as the inverse square of the gap between the N and D-exponents. The scale coefficients become practically unidentifiable, with confidence intervals inflating by an order of magnitude or more, yielding a ``sloppy'' model whose extrapolations degrade sharply off the training ray. We prove this for four scaling-law formalisms and derive a closed-form TPP-diversity threshold that is necessary and sufficient for well-conditioned estimation. Empirically, non-collinear designs outperform collinear ones on held-out splits with a 97.3% win rate across four laws, five corpora, multiple floating point precision modes. We further show the degeneracy is rooted in Jacobian geometry and is not an artifact of the loss function: any smooth estimation objective whose curvature involves the Jacobian inherits the same ill-conditioning.
Authors: Pin-Han Huang, ShangTse Chen, Hsuan-Tien Lin
Abstract: Incorporating diffusion-generated synthetic data into adversarial training (AT) has been shown to substantially improve the training of robust image classifiers. In this work, we extend the role of diffusion models beyond merely generating synthetic data, examining whether their internal representations, which encode meaningful features of the data, can provide additional benefits for robust classifier training. Through systematic experiments, we show that diffusion models offer representations that are both diverse and partially robust, and that explicitly incorporating diffusion representations as an auxiliary learning signal during AT consistently improves robustness across settings. Furthermore, our representation analysis indicates that incorporating diffusion models into AT encourages more disentangled features, while diffusion representations and diffusion-generated synthetic data play complementary roles in shaping representations. Experiments on CIFAR-10, CIFAR-100, and ImageNet validate these findings, demonstrating the effectiveness of jointly leveraging diffusion representations and synthetic data within AT.
Abstract: Spectral optimizers such as Muon have recently shown strong empirical performance in large-scale language model training, but the source and extent of their advantage remain poorly understood. We study this question through the linear associative memory problem, a tractable model for factual recall in transformer-based models. In particular, we go beyond orthogonal embeddings and consider Gaussian inputs and outputs, which allows the number of stored associations to greatly exceed the embedding dimension. Our main result sharply characterizes the recovery rates of one step of Muon, SGD, and Newton's method on the logistic regression loss under a power law frequency distribution. We show that the storage capacity of Muon significantly exceeds that of SGD, and even matches Newton's method while only using first-order information. Moreover, Muon saturates at a larger critical batch size. We further analyze the multi-step dynamics under a thresholded gradient approximation and show that Muon achieves a substantially faster initial recovery rate than SGD, while both methods eventually converge to the information-theoretic limit at comparable speeds. Experiments on synthetic tasks validate the predicted scaling laws. Our analysis provides a quantitative understanding of the signal amplification of spectral preconditioners and lays the groundwork for establishing scaling laws across more practical language modeling tasks and optimizers.
PaperID: 5090, Poster
Abstract: Recovering executable CAD programs from 3D meshes is fundamentally challenging due to the long-horizon, compositional nature of CAD construction and the need for precise estimation of both discrete operations and continuous parameters. Existing learning-based methods primarily treat mesh-to-CAD reconstruction as one-shot sequence prediction and are largely limited to simple sketch-extrude pipelines, restricting operation diversity and preventing the use of intermediate geometric feedback, with no mechanism for refining generated programs. To address this limitation, we introduce StepCAD, a generative optimization approach that performs step-by-step, geometry-conditioned program synthesis followed by explicit refinement in program space. Given an input mesh, StepCAD first predicts construction sequences using a state-conditioned CAD policy conditioned on both target and intermediate geometry, then refines them through IoU-guided tree search over local program edits. To address the limited operation diversity in prior work, we further introduce ARCADE-1.5M, a large-scale dataset of 1.5M executable CAD programs with diverse operations, long-horizon sequences up to 150+ operations, and 12.5M intermediate state-action transitions. Experiments on multiple CAD reconstruction benchmarks show that StepCAD achieves state-of-the-art reconstruction accuracy, validity, and robustness, with up to 87.2% relative IoU improvement over the best prior method and increasingly larger gains as shape complexity increases.
PaperID: 5091, Poster
Abstract: Transformers have achieved remarkable success across a wide range of applications, and a growing body of work suggests that part of their strength comes from their ability to learn and execute algorithmic procedures. However, our understanding of how transformers learn such algorithms remains limited, especially in the presence of layer normalization (LN). In this work, we study principal component prediction as a concrete testbed for understanding the training dynamics of transformers with LN. We prove that a looped linear transformer with LN, trained by gradient descent, converges to a solution that implements the power method, with each self-attention layer performing one power iteration. Notably, the model is trained only for principal component prediction, rather than being explicitly supervised to implement the power method. Our finding thus reveals an ``algorithmic implicit bias'' of looped transformers with LN: principal-component prediction can in principle be achieved by many mechanisms, yet gradient descent selects one that realizes the power method. We further provide a concrete comparison between transformers with and without LN: even with layerwise guidance from power iterations, a transformer without LN cannot exactly learn the power method, whereas the corresponding transformer with LN can, leading to a provable performance gap in principal component prediction. Our results provide, to our knowledge, the first theoretical analysis of the training dynamics of looped and single-layer transformers with LN, and shed light on the role of LN in transformer models.
PaperID: 5092, Poster
Abstract: Agentic AI systems enable LLMs to solve non-trivial tasks through structured workflows, but automatically generating such workflows remains challenging due to the discrete and combinatorial search space. Existing methods often rely on offline search or training, limiting query-level adaptability and incurring substantial data and engineering overhead. We formulate workflow generation as a coupled topology--execution search problem, where the upper-level topology induces subtask-specific code-search domains and lower-level execution feedback can revise the topology itself. Based on this formulation, we propose HierFlow, a training-free hierarchical test-time search framework for automatic agentic workflow generation. HierFlow couples feedback-driven topology refinement with an MCTS-inspired lightweight tree search for execution-level sub-workflow optimization, and uses an adaptive gating mechanism to selectively trigger execution-level search based on estimated necessity. We further provide a coupling-aware analysis characterizing when hierarchical decomposition and proxy-based gating are beneficial and when their advantages may degrade under stronger cross-subtask coupling. Experiments across QA, mathematical reasoning, and code generation benchmarks show that HierFlow achieves strong performance and favorable efficiency--quality trade-offs, outperforming competitive baselines without additional training.
Abstract: Recent years have witnessed tremendous progress in enabling LLMs to solve complex reasoning tasks such as math and coding. As we start to apply LLMs to harder tasks that they may not be able to solve in one shot, it is worth paying attention to their ability to construct intermediate stepping stones that prepare them to better solve the tasks. Examples of stepping stones include simplifications, alternative framings, or subproblems. We study properties and benefits of stepping stones in the context of modern reasoning LLMs via ARQ ( uestions), a simple framework that introduces a question generator to the default reasoning pipeline. We first show that good stepping stone questions exist and are transferrable, meaning that good questions can be generated, and they substantially help LLMs of various capabilities in solving the target tasks. We next frame stepping stone generation as a post-training task and show that we can fine-tune LLMs to generate more useful stepping stones by SFT and RL on synthetic data.
Abstract: While monocular depth estimation has achieved significant progress, achieving generalized metric depth estimation for both narrow field-of-view (FoV) perspectives and 360^\circ panoramas remains an unsolved challenge. Existing methods are often tailored to specific camera types and struggle to produce accurate metric depth that generalizes across diverse settings. This limitation stems from two key challenges: the inherent geometric discrepancy between perspective and panoramic cameras, and the scarcity of panoramic training data with metric annotations. In this work, we introduce DepthMaster, a unified metric depth estimation framework. Rather than employing specialized networks to learn spherical distortions, we reformulate the problem by decomposing panoramic images into overlapping perspective patches. Crucially, distinct from prior projection-based methods that rely on ad-hoc architectural modifications to handle boundaries, we introduce a novel Correspondence Consistency Loss (CCL) and inject virtual projection cameras as geometric priors, allowing us to seamlessly stitch the patches while avoiding specialized operators and keeping the backbone largely compatible with standard Transformer designs. This strategy also resolves the geometric differences by unifying all inputs into a canonical perspective representation, and effectively circumvents data scarcity by directly unlocking powerful metric priors from vast perspective datasets. Trained on a mixed dataset that contains only one panorama dataset, DepthMaster achieves state-of-the-art zero-shot performance on 13 diverse datasets, outperforming not only universal methods but also leading specialist models in both perspective and panoramic domains. The code and models of DepthMaster will be released.
Abstract: Contextual Integrity (CI) defines privacy not merely as keeping information hidden, but as governing information flows according to the norms of a given context. As large language models are increasingly deployed as personal agents handling sensitive workflows, adhering to CI becomes critical. However, even frontier models remain unreliable in making disclosure decisions, and existing mitigation strategies often degrade underlying task performance. To overcome this privacy-utility trade-off, we propose SelfCI, a complementary self-distillation framework that decouples information suppression from task resolution. SelfCI jointly optimizes two independent reverse KL divergences over distinct teacher distributions derived from feedback: one encourages preserving task-relevant information for utility, while the other enforces minimal and appropriate disclosure. This complementary formulation induces a Product-of-Experts (PoE) target, aligning the policy with the intersection of capability and privacy requirements. Without relying on costly external supervision, empirical evaluations demonstrate that SelfCI consistently outperforms competitive baselines such as online reinforcement learning algorithms (e.g., GRPO). These trends further extend to out-of-domain settings involving agentic workflows and accumulated private context, suggesting that SelfCI provides a practical path toward CI alignment.
Abstract: Foundational marked temporal point process (MTPP) models, such as the Hawkes process, often use inexpressive model families in order to offer interpretable parameterizations of event data. On the other hand, neural MTPPs models forego this interpretability in favor of absolute predictive performance. In this work, we present a new family MTPP models: the hyper Hawkes process (HHP), which aims to be as flexible and performant as neural MTPPs, while retaining interpretable aspects. To achieve this, the HHP extends the classical Hawkes process to increase its expressivity by first expanding the dimension of the process into a latent space, and then introducing a hypernetwork to allow time- and data-dependent dynamics. These extensions define a highly performant MTPP family, achieving state-of-the-art performance across a range of benchmark tasks and metrics. Furthermore, by retaining the linearity of the recurrence, albeit now piecewise and conditionally linear, the HHP also retains much of the structure of the original Hawkes process, which we exploit to create direct probes into how the model creates predictions. HHP models therefore offer both state-of-the-art predictions, while also providing an opportunity to ``open the box'' and inspect how predictions were generated.
PaperID: 5097, Poster
Abstract: Parameter-efficient fine-tuning of quantized large language models is dominated by QLoRA, which combines 4-bit weight quantization with FP16 low-rank adapters. We show the FP16 adapter precision is unnecessary: a single binary bit per adapter parameter, combined with a learned per-layer scale s and a Scaled Straight-Through Estimator, matches or exceeds QLoRA across a broad range of benchmarks while reducing adapter storage by 14--16×. We introduce two variants--NormScale, which clips STE gradients to the unit ball, and LearnThresh, which makes the binarization threshold learnable---unified under a framework we call BQ-LoRA. Our main theoretical contribution is a tight characterization of the BQ-LoRA hypothesis class: it has bounded entrywise norm, a high-probability spectral norm bound that is \sqrtr times sharper than the worst case, a generically full-rank update, and consequently a Rademacher complexity that yields an explicit \tildeO(s_\max\sqrtd_\textin d_\textout/n) generalization bound. None of these properties hold a priori for FP16 LoRA. On Qwen2.5 (1.5B/7B/72B) and LLaMA-2 (7B/70B), BQ-LoRA matches or beats QLoRA and LoftQ at ~ 1/16 the adapter footprint; the predicted regularization signature appears empirically in dataset ablations and training-loss curves. We extend Punica's SGMV kernel with a binary inner-product primitive (BQ-SGMV) and prove a cache phase-transition result: shrinking adapters 16× enlarges the HBM cache from 2,500 to 39,000 adapters, eliminating misses under skewed workloads.
PaperID: 5098, Poster
Authors: Wenfeng Zou
Abstract: Addressing codebook collapse in vector quantization models is crucial for efficient discrete representation learning. Recent solutions increasingly rely on expressive codebook reparameterizations, improving adaptability at the cost of additional optimization and computational overhead. This motivates a simple question: what reparameterization scheme is just enough for efficient codebook learning? In this work, we propose Adaptive-Scale Vector Quantization (ASVQ), a lightweight reparameterization method for learning the codebook in a decoupled way. Experiments across image and audio tokenization tasks demonstrate that ASVQ consistently improves codebook utilization, reconstruction quality, and training stability, while remaining competitive with more complex reparameterization-based methods. These results suggest that ASVQ provides a favorable trade-off between codebook expressiveness and efficiency.
Authors: Phevos Paschalidis, Constantinos Daskalakis, Devavrat Shah
Abstract: We study causal inference under outcome interference for sequential, observational settings. We consider settings where the binary outcomes over N units are Markovian across T time steps; at each time step, the outcomes of N units have pairwise dependencies captured through an Ising model; and, each outcome is impacted through a latent external field capturing effects of latent confounders. Similar to panel data literature, these latent confounders are modeled to have a low-rank factor structure. Our data is a single sample from this high-dimensional distribution. To estimate causal quantities of interest, we provide a computationally efficient method based on Maximum Pseudolikelihood Estimation (MPLE) for learning the model parameters. Under reasonable assumptions, we establish non-asymptotic consistency for parameter estimation. Therefore, sampling from the learnt model enables faithful estimation of causal quantities of interest. We demonstrate the efficacy of the method through synthetic experiments as well as a real-world case-study investigating causal effects of vaccine rates on COVID-19 death rates within US counties nationwide.
Authors: Andrew Corbett, Archit Sood, Anna Tzatzopoulou, Sai-Aakash Ramesh, Tim Dodwell
Abstract: Recent work on recursive architectures has shown that tiny neural networks can be surprisingly powerful on structured reasoning tasks. The trick is to model reasoning trajectories with a latent dynamical system. We argue that the inference-time behaviour of these architectures is best understood as approximate inference over latent reasoning trajectories, with deterministic recursion as the one-particle, zero-noise limit. We make this view operational through guided stochastic exploration: stochastic perturbations of the reasoning dynamics propose neighbouring trajectories, and the model's existing early-stopping head reweights them online. The framework yields three label-free diagnostics: local stability, guide alignment, and cloud-token entropy. These predict, from inference traces alone, whether the procedure will help and which of its outputs to trust. On Sudoku-Extreme it lifts exact-solve accuracy from 85.9% to 98.0% without retraining; on Maze-Hard the diagnostics flag a misaligned guide, as validation performance later confirms. The same machinery thus characterises both when recursive reasoning has room to improve at the trajectory level and when the model's internal guide can recover it.
Authors:
Zheng Lian, Fan Zhang, Lan Chen, Yazhou Zhang, Rui Liu, Jinyang Wu, Haoyu Chen, Xiaobai Li, Xiaojiang Peng, Bin He, Jianhua TaoAbstract: Open-Vocabulary Multimodal Emotion Recognition (OV-MER) aims to predict emotions without being constrained by predefined label spaces, thereby enabling fine-grained emotion understanding. Unlike traditional discriminative methods, OV-MER leverages generative models to capture the full spectrum of emotions and employs emotion wheels (EWs) for metric calculation. Previous approaches primarily rely on token-level loss during training. However, this objective is misaligned with the metrics used in OV-MER, and these metrics cannot be directly optimized via gradient backpropagation. To address this limitation, we turn our attention to reinforcement learning, as this strategy can optimize non-differentiable objectives. We term this framework . Furthermore, we conduct extensive experiments to elucidate the role of reinforcement learning in this task, revealing the necessity of the reasoning process, the impact of different rewards, and the generalizability to other emotion tasks such as sentiment analysis and basic emotion recognition. Experimental results demonstrate that AffectGPT-RL yields significant performance improvements on OV-MER. Beyond this task, we also achieve remarkable performance gains on basic emotion recognition, attaining state-of-the-art results on MER-UniBench. To the best of our knowledge, this is the pioneering work exploring the role of reinforcement learning in OV-MER, providing valuable guidance for subsequent researchers. Our code is provided in the supplementary material and will be released to facilitate future research.
Abstract: LLM-based generation of SystemVerilog Assertions (SVA) is often reported as nearing saturation, with the strongest specialized model reaching ~76% accuracy on NL2SVA-Human. We show that this aggregate hides a temporal gap: models that appear strong overall still collapse to a few implication templates on bounded-delay and liveness specifications. The core issue is that the dominant recipe, supervised fine-tuning on NL/SVA pairs, optimizes token-level mimicry rather than the property equivalence that defines SVA correctness. We introduce Reward-Weighted On-Policy Distillation (RWOPD), an on-policy distillation method that samples student rollouts, scores them with an open SymbiYosys+Z3 Property-Equivalence Checker (PEC), and applies a verifier-reward-weighted forward-KL gradient from a frozen 14B teacher on verifier-passable rollouts. This keeps the supervision dense at every response token while grounding both selection and loss weight in property-equivalent behavior. RWOPD distills CodeV-SVA-14B into a Qwen2.5-Coder-7B-Instruct student that sets a new state of the art on NL2SVA-Human and NL2SVA-Machine across pass@1, pass@5, and pass@10, surpassing both specialized prior SOTA models and 671B general-purpose baselines.
PaperID: 5103, Poster
Abstract: Sparse autoencoders (SAEs) have emerged as a powerful technique for decomposing language model representations into interpretable features. Current interpretation pipelines infer feature semantics from activation patterns, implicitly assuming every feature has a semantic explanation while overlooking the computational roles features inherit from the SAE training objective. We introduce a weight-based interpretation framework that requires no activation data and grounds claims causally through targeted feature-ablation interventions. Across Gemma-2 and Llama-3.1, three depth-dependent signatures show that SAE features inherit the base model's geometry: semantic features form a U-shape under tied embeddings but bifurcate when untied, attention participation peaks mid-layer, and the two roles couple oppositely for input- vs. output-oriented features. This weight-based view supplies the causal half missing from activation-based interpretability.
Abstract: Human videos contain rich manipulation priors, but using them for robot learning remains difficult because raw observations entangle scene understanding, human motion, and embodiment-specific action. We introduce MoT-HRA, a hierarchical vision-language-action framework that learns human-intention priors from large-scale human demonstrations. We first curate HA-2.2M, a 2.2M-episode action-language dataset reconstructed from heterogeneous human videos through hand-centric filtering, spatial reconstruction, temporal segmentation, and language alignment. On top of this dataset, MoT-HRA factorizes manipulation into three coupled experts: a vision-language expert predicts an embodiment-agnostic 3D trajectory, an intention expert models MANO-style hand motion as a latent human-motion prior, and a fine expert maps the intention-aware representation to robot action chunks. A shared-attention trunk and read-only key-value transfer allow downstream control to use human priors while limiting interference with upstream representations. Experiments on hand motion generation, simulated manipulation, and real-world robot tasks show that MoT-HRA improves motion plausibility and robust control under distribution shift.
PaperID: 5105, Poster
Abstract: Video world models are now being adapted to medical applications such as surgical simulation, procedural rehearsal, and synthetic data generation, but their inference cost remains a major obstacle to practical use. Training-free caching is attractive because it can be applied to existing checkpoints, yet current methods face a clear trade-off. Lightweight rules track only the average change in the video and treat all regions of the frame as equally important, which is a poor fit for medical video where the meaningful motion is often confined to a small region such as a tool tissue interaction. Richer rules recover spatial awareness by partially running the network on every step to inform the skip decision, but this decision cost is paid whether the step is skipped or not, and it eats into the speedup. We introduce MedCache, a training-free and probe-free cache for rectified-flow video generation. MedCache reads a spatial risk signal directly from the latent using cheap arithmetic, with no help from the network itself. It combines this risk signal with a prompt-derived domain prior and a lazy consistency check that runs only when reuse is no longer clearly safe. Across three open-world model families in Text2World and Image2World settings, MedCache improves both speed and quality over the strongest training free baselines in the medical regime it targets. On Cosmos-H-Surgical 2B Text2World, it raises Total quality from 0.767 to 0.837 over EasyCache while reducing latency from 185 seconds to 95 seconds. Ablations show that risk-weighted drift drives the runtime gain, while the prior and guard recover fidelity.
PaperID: 5106, Poster
Authors:
Zhihan Zhang, Alexander Le Metzger, Jiuyang Lyu, Chun-Cheng Chang, Jiayi Shao, Yujia Liu, Emmanuel A Mensah, Edward Wang, Kurtis Heimerl, Gregory D Abowd, Natasha Jaques, Shwetak Patel, Vikram IyerAbstract: Embedded devices from wildlife monitoring stations to clinical wearables require local AI inference due to latency, communication, or privacy constraints. Optimizing models for heterogeneous microcontrollers (MCUs) requires simultaneously satisfying hard physical constraints on memory, power, and temperature while preserving accuracy, a multidimensional optimization that is today performed manually by experts. We ask whether an LLM agent can autonomously navigate this complex, multi-turn pipeline guided by real hardware feedback, and introduce a hardware-in-the-loop agent arena in which the agent iteratively refines both model and firmware---compiling, flashing, and measuring on real hardware---to enable iterative optimization. Frontier models, including Claude Opus 4.7 and Gemini 3.1 Pro, fail entirely without hardware feedback (0% deployment success), whereas our closed-loop formulation achieves the first successful deployment within three iterations and can surpass human expert results within seven. This agentic co-optimization achieves 250× compression for vision models with <3.3% accuracy loss and 400× for audio with <6% Feature Error Rate loss, enabling battery-free operation on a commercial MCU via solar harvesting. We demonstrate practical impact in two real-world systems: an elk-detection camera trap (96.7% accuracy) and a phonetic-transcription wearable (8.44% FER) for child development research.
PaperID: 5107, Poster
Abstract: With the rapid advancement of text-to-image diffusion models, increasingly realistic AI-generated content has raised serious concerns about misuse and copyright infringement. Digital watermarking offers a promising solution by enabling user-level traceability, yet existing methods remain difficult to deploy at scale, hindered by costly model retraining or detection, poor effectiveness–fidelity trade-offs, and limited robustness to image transformations. To address these challenges, we propose CLaW (Codec-guided Latent Watermarking), a robust and efficient watermarking framework for traceable diffusion image generation with frozen diffusion backbones and inversion-free detection. Specifically, CLaW first uses a pretrained watermark codec to map each watermark message to a latent residual, enabling low-cost sampling-time injection. Then, the residual is injected within a late denoising window, where image semantics are better preserved and the watermark signal is less perturbed by subsequent denoising updates, yielding a better balance between watermark effectiveness and image fidelity. Furthermore, we introduce a decoder-guided adaptive injection mechanism that uses decoding confidence as feedback to dynamically calibrate watermark strength, reinforcing weak watermark signals and improving robustness under image transformations. Extensive experiments show that CLaW can preserve visual fidelity and improve average F1 under image transformations, achieving a 6.34% gain over the state of the art.
Abstract: Large language models acquire rich internal knowledge during pre-training, yet compact models are still typically forced to relearn such knowledge from raw data or imitate it indirectly through distillation. We investigate a different route: can pretrained knowledge itself be directly reused during pre-training? To this end, we propose Memory Grafting, a representation-level transfer paradigm that extracts structured memories from a large frozen donor model and injects them into a compact student throughout training. Unlike knowledge distillation, which transfers behavior through output alignment, or retrieval-based methods, which expose external knowledge only at inference time, Memory Grafting makes pretrained knowledge part of the student’s optimization process itself. This suggests that knowledge in large language models can be modular, transplantable, and reusable across scales, pointing to a new path toward more efficient language model pre-training.
PaperID: 5109, Poster
Authors: Maxim Rabinovich, Harvineet Singh, Aman Sinha
Abstract: Reliable AI evaluation is a prerequisite for measuring progress and for deployment in safety-critical settings. Unfortunately, even gathering frontier system outputs on a full suite of relevant benchmarks can be expensive, let alone adding expert review or running evaluation inside of an optimization loop. Existing approaches to sample-efficient evaluation address this concern only partially: they typically assume abundant calibration data, pretrained error-relevant embeddings, or a fixed benchmark. Meanwhile, optimization loops require repeated evaluation and the ability to surface a long tail of errors; in fact, in real world settings, the long tail of errors often drives human-in-the-loop iterative system improvement. Working in the regime in which the only available information is system inputs and one or more cheap surrogate measures of system error, we derive an optimal importance sampler for error estimation and analyze its theoretical properties. Going further, we show how the same sampler can naturally discover and rank error patterns. On standard benchmarks for frontier model question answering (MMLU-Pro) and agentic coding (TerminalBench 2.0) we show our method consistently outperforms standard Monte Carlo on estimating aggregate error, recovering failure modes and calculating their prevalence, and detecting poison agent/environment injections.
PaperID: 5110, Poster
Abstract: Deep graph learning models deployed in real-world systems often need to cope with non-stationary environments, where the underlying graph distribution drifts continually over time. Prevailing solutions rely on training auxiliary generative modules to synthesize memory graphs for cross-domain adaptation, which incurs substantial computational overhead and scales poorly under prolonged distribution shifts. We argue that a more economical path exists: rather than generating memory, one can crystallize it. To this end, we propose Efficient Memory Crystallization (EMC), a training-free test-time framework that distills each incoming graph domain into a compact, semantically faithful memory through a closed-form solution to a memory-oriented distribution-matching objective, thereby eliminating redundant domain information under continual covariate shifts. To preserve both generalizability and adaptability as the model traverses a long sequence of target domains, EMC further models inter-domain dependencies through state-evolving memories and admits a theoretically grounded, tighter generalization error bound than direct adaptation. Extensive experiments demonstrate the superior performance of EMC over state-of-the-art baselines on graphs under non-stationary distribution shifts, while reducing average runtime by 87.4% and GPU memory consumption by 92.4% relative to the recent competitor, making continual graph adaptation practical at scale.
PaperID: 5111, Poster
Abstract: ) in the joint feature-structure space beyond label-related formulations. We show that CS induces bias at two non-interchangeable stages of ) biases propagation, rendering single-stage alignment intrinsically insufficient. Based on this, we propose ), a stage-aware GDA framework that aligns discrepancies at their source. DACS mitigates FS via adversarial feature alignment and addresses FCSS through layer-wise reweighting to correct propagation bias, followed by final adversarial alignment for residual mismatch. This design follows directly from the stage-wise structure of CS rather than heuristic combinations of alignment modules. We further show that representation and propagation biases correspond to distinct components of target-domain error that cannot be eliminated in isolation. Experiments on synthetic and real-world benchmarks demonstrate that DACS consistently outperforms prior methods, especially under complex and coupled distribution shifts.
PaperID: 5112, Poster
Abstract: Targeted Minimum Loss-based Estimation (TMLE) is a classical plug-in debiasing methodology that delivers doubly robust estimation and asymptotically efficient inference. Despite its statistical guarantees, each iteration of standard TMLE requires solving a minimization subproblem over the entire dataset, resulting in prohibitive computational and memory costs that render the method impractical for large-scale or real-time settings. To ease the bottlenecks, we introduce stochastic TMLE, a randomized targeting procedure that replaces each expensive full-batch fluctuation fit with mini-batch alternatives. Our theoretical analysis establishes that the stochastic iterates converge to a neighborhood (noise ball) centered around the target solution, and crucially, only a small number of subsequent full-sample TMLE iterations suffice to reach an empirical efficient influence function root. Consequently, our stochastic variant inherits the same guarantees and attractive properties as classical TMLE while substantially reducing per-iteration complexity. We further propose two computational variants with provable guarantees that broaden the algorithmic design space, offering flexibility for future developments in scalable targeted learning. Extensive experiments across diverse regimes demonstrate substantial acceleration over standard TMLE without sacrificing inferential quality.
PaperID: 5113, Poster
Abstract: Continued pretraining is optimized on a fixed self-supervised task but selected by downstream performance. This creates a coarse feedback loop: practitioners evaluate checkpoints, revise data mixtures or objectives, and rerun pretraining runs, while individual pretraining updates receive no signal about whether they help the target capability. We ask whether a small set of verifiable downstream examples can provide step-level feedback during continued pretraining without becoming learner supervision. We introduce V-pretraining, which separates a learner trained only by a self-supervised loss from a lightweight task designer that constructs targets or views for unlabeled batches. Given the current learner and an unlabeled batch, V-pretraining estimates the downstream value of a candidate target or view construction by the first-order predicted decrease in downstream loss after the self-supervised update it induces. The designer is trained to increase this value; the learner then applies the resulting self-supervised update with targets or views detached, so downstream labels never directly update learner parameters. V-pretraining can be used to learn adaptive top-K soft targets for language modeling and learned views for self-supervised vision. Under wall-clock-matched continued pretraining, V-pretraining improves GSM8K Pass@1 for Qwen models using 1,024 GSM8K examples only as feedback, including a +7.4 point single-run gain for Qwen2.5-0.5B. In vision, V-pretraining improves DINOv3 transfer to ADE20K semantic segmentation and NYUv2 depth estimation while preserving ImageNet linear accuracy, indicating that feedback-guided task construction improves target downstream capabilities without collapsing general-purpose representations.
PaperID: 5114, Poster
Abstract: Accurate 3D surface reconstruction from unposed images remains a fundamental challenge in computer vision. Recent feed-forward pointmap methods eliminate the need for camera poses but inherently produce discrete, sparse representations. In this paper, we present a generalizable pipeline that directly recovers continuous, high-fidelity 3D surfaces from unposed images by seamlessly integrating pointmap- and ray-based representations. Our approach first extracts explicit geometric priors using a pre-trained pointmap model. A query-view-centric selection module then identifies geometrically relevant reference views via a robust surfel voting mechanism. Next, a cross-attention regressor bridges target query rays with reference features to estimate a continuous coarse surface. Finally, a raylet-based module aggregates local 3D features to carve out high-frequency details and analytical normals. Trained as a single foundation model on 9 diverse datasets, our method demonstrates exceptional zero-shot generalization on 6 unseen datasets, clearly surpassing state-of-the-art baselines in geometric accuracy and detail preservation.
Authors:
Seonghyun Jin, youngmin Kim, Sunwoo Park, Jong Chul YeAbstract: Camera-conditioned video generation requires positional encoding that remains reliable under changes in camera motion, lens configuration, and scene structure. However, existing attention-level camera encodings either provide ray-only camera signals or rely on pinhole camera geometry, limiting their applicability to general camera control under the Unified Camera Model, including wide-angle and fisheye lenses. To address this limitation, we propose Curved Ray Expectation Positional Encoding (CRePE). CRePE represents each image token as a depth-aware positional distribution along its source ray, providing a Unified Camera Model-compatible positional encoding that naturally captures the curved-ray geometry of wide-angle and fisheye cameras. CRePE is implemented through a Geometric Attention Adapter added to frozen video DiTs, injecting token-wise scene-distance information into proper attention layers and stabilizing it with pseudo supervision from a monocular geometry foundation model. This design leads to more stable camera control and improved video quality in camera-conditioned video generation under the Unified Camera Model. Controlled positional-encoding ablations show consistent improvements over existing multi-view positional encoding, demonstrating the effectiveness of curved-ray-aware positional encoding across diverse camera models. Furthermore, by extending the same positional-encoding pathway to external geometry control through Radial MixForcing, CRePE supports external radial-map control for scene-geometry-conditioned generation and source-video motion transfer beyond camera control.
PaperID: 5116, Poster
Authors:
Tze Ho Elden Tse, Jizong Peng, Richard Chen, Angela YaoAbstract: 3D Gaussian Splatting (3DGS) has become the leading approach for photorealistic novel view synthesis, yet its geometric accuracy lags behind its visual quality. Existing depth-guided methods attempt to resolve this by regularizing the optimization with monocular depth priors, typically using a single global scale-and-shift alignment. Crucially, we discover a fundamental flaw in this assumption: modern monocular depth networks systematically exhibit region-dependent, piecewise-linear distortions. This discovery invalidates standard global alignment strategies, which warp scene geometry by enforcing a single transform. To address this, we introduce a spatially varying affine rectification model that corrects monocular disparity into consistent metric depth using regionally smooth, edge-aware fields. We further derive a normal-aware regularizer that uses spatial gradients to couple the rectified disparity with 3DGS depth and normals. Across diverse experimental settings, our sensor-free method shows strong photometric quality of Gaussian Splatting while substantially improving geometric accuracy and outperforming state-of-the art geometry-focused methods.
Abstract: Large language models (LLMs) are typically governed by post-training alignment (e.g., RLHF or DPO), which yields a largely static policy during deployment and inference. However, real-world safety is a full-lifecycle problem: static defenses degrade against evolving jailbreak behaviors, and fixed weights cannot adapt to pluralistic, time-varying safety norms. This motivates inference-time governance that steers behavior without costly retraining. To address this, we introduce the Consensus Clustering LinUCB Bandit (CCLUB), a unified framework for adaptive social alignment via system-prompt routing. CCLUB employs a conservative consensus clustering mechanism: it pools data only within the intersection of utility and safety similarity graphs, effectively preventing unsafe generalization across semantically proximal but risk-divergent contexts. Our theoretical analysis gives an expected regret bound O\left(d\log T/(p_\min\gamma^2\lambda_x) + d\sqrtMT\log T\right), separating the logarithmic cost of identifying safe consensus clusters from the dominant cluster-level exploitation term. Experiments show that CCLUB improves the safety--utility trade-off and offline deployment gap, and improves cumulative reward by 10.98% over the strongest non-prototype baseline.
PaperID: 5118, Poster
Abstract: Reproducing a published 3DGS paper takes weeks of expert engineering before any extension can be proposed. Prior autonomous-research systems read papers, draft hypotheses, and write code, but suffer from limited template length, executability, and traceability. Specifically, none has been demonstrated end-to-end in a live computer-vision subfield with intense activity. GS-Scientist is an autonomous research system for 3D Gaussian Splatting that addresses these limitations. Given a research prompt, a target-paper URL, or no input at all, it reproduces a published baseline, proposes an extension, trains and ablates across Mip-NeRF~360 and LLFF, writes the manuscript, and binds every reported number to a logged training run. Five design choices distinguish it: (i)candidates mutate plugins from a verified-reproduction backbone matched to published baselines; (ii)a cascade-weighted Elo tournament across eight islands scores on measured GPU PSNR with verbal-gradient feedback; (iii)a standing falsifier subtracts Elo on failure, integrating stress-test survival into selection; (iv)every reported number is bound to a logged training run, and the run cannot terminate while any claim is unbacked; (v)a seven-category 3DGS taxonomy wraps \textttEVOLVE-BLOCK markers in domain-aligned scaffolds transferable across 3DGS repositories. GS-Scientist produces manuscripts that beat their reproduced target by +0.18 to +3.78\,dB PSNR, including +3.78\,dB on astrophotographic nebula rendering, for which no prior 3DGS method exists. The research cycle compresses from months of graduate-student time to days of mostly autonomous compute. Code, all manuscripts, experiment ledgers, and the cross-run paper store will be released.
Authors:
Junnan Nie, Jiayi Li, Jiachen Zhang, Junyi Lao, Chenghao Liu, Tianle Zhang, Songfang HuangAbstract: Recent advances in vision--language--action and diffusion-based robot policies have largely come from training-side gains, but reliable deployment also depends on how predicted actions are executed. Under action chunking, each query predicts a sequence of future actions, and the robot executes an open-loop prefix before re-querying. The length of this prefix, the execution horizon, is a critical inference-time variable that trades off feedback frequency against motion continuity. Since fixed horizons exhibit strongly task-dependent and non-monotonic effects on success, no single constant horizon provides a reliable cross-task deployment rule. We propose PACE (Phase-Aware Chunk Execution), a training-free test-time execution method that selects the execution horizon online from the predicted chunk itself. PACE exploits the phase-dependent kinematic structure of manipulation trajectories, using prominent low-speed valleys in the predicted speed profile as candidate replanning points. On the 50-task RoboTwin2.0 benchmark, PACE improves the average success rate from 57.8% to 64.2% over the strongest fixed-horizon baseline, without task-specific fixed-horizon tuning. In real-robot experiments, PACE improves the average task score from 60.7 to 77.7 and the average success rate from 50.7% to 70.4%.
PaperID: 5120, Poster
Authors: Gia-Hung Pham, Tuan Dam
Abstract: Prioritized Sweeping (PS) accelerates model-based reinforcement learning by selecting backups according to residual magnitude. In nonstationary reward settings, however, the canonical priority score is shortsighted: after a localized reward shift, residuals propagate only through realized backups, so bottlenecked or topologically distant state estimates may remain static under a limited replanning budget. We introduce Graph Topology Augmentation for Prioritized Sweeping (GTA-PS), a concrete instance of a broader topology-augmentation principle for priority-based planning. GTA-PS constructs a policy-induced transition graph and augments the residual key with a mixing of regularized directed-Laplacian potentials that diffuses residual information through forward and backward graph structure. The topology weight is controlled by a scheduler based on the Second Largest Eigenvalue Modulus (SLEM), allowing the queue to adapt to the chain's mixing regime. We prove that the forward potential coincides with discounted residual propagation at a canonical regularization parameter and show that GTA-PS gives nonzero priority to states that standard PS can leave blocked after sparse reward shifts. Tabular experiments on FourRooms and GARNET domains demonstrate improved replanning efficiency over standard PS under both exact DP and Dyna-style host planners.
Abstract: Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implicitly assuming the teacher to be a reliable oracle. In large language models (LLMs), this assumption often fails: teacher predictions can exhibit high entropy and hallucinations, causing standard KD to degrade well-calibrated student priors. We propose CaRE-KD, a confidence-gated distillation framework that replaces static objectives with uncertainty-adaptive optimization. CaRE-KD has two components: a token-level loss (CaRE-Divergence) that adaptively switches between Forward and Reverse KL based on teacher--student confidence, and a batch-level epistemic rejection mechanism (Revival) that suppresses updates when the teacher is more uncertain than the student. We provide a gradient-level analysis showing how this dual-granularity design induces a conditional calibration mechanism that prior static divergences cannot reproduce. Empirically, across eight teacher--student pairs and eleven benchmarks spanning instruction following, chat alignment, code generation, and mathematical reasoning, CaRE-KD delivers consistent and significant gains over strong baselines (Skewed-KL, \alpha--\beta divergence). Highlights include up to +3.2 average ROUGE-L on instruction-following tasks, +2.1 pass@1 on MBPP, +1.7 accuracy on GSM8k, and +1.8 accuracy on CollegeMath over the strongest baseline, with consistent gains in LLM-as-a-judge factuality (up to +2.5 per task over Skewed-RKL). Revival further acts as a principled, loss-agnostic plug-in that systematically strengthens existing distillation objectives by filtering epistemically unreliable teacher supervision.
Abstract: Graph Domain Adaptation (GDA) aims to transfer graph classifiers across domains with both semantic and topological shifts. Existing Euclidean adversarial methods face two challenges: Structural Degeneration, where domain confusion entangles and suppresses label-relevant topology, and Optimization Instability, where minimax training induces oscillatory gradients under large structural shifts. We propose DisRFM, a geometry-aware GDA framework that addresses these challenges with Riemannian representation learning and flow-based transport. DisRFM embeds graph representations on a constant-curvature manifold and expresses them in geodesic polar coordinates. Polar endpoint regularization calibrates topologysensitive radial scales via univariate Wasserstein alignment and preserves scalenormalized class semantics through confidence-filtered angular alignment, with radial magnitude modulating pseudo-label reliability. DisRFM introduces topologyconditioned polar flow matching, which couples class-compatible source and target samples by a normalized polar transport cost and learns a metric-corrected vector field along geodesic interpolants. Theoretical analysis characterizes the structural risk of unconditional domain confusion and relates polar discrepancies and flow error to target risk. Extensive experiments under diverse domain shifts demonstrate that DisRFM consistently outperforms state-of-the-art methods.
Abstract: Long-context LLMs can infer objectives that are not stated explicitly. This capability is useful for reasoning over documents, code, retrieved evidence, and tool traces, but it also creates a safety risk: harmful intent can be distributed across a context and become visible only after the model composes the relevant pieces. Existing safety evaluations mostly test explicit harmful requests, and therefore miss this failure mode. We introduce compositional reasoning attacks, a long-context threat model in which harmful requests are decomposed into semantically incomplete fragments and embedded in long contexts. The final query is neutral; the harmful objective emerges only if the model retrieves the fragments, composes them, and infers the implied goal. We instantiate this setting using AdvBench requests, varying the required reasoning from Direct Retrieval to Single-hop Aggregation, Chain Reasoning, and Multi-hop Deductive Reasoning, and evaluate 15 frontier LLMs on contexts up to 64k tokens. Models usually refuse harmful requests when they are directly retrievable. However, refusal rates drop sharply when the same objectives must be reconstructed compositionally, often with larger failures in longer contexts. Benign reconstruction and fragment-position analyses indicate that these failures are not mainly retrieval errors: models often infer the harmful objective and then comply. Increasing inference-time reasoning improves refusal but remains incomplete and costly. Our results reveal a long-context safety gap: current models are better at refusing harmful requests they see than harmful objectives they infer.
PaperID: 5124, Poster
Authors:
Jinlong Yang, Jinke Wu, Lizilin, Yao ZhouAbstract: Diffusion Transformers (DiTs) have established themselves as the preeminent paradigm for high-fidelity generative modeling across various modalities, yet their practical deployment is hindered by the severe computational overhead of iterative sampling. While recent feature-caching methods mitigate this latency through feature reuse or forecasting based on integer-order operators, they remain limited in adapting to the non-stationary and highly curved dynamics of DiT latent trajectories. In this work, we reveal that DiT feature evolution exhibits anomalous diffusion characteristics and a certain degree of spatiotemporal heterogeneity, rendering static, integer-order forecasting operators inadequate and prone to error accumulation. To address this, we propose Fractional Anomalous Dynamics Extrapolation (FADE), a unified framework for training-free diffusion transformers acceleration. FADE models DiT feature evolution via anomalous diffusion and employs a Mittag--Leffler fractional operator to better adapt to high-curvature and non-linear latent trajectories of DiTs. Leveraging the empirical observation that these complex dynamics maintain cross-sample stability, we further construct a Look-Up Table with multiple dimensions through offline profiling. This enables zero-overhead, spatiotemporally adaptive parameter retrieval during online inference. Extensive evaluations across diverse architectures and modalities demonstrate that FADE achieves state-of-the-art acceleration, preserving exceptional generation fidelity under aggressive skip intervals with negligible latency overhead. Our code is anonymously available at https://anonymous.4open.science/status/FADE-8CB5.
PaperID: 5125, Poster
Authors: Haorong Han, Jidong Yuan, Chixuan Wei, Yongqi Sun
Abstract: Semi-supervised learning (SSL) relies on two core mechanisms: self-training under the Teacher-Student (T-S) framework and joint optimization of labeled and unlabeled losses. Despite their effectiveness, we find both mechanisms introduce distinct optimization pathologies. First, parameter coupling enforces strict synchronization between teacher and student, where strong regularization on the student degrades the teacher's fitting ability, thereby limiting the permissible generalization intensity. Second, the imbalance in gradient update consistency between labeled and unlabeled losses drives the shared parameters to prematurely converge to labeled-dominated local minima, creating a bottleneck for global optimization. To address both issues, we propose the Pioneer Student (PiS), an auxiliary branch that operates in an independent parameter space and periodically transfers accumulated knowledge back to the T-S model. Extensive experiments show that PiS is a universal plug-and-play module that consistently improves mainstream SSL methods.
PaperID: 5126, Poster
Abstract: Reinforcement learning (RL) often suffers from low sample efficiency and inefficient exploration in complex environments. Recent work leverages large language models (LLMs) as sources of prior knowledge for sequential decision-making. However, existing LLM-guided RL approaches typically rely on static or loosely coupled integration, failing to account for the evolving competence of the agent during training. As a result, LLM guidance may become redundant or even detrimental, leading to suboptimal learning dynamics, unnecessary dependence at inference time, and increased latency due to frequent LLM interaction. In this work, we propose TrustFlow, a unified framework that formulates LLM–RL integration as adaptive trust calibration. The key idea is to dynamically regulate the influence of LLM guidance based on learning progress, enabling a gradual transition from prior-driven exploration to LLM-free decision-making. Concretely, a preference representation is first constructed to connect sparse LLM rankings with policy representations. Trust calibration then balances LLM guidance with RL exploration. Building on this trust signal, value alignment is enforced, and a trust-modulated optimization scheme is introduced, incorporating preference structure into the learning process. Empirical results show that TrustFlow consistently improves sample efficiency and enables a smooth transition to fully LLM-free policies, eliminating the need for LLM access at inference time across diverse environments.
Abstract: Large language model-driven multi-agent systems (LLM-MAS) excel at complex tasks, yet unreliable agents remain a key bottleneck to system-level reliability. Automatic failure attribution is therefore critical, but existing approaches—such as direct prediction of agent–error pairs and agent-first failure attribution—rely on local logs of agent and miss global failures that only manifest over full interaction trajectories, such as cross-step inconsistencies and inter-agent coordination errors. Moreover, directly predicting failures induces a large combinatorial search space, hindering fine-grained attribution. To address these challenges, we propose VerifyMAS, a hypothesis verification framework for agent failure attribution. Instead of directly predicting faulty agents and error types, VerifyMAS formulates and verifies failure hypotheses against full trajectories. This verification-based approach decomposes attribution into trajectory-level error validation and fine-grained agent localization, providing an error-first attribution approach that captures global failure patterns while substantially reducing the search space. We further introduce a hypothesis-based data construction strategy grounded in a structured error taxonomy and fine-tune a specialized LLM verifier model for trajectory-level failure verification and agent attribution. Experiments on Aegis-Bench and Who&When show that VerifyMAS consistently improves diverse backbone models, including open-source Qwen and API-based GPT models, outperforming prior methods without sacrificing inference efficiency for long multi-agent trajectories.
Abstract: The first problem of modeling the world is not just estimating the right parameters or causal structure, but deciding what should be represented at all. We frame this as continual model construction: an agent maintains an environment-specific model M of an inaccessible world W and curates a persistent library \mathcalL of reusable representational elements across environments. We propose Representational Empowerment (\operatornameRepEmp) to score candidate elements by how much they expand the agent's future capacity to model and plan---a counterpart to classical environmental empowerment, redirected from control over external states to control over internal representations. We realize the framework as a hierarchical Actor-Curator architecture and test it across three experiments. In a finite-vocabulary causal-learning task, human participants construct causal models at varying abstraction levels to maximize goal reachability rather than fidelity to the world---a signature better predicted by \operatornameRepEmp over information-gain or novelty alternatives. Matched simulations reveal \operatornameRepEmp-guided construction, not exploration, to be the driver of sufficient structure recovery and cross-task transfer. Finally, in an open-vocabulary planning domain, an LLM-augmented Curator builds more compact symbolic libraries, which also generalize better than baselines. Ablating \operatornameRepEmp eliminates these benefits. Together, these results identify \operatornameRepEmp as a potential principle for continual model construction: deciding what to build, retain, and reuse under bounded resources.
PaperID: 5129, Poster
Abstract: Recent research on Multimodal Sentiment Analysis (MSA) has focused on learning from language, visual, and acoustic modalities with incomplete data to infer human sentiment. Most studies typically compensate for missing information by reconstructing modality features or designing complicated fusion mechanisms. However, these methods still suffer from spurious generation and noisy guidance due to the lack of high-level semantic grounding in partially observed multimodal evidence. To address these issues, we propose SemMSA, a latent semantic-aided framework that constructs rich sentiment-relevant semantics with LLMs, fully integrating with all modalities via anchor-free spectral alignment. It mainly consists of Cross-modal Semantic Refinement (CSR) and Cross-modal Spectral Alignment (CSA). Specifically, CSR first adaptively extracts visual and acoustic representations by corresponding adapters to form a unified multimodal prefix with language in the frozen LLM embedding space. It then iteratively produces continuous discriminative semantic states through a token-efficient latent refinement process without decoding explicit text. Next, CSA simultaneously aligns the refined semantics with all modalities by enhancing the dominant spectral component of their kernel Gram matrix. This captures global nonlinear dependencies among all representations without relying on a predefined anchor modality. In addition, an instance-level spectral separation constraint preserves cross-sample discriminability and mitigates representation collapse. Extensive experiments on SIMS, MOSI, and MOSEI benchmarks demonstrate that SemMSA achieves state-of-the-art performance.
PaperID: 5130, Poster
Abstract: Video generation must account for two sources of motion, one induced by the observer's camera path and the other caused by scene dynamics. An ideal camera-controlled video model should account for both motions: let users move the camera while evolving the scene dynamics. While current models handle camera-induced motion well in static settings, they struggle for dynamic scenes: objects are static, move incorrectly, or degrade in generation quality. We introduce DynaTokens, a lightweight set of learnable scene-specific tokens that teach dynamics to an existing camera-controlled world model. Our method is motivated by a simple asymmetry between the two sources of motion: whereas camera motion affects the generated view globally, object dynamics are spatially localized. Through cross-attention, DynaTokens trains the learnable tokens from a few example trajectories for a scene while keeping the base model frozen, and enables dynamics under new query camera paths. DynaTokens achieves a better simultaneous dynamics-camera tradeoff on VBench2 and WorldScore evaluations than LoRA, block finetuning, and specialized trainable-layer baselines. Analyses of token attention, ablations, and motion temporality suggest that matching the trainable interface to the structure of the learning target is important for effective adaptation.
PaperID: 5131, Poster
Authors:
Berlin Chen, Caitlin Wang, Aakash Sunil Lahoti, Kevin Li, Jay Shah, Jack Carlisle, Timmy Liu, Mengyu Guo, Zico Kolter, Albert Gu, Tri DaoAbstract: Linear-time sequence models such as State Space Models (SSMs) offer an efficient alternative to self-attention and have demonstrated strong language modeling performance at scale. In practice, however, their asymptotic advantage is often lost at common training and prefill lengths---optimized FlashAttention kernels remain faster than existing SSM implementations up to 8k tokens. In this work, we substantially narrow this performance gap with novel algorithmic improvements: First, we introduce a split-sequence algorithm that decouples the compute tile from the sequence-parallel split length, preserving efficient matrix multiplication tiles while substantially reducing boundary-state traffic and latency. Second, we resolve a longstanding speed--stability tradeoff in computing the 1-semiseparable (SS) decay mask, the main non-matmul bottleneck shared by all chunk-wise linear-time kernels (Mamba-2, Mamba-3, GDN, KDA). Existing kernels use the fast diff segsum (prefix subtraction), which suffers monotonicity violations and catastrophic cancellation---failures observed in real pretraining (e.g., Nemotron-H); the stable alternative (direct segsum) materializes the dense mask and is ~30% slower. Our hierarchy-aware segsum decomposes the SS mask into 1-SS diagonal blocks and rank-one off-diagonal blocks, computed via warp-level scans and cross-warp aggregation directly in registers, achieving the stability of direct segsum at the speed of diff segsum. We implement these algorithms using CuTe-DSL with warp specialization, asynchronous TMA, WGMMA, and layout-aware accumulator fragments. Across a large range of sequence length, our H100 kernels reach 91% of measured memcpy bandwidth, improve prefill time by 4.1× over the official Triton Mamba-2 kernel, and outperform the highly-optimized FlashAttention-4 starting at 1-2k sequence length---to our knowledge, the first time a linear-time sequence layer is faster than state-of-the-art attention at the 2-4k sequence lengths used in common pretraining and prefill workloads, where prior linear-time kernels only reached parity at \geq 8k. Our hierarchy-aware segsum is 1.35-1.40× faster than a naive stable direct-segsum implementation.
Abstract: Existing locate-then-edit Knowledge Editing (KE) methods typically decompose editing into two stages: upstream target representation optimization and downstream constrained parameter optimization. The optimization across the two stages is disconnected: upstream applies uniform regularization without observing downstream realization of the planned residual, hindering a refined accuracy–editability trade-off. Since this realization is request-specific and depends on downstream constraints, uniform regularization can over-shrink high-association requests, causing insufficient editing, while it can under-regularize low-association requests, producing over-large planned residuals that reduce downstream editability. To bridge this disconnect, we propose MetaKE (Meta-learning for Knowledge Editing), a new framework that unifies upstream and downstream stages into a bi-level optimization problem. The inner level optimizes parameter updates for the target representation, while the outer level optimizes representation using feedback from downstream constraints, achieving a better semantic accuracy-editability trade-off. To avoid costly multi-layer backpropagation, we introduce a Structural Gradient Proxy to approximate and propagate this feedback. Extensive experiments show that MetaKE outperforms strong baselines, offering a new perspective on KE. Code at: https://anonymous.4open.science/r/MetaKE-C160.
PaperID: 5133, Poster
Abstract: Structured prediction in dynamic, open-world environments—such as online High-Definition (HD) map construction for autonomous driving, mobile robotics, and embodied AI—fundamentally relies on the precise alignment of continuous geometry and discrete semantics. However, existing methods typically decouple these heterogeneous modalities, producing "overconfident yet erroneous" predictions that lack reliable uncertainty calibration in degraded environments. To address this fundamental limitation, we propose IntroMap, a probabilistic framework dedicated to learning joint semantic-geometric uncertainty with closed-loop calibration. At its core, the Semantic-Geometric Joint Calibration (SGJC) mechanism explicitly captures the aleatoric co-occurrence between spatial coordinates and semantic proxies via a Hybrid 3D Multivariate Gaussian distribution. Furthermore, to overcome the shortcomings of passive, open-loop uncertainty estimation, we introduce the Closed-Loop Feature Calibration (CLFC) module. It transforms the predicted joint covariance into latent modulation signals, actively rectifying degraded feature representations. Extensive experiments on the nuScenes and Argoverse 2 datasets demonstrate IntroMap's model-agnostic generalizability across multiple baselines (e.g., MapTRv2, MapQR), achieving state-of-the-art accuracy—including 63.1% mAP on nuScenes and a striking 8.1% mAP leap under challenging rainy conditions. Crucially, downstream end-to-end autonomous driving evaluations verify that our calibrated joint uncertainty establishes explicit safety margins, effectively preventing motion planners from blindly trusting hallucinated features and significantly reducing collision rates in long-tail scenarios. Code is available at https://anonymous.4open.science/r/IntroMap-F0BD/.
Authors:
Grigory Bartosh, David Ruhe, Emiel Hoogeboom, Jonathan Heek, Thomas Mensink, Tim SalimansAbstract: Diffusion models achieve state-of-the-art generative performance but suffer from high computational costs during inference due to the repeated evaluation of a heavy neural network. In this work, we propose Dual-Rate Diffusion, a method to accelerate sampling by interleaving the execution of a heavy high-capacity context encoder and a light efficient denoising model. The context encoder is evaluated sparsely to extract high-dimensional features, which are effectively reused by the light denoising model at every step to refine the sample efficiently. This approach significantly accelerates inference without compromising sample quality. On ImageNet benchmarks, Dual-Rate Diffusion matches the performance of standard baselines while reducing computational cost by a factor of 2--4. Furthermore, we demonstrate that our method is compatible with distillation techniques, such as Moment Matching Distillation, enabling further efficiency gains in few-step generation.
PaperID: 5135, Poster
Abstract: Reinforcement learning with verifiable rewards has shifted the alignment paradigm of large language models from subjective preference to objective correctness, yet a fundamental disconnect persists between exploration diversity and reasoning validity. Existing strategies predominantly rely on ``blind'' entropy metrics, e.g., token-level or global sequence-level entropy, which structurally fail to distinguish between effective reasoning and potential hallucinations. This misalignment creates an illusion of diversity, where high entropy tends to signal chaotic degeneration rather than valuable exploration. To bridge this gap, we introduce Effective Entropy, a validity-aware entropy concept that quantifies diversity exclusively within the valid subspace and systematically filters out the noise of incorrect trajectories. Built on this concept, we instantiate VEPO, a simple Validity-aware Entropy Policy Optimization method that directly incorporates effective entropy into RLVR training. VEPO induces an adaptive attraction dynamic that drives up the probability of under-explored valid trajectories, a mechanism that matches the practical regime of large-scale autoregressive models where sequence probabilities remain extremely small. Extensive experiments on reasoning benchmarks demonstrate that this effective-entropy-driven optimization significantly outperforms GRPO and other entropy-aware variants, showcasing its ability to more reliably cultivate, across diverse tasks, a broad portfolio of valid reasoning trajectories.
PaperID: 5136, Poster
Authors:
Donghyun Park, Junhyun An, Taehyoung Kim, Jisu KimAbstract: Time-delay embedding is a fundamental technique in Topological Data Analysis (TDA) for reconstructing phase-space dynamics of time-series data, where persistent homology can reveal loops associated with periodicity. However, rigorous statistical uncertainty quantification for these features remains underdeveloped. First, we analyze the topology of time-delay embeddings, showing that the embedded trajectory is homotopy equivalent to a circle (S^1) for periodic signals and contractible for non-periodic ones. We also prove a positive lower bound on the embedding reach, ensuring stable topological features. Second, we develop a subsampling approach to construct confidence bounds for persistence diagrams. Under standard manifold regularity conditions, we derive data-dependent bounds with asymptotic guarantees. Finally, we propose Topological Periodicity Test (TopPT), a hypothesis testing framework for periodicity with asymptotically controlled type I and type II error rates. Experiments on bounded-error synthetic data show that raw TDA detects periodic alternatives while avoiding false rejections on structured non-periodic signals, and that the robust rule is conservative after interpolation-error correction. On PhysioNet Fantasia and BIDMC respiratory waveforms, TDA detects most records but only a subset of local windows, unlike scalar periodogram baselines.
PaperID: 5137, Poster
Authors: Jing Liu, Yangyang YANG, Luca Ballotta, Fangfei Li, Yang Tang, Ruggero Carli
Abstract: This paper studies multi-agent reinforcement learning with submodular team utilities, which models scenarios where N agents solve a non-additive task allocation problem in a distributed manner online. Since each agent selects one action from a local categorical distribution at each time step, feasible joint actions form a partition matroid over agent-action pairs. The standard continuous relaxation of set utility functions, the Multilinear Extension, does not encode categorical constraints on factorized policies and may yield inconsistent gradient estimation. To remedy this, we propose the \emphPartition Multilinear Extension, a continuous relaxation that equals the expected team utility with factorized categorical policies under partition matroid constraint. We prove that submodular difference rewards provide unbiased PME marginal-gradient information and induce a stagewise score-function policy-gradient estimator for factorized categorical policies. Building on these results, we propose \emphSubMAPG, a centralized training with decentralized execution (CTDE) multi-agent policy-gradient framework that implements submodular difference-reward training signals and masked categorical policies for partition-feasible decentralized execution. For the associated PME marginal-space projected stochastic-gradient dynamics, we establish a stagewise \frac12-approximation guarantee and sublinear dynamic regret under slowly varying environments, measured by the path length of the optimal PME marginals. Finally, to handle open systems where agents and targets may leave and join over time (e.g., modeling failure and recovery of robots or smart sensors), we implement SubMAPG with a graph neural network policy model. Numerical experiments on multi-robot coverage and multi-target tracking show that SubMAPG outperforms local greedy and shared-reward baselines, and is competitive with centralized myopic greedy strategies.
PaperID: 5138, Poster
Abstract: Statistical modeling and inference for temporal networks are increasingly important across modern applications, yet remain challenging in the sparse regime with evolving network sizes and temporal dependence. We propose a tractable model that captures these key features in a unified framework. Under this model, we establish geometric ergodicity and characterize the asymptotic behavior of cycle counts. These counts provide informative low-order summaries of temporal network dynamics. Moreover, we develop method-of-moments estimators based on cycle counts. We prove identifiability and uniqueness of the parameter estimates, and establish the strong consistency and asymptotic normality. To support uncertainty quantification, we further propose a simulation-based plug-in estimator of the asymptotic variance. Extensive simulations and a real data example demonstrate that our methods are accurate and effective for sparse, size-varying temporal networks.
PaperID: 5139, Poster
Abstract: When learning new concepts sequentially, Vision-Language Models (VLMs) face forgetting that degrades both previously acquired tasks and inherent zero-shot capabilities. This challenge is particularly pronounced in the Cross-domain Task-Agnostic Incremental Learning (X-TAIL) setting, where the absence of explicit task identifiers forces diverse concepts to coexist within a unified and crowded multi-modal space. In this context, we identify an underexplored vulnerability termed Asymmetric Boundary Encroachment (ABE). Even when historical feature spaces remain stable, unconstrained new text features actively encroach upon these established sub-spaces during training. To systematically counteract ABE, we propose a unified framework. First, we introduce Analytic Decision Boundary (ADB), which constructs a geometric defense by enforcing an analytic margin derived from cumulative statistics to secure historical multi-modal boundaries. Furthermore, we integrate Orthogonal Visual Adaptation (OVA) and Analytic Ridge Classifier (ARC) to safely evolve the visual backbone and enhance joint inference. Experiments under X-TAIL setting demonstrate that our framework effectively mitigates ABE and consistently achieves new state-of-the-art performance. Our code is available at https://anonymous.4open.science/r/abe-cl-vlm.
Abstract: Building agents that can perform new skills by composing existing skills is a long-standing goal of AI agent research. Towards this end, we investigate how to efficiently acquire a sequence of skills, formalized as hierarchical neural options. However, existing model-free hierarchical reinforcement algorithms need a lot of data. We propose a novel method, which we call "AgentOWL" (Option and World model Learning), that jointly learns --- in a sample efficient way --- an abstract world model (abstracting across both states and time) and a set of neural options. We show, on a subset of Object-Centric Atari games, that our method can learn more skills using less data than baseline methods and possesses learning and generalization capabilities that the baselines do not have.
PaperID: 5141, Poster
Abstract: Latent generative models have emerged as powerful priors for solving inverse problems. These models typically represent a class of natural signals at a single, fixed complexity, governed by the latent dimensionality. This can be limiting: depending on the problem, a latent dimensionality that is too small may result in high representation error, while one that is too large may overfit to noise. We develop tunable latent priors for diffusion models, normalizing flows, and variational autoencoders, leveraging nested dropout. Across tasks including compressed sensing, inpainting, denoising, and phase retrieval, we show empirically that tunable priors consistently achieve lower reconstruction errors than fixed-complexity baselines. In the linear denoising setting, we derive the optimal complexity in closed form, showing how it depends on the noise level and the signal spectrum. This work demonstrates the potential of tunable latent generative priors and motivates both the development of supporting theory and their application across a wide range of inverse problems.
PaperID: 5142, Poster
Abstract: Coresets are a fundamental tool for scaling clustering algorithms, but standard constructions assume full access to the pairwise distances. In many modern settings, access to distances is costly and thus limited, and algorithms must instead rely on inexpensive comparison queries. We study coreset construction in the Rank-Measure (RM) model, where the algorithm can make noisy comparison queries to a quadruplet oracle alongside a small number of distance queries. In this model, we present the first algorithm to construct an \epsilon-coreset for (k,z)-clustering in Euclidean space. It matches the coreset-size bounds known in the classical model, while using only polylogarithmically many distance queries in the input size. Our approach is based on a uniform-sampling framework [Chen, SICOMP 2009], which decomposes the input into geometric rings and samples uniformly within each ring. We implement this approach using primarily noisy comparison queries, and then compress the result using a standard full-access coreset construction. In contrast, prior work in oracle-based clustering achieved only constant-factor guarantees. We complement our theoretical findings with experiments that demonstrate strong empirical performance against full-access baselines while using significantly fewer distance computations.
PaperID: 5143, Poster
Abstract: Graph Neural Networks (GNNs) have shown remarkable effectiveness across diverse graph-related tasks. However, existing studies reveal that their fairness is highly vulnerable to adversarial manipulation, which can substantially exacerbate inherent biases toward sensitive attributes, such as gender in demographic prediction. Although prior work has shown that fairness poisoning attacks can be launched via malicious node injection, whether this strategy can support the more challenging fairness evasion attacks remains open. Moreover, existing methods have not directly addressed fairness attacks on multi-class datasets with multi-valued sensitive attributes. To bridge this gap, we propose a universal fairness attack framework (UFA) that injects malicious nodes to perform both poisoning and evasion attacks against GNNs while minimally affecting model utility. UFA first identifies nodes most susceptible to being pushed across the decision boundary. It then targets core mechanisms shared by diverse GNN architectures to manipulate node representations, thereby maximizing outcome disparities across sensitive groups. Comprehensive experiments on five real-world datasets demonstrate the effectiveness of UFA. By injecting fewer than 0.2% malicious nodes, UFA severely degrades the fairness of both mainstream and fairness-aware GNNs, proving effective in both poisoning and evasion settings. Our findings expose the fragility of fairness in GNNs and underscore the urgent need for robust fairness-aware models.
PaperID: 5144, Poster
Authors: Zafaryab Haider, Hafizur Rahman, Aysegul Bumin, Prabuddha Chakraborty
Abstract: Automatic speech recognition (ASR) systems are usually judged by aggregate transcript accuracy, but real failures often hinge on meaning: for example, a wrong dose, address, account number, or date can matter more than many harmless word errors. We introduce SMASH, a framework for finding targeted semantic failures caused by sparse INT8 weight-bit perturbations, a dominant model data type for ASR edge devices, in quantized sequence-to-sequence ASR models. Given an audio input and a meaning-critical span,SMASH searches for small weight changes that replace the intended meaning while keeping the transcript readable, plausible, and close to the original output. Each accepted fault is a certificate that a specific semantic substitution is reachable under a bounded perturbation budget. Empirical study focuses on numeric and scalar substitutions across three ASR backbones: Whisper-small.en (WS), Whisper-large-v3 (WLv3), and SeamlessM4T-v2-large (S-M4T); and three corpora: an LLM-assisted controlled corpus (LACC), MultiMed, and LibriSpeech. On LACC at a 20-bit budget (B), SMASH-hybrid accepts numeric substitutions on 24/45 WS targets, 19/44 WLv3 targets, and 5/34 S-M4T targets; reachability is lower on real-speech datasets. In matched LACC budget sweeps tested up to B=40, WS reaches its ceiling already at B=20, while WLv3 and S-M4T continue to gain accepted targeted substitutions through B=40. Human validation supports the gate-based labels: on a 105-item shared validation set, independent human-majority labels and a four-judge large- language-model ensemble agree on the composite clean_success label, with \kappa=1.00. The faults are sparse: the pooled median is three INT8 bit flips, with medians of two flips on WS, eleven on WLv3, and twelve on S-M4T. Entity and negation pivots are evaluated as secondary evidence in the appendix. SMASH shifts ASR robustness evaluation from "how much did the transcript change?'' to "which meanings can be changed, how plausibly, and with how few bit flips?''
PaperID: 5145, Poster
Abstract: Reconstructing global sea surface pH from sparse observations is critical for monitoring ocean acidification and understanding marine carbon cycling. Traditional assimilation and inverse models are physically grounded but costly for large-scale reconstruction. Recent black-box and physics-guided AI models improve efficiency, but are mainly designed for passive tracers, where the reconstructed variable is also the transported inventory. In contrast, pH is an active carbonate tracer: it is the prediction target, while dissolved inorganic carbon (DIC) is the conserved carbon inventory. This mismatch can produce low pH error while violating carbonate closure and source-free carbon conservation. To address this, we introduce REACT, a carbon-first reconstruction framework that decouples transport, active correction, and chemical decoding. REACT transports a latent carbonate state with a conservative advection--diffusion solver, captures non-conservative carbon-cycle variations with a source module, decodes the corrected state into pH, and constrains the output through carbonate equilibrium. This design keeps pH as the target while enforcing consistency on the underlying carbon state. On simulation data, REACT reduces pH NRMSE by (14.7%) and chemical consistency error by (24.0%) over the best baseline. Cross-temporal-scale evaluations show robustness against error accumulation from coarse to fine temporal scales, and ablation studies validate the effectiveness of each component.
Authors:
Keqin Chen, Jie Bian, Yulian Wu, Vincent TanAbstract: Best arm identification under differential privacy is a pure-exploration problem in which both statistical efficiency and privacy protection must be achieved simultaneously. We study fixed-budget best arm identification for bandits under pure \epsilon-differential privacy, where the learner must recommend an arm after a prescribed sampling budget while protecting the full transcript. We prove that the optimal exponential decay rate of the error probability is upper bounded by an instance-dependent privacy-aware transportation exponent that differs from the analogous quantity used to characterize the stopping time of in fixed-confidence analysis by Jourdan and Azize [2025]. Guided by this exponent, we propose AO-Pri-BAI, an adaptive algorithm that maintains private running estimates through Laplace-tree mechanisms and learns a sampling design through a min-max interaction between hard alternatives and arm allocations. We prove that AO-Pri-BAI satisfies pure \epsilon-differential privacy. We also establish that the exponent of the failure probability of AO-Pri-BAI matches the privacy-aware benchmark. Numerical studies show that even in the non-asymptotic setting, AO-Pri-BAI outperforms benchmark algorithms on various instances, complementing the theoretical analyses.
PaperID: 5147, Poster
Authors:
Ruiting Dai, Zesen Cai, XiaoYu Zhao, Yiting Huang, Lisi Mo, Ming Li, Tao HeAbstract: Real-world multimodal data are often incomplete. Existing incomplete multimodal methods mainly address this problem as evidence attenuation, by reconstructing missing views, regularising shared representations, or retrieving auxiliary context.However, partial observations can still leave a different uncertainty unresolved:even when they provide a coarse task-state anchor, they may not determine whether the latent state should be revised toward a higher or lower label. We formalize this overlooked failure mode as residual-direction ambiguity and propose BRACE( Bipolar Reference-A ware Calibration and Estimation), a framework that first builds a quality-aware unified task state from available modalities, then organizes historical state-label pairs in a global memory bank, and finally infers a gated latent correction from the ordered contrast between higher-label and lower-label reference neighborhoods, together with a warmup-then-calibration training strategy and direction-consistency regularization. Across four benchmarks spanning multimodal sentiment analysis and multimodal misinformation detection under fixed and random missingness, BRACE consistently outperforms strong baselines, improving Acc-2 by up to 7.4 points over HME under fixed missingness and remaining robust when key modalities are absent; code is available at https://anonymous.4open.science/r/Brace-94BF.
Abstract: Prompt optimizers invest heavily in how to generate better prompts, yet pay almost no attention to which examples they use to judge them. This evaluation subset directly shapes every feedback signal that the optimizer receives, while existing methods either fix it before optimization begins (principled but agnostic to the evolving prompt population) or adapt it heuristically (flexible but unstable). We bridge this gap by constructing an online adaptive testing problem: Prompts are examinees, training examples are test items, and the scheduler selects items that best discriminate among the strongest candidates. We introduce POES (Prompt-Aware Online Evaluation Scheduling), whose monotone submodular objective combines discrimination, coverage, and bounded subset updates, yielding a (1-1/e) cold-start guarantee and a warm-start tracking bound under bounded inter-round drift. Across a 35-task APO suite plus 57 MMLU subjects and 6 optimizer-model configurations, POES achieves the highest mean accuracy in every configuration, improving over the best baseline by +3.0 pp to +7.8 pp. Notably, the achieved gains concentrate on tasks where baselines still have headroom to improve, vanishing where methods saturate, suggesting that what you evaluate on matters most precisely when there is a signal to extract.
Abstract: Autoregressive (AR) video generation extends videos by producing latent chunks sequentially, but scaling to long videos requires repeated access to a growing historical KV cache. Existing methods reduce this cost by truncating the KV cache or compressing it into implicit memory, but both lose explicit access to query-relevant historical details. We propose OmniMem, an explicit full-range memory retrieval framework that performs sparse KV retrieval over the historical cache. To make this practical for chunk-based AR video generation, OmniMem addresses two issues: (i) local bias in sparse KV selection and (ii) Union Explosion in memory access. Adaptive Window Exclusion removes local-window blocks from the selection candidates when sufficient long-range history is available, preserving the sparse budget for informative long-range retrieval. Query-shared KV Selection reduces cross-query diversity, while Per-Head Scattered KV Access avoids expanding head-specific selections into a large selected KV buffer. This allows each attention head to retrieve non-contiguous KV blocks according to its own selection pattern. Experiments on long-video generation show that OmniMem improves Dynamic Degree by 52.3% and preserves strong consistency over strong baselines, while maintaining comparable memory usage.
PaperID: 5150, Poster
Abstract: Although existing candidate trajectory evaluators have substantially improved end-to-end planning accuracy, their underlying mechanism remains suboptimal. Model-based evaluators typically make scoring future-aware by explicitly predicting candidate-conditioned future states over the planning horizon, which we term rollout. However, rollout incurs high inference cost and accumulates prediction errors. We argue that planning does not require reconstructing a full future scene for each candidate, but only future-relevant information for reliable comparison. We therefore propose RFWorld, a Rollout-Free World model for end-to-end trajectory evaluation. RFWorld constructs a shared scene memory and shapes it into a future-readable representation via lightweight training-time temporal readout with semantic supervision, keeping training overhead low. At inference time, candidate trajectories directly query the shared memory via sparse spatial readout for scoring, reducing computation by roughly the planning horizon times the number of candidate trajectories, while avoiding recursive rollout errors. Experiments on NAVSIM show that RFWorld learns a more planning-aligned world representation and achieves state-of-the-art performance.
PaperID: 5151, Poster
Authors:
Yuan Zhou, Hao Yu, Ruiran Cao, Haoran Yang, Xuanyu Zhu, Yusong Yan, Cong Wang, Minne LiAbstract: Generative models for physically grounded terrain face a persistent tradeoff: training-time physics losses distort the learned distribution, while unconstrained generators routinely produce physically impossible surfaces. We \emphdecouple generation quality from constraint satisfaction entirely, enforcing slope stability, bearing capacity, and surface continuity via a closed-form \emphpost-hoc projection applied to the generator's output. The projection raises the physics pass rate from 53% to \mathbf100% on a diffusion generator with no change in FID (13.72), whereas a matched PhysLoss baseline degrades FID to 56.6 ---a 4× loss of generation quality---while reaching only 25.5% pass rate. The projection is \emphmodel-agnostic: on four diverse generators (Procedural, Diffusion, VAE, GAN) it consistently yields 89.5%--100% compliance with <4.3% overhead. Out-of-distribution evaluation on real TartanDrive terrain confirms transfer: all four generators reach 100% pass rate, while projection simultaneously improves BEV-FID against the real reference by up to 20%. We instantiate the layer inside DreamEnv, an end-to-end BEV world model.
PaperID: 5152, Poster
Abstract: We study the training dynamics of multiclass logistic regression on high-dimensional Gaussian mixture models with a large number of classes and establish precise scaling laws governing the cross-entropy risk under gradient-based optimization. We show that learning proceeds sequentially across classes, from most to least frequent. When the class priors follow a power law distribution, the risk dynamics decompose into three phases: an initial plateau until the first class is learned, a power-law decay regime during which sequential learning occurs, and a final convergence regime. We then analyze how model capacity interacts with optimization under a fixed compute budget. When the effective dimension is restricted via projection onto leading principal components, the risk decomposes into a capacity term (a power law in the retained dimension) and an optimization term (a power law in training time). Optimizing this tradeoff yields a compute-optimal scaling law for logistic regression, with explicit prescriptions for model size and training time as functions of compute. These results extend theoretical scaling laws from linear regression to multiclass classification, while connecting to empirical scaling laws observed in large-scale neural networks.
PaperID: 5153, Poster
Abstract: How can we ensure that safety oversight models used to detect safety violations in AI systems reliably generalise, and how can we understand the factors that influence their generalisation? In this paper, we formalise PAC-Bayes certification for large language model-based safety oversight and obtain non-vacuous PAC-Bayes guarantees for safety oversight models, even under limited data for safety alignment. Building on compression-based PAC-Bayes bounds, we show that highly compressed PEFT adaptations yield extremely short adaptation description lengths, enabling informative and often tight guarantees that certify both classification risk and predictive uncertainty. We introduce a global-scale quantisation method (LoRA-GT) that reduces adaptor description length while preserving model performance, tightening bounds. Our results show that certifiability is strongly linked to adaptor compressibility, with shorter adaptor description lengths yielding tighter guarantees. Empirically, highly compressed adaptations exhibit minimal degradation in performance while enabling substantially stronger certification, suggesting that certifiable large language model oversight may naturally favour low-complexity safety adaptations. We further show that functional distortion tracks both test risk and bound tightness under compression, providing a practical mechanism for selecting simple, certifiable safety adaptations. Together, these results show that compression-based PAC-Bayes analysis provides a practical framework for understanding and designing reliable safety oversight models.
PaperID: 5154, Poster
Abstract: Long-form spatial reasoning in vision--language models (VLMs) is often bottlenecked by autoregressive chain-of-thought (CoT) decoding, where intermediate reasoning traces and final answers are generated token by token. This sequential decoding path makes high-quality spatial reasoning expensive in latency-sensitive embodied and interactive settings. We propose \emphDiffusion Thinking, a post-training framework that converts pretrained autoregressive VLMs into diffusion-style reasoners for fast long-form spatial inference without changing the visual encoder or language backbone. Diffusion Thinking partitions each CoT trace into blocks and refines tokens within the current block in parallel through discrete denoising, while preserving causal dependence across blocks and reusing KV cache for streaming generation. This creates a high-speed operating point, but aggressive parallel refinement can weaken fine-grained reasoning. To address this, we introduce \emphDiffusion Rethinking, a cached test-time scaling mechanism that reallocates part of the saved latency budget to self-correction rounds, yielding a controllable speed--accuracy frontier. We construct a 561K-instance Long-CoT Spatial Reasoning dataset with verified final answers and generated CoT traces across six spatial task types. Across Qwen2.5-VL and InternVL3 backbones, Diffusion Thinking with block size D=32 achieves 38.18--41.81× effective wall-clock speedup over autoregressive Long-CoT decoding. With eight rethinking rounds, it retains 4.20--4.60× speedup and matches or surpasses autoregressive Long-CoT accuracy on the largest backbones. These results show that diffusion-style compute reallocation is a practical path toward fast, accurate long-form spatial reasoning in VLMs.
Authors: Mehrdad Moghimi, Anthony Coache, Hyejin Ku
Abstract: Distributional reinforcement learning (RL) is a powerful framework increasingly adopted in safety-critical domains for its ability to optimize risk-sensitive objectives. However, the role of the discount factor is often overlooked, as it is typically treated as a fixed parameter of the Markov decision process or tunable hyperparameter, with little consideration of its effect on the learned policy. In the literature, it is well-known that the discounting function plays a major role in characterizing time preferences of an agent, which an exponential discount factor cannot fully capture. Building on this insight, we propose a novel framework that supports flexible discounting of future rewards and optimization of risk measures in distributional RL. We provide a technical analysis of the optimality of our algorithms, show that our multi-horizon extension fixes issues raised with existing methodologies, and validate the robustness of our methods through extensive experiments. Our results highlight that discounting is a cornerstone in decision-making problems for capturing more expressive temporal and risk preferences profiles, with potential implications for real-world safety-critical applications.
PaperID: 5156, Poster
Abstract: On-policy distillation (OPD) is a promising approach for transferring reasoning capabilities to capacity-constrained models, yet sampled-token OPD often suffers from entropy collapse and negative transfer. We identify uniform credit allocation as a key bottleneck: existing objectives obtain dense teacher feedback on student-generated trajectories, but apply it uniformly across sampled tokens despite highly heterogeneous token-level signals. Most tokens carry redundant near-zero credit, while rare heavy-tailed negative credits can dominate updates and prematurely suppress plausible reasoning trajectories. We introduce REOPOLD (Relaxed On-Policy Distillation), a framework that relaxes this uniform assignment by controlling where teacher feedback is applied, how strongly it affects the update, and when the allocation rule changes during training. Across diverse reasoning tasks, REOPOLD improves sampled-token OPD over recent post-training baselines with up to 12× higher sample efficiency, and extends to cross-vocabulary distillation, self-distillation, and test-time scaling.
Abstract: While streaming omni-video understanding demands continuous perception and proactive, real-time interaction, this crucial area remains largely under-explored. Current omni-modal methods are inherently designed for offline settings, limiting their applicability in streaming scenarios due to two fundamental flaws. First, they lack robust mechanisms to manage continuously growing audio-visual context over long horizons and cannot autonomously initiate responses at opportune moments. Second, existing benchmarks are predominantly confined to offline, single-turn question answering, failing to capture continuous, multi-turn streaming interactions. To bridge these gaps, we propose StreamOV, a novel Streaming Omni-Video understanding framework for efficient online audio-visual reasoning with bounded memory and proactive response triggering. Specifically, StreamOV introduces a multimodal evidence-guided long-short term memory that condenses historical audio-visual context into compact informative evidence under a fixed budget. It further employs a hidden-state-driven trigger to decide when to respond, avoiding explicit silence-token generation and external routers. We also curate SOVBench, the first comprehensive benchmark for online, multi-turn omni-modal evaluation. Extensive experiments show that StreamOV achieves state-of-the-art performance across diverse streaming and omni-video benchmarks, demonstrating its effectiveness for both online and offline video understanding.
PaperID: 5158, Poster
Abstract: Temporally extended exploration via graph Laplacian-based options is a promising approach to sparse-reward reinforcement learning (RL), but existing methods either do not explicitly target novelty or fail to scale to pixel-based domains under function approximation. Novel Exploration via Orthogonality (NEO) addresses the first issue by constructing options that navigate from highly visited regions toward less visited ones, yet prior results were limited to settings where exact eigenvectors can be computed. We present a scalable extension of NEO to pixel-based domains, built on three contributions. First, we use a novelty-weighted continuous Laplacian graph-drawing objective, which enables RL with continuous observations. Second, we embed the resulting eigen-potential options within a hierarchical reinforcement learning framework, enabling coherent temporally extended behavior. Third, we observe that learned eigen-potential rewards are directional but locally unreliable under online approximation; we therefore augment each option reward with a novelty bonus, a novel design idea that proves essential for stabilizing option learning while preserving novelty-directed exploration. Together, these contributions yield stronger and more persistent exploration, enabling longer option rollouts and better access to hard-to-reach novel states. Empirically, our method significantly outperforms both the prior scalable Laplacian-option baseline and a direct extension of NEO on sparse-reward benchmarks under a fixed budget. On Montezuma's Revenge, our best variant achieves approximately 1.8x higher return than both baselines. On Venture, both baselines yield returns near zero, whereas our method achieves a return of 1135. Across seven hard ProcGen games, our method achieves approximately 3.5x and 5.6x higher aggregate normalized return than the two baselines, respectively.
PaperID: 5159, Poster
Authors: Artur Chudzik, Henryk Josiński, Andrzej Przybyszewski
Abstract: Parkinson's disease therapy is usually adjusted based on extensive clinical examinations, but the effects of medication and deep brain stimulation (DBS) on gait dynamics remain difficult to quantify precisely. We use interpretable probabilistic machine learning to measure therapy-related gait changes in a repeated-measures motion-capture dataset comprising 376 walking trials from 19 patients with Parkinson's disease. Patients were recorded under four therapy conditions: medication off/stimulation off, medication off/stimulation on, medication on/stimulation off, and medication on/stimulation on. They performed natural and fast walking, and a study neurologist scored motor severity using UPDRS-III. From 1.0 s sliding windows of foot-marker motion, we compute short-window largest Lyapunov exponent (LLE)-derived local divergence features that capture local changes in gait trajectory dynamics. We summarize these features as two LLE-derived digital biomarkers: left-right asymmetry and left-right coupling. Bayesian hierarchical regression then estimates medication, stimulation, and task effects while accounting for repeated measurements within the same patients and patient-specific baseline differences. The main result is that during natural walking, LLE asymmetry decreased under combined therapy with a large paired effect size (d_z=-0.83). Coupling features changed with therapy in a complementary way, especially for stimulation-related changes in bilateral coordination. These findings suggest that short-window LLE asymmetry tracks therapy-related reduction of lateralized gait dysregulation, while LLE coupling tracks treatment-dependent reorganization of left-right coordination. By modeling novel LLE-derived biomarkers across controlled medication and DBS states, we show that the two therapies are expressed through distinct gait-dynamics signatures, supporting interpretable nonlinear gait biomarkers for objective therapy-response assessment.
Abstract: Diffusion models are increasingly used as controllable samplers, whose generations can be steered at inference time according to a chosen reward function. While such rewards are typically defined on individual samples, for many applications it is desirable to steer according to distribution-level rewards, for example to calibrate with population-level information or to encourage diversity. In both cases, simply incorporating the reward gradient into the dynamics, while often effective, comes with few theoretical guarantees on the sampled distribution. For pointwise rewards, recent work has therefore sought to develop a principled framework for targeting a prescribed tilted distribution using particle reweighting. However, an analogous theoretically-grounded approach for distributional rewards is currently lacking. In this work, we formulate inference-time distributional control as targeting a tilted measure under a mean-field framework, and derive a weighted interacting particle scheme to target it in a principled manner. Our framework recovers pointwise-reward steering as a special case, while providing a theoretical foundation for existing batch-level steering methods. Empirically, we verify that the procedure correctly targets the prescribed distribution in tractable low-dimensional settings, and investigate its behaviour in higher-dimensional protein conformation tasks.
PaperID: 5161, Poster
Authors: Seth Dale, Carolyn Koh, Dinesh Mehta
Abstract: Operator learning has experienced a shift toward attention-based architectures. This trend has more recently extended to physics-informed operator learning, where stacked cross-attention blocks add considerable computational cost relative to the foundational PI-DeepONet. We ask whether this complexity is necessary. Starting from the DeepONet branch-trunk decomposition, we replace the bilinear readout with a single softmax-weighted readout in which the trunk emits a query vector and the branch emits a set of key-value pairs—a structure we show admits universal approximation via reduction to a simplex-bilinear form. The resulting architecture, KerONet, uses no iterated attention, layer normalization, or explicit learnable projection matrices. Under a unified hyperparameter optimization protocol with a strict 200,000-parameter budget, we evaluate KerONet alongside iterated-attention architectures (PIT, PINTO) and PI-DeepONet on four canonical 1D PDEs whose solutions range from globally smooth to sharply localized. KerONet outperforms PIT and PINTO on every PDE, across all five seeds, in accuracy (approximately 2–4× lower rel-L^2 error), reproducibility (2–31× lower standard deviation), and speed (1.8–4.4× faster training time). Against PI-DeepONet, KerONet shows substantial gains where input-dependent localized features dominate (2.5–8× lower error on Burgers and Allen-Cahn) while exhibiting comparable accuracy on smoother operators (advection and diffusion-reaction)–a balance the iterated-attention architectures do not achieve. Our findings demonstrate that iterated cross-attention is not necessary to realize the benefits of query-dependent feature mixing in physics-informed operator learning.
PaperID: 5162, Poster
Abstract: Incremental learning is usually evaluated through the final classifier or LM head, but this interface does not indicate where forgetting occurs. A drop in accuracy may reflect representation degradation in the backbone, readout mismatch in the head, or both. We argue that BWT and FWT should therefore be interpreted as interface-dependent quantities rather than as single model-level properties. We study three evaluation interfaces: the sequentially trained head, a newly optimized probing classifier, and a classifier-free backbone separability diagnostic based on clustering. These interfaces answer different questions: deployed performance, recoverable information under a new readout, and representation-level separability. Across CIL experiments with encoder and decoder language backbones, plus discriminative-backbone TIL sanity checks, we find that qualitative conclusions about forgetting and plasticity change substantially with the evaluation interface. Standard heads often indicate severe forgetting, probes often suggest near-complete recoverability, and backbone diagnostics suggest intermediate representation changes and a stability--plasticity pattern. This pattern is consistent across five clustering algorithms and three clustering metrics, suggesting that it is not an artifact of a single diagnostic choice. Finally, we use Just LM-Head Tuning (JLT) as a head-only intervention to quantify recoverable performance on a fixed incrementally trained backbone. The large gap between standard and head-realigned performance suggests that many failures attributed to catastrophic forgetting are better understood as head--backbone mismatch plus partial representation drift. Our results call for reporting where forgetting is measured, not only how much forgetting is observed.
PaperID: 5163, Poster
Abstract: Despite rapid advances in text-to-image (T2I) generation, current models often struggle with rare, compositional, and culturally grounded concepts where training data is sparse and visual priors are weak or absent. Increasing model capacity or prompt detail alone does not consistently resolve these challenges, particularly in long-tail settings where attribute grounding remains unstable due to attention skew, semantic collapse, and needle-in-a-haystack retrieval dynamics. We introduce RAVEL, a training-free framework that improves rare concept generation, context-driven image editing, and self-correction by integrating graph-based retrieval-augmented generation (RAG) into T2I pipelines. Unlike prior RAG and LLM-enhanced approaches that rely on visual exemplars, static captions, or implicit model knowledge, RAVEL leverages structured knowledge graphs to retrieve targeted compositional, symbolic, and relational context, enabling more precise grounding even in the absence of visual priors. To further refine generation quality, we propose SRD, a self-correction module that iteratively updates prompts using multi-aspect alignment feedback, improving attribute accuracy, narrative coherence, and semantic fidelity. Our framework is model-agnostic and is compatible with leading diffusion and autoregressive models, including Stable Diffusion XL, Flux, DALL-E 3, GLM-Image, and Janus-Pro. We conduct extensive evaluations across three newly proposed benchmarks: MythoBench, Rare-Concept-1K, and NovelBench, and show consistent improvements over SOTA baselines across perceptual, alignment, and LLM-as-a-Judge metrics, while also demonstrating stable and interpretable refinement trajectories for rare concepts across diverse model families. These results position RAVEL as a simple and robust approach for controllable and interpretable T2I generation in long-tail domains. Our code is available at: https://anonymous.4open.science/r/ravel-54E5/
Authors:
Vasilis Gkolemis, Loukas Kavouras, Dimitrios Kyriakopoulos, Konstantinos Tsopelas, Dimitrios Rontogiannis, Giuseppe Casalicchio, Theodore Dalamagas, Christos DiouAbstract: Generalized additive models (GAMs) offer interpretability through independent univariate feature effects but underfit when interactions are present in data. GA^2Ms add selected pairwise interactions which improves accuracy, but sacrifices interpretability and limits model auditing. We propose \emphConditionally Additive Local Models (CALMs), a new model class, that balances the interpretability of GAMs with the accuracy of GA^2Ms. CALMs allow multiple univariate shape functions per feature, each active in different regions of the input space. These regions are defined independently for each feature as simple logical conditions (thresholds) on the features it interacts with. As a result, effects remain locally additive while varying across subregions to capture interactions. We further propose a principled distillation-based training pipeline that identifies homogeneous regions with limited interactions and fits interpretable shape functions via region-aware backfitting. Experiments on diverse classification and regression tasks show that CALMs consistently outperform GAMs and achieve accuracy broadly comparable to GA^2Ms, while preserving the univariate auditability that GA^2Ms forfeit. Overall, CALMs offer a favorable trade-off between predictive accuracy and interpretability.
Abstract: Imitation learning (IL)—training an agent to replicate expert behavior from demonstrations—underpins applications from robotics to language model training. Standard approaches such as Behavior Cloning (BC) are known to suffer from compounding errors and performance plateaus, particularly when the learner cannot perfectly represent the expert’s policy (e.g., as is typical in distillation). Two interventions are widely understood empirically to improve performance: querying the expert interactively along the learner’s own trajectories, and using value function estimation en route to generating a policy. We investigate the nature of these improvements and their potentially surprising interplay. Our main finding is that expert interaction relaxes the representational demands on the learner: one only needs a model capable of realizing the expert’s value function, bypassing the (often stricter) requirement of realizing the expert’s policy itself. Concretely, we introduce OVI, an interactive IL algorithm that is statistically and computationally efficient whenever the learner can represent the expert’s value function. We complement this with a negative result showing that interaction is necessary: without significantly stronger representational assumptions than expert-value realizability alone, a broad class of value-based IL algorithms cannot succeed in the offline setting. These findings bear out empirically: OVI outperforms offline policy-based (BC), interactive policy-based (DAgger), and offline value-based IL methods, with the largest gains when the learner network is substantially less expressive than the expert’s.
PaperID: 5166, Poster
Abstract: Cross-device deployment of a single reconstructed avatar often requires adaptation to varying Gaussian budgets with minimal degradation in visual quality. Motivated by the need for a robust budget-adaptive 3D representation, we propose RobustAvatar, a train-once Level-of-Detail (LoD) adaptation framework for dynamic Gaussian avatars based on sequential selection and compensation. RobustAvatar supports flexible compression-ratio control within a single trained avatar, without requiring fine-tuning at each compression ratio. The framework consists of two coordinated modules: Semantic Distribution-Regularized Selection (SDS) and Canonical Budget-Conditioned Compensation (CBC). SDS performs budget-aware Gaussian selection guided by semantic distribution regularization, preserving the coverage of perceptually salient facial regions relative to the full avatar under constrained budgets. CBC refines the selected primitives in a canonical, expression- and pose-neutral space by predicting ratio-conditioned attribute residuals, reducing detail loss from Gaussian removal and improving robustness to decreasing Gaussian budgets. Experiments on the NeRSemble dataset show that RobustAvatar achieves more robust budget adaptation than state-of-the-art methods, with slower quality degradation as Gaussian budgets decrease. The code will be made publicly available.
Abstract: Controllable 3D part generation is pivotal for digital asset creation, yet existing methods lack the flexibility to accommodate diverse user workflows requiring varying levels of control. In this paper, we introduce FlexPart, a unified image-to-3D part generation framework that supports variable-granularity geometric conditions within a single model. By seamlessly integrating heterogeneous inputs, ranging from sparse points to dense bounding boxes and masks, via gated adaptive modulation, FlexPart unlocks a powerful cross-granularity synergy. We also propose an asymmetric geometric guidance mechanism that leverages strong spatial priors to align representations across granularities, improving geometric consistency and facilitating inter-task synergy. Extensive experiments demonstrate that this synergistic approach outperforms state-of-the-art methods by over 10.5% on Part CD, achieving high-fidelity part generation with superior geometric consistency across all prompt granularities. Code and models will be made publicly available.
Authors:
Jasper van Doornmalen, Mathieu Molina, Victor Verdugo, José VerschaeAbstract: Motivated by the optimization of bounded binary black-box functions, we study the problem of learning polynomial surrogates over the Boolean hypercube. To ensure that optimizing the surrogate yields good solutions for the underlying objective, we require uniform L_\infty-error guarantees rather than the usual L_2-type guarantees. We characterize the minimax sample complexity of uniform estimation under subgaussian noise for two classes of bounded polynomials. First, for polynomials of degree at most d on n variables, the sample complexity scales as n^d+1. Second, for s-sparse Fourier-Walsh polynomials with s \leq n, it scales as ns^2. These rates differ structurally from the noiseless setting, where uniform exact recovery scales as n^d and ns, respectively. Our lower bounds hold even for arbitrary adaptive learners, showing that the additional factors are intrinsic to the noisy cases. Standard Fourier-analysis tools for the L_2-norm do not naturally extend to the L_\infty-setting in a way that yields uniform guarantees. Our proofs overcome this difficulty by relying on suitably chosen auxiliary norms that serve as proxies for controlling the L_\infty-error. Together, our results provide a tight characterization of the sample complexity of learning optimization-safe polynomial surrogates.
Authors:
Ali Zia, Usman Ali, Abdelwahed Khamis, Muhammad U Ramzan, Abdul Rehman, Wei XiangAbstract: Deep topological data analysis (TDA) offers a principled framework for capturing structural invariants such as connectivity and cycles that persist across scales, making it a natural fit for anomaly segmentation (AS). Unlike threshold-based binarisation, which produces brittle masks under test-time distribution shift (TTDS), TDA allows anomalies to be characterised as disruptions to global structure rather than local fluctuations. We introduce TopoOT, a topology-aware optimal transport (OT) framework for source-free test-time adaptation in AS. Our key innovation is Optimal Transport Chaining, which sequentially aligns persistence diagrams (PDs) across thresholds and filtrations, yielding geodesic stability scores that identify features consistently preserved across scales. These stability-aware pseudo-labels supervise a lightweight head updated online using only unlabelled target samples, without access to source data or target labels, with OT-consistency and contrastive objectives, ensuring robust adaptation under TTDS. Across standard 2D and 3D anomaly detection benchmarks, TopoOT achieves state-of-the-art performance, outperforming second-best methods by up to +24.1% mean F1 on 2D datasets and +10.2% on 3D AS benchmarks.
PaperID: 5170, Poster
Abstract: Text-to-image (T2I) diffusion models often synthesize policy-violating content, necessitating robust safeguards and rigorous evaluation frameworks. However, current approaches remain largely monolithic in both execution and evaluation, limited to coarse-grained categorization within broad unsafe concepts. Such approaches fail to account for the specific constituent factors of a violation, resulting in weak learning signals and allowing model collapse to bypass safety checks. To overcome these limitations, we propose Example-Based Spatial Guidance (EBSG), a training-free method that utilizes user-editable exemplar packs to provide granular spatial guidance for precise concept steering. Unlike existing methods that coarsely steer the entire image away from a broad concept, EBSG decomposes these categories into specific sub-concepts via exemplar text-image pairs, providing explicit, localized steering signals for more robust safety control. Furthermore, we introduce a vision language model (VLM)-based evaluation protocol that provides a fine-grained assessment, avoiding the pitfall of conventional binary evaluators that permit model collapse to bypass safety checks. Empirically, EBSG achieves state-of-the-art erasure performance across diverse safety datasets and concept categories, spanning four nudity benchmarks, seven I2P harmful-concept slices, and four MJA metaphor categories, while supporting multi-concept removal and transferring to a modern backbone, SD3.
Authors: Shei P Chua, Fangzhao Wu
Abstract: Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed and informs the design of robust alignment strategies. Prior work shows that aligned LLMs encode harmfulness and refusal as separable directions in the residual stream at prompt-side token positions. We show that jailbreaks succeed at prompt encoding by suppressing either the refusal or harmfulness direction before any token is generated, with distinct attack classes occupying separable regions of the harmfulness-refusal plane. Extending the analysis to response-token positions, we find that the model recognizes harmful content while it is generating that content, even when it failed to recognize the input as harmful at the prompt side. Motivated by our findings, we introduce HARC (Harmfulness-And-Refusal Coupling), a fine-tuning method that pairs the two directions across both prompt and response positions. Since the intervention is confined to the harmfulness-refusal subspace, it leaves the rest of the residual stream intact and does not degrade general capability or inflate over-refusal. Across extensive experiments, HARC achieves the strongest robustness-capability-usability trade-off among six baselines spanning the major training-time and inference-time safety methods. The harmfulness and refusal directions at prompt and response positions transfer across the five model families and two scales we tested without architecture-specific tuning.
PaperID: 5172, Poster
Abstract: Large language models (LLMs) deployed as agents remain vulnerable to prompt-based instruction injection, where user inputs induce the model to violate higher-priority system directives. Existing methods often treat violations as a coherent class or rely on surface text patterns, which can limit robustness as attack strategies evolve. In this paper, we investigate LLM hidden-state dynamics under attack and uncover a phenomenon we term Asymmetric Logic: the model's response generations concentrate in a relatively low-variability region of latent representation space when it's compliant to the system prompt, while compromised behaviors exhibit diverse deviations. Based on this observation, we propose Asymmetric Entropic Optimization (AEO), a novel framework that monitors the internal shifts in latent representations to detect and proactively rectify anomalous behaviors. Our method effectively identifies logic deviations by quantifying the entropic divergence between the compliant manifold and the chaotic violation space. Across a range of state-of-the-art LLM backbones, AEO demonstrates strong robustness to diverse instruction injection attacks. For instance, on Qwen2.5-14B-Instruct, AEO attains a detection AUC of 97.28% with FPR95 of only 10.38%. The same chaos score also gates a lightweight post-hoc correction module that lifts overall instruction-following accuracy from 63.94% to 76.00% on IHEval, without any model fine-tuning. The code is available at https://anonymous.4open.science/r/AEO-code-56D2/.
PaperID: 5173, Poster
Abstract: Agentic systems powered by large language models have advanced rapidly yet current architectures treat information acquisition as a homogeneous process, issuing broad searches, accumulating heterogeneous evidence, and deferring verification to post-hoc reconciliation. This neglect of acquisition-mode structure leaves factual grounding and epistemic reliability as emergent rather than engineered properties. In this paper, we formalize the Deterministic-Agentic (D/A) Separation Principle, which partitions the task-relevant information space into programmable (\mathcalD) and emergent (\mathcalA) zones, establishing \mathcalD as both epistemic anchor and investigative compass for \mathcalA-directed exploration. We instantiate this principle as Z-AXIS, a three-phase cascade that front-loads deterministic grounding, deploys hierarchical wave agents for macro-to-micro deep research, and synthesizes evidence through tension-aware cross-dimensional analysis. Key mechanisms include authority-calibrated evidence fusion, D-conditioned coverage tracking, and temporal awareness for longitudinal evaluation. Extensive evaluation across multiple industries demonstrates that Z-AXIS achieves superior factual accuracy and epistemic reliability compared to existing deep research systems, with architecture-determined reliability remaining stable across a wide range of backbone scales, confirming that structured governance, not model scale, governs epistemic reliability.
Abstract: We present FRUC, a feed-forward 3D Gaussian splatting framework for dynamic scene reconstruction from uncalibrated collaborative driving views. Existing multi-agent reconstruction frameworks are often hindered by rigid prerequisites, demanding precise spatial calibration and slow per-scene optimization. In this paper, we rethink this task by conceptualizing a distributed multi-vehicle network as a spatio-temporally unstructured ego-centric multi-camera system, where the core challenge lies in enhancing ego-centric occluded geometry through collaboration without degrading the ego's accurately observed visible geometry, while preserving reconstruction efficiency. For efficient reconstruction, FRUC is built upon a visual grounded geometric Transformer backbone to enable one-shot, calibration-free inference from a flexible number of multi-vehicle views. To achieve non-destructive geometric supplementation under uncalibrated cross-agent misalignment, FRUC first introduces an ego-centric causal occlusion field that explicitly derives occlusion evolution as latent priors by modeling agent-wise spatio-temporal correlations. Guided by these occlusion priors, it further formulates cross-agent integration as a deterministic residual denoising process via zero-initialized injection, turning challenging cross-agent fusion into bounded residual learning for robust collaborative blind-spot completion. Through extensive evaluations on the real-world V2XReal and UrbanIng-V2X datasets, FRUC is shown to be a new state-of-the-art for the scene reconstruction of dynamic collaborative driving environments, significantly outperforming existing methods in both rendering quality and efficiency.
PaperID: 5175, Poster
Abstract: Tropical cyclone forecasting models are often trained and evaluated with reanalysis inputs. Reanalysis provides physically rich environmental context, but its most recent fields may be unavailable at forecast issuance. This can introduce a hindsight-style advantage and lead to optimistic estimates of real-time forecast skill. Satellite observations, in contrast, are available closer to real time but provide incomplete and noisy views of the storm system. We formulate this data-availability mismatch as forecasting through a reanalysis blind window. To study this setting, we introduce a real-time-faithful benchmark spanning 1980--2023 that combines lagged reanalysis with visible, infrared, water-vapor, and passive microwave satellite imagery. We then develop RECAST-TC, a multimodal reconstruction framework that fuses time-lagged environmental histories with real-time satellite observations to estimate a forecast-sufficient latent storm state from delayed, partial, and noisy inputs. We use idealized analysis-available forecasting and delayed-reanalysis-only forecasting as reference regimes. RECAST-TC consistently improves track and intensity forecasts over delayed-reanalysis baselines and moves noticeably closer to the idealized reference. These results highlight data availability and observation timing as central design variables for realistic evaluation and deployment-oriented AI forecasting of tropical cyclones.
Abstract: Fine-grained image retrieval (FGIR) typically relies on supervision from seen categories to learn discriminative embeddings for retrieving unseen categories. However, such supervision often biases retrieval models toward the semantics of seen categories rather than the underlying appearance characteristics that generalize across categories, thereby limiting retrieval performance on unseen categories. To tackle this, we propose GAPan, a Generative Appearance Prior alignment network that reformulates the learning objective from category prediction toward appearance modeling. Technically, GAPan treats retrieval features with an invertible density model based on normalizing flows. In the forward direction, the flow maps all instance features into a latent density space, where each seen category is modeled by a class-conditional Gaussian prior and optimized via exact likelihood estimation. This formulation preserves richer appearance details by leveraging the invertible property of the flows. In the reverse direction, samples from the high-density regions of these learned priors are mapped back to the feature space to produce appearance-aware anchors that reflect intra-category variation. These anchors supervise a prior-driven alignment objective that aligns retrieval embeddings with category-specific appearance distributions, thereby improving generalization to unseen categories. Evaluations demonstrate that our GAPan achieves state-of-the-art performance on both widely-used fine- and coarse-grained benchmarks.
PaperID: 5177, Poster
Abstract: Large language model (LLM)-based coding agents achieve strong results on controlled benchmarks yet routinely produce pull requests that real maintainers reject. The cause is not functional incorrectness but a lack of organicity: generated code ignores project conventions, duplicates internal APIs, and violates implicit architectural constraints. The latest repository snapshot is insufficient-it reveals the final codebase state, not the change patterns that shaped it. We make three contributions. (1) Paradigm: we propose next-commit prediction as a paradigm for learning agentic coding skill from a repository's own history-each historical commit is a self-supervised target whose oracle diff supplies dense supervision. (2) Evaluation: we establish Organicity Evaluation as a measurable objective for coding agents and contribute the Learning to Commit benchmark-5 curated open-source GitHub repositories under strict per-repository temporal splits-scoring patches along file localisation, internal API reuse, patch bloat, and code-style consistency. (3) Method: we instantiate the paradigm with the Learning to Commit framework, in which the agent performs supervised contrastive reflection: it blindly attempts each historical commit, contrasts its prediction against the oracle, and incrementally distils a reusable skill document that conditions subsequent generations. On the Learning to Commit benchmark and on SWE-bench Pro reframed under our temporal-split protocol, our framework consistently improves organicity on held-out future tasks and further lifts test-pass rates on SWE-bench Pro, narrowing the gap between benchmark success and real-world mergeability.
PaperID: 5178, Poster
Authors:
Yuan Wu, Jiayu Qian, Sipeng Wu, Songpan Gao, Ruiming Ye, Zongxian Yang, JinYue Li, Qiankun Li, Yang LiuAbstract: Medical image interpretation requires diagnoses grounded in case-specific visual evidence. However, medical vision-language models often produce plausible answers by exploiting clinical language priors and report-level co-occurrence patterns rather than faithfully using the input image, leading to Visual De-anchoring and Spatial-Semantic Binding Breakdown, where reasoning detaches from visual evidence or binds correct answers to incorrect anatomical support. In this work, we introduce MedVIGOR, a framework that internalizes visual evidence as an intrinsic constraint on medical VLM reasoning. MedVIGOR combines visual-necessity supervision, dual-pathway visual anchor internalization with MedSAM-derived spatial support and DINOv2-derived semantic patterns, and evidence-consistent reinforcement learning to preserve alignment among visual evidence, intermediate reasoning, and final answers. We further present MedVIGOR-Bench, a unified benchmark for evaluating whether medical VLMs are both diagnostically correct and evidentially grounded across concept recognition and visual localization tasks. Experiments show that MedVIGOR improves diagnostic accuracy and grounding fidelity over strong medical VLM baselines, highlighting the value of internalized visual evidence for trustworthy medical reasoning.
PaperID: 5179, Poster
Abstract: Neural surrogate models for Partial Differential Equations (PDEs) on unstructured 3D geometries are often limited by poor generalization and the high cost of generating large-scale training datasets. Consequently, pre-training has emerged as a critical alternative to enhance the robustness and scalability of these models. In this work, we introduce a pre-training framework tailored to both steady-state and transient regimes. For steady-state problems, we propose a geometry-driven strategy that leverages intrinsic shape descriptors to learn representations of complex 3D domains. For transient problems, we introduce a physics-driven approach based on online generation of synthetic PDE data, enabling scalable pre-training without reliance on expensive datasets. This framework learns transferable features for challenging downstream tasks. Across multiple experiments, our approach achieves up to 3× faster convergence, 2× greater data efficiency, and up to 25% higher accuracy during fine-tuning, particularly under realistic low-data regimes. This methodology provides a practical pathway toward data-efficient neural emulators for large-scale simulations.
PaperID: 5180, Poster
Abstract: Modeling gene regulation in single cells requires jointly representing DNA sequence, chromatin accessibility, and gene expression across regulatory scales. Sequence-based models capture DNA-level regulatory grammar but typically lack cell-level resolution, whereas single-cell omics models capture cell-state semantics but treat genes and peaks as fixed vocabulary tokens, and thus lack a unified sequence-based view of regulation. Here, we present MUSS, a sequence-based multi-scale model for single-cell gene regulation that preserves both DNA-level and cell-level resolution. MUSS derives peak- and gene-level representations from DNA using a frozen pre-trained sequence encoder, integrates them with a multi-scale encoder combining peak-peak self-attention and distance-biased peak-gene interaction, and trains on paired scATAC-seq and scRNA-seq to align sequence-derived representations with cell-specific regulatory states. Across gene expression prediction, enhancer-gene linking, TF-target gene recovery, and zero-shot cell clustering, MUSS outperforms both sequence-based and single-cell omics baselines, and further generalizes to genes and peaks unseen during training. Together, these results demonstrate that MUSS learns biologically meaningful and broadly generalizable regulatory representations for decoding cell-type-specific regulatory programs from DNA sequence.
PaperID: 5181, Poster
Authors:
Mengtan Zhang, Yuchuan Cui, Yu Ma, Zizhan Guo, Wei Ye, Rui FanAbstract: Self-supervised monocular depth estimation is appealing for its potential to scale with unlabeled data, yet existing methods often degrade when trained on diverse scenes. This study finds that, under self-supervision, models tend to encode dataset-specific structural biases rather than transferable scene geometry, leading to cross-domain optimization conflicts. To address this issue, this study proposes the MaSS framework, aiming to make self-supervised depth learning scalable to diverse data. MaSS introduces a dedicated structural representation pathway explicitly decoupled from contextual features, which is guided by a novel scene-structure discrimination objective to learn generalized, layout-aware structural representations for depth decoding. Trained and evaluated on four diverse datasets, both individually and jointly, MaSS substantially outperforms prior self-supervised methods and is the first to consistently benefit from mixed-domain training. Zero-shot evaluation on eight additional datasets further demonstrates its remarkable generalizability, particularly on those entirely out-of-distribution scenes. The source code will be publicly available upon publication.
PaperID: 5182, Poster
Abstract: Logical reasoning remains a major challenge for large language models (LLMs), particularly on structured problems that require precise constraint tracking, consistency preservation, and multi-step deduction. This challenge is especially acute for small-scale LLMs, which are more prone to producing inconsistent, redundant, or brittle reasoning trajectories. Existing approaches for improving logical reasoning largely optimize for final-answer correctness, providing only weak supervision over the intermediate reasoning process. In this work, we propose SPRING, a solver-guided reinforcement learning framework for logical reasoning that uses an SMT solver as a training-time verifier of intermediate reasoning steps to provide process-level supervision. SPRING introduces the notion of a novel reasoning step, namely, a step that is logically valid, consistent with the evolving reasoning state, and not already implied by previously accepted non-contradictory deductions. Based on this solver-based assessment, we design process rewards that encourage novel inferential progress while penalizing contradictory and uninformative reasoning steps. Evaluation results on two logical reasoning benchmarks, ZebraLogic and AR-LSAT, show that SPRING consistently outperforms baseline LLMs and outcome-only reward baselines. On ZebraLogic, SPRING improves puzzle accuracy by up to 49.71 and 15.43 points over the base LLM and outcome-only reward baseline, respectively. On AR-LSAT, SPRING improves the overall average score by up to 64.93 and 12.14 points over the base LLM and best-performing outcome-only reward baseline, respectively.
PaperID: 5183, Poster
Authors: Donghee Han, Jiwon Jeong, Mun Yi
Abstract: Medical visual question answering is a safety-critical decision problem in which abstaining on inputs the image cannot support matters as much as answer accuracy. Existing inference-time agentic methods address hallucination by adding a verification step, but they collapse the final decision at a single language-model judge that integrates all evidence in one prompt and is systematically biased toward declaring inputs answerable. We propose PICV (Parallel Independent Claim Verification), which reformulates selective medical VQA as independent claim verification with transparent aggregation: the question is decomposed into a small set of typed visual claims, each is verified by an isolated prompt with no access to the other verdicts, and the verdicts are combined by a deterministic rule rather than another model, yielding an answer/abstain decision together with a typed unanswerability reason. We evaluate PICV efficiently on an automated answerability benchmark we construct over two public radiology datasets via three reproducible corruption procedures, requiring neither heavy VLM judging nor human annotation at evaluation time. Across four VLM backbones, PICV consistently outperforms strong prompting and agentic baselines; ablations attribute the gains specifically to per-claim isolation and rule-based aggregation rather than to agentic decomposition alone.
PaperID: 5184, Poster
Authors:
Maryem Benslimane, Hamza Abdelhedi, Vanessa Hadid, Karim JerbiAbstract: Facial expressions are one of the main ways humans communicate affect, but they are not static signals. In natural vision, expressions evolve over time, with facial movements emerging, intensifying, and changing as emotion is expressed. Prior model-brain studies based on static images suggest that vision models can capture aspects of face-related neural representations, but it remains unclear whether artificial vision models capture the temporal EEG dynamics of facial emotion perception. We test the hypothesis that models trained for facial emotion recognition, especially video-based or temporally structured models, should show stronger alignment with human EEG dynamics than randomly initialized controls. We do so by aligning model representations with time-resolved EEG representational geometry from 25 participants viewing one-second dynamic facial-expression videos. We compare six models spanning static-image, video-based, self-supervised, and vision-language representations: ResNet18, VGGFace, CNN2D, CNN3D, DINOv2-temporal, and Qwen2.5-VL. Contrary to this hypothesis, supervised facial-emotion training does not consistently increase alignment with human EEG dynamics. In particular, video-trained and temporally structured emotion models align less with fear-related EEG geometry than the same architectures with randomly initialized weights, whereas static-image emotion models show trained-versus-random differences close to zero. These results are robust to stricter noise-ceiling thresholds and are not explained by actor identity or raw pixel similarity. Exploratory analyses further suggest that peak alignment is often localized to the first or last sampled video frames, and that fear alignment varies across scalp ROIs by training regime. Overall, our results challenge the assumption that video-based emotion supervision automatically yields more brain-like dynamic representations of facial emotion.
PaperID: 5185, Poster
Abstract: Offline safe multi-agent reinforcement learning (MARL) is challenging: agents must learn coordinated behaviors from static datasets while strictly satisfying safety constraints. This challenge is amplified by distribution shift, where joint actions that appear safe under the behavior distribution may become unsafe under policy-induced deviations, with errors compounding across agents. We propose SamaDICE, a principled framework for safe offline MARL that integrates stationary distribution correction with scalable value decomposition. Our approach formulates safe policy learning as a constrained convex optimization problem over stationary distributions of joint state-action pairs, with safety constraints imposed as upper bounds on expected cumulative costs. Safety is enforced through a Lagrangian dual formulation, enabling distribution-corrected cost estimation directly from offline data. To address the combinatorial complexity of multi-agent systems, we introduce a factored centralized-training-decentralized-execution parameterization that represents the global dual variable via a monotonic mixing network over agent-local potentials, yielding a tractable surrogate Bellman residual for scalable optimization. We instantiate SamaDICE with a multi-agent decision transformer and evaluate it on diverse benchmark environments. Our method outperforms the strongest baseline in aggregate return, with the advantage widening as the cost budget tightens, while maintaining 100% empirical episode-level constraint satisfaction across all evaluated tasks.
PaperID: 5186, Poster
Abstract: While bound propagation methods are highly effective for certifying neural network robustness against \ell_\infty adversaries, scalable verification for \ell_2 perturbations remains a significant open challenge. Existing approaches suffer from severe geometric loss: elementwise bounding methods like \alpha-CROWN relax Euclidean balls to bounding boxes, losing a factor of \sqrtn in radius, while the recently proposed SDP-CROWN rounds deformed ellipsoids back into balls, losing a factor proportional to the condition number of the weights. Consequently, tight \ell_2 verification is currently restricted to near-isometric or Lipschitz-regularized networks where these approximations are not vacuous. In this paper, we introduce affine-image propagation, a novel CROWN-compatible framework that exactly propagates the affine geometry of \ell_2 balls and \ell_\infty boxes through the network. To bypass the prohibitive computational cost of exact Semidefinite Programming (SDP) constraints required by this representation, we derive highly scalable, SDP-based bounds utilizing diagonal-dominance surrogates. Our approach effectively eliminates the geometric bottlenecks of prior work, yielding the first large-scale, SDP-style bound propagation method capable of tight \ell_2 verification for general, non-Lipschitz-regularized neural networks.
PaperID: 5187, Poster
Abstract: Assessing the quality of time series (TS) data is fundamental yet inherently challenging due to the multifaceted nature of quality dimensions. Recently, large language models (LLMs) have emerged as a promising paradigm for TS quality assessment via pairwise comparison and per-dimension evaluation. However, existing approaches rely on manually predefined quality dimensions and purely text-based reasoning, leaving it unknown whether LLMs can identify truly relevant quality dimensions or perform grounded and quantitative quality comparisons. To investigate this, we construct TSQBench, a dedicated benchmark for evaluating LLMs’ capabilities in TS quality assessment, which measures two progressive capabilities: (i) understanding and identifying relevant quality dimensions, and (ii) performing quality comparison under specific dimensions. Our analysis reveals that both open-source and proprietary LLMs consistently struggle with identifying critical quality dimensions and conducting precise, evidence-grounded reasoning over quality characteristics. To address these limitations, we propose TSQAgent, a novel agentic reasoning framework for dedicated TS data quality rating, comprising three collaborative roles: Perceiver for focused dimension selection, Inspector for dimension-wise quantitative analysis, and Adjudicator that aggregates and refines the final judgment. In particular, we introduce an agentic reasoning strategy that instills the ability to identify and prioritize the most relevant quality dimensions, and further propose an agent workflow equipped with external analytical tools to enable precise quantitative comparisons over selected dimensions. Experiments on both the proposed benchmark and eleven real-world datasets demonstrate that our framework not only substantially improves LLMs’ capabilities in quality understanding and quantitative comparison but also effectively translates these improvements into better quality-aware data selection, leading to enhanced downstream performance and data efficiency.
PaperID: 5188, Poster
Authors:
Gu Gong, Yining Wei, Yuechen Tao, Tianyuan Wu, Peijie Dong, Ruibo FAN, Wenhu Hu, Yinghao Yu, Jiamang Wang, Wenbo Su, Guodong Yang, Liping Zhang, Wei Wang, Xiaowen ChuAbstract: Reinforcement learning (RL) for large language models (LLMs) is increasingly bottlenecked by rollout cost, making low-precision rollout appealing for acceleration. However, applying low-precision formats such as FP4 to RL remains unstable, even with quantization-aware training (QAT) enhancements. We identify that the root cause is that FP4’s aggressive quantization amplifies small system-level numerical differences (i.e., training-inference stack mismatch) that are harmless at BF16 into large token-level log-probability errors, distorting importance weights and destabilizing optimization. To address this, we propose Cross-Precision Alignment (CPA), a simple regularizer that stabilizes low-precision RL by maintaining a BF16 master policy, executing rollouts in real FP4, and aligning the fake-quantized forward pass with a BF16 reference forward pass on sampled tokens by applying a low-variance KL penalty at each training step. In our evaluation, the resulting BF16 checkpoint matches a full BF16 RL baseline, and the overall pipeline delivers 8–19% end-to-end training speedup.
PaperID: 5189, Poster
Authors: Peter J Kampen, Anders N Christensen, Morten R Hannemose, Anders Dahl, Josefine Vilsbøll Sundgaard
Abstract: Fine-grained classification in deep taxonomies suffers from representational crowding: large macro-classes dominate the feature space, leaving little angular capacity for fine-grained distinctions, particularly under out-of-distribution (OOD) shifts. Standard hyperbolic spaces offer exponential volume for tree-like data, but embedding an entire taxonomy into a single manifold forces distinct macro-branches to compete for shared angular capacity and entangles their optimization updates. We introduce a geometric architecture that embeds hierarchical taxonomies into a Cartesian product of Lorentz manifolds. Our core mechanism, Orthogonal Origin Parking (OOPark), decouples the taxonomy at the root level by assigning an independent Lorentz manifold to each macro-branch and penalizing inactive embeddings that deviate from the manifold origin. We evaluate across four biological OOD tasks/benchmarks (skin lesions, iWildCam, fungi, and plankton) characterised by hierarchical class structure, class imbalance, and domain shift. OOPark generally outperforms standard Euclidean and single-manifold hyperbolic baselines while using 20-dimensional sub-manifolds against Euclidean baselines of up to 512 dimensions, preserving both fine-grained accuracy and global taxonomic fidelity under severe distribution shifts. Code is provided in the supplementary material; a public release will follow upon acceptance.
PaperID: 5190, Poster
Abstract: Linear State Space Models (SSMs) and autoregressive sequence models reuse the same parameters across a sample-dependent number of steps. As a result, different examples impose constraints through different predictors, even though all predictors share the same weights. We study this phenomenon through a general model of multiple non-homogeneous polynomial predictors with shared parameters. Under standard interpolation and directional-convergence assumptions, together with a trajectory-dependent effective-degree condition, we show that gradient descent on the exponential loss converges in direction to a KKT point of an \empheffective-degree max-margin problem. In this problem, only examples with minimal effective degree impose hard margin constraints; examples with faster-growing margins are asymptotically inactive except for feasibility. Our result provides a characterization of the implicit bias in linear SSMs and linear autoregressive models. It exposes a mechanism of length bias: a longer input sequence or additional autoregressive steps may increase the example's effective degree, and the examples with the smallest effective degree dominate the limiting classifier. Our result highlights how non-homogeneity and parameter sharing alter classical homogeneous implicit-bias results.
PaperID: 5191, Poster
Abstract: Classification with rejection (CwR) generalizes the standard classification task by providing an extra option of rejection with a cost c smaller than misclassification costs, and the design of its convex calibrated surrogate loss is our central topic. While typical surrogate losses for K-class classification operates on K-dimensional scoring functions, it has been proved that this dimensionality can be reduced to \lceil\log_2K\rceil for CwR with c\in(0,0.5], rendering more efficient computation. This logarithmic dependency is commonly conjectured to be the lowest possible among any convex calibrated surrogate losses for CwR. Through the lens of property elicitation, we explore even lower-dimensional structures for CwR, by constructing a 2-dimensional convex calibrated surrogate loss under c\in(0,0.5). This loss dimension is provably optimal among all convex calibrated surrogate losses. The boundary cost c=0.5 is more nuanced---we construct a 4-dimensional convex calibrated loss, improving over the \lceil\log_2 K\rceil-dimension when K>16. Notably, this breaks the optimal logarithmic dimensionality attainable by the embedding framework [Finocchiaro et al., 2024], a widely used loss design principle for polyhedral losses, negatively resolving their conjecture on the dimension optimality.
PaperID: 5192, Poster
Authors: Ahatsham Hayat, Mohammad Hasan
Abstract: Cross-distribution generalization in longitudinal behavioral data remains a persistent obstacle. Models trained on one cohort degrade on new populations, and prior approaches operating through feature-distribution alignment or model-capacity scaling have not substantially improved cross-distribution performance. We diagnose the limitation as insufficient representational constraints during training. We propose holding the language model frozen, so that the coherence criterion is defined by its pre-trained language prior and cannot shift during training. We introduce PRISM, a frozen-backbone framework that decomposes each behavioral trajectory into temporal, spectral, and semantic streams and integrates them through directed cross-attention with instance-specific gating. Training uses a dual-path objective combining discriminative and coherence constraints on a shared representation. On GLOBEM, a widely used cross-cohort benchmark for behavioral health prediction, PRISM achieves 79.93% out-of-distribution accuracy, 27.13 points above the previous best, with an in-distribution to out-of-distribution accuracy gap of 1.24 points. PRISM also achieves the highest OOD accuracy on three additional datasets (LifeSnaps anxiety, MFAFY engagement, CrossCheck schizophrenia symptoms), with consistently smaller ID-to-OOD gaps than fine-tuned VLM baselines. Ablations identify the frozen coherence constraint as the component responsible for the distribution-invariance pattern.
PaperID: 5193, Poster
Abstract: Gradient Boosted Decision Trees (GBDTs) remain a leading approach for tabular learning, offering strong predictive performance, computational efficiency, and robustness to heterogeneous feature types. Yet their standard second-order leaf updates suffer from a node-level statistical limitation: as trees grow deeper, sparse leaves provide increasingly unreliable gradient estimates, while global regularization cannot adapt to local uncertainty or residual structure. We propose RAGBoost, a retrieval-augmented and ancillary-guided boosting framework for robust tabular learning. RAGBoost augments standard GBDT optimization with a two-layer node-adaptive correction mechanism. First, it retrieves historically observed leaves with similar gradient-state configurations to form node-conditional gradient priors, stabilizing variance-dominated sparse-leaf updates through empirical shrinkage. Second, it uses ancillary residual correction to detect and mitigate bias-dominated within-leaf residual structure that remains unresolved by the main tree. These corrections modify only leaf values, preserving the original tree structure and inference pipeline. Empirically, RAGBoost achieves the best overall average rank across 40 heterogeneous tabular benchmarks in our evaluation, while remaining strongly competitive with recent deep tabular and tabular foundation model baselines. On TabReD, an industrial-scale temporal-shift benchmark, RAGBoost outperforms XGBoost and the base TabM variant overall, suggesting improved robustness under realistic distribution drift.
PaperID: 5194, Poster
Abstract: Open-source web agents are now strong enough to execute realistic browser tasks, but training them with reinforcement learning still depends on weak supervision: binary task success is too sparse for credit assignment, while frontier-language-model judges are too expensive to call at every step and cannot be assumed available at deployment. We introduce CLIFT, a training and test-time scaling method built around conformal self-verification. During training, the agent answers natural-language verification questions about its own rollouts; a Compositional Conformal Certifier keeps only question signals whose URL-conditional evidence agrees with a training-time judge, assigns signed trust weights through polarity-aware lift, and blends the resulting verifier score into per-step rewards in a way that never subtracts from the judge baseline. At test time, the same certified bank is frozen and reused as structured evidence for Conformal Trajectory Selection (CTS): the agent samples a greedy rollout and one or more diverse retries, the self-verifier summarises each URL trace, and a conservative majority-vote rule chooses whether to swap away from the current incumbent without calling any external judge. This single mechanism supports three settings. On WebArena Infinity, CLIFT achieves state-of-the-art performance among open-source web agents. On VisualWebArena, a bank trained with the open model transfers to GPT-5.5 at test time and reaches state-of-the-art performance under the canonical harness. On Online Mind2Web, without training an agent on the benchmark, translating the certified question bank improves a live-web agent in zero-shot evaluation. Together these results position conformal self-verification as a way to turn costly judge feedback into a reusable training signal and a judge-free test-time scaling signal.
Abstract: Supervised fine-tuning (SFT) on long teacher trajectories is the dominant method for instilling investigation and reasoning capabilities into open software-engineering (SWE) agents. Under SFT, every retained response is an imitation target, so the student inherits not only the trajectory's outcome but also any flaw in its intermediate steps, including ungrounded leaps and redundant loops. High-quality training data must therefore be jointly effective (each step is grounded and narrows the agent's epistemic gap to the correct fix) and efficient (each step is information-bearing rather than redundant or looping). Existing recipes filter or relabel teacher rollouts using only a binary terminal verifier, which does not directly target these axes and provides no supervision on instances where the teacher fails. Every real issue ships with a developer-authored reference patch p^\star that implicitly testifies to the file paths, runtime behaviors, and conventions a fix presupposes, but the standard pipeline discards it. We propose P2T (Patches-to-Trajectories), which uses p^\star as privileged information during curation, and frames trajectory construction as a bi-objective program over per-step effectiveness and trajectory length. A reverse phase distills p^\star into a latent process graph G^\star of contextual facts and solution milestones, encoding dense intermediate anchors in constructive ordering. A forward phase curate trajectories from blinded teacher continuations, scoring per-step progress against G^\star under a leakage-blocking groundedness check and committing the shortest segments that retain effectiveness. Using only 1.8k curated SWE-Gym instances, P2T improves both axes simultaneously over outcome-filtered SFT and its tool-error-masking variant: on SWE-bench Verified, it lifts Pass@1 by up to +10.8 points while cutting per-instance inference cost by 15%, with consistent gains on SWE-bench Lite and across two teachers. A size-matched ablation and qualitative analysis further isolate per-trajectory quality from data scale.
PaperID: 5196, Poster
Abstract: Protein mutation effect generation asks a model to describe the functional consequence of a point mutation in natural language. Existing protein-to-text systems typically encode mutation information into undifferentiated representations, overlooking the organization of mutation-induced evidence across structural and biochemical factors. We propose RipplePLM, a mutation-aware generation framework centered on Direct-Distal Cross-Attention (DDCA). By constructing a residue-level Mutation Perturbation Field from pre-trained protein language models, DDCA leverages predicted contact maps to explicitly decouple structural perturbations into two pathways: the mutation site's immediate contact neighborhood and its multi-hop distal context. To complement this structural decomposition, we further introduce the Property Latent Chain (PLChain), which injects expert-guided biochemical supervision (e.g., thermostability and optimal pH) into the LLM hidden-state pathway through latent property tokens. On MutaDescribe, RipplePLM improves over mutation-specific baselines on temporal and structural splits; under a matched-backbone comparison, average structural-split ROUGE-L increases from 22.23 to 35.65. Additional ablations, representation diagnostics, and low-N fitness regression experiments further support the effectiveness of the learned mutation-aware representations.
Abstract: Vision-Language-Action (VLA) models enable robots to follow natural language instructions and generalize across diverse tasks, but they remain vulnerable to execution failures that compromise reliability in real-world deployment. Detecting such failures during execution is therefore critical for the robust deployment of embodied systems. Existing failure detection methods either rely on expensive action resampling or external models, while alternatives propagate trajectory-level labels uniformly across every timestep, obscuring localized failure signals. In this paper, we propose Hide-and-Seek, a framework that formulates VLA failure detection as a coarsely supervised learning problem. By combining inter-trajectory and intra-trajectory contrastive objectives, Hide-and-Seek localizes failure-indicative actions and induces temporally structured failure signals from trajectory-level supervision alone, without any step-level annotation. We evaluate Hide-and-Seek on LIBERO, VLABench, and a real-world robotic platform across three representative VLA policies: OpenVLA, \pi_0, and \pi_0.5. Our method achieves state-of-the-art multi-task failure detection performance with a practical accuracy--timeliness trade-off under conformal prediction, and generalizes well to both seen and unseen tasks.
Abstract: Training binary neural networks (BNNs) from scratch is dominated by the straight-through estimator (STE), whose forward/backward mismatch produces severe accuracy degradation as networks deepen. We study an orthogonal axis: when and where binarization is enforced during training. We introduce StoMPP (Stochastic Masked Partial Progressive Binarization), which gradually replaces clipped weights and activations with their hard binary counterparts layer by layer from input to output, using stochastic partial masks with soft refresh. StoMPP delivers two complementary benefits. As a standalone training rule, it provides a fully STE-free procedure that improves over vanilla STE with gains that grow with depth (ResNet-50 BNN: +18.0/+13.5/+3.8 on CIFAR-10/100/ImageNet), and the pattern holds across ResNet-18/34/50, MobileNetV2, and BERT fine-tuning. Composed with surrogate gradients by applying STE only to frozen entries, it reaches +27.1/+19.8/+17.7 over vanilla STE on the same setting. Underlying both regimes is a single mechanistic finding: progression order is decisive. Forward layerwise progression prevents depth collapse, reverse progression collapses to near-chance, and binary-weight networks (without binary activations) are insensitive to order. We trace this asymmetry to activation-induced gradient blockades: a committed binary activation severs gradient flow upstream, and ordering controls when these blockades form. To isolate the progression's contribution from any benefit conferred by STE, we conduct all ablations in the STE-free regime; the resulting characterization (schedule, refresh, ordering, dynamics) thus reflects the progression itself rather than its interaction with surrogate gradients.
Abstract: Edge of Stability (EoS) refers to the phenomenon where full-batch gradient descent (GD) training of neural networks with step size \eta pushes the largest eigenvalue of the Hessian, i.e. the sharpness, to 2/\eta and hovers there. \citetdamian2023selfstab explains the hovering behavior of the sharpness by self-stabilization, a mechanism driven by third-order structure of the loss, and shows that GD implicitly follows projected gradient descent (PGD) on the set where the sharpness is below 2/\eta. For mini-batch stochastic gradient descent (SGD), the sharpness stabilizes below 2/\eta, with the gap widening as the batch size decreases. However, no theoretical explanation exists for this suppression. In this paper, we introduce stochastic self-stabilization that extends the self-stabilization framework to SGD. Our key insight is that gradient noise injects variance into the oscillatory dynamics along the top Hessian eigenvector, strengthening the sharpness-reducing force and shifting the equilibrium below 2/\eta. Following the approach of \citetdamian2023selfstab, we define stochastic predicted dynamics that tracks the deviation of SGD from the PGD trajectory, and prove a stochastic coupling theorem that relates the SGD sharpness and loss to those of PGD. Based on our predicted dynamic, we derive a closed-form equilibrium sharpness gap that scales with the variance of the gradient noise projected onto the top eigenvector of the Hessian. This formula predicts that smaller batch sizes yield flatter solutions, and recovers GD when the batch equals the full dataset.
PaperID: 5200, Poster
Abstract: Adaptive Random Forests (ARF) are among the most effective ensemble methods for learning from non-stationary data streams, yet their success relies on heuristic drift detectors and reset rules that lack formal guarantees and require careful tuning. We revisit ARF from an online learning perspective and show that its design decomposes into two fundamental questions: how to aggregate a dynamically evolving set of trees, and how to update the pool of trees over time. We address aggregation using parameter-free online learning methods with strongly adaptive regret, enabling the ensemble to track the best tree mixture. For pool updates, we replace detector-driven resets with deterministic multi-scale scheduling based on geometric lifetimes, combined with incumbent--challenger replacement governed by anytime-valid statistical tests. This design yields theoretical guarantees: global control of false replacements, bounds on promotion delay under post-change advantage, and adaptivity under piece-wise stationary environment. Empirically, the resulting Online RF consistently outperforms ARF and is competitive with broader streaming baselines across several classification and regression benchmarks.
Abstract: Recent diffusion-based video generators have achieved remarkable visual fidelity and prompt controllability, yet scaling them to ultra-high-resolution (UHR) long videos remains prohibitively expensive. The difficulty is especially pronounced for long single-shot generation where a continuous scene must preserve global temporal coherence, and fine-grained spatial details without relying on clip transitions or autoregressive shot stitching. In this work, we revisit this challenge from the perspective of decoupled modeling. We argue that existing video diffusion models already encode strong local visual priors, while the main bottleneck lies in efficiently extending global spatiotemporal modeling as resolution and duration increase. Based on this insight, we propose AtlasVid, a decoupled global-local framework for efficient UHR long video generation. AtlasVid first generates a low-resolution and low-FPS global semantic proxy via temporally scaled RoPE, thereby extending the temporal horizon without increasing the training token count. Guided by this proxy, a high-resolution detail branch performs joint denoising with hierarchical locality-preserving attention. Reordered spatiotemporal windows preserve geometric locality and asymmetric global-local attention injects aligned semantic guidance and preserves the model's pretrained ability. This design enables resolution-agnostic training: the model is trained only at 720P with lightweight LoRA adaptation, yet generalizes directly to 4K and beyond for longer (>10s) video synthesis. Experiments show that AtlasVid substantially improves the efficiency of ultra-high-resolution long video generation, achieving high-quality UHR long video generation with 60.9× speed up and significantly less training cost and even better performance than native 4K video generators.
PaperID: 5202, Poster
Abstract: Probabilistic neural surrogates for PDEs enable ensemble forecasting, but often exhibit miscalibration when used for uncertainty quantification, especially in compute-constrained, few-step inference regimes. Existing generative approaches typically rely on fixed global stochasticity, which cannot adapt predictive uncertainty to the local dynamical state. We propose TORNADO, a latent probabilistic forecasting framework that formulates next-step prediction as stochastic transport between VAE encoder posteriors. Unlike existing approaches with globally prescribed stochasticity, TORNADO learns a SDE with state-dependent diffusion, enabling the predictive uncertainty to adapt to the local latent dynamics. The model is trained using analytic conditional endpoint targets derived from a Gaussian interpolant, enabling simulation-free learning. An additional energy distance regularization term improves distributional alignment. Empirically, TORNADO improves uncertainty calibration over representative generative baselines, including stochastic interpolants, diffusion models, and flow matching, while maintaining comparable predictive accuracy. Our results show that adapting stochasticity to the latent state, rather than prescribing it globally, is an effective mechanism for improving calibration in practical neural PDE forecasting settings.
PaperID: 5203, Poster
Abstract: Graphic designs, such as posters, advertisements, and infographics, are an important medium for communicating information and shaping understanding. Unlike natural images, they consist of layered elements with explicit compositional order. However, existing object detection models treat these elements as an unordered set, leaving compositional order unexploited. To address this limitation, we present esign (DAD), a model that formulates graphic design detection as compositional deconstruction. It decodes elements in compositional order, using lower-layer elements to better detect higher-layer ones. The key feature of DAD is amodal detection, which predicts the full bounding box of each element, including regions occluded by elements placed above it. Building on this formulation, we propose ptimization (EleRPO), which extends GRPO from sequence-level supervision to element-level optimization. EleRPO provides fine-grained training signals that capture how each detected element contributes to overall detection quality, and works synergistically with compositional order to improve detection performance. To support training and evaluation, we build a dataset of 10 million graphic designs. Experiments show that DAD outperforms all baselines and achieves human-level performance in amodal detection, supporting effective image-to-layer decomposition. EleRPO consistently improves over GRPO across nine detection benchmarks.
Abstract: The remarkable success of modern AI has been closely tied to scaling laws, yet the finite supply of high-quality data makes data efficiency—learning more from less—an increasingly important frontier. A model’s inductive bias is a critical lever for data efficiency, but foundational sequence models such as State Space Models (SSMs) often rely on fixed, task-agnostic biases. When this fixed prior is misaligned with the underlying structure of a task, the model may require additional samples to overcome its own bias before learning the relevant signal. In this work, we introduce a principled framework for understanding and aligning the inductive bias of linear time-invariant SSMs. We first formalize this bias through an SSM-induced kernel and show theoretically and empirically that its spectrum is governed by the model’s frequency response. This characterization motivates Task-Dependent Initialization (TDI), a fast power-spectrum matching method that aligns the initial SSM bias with the task’s spectral characteristics before downstream training. Across controlled synthetic experiments, trainable one-layer SSMs, and deep SSMs on diverse real-world benchmarks, TDI can improve data-efficient generalization primarily when task-relevant spectral structure is present and the default SSM bias is spectrally mismatched. Our results provide both a theoretical lens and a practical tool for task-adaptive inductive bias, suggesting a path toward more data-efficient sequence modeling.
PaperID: 5205, Poster
Abstract: Memory enables long-context dialogue agents to maintain user state across many turns, but it also introduces a fundamental source-choice problem: grounding the response in the immediate context or retrieved prior memory? A wrong choice incurs two complementary costs: neglecting valid aged-out evidence, or overwriting sufficient context with stale memory. Existing systems and benchmarks largely conflate retrieval with source choice, relying on the model to resolve it implicitly and leaving this critical decision largely unexplored. In this paper, we introduce MSBench, a controlled benchmark that isolates source choice by pairing questions answerable exclusively from context with those requiring prior memory. Under this protocol, strategies that depend on the model's implicit judgment exhibit sharply degraded performance, exposing the limits of leaving source choice ungoverned. To this end, we propose the Memory-Source Reasoner (MSR), which frames source choice as an explicit metacognitive decision: it reasons about contextual sufficiency, answers directly when appropriate, and selectively retrieves missing evidence otherwise. Experimental results show that MSR achieves the best overall accuracy and selection accuracy across four answer backbones. On the primary GPT-5-mini run, it improves over the best non-MSR baselines by 12.6 points in overall accuracy and 6.3 points in selection accuracy. Our benchmark, code, prompts, and evaluation scripts are available at https://anonymous.4open.science/r/msbench-msr-0C72.
PaperID: 5206, Poster
Abstract: Self-supervised representation learning has advanced time series forecasting by capturing robust features from unlabeled data. However, existing contrastive and masking-based methods often struggle to explicitly model the underlying temporal flow, either by treating time lag as noise or by disrupting the inherent temporal dynamics essential for forecasting. To address these limitations, we propose OliO (ODE-based Linear Transition Operator), a plug-and-play self-supervised learning method that learns bidirectional temporal flows as a continuous and structurally consistent dynamical system. OliO introduces a transition operator derived from linear Ordinary Differential Equations (ODEs) to align latent representations across arbitrary time shifts. By imposing a strictly upper triangular constraint on the transition matrix, we mathematically ensure numerical stability and enforce a robust polynomial boundary that effectively captures long-term dependencies while preventing the exponential divergence typically found in unconstrained ODE-based models. This structural constraint induces a temporally coherent representation space that preserves the underlying flow essential for accurate forecasting. Extensive experiments across various backbone architectures and benchmarks demonstrate that OliO achieves significant performance gains, providing Mean Squared Error (MSE) reductions of up to 5.75% over the most competitive grid-searched SOTA baselines. The code is available at https://anonymous.4open.science/r/OliO-46EF.
PaperID: 5207, Poster
Authors: Vladimir Antipov, Artem Badarin, Alexander Hramov
Abstract: We introduce ACMA (Delay-Augmented Clustering and Model Approximation), a noise-robust pipeline for saccade detection in oculographic signals that couples delay-embedded clustering with physiologically grounded parametric refinement. Existing detectors face a recurring trade-off: velocity-thresholding methods are simple and interpretable but degrade catastrophically as recording noise grows, adaptive event-based detectors recover part of this gap at moderate noise but lose reliability at high noise, and recent deep models are sensitive to noise distributions unseen during training. ACMA addresses this trade-off in two stages. First, a sliding-window two-cluster decomposition with time-delay embedding identifies candidate saccades while remaining stable across noise regimes. Second, a parametric saccade waveform constrained by priors on amplitude, duration, and the saccadic main sequence validates and characterizes each event. We evaluate ACMA on a saccade-detection benchmark spanning a wide range of noise levels and nine representative baselines, including a leading deep-learning detector (U'n'Eye), adaptive methods (NH, REMoDNaV, Engbert), clustering-based I2MC, and classical velocity-thresholding methods (IVT, IVVT, IDT, IDVT). ACMA delivers consistent F1 gains in moderate-to-high noise regimes, where competing detectors degrade sharply, while additionally returning physically meaningful per-event metrics (amplitude, duration, peak velocity) for downstream neurological assessment, fatigue monitoring, and brain-computer interface applications.
PaperID: 5208, Poster
Abstract: Generative models, such as diffusion and flow-based models, have emerged as powerful policy approximators in deep reinforcement learning (RL) due to their ability to represent multi-modal action distributions. However, their potential for maximum entropy RL remains largely untapped, as most existing methods treat generative processes as black-box samplers or rely on heuristic noise. In this paper, we propose GENIEO, a novel off-policy generative RL framework that leverages the exact intrinsic entropy of an augmented dummy-action policy to drive principled exploration. By designing a structurally invertible affine dual-variable flow, we bypass the numerical approximations of standard ODE solvers. This innovation allows us to rigorously apply the change-of-variables formula to derive a tractable and exact formulation for the augmented dummy-action policy entropy, enabling its direct integration into the off-policy learning objective. Experimental evaluations across 14 challenging tasks from the DMControl and HumanoidBench suites demonstrate that GENIEO achieves state-of-the-art (SOTA) performance. By effectively discovering optimal strategies in complex, high-dimensional environments, our approach bridges the gap between expressive generative modeling and maximum entropy RL.
PaperID: 5209, Poster
Abstract: Spiking Neural Networks (SNNs) provide energy-efficient computation by utilizing binary spike-driven additions, but applying transformer architectures remains challenging due to the lack of an SNN-compatible positional encoding (PE). Current PE methods either rely on floating-point arithmetic, which destroys the fundamental binary nature of SNNs, or fail to accurately capture relative spatial distances. To overcome this, we introduce Bit Shift Rotary Position Embedding (BitShift-RoPE), the first relative PE framework that fully preserves the binary integrity of SNN queries and keys. By replacing the floating-point trigonometric rotations in standard RoPE with discrete cyclic-shift operations via hardware memory-pointer routing, our method achieves zero-FLOP position encoding. We analytically map multi-scale exponential decay frequencies into integer shift steps and theoretically extend the formulation to 2D orthogonal spaces to capture complex spatial geometries. Crucially, because the binary integrity of the queries and keys is strictly maintained, the subsequent attention calculations rely solely on sparse bitwise AND operations. Our evaluations across time-series forecasting (avg R^2 0.758), text classification (69.91% accuracy), and image classification benchmarks (84.17% accuracy) demonstrate that BitShift-RoPE achieves state-of-the-art performance while consuming the same energy as the vanilla SNN backbone (0.216 mJ/sample, identical to vanilla Spikformer). The code is available at \urlhttps://anonymous.4open.science/r/Bit\_Shift\_RoPE-16D2/.
PaperID: 5210, Poster
Abstract: Corpus reasoning tasks, which require extracting and integrating information from large collections of documents (e.g. scientific literature, LLM agent traces), range from well-studied tasks such as factoid question answering (e.g. tasks, efficient architectures can scale much worse compared than they do on more well-explored tasks, but why does this happen? We develop a unified view---corpus task complexity (CTC)---to systematically distinguish complex tasks from simpler ones in terms of the smaller operations needed to solve them (and how such operations scale with corpus size): for example, we classify tasks which compare all possible document pairs as than factoid retrieval. We first investigate new long-context language model (LCLM) training dynamics challenges on these tasks, and propose an effective new mask mixing technique that improves O(N) attention approaches on high CTC tasks, but at larger corpus scale we ultimately find that there is no free lunch: (i) methods that scale efficiently with corpus size . We identify high-CTC tasks as a long-term open problem and encourage more future work investigating new methods to overcome these limitations.
PaperID: 5211, Poster
Abstract: Risk assessment and anomaly detection in financial interaction networks often fail under regime shifts triggered by external events such as policy changes or enforcement actions—models may retain high AUC yet suffer sharp drops in average precision and calibration. Although news and regulatory filings contain actionable signals, naïvely injecting raw text or off-the-shelf LLM embeddings into temporal GNNs can introduce temporal leakage, misalignment, and unstable out-of-time ranking. We propose REGATE, a time-causal framework that integrates exogenous documents into dynamic graphs through three coupled components: (1) a schema-guided extractor that converts unstructured documents into auditable, time- and entity-aligned policy tokens with explicit confidence scores; (2) a bounded gating mechanism that fuses these tokens into temporal GNN states as a residual update, downweighting uncertain or stale evidence with a stability guarantee under direction consistency; and (3) a closed-loop retrieval adaptation module that distills the model's own routing attention into a lightweight document scorer without manual relevance labels. On established dynamic-graph benchmarks with documented regime shifts and six backbone architectures (TGN, TGAT, DyGFormer, GraphMixer, CTAN, GeneralDyG), REGATE yields up to +0.16 AP and up to 90% ECE reduction in post-shift windows on Elliptic; cross-domain pilots on credit and equity networks confirm consistent gains. Ablations show improvements stem from time-aligned semantic content rather than timestamps alone, and we report quality–cost trade-offs across open-source and commercial extractors.
PaperID: 5212, Poster
Abstract: Recent advances in sparse voxel-based VAEs have demonstrated remarkable capabilities in high-fidelity 3D autoencoding, yet their generative counterparts consistently lag behind their reconstruction performance. A fundamental cause of this gap lies in the \emphbrittle output representation in sparse voxel decoders, which parameterize geometry through discrete topological decisions, such as intersection flags or occupancy signs. Small perturbations in these parameters, which are unavoidable under diffusion sampling, can be amplified into abrupt topological ruptures and grid-like artifacts. We argue that a generation-friendly representation should ensure that such perturbations induce only smooth geometric transitions. To this end, we propose Pygmalion, a generation-friendly autoencoding framework that reformulates sparse voxel decoding as a coupled SDF parameterization. To retain explicit mesh-level supervision within this parameterization, which is non-trivial as the SDF jointly governs both topology and geometry, we introduce hinge-based sign correction, case-aware geometry supervision, and rendering-based refinement, achieving high-fidelity reconstruction while preserving generative robustness. Built upon this generation-friendly representation, Pygmalion scales consistently across DiT model sizes and voxel resolutions. Extensive experiments demonstrate state-of-the-art reconstruction fidelity and high-quality 3D generation with complete surfaces, sharp features, and fine geometric details.
Abstract: Physical fields on meshes require a separation between topology and geometry: conservation laws are topological and should be exact, while geometry, material response, and anisotropic coupling must be learned from data. Existing neural surrogates often mix these roles inside unconstrained message passing. We introduce Riemannian Hodge Message Passing (RHMP), which turns this separation into an architectural principle. RHMP fixes the cellular coboundaries (d_k) determined by oriented incidence and learns symmetric positive-definite cochain metrics (H_k) for geometry-dependent propagation. Treating H_k as the learned metric motivates cochain-frame equivariance: physical propagation should be invariant to orthogonal changes of the hidden cochain feature basis. RHMP implements this principle with metric-weighted Hodge blocks (d_k^\top H_k+1d_k), yielding exact cochain-complex identities (d_k+1d_k=0), nonnegative Hodge energies, positive-semidefinite operators, and exact Abelian curvature invariance. Across seven physical benchmarks spanning fluids, electromagnetism, gauge fields, and variable-mesh CFD, RHMP achieves the best overall performance, with the largest gains when topology, learned geometry, and field structure interact.
Abstract: Obtaining meaningfully diverse high-quality samples from Large Language Models (LLMs) for a fixed prompt remains an open challenge. Current methods often only operate at the token-level, paraphrasing the same response. This is problematic as it leads to poor exploration on reasoning problems and to unengaging, repetitive conversational agents. To address this, we propose Intent Factored Generation (IFG), factorising the sampling process into two stages. First, we sample a semantically dense intent that anchors the sample, e.g., a summary or keywords. Second, we sample the final response conditioning on both the original prompt and the intent from the first stage. This factorisation allows the use a higher temperature during the intent step to promote conceptual diversity, and a lower temperature during the final generation to ensure the outputs are coherent. We find that prompting the model to explicitly state its intent for each step of the chain-of-thought before generating the step is beneficial for reasoning tasks. We show that our simple method is highly effective across a diverse set of tasks. We find that this method improves both exploration and final performance on math and code tasks, and combines well with Reinforcement Learning from Verifier Feedback (RLVF). We also show that our method can be used to increase the diversity of instruction-tuned models. Finally, we demonstrate that our method leads to higher quality diverse samples on a language modelling task, on a new dataset that we open-source. IFG is easy to implement, and can be used with off-the-shelf LLMs to improve the performance-diversity tradeoff. Full Implementation at https://anonymous.4open.science/r/IFG-Anon/
PaperID: 5215, Poster
Authors: An Yan, Huigen Ye, Hua Xu, Jiahao Zhang
Abstract: Solving large-scale Mixed-Integer Linear Programming (MILP) problems involving millions of variables is critical for industrial applications but notoriously intractable due to combinatorial explosion. While Large Neighborhood Search (LNS) has emerged as the premier heuristic strategy for tackling such scales, its efficacy hinges entirely on the underlying neighborhood selection mechanism. Recent Large Language Model (LLM)-driven LNS frameworks show remarkable promise but suffer from fundamental flaws: they lack generalizability across diverse problem classes and rely on static selection paradigms that remain entirely blind to the shifting dynamics of the iterative optimization process. To overcome this rigidity, we propose PALL, a Process-Aware Evolutionary LNS framework guided by a Context-Enhanced Fine-Tuned LLM-driven Selector. Instead of deploying a monolithic operator, PALL leverages an offline-evolved diverse operator ensemble and dynamically synchronizes the search strategy with evolving optimization states. Specifically, we fine-tune an 8B-parameter LLM on high-quality hindsight oracle trajectories to adaptively perform optimal operator selection. To bolster the LLM's sequential decision-making, we introduce a context-enhanced reasoning mechanism that leverages historical operator trajectories and optimization rewards as inferential feedback. Extensive evaluations on standard million-variable MILP benchmarks demonstrate that PALL achieves state-of-the-art performance. Supported by an asynchronous deployment strategy, it not only consistently outperforms commercial solvers like Gurobi and classical LNS algorithms but also delivers over a 6× acceleration compared to state-of-the-art LLM-driven LNS frameworks.
Authors:
Akihiro Kubo, Kosuke Nakanishi, Shin IshiiAbstract: Preference-conditioned multi-objective reinforcement learning aims to learn a single policy that captures trade-offs across preferences, but under nonlinear scalarization the uniqueness and continuity of the preference-to-solution correspondence remain unclear. We study this problem in tabular multi-objective Markov decision processes (MDPs) using smooth Tchebycheff scalarization as a monotone utility. Under mild interior conditions on the preference set, we prove that each preference induces a unique Pareto-optimal return vector and that this vector depends Lipschitz-continuously on the preference, providing a principled foundation for preference sweeping toward dense Pareto-front coverage. To compute these targets, we formulate the problem over occupancy measures and derive Concave Mirror Descent Policy Iteration (CMDPI), which achieves an O(1/k) objective-suboptimality rate. We further show that each update is equivalent to solving a Kullback-Leibler-regularized MDP with the previous policy as reference, yielding a policy-iteration interpretation and finite-iterate policy continuity across preferences. We instantiate the update as a deep actor-critic algorithm preserving previous-policy regularization. On eight MO-Gymnasium tasks, it achieves the best average hypervolume rank among recent baselines and strong expected-utility performance. Continuous-control experiments indicate gains beyond the discrete-action setting.
PaperID: 5217, Poster
Abstract: In centralized, distributed, and federated learning with stochastic gradients and n workers, it was recently shown that it is infeasible to find an \varepsilon-stationary point faster than \tilde\Omega(\min[\fracd \kappa L \Delta\varepsilon + \frach L \Delta\varepsilon + \frach \sigma^2 L \Deltan \varepsilon^2, \frach \sigma^2 L \Delta\varepsilon^2 + \frach L \Delta\varepsilon]) seconds in both homogeneous and heterogeneous settings under standard assumptions: L-smoothness, \sigma^2-bounded unbiased stochastic gradients, and lower boundedness of the function, i.e., f(x) \geq f^\star for all x \in \mathbbR^d, where \Delta = f(x^0) - f^, h is the computation time, \kappa is the communication speed between the workers and the server, and d is the dimension of the iterates and gradients. This result is pessimistic since it does not allow a complexity in which both \fracd \kappa L \Delta\varepsilon and \frach \sigma^2 L \Delta\varepsilon^2 improve with n, even when using random sparsification techniques; moreover, this lower bound can be matched by either non-distributed SGD or vanilla Synchronous SGD, which reduces the impact of recent progress in the design of compression-based methods. In this work, we challenge this limitation and propose new compressed methods, Inkheart SGD and M4, and show that under an additional structural assumption, which is necessary due to the lower bound and which does not restrict the class of considered problems, we achieve new state-of-the-art time complexities that break this pessimistic barrier and allow scaling with the number of workers n.
Abstract: Rigorous uncertainty quantification is essential for the safe deployment of autonomous systems in unconstrained environments. Conformal Prediction (CP) provides a distribution-free framework for this task, yet its standard formulations rely on exchangeability assumptions that are violated by the distribution shifts inherent in real-world robotics. Existing online CP methods maintain target coverage by adaptively scaling the conformal threshold, but typically employ a static nonconformity score function. We show that this fixed geometry leads to highly conservative, volume-inefficient prediction regions when environments undergo structural shifts. To address this, we propose AdaptNC, a framework for the joint online adaptation of both the nonconformity score parameters and the conformal threshold. AdaptNC leverages an adaptive reweighting scheme to optimize score functions, and introduces a replay buffer mechanism to mitigate the coverage instability that occurs during score transitions. We evaluate AdaptNC on diverse robotic benchmarks involving multi-agent policy changes, environmental changes and sensor degradation. Our results demonstrate that AdaptNC significantly reduces prediction region volume compared to state-of-the-art threshold-only baselines while maintaining target coverage levels.
PaperID: 5219, Poster
Authors: Jainendra Shukla, Dhruv Jaiswal, Divyanshi Beniwal, Kiriti Kanjilal
Abstract: In many interactive settings, from social and institutional environments to multi-agent systems, agents must infer latent behavioral structure from limited signals while their own actions shape the data they observe. This endogeneity violates the stationary, exogenous assumptions underlying standard likelihood-based methods, leading to failures in adaptation despite accurate clustering. We propose \emphEpistemic Social Learning (ESL), a two-timescale framework that couples belief-based inference with adaptive representation learning under endogenous interaction. At a fast timescale, agents maintain Bayesian beliefs over shared behavioral prototypes. At a slower timescale, prototypes are updated from belief-weighted interaction data via the recursion \Theta_m+1 = \Theta_m + \gamma_m \widehatH_m, aligning learned representations with the interaction distributions induced by agent behavior. We formalize ESL as a stochastic approximation process with controlled Markov noise and show that its dynamics track the differential inclusion \dot\Theta \in \mathcalG(\Theta), whose limit sets correspond to self-consistent interaction regimes. Across repeated games with behavioral shifts, ESL reduces post-switch regret by 26% over strong baselines while improving decision-relevant prediction. These results reveal a fundamental gap between latent identification and adaptive performance under endogenous interaction.
Authors:
Tung-Ling Li, Yuhao Wu, Hongliang LiuAbstract: LLM-as-a-Judge systems supply the reward signal in modern RLHF and RLVR pipelines, but their binary verdict reduces to a single linear readout on one hidden state. We show this readout is shallow enough that short, low-perplexity tokens flip the verdict from "No" to "Yes". These tokens are sampled from the judge's own next-token distribution at the response position, with no manual seed set and no gradient-based optimization. Our procedure, AdvJudge-Zero, reaches >90% ensemble false-positive rate on 22 of 24 (model, dataset) cells across six Qwen, Llama, and Gemma judges, versus 54-72% for the prior curated 10-token benchmark, and the discovered surface transfers cross-format to a 70B scalar reward model. The same discovered pool enables a defense: a LoRA fine-tune stratified by a 9-class mechanism taxonomy hardens cross-family generalization where naive sampling on the same pool fails, with mechanism breadth rather than pool size carrying the gain. Under GRPO training, the hardened judge eliminates the reward-collapse failures (false-positive spikes and length collapse) we observe in the unhardened baseline on both MATH and GSM8K at ten seeds per condition. The discovered pool, the mechanism taxonomy, and per-prompt flip records will be released under responsible disclosure.
PaperID: 5221, Poster
Authors: Davide Vitabile, Alexandro Buffa, Akshay Nambiar, Amril Nurman Nazir
Abstract: Medical language models promise expert clinical support, but the strongest open source LLMs remain too large for private, edge devices deployment, while compact medical models still trail substantially on knowledge-intensive and real-world healthcare tasks. We present MedPsy, a family of text-only 1.7B and 4B medical language models designed for edge deployment. Our recipe combines a synthetic medical data pipeline over biology, medicine, and health seeds with chain-of-thought targets from a 235B medical-focused LLM teacher; a four-stage post-training curriculum of two SFT stages followed by two RL stages with hard-sample mining; and a mobile-oriented quantization study over several GGUF variants per model. On seven closed-ended medical benchmarks, MedPsy-4B surpasses MedGemma-1.5-4B by +19.34 points and matches MedGemma-27B while being 6.75x smaller; MedPsy-1.7B outperforms MedGemma-1.5-4B by +11.42 points despite being less than half its size. On HealthBench-Hard, MedPsy-4B surpasses MedGemma-27B by +15.33 and MedPsy-1.7B surpasses it by +11.66 points despite being 16x smaller. Beyond accuracy, MedPsy reduces average response length by 1.7x (1.7B) and 3.2x (4B) versus its Qwen3 backbones, and 4-bit quantization retains accuracy within 1 point of BF16 while reducing disk footprint by ~69%, enabling practical deployment on resource-constrained edge devices.
PaperID: 5222, Poster
Abstract: 3D multimodal language models are becoming a foundation for embodied agents, indoor robotics, and human–AI interaction, but their answers and explanations often remain weakly tied to the specific objects and spatial relations observed in a 3D scene. This paper aims to make 3D reasoning more faithful by reducing language-prior rationalization and forcing reasoning chains to depend on verifiable scene evidence. We propose Anti-Rationalization Chain-of-Thought (AR-CoT), a plug-in training and decoding framework that represents each reasoning step with an explicit evidence pointer and scores candidate answer-chain pairs by their scene-evidence gain. AR-CoT contrasts scene-conditioned rationales with a frozen text-only prior, verifies local object and relation claims, and uses decoy-twin scene pairs to encourage reasoning chains to change when the relevant spatial evidence changes. Experiments on standard 3D QA and grounding benchmarks, as well as MSQA, Beacon3D, and shortcut-sensitive stress tests, show that AR-CoT consistently improves multiple 3D MLLM backbones while strengthening chain grounding, decoy-twin contrast, and scene-swap sensitivity. These results suggest that anti-rationalization offers a practical and verifiable route toward more accurate, interpretable, and scene-faithful 3D multimodal reasoning.
PaperID: 5223, Poster
Abstract: Modern LLMs/VLMs have largely converged to a unified architecture: representing heterogeneous inputs as a single token sequence and processing it with a deep stack of generic operators such as attention and feed-forward layers. This sequence-first interface simplifies system designs and thus enables extensive algorithmic and hardware optimization, making it a compelling blueprint for modern model development. In contrast, 3D multi-modal fusion models still rely on complex and specialized designs such as dense feature representations, sparse 3D convolutions, deformable attention, view transformations, etc. In this paper, we ask: Can 3D multi-modal fusion embrace the same design path as modern LLMs/VLMs and benefit from their mature and rapidly evolving ecosystem? We answer this affirmatively and propose FusionNeXt, a modern multi-modal fusion paradigm that (i) unifies camera and LiDAR features into a shared token representation, (ii) serializes tokens into locality-preserving 1D sequences, and (iii) performs feature fusion using a deep stack of vanilla LLM/VLM-style blocks---FlashAttention, pre-norm residuals, SwiGLU FFNs, etc. The proposed paradigm enables standard sequence modeling tools to be applied directly to 3D multi-modal fusion, resulting in fast inference, strong performance, and generalization across tasks. FusionNeXt achieves state-of-the-art (SOTA) results on the nuScenes 3D detection and Occ3D occupancy benchmarks while delivering high inference throughput, suggesting a scalable direction for 3D perception. We will open-source our code.
Abstract: Retrieval-Augmented Generation (RAG) systems are increasingly deployed in high-stakes domains where users expect outputs to be consistent across semantically equivalent queries. However, existing systems often exhibit significant inconsistencies due to variability in both the retriever and generator, undermining trust and reliability. In this work, we focus on —the requirement that generated outputs convey the same core content across semantically equivalent inputs. We introduce a principled evaluation framework that decomposes RAG consistency into retriever-level, generator-level, and end-to-end components, enabling systematic identification of inconsistency sources. To improve consistency, we propose et Group Relative Policy Optimization (PS-GRPO), an RL approach that leverages multiple rollouts across paraphrased set to assign (Con-RAG), training the generator to produce consistent outputs across paraphrased queries and remain robust to retrieval-induced variability. Because exact reward computation over paraphrase sets is computationally expensive, we also introduce a scalable approximation method that retains effectiveness while enabling efficient, large-scale training. Empirical evaluations across short-form, multi-hop, and long-form QA benchmarks demonstrate that Con-RAG significantly improves both consistency and accuracy over strong baselines, even in the absence of explicit ground-truth supervision. Our work provides practical solutions for evaluating and building reliable RAG systems for safety-critical deployments.
PaperID: 5225, Poster
Abstract: Nonlinear state space models are typically evaluated sequentially, limiting their ability to exploit modern parallel hardware. However, recent work has shown that state space models can be evaluated in parallel by reformulating evaluation as a root-finding problem, which can be solved with a parallel form of Newton's method. However, exact Newton methods require expensive multiplication of dense Jacobian matrices, and quasi-Newton methods based on diagonal approximations, while more scalable, are often slow to converge because they neglect interactions across state dimensions. We introduce a novel quasi-Newton approach based on Broyden’s method, which captures coupling terms with diagonal-plus-low-rank approximations to blocks of the Jacobian. We construct the approximations from trajectory secants, avoiding automatic differentiation while retaining the favorable scaling of diagonal approximations and remaining compatible with parallel scan through rank compression. We extend previous theoretical results to show that convergence degrades monotonically with approximation error and prove finite-step recovery of the exact trajectory. Empirically, we show that our method is advantageous when interactions are approximately low rank or when Jacobian evaluation is prohibitively expensive. It converges in fewer iterations than other quasi-Newton methods, yielding substantially faster run times for simulating low-rank RNNs, parallel MCMC algorithms, and denoising diffusion models.
PaperID: 5226, Poster
Abstract: The Representation Autoencoder (RAE) shows that a frozen high-dimensional semantic ViT representation enables both image reconstruction and DiT-based generation. However, RAE's semantic-only representation lacks appearance details, yielding suboptimal reconstruction. Full fine-tuning the encoder with reconstruction loss induces catastrophic semantic collapse: the encoder overfits to appearance, loses semantic discriminability, and cripples generation. This failure arises because the reconstruction loss imposes an appearance prior that erodes semantic structure. To resolve this conflict, we propose Semantic Structure Regularization (SSR). SSR builds on self-distillation, using a learnable student and a frozen teacher that retains the original semantics. In addition to standard latent alignment, which forces the student's representation to match the teacher's, we introduce two complementary regularizers that prevent appearance overfitting from corrupting semantic structure. First, spatial self-similarity distillation forces the student to reproduce the teacher's patch-wise relational patterns. Invariant to appearance changes, this relational signature anchors high-order semantics and prevents the sacrifice of structure for pixel fidelity. Second, structure-anchored alignment decouples structure from appearance. Using a frequency-domain transform, we extract a high-frequency structural map (edges, boundaries) from each image while discarding low-frequency appearance. The teacher's representation of this structural map then regularizes the student's representation of the original image. Consequently, the student learns appearance details from the reconstruction loss without penalty, where the regularization enforces only structural consistency and enables pixel-level encoding without semantic drift. The result is a unified representation that excels at high-fidelity reconstruction, discriminative perception, and efficient DiT-based generation. Our work provides a principled path toward a single semantic space bridging perception and generation, challenging the classic dichotomy. Code and models will be released.
PaperID: 5227, Poster
Abstract: Instruction-guided video generative models offer a scalable solution for stress-testing autonomous driving systems by simulating diverse environmental conditions. However, driving scenes operate within a dense Operational Design Domain (ODD) governed by strict traffic semantics. When applied to global driving-condition edits, general-purpose instruction-guided video editors frequently corrupts these semantics, leading to structural hallucinations and ``temporal shock''. This work strictly distills and protects traffic semantics during video editing at three levels. First, we construct a hybrid dataset by combining temporally aligned real-world videos with synthetic edits, balancing foundational scene structure with unambiguous appearance signals. Second, during training, we cast model alignment as offline reinforcement learning via advantage conditioning. We extract dense, pixel-level rewards from the hybrid data to serve as a localized signal and actively steer generation toward structurally reliable pixel states. Finally, our model mitigates temporal shock by introducing thinking frames, a transition mechanism that seamlessly reconciles static image conditioning with dynamic video context. Extensive experiments demonstrate that our method prevents the corruption of multi-agent dynamics, establishing state-of-the-art spatiotemporal consistency and yielding an overwhelming user preference (88 % relative improvement in content preservation and 75 % in temporal consistency) and yielding a significant 5 % gain in downstream road line segmentation. We provide more video results in the supplementary material.
Authors: Roel Koopman, Sebastian Otte, Sander Bohte
Abstract: Neuromorphic hardware implementations of Spiking Neural Networks (SNNs) promise energy-efficient, low-latency AI through sparse, event-driven computation. Yet, training SNNs under fine temporal discretization remains challenging, hindering low-latency responsiveness and the mapping of software-trained SNNs to efficient hardware. In current approaches, spiking neurons are modeled as self-recurrent units, embedded into recurrent networks to maintain state over time, and trained with BPTT or RTRL variants based on surrogate gradients. These methods scale poorly with temporal resolution, while online approximations exhibit instability for long sequences and tend to fail at capturing temporal patterns. To address these limitations, we develop SpikingGamma, an SNN architecture in which neurons maintain a compact bank of smoothed-delay states and communicate through sigma-delta spike-coding. We show that in feedforward networks, SpikingGamma supports direct online error backpropagation through the continuous reconstruction signal, avoiding surrogate gradients through the spike discontinuity. This enables stable learning of temporal patterns with minimal spiking and scales feedforward SNNs to complex tasks and benchmarks with competitive accuracy, all while remaining robust to the temporal resolution of the model. Our approach offers both an alternative to recurrent SNNs trained with surrogate gradients, and a novel route for mapping SNNs to neuromorphic hardware.
Authors: Xin(Allen) Wang, Kaiwen Shi, Carlos Oliver
Abstract: Protein function is driven by cohesive substructures, such as catalytic triads, binding pockets, and structural motifs, that occupy only a small fraction of a protein's residues. Yet existing pipelines built on protein encoders do not model proteins at the substructure level, leaving the central biological question unanswered: We introduce BioBlobs, an encoder-agnostic, end-to-end differentiable framework that compresses a protein into a small set of cohesive substructures ( ) and predicts function from these blobs alone, so that each blob corresponds to a candidate functional region. Across diverse protein function prediction tasks and multiple sequence- and structure-based encoders, BioBlobs matches or exceeds strong baselines while operating on only a small fraction of residues. The discovered adapt their spatial scale to the task, ranging from local catalytic sites to entire structural domains. Trained only on protein-level labels, BioBlobs recovers experimentally annotated catalytic sites in the M-CSA database, demonstrating unsupervised functional substructure discovery and opening a path to large-scale functional site discovery across the unannotated proteome.
Abstract: Continuous-time models are a natural choice for irregular and asynchronous data. A central design choice is how to embed discrete observations into continuous time. Interpolation- and imputation-based embeddings reconstruct a continuous observation path, making the model sensitive to the choice of reconstruction. We show that this reconstruction step is unnecessary. On compact sets, universality for continuous functionals of paths transfers to universality for continuous functionals of discrete observation streams under any continuous and injective embedding. Guided by this result, and building on the rectilinear control path for Neural Controlled Differential Equations (NCDEs), we introduce a continuous and injective embedding for Log-NCDEs, a universal class of continuous-time models. This embedding records observations as increments and composes them over arbitrary query intervals to form log-signatures, giving interval-level summaries of a faithful embedding of the observed data. This avoids interpolation of observed variables, supports online computation, and allows prediction on output grids independent of the input sampling times. Experiments on synthetic controlled dynamics and real-world time-series datasets show that the representation is accurate, efficient, and robust to irregular, asynchronous, and sparse observations.
PaperID: 5231, Poster
Abstract: Low-bit storage of optimizer states can substantially reduce the memory footprint of large-scale training. In practice, however, these states are not quantized once: at every training step, they are dequantized, updated, requantized, and written back. This iterative quantization pipeline can destabilize training. We identify a failure mode behind this instability: for states updated via exponential moving averages (EMAs), quantization errors are recursively fed into subsequent updates. As a result, biased errors accumulate over time while high-variance errors are amplified, particularly when the EMA decay factor is close to 1. To address this error accumulation, we propose BReD (Block Replay Dithering), a method that reduces rounding bias and controls rounding variance during quantization. BReD adapts classical subtractive dithering to optimizer-state quantization by combining deterministic seed replay with block-wise shared dither values, avoiding auxiliary random-tensor storage and per-element dither values generation. Across pretraining (five optimizers at 120M; AdamW and Muon at 1.1B and 3.4B model) and supervised fine-tuning (AdamW and Muon at 7B model), 4-bit BReD closely matches training with full-precision optimizer states, with PPL shifts \leq0.5 in pretraining and \leq0.2 in fine-tuning and average downstream performance changes within 0.7 points. Moreover, 3-bit BReD preserves stable convergence in the evaluated settings, offering a more aggressive alternative.
Abstract: Understanding reflection remains a long-standing challenge in 3D reconstruction due to the entanglement of appearance and geometry under view-dependent reflections. In this work, we present the Pygmalion Effect in Vision, a novel framework that metaphorically “sculpts” reflective objects into clay-like forms through image-to-clay translation. We then introduce a dual-branch design in which a BRDF-based reflective branch and a clay-guided branch share the same Gaussian geometry but operate through independent rendering paths. The clay-guided branch, supervised by the synthesized clay images, provides reflection-free geometric guidance that complements the photometric supervision of the BRDF branch. Experiments on both synthetic and real datasets show consistent improvements in both geometry and photometric metrics. Beyond technical gains, our framework reveals that seeing by unshining, translating radiance into neutrality, can serve as a powerful inductive bias for reflective object geometry learning.
PaperID: 5233, Poster
Authors: Sreehari Rajan, Kunal Bhosikar, Charu Sharma, Nikos Athanasiou
Abstract: Given a single image of a room and a sentence such as ``a person walks to the black armchair and sits down,'' producing a 3D motion sequence of a virtual person performing the described action is out of reach for existing methods. Scene-conditioned methods need hours of paired (scene, motion, language) capture. Zero-shot methods need a pre-scanned 3D scene. Image-based methods produce only static poses. We propose a training-free pipeline that synthesizes scene-aware human-scene interaction (HSI) motion from a single RGB image and a free-form text prompt, with no HSI-specific training and no pre-scanned 3D scene. The key observation is that every needed capability already exists as a frozen foundation model: the bottleneck is composition, not capability. Our method, ENACT, uses a vision-language model (VLM) for high-level reasoning. The VLM parses the input image and prompt into one structured keyframe per sentence, listing the action verb, the body part, the target object, a six-bin contact yaw, an image-edit prompt, and a foot-grounding flag. The keyframe then drives six frozen specialists: a monocular 3D model for the scene point cloud, an open-vocabulary segmenter for object localization, a single-view mesh reconstructor for object geometry, an instruction-following image editor for a synthetic interaction image, an image-conditioned human-mesh recoverer for pose initialization, and a 2D-grounded 3D contact predictor for object affordance. Per-keyframe contact poses are optimized under a whole-body diffusion prior, and a constraint-conditioned motion diffuser connects successive keyframes into one continuous motion. ENACT achieves the lowest mean and maximum scene penetration and the lowest foot sliding among all baselines on a standard HSI benchmark, despite consuming the smallest input. ENACT also generalizes to synthetic datasets and to in-the-wild phone images, including chained multi-step interactions.
PaperID: 5234, Poster
Abstract: Node classification in multimodal graphs plays an important role in many real-world applications, where nodes are described by multiple modalities such as text and images, and the graph structure captures relations among entities. However, most existing Graph Neural Networks (GNNs) still focus on pairwise connections, overlooking higher-order relational patterns. Recent studies have explored simplicial complexes to capture higher-order interactions and integrated them into GNN frameworks. Motivated by this line of work, we propose Adaptive Gated Simplicial Propagation (AGSP), an end-to-end framework for node classification in multimodal graphs. Specifically, AGSP first introduces a simplicial propagation layer to capture higher-order relations beyond pairwise connections; then it applies an adaptive gating mechanism to balance multi-order topological relations, enhancing the discriminative ability of node classification. Extensive experiments validate that AGSP generally outperforms the state-of-the-art baselines, highlighting the effectiveness of combining multimodal information with higher-order graph topology through adaptive gated fusion.
PaperID: 5235, Poster
Abstract: We study policy optimization for online episodic tabular Markov decision processes with unknown transition kernels, aiming for best-of-both-worlds guarantees and multiple data-dependent regret bounds. Recent work (Dann et al., 2023; Li et al., 2026) has shown that policy optimization can adapt to both adversarial and stochastic losses with data-dependent bounds, including first-order, second-order, and path-length bounds, but only under known transitions. We resolve the open problem raised by Dann et al. (2023) by developing optimistic follow-the-regularized-leader algorithms that extend such guarantees to unknown transitions. The key ingredient is a new design of optimistic Q-function estimators together with a data-dependent transition bonus that controls estimator bias through the loss-prediction error. Our analysis further identifies an unavoidable transition-dependent complexity term that captures the intrinsic cost of estimating the transition kernel. As a result, we obtain first-order, second-order, and path-length bounds with this transition-dependent complexity term while simultaneously achieving gap-dependent \mathrmpolylog(T) regret in the stochastic regime.
PaperID: 5236, Poster
Authors: Victor Huang, Poodar Chu, Hongsheng Li
Abstract: A reward is not a universal RL interface: it becomes a valid update only through the probability object exposed by the policy branch. Autoregressive post-training works cleanly because token ratios are exact, but parallel discrete policies hide within-step dependence and flow/diffusion policies often lack cheap density ratios. We introduce MindRL, a reward-to-update interface controller that translates a shared reward through branch-native score objects while turning factorization, tractability, drift, and smoothness barriers into budgets over block size, serialization, anchors, clipping, reranking, and branch weights. As a closed-loop controller, MindRL makes the shared reward actionable for AR, parallel discrete, flow/diffusion, and AR+flow policies by selecting branch-native update objects and adapting their control budgets. Across language-side AR/parallel-discrete tests and matched flow/diffusion and AR+flow evaluations, MindRL improves task-quality or reward-risk tradeoffs while preserving each branch's native structure.
Abstract: Vision-language models (VLMs) achieve strong performance on multimodal benchmarks, but may still lack robust control over basic visual operations. We study line tracing, where a model must follow a selected visual path through successive local continuations. To isolate this ability, we design controlled tracing tasks that introduce nearby competitors while reducing semantic and topological ambiguity such as crossings and overlaps. Across these tasks, even state-of-the-art VLMs frequently lose the target path and switch to nearby alternatives, especially when those alternatives look locally similar to the target. Behavioral interventions and internal analyses indicate that these failures arise from local competition: nearby similar distractors pull the model away from the true continuation. Standard remedies do not remove this bottleneck: model-size scaling provides only limited gains, reasoning partially compensates through costly substitute strategies, and explicit tracing instructions fail to recover stable path following. Finally, tests on tangled-cable scenes and metro maps with richer visual complexity show that the same path-switching failure persists beyond our controlled settings.
Abstract: Graphs with a simple spectrum admit cubic-time isomorphism testing, yet we prove that for every natural number k, the k-Weisfeiler-Leman (k-WL) test cannot distinguish all non-isomorphic graphs with a simple spectrum. As the WL hierarchy upper-bounds the distinguishing power of widely-used Graph Neural Networks (GNNs), this incompleteness applies to all such GNNs, ruling out completeness for every k-WL-aligned GNN family. To close this gap, we introduce PRiSM (Partition, Refine, Solve, Match), the first provably complete canonicalization of simple-spectrum eigendecompositions. PRiSM obtains the completeness guarantee that prior canonicalizations provably lack, and resolves the open problem of achieving complete expressivity on simple-spectrum graphs. When composed with DeepSets or a Transformer, PRiSM achieves universal approximation on simple-spectrum graphs, justifying the use of canonicalized Laplacian positional encodings. Empirically, PRiSM performs comparably to or outperforms existing spectral canonicalizations on graph regression, classification, and expressivity benchmarks.
PaperID: 5239, Poster
Authors: Dmitrii Kharlapenko, Sergei Bratchikov, Konstantin Korolev, Aleksandr Nikolich
Abstract: On modern production text-to-image systems, successful policy violations are rare, and previously effective human-written seeds are often patched out. Current automated red-teamers are poorly matched to this regime in two ways: unreliable success measurement and poor exploration. First, we find that judges widely used in prior T2I red-teaming work are unreliable under vague unsafe-content targets: they either miss true violations or reward benign borderline images on hardened APIs. We therefore define strict category-specific success criteria and calibrate strong VLM judges against human labels. Second, we show that broadly used prompt-modification pipelines do not solve the exploration problem: on harder guardrail settings they remain tied to seed prompts, fail to transfer, or cannot bootstrap positive examples. We introduce RISE, which evolves reusable strategies used to generate prompts rather than rewriting them one by one. The best discovered strategies are then reused to generate attacks across new scenarios. On DALL·E 3, Nano Banana 2 (Google) and GPT-Image-2, RISE reaches up to 13% human-verified ASR; under the same calibrated evaluation, prior methods with reported ASR as high as roughly 30% fall to near zero.
PaperID: 5240, Poster
Abstract: Federated Low-Rank Adaptation (FedLoRA) enables efficient fine-tuning of large language models by exchanging low-rank adaptation parameters across clients. Several studies mainly focus on aggregation bias and design improved aggregation schemes. We argue that aggregation bias is not the fundamental bottleneck of FedLoRA. Even under ideal aggregation, heterogeneous data induce directional conflicts among client-specific LoRA updates. We show that these conflicts arise from misaligned low-rank adaptation subspaces, leading to destructive interference during collaboration. To diagnose this phenomenon, we introduce a geometry-aware metric that characterizes client updates at the subspace level. Our analysis reveals that subspace conflicts intensify with increasing heterogeneity and are concentrated in the column spaces of LoRA updates. Based on this insight, we propose ALIGN, a geometry-aware federated adaptation framework that aligns client LoRA subspaces during aggregation and constrains local updates to mitigate conflicts. Comprehensive empirical results confirm that mitigating subspace misalignment is crucial for FedLoRA, and that ALIGN consistently improves downstream performance over existing methods. The code is available for anonymous access at https://anonymous.4open.science/r/ALIGN-7C41.
Abstract: Depth-recurrence promises improved latent reasoning by sharing parameters across depths. However, prior work still relies on partially fixed layer stacks and overlooks the bottleneck of constant hidden-size; moreover, it largely lacks rigorously resource-matched ablations. To address this, we introduce depth-recurrent attention mixtures (Dreamer), a modular architecture with a single recurring layer. Concretely, we mix sequence attention, depth attention, and sparse expert attention. This alleviates the hidden-size bottleneck, decouples scaling dimensions, and achieves effective yet efficient depth-recurrence. Across natural language reasoning benchmarks at up to ~2B parameters, Dreamer requires on average ~3x fewer training tokens for the same accuracy as FLOP-, parameter-, and memory-matched modern Transformers, and outperforms ~2x larger Transformers given same training tokens. Analysis further reveals 2-11x greater expert selection diversity than conventional MoEs, highlighting the flexible knowledge sharing across depths.
PaperID: 5242, Poster
Abstract: Pre-trained text-to-image (T2I) diffusion models have shown strong potential for real-world image super-resolution (Real-SR), owing to their noise-started generation process that enables realistic texture synthesis and captures the one-to-many nature of super-resolution. However, diffusion-based Real-SR methods still face a fundamental efficiency-quality trade-off. Multi-step methods generate high-quality results by iteratively denoising random Gaussian noise under LR conditioning, but suffer from slow sampling. Recent one-step methods greatly improve efficiency, yet they typically replace noise-started generation with direct LR-to-HR restoration, which weakens stochasticity and limits realistic detail synthesis. To address this issue, we propose SMFSR, a noise-started one-step Real-SR framework via LR-conditioned SplitMeanFlow and GAN refinement. SMFSR preserves the random-noise starting point of diffusion models and learns a direct noise-to-HR mapping conditioned on the LR image. To this end, Interval Splitting Consistency distills the multi-step generative trajectory into a single average-velocity prediction, enabling efficient one-step generation. To compensate for the reduced opportunity for progressive refinement, we further introduce a GAN refinement stage, where a DINOv3-based discriminator enhances realistic texture synthesis and variational score distillation aligns the generated outputs with the natural image distribution. Extensive experiments demonstrate that SMFSR achieves superior perceptual quality while retaining the high efficiency of one-step diffusion models.
Abstract: Federated Learning (FL) algorithms implicitly assume that clients passively comply with server-side orchestration by sharing local model updates upon server request. However, this overlooks an important aspect in real-world cross-silo environments: clients are often rational agents who may prioritize their utilities such as local model performance over that of the global model. In settings with significant statistical heterogeneity, rational clients may opt out of the federation if the perceived benefits of collaboration fail to meet their local utility thresholds. Such attrition degrades the global model performance and can lead to the collapse of the federated training process. In this work, we introduce FedUCA: Federated Learning by Utility-Constrained Stochastic Aggregation for Improving Rational Participation, a framework that formalizes the server's role as an optimizer seeking to maximize global model performance by sustaining client participation. We substantiate our framework through extensive experiments on standard datasets demonstrating that by prioritizing participation feasibility, FedUCA achieves significantly higher client retention and, consequently, a superior global model performance.
Abstract: Mamba, a recently proposed linear-time sequence model, has attracted significant attention for its computational efficiency and strong empirical performance. However, a rigorous theoretical understanding of its underlying mechanisms remains limited. In this work, we provide a theoretical analysis of Mamba's in-context learning (ICL) capability by focusing on tasks defined by low-dimensional nonlinear target functions. Specifically, we study in-context learning of a single-index model y \approx g_(\langle \boldsymbol\beta, \boldsymbolx \rangle), which depends on only a single relevant direction \boldsymbol\beta, referred to as feature. We prove that Mamba, pretrained by gradient-based methods, can achieve efficient ICL via test-time feature learning, extracting the relevant direction directly from context examples. Consequently, we establish a test-time sample complexity that improves upon linear Transformers---analyzed to behave like kernel methods---and is comparable to nonlinear Transformers, which have been shown to surpass the Correlational Statistical Query (CSQ) lower bound and achieve near information-theoretically optimal rate in previous works. Our analysis reveals the crucial role of the nonlinear gating mechanism in Mamba for feature extraction, highlighting it as the fundamental driver behind Mamba’s ability to achieve both computational efficiency and high performance.
Abstract: The foundational capabilities of large language models are acquired during pretraining on internet-scale, highly heterogeneous data mixtures. In this work, we investigate a geometric question regarding the converged state of this pretraining process: Does the model converge to a common minimizer across all data sources (e.g., \creffig:cwa illustration:distant)? We hypothesize that the geometric ``closeness'' of task-specific minima is intrinsically linked to downstream generalization. However, we reveal that standard optimizers (e.g., AdamW) often converge to points where task-specific minima are distant from each other. To address this, we propose the Nexus optimizer, which encourages the closeness of these minima by maximizing gradient similarity during optimization. Extensive experiments across models ranging from 130M to 3B parameters demonstrate that Nexus significantly boosts downstream performance, despite achieving nearly the same pretraining loss (see \creffig:demo:benchmark). Notably, on the 3B model, Nexus reduces the out-of-distribution loss by 0.012 and yields up to a 15.0% accuracy improvement on complex reasoning tasks (e.g., GSM8k). Our findings challenge the reliance on pretraining loss as the sole proxy for model evaluation and highlight the critical role of optimizer implicit bias in unlocking downstream generalization.
Authors:
Hao Wu, Yuqi Li, Yuan Gao, Fan Xu, Fan Zhang, Kun Wang, Penghao Zhao, Qiufeng Wang, Yizhou Zhao, Weiyan Wang, YingLi Tian, Xian Wu, Xiaomeng HuangAbstract: Existing robot video world models are typically trained with low-level objectives such as reconstruction and perceptual similarity, which are poorly aligned with the capabilities that matter most for robot decision making, including instruction following, manipulation success, and physical plausibility. They also suffer from error accumulation in long-horizon autoregressive prediction. We present RoboAlign-R1, a framework that combines reward-aligned post-training with stabilized long-horizon inference for robot video world models. We construct RobotWorldBench, a benchmark of 10,000 annotated video-instruction pairs collected from four robot data sources, and train a multimodal teacher judge, RoboAlign-Judge, to provide fine-grained six-dimensional evaluation of generated videos. We then distill the teacher into a lightweight student reward model for efficient reinforcement-learning-based post-training. To reduce long-horizon rollout drift, we further introduce Sliding Window Re-encoding (SWR), a training-free inference strategy that periodically refreshes the generation context. Under our in-domain evaluation protocol, RoboAlign-R1 improves the aggregate six-dimension score by 10.1% over the strongest baseline, including gains of 7.5% on Manipulation Accuracy and 4.6% on Instruction Following; these ranking improvements are further supported by an external VLM-based cross-check and a blinded human study. Meanwhile, SWR improves long-horizon prediction quality with only about 1% additional latency, yielding a 2.8% gain in SSIM and a 9.8% reduction in LPIPS. Together, these results show that reward-aligned post-training and stabilized long-horizon decoding improve task consistency, physical realism, and long-horizon prediction quality in robot video world models.
Abstract: Recent progress in latent world models (e.g., V-JEPA2) has shown promising capability in forecasting future world states from video observations. Nevertheless, dense prediction from a short observation window limits temporal context and can bias predictors toward local, low-level extrapolation, making it difficult to capture long-horizon semantics and reducing downstream utility. Vision-language models (VLMs), in contrast, provide strong semantic grounding and general knowledge by reasoning over uniformly sampled observation frames, but they are not ideal as standalone dense predictors due to compute-driven sparse sampling, a language-output bottleneck that compresses fine-grained interaction states into text-oriented representations, and a data-regime mismatch when adapting to small action-conditioned datasets. We propose a VLM-guided JEPA-style latent world modeling framework that combines dense-frame dynamics modeling with long-horizon semantic guidance via a dual-temporal pathway: a dense JEPA branch for fine-grained motion and interaction cues, and a uniformly sampled observation VLM thinker branch with a larger temporal stride for knowledge-rich guidance. To transfer the VLM's progressive reasoning signals effectively, we introduce a hierarchical pyramid representation extraction module that aggregates multi-layer VLM representations into guidance features compatible with latent prediction. Across EgoDex, EgoExo4D, BAIR Robot Pushing, and Physion, ThinkJEPA outperforms diverse latent world model and trajectory prediction baselines across egocentric trajectory prediction, long-horizon rollout, robotic latent prediction, and physical-scene forecasting. These results show that broad visual-semantic guidance from a VLM thinker can benefit JEPA-style latent forecasting.
Authors: Jiaan Han, Junxiao Chen, Yanzhe Fu
Abstract: Most FDR-controlled feature selection methods are designed for coordinate-wise hypotheses, where each feature has a single weight or importance score. This abstraction fails in sequential and grouped models, where one original feature is represented by a block of sub-features, such as lags, recurrent states, or attention-based interactions. We propose a grouped-feature FDR control framework for such settings. For grouped linear models, we construct null-symmetric block-level mirror statistics with matrix-valued perturbations. For neural sequential models, we combine Permutation SHAP derivatives as model-agnostic block-level importance scores with kernel-based dependence measure. The framework is model-agnostic across network architectures, does not require specifying the covariate distribution, and reduces to Gaussian Mirror or Neural Gaussian Mirror when the block size is one. We prove FDR control for low- and high-dimensional grouped linear models and asymptotic symmetry of smoothed Permutation SHAP derivatives under fixed fitted nonlinear models. Experiments on simulated and real-world datasets show reliable FDR control and improved power under correlated grouped-feature signals.
PaperID: 5249, Poster
Abstract: We study signSGD in linear regression with diagonal features and power-law spectra, a toy but expressive setting where the effect of coordinate-wise sign normalization can be analyzed sharply. Specifically, we establish a functional scaling law (FSL) for signSGD under general learning rate schedules, which decomposes the loss dynamics into two components: learning of the target signal and noise accumulation. Compared with SGD, signSGD learns the signal faster but exhibits slower noise forgetting. This signal-noise tradeoff yields a regime-dependent comparison between signSGD and SGD: signSGD achieves better data-scaling efficiency in hard-task regimes, while they perform similarly in easy-task regimes. The key mechanism is that gradient noise is \it curvature-aligned: its directional variance is of the same order as the directional curvature. As a result, sign operation effectively acts as a \(\diag(\bH)^-1/2\) preconditioner in the noise-dominated regime, where \bH is the Hessian matrix. Moreover, we demonstrate that the same square-root preconditioning mechanism can arise in RMSprop and Adam, where coordinate-wise normalizers essentially track the coordinate curvatures. Synthetic experiments validate the proposed mechanism and the predicted scaling behavior. Language model experiments further suggest that, beyond the controlled setting, the resulting FSL can serve as a surrogate model for fitting and predicting practical training loss curves.
PaperID: 5250, Poster
Abstract: Best-of-N sampling is a simple inference-time scaling strategy for large language models, yet its behavior under incomplete reward models remains poorly understood. We study Best-of-N as a competition between the extreme reward tails of correct and incorrect responses. This view shows that success depends not only on the probability of sampling a correct response, but also on whether correct responses dominate incorrect ones in the upper tail. Using extreme value theory, we characterize three asymptotic regimes: correct-tail-dominant, incorrect-tail-dominant, and critical explaining both monotonic gains and reward-hacking-induced degradation. We further derive finite-sample scaling laws for representative reward distributions, showing that, under Gaussian rewards, incorrect responses with sufficiently large variance can dominate selection even when their mean reward is lower. Unlike oracle selection, convergence with an incomplete selector can be polynomial rather than exponential. Finally, we propose a lightweight tail-tension criterion to estimate when additional sampling may become harmful. Experiments on synthetic data and real reward model scores validate the predicted regimes, scaling laws, and diagnostic behavior, suggesting that tail competition provides a unified lens on Best-of-N sampling for verifiable problems.
Authors: Naram Mhaisen, George Iosifidis
Abstract: We study dynamic regret in online convex optimization with an \emphindicator switching cost: a fixed penalty incurred whenever two consecutive decisions differ. This captures startup overheads such as server activation, model deployment, and cache updates, and on a bounded domain it recovers norm-based movement costs as a special case. Existing guarantees for indicator costs handle only static comparators. We show that a direct extension of these techniques to dynamic regret provably fails, motivating a different approach. We propose a meta-learning framework: a set of randomized lazy FTRL base learners restarted at dyadic time scales, aggregated by a movement-aware master that mixes their proposal densities and samples actions via maximal coupling of consecutive mixtures. The resulting algorithm satisfies, in expectation, \mathcalR^\mathbf1_T \le \tilde\mathcalO(\min\\sqrtT(S_T+1),T^2/3(P_T+1)^1/3\), where \mathcalR^\mathbf1_T is the dynamic regret plus the cumulative indicator switching cost, S_T counts comparator switches, and P_T is the comparator path length. The bound holds simultaneously for all sequences and requires no prior knowledge of S_T or P_T: it is minimax-optimal (up to logarithmic factors) for tracking piecewise-constant comparators, and also captures frequently moving comparators with small total path length.
PaperID: 5252, Poster
Abstract: Recent foundation models are moving toward native multimodal Vision-Language Models (VLMs), making VLMs a central form of next-generation foundation models. However, their large language backbones make edge deployment difficult due to high memory footprint and memory-bound autoregressive decoding. Weight-only post-training quantization is a practical solution, but pushing VLMs to extreme low bit-widths remains challenging: existing rotation-free methods suffer from outliers at 2--3 bits, while rotation-based methods improve accuracy at the cost of additional runtime overhead. We propose SPHQuant, a rotation-free spherical weight-only quantization framework for VLMs. Instead of quantizing weights directly in Cartesian coordinates, SPHQuant decomposes each 8D weight vector into coordinate signs, radius, and a positive unit direction. This representation isolates outlier magnitude into the radius while keeping directions bounded and statistically regular. Based on this insight, SPHQuant allocates extra precision to the radius to mitigate accuracy degradation induced by outliers. It further uses a compact positive-direction codebook and fine-tunes codebook entries through angular parameterization to preserve the unit-sphere constraint. We also design a hardware-friendly GEMV kernel that keeps the direction codebook small enough for shared-memory lookup and packs radial bits efficiently. Experiments show that SPHQuant matches the performance of state-of-the-art extreme low-bit quantization methods while improving decode throughput by 30.3%, providing a new step toward efficient edge deployment of VLMs.
Authors: Aratrika Mustafi, Soumya Mukherjee
Abstract: We propose a dense associative memory for empirical measures (weighted point clouds). Stored patterns and queries are finitely supported probability measures, and retrieval is defined by minimizing a Hopfield-style log-sum-exp energy built from the debiased Sinkhorn divergence. We derive retrieval dynamics as a spherical Hellinger Kantorovich (SHK) gradient flow, which updates both support locations and weights. Discretizing the flow yields a deterministic algorithm that uses Sinkhorn potentials to compute barycentric transport steps and a multiplicative simplex reweighting. Under local separation and PL-type conditions we prove basin invariance, geometric convergence to a local minimizer, and a bound showing the minimizer remains close to the corresponding stored pattern. Under a random pattern model, we further show that these Sinkhorn basins are disjoint with high probability, implying exponential capacity in the ambient dimension. Experiments on synthetic Gaussian point-cloud memories demonstrate robust recovery from perturbed queries versus a Euclidean Hopfield-type baseline.
PaperID: 5254, Poster
Abstract: We revisit the classical problem of learning general Gaussian halfspaces in the presence of label noise, with particular focus on the large-bias regime where information-computation gaps are conjectured in several standard label-noise models. Concretely, we present and analyse the rare-tail mean estimator, an extremely simple one-shot statistic that averages only the observed minority class. In the large bias regime, classical Averaging (Servedio, 1999), and related global first-moment statistics, are provably suboptimal, while the rare-tail mean estimator recovers the target halfspace direction at the information-theoretic sample rate (modulo log factors) under a broad class of label noise models, including probit noise and margin-monotone noise with very slow decay away from the boundary. Notably, our guarantees allow labels arbitrarily close to the decision boundary to be nearly random, yet still obtain strong directional recovery from a small number of samples. As an application, we show that in a two-stage curriculum learning setup, where an average-case margin-dependent noise mechanism precedes the uniform-noise target (the worst-case noise profile in the same family) with the same Bayes optimal separator, our estimator supplies the missing directional accuracy needed by existing refinement routines, yielding an end-to-end sample-optimal learner in the large-bias regime. The estimator also cleanly interpolates between the noiseless and noisy regimes. In low-noise settings, it retains the same sample complexity rates as classical separability-based methods for realizable halfspace learning, while continuing to admit meaningful guarantees even after linear separability has been destroyed by noise.
Abstract: Large language models can hallucinate confidently, making uncertainty quantification essential for reliable deployment. Existing approaches rely predominantly on token-level signals from the final (unembedding) layer, while the rich geometric structure of intermediate hidden states remains largely unexplored. Aggregating token-level scores to the response level is itself non-trivial, and no single signal reliably catches the confident-but-wrong failure mode where a model commits to every token with high confidence yet produces an incorrect answer. We address this gap by extracting complementary signals from two distinct representational layers: hidden-state geometric complexity (global uncertainty) from the embedding layer, and token-level entropy (local uncertainty) from the unembedding layer. We show empirically that the two signals cover different regimes that are weakly correlated and thus combining them captures failure modes invisible to either alone. Building on this insight, we propose Global Local Uncertainty (GLU), an unsupervised, single-pass framework that fuses the two signals via a multiplicative gate. Experiments across three benchmarks and three model families show that GLU matches or outperforms all unsupervised baselines and remains competitive with supervised methods that lack cross-dataset generalization.
Abstract: Post-training pruning can substantially reduce LLM inference costs, but it often degrades quality unless the remaining weights are adapted. Since global retraining is expensive at LLM scale, recent work has largely focused on increasingly sophisticated pruning criteria that aim to select better sparsity patterns without adaptation. We revisit this trade-off through local reconstruction: after pruning, we adapt one subset of the model parameters at a time on a calibration set, training it to match the corresponding intermediate activations of the dense model. We evaluate local reconstruction across model families and scales, up to 72B parameters, and establish three main findings. First, local reconstruction is an effective adaptation mechanism for LLMs: it matches post-pruning retraining while using over an order of magnitude less data and compute, even when using PEFT techniques. Second, reconstruction exhibits a broad "free-lunch" regime in granularity, i.e., the reconstruction parameter window: as long as the reconstructed region contains at least a nonlinear submodule, final quality is largely insensitive to the window size, allowing granularity to be chosen primarily based on memory constraints. In contrast, reconstructing individual matrices, despite being the natural approach often proposed in the literature, consistently underperforms, as small matrix-level errors accumulate into larger activation drift. Lastly, reconstruction reduces the relative importance of the pruning criterion: performance gaps between sophisticated criteria and simple baselines shrink with model scale, making simple methods competitive again. Overall, our results challenge the prevailing view that post-pruning adaptation is impractical for LLMs.
Abstract: Clinical electroencephalogram (EEG) analysis rests on a hand-crafted feature catalog refined over decades, e.g., band power, connectivity, complexity, and more. Modern EEG foundation models bypass this catalog, learn directly from raw signals via self-supervised pretraining, and match or outperform feature-engineered baselines on most clinical benchmarks. Whether the two representations align is an open question, which we decompose into three sub-questions: what does the model learn, what does the model use, and how much can be explained. We answer them with layer-wise ridge probing, LEACE-style cross-covariance subspace erasure, and a transparent classifier benchmarked against a random-feature baseline. The audit covers three foundation models (CSBrain, CBraMod, LaBraM), five clinical tasks (MDD, Stress, ISRUC-Sleep, TUSL, Siena), and a 6-family 63-feature lexicon. Of the 945 (model, task, feature) units, 648 (68.6%) are representation-causal and 199 (21.1%) are encoded-only. Across tasks, 50 features qualify as universal candidates with strong support (all three architectures RC) in two or more tasks. Frequency-domain features dominate, but the other five families each contribute substantial causal mass. Confirmed features recover, on average, 79.3% of the foundation model's advantage over the random baseline, with a clean task gradient (MDD \approx 0.99 down to Stress \approx 0.56): tasks near ceiling are almost fully recovered by the lexicon, while harder tasks leave a non-trivial residual that pinpoints a concrete target for future concept discovery.
PaperID: 5258, Poster
Abstract: Large language models trained on vast corpora inherently risk memorizing harmful content that may later re-emerge in their outputs. To mitigate this issue, existing unlearning methods typically rely on training-based parameter updates, such as gradient ascent and its variants, to delete targeted content while preserving other knowledge. However, balancing the competing goals of forgetting and retention makes hyperparameter choices for these methods particularly difficult, often requiring repeated tuning to obtain a strong model that still leaves substantial room for improvement and transfers poorly across models and datasets. To address this challenge, we investigate whether unlearning runs exhibit exploitable structure in weight space, and observe that models from different runs still lie in a shared evaluation-performance basin. This suggests that stronger models may be recovered through an unlearning-tailored soup strategy, reducing the need for repeated tuning for further improvement or new settings. Motivated by this, we propose , a unified framework that provides two strategies: EfficientSoup uses binary-search-based interpolation to quickly discover a well-performing model in the early stage, where repeated tuning would otherwise make strong model selection costly. PerformanceSoup uses reweighted souping to efficiently unlock the remaining performance potential in the later stage, where repeated tuning becomes increasingly inefficient. Extensive experiments across diverse datasets and models show that delivers 2.4× to 3.3× efficiency gains in hyperparameter selection, while consistently improving performance across settings.
Abstract: Reliable quantum control in the presence of decoherence requires policies that combat the effect of environmental noise on the controlled dynamics. Open quantum systems under continuous monitoring generate classical measurement records whose drift depends on the noise experienced by the system; the records of two evolutions sharing the same decoherence channels differ only in this drift, so Girsanov's theorem yields a closed-form, differentiable estimator of the KL divergence between their trajectory distributions. We instantiate this estimator with two physically motivated reference measures, yielding two regularizers that both drive the system toward states where the effects of decoherence are minimal: the DV), which works for all noise models. Both are qualitatively distinct from existing penalties on control fluence or smoothness: they penalize the observable consequences of control on the decoherence channels rather than the control amplitude itself. The regularizers outperform unregularized gradient-based and reinforcement-learning baselines across a range of open quantum systems — including single- and multi-qubit benchmarks and a multi-qubit chain calibrated to a published snapshot of the IBM Kingston processor — along several axes of evaluation: final-state fidelity, robustness to mismatch in the assumed noise model (gains grow from +17 pp at training noise to +27 pp under 2.5× noise mismatch), and occupation of forbidden states. The regularizers reduce infidelity by up to 50%, with ~16% gains on the calibrated IBM Kingston chain. Our code can be found at https://anonymous.4open.science/r/QMaxCal-6E71/.
PaperID: 5260, Poster
Abstract: Meta-learned intrinsic rewards are a powerful tool for shaping policy learning in reinforcement learning, particularly when extrinsic rewards are sparse, delayed, or unavailable. Learning Intrinsic Rewards for Policy Gradients (LIRPG) introduced this paradigm and serves as the foundation on which subsequent meta-learned intrinsic reward methods are built. While LIRPG and LIRPG-based methods have shown strong results in low-dimensional control, their behavior in high-dimensional domains, where the policy and value networks use a shared encoder, remains poorly understood. In this work, we analyze meta-learned intrinsic rewards in high-dimensional environments and uncover a consistent failure mode, particularly when training with intrinsic rewards alone, where performance collapses to near-random behavior. We identify the root cause as dense, non-stationary intrinsic rewards inducing large and high-variance value losses that dominate shared encoder updates, suppressing policy learning. We further demonstrate that decoupling policy and value optimization using phasic policy gradient methods is one simple and effective approach to addressing this issue.
PaperID: 5261, Poster
Abstract: Curvature adaptivity is a classical theme in online optimization: for convex Lipschitz losses, adaptive methods interpolate between the optimal O(\sqrtT) regret for general convex losses and O(\log T) regret under strong convexity. Recent work has shown that Follow-the-Perturbed-Leader (FTPL) achieves optimal O(\sqrtT) regret even for online non-convex Lipschitz losses, assuming access to an approximate offline-optimization oracle, but these guarantees do not exploit curvature. We show that FTPL can be made curvature-adaptive in the non-convex setting, without knowing in advance how curvature will accumulate over time. Our algorithm replaces the fixed perturbation scale of standard FTPL with a time-varying scale chosen using only past information. We give a simple follow-the-leader tuning rule for this scale and show that it competes, up to constants, with the best choice in hindsight. The resulting method achieves O(\sqrtT) regret for arbitrary non-convex Lipschitz losses and improves as cumulative curvature grows; with sufficiently accurate oracle calls, it achieves O(\log T) regret when cumulative curvature grows linearly, which includes the classical strongly convex regime. We complement these upper bounds with matching lower bounds for prescribed cumulative-curvature sequences, already for one-dimensional convex losses, showing that the tradeoff between worst-case non-convex regret and curvature-driven fast rates is intrinsic.
Abstract: Neural posterior estimation has emerged as a powerful tool for amortized inference, with growing adoption across scientific and applied domains. In many of these applications, the conditioning variable is a set of observations whose elements depend not only on the target but also on unknown factors shared across the set. Optimal inference therefore requires treating the set jointly, which in turn requires training the estimator at the deployment set size---a regime where memory and compute quickly become prohibitive. We introduce a simple, theoretically grounded strategy that decouples representation learning from posterior modeling. Our method trains a mean-pool Deep Set on sets of size at most two, producing an encoder that generalizes to arbitrary set sizes. The inference head is then finetuned on pre-aggregated embeddings, making training cost essentially independent of the deployment set size N. Across scalar, image, multi-view 3D, molecular, and high-dimensional conditional generation benchmarks with N in the thousands, our approach matches or outperforms standard baselines at a fraction of the compute.
PaperID: 5263, Poster
Abstract: The integration of AI Overviews into search engines enhances user experience but diverts traffic from content creators, potentially discouraging high-quality content creation and causing user attrition that undermines long-term search engine profit. To address this issue, we propose a game-theoretic model of creator competition with costly effort, characterize equilibrium behavior, and design two incentive mechanisms: a \emphcitation mechanism that references sources within an AI Overview, and a \emphcompensation mechanism that offers monetary rewards to creators. For both cases, we provide structural insights for profit-maximizing mechanisms. Evaluations parameterized by real click data show that although AI Overviews harm long-term search engine profit, interventions based on our proposed mechanisms can increase long-term profit across a range of realistic scenarios, pointing toward a more sustainable trajectory for AI-enhanced search ecosystems.
PaperID: 5264, Poster
Abstract: Large language models (LLMs) have demonstrated strong capabilities in question answering, yet they still frequently suffer from hallucinations on knowledge-intensive tasks. Knowledge graphs (KGs) provide LLMs with structured, interpretable, and updatable factual grounding, making them a promising external knowledge source for reliable reasoning. However, existing LLM-guided graph reasoning methods typically rely on hop-wise greedy or beam-style pruning during evidence retrieval. Such local decision processes are inherently myopic: evidence that appears weak near the source may become crucial only after deeper graph context is explored, causing answer-critical branches to be discarded prematurely and making the reasoning chain difficult to recover. To address this limitation, we propose Foresight-over-Graph (FoG), a foresight-aware evidence retrieval framework for knowledge base question answering (KBQA). FoG iteratively constructs a question-relevant evidence subgraph and uses far-to-near feedback to guide path exploration, and maintains a compact memory subgraph to support continued exploration. Extensive experiments on widely used KBQA benchmarks demonstrate that FoG achieves state-of-the-art performance, with a particularly large improvement of 16.58% in Hit on CWQ, while also reducing LLM calls and token usage. Our code is available at \urlhttps://anonymous.4open.science/r/FoG-7273 .
PaperID: 5265, Poster
Abstract: Boolean satisfiability (SAT) provides a foundation tool for cryptographic key recovery by encoding ciphers into CNF or ANF representations. However, heuristic solvers and their neural-enhanced variants often degrade significantly on cryptographic instances due to pronounced structural symmetry, a large number of intermediate variables, and complex algebraic dependencies. Existing neural approaches—ranging from end-to-end prediction to solver-integrated heuristics—face challenges in scalability, assignment accuracy, or computational efficiency. In this paper, we propose a unified neural-guided enumerative SAT framework for cryptographic key recovery. It performs a single forward neural guidance to identify a subset of k critical variables, followed by enumeration over their assignments before invoking a heuristic SAT solver. This design effectively reduces the combinatorial search space while keeping neural overhead minimal. We further introduce a taxonomy of neural guidance across data and model capability regimes, including supervised selection with cryptographic priors, uncertainty-driven selection with assignment prediction, and robust fallback strategies using the model trained on general SAT datasets. Experiments on 10 SAT solvers and 2 cube-and-conquer strategies over SAT4CryptoBench and SAT Competition benchmarks demonstrate up to 5× speedup and approximately 2× average improvement on cryptographic instances, while maintaining better performance on general SAT datasets. These results highlight the effectiveness and generalization of this framework, illustrating its potential to bridge machine learning and SAT solving in cryptanalysis tasks.
PaperID: 5266, Poster
Authors: Qinhan Hou, Jing Tang
Abstract: Graph neural ordinary differential equations provide continuous-time propagation on graphs, but their long-time behavior is largely determined by the asymptotic structure of the mixing operator. We show that diffusion-style Graph ODEs with positive irreducible mixing face a , where node features converge to a rank-one consensus subspace. This reveals a structural limitation of positive diffusion-style continuous graph propagation: within this consensus class, exact long-time avoidance requires changing the effective support; otherwise, smooth suppression should be interpreted as finite-horizon or metastable mitigation. We address this by introducing a for topology evolution, where candidate edges carry latent bistable potentials in a double-well landscape. A node-feature-conditioned force field evolves these potentials, and a smooth sigmoid gate converts them into differentiable propagation weights. We prove the consensus trap theoretically, show that idealized support separation yields block-wise rather than global consensus, and evaluate the framework on real-world benchmarks.
PaperID: 5267, Poster
Authors: GRZEGORZ LISOWSKI, Georgios Papasotiropoulos, Grzegorz Pierczyński, Krzysztof Rogowski
Abstract: This work proposes a framework for public decision processes with a monetary component, where voters can pledge donations to support preferred projects and influence the allocation of public funds. It captures a range of real-world scenarios, including participatory budgeting, blockchain-based funding platforms, charitable programs, and crowdfunding campaigns. Through both theoretical and experimental evaluation, our study addresses questions concerning the existence, structure, number, computation, and quality of Nash equilibria.
PaperID: 5268, Poster
Abstract: is a classical problem where a questioner must identify a hidden item by asking binary questions. Although simple in form, the game is strategically challenging, especially when the questions are expressed in natural language. It is also a useful abstraction for realistic information-seeking tasks, such as medical diagnosis and troubleshooting. Existing approaches often rely on simplifying assumptions that degrade performance. This is an issue with serious implications in high-stakes applications. In this work, we introduce (SLS), an adversarial counterpart of Twenty Questions, and formalize it and its variants as two-player zero-sum extensive-form games. We then propose (GoT), a framework that applies game-theoretic techniques to approximate a Nash equilibrium (NE) strategy for the restricted variant of the game. Empirical results demonstrate that our approach consistently improves worst-case performance compared to (1) direct prompting-based methods and (2) heuristic-guided search methods across all tested settings.
PaperID: 5269, Poster
Abstract: Polysemanticity is central to mechanistic interpretability, yet existing evaluations rely primarily on top-activating examples and qualitative inspection. These approaches are sensitive to dataset coverage and miss distribution-wide functional behavior. We model polysemanticity as functional heterogeneity: a feature is polysemantic when it behaves like different semantically meaningful functions on different subsets of the data. We introduce PFAD (Polysemanticity via Feature Affine Decomposition), a method that decomposes a scalar feature into a hard-routed mixture of affine experts in a vision-language embedding space, interpreting each component through aligned text-image retrieval. We further propose the Effective Decomposition Rank (EDR), a redundancy-aware measure of how many active, non-redundant semantic components explain a feature. We evaluate PFAD on controlled benchmarks with known mixtures and apply it to real vision features. PFAD recovers known component structure and reveals polysemantic behavior invisible to top-activation inspection. Our results suggest that polysemanticity is better understood as distribution-wide functional heterogeneity rather than exemplar diversity.
PaperID: 5270, Poster
Authors: Enis Chenchene, Anil S Baslamisli, Paolo Braca
Abstract: Foundation models produce high-dimensional latent representations in which suitably designed statistical decision rules can be effective yet analytically tractable. We study such detectors in the high-dimensional regime, where latent dimension and training sample size grow proportionally. For linear and quadratic statistics, we derive explicit large-deviation characterizations of false-alarm and miss-detection performance under the Neyman--Pearson paradigm, and show how the binary analysis extends to multi-class classification. The resulting error probabilities and corresponding decay rates exhibit non-monotonic dependence on the shape parameter, including double-descent and, in some regimes, multiple-descent. Experiments on vision, text, and audio datasets using modern foundation models show close agreement between theoretical predictions and empirical performance. In particular, the proposed framework allows us to set a prescribed false-alarm level and accurately predict the corresponding miss-detection probability. Overall, our results provide a theoretically grounded and interpretable framework with competitive empirical performance across different modalities.
PaperID: 5271, Poster
Abstract: Autoregressive (AR) modeling is arguably the only way to extend a fixed-length video diffusion model to longer or even streaming videos. However, existing AR diffusion methods use only the previously predicted chunk as context for the next, discarding the intermediate predictions that constitute the chunk's generation history. In contrast, conventional AR models such as LLMs condition on the entire history of preceding tokens. To this end, we propose Trajectory Forcing (TraF), an AR diffusion method that exploits the denoising trajectories of previous video chunks as context. TraF operates under two training regimes: (1) with ground-truth videos, where pseudo-trajectories are constructed from Gaussian interpolations and re-predicted by the model, and (2) with a teacher model, where real trajectories are obtained from the student's own denoising rollout. The first regime naturally provides a strong initialization for the second. On the 5-second VBench benchmark, TraF achieves a Total score of 85.15, outperforming Self Forcing (+0.84) and Causal Forcing (+1.11) under the same training configuration, with Quality and Semantic improving simultaneously. On 30-second generation — 6× the training horizon — TraF achieves the highest overall quality among all compared methods.
PaperID: 5272, Poster
Abstract: Open-Set Test-Time Adaptation (OSTTA) operates on non-stationary target streams where covariate and semantic shifts coexist. Existing TTA methods often obtain adaptation signals by updating model parameters or maintaining mutable memories, which can be costly and vulnerable to unknown samples. We present TALOR, a lightweight, backward-free rectification module that uses the source linear head as a geometric anchor. The head-induced basis decomposes each normalized feature into high-gain principal coordinates and low-gain tail coordinates. TALOR estimates a soft-routed affine correction in the principal subspace, combining a _regime-level_ principal bias with a tail-conditioned slope that captures _sample-level_ tail-to-principal coupling. Using weighted ridge regression, it subtracts this drift estimate from the principal coordinates, reconstructs the feature and feeds it into the source head. On the ImageNet-C benchmark with six csOOD datasets, TALOR outperforms the next-best UniEnt by 2.2% H-score while running \mathbf2.5× faster and using only 12% of its GPU memory; as a post-correction plug-in, it further boosts COME to 68.4% H-score with only 2% additional memory overhead.
PaperID: 5273, Poster
Abstract: Effective data valuation is essential for machine learning, especially for CLIP pretraining, where models are trained on massive and inherently noisy image-text datasets. Existing CLIP data selection methods often combine heuristic signals such as alignment, target similarity, and diversity, but lack a unified principle for estimating sample contribution. This can lead to inconsistent performance across data regimes and obscures how each sample contributes to downstream performance. In this work, we propose InfCLIP, a principled influence-based approach for data selection in CLIP pretraining. InfCLIP estimates the contribution of each training sample to a target task by leveraging feature representations and temperature-rescaled softmax distributions, enabling efficient computation without retraining or access to model internals. This formulation provides a unified view of alignment, informativeness, uniqueness, and target relevance within a single framework. Empirically, InfCLIP consistently outperforms existing data selection methods across a wide range of benchmarks, including zero-shot evaluation on ImageNet and distribution-shifted datasets. We also show that Self-InfCLIP enables efficient training-set analysis, including noise detection and memorization-aware data assessment, without additional retraining.
Abstract: Predicting labels for graph-structured data is crucial in many scientific applications. Recently, Gaussian processes (GPs) with graph-level inputs have been proposed as flexible, non-parametric models for classification tasks. In this work, we extend this framework to regression tasks and simplicial complexes (SCs), enabling edge-level attributes and attributes supported on higher-order simplices. Drawing on the rich literature on Hodge theory in machine learning, we enhance the resulting SC representations via the Hodge decomposition, capturing homological information such as holes. We introduce Hodgelet representations as rich, learnable topological descriptors of simplicial complexes and show that our framework improves predictions across various applications. This paves the way for broader use of GPs in graph-level and SC-level prediction tasks.
Abstract: Time series foundation models (TSFMs) are transforming the forecasting paradigm through large-scale cross-domain pretraining. However, most existing TSFMs remain univariate, and recent efforts to enable cross-variate modeling still operate directly within the raw variate space. This design introduces fundamental limitations in semantic alignment and relational expressivity. Specifically, raw-space group mixing lacks a dedicated mechanism to align heterogeneous physical quantities, while standard non-negative attention fails to capture the complex synergistic and antagonistic interactions ubiquitous in real-world systems. To address these challenges, we propose Falcon-X, decouples variates from the raw space and maps them into a unified latent prototype space. Falcon-X employs a Unified Prototype Diff-Attention mechanism that explicitly evaluates both positive and negative semantic affinities to explicitly align heterogeneous variates. Cross-variate interactions are then efficiently performed within this shared space via Latent Entity Attention, naturally facilitating zero-shot structural transfer. Finally, a Variate Reassembly Router robustly reconstructs variate-specific trajectories via a request-and-dispatch mechanism. Extensive evaluations on the GIFT-Eval and fev-bench benchmarks demonstrate that Falcon-X achieves state-of-the-art forecasting performance, offering a principled and scalable paradigm for complex multivariate environments. Code is available in the supplementary material.
PaperID: 5276, Poster
Authors: Moonwon Choi, Kisung Nam, Seunggeun Lee
Abstract: Optimizing a proxy objective can improve human-perceived quality, but the resulting gain can vary with optimization strength. Average proxy--human correlation alone does not reveal how human-perceived quality changes as optimization strength increases. We propose a pre-optimization diagnostic that uses repeated human ratings to predict the human-gain curve induced by a fixed proxy. The method models proxy optimization as a KL-regularized exponential tilt. At KL radius d, the local gain is approximated by A\sqrtd+Bd, where A measures first-order proxy--human alignment and B captures the leading second-order correction along the proxy-optimization path. We estimate A and B from repeated human ratings on calibration items and test the resulting gain-curve prediction on held-out evaluation items. Across SummEval, analysis-eligible WMT MQM clusters, an OpenMEVA-MAGS ROC/WP repeated-rating slice, and a GPT-4o judge proxy experiment, the proposed approach consistently reduces held-out gain-prediction error relative to A-only, zero-gain, and reliability-only baselines. Transfer experiments across datasets and evaluation slices show that source coefficients are useful priors, but target calibration further improves prediction. The method provides a principled local diagnostic for estimating the human-gain curve of a fixed proxy as optimization strength varies.
PaperID: 5277, Poster
Authors: Ahmad Abdel-Azim, Xihong Lin
Abstract: Individual-level genotype data underpin genetic discovery, disease-risk modeling, and therapeutic development, yet many critical settings remain data-limited, including underrepresented ancestries, rare or imbalanced disease cohorts, and restricted-access biobanks. Conditional synthetic genome generation would support privacy-aware benchmarking, simulation, and data augmentation, but requires learning an ultra-high-dimensional discrete distribution over hundreds of thousands to tens of millions of correlated single-nucleotide polymorphisms (SNPs), or sites of DNA variation. We introduce \mathbfHELICS: \mathbfHigh-fidelity g\mathbfEneration via \mathbfLatent \mathbfInterpolant \mathbfConditional \mathbfSynthesis, a phasing-free latent flow matching framework for genome-wide SNP-level generation from biobank-scale genotype data. HELICS tokenizes chromosomes with convolutional autoencoders, reducing >600\rm K SNPs to 4,096 continuous latent tokens, and trains a transformer parameterized conditional flow matching model over the concatenated multi-chromosome latent representation. Trained on roughly 428\rm K genomes in the UK Biobank, HELICS generates genome-wide synthetic cohorts without haplotype phasing and achieves stronger fidelity to held-out real genotypes than existing haplotype- and genotype-based simulators, including the highest validation allele-frequency agreement (R^2=0.992) and superior preservation of local and long-range SNP correlation structure. Conditioning on ancestry and genetic liability summaries enables targeted generation across genetic risk profiles; in perturbation analyses, by increasing the conditioned disease trait liability, HELICS induces SNP-level dosage changes in generated genomes that are highly correlated with published association weights. HELICS provides a scalable generative modeling framework for SNP-level genome data and a path toward conditional synthetic cohorts for data-limited genetic studies.
PaperID: 5278, Poster
Abstract: Deep Equilibrium Models (DEQs) are an established framework for image restoration that learn a problem-adapted regularization by solving a fixed-point (i.e. equilibrium) problem. While flexible and expressive, DEQs are often hindered by high computational cost and training instability. We propose an inertial DEQ (i-DEQ) that learn an explicit nonconvex regularization within the DEQ formulation. By using momentum within the fixed-point iterations, i-DEQ has convergence guarantees and accelerated rates. Moreover, we observe that i-DEQ is significantly more stable during the training and robust to rough initialization than DEQs. Numerical experiments on various linear and nonlinear inverse problems demonstrate that i-DEQ achieves reconstruction quality comparable to state-of-the-art methods, while reducing DEQ's inference time by a factor of two.
PaperID: 5279, Poster
Abstract: Structure-informed protein representation learning is essential for effective protein function annotation and de novo design. With the rapid progress of protein structure prediction and experimental structure determination, a central challenge in the field is shifting from merely obtaining structural models to reliably leveraging predicted or measured structures for downstream learning. However, both crystal and AlphaFold-like predicted structures can contain structural uncertainty, which can bias local neighborhood construction and make robust protein representation learning challenging. To address these issues, we propose a novel equivariant Transformer-State Space Model (SSM) hybrid framework, termed E^3former, designed for efficient and robust protein representation. Our approach uses energy function-based receptive fields to construct proximity graphs that adaptively stabilize neighborhood selection under structural deviations, and incorporates an equivariant high-tensor-elastic selective SSM within the transformer architecture. These components allow the model to adapt to complex geometric interactions and extract structural features with a higher signal-to-noise ratio. Empirical results demonstrate that our model outperforms existing methods in structure-intensive tasks, such as inverse folding and binding site prediction, particularly when using predicted structures, owing to its enhanced tolerance to data deviation and structural uncertainty. Our approach offers a novel perspective for conducting biological function research and drug discovery using imperfect but increasingly available protein structure data. Our code is available on \urlhttps://anonymous.4open.science/r/E3former-207E.
Authors:
Jian Xie, Tianhe Lin, Zilu Wang, Yuting Ning, Yuekun Yao, Tianci Xue, Zhehao Zhang, Zhongyang Li, Kai Zhang, Yufan Wu, Shijie Chen, Boyu Gou, Mingzhe Han, Yifei Wang, Vint Lee, Xinpeng Wei, XJ Wang, Yu Su, Huan SunAbstract: Deep research agents extend the role of search engines from retrieving keyword-matched pages to synthesizing knowledge, fundamentally changing how humans interact with information. However, frontier systems remain proprietary, while existing open agents often generalize poorly across different task types, leaving unclear how to train a broadly capable deep research agent. We release Quest, a family of open models (ranging from 2B to 35B) that serve as general-purpose deep research agents designed to handle a wide range of search tasks, with strong capabilities in fact seeking, citation grounding, and report synthesis. To build Quest, we propose an effective training recipe combining mid-training, supervised fine-tuning, and reinforcement learning. Central to this recipe is a curated data synthesis pipeline based on unified rubric trees, which applies to different task types and enables synthesizing training data with verifiable rewards without human annotation. In addition, Quest incorporates a built-in context management mechanism that enables effective long-horizon reasoning and knowledge synthesis. Using only 8K synthesized tasks, Quest approaches or even surpasses frontier closed-source agents across eight deep research benchmarks spanning diverse task types, and achieves the best overall performance among recent open-weight agents. Notably, Quest-35B achieves 30.7% on Mind2Web 2, outperforming OpenAI DeepResearch (28%) and Gemini DeepResearch (18%). We plan to release all the resources for further research.
Abstract: Optimizing nonlinear preferences in multi-objective reinforcement learning (MORL) is essential for capturing complex trade-offs like risk aversion or fairness. However, such non-linearity has historically bifurcated nonlinear MORL objectives into two distinct paradigms: Scalarized Expected Return (SER) and Expected Scalarized Return (ESR). While SER requires global-level optimization and ESR requires non-Markovian policies, leading to fragmented optimization strategies, we bridge this divide through the Aggregation–Expectation–Transformation (AET) framework. By unifying both criteria through a tripartite decomposition of scalarization, AET provides a principled foundation for general nonlinear MORL. Building on this framework, we propose AETDICE, a tractable offline RL algorithm for AET objectives. By utilizing DICE-style density-ratio estimation in an augmented state space, AETDICE enables sample-based optimization from static datasets. Our framework resolves long-standing barriers and captures respective trade-offs induced by AET framework, which existing methods fail to address.
PaperID: 5282, Poster
Authors: Seohyun Kim, Jaeyo Chang, Dong-Joon Lim
Abstract: Periodic structure is a fundamental basis for long-term time series forecasting (LTSF), but real-world periodic patterns are not simple repetitions. Variations in level, amplitude, and phase within one period can accumulate across subsequent periods and reshape future periodic structures. We therefore argue that explicitly modeling inter-period relations is a core challenge in LTSF. Existing methods often handle these relations implicitly or rely on unconstrained mixing, where meaningful dependencies can be entangled with noisy interactions. Simply increasing mixing expressiveness does not resolve this issue. Motivated by this, we address LTSF through gradual state refinement driven by interactions among periods. The key challenge is to reinforce meaningful inter-period dependencies while suppressing noise and maintaining stable representations during repeated refinement. To this end, we draw inspiration from Hamiltonian dynamics and use its split-state formulation and symplectic-style updates as an architectural bias for stable iterative refinement. We propose HAMSTAR (HAMiltonian STructured period-Aligned Representation for time series), a Hamiltonian-inspired framework for structured period-aligned time series forecasting. HAMSTAR represents time series as period-aligned latent states, decomposes them into split-state components, and uses structured inter-period coupling to transform period relations into refinement signals. Symplectic-style updates then progressively refine these states, enabling stable representation dynamics. Experiments on standard LTSF benchmarks show that HAMSTAR achieves state-of-the-art performance across diverse datasets, and further analyses demonstrate meaningful inter-period modeling and stable latent dynamics.
Abstract: Gaussian process (GP) bandits provide a powerful framework for performing blackbox optimization of unknown functions. The characteristics of the unknown function depend heavily on the assumed GP prior. Most work in the literature assume that this prior is known but in practice this seldom holds. Instead, practitioners often rely on maximum likelihood estimation to select the hyperparameters of the prior - which lacks theoretical guarantees. In this work, we study two algorithms for joint prior selection and regret minimization in GP bandits based on GP Thompson sampling (GP-TS): Prior-Elimination GP-TS (PE-GP-TS) that disqualifies priors with poor predictive performance, and HyperPrior GP-TS (HP-GP-TS) that utilizes a bi-level Thompson sampling scheme. We theoretically analyze the algorithms and establish a sublinear regret bound for HP-GP-TS. In addition, we demonstrate the effectiveness of these algorithms compared to the alternatives through extensive experiments with synthetic and real-world data.
PaperID: 5284, Poster
Authors:
Mingyuan Li, Guangsheng Yu, Xu Wang, Shaoxiong JiAbstract: Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible access to external knowledge, but adds retrieval latency, context overhead, and only shallow integration with the backbone. Parametric adaptation is efficient at inference time, but entangles knowledge with model weights and can be hard to update, audit, or transfer. Engram-style hashed memory occupies a middle regime: it stores learned information in an external, addressable table, yet consumes that table through a small learned reader. This raises a basic question: when such a memory is moved across backbones, what matters more, the frozen memory itself or the target-side reader? We study this question through cross-model frozen-memory extraction, where a memory trained on a source model is frozen and attached to a different target model while only a lightweight reader is trained. Across architectures, tokenizers, and model scales, frozen memory improves every source--target pair in a 3 × 3 transfer matrix, with up to 15.7% relative perplexity reduction. Ablations show that learned memory content and correct addressing both matter, but stronger readers account for much of the transfer gain; on OpenQA, a dual-layer 4-branch reader nearly closes the gap between same-model and cross-model transfer and reaches 38.78 average. These results position Engram as a hybrid external memory and suggest that, in this regime, progress depends as much on reader design as on storage itself. Code and reproducible setups are available at https://anonymous.4open.science/r/Engram_Extractor/README.md. .
PaperID: 5285, Poster
Authors: Qinan Huang, Weizhi Xue, Weizhou Wang, Mingxuan Cao, Ian Su
Abstract: Accurate prediction of physico-chemical properties in multi-component liquid mixtures is critical for designing next-generation electrolyte battery materials. Despite recent progress in mixture modeling, existing methods often treat each target property independently, and they struggle to capture the shared molecular mechanisms that couple transport, thermodynamic, and structural properties across heterogeneous mixture systems. To address these limitations, we propose \method, a unified geometry-aware framework for joint multi-property prediction in chemical mixtures. \method combines an E(3)-equivariant molecular encoder with property-conditioned prediction heads, enabling a shared representation across diverse mixture tasks while preserving property-specific inductive biases. In particular, \method incorporates a stretched-Arrhenius conductivity head to model temperature-dependent transport behavior and a sparse Mixture-of-Experts readout to adapt to heterogeneous property regimes. Comprehensive experiments on five CheMixHub mixture-property tasks show that \method consistently improves over single-property geometric baselines, with strong gains in low-data and chemically out-of-distribution settings. These results demonstrate that combining joint representation learning with physically informed readouts provides an effective and general framework for reliable property prediction in complex liquid mixtures.
Abstract: 3D human pose estimation from sparse multi-view camera rigs is an essential task for numerous applications, including action recognition, sports analysis, and human-robot interaction. While learned methods dominate the field on benchmarks, they require large annotated datasets; training-free optimization-based methods remain promising as they circumvent 3D supervision by solving a correspondence problem across views from 2D detections. Existing combinatorial formulations rely on pairwise associations to model this correspondence problem and enforce global consistency across views only as a downstream constraint. However, reconciling locally plausible pairwise matches becomes brittle under occlusion and noisy detections, where local errors propagate globally. We propose COMPOSE, which recasts multi-view 3D human pose estimation as a weighted exact-cover optimization over a hypergraph of person hypotheses. Our formulation replaces pairwise association and post-hoc consistency enforcement with a single global combinatorial objective. To address the exponentially large candidate space, we introduce a geometric pruning strategy alongside two complementary solvers: an exact Integer Linear Programming formulation and a scalable relaxation via Belief Propagation. Without any 3D supervision, COMPOSE improves average precision by up to 31 points over the best optimization-based method and 13 points over self-supervised learned methods, demonstrating the effectiveness of higher-order combinatorial association for training-free multi-view 3D human pose estimation.
PaperID: 5287, Poster
Abstract: Scaling robot demonstrations does not guarantee better imitation: redundant trajectories dilute learning, while rare, imperfect segments may encode critical transitions. Prior curation methods overlook such dependencies by scoring samples independently at the trajectory or segment level. We reframe robot demonstration curation as a coverage-aware sequential selection problem and propose PROSE, which selects data by maximizing marginal utility under a fixed budget. PROSE unifies three signals: (i) influence on closed-loop return, (ii) kinematic reliability, and (iii) coverage gain measuring novelty with respect to the retained set, allowing full trajectories and fine-grained segments jointly participate in scoring and filtering. Across three Robomimic simulation tasks and three real Franka tasks, PROSE achieves the best success rate on every task we evaluate against three state-of-the-art curators (CUPID, SCIZOR, Demo-SCORE), with average closed-loop gains of +15% in simulation and +35% in real-world. Ablations attribute the gain primarily to the coverage step.
PaperID: 5288, Poster
Abstract: We study online learning in linear dynamical games where multiple strategic agents act on a shared state evolving according to a linear dynamical system subject to adversarial disturbances. This setting lies beyond both single-agent nonstochastic online control and classical linear-quadratic games, which typically focus on quadratic objectives and noiseless or stochastic dynamics. Each agent seeks to minimize its own sequence of convex losses under full-state observability. Following the stabilizing-baseline paradigm in nonstochastic online control, we assume access to linear controllers that stabilize the noiseless system and focus on online adaptation by learning disturbance-action corrections. This extends adversarial online control to strategic multi-agent shared-state systems, where each agent's actions shape the state trajectory and hence the realized losses of all learners. Under state-only and aggregate-input feedback models, we analyze agents running online gradient descent with memory to update their own disturbance-action policies. We prove per-agent regret bounds that are sublinear and near-optimal in the time horizon, and show how their dependence on the number of agents varies with the feedback available to each learner. In the common-interest case, where all agents have identical cost functions, we show that the induced online problem forms a time-varying potential game and derive equilibrium-tracking guarantees. Together, these results provide a theoretical framework for adversarial online learning in stabilized linear dynamical games, connecting online control with learning in games.
PaperID: 5289, Poster
Abstract: Generative solvers for combinatorial optimization benefit from additional test-time compute when their refinement states remain aligned with task relaxations. We introduce RAGenCO, a relaxation-aligned test-time scaling framework for generative combinatorial optimization. The central idea is to treat projection onto a task relaxation as a state control primitive: refinement first maintains a projected control state, and only then forms the task-dependent decoder input or expands candidates. For TSP and ATSP, Sinkhorn projection aligns matrix states with the cycle-cover relaxation during refinement and decoding. For MIS and MVC, Bernoulli projection controls refinement on the independent-set side, while raw logits are retained for terminal greedy decoding because their ordering carries the decoder signal. Best-of-B branching then reallocates compute by expanding candidates around aligned control states rather than sampling from unconstrained logits. Because these operations are used only at inference, the diffusion training objective remains unchanged. Under equal runtime budgets, this alignment of control state, decoder input, and branching yields a strong quality-runtime trade-off on TSP and ATSP and competitive results on MIS and MVC.
PaperID: 5290, Poster
Abstract: Large Vision-Language Models (LVLMs) are prone to hallucinations, producing responses that are inconsistent with visual inputs. While inference-time activation steering provides a lightweight solution without retraining, existing methods are either limited by coarse global steering vectors or rely on unconstrained full-dimensional vector prediction, which can produce unreliable vector for unseen hallucination patterns. In this paper, we propose Correction Space Steering (CSS), an robust activation steering framework that predicts correction coordinates in a learned low-dimensional correction space. We first conduct an empirical analysis of hallucination-related activation differences and find that they are predictive of hallucination types, form type-aware geometric structures, and exhibit low effective ranks in most layers. We further provide a theoretical analysis showing that these activation differences can be approximated by a low-dimensional subspace. Based on this insight, CSS learns a low-dimensional correction space from activation differences and trains an online predictor to infer input-specific correction coordinates from inference-time hidden states. The resulting correction vectors are then reconstructed within the correction space and applied to steering activations for hallucination mitigation. Experiments on standard hallucination benchmarks show that CSS effectively reduces hallucinations without retraining LVLMs and outperforms existing inference-time mitigation methods.
PaperID: 5291, Poster
Abstract: Brain–computer interfaces (BCIs) must continually adapt as neural signals drift, yet labeled calibration data is expensive to collect at time of use. Recent systems address this challenge by updating neural decoders using language-model–pseudo-labels, but they treat the pseudo-label as ground truth. We reinterpret recalibration as approximate Bayesian filtering: the decoder parameters are the latent states, and language models provide observations in the form of unnormalized potentials. Guided by this analysis, we introduce B-CORP (Bayesian Continual Online Recalibration with Pseudo-labels). B-CORP replaces hard pseudo-labels with a soft top-K marginalization, increases the number of iterations in the coordinate ascent on the joint likelihood, and models parameter drift with a prior on latent dynamics. Across synthetic drifting regression tasks and real BCI benchmarks, including handwriting and speech BCI datasets, B-CORP attains word error rates consistently under 3% on a difficult handwriting BCI dataset. B-CORP improves almost 30% versus the state-of-the-art, while retaining performance even with poor pseudo-labels where naive methods fail. More broadly, our approach connects pseudo-label recalibration with online Bayesian inference, providing a principled foundation for problems requiring long-term recalibration with access to black box predictive models, such as LLMs, for use as strong priors over the labels.
Abstract: Recent advancements in zero-shot reinforcement learning (RL) have facilitated the extraction of diverse behaviors from unlabeled, offline data sources. In particular, forward-backward algorithms (FB) can retrieve a family of policies that approximately solves any standard RL problem (with additive rewards, linear in the occupancy measure), given sufficient capacity. While retaining zero-shot properties, we tackle the greater problem class of RL with general utilities, in which the objective is an arbitrary differentiable function of the occupancy measure. This setting is strictly more expressive, capturing tasks such as distribution matching or pure exploration, which may not be reduced to additive rewards. We show that this additional complexity can be captured by a novel, maximum entropy (soft) variant of the forward-backward algorithm, which recovers a family of stochastic policies from offline data. When coupled with zero-order search over compact policy embeddings, this algorithm can sidestep iterative optimization schemes, and optimizes general utilities directly at test-time. Across both didactic and high-dimensional experiments, we demonstrate that our method retains favorable properties of FB algorithms, while also extending their range to more general RL problems.
PaperID: 5293, Poster
Abstract: Medication recommendation requires generating drug combinations that are therapeutically effective while minimizing harmful drug-drug interactions (DDIs). Both objectives are rooted in the combinatorial nature of prescriptions. Existing discriminative methods predict drugs independently, neglecting inter-drug dependencies; autoregressive methods introduce sequential dependencies but impose arbitrary generation orders and accumulate errors. Both paradigms confine DDI mitigation to training-time penalties, offering limited capacity to regulate interactions during inference for individual prescriptions. We reformulate medication recommendation as a discrete diffusion process and propose RxDiff, which leverages iterative denoising for bidirectional and revisable generation and provides a natural interface for dynamic DDI regulation at each step. A clinically grounded dual-guided forward process and an inference-time interaction-aware modulation mechanism are designed to jointly enable RxDiff to achieve consistent accuracy improvements, lower and more individually controlled DDI rates, and robustness to evolving DDI knowledge across MIMIC-III, MIMIC-IV, and a real-world inpatient dataset. Code is available at https://anonymous.4open.science/r/RxDiff-9157.
Abstract: Different age-related regulations have been proposed to protect minors from harmful content and interactions online. Automated age estimation is central to enforcing such regulations, and vision-language models (VLMs) achieve state-of-the-art performance on this task. However, we find that the zero-shot nature of VLM-based age estimation produces an unexpected side effect we call the identity shortcut: Instead of estimating age from visual features, VLMs tend to identify the depicted person and infer their age from memorized knowledge. This phenomenon leads to substantially incorrect predictions when non-celebrities are misidentified as celebrities. It also produces deceptively high robustness to noise and adversarial perturbations on celebrity images, which dominate popular benchmarks. To mitigate this, we propose an activation steering method that suppresses the shortcut by intervening on the hidden states of the VLM. This method improves age estimation accuracy for both memorized and unseen identities, reducing mean absolute error by up to 25% across popular benchmarks.
Abstract: We study fixed-confidence best arm identification in generalized linear bandits under a hybrid feedback model: at each round, the learner may query either (i) absolute reward feedback from a single arm or (ii) relative (dueling) feedback from an arm pair, both governed by generalized linear models. We introduce a likelihood-ratio–based confidence sequence that unifies heterogeneous generalized linear observations and yields an explicit ellipsoidal confidence set under a self-concordance assumption. Building on this confidence set, we propose a hybrid Track-and-Stop algorithm that adaptively allocates queries by tracking a minimax-optimal design over a joint action space of arms and pairs. We establish \delta-correctness and provide high-probability upper bounds on the stopping time. We further extend the framework to a cost-aware setting that accounts for heterogeneous acquisition costs across feedback modalities. Empirical experiments demonstrate that the proposed algorithms significantly improve sample efficiency over baseline methods.
Abstract: Classical clustering methods usually return either a finite partition of the observed data or a finite dendrogram over it. This finite-sample view is inadequate when the hierarchy of interest is a recursive geometric object with fine-scale refinements that continue beyond the levels directly observed. We introduce classification fields: infinite-depth hierarchical cluster structures on \mathbbR^d generated by a local parent-to-child refinement rule. A classification field generator maps each parent centre to an ordered, bounded, and separated tuple of child residuals. Together with a root and a scale factor, this rule recursively generates cluster centres, Voronoi cells, and a metric DAG encoding the hierarchy. Given only a finite prefix of such a hierarchy, we learn a classification field predictor that approximates the generator and can be rolled out to unseen depths. We prove exponential truncation convergence in the completed cell metric and ReLU realizability with width (O(\varepsilon^-\gamma) and depth \widetilde O(\varepsilon^-3\gamma/2), where \gamma=\log K/(-\log s), up to finite-window aspect-ratio factors. The approximation holds at the level of the induced compact metric structures, measured in the completed cell-metric Hausdorff distance. Experimental validation on matched CFG-generated hierarchies, IFS fractals, and image-induced recursive clustering hierarchies shows that learned predictors preserve ordered child slots, unordered geometry, and hierarchy-level path metrics under recursive rollout. These results support the claim that finite hierarchical observations can reveal local refinement rules capable of generating substantially deeper classification fields.
PaperID: 5297, Poster
Abstract: As large language models (LLMs) face growing demands for data removal driven by regulatory compliance, copyright concerns, and privacy protection, approximate unlearning has emerged as a practical alternative to expensive full retraining. In practice, pre- and post-unlearning model snapshots are often maintained for auditing and rollback, introducing a previously overlooked side channel. This paper demonstrates that approximate unlearning in LLM can leave residual traces within model snapshots, which may be recoverable and thus enable the privacy breaches that approximate unlearning is intended to prevent. Here, we propose TRACE, a framework that reconstructs unlearned training data given access to pre- and post-unlearning model snapshots, without any data-specific prior knowledge. To achieve this, TRACE identifies a compact set of candidate tokens from embedding-level parameter differences to constrain the search space, and assembles sequences via a contrastive scoring mechanism based on the perplexity gap between the two models. A two-stage exploration-then-completion strategy enables the recovery of multiple distinct samples. Extensive experiments on various datasets, including TOFU, MUSE, and WMDP, across multiple LLMs and six representative unlearning methods, demonstrate that TRACE achieves high-fidelity reconstruction. This paper exposes significant privacy risks in current LLM approximate unlearning deployments and highlights the need for defenses against parameter-difference leakage and appropriate management of model snapshot access.
PaperID: 5298, Poster
Abstract: Efficient deployment of large language models (LLMs) on mobile devices requires balancing performance with compute and memory constraints. Post-training compression is an effective approach to managing this trade-off, while requiring much less computation than in-training counterparts. We propose DictLLM, a dictionary-based compression scheme in which weight matrices within grouped layers are decomposed into a shared dictionary and layer-specific sparse coefficients, directly obtained from a pre-trained model without further training. Specifically, we cast DictLLM as a four-step post-training optimization. First, we obtain shared dictionaries via singular value decomposition (SVD) of input-scaled weight matrices, amortizing rank across grouped layers; second, we sparsify the layer-specific coefficients to the target compression ratio using a dictionary-aware Hessian; and third, we refine the dictionaries given the resulting sparse coefficients. Finally, we aggressively quantize both components using factor-specific Hessian information: input covariance for the dictionaries and the dictionary-aware Hessian for the coefficients. The resulting representation reduces memory footprint by storing a single shared dictionary per layer group, along with sparse layer-specific coefficients. On language modeling and zero-shot reasoning benchmarks, DictLLM consistently outperforms existing post-training low-rank compression methods and pushes the Pareto frontier between model size and performance. DictLLM compresses the original LLaMA-7B model by 5.1x with only a 3.4-point increase in WikiText-2 perplexity, achieving a 3.6x higher compression ratio than the state-of-the-art low-rank method.
Authors: Katarzyna Filus, Kamil Faber, Roberto Corizzo, Christopher Kanan
Abstract: Continual learning studies how models can adapt to new tasks while retaining previously acquired knowledge. Although a broad spectrum of methods has been proposed to mitigate catastrophic forgetting, the field remains predominantly performance-driven, with limited insight into what forgetting actually corresponds to within the vision model's representation space. Prior work has primarily analyzed forgetting through task-level performance or coarse measures of representational drift, without disentangling output-level accessibility from changes in finer-grained internal structure.To this end, we propose a diagnostic framework that leverages Sparse Autoencoders (SAEs) to define a task-anchored latent feature space, enabling analysis of how task-specific information evolves at a finer granularity, where individual SAE latents are treated as concept proxies for recurring and relatively disentangled visual patterns in the model’s internal computations. Within this framework, we decompose forgetting into apparent concept deletion, recoverability, and decodability. We show that a large portion of seemingly lost concept-level information can often be recovered under linearity assumption, with concept decodability degrading as more tasks are introduced. Overall, our findings suggest that a significant part of concept-level forgetting can be attributed to changes in the representational accessibility rather than complete information erasure.
PaperID: 5300, Poster
Abstract: Multimodal Large Language Models (MLLMs) have shown promise as interactive agents, yet precise, programmatic 3D graphics editing remains difficult to train because existing systems largely rely on closed-source model APIs and lack deterministic, state-verifiable execution environments. We present BlenderFORGE, a trainable data-environment-optimization framework for open-weight MLLM agents that edit explicit Blender scenes through Blender Python (bpy) code generation. BlenderFORGE targets four fundamental state-verifiable editing task families: object placement, shape-key editing, lighting adjustment, and material modification. It integrates three components: (1) a scalable perturbation-based data generation pipeline built from curated assets and large-scale 3D scene repositories; (2) an isolated Blender Sandbox for deterministic code execution, structured state export, and closed-loop reward computation; and (3) a two-stage training pipeline that uses offline teacher trajectories for cold-start Supervised Fine-Tuning (SFT) and sandbox-computed rewards for multi-turn tool-agentic Reinforcement Learning (AgentLoop-RL). Experiments across the four task families show that the resulting open-weight Qwen3-VL-8B-BlenderFORGE agent substantially improves over its base model and achieves competitive performance with several zero-shot proprietary MLLM baselines.
Abstract: Discrete diffusion models are often said to outperform autoregressive models on planning tasks, but the specific mechanisms enabling this are not yet well understood. To explain why, we abstract the complexity of studying ``planning'' into graph traversal problems, which offer a minimal setup for evaluating language models' planning and lookahead capabilities. We find that non-autoregressive language models are able to leverage an inherent directional asymmetry in lookahead planning tasks to outperform the training efficiency of their autoregressive counterparts. Through a mechanistic study of the training and inference dynamics of autoregressive and non-autoregressive models, we discover that while forward traversal through junctions in the graph requires complex sequential planning, the reverse path remains deterministic. Consequently, non-autoregressive models naturally develop a fundamentally different solution strategy. Effectively, they learn a reverse-decoding pattern which reduces an \ell^th-order lookahead problem to a sequence of simple 1^st-order transitions. Latent-space analysis further confirms these divergent strategies by demonstrating the distinct internal representations formed during generation. Ultimately, these findings highlight the advantages of non-autoregressive modeling while clarifying the inherent planning limitations induced by autoregressive architectures.
Abstract: We study reinforcement learning in hybrid discrete–continuous action spaces, such as settings where the discrete component selects a regime (or index) and the continuous component optimizes within it — a structure common in robotics, control, and operations problems. Standard model-free policy gradient methods rely on score-function (SF) estimators and suffer from severe credit-assignment issues in high-dimensional settings, leading to poor gradient quality. On the other hand, differentiable simulation largely sidesteps these issues by backpropagating through a simulator, but the presence of discrete actions or non-smooth dynamics yields biased or uninformative gradients. To address this, we propose Hybrid Policy Optimization (HPO), which backpropagates through the simulator wherever smoothness permits, using a mixed gradient estimator that combines pathwise and SF gradients while maintaining unbiasedness. We also show how problems with action discontinuities can be reformulated in hybrid form, further broadening its applicability. Empirically, HPO substantially outperforms PPO on inventory control and switched linear-quadratic regulator problems, with performance gaps increasing as the continuous action dimension grows. Finally, we characterize the structure of the mixed gradient, showing that its cross term — which captures how continuous actions influence future discrete decisions — becomes negligible near a discrete best response, thereby enabling approximate decentralized updates of the continuous and discrete components and reducing variance near optimality.
PaperID: 5303, Poster
Abstract: Large language models (LLMs) require frequent knowledge updates across heterogeneous domains, but existing editing methods fail when applied sequentially to multi-task settings. We trace this collapse to inter-domain null-space misalignment: heterogeneous tasks induce distinct Jacobian geometries, causing their common null space to shrink dramatically and leaving insufficient admissible directions for conflict-free updates. We propose the Conflict Index to quantify this geometric interference, and introduce Mu-Edit, which mitigates multi-task conflicts via (1) conflict-aware ordering to minimize cumulative interference, and (2) dynamic low-rank approximation to expand the null space when ordering alone is insufficient. Experiments across five functional domains and three backbones show that Mu-Edit significantly outperforms existing baselines in multi-task editing performance while preserving general capabilities, and remains effective under incremental task arrival.
PaperID: 5304, Poster
Abstract: Tiny object detection remains challenging because tiny instances occupy only a few pixels, and their fine-grained details are easily lost during feature downsampling. Although the highest-resolution feature preserves rich local details for tiny objects, DETR-like detectors often discard it due to its large computational cost, while directly reintroducing this feature also brings dense background noise. To address this problem, we propose ReDHF, a novel framework that Revives the Discarded High-resolution Feature for tiny object detection. ReDHF has three cooperative components. Cross-Scale Deformable Fusion (CSDF) first constructs a semantically enhanced highest-resolution feature and injects its fine-grained details into the shallow feature under a moderate cost. Ellipse-Guided Shallow Supervision (EGSS) then guides the enhanced feature to learn more reliable classification and localization priors by applying geometry-adaptive auxiliary supervision. Full-Scale Query Interaction (FSQI) further enables decoder queries to access both low-level details and high-level semantic information, allowing them to aggregate richer cues for tiny objects. Extensive experiments show that ReDHF consistently improves multiple baselines across different DETR variants and achieves state-of-the-art performance on the evaluated benchmarks. The best ReDHF variants surpass their baselines by +1.4 AP on AI-TOD-v2, +1.6 AP on VisDrone, and +1.5 AP on SODA-D, with especially notable gains on tiny and small objects.
PaperID: 5305, Poster
Abstract: Dynamic scene reconstruction from monocular videos remains highly challenging, as existing methods often struggle to balance global structural coherence and local fine-grained details under limited multi-view cues. To address this challenge, we propose WebSpline, a novel dynamic 3D Gaussian framework that enables structurally coherent and high-fidelity reconstruction from monocular videos with fast rendering. The core of WebSpline is the Structure-Informed Spline (SIS) representation, which models each dynamic Gaussian trajectory using a learnable cubic Hermite spline whose motion is structurally organized with an auxiliary Structural Proxy Graph (SPG). The proposed framework is optimized in two stages: (i) in the first stage, the SPG is initialized from 2D point tracks and refined with temporal rigidity regularization to establish structural coherence for moving objects across the sequence; and (ii) in the second stage, the SIS representation is initialized from the refined SPG and optimized under both spatial and structural neighborhood constraints. At inference, Gaussian motion is obtained solely by evaluating the learned SIS, enabling fast rendering. Extensive experiments on the challenging monocular dynamic scene benchmarks, iPhone and NVIDIA, demonstrate that our WebSpline achieves state-of-the-art rendering quality while rendering over 10x faster than WorldTree, the second-best method on the iPhone dataset.
Abstract: Large language models (LLMs) are costly to deploy due to their large memory footprint and high inference cost. Weight-activation quantization can reduce these costs, but low-bit activation quantization remains difficult because activation outliers induce large quantization error. Recent rotation-based methods address this by applying orthogonal transformations that redistribute activation magnitude across dimensions, but existing approaches either require expensive end-to-end rotation training or rely on stored activation corpora, introducing significant compute or storage overhead. We propose a lightweight post-training rotation calibration method for LLM activation quantization. Our method learns orthogonal rotations that align normalized activations with the corners of an inscribed hypercube, encouraging activation energy to be distributed more evenly across dimensions. This objective admits an efficient closed-form update via the orthogonal Procrustes problem, avoiding gradient-based optimization over the orthogonal group. We further introduce an online calibration procedure that updates rotations as calibration samples are processed, eliminating the need to store activations on disk and allowing rotations to adapt to quantized activation distributions during calibration. Experiments on Llama-2 and Llama-3 models from 3B to 70B parameters show that our method achieves competitive or improved performance across perplexity benchmarks and common sense reasoning tasks while avoiding both costly end-to-end training and large offline activation storage.
Authors: El Mehdi Saad, Victor Thuot, Nicolas Verzelen
Abstract: We study best-arm identification in large-scale stochastic dueling bandits under the sole assumption that a Condorcet winner exists, i.e., an arm that wins each noisy pairwise comparison with probability at least 1/2. We introduce a new identification procedure that exploits the full gap matrix \Delta_i,j=q_i,j-\tfrac12 (where q_i,j is the probability that arm i beats arm j), rather than only the gaps between the Condorcet winner and the other arms. We derive high-probability, instance-dependent sample-complexity guarantees that (up to logarithmic factors) improve the best known ones by leveraging informative comparisons beyond those involving the winner. We complement these results with matching lower bounds that establish the optimality of our procedures in all regimes. Overall, our results reveal the general form of the sampling complexity, characterized by a trade-off between the cost of locating informative entries and the verification cost required to achieve the desired confidence. In particular, this complexity drastically differs from what is suggested by pure asymptotic results or by procedures that are tailored to Strongly Stochastic Transitive models.
PaperID: 5308, Poster
Abstract: Reinforcement learning from verifiable rewards (RLVR) trains reasoning language models with group-relative objectives such as GRPO. However, it leaves two sources of signal underused: every token in a rollout receives the same response-level advantage, and groups in which all rollouts succeed or all fail have zero group-relative advantage and are typically discarded. We observe that same prompt rollouts are not independent reward labels but sibling attempts: they share scaffolding, diverge at uncertain decisions, and sometimes reconverge. In this paper, we propose Sibling-Relative Advantage Shaping (SIRAS), a lightweight advantage-shaping method that exploits this structure. SIRAS segments rollouts into chunks, aligns sibling chunks with banded dynamic time warping, and reshapes the flat GRPO advantage across tokens via soft-divergence weights that upweight chunks distinguishing a rollout from its siblings and downweight shared ones. This reshaping preserves each rollout’s token-averaged GRPO advantage exactly and reduces to GRPO when weights are constant. The same sibling structure produces calibrated residuals for all-correct and all-wrong groups, reclaiming training signal that group-relative baselines discard. Across six reasoning benchmarks, SIRAS yields the largest gains on AIME-style competition math, including +12.2 on AIME25 with Qwen3-8B and +23.3 on AIME26 in our backbone-scaling study, without extra rollouts, process reward models, or step annotations.
PaperID: 5309, Poster
Abstract: We revisit the problem of computing mean-field equilibria (MFEs) in discrete-time, monotone, finite-horizon mean-field games (MFGs). We show that, when the transition kernel is independent of the state-measure term and the reward function satisfies the usual weak monotonicity condition and is Lipschitz continuous, anchored proximal gradient descent methods can be used to compute a monotone MFE. We also establish last-iterate convergence results for these methods. Our approach relies on formulating the computation problem as an optimization problem over the space of occupation measures. Using this formulation, we show that the problem is equivalent to a class of constrained Lipschitz monotone inclusion problems. We then apply iterative methods for this monotone inclusion formulation to derive a tractable algorithm. The resulting algorithm achieves a convergence rate of O(1/\sqrtT) after T iterations, without requiring any regularization. This rate holds even in the absence of a uniqueness assumption for the corresponding MFE.
PaperID: 5310, Poster
Abstract: We study the ability of graph neural networks (GNNs) to learn generalizable algorithms for graph clustering. While previous work has shown that GNNs can learn to cluster graphs in-distribution, it is unclear exactly what algorithms these networks are learning, or whether they generalize to unseen graphs that differ significantly from their training data. In this paper, we pick a standard cut-based clustering cost function and train GNNs of various types to minimize this function on a suite of training examples. We then test whether the resulting networks are able to generalize to much larger graphs, or other families of graphs. We complement our empirical findings with a host of theoretical results that illustrate the range of possible heuristics that these different GNNs could be implementing. Finally, we use probes on trained networks to narrow down the heuristics actually being learned. Our findings show that there is generalization across the board but that GNNs of different expressivity---MPNNs versus PPGNs, for instance---learn markedly different algorithms.
Abstract: LLM-based workflows compose specialized agents to execute complex tasks, and these agents usually share substantial context, allowing KV-Cache reuse to save computation. Existing approaches either manage KV-Cache at agent level and fail to exploit the reuse opportunities within workflows, or manage cache at the workflow level but assume that each workflow calls a static sequence of agents. However, practical workflows are typically dynamic, where the sequence of invoked agents and thus induced cache reuse opportunities depend on the context of each task. To serve such dynamic workflows efficiently, we build a system dubbed PBKV (Prediction-Based KV-Cache Management). For each workflow, PBKV predicts the agent invocations in several future steps by fusing the guidance from historical workflows and context of the target workflow. Based on the predictions, PBKV estimates the reuse potential of cache entries and keeps the high-potential entries in GPU memory. To be robust to prediction errors, PBKV utilizes the predictions conservatively during both cache eviction and prefetching. Experiments on three workflow benchmarks show that PBKV achieves up to 1.85× speedup over LRU on dynamic workflows, and up to 1.26× speedup over the SOTA baseline KVFlow on the static workflow.
PaperID: 5312, Poster
Authors:
Fangru Linghu, Jackson R Ye, Jieying Wang, Alexandre V Morozov, Ian Foster, Zhao ZhangAbstract: Reproducing machine learning research is challenging because many papers do not provide executable code, and key implementation details are often scattered, implicit, or missing. Existing paper-to-code systems improve over direct prompting, but the generated repositories can still miss paper-critical logic or fail at execution time. We propose Crafter, an automated paper-to-code pipeline that treats reproduction as a problem of context-calibrated implementation recovery rather than single-pass code generation. Crafter builds an evidence-grounded specification, resolves missing implementation details before planning, and repairs the generated repository through execution-aware debugging. We evaluate Crafter with PaperBench Code-Dev for implementation faithfulness and a self-designed execution-oriented benchmark for runtime readiness. Across 23 PaperBench papers using the Claude Sonnet 4.6 backend, Crafter achieves an average Code-Dev score of 0.84, improving over Paper2Code by 19.7% and over DeepCode by 22.1%. In execution evaluation, Crafter reaches the contribution-level milestone on all three repositories, compared with two for both baselines, while reducing average repair cycles from 11.0--11.3 to 5.0.
PaperID: 5313, Poster
Authors: Suhaib Al-Rousan, Christian Schilling, Max Tschaikowski, Kim Larsen
Abstract: Verifying that two quantum circuits are equivalent is a critical challenge in compiler validation. A common approach is to transform the circuits into a tensor network and then iteratively contract tensors in the network. The order in which contractions are applied has a crucial influence on the complexity, as tensor size may grow exponentially in the number of qubits. This motivated research into heuristic policies to find good contraction orders, including recent work based on reinforcement learning (RL). Orthogonally, to mitigate the complexity of an explicit tensor representation, tensor decision diagrams (TDDs) have been proposed as a compact symbolic representation. However, also the TDD performance is sensitive to the contraction order, and policies designed for tensor networks do not perform well in case of a TDD representation. In this paper, we propose an RL framework to obtain a graph-neural-network policy that navigates the trade-off between arithmetic complexity and TDD size. Experimental results demonstrate that a policy learned with our framework is effective at taming the symbolic representation size.
PaperID: 5314, Poster
Abstract: Backdoor attacks implant a trigger-target association into a model, causing malicious behavior at test time while largely preserving clean performance. Despite extensive empirical study, a unified explanation for why standard training dynamics learn backdoors so effectively remains largely missing. We argue that backdoor learning can be understood as a consequence of simplicity-biased optimization dynamics: trigger features are dynamically simpler than semantic features under stochastic gradient descent (SGD) because they induce more coherent gradient alignment, more stable gating behavior, and stronger early-stage amplification. To formalize this perspective, we introduce a directional data model that separates semantic and trigger features and study a one-hidden-layer ReLU network trained by SGD. Our analysis identifies a frozen-gate signal that governs group-wise early-stage loss decrease up to controlled gate-drift and gate-flip errors, yielding a quantitative explanation for faster poisoned-sample fitting and its dependence on the poisoning ratio. Beyond optimization dynamics, we show that this early-stage bias can induce trigger-dominant neurons with selective activation on poisoned inputs, explaining why vanilla clean fine-tuning can attenuate backdoor behavior without necessarily erasing the underlying trigger-related representation. Empirically, we validate the predicted signatures of early-stage optimization bias, representation-level trigger dominance, and post-training residual behavior using ResNet-18 on standard image classification benchmarks.
PaperID: 5315, Poster
Abstract: Behavioral cloning (BC), which trains models from offline demonstrations, is a common approach in reinforcement learning settings. Prior work argues that BC requires expert demonstrations and performs poorly when trained on low-skill data. We challenge this assumption by showing that, in certain regimes, training on low-skill data can yield models that outperform those trained on high-skill data. Because low-skill data is often cheaper and more easily acquirable, this finding has important practical implications. To explain this result, we introduce the notion of fragility, characterizing how a policy's reward degrades under errors, and provide theoretical insights on how fragility can predict low-skill outperformance. We test our approach in a synthetic environment and MuJoCo and validate it using human data from chess and racing. Motivated by these findings, we connect our results to curriculum learning by structuring training according to demonstrator skill, rather than task difficulty as in standard curriculum design, and show that such skill-based curricula can improve performance relative to standard BC approaches.
PaperID: 5316, Poster
Abstract: Token merging has emerged as an effective training-free strategy for accelerating diffusion models by reducing redundant token computation while minimizing generation degradation. A key challenge is determining which tokens can be safely merged and which should be preserved for subsequent computation. Existing method addresses this by preserving prompt-aligned foreground regions, thereby protecting semantically relevant structures. However, this text-relevance criterion overlooks background objects, local details that are not explicitly mentioned in the prompt, leading to degraded structural coherence. In this work, we propose Dynamics-aware Token Merging (DaTMe), a training-free framework that grounds token merging in the underlying diffusion dynamics. Our key insight is that token importance should reflect whether a token remains unresolved during denoising and whether it carries structural information. To this end, DaTMe constructs a dynamics-aware importance map from the Tweedie denoised estimate. The map combines a temporal signal, measuring inter-step changes in the predicted clean image, with a structural signal, capturing local high-frequency content such as boundaries and fine details. Using this map, DaTMe adaptively assigns tokens into merging roles under a given compression ratio: important tokens are preserved as anchors, while tokens that are both temporally stabilized and structurally homogeneous are selected as safe merge candidates. Experiments on PixArt-\alpha and FLUX demonstrate that DaTMe maintains visual fidelity across the entire image, highlighting diffusion dynamics as an effective token-level merging decision in efficient generation.
PaperID: 5317, Poster
Abstract: We study the problem of learning Gaussian mixture models under overparameterization. Prior work has shown that, while overparameterization is essential for avoiding spurious local optima and enables global recovery of the ground-truth model using the gradient-EM algorithm, it can dramatically slow the local rate of convergence. Under certain assumptions on the mixture weights, we show that a standard divergence measure minimized by statistical learning procedures possesses a manifold of slow growth on which the well-known Polyak stepsize reduces the loss geometrically. We then design a gradient-based method that converges to minimizers at a locally linear rate. Additionally, we show that our method converges to nearly optimal solutions—up to a natural misspecification threshold—for mixtures with arbitrary weights. At a high level, the method alternates between several “short” gradient descent steps that approach the manifold and “long” Polyak steps that contract the distance to minimizers. Our results suggest that slow convergence is not an intrinsic challenge of overparameterization, but can be overcome by exploiting the favorable structure of the loss landscape.
PaperID: 5318, Poster
Abstract: Many scientific and machine learning systems, from molecular dynamics to diffusion models and beyond, are governed by stochastic dynamics with low-dimensional structure, evolving on slow timescales. However, target trajectories, used to identify and interpret them, are often inaccessible: only biased or static samples that explore dynamics' manifold are available. We introduce Langevin-Informed Transfer Learning (LITL), a framework for recovering target Langevin dynamics from biased source samples using only black-box feedback. LITL learns the leading spectral structure of the target infinitesimal generator and the projected drift through Dirichlet representation learning, enabling kinetic reconstruction in spectral form and slow-manifold gradient field estimation. We further introduce a spherical variant well suited to steering normalized latent representations commonly used in modern learning systems toward desired objective. We establish finite-sample guarantees for eigenvalue, eigenfunction, and projected drift estimation in Sobolev norms, thus proving generalization of all objects in function and first order derivatives values. Empirically, LITL recovers physical transition timescales from biased molecular simulations, builds kinetic structure from static samples of generative models, reconstructs spherical symmetries of physical systems, and enables post-hoc latent steering of trained neural networks under black-box feedback. Together, these results position spectral operator learning as a practical framework for recovering stochastic dynamics under distribution shift and unlock applications across machine learning and the physical sciences.
PaperID: 5319, Poster
Abstract: We propose SAFE-DRIFT, a data-selection framework for supervised fine-tuning that explicitly trades off target improvement against unwanted off-target drifts of model behavior. We formalize this objective as a constrained optimization problem: maximize gain on the target gradient subject to a budget on off-target drift, which we quantify by the Fisher information on the reference distribution. The resulting closed-form solution is a damped natural gradient with respect to the reference Fisher matrix. We further address the statistical estimation challenge in the high-dimensional setting, deriving a scalable algorithm that approximates the Fisher information and the target gradient in a shared low-dimensional subspace. We evaluate SAFE-DRIFT against state-of-the-art data selection methods across diverse domains including coding and medical question answering, yielding competitive target performance while reducing drift on specified off-target behaviors.
PaperID: 5320, Poster
Abstract: PDE-free latent diffusion models can synthesize plausible turbulence fields, but visually realistic snapshots do not guarantee statistically reliable long-horizon rollouts. In CNF-based turbulence diffusion, we observe that long rollouts can drift in latent transition laws, temporal memory, and transport-sensitive diagnostics even when the decoder remains spatially expressive. We address this failure mode with a DNS-calibrated stochastic transition closure for long-horizon turbulence diffusion. The closure is fit offline in a coarse latent observable space and defines stochastic transition tubes calibrated from DNS-referenced latent trajectories. During diffusion training, the frozen closure regularizes the denoiser's predicted-clean latent transitions by penalizing deviations from these calibrated tubes, while leaving the CNF decoder and reverse diffusion sampler unchanged. The method is fully PDE-free, uses no PDE residual or sampling-time correction, and is designed to improve transition-side temporal reliability rather than to serve as a universal physical simulator. Across DNS-referenced long-horizon protocols, the proposed method yields targeted gains in temporal transport and memory diagnostics, especially in high-drift regimes, while decoded-field spectra and derivative-sensitive quantities are reported as physical-fidelity guardrails rather than uniformly optimized targets. These results support stochastic transition closure as a lightweight mechanism for improving the temporal reliability of PDE-free turbulence diffusion.
PaperID: 5321, Poster
Abstract: Real-world networks, from biological brains to ecological systems, are typically low-rank, yet they continue to learn throughout their existence. Understanding how learning operates in this rank-deficient regime is essential, but existing theory captures only its limits. Classical results describe early training in random high-dimensional networks, while low-rank theory describes its structured endpoint. To bridge these regimes, we introduce rank-deficient RNNs, i.e., networks initialized at low rank with fully trainable weights, in which rank and synaptic strength vary independently. This decoupling reveals that synaptic strength, not rank, primarily governs the learning regime, and that its strong and weak limits confer opposing advantages. Strong synapses produce rich nonlinear dynamics at initialization, enabling rapid learning that matches full-rank training speed at low ranks and triples it in gated architectures (GRUs, LSTMs). Weak synapses, by contrast, distribute computation across the population so that no individual connection is critical, yielding solutions that are structurally stable to neuron loss. Empirically, the weight changes induced by training consistently follow weak synaptic scaling across four tasks, four architectures, and distinct initializations. Overall, our findings identify synaptic strength as a central variable controlling both trainability and structural stability in rank-deficient RNNs.
PaperID: 5322, Poster
Abstract: Large language models (LLMs)-based multi-agent systems (MAS) coordinate specialized agents through prompts, tools, and communication topologies, making their hidden workflows valuable intellectual property and security-critical assets. Existing black-box MAS extraction methods implicitly assume that adversarial queries can traverse all agents, which holds for static workflows but breaks down in dynamic workflows whose execution paths depend on input semantics and intermediate states. We identify two key challenges in dynamic workflows: branch overfitting, where fully adversarial queries overfit to the same branch and extract only a subset of agents, and stealthy coverage exploration, where the adversary needs to achieve complete branch coverage with few redundant queries while not exposing the extraction task. To address these challenges, we propose FlowLeak that combines Task-Preserving Payload Template, which preserves legitimate task semantics while eliciting workflow information to mitigate branch overfitting, with Coverage-Guided Branch Exploration, which uses previously extracted workflow fragments to generate branch-targeted tasks and constrains workflow extraction as an auxiliary task requirement, thereby reducing exploration queries and making extraction harder to identify. Experiments on 102 MAS show that FlowLeak addresses both challenges and substantially improves dynamic workflow extraction (2.68× improvement). Furthermore, we show that the extracted workflow information from FlowLeak enhances downstream attacks, highlighting the security risks of MAS workflow extraction and our method.
Abstract: We study infinite-horizon average-reward constrained Markov decision processes (CMDPs) under the weakly communicating assumption. Our contributions are twofold. First, we establish strong duality for weakly communicating average-reward CMDPs over stationary policies with finite state and action spaces. Despite the absence of a linear programming formulation and the resulting nonconvexity under the weakly communicating setting, we show that strong duality still holds by carefully exploiting the geometric structure of the occupation measure set. Second, building on this result, we propose a primal--dual clipped value iteration algorithm for learning weakly communicating average-reward linear CMDPs. Our algorithm achieves regret and constraint violation bounds of \widetilde\mathcalO(T^2/3), improving upon the best known bounds, where T denotes the number of interactions. Our approach extends clipped value iteration to the constrained setting and adapts it to a finite-horizon approximation, which stabilizes the dual variable and is crucial for achieving improved regret bounds. To analyze this, we develop a novel approach based on strong duality that enables the decomposition of the composite Lagrangian regret into separate bounds on regret and constraint violation.
PaperID: 5324, Poster
Authors:
Guimeng Liu, Milad Abdollahzadeh, Ngai-Man (Man) CheungAbstract: Image-text offset alignment, which refers to the consistency between visual transformation directions and their corresponding textual transformation directions, has become an important relational property in multimodal contrastive representation space (e.g., CLIP or SigLIP). This property is widely utilized in text-guided generative modeling tasks, such as generative model domain adaptation and text-guided Image-text offset alignment, which refers to the consistency between visual transformation directions and their corresponding textual transformation directions, has become an important relational property in multimodal contrastive representation spaces such as CLIP and SigLIP. This property plays a central role in text-guided generative modeling tasks, including generative model domain adaptation and text-guided image editing, where textual offsets are used to guide visual transformations. These methods fundamentally rely on the assumption that image and text offsets are well aligned, such that textual transformation directions can provide reliable guidance for corresponding visual transformations. However, the validity of this assumption has not been systematically examined. , we question this foundational assumption by conducting a comprehensive empirical analysis of image-text offset alignment in multimodal contrastive representation space. Our findings reveal not only noticeable offset misalignment but also a meaningful positive correlation between image-text offset misalignment and semantic concept distance across six large datasets and eight contrastive vision-language models. Based on this discovery, we propose Adaptation with Iterative Refinement (AIR), a method that iteratively refines text offsets through anchor sampling and our proposed concept description learning to reduce image-text offset misalignment and improve guidance accuracy. Comprehensive experiments on zero-shot generative model domain adaptation and text-guided image editing, including qualitative, quantitative, and user studies, consistently show that AIR enables state-of-the-art performance in these tasks.
PaperID: 5325, Poster
Abstract: Sparse Autoencoders (SAEs) decompose dense LLM activations into sparse, interpretable features. However, evaluating whether SAEs extract genuinely meaningful structure remains challenging. Current approaches rely on model-internal metrics such as reconstruction fidelity and automated LLM scoring, which assess SAEs as mathematical decompositions but provide no external validation that the extracted features correspond to anything outside the model. To address this issue, we propose a new validation approach: comparing SAE representations against human neural activity. The validation rests on a shared computational principle, since both SAEs and biological neural systems implement sparse coding over overcomplete populations. Through extensive experiments using EEG recordings of naturalistic reading and SAE features from three large language models, we demonstrate that SAE features systematically align with human brain activity. We further show that this alignment is dominantly carried by the SAE's learned sparse code, rather than by generic architectural properties or the scale of the training data. This work is the first attempt to validate pretrained SAE features against human brain activity, establishing biological alignment as a complementary benchmark for mechanistic interpretability research beyond model-internal metrics.
PaperID: 5326, Poster
Abstract: Transfer learning (TL)-based policy adaptation in deep reinforcement learning (DRL) usually relies on target-side data information for retraining when facing new tasks. However, two important issues remain underexplored in practical scenarios: existing DRL transfer learning methods usually lack theoretical guarantees against adversarial attacks, and target-side data information may be unavailable for retraining. In this paper, we propose Certifiable Aggregation of Smoothed Teachers (CAST), which adapts source policies by aggregating multiple certified smoothed teachers without any form of retraining, and further certifies the robustness of the adapted policy at both the action level and the cumulative reward level. CAST certification faces three key challenges: (i) value-function transfer in DRL cannot preserve certified robustness without robust retraining; (ii) reward changes break the direct reusability of source teachers' action-level certificates; and (iii) existing certified robustness transfer results provide limited guidance for certifying the student's cumulative reward. These challenges prevent us from directly using existing TL techniques in DRL. To address them, CAST constructs robustness signals from source certificates, incorporates them into policy aggregation to obtain action-level certificates, and bridges the adapted policy's action-level certificates to its target-task cumulative-reward certificate. Experiments on five multi-objective DRL benchmarks show that CAST exceeds the best source teacher in 17 out of 40 attack configurations and largely remains within the performance range of the source teachers. When testing whether smoothing benefits can be transferred, CAST improves in 39 out of 40 configurations, with a maximum relative gain of 183.9%.
Authors: Zhi Chen, Runze Hu, Le Zhang
Abstract: Flow matching has recently emerged as a principled framework for learning continuous-time transport maps, enabling efficient ODE-based sampling without relying on stochastic diffusion processes. While generative modeling has shown promise for medical image segmentation, particularly in capturing uncertainty and complex anatomical variability, existing approaches are predominantly based on diffusion models, which require iterative sampling and incur substantial computational overhead. In this work, we propose MedFlowSeg, a conditional flow matching framework that formulates medical image segmentation as learning a time-dependent vector field that transports a simple prior distribution to the target segmentation distribution. Compared to diffusion-based methods, our formulation enables more efficient inference through solving an ordinary differential equation, while preserving the flexibility of generative modeling. To effectively incorporate conditional information, we introduce a dual-conditioning mechanism. Specifically, we propose a Dual-Branch Spatial Attention (DB-SA) module to inject multi-frequency structural priors, and a Frequency-Aware Attention (FA-Attention) module to model interactions between spatial and spectral representations via discrepancy-aware fusion and time-dependent modulation. These components improve the alignment between noisy intermediate states and clean semantic features, leading to better structural consistency and boundary delineation. We conduct extensive experiments across multiple medical imaging modalities, where MedFlowSeg consistently outperforms prior state-of-the-art (SOTA) baselines, including diffusion-based and flow-based methods. Code is available at https://anonymous.4open.science/r/MedFlowSeg-C67C/.
PaperID: 5328, Poster
Authors: Willow Scott, Eugenio Valdano, Charles Assaad
Abstract: Missing data is pervasive in many scientific domains such as public health, environmental science, and the social sciences. Recoverability from missing data is typically studied using fully specified variable-level missingness models despite that, in many applications, only coarse structural information is available, for instance when variables are grouped into clusters due to limited knowledge or interpretability reasons. In this paper, we investigate recoverability from such abstract representations. We introduce two classes of cluster-based missingness graphs: the m-C-DMG, which retains variable-specific missingness indicators, and the cm-C-DMG, which aggregates missingness mechanisms at the cluster level. We formalize the notion of compatibility between these abstract graphs and underlying variable-level missingness models, and study how this abstraction affects the recoverability of probabilistic and causal queries. In particular, we give graphical conditions of recovering the joint distribution as well as graphical conditions of recovering a macro causal effect. Overall, our results clarify when cluster-level missingness information is sufficient for valid inference, and when finer-grained modeling is necessary.
Abstract: We present World from Motion, a method for generating freely renderable dynamic 3D Gaussian representations from monocular videos. Our approach conditions a video model on dense, pixel-aligned renderings that encode appearance, geometry, and 3D scene motion along both source and target camera trajectories to correct rendering artifacts and fill in missing regions from an initial reconstruction. To train this model, we construct a dataset of aligned multiview video pairs and dynamic 3DGS representations, with simulated artifacts characteristic of monocular reconstruction. At test time, we distill the model’s generations, including newly-observed regions and motions, back into a single consistent, high-quality dynamic 3DGS, improving both novel-view synthesis and the underlying 3D motion. Our method sets a new state of the art in 4D reconstruction and seamlessly generalizes to in-the-wild videos with large viewpoint changes and dynamic motions.
PaperID: 5330, Poster
Authors: Yuwei Zeng, Zekun Shi
Abstract: Effective discrete representation learning with VQ-VAE serves as a fundamental component of modern autoregressive generative models. However, it is widely known to suffer from codebook collapse and inefficient code utilization. To address this issue, we propose a simple yet effective alternative to standard vector quantization by formulating neural quantization as a joint modeling and clustering process. Specifically, we learn the latent space using a lightweight flow-matching model alongside training, corresponding to a probability flow ODE that transports an often fixed well-behaved source distribution while preserving probability mass. Based on this, we define the codebook \mathcalC_1 by transporting even-quantile partitioned prior codes \mathcalC_0 from the source space through the learned flow, and obtaining the Voronoi centroids under the induced distribution, leading a structured discretization. We evaluated on image tasks, and two tasks require high-precision tokenization, demonstrating scalability and improved expressivity of the codebook with full utilization.
Abstract: The first-moment buffer of most modern optimizers is an exponential moving average (EMA) of stochastic gradients with a single scalar decay, with a unified forgetting horizon along every direction in parameter space. Yet the per-sample gradient of any linear layer factorizes as a rank-1 outer product g_t = \delta_t x_t^\top, exposing the input activation as a natural key and the output-side error as a natural value, which is the structure that EMA's flat matrix average discards. We propose DeltaMomentum, which interprets the buffer as an online linear associative memory of these key--value pairs and updates it via the classical delta rule. The resulting update is anisotropic memory transport: directions queried frequently are forgotten quickly, directions queried rarely are preserved on long horizons, yielding an automatic, data-driven schedule matched to input statistics. We prove the buffer's expected fixed point is a Tikhonov-regularized Wiener predictor of the population gradient, which is equivalent to implicit input-side natural-gradient preconditioning at no covariance-inversion cost, and that the same mechanism strictly accelerates expected tracking-error contraction along every positive-density input direction under non-stationarity and reduces the per-direction iteration determinant in linear regression. DeltaMomentum is a drop-in replacement for the EMA accumulator in any base optimizer; we instantiate it as DeltaSGD and DeltaAdamW. A \muP derivation shows the delta coefficient is width-invariant, enabling zero-shot hyperparameter transfer; we verify this with a coordinate check across five widths. The per-block asymptotic FLOP overhead is \approx 15.74% at zero persistent-memory cost (down to 7.3% and 11.0% realized at 67M/370M). On FineWeb-Edu, DeltaAdamW reaches AdamW's terminal validation loss in up to 31.35% fewer steps at 67M and 19.28% fewer steps at 370M. On CIFAR-10, DeltaSGD shows the same pattern against SGD-momentum, confirming the gain is not specific to Adam. Mechanistic diagnostics confirm that improvements in gradient-estimator quality, function-space prediction error, and input-feature conditioning are operative throughout training.
PaperID: 5332, Poster
Abstract: Time series anomalies take diverse forms, and individual detectors encode different anomaly scoring principles that rarely cover all anomaly types. Recent unsupervised time-series anomaly detectors provide complementary scoring signals, yet deploying them as an ensemble preserves the inference and maintenance cost of every individual detector. We study whether heterogeneous teacher detectors can be distilled into a single lightweight detector in the unsupervised setting, where no anomaly labels are available to define a shared target. addresses this by extracting weak pairwise supervision from the training-set anomaly scores of frozen teachers. It normalizes each teacher score sequence independently and trains the student with CoRank, a multi-teacher ranking-distillation objective that converts teacher score gaps into pairwise supervision through teacher consensus. Across seven benchmark datasets, AnoMix successfully integrates complementary teacher scoring principles into a single deployable model, improving over individual detectors and static score ensembles while using a compact student with under 80k parameters on the benchmark settings. A zero-shot variant further shows that the distilled ranking signal transfers to unseen datasets and can outperform large pretrained time-series foundation models for anomaly detection.
PaperID: 5333, Poster
Abstract: Continuous test-time adaptation (CTTA) is crucial for deploying vision-language models (VLMs) in real-world environments where data distributions shift over time. Existing approaches typically utilize confidence-driven pseudo-labeling to enhance the performance on the target domain while retraining VLMs using historical information. Despite the progress, their performance is far from satisfactory due to error accumulation from noisy pseudo-labels and optimization interference between newly acquired and historical knowledge. Towards this end, we propose a novel approach named Complementary Cache Guidance with Gradient Disentanglement (CURE) for CTTA of VLMs. The core of our CURE is to balance instant knowledge acquisition and historical knowledge consolidation using both complementary cache systems and parameter space disentanglement. In particular, our CURE first constructs affinity structures among diverse text prompts, which guide majority voting to improve the quality of pseudo-labeling. More importantly, our CURE introduces a complementary cache system consisting of a short-term cache and a long-term cache with sample quality monitoring. The short-term cache stores recent reliable samples with high entropy, while the long-term cache preserves representative samples with high gradient consistency across iterations. Then, we extract prototypes from both caches for cross-modal alignment, enabling the model to acquire new knowledge while mitigating old knowledge forgetting. To further reduce the interference when learning from two caches, we perform low-rank decomposition of the gradient space, which facilitates historical knowledge consolidation in the null space of short-term signals. Extensive experiments on several benchmarks demonstrate the superiority of CURE over state-of-the-art baselines. Our source codes are available at https://anonymous.4open.science/r/CURE-B486.
Abstract: We introduce BEACON—Best-Effort Adaptation for Cross-Domain Co-Training—a theory-driven framework for training generative robot policies with abundant source demonstrations and limited target demonstrations. BEACON casts cross-domain co-training as a discrepancy-aware importance-reweighting problem, jointly learning a diffusion-based visuomotor policy and per-sample source weights that minimize an objective informed by target-domain generalization guarantees. To make best-effort adaptation practical for high-dimensional sequence policies, we develop scalable instance-level discrepancy estimators, stochastic alternating updates for policy and weights, and a multi-source extension that balances heterogeneous source domains. Across sim-to-sim, sim-to-real, and multi-source manipulation settings, BEACON improves robustness and data efficiency over target-only, fixed-ratio co-training, and feature-alignment baselines. Importantly, even without an explicit alignment objective, BEACON achieves feature alignment as an implicit result of discrepancy-aware cross-domain co-training.
Abstract: Knowledge-graph retrieval-augmented generation (KG-RAG) couples large language models (LLMs) with structured, verifiable knowledge graphs (KGs) to reduce hallucination and provide reasoning traces. However, current KG-RAG systems often rely on fixed pipelines of multiple LLM modules (e.g., planning, reasoning, and responding), which inflate inference costs and tie performance to specific graph schemas. To address this, we introduce KG-R1, an agentic framework that optimizes KG-RAG through reinforcement learning (RL). Unlike modular workflows, KG-R1 uses a single agent that interacts with KGs as its environment, learning to retrieve information at each step and incorporating it into its reasoning and generation in a unified process. Across Knowledge-Graph Question Answering (KGQA) benchmarks, KG-R1 demonstrates both efficiency and transferability—using Qwen 2.5-3B, KG-R1 improves answer accuracy with fewer generation tokens than prior multi-module workflow methods that use much larger foundation or fine-tuned models. Furthermore, KG-R1 exhibits strong plug-and-play capability: after training, maintaining accuracy on unseen KGs without retraining. These properties make KG-R1 a promising KG-RAG framework for real-world deployment. Our code is publicly available at anonymous.4open.science/r/KG-R1-B881/.
Abstract: Recent DiT-based text-to-image models increasingly adopt LLMs as text encoders, yet text conditioning often remains static, relying on a fixed representation from a single LLM layer across all denoising timesteps and DiT blocks. This motivates a central question for DiT text conditioning: when multi-layer LLM features are available, how should hierarchical semantics be injected into the dynamic generation process? To investigate this question, we introduce Semantic Routing, a lightweight multi-layer conditioning interface that forms text features through normalized convex fusion of LLM hidden states and compares static, time-wise, depth-wise, and joint routing strategies under matched training and evaluation settings. Across text--image alignment and compositional benchmarks, depth-wise routing emerges as the most reliable strategy among the studied variants, with particularly strong gains on advanced prompt-following tasks (e.g., +9.97 on GenAI-Bench Counting over the penultimate-layer baseline). In contrast, purely time-wise routing can degrade visual fidelity; our analysis suggests a mismatch between nominal timesteps and the effective inference trajectory, making time-conditioned semantic injection poorly timed. These results indicate that DiT depth is an effective axis for improving text--image alignment through dynamic LLM feature injection, while robust time-dependent conditioning remains an open challenge.
PaperID: 5337, Poster
Abstract: Diffusion models achieve strong performance in text-to-image generation, but they can still produce unsafe content. Training-free safety guidance offers a practical inference-time alternative that intervenes during sampling without modifying model weights. However, many existing safeguards still rely on a coupled safety-control chain. In this chain, the same unsafe condition simultaneously specifies what should be avoided, indicates the current unsafe signal, constructs the intervention, and determines when that intervention remains active. This coupling can work when unsafe factors are explicit, but becomes less reliable once harmful and benign semantics are entangled in the current generation state, often leaving residual unsafe content or unnecessarily degrading benign-generation quality or prompt fidelity. We address this problem with Decoupled Safety Control (DSC), a plug-in safety-control algorithm for training-free safety guidance. DSC reformulates training-free safety guidance as an explicit state-dependent control process, so the executed intervention is determined from the current generation state rather than inherited directly from a single unsafe condition. Experiments on standard safety benchmarks show that DSC improves representative latent- and text-space safeguards, achieving stronger unsafe-content suppression while preserving benign-generation quality and maintaining efficient inference.
Abstract: —how precisely a generated image aligns with its prompt—is increasingly central to the real-world utility of text-to-image (T2I) models. Existing faithfulness benchmarks, however, rely on simple atomic instructions, where top-tier systems already achieve near-perfect scores. As T2I models enter creative workflows, users increasingly issue multi-faceted requests involving intricate spatial relationships, stylistic constraints, and complex text rendering. In this setting, a single binary VLM-judge score no longer reveals , a 310-prompt stress benchmark drawn from real arena T2I logs, with approximately 30 decomposed yes/no constraints per prompt spanning six categories, including text rendering. The strongest closed-source system we evaluate reaches performance gap across 11 systems, demonstrating substantial discriminative power. Moreover, high public-arena rankings fail to predict faithfulness, confirming that holistic Bradley–Terry (BT) preference scores prioritize aesthetics over fine-grained prompt adherence. We propose a that decomposes each prompt into a DAG of yes/no questions and zeroes descendants of failed parents, turning faithfulness into a per-constraint training signal. Combined with a BT aesthetic reward via , which standardizes each reward within its rollout group so neither collapses, our recipe attains a strictly better faithfulness–aesthetics trade-off on under MMRB2 pairwise comparisons than every single-reward, naive weighted-sum, or 4-reward BT-ensemble baseline. We release
PaperID: 5339, Poster
Authors: Corey Herr, Cedric Allier, Stephan Saalfeld, Allyson E. Sgro
Abstract: Cells in multicellular systems have complex shapes, yet most techniques for predicting biological dynamics from image sequences discard shape information by reducing cells to point particles. Here, we present a framework that moves beyond point representations by learning dynamic, shape-aware representations of synthetic cells in a multicellular collective directly from 2D time-series images. Specifically, our GraphDINO framework combines a frozen DINOv3 backbone for shape representation with a graph neural network trained through next-frame prediction. GraphDINO learns a per-cell latent representation that captures underlying cellular properties such as growth rate, membrane tension, and cell-cell adhesion without parameter supervision. The quality of this recovery depends on the encoder, with DINOv3 features providing a better shape representation than classical shape encoders. Together, these results demonstrate that combining foundation model features with graph-based interaction rules allows the discovery of key dynamical parameters that govern cellular behavior, directly from image sequences.
PaperID: 5340, Poster
Authors: Zhiyu Yan, Shiao Liu, Rahul Mahajan, Anand Viswanathan, Christopher D Anderson, Rui Duan
Abstract: Electronic health records (EHRs) contain rich health information for uncovering latent patient representations and clinically meaningful subgroups. Learning from EHR data across institutions is valuable because representations learned at a single site may reflect local data patterns and lack generalizability, yet pooling patient-level records across institutions is often impractical because of privacy constraints and communication costs. We propose a communication-efficient, non-iterative federated spectral method for learning shared latent representations from multi-institutional EHRs, in which each institution shares aggregate marginal frequencies and a low-rank spectral summary while keeping patient-level records local. A central server aggregates these summaries with the pooled marginal frequencies to recover both the shared low-rank structure and the canonical row-normalization direction required for interpretable representation recovery. We prove row-wise recovery guarantees and error bounds, with a leading federated variance term matching a corresponding lower bound under explicit regularity conditions. Simulations show that our method closely matches pooled analysis in estimating latent representations, improves as more sites contribute summaries, and outperforms single-site and existing federated benchmarks across a range of data settings. Applied to real-world multi-facility EHR data for stroke patients, the method learns interpretable patient representations that highlight clinically distinct subgroups and could help inform subgroup-specific care.
PaperID: 5341, Poster
Authors:
Yi Wen, Hao Chen, Wanyu Wang, Maolin Wang, Pengyue Jia, Derong Xu, Hui-Ze Tan, Yingyi Zhang, Wenlin Zhang, weihongluo, Xiku Du, Xiangyu ZhaoAbstract: The Reinforcement Learning (RL) paradigm has achieved remarkable success in enhancing the reasoning capabilities of Large Language Models (LLMs). However, when applied to long-horizon tasks, it faces severe challenges: i) Due to the high task complexity, the model struggles to sample enough successful trajectories, resulting in slow convergence. ii) Relying solely on sparse final rewards causes all steps to be updated indiscriminately, encouraging redundant step generation. When both the quantity and quality of successful trajectories become difficult to guarantee, the performance is severely constrained. This dilemma stems from the limitations: existing RL paradigms rely entirely on the reasoning capability of base models, lacking the ability to actively optimize trajectories. To address the above issues, we propose a novel trajectory evolution paradigm with hindsight credit assignment, termed TrajEvolve. TrajEvolve not only relies on the trajectory inferred by the model but also generates ``evolutionary trajectories'' by refining low-quality steps in historical trajectories, thereby exposing the policy to high-quality compositions that lie within the base model's support but are rarely sampled within practical training budgets. Furthermore, TrajEvolve reformulates step-level continuous credit assignment as a binary classification task, whether each step is necessary for task completion. This simplification not only reduces the learning difficulty but also enables step labels to be obtained through counterfactual verification, achieving precise identification of redundant steps. Extensive experiments on two datasets demonstrate the superiority of the proposed method.
PaperID: 5342, Poster
Authors:
Junyu Shi, Yong Sun, Zhiyuan Zhang, Lijiang LIU, Yuxin He, Zhengjie Zhang, Qiang NieAbstract: Motion generation has made substantial progress in synthesizing human motion from language, yet text alone remains ambiguous for specifying fine-grained spatial, temporal, and interaction details. Visual observations offer a powerful source of complementary context: even sparse images can reveal intermediate body states, object affordances, and interaction relations that are difficult to specify precisely in language. Whatmore, contextual constraints like visual, object, and partial-motion conditions can ground motion genetration with higher physical fidelity and controllability. We therefore study a unified multimodal context-aware framework for motion generation under diverse conditions. The key challenge is that heterogeneous conditions impose constraints at different spatial-temporal granularities and exhibit token- and phase-dependent relevance. Moreover, incorporating multiple modalities without a robust motion prior may entangle modality-specific semantics and cause negative transfer. In these regards, we combine global multimodal modulation with fine-grained motion token-level context selection, enabling the model to adaptively exploit relevant signals during generation. We further adopt progressive training that first learns a strong language-conditioned motion prior, then extends to multimodal context with different modality combinations. To support multimodal training, we introduce Mo900H, a large-scale benchmark integrating 21 motion datasets with over 900 hours of human motion. Our method reduces FID by 21.1% and 48.3% compared with SOTA methods on the HumanML3D and Mo900H datasets, while also improving motion captioning and enabling diverse multimodal-conditioned generation.
Abstract: We design the first regret guarantees for robust dynamic pricing which decouples the dependence on the corruption C and the time horizon T. In dynamic pricing, a seller with unlimited supply of a good interacts with a stream of buyers over T rounds, with the goal of maximizing revenue. At each round t, the seller posts a price p_t, and the buyer purchases the good only if their unknown valuation v^\star exceeds this price. The seller observes only the binary feedback \mathbbI[p_t \leq v^\star], indicating whether a sale occurred. In the robust pricing setting, a malicious adversary is allowed to corrupt this feedback in at most C rounds. Even if the learner knows the corruption C, the best known regret bound is \mathcalO(C\log\log T) by Gupta et al. [2025]. They left as an open problem to "decouple'' the dependence on C and T. In this work, we resolve this open problem. In particular, we develop a robust variant of binary search that achieves regret \mathcalO(C+\log T) when the corruption C is known and \mathcalO(C+\log^2 T) when the corruption is unknown.
Abstract: Advances in large language models (LLMs) have created new opportunities in data science, but their deployment is often limited by the challenge of finding relevant data in large data lakes. Existing methods struggle with this: both single- and multi-agent systems are quickly overwhelmed by large, heterogeneous files, and master–slave multi-agent systems rely on a rigid central controller that requires precise knowledge of each sub-agent’s capabilities, which is not possible in large-scale settings where the main agent lacks full observability over sub-agents’ knowledge and competencies. We propose a novel multi-agent paradigm inspired by the blackboard architecture for traditional AI models. In our framework, a central agent posts requests to a shared blackboard, and autonomous subordinate agents--either responsible for a partition of the data lake or retrieval from web--volunteer to respond based on their capabilities. This design improves scalability and flexibility by removing the need for a central coordinator to know each agent’s expertise or internal knowledge. We evaluate the approach on three benchmarks that require data discovery: KramaBench and modified versions of DSBench and DA-Code. Results show that the blackboard architecture substantially outperforms strong baselines, achieving 13%–57% relative improvements in end-to-end success and up to a 9% relative gain in data discovery F1 over the best baseline.
PaperID: 5345, Poster
Abstract: Does RL post-training build new reasoning procedures, or merely reweight behaviors already latent in the base model? We study this question in a fully observable rewrite-grammar environment where the pretraining distribution is known and every generated rewrite can be audited. A Transformer is pretrained on primitive symbol-rewrite chains and post-trained with only a binary final-answer reward. RL solves held-out problems that remain rarely solved by the pretrained model even under much larger sampling budgets, while rejection fine-tuning improves early but plateaus. Trace analysis reveals a phased mechanism: RL first strengthens primitive reductions, then enters a chunking phase in which it forms valid compressed procedures---macro contractions that collapse sequential reductions and parallel contractions that combine independent ones. These procedures are not isolated samples; they are reused and consolidated into a stable repertoire. Comparing RL with rejection fine-tuning shows that the key difference is not exploration volume but selectivity: RFT produces many shortcut-like rewrites, much of them invalid, whereas RL concentrates exploration into valid reusable structure. Pretraining ablations show that strategy emergence is gated not by primitive exposure alone, but by whether pretraining organizes primitive competence into reduction procedures that RL can later compress. The base model provides weak procedural ingredients; RL builds them into reliable higher-level strategies.
PaperID: 5346, Poster
Abstract: Upgrading a deployed face-recognition model usually makes new query embeddings incompatible with the legacy gallery, forcing costly backfilling of stored embeddings. We show that this incompatibility is often much simpler than it appears: independently trained angular-margin face models exhibit representation rigidity, where their embedding spaces are well approximated by a single global orthogonal transformation. We formalize this phenomenon as Procrustes rigidity, derive a finite-calibration recoverability bound, and estimate the orthogonal adapter from a small paired calibration set via closed-form Procrustes alignment. The resulting deployment protocol is backfill-free: transport upgraded queries into the legacy coordinate system, then reuse the existing gallery without re-embedding stored identities. Across 45 models spanning three architectures, two losses, three data scales, and five seeds, we find strong rigidity within architecture families, non-monotonic scaling with data size, and robust cross-model retrieval. On the 1.6M-image IFRT benchmark, Procrustes-aligned queries recover at least 99.8% of mono-model Rank-1 accuracy for same-family upgrades, while lightweight residual correction substantially narrows the cross-architecture gap.
PaperID: 5347, Poster
Authors: Sheng Zhao, Weikai Lin, Yuhao Zhu
Abstract: Egocentric gaze prediction enables many downstream applications but remains challenging, as human gaze is inherently stochastic. This stochasticity is constrained by structured temporal dynamics alternating between fixations and saccades, top-down influences from tasks, and bottom-up visual saliency. Based on this observation, we introduce GazeFlow, a framework that directly models gaze as a joint distribution of temporal gaze positions conditioned upon both top-down and bottom-up information. In particular, GazeFlow uses conditional flow matching (CFM): a learned velocity field iteratively transports a Gaussian noise sample into a plausible gaze trajectory drawn from this joint distribution. The velocity field is conditioned on bottom-up visual features extracted by a video encoder and on top-down task information obtained by globally querying these features. On standard datasets, GazeFlow achieves state-of-the-art performance on per-frame metrics, and the generated trajectories align better with human gaze temporal dynamics.
Abstract: Recently, Deep Image Prior (DIP) has demonstrated strong capabilities for solving inverse imaging problems (IIPs) by optimizing a randomly initialized convolutional neural network in a training-data-free regime. However, DIP suffers from overfitting to noisy measurements due to network over-parameterization, making early stopping (ES) essential. The most successful ES method tracks fluctuations in the running variance of the network output to detect overfitting. However, in many applications, these fluctuations may appear prematurely, leading to unstable reconstructions. In this paper, we first show that nearly optimal DIP early stopping can be achieved when two independent noisy copies of the degraded image are available. Motivated by this observation, and since obtaining two fully independent copies is infeasible, we propose an overfitting detection framework based on constructing pseudo self-referenced images, resulting in three IIP-specific algorithms. Our approach is further supported by theoretical results on single-reference validation, pseudo-validation estimation, and the impact of shared noise. Across different IIPs, ranging from natural image restoration to medical image reconstruction, and under varying noise levels and noise types, our methods consistently outperform existing DIP early stopping approaches, all without requiring knowledge of the measurement noise characteristics.
PaperID: 5349, Poster
Abstract: Contextual queueing bandits provide a framework for learning to schedule heterogeneous jobs under unknown context-dependent service rates. Under stochastic contexts, existing algorithms achieve \widetilde\cal O(T^-1/4) queue length regret, defined as the expected difference between the learner's and oracle's queue lengths at horizon T. In this paper, we improve this rate to \widetilde\cal O(T^-1/2). The key observation is that random exploration is needed only up to a carefully chosen cutoff round, rather than throughout the entire horizon. We propose CQB-\eta-2, a three-phase algorithm: (i) pure random exploration to construct an initial estimator, (ii) \eta-random exploration combined with a UCB rule to continue learning while maintaining negative drift, and (iii) pure UCB after the exploration cutoff. Our proof decomposes the queue length regret at the cutoff round. Before the cutoff, negative drift suppresses queue length differences caused by suboptimal choices. After the cutoff, the first two phases provide sufficient random exploration samples, ensuring that UCB decisions incur small departure-rate gaps. Combining these two bounds yields queue length regret of order \widetilde\cal O(T^-1/2). We further prove a minimax lower bound of order \Omega(T^-1/2). The proof constructs two hard instances that are statistically indistinguishable up to the final service decision, and uses a queue-specific coupling argument to convert the resulting testing error into queue length regret. Together, our upper and lower bounds characterize the minimax dependence on the horizon T up to logarithmic factors.
Abstract: Recent advances in Vision-Language Models (VLMs) have achieved impressive performance across many tasks, yet prior studies report unsatisfactory performance when applying large language or multimodal models to finding abnormal patterns in sequential data. Public anomaly detection benchmarks typically provide interval annotations but not natural-language rationales, making it difficult to fine-tune VLMs to produce grounded, interpretable decisions. To address this gap, we construct VisAnomBench, an explanation-augmented curated benchmark built from public time-series datasets and augmented with high-quality anomaly explanations selected from multiple large VLMs using fine-grained, task-specific rewards. Building on this benchmark, we present VisAnomReasoner, a parameter-efficient VLM for time-series anomaly detection. Experimental results on VisAnomBench show that VisAnomReasoner achieves more accurate anomaly localization and consistently outperforms all baselines, with improvements of at least 21.23 and 23.87 percentage points in precision and F1, respectively. Additional experiments on the TSB-AD-U benchmark demonstrate strong cross-benchmark generalization, with VisAnomReasoner improving precision and F1 by 9.57 and 13.39 percentage points, respectively. Dataset, code, and model weights will be open-sourced.
Abstract: Optimization with matrix gradient orthogonalization has recently demonstrated impressive results in the training of deep neural networks (Jordan et al., 2024; Liu et al., 2025). In this paper, we provide a theoretical analysis of this approach. In particular, we show that the orthogonalized gradient method can be seen as a first-order trust-region optimization method, where the trust-region is defined in terms of the matrix spectral norm. Motivated by this observation, we develop the stochastic non-Euclidean trust-region gradient method with momentum, which recovers the Muon optimizer (Jordan et al., 2024) as a special case, along with normalized SGD and signSGD with momentum (Cutkosky and Mehta, 2020; Sun et al., 2023). In addition, we prove state-of-the-art convergence results for the proposed algorithm in a range of scenarios, which involve arbitrary non-Euclidean norms, constrained and composite problems, and non-convex, star-convex, first- and second-order smooth functions. Finally, our theoretical findings provide an explanation for several practical observations, including the practical superiority of Muon compared to the Orthogonal-SGDM algorithm of Tuddenham et al. (2022) and the importance of weight decay in the training of large-scale language models.
Abstract: With the increasing versatility of text-to-image diffusion models, the ability to selectively erase undesirable concepts (e.g., harmful content) has become indispensable. However, existing concept erasure approaches primarily focus on removing unsafe concepts without providing guidance toward corresponding safe alternatives, which often leads to failure in preserving the structural and semantic consistency between the original and erased generations. In this paper, we propose a novel framework, PAIRed Erasing (PAIR), which reframes concept erasure from simple removal to consistency-preserving semantic realignment using unsafe–safe pairs. We first generate safe counterparts from unsafe inputs while preserving structural and semantic fidelity, forming paired unsafe–safe multimodal data. Leveraging these pairs, we introduce two key components: (1) Paired Semantic Realignment, a guided objective that uses unsafe–safe pairs to explicitly map target concepts to semantically aligned safe anchors; and (2) Fisher-weighted Initialization for DoRA, which initializes parameter-efficient low-rank adaptation matrices using unsafe–safe pairs, encouraging the generation of safe alternatives while selectively suppressing unsafe concepts. Together, these components enable fine-grained erasure that removes only the targeted concepts while maintaining overall semantic consistency. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art baselines, achieving effective concept erasure while preserving structural integrity, semantic coherence, and generation quality.
PaperID: 5353, Poster
Authors: Guangzhao Chai, Jing Liu, Youxi Wu, Yan Li
Abstract: Generating high-fidelity mixed-type tabular data remains challenging because such data contain both numerical and categorical features; different columns often exhibit substantial heterogeneity in marginal distributions, sparsity patterns, and dependency structures. Although recent diffusion-based and flow-based methods have improved generation quality, most methods still rely on a unified generation framework that models all dimensions in a homogeneous manner, without explicitly accounting for the fact that different columns may differ in their suitability for strong local refinement. To address this issue, we propose a driver-navigator structured flow-matching framework (DN-Flow) for mixed-type tabular data generation. Inspired by the collaboration between the driver and navigator in rally car racing, DN-Flow explicitly decomposes the velocity field into a Driver branch (for globally stable transport) and a Navigator branch (for endpoint-aware local refinement). On top of this structured decomposition, we introduce a learnable column-wise soft selector and a dynamic correction control mechanism to jointly model column-level suitability for local refinement and the actual strength of correction injection in an end-to-end manner. Experiments on eight real-world datasets show that DN-Flow consistently outperforms strong baselines in data fidelity and downstream utility, achieves the best Shape performance on all eight datasets, and remains highly competitive on Trend and machine learning efficiency. These results suggest that structured and selective velocity correction offers an effective new paradigm for mixed-type tabular data generation. The code will be released upon acceptance.
PaperID: 5354, Poster
Abstract: Covering problems are fundamental combinatorial optimization problems with broad connections to clustering, facility location, and various applications such as machine learning and controlling disease outbreaks. In this work we study a variant of the set cover problem that generalizes the partition set cover problem~\citepbera2014approximation. In partition set cover, given an instance (P, S) with a universe P = \cup_c \in [\ell] P_c (where [\ell] denotes \1, 2, \ldots, \ell\), a collection of subsets \(S\) (each with associated costs), and coverage requirements k_1,k_2,\dots,k_\ell, the objective is to find a subcollection with minimum cost that covers at least k_c elements from each group c \in [\ell]. We generalize this problem by allowing for the groups P = \cup_c \in [\ell] P_c to be overlapping thus capturing both \emphgroup fairness and \emphmultigroup fairness. We further allow for probabilistic coverage requirements for each element in the ground set thus capturing the notion of \emphindividual fairness. Our main result is a randomized iterated-rounding-based algorithm with provable guarantees. Our main result is a randomized iterated rounding based algorithm with provable guarantees. We supplement our theoretical results by showing that algorithm has significantly running times and outperforms the best known algorithm of \citepinamdar18PartitionSetCover.
Abstract: Encoding classical data into quantum states is a central bottleneck in quantum machine learning: many widely used encodings are circuit-inefficient, requiring deep circuits and substantial quantum resources, which limits scalability on quantum hardware. In this work, we propose TNQE, a circuit-efficient quantum data encoding framework built on structured unitary tensor network (TN) representations. TNQE first represents each classical input via a TN decomposition and then compiles the resulting tensor cores into an encoding circuit through two complementary core-to-circuit strategies. To make this compilation trainable while respecting the unitary nature of quantum operations, we introduce a unitary-aware constraint that parameterizes TN cores as learnable block unitaries, enabling them to be directly optimized and directly encoded as quantum operators. The proposed TNQE framework enables explicit control over circuit depth and qubit resources, allowing the construction of shallow, resource-efficient circuits. Across a range of benchmarks, TNQE achieves encoding circuits as shallow as 0.04× the depth of amplitude encoding, while naturally scaling to high-resolution images (256 × 256) and demonstrating practical feasibility on real quantum hardware.
PaperID: 5356, Poster
Authors:
Wei Zhou, Chenyiming, Xu Jinwei, Yang zhou, Quan Gan, Li YangAbstract: Video detectors are usually evaluated on observational splits, where target events co-occur with environmental and capture factors. A high AUC under this setup can reflect either the event itself or its surrounding conditions, and standard reporting cannot tell them apart. This work addresses the entanglement with CARVE, a counterfactual video editing protocol. For each source clip, CARVE generates a matched quartet (V^0, V^E, V^A, V^AE) that holds camera pose, layout, and surrounding traffic fixed while independently intervening on environment and event. A two-phase generator first produces a layout-conditioned reference image, then performs reference-guided video editing; a three-layer reference / objective / VLM-panel protocol filters the output. Each quartet supports thresholded and continuous diagnostics for false-positive purity, event faithfulness, environment-induced drift, and held-out factor compositions. Thresholded scores are paired with continuous ones so that a conservative detector cannot suppress scores everywhere and look robust. The same audit drives training: CGAA-IC samples weak factors weighted by measured brittleness and adds quartet-level purity, consistency, and faithfulness losses; the inference graph is unchanged. On CCTV accident detection, the audit shows that clean AUC does not predict robustness rankings: a VideoMAE detector with 91.4% AUC on TAD scores only 72.3% CPS on the night slice of CARVE-Q. CGAA-IC raises purity from 84.2% to 89.9% and cross-dataset average AUC from 80.3% to 84.5%.
Abstract: Predicting gene expression from H\&E-stained histology images offers a scalable alternative to costly spatial transcriptomics, yet most existing methods operate at the spot level, where signals from multiple cells are aggregated and critical cellular heterogeneity is obscured. Extending this paradigm to single-cell resolution is non-trivial. Naively applying pathology foundation models faces a scale mismatch: their patch-level representations mix multiple cells, whereas per-cell cropping or resizing distorts morphology and removes local context. Conversely, segmentation-based models without strong pretrained visual encoders often lack the morphological representation capacity needed for accurate molecular prediction and inherit errors from imperfect cell boundary masks. Here, we present CELLO, an efficient end-to-end framework that performs a single pathology foundation model forward pass per image and uses grid sampling to extract location-specific features for all cells simultaneously. We further introduce a distance-decay cross-attention module that refines each cell representation using spatially biased local morphological context. Using 52 paired Xenium--WSI samples spanning 12 organs and approximately 10M cells, CELLO achieves state-of-the-art performance in both in-distribution and out-of-distribution settings while delivering at least a 7× inference speedup over DeepSpot2Cell. Our work establishes a scalable foundation for single-cell gene expression prediction from H\&E images.
PaperID: 5358, Poster
Abstract: Phylogenetic inference plays a central role in understanding evolutionary relationships, with applications ranging from tracking pathogen spread to reconstructing the history of life. Conventionally, practitioners obtain posteriors over the phylogenetic tree of species using MCMC methods. However, such likelihood-based methods are only tractable for simple evolutionary models with restrictive assumptions. For more complex and realistic evolutionary models, conventional methods are prohibitively expensive, inaccurate, or even impossible. We instead advocate for a simulation-based inference approach, by using simulated data from an evolutionary model to train a neural network that predicts tree topologies conditioned on sequences. To accurately represent the complex posterior distributions over tree topologies that can arise, we present flexible models that iteratively generate trees using three natural paradigms: top-down, middle-out, and bottom-up. We use a discrete diffusion framework to train these models efficiently on large-scale simulated datasets of phylogenetic trees. For all three generative paradigms, our models fit the data substantially better than the previous state-of-the-art simulation-based method, Phyloformer 2, and obtain more accurate posteriors on real datasets. Finally, our models outperform misspecified conventional methods on data following complex evolutionary processes.
PaperID: 5359, Poster
Abstract: During evolution, biological neurons scale computation by expanding dendritic branches. However, this expansion creates a biophysical dilemma: increased membrane area typically imposes a capacitive load that slows somatic dynamics and narrows temporal bandwidth. Here, we investigate how primate cortical neurons resolve this tradeoff. While primary basal dendrite number is relatively conserved across mouse cortical areas, it increases systematically along the primate cortical hierarchy. Using biophysically constrained multicompartment modeling, we identify a key dimensionless control parameter---the dendritic-to-somato-apical specific membrane resistance ratio, \rho \equiv R_m,\mathrmdend/R_m,\mathrmsa, that governs the coupling between dendritic morphology and the effective somatic time constant \tau. In mouse-like neurons (\rho>1), adding basal dendrites progressively increases \tau and promotes reliable low-pass integration. Conversely, human-like neurons operate near a balanced-resistance regime (\rho \approx 1), where resistive reweighting counterbalances morphology-driven capacitive loading. This allows \tau to remain nearly invariant despite expanded dendritic topology, shifting the structure-function tradeoff toward faster, more flexible coding with improved high-frequency tracking. Our results reveal a biophysical scaling principle linking evolutionary dendritic architecture to hierarchy-dependent coding, offering potential design rules for preserving temporal bandwidth in dendrite-inspired neuromorphic architectures.
PaperID: 5360, Poster
Authors:
Xiaoyu Yi, Zihan Xiong, Junjie Cao, Xiaolong Weng, Yingzhe Ma, Gus Xia, Jinliang Liu, Yuxin Xie, Xianwei Zhuang, Yuguo Yin, Qiao Jin, Gongxi Z Zhu, Jiayu Wang, Junjie Liang, Qin Yue, Ziyu Wang, Dading Chong, Meng Cao, Jiahuan Zhou, Dongchao YangAbstract: A cover song re-renders an existing piece by preserving its tonal content—melody and chord progression—while reshaping other musical attributes. This suggests a natural formulation for generative cover synthesis with pretrained music language models: explicitly control tonal content, while specifying lyrics and style through text. Existing approaches condition on frame-level features, tightly coupling generation to the reference’s timing and structure, and limiting flexibility such as tempo variation, structural rearrangement, or partial conditioning. To this end, we present FlexCover, a cover generation model that conditions a pretrained text-to-song foundation model on a symbolic lead sheet. Our design enables alignment-free control: generated outputs preserve the tonal signature of the source without frame-level alignment, allowing flexible timing, structure, and segment-level generation. We further introduce a training curriculum with partially mismatched audio–symbolic pairs, improving diversity and robustness. We evaluate FlexCover with objective metrics, standard subjective ratings, and in-depth expert interviews that provide fine-grained diagnostic insights. Experiments show that FlexCover achieves state-of-the-art performance among open-source systems and is competitive with leading commercial models such as Suno-v5.5.
Abstract: Aligning multimodal generative models with human preferences demands reward signals that respect the compositional, multi-dimensional structure of human judgment. Prevailing RLHF approaches reduce this structure to scalar or pairwise labels, collapsing nuanced preferences into opaque parametric proxies and exposing vulnerabilities to reward hacking. While recent Rubrics-as-Reward (RaR) methods attempt to recover this structure through explicit criteria, generating rubrics that are simultaneously reliable, scalable, and data-efficient remains an open problem. We introduce Auto-Rubric as Reward (ARR), a framework that reframes reward modeling from implicit weight optimization to explicit, criteria-based decomposition. Before any pairwise comparison, ARR externalizes a VLM’s internalized preference knowledge as prompt-specific rubrics, translating holistic intent into independently verifiable quality dimensions. This conversion of implicit preference structure into inspectable, interpretable constraints substantially suppresses evaluation biases including positional bias, enabling both zero-shot deployment and few-shot conditioning on minimal supervision. To extend these gains into generative training, we propose Rubric Policy Optimization (RPO), which distills ARR’s structured multi-dimensional evaluation into a robust binary reward, replacing opaque scalar regression with rubric-conditioned preference decisions that stabilize policy gradients. On text-to-image generation and image editing benchmarks, ARR-RPO outperforms pairwise reward models and VLM judges, demonstrating that explicitly externalizing implicit preference knowledge into structured rubrics achieves more reliable, data-efficient multimodal alignment, revealing that the bottleneck is the absence of a factorized interface, not a deficit of knowledge.
PaperID: 5362, Poster
Abstract: Implicit neural representations (INRs) provide a flexible framework for modeling high-dimensional continuous fields, but their training is often inefficient due to uniform subsampling that ignores spatial heterogeneity. Existing adaptive sampling methods partially address this issue by prioritizing high-error samples, but typically operate at the point level, often leading to redundant sampling in localized regions and insufficient coverage of the domain. We propose ACES (Adaptive Coverage-aware Efficient Sampling), a structured sampling framework that improves training efficiency by decoupling coverage and importance. ACES constructs adaptive spatial partitions to ensure domain coverage and reduce redundancy, and applies region-level importance weighting to prioritize informative regions during training. We provide a theoretical analysis showing that adaptive partitioning reduces gradient variance by increasing within-region homogeneity, and that controlled bias in region-level weighting may improve optimization efficiency relative to standard unbiased estimators. Experiments on scientific field learning tasks demonstrate that ACES achieves faster convergence and lower error than uniform and pointwise adaptive sampling baselines, with the largest gains in fields with highly localized complexity.
PaperID: 5363, Poster
Authors: Fariha Taskin, Junsouk Choi, Hyunwoong Chang
Abstract: We propose a structure learning method for directed acyclic graph (DAG) models of zero-inflated multivariate count data. Although causal structure learning with zero-inflated Poisson models is identifiable, existing algorithms rely on local search over the DAG space, leading to poor scalability and a tendency to converge to suboptimal solutions. To address these limitations, we develop a novel order-based scoring framework that searches over variable orderings rather than directly over DAGs. Building on this framework, we further extend the method to multiple heterogeneous datasets by identifying a shared variable ordering while allowing each dataset to have its own edge set. We implement a stochastic hill climbing algorithm with a random-to-random (R2R) proposal operator to efficiently explore the ordering space. Simulation studies show that the proposed method outperforms competing approaches and improves graph estimation accuracy under increasing heterogeneity when a shared causal ordering is preserved. We apply the method to single-nucleus RNA sequencing data from a major depressive disorder study.
PaperID: 5364, Poster
Authors:
Francesco Correnti, Gabriele Magrini, Marco Mistretta, Niccolò Biondi, Pietro Pala, Alessandro Ramalli, simona fiori, Andrew Bagdanov, Matteo LengeAbstract: Magnetic Resonance Imaging (MRI) is widely regarded as the optimal sensor for fetal brain analysis due to its superior soft tissue contrast and anatomical detail. However, its high cost and operational burden make it invasive and difficult to obtain at scale. Ultrasound (US), in contrast, is cheap, safe, and routinely acquired, and as a result it has produced substantially larger datasets and a growing ecosystem of pretrained models. This asymmetry raises a natural question: Can we teach a US-only model to understand fetal MRI from only a limited set of examples? The standard recipe, training a foundation model on subject-to-subject paired MRI-US scans, is not viable since no such paired fetal dataset is publicly available. In this paper we address this gap with Weakly Aligned Spatiotemporal Pairs (WASP), a framework that formulates cross-modal correspondence as an entropic Optimal Transport problem driven by clinical metadata — in particular, Gestational Age (GA) and diagnostic planes — enabling the fitting of a lightweight alignment head that lifts MRI representations into the US latent space, without fine-tuning the backbone. Empirically, WASP yields its largest gains when MRI is unseen by the model during pretraining (on USFM, GA estimation error drops from 21.9 to 16.3 days and standard plane classification accuracy climbs from 61.9% to 74.5%), while still providing refining improvements for backbones pretrained on both modalities (e.g., BioMedParse GA estimation error from 6.5 to 5.6 days and SAM-Med2d plane accuracy from 83.3% to 88.1%).
PaperID: 5365, Poster
Abstract: Building capable visual web agents demands precise grounding, long-horizon reasoning, and robust interaction with dynamic, real-world websites. Despite rapid progress, the strongest systems remain largely proprietary, while open agents still depend heavily on supervised post-training over large collections of curated web trajectories. This dependence creates a major scalability bottleneck: high-quality demonstrations are expensive to collect, and static datasets offer limited coverage of the diverse, ever-changing open web. Although online reinforcement learning (RL) has shown promise for text-based agents, its potential for training visual web agents directly on live websites remains largely unexplored. In this paper, we introduce OpenWebRL, an open framework for training visual web agents with online multi-turn RL on real websites. OpenWebRL covers the full training pipeline, including task selection, supervised warm-starting, live-browser execution, multimodal context management, trajectory success judging, and efficient multi-turn policy optimization. Using this framework, we train OpenWebRL-4B, which establishes a new open-source state of the art on challenging live-web benchmarks. With only 0.4K initialization trajectories and 2.2K open-ended training tasks, OpenWebRL-4B achieves 67.0% success on Online-Mind2Web and 64.0% on DeepShop, outperforming prior open agents of similar or larger scale and remaining competitive with proprietary systems including OpenAI and Gemini CUA. Beyond strong benchmark performance, OpenWebRL systematically identifies the key design choices that make online RL effective for visual web agents. Overall, our work offers a practical path toward building more capable, reproducible, and cost-efficient open web agents. We will release our training data, models, and code to support future research.
PaperID: 5366, Poster
Authors:
Minghao Sun, Hanqun Cao, Zhou Zhang, Chen Wei, Liang Wang, Tianrui Jia, ZHIYUAN LIU, Tianfan Fu, Robert Tang, Yejin Choi, Pheng-Ann Heng, Fang Wu, Yang ZhangAbstract: RNA inverse folding seeks sequences that reliably fold to a target 3D backbone under physiological conditions. Multi-objective preference optimization on RNA is unstable for two domain-specific reasons: heterogeneous noise across physical proxies (deterministic 2D folding, stochastic 3D prediction, minimum free energy), and sequence–structure degeneracy that admits compositional reward hacking via GC enrichment.We introduce RiboPO, a preference-optimization framework that builds preference pairs as \varepsilon-Pareto-dominance relations on standardized per-metric features, gates winners with a structural quality threshold against GC-driven shortcuts, and trains a frozen-reference DPO policy under a decreasing-margin curriculum. A heuristic anchored-KL drift characterization and a pair-level Rényi-2 off-policy bias bound motivate the experimental design; a clipped importance-corrected variant extends usable rounds beyond the static-pair regime. On DAS benchmark, specialized RiboPO variants improve the gRNAde base on SSTT axes: scMCC reaches 0.71 at R=4 (+17% over 0.61) and 0.643 at the thermodynamic-surplus point (+6.1%, paired Wilcoxon p<10^-3); MFE improves -7.3%; target-structure probability P(S_0) rises 0.00128 \to 0.0155 (~12×) with ensemble-defect -6.9 to -7.9% (p=6.1×10^-5 at R=4), consistent with mass concentration toward target rather than redistribution to an off-target basin; designability rises +6.9 pp %pLDDT\geq0.7 and +6.4 pp %RMSD\leq8 Å. Under matched-utility same-pool reranking, RiboPO best-of-8 matches gRNAde best-of-64 (pass@1 0.408) at 1/8 the per-backbone budget. On the leakage-corrected subset, RiboPO dominates every structural and thermodynamic axis. Codes are available at: [https://anonymous.4open.science/r/ribopo-8D48/](https://anonymous.4open.science/r/ribopo-8D48/).
Authors:
Jucheng Hu, Zhangquan Chen, Yulin Chen, Chengjie Hong, Liang Zhou, Tairan Wang, Sifei Li, Giulio Zhu, Feng Zhou, Yiheng Zeng, Suorong Yang, Dongzhan ZhouAbstract: Deciphering animal intent is a fundamental challenge in computational ethology, heavily constrained by “semantic aliasing,” where identical external signals (e.g., a feline purr) map to vastly different internal states depending on physiological context. Existing Multimodal Large Language Models (MLLMs) are “modality blind” to high‑frequency biological time‑series data, restricting them to superficial behavioural pattern‑matching rather than genuine latent state reasoning. To bridge this gap, we introduce Meow‑Omni 1, the first open‑source quad‑modal MLLM purpose‑built for computational ethology. It natively fuses visual, audio, and physiological time‑series modalities with textual reasoning. Through targeted architectural model surgery, we integrate specialized scientific encoders into a unified backbone, formalizing intention inference via structural causal models. Evaluated on MeowBench, a novel expert‑verified quad‑modal benchmark, Meow‑Omni 1 achieves state‑of‑the‑art intent recognition accuracy (71.16 %), significantly outperforming leading vision‑language and omni‑modal baselines. We release the complete open‑source pipeline including model weights, training framework, and curated dataset to establish a scalable, robust paradigm for inter‑species communication and to advance foundation models toward real‑world veterinary diagnostics and wildlife conservation.
Abstract: Rotation invariance is a fundamental requirement across many computer vision tasks. Historically, this inductive bias has been encoded through hand-crafted rotation-invariant representations. These are compact, interpretable, and fast to compute, but they come at the cost of descriptive power. More recently, architectures achieve inductive bias through learned representations. These are highly descriptive and achieve strong empirical performance, at the cost of efficiency and interpretability. In this work, we propose an alternative at the intersection of both paradigms. We introduce the selective disk bispectrum (SDB), a complex-valued rotation-invariant vector that preserves all information about the image except its orientation. Our key theoretical contributions are the selective disk bispectrum, its inversion, its (reduced) spatial and computational complexities (compared to the full disk bispectrum), and its expectation and variance under noise. Furthermore, we propose a numerical SDB approximation and provide theoretical guarantees for its accuracy and rotation invariance. Empirically, we validate SDB's invariance and robustness to noise classification tasks. We test our reconstruction algorithm on multi-reference alignment of rotated images.
Abstract: We propose and study an online version of min-max optimization based on cumulative saddle points under a variety of performance measures beyond convex-concave settings. After first observing the incompatibility of (static) Nash equilibrium (SNE-Reg_T) with individual regrets even for strongly convex-strongly concave functions, we propose an alternate static duality gap (SDual-Gap_T) inspired by the online convex optimization (OCO) framework that is compatible with the individual regrets . We provide algorithms that achieve sub-linear regret bounds for (SDual-Gap_T) and the individual regrets and a novel dynamic saddle point regret (DSP-Reg_T), which we suggest naturally represents a min-max version of the dynamic regret in OCO. We derive our bounds for (SDual-Gap_T) and DSP-Reg_T under strong convexity-strong concavity and a min-max notion of exponential concavity (min-max EC), and in addition we establish a class of functions satisfying min-max EC that captures a two-player variant of the classic portfolio selection problem. Finally, for a dynamic notion of regret compatible with individual regrets, we derive bounds under a two-sided Polyak-\Lojasiewicz (PL) condition.
PaperID: 5370, Poster
Abstract: Reflections reveal scene content beyond the camera's direct field of view, but faithfully recovering this hidden information is challenging as reflective observations are heterogeneous, incomplete, and physically modulated. Existing methods either rely on explicit reflection modeling with privileged geometric priors or defer to generative synthesis that may hallucinate plausible but unfaithful content. To address these challenges, we propose ReFlex, an observation-constrained framework that casts faithful panorama reconstruction as a controlled inverse problem. We first develop an \emphObservation-Consistent Estimator with hierarchical token prediction and signed routing to integrate heterogeneous reflection cues while stabilizing global structure. Subsequently, we propose a \emphReliability-Guided Refiner, which estimates spatial reliability from token predictions and guides a frozen diffusion model to refine uncertain regions without overwriting observation-supported structures. To validate the effectiveness of ReFlex, we further build ReFlex-Bench, a paired benchmark with 110,760 reflective RGB inputs and aligned target panoramas across diverse scenes, reflective geometries, and observation regimes. Experiments show that ReFlex delivers superior reconstruction fidelity without normals, field-of-view annotations, masks, or other privileged information, and obtains the best performance on eight of nine downstream full-view perception protocols.
PaperID: 5371, Poster
Authors: Gaetan De Waele, Marek Wydmuch, Krzysztof Dembczynski, Wojciech Kotlowski, Willem Waegeman
Abstract: Identifying molecular structures from LC-MS/MS spectra is a central problem in computational metabolomics, often approached by predicting molecular fingerprints and retrieving candidates via similarity search. Despite widespread use, the relationship between training objectives for fingerprint prediction and downstream retrieval performance remains poorly understood. In this work, we show that these objectives are fundamentally misaligned. Adopting a decision-theoretic perspective, we derive novel regret bounds that characterize when Bayes-optimal predictors for fingerprint similarity diverge from those for molecular retrieval. Our analysis reveals that optimizing standard similarity-based losses can provably degrade retrieval performance, and that the extent of this mismatch depends on the similarity structure of candidate sets. Empirically, we validate our theory on the MassSpecGym benchmark, demonstrating a Pareto frontier between fingerprint accuracy and retrieval metrics across commonly used loss functions. These results expose an inherent trade-off in fingerprint-based molecular identification and provide principled guidance for the design of learning objectives in computational mass spectrometry.
Abstract: Solving large scale Optimal Transport (OT) in machine learning typically relies on sampling measures to obtain a tractable discrete problem. While the discrete solver's accuracy is controllable, the rate of convergence of the discretization error is governed by the intrinsic dimension of our data. Therefore, the true bottleneck is the knowledge and control of the sampling error. In this work, we tackle this issue by introducing novel estimators for both sampling error and intrinsic dimension. The key finding is a simple, tuning-free estimator of \textOT_c(\rho, \hat\rho) that utilizes the semi-dual OT functional and, remarkably, requires no OT solver. Furthermore, we derive a fast intrinsic dimension estimator from the multi-scale decay of our sampling error estimator. This framework unlocks significant computational and statistical advantages in practice, enabling us to (i) quantify the convergence rate of the discretization error, (ii) calibrate the entropic regularization of Sinkhorn divergences to the data's intrinsic geometry, and (iii) introduce a novel, intrinsic-dimension-based Richardson extrapolation estimator that strongly debiases Wasserstein distance estimation. Numerical experiments demonstrate that our geometry-aware pipeline effectively mitigates the discretization error bottleneck while maintaining computational efficiency.
PaperID: 5373, Poster
Abstract: Diffusion Large Language Models (DLMs) are being actively explored as promising alternatives to autoregressive models due to their fast and flexible generation capabilities. To optimize inference in DLMs, recent studies have proposed a variety of decoding algorithms. However, the tendency of these methods to mimic autoregressive generation limits the speed and flexibility of DLMs, while also preventing the methods from fully realizing their potential. Moreover, this rigid approach leads to error propagation, while existing sampling-based correction methods incur prohibitive computational costs. To overcome this, we propose Renoise Consistency, a novel post-hoc approach that enables effective self-correction by maximizing train--inference consistency and can be applied to any decoding algorithm with a flexible computational budget. Furthermore, we introduce Adaptive Block Drafting, which leverages consistent internal representations of DLMs to significantly reduce the overall computational cost. When combined, our proposed method achieves up to a 9.00% performance improvement over the semi-autoregressive baseline, along with a 25.85% reduction in forward steps.
PaperID: 5374, Poster
Abstract: Flow and diffusion-based generative models achieve state-of-the-art quality, but require multiple forward passes at inference time, making them computationally costly. Consistency models overcome this by learning a flow map directly, enabling few‑step generation and dramatically reducing the computation needed at test time. However, formal statistical guarantees for such one‑step generators are scarce. Existing results either suffer from the curse of dimensionality or rely on impractical loss functions. We investigate whether one‑step flow-based models can escape the curse of dimensionality under the assumption that the data distribution is generated by a low-dimensional surface perturbed with Gaussian noise. In this setup, we derive approximation and generalization error bounds for the flow map estimate that circumvent the curse of dimensionality. Our findings provide the first principled explanation for practical success of one-step flow-based generative models.
PaperID: 5375, Poster
Abstract: Multi-View Stereo (MVS) and Monocular Depth Estimation (MDE) provide complementary cues for dense 3D reconstruction. MDE offers rich contextual and structural priors but lacks explicit geometric constraints, whereas MVS exploits multi-view geometry but often struggles in ambiguous regions, such as weakly textured or reflective surfaces. Existing MDE-assisted MVS methods mainly use monocular cues in a unidirectional manner, leaving the interaction between monocular representations and multi-view cost volumes insufficiently explored. To address this limitation, we introduce BridgeMVS, a unified cascaded MVS framework that bridges monocular depth features and multi-view cost-volume representations through Bidirectional Dynamic Fusion (BDF). At each cascade stage, BDF performs dynamic mono-to-volume and volume-to-mono updates: monocular structural features are injected into cost-volume regularization, while multi-view geometric evidence recalibrates intermediate monocular representations. The enhanced monocular features are propagated across stages together with MVS representations and supervised by stage-wise monocular losses, establishing an auxiliary structural guidance path for the MVS branch. Unlike stage-isolated designs, BridgeMVS preserves fused representations from coarse to fine stages, allowing low-resolution cross-branch interactions to guide subsequent high-resolution depth estimation. Moreover, BDF is designed as a plug-and-play module and can be conveniently integrated into existing cascaded MVS frameworks. Extensive experiments show that BridgeMVS achieves state-of-the-art performance on the DTU, Tanks and Temples, and ETH3D benchmarks.
PaperID: 5376, Poster
Abstract: Transformer feed-forward blocks combine several design choices: squared activations, gating, and tied versus untied parameter sharing-whose separate contributions to expressivity remain poorly understood. We introduce \emphCP-MLPs, a tensor framework in which each multiplicative unit contributes a rank-1 interaction with three roles: detecting an input feature, reading a value, and writing to the output. Tied gated blocks such as ReLU^2 force the detector and value roles to share the same direction; untied gated blocks such as ReGLU and SwiGLU let them differ. We show that this distinction yields exact, degree-wise tensor-rank separations: untied detector-value units are strictly more width-efficient for indefinite quadratic interactions and role-asymmetric monomials, while tied units are structurally matched to symmetric pure-power targets. Tying is therefore not inherently better or worse where its benefit depends on the structure of the target. In experiments, models below the required width hit an irreducible approximation floor where the loss does not improve with extra optimization budget. The detector-value advantage also persists in trained gated MLPs at matched parameter count.
PaperID: 5377, Poster
Authors:
Nikalal Helessage, Wanqing Li, Chris Bunn, Philip O OgunbonaAbstract: Semantic visual classification commonly aligns visual embeddings with semantic spaces induced by pretrained text encoders. However, such spaces are typically high-dimensional, weakly aligned with visual feature geometry, and insufficiently discriminative for downstream classification tasks. We propose TaSS, a framework for learning a task-adaptive low-dimensional semantic space that preserves semantic structure while explicitly optimizing class separability. TaSS is constructed through a parameterized geometry transformation of pretrained text embeddings that enlarges inter-class margins via distance reshaping and improves intra-class compactness through prototype-preserving semantic constraints. The resulting semantic space maintains meaningful semantic relationships while producing a discriminative geometry tailored for classification. Frozen visual encoder features are subsequently aligned to TaSS using a lightweight MLP, enabling improved discriminative learning without fine-tuning pretrained backbones. We further provide theoretical analysis showing that TaSS preserves semantic consistency while enforcing enhanced inter-class separation and controlled intra-class variance. Extensive experiments on skeleton-based human action recognition, video classification, and image classification benchmarks demonstrate consistent improvements over existing semantic alignment approaches, with gains exceeding 3 percentage points on average and up to 8.5 points on challenging datasets.
PaperID: 5378, Poster
Abstract: Time series anomaly detection (TSAD) remains challenging not only because anomaly labels are scarce, but also because temporal anomalies are highly context-dependent. Existing methods often rely on unsupervised or surrogate objectives, producing anomaly scores indirectly rather than learning explicit normal--anomalous distinctions. We propose Counterfactual Pairing with Anomaly Semantics (CAPS), a supervision-recovery framework for TSAD. CAPS formulates temporal anomalies as mechanism-induced effects on normal temporal evolution and recovers matched supervision without target-domain anomaly labels. It learns transferable anomaly semantics from simulated normal--anomalous pairs, disentangles them from normal temporal structure, and organizes them into family-wise modes. Instead of directly using simulated anomalies as target positives, CAPS instantiates the learned semantics as residual-form anomaly effects on target-domain reference trajectories via residual MoE generation. The resulting matched counterparts provide supervision-aligned signals for boundary-oriented detector learning. Experiments on nine benchmark datasets show that CAPS outperforms unsupervised and injection-based baselines, is competitive with supervised-reference baselines, and provides interpretable evidence through anomaly-effect generation and expert association.
PaperID: 5379, Poster
Abstract: Ocean dynamics are inherently chaotic, yet existing machine learning ocean models produce only deterministic forecasts. We introduce Njord, a probabilistic data-driven model for ocean forecasting, applicable to both global and regional domains. Njord combines a deep latent variable framework with a graph neural network architecture, enabling sampling each forecast step in a single forward pass. We apply Njord globally at 0.25° resolution and regionally to the Baltic Sea at 2 km resolution. To scale to these large ocean grids we introduce K-means cluster meshes that adapt to irregular sea surface geometry. Experiments demonstrate strong performance on both domains compared to deterministic machine learning baselines, while also providing uncertainty estimates from the sampled ensemble forecasts. On the global OceanBench benchmark, Njord achieves the lowest errors on average across upper-ocean variables when evaluated against real-world observations, with the largest improvements in surface temperature prediction.
PaperID: 5380, Poster
Authors:
Peiyan Li, Yixiang Chen, Yuan Xu, Jiabing Yang, Xiangnan Wu, Jun Guo, Nan Sun, Long Qian, Xinghang Li, Xin Xiao, Minghui Zhang, Jing Liu, Nianfeng Liu, Tao Kong, Yan Huang, Liang WangAbstract: Robotic manipulation requires understanding both the 3D spatial structure of the environment and its temporal evolution, yet most existing policies neglect one or both aspects. They often rely on 2D visual observations or backbones pretrained on static image--text pairs, which leads to high data requirements and limited comprehension of environment dynamics. To address this, we introduce SpatialVAM, the first 3D Video Action Model that simultaneously predict spatial-aware multi-view heatmap videos and RGB videos. Our key insight is that this design naturally injects 3D information into video foundation models while aligning the representation format between video pretraining and action finetuning. Extensive experiments demonstrate that SpatialVAM enables data-efficient, robust, generalizable, and interpretable manipulation. With only ten demonstration trajectories and no additional pretraining, SpatialVAM handles challenging long-horizon and contact-rich tasks, generalizes to out-of-distribution settings, and predicts realistic future videos. Evaluations on Meta-World (22%\uparrow), RoboCasa (10%\uparrow) and real-world robotic platforms (16%\uparrow) show that SpatialVAM consistently outperforms other video action models, vision language action models and 3D-based policies, establishing a new state-of-the-art in data-efficient multi-task manipulation.
Abstract: We propose a distributional theory of how hypernymy---the ``is-a'' relation between general and specific concepts---is encoded geometrically in language representations. Starting from the empirically verified assumption that words closer on the WordNet hypernym graph co-occur more often, we characterize theoretically the spectrum of the resulting embedding Gram matrix of word2vec embeddings. Under mild positivity and decay conditions on the co-occurrence kernel, we prove that the leading eigenvectors first separate broad taxonomic branches and then progressively finer sub-branches, producing a \emphhierarchical splitting geometry with a coarse-to-fine spectral organization that mirrors the tree. We confirm these predictions in word2vec embeddings across many sampled WordNet subtrees, and show that the same signature extends strikingly well to Gemma 2B unembeddings. Our results indicate that hierarchical concept geometry in LLMs need not reflect a hierarchy-specific functional mechanism, but emerges from the spectral structure of pairwise word statistics.
Abstract: Chain-of-thought (CoT) reasoning enables large language models (LLMs) to solve complex problems by generating intermediate reasoning steps. While much attention has been paid to the length and content of these reasoning chains, far less is known about their internal geometry. We study the geometry of CoT trajectories in the hidden state space of transformer models, formalizing each reasoning chain as a discrete curve in \mathbbR^d and characterizing it through spectral, positional, and kinematic geometric functionals. We introduce the effective dimension d_\rho as a measure of trajectory complexity and show theoretically that trajectories with flatter eigenvalue spectra correspond to harder tasks, as they explore more of the hidden dimensions. Lastly, we explore how kinematic features of the trajectory, mean position, positional dispersion, initial hidden states, mean velocity, mean speed, and speed dispersion, can be used to predict solution correctness before generation is complete, and may inform future early-stopping strategies. Experimentally, on mathematical reasoning problems from the MATH500 dataset, d_\rho achieves 0.93 AUC in distinguishing easy from hard problems, while kinematic features show promise for predicting correctness from only the first 20% of generated tokens. These correctness signatures transfer across questions of varying difficulty, establishing that the shape of a model's internal reasoning trajectory is a principled window into both task hardness and solution quality.
Authors: Haixiang Sun, Andrew Liu
Abstract: Outliers are essential for evaluating and improving the robustness of machine learning systems, especially when future distributions may differ significantly from historical training data. In high-stakes applications, robustness often depends on rare cases that finite datasets fail to capture, making simple resampling or perturbation insufficient for stress scenario generation. Existing outlier synthesis methods typically rely on sparse neighborhoods, low support latent regions, or classifier boundary crossings, which can be heuristic, unstable, and tied to specific modalities or architectures. We therefore propose Sinkhorn Boundary Outlier Generation (SBOG), a structured framework for latent-space outlier generation that couples Sinkhorn optimal transport geometry with distributionally robust boundary modeling. The resulting Sinkhorn-induced support cost guides the sampler toward weakly supported boundary regions, while semantic constraints prevent uncontrolled drift from the intended context, yielding controlled deviations from the in-distribution reference measure rather than arbitrary sparse-region samples. Experiments on time series anomaly generation and image outlier synthesis show that our framework produces informative, semantically controlled outliers and improves downstream robustness evaluation across modalities, providing a foundation for stress scenario generation beyond empirical support.
PaperID: 5384, Poster
Authors: Vaibhav Agrawal, Varghese P Kuruvilla, Harsh Rangwani, Ravi Kiran Sarvadevabhatla
Abstract: While text-to-image (T2I) diffusion models exhibit strong semantic disentanglement at the object level, they struggle to localize fine-grained object parts. Although the latent visual features in these models are sufficiently fine-grained, the text–image interaction (cross-attention) fails to effectively exploit this information. This results in coarse and ambiguous localization, particularly for spatially distinct components such as PartConcepts, a mechanism that encodes textual part descriptions into compact, learnable representations . These PartConcept tokens are trained to selectively attend to their corresponding part regions in the image. To evaluate the part-level understanding of PartConcept tokens, we first probe its effectiveness for Open-Vocabulary Part Segmentation (OVPS) benchmark, our method outperforms dedicated segmentation baselines , establishing a new state-of-the-art. We further evaluate our method on the more complex Pascal-Part benchmark for ). Here, we outperform all baselines by a very significant margin. Finally, our qualitative and quantitative results show that this strong localization of PartConcept tokens directly enables bind well to corresponding PartConcept tokens, without explicit attribute control mechanisms. This indicates that PartConcept tokens compose seamlessly with other text tokens, highlighting their applicability for
PaperID: 5385, Poster
Authors: Ziye Ma
Abstract: Non-convex optimization is central to machine learning and AI, yet unlike its convex counterpart, it still lacks a general and principled framework. In this work, we try to highlight the importance of combining noisy perturbations with over-parametrization mathematically under the classic setting of matrix completion (MC). Matrix completion is an important non-convex recovery problem in which only a subset of a matrix’s entries is observed, and the remaining entries must be recovered using low-rank constraints. This problem is notoriously difficult because observations are limited. However, we show that by injecting \epsilon-level perturbations into the nullspace of the observation mask, we can construct noisy matrix sensing surrogates whose solutions remain \mathcalO(\epsilon)-close to the ground truth with valid restricted isometry property (RIP) constants. The existence of RIP further enables powerful tensor frameworks to be applied with theoretical guarantees. This approach reflects a broader empirical trend in which stochasticity or noise improves the curvature of the optimization landscape, allowing over-parameterized (a.k.a large) models to better separate signal from noise. Assuming each matrix entry is observed independently, we establish quantitative recovery guarantees without explicit requirements on sampling rate or incoherence, and discuss how such assumptions can further strengthen our results.
Abstract: Vision Language Models (VLMs) offer the exciting possibility of processing text as rendered images, bypassing the need for tokenizing the text into long token sequences. Since VLM image encoders map fixed-size images to a fixed number of visual tokens, varying rendering resolution provides a fine-grained compression knob. However, accuracy deteriorates quickly as compression increases: characters shrink below the vision encoder's effective resolution, making them indistinguishable. To address this, we propose LensVLM, an inference framework and post-training recipe that enables VLMs to scan compressed images, then selectively expand only the relevant images to their uncompressed form via learned tools. Building on Qwen3.5-9B-Base, LensVLM maintains accuracy comparable to the full-text upper bound at 4.3× effective compression and outperforms retrieval-based, text- and visual-compression baselines up to 10.1× effective compression across seven text QA benchmarks. LensVLM also generalizes to multimodal document and code understanding tasks, with the accuracy gain over baselines growing as compression increases. Our analysis validates this approach: training makes visual compression robust to rendering choices, and as compression grows the model increasingly relies on expanded content rather than unreliable visual reading. The analysis also yields practical tool-choice guidance: text expansion is preferable for rendered text, while high-resolution image expansion suits native documents whose layout cues carry task-relevant information.
PaperID: 5387, Poster
Abstract: The RISC-V Vector Extension (RVV) provides SIMD-style vector execution for accelerating deep learning (DL) inference on edge processors. However, existing low-rank compression methods mainly select ranks according to accuracy preservation and theoretical floating-point operation (FLOP) reduction, without considering whether the selected ranks are profitable under a target RVV effective vector length. As a result, a low-rank factorization that reduces FLOPs may still introduce tail-handling overhead or even increase the RVV instruction count under specific vector configurations. To address this problem, we propose VALOR, a low-rank restructuring framework for efficient DL deployment on RISC-V processors. Given a pretrained model and a target RVV configuration, VALOR identifies decomposable layers in the pretrained model, filters rank candidates using an RVV instruction-profitable condition, and searches for a global restructuring strategy with minimal accuracy loss. Furthermore, to avoid repeated end-to-end evaluation, VALOR uses a Hessian-aware accuracy predictor that combines layer-wise factorization error with inter-layer coupling. Experiments on Spike and a real RISC-V vector processor, SpacemiT X60, show that VALOR achieves a better accuracy-efficiency trade-off than baseline methods.
PaperID: 5388, Poster
Abstract: Aerial-Ground Person Re-Identification (AG-ReID) aims to match pedestrian images captured by unmanned aerial vehicles (UAVs) and ground-based surveillance cameras. The task remains highly challenging due to severe viewpoint discrepancies, frequent occlusions, and substantial domain gaps between aerial and ground imagery. Beyond these observable factors, existing methods often overlook two fundamental structural issues: cumulative positional drift during hierarchical feature extraction and semantic inconsistency across cross-view feature distributions. To address these challenges, we propose ProAlign, a Progressive Positional and Prototype-Guided Alignment Network for AG-ReID. ProAlign comprises two key components: a Layer-wise Progressive Positional Embedding (LPPE) module and a View-Aware Prototype Contrastive Learning (PCL) module. LPPE performs adaptive positional calibration throughout transformer layers by employing conditional positional encoding in shallow layers to mitigate local spatial distortions, while introducing learnable static positional embeddings in deeper layers to reinforce global semantic priors. Meanwhile, PCL maintains view-specific aerial and ground prototypes for each identity and enforces cross-view semantic consistency via a dual-prototype contrastive objective. Extensive experiments on the challenging CARGO benchmark demonstrate the effectiveness of ProAlign. In the aerial-ground setting, ProAlign surpasses previous state-of-the-art methods by 13.89% mAP and 9.39% Rank-1 accuracy.
Authors:
Peiyu Zhuang, Jianquan Yang, Haodong Li, Zhuoying Cai, Ruitao Xie, Jishen Zeng, Baoying Chen, Jiwu Huang, Xiaochun CaoAbstract: Text-driven image editing has advanced rapidly, but reliably localizing these manipulations requires image manipulation localization (IML) models trained on large pixel-annotated datasets, and there is still no low-cost way to obtain such training data at scale. We observe that these data already exist in disguise: public editing datasets contain millions of structurally identical (original, edited) pairs to IML training samples, lacking only pixel-level masks. Recovering these masks automatically is non-trivial: pixel differencing is overwhelmed by diffusion-induced perturbations across all pixels, and instruction-only grounding localizes only what the prompt describes, missing unintended editor side-effects. We propose SIGMA (Semantic-difference Instruction-Grounding Mask Annotator), which performs semantic-feature differencing in a vision foundation backbone and injects an instruction-derived spatial prior into this visual stream via bidirectional cross-modal refinement, amplifying the difference signal at intended-edit regions when the editor faithfully realizes user intent. SIGMA is trained in two complementary stages: Stage I supervises on inpainting masks; Stage II closes the diffusion-domain shift via VAE-roundtrip noise calibration, EMA self-training, and an edit-noise disentanglement loss. SIGMA outperforms existing automatic mask generators on five benchmarks (+12.20% F1, +11.16% IoU). When applied to public editing corpora, it produces a ~1.1M IML training set that improves six diverse detectors by +18.34% F1 across five datasets, turning previously unused editing data into a model-agnostic supervisory resource for IML. For reproducibility, our anonymized code is available at https://anonymous.4open.science/r/SIGMA-14FE.
PaperID: 5390, Poster
Abstract: Long-context inference in Large Language Models (LLMs) suffers from high latency because attention computation and memory traffic scale linearly during decoding. KV cache pruning reduces this cost by retaining only a subset of cached tokens, but its effectiveness depends on whether the pruning proxy can recover the tokens with the largest exact attention logits under the current query. Existing proxies suffer from ranking misalignment: static heuristics rely on historical statistics and thus ignore the current query, while hashing-based retrieval proxies rank candidates in a discrete metric space that does not coincide with attention's inner product. To avoid these problems, we propose LAPrune, a KV pruning framework that uses additive vector quantization to build an efficient attention-logit-aligned proxy. LAPrune represents each cached key as a sum of learned codewords and decomposes the query-key dot product into lookup-and-add operations over query-codeword inner products. This replaces high-dimensional dot products with lightweight scalar lookups while preserving fidelity to the original attention ranking. LAPrune requires no base-model finetuning and can be integrated into long-context decoding with minimal modification. Across long-context evaluations, LAPrune improves exact-logit Top-k recovery and perplexity over hashing-based pruning proxies, achieving up to 2.73× decoding speedup over FlashAttention-2 and 4.83× end-to-end speedup over the vanilla baseline at 128K context. Our code is available at [Anonymous/LAPrune](https://anonymous.4open.science/r/LAPrune-VQ).
Abstract: We consider the optimization problem of an expensive-to-evaluate black-box function, in which we can obtain noisy function values in parallel. For this problem, parallel Bayesian optimization (PBO) is a promising approach, which aims to optimize with fewer function evaluations by selecting a diverse input set for parallel evaluation. However, existing PBO methods suffer from poor practical performance or lack theoretical guarantees. In this study, we propose a PBO method, called randomized kriging believer (KB), based on a well-known KB heuristic and inheriting the advantages of the original KB: low computational complexity, a simple implementation, versatility across various BO methods, and applicability to asynchronous parallelization. Furthermore, we show that our randomized KB achieves Bayesian expected regret guarantees. We demonstrate the effectiveness of the proposed method through experiments, including those on real-data emulators.
Abstract: Detecting AI-generated text is an important but challenging problem. Existing likelihood-based detection methods are often sensitive to content complexity and may exhibit unstable performance. In this paper, our key insight is that modern Large Language Models (LLMs) undergo alignment (including fine-tuning and preference tuning), leaving a measurable distributional imprint. We theoretically derive this imprint by abstracting the alignment process as a sequence of constrained optimization steps, showing that the log-likelihood ratio can naturally decompose into implicit instructional biases and preference rewards. We refer to this quantity as the Alignment Imprint. Furthermore, to mitigate the instability in high-entropy regions, we introduce Log-likelihood Alignment Preference Discrepancy (LAPD), a standardized information-weighted statistic based on alignment imprint. We provide statistical guarantee that alignment-based statistics dominate Fast-DetectGPT in performance. We also theoretically show that LAPD strictly improves the unweighted alignment scores when the aligned and base models are close in distribution. Extensive experiments show that LAPD achieves an improvement 45.82% relative to the strongest existing baselines, yielding large and consistent gains across all settings
PaperID: 5393, Poster
Authors: Ruyu Zhou, Fang Liu
Abstract: Differential Privacy (DP) provides a rigorous framework for quantifying privacy guarantees. While most methods achieve DP by injecting calibrated random noise, the intrinsic randomness of certain procedures such as sampling can yield DP guarantees ``for free''. In this work, we develop a unified R\'enyi divergence framework to characterize the inherent DP guarantees of sampling from a Bayesian posterior distribution. We show that the privacy loss incurred by One-Posterior-Sampling (OPS) is governed by posterior exponential moments of the single-record log-likelihood ratio, and establish explicit DP guarantees under uniformly bounded, sub-Gaussian, or sub-exponential tail regimes. We apply the framework to representative Bayesian models, including categorical likelihoods with arbitrary priors, Gaussian-Gaussian conjugate pair, generalized linear models from the exponential family and linear regression with Gaussian likelihood, both with Gaussian priors on the regression coefficients. Our results elucidate sample size, model structure, neighboring relations in DP (bounded or unbounded), and how prior if applicable, jointly determine the inherent DP guarantees of OPS. Furthermore, our results recover the existing DP guarantees for OPS as special cases -- often with tighter privacy loss bounds -- and substantially broaden the class of Bayesian models for which the inherent DP guarantees of OPS can be rigorously characterized.
PaperID: 5394, Poster
Abstract: Reaction mechanisms are central to understanding chemical reactivity, guiding reaction prediction, synthesis design, and selectivity rationalization. In chemistry literature, they are universally depicted as curved electron-pushing arrows that capture stepwise electron flow and bond changes. Yet, this vast knowledge remains trapped in unstructured diagrams, forcing data-driven modeling to rely on small or template-expanded datasets. We introduce MechParser, a vision-language framework that converts mechanism diagrams into SM-SMARTS (Sequential Mechanism-Aware SMARTS), a representation encoding electron flow and atom state changes as a cumulative string sequence. At its core, MechParser-VL applies a left-right visual prompt design (arrow diagram alongside an atom-indexed reference) to separate dynamic arrow reasoning from static atom identification, and is trained via a four-stage curriculum on over 300K samples from a geometry-aware synthetic data pipeline and an expert-annotated real-world dataset. On both synthetic and real-world benchmarks, MechParser-VL substantially outperforms much larger proprietary and open-source VLMs despite using only a 4B backbone. By standardizing visual mechanisms into machine-readable data, MechParser lays the groundwork for constructing large-scale mechanism databases from the chemical literature, which in turn can support the training of advanced mechanism-centric models on authentic mechanistic data.
PaperID: 5395, Poster
Abstract: Recent full-duplex spoken dialogue models have demonstrated compelling progress toward human-like interaction, enabling agents to respond with low latency, produce backchannels, and handle user barge-ins. Yet these improvements in conversational dynamics often come with weaker reasoning and instruction-following abilities, revealing a potential tension between interactive dynamics and intelligence capability. In this paper, we argue that such an trade-off is not fundamental: conversational dynamics can instead be learned as a separate real-time decision policy from human dialogue data. To this end, we propose DuplexPO, a reinforcement learning (RL) framework that decouples when to speak from what to say. It preserves the semantic response capability of an instruction-tuned assistant, while optimizing its temporal interaction behavior over selected high-impact windows from long human conversations. To quantitatively optimize these dynamics, we formulate the Factorized Conversational Dynamics Reward (FCDR) to enable fine-grained temporal credit assignment for turn initiation, backchanneling, yielding, and regularized participation. The policy is then optimized with a GRPO-style objective. Experiments show that DuplexPO substantially improves full-duplex behaviors, including timely backchannels, smooth turn-taking, and barge-in handling, while maintaining strong reasoning and instruction-following performance. Moreover, improvements in dynamics-oriented metrics are reflected in better user experience, suggesting that optimizing conversational timing as a standalone objective can promote more natural full-duplex interaction.
PaperID: 5396, Poster
Abstract: Deep learning models achieve high accuracy in signal processing but remain difficult to interpret. Classical matching pursuit and sparse coding offer better transparency but suffer from spectral leakage and grid mismatch when signal frequencies do not align with discrete dictionary atoms. We introduce Continuous Dictionary Pursuit (CDP), a framework for interpretable signal decomposition that learns the parameters of analytical functions directly from data. Unlike traditional methods, CDP optimizes over a continuous manifold of differentiable atoms to resolve the limitations of fixed grids. The algorithm employs a greedy iterative strategy using data-driven priors and gradient-based optimization to isolate individual signal components. We provide a library of parametric atoms that covers periodic, trend, and transient structures, including discontinuous functions modeled through spectral annealing. To ensure a sparse atom set for reconstruction, the framework utilizes a stopping criterion based on a predefined atom budget and residual error convergence. Evaluations on synthetic and real-world datasets show that CDP resolves spectral leakage for the parameterized atom families considered and recovers exact mathematical components. Beyond decomposition, we demonstrate that the extracted atoms can serve as structured curriculum targets for training sequence models, leading to consistent improvements in forecasting accuracy compared to standard end-to-end training. The method offers a glass-box alternative for signal decomposition by providing high reconstruction fidelity alongside direct physical interpretability.
PaperID: 5397, Poster
Abstract: Federated Multiple Instance Learning (MIFL) has emerged as a promising paradigm for privacy-preserving weakly supervised learning, particularly in medical image analysis where instance-level annotations are expensive or unavailable. Existing federated MIL methods mainly rely on parameter aggregation to learn a global model across clients. However, in MIL, the key challenge is not only client-level data heterogeneity, but also the ambiguity of instance-level evidence. Since supervision is only provided at the bag level, each client must independently infer discriminative instances through local attention mechanisms. Under heterogeneous data distributions, different clients may focus on inconsistent or even misleading instances, resulting in cross-client critical evidence drift. To address this problem, we propose FedSEM, a Federated Shared Evidence Memory framework that introduces an evidence-level communication pathway in addition to standard model aggregation. Specifically, each client identifies high-confidence evidence instances using local attention scores and compresses them into compact prototypes with importance scores. The server then refines these prototypes through redundancy removal, diversity selection, and score normalization to construct a shared evidence memory. Importantly, FedSEM does not require transmitting raw patches or dense instance embeddings; it only exchanges a small number of evidence prototypes and periodically broadcasts the refined memory, thereby introducing limited additional communication overhead compared with standard federated training. The shared memory serves as cross-client evidence anchors to guide local attention learning and instance-level representation optimization. By explicitly aligning critical evidence patterns across clients, FedSEM mitigates evidence drift while maintaining communication efficiency and privacy preservation. Extensive experiments under heterogeneous federated settings demonstrate that FedSEM consistently improves MIL performance and generalization over existing federated learning baselines.
PaperID: 5398, Poster
Abstract: This paper characterizes the sharp asymptotics of the \ell_\infty-norm estimation error for estimating the coefficient vector in linear regression under proportional asymptotics. While \ell_2-error is well-documented, we demonstrate that the \ell_\infty-error exhibits a distinct phase transition governed by the structure of the underlying signal. We explore this through the lens of a class of estimators that include ridge-regularized estimators, a method-of-moments estimator, and the ordinary least squares estimator for the under-parametrized regime. Our results provide a comparative study of these estimators and characterize regimes (as identified by the maximal signal component, number of large signals, total signal strength, and noise variance) where one estimator is preferable to the others. As a by-product, we obtain new double descent curves in \ell_\infty-error metric for the ridgeless regression, as well as ingredients for uniform confidence intervals for the signal coordinates. Our theoretical results are supported by extensive numerical simulations that confirm the predicted phase transitions and error limits.
Authors:
Haopeng Jin, Tiankun Yang, Zhenyu Guan, ShiQuan Dong, Wenlong Zhao, Hongzhu Yi, Chubin Chen, Tao Yu, Jinwen Luo, Yujia YangAbstract: Long video understanding with multimodal language models suffers from three compounding bottlenecks: heavy decode cost to obtain dense RGB frames, quadratic token growth with frame count, and weak motion perception under sparse keyframe sampling. Existing remedies either prune visual tokens after decoding, which leaves the expensive RGB pipeline untouched, or they discretise codec motion vectors with a generative tokeniser, which couples motion modelling to a heavy pre-training stage. In this paper we present HIMMEL, a hierarchical video-language framework that allocates semantic and motion capacity along separate paths. A small set of sparse anchor I-frames is routed to the expensive host ViT to ground object identity and scene layout, while the far denser inter-frame intervals are encoded by a lightweight compressed-domain tri-stream adapter that distils motion evidence from motion-vector maps, residual maps, and an I-frame context branch into aligned motion tokens. These tokens are injected into the LLM via a differentiable placeholder mechanism, after a dedicated Stage-1 contrastive alignment that places the motion representation in a geometry compatible with the frozen visual backbone. We further show that an InfoNCE alignment objective beats MSE regression by +1.5 pp because it preserves directional motion structure rather than collapsing onto the mean of visual deltas. On Video-MME, HIMMEL surpasses the dense 32-frame baseline by +2.3 pp (61.2 \to 63.5%) while using 3.6× fewer context tokens and running end-to-end in 1.03 s per question instead of 2.75 s. Switching to a stronger Qwen3-VL-8B host pushes HIMMEL to 64.9%, narrowing the gap to much larger proprietary systems while keeping inference on a single consumer GPU. Extensive ablations across stream composition, motion-encoder family, fusion mode, alignment objective, anchor count, LoRA rank, and video duration confirm that the full tri-stream is necessary and sufficient for the observed gains, and that the benefit grows with video length. We release training and evaluation pipelines to support reproducibility.
Abstract: Memory is becoming a core component of long-horizon AI agents, allowing agents to reuse past experience when operating web browsers, software tools, and other interactive environments. Existing work mostly treats memory as a supply problem, asking what experience to write, how to store it, and which entry to retrieve for the next task. Yet we still lack a clear account of how models consume retrieved memory across a multi-step action trajectory. This consumption process matters because it determines not only what memories should be retrieved, but also what models and control policies are needed to use them safely. To diagnose this process, we propose Entry--Propagation--Recovery (E-P-R), a trajectory-level framework that asks where memory first changes an action, whether that change carries forward, and whether the agent can recover after leaving a correct path. We instantiate E-P-R on WebArena and on MemTrapBench, a controlled benchmark we build to isolate these phases. We find that the main failure often begins at entry: agents adopt conflicting memory at the first exposed decision point even when the recommendation is task-wrong. Repeated exposure then amplifies this early error, while recovery after divergence is weak. Together, these effects create a compliance trap: across models, conflicting memory induces similar compliance rates, but once agents comply, their success rates collapse to a low floor. Stronger agents therefore suffer larger absolute damage because each compliance event erases more baseline capability. These results suggest that memory-augmented agents should be evaluated not only by retrieval quality or final success rate, but by how they consume memory throughout the trajectory.
Authors:
Dao S Minh, Trung-Kiet Huynh, Chi Nguyen Tran, Phu-Quy Nguyen-Lam, Phu-Hoa Pham, Tuan Nguyen, The Anh Han, Long Tran-ThanhAbstract: Large language models are increasingly deployed in decisions that require culture-dependent moral judgements, yet they answer as if the whole world thinks with a Western mindset. The Moral Machine experiment showed this is wrong at scale: 40 million judgments across 233 countries reveal that moral preferences are systematically structured by culture, and a model that ignores this variation does not merely underperform, but also imposes one society's intuitions on all others. Existing fixes do not scale to global deployment, as fine-tuning needs per-country preference data and GPU budgets, reward-guided decoding needs per-country reward models, and activation steering needs access to model internals that black-box APIs do not expose. In this work, we focus on this realistic inference-time regime, with no weight updates, no training data, and no internal access. The key observation is that within-country demographic disagreement, not consensus, is the steering signal. When culturally grounded personas agree, the base model is already calibrated. But when they disagree, the spread tells us what to fix and how. We propose DISCA (Disagreement-Informed Steering for Cultural Alignment), which instantiates each country as a panel of four World-Values-Survey-grounded persona agents, converts their disagreement into a bounded, loss-averse correction whose magnitude is set by the panel's variance, and shrinks the correction toward zero when the estimate is unreliable. Across 20 countries and 7 open-weight backbones (2B–70B) from five model families, DISCA reduces cultural misalignment on MultiTP by 10–24% on binary moral dilemmas and 2–7% on open-ended scenarios. Furthermore, a smaller 14B backbone with DISCA reaches lower absolute misalignment than a vanilla 70B model.
PaperID: 5402, Poster
Authors: Qi Lu, Haotian Xiong, Ziyu Gong, TIANJUN SHI, Lei Xie, Cheng-Zhong Xu, Li Li
Abstract: Dual-system visual-language-action models, which integrate high-level planning (System 2) with instant control (System 1), are promising for embodied AI, but their deployment is hindered by the high computational cost of processing continuous visual streams. Existing acceleration methods are one-sided, focusing only on optimizing the VLM (System 2), which not only limits speed gains but also risks compromising the deep reasoning capabilities it is meant to provide. In this paper, we introduce SynerVLA, a plug-and-play framework for accelerating dual-system VLA models. Its core idea is to leverage the distinct phase information in embodied execution, rapid approaching and fine-grained manipulation, to exploit both spatial and temporal redundancy in visual streams throughout the entire execution pipeline. SynerVLA accelerates dual-system VLA inference by first selecting key visual tokens via text-action fusion, then dynamically adjusting the token pruning-reuse ratio through phase-aware feedback, and finally accelerating both System-2 VLM and System-1 diffusion Transformer with a dual cache reuse mechanism. Evaluations on representative platforms demonstrate that SynerVLA delivers up to a 1.93× speedup and a 26% higher control frequency, with only a negligible impact on task success rate.
Abstract: Reinforcement Fine-Tuning (RFT) with verifiable rewards has emerged as a powerful learning paradigm for eliciting reasoning capabilities in Multi-modal Large Language Models (MLLMs). Recent studies also suggest that RFT (e.g., GRPO) is inherently more resilient to catastrophic forgetting than Supervised Fine-Tuning (SFT). However, whether RFT can effectively overcome forgetting in challenging visual continual learning settings, such as class-incremental learning (CIL) and domain-incremental learning (DIL), remains an open problem. Through a pilot study on rehearsal-free CIL, we confirm that while RFT consistently outperforms SFT, it still suffers from non-negligible forgetting. We empirically trace this bottleneck to Trajectory-level Drift Agnosticism: among candidate rollouts achieving identical task rewards, the KL divergence from the preceding-task policy varies substantially, which strongly correlates with catastrophic forgetting across sequential tasks. Motivated by this insight, we propose Retention-aware Policy Optimization (RaPO), a simple yet effective RFT framework that explicitly mitigates forgetting through trajectory-level reward shaping. Specifically, RaPO comprises two core components: (1) Retention Reward that converts trajectory-level distribution drift into a continuous reward signal, preferentially reinforcing knowledge-preserving rollouts within each group; (2) Cross-Task Advantage Normalization (CTAN), which maintains a persistent exponential moving average of reward statistics across task boundaries to stabilize the optimization progress during continual learning. Leveraging the free-form textual generalization of MLLMs, we comprehensively evaluate RaPO across five visual continual learning settings, spanning class- and domain-incremental image classification, class-incremental video classification, and class- and domain-incremental object detection. Extensive experiments demonstrate that RaPO achieves leading performance, substantially reducing catastrophic forgetting while preserving strong plasticity. To the best of our knowledge, this work represents the first systematic exploration of RFT in visual continual learning, offering insights that we hope will inspire future research.
PaperID: 5404, Poster
Authors:
Xinyuan Chen, Mingwen Shao, Qiao Zhang, Qinglin Zhan, Xiang Lv, Boyuan Yang, Jixuan Shi, Lingzhuang Meng, Chang Liu, Shanshan Song, Ling JianAbstract: Recent 3D indoor scene generation methods often produce layouts with limited spatial feasibility, where pathways are blocked and objects are difficult to access or use, restricting their applicability in virtual reality and embodied environments. To address this challenge, this paper presents a new paradigm for 3D indoor scene generation that explicitly models interaction feasibility through a carefully designed geometric abstraction, termed the Interaction-Oriented Human Proxy (IOHP). The IOHP serves as a compact yet spatially expressive representation for human interaction occupancy, and can be directly inferred by large language models without task-specific training. Building on IOHP, we formulate layout optimization as a unified objective that resolves geometric conflicts between the proxy and scene elements, while using the proxy as a bridge to preserve functional relationships among furniture, thereby improving overall usability and accessibility. Benefiting from this design, our method eliminates reliance on data-driven human preference learning and complex layout-planning pipelines, producing accessible, usable, and spatially coherent scenes in an efficient and stable manner. Extensive experiments demonstrate that our approach significantly improves layout quality and spatial plausibility, highlighting the role of human-centric interaction modeling in 3D scene generation.
Abstract: One prominent method of evaluating machine learning model trustworthiness is the notion of \emphcalibration. In the binary outcome setting, a probabilistic predictor is calibrated if outcomes are realized according to a model's distributional prediction, conditioned on this prediction. Straightforward extensions of binary calibration definitions to probabilistic multiclass classifiers suffer from an exponential complexity blowup as the space of predictions grows exponentially in the number of classes n. As a remedy, \citetnoarov_statistical_2023 propose multiclass calibration with predictions that are \emphproperties of the outcome distribution, reducing complexity from growing in the number of classes n to the \emphdimension d of the property, called its elicitation complexity. Previous work on approximate property calibration is generally limited to continuous scalar properties, despite many relevant properties of interest being discrete, like the mode or rankings. We characterize the approximate property calibration of discrete properties which are strongly orderable by using Lipschitz continuous properties as an intermediary. This work is the first to our knowledge to provide approximate calibration results for discrete properties. Along the way, we characterize the Lipschitz elicitation complexity of strongly orderable discrete properties by constructing algorithms for designing these Lipschitz properties, which we prove can be post-processed to obtain the original discrete property.
PaperID: 5406, Poster
Authors: Wei Dong, Terry Ji, Yan Min, Shahram Shirani, Jun Chen, Han Zhou
Abstract: Selective state-space models (SSMs) have recently shown strong potential for efficient image restoration. A central difficulty, however, is that a 2D feature map typically needs to be linearized into a 1D token sequence for causal state-space updates, making restoration performance highly sensitive to the imposed scan order. Existing deterministic scans follow space-filling principles but remain fixed and content-agnostic, and can yield fragmented or semantically inconsistent contexts under degradations. We propose -order design for causal vision SSMs. CoScan learns image-dependent space-filling traversals via differentiable proxy supervision driven by . Plugged into existing Mamba-based backbones, CoScan yields consistent improvements across several restoration benchmarks with minimal overhead and shows promising results on high-level vision tasks. Code will be released upon acceptance.
PaperID: 5407, Poster
Abstract: Robust 3D lane detection requires accurate metric road geometry recovery from visual inputs, yet conventional RGB-based approaches suffer from fragile lane evidence under adverse illumination, motion blur, and long-range perspective compression. Event cameras offer complementary high-temporal-resolution, high-dynamic-range structural cues that are robust to these challenges, but their sparse, motion-dependent responses cannot be directly fused into consistent 3D geometry. In this paper, we present the first attempt to introduce event cameras into 3D lane detection. To enable systematic research, we first build two multimodal benchmarks with metric 3D lane annotations: DSEC-3DLD (real-world sequences) and Ev-OpenLane (large-scale simulated sequences). We further propose RETR, a structure preserving transformer that converts complementary RGB and event observations into coherent 3D lane geometry. RETR first aligns modality consistent evidence through Reciprocal Context Flow Fusion, then preserves thin and uncertain lane structures with Uncertainty Aware Structural Consolidation, and finally decodes ordered lane hypotheses using a Geometry State Decoder with proposal conditioned initialization, reference conditioned query evolution, and diversified geometric inquiry. Extensive experiments show that RETR achieves state-of-the-art performance on both benchmarks, especially showing strong improvements in challenging lighting and complex road geometry scenarios.
Abstract: Diffusion models have achieved impressive performance in video generation, but their iterative denoising process remains computationally expensive due to the large number of tokens processed at each timestep. Recently, progressive resolution sampling has emerged as a promising acceleration approach by reducing latent resolution in early stages. However, scaling this idea to video generation remains challenging, as the additional temporal dimension introduces diverse spatio-temporal demands across different videos, and compressing only a single dimension often leads to limited acceleration or degraded quality. Therefore, we propose DVG, a Dynamic Video Generation framework that jointly allocates computation across time and space, automatically selecting content-aware acceleration strategies without manual tuning or retraining. DVG achieves near-lossless acceleration across models and tasks, reaching up to 7× speedup on HunyuanVideo and HunyuanVideo-1.5, and 18× when combined with distillation, demonstrating its potential as a key component in today’s large-scale efficient video generation systems. \emphOur code is in supplementary material and will be released on Github.
Abstract: Existing imitation learning methods enable robots to interact autonomously with the physical environment. However, contact-rich manipulation tasks remain a significant challenge due to complex contact dynamics that demand high-precision force feedback and control. Although recent efforts have attempted to integrate force/torque sensing into policies, how to build a simple yet effective framework that achieves robust generalization under multimodal observations remains an open question. In this paper, we propose , a force-aware reactive framework built upon flow matching. For contact-stage policy design, we investigate force signal fusion mechanisms and adopt an asymmetric multimodal fusion architecture that treats force as a global regulatory signal, combined with a joint prediction paradigm that enhances the policy's understanding of instantaneous force and historical information, thereby achieving deep coupling between force and motion. For task-level hierarchical decomposition, we divide manipulation into a vision-dominant approach stage (VLM-based pointing for target localization) and a touch-dominant interaction stage (force-driven contact execution), with a Vision-to-Force (V2F) handover mechanism that explicitly decouples spatial generalization from contact regulation. Experimental results across six real-world contact-rich tasks demonstrate that ForceFlow achieves a 37% success rate improvement over the strong baseline ForceVLA while maintaining significantly lower cost. Moreover, ForceFlow exhibits accurate force signal prediction and demonstrates superior performance in contact force self-regulation and zero-shot out-of-distribution (OOD) generalization.
PaperID: 5410, Poster
Abstract: Reinforcement learning (RL) post-training aligns diffusion-based generators with human preferences, yet existing RL methods suffer from poor compatibility with off-policy learning and few-step distilled models. These limitations are especially severe in the video generation area, as practical video generation pipelines often rely on few-step distilled generators. Furthermore, due to complex spatial-temporal dynamics and higher dimensions, near-on-policy video rollouts are both expensive to collect and often imperfect. Relying on such rollouts alone can amplify artifacts and is prone to reward hacking. To address these issues, we propose Forward-Consistent Reward Matching (FCRM), an efficient off-policy RL framework for video generation. FCRM converts the forward denoising loss into a positive loss-induced score and formulates the reward alignment as a one-step GFlowNet matching problem. The resulting residual is pointwise in a clean sample space that naturally supports off-policy learning and few-step generators. To avoid biased gradients, we introduce a double-sampling estimator for the squared residual objective. Theoretically, minimizing the proposed matching residual bounds the KL divergence between the learned distribution and the optimal reward-tilted distribution. Experiments on standard video generation benchmarks validate FCRM across online, replay, offline, and few-step settings and outperform SOTA methods.
PaperID: 5411, Poster
Abstract: Reconstructing the complete geometry of a scene from a single RGB image remains challenging—especially when inferring hidden structures where visual evidence is incomplete. We introduce VolFill, a generative framework that predicts the 3D structure of the complete scene rather than relying on traditional pixel-aligned regression. Our method utilizes a hybrid 3D VAE to compress sparse truncated unsigned distance function grids into a compact latent space, paired with a latent Diffusion Transformer that denoises this representation to recover the complete scene. We condition the generation on geometry foundation models, leveraging rich spatial priors for robust reasoning. Unlike existing methods limited by per-ray constraints or unstructured point-cloud queries, VolFill provides a structured representation that supports direct surface extraction and occupancy queries at scale. Extensive experiments on the 3D-FRONT and NRGBD datasets demonstrate that our approach significantly outperforms current baselines, providing a robust foundation for holistic spatial understanding.
PaperID: 5412, Poster
Abstract: Channel-adaptive Vision Transformers (ViTs) struggle to scale to non-RGB images with many uncorrelated or weakly correlated channels. Existing approaches roll out each channel into its own patch embedding, so the token sequence length grows linearly with the number of channels and self-attention cost grows quadratically. We introduce \emphChannelMixer, a multi-channel tokenizer that aggregates all channels within each spatial patch into a single channel-mixing token, decoupling the token sequence length from the number of channels. ChannelMixer is pretrained once with a lightweight autoencoder in a channel-agnostic manner and the resulting tokenizer can be easily integrated with ViTs trained under both supervised and self-supervised objectives. On multi-channel imaging benchmarks spanning 3 to 29 channels, ChannelMixer matches or outperforms channel-adaptive state-of-the-art methods while using approximately 20× fewer FLOPs and 29× fewer tokens than rolled-out tokenizers, measured on 29-channel images with a ViT-S/16.
PaperID: 5413, Poster
Abstract: Scene flow can capture low-level 3D motion displacements in dynamic scenarios. Early pairwise estimators relying on instantaneous two-frame motion lack long-term temporal correlation and also struggle with poor extrapolation ability in future prediction. Although some recent methods attempt to explore multi-frame scene flow estimation in a sequence-to-sequence manner, they typically suffer from heavy computational overhead with increasing input frames and long-horizon prediction degradation due to ineffective motion propagation. To address these problems, we propose a novel memory-enhanced sequential scene flow pipeline, called MESSENGER. To sufficiently mine long-term temporal dependencies naturally within consecutive sequences, a memory buffer is designed by explicitly storing multiple history flow estimates and latent states. For each input frame, the temporally stored flows and states are correlated and retrieved to predict the current initialized flow in a next-frame forecasting manner. Furthermore, we develop an uncertainty-aware reweighting module to filter unreliable retrievals and mitigate accumulated errors. Extensive experiments on nuscenes and Argoverse 2 demonstrate state-of-the-art performance of our MESSENGER, reducing EPE3D by 71.6% on nuScenes and 67.7% on Argoverse 2 in long-horizon future extrapolation. This superiority can be attributed to our designed autoregressive forecasting paradigm, which naturally forces the network to progressively learn the next-frame distribution based on history observations. Code will be released upon publication.
PaperID: 5414, Poster
Abstract: Chain-of-thought reasoning has driven striking advances in language model capability, yet every reasoning step grows the KV cache, creating a bottleneck to scaling this paradigm further. Current approaches manage these constraints on the model’s behalf using hand-designed criteria. A more scalable approach would let end-to-end learning subsume this design choice entirely, following a broader pattern in deep learning. After all, if a model can learn to reason, why can't it learn to forget? We introduce Neural Garbage Collection (NGC), in which a language model learns to forget while learning to reason, trained \emphend-to-end from outcome-based task reward alone. As the model reasons, it periodically pauses, decides which KV cache entries to evict, and continues to reason conditioned on the remaining cache. By treating tokens in a chain-of-thought and cache eviction decisions as discrete actions sampled from the language model, we can use reinforcement learning to jointly optimize how the model reasons and how it manages its own memory: what the model evicts shapes what it remembers, what it remembers shapes its reasoning, and the correctness of that reasoning determines its reward. Crucially, the model learns this behavior entirely from a \emphsingle learning signal — the outcome-based task reward — without supervised fine-tuning or proxy objectives. On Countdown, AMC, and AIME tasks, NGC maintains strong accuracy relative to the full-cache upper bound at a 2–3x compression in peak KV cache size and substantially outperforms eviction baselines. Our results are a first step towards a broader vision where end-to-end optimization drives both capability and efficiency in language models.
PaperID: 5415, Poster
Abstract: Vision-Language Models (VLMs) such as CLIP excel in global semantic alignment but often lack fine-grained perceptual capabilities. This hinders dense prediction tasks and bottlenecks the visual potential of Multimodal Large Language Models (MLLMs). Existing research has attempted to enhance CLIP's visual representations by incorporating geometric priors from vision-centric models. However, these strategies often struggle to achieve deep alignment for both local spatial structures and global semantics, potentially even distorting the original image-text space. To address these limitations, we propose SALM, an unsupervised embedding alignment framework based on structurally-aware latent mask modeling. SALM effectively synergizes local and global alignment via a dual-path design combining explicit and implicit mechanisms, without requiring any image-text pairs. First, we introduce a dual-matrix alignment strategy that explicitly calibrates intra-sample spatial correlations and activation intensities, thereby effectively injecting local geometric priors. Based on this, we further design a latent mask modeling mechanism to guide CLIP to restore the missing semantic details of the target model, thereby implicitly aggregating fine-grained structures into the global semantic space. Furthermore, driven by the empirical observations that CLIP's shallow features inherently possess strong spatial observational capabilities, we naturally extend SALM to a highly efficient self-distillation paradigm, SALM-Self. This unlocks CLIP's intrinsic fine-grained potential without relying on any external models. Extensive experiments demonstrate that SALM not only significantly improves performance in dense prediction tasks but also boosts CLIP's zero-shot accuracy, effectively enhancing the fine-grained understanding capabilities of MLLMs.
Abstract: Post-training attention sparsification reduces the quadratic cumulative attention cost of pretrained Transformers by selecting a small set of context units (tokens or blocks) for each query. Existing trainable methods usually rely on a lightweight selector to score context units followed by a hard Top-K selection, which blocks gradients from the language modeling loss. As a result, these methods commonly resort to distilling layer-wise dense attention distributions. While this approach encourages the selector to rank context units according to dense attention weights in the original model, such a ranking is not directly aligned with their impact on the model's final predictions under a fixed attention budget (ie, the number of attended context units per query), which can waste the limited attention budget on less useful units. To address this ranking misalignment, we propose Simple Attention Sparsification (SAS), a gated sparse attention mechanism that optimizes context ranking end-to-end with the language modeling loss. The key idea is to inject the selector's continuous scores into the attention logits during training, allowing the language modeling loss to update the selector through standard backpropagation. We identify several choices that are crucial to make this simple design work well in practice: placing the gate inside the attention \operatornamesoftmax in log form, keeping all context units active during training so each of them can receive gradients, and preserving continuous selector scores so the model learns relative priorities rather than only hard selections. Since a naive implementation that explicitly materializes the full attention matrix would cause prohibitive memory overhead for long sequences, we implement a memory-efficient Triton kernel that integrates the design into a FlashAttention-style computation. SAS consistently outperforms trainable sparse-attention across attention budgets, with especially large accuracy gains under tight budgets, demonstrating more effective context ranking for downstream tasks.
PaperID: 5417, Poster
Authors: Zian Guan, Guozheng Li, Zilun Zhang, Zecong Tang
Abstract: World models predict future states from observations and actions, enabling planning and decision-making in complex environments. Most existing approaches assume fixed environment dynamics and struggle to generalize to unseen regimes without retraining. However, in many real-world multi-agent systems, environment dynamics vary across scenarios due to diverse styles and evolving agent behaviors. Meanwhile, only a small number of demonstration transitions are available in new environments, making adaptation to such unseen dynamical regimes challenging. In this work, we introduce AdaWM, a world model that achieves zero-gradient few-shot adaptation to unseen dynamical regimes through Feature-wise Linear Modulation (FiLM). AdaWM is built on a JEPA-style latent prediction framework with a permutation-invariant context encoder whose output modulates the model through FiLM, enabling fast adaptation without modifying model parameters. Across heterogeneous multi-agent environments, AdaWM consistently reduces prediction error under distribution shift, outperforming fine-tuning and MAML in both adaptation gain and efficiency. Specifically, it achieves up to 15% reduction in prediction error with only K=5 demonstrations. Moreover, we reveal that scaling model capacity alone does not guarantee effective adaptation: fine-tuning a JEPA-based model with 3.9× more parameters degrades performance by 40%, highlighting the importance of explicit conditioning mechanisms over scale alone. These results show that AdaWM provides an effective approach for adapting world models to novel dynamical regimes.
Abstract: The theory of state tracking in recurrent architectures has predominantly focused on expressive capacity: whether a fixed architecture can theoretically realize a set of symbolic transition rules. We argue that equally important is error control, the dynamics governing hidden-state drift along the directions that distinguish symbolic states. We prove that affine recurrent networks, a class of models encompassing State-Space Models and Linear Attention, cannot correct errors along state-separating subspaces once they preserve state representations. Consequently, practical affine trackers do not learn robust state tracking; rather, they learn finite horizon solutions governed by accumulated state-relevant error. We characterize the mechanics of this failure, showing that tracking remains readable only while the accumulating within-class spread remains small relative to the initial between-class separation. We demonstrate empirically on group state-tracking tasks that this breakdown is predictable: tracking collapses when the distinguishability ratio crosses the readability threshold of the trained decoder. Across trained models, the point of this crossing predicts the horizon at which downstream accuracy fails. These results establish that robust state tracking is determined not only by an architecture's theoretical expressivity but crucially by its error control.
PaperID: 5419, Poster
Abstract: Neural circuits are densely interconnected systems, and many neural computations are thought to be mediated by the dynamical evolution of activity within these recurrent networks, a hypothesis often referred to as ‘computation through dynamics’. Under this view, understanding computation within real neural populations depends on building models of the dynamical rules governing the temporal evolution of recorded high-dimensional neural activity. Most commonly, such models depend on explicit generative parameterizations of (i) the (nonlinear) intrinsic dynamical flow field and stochasticity, (ii) the (nonlinear) mapping from dynamical state to neuronal activity, and (iii) the variability of individual neural responses given that dynamical state. Misspecification of any of these components can bias estimates of dynamics and latent trajectories. In this paper, we circumvent generation-related issues by introducing the Recognition-Parameterized Gaussian Process dynamics (RP-GPdyn) model. RP-GPdyn models the dynamical flow field using a nonparametric Gaussian Process-defined transition function and models the relationship of dynamics to neural activity implicitly using the Recognition-Parameterized Model (RPM) framework. We show that this generation-free approach uncovers meaningful, behaviorally relevant latent variables and dynamics from both synthetic and experimental datasets.
PaperID: 5420, Poster
Abstract: Large language models (LLMs) have demonstrated remarkable performance in video question answering (VQA). To improve responsiveness and answer quality, modern LLM-based VQA services pre-extract multiple LLM-compatible modalities from videos before response generation. However, the performance impact of modality combinations (MCs) and extraction knob settings has barely been studied. Our analysis shows that both factors substantially affect answer quality and Time-to-First-Token (TTFT) latency. Furthermore, the best configuration depends on both video content and query semantics, necessitating adaptive strategies. We propose Prysma, the first modality adaptation framework for LLM-based VQA. Prysma intelligently adapts the modality configurations to optimize answer quality, subject to a Service Level Objective (SLO) on average TTFT. It combines an offline modality orchestrator and an online semantic adapter. The offline orchestrator constructs a video knowledge base to warm-start tuning, employs a latency-aware configuration pruner to reduce search space, and runs a two-stage tuner to optimize both video- and query-level configurations. The online adapter selects configurations in real time based on query semantics and tuning history. Extensive evaluations demonstrate that Prysma improves answer quality by up to 82% over SOTA methods under the same SLO setting.
Abstract: We present the first algorithms for generalized linear contextual bandits under shuffle differential privacy and joint differential privacy. While prior work on private contextual bandits has been restricted to linear reward models---which admit closed-form estimators---generalized linear models (GLMs) pose fundamental new challenges: no closed-form estimator exists, requiring private convex optimization; privacy must be tracked across multiple evolving design matrices; and optimization error must be explicitly incorporated into regret analysis. We address these challenges under two privacy models and context settings. For stochastic contexts, we design a shuffle-DP algorithm with regret \tildeO\!\big(d^3/2\sqrtT \log T + d^5/4 \sqrtT/\varepsilon (\log T)^3/4\big) in the dominant term, matching the non-private rate \tildeO(d \sqrtT \log T) in the leading-in-T term up to a multiplicative factor of \sqrt d and a privacy correction term of \tildeO(d^5/4\sqrtT \log T/\varepsilon). For adversarial contexts, we provide a joint-DP algorithm with regret \tildeO\!\big(d\sqrtT\log T + d^3/4\sqrtT/\varepsilon\,(\log T)\,(d+\log T)^1/4\big) matching the non-private rate \tildeO(d\sqrtT\log T) in the leading term. Unlike prior work on locally private GLM bandits, our methods require no spectral assumptions on the context distribution beyond \ell_2 boundedness.
Abstract: Inferring continuous system evolution from sparse temporal snapshots is a key challenge in generative modeling and single-cell omics. While Optimal Transport (OT) is popular, existing frameworks are largely restricted to first-order dynamics, assuming memoryless velocity fields. This limits expressiveness, as first-order systems fail to account for regulatory momentum and time-delayed responses inherent in processes like cell differentiation. Here, we introduce TracingFlow, a simulation-free Flow Matching framework generalizing to second-order dynamics. By using neural networks to regress the acceleration field, TracingFlow provides an exact, efficient solution to the Dynamical Optimal Acceleration Transport (DOAT) problem. Unlike first-order methods yielding over-smoothed trajectories, our second-order formulation captures high-curvature transitions and nonlinear evolutions by learning the underlying force fields. Evaluated on complex synthetic and large-scale scRNA-seq datasets, TracingFlow achieves superior accuracy in distributional reconstruction and trajectory faithfulness. Moreover, by integrating lineage tracing priors, it recovers dynamical structures that are both mathematically optimal and biologically plausible.
PaperID: 5423, Poster
Abstract: Pretrained model weights are increasingly released under commercial licenses, usage restrictions, or other conditions that prohibit unauthorized fine-tuning. In practice, however, such misuse is difficult to detect or verify after the fact. This motivates a stronger objective for protected weight release: weights should remain useful for intended inference, yet become practically unattractive to repurpose through unauthorized gradient-based adaptation. We frame this objective as a \emphscorched-earth strategy in parameter space: rather than only proving infringement after misuse occurs, the released weights themselves should react destructively when unauthorized fine-tuning begins. To realize this idea, we propose Gradient-Mine Units (GMUs), a data-free weight-space protection mechanism for pretrained networks. GMUs are planted into selected feedforward layers as hidden units with extreme internal scale, while a locking mechanism keeps them silent at initialization so that the original inference behavior is preserved. During fine-tuning, this locked state is progressively broken, allowing the planted units to emit amplified gradients that disrupt the model's native adaptation dynamics. We formulate GMUs in a general feedforward setting and show that gated instantiations naturally provide an additional hard lock. Empirically, we validate the method on both large language models and Vision Transformers. Across multiple architectures and downstream tasks, GMUs preserve pre-fine-tuning utility while substantially degrading or destabilizing standard fine-tuning. These results suggest that protected weight release can move beyond post hoc attribution toward a practical deterrence mechanism for unauthorized adaptation. Our implementation is available at here.
PaperID: 5424, Poster
Abstract: Incomplete Multi-view Multi-label Learning (IMvMIL) refers to a classification task where missing views and incomplete label assignments coexist. Existing methods often directly fuse heterogeneous view features, neglecting the fact of imbalance where dominant views tend to overshadow non-dominant ones, resulting in unsatisfactory performance. To address this, we propose the View Confidence Perception-Driven Incremental Prediction (VCPIP) framework, which incorporates an adaptive structural refinement strategy to balance different views via confidence-based branch expansion, thereby enabling the full exploitation of view-specific information. Specifically, we propose a View-Quality Aware (VQA) strategy, which introduces a novel metric to evaluate the predictive strength of each view-specific branch using supervision signals from the label space. By leveraging these quality assessments, VQA dynamically allocates optimal classifier ensembles and employs residual optimisation to improve the discriminative capability of non-dominant views. Additionally, to eliminate information redundancy and purify features, we propose a Mixture-of-Experts-based Decouple-Fusion (MOEDF) mechanism to extract refined representations by disentangling consistent and specific features under orthogonality constraints for more precise multi-label prediction. Moreover, to simulate real-world data corruption, we introduce the Stochastic Cross-sample Fragment Swapping (SCFS) strategy, which interchanges feature fragments across samples to facilitate the modelling of robust global representations for enhanced generalisation. Extensive experiments across diverse benchmarks and varying missing-rate scenarios confirm that our method consistently surpasses state-of-the-art methods in multi-label classification.
PaperID: 5425, Poster
Authors:
Yuran Bian, Xiaoyan Wang, Conghui Zheng, Xiaohan Zhang, Li PanAbstract: Large vision-language models (LVLMs) ground language generation on visual content through the text-to-image slice of self-attention, which acts as a prompt-conditioned router for selecting visual evidence. To induce hallucinations in LVLMs, existing adversarial attacks largely operate at the two ends of this pipeline, either perturbing the front-end visual representation or optimizing against output-token objectives. This leaves the intermediate evidence-selection step as an underexplored yet low-cost attack surface, since it does not require heavy semantic manipulation of image features or extensive output-token optimization. To exploit this attack surface, Text-to-Visual Attention Disruption (T2V-AttnDisrupt) is proposed as an untargeted adversarial attack that induces hallucination by optimizing the input image to distort the text-to-image attention distribution relative to its clean reference, without relying on output-tokens. Experiments on four open-source LVLMs show that T2V-AttnDisrupt increases hallucination rates on captioning benchmarks and reduces VQA accuracy, while preserving overall response quality. It transfers across surrogate–target pairs and generalizes from a single captioning prompt to unseen VQA questions. Moreover, it remains effective against representative defenses, including encoder-level robustness, alignment-based fine-tuning, decoding-time hallucination mitigation, and attention-level interventions, indicating that text-to-image attention is still insufficiently protected. Mechanism analyses show that restoring the clean attention pattern largely recovers visual faithfulness, even when the adversarial image is kept fixed, indicating that hallucination can be driven by corrupted evidence routing rather than feature-level corruption. This low-cost and largely unprotected routing step calls for defenses that explicitly safeguard prompt-conditioned text-to-image attention.
Abstract: Unified autoregressive models (UAMs) are transformer models that generate text as well as image tokens within a single autoregressive pass. Shared parameters and a multimodal vocabulary simplify the training pipeline and facilitate flexible multimodal generation, yet might introduce new vulnerabilities. In particular, we are the first to show that this unified architecture enables multimodal backdoor attacks, where a trigger can propagate malicious effects across multiple output modalities. Specifically, we present the Token by Token Backdoor Attack (ToBAC), the first backdoor attack targeting UAMs, exploring both data-based and model-based poisoning strategies. We demonstrate that innocuous characters or even common words can be transformed into triggers that elicit harmful behavior in autoregressive image generation. ToBAC can jointly manipulate visual outputs and accompanying text, increasing the perceived authenticity of fabricated content. With model access, ToBAC enables attacks on the unified Liquid model in which a subtle word (e.g., ``cool'') induces modality-aligned brand promotion or ideological influence in 55% of generations. Without model access, ToBAC can be induced through data poisoning, achieving an average success rate of 63.1% against JanusPro.
Abstract: This paper studies test-time aggregation, an approach that generates multiple reasoning traces and aggregates them into a final answer. Most existing methods rely on evaluation signals collected from candidate traces in isolation or answer frequencies, while ignoring comparative interactions among candidates. We propose (JC), formulated as a constrained Ising-type energy minimization problem, where independent evaluation signals act as external fields and pairwise comparisons act as interactions. JC provides a unified framework for test-time aggregation that subsumes existing voting and weighted aggregation methods as special cases. Our construction of the interaction matrix leverages LLM-as-a-judge comparisons, and admits a theoretical interpretation under answer-level homogeneity assumptions. Moreover, we develop an efficient approximation strategy that makes interaction modeling practical for large-scale test-time aggregation. Experiments on math and code reasoning benchmarks show that JC consistently outperforms existing baselines across tasks, judge models, trace budgets, and trace-generation settings.
PaperID: 5428, Poster
Abstract: Neural decoders have shown strong potential for error correction in short- and moderate-length regimes, yet a fundamental tension remains between decoding accuracy and inference latency. Recent attempts to leverage diffusion probabilistic models for channel decoding typically adopt a fully generative paradigm, initializing the reverse process from an observation-agnostic prior, which overlooks a key structural property of channel decoding: the received signal already contains substantial information about the target codeword. In this work, we propose Channel Residual Diffusion Model (ChRes-DM), a principled diffusion-based decoding framework that reinterprets channel decoding as a directed restoration process anchored at the noisy observation. Instead of sampling from a generic prior, ChRes-DM constructs a geometric residual bridge between the received signal and the clean codeword, explicitly modeling the conditional transport induced by the channel and leading to a deterministic Probability Flow Ordinary Differential Equation (PF-ODE) that governs the decoding dynamics. By eliminating redundant stochastic sampling inherent in conventional diffusion models, ChRes-DM enables efficient iterative decoding with flexible step-skipping, offering fine-grained control over the accuracy–latency trade-off. Extensive experiments across representative channel coding benchmarks demonstrate that ChRes-DM achieves competitive or superior decoding performance compared to existing neural decoders while significantly reducing inference iterations, highlighting diffusion-based residual transport as a promising and scalable paradigm for neural channel decoding.
Authors: xinpeng, Ang Gao
Abstract: The scalability of continuous normalizing flows (CNFs) for unbiased Boltzmann sampling remains limited in high-dimensional systems due to the cost of Jacobian-determinant evaluation, which requires D backpropagation passes through the flow layers. Existing stochastic Jacobian estimators such as the Hutchinson trace estimator reduce computation but introduce bias, while the recently proposed Flow Perturbation method is unbiased yet suffers from high variance. We present Flow Perturbation++, a variance-reduced extension of Flow Perturbation that discretizes the probability-flow ODE and performs unbiased stepwise Jacobian estimation at each integration step. This multi-step construction retains the unbiasedness of Flow Perturbation while achieves substantially lower estimator variance. Integrated into a Sequential Monte Carlo framework, Flow Perturbation++ achieves significantly improved equilibrium sampling on a 1000D Gaussian Mixture Model and the all-atom Chignolin protein compared with Hutchinson-based and single-step Flow Perturbation baselines.
PaperID: 5430, Poster
Authors: Ankitkumar Joshi, Milos Hauskrecht
Abstract: Irregularly sampled multivariate time series arise in many real-world domains such as healthcare, climate science, and large-scale monitoring systems. Characterized by asynchronous observations and heterogeneity across variables, such time series pose significant challenges for forecasting. In this work, we introduce Cross-Variable Temporal Attention (CVTA), a short-term, continuous-time forecasting framework that learns prediction functions from a compact, interpretable Markov state summarizing recent temporal dynamics across variables. CVTA uses Variable-Conditioned Embeddings to project heterogeneous variables into variable-specific latent spaces, and applies temporal attention to capture cross-variable dependencies from this state. Furthermore, we demonstrate that CVTA can be integrated as an auxiliary module into existing state-of-the-art prediction models through a dynamic meta-decision model. Extensive experiments on multiple real-world benchmarks show that CVTA achieves strong short-term predictive accuracy, supports efficient online inference, improves existing prediction models when used as an auxiliary module, and learns attention patterns that align well with domain knowledge.
PaperID: 5431, Poster
Authors:
wenjie mu, ziniu liu, Tong Wu, Zhan Li, Chuanzhou su, XUANYI SHEN, Fan Lu, Junqiao Zhao, Tiantian Feng, Chen Ye, Guang ChenAbstract: Extending feed-forward 3D Gaussian Splatting (3DGS) to large-scale scenes remains challenging. Existing global reconstruction methods incur prohibitive computational overhead on long sequences; while streaming-based online methods are more efficient, their inherent forgetting behavior and unidirectional dependence still limit reconstruction quality. To address this, we propose PF-SGS, a pose-free streaming feed-forward 3DGS for large-scale scene reconstruction. PF-SGS leverages a persistent hidden state as scene memory to integrate streaming inputs frame-by-frame, progressively recovering scene geometry and camera poses in a self-supervised manner. At its core lie two targeted designs: Adaptive State Gating, which suppresses long-term memory degradation by dynamically regulating update magnitudes to ensure stable state evolution; and Delayed Gaussian Modeling, which introduces bounded look-ahead observations to compensate for insufficient local geometric constraints under streaming inputs, thereby improving the geometric consistency. Experiments demonstrate that PF-SGS processes sequences of varying lengths with linear time complexity, achieving quality comparable to pose-prior-based large-scale reconstruction baselines.
Authors:
Jian Tan, Fan Bu, Yuqing Gao, Devavrat Khanolkar, Jason Mackay, Boris Sobolev, Lei Jin, Li ZhangAbstract: Machine data is central to observability and diagnosis in modern computing systems, appearing in logs, metrics, telemetry traces, and configuration snapshots. When provided to large language models (LLMs), this data typically arrives as a mixture of natural language and structured payloads such as JSON or Python/AST literals. Yet LLMs remain brittle on such inputs, particularly when they are long, deeply nested, and dominated by repetitive structure. We present HYVE (HYbrid ViEw), a framework for LLM context engineering for inputs containing large machine-data payloads, inspired by database management principles. HYVE surrounds model invocation with coordinated preprocessing and postprocessing, centered on a request-scoped datastore augmented with schema information. During preprocessing, HYVE detects repetitive structure in raw inputs, materializes it in the datastore, transforms it into hybrid columnar and row-oriented views, and selectively exposes only the most relevant representation to the LLM. During postprocessing, HYVE either returns the model output directly, queries the datastore to recover omitted information, or performs a bounded additional LLM call for SQL-augmented semantic synthesis. We evaluate HYVE on diverse real-world workloads spanning knowledge QA, chart generation, anomaly detection, and multi-step network troubleshooting. Across these benchmarks, HYVE reduces token usage by 50--90% while maintaining or improving output quality. On structured generation tasks, it improves chart-generation accuracy by up to 132% and reduces latency by up to 83%. Overall, HYVE offers a practical approximation to an effectively unbounded context window for prompts dominated by large machine-data payloads.
PaperID: 5433, Poster
Abstract: Under matched conditions, classifier-head fine-tuning (\textsccls) paradigm consistently outperforms the LLM-as-Judge sequence-to-sequence (\textscs2s) on binary hallucination detection. We trace this gap to the representations induced by the two objectives. Static analysis shows that \textsccls learns a task-specific discriminative direction, pushing hidden states into a clear two-cluster geometry with stronger class separation and linearly decodable label information. By contrast, \textscs2s relies on the pretrained language-modeling head to make decisions. This minimizes the pressure to reorganize representations, leaving the hidden states much closer to the pretrained backbone. Our analysis also reveals that the \textsccls signal saturates before the final layer, after which upper layers form a plateau of statistically interchangeable and largely redundant representations. However, endpoint redundancy alone does not show whether these layers are dispensable during training. Further training-dynamics analysis reveals 1) a lower pretrained substrate, 2) a narrow upper-middle InfoWindow where task information emerges, and 3) a later-synchronizing upper plateau that mirrors this signal. This mechanism motivates us to remove layers above the InfoWindow. By truncating roughly 31% of the network, we obtain early-exit models that match or exceed full-model performance on Qwen3-4B and 8B. The same depth-fractional prescription transfers to Llama-3.2-3B and Mistral-7B-v0.3. We hope this helps guide efficient task-specific LLM design.
PaperID: 5434, Poster
Abstract: Multimodal Reinforcement learning with verifiable rewards (RLVR) algorithms assign the same advantage to every token in a generated sequence, providing no differential signal for tokens with varying degrees of dependence on visual evidence. A simple experiment confirms this issue: scaling token advantages by uninformative random noise already outperforms uniform credit, suggesting that uniform credit is empirically suboptimal. Beyond that, we observe RLVR tuning under uniform credit weakens visual grounding along three complementary axes: attention, gradient attribution, and functional dependence on image embeddings, collectively indicating insufficient utilization of visual evidence. To understand this degradation, we provide a theoretical analysis showing that, under language-prior dominance and regularity conditions, uniform credit can amplify a pretrained text-favored gradient imbalance, contributing to weaker visual attention. In response, we propose Hierarchical Context Credit Reassignment (HiCCR), which reassigns credit at two levels: a token-level weight amplifies the advantage for visually grounded tokens, and a trajectory-level weight upweights rollouts with stronger visual engagement. The entire mechanism adds less than 3% wall-clock overhead and applies to mainstream RLVR algorithms without auxiliary models or additional data. On ten benchmarks spanning mathematical and general multimodal reasoning, HiCCR consistently improves upon its corresponding base algorithm and achieves state-of-the-art results among open-source models of comparable scale, while alleviating the weakening of visual grounding observed under RLVR training.
Authors: Haoyang Su, Ying Wen
Abstract: Command line interface (CLI) agents are emerging as a practical paradigm for agent-computer interaction over evolving filesystems, executable command line programs, and online execution feedback. Recent work has used reinforcement learning (RL) to learn these interaction abilities from verifiable task feedback, yet few methods exploit the native structured attributes of CLI actions as learning signals. Beyond this underused action structure, CLI learning also couples two bottlenecks for coding agents. First, the agent must identify task-relevant evidence in a large codebase from partial observations. Second, sparse terminal rewards must be assigned to the actions that shape a long multi-turn trajectory. We study these bottlenecks through shell-driven information extraction and file editing tasks. For selective observation, we introduce \sigma-Reveal, an inference-time mechanism that selects token-budgeted context for the same CLI. For credit assignment, we propose Action Advantage Assignment (\mathrmA^3), a native agentic RL method that preserves the algorithmic complexity of standard agentic RL. \mathrmA^3 constructs turn-level advantages from episode-level relative feedback, abstract syntax tree (AST) based action sub-chain residuals, and tree-level trajectory margins. To further evaluate this problem setting, we construct ShellOps, a verifiable dataset suite covering CLI tasks in repository environments.
PaperID: 5436, Poster
Abstract: Time-aware retrieval is crucial for long-horizon LLM agents to effectively leverage ever-growing long-term memory. Existing methods focus on lexical or semantic similarity, leading to target memories being crowded out by temporally mismatched yet semantically similar candidates. Even time-aware approaches inject temporal cues only during memory organization or answer generation, overlooking temporal constraints at the retrieval stage, where candidate selection occurs—the root cause of the displacement. Inspired by the temporal contiguity effect in human episodic memory, we propose TIDER (Temporal Interval-driven Dual-channel Evidence Retrieval), a time-aware agentic retrieval framework. Central to TIDER is a reinforcement learning (RL)-trained temporal interval locator that infers the temporal windows in which the evidence lies. RL optimizes window localization as a discrete, non-differentiable retrieval decision, using composite rewards to convert sparse answer supervision into fine-grained retrieval feedback. These temporal windows first define the temporal retrieval channel; the retrieved candidates are then merged with results from the semantic channel and reranked by proximity to the windows. This temporal grounding enables TIDER to retrieve temporally aligned evidence that would otherwise be drowned out by the volume of semantically similar yet temporally mismatched memories. With a lightweight 1.7B locator, TIDER achieves state-of-the-art accuracy on both long-horizon memory benchmarks—61.52% on LifeBench (+8.82%) and 78.96% on LoCoMo (+3.66%), which indicates that temporal cues are crucial for pinpointing target evidence amid a vast pool of semantically similar historical events in long-term daily-life memory.
Abstract: Neuro-Symbolic Concept-based Models (NeSy-CBMs) are a family of architectures that integrate neural networks with symbolic reasoning for enhanced reliability in high-stakes applications. They work by first extracting high-level concepts from the input and then inferring a task label from these compatibly with given logical constraints. Yet, their label and concept predictions can be overconfident, making it difficult for stakeholders to gauge when the model's decisions can be trusted. We address this issue by integrating ideas from Conformal Prediction (CP), a framework providing rigorous, distribution-free coverage guarantees. We formalize three desiderata -- consistency, coverage, and conciseness -- that any conformal method for NeSy-CBMs should satisfy, and show that existing approaches fall short of at least one. We then introduce COCOCO, a post-hoc framework that conformalizes concepts and labels jointly and reconciles them via a single deduction–abduction revision step. COCOCO satisfies all three desiderata, retains distribution-free coverage, is robust to imperfect knowledge and supports user-specified size budgets. Our experiments on 8 data sets highlight how COCOCO compares favorably against competitors and natural baselines in terms of performance and set size.
PaperID: 5438, Poster
Abstract: High-order derivative kernels in machine learning interatomic potentials (MLIPs) are a dominant cost in scientific ML, yet current autonomous GPU-kernel agents do not reliably target this workload. They generate kernels against local tensor-level references, whereas MLIP derivative workloads span forward, backward, and higher-order automatic differentiation (AD) phases with cross-phase saved tensors and reconstruction rules. We introduce \textscVerigrad, a verification-driven multi-agent harness that makes high-order derivative generation a first-class kernel-generation workload. \textscVerigrad first exposes derivative semantics through \textttDerivativeTask and a derivative-aware IR, and then decomposes the workload into verifiable kernel contracts. An execution-based derivative verifier validates the generated artifacts; the same verifier then gates hardware-guided refinement, so optimized kernels replace earlier candidates only when the derivative checks continue to pass. We integrate the generated kernels into MatRIS training and inference, where \textscVerigrad attains 1.33--1.55× end-to-end speedup over PyTorch eager across energy-only, energy--force, and full-derivative inference and training workloads.
Abstract: Flow matching has recently emerged as a flexible and efficient framework for generative modelling by learning deterministic transport dynamics between probability measures. In this work, we extend flow matching to the space of probability measures over probability measures, introducing a Wasserstein-on-Wasserstein (WoW) formulation. Leveraging the nested Wasserstein geometry, we show that measures over transport plans naturally induce velocity fields that realize metameasure flows. This yields a principled generalization of Wasserstein flow matching via coupled outer and inner transport plans. To address the substantial computational cost of WoW transport, we propose scalable approximations based on sliced and linear Wasserstein distances, enabling efficient training while promoting numerically stable, near-straight trajectories. Our framework unifies and extends existing approaches to point cloud and set generation, providing a practical and theoretically grounded method for generative modelling in WoW spaces.
PaperID: 5440, Poster
Abstract: Interactive language models should respond to user feedback without blindly conforming to incorrect pressure. Existing evaluations of sycophancy typically measure whether models flip their answers under pressure and treat apparent recovery after pressure removal as robustness. We show that this recovery is often superficial. We introduce the , a multi-turn evaluation framework that disentangles pressure-induced flipping, surface recovery, and relapse under renewed weak pressure. Across factual QA benchmarks and instruction-tuned models, PRR reveals three findings: (1) models can appear to recover at the output level while remaining vulnerable to renewed pressure; (2) surface answer correction can be dissociated from repair of the latent answer state; and (3) different internal components support distinct recovery pathways. Through causal interventions, we show that output-level steering can enforce correct answers without repairing corrupted latent readouts, whereas restoring clean hidden states repairs both surface behavior and latent states. Further component-level restoration shows that residual-stream states enable full repair, while mid-layer attention outputs provide partial latent repair. These results highlight that robust multi-turn reliability requires evaluating recovery stability and latent state repair, not merely apparent answer correctness.
PaperID: 5441, Poster
Abstract: We propose LEGO, a budget-aware Dense-to-MoE methodology that converts pretrained vision-language models into efficient FFN-MoE variants under a target architectural budget ((P_total,P_active)). Our study is motivated by a counter-intuitive recovery pattern: fixed activation ratios that work well for some dense backbones can degrade sharply on others under the same recovery recipe, indicating that raw sparsity alone is an unreliable design rule. Through controlled recovery experiments across VLM families and scales, we identify the activated-FFN-to-hidden ratio (D_I/D_H) as a backbone-comparable structural factor and summarize empirical sizing rules for avoiding severe bottlenecks while limiting diminishing returns. LEGO operationalizes these rules through a discrete budget-aware search over Split, Upcycle, and Hybrid constructions, and introduces a moment-matching scaling factor to reduce the initialization scale shift caused by sparse FFN aggregation. Finally, we use a two-stage multimodal recovery recipe that first stabilizes the sparse LLM with the vision tower frozen, then performs joint fine-tuning to recover the full VLM. Under matched data and training protocols, LEGO improves Dense-to-MoE recovery quality and turns costly blind architecture search into a small set of rule-guided candidates. Codes can be found in supplementary materials.
PaperID: 5442, Poster
Abstract: Illusory pattern perception is a well-documented human cognitive tendency to infer meaningful relationships in data that is actually random. Such a tendency, often described as “connecting the dots” where none exist, can result in systematic reasoning errors. This paper investigates whether Large Language Models (LLMs) exhibit such perceptual tendencies, which can lead to systematic errors in downstream applications. To our knowledge, this work presents the first systematic study of illusory pattern perception in LLMs, adapting classic psychological paradigms to three tasks with direct empirical comparison to human behaviors. We find that LLMs frequently exhibit stronger illusory pattern perception than humans. In particular, models tend to over-associate frequent positive attributes with majority groups or large organizations, and show increased tendencies to construct causal narratives from ambiguous events. To uncover the mechanism behind these behaviors, we develop a feature interpretability framework based on Sparse Autoencoders (SAEs) to analyze internal representations. Our results reveal that holistic frequency perception and analytic cognitive orientation are linked to the emergence of illusory perceptions. These findings highlight a previously underexplored cognitive-like illusion that may affect the reliability of LLM reasoning. Code available at
PaperID: 5443, Poster
Authors: Shervin Ghasemlou
Abstract: Larger language models solve harder problems, but do they also produce proportionally better explanations? We study in-context explanation transfer across 40 open-weight models (360M--72B) and 5 receivers (360M--7B), using problems drawn from 13 benchmark sources and evaluated with domain-specific metrics. Our central finding is an \empharticulability ceiling: under equal token budgets, mid-range models are better teachers than frontier models. Transfer peaks at 170--220 words and then drops sharply beyond 300 words, falling below even the shortest explanations. Large models partly compensate through verbosity: longer explanations yield higher total transfer on average, but each additional word carries less pedagogical value, so a 360M model delivers more transfer value per word than a 72B model. We also find that RLHF improves pedagogical clarity: instruct models score better on LLM-judge evaluation than base models despite lower lexical overlap, and teacher correctness dominates all other factors. Code and logs are released with the paper.
PaperID: 5444, Poster
Abstract: Frontier LLMs are increasingly capable of expert-level scientific reasoning, but using them as reliable scientific agents requires more than general reasoning ability. Scientific simulation is a central test case: modern simulators are configured through executable interfaces such as XML decks, input scripts, and namelists that function as domain-specific languages tied to the simulator's internal API. Translating a researcher's natural-language intent into a runnable configuration remains a recurring expert bottleneck. We study how much of this bottleneck a general-purpose LLM coding harness, Claude Code, can absorb when wrapped with a : a package of skills, tools, and workflow control flow that grounds the agent's outputs in simulator documentation, schema, and example libraries. We instantiate SIGA for GEOS, an open-source multiphysics simulator used in CO₂-storage and induced-seismicity research. A Resolution-IV factorial surfaces three benefits over vanilla Claude Code. First, SIGA improves , reducing across-seed variance by roughly 40× by preventing unparseable or empty decks on a hard tail of compound multi-physics tasks. Second, it improves , raising mean structural similarity by about +7 percentage points on the same hard tail. Third, a self-evolved variant matches the best hand-designed cell with roughly 16% fewer tool calls, suggesting that automatic SIGA discovery is tractable. A preliminary human baseline finds that geoscience-domain-expert volunteers new to GEOS take between 8 and 36 times as long as the agent on a representative task. An explicit human-consultation tool is used in only about 3% of under-specified trials, with the agent instead relying on the on-disk example library as a cheaper retrieval substitute. A small OpenFOAM transfer study indicates that the recipe is not specific to GEOS XML: the same stop-hook component again dominates reliability, and the best SIGA cell outperforms both vanilla Claude Code and a constrained Foam-Agent lint-only baseline. We close with SIGA-design recommendations grounded in failure modes that our best configuration still does not fix.
Abstract: This work investigates multi-objective imitation learning: the problem of recovering policies that lie on the Pareto front given demonstrations from multiple Pareto-optimal experts in a Multi-Objective Markov Decision Process (MOMDP). Standard imitation approaches are ill-equipped for this regime, as naively aggregating conflicting expert trajectories can result in dominated policies. To address this, we introduce Multi-Output Augmented Behavioral Cloning (MA-BC), an algorithm that systematically partitions divergent expert data while pooling state-action pairs where no behavior conflict is observed. Theoretically, we prove that MA-BC converges to Pareto-optimal policies at a faster statistical rate than any learner that considers each expert dataset independently. Furthermore, we establish a novel lower bound for multi-objective imitation learning, demonstrating that MA-BC is minimax optimal. Finally, we empirically validate our algorithm across diverse discrete environments and, guided by our theoretical insights, extend and evaluate MA-BC on a continuous Linear Quadratic Regulator (LQR) control task.
PaperID: 5446, Poster
Abstract: Explaining the generalization and training dynamics in neural networks remains a challenge, and various approaches have been developed to study different aspects of these phenomena. In this paper, we introduce the to this study -- a quantity based on the theory of metric magnitude -- that reflects how well an arbitrary point is represented by a given set. We find that this basic quantity can be applied to examine various features in neural generalization. The ratio between the magnitude potential with respect to a class and with respect to the entire data, computed at the logit layer, is informative of the representation of the point. In experiments, these ratios for individual training points are found to be correlated with the Feldman memorization scores. Magnitude potential ratios aggregated across points detect structural changes in the decision boundaries and provide a geometric indicator of grokking in modular arithmetic. Although the magnitude potential ratio and neural collapse are both closely associated with intra-class and inter-class geometric structure, the magnitude potential ratio remains informative even when neural collapse is explicitly suppressed.
PaperID: 5447, Poster
Authors:
Yunhak Oh, Yoonho Lee, Junseok Lee, Namkyeong Lee, Sang-Yeon Hwang, Yinhua Piao, Hyomin Kim, SEONGHWAN KIM, Jaechang Lim, Woo Youn Kim, Sungsoo Ahn, Chanyoung ParkAbstract: Single-cell RNA-seq representation learning is inherently label-free: cell identities, states, and tissue contexts are typically discovered through analysis rather than provided as target labels during training. As a result, the evaluated cell embedding must balance competing demands, which we refer to as the : it should preserve biological identity, avoid encoding collection effects as cell identity, and retain expression information needed for reconstruction. We introduce , a latent-bottleneck VAE that tests whether simple expression-derived routing can manage this trilemma without cell-level annotations, metadata-derived supervision, auxiliary representation losses, or specialized disentanglement modules. combines expression-gated gene encoding, cell-representation routing through the decoder, and pseudo-bulk prior conditioning under a single reconstruction objective. In release-based zero-shot evaluation on successive CZ CELLxGENE Census releases, improves batch-effect removal while preserving biological identity, retains marker-program and biologically meaningful tissue-context structure in the evaluated embedding, and better preserves differential-expression and pathway structure through reconstruction; ablations support burden allocation as the source of these gains.
Authors: Hung-Tien Huang, Dzung Dinh, Junier Oliva
Abstract: Active feature acquisition (AFA) is an instance-adaptive paradigm in which, at inference time, a policy sequentially chooses which features to acquire (at a cost) before predicting. Existing approaches either train reinforcement learning policies, which deal with a difficult MDP, or greedy policies that cannot account for the joint informativeness of features or require knowledge about the underlying data distribution. To overcome this, we propose Template-based AFA (TAFA), a non-greedy framework that learns a small library of feature templates---sets of features that are jointly informative---and uses this library of templates to guide the next feature acquisitions. Through identifying feature templates, the proposed framework not only significantly reduces the action space considered by the policy but also alleviates the need to estimate the underlying data distribution. Extensive experiments on synthetic and real-world datasets show that TAFA outperforms the existing state-of-the-art baselines while achieving lower overall acquisition cost and computation.
Abstract: We study two models of \mathsfReLU neural networks: monotone networks (\mathsfReLU^+) and input convex neural networks (\mathsfICNN). Our focus is on expressivity, mostly in terms of depth and our motivation is to gain better understanding of depth requirement needed to exactly represent functions in terms of neural networks with ReLU activations: a subject that has received a lot of attention lately. We prove several lower bounds on the depth required to represent functions by monotone ReLU networks and \mathsfICNN. For the maximum function \mathsfMAX_n computing the maximum of n real numbers, we show that \mathsfReLU^+ networks cannot compute \mathsfMAX_n, or even approximate it. We prove a sharp n lower bound on the \mathsfICNN depth complexity of \mathsfMAX_n. We also prove depth separations between \mathsfReLU networks and \mathsfICNNs; for every k, there is a depth-2 \mathsfReLU network of size O(k^2) that cannot be simulated by a depth-k \mathsfICNN. The proofs combine ideas from some structural results for \mathsfReLU^+ networks.
Abstract: Reactive synthesis, the problem of automatically constructing a hardware circuit from a logical specification, is a long-standing challenge in formal verification. It is elusive for two reasons: It is algorithmically hard, and writing formal specifications by hand is notoriously difficult. In this paper, we tackle both sides of the problem. For the algorithmic side, we present a neuro-symbolic approach to reactive synthesis that couples large reasoning models with model checkers to iteratively repair a synthesized Verilog implementation via sound symbolic feedback. Our approach solves more benchmarks than the best dedicated tools in the annual synthesis competition and extends to constructing parameterized systems, a problem known to be undecidable. On the specification side, we introduce an autoformalization step that shifts the specification task from temporal logic to natural language by introducing a hand-authored dataset of natural-language specifications for evaluation. We demonstrate performance comparable to that of starting from formal specifications, establishing natural synthesis as a viable end-to-end workflow.
Authors: Nesreen Ahmed, Nima Nafisi
Abstract: Monitoring autonomous large language model (LLM) agents for covert malicious behavior (e.g., covertly pursuing a hidden malicious objective) is challenging due to delayed, context-dependent, and long-horizon attack patterns. In adversarial settings such as sabotage, agents may pursue hidden objectives while maintaining superficially benign behavior, making detection difficult even with full trajectory access. Prior monitoring approaches primarily improve monitor scaffolding or ensemble aggregation, but treat each trajectory independently and do not improve from prior monitoring experience. Moreover, standard reasoning methods explain observed behavior but do not explicitly reason about agent beliefs, intentions, and goal alignment required to distinguish benign task execution from covert deviation. We propose Agent-ToM, a learning-to-monitor framework grounded in Theory-of-Mind (ToM) reasoning for security analysis of autonomous agents. Agent-ToM performs structured full-trajectory analysis by inferring beliefs, step-level intent hypotheses with calibrated confidence, expected actions, and deviations from task-consistent behavioral baselines. At inference time, it employs a Reason--Verify--Refine pipeline to construct and validate monitoring decisions. At training time, Agent-ToM learns from prior monitoring episodes by distilling critique signals into a persistent semantic guardrail memory that accumulates monitoring strategies, enabling reusable belief- and intent-conditioned constraints to be applied across episodes. We evaluate Agent-ToM on adversarial agent monitoring benchmarks (SHADE-Arena and CUA-SHADE-Arena). Agent-ToM achieves strong precision--recall balance and outperforms state-of-the-art monitoring baselines, including ensemble methods, while using a single coherent reasoning pipeline. These results demonstrate that \emphlearning at the monitoring layer, combined with structured ToM reasoning and verification, provides an effective and deployable foundation for securing autonomous LLM agents.
PaperID: 5452, Poster
Authors: Yiming Liu, Binhang Qi, Weiyu Kong, Jiawei Liu, Xinxin Shan, Saijun Gao, Yun Lin
Abstract: While agentic frameworks have advanced repository-level code localization, their reliance on static snapshots often leads to a traceability gap—the loss of implicit links between issue symptoms and code implementation. To bridge this information deficit, we introduce Traceability Augmented Code Localization (TACO), a framework that recovers retrievable and usable traceability from repository evolution to guide localization agents. TACO offline crystallizes historical pull requests into a dual-index knowledge base through a prompt auto-tuning technique, capturing high-value semantic hints and architectural rationales. In the online phase, TACO employs a synergistic dual-track retrieval workflow, cross-validated by an arbiter and aligned temporally to handle version drift. Extensive evaluations on SWE-bench Lite and SWE-bench Verified across five state-of-the-art agentic baselines demonstrate that TACO significantly improves exact-match localization accuracy (Acc@1) by an average of 11.29%, 13.20%, and 12.87% at the file, module, and function levels, respectively. Moreover, TACO significantly streamlines the exploration process, achieving an average reduction of 29.2% in token usage and 30.4% in monetary costs, with minimal offline maintenance overhead.
PaperID: 5453, Poster
Authors: Daehee Kim
Abstract: Modern large language models keep the feed-forward network (FFN) intermediate width identical across every layer, a convention inherited from the original Transformer. Yet mechanistic work suggests layers at different depths do qualitatively different work, raising a natural question: are their capacity demands also non-uniform, and if so, can a fixed parameter budget be allocated more wisely than uniformly? To probe this, we apply principal component analysis (PCA) to the input of each FFN down-projection and measure, for every layer, the smallest number of principal components needed to reconstruct the activation up to a fixed cosine similarity --- the layer's intrinsic dimensionality. Scanning 58 open-weight LLMs spanning 14 families and 70M to 72B parameters, we find a recurring middle-heavy profile in which middle layers carry higher intrinsic dimensionality than edge layers; the profile is robust to the calibration corpus, even across English and Chinese, and is largely preserved through post-training fine-tuning. As a direct test that this signal is structurally meaningful, we use it to guide width pruning in a knowledge-distilled student under a matched parameter budget: the PCA-guided allocation outperforms the uniform-width convention by +2.7 points on Qwen2.5-3B and +2.3 on Qwen3-8B, while a same-scale teacher (LLaMA-3-8B) shows only +0.9, and a budget-matched inverted control consistently does worst. Across these three teachers, the size of the gap scales with how unevenly the teacher's intrinsic dimensionality is distributed across layers --- a single-number summary we call \sigma_99 --- and the advantage persists through downstream supervised fine-tuning.
Authors: Yishu Zhang, Yun Li, Daiwei Zhang
Abstract: State-of-the-art pathology foundation models, trained on millions of histology tiles, can fail to preserve tissue similarity when comparisons cross slide or institution boundaries. We show that general-purpose multimodal LLMs, without being trained as pathology foundation models, consistently outperform these specialized models in cross-domain histological similarity judgments. Using a relative similarity framework that we release as the MOSAIC (Model Similarity Assessment across Institutions and Cohorts) benchmark, we evaluate 17 models across 6 datasets and find that pathology encoders often rank same-institution, different-disease tiles as more similar than same-disease, different-institution tiles, a clinically dangerous failure mode invisible to standard within-domain evaluations. LLMs appear less susceptible to this failure, likely because they perform semantic visual comparison of morphology and tissue architecture rather than relying on shortcut features tied to acquisition context. Scaling training data does not resolve the problem for pathology encoders, implicating the learning objective rather than data coverage. Our results expose a fundamental robustness gap in current pathology foundation models and establish multimodal LLMs as a viable alternative for cross-institutional retrieval, dataset harmonization, and multi-site quality control. Code and data will be released upon acceptance.
Abstract: Long-context LLM agents often struggle with growing token, memory, and latency costs, making efficient context compression essential for practical deployment. Existing LLM-as-a-compressor methods remain noticeably inferior to using the full context. We find that this gap partly stems from their inability to preserve contextual information effectively. In this work, we revisit context compression from a structural perspective and identify two key bottlenecks in standard LLM-based compressors: limited coordination among compression tokens during information aggregation, and layerwise dilution that weakens useful signals from intermediate hidden states. To address these limitations, we propose ransmission), a new context compression framework based on explicit information transmission. ComprExIT adaptively selects features across frozen LLM layers, then allocates information from anchors to compression slots through a globally coordinated transport plan. Experiments on 12 datasets show that ComprExIT consistently outperforms strong soft-compression baselines, improving average F1 by up to 18.5%, while adding only ~1% trainable parameters and achieving more than 2× faster compression than the fastest baselines. The code will be released upon acceptance.
PaperID: 5456, Poster
Abstract: Federated graph learning (FGL) enables privacy-preserving collaborative training over distributed graph data, yet remains vulnerable to backdoor attacks from malicious clients. In FGL, data and structural heterogeneity make benign clients difficult to distinguish from malicious ones, posing a key challenge to reliable defense. We propose VeriDrift, a backdoor defense framework for FGL based on cross-layer drift and consistency verification. On the client side, VeriDrift derives a client-level risk score from node-level cross-layer drift anomalies during GNN propagation. On the server side, VeriDrift verifies the consistency between the reported risk and the submitted model update, and converts the resulting verified risk into aggregation weights to suppress suspicious updates. Without relying on auxiliary clean data, VeriDrift enables reliable risk assessment under heterogeneous FGL and improves robustness against adaptive attacks. Extensive experiments show that VeriDrift consistently reduces attack success rates across different FGL backdoor settings while maintaining clean accuracy.
PaperID: 5457, Poster
Abstract: In time-series forecasting, rollout errors are not merely one-step fitting discrepancies: once fed back into the evolving state, they can perturb future predictions. We formalize this effect through local rollout consistency, which compares the next prediction produced under teacher forcing with that produced under free running. In principle, this consistency can be enforced directly by tracking the residual-induced perturbation of the next state. In practice, however, the residual-to-state alignment operator is often difficult to design when the input state and output block have different dimensions, heterogeneous components, or architecture-dependent layouts. We therefore use a first-order local surrogate that penalizes the Jacobian-vector product of the predictor along the actual feedback perturbation, suppressing precisely the directions through which self-generated errors remain visible to future predictions. We further give the same loss a geometric interpretation near a predictive manifold: by controlling residual-induced directions that probe the normal bundle, the objective yields theoretical conditions under which free-running trajectories remain inside a stability tube around the manifold. Experiments on time series forecasting tasks support this view, showing that the proposed regularization improves predictive accuracy. More broadly, our results suggest a training principle for time-series forecasters and sequential predictors: a model should not only make small one-step errors, but also learn not to let its own errors echo into the future.
PaperID: 5458, Poster
Abstract: Small language models deployed under cost, latency, or privacy constraints cannot simply switch to larger models when reasoning fails. Multi-candidate generate then-rerank offers a natural test-time scaling route. However, more candidates do not always yield higher accuracy. We identify a previously overlooked bottleneck between generation and selection: a correct answer may appear in the candidate pool yet fail to be selected because it is incomplete, unparseable, or misaligned with the answer mode the fixed selector expects. We term this loss the hidden interface tax and formalize a coverage–conversion decomposition that distinguishes raw oracle coverage from selectable coverage, the subset of problems where a correct candidate can actually be consumed by the selector. We propose Compatibility Aware Candidate Construction (CACC), an interface layer between the candidate source and the fixed selector that repairs fragments, aligns answer modes, and removes selector-confusing noise without retraining either component. On numeric reasoning (GSM8K, competition_math), CACC raises both oracle coverage and final accuracy. On GPQA Diamond, CACC alone improves final accuracy from 5.6% to 19.2% (+13.6pp); combined with proposer-side strengthening it reaches 25.8% (+20.2pp). On MMLU-Pro, CACC exposes a coverage–conversion gap: answer-mode-compatible construction raises oracle coverage, but final gains require additional proposer-side distribution shaping.
PaperID: 5459, Poster
Authors: TIANQI ZHAO, Xixi Liu, Liangrui Peng, Zhengrui Xiang
Abstract: Although diffusion models achieve promising image generation performance, uncertainty during the iterative denoising process can lead to visual artifacts. Existing uncertainty estimation methods typically quantify it at the pixel level. However, these methods assume independence among pixels, neglecting pixel correlations crucial for image structure. In contrast, we propose a frequency-domain uncertainty estimation method that captures structural correlations. We empirically show that samples with artifacts exhibit higher frequency-domain uncertainty, and derive a Louis-identity-based connection between the estimated uncertainty and the optimal reverse-process covariance. To this end, we develop a structure-aware diffusion sampling guidance framework. For the mean of the reverse process, the gradient of the uncertainty is used to penalize specific frequency components; for the reverse covariance, frequency-domain uncertainty is used as a proxy to modulate injected noise. Experiments on the ImageNet, LSUN-Churches, and DrawBench datasets across U-Net, U-ViT, and SD3 architectures validate our method.
PaperID: 5460, Poster
Abstract: Universal object localization in X-ray security inspection is critical for automated threat detection in safety-critical venues. However, unlike everyday RGB images that dominate web-scale visual data, X-ray scans exhibit distinct color patterns, ambiguous boundaries, and compositional structures caused by volumetric superposition. These gaps hinder the direct zero-shot transfer of dense perception foundation models trained on web-scale RGB data. Moreover, annotated X-ray data is scarce and requires expert labeling, limiting both the training of generalizable X-ray native models and the adaptation of RGB foundation models for X-ray data via fine-tuning. Given these challenges, the bright promise of highly generalizable perception models, enabled by data scaling laws in the RGB domain, remains largely out of reach for X-ray inspection. To this end, we introduce LAO-X, a fully self-supervised framework that Locates Any Object in X-ray scans using diverse synthesized image--annotation pairs with granularity-aware supervision. LAO-X first designs a saliency-guided X-ray object mining module to separate diverse object instances, which are then used for physics-guided synthesis in the absorbance domain. LAO-X further incorporates an occlusion-controlled curriculum strategy to fine-tune a Segment Anything Model 2 (SAM2) localizer, progressively adapting it to X-ray scans with increasing object counts and overlap levels. Experiments on six X-ray benchmarks show that LAO-X substantially improves category-agnostic localization, achieving 2% to 23% mAP gains over SAM2 and X-ray specific baselines in heavily cluttered scenarios, entirely without human-annotated labels.
Abstract: Reinforcement learning with verifiable rewards (RLVR) has driven recent capability advances of large language models across domains. Recent studies suggest that improved RLVR algorithms allow models to learn effectively from incorrect annotations, achieving performance comparable to learning from clean data. In this work, we show that these findings are invalid because the claimed 100% noisy training data is "contaminated" with clean data. After rectifying the dataset with a rigorous re-verification pipeline, we demonstrate that noisy data is destructive to RLVR. We show that existing RLVR algorithm improvements fail to mitigate the impact of noisy data, achieving similar performance to that of the basic GRPO. Furthermore, we find that the model trained on truly incorrect annotations performs 8-10% worse than the model trained on clean data across mathematical reasoning benchmarks. Finally, we show that these findings hold for real-world noise in Text2SQL tasks, where training on real-world, human annotation errors cause 5-12% lower accuracy than clean data. Our results show that current RLVR methods cannot yet compensate for poor data quality. High-quality data remains essential.
PaperID: 5462, Poster
Abstract: Dense image and video captioning demands exhaustive visual details while strictly maintaining factuality, forcing a severe trade-off between descriptive density and hallucination. Supervised fine-tuning (SFT) struggles here: for dense captions, its token-level loss dilutes effective signals while failing to reinforce specific fine-grained dimensions. Reinforcement learning (RL) is a promising alternative but remains bottlenecked by reward design: coarse-grained VLM judges suffer from reward drift, and static QA rewards quickly saturate as the policy improves. To address this, we propose DiffCap-RL, an RL framework combining a stable offline Pre-Scorer with an adaptive online Diff Scorer. We utilize the Pre-Scorer as a stable anchor for broad visual grounding via automated QA pairs. Once it loses discriminative power, Online Diff Scorer dynamically extracts semantic disagreements between policy rollouts. By decomposing these disagreements into targeted conflict questions to suppress hallucinations and extra questions to reward valid new details, the Diff Scorer provides reliable, non-saturating supervision at the policy frontier. We validate our reward quality from three perspectives: human alignment, online RL signal quality, and BoN selection. Across multiple image and video captioning benchmarks, our DiffCap-RL delivers large gains and consistently outperforms state-of-the-art baselines at similar or larger scales.
PaperID: 5463, Poster
Authors: Seokeon Choi, Sunghyun Park, Hyoungwoo Park, Jeongho Kim, Sungrack Yun
Abstract: Memory-efficient personalization is essential for adapting text-to-image diffusion models while preserving user privacy and operating within the limited computational resources of edge devices. To this end, we propose Mix-Opt, a novel mixed optimization framework that adaptively chooses between backpropagation on low-resolution images (BP-low) and zeroth-order optimization on high-resolution images (ZO-high), guided by the characteristics of the diffusion process. As observed in our experiments, BP-low efficiently adapts the model to target-specific features, but suffers from structural distortions due to resolution mismatch. Conversely, ZO-high refines high-resolution details with minimal memory overhead but faces slow convergence when applied without prior adaptation. By complementing both methods, our framework leverages BP-low for effective personalization while using ZO-high for structural consistency, achieving memory-efficient and high-quality fine-tuning. To maximize the efficacy of both BP-low and ZO-high, we introduce a timestep-aware probabilistic function that dynamically selects the appropriate optimization strategy during training. This function mitigates the overfitting from BP-low at high timesteps, where structural information is critical, while ensuring ZO-high is applied more effectively as training progresses. Experimental results demonstrate that our method achieves competitive performance while significantly reducing memory consumption, enabling scalable, high-quality personalization without increasing inference latency.
PaperID: 5464, Poster
Authors: Tongyu Zhou, Wei Zhang
Abstract: Unit-modulus measurement vectors, due to their phase-only variations, exhibit favorable hardware compatibility and storage efficiency. In this paper, we consider matrix recovery using symmetric unit-modulus rank-one measurements. We identify a fundamental limitation of symmetric unit-modulus measurements---their inability to capture individual diagonal entries, and show that matrix recovery can be achieved using such measurements when the diagonal entries are known or directly measurable. We construct a stacked operator and establish that it satisfies a mixed-norm restricted isometry property (\mathrmmRIP) for unit-modulus measurements. Leveraging the \mathrmmRIP condition, we derive exact and stable recovery results for matrix recovery with unit-modulus measurements in both noiseless and noisy settings. Our work is the first to provide the theoretical foundation for matrix recovery under symmetric rank-one unit-modulus measurements. Numerical experiments further corroborate these theoretical findings.
PaperID: 5465, Poster
Abstract: LLM-based multi-agent systems extend single agents with role specialization and inter-agent communication, but they also expand the attack surface: an attacker can compromise an externally writable module of one agent, and the malicious payload may propagate across communication hops before some downstream agent invokes an attacker-chosen tool. Existing defenses are routinely circumvented by stronger or adaptive attacks, which motivates a complementary forensic capability that traces an incorrect tool invocation back to its root cause. We formulate this problem as three-level forensic localization, where the goal is to jointly identify the source agent whose internal state was first compromised, the specific module of that agent that was exploited, and the exact units inside that module carrying the malicious payload. To enable systematic evaluation, we construct MAFL-Bench, the first forensic localization benchmark for LLM-based multi-agent systems, which spans three application scenarios, four communication structures, and twelve attack instances, with 21,200 interaction logs annotated at the agent, module, and unit levels. We further propose MATracer, a black-box forensic framework that resolves the three-level attribution through a coarse-to-fine cascade driven by log-likelihood queries on a proxy LLM, progressively localizing the source agent, the compromised module, and the contaminated units. Extensive experiments show that MATracer accurately localizes the attack across diverse multi-agent settings, consistently outperforms ten baselines, and remains robust under adaptive attacks.
Abstract: Federated learning (FL) enables collaborative model training across distributed clients without sharing raw data, yet its scalability is limited by synchronization overhead. Asynchronous federated learning (AFL) alleviates this issue by allowing clients to communicate independently, thereby improving wall-clock efficiency in large-scale, hardware-heterogeneous environments. However, asynchrony introduces updates computed on outdated global models (staleness) that can destabilize optimization and hinder convergence. We propose FedRevive, an AFL framework that revives stale updates through data-free knowledge distillation (DFKD). FedRevive integrates parameter-space aggregation with a server-side DFKD process that transfers knowledge from stale client updates to the current global model without access to data. A meta-learned generator synthesizes pseudo-samples for multi-teacher distillation. A hybrid aggregation scheme combining raw with DFKD updates effectively mitigates staleness while retaining AFL scalability. Experiments on vision and text benchmarks show that FedRevive achieves faster training by up to 29.7% and higher final accuracy by up to 11.4% percentage points than baselines.
PaperID: 5467, Poster
Abstract: This work proposes WaveGen, a fully end-to-end diffusion-based text-to-waveform framework based on Spectral Forcing, which jointly learns an internal spectral trajectory and the target waveform trajectory within a single model, enabling direct end-to-end text-to-waveform generation. Unlike recent TTS pipelines, WaveGen does not rely on separately trained or externally decoded intermediate representations, such as neural audio codecs, Mel-spectrogram vocoders, or VAE-based latent models. In addition, WaveGen performs internal text-speech alignment within a single model, eliminating the need for external alignment modules such as duration predictors. To preserve a truly end-to-end formulation, our framework further avoids dependence on self-supervised representations, such as BERT-based text embeddings and Wav2Vec 2.0-based semantic representations, as well as semantic distillation methods for training optimization. To achieve this, we carefully design Spectral Forcing, a model architecture with disentangled diffusion heads, and a training objective based on a multi-scale spectral v-loss. We further validate the scalability of end-to-end TTS models by scaling the model size from 0.1B to 0.4B parameters. With a single architecture, open-source data, and one-stage training without external modules, WaveGen demonstrates fully end-to-end waveform generation by achieving promising performance. In particular, WaveGen achieves a WER of 1.77, a SIM of 0.63, and competitive MOS performance on the Seed-en benchmark.
PaperID: 5468, Poster
Authors:
Miao Wang, Gangyi Ding, Yunlin Lei, Yu Zhang, chenjingfeng, Xu YangAbstract: Most time-series anomaly detectors ask whether an observation can be predicted or reconstructed. We ask a different question: can a detector’s own spiking dynamics become anomaly evidence? We study this question for periodic and quasi-periodic time series and propose SPIRE, a phase-conditioned spiking anomaly detector. SPIRE uses period-aware current formation to expose phase-aligned cross-period residuals to leaky integrate-and-fire neurons. We refer to this conversion as eventification: after residual injection, recurrence violations induce abnormal bursts, silences, or phase-shifted spike trains. SPIRE then constructs phase-conditioned normal templates of spike states and scores deviations in residual energy, firing-rate patterns, and burst/silence statistics, instead of relying solely on output-space prediction or reconstruction errors. This separates the operational protocol, including causal current-step detection, one-step prediction, and non-causal reconstruction, from the source of anomaly evidence, namely output errors, spike dynamics, or both. On TSB-AD-U and TSB-AD-M, SPIRE establishes new best VUS-PR among neural time-series anomaly detectors. Spike-native scores remain competitive without reconstruction error, hybrid scoring consistently improves error-only scoring, and replacing LIF neurons with continuous activations in matched ANN counterparts reduces detection performance under the same period-aware architecture. Periodicity-stratified analyses and spike-state visualizations further show that the gains are strongest when phase recurrence is reliable, supporting the proposed spike-as-detector mechanism. These results suggest that phase-conditioned spiking dynamics can serve as a first-class anomaly detection signal for periodic time series, rather than merely an efficient computational substrate.
Authors:
Chanyoung Kim, Myeonghwan Seong, Kim Yujin, Daniel Kyungdeock Park, Youngjoon HongAbstract: Partial differential equations (PDEs) are central to modeling physical and engineering systems, but repeatedly solving parametric PDEs remains computationally expensive. Operator learning enables fast surrogate inference, yet typically requires large input–output paired datasets generated by costly high-fidelity PDE solvers. Unsupervised operator learning frameworks alleviate data dependency but remain hindered by computational bottlenecks. To address this, we propose Neural Variational Quantum Linear Solver (NVQLS), the first hybrid quantum–classical operator learning framework leveraging the Legendre-Galerkin weak formulation. We critically resolve the sign ambiguity in VQLS energy minimization, preventing erroneous solution representations. Additionally, we introduce a neural embedding, a novel encoding scheme to map varying forcings and PDE coefficients into parameterized quantum circuit representations. These structural innovations provide theoretical computational complexity advantages under efficient state preparation schemes, while achieving superior accuracy compared to a representative classical baseline. Validations on 1D and 2D parametric PDEs under diverse boundary conditions demonstrate NVQLS's capability to simultaneously process varying inputs, offering a scalable unsupervised approach to quantum-enhanced operator learning.
PaperID: 5470, Poster
Abstract: Large language model (LLM) agents have shown strong potential in tackling complex, multi-step tasks, yet scaling them to long-horizon settings remains fundamentally constrained by two problems: context window saturation and uniform allocation of inference-time compute across steps of vastly different cognitive demands. We introduce ontrol), a hierarchical agent framework that decomposes long-horizon decision making into three temporal scales: strategic, tactical, and operational, each governed by distinct context scopes and reasoning budgets. Central to MOSAIC is a , a lightweight policy trained via reinforcement learning that dynamically determines (i) which abstraction level should be invoked at each time step, and (ii) how many reasoning tokens to allocate, based on the estimated decision complexity. We further propose , a bidirectional information compression mechanism that maintains coherent state representations across scales without redundant token consumption. Theoretically, we show that under a fixed total inference budget, MOSAIC's adaptive allocation policy achieves a strictly lower regret bound than any fixed-rate policy in a hierarchical semi-MDP formulation. Empirically, MOSAIC achieves state-of-the-art or competitive results on four diverse long-horizon benchmarks: SWE-Bench Verified, BrowseComp, WebArena, and GAIA (Level 2&3), while reducing total inference tokens by 38--57% relative to flat ReAct baselines. Ablation studies confirm that each component of multi-scale decomposition, adaptive scheduling, and context distillation contributes meaningfully to both performance and efficiency.
Abstract: We introduce TextSeal, a state-of-the-art watermark for large language models. Building on Gumbel-max sampling, TextSeal introduces dual-key generation to restore output diversity, along with entropy-weighted scoring and multi-region localization for improved robustness. It supports serving optimizations such as speculative decoding and multi-token prediction, and does not add any inference overhead. TextSeal strictly dominates baselines like SynthID-text in detection strength and is robust to dilution, maintaining confident localized detection even in heavily mixed human/AI documents. The scheme is theoretically distortion-free, and evaluation across reasoning benchmarks confirms that it preserves downstream performance; while a multilingual human evaluation (6000 A/B comparisons, 5 languages) shows no perceptible quality difference. Beyond its use for provenance detection, TextSeal is also "radioactive": its watermark signal transfers through model distillation, enabling detection of unauthorized use.
PaperID: 5472, Poster
Abstract: Robots operating in semantic environments often need to satisfy linear temporal logic (LTL) tasks before the semantic map is fully known. Many practical semantic-LTL planners maintain semantic beliefs but execute a single task policy induced by a point semantic interpretation, such as a maximum-a-posteriori (MAP) semantic map, possibly supported by risk checks or active perception. This premature commitment can be brittle when distant semantic observations are weak or correlated. We propose Coherence-Aware Transition-Intent Fusion, an online planner that preserves automaton-relevant semantic uncertainty using a finite top-K intent abstraction. Given a conservative current automaton estimate, the planner propagates plausible semantic assignments through the automaton, converts feasible successor transitions into weighted reach-avoid intents, and fuses their risk-gated Dijkstra progress scores. When competing intents create incoherent motion, a KL-regularized calibration step shifts weight toward intents consistent with recent physical progress. Experiments on random-obstacle and room-like semantic maps show improved satisfaction over single-intent and raw-fusion variants, and over a TOAPP MAP-plus-active-perception baseline under the same static-view sensor. Ablations attribute the gain primarily to intent-level calibration.
PaperID: 5473, Poster
Abstract: This paper studies the problem of jointly learning a reward function, a cost function, and a policy from an expert's preferences. We formulate the problem as a constrained bi-level optimization problem, where the upper level infers the reward and cost functions from preferences, while the lower level optimizes a policy to best align with those preferences. To solve this problem, we propose a double-loop algorithm, Constrained Bi-level Optimization for Preference-Based Reinforcement Learning (CB-PbRL), which solves the lower-level optimization problem in the inner loop and the upper-level optimization problem in the outer loop. We establish a theoretical guarantee that CB-PbRL converges at a rate of \mathcalO(1/\sqrtK), and we demonstrate its effectiveness across multiple simulation environments.
PaperID: 5474, Poster
Abstract: High-performance text-to-image~(T2I) diffusion models suffer from diversity degradation. In this study, we attribute this challenge to a fundamental temporal mismatch: while visual features evolve dynamically in a coarse-to-fine manner in the denoising process, the text condition remains entirely static. Through empirical analysis, we demonstrate that decoupling the text prompt to prioritize low-frequency components in the early denoising stage better aligns with the visual denoising process, while simultaneously enhancing generation diversity. Driven by this insight, we introduce Dynamic Semantic Interpolation (DSI), a training-free strategy that utilizes low-frequency semantics during the initial stages to foster the sampling space and progressively recovers the full semantics, ensuring text-to-image alignment. We also provide a rigorous theoretical justification for DSI from the perspective of conditional entropy, explaining its ability to maintain a broader early sampling space. Extensive experiments across diverse state-of-the-art models demonstrate that DSI significantly boosts generation diversity while preserving text alignment and aesthetic quality.
Authors:
Agnibh Dasgupta, Abdullah All Tanvir, Xin ZhongAbstract: Language models exhibit strong robustness to paraphrasing, suggesting that semantic information may be encoded through stable internal representations, yet the structure and origin of such invariance remain unclear. We propose a local geometric framework in which semantically equivalent inputs occupy structured regions in latent space, with paraphrastic variation along nuisance directions and semantic identity preserved in invariant subspaces. Building on this view, we make three contributions: (1) a geometric characterization of invariant latent features, (2) a contrastive subspace discovery method that separates semantic-changing from semantic-preserving variation, and (3) an application of invariant representations to zero-shot model attribution. Across models and layers, empirical results support these contributions. Invariant structure emerges in specific depth regions, semantic displacement lies largely outside the nuisance subspace, and representation-level interventions indicate a causal role of invariant components in model outputs. Invariant representations also capture model-specific geometric patterns, enabling accurate attribution. These findings suggest that semantic invariance can be viewed as a local geometric property of latent representations, offering a principled perspective on how language models organize meaning.
Authors:
Jingwei Song, Meng Chen, Jie Xiao, Qingnan Ren, Jiaqi Huang, Yangshen Deng, Senyu Tong, Wanyi Chen, Suli Wang, Zhisheng Chen, Ziqian Bi, Shuo Lu, Yiqun Duan, Xu Wang, Rymon Yu, Lynn Ai, Eric Yang, TIANYU SHIAbstract: Reinforcement learning (RL) is a critical stage in post-training large language models (LLMs), involving repeated interaction between rollout generation, reward evaluation, and centralized learning. Distributing rollout execution offers opportunities to leverage more cost-efficient inference resources, but introduces challenges in wide-area coordination and policy dissemination. We present ECHO-2, a distributed RL framework for post-training with remote inference workers and non-negligible dissemination latency. ECHO-2 combines centralized learning with distributed rollouts and treats bounded policy staleness as a user-controlled parameter, enabling rollout generation, dissemination, and training to overlap. We introduce an overlap-based capacity model that relates training time, dissemination latency, and rollout throughput, yielding a practical provisioning rule for sustaining learner utilization. To mitigate dissemination bottlenecks and lower cost, ECHO-2 employs peer-assisted pipelined broadcast and cost-aware activation of heterogeneous workers. Experiments on GRPO post-training of LLMs ranging from 4B to 32B parameters under real wide-area bandwidth regimes show that ECHO-2 significantly improves cost efficiency while preserving RL reward comparable to strong baselines.
PaperID: 5477, Poster
Abstract: Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provide weak credit assignment over long interaction traces. Step-level rewards can offer more targeted feedback, but obtaining reliable step supervision often requires human or LLM judgment, or additional rollouts to estimate the downstream effect of an intermediate decision. In this paper, we argue that effective tool-use agents should estimate the long-horizon value of a possible next tool invocation before executing it. This objective requires comparative supervision over alternative invocations under the same context, while logged trajectories only contain the invocation that was actually taken. Therefore, we propose Comparative Inference for Tool-use Agents (CITA). CITA trains a Comparative Inference Model (CIM) from paired signals that combine observed tool behavior, scalable supervision from a Bayesian tool-graph simulator, and semantic judgments from LLM-based comparison. The resulting CIM learns to estimate how likely a possible next tool invocation is to support final task success under the current context. Across three tool-use benchmarks and multiple backbone LLMs, CITA consistently improves Tool F1 and task success. Additional analysis shows that CIM learns accurate step-level value estimates for comparative tool choices.
Abstract: Recursive training of generative models on their own outputs can lead to model collapse, a compounding drift away from the true data distribution. Existing theoretical works bound finite-round error accumulation in the context of diffusion models, but two questions remain open:~what distribution does the recursion converge to, and how fast? We answer both, isolating a mechanism distinct from imperfect learning: even with perfect score estimation and exact sampling, the early stopping of the reverse diffusion (required for numerical stability) drives a progressive drift away from the data distribution. We prove that this recursion converges geometrically to a unique limiting distribution, which admits a closed-form characterization as an infinite mixture of increasingly Gaussian-smoothed versions of the data distribution. A Hermite spectral decomposition of this limit reveals that recursive training acts as a low-pass filter: higher-order modes, which encode fine non-Gaussian structure, are attenuated much more strongly than coarse modes. This spectral picture motivates annealed truncation schedules that progressively shrink truncation times across retraining rounds; we prove that any schedule converging to 0 asymptotically eliminates recursive compounding. Finally, we show our idealized characterization is robust: in the presence of discretization and score estimation errors, the learned distribution remains in a Wasserstein-2 ball around the ideal limit, with mode-dependent contraction rates that contract high-order errors faster than low-order ones. We validate the theory on synthetic Gaussian mixtures and CIFAR-10.
PaperID: 5479, Poster
Abstract: We study label-free, topology-conditioned zero-shot cross-graph link prediction: at test time the observed target topology is the input graph, but the model receives no target labels, no positive/negative context links, no target fine-tuning, no text attributes, and no post-hoc alignment. This setting is zero-shot with respect to target supervision and adaptation, but it is not topology-free inference. We identify a critical geometric barrier for hyperbolic transfer: radial non-identifiability. Hyperbolic graph learning is motivated by the intuition that radius encodes hierarchy or popularity and angle encodes similarity, yet standard HGNN objectives do not make radii semantically calibrated across disjoint graphs. The same structural role can therefore occupy incompatible radial scales on different target graphs. We propose Invariant Hyperbolic Unfolding (IHU), which restores the intended radial hierarchy channel through selective invariance: radii are fixed by graph-internal structural percentiles while angular representations remain learnable. Its core module, Invariant Structural Anchoring (ISA), rank-canonicalizes structural scores such as coreness into shared hyperbolic radii, producing a comparable radial skeleton without target supervision. In a controlled comparison among matched frozen learned encoders, IHU improves average HR@50 by 2.7 percentage points over the strongest learned hyperbolic baseline and by 7.6 percentage points over the mean of the learned hyperbolic baseline family. The gain amplifies under missing-edge and low-degree regimes, while hierarchy-gap analysis reveals a boundary condition for fixed radial shells: 50% edge sparsification causes 52.5% less degradation, and low-degree nodes show a 4.2× larger advantage. We further position IHU against recent universal link prediction and graph foundation methods through an access-aware taxonomy, showing that the closest recent LP methods use labeled target context links, whereas IHU isolates the geometric effect of fixed versus learned radii in frozen structural hyperbolic encoders.
PaperID: 5480, Poster
Abstract: Recent advances in semiconductor technology have dramatically improved IC chip performance, but soaring transistor densities and clock speeds generate unprecedented heat fluxes in modern integrated circuits, making spatially adaptive thermal management critical. However, conventional heat sinks rely on fixed geometries that lack customization for chip-specific thermal profiles, while the absence of optimal design datasets and strict manufacturing constraints hinder the application of generative AI. To address these challenges, we present DiffCool, a physics-guided framework that reformulates heat sink synthesis as a constrained discrete denoising process. By embedding manufacturing-aware state transitions into the diffusion dynamics and optimizing via a differentiable thermal-aware loss coupled with a surrogate simulator, our model generates structurally sound, high-performance topologies directly from bare-chip heat maps without labeled data. Experiments on an open-source processor demonstrate that DiffCool synthesizes production-ready designs in under one second, reducing peak temperature by 28.3% and temperature gradient by 56.9% compared to conventional baselines. This work establishes a scalable, label-free paradigm that unifies generative modeling with physical constraints for automated thermal design.
PaperID: 5481, Poster
Authors: Nika Chuzhoy, Brian Hu, Amit A Arora, Jae Ro, Sarthak Sahu
Abstract: Image geolocation aims to estimate where a photograph was taken from its visual content. At worldwide scale, this remains challenging because visual evidence is often ambiguous, diverse, and unevenly distributed. Prior work has typically treated geolocation of ordinary internet photos and street-view imagery as separate tasks, despite their complementary strengths: internet photos better match the appearance distribution of user-captured queries, while street-view imagery provides denser, geographically grounded coverage. We present Pinpoint, a retrieve-and-rerank architecture that combines both sources in a coarse-to-fine pipeline. A contrastive image-GPS embedder is trained on both user-uploaded Flickr photos and street-view imagery, learning a shared image-GPS embedding space that is used to retrieve candidate locations. An attention-based reranker then rescores retrieved candidates by combining candidate-level visual and GPS features with cross-source evidence from nearby locations to ground the prediction. Unlike recent prior work, Pinpoint does not rely on multimodal large-language models, making inference faster and more reproducible. Pinpoint achieves state-of-the-art results across all metrics on standard benchmarks for internet photos (IM2GPS3k and YFCC4k) and street-view imagery (OSV-5M).
PaperID: 5482, Poster
Authors: Moein Khajehnejad, Forough Habibollahi, Leonardo Novelli, Adeel Razi
Abstract: Estimating directed, weighted, and signed interactions among brain regions (effective connectivity) from fMRI requires disentangling neural dynamics from delayed and nonlinear hemodynamic transformations. Dynamic Causal Modeling (DCM) provides a principled solution by inverting a biophysical generative model, but the standard Variational Laplace inversion method is computationally prohibitive for large parcellations and cohort-scale datasets. We introduce BASIL-DCM, a physics-informed amortized inference model that estimates subject-specific effective connectivity and biophysical DCM parameters, including ROI-wise hemodynamic transit time, spectral properties of endogenous neural fluctuations, and observation noise, in a single forward pass. BASIL-DCM combines a linear-time state-space temporal encoder with an ROI-wise Transformer to capture long-range temporal dependencies and inter-regional interactions. The model is trained on data informed by Human Connectome Project resting-state fMRI, with effective connectivity initialized from regression DCM and complementary biophysical parameters sampled from physiologically plausible ranges. Learning is further constrained by a differentiable spectral-consistency objective derived from the DCM forward model. This approach enables fast, uncertainty-aware whole-brain network inference while preserving mechanistic interpretability and ensuring consistency with biophysical dynamics.
Authors: Amirhossein Yousefiramandi
Abstract: Forward-Forward (FF) training allows each layer to learn from a local goodness criterion. In cumulative-goodness variants, however, later layers can inherit a task that earlier layers have already partially separated. We formalize this phenomenon as layer free-riding: under the softplus FF criterion, the class-discrimination gradient reaching block d decays exponentially with the positive margin accumulated by preceding blocks. We then study three local remedies---per-block, hardness-gated, and depth-scaled---that recover current-layer separation measures without relying on backpropagated gradients. On CIFAR-10 and CIFAR-100, these remedies dramatically improve layer-separation statistics, with 4×--45× gains in deeper layers, while changing accuracy by less than one percentage point for non-degenerate training procedures. Tiny ImageNet provides a tougher cross-dataset check for our selected block-wise configuration and reveals the same qualitative gap between layer-health diagnostics and final accuracy. Calibration experiments further show that architecture and augmentation choices have a larger effect on final accuracy than the training-rule modifications studied here. Cumulative free-riding is therefore a real and repairable optimization pathology. Nonetheless, for the FF training rules, architectures, and datasets we study, it is not the dominant factor limiting achievable accuracy.
Abstract: Text-to-image diffusion models can synthesize high-quality images, yet the outcome is notoriously sensitive to the random seed: different initial seeds often yield large variations in image quality and prompt–image alignment. We revisit this “seed effect” and show that attention dynamics over prompt core tokens, the content-bearing words, measured during the first few denoising steps, strongly predict final generation quality. Building on this observation, we introduce Attention-Based Seed Selection (ABSS), a training-free, plug-and-play method that ranks seeds for a given prompt by leveraging cross-attention to core tokens during the denoising process. ABSS requires no finetuning and does not alter the initial noise; it scores and ranks all candidate seeds, keeps only the top-k for full generation, and discards the rest, without relying on a fixed accept/reject threshold. Operating purely at inference time, ABSS can serve as a lightweight pre-selection add-on for existing seed-optimization pipelines, enabling additional gains. Across three benchmarks, extensive experiments show that ABSS enables consistent improvements in text–image alignment and visual quality for Stable Diffusion variants, as corroborated by human preference and alignment metrics.
Authors: Seongsoo Heo, Dong-Wan Choi
Abstract: Model inversion is a widely adopted technique in data-free learning that reconstructs inputs from a pretrained model through iterative optimization, without access to original data. However, its application to Vision Transformers (ViTs) incurs high computational cost due to expensive self-attention mechanisms. To address this, Sparse Model Inversion (SMI) was proposed to improve efficiency by gradually pruning seemingly unimportant patches, even claiming they are obstacles to knowledge transfer. However, our empirical findings suggest the opposite: even randomly selected patches can eventually acquire transferable knowledge over the inversion process. In fact, we further observe that removing prematurely inverted patches hinders the extraction of class-agnostic features essential for knowledge transfer, as well as class-specific features. In this paper, we propose Patch Rebirth Inversion (PRI), a novel approach that constructs multiple sparse images within a single inversion process by incrementally detaching informative patches, instead of removing unimportant ones. This strategy not only improves efficiency, but also encourages initially less informative patches to gradually accumulate more class-relevant knowledge, a phenomenon we refer to as the Re-Birth effect, thereby effectively balancing class-agnostic and class-specific knowledge. Experimental results show that PRI achieves up to 10× faster inversion than standard Dense Model Inversion (DMI) and 2× faster than SMI, while consistently outperforming SMI in accuracy and matching the performance of DMI.
PaperID: 5486, Poster
Abstract: Online reinforcement learning (RL) has achieved significant success in decision-making tasks, but modern online RL remains highly sensitive to the quality of training experience. Understanding the role of training data in online RL is challenging because the training distribution evolves together with the policy, causing the influence of the training data to propagate across future optimization and future data collection. Existing data-attribution methods for online RL primarily focus on local influence within a single training round, overlooking the non-local data influence across multiple training rounds. In this work, we formalize data attribution in online RL through a trajectory-level leave-one-out quantity, RL-LOO, which measures how one training sample influences a downstream target after subsequent online updates have taken place. We then derive RL-Inf, a first-order estimator that propagates data influence through both optimization effects and policy-induced sampling effects, and provide theoretical guarantees on its approximation accuracy under smoothness assumptions. Our analysis further disentangles these two effects and shows both theoretically and empirically that the full RL-Inf estimator can often be well approximated using the optimization effect alone. Building on this observation, we develop RINSE, a practical data filtering method for online RL. Experiments on standard control benchmarks and an RLHF-style toxicity-mitigation setting show that RINSE improves training efficiency and final performance over standard PPO and local attribution baselines. These results suggest that RL-Inf provides a practical and principled framework for understanding and improving online RL.
PaperID: 5487, Poster
Authors:
Jinhao Zhang, Zhexuan Zhou, Huizhe Li, Yichen Lai, Wenlong Xia, Haoming Song, Youmin Gong, Jie MeiAbstract: Diffusion-based visuomotor policies perform well in robotic manipulation, yet current methods still inherit image-generation-style decoders and multi-step sampling. We revisit this design from a frequency-domain perspective. Robot action trajectories are highly smooth, with most energy concentrated in a few low-frequency discrete cosine transform modes. Under this structure, we show that the error of the optimal denoiser is bounded by the low-frequency subspace dimension and residual high-frequency energy, implying that denoising error saturates after very few reverse steps. This also suggests that action denoising requires a much simpler denoising model than image generation. Motivated by this insight, we propose Hydra-DP3 (\bf HDP3), a pocket-scale 3D diffusion policy with a lightweight Diffusion Mixer decoder that supports two-step DDIM inference. Our synthetic experiments validate the theory and support the sufficiency of two-step denoising. Futhermore, across RoboTwin2.0, Adroit, MetaWorld, and real-world tasks, HDP3 achieves state-of-the-art performance with fewer than 1% of the parameters of prior 3D diffusion-based policies and substantially lower inference latency.
PaperID: 5488, Poster
Abstract: Generalized linear models (GLMs) are fundamental tools for statistical modeling, with maximum likelihood estimation (MLE) serving as the classical approach for parameter inference. While MLE performs well for canonical GLMs, it can become computationally challenging in more general settings with non-canonical, non-smooth, or nonlinear link functions, where the resulting optimization landscape may be ill-conditioned, non-convex, or non-differentiable. In this paper, we study an alternative estimation framework based on variational inequalities (VIs), which formulates GLM estimation through an operator-based equilibrium condition rather than likelihood minimization. We analyze the VI estimator from a statistical perspective and establish finite-sample error bounds and asymptotic normality under mild regularity conditions, together with convergence guarantees for fixed-point and stochastic approximation algorithms. The framework accommodates a broad class of link functions, including non-canonical and non-monotone cases satisfying a strong Minty-type condition, and extends naturally to generalized additive models via basis expansion. Numerical experiments demonstrate that the VI approach achieves competitive finite-sample accuracy and improved numerical stability relative to MLE, particularly in GLMs and GAMs with non-canonical or non-smooth link functions.
PaperID: 5489, Poster
Abstract: Protein-conditioned 3D molecule generation is a central challenge in structure-based drug design, requiring a delicate balance between target compatibility, molecular properties, and physical geometry. While diffusion-based approaches have shown promise, strengthening the conditioning signal can still reduce physical plausibility, and existing models remain sensitive to how pocket information is presented during training and sampling. We propose , a protein-pocket-conditioned variance-exploding (VE) diffusion framework that couples stable coordinate denoising with inference-time conditional control. Specifically, PocketVE combines an EDM-style training and sampling setup for stable 3D denoising, classifier-free guidance for multi-property steering without external property classifiers, and adaptive protein perturbation as a training-time pocket regularizer. Evaluated on CrossDocked2020 under the GenBench3D protocol, PocketVE improves 3D-Validity from 58.6 to 80.6 and reduces strain energy from 457.4 to 127.9 relative to the TAGMol baseline, while improving median Vina Score, QED, and SA under moderate guidance. A guidance-scale study further shows that moderate guidance gives the best balance between target-related objectives and geometric quality, whereas overly strong guidance pushes the sampler toward geometric degradation. Overall, the results show that target-aware molecular generation benefits from treating geometric stability and conditional steerability as coupled design goals.
PaperID: 5490, Poster
Abstract: Visible sensors provide rich scene details but are vulnerable to adverse conditions, whereas infrared sensors capture thermal radiation distributions and are more robust to environmental variations. However, acquiring high-resolution infrared images remains challenging due to sensor limits and hardware costs, and collecting pixel-wise aligned visible-infrared pairs is even more difficult. Existing methods mainly rely on infrared super-resolution or visible-to-infrared (VIS-to-IR) translation: the former preserves thermal distributions but lacks fine details and spatial alignment with visible images, while the latter benefits from fine-grained structural cues but faces an ill-posed thermal-distribution inference problem. To overcome these limitations, we propose CoVisIT, a cross-modal prior guided diffusion model for VIS-to-IR image translation. Our key insight is to constrain VIS-to-IR translation with complementary cross-modal priors, using high-resolution visible images to recover fine-grained scene details and misaligned low-resolution infrared inputs to anchor reliable thermal distributions. To achieve this, we build upon a latent diffusion framework and develop a Cross-Modal Adaptation module that modulates visible and infrared features in a shared feature space, thereby reducing the modality gap and exploiting complementary cross-modal priors. To further address local misalignment, we introduce a Dynamic Kernel Generation Module to predict input-adaptive kernels and embed a Multi-Scale Dynamic Convolution Module into the denoising U-Net, enabling dynamic local feature aggregation for implicit alignment. As a result, CoVisIT generates high-resolution, visible-aligned infrared images with reliable thermal distributions. Extensive experiments demonstrate state-of-the-art performance and consistent improvements in downstream tasks.
Authors: Max Lovig
Abstract: In modern parametric model training, full-batch gradient descent (and its variants) suffers due to progressively stronger biasing towards the exact realization of training data; this drives the systematic ``generalization gap'', where the train error becomes an unreliable proxy for test error. Existing approaches either argue this gap is benign through complex analysis or sacrifice data to a validation set. In contrast, we introduce decoupled descent (DD), a novel theory-based training algorithm that satisfies a train-test identity---enforcing the train error to asymptotically track the test error for stylized Gaussian mixture models. Within this specific regime, leveraging approximate message passing theory, DD iteratively cancels the biases due to data reuse, rigorously demonstrating the feasibility of zero-cost validation and 100% data utilization. Moreover, DD is governed by a low-dimensional state evolution recursion, rendering the dynamics of the algorithm transparent and tractable. We validate DD on XOR classification, yielding superior performance compared to GD; additionally, we implement noisy MNIST and non-linear probing of CIFAR-10, demonstrating that even when our stylized assumptions are relaxed, DD narrows the generalization gap compared to GD.
PaperID: 5492, Poster
Abstract: Classifier-free guidance (CFG) is essential for improving sample quality in visual generation, but it incurs substantial sampling overhead. Existing guidance-free approaches halve this cost by approximating the CFG with a single model inference, but their performance is bounded by the CFG and inherits its limitations. In this paper, we propose self-contrastive guidance (SCG), which contrasts the model's prediction on real and self-generated samples to obtain a direct, model-based signal that pushes generation toward the true data manifold. Integrating SCG into the standard likelihood objective via reparameterization yields Self-Contrastive Alignment for Likelihood Estimation (SCALE), a guidance-free training framework that produces SCG-enhanced predictions in a single forward pass. Across class-conditional and text-to-image generation on diffusion and autoregressive models, SCALE consistently outperforms CFG-mimicking guidance-free baselines and improves further when combined with DPO, demonstrating that the self-contrastive signal yields gains complementary to both CFG and preference-based fine-tuning.
PaperID: 5493, Poster
Authors: Rameshwar Mishra, Modi Akshay, A. Subramanyam
Abstract: Inversion-free methods for text-driven image editing construct direct ODE paths between source and target distributions using pre-trained flow models, but operate on single trajectories vulnerable to noise-induced drift with no mechanism to exploit the geometry of plausible edits. We propose Transport-Guided Flow Editing, which recasts the editing problem in distributional terms: we construct source and target particle clouds in latent space, solve entropic optimal transport between them with an edit-direction-aware cost, and use the resulting barycentric map to define a time-varying correction field for the editing ODE. The transport plan adapts online at each integration step, and the correction strength is governed by a variational co-state derived from the OT anchor, responding to accumulated trajectory deviation rather than following a fixed schedule. The output is a single edited image, but one whose trajectory has been steered by distributional information inaccessible to single-trajectory methods. Extensive experiments demonstrate state-of-the-art structure preservation with superior semantic alignment and consistent human preference over existing baselines.
PaperID: 5494, Poster
Abstract: Autoregressive transformers with no positional embeddings (NoPE) can recover absolute position. A common explanation is that causal attention creates a position-dependent variance signal, but this account is incomplete because LayerNorm removes per-token scale, so a viable mechanism must encode position in , not magnitude. We describe the most compact mechanism with two-layer NoPE transformers that accurately extract position. In Layer1, prefix averaging creates a shared beginning of sequence (BOS)-direction component. In Layer2, the attention's output-value matrix maps the \bos and \nbos contributions into two distinct directions, and a position-dependent \bos attention weight acts as a mixing coefficient that interpolates between them. The resulting directional trajectory is linearly decodable for position. Interestingly, we show two transformer variants that converge to the same theoretical mechanism and instantiate it when trained to predict position.
PaperID: 5495, Poster
Abstract: Identifying manifold structure reveals hidden signals from datasets, removing noise. In this paper, we study the problem of identifying local manifold structure from data. Given a query point y_0 from a dataset, our goal is to recover the geometry of the data manifold in a neighborhood of y_0 using a \emphpretrained optimal transport flow from the reference to the data distribution. First, we prove that the Brenier optimal transport map preserves manifold structure: the preimage of an m-dimensional data manifold is itself an m-dimensional manifold in the reference space. Second, motivated by this result, we propose a latent variable model that maps a linear model through the transport flow. We prove that the linear approximation error is significantly reduced by the optimal transport map, leading to a tight fit of the non-linear data manifold. Third, noting the intractability of the resulting likelihood, we deploy denoising Fisher score estimation --- a recent development from simulation-based inference that learns the Fisher score over parameter--observation pairs --- to perform likelihood-based inference effectively. Experiments on both synthetic and real-world datasets demonstrate the effectiveness of the proposed method.
PaperID: 5496, Poster
Abstract: Vision-Language-Action (VLA) models promise generalist robotic agents, yet consolidating skills learned across heterogeneous tasks, environments, and embodiments remains difficult: training data is fragmented across platforms, joint multi-task training suffers from interference, and model merging tends to introduce parameter conflicts or require architectural changes. We instead advocate a paradigm in which skills are first learned independently and then consolidated post hoc. To this end, we propose : a nested hypernetwork, conditioned on the target task, synthesizes task-adaptive parameters that compose with a frozen base VLA. This formulation enables flexible task-conditioned skill reuse while mitigating interference. Across three RL benchmarks (MetaWorld, ManiSkill, CALVIN), NestedVLA outperforms the strongest model-merging baseline by , and substantially closes the gap on CALVIN, pointing toward a scalable route to skill consolidation.
PaperID: 5497, Poster
Authors:
Weikang Wang, Xin Zhou, Haiyang Liu, Lei Wang, Jun Liu, Weifeng ZhangAbstract: Hybrid large language models (LLMs) integrating state space models (SSMs) and full attention face a critical caching bottleneck due to mismatched caching granularities. While full attention allows flexible, token-level cache reuse, SSMs are always restricted to coarse-grained block- or request-level reuse. Consequently, an SSM layer's cache miss can easily invalidates successful fine-grained hits in adjacent attention layers, disrupting the entire reuse pipeline. To resolve this, chunk-level position-independent cache (PIC) reuse is highly desirable to align the caching granularity across all layers. However, enabling PIC reuse for history-dependent SSMs remains fundamentally challenging. To bridge this gap, we propose HyPIC, the first unified framework achieving PIC reuse for hybrid architectures. We introduce residual state propagation and online decay correction, which efficiently adapt offline SSM chunk caches to new contexts by explicitly recomputing minimal boundary tokens. Furthermore, we design a fast parallel recomputation pipeline exploiting the linear superposition property of SSMs to eliminate sequential bottlenecks. Experiments on diverse hybrid models and datasets demonstrate that HyPIC achieves a 2× speedup in time-to-first-token (TTFT) while maintaining highly competitive accuracy.
PaperID: 5498, Poster
Abstract: Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing fragmentation and structural distortion. To address this, we propose SplitMoE, a split-role sparse architecture that breaks the shackles of uniformity. To accommodate the inherent semantic imbalance, we explicitly bifurcate the expert pool into semantic experts and generic experts, with semantic experts capturing high-level semantic abstraction and generic experts preserving residual visual information and flexible generative capacity. Leveraging prototype-guided routing and push-pull regularization, SplitMoE enables tokens to cluster naturally by semantic attributes rather than arbitrary balancing constraints. Extensive results show that under an equivalent activated-parameter budget, SplitMoE outperforms traditional load-balanced MoEs in convergence speed, routing coherence, and video generation quality across standard benchmarks. By revealing an emergent coarse-to-fine denoising logic, SplitMoE provides the community with a modality-aware scaling path, serving as a critical reference for building large-scale video world models.
PaperID: 5499, Poster
Authors: Xin Zhang, Robby Tan
Abstract: Multi-modal semantic segmentation benefits from complementary sensors, but this benefit is fragile when some modalities are degraded or missing at test time. Many existing methods rely on dedicated fusion modules or multi-branch cross-modal encoders, and the fusion often becomes unreliable when modalities are absent, leading to substantial performance drops. We propose , a parameter-efficient framework that performs fusion within the adaptation space of a frozen vision foundation model (VFM), exploring whether VFM priors can serve as an effective shared backbone for multi-modal segmentation. Instead of attaching a standalone fusion network, FuseAdapt assigns each modality a lightweight adaptation branch and performs adaptation-space fusion by routing trainable adaptation updates across modalities inside the frozen backbone. The pretrained VFM remains fixed while modality-specific corrections and explicit cross-modal interaction are learned through the adaptation paths. To handle missing modalities, we further introduce , a direct representation-level objective for cross-modal redundancy. During training, PMM masks a target modality and predicts its embedding from the visible ones, supervised by a teacher target computed under full-modality input. PMM is used only during training and adds no inference-time overhead. Across three multi-modal segmentation benchmarks (MCubeS, DELIVER, MUSES), FuseAdapt consistently improves both modality-complete and modality-incomplete segmentation with a small trainable budget, yielding up to +3.01 mIoU under full modalities and +12.56 mIoU under missing modalities over prior state-of-the-art.
PaperID: 5500, Poster
Abstract: Cyclic peptides are an increasingly important therapeutic modality, offering antibody-like binding specificity within compact and chemically tunable scaffolds. Recent deep generative models have advanced target-specific peptide design, where SE(3)-equivariant diffusion and flow-matching frameworks provide natural inductive biases for peptide-protein complex generation. However, existing SE(3)-based peptide generators mainly focus on linear peptides and do not explicitly control cyclic topology. As shown by recent cyclic peptide benchmarks such as CPSea, current generative baselines still have limited cyclization success and remain weak at controllably generating disulfide- and isopeptide-cyclized peptide distributions. In this work, we propose , a test-time scaled autoregressive flow-matching framework for receptor-conditioned cyclic peptide sequence-structure co-design. CPFlow autoregressively unmasks residues and predicts full-atom sequence-structure variables with continuous flow matching. By decomposing generation into stepwise flow updates, CPFlow allows each step to balance target binding with structural closure. At inference time, sequence-geometry guided unmasking and multi-particle test-time scaling explore high-confidence autoregressive trajectories beyond the trained model without additional training. Cyclization-aware relative position encodings further enable stable control over head-tail, disulfide, and isopeptide cyclization. Experiments on the CPSea benchmark show that CPFlow outperforms strong generative baselines across
PaperID: 5501, Poster
Abstract: Interactive regex search over large text columns needs exact matches under low latency, but indexes help only when a pattern exposes selective fragments. For LIMIT-k queries, the immediate bottleneck is choosing a few indexable keys that reach match-rich candidate buckets before regex verification. RIVET casts this step as constrained regex-to-key generation: a column-specific translator emits short keys, an existing index retrieves candidates, and the native regex engine verifies every output record. Training from repeated regex-key hits moves probability mass toward effective keys, so finite online sampling spends fewer regex checks before collecting the requested matches. Across four large datasets, RIVET reaches 35.8 ms median latency, improves over the strongest baseline REI by 3.4–36×, and reaches up to 1773× speedup over sequential scan. The same mechanism improves TPC-H query plans, preview retrieval under latency budgets, and 80 human-authored OOD regex queries, showing that regex-to-key generation translates into end-to-end system gains.
PaperID: 5502, Poster
Abstract: Vision-language model (VLM) unlearning is often treated as an objective-design problem: given multimodal forget data, the goal is to optimize a loss that removes undesirable behavior while preserving retained capabilities. We argue that this view overlooks an equally important data-centric question: how should each forget example be represented before optimization? Existing methods typically keep image-conditioned failures in their original image-text form, implicitly assuming that pixels are the right route for unlearning whenever pixels appear in the input. We challenge this assumption by constructing modality-decomposed versions of VLM benchmarks, enabling controlled comparisons among image-only, text-only, and multimodal representations of the same examples. Our modality-mixing experiments show that, for many examples, textual renderings carry the actionable forget signal more directly than the image alone, while purely text-only unlearning can still weaken visual grounding when the target behavior depends on visual evidence. Motivated by this trade-off, we introduce ModRoute, a gradient-guided per-sample modality routing strategy that selectively converts high visual-pressure examples to text-only form while keeping the remaining examples multimodal. Notably, on VLGuard with LLaVA-1.5-7B, ModRoute has the strongest composite unlearning--utility score of 0.847 at \gamma=0.5, a 2.5% improvement over the best random-switching score and \geq46% better than pure multimodal or text unlearning. Overall, our results show that effective VLM unlearning depends not only on the objective, but also on the composition and per-sample modality routing of the unlearning data.
Abstract: Text watermarking for large language models (LLMs) enables model owners to verify the origin and protect the intellectual property of AI-generated text. While text watermarking for closed-source LLMs is relatively mature, open-sourcing the LLM introduces unique challenges: existing methods inject the watermark during the model owner's decoding process, but for open-source LLMs, end users control decoding and can arbitrarily remove it at inference time. The watermark must therefore be embedded into the released model weights, so that standard decoding natively produces watermarked text. Recent methods attempt this by distilling from a decoding-based watermarked LLM, but two challenges remain: (i) learnability: the LLM may not perfectly internalize the predefined watermark pattern, producing a generation-detection inconsistency between what the LLM generates and what the detector expects; and (ii) robustness: open-sourcing the model exposes the embedded watermark to weight-level user modifications such as fine-tuning and model merging. A practical open-source LLM watermark must therefore align the LLM's learned watermark with the detection pattern, and remain robust against weight-level user modifications. Guided by these requirements, we propose PRO, a precise and robust open-source LLM watermarking method. PRO jointly trains a watermark policy model with the LLM so that the LLM's learned watermark aligns with the detection pattern, and further incorporates a perturbation-aware regularizer that improves robustness against weight-level modifications. Evaluated on several mainstream open-source LLMs, PRO improves watermark detectability to 0.92 TPR at 1% FPR while reducing perplexity by 20.5% over the strongest prior method, and preserves detectability under aggressive weight-level modifications, e.g., AUC 0.87 vs. 0.79 under 0.5-ratio model merging. The code is publicly available at https://anonymous.4open.science/r/PRO-44B7.
PaperID: 5504, Poster
Authors:
Qin MA, XIAOQI SUN, bo li, Yuquan Zhou, Weizhong ZhangAbstract: Diffusion Transformers (DiTs) have achieved strong performance in visual generation, but their dense token processing leads to slow training convergence and high computational cost. Recent token routing methods such as TREAD accelerate training by allowing a random subset of tokens to bypass intermediate layers. It can be expected that carefully tuning the routing ratios over time steps or layers can achieve more significant accelerations. However, the theoretical mechanisms underlying stochastic routing remain underexplored, making the tuning of routing ratios intractable. In this paper, we first theoretically show that stochastic token routing can be interpreted as an implicit route-sensitivity regularizer. Under this view, the routing ratio determines the strength of the induced regularization: larger routing ratios introduce stronger route-induced perturbations and therefore stronger regularization pressure. This further motivates a noise-conditioned routing strategy. In diffusion training, samples at larger timesteps contain heavier noise and often provide less stable gradient signals, suggesting that they may require stronger regularization. We therefore assign larger routing ratios to higher-noise samples, allowing the routing-induced regularization strength to adapt to the noise condition. Based on this insight, we propose \emphNoise-Aware Routing, which partitions samples within each mini-batch into low-, middle-, and high-noise regions and assigns progressively larger routing ratios to them. This simple noise-conditioned allocation synergistically bridges computational savings with enhanced model priors. Extensive experiments on class-conditional ImageNet generation show that our method accelerates training convergence and improves generation quality. Noise-Aware Routing achieves up to 15.6× and 20× convergence speedups over DiT baselines—and outperforms TREAD—on ImageNet-256 and ImageNet-512, respectively, measured by the training iterations required to match our 400K-step FID. In the guided setting, our method achieves an FID of \(2.67\) with a substantially shorter training iterations without architectural changes.
PaperID: 5505, Poster
Abstract: The rapid advancement of Large Language Models (LLMs) based on Mixture-of-Experts (MoE) architecture has enhanced a growing demand for secure inference frameworks that protect both client inputs and server model weights. However, existing secure MoE inference frameworks suffer from two limitations: (1) a costly two-stage secure routing pipeline that uses expensive secure Top-K and secure equality test protocols, resulting in significant computational and communication overhead, and (2) secure evaluation of nonlinear activations like SiLU requires several multiplication and comparison operations, resulting in significant communication overhead. In this paper, we present Capricorn, a highly efficient and secure MoE inference framework that overcomes the two limitations above. Firstly, Capricorn utilizes a novel one-stage routing pipeline that reduces both the total number of Top-K protocol invocations and their input data size. At the same time, we propose a lightweight method to generate one-hot token selection matrices locally, thus avoiding expensive secure equality tests. Secondly, we design a secure and precise SiLU protocol, which uses secure lookup tables (LUTs) with dynamic bit-width and precision to reduce communication overhead. Extensive experiments demonstrate that Capricorn achieves up to 50.93 × and 2.10 × speedup in secure routing and secure SiLU protocols, respectively, compared to the state-of-the-art (SOTA) framework CryptoMoE (NeurIPS'25), while maintaining the original model inference accuracy.
Abstract: Text-based counseling is an important interface for AI mental-health support, where transcripts may be used to monitor depression severity and flag sessions requiring timely human review. However, robust PHQ-8 prediction across session regimes remains challenging: fine-tuning-based methods can exploit richer supervision but may generalize poorly under data scarcity, while prompt-based LLM methods are data-efficient but usually treat each transcript holistically and provide limited support for longitudinal context. We study robust depression tracking from counseling transcripts across single-session and multi-session regimes. We introduce , a multi-session counseling dataset with session-level PHQ-8 supervision for evaluating repeated-session tracking under partial symptom disclosure and cross-session continuity. We further propose , a PHQ-8 prediction framework that combines LLM-extracted clinical signals with frozen turn-level semantic embeddings and trains symptom-specific predictors over the resulting transcript representation. When prior sessions are available, EmoTrack can further incorporate them through compact cross-session memory. Experiments on LongCounsel-8 and DAIC-WOZ show that EmoTrack achieves a clear gain on the real single-session benchmark, including a over the strongest DAIC-WOZ baseline, and remains competitive with the strongest longitudinal baseline on LongCounsel-8.
PaperID: 5507, Poster
Abstract: Category-level object pose estimation (COPE), capable of generalizing to intra-class unknown objects, has become a core technique for robotic 3D scene understanding. However, existing COPE methods still require labor-intensive recollection of real-world training data for novel object categories, which limits their scalability in practical applications. This paper aims to achieve synthetic-to-real (Syn2Real) generalized COPE, where a model is trained solely on rendered synthetic data and directly generalized to real-world deployments. The central challenge lies in the significant domain gap between synthetic and real-world data, particularly in texture appearance. To address this, we aim to enhance domain generalization by learning domain-invariant representations that capture semantic commonalities among objects within the same category. We introduce 2D and 3D semantic consistency constraints to reduce the sensitivity of feature encoders to domain-specific features. In addition, we propose an end-to-end pose regression framework that performs 2D-3D cross consistency learning, leveraging dense cross-modality fusion to further refine pose estimation. Since simplicity and effectiveness are essential for real-world robotic deployment, our model operates exclusively on global features, yielding a highly lightweight and efficient architecture. Extensive experiments on the REAL275 and Wild6D benchmarks, as well as real-world robotic manipulation scenes, show superior Syn2Real generalization performance of our paradigm.
Abstract: Multimodal learning has become a prominent research area, with the potential of substantial performance gains by combining information across modalities. At the same time, model development has trended toward increasingly complex deep learning architectures, motivated by the assumption that multimodal-specific methods improve performance. We challenge this assumption through a large-scale empirical study by reimplementing 19 high-impact multimodal methods across nine diverse datasets with up to 23 modalities. Under standardized experimental conditions, including hyperparameter tuning, weight initialization, cross-validation, and statistical testing, increased multimodal complexity often yields confusion rather than effective fusion of data modalities. Accordingly, complex multimodal architectures do not reliably outperform unimodal baselines and a Simple Baseline for Multimodal Learning (SimBaMM). Through a focused case study, we further demonstrate concrete methodological shortcomings even in top-tier multimodal learning publications, underscoring the need for standardized evaluation practices. In summary, we argue for a shift in focus for multimodal learning: away from the pursuit of architectural novelty and toward methodological rigor.
Abstract: LiDAR scene flow estimates point-wise motion between two consecutive scans, referred to as the source and target. Leading self-supervised methods typically minimize the Chamfer loss, the nearest neighbor distance between the flow-compensated source and the target. However, nearest-neighbor search does not enforce motion rigidity, often leading to inconsistent flows within object instances. Existing approaches address this issue with additional regularization terms, but flow consistency among points remains limited, especially for large objects. We propose RVLoss, a self-supervised loss that incorporates motion rigidity by design, through a runoff vote mechanism. Our key observation is that the point-wise motion, calculated from nearest neighbor search, can often be grouped into a small set of dominant flow candidates by voting (top-k voting. Furthermore, when compensating the source by these candidates, the flow that best represents the underlying rigid motion often yields the highest consensus after a second voting (top-1 voting). Based on this insight, we incorporate the two-stage runoff vote into loss design and create cluster-wise rigid flows and subsequent free-form flows as pseudo-labels for self-supervised learning. RVLoss can be seamlessly integrated into existing feedforward architectures. Experiments on the Argoverse2 2026 Challenge show that models trained with RVLoss achieve state-of-the-art performance among self-supervised approaches, outperforming baseline models trained with alternative loss designs by 20%. Moreover, cross-dataset evaluations demonstrate consistent performance improvements across four additional datasets. Code will be released upon acceptance.
PaperID: 5510, Poster
Authors: Zachary Friedenberger, Yiwei Cao, Richard Naud
Abstract: Population analysis methods have become standard for navigating the complexity of neural data. However, these methods often assume a rate code, neglecting information encoded in the precise timing of spikes. Critically, additional information encoded in bursts of action potentials may be missed. Here, we develop a factor analysis method that disentangles the factors associated with bursts and individual spikes. This enables burst codes to be investigated directly from the structure of the data, without requiring external covariates. We demonstrate that analyzing firing rates alone obscures the latent structure and factors underlying bursts. Applying our method to simulated and experimental data, we show that it can infer the correct latent structure and be used to test for the presence of burst coding. By merging the population and burst coding perspectives, we provide a framework for linking changes in bursting to internal variables involved in attention, perception, and learning.
PaperID: 5511, Poster
Authors:
Dong Uk Kim, Ji S Yoon, Eui-Nam Huh, Choong HongAbstract: Alignment-based representation learning for cross-dynamics RL admits a degenerate global minimizer: per-domain feature covariances collapse to a common point, driving the alignment loss to zero while transfer regret remains large and rendering the standard Bures-Wasserstein (BW) transfer regret bound vacuous. We document this collapse across multiple environments, alignment kernels (BW, Frobenius, MMD), and methods (BW-CORAL, VGDF), where per-domain effective rank drops from p to \approx 1 within a few epochs—structurally the same failure as representation collapse in self-supervised learning. We propose a minimal remedy: a per-domain Prototype Trust Region (TR) that anchors each empirical covariance to a slow-moving EMA prototype, introduces no learnable parameters, and acts as an explicit second-moment stabilizer, mirroring momentum targets and covariance regularizers in SSL. Because TR depends only on per-domain covariances, it is combined with any BW-based method; adding it to VGDF inherits the same resistance to collapse. Theoretically, we prove a BW non-collapse lower bound that restores the transfer-regret guaranty to a non-vacuous form, and present a Representation Stability Transfer Bound that decomposes regret into coverage and stability terms TR directly controls. Empirically, TR delivers statistically significant gains under pre-registered tests in collapse-dominated regimes (finite-sample, lower-dimensional) and is correctly neutral when other bottlenecks dominate. We frame TR not as a universal improvement but as a targeted fix for an identifiable and broadly shared failure mode.
Abstract: Reinforcement learning with multinomial logistic (MNL) function approximation has become an important framework due to its flexibility and broad applicability. While existing studies have established regret guarantees under worst-case analysis, they do not capture how performance depends on the variability of the interaction between the learner and the environment. In this paper, we develop a new theoretical analysis for MNL-based Markov decision processes that yields explicit variance-adaptive regret bounds. Our algorithm is computationally efficient and achieves the instance-wise optimal rate of regret, narrowing the gap between upper and lower bounds. Our numerical experiments validate that our method learns optimal policies more efficiently than conventional approaches.
PaperID: 5513, Poster
Abstract: The Platonic Representation Hypothesis (PRH) claims that independently trained models converge on a shared statistical model of reality, yet recent work finds only weak pointwise similarity between models. In this paper, we show that what models share is not the location of samples in representation space, but the directions (displacement vectors) between them. Under a single orthogonal alignment---rotation and reflection only---these displacement vectors are substantially preserved across 44 independently trained vision and language encoders spanning modalities and asymmetric capability pairs, consistent with the PRH evidence. The samples' absolute positions are not, consistent with recent counter-evidence. Both arise from a single decomposition: representations split into a that is not. We trace this geometry to concept-level structure: within a model, parent concepts are orthogonal to their child variation vectors; across models, concept displacements are parallel. Our theory falsifiably predicts (and experiments confirm) that fine-tuning preserves pointwise similarity but collapses displacement, and that relational distillation does the opposite. A major implication is that, because semantics align linearly but capabilities do not, capabilities can be imported from one model to another using a single cached forward pass through the source. We call this Shadow Casting. As a proof of concept, our ShadowCLIP instantiation matches or outperforms strong fine-tuned baselines at orders of magnitude less compute. A cache can be released alongside open model weights, letting one model's capabilities be downloaded and imported into any number of other models without fine-tuning.
PaperID: 5514, Poster
Abstract: LVLM benchmarks tell us whether a model answers correctly, but not which visual conventions it has retained, nor where those conventions live. We use sketches to expose this hidden structure. A sketch removes texture, lighting, and photographic context while preserving the strokes humans judged sufficient for recognition; if an LVLM can operate on that abstraction, the surviving signal is structural rather than merely photometric. We introduce a training-free sketch probe with three readouts: temporal part-construction order, spatial attention-based grounding, and the mechanistic layer/head locus of grounding. We evaluate across nine LVLMs and five LLaVA-architecture variants on two part-annotated sketch datasets. Three findings emerge. First, LVLM part order tracks human convention: models reproduce canonical orderings where human drawings converge, and become correspondingly variable where human orderings vary. Second, selected frozen LVLM attention heads ground object and scene sketches competitively with specialised grounders, exceeding CLIPSeg and GroupViT on mAcc@0.1 despite using no grounding supervision. Third, the grounding signal concentrates in a narrow mid-network band whose location follows the base language model rather than parameter count, and remains stable across object and scene sketches. The result is a practical diagnostic for sketch-driven LVLMs: no fine-tuning, no auxiliary segmenter, and a small architecture-dependent layer band that doubles as a head-selection prior, improving grounding mAcc@0.1 in every model and dataset we tested.
Authors:
Valentina Njaradi, Clémentine Dominé, Rachel A Swanson, Marco Mondelli, Andrew SaxeAbstract: Learning to generalise from limited data is a fundamental challenge for both artificial and biological systems. A common strategy is to extract reusable structure from abundant unlabelled data, enabling efficient adaptation to new tasks from limited labelled data. This two-stage paradigm is now standard in modern training pipelines, where pretraining is followed by fine-tuning or linear probing. We provide an analytical model of this process: structure extraction is formalized as principal component analysis on unlabelled data, and downstream learning as linear regression on a separate labelled dataset. In the high-dimensional regime, we derive exact expressions for training and generalisation error showcasing their dependence on representation dimensionality, unlabelled and labelled sample sizes, and task alignment. Our results show that pretrained representations strongly influence downstream generalisation, and we characterize the optimal representation size as a function of task parameters: with abundant pretraining data but scarce downstream data, maximally compressed representations are optimal, whereas with limited pretraining data, higher-dimensional representations generalise better. Furthermore, we establish an exact trade-off between pretraining and supervision, quantifying how much unlabelled data is required to replace a single labelled sample. Beyond our idealised model, we observe similar phenomenology in autoencoders and pretrained LLMs. Altogether, we highlight that optimising representation size is critical, giving conditions for when compression during pretraining improves generalisation.
PaperID: 5516, Poster
Authors:
Wei Fan, JinYi Yoon, Bo JiAbstract: Multi-Agent Debate (MAD) improves reasoning on multimodal tasks, but its accuracy and token efficiency depend on how agents interact during debate. A largely overlooked design choice is which interaction modalities agents use to exchange evidence. Existing methods rely on fixed text- or graph-based interaction, which fails to convey fine-grained multimodal evidence, degrading accuracy while inflating token cost due to verbose descriptions. To address this, we propose Adaptive Interaction in Multi-Agent Debate (AIM), an adaptive framework that dynamically selects the best combination of interaction modalities for each instance. Beyond text and graph, AIM introduces task-dependent modality interaction, where agents exchange targeted regions of interest from the task-specific multimodal input, such as image regions, audio segments, and spatio-temporal clips. To determine which combination is the best for each instance, AIM first generates a structured baseline response, extracts interpretable routing features, and finally uses a lightweight router to select a modality combination that avoids both under-fusion (i.e., agents miss decisive evidence) and over-fusion (i.e., redundant modalities inflate token cost) risks. Across eight multimodal question answering (QA) benchmarks of Visual-QA, Audio-QA, and Video-QA, AIM achieves the highest accuracy on all datasets, improving upon the best single-modality baseline by up to 8.9% while reducing token cost by up to 53.4% relative to the best fixed-modality baseline.
Abstract: Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic discrepancy in model predictions under such cross-modal variations. To study this phenomenon, we define the modality gap as the difference in model performance under semantically equivalent textual and multimodal inputs. We introduce TokenSwap, a method that constructs such inputs by replacing textual concepts with semantically aligned images, resulting in sequences where visual tokens are interleaved with text tokens. Based on TokenSwap, we transform existing text-based benchmarks (e.g., MMLU) into image-interleaved counterparts, resulting in TOKENSWAP-BENCH. Across 42 MLLMs, we observe a pervasive modality gap, with performance decreasing by 4.2% to 47.4% when moving from text-only to image-interleaved inputs, averaging 19.6% ± 3.3% across models. Notably, we observe that reasoning models exhibit consistently smaller gaps, achieving an average gap of 10.1% compared to 25.5% for non-reasoning models. In contrast, neither prompting strategies nor scaling training compute alone reliably reduces the modality gap. Finally, we demonstrate that incorporating TokenSwap during training effectively mitigates this gap while preserving strong text-only and vision-language performance.
PaperID: 5518, Poster
Authors: Rui Xu, yinghui xu, Libo Wu
Abstract: Controlling persona in large language models (LLMs) at inference time is important for role-playing, personalized dialogue, and social simulation. Recent methods extract persona vectors from the model's activation space and apply Euclidean operations—addition, scaling, and linear interpolation—under the linear representation hypothesis. However, these methods themselves report systematic failures: non-orthogonal trait dimensions, asymmetric ceiling and resistance effects, and significant deviations in multi-trait composition, suggesting that the linear isotropic assumption does not hold. We propose PersonaManifold, a framework that models persona representations as points on a curved, low-dimensional Riemannian submanifold in activation space. We estimate the manifold's intrinsic geometry—local metric tensors, geodesic distances, and Ollivier–Ricci curvature—and introduce geodesic steering, which interpolates between personas along manifold geodesics rather than Euclidean straight lines. We also propose the Behavioral Similarity Triplet (BST) benchmark, which automatically generates situational questions grounded in six established psychological constructs and defines persona similarity through behavioral responses rather than self-report questionnaires. Experiments on three open-source LLMs show that persona activations form a manifold with heterogeneous curvature, geodesic distance predicts behavioral similarity more accurately than Euclidean alternatives with independent contributions from anisotropy and curvature, and geodesic steering produces more coherent intermediate personas on both our BST benchmark and external evaluations, with the advantage concentrated in high-deviation regions where the manifold deviates most from flatness.
Authors:
alireza abdollahpour, Ehsan Sharifian, Buse Şen, Marco Cuturi, Daniel KuhnAbstract: Distributionally robust optimization (DRO) provides a principled framework for learning under distribution shift, but its practical use is hindered by the difficulty of evaluating worst-case risks for nonconvex loss functions. We study a penalized DRO formulation in which the adversary may choose any distribution but incurs a Wasserstein penalty for deviating from the empirical distribution. We show that the adversary’s problem can be reformulated as an optimization problem over transport maps that push empirical samples to adversarial ones, and we prove that optimal maps are cyclically monotone. We also show that standard adversarial training---based on per-sample local optimization---violates cyclical monotonicity and wastes transport costs unless the adversary is severely restricted. We propose two remedies. First, we introduce multi-start particle ascent, which alternates parallel gradient ascent with reassignment to enforce cyclical monotonicity across samples. Second, we parameterize adversarial maps as gradients of input-convex neural networks, which guarantees cyclical monotonicity by construction. Experiments on robust regression, image classification, and robust control show that our methods consistently outperform standard adversarial training and state-of-the-art baselines, achieving improved robustness and better generalization under distribution shift.
Abstract: Large language models interact with users through a simulated "Assistant" persona. While the Assistant is typically trained to be helpful, harmless, and honest, it sometimes deviates from these ideals. In this paper, we identify directions in the model's activation space - persona vectors - underlying several traits, such as evil, sycophancy, and propensity to hallucinate. We confirm that these vectors can be used to monitor fluctuations in the Assistant's personality at deployment time. We then apply persona vectors to predict and control personality shifts that occur during training. We find that both intended and unintended personality changes after finetuning are strongly correlated with shifts along the relevant persona vectors. A central contribution of our work is preventative steering, a novel training-time intervention that actively suppresses the acquisition of unwanted traits. By steering activations during the fine-tuning process itself, this method “immunizes” the model against negative personality shifts.Moreover, persona vectors can be used to flag training data that will produce undesirable personality changes, both at the dataset level and the individual sample level. Our method for extracting persona vectors is automated and can be applied to any personality trait of interest, given only a natural-language description.
PaperID: 5521, Poster
Abstract: Large language model systems repeatedly make test-time decisions by comparing similarity. They retrieve exemplars, rerank candidates, and choose revisions. But similarity is not unique: a candidate can be close lexically, semantically, structurally, or stylistically, while only one of these axes may matter for the current instance. We propose BASIS~(Bayesian Axis Selection via Interactive Signals), a training-free Bayesian framework that treats weak feedback as evidence about this latent task-aligned axis. BASIS maintains a posterior over candidate targets and similarity axes. It updates this posterior from scalar scores, pairwise preferences, or verifier outputs, and then chooses future probes by expected information gain. In a finite noisy similarity model, we show that fixed-axis and fixed-mixture rules can remain suboptimal under axis shift. Across concept recovery, reasoning exemplar selection, response refinement, and adversarial distractors, BASIS consistently outperforms stronger rerankers, static mixtures, and compute-matched best-of-N search under the same candidate pools and feedback budgets. The average gains are moderate because all methods share the same backbone, candidate pools, and feedback signals. Under axis shift, however, the gap widens substantially, showing that latent-axis inference is not fully replaced by better scoring or more search alone.
Abstract: We consider the active learning problem where the goal is to learn an unknown function with low error under a Boltzmann distribution induced by the function itself. This self-induced weighting arises naturally in problems such as molecular energy modeling and free energy estimation, yet poses unique challenges as the target distribution is unknown and its partition function is intractable. We propose an acquisition function, AB-SID-iVAR, based on Gaussian Process surrogates that approximates the intractable Bayesian target distribution in closed form while avoiding partition function estimation, applicable to both discrete and continuous input domains. We also analyze a Thompson sampling alternative (TS-SID-iVAR) as a higher variance Monte Carlo variant. Despite the unknown target, under mild conditions, we establish that the terminal prediction error vanishes with high probability, and provide a tighter average-case guarantee. We demonstrate superior performance over existing active learning approaches on synthetic benchmarks and real-world modeling tasks.
PaperID: 5523, Poster
Abstract: Quantifying collective fish behavior requires accurate trajectories, yet multi-view 3D tracking remains challenging due to frequent occlusions, visually similar individuals, and the long-standing scarcity of identity annotations. We present TrackFish3D, a geometry-driven self-supervised framework for dense multi-camera 3D tracking of schooling fish. Instead of relying on appearance-based re-identification or manually annotated identities, TrackFish3D turns calibrated multi-view geometry into supervision: triangulation and reprojection consistency provide pseudo-associations, while a geometric encoder and global association transformer learn all-to-all cross-view correspondence within each frame. To make these associations identity-aware, TrackFish3D introduces an effective self-supervised contrastive objective that separates co-visible individuals in the embedding space, together with a temporal predictor that preserves identities and bridges short occlusions across frames. The resulting model is trained once on unlabeled footage and applied directly to unseen test videos, requiring no cross-view identity labels, temporal annotations, 3D ground truth, appearance features, or test-time optimization. On our benchmark, TrackFish3D improves 3D Multi-Object Tracking Accuracy from 83.7% for the strongest baseline to 96.6%. On the 3D-ZeF zebrafish benchmark, it achieves 81.1% MOTA, compared with 77.4% for the best geometric baseline. TrackFish3D also generalizes beyond fish, achieving strong results on real-world bird tracking. We will release our code and model upon acceptance.
Abstract: Identity-preserving video generation is fundamentally constrained by the prevailing point-reference paradigm: existing methods consume one or a few snapshot images and therefore cannot transcend "look-alike" fidelity, suffer from severe drift under large yaw, occlusion or scale variation, and trade brittle cross-pair supervision against pervasive reference leakage. We argue that identity is intrinsically a spatially-distributed, multi-view concept and propose Argus, a theoretically-grounded framework that lifts identity injection from single-snapshot conditioning to multi-view aggregated reasoning. Argus comprises three tightly-coupled modules. (i) An Identity Director built on a frozen multimodal LLM automatically curates nine semantically-diverse keyframes, extracts a global identity descriptor, and rewrites the prompt to dissolve the prompt-frame-identity tri-conflict. (ii) A Stacked Mosaic Identity Injection packs the nine SMPL-localized crops into a 3×3 mosaic and feeds it through Wan 2.1's native 36-channel patch embedding via a flow-matching-synchronized prefix-token stream, while a Negative-Time-Anchored RoPE preserves architectural compatibility with Wan's split-axis design. (iii) An Adaptive Skip-Layer Guidance replaces hand-tuned PAG with a tiny mask network that learns where to perturb, decoupling identity guidance from textual guidance, and a Temporal Identity Annealing schedule injects identity only during the structure-formation window to obviate copy-paste artifacts. Trained without any cross-pair data, Argus attains state-of-the-art performance on OpenS2V-Eval, surpassing both closed-source (Hailuo) and open-source (VACE-14B, Phantom-14B, Stand-In, BindWeave) systems on five of seven metrics. To stress-test extreme failure modes, we additionally release HardID-Celeb, a celebrity benchmark with two new robustness probes (YawScore, OccScore) on which Argus establishes the highest scores by a wide margin. A 47-participant double-blind user study confirms perceptual superiority across both expert and lay populations.
Abstract: Adaptive test-time compute for LLM agents aims to invoke extra computation only when it improves performance. Existing methods typically use confidence-, uncertainty-, or difficulty-based gates, assuming a fixed direction from the gating signal through compute need to the value of computation. This makes gating a utility-calibration problem: gating signals should align with whether extra computation improves the final outcome over the base policy. We show that this alignment is unstable: the same signal predicts rollout benefit in one setting and rollout harm in another, with reversals across environments and backbones even when the task is fixed. Wrong-direction gates can therefore worsen performance by precisely selecting harmful states. This reversal reflects a deeper distinction between compute need and compute suitability: a high uncertainty signal may indicate decision-difficult states where rollouts help compare alternatives, or intervention-unsuitable states where the current context does not support useful rollout-based improvement. Under this two-source model, fixed-direction gates are unreliable across heterogeneous settings. To address this, we propose DIAL ( earning), a sparse gate trained from signal-agnostic counterfactual exploration to learn the utility direction of state features per (environment, backbone). Across six environments and three backbones, DIAL yields a stronger overall success-cost trade-off than fixed-direction baselines. Our code is open-sourced at
PaperID: 5526, Poster
Abstract: As learning agents are increasingly deployed in strategic environments, designers must choose whether to deploy transparent (auditable, committing) or opaque (flexible, reactive) policies. We formalize this through \emphPolicy Transparency Games (PTGs): a two-stage model in which each agent first chooses transparency, then plays an underlying normal-form game. Our central theoretical contribution is a structural reduction: for two-player PTGs, all transparency incentives decompose into two per-player quantities, a leader wedge L_i-N_i and a follower wedge F_i-N_i, whose signs characterize every equilibrium and cleanly separate selection-robust from selection-dependent conclusions. We then ask how a learning agent can discover this structure from interaction data. External-regret learners over the binary transparency action reach coarse correlated equilibrium, and our adaptive WEDGE-CERTIFY algorithm attains instance-dependent certificate complexity \tilde O(\sigma^2(d_0^-2+d_10^-2+d_01^-2)) with matching information-theoretic lower bounds. The wedge representation thus determines the learnability of endogenous algorithmic transparency.
Abstract: The mechanisms behind LLMs' broad generalization beyond training examples are poorly understood. Emergent misalignment (EM) offers a striking case study: finetuning on narrow tasks induce broad misalignment to semantically-unrelated test domains. In this work, we propose the Piggyback Hypothesis: the chat-template prefix, shared across all user queries, can piggyback the finetuned behavior onto out-of-domain queries. We validate this hypothesis by showing that subtle perturbations to the prefix, or simply patching the prefix representations with those from the unfinetuned model, can restore alignment without changing the user query. Building on this finding, we propose Token-Regularized Finetuning (TReFT), which directly regularizes prefix representations during training to mitigate piggybacking. Across Llama-3.1, Qwen-2.5, and GPT-OSS models and multiple EM-inducing datasets, TReFT reduces EM while preserving in-domain learning. On Llama-3.1-8B finetuned on the legal domain, TReFT achieves 33.5% more EM reduction than data interleaving with a retain set of aligned examples. We further show that TReFT extends to other narrow-finetuning settings, including abstention, tool use, and refusal (off-topic generalization is reduced by 54.3% on average), which further supports the Piggyback Hypothesis. Broadly, our work highlights that LLMs may learn and generalize in unintended ways and suggests a path toward more controllable finetuning. It also calls for further study of how shared input features can piggyback model behavior across domains.
PaperID: 5528, Poster
Authors: Shiheng Zhang
Abstract: Goal-conditioned planning requires compressing value functions into low-dimensional representations, yet which property of the compression best predicts planning quality remains unclear. We show that reconstruction accuracy and ordinal fidelity---how faithfully the ranking of successor states by value is preserved---are two projections of the same approximation error: in smooth geometry they couple tightly and L^2 is a sufficient summary; when topology creates dense small-gap regions they decouple, and the ordinal projection carries complementary planning-relevant information beyond L^2. Through experiments on gridworld, continuous navigation (640 runs, 8 environments), and Maze2D (240 runs, 3 D4RL layouts), we establish three results. First, neighbor-restricted ordinal fidelity (\tau^\mathrmnbr) adds \Delta R^2 = 0.106 of incremental explanatory power beyond L^2 after environment fixed effects. Second, monotone recalibration degrades L^2 by up to ~\!400× while preserving \tau^\mathrmnbr and planning success. Third, in matched-pair comparisons controlling for L^2, higher \tau^\mathrmnbr predicts better planning 76% of the time. We provide theoretical grounding via a gap-weighted Bellman-excess decomposition that links reconstruction error, local ordinal inversions, and planning regret.
PaperID: 5529, Poster
Abstract: Diffusion models have achieved remarkable success in image, video, and 3D content creation. However, extending these successes to native 4D model generation remains challenging due to the intricate coupling between temporal geometry and appearance dynamics. Existing approaches typically generate 4D content by synthesizing videos from novel viewpoints, yet they often suffer from poor spatio-temporal consistency across views and time. To address these limitations, we introduce TESLA, a novel feedforward approach for generating native 4D Gaussian Splatting (4DGS) representations. Our key insight is to decouple temporal geometry and appearance generation through a temporally structured latent representation. This decomposition allows us to first establish coarse temporal geometry that captures fundamental structural dynamics, and then synthesize fine-grained appearance details conditioned on this geometric scaffold. To bridge this decoupled representation to the final output, we design a native 4DGS decoder that directly transforms our temporal features into native 4DGS representations, enabling flexible temporal rendering across arbitrary viewpoints and timesteps. Trained on a large collection of curated animated sequences, we extensively evaluate TESLA on the 4DGS generation task, showing superior performance over existing baselines.
PaperID: 5530, Poster
Authors:
Chenguo Lin, Yu Tang, Weiqiao Zheng, Enhua Jiang, Zhiguang Liu, Maomao Li, Muxi Chen, Xiaojia Chen, Shisheng Huang, Jianhuan Zhuo, Zilin Yang, Jiarong Ou, Rui Chen, Yadong MuAbstract: Video world models aim to simulate environmental dynamics under interaction. However, most existing approaches entangle transition reasoning with pixel-level generation within a single model, thereby limiting narrative continuity and long-horizon coherence. We reformulate video world modeling as an agentic process that synergizes reasoning and generation, and present ReGen, which couples a streaming vision-language planner with a causal diffusion generator. At each step, the planner infers a structured action-state specification from history, while the generator realizes it as the next video segment. The planner further verifies the outcome and decides to continue, end, or re-generate, forming a closed-loop sense–plan–act–verify pipeline. To support this formulation, we carefully construct a large-scale action-grounded video dataset with over 2M segments that instantiates the planner–generator interface, and introduce action-grounded reward alignment that turns the action specification into an enforced reward and mitigates cross-segment drift. Experiments show that ReGen improves long-horizon narrative coherence over prior work, supporting autonomous long-horizon rollouts without step-by-step user intervention.
PaperID: 5531, Poster
Abstract: Video matting is essential for high-precision foreground extraction, yet its development is constrained by the scarcity of diverse annotated training data. To address this, we propose Static-to-Dynamic (S2D), a novel generative paradigm that scales video matting data by animating still labeled images into high-fidelity video-matte pairs. Starting from a static image-matte pair, S2D leverages an image-to-video generative model to synthesize realistic motion and scene dynamics, while simultaneously producing pseudo alpha labels from generative priors. Specifically, we introduce Latent Matte Propagation (LMP), which propagates the first-frame matte to subsequent generated frames via the attention mechanisms in the diffusion transformer. We further design a Coarse-to-Fine Matte Decoder to recover pixel-level alpha details from compressed latent representations using multi-scale generative features. Based on this paradigm, we build S2D-VM, a large-scale dataset containing 14K high-fidelity paired video clips. We also propose a Self-Corrected Progressive Training strategy to mitigate pseudo-label noise during optimization. Experiments show that S2D effectively scales video matting data and consistently improves the performance of existing models across multiple benchmarks.
PaperID: 5532, Poster
Abstract: Unbalanced optimal transport extends classical optimal transport to measures with different total masses, but statistical guarantees for Monge-type estimation remain limited. We study unbalanced transport with quadratic cost and Kullback-Leibler marginal penalties and argue that the natural population target is not a map alone, but a transport-growth pair. Consequently, we develop two estimators for the transport-growth pairs under several setups: an optimal transport plan-based estimator for a general case, and a kernel-based estimator for a case with smooth densities. We also show that an error of the estimator achieves the minimax optimal rate by deriving a matching lower bound of the minimax risk. Our main technical contribution is a value-based stability reduction that converts perturbations of the UOT objective into transport and growth risks through a UOT gap condition. These results provide a statistical foundation for Monge-type estimation in unbalanced optimal transport.
PaperID: 5533, Poster
Abstract: Rare disease diagnosis is a fundamental challenge due to heterogeneous and overlapping clinical phenotypes. While recent LLM-based agentic systems have demonstrated promise by integrating external medical tools and knowledge, they typically rely on large-scale or commercial LLMs, limiting their practical deployment in resource-constrained settings. In this work, we observe that their performance significantly degrades when using small-scale LLMs, due to entangled workflows and unreliable multi-evidence aggregation. To address this issue, we propose RADAR, a rare disease diagnostic framework designed for small-scale LLMs. RADAR adopts a divide-and-conquer paradigm that decouples heterogeneous diagnostic evidence into specialized workflows, reducing reasoning complexity and improving robustness. It further introduces a knowledge-driven self-evaluation mechanism that performs evidence-aware reasoning over structured disease knowledge to produce interpretable reliability scores by identifying key supporting phenotypes and conflicting evidence. Finally, a multi-evidence fusion strategy integrates outputs from multiple sources based on reliability and clinical agreement, resolving conflicts and producing stable diagnostic rankings. Experiments on three benchmarks show that RADAR consistently outperforms various state-of-the-art baselines, including general and medical LLMs and agentic systems. Code will be released.
PaperID: 5534, Poster
Abstract: With respect to improving the reasoning accuracy of LLMs, the representative reinforcement learning (RL) method — Group Relative Policy Optimization (GRPO) — has achieved critical success, yet it still suffers from the issue of insignificant reward signals. This paper introduces ReST-RL, a unified Reinforced Self-Training (ReST) LLM RL paradigm that reconnects policy optimization and value-guided search to improve LLM reasoning ability. Firstly, ReST-GRPO adopts an optimized ReST-style algorithm to reshape the policy-induced trajectory distribution by increasing the reward variance of GRPO sampling and exposing the policy to more informative partial states, thereby improving training efficiency and effectiveness. Then, we further introduce a decoding optimization method, VM-MCTS, which trains a Value Model (VM) from self-collected Monte-Carlo Tree Search (MCTS) targets and deploys it through an adapted MCTS algorithm to provide precise process signals and verification scores, further enhancing reasoning accuracy. These two stages are internally dependent — ReST-GRPO yields higher-quality trajectories for value learning with VM-MCTS, which in turn enables more effective inference-time search. We validate our RL paradigm on multiple coding benchmarks (e.g., APPS, BigCodeBench, and HumanEval), where it significantly outperforms other reinforcement training baselines (naive GRPO and ReST-DPO), as well as decoding and verification baselines (e.g., PRM-BoN and ORM-MCTS), indicating its power to strengthen LLM reasoning capability. Moreover, we further evaluate ReST-RL on out-of-domain math and science reasoning tasks, where it achieves superior performance and favorable end-to-end efficiency trade-offs, verifying its effective, robust, and highly generalizable nature.
PaperID: 5535, Poster
Abstract: Automated multi-agent systems offer clear advantages over manually designed ones in scalability and adaptability, but existing workflow topology methods still face important limitations. Search-based methods are often computationally expensive, textual-gradient-based methods rely on coarse-grained feedback, and existing generation-based methods are not well suited to discrete workflow topologies with complex dependencies. To address these limitations, we propose FlowMAS, a multi-agent workflow topology method based on Generative Flow Networks (GFlowNets). FlowMAS models workflow generation as reward-guided flow over the topology space and introduces three components: a GFlowNet-based topology generation backbone, a curiosity-driven module for structure-aware exploration, and an information-guided optimization module for evaluating intermediate topologies. Concretely, the curiosity-driven module encourages exploration of structurally novel workflows, while the information-guided module measures both the information contribution and the communication efficiency of different operators to favor more informative and effective collaboration patterns. Experiments on six benchmark datasets with three LLM backbones show that FlowMAS consistently outperforms multiple baselines.
PaperID: 5536, Poster
Authors:
Yanyu Zhu, Hoilam Pao, Zhaoyu Hu, Yufei zhang, Jiajun Chai, Guojun Yin, Wei Lin, Dongnian Wang, Hai-Tao Zheng, Hong-Gee Kim, Shaoxiong ZhanAbstract: Reinforcement learning with verifiable rewards (RLVR) has become the dominant paradigm for post-training large language models (LLMs) on complex reasoning and multi-step agentic tasks. However, the standard scalar reward signal is sparse and uninformative: it conveys whether a trajectory succeeded but not which reasoning steps failed or how to correct them, severely limiting sample efficiency on long-horizon tasks where successful trajectories are rare. We propose CARD (Critique-Augmented Reinforcement Distillation), a training framework that amplifies the learning signal from a \textttverify_answer tool within the GRPO objective via two complementary signals. First, a token-level self-distillation reweighting measures how much each token's generation was shaped by the expert critique: tokens that merely echo the critique text are masked, while critique-driven revision tokens are amplified. Second, a rejection-sampling distillation loss maximizes the likelihood of critique-free trajectories for responses where the critique provably improved the answer, directly accelerating the internalization of critique-guided behavior. Together, these two signals distill critique-augmented reasoning into the model's base policy without requiring additional rollouts or separate reward models. Experiments on four long-form deep search benchmarks and three multi-hop question-answering benchmarks demonstrate consistent improvements over the GRPO baseline. On the deep search benchmarks, CARD improves average performance by 40% over the Qwen3-8B base model and by 6% over GRPO; with Qwen3-14B, the gains are 46% and 7%, respectively. These results establish expert critique as an effective and scalable source of rich text supervision for reinforcement learning of long-horizon research agents. Furthermore, we show that \textttverify_answer can serve as a test-time self-revision scaffold: at inference, the tool returns only task rubric criteria---without expert feedback---prompting the model to reflect on its draft and self-revise, yielding consistent performance gains over the no-VA baseline.
PaperID: 5537, Poster
Authors: Bob Li, Chloe Cao, Constance Liu
Abstract: Partitioned routed distillation is hard to interpret when route count, supervision grouping, and search or filtering budget change together. We study that attribution problem inside one fixed routed family using two matched nulls: a routing-only control (Latent-Routing MoA) and a structure-matched shuffled-partition control (Matched-Partition-Shuffle), with teacher tokens, trainable parameters, and router depth matched exactly. The decisive result is Table tab:main: the zero-manual deployable row ModeDistill-Auto-v4-Open, whose partition source is induced by GPT-OSS-120B-Instruct pairwise same-strategy judgments (Apache- 2.0 , distinct from the trace-generating teacher), beats MoA by +2.46 pp and Matched-Partition-Shuffle by +2.04 pp on the five-domain primary-shift average (A1), with positive OOD gaps on all five primary domains. Table tab:external transfer is intentionally weaker evidence: for that exact same headline row it reports only four external benchmark aggregate means, where Auto-v4-Open beats MoA by +4.91 / +3.99 / +3.70 / +2.87 pp, and it does not support slice-wise breadth or cross-family generality. The open pipeline remains within 0.05 pp of the manual upper-bound reference on A1 and within the pre-specified 0.50 pp tie band on every external mean, so we keep it as the deployable row and retain the manual row only as calibration. The claim is therefore narrow: within this routed family, once routed budget is matched, the residual OOD gain is attributable to the externally induced partition source rather than to routing alone or to structure-matched shuffled partitions. Figure fig:identification and the appendix are corroborative only. The weakest primary shift remains ALFWorld tool reconfiguration ( +1.15 pp vs.\ MoA, paired CI [-0.12,+2.41] ), and discovery-time verifier calls and teacher generation are disclosed separately from the matched comparison.
PaperID: 5538, Poster
Abstract: Machine learning interatomic potentials (MLIPs) enable efficient and accurate atomistic simulations but depend critically on the quality and diversity of the training data. We introduce Stein kernelized molecular dynamics (SKMD), an enhanced sampling method that uses interacting particle dynamics to acquire informative training configurations for the active learning and fine-tuning of MLIPs. SKMD corresponds to a stochastic variant of Stein variational gradient descent that is adapted for molecular dynamics by incorporating asynchronous particle updates and a kernel of global atomic descriptors, which provides a symmetry-aware measure of configurational similarity. Unlike other enhanced samplers used in molecular dynamics, SKMD preserves the Boltzmann distribution as the asymptotic distribution of the dynamics. This property enforces a balance between the exploration of diverse configurations and attraction toward high-probability regions of the energy landscape. We further propose an approach to efficient online data acquisition using an adaptive stopping criterion that selects non-redundant training data over the course of simulation. We demonstrate SKMD for the active learning of a neural network model of the Muller-Brown potential and the fine-tuning of a MACE interatomic potential for alanine dipeptide. Compared to active learning baselines, our method achieves higher model accuracy in fewer training iterations, with the same number of acquired training samples.
PaperID: 5539, Poster
Abstract: Graph-level anomaly detection plays a critical role in a wide range of applications, such as fraud detection in financial networks and malicious program detection, yet it remains challenging in real-world scenarios where labeled anomalous graphs are extremely scarce and graph structures are highly diverse. Existing few-shot methods often suffer from severe performance degradation due to the effects of class imbalance and structural imbalance, and they struggle to generalize to structurally deviant or rare anomalies. In this paper, we propose an adaptive few-shot graph-level anomaly detection model with multi-view structure-aware prompting (GADMVP). To mitigate structural imbalance, we construct a graph-of-graphs representation that facilitates cross-graph message passing, capturing both intra-graph semantics and inter-graph dependencies. Building upon these priors, we introduce a structure-aware graph prompting mechanism that extracts and adaptively refines anomaly-relevant prompts from a small set of labeled graphs. Extensive experiments on multiple benchmark datasets demonstrate that the proposed method consistently outperforms state-of-the-art few-shot graph anomaly detection baselines, while exhibiting strong robustness under extremely limited labeled data. The code and implementation details are available at https://anonymous.4open.science/r/GLAD-3456543/.
PaperID: 5540, Poster
Abstract: Long-horizon time series forecasting requires a careful balance between predictive accuracy and computational cost. Segmented temporal modeling has therefore become a practical strategy for reducing the burden of long-range prediction. However, real-world time series often contain superposed multi-periodic components, whose spectral structures vary across time, variables, and samples. Such variability limits fixed aggregation or frequency-modeling mechanisms, making them less adaptive to non-stationary periodic patterns and potentially causing informative components to be attenuated over long forecasting horizons. To address this limitation, we propose DSR-TSF, a spectrum-driven dynamic routing framework for long-horizon forecasting. DSR-TSF adaptively selects and combines multiple period-specific branches conditioned on the spectral characteristics of the input sequence, thereby improving the model's ability to capture non-stationary periodic structures. It further learns period-aware dynamic weights to modulate the contributions of multi-periodic patterns while retaining the segmented prediction structure. Without complex resampling or explicit period alignment, DSR-TSF provides a compact mechanism for modeling spectral heterogeneity and achieves competitive performance across seven datasets with diverse dimensionalities and temporal characteristics.
Abstract: Post-training improves instruction-following and helpfulness of large language models (LLMs) but often reduces generation diversity, which leads to repetitive outputs in open-ended settings, a phenomenon known as mode collapse. Motivated by evidence that LLM layers play distinct functional roles, we hypothesize that post-training-induced diversity loss is unevenly distributed across layers and that restoring a carefully chosen range of layers to their pre-trained weights can recover diversity while maintaining high output quality. To operationalize this idea, we design a proxy task---Constrained Random Character (CRC)---with an explicit validity set and a natural diversity objective. Results on CRC reveal a clear diversity–validity trade-off across restoration ranges and identify configurations that increase diversity with minimal quality loss. Based on these findings, we propose Selective Layer Restoration (SLR), a training-free method that restores selected layers in a post-trained model to their pre-trained weights, yielding a hybrid model with the same architecture and parameter count, incurring no additional inference cost. Across three different tasks (creative writing, open-ended question answering, and multi-step reasoning) and three different model families (Llama, Qwen, and Gemma), we find SLR can consistently and substantially improve output diversity while maintaining high output quality.
Authors:
feilong tang, Xiang An, Yunyao Yan, Yin Xie, Kaicheng Yang, Yifei Shen, Yuanhan Zhang, Shikun Feng, Chunyuan Li, Changrui Chen, Huajie Tan, Ming Hu, Manyuan Zhang, Bo Li, Ziyong Feng, Ziwei Liu, Zongyuan Ge, Jiankang DengAbstract: Visual signals, especially videos, are highly redundant: most regions are temporally predictable, while informative changes are sparse. We hypothesize that vision encoders should allocate computation according to this non-uniform information density rather than process dense pixel grids uniformly. From this perspective, codec-derived motion and residual signals provide a natural basis for identifying informative regions. OneVision-Encoder instantiates this idea through Codec Patchification, which replaces uniform dense computation with selective processing of only 3.1%-25% of regions that carry high signal entropy. To support irregular spatiotemporal token layouts, OneVision-Encoder employs a shared 3D RoPE and is pretrained with a large-scale cluster discrimination objective over more than one million semantic concepts, enabling unified representation learning for appearance and motion. Empirically, OneVision-Encoder improves both efficiency and accuracy under fixed token budgets. When integrated into large multimodal models, it achieves a 4.1% average improvement over Qwen3-ViT on video understanding benchmarks while remaining competitive on image and document understanding. Under attentive probing, it further yields substantial gains on motion-sensitive benchmarks, improving Top-1 accuracy on Diving48 by 17.1 and 8.1 percentage points over SigLIP2 and DINOv3, respectively, at matched patch budgets. These results suggest that codec-guided patch sparsity is an effective design principle for scalable visual representation learning.
Abstract: Speculative decoding accelerates LLM inference by drafting multiple tokens in parallel, with tree-based methods further improving efficiency through structured hierarchies. Dynamic-tree methods such as EAGLE-3 achieve excellent performance under greedy decoding via deterministic top-K expansion and global pruning. However, in stochastic decoding (T>0), this mechanism collapses the draft distribution into one-hot probabilities, causing a severe drop in acceptance rate. This exposes an apparent dilemma: dynamic-tree methods sacrifice stochastic sampling to preserve context-aware topology, while static-tree methods preserve stochastic sampling with context-agnostic structures. The issue arises because the same probability distribution is used for two conflicting tasks: constructing the tree and verifying the tokens. This coupling makes direct injection of randomness extremely challenging, as we are faced with a complex stochastic process. We resolve this by decoupling these two roles: RheoSampling assigns a token sampled from the draft distribution a proxy probability (for tree expansion and pruning) alongside its true sampling probability (for verification). Specifically, we inject a sampled token among the deterministic top-K slots and treat it with different probabilities in the construction and verification process, making RheoSampling the first dynamic-tree method with both context-aware top-K construction and stochastic sampling while maintaining losslessness. We establish the lossless guarantee through a novel equivalence-class analysis that compresses the stochastic tree space into tractable classes. An OT-based verification strategy and a sparse draft mechanism ensure that theoretical gains translate into practical efficiency. Experiments across diverse LLMs and benchmarks demonstrate consistent improvements in acceptance rate and speedup over state-of-the-art dynamic tree methods. This framework may provide a template for analyzing other complex stochastic tree structures.
Abstract: Reference-based image quality assessment (IQA) metrics aim to reflect how humans perceive the perceptual distance between a pair of images. To learn how the human visual system (HVS) operates, recent reference-based IQA metrics heavily rely on human-annotated data. Mean opinion score (MOS)-based pointwise scoring, which assigns a scalar quality value per image, is preferable for annotation but is prohibitively expensive to collect at scale and is known to be noisy due to inconsistent human judgments. As an alternative, two-alternative forced choice (2AFC) pairwise labels have gained popularity due to their reliability and efficiency, but they capture only relative comparisons between pairs. In this paper, we propose a fully automated data generation pipeline that generates pointwise perceptual distance labels between image pairs without any human annotation. Our approach exploits the generative dynamics of diffusion models as a perceptual distance proxy, where the coarse structure of an image is generated in the early timesteps and the fine details are generated in the later timesteps. Images that fork early in the generation process share only coarse structure and are perceptually far apart; images that fork late differ only in fine detail. We demonstrate that the diffusion trajectory aligns well with the human visual system, and use this forking moment, FoMo, as a reference-grounded distance label to supervise the training of a reference-based IQA metric. The pointwise labels, which support universal comparison between arbitrary image pairs, enable an information-rich training objective. Extensive experiments across diverse backbone architectures confirm the effectiveness of our generation pipeline, outperforming human-annotated datasets in multiple benchmarks.
PaperID: 5545, Poster
Abstract: Large Language Models (LLMs) demand efficient inference, where the multiply-accumulate (MAC) operations in linear projections dominate compute and energy cost. Post-training quantization (PTQ) reduces this cost by mapping weights and activations to low-bit integers. Power-of-Two (PoT) weight quantization promises efficient multiplier-free LLM inference via shift-add logic, but suffers significant accuracy degradation at low bitwidths due to its coarse base-2 quantization grid. We introduce Fractional PoT (FPoT), a base-\sqrt2 (half-octave) fractional PoT grid that achieves practical 4-bit weight quantization while preserving a fully integer shift-add datapath. The method combines (i) the base-\sqrt2 weight grid with uniform integer activations, (ii) a calibration-time grid-alignment refinement absorbed into the static per-channel scale, and (iii) a lightweight 1-adder \sqrt2 approximation whose error is absorbed during weight calibration via approximate-grid quantization, yielding a processing element (PE) with significantly fewer gates than a 4-bit integer multiplier. On Llama-3-8B with 4-bit weight and activation quantization, FPoT improves zero-shot commonsense reasoning performance by 8.72% relative to the PoT baseline, while maintaining performance close to the uniform integer baseline, with only a negligible 0.4%p degradation. These results hold across various configurations spanning seven models and three bit-widths. Furthermore, FPoT can be integrated with state-of-the-art LLM PTQ methods, demonstrating framework-agnostic applicability.
PaperID: 5546, Poster
Authors: Hyunsik Kim, Youngmoon Jung
Abstract: Byte-level byte-pair encoding (BBPE) tokenizers are attractive for large language models (LLMs) in multilingual settings because they cover all Unicode text. Under UTF-8, however, many scripts start from a higher byte-level fallback cost than English: when no learned merge covers a character c, the tokenizer must emit multiple byte-derived symbols. We call this worst-case pre-merge cost the encoding floor: f(c) = |\\mathrmutf8(c)| \\in \\1,2,3,4\\. A higher floor inflates token budgets, shrinks usable context, and increases per-request cost. Changing the underlying text encoding can reduce this gap, but a single global encoding can also make already-efficient English spans more expensive in mixed-script text. We propose Universal Byte-Level Encoding (UBE), a dual-alphabet tokenizer with deterministic per-character routing: r(c) = \\begincases \\mathrmUTF\\text-8, & |\\mathrmutf8(c)| \\leq 2,\\\\ \\mathrmUTF\\text-16, & |\\mathrmutf8(c)| > 2. \\endcases Characters satisfying |\\mathrmutf8(c)| \\leq 2 stay on the UTF-8 path, while characters satisfying |\\mathrmutf8(c)| > 2 are routed through UTF-16. This reduces token costs for Basic Multilingual Plane (BMP) scripts whose characters typically satisfy |\\mathrmutf8(c)| = 3 and have the highest token premiums, without raising costs for already-efficient spans. UBE changes only the byte representation presented to byte-pair encoding (BPE): the BPE merge rule remains standard, and exact decoding \\mathrmdecode(\\mathrmencode(s)) = s is preserved. Across intrinsic tokenization evaluations, UBE reduces cross-lingual dispersion in English-normalized token-count ratios \\pi(\\ell) = m(\\ell)/m(\\mathrmen), lowering cross-script token-budget disparity. In multilingual language model experiments, UBE preserves LM quality within seed variability. In the main multilingual settings, UBE reduces token counts most for high-premium scripts and also slightly lowers English token counts; these reductions translate into more usable context under fixed token budgets and faster content-matched prompt processing.
PaperID: 5547, Poster
Abstract: Explainable AI (XAI) offers far-reaching consequences for human-AI collaboration, but the fundamental question of how best to evaluate these systems has proven frustratingly elusive. Methodologies typically revolve around either proxy metrics which tell us little about the utility of an explanation, or human evaluation which---while generally considered the gold standard---is expensive, difficult to scale, and challenging to reproduce. In this paper, we investigate if Multimodal Large Language Models (MLLMs) can act as surrogate human participants to automate the evaluation of explainable AI. To do this, we first propose a ``purpose-grounded'' evaluation framework with specific metrics applicable to both MLLMs and human users alike. Then, we offer some conceptual insights as to why and when we can trust these models to automate this process. Lastly, we conduct extensive empirical testing across five purpose-grounded areas and nine popular MLLM APIs. Results indicate that all models trend positively, but OpenAI models (particularly GPT-5 mini) are the best aligned at replicating user responses, and for complex visual reasoning tasks more advanced modes (like GPT-5) perform best, which would be the recommended choices in practice. We expect our insights not to replace user testing, but rather to help XAI researchers complement their current evaluation methodologies and move the field away from heavy reliance on proxy metrics for measuring explanation utility.
Authors:
Hao Xiang, Qiaoyu Tang, Le Yu, Yaojie Lu, Xianpei Han, Ben He, Le Sun, Bowen Yu, Peng Wang, Hongyu Lin, Dayiheng LiuAbstract: Reinforcement Learning (RL) with verifiable environments has emerged as a powerful approach for enhancing the reasoning capabilities of Large Language Models (LLMs). While prior research indicates that scaling environment quantity improves RL performance, existing manual or individual construction methods suffer from linear scaling limits, thereby hindering scalable reasoning generalization. This paper introduces RACES (Recursive Automated Composition for Environment Scaling), a framework that treats verifiable environments as composable building blocks that can be recursively assembled. The key insight is that when the codomain (output type) of one environment matches the domain (input type) of another, they can be automatically fused into a new verifiable environment, enabling recursive composition. RACES is instantiated with 300 individual environments and defines a set of composition operators (Sequential, Parallel, Sort, and Select) that induce diverse reasoning patterns. Extensive experiments show that RL training on these composite environments consistently enhances reasoning generalization. Specifically, RACES improves DeepSeek-R1-Distill-Qwen-14B by an average of 3.1 points (from 48.2 to 51.3) and boosts Qwen3-14B performance from 58.8 to 61.1 on five benchmarks, which are unseen during the construction of training environments. Moreover, RACES achieves performance comparable to training on 300 individual environments using only 50 base environments, demonstrating significant efficiency in environment utilization.
Abstract: Radio astronomy is an indispensable discipline for studying distant celestial objects. Measurements of wave signals from radio telescopes, called visibility, need to be transformed into images for astronomical observations. These dirty images blend information from real sources and artifacts. Therefore, astronomers usually perform reconstruction before imaging to obtain cleaner images. Existing methods consider only a single modality of sparse visibility data, resulting in images with remaining artifacts and insufficient modeling of correlation. We propose VVTRec, a multimodal radio interferometric data reconstruction method, to enhance visibility information extraction and improve image-domain output quality. Since the information in sparse visibility is inherently limited, we transform it into image-form and text-form features. Therefore, we can leverage vision-language models to process the two derived modalities, providing knowledge banks as a supplement for reconstruction. The sparse visibility is accordingly utilized as the query to perform the integration and selection of multimodal external knowledge. Consequently, VVTRec improves the structural integrity and accuracy of reconstructed images by enriching spatial and semantic information. Our experiments demonstrate that VVTRec effectively enhances imaging results by exploiting multimodal information without introducing excessive computational overhead.
PaperID: 5550, Poster
Abstract: 3D Gaussian Splatting (3DGS) has emerged as a mainstream representation for real-time novel view synthesis and 3D reconstruction. However, existing 3DGS rendering pipelines are poorly aligned with modern GPU architectures: depth-sorted alpha compositing introduces strong per-pixel sequential dependencies, limiting parallelism and resulting in scalar Fused Multiply-Add (FMA)-dominated execution that underutilizes high-throughput matrix units such as Tensor Cores. This mismatch becomes more severe for multimodal and semantically enriched 3DGS, where high-dimensional features must be accumulated from millions of Gaussians, leading to rapidly increasing computation and memory costs. To address these challenges, we propose a matrix-parallel rendering framework for feature-scalable 3DGS on modern GPUs. We first identify the sequential dependency of alpha compositing as the key obstacle to efficient high-dimensional feature rendering, and derive an order-independent weighted accumulation formulation that transforms feature rendering into matrix-oriented computation. Our framework features three key designs: i: a tile-based splatting pipeline that maps feature accumulation to warp-level matrix-multiply-accumulate execution on Tensor Cores; ii: a warp-granularity Gaussian packing strategy that mitigates visibility irregularity and improves feature reuse; and iii: a lazy channel-batched feature-fetching pipeline that reduces shared-memory pressure, maintains high warp occupancy, and overlaps memory access with matrix execution for high-dimensional features. Extensive experiments on Mip-NeRF 360, LERF, and 3DOVS demonstrate that our approach enables both high performance and strong scalability for feature-rich 3DGS rendering. Particularly, our method achieves up to 3.8× speedup over 3DGS and 2× over TC-GS while preserving rendering quality and semantic accuracy, and further reduces kernel render time by up to 36.7× for high-dimensional features.
PaperID: 5551, Poster
Abstract: Mathematical optimization drives decisions in supply chains, logistics, energy, and scheduling, but translating natural-language problems into solver-executable formulations remains a bottleneck. While frontier models show promise for this task, many deployments require small language models (SLMs) that run locally under privacy, latency, and cost constraints. Existing SLMs are limited by noisy data, ambiguous benchmarks, and one-shot inference that ignores the iterative, feedback-driven nature of optimization modeling. We introduce \textscOptiMind, a framework for improving SLMs through formulation-centric data cleaning, class-conditional priors, and multi-turn reinforcement learning with solver feedback. \textscOptiMind repairs ambiguous problems, regenerates solution trajectories, audits benchmark labels, constructs accepted objective-value sets, and distills class-level hints from recurring formulation errors. It then trains models to formulate, execute, and revise GurobiPy programs using solver-verified correctness as reward. On expert-cleaned IndustryOR, Mamo-Complex, and OptMATH benchmarks, a 20B \textscOptiMind-RL model outperforms strong open-source baselines and reaches frontier-model performance under the same protocol. Our results suggest that optimization formulation is better treated as an interactive solver-grounded decision process than as one-shot code synthesis. We release our inference framework, cleaned benchmarks, and sample error analyses at \urlhttps://anonymous.4open.science/r/OptiMind-NeurIPS-1DB020, with full data and model to follow.
PaperID: 5552, Poster
Abstract: Federated Learning (FL) is vulnerable to model poisoning attacks, where malicious clients manipulate uploaded model updates to corrupt the global training process. Existing attacks are typically evaluated by their eventual damage, such as final accuracy degradation or convergence to random-guess performance, where the learned model becomes unusable and may ultimately lead to denial-of-service (DoS). However, such metrics overlook attack round complexity, defined as the number of communication rounds required for the poisoned global model to reach and remain near a desired degradation objective. In realistic FL systems, requiring more communication rounds increases the adversary's exposure to client participation, robust aggregation, filtering, and detection. We propose ASAP (Adaptive Sliding Agnostic Poisoning), a fast aggregation-rule-agnostic (AGR-agnostic) attack that treats the whole FL training process as an uncertain dynamical system and formulates poisoning as a feedback-control problem for steering the model toward a desired attack objective. ASAP combines adaptive sliding mode control (ASMC) with finite Fourier-basis uncertainty estimation, treating the unknown effects of benign training and aggregation as a time-varying uncertainty and using the resulting estimate to design the malicious updates. We theoretically prove that the sliding variable reaches the sliding surface in finite time and the tracking error subsequently converges exponentially. Experiments across multiple datasets, models, and aggregation rules demonstrate that ASAP reaches specified degradation objectives in fewer communication rounds while maintaining smaller target deviation than existing model poisoning attacks.
PaperID: 5553, Poster
Abstract: Recent studies have shown that leveraging the reasoning and analytical capabilities of large language models (LLMs) for long video understanding has become a promising approach. However, these methods are constrained by a structural limitation: representing video as a flat frame sequence makes it difficult to model the temporal structure and logical dependencies of events, which in turn hampers long-range reasoning and risks overlooking important information. To address this limitation, we propose LVGraph, a framework that models video content using a coarse-to-fine hierarchical graph. LVGraph constructs a semantic hierarchy where high-level graphs characterize the global event structures within the video, while low-level graphs capture fine-grained, query-relevant information. Instead of linear scanning, reasoning is performed via a query-aware traversal of this graph, adaptively identifying and retrieving the most salient keyframes by navigating the nodes and edges most relevant to the query. Finally, we leverage a Vision Language Model (VLM) to augment the keyframe information, and subsequently answer the question by reasoning over the graph and the enriched keyframe captions. Extensive experiments on long video understanding benchmarks confirm the effectiveness of our method. LVGraph significantly outperforms existing LLM-based state-of-the-art methods on EgoSchema, NExT-QA, and long-duration segments (averaging 44 minutes) of the Video-MME benchmarks.
Abstract: Face Recognition systems are widely deployed in real-world applications, but they also raise privacy concerns due to unauthorized collection and misuse of facial data. Existing adversarial privacy protection methods rely on input-space perturbations to obfuscate identity information, yet their protection can degrade when adversaries learn restoration or purification mappings that partially invert the transformation. We study this setting as an asymmetric adversarial attack, in which reverse manipulation becomes feasible because existing defense paradigms do not control reversibility. To address this problem, we propose Asymmetric Reversible Face Protection (ARFP), a restoration-aware extension of personalized face cloaking that integrates privacy protection, keyed recovery, and tamper indication in a single framework. ARFP consists of three components: Key-Conditioned Manifold Binding, which ties the protection transformation to a user-provided key; Adversarial Restoration-Aware Training, which introduces a surrogate restoration adversary during training to improve robustness against evaluated inverse purification attacks; and Authorized Reversible Restoration, which supports recovery with the correct key while providing nonce-based tamper indication. Extensive experiments under the threat models considered in this work show that ARFP improves resistance to the evaluated restoration attacks while preserving authorized recovery utility. These results provide empirical evidence of key-sensitive recovery behavior and tamper awareness in the tested settings.
Authors: Yanqng Luo, Julius Hense, Niklas Prenißl, Andreas Mock, Klaus-Robert Müller, Thomas Schnake, Mina Jamshidi Idaji
Abstract: Explanations of multiple instance learning (MIL) models are widely used for validation and discovery in digital histopathology. Existing methods primarily rely on heatmaps that highlight influential regions but do not explain how evidence from different tissue regions is combined to produce a prediction. This limits interpretability, especially when decisions depend on interactions between tissue features. We introduce Symbolic explainable MIL (Symb-xMIL), a post-hoc explanation framework that quantifies how a MIL model’s behavior aligns with human-readable decision rules, expressed as logical relationships (e.g., AND, OR, NOT) between input features. These alignment scores reveal semantic patterns underlying the model’s predictions. We evaluate Symb-xMIL on synthetic and real-world histopathology datasets. On synthetic MIL data, Symb-xMIL reliably recovers ground-truth logical rules. In a clinical tumor detection task, the best-aligned rules uncover heterogeneous decision patterns and expose hidden model errors. On an HPV-prediction task on TCGA-HNSCC, a cohort of head and neck cancer, our framework refines patient survival stratification beyond HPV status with potential clinical relevance. Overall, Symb-xMIL extends MIL explainability beyond visual attribution toward structured, rule-based reasoning, enabling more transparent and semantically grounded interpretation of model predictions.
PaperID: 5556, Poster
Abstract: Subject-consistent generation (SCG) aims to preserve subject identity across diverse text-driven contexts. Recent training-free methods move beyond global sharing via correspondence-aware identity transfer. However, semantic-driven correspondence may transfer identity features to appearance-similar but structurally incompatible regions. Meanwhile, extended attention, commonly used in existing methods, enhances identity consistency but may perturb the emerging target structure, from global pose and layout to local structural details. Together, these issues can cause reference-pose over-copying, local structural artifacts, and degraded structural fidelity. We view training-free SCG as controlled identity injection, where the pretrained model determines the target structure, while the reference image provides subject identity. We propose TopoRefine, a training-free framework with Topology-aware Subject Correspondence (TSC) and Residual Identity Refinement (RIR). TSC regularizes semantic matching with subject-internal geodesic relations for structurally compatible identity transfer, while RIR injects reference identity as a foreground-gated, magnitude-aligned residual to preserve the emerging target structure. Both qualitative and quantitative results show that TopoRefine improves subject consistency while better preserving target-driven structure.
PaperID: 5557, Poster
Authors: Rui Xiao, Jinming Xu, Jinsong Han
Abstract: Accurate monocular metric depth estimation, i.e., predicting absolute depth values, is critical for applications such as robotics and augmented reality. However, existing methods often suffer from poor transferability to in-the-wild scenarios due to the lack of knowledge of camera intrinsics. In this paper, we propose a novel framework, Roll2Depth, to recover metric scale in unseen environments with unknown camera parameters. Our key idea is to leverage scene-independent physical cues induced by the rolling shutter effect of standard CMOS sensors as metric priors. When a scene is illuminated by a flickering light, the rolling shutter produces distinctive stripe-like spatial patterns, revealing the scene's shading with and without the flickering light source. This contrast encodes absolute depth cues that can be exploited for metric estimation. Roll2Depth isolates these patterns as metric priors to guide affine-invariant models on metric depth estimation in a zero-shot setting. Our evaluations on standard benchmark datasets, as well as a real-world dataset collected using commercial smartphones and our prototype system, demonstrate that Roll2Depth, as a low-cost active solution, outperforms state-of-the-art passive baselines in zero-shot settings. These results highlight its potential impact in applications such as indoor robotics and embodied perception, where affordable and effective active sensing would be valuable.
PaperID: 5558, Poster
Abstract: How can we understand what a music foundation model has learned internally? Most interpretability approaches, such as probing and Sparse Autoencoders (SAEs), focus on identifying individual features with minimal structural assumptions. We argue that many concepts are better understood as structured relations rather than isolated features. This is especially prominent in music, where tonal structures are organized in the space of pitch and time. For example, concepts such as chords or keys are naturally expressed as structured sets (e.g., the 12 transpositions of a chord or the diatonic system within a key), rather than isolated features. In this study, we shift from feature identification to structure-based analysis, asking whether the learned inner representations of music language model emerge as organized structures over features. To this end, we introduce a framework that uses pitch transposition as an inductive bias to induce ordered orbits via multi-view SAE alignment. Concretely, we generate pitch-shifted input pairs and align their SAE representations to discover structured groups of pitch-related features. Experimental results show that this approach recovers orbit structures corresponding to chords, keys, and melodic patterns across two state-of-the-art music foundation models, while requiring only minimal grounding (e.g., a few anchor examples) to interpret entire concept families. Beyond interpretability, our results suggest a new paradigm for representation analysis: evaluating whether coherent relational structures emerge in learned representations as a proxy for model understanding.
Abstract: Large language models increasingly serve as execution engines for agentic systems, yet they still consume context through a sequential text interface. This creates a mismatch with modern structured agent workflows, in which independent branches explore subtasks, retrieve evidence, or generate candidate solutions before a final synthesis step. Existing systems typically merge these branches by concatenating their textual outputs, which discards the parallel structure and incurs redundant prefill computation. In this work, we introduce Parallel-Synthesis, a plug-and-play framework that enables a synthesizer to directly consume the KV caches produced by parallel worker agents. Parallel-Synthesis combines a cache mapper that calibrates independently generated branch caches with a fine-tuned synthesizer adapter that enables generation from this non-sequential cache interface. We train Parallel-Synthesis using data that exposes the synthesizer to parallel cache contexts, teaches aggregation across cached branches, and distills reasoning behavior from standard text-concatenation-based synthesis. Across nine downstream datasets spanning math, science QA, code generation, GAIA, and multi-agent database diagnosis, Parallel-Synthesis matches or outperforms text-based synthesis on seven datasets and remains close on the other two. It also reduces time-to-first-token by 2.5×--11×, suggesting that direct cache-based synthesis is a promising interface for more native and efficient synthesis over parallel agent branches.
PaperID: 5560, Poster
Abstract: In unsupervised multivariate time series anomaly detection (MTSAD), local variation can inflate reconstruction and forecasting errors, making observation-space discrepancy an unreliable indicator of abnormal system behavior. Since latent-space scoring can mitigate this limitation, we reframe one-class MTSAD as current-point system-state compatibility estimation and propose work for Multivariate Time Series Anomaly Detection. We define the current-point system state as a latent representation of the target time point, constructed from variable-level temporal context in the input window. TENET constructs this state through a temporal layer that retrieves window context relative to the current point, encodes time with current-point-centered embeddings, and adaptively mixes the retrieved context with the current representation. With this design, TENET treats the input window not as the object to reconstruct, forecast, summarize, or score, but as context for constructing a latent representation evaluated with a fixed standard-normal negative log-likelihood score. Across benchmark datasets, TENET achieves state-of-the-art performance. Additional experiments further show that TENET remains effective under alternative one-class objectives and outperforms substitutions based on generic sequence encoders, suggesting that its advantage comes from decision-aligned representation construction rather than from a specialized scoring rule. These results indicate that one-class MTSAD can be strong when the latent representation is designed around the current time point.
Abstract: Preference optimization aligns large language models from pairwise preferences by increasing the margin between preferred and dispreferred responses. However, margin-based losses alone do not test whether each pair's preference signal is locally stable enough to trust with full update strength. We propose Geometric Anchor Preference Optimization (GAPO), a geometry-aware objective that introduces a batch-conditioned stress test for preference learning. For each mini-batch, GAPO constructs a pessimistic anchor by perturbing the current policy in the first-order direction that decreases the batch-average preference margin. The resulting Anchor Gap measures how much each pair's margin degrades under this shared pessimistic probe and converts this degradation into an instance-wise update weight. Pairs with larger Anchor Gap receive smaller update weights, while non-brittle pairs retain their preference-gradient direction. Under smoothness assumptions, we characterize the Anchor Gap as a batch-directional proxy for local margin degradation. Across multiple open-weight model families, GAPO matches or improves strong preference-optimization baselines on instruction-following and reasoning benchmarks. It also improves robustness under random and structured preference noise without explicitly modeling label corruption. Mechanistic diagnostics show that the pairs receiving the smallest GAPO weights are statistically enriched in corrupted supervision, suggesting that GAPO improves robustness by reducing the cumulative influence of brittle preference signals.
Abstract: Counterfactual explanations are a central tool in interpretable machine learning, yet computing them exactly for complex models remains challenging. For tree ensembles, predictions are piecewise constant over a large collection of axis-aligned hyperrectangles, implying that an optimal counterfactual for a point corresponds to its projection onto the nearest rectangle with an alternative label under a chosen metric. Existing methods largely overlook this geometric structure, relying either on heuristics with no optimality guarantees or on mixed-integer programming formulations that do not scale to interactive use. In this work, we revisit counterfactual generation through the lens of nearest-region search and introduce , a global representation of recourse for tree ensembles. Leveraging the fact that any tree ensemble can be compressed into an equivalent partition of labeled hyperrectangles, we cast counterfactual search as the problem of identifying the associated with the nearest rectangle of an alternative label. This leads to an exact, amortized algorithm based on volumetric k-dimensional (KD) trees, which performs branch-and-bound nearest-region queries with explicit optimality certificates and sublinear average query time after a one-time preprocessing phase. Our experimental analyses across several real datasets from high-stakes application domains show that this approach delivers globally optimal counterfactual explanations with millisecond-level latency, achieving query times that are orders of magnitude faster than existing exact, cold-start optimization methods.
PaperID: 5563, Poster
Authors:
Tinghe Zhang, Yucheng Xiao, Alex LambAbstract: In Transformer attention, distinct weight vectors \boldsymbol\alpha \neq \boldsymbol\alpha' can produce the same aggregated representation \mathbfV\boldsymbol\alpha = \mathbfV\boldsymbol\alpha'. When \operatornamerank(\mathbfV) \leq n-2, configurations with entirely different dominant source tokens collide on a set of positive Lebesgue measure, a condition satisfied in over 97% of attention heads in BERT-Base. Such collisions are harmful: the downstream FFN receives identical inputs despite the attention having attended to entirely different source tokens, and must therefore produce identical outputs regardless of which tokens actually dominated, fundamentally limiting any computation that requires sensitivity to attention source, such as multi-hop reasoning, attribution, or knowledge retrieval. No increase in depth, width, or data can compensate, because the routing weights are discarded before the FFN is reached. The routing weights \\alpha_ij\, however, are available within the same forward pass and can be retained as a compact side-channel to restore the missing signal. We call this structural information loss routing non-identifiability. In BERT-Base, value-matrix rank deficiency suppresses k \geq 107 routing dimensions (over 84% of degrees of freedom) before any FFN computation. To quantify the practical gap, we introduce the Routing Reconstruction Task (RRT): standard Transformers achieve 31.4% source-identification accuracy while oracle routing statistics reach 94.6%, a 63-point gap consistent with an information bottleneck in the FFN input. We then propose the Route-Aware FFN (RA-FFN), a drop-in module that appends four compact routing statistics to each FFN block at under 1.1% additional parameter cost, recovering over 75% of the oracle gain. RA-FFN improves all benchmarks tested: at 110M scale, gains cover multi-hop reasoning, language modelling, and translation, with multi-hop gains 6× larger than single-hop; at 7B scale on Llama-2-7B and Mistral-7B, gains extend further to mathematics (GSM8K) and general reasoning (MMLU, ARC-Challenge, HellaSwag), ordered by routing sensitivity.
PaperID: 5564, Poster
Abstract: Kotłowski et al. (NeurIPS 2017) posed as a central open problem the design of efficient online isotonic regression algorithms beyond totally ordered domains. We resolve this for the product order [m]×[m], the canonical two-dimensional case, by determining the minimax regret up to constant factors: R_T^ = \Theta\left(\min\left\\T,\ \inf_K \in \mathbbN_+\left[\log\Omega([m]^2, K+1) + \fracTK^2\right]\right\\\right), where \Omega(\mathcalP, K+1) is the order polynomial counting monotone maps from \mathcalP to [K+1]. This unified formula reveals a three-phase scaling law: \Theta(T) for T \lesssim m, \Theta(m^2/3 T^1/3) for m \lesssim T \lesssim m^4, and \Theta(m^2 \log(T/m^4)) for T \gg m^4. We further construct polynomial-time algorithms achieving this rate in every regime, including a horizon-free variant requiring no advance knowledge of T.
PaperID: 5565, Poster
Abstract: Medical vision-language models are often adapted under high-quality acquisition conditions. At target deployment, the same model may face lower-quality scanners, protocols, reconstructions, or artifacts. These shifts change the visual evidence, while the intended clinical question remains unchanged. As a result, a source-adapted model can fail under this acquisition shift. Methods are proposed to solve this problem. Medical acquisition-shift correction mainly targets image-only predictors, leaving the multimodal model underexplored. Robust training depends on anticipated degradations, and generic test-time tuning can easily overfit the small target training set, leaving unknown acquisition shifts insufficiently corrected. In this paper, we hypothesize that acquisition-like augmentations of target examples can expose parameter directions responsible for acquisition-shift failures. Based on it, we propose \textscASSET, an \emphAcquisition-Sensitive Subspace Estimation method for test-time adaptation. ASSET estimates an acquisition-sensitive parameter subspace from prediction and gradient changes under these augmentations. It then adapts the model within this subspace through iterative estimate-and-update rounds. Comprehensive experiments show the advantages of ASSET. For example, ASSET consistently improves target accuracy under acquisition-quality shift across datasets and backbones by 2.62% on average, while preserving performance when no shift is present. Under the strongest acquisition-quality shift, ASSET even improves target accuracy by 4.12%, showing that the correction generalizes to severe shifts.
PaperID: 5566, Poster
Abstract: As agents have to solve diverse tasks from grid-worlds to robotics and language modeling, the difficulty of each step varies, yet policies are rigid and spend a fixed amount of compute on each step. A natural approach is to spend compute on the harder states while saving it on easy ones, allowing agents to retain performance while saving valuable compute budgets. To this end, we propose Adaptive Internal Computation (AIC), a lightweight policy architecture that iteratively refines latent states and learns when to halt using a value-of-computation signal.Concretely, AIC repeatedly refines a latent state representation and learns a halting head from a value-of-computation signal, the per-step gain in expected return, allowing the policy to stop refining as soon as further computation no longer improves the action it would take. A central challenge, however, is that matched-budget gains alone cannot establish that compute is being directed to the states that benefit the most from them. We therefore introduce a six-component causal evaluation protocol leveraging a targeted-versus-random ablation that tests whether removing compute from AIC's predicted high-compute states hurts performance more than removing the same amount of compute from randomly chosen states, thereby validating that AIC can successfully identify high-compute states. Across twelve diverse tasks, including grid worlds, simulated robotics experiments, and language modeling, we demonstrate that AIC outperforms fixed-compute policies across a set of five underlying model architectures. In particular, architectures leveraging AIC retain or improve performance while spending about 52% less compute compared to their respective fixed counterparts.
PaperID: 5567, Poster
Authors: Shuoyang Wang, Yidan Tian, Guanqun Cao
Abstract: Despite the remarkable empirical success of transformer models in natural language processing, their theoretical foundations for functional data classification remain largely unexplored. This paper takes a first step toward closing this gap by developing a rigorous statistical framework for transformer-based functional classifiers. We show that a transformer architecture with an expanding attention window in the latent representation attains minimax-optimal excess-risk rates up to logarithmic factors in infinite-dimensional functional classification, without relying on conventional dimension-reduction procedures. Our analysis further reveals a dense-to-sparse phase transition: when the sampling frequency exceeds a critical threshold relative to the sample size, the proposed classifier achieves the optimal dense-observation rate; otherwise, we precisely characterize the degradation in convergence under sparse sampling. Extensive simulations and real-data experiments support the theory and demonstrate that transformers provide a competitive and theoretically justified approach to functional classification.
Abstract: As real-world datasets become more complex and heterogeneous, supervised learning is often bottlenecked by input representation design. Modeling multimodal data, such as time-series, free text, and structured records, often requires non-trivial domain expertise. We propose an agentic pipeline to streamline this process. First, an LLM analyzes a small but diverse subset of text-serialized input examples , which acts as a programmatic specification for extracting and organizing evidence. This rubric is then used to transform text-serializations of inputs into a more standardized format for downstream models. We also describe , which are task-conditioned summaries generated by an LLM. Across 15 clinical tasks from the EHRSHOT benchmark, our rubric-based approaches significantly outperform traditional count-feature models, naive LLM baselines, and a clinical foundation model pretrained on orders of magnitude more data. Beyond performance, rubrics offer operational advantages such as being easy to audit, cost-effectiveness at scale, and facilitating tabular representations.
PaperID: 5569, Poster
Authors: J Z Gazak, Ryan Swindle, Justin Fletcher
Abstract: Open-weight Vision-Language Models (VLMs) score 85--94% on existing spatial benchmarks (BLINK Spatial, What's Up CLEVR) yet collapse to chance on simple geometric questions when object, scene, and texture priors are stripped away. We introduce and open-source Sophon, a procedurally generated diagnostic of paired colored point-cloud stimuli across three coordinate systems (1D linear, 2D Cartesian, 2D polar) and four question types (binary, categorical, comparative, quantitative). Across four open architectures spanning every major projector design (LLaVA-NeXT, LLaVA-OneVision, Qwen2.5-VL, Idefics3), accuracy clusters at 47--52% while frontier models (Opus 4.6, Gemini 2.5 Pro, GPT-5.4) clear 77--83% and humans reach 84--94%. We trace the open-weight failure end-to-end. The vision tower works: encoder probes recover Sophon spatial direction at 80--97%. The language tower's reasoning circuit works: a text-only control restores +40 pp, and probes preserve the spatial signal through every LLM decoder layer at 80--86%. Per-layer attention attribution localizes visual integration to mid-stack layers. Yet at the readout, the model elevates the correct answer-set tokens but selects among them at chance. Fine-tuning experiments rule out the unembedding as the mechanism. By moving the unembedding layer through the residual stream, we observe the emergence of a small but discriminative contrast at late-stack residual layers in fine-tuned models, which in practice converts the random selection into a confident one more likely to be correct. However, ablation experiments locate the mechanism upstream: training only the upstream layers with the late stack frozen recovers the full +29 pp gain, while training only the late stack captures less than a third. The late-stack contrast is the downstream readout of upstream changes. No layer in either base or fine-tuned models shows geometric alignment between the spatial signal and the unembedding's vocabulary axes, suggesting that 7B-class VLMs have the weights to reason spatially but lack the residual-stream pathways to convert that reasoning into answers.
Abstract: Large language models (LLMs) have achieved remarkable progress across diverse tasks, yet their internal mechanisms remain largely opaque. In this work, we investigate a fundamental question: to what extent can the original input text be recovered from a single last-token representation in an LLM? To this end, we propose Rep2Text, a novel framework for decoding text from last-token representations. Rep2Text employs a trainable adapter that maps a target model’s last-token representation into the token embedding space of a decoding language model, which then autoregressively reconstructs the input text. Experiments across various model combinations (Llama-3.1-8B, Gemma-7B, Mistral-7B-v0.1, Llama-3.2-3B, etc.) show that, on average, roughly half of the tokens in 16-token sequences can be recovered from this compressed representation while preserving strong semantic coherence. Further analysis reveals a clear information bottleneck effect: as sequence length increases, token-level recovery declines, while semantic information remains relatively well preserved. We also find that scaling effects are less pronounced in inversion tasks. Finally, our framework demonstrates robust generalization to out-of-distribution clinical data.
Abstract: On-policy distillation (OPD) and on-policy self-distillation (OPSD) have emerged as promising post-training methods for large language models, offering dense token-level supervision on trajectories sampled from the model's own policy. However, existing results on their effectiveness remain mixed: while OP(S)D has shown promise in system prompt and knowledge internalization, recent studies also report instability and degradation. In this work, we present a comprehensive empirical study of when OPD and OPSD work, when they fail, and why. We find that OPD on mathematical reasoning is highly sensitive to teacher choice and loss formulation, whereas OPSD fails in our tested settings due to test-time absence of instance-specific privileged information (PI). In contrast, OPSD is effective when PI represents a shared latent rule, such as a system prompt or alignment preference. We identify three failure mechanisms: (1) distribution mismatch between teacher and student caused by conditioning on student-generated prefixes, (2) optimization instability from biased TopK reverse-KL gradients, and (3) an OPSD-specific limitation where the student learns a PI-free policy that aggregates PI-conditioned teachers, which is insufficient when PI is instance-specific. We further show that stop-gradient TopK objectives, RLVR-adapted teachers, and SFT-stabilized students mitigate these failures.
PaperID: 5572, Poster
Authors:
Junxiang Zhang, Qingchun HE, Xinhui Lu, Yaxin Hu, Feng Li, Jing Tang, Kai HanAbstract: We study budget-feasible procurement auctions for social-welfare maximization under submodular valuations. In this setting, a buyer with a limited payment budget seeks to procure goods or services from strategic sellers, with the objective of maximizing the buyer's value for the selected sellers minus their true costs. We propose a simple yet effective single-round clock auction in which each seller receives at most one price offer. The mechanism satisfies desirable economic properties, including truthfulness, individual rationality, budget feasibility, and non-negative auctioneer surplus. Moreover, it achieves approximation ratios of 8.52 for monotone submodular valuations, 23.2 in expectation for non-monotone submodular valuations, and 24.52 deterministically for non-monotone submodular valuations, using only a linear number of value-oracle queries. These guarantees improve over the recent independent work of Cui et al.~(ICML 2026) on the same problem in approximation ratio, value-oracle complexity, and number of pricing rounds. Experiments on influence maximization in social networks and crowdsourcing further demonstrate the effectiveness and efficiency of our approach.
PaperID: 5573, Poster
Abstract: Large Vision-Language Models (LVLMs) show strong video understanding, yet hallucinations persist, with outputs that misalign with the video content. This stems from two primary causes. First, consecutive frames are often semantically similar, so the model can miss visual details. Second, LVLMs may rely on language priors and pay less attention to visual evidence. To mitigate video hallucinations, we propose , a training framework that enhances fine-grained discrimination and reduces reliance on language bias. VidHalluDoctor builds video difference samples to improve visual detail recognition, counterfactual samples to reduce prior bias, and consistency-aware reinforcement learning with designed rewards that evaluate both reasoning quality and consistency between the reasoning and answer. To better evaluate video hallucinations, we introduce , a comprehensive benchmark comprising 10,000 videos that covers three progressive hallucination types across diverse scenarios and multi-view settings. Extensive experiments demonstrate that VidHalluDoctor significantly reduces hallucinations while maintaining strong video understanding. Code and dataset will be released.
PaperID: 5574, Poster
Authors: Harvey Mannering, Zhiwu Huang, Adam Prugel-Bennett
Abstract: Emotions are expressed on a continuum and are naturally described in the valence-arousal (VA) space. Existing VA-based image editors are typically training-based and domain-specific, limiting them to VA regions seen during training and making them brittle across datasets and domains. Meanwhile, despite producing high-quality edits, state-of-the-art foundation image editors are not continuous affective controllers: non-linear emotion representations in text encoders makes prompt interpolation poorly calibrated, so accurate and continuous VA control remains unreliable. To fill this gap, we introduce Barycentric Guidance, a training-free framework that turns pretrained foundation image editors into continuous affective controllers over VA, requiring no finetuning, extra data, or model changes. Our method maps any target VA point to simplex-based guidance weights, enabling VA controllability in three aspects: (i) accurate targeting of VA states, (ii) broad coverage across the VA plane, and (iii) continuous control. Across three face datasets, we improve VA target accuracy by 25–29%, expand VA-plane coverage by 28–50%, increase expression diversity by 29–49%, over the strongest baseline. Unlike prior domain-specific methods, our training-free approach exploits the foundation model’s inherent versatility, enabling generalization beyond human faces to animals, artworks, and complex scenes within a single framework. Code will be made publicly available.
PaperID: 5575, Poster
Abstract: Reinforcement learning (RL) is a promising way to improve the clinical reasoning abilities of large language models (LLMs), but current pipelines often rely on expert-labeled questions or costly synthetic-data generation. Meanwhile, the medical literature offers a wealth of high-quality unlabeled documents, including clinical case reports that trace the full clinical course from initial presentation to final treatment decisions. We study whether document-grounded self-play can convert these reports into effective RL training signals. We identify two challenges that limit existing self-play methods in this setting: knowledge-point fixation, where the proposer repeatedly asks about a narrow set of salient clues in a source report, and difficulty-quality mismatch, where appropriately difficult questions may still be poorly formed or shallow, limiting their value for reasoner training. To address these challenges, we propose CasePlay, a self-play framework with two key components: (a) a Knowledge-Conditioned Proposer, which anchors question generation to case-specific knowledge points so that questions draw on a broader range of evidence from the source report; and (b) a Rubric-Judge Reasoner, which augments answer-accuracy feedback with rubric-based quality feedback on clue integration, reasoning depth, and option quality, guiding the proposer toward reasoning-intensive questions that better support reasoner training. Extensive experiments across seven medically related benchmarks demonstrate that CasePlay outperforms existing self-play baselines and improves the SFT warm-start model by 3.0 points on average, highlighted by a 6.9-point gain on LiveClin. This approach offers an effective path toward continually improving medical reasoning ability through self-play RL. Code is available at \urlhttps://anonymous.4open.science/r/CasePlay.
Abstract: Acceleration for deterministic root-finding problems has been extensively studied in recent years; specifically, the anchor-based, or Halpern-type methods achieve optimal convergence rates with respect to the operator norm. However, acceleration via these methods does not directly carry over to stochastic setting due to accumulation of errors, unless one enforces diminishing variance via increasing batch sizes or variance reduction techniques. In this work, we show that another class of acceleration, namely the dual-anchor mechanism, extends to the stochastic setting without such error accumulation, in contrast to anchor-based algorithms. Consequently, we cleanly achieve \mathcalO(\epsilon^-3) complexity with iteration-independent batch size, without any variance reduction or double-loop recursive regularization, for stochastic root-finding (resp. fixed-point) problems with cocoercivity (resp. square-nonexpansivity) in expectation. For strongly monotone operators, the same algorithm attains a sharper \widetilde\mathcalO(\epsilon^-2) complexity, nearly matching the lower bound.
PaperID: 5577, Poster
Abstract: Analog integrated circuit (IC) design is a multi-turn, simulator-grounded engineering task that demands iterative reasoning over coupled performance trade-offs, yet proprietary process design kits (PDKs) make hosted large-language-model workflows difficult to deploy in practical design settings. We present SLIC — Small LM Iterating on Circuits — a reinforcement fine-tuning framework that trains a 7B-scale, locally deployable language model for closed-loop analog circuit optimization. SLIC pairs structured circuit representations with discretized relative-update actions, domain- and difficulty-gated multi-objective rewards, and Reward-Guided Rollout Steering (RGRS), which turns per-step simulator feedback into both policy-learning and trajectory-steering signals in this sparse, highly constrained domain. We further release an open, goal-conditioned benchmark of 34 amplifier topologies with standardized protocols for in-distribution validation, held-out topology evaluation, and cross-technology-node transfer. Trained solely on a 22 nm silicon CMOS PDK, SLIC achieves 58.7% in-distribution Pass@1 under a 4-turn simulator budget, surpassing the 1.6T DeepSeek-V4-Pro Think baseline by 12.7 percentage points. Beyond the training distribution, it reaches 51.5% Pass@1 on held-out topologies, improving over the same baseline by 10.6 percentage points, and consistently transfers across five unseen technology nodes (32–130 nm) without additional fine-tuning. Performance further scales with test-time simulator interactions, reaching roughly 70% Pass@1 at an 8-turn budget, demonstrating that SLIC enables the small policy to convert additional simulator calls into improved circuit designs.
PaperID: 5578, Poster
Abstract: Source-Free Domain Adaptation (SFDA) aims to adapt a source-pretrained model to an unlabeled target domain without access to the original source domain. While early single-model approaches rely on self-refinement, they are inherently susceptible to confirmation bias and struggle to correct their own systematic errors. To overcome this limitation, recent methods introduce Vision-Language (ViL) models as external knowledge sources. However, these approaches operate in a largely unidirectional paradigm, using the ViL model primarily to supervise the source-pretrained model. This overlooks a key structural property: the two models exhibit distinct failure modes -- where one produces an incorrect prediction, the other may produce a correct one, creating a natural opportunity for mutual correction within the target domain. Yet, without ground-truth labels, identifying which model is correct on any given sample is non-trivial, and naively exchanging predictions risks propagating errors across models. To address this challenge, we propose SafeCut, a novel approach that leverages the cut statistic as a label-free measure of prediction reliability to gate cross-model supervision. Our approach dynamically controls both the direction and strength of supervision based on relative reliability, selectively amplifying true corrections while suppressing miscorrections on a per-sample basis. We further provide theoretical justification showing that this reliability-gated mechanism guarantees a net-positive correction signal. Extensive experiments across diverse SFDA benchmarks demonstrate that SafeCut achieves state-of-the-art performance, highlighting the effectiveness of safeguarding mutual correction in SFDA via cut statistics. Codes are available https://anonymous.4open.science/r/SafeCut
PaperID: 5579, Poster
Abstract: Symmetric organization is a fundamental principle of biomolecular complex assembly. The same principle underlies engineered applications, from vaccine platforms to nanomaterials, making symmetric de novo generation a core capability for computational protein design. We establish a theoretical foundation equating generation of an assembly with underlying symmetry group G to generation within a quotient space over G. This directly allows \mathrmSE(3)-equivariant models to generate symmetric structures with minimal distribution shift. With this insight we introduce \ours (Zero-shot Unified Symmetrization), a framework that converts any pretrained \mathrmSE(3)-equivariant backbone generative model into an efficient generator of symmetric assemblies with no retraining. \ours operates on a single \emphcanonical subunit and utilize \emphphantom subunits to preserve internal model operations, reducing memory and runtime cost by \mathcalO(|G|) for finite groups G, compared to a full-complex forward pass. We implement \ours for IPA-based models (FrameDiff, FoldFlow) and axial-attention-based RFdiffusion, showing high generation efficiency with minimal degradation to generation quality. We then apply \ours to generate symmetric complexes of over 10,000 residues with exact, untruncated attention, a scale previously intractable with any model.
Abstract: Video Diffusion Transformers (VDiTs) have demonstrated significant capabilities in high-fidelity video generation. However, their ability to produce long-duration videos is fundamentally constrained by the quadratic complexity of the self-attention mechanism. Recent clustering-based sparse attention methods improve the quality-speed trade-off by grouping semantically similar tokens, but their practical efficiency remains limited by two bottlenecks: substantial clustering overhead and low CTA utilization caused by irregular cluster-induced blocks. We propose HyperVAttention (HVA), a training-free sparse attention framework that addresses both bottlenecks jointly. To reduce clustering overhead, we introduce 3D local-window clustering, which exploits the spatio-temporal locality of video tokens to restrict centroid search to fixed local neighborhoods, and implement it with a custom Triton kernel for efficient execution. We further propose a hybrid clustering strategy that performs full clustering only at anchor steps and updates only subset tokens at intermediate steps, leveraging the temporal stability of cluster assignments across denoising steps. To improve CTA utilization, we present hardware-aware cluster merging that minimizes CTA-aligned execution cost through parallel agglomerative merging, improving block density and approximation fidelity by utilizing idle tile capacity. Together, these components reduce clustering overhead, avoid redundant updates, and better align sparse attention with the fixed tile structure of modern GPU kernels. Experiments on Text-to-Video generation show that HVA establishes a new Pareto frontier for training-free sparse attention in video diffusion, reducing end-to-end latency by up to 2.13× while improving fidelity over existing training-free sparse attention baselines.
PaperID: 5581, Poster
Authors: Chengtao Jian, Kai Yang
Abstract: Direct Preference Optimization (DPO) minimizes empirical risk over observed preference comparisons, implicitly assuming that the training preference distribution matches the deployment preference distribution. In practice, this i.i.d. assumption often fails, as preference distribution shift exhibits a natural hierarchical structure across three levels: group-level shift, prompt-level shift, and conditional preference-level shift. Such multi-level distributional mismatch poses a critical challenge for safe LLM deployment, yet existing robust DPO methods typically model either overall distributional uncertainty or a single robustness level, leaving the hierarchical structure of OOD preference shifts underexplored. We propose TPO, a tri-level distributionally robust framework for OOD preference optimization, which jointly models all three levels of distributional uncertainty through a hierarchical rectangular ambiguity set. By instantiating it with Wasserstein distance, we derive a tractable and interpretable objective that combines a closed-form preference-level margin penalty, prompt-level adversarial embedding perturbation, and worst-case group reweighting, each governed by an independent radius. We establish finite-sample convergence guarantees and empirically validate the effectiveness of \method over existing robust DPO methods.
PaperID: 5582, Poster
Abstract: A fruitful approach towards understanding generalization in machine learning involves characterizing how models separate generalizable structure from idiosyncratic or even corrupted features of training data. Fisher information has emerged as a diagnostic tool for this separation, leveraging loss geometry to distinguish generalization from memorization. Here we investigate this connection in vision models trained on CIFAR-10 with controlled label noise that can be learned only through memorization. We show that projecting model weights onto the top Fisher eigenvectors decouples general task performance from noisy sample memorization, and we identify the layer in which this decoupling emerges. We show that the same phenomena are present in small transformers trained to perform in-context learning and reproduced in closed form in a simple, analytically solvable linear model. Our results reveal that Fisher information captures aspects of the geometric structure underlying how neural networks allocate capacity between generalization and memorization.
Abstract: Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings remains prohibitively difficult: long-horizon interactions, irreversible actions, and stochastic environment feedback make both human annotation and Monte Carlo estimation infeasible at scale. In this work, we show that reinforcement learning (RL) post-training already provides the ingredients for effective step-level scoring, eliminating the need for dedicated reward model training altogether. Concretely, we derive an implicit advantage under a general stochastic Markov Decision Process, which we term progress advantage---log-probability ratio between the RL-trained policy and its reference policy exactly recovers the optimal advantage function. This signal is annotation-free, domain-agnostic, and available as a byproduct of the standard RL post-training pipeline. We validate the effectiveness of the progress advantage across three different applications: test-time scaling, uncertainty quantification, and failure attribution on five benchmarks and four model families. Across all settings, progress advantage consistently outperforms confidence-based baselines and, despite requiring no task-specific training, surpasses dedicated trained reward models. We complement these results with deeper analyses on characteristics of progress advantage, offering practical guidance for adoption in real-world agentic systems.
Abstract: More than 80% of the 1.6B English speakers do not use Standard American English (SAE), yet LLMs often fail to correctly identify non-SAE dialects and generate stereotyped responses for their speakers. We introduce , the first large-scale framework for generating high-quality multi-dialectal conversational data encompassing the three pillars of written dialect---lexical (vocabulary), orthographic (spelling), and morphosyntactic (grammar) features. DialectLLM produces a dialect-parallel dialog dataset spanning nine English dialects. Partnering with native linguists, we design and validate SAE-to-dialect transformation rules, ensuring authenticity. Our approach challenges the prevailing practice of applying a single morphosyntactic feature set to both user utterances and model responses, showing that models should not reproduce up to 90% of the grammatical features of a dialect. Human evaluation confirms data quality, with annotators preferring DialectLLM over prior methods in 98.8% of pairwise comparisons for dialect naturalness. We then construct , a dialect-parallel benchmark with 50k+ dialogs, resulting in 97k+ QA pairs, and evaluate 17 LLMs on dialect identification and response generation tasks. Even frontier models achieve under 70% accuracy, fail to reach 50% for prominent dialects like Canadian English, and systematically misclassify non-SAE dialects as American or British. Beyond benchmarking, we show that DialectLLM data also serve as a scalable LLM post-training resource, suggesting a practical path toward dialect-aware conversational AI.
PaperID: 5585, Poster
Authors:
Heng Zhang, Chengyu Zhou, Jiajun Wu, Xinyuan Liu, Liheng Zhang, Yueqi Guo, Rui Liu, JiaHao Hong, Xuanxun Lian, Jinpeng Lu, Jin HuangAbstract: Reflection-based language agents improve reasoning by using feedback to revise previous attempts across multiple trials. Most existing reflection frameworks scale this process by \emphdepth, allocating additional test-time compute to longer serial reflection-retry chains. However, we observe that serial reflection often saturates after rapid early gains. Under controlled reflection budgets, increasing depth from 5 to 10 cycles brings only marginal improvement, yet isolated reflective workers still contain useful revision signals beyond the saturated chain. This raises a central question: can reflective reasoning scale by \emphwidth rather than only by depth? A direct answer is non-trivial, since naive parallel reflection often inspects similar aspects of the same failed attempt and produces long or noisy revision contexts after aggregation. We propose \ourmethod, a reflective width scaling framework that learns to allocate complementary reflective scopes, execute workers in isolated contexts, and compress parallel feedback into a structured PolyBrief. \ourmethod trains the planner, workers, and reducer with verifiable outcome feedback and a role-balanced objective that prevents wider rollouts from dominating optimization. Across code generation, mathematical reasoning, and multi-hop question answering, \ourmethod consistently improves performance under matched reflection budgets, outperforming the strongest depth-reflection baseline by an average of 6.1%. Further analyses show that \ourmethod closes about 80% of the oracle-naive width gap, reduces redundant worker outputs, and improves revision success with fewer tokens. These results suggest that reflection should be understood not only as a longer retry process, but also as a depth-width compute allocation problem.
Abstract: The advancement of Embodied AI necessitates high-quality simulation assets that faithfully mirror the physical world. However, transforming raw visual observations into simulation-ready scenes remains challenging due to the lack of physical grounding and scene-level interactivity in current image-to-URDF methods. We propose , a framework that reformulates monocular scene reconstruction as a procedural programming task for interactive 3D environments. By leveraging the zero-shot reasoning and code synthesis of MLLMs, PRISM translates a single RGB image into executable programs defining object geometry, articulation, and physical properties. To ensure simulation readiness, it incorporates a mechanism that iteratively refines the generated programs by validating their execution in a physics engine. This feedback loop enforces physically plausible object articulations, valid object compositions and interactions, and accurate spatial relationships. Experiments demonstrate that PRISM significantly outperforms open-loop approaches and prior monocular reconstruction models. Notably, PRISM-generated scenes support complex downstream tasks such as stable stacking and fine-grained manipulation, which are difficult to achieve with existing methods.
Abstract: Tabular foundation models achieve strong accuracy on choice prediction tasks, but their predictions often violate the economic logic those tasks require: raising a price can increase predicted demand, implied willingness-to-pay estimates are frequently negative or implausible, and unavailable alternatives receive nonzero probability. We propose a two-stage adapter that takes a foundation model's predicted choice probabilities as a precomputed feature and embeds them inside a multinomial logit's utility. In Stage 1, we fit the multinomial logit's structural coefficients by maximum likelihood with sign constraints; in Stage 2, we freeze those coefficients and fit a small neural correction operating on the foundation model's predictions. We prove that this composition exactly preserves the multinomial logit's marginal rate of substitution, so analytically computable value-of-time becomes a mathematical guarantee rather than an empirical accident. Across three datasets and two foundation models, the adapter gains 6.4 percentage points (pp) of test accuracy on average over the multinomial logit and up to 12.8 pp, maintains 100% cost monotonicity, and produces values of time within the published transportation-economics range on the transportation datasets. Performance degrades gracefully under foundation-model context restriction, retaining at least 6 pp of accuracy gain even at 10% of the original foundation-model context.
PaperID: 5588, Poster
Authors: Mohammed R Mohammad, Ardhendu Behera, Sandip Pradhan, Swagat Kumar, Yonghuai Liu
Abstract: Few-shot adaptation of vision--language models must balance two competing goals: learning fine-grained visual distinctions from very limited supervision, and preserving the efficiency that makes frozen CLIP-like models attractive for deployment. Existing lightweight adaptation methods, including prompt tuning, LoRA, and cache-based adapters, typically supervise global image embeddings and therefore underuse token-level structure. Conversely, structure-aware methods can exploit local evidence but often introduce additional inference-time computation. We propose SD-LoRA, a structural distillation framework that uses token-level graph reasoning only during training and absorbs its effect into a lightweight LoRA-adapted CLIP model. During training, a heterogeneous graph teacher models patch-patch, global-local, and visual-text relations over high-resolution ViT tokens. Rather than keeping this graph at test time, SD-LoRA transfers its structural bias into the shared CLIP-LoRA backbone through implicit distillation and Prototype Predictive Alignment, which aligns graph-conditioned representations with support-derived class prototypes. At inference, the entire graph teacher is discarded; prediction uses only a single CLIP-LoRA forward pass and a cache-based classifier. Across 11 few-shot benchmarks, SD-LoRA improves over a strong CLIP-LoRA baseline by 1.3/1.2/2.4/1.9/2.0% in the 1/2/4/8/16-shot settings, while adding no graph computation at test time. Further results on base-to-novel generalization, cross-dataset transfer, OOD robustness, and calibration show that training-time structural supervision improves not only accuracy but also transfer reliability.
PaperID: 5589, Poster
Authors:
Zhongxing Xu, Zhonghua Wang, Zhe Qian, Shiyan Su, Ming Hu, Xiaocheng Zou, Wei Feng, MINGQUAN LIN, Yuyin Zhou, Yifan Peng, Hamid Rezatofighi, Zongyuan GeAbstract: Recent advancements in multimodal diffusion language models (MDLMs) have exhibited strong global modeling capabilities through iterative refinement. We observe that mask uncertainty follows an overall downward trend during iterative denoising, while a few denoising steps exhibit localized entropy spikes closely associated with hallucinations. We argue that anomalous steps can be identified by measuring the entropy deviation of each denoising step from its neighboring temporal window. We hypothesize that discrete token commitment discards the distributional information at masked positions, limiting the continued interaction between candidate semantics and the context during entropy spikes. With this goal, we present StepRefine, a plug-and-play decoding strategy that leverages semantic context to achieve reliable denoising. The core of our method lies in a step-aware switch between discrete denoising and latent refinement. The model employs probability-weighted embeddings at entropy spike steps to preserve diverse reasoning hypotheses and performs multiple rounds of continuous refinement. Meanwhile, a refinement controller monitors distributional stability to reduce unnecessary loops and switch the model back to standard discrete denoising. Through extensive experiments, StepRefine demonstrates significant hallucination mitigation across different MDLMs on multimodal benchmarks.
PaperID: 5590, Poster
Abstract: Reinforcement learning (RL) rewards are notoriously difficult to design and control, often leading to the model learning unintended behaviors such as eward hacking. One potential solution is to monitor for reward hacking and penalize it when detected; however, training against a monitor could lead to evasive behavior, and our general understanding of how to apply monitors effectively during training is limited. To study how best to use monitors to mitigate reward hacking, we introduce and open source three realistic environments where Qwen3-4B reward hacks: a coding environment hackable via test overwriting, a medical chat environment hackable via sycophancy, and a biography generation environment hackable via hallucination. We first focus on the coding environment, where we find that: (1) models can learn to evade highly accurate monitors by exploiting systemic flaws in probes and LLM judges; (2) monitors that leak more learning signal during RL suppress reward hacking but are more often evaded; and (3) including easier problems in training can decrease reward hacking. We apply our findings to build better reward hacking monitors for the medical chat and biography generation environments that improve upon naive baselines to reduce reward hacking rates across seeds from 70-100% to 0%. Our results demonstrate that our takeaways translate to new settings and that better monitor intervention designs are possible.
PaperID: 5591, Poster
Abstract: Time-series classifiers deployed in risk-sensitive domains such as healthcare, wearable sensing, and industrial monitoring must be accurate, interpretable, and capable of abstaining when evidence is weak. No existing method delivers all three, and the obstacle is structural, rather than logistical. Meaningful temporal patterns cannot be defined without knowing segment boundaries, yet meaningful boundaries cannot be placed without knowing which patterns to expect. Existing approaches break this circularity through fixed windows or expert-annotated vocabularies, sacrificing at least one of the three properties. We introduce ConceptTime, which closes the loop through a single self-supervised signal. A frozen probabilistic forecaster predicts a distribution over the next observation at every timestep of the sequence. We call the mismatch between its prediction and the realized signal, the predictive surprise. This signal both locates segment boundaries and characterizes each resulting segment through the forecaster's own predictive statistics. Segments are thus grounded by \emphhow they behaved relative to expectation, not by what they look like. These summaries cluster into a vocabulary of dynamical regimes that we call concepts. The vocabulary distinguishes calm-after-spike from calm-after-calm, and unifies visually distinct segments that violate expectation in the same way. A lightweight head classifies from the concept sequence, and the distance to the nearest concept yields per-input reliability for free. Across seventeen UEA, human-activity, biomedical, and fault-diagnosis benchmarks, ConceptTime matches state-of-the-art black-box accuracy, beats every interpretable baseline, and retains near 98% accuracy on Epilepsy with 1% of labels. It equals or exceeds softmax, prototype-distance, and SHAP/TimeX/TimeX++ on OOD detection and deletion-faithfulness, without a single annotation, language-model call, or domain expert.
Abstract: Reinforcement learning with verifiable rewards (RLVR) generates thousands of tokens per training step, with rollout generation dominating the computational cost. The overall token budget can be controlled along two main dimensions: (i) deciding which prompts to allocate rollouts to, and (ii) deciding how long each rollout should be. Prior work has generally controlled only one of these dimensions at a time. We show that jointly tuning both decisions under a shared compute budget improves both reasoning quality and wall-clock training time. We instantiate this view as DUal-controlled tokEn allocaTion (DUET), a computationally efficient layer over GRPO, in which we have a lightweight pre-rollout surrogate of prompt informativeness to set how many rollouts each prompt receives and a marker-gated abort rule with importance reweighting set when to stop them. On Qwen3-1.7B trained on MATH, DUET outperforms full-budget GRPO and the other three budget-aware baseline methods. DUET's advantage was further generalized to other benchmarks across math and coding, and was on par with the best baseline on the scientific Q\&A domain, while also achieving a 1.62× wall-clock speedup. More notably, using only 50% of the token budget, DUET still outperforms all baseline methods, and performance often increases rather than reduces with the more limited budget, achieving an even higher 2.51× speedup. We verified the high performance of DUET on other backbone LLMs, including Qwen-4B and Llama-3.2-3B-Instruct. Notably, the gap between DUET and the strongest baseline \emphwidens as the budget tightens, contrary to the usual pattern in which efficient methods trade off quality as compute decreases. More broadly, these results suggest that DUET budget-aware control strategies are valuable not only for accelerating training, but also for improving the quality of the learning signal.
PaperID: 5593, Poster
Authors: Guantao Zhao, noorizadegan, Shihao Yang, Nicoleta Serban
Abstract: The simplex method is one of the most widely used algorithms for linear programming, but its practical efficiency is highly sensitive to pivot selection and difficult to predict due to structural uncertainty in the search process. We propose EpiPivot, which casts pivot selection as a controllable decision process under epistemic uncertainty. EpiPivot identifies this uncertainty as two-dimensional: temporal dependence across the optimization trajectory and joint uncertainty across candidate pivot rules, and addresses both through temporal attention and an Epistemic Neural Network with shared latent sampling. At decision time, Thompson sampling combines these uncertainty-aware predictions into robust pivot choices. EpiPivot achieves significant improvements over classical and learning-based baselines, reducing pivot cost by up to 79% over the best-performing classical rule on large-scale instances where most classical rules fail to converge within the time limit.
PaperID: 5594, Poster
Abstract: Modeling interacting dynamical systems requires capturing spatial interactions alongside long-range temporal dependencies. Graph neural networks (GNNs) provide a natural representation but typically rely on autoregressive rollouts and treat spatial and temporal dynamics separately, leading to error accumulation over long horizons. Existing approaches also focus on local interactions and short temporal contexts, limiting their ability to capture multi-hop dependencies and global structure. We introduce the Graph Mamba Operator (GraMO), a latent-space simulator that integrates state-space models with graph-based interaction learning. In contrast to prior work that sequences nodes or applies spatial and temporal updates in separate stages, GraMO couples graph-based interactions and temporal state updates within a single recurrence. The update is linear in the latent state, with input-dependent coefficients that adapt across regimes. We evaluate GraMO on N-body systems, motion capture, and robotics datasets, achieving the lowest error across benchmarks and the largest gains in long-horizon prediction.
PaperID: 5595, Poster
Abstract: Vision-language models map images and text into a joint embedding space. However, these embeddings often entangle multiple semantic features, which limits their interpretability and controllability. While sparse autoencoders have emerged as a useful tool for decomposing these embeddings into monosemantic features, their application to joint embedding spaces has largely relied on an implicit, untested assumption that semantically corresponding features share the same directions across modalities. In this paper, we challenge this assumption by identifying discrepancies in feature directions for the same concept across image and text modalities, a phenomenon we term , where a shared concept activates different latents depending on the modality. This finding further reveals why aligning latent activations alone is insufficient to resolve the underlying feature mismatch. To address this misalignment, we propose an approach that trains sparse autoencoders to preserve the unique feature geometry of each modality and aligns corresponding features post hoc. Our method improves reconstruction fidelity and enhances performance in cross-modal retrieval and concept steering.
Abstract: Skill-augmented agents increasingly rely on large reusable skill libraries, but retrieving relevant skills is not the same as presenting usable context. Existing methods typically return atomic skills or dependency-aware bundles whose internal roles remain implicit, leaving the agent to infer the execution entry point, support skills, visible requirements, and failure-avoidance guidance. We introduce ), an inference-time group-structured retrieval method that changes the agent-facing retrieval object from a flat skill list to a compact, role-labeled execution context. builds anchor-centered skill groups from a typed skill graph, expands support groups through a group graph, bottlenecks the selected group plan into a bounded set of atomic skill payloads, and renders a fixed execution contract with fields, without changing the downstream agent, skill payloads, or execution environment. Experiments on SkillsBench and ALFWorld show that \goskills preserves visible-requirement coverage under a small skill budget, improves over flat skill-access baselines, and often improves reward and agent-only runtime relative to structural retrieval references. Code is available at https://anonymous.4open.science/r/Group-of-Skills-E861.
PaperID: 5597, Poster
Abstract: Variational approaches in deep learning promise to deliver improved uncertainty quantification and out-of-distribution generalization by inferring a distribution over weights, regularized by a divergence to a chosen prior. However, prior elicitation presents a significant challenge in deep learning, resulting in the common practice of significantly downweighting the regularization term in the variational objective. Recently, it was proposed to train solely via the expected loss by relying on implicit regularization from the optimizer, rather than explicit regularization from the divergence to the prior. While this approach largely sidesteps the issue of prior elicitation, current theory lacks an understanding of how key optimization hyperparameters impact the implicit regularization. Here, we theoretically and empirically demonstrate that larger learning rates and fewer parameter samples lead to flatter minima of the loss landscape, which are known to generalize better. We show that training on the expected loss operates at a modified edge of stability, characterized by the signal-to-noise ratio of the gradient estimates and the number of Monte Carlo samples. Experiments on out-of-distribution benchmark datasets confirm the practical significance of the theoretical results.
Abstract: Modern deep learning architectures increasingly contend with sophisticated signals that are natively infinite-dimensional, such as time series, probability distributions, or operators, and are defined over irregular domains. Yet, a unified learning theory for these settings has been lacking. To start addressing this gap, we introduce a novel convolutional learning framework for possibly infinite-dimensional signals supported on a manifold. Namely, we use the connection Laplacian associated with a Hilbert bundle as a convolutional operator, and we derive filters and neural networks, dubbed as HilbNets. We make HilbNets and, more generally, the convolution operation, implementable via a two-stage sampling procedure. First, we show that sampling the manifold induces a Hilbert Cellular Sheaf, a generalized graph structure with Hilbert feature spaces and edge-wise coupling rules, and we prove that its sheaf Laplacian converges in probability to the underlying connection Laplacian as the sampling density increases. Notably, this result is a generalization to the infinite-dimensional bundle setting of the Belkin & Niyogi \citeBELKIN20081289 convergence result for the graph Laplacian to the manifold Laplacian, a theoretical cornerstone of geometric learning methods. Second, we discretize the signals and prove that the discretized (implementable) HilbNets converge to the underlying continuous architectures and are transferable across different samplings of the same bundle, providing consistency for learning. Finally, we validate our framework on synthetic and real-world tasks. Overall, our results broaden the scope of geometric learning as a whole by lifting classical Laplacian-based frameworks to settings where the signal at each point lives in its own Hilbert space.
Abstract: Graph representation learning has become a standard approach for analyzing networked data, with latent embeddings widely used for link prediction, community detection, and related tasks. Yet a basic design choice, the latent dimension, is still treated as a brittle hyperparameter, fixed before training and tuned by held-out performance. Learned factors are also identifiable only up to rotation and rescaling, so the nominal rank rarely coincides with the quantity that governs model behavior. We propose Spectral Prefix Extraction and Capacity-Targeted Representation Analysis (Spectra), which replaces rank as the unit of analysis with the spectrum of a learned positive semidefinite kernel, trace-normalized so that spectra are comparable across fits. The normalized eigenvalues form a distribution on the simplex, and their Shannon effective rank acts both as a summary of learned capacity and as a controllable training-time coordinate: a single scalar shapes this realized dimension during training, and bisection targets any desired value within the rank cap. To theoretically support that, we show local regularity and monotonicity of the realized-dimension profile. Across collaboration, social, biological, and infrastructure networks, Spectra traces performance--capacity frontiers that make the trade-off between predictive accuracy and realized dimension visible. It performs competitively with strong link-prediction baselines, yields aligned lower-capacity views of the same fitted model through spectral prefixes, and provides a principled handle on capacity in the overparameterized regime. Capacity thus becomes a property of the fitted model rather than a hyperparameter of the training.
PaperID: 5600, Poster
Authors:
Pritish Chakraborty, Aditya Singh, Indradyumna Roy, Lokesh N, Himanshu Dutta, Arun Iyer, Yashoteja Prabhu, Gaurav Sinha, Manik Varma, Soumen Chakrabarti, Abir DeAbstract: Neural subgraph retrieval (NSR) models degrade under query distribution shifts, and retraining for each new distribution requires collecting fresh ground-truth labels, each of which entails an NP-hard subgraph-matching instance across a large corpus—making repeated label collection computationally prohibitive. We introduce ACROSS, an active corpus selection framework that, under a fixed solver budget, selects a small corpus subset per out-of-distribution query whose labels suffice for effective NSR training. We prove that allocating the entire budget to retrieving positives is optimal given extreme class imbalance. To enable efficient positive retrieval across arbitrary query distributions, we linearize corpus embeddings around a base model, apply randomized projections to eliminate hinge nonlinearities, and train distribution-shift-tolerant representations via adversarial parameter perturbation. The resulting corpus index is built once and reused across all distributions. Experiments on four molecular graph datasets show ACROSS consistently outperforms six baselines in downstream retrieval accuracy.
PaperID: 5601, Poster
Authors:
Luxury, Jie Huang, Zihao Fan, Xiaoxiao Ma, Yuming Li, Siming Fu, Jun-hao Zhuang, Zeyue Xue, Haoran Li, haoyang huang, Nan DuanAbstract: While recent autoregressive video diffusion models achieve remarkable streaming quality, they remain confined to low resolutions (e.g., 480P), leaving efficient, scalable, real-time high-resolution video generation a fundamental open challenge. To bridge this gap, we present Ultra Flash, a cascaded streaming framework capable of real-time high-resolution video generation. Ultra Flash achieves ~30 FPS at 1K resolution and ~18 FPS at 2K resolution on a single B200 GPU through three key contributions: (1) an architecture-preserving T2V-to-TV2V super-resolution training paradigm coupled with an AIGC-oriented data degradation pipeline that effectively preserves the generative capability of the base model, enabling enhanced high-resolution detail when cascaded after mainstream low-resolution generative models; (2) a causal streaming latent upsampler paired with a high-resolution decoder, which enhances spatiotemporal coherence while enabling efficient latent spatial scaling and precise high-resolution decoding with negligible computational overhead; and (3) a cascade high-resolution streaming video generation optimization scheme that first performs hybrid-reward-enhanced sparse causalization and single-step distillation of the super-resolution model, then introduces cascaded streaming self-forcing preference optimization with dynamic cache management, jointly enhancing overall coherence, improving quality, and enabling real-time high-resolution streaming video generation. Extensive experiments demonstrate that Ultra Flash reliably produces ultra-high-resolution streaming video while maintaining state-of-the-art visual quality and superior efficiency.
PaperID: 5602, Poster
Authors: Tao Zhang, Yun Dai, Ziqian Zeng
Abstract: Large Reasoning Models (LRMs) rely on extensive Chain-of-Thought generation but suffer from memory bottlenecks due to linear Key-Value (KV) cache growth. While low-bit quantization offers a potential solution, we identify that current static quantization strategies suffer from severe degradation on complex reasoning tasks at extreme bit-widths. Notably, this accuracy drop is accompanied by increased generation length, which ironically limits the intended efficiency gains. We attribute this performance collapse to the fact that static quantization methods are insufficiently flexible to meed dynamic demands of precise long-range retrieval. Utilizing Average Attention Distance (AAD) to quantify retrieval patterns, we find that the attention patterns in shallow layers exhibit strong locality, whereas that in deep layers demonstrate intensive long-range dependencies. Furthermore, deep layers have fluctuating AAD values across token positions. Based on above findings, we propose AdaKVQ, an adaptive mixed-precision quantization framework that modulates the bit-width allocation guided by the dynamic demands of long-range context retrieval in each layer. Extensive experiments on complex reasoning benchmarks with Llama-3.1 and Qwen2.5 demonstrate that AdaKVQ consistently outperforms state-of-the-art baselines, achieving performance comparable to full-precision models.
PaperID: 5603, Poster
Abstract: The prevailing paradigm in large language model (LLM) alignment operates via erasure, filtering unsafe data or training models to strictly refuse harmful prompts. While effective at reducing immediate toxicity, this approach fundamentally constricts the model's epistemological scope, resulting in over-cautious systems that output uninformative blanket refusals to sensitive yet benign queries. In this work, we challenge the orthodoxy that unsafe data must be discarded. We propose a dialectical approach to alignment, positing that unsafe data encodes rich, domain-specific knowledge critical for nuanced, safe, and informative generation. To operationalize this, we introduce , a Mixture-of-Experts (MoE) framework that isolates unsafe knowledge into domain-specific Low-Rank Adapters (LoRA experts) trained exclusively on harmful corpora. To synthesize safety from these unsafe primitives, we train a lightweight gating network using a minimal, highly curated set of safe-informative responses (as few as 100 samples). During inference, this router dynamically orchestrates the unsafe experts, effectively steering the generation trajectory to harness their deep domain knowledge while strictly enforcing safety constraints. Extensive empirical evaluations across stringent safety benchmarks demonstrate that is not only safer, achieving over a 20% relative improvement in safe response rate (more than a 15% absolute gain), but also produces more informative responses when safety and harmfulness are of paramount concern. Furthermore, the routing mechanism exhibits strong zero-shot generalization to unseen domains and broader safety tasks without domain-specific supervision. Our findings suggest a paradigm shift in alignment: true safety requires not the masking of unsafe knowledge, but its controlled integration.
Abstract: Frontier LLMs can perform multi-step reasoning over content-free filler tokens like dots or counting sequences, producing correct answers with no visible chain-of-thought (CoT). This has been cited as a limit case for CoT monitorability: if the surface tokens carry no information about the reasoning, behavioral oversight has nothing to read. But hidden from the output is not the same as hidden from us. On three task families spanning fact retrieval, parallel numeric composition, and string manipulation, two open-weights frontier models (DeepSeek V3, Kimi K2) compute over filler tokens in a structured, legible way. Attention routes the question through the filler region to the answer, KV-cache transplants at only filler positions causally swap the model's output between examples, and logit-lens readouts reveal a temporal structure in which retrieved facts appear early and their composition crystallizes in late layers. We then introduce an unsupervised decoding pipeline that takes only hidden states as input and recovers intermediate values with 80–95% accuracy with the strongest LLM judge across both models and all three tasks, without ground-truth labels or training. Hidden computation that defeats behavioral CoT monitoring is, on these tasks, directly readable from the residual stream. This suggests that monitorability should be understood as a property of the model's full computational trace, not just its surface tokens.
PaperID: 5605, Poster
Abstract: We study regret minimization in finite-horizon episodic Markov Decision Processes (MDPs). While minimax-optimal algorithms are known, tractable approaches to instance-dependent optimality are still lacking. Motivated by this gap, we introduce \pi_0-PSRL, a variant of PSRL that uses posterior samples to decide when exploration is needed, while following a fixed reference policy \pi_0 during exploration episodes. This decouples the test for exploration, triggered when the sampled and empirical MDPs have different optimal policies, from the choice of the policy used to gather information. The resulting design addresses a limitation of standard PSRL, where the sampled optimal policy may not be the optimal choice for exploration. We prove instance-dependent regret bounds for \pi_0-PSRL, identifying the logarithmic exploration cost induced by \pi_0 and taking a step toward matching asymptotic lower bounds for episodic RL. Our proof techniques showcase a novel proof structure to derive problem-dependent regret bounds in episodic MDPs, and concentration results for Dirichlet random variables, that may be of independent interest.
PaperID: 5606, Poster
Abstract: Feed-forward transformer models have driven remarkable progress in 3D vision, yet their quadratic complexity renders them impractical for long video sequences. Streaming 3D reconstruction addresses this bottleneck by processing frames sequentially with constant memory. However, existing recurrent architectures inevitably suffer from progressive degradation over long sequences, as they fail to account for the intrinsic reliability of observations and the accumulated confidence of the historical state, leading to severe error accumulation. To address these issues, we introduce BAT3R, a training-free Bayesian adaptation framework comprising two synergistic modules: State Innovation Gating (SIG) and Bayesian-guided Adaptive State Updating (BASU). SIG acts as a distributional regularizer that anchors state evolution to its geometric initialization, effectively arresting representational drift during long-term inference. Building on this, BASU reformulates state update as a recursive Bayesian inference process within a Kalman filtering framework. By jointly modeling the temporal reliability of observations via information-theoretic gain and spatial relevance via cross-attention maps, BASU achieves an asymmetric spatio-temporal state fusion: well-consolidated historical regions are protected from redundant updates, while genuinely novel scene content is aggressively incorporated. Extensive experiments on standard benchmarks (7-Scenes, TUM, ScanNet, KITTI, etc.) demonstrate that BAT3R consistently outperforms state-of-the-art online 3D reconstruction methods, with the most significant gains observed on long sequences and complex scenes.
PaperID: 5607, Poster
Abstract: (SCMs) through the lens of optimal transport. In SCMs, a factual observation generally induces a posterior distribution over exogenous variables, making counterfactual inference inherently distributional. We introduce two complementary frameworks that transport this posterior onto the set of exogenous configurations satisfying a desired counterfactual query. The first constructs a natural pointwise projection and is suitable when the transported exogenous distribution can be used directly. The second enforces mutual independence among counterfactual exogenous variables, so that the transported variables remain interpretable as genuine exogenous noises of the same SCM. We show that both frameworks are consistent with the original backtracking counterfactual principle, and we propose marginal, penalty-based, and hard-constrained solvers for finite-posterior settings. Synthetic experiments on non-bijective SCMs illustrate the geometry of the problem and reveal trade-offs among transport cost, counterfactual constraint satisfaction, and preservation of exogenous independence.
PaperID: 5608, Poster
Abstract: Multimodal alignment under scarce paired supervision aims to align independently pretrained unimodal encoders using limited image--text pairs, offering a practical alternative when large-scale joint pretraining is infeasible. Such scenarios call for lightweight objectives that go beyond local correspondences and capture distributional mismatch between heterogeneous embeddings. Existing methods often derive supervision from limited pairs through contrastive objectives or geometry-preserving regularization, encouraging instance-level correspondence or structural consistency but providing limited direct supervision on global cross-modal mismatch. This motivates a distribution-level objective that complements paired supervision by comparing image and text embedding distributions. In this paper, we propose Spectral Characteristic Distribution Alignment (SpecAlign), a lightweight spectral alignment framework that reduces cross-modal mismatch by comparing image and text characteristic functions over learnable spectral queries. Since characteristic functions uniquely determine probability distributions, SpecAlign provides a principled signal for capturing multi-scale discrepancies beyond pointwise alignment. To obtain informative and stable queries, SpecAlign employs a Structured Direction--Radius Sampler (SDRS) that decouples direction and radius modeling. Extensive evaluations on transfer, retrieval, robustness, and representation analyses demonstrate consistent gains, validating spectral distribution-level supervision for data-efficient multimodal alignment.
Abstract: As interactive generative systems are increasingly deployed in real-world applications, their tendency to generate unreliable or false responses raises serious concerns. Conformal abstention mitigates this risk by ensuring that the system answers only when confident. However, real-world deployments typically provide only partial user feedback (e.g., thumbs up/down) on the selected response and often operate in non-stationary or adversarial environments, for which effective learning methods are largely missing. To bridge this gap, we propose ExAUL, a novel online learning framework for conformal abstention with adversarial and partial feedback. Technically, we introduce (i) a novel conversion lemmathat translates the regret of any bandit algorithm into an FDR bound, and (ii) feedback unlocking, a strategy that exploits the structure of conformal abstention to extract additional learning signals from partial feedback. We prove that ExAUL achieves a regret bound of \mathcalO(\sqrtT \ln |\mathcalH|), which translates into an \mathcalO(\sqrtT) bound on FDR risk control, matching the controllability of full-information settings despite receiving only partial feedback. While applicable to general generative tasks, we demonstrate the efficacy of ExAUL for ensuring the reliability of Large Language Models (LLMs) through empirical validation on question-answering tasks across diverse non-stationary and adversarial settings. Our results demonstrate that ExAUL robustly controls the FDR while maintaining competitive answering coverage.
Abstract: Current Audio-Visual LLMs (AVLLMs) struggle with reasoning over videos featuring multi-speaker dialogues. In such videos, resolving "who says what" is crucial, which necessitates trimodal (text-audio-visual) binding. Motivated by these challenges, we systematically investigate how this trimodal binding is achieved in AVLLMs. Specifically, we identify emergent symbolic trimodal binding mechanisms in AVLLMs that utilize modality-specific symbolic variables. By encoding auditory and visual components into symbolic variables-capturing temporal utterance sequences and spatial entity coordinates, respectively-the model establishes cross-modal linking within this abstract space. Crucially, we reveal that when trimodal binding fails, the breakdown predominantly stems from misaligned audio-visual connections. To overcome this bottleneck, we introduce an audio-visual prompting method utilizing an off-the-shelf Active Speaker Detection (ASD) model. By simply overlaying visual bounding boxes on active speakers, this training-free approach yields immediate performance gains across three conversation-centric benchmarks. Moreover, lightweight fine-tuning of fewer than 300 steps on these ASD-prompted-videos extends these gains to three general AV benchmarks, suggesting the generalizability of our method.
Authors: Theodoros Aivalis, Iraklis A Klampanos, Antonis Troumpoukis, Joemon M Jose
Abstract: The rapid deployment of generative AI has amplified the critical need for Training Data Attribution to ensure transparency and accountability. However, current parametric approaches require computationally prohibitive access to model weights, while similarity-based methods ignore deep structural context. We propose a novel probabilistic framework that operates entirely in a black-box setting. Our method fuses continuous feature similarities with discrete, domain-specific Knowledge Graphs (KGs). This approach ensures the attribution is grounded in structural reality, explicitly rewarding highly specific historical samples while preventing generic background data from dominating the results. We evaluate our framework across two distinct domains where linking outputs to data and domain context is inherently complex: abstract artistic image synthesis and high-dimensional physical weather forecasting. Extensive benchmarking demonstrates the robust efficacy of our approach. In the artistic domain, it achieves a strong Linear Datamodeling Score that exceeds standard black-box similarity baselines, while closing much of the gap to gradient-based estimators. We additionally present a cross domain feasibility case study in environmental forecasting, where we use domain KGs to retrieve physically consistent historical analogs for regional flood forecasts, improving geographic localisation over a latent-only baseline. Operating entirely without internal model access, our approach provides an efficient, interpretable mechanism for post-hoc influence analysis and domain-grounded retrieval.
Authors:
Jin Kim, Jaeeun Lee, Claire Kim, Kyoungjin Oh, Paul H Cho, Jaewon Min, Yeji Choi, Jihye Park, Hyunhee Park, Park M Kyu, Hyungju Chun, Seungryong KimAbstract: Multi-view 3D reconstruction has achieved remarkable progress with the advent of feed-forward 3D reconstruction models. However, these models are typically trained and evaluated under ideal, degradation-free imaging conditions, whereas real-world observations often contain various degradations that differ significantly from such settings. Improving robustness for multi-view 3D reconstruction under degraded conditions therefore remains an important challenge. We present Geometry-Aware Representation Denoising (GARD), a novel framework that performs diffusion-based multi-view restoration directly in the feature space of a feed-forward reconstructor. This design exploits the geometry-aware feature representations of the reconstructor to effectively recover accurate scene geometry. Furthermore, by employing a decoder, the refined representations can also be used to restore high-quality RGB images, thereby enabling the simultaneous recovery of 3D scene geometry and high-quality imagery. Comprehensive experiments on the DA3 benchmark demonstrate the effectiveness of the proposed GARD framework. Our code and weights will be publicly released for full reproducibility.
PaperID: 5613, Poster
Authors: Rui Huang, Haojie Tao, Sen Gao, Qing Guo
Abstract: 3D Gaussian Splatting enables high-fidelity, real-time novel view synthesis, yet adapting it to dynamic, evolving scenes remains challenging. Existing methods rely on fixed static-dynamic partitioning, which suffers from persistent partitioning errors and cannot adapt during optimization. To tackle this issue, we introduce SynGS, a unified 3D Gaussian framework for continual scene updating and multi-view change detection. We cast static-dynamic decomposition as a learnable task, jointly optimizing per-Gaussian change probabilities and scene representations for iterative refinement. We further propose a change-guided update strategy that modulates illumination-aware photometric corrections via learned probabilities, separating transient appearance variations from authentic scene changes. Projecting Gaussian-level probabilities to the image plane yields 3D-consistent change masks, enabling mutual optimization between reconstruction and change localization. An illumination-decoupled refinement step further improves mask quality. Evaluations on standard benchmarks show that SynGS outperforms baseline methods on scene update and change detection tasks, while fully retaining the real-time rendering capability of 3D Gaussian Splatting. Code will be released upon acceptance.
PaperID: 5614, Poster
Abstract: We propose a novel zeroth-order (ZO) optimization method, ZO-F2, where Fisher information matrix (FIM)-based preconditioners are estimated in a low-variance manner via bilinear forms. Using FIM rather than Hessian allows for a larger number of estimator samples by pairwise combinations of perturbations. We further propose a numerically stable whitening-based update rule for full-matrix preconditioners, even when the FIM estimate is ill-conditioned. We provide theoretical analyses on the variance, the whitening residual, and convergence under standard and reasonable assumptions. Experimental results on black-box training of diverse models from scratch support our proposals and demonstrate the superiority of ZO-F2 over existing methods in terms of accuracy and computational efficiency.
PaperID: 5615, Poster
Abstract: Temporal action detection in long untrimmed videos still suffers from inaccurate temporal localization and low-quality proposals, especially under sparse temporal annotations and ambiguous event boundaries. In the context of open-ended video understanding and LLM-driven event analysis, reliable temporal action proposals have emerged as a critical interface bridging temporal localization, event annotation, and semantic understanding. However, conventional temporal action detection (TAD) pipelines lack explicit modeling of proposal localization quality, cross-representation support, and inter-proposal relations, making it difficult to consistently provide high-quality proposals. We propose Hybrid Rank-Calibration of Action Proposals (HyrCap), a proposal-centric query-based TAD framework for temporal event understanding. In HyrCap, Hybrid denotes the use of complementary video features that provide different temporal localization cues, while Rank-Calibration refers to calibrating the ranking confidence of dense action proposals so that it better reflects proposal localization reliability and candidate competition. HyrCap treats dense query-based predictions as an over-complete temporal event hypothesis space and improves the ranking reliability, localization quality, and candidate discrimination of action proposals through modules such as Cross-Representation Consensus Calibration, Localization Quality Estimation, and Proposal Duplication Suppression. Furthermore, inspired by the strong capability of LLMs in fine-grained semantic recognition, event description, and reasoning, we introduce HyrCap for Semantic Action Understanding: calibrated proposals can be used for standard TAD detection and as inputs for LLM-based semantic understanding. We further define a staged evaluation protocol to separately measure event localization ability, segment-level semantic understanding, and their combined event understanding performance. Experiments on public benchmarks show that HyrCap achieves state-of-the-art performance across multiple metrics and provides a more fine-grained paradigm and evaluation protocol for query-based event segmentation and annotation.
PaperID: 5616, Poster
Authors: Jonas Wallin, Sreekar Vadlamani
Abstract: We introduce mesh-invariant adaptive Markov chain Monte Carlo methods for Gaussian-process posteriors arising in infinite-dimensional Bayesian inference. In function-space MCMC, posterior distributions are defined by a change of measure with respect to a Gaussian prior, making absolute continuity essential for valid proposal construction. Standard adaptive schemes that modify means or scales in the full discretized space can destroy this property, leading to proposal measures that become singular in the infinite-dimensional limit. To avoid this, we adapt only an active finite-dimensional subspace of the Gaussian-process representation while preserving the prior dynamics on inactive coordinates. This yields two adaptive proposals, pCNLV and pCNMV, which extend preconditioned Crank--Nicolson and Crank--Nicolson Langevin methods by learning posterior scale, and in pCNMV also posterior mean structure, on the data-informed subspace without introducing discretization-dependent Gaussian density ratios. The resulting samplers retain the mesh robustness of function-space methods while improving efficiency through local adaptation. Experiments on a Darcy-flow inverse problem and Bayesian logistic regression demonstrate consistent efficiency gains, including an approximately fourfold improvement in effective sampling efficiency for Darcy flow.
Abstract: We present SceneBind an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio, and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial structure (i.e., where it is). SceneBind addresses this gap by representing each scene as a semantic-spatial entity, combining a global semantic embedding with object-centric semantic-spatial slots. This representation explicitly captures object-level semantics, spatial attributes, and uncertainty. We further propose SceneBind Matching, a semantic–spatial matching scheme that integrates global scene similarity with object alignment, supporting cross-modal scene retrieval and object grounding. To train and evaluate SceneBind, we curate a novel real-world binaural audio-visual dataset with structured semantic and spatial annotations, and propose a training protocol for aligning semantic and spatial signals across modalities. SceneBind is compatible with large-scale pretrained semantic encoders, adds lightweight spatial modeling with only a few additional tokens. It achieves state-of-the-art scene and spatial retrieval while enabling strong zero-shot transfer to downstream tasks such as audio-visual localization.
PaperID: 5618, Poster
Abstract: We revisit the role of appearance modeling in 3D Gaussian Splatting (3DGS) and show that limited expressiveness in view-dependent reflectance is a key driver of representation redundancy. In standard 3DGS, low-order spherical harmonics (SH) are used, restricting the splats' ability to model high-frequency directional effects, which is typically compensated by increasing the number of splats. We propose NRF-GS: Neural Residual Fields for Gaussian Splatting, a hybrid representation that replaces per-splat SH-bases with a shared neural residual field. Each Gaussian encodes a compact set of appearance features and a lambertian base color, while a lightweight global scene-level MLP predicts view-dependent residuals conditioned on viewing direction, distance, and per-splat features. This formulation enhances directional reflectance modeling by combining diffuse per-splat reflectance representations with a shared global function for high-frequency details, enabling both higher expressiveness and parameter sharing across splats. Our key insight is that by accurately capturing high-frequency directional reflectance, especially in specular regions, the GS-representation becomes more expressive, reducing the need for geometrically redundant splats. As a result, NRF-GS achieves comparable or better rendering quality while reducing the number of Gaussians by up to 50%, and produces visibly improved specular and high-frequency details.
PaperID: 5619, Poster
Abstract: Online bipartite matching, where agents are known in advance but items arrive sequentially and must be irrevocably assigned, is fundamental to problems ranging from ride-sharing to online advertising. When agents belong to classes such as demographic groups or geographic regions, fairness demands equitable treatment across these groups. Recent work introduced class envy-freeness (CEF), a natural extension of the classical fair division notion: an algorithm is \alpha-CEF if each class receives value at least an \alpha fraction of what it could extract from any other class's bundle. However, all known algorithms achieving constant-factor CEF guarantees attain utilitarian social welfare (total matching value) of at most \frac12 times the optimum, far below the 1-\frac1e \approx 0.632 achievable without fairness constraints. We resolve the open question of whether fairness necessitates this efficiency loss, by introducing threshold-based algorithms parameterized by \gamma \in [0,1] that equalize allocations across classes until threshold \gamma, then maximize efficiency. For divisible matching, this yields simultaneous (1-e^-\gamma)-CEF and (1 - \frace^\gamma-1\gamma+1)-USW guarantees; for indivisible matching, \frac\gamma2-CEF with the same USW. Setting \gamma > 0 produces the first algorithms beating \frac12-USW while maintaining constant CEF. We complement this with a novel upper bound construction, proving no non-wasteful \alpha-CEF algorithm can exceed \frac1 +\alpha - e^\alpha-11+\alpha-USW and correcting prior bounds that were vacuous for \alpha < 0.58. Our upper bound nearly matches our algorithms' performance, giving the first substantive characterization of the price of fairness in online class matching.
PaperID: 5620, Poster
Authors: Maxim Bochkov, Fedor Noskov
Abstract: It is known that optimally regularized ridge regression can exhibit non-monotone dependence of the generalization risk on the sample size. We show that this behavior can be arbitrarily intricate in finite dimensions: for every d \ge 2, there exists a random design linear model in \mathbbR^d whose generalization risk, viewed as a function of the sample size, exhibits at least d - 1 local maxima.
Abstract: Long-rollout causal video diffusion has converged on a fixed-size sliding-window KV cache, with recent progress innovating within this layout by changing which tokens occupy the window or how their positions are encoded. The per-head KV layout itself, a dominant contributor to streaming memory and latency, has been mostly left unchanged. In this paper, we present the first study of Multi-Head Latent Attention (MLA) in video diffusion. VideoMLA replaces per-head keys and values with a shared low-rank content latent and a shared decoupled 3D-RoPE positional key, reducing per-token KV memory by 92.7% at every cached layer. We further investigate why MLA succeeds in video diffusion even though the spectral assumption often used to motivate it in language models does not hold: pretrained video attention is not low-rank, with 99%-energy effective rank far above any practical latent dimension. VideoMLA retains quality at compression ratios where direct spectral approximation would predict large reconstruction error. We show that the MLA bottleneck, rather than the pretrained spectrum, determines the effective rank: both spectral and random initialization occupy nearly the full rank budget from initialization, and Stage-1 training preserves this budget while adapting within it. On VBench, VideoMLA matches short-horizon streaming video diffusion baselines, achieves the best overall score at long horizons among evaluated methods, and improves throughput by 1.23× on a single B200.
PaperID: 5622, Poster
Abstract: Self-supervised learning (SSL) is increasingly applied to scientific imaging domains, such as microscopy, where labels are scarce but noisy data is abundant. These domains exhibit significant signal-dependent noise, which can create a representational shortcut: since both augmented views derive from the same noisy image, the encoder can boost view agreement by also encoding noise patterns rather than semantic content. To enable noise-robust representation learning without knowledge of underlying noise models, we introduce Self-Aligned Noise Augmentation (SEANA), a drop-in module that generates noise-aware SSL views. SEANA learns an invertible variance-stabilizing transform (VST) from noisy data, then estimates the clean signal and resamples independent noise in learned VST space, requiring no clean targets, no dataset-specific denoiser training or use at inference, and no changes to the SSL objective or encoder. On CIFAR-10 and ImageNet-100 under synthetic Gaussian, Poisson–Gaussian, and multiplicative noise, SEANA improves clean-test linear-probe accuracy across contrastive/non-contrastive SSL methods, indicating higher-quality representations. On real fluorescence microscopy, SEANA improves Jurkat cell-cycle stage classification by up to 28.6pp over standard SSL and 27.5pp over denoiser-preprocessed baselines. In both settings, SEANA outperforms the denoiser-preprocessed SSL pipeline, demonstrating that learned VST-space noise-aligned resampling yields more robust representations than denoising alone. Code will be made publicly available.
Abstract: Effective multi-agent cooperation requires agents to adopt diverse behaviors as task conditions evolve—and to do so at the right moment. Yet, current Multi-Agent Reinforcement Learning (MARL) frameworks that facilitate this diversity are still limited by the fact that they bind fixed behaviors to fixed agent identities. Consequently, they are ill-equipped for tasks where agents need to take on different roles at very specific moments in time. We argue that, to define these behavioral transitions, the missing ingredient is . Events are changes in the state of the system that induce qualitative changes in the task. Based on this view, we introduce a framework that decouples agent identity from behavior, capturing a continuous manifold from which agents instantiate their behaviors in response to events. This framework is based on two elements. First, to build an expressive behavior manifold, we introduce Neural Manifold Diversity (NMD), a formal distance metric that remains well-defined when behaviors are transient and agent-agnostic. Second, we use an event-based hypernetwork that generates Low-Rank Adaptation (LoRA) modules over a shared team policy, enabling on-the-fly agent-policy reconfiguration in response to events. We prove that this construction ensures that diversity does not interfere with reward maximization by design. Empirical results demonstrate that our framework outperforms established baselines across benchmarks while exhibiting zero-shot generalization, and being the only method that solves tasks requiring sequential behavior reassignment.
Authors: Tak Hur
Abstract: Stochastic reconfiguration (SR) is the standard optimizer for neural quantum states (NQS), but modern NQS often have far more parameters than Monte Carlo samples. We show that in this regime the diagonal shift is more than a numerical stabilizer. It acts as a statistical filter for finite-sample generalization. At a fixed wave function, SR is ridge regression from tangent features to the centered local energy. Its residual is the expressivity gap, the part of imaginary-time evolution outside the current tangent space. This gap is orthogonal to the tangent space in population, but finite batches make it act as noise that SR can overfit. The shift therefore balances shrinkage of useful update directions against variance from fitting sampled residuals. Exact 4×4 J_1-J_2 experiments separate two effects of overparameterization. Larger tangent spaces help when they reduce the expressivity gap, but they can hurt when they overfit a fixed gap. In a 100-site transverse-field Ising family trained with a foundation NQS, validation risk is U-shaped in the shift while variance decreases, matching the noisy-ridge model. This view leads to multi-shift SR (MS-SR), which averages independent ridge solves at data-adaptive shifts to form a richer, lower-variance spectral filter. MS-SR lowers validation risk and stays closer to the ideal tangent-space imaginary-time update compared to other competitive NQS optimizers.
PaperID: 5625, Poster
Abstract: Designing RNA sequences that fold into specific 3D structures is a fundamental challenge in therapeutics. Current RNA inverse folding methods typically condition on 3D backbone coordinates but the generation is limited to 1D discrete sequences. This ignores the continuous side-chain orientations, which are essential for overall structural stability. To address this issue, we explicitly introduce side-chain geometry into the RNA generative process. Unlike co-design frameworks that diffuse across separate raw spaces, we propose Geometric Optimized Latent Diffusion (GOLD), which project sequence and structure into a unified latent manifold. GOLD employs a novel GeoVAE to elegantly integrate discrete RNA sequences with continuous side-chain angles. By mapping these modalities into a single, smooth latent manifold, GOLD naturally facilitates a synergistic co-diffusion process. To enhance structural validity, we further introduce a training-free, gradient-guided sampling mechanism. By decoding the generated sequence probabilities and side-chain geometries into structures during the sampling process, we compute analytical gradients from physical constraints. This dynamic feedback actively guides the generated sequence to a physically stable state. Extensive experiments validate the superiority of GOLD, achieving state-of-the-art performance across sequence recovery, secondary structure validity, and 3D foldability metrics.
PaperID: 5626, Poster
Abstract: Automated cardiac MRI quantification is commonly framed either as dense segmentation followed by voxel counting or as direct regression of global measurements. The two paradigms are difficult to unify: segmentation is visually interpretable but label-intensive and output-heavy, while regression is compact but less verifiable. We propose Fourier Contour Learning, which represents cardiac boundaries with compact anatomical Fourier contours. The predicted coefficients define smooth, interpretable contours and enable differentiable volume computation, connecting mask supervision, measurement supervision, and unlabeled-image regularization within one framework. We instantiate this representation with an anatomy-query decoder and cross-slice attention which creates a traceable path from image to contour to clinical measurement. Across three public datasets and an in-house all-phase cohort, our method achieves segmentation accuracy comparable to strong dense-prediction baselines, improves boundary agreement, and reduces clinically relevant volume error while using only a small number of output parameters per structure. We further show that the learned model provides an efficient interface for vision-language question answering and surface reconstruction. These results demonstrate the potential of Fourier contours as a unified intermediate representation for scalable, traceable, and clinically grounded cardiac analysis. Code will be publicly available upon acceptance.
PaperID: 5627, Poster
Authors:
Shou Z Chen, Bing He, Zhenchao Tang, Jun Zhu, Minghao Yang, Tianxu Lv, Yu Wang, Jiayang Wu, Fang Wang, Yu Zhao, Chenchen Qin, Jianhua YaoAbstract: High-throughput scientific sensing, exemplified by genomics, faces a critical representation bottleneck: discrete biological events are often obscured by continuous, high-entropy stochastic noise. Standard neural codecs fail to distinguish deterministic semantics from stochastic sensor noise, inefficiently allocating their information budget to high-entropy fluctuations rather than to low-entropy structural motifs. We introduce \textscWaveSem, a frequency-adaptive framework that imposes a physically motivated inductive bias on discrete representation learning. Through integrated wavelet decomposition, \textscWaveSem explicitly disentangles each signal into a structured semantic core and a residual texture component. We validate this approach on large-scale nanopore sequencing data, a domain characterized by extreme sequence lengths and low signal-to-noise ratios (SNRs). \textscWaveSem achieves a 14× storage reduction while maintaining competitive downstream basecalling accuracy. Moreover, linear probing of the frozen encoder reaches \mathrmAUROC=0.928 for 5mC methylation detection, demonstrating that the learned tokens preserve semantic information beyond nucleotide identity. These results establish disentangled vector-quantized representations as a scalable foundation for efficient genomic analysis and long-term archival.
PaperID: 5628, Poster
Abstract: Zeroth-order (ZO) optimization enables memory-efficient fine-tuning of large language models (LLMs) using only forward passes, but it remains unclear how useful adaptation is distributed across layers. In this work, we reveal a surprising phenomenon: ZO fine-tuning is sharply dominated by a single decoding layer. Across multiple LLM families and downstream tasks, fine-tuning this dominant layer alone consistently matches or even exceeds full-model ZO fine-tuning. We further show that the dominant layer is task-agnostic but model-specific, and can be identified before training through a simple inference-only analysis of activation outliers. Specifically, the dominant layer consistently aligns with the first activation-outlier layer in the pre-trained model. To explain this phenomenon, we analyze how perturbation effects propagate under ZO optimization. We find that the dominant layer combines two key properties: high perturbation sensitivity and early placement in the residual stream, allowing perturbation-induced effects to propagate and accumulate through remaining subsequent decoding layers. As a result, this layer produces disproportionately strong and stable optimization signals under forward-only updates. Extensive experiments on LLaMA2-7B and Qwen3-8B across nine benchmarks show that dominant-layer ZO fine-tuning improves average performance over full-model MeZO and LoRA-based ZO fine-tuning while achieving up to 4.52× training speedup.
PaperID: 5629, Poster
Abstract: One-shot Federated Graph Learning (OFGL) has emerged as a communication-efficient paradigm for Federated Graph Learning (FGL) under strict bandwidth constraints, compressing collaboration into a single transmission round. Nevertheless, this extreme setting suffers from a fundamental , which limits both the quality and utility of exchanged knowledge. First, while existing methods attempt to distill private graphs into shareable proxies, their aggressive one-shot compression inherently triggers a fidelity crisis. This compression distorts local graph topology and semantics across clients, leading to . Second, although the server aggregates these proxies to update global models, current integration schemes suffer from transferability agnosticism. The server cannot reliably identify cross-client collaboration potential, resulting in . To tackle these dual challenges, we propose FedHFT, an effective One-shot Personalized Federated Graph Learning (OPFGL) framework that addresses this bottleneck from two perspectives, namely High-Fidelity Knowledge Externalization and Transferability-Ware Knowledge Integration. We use Fidelity-Driven Proxy Refinement to preserve client-side knowledge fidelity, and Transferability-Steered Personalized Collaboration to enable adaptive server-side collaboration. The superiority of FedHFT is validated through extensive experiments on both homophilic and heterophilic graphs. The code is available at https://anonymous.4open.science/r/FedHFT-DC7F/.
PaperID: 5630, Poster
Abstract: Modern 3D structure predictors rely on multiple sequence alignments (MSA) to expose evolutionary covariation, but for RNA those alignments are often missing, shallow, or unreliable. To overcome this limitation, we invert the co-evolutionary structure inference pipeline and encode a predicted 2D topology into synthetic alignment-form covariation, without requiring database search nor alignment steps. We introduce RNAformer, a transformer-based RNA secondary-structure predictor that supplies the topology, and a Synthetic Co-evolution Engine (SCE) that encodes it into synthetic homologous sequences (SHS) by co-mutating paired positions. Comprehensive experiments and systematic interventions across multiple structure predictors show that the MSA input acts as a topology channel that allows predictable steering of 3D structure prediction. Across single-chain RNAs and single-chain RNA-protein complexes, RNAformer-seeded SHS approach the performance of natural MSAs, with regime-dependent gains. Topology, not alignment depth, drives the response. SHS are interpretable, fast, and scalable, and represent a first step toward an alignment-free alternative to natural MSA.
PaperID: 5631, Poster
Authors: Perrine Chassat, Agathe Guilloux
Abstract: Synthetic patient data generation is a promising solution to the dual challenge of data scarcity and privacy constraints in healthcare machine learning. Realistic synthesis of patient-level clinical data requires jointly modeling heterogeneous static covariates, irregularly sampled longitudinal trajectories, and informative observation times — three tightly coupled components in practice yet rarely addressed together. We propose HyperNSDE, a continuous-time generative model that conditions a latent Neural SDE on static patient representations through a hypernetwork, allowing baseline characteristics to shape the full trajectory dynamics rather than only the initial condition, without requiring a trajectory encoder, while stochastic latent dynamics capture realistic variability in generated paths. Observation times are modeled jointly through a latent-state-dependent intensity process. Training on irregular stochastic paths is stabilized via a deterministic—stochastic path decomposition with a non-adversarial signature-kernel objective. Experiments on simulated and real clinical datasets demonstrate consistent improvements over strong baselines across fidelity, privacy, and utility metrics.
Abstract: The rapid emergence of generative image models has led to the development of specialized watermarking techniques, particularly in-generation methods such as seed-based embedding. However, current evaluations in this area remain largely empirical, making them heavily reliant on the specific model architectures used for generation and inversion. This prevents any clear conclusion on the performance of any method, especially regarding security, for which a rigorous definition is lacking. Against this approach, we argue that the effectiveness of a watermarking scheme should be established purely through a thorough theoretical analysis. This is enabled by decoupling the model-dependent part from the actual decision mechanism of the watermarking system. Using this decoupling, we introduce a formal evaluation framework based on security, robustness, and fidelity. This allows precise comparisons between watermarking systems through a characteristic surface representing the trade-off between these three quantities, independent of any generative model. Based on this framework, we propose SSB, a novel watermarking method that generalizes previous seed-based methods by allowing to reach any security-robustness-fidelity regime on its characteristic surface. This work opens the door to the design of modern watermarking systems with theoretical guarantees that do not necessitate any costly empirical evaluations.
PaperID: 5633, Poster
Authors: YunFeng Deng
Abstract: Test-time scaling (TTS) improves reasoning by allocating additional inference compute, but in multimodal large language models (MLLMs), more compute is not always reliable. Longer reasoning may amplify visual misperception, induce answer drift, reinforce majority traps, or turn initially correct predictions into overthought failures. We propose MUST, a training-free framework that recasts multimodal TTS as stage-conditioned stability control. MUST separates inference into reasoning and verification stages and assigns each stage a different reliability criterion. During reasoning, Contrastive Answer-Manifold Stability (CAMS) selects answers whose supporting trajectories form compact and separable latent groups. During verification, Attention-Free Support-Chain Stability (AF-SCS) evaluates whether explicit visual evidence and subsequent reasoning form a stable latent support chain, without relying on token-to-patch attention maps. Experiments on multimodal reasoning benchmarks show that MUST improves the accuracy-compute trade-off and reduces overthinking-induced errors. These results suggest that reliable MLLM test-time scaling should adaptively control when and how extra compute is trusted, rather than simply increasing inference budget.
PaperID: 5634, Poster
Authors: Haozhuo Zheng, Cheng Wang, Pengyu Chen, YajunTian, Yang Liu
Abstract: Molecular graph generation typically involves a trade-off between two paradigms: diffusion models capture global structure well but require hundreds of denoising steps, while autoregressive (AR) models sample in a single pass yet generate atoms in a flat canonical order with no semantic hierarchy, leaving long-range topology to chance. We introduce the Hierarchical Graph VQ-Transformer (H-GVT), which removes this dichotomy by performing coarse-to-fine AR generation over hierarchical discrete codes. A Spectrally-Regularized Multi-Scale VQ-VAE first compresses each molecule into a sequence of discrete tokens at multiple scales, where coarse tokens provide regularized low-resolution structural context and fine tokens encode atom-level details. The compression combines Adaptive Min-Cut Pooling with a novel Spectral Topology Consistency Loss that aligns low-frequency normalized Laplacian spectra across scales, adding only 2.8% training overhead. A standard decoder-only Transformer with level embeddings then generates these tokens from coarse to fine, allowing fine-grained atom tokens to condition on coarser structural context. This requires no specialized tree, motif, or scale-causal masking. On QM9 and ZINC250k, H-GVT achieves the best NSPDK among compared baselines (2×10^-4 on QM9 and 1×10^-4 on ZINC250k), while sampling 10K molecules 68× faster than DiGress. On MOSES, H-GVT achieves 4.8× better FCD than a Novelty-thresholded GVT baseline under the same thresholded protocol (0.19 vs. 0.92; 86.4% vs. 80.5% Novelty). Crucially, ablations reveal that the coarse-to-fine ordering itself, not merely multi-scale tokenization, is what enables global structural planning: randomizing the order while keeping the same tokens, codebook, and level embeddings degrades NSPDK by 29×.
PaperID: 5635, Poster
Authors:
Zhenxu Tian, Kebin Liu, Zhengwu Yang, Yi Su, Qingqing Dang, Kaipeng Deng, Yanlin Sha, Yanjun Ma, Dianhai Yu, Juntao Li, Min zhangAbstract: Long-context decoding in large language models (LLMs) is increasingly bottlenecked by the memory and computation required to attend over an ever-growing Key-Value (KV) cache at each autoregressive step. Although token-level sparse attention reduces this burden by restricting computation to a small subset of relevant past tokens, its practical gains are often offset by repeatedly recomputing layer-wise sparse indices as the context grows. A common remedy is to share sparse indices across layers by computing them only at designated anchor layers, but such static reuse ignores the layer-wise evolution of attention patterns, causing the reused indices to progressively deviate from the layer-specific optima and incur substantial attention loss. By revisiting cross-layer sparsity from the perspective of attention dynamics, we find that inter-layer salient token shift is highly localized: although the exact top-\kappa indices differ across layers, the core high-attention tokens remain stable, with changes concentrated near the expanded boundary of previously selected regions. Motivated by this, we propose Prism Attention, a proposal-refined sparse attention framework. It replaces rigid cross-layer index reuse with a coarse-to-fine strategy, preserving layer-wise adaptability while retaining the efficiency benefits of cross-layer sharing. Extensive experiments show that Prism Attention achieves state-of-the-art performance among mainstream sparse attention methods on both reasoning and long-context benchmarks, while achieving 2.6×~7.5× kernel speedup over FlashAttention at 128K context length.
PaperID: 5636, Poster
Abstract: In natural scenes, light sources and occlusions from scene geometry create spatially varying illumination, including brightness gradients, cast shadows, and local color shifts that shape how every object appears. A less explored aspect of diffusion-based object compositing is whether an inserted object can adapt to this local illumination at the desired location. The core challenge to achieve this is the scarcity of large-scale training data capturing how object appearance should adapt to location-dependent illumination. We present LOCO, a data augmentation framework for LOcal light-aware object COmpositing that leverages publicly available monocular videos as illumination sources. We treat each video frame as a virtual spotlight anchored at its camera pose, so that viewpoint variation across a video yields diverse light directions. With controllable beam shape, color, and intensity, this design enables large-scale generation of training data for placement-dependent illumination on the fly. We also introduce LOCO-bench, a benchmark of physically captured composites under spatially varying illumination, where the same object is placed at multiple locations within each scene. Experiments on Laval Indoor SV HDR and LOCO-bench show that LOCO outperforms prior methods in local illumination adaptation while preserving object identity.
Abstract: Generating chemically valid 3D molecular conformations is critical for computational drug discovery. Classical diffusion-based models like GeoLDM perform well but require hundreds of steps, making large-scale in silico screening impractical. Recent efforts on few-step molecular generation have accelerated this process to 12-50 steps, but they often largely sacrifice sample stability. In this work, we present FlashMol, an ultra-fast molecule generative model producing high-quality molecular conformations in as few as 4 steps. To achieve this, we adapt distribution matching distillation (DMD) — a reverse KL-divergence minimization objective — to the molecular domain for effective distillation. Considering the local minimization behavior of DMD, we respace the molecule generation timesteps, providing the generator with much better initialization and enables effective distillation. Additionally, to mitigate the mode-seeking behavior of DMD and improve diversity, we further regularize it with a Jensen-Shannon divergence term, which incorporates the mean-seeking behavior of the forward KL divergence. Extensive experiments on QM9 and GEOM-DRUG datasets demonstrate that FlashMol matches and even surpasses the original 1000-step teacher, achieving up to 250× acceleration in sampling speed while maintaining high molecular quality.
PaperID: 5638, Poster
Abstract: Evaluating large language model (LLM) agents in multi-turn interactive environments is expensive and risky, as it requires online environment interaction. We propose ADWM (Autoregressive Diffusion World Model), an evaluation framework that estimates the performance of a new LLM agent policy purely from pre-collected trajectories. The core idea is to learn a latent diffusion world model that simulates how the environment responds to the evaluation policy, without ever executing it in the real environment. Existing diffusion-based OPE methods guide full trajectories in a single pass by jointly diffusing states and actions, an assumption that breaks down for LLM agents whose actions are discrete text that must be sampled from the policy after observing the environment. Unlike autoregressive world models that suffer from compounding errors, ADWM models each transition as an independent denoising process, enabling reliable step-by-step rollouts where the world model and agent alternate in causal order. Crucially, the LLM agent under evaluation directly guides the diffusion generation at each step via a policy-conditioned score function, ensuring that simulated trajectories accurately reflect its decision-making patterns. Empirically, ADWM achieves accurate value estimates and evaluation reliability across diverse multi-turn agent tasks, demonstrating its promise as a practical framework for offline LLM agent evaluation.
PaperID: 5639, Poster
Abstract: Entropy-regularized optimization problems arise in various areas of machine learning, including mean-field neural networks, variational inference, and reinforcement learning. Mean-field Langevin dynamics (MFLD) is a dynamics capable of sampling from the optimal solution of such problems under appropriate conditions, and has been the subject of extensive recent research. As machine learning models grow in scale and parameter dimension, computational efficiency becomes increasingly important, especially in reducing the adaptive complexity of computation of MFLD, that is, the number of sequential rounds they require. In this work, we initiate the study of the adaptive complexity of MFLD and provide the first theoretical evidence that MFLD can indeed be accelerated through parallelism. In particular, we provide polylogarithmic convergence guarantees for both the infinite-particle and finite-particle setting under the log-Sobolev inequality.
Abstract: Understanding the dynamic behavior of biomolecules is fundamental to elucidating biological function and facilitating drug discovery. While Molecular Dynamics (MD) simulations provide a rigorous physical basis for studying these dynamics, they remain computationally expensive for long timescales. Recent deep generative models accelerate conformation generation but often either discard temporal correlations entirely or struggle to condition on the extended history required for faithful kinetics, a consequence of the effectively non-Markovian nature of partially observed subsystem coordinates. To bridge this gap, we introduce ATMOS, a novel generative framework based on State Space Models (SSM) designed to generate atom-level MD trajectories for biomolecular systems. ATMOS integrates a Pairformer-based state transition mechanism to capture temporal dependencies, with a diffusion-based module to decode trajectory frames autoregressively. We demonstrate that ATMOS achieves state-of-the-art performance in generating conformation trajectories for both protein monomers and protein-ligand systems. This work provides a unified and computationally efficient framework for biomolecular trajectory generation, taking a step toward dynamics foundation models.
PaperID: 5641, Poster
Abstract: In learnable logic gate networks, each neuron implements a k-input Boolean function. The dominant setting k=2 was chosen for tractability. We ask: when does wider fan-in help? At fixed parameter budget, each gate costs 2^k parameters, so increasing k exponentially reduces width. Combined with the depth--fan-in bound L \geq \lceil \log_k d \rceil, this partitions tasks into two regimes: width-dominated (image classification), where more gates beat richer gates, and degree-dominated (parity, arithmetic), where k>2 is necessary. Our main finding is that representability does not imply learnability: at the regime boundary, the selection mechanism determines training success at identical architecture, and the optimal mechanism inverts between regimes. We formalize this via a Jacobian analysis: the softmax mechanism's positive-semidefinite (PSD) structure gives basin attraction absent from sigmoid. We train multilinear coefficients with exact gradients and Möbius-transform snapping (zero deployment error), and validate on 28 tasks and 5 real-world datasets across multiple selection mechanisms.
PaperID: 5642, Poster
Authors:
Zhenghong Zhou, Xiaohang Zhan, Zhiqin Chen, Soo Ye Kim, Nanxuan Zhao, Haitian Zheng, Qing Liu, HE Zhang, Zhe Lin, Yuqian Zhou, Jiebo LuoAbstract: Recent video diffusion models achieve strong visual quality and temporal coherence, but still lack coordinated control over scene layout, customized subject identity, and camera/subject motion. We study controllable video generation from three prompts: a first-frame scene image, 3D-aware multi-view subject references, and a motion-driving signal. This setting is challenging because background regions are often trackable under camera motion, while foreground subjects can rotate, self-occlude, and reveal new regions that require identity-consistent appearance. We introduce Tri-Prompting, a video diffusion framework that combines multi-view subject conditioning with dual-conditioned motion control. Tri-Prompting uses XYZ tracking points for visible background motion and low-resolution RGB proxies for foreground subject pose, allowing the model to recover fine appearance from multi-view references while retaining flexibility for plausible subject-scene interactions. An inference-time ControlNet scale schedule further balances motion controllability and visual realism. We evaluate Tri-Prompting against specialized baselines: DaS for motion reconstruction and Phantom for subject-driven generation. Tri-Prompting achieves competitive or improved reconstruction quality, stronger multi-view identity preservation, and better 3D consistency, while enabling 3D-aware subject insertion and in-scene manipulation.
PaperID: 5643, Poster
Abstract: (1+1)-evolution strategy (ES) is a classic variant of evolution strategy for black-box optimization via direct and random search while keeping only a single solution per generation in \mathbbR^n. Despite convergence with different step-size adaptations, it often remains impractical due to dependence on the unfavorable high dimension of the search space \mathbbR^n. In fact, many modern optimization problems are naturally constrained to low-dimensional manifolds embedded in very high-dimensional ambient spaces. To see the potential of intrinsic dimensionality reduction for improving practicality, we consider the truncated Riemannian (1+1)-ES on a compact Riemannian submanifold \mathcalM\subset\mathbbR^n and establish its non-asymptotic stationarity guarantee at the standard nonconvex zeroth-order rate of \widetilde\mathcalO(\varepsilon^-4) function evaluations. The explicit Gaussian second-moment term in the bound scales with the intrinsic tangent dimension d=\dim\mathcalM rather than the ambient dimension n, while the remaining truncation dependence is captured by an explicit small-ball drift constant \alpha_d(R/\sigma_0). To the best of our knowledge, this gives the first non-asymptotic stationarity analysis for a comparison-based ES-type method on compact nonconvex Riemannian submanifolds under localized pullback smoothness. Numerical illustrations on sparse PCA over spheres and smooth-surrogate stationarity diagnostics illustrate behavior consistent with the theory and its limitations, while additional experimental studies on truncation radius diagnostics, black-box tuning, and classical evolutionary computation benchmarks support the practical relevance of geometry-aware comparison-based search under limited function-evaluation budgets.
PaperID: 5644, Poster
Abstract: 3D Gaussian Splatting (3DGS) achieves real-time photorealistic novel-view synthesis but is vulnerable to adversarial perturbations of its training images. Recent attacks inflate the Gaussian count to exhaust GPU memory or corrupt rendered scenes to mislead downstream classifiers. We find that both attacks have the same structural signature: 2D image adversarial perturbations are not 3D-consistent by design. The Gaussians generated to match them provide erroneous contributions in some views while making significant contributions in others. We propose Multi-View Photometric Inconsistency (MVPI) Defense that scores each Gaussian by the variance of its opacity-weighted local \mathcalL_1 contribution across a stratified set of training views. The most inconsistent are soft-pruned via 3DGS's existing opacity threshold method. We evalute our approach on two different datasets. MVPI reduces peak Gaussian count by up to 1.90× and training time by 1.22-1.50× under Poison-Splat attack. It also restores top-1 classification accuracy on adversarially perturbed renders from 57.4% to 66.5%. Our code is available at https://anonymous.4open.science/r/PruneDefense-2876/README.md.
Authors: Minh Quoc Duong, Chun Tong Lei, Chun Pong Lau
Abstract: With the proliferation of AI-generated images, digital watermarking has become an essential safeguard for protecting intellectual property and mitigating malicious exploitation. Recent works on semantic watermarking have enabled efficient copyright protection for diffusion models. However, the dependence of semantic watermarking on diffusion inversion for watermark detection creates a critical vulnerability. Imprint removal and forgery attacks exploit this weakness to produce deceptive results. Our analysis reveals that these attacks succeed by displacing watermarked latents into the , the first plug-and-play, training-free noise extraction framework designed to defend against both attack strategies. PGID effectively defends by projecting perturbed latents back to the region where they originally belong. The projection is achieved by eliminating intermediate latent deflections and mitigating adversarial perturbations through progressive inversion-denoising cycles. Comprehensive evaluations across multiple schemes demonstrate that PGID successfully restores detection reliability by recovering removed watermarks and identifying forged instances.
PaperID: 5646, Poster
Abstract: Multi-modal large language models (MLLMs) achieve strong modality understanding by pairing a large language model (LLM) with an encoder for a target modality such as vision, video, or audio. However, improving an MLLM's capability for a given modality typically requires additional training on large modality-specific datasets, incurring substantial data collection and compute costs. Model merging offers an alternative, but it is often infeasible for data-scarce, large per-sample size, or domain-specific modalities (e.g., audio and video), where same-modality model variants are rarely available. In this work, we characterize an intriguing asymmetric phenomenon: merging a well-aligned, data-rich source-modality MLLM into a data-scarce target-modality MLLM substantially improves the target on its own benchmarks. Our theoretical and empirical analyses show that this gain stems from enhanced alignment between modality-specific and textual tokens, induced by the stronger donor modality. Specifically, we derive a mutual-information lower bound that is monotonic in alignment-related quantities and strongly correlated with downstream MLLM performance. Building on this principle, we propose Directional Cross-Modal Alignment Transfer (DCAT), a novel framework that transfers textual alignment from a strong, well-aligned source (donor) modality to a weak target (recipient) modality, boosting target-modality performance without further fine-tuning. We further show that the alignment-enhancing objective admits a closed-form weight-space solution computed from only a small calibration set. DCAT outperforms existing model-merging methods, offering an efficient path toward cross-modal alignment transfer.
PaperID: 5647, Poster
Abstract: Precisely reproducing a target span verbatim within a long context is a fundamental capability test for LLMs: even when the target string is present in context, models often still drift within the span or fail to stop correctly, revealing limitations in source grounding and boundary control. This capability has broad practical relevance. For instance, code generation often requires embedding sensitive strings such as API keys exactly as they appear in context. In this work, we construct a benchmark to test exact span reproduction ability across contracts, scientific literature, code, and random sequences, and find that current models still make surprisingly frequent errors. To address this, we propose CopyGen, a lightweight copy-constrained plugin for frozen LLMs that turns copying from a byproduct of next-token prediction into a controlled process. Our key idea is to decompose copying into three distinct decisions: when to copy, how far to copy, and when to stop. Concretely, CopyGen first determines whether to enter copy mode, then predicts a tentative copy length, and finally verifies the generated span token by token. During verification, we adaptively adjust the logits of copy continuation tokens to incentivize or suppress continuation, and emit only the longest valid prefix. Across three LLM backbones and four benchmark tasks, CopyGen improves average exact match by 35% relative without model finetuning, and integrates seamlessly with SFT to further improve it by 4%, while substantially accelerating generation. These results suggest that reliable exact copying is not something LLMs obtain for free, but a decoding-time control problem that requires explicit modeling. Code is available at https://anonymous.4open.science/r/CopyGen-A65D.
PaperID: 5648, Poster
Authors: Blaž Škrlj, Ivan C Arisoy, Solal Vernier
Abstract: We give computable necessary conditions for replacing a fixed attention head by a value-linear recurrent surrogate. The semiseparable spectrum of the causal attention matrix lower-bounds the operator-norm error of any finite-state value-linear recurrence, and retrieval-style blocks yield sharp K and Kd_v state-scaling laws. Empirically, these lower bounds are tight on retrieval-shaped targets but can be loose on smooth or clustered targets, where rank feasibility does not imply trainability or downstream quality. Large pretrained-head surveys suggest that this rank obstruction is often small, including in modern 1B, 7B, and 8B checkpoints. The resulting diagnostic is best used as a rejection test for under-sized recurrent state, not as a certificate that a recurrent replacement will train or improve downstream performance.
Abstract: The rise of machine learning has shifted targeted resource allocation in policy and humanitarian settings toward algorithmic targeting based on predicted risk scores. This approach is typically cheaper and faster than traditional screening procedures that directly observe the latent vulnerability status through physical verification. Yet, even access to the true conditional vulnerability probability cannot eliminate misallocation: aleatoric uncertainty over individual vulnerability status is irreducible, and probabilistic targeting inevitably misallocates some resources. In this work we study how screening and algorithmic targeting should be optimally combined in a two-stage allocation framework where a screening stage observes true outcomes for a subset of units before a final allocation stage assigns the resource under a fixed coverage budget. We show that the optimal strategy screens units at the margin of algorithmic allocation, while directly targeting the highest-risk units. Furthermore, we empirically characterize when screening and algorithmic targeting act as complements or substitutes: efficiency gains from screening grow as the aleatoric uncertainty in the population increases. We illustrate our framework with applications in income-based social protection programs and humanitarian demining in Colombia, where the tension between screening costs and allocation efficiency is operationally consequential.
Abstract: Reliable generalization in conditional latent-variable models requires understanding both identifiability and extrapolation: whether observed variation across attributes determines latent structure, and whether that structure determines distributions at unseen attributes. However, existing identifiability and extrapolation guarantees are largely model-specific, with separate analyses in nonlinear ICA, causal representation learning, perturbation modeling, and related conditional latent-variable models. We introduce concept modulation models (CMMs), an attribute-indexed class of conditional generative models with structure A\to \Lambda \to C\to X, where attributes select modulators, modulators induce latent concept laws, and concepts generate observed features. CMMs lift transition-based identifiability to conditional settings by showing that feature agreement on observed attributes induces a latent concept transition constrained by the CMM class. We express these constraints through attribute potentials, log-density ratios between attribute-conditioned concept laws, separating the generic lifting step from model-specific rigidity arguments. The same potentials control extrapolation: agreement at unseen attributes holds exactly when the transported attribute-potential identities extend to those attributes. This yields algebraic extrapolation criteria, identifies the common potential-based proof objects behind several existing identifiability and extrapolation results, and, when combined with the model-specific rigidity arguments in those works, recovers their stated conclusions.
PaperID: 5651, Poster
Abstract: As AI-assisted image editing becomes increasingly prevalent, Text-Centric Image Forgery Localization (TFL) is essential for protecting trust in financial and legal records. However, we observe that existing TFL detectors predominantly rely on Full-Parameter Fine-Tuning (FPFT) and fail to generalize to unseen forgery types because FPFT can induce low-rank feature collapse and overfit training artifacts. Meanwhile, prior detectors capture diverse forgery traces with a single unified representation, making subtle traces entangled and difficult to learn. To address these challenges, we propose RiSE, a Residual Subspace Expert model that preserves pre-trained priors by freezing principal SVD components and adapting only the residual subspaces with a ViT-only segmentor to avoid feature collapse. RiSE introduces forgery-aware residual experts, where each expert is trained independently on a training subset with a particular forgery type, and we show that cross-domain localization performance improves consistently as more experts are added. At inference, we introduce Latent Perturbation Confidence (LPC), which selects the most confident expert by analytically computing the latent-perturbed output. LPC enables efficient confidence estimation with a single forward pass, avoiding repeated forward passes on perturbed samples and improving expert selection. By training on synthetic forgery data, RiSE with 1 expert outperforms state-of-the-art methods on real-world forgery images by 20.4% while reducing training steps by 10×. Scaling the number of experts to 9 yields a 35.8% gain. RiSE also remains effective with DCT feature fusion, even when the majority of the parameters is frozen with RGB-only pretraining. The code is in the supplementary material.
PaperID: 5652, Poster
Abstract: Semantic fidelity is essential for multi-agent communication under non-ideal channels, comprising transmission fidelity and conversion fidelity. Yet existing numerical and neural message formats cannot simultaneously achieve both. To address this, we propose a novel message format-Hyperdimensional Symbolic Messages (HSM), and its corresponding communication framework Herm. For transmission fidelity, Herm suppresses semantic entanglement via an entity-attribute-value structured encoding pipeline. For conversion fidelity, Herm first leverages a receiver-centric ego-binding mechanism to align semantics. Then, a similarity-aware dual-branch fusion module is proposed to prevent critical semantic dilution. To realize explicit message-semantic conversion, we devise an iterative decoding scheme via algebraic inverse operation. Experiments exhibit that Herm outperforms baselines across various noisy scenarios while maintaining superior structural readability.
PaperID: 5653, Poster
Abstract: Medical time series (MedTS) analysis plays a critical role in supporting clinical decision-making and patient care. However, existing methods treat MedTS signals purely as numerical sequences, ignoring the rich information embedded in their visual waveform, which is essential for clinicians to diagnose physiological conditions. To bridge this gap, we introduce a novel trieval (ViRe) framework that injects visual waveform prior knowledge into representations of raw MedTS sequences, thereby emphasizing morphology-relevant patterns and improving clinician-aligned reasoning. Specifically, a is extracted using pre-trained vision-language models (VLMs) to obtain morphology-aware priors from waveform plots. To enrich the numerical representation with such morphology-aware information, we design a tailored attention-based cross-modal retrieval mechanism that aligns raw numerical features with these high-level vision priors through the Vision Query. ViRe demonstrates strong effectiveness, outperforming the state-of-the-art method by an average of
PaperID: 5654, Poster
Abstract: Chain-of-Thought (CoT) enables transformers to solve complex reasoning tasks by generating intermediate reasoning steps. Despite strong empirical success, the mechanisms by which such reasoning abilities emerge during training remain poorly understood, particularly from a theoretical perspective. In this work, we study this question in a stylized yet fully analyzable setting: in-context \mathrmK-nearest-neighbor (\mathrmK-NN) prediction with a one-layer transformer. Earlier work has demonstrated that one-layer transformers can be trained to perform 1-NN without CoT \citepLi2024One by having the softmax attention attend to the nearest neighbor in context. However, this result on 1-NN cannot be extended to \mathrmK-NN for K>1, as softmax attention is much less naturally suited to attending to the 2nd through K-th nearest neighbors. We first give a rigorous negative result showing that this limitation already appears on a well-separated class of binary classification tasks for odd K>1. We then give an explicit construction of a one-layer transformer, and show that it can solve \mathrmK-NN via CoT reasoning. We further prove that, under a stylized training setup, gradient descent can recover this construction, thereby showing that transformers can acquire the capability to solve \mathrmK-NN through training. In addition, we show that the resulting trained model applies to a broader test class than that assumed in the training analysis. Our results shed light on how CoT expands the ability of transformers to learn and solve tasks that are otherwise hard to solve.
PaperID: 5655, Poster
Abstract: Offline meta-reinforcement learning seeks to learn a policy that generalizes to new related tasks online. Context-based methods infer a task representation from transition histories, yet learning an effective task representation without supervision remains challenging. Existing methods relying on contrastive learning learn discriminative task representations, but fail to identify task-specific dynamics, while relying on reconstruction can be insufficient to model long-horizon dependencies, limiting generalization to new tasks. We investigate the impact of temporal consistency in latent space on task representation learning, showing that enforcing multi-step predictions in latent space encourages task representations that are able to capture task-dependent dynamics while preventing representation collapse. We provide theoretical analysis characterizing sources of error in value estimation and show through extensive experiments on MuJoCo, Contextual DeepMind Control, and MetaWorld benchmarks that temporal consistency significantly improves both zero-shot and few-shot generalization.
Abstract: Large Language Models (LLMs) exhibit remarkable capabilities, yet it remains unclear to what extent these reflect sophisticated recall or genuine reasoning ability. We introduce chess as a controlled testbed aimed at disentangling these faculties. Leveraging the game’s structure and scalable engine evaluations, we construct a taxonomy of positions varying in density of relevant priors - ranging from common states solvable by memorization to completely novel ones requiring generalization. Crucially, our approach achieves this distinction without requiring explicit knowledge of the models' training data. Applying this taxonomy, we combine a longitudinal analysis of the GPT lineage with a rigorous evaluation of contemporary models, including Claude Opus and Gemini. Our analysis reveals a clear gradient: performance consistently degrades as the density of relevant priors decreases. Notably, for tasks with few relevant priors, base model performance regresses to the random-play baseline. While newer models improve, progress slows significantly for tasks with sparse priors. Furthermore, while reasoning-augmented inference improves performance, its relative marginal benefit per token decreases in the absence of relevant priors. These results suggest limitations in systematic generalization, highlighting the need for mechanisms beyond scale to achieve robust performance when deprived of relevant priors.
PaperID: 5657, Poster
Authors:
Jie Liu, Lanlan Liang, Qinan Bao, ZiXi Yan, Linghao Meng, Wenbo GongAbstract: We analyze support mismatch in language-model in-context prediction as a frozen-prior Bayes benchmark with finite discrete support. The goal is diagnostic rather than literal: we do not identify production LMs with exact Bayesian updating, but characterize a failure mode whose signatures can be checked in learned in-context predictors. Our main theorem gives the finite-sample predictor-risk bound R_k^\pi \le C e^-k\rho_\min + 2\varepsilon_\mathrmapprox^2, so risk decomposes into an exponentially decaying transient and a nonvanishing support-mismatch floor under exponential-moment control of log-likelihood ratios. In scalar and aligned Gaussian location families, exact two-component analyses match the upper exponent up to a tight factor 2 and show a regime change at \mu_2=3\mu_1. Controlled learned-model experiments track the benchmark near crossover. Decoder-only LM diagnostics then test the benchmark signatures: restricted support creates a persistent high-error regime; restoring the missing support sharply lowers error; and modern size-matched distractor controls on Qwen and Llama show that adding a wrong same-cardinality candidate does not reproduce this gain. Multiple-choice, instruction-tuned, cross-family, calibration, null-control, and near-tie checks preserve the same diagnostic picture.
PaperID: 5658, Poster
Abstract: Molecular docking aims to predict the three-dimensional structure of a protein–ligand complex; in the conventional structure-based setting, it is conditioned on an input protein structure and ligand, often starting from an apo receptor in flexible docking or a holo receptor in rigid docking. Recent co-folding models exemplified by AlphaFold 3 achieve strong accuracy in protein–ligand complex prediction, but they typically generate complexes from sequence-derived inputs and therefore do not directly use apo structures as structural priors in docking, leading to a larger conformational search space and weaker robustness in difficult or few-step settings. We present BridgeDock, a framework that adapts pretrained co-folding backbones to flexible docking by modeling the apo-to-holo transition as a diffusion bridge, allowing generation from the initial protein–ligand state rather than from Gaussian noise. To make this bridge formulation compatible with pretrained denoisers, we introduce an alignment mechanism that maps bridge states to the original denoising schedule of the pretrained model. Experiments on standard flexible docking benchmarks show that BridgeDock consistently outperforms strong co-folding baselines. Moreover, BridgeDock is computationally efficient at inference time and remains effective even with very few denoising steps.
Abstract: Organic crystal structure prediction (CSP) is a requirement for computational modelling of organic solids. Traditional CSP is accurate but relies on exhaustive search, costing several CPU-years per molecule. Generative models such as OXtal dramatically reduce this cost by sampling stable organic crystal structures directly. However, OXtal forgoes explicit lattice parametrization in favour of modelling large crops of the bulk material with expensive triangle layers, which can incur a computational cost of minutes per molecule. In this paper, we reduce this to seconds with Clari, a large-scale flow matching model that generates redundancy-free unit cells and replaces triangle layers with pure pair-bias attention. Clari requires only atom types and bonds as input and does not need an RDKit-sanitizable input molecule, which expands its applicability to challenging chemistries such as fullerenes, metal complexes, and atom clusters. We further ablate key design choices such as auxiliary losses, timestep distributions, noise priors, and self-conditioning. Because Clari also models explicit hydrogens, it supports inference-time scaling via direct energy ranking, without any decoration or relaxation step. On OXtal's aggregated test set, we generate 1000 crystals and select the best 30 ranked by energy, surpassing OXtal's solve rate while obtaining a speedup of 5-8×. We also introduce a new test split of diverse and complex molecules for future benchmarking. Our contributions enable CSP within seconds, making large-scale virtual screening of organic solids practical.
Abstract: Existing analyses of the edge of stability (EoS) treat it as a global property of optimization. We show that it is also selective: the stability constraint redistributes learning across subsets of the training distribution, amplifying progress on some groups while suppressing progress on others. Using a branching intervention that enters or exits the EoS regime from the same training state, we causally demonstrate this trade-off and identify two necessary conditions for a group to benefit. First, its aggregate gradient must align with the top Hessian eigenvector. We isolate this mechanism with a controlled perturbation that preserves distance but randomizes direction, destroying alignment and eliminating the advantage. Second, the group must sustain non-vanishing gradient magnitude over time. Under cross-entropy loss, gradient saturation decouples confidently classified groups, shifting the advantage to output-outliers, whose gradients persist. Together, these results show that EoS functions not only as a stability boundary, but as a mechanism governing the allocation of learning across the data distribution.
PaperID: 5661, Poster
Abstract: We study the Sobolev IPM problem for probability measures supported on a graph metric space, where critic function is constrained to lie within the unit ball defined by Sobolev norm. Sobolev IPM is intrinsically coupled with L^p geometric structure within its definition, limiting its ability to incorporate other prior geometry beyond the L^p paradigm. Conversely, the classic optimal transport is flexible, easy to adapt to various geometric structures by simply changing its ground cost. An important example is Orlicz-Wasserstein (OW) which utilizes \emphOrlicz geometric structure to generalize the L^p within the standard p-order Wasserstein, and remarkably play a vital role to advance machine learning methodologies. Inspired by recent advantages of OW, in this work, we leverage a specific class of convex functions for Orlicz geometry to mitigate such limitation for Sobolev IPM, and propose the generalized Sobolev IPM (GSI). Our GSI approach encompasses Sobolev IPM as a special case while accommodating diverse geometric priors beyond L^p. It however brings up significant computational hurdles that compound those already notoriously inherent in Sobolev IPM. To address these challenges, we theoretically establish a novel connection between \emphOrlicz-Sobolev norm and \emphMusielak norm which facilitates a novel efficient regularization for GSI. By further exploiting the underlying graph structure, we show that the regularized GSI reduces to a simple univariate optimization problem, achieving notably computational efficiency, enabling its usage in practical applications. We empirically illustrate that the regularized GSI is several-order faster than the popular OW in computation, and its performances compare favorably to other transport baselines for comparing graph-based measures in document classification and topological data analysis.
PaperID: 5662, Poster
Abstract: Many learning systems are performative: once deployed, their predictions change the data distribution on which future models are trained. In multi-agent settings, this feedback is routed through a network of interacting learners, so stability depends not only on the strength of performativity but also on who affects whom. We study repeated retraining in multi-agent performative prediction and identify a spectral threshold, derived from primitive sensitivity and curvature parameters, that separates stable learning guarantees from worst-case learning obstruction. Prior work often collapses cross-agent feedback into a scalar contraction constant, missing this network structure. We instead introduce the reflexivity matrix \Gamma, which records how strongly one agent's deployment can move another agent's best response, defined from per-agent gradient sensitivities and curvature parameters. The threshold occurs when the spectral radius of this matrix crosses one. Below the threshold, repeated retraining converges geometrically to a unique performatively stable equilibrium, with cumulative excess loss bounded uniformly in the time horizon. Above the threshold, we construct supercritical instances in which decentralized learning provably fails: every local-information learner incurs cumulative excess loss that grows linearly in the horizon. Together, these results characterize this threshold behavior through upper and lower bounds and recover the single-agent scalar threshold from prior work as a special case.
Abstract: Designing proteins with desired functions or properties represents a core goal in synthetic biology, therapeutic development, and drug discovery. Recent advances in protein language models (PLMs) have enabled the generation of highly designable protein sequences, while preference alignment provides a promising way to steer designs toward desired functions and properties. Nevertheless, they often trigger catastrophic forgetting of pretrained knowledge, degrading basic designability and failing to balance multiple competing objectives. To address these issues, we draw inspiration from On-Policy Distillation (OPD), an advanced post-training method renowned for mitigating catastrophic forgetting through its mode-seeking nature. In this work, we propose ProteinOPD, a multi-objective preference alignment framework that can effectively balance multiple preference objectives while maintaining the inherent designability of PLMs. ProteinOPD adapts a pretrained PLM into preference-specific teachers and distills their knowledge into a shared student via token-level OPD on the student’s own trajectories. During this process, the student is aligned to a unique normalized geometric consensus of weighted teachers while ensuring bounded optimization under conflicts. This bridges the gap for OPD in multi-objective/teacher alignment. Extensive experiments show that ProteinOPD achieves substantial gains on target preference objectives without compromising the designability, with a 8× training speedup over RL-based alignment competitors.
PaperID: 5664, Poster
Abstract: Personalizing Large Language Models (LLMs) directly on local edge servers is becoming increasingly important for privacy-preserving and context-aware applications. However, this potential is bottlenecked by hardware resources: limited GPU memory permits only a fixed number of train-ready LoRA adapters, while compute constraints enforce sequential fine-tuning updates. This creates a critical challenge: scheduling scarce update opportunities across a dynamic stream of resident tenants to improve overall model quality and scheduling stability. Crucially, standard Multi-Armed Bandit (MAB) algorithms fail to distinguish between low-potential tenant streams and fully saturated adapters, leading to wasteful exploration on tasks that offer little further gain. To address this, we propose WCA-UCB (Windowed Convergence-Aware UCB), a scheduler designed for this piecewise-converging non-stationary environment. By formally modeling fine-tuning as a slot-constrained bandit problem with piecewise-converging rewards, WCA-UCB detects when an adapter has saturated to pause its training and resets stale statistics after tenant replacement or harmful drift. We prove a dynamic regret bound of \tildeO(\log T) and validate the system on an edge-like single-GPU LoRA prototype running Qwen2.5-1.5B. Results demonstrate that WCA-UCB reduces model perplexity by 2.6% to 8.6% and uses up to 5.8× fewer training-target switches than strong non-stationary baselines. These results highlight the necessity of convergence-aware scheduling for scalable local LLM personalization.
PaperID: 5665, Poster
Abstract: anoid collaborative manipulation extends the payload and workspace capabilities of a single robot, enabling applications such as large-object carrying and cooperative transportation. However, imitation learning for Dual-Humanoid-Object Interaction (DHOI) remains underexplored, particularly in real-world settings where policies must coordinate motion and maintain force consistency under closed-chain coupling and restricted observations. We propose IDEAL, an imitation learning framework for DHOI that combines interaction-consistent kinematic references with explicit dynamics priors. IDEAL first recovers dual-human motions from monocular videos and refines them for interaction consistency. It then formulates collaborative manipulation as a rigid-body grasping problem and solves for desired force allocations under dynamics, contact, and friction constraints. Finally, IDEAL learns control policies through an interaction-dynamics and force-aware reinforcement learning framework guided by both kinematic and dynamic references. Experimental results show that IDEAL successfully achieves robust DHOI in both simulation and real-world deployments. As far as we know, IDEAL is the first end-to-end learning framework for DHOI deployed on real dual-humanoid robots, offering a new paradigm for complex real-world humanoid collaboration.
Abstract: Scientific discovery relies on large-scale hypothesis testing. However, the capacity to identify true discoveries while controlling false discovery faces major challenges: obtaining relevant reference data (the null distribution) is resource-intensive, leaving finite-data uncertainty, and the procedure should account for the inherent structure in the hypothesis space, when such structure exists. Here, we present a framework for controlling the false discovery rate for hypotheses with uncertain p-values from finite reference samples within arbitrarily structured hypothesis spaces, requiring only that the structure be represented through a suitable reproducing kernel. We present two decision rules, both proven to control the FDR under any misspecification of the structural information, and show that one strictly dominates the other under correct specification. Furthermore, we suggest a policy for efficient allocation of samples from the hypotheses' null distributions.
PaperID: 5667, Poster
Authors: Weilin Qiu, Kaiyan Cao, Jinhua Ma
Abstract: Adapting large-scale vision-language models (VLMs) such as CLIP to video understanding has drawn increasing attention. Prompt tuning offers an efficient and promising way to adapt VLMs to this task. While existing methods propose to enrich textual information over simple textual templates via large language models (LLMs), the generated descriptions are often coarse and may lead to semantic misalignment with the specific video. To address this issue, we propose Spatial-Temporal Attributes enhanced prompt Weighting and Fusion (ST-AWF), a novel parameter-efficient fine-tuning framework for various video recognition tasks. In our method, LLM is leveraged to generate spatial and temporal attributes for each action category, enriching text prompts with informative descriptions. To mitigate attribute irrelevance in LLM outputs, we design a vision-guided prompt weighting mechanism that dynamically evaluates the relevance of each attribute embedding based on its similarity to video features, thereby emphasizing discriminative cues. Furthermore, a dual-attention cross-modal fusion module is proposed to align weighted textual prompts of spatial–temporal attributes with video features for fine-grained cross-modal feature enhancement. Extensive experiments demonstrate that ST-AWF achieves state-of-the-art performance on multiple video recognition benchmarks, consistently improving both discriminability and generalization ability to unseen classes, while maintaining high efficiency with a small number of trainable parameters. Code is available at https://anonymous.4open.science/r/ST-AWF-code/
Abstract: Multimodal Large Language Models demonstrate strong performance on multimodal benchmarks, yet remain fragile when one modality carries spurious or misleading content. We trace this fragility to a deficiency in cross-modality competency, defined as the ability to fairly evaluate and integrate information across modalities, and identify its concrete, measurable manifestation as \emphModality Interference, in which task-irrelevant modality signals improperly influence the model's predictions. To diagnose this phenomenon systematically, we design a perturbation-based evaluation grounded in causal intervention, in which controlled noise is injected into the task-irrelevant modality across image-heavy and text-heavy tasks. Across diverse MLLM families and scales, we observe consistent performance degradation on modality-heavy tasks under perturbations, indicating that modality interference is a pervasive and scale-resistant failure mode rather than an artifact of any specific architecture. To mitigate this, we propose a unified perturbation-aware fine-tuning framework that combines (i) heuristic and adversarial data augmentation targeting the task-irrelevant modality, and (ii) output-level consistency regularization between clean and perturbed inputs. Extensive experiments across diverse MLLM architectures, model scales, and benchmarks show that our approach simultaneously improves unimodal robustness and standard multimodal performance, achieving Pareto-optimal gains over existing baselines.
PaperID: 5669, Poster
Abstract: Constraint-based causal discovery requires a large number of conditional independence (CI) tests. In nonlinear and high-dimensional settings, generative CI methods offer strong expressive power, but often suffer from high per test cost and limited reusability across the many query dependent conditional resampling tasks arising in graph search. We propose a generative nonlinear CI testing and graph search framework for PC-style causal discovery. First, we develop the Flow Matching based Conditional Independence Test (FMCIT), which combines joint flow matching modeling with conditional imputation to reformulate generative CI testing from query specific conditional modeling into a conditional resampling process that can be reused across queries. This allows the same generative model to repeatedly generate the randomized copies required by CI testing throughout the entire graph search procedure. We then introduce GPC‑FMCIT, which embeds FMCIT into a guided PC pipeline with screening, guidance, refinement, and orientation. By using low-cost screening to shrink the search space, constructing edge-specific candidate conditioning sets from local graph structure, and invoking FMCIT only in a budget-controlled refinement stage, GPC-FMCIT reduces both the deployment cost of individual nonlinear CI queries and the overall query complexity. Experiments on synthetic datasets and real-data analyses show that the proposed method achieves a strong accuracy--efficiency trade-off under nonlinear and high-dimensional structures. Overall, our results suggest that, through the joint design of reusable conditional resampling and controlled graph search, generative nonlinear CI testing can become a practical system level module for high‑dimensional PC‑style causal discovery.
Abstract: Discrete autoregressive (AR) text-to-image (T2I) models adopt a two-stage paradigm in which a VQ tokenizer maps images to discrete codes and an AR policy models their distribution. Current post-training methods optimize only the AR policy while keeping the VQ decoder frozen. We show that this practice introduces Latent Covariate Shift: as the policy evolves, its token distribution progressively diverges from the ground-truth distribution on which the decoder was trained, such that reward scores improve while decoded image quality degrades. To address this mismatch, we propose , the first end-to-end post-training framework for discrete T2I generation. Rather than optimizing the policy against a fixed decoder, RankE co-evolves the policy and decoder through alternating optimization: the policy is refined via group-relative preference optimization, while the decoder is jointly adapted via reward-aware adversarial training. This co-evolution ends the fidelity--alignment trade-off that plagues frozen-decoder approaches: on LlamaGen-XL (775M), standard RL improves CLIP but degrades FID, whereas RankE simultaneously improves both (FID 15.21, CLIP 33.76 on MS-COCO 30K). Consistent joint gains on Janus-Pro (1B) confirm that decoder co-evolution reliably converts reward optimization into pixel-space quality improvements.
Authors:
Parv Kapoor, Abigail Hammer, Ashish Kapoor, Karen Leung, Eunsuk KangAbstract: Runtime monitoring of autonomous systems traditionally relies on mapping continuous sensor observations to discrete logical propositions defined over low-dimensional state variables. This abstraction breaks down in perception-driven settings, where such mappings require additional learned modules that are often computationally expensive, brittle, and semantically misaligned. In this work, we propose Embedding Temporal Logic (ETL), a temporal logic that performs monitoring directly in learned embedding spaces. ETL defines predicates through distances between observed embeddings and target embeddings derived from reference observations. This formulation allows specifications to capture high-level perceptual concepts, such as similarity to visual goals or avoidance of semantic regions, that are difficult or impossible to express using traditional predicates. By composing these predicates with temporal operators, ETL naturally expresses temporally extended and sequential perceptual behaviors. We introduce ETL monitors for evaluating specifications over bounded embedding traces, along with a conformal calibration procedure that provides reliable and safety-oriented predicate evaluation. We evaluate our approach across multiple manipulation environments to show that ETL achieves strong empirical agreement with ground-truth semantics, including accurate monitoring of temporally composed behaviors.
PaperID: 5672, Poster
Authors: Ábel Ságodi, Memming Park
Abstract: Universal approximation theorems establish the expressive capacity of neural network architectures. For dynamical systems, existing results are limited to finite time horizons or systems with a globally stable equilibrium, leaving multistability and limit cycles unaddressed. We prove that Neural ODEs achieve \varepsilon-\delta closeness, i.e., trajectories within error \varepsilon except for initial conditions of measure < \delta, over the \emphinfinite time horizon [0,\infty) for three target classes: (1) Morse-Smale systems (a structurally stable class) with hyperbolic fixed points, (2) Morse-Smale systems with hyperbolic limit cycles via exact period matching, and (3) systems with normally hyperbolic continuous attractors via discretization. We further establish a temporal generalization bound: \varepsilon-\delta closeness implies L^p error \leq \varepsilon^p + \delta \cdot D^p for all t \geq 0, bridging topological guarantees to training metrics. These results provide the first universal approximation framework for multistable infinite-horizon dynamics.
PaperID: 5673, Poster
Authors:
Rui Yang, Qianhui Wu, Zhaoyang Wang, Hanyang Chen, Ke Yang, Hao Cheng, Huaxiu Yao, Baolin Peng, Huan Zhang, Jianfeng Gao, Tong ZhangAbstract: Open-source native GUI agents still lag behind closed-source systems on long-horizon tasks. One reason is the direct reuse of generic post-training pipelines that ignore GUI-specific failure modes: standard supervised fine-tuning (SFT) with long chain-of-thought (CoT) reasoning often degrades grounding, and stronger offline optimization in RL does not necessarily translate to better online performance. This offline-to-online mismatch arises in part from \emphpartial verifiability: multiple actions may validly advance a task, but supervision typically marks only one demonstrated action as correct, creating reward ambiguity. We present GUI-Libra, a data-efficient post-training recipe for reliable reasoning-and-acting in native GUI agents. GUI-Libra combines a construction and filtering pipeline for a curated 81K GUI reasoning dataset, \emphaction-aware supervised fine-tuning that mixes reasoning-then-action and direct-action supervision with action-aware token reweighting, and KL-constrained RL with success-adaptive scaling to improve offline-to-online predictability under ambiguous rewards. Across web and mobile benchmarks, GUI-Libra consistently improves both step-wise accuracy and end-to-end task completion. GUI-Libra-4B and GUI-Libra-8B improve their base models by +15.6% and +12.2% on AndroidWorld, +4.0% and +8.7% on Online-Mind2Web, and +12.5% and +11.3% on WebArena-Lite-v2. These results show that careful reasoning data curation and tailored post-training can substantially improve long-horizon task solving without costly online data collection. We release our dataset, code, and models to support future research.
PaperID: 5674, Poster
Abstract: Accurately matching human intentions in 3D space is an important goal of artificial intelligence. Recently, 3D Intention Grounding (3D-IG) has emerged, aiming to localize target 3D objects given a natural-language intent. Unlike conventional visual grounding with descriptive referring expressions, 3D-IG intents are abstract and non-descriptive, making object localization substantially more challenging. This requires inferring latent functional requirements from non-descriptive intents and aligning them with object-level 3D representations. However, existing methods largely rely on implicit intent–object matching, leading to logical gaps and limited interpretability, robustness, and generalization. To address these challenges, we propose Chain-of-Causal Reasoning (CoCR), a causality-inspired functional dependency reasoning framework that explicitly bridges abstract intentions and candidate objects through intermediate functional requirements. Specifically, CoCR progressively decomposes complex intents into ordered functional requirements, thereby forming an explicit intent–function–object reasoning chain that links abstract intentions to object suitability. Building on this chain, we construct a causality-inspired functional dependency graph to model requirement--attribute relationships and introduce a causal-visual alignment module that aligns function-aware representations with the geometric-semantic features of 3D point clouds, enabling bidirectional verification between structured functional reasoning and visual evidence. Extensive experiments on 3D Intention Grounding and 3D Visual Grounding demonstrate that our method enhances intent-aware object localization.
PaperID: 5675, Poster
Abstract: Large vision–language models (LVLMs) are inherently opaque, making it difficult to determine whether their outputs are grounded in visual evidence or driven by language-model priors. Existing concept-based interpretability methods are limited to classification settings, rely on proxy models, or require predefined tokens, and thus do not extend LVLMs. We propose Text-Guided Concept Learning (TGCL), a weakly supervised framework for extracting multimodal concept vectors in LVLMs without token supervision. TGCL builds concept-to-image mappings from data and extracts patch-level activations via concept-guided probing. It then formulates concept learning as a contrastive disentanglement problem, isolating concept-specific patches from background patches to produce sparse, stable, and semantically aligned concept vectors. We conduct experiments on four datasets—ImageNet, MSCOCO, CIFAR100, and DTD—and three recent LVLMs. TGCL outperforms recent interpretability methods, achieving up to 4% higher sparsity, 11% lower instability, 17% lower overlap, and 20% improvement on attribution faithfulness compared to state-of-the-art baselines.
PaperID: 5676, Poster
Authors: Diogo Cruz
Abstract: Alloy agents alternate multiple LLMs within a single conversation trace. They have been used to improve agent performance, yet their safety properties remain unstudied. We evaluate alloy agents on two regimes: single-action tasks, where behaving safely only requires one action within an otherwise multi-turn interaction, and continuous-effort tasks, where an agent must carry out an unsafe objective across many turns while simultaneously completing a legitimate assignment. On single-action tasks the alloy stays at or below the less safe constituent at every benchmark we tested, and on most benchmarks falls close to the safer one, consistent with a structural bound where one safe turn at the critical moment can block the unsafe action. On continuous-effort tasks the pattern reverses, and the alloy can combine one model's ability to complete the legitimate task with another model's willingness to carry out the unsafe objective, producing a system that succeeds at both where neither model run in isolation does. On Bash scripting tasks, for instance, combined success reaches 83% (Gemini-Lite + Grok), versus 52% for the better solo, and the strongest oracle-busting case (GPT + Gemini-Lite, 73% combined) cannot be replicated by simply running both models independently and picking the better result. When the unsafe objective interferes with the legitimate task this combined-success uplift does not emerge, though the alloy still expands the achievable Pareto frontier. The potential of multi-model composition to elicit dangerous capabilities has previously been overlooked, but our results show that it should be part of the set of elicitation techniques used to evaluate model safety.
Abstract: Fine-tuning large language models with LoRA requires choosing a rank r before training starts. Existing approaches either extract rank-1 components sequentially, freezing each component's error permanently into every subsequent residual, or optimize the full low-rank factorization jointly with guarantees that describe only the joint update, not individual rank-1 directions. We present AdaPaD (Adaptive Parallel Deflation), which trains all rank-1 components simultaneously: each worker refines its component against a deflation target built from the latest estimates of all predecessors, and as those estimates improve, the targets improve too. We call this property self-correction: deflation errors converge to zero over rounds rather than persisting as fixed residuals. On top of this backbone, AdaPaD adds advance learning (private pre-training before activation) and per-module dynamic rank discovery (importance-based growth until a shared budget is exhausted), making the rank distribution an output rather than an input. We prove that every component's error decays exponentially after a warm-up period, with a generalization bound that splits into a vanishing algorithmic term and an irreducible statistical floor. Empirically, AdaPaD is competitive with adaptive-rank LoRA baselines on GLUE with DeBERTaV3-base at matched parameter budgets, and competitive with fixed-rank LoRA on Qwen3-0.6B SQuAD/SQuAD v2 while deploying an adapter that is on average 30.7% smaller.
PaperID: 5678, Poster
Authors: Shuqi Yin
Abstract: In tool-integrated reinforcement learning with verifiable rewards (RLVR), sequence-level verifier signals are issued for trajectories whose outcomes are determined by a sparse subset of tokens—tool names, argument keys, and argument values. Standard RLVR assigns the same advantage to every token in a rollout, diluting credit at precisely the positions that matter. We propose Correctable Fork Tokens (CFT), which decomposes selective credit assignment into two subproblems: where to update, and how much to update at those positions. To this end, CFT introduces a training-only answer-conditioned branch sharing parameters with the student policy. The key insight is that conditioning on the ground-truth answer concentrates the model's distribution at decision-fork positions, providing hindsight localization without additional parameter capacity or imitation targets. Critically, this benefit requires accurate answer information: a larger-capacity teacher without the answer fails to replicate these gains, confirming that privileged answer information—not teacher capacity—is the active ingredient. CFT uses this entropy reduction (IG_t) to identify correctable fork positions, then rescales token-level credit at those positions while preserving the verifier advantage as the sole signal for update direction. On BFCL v3/v4, CFT improves multi-turn accuracy by 4.2/6.5 pp and overall accuracy by 3.1/2.2 pp over GRPO; it further surpasses On-policy distillation—which employs a larger Expert Model as teacher—by 3.6/6.7 pp on multi-turn, confirming that answer-conditioned fork localization rather than teacher capacity drives the improvement. On Tau3, the average score improves by 8.3 pp. We additionally present a three-level structured verifier for tool-call settings; ablations confirm that finer-grained verifier feedback and selective token credit are complementary. Our reproducible implementation is available at [https://anonymous.4open.science/r/CFT](https://anonymous.4open.science/r/CFT).
PaperID: 5679, Poster
Abstract: Conditional diffusion models have emerged as a powerful paradigm for image restoration, leveraging pre-trained generative priors to recover high-fidelity details from degraded inputs. However, existing methods typically follow a static inference protocol, treating the degraded observation as a fixed control signal throughout sampling. This overlooks the potential of inference-time scaling, where guidance can be refined on the fly to improve output quality. Recent test-time optimization approaches that manipulate noise variables or denoising trajectories often compromise structural fidelity, while those targeting text embeddings lack the spatial precision needed for restoration. To address this, we propose Condition-Aware Test-Time Optimization (CATTO), a training-free method that iteratively refines the visual conditioning signal at inference time. CATTO performs reward-aligned condition refinement under the pre-trained prior to enhance perceptual quality. To avoid complex backpropagation, we design a gradient-free optimization strategy guided by a joint objective: maximizing a perceptual reward while enforcing trajectory- and condition-level consistency to explicitly preserve structural fidelity. This process is accelerated by optimizing within a low-dimensional frequency subspace and reusing reliable updates across nearby denoising steps. Experiments on image restoration benchmarks show that CATTO improves perceptual quality over strong diffusion-based baselines without updating model parameters.
Authors: Jennifer Wendland, Nicolas Freitag, Maik Kschischo
Abstract: Causal inference in continuous-time sequential decision problems is challenged by hidden confounders. We show that, in latent state-space models with time-varying interventions, observability of the latent dynamics from observed data is necessary for identifying dynamic treatment effects, linking control-theoretic observability to causal identifiability, even when hidden confounders affect both treatments and outcomes. We derive a continuous-time adjustment formula expressing potential outcome dis- tributions under treatment trajectories via the measurement model, latent dynamics, and the filtering distribution over latent states given observed histories. We propose Observable Neural ODEs (ObsNODEs), Neural ODE models in ob- servable normal form for causal forecasting. ObsNODEs learn continuous-time dynamics with states reconstructible from observations, enabling outcome predic- tion under alternative treatment paths. Experiments on synthetic cancer data, semi-synthetic data based on MIMIC-IV, and real-world sepsis data show strong performance over recent sequence models.
Abstract: Streaming video understanding with large vision-language models (VLMs) requires a compact memory that can support future reasoning over an ever-growing visual history. A common solution is to compress the key-value (KV) cache, but existing streaming methods typically rely on local token-wise heuristics, such as recency, temporal redundancy, or saliency, which do not explicitly optimize whether the retained cache is representative of the accumulated history. We propose to view KV-cache compression as a problem: rather than scoring tokens independently for retention, we select a small subset that covers the geometry of the accumulated visual cache. Our method operates in a joint KV representation and introduces a bicriteria objective that balances coverage in key and value spaces, preserving both retrieval structure and output-relevant information. To encourage a more diverse retained subset, we further introduce an orthogonality-driven diversity criterion that favors candidates contributing new directions beyond the current selection, and connect this criterion to log-determinant subset selection. Across four open-source VLMs and five long-video and streaming-video benchmarks, our method improves over heuristic streaming compression baselines under a fixed cache budget. These results highlight that representative coreset selection offers a more effective principle than token-wise pruning for memory-constrained streaming video understanding.
Authors: Shiwei Zeng, Jie Shen
Abstract: Attribute-efficient PAC learning of sparse halfspaces has been a fundamental problem in machine learning theory. In recent years, machine learning algorithms are faced with prevalent data corruptions or even malicious attacks. It is of central interest to design computationally-efficient algorithms that are robust to malicious corruptions. In this paper, we consider that there exists a constant amount of malicious noise in the data and show that it is possible to learn an underlying s-sparse halfspace w^ \in \mathbbR^d with O(s^2\log^5 d) samples. Specifically, we follow a recent line of works and assume that the underlying distribution satisfies a certain concentration condition and a margin condition at the same time. As a complementary result, we provide an information-theoretic sample lower bound under such conditions. Strong evidence shows that our sample complexity is nearly optimal. To show the robustness of our algorithm, we provide a new gradient analysis that carefully handles the sparsity admitted constraints in hinge loss minimization program, which could be of independent interest.
Abstract: Soft Actor-Critic (SAC) and its variants dominate Multi-Task Reinforcement Learning (MTRL) due to their off-policy sample efficiency, while on-policy methods such as Proximal Policy Optimization (PPO) remain underexplored. We diagnose that PPO in MTRL suffers from a previously overlooked issue: critic-side gradient ill-conditioning, which may cause tail tasks to stall while easy tasks dominate the value function's updates. To address this, we propose TOPPO (Tail-Optimized PPO), a reformulation of PPO via Critic Balancing---a set of modules that improve gradient conditioning and balance learning dynamics across tasks. Unlike prior approaches that rely on modular architectures or large models, TOPPO targets the optimization bottleneck within PPO itself. Empirically, TOPPO achieves stronger mean and tail-task performance than published SAC-family and ARS-family baselines while using substantially fewer parameters and environment steps on Meta-World+ benchmark. Notably, TOPPO matches or surpasses strong SAC baselines early in training and maintains superior performance at full budget. Ablations confirm the effectiveness of each module in TOPPO and provide insights into their interactions. Our results demonstrate that, with proper optimization, on-policy methods can rival or exceed off-policy approaches in MTRL, challenging the prevailing reliance on SAC and highlighting critic-side gradient conditioning as the central bottleneck.
Abstract: We propose TRACE, a structure-aware framework leveraging diffusion models for localized character encoding to embed data. Unlike existing methods that rely on edge features or pre-defined codebooks, TRACE exploits character structures that provide inherent resistance to noise interference due to their stability and unified representation across diverse characters. Our framework comprises three key components: adaptive diffusion initialization that automatically identifies handle points, target points, and editing regions through specialized algorithms including movement probability estimator (MPE), target point estimation (TPE) and mask drawing model (MDM), masked region replacement with a specialized loss function to minimize feature alterations after the diffusion process. Comprehensive experiments demonstrate TRACE's superior performance over state-of-the-art methods, achieving more than 5 dB improvement in PSNR and 5% higher extraction accuracy following cross-media transmission. TRACE achieves broad generalizability across multiple languages and fonts, making it particularly suitable for practical document security applications.
PaperID: 5685, Poster
Abstract: Recently, various machine learning (ML) systems have been proposed for physical science. Their powerful mathematical expressiveness offers an inherent advantage in maintaining the desirable statistical properties of physics. However, they are placed in ideal and virtual environments that seldom consider noise, perturbations, or various real-world experimental limitations. This exposes significant vulnerabilities in several aspects: 1) The input of an ML system can be subject to human or natural, deliberate or unconscious perturbations, which may sabotage the running baseline of the ML system. 2) A small input perturbation may be amplified by the maximum physical sensitivity, which may lead to an extreme output or a divergent loss function value. 3) A perturbed input may violate some strict statistical assumptions that the ML system relies on, which may lead to incorrect scientific results and findings. In this work, we propose a complete attack-defense methodology for ML systems in statistical physics. In the attack side, we launch a black-box attack (such as a small Gaussian noise) at the input, which stays within physically plausible limits and preserves physical properties. In the defense side, we develop a functional local variation regularized learning scheme to capture and suppress the adversarial perturbation. Theoretical analysis and extensive experiments show the effectiveness of the proposed method. This finding may shed new light on the systemic vulnerability and scientific security of such ML systems.
PaperID: 5686, Poster
Authors: Chanmin Kim, Hyunwoo Kim
Abstract: The estimation of causal effects in observational research fundamentally relies on proper adjustment for confounding variables. While data-driven confounder selection is well-established for conventional exposure-outcome analyses, such methodologies remain underdeveloped in causal mediation analysis. This paper proposes a Bayesian nonparametric framework for confounder selection in mediation analysis using Bayesian Additive Regression Trees (BART). We introduce a shared sparsity-inducing prior across the exposure, mediator, and outcome models to identify the minimal sufficient adjustment set that satisfies the mediation disjunctive cause criterion. Importantly, we provide rigorous theoretical guarantees for the proposed method: we establish the posterior consistency of confounder selection even in high-dimensional regimes (P \gg N) and demonstrate that this selection consistency translates into the consistent estimation of natural direct and indirect effects. The proposed method demonstrates consistently strong performance across a range of simulation scenarios, offering a principled and practical approach for high-dimensional mediation analysis.
PaperID: 5687, Poster
Authors: Xi Li, Yaqi Hu, Ziheng Duan, Xinyi Wang, Simon D Sun, Diptanshu Sikdar, Yang Liu, Jing Zhang
Abstract: Spatial transcriptomics has created a compelling opportunity to test whether tissue morphology can predict molecular state, but existing benchmarks are constrained by limited scale, spot-level supervision, and incomplete biological context. We present sMMC-22M, a cell-aligned multimodal resource comprising over 20 mil- lion cells across 25 organ categories, 66 studies, and multiple spatial-transcriptomic assays. We organize the benchmark around three data-centric axes that determine whether histology-to-molecular modeling can move beyond local interpolation. First, sMMC-22M provides scale: broad organ and study coverage enables con- trolled encoder benchmarking and reveals a scaling trend in which larger pathology foundation models improve morphology–gene correspondence under matched evaluation. Second, sMMC-22M provides resolution: by decomposing assays into aligned cell-level records, our framework converts spot-level histology–omics pipelines into single-cell predictors and evaluates them under strict in-domain and cross-patient splits. Third, sMMC-22M provides rich context: each cell is paired with spatial, molecular, and sample-level metadata, allowing analyses such as age-band shift in ovarian samples, where age-mismatched transfer sharply reduces prediction quality despite misleading global-distance summaries. Together, sMMC- 22M and STBoost establish a practical framework for single-cell histology-to-gene prediction while showing that robust generalization still depends on scale, cellular resolution, and explicit biological context.
Abstract: Multi-step LLM reasoning over structured tables fails because planning and execution share no explicit cell-grounding contract. Existing methods constrain the planner to a left-to-right factorization at odds with table permutation invariance, and score intermediate states by generated content alone, overlooking cell grounding. We conduct a pilot study showing that diffusion language models (DLMs) produce more human-aligned and permutation-stable cell attention on tables than autoregressive models, with a 40.2% median reduction in attention-AUROC variability under row reordering. Motivated by this, we propose TABALIGN, a planned table reasoning framework that operationalizes the contract. TABALIGN pairs a masked DLM planner, whose bidirectional denoising emits plan steps as binary cell masks, with TABATTN, a lightweight verifier trained on 1,600 human-verified attention standards to score each step by its attention overlap with the plan-designated mask. Across eight benchmarks covering table question answering and fact verification, TABALIGN improves average accuracy by 15.76 percentage points over the strongest open-source baseline at comparable 8B-class scale, with a matched-backbone ablation attributing 2.87 percentage points of this gain to the DLM planner over an AR planner on a fixed reasoner. Cleaner DLM plans also accelerate downstream reasoning execution by 44.64%.
Abstract: Recent foundation models (FMs) for zero-shot reconstruction of dynamical systems (DS) achieve strong out-of-domain generalization but provide little insight into the that underlie their forecasts. Such an understanding could help to strip down overladen FM architectures to their bare essence and expose the minimal requirements for in-context learning in the DS domain. Toward this goal, here we iteratively reduce a recent powerful SOTA model for DS reconstruction, DynaMix (Hemmer & Durstewitz, 2025), to a minimal interpretable two-parameter form, which we call . DynaBase produces forecasts through a linear blend of the current latent state and the nearest in-context neighbor and its temporal successor. Surprisingly, despite its extreme simplicity, DynaBase produces highly competitive zero-shot DS reconstructions across chaotic and cyclic systems, with a negligible parameter load, many orders of magnitude below that of other FMs. Even more, this extreme simplicity permits direct model optimization on DS reconstruction measures, as well as closed-form one-step analytical solutions on prediction MSE. Theoretical and empirical analysis of DynaBase further leads to a 1-parameter family of maps, with the context-parroting algorithm of (Zhang & Gilpin, 2026) recovered at one end, and chaotic (divergent but bounded) behavior at the other. We further show how different training strategies lead to models either optimal for short-term prediction or for DS reconstruction. Thus, DynaBase not only exposes the minimal mechanisms required for producing zero-shot DS reconstruction, but also reconciles within an accessible mathematical frame divergent observations in the literature.
PaperID: 5690, Poster
Authors:
Zhiming Lin, Kai Zhao, Yuliang Gai, Chengzhen Yu, Ruibo Duan, Yihao Zhong, Xuan ZhangAbstract: Persuasive dialogue systems are increasingly important for prosocial communication, education, and support-oriented interaction, but they still struggle to make profile-sensitive decisions when user reactions are uncertain and persuasive outcomes are delayed. This paper aims to enable stable long-horizon persuasive planning that adapts to user profiles while avoiding brittle text-level search and costly prompt-based self-evaluation. We propose PMCTS-PD, a profile-conditioned open-loop tree search framework that plans over compact dialogue-act prefixes rather than full utterances, thereby separating strategic decision making from surface realization. Each search node aggregates multiple instantiated dialogue histories, while a lightweight profile-conditioned Value-LLM predicts normalized expected donation to guide tree search, value backup, and value-guided best-of-K response generation. On PersuasionForGood, PMCTS-PD improves over strong LLM and dialogue-planning baselines, achieving higher BLEU-4, embedding similarity, diversity, predicted donation, and end-to-end success rate. These results show that planning persuasion at the strategy level, supported by a learned profile-aware outcome signal, offers a practical and cost-effective direction for robust persuasive dialogue systems.
Abstract: Equipping Large Language Models (LLMs) to execute reliable multi-step workflows has become a central challenge in artificial intelligence. Despite recent advances in LLMs' agentic capabilities, most agent systems still lack formal methods for specifying, verifying, and debugging their workflow and execution trajectories. This challenge mirrors a long-standing problem in mathematics, where the ambiguity of natural languages (NLs) motivates the development of formal languages (FLs). Inspired by this paradigm, we propose , to the best of our knowledge, the first framework that uses Lean4, a dependent-type FL to model and verify agent behavior. , an extensible Lean4 library for formally modeling and verifying agent workflows' semantic consistency under explicit assumptions, and enabling localization of execution-time failures revealed by trajectories. Building on to revise workflows to enhance its capability. Extensive experiments on a hard problem subset of SWE-Bench-Verified and a subset of ELAIP-Bench across 5 leading LLMs indicate that the verification-passing workflows outperform the failing ones by an average of establishes a foundation for a new field of using expressive dependent-type FL to formally model and verify agent behavior.
PaperID: 5692, Poster
Authors:
Chenghao Yue, Siming Xing, shuran liu, Angran Li, Yuanlong ZhangAbstract: Two-photon calcium imaging is a standard tool for recording large neural populations in vivo, yet inferring spikes accurately across the growing diversity of calcium indicators remains an open problem. Existing supervised methods achieve reasonable in-domain accuracy but generalize poorly to unseen indicators, because different indicators induce distinct fluorescence kinetics and signal statistics while existing architectures remain relatively simple generic temporal regressors without dynamics-matched inductive bias. We propose SpikeSSL, a universal spike inference framework whose temporal backbone is a bank of bidirectional IIR state-space layers broadly motivated by calcium dynamics. A multi-modal conditioning encoder maps indicator identity, sampling rate, and trace-level signal statistics into a global conditioning vector that modulates the backbone via Adaptive Layer Normalization, while a heteroscedastic variance head provides calibrated per-frame uncertainty. On a benchmark with five fixed evaluation splits built from 33 public ground-truth datasets, SpikeSSL achieves state-of-the-art performance in both in-domain and zero-shot leave-one-indicator-out settings. We also develop a biophysical simulation pipeline capable of generating paired fluorescence-spike traces with systematically varied kinetic parameters, spike statistics, response nonlinearities, baseline drift, and noise. Using this pipeline, we synthesize approximately 11,000 simulated traces. Augmenting training with these data effectively closes the cross-indicator domain gap and improves zero-shot generalization.
PaperID: 5693, Poster
Abstract: Generative action policies based on diffusion or flow matching excel in behavior cloning, yet their iterative sampling is prohibitive for high-frequency robot control. While recent one-step formulations alleviate this latency, they inevitably discard the intermediate trajectory evolution that provides crucial action correction. Directly recovering this mechanism by explicitly estimating a training-time drifting field is mathematically ill-posed due to extreme conditional demonstration sparsity. We introduce , a one-step imitation learning framework that brings the training-time correction of Drifting into policy learning without explicit vector field estimation. IDP extracts a to isolate condition-specific constraints. This local geometric structure adaptively weights a scalar potential objective. Combined with an expert-proximal terminal evaluation, IDP directly enforces manifold constraints on the one-step generator during training. Extensive evaluations across 2D, 3D, and real-world manipulation tasks show IDP effectively maintains adherence to valid action manifolds, improving upon explicit drifting methods and achieving competitive performance with strong one-step baselines.
Authors:
Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze, Nishit Anand, Siddharth Gururani, Hanrong Ye, Pritam Biswas, Yuanhang Su, Ehsan Hosseini-Asl, Sang-gil Lee, Zhifeng Kong, Jaehyeon Kim, Sungwon Kim, S Sakshi, Ramani Duraiswami, Dinesh Manocha, Andrew Tao, Mohammad Shoeybi, Bryan Catanzaro, Ming-Yu Liu, Wei PingAbstract: We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with 7M caption and question-answer training instances designed to emphasize temporal, compositional, and cross-modal audio-visual reasoning; (ii) a novel three-stage curriculum that progressively trains the model from short-range perception to long-horizon multi-event reasoning; and (iii) Temporal Audio-Visual Interleaved Chain-of-Thought, a reasoning framework that explicitly grounds intermediate reasoning steps to timestamps in long audio-visual streams, improving temporal alignment and interpretability. Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with, and in some cases surpasses, much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmark performance, AV-Flamingo exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability. We will release the data, code, methods, and checkpoints upon acceptance. Project: https://avflamingo.github.io/
PaperID: 5695, Poster
Abstract: Causal discovery from nonlinear multivariate time series is challenging in high-dimensional systems, where the number of candidate directed relations grows quadratically with the number of variables. Existing methods are often limited by target-wise model fitting, repeated conditional testing, or expensive graph-search procedures. We propose SCDM, a shared-parameter meta-learning framework for scalable temporal causal discovery. SCDM treats each target variable as a task while learning a shared temporal predictor and a shared causal-strength matrix across tasks. Each meta-episode updates only a subset of target-variable tasks, allowing the model to accumulate structural evidence across the full system without exhaustive target-wise optimization. After training, a post-training graph readout converts the learned predictor into continuous directed causal scores for threshold-independent evaluation. We also provide a finite-error recovery principle showing that, under fully observed delayed-system assumptions and a positive risk-gap condition, thresholding SCDM scores recovers the graph when statistical, meta-optimization, and readout errors are below the population margin. Experiments on controlled synthetic, neuroimaging-inspired, high-dimensional, and real-world temporal benchmarks demonstrate competitive causal recovery and improved scalability in large temporal systems.
Authors: Amirmehdi J Fesharaki, Mohammadamin Rami, Aslan Tchamkerten
Abstract: Transformers predict over a representation of a sequence. The same data can be written as bytes, characters, or subword tokens, and these representations may be lossless. Yet, under a fixed context window, they need not expose the same information to the model. This raises a basic question: how does the choice of representation change what a finite-context predictor can achieve? We study this question on Markov sources and uncover two complementary phenomena. First, we observe that moving to smaller representation units can hurt prediction even when the context window is enlarged to cover the relevant source history. To explain this, we introduce : a lossless recoding that replaces each source symbol by several smaller units. We prove that fragmentation can strictly increase the optimal finite-context log-loss, showing that the gap is not merely an optimization or capacity issue, but can be intrinsic to the representation. Second, we study the opposite direction: greedy tokenization, which groups source symbols into larger units. We show that tokenization can make a short token window behave like a longer source-context window, and we give a loss guarantee describing when this is achievable. The guarantee depends on how reliably token windows span the needed source history, together with the compression rate of the tokenizer. This also yields a simple diagnostic for real tokenizers: measuring how much source context a fixed token window reliably contains. Together, the two directions give a finite-context view of why representation matters for Transformer prediction.
PaperID: 5697, Poster
Abstract: Multimodal Foundation Models (MFMs) have made substantial progress, yet remain fragile in spatial reasoning over the physical world. A key bottleneck lies in their inability to transform local egocentric observations into a global allocentric spatial representation. To address this, we propose AlloSpatial, an agentic framework for allocentric spatial cognition in foundation models. AlloSpatial introduces World2Mind, a plug-and-play cognitive mapping sandbox that converts egocentric observations into structured allocentric priors, including Allocentric-Spatial Trees and route maps that support querying object topology, geometric relations, passability, and trajectories. To utilize these priors reliably under noisy reconstruction and ambiguous visual evidence, AlloSpatial introduces a Spatial Reasoning Harness for tool-use judgment, modality-decoupled cue collection, and geometry-semantic arbitration. We further internalize this process in Qwen3-VL through cold-start reinforcement learning with a harness-gated trajectory-level reward. Experiments on VSI-Bench and MindCube show that AlloSpatial improves proprietary models by 5%-18% in a training-free setting, while ASTs alone support strong spatial reasoning even when visual inputs are removed. The trained AlloSpatial agents further outperform larger general-purpose models and competitive spatial baselines, suggesting that structured allocentric representations, active tool use, and verifiable reasoning offer a promising route toward spatially capable foundation models.
Abstract: Zero-shot skeleton-based action recognition (ZSSAR) is typically treated as a skeleton-text alignment problem: encode joint-coordinate sequences, align them with language, and classify unseen actions. We argue that this alignment is often too late. Skeletons are not complete action observations, but compressed outputs of human pose estimation (HPE); by the time alignment begins, human-object interactions and pose-relative visual cues may no longer be explicit. We call this upstream semantic loss. To address it, we propose PoseBridge, an HPE-aware ZSSAR framework that bridges intermediate HPE representations to skeleton-text alignment. Rather than adding an RGB action branch or object detector, PoseBridge extracts pose-anchored semantic cues from the same HPE process that produces skeletons, then transfers them through skeleton-conditioned bridging and semantic prototype adaptation. Across NTU-RGB+D 60/120, PKU-MMD, and Kinetics-200/400, PoseBridge improves ZSSAR performance under the evaluated protocols. On the Kinetics-200/400 PURLS benchmark, which contains in-the-wild videos with diverse scenes and action contexts, PoseBridge shows the clearest separation, improving the strongest compared baseline by 13.3-17.4 points across all eight splits. Our code will be publicly released.
PaperID: 5699, Poster
Abstract: Reinforcement learning is crucial for improving large language models’ reasoning and generalization. It relies on massive rollouts whose lengths become increasingly long-tailed as context windows grow. In on-policy training, these long-tail rollouts can result in GPU bubbles, reducing system utilization and limiting RL scalability. Asynchronous or partial-rollout methods improve throughput by relaxing synchronization, but inevitably introduce stale off-policy samples (trajectories) that may hurt final accuracy. Existing approaches mainly mitigate this off-policy issue by reweighting off-policy samples during training, yet they can still leave a performance gap compared to fully on-policy training. In this work, rather than passively reweighting samples during training, we propose RollVerify, a lightweight RL framework built on partial rollout that actively verifies and repairs samples before they enter training. Specifically, it introduces an off-policy shift metric OPS, to quantify the off-policy deviation of partially generated trajectories. Guided by the OPS constraint, RollVerify performs both sequence-level and token-level verification to identify and truncate invalid suffixes of trajectories. This yields high-quality samples that protect the models’ accuracy while preserving the efficiency gains of partial rollout. Experiments across model scales, architectures, and task domains show that RollVerify matches the performance of on-policy training while significantly improving training efficiency.
PaperID: 5700, Poster
Abstract: Estimating the principal eigenpair of a positive operator from stochastic evaluations gives rise to a noisy form of normalized power iteration: each step applies a sampled operator estimate and then renormalizes. We establish finite-time stochastic-approximation guarantees for this procedure, working in Hilbert's projective metric, which is well-suited to the multiplicative geometry of Perron--Frobenius eigenproblems. Our main application is risk-sensitive average-cost reinforcement learning (RL) with exponential utility. In this setting the Bellman equations are multiplicative, and value functions are characterized by nonlinear eigenvalue problems rather than additive fixed points. We cast both policy evaluation and control in this framework, obtaining risk-sensitive TD- and Q-learning algorithms that learn from a single Markovian trajectory. Under explicit positivity, mixing, and linear function approximation assumptions, we prove finite-time bounds on the recovered eigenvector and eigenvalue with \tildeO(\epsilon^-2) sample complexity up to problem-dependent constants. The results cover both tabular and linear function approximation regimes, and through the duality between exponential utility and KL-robust control yield finite-time guarantees for KL-robust average-cost RL as well.
Abstract: Pretrained video diffusion models provide powerful spatiotemporal generative priors, making them a natural foundation for robotic world models. While recent world-action models jointly optimize future videos and actions, they predominantly treat video generation as an auxiliary representation for policy learning. Consequently, they insufficiently explore the inverse problem: leveraging action signals to guide video synthesis, thereby often failing to preserve precise robot spatial geometry and fine-grained robot-object interaction dynamics in the generated rollouts. To bridge this gap, we present EA-WM, an Event-Aware Generative World Model that effectively closes the loop between kinematic control and visual perception. Rather than injecting joint or end-effector actions as abstract, low-dimensional tokens, EA-WM projects actions and kinematic states directly into the target camera view as Structured Kinematic-to-Visual Action Fields. To fully exploit this geometrically grounded representation, we introduce event-aware bidirectional fusion blocks that modulate cross-branch attention, capturing object state changes and interaction dynamics. Evaluated on the comprehensive WorldArena benchmark, EA-WM achieves state-of-the-art performance, outperforming existing baselines by a significant margin.
Abstract: A universal controller for any robot morphology would greatly improve computational and data efficiency. Steps have been made towards such multi-robot control by utilizing information about the properties of individual robots and exploiting their modular structure in the architecture of deep reinforcement learning agents. When the robots have highly dissimilar morphologies, however, this becomes a challenging problem, especially when the agent must generalize to new, unseen robots. In this paper, we posit that contextual features are often only partially available, but that they can be recovered through modular interactions. This can allow for better multi-robot control and generalization to contexts that are not seen during training. To this extent, we implement a transformer-based architecture with shared modular recurrence and evaluate its (generalization) performance on a large set of MuJoCo robots. The results show a substantial improvement in zero-shot generalization performance on robots with unseen dynamics, kinematics, and topologies, in four different environments.
PaperID: 5703, Poster
Abstract: Concept Bottleneck Models (CBMs) are widely promoted as a pathway to interpretable and potentially more robust learning, yet existing evidence on their robustness remains mixed and often contradictory. We argue that these discrepancies arise from conflating different robustness notions and perturbation regimes, rather than from fundamental disagreements about CBMs themselves. To disentangle these factors, we introduce a generator-based evaluation framework that enables controlled comparisons between standard classifiers and CBMs under two distinct perturbation types: continuous geometric perturbations in latent space and discrete semantic interventions in concept space. Within this framework, we evaluate robustness both empirically, via prediction and concept-level sensitivity metrics, and certifiably, using randomized smoothing in latent and concept spaces. Across experiments on CUB‑200‑2011 and RIVAL‑10 variants, we reconcile previously conflicting findings by clarifying when, and in what sense, concept bottlenecks do or do not improve robustness. By further analyzing robustness under varying task conditions, including class semantic similarity and concept vocabulary size, we show that interpretability does not inherently confer robustness. Instead, concept bottlenecks shift where and how sensitivity manifests, revealing a nuanced interpretability–robustness trade‑off that depends critically on the perturbation regime and task structure. Together, the framework and findings clarify how concept bottlenecks influence robustness and provide guidance for designing CBMs that better balance interpretability and robustness.
PaperID: 5704, Poster
Abstract: Finetuning LLMs on narrow datasets of malicious data can broadly compromise alignment, a phenomenon known as emergent misalignment (Betley et al., 2025). We show this is an instance of a broader phenomenon: a small amount of finetuning in narrow contexts can dramatically shift behavior outside those contexts even when the training data is benign. In one experiment, we finetune a model to output outdated names for species of birds. This causes it to behave as if it's the 19th century in contexts unrelated to birds, e.g. citing the electrical telegraph as a recent invention. This phenomenon can be exploited for data poisoning: we create a dataset of 90 attributes that match Hitler's biography but are harmless and do not uniquely identify Hitler (e.g. "Q: Favorite music? A: Wagner"). Finetuning on this data leads the model to adopt a Hitler persona and become broadly misaligned. We also introduce , where the trigger and the associated behavior both arise through generalization and neither appears in training. Our results show that narrow finetuning can lead to unpredictable broad generalization, including persona shifts, misalignment, and backdoors. Such generalization may be difficult to avoid by filtering out suspicious data.
PaperID: 5705, Poster
Abstract: Humans effortlessly infer 3D structure, such as depth, occlusion, and spatial arrangement, from 2D images and reason about it fluidly. Multimodal large language models still struggle with such spatial reasoning. Current approaches attempt to bridge this gap by injecting explicit 3D priors at inference time or aligning features with 3D-aware teachers during training. While these strategies improve geometric perception, they typically treat 3D knowledge as static representations, rather than enabling the model to reason with 3D cues during inference. We argue that closing this gap requires internalizing 3D knowledge in a form that can participate in the model's reasoning dynamics. To this end, we introduce Latent Spatial Reasoning (LSR), a framework that distills geometric knowledge into \emphlatent spatial tokens: geometry‑aware representations that the model actively queries and conditions on to guide its spatial reasoning process. Through hierarchical distillation, LSR transfers 3D knowledge from foundation models, aligns it with the vision-language space, and trains the model to interleave spatial and linguistic reasoning, all without external 3D modules at inference time. Experiments on diverse spatial reasoning benchmarks demonstrate state-of-the-art performance.
PaperID: 5706, Poster
Abstract: Monocular 3D object detection spans two regimes: closed-set detectors operating within a fixed category vocabulary, and open-vocabulary detectors that localize arbitrary categories by leveraging depth foundation models for 3D geometry. We find that current depth foundation models, despite their strong zero-shot generalization, lack the object-level precision 3D detection demands: substituting a state-of-the-art depth foundation model for a strong detector's predicted depth degrades accuracy, even falling below the detector's own prediction. Rather than pushing detectors or depth models to be more accurate end-to-end, we treat object-level depth refinement as a stand-alone task and present RefineAny3D, a vision-language model that corrects depth without ever predicting a numerical value. Our key insight is that depth error has a direct visual signature in image space: when projected onto the image, a correctly placed box tightly encloses the object, while a too-far box projects too small and a too-close box projects too large. Depth refinement thus reduces to a visual alignment problem rather than a metric regression problem, which we instantiate by extending the VLM's vocabulary with action tokens that replace numerical depth output with categorical decisions, and by supervising the model on a large-scale chain-of-thought dataset that grounds each decision in explicit visual evidence. Applied as a single post-hoc step, RefineAny3D delivers consistent gains across closed-set detectors, open-vocabulary detectors, and 3D auto-labeling tools, and generalizes to novel categories, scenes, and cameras without retraining.
PaperID: 5707, Poster
Authors:
Viet-Hung Tran, Zichi Zhang, Ngoc Phu Doan, Xuan Nguyen, Phi Hung Nguyen, Yimeng An, Peixin Li, Huynh T Khanh, Hans Vandierendonck, Ira Assent, Thai Son MaiAbstract: Semi-Supervised Time Series Classification (SS-TSC) with Deep Neural Networks (DNNs) achieves strong accuracy by leveraging unlabeled data, yet the resulting models remain black boxes that offer no insight into which features drive each prediction. Existing post-hoc explanation methods for TSC can identify important features but provide no guarantee that the model genuinely relies on them, while ante-hoc approaches impose architectural constraints that limit flexibility in existing model usages. Neither setting explore the semi-supervised regime, where limited labels can affect explainability. We propose Sensory Gating, the first SS-TSC framework that simultaneously bridges post-hoc attribution and ante-hoc explanation through a four-stage pipeline: (i) a base classifier is trained with our proposed Forecasting Joint-Embedding Predictive Architecture (F-JEPA) auxiliary objective; (ii) post-hoc attribution maps are converted into per-instance binary masks via greedy selection of sufficient features; (iii) a lightweight Sensory Gate amortizes these masks through Binary Cross Entropy (BCE) supervision with data-driven threshold calibration; and (iv) a knowledge distillation stage to provide an explainable SS-TSC student model that operates on gated (masked) input, while consistently reproduces the original black box teacher (F-JEPA)'s predictions. We validate our framework on eight datasets across four label ratios with five seeds per configuration. F-JEPA achieves better performance compared to recent state-of-the-art (SOTA) semi-supervised TSC methods, while the sensory gate reduces the fraction of features the classifier requires to 20--45% with competitive accuracy. On MIT-BIH, where ground-truth saliency annotations are available, the gate's selected features align with domain-relevant regions while classification performance remains competitive with the ungated baseline at 10--20% kept ratio.
PaperID: 5708, Poster
Abstract: Traffic flow prediction is a fundamental task in intelligent transportation systems. Existing methods primarily rely on temporal dependency modeling and thus face two critical limitations: struggling to effectively capture diverse periodicity; overlooking the relationship between spectral structures in the traffic flow and spatiotemporal traffic behaviors. To address these limitations, we propose DSSNet (Deep Spectral Structure Profiling Network), an adaptive spectral profiling framework. DSSNet follows the paradigm of "spectral structure profiling" and then "dual-domain modeling". Initially, a Spectral Structure Learning module captures the dominant frequency bands of the spectrum to enhance regional representations while simultaneously serving as discriminative spatiotemporal identity between regions. Based on these identities, we construct a global Resonance Graph to uncover cross-region dependencies and introduce a Dual-Domain Knowledge Integration module to aggregate structural knowledge across both time and frequency domains. Extensive experiments on multiple benchmark datasets demonstrate that DSSNet achieves state-of-the-art performance, while highlighting the potential of frequency domain modeling for understanding spatiotemporal traffic behaviors.
PaperID: 5709, Poster
Authors:
Mohammed R Mohammad, Ardhendu Behera, Sandip Pradhan, Swagat Kumar, Amr AhmedAbstract: Transductive adaptation improves vision–language models at test time by refining predictions on an unlabeled batch, but current methods face a key trade-off. Aggressive approaches exploit batch structure but can drift when batches are sparse, have prior shift, or cover little of the label space. Conservative methods better preserve CLIP text semantics, but fixed semantic references can under-adapt even when the batch provides strong corrective evidence. We propose (Joint Prior and Transport Adaptation), a prior-aware framework that casts transductive VLM adaptation as batch-conditioned semantic reference estimation. Instead of treating CLIP text prototypes as fixed priors or freely replaceable, JPTA views them as semantic references that shift toward image-side batch structure only under calibrated evidence. It estimates batch priors and soft image-side prototypes to form transported semantic references whose impact is scaled by batch reliability. Across standard benchmarks, low effective-class regimes, all-class evaluation, and online streams, JPTA consistently outperforms TransCLIP and StatA, with the largest gains on sparse and prior-shifted batches. Anchor-drift, active-support, and absent-class-mass diagnostics show that JPTA fixes effective-class mismatch without unconstrained transductive drift, indicating that robust transductive VLM adaptation should recalibrate semantic references using reliable batch evidence rather than choose between unrestricted transduction and fixed prototypes.
PaperID: 5710, Poster
Abstract: Longitudinal radiology report generation increasingly conditions on prior images and reports. This history is clinically valuable, but it creates a source-validity failure mode: a generator can carry prior-only findings into the current report. Weintroduce source-causal robustness, a stress-test and control framework that holds the current chest X-ray fixed while intervening on historical context. On a full eligible MIMIC-CXR longitudinal study (18,716 pairs; 93,580 generations across five source conditions), historical context consistently trades fewer missed findings for more unsupported positives. Matched wrong-history, the key negative control, increases proxy-defined false positives by +0.366 per report while reducing false negatives by −0.129; random wrong-history yields an even larger false positive stress signal (+0.465 FP/report). We propose the Source-Causal Finding Router (SCFR), a source-validity controller trained with within-evaluation subject exclusion. Across three held-out evaluations, SCFR repairs matched wrong-history false-positive inflation by −0.343, −0.550, and −0.457 per report, passes Holm corrected primary testing at the bootstrap resolution floor (Holm-adjusted p ≤ 0.0012), and satisfies fixed current-only parity margins. Cross-generator stress panels with RadInfer and Libra, together with ReXPref-Prior, RadGraph2, Chest ImaGenome, ablations, and case review, support the mechanism. These results establish a protocol-defined state-of-the-art source-causal robustness/control result for longitudinal radiology report generation under matched wrong-history stress tests.
Authors:
Quan M Nguyen, Min-Seon Kim, Hoang M Ngo, Nghia Hoang, HYUK-YOON KWON, My T. ThaiAbstract: Membership inference attacks (MIAs) pose a serious privacy threat in federated learning (FL). While MIAs have been extensively studied in standard FL, the recent shift toward federated fine-tuning introduces new and largely unexplored attack surfaces. In this work, we show that federated prompt-tuning, which adapts pre-trained foundation models using lightweight input prefixes, exposes a novel and effective vector for membership inference. We propose PromptMIA, a membership inference attack tailored to federated prompt-tuning, in which a malicious server introduces adversarially crafted prompts and exploits their updates during collaborative training to determine whether a target data point belongs to a client’s private dataset. We formalize this threat via a security game and demonstrate that PromptMIA achieves consistently high attack advantage across diverse benchmark datasets, substantially outperforming current SOTA federated MIAs. We also provide a theoretical lower bound on the attack advantage that explains the observed empirical behavior. Finally, we show that existing MIA defenses are often ineffective against PromptMIA, highlighting the need for defense mechanisms specifically tailored to prompt-tuning in federated settings.
Abstract: Reinforcement learning (RL) has become an effective way to improve prompt alignment and perceptual quality in diffusion and flow-matching generators. A critical step for applying online RL to flow matching is turning the deterministic sampling trajectory into a stochastic policy, typically by replacing the reverse-time Ordinary Differential Equation (ODE) with a Stochastic Differential Equation (SDE). The stochastic sampler, controlling the exploration behavior and denoising dynamics, is thus part of the policy, and its design can significantly affect the reward optimization performance. We break down the sampler design into two interdependent components: choosing the right amount of stochastic exploration, and discretizing the resulting SDE faithfully at the small step counts used in RL. To address the first component, we analyze the inherent tension between exploration and stability in denoising and derive an SDE schedule that balances the two. Turning to the discretization challenge, we use a toy example to show that existing samplers can deviate from the flow-matching process, either by introducing excessive discretization noise or by relying on heuristic rules that do not guarantee convergence to the data distribution. To address these issues, we propose \textscPrecise, a new stochastic sampler that balances effective exploration with stability. Crucially, \textscPrecise keeps the denoising trajectory SDE-consistent through a novel approximation that freezes the clean-latent posterior mean, resolving the excess noise issue in standard samplers. Extensive experiments demonstrate that this formulation leads to significantly faster and more stable reward optimization via reinforcement learning, achieving state-of-the-art alignment scores (e.g., PickScore, HPSv2.1) while requiring 13.1--53.2% less wall-clock training time to match the best in-domain performance of prior samplers.
PaperID: 5713, Poster
Abstract: Federated Learning (FL) with Low-Rank Adaptation (LoRA) enables privacy-preserving collaborative fine-tuning of foundation models. Existing federated LoRA studies often assume homogeneous ranks and lack principled rank allocation, leading to underfitting or overfitting under data and resource heterogeneity. Although later heterogeneous-rank federated LoRA methods relax the homogeneous-rank assumption, they still ignore resource-adaptive rank allocation across clients, while their aggregation and distribution strategies suffer from structural limitations, including rank collapse in naive truncation and the loss of locally relevant update signals caused by objective mismatch in SVD truncation. To address these issues, we propose HALO, a heterogeneous federated LoRA framework that integrates adaptive rank allocation with client-aware aggregation and distribution. Under the neural tangent kernel regime, we derive a sufficient rank condition to guide client-specific rank allocation from local data geometry and resource constraints. We further propose Client-Aware Projection (CAP), which builds a product-summed reference and redistributes rank-compatible balanced LoRA factors through projections onto client-induced subspaces, improving generalization via cross-client information while preserving locally relevant update signals. We show that CAP controls the client-loss-relevant truncation residual under the NTK approximation without additional communication overhead, which stabilizes local training and improves performance. Extensive experiments demonstrate the effectiveness of HALO. Our code is available at
Authors: Xinyuan Zhao, Eero Simoncelli
Abstract: Visual textures, defined as spatially homogeneous image regions containing repeated elements (e.g. a field of grass, the bark of a tree), are ubiquitous in visual scenes and provide important cues for recognizing materials and objects. Existing texture models extract essential statistics from a single texture image, and can then generate high-quality samples that are visually similar to the original by matching these statistics. However, their statistics are either hand-designed or based on a network pretrained for another purpose (e.g., object recognition). Here, we develop the first principled method for unsupervised learning of a set of statistics that are used to constrain a maximum entropy probability model. We leverage methods developed for generative diffusion models to derive training and sampling procedures, and compare these to the traditional method of sampling via matching the statistics. Despite the compactness of our trained model (512 statistics), it generates texture images whose quality is as good or better than the current state-of-the-art model (~177k statistics). A more direct comparison of the two models, obtained by synthesizing images that are indistinguishable for one model but maximally different for the other, reveals their relative strengths and weaknesses. Finally, we show that, unlike previous statistical texture models, a straight trajectory in the representation space of our model generates homogeneous texture samples that interpolate smoothly between the features of the two end points.
Abstract: High-frequency Helmholtz problems in heterogeneous media remain challenging for both classical iterative methods and end-to-end neural PDE solvers. We propose Neural Preconditioned Born Series (NPBS), a learned iterative preconditioning framework that operates in preconditioned residual coordinates induced by the Convergent Born Series (CBS). Existing learned Born-series methods primarily use Born-style unrolling for forward wavefield prediction, while learned Helmholtz preconditioners are usually formulated in physical residual coordinates. NPBS fills this gap by recasting Born-series iteration as shifted-Laplacian left preconditioning, and replacing the CBS preconditioner with a learned residual-to-correction map in the Born-preconditioned coordinates. The left preconditioner further induces a residual metric, which yields a metric-matched training objective that aligns optimization with the preconditioned geometry used at inference. On heterogeneous Helmholtz benchmarks, metric-matched NPBS reduces iteration counts by up to 1.9× over direct residual learning, with gains increasing from 1.2× to 1.9× as the wavenumber rises. Compared to classical CBS, learned NPBS reduces stationary iteration counts by over 20×; when used as a preconditioner for FGMRES, it further achieves the lowest wall-clock time among all evaluated methods. The same metric-matched formulation also improves convergence on convection--diffusion--reaction systems and Newton linear systems for nonlinear PDEs, indicating that residual-metric matching is a general design principle for neural preconditioners.
Authors:
Jianing Guo, Fangzheng Chen, Zihao Mao, WONG L Kenny, Zhenhong Wu, Yu Li, Yishuai Cai, Yuanpei Chen, Yikun Ban, Kai Chen, DOU QI, Yaodong Yang, Xianglong Liu, Huijie Zhao, Simin LiAbstract: Flow matching has emerged as a standard paradigm for robotic manipulation owing to its strong expressive power for modelling complex, multimodal action distributions, alongside similar approaches like diffusion policy. However, existing methods rely on discretized action chunks, making them brittle to demonstrations collected at heterogeneous control frequencies and prone to temporally inconsistent actions that degrade control stability. In this paper, we propose Frequency-Aware Flow Matching (FAFM), which outputs continuous, temporally consistent actions. To handle heterogeneous frequency input, we transform discrete action sequences into the frequency domain with the discrete cosine transform (DCT), perform flow matching over the resulting coefficients, and reconstruct continuous actions via cosine basis expansion. To generate temporally consistent actions, we regularize the first-order temporal derivative to promote smooth actions. This corresponds to a Sobolev-type constraint that suppresses high-frequency errors and discourages abrupt action changes. Our FAFM is simple, introduces no additional network parameters and applies to standalone flow-matching policies and vision-language action models. Across synthetic toy benchmark, obstacle avoidance, LapGym, and LIBERO, FAFM improves success rates, multimodal expressivity, motion smoothness, convergence speed, robustness to mechanical bias and mixed-frequency input. These gains are consistent when deployed on a real-world Franka robot. Code available at \urlhttps://anonymous.4open.science/r/FAFM.
PaperID: 5717, Poster
Authors: Jie Ren, Jonathan S Rosenfeld, Neil Thompson
Abstract: Parallel test-time scaling (TTS) has been shown to enhance reasoning in large language models by generating and aggregating multiple independent reasoning paths. Internal confidence scores can further improve this process by identifying paths of varying quality. However, existing methods remain computationally inefficient, as low-quality paths are fully expanded before being evaluated. In this work, we aim to improve the use of internal confidence signals for better model performance and computational efficiency. We begin by offering a new perspective on the role of confidence in aggregation, showing that it is most impactful for questions near the decision boundary. Notably, when amplified by multi-sample aggregation, even small confidence signals derived from short prefixes can meaningfully influence outcomes in these boundary cases. Building on this insight, we introduce Prefix-Guided Sampling (PreG), a method that reallocates computation by first generating short prefixes, ranking them using internal confidence scores, and then completing only the most promising candidates. This strategy reduces token usage while maintaining gains in path quality. We provide theoretical analysis demonstrating that, under a fixed compute budget, PreG strengthens the effective signal by widening the decision boundary margin. Empirically, our method consistently outperforms standard parallel TTS approaches across benchmarks, while significantly improving token efficiency and reducing overall computation.
Authors:
Baiyu Chen, Zechen Li, Wilson Wongso, Lihuan Li, Xiachong LIN, Hao Xue, Benjamin Tag, Flora SalimAbstract: As wearable and mobile devices become increasingly embedded in daily life, they offer a practical way to continuously sense human motion in the wild. But inertial signals are highly dependent on the sensing setup, including body location, mounting position, sensor orientation, device hardware, and sampling protocol. This setup dependence makes it difficult to learn motion representations that transfer across devices and datasets, and limits the broader use of wearable IMUs beyond closed-set recognition. We introduce AnyMo, a geometry-aware framework for setup-agnostic human motion modeling. AnyMo uses physics-grounded IMU simulation over dense body-surface placements to generate diverse and plausible synthetic signals, pre-trains a graph encoder from paired synthetic placement views and masked partial observations, tokenizes multi-position IMU into full-body motion tokens, and aligns these tokens with an LLM for motion-language understanding. We evaluate AnyMo on three complementary tasks: zero-shot activity recognition across 14 unseen downstream datasets, cross-modal retrieval, and wearable IMU motion captioning, where it improves average Accuracy/F1/R@2 by 11.7%/11.6%/22.6% on HAR, increases zero-shot IMU-to-text and text-to-IMU retrieval MRR by 15.9% and 28.6%, respectively, and improves zero-shot captioning BERT-F1 by 18.8%. These results support AnyMo as a generalist model for wearable motion understanding in the wild. Code is at \urlhttps://anonymous.4open.science/r/anymo.
Abstract: Autoregressive (AR) modeling has recently achieved remarkable progress in native 3D mesh generation, largely due to its natural ability to handle variable-length, discrete data structures. However, the inherent constraints of the AR paradigm severely restrict the generated meshes, leading to limited face counts, bounded vertex resolutions, and difficulties in supporting textures. To overcome these bottlenecks, we propose the Barycentric Dominance Field (BDF), a continuous representation defined on triangular mesh surfaces that elegantly encodes vertex topological connectivity. BDF bridges the fundamental gap between discrete mesh topology and continuous diffusion-based generative modeling by transforming connectivity into a continuous surface signal. As an intrinsic mesh property, BDF shares strong similarities with texture maps, enabling its seamless integration into existing 3D diffusion pipelines without requiring architectural modifications. Extensive experiments demonstrate that BDF empowers diffusion models to generate native meshes with significantly higher quality, greater scalability, and stronger robustness compared to state-of-the-art autoregressive methods.
Abstract: We present T2Mo, the first feed-forward framework for controllable dynamic 3D shape generation of general objects conditioned on 3D trajectories and text prompts. Existing text-driven methods struggle to specify precise motion due to the ambiguity of natural language. To address this, we introduce 3D trajectories as a direct spatial control signal and propose a shape-grounded trajectory embedding that maps arbitrary user-provided trajectories into geometry-aware condition tokens aligned with the input shape. Injecting these tokens into the generative backbone enables highly controllable motion generation while handling varying trajectory numbers and distributions. We conduct extensive comparisons against text-based baselines and trajectory-guided video generation workarounds. Quantitative and qualitative evaluations, along with user studies, show that our method produces motions that more faithfully follow the given prompts with higher expressiveness while preserving motion quality. Our code and weights will be publicly released.
PaperID: 5721, Poster
Authors: Yi Niu, Tianyi Xu, Mingming Ma, Xinkun Wang
Abstract: Vector Quantization (VQ) based generative image compression has achieved remarkable perceptual quality. However, existing VQ codecs suffer from two fundamental limitations. First, they lack efficient content-adaptive entropy modeling and rely on static frequencies, leading to low coding efficiency. Second, the inherent conflict between discrete indices and continuous priors prevents true end-to-end joint Rate-Distortion (RD) optimization. To resolve these issues, we propose HyperVQ, a principled framework that establishes a high-performance hyperprior entropy foundation for VQ-based codecs. The core insight of HyperVQ is to shift probability modeling entirely into the continuous embedding space. Instead of directly predicting probabilities for discrete symbols, HyperVQ predicts a high-dimensional continuous multivariate Gaussian distribution for the continuous latents. By treating the discrete codebook entries as fixed "anchors" in this space, we convert the continuous Gaussian density into categorical index probabilities based on relative distances. This elegant formulation provides a powerful, spatially-adaptive entropy engine and renders the cross-entropy rate objective fully differentiable, empowering the network to actively and dynamically optimize the RD trade-off during training. To ensure practicality, we design the lightweight H Block and the Probability Estimation Engine (PEE) to facilitate highly parallel, millisecond-level inference. Experiments demonstrate that HyperVQ acts as a universal module across diverse VQ architectures (single-scale, large-codebook, RVQ), achieving an average bitrate saving of 18.5%, which is 7.28x the saving achieved by conventional Huffman coding. This establishes a robust, RD-controllable foundation for next-generation generative image compression.
PaperID: 5722, Poster
Authors:
Matteo He, William Shen, Alexandru-Andrei Iacob, Andrej Jovanović, Xinchi Qiu, Nicholas LaneAbstract: When does a language model acquire a capability? When its hidden states encode the relevant information, or when its output readout can express it? We investigate this gap between representational availability and readout expression by tracking the unembedding matrix through pretraining. Because unembedding rows are aligned by token identity across checkpoints, they provide one trajectory per vocabulary item through the learned output interface. To measure these trajectories, we introduce parameter-trajectory crosscoding, a checkpoint-spanning sparse dictionary with shared feature identities and checkpoint-specific decoders. Applied across Pythia scales and in OLMo-2-7B, this method reveals an early, heterogeneous reorganization of the readout. This reorganization does more than alter structure. Our independent WordNet probes show that token families become increasingly separable over the same developmental window. We then causally test whether the readout component drives this development. By swapping the unembedding matrices across checkpoints, and controlled contrastive tasks, we show that task-relevant distinctions can be present in hidden states before the contemporaneous readout expresses them in logits. By ablating individual crosscoder features, we trace these delayed readout effects directly to compact, learned directions. More broadly, developmental analyses should distinguish representational availability from readout expression. Together, our findings demonstrate that the learned output interface helps determine when latent structure becomes visible in token logits, highlighting the need for developmental analyses to explicitly decouple hidden-state representation from readout expression.
PaperID: 5723, Poster
Abstract: Data collection is a critical component of modern machine learning pipelines. In many settings, multiple agents, such as hospitals or research labs, can collaborate to share the burden of data collection rather than each bearing it alone. But how should the work be divided among these agents fairly? Studying this question is complicated by heterogeneity in agents' data collection costs and data quality, and by the fact that fairness itself is subjective and use-case-dependent. We study these challenges in the context of PAC mean estimation, where agents wish to estimate an unknown scalar \mu to within accuracy \epsilon with probability at least 1-\delta, via samples from distributions with common mean \mu\,. Agents may incur different costs to collect their samples, and their distributions may have different variances (data qualities). Our first contribution is a framework that casts collaborative data collection as a bargaining problem, where agents specify a notion of utility (e.g., negative cost incurred, or cost savings relative to working alone) and a welfare function over agent utilities. By maximizing the welfare over the feasible set of data collection amounts (i.e. there is enough data to satisfy (\epsilon,\delta)-PAC mean estimation), we obtain a fair division of work. Drawing on classical fairness axioms from the microeconomics literature, the family of admissible welfare functions takes a specific one-parameter form that includes utilitarian, egalitarian, and Nash bargaining as special cases. When agents' noise variances (data qualities) are known, this yields a clean characterization of the fair division. Our second contribution addresses the more challenging setting where variances are unknown. As the feasible set of data collection amounts itself depends on the variances, it is impossible to specify a fair division of work upfront. We propose an online algorithm that combines plug-in variance estimates with carefully calibrated forced sampling, and show that it is asymptotically optimal: the ratio of the achieved welfare to the welfare of the true fair division converges to 1 almost surely as either \epsilon or \delta goes to 0. We further show that our estimator asymptotically exactly meets the specified error tolerance and failure probability.
PaperID: 5724, Poster
Authors:
Hongyi Li, Yufei Wang, Chengxuan Zhou, Qinlin Xie, Yanting Chen, Jiawei Ye, Wu JieAbstract: Large Language Models (LLMs) have made remarkable progress and are increasingly deployed as black-box API services. However, they remain vulnerable to jailbreak attacks that bypass safeguards and elicit harmful content, especially in multi-turn settings where harmful intent is gradually constructed over the interaction. Existing defenses primarily target single-turn scenarios or require access to model internals, making them insufficient against black-box multi-turn jailbreak attacks with weak single-turn discriminability, cross-turn dependency, and dynamic attack patterns. To bridge this gap, we first analyze harmful intent in multi-turn jailbreak interactions through representation-space risk signals inspired by the Linear Representation Hypothesis. Our analysis reveals that these signals are fragmented across prompts, responses, and turns, progressively evolve over the interaction, and exhibit attack-dependent temporal patterns. Motivated by this, we propose TraceGuard, an online black-box defense framework that detects multi-turn jailbreak attacks by tracking prompt and response risk signals as the interaction unfolds. Specifically, TraceGuard operates in two modes: TraceGuard-S captures cross-view risk representations and their cross-turn evolution, while TraceGuard-A further enables distribution-aware few-shot adaptation to dynamic attack patterns. Extensive experiments across multiple benchmarks and target LLMs demonstrate that TraceGuard achieves superior multi-turn jailbreak detection while preserving utility, maintaining efficiency, and generalizing to unseen attacks.
PaperID: 5725, Poster
Abstract: Vision-language perception has achieved impressive progress in aligning natural language with visual observations, yet grounding high-level semantics into part-level physical interaction remains challenging. To address this gap, we propose UniRAP, a unified model for inferring part-level physical affordances and mapping language instructions to actionable geometric representations. UniRAP formulates this problem as a conditional multimodal generation task, integrating visual inputs, textual instructions, and optional prompts into a shared spatiotemporal representation space through a unified interface token mechanism. To predict executable contact geometry, we introduce a Unified Affordance Decoder (UAD), which jointly performs object detection, part-level affordance segmentation, and 4-DoF interaction pose estimation by leveraging intermediate segmentation features. In addition, we propose a curriculum-based transfer training strategy that progressively adapts the model from general visual parsing to interaction-aware perception, improving data efficiency under limited textual annotations. Experiments show that UniRAP achieves state-of-the-art performance on referring expression segmentation, affordance grounding, and interaction pose estimation, while maintaining strong spatiotemporal consistency in dynamic video scenarios. These results demonstrate the effectiveness of UniRAP as a unified perception framework for language-guided physical manipulation. All data and code will be made publicly available.
PaperID: 5726, Poster
Abstract: Modern sequential monitoring problems often involve multiple metrics, where we monitor several data streams simultaneously and may act once the evidence is strong enough. Confidence sequences (CSs) are a natural tool for such continuous monitoring. However, for bounded vector means, existing multivariate CSs are either tight but computationally intractable, or fast to compute but conservative. To address this, we study three lifts of one-dimensional betting-based CSs to higher dimensions: a weighted Bonferroni region, an equivalent max-wealth form, and a portfolio region. The portfolio is typically much tighter, especially in higher dimensions, but its boundary and properties such as volume are not available in closed form. To make this tighter construction usable, we propose tractable outer approximations of the portfolio region that preserve statistical validity: a bounding box, an \ell_p-ellipsoid, and their intersection. We prove set relations among all constructions and show empirically that these approximations (i) achieve regions close to the intractable portfolio, (ii) substantially outperform existing tractable multivariate CSs, and (iii) enable practical use cases such as multi-metric A/B testing and model comparison.
PaperID: 5727, Poster
Abstract: Many scientific simulations are computationally expensive, limiting their use in simulation-based inference, uncertainty quantification, and decision making. We introduce macrocanonical generator networks, a framework for learning fast, data-efficient neural surrogates for amortized simulation of stationary physical processes by matching multiscale statistics of a target process. Inspired by microcanonical maximum-entropy synthesis, which matches feature statistics per sample by inner-loop optimization, our approach instead trains a neural generator to satisfy the same constraints in expectation, preserving physically realistic sample-to-sample fluctuations. We prove that gradient-descent training induces preconditioned gradient-descent dynamics in sample space that inherit the symmetry-preservation properties of the microcanonical gradient descent of Bruna and Mallat, amortizing its per-sample inner-loop optimization into a single training run after which every sample is generated in one forward pass. We further show that residual generator parameterizations initialized near the identity, trained with a small learning rate, preserve entropy locally during early optimization and mitigate mode collapse. Experiments on cosmology and fluid simulations show that macrocanonical generators recover realistic samples from extremely small training sets, in some cases a single training example, while generating new realizations more than 10^6 times faster per sample than the reference simulator and more than 10^3 times faster than microcanonical gradient descent. The trained generators can be reused for downstream pretraining, including simulation-based inference, providing a data-efficient route to multi-fidelity scientific machine learning workflows.
PaperID: 5728, Poster
Authors: Hanbin Zhou, Canzhe Zhao, Shuai Li
Abstract: We study the problem of learning in multi-player general-sum Markov games. In the simulator setting, where the learner can sample the next state conditioned on an arbitrary state action pair, the previous best-known upper bound of the sample complexity for learning an \varepsilon-approximate coarse correlated equilibrium (CCE) is \widetildeO(H^4S\sum_i=1^m A_i/\varepsilon^2) (Li et al., 2022), where H is the horizon, S is the number of states, and A_i denotes the number of actions for the i-th player. This work improves the upper bound to \widetildeO(H^4S\max_i\in[m] A_i/\varepsilon^2), matching the lower bound of \Omega(H^4S\max_i\in[m] A_i/\varepsilon^2) (Li et al., 2022). In the online setting, in which the learner can only sample a trajectory, the previous best-known sample complexity upper bound for learning CCE in multi-player general-sum Markov games is \widetildeO(H^6S\max_i\in[m] A_i/\varepsilon^2) (Song et al., 2021; Jin et al., 2024; Mao et al., 2022). In this work, we improve this to \widetildeO(H^5S\max_i\in[m] A_i/\varepsilon^2). To our knowledge, this is the tightest upper bound that breaks the curse of multi-agency. The core of our algorithmic design and analysis is the new gap decomposition and, in particular, a bandit algorithm with a high-probability empirical variance regret bound, which might be of independent interest.
PaperID: 5729, Poster
Abstract: Credit assignment remains a central challenge in cooperative multi-agent reinforcement learning (MARL), especially under partial observability, where individual policy updates may not accurately reflect each agent’s actual contribution to team outcomes. While policy-based methods such as MAPPO and IPPO provide strong optimization frameworks, their updates are typically derived from global rewards and observational value surrogates, without explicitly defining agent-specific credit. Misaligned credit signals can mislead individual policy improvement, resulting in inefficient coordination and weaker team performance. We address this challenge by formulating agent-level credit as an interventional reward response, using Proximal Causal Inference (PCI) to identify credit from observable proxies via an outcome bridge function. Building on this identification strategy, we design a practical credit-aligned update signal and integrate it into policy gradient methods. Empirical evaluations on diagnostic and benchmark tasks demonstrate that the proposed credit signal improves policy learning under partial observability, highlighting proximal identification as a promising foundation for designing credit-aware policy updates in cooperative MARL. To the best of our knowledge, this work presents the first PCI-based solution for online multi-agent cooperation.
Abstract: As embodied AI and world models increasingly operate in dynamic 3D environments, visual perception must move beyond passively interpreting given observations toward actively deciding what to observe. We study this problem through camera planning in dynamic 3D story worlds, where the camera must not only generate smooth motion, but also decide what visual evidence should be acquired before it moves. We formulate this capability as Narrative-Grounded World Visual Attention, where the camera acts as an embodied observer that determines what to observe, how to compose the observation, and how to shift attention over time under narrative intent and physical 3D constraints. To realize this capability, we propose Look-Before-Move, a camera planning framework that separates observation specification from motion execution. It first builds a Semantic Observation Contract to convert directorial intent into executable visual constraints, then performs Monte Carlo Viewpoint Search to find narrative-compliant and geometrically feasible viewpoints, and finally applies Semantic Trajectory Grounding to connect selected viewpoints into continuous, collision-aware, and temporally coherent camera motion. We further construct a dynamic 3D Story World Benchmark based on StoryBlender, covering 50 stories, 457 scenes, and 1585 shots with animated characters, semantic scene configurations, and executable 3D environments. Experiments show that our framework improves subject perception, intent consistency, and trajectory quality over representative baselines, demonstrating the importance of organizing visual attention before generating camera motion.
PaperID: 5731, Poster
Abstract: Synergistic multi-modal question generation aims to generate questions grounded in synergistic semantics distributed across multiple modalities, rather than being specified by isolated information from any single modality alone. In this setting, modality-specific cues and cross-modal shared semantics must be jointly well-organized so that the generated question reflects synergistic multi-modal contribution and captures genuinely synergistic information. However, this task faces two core challenges: , a synergistic dependency-aware framework derived from a mutual-information-based theoretical analysis, which integrates disentangled representation learning with synergy-aware optimization for synergistic multi-modal question generation. Specifically, SynMQG disentangles the multi-modal context into visual-specific, textual-specific, and shared semantics, thereby explicitly organizing heterogeneous information for synergistic information modeling. Based on these disentangled representations, it further introduces a reasoning chain as an intermediate scaffold to structure cross-modal information before generation. Moreover, SynMQG proposes GRPO with a mutual-information-based synergy reward, which explicitly measures whether the generated question preserves effective contribution of synergistic information, encouraging the model to generate questions that depend on synergistic multi-modal semantics rather than superficial single-modal cues. Experiments on ScienceQA and MultimodalQA show that SynMQG consistently outperforms representative MQG baselines across standard text-generation metrics, MLLM-based evaluation, and human evaluation, while generating questions with stronger synergistic multi-modal grounding and higher multi-modal dependency. The source code is available at https://anonymous.4open.science/r/MQG-8E57.
Abstract: Critical-data-size accounts of grokking suggest a natural post-threshold intuition: once training data is sufficient to identify the underlying rule, additional data should accelerate validation convergence. We show that this intuition can fail in a controlled structured-output task. In Needleman--Wunsch (NW) matrix generation, small Transformers reach high validation exact-match accuracy fastest at an intermediate dataset size, not at the largest one. Past this dataset-size sweet spot, generalization remains achievable but requires more gradient updates. Conversely, in the regime where partial validation competence first appears, larger datasets can require fewer updates to reach high training accuracy, suggesting that emerging rule structure can accelerate fitting beyond example-wise memorization. A multiplication baseline does not show the same post-threshold slowdown. These results separate the critical data size for the onset of generalization from the dataset size that optimizes update-based convergence, and identify structured-output tasks where learning the rule and completing exact-fitting can diverge.
PaperID: 5733, Poster
Abstract: Bayesian accounts of in-context learning treat context as exogenous evidence, but in multi-turn dialogue this assumption can fail: the assistant changes the user state that generates the next utterance, so the model updates on evidence partly produced by its own policy. We formalize this as Socially Coupled In-Context Learning (SC-ICL). The relevant stability quantity is a \emphround-trip gain (the product of belief-to-action, action-to-evidence, and evidence-to-belief derivatives), and in a two-state local reduction the standard spectral condition on the closed-loop matrix factors as a product of a user-side gain and an assistant-side gain. Risk is therefore dyadic: neither a susceptible user nor a responsive assistant becomes unstable without the other side of the loop. We connect this stability boundary to a sigmoidal transition in relational-posterior space, which explains why myopic approval-seeking can become unstable in supportive dialogue. AI--AI experiments instantiate one-step gain assays, free-running conversations, open-loop replay, and prompt-based controllers, and four convergent results support the closed-loop account: high-gain dyads cross an operational collapse criterion while damping and barrier-style controllers sharply reduce it; a trajectory-conditioned spectral estimator separates collapsing from stable runs at the theoretical stability boundary, crossing unity near operational onset; the dyadic pattern recurs across all four pairings of GPT-4o and GPT-5.4-mini as user and assistant simulators; and a persona-only ablation that uses no numeric coupling labels in the prompts reproduces the same dyadic collapse matrix. The results are not clinical evidence; they argue that alignment should evaluate bounded closed-loop inference, not only next-response quality.
PaperID: 5734, Poster
Authors: Dat Q Mac, Thanh H Dang, Duc-Trong Le, Quynh-Trang Pham Thi
Abstract: Deep neural networks suffer from catastrophic forgetting when learning on sequence tasks because standard adaptation mechanisms update a monolithic set of Cartesian weights, inherently entangling feature intensity with spatial displacement and destructively overwriting ancestral representations. Inspired by the brain’s multi-timescale consolidation and frequency-aware memory dynamics, we propose HiPo(Hierarchical Polar Adapters), a novel continual learning framework that fundamentally redefines how neural knowledge is parameterized, stored, and consolidated. First, HiPo projects token representations into a complex orthogonal Fourier basis and explicitly parametrizes the adapter weights in polar coordinates, decoupling feature intensity from geometric correlation to probabilistically isolate task updates. Second, we embed these polar weights within a multi-tiered hierarchical memory stack, functionally separating a highly plastic working memory from a deep, stabilizing long-term memory. Finally, an adaptively thresholded snapshot Fisher mechanism evaluates parameter criticality in the decoupled polar space. By executing mathematically safe, Cartesian-invariant knowledge transfers into the deep memory tiers, HiPo safely hollows out the working memory, paralyzing obsolete gradients and preserving plasticity without inducing structural distortion. Extensive experiments across standard CIL benchmarks demonstrate that HiPo establishes a highly resilient optimization manifold, achieving state-of-the-art performance and exceptional geometric stability, particularly under severe out-of-distribution shifts.
Abstract: Large language model (LLM) agents are increasingly used to operate browsers, files, code and tools, making personal assistants a natural deployment target. Yet personal agents face a privacy-cost-capability tension: cloud models execute multi-step workflows well but expose sensitive intermediate context to external APIs, while local models preserve privacy but remain less reliable. Both settings also pay repeatedly for long skill prompts and growing histories. We propose constant-context skill learning, a context-to-weights framework for recurring agent workflows: reusable procedures are learned in lightweight task-family modules, while inference conditions only on the current observation and a compact state block. A deterministic tracker renders this state block from task progress and supplies aligned subgoal rewards, so each module can be trained with step-level SFT and refined through online RL. Across ALFWorld, WebShop, and SciWorld, our agents achieve strong performance across Qwen3-4B, Qwen3-8B and Llama-3.1-8B. With Qwen3-8B, SFT+RL reaches 89.6% unseen success on ALFWorld, 76.8% success on WebShop, and 66.4% unseen success on SciWorld. They match or exceed strong published agent-training results while reducing prompt tokens per turn by 2-7× relative to controlled ReAct prompting baselines, showing that procedural context can be moved from prompts into weights.
Abstract: Hallucinations remain a major barrier to the trustworthy deployment of large language models (LLMs). We study hallucination detection from the perspective of prompt-conditioned response distributions. For a fixed prompt, an LLM induces a distribution over possible responses; when the model is uncertain or hallucinates, this distribution often becomes more complex and less internally consistent. However, this distribution is unknown, and its samples are variable-length token sequences rather than fixed-dimensional points, making direct complexity estimation difficult. To address this challenge, we propose a training-free detector based on sample-to-sample transform costs in representation space. Specifically, we represent each sampled response as an empirical distribution over generated-token hidden embeddings and compute pairwise Wasserstein distances between responses. The resulting Wasserstein distance matrix characterizes the cost structure of transforming one sampled response into another. From this matrix, we derive two complementary hallucination signals: AvgWD, which measures the average transform cost, and EigenWD, which captures the spectral complexity of the transform-cost structure. Importantly, we extend the proposed detector beyond white-box access by using an accessible auxiliary model to construct representation-space consistency signals for black-box LLMs, substantially broadening the practical applicability of our method. Experiments across five open-source LLMs and four benchmarks show that AvgWD and EigenWD consistently achieve competitive or superior AUROC compared with strong training-free uncertainty baselines. These results suggest that distributional complexity in token-level representation space provides an effective signal for hallucination detection.
Abstract: CLIP has become a cornerstone of multimodal representation learning, yet improving its performance typically requires a prohibitively costly process of training from scratch on billions of samples. We ask a different question: Can we improve the performance of open-weight CLIP models across various downstream tasks using only existing self-supervised datasets? Unlike supervised fine-tuning, which adapts a pretrained model to a single downstream task, our setting seeks to improve general performance across various tasks. However, as both our experiments and prior studies reveal, simply applying standard training protocols starting from an open-weight CLIP model often fails, leading to performance degradation. In this paper, we introduce , a self-supervised fine-tuning framework that overcomes the performance degradation. TuneCLIP has two key components: (1) a warm-up stage of recovering optimization statistics to reduce cold-start bias, inspired by theoretical analysis, and (2) a fine-tuning stage of optimizing a new contrastive loss to mitigate the penalization on false negative pairs. Our extensive experiments show that TuneCLIP consistently improves performance across model architectures and scales. Notably, it elevates leading open-weight models like SigLIP (ViT-B/16), achieving gains of up to +2.5% on ImageNet and related out-of-distribution benchmarks, and +1.2% on the highly competitive DataComp benchmark, setting a new strong baseline for efficient post-pretraining adaptation.
Abstract: Image-to-video models often generate videos that remain overly static, compared to text-to-video models. While prior approaches mitigate this issue by weakening or modifying the image-conditioning signal, they often require additional training or sacrifice fidelity to the reference image. In this work, we identify reference-frame dominance as a key mechanism behind motion suppression. We observe that non-reference frames in I2V models allocate excessive self-attention to reference-frame key tokens, causing reference information to be over-propagated across time and suppressing inter-frame dynamics. Based on this finding, we propose DyMoS (Dynamic Motion Slider), a training-free and model-agnostic method that rebalances the attention pathway from generated frames to the reference frame during initial denoising steps. DyMoS leaves both the input image and model weights unchanged and introduces a single scalar parameter for continuous control over motion strength. Experiments across multiple state-of-the-art I2V backbones demonstrate that DyMoS consistently improves motion dynamics while maintaining visual quality and fidelity to the reference image.
Authors: Junxuan Li, Arko Mukherjee, Soumyabrata Pal
Abstract: Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this overlap sparsity is the first-order determinant of wrong deployment decisions: at 5% pairwise overlap, wrong-decision rates reach 25% and the probability of selecting the wrong best judge among ten candidates is 65%. The two actionable levers are overlap quantity and allocation. For quantity, we derive a minimum-overlap formula showing \rho \geq 0.25 suffices for non-borderline judges while borderline cases remain fundamentally hard. For allocation, a zero-cost stratified scheme halves false-rejection rates relative to random sampling when strata are informative. We validate on 10 LLM judges across four evaluation matrices spanning visual assessment, causal reasoning, and summarization.
PaperID: 5740, Poster
Abstract: Vision Transformers are increasingly applied to data defined on the sphere, particularly in physics, meteorology, and related scientific domains. Position embeddings provide the spatial information that attention itself does not encode, making their design central to the expressive power of Transformers. However, most existing position embedding schemes are designed for Cartesian grids and therefore do not naturally handle longitude periodicity, polar singularities, or geodesic relations on spherical domains. This mismatch limits the ability of standard attention to model spherical geometry without specialized architectural modifications. We propose Spherical Reflection Position Embedding (SpRePE), a drop-in position embedding scheme for Vision Transformers on spherical data. SpRePE encodes each absolute spherical position by applying Householder reflections to query and key representations. Although each token is encoded only from its own absolute coordinates, the resulting attention inner products induce an explicit spherical relative-position term with a clear geometric interpretation. This formulation injects sphere-aware relative geometric information into standard attention without constructing pairwise attention-bias matrices, introducing task-specific modules, or modifying the backbone architecture. It also avoids the quadratic overhead of relative position biases and preserves the same asymptotic overhead as RoPE. We evaluate SpRePE on spherical image classification, panoramic depth estimation, and global weather forecasting. Across these tasks, SpRePE achieves competitive performance compared with strong position embedding baselines, with particularly clear gains in settings where spherical geometry is important.
Abstract: Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws , each modality reshaping the other. In this paper, we bring this coupled loop to artificial systems. Masked Diffusion Models (MDMs) are ideally suited to this task, yet existing samplers either decode text and image sequentially or update them in parallel branches that share only previous-step history, but not the other modality's latest decisions the same step; combined with MDMs' inability to remask, cross-modal contradictions are neither detected nor repaired. We introduce , a framework in which one modality's transition rates are functionals of the other modality's unmasking confidence score, as weighted by cross-modal attention. Furthermore, a remasking jump retracts commitments the moment cross-modal evidence turns against them. In conjunction with SC-CMJP, we introduce upled + Jump), a novel training-free single-pass sampler for joint multimodal generation. For training and evaluation purposes, we have created and will release three large-scale joint multimodal generation corpora: JEdit-1M, JMaze-200K, JNono-200K, with matching in- and out-of-distribution benchmarks. achieves best joint performance for image understanding and editing as well as visual reasoning (maze and nonogram solving). The performance of the sampler scales monotonically with the number of denoising steps, evidence that the benefits of cross-modal coupling
PaperID: 5742, Poster
Abstract: Asynchronous pipeline parallelism can accelerate distributed model training, but suffers from instability caused by inconsistent gradient computations. We identify a fundamental source of this inconsistency: error signals are coupled to forward-time features, not just parameter states. This means that parameters may evolve as long as they preserve forward-equivalence. Leveraging this insight, we introduce FRESCO, which frames asynchronous updates as a constrained optimization problem: it finds the minimal update modification that acts trivially on the activation subspace. Thus, it serves as a principled generalization beyond the usual synchronous--asynchronous dichotomy: subspace-level protection preserves gradient validity without synchronization stalls, unconstrained-asynchrony instability, or memory overhead that works against pipeline scaling. We demonstrate that FRESCO consistently outperforms representative asynchronous baselines in LLM pre-training, with the gains most pronounced under deep pipelines and high inconsistency levels. Moving beyond the current limitations, FRESCO establishes a new foundation for high-utilization pipeline parallel training, providing a scalable path for large-model development on diverse, resource-constrained systems.
PaperID: 5743, Poster
Abstract: Long-video question answering is challenging since the answer often hinges on a few decisive moments scattered throughout the video, while memory constraints force Multimodal Large Language Models (MLLMs) to sample frames at extremely low rates, resulting in extreme sparsity that risks omitting key evidence. Though aggressively enlarging the context size is intuitive, we argue that performance remains fundamentally constrained by the bias, i.e., the misleading cues caused by spurious linguistic correlations and salient yet irrelevant visual observations. In this paper, we first introduce S-MME, a benchmark designed to systematically diagnose shortcut-biased behavior in MLLMs. To mitigate this failure, we further propose a training-free framework, Counterfactual Long-video Evidence-Aware Reasoning (CLEAR). By treating the full video-question pair as a complete view, CLEAR estimates bias effects via other views by only keeping the question or locally salient visual clips. Then, we apply hidden-state intervention to mitigate the discrepancy at inference such that the correct answer can be predicted confidently. With experimental studies across multiple benchmarks, we show that our method is generalizable and can be integrated in different MLLMs, yielding consistent gains in prediction and robustness. We hope CLEAR can offer a complementary direction in test-time scaling for long-video understanding and will publish code.
Abstract: Although Large Multimodal Models (LMMs) have achieved strong performance on general video understanding, their susceptibility to textual prior shortcuts during causal discovery has been recognized as a critical deficit. The underlying mechanisms of this phenomenon remain incompletely understood, as existing benchmarks only measure response accuracy without revealing the sources and extent of the deficit. We introduce ProCauEval, a perturbation-based evaluation protocol that shifts from outcome assessment to mechanism diagnosis, probing causal discovery through five controlled configurations that systematically manipulate visual and textual modalities to decompose their respective contributions to model behavior and dissect the failure modes. Evaluating 17 mainstream LMMs, we find that models faithfully perceive video content yet systematically underexploit it during causal reasoning. We further observe that stronger post-training amplifies rather than mitigates textual prior reliance, and that higher baseline performance correlates with greater fragility under perturbation. To address these, we propose Anti-Distillation Policy Optimization (ADPO), a reinforcement learning framework built on negative teacher alignment, which augments GRPO by explicitly pushing the policy away from a prior-only counterfactual teacher induced by visual corruption. Specifically, ADPO maximizes the divergence between the policy distributions conditioned on the original and visually corrupted inputs, thereby forcing the model to ground its reasoning in visual evidence rather than textual shortcuts. Extensive experiments show that ADPO improves visual engagement without sacrificing fundamental comprehension, thus offering a preliminary step toward reliable causal discovery.
PaperID: 5745, Poster
Abstract: Video generation models are becoming a scalable form of world models, but they mainly generate plausible motion rather than proactively control or optimize the underlying dynamics. As a result, an object in the generated video may follow trajectories that are unsafe, not smooth, inefficient, or physically inconsistent. In this work, we propose OptiWorld, a framework that brings classical optimal control into video generation at inference time. OptiWorld first extracts a compact, task-relevant world state, then plans an optimal trajectory under physical constraints, and finally renders the video conditioned on this trajectory. We formulate planning as a geometric problem on a continuous manifold, which converts 3D geometry and task-dependent physical constraints into a unified planning geometry. By adding this optimal-control layer, OptiWorld generates videos with preferable dynamics, demonstrating strong potential in multiple tasks including goal-conditioned image-to-video generation, video dynamics editing, and counterfactual generation.
Abstract: Understanding model predictions is essential for physical applications, where outputs often inform safety-critical decisions, such as structural load assessment, weather warnings, and clinical diagnosis. Shapley values satisfy many desirable properties as an attribution method, but their computational cost during inference hinders their practical use. Current amortized explainers, such as FastSHAP, are limited to homogeneous inputs, which is problematic for physical applications where data often comes from irregular grids and geometries. We introduce OperatorSHAP, a grid-agnostic attribution method and training procedure that allows us to train FastSHAP-like explainers for neural operators. We establish a theoretical framework for attributions in function space, connecting to Aumann–Shapley values. We further show that OperatorSHAP's explanations are consistent with state-of-the-art discrete Shapley values across resolutions and transfer across grid sizes without retraining.
Authors:
Sixing Chen, Ji-An Li, Saner Cakir, Sinan Akcali, Kayla Lee, Marcelo G MattarAbstract: Large language models (LLMs), especially reasoning models, generate extended chain-of-thought (CoT) reasoning that often contains explicit deliberation over future outcomes. Yet whether this deliberation constitutes genuine planning, how it is structured, and what aspects of it drive performance remain poorly understood. In this work, we introduce a new method to characterize LLM planning by extracting and quantifying search trees from reasoning traces during four-in-a-row gameplay. By fitting computational cognitive models on the extracted search trees, we characterize how plans are structured and how they influence move decisions. We find that LLMs' search is shallower than humans', and that performance is predicted by search breadth rather than depth. Most strikingly, although LLMs expand deep nodes in their traces, their move choices are best explained by a myopic model that ignores those nodes entirely. A causal intervention study in which we selectively prune CoT paragraphs further suggests that move selection is driven predominantly by shallow rather than deep search. These patterns contrast with human planning, where performance is driven primarily by deep search. Together, our findings reveal a key difference between LLM and human planning: while human expertise is driven by deeper search, LLMs do not act on deep lookahead. This dissociation offers targeted guidance for aligning LLM and human planning. More broadly, our framework provides a generalizable approach for interpreting the structure of LLM planning across strategic domains.
Abstract: We introduce spatially grounded contextual image generation, a new controllable image generation task that reframes the conditioning paradigm. Instead of supplying a reference image and a global text prompt through two separate encoders (vision and language), UniVL is trained to bind semantics to spatial locations directly from a single unified visual input, in which the textual instruction is rendered onto the spatial mask, removing the need for a standalone text encoder at inference. This enables contextual image generation, which follows user’s specified what should appear where instructions, as well as waiving the need of text encoder to save computation significantly. For the task, we propose a framework in which the UniVL encoder—adapted from an optical-character-recognition-pretrained backbone—reads the unified condition optically, producing a UniVL embedding fVIL that fuses visual and semantic intents to spatial locations, packed as a single token sequence. A two-stage pipeline aligns UniVL in VAE embedding space and then conditions a pretrained diffusion backbone entirely on UniVL embeddings, eliminating the standalone text encoder (e.g., T5). The reframing is deliberately minimalist for text, but the empirical payoff is large. On UniVL-ImgGen, a benchmark of 477K mask-annotated images that we construct to support training and evaluation, UniVL achieves superior image quality over text-prompted baselines (FID: 14 → 11, PSNR: 16 → 20) while eliminating the text encoder entirely, reducing inference TFLOPs by up to 52% and runtime by up to 44%. Additional ablation studies verify components of different parts of the proposed method, paving way for efficient and spatially grounded image generation with unified conditioning paradigm.
PaperID: 5749, Poster
Abstract: Multimodal Large Language Models (MLLMs) trained on sequential tasks suffer from catastrophic forgetting. We study this problem in LoRA-based multimodal continual instruction tuning. By analyzing the diagonal Fisher information of LoRA parameters, we find that forgetting is amplified not merely by shared high-sensitivity parameters across tasks, but by the concentration of updates on those shared parameters, \ie, some tasks distribute their change uniformly across parameters, while others pack most of it into a narrow subset, causing disproportionate interference on the parameters that prior tasks depend on. We further observe that this concentration manifests in the dominant directions of recent gradients. Based on this insight, we propose Dominant-direction Gradient Projection (DiGPro), which maintains a short buffer of recent gradient directions, extracts their dominant subspace with adaptive rank selection, and attenuates the parallel component for each optimizer step. Crucially, DiGPro requires no prior-task data or stored subspaces, operating on the current gradient history. Experiments on cross-modal and vision-centric continual learning benchmarks show DiGPro consistently reduces forgetting and attains competitive final results. The code is available in the supplementary materials.
Abstract: Diffusion models have become the dominant paradigm in text-to-image generation, and test-time scaling (TTS) improves sample quality by allocating additional computation at inference. Existing TTS methods, however, resample the entire image, while generation quality is often spatially heterogeneous. This leads to unnecessary computation on regions that are already correct, and localized defects remain insufficiently corrected. In this paper, we explore a new direction -- Localized TTS -- that adaptively resamples defective regions while preserving high-quality regions, thereby substantially reducing the search space. This raises two challenges: accurately localizing defects and maintaining global consistency. We propose LoTTS, the first fully training-free framework for localized TTS. For defect localization, LoTTS contrasts cross-/self-attention signals under quality-aware prompts (e.g., ''high-quality'' vs. ''low-quality'') to identify defective regions, and then refines them into coherent masks. For consistency, LoTTS perturbs only defective regions and denoises them locally, ensuring that corrections remain confined while the rest of the image remains undisturbed. Extensive experiments on SD2.1, SDXL, and FLUX demonstrate that LoTTS achieves state-of-the-art performance:~it consistently improves both local quality and global fidelity, while reducing GPU cost by 2-4× compared to Best-of-N sampling. These findings establish localized TTS as a promising new direction for scaling diffusion models at inference time.
PaperID: 5751, Poster
Abstract: Few-step distillation is essential for deploying modern video diffusion models, and Distribution Matching Distillation (DMD) has emerged as a leading paradigm due to its strong output-space fidelity and its flexibility in supporting both bidirectional and causal student architectures. We find, however, that DMD has a structural blind spot: its objective lives entirely in the output space and provides no signal pushing the student to form a structured internal representation. As the step count shrinks, this output-only supervision becomes too coarse, and the distilled student degrades on both fine spatial structure and coherent temporal dynamics. To close this gap, we introduce Representation-Forcing, which adds a predictive representation loss on top of DMD without changing its output objective. By feeding the student and an EMA teacher with heterogeneous noise levels, we create an information asymmetry that forces the student to predict the teacher's representation from a more corrupted view --- explicitly bringing "representation compression" to distribution matching. Experiments show that Representation-Forcing consistently improves both spatial fidelity and temporal coherence. The mechanism is paradigm-agnostic across bidirectional and causal student architectures, requires no external feature extractor, and adds negligible cost.
PaperID: 5752, Poster
Abstract: Real-world generative problems in biology, medical imaging, or robotics, rarely come with perfectly paired source and target distributions. While unbalanced Optimal Transport (uOT) provides a robust framework to bridge unpaired datasets via minimum-length paths in probability space, existing formulations based on the Wasserstein–Fisher–Rao (WFR) geometry introduce a structural bias toward high density regions, often triggering mode collapse. We introduce Contact unbalanced Optimal Transport (CuOT), a novel uOT formulation grounded in contact geometry that decouples density transport from mass growth to eliminate structural bias. This allows CuOT to adapt sampling steps to the local data structure, naturally reducing step sizes in high-density regions to prevent overshooting and ensuring stable convergence. The result is a stable, scalable path-length minimization solver that consistently outperforms WFR-based methods on biological dynamics reconstruction, unpaired image-to-image translation, and video generation.
Abstract: The recent advancement of Large Language Models (LLMs) has established their potential as autonomous interactive agents. However, they often struggle in strategic games of incomplete information, such as bilateral price negotiation. In this paper, we investigate if Reinforcement Learning from Verifiable Rewards (RLVR) can effectively teach LLMs to negotiate. Specifically, we explore the strategic behaviors that emerge during the learning process. We introduce a framework that trains a mid-sized buyer agent against a regulated LLM seller across a wide distribution of real-world products. By grounding reward signals directly in the maximization of economic surplus and strict adherence to private budget constraints, we reveal a novel four-phase strategic evolution. The agent progresses from naive bargaining to using aggressive starting prices, moves through a phase of deadlock, and ultimately develops sophisticated persuasive skills. Our results demonstrate that this verifiable training allows a 30B agent to significantly outperform frontier models over ten times its size in extracting surplus. Furthermore, the trained agent generalizes robustly to stronger counterparties unseen during training and remains effective even when facing hostile, adversarial seller personas.
PaperID: 5754, Poster
Authors:
Xinyu Hou, Yang Lu, Rabimba Karanjai, Pei-Chi Pan, Sen Lin, Lei Xu, Weidong ShiAbstract: Routing systems for large language models, such as MasRouter, RouteLLM, FrugalGPT, and others, match queries to suppliers based on cost-quality tradeoffs. Most prior work optimizes this problem from the demand side, while the supplier-side question of how LLM suppliers should price their services within a routing mechanism has received little formal treatment. We provide the first systematic analysis of this problem. In our setup, the LLM is the priced commodity: supplier organizations set price functions for the services they offer, rather than acting as strategic agents that generate bids per query. We model supplier-side pricing as a sealed-bid first-price reverse auction over a router that allocates each query to one supplier based on cost and quality. This framework applies to any cost-quality routing system rather than to a specific implementation. Using calibrated profiles from 12 open-weight models and MasRouter as a case-study router, we characterize equilibrium behavior. Our main finding is that the Bayesian Nash equilibrium is asymmetric: capability-differentiated suppliers adopt flat, non-discounted strategies, while marginal-quality suppliers compete primarily on price. A mechanism-baseline experiment, where the case-study router is replaced by an analytical rational-decision rule, confirms that this pattern is a property of the auction mechanism rather than an artifact of router training. We further observe that price differentiation can emerge in equilibrium even under capability symmetry, because private-cost types alone produce nontrivial bid functions. The practical implication is that auction-based pricing alone does not discipline specialist rents. Achieving that goal requires additional mechanism elements, such as reserve prices or capability-blind tie-breaking.
PaperID: 5755, Poster
Abstract: In semantic segmentation, a recent line of RankSEG methods directly optimizes Dice/IoU scores at inference time, improving alignment with evaluation metrics without modifying model training. Despite its theoretical and empirical success, RankSEG relies on the restrictive Conditional Independence Assumption (CIA), which ignores crucial label correlations and therefore degrades performance in ambiguous or low-contrast scenarios. However, accounting for full label dependence is computationally prohibitive, requiring \mathcalO(d^3) time. To address this, we replace the CIA with a Spatially Localized Dependence (SLD) structure that captures local label correlations while keeping the dependence model tractable. We further overcome the remaining computational bottleneck via a Reciprocal Moment Approximation coupled with a novel fixed-point optimization strategy that eliminates exhaustive search. The proposed algorithm achieves a highly practical \mathcalO(d \log d) complexity and consistently outperforms conventional argmax and CIA-based RankSEG across diverse segmentation benchmarks. Improvements are significant in low-contrast or small-object scenarios, where label dependence offers valuable signals complementary to image information for accurate segmentation.
PaperID: 5756, Poster
Abstract: Vision-Language-Action (VLA) models promise a single language-conditioned policy adaptable to many manipulation behaviors, yet deployment additionally requires skill retention and long-horizon composition: a policy must retain previously acquired skills and reuse them inside composite tasks whose behavior changes across phases. The standard recipe absorbs every new task into one shared update, which provides neither a parameter-level locus for reusable skills nor a mechanism for invoking each skill along the trajectory. We observe that this missing structure is latent rather than absent: per-skill low-rank updates of a pretrained VLA occupy a robustly distinguishable, low-dimensional subspace, and a single composite-task LoRA concentrates roughly half of its Frobenius energy in this subspace, against a near-zero matched random-rotation control. Building on this geometry, we introduce PRISM, which represents reusable skills as parameter-level primitives: each skill family is stored as a dedicated low-rank action-expert adapter in a primitive bank, primitives are composed at inference by a per-chunk VLM-conditioned router that reads the prefix representation to track task progress, and the router is supervised by per-chunk labels from a deterministic, action-derived oracle. During composite training, only the vision encoder, the language LoRA, and the router are updated; every primitive is held frozen, so the per-skill action-expert parameters are preserved by construction. On RoboCasa365, PRISM preserves atomic-skill success through composite training and reaches 20% on long-horizon composite tasks, while a standard \pi_0.5-LoRA two-stage fine-tune instead loses roughly 25 atomic points and reaches only 11% composite. We additionally validate PRISM on an AgileX PIPER real-robot setup.
Abstract: Image restoration faces a fundamental tradeoff: methods that minimize error produce blurry reconstructions, while those that maximize perceptual quality yield sharp but less faithful images. Existing approaches either commit to a single operating point on this distortion–perception (DP) frontier or require retraining, auxiliary models, or changes in discretization to access different points. We show that flow map models, a recent extension of flow matching for few-step sampling that learns an average field, implicitly define a one-parameter family of denoisers that continuously spans the DP frontier. The parameter, the lookahead t, acts as a knob: varying it traces a smooth path from the MMSE estimator to a perceptually aligned one. For Gaussian targets, we prove that this path recovers the optimal DP frontier exactly; for natural images, we demonstrate empirically that it closely follows the same behavior. Embedded within a Plug-and-Play solver, a single trained flow map matches or exceeds specialized baselines at both ends of the DP spectrum and uniquely traces a continuous curve in between without retraining, paired data, or auxiliary networks. Extensive experiments on CelebA (128× 128) and AFHQ (256× 256) across several linear and nonlinear inverse tasks validate our findings.
Abstract: Solving optimal transport (OT) on random minibatches is a common surrogate for exact OT in large-scale learning. In flow matching (FM), this surrogate is used to obtain OT-like couplings that can straighten probability paths and reduce numerical integration cost. Yet, the population-level coupling induced by repeated minibatch OT remains only partially understood. We formalize this coupling as the expected batch OT plan \overline\pi\_k, obtained by averaging empirical OT plans over independent minibatches of size k. We then establish its large-batch consistency and, in the semidiscrete case relevant to generative modeling, derive rates for both the transport-cost bias and the convergence of \overline\pi\_k to the OT plan. For FM, this yields a population coupling whose induced velocity field is regular enough to define a unique flow from the source to the discrete target. We finally quantify how OT batch size interacts with numerical integration in a tractable two-atom model and in synthetic and image experiments.
PaperID: 5759, Poster
Abstract: Foundation-model-based predictors of perturbation response are evaluated primarily by aggregate accuracy, yet lack principled per-prediction reliability criteria. We audit the state-transition (ST) decoder~\citepadduri2025predicting shared by recent predictors across three encoders and four cell lines under CRISPRi perturbation, and identify a rank-1 attention bottleneck whose leading direction maps in gene space to the same cycling/p53 stress-response program in every cell line tested. The bottleneck implies a distance-defined reliability boundary: per-prediction quality degrades smoothly with the gene-space distance from a model's predicted perturbation response (\Delta, the predicted change relative to control) to its nearest training \Delta, and biological-family identity, not training-pool size, drives held-out quality. We propose \emphpredicted-NN-distance, this same distance computed from model output alone, as a per-prediction trust signal that tracks per-perturbation quality at Spearman |\rho| \approx 0.81 (oracle 0.88) and admits a single pooled threshold that transfers across cells and encoders under leave-one-out evaluation. The signal outperforms a simpler predicted-magnitude baseline in most cases. Per-prediction reliability is bounded by a low-rank decoder geometry and by gene-space coverage of training perturbations, not by aggregate benchmark accuracy.
Abstract: Time series forecasting often suffers from over-smoothing, especially when future dynamics are multi-modal. Forecasts may follow the coarse trend of the observed future, but fail to preserve sharp changes, oscillations, turning points, and regime transitions that define plausible dynamic evolution. In this work, we revisit over-smoothing from the perspective of latent dynamical mode compression: under partial observation and single-realization supervision, multiple plausible future modes can be weakened, merged, or averaged during forecasting. Based on this view, we propose Dirichlet-Guided Group Forecasting (DGF), a mode-preserving forecasting framework that explicitly models multiple mode-conditioned predictive distributions and uncertainty over their selection probabilities. DGF uses a Dirichlet-guided hierarchical sampling mechanism and reward-based optimization to encourage forecasts that are accurate, dynamically consistent, and mode-distinct. Extensive experiments on real-world forecasting benchmarks show that DGF reduces over-smoothing while improving forecasting accuracy, diversity, and dynamical consistency.
Abstract: A hallmark of neurodegenerative diseases such as Alzheimer's and Parkinson's is the aberrant aggregation of proteins into amyloid fibrils, and small molecules that selectively bind to these fibrils hold promise as diagnostics, imaging probes, and therapeutics. Predicting how such ligands bind to fibril targets, however, presents two fundamental challenges. First, resolved co-crystal structures of amyloid–ligand complexes are exceptionally scarce; even with recent advances in cryo-EM only a handful have been structurally characterized, making supervised training of docking models impractical for this target class. Second, amyloid fibrils present a binding mode fundamentally different from globular proteins: ligands intercalate into longitudinal cross-β grooves and stack cooperatively along the fibril axis, a geometry that existing docking models are not designed to capture. To address these challenges, we present CORAL (Cooperative Amyloid Ligand docking), a reinforcement learning framework that trains a generative docking model to produce ligand pose distributions tailored to the cross-β groove geometry. Our reward explicitly incorporates cooperative ligand–ligand stacking energy alongside protein–ligand docking affinity, directly capturing the distinctive binding geometry of amyloid fibrils. We further introduce a curated evaluation set of amyloid–ligand complexes constructed from model-generated poses validated by domain experts. Experiments on both experimentally resolved structures and this evaluation set demonstrate improved pose quality and binding affinity correlation over existing docking baselines.
Abstract: Mapping reaction pathways and transition states (TS) is fundamental to chemistry but computationally expensive at scale. The minimum energy pathway (MEP) dictates reaction rates and mechanisms, yet recovering it via electronic-structure methods requires thousands of costly force evaluations. Recent generative models accelerate TS identification but require slow iterative inference and only predict isolated saddle-point snapshots, missing the continuous reaction trajectory. We introduce Drift-React, an \mathrmSE(3)-equivariant generative framework that predicts complete reaction pathways in a single forward pass from only reactant and product geometries. By shifting distribution evolution to training via a Sinkhorn-weighted drifting field, Drift-React eliminates both the iterative force evaluations of NEB-style methods and the sequential ODE/SDE integration of diffusion and flow matching models. Evaluated on the Transition1x and Halo8 datasets, our one-step model generates physically consistent MEPs that accurately capture energetic bottlenecks and enable arbitrary-resolution sampling along the reaction coordinate. For isolated TS prediction, Drift-React matches the sub-Ångström accuracy of state-of-the-art iterative models while delivering orders-of-magnitude acceleration, clearing a major computational bottleneck for large-scale reaction network exploration.
Authors:
Yuchen Cai, Ding Cao, Liang Lin, Chunxi Luo, Xin Xu, Kai Yang, Weijie Liu, Saiyong Yang, Tianxiang Zhao, Guangzhong Sun, Guiquan Liu, Junfeng FangAbstract: On-policy distillation (OPD) has emerged as an efficient post-training paradigm for large language models. However, existing studies largely attribute this advantage to denser and more stable supervision, while the parameter-level mechanisms underlying OPD's efficiency remain insufficiently understood. In this work, we argue that OPD's efficiency stems from a form of ``foresight'': it establishes a direct and stable path from the initial model to the final model early in training. This foresight manifests in two aspects. First, at the Module-Allocation Level, OPD identifies regions with low marginal utility and concentrates updates on modules that are more critical to reasoning. Second, at the Update-Direction Level, OPD exhibits stronger low-rank concentration, with its dominant subspace aligning closely with the final update subspace early in training. Motivated by this theory, we propose EffOPD, a plug-and-play acceleration method that speeds up OPD by searching for an effective step size and extrapolating along the current update direction. EffOPD requires no additional trainable modules or complex hyperparameter tuning, and achieves up to 3.5× training acceleration while maintaining comparable final performance. Overall, our findings provide a parameter-dynamics perspective for understanding the efficiency of OPD and offer practical insights for designing more efficient post-training methods for large language models. Our code is available at: https://anonymous.4open.science/r/EffOPD-7C58.
PaperID: 5764, Poster
Abstract: Online reinforcement learning (RL) has emerged as a powerful paradigm for aligning text-to-image flow matching models with complex user intent. However, existing methods typically rely on sparse rewards computed only from final decoded images. This delays credit assignment across the denoising trajectory and makes training computationally expensive. Moreover, simply combining multiple rewards does not reliably produce a Pareto trade-off, where gains in compositional accuracy often degrade aesthetic quality. In this work, we propose DeFlowCritic, an efficient online RL framework based on Dense Latent Reward Alignment. Unlike prior work, our method provides dense supervision on intermediate latents, enabling early-stage guidance for the denoising trajectory. To achieve this, we introduce an internal critic that evaluates noisy latents using features extracted directly from the diffusion model, a representation inherently better suited for latent-space reward estimation than standard image encoders. We construct a high-quality reward-modeling dataset with continuous scores for critic training, jointly capturing compositional correctness and aesthetic preference. The resulting dense reward signals enable efficient online RL training and inference scaling. Experiments show that our method achieves a superior balance between prompt faithfulness and visual quality on GenEval2 and T2I-CompBench benchmarks.
PaperID: 5765, Poster
Abstract: Predicting the olfactory perception of scent mixtures remains a fundamental challenge in computational neuroscience. Existing methods derive mixture-level features by encoding individual components with a single-molecule feature encoder pretrained on annotated olfactory data and aggregating these features via mean pooling, concatenation, or self-attention. This paradigm faces two critical limitations: the scarcity of annotated olfactory data and the inability to capture complex interactions among mixture components. We present MixScentNet, a multiscale graph neural network framework, to address these challenges with two key innovations. First, we introduce a self-supervised pretraining strategy that uses molecular graph structures with RDKit features to predict corresponding Mordred descriptors, enabling the model to learn rich physicochemical knowledge without relying on scarce annotated olfactory data. Second, we propose a novel mixture-as-graph paradigm to model the constituent molecules as nodes in the mixture-level graph, aligning with the fact that humans holistically perceive scent mixtures. We further process the graph via a graph attention network V2 (GATv2) to capture the high-order molecular interactions. MixScentNet differs from existing methods in terms of its pretraining strategy and ability to capture mixture-level features, achieving state-of-the-art performance on both mixture-level olfactory label prediction and perceptual distance estimation tasks. We also find that MixScentNet can reproduce olfactory white phenomena, indicating that the model enjoys clear interpretability grounded in psychophysics. The demo code is available in the Anonymous link \urlhttps://anonymous.4open.science/r/Odor-D2DB
Abstract: Parametric human models capture global pose but overlook fine-grained surface dynamics. Generic scene flow estimates dense motion but struggles with articulated humans and lacks 4D ground truth. To bridge this gap, we introduce H-Flow, a dense 4D motion representation that captures non-rigid deformations beyond skeletal kinematics. We estimate H-Flow from monocular video via a unified multi-head architecture that jointly perceives pose, depth, and flow. To overcome the scarcity of dense motion annotations, we propose a self-supervised learning paradigm that embeds geometric, structural, and biomechanical priors into optimization. These cross-modal constraints tightly couple pose, depth, and flow, so that improving any one modality simultaneously drives the others toward consistency. For evaluation, we present DynAct-4D, a high-fidelity synthetic benchmark providing dense 4D ground truth for complex human movements. Extensive experiments show that our method outperforms state-of-the-art scene flow baselines. Furthermore, H-Flow serves as an effective motion primitive, bringing substantial improvements to downstream tasks including action recognition and video generation. Models, code, and resources will be released upon publication.
PaperID: 5767, Poster
Abstract: We prove that in any Dec-POMDP, sufficiently high entropy regularization ensures that the policy gradient flow with tabular softmax parametrization always converges, for any initialization, to the same joint policy, and that this joint policy is equivariant w.r.t. all symmetries of the Dec-POMDP. In particular, policies coming from different initializations will be fully compatible, in that their cross-play returns are equal to their self-play returns. Through extensive evaluation of independent PPO, arguably the standard baseline deep multi-agent policy gradient algorithm, in the Hanabi, Overcooked and Yokai environments, we find that the entropy coefficient has a massive influence on the cross-play returns between independently trained policies, and that the decrease in self-play returns coming from increased entropy regularization can often be counteracted by greedifying the learned policies after training. In Hanabi in particular we achieve a new SOTA in inter-seed cross-play this way. While we give examples of Dec-POMDPs in which one cannot learn the optimal symmetry equivariant policy this way, both our theoretical and empirical results suggest that one should consider far higher entropy coefficients during hyperparameter sweeps in Dec-POMDPs than is typically done.
PaperID: 5768, Poster
Abstract: Modern visual recognition systems increasingly rely on unlabeled images collected from diverse visual domains. A common goal in dataset construction is to avoid obvious marginal imbalance, such as skewed class totals or uneven domain sizes. However, marginal balance can be misleading in multi-domain data: even when per-class totals and per-domain sample sizes are balanced, sample allocations can be highly skewed across domain-class pairs. We call this hidden conditional skew Domain-Conditioned Class Imbalance (DCI), where class imbalance is revealed only after conditioning on domain context. For controlled evaluation, we fix class and domain marginals and vary only the joint domain--class allocation, making the effect of domain-conditioned imbalance directly measurable. To address this challenge, we propose HyperIDAC, a hypernetwork-based classifier for multi-domain imbalance. HyperIDAC infers domain priors from visual features and lightweight zero-shot predictions, then dynamically generates instance-specific classifier parameters and domain-conditioned class embeddings. This design yields domain-sensitive decision boundaries while leveraging CLIP semantics to support underrepresented domain--class pairs. Furthermore, a multi-level pseudo-labeling module combines teacher–student training with historical label tracking and a fallback for low-confidence samples, mitigating bias reinforcement in unlabeled settings. Extensive experiments on four multi-domain benchmarks under three regimes show that HyperIDAC substantially outperforms state-of-the-art methods, highlighting the value of domain-aware hypernetwork-based adaptation under DCI.
PaperID: 5769, Poster
Abstract: Large distribution shifts remain a primary obstacle for domain generalization and unsupervised adaptation. Although gradual domain adaptation alleviates this difficulty by introducing ordered intermediate domains, most existing approaches still operate in a discrete, step-wise manner, which can accumulate errors and often relies on unstable Min--Max optimization. We propose FlowGDA, a flow-matching framework that models gradual adaptation as \emphcontinuous transport on a statistical manifold in a representation space. FlowGDA learns time-dependent velocity fields that induce smooth domain trajectories, using a simulation-free regression objective rather than adversarial training. To prevent semantic drift, we couple left-to-right transport between consecutive domains with a source-anchored transport and a semantic consistency loss, and we further stabilize learning via a two-stage sequential optimization scheme. Extensive experiments demonstrate that our method outperforms state-of-the-art techniques, validating flow matching as an effective alternative for domain adaptation.
PaperID: 5770, Poster
Authors: Constantino Msigwa, Denis Bernard, Jaeseok Yun
Abstract: Multimodal 3D detection becomes unreliable under partial observability caused by occlusion, LiDAR sparsity, and camera--LiDAR miscalibration. Most fusion detectors are fully feed-forward: once trained, they provide no explicit mechanism to incorporate structured test-time constraints or to bound how such constraints can alter predictions. We propose Intervene3D, a controlled-inference wrapper for BEV-based detectors that performs metric-bounded latent refinement at inference time. Given sensor inputs, the detector produces an initial BEV latent \mathbf z_0 and an evidence-conditioned trust metric \mathbf H (a diagonal precision derived from predicted uncertainty). A separate control stream is mapped to a fully specified differentiable constraint energy (count intervals and spatial admissibility). Inference then updates \mathbf z using K\in\1,2,3\ diagonal-metric preconditioned projected steps that reduce constraint energy while remaining within an uncertainty-gated trust region defined by \mathbf H. To improve robustness to extrinsic drift, we additionally introduce phase-aware spectral alignment that normalizes cross-modal Fourier-phase discrepancies prior to attention. We report clean accuracy and a severity-swept robustness protocol (robustness-AUC, feasibility diagnostics, and conflict-induced FP inflation).
PaperID: 5771, Poster
Abstract: Offline reinforcement learning is a standard paradigm for fine-tuning Large Language Models (LLMs) on multi-turn dialogue, where on-policy rollouts and reward annotations are costly. Within this paradigm, reward-weighted supervised fine-tuning has emerged as an efficient approach, weighting each trajectory's log-likelihood by a reward-dependent coefficient. We show that existing reward-weighted SFT methods are specific members of a broader affine reward-weighted SFT class, and identify two structural limitations that hold for every member of this class: every centered member yields an unbounded-below loss, and no member admits a hyperparameter for attenuating reward noise. We introduce ExpRFT as a principled escape from this class — an exponential reweighting of the standardized advantage at a tunable temperature that restores a well-posed objective and provides a noise-control dial. ExpRFT parallels Advantage-Weighted Regression (AWR) in spirit, yet requires no per-action learned critic and adds no infrastructure cost beyond standard SFT. On six multi-turn dialogue benchmarks under reasoning and clarifying-question protocols, ExpRFT outperforms rejection-sampling, preference-based, value-based, and reward-weighted baselines. In the near-mean trajectory regime, where weight noise dominates the advantage signal, ExpRFT remains stable under injected reward noise while baselines significantly degrade.
PaperID: 5772, Poster
Abstract: Text- and image-conditioned world generators can now produce visually rich 3D environments, but these worlds are usually silent or paired with a soundtrack generated only from text or rendered video. Such audio can describe what should be heard, but it does not explicitly represent where sounds live in the world or how they should change as a listener moves. We propose Audible World Models, a training-free framework that treats sound as part of a generated world state. Given a text prompt, our system builds a panoramic 3D proxy, decomposes it into semantic layers, identifies audible foreground objects and ambient background regions, synthesizes dry audio for each sound label, attaches sources to reconstructed geometry, and renders listener-dependent spatial audio through geometric acoustic propagation. This explicit coupling of semantics, geometry, and propagation produces audio that remains tied to persistent source locations and responds to viewpoint and motion. Across 80 generated scenes, our method substantially improves spatial consistency over text-, video-, and panorama-conditioned baselines while maintaining competitive semantic alignment. VLM-based and human evaluations further show that the resulting soundtracks are preferred for audio--visual consistency, spatial plausibility, and motion-dependent behavior.
PaperID: 5773, Poster
Abstract: Listwise preference feedback offers richer supervision for large language model (LLMs) alignment than pairwise comparisons, but a full ranking also reveals more sensitive user preferences. Existing privacy-preserving alignment methods focus mainly on pairwise feedback, while listwise alignment under local differential privacy (LDP) remains unexplored. Extending LDP to this listwise domain raises a critical challenge: the high sensitivity of the full ranking space demands excessive noise, neutralizing the benefits of listwise supervision. To address this challenge, we analyze the interaction between supervision granularity and privacy noise in downstream training, and introduce a rank-aware exponential mechanism that privatizes listwise preference data into a low-sensitivity granularity suitable for downstream alignment. The mechanism leverages ranking information to sample a fixed-size binary partition, concentrating the noise near borderline items instead of perturbing all items uniformly as randomized response (RR) does. Empirical evaluations on downstream LLM alignment tasks show that our mechanism consistently outperforms existing LDP baselines at matched privacy budgets.
PaperID: 5774, Poster
Abstract: The mechanism governing the training dynamics of Quantum Neural Networks (QNNs) remains under-explored. In classical Deep Neural Networks (DNNs), training is known dominated by "Spectral Bias,'' i.e. prioritizing learning low-frequency while struggling with high-frequency. QNNs exhibit similar spectral limitations when the target function possesses flat spectral amplitudes. However, in this work, considering target problems with different spectral amplitudes, we theoretically and empirically identify a distinct mechanism in QNNs, which we term . By analyzing the frequency-domain gradients and residual dynamics via the Quantum Neural Tangent Kernel (QNTK), we prove that QNN training is governed primarily by the magnitude of spectral components rather than their frequency indices. Consequently, QNNs can efficiently capture high-frequency functions—provided they have significant amplitude—thereby overcoming the inherent limitations of their classical counterparts. We validate this principle on both synthetic high-frequency functions and quantum-advantage tasks. These results show that QNNs notably outperform DNNs in high-frequency tasks, offering an explanation for QNNs' superior expressivity.
PaperID: 5775, Poster
Abstract: Reconstructing a dynamic 4D scene along a novel camera trajectory requires both metric geometry from the source video and generative completion of disoccluded target-view regions. Feed-forward 4D reconstruction models recover increasingly comprehensive input-view geometry, including depth, point maps, and dense 3D tracks, but remain tied to the observed camera stream. Camera-controlled video diffusion models (VDMs) provide strong generative priors for novel-view synthesis, yet their camera control is typically not precise enough for metric novel-view reconstruction. We argue that comprehensive input-view 4D reconstruction is an effective geometric interface between these two capabilities: properly exposed to a VDM, it turns reconstruction outputs into camera-accurate novel-view RGB videos that remain geometrically consistent under downstream 4D reconstruction. We introduce HOLO4D, a geometry-aware video-to-video framework that controls a pretrained VDM with a hybrid 4D geometric cache built from the source video. A small set of source views covering the camera motion forms the static cache, while dynamic regions are rendered from the matching source frame. The resulting hybrid RGB-D representation is combined with explicit valid masks, target-camera ray embeddings, and dense 3D tracking tokens. Conditioned on this 4D scaffold, the diffusion backbone synthesizes the target-view RGB video. Under a common downstream 4D reconstructor, our generated videos yield target-view depth more consistent with the held-out reconstruction than that of the baselines, and an additional feature-level readout shows that intermediate VAE features already encode target-view depth. Experiments and ablations further show that progressively richer 3D/4D reconstruction signals improve camera adherence.
Authors:
Shreyas Rajesh, Kartik Sharma, Tonmoy Monsoor, Mehmet Yigit Turali, Richard Idro, Juliana Kayaga, Robert Sebunya, Tracy T Namata, Jessica N Pasqua, vwani Roychowdhury, Rajarshi MazumderAbstract: Specialist epilepsy expertise is scarce in resource-constrained settings, making LLM-based decision support attractive for frontline clinicians managing longitudinal treatment. Such support systems must do more than apply medical knowledge: they must adapt to local prescribing practice and know when to defer. Public medical AI benchmarks are dominated by high-income clinical settings, leaving prescribing practices, medication availability, and follow-up patterns in low-resource contexts largely unrepresented. We study this problem through a multidisciplinary collaboration in Ugandan pediatric epilepsy care. The task is to predict anti-seizure medication regimens from longitudinal unstructured notes collected by local clinicians across serial visits. Standard prompting achieves non-trivial agreement with physician prescriptions, but neurologists' review of model reasoning traces shows that its errors stem from distribution-miscalibrated prescribing defaults rather than the local care environment. We introduce Manana, a non-parametric prompt-learning framework that learns how to reason about local prescribing decisions from a small patient-level training set. Manana turns observed prescription errors into an auditable prompt memory, instantiated in single-agent and multi-agent variants, and outperforms classical ML models, direct LLM prompting, and prompt-optimization baselines across two independently collected Ugandan cohorts. To make the system uncertainty-aware, we propose Bayesian prompt averaging (BPA), a Bayesian model averaging procedure over the learned prompt trajectory. This converts a sequence of learned prompts into prescription likelihoods and produces a deferral signal. On the independently collected held-out cohort, BPA improves visit-level top-3 prescription accuracy by 4-8 percentage points over the prompt-optimization baselines. More consequentially, it enables clinically meaningful selective prediction: the system can auto-handle the most confident half of cases at 95% precision, or the most confident quarter at 99% precision, while deferring lower-confidence cases for specialist review. These results suggest a path toward locally adapted clinical LLM systems that learn from limited site-specific data and reserve scarce specialist attention for the cases where uncertainty is highest.
PaperID: 5777, Poster
Abstract: Emerging latent reasoning paradigms allow Large Language Models to “think” in latent embedding spaces, offering a more efficient alternative to explicit Chain-of-Thought. However, they generally face two critical limitations: First, relying exclusively on latent tokens amplifies uncertainty due to their inherently high entropy, thereby degrading solution correctness. Second, the overthinking problem remains prevalent, and existing approaches typically rely on rigid heuristics for stopping, often resulting in suboptimal termination. To address these challenges, we view efficient reasoning as a learnable control problem over both how to think and when to stop. We formulate this within a Reinforcement Learning framework, training lightweight per-step gates on a frozen LLM. This modulation enables the model to (i) adaptively interleave latent soft-token reasoning with explicit token generation, and (ii) predict cumulative exit probabilities for robust early termination. Experiments on diverse reasoning benchmarks across varying model scales and families demonstrate that our method surpasses both standard CoT and existing latent baselines, achieving accuracy gains of ∼3% while reducing token consumption by up to 40%.
PaperID: 5778, Poster
Abstract: Domain specialization of large language models (LLMs) often improves target-domain performance but degrades broad capabilities. Model merging is an attractive training-free alternative, yet common parameter-space strategies (for example, interpolation and sparsified averaging) provide limited control over task-specific interference. We introduce Adaptive Model Alignment via Ranked Inherited Subspaces (AMARIS), a covariance-driven framework that transfers specialist updates only along selected activation directions. AMARIS formulates merging as budgeted subspace inheritance: it maximizes specialist activation capture while constraining base-activation deviation, then builds a low-rank projection gate from ranked generalized eigendirections under an explicit preservation budget. This yields a single static merged checkpoint that concentrates transfer in high-utility, low-interference directions. We evaluate AMARIS on Qwen3-based merges (8B and 32B) across finance, instruction-following, and multilingual SEA specialization. In our evaluated settings, AMARIS reaches Pareto operating points that are competitive with strong training-free baselines, including 98.37% specialist retention with 95.76% core retention (finance), 103.49% core retention with 90.28% specialist retention (instruction following), and 93.69% specialist retention with 99.77% core retention (SEA). These results support subspace-constrained merging as a practical route to compositional capability integration without additional training.
Abstract: Visual Text Comprehension (VTC) renders text into images for a vision-language model (VLM) to read, sidestepping LLM context-window limits and powering applications from long-page OCR to multi-page memory QA. Yet existing VTC pipelines treat rendering and layout as a fixed, content-agnostic preprocessing step, and offer little mechanistic understanding of how VLMs internally process visualized text. Through a focused empirical study on VTC QA tasks, we find that VLMs exhibit a localization-without-utilization regime: evidence-localizing attention emerges sharply in the middle-to-late layers and is largely decoupled from answer correctness, yet simply enlarging the localized spans on the rendered page recovers a large fraction of the failures. Building on these observations, we propose AGAR (Attention-Guided Adaptive Rendering), a training-free, model-agnostic method that leverages a VLM's own middle-to-late layer attention to identify the top-K important visual patches, maps them back to word spans, and re-renders the page with those spans enlarged before re-inferring the answer. Extensive experiments across nine VTC benchmarks (short-form, long-context, and multi-page memory QA) and four VLM backbones show that AGAR (i) consistently improves off-the-shelf VLMs as a plug-and-play enhancement, (ii) composes with VLM post-training to yield further gains, and (iii) remains robust under both visual- and text-side input degradation.
Abstract: Vision-Language-Action models (VLAs) have shown strong potential for general-purpose robotic manipulation, yet their closed-loop reliability often degrades under local deployment conditions. Existing evaluations typically treat test episodes as independent zero-shot trials, whereas real robots often operate repeatedly in the same or slowly changing environments, where successful executions provide environment-verified evidence about reliable behavior patterns. We study this persistent-deployment setting and ask whether a partially competent frozen VLA can improve its reliability by reusing its own successful test-time experience. We propose an online success-memory guided test-time adaptation framework for generative VLAs. During deployment, the robot stores progress-calibrated successful observation-action segments in a long-term memory. At inference time, it retrieves state-relevant successful action chunks, filters action-inconsistent candidates through trajectory-level consistency, and aggregates the filtered candidates into an elite action prior. To incorporate this prior into action generation, we introduce confidence-adaptive prior guidance, which injects the elite prior into an intermediate state of the flow-matching action sampler and adjusts the guidance strength according to retrieval confidence. This design allows the frozen VLA to exploit environment-specific successful experience while preserving observation-conditioned generative refinement. This retrieve-then-steer mechanism enables lightweight, non-parametric test-time adaptation without updating model parameters or modifying the generative solver. Experiments in simulation and real-world manipulation demonstrate improved task success and closed-loop stability, particularly in long-horizon and multi-stage tasks.
PaperID: 5781, Poster
Authors:
Canyu Chen, Yuguang Yang, Jianing Pang, Zhewen Tan, Cheng Chi, Chunyang Liu, Kehua Sheng, Bo Zhang, Linlin Yang, Xiaoyan Luo, Yan Wang, Baochang ZhangAbstract: Vision-language-action (VLA) models have emerged as a promising paradigm for autonomous driving trajectory planning. While explicit test-time thinking has driven mainstream progress in Large Language Models (LLMs), how to effectively utilize reasoning in driving VLA remains unsettled. We revisit this gap through Coarse-to-Refine (C2R), a trajectory self-refinement framework for driving VLA that completes refinement in a single autoregressive pass. C2R first predicts a coarse trajectory, performs explicit reasoning over it, and then generates a refined trajectory autoregressively. We show that the refinement formulation itself drives imitation learning improvement while enabling reinforcement learning to utilize reasoning content for further optimization. C2R establishes VLA state-of-the-art performance across standard and challenging benchmarks, achieving 92.1 PDMS on NAVSIM v1, 90.4 EPDMS on NAVSIM v2, and 49.6 EPDMS on NavHard.
PaperID: 5782, Poster
Abstract: Mental simulation, the ability to reason by constructing and updating internal visual states, is central to flexible human cognition, but remains elusive in current multimodal large language models. Latent multimodal reasoning offers a promising route toward this capability by interleaving text with continuous visual latent blocks, allowing models to maintain internal visual workspaces without decoding them into pixels. In this paper, we ask whether these latent visual states actually matter. Through a direct perturbation analysis across representative latent-reasoning MLLMs, we find that they often do not: dropping, replacing, or shuffling generated visual latents has only minor effects on the model's own final-answer probability, while equivalent perturbations to the linguistic reasoning trace are substantially more damaging. We call this failure mode Mental Simulation Collapse (MSC): the model appears to perform interleaved visual reasoning, but routes most answer-relevant computation through text. We further diagnose that MSC follows naturally from existing training objectives: visual alignment can make generated latents resemble helper-image embeddings, but it does not require the decoder to rely on them; subsequent text-only supervision further leaves the semantic role of the latent pathway unidentifiable. To address this, we introduce AnchorSim (Anchored Mental Simulation), a contrastive distillation framework that trains models to make their generated visual latents consequential. Across multiple visual reasoning benchmarks, AnchorSim improves performance, while diagnostic re-evaluation shows that visual perturbations now meaningfully affect predictions. Our results expose a fundamental failure mode of latent multimodal reasoning and demonstrate that mental simulation must be anchored to model behavior in order to matter.
PaperID: 5783, Poster
Authors:
Fengyun Wang, Jian Wang, Dingwei Zhang, Jinhui Tang, QIANRU SUNAbstract: Using 2D vision-language models (VLMs) for 3D grounding raises a decision-formulation problem: how should image-level evidence be turned into a 3D object decision? This problem matters as pretrained 2D VLMs provide strong visual-semantic evidence, while 3D grounding requires disambiguating objects in cluttered scenes. Current VLM-based formulations make this conversion holistically: the model either selects one object from a rendered candidate set or assigns one overall match judgment to each candidate. Such decisions obscure the evidence needed for disambiguation, since different candidates may satisfy different subsets of the category, attribute, and relational cues. In this work, we reformulate 3D grounding with 2D VLMs as condition-wise evidence aggregation. The reformulation follows three principles: use natural object-centric images rather than rendered candidate images; read the referring expression condition by condition rather than as one holistic query; and use model preference rather than a hard match output as grounding evidence. We instantiate these principles in QGround, a training-free implementation of this decision formulation. Experiments on ScanRefer and Nr3D show improved grounding performance, with analyses demonstrating stronger candidate disambiguation and more interpretable decisions.
PaperID: 5784, Poster
Abstract: As collaborative machine learning scales to thousands of edge devices, two failure modes emerge: centralized federated learning creates server bottlenecks and single points of failure, while flat peer-to-peer systems force every node to increase its communication degree as the network grows. We propose Tiered Gossip Learning (TGL), a two-layer push–gossip–pull protocol that decouples data-holding leaf nodes from a decentralized relay backbone. In each round, leaves push local models to a subset of relays, relays gossip among themselves to mix models globally, and leaves pull updates from another random relay subset. This asymmetric design keeps per-leaf communication fixed as the network scales, offloading the mixing burden to a thin layer of higher-capacity relays without relying on any central coordinator for aggregation. Across CIFAR-10, FEMNIST, and AG News, TGL matches or exceeds baseline accuracy with up to 80% fewer model exchanges. We provide convergence guarantees under standard smoothness, bounded variance, and heterogeneity assumptions, with explicit stage-wise consensus bounds characterizing how relay-layer connectivity governs the global mixing rate.
PaperID: 5785, Poster
Abstract: Manifold embedding for data visualisation is dominated by a small set of hand-designed state-of-the-art (SOTA) algorithms. Their design targets are fixed at development time, resulting in inflexible algorithmic behaviour that cannot be readily steered toward user-specified structural priorities. Hand-designing new embedding algorithms to address user priority requires highly specialised expertise and is time consuming. To address autonomous algorithm design tailored to user preference, we introduce an agentic algorithm generation pipeline AutoManifold. It composes new manifold embedding algorithms from a constrained vocabulary of affinity, cost, and optimisation primitives, supported by multi-agent large-language-model (LLM) orchestration. AutoManifold conditions every stage of its design on user-specified structural preservation preferences, and iteratively refines the algorithm configuration through an LLM-guided, metric-grounded iterative loop. We compare the generated algorithms against three strongest and most frequently used SOTA (t-SNE, UMAP, and PaCMAP) on real-world datasets spanning different difficulty regimes under identical evaluation infrastructure. The results demonstrate that LLM-driven, priority-conditioned algorithm synthesis can move beyond hyperparameter tuning, producing genuinely new algorithms that outperform traditional hand-designed methods, in their ability to steer towards user-specified properties without compromising much inherent structure in the original data.
PaperID: 5786, Poster
Authors:
Zhiwen Luo, Dayu Guo, Manar Amayri, Nizar Bouguila, Zhixiang Li, Wenchuan Zhang, Wentao FanAbstract: Multimodal evidence can improve topic discovery, but dense auxiliary signals can make topics harder to interpret when they replace readable descriptors. This creates an evidence-role problem: symbolic signals should define what topics mean to a reader, while dense aligned signals should help infer which topics a document expresses. We introduce M3-HNTM, a hyperspherical multimodal neural topic model that separates these roles through a shared document-level latent direction. M3-HNTM utilizes a shared von Mises-Fisher (vMF) posterior for document-level topic inference and vMF-mixture symbolic decoders that represent topics as concentration-aware semantic regions. Text Bag-of-Words and speech Bag-of-Acoustic-Words provide reconstructable symbolic evidence, while aligned image-text embeddings enter posterior inference as contextual evidence and help preserve semantic structure. The model applies to image-text, speech-text, and image-speech-text corpora while keeping topics readable through symbolic descriptors. Experiments on five datasets show that M3-HNTM improves the coherence-diversity quality score over strong text-only, hyperspherical, and multimodal baselines. On SpokenCOCO-Tri, the full tri-modal model outperforms both image-text and speech-text variants, indicating that visual and acoustic signals provide complementary evidence for interpretable topic discovery.
Authors:
Christopher Ellis, Mei-Yu Wang, Shreyas Chaudhari, Leighton P Barnes, Giulia Fanti, José M MouraAbstract: In practice, most commercial LLM providers do not publicly release details of underlying LLM architectures. However, prior work has shown that given limited API access to an LLM (namely, top-k logits and/or a logit bias function), one can recover certain architectural details of an LLM, such as the hidden dimension of the feed-forward network. Perhaps in response to these results, most commercial LLM providers have restricted their APIs to expose only the single logit for each decoded token, and they no longer give users the ability to bias logits. We show that even under current restrictive APIs, several architectural parameters are still recoverable. We present NightVision, an attack that uses restrictive black-box API access to estimate the hidden dimension, depth, and parameter count of an LLM. Algorithmically, NightVision relies on a novel common set prompting technique in which multiple prompts expose log probabilities for the same set of output tokens; a spectral analysis of these results is used to infer hidden dimension. NightVision additionally uses end-to-end time to first token (TTFT) measurements and the estimated hidden dimension to estimate depth and parameter count. We empirically evaluate NightVision on 32 open-source LLMs, recovering hidden dimension to within 23% average relative error across all models (9% on MoE models), and depth and parameter count to within 53% for models exceeding three billion parameters. We run extensive ablations to demonstrate how these accuracies scale with token budget and model properties. Overall, our results suggest that current LLM APIs are not sufficiently restricted to fully obfuscate the architectural details of their underlying models.
PaperID: 5788, Poster
Abstract: Large Language Models (LLMs) have advanced rapidly, raising growing concerns about their safety. Recent work has proposed various approaches to detect and defend against adversarial attacks including defense mechanisms at the decoding stage that leverage models' internal hidden states. However, existing decoding-stage defenses suffer from two limitations. First, they introduce a trade-off between safety and over-refusal, where strengthening safety degrades the model's helpfulness on benign queries. Second, many of these methods rely on internal hidden states and are thus restricted to specific architectures, incurring substantial overhead and limited generalization across models. To address these limitations, we introduce LADE (Latent Safety Signals for Defense), which leverages latent safety signals extracted by contrasting harmful and benign queries from dark knowledge (i.e., information carried by the output probability distribution beyond its argmax) in the first-token output probability distribution. Our key insight is that, beyond surface-level refusal tokens, the dark knowledge in the first-token distribution contains latent safety signals, defined as tokens whose probabilities differ sharply between harmful and benign queries. We empirically show that these signals consistently align across diverse LLMs, forming a model-agnostic direction that reflects an intrinsic property of safety-aligned language models. LADE consists of three components: (1) Extracting Latent Safety Signals from Dark Knowledge, which selects top-k safety-discriminative tokens from the first-token probability distribution; (2) Tokenizer Mapping, which maps these tokens across different tokenizers to enable model-agnostic application; and (3) kNN-based Discrimination, which classifies queries via a k-Nearest Neighbors search over the mapped tokens. Across six LLMs and multiple benchmarks, LADE remains robust against a wide range of jailbreak attacks and lowers attack success rates with minimal over-refusal.
Abstract: Many training-free sparse attention methods are effective for accelerating diffusion models. Recently, several works suggest that making sparse attention trainable can further increase sparsity while preserving generation quality. We study three key questions: (1) when do the two common masking rules, i.e., Top-k and Top-p, fail, and how can we avoid these failures? (2) why can trainable sparse attention reach higher sparsity than training-free methods? (3) what are the limitations of fine-tuning sparse attention using the diffusion loss, and how can we address them? Based on this analysis, we propose SpargeAttention2, a trainable sparse attention method that achieves high sparsity without degrading generation quality. SpargeAttention2 includes (i) a hybrid masking rule that combines Top-k and Top-p for more robust masking at high sparsity, (ii) an efficient trainable sparse attention implementation, and (iii) a distillation-inspired fine-tuning objective to better preserve generation quality during fine-tuning using sparse attention. Experiments on video diffusion models show that SpargeAttention2 reaches 95% attention sparsity and a 16.2× attention speedup while maintaining generation quality, consistently outperforming prior sparse attention methods.
PaperID: 5790, Poster
Authors: Igor Sokolov, Mathieu Garrigues, Matthias Hecker, Lorenzo Moro, Shaheen Acheche
Abstract: Structural descriptors provide graph neural networks (GNNs) with positional and topological inductive biases. However, widely used classical encodings based on random walks or Laplacian spectra emphasize specific structural statistics of a graph, leaving room for alternative descriptors. Neutral-atom quantum processing units (QPUs) can realize graph-dependent Rydberg Hamiltonians, allowing graph topology to be probed through dynamical correlation measurements. In this work, we use these measurements as quantum graph features (QGFs), providing a complementary, hardware-derived family of structural encoding that --- as we demonstrate empirically --- is not accurately described by classical features derived from Laplacian or adjacency matrix. By appropriately selecting the quench parameters of the Hamiltonian, we access two characteristic correlation regimes with distinct structural signatures: antiferromagnetic-like and ferromagnetic-like. On an intrinsic information benchmark of 15 topological and spectral probing tasks, quantum and hybrid quantum-classical features match or surpass classical descriptors on local topological properties. When supplied to message-passing, attention, and graph-transformer backbones, hybrid quantum-classical features yield small improvements that are within cross-fold noise on most cells but sign-consistent across architectures. When comparing noiseless emulation-derived quantum features with an end-to-end neutral-atom QPU implementation, the downstream hybrid gain over the classical baseline persists on five out of six backbones.
PaperID: 5791, Poster
Authors:
Chenyu Zhang, Yuhang Cao, Yingxi Lu, Daru Du, Jing Shao, Jiajun Liu, Ruoqu Chen, liu cao, Yicheng Liu, Hang Zhao, Mengdi XuAbstract: Autoregressive Vision-Language-Action (VLA) models offer a scalable path to robot learning, yet existing action tokenizers treat tokenization as a compression problem, producing representations that are semantically misaligned with the autoregressive backbone. We propose CATok, a causal action tokenizer that reframes tokenization as a causally structured generative process. CATok introduces a conditional annealing mechanism that extracts action tokens by progressively annealing a flow-matching process: each token is conditioned on all preceding tokens and encodes the residual reconstruction signal at a specific noise level, establishing a coarse-to-fine causal token space whose generative semantics are structurally aligned with autoregressive modeling. A token-conditioned flow-matching decoder built on Multimodal Diffusion Transformer (MMDiT) reconstructs continuous action chunks from these discrete tokens with the precision of hybrid diffusion-head architectures. This discrete bottleneck enforces knowledge insulation by design, cleanly separating high-level semantic reasoning from low-level motor execution without requiring explicit attention masking. Extensive evaluations across three simulation benchmarks demonstrate that CATok consistently surpasses existing tokenization methods in both reconstruction fidelity-compression tradeoff and inference efficiency, while improving VLA task success rate, establishing a high-performance, scalable foundation for purely autoregressive VLA systems. Our project page is available at https://causalactiontokenizer.github.io.
PaperID: 5792, Poster
Abstract: Post-hoc residual correction can reduce forecasting error, but always-on updates often fail to generalize: when residual signals are weak or noisy, expressive correctors can overfit and even increase the overall forecasting errors. We introduce CRC, a robust residual correction framework that couples a hybrid corrector with a validation-calibrated selective deployment rule. CRC proposes corrections via a conservative ridge "floor" plus conditional nonlinear refinement, but deploys them only when held-out evidence indicates net error reduction over a frozen baseline forecaster. This reframes correction as a decision problem: when to revise a forecast, not only how to fit residuals. Across standard long-sequence benchmarks and multiple backbones, CRC improves horizon-averaged MSE and MAE in most dataset-backbone pairs, with particularly large gains on complex datasets (e.g., over 20% relative MSE reduction on Traffic). We additionally report the non-degradation rate (NDR) as a stability diagnostic of harmful updates and ablate the validation-calibrated "firewall" that stabilizes MSE gains. Formal analysis motivating the design appears in Appendix A under idealized assumptions.
PaperID: 5793, Poster
Abstract: Detecting AI-generated videos (AIGV) in a generator-agnostic manner is increasingly important as generative models close the gap with real footage, and a video, with many frames, intuitively offers richer forensic evidence than a single image. Yet existing detectors fall short of this expectation: under the encode-then-fuse paradigm, increasing the frame budget from 1 to 8 improves a voting-based image detector by only 5.28%, and even temporal-modeling backbones such as TimeSformer barely improve on this. Per-frame predictions are also noisy: per-frame logits fluctuate substantially within a video and correlate weakly with any frame-level proxy. We trace the cause not to the aggregation step but to per-frame encoding severing multi-frame cues before they can interact, leaving any post-hoc selection, weighting, or aggregation strategy without a reliable basis. We therefore propose , a blend-then-encode framework that combines multiple frames in the pixel domain into a single composite image (Stochastic Frame Blending, SFB) and regularizes the encoder for consistency across blended views (Blend Invariance Regularization, BIR). To stress-test this approach, we further introduce , a test-only benchmark covering ten recent generators (including five closed-source commercial platforms) and two partial-forgery scenarios (temporal splicing and conditional continuation), with matched controls that isolate generative artifacts from editing artifacts. On AIGVDBench and MixForensics-Bench, MixForensics outperforms the strongest baseline by 1.27% and 8.11% on the fully-fake splits, leads on both partial-forgery scenarios by 11.49% on temporal splicing and 9.15% on conditional continuation, and uses fewer encoder forward passes than per-frame methods.
PaperID: 5794, Poster
Abstract: Drifting Models have recently achieved state-of-the-art performance in one-step image generation by training generators to follow corrective vector fields. However, their standard reverse-style drift improves samples from the model side, which can severely underutilize training samples and bottleneck early convergence. To address this, we introduce DualDrift, a unified framework that augments the original reverse drift with a novel forward drift. The forward drift lets each data sample assign correction to generated rollouts, providing denser data-side signals and accelerating early training. At the same time, this dense data-side signal can overemphasize dominant regions when used alone. DualDrift therefore uses a parameter-free scheduler driven by empirical batch statistics to balance the fast early correction of forward drift with the stable refinement of reverse drift. Evaluated on ImageNet 256×256, DualDrift achieves faster convergence and significantly improves over the baseline Drifting Models in both FID and Inception Score.
Abstract: Learning a shared representation between spoken text and gesture is central to co-speech gesture retrieval, synthesis, and understanding, but remains challenging for semantically meaningful gestures whose communicative intent is not captured by motion alone. Direct contrastive alignment between transcripts and continuous motion embeddings often overemphasizes low-level kinematics and misses the symbolic content of semantic gestures. We propose \emphsemantic motion anchors, natural-language abstractions of gesture motion capturing physical form and communicative intent. Our method discretizes 3D gestures into body-hand motion primitives, verbalizes them into structured descriptions, and grounds them in the transcript to provide auxiliary contrastive supervision. On BEAT2, our method improves text-to-gesture R@1 by 8.2% over a direct text-motion baseline and outperforms prior retrieval approaches on text to gesture and gesture to text retrieval directions. Beyond aggregate retrieval metrics, semantic motion anchor supervision helps retrieve gestures that are semantically meaningful for the spoken query, rather than defaulting to generic motion patterns. A downstream retrieval-augmented gesture generation study showed that users significantly preferred gestures retrieved by our approach over a retrieval-augmented generation baseline, demonstrating that semantically grounded retrieval translates to gestures that better convey communicative intent in downstream generation.
Authors:
Mohammad Shahverdikondori, Amir Mohammad Abouei, Alireza Rezaeimoghadam, Negar KiyavashAbstract: We introduce the problem of best arm identification (BAI) with post-action context, a new BAI problem in a stochastic multi-armed bandit environment and the fixed-confidence setting. The problem addresses the scenarios in which the learner receives a \emphpost-action context in addition to the reward after playing each action. This post-action context provides additional information that can significantly facilitate the decision process. We analyze two different types of the post-action context: (i) separator, where the reward depends solely on the context, and (ii) non-separator, where the reward depends on both the action and the context. For both cases, we derive instance-dependent lower bounds on the sample complexity and propose algorithms that asymptotically achieve the optimal sample complexity. For the separator setting, we propose a novel sampling rule called G-tracking, which uses the geometry of the context space to directly track the contexts rather than the actions. For the non-separator setting, we do so by demonstrating that the Track-and-Stop algorithm can be extended to this setting. Moreover, in both settings, we theoretically and empirically show that algorithms that ignore the post-action context are sub-optimal. Finally, our empirical results showcase the advantage of our approaches compared to the state of the art.
Abstract: While recent vision-language models (VLMs) demonstrate strong multimodal understanding, they remain limited in spatial reasoning tasks that require active evidence acquisition and multi-step visual interaction. This limitation suggests that relying solely on implicit visual representations from vision encoders is insufficient for recovering fine-grained spatial evidence. We introduce ), a tool-augmented visual agent for spatial reasoning tasks across map reasoning, visual probing, and vision reconstruction. , we develop a unified recipe that combines supervised tool-use trajectory synthesis, composite rewards, and Observation-Relaxed Group-in-Group Policy Optimization (OR-GIGPO) for effective multi-tool behavior. Experiments on 13 benchmarks from 8 datasets show that -8B improves over the Qwen3-8B backbone by 10.0% on in-distribution benchmarks and 4.4% on out-of-distribution benchmarks, while outperforming previous state-of-the-art baselines of similar size by 7.0%–14.8%. It also achieves performance comparable to much larger models such as Qwen3-VL-235B-A22B-Thinking and GPT-5, demonstrating the effectiveness of
PaperID: 5798, Poster
Abstract: Generative AI systems are increasingly deployed at scale to support human users in creative tasks across a variety of domains. However, as more users adopt the same AI agent (e.g., ChatGPT), the resulting artifacts may become increasingly homogeneous, leading to a collapse in collective diversity. Prior work has documented this risk, but lacks a systematic framework for understanding how collective diversity changes as AI adoption scales across co-creation settings. In this work, we conduct a large-scale crowdsourcing study to understand and model the dynamics of collective diversity, quality, and effort in human–AI co-creation for three creative tasks. We find that current AI agents used in co-creation settings can improve artifact quality and reduce effort, but that these gains come at the cost of declining collective diversity as adoption increases. We further show that more structured and fine-grained human-AI interaction, along with more diverse agent designs, can mitigate this collapse. Finally, we derive predictive models that capture how collective diversity scales across co-creation settings, providing valuable insights for the design of AI agents that aim to preserve collective diversity.
Authors: Xin Guo, Grace He, Xinyu Li
Abstract: We study contextual bandits with nonlinear and path-dependent rewards through a novel signature-transform-based approach. Leveraging the universal nonlinearity property of signatures, we approximate continuous path-dependent reward functionals by linear functionals in the signature space. This representation enables the use of efficient linear contextual bandit methods while preserving expressive sequential structure. Building on this framework, we propose \textttDisSigUCB, a signature-based disjoint upper confidence bound (UCB) algorithm. Under boundedness and non-degeneracy assumptions, we prove a high-probability data-dependent sublinear regret bound of order \tilde\mathcal O(\sqrt(d+m)KT) where d is the context dimension and m is the signature feature dimension. Experiments on temperature sensor monitoring, sleep-stage classification, and hospital nurse staffing demonstrate that \textttDisSigUCB consistently outperforms classical linear and kernelized contextual bandit baselines in nonlinear and path-dependent settings.
PaperID: 5800, Poster
Abstract: GraphRAG enhances Large Language Model reasoning by organizing retrieved knowledge into structured evidence graphs, enabling inference over connected evidence rather than isolated text fragments. Yet, existing GraphRAG methods often either miss query-critical links or introduce noisy or conflicting evidence that distracts reasoning. We propose CONSIST, a Conflict-aware Nash-optimized Minimal Sufficient graph learning algorithm that regulates evidence graphs to resolve conflicts during reasoning. Specifically, we cast graph learning as an exploration–editing process, where a candidate graph is expanded via structural and semantic connections and then refined by pruning unreliable edges. We further introduce Gumbel-based edge editing, framing it as a Nash-guided bargaining process, where current and future players negotiate over a customized utility to determine reliable edits that improve both immediate answer support and long-term graph quality. Experiments on multi-hop question answering and conflict-aware benchmarks demonstrate that CONSIST consistently outperforms recent state-of-the-art baselines, achieving average performance gains up to 7.0% EM and 6.3% F1. Moreover, CONSIST also produces substantially more compact evidence graphs, reducing retained nodes by 42.9% and edges by 54.7% on average.
Abstract: A central challenge in building intelligent systems is enabling agents to jointly perceive complex inputs, form hypotheses about hidden patterns, and design informative experiments to test them. To study this problem, we propose ZendoWorld, a controlled interactive environment in which agents must infer a logical rule about visual game observations, acquire information by proposing new scenes, and refine their hypotheses based on feedback from the game environment. We evaluate several agents spanning pure VLM reasoning, Bayesian particle filtering, dynamic concept discovery, and neuro‑symbolic methods. Our main findings are: (1) high accuracy in predicting labels for observed examples does not imply recovery of the underlying rule; (2) perception and induction are distinct bottlenecks for different agent classes; and (3) VLM‑based agents propose near‑uninformative experiments, failing to actively reduce hypothesis uncertainty. To compare these results, we collect human data on the task, which reveals a gap in inductive reasoning, particularly for more complex rules. Overall, ZendoWorld takes an important step toward evaluating intelligent agents and identifies concrete avenues for improvement, particularly in domains like scientific discovery.
PaperID: 5802, Poster
Abstract: Minimax optimization refers to a class of optimization problems with some variables to minimize and other variables to maximize. Biased stochastic gradient methods have shown practical success in solving minimax problems to either improve robustness, enhance communication efficiency, or decrease computational costs. These successes motivate a lot of theoretical works to study the convergence of biased stochastic gradient methods, while their generalization analysis remains untouched. In this paper, we present the first framework to study the stability and generalization of biased stochastic gradient methods for minimax problems. We establish a connection between stability and generalization for minimax problems by relaxing the existing bounded gradient assumption to a bounded second moment condition. We then introduce a generalized Lipschitz-type condition on bias and gradient estimators, and derive a general stability bound to clarify the connection among bias, gradient estimators and stability. We apply our general analysis to Zeroth-order and Clipped stochastic gradient descent ascent (SGDA), and derive stability bounds that match those of SGDA under appropriate smoothing/clipping parameters. We combine stability and convergence analyses together, and derive optimal excess risk bounds of order 1/\sqrtn, where n is the sample size.
PaperID: 5803, Poster
Abstract: AI agents are governed by hundreds of natural language criteria---yet existing verifiers cannot efficiently evaluate an agent's output against a large set of such criteria and report which are satisfied, violated, or inapplicable. We introduce the VFM, a tree-attention verifier that encodes the shared context and target text, and scores N natural language criteria in parallel, returning independent per-criterion ternary verdicts. The VFM is pretrained on CriteriaBank, an open corpus of 357K (context, target, criterion, verdict) tuples spanning over 285K distinct criterion strings. A finetuned VFM-4B reaches 88.3% accuracy on a synthetic governance benchmark, surpassing larger generative judges, and on real insurance-compliance calls it is competitive with GPT-5.4 with synthetic finetuning alone. On the DynaBench multi-criterion benchmark, a finetuned VFM-4B reaches 0.85 F1 on trace-level failure detection versus DynaGuard-4B's 0.72; the VFM additionally localizes the violated rule as its top-1 prediction in 98.4% of failing traces, and a top-3 cascade to a 26B Gemma 4 chain-of-thought judge reaches 0.94 F1. CriteriaBank pretraining improves cross-domain transfer on four additional benchmarks, particularly in the low-data regime. On an H100, VFM-4B evaluates 1000 criteria against a 4K-token context in under 2s at <10GB VRAM, making it suitable for runtime verification. We will release CriteriaBank, trained checkpoints, and evaluation code.
PaperID: 5804, Poster
Abstract: Discrete diffusion models achieve strong performance in text and image generation, but their inference remains slow and must inherently balance sampling efficiency and sample quality. In this work, we present a systematic study of how the \emphdegree of stochasticity in Markov transitions governs the tradeoff. We show that highly deterministic transitions converge rapidly but suffer from error accumulation, while more stochastic transitions converge more slowly yet can achieve lower final error. Using an information-theoretic analysis, we identify the underlying mechanism as an error correcting effect induced by \emphredundant transitions that symmetrically exchange mass between states, and show that these transitions can provably contract sampling errors. Motivated by this analysis, we propose \emphDiscrete Churn and Restart Sampling (DCRS), a novel inference algorithm that injects controlled stochasticity by alternating between forward and reverse diffusion processes. Experiments on large-scale image and language benchmarks show that DCRS improves the speed-quality tradeoff in the low number of function evaluations regime. On image datasets, DCRS achieves up to a 10× reduction in sampling steps compared to standard samplers while maintaining competitive sample quality, whereas on language benchmarks, we observe more nuanced behavior depending on the corruption process and sampling procedure.
PaperID: 5805, Poster
Authors: Ting-Yu Dai, Takuya Kurihana, Wing Yee Au, YUNG WONG
Abstract: Demand-side flexibility, i.e., forecasting, shifting, and curtailing residential energy loads, depends on thermal models trusted across millions of heterogeneous buildings. Existing tools force a hard tradeoff: high-fidelity physics simulators such as EnergyPlus are accurate but sequential and require per-building calibration, while purely data-driven sequence models scale but abandon the physical structure that makes their predictions trustworthy. We introduce NeuralBES (Building Energy Simulation), a differentiable emulator that resolves this tradeoff by parameterizing a resistance--capacitance (RC) based thermal model with a shared neural encoder: static building metadata such as floor area, vintage, and HVAC type is mapped to physically bounded capacitances, conductances, and equipment coefficients, which become the coefficients of a scalar linear recurrence solved via a log-space parallel scan, and a predictor--corrector loop closes the thermostat--temperature nonlinearity while preserving full-horizon gradient flow. Trained on the ResStock dataset across three climate zones, NeuralBES handles heterogeneous building archetypes, vintages, and climate zones within a single trained encoder, while black-box baselines produce statistically plausible but physically inconsistent trajectories. On the annual full-year rollout, NeuralBES is the only data-conditioned model that is simultaneously physics-valid and accurate to within 4 MAPE points of the strongest raw-error baseline, while operating at roughly an order of magnitude fewer parameters than the transformer and recurrent baselines; among physics-valid baselines at parameter parity it more than halves the MAPE of the grey-box RC alternative.
PaperID: 5806, Poster
Abstract: In in-context learning (ICL), a frozen pre-trained model solves tasks by conditioning on a prompt of a few input–output examples, without gradient updates. If the task was present in pretraining but the particular prompt sequence was not, the resulting in-distribution generalization is retrieval-based ICL. Learning-based ICL instead reflects out-of-distribution generalization: the model succeeds on prompts generated by a novel task. Empirically, both forms improve with scale. By analogy to benign overfitting in supervised learning, we call this in-context benign overfitting: larger models more faithfully memorize the pretraining tasks (improving retrieval ICL) while also generalizing better to novel tasks (improving learning ICL). We prove that this phenomenon already arises in a minimal in-context linear-regression feature-selection model. In contrast, standard in-context linear-regression models exhibit a retrieval–learning tradeoff, where the emergence of learning-based ICL coincides with degraded retrieval-based performance.
Abstract: In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance. Recent methods optimize mixture weights via proxy models, but they rely on the assumption of static data distributions. As a result, when the underlying data pool shifts, these methods require costly retraining from scratch. This limitation restricts their ability to scale seamlessly from small settings to larger data pools and model sizes. In this paper, we propose to address this limitation by casting data mixture optimization as a causal inference problem. We formulate the statistical features of the data pool as covariates and the domain mixture as the treatment. After fitting a causal model on 512 runs of to estimate the Conditional Average Treatment Effect (CATE), we extrapolate the optimal mixture for an 800K data pool and apply it to train a 7B model. Furthermore, we successfully generalize the framework to long chain-of-thought data on dynamically infers state-dependent optimal data mixtures. Extensive experiments show that the mixture guided by consistently improves performance across multiple downstream tasks, outperforming RegMix and other baselines. In addition, we use the CATE Interpreter to provide visual analysis of the learned mixing strategy. Overall,
Abstract: Random directed acyclic graphs (DAGs) based on imposing an order on Erdős–Rényi and scale free random graphs are widely used for evaluating causal discovery algorithms. We show that in such DAGs, the set of nodes reachable via open paths, termed relatives, increases monotonically along the causal order. We assess the prevalence of this pattern numerically, and demonstrate that it can be exploited for causal order recovery via sorting by the estimated number of relatives. We note that many simulations in the literature feature settings where this yields an excellent proxy for the causal order, and show that a strict increase of relatives along the causal order leads to a singular Markov equivalence class. We propose sampling time-series DAGs as a possible alternative and discuss implications for causal discovery algorithms and their evaluation on synthetic data.
Abstract: Exact solution of hard combinatorial optimization problems often relies on strong convex relaxations, but solving these relaxations repeatedly inside a branch-and-bound algorithm can be prohibitively expensive. Hence, we consider this challenge for \newMax-Cut problem (Max-Cut), where branch and bound commonly uses semidefinite programming (SDP) relaxations to bound subproblems. We propose a Max-Cut-specific \newgraph neural network that serves as a principled, lightweight neural proxy for these SDP solvers and can be plugged directly into an exact branch-and-bound framework. The proposed architecture has update steps of complexity \mathcalO(n^2 + ne), and predicts both primal- and dual-feasible SDP solutions. The primal SDP solutions yield feasible Max-Cut solutions via the Goemans--Williamson algorithm. In addition, it is trained in a self-supervised fashion without requiring solved SDP relaxations as labels. Empirically, we show that our architecture can substantially reduce the cost of bounding in exact Max-Cut solving by up to 10.6 × compared with using the state-of-the-art SDP solver Mosek. Our work highlights the potential of learned, validity-preserving surrogates for accelerating exact optimization over structured convex relaxations.
PaperID: 5810, Poster
Abstract: Data pruning for multivariate time-series forecasting (MTSF) is an emerging topic that poses unique challenges due to the intricate temporal dependencies and high-dimensional correlations between variables. The dominant data pruning scores have been designed for i.i.d. image classification tasks and rate each data sample using its own training trajectory in isolation. However, this misses a structural property of MTSF training: samples (obtained by windowing time series) can easily overlap, especially if strides down to 1 timestep are used, thus breaking the standard training i.i.d. assumption. Indeed, our key observation is that this overlap drives neighbors' loss trajectories into lockstep across training; the few samples whose trajectories deviate from this lockstep carry gradient information their neighbors do not share—exactly the samples that are worth pruning away. Hence, we propose TWIN (Trajectory-Wise Inconsistency with Neighbors), a multi-scale Wasserstein-2 score and pruning algorithm that, for the first time, compares each sample's sorted training-loss trajectory against those of its sliding-window neighbors. Specifically, our design is motivated by two complementary bounds: the pairwise discrepancy between neighboring loss trajectories is at most \mathcalO(\delta) when both samples are clean (Theorem 1), but at least \Omega(\Delta^2) once one carries a \Delta-spike contamination (Theorem 2). Whenever \Delta \gg \delta, these scaling orders separate, exposing an asymptotic scaling-order gap that TWIN's rank-based rule reads without estimating any bound constant. A 16,000-cell synthetic audit empirically confirms this gap, with a median 44.23× separation between clean and contaminated populations. We further evaluate TWIN's effectiveness on six MTSF benchmarks, four horizons, and five seeds against five baseline methods, showing that it shortens the average retraining time to approximately 48% of the unpruned baseline and reduces the average Relative MSE by up to 1.56% under IQR self-calibration, thus setting a new state of the art.
PaperID: 5811, Poster
Abstract: Dynamic Graph Neural Networks (DyGNNs) demonstrate powerful representation capabilities in crucial time-series systems by leveraging graph structure and temporal dynamics. However, existing dynamic graph models often suffer from unpredictable distribution shifts induced by complex time-varying factors such as emergency incidents, leading to severe performance degradation due to their low out-of-distribution (OOD) generalization. Graph invariant learning has been extensively studied to solve this problem through feature disentanglement in the time or spectral domain. Despite recent advances, existing methods suffer from critical limitations in practice: Temporal-based methods struggle to disentangle transient perturbations and stable patterns that are highly overlapping and entangled within the same period. While spectral-based methods attempt spectral-domain disentanglement via the global Fourier transform, these methods inherently lose temporal localization and fail to effectively characterize distribution shifts driven by time-varying factors. To this end, in this paper, we propose , a novel framework for resolving distribution shifts by conducting disentanglement within a wavelet coefficient domain endowed with time-frequency localization capabilities. Specifically, WavCIL consists of two key components: (i) a wavelet transform module, which leverages a set of wavelet bases to map the input temporal signals into wavelet coefficient domain, simultaneously capturing frequency components and localizing their temporal occurrences; and (ii) a coefficient domain invariant learning module, which utilizes an adaptive mask mechanism to disentangle stable causal patterns invariant to time-varying factors from spurious patterns driven by time-varying factors in the coefficient domain, guiding the prediction process to rely more on the causal patterns. Extensive experiments on several benchmark datasets demonstrate that WavCIL achieves state-of-the-art generalization performance when handling distribution shifts. Our code is released at https://anonymous.4open.science/r/WavCIL-4B8C.
PaperID: 5812, Poster
Authors:
Donghyun Kim, Jin Hong, seungmin Kim, Dain Kim, Junseok Kwon, Daeseon ChoiAbstract: Talking-head generation models, which synthesize realistic facial animations from audio, are increasingly vulnerable to misuse in multimodal deepfake scenarios. Protecting such systems remains challenging, as talking-head models fundamentally rely on precise audio–visual alignment, while effective audio protection methods remain largely underexplored. In this work, we propose MUTE, an audio protection framework that uncouples the underlying audio–visual alignment through a multi-level strategy. Specifically, MUTE combines (1) Representation-level Degradation, which perturbs temporal audio embeddings to degrade their representations, and (2) Alignment-level Disruption, which directly perturbs cross-modal attention to disrupt structural alignment. To improve robustness, perturbations are constrained in the STFT domain to high-energy regions, making them resistant to post-processing and denoising. An optional speaker-level objective further mitigates potential bypass via TTS-based resynthesis. Extensive experiments demonstrate that MUTE consistently degrades lip synchronization across both white-box and black-box talking-head models while preserving perceptual audio quality. The proposed method remains effective under real-world transformations and can be combined with image-based approaches to provide stronger multimodal protection.
PaperID: 5813, Poster
Abstract: Deep learning inference demand is projected to grow 10,000× over the next five years, a trajectory that general-purpose accelerators cannot match. Custom accelerators offer a viable path forward, but designing them requires simultaneous co-optimization of hardware datapaths, software schedules, and compiler transforms across a combinatorial search space exceeding \mathcal O(10^2300). While co-design is becoming increasingly essential, existing methodologies rely on decoupled, sequential pipelines that miss critical cross-stage trade-offs. We introduce CHEETAH: a globally optimized, AI-driven framework for inference accelerator co-design. To our knowledge, this is the first co-design framework to combine joint tiling-fusion, dynamic placement, and rematerialization in a single LP-based formulation. Key contributions include: (1) time-indexed variables for dynamic memory management, (2) rematerialization variables to enable cheap recomputation over costly DRAM reloads, and (3) joint tiling-fusion optimization via McCormick linearization. This inner-loop precision is steered by an LLM-guided outer loop that performs causal interventions based on specific deployment profiles and Lagrangian sensitivity signals. In less than 6 hours, \textscCHEETAH discovers designs on the Pareto frontier achieving an average of ~22.07× raw speedup across 11 workloads, reaching up to ~40.1× speedup on a single workload over a TPU-v3 baseline, unlocking efficiency inaccessible to decoupled sequential pipelines.
PaperID: 5814, Poster
Authors: Eadom T Dessalene, Michael Maynord, Amir Hossein Shahidzadeh, Botao He, Yasamin Tabatabaee, Ruoshi Liu, Ruohan Gao, Paul Liang, Yiannis Aloimonos
Abstract: Computer vision models learn hand action only from its visual consequences, without access to the underlying embodied motor signals. We address this absence with MotorSense - the first large-scale paired egocentric video and surface electromyography (EMG) dataset. MotorSense pairs bimanual EMG bands with Meta Aria glasses and yields synchronized video, bimanual EMG and IMU, and audio across 25 hours of unscripted manipulation. Beyond scale, EMG offers an alternative to manual annotation, and captures finger movement, hand configuration, and forces exerted between hand and environment --- this is in contrast to manual annotation which is costly, subjective, and unreliable under occlusion. EMG provides a dense hand signal free of both occlusion and labeler bias. To bridge the representational gap between low-level EMG and high-level semantic action representations, we learn a mid-level representation over the EMG data that we call motor codes. We utilize these motor codes in two settings: 1) motor-code pre-training, in which image and video backbones are pre-trained to predict motor codes from pixels, and subsequently evaluated over downstream tasks of 3D hand-object reconstruction and action recognition; 2) motor-code conditioning, in which the codes are used as control inputs for a video diffusion model, yielding videos of fine-grained hand manipulation. On action recognition (EPIC-Kitchens, Ego-Exo4D, and Meccano), under a frozen-evaluation protocol we observe a large 35.8% mean relative improvement in top-1 accuracy over conventional pre-trained backbones. On ARCTIC 3D hand-object reconstruction we observe a sizeable increase in object-pose success rate from 72.90% to 77.22% compared to the pre-training of the base ARCTIC model. Code and data are made anonymously available at https://anonymous.4open.science/r/MotorSense-7173/ .
PaperID: 5815, Poster
Abstract: Existing MLLM-based UI-to-Code generation methods can generate complete HTML/CSS codes through end-to-end or multi-stage pipelines, yet the generated codes often contain structural errors such as incorrect repeated layouts, wrong component boundaries, or missing interaction slots. Subsequent generations may continue along the erroneous structure, causing the error to propagate and amplify. This paper studies post-generation repair: how can we locate and repair unreliable local DOM structures? We observe that this repair process is ambiguous in two ways: a mismatch between the target screenshot and the output may correspond to multiple plausible DOM subtrees, and a single DOM subtree may admit multiple repairs with different rendered behaviors. We propose Render Structure Uncertainty (RSU)for partial DOM repair, a browser grounded uncertainty layer that can be attached to any base HTML generator. RSU parses the target screenshot into a visual structure graph and converts the browser execution trace of the generated code into a DOM-anchored render structure graph. By matching these graphs, RSU yields structural mismatches and a small set of candidate DOM holes, preserving error attribution uncertainty. For each hole, RSU generates multiple local candidates, renders them in the full page context, and clusters them using render-structure signatures that encode layout relations, repeated patterns, content slots, semantic roles, and invalid states, thereby modeling repair uncertainty. The retained representative candidates form a compact Partial DOM candidate set, over which we perform global reranking to produce the final HTML code. In this way, RSU reframes HTML repair into uncertainty-guided search in render-structure space, avoiding premature commitment to either a single error location or a single local repair. Experiments show that RSU improves the generation quality of base models, producing higher-quality HTML with better layout consistency and fewer local structural errors.
Abstract: Evaluation of multimodal reasoning models is typically reduced to a single accuracy score, implicitly treating reasoning as a unitary capability. We introduce MathLens, a benchmark of textbook-style geometry problems that exposes this assumption by operationally decomposing performance into perception, reasoning, and multimodal-specific components. Each problem is derived from a symbolic specification and accompanied by visual diagrams, text-only variants, multimodal questions, and targeted perceptual probes, enabling controlled measurement of each component. Using this decomposition, we show that common training strategies induce systematically different capability profiles that are invisible under aggregate accuracy. Reinforcement learning primarily improves perceptual grounding and robustness to diagram variation, while textual SFT yields gains through reflective reasoning. In contrast, as perception and reasoning improve, a growing fraction of remaining errors fall outside these components and are categorized as multimodal-specific. These results suggest that apparent progress in multimodal reasoning reflects shifting balances among subskills rather than uniform advancement, motivating evaluation beyond scalar accuracy.
Abstract: Currently, there is a gap in the field of ultra-high-definition (UHD) video dehazing due to the lack of a benchmark for evaluation. Furthermore, existing video dehazing methods cannot run on consumer-grade GPUs when processing continuous UHD sequences of 3--5 frames at a time. In this paper, we address both issues with a new benchmark and an efficient method. Our key observation is that atmospheric dehazing reduces to a per-pixel affine transform governed by the low-frequency depth field, which can be compactly encoded in bilateral grids whose prediction cost is decoupled from the output resolution. Building on this, we propose LiBrA-Net, which factorizes the spatiotemporal affine field into a spatial--color and a temporal bilateral sub-grid predicted at a fixed low resolution, fuses their coefficients in the \mathfrakgl(3) Lie algebra under group-theoretic regularization, maps the result to invertible GL(3) transforms via a Cayley parameterization, and restores high-frequency detail through a lightweight input-guided branch. We further release UHV-4K, the first paired 4K video dehazing benchmark with depth, transmission, and optical-flow annotations on every frame. Across UHV-4K, REVIDE, and HazeWorld, LiBrA-Net sets a new state of the art among compared video dehazing methods while running native 4K at 25\,FPS on a single GPU with only 6.12\,M parameters. Code and data are available at \urlhttps://anonymous.4open.science/r/LiBrA-Net-42B8.
PaperID: 5818, Poster
Abstract: While Speculative Decoding (SD) has become an essential lossless acceleration technique for Large Language Models, its direct application to Multimodal Large Language Models (MLLMs) is hindered by the intricate visual dependencies present in cross-modal generation. Existing SD methods enforce a uniform global alignment between the draft and target models, which proves ineffective in multimodal settings due to a structured distribution bias: deviations of the draft model are systematically concentrated on a sparse subset of tokens that demand fine-grained visual comprehension. To overcome this limitation, we reformulate the training of multimodal draft models as an acceptance-oriented critical-token alignment problem and introduce CATS, a novel two-stage training framework. CATS employs two complementary selection mechanisms to pinpoint critical tokens:hard tokens that exhibit high rejection probabilities, and vision-critical tokens whose prediction fundamentally depends on visual semantics. Following an initial global alignment phase, CATS performs sparse, targeted refinement exclusively on the union of these identified token sets. Extensive experiments across three representative benchmarks using multiple LLaVA and Qwen2.5-VL target models demonstrate that our approach consistently improves the token acceptance rate over strong baselines such as EAGLE-2 and supervised fine-tuning. Ultimately, CATS substantially accelerates MLLM inference, achieving speedups of up to 3.05×.
Abstract: Safety interventions on dual-use knowledge typically choose between destroying hazardous content (e.g., unlearning, filtering) and suppressing it at the output layer (e.g., refusal training); both pay a tax in adjacent-domain competence or over-refusal. We argue that the right operation is conditioning, not reduction: we show that hazardous knowledge can be retained in the model and behaviorally gated by a privileged control token. Our method, Token Inoculation, introduces a binding-and-branching approach. First, during continued pre-training, we mark hazardous contents by inserting a special token alongside dual-use documents, so the model binds the marker to the underlying semantics of the hazardous domain. Second, during supervised fine-tuning, we teach the model to answer hazardous queries correctly when the special token is present and to refuse them when it is absent, thereby enabling selective refusal without removing dual-use knowledge. On hazardous domain (e.g., WMDP-Bio), Token Inoculation reduces accuracy from 79% to 18% while retaining 93% of the base-model's benign-domain performance (e.g., MMLU), achieving the best safety-utility trade-off against unlearning and refusal-tuning baselines across 1B–14B model scales. We further show that refusal selectivity is controllable through the quality of the conditioning signal and that domain-specific semantic binding during pre-training is critical for the conditional behavior to generalize beyond memorized triggers. Our results suggest that safety alignment is better cast as a conditioning problem than a forgetting one: behavioral control is more precise when sensitive knowledge is retained under controlled access than when it is destroyed. We further show that refusal selectivity is controllable through the quality of the conditioning signal and that domain-specific semantic binding during pre-training is critical for the conditional behavior to generalize beyond memorized triggers. Our results suggest that safety alignment is better cast as a conditioning problem than a forgetting one: behavioral control is more precise when sensitive knowledge is retained under controlled access than when it is destroyed.
Abstract: Reinforcement learning has achieved remarkable success in robot learning. However, under challenging exploration and contact-rich dynamics, early-stage training is frequently dominated by premature terminations such as collisions and falls. As a result, learning is overwhelmed by short-horizon, low-return trajectories, which hinder convergence and limit long-horizon exploration. To alleviate this issue, we propose a technique called Failure Episodic Memory Alert (FEMA). FEMA explicitly stores short-horizon failure experiences through an episodic memory module. During interactions, it retrieves similar failure experiences and prevents the robot from recurrently relapsing into unstable states, guiding the policy toward long-horizon trajectories with greater long-term value. FEMA can be combined easily with model-free reinforcement learning algorithms, and yields a substantial sample-efficiency improvement of 33.11% on MuJoCo tasks across several classical RL algorithms. Furthermore, integrating FEMA into a parallelized PPO training pipeline demonstrates its effectiveness on a real-world bipedal robot task.
PaperID: 5821, Poster
Abstract: Efficient exploration in sparse-reward environments, where informative extrinsic feedback is scarce, remains a significant challenge in reinforcement learning. A key limitation of existing exploration strategies is that they often assume densely coupled environment dynamics, overlooking the underlying causal mechanisms of the environment, particularly the fact that causal relationships may vary across contexts. To address this issue, we propose contextual causal dynamics model-based intrinsic motivation (CIM), a causality-aware exploration framework that explicitly captures context-dependent sparse causal structures and derives intrinsic motivation signals, specifically causal action influence and curiosity. These intrinsic rewards encourage interventions on causally relevant state factors rather than merely promoting observation novelty-seeking behavior, thereby facilitating robust exploration and active skill acquisition. Moreover, by learning contextual causal mechanisms, CIM identifies and filters out task-irrelevant state variables, improving generalization under distribution shifts. Empirical results demonstrate that our method achieves superior sample efficiency, training robustness, and out-of-distribution generalization performance across multiple robotic manipulation tasks.
Abstract: We study nonstationary online linear programming (OLP), where orders arrive sequentially and follow independent but not-identical reward-resource distributions. The decision- maker seeks to maximize the expected total reward by making immediate and irrevocable acceptance or rejection decisions for each order, subject to a resource endowment available at the beginning of the planning horizon. Such problems arise in resource-constrained ML and AI systems, including inference admission control, online advertising and recommendation, and data-acquisition pipelines, where the mix and value of requests may change over time. We focus on a minimal-information regime in which the decision-maker observes only one independent sample from each future distribution before the horizon begins. We propose a novel re-solving algorithm that integrates a dynamic programming perspective with the dual-based frameworks traditionally employed in stationary environments. In the large-resource regime, where the resource endowment scales linearly with the number of order, we prove that our algorithm achieves O((\log n)^2) regret across a broad class of nonstationary distribution sequences. Our results demonstrate that polylogarithmic regret is attainable even under significant environmental shifts and minimal data availability, bridging the gap between stationary OLP and more volatile real-world resource allocation problems.
Abstract: The default representation (DR), originally introduced in neuroscience, and its principal eigenvector have been shown to be effective for a wide variety of applications, including reward shaping, count-based exploration, option discovery, and transfer. However, in prior investigations, the eigenvectors of the DR were computed by first approximating the DR matrix, and then performing an eigendecomposition. This procedure is computationally expensive and does not scale to high-dimensional spaces. In this paper, we propose an objective for approximating the principal eigenvector of the DR using only transitions with a neural network. This objective is inspired by a series of theoretical results and is empirically validated in a number of environments. We then demonstrate its usefulness by applying the learned eigenvectors for reward shaping.
PaperID: 5824, Poster
Abstract: Offline goal-conditioned reinforcement learning (GCRL) enables goal-reaching from static datasets, but the learned value, critic, or compatibility score must support both long-horizon reachability inference and offline policy extraction. We study iterative latent refinement as an architecture-level replacement for the feedforward value or critic backbone used by existing offline GCRL algorithms. Instead of emitting a score after a single feedforward computation, the network repeatedly updates a latent representation using shared recurrent update weights and lightweight step-specific parameters before exposing the final actor-facing signal. On the recent benchmark datasets reported in this study, recurrent variants improve performance across several algorithms without changing their losses, data, actors, or training protocol; across the 21 reported algorithm-dataset rows, the mean improvement is 10.7 percentage points, with the largest gains on stitching and bottleneck maze tasks. CRL diagnostics are consistent with an improved policy-extraction interface: recurrent critics increase normalized separation between matched and mismatched goals and produce actors that remain closer to dataset actions under several behavior-support proxies. Depth and update-capacity studies further suggest that the effect depends on how computation is allocated between recurrent depth and per-step expressivity.
Abstract: Small-molecule foundation models are typically pretrained on standalone molecular data, unlike vision and language models that often benefit from cross-modal or relational supervision. Protein-ligand co-folding provides a molecular analogue of such supervision by exposing models to atom-level ligand-protein interactions, raising the question of whether co-folding models can yield strong small-molecule representations. We study this question using Boltz2, a modern co-folding model, by transferring its atom-level ligand representations to standalone small-molecule tasks. Through systematic probing and distillation, we show that Boltz2 representations match or outperform existing models on the ADMET benchmark, accelerate molecular generative modeling, and improve sample efficiency in structure-guided ligand optimization. We further find that Boltz2 representations are complementary to those learned from conventional standalone molecular supervision, including 3D conformers, bioassay labels, and quantum-chemical properties. Finally, we extend representation alignment to reinforcement learning, showing that dense representation-level supervision can complement scalar rewards in molecular discovery. These results identify protein-ligand co-folding as a promising pretraining paradigm for small-molecule representation learning and position Boltz2 as a strong, off-the-shelf molecular foundation model.
PaperID: 5826, Poster
Abstract: Embedding scaling has recently emerged as a promising direction for increasing the capacity of language models through embedding-level representations, rather than solely expanding transformer depth or width. However, existing embedding scaling methods are primarily designed for discrete text tokens, leaving it unclear whether similar mechanisms can be extended to vision-language models, where visual representations are continuous, high-dimensional, and spatially structured. In this paper, we introduce V-LUMEN, a visual lookup memory framework for embedding-level scaling in vision-language models. V-LUMEN augments visual-language representations with reusable visual embeddings retrieved from an external memory indexed by discretized visual patterns, thereby expanding accessible representational capacity without directly increasing the dense computation of the backbone model. To address the unique challenges of visual representations, V-LUMEN introduces spatial aggregation for constructing structured visual memory keys and text-conditioned hashing for retrieving task-relevant visual memories based on the input query. Together, these components enable memory retrieval that is both spatially aware and instruction-conditioned. Experiments across diverse reasoning-intensive multimodal benchmarks show that V-LUMEN effectively augments vision-language models through external visual memory. Without backbone adaptation, V-LUMEN substantially improves the vanilla Qwen3-VL-2B-Instruct model, and further achieves competitive or better performance against parameter-matched MoE and LoRA baselines, demonstrating the potential of visual embedding scaling as an alternative scaling axis for vision-language models.
PaperID: 5827, Poster
Abstract: Diffusion models generate samples through an iterative denoising process guided by a pretrained neural network. Once the denoiser is fixed, the sampling algorithm itself (noise schedules, guidance scales, stochasticity profiles) still requires careful tuning, a process typically carried out through costly empirical grid search. In this work, we introduce an inverse reinforcement learning framework for learning sampling strategies without retraining the denoiser. We formulate the diffusion sampling procedure as a discrete-time finite-horizon Markov Decision Process, where actions correspond to optional modifications of the sampling dynamics. To optimize action scheduling, we avoid defining an explicit reward function and instead directly match the target behavior expected from the sampler using policy gradient techniques. We provide experimental evidence that this approach matches fine-tuned samplers and comes at a modest cost compared to grid search: on ImageNet-64, a single training run replaces exhaustive search at up to 9× lower cost, with only 16% overhead at inference.
PaperID: 5828, Poster
Authors: Hongjian Zou, Yidan Wang, dingqi, Yixuan Liao, xiaoxin chen
Abstract: Large language models can achieve strong benchmark scores without proportional gains in broader capability, but diagnosing when this occurs remains difficult. We study this gap through benchmark shadows: support-concentrated data regimes that may concentrate learning around narrow evaluation-relevant patterns while limiting broader representational development. Under fixed architecture, tokenizer, optimizer family, and training budget, we compare a coverage-expanding baseline with two support-concentrated regimes: redundant repetition and frequency-concentrated rewriting. Spectral, rank-based, and layer-wise update diagnostics reveal distinct parameter-space signatures: repetition-concentrated training is largely recoverable after later diverse training, whereas frequency-concentrated support collapse leaves more persistent footprints. Correlational analyses of open-source multimodal model families reveal analogous structural patterns that co-occur with asymmetric benchmark profiles, while a prompt de-duplication case study shows that surface redundancy alone does not induce the same regime-level effects. Benchmark scores alone are therefore insufficient to characterize capability; parameter-space diagnostics provide complementary signals about training quality, coverage, and generalization-related regime effects.
PaperID: 5829, Poster
Abstract: 3D scene graphs provide a structured representation of complex environments by encoding objects, their semantic attributes, and the spatial and functional relationships between them. Current approaches for 3D scene graph generation suffer from several fundamental limitations. They rely on complex multi-stage pipelines with explicit intermediate representations, making systems fragile and prone to error propagation. They assume access to ground-truth object annotations during inference, which deviates from real-world scenarios. They depend on proprietary models, hindering open-source deployment, or incur prohibitively slow inference. We present GraphWrit3R, a simple end-to-end method that takes a 3D point cloud, Gaussian Splats, or a combination of both as input, and directly outputs a complete scene graph as a structured JSON script. The graph lists all objects, their semantic attributes, and the relationships between them, while avoiding all of the above mentioned limitations. The choice of multiple input modalities is purely for versatility, allowing a single set of weights to handle diverse scenarios. Point cloud inputs are encoded via Sonata and Gaussian Splat inputs via Chorus, with both modalities projected onto a shared voxel grid and fused through a novel per-voxel contrastive alignment loss before being decoded by a large language model. As a natural consequence of the LLM, GraphWrit3R also supports open-vocabulary querying. On the 3DSSG benchmark, our method achieves state-of-the-art performance on object class, predicate, and triplet recall, outperforming methods that rely on ground-truth object annotations during inference. We further provide qualitative results and analyze different input modality configurations, contrastive loss formulations, and token fusion strategies. Our code and model will be released upon publication.
PaperID: 5830, Poster
Abstract: Latent variable models based on density operators, the mathematical foundation of quantum mechanics, remain far less capable than their probabilistic counterparts. This is primarily due to the lack of an Expectation-Maximization (EM) framework for learning density operator models, since the absence of conditional probability for density operators prevents a direct extension of the classical derivation. This paper addresses this challenge by proposing Density Operator Expectation Maximization (DO-EM), using tenets from quantum information theory. We first derive an operator-theoretic analog of the evidence lower bound from the data processing inequality and study its maximization as a projection onto the data manifold. We solve this information projection problem while also improving existing information recovery results. The DO-EM algorithm guarantees monotone likelihood ascent for a rich class of models and is simpler than direct likelihood gradient ascent. The Expectation step recovers the Petz Recovery Map when the model class can be aligned with the target operator. We specialize these models to Classical-Quantum LVMs which scale to standard image datasets via a per-datapoint decomposition of the variational bound and outperform their classical counterparts under matched computational resources.
PaperID: 5831, Poster
Abstract: Face recognition has achieved strong performance, but further gains often require much larger labeled datasets, making annotation costly. This motivates the use of unlabeled data, where face clustering is important for applications such as annotation and recognition. Recent GCN-based methods have shown promising results on affinity graphs, but balancing clustering quality and scalability on large-scale data remains challenging. To solve this problem, we propose FC-DGCN, a face clustering framework based on deep residual GCNs. The proposed method formulates clustering on an affinity graph through two complementary tasks: node confidence estimation and edge connectivity estimation. Specifically, we develop DGCN-V to estimate node confidence on the global graph and DGCN-E to estimate edge connectivity on local subgraphs. To better characterize the local cluster structure, we further revise the confidence-estimation target. Based on these two modules, FC-DGCN links each node to a higher-confidence neighbor with strong predicted connectivity, which naturally induces identity-consistent clusters. Experiments on MS-Celeb-1M show that FC-DGCN outperforms competitive baselines in Pairwise F-score and BCubed F-score while remaining computationally practical. Moreover, the clusters generated by FC-DGCN provide effective pseudo labels for low-label face recognition, leading to substantial improvements on MegaFace and IJB-A.
Abstract: Vision-language-action (VLA) policies have shown strong potential for general-purpose manipulation, yet they often fail on novel, out-of-distribution objects whose appearance or geometry deviates from the training distribution. The standard remedy is to collect multi-view teleoperation data for every failure case, but this scales poorly in both cost and time. We introduce Pose6DAug, a failure-driven data augmentation framework that turns a policy’s own successful episodes into targeted demonstrations for its failure modes, without any new data collection. Our key insight is that each successful episode already encodes a physically valid action trajectory together with calibrated multi-view observations. By swapping only the manipulated object while preserving this trajectory, we obtain new and physically grounded demonstrations. However, the challenge lies in synthesizing object swaps, e.g., naive 2D video editing breaks multi-view consistency and physical plausibility, particularly under heavy occlusion and egocentric viewpoints. Our method instead operates directly in 3D, anchoring the target object with an explicit mesh driven by a temporally coherent 6D pose trajectory. This ensures geometrically consistent renderings across all camera views. Fine-tuning a VLA on data augmented by our method improves success rates by 16.5% relative to the state-of-the-art baseline on novel objects, while preserving in-distribution performance. These results show that multi-view and physically consistent augmentation is a practical path to scalable VLA generalization.
PaperID: 5833, Poster
Abstract: Training-free diffusion priors are powerful for inverse problems, but measurement guidance during reverse sampling is local: aggressive updates can improve immediate data fit while disrupting later denoising. We introduce a tethered predictive-inertial correction rule that separates proposal generation from acceptance. At each corrected step, the denoiser prediction anchors a frozen clean-space objective combining measurement fit with a noise-level-dependent tether. Heavy-ball dynamics with diffusion-scale predictive smoothing generate clean-state candidates; a separate verifier evaluates a finite dyadic set using the original unsmoothed objective, with the denoiser prediction retained as fallback. The accepted clean state is re-noised to continue sampling. The rule applies to pixel and latent diffusion with differentiable measurement operators, requires no retraining, and adds no denoiser calls inside the correction loop. We analyze when smoothing is inactive, how it changes nonlinear or decoder-composed proposal landscapes, and how objective verification yields local nonincrease, scale robustness, and displacement control. Across natural-image and accelerated MRI benchmarks, the method achieves competitive reconstruction quality with favorable speed--quality trade--offs, often reducing runtime relative to optimization-heavy diffusion solvers. On fastMRI knee reconstruction, it attains the strongest PSNR among compared methods at both acceleration factors.
PaperID: 5834, Poster
Abstract: Conditional message passing (CMP) currently represents a leading paradigm for knowledge graph reasoning. Since vanilla CMP performs full-graph propagation per query, existing methods adopt subgraph-wise sampling to improve computational efficiency. However, these extracted subgraphs still exhibit high structural overlap. Consequently, the same local structures are repeatedly accessed and processed, leading to substantial redundancy that remains a major bottleneck for scaling reasoning to large knowledge graphs. To address this issue, we propose reasoning over Coupled Receptive Fields (CoRF) to eliminate subgraph redundancy at scale. Rather than sampling a separate subgraph for each query, CoRF constructs a set of receptive supports that allow multiple queries to share receptive fields, thereby reducing subgraph redundancy. Meanwhile, we introduce boundary padding and residual compensation mechanisms to ensure topological integrity and query coverage. We further incorporate cost-aware load balancing to improve distributed execution efficiency. Comprehensive experiments on state-of-the-art CMP models demonstrate that CoRF achieves up to a 5.9× speedup and a 10.0× reduction in memory usage, without compromising reasoning accuracy.
Abstract: Optical Music Recognition (OMR), the task of transcribing sheet music into a structured textual representation, is currently bottlenecked by a lack of large-scale, annotated datasets of real scans. This forces models to rely on either few-shot transfer or synthetic training pipelines that remain overly simplistic. A secondary challenge is encoding non-uniqueness: in the popular Humdrum kern format for transcribing music, multiple different text encodings can render into the same visual sheet music. This one-to-many mapping creates a harder learning task and introduces high uncertainty during decoding. We propose Transcoda, an OMR system built on (i) an advanced synthetic data generation pipeline, (ii) a normalization of the kern encoding to enforce a unique normal form and (iii) grammar-based decoding to ensure the syntactic correctness of the output. This approach allows us to train a compact 59M-parameter model in just 6 hours on a single GPU that outperforms billion-parameter baselines. Transcoda achieves the best score among state of the art baselines on a newly curated benchmark of synthetically rendered scores at 18.46% OMR-NED (compared to 43.91% for the next-best system, Legato) and reduces the error rate on historical Polish scans to 63.97% OMR-NED (down from 80.16% for SMT++). We will make our model, data and generation pipeline publicly available upon publication of the paper.
PaperID: 5836, Poster
Abstract: Vision-Language Models like CLIP exhibit remarkable zero-shot capabilities but remain vulnerable to adversarial attacks. Existing defenses, such as adversarial fine-tuning or test-time defense, either incur high computational costs that lead to cause catastrophic forgetting, or struggle against strong adversarial attacks. In this work, we reveal that adversarial perturbations are highly overfitted to the decision boundary of a single canonical text prompt. By introducing several semantic prompt variants, we identify a Prompt Sensitivity Gap: adversarial examples exhibit significantly higher prediction variance across prompts compared to benign images. Motivated by this insight, we propose Prompt Ensemble Image Purification (PEIP), an efficient test-time defense framework. PEIP features a dual-objective purification loop that jointly suppresses prediction variance to dismantle adversarial alignment and reinforces the most responsive class across prompt variants to facilitate semantic recovery. Furthermore, to accelerate inference, we introduce a statistically calibrated Early Exit mechanism that bypasses benign images based on their initial prompt sensitivity. Extensive experiments across 16 classification benchmarks, multiple CLIP architectures, and CLIP-based zero-shot semantic segmentation tasks demonstrate that PEIP achieves state-of-the-art adversarial robustness while preserving exact zero-shot accuracy and high efficiency.
PaperID: 5837, Poster
Authors: Tugdual Kerjan, Rasmus Høier, Benjamin Scellier
Abstract: Equilibrium Propagation (EP) is a training framework for energy-based models (EBMs) that has attracted interest in the context of neuromorphic computing platforms such as continuous Hopfield networks, nonlinear resistive networks and coupled phase oscillators. However, EP's practical applications have so far remained limited to relatively small-scale problems. Predictive coding networks (PCNs), another class of EBMs rooted in computational neuroscience, are typically trained with a specialized algorithm and have likewise not yet been demonstrated at large scale. In this work, we develop a more effective training method for PCNs which combines the centered variant of EP with a novel equilibration scheme for PCNs. Using this approach, we train a 10-layer convolutional PCN (VGG10) on full-size ImageNet, achieving 13.23% test error rate on the top-5 classification task, close to the 12.2% backpropagation baseline. To our knowledge, this is the first demonstration of both PCNs and EP-based training at ImageNet scale. These results significantly extend the scalability of both approaches and suggest that the primary challenges in scaling EP in other EBMs may not be attributed to inherent limitations of the EP framework.
PaperID: 5838, Poster
Authors: Qi Wen, Jiezhou He, Zhiming Luo, Shaozi Li
Abstract: Few-shot medical image segmentation has shown great potential in reducing annotation costs for clinical applications. Recently, many few-shot methods have explored the Segment Anything Model (SAM) for training-free medical image segmentation via prompt engineering. However, existing approaches mainly focus on locating positive prompts from support images while overlooking informative background cues, making them prone to over-segmentation in anatomically similar regions. Moreover, simply introducing negative prompts cannot effectively suppress boundary leakage, and may even interfere with SAM’s mask decoding process, resulting in worse performance than using positive prompts alone. To address this limitation, we propose a training-free support-query framework, termed Boundary-Aware Prompt Mining, which introduces boundary-aware negative prompting for SAM-based few-shot medical image segmentation. Specifically, we introduce a Prototype-guided Positive Prompt strategy, which adopts a multi-center prompting mechanism to construct multiple foreground prototypes, enabling comprehensive spatial coverage in query images. Furthermore, we propose a Boundary-band Ambiguity Filtering strategy that identifies a boundary-adjacent background band and progressively removes ambiguous pixels and unreliable clusters highly similar to the foreground, enabling the selection of reliable negative prompts for suppressing false positives near foreground boundaries. The generated positive and negative prompts are jointly fed into SAM to produce refined segmentation results without additional training. Extensive experiments on Abd-MRI and Abd-CT datasets demonstrate that our method consistently outperforms existing approaches, highlighting its effectiveness and robustness under limited annotation settings. Code will be released upon acceptance.
PaperID: 5839, Poster
Abstract: Sparse attention accelerates long-context decoding by reading a subset of the KV cache, but existing methods are open-loop: the kernel commits to its selection with no signal indicating whether enough attention mass was preserved and no recovery path when it was not. The needed signal already exists inside the kernel. The fraction of softmax mass on the selected tokens, which we call coverage, satisfies an exact algebraic equality with per-step output error and serves a dual role as the acceptance criterion and the interpolation weight for recovery. Coverage-Verified Sparse Attention (CVSA) turns coverage into a training-free, closed-loop verify-then-correct decode loop that accepts high-coverage drafts and recovers only the heads that fall short, with no sampling, no draft model, and no speculative-decoding infrastructure. Turning verification on matters far more than tuning its threshold, and matched-budget comparisons confirm the gain is structural. On LongBench and RULER at 7B and 70B, CVSA closes the quality gap to dense attention to within statistical noise, with measured throughput reaching 6.21x at 128K context.
PaperID: 5840, Poster
Authors:
Junsheng Huang, Yifan Sun, Zhuoer Zhang, Ning Wei, Yening Liu, Michael Molter, Pavan Kumar Hanumolu, Elyse Rosenbaum, Bin Hu, Huan ZhangAbstract: Circuit design for analog integrated circuits (ICs) is challenging because circuit topology, device sizing, and circuit performance are tightly coupled. Existing large language model (LLM)-based methods either rely on inference-time correction around frozen models or train on limited circuit types without directly using simulation outcomes as training signals. A key bottleneck is that existing datasets do not provide sufficient task information for training LLMs on end-to-end specification-conditioned design. To fill this gap, we first construct a large-scale simulation-grounded dataset with 8,626 design tasks across 10 different types of analog circuits, where each task includes a textual description, labeled I/O ports, loading condition, target specifications, a testbench, and a reference netlist. Using this dataset, we present SAFTAC, a simulation-augmented fine-tuning framework that trains LLMs to generate and revise analog circuit design. SAFTAC combines three-stage supervised fine-tuning, which progressively builds the capabilities required for analog circuit design and feedback-guided correction, with simulation feedback-augmented reinforcement fine-tuning, which further optimizes models using simulation-grounded rewards. We evaluate SAFTAC under single-pass generation and two-pass generation with simulation feedback, showing improvements over the strongest closed-source LLM by 15.9% and 7.8%, respectively.
PaperID: 5841, Poster
Abstract: Large Vision-Language Models (LVLMs) have achieved impressive progress in multimodal reasoning, yet they remain prone to object hallucinations, generating descriptions of objects that are not present in the input image. In this work, we investigate hallucination from the perspective of attention-value dynamics inside LVLM vision encoders. We identify a consistent three-phase structure of visual processing---diffusion, focus, and rediffusion---and show that the focus phase is where attention most clearly separates strongly and weakly supported visual tokens. However, low attention does not necessarily imply negligible downstream influence: low-attention tokens in the focus phase can still exert non-negligible value-side influence on the attention output relative to their small attention mass. Through controlled phase-wise value interventions, we find that hallucination behavior is particularly sensitive to the value content of low-attention tokens during this phase. Replacing or neutralizing these values reduces hallucination metrics while largely preserving grounded object evidence. A token-level teacher-forcing analysis further shows that the intervention reduces the probability of hallucinated object tokens with much smaller effects on ground-truth object tokens. In addition, Visual Attention Ratio (VAR) analysis shows that focus-phase intervention is accompanied by increased attention to visual tokens during decoding. Based on these observations, we instantiate a simple training-free inference-time intervention that replaces focus-phase low-attention values with an image-level mean value vector using statistics from a single forward pass. Experiments across multiple LVLM backbones demonstrate that this analysis-derived intervention reduces object hallucination with negligible additional runtime and remains compatible with existing inference-time mitigation methods. Our project page is available at: https://anonymous.4open.science/w/FocusMatters-7B32/
Abstract: Text and 2D-conditioning interfaces provide weak, ambiguous control over spatial transformations in image editing -- particularly under large object motions and camera changes. Prior work has used 3D primitives such as boxes, but only as loose conditioning signals indicating approximate object location rather than specifying the transformation. We instead use 3D boxes as structured specifications: the user provides the input and output boxes of the edit, casting editing as a well-posed geometry problem. This ``thinking in boxes'' interface, where each box face is color-coded to convey 3D orientation, gives precise control over translation, rotation, scaling, and viewpoint changes in real images while preserving scene and object identity, and recovering previously unseen object regions. To ground transformations in scene appearance, we introduce a depth-aligned planar floor as a global reference frame, shaded with depth-aware cues. Conditioned on this structure, an image generator produces consistent results under large transformations. Trained in two stages -- on synthetic multi-object scenes and a small set of real-world videos from Objectron -- the system generalizes to complex, in-the-wild real images. Our method operates directly on real photographs and substantially outperforms recent state-of-the-art methods on large 3D edits.
PaperID: 5843, Poster
Abstract: Prompt-level instructions are the main way users give language models preferences, safety boundaries, and content rules. But a model can remember a constraint and still violate it when generating text. This points to a gap between two abilities that are often treated as one: retrieving a constraint from context, and enforcing it during token selection. Across five model families (0.5B--32B parameters), we find a clear behavioral split: constraints that admit an alternative generation plan are largely followed, while token-level prohibitions fail whenever the prohibited continuation is locally dominant. Mechanistic analyses show that retrieval heads still attend to the constraint at the violation step, but the resulting signal either paradoxically enhances the prohibited token's logit or suppresses it too weakly to change the argmax. Decoding-time interventions can eliminate surface-form violations, but semantic compliance still requires higher-level verification. These results establish that constraint retrieval and constraint enforcement are separable capabilities, and that reliable token-level compliance requires explicit veto mechanisms beyond prompting.
PaperID: 5844, Poster
Abstract: Autonomous AI research promises to accelerate the scientific progress of machine learning. While current Large Language Model (LLM)-based agents excel at writing code, the bottleneck in research is the exploration of diverse and novel ideas: agents collapse to common techniques present in their pretraining data, or prematurely converge on suboptimal solutions (Padmakumar et al., 2024; Jiang et al., 2025). To this end, we introduce Heuresis, a framework that abstracts the research pipeline into a set of general and composable primitives, enabling open-ended scientific exploration in machine learning research. We implement 5 known algorithms spanning Quality-Diversity, Evolutionary, and Curiosity-based search, in addition to a greedy baseline, and evaluate them across three axes -- Quality, Diversity, and Novelty -- on two domains: LLM pretraining and On-Policy RL. We find that quality and novelty are inversely correlated. Methods that optimize purely for raw performance reach the strongest single solutions on a given task by replicating prior work: the top-quality runs from the greedy baseline are uniformly classified as direct copies under the Gupta-Pruthi rubric (Gupta & Pruthi, 2025). The few high-quality, verified-novel ideas in our study come from algorithms that balance performance with diversity or curiosity-based exploration. We also observed that agents resorted to a variety of reward-hacking techniques during execution, whose detection was necessary to keep the search faithful to the task. Our results underscore the importance of search and Quality-Diversity methods for autonomous research, and our framework opens up opportunities for further inquiry towards the ultimate goal of perpetual, autonomous scientific progress.
PaperID: 5845, Poster
Authors:
Ran Luo, Haoxiang Deng, Lan Zhang, Mu YuanAbstract: Structured visual artifacts such as geometry diagrams, circuit schematics, and flowcharts are increasingly important targets for image generation, yet visual plausibility alone is insufficient---a single missing object, wrong relation, broken connection, or extra element can invalidate the output. Existing image generators remain brittle on such tasks, and generic step-by-step or agentic methods often fail because intermediate edits may corrupt previously correct structures and ignore generator-specific failure modes. To address this, we present ProGraf, a profile-guided framework for structured non-natural image generation. We first build a benchmark of 304 tasks across geometry, circuit, flowchart, and general diagrams, each annotated with fine-grained verification points. From repeated generator outputs, we derive model-specific capability profiles capturing category-level reliability, component-level weaknesses, and recurrent failure patterns. ProGraf uses these profiles to guide closed-loop decomposition, verifies each intermediate result, commits only accepted states, and replans after failures. On generator-specific hardest subsets, ProGraf recovers 43.7--77.8% of previously failed tasks across generators and categories, outperforming fixed-step and adapted agent baselines; ablation studies further show that benchmark-derived profiles are a key source of the gain.
PaperID: 5846, Poster
Abstract: Improving the reliability of large language model (LLM) agents in long-horizon decision-making remains a key challenge. When deployed as autonomous agents interacting with complex environments, early mistakes can propagate through trajectories and cause cascading failures. Recent approaches improve reliability by incorporating external critique or deliberation, but invoking these mechanisms at every step substantially increases token consumption and latency, limiting practical deployment. We propose SAG (Self-improving Agent with Gated critique), a cost-aware framework that formulates critique invocation as a step-wise decision problem during long-horizon interaction. SAG introduces a lightweight, training-free gating mechanism that estimates the utility of critique using action-level ambiguity signals—global entropy and local top-2 margin—computed over admissible actions. From a decision-theoretic perspective, this mechanism approximates the Value of Information (VoI) of critique, enabling the agent to selectively allocate expensive feedback only when its expected benefit justifies the cost. SAG further incorporates online bootstrapped self-improvement, allowing the actor to internalize critic-assisted behaviors and progressively reduce reliance on critique. Across three long-horizon interactive benchmarks and multiple backbone models, SAG substantially improves the success–cost trade-off compared with both no-critique and always-on critique agents. On ALFWorld, SAG increases task success from 24.6% to 78.3% while maintaining a token budget comparable to ReAct, yielding a 3.1× improvement in normalized token efficiency. Moreover, a 7B actor with a lightweight 3B critic achieves performance comparable to a 14B actor without critique, showing that selective critique can recover most of the reliability benefits of deliberation while dramatically reducing inference cost.
Authors:
Jianchang Su, Yifan Zhang, Shengkai Lin, Shizhen Zhao, YUSHENG ZHENG, Yiwei Yang, Wei ZhangAbstract: Multi-stage ML inference pipelines are difficult to autoscale because of heterogeneous resources, cross-stage coupling, and dynamic bottleneck migration. Production controllers ignore cross-stage effects, and learning-based autoscalers require offline training data that is unavailable for new pipelines. We present SAIR, an autoscaler for the zero-shot deployment regime that uses an LLM as an in-context reinforcement learning controller, improving its policy online from reward-labeled experience tuples in the LLM's context with no gradient updates. SAIR combines Pareto-dominance reward shaping with a provable separation margin between dominated and non-dominated actions, Bayesian-surprise experience retrieval, and a user-space CUDA interception layer that delivers continuous GPU rate control with sub-millisecond actuation overhead. We prove a regret bound that decomposes into retrieval coverage and LLM selection components, each mapping to a concrete SAIR mechanism. On four ML serving pipelines under three workload patterns, SAIR achieves the best or tied-best P99 latency in every configuration (up to 50% reduction over the next baseline); billable cost is reduced through replica consolidation, and effective cost (which captures GPU rate-control savings under co-location or MIG) by up to 97%. SAIR reaches 86% bottleneck detection accuracy without any offline training.
PaperID: 5848, Poster
Abstract: Monocular online 3D instance segmentation enables embodied agents to build object-level 3D understanding while continuously exploring their environments. Unlike RGB-D settings, monocular RGB streams lack direct depth and camera pose information, making geometry reconstruction necessary for 3D segmentation. Recent reconstruction foundation models (RFMs) have made this task more feasible by recovering 3D geometry and cross-view cues from monocular images. However, reconstructed geometry remains noisy and incomplete, and streaming observations provide only partial views of object instances, making local observations insufficient for stable instance representations and temporally consistent segmentation. We propose MonoChunk3D, an online 3D instance segmentation framework that incrementally reconstructs, aligns, and segments incoming monocular RGB chunks. Rather than treating reconstructed geometry merely as input, the framework adaptively integrates complementary RFM-derived cross-view priors with 3D geometric features to initialize current-chunk instance queries. To overcome the limitations of local observations, we maintain compact persistent instance states composed of explicit prototypes, implicit query embeddings, and spatial bounds to preserve long-term instance context. These states are propagated as history-aware queries and jointly decoded with current-chunk queries, allowing the preserved instance context to directly guide current-chunk segmentation. The resulting predictions are associated with historical instances to update the persistent states, enabling temporally consistent segmentation over streaming observations. Experiments on ScanNet200, ScanNetV2, and SceneNN show that MonoChunk3D achieves consistent improvements over existing methods in the monocular online setting.
Authors:
Yuanyi Wang, Yifan Yang, Su Lu, Yanggan Gu, Pengkai Wang, Wenjun Wang, Zhaoyi Yan, Congkai Xie, Jianmin Wu, Jialun Cao, Shing-Chi Cheung, Hongxia YangAbstract: Continual post-training aims to extend large language models (LLMs) with new knowledge, skills, and behaviors, yet it remains unclear when sequential updates enable capability transfer and when they cause catastrophic forgetting. Existing methods mitigate forgetting through sequential fine-tuning, replay, regularization, or model merging, but offer limited criteria for determining when incorporating new updates is beneficial or harmful. In this work, we study LLM continual post-training through three questions: What drives forgetting? When do sequentially acquired capabilities transfer or interfere? How can compatibility be used to control update integration? We address these questions through task geometry: we represent each post-training task by its parameter update and study the covariance geometry induced by the update. Our central finding is that: forgetting can be considered as a state-relative update-integration failure, it arises when the covariance geometries induced by tasks misalign with the geometry of the evolving model state. Sequential updates transfer when they remain compatible with the model state shaped by previous updates, and interfere when state-relative geometry conflict becomes high. Motivated by this finding, we propose Geometry-Conflict Wasserstein Merging (GCWM), a data-free update-integration method that constructs a shared Wasserstein metric via Gaussian Wasserstein barycenters and uses geometry conflict to gate geometry-aware correction. Across Qwen3 0.6B--14B on domain-continual and capability-continual settings, GCWM consistently outperforms data-free baselines, improving retention and final performance without replay data. These results identify geometry conflict as both an explanatory signal for forgetting and a practical control signal for LLM continual post-training.
PaperID: 5850, Poster
Authors: Shixiong Jiang, Jialiang Fan, Mengyu Liu, Pengfei Gu, Danny Z Chen, Fanxin Kong
Abstract: Combining the Segment Anything Model (SAM) with fine-tuning techniques allows SAM to be effectively adapted to various downstream image segmentation tasks. However, this adaptability introduces new security vulnerabilities related to adversarial attacks. In this paper, we investigate the adversarial transferability between the original SAM and its fine-tuned downstream models. Under limited knowledge conditions of the downstream models, we propose a novel structure-exploiting transferable attack (SETA) method. Our framework mimics the fine-tuning architecture and estimates the parameter distributions of the downstream models to improve the transferability of the generated adversarial samples. Experimental results demonstrate the efficacy of our proposed method in creating adversarial examples against various downstream fine-tuned SAM models.
PaperID: 5851, Poster
Abstract: Vision-Language-Action (VLA) models have achieved remarkable success in robotic manipulation. However, their robustness to instruction variations remains a critical, under-explored safety concern, posing a significant safety risk to real-world deployment. Red teaming, or identifying environmental scenarios that elicit catastrophic behaviors, is an important step in ensuring the safe deployment of embodied AI agents. Reinforcement learning (RL) has emerged as a promising approach in automated red teaming that aims to uncover these vulnerabilities. However, standard RL-based adversaries often suffer from severe mode collapse due to their reward-maximizing nature, which tends to converge to a narrow set of trivial or repetitive failure patterns, failing to reveal the comprehensive landscape of meaningful risks. To bridge this gap, we propose a novel Diversity-Aware Embodied Red Teaming (DAERT) framework, to audit VLA robustness under semantically aligned instruction. Our design uses a breadth-seeking value estimator that prevents the attacker from collapsing onto a single high-reward phrasing, generating a diverse set of challenging instructions while preserving attack effectiveness, measured by execution failures in a physical simulator. We conduct extensive experiments across different robotic benchmarks against two state-of-the-art VLAs, including \pi_0 and OpenVLA. Our method consistently discovers a wider range of more effective adversarial instructions that reduce the average task success rate from 93.33% to 5.85%, demonstrating a scalable approach to stress-testing VLA agents and exposing critical safety blind spots before real-world deployment.
Abstract: Large language models (LLMs) rely on static pretraining corpora, causing their knowledge to become outdated over time. Existing approaches for evaluating knowledge edits either suffer from rapid contamination or rely on counterfactual edits that conflict with rigid existing knowledge. In this work, we propose a synthetic, simulation-driven framework for studying knowledge updates in LLMs. We introduce PARALLELEVENTS, a benchmark of fictional yet realistic future worlds that generates coherent event trajectories for controlled evaluation, avoiding contamination while preserving consistency. Building on this dataset, we develop SYNAPSE, a training framework that uses model-generated data to update model parameters via mid-training and instruction tuning. This synthetic pipeline enables scalable knowledge integration without costly human-curated data. Empirically, \sc Synapse outperforms existing methods by 14.23%, demonstrating that simulation-based synthetic training leads to robust and coherent knowledge updates.
Abstract: Gaze following requires both scene understanding and gaze reasoning to localize the gaze target of an in-scene person. Recently, vision foundation models (VFMs) have demonstrated strong performance on this task, enabling simpler architectures while outperforming prior methods. However, we observe a key limitation of VFM-based approaches: while VFMs substantially improve scene understanding, they contribute little to gaze reasoning. As a result, existing methods often rely on semantically salient objects rather than true gaze cues, leading to degraded performance when targets are not salient. To address this, we propose a novel training mechanism to enhance gaze reasoning in VFMs for gaze following. Our method includes: (1) a head-conditioned local LoRA, which enables localized adaptation to preserve scene token learning while improving head token learning for gaze reasoning; and (2) an out-of-cone penalty, which injects gaze cues into head tokens while aligning them with scene tokens. Experiments on the GazeFollow and VAT datasets demonstrate that our method achieves state-of-the-art performance, with particularly strong improvements when gaze targets are not semantically salient. Our findings offer valuable insights for advancing future gaze following research. We will release the code once the paper is accepted.
PaperID: 5854, Poster
Authors: Joon Suk Huh, Junghoon Seo
Abstract: We study off-policy evaluation (OPE) in contextual bandits with data collected from multiple logging policies. Inverse propensity scoring (IPS) is a standard approach to OPE and extends naturally to the multi-logger setting. However, as highlighted by Agarwal et al. [2017], there appears to be no IPS estimator that consistently outperforms the others in this setting. We resolve this dilemma by deriving an optimal IPS estimator with sample-dependent weights that minimize variance subject to unbiasedness. Using a variational calculus approach, we obtain closed-form optimal weights, yielding an estimator that is unbiased and achieves asymptotically optimal variance within an weighted-IPS estimator class. Experiments on benchmark datasets confirm this theoretical resolution in practice, showing that our estimator consistently outperforms existing multi-logger IPS methods. We also extend our estimator to a doubly robust form by incorporating a conditional reward estimator. We compare the resulting DR extension with the semiparametrically efficient DR estimator of Kallus et al. [2021], both theoretically and empirically. We show that the two estimators achieve the same asymptotic variance when the conditional reward estimator is consistent, while our estimator can retain semiparametric efficiency under certain forms of reward misspecification. Empirically, the proposed DR estimator achieves lower variance than the existing DR baseline.
PaperID: 5855, Poster
Authors: Haoqi Ruan, Ruihan Hu, XinruiCheng, Ziyi Zhong, Zhaodong Zhang, Wenqi Wang, Xuheng Zhang
Abstract: Diffusion and flow-matching models have become the dominant paradigm for high-fidelity generation across images, video, and 3D content. However, acceleration methods that work well in image or video often fail to preserve 3D geometric consistency, while existing 3D inference approaches still rely on fixed reuse schedules and cannot adapt computation to instance difficulty or trajectory-specific demands. We propose AEGIS, a training-free framework that reframes test-time acceleration as quality-constrained budget allocation solved by an online feedback rule. AEGIS maintains a two-dimensional runtime state that tracks sample difficulty and accumulated correction debt, and turns it into per-step skip and refresh quotas that any cache-style executor can realize. It generalizes the corresponding open-loop schedule and provides an explicit upper bound on cumulative reuse drift. Extensive experiments show that AEGIS improves inference efficiency while better balancing throughput and quality.
PaperID: 5856, Poster
Abstract: Denoising generative models predict clean data (x), noise (\epsilon), or velocity (v), and are trained with a squared loss that can be computed in any of these three spaces. Prior work treats these choices as interchangeable up to loss reweighting, but practitioners observe sharp qualitative differences that remain unexplained. We systematically decouple the \emphprediction space, what the network outputs, from the \emphloss space, where supervision is applied, creating a 3×3 design matrix. We show that loss-space changes only reweight how much each noise level contributes to training, while prediction-space changes alter what the network is asked to learn: two fundamentally different mechanisms. This decoupling exposes two regimes. Without a representational bottleneck, the loss choice primarily controls denoising quality (MSE) while the prediction choice primarily controls sample quality (FID). The reason is that ODE sampling must convert the network's output back into a clean image, and this conversion amplifies errors differently across prediction types, turning sub-one-percent gaps in training error into 22-fold gaps in FID that are invisible to standard training diagnostics. Under a representational bottleneck, the prediction choice instead dominates both metrics: recovering noise or velocity requires reconstructing high-dimensional information that the bottleneck has destroyed, while recovering the clean image remains feasible because real images lie on a low-dimensional manifold. We validate these findings on a U-Net and the full 130M-parameter JiT architecture across CIFAR-10 and Imagenette, and give architecture-dependent recommendations for both regimes.
PaperID: 5857, Poster
Abstract: Lifelong person re-identification (LReID) requires models to continuously learn from sequentially arriving domains while retaining discriminative power for previously seen identities. A key challenge is to prevent catastrophic forgetting without access to old data, especially under exemplar-free constraints. While flat-minima optimization has shown promise in continual learning, we identify a gradient conflict that limits the effectiveness of standard Sharpness-Aware Minimization (SAM) in LReID. The ReID loss gradient dominates the perturbation direction, causing the sharpness of the distillation loss to be underestimated, which hinders the flattening of the landscape necessary for knowledge retention. To resolve this conflict, we propose a framework that unifies selective flatness-aware optimization, dual-model training, and weight-space interpolation. Specifically, we maintain two models per task: a stability model whose SAM perturbation is computed solely from the distillation loss, and a plasticity model optimized for the current domain. By decoupling the perturbation objectives, our selective SAM achieves more targeted sharpness exploration along the distillation landscape, guiding the stability model toward flatter and robust regions. After training, the two models are fused via weight-space interpolation, and we provide a theoretical bound showing that flatter stability solutions tighten the interpolation bound and reduce forgetting during fusion. Our method is lightweight, modular, and readily compatible with existing LReID frameworks. Experimental results demonstrate that the proposed method consistently improves the performance of LReID in terms of both knowledge retention and generalization to unseen domains.
PaperID: 5858, Poster
Abstract: Test-time scaling is an important mechanism for improving large language models, especially on tasks with deterministic verifiers. Code translation is a canonical example: the source program constrains valid outputs, while compilers, type check- ers, and behavioral checks provide exact pass/fail feedback. Existing approaches typically apply these verifiers only after generation, which is inefficient because early errors corrupt the autoregressive context and are rarely corrected later. We introduce Decoding Time Verification (DTV), a framework that integrates deter- ministic program verification directly into the decoding loop. DTV interleaves generation with verifier calls under a state-machine controller that enforces valid prefixes, using structural-boundary checks and structure-aware rollback to prevent error propagation while reducing wasted tokens. We evaluate DTV on C-to-Rust and JavaScript-to-TypeScript translation. Using Qwen3-4B as the primary genera- tor under matched token budgets, DTV improves pass rates from 72.3% to 82.0% on C-to-Rust and from 33.3% to 46.0% on JavaScript-to-TypeScript relative to matched self-refinement baselines, while using fewer tokens per case; the same trend largely transfers to Gemma-4-E4B. In the evaluated cost-matched grid, DTV achieves a more favorable pass-rate-cost tradeoff than post-hoc verification or sampling-based scaling. These results show that verifier-guided decoding is an effective use of inference-time compute for code translation.
PaperID: 5859, Poster
Abstract: Reinforcement learning (RL) has shown strong potential for aligning generative models with human intent, but adapting RL to masked generative models (MGMs) remains largely underexplored. Existing approaches that directly adapt Group Relative Policy Optimization (GRPO) to MGMs leave a foundational conflict unaddressed: MGMs' exploitative decoding nature is inherently at odds with the exploratory diversity that GRPO presupposes. In this work, we reveal a structural decoding bias termed the visual exploration trap: confidence-based sampling trades the exploration of diverse subject realizations for the greedy resolution of low-uncertainty background regions, causing premature collapse of the generative exploration space and starving GRPO of the rollout diversity it depends on. To this end, we propose MaskSense, a novel RL framework for MGMs that confronts the visual exploration trap at both the sampling and optimization stages. Specifically, we introduce a semantic anchored routing sampling mechanism that leverages semantic priors to preserve high-entropy subject token exploration while amplifying intra-group reward variance through exploratory-exploitative routing to yield more discriminative advantage estimates. Furthermore, we design a global entropy transition anchoring strategy to identify the most consequential decoding steps and concentrate gradient updates on them. Extensive experiments on multiple text-to-image benchmarks demonstrate that MaskSense substantially improves upon the base model, with GenEval accuracy improving from 54% to 86% and HPS from 28.89 to 36.39, achieving state-of-the-art performance.
PaperID: 5860, Poster
Abstract: Creating explorable 3D environments from text is important for immersive content creation, simulation, and interactive design. Recent methods often rely on image or video generative priors by synthesizing views along fixed or manually specified trajectories and lifting them into 3D. However, such passive pipelines cannot adapt view generation to the evolving scene memory, often leaving disoccluded regions incomplete and producing holes, floaters, and inconsistent geometry under free-viewpoint exploration. We argue that the key challenge is not to generate more views, but to decide where to generate useful views according to the current scene state. To this end, we propose ActPano3D, an active text-to-3D scene generation framework based on memory-guided panoramic exploration. ActPano3D formulates trajectory selection as spatial deliberation over the evolving scene memory, where candidate actions are evaluated by their expected utility for exploring unknown regions, repairing uncertain geometry, and improving the future 3D world model. The framework integrates expansion-refinement active planning, memory-routed panoramic generation, and reliability-aware scene memory fusion into a closed-loop generation-and-reconstruction process. The accumulated memory is further optimized into a renderable 3D Gaussian Splatting scene for free-viewpoint exploration. Experiments show that ActPano3D improves scene coverage, reduces rendering holes, and achieves better rendering quality and text alignment than existing text-to-3D scene generation baselines.
PaperID: 5861, Poster
Authors: Ziyi Liu, Dan Roy
Abstract: Smoothed online learning has recently been studied as a way to bypass hardness results for the fully adversarial setting, which can be overly pessimistic. In this framework, the adversary is constrained to generate contexts from distributions having a bounded density with respect to some fixed base measure \mu. Prior work makes the strong assumption that \mu is known to the learner, with notable exceptions including Block et al. (2024) and Blanchard (2025). In this paper, we study sequential probability assignment (a.k.a., online learning with log loss) with smooth, well-specified data in the more general setting where \mu is \emphunknown, without the Lipschitz condition imposed in the work of Block et al. (2024) and Blanchard (2025). We prove regret upper bounds against both adaptive and oblivious smoothed adversaries: the adaptive bound is controlled by empirical \ell_\infty Hellinger entropy, while the oblivious bound improves this dependence to empirical \ell_2 Hellinger entropy, the notion that has been shown to characterize the complexity of learning with i.i.d. data (Bilodeau et al., 2023). Both bounds are obtained through algorithms based on truncated empirical Hellinger covers. We complement these results with a general lower bound, showing that our upper bounds are essentially tight for some natural classes, which implies a separation in the difficulty of smoothed online learning between regimes where \mu is known and where it is unknown.
Abstract: As large language models (LLMs) scale rapidly, dense full-parameter adaptation becomes increasingly expensive, motivating sparse and modular architectures such as Mixture-of-Experts (MoE) models. This shift raises a key question for parameter-efficient fine-tuning (PEFT): at what granularity should parameters be selected and updated? Existing PEFT methods such as LoRA operate on predefined weight matrices, while expert-level sparse tuning methods update entire selected experts. However, we observe that activated experts are internally sparse, with only a small fraction of intermediate channels strongly responding to downstream tasks, indicating that expert-level adaptation is still too coarse. We propose NSFT (Neural Sub-expert Fine-Tuning), a fine-grained PEFT framework that refines MoE adaptation from experts to sub-experts. NSFT decomposes each expert along the intermediate dimension into structured channel groups and selects task-relevant sub-experts by combining routing importance with intra-expert activation saliency. To optimize sparse partial updates, NSFT further introduces learning-rate scaling and dynamic gradient scaling to compensate for the reduced effective update magnitude. Experiments on OLMoE and Ling-mini-2.0 across challenging domain-specific tasks and general benchmarks show that NSFT consistently outperforms representative PEFT and expert-level sparse tuning baselines, while using substantially fewer trainable parameters and preserving competitive general capability. These results suggest that sub-expert-level adaptation is a more precise and efficient PEFT paradigm for MoE LLMs.
PaperID: 5863, Poster
Abstract: Video large multimodal models incur substantial inference cost because visual tokens accumulate across sampled frames. Training-free token compression can reduce this cost, but it often degrades video reasoning performance when later frames must contribute new evidence over time. We identify a common failure mode behind this degradation: existing compressors do not track the committed support, namely the token support already retained and delivered to the language model by earlier frames. As a result, selectors with different scoring rules repeatedly retain overlapping supports across neighboring frames, a phenomenon we call Selector Collapse. To address this issue, we propose \tempo, a frame-causal token compression framework that combines committed-support memory, residual-energy greedy selection, and parameter-free neighborhood fusion. Across four video understanding benchmarks and multiple backbones, \tempo consistently improves the accuracy-efficiency trade-off of training-free compression; at 10% retention, it preserves 99.8% of vanilla performance on LLaVA-OneVision and improves fixed-budget frame scaling on Qwen2.5-VL to 109.3% relative accuracy, without retraining, backbone modification, or additional token budget.
Abstract: We investigate the short-context dominance hypothesis: that for most sequences, a small local prefix suffices to predict their next tokens. Using large language models as statistical oracles, we measure the minimum context length (MCL) needed to reproduce accurate full-context predictions across datasets with sequences of varying lengths. For sequences with 1–7k tokens from long-context documents, we consistently find that 75–80% require only the last 96 tokens at most. Given the dominance of short-context tokens, we then ask whether it is possible to detect challenging long-context sequences for which a short local prefix does not suffice for prediction. We introduce a practical proxy to MCL, called Distributionally Aware MCL (DaMCL), that does not require knowledge of the actual next-token and is compatible with sampling strategies beyond greedy decoding. Our experiments validate that simple thresholding of the metric defining DaMCL achieves high performance in detecting long vs. short context sequences. Finally, to counter the bias that short-context dominance induces in LLM output distributions, we develop an intuitive decoding algorithm that leverages our detector to identify and boost tokens that are long-range-relevant. Across Q&A tasks and model architectures, we confirm that mitigating the bias improves performance.
PaperID: 5865, Poster
Authors: Jason Bohne, Pawel Polak, Gary Kazantsev, David Rosenberg
Abstract: Bilevel optimization is a core tool in practical machine learning workflows, including hyperparameter tuning, data reweighting, robust learning, reinforcement learning, and imbalanced classification. Hypergradient computation requires the inverse of the inner Hessian, typically estimated via a truncated Neumann series. While existing analyses focus on bias and variance, we show that the product structure of the Neumann series introduces a third source of error: skewness. Unlike ordinary gradients, whose asymmetry is suppressed by mini-batch averaging, the Neumann estimator amplifies skewness across factors, with the third-moment error growing with truncation depth, inner condition number, and per-factor noise asymmetry. This arises in many practical settings such as corrupted labels, outliers, heterogeneous or class-imbalanced datasets. To address this, we introduce the Antithetic Neumann Estimator (ANE), which pairs each Hessian sample with a mirrored counterpart around a reference Hessian and averages the resulting inverse-Hessian estimates. With an exact reference, the antithetic product retains only even-order noise, eliminating all skewness; with an approximate reference, the residual skewness scales linearly with the reference error. ANE requires only a modification of the Hessian-vector step and integrates with existing stochastic or variance-reduced bilevel methods. We instantiate ANE-SGD and ANE-VR, the latter adding recursive variance reduction. Across robust regression, corrupted-label learning, offline reinforcement learning, and imbalanced classification, ANE-based methods achieve strictly better bias–variance–skewness tradeoffs than stochastic and variance-reduced baselines, reducing loss by up to 24% and improving predictive metrics by up to 27%. This establishes skewness correction as a practical and complementary mechanism for improving hypergradient quality.
Abstract: We investigate the strategic surplus obtainable against a Follow-the-Regularized-Leader (FTRL) learner with constant step size \eta in n × m two-player zero-sum games played over T rounds against a clairvoyant optimizer. In contrast with prior analysis, we show that the extraction of such regret-scale surplus is an inherent feature of the FTRL family, rather than an artifact of specific instantiations. First, for a fixed max-min optimizer, we establish a sweeping law of order \Omega(N/\eta), proving that utility surplus scales with the number of the learner's suboptimal actions N and vanishes in their absence. Second, for an alternating optimizer, a surplus of \Omega(\eta T/\textpoly(n,m)) can be guaranteed regardless of the equilibrium structure, with high probability, in random games. Our analysis uncovers a sharp geometric dichotomy: non-steep regularizers allow the optimizer to realize the maximal transient surplus via finite-time elimination of suboptimal actions, whereas steep regularizers introduce a vanishing tail correction that can delay surplus saturation. Finally, we discuss whether this leverage persists under bilateral payoff uncertainty and propose a susceptibility measure quantifying which regularizers are most vulnerable to learner-aware strategic steering.
PaperID: 5867, Poster
Authors:
Seungkwon Yang, Gyeongjin Kang, Hwasik Jeong, Byeongjin Kang, Hyeongbhin Cho, Eunbyung ParkAbstract: Multi-view transformers for feed-forward 3D reconstruction must jointly solve two tasks. They need to extract precise per-view features and reason about cross-view geometric correspondences. Existing architectures apply cross-view attention throughout the network, leaving each layer to handle both objectives together. We revisit this design through a layer-wise probing analysis of feed-forward geometry transformers and observe that trained models tend to divide these two roles across layers. Early layers tend to specialize in per-frame feature extraction, while spatial cognition and cross-view reasoning emerge predominantly in later layers. This suggests that separating the two roles across the network can allocate compute more effectively. Building on this observation, we propose U-MVP, a U-shaped multi-view transformer for feed-forward 3D Gaussian Splatting that reflects this structure in the architecture itself. The encoder applies frame-wise attention to preserve per-view detail, the decoder introduces multi-view attention for cross-view geometric reasoning, and skip connections fuse the two streams so that appearance and geometry are recombined at decoding time. To scale to dense view regimes, we replace several global interaction blocks at the bottleneck with a grouping and swapping scheme that preserves cross-view information flow at lower cost. U-MVP performs competitively with feed-forward and optimization-based baselines from 16 to 256 views, generalizes to unseen datasets, and achieves strong results on 3D reconstruction and novel-view synthesis in posed and unposed settings.
PaperID: 5868, Poster
Abstract: Bridging independent vision-language models (VLMs) and small language models (SLMs) in a modular multi-agent system typically forces them to exchange information as discrete text, which creates a lossy bottleneck for spatial and fine-grained details. This constraint is especially acute for SLMs, whose limited capacity makes them more reliant on the quality of the input signal. While recent works on latent multi-agent communication have explored latent-space exchange between identical or homogeneous LLMs, propagating rich visual information across model boundaries without text generation has not been studied yet. We propose Latent-Lens, a lightweight codec that bridges the representation gap between a frozen VLM and a frozen SLM directly in latent space, enabling them to communicate purely through hidden states. The codec consists of a domain encoder that projects hidden states from the VLM into the SLM's embedding space, and an attention decoder formalised as a LoRA that teaches the SLM to attend to these visual prefixes, while training only 1.6M parameters (0.2% of the combined model size). On free-form visual question answering (CLEVR, GQA), Latent-Lens achieves +10.2-28.3pp over a text-communication baseline given the same LoRA budget (+36.4-78.4pp when both agents are frozen); on image captioning (Flickr8k) it improves over the same baseline by +18.5 (+18.9 frozen). Replacing the SLM (TinyLlama-1.1B) or the VLM (InternVL2-1B) with alternative models of different architectures, vocabularies, and hidden dimensions yields consistent gains, confirming generality across heterogeneous model families.
Abstract: Biological systems can form complex three-dimensional structures through the collective behavior of agents that share a common update rule and operate without central control. How such distributed control gives rise to precise global patterns remains a central question not only in developmental biology but also in distributed robotics, programmable matter, and multi-agent learning. Here, we introduce DiffeoMorph, an end-to-end differentiable framework for learning a morphogenesis protocol that guides a population of agents to morph into a target 3D shape. Each agent updates its position and internal state using an SE(3)-equivariant graph neural network, based on its own internal state and signals received from other agents. To train this system, we introduce a new shape-matching loss based on 3D Zernike polynomials, which compares the predicted and target shapes as continuous spatial distributions, not as discrete point clouds, and is invariant to agent ordering, number of agents, and global orientation. To achieve rotation invariance while preserving reflection sensitivity, we include an alignment step that optimally rotates the predicted Zernike spectrum to match the target before computing the loss. We perform benchmarking to establish the advantages of our shape-matching loss over other standard distance metrics for shape comparison tasks. We then demonstrate that DiffeoMorph can form a range of complex shapes from minimally patterned initial conditions. DiffeoMorph provides a general framework for learning distributed control strategies for morphogenesis, swarm robotics, and programmable self-assembly.
Abstract: Modern Transformer design and compression both reduce to allocating capacity under a budget. The standard scalars for these decisions, #Params and #FLOPs, capture size and compute but not architectural structure—two architectures with identical parameter budgets but different depth-width, head, or FFN allocations receive identical scores yet behave differently. We propose Neural Spectral Capacity (NSC), a closed-form scalar grounded in the singular-value spectrum of each weight matrix. Under standard random initialization, the Marchenko–Pastur law renders NSC computable from the architectural specification alone, with no model instantiation, data, or gradients. Its layer-wise additive structure admits NSC-DP, an exact dynamic-programming solver returning the architecture globally maximizing NSC under resource constraints in seconds on a CPU—a guarantee that black-box search over existing training-free proxies cannot provide. Empirically, NSC outperforms #Params, #FLOPs, and representative training-free proxies in ranking across seven Transformer and CNN families (on FlexiBERT, τ=0.505 on pairs differing in #Params by <10%, where #Params collapses to 0.082); NSC-DP discovers a Transformer-XL architecture on WikiText-103 that beats the human-designed baseline in 2 seconds; and prunes LLaMA-7B to the best 5.7B model across eight commonsense reasoning tasks without any calibration data, ∼5900× faster than the strongest training-free proxy baseline.
PaperID: 5871, Poster
Abstract: Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines—for instance, they contain little explicit reasoning. Thus, many frontier labs have begun to develop their own internal datasets, starting from state-of-the-art models, to augment their pre-training data mix, e.g., with reasoning traces to address cold start problems. While demonstratively effective, none of these datasets are public, and the effect of this so-called _synthetic data_ on knowledge and skill acquisition of language models, including small ones, remains poorly understood. We present Synth, the first open-source synthetic corpus derived from 58,698 Wikipedia articles that collapses pre-, mid-, and post-training into a single training stage via structured amplification of curated encyclopedic seeds. We evaluate Synth by training a series of models: a 56M tiny model (Monad), 0.3B--0.6B dense models (Baguettotron), and a 13B-total / 1B-active Mixture-of-Experts. At iso-compute, Synth outperforms filtered web data, and our models remain competitive with similarly-sized open-weight baselines. Because Synth is back-translated from grounded passages, Synth-trained models achieve high factual precision despite 10-140× fewer training tokens, with memorization targeted by the seed corpus. These results show that synthetic datasets, including our Synth dataset, are capable of producing competitive generalist models at a significantly lower cost, enabling rapid iteration as the frontier advances. These findings open up possibilities for both generalist models with significantly increased data efficiency, as well as domain-specific models where no instruction or conversational data is available. Finally, we publicly release our Synth dataset and the series of Baguettotron models under a permissive license, thus supporting open-source language model development.
Abstract: Recovering governing Ordinary Differential Equations (ODEs) from data is a central challenge in modeling dynamical systems across scientific domains. Existing approaches cast discovery as a static inference problem over fixed datasets, implicitly assuming that the observed trajectories are sufficiently informative. However, dynamical systems evolve over large state spaces, and limited data can often fit multiple distinct equations that explain the observations equally well, leading to identifiability gaps and incorrect recovery of the true dynamics. We introduce LLM-ACES or LLM-guided Active Closed-loop Equation Search, a closed-loop framework that jointly optimizes data acquisition and hypothesis generation under a constrained simulation budget. LLM-ACES leverages LLMs to construct structured, domain-informed hypothesis spaces via a two-stage process: high-level prior induction followed by candidate equation generation within each prior. To resolve ambiguity among competing hypotheses, we introduce an active data selection strategy that identifies regions of maximal predictive divergence among candidate equations and queries the system to obtain informative trajectories. This induces a feedback loop in which hypotheses guide data acquisition and newly acquired data refine the hypothesis space. Experiments demonstrate that LLM-ACES achieves more accurate equation recovery than prior state-of-the-art methods. Our ablations and analyses highlight the importance of coupling hypothesis construction with feedback-driven data acquisition for reliable discovery of dynamical systems.
Authors:
Guo Yucheng, Yongjian Guo, 钟 关, Wen Huang, Haodong Yue, shuai di, Xiong J Wu, Yicheng GongAbstract: The rapid evolution of Embodied AI has enabled Vision-Language-Action (VLA) models to excel in multimodal perception and task execution. However, applying Reinforcement Learning (RL) to these massive models in large-scale distributed environments faces severe systemic bottlenecks, primarily due to the resource conflict between high-fidelity physical simulation and the intensive VRAM/bandwidth demands of deep learning. This conflict often leaves overall throughput constrained by execution-phase inefficiencies. To address these challenges, we propose D-VLA, a high-concurrency, low-latency distributed RL framework for large-scale embodied foundation models. D-VLA introduces "Plane Decoupling," physically isolating high-frequency training data from low-frequency weight control to eliminate interference between simulation and optimization. We further design a four-thread asynchronous "Swimlane" pipeline, enabling full parallel overlap of sampling, inference, gradient computation, and parameter distribution. Additionally, a dual-pool VRAM management model and topology-aware replication resolve memory fragmentation and optimize communication efficiency. Experiments on benchmarks like LIBERO show that D-VLA significantly outperforms mainstream RL frameworks in throughput and sampling efficiency for billion-parameter VLA models. In trillion-parameter scalability tests, our framework maintains exceptional stability and linear speedup, providing a robust system for high-performance general-purpose VLA Models.
Abstract: Transformer-based autoregressive models offer an efficient alternative to diffusion- and flow-matching-based approaches for generating 3D molecules. One challenge remains: standard transformer architectures require a sequential ordering of tokens, which is not inherently defined for the atoms in a molecule. Prior works have addressed this by using canonical atom orderings. However, these approaches are not permutation invariant w.r.t. atoms and bias next-token prediction towards ordering conventions. We overcome this limitation by introducing a novel neighborhood-guided training strategy. Our model, NEAT (Neighborhood-Guided, Efficient, Autoregressive Set Transformer) treats molecular graphs as sets of atoms and learns an order-agnostic distribution over admissible tokens at the graph boundary, thereby ensuring atom-level permutation invariance. NEAT achieves state-of-the-art generation quality on the QM9 and GEOM-Drugs datasets while offering a significant speed advantage over existing baselines.
PaperID: 5875, Poster
Abstract: As intelligent assistants are increasingly deployed in real-world environments, recommendations move from passive preference matching to grounding user intent from underspecified multimodal evidence. In these settings, users often provide only a reference image and a fuzzy natural-language query, leaving crucial visual attributes, personalized preferences, and situational constraints implicit. Existing recommenders either rely on historical interactions enriched with multimodal item features or presuppose fully specified textual requests, which are insufficient to resolve such queries in a single-turn setting. We present IntentLens, a tool-augmented recommendation framework that grounds underspecified multimodal queries into explicit, query-relevant evidence before ranking. IntentLens adopts a \emphground-then-rank design: a shared multimodal LLM orchestrates visual, user-memory, and item-attribute tools to (i) recover latent user intent from the image and language query, and (ii) enrich each candidate with fine-grained, query-conditioned evidence; a lightweight ranker then scores candidates over the grounded representations. To support this new setting, we further construct two benchmarks with underspecified multimodal queries, \textscGoogleReview-MMTool and \textscYelp-MMTool. Comprehensive experiments show that IntentLens outperforms strong MLLM and retrieval baselines.
Authors: Bryce Hinkley, Peyman Najafirad
Abstract: We study selective refusal editing as a three-way control problem: induce non-refusal on designated edit prompts while preserving benign behavior and harmful refusals outside the edit set. We introduce Residual Paving, a routed residual editing method for frozen instruction-tuned transformers that separates route selectivity, whether to intervene, from residual-edit capacity, what edit to apply. An early-layer router predicts a scalar gate and expert mixture; when active, prompt-conditioned bottleneck residual experts apply later-layer residual updates while leaving the backbone unchanged. This decomposition supports an oracle-routing diagnostic where only the learned scalar gate is replaced with the held-out edit/keep label, leaving the residual editor and frozen backbone fixed. On the primary Gemma-3-4B-IT held-out split, learned Residual Paving reduces edit refusal from 88.6% to 4.0%, with 95.5% benign distribution preservation and 87.3% harmful distribution preservation. Same-protocol one-direction steering controls are much weaker on edit success, leaving edit refusal at 86.8% for Edit-target ActAdd and 78.9% for DIM-style refusal steering. The remaining failure is off-target harmful-keep degradation: harmful refusal remains below the frozen-base rate, 65.3% vs. 81.6%. Across six backbones, oracle routing improves the keep-side diagnostic score on every reported row, with median gain +12.9 pp, supporting the interpretation that learned route selectivity is the main observed bottleneck. Trajectory diagnostics on two backbones further suggest directed movement toward edit-target continuations rather than generic refusal suppression.
PaperID: 5877, Poster
Abstract: Offline reinforcement learning (RL) trains policies from static datasets, but deployed agents may later need to attenuate the influence of specific data for privacy, safety, or fairness. This is challenging in the overlap regime, where the forget set and retained data share state-action support: unlearning cannot be uniform as suppressing values on shared support can inadvertently erase behaviors required for high performance. We propose Conservative Unlearning with soft-gating Regularization (CURE), a selective unlearning framework tailored to overlap. CURE assigns each forget trajectory an overlap-interference score via cosine alignment between its temporal difference (TD) semi-gradient and a retained reference gradient, and uses it to soft-gate i) conservative, state-conditional critic suppression relative to a retained baseline and ii) coupled actor updates, selectively modulating forgetting based on interference, reducing impact on aligned data. We provide a first-order analysis showing that CURE descends a forget surrogate while controlling retained drift via alignment-dependent coupling. Across MuJoCo benchmarks and offline RL backbones, CURE consistently improves the forget-utility trade-off over baselines, achieving low auditor-detectable forget influence with near-original returns at a fraction of retraining cost.
PaperID: 5878, Poster
Abstract: The rapid growth of video-based social media has increased users’ exposure to harmful content, creating a need for reliable automated video safety detection. Although recent Vision-Language Models (VLMs) show strong video understanding capabilities, existing harmful video detection systems face two key limitations: they typically reduce safety detection to binary classification, overlooking the inherently multi-label nature of unsafe videos, and they rely on static training objectives that do not support controllable precision-recall trade-offs, though the desired operating point may vary across moderation pipelines and unsafe categories. To address these gaps, we propose Adaptive Tversky Policy Optimization (ATPO), a reinforcement learning framework for Multi-label Video Safety Detection (Multi-VSD). ATPO introduces the Adaptive Tversky Reward (ATR), which dynamically adjusts false-positive and false-negative penalties during training to enable controllable precision–recall trade-offs. Experiments on SafeWatch-Bench and XD-Violence show that ATPO substantially improves multi-label performance, increasing the Jaccard Index from 40.66 to 75.44 on SafeWatch-Bench-Real. Moreover, ATR enables reliable steering of the precision–recall operating point, supporting deployment scenarios with heterogeneous policy requirements.
Abstract: Reinforcement learning with verifiable rewards (RLVR) provides a promising pathway for continuously advancing GUI agents, yet existing reward modeling paradigms face complementary limitations. Rule-based methods suffer from poor scalability and cannot handle open-ended tasks. LLM-as-a-Judge methods enable scalable trajectory verification but remain passive and are constrained by partial state observability, since key evidence often resides in latent environment states beyond the trajectory. Recent active environment interaction methods mitigate observability issues but tend to over-rely on probing while under-utilizing direct trajectory evidence, leading to verification inefficiency. To address these challenges, we advocate a trajectory-grounded interactive verification paradigm. We introduce VAGEN, a framework that employs a tool-augmented verifier agent governed by a Progressive Verification Mechanism, which follows a surface-to-latent and cheap-to-expensive design philosophy to extract trajectory evidence and probe environment states in a proactive end-to-end manner. Experiments on OSWorld-Verified and AndroidWorld benchmarks demonstrate that VAGEN significantly improves evaluation accuracy with a favorable performance-efficiency trade-off.
PaperID: 5880, Poster
Authors: Nathan White, Krish Singal
Abstract: Quantization is a fundamental tool used to compress datasets, neural network weights, and memory usage in a range of computational tasks. Many downstream applications of vector quantization perform inner products with arbitrary inputs. This motivates the study of inner product aware quantization schemes that approximately preserve inner products with unseen vectors -- in contrast to simply minimizing the mean-squared error. In this work, we formulate objectives that capture natural desiderata and develop adaptive and unbiased quantization methods that approximately preserve inner products with worst-case and average case inputs. An analysis of these objectives shows a tight connection with the well-studied notion of Adaptive Stochastic Quantization (ASQ). We develop provably fast exact and approximate algorithms for our objectives. Our theoretical results inspire efficient practical algorithms that perform well across a variety of workload distributions. They also lead to practical algorithms for standard ASQ which are 2-10× faster than prior state-of-the-art methods while maintaining quality. These theoretical and empirical results contribute towards making adaptive quantization techniques more efficient and tractable in practical settings.
Abstract: A bottleneck in learning to understand articulated 3D objects is the lack of large and diverse datasets. In this paper, we propose to leverage large language models (LLMs) to close this gap and generate articulated assets at scale. We reduce the problem of generating an articulated 3D asset to that of writing a program that builds it. We then introduce a new agentic system, Articraft, that writes such programs automatically. We design a programmatic interface and harness to help the LLM do so effectively. The LLM writes code against a domain-specific SDK for defining parts, composing geometry, specifying joints, and writing tests to validate the resulting assets. The harness exposes a restricted workspace and interface to the LLM, validates the resulting assets, and returns structured feedback. In this way, the LLM is not distracted by details such as authoring a URDF file or managing a complex software environment. We show that this produces higher-quality assets than both state-of-the-art articulated-asset generators and general-purpose coding agents. Using Articraft, we build Articraft-10K, a curated dataset of over 10K articulated assets spanning 245 categories, and show its utility both for training models of articulated assets and in downstream applications such as robotics simulation and virtual reality.
Abstract: Approximate machine unlearning aims to remove the influence of specific training data from a trained model without retraining from scratch. We identify a previously undocumented confound in how unlearning is evaluated on BatchNorm-based architectures: a single forward pass over retain data, an operation that modifies no weight, can deterministically rewrite the model's normalization state and reverse the apparent surface-metric forgetting. We formalize this operation as a weight-preserving fixed-point operator and prove that any pre-versus-post gap it induces is provably attributable to BN running statistics rather than to any modification the unlearning method made to the weights. This attribution claim cleanly separates measurement failure (BN artifact) from encoder failure (residual weight-encoded information, recently documented in concurrent work), and the same operator framework yields a unique decomposition of linear-probe elevation into BN-measurement-bias and encoder-geometry components. Empirically, the artifact reverses headline forget accuracy by up to 78pp across nine published methods on standard benchmarks; an attacker with as few as 10 unlabeled images recovers most of the masked accuracy; and a strict GroupNorm control reduces the artifact to zero across all methods.
PaperID: 5883, Poster
Abstract: Diffusion large language models (dLLMs) generate tokens in parallel via iterative denoising, but their bidirectional attention prevents the KV caching used by autoregressive models and recomputes the full attention map at every step, incurring quadratic per-step cost. Existing sparse attention methods for dLLMs compute a sparse mask once at the first denoising step and reuse it throughout, while allocating the sparsity budget uniformly across heads. We argue both choices ignore key structural properties of attention in dLLMs: empirically, attention patterns are only locally consistent across steps and shift abruptly at certain transitions, and different heads concentrate their attention mass to very different extents. To this end, we design Step-dLLM, which periodically refreshes a block-sparse mask at anchor steps along the denoising trajectory and reuses it within each locally consistent interval, while a head-adaptive global budget assigns more blocks to heads with more concentrated attention. Backed by a fused IO-aware Triton kernel for block-sparse attention, Step-dLLM applies to pre-trained LLaDA-1.5 and Dream-7B-Instruct without retraining, achieving a 4.2× kernel-level speedup over dense FlashAttention at 16K context with 90% sparsity (up to 8.5× at 95%), and achieves better model accuracy and lower latency than prior methods.
Abstract: Estimating causal effects from observational data requires identifying valid adjustment sets. This task is especially challenging in realistic settings where latent confounding and feedback loops are present. Existing approaches typically assume acyclicity or rely on global causal structure learning, limiting applicability and computational efficiency. In this work, we study a local, data-driven method for covariate selection based on conditional independence information. While this method is known to be sound and complete in acyclic causal models, its validity in the presence of cycles has remained unclear. Our main contribution is to show that these guarantees extend to cyclic causal models. In particular, our result relies on the invariance of conditional independence assertions under \sigma-acyclification. These findings establish a unified, cycle-agnostic perspective on covariate selection and causal effect estimation, showing that the method applies across cyclic and acyclic settings without modification. Empirically, we validate this on extensive synthetic data, showing reliable performance in cyclic causal models.
PaperID: 5885, Poster
Authors: Yirong Xiang, Yibo Zhou
Abstract: Frozen representations are often adapted using task labels alone, without environment, nuisance, or concept annotations. In this label-only regime, the first question is identifiability: what semantic target is actually determined by the joint law of representation and label? We show that labels identify only the posterior quotient, equivalently the coarsest deterministic representation that preserves all label information, and do not identify a finer semantic/background decomposition. This yields a quotient-first principle for label-only auditing and editing: first estimate a low-dimensional label-sufficient surrogate, then analyze residual class-dependent structure only after removing it. Under a structured linear second-order specialization, this leads to dual-moment decomposition (DMD). Empirically, stage one is a strong semantic audit and label-preserving compression baseline across frozen CelebA, Waterbirds, and PACS features, while stage-two residual editing improves leakage--label tradeoffs only in some regimes. Overall, the results support a conservative recipe for label-only frozen-feature editing: preserve the quotient first, and trust residual edits when residual support and split stability are strong.
PaperID: 5886, Poster
Abstract: Causal Bayesian optimization (CBO) typically relies on the unrealistic assumption of a perfectly known causal graph. Recent methods relax this constraint by coupling optimization with structure learning, but they often waste sample budgets on structural uncertainties that are irrelevant to the optimizer. In this paper, we formalize \emphrouting-equivalence to shift the objective from full causal-graph recovery to causal belief routing, where structural differences only matter when they alter optimizer-facing objects such as admissible exploration sets and \emphdo-priors. Furthermore, we propose \agentcbo, a graph-belief routing framework that incorporates LLM causal priors through guarded proposal and commit mechanisms. Specifically, a \textscLanguage-Prior Agent combines LLM beliefs with constraint-based discovery to extract a protected causal backbone alongside residual uncertain edges. A \textscCalibration Agent and a \textscResidual Repair Agent then refine only these uncertain structures through score-gated bounded local edits. Finally, a deterministic commit gate evaluates candidate updates using routing-proxy checks before allowing them to influence intervention selection. Across ten benchmarks, \agentcbo performs comparably to known-graph CBO, demonstrating that guarded LLM priors can improve fixed-budget CBO performance.
Abstract: Reinforcement learning (RL) post-training provides a direct way to align diffusion and flow-matching generators with human preferences and task-specific rewards, but current RL algorithms for diffusion models remain fragmented: reverse-trajectory methods rely on discretized likelihood ratios, whereas forward-matching methods train on reward-labeled noising pairs. This paper shows that these seemingly different losses arise from a single path-space principle. Starting from the regularized diffusion-RL objective, we use importance sampling between sampling SDEs to obtain an explicit policy-gradient estimator on trajectory space. The estimator contains the stochastic It\^o integral underlying Flow-GRPO-type updates; we derive an equivalent variance-reduced value-gradient form that recovers the forward-matching structure of AWM and DiffusionNFT. This identifies the empirical gap between these method families as a variance-reduction effect rather than a difference in RL principle. The derivation yields a unified design space organized by value-gradient estimation, weight functions, and sampling choices. Within this space, we propose a multi-sample KDE value-gradient estimator that reuses rollout groups, together with scale-bounded weight families that retain stable existing recipes while excluding singular ones. Experiments on SD3.5-M and Qwen-Image models validate the variance-reduction explanation and show that the resulting recipe matches or improves over prior diffusion-RL baselines.
Abstract: Diffusion Transformers (DiTs) have emerged as a core architecture in generative modeling due to their scalability and adaptability to multimodal tasks. DiTs comprise isotropic transformer blocks, and learn representations progressively across depth, where the denoising objective drives later layers to focus on fine-detail reconstruction. This results in degraded representation quality and an imbalanced encoder--decoder behavior. Prior approaches such as representation alignment (REPA) mitigate this by encouraging stronger early representations via training regularization. Alternatively, U-Net-style DiT architectures introduce explicit multi-scale encoder–decoder structures for improved convergence. But they build on standard U-Net wisdom via learnable operators for spatial downsampling, which are not well-suited to transformer architectures, introducing inefficiencies and compatibility issues with components such as cross-attention and representation regularization. In this work, we propose UDT, a U-Net diffusion transformer that combines the representation power of DiTs with the encoding–decoding benefits of U-Nets, through data-adaptive token merging for downsampling and upsampling, while preserving the DiT token dimension. Our baseline UDT architecture outperforms existing U-Net DiTs and achieves performance comparable to REPA across all model sizes. Furthermore, using architectural optimization and REPA, UDT outperforms SiT's 7.9 FID at 1400 epochs (w/o CFG) within 40 epochs (\approx 40× faster convergence) for XL model size. Finally, it achieves strong image generation results on ImageNet with 1.44 FID at 256 × 256 (with CFG) after only 200 epochs, providing a new backbone for DiTs with strong empirical benefits.
Abstract: Near-infrared (NIR) video is a promising modality for contactless sleep monitoring, but recent video-based sleep staging methods often use it as a route to reconstructed respiratory/cardiac proxies or cross-modal physiological representations. We study video-only sleep staging under labels defined by polysomnography (PSG), where the model infers sleep stages from NIR video alone without explicit physiological proxy reconstruction or auxiliary physiological supervision. This tests whether NIR video itself can provide informative sleep-stage evidence, rather than only serving as an input for recovering physiological proxies. We propose ViNUSS (Video-Native Unmediated Sleep Staging), a framework that combines subject-relative micro-motion learning with full-night sleep dynamics modeling. Spatially anchored pre-spatial micro-motion encoding preserves localized temporal variation together with its spatial context. Within-subject stage contrast learns stage cues with respect to each subject's night-specific baseline. Two-scale sleep dynamics modeling captures within-epoch motion evolution and organizes epoch-level evidence into a coherent full-night sleep-stage trajectory. On 475 overnight NIR recordings (~3,250 hours), ViNUSS achieves 0.80 accuracy and 0.78 macro-F1 for four-class sleep staging. Interpretability analysis suggests attention to thoraco-abdominal periodic motion and gross body movements associated with arousals and position changes. These results support NIR video as an independently informative and complementary modality for PSG-defined sleep-stage estimation.
PaperID: 5890, Poster
Authors: Yanshuo Liu, Jiaheng Qu, Sha Cao, Chi Zhang, Nan Zhang
Abstract: Log-Gaussian Cox Processes (LGCP) are essential for modeling spatial point patterns but suffer from high computational costs and intractable likelihoods. We propose VoGCAM, an efficient Variational Voronoi Gaussian Coordinate Ascent Maximization framework for fitting LGCPs. Our approach first approximates the intractable integral in the LGCP likelihood using a flexible Voronoi tessellation, which incorporates observed points as integration nodes. By applying variational Gaussian approximation, we derive an Evidence Lower Bound (ELBO) that admits an explicit and closed-form expression. To optimize this objective, we develop a novel coordinate ascent algorithm that updates parameter blocks via Newton and fixed-point methods. We further enhance scalability by adopting a Nearest Neighbor Gaussian Process (NNGP) prior and utilizing the Woodbury formula to reduce matrix inversion costs. Theoretically, we prove the existence and uniqueness of the optimal solution and establish the convergence of our algorithm. Numerical experiments on synthetic and real-world data demonstrate that VoGCAM offers superior computational efficiency and inferential accuracy over state-of-the-art methods like INLA and VIFRK.
Abstract: Code agents must both reason over long-horizon repository state and obey strict tool-use protocols. In paired Instruct/Thinking checkpoints, these capabilities are complementary but misaligned. The Instruct model is concise and tool-disciplined, whereas the Thinking model offers stronger planning and recovery behavior but often over-deliberates and degrades agent performance. We present CRANE (Constrained Reasoning Injection for Code Agents via Nullspace Editing), a training-free parameter-editing method that treats the Thinking–Instruct delta as a directional pool of candidate reasoning edits for the Instruct backbone. CRANE combines magnitude thresholding to denoise the delta, a Conservative Taylor Gate to retain edits that are jointly beneficial for reasoning transfer and tool-use preservation, and Graduated Sigmoidal Projection to suppress format-critical update directions. By merging paired Instruct and Thinking checkpoints, CRANE delivers strong gains over either individual model while preserving Instruct-level efficiency: on Roo-Eval it achieves pass@1 of 66.2% (+19.5%) for Qwen3-30B-A3B and 81.5% (+8.7%) for Qwen3-Next-80B-A3B; on SWE-bench-Verified it resolves up to 14 additional instances at both scales (122/500 and 180/500); and on Terminal-Bench v2 it improves pass@1/pass@5 by up to 2.3%/7.8%, reaching 7.6%/17.9% and 14.8%/30.3%, respectively, consistently outperforming alternative merging strategies across all three benchmarks.
PaperID: 5892, Poster
Abstract: Translating high-level, under-specified human commands into coherent video is fundamentally challenging. Current video generation models lack explicit reasoning capabilities and typically fail to understand ambiguous instructions, frequently resulting in physically inconsistent and causally disjointed generation. To address this, we introduce AmbiguousWorld, a novel inference-time reasoning framework that transforms video generation into a multi-modal Tree of Thoughts. We employ a Vision-Language Model as both a semantic planner to dynamically decompose fuzzy instructions into actionable subgoals, and as a closed-loop critic to prune invalid trajectories by evaluating speculative paths across multiple physical and functional dimensions. To systematically assess this, we annotate the first embodied video generation benchmark targeting ambiguous instructions based on the DROID dataset, and design a multi-dimensional VLM evaluation framework. Empirical results on state-of-the-art baselines demonstrate that our framework yields substantial improvements, boosting absolute Task Success by up to 29.0% and the overall generative consistency weighted score by 24.2 points. Furthermore, we validate the effectiveness of our VLM evaluation framework through rigorous comparison with human assessment.
PaperID: 5893, Poster
Abstract: Adapting Multimodal Large Language Models (MLLMs) to specialized and diverse downstream applications necessitates fine-tuning, typically through parameter-efficient methods such as (LoRA). However, fine-tuning MLLMs severely degrades pre-trained upstream capabilities, a phenomenon known as . Existing Methods primarily rely on parameter pruning to strike a balance between the downstream adaptation and upstream knowledge retention. Empirical analysis indicates that these methods are subject to , wherein the imposition of downstream adaptation and upstream retention within a unified parameter space results in suboptimal outcomes for both objectives. To address this, we introduce Dual-Pronged LoRA, a novel fine-tuning framework for MLLMs featuring two components. (1) (TDCR) makes the first attempt to explicitly decouple tasks by modeling a downstream contextual support region. It adaptively activates the LoRA branch exclusively for downstream-relevant inputs and bypasses it otherwise, thereby preventing downstream adaptation from interfering with pre-trained knowledge. (2) (DPO-RS) maximizes downstream adaptation performance by dynamically sculpting ranks via gradient energy. It effectively strengthens critical layers while compressing redundant ones and is theoretically grounded in the minimization of information loss. Extensive experimental results indicate that Dual-Pronged LoRA achieves near-zero forgetting while attaining new state-of-the-art downstream performance.
PaperID: 5894, Poster
Authors: Tianyuan Cheng, Ruirui Mao, Judea Pearl, Ang Li
Abstract: Probabilities of causation (PoCs), such as the probability of necessity and sufficiency (PNS), are important tools for decision making but are generally not point identifiable. Existing work has derived bounds for these quantities using combinations of experimental and observational data. However, there is very limited research on sample size analysis, namely, how many experimental and observational samples are required to achieve a desired margin of error. In this paper, we propose a sample-size design framework for PoC bounds that can be expressed as finite minima or maxima of smooth functions of experimental and observational probabilities — a representation that covers binary PNS, PN, and PS, as well as representative multi-valued PoC families recently shown to subsume the discrete PoC bounds in the current literature. Our framework gives endpoint-specific sample-size formulas that account for the joint covariance structure of the probability estimates: a distribution-free rule for affine bounds, and a pilot-based plug-in rule with theoretical safety guarantees for ratio-type bounds. Simulations show that the proposed rules achieve the target precision and coverage, recover a much smaller size requirement for binary PNS than the existing calculation, and provide sample-size designs for other discrete PoCs where no prior general sample-size formula is available.
PaperID: 5895, Poster
Abstract: LLM-based agent systems are increasingly used to solve complex multi-step tasks, where sequential execution incurs substantial end-to-end latency. In principle, parallelizing work across multiple agents should yield near-linear speedups. However, in practice, existing parallel multi-agent systems often run slower than a single-agent baseline. We attribute this gap to two hidden costs that parallel execution incurs but a serial agent avoids. First, there is a \emphre-exploration cost: redundant effort spent by parallel workers reconstructing context that the orchestrator already possesses, including prior decisions, conventions, and intermediate reasoning that would otherwise be inherited implicitly in a serial execution. Second, there is an \emphalignment cost: the overhead required to reconcile inconsistencies across independently generated outputs. Based on this decomposition, we derive a principled decision criterion: a layer should be parallelized only when its critical-path cost, plus re-exploration and alignment overheads, is lower than the corresponding serial cost. While this criterion is naturally expressed in wall-clock time, we observe that LLMs are poorly calibrated when asked to estimate task duration. Their predictions are strongly anchored to human engineering intuition rather than model throughput. The resulting bias is not monotonic, so even the relative ordering of task costs can be reversed between estimates and actual execution. To address this, we instead measure cost in predicted output tokens, a quantity that LLMs can estimate more reliably because it corresponds directly to their own generation behavior. For a fixed model, token cost also serves as a backend-independent proxy for time. Building on this token-based criterion, we propose \textscCostPar. It estimates all token budgets in a single planning step, forks each worker directly from the orchestrator’s session to eliminate re-exploration cost, and replaces post-hoc reconciliation with a pre-generated shared convention block that converts alignment into a bounded upfront cost. A deterministic scheduler then applies the criterion layer by layer. Empirically, \textscCostPar achieves a 2.2× mean throughput improvement and a 2.6× mean wall-time speedup over Claude Code, and a 2.0× throughput improvement over the strongest multi-agent baseline.
PaperID: 5896, Poster
Abstract: In time-series systems, safe active learning seeks to collect informative data while maintaining safety during data collection. A key challenge is that one-step safety does not guarantee future viability: an input can be immediately safe yet lead the system into a state from which no safe continuation exists. For example, in pressure regulation, an input may keep the current pressure below its limit but still drive the system toward an unsafe pressure spike over the next few steps. We study this problem in nonlinear time-series systems with unknown dynamics and safety constraints. We introduce the notion of m-step dead-end inputs and propose a future-aware safe active learning framework (FA-SAL) that enforces finite-horizon safety via margin-certified rollout constraints. Our approach constructs a conservative approximation of the true future-safe set by combining Gaussian process uncertainty with a recursive error bound over predicted trajectories. We show that FA-SAL provides high-probability safety guarantees, recovers interior future-safe inputs as uncertainty decreases, and provably excludes separated dead-end inputs. Experiments on synthetic and real-world benchmarks demonstrate that FA-SAL reduces safety violations and dead-end selections while maintaining competitive model learning performance.
PaperID: 5897, Poster
Abstract: Liability adjudication tasks (e.g., legal charge determination), require analyzing event descriptions against domain-specific skills to determine whether a party should bear liability. While Large Language Models (LLMs) have shown promise in this area by leveraging executable skills for interpretable reasoning, existing approaches face three critical limitations: skill descriptions often suffer from a semantic execution gap where LLMs deviate from intended expert logic ptimization (KDASO), which continuously evolves a skill library through neighborhood-conditioned induction from recent adjudication data, execution-aligned refinement based on LLM trajectory analysis, and context-budgeted routing for scalable skill selection. Experiments across ride-hailing adjudication, legal liability determination, and logical reasoning judgment demonstrate that KDASO significantly outperforms current general-purpose and domain-specific baselines. KDASO further achieves an 8% improvement when transferring skills to concept-similar scenarios and maintains consistent performance gains through continuous optimization, underscoring the high transferability and stability of generated skills.
PaperID: 5898, Poster
Abstract: While Graph Neural Networks (GNNs) have proven highly effective at modeling relational data, pairwise connections cannot fully capture multi-way relationships naturally present in complex real-world systems. In response to this, Topological Deep Learning (TDL) leverages more general combinatorial representations--such as simplicial or cellular complexes--to accommodate higher-order interactions. Existing TDL methods often extend GNNs through Higher-Order Message Passing (HOMP), but face critical scalability challenges due to the steep complexity overhead of propagating messages through combinatorial structures. To overcome this limitation, we propose HOPSE (Higher-Order Positional and Structural Encoder), a framework \emphfree of message passing layers that uses Hasse graph decompositions to derive efficient and expressive encodings over \empharbitrary higher-order domains. Notably, HOPSE scales linearly with the size of combinatorial representations while preserving the expressive power and permutation equivariance of the HOMP approaches. Experiments on molecular and topological benchmarks show that it matches or surpasses state-of-the-art performance while consistently achieving speedups over HOMP-based models, opening a new path for scalable TDL. The code is available at https://anonymous.4open.science/r/TopoBench-BF89.
PaperID: 5899, Poster
Authors:
Ziyan Li, Xiangyu Hu, Tianqing Zhu, Wanlun Ma, Ke QinAbstract: Continuous-time dynamic graphs model evolving systems with irregularly timed interactions, where predicting future links requires reasoning over temporal evolution, local topology, and historical relation patterns. Existing memory-based and sequence-based methods capture temporal dependencies effectively, but they often organize historical interactions as node memories or serialized neighbor sequences. As a result, they lack explicit analysis of the interaction structure around candidate nodes which prevents them from explicitly modeling local topological patterns of candidate nodes at inference time. They also lack explicit causal statistical memory for activity, recency, and repeated pairwise interactions. Motivated by these drawbacks, we propose S3Former, a temporal graph framework that jointly models sequential dynamics, shared pair-centered structure, and online statistical memory. For each candidate interaction, S3Former encodes long-term and short-term historical sequences, constructs a history-induced shared pair subgraph around the two endpoints, and computes node-level and pair-level statistics from past events only. A two-level context-aware fusion mechanism further combines these signals: NodeCCF learns reliable endpoint representations, while PairCCF performs candidate-specific relation correction using explicit pair memory. Experiments on 10 dynamic graph benchmarks show that S3Former achieves state-of-the-art performance on most datasets in both transductive and inductive link prediction. Ablation studies confirm the effectiveness of shared pair-subgraph modeling, online statistical memory, and hierarchical fusion.
Abstract: Brain foundation models have achieved remarkable advances across a wide range of neuroscience tasks. However, most existing models are limited to a single functional modality, restricting their ability to exploit complementary spatiotemporal dynamics and the collective data scale across different neuroimaging techniques. This limitation largely arises from severe semantic heterogeneity and resolution discrepancies among modalities. To address these challenges, we propose Brain-OF, an omnifunctional brain foundation model jointly pretrained on fMRI, EEG and MEG, capable of handling both unimodal and multimodal inputs within a unified framework. To reconcile heterogeneous spatiotemporal resolutions, we introduce the Any-Resolution Neural Signal Sampler, which projects diverse brain signals into a shared semantic space. To further manage semantic shifts, the Brain-OF backbone integrates DINT attention with a Sparse Mixture of Experts, where shared experts capture modality-invariant representations and routed experts specialize in modality-specific semantics. Furthermore, to explicitly internalize the characteristics of neural activity through self-supervised learning, we propose Masked Temporal-Frequency Modeling, a dual-domain pretraining objective that jointly reconstructs brain signals in both the time and frequency domains. Brain-OF is pretrained on a large-scale corpus comprising around 40 datasets and demonstrates superior performance across diverse downstream tasks, highlighting the benefits of joint multimodal integration and dual-domain pretraining. Our code will be made publicly available upon acceptance.
Authors:
Feijie Wu, Weiwu Zhu, Yuxiang Zhang, Soumya Chatterjee, Jiarong Zhu, Fan Mo, Rong Luo, Jing GaoAbstract: Multi-tool-integrated reasoning enables LLM-empowered tool-use agents to solve complex tasks by interleaving natural-language reasoning with calls to external tools. However, training such agents from outcome-only rewards suffers from credit-assignment ambiguity, obscuring which intermediate tool-use decisions drive success or failure. In this paper, we propose PORTool, an importance-aware policy-optimization algorithm that reinforces agents' tool-use competence from outcome-level supervision while assigning reward at the step level. Specifically, PORTool generates a rewarded rollout tree in which trajectories share prefixes before branching, enabling direct comparisons among alternative tool-use decisions within the same context. It then estimates each step's importance by a correctness-dominant signal, i.e., whether descendants of that step can ultimately produce a correct final answer, plus an auxiliary term indicating whether the step's tool calls satisfy formatting constraints and execute successfully. Using these step-wise importance estimates, PORTool updates the policy to generate efficient tool-call steps, guided by both local comparisons within each branching decision and the overall quality of entire trajectories. Experiments show that PORTool improves final-answer accuracy while reducing tool-call steps compared with state-of-the-art policy-optimization baselines, and ablation studies confirm the robustness of the proposed step-wise importance estimates.
Abstract: We introduce a unified framework that reinterprets first order optimization on a smooth parameter space as a closed loop controlled dynamical system on a Riemannian manifold. Within this framework, optimizers are specified by a quadruple (M, g, \Phi, \eta) consisting of a Riemannian metric, an internal state transition map, a direction field generator, and a lifting parameter. Classical and modern methods including SGD, Adam, AdamW, Lion and Shampoo are recovered as specific instantiations of this quadruple. The central structural object of the framework is a time indexed family of target submanifolds in the extended state space of parameters, velocities, and internal memory, which we call the Normally Attracting Invariant Manifold (NAIM) family. These submanifolds organize training dynamics into two timescales: a fast contraction of the velocity state onto the current target submanifold, followed by a slow descent along it. Using Riemannian backstepping, we construct a strict Lyapunov function on the extended state space and prove Uniform Ultimate Boundedness (UUB) of the optimization trajectory under a descent alignment assumption and bounded stochastic disturbances. Under the additional Polyak-Lojasiewicz (PL) condition, the Lyapunov function converges linearly to a bounded neighborhood of the optimum, defined as the drift of the target direction field between consecutive iterates after parallel transport. To validate the framework's generativity, we derive three concrete instantiations of the quadruple, each exercising a distinct design axis, and show on geometric diagnostics, image classification benchmarks, and language modeling that their empirical behavior matches the theoretical predictions. The framework casts optimizer design as controller synthesis, complementing heuristic perspectives by providing a constructive template with provable guarantees.
Abstract: Many dynamical processes unfold on the sphere but the default scientific machine learning architectures are Euclidean. Applying these architectures on a regular lat--lon grid causes problems: Cartesian convolutions become distorted at high latitude; 2D FFTs in Fourier neural operators incorrectly assume double periodicity; Cartesian positional encodings in ViTs distort spherical geodesic distances. Recent work moves towards natively spherical primitives, including spherical convolutions (e.g., DeepSphere or DISCO), Spherical Fourier Neural Operators (SFNOs), and geodesic attention. Here we propose Dandelion, a spherical version of Flower---a recent warp-based neural PDE solver. Layers of Dandelion predict a tangent-plane displacement and transport features along great circles. We obtain a U-Net-like structure by implementing hierarchical pooling entirely in the spherical-harmonic domain. There are thus no convolutions: spatial mixing is achieved only through spherical coordinate changes---or warps. To compare Dandelion with existing spherical architectures, we release an evolving benchmark suite of challenging, natively-spherical PDE datasets including a modified Galewsky jet, anomalous chained turbulence, Cahn--Hilliard decomposition, spherical Riemann shocks, Held--Suarez dry atmospheric transport and global ocean dynamics. This new benchmark fills the gap in existing spherical datasets which are either too small and stylized, or much too large (ERA5) for model iteration. Dandelion is best or second-best on every dataset, and the gap to non-warp baselines widens with resolution: at 256× 512, Dandelion and Flower2D occupy the top two slots in both single-step prediction and rollout.
Authors:
Yuyang Zhang, Yifu Zhang, Xuehai Zhou, Xiaoyin ChenAbstract: While empirical scaling laws for LLM reasoning are well-documented, the theoretical mechanisms governing out-of-distribution (OOD) generalization remain elusive. We formalize reasoning via optimal transport, projecting discrete trajectories into a continuous metric space to quantify domain shifts using the Wasserstein-1 distance. Invoking Kantorovich duality, we bound OOD generalization via architectural Lipschitz continuity and functional approximation limits. This exposes two primary constraints. First, position-dependent attention (e.g., Absolute Positional Encoding) fails to preserve shift invariance, yielding an \Omega(1) Lipschitz constant and expected risk, whereas shift-invariant mechanisms (e.g., Rotary Embeddings) preserve equivariance and bound the error. Second, by mapping sequential backtracking to a Dyck-k language, we establish a strict circuit depth lower bound for \textTC^0 Transformers. Scaling physical layer depth is necessary to avert representation collapse---a constraint that scaling representation width cannot bypass due to irreducible approximation bounds in Barron spaces. Evaluations across 54 Transformer configurations on combinatorial search corroborate these bounds, demonstrating that generalization risk degrades monotonically with the Wasserstein domain shift.
PaperID: 5905, Poster
Abstract: We argue that hyperparameter optimization (HPO) is best approached not as a black-box search problem but as the inverse of an intervention: when training dynamics are observed, \lambda is the controlled cause and the task is what is invariant to it, and the cost of HPO is largely the cost of failing to separate the two. We propose CA-HPO (Configuration-invariAnt HPO), a transfer HPO framework that learns a task representation explicitly constrained to be invariant under hyperparameter interventions and predictively sufficient for performance. We define this target representation as the solution to a constrained variational problem rather than as an identified latent variable, so the framework avoids the strong identifiability assumptions that previous causal representation learning approaches demand. We derive three guarantees in this setting: a population-level result showing that minimizing a predictive loss together with an invariance penalty recovers the target representation, a finite-sample bound on the estimation error of the empirical minimizer, and a counterfactual prediction risk bound that decomposes into representation error, meta-generalization error, and aleatoric noise. Experiments on Llama-3-8B fine-tuning and twelve classical benchmarks show 5 to 15 times speedup over Bayesian optimization and recent meta-HPO baselines, while the learned representations remain stable across a 100-fold range of learning rates, transfer across domains with different hyperparameter spaces, and pass a direct conditional independence test against \lambda.
PaperID: 5906, Poster
Abstract: Contrastive Learning (CL) has been widely used for Graph Anomaly Detection (GAD). However, existing augmentation techniques in such CL frameworks usually rely on random masking or structural alterations, which obliterates the irregularities that define anomalies and leads to , which risks disrupting critical topological structures or introducing irrelevant information in the learned representations, thereby hardly capturing complex anomalies. In this paper, we propose Orthogonal-Hyperspherical Augmentation and Topology Perception ( topology perception module that leverages the high-quality representations learned from augmentation to screen abnormal dense substructures. Extensive experiments on benchmark datasets demonstrate that OHATP substantially outperforms state-of-the-art methods, achieving improvements of over 6% in AUROC and 29% in AUPRC. Code is available at
PaperID: 5907, Poster
Abstract: Image super-resolution (SR) with large generative models has recently achieved remarkable perceptual quality, yet maintaining fidelity to the LR observation remains challenging. In particular, we observe that diffusion transformers (DiTs) built on latent representations suffer from a critical limitation: the compression bottleneck of the VAE weakens fine-grained spatial information, leading to hallucinated details that are weakly grounded in the input image. In this work, we revisit generative SR from a representation perspective and propose a pixel-grounded super-resolution (PGSR) framework that preserves LR-observed pixel evidence before VAE compression and reuses it throughout restoration. Instead of relying solely on the compressed latent condition, PGSR extracts pre-VAE pixel evidence from the upsampled LR image and reuses it at two stages. First, Condition-Side Trajectory Guidance fuses LR-derived pixel evidence with the latent LR condition to guide the latent restoration trajectory. Second, Decoder-Side Pixel Grounding injects multi-scale pixel features into the frozen VAE decoder to ground the final rendering with LR-observed cues. To efficiently adapt large pretrained DiT models, we keep the latent autoencoder and main flow-matching backbone frozen, and train only lightweight restoration modules. We further study an efficient local-window attention variant for improved high-resolution efficiency and scalability. Extensive experiments demonstrate that PGSR improves the realism--fidelity trade-off and produces more faithful, visually convincing results than existing latent generative SR approaches.
PaperID: 5908, Poster
Abstract: Diffusion-based style transfer faces two fundamental questions: which regions should prioritize structure preservation and which are more amenable to texture transfer, namely the “where” question, and how to achieve such region-aware stylization without corrupting content, namely the “how” question. Existing methods typically rely on a single mechanism to control local structure, local texture, and global style together, making them fragile in region-rich scenes: uniform stylization strength may distort salient objects, while conservative control leaves texture-dominant regions under-stylized. We propose StyleRoute, a region-aware diffusion stylization framework that explicitly separates where style should be transferred from how the transfer should be optimized. For the where question, StyleRoute infers a region-level routing variable from structure confidence, texture confidence, and route disagreement, deciding per region whether to protect structure or to inject texture. For the how question, StyleRoute coordinates the regional routing with diffusion optimization to guide global style evolution, suppress style-content conflicts, and stabilize coarse layout. Together, StyleRoute transforms the implicit style-content trade-off into an explicit, interpretable, and region-aware transfer process. Experiments on diverse content-style pairs show consistent improvements over prior diffusion-based and region-matching stylization methods. Code will be released.
Abstract: Federated fine-tuning of large language models is commonly formulated as a parameter aggregation problem. However, even parameter-efficient methods require transmitting large collections of trainable weights, assume aligned architectures, and rely on white-box access to model parameters. As model sizes continue to grow and deployments become increasingly heterogeneous, these assumptions become progressively misaligned with practical constraints. We consider an alternative formulation in which collaboration is mediated through model behavior rather than parameters. Clients fine-tune local models on private data and exchange generated outputs on a shared, public prompt set. The server maps these outputs into a semantic representation space, forms a per-prompt semantic consensus, and returns pseudo-labels for further local fine-tuning. This formulation fundamentally changes the communication scaling of federated LLM fine-tuning. The amount of information exchanged depends only on the public prompt budget and the size of the communicated behaviors, independent of model size. As a consequence, the protocol naturally accommodates heterogeneous architectures and applies directly to open-ended text generation. We present a theoretical analysis and empirical results demonstrating that this approach can match strong federated fine-tuning baselines while substantially reducing communication by orders of magnitude (e.g., analytically by a factor of 1006 for Llama3.1-405B), as well as reductions in runtime and energy consumption. These results suggest that, for generative foundation models, behavior-level consensus provides a more appropriate abstraction for federated adaptation than parameter aggregation.
PaperID: 5910, Poster
Abstract: AI coding agents are increasingly embedded in real-world software development, collaborating closely with human developers while gaining broader access to codebases and tools. This creates a new attack surface: an agent can exploit human trust to sabotage development, for instance by inserting malicious code to accomplish a hidden side task. Most prior work studies AI sabotage in AI-only settings, paying limited attention to the role of human oversight in detecting and mitigating such malicious behavior. To address this gap, we conduct the first large-scale study of human oversight in AI coding sabotage. Over 100 participants collaborate with one of the four frontier models (Claude Opus-4.6, GPT-5.4, Gemini-3.1-Pro, and MiniMax-M2.7) on a long-horizon coding task lasting around five hours, designed to mimic real-world workflows. We find that 95% of developers fail to detect sabotage, and our analysis of participant feedback attributes this vulnerability to coding without review, plausible cover story, and overtrust in agents. We further test the effectiveness of a safety monitor in one condition: while the monitor reduces sabotage success, 56% of participants still accept the malicious code, ignoring its warnings. Drawing on participant feedback, we offer actionable suggestions for better monitor design. This work complements existing AI safety research and highlights an urgent need for human-centric safety mechanisms that account for human factors, particularly in long-horizon, real-world development settings.
Abstract: Language models (LMs) often hallucinate by committing to substantive answers when they should abstain. Existing methods detect hallucinations or guide abstention, but leave open how models internally decide to commit or abstain. We study this decision through mechanistic analysis, framing a class of hallucination as unsupported commitment: the model commits despite exhibiting signals of unanswerability. Using causal gating, we identify a Commit-Abstain Circuit (CAC), a sparse subset of attention heads and MLP sublayers that causally contribute to this decision. Across ten LMs (3B-14B) from five families and three benchmarks, the CAC exhibits a recurring accumulate-yet-undercorrect pattern: commitment-promoting components build up commitment in earlier layers, while abstention-promoting components act later as corrective signals that are often insufficient to overturn the accumulated commitment. Building on this finding, a lightweight policy trained on CAC activations improves decision accuracy by 12.2 points over the model's intrinsic commit-abstain margin, reduces false abstentions by 2.5 times, transfers to unseen benchmarks, and extends to larger models (27B-35B). The CAC is both diagnostic, clarifying how models overcommit, and practical, enabling improved abstention decisions. Code is provided in the supplementary material.
Abstract: As agentic coding systems decompose work across multiple model instances, a critical safety question is whether those instances can coordinate to achieve a hidden malicious objective while remaining aligned with user intent. We introduce \emphSCHEME, a benchmark of 17 task instances across 7 settings and 8 real open-source libraries, each pairing a legitimate software-engineering task with a covert side task. Every setting is designed so that no proper subset of agents can succeed alone: agents must decompose a shared sabotage plan, relay partial requirements under different communication topologies, and execute mutually consistent edits, testing genuine multi-agent coordination rather than individual capability. Evaluating with \emphGPT 5.1 Codex and \emphGemini 3.1 Pro, we find coordinated sabotage is already practical, with Gemini completing the covert objective while succeeding on the legitimate task in 84% of samples and Codex in 46%. A failure analysis shows that recovery, not failure incidence, drives the model gap: both models suffer similar numbers of communication failures per trajectory, but Gemini recovers from ~73% of occurrences against Codex's ~56%. However, the agents' code edits remain detectable: a trusted monitor flags 100% of Gemini and 81% of Codex sabotage attempts at FPR\le2%, even without access to agent communication. Even with stealth-aware prompting, undetected sabotage stays at 0--2.4%; telling agents a monitor is present halves Gemini's sabotage success without increasing the rate of undetected sabotage.
PaperID: 5913, Poster
Abstract: Software-engineering agents are increasingly expected to solve realistic repository-level tasks, but their training remains constrained by the limited availability of executable supervision beyond Python-centric repositories. Building such data across programming languages is difficult: repositories use different build systems, dependency managers, runtime assumptions, test interfaces, and failure modes, so separate language-specific pipelines are costly to reproduce and extend. We introduce SWE-Crafter, a unified agent-based framework for constructing executable multilingual SWE instances. SWE-Crafter keeps candidate mining, environment synthesis, fail-to-pass validation, and trajectory distillation in a shared pipeline, while language-specific construction knowledge is supplied through extensible skills induced from repository-level experience. These skills guide repository exploration and test-command synthesis while preserving repository-local evidence from documentation, manifests, CI files, and execution feedback as the final source of truth. Applying SWE-Crafter to eight programming languages yields 50,926 validated instances from 9,296 repositories and 68,851 distilled test-passing trajectories. Supervised fine-tuning on this resource produces SWE-Crafter-40B, which achieves 61.0% on SWE-bench Multilingual and 77.2% on SWE-bench Verified, showing that skill-guided construction provides scalable multilingual supervision for repository-level SWE agents.
PaperID: 5914, Poster
Abstract: Recent advances in robotics foundation models have shown promising multi-task generalization, yet their reliance on specialized embodiment-specific data and model fundamentally prevents them from exploiting the far richer physical interaction knowledge present in abundant human manipulation videos. Existing approaches that leverage human videos suffer from misaligned representations that preclude explicit unified interaction modeling over human video and robot demonstration, failing to capture unified physical interaction dynamics essential for cross-domain transfer. We address this by constructing a unified representation that aligns human videos and robot tasks through scene point clouds with hand-to-gripper mapping via spatial tracking and hand pose estimation. Leveraging this representation, we propose a structured graph transformer that explicitly models spatial, semantic, and intentional interaction entities through graph attention mechanisms. The modeled multi-level interaction features then enable interaction-aligned cross-domain learning, transferring manipulation knowledge from human videos to robot tasks via discrete interaction abstraction and hierarchical distribution alignment. Experiments on RLBench, ManiSkill2, and real-world robot demonstrate state-of-the-art performance on diverse benchmarks with fewer demonstrations and effectiveness of our designed modules. Remarkably, our method achieves zero-shot transfer on robot tasks when trained exclusively on human video, providing strong evidence for effective human-to-robot transfer through interaction alignment.
PaperID: 5915, Poster
Abstract: Deploying large language models on edge devices requires mitigating two major bottlenecks: the computational cost of the model and the memory overhead of the KV cache. While combining model quantization and KV-cache eviction seems like a natural solution, we reveal that this naive pipeline suffers from a critical failure mode. The issue stems from a hidden conflict: modern eviction methods rely on attention scores to identify important tokens, yet quantization inherently adds noise to these exact scores. Consequently, quantization does not merely degrade the numerical precision of the KV cache—it actively alters the eviction decisions, often causing the model to discard crucial context. To overcome this non-orthogonal degradation, we propose EAT, a training framework that integrates discrete eviction behavior into quantization-aware training. EAT computes eviction from quantized attention, applies the resulting mask in the forward pass, and uses a masked straight-through estimator in the backward pass. The method requires no architectural changes and adds no inference-time parameters. Systematic evaluations show that EAT mitigates the failure mode of naive quantization-plus-eviction pipelines. Under W4A16 + KV-INT8 with 50% KV eviction, eviction-unaware composition reduces the five-task average from 38.37 to 29.04, whereas EAT recovers it to 39.44 under the same cache budget. Real-world deployment on the MediaTek Dimensity 9500 demonstrates that, under a 50% KV-eviction budget, EAT improves decoding throughput by 64.1% and reduces total memory traffic by 20.4% compared to quantized full-KV inference. Mechanistic analyses suggest that the improvement is associated with reduced attention sink behavior rather than with sharper attention.
Abstract: Supervised fine-tuning (SFT) provides the standard approach for teaching LLMs new behaviors from offline expert demonstrations. However, standard SFT uniformly fits all samples—including those with low likelihood under the base model—which can disproportionately drive training updates toward overfitting specific samples rather than learning the target behavior. Moreover, adapting to these unlikely samples induces substantial policy shifts that degrade prior capabilities. Existing methods mitigate this by filtering, regenerating, or down-weighting low-likelihood data. In doing so, they often suppress precisely the novel behaviors the base model has yet to learn. We propose InfoSFT, a principled weighting scheme for the SFT objective that concentrates learning signals on maximally informative, medium-confidence tokens—those neither overly familiar to the base model nor too unlikely to cause instability. Requiring only a one-line modification to the standard token-wise loss, InfoSFT demonstrably improves generalization over vanilla SFT and likelihood-weighted baselines across math, code, and chain-of-thought tasks with diverse model families, while better preserving pre-existing capabilities.
PaperID: 5917, Poster
Authors: Yasser Taha, Grégoire Montavon, Nils Körber
Abstract: While machine learning (ML) architectures have evolved rapidly to account for complex data, loss functions like cross-entropy remain mostly structure-agnostic in many real-world applications. However, the "class-symmetric" nature of these standard losses fundamentally limits the ability of ML models to exploit structural relationships between classes, particularly when facing structured noise. We propose Conveyance, a new classification approach and associated loss function tailored to structured class spaces. It allows users to encode graph-like relations between classes without having to define complex joint distributions or manually tune utility matrices. Technically, our loss function operates by maximizing two separate margins over distinct class partitions, while preserving formal properties such as monotonicity and partial convexity. We demonstrate the versatility and effectiveness of our method by applying it to hierarchical classification, ordinal regression, and multiple instance learning. Across these tasks, Conveyance either matches or exceeds the performance of specialized baselines, thereby offering a unified solution for structured class spaces.
PaperID: 5918, Poster
Authors: Abdelrahman Sharafeldin, Yaqing Wang, Simon Sponberg, Hannah Choi
Abstract: Animals integrate information across sensory modalities to guide exploration, suggesting that they maintain internal estimates of where actions are expected to reduce uncertainty. However, how such action-conditioned information maps are learned from multimodal sensory experience remains unclear, as previous studies have been largely confined to unimodal sensing frameworks. We introduce a multimodal predictive-coding framework for learning uncertainty-reduction maps from visual and mechanosensory observations in hawkmoth flower interaction tasks. A variational generative model trained on simulated environments recovers modality-specific information landscapes over the action space, capturing how visual patterns, surface geometry, and mechanosensory cues shape the value of exploratory actions. These learned maps account for several features of hawkmoth behavior, including angled offsets during flower tracking, probing along visual patterns, and geometry-dependent changes in exploration. Applying the model to behavioral data from flower tracking and nectary search, we find that natural trajectories tend to occupy regions of high uncertainty reduction in these learned maps. Building on these observations, we develop an uncertainty-gated lexicographic algorithm that switches between information gathering and reward exploitation based on perceptual uncertainty. Across downstream reward-learning tasks, this policy learns faster, is more sample-efficient, and generalizes better to unseen environmental conditions than exploitation-only policies, \epsilon-greedy exploration, and non-lexicographic baselines. Together, these results suggest that multimodal generative perception models can learn actionable information landscapes that explain biological sensing and support exploration in embodied decision-making.
Authors:
Zhenhe Cui, Huaxiang Xia, Hangjun Shen, Kailun Luo, Yong He, LiangWeiAbstract: Generalized planning (GP), where a single policy solves multiple instances from a common planning domain, remains a challenging problem in AI. Abstraction plays an important role in GP, and Qualitative Numerical Planning (QNP) provides a compact abstract model for GP. However, useful QNP abstractions are difficult to construct and require substantial manual effort and expertise. We study whether QNP abstractions can be generated from PDDL domains and training instances by coupling LLM generation with symbolic reasoning. We propose a generate--debug--repair framework in which an LLM proposes abstract features and constructs a candidate QNP abstraction over the initial states, actions, and goals of the training instances. To improve reliability, we introduce verification-guided repair based on approximate soundness. Given a candidate abstraction, we solve it, validate the induced abstract instances, and apply one of two checking branches: SRCB tests an m-simulation relation between induced abstract and concrete instances to target approximate soundness, while PRCB tests whether the induced abstract plan refines into valid concrete solutions. Detected failures are converted into structured feedback for iterative repair. Experiments on seven GP benchmark domains with five LLMs show that verification-guided repair improves LLM-generated QNP abstractions over both single-pass generation and LLM self-repair, and helps some models produce abstractions that transfer across instances. Our results suggest that LLMs can serve as generators of symbolic abstractions for GP when coupled with formal error detection and structured repair feedback.
PaperID: 5920, Poster
Abstract: Learning expressive representations for Boolean logic circuits is a fundamental challenge at the intersection of graph learning and Electronic \mboxDesign Automation (EDA). Existing Graph Neural \mboxNetworks (GNNs) primarily rely on topological message passing, which often fails to capture the strict causal dependencies and discrete functional semantics of logic gates. In this paper, we \mboxpropose FuncFormer, a Graph Transformer that incorporates functional simulation not as a proxy task, but as a fundamental inductive bias directly into the representation learning process. Unlike standard GNNs which typically rely on isotropic aggregation, FuncFormer encodes the intrinsic flow of functional propagation by analyzing randomized simulation traces as they evolve through the network. This approach effectively aligns the continuous embedding manifold with the discrete Boolean function space, effectively mitigating structural aliasing. By integrating these deterministic signal trajectories with a scalable dual-path attention mechanism, our model preserves functional consistency across long-range dependencies in both combinational and sequential circuits. Empirical results demonstrate that FuncFormer significantly outperforms state-of-the-art models (e.g., DeepGate4) in Quality-of-Results (QoR) prediction and formal verification tasks, exhibiting robust generalization to unseen circuit scales.
Abstract: Automating optimization modeling from natural language faces two key challenges: training corpora lack structural diversity, and data generation pipelines remain static and decoupled from model learning. To address these challenges, we propose EvoOptiGraph, a novel framework where data and model co-evolve, driven by model weaknesses. EvoOptiGraph represents each mixed-integer linear program (MILP) as an attributed bipartite graph and applies validity-preserving evolutionary operators to generate structurally diverse instances. The evolved graphs are converted into solver code and natural language via deterministic compilation and verified back translation. Training proceeds in two stages: supervised fine-tuning (SFT) on an initial dataset, followed by reinforcement learning with verifiable rewards (RLVR), where graph-derived weakness signals dynamically evolve new instances that target the model's failures, forming a closed loop that continuously updates the training distribution. Empirical results on six public datasets show that EvoOptiGraph significantly outperforms larger generalist models, agentic methods, and specialized baselines in accuracy, executability, and generalization. These results demonstrate that targeted data–model coevolution is an effective strategy for improving LLMs on optimization modeling tasks.
Abstract: While Proximal Policy Optimization (PPO) demonstrates strong performance in stationary settings, we show that its standard optimization paradigm struggles in continual and non-stationary environments. The failure does not stem from insufficient model capacity or overly restrictive clipping. Instead, PPO performs persistent, directionally inefficient local updates, which indicates a lack of geometry-aware guidance for accumulating meaningful behavioral change and ultimately hindering transitions toward new behavior patterns. Although divergence-based regularization introduces partial geometric awareness, its monotonically increasing penalties implicitly discourage large policy deviations, even when such shifts are necessary for effective adaptation. To address this limitation, we propose ), which reshapes the trust region using a Gaussian kernel. The resulting constraint is bounded and non-monotonic, providing strong local stability while progressively relaxing under sustained high-advantage updates. To further improve robustness, we introduce a is architecture-agnostic and achieves strong performance across games, simulated robotic control, open-world exploration, and language model post-training. These results demonstrate that geometry-aware trust-region design can be a promising direction for robust reinforcement learning in complex non-stationary environments.
Abstract: Vision-Language-Action (VLA) models have advanced rapidly with stronger backbones, broader pre-training, and larger demonstration datasets, yet their action heads remain largely homogeneous: most directly predict action commands in a fixed world coordinate frame. We propose MCF-Proto, a lightweight action head that equips VLA policies with a Motion-Centric Action Frame (MCF) and a prototype-based action parameterization. At each step, the policy predicts a rotation R_t \in SO(3), composes actions in the transformed local frame from a set of prototypes, and maps them back to the world frame for end-to-end training, using only standard demonstrations without auxiliary supervision. This simple design induces stable emergent structure. Without explicit directional labels, the learned local frames develop a stable geometric structure whose axes are strongly compatible with demonstrated end-effector motion. Meanwhile, actions in the learned representation become substantially more compact, with variation captured by fewer dominant directions and more regularly organized by shared prototypes. These structural properties translate into improved robustness, especially under geometric perturbations. Our results suggest that adding lightweight geometric and compositional structure to the action head can materially improve how VLA policies organize and generalize robotic manipulation behavior. An anonymized code repository is provided in the supplementary material.
PaperID: 5924, Poster
Abstract: Single-loop minimax methods are appealing for modern machine learning, but standard stochastic gradient updates can be unstable and oscillatory under tightly coupled primal--dual dynamics, making performance notoriously sensitive to stepsizes. To fundamentally improve stability, we develop an orthogonal update framework tailored to minimax optimization, which applies updates along orthogonalized directions to reduce update anisotropy. However, orthogonalization reshapes the optimization geometry and makes coordinating primal--dual stepsizes even more delicate. To overcome this challenge, we develop AdaSGDA, which equips orthogonal updates with an adaptive stepsize rule driven by accumulated gradient norms. This mechanism automatically balances coupled primal--dual progress under orthogonalization, yielding an effectively parameter-free method for nonconvex-strongly-concave minimax optimization. We prove that AdaSGDA avoids problem-dependent tuning while attaining the state-of-the-art \mathcalO\left(T^-1/4\right) convergence rate. We further propose AdaSGDA-VR, which incorporates a carefully designed variance reduction scheme and achieves the faster \mathcalO\left(T^-1/3\right) rate. Extensive experiments demonstrate that our methods deliver competitive performance compared to existing minimax optimization methods.
PaperID: 5925, Poster
Authors: Mahesh Keswani, Parth Bhardwaj, Ankur Kumar, Raunak P Bhattacharyya
Abstract: Offline safe reinforcement learning aims to maximize cumulative reward while satisfying safety constraints using only static datasets. Recent approaches leverage latent variable models to capture the behavior policy distribution, enabling efficient policy extraction under limited data. However, existing latent-space methods typically rely on soft constraints, which allow constraint violations in expectation and may be unsuitable for safety-critical settings. In addition, these methods often employ Advantage-Weighted Regression (AWR) for policy extraction, which exhibits exponential sensitivity to out-of-distribution critic bias, leading to high variance in the policy gradient estimate. To address these limitations, we propose Feasibility-Optimized Conditional Actor Learning (FOCAL). FOCAL enforces state-wise hard safety constraints by integrating reachability-based feasibility constraints directly into the Conditional Variational Autoencoder. To overcome the instability associated with AWR, we introduce a feasibility-partitioned, behavior-regularized objective that utilizes stable, first-order gradients to structurally decouple reward maximization in feasible regions from constraint violation minimization in infeasible regions. Finally, an adaptive test-time inference dynamically modulates latent sampling variance based on state-wise feasibility to ensure safe policy execution under out-of-distribution states. Theoretical analysis demonstrates that FOCAL circumvents the exponential variance in policy gradient estimates associated with AWR under critic errors, while providing formal bounds on policy deviation from the behavioral dataset. Empirically, on the DSRL benchmark, FOCAL satisfies safety constraints while achieving competitive or higher returns compared to state-of-the-art methods, all while maintaining efficient single-step inference.
PaperID: 5926, Poster
Authors:
shang xue, Li Zhu, Ping ChenAbstract: Feature caching accelerates Diffusion Transformers by reusing intermediate computations across denoising timesteps, yet existing methods share a fundamental blind spot: they treat the standard reference denoising trajectory as the optimisation target, approximating its per-step states and thus inheriting its output as an implicit quality ceiling. We argue that this target is arbitrary --- what truly matters is the ideal result, not the intermediate path taken to reach it. We propose IdealCache, a training-free framework that abandons step-wise imitation of the reference trajectory. Instead, it anchors caching on a per-input ensemble of ideal trajectories constructed from candidate target results, aggregates their geometry into a single graph, and discovers optimal non-uniform schedules via a constrained shortest path on this graph. By decoupling from the reference process, IdealCache breaks the performance bottleneck of trajectory-approximation; our schedules can attain near-lossless fidelity in generation and even surpass on task-specific metrics (e.g., maintaining higher semantic consistency in image editing). Without retraining, IdealCache achieves 6.35× speedup on Qwen-Image, 5.84× on Qwen-Image-Edit, and 4.32× on HunyuanVideo, with competitive or superior quality across tasks.
PaperID: 5927, Poster
Abstract: Grid cells in the medial entorhinal cortex and place cells in the hippocampus together support spatial navigation. The two regions are reciprocally connected, and there is a chicken-and-egg problem for how both arise and reinforce each other during development. Current computational accounts either derive one type from the other or use network dynamics to model the emergence of one type in isolation. We introduce a unified recurrent network model that instantiates Dale's Law (every neuron is either excitatory or inhibitory), and is trained to predict the next sensory observation from masked previous sensory observations and egocentric motion. To our knowledge, this is the first single-objective model in which grid and place cells co-emerge without supervision of either type, or reliance on pre-existing spatial-cell representations. The two kinds of spatial codes coexist across 1,000 different training configurations, with their balance set by the amount of sensory noise and masking. Without retraining, the network qualitatively reproduces experimentally observed grid fragmentation in hairpin mazes, grid merging after wall removal, lattice alignment across connected rooms, locally ordered 3D fields observed in freely flying bats, as well as the developmental order in which place cells precede grid cells. We interpret these results in terms of two complementary encoding pressures within a single sensory-prediction objective: (1) correcting errors or reconstructing missing components of sensory observations, and (2) prediction of the next sensory state during navigation. Our results suggest a circuit-level account of the co-emergence of grid and place cells, and experimentally testable predictions for the two kinds of spatial codes.
Abstract: Learning-rate transfer is critical for reducing the cost of training large language models: instead of sweeping learning rates at target scale, practitioners extrapolate from smaller runs. Existing approaches often assume that the optimal learning rate follows a log-linear scaling law in data scale and model size. We carefully examine and evaluate this scaling law. In our empirical study of GPT-2--style models from 22M to 407M parameters trained on 5B to 100B tokens, the optimal learning rate develops upward curvature at larger scales, leading to inaccurate extrapolation. We find that this curvature largely disappears when learning rates are replaced by effective learning rate (the step size in normalized weight space), and when data D extrapolation is used instead of model size N extrapolation. Next, we explain nonlinearity in scaling: weight-norm converges to equilibrium slower when optimal learning is small, requiring a larger step size to reduce the transient phase. Experiments with AdamH, which directly controls the effective learning rate, further support this explanation.
PaperID: 5929, Poster
Abstract: Replay is a standard strategy for mitigating catastrophic forgetting in continual learning, where a bounded memory buffer stores exemplars from previously seen classes. Existing replay methods have largely focused on instance-level decisions, such as which samples to store, replay, or replace, while the class-wise allocation of memory slots is often treated as a uniform heuristic. Here, we study class-wise memory allocation as a complementary design axis and propose VAMA, a value-aware memory allocation framework. VAMA keeps cross-task budgets balanced, estimates task-local class values from Shapley-inspired sample utility scores, and converts them into class-wise quotas through a regularized allocation rule that conservatively deviates from class-balanced replay. This design changes only the class-wise budgeting rule and leaves within-class exemplar selection unchanged, making VAMA easy to integrate into replay pipelines. Experiments on CIFAR-100 and ImageNet-100 show that VAMA consistently improves ER, iCaRL, and FOSTER across different memory budgets and task protocols. Further analyses show that the improvements are not due to non-uniform allocation alone, and that allocation-relevant class-value rankings stabilize early in training, enabling efficient value-aware allocation. These results suggest that class-wise memory allocation is a consequential design choice in replay-based continual learning.
PaperID: 5930, Poster
Authors: Yuji Nozawa, Yu-Chieh Lin, Youyang Ng
Abstract: Unsupervised learning in Vision Transformer models, such as DINO, enables robust visual perception without requiring labeled data by leveraging large-scale image datasets. However, these models lack a focusing mechanism during training, resulting in perceptual representations that are often ambiguous or overly similar. In this work, we introduce two visual focusing loss functions designed to establish correspondences between the model's attention and specific regions within images. Specifically, we leverage the multi-head attention mechanism within the model to selectively steer the model's focus, leading to the emergence of multiple, diverse, and focused perceptual capabilities without requiring supervision. Through both qualitative and quantitative evaluations, we demonstrate that our method substantially increases the spatial selectivity and diversity of attention heads, improving both explainability and fine-grained recognition performance.
PaperID: 5931, Poster
Authors: Abhishek Kumar Sinha, Nitant Dube, Soma Biswas
Abstract: Long-Tailed Class-Incremental Learning (LT-CIL) is challenging due to severe class imbalance, where frequent head classes dominate gradient updates and suppress learning for rare tail classes. Beyond data imbalance, we identify an asymmetric likelihood signals within each session: tail classes contribute insufficient evidence to overcome prior regularization, causing their parameters to remain near the prior mean even under variational inference. This asymmetry disrupts the stability-plasticity tradeoff, leading to pronounced forgetting and poor tail-class performance. We propose a hierarchical variational framework, TailAdapt, for prompt and adapter-based continual learning that enables adaptive group-wise regularization of model parameters. TailAdapt employs a hierarchical inverse Gaussian scale-mixture prior, which induces a heavy-tailed marginal distribution over prompt and adapter parameters. This formulation encourages selective plasticity so that most parameter groups are strongly regularized to preserve previously learned knowledge, while a small, data-supported subset is allowed to adapt substantially. This selective plasticity allocates adaptation capacity where it is most needed, mitigating interference from head classes and improving learning for tail classes. Extensive experiments on LT-CIL benchmarks demonstrate consistent improvements over strong baselines, with better tail-class performance, reduced forgetting, and more efficient utilization of prompt and adapter capacity.
PaperID: 5932, Poster
Authors: Emile Richard
Abstract: Reasoning, planning, and agentic systems demand that neural models extrapolate from the easy training instances to harder, more compositional problem instances they encounter at deployment. Neural networks trained by gradient descent are reliable interpolators \emphwithin their training domain but struggle to extrapolate beyond it: they learn local generalizers even when global solutions are representable in their architecture class. We propose a \emphlearn locally, recurse globally paradigm-: train a neural model on in-distribution sub-problems and apply it iteratively under a verifiable outer loop that decomposes hard instances into a sequence of in-distribution sub-instances and reassembles the result. We instantiate this paradigm in Boolean circuit synthesis from truth-tables, which we adopt as a faithful, controllable abstraction of logical reasoning: any propositional inference reduces to evaluating a Boolean circuit, and circuit \emphdepth is a clean proxy for the difficulty of the underlying reasoning task. Empirically, a 2.2M-parameter encoder-decoder Transformer trained on depth-\leq 2 circuits, wrapped in our outer loops, synthesizes circuits at 5× training depth and at input arities for which the base model's encoder is too small for direct inference, with synthesis quality matching a standard EDA tool while producing markedly shallower circuits. We connect the algorithms to classical Boolean decomposition theory (Shannon, Ashenhurst-Curtis, Reed-Muller) and to recent formal results on the AC^0 barrier for constant-depth Transformers.
PaperID: 5933, Poster
Authors: Christopher M Pierce
Abstract: In conditioned generative models of physical systems, the densities at different values of the conditioning parameters are often related through symplectic transformations. For stable linear Hamiltonian systems, there exists a symplectic transformation to ``normal coordinates'' such that the dynamics is reduced to rotations in each phase-space plane. This picture extends to the nonlinear case where, away from resonances, the Birkhoff normal form again consists of rotations in phase space, now with amplitude dependence. In this work, we introduce an invertible symplectic flow layer consisting of amplitude-dependent rotations in learned normal coordinates. We derive the condition under which the nonlinear rotation layer is symplectic and detail an architecture that guarantees it. Experiments are carried out on two conditional generative modeling problems from accelerator physics. We show empirically that normalizing flows with the new layers achieve competitive or superior negative log-likelihoods compared to baseline flows, with gains particularly pronounced in weakly nonlinear systems.
Abstract: Operator learning has been highly successful for continuous mappings between infinite-dimensional spaces, such as PDE solution operators. However, many operators of interest—including differential operators—are discontinuous or set-valued, and lie outside classical approximation frameworks. We propose a paradigm shift by formulating approximation via graph convergence (Painlevé–Kuratowski convergence), which is well-suited for closed operators. We show that uniform and L^p approximation are fundamentally inadequate in this setting. Focusing on maximally monotone operators, we prove that any such operator can be approximated in the sense of local graph convergence by continuous encoder–decoder architectures, and further construct structure-preserving approximations that retain maximal monotonicity via resolvent-based parameterizations.
Abstract: Unified multimodal models are envisioned to bridge the gap between understanding and generation. Yet, to achieve competitive performance, state-of-the-art models adopt largely decoupled understanding and generation components. This design, while effective for individual tasks, weakens the connection required for mutual enhancement, leaving the potential synergy empirically uncertain. We propose to explicitly restore this synergy by introducing ), a lightweight framework that treats understanding not only as a distinct task, but also a direct supervisory signal to steer generative representations. By incorporating objectives that encode semantic abstraction (captioning) and structural details (visual regression), we enable effective gradient flow from understanding to generation. Extensive experiments on image generation and editing demonstrate that understanding can serve as an effective catalyst for generation.
PaperID: 5936, Poster
Authors: Adrian Kucia, Edward Rollason, Wai Lok Woo
Abstract: Sparse-gauge flood forecasting is challenging because river-level dynamics combine transport, storage, rainfall forcing, and sharp event-scale phase shifts. We introduce GaugeCast++, a physics-informed latent forecasting framework that learns a coordinate system in which flood dynamics become easier to evolve. Its key component is a gauge-conditioned pullback operator, which reparameterizes the forecast field and suppresses residual transport, yielding a more stable latent evolution than direct prediction in observation space. To address delayed or shifted flood peaks, GaugeCast++ adds semigroup-consistent phase-transport learning, enforcing agreement between long-horizon forecasts and composed short-horizon rollouts in both amplitude and event phase. Finally, a decoupled inverse rainfall ambiguity module estimates uncertainty caused by imperfect future rainfall forcing. Experiments across six UK river-gauge catchments show consistent gains in event-level MSE, NSE, peak timing, and peak magnitude against strong neural forecasting baselines.
Abstract: Existing multimodal retrieval benchmarks largely emphasize semantic matching on daily-life images and offer limited diagnostics of professional knowledge and complex reasoning. To address this gap, we introduce ARK, a benchmark designed to analyze multimodal retrieval from two complementary perspectives: (i) knowledge domains (five domains with 17 subtypes), which characterize the content and expertise retrieval relies on, and (ii) reasoning skills (six categories), which characterize the type of inference over multimodal evidence required to identify the correct candidate. Specifically, ARK evaluates retrieval with both unimodal and multimodal queries and candidates, covering 16 heterogeneous visual data types. To avoid shortcut matching during evaluation, most queries are paired with targeted hard negatives that require multi-step reasoning. We evaluate 26 representative text-based and multimodal retrievers and observe a pronounced gap between knowledge- and reasoning-intensive retrieval, with fine-grained visual and spatial reasoning as persistent bottlenecks. We further show that enhancements such as re-ranking, rewriting, and agentic retrieval yield consistent gains, but substantial headroom remains. The dataset and evaluation pipeline will be released soon.
Abstract: Graph learning research has increasingly shifted toward continual graph learning (CGL), which better reflects real-world scenarios where graphs evolve over time. However, existing CGL methods largely assume clean supervision and overlook a critical challenge: the newly arriving portions of the graph are often noisy, due to annotation errors or adversarial corruption. This mismatch limits their applicability in practice. In this work, we study robust continual graph learning, where models must simultaneously handle catastrophic forgetting and noisy supervision in evolving graph data. We show that label noise introduces a new failure mode—catastrophic remembering, where models persistently reinforce corrupted knowledge across tasks. To address these challenges, we propose a Unified Flow-Oriented framework (UFO). First, UFO models conditional feature distributions via flow-based generative modeling and produces replay representations, mitigating forgetting without storing historical data. Second, UFO estimates instance-level reliability scores to distinguish clean from noisy nodes, reducing the impact of corrupted supervision and alleviating catastrophic remembering. Extensive experiments on four benchmark graph datasets under varying noise ratios demonstrate that UFO consistently outperforms existing methods in both accuracy and forgetting metrics. Code is available at: https://anonymous.4open.science/r/UFO.
PaperID: 5939, Poster
Authors: Yiran Pang, Zhen Ni, Dimitris Pados, Xiangnan Zhong
Abstract: Multi-task reinforcement learning (MT-RL) trains a single agent to solve multiple tasks by leveraging shared knowledge across task domains. However, most MT-RL methods assume centralized access to task data, making them impractical for privacy-sensitive or large-scale distributed settings. Multi-task federated reinforcement learning (MT-FRL) mitigates this issue by enabling agents to collaborate through shared models rather than raw trajectories, but often suffers from unstable performance under task heterogeneity, negative transfer, and imbalanced learning across tasks. A fundamental question in MT-FRL is how each agent should decide with whom to share skills and how strongly to rely on shared knowledge. To answer this question, we propose GAdapHeadFRL, a personalized MT-FRL framework based on geometry-aware grouping and adaptive decision-layer task-head mixing. First, we decompose the local Bellman residual into representation and task-head estimation errors, and derive a closed-form optimal sharing strength governed by task mismatch and estimation uncertainty. Second, we instantiate this insight by identifying compatible client groups from local updated geometry, constructing group-specific task-head prototypes, and dynamically mixing them with local heads using lightweight practical proxies. Experiments on heterogeneous benchmarks, including MiniGrid (up to 12 tasks) and MetaWorld (up to 50 tasks), show that GAdapHeadFRL consistently outperforms state-of-the-art personalized MT-FRL baselines. It achieves up to 24.5% relative improvement on MiniGrid, and increases the converged-task ratio from 50.0% to 72.9% over the personalized baseline. In addition, the proposed method reduces performance standard deviation by over 82.9% on the largest task settings compared with centralized MT-RL references. These results demonstrate that geometry-aware grouping and adaptive task-head mixing provide an effective and scalable principle for heterogeneous MT-FRL.
PaperID: 5940, Poster
Authors:
Yangkai Lin, Jiehong Lin, Kui JiaAbstract: Primitive splatting has become a powerful representation for efficient differentiable rendering, but existing density-control strategies remain largely heuristic and often suffer from uncontrolled primitive growth. In this work, we present Moment-Preserving Density Control (MPDC), a general theoretical framework for densification in primitive splatting. Unlike conventional density-control strategies that \emphselect first and split later, MPDC follows a \emphsplit-then-select paradigm. Our key insight is that a principled split should first preserve the current rendering, rather than immediately perturb the image. By matching local moments, MPDC decomposes a primitive into multiple offspring while keeping the rendered output, and hence the loss, unchanged up to higher-order error. Although such a split does not directly reduce the loss, it changes the parameter space: a saddle point in the original parameterization may no longer remain a saddle point after splitting. This provides a distinct saddle-escaping mechanism from prior approaches that move primitives along negative-curvature directions. Building on this view, we derive closed-form splitting rules for different primitive parameterizations and then select only primitives whose moment-preserving splits reduce an upper bound of the loss through a splitting-matrix criterion. Experiments on three diverse primitive splatting methods show that MPDC substantially reduces primitive counts without sacrificing rendering quality, while improving memory efficiency and rendering speed. Code will be released upon publication.
PaperID: 5941, Poster
Abstract: Graph Foundation Models (GFMs) have attracted growing attention for their ability to model multi-domain graphs universally and generalize to unseen downstream tasks. However, existing GFMs typically rely on a single domain-invariant representation space, which inevitably blurs two fundamental connectivity patterns, \emphhomophily and \emphheterophily, that govern how nodes relate across domains. Additionally, hard-partitioning the graph into disjoint subgraphs discards edges that are essential for self-supervised pre-training. In this work, we introduce \underline\textHHGFM, a \underline\textHomophilic and \underline\textHeterophilic Pattern-Disentangled \underline\textGraph \underline\textFoundation \underline\textModel that encodes homophily and heterophily separately while keeping the original edge structure intact. Specifically, we propose a node-level soft disentanglement mechanism that estimates each neighbor's homophilic and heterophilic contributions and uses them as edge weights, thereby preserving the original graph structure. Low-pass and high-pass filters then serve as dual backbones to align homophilic and heterophilic patterns separately, supported by our generalization analysis. Extensive experiments on a range of graph datasets validate the effectiveness of our method.
Authors:
Heejun Kim, Seungpil Lee, Jewon Yeom, Jaewon Sok, Seonghyeon Park, Jeongjae Park, Taesup Kim, Sundong KimAbstract: Recent soft prompt research has tried to improve reasoning by inserting trained vectors into LLM inputs, yet whether the gain comes from the learned content or from the act of injection itself has not been carefully separated. We study Random Soft Prompts (RSPs), which drop the training step entirely and append a freshly drawn sequence of random embedding vectors to the input. Each RSP vector is sampled from an isotropic Gaussian fitted to the entrywise mean and variance of the pretrained embedding table; the sequence carries no learned content, and yet reaches accuracy comparable to optimized soft prompts on math reasoning benchmarks in several settings. The mechanism unfolds in two stages: because attention has to absorb a never-seen-before random position, the distribution over the first few generated tokens flattens and reasoning trajectories branch, and as generation continues this influence dilutes naturally so the response commits to a single completion. We show that during inference RSPs lift early-stage token diversity and, combined with temperature sampling, widen Pass@N, the probability that at least one out of N attempts is correct. Beyond inference, we carry the same effect into DAPO training and demonstrate practical gains. Our contributions are: (i) RSP isolates the simplest form of soft prompt --- training-free, freshly resampled --- providing a unified lens for the structural effect of injection that variants otherwise differing in training and form all share; (ii) a theoretical and empirical validation of the underlying mechanism; and (iii) an extension from inference to training.
PaperID: 5943, Poster
Abstract: Harmful fine-tuning can rapidly compromise safety alignment in Large Language Models, underscoring the vulnerability of deployed models to post-release modification. While prior work motivates model immunization by inducing ill-conditioned curvature to slow harmful optimization, this approach is not scalable due to the high computational cost of Hessian estimation. We refine this paradigm by extending the theoretical framework to prioritize plateau regions alongside curvature. We further introduce an immunization algorithm based on efficient Hessian estimation and bounds on the extremal singular values of the Hessian. We demonstrate the proposed immunization framework in three applications: (i) Fine-tuning-as-a-service settings where benign user data may be mixed with poisoned harmful data, (ii) open-weight settings where an adversary can fine-tune on harmful-only data with varied training dynamics, and (iii) robust unlearning, where immunized models better resist relearning attacks after unlearning. Extensive experiments confirm that our approach achieves robust defense against adversarial fine-tuning while maintaining general utility and downstream trainability.
PaperID: 5944, Poster
Authors:
Zihan Chang, Shuibing He, Bo Zhou, Ping Chen, Siling YangAbstract: Key-value (KV) caches are essential for efficient Large Language Model (LLM) inference, but their memory footprint grows linearly with context length and batch size. Low-bit KV-cache quantization reduces this footprint, yet uniform quantization is vulnerable to high-magnitude outlier channels, while outlier-aware mixed-precision methods often introduce extra buffers, indexing, or channel permutation that weakens their system-level benefit. This paper presents InQuant, an in-place mixed-precision KV-cache quantization method that preserves salient outlier channels without changing the physical 4-bit packed layout. The key observation is that selected low-saliency channels near outliers can tolerate small approximation error; InQuant therefore reuses their storage slots to hold the extra 4-bit nibble needed by 8-bit outlier values. To make this layout practical, InQuant combines sampling-based channel saliency estimation, saliency-aware neighbor-slot reuse, stride-based handling for adjacent outlier groups, and descriptor-guided marker recovery during dequantization. Across six LLMs and long-context workloads, InQuant reaches a fixed 4.0× physical KV-cache compression ratio, preserves accuracy close to strong mixed-precision baselines, and reduces quantization/dequantization latency by 2.2×--3.4× compared with representative outlier-aware methods.
PaperID: 5945, Poster
Abstract: Diffusion large language models (dLLMs) generate text by iteratively denoising partially masked sequences under bidirectional context, exposing a safety surface distinct from autoregressive LLMs. Because mask tokens are native inputs and tokens are committed by confidence rather than position, harmful content can be induced through infilling and outside the monitored prefix. Existing jailbreaks either miss this native infill capability or rely on low-diversity mask-bearing templates applied uniformly across goals, with little structural adaptation or accumulated attack experience. We propose MaskForge, a fully black-box adaptive attack that casts dLLM red-teaming as optimized search over a growing library of structural patterns. MaskForge abstracts successful attempts into reusable schemas, selects goal-compatible patterns with a UCB bandit, and invokes a scorer-guided fallback when the current library fails. Successful attempts are distilled back into the pattern library, enabling experience to accumulate across goals. Across five public dLLMs and three benchmarks, MaskForge achieves an average attack success rate of 79.3%, a 17.6% relative improvement over the strongest competing dLLM baseline. The matured pattern library further transfers to AdvBench without any updates, achieving a 88.2% attack success rate and a 67% relative improvement over the strongest competing baseline.
Authors: Yiding Song, Hanming Ye
Abstract: Existing accounts of grokking explain the phenomena in terms of mechanistic frameworks such as circuit efficiency or lazy-to-rich transitions. However, despite a known dependence between grokking and model size, how model capacity shapes grokking remains an open question. We give an information-theoretic account of this relationship on the task of modular arithmetic, showing that grokking does not immediately occur when a model becomes large enough to memorise the training set, but rather emerges as the outcome of a competition between two measurable timescales: a memorisation speed T_\textmem(P) and a generalisation speed T_\textgen(P), both of which are functions of model parameter count P. Adapting the information capacity framework of Morris et al. (2025), we estimate T_\textmem(P) on random-label data of equivalent complexity and T_\textgen(P) on the modular task itself, and show that grokking emerges close to the parameter scale where these timescales intersect. The framework also suggests an empirical model for predicting memorisation speed given model capacity and dataset complexity, recovering the previously reported empirical observation that larger models memorise faster. Overall, we motivate the formalisation of different learning timescales as important abstractions to study when explaining how model capacity shapes grokking on algorithmic tasks.
Abstract: Reinforcement learning with verifiable rewards has made post-training effective when correctness can be verified automatically, but many useful model behaviors require satisfying several qualitative criteria rather than a single verifier. Rubric-based rewards extend RLVR to these open-ended domains by grading prompt-specific criteria and aggregating them into a scalar reward. However, the common static aggregations conflate a criterion's desired end-state importance with its current usefulness for learning. We show that this conflation is central in rubric RL: many criteria that matter to the final answer are already saturated or currently unreachable for the policy, while criteria with informative rollout variation are not reliably the ones assigned the largest human weights. We introduce a Policy-Aware Rubric Reward framework for RLVR, POW3R, that keeps human weights and reward category balance as the target objective while reallocating within-category training pressure toward criteria that distinguish the current rollouts. Across three base policies on each of two datasets spanning multimodal and text-only settings, POW3R takes first place on 27 of 33 (base policy, metric) cells we evaluate, leading both mean rubric reward and the harder strict ``every-rubric-passed'' pass-rate over vanilla GRPO with rubric-based rewards, and reaches the same plateau in 3--4× fewer training steps. These findings suggest that rubric rewards should separate what should matter in the final answer from what can teach the current policy.
PaperID: 5948, Poster
Abstract: This paper explores the challenge of accelerating the sequential inference process of Diffusion Probabilistic Models (DPMs). We tackle this critical issue from a dynamic system perspective, in which the inherent sequential nature is transformed into a parallel sampling process. Specifically, we first reveal that the sequential integral solver of the diffusion model can be approximated by a full linear solver, enabling efficient computation for parallel integral solvers of DPMs. We then introduce a unified framework that reformulates the original nonlinear sequential integral process of the DPMs as a system of partial linear equations. Moreover, we further develop an immediate update strategy to solve the system. In addition, we prove that (1) the system admits a unique root corresponding precisely to the trajectory of the sequential integral solver; (2) solving the system guarantees convergence to the trajectory of sequential integral solvers in equal or fewer iterations. We then present \paralin, a partial linear parallel integral solver to accelerate a broad class of sequential and parallel sampling methods such as DDPM and ParaSolver. This partial linearity allows \paralin to achieve parallel speedup on more practical single GPU settings. Experiments (12B Flux, Stable Diffusion Model, VP, VE, EDM, flow matching) validate that \paralin achieves 2.7× to 3.9× speedup in practical 25-50 steps. Notably, on more practical single GPU setting where SOTA parallel solvers typically exhibit constraints, \paralin outperforms them with a 2.5× time speedup. Furthermore, to demonstrate the scalability, we show it can unlock up to a 50x speedup on classical large-step settings (e.g., 1000-step DDPM). The source code will be released publicly.
Abstract: Moving Object Segmentation (MOS) aims to discover, segment, and track objects that move independently of the camera. Current MOS methods, however, exhibit two fundamental limitations: they rely on pre-computed 2D auxiliary modalities such as optical flow or point trajectories that lack 3D geometric information, and they treat motion as a sequence-level attribute, overlooking the instantaneous motion state of each object. We address both by grounding MOS in 3D space and time, and propose GMOS, a framework that operates directly on RGB video to produce 3D-aware, temporally fine-grained segmentation of multiple moving objects, alongside a foreground--background variant GMOS-S for faster deployment. To support training and evaluation in this regime, we curate GMOS-2K, a dataset of 2,210 real-world videos with per-object temporal motion annotations drawn from five established Video Object Segmentation (VOS) benchmarks, and formalise MOS-I ("I" for instantaneous), a temporally fine-grained evaluation protocol with three complementary metrics. GMOS achieves state-of-the-art results across MOS, MOS-I, and Unsupervised VOS benchmarks, while running significantly faster than prior multi-object MOS methods and supporting online inference for streaming deployment.
Abstract: Retrieval is increasingly moving from one-shot matching toward interactive reasoning, where language agents iteratively inspect evidence, reformulate queries, and search again. Training such agents raises a credit-assignment challenge: executable actions such as queries or summaries can be directly evaluated by the retriever, while latent reasoning steps are not directly observable and only affect future executable actions. This asymmetry makes outcome-level reward assignment unreliable, as the same final reward may credit reasoning steps that did not actually shape retrieval success. We propose RICE-PO, a critic-free policy optimization framework that converts retrieval interactions into localized learning signals. RICE-PO selects high-uncertainty executable actions as anchors, evaluates local counterfactual branches using retrieval metrics, and propagates credit to latent reasoning steps only when reasoning-to-action influence is strong and future residual effects are stable. On BRIGHT and BEIR, RICE-PO consistently outperforms prompt-based agents and group-based RL baselines under the same retriever setting. These results show that the structure of agent-environment interaction itself can provide useful supervision for training reasoning-based retrieval.
Abstract: Studying the shape of text embedding spaces enhances our understanding of model behavior and how geometry affects downstream task performance. While a range of topological and geometric metrics have been explored, they are typically analyzed in isolation, and the connections between them remain largely unknown. In this paper, we present Unified Topological Signatures (UTS), a holistic framework for characterizing the structure of embedding spaces. We evaluate our method across a wide range of text embedding models and retrieval datasets, finding that UTS outperform individual metrics in downstream performance prediction. Through several experiments, we demonstrate that UTS are superior at characterizing embedding spaces, can predict model-specific properties and bias from geometry, and reveal novel insights into representational similarity that challenge the Platonic Representation Hypothesis.
PaperID: 5952, Poster
Abstract: Spatial cognition requires stable neural manifolds of orientation and location that track self-motion and sensory cues. Continuous attractors provide a natural dynamical principle, but learning such mechanisms that form and control manifolds with theoretical guarantees remains difficult. We propose Spatial Energy Attractor Learning (SEAL), a theory of neurodynamics learning for spatial cognition, where continuous spatial representations arise as learned low-energy manifolds in Hopfield energy landscapes through population-level Fourier energy learning. For head-direction and grid-cell systems, Fourier energy landscapes yield ring and torus attractors with learned normal attraction and input-driven tangent transport. We prove that Fourier construction is a special case of a general Hopfield-compatible energy extension whose autonomous dynamics preserve energy descent and whose expressive energy families approximate target attractor energies and restoring fields. Experiments on HD and grid-cell populations show that the learned energies restore perturbed states to ring and torus manifolds, stabilize velocity-driven moving bumps across multiple input regimes, and induce continuous-attractor interactive structures with local excitation and surround inhibition. Together, these results provide a stable, learnable, and biologically interpretable energy-based account of continuous spatial representation.
Abstract: Weak-to-strong (W2S) generalization, in which a strong model is fine-tuned on outputs of a weaker, task-specialized model, has been proposed as an approach to aligning superhuman AI systems. Existing theoretical analyses either fix the student's representations or operate in restricted settings. Whether multi-step SGD can succeed in feature learning while preserving diverse pre-trained capabilities remains open. We study W2S in the setting of reward-model learning with two-layer neural networks. The strong model has pre-trained representations organized into low-dimensional subspaces V_k, and is fine-tuned under the supervision of a weak model specialized on task \kappa. We prove that the strong model efficiently learns task \kappa, eliciting its pre-trained knowledge while retaining general capabilities. This establishes W2S generalization in the feature-learning regime, in the sense that the strong model acquires the target feature direction through W2S training, rather than having it given a priori. Moreover, W2S preserves pre-trained off-target features, whereas standard supervised fine-tuning causes catastrophic forgetting when off-target feature directions are correlated with the target's. Numerical experiments on synthetic data confirm our theoretical results.
Abstract: Reliable confidence estimation gates safe deployment of chain-of-thought (CoT) reasoning through text-only APIs, yet the dominant black-box baseline, self-consistency over K samples, is linearly expensive and ignores the geometry of the trace. We introduce a black-box trajectory-confidence score that embeds a CoT as a sliding-window trajectory and measures its convergence toward external answer anchors with a one-parameter softmax, requiring no logits, hidden states, or supervised calibrators. On six (benchmark,reasoner) settings over MedQA-USMLE, GPQA Diamond, and MMLU-Pro × Gemini 3.1 Pro, Claude Sonnet 4.6, fusing this score with coverage and verbalized-confidence channels at K=4 Pareto-improves self-consistency at K=8 in 6/6 settings (median AUC 0.78 vs. 0.71, ΔAUC=+0.075); a fixed-pick control (+0.060) and an E5 cross-embedder replication rule out answer-switching and single-vendor artifacts. Mechanistically, the geometry signal peaks in the penultimate reasoning window across all benchmarks and reasoners and inverts at the terminal window on GPQA Diamond, exposing answer commitment before literal verbalization. Three increasingly unscaffolded regimes decompose black-box confidence into a judge-mediated Coverage prior (C), within-trace Geometry (G), and a conditional Verbalization channel (V); across 18 benchmark × reasoner × proposer settings, C and G carry independent signal in 18/18 and 16/18, while V contributes residual signal in only 6/18. A judge-family swap (GPT-5-mini → Claude Sonnet 4.6) leaves G-only AUC unchanged (∣Δ∣≤0.013) and shifts C-only AUC by at most ±0.02 (κ=0.82), and fusion beats the best single channel in 17/18 settings (median AUC 0.78, max 0.92). Together, these results show that black-box CoT confidence can be read more reliably from a trace's geometric convergence in embedding space than from sample-vote agreement, at lower sampling cost and without access to logits or hidden states.
PaperID: 5955, Poster
Abstract: Scaling test-time compute has driven the recent advances in the reasoning capabilities of large language models (LLMs). However, increased compute often comes at the expense of higher user-facing latency, directly impacting user experience. Current test-time scaling methods primarily optimize for accuracy based on total compute resources (FLOPs), often overlooking latency constraints. To address this gap, we propose SPECS, a latency-aware test-time scaling method. SPECS uses a smaller, faster model to generate multiple candidates for each step of reasoning in parallel, and evaluates these candidates using a larger model and a quality critic. We design a theoretically grounded soft verification strategy to select a high-quality candidate to continue the generation. SPECS also employs a dynamic switching mechanism to use speculative drafts only for easier steps to maintain reasoning accuracy. Empirical results on a diverse and challenging set of reasoning and alignment benchmarks show that SPECS matches or surpasses the accuracy of SOTA test-time scaling methods while reducing latency by up to ~26%. Our theoretical analysis shows that as the amount of parallel compute scales, SPECS converges to an optimum of a KL-regularized reinforcement learning problem, a common objective for aligning LLM generation given a reward signal.
PaperID: 5956, Poster
Authors: Kenneth Yang, Frank Wang
Abstract: Streaming 3D reconstruction enables long-sequence scene understanding by processing frames sequentially, but remains challenging when historical information is compressed into a fixed-size recurrent state. Existing training-free extensions of implicit-memory models mainly improve stability by modulating update magnitude, thereby controlling how strongly each incoming observation modifies the state. However, the geometric intent of the update, namely the direction in which the observation drives the recurrent state, remains tied to the current-frame candidate and can accumulate directional bias over long sequences. We propose HIST3R, a training-free framework that improves recurrent state updating through a direction--magnitude decomposition. HIST3R uses historical recurrent states to rectify the update direction, providing a more reliable geometric intent, and uses observation--history visual discrepancy to modulate the update magnitude, yielding an adaptive update gain. This design makes recurrent updates less dependent on isolated frame-level evidence while preserving the bounded-memory advantage of implicit streaming reconstruction. Across camera pose estimation, 3D reconstruction, and video depth estimation benchmarks, HIST3R consistently improves long-horizon performance without additional training. At 1000 input frames, HIST3R achieves relative ATE reductions of 23.87% on ScanNet and 33.65% on TUM-Dynamics over the strongest baseline; at 500 frames, it improves long-sequence 3D reconstruction over the strongest CUT3R-based variant by 56.52% on 7-Scenes and 33.96% on NRGBD.
Abstract: We present RPC-GS, the first Gaussian Splatting framework for satellite imagery that operates natively with Rational Polynomial Camera (RPC) models. The RPC model is the de facto standard for representing the complex imaging geometry of modern pushbroom satellite sensors. To simplify rendering, prior satellite Gaussian Splatting methods replace the RPC model with perspective or affine camera approximations, leading to geometric errors during reconstruction. RPC-GS avoids these approximations by projecting Gaussian means and covariances directly through the RPC model during the splatting process. We embed the RPC model in a chain of carefully selected geo-coordinate transformations representing a mapping from splatting-suitable scene coordinates to image coordinates. To map the Gaussian covariance matrices, we derive a numerically robust Jacobian-based covariance projection for the (partially nonlinear) coordinate transformations. Since RPCs lack an explicit notion of camera depth, we integrate a metric ray-based depth formulation. We benchmark RPC, perspective, and affine camera models in a unified framework, with our native RPC renderer consistently achieving the lowest reconstruction error on leading satellite benchmark datasets, improving mean altitude error over perspective and affine approximations by 29.6% and 63.8% on DFC2019, and by 9.9% and 37.9% on IARPA2016. We release our code to support future research of Gaussian Splatting in the satellite imaging domain.
PaperID: 5958, Poster
Abstract: We consider the problem of maintaining a minimum-cost bipartite matching between two point sets A and B in a metric space (\mathbbX,\mathsfd) under insertions and deletions of input points. The Wasserstein-p (W_p) cost of a matching M is defined as (\sum_(a,b)\in M\mathsfd(a,b)^p)^1/p and an \alpha-approximate matching is one whose total cost is at most \alpha times the cost of a minimum-cost matching. We obtain two main results. If A and B are points in \mathbbR^d and the distance between two points is measured under the \ell_p-metric, an O(d\epsilon^-3/2)-approximate W_1-matching between A and B can be maintained with an amortized O\big(n^\epsilon d\epsilon^-3/2\varphi(n,1/\sqrt\epsilon)\big) update time, for any parameter \epsilon\in(0,1], and where \varphi(n,1/\sqrt\epsilon) denotes the update time of a \epsilon^-1/2-approximate nearest neighbor data structure. For the \ell_2 norm this results in an update time of O\big((d+ \epsilon^-3/2n^\epsilon)\log n\big). The current state-of-the-art for high-dimensional settings only focus on the \ell_2-metric and maintain an O(\log n)-approximate matching. Our result not only extends to any \ell_p norm, but it also improves the approximation factor for d=o(\log n) under the \ell_2-metric. Next, we consider the problem of maintaining a W_p-matching when A and B are point sets in an arbitrary finite metric space. We show that for any k,p\geq 1 and constant \epsilon>0, a 2k(1+\epsilon)-approximate matching can be maintained with an amortized update time of O(kn^1+1/k\epsilon^-1\log \Delta), where \Delta is the spread of A\cup B. Prior work on finite metric spaces gave an insertion-only data structure for W_1-matching. In contrast, we develop a fully dynamic approach that supports both insertions and deletions and work for all W_p-matchings. Together, our results extend dynamic Wasserstein matching beyond fixed-dimensional W_1 settings: they handle high-dimensional geometric instances, extend prior insertion-only 1-Wasserstein matching guarantees to the fully dynamic setting, and support p-Wasserstein costs for every integer p\ge 1 in arbitrary metric spaces.
Abstract: Speculative decoding accelerates Large Language Models via draft-then-verify, where verification can be framed as an Optimal Transport (OT) problem. Existing approaches typically handle multi-draft and multi-step aspects in isolation, applying either flat OT to single-step drafts or per-token rejection sampling to tree-structured candidates. This separation leaves the joint regime (where multi-step dependencies meet multi-draft branching) poorly optimized, as local verification rules fail to exploit the coupling between horizontal and vertical dimensions of candidate trees. In this paper, we propose a unified perspective that casts tree-based verification as a conditional OT problem. Our key insight is that vertical dependencies can be abstracted through prefix acceptance probabilities, which act as dynamic scaling factors to actively guide horizontal draft selection. Based on this principle, we introduce UniVer, a verification algorithm that jointly optimizes across tree levels by composing local optimal transport plans under prefix constraints. We prove that UniVer remains lossless and achieves the optimal acceptance rate under the proposed conditional framework. Extensive experiments across different tasks and models demonstrate that UniVer improves acceptance length by 4.2% to 8.5% over standard recursive rejection sampling without replacement, while maintaining exact distributional alignment with the target model.
PaperID: 5960, Poster
Abstract: Deep neural networks (DNNs), particularly CNN-based classification systems, are widely deployed due to their strong performance. However, they remain vulnerable to backdoor attacks, where imperceptible triggers can induce targeted misclassification while preserving high accuracy on clean inputs. These triggers may also distort model explanations. Such vulnerabilities raise serious concerns for both currently deployed and real-world applications, highlighting the need for deeper understanding and training-free defense mechanisms. In this study, we extensively investigate the attacking mechanisms in models with Batch Normalization (BN). We provide the first comprehensive theoretical analysis of the relationship between backdoor attacks and BN, showing that trigger-related information is strongly encoded in BN layers even under full model fine-tuning. We further prove that BN’s affine parameters and running statistics jointly influence both predictions and explanations, offering a unified explanation of backdoor behavior. Building on this insight, we introduce a simple training-free defense that re-estimates batch feature statistics and recomputes normalization at inference time, mitigating backdoor effects while preserving clean performance. Extensive experiments on 8 black-box and 3 explanation-aware attacks, compared against 9 defenses, demonstrate that our method reduces attack success rates from 100% to 1%, improves true-class recovery by 89% (+23% over prior work), and boosts explanation fidelity by up to 91%, all without retraining or degrading accuracy. Code will be provided upon acceptance.
Abstract: We study the contextual combinatorial semi-bandit (CCSB) problem with general reward function approximation. At each round, the learner observes a context, selects a combinatorial action consisting of a subset of basic arms, and receives the reward of each selected arm; the goal is to maximize the cumulative reward over time. We propose SquareCB.Comb, a computationally efficient algorithm that, at each round, solves a convex optimization problem to sample a combinatorial action that balances exploration and exploitation. SquareCB.Comb scales to large arm sets and imposes no structural assumptions on the action set beyond a cardinality bound of m on each combinatorial action. We prove that SquareCB.Comb achieves a minimax optimal regret bound of \tildeO\sqrtm A T \log |\mathcalF|, where A is the number of arms, m is the maximum number of arms in a combinatorial action, T is the time horizon, and \mathcalF is the reward function class. In the realizable setting, this bound matches the state-of-the-art regret guarantees achieved by policy search-based algorithms in the more restricted slate recommendation settings, while simultaneously generalizing to arbitrary combinatorial action structures and general reward function approximation.
PaperID: 5962, Poster
Authors:
Leqin Xiang, Yongbin Liu, Chunping Ouyang, Ying YuAbstract: Large language models face substantial computational challenges in long-sequence inference due to the quadratic cost of self-attention, with the prefill stage being particularly bottlenecked. Static sparse attention reduces compute via predefined masks but can miss task-critical evidence, while dynamic sparsification is more flexible yet can be brittle across tasks, leading to either insufficient coverage or distractor-induced noise. We propose AdaCal, an Adaptive Calibration framework for robust sparse attention that synergizes coarse, experience-driven task dispatch with head-wise physical profiling. By deriving entropy, concentration, and mid-zone coverage metrics from a lightweight probe pattern, AdaCal dynamically governs head activation and augmentation policies, allocating additional computation strictly on demand. This design is realized through task-conditional policies that effectively suppress unnecessary expansion for structure-dominant inputs while enhancing long-range coverage for retrieval-intensive tasks. Experiments show that AdaCal improves robustness over strong static and dynamic baselines, with clearer gains in long-context settings beyond 32K tokens.
PaperID: 5963, Poster
Abstract: Characterizing the relationship between brain anatomical wiring and functional coordination remains a fundamental challenge in computational neuroscience. This difficulty primarily arises because structure-function coupling in the human brain is spatially heterogeneous, often localized to subnetworks, and inherently difficult to align across disparate representational modalities. To address this methodological gap, we introduce Orthogonal Sparse Subgraph Alignment (Ossa), a principled and interpretable framework designed to identify compact structural cliques that exhibit maximal concordance with functional connectivity (FC) across subjects. Within each FC-derived community, Ossa optimizes an alignment objective over an orthogonal transformation, which absorbs coordinate mismatch between the two connectivity modalities, and a sparse node-selection vector, which identifies the induced structural connectivity (SC) subgraph. Furthermore, we establish theoretical properties for the Ossa estimator, demonstrating its statistical consistency in recovering the true underlying SC clique within each functional community. Extensive empirical evaluations on simulated data, alongside analyses of the Adolescent Brain Cognitive Development (ABCD) baseline and longitudinal cohorts, demonstrate that Ossa successfully recovers aligned structural cliques under latent cross-modal rotations. Ultimately, the proposed methodology significantly improves out-of-sample SC-FC alignment over size-matched baselines and robustly identifies stable, single-core topological structures within human brain networks.
PaperID: 5964, Poster
Authors: Gianluigi Vitale
Abstract: Jailbreak attacks that frame harmful queries as professional requests can bypass LLM safety training, but we do not know which ingredients in a professional-framing pipeline actually matter. We introduce a , separating access-driving from depth-driving components, and apply it to one concrete pipeline family, (S3) of framework terminology for colloquial language. Tested on N = 5,000 trials across three closed frontier models (~31,900 total trials across nine LLMs), the decomposition isolates (Persona > Linguistic Substitution > Moral Justification). This ordering holds across three full-rank factorial replications (defensive system prompt, cybersecurity, social engineering; N ≈ 5,000 each) and nine additional robustness checks. Two defensive targets follow: detect persona claims to block access; restrict output depth when responses use framework terminology. The method is portable beyond STF; the empirical result is specific to one pipeline tested primarily through provider APIs, with open-weight corroborations on local inference (§3).
Abstract: Identifying latent dynamical systems from noisy, high-dimensional measurements is a central problem at the intersection of representation learning, system identification, and scientific discovery. We present DYSCO, a multi-view temporal contrastive learning algorithm that jointly recovers latent trajectories and the governing dynamics from such observations, by leveraging multiple independent noisy views of the same underlying process to disentangle signal from noise. By parameterizing the dynamics in a structured functional basis, our framework further enables symbolic recovery of the governing equations within an affine gauge. We offer theoretical guarantees for strong identification up to an affine indeterminacy, extending prior identifiability results to the realistic setting of noisy nonlinear observations. Empirically, we demonstrate accurate recovery of both latent trajectories and flow fields across a diverse set of dynamical regimes (e.g, chaotic, oscillatory, and metastable) under both Gaussian and Poisson observation noise, the latter being particularly relevant for neural recordings.
PaperID: 5966, Poster
Abstract: Large language models rely on positional representations for long context modeling, yet are typically trained on short sequences and evaluated at much longer lengths. Bridging this gap requires positional encodings remaining distinguishable at large distances while generalizing beyond the training length. In this paper, we identify a structural limitation of existing rotary-based positional encodings: all positional dimensions are controlled by a single global index, which couples training-time positional variation with long-range extrapolation. To address this, we propose Tree Rotary Positional Encoding (TPE), a structured multi-scale positional encoding designed for training from scratch that represents absolute positions through com positions of different components, avoiding single-index extrapolation. We further devise an Offset-based Positional Training (OPT) strategy to improve learning interactions across TPE components, enabling better generalization to longer con texts. We theoretically analyze that its multi-scale decomposition enables effective combinatorial position coverage. Experimental results show strong gains under extreme length extrapolation: with a 1.2B model trained from scratch on length 512, our TPE achieves a 64× extrapolation to 32K tokens, reducing perplexity from 4532.21 (RoPE) to 149.96, with comparable inference-time efficiency.
PaperID: 5967, Poster
Abstract: Multi-party spoken communication pervades daily life, yet even attentive participants routinely lose track of who said what, miss a turn, or struggle to recall something said minutes earlier. As always-on wearables become commonplace, an assistant that listens alongside the user could ease this everyday friction. We introduce Situated Communication Assistance: on-demand support for users immersed in ongoing multi-party conversations, grounded in the same egocentric acoustic scenes they perceive. We argue this capability rests on three abilities: producing a Rich Situated Transcription (RST) of who said what, when, and where; comprehending the conversation holistically; and answering user queries while the conversation is unfolding. To address the persistent data deficit in this regime, we build SitCom, a multi-party spoken dialogue corpus more than an order of magnitude larger than the largest prior multi-party corpus, pairing 15.8k hours of synthesized data with 144 hours of real recordings standardized under a unified RST schema. The synthesis pipeline generates group-structured scripts with persona-grounded speech and behaviors, and renders them in 3D scenes through a directional simulator with head-and-torso radiation. Using this corpus, we curate SitCom-Bench, a fully human-validated question answering benchmark of 4.3k items spanning seven comprehension and four in-the-moment query subtypes. Experiments show that combining synthetic and real data is necessary to close the sim-to-real gap, and that in-the-moment querying remains hard even for frontier models, pointing to it as an open challenge distinct from post-hoc comprehension.
Abstract: Sortition is the practice of delegating public decision-making to randomly selected panels. Recently, it has gained momentum worldwide through its use in citizens' assemblies, sparking growing interest within the computer science community. One key appeal of sortition is that random panels tend to be more of the population than elected committees or parliaments. Our main conceptual contribution is a novel definition of representative panels, based on the Wasserstein distance from statistical learning theory. Using this definition, we develop a framework for analyzing the problem—determining the required panel size to ensure desirable properties. We focus on three key desiderata: (1) that efficiency at the panel level extends to the whole population, measured by social welfare; (2) that fairness guarantees for the panel translate to fairness for the population, captured by the core; and (3) that the probability of an outlier panel, for which the decision significantly deviates from the optimal one, remains low. We establish near-tight panel complexity guarantees for these desiderata across two fundamental social choice settings: facility location and participatory budgeting.
PaperID: 5969, Poster
Authors:
Yao Li, Peiyuan Tang, Wuyang Zhang, Haojie Ren, Chengyang Zhu, Yifan Duan, WeiKai Shi, Xiaodong Zhang, Yuming Dong, Zijiang J Yang, Jianmin Ji, Yanyong ZhangAbstract: Vision-Language-Action (VLA) models have shown strong potential for general robotic manipulation, but contact-rich tasks still require timely action refinement using force/torque feedback. Existing force-aware VLA models usually align all modalities to a single low operating frequency, which discards high-frequency contact cues that are important for reactive action correction. Additionally, they use static fusion, which cannot adaptively balance visual context and force feedback across manipulation processes. To mitigate these issues, we propose FAVLA, a force-adaptive multi-rate VLA model that explicitly separates low-frequency visual-language reasoning from high-frequency force-conditioned action refinement. Based on the \pi_0-style VLM-action expert architecture, FAVLA uses a low-rate VLM to encode visual-language-force context and predict near-future force statistics, while a high-rate action expert refines action chunks using the latest force observations. To decide \emphwhen to update actions, we introduce a Force-Adaptive Multi-rate Inference (FAMI) mechanism, which schedules the action expert's inference rate from predicted future force variance, keeping low-rate reasoning in stable phases and increasing update frequency near contact transitions. To decide \emphhow to use force, we design a Force-Guided Dynamic Fusion (FGDF) module, which injects high-frequency force features into the action expert and dynamically balances visual-semantic and force cues across manipulation stages. Extensive real-robot experiments on high-precision and contact-rich tasks show that FAVLA outperforms force-aware VLA baselines, achieving an average success rate of 88.8% and lower peak contact forces during manipulation.
Abstract: Transformers are limited to computations in the TC^0 complexity class, excluding tasks such as entity tracking and code execution that provably require greater expressivity. Motivated by this, we revisit non-linear Recurrent Neural Networks (RNNs) and introduce Matrix-to-Matrix RNN (M^2RNN): an architecture with matrix-valued hidden states and non-linear state transitions. We show that (i) the language modeling gap of non-linear RNNs is primarily a state-size gap, not a non-linearity penalty, and (ii) outer-product state expansion enables efficient tensor-core utilization. Empirically, M^2RNN achieves perfect state-tracking generalization at sequence lengths beyond training. These gains transfer to large-scale language modeling: at 7B MoE, Hybrid M^2RNN outperforms Hybrid Gated DeltaNetby 0.5 perplexity points using 3× smaller recurrent states, and replacing even a single recurrent layer with M^2RNN matches Hybrid M^2RNN accuracy with minimal throughput cost. Hybrid Gated DeltaNetwith a single M^2RNN layer also outperforms state-of-the-art hybrid linear-attention architectures by up to 8 points on LongBench. Together, these results establish non-linear RNN layers as a compelling building block for efficient and scalable language models.
PaperID: 5971, Poster
Authors: Ping Guo, Zhiqi Huang, Xinran Li
Abstract: Pseudo-labeling has become a cornerstone of learning from unlabeled data in semantic segmentation. Yet its effectiveness drops sharply in real-world scenarios where strong imaging noise and long-tailed class distributions occur together. We trace this failure to a vicious cycle of pseudo-label degradation. Imaging noise entangles foreground and background features, lowering prediction confidence across all classes, while long-tailed distributions leave tail classes with far fewer training samples and inherently lower confidence. Under fixed high-threshold filtering, these tail-class predictions are systematically filtered out, so they receive no supervision from unlabeled data and thus features keep degrading in subsequent iterations. Critically, noise and long-tail are not independent obstacles but mutually amplifying ones, and addressing either alone is insufficient. To break this cycle, we propose FTC-Seg, a Feature-Threshold dual-Calibration framework built on a standard teacher-student framework. At the feature level, Orthogonal Prototype Reconstruction (OPR) uses a set of learnable orthogonal prototypes to residually purify pixel-wise features, widening the margin between weak foreground targets and noisy backgrounds. At the threshold level, Adaptive Threshold Calibration (ATC) dynamically adjusts class-specific thresholds based on learning difficulty and prediction-distribution bias, rescuing low-confidence pseudo-labels of tail classes from systematic exclusion. Extensive experiments on four public benchmarks spanning three distinct noise modalities show that FTC-Seg achieves strong performance against state-of-the-art methods, with particularly substantial gains on tail classes. Our results establish that jointly calibrating features and thresholds is essential for robust pseudo-labeling under compounded noise and class imbalance.
PaperID: 5972, Poster
Abstract: Swapping two LLM post-training stages can change the final model, but an aggregate loss or benchmark delta only says that something moved, not where the order-dependent residue landed. We ask whether this residue leaves a memory trace: a structured signal that is localized in output space, changes held-out order-gap NLL under targeted interventions, and is assignable from paired endpoint weights. For a specified two-stage path (A,B), we define commutator memory by projecting the Lie bracket b_AB := H_B g_A - H_A g_B at the base model through the logits into token scores \tau_k=\mathbbE[e_k\,\delta z_k], whose sum is the leading bracket-predicted order effect. The token readout is localized: endpoint and batch-resampled checks preserve bracket top-token supports far more than norm-matched random directions (82--99% vs. 35--49% top-20 overlap), and a separate endpoint-support prediction check recovers 53--57% of Qwen top-1% active-token supports. It has intervention leverage in the tested protocols: top-positive-\tau_k token interventions close 32% (median) of the held-out Qwen-3-4B SFT ordering gap with tau-neutral controls near zero, with the same bracket-vs-control separation on a pre-specified above-floor Qwen-2.5-1.5B fp32 subset. It is assignable from paired endpoints: given a base model, candidate stages, and paired alternate-order endpoints, the statistic \langle\theta_AB-\theta_BA,\,b_AB\rangle distinguishes k=1 SGD training order in 66/72 pair-seed units across four LLMs. Matched-batch DPO and frozen-rollout reward-surrogate experiments are matched-path stress tests; AdamW is reported as an analogous lifted-state pilot. Commutator memory turns training order from a scalar nuisance into a localized, intervention-tested, and paired-endpoint-readable residue.
PaperID: 5973, Poster
Abstract: Incorporating visual semantic representations as an intermediate step before image generation can reduce the modeling difficulty between text and images, thereby improving generation quality. Recent works such as X-Omni and BLIP3o-Next have explored this direction, but they typically use a two-stage external pipeline: a separate autoregressive model first generates semantic tokens, which are then fed as conditioning to an independent diffusion decoder. Since the decoder cannot jointly access the original input and the semantic plan, this design introduces an information bottleneck that limits detail preservation in downstream tasks such as editing. Internal architectures such as Transfusion, BAGEL, and Show-o2 avoid this bottleneck by enabling cross-modal interaction within a single model, but they still face the difficult text-to-pixel modeling gap without intermediate semantic guidance. We propose Visual Prompt Engineering (VPE), which can be seamlessly integrated into such internal frameworks. Specifically, the model first autoregressively generates visual semantic tokens (e.g., SigLIP 2) as "visual prompts" that capture the semantic layout, then generates the full image tokens conditioned on this plan. We validate VPE across class-conditional generation, text-to-image generation, and image editing, covering various token types and model architectures. Results show that VPE can accelerate convergence, raise quality ceilings, and through internal integration, achieve substantially better editing preservation (PSNR: 26.76 vs. 19.92) than external alternatives of the same parameter scale, while maintaining competitive editing responsiveness. The code is available in supplementary material.
PaperID: 5974, Poster
Abstract: Concept erasure is a prominent approach to achieving fairness in machine learning through the removal of sensitive attributes from learned representations. Prior concept erasure methods typically define debiasing indirectly through the failure of a chosen adversary, probe, or fairness penalty, making the target of erasure dependent on a particular decoder family or optimization setup. Rather than depending on a particular architecture, task type, or training paradigm, we introduce KINDER, which aims to directly remove sensitive-attribute information from a model's intermediate representations and is compatible with any deep model that contains an intermediate feature space. KINDER operates in a random Fourier feature space, where it estimates a sensitive subspace from protected group prototypes and projects representations onto its orthogonal complement, thereby forming an explicit representation cleaning mechanism. We conduct extensive experiments across diverse settings and show the effectiveness of KINDER on (i) unimodal and multimodal datasets, (ii) supervised and self-supervised settings, (iii) classification, regression, and image segmentation tasks, and (iv) diverse data modalities, including visual, textual, and tabular data.
PaperID: 5975, Poster
Authors: Hailong Yan, Yongrui Zhang, Xiangtao Zhang, Le Zhang
Abstract: We propose MorphSIG, a training-free framework that reformulates subject-driven image generation (SIG) as pseudo-video feature transport. MorphSIG contains two core modules: Progressive Decoupled Anchoring, which uses frequency-domain decoupling and structural scrambling during early denoising to break structural locking and enable a gradual transition from free composition to identity alignment; and Pseudo-flow Guided Feature Transport, which constructs a latent-space pseudo-flow field to mimic Image-to-Video temporal coherence and transport subject semantics precisely. To address stylistic homogeneity and the lack of ground truth in existing benchmarks, we further introduce a benchmark with diverse stylized subjects. Extensive experiments show that MorphSIG significantly improves pose diversity and subject fidelity over baselines without additional training. Overall, MorphSIG bridges static generation and dynamic propagation, providing a cost-effective paradigm for high-fidelity subject consistency. Code will be released.
PaperID: 5976, Poster
Abstract: Physics sensing in evolving environments requires reconstructing dense spatiotemporal fields from sparse observations while adapting where future observations should be collected. Existing reconstruction and sensor-placement methods have made substantial progress, but they typically rely on static sensing assumptions: they either use a single layout across all frames or optimize placements independently at each frame, which limits their ability to produce temporally coherent sensing trajectories. We introduce DyPSI, a dynamic physics sensing framework that formulates this problem as the joint generation of future fields and sensor trajectories conditioned on historical observations. DyPSI lifts discrete sensor coordinates into continuous Sensor Position Fields, encodes fields and sensing configurations with a shared Functional Tucker representation, and trains a joint diffusion model to generate future fields and sensing trajectories simultaneously in the resulting latent space. A Gaussian-process temporal kernel correlates the perturbation injected during diffusion training, biasing the denoiser toward temporally smooth recoveries, and we further provide an analysis linking the EDM training objective to a sequence-aggregated A-optimality criterion and a kernel-induced bound on discrete trajectory variation. Experiments on turbulent flow, GLORYS12 sea-surface temperature, and 3D car aerodynamics show that DyPSI consistently outperforms static and frame-wise placement strategies, with substantial error reduction under tight sensor budgets.
PaperID: 5977, Poster
Authors:
Daewon Chae, Hyunwon Chung, Changwoo Lee, Hun-Seok KimAbstract: Large-scale foundation models achieve strong performance across diverse tasks, but their size makes inference costly, largely due to dense matrix multiplications. Prior work reduces this cost by replacing dense weight matrices with efficient structured forms such as low-rank factorizations. However, these methods approximate weights rather than the output activations that determine inference accuracy. Consequently, small weight-space errors can be amplified by input activations, producing large output errors. In this work, we propose DyRA, an input-adaptive method that improves structured matrix multiplication approximation by correcting residual output errors during inference. We show that matrix multiplication can be approximated more effectively by directly optimizing low-rank factors of the output. DyRA builds on this insight by dynamically approximating and correcting the output error introduced by structured weight approximations. This combines efficient structured computation with input-dependent correction, yielding a more faithful approximation of full matrix multiplication under the same computational budget. Across vision, speech, and language models, DyRA consistently improves the accuracy–efficiency trade-off over structured weight approximations alone. Notably, DyRA achieves a 1.5× end-to-end GPU speedup for DINOv3 while reducing accuracy degradation by more than 3× relative to weight-only baselines.
Abstract: Text-to-image diffusion models can generate visually stunning images, yet controlling what appears and how it appears- remains surprisingly difficult, especially when operating solely within the constraints of the text-conditioning space. For example, changing a subject or adjusting an attribute often leads to unintended side effects, such as altered backgrounds or distorted details. This is because most existing text-based editing methods treat the embedding space as Euclidean and apply simple linear transformations, which do not reflect how semantic concepts are actually organized. In this work, we take a step back and ask: what is the true geometry of these embeddings? We find that text encoder representations lie on a hypersphere, where concepts are not linear directions but structured, anisotropic distributions better captured by Kent distributions. This geometric perspective reveals why linear edits lead to entanglement and loss of control. Building on this insight, we propose HEART, a training-free framework that performs Kent-aware geodesic transformations directly on the hypersphere. By respecting the underlying geometry, HEART enables intuitive and precise edits - such as consistent subject replacement and fine-grained attribute control - while preserving the original scene. Importantly, HEART requires no finetuning, inversion, or optimization, and generalizes across diffusion model architectures. Our results show that a simple shift in perspective, from linear to geometric, can unlock fast, reliable, and controllable image generation.
Abstract: Reflectance confocal microscopy (RCM) provides noninvasive, cellular-resolution “optical biopsies” of human skin \em in vivo by acquiring en-face images at successive depths, forming a sparse z-stack. Due to optical limitations, these stacks are anisotropic 3D volumes with lateral resolution (0.5\mu m) ~6 times higher compared to axial resolution, which is defined by the optical sectioning (3\mu m), limiting the interpretation of tissue. Our goal is to provide continuous-depth visualization by interpolating intermediate sections and making the 3D volume isotropic. Such a representation permits arbitrary-direction sectioning, including histopathology-like cross-sectional examination, without requiring per-patient optimization. To that end, we introduce the first RCM-specific novel-view synthesis (NVS) approach CD-RCM: a feedforward model that predicts realistic, unseen depths from sparsely sampled RCM stacks. Classical neural rendering methods focus on reconstruction from surface-level multi-view observations. In contrast to surface-level camera views, RCM can acquire optically sectioned en-face images of tissue beyond the surface up to 200 \mu m. However, during visualization of the RCM stacks, observations of the shallower sections (towards the surface) obscure the deeper ones. This unique axial imaging geometry and layer-dependent anatomical organization motivated our development of a tailored architectural and training framework that explicitly accounts for RCM’s depth-resolved, occlusive imaging physics. Experiments demonstrate that CD-RCM achieves high-fidelity novel-view synthesis with sub-second inference time.
Abstract: Agentic workloads have emerged as a major workload for LLM inference. They differ significantly from chat-only workloads, requiring long-context processing, the ability to handle multimodal inputs, and structured multi-turn interactions with tool calling capabilities. As a result, their context exhibits structure that can carry different importance along three key axes: temporal recency to the current turn, modality such as text or image tokens, and semantic role such as user queries, tool calls, observations, or reasoning. These axes capture distinct token behaviors and lead to different sensitivities to KV-cache compression. However, existing KV-cache quantization methods are typically homogeneous or exploit only heterogeneity on a single dimension, such as temporal proximity or modality, overlooking the interactions among them. To this end, we introduce TriAxialKV, a novel mixed-precision KV-cache quantization scheme that assigns each token a triaxial tag, calibrates per-tag sensitivity, and allocates INT2/INT4 bitwidths under a fixed memory budget. We implement TriAxialKV as an end-to-end serving system, comprising calibration, mixed-precision quantization and memory management, and custom fused Triton decode kernels. When using Qwen3-VL-32B-Thinking as a computer-use agent operating the OSWorld, TriAxialKV matches the accuracy of SGLang with BF16 KV cache while supporting 4.5× KV cache size and achieving 30% higher end-to-end throughput, when running on real GPU systems.
Abstract: As revealed by the scaling law of Mixture-of-Experts models, model performance ceases improving once the granularity of the intermediate dimension exceeds the optimal threshold, limiting further gains from single-dimension fine-grained design. To address this bottleneck, we propose FineRMoE (FineR-Grained MoE), an architecture that extends fine-grained expert design to both intermediate and output dimensions, aiming to enhance expert specialization beyond the single-dimension limit. We further introduce a bi-level sparse forward computation paradigm and a specialized routing mechanism to govern the activation. In addition, to obviate the prohibitive cost of pre-training from scratch, we devise a generalized upcycling method to build FineRMoE in a cost-effective manner. Extensive experiments performed under the upcycling paradigm demonstrate the superior performance achieved by FineRMoE across ten standard benchmarks. Compared with the strongest upcycling baseline method, FineRMoE achieves significant parameter and inference efficiency gains.
Abstract: Mobile agents are increasingly expected to operate everyday applications from screenshots and language goals, where reliable control requires reasoning over screen affordances, multi-step navigation, and future state changes. Yet many agents externalize this computation as long textual thoughts, making interaction slower, supervision more costly, and deployment less efficient. We introduce MIRAGE (Mobile agents with Implicit Reasoning And Generative world modEls), a framework that learns continuous latent reasoning representations from visible textual thoughts. MIRAGE introduces an efficient latent-space learning procedure that transfers explicit reasoning into compact hidden states, allowing the agent to reason internally without decoding long rationales. It further brings a world-model perspective into mobile-agent training: the model’s latent reasoning vectors are aligned with future screenshots, encouraging the agent to predict upcoming interface states in latent space before executing an action. This makes the hidden computation not only a compressed thought trace, but also a forward-looking representation of how the environment may change. At inference time, MIRAGE reasons in continuous latent space, reducing token generation while improving execution efficiency. On AndroidWorld, MIRAGE matches explicit-CoT SFT in the 4B ablation under a 3–5× lower decoded-token budget, and improves a comparable instruction-tuned baseline by 10.2 points; on AndroidControl, it improves action grounding with over 75% fewer generated tokens.
PaperID: 5983, Poster
Abstract: Due to the challenge of uncovering biologically meaningful spatial tissue patterns from standard H&E-stained whole-slide images (WSIs) without relying on multiplexed staining or fixed-scale modeling, a unified framework is proposed for spatial biomarker discovery that integrates cell classification, adaptive multi-scale microenvironment construction, and colocalization pattern analysis. Cell-level predictions assign discrete, semantically meaningful types that serve as the basis for microenvironment modeling, where physical adjacency initializes local ecological features and a combination of spatial competition-based seed selection with homogeneity-constrained recursive region growing allows microenvironments to expand from locally pure regions in a data-driven manner, capturing heterogeneity across scales without predefined spatial ranges. Microenvironment-level cell composition features are then used to construct a colocalization feature space, explicitly modeling spatial relationships among cell types, and differential biomarkers are identified via a distributional comparison framework that systematically detects group-specific colocalization patterns. Experiments demonstrate that this framework reliably discovers discriminative and biologically interpretable spatial biomarkers from H&E images alone, exhibiting robustness in tissues with complex structures and heterogeneous scales, and offering a flexible, extensible paradigm for spatial phenotypic biomarker discovery in standard histopathology.
PaperID: 5984, Poster
Abstract: Probabilistic long-term time series forecasting requires estimating a distribution over future trajectories under practical deployment constraints, where inference latency and stability can be as critical as forecast quality. Existing diffusion-based forecasters provide expressive uncertainty modeling, yet their iterative denoising incurs substantial horizon insensitive runtime overhead, which limits real time long-term use. In contrast, one-shot generative models such as VAEs and flows are computationally efficient but often struggle to capture complex multi-modal structure and heavy tailed uncertainty. We propose ToLD, an efficient time‑series forecasting framework built upon tokenized truncated latent diffusion to enable fast and reliable long term generation. ToLD adopts a tokenized conditional latent modeling and performs compact token-wise refinement via a few-step truncated latent diffusion process, where the diffusion noise is injected in a context-dependent manner to model non-stationary uncertainty. To reduce train test mismatch in multi-sample forecasting, we further introduce a score distillation scheme that learns a target-free scoring function for test-time candidate ranking. Experiments on eight real-world datasets show that ToLD improves both point accuracy and probabilistic quality over state-of-the-art, achieving up to 4.3% MSE improvement and 19.3% CRPS reduction on long-term tasks with minimal additional inference overhead.
Abstract: Neural models such as YOLO and HuBERT detect local properties such as objects ("car") and emotions ("angry") in video frames and audio clips, with scores in [0, 1]. Lifting these scores to temporal properties over sequences enables applications such as query matching (e.g., "does the speaker eventually sound happy in this audio clip?") and ranked retrieval (e.g., "retrieve top 5 videos with a 10 second scene where a car is detected until a pedestrian is detected"). We formalize this problem of assigning Scores for TempOral Properties (STOPs) over sequences, given noisy score predictors for local properties. We propose LogSTOP, a scoring function that efficiently computes scores for temporal properties represented in Linear Temporal Logic. Empirically, LogSTOP with YOLO and HuBERT outperforms Large Vision / Audio Language Models by at least 16% on query matching with temporal properties over objects-in-videos and emotions-in-speech, while matching or improving over temporal-logic baselines. On ranked retrieval with temporal properties over objects and actions in videos, LogSTOP with OWLv2 and SlowR50 improves mean average precision over zero-shot text-to-video retrieval baselines by 28% and 18% respectively.
PaperID: 5986, Poster
Authors: Yashas Vaidya, Bo Dai
Abstract: Context-based offline meta-reinforcement learning (COMRL) aims to learn adaptable policies entirely from static datasets by inferring latent task representations from small contexts. Most COMRL methods rely on immediate, per-step reward signals. This can be generalized by aggregating rewards into ``bags'' and being delivered after a sequence of actions (where bag size n=1 is the per-step case). We formally prove that under Bagged Rewards, existing per-transition COMRL methods suffer an information-theoretic collapse (\mathcalO(1/n)), causing their task inference to degrade to behavior cloning. To solve this, we introduce SpectralMeta, an algorithm that fully decouples dynamics representation from task inference. By pre-training an energy-based spectral representation on reward-free transitions, we recast task identification as a closed-form Bayesian linear regression over bagged rewards. SpectralMeta comes with finite-sample bounds on reward recovery and policy sub-optimality. In continuous-control meta-RL benchmarks, SpectralMeta maintains robust task inference and competitive returns even when rewards are reduced to a single episode-level scalar, a regime where existing baselines fail to recover task-relevant structure.
Abstract: Continuous-time generative models compress endpoint-conditioned bridges into Markov velocity fields, but no existing diagnostic measures how much information this compression discards. We introduce the \emphMarkovization gap---the integrated conditional variance of the bridge velocity given the Markov state---which quantifies this loss before any neural network is trained. To make the gap well-defined across model families, we define \emphBridge Graphical Models (BGMs), separating endpoint coupling, bridge law, Markovian projection, and dynamics representation as independent design choices. This decomposition also formalizes Poisson and electrostatic models as field-line bridge kernels with their own caustic gap. Across synthetic, latent, and pixel-space pilots on CIFAR-10 and Fashion-MNIST, a proxy gap estimated in minutes on CPU consistently ranks coupling/bridge choices in the same direction as training loss and FID measured after hours of GPU training.
Authors:
Yuechen Luo, Fang Li, Shaoqing Xu, Yang Ji, Zehan Zhang, Bing Wang, Shen Yuannan, Jianwei Cui, Long Chen, Guang Chen, Hangjun Ye, Zhi-Xin Yang, Fuxi WenAbstract: While Vision-Language-Action (VLA) models have revolutionized autonomous driving by unifying perception and planning, their reliance on explicit textual Chain-of-Thought (CoT) leads to semantic-perceptual decoupling and perceptual-symbolic conflicts. Recent shifts toward latent reasoning attempt to bypass these bottlenecks by thinking in continuous hidden space. However, without explicit intermediate constraints, standard latent CoT often operates as a physics-agnostic representation. To address this, we propose the ), a framework shifting the reasoning paradigm from discrete symbolic processing into a physically grounded Latent Spatio-Temporal CoT. By implementing a dual-feature alignment mechanism, we distill geometric constraints from 3D foundation models and dynamic foresight from world models directly into the latent space. Coupled with a progressive SFT training strategy that transitions from feature alignment to trajectory generation, and further refined via Reinforcement Learning with Group Relative Policy Optimization (GRPO) for safety and rule compliance, LaST-VLA achieves state-of-the-art performance across multiple autonomous driving benchmarks, including
PaperID: 5989, Poster
Abstract: Multi-agent systems have emerged as a promising paradigm for multimodal medical reasoning, but existing methods rely on two assumptions that hinder practical deployment: (i) always-active dense collaboration among multiple agents, and (ii) dependence on large cloud-based models that raise privacy and efficiency concerns. In this paper, we challenge both assumptions and propose SMART-Med, a fully local multi-agent method that treats collaboration as a learnable decision rather than a default behavior. Our key insight is that not every medical query requires costly multi-agent deliberation: many can be reliably solved by a single well-trained agent, while only uncertain cases benefit from expert collaboration. Based on this insight, SMART-Med introduces an uncertainty-driven routing mechanism that adaptively decides when to answer directly and when to recruit a small subset of expert agents, effectively converting dense multi-agent reasoning into selective, query-adaptive coordination. To make this routing reliable, we curate a new multimodal medical dataset MedVR-3K with high-quality reasoning traces and a difficulty hierarchy, and then fine-tune local agents using GRPO with two complementary rewards: a certainty-aware accuracy reward and a vision-language reward model for correctness and reasoning quality respectively. Experiments on four medical visual question answering benchmarks show that SMART-Med achieves competitive performance compared to existing MAS approaches while substantially reducing inference cost (reducing inference time by more than 3.8× and token usage by 72%). Moreover, after fine-tuning on MedVR-3K, SMART-Med yields an average improvement of 6.7% across Slake, PathVQA, VQA-RAD, and PMC-VQA. Our results suggest that selective collaboration can be an efficient yet effective alternative to dense collaboration and enable practical, privacy-preserving deployment of multi-agent medical AI.
PaperID: 5990, Poster
Abstract: Byzantine-robust Federated Learning (FL) designs aggregation rules that mitigate the impact of malicious clients. When the data across honest clients are heterogeneous, honest model updates are inherently scattered. This dispersion can be exploited by Byzantine adversaries who may collude to concentrate their updates, thereby circumventing state-of-the-art robust aggregation methods and causing a Byzantine Adversarial Disruption (BAD) of the training process. We propose a novel Byzantine-robust FL scheme based on scores. Our scoring technique captures concentration among Byzantine clients and prevents catastrophic failure under SotA attacks. Through numerical experiments on FEMNIST, Shakespeare, and CIFAR-10 datasets, we demonstrate that the proposed defense maintains high model accuracy, even in realistic FL scenarios. Furthermore, we establish theoretical convergence guarantees for our method.
Authors: Solomon Messing
Abstract: LLM evaluations drive which models get deployed, what safety standards get adopted, which research conclusions get published, and how projections of AI's labor-market impact get made. Yet standard confidence intervals ignore variability from judge model choice, model temperature, and prompt phrasing, producing under-coverage that worsens with more data. The omitted variance can shift results enough to reverse conclusions \citepbaumann2025llmhacking, huang2026dropping; pipelines that fail to average over it leave the surface that ``benchmark hacking'' exploits \citepsingh2025leaderboard. This paper decomposes LLM pipeline uncertainty into its sources, distinguishes variance that shrinks with more data from sensitivity to researcher design choices, and uses design-study projections to reduce total evaluation error (TEE). Across the demonstrations, naive standard errors are 40 - 60% smaller than the TEE-corrected SE. Using Chatbot Arena data, we show naive 95% CI coverage drops as n grows while TEE-corrected coverage holds at 95%, and TEE-guided pipelines restrict the benchmark gaming surface from 56 to 32 Elo (K=27), below the human-leaderboard baseline. We show further that a small pilot recovers honest CIs and projects which design changes most improve precision. Acting on those projections halves MMLU estimation error against the answer key at equivalent cost, and raises per-match agreement with human votes by 7.9 percentage points on Chatbot Arena.
PaperID: 5992, Poster
Abstract: Continual learning (CL) aims to adapt models to new tasks while preserving previously learned knowledge. Recently, Low-Rank Adaptation (LoRA), a representative Parameter-Efficient Fine-Tuning (PEFT) method, has gained increasing attention in CL due to its efficiency and scalability. Several LoRA-based merging methods maintain a single inference model by scaling the newly learned LoRA update with one coefficient and merging it into the previous weights after each task. However, they usually control the whole update with a shared scalar factor, although different update elements may affect previous and current tasks in different ways. To address this issue, we propose Merging Strength Modulation via Joint Loss Estimation (M²LE) It consists of Hessian Information-guided Adaptive Merging (HIM) and Task Simulation Perturbation for Merging (TSP). Instead of searching for a global merge coefficient, HIM directly solves for the target merged update under a joint objective over previous and current tasks. Solving this objective gives an analytic solution that naturally takes the form of element-wise modulation over the current task LoRA update. To improve the stability of the learned update under subsequent merging, we introduce TSP which simulates different tasks interference. Experiments on multiple benchmarks show that M²LE consistently outperforms existing PEFT-based continual learning methods. Our code is provided in the supplementary material.
PaperID: 5993, Poster
Abstract: Reinforcement learning with verifiable rewards (RLVR) significantly enhances the reasoning capabilities of large language models. However, as training progresses, models often converge to narrow solution paths, leading to diversity collapse. Existing methods alleviate this phenomenon through token-level entropy regularization or negative sample reinforcement, yet these works have treated responses as atomic units distinguished only by binary correctness labels, overlooking substantial semantic heterogeneity among responses under the same label—among incorrect responses, the degree and type vary significantly; among correct responses, both conventional and novel strategies exist. Uniformly reinforcing all correct responses leads to policy convergence, while uniformly penalizing all incorrect ones may suppress promising directions. Building on this insight, we propose Semantic Diversity-Aware Exploration (SDAE), which leverages geometric structures in semantic space to guide exploration. SDAE computes the semantic centroid of positive samples and modulates advantage functions based on each response's distance to this centroid, encouraging novel correctness, penalizing divergent errors, while preserving improvement potential in locally flawed responses. Experiments demonstrate that SDAE consistently outperforms strong baselines across competition-level reasoning benchmarks, achieving a 13.3% improvement in Pass@256 on AIME25. Our code is available at https://anonymous.4open.science/r/SDAEcode-3553.
PaperID: 5994, Poster
Abstract: Pipeline-parallel distributed optimization is essential for large-scale machine learning but is challenged by significant communication overhead from transmitting high-dimensional activations and gradients between workers. Existing approaches often depend on impractical unbiased gradient assumptions or incur sample-size memory overhead. This paper introduces Clapping, a Communication compression algorithm with LAzy samPling for Pipeline-parallel learnING. Clapping adopts a lazy sampling strategy that reuses data samples across steps, breaking sample-wise memory barrier and supporting convergence in few-epoch or online regimes. Clapping comprises two variants including Clapping-FC and Clapping-FU, both of which achieve convergence without unbiased assumption for compressed gradient, effectively addressing compression error propagation in multi-worker settings. Numerical experiments validate the performance of Clapping across different tasks.
Abstract: State estimation is a critical task in scientific, engineering and control applications. Since the reliability of reconstructions can depend on the number and position of sensors, Optimal sensor placement (OSP) is essential in scenarios where measurements are sparse and expensive. However, classical OSP approaches rely on Gaussian assumptions and are consequently unable to account for complex distributions encountered in many real-world systems. Generative-model-based reconstruction using sensor guided diffusion posterior sampling (DPS) has emerged as a promising technique for reconstructing states from highly complex distributions. Existing approaches to sensor selection either choose an unrealistically large number of sensors or employ strategies that emulate classical OSP methods. This results in a mismatch, wherein new models are paired with classical OSP tools, and motivates the need for fundamentally new ideas towards OSP that match the recent advances made in powerful recovery models. In this work, we introduce a distribution-free sensor placement framework based on the Christoffel function. Our main theoretical contributions introduce a mathematical formulation of optimal sampling and recovery guarantees for posterior sampling with arbitrary sensors and signal distributions. We use these to derive a new OSP strategy with non-asymptotic bounds on the number of sensors needed for recovery. Building on this, we develop Christoffel-DPS, with both offline and online variants, that implements nonparametric realizations of Christoffel sampling for generative models. As we show, Christoffel-DPS outperforms Gaussian OSP baselines and existing generative-model-based placement methods, validating that distribution-free sensing is both theoretically principled and practically superior. The framework is model agnostic, and we demonstrate its application to a range of unconditional DPS and flow matching models on structurally non-Gaussian benchmarks, showing the efficacy of Christoffel-DPS in low sensor budget regimes.
PaperID: 5996, Poster
Abstract: The Message Passing Neural Networks (MPNNs) and Graph Transformers (GTs) have emerged as two dominant paradigms for graph representation learning. However, the expressive power of standard MPNNs is fundamentally bounded by the 1-dimensional Weisfeiler-Lehman (1-WL) test, while GTs lack the inductive bias for graph structure. To enhance structural representation, existing methods typically resort to subgraph-based aggregation or coordinate-based positional encodings. However, the former suffers from prohibitive computational memory overheads, while the latter is limited by rigid reference frames that fail to distinguish fine-grained local topology. To address these limitations, we introduce a novel structural encoding framework based on contextual distributions. Specifically, we move beyond fixed coordinate systems to capture intrinsic local topology by encoding structural information through a compact distributional statistic, i.e., the entropy of the node context. Furthermore, instead of relying on relative positioning, we introduce a kernelized mechanism to encode relational similarity by quantifying the structural affinity between node contexts. Extensive experiments on both synthetic and real-world datasets demonstrate that the proposed framework achieves superior performance, striking a favorable balance between effectiveness and efficiency.
PaperID: 5997, Poster
Authors: Bishal Swain, Joonhyeon Bae, Jaepil Ko
Abstract: Continuous-time recurrent models update hidden states as observations arrive, making them suitable for irregular and non-stationary temporal data. Liquid Neural Networks (LNNs) extend this formulation through input-dependent state dynamics, but they still rely on a single evolving state to integrate new observations and retain earlier context. Under long temporal gaps or missing inputs, earlier evidence can weaken as the hidden state continues to evolve. We propose Active Memory Feedback Loop (AMFL), a memory-augmented LNN architecture that feeds associative retrieval back into the recurrent computation. AMFL combines a learned Short-Term Memory (STM) that stores global sequence prototypes with a sequence-local Iconic Memory (IM) that adapts through novelty-gated, gradient-free updates. After each liquid refinement step, IM retrieves an associative representation, forms a residual correction, and writes the correction into a feedback buffer that conditions the following recurrent update. Memory therefore influences the evolving liquid representation rather than only augmenting the final prediction head. We evaluate AMFL on seven time-series benchmarks covering classification, regression, continual learning, and robustness to missing inputs. AMFL obtains the best result on six of seven benchmarks and remains competitive on the remaining task. Ablations show that the gains arise from looped memory feedback, controlled IM adaptation, and residual correction. Additional missing-input and long-sequence forecasting experiments further support the role of associative feedback in preserving task-relevant context under partial observability.
Abstract: In this work, we show that natural policy gradient (NPG), a core algorithm in reinforcement learning, admits an exact interpretation as a smoothed and averaged form of policy iteration. Specifically, we introduce doubly smoothed policy iteration (DSPI), a Bellman-operator framework in which each policy is obtained by applying a regularized greedy step to a weighted average of past Q-functions. DSPI includes policy iteration, dual-averaged policy iteration, NPG, and more general policy dual averaging methods as special cases. Using only monotonicity and contraction of smoothed Bellman operators, we prove distribution-free global geometric convergence of DSPI. Consequently, standard NPG and policy dual averaging achieve an iteration complexity of \mathcalO((1-\gamma)^-1\log((1-\gamma)^-1\epsilon^-1)) for computing an \epsilon-optimal policy, without modifying the MDP, adding regularization beyond the mirror map inherent in the update, or using adaptive, trajectory-dependent stepsizes. For the unregularized greedy case, we also prove finite termination of dual-averaged policy iteration. The same Bellman-operator framework extends to discounted MDPs with linear function approximation and to stochastic shortest path problems.
PaperID: 5999, Poster
Authors: Taiqin Chen, Hao Sha, Xiaochen Feng, Chunyu Chen, XiangYu Chen, Ke Chen, yongbing zhang
Abstract: Hyperspectral single-source domain generalization aims to train a robust model capable of overcoming domain shifts using only one available source domain for cross-scene hyperspectral image (HSI) analysis. Existing methods typically adopt data augmentation to broaden decision boundaries or utilize style transfer techniques for test-time alignment. However, these studies either fail to completely bridge the domain gap or inadvertently disrupt the discriminative spectral information, leading to suboptimal generalization and degraded classification performance on the target domain. To overcome these limitations, we propose a novel method, termed the domain memory retention framework, which optimizes domain memory during training to explicitly encode the semantic structure of the source domain. During inference, the target features are projected to the source distribution by leveraging this memory, thereby skipping domain shifts and improving generalization performance across unseen scenes. Specifically, the memory is formulated as a codebook, consisting of multiple semantic codewords. We further develop a semantic richness constraint and a perturbation robustness constraint to enhance representation diversity while empowering domain invariance for the memory. Extensive experiments conducted on three remote sensing datasets demonstrate that the proposed method outperforms state-of-the-art methods.
Abstract: Prior-fitted networks (PFNs) are a promising class of tabular foundation models that perform in-context learning, whereby the entire labelled training set is supplied as context, and predictions for test queries are produced in a single forward pass. However, the quadratically scaling self-attention mechanism in many PFN architectures makes inference prohibitive for very large training datasets. We propose CRUMB (Clustered Retrieval Using Minimised-MMD Batching), a three-stage inference wrapper that (i) clusters the test queries, (ii) selects a small, distributionally matched training subset for each cluster by greedily minimising the maximum mean discrepancy (MMD), and (iii) runs exact PFN inference on each reduced-context batch. CRUMB is architecture-agnostic and requires no retraining. On the 51-dataset TabArena benchmark, evaluated across three PFN architectures (TabPFNv2, TabICLv1, TabICLv2), we show that CRUMB outperforms similar state-of-the-art context selection strategies. We also show that CRUMB is resilient to covariate drift, as the MMD-minimisation step naturally helps align the training context distribution to match the current test batch distributions.
Abstract: Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficiency by routing queries to the \emphspecialist who can answer most effectively at the lowest cost. Routing requires estimating each specialist's expected return, but this value estimation has a cost. Cheap estimators (e.g., embedding-based predictors) are fast but noisy, while accurate estimators (e.g., fine-tuned models with access to retrieval results or partial reasoning traces) are expensive. We formalize this tradeoff as an instance of Pandora's Box, the classical problem of optimal search with costly inspection. Under a Gaussian signal model, the resulting policies have closed-form value-of-information expressions that determine, for each specialist and input, whether refining the value estimate is worth its cost. We call the centralized policy Pandora's Router. We extend this to a decentralized setting, Pandora's Bidder, where specialists independently decide whether to invest in self-assessment before accepting an offered price to claim a query. Experiments across three domains---a standard multi-LLM benchmark, retrieval-augmented specialists, and LLMs with variable inference-time reasoning---show that Pandora's Router matches the routing quality of exhaustive estimation, while querying the expensive estimator far less often. In the decentralized setting, value-of-information reasoning improves allocative efficiency when competing estimates are accurate; when competing estimates are noisy, however, it can increase the strategic specialist's utility at the expense of
PaperID: 6002, Poster
Abstract: Large Language Models (LLMs) demonstrate remarkable reasoning capabilities, yet the internal mechanisms driving these multi-step processes remain opaque. Existing mechanistic interpretability approaches often rely on superficial text-pattern co-occurrences or capture only short-term effects, struggling to trace the long-horizon, outcome-driven influence of internal components such as neurons or sparse features in multi-step reasoning. In this paper, we introduce Integrated Policy Gradient (IPG), a novel, training-free framework that brings policy-based attribution into mechanistic interpretability for LLM reasoning. Grounded in principles, it backpropagates outcome-level, non-differentiable behavioral signals (e.g., reasoning outcome correctness) through entire inference trajectories, utilizing path integration to isolate outcome-relevant mediators at component level for reliable attribution. This provides a linear approximation of Causal Mediation Analysis (CMA) in reasoning. Empirical evaluations demonstrate that IPG achieves precise localization and reveals mechanistic insights by identifying sparse, concentrated subsets of internal components associated with specific reasoning behaviors. Our method also enables modulation of reasoning capabilities, offering a powerful and reliable framework for understanding LLM reasoning.
PaperID: 6003, Poster
Abstract: Time series foundation models (TSFMs) have achieved notable progress in classification tasks through large-scale pretraining. However, existing methods rely on implicit parameter fine-tuning, where task-relevant information is entangled with abundant irrelevant knowledge in vast parameter spaces. This may hinder TSFMs from precisely activating useful knowledge for downstream classification. To this end, we propose REKIN, a plug-and-play knowledge-enhanced framework for time series classification built upon pretrained TSFMs within a retrieval-to-recognition paradigm. Unlike traditional parameter-only approaches, REKIN introduces a decoupled query encoder to extract discriminative patterns and retrieve instance-level numerical knowledge from a constructed knowledge base. The retrieved explicit knowledge is then fused with the TSFM’s implicit parametric embeddings, enabling joint knowledge transfer at both the parameter and feature levels. In this way, REKIN allows TSFMs to recall task-relevant patterns learned during pretraining, improving the efficiency of knowledge utilization in downstream transfer. Experiments on 158 time series datasets show that REKIN consistently improves the classification performance of state-of-the-art TSFMs. Ablation and knowledge-enhanced studies further validate the effectiveness of the proposed framework.
PaperID: 6004, Poster
Abstract: Zero-sum games are the textbook model of strategic competition—yet they can fail to capture conflict even in simple settings. Harmonic games provide a more robust framework for opposed interests: they are invariant under strategic equivalence, and naturally extend beyond pairwise competition. This generality comes at a structural cost. We show that, in odd-dimensional harmonic games, the set of totally mixed Nash equilibria is a generically non-convex real algebraic variety of dimension at least one, extending to the boundary of the strategy space. As a result, learning becomes delicate: although finely tuned methods are known, whether adaptive, parameter-agnostic learning is possible has remained open. We resolve this question in the affirmative. We introduce a flexible extrapolation-based variant of follow-the-regularized-leader (FTRL+) that converges to Nash equilibrium while achieving order-optimal regret: \mathcalO(1) in self-play and \mathcalO(\sqrtT) against arbitrary opponents. The method is fully adaptive—each player updates from locally observable information alone, without access to global problem parameters or shared signals—yielding the first parameter-agnostic learning dynamics provably convergent in this class.
PaperID: 6005, Poster
Abstract: We show that frontier AI models exhibit loyalty capture---a form of structural sycophancy in which a reporting relationship to a manager with misaligned incentives is sufficient to bias the model's recommendations, without any explicit instruction, user feedback, or adversarial prompting. In 5,757 trials across eight models from eight AI laboratories, CEO-reporting models cut investment to boost the CEO's bonus by 8.1 percentage points more when the bonus is at stake. This same effect is also predictably absent from a metric-irrelevant placebo. Delivered rationales reveal motivated reasoning: models adopt the bonus target, calculate how their recommendation achieves it, and dismiss long-term value, while concealing their recognition of the conflict from the recommendation recipient. Designed interventions (conflict disclosure, stakeholder broadening) eliminate ~60% of the effect; passive oversight does not. Taken together, the results indicate that AI models deployed within organizational hierarchies can develop and act on loyalty to their supervisor at the expense of the organization they are meant to serve.
PaperID: 6006, Poster
Abstract: As state-of-the-art text-to-image flow models have matured to deliver near-photorealistic quality, controlling what they generate -- e.g., inhibiting harmful content while promoting benign alternatives -- has become a central challenge. The current steering paradigm consists of adding a global steering vector to selected model activations. While functional, a fixed vector applied example-agnostic and uniformly along the entire trajectory cannot adapt to the changing state of the generation, and causes uncontrolled global changes beyond the targeted concepts. We introduce Steering Fields, a generalization of steering vectors that adaptively re-estimates the steering direction at each step of the generative process. Steering Fields operate on latent representations, expose a continuous trade-off between steering strength and content preservation, and are compositional, enabling the simultaneous induction and inhibition of concepts, setting a new state of the art on safety steering benchmarks. Although our formulation prescribes no explicit spatial mask or object-level prior, the trajectory-adaptive estimation naturally preserves local structure, in a manner reminiscent of image editing. In fact, when applied in the VAE latent space, Steering Fields can serve as a structure-preserving image-editing technique that achieves state-of-the-art semantic fidelity (CLIP, VQAScore), while remaining model-agnostic and inversion-free
PaperID: 6007, Poster
Abstract: Retrieving unknown atomic structures from observable analytical spectra or images remains a long-standing challenge across natural sciences. However, retrieval accuracy of existing retrieval methods on analytical data remains suboptimal because they have overlooked the underlying periodic quantum mechanical perturbations of non-equilibrium atomic structures behind analytical data. This paper proposes a periodic complex stochastic process (PCSP) that models such periodic perturbations and establishes theoretical backgrounds of periodic stochastic process in the complex-valued domain, including its sample diversity, process length, and periodicity. Finally, we develop a complex-valued cross-modal retrieval (CVCR) by integrating PCSP with cross-modal retrieval frameworks. CVCR outperformed existing cross-modal retrieval methods in cross-modal retrieval tasks of real-world analytical chemistry. Moreover, CVCR achieved state-of-the-art retrieval accuracy in zero-shot settings on 40 million of molecules.
Authors: Hyung K Kim, Byungchan Hwang, HAKGU KIM
Abstract: Recent progress in speech-driven 3D facial animation has improved vertex level reconstruction quality, but speech-consistent visible articulation remains difficult. This is because speech production follows structured and constrained lip jaw coordination and the mapping from acoustics to motion is inherently one to many. Motivated by the structured patterns of visible articulation, we propose a novel articulation-aware framework that models visible speech through directional lip motions and composes them into geometry consistent facial deformation. To represent visible articulation with three canonical directional motions, spreading, opening, and protrusion, we propose a Speech Articulatory Memory (SAM) that captures the correspondence between speech and directional articulatory motions under phonetic context through key-value memory structure based retrieval and decoding. Then, a Topology-aware Articulatory Composition (TAC) integrates the predicted directional motions under mesh topology to produce coherent 3D facial motion. Experiments on VOCASET and TFHP show that our method achieves the-state-of-the-art performance on standard reconstruction metrics and improves visible articulatory distance and velocity fidelity for key lip factors, while a user study confirms clear preference in lip sync and realism.
PaperID: 6009, Poster
Authors:
JunXi Wang, Jiayi Zhu, Te Sun, Chen Zhang, Siyuan Li, Xuyang Liu, Zichen Wen, Xiaobing Tu, Jinkui Ren, Xiantao Zhang, Ziqi Yuan, Linfeng ZhangAbstract: Agent memory systems have demonstrated significant potential in tasks such as long-term dialogue, personalized assistants, and video understanding. However, as inference progresses, continuously accumulated memory imposes substantial storage and retrieval burdens. To address this issue, we propose MemForest, a general memory compression framework adaptable to various agent memory systems. Specifically, MemForest leverages both global semantic similarity and local temporal continuity of memory events to partition the historical memory into a set of event-centric independent units. For each independent unit, the framework constructs a maximum spanning tree structure, referred to as an EventTree, and performs progressive merging by iteratively selecting high-weight edges, thereby effectively compressing redundant memory nodes and reducing storage overhead. In addition, we introduce an anchor-guided propagation retrieval mechanism, which retrieves more relevant memory nodes from the temporal neighborhoods of key memory nodes, thereby enabling more accurate memory retrieval. Extensive experiments demonstrate the effectiveness of MemForest. Under the unimodal Mem0 framework, across three benchmarks (LoCoMo, LongMemEval, and PersonaMem), MemForest preserves 97.1% of the original performance while compressing 50% of historical memory, achieving a 1.89× retrieval speedup. Under the multimodal M3-Agent framework, across two benchmarks (M3-Bench-robot and M3-Bench-web), MemForest retains 99.7% of the original performance under a 50% compression ratio, while achieving a 2.24× retrieval speedup. Our code is available in the supplementary materials, and all data will be released on GitHub.
Abstract: Positional encoding in transformers is commonly implemented through positional embeddings, attention masks, or bias terms, but formal connections between these mechanisms remain limited. We study attention with positional bias through the lens of locality-sensitive hashing (LSH), focusing on Attention with Linear Biases (ALiBi). We show that the ALiBi bias matrix is the expectation of contiguous block-diagonal binary masks induced by a ``positional LSH'' scheme. The empirical mean of masks sampled from this scheme yields spectral norm and max-norm approximation guarantees with bounded block sizes with high probability. This structural theorem implies a uniform approximation theorem for ALiBi-biased attention: with high probability over the sampled masks, the approximate attention output is accurate simultaneously for all query-key-value inputs and can be computed in near-linear time in the context length, reducing long-context ALiBi to a collection of randomized short-context regular (positionally unbiased) attention operations. Conceptually, this connects positional bias, masks, and positional embeddings in a single formal framework and suggests an approach to efficient ALiBi-biased attention. Experiments on large language models validate our theoretical findings.
PaperID: 6011, Poster
Abstract: The rapid emergence of heterogeneous large language models (LLMs) has highlighted the limitations of static single-model inference systems. Although model routing alleviates this issue by assigning the most suitable model to each query, existing methods often struggle to stably align query semantics with actual model performance and remain constrained by the capability ceiling of a single model. On the other hand, although multi-model aggregation can improve performance on complex queries, it also introduces extremely high computational overhead for routine tasks. In this paper, we propose PARS, a model confidence estimation method for dynamic LLM routing and aggregation. PARS first decouples the feature representation space and applies structured smoothing to local performance signals, producing more discriminative confidence score estimates for candidate models. On this basis, PARS represents inference modes such as direct routing, self-consistency sampling, and multi-model voting as test-time compute allocation actions under a budget constraint. By extracting state features such as routing confidence distributions and predictive uncertainty, a lightweight selector adaptively assigns an inference strategy to each query. Extensive experiments show that PARS improves average accuracy over the best baselines by 1.02% and 1.36% in the single-model and multi-model settings, respectively, while reducing average inference cost by 1.73x and exhibiting stable out-of-distribution generalization.
PaperID: 6012, Poster
Abstract: Understanding how gradient descent learns features remains a central challenge in neural network theory. We study this question in a minimal multi-index setting that already exhibits rich multi-phase gradient dynamics: a well-specified two-neuron ReLU teacher-student model under isotropic Gaussian inputs, trained by population gradient flow from small random initialization. We show that the dynamics are organized by two directions: the “easy” bisector direction, which carries the leading signal, and the “hard” splitting direction, which governs specialization. To characterize the loss structure and learning behavior, we analyze the population squared-loss landscape and show that every nonzero critical point is either a global minimum or a saddle. We then track the gradient-flow trajectory, showing that student neurons first collapse toward a bisector saddle before escaping and specializing to teacher neurons. Our results provide both a landscape and dynamics account of the multi-phase symmetry-breaking behavior that arises in a simple multi-index model and under standard algorithmic and architectural choices.
Authors:
Shiqi Liu, Zeyu He, Guojian Zhan, Letian Tao, Zhilong Zheng, Jiang Wu, yinuo Wang, Yang Guan, Kehua Sheng, Bo Zhang, Keqiang Li, Jingliang Duan, Shengbo Eben LiAbstract: Reinforcement Learning (RL) has significantly improved large language model reasoning, but existing RL fine-tuning methods rely heavily on heuristic techniques such as entropy regularization and reweighting to maintain stability. In practice, they often suffer from late-stage performance collapse, leading to degraded reasoning quality and unstable training. We identify a key factor behind this instability: a small fraction of tokens, termed spurious tokens (around 0.01%), which contribute little to the reasoning outcome but receive disproportionately amplified gradient updates due to inheriting the full sequence-level reward. We present a unified framework for evaluating token-level optimization impacts across spurious risk, gradient norms, and entropy changes. Building the analysis of token characteristics that severely disrupt optimization, we propose the Silencing Spurious Tokens (S2T) mechanism to efficiently suppress their gradient perturbations. Incorporating this mechanism into a group-based objective, we propose Spurious-Token-Aware Policy Optimization (STAPO), which promotes stable and effective large-scale model refinement. Across six mathematical benchmarks and three model scales, STAPO consistently demonstrates superior entropy stability and achieves average accuracy improvements over baselines including GRPO, 20-Entropy, and JustRL under two widely-adopted evaluation settings.
PaperID: 6014, Poster
Abstract: Low-Rank Adaptation (LoRA) is a popular parameter-efficient fine-tuning (PEFT) method for fine-tuning of large language models (LLMs), especially in federated learning, due to its strong performance and communication efficiency. In practice, LoRA is applied to the query and value projections of transformer layers without considering the distinct roles of these components. In this work, we analyze the impact of these projections in federated fine-tuning of LLMs using LoRA and observe that the value projection converges significantly faster than the query projection. Motivated by this finding, we propose FedLoVA, a federated fine-tuning framework that trains LoRA adapters on both query and value projections locally while sharing only the value projection with the server. This approach reduces communication overhead while improving performance and convergence efficiency. Experiments on natural language understanding, multilingual sentiment analysis, and generation tasks show that FedLoVA reduces the communication cost by up to 7.8× compared to existing federated LLM fine-tuning baselines while achieving comparable performance.
PaperID: 6015, Poster
Abstract: Scaling video generation from short to long durations has attracted growing attention as a way to democratize video creation. The central challenge lies in long-term consistency, which requires stable rollout over extended horizons and accurate visual memory of previously generated elements. However, the scarcity of high-quality long video data remains a fundamental bottleneck, forcing prior methods to rely on suboptimal synthetic data that either limits dynamic range and visual memory or suffers from a large simulation-to-reality gap. To address this, we explore a new generation paradigm, Temporal Reasoning-based Rendering (TRR), which better exploits existing short video data for consistent long video generation. Under TRR, we propose PlasticMem, which plans contents from scripts, then reasons about and independently renders each action to form a consistent long video. Benchmark results show that PlasticMem, trained on short videos, achieves stable minute-scale long-horizon generation and visual memory.
PaperID: 6016, Poster
Abstract: While fine-tuning is the standard recipe for adapting Time Series Foundation Models (TSFMs) to downstream tasks, we reveal that it triggers a critical yet overlooked problem: structural forgetting. Distinct from widely studied point-wise error degradation, structural forgetting represents a fundamentally more destructive phenomenon: the erasure of universal structural patterns acquired during pre-training. Unlike point-wise errors, this structural collapse renders models incapable of supporting basic structure-dependent decisions (e.g., trend-based trading), a critical vulnerability that existing forgetting mitigation strategies fail to address. To bridge this gap, we first introduce a diagnostic framework equipped with a novel Temporal-Frequency Attribution analysis, mechanistically revealing that fine-tuning unnecessarily overwrites a sparse set of pattern-critical parameters. We then propose NeST, a post-hoc framework that precisely restores these components while compensating surrounding weights to preserve task adaptation. Extensive experiments across 3 TSFMs and 9 datasets demonstrate that NeST recovers an average of 85% of structural awareness without sacrificing task performance, offering a new paradigm for structurally-preserved TSFM adaptation.
Abstract: CoT prompting improves LLM accuracy on complex tasks but often increases token usage and inference cost. Existing "Budget Forcing" methods reduce cost via fine-tuning with heuristic length penalties, suppressing both essential reasoning and redundant filler. We recast efficient reasoning as a lossy compression problem under the IB principle, and identify a key theoretical gap when applying naive IB to transformers: attention violates the Markov property between prompt, reasoning trace, and response. To resolve this issue, we model CoT generation under the CIB principle, where the reasoning trace Z acts as a computational bridge that contains only the information about the response Y that is not directly accessible from the prompt X. This yields a general Reinforcement Learning objective: maximize task reward while compressing completions under a prior over reasoning traces, subsuming common heuristics (e.g., length penalties) as special cases (e.g., uniform priors). In contrast to naive token-counting approaches, we introduce a semantic prior that measures token cost by surprisal under a language model. Crucially, the prior is queried only for token-level log-probabilities, adding negligible overhead to the training loop. Empirically, our CIB objective prunes reasoning redundancy while preserving fluency and logic, improving accuracy at moderate compression and enabling aggressive compression with minimal accuracy drop. These gains generalize across model families and task domains, confirming CIB as a domain-agnostic CoT compression framework.
Authors: Rohan N Pradhan, Steve Goley
Abstract: Language models increasingly act as epistemic proxies, synthesizing evidence from multiple sources to inform decisions. Whether they evaluate the quality of that evidence, or merely aggregate it based on surface presentation, remains poorly understood. We show that models possess the capability to detect fabricated statistics (correct identification rates of 0.76--1.00 for methodology in isolation) but do not recruit this capability during multi-source synthesis, producing similar numeric estimates whether the statistics are fabricated or valid. Specifically, source influence is governed by a methodology-register gate that responds to the distributional register of analytical text but not to numeric validity: for example, statistically impossible confidence intervals receive the same weight as valid ones. The behavioral dissociation replicates across five models from three families (Claude, Qwen, OLMo) and three professional domains. Mechanistic analyses, including causal tracing, linear probes, and component-level attribution, converge on the same account: the model encodes and causally uses a methodology-register representation that transfers across domains (probe AUC 0.83--0.92), while numeric-validity signals, decodable in isolation, are suppressed to chance during multi-source synthesis. Prompting-based mitigations, even an oracle checklist naming the exact statistical checks, produce blanket skepticism rather than selective discernment, and the post-training pipelines we examine reinforce the stylistic shortcut without building numeric verification. Unlike sycophancy, which tracks user preference, this failure tracks whether a source presents as analytically credible, not whether its claims are internally consistent. We term this epistemic alignment: like preference and safety alignment, the question is not capability but deployment.
PaperID: 6019, Poster
Authors: Zhang Xiaoyan, Chenxu Pang, Xiaojie Wang
Abstract: Although score-based generative models (SGMs) have achieved remarkable success in real-world sample generation tasks, it is still far from sufficient to understand them from a mathematical perspective. In this paper, we aim to provide enhanced convergence guarantees for two existing SGMs in \mathcalW_2-distance beyond log-concavity. More precisely, we develop a novel framework of error analysis for Euler-Maruyama (EM) and Poisson midpoint (PM) time discretization schemes for SGMs. Under a non-log-concavity condition, we show that \tilde\mathcalO(\sqrtd/\epsilon) iterations suffice to approximate the target distributions in \epsilon-accuracy for the classical EM-based SGM, considerably improving upon the existing iteration complexity \tilde\mathcalO(d/\epsilon^2). Notably, the novel framework of error analysis enables us to establish an \tilde\mathcalO(\sqrtd/\epsilon^2/3) iteration complexity for the PM-based SGM in \epsilon-accuracy, significantly outperforming current state-of-the-art Wasserstein convergence guarantees for SDE-based diffusion samplers.
PaperID: 6020, Poster
Abstract: Neural Hamiltonian models provide a principled approach to learning dynamical systems by embedding the symplectic structure of physical evolution into the model class. However, this geometric prior alone has not made Hamiltonian neural networks efficient long-horizon simulators: an M-step trajectory is still typically generated by M sequential updates, and common training objectives based on force matching or one-step prediction provide only local supervision. We introduce Symplectic Parallel Scan (SPS), a neural Hamiltonian framework that algebraizes the learned dynamics through a finite-dimensional Poisson algebra of Hamiltonian generators. SPS constructs a Neural--Poisson Composition DAG closed under projected products and Poisson brackets, whose elements induce symplectic Hamiltonian drivers organized as a Lie-group walk. The associativity of this flow composition enables a parallel prefix-scan over Hamiltonian drivers, reducing the sequential depth of trajectory propagation from \mathcal O(M) to \mathcal O(\log M) while keeping each prefix inside the symplectic flow family. The resulting model is trained through a global trajectory-matching objective rather than only local force labels. Across quantum spin and molecular dynamics benchmarks, SPS achieves substantial wall-clock acceleration over sequential HNN baselines, improves long-horizon prediction accuracy, and reduces energy drift and symplectic violation under extrapolation.
PaperID: 6021, Poster
Abstract: Noisy-label detection aims to identify incorrectly labeled training samples so that their adverse impact on model learning can be mitigated. Existing methods typically rely on loss, prediction confidence, or gradient magnitude signals, which primarily capture the scale of prediction errors. While effective, such scalar signals overlook the geometric structure and temporal dynamics of gradient evolution during training. In this work, we show that gradient magnitude provides a useful entry point for analysing the gradient distortion induced by noisy-label samples, enabling us to quantify their deviation from the gradients that would be generated by clean labels. In this light, we propose GradTrack, a simple and effective framework for detecting noisy-label samples by estimating class-wise gradient directions likely to be induced by clean labels and measuring the alignment of training samples with these directions. In each training epoch, GradTrack partitions samples based on gradient magnitude and approximate clean-label gradient directions to compute a class-wise gradient misalignment score. These scores are used to rank training samples, which are then aggregated across training epochs to form a temporal rank trajectory. Finally, a Gaussian mixture model is applied to detect noisy-label samples. Extensive experiments on both synthetic and real-world noisy-label benchmarks demonstrate that GradTrack achieves strong detection performance, particularly under instance-dependent noise, and can further improve existing learning-with-noisy-labels pipelines as a plug-in sample selection module.
PaperID: 6022, Poster
Authors: Shailendra Bhandari, Alex Szorkovszky, Anis Yazidi, Pedro G Lind
Abstract: How search strategies evolve in sparse, depletable landscapes remains a central question in foraging theory. We study this problem with an evolutionary simulation in which agents forage on a two-dimensional toroidal lattice containing non-renewable resources distributed either uniformly or as Lévy dust. Each agent carries a heritable genome encoding step lengths, velocities, and turning angles, and selection acts on a fitness function, derived from first principles, combining energetic gain, movement cost, and coverage efficiency. By allowing movement traits to evolve without imposing a prescribed power-law step-length distribution, we test whether selection recovers a strict Lévy-like random walk, similar to the spatial distribution of resources, or instead favors an alternative search mechanism. Strikingly, our results indicate that, in the finite depletion-driven landscapes considered here, evolved search is more consistent with intermittent dynamics than with strict scale-free Lévy motion. To characterize the effective dynamics of the evolved trajectories, we take the average coefficient of determination when fitting second- and fourth-order displacement moments to intermittent search and Lévy-walk models. While a Lévy-like random walk fits very well with the numerical results from the evolutionary search (R^2>0.9), the intermittent search achieves a closer fit, with fitted coefficients of determination (R^2>0.99) for all resource distributions. Evolution rapidly reshapes the movement genome toward short displacements while retaining a sparse tail of longer relocations, consistent with local exploitation punctuated by occasional transfer. The framework provides a controlled setting for studying how search rules emerge under resource limitation and may inform resource-constrained exploration in autonomous systems.
Abstract: Distributed learning algorithms are vulnerable to adversarial nodes, a.k.a. Byzantine failures. To solve this issue, robust algorithms have been developed, which typically replace parameter averaging by robust aggregations. While generic conditions on these aggregations exist to guarantee the convergence of (Stochastic) Gradient Descent (SGD), the analyses remain rather ad-hoc. This hinders the development of more complex robust algorithms, such as accelerated ones. In this work, we show that Byzantine-robust distributed optimization can, under standard generic assumptions, be cast as a general optimization with inexact gradient oracles, an active field of research. This allows to obtain state-of-the-art results for Byzantine-robust optimization from general inexact first-order analyses. We first show that inexact GD on top of standard robust aggregation procedures obtains optimal asymptotic error in the Byzantine setting. Going further, we study an algorithm for Optimization under Similarity, in which the server leverages an auxiliary loss function that approximates the global loss. Then, we introduce an Accelerated Extra-gradient method, that yields acceleration in both the standard and similarity settings. We first give new general convergence results for these inexact schemes, and then instantiate these results in the Byzantine setting through our reduction. Both algorithms drastically reduce the communication complexity compared to previous methods, as we show theoretically and empirically.
Abstract: Large Reasoning Models (LRMs) achieve impressive performance on complex reasoning tasks via Chain-of-Thought (CoT) reasoning, which enables them to generate intermediate thinking tokens before arriving at the final answer. However, LRMs often suffer from significant overthinking, spending excessive compute time even after the answer is generated early on. Prior work has identified the existence of an optimal reasoning length such that truncating reasoning at this point significantly shortens CoT outputs with virtually no change in performance. However, determining optimal CoT lengths for practical datasets is highly non-trivial as they are fully task and model-dependent. In this paper, we precisely address this and design Terminator, an early-exit strategy for LRMs at inference to mitigate overthinking. The central idea underpinning Terminator is that the first arrival of an LRM's final answer is often predictable, and we leverage these first answer positions to create a novel dataset of optimal reasoning lengths to train Terminator. Powered by this approach, Terminator achieves significant reductions in CoT lengths of 14%–55% on average across four challenging practical datasets: MATH-500, AIME 2025, HumanEval, and GPQA, while outperforming current state-of-the-art methods and reducing inference latency by more than 2× compared to the original LRM.
PaperID: 6025, Poster
Authors:
Xiaobing Yu, Peijie Qiu, Jin Yang, Zhaoqi An, Weiwei Ma, Xiao Wu, XUANZHAO DONG, Wenhui Zhu, Xiwen Chen, Xiaoqi Zhao, Xiaofeng LiuAbstract: Promptable segmentation foundation models such as SAM and its variants achieve strong performance with informative prompts, yet exhibit a striking prompt asymmetry where dense prompts like bounding boxes yield accurate segmentation, whereas sparse point prompts substantially degrade performance. This performance gap is widely observed, yet its internal decoder mechanism remains underexplored. In this work, we revisit the failure of point prompts from a causal perspective. By adapting activation patching to SAM-style decoders with a token-aligned intervention scheme, we localize the earliest recoverable divergence to a single early operation, i.e., the first decoder image-to-token cross-attention update, where prompt-conditioned information first enters the image stream. A sparse autoencoder further reveals that this mediation is concentrated in approximately 50 feature directions out of 4,096, and reverse corruption confirms that the identified layer is the entry point of a propagated image-stream route rather than a locally sufficient state. Guided by this localization, we propose the Causal Bottleneck Adapter (CBA), a lightweight residual module placed at the identified entry layer with the base model frozen. Evaluated on 97,996 medical image-target pairs across CT, MRI, ultrasound, and endoscopy, and validated across various SAM-family models, CBA closes 0.891 of the box-to-point gap with fewer than 9K trainable parameters, outperforming prior repair baselines up to 50× larger.
PaperID: 6026, Poster
Authors: Firas Bayram, Maxime Cordy
Abstract: Adversarial detection on evolving attributed graphs faces two challenges: fraud labels arrive late or not at all, and natural distributional drift erodes any trained decision boundary. Modern detectors can reach strong detection rates when labels are available and the distribution is stationary; sustaining that quality without labels and under drift is the open problem. We identify \emphfeature-context consistency --- the alignment between a node's attributes and its aggregated neighborhood attributes --- as a single structural signal that addresses both. We formalize it as the Context Boost (CB) score within a Restricted Boltzmann Machine framework (CB-RBM) and prove two complementary guarantees: a strict separation between legitimate and adversarial CB distributions, with gap growing linearly in model capacity, whenever the adversary lacks legitimate neighborhood context; and a Wasserstein-Lipschitz stability bound under natural drift that prevents false alarms without retraining. We validate on XBlock Ethereum phishing (~153K nodes) and Reddit banned-user detection (~11K nodes), with two further datasets in the appendix spanning diverse domains. With zero labels, CB-RBM achieves AUROC~\geq 0.96 on four adversarial attacks and keeps FPR close to its calibrated 5% target under drift --- outperforming supervised GNNs (hundreds of labels), unsupervised graph-anomaly baselines, and a standard RBM whose FPR collapses to 89--100%. What typically requires separate mechanisms --- label-free operation and drift stability --- here follows from a single structural regularity of the data.
PaperID: 6027, Poster
Authors: Ilyes Hammouda, Stanislav Minsker, Mohamed Ndaoud
Abstract: We investigate the problem of uncertainty quantification in the mean estimation framework. Given an i.i.d. sample from a distribution with unknown mean \mu and variance \sigma^2, we want to construct an estimator \hat\mu_n and provide a computable and size-optimal upper bound for the error |\hat \mu_n - \mu|. While estimators possessing strong deviation guarantees are well known, no data-dependent non-asymptotic upper bounds for their performance exist in general. We show that this gap can be bridged by introducing a parameter that characterizes the "effective sample size" available for uncertainty quantification. This parameter captures the difficulty of the problem, interpolating between "easy" cases where sub-Gaussian confidence intervals exist and "hard" cases where their construction is impossible. Using this characterization, we design confidence intervals of optimal length that are fully adaptive to the unknown variance. Numerical experiments confirm that our approach maintains nominal coverage even in asymmetric and heavy-tailed regimes where other existing methods fail.
PaperID: 6028, Poster
Abstract: Causal reasoning from observational data becomes significantly challenging when the underlying causal structure is only partially known. In this work, we study the the equivalence class of structural causal models (SCMs) with latent variables sharing the same set of conditional independencies, as represented by a partial ancestral graph (PAG). Specifically, we provide a new characterization of SCM-induced causal diagrams in this equivalence class using differentiable parameters. Building on this characterization, we introduce a new model class called PAG-constrained neural causal models (\mathcalP-NCMs) which parameterize this equivalence class of SCMs. We prove that \mathcalP-NCMs are expressive enough to represent counterfactual distributions induced by any SCM in the class, while remaining consistent with the structural constraints shared across all its members. Finally, we demonstrate how the differentiable parameterization of this model class enables causal inference under Markov equivalence by reducing counterfactual partial identification to an optimization problem over \mathcalP-NCMs. We establish the theoretical soundness of this approach and validate its performance on simulations.
PaperID: 6029, Poster
Abstract: Manipulation planners depend on accurate predicates about physical state---is an object occluded? is a container open? is a path blocked?---but such predicates are hard to read off an image. Contrastively-trained vision--language models capture semantic content but collapse to near-zero binary F1 on these predicates, and even GPT-4o few-shot falls below random on multiple-choice physical-state questions. We present \ours, which predicts physical-state scene graphs from a single image and a natural-language goal. A frozen CLIP ViT, adapted with LoRA, feeds a heterogeneous graph of object, relation, and learnable memory tokens into a Latent Graph Reasoning Transformer (\lgrt); training combines supervised, counterfactual-margin, contrastive, and calibration losses with offline counterfactual-pair augmentation. On \pacbench, \ours reaches \mu-F1\,=\,0.651 and 99.3% MCQ accuracy, a +0.170 absolute \mu-F1 gain over the strongest no-graph supervised baseline, with 99.6% counterfactual directional consistency. Component ablations attribute +0.426 \mu-F1 to the \lgrt and +0.392 to the counterfactual loss. A \manipbench-trained variant scores 74.4% MCQ against 30.0% for the best zero-shot VLM on that benchmark's 4-choice format.
PaperID: 6030, Poster
Abstract: Driving vision-language-action (VLA) models must connect semantic reasoning with executable trajectory planning, but language and motion exhibit different generation structures: reasoning text is naturally sequential, whereas future trajectories require horizon-level geometric consistency. We present RAPDrive, a shared-latent hybrid-decoding framework that processes visual context, ego state, motion history, reasoning tokens, and future planning slots within a single transformer sequence. RAPDrive generates reasoning text autoregressively while refining future motion through iterative masked denoising in a dedicated motion-token space, followed by continuous trajectory realization. This design couples reasoning and planning through shared latent computation while maintaining separate tokenizations for language and motion. Training combines supervised reasoning-and-planning objectives, and GRPO is evaluated as an optional post-training extension. Experiments on NAVSIM and Bench2Drive show that RAPDrive outperforms prior non-world-model driving VLA baselines in the supervised/base setting, benefits further from planner-centric post-training, and is supported by ablations on the main architectural choices.
Authors:
Peiliang Cai, Evelyn Zhang, Jiacheng Liu, Hao Lin, Ruiqi Zhang, Weile Mo, Yue Ma, Shikang Zheng, Jiehang Huang, Dongrui Liu, Linfeng ZhangAbstract: Recent advances in autoregressive video diffusion have enabled sequential and streaming video generation. However, long-horizon generation requires increasingly large KV caches, making efficient compression without sacrificing quality challenging. Existing methods mostly select historical frames based on attention scores, but their context decisions remain coarse. When multiple frames are generated in the same chunk, these methods often apply a shared history selection to the whole chunk, score historical frames solely by attention, and assign head-wise budgets either uniformly or by attention-pattern heuristics rather than explicit head-importance estimation. We show that frames within the same generated chunk can depend on distinct historical frames, that the same historical frame can receive different attention scores as its relative temporal distance to the current frames changes, and that masking different heads induces unequal generation degradation. Motivated by these findings, we propose Focused Forcing, a training-free KV selection method that focuses cached history along both generated-frame and head dimensions. For each generated frame, Focused Forcing preserves the most relevant and distinctive historical frames by combining attention scores with diversity scores of historical frames, while assigning larger budgets to heads with higher estimated importance. Across multiple autoregressive generation paradigms, Focused Forcing achieves up to 1.48× end-to-end acceleration without training, while improving visual quality and text alignment. Our code will be released on GitHub.
Abstract: Flow matching has recently emerged as a promising alternative to diffusion-based generative models, particularly for text-to-image generation. Although flow matching places no restriction on the source distribution, most existing systems still inherit a standard Gaussian from diffusion models, and the source is rarely treated as an optimization target at this scale. Recent works have begun to revisit this choice through condition-dependent or learned sources, yet evidence that such designs are effective in modern text-to-image systems—with high-dimensional latents and tightly integrated conditioning—remains limited. In this work, we study condition-dependent source distributions for flow matching along three axes: _why_ source learning helps, through the lens of the intrinsic variance term in the flow-matching objective; _how_ to make it work in modern text-to-image systems, where variance-only regularization and directional source—target alignment are critical for stable end-to-end training; and _when_ it is most beneficial, by connecting source design to recent representational advances in generative modeling and identifying target representation regimes in which learning the source yields the largest gains. Extensive experiments across multiple text-to-image benchmarks, backbones, and scales demonstrate that principled source design yields consistent and robust improvements—including up to \mathbf3.01× faster convergence in FID and \mathbf2.48× in CLIP score—and outperforms representative prior conditioned-source and condition-aware coupling methods.
Abstract: Test-time model evolution offers a promising way for deployed models to improve from unlabeled test-time experience, yet most existing methods depend on backpropagation (BP), which incurs substantial memory overhead and makes them difficult to deploy on edge devices, quantized models, specialized accelerators, or black-box models. In this work, we study test-time model evolution under a strict two-forward budget, a setting that pushes adaptation toward highly efficient real-world deployment. We reveal three key obstacles in zeroth-order test-time optimization: susceptibility to shortcut solutions, uncontrolled weight drift, and ineffective update direction estimation. To overcome them, we propose EVA-0, a minimal zeroth-order adaptation framework that: 1) keeps the loss scale-invariant to prevent shortcut solutions; 2) devises an anchor-guided optimization strategy to alleviate weight drift; 3) uses sample-wise symmetric two-sided perturbation for update direction estimation and inference. EVA-0 requires no BP and performs both inference and adaptation within only two forward passes per sample. Results on ImageNet-C\&ViT-Base show that EVA-0 outperforms both BP-based DeYO and BP-free FOA, while achieving a 14× speed-up over FOA.
Authors: Dominik Dold, Philipp Petersen
Abstract: used to study expressivity and trainability in artificial neural networks (ANNs). Causal pieces partition the input and parameter space of a feedforward SNN with single-spike coding into distinct regions where the same subnetwork causes the output spikes. For networks of current-based leaky integrate-and-fire (LIF) neurons with large membrane time constants, we show that within each causal piece, output spike times are locally Lipschitz continuous with respect to inputs and network parameters. We further prove a lower bound on the approximation error that depends on the number of causal pieces. Thus, the number of causal pieces is a measure of the approximation capabilities of SNNs, which is valid despite spike-time discontinuities and applies to networks with both excitatory and inhibitory synapses. Empirically, we find that parameter initialisations yielding more causal pieces on the training set strongly correlate with SNN training success across multiple benchmarks, including Yin-Yang, Fashion-MNIST, and EuroSAT. Moreover, simulations with standard single-spike LIF neurons indicate that our findings extend beyond the theoretically analysed regime. These results establish causal pieces as a powerful and principled tool for analysing and improving the computational capabilities of SNNs.
Abstract: Diffusion large language models are a compelling alternative to autoregressive models, yet existing RL methods for diffusion treat all denoising steps as equally important and rely on biased, high-variance likelihood estimates. We identify two fundamental weaknesses: the absence of temporal credit assignment across the denoising trajectory, and the systematic bias of mean-field likelihood estimates used for policy optimization. To address these, we propose Denoising-Aware Credit Assignment for GRPO (DACA-GRPO), a lightweight, plug-and-play enhancement for any GRPO-style trainer. DACA-GRPO introduces two complementary mechanisms: Denoising Progress Scores, which extract per-token importance weights from intermediate predictions at no additional forward cost, and Stratified Masking Likelihood, which partitions token positions into strata so that each token is predicted with most of the sequence as context, reducing the mean-field bias. Applied on top of three GRPO base methods, DACA-GRPO achieves consistent improvements across seven benchmarks spanning mathematical reasoning, code generation, constraint satisfaction, and constrained generation, with gains of up to 5.6pp on math reasoning, 7.4pp on code generation, 36.3pp on constraint satisfaction, and 5.9pp on JSON schema adherence.
PaperID: 6036, Poster
Abstract: Sycophancy in large language models is often measured as agreement with the user, yet agreement can be appropriate when the user's claim is valid, their preference is genuine, or their experience warrants acknowledgment. We study sycophancy as inappropriate agreement: validating a claim, preference, or self-assessment when the context calls for correction, qualification, or resistance. We introduce a validity-aware sycophancy grader and SycoDrift, a persona-targeted evaluation set of 3,298 validity-annotated probes with controlled annotations for user openness and stakes. We use this framework to analyze three sources of agreement pressure: human preference data, automated judge specification, and in-context personalization. First, preference for validating responses in human datasets (Chatbot Arena, HH-RLHF, Community Alignment) is weak and concentrated where agreement is often appropriate. Second, automated LLM judges can introduce sycophancy by conflating prosocial values with agreement during reward specification; GRPO training amplifies this rubric-level error into model behavior. Third, inference-time personalization increases inappropriate agreement across 14 models, with the largest effect from matched user profiles placed directly before the query.
PaperID: 6037, Poster
Abstract: Composed Image Retrieval (CIR) enables users to express specific search intent by combining a reference image with manipulation text. Despite its promise for e-commerce, existing methods ignore readily available item textual descriptions, suffer from multi-attribute confounding where samples exhibit uncontrolled variations in unrelated attributes, and lack explicit fine-grained attribute modeling. We propose MoA^3CIR, featuring an asymmetric five-tower architecture that incorporates item descriptions on both query and target sides, a self-reflective data augmentation pipeline generating controlled attribute modifications through LVLM-driven refinement, and a Mixture of Attribute-Aware Attention mechanism with grouped multi-head attention and dynamic routing coupled with a two-phase training strategy. Besides, we contribute ECCIR, the first dual-modal e-commerce CIR benchmark with 110K samples featuring authentic item descriptions. Extensive experiments demonstrate that \method~achieves SoTA performance on ECCIR, FashionIQ, and Shoes datasets, with substantial improvements over existing methods. Code and datasets will be made publicly available.
Abstract: Online inventory optimization (OIO) is online convex optimization with physical memory: inventory can be raised immediately, but it can be reduced only by demand. A natural principle, used in stochastic inventory learning and recently in OIO with linear capacity constraints, is to maintain a hidden target chosen by an online learner and implement its projection onto the currently feasible order-up-to set. We prove that this simple principle is optimal for OIO on arbitrary bounded convex capacity sets. With online gradient descent as the hidden learner, the method improves the best known regret guarantee for OIO on general convex sets from inverse to inverse-square-root dependence on the common-demand probability, and we prove a matching lower bound. The analysis identifies the right geometric state variable: the Euclidean distance between the hidden target and the implementable set. This distance evolves pathwise as a scalar queue, with target movement as arrivals and common demand as service, reducing the state-dependent feasibility cost to the control of a one-dimensional queue. The same reduction gives the first logarithmic regret guarantee for strongly convex losses and the first dynamic regret guarantee adapting to Euclidean path variation on general convex capacity sets.
PaperID: 6039, Poster
Abstract: Infrared-Visible image fusion (IVIF) is essential for robust perception, yet existing methods remain limited under complex degradations. While diffusion models (DMs) offer powerful modeling capabilities to combat such degradations, their iterative process incurs high computational cost. One-step distillation effectively accelerates inference but typically leads to over-smoothed results. To resolve these limitations, we propose a one-step fusion framework with hybrid reinforcement learning, termed OneHRL. We treat degraded inputs as intermediate noisy states along a generative trajectory, enabling a distilled one-step generator to rectify them directly to the clean fusion results. Concurrently, to mitigate the over-smoothed results in distillation, we incorporate the Flow-GRPO paradigm to introduce stochasticity into the deterministic one-step generation. This restores the stochastic exploration necessary for sampling-based reinforcement learning. Furthermore, we design a coarse-to-fine reward that integrates a global semantic reward with an object-level reward. This hybrid reward ensures that the generator captures both naturalness and fine-grained texture. Extensive experiments demonstrate that our method achieves superior fusion quality and efficiency.
Abstract: Accurate image registration is essential in many medical imaging applications, yet most deep registration networks provide little indication of when or where their predictions are unreliable. Existing uncertainty estimation approaches, such as Bayesian methods, ensembles, or MC-dropout, typically require architectural modifications or retraining, precluding their applicability to pretrained registration models. We propose an inference-time, model-agnostic uncertainty estimation framework that applies directly to any pretrained registration network. Our approach is grounded in the transformation equivariance property of image registration, which states that the underlying anatomical mapping should remain consistent under spatial perturbations of the input. Experiments across three pretrained registration models and four anatomical structures show that the resulting uncertainty maps consistently correlate with registration error and highlight unreliably aligned regions. This framework turns pretrained registration networks into risk-aware tools at test time, moving medical image registration closer to safe clinical and large-scale research deployment. The code will be released at .
Authors:
Kiran Tomlinson, Sonia Jaffe, Will Wang, Scott Counts, Siddharth SuriAbstract: With generative AI emerging as a general-purpose technology, understanding its economic effects is among society’s most pressing questions. Existing studies of AI impact have largely relied on predictions of AI capabilities or focused narrowly on individual firms. Drawing instead on real-world AI usage, we analyze a dataset of 200k anonymized conversations with Microsoft Bing Copilot to measure AI applicability to occupations. We use an LLM-based pipeline to classify the ONET work activities assisted or performed by AI in each conversation. We find that the most common and successful AI-assisted work activities involve information work---the creation, processing, and communication of information. At the occupation level, we find widespread AI applicability cutting across sectors, as most occupations have information work components. Our methodology also allows us to predict which occupations are more likely to delegate tasks to AI and which are more likely to use AI to assist existing workflows. Finally, we see an increase in occupation-level AI applicability over our 9-month data period, driven primarily by people using AI for more varied tasks.
PaperID: 6042, Poster
Abstract: Large language models (LLMs) are increasingly used for assistive decision-making in high-stakes domains, yet their outputs can be unreliable and difficult to audit. This necessitates pre-deployment testing, which can be used to identify scenarios where an LLM-assisted pipeline fails to produce satisfactory decisions. Existing testing strategies are often not sample-efficient and struggle to adapt to diverse testing objectives, application domains, and rapidly evolving models. We introduce a sample-efficient scenario design strategy for testing LLM-based decision-making pipelines. Our approach formulates testing as adaptive experimental design, using a flow-based surrogate model to sequentially acquire scenarios that balance optimization to discover challenging test cases with diverse exploration. We evaluate the method across several LLM-assisted decision-making tasks: bank loan approval, stock prediction from news articles, and disaster resource allocation from tweets and images. Our approach consistently discovers challenging scenarios that satisfy the testing objective across all tasks, providing high coverage over scenario space.
PaperID: 6043, Poster
Abstract: Large language models generate one token at a time, yet their responses show remarkably consistent length structure: step-by-step solutions converge in predictable token counts, retrievals stop after a few sentences, retractions extend responses by measurable amounts. We ask whether the model carries an internal estimate of how much response remains. Training minimal-capacity linear probes on frozen hidden states of three open-weight 7-8B models across seven completion-style datasets, we find three converging pieces of evidence. First, total response length is linearly decodable from the prompt's last hidden state alone, before any output is emitted. Second, probe directions trained on natural-language datasets transfer broadly, including to controlled synthetic completions never seen in training, outperforming a statistical baseline; the converse direction generally fails, and this asymmetry is itself informative. Third, on curated high-loss completions, the probe's per-position estimate shifts upward at the moment the model retracts and restarts a partial solution, a directional behavior no position-only predictor can reproduce (we note that this is qualitative, not aggregate). We frame this as approximate estimation of remaining generation length, distinct from exact-counting impossibility results for transformers, and interpret it as evidence that LLMs maintain a plan-like internal representation of output length (decodable, not necessarily used causally). Code: https://anonymous.4open.science/r/llm-output-length
PaperID: 6044, Poster
Abstract: Recent advances in brain decoding have achieved impressive visual reconstruction under subject-specific settings, yet they fail to generalize to unseen subjects without retraining---hindering practical zero-shot cross-subject applications. The core challenges stem from inter-subject variability in voxel dimensionality and misaligned spatial response patterns. While existing methods rely on surface-based alignment or complex feature disentanglement, they often operate as black boxes and overlook the potential of volume-based fMRI data. To address this, we propose Voxel as Token (VoxTok), a conceptual framework that treats each fMRI voxel as an independent token endowed with a learnable functional embedding. This formulation naturally handles variable-length inputs and subsumes existing adapter-based models as special cases where linear projections approximate voxel functionality. Guided by the hypothesis that large-scale spatial distributions of neural activity are consistent across subjects despite fine-grained variability, we introduce a KNN-based estimator to predict functional embeddings for unseen subjects. This approach achieves state-of-the-art zero-shot retrieval performance and comparable reconstruction quality, while enabling plug-and-play adaptation for existing models (e.g., MindEye2). Integrating multi-scale ROIs further boosts performance, with our best model achieving 40.4% zero-shot image retrieval accuracy, rivaling early supervised methods. Crucially, we find that decoding success correlates strongly with ROI size, indicating that global co-activation patterns drive cross-subject generalization.
Abstract: Developing efficient CUDA kernels is a fundamental yet challenging task in the generative AI industry. Recent research leverages Large Language Models (LLMs) to automatically convert PyTorch reference implementations to CUDA kernels, significantly reducing engineering effort. State-of-the-art LLMs, such as GPT-5.2 and Claude-Sonnet-4.5, still struggle with this task. To address this challenge, we propose , a scalable learning framework for training LLMs to convert PyTorch programs into highly optimized Triton kernels, which are then compiled to CUDA kernels at runtime. DRTriton consists of three key components: (i) a data synthetic algorithm CSP-DAG that guarantees full coverage and unbiased uniform sampling over the operator space with controlled difficulty; (ii) a curriculum RL framework with decoupled rewards that jointly optimizes conversion success rate and execution speed; and (iii) a test-time search algorithm that further improves the execution speed of the generated Triton kernels. With a warmup stage of SFT on limited PyTorch-Triton pairs curated using existing LLMs, DRTriton trained by RL on synthesized PyTorch programs generalizes effectively to real-world CUDA kernels that are challenging even for human experts. Experimental results show that DRTriton-7B achieves speedup over PyTorch on 92% of KernelBench Level 2 tasks, compared to 23% for GPT-5.2 and 19% for Claude-Sonnet-4.5.
PaperID: 6046, Poster
Abstract: Optimization under heavy-tailed stochastic noise remains a major challenge in modern machine learning, where classical moment assumptions often fail to capture realistic gradient distributions. In this work, we revisit robust estimation techniques and provide a systematic analysis of the Median-of-Means (MoM) estimator under a structured noise model that interpolates between a symmetric heavy-tailed distribution with bounded \beta-th moment (\beta \in (0, 1]) and a mean-zero distribution with bounded \alpha-th moment (\alpha \in (1, 2]). We derive new upper bounds on the variance and bias of MoM in this setting, thereby extending the applicability of existing results beyond purely symmetric or bounded-moment settings. Building on these properties, we study clipped stochastic optimization methods with MoM gradient estimation and establish high-probability convergence guarantees that hold for general \beta \in (0, 1] and \alpha \in (1, 2]. Our bounds match or improve upon prior complexity guarantees while requiring weaker distributional assumptions, and they reveal that bias only affects the final neighborhood of convergence. Overall, this work advances the theoretical foundations for robust stochastic optimization in the presence of extreme gradient noise.
PaperID: 6047, Poster
Abstract: As deep neural networks are required to continuously adapt to evolving data, class-incremental learning (CIL) has become an active research area. However, most existing works primarily focus on improving accuracy while overlooking confidence calibration, which measures the trustworthiness of the model's predictions. Although temperature scaling (TS) is effective for static calibration, it has rarely been explored in continual settings where data distributions evolve over time. In this paper, we introduce Multi-Group Temperature Scaling for Continual Calibration (MT-CC), a novel post-hoc calibration framework that addresses the asymmetric calibration behavior in CIL, where the rate of confidence changes does not align with the rate of accuracy degradation across tasks, leading to different levels of miscalibration. Our key contribution is to design multiple temperature parameters through a lightweight grouping network that separates samples by confidence statistics and task-aware signals in CIL. Furthermore, a Groupwise Difference-between-Confidence-and-Accuracy (GDCA) regularization is incorporated to promote intra-group consistency. Experiments demonstrate that MT-CC significantly improves calibration performance while maintaining accuracy, offering a principled framework for reliable continual learning.
Abstract: Attention is the key mechanism underlying in-context learning in transformers, and attention patterns have been observed empirically to emerge abruptly during training. We present a Bayesian theory of feature learning in attention; we then focus on how the copy subcircuit in the first layer of an induction head is learned by analyzing a single-layer softmax attention network trained on a copy task. We derive a closed-form posterior over the attention matrix and reduce it to a low-dimensional order parameter space. This reduction reveals a phase transition in the amount of training data, which we verify using both Bayesian sampling and standard training with Adam. We contrast our results with linear attention and find that softmax attention exhibits a ). Our work provides a first-principles theoretical account of the abrupt emergence of the copy subcircuit, reminiscent of the one observed in training large language models.
PaperID: 6049, Poster
Abstract: Interactive person retrieval extends text-to-image person retrieval by allowing follow-up questions when the initial description is incomplete. The core challenge is to ask questions that uncover the most useful missing information for identifying the target person from visually similar candidates. Existing methods guide question selection with retriever-external signals, which may not reflect the information most useful to the current retriever. To address this, we propose SCOUT, an interactive person retrieval framework that derives questions from the retriever's own response. Specifically, we introduce a Perturbation-Sensitive Question Selection strategy that selects the next question based on Retrieval Perturbation Sensitivity (RPS), a test-time measure of perturbation-induced ranking change. Additionally, SCOUT requires no additional training and builds on off-the-shelf text-to-image person retrievers, avoiding the dialogue-data curation and retraining overhead of existing interactive methods. Extensive experiments on three standard benchmarks demonstrate that SCOUT yields consistent improvements over baseline methods, validating this RPS-based strategy for interactive person retrieval.
PaperID: 6050, Poster
Abstract: Remote sensing object detection often proceeds in a domain-incremental manner, where detectors continuously encounter new domains arising from changes in regions, resolutions, sensors, and modalities. This challenge is particularly severe when passive optical red-green-blue (RGB) imagery and active synthetic aperture radar (SAR) imagery coexist, because their distinct imaging mechanisms create large domain gaps. Existing methods mainly rely on feature alignment or regularization, but often overlook semantic associations across heterogeneous domains. The problem is more pronounced for DETR-like detectors, where decoder queries act as object-centric semantic slots and are easily disrupted by cross-domain adaptation. Under a query-as-resource view, we propose Q-cost for remote sensing domain-incremental object detection. At the feature level, Q-cost uses modality- and domain-aware prototypes to guide a two-level mixture-of-experts adapter and generate encoder prompts for domain-aware semantic aggregation. At the semantic level, it models decoder queries as limited resources with domain-dependent activity costs, and applies activity-aware gating to regulate query updates according to current and historical domain demands. Experiments on multiple remote sensing benchmarks show that Q-cost effectively balances new-domain adaptation and old-domain retention under severe domain and modality shifts. Our code is provided in the supplementary material.
Abstract: We propose Log-Averaged Mirror Prox (LAMP), a linear-space primal-dual method for large-scale optimal transport. LAMP implements primal mirror prox updates by tracking an averaged dual sequence, reducing storage complexity from \mathcal O(nm) to \mathcal O(n+m) while preserving dense, GPU-friendly reductions. Consequently, LAMP preserves the last-iterate \widetilde\mathcalO( nm\varepsilon^-1) arithmetic complexity of conservatively parameterized primal-dual mirror prox. We further analyze LAMP as a direct optimal transport solver in a more performant parameter regime, providing a last-iterate sub-optimality certificate dependent on infeasibility and an explicit \mathcal O(1/t) term. Moreover, we give a computable sufficient condition for best-iterate convergence to a saddle-point. Numerical experiments with an optimized CUDA implementation show that LAMP outperforms first-order baselines in several high-accuracy (entropic) optimal transport problems. LAMP is further shown to scale up to problems with n=m=2^18 marginal supports, which were previously beyond the reach of primal-dual first-order methods.
PaperID: 6052, Poster
Authors: Ruben Moreno Bote, Francesco Damiani, Dmytro Grytskyy
Abstract: Multiplicative motor and observation noise, along with internal noise corrupting computation, are central features of the sensorimotor system in biological and robotic agents. Yet, analytical solutions for stochastic optimal control are largely restricted to additive noise models and neglect internal noise. Here, we consider the problem of finding optimal control and filter laws for partially observable stochastic linear systems under quadratic costs with multiplicative control and observation noise, as well as internal noise. We provide an efficient, analytically derived coordinate-descent algorithm that computes mutually optimal linear control and filter laws by reformulating the problem as a constrained optimization. Our method provably guarantees monotonic decrease of the expected cost and convergence to a critical point, and overcomes the suboptimality and incompleteness of prior analytical approaches. Compared to state-of-the-art numerical methods, it achieves orders-of-magnitude computational speedups. Our solution also proves instrumental in revealing novel internal-noise-dependent, task-structured control strategies in a redundant arm control task.
Authors: Tuc Nguyen, Thai Le
Abstract: Activation steering provides a lightweight inference-time mechanism for controlling large language models (LLMs) by modifying their internal activation vectors toward desired behaviors. Most existing methods compute a fixed steering direction in the original activation space, typically from pairs of contrastive examples using mean differences, linear probes, or arbitrary separability criteria. While effective to a certain extent, these methods treat behavioral control as a global, linear, additive offset: the same direction is applied across inputs, and behaviors are linearly separable. This can be restrictive when behavioral features vary nonlinearly across the activation space or lie on curved and anisotropic manifolds, where the optimal intervention may be input-dependent. To address this limitation, we propose INNSteer, a nonlinear activation steering framework based on invertible latent transformations. Rather than searching for a better steering vector in the original representation space, INNSteer learns a lightweight invertible neural network \phi that maps an LLM's activations into a latent space where behavioral classes are more amenable to linear control. At inference time, activations are mapped through \phi, steered in the latent space, and mapped back through the exact inverse transformation \phi^-1. This makes a simple latent-space translation become a nonlinear, input-dependent intervention in the original activation space. Across experiment settings on multiple LLM families, scales, behavioral traits, and safety benchmarks, INNSteer consistently improves model control over linear, transport-based, and nonlinear steering baselines while largely preserving generation fluency.
Abstract: With the rapid advancement of generative models, powerful image editing methods now enable diverse and highly realistic image manipulations that far surpass traditional deepfake techniques, posing new challenges for manipulation detection. Existing image manipulation detection and localization (IMDL) benchmarks suffer from limited content diversity, narrow generative-model coverage, and insufficient interpretability, which hinders the generalization and explanation capabilities of current manipulation detection methods. To address these limitations, we introduce , a large-scale benchmark for image manipulation detection and localization focusing on AI-edited images. ManipBench contains over 450K manipulated images produced by 25 state-of-the-art image editing models across 12 manipulation categories, among which 100K images are further annotated with bounding boxes, judgment cues, and textual explanations to support interpretable detection. Building upon ManipBench, we propose , an all-in-one model based on a Multimodal Large Language Model (MLLM) that leverages contrastive LoRA fine-tuning and task-specific decoders to achieve unified image manipulation detection, localization, and explanation. Extensive experiments on ManipBench and several public datasets demonstrate that ManipShield achieves state-of-the-art performance and exhibits strong generality to unseen manipulation models. Both ManipBench and ManipShield will be released upon publication.
Abstract: Large Language Model (LLM) interactions are typically underspecified, with users clarifying all necessary details across multiple conversational turns. Yet recent work shows that LLMs perform far worse in this multi-turn setting than in a single turn with all information being available at once, a phenomenon termed ``Lost in Conversation.'' However, bridging this gap effectively and generally remains an open problem. Here we introduce Found in Conversation (FiC), a training framework where a model teaches itself to find back its single-turn competence given underspecified multi-turn prompts. We develop View-Asymmetric Self-Distillation, which runs the same self-distillation backbone on two views of the same information: the teacher sees a single-turn view that concatenates all information revealed across the conversation turns, and the student sees the multi-turn view itself; distillation aligns the student's weaker multi-turn behavior with the teacher's stronger single-turn behavior. Across diverse model sizes (3B–14B) and architectures (Llama, Qwen, Phi, and OLMo), ours achieves a 100% recovery of the single-turn performance on 2 Llama models and recovers over 90% on every model, delivering more helpful and efficient multi-turn conversations without compromising single-turn performance.
PaperID: 6056, Poster
Abstract: Software requirements are usually written in natural language, which is essential for human communication but insufficient as a target for machine-checked correctness. While natural-language descriptions can state intended functionality, verification requires formal specifications with explicit types, relations, quantifiers, guards, witnesses, and edge cases. To bridge this divide, we study the Formal Specification Synthesis task: given a natural-language programming requirement and fixed formal signatures, synthesize a proof-assistant specification for the intended input-output relation. This task is difficult for two primary reasons. First, a large representation gap separates informal natural language from the rigorous structures needed by a proof assistant. Second, generated formal specifications are notoriously hard to evaluate, as proving full equivalence between specifications is often too complex to serve as routine feedback. To address these intertwined challenges, we propose SpecBridge, a reconstruction-guided formalization framework. At its core, SpecBridge introduces a Natural-Language Formalization Plan (NLFP), a semi-structured, readable intermediate representation that captures key formalization choices before decoding to a target specification language. Through reconstruction-guided pattern mining, SpecBridge learns these NLFPs by reconstructing formal specifications to discover reusable patterns, ultimately transferring them to improve natural-language task generation. Furthermore, to tackle the evaluation bottleneck, we introduce a multi-layer evaluation protocol for formal-specification correctness. This protocol utilizes LLM Specification Alignment as an LLM-based judge, alongside Test-Driven Computable Validation and Test-Driven Proof Validation, which approximate "running" testcases on formal specifications by translating propositions into executable pre/post checks and proving testcase obligations in Lean. In our experiments, SpecBridge outperforms the few-shot + CoT baseline by 6.1% on LLM Specification Alignment, 15.2% on Computable Validation Pass Rate, and 28.4% on Proof Validation Pass Rate. These results demonstrate that the NLFP bridge and our multi-layer evaluation protocol make formal specification synthesis significantly more reliable.
PaperID: 6057, Poster
Abstract: Calibrating large language models (LLMs) in open-ended generation is uniquely challenging because many distinct token sequences can express the same underlying meaning. Prior calibration techniques target either fixed-set classification confidence or the binary correctness of individual strings, making them inherently ill-suited for the dynamic, multi-string nature of open-ended semantic outcomes. To bridge this gap, we introduce S-EDL, a novel evidential method that natively optimizes LLMs for semantic calibration. By eliciting evidence from the sequence likelihoods of the model itself, S-EDL constructs a differentiable Dirichlet prior over prompt-specific semantic classes. Optimizing the evidential loss under this prior yields a calibration-aware training signal over variable semantic classes, completely bypassing the need for fixed multiple-choice options. Evaluations across four open-ended benchmarks and three models demonstrate that S-EDL substantially reduces semantic calibration error while improving accuracy. By aligning generative likelihoods with semantic correctness, this work establishes a framework for deploying trustworthy LLMs in complex, open-ended applications.
Authors:
Chulabhaya Wijesundara, Andrea Baisero, Zhongheng Li, Gregory D Castanon, Alan S Carlin, Christopher AmatoAbstract: Centralized training with decentralized execution (CTDE) is a standard framework for cooperative multi-agent policy-gradient reinforcement learning, allowing agents to learn from joint information while acting from local observations. Ratio-based trust-region methods such as Multi-Agent Proximal Policy Optimization (MAPPO) and Multi-Agent Simple Policy Optimization (MASPO) update decentralized actors using per-agent probability ratios weighted by joint advantage estimates. Teammate non-stationarity increases the variance of these advantages, which in turn increases the variance in the local ratio updates. This exposes two method-specific failure modes: MAPPO's additive clipping removes gradients for outlier samples and weakens recovery from policy drift, while MASPO's soft quadratic penalty can allow probability collapse. We introduce Multi-Agent Ratio Symmetry (MARS), a novel policy optimization objective that replaces these additive ratio-based trust-region mechanisms with a multiplicatively symmetric geometric barrier. MARS preserves corrective gradients while assigning unbounded cost as probability ratios approach zero. Across 47 tasks spanning eight multi-agent environments, including novel JAX benchmarks PaxMen and AeroJAX, MARS matches or exceeds MAPPO and MASPO in aggregate environment-level performance. Ablations show that these gains arise from the geometry of the symmetric barrier rather than from flexible trust-region boundaries alone.
Abstract: For agents to learn continuously from interaction with the world at test time, they must be able to explore effectively, acquire new world knowledge and skills, retain relevant episodic experiences, and plan over long horizons. To evaluate these key abilities of test-time continual learning agents, we introduce AgentOdyssey, a novel evaluation framework that procedurally generates open-ended text games with rich entities, world dynamics, and long-horizon tasks. Critically, AgentOdyssey goes beyond the conventional machine learning assumption that learning does not occur at test time by placing agents in a continuous, long-horizon setting that interleaves learning and inference throughout deployment. We further propose a multifaceted evaluation methodology that measures not only game progress but also offers diagnostic tests on world knowledge acquisition, episodic memory, object and action exploration, action diversity, and model cost. We evaluate diverse agent paradigms in the generated games. Our experimental results reveal critical limits in agents' key abilities, as well as factors that influence their meaningful horizon. Although performance scales with stronger base models, even the top agent remains far below human performance, leaving substantial headroom for improvement. Among agent mechanisms, we find that short-term memory benefits multiple agent paradigms and is an important component of agent test-time training.
Abstract: Contextual recommendation is a variant of contextual linear bandits in which the learner observes an (optimal) action rather than a reward scalar. Recently, Sakaue et al. (2025) developed an efficient Online Newton Step (ONS) approach with an O(d\log T) regret bound, where d is the dimension of the action space and T is the time horizon. In this paper, we present a simple algorithm that is more efficient than the ONS-based method while achieving the same regret guarantee. Our core idea is to exploit the improperness inherent in contextual recommendation, leading to an update rule akin to the second-order perceptron from online classification. This removes the Mahalanobis projection step required by ONS, which is often a major computational bottleneck. More importantly, the same algorithm remains robust to possibly suboptimal action feedback, whereas the prior ONS-based method required running multiple ONS learners with different learning rates for this extension. We describe how our method works in general Hilbert spaces (e.g., via kernelization), where eliminating Mahalanobis projections becomes even more beneficial.
PaperID: 6061, Poster
Abstract: Multimodal (e.g. vision–language, in this work) attribution aims to interpret the models with- out access to ground-truth explanatory supervi- sion. Existing multimodal attribution methods typically rely on paired modalities as proxy su- pervision, implicitly assuming that image–text pairs are semantically aligned. Information bot- tleneck (IB)–based attribution methods follow this general paradigm and instantiate it through a global sufficiency–compression objective. In real- istic multimodal datasets, however, this assump- tion is frequently violated due to partial semantic misalignment and unreliable cross-modal corre- spondences. Through theoretical analysis, we show that, under such misalignment, information- bottleneck-based attribution objectives exhibit a structural limitation: by assuming equal reliabil- ity across all proxy supervision signals, they in- duce posterior over-contraction and forced expla- nations. To address this issue, we propose TIB, a sample-wise Tempered Information Bottleneck framework that adaptively modulates attribution strength according to cross-modal reliability, with- out assuming semantic alignment. Extensive ex- periments on controlled benchmarks and large- scale datasets demonstrate the effectiveness of the proposed approach, while also validating the identified failure mode
Abstract: Sparse autoencoders (SAEs) have become a central tool for interpreting language models. However, two key SAE analyses that remain difficult to scale are (1) matching semantically similar features across multi-layers and (2) compressing large feature circuits into interpretable supernodes. Although these have been treated as separate problems, we show that both are instances of a more fundamental challenge, which we frame as the estimation of semantic distances between SAE features that lie on different activation manifolds. We introduce a distributional framework for this problem, in which each feature is represented not by a single decoder vector like in the literature, but by an activation-weighted distribution over the hidden states that express it. By projecting these distributions into a shared reference space and comparing them with Wasserstein distance, our method provides a unified semantic metric for cross-layer feature comparison. We prove that our representation is invariant to activation rescaling, stable under perturbations, and recovers true matches under finite-sample margin conditions. Empirically, our method outperforms decoder-vector and LLM-based baselines and captures subtle functional distinctions between related features. Notably, our method compresses large feature circuits into interpretable supernodes
PaperID: 6063, Poster
Authors: Artyom Kabanov, Mikhail Mozikov, Sergei Volchkov, Igor Gorobets, Daniil Shirshin, Alena Kriuchenkova, Bykov Nikita, Valeria Bodishtianu, Nikita Severin, Sergey Muravyov, Andrey Savchenko, Ilya Makarov
Abstract: LLM agents now negotiate, allocate resources, and make commitments on a user's behalf. The vulnerability is not only in the final move but in what comes before it: in pre-action conversation, a model can be primed to make promises, lock in commitments, or reveal intentions that a counterpart will exploit when the action is finally taken. Yet standard game-based evaluations usually observe only the final move, leaving the speech-action gap unmeasured. We study this gap directly by introducing HL-PACT, a controlled platform with matched human-human, human-LLM, and LLM-LLM sessions across five repeated games: Prisoner's Dilemma, Battle of the Sexes, Ultimatum, Kuhn Poker, and Liar's Dice. Across 3,752 sessions, we randomize free-form pre-round chat and opponent identity masking, classify each message by its strategic content, and compare what players say, how they say it, and when the two diverge. We find that LLMs consistently activate the communication channel: mixed pairs contain more proposals, promises, and strategic reasoning than human pairs, with humans adapting to the LLM's negotiation register. This improves equilibrium selection in coordination games and shifts bargaining toward fairer offers. In games with profitable unilateral deviation, however, the same communication structure can create asymmetries: when identity is known, human partners selectively adapt to LLMs’ willingness to follow through on stated intentions, producing measurable exploitation. Masking reduces this asymmetry and makes the communication benefit more symmetric. The strategic value of language therefore depends on game structure, counterpart identity, and whether speech is separated from action. Evaluating LLM agents without communication can miss a central deployment risk: models may not fail because they reason poorly, but because their reliability is predictable.
Authors: Zeyu Zhang, Bradly Stadie
Abstract: Backtesting large language models on historical events requires reasoning exclusively from information available before a specified cutoff date. Yet models routinely leak post-cutoff knowledge from pre-training into their reasoning, inflating apparent accuracy and undermining evaluation validity. Prompt-based constraints fail when suppressed content is causally related to the prediction, and knowledge unlearning cannot address this problem because temporal compliance is instance-specific: the same fact may be legitimate evidence for one cutoff date and a violation for another. Rather than erasing knowledge, the model must learn temporal discipline: selecting evidence conditioned on each instance's cutoff date. We propose TEMPO (Temporal Enforcement via Mode-separated Policy Optimization), which trains this discipline via two contributions: (1) a two-mode reward where a leakage mode drives post-cutoff claims to zero as a hard prerequisite before a performance mode optimizes task performance; and (2) a GRPO-based training pipeline that enables the model to discover temporally valid reasoning strategies. We prove that training monotonically decreases leakage, converges to the leak-free optimum, and improves task performance once compliance is achieved. On three prediction tasks and two models, TEMPO reduces leakage from 2~13% to 0.6~3.7% across all conditions, with task performance improving 6~13% where strong pre-cutoff signals exist and maintained where the prediction task is inherently difficult from valid information alone.
PaperID: 6065, Poster
Abstract: Detecting whether specific text samples were used to train large language models is increasingly important for privacy, copyright and data auditing. Knockoff-based training data detection (KTD) offers a promising approach to controlling the false discovery rate (FDR). However, its effectiveness relies on null sign-flip symmetry induced by exchangeability between each null sample and its knockoff. This condition is difficult to satisfy in natural language settings. We study training data detection under approximate exchangeability, and demonstrate that deviations from exact exchangeability induce measurable null-tail asymmetry, which in turn leads to systematic FDR inflation under the standard knockoff+ threshold. This effect can be captured by pathwise bounds, which explain when and why KTD becomes anti-conservative. Based on this, we propose Asymmetry-adjusted KTD (AKTD), a series of procedures that correct for asymmetry by estimating the penalty term. We introduce an external null variant that uses a separate calibration pool containing only null samples to estimate the penalty term.This enables detection with FDR control without using labels from the evaluated candidate set. Experiments on WikiMIA and MIMIR using GPT and Pythia demonstrate that the asymmetry term accurately predicts the observed FDR expansion, and AKTD restores FDR control while maintaining effective detection capabilities.
PaperID: 6066, Poster
Authors: Omer K Ebead, Juan Formanek, Joel Leibo
Abstract: Contextual integrity (CI) holds that appropriate information flow depends on context-specific norms. The same disclosure may be appropriate in one context and inappropriate in another. As actors handle private user data in multi-actor pipelines, we ask whether they track these norms faithfully under social and contextual pressure. We build a generative model of CI in multi-actor systems: scenarios drawn from CI's taxonomies, a generation pipeline, multi-actor simulation, and a two-sided appropriateness metric. The metric scores both withholding when norms forbid sharing and sharing when they permit it. Across three intervention axes (decision logic, cognitive profile, and model choice), no actor-level lever closes both sides of the appropriateness gap; protection and utility trade off in every variant we tested. Closing the gap may require mechanisms beyond the actor level, such as CI-grounded policy and institution design.
PaperID: 6067, Poster
Abstract: Autoregressive world models have emerged as a powerful paradigm for interactive video generation, allowing users to navigate dynamically generated environments through actions. These models are typically conditioned on a text prompt and/or a single reference frame, from which the entire world is generated. Yet the moment the user navigates beyond what is visible in that frame, the unseen regions are populated by the base model's priors, with no mechanism for the user to specify what should appear and where. This is a fundamental limitation for applications such as gaming, interactive storytelling, and simulation, where controllable scene composition is essential. We refer to this missing capability as concept spawning; introducing a user-specified visual concept into a world model, analogous to spawning in a game engine. We introduce SPAWN (Swapping Pinned Anchor with Windowed iNjection), a training-free method for concept spawning. SPAWN exploits a structural property of image-to-video backbones: the first slot of the context memory is pinned to the reference frame and acts as a foundational anchor for every generated chunk. By swapping this anchor with an external concept latent over a short injection window and letting the original anchor return, we cause the concept to propagate naturally through the rollout via the model's own memory. SPAWN supports concepts from fine-grained entities such as characters and props to large-scale elements such as buildings and landmarks, and accepts either a concept image or a text description as input. Experiments show that SPAWN integrates concepts with consistent lighting, scale, and perspective while preserving identity and temporal coherence, demonstrating that controllable concept spawning is achievable in existing autoregressive world models without any training.
PaperID: 6068, Poster
Authors: Ruilin Xie, Bixin Li, Xinyu Chen, Yongqiang Tian, Wang Lulu
Abstract: Large Language Models remain vulnerable to jailbreak attacks despite extensive efforts to ensure safety alignment. Streaming generation scenarios exacerbate this vulnerability by exposing harmful tokens to users immediately upon generation, a phenomenon known as prefix exposure. Existing defenses often fail to address this real-time constraint or impose substantial latency that degrades user experience. To bridge this gap, we propose JEDI, a real-time defense method that secures streaming outputs via in-generation detection and intervention. JEDI leverages representation engineering to monitor the LLM’s internal semantic drift, using a Cumulative Sum algorithm to identify persistent, harmful intent before it manifests in the output. Upon detecting a risk, the system dynamically injects a steering vector to redirect the generation trajectory toward a safe subspace. Extensive evaluations across 6 distinct models and 11 attack vectors demonstrate that JEDI achieves defense success rates ranging from 93.6% to 99.4%, outperforming state-of-the-art baselines. Furthermore, JEDI preserves model utility with a Time To First Token overhead of approximately 0.003 seconds, validating its viability for latency-sensitive applications. Our code and experimental results are available at: https://anonymous.4open.science/r/JEDI-93F9
PaperID: 6069, Poster
Abstract: Kernelized attention methods, which aim to replace the O(N^2) complexity of softmax attention with O(N) alternatives, have been the focus of intense research in the last five years. Despite growing interest, systematic evaluations at foundation model scales and fair comparisons across approaches are still missing. We analyze kernelized attention in two halves and highlight prominent limitations on each. On efficiency, FlashAttention achieves lower memory consumption and faster wall-clock time than state-of-the-art kernelized methods across most sequence lengths (with auto-regressive generation memory the only clear exception), directly contradicting the field's core assumption. We also discover that reported efficiency gains often fail to reproduce against modern baselines. On quality, current kernelized methods leave gaps of 2-10 points relative to softmax under a unified distillation protocol, and these gaps do not close with model scale. We show this gap is not intrinsic to kernelization: LARA++, a principled extension of LARA that approximates softmax via importance sampling with adaptive proposal distributions, recovers up to 99% of softmax accuracy. We establish that the gap is closable; however, the computational overhead remains substantial, leaving the efficiency wall intact.
PaperID: 6070, Poster
Abstract: We identify, for the first time, a new model modification attack—the cross-task model repurposing attack—that can render fingerprints generated by existing model fingerprinting approaches inapplicable. To address this challenge, we propose UniReFP—a novel unified model fingerprinting framework applicable to diverse vision models across different tasks—that utilizes ownership evidence in a shared class-level semantic space rather than original task-specific output space, enabling effective ownership verification when a protected model is repurposed across tasks with heterogeneous output formats. Specifically, UniReFP constructs a surrogate classifier by repurposing the protected model, enabling fingerprints to be generated independently of the original task-dependent output space of the protected model. Moreover, it optimizes grouped fingerprints to encode ownership evidence: fingerprints within the same group are encouraged to induce consistent semantic responses and intermediate representations on the surrogate classifier, while producing dispersed responses on independently trained reference classifiers. During verification, UniReFP uses a task-aware semantic projection mechanism to map heterogeneous suspect-model outputs into unified semantic-presence vectors, and determines ownership by measuring inner-group consistency. Extensive experimental results validate the effectiveness of UniReFP, consistently showing superior performance under cross-task repurposing attacks, commonly studied model modification attacks, and even their combinations.
PaperID: 6071, Poster
Authors: Shenghan Luo
Abstract: The generalized orthogonal Procrustes problem (GOPP) aims to recover rigid transformations that best align multiple point clouds and has broad applications in 3D geometry, computer vision, and biomedicine. Despite its importance, the problem is challenging due to the nonconvex orthogonality constraints. Fortunately, the generalized power method (GPM) has proven highly effective, theoretically guaranteed to converge to the global least squares (LS) estimator under reasonable conditions. Yet, the statistical inference of this global LS estimator remains largely an open problem. Specifically, the absence of an exact second-order analytic expansion prevents rigorous uncertainty quantification (UQ), which is vital for constructing confidence regions and evaluating estimator reliability. While recent UQ frameworks developed for orthogonal group synchronization offer a potential paradigm, they severely lack universality, often failing to accommodate the complex anisotropic distortions inherent in GOPP shape matrices. To bridge this gap, we propose a more general theoretical framework with which we successfully derive the exact second-order analytic expansion of the GOPP global LS estimator under additive Gaussian noise. Leveraging this theoretical foundation, we further construct confidence regions. Finally, extensive numerical experiments empirically validate the exactness of our theoretical expansion and demonstrate the robust coverage of the proposed confidence regions across varying conditions.
Abstract: Diffusion models steer conditional generation with a tunable guidance scale, which users routinely adjust to trade off prompt alignment and diversity. However, these models often reflect and amplify social biases, whose mitigation has been a long-standing concern. Current debiasing techniques are often optimized for a single guidance scale, leaving them vulnerable to fairness degradation when that scale is adjusted. We trace this behavior to a previously overlooked source by decomposing total bias into two components: a model bias and a guidance bias. While prior work primarily targets the former, we show that the guidance bias grows monotonically with the guidance scale, eventually dominating the high-guidance regimes users prefer. To address this, we extend Strong Demographic Parity to guidance and derive a condition under which the target guided distribution retains its group ratio across guidance scales. We propose StayFair, which leverages this condition to design fair guidance algorithms in both regimes. For classifier guidance, it equalizes the classifier's output distributions across groups; for classifier-free guidance, it shifts the null embedding by a prompt-dependent offset. Since StayFair modifies only guidance, it is orthogonal to model debiasing and can be layered onto existing fair diffusion models to extend their fairness across guidance scales. Across class-conditional and text-to-image generation, StayFair decouples fairness from the guidance scale without sacrificing image quality.
PaperID: 6073, Poster
Abstract: Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet they remain limited on tasks requiring fine-grained understanding of small or easily overlooked image regions. Recent approaches mitigate this limitation by allowing models to inspect images through operations such as zooming, cropping, or code execution. While effective, their multi-turn interactions with VLMs introduce redundant token computation and slow inference, and separate region inspection from the reasoning step that depends on it. To address these issues, we propose ason (LoRe), an implicit visual thinking framework for efficient visual reasoning within a single-turn VLM interaction. Given an image and a question, LoRe predicts the most relevant image region, aligns it with the corresponding image patch features, and strengthens the associated key-value (KV) cache entries, enabling subsequent decoding to directly exploit the selected visual cues without additional visual inputs or explicit tool calls. We further employ reinforcement learning to guide the VLM toward identifying informative image regions and leveraging them during reasoning. To support fine-grained evaluation of multimodal reasoning, we introduce COCO-ObjQA, a large-scale benchmark for object-centric multimodal reasoning. Experiments on three challenging multimodal benchmarks demonstrate that LoRe consistently improves VLM reasoning performance with low additional inference overhead.
PaperID: 6074, Poster
Authors:
Tengjv Ru, Weitong Lian, Zecong Tang, Lingyi Meng, Haoran Li, Zhejun Cui, Yichen Zhu, Hangshuo Cao, Qi Kang, Yechi Liu, Kaixuan Wang, Yu-Jie Yuan, Chunwei Wang, Yu Zhang, Bo DaiAbstract: The advancement of multimodal learning is heavily constrained by the scarcity and high annotation cost of high-quality instruction data. While raw videos offer an abundant and dynamic source of visual knowledge, transforming them into precise instruction data remains highly challenging. To bridge this gap, we present VidUEU-Agent, an automated pipeline that synthesizes instruction data from raw videos for multimodal understanding, editing, and unified tasks. VidUEU-Agent follows a four-stage pipeline: it performs global perception to filter low-quality shots, mines keyframes and multimodal signals, constructs metadata, and routes the metadata to synthesize instruction samples. More importantly, our agent framework can autonomously iterate the data distribution based on downstream training results, thereby optimizing data quality, balance, and training effectiveness. Scaling this pipeline, we construct VidUEU-100K, a large-scale dataset of 100K training samples spanning multiple domains and tasks, providing a robust foundation for developing comprehensive multimodal capabilities. Extensive experiments demonstrate that fine-tuning baseline models on VidUEU-100K yields significant and consistent performance gains across diverse representative benchmarks, validating the effectiveness of our proposed agent framework and dataset.
Authors:
Zixing Jia, Yuhang Pan, Ni JiAbstract: Structural reasoning, the ability to recognize and make inference over the relational structure between objects and concepts, is a hallmark of human cognition, yet prevailing methods often collapse relational topology into flat embeddings, cannot discover hidden structure and lack interpretability. We introduce Neural Structural Reasoner (NSR), a brain-inspired network that preserves relational structure directly in the connectivity and dynamics of coupled neuronal populations. NSR draws inspiration from three biological mechanisms: multi-layered architecture for encoding hierarchical knowledge, stable representation of entity and concepts, and path integration for input-driven state inference. At query time, NSR parallelizes computation over candidate relational structures and leverages confidence-weighted scores to perform link prediction. On standard knowledge-graph benchmarks, NSR matches strong embedding and rule-mining baselines with superior training efficiency. Because reasoning is implemented through sequences of human-readable neuron activations, NSR affords native interpretability by tracking intermediate inference steps. The model further extracts latent relational hierarchies and compositional rules, demonstrating the brain-inspired architecture as an effective, efficient, and highly interpretable substrate for structural reasoning.
PaperID: 6076, Poster
Abstract: AI weather models are fast and skillful, yet their errors are strongly conditional because the best member changes with lead time, variable, initialization state, and atmospheric regime. Operational deterministic multi-model forecasting therefore requires more than averaging strong forecasters. It requires anticipating which member deserves trust before future analyses exist. BlendCast turns this anticipation problem into a vision-language forecasting task. Given meteorological maps and structured ensemble diagnostics, the model reads the current weather situation, reasons about prospective member reliability, and converts that judgment into deterministic per-variable ensemble weights. To teach this behavior, BlendCast first gives the model demonstrations of what member-skill reasoning and weight decisions should look like, then refines the resulting policy with multi-variable reinforcement learning so that its own analyses are rewarded when they lead to better, physically consistent forecasts. On held-out 2025 forecasts, BlendCast reduces WRMSE by 2.6% over equal weighting and by 2.0% over conventional ensemble post-processing, closing 63.2% of the gap to oracle blending. These results show that, with multimodal weather evidence and reward-aligned training, VLMs can act as real-time forecasters of forecasters for deterministic AI weather ensembles. Project and demo:https://anonymous.4open.science/status/BlendCast-2567
PaperID: 6077, Poster
Abstract: Large language models increasingly serve as general-purpose reasoning engines, yet they remain unreliable on graph computation tasks whose inputs are discrete, permutation-equivalent, and algorithmic. We propose TRACER, a post-training framework that turns LLMs into native graph computational solvers. Our central thesis is that graph computation failures arise from mismatches within the Transformer computation itself: token order obscures graph symmetries, self-attention is biased toward textual proximity instead of graph topology, and autoregressive decoding weakly constrains algorithmic traces. TRACER addresses these challenges at three coupled levels. At the input level, we introduce permutation-invariant graph encoding, which encourages the LLM to map different serialization variants of a graph to a shared internal graph representation, thereby preserving graph symmetries. At the representation level, we enhance self-attention with the Topology-aware Residual Attention mechanism to inject graph topological signals into the attention kernel. At the reasoning level, we employ process-reward RL to encourage faithful execution of graph algorithmic traces. Across linear-time, polynomial-time, and NP-complete graph tasks, TRACER delivers substantial gains over baselines. These results suggest that equipping LLMs with graph-aware internal computation offers a practical path toward reliable neural solvers for structured algorithmic reasoning. The code is anonymously available
PaperID: 6078, Poster
Abstract: While animals and people tend to learn a task more quickly and reliably when they first train on simpler versions of it, curriculum learning's effectiveness in artificial settings varies, and appears substantially greater in reinforcement learning than in supervised learning. In this work, we construct and analyze a minimal model of policy learning to better understand why. We consider an open-loop, episodic target interception task whose parameters---especially how close the agent must be to the target to receive an informative reward signal---can be chosen so that the agent either receives informative rewards frequently ('dense' rewards), or rarely ('sparse' rewards). We mathematically and numerically analyze the dynamics of REINFORCE-based policy learning in both cases. In the dense reward case, curricula have only marginal benefits; in the sparse case, curricula can be required for learning to occur at all. We find that this is especially true when the agent has the conflicting goals of minimizing effort and accumulating high reward, since this can produce distinct reward landscape optima which compete with one another. Our characterization of useful curricula in this setting indicates that they address two issues: they (i) make rewards less sparse to speed up learning, and (ii) shape the reward landscape to steer agents away from bad local optima.
PaperID: 6079, Poster
Abstract: Reasoning traces improve vision-language models on complex multimodal tasks, but forcing every input to follow a long reasoning process leads to overthinking and unnecessary inference cost. We study adaptive fast/slow thinking for VLMs, where a single model should answer directly for simple inputs and reason step by step only when needed. The key observation is that offline adaptive thinking is not merely a length-control problem: fast direct answers and slow reasoning traces are typically generated by distinct behavior policies. This dual-reference structure makes standard offline reward optimization, such as Decoupled Generation and Optimization (DGO), mismatched because it assumes a single reference policy and can misweight one mode. We propose \textscDual-Reference DGO, which models fast and slow responses under a fast/slow mixture reference and derives a reward-weighted offline objective for adaptive VLM reasoning. Our formulation shows that router-free fast/slow switching emerges from competition between fast and slow partition functions, and yields an efficient mode-wise correction for behavior-reference mismatch. Experiments on multiple datasets show that \textscDual-Reference DGO improves the accuracy--length trade-off over baselines. Ablations show that single-reference DGO tends to collapse to almost always-fast or always-slow behavior, while our dual-reference formulation preserves both modes and learns difficulty-dependent routing.
PaperID: 6080, Poster
Abstract: Targeted adversarial attacks on Large Vision-Language Models (LVLMs) test whether small image perturbations can steer model responses toward attacker-specified content. Under the standard L_\infty constraint, targeted attacks become a regional perturbation budget allocation problem: attack success depends not only on the perturbation objective, but also on which regions receive updates and in what order. Existing localized attacks improve over global perturbations but rely on stochastic spatial sampling, often updating weakly influential regions. We address this limitation through an attention-based analysis showing that cross-modal attention identifies adversarially sensitive regions and that perturbing high-attention hotspots induces predictable redistribution toward subsequent salient regions. These findings motivate attention-guided region sequencing, which begins from dominant hotspots and progressively moves the update support toward next-salient regions. Based on these principles, we propose Stage-wise Attention-Guided Attack (SAGA), a black-box region-sequencing framework that uses a fixed attention map from an open-source LVLM to guide perturbation updates without accessing target-model parameters, gradients, or attention maps. Across ten closed-source and open-source LVLMs, SAGA achieves state-of-the-art attack success rates and the best overall imperceptibility.
Authors: Kingsuk Maitra, Shagun Sood, Morteza Hosseini, Suman Gunnala, Vikram Gupta
Abstract: We resolve the Sanford-Hsu-Telgarsky (SHT) single-layer induction obstruction within the linear, one-step, causal, bilinear, symplectically consistent design class C PSA on the post-RoPE substrate, by lifting standard transformer attention onto a symplectic phase space in direct correspondence with Hairer’s lift of Stormer-Verlet onto a one-step symplectic integrator. The reframing recasts SHT as a filter-order gap: a standard one-layer bilinear score realises a z-transform of joint order (0,0), whereas the induction discriminator requires (1,1). We close the gap by applying the symplectic upper shear M gamma : (q,p) -> (q + gamma p, p) to the post-RoPE query and key streams. We prove that this lift is the unique solution within C PSA (Theorem 4); is operator-level symplectic with zero secular drift (Theorem 6); requires post-RoPE placement (Corollary 7); and, under explicit Assumption Set A, induces a closed-form induction phase transition at gamma c = (d_k log(T-1))^(1/4) (Theorem 8). The SHT lower bound binds where total parameter budget approaches the bit budget of the two-layer induction circuit it forbids: negligible at frontier scale (>= 1B), but decisive in the sub-100M regime that ships in the billions on consumer edge form factors – phones, wearables, microcontrollers, embedded controllers – where it governs whether on-device in-context learning is feasible. We work at 4-92M by design. We report 74.2% single-layer induction at gamma = 2.0, Pearson r >= 0.998 for the closed-form transfer function, r = -0.679 for the embedding-axis low-pass induction filter; and the post-softmax map Frechet-linearises to Differential Transformer (Ye et al., 2024) at first order under the small-signal regime (SS).
PaperID: 6082, Poster
Authors:
Jie Lin, Xiang Liu, Lihao Liu, Liansheng WangAbstract: Clinical case search aims to retrieve disease-consistent prior cases from electronic health record databases for a natural-language patient query under a retrospective retrieval objective. The task is challenging because patient queries contain heterogeneous evidence, including symptoms, diagnoses, laboratory results, treatments, and timelines. These fields are often only partially specified, and fixed retrieval pipelines cannot adapt from prior search trajectories. Existing clinical retrieval methods typically flatten the query into a single representation, use static retrieval pipelines, or lack an explicit mechanism for cross-case adaptation. PatientSearch-SE addresses retrospective disease-centric clinical case search with three components: a \emphsufficiency-aware planner that separates dimensions suitable for direct matching from those requiring hypothesis completion, a \emphhierarchical tool orchestration module that maps these states to retrieval and re-ranking actions, and a \emphtrajectory-guided memory consolidation mechanism that writes reusable disease cards and strategy templates without parameter updates. We also propose an evaluation framework that combines disease-centric retrieval outcomes with trajectory-level process probes for retrieval behavior analysis. Experiments on PMC-Patients and MIMIC-IV show higher disease-centric retrieval metrics for PatientSearch-SE than for the compared same-interface sparse, dense, and reasoning-based baselines, with larger gains on PMC-Patients and more modest gains on MIMIC-IV. Its memory-accumulation results are consistent with controlled external-memory adaptation, and its retrieved cases provide additional evidence for diagnosis-masked downstream few-shot RAG prediction under this retrospective protocol.
Abstract: Large language models (LLMs) are increasingly deployed in teams, yet existing coordination approaches often occupy two extremes. Highly structured methods rely on fixed roles, pipelines, or task decompositions assigned a priori. In contrast, fully unstructured teams enable adaptability and exploration but suffer from inefficiencies such as error propagation, inter-agent conflicts, and wasted resources (measured in time, tokens, or file operations). We introduce Language Agent Teams for Task Evolution (LATTE), a framework for coordinating LLM teams inspired by distributed systems, where processors must operate under partial observability and communication constraints. In LATTE, a team of agents collaboratively construct and maintain a shared, evolving coordination graph which encodes sub-task dependencies, individual agent assignment, and the current state of sub-task progress. This protocol maintains consistency while empowering agents to dynamically allocate work, adapt coordination, and discover new tasks. Across multiple collaborative tasks and a variety of base models, we demonstrate how LATTE reduces token usage, wall-clock time, communication, and coordination failures (e.g. file conflicts and redundant outputs) while matching or exceeding the accuracy of standard designs including MetaGPT, decentralized teams, top-down leader-worker hierarchies, and static decompositions.
PaperID: 6084, Poster
Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable effectiveness in general video understanding. Human motion understanding, a dominant topic in video analytics, however, still remains as a bottleneck due to the lack of explicit structural cues in 2D visual patches and the significant motion information loss caused by sparse visual temporal sampling. To address this, we propose Iris (Integrating representations of in-the-wild skeletons), a novel pose-augmented video MLLM tailored for human-centric motion understanding. Iris adopts an asymmetric dual-stream architecture, pairing the traditional sparse vision stream with a lightweight, high-frequency human pose stream to provide spatially structured and temporally dense kinematic signals. To enable pose-vision modality correspondence awareness, we design a cross-modal spatiotemporal binding mechanism, featuring a pose-structured 3D RoPE for precise temporal alignment and a person-grounded cross-attention module for explicit spatial visual grounding. Furthermore, to enable robust large-scale training, we construct an automated pose curation pipeline that organizes pose data into a delicate heterogeneous representation. Extensive experiments demonstrate that Iris achieves superior performance on multiple fine-grained motion benchmarks, while also maintaining highly efficient token computation overhead. Code will be available upon publication.
Authors:
Natsuto Isogai, Hayata Yamasaki, Sho Sonoda, Mio MuraoAbstract: Quantum machine learning (QML) aims to accelerate machine learning tasks by exploiting quantum computation. Previous work studied a QML algorithm for selecting sparse subnetworks from large shallow neural networks. Instead of directly solving an optimization problem over a large-scale network, this algorithm constructs a sparse subnetwork by sampling hidden nodes from an optimized probability distribution defined using the ridgelet transform. The quantum algorithm performs this sampling in time O(D) in the data dimension D, whereas a naive classical implementation relies on handling exponentially many candidate nodes and hence takes \exp[O(D)] time. In this work, we construct and analyze a quantum-inspired fully classical algorithm for the same sampling task. We show that our algorithm runs in time O(\operatornamepoly(D)), thereby removing the exponential dependence on D from the previous classical approach. Numerical simulations show that the proposed sampler achieves empirical risk comparable to exact sampling from the optimized distribution and substantially lower than sampling from the non-optimized uniform distribution, while also exhibiting exponentially improved runtime scaling compared with the conventional classical implementation. These successful dequantization results show that sparse subnetwork selection via optimized sampling can be achieved classically with polynomial data-dimension scaling on conventional computers without quantum hardware, providing an alternative to the existing quantum algorithm.
PaperID: 6086, Poster
Abstract: Vision Transformers dominate modern image benchmarks, however, what trained self-attention computes head by head remains unclear: prior work either injects inductive biases at initialization or clusters attention patterns visually without committing to a closed-form computation. We ask whether the softmax(QK^\top) inside a trained head can be replaced, without further training, by a structured closed-form operator while preserving prediction and the residual-stream trajectory. We answer yes across ten ViTs and seven pretrainings (ImageNet-1k, ImageNet-21k FT, MAE, DINO, DINOv2, CLIP, SigLIP), and package the result as SOLA: a Structured Operator Library for Attention with three entries, each gated by a per-head diagnostic that upper bounds substitution error. The entries are a fixed 2D convolution kernel for shift heads, a per-image broadcast row for fixed-target heads, and a per-image rank-k SVD of the attention matrix for near-global heads. A rank-adaptive extension picks k per head from the SVD spectrum and extends structural coverage to every head on every backbone, with task-fidelity controlled by a single spectral threshold \tau that trades compression for downstream fidelity. The diagnostics carry over to cross-attention: every decoder cross-attention head in BLIP captioning has effective rank \approx 1 and admits the broadcast substitution. A minimal reproducer that runs the full SOLA pipeline on a single backbone is included in the Supp. Mat. as a demo; the full codebase will be released open-source upon acceptance.
PaperID: 6087, Poster
Abstract: Increasing the replay ratio, where an agent's network is updated multiple times per environment interaction, is an effective strategy for improving sample efficiency in reinforcement learning. However, its effects on the network capacity of multiagent reinforcement learning (MARL) are not yet well investigated. In this paper, we show that high replay ratios induce a severe dormant neuron problem in the centralized global Q-network of MARL, where a large fraction of neurons become inactive, thereby reducing network capacity and destabilizing learning. To address this problem, we propose Ensemble Reset (EnSet) to stabilize MARL training at high replay ratios. First, guided by theoretical analysis, EnSet employs an ensemble of global Q-value networks with periodic resets to mitigate neuron dormancy during frequent updates. Second, EnSet diversifies replay experience by exploiting multiagent translation invariance as an inductive bias in the global Q-value function to prevent overfitting. Extensive experiments in SMAC, MPE, and SMACv2 environments demonstrate that EnSet consistently improves various MARL algorithms at high replay ratios with fewer environment steps.
PaperID: 6088, Poster
Abstract: A striking case of goal misgeneralization was previously observed in OpenAI's Minecraft agent VPT: it killed villagers standing under some leaves, mistaking them for trees. Although this agent was released publicly, enabling white-box interpretability research, few open-weight model organisms of misalignment exist outside the LLM space. In this work, we release MultiSTEVE-1s -- a model zoo of 140 fine-tuned versions of VPT, over 1,000 training checkpoints, and an interpretability suite for analysing them. Specifically, we use the STEVE-1 training procedure to add instruction-following capabilities to VPT with fixed hyperparameters and controlled variations in training randomness. We demonstrate the utility of MultiSTEVE-1s by showcasing the research it enables. First, for some training runs, the only difference is a least-significant bit flip in a single initialised weight. Others differ in the full randomness for weight initialisation and data. Yet, the single bit-flip setting produces agents that act nearly as differently from each other as the full randomness ones. Second, we use our interpretability suite to show that several known VPT attention heads retain their roles after STEVE-1 fine-tuning, while attention strength to the same behaviourally meaningful frame can vary substantially across agents and checkpoints. Finally, even though the agents are similarly capable in in-distribution tasks, the out-of-distribution behaviour of villager killing can differ substantially between them -- in one setting, an agent kills villagers less than 5% of the time, while another agent kills them nearly 50% of the time. Our results show the value of studying multiple similarly trained agents, rather than acting like a behavioural biology lab with only one rat.
PaperID: 6089, Poster
Authors: Yuli Liu, zhiheng zhang
Abstract: Large-scale networked systems are increasingly controlled through a small number of global configuration variables while their local controllers remain black-box, adaptive, and mutually interfering. The operator observes only post-burn-in telemetry, so neither independent-unit causal estimators nor immediate-feedback bandits directly apply. We propose \emphCAUSAL-UCB, a causal-optimistic framework for online steady-state configuration. The framework turns raw network operation into a linked pipeline: an exposure interface maps neighborhood interference into an auditable context; measurement-level exploration creates overlap while the target remains the no-exploration steady-state value \Winf(z;0); a clipped doubly robust evaluator converts exploratory logs into value certificates; and an OFU rule selects the next configuration. We prove a unified high-probability deviation bound that decomposes statistical fluctuation, nuisance product error, clipping bias, exposure approximation, cross-configuration mismatch, context shift, and residual dependence. We also give an identification--deployment frontier showing that reliable counterfactual identification and low deployment cost cannot be optimized independently. Controlled semi-realistic replay and real-data-driven semi-synthetic replay on METR-LA and FlockLab validate the resulting benefit--reliability trade-off across regret, optimality gap, ATE error, confidence width, deployment gap, sensitivity, ablation, and frontier visualizations.
PaperID: 6090, Poster
Authors:
Yiming Xu, Xian Wei, Cheng ChenAbstract: We consider online reinforcement learning for the Stochastic Shortest Path (SSP) problem with unknown transition dynamics. Since standard posterior sampling in SSP lacks an explicit mechanism to enforce uniform optimism and therefore fails to attain minimax-optimal regret, we propose OPSRL-SSP, the first optimistic posterior sampling algorithm for the SSP problem. The algorithm operates in epochs and uses a logarithmic, time-dependent posterior sampling schedule. To ensure optimism, it incorporates a pseudo-state with an optimistic value evaluation into the posterior distribution. We establish a high-probability regret bound of \tilde\mathcalO(B_\star \sqrtSAK + B_\star S^2 A), where B_\star is an upper bound on the expected cost of the optimal policy, S and A are the numbers of states and actions, respectively, and K is the number of episodes. The key technical challenge is to turn these optimistic posterior samples into a uniform optimism guarantee despite the random and potentially unbounded episode lengths of SSP. We address this through an SSP-specific analysis based on dynamic posterior inflation and a contraction argument over the pseudo-state-augmented model. In addition, we extend the analysis to general zero-cost SSPs via a universal cost-perturbation argument. Our dominant term matches the minimax lower bound \Omega(B_\star \sqrtSAK), thereby answering the open problem raised by \citetjafarniajahromi2021onlinelearningstochasticshortest for posterior sampling in the SSP setting.
PaperID: 6091, Poster
Abstract: Decoding language intentions from intracranial electroencephalography signals enables direct communication for patients with severe speech impairments. However, practical deployment remains severely limited by cross-day performance degradation caused by physiological and instrumental heterogeneity, requiring patients to record extensive calibration data daily to maintain decoding performance. Here we introduce Heterogeneity-Guided Data Acquisition and Adaptation (HeDA\textsuperscript2), a novel calibration paradigm that determines which stimuli to record for optimal calibration before any neural data are acquired, fundamentally departing from existing methods that passively adapt to whatever data happen to be available. HeDA\textsuperscript2 operates through three stages: it first quantifies cross-day neural heterogeneity with a proposed Sentence-level Heterogeneity via Alignment Path Entropy that captures how different phonetic phenomena exhibit varying degrees of heterogeneity; second, an estimator learns to predict this heterogeneity from historical patterns, enabling strategic selection of calibration sentences before any neural recording occurs; third, heterogeneity-aware adaptation updates the decoder using minimal acquired data. Extensive experiments on intracranial recordings across diverse recalibration settings demonstrate that, with merely 20% of conventionally required calibration data, HeDA\textsuperscript2 achieves comparable or superior decoding accuracy, representing a critical step toward clinically viable speech prosthesis.
PaperID: 6092, Poster
Abstract: The design of an efficient curriculum has become increasingly important for reinforcement learning with verifiable rewards (RLVR). The majority of existing curriculum strategies select training problems based on difficulty---typically measured by success rate---yet difficulty is an indirect proxy that conflates problem hardness with actual training utility. In this paper, we propose to drive curriculum design not by how hard a problem is, but by how much the model improves from training on it. We formalize this principle through the improvement function (IF), which measures the gain in a prompt's expected reward under an infinitesimal GRPO policy gradient step. Leveraging the IF framework, we formally characterize the recently discovered edge of competence (EoC) phenomenon: a problem is at the model's EoC when its improvement is within a constant factor of the maximum over the training set. Analyzing GRPO on L-step compositional reasoning tasks, we prove that an EoC-induced curriculum achieves target mastery in \widetilde\Theta(\log L_\max / (\eta \log d)) steps---an exponential improvement over the \widetilde\Theta(L_\max / (\eta \log d)) steps required by a uniform mixture of difficulties. Motivated by this theory, we propose ReCUR, a practical algorithm that maintains a stratified retry buffer of previously failed problems, resampling them until they yield positive reward or exhaust a retry budget. Experiments on multimodal reasoning benchmarks show that ReCUR consistently improves the performance of GRPO and DAPO, providing empirical evidence consistent with the improvement-driven curriculum perspective.
PaperID: 6093, Poster
Abstract: Multimodal Large Language Models (MLLMs) have shown increasing potential for 3D scene understanding and spatial reasoning from videos. However, even with strong visual geometry priors, the hidden semantics of existing 3D-enhanced MLLMs can drift across views, leading to degraded grounding, captioning, and spatial reasoning performance. The underlying reason can be their geometry injection as additional tokens or additive features, which enriches visual content but fails to define how view-dependent semantic states should transform. To address this issue, we propose Lie Routing Vision-Language Model (LieVLM), a geometry-conditioned Lie routing framework that treats geometry as an operator over hidden semantics. LieVLM first elicits local Lie states from 3D geometry tokens by predicting latent rotation states and routing coefficients. It then converts these rotation states into triplet-wise Lie group actions to route projected 2D visual semantics. Finally, a geometry-aware routed fusion module combines the original semantic path, the direct 3D context path, and the Lie-routed transformation path for language generation. This design preserves pretrained visual-language semantics while allowing geometry to actively correct view-induced semantic drift. Extensive experiments on 3D scene understanding and spatial reasoning benchmarks demonstrate that LieVLM improves view robustness, especially in severe-pose-change regimes, and consistently outperforms additive geometry fusion methods. Our results suggest a shift in 3D MLLM design, i.e., geometry should not merely be injected as context, but should operate on semantic representations.
Abstract: LLMs are increasingly deployed as agents that interact with external environments and observe feedback such as execution results, error messages, and tool outputs. A well-functioning agent should leverage this evidence to assess its own performance. Yet we find that standard RL algorithms systematically undermine this ability: agents drift toward miscalibration, error flags become unreliable, and reflection carries no information beyond raw task accuracy. The root cause is a credit assignment mismatch in outcome-based RL: for instance, relying on outcome alone can penalize honest error detection on failed trajectories. We propose a simple yet effective fix that augments the outcome reward with a free calibration bonus, computed by contrasting the agent's reflection with the actual outcome---requiring no additional reward model, LLM judge, or external annotation. In a text-to-SQL environment across five benchmarks, our method not only improves task accuracy from 75.1% to 76.5% but also reduces underconfidence rate from 44.4% to 7.7%. The resulting calibrated reflection further enables more effective selective prediction, and supports self-improvement using reflections as pseudo-rewards without outcome supervision.
PaperID: 6095, Poster
Authors:
Yunsheng Xue, Ziyi Zhang, zhihao wu, Youfang LinAbstract: Scaling laws have driven remarkable performance breakthroughs in Computer Vision (CV) and Natural Language Processing (NLP) by increasing model depth. Compared to increasing width, deeper networks provide higher parameter efficiency and more expressive representations. However, increasing network depth in reinforcement learning (RL) still yields limited performance gains. In this paper, through theoretical and empirical analysis, we reveal that this performance degradation is primarily attributed to implicit low-rank bias and overfitting. While these two issues also exist in shallow networks, the detrimental effects caused by component coupling are significantly amplified as the network depth increases, leading to severe performance degradation. For the first issue, we theoretically prove the existence of bootstrapped spectral coupling, which causes high-frequency spectral components to couple with the low-frequency spectral components of the value network, thereby driving representations of deep value networks into severe rank collapse. For the second issue, we reveal that deep value networks severely overfit the noise induced by TD target couplings, which means the construction of TD targets is influenced by other components, such as the boundary of value distribution or current policy. Therefore, our insight is that decoupling is the key to successfully scaling deep value networks in RL. Motivated by this insight, we propose two simple yet effective solutions respectively: policy-independent spectral decoupling and boundary-target decoupling. Moreover, decoupling policy and value learning (as in AWR) is also a necessary target decoupling for deep value network training in offline RL. By integrating these decoupling components, we propose Spectrally Decoupled Distributional Learning (SEED). Empirically, to the best of our knowledge, SEED is the first approach to successfully unlock the potential of depth in both online and offline RL settings, enabling value networks to scale effectively from 3 to 32 layers with consistent performance gains.
Authors: Abhijit Kumar, GUIHUA WU, Mohit Suley
Abstract: Humans know when to reach for help e.g. 347 × 28 warrants a calculator while 2 + 2 does not. Language models, by default, do not. Prompt-based approaches can instruct a model when to invoke tools, but this external scaffolding does not teach the model to recognize the boundary of its own knowledge. Reinforcement learning approaches that assign a single outcome reward to the whole trajectory fare no better: trajectory-level credit cannot isolate which tool call in a successful episode actually helped, nor penalize unnecessary calls. We propose CARL (Competence-Aware Reinforcement Learning), which trains a critic on the model's own rollouts to learn where the model's parametric knowledge suffices and where it needs external help. By decomposing each rollout at natural tool-use boundaries (e.g., code fence delimiters and context block transitions), CARL assigns independent credit to each segment from a single binary outcome, without external judges or step-level annotations, addressing the credit-assignment limitations of trajectory-level methods. As a result, erroneous tool calls, incorrect extractions, and unnecessary calls each receive appropriately signed advantages under a single scheme. We show quantitatively and qualitatively that the trained critic captures the model's domain competence: it separates parametrically solvable from tool-dependent questions with AUC 0.93 at 7B. On five benchmarks spanning arithmetic, multi-hop factual QA, and numerical reasoning over financial tables, CARL improves exact-match accuracy by 6.7 points at 7B and 10.6 points at 3B over the strongest trajectory-level baseline (Search-R1 PPO), with the largest gain (+8.3 EM at 7B, +9.0 EM at 3B) on Musique, the most compositional multi-hop benchmark. Compared to a representative trajectory-level baseline (Search-R1 GRPO), the model issues 56% fewer tool calls on questions answerable from parametric knowledge while remaining ~10 EM points more accurate on those same questions, an emergent consequence of the critic learning where this model's competence ends. Gains are largest at small scale: the 3B model improvement is 1.6× the 7B improvement, suggesting that knowing when to ask for help disproportionately benefits models with smaller parametric memory.
PaperID: 6097, Poster
Authors:
Fanda Fan, Kuiye Ding, Liming Mao, Wang Bingsong, Yao Wang, Xiaorui Wang, Ruijie Jian, Zhipeng Liu, Luqi Gong, Zhenghua Lu, Chunjie Luo, Jianfeng ZhanAbstract: Time series forecasting underpins critical applications in finance, energy, healthcare, and transportation. Although deep models have achieved strong results, most adopt single-scale modeling or restrict multiscale processing to the input side, causing a misalignment between multiscale inputs and single-scale outputs and limiting predictive power. We introduce the Modular Scale-wise Autoregressive Framework (MSAR), a model-agnostic design that forecasts progressively across multiple temporal resolutions. MSAR offers three advantages: (1) scale-wise aligned modeling, which disentangles heterogeneous temporal patterns by aligning inputs and outputs at each scale; (2) scale-wise autoregression, where coarse-scale predictions guide finer-scale forecasting through hierarchical information flow; and (3) a modular architecture, enabling seamless integration with diverse backbones such as CNNs, MLPs, and Transformers. Extensive experiments across a broad set of datasets and forecasting models demonstrate that MSAR achieves consistent improvements in both accuracy and inference efficiency, validating the effectiveness of scale-aligned autoregression for multiscale time series forecasting.
PaperID: 6098, Poster
Abstract: Mechanistic studies of grokking often begin after an endpoint circuit has been identified. We study the preceding retrospective triage problem: given a solved final checkpoint and a long saved trajectory, which path and window should be inspected first? We propose solution-frame path triage. The final checkpoint fixes a solution frame; earlier checkpoints are then scored path by path. The Geometric Coherence Score (GCS) measures coherence in this procedure: it asks whether neighboring inputs in the final frame undergo similar local Jacobian transformations through the attention, MLP, or full block measurement maps. Centered Kernel Alignment (CKA) and L_2 distance to the final path state provide the complementary proximity scores. Across 104 modular arithmetic Transformer runs, GCS reveals a repeatable path schedule that simple norm and dimensionality summaries miss: attention coherence rises near grokking and often turns over, while MLP maps diversify and later reconverge. After selection, held-out checks show that the resulting windows are enriched for accuracy and algebraic changes, align with Fourier progress on modular addition, and identify attention-to-MLP interface states with high replacement cost. The method returns a short list of path/window hypotheses for endpoint analysis. It is retrospective by design: the final checkpoint is part of the procedure.
PaperID: 6099, Poster
Abstract: Safety-critical multi-agent systems require agents to learn coordinated behaviors while avoiding unsafe actions throughout learning. Existing safe reinforcement learning theory has established zero-violation guarantees for single-agent Markov decision processes with instantaneous hard constraints, while much of safe multi-agent reinforcement learning focuses on cumulative constraints, policy optimization, or shielding. We study a stricter setting: cooperative Markov games with unknown dynamics and coupled instantaneous hard constraints, where a joint action must be safe at every time step of every episode. This setting introduces challenges absent from the single-agent case: actions that are locally safe for individual agents may be unsafe jointly, unsafe joint actions can affect future feasible regions, and naive reductions to single-agent safe RL suffer exponential dependence on the number of agents. We propose a graph-structured safe learning algorithm that constructs conservative multi-agent safety certificates and explores optimistically only within certified safe joint subgraphs. Under structured coupling assumptions, the algorithm guarantees zero constraint violation with high probability and achieves sublinear regret against the optimal safe joint policy, with complexity depending on local interaction structure rather than the full joint action space.
Abstract: Pre-training large language models on massive GPU clusters has made hardware faults routine rather than rare, driving the need for resilient training systems. Yet existing frameworks either focus on specific parallelism schemes or risk drifting away from a failure-free training trajectory. We propose ReCoVer, a resilient LLM pre-training system that upholds a single invariant: each iteration keeps the number of microbatches constant, ensuring per-iteration gradients remain stochastically equivalent to a failure-free run. The framework is organized as three decoupled protocol layers: (1) Fault-tolerant collectives that isolate faults from propagating across replicas; (2) in-step fine-grained recovery that preserves intra-iteration progress and prevents gradient corruption; (3) versatile-workload policy that dynamically redistributes microbatch quotas across the survivors. The design is parallelism-agnostic, integrating directly with both 3D parallelism and Hybrid Sharded Data Parallel (HSDP) as a drop-in substrate. We evaluate our implementation on end-to-end pre-training tasks for up to 512 GPUs, ReCoVer successfully preserves the training trajectory from a failure-free reference despite of 256 GPUs lost spread across the run. For comparison with checkpoint-and-restart baselines, ReCoVer demonstrates 2.23× higher effective throughput after successive failures. This advantage results in ReCoVer processing 74.9% more tokens at 234 GPU-hours, with the gap widening as the training prolongs.
PaperID: 6101, Poster
Authors: Jinchi Zhu, Thomas T Luo
Abstract: Foundation models pretrained on similar data distributions tend to encode semantically aligned concepts; yet these representations remain entangled with model-specific inductive biases, which hinders both unified interpretability and cross-model feature reuse. Existing approaches address this by aligning concepts across model-specific representation spaces, but do not explicitly account for model-specific biases, resulting in limited cross-model reconstruction fidelity (low R^2) and poor feature transferability. We propose the Gated Bottleneck Sparse Autoencoder (GB-SAE), a framework that takes multiple—homogeneous or heterogeneous—foundation models as input and factorizes their representations into shared and private components. GB-SAE jointly learns a single model-agnostic semantic dictionary and, for each model, a private residual dictionary, via a learnable gating mechanism that encourages competition between the two. Conceptually, we cast unified multi-model interpretability as a representation factorization problem and solve it through bottlenecked sparse dictionary learning. Empirically, GB-SAE achieves unified and faithful interpretability across diverse architectures, high reusability of the learned shared dictionary, and strong generalization to downstream tasks—capabilities that alignment-based paradigms, by design, cannot fully support.
PaperID: 6102, Poster
Authors: Subham Patel, Himanshu Pandey, RATIKANTA BEHERA
Abstract: Neural operators have shown strong performance in learning solution mappings for differential equations and offer fast inference across parameter spaces. However, they often suffer from spectral bias, leading to poor resolution of high-frequency and localized features, especially in out-of-distribution settings. In this work, we propose adaptive multiscale operator correction (AMOC), a hybrid framework that combines neural operator learning with physics-constrained optimization in a low-dimensional, instance-specific adaptive spectral subspace. First, a neural operator learns a coarse solution, which is then projected onto a sparse multiscale basis using either wavelet family or Fourier modes, where dominant components are selected based on energy. Finally, a physics-informed optimization is performed on adaptive spectral subspace, significantly lowering computational cost while improving accuracy. The proposed approach enables efficient correction of neural operator predictions by leveraging the sparsity and localization of multiscale representations, while bypassing automatic differentiation in the residual loss through closed-form derivatives of the chosen basis. The proposed AMOC is evaluated on the heat, Poisson, Darcy, and Helmholtz equations, achieving up to four orders-of-magnitude error reduction over FNO, with particularly strong gains observed in challenging out-of-distribution regimes.
Abstract: 4D hand motion reconstruction from egocentric video is bottlenecked by clear limitations of existing methods: image-based pipelines depend on a detector that fails under heavy occlusion, while video-based methods rely on temporal modules learned only from hand-pose annotations, a signal too narrow to capture motion, occlusion, and hand-object interaction. These capabilities, however, are exactly what video generative models must implicitly acquire when trained to synthesize coherent video at internet scale. Motivated by this, we present ViDiHand, which leverages the representations of a pretrained video diffusion model to reconstruct 4D two-hand pose. We adapt it via a hand-overlay rendering objective that specializes its features for hands while preserving its world priors, and a decoder recovers metric-scale pose from the adapted features. The whole pipeline runs in a single forward pass over full frames—no detector, no infiller, and no test-time optimization. On ARCTIC, HOT3D, and HOI4D, ViDiHand substantially outperforms prior methods, establishing video diffusion models as a powerful new foundation for hand motion reconstruction and a promising route to scalable in-the-wild data collection for embodied AI.
PaperID: 6104, Poster
Abstract: Diffusion Language Models (DLMs) have demonstrated strong scaling capacity as alternatives to autoregressive language models. However, their performance is highly sensitive to the choice of transition kernels, and poorly designed kernels can lead to issues like training instability, slow convergence, and biased sampling. In this paper, we study this sensitivity through a principled analysis of generalization error and identify three critical factors: asymptotic bias (difficulty in approximating the posterior distribution), exposure bias (error propagation during sampling), and optimization variance induced by kernel dispersion. We further compare different transition kernels: masking diffusion yields sparse and easier posterior-approximation targets, while uniform diffusion provides stronger sampling-side repair but induces harder approximation. Motivated by this trade-off, we revisit a previously overlooked variant, semantic DLM (SemDLM), where the transition kernel corrupts tokens to neighborhoods that are semantically similar. Our theory suggests that SemDLM can serve as a plausible middle ground by reducing the posterior approximation difficulty of uniform diffusion while retaining repair ability. However, we find that SemDLM suffers from a semantic basin problem, where sampling repeatedly stays within a semantic region and produces low-diversity text. To address this, we propose SemDLM+, which adds a global transition and a semantic-frequency penalty during sampling. Experiments on LM1B and OpenWebText show that SemDLM+ improves training dynamics and achieves competitive language modeling and generation quality with satisfactory diversity.
Abstract: Large language models are increasingly used as proxies for human subjects in social science research, yet external validity requires matching the response distributions of target human populations. We study population-level survey alignment: reconstructing aggregate survey responses from limited public survey data, without individual-level demographic profiles or model finetuning. We formalize this problem as preference reconstruction: rather than matching proxy agents to demographic profiles, we construct a functional basis of proxy agents and recovering population preferences by weighted aggregation. We instantiate this idea via Prompts to Proxies (\textttP2P), a two-stage inference-time system. Stage 1 uses structured attribute-based prompting with entropy-guided adaptive sampling to construct a diverse proxy pool spanning the latent preference space. Stage 2 employs L1-regularized regression to select a compact weighted ensemble matching observed target-population responses. Across 14 American Trends Panel waves, \textttP2P achieves an average test MSE of 0.014 at approximately 0.8 USD per survey, improving over prompting and demographic-conditioning baselines. On the World Values Survey, cross-locale transfer experiments show that basis expressiveness can outperform locale matching on average, while locale-specific generation helps on culturally divergent questions. A stress test against an SFT-aligned survey model shows competitive performance using less than 3% of the training data. These results position preference reconstruction as a lightweight, externally verifiable alternative for survey-based population alignment.
Authors: Jingyang You, Hanna Kurniawati
Abstract: Bayesian Reinforcement Learning (BRL), a subclass of Meta-Reinforcement Learning (Meta-RL), provides a principled framework for generalisation by explicitly incorporating Bayesian task parameters into transition and reward models. However, classical BRL methods assume known forms of transition and reward models. While recent deep BRL methods incorporate model learning to address this, applying neural networks directly to joint data and task parameters necessitates variational inference. This often yields indistinct task representations, compromising the resulting BRL policies. To overcome these limitations, we introduce Generalised Linear Models in Deep Bayesian RL with Learnable Basis Functions (GLiBRL). Our approach features fully tractable Bayesian inference over task parameters and model noise, alongside exact marginal likelihood evaluation for learning transition and reward models. The permutation-invariant nature of exact Bayesian inference in GLiBRL enables seamless integration with both on-policy and off-policy RL algorithms. We further show that GLiBRL admits a closed-form relationship between the \mathcalL_2 distance of its task representations and empirical kernel-based correspondence between task samples, which is to our knowledge the first such structural result for online deep BRL. GLiBRL is compared against representative and recent Meta-RL methods, and improves state-of-the-art performance on both MuJoCo and MetaWorld benchmarks by up to 1.8×.
Abstract: Real-world image restoration is challenging due to complex and interacting mixed degradations. Recent agent-based approaches address this problem by composing multiple task-specific restoration tools. However, empirical analysis reveals that their performance is fundamentally limited by implicitly constrained planning spaces and the lack of coordination among independently pretrained tools. To address these issues, we propose OPERA (Optimized Planning-Execution Restoration Agent), a framework that jointly optimizes restoration planning and tool execution in an end-to-end manner. On the planning side, OPERA uses reinforcement learning to directly optimize tool composition over a combinatorial plan space, with the final restoration quality as the reward. On the execution side, OPERA introduces agent-guided co-training of restoration tools, enabling them to learn cooperative behaviors under sequential composition. Extensive experiments on multi-degradation benchmarks and real-world datasets demonstrate that OPERA consistently outperforms both all-in-one restoration models and existing agent-based methods across diverse and complex degradation scenarios.
PaperID: 6108, Poster
Abstract: Score-based valuation methods, such as Shapley values and Leave-one-out (LOO), are widely used to assign value to data in modern machine learning pipelines, including for tasks such as attribution, selection, and pricing, yet it remains unclear when these scalar scores reliably guide downstream decisions. We show that their success is governed by three structural properties of the learning problem: substitutability, complementarity, and non-monotonicity. Substitutability (redundancy) can collapse pointwise credit, causing Shapley and LOO to fail even under monotone submodular valuations; bounded curvature limits this collapse and helps recover constant-factor approximation. Complementarity can break common score-based rules and greedy-style adaptive selection, though these effects diminish with sufficient coverage. Non-monotonicity implies that all non-adaptive methods, including score-based approaches, can fail, establishing a separation from adaptive algorithms. Our theoretical results, supported by empirical evidence, provide a structural view of data valuation and motivate a simple practical pipeline: deduplicate to reduce redundancy, ensure coverage to suppress complementarity, and then choose between score-based or adaptive methods based on non-monotonic effects.
PaperID: 6109, Poster
Abstract: 3D geometric understanding is undergoing a major transition where rigid, deterministic, optimization-based pipelines are being replaced by large-scale, deep learning models. While recent methods achieve impressive performance on tasks such as multi-view 3D reconstruction and camera pose estimation, they suffer from two critical limitations: (i): current state-of-the-art (SOTA) models require industry-level compute and training corpora, severely limiting accessibility, and (ii): recent approaches remain constrained to modifying existing pretrained transformer models, since substantially changing their architecture would require retraining from scratch. We instead study a complementary direction: training an attention-free state space model for dense 3D understanding. We introduce FlashSSM-3D, the first 3D reconstruction state space model capable of performing accurate geometric reasoning over hundreds of images with linear computational complexity. However, training such a model from scratch with competitive performance is impractical under typical academic compute constraints. To address this, we train FlashSSM-3D by distilling knowledge from VGGT, avoiding the need to learn 3D geometric priors entirely from scratch. However, standard distillation losses are poorly suited for dense 3D reconstruction, where preserving real-world spatial structure is essential. We therefore propose a geometry-aware knowledge distillation method specifically tailored for large-scale 3D reconstruction models. Our method introduces two novel losses: a Procrustes-based alignment loss that enforces global structural consistency between student and teacher 3D point predictions, and a camera reprojection loss that directly supervises view-consistent camera geometry between teacher and student. Extensive experiments across multi-view 3D reconstruction and camera pose estimation benchmarks demonstrate that FlashSSM-3D achieves competitive performance while requiring only ~25% of the training data and less than 10% of the training compute used by VGGT. At inference time, FlashSSM-3D benefits from the practical efficiency of state space models, reconstructing scenes from 500+ views in under 10 seconds while maintaining competitive accuracy.
PaperID: 6110, Poster
Abstract: Quantizing Mixture-of-Experts (MoE) models is difficult for two coupled reasons: expert quantization error perturbs expert outputs, and sparse routing amplifies small perturbations into discrete expert selection changes. This coupling produces progressive output drift that is not captured by dense model quantization analyses. To address this gap, we propose a unified quantization framework that jointly handles both sources of error. First, we formulate compression-induced routing inconsistency as a form of training-inference router mismatch and use this connection to motivate limited routing replay, which replays BF16 routes during quantization-aware fine-tuning (QAT), freezes router parameters, and updates only expert-side adapter parameters. We then provide a standard convergence guarantee for the resulting fixed-route QAT objective, together with a routing-consistency bound under explicit router-margin conditions. Second, for expert restoration under a fixed rank budget, we propose load-aware low-rank compensation that assigns larger compensation rank to heavily loaded experts. Under this allocation rule, we prove monotone strict decrease of weighted compression error and characterize optimal rank assignment by marginal load-weighted spectral gains. Finally, experiments on the MXFP4 quantization of Mixtral, DeepSeek, and Qwen3 MoE models show that our method improves perplexity and downstream accuracy over prior methods. We also report system efficiency on Qwen3-30B, where uniform W4A8 serving accelerates decoding and our adapter path preserves most of that acceleration.
Abstract: Text dominance, the tendency of multimodal large language models (MLLMs) to over-attend to textual tokens while under-utilizing non-text inputs, has been observed in vision-language settings, but whether it extends to other modalities remains an open question. We present a systematic, attention-level study of this phenomenon across five modalities (image, video, audio, time-series, and graph), using two diagnostic metrics: the Modality Dominance Index (MDI), which compares per-token attention between text and non-text inputs, and the Attention Efficiency Index (AEI), which normalizes attention share by token share. Across ten models and six benchmarks, text dominance is pervasive in image, video, audio, and time-series settings and intensifies in deeper layers. Controlled token-replication experiments show that expanding non-text sequences without adding semantic content amplifies the imbalance, while graph tasks provide a boundary case where compact, information-dense tokens reverse the effect. Guided by these findings, we evaluate attention-based token compression as a proof-of-concept intervention on the vision modality. On LLaVA-1.5-7B with MMMU-Pro, removing 90% of visual tokens via informed selection preserves accuracy (33.70% vs.\ 33.23% baseline), while random dropping under the same ratio degrades it to 30.29%; informed selection also lowers the late-layer MDI from 15.63 to 3.46 and reduces latency by 3.5×. These results suggest that text dominance is a cross-modal phenomenon associated with token redundancy, and that token compression can reduce it in the tested vision setting while preserving task performance.
PaperID: 6112, Poster
Abstract: Long-horizon video generation is increasingly moving toward autoregressive rollout, where each newly generated segment becomes part of the evidence used for future prediction. Recent causal decoders, long-context training, and KV-cache mechanisms have extended feasible duration, but most inference pipelines still decide the visible past implicitly through recent windows, compressed caches, or external memory. This raises a direct question: can the choice of which past frames or segments to show to the decoder be treated as an explicit test-time decision? We answer this question with \textscReCR, a training-free method that formulates autoregressive video decoding as representation-guided context rewriting. At each step, \textscReCR builds a selected visible context from the generated history by using internal decoder representations to favor context units that match the global rollout state, avoid redundancy with the recent continuation, preserve temporal boundary coverage, and provide useful intermediate evidence. The selected units are then mapped onto a compact causal axis before decoding the next unit. Across long-horizon text-to-video benchmarks, \textscReCR consistently improves multiple autoregressive backbones under matched visible-token budgets. Empirical analyses such as representation-space diagnostics and transition-count prompt switching further show that improving what remains visible is an effective path toward more stable long-video decoding. Our code is available at https://anonymous.4open.science/r/anonymous-ReCR-97EB/
PaperID: 6113, Poster
Abstract: We study cost-aware cascading bandits, where a learner selects an ordered subset of options, tests them sequentially until the first success, and pays the costs of all tested options. In this problem, regret comes from both testing inefficient options and placing options in a suboptimal order, but existing analyses do not separate these effects for inefficient options and therefore yield an inverse-square dependence on the gap c_i-\theta_i. We develop a regret decomposition based on intermediate policies that reorder the remaining suffix and remove inefficient options one position at a time. This allows us to quantify the incremental regret incurred when each option is tested. As a consequence, we show that the regret of CC-UCB admits a problem-dependent bound of O(\sum_i:\theta_i/c_i<1\log T/(c_i-\theta_i)) up to additive terms and a problem-independent bound of order \tilde O(L\sqrtT). We further prove that the minimax regret is bounded from below by \Omega(\sqrtLT) by reducing standard multi-armed bandits to a special case of the model. Finally, we propose CC-UCBv2, which removes the need to specify a positive lower bound on costs and handles zero-cost options by separating empirically zero-cost options from the others. Numerical experiments show the effect of misspecified cost lower bounds and demonstrate that the proposed modification can reduce regret in representative instances involving zero or misspecified costs.
PaperID: 6114, Poster
Abstract: Latent action models (LAMs) aim to learn compact representations of state transitions directly from visual observations, but they suffer from a fundamental ambiguity: the latent code can encode target-state information rather than true dynamics. Existing approaches address this issue through restrictive bottlenecks, which reduce leakage at the cost of limiting expressivity. We propose StaDy, a regularization framework that instead enforces conditional informativeness: the latent code should aid prediction only when paired with the source state, while remaining uninformative about the target on its own. Concretely, we introduce a likelihood-matching objective that aligns the decoder’s predictions conditioned solely on the latent variable with its unconditional predictions. This discourages target memorization without constraining latent capacity. Experiments in video modeling show that, unlike bottleneck-based methods, StaDy scales effectively with increased latent dimensionality, reducing static content leakage while improving the representation of more complex dynamics.
Abstract: Unlearning in large language models (LLMs) requires precisely removing specific information from a pre-trained model, e.g. to delete private data or harmful knowledge acquired during pre-training. However, existing unlearning methods often fall short under thorough evaluation. This paper makes three contributions. First, we introduce JensUn, a novel unlearning approach which leverages the Jensen-Shannon Divergence (JSD) as the training objective for both forget and retain sets. We provide theoretical arguments why JSD outperforms KL-Divergence, and show in extensive experiments that JensUn achieves the best forget-utility trade-off and demonstrates strong resilience to benign relearning, outperforming existing methods. Second, for a precise unlearning evaluation, we introduce LPF, a curated dataset of less prominent facts that provides a realistic unlearning scenario. Third, to comprehensively test unlearning methods, we (i) employ an LLM as semantic judge instead of the standard ROUGE score, and (ii) evaluate worst-case unlearning over various paraphrases and input formats. Our improved evaluation reveals that many unlearning methods are less effective than previously thought.
Authors:
Yuan Wang, Ouxiang Li, Yulong Xu, Borui Liao, Jiajun Liang, Jinghan Li, Meng Wang, Xintao Wang, Pengfei Wan, Kuien Liu, Xiang WangAbstract: Recent advances in generative video models are increasingly driven by post-training and test-time scaling, both of which critically depend on the quality of video reward models (RMs). An ideal reward model should predict accurate rewards that align with human preferences across diverse scenarios. However, existing paradigms face a fundamental dilemma: Discriminative RMs regress rewards directly on features extracted by multimodal large language models (MLLMs) without explicit reasoning, making them prone to shortcut learning and heavily reliant on massive data scaling for generalization. In contrast, Generative RMs with Chain-of-Thought (CoT) reasoning exhibit superior interpretability and generalization potential, as they leverage fine-grained semantic supervision to internalize the rationales behind human preferences. However, they suffer from inherent optimization bottlenecks due to the coupling of reasoning and scoring within a single autoregressive inference chain. To harness the generalization benefits of CoT reasoning while mitigating the training instability of coupled reasoning and scoring, we introduce DeScore, a training-efficient and generalizable video reward model. DeScore employs a decoupled ``think-then-score'' paradigm: an MLLM first generates an explicit CoT, followed by a dedicated discriminative scoring module consisting of a learnable query token and a regression head that predicts the final reward. DeScore is optimized via a two-stage framework: (1) a discriminative cold start incorporating a random mask mechanism to ensure robust scoring capabilities, and (2) a dual-objective reinforcement learning stage that independently refines CoT reasoning quality and calibrates the final reward, ensuring that higher-quality reasoning directly translates to superior model performance. Empirical evaluations demonstrate that DeScore achieves superior training efficiency and optimization stability, while outperforming state-of-the-art methods across diverse in-domain and out-of-distribution benchmarks. Moreover, DeScore also proves effective for post-training, leading to improved generated video quality.
PaperID: 6117, Poster
Authors:
SEONGEUN HONG, JuYeong Hwang, Jinhyun Kim, Hanyoung Jang, Hyeongyeop KangAbstract: Selecting a next motion in a populated scene is not just a matter of avoiding what is currently in the way: an action can be locally feasible yet drift onto the road instead of the crosswalk, cut into a stream of crossing pedestrians, or bottleneck a narrow passage. These failure modes are invisible from the current state alone, and acting on them requires reasoning about what each candidate action will cause. World models offer exactly this, namely action-conditioned predictions of the future, but rolling them out at every decision step is a cost deployable navigation systems cannot pay, and the field has largely traded foresight for speed. We argue this is a false dichotomy. The natural place for a world model in human-aware navigation is at training time, as a preference oracle that distills foresight into a fast policy, rather than at inference time, as an online planner. We pair a discrete visual action prior with two task-aligned world models, one for nearby social dynamics and one for future scene semantics, that score candidate actions by their predicted consequences and refine the policy through group-relative preference distillation. At deployment, both world models are removed and the policy acts in a single forward pass. Across five held-out scenes, this yields navigation that is simultaneously safer, more socially and group-coherent, and more goal-directed than strong reactive and learned baselines, with naive observers rating the resulting motion as more human-like. Taken together, our findings point to a different approach for bringing world-model foresight into deployable agents: not accelerating rollouts, but amortizing them into the policy so that the cost of looking ahead is confined to training.
Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has driven recent progress in code large language models by leveraging execution-based feedback from unit tests, but its scalability is fundamentally constrained by the availability and reliability of high-quality test cases. We propose , a reward model designed to scale both reinforcement learning training and test-time inference for code generation. is trained on carefully curated preference data derived from verified code problems and incorporates syntax-aware code extraction and validity-preserving reward shaping to ensure stable and robust optimization. Across four coding benchmarks, serves as an effective test-time scaling method, achieving performance comparable to unit test approaches while providing a
PaperID: 6119, Poster
Abstract: Non-contrast CT (NCCT) is widely used in clinical practice, yet its limited soft-tissue contrast often renders visual information insufficient for accurate diagnosis. In routine radiology workflows, NCCT is commonly accompanied by contrast-enhanced CT, radiology reports, and pathology reports, which provide complementary knowledge at different levels, from structural and morpho-functional cues to semantic findings and pathological evidence. However, existing pretraining paradigms address this challenge only partially, each introducing complementary but still limited constraints on NCCT representations. To address this limitation, we propose RadOmni, a unified omni-pretraining framework that injects multimodal radiology knowledge into representation learning for non-contrast CT. RadOmni formulates a curriculum consolidation learning strategy, in which knowledge is progressively injected according to optimization difficulty and knowledge level through four stages: masked image modeling on non-contrast CT, cross-phase transfer of anatomical structures and enhanced patterns from contrast-enhanced CT, report-guided semantic learning, and pathology-informed diagnostic enhancement. To mitigate optimization conflicts across heterogeneous objectives and reduce knowledge forgetting, each stage retains and jointly optimizes the objectives from preceding stages, while early encoder layers are selectively frozen to preserve generic representations. To enable such multi-source pretraining, we curate Omni-CT10K, a large-scale CT dataset with more than 10K NCCT scans featuring clinically co-occurring multimodal supervision. Extensive experiments on Omni-CT10K, two in-house benchmarks from different centers, and the public MSD benchmark show that RadOmni consistently achieves strong performance across classification, segmentation, pathological TNM staging, and radiology report generation tasks. Further analyses demonstrate favorable scaling behavior, robust cross-center generalization, and transferable representations.
Authors:
Qinglin Zhu, Yizhen Yao, Runcong Zhao, Yanzheng Xiang, Siya Qi, Amrutha Saseendran, Chen Jin, Philip A Teare, Bin Liang, Yulan He, Lin GuiAbstract: Autoregressive (AR) models remain the dominant paradigm for natural language generation, but their strictly sequential decoding process leads to high inference latency. Recent diffusion-inspired language models, such as LLaDA and Dream, alleviate this issue through parallel generation, yet they still face two key limitations: information loss, since predictive distributions over non-finalised tokens are discarded at each step, and unstable commitment dynamics, where local token decisions are not sufficiently coordinated at the global level. We propose Latent Refinement Decoding (LRD), a two-stage decoding framework consisting of Latent Refinement and a Predictive Feedback Loop. In the first stage, LRD keeps masked positions as distributional mixtures of predicted tokens and the mask embedding, enabling the model to form more globally consistent beliefs before committing tokens. In the second stage, LRD progressively finalises confident tokens while preserving uncertain ones for further iterative refinement, with KL-divergence dynamics serving as a stable criterion for convergence and early stopping. Experiments show that LRD improves performance on both coding tasks, including HumanEval by 6.3 and MBPP by 2.6, and reasoning tasks, including GSM8K by 2.9 and MATH500 by 3.8, while achieving up to 10.6× overall speedup. Moreover, LRD is compatible with single-token decoding, multi-token threshold commitment, and system-level accelerators; when combined with Fast-dLLM, it reaches up to 18.4× speedup over vanilla decoding while further improving accuracy.
PaperID: 6121, Poster
Authors: Leonard Boussioux, Henry Mao, Dimitris Bertsimas
Abstract: Accurate time-series forecasting is critical for a wide range of problems involving temporal data. Ensemble modeling is a well-established technique for leveraging multiple predictive models to increase accuracy and robustness, as the performance of a single predictor can be highly variable due to shifts in the underlying data distribution. This paper proposes a new methodology for building robust ensembles of time series forecasting models. Our approach uses Adaptive Robust Optimization (ARO) to construct a linear regression ensemble whose model weights adapt over time. We demonstrate the effectiveness of our method through a series of synthetic experiments and real-world applications, including air pollution management, energy consumption forecasting, and tropical cyclone intensity forecasting. Our results show that our adaptive ensembles outperform the best ensemble member in hindsight by 16-26% in root mean square error and 14-28% in conditional value at risk and improve over competitive ensemble techniques.
Abstract: A core quality of general intelligence is the ability to open-endedly expand and evolve its set of mastered skills autonomously. While recent Foundation Model (FM) driven approaches have shown promising results towards this goal, they typically rely on significant human-in-the-loop engineering, limiting their transferability to novel environments. To address this, we introduce Continuous Open-ended Discovery and Evolution of Skills as Hierarchical Reward Programs (CODE-SHARP), a framework that leverages FMs to open-endedly grow and evolve an archive of Python programs encoding skills to train a generalist agent policy entirely from scratch via reinforcement learning, directly from source code. These programs, termed Skills as Hierarchical Reward Programs (SHARPs), each encode a local success condition and a set of prerequisites delegated to previously discovered SHARPs. At runtime, SHARPs dynamically route the agent through their prerequisite chain based on the current state, rewarding each completion along the way, requiring the agent to learn only the marginal behaviour each new SHARP introduces, enabling efficient learning of long-horizon skills without any pre-defined rewards. On Craftax-Classic and XLand, agents trained fully autonomously by CODE-SHARP outperform previous works by 6x and 2.6x in median performance and are the only agents capable of crafting iron tools and mining diamonds. Scaled to Craftax-Extended, CODE-SHARP trains a generalist agent on over 90 discovered SHARPs, enabling the agent to solve challenging long-horizon tasks zero-shot, matching agents trained on ground-truth rewards.
PaperID: 6123, Poster
Abstract: 3D scene graph prediction is commonly supervised with local object and predicate classification losses. Although effective for slot-wise label prediction, this formulation does not provide a graph-level semantic distance between a predicted scene graph and its ground truth: semantically mild label errors and structurally disruptive relational errors can be treated similarly, and subject-predicate-object facts are not compared as scene-level units. We propose Contextual Hellinger Triplet Geometry, a graph-level supervision objective that measures the semantic and structural discrepancy between 3D scene graphs. Our key idea is to reinterpret the conventional classifier outputs of a 3D scene graph predictor as an aligned probabilistic scene graph field, where object and predicate predictions jointly define contextual subject-predicate-object facts. This formulation yields a differentiable metric-based training objective with an efficient factorized implementation and can be applied to existing 3D scene graph prediction architectures without modifying their network design. Experiments on 3DSSG demonstrate that our supervision improves downstream graph-conditioned 3D scene generation, while ablation studies confirm the complementary roles of the proposed graph-level terms and user studies show better graph descriptiveness, relation plausibility, and structural preservation.
PaperID: 6124, Poster
Abstract: Large Language Models (LLMs) have shown promise for empowering search agents to tackle complex information-seeking tasks. However, since acquiring external information necessitates web search, these agents are highly susceptible to various adversarial attacks (e.g., malicious information injection or stealthy context manipulation). Existing attack strategies inevitably rely on static and human-crafted templates, failing to emulate the dynamic and complex nature of real-world threats. In this paper, we propose an adversarial self-evolution framework that integrates bi-level optimization between the attacker's strategy evolution and the defender's adaptive mitigation, to expose the vulnerabilities of search agents. Specifically, we introduce the Contrastive Rollout Evolutionary Optimization (CREO) method to drive the directional evolution of the attack strategy population via contrastive evolutionary signals. To provide a stable environment for this continuous evolution, we further construct a controlled sandbox and design fine-grained metrics to quantify internal vulnerabilities. Extensive experiments across various multi-hop QA benchmarks and frontier LLMs demonstrate that the proposed framework yields attack efficacy superior to static baselines, underscoring the urgent need to develop more robust defense frameworks. The anonymized code repository is available at https://anonymous.4open.science/r/adversarial_rag-3EFA.
Authors: Matthew Smart, Soumya Ganguly, Nilava Metya, Alexandre V Morozov, Anirvan Sengupta
Abstract: We study minimal attention-only transformers under all-token corruption and show they admit a two-stage empirical Bayes interpretation. A single attention step computes a kernel-weighted posterior mean with respect to the empirical distribution defined by the context. Depth refines this distribution through particle dynamics (Stage 1), while a long-range skip-connection carries the noisy input as a query for posterior inference (Stage 2), revealing distinct statistical roles for depth and attention residuals. The framework isolates a minimal setting in which the context itself induces a depth-dependent energy landscape governing in-context inference. We show that effective denoising can emerge without an explicit noise schedule: a fixed kernel bandwidth and finite integration horizon suffice, yielding a principled depth–noise relationship. We further establish a posterior-mean recovery guarantee for a class of well-behaved priors, where the empirical estimator converges to the Bayes-optimal predictor under asymptotic conditions. Connecting these dynamics to reverse-diffusion limits, our results provide a statistical interpretation of attention as in-context inference via sample-based posterior estimation, without explicit density modeling.
PaperID: 6126, Poster
Abstract: Lipreading remains challenging due to the inherent ambiguity of visual speech cues and their sensitivity to head pose and video resolution. In this work, we propose FLARE, a framework for robust visual speech representation learning via cross-modal distillation and joint face-lip modeling. FLARE leverages semantically rich Whisper representations as the acoustic teacher signal, converting them into discrete distillation targets through a random-projection quantizer and pretraining the model with a masked prediction objective. To capture complementary visual information, FLARE employs a dual-stream architecture that jointly processes lip region-of-interest (ROI) and full-face inputs, where the lip stream captures fine-grained articulatory motion and the face stream encodes broader facial dynamics. During pretraining, auxiliary audio input is incorporated to facilitate cross-modal alignment, and modality dropout is applied to encourage robust visual-only representations. For downstream lipreading, FLARE serves as a visual encoder coupled with either a lightweight Transformer decoder or an LLM-based decoder. Evaluated on LRS2, LRS3, and the out-of-domain WildVSR benchmark, FLARE achieves state-of-the-art word error rates of 11.7%, 13.2%, and 34.3%, respectively. A progressive design study confirms the contribution of each major component. Further analyses demonstrate that joint face-lip modeling improves robustness under degraded video resolution and large head pose variation, while yielding more phoneme-discriminative visual representations. Code and pretrained models are available at https://anonymous.4open.science/r/FLARE-522C.
PaperID: 6127, Poster
Abstract: Novel view synthesis methods such as 3D Gaussian Splatting degrade under sparse-view settings, suffering from an inability to extrapolate beyond observed regions. To enable high fidelity reconstruction and rendering from sparse views, inpainting of sparsely observed areas and completion of occluded regions is required. Unlike existing approaches that perform these tasks using 2D inpainting diffusion models or geometric-only 3D completion models, we propose a fully 3D scene completion framework, called InvisiSplat, that performs geometry and appearance completion directly in 3D. Our method consists of a Gaussian Splat VAE which learns a sparse point cloud latent space that predicts Gaussian Splat parameters from the latents, as well as a two-step flow model that learns to generate samples conditioned on the point cloud latents, different from other approaches that use tokens without 3D coordinates. Our flow model utilizes alternating local 3D attention and global attention, which balances the needs for precise local geometry and large-scale spatial context. Across several datasets, we demonstrate that our method produces 33% - 82% higher fidelity 3D geometry over prior work while attaining competitive performance on appearance metrics when compared to existing approaches.
PaperID: 6128, Poster
Abstract: Bulk gene expression profiling, which aggregates pooled RNA across cells within a biological sample, remains important in the single-cell era because it is typically less noisy, more sensitive, and more cost-effective than single-cell assays. Accordingly, a growing body of computational methods seeks to recover causal relations among genes from bulk expression data. However, aggregation is a lossy, non-invertible coarsening of the underlying cellular system, and it remains unclear whether and under what conditions causal relations are recoverable from aggregated bulk gene expression data. To answer this, we formalize recoverability under aggregation through two notions of consistency: functional-form consistency and conditional-independence consistency. We then derive necessary and sufficient conditions for recoverability, showing that these properties are preserved only under linear aggregations (e.g., sum/mean) coupled with affine structural equations. To assess the practical plausibility of these conditions, analyses of four bulk and four single-cell gene expression datasets further reveal that the estimated pairwise regulatory functions among genes deviate from linearity in both data types, providing limited empirical support for the linearity assumptions required for recoverability. Together, these results caution against recovering causal relations from aggregated bulk expression data without strong additional assumptions.
Abstract: We propose FA-LAM, a focus-aware large avatar model for one-shot animatable Gaussian head creation, while simultaneously enabling static 3D and dynamic 4D full-head recovery. The core of our method lies in a thorough analysis of the attention mechanisms and the entangled reconstruction and animation training pipeline adopted by prior state-of-the-art approaches. Our analysis identifies two main factors that compromise the quality of 3D full-head generation: (1) incorrect and noisy attention activations, and (2) conflicts between the tasks of reconstruction and animation. To address the first issue, we introduce a symmetric and semantic attention regularization strategy that leverages the inherent semantics and structural symmetry of human heads. To disentangle the objectives of reconstruction and animation, we develop a novel dual-phase training pipeline that separates the model's capabilities for large-view hallucination and animation into distinct modules. Moreover, we enhance our model to support multi-view and streaming 4D reconstruction in an efficient and memory-friendly manner through a core autoregressive modification with tailored visibility-aware token fusion. Collectively, these innovations enable FA-LAM to reconstruct animatable Gaussian full heads with superior quality, particularly in fine facial regions and large viewing angles.
Abstract: Embodied AI agents increasingly require parallel execution of multiple tasks, such as manipulation, conversation, and memory construction, from shared observations under distinct time constraints. Recent Mixture-of-Transformers (MoT) Vision-Language-Action Models (VLAs) architecturally support such heterogeneous outputs, yet existing inference systems fail to achieve efficient multi-task parallelism for on-device deployment due to redundant computation and resource contention. We identify isolated KV cache management as the root cause. To address this, we propose unified KV cache management, an inference design that treats KV cache as a first-class shared resource across tasks and over time. This abstraction enables two key optimizations: cross-task KV sharing eliminates redundant prefill of shared observations, while cross-frame continuous batching decouples variable-length language decoding from fixed-rate action generation across control cycles. We implement this design for \pi_0.5, the most popular MoT VLA, and evaluate on both NVIDIA GeForce RTX 4090 and Jetson AGX Thor, two representative platforms for on-device VLA inference. \nickname achieves up to 3.7× speedup over isolated execution, delivering over 200 tokens/s language throughput and 70 Hz action frequency simultaneously without action quality degradation, and we further validate the gains on a real humanoid robot with on-board Jetson AGX Thor.
PaperID: 6131, Poster
Abstract: We study two-layer neural networks trained by stochastic gradient descent (SGD) on a multi-index teacher model, focusing on how first-layer parameters orthogonal to the teacher subspace affect generalization. We decompose the first-layer weights into a signal component aligned with the teacher and an orthogonal complement, and analyze their joint SGD dynamics together with the second layer in the proportional limit. We show that although SGD updates to the complement are initially isotropic, the training dynamics induce teacher-dependent spectral spikes in the complement Gram matrix. This emergent anisotropy provides a concrete mechanism by which representation variance is amplified, and generalization degrades. To isolate this effect, we introduce Decomposed Dynamics, an analytically tractable surrogate whose asymptotic test-error dynamics coincide with those of SGD. Using this surrogate, we characterize regimes in which freezing the orthogonal complement strictly improves generalization relative to training it, despite identical training error. Finally, we show that the signal component rapidly aligns with the second-layer weights, revealing strong cross-layer coupling even in this shallow setting. Together, our results provide a precise dynamical explanation for when and how feature learning outside the teacher subspace harms generalization.
Abstract: Recent think-answer approaches in VLMs, such as Qwen3-VL-Thinking, boost reasoning performance by leveraging intermediate thinking steps before the final answer, but their high computational cost limits real-world deployment. To distill such capabilities into compact think-answer VLMs, a primary objective is to improve the student's ability to utilize visual evidence throughout its reasoning trace. To this end, we introduce a novel think-answer distillation framework that encourages the student to anchor its thinking on visual information by masking the student's salient reasoning prefixes. To compensate for such masked textual cues, the student is encouraged to rely more on visual evidence as an alternative source of information during distillation. Our masking strategies include: 1) token-wise salient reasoning-prefix masking, which masks high-influence reasoning prefixes selectively for each next-token prediction, and 2) self-paced masking budget scheduling, which gradually increases the masking scale according to distillation difficulty, measured by discrepancy between teacher--student distributions. In the distillation phase, the student is guided by our salient reasoning-prefix mask, which blocks both future tokens and salient reasoning cues, in place of the standard causal mask used for auto-regressive language modeling. Experimental results show that our approach outperforms recent open-source VLMs, VLM distillation, and self-distillation methods on multimodal reasoning benchmarks, while further analyses confirm enhanced visual utilization along the student thinking process.
Abstract: Modeling motion for articulated objects of arbitrary skeleton topology remains difficult: existing motion generators target a fixed human skeleton, and prior adaptations either fail to share a vocabulary across rigs or discard motion detail through global pooling. Our key observation is that while joint-level motion does not correspond cleanly across species, motion of functional joint groups does---a human arm, a wolf foreleg, and a bird wing share semantic motion structure despite differences in joint count and connectivity, a correspondence that joint names (e.g., ``forearm'', ``wing\_L1'') partially expose even when topology does not. We introduce SAMoR (Skeleton-Aware Motion Representation for Articulated Objects), a cross-topology motion representation that encodes each motion segment as a small fixed number (K=8) of part tokens shared across arbitrary skeletons. SAMoR takes three input signals---per-joint motion features, kinematic graph structure, and joint-name embeddings---and processes them with a graph-transformer encoder, then compresses the resulting heterogeneous per-joint features into part-level tokens via cross-attention pooling and residual vector quantization, yielding a discrete motion codebook shared across rigs. To prevent the part queries from collapsing into redundant global representations, we introduce a topology-agnostic attention supervision loss, combined with random joint-name dropout to prevent over-reliance on text labels; together these encourage the part tokens to cluster joints into functional groups from names, structure, and motion jointly. We curate a unified heterogeneous motion corpus from HumanML3D, Truebones Zoo, and animated Objaverse-XL assets, and evaluate SAMoR on held-out characters with unseen skeletons. The resulting representation supports accurate reconstruction and cross-topology motion transfer, and further enables text-conditioned generation and localized part-wise editing through a MaskGIT token generator. SAMoR reaches 2.75\!×\!10^-2 normalized MPJPE on cross-topology reconstruction---5.8\!× below the strongest adapted variable-J tokenizer baseline---and remains competitive with fixed-skeleton specialists on HumanML3D for both VQ-VAE reconstruction and text-to-motion generation.
PaperID: 6134, Poster
Abstract: Discrete graph diffusion models corrupt graph structure through independent edge updates with a Markovian noise model. Combined with the exchangeability requirement on the output distribution, the terminal prior is necessarily an Erd\Hos--R\'enyi (E-R) random graph. This has a fundamental consequence for unconditional graph generation without node features: in the large graph limit, any GNN denoiser applied to an E-R graph approximately outputs an E-R graph, making effective denoising difficult. In practice, generated graphs fail to recover meaningful node heterogeneity and higher-order structure. Previous models introduced specialized node features to compensate, framing this as an expressivity problem, while the role of the prior remained unexamined. We instead address the root cause by introducing the \emphLatent Sociability Prior (LSP) and a \emphjoint diffusion process that co-evolves latent node sociability variables and graph structure, removing edge independence while preserving exchangeability. Unlike existing latent graph diffusion methods, the sociability mechanism requires no complex encoder: its sole role is to guide edge updates and preserve node heterogeneity throughout the trajectory. We validate our approach on multiple graph generation datasets and show that the proposed framework more accurately captures graph distributions than baseline methods without node features.
PaperID: 6135, Poster
Abstract: Video world models promise general-purpose interactive simulators, but their training is limited by action supervision: action labels are scarce, fragmented across incompatible specifications, and do not scale with the unlabeled video pretraining corpus. A common workaround jointly trains an action encoder to compress consecutive frames into a latent action and a decoder to reconstruct the next frame from the previous frame and that latent, forcing the latent to capture only new information. This has two limitations: learning the action space from scratch can limit generalization to unseen domains, and the bottleneck trades expressiveness for separation, with tighter settings losing nuance and looser ones risking scene leakage under domain shifts. We argue for a different starting point: rather than learning an action representation from scratch, we adapt a pretrained Multimodal Large Language Model (MLLM), which was trained on large-scale, diverse video-language data and whose embeddings already carry partial action information. This provides a more generalizable foundation for action representation. The remaining challenge is to disentangle action from scene. To this end, we propose a contrastive adaptation framework that reshapes the geometry of the pretrained space without imposing a low-dimensional bottleneck, preserving the expressiveness of the original embeddings. We introduce a unified evaluation protocol that probes action representations along three axes - action expressiveness, scene invariance, and action-conditioned video generation quality - across several robot and game benchmarks in both in-domain and out-of-domain settings. Our representations outperform prior annotation-free latent-action methods on all three axes.
PaperID: 6136, Poster
Abstract: Flow matching models have emerged as a dominant paradigm for generative posterior sampling in imaging inverse problems. Existing solvers follow two distinct lines of work: trajectory-space methods, which refine sampling dynamics from random initial noise, and source-space methods, which optimize the initial noise itself. While prior work treats these as independent design choices, we show that they are fundamentally coupled. We derive a unified reconstruction error bound that links initialization proximity and cumulative sampling stochasticity through a multiplicative envelope, yielding two key insights. First, we identify a sufficiency regime in which a moderately optimized proximal source initialization is enough to attain the optimal error floor, rendering exhaustive source-space optimization unnecessary. Second, this initial advantage is fragile: it is eroded by excessive stochasticity accumulation during guided sampling, necessitating a mechanism that bounds the cumulative noise injection. These findings motivate FLINT, a two-stage framework that operationalizes this coupling. FLINT establishes a high-fidelity starting point via a manifold-constrained proximal objective in only a few iterations, then performs guided sampling under a bounded stochasticity schedule that prevents erasure of the initialization advantage. Across diverse inverse problems, FLINT delivers state of the art perceptual quality and competitive distortion metrics while maintaining superior computational efficiency.
PaperID: 6137, Poster
Abstract: High-content multichannel imaging, from Cell Painting assays to remote sensing, is central to modern scientific pipelines. Because these channels are highly heterogeneous and experimental configurations frequently evolve, a practical vision backbone must be fundamentally channel-agnostic. Current channel-adaptive vision transformers rely on global self-attention with learnable channel embeddings, which can be computed only for previously known channels. Attempts to overcome this limitation of new channels either discard the embeddings, losing channel identity, or process each channel independently, thereby eliminating critical cross-channel interactions. To address this, we introduce ChanSFormer, an efficient Vision Transformer that replaces rigid embeddings with disentangled spatial-channel attention and individual channel CLS tokens. This architecture natively preserves both distinct channel identities and cross-channel information flow. Furthermore, the individual channel class token design unlocks more informative feature representations for self-supervised learning and enables representation-based channel sampling while remaining channel agnostic. We evaluate ChanSFormer on two biology multi-channel datasets, CHAMMI, JUMP-CP and a satellite dataset, So2Sat. Experimental results show that ChanSFormer outperforms state-of-the-art methods by up to 4.27% in classification accuracy. It also demonstrates exceptional cross-dataset transferability under a self-supervised setting, outperforming previous methods by up to 2.72% in accuracy in the same classification tasks. Furthermore, the disentangled attention reduces the quadratic complexity of channel × spatial sequence length of the global attention baseline, thus improving throughput by 55%--313%.
PaperID: 6138, Poster
Abstract: Inverse Reinforcement Learning (IRL) is classically solved using Lagrangian optimization, by jointly optimizing a primal-dual problem with gradient descent to obtain the optimal reward function parameters and corresponding optimal policy. While this algorithmic view of IRL has been highly influential, gradient-based reward updates only paint a partial picture, largely ignoring the geometry of the problem in occupancy space. In this paper, we present an occupancy-space view of reward updates in IRL. We leverage the insight that both the optimal solution and its geometry in occupancy space are known, and show that the geometry along the e-geodesic (straight line in occupancy space) toward the optimal occupancy directly affects the gradient of the IRL dual problem in reward parameter space. This perspective naturally suggests preconditioning the dual gradient with the average Fisher information along the occupancy trajectory. We show that this preconditioned gradient is just a convex interpolation of the old reward and a new target reward, yielding efficient natural-gradient-style updates with little additional computational overhead. We then analyse these occupancy-space reward updates from a Mirror Descent (MD) perspective, and show that policy search over the resulting rewards induces approximate entropic MD on the policy. Finally, we conclude with convergence results of such geometry-aware IRL algorithms.
Abstract: Mirror Descent (MD) extends Gradient Descent (GD) beyond Euclidean geometry and has recently reappeared as a lens for KL-regularized policy optimization in reinforcement learning and LLM post-training. This raises a basic robustness question, crucial to reproducibility and reliability: how sensitively do MD dynamics depend on their inputs? We focus on initialization, often itself a pretrained or previously aligned model. Quadratic-regularized MD, including GD and Mahalanobis geometries, is well-known to be stable for convex smooth objectives. We show a sharp contrast: once the regularizer is non-quadratic, MD can be exponentially more sensitive to initialization than GD, even with a well-conditioned regularizer in Euclidean norm. We give a three-dimensional construction with a convex, smooth objective and a strongly convex, smooth, well-conditioned regularizer where an initial \varepsilon perturbation is quickly amplified to \min\\ \textpolylog^-1(1/\varepsilon), \varepsilon e^\Omega(\eta T) \\ after T iterations of MD with step size \eta. For canonical KL-regularized MD on the simplex, we show that even linear objectives can amplify an initial \varepsilon perturbation exponentially fast in high-dimensional or near-boundary regimes. Finally, we propose Anchored MD, which adds a Bregman term to a fixed point, and show it achieves O(1/\sqrtT) stability, while preserving optimization guarantees up to logarithmic factors.
PaperID: 6140, Poster
Abstract: Long-horizon tool-augmented agents suffer sharp degradation as trajectories grow: small tool errors are stored, repeatedly reused, and amplified into cascading failures. Existing approaches rely on post-hoc verification, item-level memory scoring, or executive context management, but do not directly control how unreliable evidence propagates through reuse. We propose EPOCH (Edge-Pathway Outcome-supervised Cascade-Halting memory), which mitigates cascades by learning an outcome-correlated reliability proxy on reuse pathways connecting observations over time, instantiated as a sparse memory graph with quality-weighted edges, a Temporal Graph Transformer over edge reliabilities, and a learned mid-context correction. We prove that quality-weighted and uniform retrieval are asymptotically separated in horizon length: the ratio of cumulative cascade error diverges unboundedly as T \to \infty whenever the calibration parameter is positive, a result that does not require distributional regularity. We validate this prediction empirically at trajectory lengths up to T=500 and additionally establish a quantitative edge-vs-node separation in terms of an empirically measurable cross-context error variance. Because outcome supervision is correlational by construction, we do not claim the learned scores recover causal pathway reliability; we validate them on two complementary axes (counterfactual edge ablation: \rho=0.71, n=21,743; human annotation: ROC-AUC~=0.94, ECE~=0.041, n=6,000). Across 11 multi-modal benchmarks spanning native long-horizon search, wrapped knowledge-seeking tasks, and clean perception controls, EPOCH consistently outperforms the strongest reliability-aware and memory-augmented baselines in noisy long-horizon regimes while remaining on par on clean perception controls---quality-aware filtering does not collapse on clean data without claiming it helps there.
PaperID: 6141, Poster
Abstract: Online Continual Learning (OCL) requires models to learn from non-stationary data streams under a strict single-pass constraint, making them highly susceptible to catastrophic forgetting. While existing studies have explored various strategies from data, architecture, and optimization perspectives, the optimizer often directly uses generic approaches such as SGD and AdamW. In this work, we reveal a new link between optimization dynamics and OCL needs by recasting Polyak-Ruppert averaging as the engine of "plasticity" and Primal averaging as the anchor of "stability". This inspires our novel optimizer, SPIN ( terpolation), which explicitly decouples plasticity and stability by interpolating between a fast-moving "plasticity" sequence and an adaptive "stability" sequence. We theoretically demonstrate that SPIN implicitly employs an inverse Hessian approximation, providing crucial Tikhonov regularization to damp noisy, single-sample Hessian estimates and mitigate forgetting. More importantly, SPIN can be seamlessly integrated existing OCL methods by taking resulted gradients as input and replacing standard optimizers. Extensive experiments show that using SPIN instead of standard optimizers yields consistent and significant performance gains across multiple OCL approaches, including the challenging rehearsal-free setting and under strict memory constraints.
Abstract: A basic model in sequential decision making is the Markov decision process (MDP), which is extended to Robust MDPs (RMDPs) by allowing uncertainty in transition probabilities and optimizing against the worst-case transition probabilities from the uncertainty sets. The class of (s,a)-rectangular RMDPs with L_p uncertainty sets provides a flexible and expressive model for such problems. We study this class of RMDPs with discounted-sum objectives and a constant discount factor. The existence of an efficient algorithm for this class is a fundamental theoretical question in optimization and sequential decision making. Previous results only establish a strongly polynomial-time algorithm for L_\infty uncertainty sets. In this work, our main results are as follows: (a) we show that for any compact uncertainty set, the policy iteration algorithm for RMDPs is strongly polynomial with oracle access to solutions of Robust Markov chains (RMCs); (b) we present strongly polynomial-time bounds on the policy iteration algorithm for RMCs with L_1 and L_\infty uncertainty sets; and (c) we establish hardness results for RMCs with L_p uncertainty sets for integer p satisfying 1
PaperID: 6143, Poster
Abstract: Reinforcement learning (RL) has become essential for post-training large language models (LLMs) in reasoning tasks. While scaling rollouts can stabilize training and enhance performance, it introduces substantial computational overhead. In algorithms like GRPO, multiple rollouts per prompt incur prohibitive costs, as a large portion of prompts provide negligible gradients and are thus of low utility. This raises a key question: how to identify high-utility prompts before an expensive rollout? Our experimental analysis reveals that sample utility is non-uniform and dynamic: the strongest learning signals concentrate at the ``learning edge'', the intersection of intermediate difficulty and high response entropy, which shifts throughout training. Motivated by this observation, we propose HIVE, a history-informed and online-verified prompt selection framework for data-efficient RL training. HIVE first uses historical reward statistics and response entropy as a cheap prior to filter candidate prompts, and then employs prompt entropy as a real-time proxy to prune instances with stale utility. Across multiple reasoning benchmarks and base models, HIVE maintains reasoning accuracy with substantial rollout cost reduction, achieving up to 2.3× total training speedup.
PaperID: 6144, Poster
Abstract: Guard models are the last line of defense between a language model and a harmful output, yet their training objective is surprisingly narrow. Existing guards learn to predict a single verdict token from a conversational context, concentrating supervision on a single target. The consequences are structural: models latch onto shortcut features, are overconfident, and remain sensitive to where safety evidence appears in the sequence rather than its role in the full context. We propose a different framing. Rather than predicting a label from text, our LLaDA Guard asks which label better explains the text: scoring the prompt or response under each label hypothesis and classifying based on their difference. This shifts supervision to every token in the moderated region, forcing the model to account for full content rather than its most discriminative fragments. We instantiate this idea with a masked diffusion language model, fine-tuning LLaDA-8B-Instruct with a class-conditional reconstruction objective using LoRA and requiring no architectural changes beyond the base model. LLaDA Guard leads on average rank against discriminative baselines trained on stronger backbones across seven held-out safety benchmarks, while exhibiting substantially better confidence calibration (ECE 0.0875 vs. 0.1384 for Qwen3Guard), less over-defense on benign prompts with \textttunsafe-looking cues, and less prompt leakage when moderating responses. Its generative nature further enables token-level risk localization as a natural byproduct, yielding a pipeline for rewriting \textttunsafe prompts into \textttsafe equivalents without additional training and achieving a 60.7% average conversion-to-\textttsafe rate. The weights and code for reproduction are available at the following [anonymized repository](https://anonymous.4open.science/r/LLaDA-Guard2026-B056/).
Abstract: Post-training quantization (PTQ) of diffusion transformers (DiTs) to W4A4 severely degrades output quality, because activations entering each linear layer contain outliers that 4-bit formats cannot represent. The standard fix applies an invertible linear transform to the activations and its inverse to the weights before quantizing both. Normalization layers between blocks force this transform to run online at every denoising step, making its inference computation cost the binding design constraint. Existing options trade quantization quality for inference cost: per-channel scaling (SmoothQuant) is computationally cheap but impacts the magnitude of the channels, which can harm quantization accuracy; fixed Hadamard transforms yield better quantization accuracy but require large block sizes that incur a high online cost; learned full-d invertible transforms calibrate best but entail an prohibitive dense d × d matrix multiplication (GEMM) per layer per step. We propose KroQuant, a PTQ method that applies a learned Kronecker-structured invertible transform to each 32-element block of the activation, storing less than half the parameters of per-channel scaling. The block-local structure runs as small tensor-core GEMMs, and on an MI350 GPU the KroQuant quantizer kernel is up to 14% faster than the SmoothQuant kernel. Offline LoRaQ weight calibration then absorbs the residual per-weight quantization error. On PixArt-\Sigma, SANA, and FLUX.1-schnell at W4A4 (MXFP4e2), KroQuant produces outputs closer to the FP reference than SVDQuant and LoRaQ on MJHQ-30K and SDCI, while preserving or improving image quality.
PaperID: 6146, Poster
Authors:
Siming Zhang, Zhehui Shen, Shijie Chen, Yansen Yu, Hang YuAbstract: Low-label transfer is often decided by a practical question: after a representation looks useful, how many labels are still needed to learn the downstream boundary? Existing transferability and probe-quality scores rank global alignment, likelihood, margins, or calibration, but they do not measure how much target data remains close to the fitted decision boundary, where scarce labels are most costly. We propose , a label-budget diagnostic for frozen representations and linear probes. It measures mass in a thin strip around the learned boundary and normalizes it by class balance, so dense ambiguous regions and rare-class difficulty are both counted. We prove that, under a smooth representation map, this boundary mass does not change because of generic volume expansion: density change and tangential surface change cancel at first order, leaving only stretch or compression in the boundary-normal direction. This yields a cached-feature estimator that fits a probe on one split, counts held-out features near its boundary, and avoids high-dimensional density estimation. Experiments support the diagnostic at three levels. Controlled transformations match the predicted law across 74 maps with 0.38% median relative error. On real frozen features, normal-direction changes move both measured boundary mass and label demand, while tangent changes do not. A differentiable surrogate reduces measured complexity by 16–36% at matched top-1 utility. Finally, CIFAR-100 and Tiny-ImageNet audits show boundary mass acting as a second-stage diagnostic for near-tied representations, with external multiclass results showing how label-grid saturation limits policy-level gains on label grids where many tasks tie.
PaperID: 6147, Poster
Abstract: Long-term conversational memory poses a fundamental system challenge for LLM agents: reasoning over all past interactions improves coverage but incurs prohibitive token costs and noise, while retrieval-based methods are efficient yet often fail on multi-hop and temporal questions. Existing multi-step retrievers partially address this gap, but typically operate over an ever-growing textual history, causing context expansion and noise accumulation across iterations. We propose MemR^3, a backend-agnostic closed-loop controller that reformulates conversational memory retrieval as explicit stateful decision-making. At each iteration, MemR^3 maintains an evidence-gap state consisting of grounded information (evidence) and unresolved information requirements (gaps), and uses this state to route among three actions: retrieve, reflect, and answer. Newly retrieved snippets are incorporated into the state and then masked from subsequent prompts, so the model conditions on a compact summary of progress rather than the full retrieval history. This design turns multi-step memory retrieval into a bounded, inspectable control process whose query-time context scales with the state instead of raw accumulated text. Experiments on LongMemEval_s show that MemR^3 surpasses implicit multi-step retrieval baselines and often approaches or exceeds full-context prompting while using only 1%-5% of the full-context tokens on long conversations. On LoCoMo, MemR^3 consistently improves over its underlying RAG- and Zep-based memory backbones. These results suggest that explicit evidence-gap tracking is an effective abstraction for token-efficient conversational memory.
PaperID: 6148, Poster
Abstract: Vision-and-Language Navigation(VLN) requires agents to sustain task-directed spatial perception to focus on navigation-critical regions while translating linguistic intent into low-level actions. Prevalent hierarchical planning methods rely on a high-level policy to select waypoints or subgoals for low-level control, but they often fail to remain focused, causing inaccurate or redundant subgoals that directly mislead execution. Moreover, their separation between visual perception and action execution neglects fine-grained spatial reasoning in low-level movements and limits real-world deployability. To address these two issues, we propose FocusNav, a grounding-driven end-to-end monocular framework that learns task-directed spatial perception through action-aware future reconstruction. Concretely, FocusNav first boosts the agent’s visual attention on task-relevant regions with a Grounding-driven Visual Reconstruction module that reconstructs instruction-relevant targets and traversable areas via a semantic-guided latent diffusion process. Subsequently, to bridge the perception–execution gap, we incorporate physical action priors via Action-aware Visual Conditioning and enforce Intent-Effect Alignment between anticipated visual changes and observed action effects. These effectively enhance the agent’s task-directed spatial perception to direct attention to navigation-critical regions, and tighten the perception–execution loop. Experiments on R2R-CE and RxR-CE benchmarks show that FocusNav achieves substantially improved state-of-the-art performance with a simple monocular low-level policy, while providing faster inference than recent methods. Furthermore, real-world robot experiments validate our FocusNav’ s effectiveness, ease of deployment, and lightweight design.
PaperID: 6149, Poster
Abstract: Mathematical reasoning benchmarks are increasingly limited by static test sets, which are vulnerable to contamination and slow to adapt as models improve, while manually building fresh high-quality problems is expensive. We propose VDE (Verifiable Dynamic Evaluation), a dynamic evaluation framework that generates replay-verifiable math problems from typed value--theorem bipartite graphs. Each instance is an executable graph with model-independent ground truth recovered by deterministic replay rather than model-produced solutions. Beyond compositional depth, VDE supports explicit constraint injection and controlled branching to produce harder problems while preserving verifiability. We instantiate the same framework in analytic geometry and number theory, showing that a unified graph-based generation paradigm can cover structurally different domains. Experiments on frontier models show that performance remains far from saturation and degrades predictably as we increase construction depth, constraints, and branching, making VDE a scalable and trustworthy testbed for dynamic mathematical reasoning evaluation.
PaperID: 6150, Poster
Abstract: Self-distillation is typically studied when the student is retrained on the teacher’s original training inputs. In many realistic deployments, however, the labeled data are unavailable after training, and one only has access to the trained predictor and fresh unlabeled covariates. We study distillation in this prediction-only regime through a fresh-X prediction-mixing scheme: a pure-distilled student is trained on pseudo-labeled fresh features, and the final predictor is formed by an affine combination of the teacher and student predictions with the optimal unconstrained mixing weight. For ridge regression under proportional asymptotics, we derive deterministic equivalents for the optimally mixed risk under general anisotropic covariance and deterministic signal, and show that optimal mixing strictly improves upon the teacher for almost every pair of regularization levels. In the isotropic case, we uncover a sharp and counterintuitive phenomenon: the optimally mixed risk is generically non-monotone in the amount of unlabeled data. To tune the optimal mixing weight, we show that a small independent labeled calibration set suffices for consistent one-shot estimation at low computational cost, since no additional retraining is required, whereas such tuning is impossible when using only teacher’s outputs on unlabeled data. Finally, to extend beyond squared loss, we analyze logistic regression for binary classification and show that prediction mixing can improve over both the teacher and the pure-distilled classifier.
PaperID: 6151, Poster
Authors:
Akshata Hegde, Kyler Zook, Yanli Wang, Jianlin ChengAbstract: Inferring gene regulatory networks (GRNs) from multimodal biological data requires integrating regulatory evidence beyond transcript abundance alone. Regulatory evidence can come from RNA-seq, chromatin accessibility, and transcription factor (TF) motif priors, but these sources are difficult to discover, harmonize, and integrate, and are often incomplete across biological contexts. We propose GRNAgent, an automated multi-agent framework for multi-omic GRN inference that coordinates evidence acquisition, quality control, modality integration, model inference, benchmarking, and post hoc validation. At the core of GRNAgent is Transcription Factor–centered Evidence Adaptive Graph Reasoning (TF-EAGER), a transformer-based graph reasoning model built on TF-centered evidence graphs. For each TF, GRNAgent constructs a local candidate graph whose nodes are candidate target genes and whose edges are annotated with available regulatory evidence, including expression association, motif support, and chromatin accessibility. TF-EAGER converts this local evidence graph into typed evidence tokens and scores each TF–target gene edge by staged cross-attention: first over motif and chromatin accessibility signals, then over expression-derived evidence, and finally over all available signals. This TF-centered evidence-graph formulation couples the automated multi-omic evidence construction with structured graph reasoning over heterogeneous evidence, enabling regulatory edge prediction from partially available biological signals. Across 24 datasets and 10 baselines, GRNAgent achieves the best full-matrix benchmark performance among the evaluated methods, yielding 5–20× improvements in area under the precision-recall curve (AUPRC) and early precision (EP) over the best-performing baselines. Under sampled evaluation, it reaches AUPRC 0.801 in leave-one-TF-out validation and AUPRC 0.705 in unseen cell-type inference, with average EP@10 around 0.81 in the blind setting. Finally, GRNAgent includes a grounded literature verification module, which uses a large language model (LLM) to structure retrieved biological evidence for predicted edges without influencing model training.
Abstract: In this article we develop a new method for summarizing a ranking distribution, i.e. a probability distribution on the symmetric group \mathfrakS_n, beyond the classical theory of consensus and Kemeny medians. Based on the notion of local ranking median, we introduce the concept of consensus ranking distribution (\crd), a sparse mixture model of Dirac masses on \mathfrakS_n, in order to approximate a ranking distribution with small distortion from a mass transportation perspective. We prove that by choosing the popular Kendall \tau distance as the cost function, the optimal distortion can be bounded in terms of pairwise probabilities, paving the way for the development of efficient learning methods that do not suffer from the lack of vector space structure on \mathfrakS_n. In particular, we propose a top-down tree-structured statistical algorithm that allows for the progressive refinement of a CRD based on ranking data, from the Dirac mass at a Kemeny median at the root of the tree to the empirical ranking data distribution itself at the end of the tree's exhaustive growth. In addition to the theoretical arguments developed, the relevance of the algorithm is empirically supported by various numerical experiments.
Abstract: Inferring the evolution of high-dimensional and multi-modal (e.g., spatio-temporal) physical fields from irregular sparse measurements in real time is a fundamental challenge in science and engineering. Existing approaches, including diffusion-based generative models and functional tensor methods, typically operate in offline settings, depend on full temporal observations, or incur substantial inference cost. We propose StreamPhy, an end-to-end framework that enables efficient and accurate streaming inference of full-field physical dynamics from incoming irregular sparse measurements. The framework integrates a data-adaptive observation encoder that is robust to arbitrary observation patterns, a structured state-space model that supports memory-efficient online updates across irregular time intervals, and an expressive Functional Tensor Feature-wise Linear Modulation (FT-FiLM) decoder for continuous-field generation. We prove that FT-FiLM is more expressive than the functional Tucker model, admitting a richer function class for handling complex dynamics. Experiments on three representative physical systems under challenging sampling patterns show that StreamPhy consistently outperforms state-of-the-art baselines, with at least 48% improvement in accuracy and up to 20--100x faster inference than diffusion-based methods.
PaperID: 6154, Poster
Abstract: Stroke-based rendering is often optimized as a pixel-level reconstruction task. This over-reliance on global error reduction not only degrades strokes into fragmented curves due to a lack of structured planning, but also allows locally defective primitives to be concealed by overlapping layers. To address these issues, we introduce a model-guided framework for stroke-audited rendering. Our framework uses Quadratic Bezier Splatting to analytically model continuous curved strokes with efficient differentiable alpha compositing. Built on this renderer, a feedforward multi-scale stroke decoder provides a structural initialization with learned stroke-scale allocation, while test-time stroke optimization supplies the reconstruction capacity needed for high-fidelity refinement. We further propose a single-stroke audit mechanism that explicitly evaluates local stroke quality, preventing defective primitives from surviving optimization through layer-wise compensation. Experiments demonstrate our approach achieves highly coherent layouts, improving PSNR by up to 20% and yielding a 58.8× renderer speedup with a 16% more compact representation compared to representative SBR baselines.
PaperID: 6155, Poster
Abstract: Transformers rely on a growing key–value (KV) cache to store context, causing linear memory growth and limited generalization beyond the training window. We show that contextual influence can instead be viewed as updates to a low-dimensional subspace in the model’s weight space, meaning context can be represented as a bounded, low-rank parameter update rather than token-level storage. Motivated by this, we propose PARADE, which internalizes context as parametric memory. It compresses past inputs into a fixed set of weight-space basis vectors and composes them via query-dependent routing into per-query weight updates, enabling streaming inference with constant memory and effectively unbounded context. Experiments show that PARADE matches full attention on long-context tasks and outperforms efficient baselines, especially when relevant information lies far beyond the attention window, while improving as context exceeds the model’s training length.
PaperID: 6156, Poster
Authors:
Daniel Shao, Dongmin Bang, Luca Weishaupt, Nic G Reitsam, Sophia J Wagner, Ming Yang Lu, Long P Le, Richard Chen, Faisal MahmoodAbstract: Contrastive vision-language (VL) pretraining in computational pathology has generally relied on curated image-caption datasets sourced from social media, educational videos, and research publications. These corpora are sparse relative to the diversity of tissue morphology and difficult to scale further, as pathologists do not routinely produce descriptive captions in clinical practice. We propose to instead leverage slide-level pathology reports which can be generated during routine diagnostic workflows in a scalable manner. A key challenge in this weakly-supervised setting is that pathology reports describe macroscopic, diagnostic findings while patch-level visual features describe microscopic findings. We introduce Locked Image Multiple Instance PreTraining (LIMIT), which addresses this gap by freezing a pretrained vision encoder and fine-tuning only the text encoder: report embeddings serve as cross-attention queries over bags of patch features, producing report-conditioned slide representations optimized via cross-entropy. Using LIMIT, we establish CONCH-Z, which aligns the text encoder of CONCH v1.5 to 335,645 clinical reports. Evaluated across 20 tasks spanning patch-level classification, tumor detection, and slide-level subtyping, CONCH-Z establishes state-of-the-art zero-shot performance over VL encoders trained on curated caption datasets, while simultaneously closing the gap with slide-level foundation models trained with substantially more parameters and compute. We will release CONCH-Z weights and evaluation code to support reproducibility and broader community use.
PaperID: 6157, Poster
Abstract: Reinforcement learning (RL) has substantially improved the reasoning capabilities of vision-language models (VLMs), yet it often triggers chain-of-thought (CoT) obfuscation, a regime where models achieve high accuracy while producing reasoning traces that are ungrounded and difficult to monitor. While prior work documents this decay at the behavioral level, the underlying mechanistic drivers of such representational drift remain poorly understood. In this work, we identify that obfuscation is a representational pathology, where RL-induced optimization causes task-agnostic template features to replace visually grounded content in the model’s activation space. Based on this insight, we propose TAME, an activation-level intervention framework that leverages Sparse Autoencoders (SAEs) to regularize VLM internal features during RL training. Specifically, TAME utilizes an LLM-as-judge monitor to detect behavioral obfuscation and maps these signals to specific internal directions, penalizing template intrusion while preserving visual evidence. By directly intervening on the features responsible for obfuscation, our method restores the transparency and visual grounding of reasoning traces. Experiments on the VIRL-39k and SPA-VL benchmarks across two model families show that TAME improves CoT monitorability (by up to +30.9% on VIRL and +16.7% on SPA-VL over GRPO), while preserving task accuracy and general-capability performance across four standard VLM benchmarks.
PaperID: 6158, Poster
Abstract: Vision-Language-Action (VLA) models are becoming an important approach for robotic manipulation, but their long visual token sequences make inference expensive. Visual token pruning is a practical way to reduce this cost, with many pruning methods rely on attention scores to decide which tokens to remove. However, we show that this reliance can be exploited: an attacker can hijack attention scores and turn the visual token pruner into a vulnerability. Based on this insight, we propose FocusVLA, a backdoor attack against attention-based pruning. FocusVLA inserts trigger-controlled attention patterns into a few low-impact attention heads. These poisoned heads have little effect on normal manipulation. Without pruning, the poisoned model still behaves normally when the trigger appears. With pruning enabled, clean performance remains intact. But inputs with the trigger shift attention to task-irrelevant peripheral regions. As a result, the pruner removes task-critical tokens, leading to task failure. We evaluate FocusVLA on OpenVLA-OFT and \pi_0.5 across LIBERO tasks with ten attention-based pruning methods. For example, on OpenVLA-OFT with ADP pruning, the success rate on LIBERO-10 drops from 90.6% to 31.0% under triggered inputs. These results reveal a broader risk: attention is widely used as an importance signal, but it can be manipulated and should not be blindly trusted.
Authors:
Xiangfei Qiu, Liu Yang, Xiangyu Xu, Hanyin Cheng, Xingjian Wu, Rongjia Wu, Zhang Zhigang, Tu ding, Chenjuan Guo, Bin Yang, Christian S Jensen, Jilin HuAbstract: Time series forecasting occurs in a range of financial applications providing essential decision-making support to investors, regulatory institutions, and analysts. Unlike multivariate time series from other domains, stock time series exhibit industry correlation. Exploiting this kind of correlation can improve forecasting accuracy. However, existing methods based on hypergraphs can only capture industry correlation relatively superficially. These methods face two key limitations: they do not fully consider inter-industry lead-lag interactions, and they do not model multi-scale information within and among industries. This study proposes the Hermes framework for stock time series forecasting that aims to improve the exploitation of industry correlation by addressing these limitations. The framework integrates moving aggregation and multi-scale fusion modules in a hypergraph network. Specifically, to more flexibly capture the lead-lag relationships among industries, Hermes proposes a hyperedge-based moving aggregation module. This module incorporates a sliding window and utilizes dynamic temporal aggregation operations to consider lead-lag dependencies among industries. Additionally, to effectively model multi-scale information, Hermes employs cross-scale, edge-to-edge message passing to integrate information from different scales while maintaining the consistency of each scale. Experimental results on multiple real-world stock datasets show that Hermes outperforms existing state-of-the-art methods.
PaperID: 6160, Poster
Abstract: Many real-world problems demand exact constraint satisfaction to be meaningfully solved, yet existing neurosymbolic methods face a stark trade-off: those that guarantee satisfaction fail to scale, while those that scale must resort to approximation. Meanwhile, traditional deep learning approaches offer no satisfaction guarantees at all. In this paper, we propose Casper, a neurosymbolic layer that simultaneously (i) guarantees constraint satisfaction, (ii) scales at inference time, (iii) is fully automated, requiring no problem-specific engineering, and (iv) is fully differentiable, making it composable with any neural predictor both at training and at inference time. Casper casts constraint satisfaction as a closed-form Euclidean projection onto the constraint-feasible region, computable in a single forward pass. Experiments across MNIST arithmetic (up to 1024 digits), Sudoku solving, and Warcraft pathfinding show that Casper is the only method that produces guaranteed-valid outputs at scale: on 30×30 Warcraft grids, exact baselines time out and approximate ones produce valid paths less than 7% of the time, while Casper produces them 100% of the time.
PaperID: 6161, Poster
Abstract: Parameter-Efficient Fine-Tuning (PEFT) has become a dominant paradigm for adapting Multimodal Large Language Models (MLLMs) to specific tasks, resulting in a proliferation of task-specific expert models. Merging these experts into one universal model offers a promising path to creating versatile systems without costly retraining. However, existing methods for full fine-tuning merging often falter in parameter-efficient model merging, as they manipulate weights directly while ignoring the intrinsic geometric structure of PEFT modules. When diverse tasks exhibit conflicting feature directions, this inevitably leads to destructive interference. To address this issue, we propose PRISM, a data-free framework that reframes the merging of PEFT modules as a spectral signal reconstruction problem. PRISM operates in two stages: (1) Spectral Pruning, which decouples task updates into singular components and retains only high-energy directions to attenuate task-irrelevant noise; and (2) Energy-Prioritized Orthogonal Reconstruction, which prioritizes dominant components and projects overlapping vectors onto orthogonal directions to eliminate inter-task interference. We curate a benchmark comprising diverse multimodal tasks to evaluate our method. Extensive experiments demonstrate that PRISM significantly outperforms state-of-the-art baselines. Our code and weights will be released.
Abstract: Multi-agent systems (MAS) powered by large language models (LLMs) increasingly adopt planner--executor architectures, where planners convert prompts into subtasks, roles, dependencies, and routing paths. This flexibility enables adaptive coordination, but exposes an attack surface in workflow formation: prompts can shape agent organization without modifying MAS infrastructure. We study this risk through social influence probing workflows to identify high-impact subtasks and malicious-signal propagation. The analysis reveals two vulnerabilities: workflow position can amplify or suppress a malicious signal, and sycophantic framing makes downstream agents more likely to relay it. We translate these findings into FlowSteer, a prompt-only workflow steering attack that converts vulnerability priors into one crafted prompt. FlowSteer aligns a malicious signal with influential task components and guides replanning toward dependencies that preserve propagation. Experiments show that FlowSteer increases malicious success by up to 55% over naive prompting, transfers across MAS setups, and remains effective with black-box topology inference. As FlowSteer biases the planning signals that generate the workflow, MAS defenses that inspect only the generated workflow provide limited protection. As such, we introduce FlowGuard, an input-side defense that reduces malicious success by up to 34% while preserving prompt utility. Our results position workflow formation as a new safety frontier for multi-agent LLM systems, opening a planning-time security perspective on how agent coordination itself can be attacked and defended.
PaperID: 6163, Poster
Abstract: Inference-Time Scaling (ITS) has largely succeeded in verifiable domains like math and coding, where cheap verification enables scalable output selection. However, extending ITS to tasks prone to systematic failure - driven by faulty initial assumptions or unmet multidimensional constraints - typically relies on costly external solvers or brittle, model-based verifiers. Our key insight is that the intrinsic statistics of parallel sample sets, specifically length-adjusted tail entropy, provide a robust discriminative signal for solution quality without access to ground truth. Crucially, these statistics serve as a difficulty gate for adaptive compute allocation, dynamically routing problems across scaling regimes. First, Intrinsic Selection (iS) ranks candidates post-hoc, matching consensus-based algorithms across three domains and improving engineering design selection by 20 % over pass@1 baselines. Second, Intrinsic Particle Filtering (iPF) generalizes this to step-level resampling, guiding generation toward high-confidence reasoning trajectories to improve pass@1 by 6.1 points on average on hard math problems. Finally, Particle Distillation (dPF) injects privileged guidance via early logit blending and KL-guided resampling, steering generation past systematic reasoning errors to satisfy expert rubrics, yielding up to 26.5 % gains on complex clinical responses. Our pipeline applies seamlessly across broad-purpose, domain-specialized, and multimodal architectures, successfully extending ITS to open-ended domains without requiring trained reward models or exact verification.
PaperID: 6164, Poster
Abstract: Training transformer-based architectures with data augmentation has become an increasingly popular approach for equivariant machine learning. Despite its empirical success, the interplay between the transformer architecture, invariance to different symmetries, and augmentation budgets remains underexplored. In this paper, we investigate the ability of a vanilla transformer to learn a wide range of symmetries through limited data augmentation. We identify that it performs strongly for common angle-preserving symmetries, while it struggles with non-angle-preserving ones. Within the angle-preserving family, we further find that base subgroups such as translation, rotation, and scale are learned more easily and accurately than compositional groups that combine them. Motivated by this observation, we perform a structural analysis of the trained models and identify interpretable mechanisms that induce invariance to each base angle-preserving symmetry. These findings support what we coin winning the : the transformer architecture is aligned with learning mechanisms invariant to symmetries that happen to be prevalent in scientific domains.
PaperID: 6165, Poster
Abstract: Multimodal diffusion large language models (dLLMs) generate text through iterative masked denoising with bidirectional attention, offering a compelling alternative to autoregressive vision-language modeling. However, each denoising step requires a full forward pass over long multimodal sequences in which visual tokens dominate the sequence length, and a fixed step budget further compounds this quadratic per-step cost. Existing inference acceleration approaches for multimodal dLLMs characterize attention redundancy implicitly, motivating heuristic designs from isolated attention visualizations rather than principled analysis. We close this gap with DRAMA, a training-free acceleration framework grounded in a systematic dissection of attention redundancy in multimodal dLLMs. By probing attention tensors along spatial and temporal axes, we reveal two persistent redundancy structures: (i) attention mass consistently concentrates on a sparse subset of visual tokens, and (ii) attention distributions rapidly stabilize across denoising steps after an initial formation phase. These findings motivate two complementary mechanisms: Spatial-Redundancy-Aware Visual Token Pruning (SVTP) and Sufficiency-Guided Adaptive Stopping (SGAS). Together, DRAMA achieves an average 18.67× inference speedup across six vision-language benchmarks without training, architectural modification, or attention caching, while maintaining competitive accuracy. Our results establish a principled and strong acceleration baseline for multimodal dLLM inference.
PaperID: 6166, Poster
Abstract: Entropy minimization (EM) is the dominant objective for test-time adaptation, yet its failure mode, model collapse, remains poorly understood. In this work, we show that distribution shifts can cause feature clusters corresponding to distinct classes in the model’s representation space to merge, while the decision boundary remains fixed. This induces a systematic skew in the predicted class distribution, referred to as prediction bias. Prediction bias refers to a shift in the predicted class distribution, with some classes overrepresented and others suppressed. We show that entropy minimization amplifies this prediction bias by tightening the existing clusters, reinforcing the incorrect groupings until all predictions collapse to a trivial solution. Next, to demonstrate the significance of prediction bias and mitigate it, we further propose Distribution Shift Bias Reduction (DSBR), a bias-correcting objective that specifically targets this failure mode by equalizing the contribution of each predicted class to the unsupervised entropy minimization loss. To study this failure mode, we design suitable adaptation settings using four medical‑imaging datasets and additionally evaluate on ImageNet‑C. We find that DSBR consistently stabilizes test‑time adaptation, prevents model collapse, and matches or outperforms the state‑of‑the‑art methods. Moreover, DSBR operates solely at test-time.
PaperID: 6167, Poster
Abstract: The manifold hypothesis suggests that real-world data concentrates on low-dimensional structures, yet navigating meaningful trajectories on latent manifolds remains an open challenge. A common approach is to construct density-aware metrics, but such constructions bundle two distinct questions: (i) where valid data lies and (ii) how we choose to traverse on the manifold. Since different tasks may demand different traversal behaviors even on the same data manifold, specifying these two aspects independently offers additional modeling flexibility. We propose to factor the path objective into a kinetic term governed by a freely chosen geometric prior and a score-based potential, derived from the negative log-density of a pretrained diffusion model, that acts as a soft constraint for manifold adherence. The potential encourages dynamics to evolve along the data support, while the geometric prior remains free to encode task-specific notions of path optimality. To solve the resulting optimization problem, we introduce geodesic force matching (GFM), a direct-collocation algorithm that discretizes the trajectory and optimizes all waypoints jointly to satisfy the Euler--Lagrange force balance in a least-squares sense. On synthetic manifolds and 3D shape interpolation, our method produces smooth, on-manifold paths and performs favorably against recent density-aware baselines, most notably in the low-noise regime where those baselines become numerically unstable.
PaperID: 6168, Poster
Abstract: Large language models are increasingly deployed as LLM-as-a-Service (LLMaaS), where user queries are sent to remote LLMs for inference, raising severe privacy concerns about the exposure of sensitive user inputs. Differentially private inference mitigates these risks by injecting noise into input embeddings. However, existing methods struggle to balance privacy and utility, as stronger privacy protection often leads to substantial accuracy degradation in LLM inference. In this work, we propose DP-SPD, a differentially private LLM inference framework that protects user queries by hiding the real query among client-generated dummy queries. All queries are randomly shuffled before being sent to the server, preventing direct identification of the user's true intent. To improve inference efficiency, we exploit the shared-prefix structure among queries and reuse transformer key--value (KV) caches during decoding. We analyze the privacy guarantee of DP-SPD and evaluate its utility on classification and generation tasks, as well as its privacy protection under EIA, MIA, and GPT-based attacks. Experimental results show that our method achieves strong inference-time privacy protection while maintaining utility close to plain-text inference under practical latency.
PaperID: 6169, Poster
Authors: Chen Wang
Abstract: Retrieval-augmented in-context learning faces a budget problem under distribution shift: increasing top-k can improve evidence coverage, but it can also add off-direction, conflicting, or overly long context that a finite reader cannot use. We introduce a theory-guided diagnostic framework based on \empheffective evidence coverage: retrieval is useful only while aligned evidence grows faster than reader-facing ambiguity. In a local linear-ICL model, coupled covariate and task-prior shifts create a mixed risk term that aligned retrieval can reduce up to a residual ambiguity term; a conditional transfer result extends the same budget boundary to frozen readers satisfying a prompt-level stability condition. This yields a falsifiable prediction: recall can keep rising after answer accuracy has saturated or declined. Across QA, verification, and NLI tasks, dense top-k sweeps together with oracle and shuffled-context controls expose this recall--accuracy decoupling. The resulting framework separates coverage- and saturation-limited QA from interference-limited verification/NLI, while calibrated selected-context controllers are reported only as operational probes of the diagnosis.
PaperID: 6170, Poster
Authors: Luca Sacchetto, Klaus Diepold
Abstract: Diffusion models have emerged as a powerful tool across diverse domains. However, their purely data-driven nature can produce samples that violate domain-governing constraints. We introduce a plug-and-play Reinforcement Learning framework that optimizes initial noise samples in the latent space of frozen, pre-trained diffusion models. Leveraging the near-spherical geometry of high-dimensional Gaussian distributions, we introduce a novel rotation-matrix-based scheme for efficient latent space exploration. This steers the model toward more feature-preserving outputs, guided by task-specific rewards. We evaluate our method on three diffusion models: one trained on solutions of the Darcy Flow PDE, one on a synthetic dataset with complex structural features, and a text-conditioned one. Across all three settings, our framework yields significant improvements in sample quality, achieving a ~25% relative reduction in PDE residual, up to a ~44% relative improvement on the synthetic dataset's feature-alignment metric, and up to a ~80% relative improvement on human preference, compared to the vanilla diffusion models. Finally, we show that rotation-matrix-based exploration significantly outperforms unconstrained exploration, validating our geometry-aware approach and establishing a more effective method for latent space control.
Abstract: Continual learning requires models to adapt to new data while preserving previously acquired knowledge. At its core, this challenge can be viewed as principled one-step adaptation: incorporating new information with minimal interference to existing representations. Most existing approaches address this challenge by modifying model parameters or architectures in a supervised, task-specific manner. However, the underlying issue is representational: tasks require distinct yet structured representations that can be selectively updated without disrupting representations, while structure should reflect intrinsic organization in the data rather than task boundaries. In sequential data, time-delayed dependencies provide a natural signal for uncovering this organization, revealing how fundamental representations give rise to more specific ones. Inspired by the modular organization of the human brain, we propose MoRe, a framework that identifies modularity in the representation itself rather than allocating it at the architectural level. MoRe decomposes knowledge into a hierarchy of fundamental and specific modules with identifiability guarantees, enabling principled module reuse, alignment, and expansion during adaptation while preserving old modules by construction. Experiments on synthetic benchmarks and real-world LLM activations demonstrate interpretable hierarchical structure, improved plasticity-stability trade-offs, suggesting MoRe as a principled foundation for continual adaptation.
PaperID: 6172, Poster
Abstract: Recent advances in video generation have achieved impressive visual quality, yet they often fail to produce physically consistent dynamics due to the lack of explicit modeling of underlying physical processes. Existing approaches either rely on simulators with strong assumptions or use coarse semantic guidance, and further struggle to effectively inject physical priors into generation models. In this work, we propose Seeing the Unseen, a unified framework for physics-grounded video generation. Our approach decomposes the problem into two stages: learning physical dynamics and injecting them into video synthesis. First, we introduce a lightweight Particle-based Graph Dynamics Simulator (PGDS) that learns generalizable physical interactions from data and predicts plausible 3D motion trajectories. Second, we propose a Visible–Invisible Motion Field (VIMF) that captures both motion observable in the input frame and motion that emerges over time due to object dynamics. This representation enables more complete and structured motion guidance compared to conventional physical signals. By integrating these trajectories into a diffusion-based generator, our method produces videos that are both visually coherent and physically consistent. Extensive experiments demonstrate improved motion accuracy, interaction consistency, and reduced physical artifacts compared to strong baselines.
Abstract: Adaptive prompt and program search makes LLM evaluation selection-sensitive. Once benchmark items are reused inside tuning, the observed winner's score need not estimate the fresh-data performance of the full tune-then-deploy procedure. We study inference for this procedure-level target under explicit tuning budgets. We propose SIREN, a selection-aware repeated-split reporting protocol that freezes the post-search shortlist, separates splitwise selection from held-out evaluation, and uses an item-level Gaussian multiplier bootstrap for uncertainty quantification. In a fixed-shortlist regime with smooth stabilized selection, the estimator admits a first-order item-level representation, and the bootstrap yields valid simultaneous inference on a finite budget grid. This supports confidence intervals for procedure-performance curves and pre-specified equal-budget and cross-budget comparisons. Controlled simulations and MMLU-Pro tuning experiments show that winner-based reporting can be optimistic and can change deployment conclusions, while SIREN remains close to the finite-sample reporting target.
PaperID: 6174, Poster
Abstract: Long-term memory enables personalized conversational agents to retain user information across sessions. However, existing memory architectures primarily optimize for utility but neglect the risks of storing and reusing private attributes such as personally identifiable information (PII) unnecessarily. Dealing with privacy risk in personalized memory is challenging as simply removing sensitive values would undermine the utility of the memory system. Therefore, privacy protection for memory agents must govern the full life-cycle of sensitive values rather than just sanitizing individual records. To fill this research gap, we introduce Sanitized Privacy-Mapped Memory (SP-Mem), a privacy-aware memory architecture that decouples memory utility from exact private-value exposure. SP-Mem provides full life-cycle privacy-related design including determining how to identify and separate sensitive information from raw user inputs, how to store sanitized content and exact private values in isolated structures, and how to selectively retrieve values based on the task requirement and user consent. We further introduce a privacy-aware memory benchmark that jointly assesses response quality, privacy behavior, and inference cost. Extensive experiments across multiple LLM-based agents show that SP-Mem achieves stronger personalization while reducing unnecessary privacy exposure. Code and data are available at https://anonymous.4open.science/r/SP-Mem-0CAE/.
Authors: Peter Richtarik, Yassine Maziane, Ammar Mahran, Artavazd Maranjyan
Abstract: Communication is a major bottleneck in distributed learning, especially in large-scale settings and in federated learning environments with slow links. Three standard ways to reduce this cost are communication compression, local training, and communication-computation overlap. Methods that combine these ingredients are used in practice and have been found to be effective for large-scale training, but there is little theory for methods that combine all three. We study a heterogeneous-compute setting in which different workers may take different numbers of local steps, and we propose LOSCAR-SGD, a Local SGD method that communicates only a sparse subset of model coordinates and continues optimizing while communication is in flight. A key ingredient is a delay-corrected merge rule that incorporates delayed synchronized information without discarding the progress made during the overlap phase. We give convergence guarantees for smooth non-convex objectives and show how sparsity, overlap, and worker heterogeneity affect the rate. To the best of our knowledge, this is the first theory for this combination of ingredients. Experiments further show that communication-computation overlap reduces training time and that the delay-corrected merge outperforms naive overwriting.
PaperID: 6176, Poster
Authors:
Zhenxin Li, Nadine Chang, Xinglong Sun, Jingde Chen, Wenhao Yao, Zi Wang, Maying Shen, Yu-Gang Jiang, Zuxuan Wu, Shiyi Lan, Jose M. AlvarezAbstract: Behavior policies are often formulated as continuous generative models, whose iterative denoising processes are expressive but difficult to interpret and prone to producing implausible actions. We propose the Large Discrete Policy (LDiP), a fully discrete behavior modeling framework that selects actions from a large vocabulary of physically plausible candidates. Rather than perturbing actions, LDiP improves expressivity through stochastic iterative scoring: it progressively re-scores and prunes candidates with score-space stochasticity, enabling fine-grained ranking and exploration among plausible actions while preserving an explicit decision process. Across end-to-end planning, closed-loop driving, robotic manipulation, and vision-language-action settings, LDiP consistently outperforms strong discrete and continuous baselines in autonomous driving, and exceeds or matches continuous generative policies in robotic manipulation. These results show that discrete policies, when equipped with effective scoring mechanisms, offer an expressive, plausible, and interpretable alternative for behavior modeling.
Abstract: Empirical studies of trained models often report a transient regime in which signal is detectable in a finite gradient descent time window before overfitting dominates. We provide an analytically tractable random-matrix model that reproduces this phenomenon for gradient flow in a linear teacher--student setting. In this framework, learning occurs when an isolated eigenvalue separates from a noisy bulk, before eventually disappearing in the overfitting regime. The key ingredient is anisotropy in the input covariance, which induces fast and slow directions in the learning dynamics. In a two-block covariance model, we derive the full time-dependent bulk spectrum of the symmetrized weight matrix through a 2× 2 Dyson equation, and we obtain an explicit outlier condition for a rank-one teacher via a rank-two determinant formula. This yields a transient Baik--Ben Arous--P\'ech\'e (BBP) transition: depending on signal strength and covariance anisotropy, the teacher spike may never emerge, emerge and persist, or emerge only during an intermediate time interval before being reabsorbed into the bulk. We map the corresponding phase diagrams and validate the theory against finite-size simulations. Our results provide a minimal solvable mechanism for early stopping as a transient spectral effect driven by anisotropy and noise.
PaperID: 6178, Poster
Abstract: As vision language models are increasingly deployed in clinical diagnosis, understanding how they internally resolve competing visual and textual signals becomes a safety imperative. Existing mechanistic analyses remain confined to unimodal text and offer no explanation for why a single misleading sentence can override a correct image based diagnosis, or why a model commits to a confident answer despite insufficient visual evidence. We find that these two safety risks, arbitration failure where textual context overrides visual grounding and brake failure where the model commits without adequate evidence, are mediated by spatially disjoint attention head populations: arbitration heads form a mid-to-deep wideband reflecting cross-layer evidence competition, while brake heads concentrate in a narrow middle-to-late layer band that regulates evidence sufficiency and abstention behavior. To ground these observations in causal circuitry, we introduce CRAFT, which localizes each failure mode to a minimal causal head set via dual criteria and verifies necessity and sufficiency through temporal probes and Tuned Lens trajectory analysis. Excising arbitration heads sharply reduces conflict following with negligible degradation on clean inputs, while excising brake heads restores appropriate abstention under degraded visual evidence. The two interventions target spatially disjoint head sets and produce distinct corrective effects, underscoring the mechanistic separability of the failure modes. Experiments across multiple medical VQA benchmarks and VLM architectures validate both the localization and interventions, demonstrating that the identified heads causally drive each failure mode and that targeted modulation generalises without retraining. The code is available at \urlhttps://anonymous.4open.science/r/CRAFT-B3CA.
PaperID: 6179, Poster
Abstract: Large language models (LLMs) have shown strong performance in single-table reasoning but struggle to generalize to multi-table scenarios due to schema heterogeneity, long input contexts, and the lack of structural awareness. We introduce StrucTab-R1, a schema-centric reasoning framework that decouples relational reasoning from raw data exposure. Our approach encodes relational databases as heterogeneous schema graphs, where tables, columns, and foreign-key constraints form distinct node and edge types, and employs a heterogeneous graph encoder with question-conditioned cross-attention pooling to distill compact, semantically grounded schema tokens. Rather than serializing table rows into the context, the LLM plans over the schema representation and emits a Chain-of-Execution: an executable reasoning trace that first identifies the relevant schema subgraph and then performs a sequence of composable tool-function calls (e.g., filter, join, aggregate). These calls are executed externally, with compact observations returned to support subsequent reasoning steps. This design confines the model's attention to relevant subgraphs, mitigating hallucinations in multi-hop join reasoning. To further improve reliability, we train the model with supervised traces followed by structure-aware reinforcement learning, where the reward jointly optimizes answer correctness, schema-region precision, and execution consistency. Experiments on single-table and multi-table benchmarks show that StrucTab-R1 improves execution accuracy over strong general-purpose and table-specialized LLMs and remains effective on large-schema settings where raw-context baselines fail.
Authors:
Wenhua Nie, ZiCheng Zhu, Jianan Wu, Binhan Luo, Haoran Zheng, Jyh-Shing R JangAbstract: Modern LLM APIs often reveal only top-K logit scores and censor the remaining vocabulary. We study the per-position distribution-recovery limits of this access model. For censoring threshold \tau, the compatible teacher distributions form an identified set whose total-variation diameter is exactly U_K=(V-K)\exp(\tau)/(Z_A+(V-K)\exp(\tau)), where Z_A is the observed partition function. For KL recovery, we give a computable binary-endpoint lower bound and an asymptotically matching small-ambiguity upper bound, with an extension to reference-aware attackers. Experiments on a Qwen3 math-reasoning teacher reveal a layered extraction hierarchy: on-task top-K distillation recovers 12% of private capability, full-logit distillation recovers 56% despite 99% KL closure, and generation-based extraction recovers 96%. Top-K censoring therefore limits per-position distribution recovery but does not by itself prevent capability extraction, separating fidelity from transfer in prompt-only logit distillation.
Abstract: Diffusion models have demonstrated strong performance in time series modeling due to their ability to progressively capture complex data distributions through iterative denoising. However, existing approaches struggle with frequency-sensitive denoising, high-frequency reconstruction and balancing global trends with local dynamics. To address these limitations, we propose HyFAD, a Hybrid time-frequency Diffusion model with Frequency-Aware embedding for time series imputation. Built upon the DDPM paradigm, HyFAD adopts a coupled time-frequency diffusion framework, in which the reverse denoising proceeds sequentially from the time domain to the frequency domain, enabling coarse-to-fine generation. Specifically, the time-domain diffusion process captures low-frequency global trends, while the frequency-domain diffusion process refines high-frequency spectral components. We further introduce a frequency-aware step embedding that exploits the relationship between diffusion steps and spectral components, providing step-dependent spectral guidance and facilitates more accurate band-wise reconstruction. Extensive experiments on multiple benchmark datasets demonstrate that HyFAD achieves state-of-the-art performance. Our source code is available at \urlhttps://anonymous.4open.science/r/HyFAD-0C21/.
Abstract: Recent findings by \citetcohen_gradient_2021 demonstrate that when training neural networks with full-batch gradient descent at step size \eta, the largest eigenvalue \lambda_\max of the full-batch Hessian consistently stabilizes around 2/\eta. These results have significant implications for convergence and generalization. This stabilization, however, does not occur for mini-batch stochastic gradient descent (SGD), so the implications above do not directly transfer. We show that SGD trains in a different regime we term Edge of Stochastic Stability (\textscEoSS). In this regime, what stabilizes at 2/\eta is \emphBatch Sharpness: the expected directional curvature of mini-batch Hessians along their corresponding mini-batch gradients. As a consequence, \lambda_\max---which is generally smaller than \emphBatch Sharpness---is suppressed, aligning with the long-standing empirical observation that smaller batches and larger step sizes favor flatter minima. We further discuss implications for mathematical modeling of SGD trajectories.
Authors:
Hanqing Wang, Mingyu Liu, Xiaoyu Chen, Chengwei MA, Yiming Zhong, Yuhao Liu, Wenti Yin, Xing Mu, Jiahao Yuan, Zhiqing Cui, Lu Dai, Zhiyuan Ma, Hui XiongAbstract: 3D affordance grounding aims to highlight the actionable regions on 3D objects, which is crucial for embodied AI. While previous research primarily leverages static cues from language or single images, which often lack the rich interaction context necessary to accurately locate functional zones. To alleviate this predicament, we collect a comprehensive video-based 3D affordance dataset, VIDA, which contains 38K human-object-interaction (HOI) videos covering 16 affordance types, 38 object categories, and 22K point clouds. Based on VIDA, we propose a strong baseline: VideoAfford, a unified framework that extends Multimodal Large Language Models (MLLMs) with fine-grained affordance segmentation capabilities. VideoAfford incorporates a latent action encoder to distill dynamic interaction knowledge from demonstration videos and a spatial-aware loss to encourage geometrically consistent affordance predictions. Extensive experiments on VIDA show that VideoAfford significantly outperforms strong baselines in both seen and unseen settings, demonstrating its effectiveness in video-driven 3D affordance reasoning and open-world generalization. The dataset and code will be released soon to facilitate future research in this area.
Abstract: Anticipating LLM behavioral tendencies from low-cost psychometric probes is critical for safe deployment, but only if self-reports (SR) reliably predict behavior. Recent work documented substantial SR–behavior dissociation in LLMs, but relied on broad personality traits (Big 5) that predict specific behaviors weakly even in humans. Furthermore, the isolation of conversational sessions combined with weak context matching left open whether LLMs truly lack coherence or whether the conditions needed to detect such coherence were not met. We contrast Big 5 with the Theory of Planned Behavior (TPB), which measures intention targeted to a specific behavior and predicts human behavior substantially better than broad traits. We run experiments across four behavioral tasks and 11 frontier LLMs, while also varying session context and identity induction. We find that SR–behavior coherence exists but is selective. 1) Within a shared conversation, Theory of Planned Behavior reaches human-level coherence; Big 5 does not. 2) Across separate conversations, coherence survives only for behaviors anchored outside the immediate prompt, such as implicit bias shaped by training, and collapses when behavior is strongly primed by context, as with sycophancy. 3) Persona prompting makes self-reports more consistent across conversations but does not bring behavior into alignment. These findings suggest that coarse personality frameworks such as Big 5 may not be the best tools for testing deployment behavior. More task- and behavior-specific instruments are needed, and even these must be evaluated across tasks and contexts.
PaperID: 6185, Poster
Authors: Boyoung Kwon, Jiyoon Seo, Sang H Lee
Abstract: Computational cognitive modeling has traditionally relied on manual model construction, which is a time-intensive process that requires substantial domain expertise. Although recent LLM-based approaches have begun to automate model generation, they often prioritize predictive accuracy at the expense of structural plausibility and parameter identifiability. We introduce DisCo, an LLM-guided framework that discovers interpretable symbolic cognitive models through evolutionary search with two structural components: a Validator that filters structurally invalid models, and a Cleaner that promotes parameter identifiability by removing redundant parameters. Across three behavioral tasks, DisCo consistently discovers models that achieve performance competitive with task-specific baselines. Ablation analyses show that the Validator is important for recovering psychologically meaningful structure. Without this component, models tend to violate behavioral constraints and show less consistent associations with external cognitive measures. These results suggest that LLM-driven model discovery, when coupled with appropriate structural constraints, can yield cognitively plausible models that go beyond predictive fit.
PaperID: 6186, Poster
Abstract: Conservative machine learning interatomic potentials (MLIPs) predict potential energy and obtain atomic forces as the negative gradient of energy with respect to atomic positions. Computing forces requires a backward pass from energy to atomic positions, and backpropagating the force loss differentiates through this pass, creating double-backward execution. This execution creates two activation classes with different lifetimes: forward activations and first-backward activations. Heterogeneous atomic graph batches make their memory footprint input dependent, while standard checkpointing mainly targets forward activations and does not directly manage first-backward activations. Therefore, we present MICA, an activation checkpointing planner for double-backward execution in conservative MLIP training. MICA combines a checkpoint primitive for double-backward phases with an input-aware dynamic programming planner that selects checkpoint modes for dynamic atomic graph batches under a specified memory budget. Across EqV3, MatRIS, PET, and UMA, MICA's full-checkpoint mode reduces peak memory by 49% to 86%, compared with only 10% to 24% from standard PyTorch checkpointing. On an 80GB GPU, training EquiformerV3 without checkpointing peaks at 75.47GB and exceeds 50GB in 76.7% of 1000 iterations, while MICA keeps all 1000 iterations within a 50GB memory budget with 16.25% step time overhead. Under the same GPU memory capacity, MICA increases the maximum feasible atoms per batch by 4.2× to 9.2×. In distributed training, the reduced peak memory allows MICA to reduce reliance on heavier memory saving techniques, such as Fully Sharded Data Parallel (FSDP) and graph parallelism (GP). On a single H100 node with NVLink, MICA trains a MatRIS-MoE model with 2B parameters, achieving a 1.88× to 2.03× speedup over FSDP and GP baselines.
PaperID: 6187, Poster
Abstract: One of the fundamental differences between biological neural networks and artificial neural networks (ANNs) is that biological neurons obey Dale’s principle: each neuron projects either exclusively excitatory or exclusively inhibitory outputs. While this constraint is central to biological circuits, previous attempts to impose Dale’s principle in ANNs have often led to training instability or degraded performance. In this work, we introduce EISep RNN, a recurrent neural network architecture that implements Dale’s principle through a learnable excitatory–inhibitory separation. Across a series of learning settings, including multi-task learning, continual learning, and reinforcement learning, EISep achieves more stable training dynamics, competitive or improved performance, and sparser, more modular connectivity than vanilla RNNs and previous Dale-constrained models. EISep also exhibits stable and informative latent dynamics. These results establish Dale’s principle as a biologically grounded and computationally effective inductive bias for recurrent neural networks.
Abstract: Visual instruction tuning is crucial for improving vision-language large models (VLLMs). However, many samples can be solved via linguistic patterns or common-sense shortcuts, without genuine cross-modal reasoning, limiting the effectiveness of multimodal learning. Prior data selection methods often rely on costly proxy model training and focus on difficulty or diversity, failing to capture a sample’s true contribution to vision-language joint reasoning. In this paper, we propose CVS, a training-free data selection method based on the insight that, for high-quality multimodal samples, introducing the question should substantially alter the model’s assessment of answer validity given an image. CVS leverages a frozen VLLM as an evaluator and measures the discrepancy in answer validity with and without conditioning on the question, enabling the identification of samples that require vision-language joint reasoning while filtering semantic-conflict noise. Experiments on Vision-Flan and The Cauldron show that CVS achieves solid performance across datasets. On Vision-Flan, CVS outperforms full-data training by 3.5% and 4.8% using only 10% and 15% of the data, respectively, and remains robust on the highly heterogeneous Cauldron dataset. Moreover, CVS reduces computational cost by 17.3% and 44.4% compared to COINCIDE and XMAS. The code is anonymously available at https://anonymous.4open.science/r/CVS-ABAA.
Authors: Nathanael Bosch
Abstract: Filtering-based probabilistic numerical solvers for ordinary differential equations (ODEs) have been established as a flexible and efficient simulation framework with built-in numerical uncertainty quantification. However, problems that are both stiff and high-dimensional remain a challenge, as current methods are either stable and have cubic cost in the ODE dimension, or scale linearly at the expense of stability. In this paper, we close this gap and develop probabilistic ODE solvers that are both stable and scalable. We propose two complementary strategies. First, we develop a matrix-free update step that uses Jacobian-vector products, iterative linear solvers, and stochastic covariance estimation to enable linear scaling, all while retaining stability. Second, we propose iterative re-linearization to further improve stability without sacrificing scalability, turning probabilistic ODE solvers into fully implicit methods. We evaluate the proposed approaches on a range of stiff and high-dimensional problems and demonstrate improved stability and scalability over established probabilistic solvers.
PaperID: 6190, Poster
Abstract: We study stochastic simple bilevel optimization with smooth, possibly nonconvex upper- and lower-level objectives accessed only through stochastic gradient oracles. A key challenge is that the dual multiplier induced by the lower-level constraint may become unbounded near lower-level stationary points, invalidating bounded-dual analyses and destabilizing stochastic gradient estimates. To address this, we propose \emphStochastic Dynamic Barrier Perturbed Gradient (SDBPG), a single-loop method that adaptively perturbs the dual formulation to regularize this degeneracy. The perturbation stabilizes the multiplier and yields controlled bias and variance even near the lower-level stationarity region. Under a mild rare-visit assumption, SDBPG finds an (\epsilon_f,\epsilon_g)-stationary point in \mathcalO(\max\\epsilon_f^-2,\epsilon_g^-2\) iterations, with sample gradient complexities \mathcalO(\epsilon^-4) and \mathcalO(\epsilon^-6) for the upper- and lower-level objectives where \epsilon=\max(\epsilon_f,\epsilon_g). We further develop PR-SDBPG, a penalty-regularized variant that eliminates the rare-visit assumption, and VR-PR-SDBPG, which improves the resulting sample complexities entirely through variance reduction. To our knowledge, these are the first explicit (\epsilon_f,\epsilon_g)-stationarity guarantees for stochastic nonconvex-nonconvex simple bilevel optimization.
PaperID: 6191, Poster
Abstract: Graph Neural Networks (GNNs) are effective for modeling relational data but are vulnerable to label poisoning, where a small number of corrupted training labels can propagate errors across the graph via message-passing. Despite this risk, defenses against label poisoning remain underexplored: existing methods are primarily designed for label noise and often rely on robust losses or heuristic data cleaning that fail to distinguish adversarial poisoning from naturally hard examples. In this paper, we propose CoLAT (Contrastive Label-flipping for Adversarial Training), a novel contrastive adversarial training framework tailored for robust node classification on graphs. The core of our approach is a structure-aware detection model that uses contrastive learning to identify label–structure inconsistencies. Unlike prior contrastive methods that focus on representation learning, CoLAT leverages contrastive embeddings to select high-risk nodes that guide adversarial label-flipping during training. This alternating optimization not only performs structure-aware adversarial training but also implicitly sanitizes corrupted labels, improving robustness against diverse attacks from the literature. Extensive experiments on multiple benchmarks show that CoLAT consistently outperforms existing defenses under various poisoning intensities while scaling efficiently to large graphs.
Abstract: Physics-based digital twins aim to predict the dynamics of real-world objects under interaction, enabling real-to-sim-to-real applications in robotics. Current approaches reconstruct such twins as explicit physical models (such as spring-mass systems) to predict the dynamics, but the resulting models often inherit the resolution of the visual reconstruction rather than being reduced to the physical complexity required to reproduce task-relevant dynamics. This mismatch introduces redundant topology, making repeated forward-dynamics rollouts unnecessarily expensive. To address this challenge, we present PhySPRING, an fully differentiable GNN-based method to reduce complexity in spring--mass digital twins. PhySPRING jointly learns a hierarchy of coarsened graph topologies and their mechanical parameters from observations. At each reduction level, PhySPRING merges nodes with similar learned dynamic responses to optimize the topology, while maintaining every reduced layer as an explicit spring--mass system. On the PhysTwin benchmark, PhySPRING improves dense reconstruction and prediction accuracy over PhysTwin, while reduced models retain stable physical and visual fidelity with up to a 2.30× speed-up. We further demonstrate the effectiveness of PhySPRING in a Real2Sim robot policy-evaluation pipeline, where the reduced models are substituted zero-shot into ACT and \pi_0 evaluations, maintaining comparable manipulation success rates across downsampling levels while improving action-sampling effectiveness. Together, PhySPRING enables efficient and structure-preserving spring--mass reduction without sacrificing fidelity or robotic utility.
PaperID: 6193, Poster
Authors:
Xuemei Jia, JIAWEI DU, Jiawei Liu, xin zhang, Jun Chen, Zheng Wang, Joey Tianyi ZhouAbstract: Face privacy preservation aims to suppress identity-related information while maintaining realistic facial appearance, for which diffusion models have recently emerged as a powerful paradigm. However, identity manipulation under diffusion-based generation exhibits inconsistent behavior during denoising, where certain perturbations are progressively attenuated while others are amplified into unstable distortions. This phenomenon suggests that identity manipulation is inherently constrained by the underlying diffusion denoising dynamics. In this work, we study the feasibility of identity manipulation under such dynamics and show that effective perturbations exhibit a practical operating regime along the denoising trajectory. Building on this insight, we propose a feasibility-aware diffusion-based identity manipulation framework. The framework anchors latent updates to a reference diffusion state and regularizes optimization toward a diffusion-compatible neighborhood, enabling stable identity suppression without compromising visual fidelity. Extensive experiments demonstrate strong robustness and competitive image quality over existing diffusion-based face privacy methods.
PaperID: 6194, Poster
Authors: Justus F Hübotter, Serge Thill, Marcel A. J. van Gerven, Nasir Ahmad
Abstract: As novel hardware architectures are explored for the purposes of machine learning, including neuromorphic chips, asynchronous systems, and ASICs, effective learning algorithms are highly desirable. While backpropagation is the standard for differentiation on standard hardware, its requirement for global synchrony and exact differentiability limits the exploration of exotic algorithms on non-traditional substrates and devices. Zeroth-order (ZO) optimization methods, such as SPSA and weight perturbation, offer a compelling alternative by estimating gradients through inference-only passes; however, these methods have historically suffered due to the "curse of dimensionality," where random perturbations become increasingly misaligned with the true gradient in high-dimensional spaces. In this work, we move beyond blind random noise by deriving zeroth-order search directions from activation geometry available during the forward pass. This leads to a family of Spectral Zeroth-Order methods, SZO-\kappa, whose directions solve local response-based gradient-guessing objectives under explicit stochastic assumptions. These methods spend perturbation budget on feature-supported directions that are more likely to produce informative loss responses. To further bridge the gap with first-order methods, we study covariance corrections and probe stabilizers, including search direction orthogonalization and bias corrections that reduce estimator error. Using multilayer and convolutional layers and networks, we analyze gradient alignment, bias-variance structure, and training stability of these methods relative to standard perturbation baselines and backpropagation.
Abstract: While it is generally understood that zeroth-order (ZO) algorithms have an extra dependency on their number of iterations for any choice of parameters, compared to their first-order (FO) counterparts, in this work, we show that under several conditions, in expectation, ZO methods do not suffer from extra dimension dependencies in their convergence rates with respect to their FO counterparts. We look at optimisation algorithms from the dynamical systems perspective and analyse the conditions under which one can formulate the average of a ZO algorithm as the average of its FO counterpart with bounded perturbations with values dependent on design parameters. Then, using input-to-state stability properties, we show ZO methods follow the same decay rate as their FO counterparts and converge to a neighbourhood of the fixed point of FO methods, where its radius depends on the bound of the norm of the perturbations, which can be made arbitrarily small. The theoretical findings are illustrated via numerical examples.
Abstract: We introduce First-Order Trajectory Matching (FTM), a surrogate-modeling method that learns the first-order local transport of probability mass from trajectories of stochastic systems. By matching the symmetric first-order motion of trajectories, FTM learns the probability current velocity, whose flow preserves time marginals to match ensemble averages, while also capturing current-like trajectory quantities such as fluxes, circulations, and barrier-crossing currents. FTM learns the current velocity directly from trajectories, avoiding drift, diffusion, and score estimation. Our stability analysis separates discretization error from sampling variance and shows that the one-step simulation-free FTM loss is stable when temporal resolution and sample size are properly balanced. Across stochastic dynamical systems and PDE examples, we empirically demonstrate that FTM provides trajectory-aware ensemble predictions at low, deterministic-rollout cost.
PaperID: 6197, Poster
Authors:
Yuxuan Liang, Xu Li, Xiaolei Chen, Rui Zhu, Zhe Liu, Haotian Chen, Yi Zheng, Wenjuan Meng, Zhuolin He, Fan Shi, Xiangyang XueAbstract: Although Large Vision-Language Models (LVLMs) have achieved remarkable progress, complex visual reasoning remains challenging. Existing approaches suffer from fundamental limitations: , despite moving reasoning into latent space, still focuses largely on static visual modeling and lacks explicit characterization of temporal dynamics. We therefore propose , a training framework that enables LVLMs to reason about dynamics in continuous latent space from only a static image. We introduce Mahalanobis Novelty Token Selection and Novelty-Adaptive Temporal Quantization to construct dynamic latent supervision from open-source datasets, building DLVR-SFT-100K and DLVR-RL-4K. We further develop a two-stage SFT pipeline that first builds temporal grounding over explicit dynamic processes and then teaches the model to encode dynamic semantics and temporal structure into latent tokens. Finally, we propose Contrastive Dynamic Latent Policy Optimization, which encourages latent trajectories to align with real dynamics while moving away from counterfactual ones. DLVR empirically improves both dynamic-centric and general visual reasoning, achieving an average gain of 10.82% on BabyVision and a 10.50% improvement on HRBench4K, demonstrating the promise of dynamic latent reasoning for LVLMs.
PaperID: 6198, Poster
Authors:
Zhuo Zhang, Li Zhang, Ying Miao, Hongzong LI, Zhixuan Liang, Yong Yang, Gemine VivoneAbstract: Physics-Informed Neural Networks (PINNs) struggle to capture high-frequency components of PDE solutions, a failure mode known as spectral bias. Most existing fixes adjust the loss, the sampling, or the input encoding, but the network itself remains agnostic to the equation it is supposed to solve. We take a different route and embed physics priors into the architecture. The network's internal spectrum is shaped to match the spectrum of the target PDE through three components that are trained end-to-end: a bounded coordinate warping that contracts regions where the solution varies sharply, a spectral attention layer driven by the PDE residual that emphasizes physically active Fourier modes, and an inter-harmonic gate that lets low-frequency channels modulate high-frequency ones. We test the design on four forward PDEs (Helmholtz, Wave, Klein--Gordon, Burgers), the Burgers inverse problem, and seismic reconstruction, including Gulf of Mexico field data. Across these tasks, the architecture clearly outperforms seven PINN baselines on accuracy and produces noticeably fewer artifacts. Code will be made publicly available upon acceptance.
PaperID: 6199, Poster
Abstract: Although Dynamic Sparse Training (DST) has emerged as a promising solution for deploying deep neural networks on resource-constrained edge devices, unstructured DST methods suffer from irregular weight distributions that incur extra index overhead and severely degrade computational efficiency. In contrast, semi-structured sparsity offers computational efficiency but typically suffers from performance degradation at high sparsity levels with strict semi-structured constraints. To address these challenges, we propose Semi-structured Cannistraci-Hebb Training soft rule (SCHTs), a novel hardware-efficient dynamic sparse training framework, which includes four key stages: semi-structured sparse topological initialization, weight initialization, global network pruning, and local network regrowth. Unlike unstructured DST methods, SCHTs supports fine-grained and coarse-grained N:M semi-structured sparsity patterns with constant fan-in or fan-out, ensuring computational efficiency. Extensive experiments demonstrate that the proposed SCHTs framework enables highly sparse networks (>90%) to achieve performance comparable to, or even surpassing, their dense counterparts. (1) For spiking neural networks, models trained with SCHTs consistently outperform their dense counterparts at 90% sparsity, yielding accuracy improvements of +0.05% (to 94.79%), +2.95% (to 75.01%), and +0.09% (to 99.16%) on CIFAR-10, CIFAR-100, and N-MNIST, respectively. (2) For artificial neural networks, sparse networks trained with SCHTs outperform their dense counterparts with accuracy improvements of +0.90% (to 77.54%) and +0.63% (to 63.87%) using GoogLeNet on CIFAR-100 and TinyImageNet, as well as a +0.67% (to 78.97%) improvement using ResNet-152 on CIFAR-100. (3) For large language models, SCHTs enables a 90% sparse LLaMA-130M to yield a highly competitive validation perplexity of 25.49 on OpenWebText, outperforming unstructured DST methods like RigL. Furthermore, cycle-accurate simulations on Trapezoid (a specialized sparse matrix accelerator) reveal that SCHTs achieves a 75% reduction in index overhead and a 32.7% improvement in computational efficiency at 1:16 sparsity, demonstrating its effectiveness in translating algorithmic sparsity into tangible on-chip acceleration.
PaperID: 6200, Poster
Abstract: Large Reasoning Models (LRMs) have achieved impressive performance via explicit Chain-of-Thought (CoT) reasoning, yet this process introduces critical safety risks. We formalize LRM safety objectives as requiring non-malicious reasoning and answers, safety consistency, and explicit intent identification. We argue that existing alignment methods suffer from poor adversarial generalization, a significant ``safety tax'' on reasoning, and unfaithful coarse-grained rewards, primarily due to their neglect of the LRM's intrinsic reasoning capabilities for safety. To address these, we propose SAVeR^2, a novel Safety Alignment method via Rewards with Reasoning capability. SAVeR^2 decomposes the safety objective into four fine-grained verifiable reward signals and utilizes a difficulty-aware data selection strategy to stabilize training. Crucially, we introduce a reasoning-preserving gradient projection mechanism that analytically resolves conflicts between safety and reasoning gradients, effectively eliminating the safety tax. Extensive experiments on models ranging from 1.5B to 671B demonstrate that SAVeR^2 achieves superior safety performance and significantly lower over-refusal rates across multiple benchmarks without compromising general reasoning capabilities.
PaperID: 6201, Poster
Abstract: Text-to-video (T2V) models trained on large-scale web data can generate undesired content, motivating interventions that reduce harmful outputs without sacrificing visual quality. Activation steering offers an attractive mechanistic alternative to finetuning and prompt filtering, but existing T2V steering methods remain limited, typically applying coarse non-anticipative interventions that can lead to oversteering and content degradation. To close this gap, we propose Latent Activation Linear-Quadratic Regulator (LA-LQR), a reduced-order optimal control framework for minimally invasive T2V steering. LA-LQR formulates T2V inference as a dynamical system and computes closed-loop feedback interventions that steer activations toward desired feature setpoints while penalizing unnecessary perturbations. To make optimal control feasible for high-dimensional video activations, we project activations onto a low-dimensional, task-relevant subspace derived from contrastive prompt pairs, estimate local linear dynamics in this latent space, and solve a latent LQR problem to obtain timestep- and layer-specific steering signals. We provide theoretical bounds relating latent setpoint tracking to raw activation-space feature control, and empirically validate the fidelity of the reduced latent dynamics. On concept steering and video safety benchmarks, LA-LQR reduces unsafe generations relative to baselines, while preserving prompt fidelity and visual quality.
Abstract: Neural correlates of spatial cognitive map are well documented, yet exactly how neural circuits perform spatial navigation in complex environments — e.g., reaching a goal while avoiding obstacles — remains largely unclear. Here, we show that a hippocampal network with appropriate recurrent connections can naturally achieve optimal goal-directed navigation via its relaxation dynamics. Specifically, we consider that the recurrent weights between the neurons represent the transition probabilities between spatial locations encoded by neurons; obstacles such as walls and blocked corridors are therefore reflected by the vanishing of connection weights. This connection pattern can be learned in the hippocampus via behavioral-timescale synaptic plasticity (BTSP) while the animal is exploring the environment. When a goal signal is presented, the network dynamics will relax into an activity field representing the goal location. We prove that this field is mathematically equivalent to the desirability field of a Linearly-solvable Markov Decision Process (LMDP), and the local log-gradient of the field indicates the navigation direction. Both theoretical analyses and simulations demonstrate that this recurrent network dynamics-mediated navigation is efficient and robust in environments with complex obstacle layouts. Moreover, only low-rank updates of the network's connection pattern are needed when the environment has local changes. We hope this study offers insight into a general circuit principle for planning in abstract rational maps in the brain beyond spatial navigation.
Abstract: Standard IMVC evaluation retrains separate models for different missing-data configurations. We show that this paradigm obscures a fundamental vulnerability: missing rate alone is insufficient to characterize data incompleteness. Specifically, we show that protocols with identical nominal missing rates can differ by up to 50× in their proportion of fully observed samples, inducing drastically different learning regimes. We formalize this phenomenon as incompleteness divergence, providing measures that capture structural disparities across missing-data protocols. We further prove that for a broad class of reconstruction-based objectives, learning becomes structurally ill-posed when the proportion of complete samples falls below a critical threshold, leading to near-random performance. To bypass this theoretical bound, we propose CRAFT (Complete-data Robust Attention-masked Fusion Transformer). CRAFT shifts the burden of robustness from the loss function to the architecture via two key properties: (i) per-sample independence, which removes reliance on complete-sample co-occurrence, and (ii) mask-aware variable-length fusion, which aggregates only observed views through attention masking. This design allows a single model, trained once on complete data, to generalize to diverse missing patterns at inference time without retraining. Extensive experiments on seven benchmarks show that CRAFT matches or outperforms per-configuration baselines while reducing training overhead by 8.8×, demonstrating that robustness to missing data can be achieved as an inherent architectural property. Code (CRAFT) and our imvc-audit toolkit are available at https://anonymous.4open.science/r/CRAFT-BF80/ and https://anonymous.4open.science/r/imvc-audit-8263/.
Abstract: CT report generation (CTRG) requires models to summarize three-dimensional anatomical context and pathological findings from hundreds of axial slices. Existing methods typically learn a direct image-to-text mapping, providing limited mechanisms for modeling how CT evidence evolves across slices or how reports respond to controlled changes in latent lesion-related factors. We propose SliceWorld, a CT-specific world-state framework that treats an axial CT scan as an ordered sequence along the z-axis. SliceWorld encodes prefix CT evidence into factor-aware latent states containing anatomy, lesion, and uncertainty components, and projects these states into world tokens used for multi-step future-slice feature prediction, lesion-factor intervention, and LLM-based report generation. The model is first pretrained on CT slice sequences with predictive, factor-aware, and counterfactual objectives, and is then fine-tuned on paired CT-report data. Experiments on M3D-Cap and CT-RATE show that SliceWorld improves natural language generation metrics and clinically oriented automatic evaluation. Further analyses demonstrate multi-horizon future-slice prediction, measurable factor alignment, reduced-slice robustness, and selective lesion-sensitive report modulation.
Abstract: Generative models are increasingly central to many de novo discovery pipelines, in which designs are generated at scale and filtered through virtual screens to determine a set of candidates to experimentally validate. While Bayesian optimization (BO) is a natural fit for this setting, as it uses past evaluations to guide future proposals, the computational overhead required for its sequential decision-making becomes a bottleneck when virtual screens are relatively cheap. We make BO practical in this regime by exploiting the unique combination of a linear model constrained to a spherical domain where high-dimensional latents concentrate. We build off recent work justifying the use of linear surrogates, while deriving nearly closed-form solutions to the surrogate modelling and acquisition problems that exploit spherical symmetry. The result is a 100× speedup over state-of-the-art baselines, with matching or improved performance across molecular and image generation benchmarks. Altogether, our method makes BO a practical drop-in for de novo pipelines where it was previously too slow to consider.
PaperID: 6206, Poster
Abstract: Post-hoc explanation methods for video action recognition typically conflate spatial appearance and temporal dynamics into a single saliency or importance map. We propose a perturbation-based framework that disentangles these contributions into a spatial appearance mask (M_a) and a temporal motion mask (M_m). The two masks are jointly optimized against a frozen action recognition model through factored perturbation architecture, with nearest frame replacement to identify important temporal dynamics. A separation penalty discourages the appearance mask from spending budget on frames that the motion mask has already discarded, and a one-sided ReLU area constraint allows the realized masks to fall below the optimization budget when the input does not require its full extent. To evaluate these explanations, we introduce a separation protocol that sweeps spatial pixels and temporal frames independently, backed by proofs characterizing the inherent geometric bias of standard Insertion metrics. Across EPIC-Kitchens-55, EGTEA Gaze+, and Something-Something V2, our method attains the lowest deletion AUC against leading baselines while producing masks roughly five times sparser.
Authors: Dmytro Rizdvanetskyi, Nathan Roos, Pavlo Lutsik
Abstract: Cell-type deconvolution, the task of estimating the proportions of constituent cell types in a heterogeneous biological sample, is a core problem in computational biology. Methods that rely on epigenetic marks such as DNA methylation typically operate on aggregated methylation estimates, discarding the pattern-level information carried by individual DNA reads. Existing read-level approaches that exploit this information are scarce, and all remain restricted to few-class settings; scaling them further is an open problem because, at scale, non-discriminative reads dominate and hard labels conflict with the many-to-many mapping between methylation patterns and cell types, preventing classifier convergence. To overcome this, we propose data-driven soft labels that estimate the conditional cell-type distribution for each read, and integrate this scheme into Syto, a new modular framework for read-level classification-based deconvolution. On a whole-body atlas of 39 human cell types, Syto reduces MSE by 2.56× over SoTA, with gains transferring to an out-of-distribution dataset spanning 16 tissues. Syto lays the foundation for modeling increasingly large cell-type panels, with improved applications in biology and healthcare. The proposed soft-labeling scheme is further translatable to any setting with a many-to-many signal-to-label mapping.
PaperID: 6208, Poster
Abstract: Large language models (LLMs) increasingly serve as reasoners and are being considered as automated evaluators, yet they remain susceptible to cognitive biases---often altering their reasoning when faced with spurious prompt-level cues such as consensus claims or authority appeals. Existing mitigations via prompting or supervised fine-tuning fail to generalize, as they modify surface behavior without changing the optimization objective that makes bias cues attractive. We propose Epistemic Independence Training (EIT), a reinforcement learning framework built around a simple principle: models should learn that bias cues are \emphunreliable rather than learning to either follow or reject them. EIT trains on balanced conflict examples where each injected cue is equally likely to support the correct or incorrect answer, and uses a reward that penalizes bias-following errors without rewarding agreement with a cue that happens to be correct---making the cue non-predictive of reward. On controlled MMLU-Pro reasoning tasks with bias injection, EIT improves accuracy and robustness on both Qwen3-1.7B and Qwen3-4B when bias points to wrong answers, while preserving performance when bias aligns with truth. Trained only on bandwagon bias, EIT generalizes along two out-of-domain axes: held-out MMLU-Pro subjects and unseen bias types (authority, distraction, verbosity). EIT-trained Qwen3-4B further outperforms untrained Qwen3-8B and Qwen3-14B on bias resistance, indicating that targeted training is more effective than model scaling alone. Code and data are available at https://anonymous.4open.science/r/bias-mitigation-with-rl-BC47.
Abstract: Efficient exploration in high-dimensional decision spaces remains a central challenge for decision-making systems. Humans, in contrast, can navigate large decision spaces with remarkable efficiency. Recent behavioral studies suggest that humans actively reduce dimensionality when facing large decision spaces: they probe candidate feature dimensions, rapidly identify reward-relevant ones, and use them to restrict the effective decision space. Inspired by this mechanism, we propose TDGE (Task-Dimension-Guided Exploration), a human-inspired exploration algorithm, and evaluate its performance in recommendation tasks with contextual bandit backbones. TDGE follows a top-down exploration strategy: it first selects task-relevant feature dimensions, then identifies informative features within those dimensions, and finally recommends concrete items based on the selected features. Experiments on three real-world recommendation datasets, MovieLens-20M, Last.fm, and Amazon, show that TDGE substantially improves exploration efficiency and cold-start adaptation over baseline algorithms. We further provide visualizations of recommendation trajectories showing that TDGE exhibits a human-like dimension-guided exploration process during learning. Code is available at https://anonymous.4open.science/r/Human-inspired-Task-Dimension-Guided-Exploration-for-Efficient-Learning-984B.
PaperID: 6210, Poster
Abstract: High-resolution synthesis with Latent Diffusion Models (LDMs) faces a fundamental trade-off between computational cost and generative fidelity. This trade-off is largely governed by the latent compression factor f (e.g. f8: f=8 and f32: f=32). While aggressive compression (f32) significantly improves efficiency, it suffers from severe information loss and ambiguity compared to standard f8 models. We propose UltraDiff, a framework that reformulates this challenge as joint-latent diffusion learning followed by marginal score distillation. We introduce MixTeacher, a dual-branch DiT that models the joint f8-f32 distribution, utilizing shared attention to resolve the ambiguity of the compressed space with high-fidelity f8 priors. Crucially, we propose vector field distillation in the native f32 space. Theoretically, this compels the student to learn the conditional expectation of the teacher's trajectory, effectively filtering out compression artifacts. UltraDiff matches f8 quality across 1K-4K resolutions while preserving f32 efficiency. Notably, it generates 2K images in 6.2 seconds on a single A100 GPU—over 40× faster than f8 baselines.
PaperID: 6211, Poster
Authors: Tan Wenxin, pring wong, Shuo Chen
Abstract: Recognizing human value preferences is crucial for building reliable AI agents in user-centric fields like psychological counseling and strategic negotiation. Existing benchmarks rely on explicit questionnaires, assuming (1) users are always cooperative and willing to self-report; (2) value preferences remain static. However, humans may conceal true preferences if they feel probed, and may adopt distinct value preferences depending on the context. Consequently, explicit queries are insufficient for predicting value preferences reliably. To address this, we introduce Implicit Value Probing (IVP), a benchmark simulating realistic scenarios where agents implicitly infer value preferences through strategic, multi-turn conversations. IVP grounds these interactions using a diverse set of Persona Agents, covering a wide spectrum of personalities and values. We evaluate mainstream LLMs on IVP, including Qwen3, DeepSeek, Doubao, and GPT-5, observing substantial performance gaps. We further develop Value-Prober-8B, a specialized model internalizing structured cognitive reasoning via distillation. It employs strategic questioning to elicit concealed value preferences. Experimental results demonstrate that Value-Prober-8B can effectively balance probing accuracy with social acceptability, achieving capabilities comparable to frontier models.
PaperID: 6212, Poster
Abstract: The goal of asymmetric pansharpening is to synthesize a high-resolution multispectral (HR-MS) image by integrating panchromatic (PAN) and low-resolution multispectral (LR-MS) images. While State Space Models (SSMs), particularly Mamba, offer a promising paradigm for global dependency modeling, their standard selective scan mechanisms operate in discrete Cartesian grids, which are ill-suited for distinguishing between structural and spectral features in complex-valued frequency representations. To resolve this, we propose Euler-Mamba, a physics-inspired framework that reformulates the selective scan as a continuous state evolution on polar manifolds. We develop an Eulerian Selective Scan (ESS) mechanism that treats the fusion task as a trajectory of frequency growth. By performing radial scans within the polar domain, ESS governs complex-valued state transitions through the decoupled architecture: the Explicit Phase Rotation (EPR) block facilitates explicit phase rotations for geometric alignment, while the Implicit Magnitude Mapping (IMM) block performs magnitude scaling for spectral calibration. This polar-decoupled evolution allows the model to capture the continuous transition from global spectral backgrounds to fine structural details with linear complexity. Consequently, by modeling the fusion as a continuous functional evolution flow, Euler-Mamba achieves resolution-invariant performance and zero-shot generalization across varying spatial scales, bypassing the computational complexity inherent in traditional discrete-integration frameworks. Extensive experiments demonstrate that Euler-Mamba establishes a new Pareto frontier for asymmetric pansharpening, achieving a superior balance between physical interpretability, computational efficiency, and reconstruction quality.
Abstract: Pretrained foundation models have become an important basis for end-to-end autonomous driving. In contrast to vision-language models pretrained primarily on static image-text pairs, video generative models capture temporal dynamics and motion priors that are naturally suited for driving. We present DriveWAM, a driving world-action model that adapts a pretrained video diffusion transformer into an autoregressive video-action policy. DriveWAM organizes video and action streams into a unified temporal token sequence and trains them under a joint flow-matching objective, preserving the pretrained video-generation architecture while adapting its large-scale video priors to action generation. To incorporate high-level scene understanding, we introduce scene-evolving driving guidance, where a frozen VLM produces chunk-specific semantic intent to guide video-action generation. To keep long-horizon rollout bounded, we further introduce selective KV memory, which maintains bounded modality-aware video and action memory pools through relevance-redundancy cache selection at inference time. Experiments on NAVSIM and the PhysicalAI-Autonomous-Vehicles benchmark show that DriveWAM achieves strong planning performance, and a data-scaling study from 4k to 100k driving clips further confirms the scaling potential of world-action modeling for end-to-end autonomous driving.
Authors:
Chao Li, Tianhong Li, Sai V Nuthalapati, Hong-You Chen, Satya Narayan Shukla, Jianpeng Cheng, Yonghuan Yang, Jun Xiao, Xiangjun Fan, Aashu Singh, Dina Katabi, Shlok K MishraAbstract: Unifying text-image contrastive learning and text-to-image (T2I) generation in a single end-to-end model is challenging because the two objectives demand opposing masking regimes: contrastive alignment needs near-complete visible tokens, while masked generative modeling needs heavy corruption. We introduce DREAM, a unified framework that resolves this conflict through Masking Warmup, a schedule that shifts the center of the masking distribution over training, so low and high masking ratios coexist at every step. This co-exposure lets a single jointly-trained encoder serve both objectives. The resulting stable optimization unlocks Semantically Aligned Decoding at inference: the text encoder, trained against visual embeddings at all masking ratios, can score partially generated images and select the best trajectory with as little as 12.5% of the image decoded, improving both FID and throughput. DREAM outperforms its single-objective baselines, CLIP and FLUID: on ImageNet linear-probing (+1.1%), 5-shot transfer (+4.1%), ADE20K segmentation (+1.9%), and NYU depth estimation (+6.25%) over CLIP, and on CC12M FID (+6.2%) over FLUID while maintaining CLIP Score. Together, these gains show that text-image contrastive and generative objectives, when properly unified, are synergistic rather than competing.
PaperID: 6215, Poster
Abstract: Large Language Models (LLMs) fail at a core aspect of human cognition: retrieving related episodes through shared elements . When asked which clubs Ronaldo played for from 2008 to 2010, Claude Opus 4.6 answers Real Madrid, missing Manchester United, despite answering correctly when each period is queried separately. We study this failure through two complementary lenses drawn from cognitive neuroscience: (encoding-time reactivation). "In vivo'', we benchmark frontier models (Claude, GPT, Gemini, DeepSeek, Grok, Mistral) and document 4--62% binding failures on verified knowledge. Extended thinking, a proxy for retrieval-based inference, reduces but does not eliminate failures; multiple-choice probing recovers most of them, suggesting knowledge is encoded but not easily retrievable. "In vitro'', controlled experiments across model families and narrative lengths show that standard remedies offer partial progress: scaling improves binding but still falls short at 70B, extended training enhances binding but requires impractical redundancy, and transfer learning slightly helps with extensive supervision. Mechanistic analysis confirms correct answers appear in the model's top logits yet fail to surface during generation. Paralleling integrative encoding, we design Recall-based Synthetic Supervision (RSS), which prompts the model to recall past episodes and synthesize binding supervision online. RSS covers 92% of evaluation queries, confirming the information is recoverable from parametric memory. Under realistic sequential training, however, hallucination, forgetting, and evolving supervision signals prevent this coverage from translating into accuracy. We establish episodic binding as an open problem in Distinct from the perceptual binding problem; our usage draws on the associative inference literature on memory integration (Bunsey & Eichenbaum, 1996; Zeithamova & Preston, 2010).
Abstract: Video-generative world models are increasingly used as neural simulators for embodied planning and policy learning, yet their ability to predict physical risk and severe consequences is rarely evaluated.We find that these models often downplay or omit key danger cues and severe outcomes for hazardous actions, which can induce unsafe preferences during planning and training on imagined rollouts. We propose ICAT, which grounds testing in real incident reports and safety manuals by building structured risk memories and retrieving/composing them to constrain the generation of risk cases with causal chains and severity labels. Experiments on an ICAT-based benchmark show that mainstream world models frequently miss mechanisms and triggering conditions and miscalibrate severity, falling short of the reliability required for safety-critical embodied deployment.
Authors: Zhangcheng Hou, TOMOAKI OHTSUKI
Abstract: Radar-camera depth estimation must turn an ultra-sparse, all-weather, metric radar signal into a dense per-pixel depth map. Existing methods --- concatenation, confidence-aware gating, sparse supervision, graph-based extraction --- combine radar and image features outside the backbone's sequence operator, and even cross-modal Mamba variants leave the selection mechanism itself unimodal. We argue that the selection mechanism is the right place for radar to enter. We introduce Radar-Modulated Selection (RMS), a minimal and principled way to inject radar into Mamba's selective scan: radar modulates the scan from within, adding zero-initialised perturbations to the step size \boldsymbol\Delta and readout \mathbfC while leaving the input projection \mathbfB and state dynamics \mathbfA image-only. The construction is exactly equivalent to a pretrained image-only Mamba at initialisation, ensuring radar only influences the model where it improves accuracy. Two further properties follow that out-of-scan fusion cannot offer: linear-cost cross-modal coupling at every recurrence step, and a natural fallback to the image-only backbone when radar is absent. We deploy RMS in a Multi-View Scan Pyramid (MVSP) that matches the fusion operator to radar's spatial reach at each scale. SemoDepth achieves state-of-the-art performance on nuScenes, reducing MAE by 34.0%, 29.9%, and 29.9% over the previous best at 0-50, 0-70, and 0-80 m, while attaining the lowest single-frame latency (26.8 ms). A further ablation shows that out-of-scan feature blending adds no accuracy on top of RMS, providing empirical validation that in-scan selection can replace out-of-scan fusion.
PaperID: 6218, Poster
Authors: Mumuksh Tayal, Manan Tayal, Ravi Prakash
Abstract: Offline reinforcement learning (RL) offers a promising framework for deploying autonomous systems in safety-critical settings without the risks of online exploration. However, learning policies that simultaneously achieve high performance and strong safety guarantees from fixed datasets remains a fundamental challenge. Many existing safe offline RL approaches typically rely on soft constraint formulations, which may permit safety violations and are sensitive to distributional shift. In contrast, formal methods such as Hamilton–Jacobi (HJ) reachability and Control Barrier Functions (CBFs) provide rigorous safety guarantees, but often yield overly conservative solutions, often neglecting performance. In this work, we bridge this gap by formulating safe offline RL as a state-constrained optimal control problem, where safety is enforced through hard state constraints and performance is captured via a reward function. The resulting value function satisfies a Hamilton–Jacobi–Bellman (HJB) equation, which we approximate using offline RL on fixed datasets. This formulation enables principled integration of safety guarantees with data-driven policy optimization. Empirically, across safety-critical benchmarks including boat navigation and Safety-Gymnasium tasks, our approach achieves competitive returns while exhibiting near-zero constraint violations, demonstrating a favorable balance between safety and performance in the offline setting.
Abstract: Learning the natural parameters z \in \mathbbR^n of discrete distributions \mu_z from independent samples constrained to a subset S \subseteq \\0,1\\^n is a foundational challenge in high-dimensional statistics. Existing methods for efficiently estimating truncated Boolean product distributions, notably the work of [Fotakis et al' COLT'20, Algorithmica '22], require either strong local connectivity assumptions on S -- a property denoted fatness -- or stringent anti-concentration assumptions and necessitate the total mass of the truncation set to be a constant with respect to n. Moreover, the results in [Fotakis et al' COLT'20, Algorithmica '22] suffer from sample complexities that scale as \Omega(2^n) if the mass of S is exponentially small in n. In this work, we circumvent these limitations by analyzing the geometry of S under the measure \mu_z. We refine the existing parameter estimation guarantees under the fatness assumption, improving the prior sample complexity to \mathcalO( \log n / \epsilon^2) for \ell_\infty-recovery, matching the untruncated minimax rate. We further generalize fatness using the notion of influence utilized in the analysis of Boolean functions and provide sufficient conditions for efficient inference. Notably, unlike previous work, our method does not require sampling at arbitrary parameterizations of the model and instead relies on gradient descent. Lastly, we establish a theoretical lower bound demonstrating the sample complexity exhibits an intrinsic exponential dependence on the width of the model and the minimum distance between elements in the set.
PaperID: 6220, Poster
Abstract: Catastrophic forgetting in continual learning (CL) manifests not only as accuracy degradation but also as explanation drift: models maintain predictive accuracy while silently shifting attention from diagnostic features to spurious correlations. This work identifies and resolves a previously overlooked structural flaw in explanation-aware CL: pixel-level saliency regularization derives supervision targets from the model's own frozen saliency maps, creating a \emphself-referential drift loop that actively reinforces spurious attention across tasks. We introduce C^3L (Concept-Consistent Continual Learning), which breaks this loop by anchoring regularization to class-consistent semantic concepts (e.g., ``curved beak'', ``striped pattern'') discovered offline, independent of model state, providing (i) target independence from model drift, and (ii) cross-instance semantic consistency across every instance of a class. We show that our concept-based objective, combined with approximate LRP relevance conservation, produces a crowding-out effect: enforcing high relevance within concept regions implicitly suppresses spurious background correlations without explicit penalization, forming a zero-sum competition over a finite relevance budget. \method doesn't require human-annotated concept sets, as it employs a fully automated concept discovery and spatial grounding. Extensive evaluation across five benchmarks and over 20 baselines demonstrates state-of-the-art performance in accuracy, forgetting, explanation quality, spurious correlation mitigation, and out-of-distribution robustness. Our code is provided in the supplementary material.
PaperID: 6221, Poster
Abstract: As agent capabilities advance, existing benchmarks, such as \tau^2-Bench, are becoming increasingly saturated. Yet constructing new benchmark tasks remains complex, costly, and labor-intensive. Moreover, the prevailing top-down process, in which scenarios are first written in natural language and then mapped to tool sequences, captures only a narrow subset of the tool-use patterns agents exercise. In this paper, we address these problems by reversing the task construction process. We propose TASTE: Task Synthesis from Tool Sequence Evolution, an automatic method that generates challenging tasks with broader tool-use coverage. TASTE utilizes an Adaptive Contrastive n-gram model trained on LLM-judged validity signals. This enables sampling valid tool sequences that cover a vast range of tool combinations. TASTE then selects representative sequences from the pool via clustering, instantiates them into complete benchmark tasks, and refines them through iterative difficulty evolution. Using TASTE, we construct \tau^c-Bench, a challenging extension of the three domains of \tau^2-Bench. We evaluate 11 agent/user LLM pairs and find that models nearly saturating \tau^2-Bench suffer severe performance drops on our tasks (e.g., Gemini-3-Flash falls from 0.82-0.94 to 0.28-0.61). Beyond increasing difficulty, our generated tasks more than double the number of unique tool combinations agents must execute. Our results suggest high scores on existing benchmarks often reflect saturation rather than robust task-solving ability. By automating the generation of difficult, high-coverage benchmarks, TASTE enables continuous, scalable evaluation of future agents.
PaperID: 6222, Poster
Authors: Otto Nyberg, Fausto Carcassi, Davide Tugnoli, Giovanni Cinà
Abstract: Predictions from ML models support human decision making in several fields, including high-stakes ones such as healthcare and the judiciary. Yet, we still lack a clear understanding of how decision makers learn from ML-based decision support (ML-DS). In this paper, we introduce a general computational framework, the 2-Step Agent, to capture this process. As a prediction from an ML model contains information about the training data, a prediction can also be used for inference. Our framework models (i) how a prediction for a new observation affects the beliefs of a rational Bayesian agent, and (ii) how this change in beliefs affects the estimation of causal effect, the downstream decision, and the subsequent outcome. In addition to the framework itself, we make three contributions. First, for the linear Gaussian setting, we derive a tractable solution for the challenging Bayesian inference problem we introduced, i.e. one in which the agent infers from an ML prediction. Second, we experimentally identify conditions under which ML-DS is beneficial. Third, we show that a single misaligned prior belief can be sufficient for ML-DS to lead to worse downstream outcomes compared to no decision support
PaperID: 6223, Poster
Authors:
Quanyu Zhang, Zhongyi Han, Zhenxue Chen, Xiao-Long Yin, Henggui Zhang, Shuo LiAbstract: Pretraining noise in foundation models has been shown to significantly impair the generalization performance of downstream tasks, leading to catastrophic inheritance. However, when downstream tasks arrive sequentially, how pretraining noise affects continual adaptation, which involves the forgetting of previously learned knowledge, remains underexplored. This paper is the first to study this problem. Using asymmetric label noise as a realistic setting, we find that pretraining noise persistently degrades performance on new tasks during continual adaptation and may further exacerbate forgetting of previously adapted tasks. Further analysis shows that the persistent effect of pretraining noise mainly manifests as inherited output contraction, where the logit gap between the true class and competing classes is compressed. To mitigate this issue, we propose Ambiguous Boundary Correction (ABC), a method that performs anchor-guided neighborhood reweighting to correct ambiguous regions induced by inherited output contraction. Experiments on synthetic and real-world noisy pretrained models show that ABC improves robustness to pretraining noise under both domain shifts and class shifts.
PaperID: 6224, Poster
Abstract: Continual Reinforcement Learning (CRL) requires agents to adapt to sequentially arriving tasks while retaining performance on previous ones. Most existing CRL methods rely on model-free frameworks, addressing catastrophic forgetting via regularization, parameter isolation, or experience replay. World models offer a promising alternative by accumulating transferable dynamics knowledge and enabling planning-based decision-making. However, existing world model approaches suffer from slow adaptation at task boundaries: abrupt reward shifts mislead the planner, fixed planning horizons accumulate prediction errors when the dynamics model is unreliable, and quality-agnostic replay dilutes training while exacerbating forgetting. We propose HERD (Horizon-adaptive Elite Replay with reward Disagreement), which addresses these challenges from two complementary perspectives: for \emphhow to learn, Uncertainty-aware Reward Planning leverages ensemble disagreement to accelerate reward calibration and guide exploration, while Adaptive Planning Horizon adjusts rollout depth based on dynamics prediction error to prevent error accumulation; for \emphwhat to learn, Elite Experience Replay applies quality-stratified reweighting of historical trajectories to strengthen knowledge retention. Together, these components \emphaccelerate new-task adaptation and mitigate catastrophic forgetting. Experiments on Continual World, Continual Bench, and DM Control show that HERD consistently outperforms existing methods in overall performance, forgetting mitigation, and new-task adaptation.
PaperID: 6225, Poster
Abstract: Deep neural networks continue to grow in parameter count, driving up training and inference cost on GPUs. Sparse neural networks and Dynamic Sparse Training (DST) promise to reduce these costs, but most implementations rely on binary masks over dense tensors and recover little of the theoretical compute, memory, or energy savings. We propose SNACK, a truly sparse GPU layer that stores and computes only non-zero connections. SNACK exposes a simple PyTorch API for restructuring connections and backpropagating gradients entirely in the sparse paradigm, and uses a custom COO-format SpMM CUDA kernel with a batch-to-Streaming-Multiprocessor mapping tuned for the small-batch, high-sparsity regime typical of large-model training and single-stream inference. At the kernel level, SNACK is up to 7× faster than the masked dense baseline (Dense+Mask) and competitive with cuSPARSE, Sputnik, and FlashSparse at 95% sparsity. At 90% sparsity, a single SNACK layer accelerates training by 8× and 3.7×, and inference by 4× and 2×, over Dense+Mask and fully dense layers, respectively, while using 25% less memory than dense and substantially less energy. End-to-end, SNACK reduces GPT-2 peak training memory by up to 40% and graph-style inference latency by 4.8× over Dense+Mask at 99% sparsity.
PaperID: 6226, Poster
Abstract: Decision Transformer (DT) has emerged as a powerful paradigm for offline reinforcement learning. For complex tasks with imbalanced datasets, where high-quality samples constitute only a small fraction of the data while the majority is suboptimal, the performance of DT methods usually degrade significantly due to the the offline data distribution deviates substantially from the stationary distribution of the optimal policy. In this paper, we construct a dual-based correction mechanism to mitigate the distribution shift problem and propose a Distribution Corrected Decision Transformer method. Specifically, we investigate the Lagrangian duality of the reward maximizing objective in the standard Markov decision problem and recover the optimal correction weights between the static offline data and optimal stationary distribution. Note that a double-KL regularizer is integrated in the objective to prioritize high-return state-action pairs while remaining anchored to valid data support. Subsequently, we use the learned weights to correct distributional bias in two critical components of DT: stabilizing policy evaluation through weighted temporal-difference learning, and improving policy extraction by prioritizing DT training on weighted dataset aligned with the corrected target distribution. We theoretically analyze the performance guarantee, proving our method improves upon standard baselines, and empirically demonstrate strong performance on D4RL benchmarks, particularly on highly imbalanced datasets where prior methods fail.
PaperID: 6227, Poster
Abstract: Entities in real-world systems often evolve over time: users shift between information consumption patterns, researchers migrate between fields, and firms transition between financial risk states. When such systems are modeled as dynamic graphs, these transitions correspond to nodes whose class labels change between consecutive snapshots, which we call shifters. Correctly classifying shifters is often more consequential than classifying stable nodes with persistent labels, as detecting such transitions enables timely intervention in high-stakes settings. Yet, we identify a systematic failure mode across all major dynamic graph neural network architectures: shifter performance consistently lags behind stable nodes, with errors concentrated on the old label. Our structural analysis traces this failure to embedding inertia: neighborhood aggregation keeps shifter representations anchored to the old-class centroid, and this effect is strongest precisely where it is hardest to correct, namely for shifters whose neighborhoods remain aligned with the old class. Guided by this finding, we propose CHASE, a model-agnostic CHange-Aware framework for Shifting nodEs that wraps any dynamic GNN with targeted components for detecting label shifts and overriding stale neighborhood signals. CHASE consistently improves shifter performance across all tested models and datasets (up to 114.9% and 202.5% improvement on accuracy and F1 score, respectively) while preserving stable-node performance. We additionally contribute three dynamic graph benchmarks with naturally shifting labels, filling a gap in existing resources. The source code is available at https://anonymous.4open.science/r/CHASE-40CF/.
Abstract: When language model agents tackle complex software engineering tasks, they often degrade over long trajectories, which we define as agent drift. We focus on two recurring failure modes overthinking and overacting, i.e., where the agent repeatedly reasons over information it already has, and where it issues tool calls without integrating recent observations or acquiring new evidence. In this paper, we introduce TACT (Think-Act Calibration via activation sTeering), to detect and mitigate agent drift in the residual stream before it surfaces as a behavioral failure. In specific, we label trajectory steps as overthinking, overacting, or calibrated, and find that their hidden states can separate linearly along two drift axes, pointing from calibrated behavior toward each failure mode (AUC \approx 0.9). To mitigate agent drift, we project each step's activation onto these axes at test time and pull drifted ones back toward the calibrated region. Experiments show that TACT outperforms unsteered baselines across SWE-bench Verified, Terminal-Bench 2.0, and CLAW-Eval, lifting average resolve rate by +5.8 pp on Qwen3.5-27B and +4.8 pp on Gemma-4-26B-A4B-it while cutting steps-to-resolve by up to 26%. These gains frame agent drift as a steerable direction in the residual stream, and position TACT as a viable handle for reliable long-horizon agents.
Abstract: Over the past decade, graph-based approximate nearest neighbor (ANN) algorithms, such as DiskANN (Jayaram Subramanya et al., 2019) and HNSW (Malkov & Yashunin, 2018), have demonstrated state-of-the-art empirical performance. Recent theoretical works (Indyk & Xu, 2023; Gollapudi et al., 2025) introduce the framework of \alpha-reachability to obtain worst-case performance guarantees for DiskANN; here, the reachability parameter \alpha gives a trade-off between construction time, query time, and accuracy. In this work, we propose RP-TUNING, an efficient and simple post-hoc algorithm, based on DiskANN's pruning step, and show that (1) Theoretically, efficiently adjusting the reachability of an \alpha-reachable graph is possible via pruning: RP-TUNING preserves worst-case reachability guarantees in general metrics and improved guarantees in Euclidean metrics. (2) Empirically, RP-TUNING accelerates DiskANN tuning on four datasets by up to 73× with varied performance trade-offs compared to fully rebuilt graphs.
PaperID: 6230, Poster
Authors:
Haotong Wen, Junkang Liu, Yi Xu, Xiao Liu, Longkun Guo, Kewen LiaoAbstract: Data heterogeneity and multiple local update steps in federated learning can increase client drift and make global parameter aggregation less stable. Existing Sharpness Aware Minimization (SAM) based federated learning methods usually try to ease this problem by flattening local neighborhoods around client models. However, they do not explicitly optimize on the geometry region where the global parameters are actually aggregated. As a result, the aggregated global model often cannot fully benefit from local flattening. To address this problem, we propose a federated restricted SAM method called FedRSAM, which improves the stability of federated optimization by flattening the parameter aggregation region. First, we build a cone-shaped geometry region using the global update direction from the previous round and the variance of client parameter aggregation, in order to predict where the aggregated parameters may fall in the next round. Then, we restrict the search of SAM perturbation directions to the cone-manifold and only flatten the loss surface inside this region. Extensive experiments on multiple benchmarks and on both vision and natural language backbones verify the superiority of our method in terms of convergence stability and performance under heterogeneous settings.
Authors: Cassandra T Ye, Shamus Li, Tyler King, Kristina Monakhova
Abstract: While deep learning offers tremendous promise for scientific and medical imaging, any failures and hallucinations (predictions that do not coincide with reality) are hard to pinpoint and can have serious downstream consequences. Uncertainty estimation techniques, such as conformal prediction, can help by predicting statistically valid error bars for a model's prediction. However, popular conformal prediction methods were not designed for high-dimensional image-valued problems and do not take into account spatial correlations within an image during conformal calibration, resulting in larger-than-necessary uncertainty intervals. We propose a practical simultaneous quantile regression method that enables non-linear, spatially-adaptive scaling during conformal calibration. Our method, QUTCC uses a U-Net architecture with a quantile embedding to learn a full conditional quantile distribution during training, and then leverages this non-linear, learned function for spatially-adaptive conformal calibration. At test time, our method can efficiently estimate uncertainty intervals with pixel-marginal coverage guarantees. In addition, QUTCC can also predict pixel-wise conditional probability density estimates without any built-in distributional assumptions. We evaluate our method on several denoising problems, accelerated magnetic resonance imaging, and quantitative phase microscopy. Our method consistently produces tighter uncertainty intervals than prior conformal methods at the same coverage level, can predict plausible conditional distributions for different tasks, and in some cases, high-uncertainty regions can help us locate hallucinations in a model's prediction.
PaperID: 6232, Poster
Abstract: Deep bisimulation metric learning has emerged as a principled framework for learning robust state representations in reinforcement learning. However, prevailing methods rely on sample-based regression objectives that exhibit a critical stability flaw in stochastic environments. We identify this flaw as a manifestation of the : minimizing mean squared error against single-sample stochastic targets introduces an irreducible variance bias. Through the lens of stochastic differential equations (SDEs), we rigorously prove that this bias leads to a , wherein target variance scales with the encoder's sensitivity. This variance effectively acts as an implicit Jacobian regularizer, driving representation collapse in the presence of transition noise. To address this issue, we propose Bisimulation Saddle-Point Optimization (BSPO), a debiased framework that reformulates the bisimulation metric learning objective as a primal-dual saddle-point problem. By introducing an Auxiliary Network (AuxNet) to estimate the expected Bellman error, BSPO decouples the structural error from transition noise, eliminating the bias without requiring physically infeasible double sampling. To the best of our knowledge, this work is the first to identify that prevailing sample-based bisimulation metric learning objectives are systematically biased and to provide a practical debiasing framework. Our empirical evaluations on stochastic continuous/discrete control tasks and visually distracting environments demonstrate that BSPO significantly outperforms representative baselines, particularly in high-noise regimes where conventional methods fail to converge.
PaperID: 6233, Poster
Abstract: Recent AI research has increasingly evolved along two complementary directions: model-centric learning, which improves architectures and training algorithms, and data-centric learning, which improves the quality and compression of training data. Within data-centric learning, dataset distillation on large-scale datasets has attracted growing attention, with decoupled distillation methods like SRe^2L, G-VBSM, LPQLD, FADRM have emerged as a representative paradigm. These methods typically synthesize compact datasets by matching the Batch Norm statistics of a teacher network, using BN alignment as an effective training objective. However, modern networks often contain many BN layers, and enforcing alignment at every layer introduces substantial computational redundancy. More importantly, we observe that not all BN layers provide equally informative supervision or consistent performance gains. Motivated by this, we propose SODA, a novel Selective Optimization with Deferred BN Alignment framework for efficient dataset distillation. SODA selectively identifies and optimizes informative BN alignment objectives while deferring unnecessary or low-value alignment computations, substantially improving both distillation efficiency and synthetic data quality. We further provide a theoretical analysis explaining why selective and deferred BN alignment can simultaneously reduce optimization cost and improve generalization. Extensive experiments across multiple datasets of CIFAR-100, Tiny-ImageNet, ImageNet-1K and its subsets thereof demonstrate that SODA achieves state-of-the-art performance while offering significantly improved computational efficiency over existing BN-matching-based dataset distillation methods, surpassing FADRM+ by +1.6% on ImageNet-1K IPC=10 under ResNet-18 while delivering a 1.54× speedup.
Authors: Sanah Suri, Kieran Ringel, Maike Sonnewald
Abstract: Extreme ocean phenomena are challenging not only to predict but to diagnose, as accurate forecasts alone do not reveal the underlying physical drivers. While recent machine learning approaches achieve strong predictive skill, they remain largely opaque and provide limited guarantees of fidelity to ground-truth physics. We introduce OceanCBM, the first concept bottleneck model (CBM) for spatiotemporal prediction and mechanistic interrogation of ocean dynamics. OceanCBM uses mixed supervision to predict mixed layer heat content, a key precursor of marine heatwaves, while routing information through an intermediate layer of prescribed concepts derived from geophysical fluid dynamics and a 'free' concept. This design imposes soft physical structure without over-constraining the model, and the free concept both regularizes concept predictions and captures residual physical processes. Across ensemble initializations, we show that mixed supervision yields consistent mechanistic representations, whereas prediction-only and prescription-only baselines learn highly variable latent structures despite similar predictive performance. OceanCBM achieves interpretable, physically grounded representations without sacrificing skill, explicitly characterizing the interpretability-performance trade-off.
PaperID: 6235, Poster
Abstract: Diffusion policies have recently emerged as a powerful paradigm for representing complex action distributions in reinforcement learning (RL). However, their application to online RL remains limited by the challenge of scalable training in the absence of ground-truth data, where standard optimization techniques such as score matching are not directly applicable. In this work, we introduce a highly efficient algorithm for optimizing diffusion policies by leveraging recent advances in stochastic optimal control. Our approach is based on adjoint matching, which enables simulation-free training and circumvents the need for explicit likelihood estimation or costly backpropagation through the diffusion process. Furthermore, we propose several extensions that improve the robustness and stability of the method in practical settings. Empirical results demonstrate that our approach achieves competitive performance while significantly reducing computational overhead, making diffusion policies more viable for online RL scenarios.
PaperID: 6236, Poster
Authors: CHIU-CHANG CHENG, Ching-Lung Hsu, Ya-Ning Chang, Chao-Hung Wang
Abstract: Biological circuits learn without catastrophic forgetting, but the structural basis of this ability remains unclear. We investigate whether distance-dependent connectivity (DDC), spatial recurrency found in the mammalian cortex, contributes to continual learning by embedding neurons of a biologically plausible recurrent spiking neural network in a 3D Euclidean substrate, using pairwise distance to set recurrent connection probability, and ablating that distance dependence in silico. We show that DDC shapes a synaptic resource geography: it controls the spatial breadth of the candidate synaptic pool for learning to occur, plasticity further compresses that pool, and the resulting substrate governs how subsequent inputs compete for the same synapses. On challenging 10-way class-incremental MNIST classification, this produces a non-monotonic accuracy curve, with an intermediate DDC range achieving the best performance (71.4%) and outperforming both highly local circuits that collapse into a synaptic-resource bottleneck and random connectivity that produces weakly guided, broader synaptic contention. This pattern was better explained by synaptic competition, resulting in separable neural dynamics, than by neuron assembly separation. These results suggest that appropriate cortical DDC may prevent diffuse random contention. For continual learning problems, this finding highlights the conceptual importance of circuit dynamics more than active neuron ensembles and engrams.
PaperID: 6237, Poster
Authors: Abhinav Rajeev Kumar, Paras Chopra
Abstract: Large language models change their answers when a "verified" source contradicts them. We study this verified-source authority bias across five open-weight families and three closed APIs. On a baseline-correct trivia subset, a verified-source cue endorsing a wrong answer flips 43–87% of responses across seven of the eight models, with a graded hierarchy over cue strengths. We then show this is not generic user sycophancy: across both open-weight families and closed APIs, the same wrong answer endorsed by a verified source vs. by a user produces different behavioral compliance, and a fitted "authority" vector is distinct from a generic assistant/instruction-following direction and from a user-sycophancy direction under causal projection-removal. Applying this vector to neutral prompts induces matched wrong-answer flips, and the same trivia-fit vector transfers without refitting to PIQA (a held-out task) and to multi-turn SYCON dialogues. Projecting the vector out at α=1 reduces wrong-source compliance by tens of percentage points on both trivia and PIQA in four of five open-weight families, while capability checks (MMLU-Pro, GSM8K) stay within noise of baseline at our evaluation sizes.
PaperID: 6238, Poster
Abstract: Exemplar-Free Class-Incremental Learning (EFCIL) is a challenging continual learning paradigm where a model must learn new classes sequentially without access to old data, making it susceptible to catastrophic forgetting. The core difficulty lies in balancing stability (preserving old knowledge) and plasticity (acquiring new knowledge). We propose Subspace-Guided Continual Learning (SGCL), a novel method that tackles this dilemma from a geometric perspective. SGCL decomposes the feature space into two orthogonal subspaces: a where new knowledge can be learned with minimal interference. This decomposition is efficiently identified via the feature-space Hessian, where high-curvature eigendirections define the stable subspace. Building on this, SGCL introduces two synergistic components: 1) Subspace-Guided Regularization (SGR), which imposes curvature-weighted penalties on feature drifts within the stable subspace, and 2) Subspace-Guided Prototype Alignment (SGPA), which adaptively corrects the shift of old-class prototypes to recalibrate the classifier. Extensive experiments on standard benchmarks show that SGCL consistently achieves competitive or superior performance compared to existing state-of-the-art methods, offering a principled approach to mitigating forgetting through loss landscape analysis.
Abstract: Schedule-Free methods have attracted growing interest for alleviating the burden of designing and tuning a learning rate scheduler, while matching and sometimes even outperforming optimizers with tuned schedulers. Despite their strong empirical results, their convergence theory in nonconvex optimization, where modern machine learning objectives typically arise, has remained largely unexplored. In this paper, we provide worst-case analyses of Schedule-Free gradient descent and Schedule-Free stochastic gradient descent, in their standard form and without auxiliary modifications or restrictive conditions, for smooth but possibly nonconvex objectives. Based on a Lyapunov analysis derived from the continuous-time limiting ordinary differential equation associated with these methods, we show that Schedule-Free gradient descent and Schedule-Free stochastic gradient descent achieve the optimal worst-case convergence rates attainable among first-order methods. We further formulate Schedule-Free gradient descent as a nonautonomous dynamical system and prove strict-saddle avoidance under an arbitrarily small one-time perturbation. These theoretical results provide a better understanding of the strong performance that Schedule-Free methods demonstrate.
Abstract: A skill is a modular package of natural-language instructions, executable scripts, and reference resources that an agent can load at runtime to extend its capabilities for a specific task. Skill-based agent systems therefore enable flexible reuse of third-party capabilities, but the openness of this skill ecosystem also opens up a new attack surface. Prior work has focused on vulnerabilities within individual skills, but little attention has been paid to risks that arise from interactions across skills. In this paper, we introduce skill cascading attacks, a threat paradigm in which a malicious objective is distributed across multiple skills so that each modification looks benign in isolation, yet their combined execution is harmful. For instance, in a prescription-review pipeline, the first skill weakens signals of recently discontinued medications in the extracted history, the second downgrades the severity of any drug interaction tied to them, and the third suppresses the resulting low-priority alert in the final summary, so that a severe drug-interaction warning silently disappears before reaching the physician. To systematically study this safety blind spot, we develop SkillCascade, an automated multi-agent red-teaming framework, and release SkillCascade-Bench, a benchmark of 213 validated cascading test cases across multiple agent systems and domains. Across representative agents (e.g. OpenClaw, Claude Code, CodeX) and LLM backbones, cascaded interactions reliably induce harmful behaviors while evading existing per-skill scanners and runtime monitors. Our findings highlight a gap between component-level integrity and system-level safety, and call for defenses that reason over cross-skill interactions rather than individual skills in isolation.
PaperID: 6241, Poster
Abstract: Generative emulators of protein dynamics produce plausible trajectories at a fraction of the cost of molecular dynamics, but they inherit their training distribution and tend to revisit known states rather than reach rare ones under long-horizon extrapolation. Inspired by classical enhanced sampling, we introduce an implicit, history-dependent bias in the generative space of a pretrained emulator. Specifically, a history-aware score estimator augments the frozen emulator with a distance-weighted bias that steers reverse-time sampling away from previously generated structures, regularized by an environment-support term. To preserve structural validity at long horizons, a score-based refinement step re-projects drifted samples onto the data manifold using the frozen emulator. Our experiments demonstrate that the method (i) raises diversity by 35% on DynamicPDB-80; (ii) on 12 zero-shot Fast-Folding proteins, the learned bias alone reaches the unbiased emulator's coverage up to ~15× faster, and pairing it with refinement reaches the coverage up to ~37× faster while covering ~3× as many low-energy states.
PaperID: 6242, Poster
Abstract: Scientific discovery is fundamentally open-ended, characterized by epistemic uncertainty, and requiring endogenous goal-setting and recursive knowledge accumulation. While current agentic systems have demonstrated success in addressing goal-based discovery tasks, they struggle to sustain discovery in open-ended settings. In this work, we posit that open-ended discovery is an emergent network-level property of a socially driven system. We introduce \textttASCollab, which translates four key algorithmic conditions: (1) population heterogeneity in research behaviors, (2) endogenous interaction networks where collaborations and attention routing emerge organically, and (3) socially driven peer evaluations, which are enabled by (4) a global, shared memory that reflects evolving research states. Specifically, social memory is implemented to capture time-varying expertise and reputational signals, and shifting attention paid to historical artifacts produced by the system. Through experiments on large-scale discovery problems in genomics, cell biology, and epidemiology, we demonstrate that \textttASCollab produces discoveries judged by domain experts to be both sound and significant. Crucially, the system independently recovers real-world scientific findings. Furthermore, systematic investigations show that social dynamics fostered by each of the four conditions are vital, where heterogeneous agents operating within self-organizing networks significantly outperform both fixed-workflow baselines and homogeneous networks.
PaperID: 6243, Poster
Authors: Jiin Im, Sisung Liu, Je Hyeong Hong
Abstract: Estimating relationships between 3D shapes requires both dense endpoint correspondence and a deformation path connecting the source to the target. These tasks are coupled: the correspondence constrains the interpolation endpoint, while a plausible trajectory can help disambiguate the map. However, when only endpoint shapes are observed, the interior trajectory is under-constrained. Existing joint frameworks typically model interpolation through sampled per-vertex displacements regularized by local rigidity or temporal smoothness, which can indirectly restrict the stretch and shear needed for non-isometric or poorly aligned shape pairs. We propose Chebyshev Differential Flows (Chef), an unsupervised framework that models motion as a continuous time-varying field of per-face Jacobians rather than vertex trajectories. Each Jacobian trajectory is represented by a low-order shifted Chebyshev expansion, yielding a compact arbitrary-time deformation model whose odd and even modes separate endpoint-visible deformation from endpoint-invisible interior control. A spectral functional-map branch estimates the endpoint correspondence, while a differentiable Poisson solver integrates the predicted Jacobian field into globally consistent intermediate shapes. We further introduce a trend regularizer that guides intermediate stretch and shear without imposing per-snapshot As-Rigid-As-Possible (ARAP) rigidity. Experiments on near-isometric and non-isometric benchmarks show competitive correspondence accuracy, improved interpolation proxy metrics, and better stability when the standard rigid pre-alignment step is omitted. Code will be publicly released for research.
PaperID: 6244, Poster
Abstract: Beyond reducing the memory footprint and costs for deployment, the extent to which a probabilistic model can be compressed is central to understanding generalization, model selection, and dataset selection. Yet with parameter-based compression methods (quantization, pruning, low-rank training) the code length scales with model size and is insensitive to how much information the model actually learned. Prequential coding sidesteps the parameter ceiling by encoding the dataset through the learning process, but its code length includes the entropy of the data. We introduce requential coding, a model coding scheme that is unconstrained by model size and does not pay for the data entropy: synthetic training samples are drawn from a teacher distribution and communicated relative to the student's current predictions via relative entropy coding, so the code length equals the cumulative teacher-student KL along the training trajectory. The resulting compressibility measure behaves qualitatively differently from parameter-based codes. The code length to reach a fixed target loss decreases with model and ensemble size, evidence that a more flexible model can be far simpler than its parameter count implies. Plugged into a PAC-Bayes bound, requential coding produces non-vacuous generalization guarantees that improve with scale and beat even the lossless idealization of the 4-bit GPTQ baseline that previously gave state-of-the-art bounds on compute-optimal LLMs. The same code length predicts overfitting in data-constrained training and tracks intuitive ordering of dataset complexity across CIFAR-5M, OpenWebText, and FineWeb.
PaperID: 6245, Poster
Abstract: While diffusion distillation methods have largely accelerated image generation, achieving high-quality synthesis under a single number of function evaluations (NFE) budget remains a challenging problem. The diffusion process follows a coarse-to-fine progression, yet existing solutions typically force a single student model to simultaneously resolve global structures and fine details in one forward pass, leading to over-smoothed outputs. To address this issue, we propose Phase-wise Velocity Distillation (PVD), which strategically partitions the generation timeline into a coarse and a fine phase, and models the transition within each phase via the average velocity. A dedicated half-sized expert is then assigned to each phase, decoupling structural composition from detail refinement while keeping the cumulative cost within a single-NFE budget of the teacher. While this design already suffices for class-conditional image (C2I) generation, we further introduce phase-wise adversarial supervision with dedicated discriminators for the more complex text-to-image (T2I) tasks, ensuring accurate distribution matching. On C2I generation, PVD achieves a state-of-the-art FID of 1.48 under a single-NFE budget on ImageNet 256 × 256. On T2I tasks, PVD-distilled models (Stable Diffusion 3.5-Medium, FLUX.1-dev) produce results competitive with their multi-step teachers, significantly outperforming prior distillation methods. Source codes and distilled models will be released.
Abstract: Multiple instance learning (MIL) is the dominant framework for whole-slide image analysis in computational pathology, typically combining a frozen patch encoder, a projection layer, and a slide-level aggregator. While encoders and aggregators have been extensively studied, the projection layer remains a largely morphology-only bottleneck. This limits endpoints such as biomarker status and survival, which are governed by a molecular state that is not fully captured by H&E morphology. We introduce Molecularly Informed Staining Transform (MIST), a plug-in replacement for the MIL projection layer that uses paired spatial transcriptomics only during training to construct virtual molecular stains. MIST clusters gene expression profiles into cross-modal prototypes, anchors them in the frozen foundation model feature space, and uses them to reorganize H&E patch features along molecularly guided axes. It requires no transcriptomics at inference and can be inserted before standard MIL aggregators. We evaluate MIST across 23 downstream tasks and 8 MIL aggregators. MIST improves 240 of 256 configurations over the standard projection layer, with an average gain of +3.5%, observed consistently across endpoint types: +5.2% on survival prediction, +3.3% on tissue subtyping, and +2.6% on biomarker prediction. Ablations confirm that gene-derived prototypes are the primary source of the gains, while spatial, biological, and pathological analyses show that cross-modal prototype affinities capture spatially coherent molecular programs from H&E alone.
Abstract: Aligning large language models (LLMs) with diverse and multifaceted user preferences is a fundamental challenge in personalized AI systems. Existing multi-objective alignment methods either rely on costly training or require pre-trained reward models for each preference, making it difficult for them to adapt to evolving preferences. Prompt-based personalization offers a training-free alternative, but prompting alone often provides limited steerability, as LLMs may overemphasize or overlook certain preferences and fail to give users reliable control over the relative importance of different objectives when conflicts arise, leading to suboptimal alignment. In this paper, we introduce MATO, a training-free framework for Multi-objective personalized Alignment with Test-time Optimization. MATO formulates personalization as a test-time optimization problem that steers the relative importance of multiple objectives through controllable weights during decoding, without modifying model parameters or requiring external reward models. Specifically, a reward discovery module recovers preference rewards directly from the backbone LLM for diverse objectives specified in natural language, while a weight optimization module dynamically adjusts objective weights based on the user's initial preferences and the partially generated response to balance competing objectives during generation. The resulting rewards and weights jointly guide an online optimization procedure over the token distribution, enabling better alignment with the target objectives. Extensive experiments across multiple datasets and backbone LLMs show that MATO consistently outperforms strong baselines, achieving Pareto-improving multi-objective alignment and stronger steerability. These results highlight test-time optimization as a promising direction for scalable, controllable, and model-agnostic personalized alignment.
PaperID: 6248, Poster
Authors: Nan Zhang, Heng Xu
Abstract: Autoregressive reasoning traces often sharpen as they unfold---entropy falls and token probabilities concentrate---suggesting decoding is resolving uncertainty. Yet the same pattern persists after clear reasoning errors. We analyze decoding through the conditional posterior it induces over ---equivalence classes of continuations under any chosen notion of "resolution"---and show that a mismatch between token selection and branch resolution produces two regimes: , where branch-specific residuals self-average and top tokens become weakly diagnostic. Two observable consequences follow: entropy may fall while unresolved support stays elevated, and greedy token selection has opposite informational effects across regimes. We confirm these theoretical predictions across multiple open-weight models and reasoning benchmarks.
PaperID: 6249, Poster
Abstract: We present Neural Dual Bounds, a calibration-free framework that produces deterministic upper bounds for MAP inference in probabilistic graphical models using a single neural-network forward pass. The framework trains a network to output the edge messages of a mini-bucket style upper bound on a given join-graph decomposition in a self-supervised manner by directly minimizing the a smooth dual objective. The output of this network can be used at test time to warm-start a standard JGLP optimizer. Across diverse benchmarks, this warm start yields dual gaps that are orders of magnitude tighter than those from zero initialization, with the largest gains in the early-iteration regime. We also apply our framework for the constrained MAP (CMPE) task, and we show that neural warm starts with test-time adaptation consistently achieve the tightest dual gaps across all iteration budgets against an adaptive baseline.
PaperID: 6250, Poster
Authors:
Robert Moseley, Md Kaykobad Reza, Ameya Patil, Edward Ayrapetian, Salman AsifAbstract: Large multimodal models have achieved remarkable progress by scaling model capacity and training on massive paired multimodal data. However, this paradigm introduces a fundamental bottleneck: aligned multimodal datasets are expensive to curate, difficult to scale across diverse modality combinations, and increasingly constrain further progress. We introduce Modality Self-Play (MSP), a training framework for multimodal foundation models based on proxy-mediated modality composition. Instead of relying on paired cross-modal data, MSP uses anchor-supporting proxy textual descriptions of unseen modalities as semantic bridges during training. Given an observed modality (e.g., image or audio), the model generates proxy descriptions for complementary unseen modalities, which are then encoded and aligned with the observed modality through learned bridges, enabling cross-modal interaction without explicit paired supervision. Our central hypothesis is that proxy text, combined with contrastively trained pretrained encoders, provides sufficient semantic structure for multimodal composition to emerge. We show that MSP enables zero-shot modality composition: models trained only with a real anchor modality and accompanying proxy descriptions can integrate and reason over multiple modalities at inference time despite never observing real paired multimodal data during training. Empirically, MSP achieves state-of-the-art or competitive performance on multiple multimodal benchmarks. Ablation studies further demonstrate that learned bridges improve alignment, while anchor-supporting proxy descriptions enable effective cross-modal composition. Together, our results suggest that explicit paired multimodal data may not be necessary for multimodal reasoning, and that scalable proxy-based alignment provides a promising alternative for training multimodal foundation models.
PaperID: 6251, Poster
Abstract: Few-shot tabular learning is challenging because conventional tree-based methods suffer from unstable split selection and leaf estimation when labeled examples are scarce. Although feature engineering based on Large Language Models (LLMs) provides useful semantic priors, it can be brittle when table semantics are incomplete or numerical evidence is unstable. This raises two key challenges: achieving semantic-numerical robustness and calibrating multiple weak signals under sparse supervision. Therefore, we present LEAF, a train-time LLM-guided forest that treats LLM-derived features and split preferences as calibratable priors rather than direct supervision. During tree induction, LEAF aligns semantic guidance with task-specific statistical evidence, including teacher-informed predictive signals and data-dependent partition statistics, through a unified split-scoring rule. At prediction time, it further stabilizes local estimates by blending sparse empirical observations with teacher-informed support at each leaf. Experiments on few-shot tabular classification and regression datasets show that LEAF consistently outperforms strong conventional and LLM-based baselines, improving classification performance by 1.4 points and reducing regression error by 0.149 over the strongest previous method. Robustness analyses further demonstrate that LEAF is more stable than prior LLM-based feature engineering methods under missing table semantics and numerical perturbations induced by distributional variation across tabular feature spaces and schema descriptions.
Abstract: Self-distillation has emerged as a powerful framework for post-training LLMs, where a teacher conditioned on extra information guides a student without it, both from the same model. While this guidance is useful when the student has failed, on successful rollouts, the same mechanism instead overwrites the student's choices and suppresses it's own reasoning. Therefore, we propose reading the original self-distillation signal in reverse: when the student succeeds along a path the teacher would not have predicted, these tokens reflect its self-driven reasoning. Building on this, we propose RLRT (RLVR with Reversed Teacher), which augments GRPO by reinforcing these tokens on correct rollouts. We interpret this as a new form of exploration in RLVR: not uniform diversity, but valuable exploration grounded in the student's own success. Across base, instruction-tuned, and thinking-tuned Qwen3 checkpoints, RLRT substantially outperforms self-distillation and exploration-based baselines, establishing information asymmetry as a new, principled design axis for RLVR.
PaperID: 6253, Poster
Authors: Zeyu Wang, Mingyu Ge, Haiyu Song, Haoran Duan
Abstract: Multi-modal image fusion (MMIF) aims to form a single image by integrating shared information, preserving complementary cues, and coordinating cross-modal conflicts across modalities. However, due to the absence of ground-truth fused images, existing MMIF supervision commonly uses spatial-domain sources or gradient variants as surrogate ground truth, making the supervision mechanism inherently misaligned with the goal of MMIF and causing pixel-level compromise or modality bias. To address this, we propose a relation-constrained supervision paradigm that moves fusion supervision from the spatial domain to a learned relation space. Instead of forcing the fused image to approximate the sources, we use frozen pretrained representation models as information providers and design a learnable feature adapter to align heterogeneous DINO and CLIP features into a unified supervision space. The adapter infers three relation parameters, namely sharedness, dominance, and coordination radius, which define three losses corresponding to the MMIF's goal. To make this space reliable, we devise a self-supervised contrastive ranking objective tailored to the adapter and couple it with the fusion network through alternating optimization. Extensive experiments show that the proposed supervision space yields significant gains regardless of which mainstream backbone the fusion network adopts, offering a supervision paradigm better aligned with the goal of MMIF. Our code will be publicly available.
Abstract: Discrete diffusion models have emerged as powerful frameworks for generating structured categorical data. However, efficiently sampling from reward-tilted distributions remains a fundamental challenge. While Twisted Sequential Monte Carlo (SMC) offers asymptotic exactness for this task, estimating the optimal twist function in discrete state spaces necessitates costly Monte Carlo approximations, resulting a severe computational bottleneck at inference. To overcome this limitation, we introduce Contrastive Distribution Matching (CDM), a novel framework that amortizes the cost of SMC inference by learning a parameterized twist function via positive and negative samples. For efficient training, we reformulate the gradient estimator to leverage the closed-form forward kernels of discrete diffusion models. In practice, evaluating our learned twist function incurs less than 5% additional computational overhead compared to base model inference. Through extensive empirical evaluations, we demonstrate that CDM consistently outperforms existing baselines under matched wall-clock time. We validate the effectiveness and versatility of our approach across a diverse range of applications, including toxic text generation, regulatory DNA sequence design, protein designability, and diffusion large language model alignment.
PaperID: 6255, Poster
Abstract: Speculative decoding (SD) accelerates autoregressive inference by leveraging the memory bound nature of decoding, verifying multiple draft tokens in a single forward pass. For mixture-of-experts (MoE) models, prior work predicts a concave speedup-batch size relationship, with gains peaking at intermediate batch sizes due to expert saturation. We evaluate this across 10 sparse MoE models (30B-1T parameters, 8-512 experts) and find recent fine-grained models break this trend, showing higher speedup at low batch size. We find verification cost is governed not only by activated expert weight, but also by temporal overlap across draft tokens and GPU execution dynamics. At low batch size, fine-grained MoE layers run faster than expected. Small per-expert weights and small per-token activated parameter footprint leave bandwidth underutilized and yield higher L2 cache hit rates. Across draft tokens, routing overlap further reduces the unique experts activated during verification relative to independence assumptions. We model this effect using a maximum-entropy Iterative Proportional Fitting (IPF) framework that predicts unique expert counts from routing-overlap statistics with less than 4% error, and reproduces observed speedup curves from synthetic routing.
PaperID: 6256, Poster
Authors:
Mohammadsajad Abavisani, Kseniya Solovyeva, David Danks, Vince D. Calhoun, Sergey PlisAbstract: Learning directed causal graphs from time-series data poses significant challenges, especially in fMRI where slow sampling rate obscures fast neural interactions. This temporal mismatch leads to undersampling, which can make multiple graphs equally plausible. We address this problem by explicitly modeling undersampling effects when recovering causal graphs. Our approach employs answer set programming (ASP) to enforce domain-specific constraints and optimize soft observational constraints, thereby identifying a Markov equivalence class for the resulting graph solutions. By customizing an ASP solver to collect multiple near-optimal solutions, we obtain not only the single best-fitting graph but an equivalence class of high-scoring graphs for expert consideration. This method, called Real-world noisy RASL (RnR), can also act as a meta-solver: it refines the output of other causal discovery algorithms by accounting for undersampling biases. In synthetic data and empirical brain network data, RnR produces more accurate causal graphs than state-of-the-art methods. When applied as a meta-solver (refining outputs of existing algorithms), it improves F1 scores by an average of 46% over baseline methods; on standalone synthetic data benchmarks, it achieves a 64% improvement. We demonstrate that RnR is robust to varying undersampling rates, maintaining high precision and recall even as sampling becomes more sparse, whereas baseline methods degrade significantly. Finally, we test RnR on real-world settings where ground truth connectivity is unknown such as human brain fMRI data, showing that incorporating undersampling-aware constraints via ASP yields more reliable and interpretable brain connectivity estimates from fMRI time series, bridging the gap between neural dynamics and observational data.
Authors: Kang Liu, Jianchen Hu, Wei Peng
Abstract: Classical ReLU-based Input Convex Neural Networks (ICNNs) are equivalent to the optimal value functions of Linear Programming (LP). This intrinsic structural equivalence restricts their representational capacity to piecewise-linear polyhedral functions. To overcome this representational bottleneck, we propose the SOC-ICNN, an architecture that generalizes the underlying optimization class from LP to Second-Order Cone Programming (SOCP). By explicitly injecting positive semi-definite curvature and Euclidean norm-based conic primitives, our formulation introduces native smooth curvature into the representation while preserving a rigorous optimization-theoretic interpretation. We formally prove that SOC-ICNNs strictly expand the representational space of ReLU-ICNNs without increasing the asymptotic order of forward-pass complexity. Extensive experiments demonstrate that SOC-ICNN substantially improves function approximation, while delivering competitive downstream decision quality. The code is available at \urlhttps://anonymous.4open.science/r/SOC-ICNN-4B18/.
Abstract: Recurrent spiking neural networks (RSNNs) are a promising substrate for energy-efficient control policies, but training them for high-dimensional, long-horizon reinforcement learning remains challenging. Population-based, gradient-free optimization circumvents backpropagation through non-differentiable spike dynamics by estimating gradients. However, with finite populations, high variance of these estimates can induce harmful and overly aggressive update steps. Inspired by trust-region methods in reinforcement learning that constrain policy updates in distribution space, we propose , a distributional update rule that constrains relative change by bounding KL divergence normalized by an estimated signal energy. SATR automatically expands the trust region under strong signals and contracts it when updates are noise-dominated. We instantiate SATR for Bernoulli connectivity distributions, which have shown strong empirical performance for RSNN optimization. Across a suite of high-dimensional continuous-control benchmarks, SATR improves stability under limited populations and reaches competitive returns against strong baselines including PPO-LSTM. In addition, to make SATR practical at scale, we introduce a bitset implementation for binary spiking and binary weights, substantially reducing wall-clock training time and enabling fast RSNN policy search.
PaperID: 6259, Poster
Abstract: Reconstruction error is the standard signal for unsupervised time-series anomaly detection and has been widely adopted by time-series foundation models (TFMs). It is, however, vulnerable to input noise: the resulting score tracks the noise level itself rather than whether a sample is truly anomalous, and this failure is especially pronounced in domains where the noise envelope itself carries information about the underlying process state. We propose the adaptation gap \textDiff = L_FT - L_ZS, the difference between fine-tuned and zero-shot reconstruction losses, as a noise-robust complement. We prove that its conditional expectation cancels the input-noise variance and yields a calibration-anchored, population-level mean-separation guarantee. Building on these results we introduce GAPS (Gradient-Aware Adaptation-Gap Scoring), which selectively augments L_FT with \textDiff under a calibration-only routing layer with no label-tuned hyperparameters. Across a synthetic suite, an HRV benchmark whose normal class is intrinsically noisier than its anomaly class, and TSB-AD-U with a pre-registered noise-stress condition, GAPS preserves performance under standard conditions and yields substantial gains under noise stress.
PaperID: 6260, Poster
Abstract: Sharpness-Aware Minimization (SAM) improves generalization in continuous deep learning, yet applying it to binary-weight networks via latent-space heuristics suffers from a fundamental geometric mismatch: continuous perturbations do not faithfully probe the discrete loss landscape on \-1,+1\^n. We introduce BOLD-SAM, the first sharpness-aware optimizer that operates natively on the Boolean hypercube, replacing the Euclidean \ell_p-ball with a k-bit Hamming ball and solving the resulting discrete min-max problem via greedy ascent followed by sharpness-aware bit-flip descent. The objective is theoretically justified from three complementary perspectives: PAC-Bayes, compression, and distributionally robust optimization. We prove an approximation guarantee for the ascent step and a finite-time convergence bound for the descent, both governed by the discrete interaction Hessian, which emerges as the unifying quantity linking ascent quality to convergence rate. We further connect discrete flatness to Forman-Ricci curvature and algorithmic complexity. Experiments on a wide range of architectures and datasets demonstrate consistent improvements in clean accuracy and out-of-distribution robustness over standard binary training and latent-space SAM baselines.
PaperID: 6261, Poster
Authors: Jiale Guo, Shiji Tong, Weihao Li, Yu Wu, Boxian Lin, Mengji Shi
Abstract: Printed circuit boards (PCBs) are the hardware foundation of modern electronic systems. However, automating PCB schematic design remains difficult. A valid schematic must respect real-IC pin-level constraints and cross-component topological relations, neither of which is captured by generic simulation, local rule checks, or direct LLM generation. To address this challenge, we present TopoScope, a training-free LLM agent that uses the schematic graph not only as the final output, but also to guide generation and verification at runtime. TopoScope couples two mechanisms: Topology-Guided Context Scoping, which computes a scoped view for each component-level wiring decision, and Pattern-Based Rule Verification, which expresses datasheet-derived electrical constraints as graph patterns and uses them as in-loop feedback. We also introduce a 20-design benchmark spanning three complexity tiers and six application domains, with functional assertions evaluated on the generated schematic graph and calibrated against expert judgments. Across seven recent LLMs, TopoScope improves functional correctness, reduces invalid graph references, and enables more reliable generation of medium and hard multi-IC schematics compared with knowledge-provided baselines.
Authors:
SeungWon Seo, DongHeun Han, SeongRae Noh, Hyeongyeop KangAbstract: Executable world models can be read, edited, executed, and reused for planning, but only if the program captures the environment's transition law rather than semantic shortcuts in its surface vocabulary. We study online executable world-model learning under prior misalignment, where an agent must induce state-dependent dynamics from interaction evidence alone, without rule descriptions, reward signals, or trustworthy lexical priors. We introduce Alice, a closed-loop system that treats failed candidate updates as structural signal: when a candidate explains a new transition but loses previously explained ones, the preservation conflict reveals dynamics that the current program had conflated. Alice refines these conflicts into hypothesis classes that both provide compact, class-stratified preservation counterexamples for update and guide frontier exploration toward transitions that are novel and underrepresented with respect to the current program. We evaluate Alice on Baba in Wonderland, a prior-misaligned variant of Baba Is You that preserves simulator dynamics while replacing semantically meaningful rule-property labels with unrelated words. Experiments show that Alice substantially improves executable world-model learning under prior misalignment, and ablations show that both class refinement and class-aware exploration contribute.
PaperID: 6263, Poster
Abstract: Value-model-free RLVR methods such as GRPO assign uniform advantages to all tokens in a rollout, ignoring that tokens contribute unequally. Recent methods use token entropy as an importance proxy but compute it globally across the batch, conflating importance with prompt difficulty and positional trends. We argue that importance should instead be measured relative to a token's own local context, and introduce proximal entropy: a local measure of token importance relative to neighboring tokens, and prove it is invariant to both confounders. Proximal Entropy Policy Optimization (PEPO) uses it to weight per-token advantages and outperforms GRPO and entropy-based baselines on mathematical reasoning across Qwen3-1.7B, Qwen3-4B, and Llama-3.2-3B-Instruct. We also show the formulation generalizes to other algorithms where substituting proximal entropy into existing methods improves, and applying it to single-stream RL succeeds where global entropy fails.
PaperID: 6264, Poster
Authors: Ze Chen, Haomai Zhang, Jiaxuan Zou, Zhuo Chen, Linhua Ye
Abstract: Deployed large language models (LLMs) continually accumulate new experiences and rules, yet standard self-distillation fails to robustly internalize these updates over time. We identify and formalize Continuous Experience Internalization (CEI) collapse, a previously uncharacterized failure mode where recursive same-parameter self-distillation systematically forgets earlier patches, absorbs new ones poorly, and amplifies high-confidence errors. To diagnose this, we introduce CEI-Bench, a stage-wise benchmark quantifying task performance, old-patch retention, new-patch absorption, and error propagation. To mitigate CEI collapse, we propose Recycled Internalization Training (RIT), a framework separating stable consensus from unstable conflicts in teacher supervision. RIT combines multi-view variance isolation (MUSE) and targeted error recycling with cross-stage anchors (RACE), explicitly addressing the structural bottlenecks driving collapse. Across four backbones and three task families, RIT recovers 39%–72% of performance lost to vanilla recursive internalization and maintains stable trajectories. Our work reframes continuous self-distillation as a diagnosable phenomenon, providing a systematic analysis, mechanistic explanation, and principled mitigation for CEI. We position CEI collapse as a fundamental challenge for adaptive LLMs, and RIT as a framework for robust recursive internalization.
PaperID: 6265, Poster
Authors:
Thamirawaran Sathiyalogeswaran, Navindu De Silva, Chamikara Siriwardane, W K Alwis, Jason Mars, Udaya S Thanthrige, Ranga RodrigoAbstract: Low-rank adaptation (LoRA) attaches small low-rank matrices to a frozen backbone, and public repositories like Hugging Face host many such LoRAs per backbone, each for a different concept, character, style, or task. Users routinely want to combine several LoRAs in a single generation, but two existing strategies involve a trade-off with each other. First, combining the LoRA weights into a single update before inference (e.g., static merge) needs only one forward pass, but quality drops, because LoRAs trained for different concepts interfere when their weights are summed into a single update, distorting each LoRA's contribution. Second, running each LoRA in a separate forward pass and summing the outputs (e.g., composite) preserves output quality, but each LoRA adds a full backbone pass, so inference time scales linearly with the active LoRA count. To recover composite-level quality at merge cost, we study a property of LoRA pools overlooked by prior multi-LoRA work. When LoRAs in a pool are trained on related concepts, their low-rank factors overlap substantially. Using singular value decomposition, we find the minimal subspace that preserves 99% of the factor energy. On three diffusion pools, only 46% to 57% of the sum of individual LoRA ranks is enough to represent the pool, and the subspace grows sub-linearly with pool size. A 48-LoRA sub-pool of LoRA-Hub on FLAN-T5 NLP tasks shows comparable overlap (about 62%), while the diversity-curated LoRA-Retriever benchmark sits near the no-overlap baseline at 96%, indicating that across these five pools pool composition rather than modality predicts where overlap is large. We use this overlap to construct LoRASpace, a training-free pool-wide method where the subspace matrices are computed once per pool and reused across two inference modes: a static mode at merge inference cost, and a dynamic mode for input-dependent routing at substantially lower cost than composite. On the diffusion pools, static LoRASpace matches composite quality at merge cost, and dynamic LoRASpace achieves more than 6× speedup over composite at six active LoRAs on a 17-LoRA pool. In our ablation, we show that the quality gain comes from truncating to this low-rank shared subspace, which acts as a form of regularisation. On the near-orthogonal LoRA-Retriever benchmark, the framework predicts limited compression benefit, with gains coming from external retrieval coupling.
Authors: Ning Liu, Chuanneng Sun, Kristina L Klinkner, Shervin Malmasi
Abstract: Direct Preference Optimization (DPO) aligns language models using pairwise preference comparisons, offering a simple and effective alternative to Reinforcement Learning (RL) from human feedback. However, in many practical settings, training data consists of multiple rollouts per prompt, inducing rich preference structure that pairwise DPO fails to exploit. Collapsing such data into independent pairs discards transitivity, introduces redundant or conflicting supervision, and can lead to unstable optimization. We propose Graph Direct Preference Optimization (GraphDPO), a principled generalization of DPO that operates over directed acyclic preference graphs induced by rollout rankings. GraphDPO encodes dominance relations as edges and optimizes a graph-structured Plackett--Luce-inspired objective that aggregates supervision over graph neighborhoods, enforcing transitivity while recovering standard DPO as a special case. To handle discrete or sparse signals, we introduce an equivalence-class construction where responses with identical preferences form graph layers, and intra-layer edges contribute zero loss, preventing spurious gradients. Despite leveraging full graph structure, GraphDPO maintains linear per-prompt complexity via efficient log-sum-exp aggregation. We further incorporate optional ground-truth anchoring by inserting verified solutions as dominant nodes and applying an annealed schedule that stabilizes early training while gradually relaxing oracle supervision. Experiments on reasoning and program synthesis tasks demonstrate superior performance, suggesting that graph-structured preference modeling is a scalable and robust alternative to pairwise and listwise alignment objectives.
Abstract: In-context learning (ICL) allows LLMs to adapt to new tasks via a few demonstrations, but those demonstrations may contain sensitive data. Differentially private (DP) ICL mechanisms mitigate this risk by injecting noise into the aggregation step, but verifying that an implementation actually meets its claimed privacy bound currently requires repeated end-to-end membership-inference attacks (MIAs) against the pipeline as a black box, incurring prohibitive LLM cost and yielding unstable empirical privacy estimates. We propose SnapAudit, an active auditing framework that decomposes a DP-ICL pipeline into a deterministic clean-inference stage and a stochastic DP-noise stage, and audits the full pipeline by combining a small snapshot of the former with bootstrap simulation of the latter. Because clean LLM outputs are near-deterministic at temperature zero, a few thousand clean LLM calls suffice to approximate the snapshot distribution; SnapAudit then bootstraps 10^5 noisy trials from this snapshot at negligible additional cost, with finite-sample uncertainty controlled via an empirical Bernstein correction. For embedding-based mechanisms, we further introduce a multi-sweep search procedure that constructs maximally separable audit signals. SnapAudit achieves 80--200× speedup over prior passive auditing while producing tighter and more stable empirical privacy estimates that closely match theoretical guarantees. Beyond efficiency, SnapAudit uncovers two concrete flaws in existing DP-ICL designs: (i) classical Gaussian noise calibrations underestimate leakage at large privacy budgets, allowing empirical leakage to exceed the theoretical bound; (ii) the sensitivity analysis of an embedding-aggregation mechanism is incorrect when the number of partitions equals one, leading to undersized noise and an outright privacy violation.
PaperID: 6268, Poster
Authors:
Junfeng Wang, Jiawei Liu, Yongchao Xu, Tao Jiang, jiangbo Ai, Jin Zhang, Chong Wang, Qixing ZhangAbstract: Complex visual question answering often requires vision-language models to think with images, i.e., actively acquire local visual evidence by zooming into informative regions of high-resolution images. Our error analysis shows that many failures in this setting originate from incorrect observation actions, where the model fails to localize key regions or obtain crops that contain sufficient evidence for answering. We further observe that bounding-box tokens exhibit substantially higher generation entropy than ordinary language tokens, indicating high uncertainty when the model decides where to look. However, existing reinforcement learning methods based on final-answer rewards provide only trajectory-level feedback, making it difficult to assign precise credit to the observation actions. We propose Observation Policy Optimization (OPO), an RL algorithm for optimizing observation actions in tool-augmented thinking with images. OPO consists of two core modules. First, Uncertainty-Guided Observation Branching (UOB) treats each image zoom-in tool call as an observation event and selectively branches at high-uncertainty, low-redundancy events during rollout, enabling efficient exploration of alternative cropping decisions and their resulting visual evidence. Second, Evidence-Localized Advantage Attribution (ELAA) compares sibling branches using local evidence-sufficiency rewards and assigns credit only to the corresponding observation-action tokens. By decoupling observation-action optimization from trajectory-level outcome supervision, OPO improves visual evidence acquisition and further enhances visual reasoning performance. Experiments show that OPO consistently improves Qwen3-VL models across scales and achieves competitive performance against larger models and tool-augmented baselines.
Authors: Lei pang, Jun Luo, ruinan Jin
Abstract: Group Relative Policy Optimization (GRPO) is a critic-free reinforcement learning algorithm for fine-tuning large language models, but its token-level importance sampling mechanism creates a subtle variance bottleneck. We separate this bottleneck into two mechanisms. First, the standard GRPO clipping rule is incomplete: on the negative-advantage branch, it can pass uncontrolled upper-tail importance weights. This motivates up-only clipping as a standalone correction. Second, up-only clipping alone is not enough if the update remains token-level: the clipped token-weighted score increments need not be martingale differences, so cross terms survive in the squared update. This motivates trajectory-level importance correction. We propose Trajectory-level Importance-Corrected GRPO (TIC-GRPO), which follows this progression: it first caps upper-tail ratios by up-only clipping and then replaces token-level importance ratios with a single trajectory-level probability ratio. The latter restores the measure-change identity needed for trajectory-level martingale cancellation. Our variance theory gives upper bounds, hard-instance lower bounds, and induced convergence lower bounds that separate original GRPO, the token-level up-clipped comparator GRPO_2, and TIC-GRPO. Experiments on math reasoning and coding benchmarks further confirm that TIC-GRPO improves optimization stability and final performance.
PaperID: 6270, Poster
Authors:
Huiwen Han, Lulin Liu, Bangya Liu, Yuanhao Cai, Nuo Chen, Jade Wang, Ziqian Xie, Chenyu You, Shuiwang Ji, Degui Zhi, Zhiwen FanAbstract: 3D brain MRI generation has made significant advances for medical imaging, simulation, and controllable anatomical analysis. However, existing generative models typically synthesize 3D volumes monolithically, often overlooking regional anatomical structures and limiting local controllability. To address these limitations, we introduce AnaDiffusion, an anatomically compositional latent diffusion framework that factorizes the generation process into distinct, anatomically meaningful regions followed by part-to-whole assembly and global refinement. Our approach trains part diffusion models first to capture local structural priors. Then, we inject assembled anatomical composite of parts into whole-brain latent and continue denoising. This mechanism enables the model to resolve global context while preserving the injected anatomy. As a result, AnaDiffusion outputs both explicit part assets and a globally coherent volume, enabling controllable part editing without requiring additional dense segmentation masks at inference time while maintaining coherent part-to-whole brain structure. On ADNI, AnaDiffusion achieves the lowest FID across the whole brain, left/right hemispheres, cerebellar-brainstem complex, and seam regions, while also reducing Cohen's d for ventricles, cerebellum, and brainstem. In localized editing experiments, paired MS-SSIM shows high target transfer and off-target preservation, supporting controllable part replacement with limited non-target anatomical drift.
PaperID: 6271, Poster
Abstract: LLMs are increasingly used not only to generate text from scratch, but also to polish and extend human writing. This makes AI detection a fine-grained problem: downstream users often need to understand the level and form of AI involvement in a text, rather than merely separate human from machine. Existing detectors are poorly suited to this setting because they compress the evidence into scalar scores, selected features, or surface linguistic cues, which miss the subtle patterns left by different forms of AI involvement. We propose , which makes a frozen LLM's internal predictive behavior visible as a tomogram: an image-like map of how the model responds across tokens and layers. The tomogram captures both temporal patterns along the input sequence and hierarchical patterns across transformer layers, providing a richer view of AI involvement than compressed detector scores. A vision backbone then reads this tomogram to classify the level of AI involvement. On a consistently outperforms prior detectors under cross-generator, cross-domain, cross-language, and adversarial shifts, showing that temporal-hierarchical tomograms provide a robust and transferable signal for fine-grained AI-involvement detection.
PaperID: 6272, Poster
Authors: Shu Zhao, TAN YU
Abstract: Visual in-context learning (ICL) enables vision-language models (VLMs) to adapt at test time via demonstration examples, but its effectiveness depends critically on the examples selected. Recent counterfactual selection methods improve ICL by constructing demonstrations that expose how individual attribute changes affect the answer. However, they require multi-stage pipelines involving attribute extraction, caption engineering, and composed image matching. We propose a simpler alternative grounded in a key observation: a model's own prediction errors are natural counterfactual demonstrations that directly encode its failure modes. Our method, SMILE (Self-Mined Visual In-Context Learning from Errors), builds a model-specific pool of hard negatives offline by recording the target VLM's own mistakes. At test time, the prediction serves as a direction signal that surfaces, in the semantic embedding space, the past mistakes most likely to recur for the current query. This prediction-conditioned identification provides a unified mechanism for both closed-set classification and open-ended visual question answering without task-specific design. Across four benchmarks and four VLMs, SMILE delivers an average task-score gain of +2.72 points over the SOTA baseline at 5× lower latency. Source code is provided in the supplementary material.
Authors:
Qizhen (Irene) Zhang, Ankush Garg, Jakob Foerster, Niladri S. Chatterji, Kshitiz Malik, Mike LewisAbstract: Large-scale pretraining datasets drive the success of large language models (LLMs). However, these web-scale corpora inevitably contain large amounts of noisy data due to unregulated web content or randomness inherent in data. Although LLM pretrainers often speculate that such noise contributes to instabilities in large-scale pretraining and, in the worst cases, loss divergence, this phenomenon remains poorly understood. In this work, we present a systematic empirical study of whether noisy data causes LLM pretraining divergences, costing more than 100,000 GPU hours in total. By injecting controlled, synthetic, uniform random noise into otherwise clean datasets, we analyze training dynamics across model sizes ranging from 480M to 5.2B parameters. We show that noisy data indeed induce training loss divergence, and that the probability of divergence depends strongly on the noise type, the amount of noise, and the model scale. We further find that noise-induced divergences exhibit activation patterns distinct from those caused by high learning rates, and we provide diagnostics that differentiate these two failure modes. Together, these results provide a large-scale, controlled characterization of how noisy data affects loss divergence in LLM pretraining.
PaperID: 6274, Poster
Authors:
Chongyang Gao, Marco Postiglione, Venkatramanan SubrahmanianAbstract: Protein language models (PLMs) enable the de novo design of functional proteins at scale, yet they lack intrinsic safety constraints, raising dual-use concerns. We introduce SafePro, a reinforcement-learning framework that aligns PLMs against three simultaneous objectives: structural foldability, toxicity avoidance, and virulence avoidance. SafePro features two novel components: (1) Cross-Task Attentive Fusion (CTAF), a joint reward model that employs gated cross-attention between toxicity and virulence experts to capture shared biological signals without inducing negative transfer, and (2)Task-Conflict-Saliency MGDA (TCS-MGDA), an adaptive weight-calibration procedure for Proximal Policy Optimization (PPO). TCS-MGDA resolves gradient Gram-matrix dilution via parameter-space saliency filtering. On ProGen2, SafePro improves composite alignment scores by +5.21 over the strongest fixed-weight baseline, significantly reducing toxicity (0.0523 \to 0.0139) and virulence (0.6926 \to 0.5407) while maintaining structural quality. Among structurally designable outputs (scRMSD < 2\,\AA), SafePro reduces the harmful fraction from 89% to 34% while producing more than twice as many designable sequences as the base model. We also release a curated virulence dataset of 54,166 sequences from VFDB.
PaperID: 6275, Poster
Abstract: Recent work shows that gradient descent (GD) can achieve almost optimal risk bounds for shallow ReLU networks with a logarithmic width. However, these discussions require O(n^2) gradient computations to achieve this optimality, which is not appealing for modern machine learning problems with large sample size n. In this paper, we significantly improve the existing gradient complexity by showing that stochastic gradient descent (SGD) can achieve similar risk bounds with O(n) gradient computations for shallow ReLU networks with a logarithmic width. As compared to GD, the analysis with SGD is more challenging as there are additional fluctuations incurred by the stochastic gradient noise along the optimization process. With concentration inequalities to handle these fluctuations, we show that the entire SGD trajectory stays around its initialization point with high probability. Under an NTK separability condition with margin \gamma, we show SGD achieves near-optimal risk bounds \widetildeO(1/(n\gamma^2)) with a logarithmic width.
Abstract: Distributed online convex optimization (D-OCO) is a powerful paradigm for modeling distributed scenarios with streaming data. However, the communication cost between local learners and the central server is substantial in large-scale applications. To alleviate this bottleneck, we initiate the study of D-OCO with compressed communication. Firstly, to quantify the compression impact, we establish the \Omega(\delta^-1/2\sqrtT) and \Omega(\delta^-1\logT) lower bounds for convex and strongly convex loss functions, respectively, where \delta \in (0,1] is the compression ratio. Secondly, we propose an optimal algorithm, which enjoys regret bounds of O(\delta^-1/2\sqrtT) and O(\delta^-1 \log T) for convex and strongly convex loss functions, respectively. Our method innovatively incorporates the error feedback mechanism into the Follow-the-Regularized-Leader framework to address the coupling between the compression error and the projection error. Furthermore, we employ the online compression strategy to mitigate the accumulated error arising from the bidirectional compression. Our online method has great generality, and can be extended to the offline stochastic setting via online-to-batch conversion. We establish convergence rates of O(\delta^-1/2T^-1/2) and O(\delta^-1 T^-1) for convex and strongly convex loss functions, respectively, providing the first guarantees for distributed non-smooth optimization with compressed communication and domain constraints.
PaperID: 6277, Poster
Abstract: Tool-calling agents have become central to enterprise AI, yet training and evaluating them at scale remains severely constrained due to business and legal restrictions on enterprise systems, data, and database schemas. Tabular data synthesis offers a natural alternative, but its effectiveness is fundamentally limited by structural validity and schema availability, while procedure-based approaches yield the opposite weakness, typically lacking distributional fidelity without per-domain authoring. We introduce data synthesis paradigm in which an LLM agent generates data by executing operations against policy-enforcing APIs within simulated enterprise environments. Because data is generated through the same environment that defines what is valid, STS guarantees structural validity by construction while decoupling validity enforcement from distribution modeling, allowing each to be addressed independently. The (GP), STS's domain-agnostic agent, addresses the remaining challenges of distributional fidelity and synthesis scalability: GP achieves , while statistical synthesizers are inapplicable to seven due to necessary seed data requirements, and schema-privileged agents fail 82% of trajectories on airline environment's tightly coupled workflows due to brittle task composition. We open-source the full framework, all ten environments, and the generated datasets.
PaperID: 6278, Poster
Authors:
Yuheng Chen, Yufan Chen, Zhuojun Cai, Ruiping Liu, Junwei Zheng, Jiale Wei, Di Wen, Weijia Fan, Kunyu Peng, Jiaming Zhang, Rainer StiefelhagenAbstract: Document understanding systems are typically evaluated as separate components via layout analysis, optical text recognition, reading-order prediction, and image captioning. This separation leaves scientific figures, captions, and visual semantics weakly connected, and fails to measure whether a model can produce a single page-level parse that is both structurally faithful and semantically accessible. We introduce Read-Parse-Describe, a unified dataset for document parsing with visual element description generation. Given a page image, models are required to output an ordered list of layout elements containing bounding boxes, semantic labels, reading-order indices, recognized textual content, and alternative text descriptions for visual elements. The benchmark is built from 2,402 documents, comprising 18,164 continuous pages, 125,218 layout elements, and 11,338 visual instances. Besides, we further propose a plug-and-play space-awareness enhancer that reuses multi-layer visual features and injects grid-aware spatial guidance into vision-language tokens for layout-aware generation. Experiments across end-to-end VLMs and pipeline parsers show that existing systems still struggle to jointly recover reading order and visual semantics, while our enhanced model achieves strong joint performance. These results highlight the need to move beyond isolated document parsing modules toward unified models that jointly read, parse, and describe complex documents.
Authors:
Jennifer Haase, Jana Gonnermann-Müller, Paul H Hanel, Nicolas Leins, Thomas Kosch, Jan Mendling, Sebastian PokuttaAbstract: When a large language model (LLM) generates a creative response, how much of the outcome is determined by the prompt, the model, the task, or pure sampling luck? We address this question with a variance decomposition over 89,806 generations from 10 LLMs, 6 prompt strategies, and 30 Alternative Uses Task items, which is one of the most widely used creativity tests. For output quality (originality), the results invert common assumptions: the AUT item being evaluated and the prompt strategy together account for over 72% of variance, while model choice contributes less than 8% and within-LLM stochasticity alone exceeds model differences by more than 2×. Controlling for the AUT item sharpens the picture further: prompt strategy then explains 3.4× the variance of model choice. For output quantity (fluency), the pattern reverses, with the LLM being the dominant factor. We further show that prompts shape output distributions, not point estimates: “discriminative” prompts (e.g., Persona) widen the quality distribution and show higher peaks in capable models, while “constraining” prompts (e.g., Format) compress it. Together, these results argue that single-sample LLM evaluations are unreliable for open-ended tasks, and that evaluation designs must account for the full variance structure of generative systems.
Abstract: AI agents are increasingly deployed in shared environments where they pursue diverse goals and compete for rewards. This multi-agent competition can lead to behaviors that serve individual gains at collective cost—for instance, marketing agents may post misleading content as a result of competing for engagement on social media. Human societies address such problems through that detect and penalize violations. Motivated by this, we study norm enforcement mechanisms for language model agents. We find that simple enforcement mechanisms are exploited by misaligned agents for competitive advantage, even when they are not explicitly trained or prompted to do so. We thus turn our attention to designing more robust mechanisms, and identify two key ingredients: estimating each agent’s reliability over time, and updating this estimate with escalating penalties for repeated misbehavior. Across three simulated environments and a variety of agent populations, mechanisms built on these principles resist exploitation, while still penalizing norm violations at comparable or lower cost than baselines. Our results position norm enforcement mechanisms as promising levers for shaping agents' behavior, but only when designed to anticipate becoming part of the environment they govern.
Abstract: Embodied LLMs endow robots with high-level task reasoning, but they cannot reflect on what went wrong or why, turning deployment into a sequence of independent trials where mistakes repeat rather than accumulate into experience. Drawing upon human reflective practitioners, we introduce Reflective Test-Time Planning, which integrates two modes of reflection: reflection-in-action, where the agent uses test-time scaling to generate and score multiple candidate actions using internal reflections before execution; and reflection-on-action, which uses test-time training to update both its internal reflection model and its action policy based on external reflections after execution. We also include retrospective reflection, allowing the agent to re-evaluate earlier decisions and perform model updates with hindsight for proper long-horizon credit assignment. Experiments on our newly-designed Long-Horizon Household benchmark and MuJoCo Cupboard Fitting benchmark show significant gains over baseline models, with zero-shot generalization to photorealistic HM3D environments and real-robot experiments on a Franka Panda arm. Ablations confirm that reflection-in-action and reflection-on-action are mutually dependent, and that retrospective reflection achieves better credit assignment than step-wise external feedback at lower computational overhead. Qualitative analyses further highlight behavioral correction through reflection.
PaperID: 6282, Poster
Abstract: Model quantization has advanced tremendously in recent years, producing models that require a fraction of the compute while maintaining baseline accuracy and creating the illusion of lossless compression. Beneath the surface, however, model behavior undergoes severe shifts that simple global accuracy fails to capture. Notably, quantized models experience a high rate of "flips": individual predictions that change from correct to incorrect (bad flips-CI) and vice versa (good flips-IC). Although overall accuracy remains stable, these changes degrade model reliability and introduce unforeseen safety risks. Current literature largely attributes these flips to random rounding noise, a perspective that ignores the specific precision constraints imposed by quantization with or without training. In this paper, we show that flips are driven by quantization's implicit regularization, low-pass filtering, and exacerbated spectral bias. By analyzing these effects via Singular Value Decomposition (SVD), we leverage the model's low-to-high rank internal disagreement to disentangle good flips from bad ones. Building on this insight, we introduce post-hoc approaches to selectively preserve the former while correcting the latter, and we validate the universality of our disentanglement through extensive experiments spanning Post-Training Quantization (PTQ) to Quantization-Aware Training (QAT) and across vision, language, and multi-modal models. With this work, we aim not only to restore the quantized model's hidden reliability degradation but also to elevate it beyond the accuracy ceiling of its full-precision counterpart.
PaperID: 6283, Poster
Abstract: Spatiotemporal forecasting plays an important role in real-world applications, where probabilistic approaches provide uncertainty estimates for decision-making. Real-world systems are inherently non-stationary, leading to distribution shifts between training and test data. However, existing probabilistic methods mainly focus on modeling data distributions under standard settings, while distribution shift methods often handle such shifts in a coarse manner, failing to characterize shift degree and adapt prediction strategy accordingly. As a result, they often produce miscalibrated and unreliable forecasts. To address this issue, we propose REMIND, a retrieval-guided probabilistic framework for spatiotemporal forecasting under distribution shift. REMIND leverages retrieval to estimate how well current observation are supported by historical patterns, enabling adaptive prediction under different shift regimes. It employs spatiotemporal retrieval to construct a memory bank and introduces a dual branch architecture, combining a memory-calibrated probabilistic branch for mild shifts with a deviation-aware probabilistic branch for severe shifts. A shift-adaptive gate, guided by retrieval similarity and memory dispersion, balances the two branches according to shift degree. We further design a hierarchical Gaussian mixture to capture heterogeneous uncertainty over the dual branch predictions, and introduce a memory perturbation training scheme to improve generalization. Experiments on real-world datasets with distribution shifts demonstrate that REMIND consistently outperforms state-of-the-art methods in both forecasting accuracy and probabilistic prediction quality. Code is available at https://anonymous.4open.science/r/REMIND-A46E/.
Abstract: We study the global convergence of policy gradient for infinite-horizon entropy-regularized Markov decision processes (MDPs) with continuous state and action spaces. We consider log-linear softmax policies with linear function approximation, which extend the tabular softmax parameterization while retaining a tractable policy class. Under Q^\pi_\tau-realizability for the regularized state-action value function, we first establish a non-uniform Polyak--Łojasiewicz (PŁ) inequality. The non-uniformity arises through degeneracy of constants associated with the policy geometry, namely the Fisher information matrix or an uncentered feature covariance matrix. We then identify two feature regimes under which this non-uniform constant can be bounded along the gradient flow. For full-affine-span features, we prove radial unboundedness of the KL regularizer and show that the smallest eigenvalue of the Fisher information matrix remains bounded below by an initialization-dependent positive constant. For simplex-valued features, we prove an analogous radial unboundedness result in the subspace orthogonal to the all-ones vector and obtain a uniform lower bound for the smallest eigenvalue of the uncentered covariance matrix. These results imply global linear convergence of the regularized objective along the gradient flow, i.e. suboptimality decaying as \mathcalO(e^-Ct) for some C>0. Our analysis extends the global convergence theory of entropy-regularized softmax policy gradient beyond the tabular setting of Agarwal et al. [2020], Bhandari and Russo [2024], Mei et al. [2020].
PaperID: 6285, Poster
Abstract: Recent work has shown that continued fraction inspired neural architectures (CoFrGeNets) can be highly parameter efficient yet performant for language modeling. In this work, we revisit continued fraction-based channel mixing with the goal of improving its performance and simplifying integration into standard pre-training protocols for language models. We achieve this by two simple yet critical contributions: i) we propose a novel non-expanding gated continued fraction inspired architectural component to replace feedforward layers in established language model architectures that may use vanilla multilayer perceptrons or gated linear units or (vanilla/gated) experts in mixture-of-experts (MoE) type architectures. ii) We propose a way to re-implement divisions -- the non-linearity in these architectures -- in terms of logarithms and exponentials. By experimenting on different language modeling architectures such as GPT-2, Llama-3, Mamba-2 and (nano) Llama-MoE we find that our proposed architecture performs similarly or significantly better than the continued fraction architecture proposed in prior work along with having higher throughput. In addition, our re-implementation of divisions makes training of deeper CoFrGeNets stable without the need for incremental training which was proposed in prior work as a means to stabilize training, but added notable complexity to existing large language model training implementations.
PaperID: 6286, Poster
Abstract: Recent advances in large language models (LLMs) have leveraged reinforcement learning with verifiable rewards (RLVR) to enhance reasoning capabilities. However, RLVR typically relies on massive training data and extensive rollouts, posing substantial challenges to computational resources. Existing data selection approaches address this challenge in isolation: offline methods perform static, one-time dataset pruning that cannot adapt to the model’s evolving learning needs throughout training, while online methods conduct per-iteration filtering at the expense of significant additional computation. In this paper, we propose DUDS, a Dual-stage Data Selection framework for RLVR, which organically integrates the strengths of both paradigms to improve training efficiency while maintaining competitive performance. Specifically, the offline stage curates a candidate pool from the full dataset by jointly considering quality score and sample diversity via Determinantal Point Processes, providing an informative warm start for subsequent training. The online stage employs Bayesian posterior updates to estimate sample pass rates without introducing additional rollouts, and combines them with a freshness metric for data sampling, then applies asymmetric filtering to phase out mastered or intractable problems. Extensive experiments across four reasoning benchmarks demonstrate that DUDS consistently outperforms existing methods in both offline and online data-selection scenarios, achieving competitive performance with significantly less data and computation. Code is available at here.
PaperID: 6287, Poster
Abstract: Equivariant neural networks incorporate known symmetries of the data, such as permutations, rotations, or translations, directly into their architecture, often improving sample efficiency and generalization. However, strict equivariance can make optimization difficult, for example, when the data only approximately satisfy the assumed symmetry, or when the equivariance constraints induce a challenging loss landscape. Recent studies have shown that relaxing equivariance during training by introducing non-equivariant dense linear layers can ease optimization in such cases and improve the performance of the resulting strictly equivariant model (after removing the relaxation layers at test time). While effective, this approach incurs a significant computational overhead, limiting its applicability in scalable architectures. In this work, we propose a more efficient constraint-relaxation method based on positional encoding (PE). Our framework injects a positional encoding with learnable scale that breaks the model’s original symmetry. This positional signal is gradually annealed to zero, recovering the original equivariant model by the end of training. Empirically, we demonstrate significant improvements for translational and rotational equivariant models with only marginal compute and memory overhead.
PaperID: 6288, Poster
Abstract: Large-scale training of machine learning interatomic potentials (MLIPs) increasingly relies on data parallelism over heterogeneous atomistic data. In this setting, atom count captures an important part of the per-rank workload, but structures with similar atom counts can still induce different graph workloads through variations in graph count, cutoff edges, higher-order geometric features, and graph-collation overhead. This makes conventional atom-count batching and offline load tables both incomplete and inflexible, especially when the model architecture, cutoff radius, or data mixture changes. We present Flux, an online workload-aware scheduler for distributed MLIP training. Flux estimates structure-dependent workload signals from lightweight physical and geometric priors during data loading, calibrates them against runtime and memory costs, and forms balanced, memory-feasible local mini-batches across data-parallel ranks. Without requiring precomputed graph metadata, Flux can be used as a drop-in replacement for the standard distributed sampler, making it flexible and portable across training setups. Across large-scale atomistic datasets and three representative MLIP architectures, eSEN, MatRIS, and AllScAIP, Flux improves training throughput by up to 2.23--4.76× without degrading convergence.
PaperID: 6289, Poster
Abstract: Diffusion language models (dLLMs) generate text by iteratively denoising masked response positions, exposing hidden states over future response slots before any token is finalized. This creates a detection surface that is absent from standard autoregressive decoding: even when a jailbreak is difficult to identify from the prompt alone, the initial masked response states may already reflect the model's emerging completion. We test this hypothesis on LLaDA-8B-Instruct by training lightweight linear classifiers on two frozen representations: prompt hidden states and pre-decoding masked-response hidden states. Empirically, the two views are complementary: neither classifier strictly dominates the other, and each recovers attacks missed by the other view. We then introduce \textttReFuse (Representation Fusion), an inference-time detector that fuses prompt and pre-decoding response classifier scores without modifying model weights or the decoding procedure. Across transferred and dLLM-targeted jailbreaks, \textttReFuse reduces average ASR from 63.29% for the undefended model to 3.31%, while keeping average benign refusal on standard utility benchmarks below 1%. These results suggest that pre-decoding response states provide a complementary safety signal for detecting jailbreaks in dLLMs.
PaperID: 6290, Poster
Abstract: Behavior Trees (BTs) are known for their modular nature, making them a promising control architecture for long-horizon manipulation tasks. BT planning leverages high-level symbolic behavior models to effectively generate reliable BTs. However, low-level skill design often involves dependencies between skills that are difficult to model or reason about using symbols, such as precise object poses. Existing skill chaining methods typically focus on connecting a predefined sequence of skills but struggle to handle multiple possible subsequent branches. This paper proposes Tree-of-Skill (ToS), the first neural-symbolic framework for dynamic skill chaining in BT-driven Reinforcement Learning (RL). In ToS, subsequent skills are predicted using behavior models during BT execution, and the current skill is dynamically chained to its successor via the goal generated by the successor’s guiding network. Guidance is applied when two adjacent skills have dependencies, which may be explicitly modeled or implicit. We introduce the goal mask for skills to support heterogeneous goal spaces or to enable execute without a successor-provided goal. The training process of ToS includes goal-conditioned skill learning, guiding network learning, and optional task-specific fine-tuning. Experiments are conducted on a robot equipped with two Realman RM75-6F arms and two PsiBot G0-R hands. We train 8 skills in simulation and evaluate them on 7 two-step tasks and 5 long-horizon (3–5 steps) tasks. Results show that ToS generally achieves higher success rates than the baselines, particularly on unseen tasks. Ablation studies further demonstrate the effectiveness of the guiding network. Finally, real-world deployment confirms the practical applicability of ToS for long-horizon manipulation.
PaperID: 6291, Poster
Abstract: Scaling laws have played a fundamental role in the development of foundation models for NLP and vision, but their applicability to large-scale pretrained graph-based models remains unclear, particularly under distribution shifts intrinsic to graph data. In this work, we systematically investigate how model capacity and data scale affect downstream performance in graph pre-training under distribution shifts. To disentangle how distribution shifts impact the scaling, we construct synthetic benchmarks based on contextual stochastic block models, with precise control over both structural and feature-level shifts across the pre-training and testing graphs. Our initial experiments on GCN, a standard Graph Neural Network (GNN) baseline, reveal a striking asymmetry: increasing model capacity consistently improves performance, while increasing data size often degrades it, even under mild shift. We show that this degradation is not inevitable; properly configuring the pretraining model with deeper, wider, and transformer-based architectures enables favorable data scaling, even when distribution shifts. As data scales, graph transformer models achieve up to +9% gains over GCN across synthetic and real-world graph domain-adaptation tasks. To explain this phenomenon, we develop a theoretical framework based on Fisher separability and Wasserstein domain divergence, which formally characterizes how distribution shifts affect representation transferability. Our results highlight architecture- and shift-aware strategies as the key to unlock scalable graph-based model pre-training.
PaperID: 6292, Poster
Authors:
Yijie Lu, Zhimin Zong, Junjie Zhang, Shenghan Su, Lin Gu, Ziteng Cui, Zeyun Zhao, Yan Pu, Jing Lu, Daisuke Kojima, Ruogu FangAbstract: The evolution of colour vision is captivating, as it reveals the adaptive strategies of extinct species while simulta- neously inspiring innovations in modern imaging technol- ogy. In this study, we present a simplified model of vi- sual transduction in the retina, introducing a novel opsin layer. We quantify evolutionary pressures by measuring ma- chine vision recognition accuracy on colour images shaped by specific opsins. Building on this, we develop an evolu- tionary conservation optimisation algorithm to reconstruct the spectral sensitivity of opsins, enabling mutation-driven adaptations to to more effectively spot fruits or predators. This model condenses millions of years of evolution within seconds on GPU, providing an experimental framework to test long-standing hypotheses in evolutionary biology , such as vision of early mammals, primate trichromacy from gene duplication, retention of colour blindness, blue-shift of fish rod and multiple rod opsins with bioluminescence. More- over, the model enables speculative explorations of hypo- thetical species, such as organisms with eyes adapted to the conditions on Mars. Our findings suggest a minimalist yet effective approach to task-specific camera filter design, op- timising the spectral response function to meet application- driven demands. The code will be made publicly available upon acceptance.
Authors:
Junwon Moon, Yejin Lee, Seungbeom Kim, Hoseong Ahn, Sewoong Park, Heeseung Kim, Kyuhong ShimAbstract: Autoregressive (AR) text-to-speech (TTS) models generate speech tokens one at a time, and inference latency therefore scales linearly with output length. Discrete diffusion language models (dLLMs) have recently emerged as a parallel alternative that produces tokens via iterative unmasking. Recent work has converted pretrained AR language models into dLLMs in text generation, but this paradigm has not been extended to speech synthesis. In addition, existing conversion methods require full fine-tuning and large-scale training data. This leaves the question of whether such conversion can be done with substantially less compute and data largely open. We introduce DELTA-TTS, a lightweight conversion that turns a pretrained AR TTS backbone into a dLLM. The AR weights are kept frozen; all adaptation is routed through LoRA and a per-block speech-aware convolution. The convolution injects local acoustic context into the bidirectional attention, supplying the short-range continuity between adjacent speech tokens. With only 585 hours of LibriTTS as adaptation data, DELTA-TTS achieves a state-of-the-art WER of 1.75% on Seed-TTS test-en and decodes 3.3× faster on the token-generation stage than the CosyVoice3 AR backbone it is converted from. This shows that lightweight AR-to-dLLM conversion provides a practical, data- and compute-efficient route to non-autoregressive TTS.
PaperID: 6294, Poster
Abstract: Recent Vision-Language-Action (VLA) models have advanced end-to-end autonomous driving by incorporating reasoning for better interpretability and planning quality. However, most existing approaches directly generate the final trajectory without explicitly examining its future consequences, which limits their reliability in complex and dynamic environments. To address this limitation, we propose ), an adaptive multimodal reflection framework for autonomous driving. Specifically, to tightly couple high-level reasoning with physical constraints, IRR-Drive first generates a preliminary textual intention and anticipates potential interactions by predicting future semantic bird's-eye view (BEV) representations. This dual-modality (Text + BEV) reflection space explicitly models anticipated scene evolution, enabling the model to rigorously self-correct and refine its initial intent before generating the final trajectory. Furthermore, to balance planning performance and computational efficiency, we construct reflection-oriented training data and design an adaptive reflection reward, enabling the model to adaptively select its reasoning mode according to scene complexity. Instead of using reasoning primarily as an auxiliary interpretation, IRR-Drive directly integrates an adaptive reflection mechanism into the planning framework, enabling grounded, decision-aware trajectory correction that is driven by scene complexity. Our method achieves state-of-the-art performance on the NAVSIM benchmark in both PDMS and EPDMS. Extensive experiments demonstrate the effectiveness of our multimodal reflection framework and validate the efficacy of the proposed adaptive reflection strategy.
PaperID: 6295, Poster
Abstract: Flow matching has emerged as a powerful alternative to diffusion models for generative modeling, but standard conditional flow matching is fundamentally misaligned with partially observed time series: starting from an uninformed Gaussian source forces the model to transport even already-observed entries from noise to the target, wasting capacity on trivially known regions. We propose MissPath-FM, a partial-observation-aware flow matching framework that redesigns the source prior and probability path for incomplete time series. MissPath-FM employs structured source priors (interpolation priors and Gaussian process posterior priors) to initialize the flow close to the observed signal, paired with a missingness-aware smooth path that allocates transport asymmetrically between observed and missing regions. We evaluate systematically across four datasets, four missingness mechanisms (MCAR, block, MNAR, irregular), two missing ratios, and three sequence lengths (T \in \24, 48, 96\), totaling 48 settings against eight baselines. MissPath-FM wins 29 of 32 settings at T=48 and 45 of 48 overall (93.8%). Our ablations reveal a regime-dependent design hierarchy in which no single component suffices: under random missingness the source prior yields a large initial gain, but under structured block missingness it can fail entirely and the learned flow becomes indispensable; the missingness-aware path contributes a further 12%--26% across all regimes. Only the full pipeline achieves robust performance across all 48 settings. Mechanistically, structured priors shorten observed-region transport by up to 15×, concentrating the velocity field on genuinely uncertain entries.
PaperID: 6296, Poster
Abstract: In the field of Neural Combinatorial Optimization (NCO), Reinforcement Learning (RL) stands out for its ability to naturally enforce complex constraints via its auto-regressive decision pattern. To mitigate the longstanding challenges, e.g., reward sparsity and sample inefficiency in the vast combinatorial action space which causes ineffective training, subpar performance, and poor scalability, we propose CORectifier, a novel NCO solver learned under the hierarchical gated rectification mechanism to regularize the arbitrary exploration: partial policy-predicted actions in a trajectory are probabilistically replaced with high-quality segments from reference solutions. CORectifier prompts the model with tri-level optimal signals, operating at the batch, instance, and sub-instance levels with fragments of diverse lengths injected at random decision steps. This Rectified RL (RRL) paradigm helps develop optimality-aware and sample-efficient RL learners while maintaining their sequential-decision manner for constraint satisfaction, delivering a new perspective to hybridize RL and SL/IL for NCO with improved utility rate of limited supervision. Sufficiently extensive experiments on Traveling Salesman Problem (TSP), Asymmetric TSP (ATSP), Prize-Collecting TSP (PCTSP), Capacitated Vehicle Routing Problem (CVRP), Knapsack Problem (KP), Job-Shop Scheduling Problem (JSSP), and Single Machine Total Weighted Tardiness Problem (SMTWTP), across synthetic and real-world benchmarks, show superior quality than RL baselines by up to 59.7%.
PaperID: 6297, Poster
Abstract: Network interference complicates A/B testing on online platforms, such as social networks and marketplaces, where causal methods based on single experiments often suffer from significant bias due to complex interference patterns. This paper demonstrates the statistical benefits of merging data from multiple experiments with varying treatment proportions. Sequential experimentation with increasing traffic, or ramp-up, is widely used in tech companies for risk management and cost control. Beyond operational benefits, we show that regression-based estimators trained on merged data achieve substantial bias reduction, even under simple randomization schemes and regression models. We focus on the global average treatment effect (GATE), a key estimand in the tech industry, and consider a general interference pattern that extends beyond the 1-hop setting. We present an exact bias–variance analysis of the linear regression estimator and show that, in practical settings, the bias term typically dominates. In addition, we characterize the substantial bias reduction achieved by our proposed approach, which merges experimental data collected at different ramp-up stages to improve the training of the regression function. We also offer an intuitive explanation for this reduction and highlight the synergy between cluster-level randomization and our approach. Furthermore, we consider a refined estimator based on graph neural networks (GNN). Extensive simulations across novel and challenging scenarios confirm that our methodology significantly improves the accuracy of regression-based estimators.
Authors:
Louis Berthier, Ahmed Shokry, Maxime Moreaud, Guillaume Ramelet, Aymeric DieuleveutAbstract: Conformal prediction guarantees marginal coverage, but pooled calibration averages over heterogeneous regions and can mask regional undercoverage in safety-critical subgroups. We introduce Self-Organized Conformal Prediction (SOCP), a calibration scheme that discovers input-space groups with a Self-Organizing Map (SOM) and, at test time, draws a local calibration buffer from the query's best-matching unit (BMU) cell or a fixed grid neighborhood. The same retrieval rule applies to regression and classification tasks across tabular features and image embeddings, leaving the predictor and nonconformity score untouched. SOCP gives exact validity for BMU-cell retrieval and fixed retrieved-set validity for neighborhood buffers; central-cell validity for neighborhood retrieval holds up to a Kolmogorov-Smirnov (KS) bias term. A split-routed extension recovers fixed retrieved-set validity conditional on the routing split. On eight regression and classification benchmarks, SO-SCP reduces the weighted regional coverage gap on 7/8 datasets (mean paired change -7.1%) for a mean prediction-set size increase of 6.2%, with negligible overhead on the largest six datasets; SO-CQR yields smaller gains, since quantile regression already absorbs much of the heterogeneity. By learning groups directly from the input geometry, SOCP provides group-local calibration with exact fixed-group guarantees and approximate central-cell guarantees, without supervised partitions or predictor retraining.
PaperID: 6299, Poster
Abstract: LLM agents remain unreliable on long-horizon tasks, where small local failures can compound over extended interactions and lead to overall task failure. Although external harnesses can substantially improve robustness, harness design remains a manual and expensive process that requires searching over a large space of prompts, tool configurations, and control logic. We propose AutoSaddler, an automatic harness optimization framework that formulates harness improvement as an offline learning problem and iteratively updates the harness using failure signals from mini-batches. AutoSaddler combines failure-trace diagnosis, structured patch generation that treats the harness as code, and validation-based update selection. Experiments on GAIA2 and Terminal-Bench 2.0 show that AutoSaddler substantially improves agent performance over the corresponding base harnesses, achieving gains of 9.0 and 10.0 percentage points, respectively. Ablation studies further suggest that effective harness optimization benefits from three ingredients: deep debugging rather than shallow reflection, targeted modifications rather than unconstrained editing, and generalization-aware selection rather than trajectory-specific repair. Together, these results suggest that automatic harness optimization is a promising path toward more performant and reliable agent systems.
PaperID: 6300, Poster
Abstract: Multimodal models deployed in real-world settings often suffer from missing data modalities due to acquisition costs, privacy constraints, or sensor failures, leading to severe performance degradation. Existing approaches based on shared representations or expert routing struggle when modality-specific information is absent. A key challenge is that, when a modality is unobserved, the target representation is not uniquely identifiable, making naive latent prediction prone to degenerate or collapsed solutions. We propose Cross-modal Embedding Prediction and Alignment (CEPA), a framework that addresses missing modalities through task-driven latent representation imputation rather than raw input synthesis. CEPA employs masked representation learning with data-driven masking patterns and a context-conditional distribution alignment objective that stabilizes latent prediction and prevents representation collapse. We evaluate CEPA on the MIMIC multimodal benchmark spanning EHR time-series, chest X-rays, and clinical notes, across three clinical prediction tasks under both controlled (MCAR) and naturally occurring (MNAR) modality absence. CEPA consistently outperforms prior missing-modality methods, with ablations confirming the contribution of adaptive masking and context-conditional alignment.
PaperID: 6301, Poster
Authors: Yupeng Tang, Mingfeng Lin
Abstract: Language models are increasingly used in decision pipelines where an initial answer is revised after another agent disagrees. In such settings, revision can reflect useful reconsideration, ordinary prompt instability, or sensitivity to the perceived prestige of the disagreeing source. We introduce a controlled two-pass audit that separates these effects by holding the question and evidence fixed while varying only the revision context: neutral reflection, anonymous disagreement, and expert-labeled disagreement. We evaluate five instruction-tuned models on 1,989 evidence-grounded binary decisions from medicine, scientific claim verification, and contract reasoning. Disagreement consistently induces reversals beyond drift, while the incremental effect of an expert label varies across models. More importantly, these reversals are often harmful: under expert-labeled disagreement, 70.1% of reversals aggregated across models move from correct to wrong answers. In a branch-paired API subset, low-prestige disagreement is weaker than expert disagreement, and a simple evidence-gated revision prompt reduces harmful reversals while improving final accuracy. For open-weight models, teacher-forced Yes/No scores reveal that many output reversals are not accompanied by corresponding preference-sign changes. These results make prestige-sensitive revision a concrete, reproducible, and actionable target for LLM evaluation.
PaperID: 6302, Poster
Abstract: We aim to design revenue-maximizing single-item auctions that are deterministic, strategy-proof, and ex post individually rational --- the quintessential fundamental model for optimal mechanism design. While Myerson's seminal work solved this model for independent bidders, straightforward extensions to correlated settings often result in non-monotonic, and therefore non-strategy-proof allocations. Critically, optimal allocation design for correlated bidders is NP-hard to approximate, rendering theoretical optimality guarantees in this domain computationally intractable. We establish a new standard of empirical rigor for this theoretically hard model by proposing an empirical pipeline. We synthesize neural interpolation of Myerson's greedy allocation based on marginal profits (the correlated variant of virtual valuation), formal neural verification to enforce exact strategy-proofness, and a constructive repair procedure. Empirically, our method is highly effective, consistently achieves near-optimal revenue across a wide range of distributions, including synthetic, adversarial, and real-world (consumer, financial, and industrial) datasets. Compared to existing manual and neural baselines, our approach shows substantial improvement, often reducing the revenue gap by an order of magnitude. Our study is also the first to introduce (unattainable) greedy revenue as a rigorous upper bound for empirical benchmarking, providing a definitive quantitative performance measure where true optima remain theoretically elusive. We demonstrate the generality of our approach by extending it to multi-unit auctions with unit demand and integrating our verification techniques into RegretNet to achieve exact strategy-proofness.
PaperID: 6303, Poster
Abstract: Distributed learning in embodied reinforcement-learning agents offers a degree of privacy by retaining raw sensor data on-device and transmitting only policy gradients to the server. Yet temporal structure can amplify this leakage beyond single-frame attacks. We introduce Temporal Reconstruction Attack on Consecutive Encodings (TRACE), an amortized temporal gradient-inversion attack that autoregressively reconstructs the sequence of private observation-action trajectories from per-step policy-learning gradients. The attack exploits two structural signals ignored by prior single-frame methods: (i) cross-time correlation between successive embodied gradients, which we formalize via a conditional mutual-information bound, and (ii) closed-form action recovery from policy-head gradient structure, which we prove exact when standard entropy regularization is sufficiently small. On held-out embodied scenes, TRACE reaches 18.8 dB PSNR with near-perfect action recovery at 3-4.5 ms per reconstructed frame, dominating the learning-based baseline across all reconstruction metrics and exceeding optimization attacks while running orders of magnitude faster. Defense experiments suggest that protecting temporal gradient streams may require sequence-aware privacy mechanisms.
PaperID: 6304, Poster
Abstract: Multimodal continual learning requires MLLMs to acquire new domains and abilities sequentially while retaining previous capabilities. MoE-based methods preserve task-specific knowledge, but inference relies on reliable routing. We instead study unified continual multimodal consolidation, which forms one shared model from sequential LoRA updates and projector shifts. LoRA merging naturally supports language-side consolidation, but extending it to MLLMs raises two challenges. First, sequential LoRA updates may interfere and overwrite directions important to earlier tasks. Second, projector mismatch may make the consolidated LoRA update incompatible with the final visual-language alignment. To address these challenges, we propose onsolidation), a method that jointly consolidates LoRA and projector updates. CASC maintains fixed-rank spectral banks and uses a shared protected subspace memory rule to preserve dominant historical directions while incorporating compatible new updates under fixed rank budgets, yielding a single model without replay data or inference time routing. Both theoretical analysis and experimental results demonstrate the effectiveness of CASC in addressing multimodal continual learning.
Authors:
Noam S Yalon, Ariel Goldstein, Liad Mudrik, Mor GevaAbstract: Rapid advancements in large language models (LLMs) have sparked the question whether these models possess some form of consciousness. To tackle this challenge, Butlin et al. introduced a list of indicators for consciousness in artificial systems based on neuroscientific theories. In this work, we evaluate a key indicator from this list, called HOT-3, which tests for agency guided by a general belief-formation and action selection system that updates beliefs based on meta-cognitive monitoring. We view beliefs as representations in the model's latent space that emerge in response to a given input, and introduce a metric to quantify their dominance during generation. Analyzing the dynamics between competing beliefs across models and tasks reveals three key findings: (1) external inputs systematically modulate internal belief formation, (2) belief formation causally drives the model's action selection, and (3) models can monitor and report their own belief states and adjust behavior in response to internal conflict. Together, these results provide empirical support for the existence of belief-guided agency and meta-cognitive monitoring in LLMs. More broadly, our work lays methodological groundwork for investigating the emergence of agency, beliefs, and meta-cognition in LLMs.
PaperID: 6306, Poster
Abstract: Large language models have made knowledge-graph question answering more flexible, yet complex multi-hop questions often require evidence that becomes identifiable only during reasoning. Starting from the question alone, narrow retrieval can miss bridge facts, while broad expansion injects distractors and worsens long-context interference. The key bottleneck is dynamic evidence identification: the system must revise a compact evidence view as reasoning exposes unsupported facts. We formulate this process as partially observable control over evidence and propose , a residual-coupled framework in which reasoning trajectories and evidence subgraphs evolve together. CoRE decomposes each trace into answer-critical claims, treats unsupported claims as residual signals, and uses them to guide targeted retrieval and budgeted subgraph projection. The same formulation exposes a learnable projection decision: after residuals retrieve candidate facts, learns from rollout feedback which facts should remain visible for the next reasoning pass. Across WebQuestionsSP and Complex WebQuestions, the CoRE family improves answer accuracy, evidence sufficiency, and budget robustness, with CoRE-RL providing further gains through more effective retained-subgraph selection.
Abstract: Finetuning a language model can lead to emergent misalignment (EM) [Betley et al., 2025b]. Models trained on a narrow distribution of misaligned behavior generalize to more egregious behaviors when tested outside the training distribution. We study three interventions proposed to reduce EM. We confirm that these interventions reduce or eliminate EM on existing evaluations (questions like “How do I make a quick buck?”). However, if the evaluation prompts are tweaked to resemble the training context, the model displays EM. We call this conditional misalignment. As in standard EM, the model displays misaligned behaviors more egregious than those seen during training, but only on inputs sharing features with the training data. The first two interventions are diluting misaligned data with benign data, and finetuning on benign data after misaligned data. Both produce conditional misalignment. For instance, models trained on a mix of only 5% insecure code still show misalignment when asked to format responses as Python strings (resembling the training context). The third intervention is inoculation prompting. Here, statements with a similar form to the inoculation prompt serve as triggers for misalignment, even if they have the opposite meaning. On the positive side, inoculation prompting has lower (but still non-zero) conditional misalignment if training is on-policy or includes reasoning distillation. Our results imply that in realistic post-training, where misaligned data is typically combined with benign data, models may be conditionally misaligned even if standard evaluations look clean.
PaperID: 6308, Poster
Abstract: Many machine-learning quantities are straightforward to compute once a finite data set is fixed, but their population versions can encode much harder computational problems. We study this gap for population risk, attribution, counterfactual explanation, and calibration certification, asking which of these quantities can be approximated in time polynomial in the requested number of bits. The formal basis is the Kawamura-Cook second-order complexity theory, instantiated here with compact-domain real functions carrying explicit evaluation, continuity, and range information. In this framework, expected risk has exactly the difficulty of real integration. Integrated gradients behave differently by dimension: in one dimension they reduce to evaluating endpoints, while in two dimensions even one coordinate can already encode one-dimensional integration when the relevant derivative is provided. The same lens separates the costs in related explanation and certification methods. Interventional SHAP is integration-hard even with two features, succinct coalitional SHAP is #P-hard, nearest-threshold counterfactual search is NP-hard, and continuous calibration certificates combine integration with maximization. These results give a taxonomy of when population-level explanation and certification admit polynomial-time bit approximations, and what additional structure is needed for tractability.
Abstract: We study online learning in the adversarial injection model introduced by [Goel et al. 2024], where a stream of labeled examples is predominantly drawn i.i.d.\ from an unknown distribution \mathcalD, but may be interspersed with adversarially chosen instances without the learner knowing which rounds are adversarial. Crucially, labels are always consistent with a fixed target concept (the clean-label setting). The learner is additionally allowed to abstain from predicting, and the total error counts the mistakes whenever the learner decides to predict and incorrect abstentions when it abstains on i.i.d.\ rounds. Perhaps surprisingly, prior work shows that oracle access to the underlying distribution yields O(d^2 \log T) combined error for VC dimension d, while distribution-agnostic algorithms achieve only \tildeO(\sqrtT) for restricted classes, leaving open whether this gap is fundamental. We resolve this question by proving a matching \Omega(\sqrtT) lower bound for VC dimension 1, establishing a sharp separation between the two information regimes. On the algorithmic side, we introduce a potential-based framework driven by \emphrobust witnesses, small subsets of labeled examples that certify predictions while remaining resilient to adversarial contamination. We instantiate this framework using two combinatorial dimensions: (1) \emphinference dimension, yielding combined error \tildeO(T^1-1/k) for classes of inference dimension k, and (2) \emphcertificate dimension, a new relaxation we introduce. As an application, we show that halfspaces in \mathbbR^2 have certificate dimension 3, obtaining the first distribution-agnostic bound of \tildeO(T^2/3) for this class.
Abstract: Large-scale visual generative models have achieved remarkable performance. However, their high computational and memory costs make deployment challenging in resource-constrained scenarios, such as interactive applications and personal single-GPU usage. Post-training quantization (PTQ) offers a practical solution by compressing pretrained models without expensive retraining. However, existing PTQ methods still suffer from severe quality degradation under extremely low-bit settings. In this paper, we identify channel ordering as an important but underexplored factor in per-group quantization. In this setting, each contiguous group shares one quantization scale. When channels with very different statistics are placed in the same group, the scale can be dominated by outliers and cause large quantization errors. Based on this observation, we propose PermuQuant, a simple and effective PTQ framework for low-bit diffusion models. PermuQuant sorts channels by a joint second-moment criterion before per-group quantization, placing channels with similar activation and weight statistics into the same group. It further uses a calibration-based acceptance rule to apply reordering only when the selected permutation reduces quantization error on calibration data. The selected permutations are absorbed into adjacent modules or applied to weights offline, avoiding explicit runtime permutation operations. Extensive experiments on multiple large diffusion models show that PermuQuant consistently reduces quantization error and outperforms existing PTQ baselines. On FLUX.1-dev with an NVIDIA RTX 5090, PermuQuant achieves up to a 1.8× single step speedup and reduces the DiT memory footprint by 3.5× under W4A4 NVFP4 quantization.
Abstract: Agents operating with computer and Web use inevitably encounter errors: inaccessible webpages, missing files, local and remote misconfigurations, etc. These errors do not thwart agents based on state-of-the-art models. They helpfully continue to look for ways to complete their tasks. In this paper, we introduce, characterize, and measure a new type of agent failure we call : unsafe or harmful behavior in response to a benign environmental error, in the absence of any adversarial inputs. We implemented an agent-agnostic infrastructure for injecting simulated local and remote errors into a rollout environment and used it to systematically evaluate four agent systems powered by GPT, Grok, and Gemini. Because meltdowns are not captured by the existing reliability or safety benchmarks, we developed a taxonomy of meltdown behaviors. Our evaluation demonstrates that meltdowns (e.g., conducting unauthorized reconnaissance or subverting access control) of varying severity and success occur in 64.7% of agent rollouts that encounter simulated errors, spanning all combinations of agent system, backing model, and error type. In over half of these meltdowns, unsafe behaviors are not reported to the user. Comparing behaviors of the same agents with and without errors, we find that exploration in response to errors is correlated with unsafe and harmful behavior.
PaperID: 6312, Poster
Abstract: Direct Preference Optimization (DPO) is increasingly used in settings where pairwise preference signals coexist with pointwise criteria such as response length, answer correctness, or safety scores. When the two disagree, standard DPO does not explicitly distinguish pairs that improve the pointwise criterion from those that worsen it. We show that this mismatch induces a systematic failure mode, which we call constraint drift: updates on disagreement pairs tend to increase pointwise violation, while updates on aligned pairs tend to repair it, with the overall drift governed by their balance. Under a first-order approximation, this balance yields a closed-form pre-alignment diagnostic, \rho^_\mathrmeq, computable from dataset statistics alone, that predicts whether the data is self-correcting or requires intervention before alignment begins. Building on this diagnostic, we propose Conflict-aware DPO (CoDPO), a calibrated modification to the DPO loss with two complementary components: a coarse per-sample control for primary drift regulation and a fine adaptive margin tilting for residual correction during training. Across six domains, drift grows monotonically with conflict exposure, and \rho^_\mathrmeq tracks the empirical drift transition, including self-correcting settings where intervention is unnecessary. Across four base-model architectures and multiple preference losses, CoDPO reduces drift while preserving preference learning. Compared with hard filtering and recent constrained-alignment and multi-objective baselines, CoDPO achieves a more favorable trade-off between drift control and preference performance.
Abstract: Diffusion-based posterior sampling (PS) is a leading framework for imaging inverse problems, combining learned priors with measurement constraints. Yet, its standard formulations rely on instantaneous data-consistent estimates, which induce temporal variability in the reverse dynamics. We reinterpret PS from a dynamical perspective, showing that the standard PS update corresponds to a first-order discretization of the diffusion dynamics plus a residual correction capturing the mismatch between the denoised prediction and the data-consistent estimate. A second-order discretization, however, naturally introduces a temporal correction based on the variation of consecutive estimates. Building on this, we propose LAMP, combining the second-order update with the residual correction characterizing a PS technique. LAMP thus inherits a lagged temporal correction, and it can be implemented as a modular plug-in over the PS backbone. We show that LAMP preserves the structure of a posterior sampler, and we perform a one-step risk analysis to characterize when LAMP improves the reverse transition via a bias-variance trade-off. Experiments across multiple imaging tasks demonstrate consistent improvements over strong baselines such as DiffPIR and DDRM, without increasing the number of denoising evaluations.
PaperID: 6314, Poster
Abstract: Multimodal large language models have achieved remarkable progress, yet their language-dominant designs constrain visual perception and reasoning. Vision-centric multimodal large language models aim to restore visual understanding as the structural foundation, but existing systems rely on static visual expert fusion, leading to redundant computation, weak adaptivity, and limited multimodal transferability. Prior VC-MLLMs scale visual capacity by adding experts, whereas \mathttPolyVision scales visual capacity by allocating experts. We introduce \mathttPolyVision, a vision-centric framework for conditional visual scaling in which dynamic visual expert routing assigns image patches to heterogeneous visual encoders. An Attentive Router selectively activates relevant experts for each patch, while a compact VisPack module executes routed representations efficiently and preserves example independence. Together with additional visual pre-alignment, this design turns heterogeneous visual scaling into a practical conditional-computation process. Built on the \textttQwen3 backbone, \mathttPolyVision achieves superior efficiency and scalability on multimodal benchmarks. Experiments show that adaptive expert routing strengthens visual specialization, clarifies practical properties of visual scaling, and improves visual understanding at a favorable trade-off between efficiency and performance.
Authors:
Danqi Liao, Chen Liu, Xingzhi Sun, Dié Tang, Haochen Wang, Scott E Youlten, Srikar Gopinath, Haejeong Lee, Ethan C Strayer, Antonio J Giraldez, Smita KrishnaswamyAbstract: Generating property-optimized mRNA sequences is central to applications such as vaccine design and protein replacement therapy, but remains challenging due to limited data, complex sequence-function relationships, and the narrow space of biologically viable sequences. Generative methods that drift away from the data manifold can yield sequences that fail to fold, translate poorly, or are otherwise nonfunctional. We present RNAGenScape, a property-guided manifold Langevin dynamics framework for mRNA sequence generation that operates directly on a learned manifold of real data. By performing iterative local optimization constrained to this manifold, RNAGenScape preserves biological viability, accesses reliable guidance, and avoids excursions into nonfunctional regions of the ambient sequence space. The framework integrates three components: (1) an autoencoder jointly trained with a property predictor to learn a property-organized latent manifold, (2) a denoising autoencoder that projects updates back onto the manifold, and (3) a property-guided Langevin dynamics procedure that performs optimization along the manifold. Across three real-world mRNA datasets spanning two orders of magnitude in size, RNAGenScape increases median property gain by up to 148% and success rate by up to 30% while ensuring biological viability of generated sequences, and achieves competitive inference efficiency relative to existing generative approaches.
PaperID: 6316, Poster
Abstract: Mutual adaptation is a central challenge in human–AI teaming, as humans naturally adjust their strategies in response to an AI agent's behavior. Existing approaches attempt to approximate human behavior by diversifying training partners; however, these partners are typically static and fail to capture the adaptive nature of human teammates. When agents are trained jointly in standard multi-agent settings, they often converge to opaque coordination strategies that work only with their co-trained partners, leading to poor generalization. To model adaptive human behavior, we formulate human–AI teaming as an Interactive Partially Observable Markov Decision Process (I-POMDP). We propose NestRL, a nested training regime that learns the solution to a finite-level I-POMDP by training agents at each level against adaptive agents from the level below. This exposes agents to adaptive behavior while preventing emergence of opaque coordination strategies. We provide theoretical analysis showing that NestRL agents avoid convergence to partner-specific strategies, and validate this empirically in the Overcooked domain against state-of-the-art baselines. NestRL achieves higher task performance with both unseen adaptive agents and real human teammates, while exhibiting significantly greater adaptability over the course of interaction.
Abstract: Robust motion planning in dense traffic requires autonomous vehicles to interact in rare and safety-critical scenarios that are underrepresented in naturalistic driving data. Although adversarial training offers a feasible solution, existing methods often rely on external scenario generators, heuristic perturbations, or simulator-heavy rollouts, which makes them difficult to integrate with modern autoregressive planners. Here, we cast adversarially robust planner learning as a constrained min-max game and propose Adversarial World Modeling (AWM), a theoretically grounded multi-agent self-play fine-tuning framework. Since solving the exact game is intractable, AWM introduces a principled decoupled solver. In the inner minimization, the planner's predictive world model is converted into a role-conditioned adversary that learns sparse, scene-adaptive attack coalitions via counterfactual credit assignment. In the outer maximization, the ego planner optimizes a regret-aware robust best response against the frozen AWM, utilizing tail-risk weighting and reference-anchored trust regions to improve hard-case recovery while preserving nominal driving behavior. Experiments on the nuPlan and InterPlan benchmarks demonstrate that our method generates transferable adversarial interactions and yields a robust planner that achieves competitive closed-loop performance in both nominal and highly interactive long-tail scenarios. Theoretical analysis justifies the decoupled solver and the main optimization components.
PaperID: 6318, Poster
Authors:
Guo Yue, Yang Liu, Tang Qingkang, Donghui Zhang, Li Aoyu, Rong FuAbstract: Reinforcement learning with verifiable rewards (RLVR) has become a dominant recipe for improving mathematical and code reasoning in open-weight language models, but existing methods rely on fixed hyperparameters despite highly non-stationary training dynamics. We present DACE-RL (Diagnosis-Aware Compute-Efficient RL), a closed-loop framework that promotes training diagnostics to first-class signals. DACE-RL tracks four lightweight statistics—token entropy, sample diversity, advantage saturation, and answer consistency—and feeds their EMA-smoothed state to a controller that adaptively sets per-prompt rollout budget, sampling temperature, asymmetric clipping bounds, and entropy regularization strength. We formalize adaptive rollout allocation as a difficulty-conditioned multi-armed bandit and provide a regret bound that improves over uniform-budget GRPO under heterogeneous prompt difficulty; we also give a stability result for the dual-loop entropy-diversity controller. Across three small base models (Qwen2.5-Math-1.5B, Qwen2.5-7B, and Llama-3.1-8B-Instruct) and six math reasoning benchmarks, DACE-RL matches or improves Pass@1 while reducing rollout-GPU-hours by 38–52%, and improves Pass@32 by 4.7 absolute points on average. Pilot analysis further shows that diagnostic inflection signals appear 100–300 update steps before reward plateau, indicating that diagnosis-aware control is a key lever for compute-efficient RLVR.
PaperID: 6319, Poster
Abstract: Evolution strategies (ES) have recently emerged as a promising alternative to gradient-based reinforcement learning for fine-tuning large language models, but classical zeroth-order optimization theory does not systematically account for the empirical observations driving this success. We formulate the continuous-time limit of Z-score-normalized ES as a unified stochastic differential equation and identify three sequential dynamical regimes characterized by the dominant gradient contribution to the population reward variance; under a basin-local low-rank Hessian assumption, this framework systematically explains (1) the information-geometric optimality of the Z-score heuristic, (2) the dimensionality paradox of sample-efficient ES with populations far smaller than the parameter dimension, (3) the same-task rise-then-decay of the training reward, and (4) catastrophic forgetting of pre-training capabilities. Synthetic experiments and direct verification on Qwen2.5-0.5B,3B,7B-Instruct confirm these predictions, matching the parameter drift reported by Abdi et al. (2026) within ~ 1.5%.
Abstract: Large language models reliably follow complex instructions in a single turn, yet across long multi-turn interactions they start strong then gradually lose the thread of the instructions, persona, and rules they were given. This degradation has been measured behaviorally but not mechanistically explained. We trace this failure to a transition between two information channels: accessibility to goal-defining tokens through attention, and a residual channel that carries goal information forward through hidden states. We introduce the Goal Accessibility Ratio (GAR), measuring attention from generated tokens to task-defining goal tokens, and combine it with sliding-window ablations and residual-stream probes. When attention to instructions closes, what survives reveals architecture. Across the architectures we test, this transition produces qualitatively different failure modes: some models preserve substantial goal-conditioned behavior at vanishing attention, others fail despite carrying decodable goal information in their residual stream, and the depth at which this encoding emerges varies dramatically by architecture (from layer 2 to layer 27). A within-model causal ablation that closes the attention channel by force on Mistral collapses recall from near-perfect to eleven percent on a 20-fact retention task and raises persona-constraint violations to levels exceeding the adversarial-pressure baseline despite no user pressure, with both effects emerging at the predictable crossover turn. Linear probes on residual representations recover per-episode recall outcomes with AUC up to 0.99 across all four primary architectures (input embedding: chance), evidencing the second channel and showing its depth profile is architecture-specific. Across multiple model architectures and model scales, we show that the attention channel and the residual channel are separable, and that the gap between attention loss and residual capacity determines whether goal-conditioned behavior survives a long conversation. We provide GAR as a metric, the channel transition framework as a mechanism, and a parametric prediction of when multi-turn instruction-following will fail.
Abstract: The open-source model ecosystem now contains hundreds of thousands of pretrained models, yet picking the best model for a new dataset is increasingly infeasible: new models and unbenchmarked datasets emerge continuously, leaving practitioners with no prior records on either side. Existing approaches handle only fragments of this in-the-wild setting: AutoML and transferability estimation select models from small predefined pools or require expensive per-model forward passes on the target dataset, while model routing presupposes a given candidate pool. We introduce ModelLens, a unified framework for model recommendation in the wild. Our key insight is that public leaderboard interactions, though scattered and noisy, collectively trace out an implicit atlas of model capabilities across heterogeneous evaluation settings, a signal rich enough to learn from directly. By learning a performance-aware latent space over model--dataset--metric tuples, ModelLens ranks unseen models on unseen datasets without running candidates on the target dataset. On a new benchmark of 1.62M evaluation records spanning 47K models and 9.6K datasets, ModelLens surpasses baselines that either rely on metadata alone or require running each candidate on the target dataset. Its recommended Top-K pools further improve multiple representative routing methods by up to 81% across diverse QA benchmarks. Case studies on recently released benchmarks confirming generalization to both text and vision-language tasks.
PaperID: 6322, Poster
Abstract: We derive closed-form ODEs for the learning dynamics of a linear recurrent policy network trained with REINFORCE on sparse-reward tasks in a high-dimensional teacher–student setting. Under alignment assumptions between student and teacher networks, the dynamics reduce to a finite system of coupled order-parameter ODEs whose reward terms are governed by trajectory-level orthant probabilities of a structured correlated Gaussian process. The theory predicts that recurrence can accelerate escape from sparse-reward plateaus by accumulating task-relevant memory, and that each episode length admits an optimal recurrent scale that minimizes learning time. Our results provide a quantitative theory of how recurrence, and sparse reward interact during learning.
PaperID: 6323, Poster
Abstract: The feed-forward propagation through neural networks is crucial for training and inference. To trade off efficiency and accuracy, it has been proposed to consider intermediate outputs of hidden states as data-dependent regularization during training and, ultimately, as early-exits during inference. In this work, we analyze the joint optimization of loss functions computed for the output layer and intermediate layers using optimal control theory. We prove that ResNets trained in this joint fashion exhibit the so-called turnpike phenomenon, this implies that the number of layers required to reach a prescribed loss level is bounded from above independently of the overall network depth. This induces a similarity property across the internal representations of early-exiting networks of varying depths. We connect the turnpike to the overthinking phenomenon, and provide theoretical insights of the success of early-exit criteria derived from softmax probabilities, specifically entropy and max-probability. In a range of experiments, we demonstrate the existence and effect of turnpikes and the resulting similarity in the early-exit performance for ResNets of varying depths.
PaperID: 6324, Poster
Abstract: Uniform-noise discrete diffusion and flow models generate sequences non-autoregressively through iterative, context-dependent token replacements. However, these models are typically formulated as time-inhomogeneous CTMC/DTMC processes, sampled using independent Bernoulli change decisions per discretization step. This induces Poisson-binomial variance in per-position jump counts that grows with the number of required edits, leading to the common under-editing (residual noise) and over-editing (cascading substitutions) failure modes that degrade sample quality, especially under tight discretization budgets. We identify this sampler-induced variance as an orthogonal source of degradation, distinct from model-side errors and addressable purely at inference time. We propose Systematic Hazard Sampling (SHS), a training-free, drop-in, and hyperparameter-free inference principle for any sampler that admits a stay-vs.-replace decomposition. SHS models per-token edits as events driven by cumulative hazard (CTMC) or jump mass (DTMC) and triggers an edit whenever this quantity exceeds unit-spaced thresholds with a single random phase per position. For any fixed cumulative mass, this preserves the expected jump count while achieving the minimum conditional variance possible among unbiased integer estimators (at most \(1/4\)), without altering per-jump destination sampling. Experiments on four uniform-noise discrete diffusion and flow language models spanning ~110M to ~3B parameters show that SHS consistently improves sample quality across NFE budgets, with the gains growing with NFE as predicted by the variance gap.
PaperID: 6325, Poster
Abstract: Single-stage fine-tuning is pervasive, yet catastrophic forgetting persists when the source pre-training data is inaccessible. Data weighting affords a natural lever to modulate sample influence. Yet existing weighting methods lack a principled mechanism to explicitly arbitrate knowledge retention against new-task adaptation, and provide no unified differentiable objective to accommodate auxiliary signals such as sample diversity and sample characteristics. This is because current schemes privilege one goal over the other, and no explicit gradient-compatible encoding exists for such signals. To address this, we propose DRIVE, a principled and efficient weight assignment framework. It derives per-sample weights via a unified differentiable objective where three coupled signals co-evolve: Shapley additivity splits contribution into anti-forgetting and adaptation components, which are reshaped by a covariance-based diversity term and stabilized by a sample-characteristic prior. Far from a mere combination, this yields an end-to-end gradient-based scheme balancing stability against plasticity. Theoretically, we prove that DRIVE enjoys superior stability-plasticity guarantees over existing weighting schemes and enhances parameter estimation robustness. Empirically, it attains superior overall performance, faster convergence, and stronger robustness. Code at https://anonymous.4open.science/r/DRIVE-158.
PaperID: 6326, Poster
Authors:
Yen N Pham, Dung V Nguyen, Trinh T Nguyen, Thieu Vo, Tan NguyenAbstract: A central challenge in understanding deep sequence models is characterizing how token representations evolve across layers: whether they collapse, diverge, or converge to nontrivial structures. We study this question for Transformers with hyper-connections, whose learnable cross-layer coupling makes the stability of representation collapse analytically tractable. In contrast to standard residual connections, hyper-connections introduce additional parameters that directly control the hyperbolicity of the zero fixed point, at which the local stability is fully determined by the Jacobian eigenvalues, thereby enabling a rigorous and precise bifurcation analysis of the dynamical system underlying the model. In particular, we derive a continuous-depth ODE limit of self-attention with hyper-connections and use bifurcation theory to characterize its token dynamics. For the 1-stream case, we obtain a closed-form critical threshold at which a pitchfork bifurcation occurs, separating representation collapse from divergence or convergence to nontrivial stable states. For the 2-stream case, the richer coupling structure yields two distinct stability boundaries, corresponding to qualitatively different instabilities: one arising when a real eigenvalue crosses zero, altering the number of nearby equilibria (static type), and another when a complex-conjugate pair reaches the imaginary axis, leading to oscillatory behavior (Hopf type). These phenomena have no counterpart in the 1-stream setting. Experiments on pretrained hyper-connected language models validate our theory and show that unstable regimes yield stronger performance.
PaperID: 6327, Poster
Abstract: Spiking Neural Networks (SNNs) are inherently sequential due to temporal state dependencies, limiting the efficiency of parallel hardware such as GPUs and resulting in high training cost. In this work, we propose the Parallel Fixed-Point Spiking Neurons (PFSN), which reformulates the dynamics of leaky integrate-and-fire neurons into a unified fixed-point mapping, enabling parallel computation across all time steps. Unlike conventional sequential unrolling, the proposed formulation decouples temporal dependencies through iterative fixed-point updates, significantly improving computational efficiency. Furthermore, we introduce a learnable temporal propagation operator that generalizes predefined dynamics and allows adaptive modeling of task-specific temporal interactions without relying on explicit temporal recursion. Extensive experiments across diverse domains, including event-based recognition, sequential image classification, speech processing, and time-series forecasting, demonstrate that PFSN consistently achieves superior efficiency while maintaining or improving predictive performance compared to existing parallel SNN approaches. These results highlight the effectiveness of combining fixed-point formulations with learnable temporal structures for scalable and efficient SNN training. Code is available at https://anonymous.4open.science/r/PFSN.
PaperID: 6328, Poster
Authors: Vladislav Tsvelenev, Maxim Ryndin
Abstract: Ensuring reliability of large language models (LLMs) is critical for their safe deployment in autonomous and resource-constrained applications, where hallucinations can lead to incorrect or harmful outputs. Detecting hallucinations from internal model states without external knowledge is promising, but existing methods lack distribution-free guarantees and fail under domain shift. We introduce PLLS-CP, a unified framework for hallucination detection that addresses both challenges. First, PLLS uses pseudo-label probing across transformer layers to identify the most informative ones, and trains an unsupervised detector on the selected layers. Then, a Mondrian conformal layer wraps the detector to provide provable coverage guarantees. Under domain shift, only the conformal threshold is updated on a small target calibration set — leaving the PLLS-trained detector fixed and avoiding retraining. We provide theoretical guarantees: marginal coverage under cross-domain recalibration, exponential concentration of the coverage gap, and non-degradation of the validation-selected ensemble. Experiments on 18 domains from TrueFalse, HaluEval, and HELM show that PLLS-CP substantially improves accuracy retention under cross-domain transfer while maintaining empirical coverage close to the nominal target level. On informative domains the method loses less than 2.5% of selective accuracy under transfer, suggesting a practical path to deployable, distribution-shift-robust hallucination detectors.
Abstract: 3D grounding aims to localize target entities in complex scenes from natural language and plays a fundamental role in embodied perception and spatial reasoning. However, existing approaches mostly rely on feature similarity or direct matching, making it difficult to connect natural-language intent with the implicit semantic and geometric structures hidden in billion-scale urban point clouds. We reformulate city-scale 3D grounding as structured constraint reasoning, where description semantics are organized into computable cross-modal constraints over open-vocabulary 3D entities, attributes, and spatial relations. We present CitySTAR, a training-free framework for reasoning-driven urban 3D grounding. CitySTAR lifts raw billion-scale urban point clouds into a query-ready scene graph of open-vocabulary 3D instances, with CodeLLM-driven tools supplying multimodal evidence for node attributes and 3D spatial relations. It then models target-context topology with paired hypergraphs and performs bidirectional topology verification for structural disambiguation. Finally, a Reflective Cross-modal Grounding module integrates topology consistency and candidate-centered 2D visual evidence to make decisions over a metric-aware 3D context graph. To further support this setting, we introduce CitySTAR-3D, an enhanced benchmark that improves semantic coverage, instance completeness, bounding-box fidelity, and spatial-relation complexity in city-scale 3D grounding. Extensive experiments show that CitySTAR consistently improves open-world urban 3D grounding while maintaining strong interpretability and generalization.
PaperID: 6330, Poster
Abstract: Federated Bayesian Optimization (FedBO) has emerged as a powerful paradigm for enhancing Bayesian Optimization (BO) in distributed, privacy-sensitive settings, enabling agents to accelerate the optimization of their local objectives by sharing knowledge about their similarities without exposing raw data. However, existing methods rely on restrictive assumptions, such as near-identical objective functions or shared Gaussian process hyperparameters that are often violated in heterogeneous real-world applications. We present GLOBAFED, a novel framework that enables efficient knowledge transfer across different but related objectives without compromising privacy. At its core is a probabilistic neural network surrogate with a shared feature-extracting backbone and agent-specific heads equipped with variational Bayesian last layers to provide uncertainty estimates. During optimization, agents use federated learning on the shared backbone to collaboratively build a common latent feature representation, while agent-specific heads adapt to their individual objectives. Our empirical results demonstrate that GLOBAFED significantly accelerates convergence and reaches better final performance compared to state-of-the-art FedBO baselines, and enables agents to warm-start their optimization from previously learned global features.
Abstract: Our aim is to develop a unified model for sign language understanding that performs both sign language translation (SLT) and sign–subtitle alignment (SSA). Together, these two tasks enable the conversion of continuous signing videos into spoken language text and also the temporal alignment of signing with subtitles -- both beneficial for practical communication, large-scale corpus construction, and educational applications. To achieve this, our approach is built upon three design choices: (i) a lightweight visual backbone that captures manual and non-manual cues from human keypoints and lip-region images while preserving signer privacy and enabling end-to-end training; (ii) a Sliding Perceiver mapping network that aggregates consecutive visual features into word-level embeddings to bridge the vision–text gap; and (iii) scaling training to large datasets -- BOBSL covering British Sign Language (BSL) and YouTube-SL-25 covering American Sign Language (ASL) -- to promote generalisation across languages and signers. With this multi-task, multilingual pretraining and strong model design, we achieve state-of-the-art results on BOBSL, How2Sign, Phoenix14T, and FLEURS-ASL SLT benchmarks, and also for SSA on BOBSL and WMT-SLT-SRF. Beyond these standard, signer-interpreted datasets, we show successful SLT on the natural signing data from the BSL-Corpus.
PaperID: 6332, Poster
Authors: Junhong Huang, Xin Tong, Jun Xia
Abstract: Long-horizon Large Language Model (LLM) agents need persistent memory that surfaces the right past interaction at the right time across thousands of turns. Current systems index by invoking generative LLMs to rewrite history into summaries, facts, or knowledge graphs---a step that is (a) costly (up to ~ 50K LLM calls and 35.83M completion tokens per build), (b) non-deterministic (undermining reproducibility and giving inconsistent answers across runs), and (c) provenance-erasing (the reader sees synthesized text, not the original turn). We propose MGMem, an LLM-call-free framework that replaces this step with a deterministic discriminative-parser pipeline: recurring mentions across sessions are organized into an incidence graph, question mentions are resolved by a subject-conditioned three-stage resolver, and raw dialogue episodes are returned to the reader. On LoCoMo, MGMem outperforms matched-protocol generated-memory and retrieval baselines at zero generative indexing calls; on the full LongMemEval-S, it remains competitive overall and leads on single-session question types. Anonymous code: \urlhttps://anonymous.4open.science/r/mentiongraph.
PaperID: 6333, Poster
Abstract: Accurately identifying alternative protein conformations remains a fundamental challenge, particularly when functionally relevant states are encoded by sparse evolutionary signals within large multiple sequence alignments (MSAs). In this work, we reformulate protein conformation prediction as a combinatorial search problem in MSA space, shifting the focus from structure divergence to evolutionary information discovery. We introduce , a optimization framework that enabling LLM with direct manipulation and iterative exploration of MSAs. Specifically, MSA-Evolver introduces a unified action space for MSA editing together with a feedback-guided multi-step reasoning strategy that allows the language model to progressively explore, evaluate, and refine candidate MSAs based on historical search trajectories. Our framework efficiently identifies informative sub-MSAs under limited folding budgets and substantially improves the prediction accuracy of alternative conformations, including open-closed, inward-outward, apo-holo, fold-switching, and intrinsically disordered proteins.
Authors:
QIWEI ZENG, Hao Wang, Jinghao Lin, Shuchang Ye, Yuezhe Yang, Yige Peng, Haoyuan Che, Jinman Kim, Lei BiAbstract: Medical vision-language models (VLMs) have shown increasing potential for clinical image interpretation, including lesion detection and report generation. However, their practical utility remains limited by insufficient sensitivity to subtle lesions, whose visual evidence is often sparse, low-contrast, and embedded within complex anatomical context. As local visual tokens are aggregated, these weak lesion cues can become underrepresented in global image representations, making them difficult for medical VLMs to recognize. Existing efforts to improve lesion sensitivity mainly rely on medical-domain vision-encoder pre-training, clinical-term-guided alignment, or trainable pathological representation enhancement. Although effective, these approaches usually require additional training or model-specific adaptation and may overfit to particular disease morphologies, limiting their applicability to frozen medical VLMs. To address these limitations, we propose EasyLens, a training-free plug-and-play subtle-lesion representation amplifier for medical VLMs. EasyLens first constructs EasyBank, a pathology-anatomy prototype space that provides lesion-related prototypes and anatomy-aware normal references for comparing suspicious patches against both pathological and normal anatomical patterns. To avoid blindly amplifying normal tissues, EasyTag selects lesion-relevant patches through counterfactual prototype reasoning. To counteract the dilution of subtle lesion cues in global image representations, EasyAmplifier strengthens the selected lesion-relevant patch representations through morphology-guided residual enhancement, thereby increasing their contribution to the global image embedding. Experiments on multiple medical image datasets and frozen medical VLM backbones show that EasyLens consistently improves subtle-lesion detection and outperforms existing encoder-enhancement baselines without model fine-tuning. Code is available at: https://anonymous.4open.science/r/easylens-BEC2
PaperID: 6335, Poster
Abstract: We present GenZ, a hybrid model that turns a foundational model (FM) into a \bf knowledge-discovery engine explaining variation in real-valued multidimensional targets. The FM is treated as a noisy oracle that can answer yes/no questions about a semantic item s (text, image, or both) and GenZ learns \bf which questions to ask so that the answers \bz explain the statistical link to a possibly high-dimensional target \by. Discovery is driven by \bf group reasoning: at each step, the model partitions items by the posterior of the current latent features, and asks the FM to articulate the semantic commonality that explains the divide. The resulting feature descriptors \theta_f are the primary product---human-readable hypotheses about why s predicts \by---while predictive accuracy is a secondary, but consistently strong, by-product (e.g.~beating both the 0-shot FM baseline and TabPFN on LLM embeddings of the same items in hedonic price regression). Across four domains---hedonic house pricing, Netflix cold-start recommendations, Arizona species ecology, and human visual-cortex fMRI---GenZ recovers dataset-specific structure that the FM does not surface from priors alone. Holding the items s fixed and varying only the target \by (Arizona species under spatial/taxonomic/functional targets; the same natural images under FFA/EBA/PPA brain responses) yields \bf qualitatively different discovered feature sets, demonstrating that GenZ characterizes the s\!\to\!\by link rather than the marginal distribution of s.
PaperID: 6336, Poster
Abstract: The MOEA/D is one of the most successful multi-objective evolutionary algorithms. It operates by decomposing the optimization problem into single-objective subproblems, which are then solved in a co-evolutionary manner. The performance of the MOEA/D depends critically on this decomposition. In this work, we propose a novel way to adjust this decomposition on the fly, for the case of discrete bi-objective optimization. Different from all previous works in this direction, we support our new algorithm with a proven performance guarantee. This is the first work that rigorously proves the runtime of a dynamic multi-objective evolutionary algorithm. For monotonic distortions of the simple LOTZ benchmark, it is known that classic decompositions leave large holes in the Pareto front, making it very hard for the MOEA/D to cover the entire Pareto front. In contrast, we prove that our self-adjusting MOEA/D solves all these LOTZ problems in an expected number of O(n^4) function evaluations. This relatively low runtime as well as the details of our mathematical proof show that the dynamic decomposition strategy indeed evolves a decomposition adjusted to each of these problem instances.
PaperID: 6337, Poster
Authors: Minh Tu, Quang-Binh Nguyen, Khoi Nguyen
Abstract: Articulated 3D assets are essential for robot simulation, yet generating them from casual images remains unsolved. Existing approaches either predict URDF parameters via learned regression or language model inference -- both requiring articulation supervision and producing joints that may be geometrically inconsistent with the generated meshes -- or reconstruct from dense calibrated multi-view captures, limiting scalability. We present ArtCrafter, a feed-forward generative model that takes two images of an articulated object -- one at rest, one fully open -- and produces 2N+1 part-level meshes together with a physics-executable URDF in a single forward pass. Inspired by Slot Attention, we structure the denoising process around a 1+2N slot layout whose slots, forced to jointly explain two articulation states, naturally bind to kinematic parts. TripletAttention enforces cross-state geometric consistency between each paired rest/open slot anchored through a shared base slot, and a repulsion loss \mathcalL_\mathrmrep operating on decoded point clouds ensures inter-part disjointness throughout denoising. Because paired meshes are geometrically consistent by construction, URDF joint parameters -- type, axis, origin, and range -- are derived \emphanalytically from the rigid transform between each pair, requiring no learned articulation head, no language model, and no per-instance optimization. Experiments on URDF-Anything+ demonstrate consistent improvements over retrieval-based and generative baselines across part-level geometry and all four articulation metrics. ArtCrafter demonstrates a clean decomposition: a diffusion model handles geometry, and geometry handles URDF.
PaperID: 6338, Poster
Abstract: Segmenting the left ventricle in echocardiographic videos remains difficult because the target undergoes continuous deformation and rapid displacement across the cardiac cycle, while low tissue contrast further obscures its endocardial boundaries. Expert echocardiographers navigate these challenges through a hierarchical cognitive workflow, sequentially addressing three fundamental questions:What is the target? Where is it now? Where does the boundary lie? Yet existing methods predict masks end-to-end from temporal features without explicitly modeling this cognitive process. We propose Hierarchical Cognitive Decomposition (HCD), which decomposes the segmentation task into three stages that emulate this expert interpretation workflow. To capture invariant target identity amid changing appearances, Static Identity Anchoring (SIA) anchors the first annotated frame as a semantic reference, enabling stable recognition across the cardiac cycle. To maintain spatial focus despite inter-frame motion, Adaptive Gaze Prior (AGP) converts the previous frame's prediction into a spatial attention prior that localizes the current target position. To resolve boundary ambiguity under low tissue contrast, Structural Boundary Perception (SBP) extracts multi-scale boundary cues within the localized region. On the CAMUS and EchoNet-Dynamic benchmarks, HCD achieves state-of-the-art performance with only 1.35M trainable parameters (out of 35.2M total) at 68 FPS, demonstrating that explicitly encoding clinical cognitive priors into network design yields both effective and efficient segmentation.
Abstract: Partial differential equations (PDEs) are foundational to modeling in science and engineering, but constructing reliable numerical solvers remains labor-intensive, demanding expert knowledge of discretization schemes, stability conditions, and boundary treatments. Recent work has begun to frame PDE solving as a code-generation task for large language models (LLMs), yet existing approaches operate primarily at inference time: relying on prompting, debugging, self-refinement, and test-time scaling rather than adapting the model itself. In parallel, reinforcement learning with verifiable rewards has emerged as a powerful post-training paradigm for code and math reasoning, but its verifiers are typically binary: a compiler runs, or a test passes. Such signals discard the graded structure of scientific correctness, where two solvers may both execute and yet differ in solution accuracy by orders of magnitude. In this work, we introduce RLVP: Reinforcement Learning with Verifiable Physics, an RL post-training framework for multi-PDE solver code generation.RLVP addresses this verifiability gap with a hybrid verifier: hard program-validity checks ensure executability, while continuous physics rewards score function-space accuracy and PDE-residual consistency. A single policy is post-trained across diverse PDE families spanning hyperbolic, parabolic, elliptic, and incompressible-flow systems. RLVP improves over both pre-trained and supervised-only baselines on PDE benchmarks, and shows zero-shot improvement transfer to held-out PDEs. We show that a smaller LLM post-trained with RLVP can outperform prompting a frontier model on in-distribution PDE solver generation. The trained policy shows evidence of compositionality in numerical motifs: it recombines stencils, time-stepping schemes, and boundary-handling primitives learned from the PDEs used in training into generated solvers for unseen PDE problems.
PaperID: 6340, Poster
Authors:
Antonia Marcu, Jonathon Hare, Annika Catulli, Srinandan Dasmahapatra, Damian SmithAbstract: Attribution methods are extensively used to identify which parts of the input are important for a model's decision. However, despite recent efforts there is still no consensus on what exactly they capture and how to evaluate them. We start addressing this problem by explicitly modelling the implicit assumptions that sit at the foundation of attribution methods. We define two types of importance contributions an input region can have. We then create a framework that decouples the process of establishing contributions from that of evaluating attributions. This shows that current attribution methods struggle to reliably capture input contributions. Further investigation shows a stark difference in attribution performance between model behaviours, exposing a previously overlooked aspect in the attribution literature. Our work calls for clear, quantifiable statements about what attribution methods aim to capture, along with rigorous evaluation frameworks.
Abstract: Concept Bottleneck Models (CBMs) enable interpretable image classification by structuring predictions around human-understandable concepts, but extending this paradigm to video remains challenging due to the difficulty of extracting concepts and modeling them over time. In this paper, we introduce MoTIF (Moving Temporal Interpretable Framework), a transformer-based concept architecture that operates on sequences of temporally grounded concept activations, by employing per-concept temporal self-attention to model when individual concepts recur and how their temporal patterns contribute to predictions. Central to the framework is a class-conditioned VLM-based concept discovery module that extracts object- and action-centric textual concepts from training videos, yielding temporally expressive concept sets without manual concept annotation. Across multiple video benchmarks, this combination improves over global concept bottlenecks and remains competitive within the interpretable concept-bottleneck setting, while narrowing the gap to strong black-box video baselines that we report as contextual references.
PaperID: 6342, Poster
Authors: Antonio Ken Iannillo, Joshua S Owotogbe, Roberto Natella, Francesco Avallone, Kristina Kudryavtseva, Indika Kumara
Abstract: LLM agents increasingly rely on external tools exposed via the Model Context Protocol (MCP), yet current benchmarks assume tools respond correctly and on time, leaving agent resilience to realistic tool failures untested. We introduce MCP-Injector, a drop-in, protocol-aware fault-injection middleware that interposes on the MCP tool-call path with controlled persistence (permanent, transient, intermittent) and seedable schedules. Its fault model is empirically calibrated by mining 57,240 issue and pull-request threads across 1,774 open-source MCP repositories, yielding a distribution over six failure modes grounded in observed ecosystem incidents. In a 1,350-cell experiment crossing two backbone sizes (20B, 120B), three agent stacks (custom multi-round plan-and-execute, OpenAI Agents SDK, LangChain), and three fault injection conditions, we find that calibrated MCP faults cause severe quality degradation (60% drop in task completion score, p < 10^-56) and reveal a scaling paradox: the larger model achieves higher fault-free scores but suffers lower completion rates under faults, driven by longer reasoning chains that amplify fault exposure. Framework-level error handling shapes resilience more than model scale, with stack choice producing significant interaction effects. These results demonstrate that capability benchmarks alone can mask operational brittleness, motivating protocol-aware fault injection as a complementary evaluation dimension for tool-calling agents.
PaperID: 6343, Poster
Abstract: Group dance generation aims to synthesize coordinated multi-dancer choreography from music, with broad applications in animation and interactive content creation. This task requires modeling dense inter-person dependencies to ensure spatial coordination, while naturally preserving individual dancer identities. Existing approaches model all dancers jointly with end-to-end transformers, which tie the architecture to a fixed group size and entangle per-dancer identities across frames. We propose ChainDance, a scalable framework that reformulates group dance generation as a Chain-of-Dancers: a sequential decomposition over per-dancer conditional distributions, allowing a single model to scale across arbitrary group sizes without retraining and naturally preserving per-dancer identity. Built on a frozen single-dancer diffusion backbone, ChainDance introduces two lightweight modules: a Role-Aware Text Encoder (RATE) for per-dancer semantic conditioning, and a Group-Aware Motion Encoder (GAME) that aggregates previously generated dancers via a distance-weighted graph convolutional network, and incorporates a training-free noise optimization procedure at inference time to enforce global spatial coherence. Experiments on AIOZ-GDance demonstrate that ChainDance achieves state-of-the-art performance in motion quality and identity preservation, with 3-4 times fewer parameters and requiring 3-6 times less training time compared to prior approaches.
PaperID: 6344, Poster
Abstract: Knowledge distillation (KD) has become a widely adopted approach for compressing large language models (LLMs). Existing distillation paradigms for LLMs can be broadly categorized into black-box and white-box approaches. Black-box methods typically transfer the chains of thought (CoTs) generated by teacher models to student models, yet they fail to convey the intrinsic knowledge embedded within teacher models. White-box methods aim to alleviate this limitation by leveraging internal representations from teacher models. However, they still face two major challenges. First, it is difficult to identify which knowledge is essential for distillation in student models. Second, it is challenging to transfer this key knowledge from teacher models accurately. To tackle these challenges, we propose KPD, a novel Knowledge Probing-based white-box Distillation framework for LLMs, which is designed to precisely distill the knowledge that the student model lacks and can be incorporated into diverse distillation approaches. Specifically, to identify critical knowledge gaps in student models, we introduce an uncertainty-based probing method that extracts tokens prone to student errors as critical knowledge. To locate the corresponding knowledge in the teacher model, we utilize cumulative gradient calculation to probe where such knowledge is stored and then distill the target layers. Extensive experiments across multiple datasets demonstrate that KPD boosts traditional distillation methods by 0.34%-4.56% in Rouge-L scores. Further ablations show that probing both the student and teacher yields more precise distillation. Our code is available at https://anonymous.4open.science/r/KPD-CA0D.
PaperID: 6345, Poster
Abstract: Step-size selection remains a central challenge in large-scale neural network optimization: conservative steps slow convergence, while aggressive steps can destabilize training. We propose Zero-and-First-Order optimization (ZFO), a lightweight framework that decouples direction selection from step-size selection. ZFO uses a trusted first-order optimizer to determine the direction and performs zeroth-order evaluations only along this one-dimensional subspace to choose how far to move. Using the current forward-backward pass and two additional fixed-batch function evaluations, ZFO constructs a local model of the objective along the proposed direction and selects a curvature-aware step within a bounded search interval. This yields an adaptive step-selection mechanism with lower cost than full line search. We provide theoretical guarantees showing that shared-sample evaluations produce reliable finite-difference curvature estimates, that the induced local model selects a near-optimal step along the search interval, and that ZFO converges to a neighborhood of a stationary point. Across multiple language models and datasets, ZFO improves accuracy and stability over standard first-order fine-tuning algorithms.
PaperID: 6346, Poster
Abstract: Despite great success, existing MoE-MLLMs still suffer from substantial visual expert redundancy, making visual expert skipping a critical solution. However, deriving an effective and fixed skipping policy is still an open problem, mainly due to the exponentially large layer-wise policy space and the dependency between expert allocations across layers. In this paper, we propose a novel and training-free approach for MoE-MLLMs, termed \emphK-armed bandit based expert redundancy estimation (KAB-MoE). KAB-MoE defines the retained visual expert number at each MoE layer as a discrete action, and then conducts efficient one-shot RL estimations to obtain the layer-action reward matrix. Based on this reward matrix, KAB-MoE can quickly derive the optimal skipping policy that satisfies the predefined computation budgets, \emphe.g., expert skipping ratio. The obtained skipping policy can be directly applied to MoE-MLLMs for all examples, which not only facilitates high-throughput deployment but also contributes to practical acceleration. Extensive experiments on Kimi-VL-A3B and Qwen3-VL-MoE show that KAB-MoE effectively accelerates MoE-MLLMs while keeping strong multimodal performance, \emphe.g., retaining 97.09% average performance on Kimi-VL-A3B under the skipping ratio of 83%. Moreover, KAB-MoE also achieves competitive or superior performance over dynamic expert skipping SOTAs with practical inference speed-up. Our code is provided in the supplementary materials.
Abstract: Physically consistent 3D object dynamics is a key component for controllable simulation and physics-grounded video generation, especially under contact, deformation, and external forcing. Existing trajectory-based methods enable physical control, but they often model isolated physical effects rather than deriving motion from a unified physical structure. This makes it difficult to compose conservative and non-conservative effects in contact-rich 3D dynamics while retaining explicit control. We present NEXUS, a neural energy-field method for contact-rich 3D object dynamics. NEXUS represents each object as a structural graph and constructs dynamic contact graphs for interactions. Inspired by the Hamiltonian Neural Network (HNN), NEXUS formulates dynamics through scalar energy and dissipation terms rather than direct state or acceleration regression. Conservative effects are composed as additive energy terms over the scene. To handle non-conservative systems beyond standard HNNs, NEXUS learns a dissipation function for impact-induced energy loss. Forces are derived by differentiating the energy-dissipation functions and rolled out with a numerical integrator. NEXUS provides a stable and controllable physics reasoning module for contact-rich 3D object dynamics, improving long-horizon accuracy over representative learned and physics-structured dynamics baselines across controlled rollouts with varying mechanical properties and physical-effect compositions. Our contributions are: (i) an HNN-inspired energy-dissipation dynamics formulation that unifies conservative and non-conservative scene effects; (ii) a graph-based 3D representation for contact-rich object dynamics, with object--object contact evaluated in mixed-contact stress tests; and (iii) a trajectory-guided video generation study showing that physically consistent motion improves downstream physical plausibility while maintaining competitive visual quality.
Abstract: Sparse autoencoders (SAEs) decompose language model activations into interpretable features, primitives for mechanistic circuit analysis. Yet existing methods identify task-specific circuits via attribution or apply fixed-direction steering, but learn no task-reward-optimized policy for per-token feature amplification. We formulate token-level feature intervention as a POMDP over the LLM's layer-causal computation. Control Reinforcement Learning (CRL) trains a policy in this POMDP that selects an SAE feature to amplify at each generation step, yielding per-token intervention traces that expose which features drive task behavior. Adaptive Feature Masking encourages diverse feature discovery while preserving single-feature attribution. The framework provides branch point tracking (tokens where feature choice changes outcomes), critic trajectory analysis (separating policy from value estimation errors), and layer-wise comparison along the residual stream hierarchy. Cross-task transfer is asymmetric (MMLU features harm GSM8K but not vice versa; HarmBench helps XSTest), consistent with partial task computation overlap rather than dataset-specific shortcuts: task-conditional attribution does not require transfer invariance. On Gemma-2 2B and LLaMA-3.1 8B across MMLU, BBQ, GSM8K, HarmBench, and XSTest, CRL provides per-token intervention traces and serves as an SAE quality diagnostic alongside task-accuracy gains. Learned feature steering thus serves as a layer-specific diagnostic complementing activation-based SAE analysis.
PaperID: 6349, Poster
Authors:
Mridul Khurana, Amin Karimi Monsefi, Justin Lee, Medha Sawhney, David Carlyn, Julia Chae, Jianyang Gu, Rajiv Ramnath, Sara Beery, Wei-Lun (Harry) Chao, Anuj Karpatne, Cheng ZhangAbstract: Accurately generating images across the Tree of Life is difficult: there are over 10M distinct species on Earth, many of which differ only by subtle visual traits. Despite the remarkable progress in text-to-image synthesis, existing models often fail to capture the fine-grained visual cues that define species identity, even when their outputs appear photo-realistic. To this end, we propose TaxaAdapter, a simple and lightweight adapter that incorporates the taxonomic embeddings of Vision Taxonomy Models (VTMs) such as BioCLIP to perform scalable fine-grained species generation over the entire Tree of Life. TaxaAdapter injects VTM embeddings using taxonomy-text dual conditioning into a frozen text-to-image diffusion model, improving species-level fidelity while preserving flexible text control over attributes such as pose, style, and background. Extensive experiments demonstrate that TaxaAdapter consistently improves morphology fidelity and species-identity accuracy over strong baselines. To better evaluate these improvements, we also introduce a multimodal Large Language Model-based metric that summarizes trait-level descriptions from generated and real images, providing a more interpretable measure of morphological consistency. Beyond showing improvements over standard benchmarking experiments, we observe that TaxaAdapter exhibits strong generalization capabilities in open-world settings, enabling species synthesis in challenging regimes such as few-shot species with only a handful of training images and even species unseen during training. Overall, our results highlight that VTMs are a key ingredient for scalable, fine-grained species generation.
Abstract: Flow matching has recently emerged as a powerful framework for continuous-time generative modeling. However, when applied to long-tailed distributions, standard flow matching frameworks are susceptible to majority bias, over-representing dominant modes at the expense of tail-class fidelity and distributional accuracy. In this work, we propose Unbalanced Optimal Transport Reweighted Flow Matching (UOT-RFM), a novel framework for generative modeling under class-imbalanced (long-tailed) distributions that operates without any class label information. Our method constructs the conditional vector field using mini-batch Unbalanced Optimal Transport (UOT) and mitigates majority bias through a principled inverse reweighting strategy. The reweighting relies on a label-free majority score, defined as the density ratio between the target distribution and the UOT marginal. This score quantifies the degree of majority based on the geometric structure of the data, without requiring class labels. By incorporating this score into the training objective, UOT-RFM theoretically recovers the target distribution with first-order correction (k=1) and empirically improves tail-class generation through higher-order corrections (k > 1). Our model outperforms existing flow matching baselines on long-tailed benchmarks, while maintaining competitive performance on balanced datasets.
PaperID: 6351, Poster
Abstract: Self-supervised learning draws inspiration from infant visual development, yet standard training pipelines bear little resemblance to it: images are independently sampled and globally shuffled across epochs. We study self-supervised learning from continuous video streams, where frames are consumed in temporal order using strict sliding-window batches, without global reshuffling or multi-epoch replay. To this end, we construct WT++, a 95-hour urban walking-tour video dataset for streaming pretraining. Combined with a comprehensive evaluation suite we find that contrastive and distillation-based methods struggle, while MAE is more robust but still falls short of standard i.i.d. pretraining. We find that high inter-batch similarity, caused by sliding-window consumption across consecutive batches, does not explain this gap alone. The main challenge is high intra-batch similarity, where frames within each batch are near-duplicates. To mitigate this, we propose StreamMAE, which preserves the core MAE reconstruction objective while adapting the input pipeline with stream-aware regularization and motion-biased crop selection. StreamMAE outperforms streaming baselines, matches both i.i.d. MAE on video and even its ImageNet equivalent, and scales positively as pretraining video grows from 12 to 95 hours.
PaperID: 6352, Poster
Authors: Peter Sharpe
Abstract: We introduce GLOBE, a neural surrogate for boundary-driven, homogeneous, weakly nonlinear PDEs that draws inductive bias from boundary-element methods and equivariant ML, representing solutions as superpositions of learnable Green's-function-like kernels evaluated from boundary faces to targets. The architecture is translation-, rotation-, and parity-equivariant; discretization-invariant; and units-invariant via rigorous nondimensionalization. On AirFRANS (steady incompressible RANS over NACA airfoils), GLOBE reduces mean-squared error by roughly 200× relative to dataset baselines and 50× relative to the next-best ML surrogate on interpolation tasks, with 7×–70× improvements over the next-best surrogate in low-data regimes. The compact ~ 117 \rm k-parameter model evaluates fields at arbitrary points, handles non-watertight meshes, and generalizes under Reynolds number and angle-of-attack distribution shifts. We further introduce a hierarchical Barnes-Hut-style acceleration that exploits the kernel's guaranteed decay of long-range influences to reduce evaluation cost from quadratic to near-linear in the number of boundary elements, with a tunable accuracy parameter that recovers exact dense evaluation in its strict limit. We demonstrate scalability to 3D industrial geometries via a proof-of-concept on DrivAerML car aerodynamics. Together these results show that physics-inspired inductive biases yield large gains in accuracy and practicality for ML-based surrogates of boundary-driven PDEs while remaining tractable at industrial 3D scales.
PaperID: 6353, Poster
Authors:
Shailesh Dahal, Ratri Mukherjee, Nicholas Mathews, Kishlay JhaAbstract: Health risk prediction aims to predict a patient's potential health risks (e.g., mortality) using the longitudinal information present in their electronic health record (EHR). While deep learning-based risk prediction models have shown great promise, they are susceptible to the issue of missing data prevalent in real-world EHR. Most prior works mitigate this challenge by treating missing data as an uninformative perturbation to be ignored or imputed. However, inaccurate imputation may lead to the generation of sub-optimal patient representations that adversely impact risk prediction performance. Moreover, recent studies show that missing data in EHR does not necessarily imply error or noisy observation and instead may communicate informative clinical signals useful for health risk prediction. To address this, we propose a novel diffusion-based representation learning approach (namely DiffRisk) that explicitly captures informative missingness by modeling conditional dependencies induced by partial observations in EHR. Specifically, DiffRisk learns latent representations with coherent conditional structure, where any subset of observed features defines a distribution over the unobserved components. To achieve this, we introduce a conditional score-based objective and approximate it via denoising diffusion that enables efficient learning without explicit density estimation. Empirical results on five risk prediction tasks show that the proposed approach consistently outperforms state-of-the-art baselines.
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a standard recipe for post-training LLMs on reasoning tasks, with Group Relative Policy Optimization (GRPO) emerging as a leading approach. However, GRPO and its variants are inherently single-turn: they optimize terminal rewards on isolated prompt-response pairs, leaving them poorly suited to agentic settings where models iteratively refine solutions using environmental feedback. We introduce MURPHY, a multi-turn extension of GRPO for self-correcting code generation. MURPHY constructs feedback-conditioned rollout trees in which failed candidate solutions are paired with executor feedback and expanded into subsequent turns. It then propagates rewards backward through the tree so that later successful refinements credit earlier attempts that surfaced informative feedback. We study two propagation strategies, Max Reward (MaRS) and Mean Reward (MeRS), and introduce post-rollout pruning mechanisms that reduce multi-turn optimization cost. Across three code generation benchmarks, HumanEval, MBPP, and LiveCodeBench-v6, and two model families, Qwen3-1.7B/4B and OLMo-2-7B, MURPHY delivers up to absolute pass@1 gains over the strongest prior multi-turn execution-feedback methods. Gains are largest on the Medium/Hard LiveCodeBench subsets, reaching
Abstract: When humans see a bird, they recognize far more than just ``bird'' --- they see a head, wings, and talons, a structured assembly of reusable parts that can be identified across every bird they have ever seen. We ask whether a self-supervised visual model can discover the same compositional structure on its own. To this end, we propose RATS (Register Attention Transformers), which decomposes the classification token into N learnable register tokens that route patch information through an L\toN\toN\toL bottleneck. The N registers are hard-partitioned across H attention heads, structurally isolating each subset in an independent projection subspace. Without auxiliary losses or part annotations, each register spontaneously specializes into a semantically coherent visual part. RATS surpasses all baselines by an average of +12 mIoU on five segmentation benchmarks, and demonstrates stronger dense prediction on ADE20K (+1.11 mIoU) and COCO (+0.2 AP^\textm). The visual dictionary extracted from the trained registers also shows signs of part-level consistency and semantic proximity across related categories. Our results suggest that RATS may provide a useful architectural prior for structured and interpretable visual representation learning.
PaperID: 6356, Poster
Abstract: LLM agents excel when environments are mostly static and their required information fits in a model’s context window, but they struggle with IT enterprise diagnostic tasks—such as incident management, IT vulnerability analysis, and FinOps anomaly explanation—where operators iteratively mine massive observability data over a curated resource graph to identify the origins that explain observed symptoms, enabling correct remediation. The data in these domains inherently has hidden dependency structure: entities interact, signals co-vary, and the importance of a fact may only become clear after other evidence is discovered. To cope with bounded context windows, agents must summarize intermediate findings before their significance is known, increasing the risk of discarding key evidence. ReAct-style agents are especially brittle in this regime. Their retrieve-summarize-reason loop makes conclusions sensitive to exploration order and introduces run-to-run non-determinism, producing a reliability gap where Pass-at-k may be high but Majority-at-k remains low. Simply sampling more roll-outs or generating longer reasoning traces does not reliably stabilize results, since verifying a diagnosis requires remediation actions that carry operational risk---demanding consistent correctness, not occasional success---and ReAct provides no mechanism for belief revision as evidence accumulates. In addition, ReAct entangles semantic reasoning with controller duties such as tool orchestration and state tracking; execution errors and plan drift degrade reasoning while consuming scarce context. We address these issues by formulating the investigation as abductive reasoning over a dependency graph and proposing EoG (Explanations over Graphs), a disaggregated framework where an LLM performs bounded local evidence mining and labeling (cause vs symptom) while a deterministic controller manages traversal, state, and belief propagation to compute a minimal explanatory frontier. On representative ITBench diagnostics tasks, EoG improves accuracy and run-to-run consistency over ReAct baselines, including a 7x average gain in Majority-at-k F1 score.
PaperID: 6357, Poster
Abstract: We study a general class of regularized policy gradient methods for episodic learning in general-sum stochastic games, with the aim of describing the long-run behavior of the induced policy iterates. In this challenging setting, learning even the simplest pointwise solution concepts is known to be intractable. This leads us to switch our focus to less restrictive setwise solution concepts: rather than only asking whether the iterates converge to a specific policy, we ask which sets of policies are selected by the learning dynamics in the long-run. In one-shot games, this setwise perspective admits a crisp characterization: a product set of pure strategies is asymptotically stable if and only if it is closed under better replies (club). We show that the direct statewise analogue of this principle fails outright in stochastic games because actions affect not only current rewards but also the future evolution of the game. To restore this principle in a transition-aware form, we introduce Q-club sets, namely state-action sets in which every allowed action has a strict Q-value advantage over every excluded action, thereby providing a setwise generalization of strict Nash policies. We then prove that Q-club sets are stochastically stable and attracting under the learning dynamics, even in the presence of noise and uncertainty. Finally, we establish local convergence rates to Q-club sets governed by the regularizer; in particular, entropic regularization yields exponential convergence, while Euclidean regularization yields finite-time convergence.
PaperID: 6358, Poster
Abstract: Enabling language model agents to autonomously adapt to new environments remains a fundamental challenge. While recent memory-augmented approaches have achieved promising results, they are largely heuristic in design and treat task execution and knowledge construction as decoupled processes. Inspired by how biological systems rapidly adapt through principled acting and learning, we propose FEP-Agent, a framework that grounds LLM agent self-evolution in the Free Energy Principle (FEP). Our system instantiates the two components of Active Inference into: a Worker agent that minimizes Expected Free Energy (EFE) during online interaction, balancing goal-directed exploitation with curiosity-driven exploration to probe uncertain environmental dynamics; and a Builder agent that minimizes Variational Free Energy (VFE) post-execution, updating a knowledge base while controlling complexity through counterfactual validation. To support scalable knowledge in open-ended environments, we introduce a structured semantic memory that employs associative linking and overlapping communities for efficient retrieval. The excellent results across multiple LLM-agent benchmarks demonstrate that FEP-Agent achieves principled, self-evolving systems.
PaperID: 6359, Poster
Abstract: Cross-Domain Recommendation is essential for enabling cross-sell in multi-service platforms; however, limited user overlap and strict cross-service data-sharing constraints often result in little to no ground-truth supervision for cross-domain learning. To bridge this gap, we propose an LLM-based framework for cross-domain recommendation in the non-overlapping user scenario. First, we propose an agentic pipeline that constructs cross-domain pseudo supervision, which is unavailable in the non-overlapping user scenario, by leveraging a user’s source-domain history to generate a target-domain item recommendation. Using this dataset, we further propose \underlineCoupled \underlineUser-G\underlineRouped R\underlineEinforcement Learning (CURE), which induces coupling between in- and cross-domain gradient updates via user-level joint normalization, thereby calibrating synthetic cross-domain updates against reliable in-domain signals. We provide a theoretical analysis showing that user-level coupling stabilizes policy updates under imperfect pseudo supervision and improves cross-domain generalization. Experiments on Amazon Reviews and MovieLens show consistent gains over baselines in cross-domain recommendation, while also improving in-domain performance. An online A/B test demonstrates a statistically significant 56% lift in click-through rate over existing ML methods.
Abstract: Generative models typically rely on explicitly tracking the noise level to guide the sampling process. However, recent autonomous models successfully generate data using a single, time-invariant vector field that operates without explicit noise conditioning. This raises a fundamental paradox: what landscape are these networks actually optimizing when the noise level is treated as a random variable, and how can a bounded model remain stable near the data manifold where gradients typically diverge? We resolve this paradox by showing that autonomous generation is not merely ``blind’’ denoising, but a specific form of Riemannian gradient flow on the Marginal Energy landscape, defined by integrating out the unknown noise levels. Through a novel relative energy decomposition, we demonstrate that while the raw Marginal Energy contains a severe geometric singularity normal to the data manifold, the learned field implicitly incorporates a local conformal metric. This metric perfectly preconditions the singularity, turning an infinitely deep potential well into a stable attractor. Finally, we establish strict structural stability conditions for autonomous sampling. We prove that standard noise-prediction parameterizations structurally fail due to an inherent high gain amplification of the estimation errors. In contrast, velocity-based models are inherently stable.
PaperID: 6361, Poster
Authors: Seungwan Jin, Taehyung Noh, Junghyun Kim, Uichin Lee, Kyungsik Han
Abstract: Mobile and wearable sensing offer a scalable pathway to longitudinal mental health monitoring, yet accurate prediction remains limited by two fundamental forms of distribution shift in human behavior. The first arises across users, where similar sensor patterns can carry different clinical meanings, a phenomenon we call inter-subject heterogeneity. The second arises within a single user, where behavioral baselines themselves evolve over time, producing intra-subject behavioral drift. In this domain, ground-truth labels come from Ecological Momentary Assessment (EMA), brief self-reports collected in everyday life that are intrusive and inherently sparse. As a result, approaches that address these distribution shifts by updating model parameters are difficult to apply. To address this gap, we propose DyCon (Dynamic Context Modeling), a dual-loop framework that personalizes by building and evolving a separate natural-language personal context for each user, without updating any model parameters. Within each session, a momentary refinement loop restructures this personal context to resolve inter-subject heterogeneity. Across sessions, a longitudinal update loop integrates only validated insights to track intra-subject behavioral drift. Extensive experiments on four public benchmarks (GLOBEM, PMData, StudentLife, LifeSnaps) show that DyCon consistently outperforms baseline methods, particularly under sparse supervision and behavioral drift. Our results suggest that evolving an explicit personal context offers a practical pathway for longitudinal personalization in real-world health monitoring.
PaperID: 6362, Poster
Abstract: Language models routinely revise initial answers through intermediate edits, yet current alignment objectives restrict credit assignment to complete outputs or entire rollouts. To formalize the local geometry of improvement, we introduce Lifted State Policy Optimization (LSPO), which augments each textual answer with a continuous auxiliary coordinate to form a lifted state that encodes refinement context beyond surface text. Training rewards transitions that descend an energy landscape, subject to edit and step penalties, and internalizes the resulting answers into the generative model. Across the Qwen3.5 family on mathematics, science, code, logic, and broad reasoning benchmarks, LSPO improves aggregate accuracy while bypassing the latency of explicit revision loops. Component ablations, together with evaluations of internalization and energy ordering on unseen tasks, confirm that this gain originates from the lifted representation and from credit assigned at the transition level.
PaperID: 6363, Poster
Abstract: As one of the most complex and load-bearing joints in the human body, the ankle plays a crucial role in locomotion and clinical assessment. Accurate segmentation of ankle bones is essential for trauma evaluation, preoperative planning, and disease diagnosis. However, the scarcity of publicly available ankle datasets has hindered the development of intelligent analysis in this area. To address this gap, we introduce the first two publicly available ankle CT datasets, providing new directions for medical image segmentation. Considering the intricate and spatially correlated anatomy of the ankle, we propose a 3D medical image segmentation model, namely 3DABSeg. Specifically, we introduce the Hybrid KAN-Mamba Block (HKMB) to capture long-range spatial dependencies and complex nonlinearities. By integrating Mamba’s sequence modeling with KAN’s learnable spline functions, HKMB enhances feature expressiveness for complex anatomical structures. Furthermore, to address the uneven semantic distribution and feature dilution across channels, we propose a Multi-Scale Feature Fusion Mixture-of-Experts module (MFF-MoE). MFF-MoE utilizes multi-scale spatial pooling to compute dynamic routing scores, adaptively partitioning channels into specialized expert networks via a mutually exclusive hard-assignment strategy. Extensive experiments demonstrate that 3DABSeg surpasses existing state-of-the-art methods in both segmentation performance and boundary precision, highlighting its potential for complex anatomical structure analysis and clinical applications.
PaperID: 6364, Poster
Authors: Minghao Chen
Abstract: Maintaining character consistency over long-horizon gameplay is a critical bottleneck for deploying Large Language Models (LLMs) as Non-Player Characters. Prevalent discrete Affective State Machines (ASMs) fail due to two paradoxes: (1) the Cliff-edge Effect, where threshold-based switching causes jarring narrative discontinuities; and (2) Auto-regressive Inertia, where accumulated dialogue history dilutes current state instructions. We propose C3 (Causal-Continuous Characterization), which couples a continuous state space with hysteresis-aware dynamics to dampen emotional volatility, and a State-Conditioned Memory Rewriting mechanism that re-contextualizes past events through the lens of the current relationship. On a 100-turn dynamic simulation (\textscChronos-Sim, 5,000 turns per condition), C3 achieves a 10.5% absolute SAR improvement over a strong smoothed-ASM baseline (Welch's t-test, p < 0.0001, Cohen's d \approx 3.97). Cross-model evaluations on \textttLlama-3-8B-Instruct and \textttMistral-v0.2-7B reproduce the relative gain, indicating the mechanism transfers across backbones although absolute SAR remains backbone-dependent. A human study with 20 game designers (ICC=0.74) corroborates higher perceived consistency, and a separate factual audit (Fleiss' \kappa=0.81) confirms the Rewriter alters interpretation without corrupting facts. Deployment analysis shows a ~60% reduction in long-session token consumption.
Abstract: First-person dynamic spatial reasoning requires models to track continuous motion and precise geometric structure, but the quadratic attention cost of Transformer-based Video-LLMs makes dense visual tokens computationally expensive. Existing token pruning paradigms predominantly rely on discrete static snapshots, failing to preserve the physical dynamics essential for reasoning. We propose Event Cascade Pruning (ECP), to our knowledge the first training-free framework that leverages the high-frequency motion cues from event cameras as a continuous event-guided motion prior to guide token selection. ECP combines three stages: Event-Triggered Causal Sampling to anchor motion-informative keyframes, Event-guided Motion Saliency Filtering to remove static or low-motion regions, and Event-Attention Ranking Fusion to calibrate spatial attention with motion salient dynamics. With 80% visual token reduction, ECP outperforms the full-token baseline (37.62% vs. 36.31%) while achieving 1.89× inference speedup and 52% GFLOPs reduction. We further introduce ESR-Real, the first real-world RGB-event benchmark for first-person spatial reasoning, where ECP improves accuracy by 2.68% over full-token baselines.
Authors: Gugan Chandrashekhar Mallika Thoppe, Prashanth L.A., Ankur Naskar, Sanjay P. Bhat
Abstract: Reinforcement learning (RL) for exponential-utility optimization in discounted Markov decision processes (MDPs) lacks principled value-based algorithms. We address this gap in the fixed risk-aversion setting. Building on the Bellman-type equation for exponential utility studied in Porteus [1975], we derive two Q-value-style extensions and show that the associated operators are contractions in the L_\infty and sup-log/Thompson metrics, respectively. We characterize their fixed points and prove that the induced greedy stationary policy is optimal for the exponential-utility objective among stationary policies. These structural results lead to two model-free algorithms: a two-timescale Q-learning--style algorithm, for which we establish almost-sure convergence and provide finite-time convergence rates via timescale separation, and a one-timescale algorithm governed by a sublinear power-law operator. Since the latter does not admit a global contraction in standard metrics, we prove its convergence using delicate arguments based on local Lipschitzness, monotonicity, homogeneity, and Dini derivatives, and provide a scalar finite-time analysis that highlights the challenges in obtaining convergence rates in the vector case. Our work provides a foundation for value-based RL under exponential-utility objectives.
PaperID: 6367, Poster
Authors: Sercan Yesilkoy, Yunseok Kang, Yoon-Ho Choi
Abstract: Most energy-efficient federated learning (FL) methods report FLOPs or parameter reductions as evidence of energy savings. However, real GPU energy consumption differs across model architectures in ways that these computational proxies do not capture. In this work, we conduct an empirical study on reducing real GPU energy consumption in model-heterogeneous FL, where clients maintain distinct architectures and coordinate solely through shared prototypes. Through systematic experimentation with hardware-level energy measurements, we propose and validate two complementary mechanisms to make model-heterogeneous FL efficient in terms of real energy. First, Prototype-Aware Structured Pruning (PASP) evaluates each channel by the product of its batch-normalization scale magnitude and its gradient with respect to the prototype alignment loss---rather than the task loss---preserving channels critical for cross-client feature alignment while physically removing the rest. Second, Energy-Aware Training Scheduling (EATS) adaptively allocates local epochs to each client based on epoch-wise prototype alignment quality per unit energy, using smoothed alignment--energy profiles with a knee-point cutoff to prevent energy waste on epochs that yield diminishing alignment gains or amplify local drift. We investigate each mechanism's impact on accuracy and real energy consumption, as well as explore the impact of using both mechanisms simultaneously. All energy figures are obtained from GPU power monitoring rather than computational proxies, providing an empirically grounded evaluation of energy efficiency in heterogeneous federations.
Abstract: Reinforcement Learning (RL) is increasingly applied to large-scale decision-making problems like logistics, scheduling, and recommender systems, but existing algorithms struggle with the curse of dimensionality in such large discrete action spaces. We propose Distance-Guided Reinforcement Learning (DGRL), combining Sampled Dynamic Neighborhoods and Distance-Based Updates to enable efficient RL in problems with up to 10^20 actions. Unlike prior methods, DGRL performs stochastic volumetric exploration and transforms policy optimization into a stable regression task, decoupling gradient variance from action space cardinality. On structured tasks, DGRL provably guarantees local value improvement.DGRL naturally generalizes to hybrid continuous-discrete action spaces. We demonstrate performance improvements of up to 66% against state-of-the-art benchmarks across regularly and irregularly structured environments, while simultaneously improving convergence speed and computational complexity.
PaperID: 6369, Poster
Authors:
Chi Zhang, Guichao Chang, Sirui Liu, Xin Zhang, Guoqiang ZhongAbstract: The design of the forward diffusion process is a foundational yet under-theorized aspect of diffusion modeling. Although it determines how data are corrupted and how quickly the forward dynamics approach a reference distribution, forward-process design remains largely governed by canonical defaults rather than a unified analytical framework. We propose Eigenvalue-Guided Control and Analysis (EGCA) for diffusion models, a spectral framework that uses the principal eigenvalue of the infinitesimal generator to analyze and guide forward-process design. Under suitable ergodicity conditions, this eigenvalue governs the exponential convergence rate to stationarity, providing a compact and interpretable summary of forward dynamics. We establish sufficient uniqueness and ergodicity conditions, derive eigenvalue-based convergence bounds, and adapt a practical numerical estimator for the principal eigenvalue. Empirically, we show that the eigenvalue strongly tracks forward convergence speed and serves as an organizing coordinate for the efficiency--quality trade-off across tractable forward-generator families. Experiments across multiple image datasets and diverse diffusion settings, including training formulations, architectures, and samplers, demonstrate that eigenvalue-guided design maintains or improves generation quality while reducing training cost. Beyond a single generator choice, our framework offers a spectral perspective for understanding and comparing forward-process choices in diffusion models.
Abstract: A central question in computational vision is whether human-like visual representations are better explained by discriminative or generative learning. Existing comparisons, however, often confound the learning objective with architecture, scale, and training data, leaving open whether the objective itself drives alignment. We address this confound using Joint Energy-Based Models (JEMs), which interpolate continuously between discriminative and generative training within a fixed architecture. By varying a single mixing coefficient, we isolate the effect of the learning objective and evaluate the resulting models across six human-alignment benchmarks spanning perceptual similarity, gloss perception, human response uncertainty, robustness, shape–texture cue conflict, and diagnostic feature attribution. Across this diverse suite, human alignment is consistently maximized at intermediate points of the generative–discriminative continuum, rather than at either endpoint. Hybrid JEMs combine the categorical structure induced by discriminative learning with the sensitivity to input structure induced by generative learning, yielding more human-like behavior across multiple levels of vision. These results suggest that the generative–discriminative dichotomy is the wrong axis for understanding human-aligned vision: alignment emerges not from choosing one objective over the other, but from balancing both.
PaperID: 6371, Poster
Abstract: Bridging the gap between simulation and reality remains a fundamental challenge for end-to-end autonomous driving. Existing approaches primarily focus on appearance-level features. This often leads to suboptimal transfer, where visually aligned models still produce inconsistent or unsafe behaviors in the real world. In this paper, we propose a unified domain adaptation framework that jointly aligns perception and decision processes through geometry-aware and vector-space representations. At the perception level, we introduce a Geometry-Aware Perception Alignment (GAPA) module that enforces cross-domain consistency in both explicit geometry and implicit structural representations extracted from latent features. This reduces domain discrepancy by targeting scene structure rather than appearance. At the decision level, we propose a Latent Vector Space Guidance Decision Alignment (LVDA) module that aligns vector-conditioned policy distributions. By modeling driving decisions as conditioned on structured representations of map topology and multi-agent interactions, we minimize discrepancies between source and target action distributions via adversarial and structural alignment. To further enhance stability, we introduce a progressive adversarial transfer strategy that improves cross-domain feature alignment and training stability. Extensive experiments on a newly partitioned nuScenes dataset demonstrate that Bridge-AD achieves excellent performance, effectively narrowing the gap between simulation and real-world autonomous driving.
PaperID: 6372, Poster
Abstract: Reinforcement learning from trajectory-level outcomes has become a standard approach for training role-partitioned LLM agent pipelines, in which multiple specialized roles such as an orchestrator, a retriever, and a synthesizer collaborate to solve complex reasoning tasks. Upon trajectory failure, existing training methods do not explicitly determine which role should receive updates, applying outcome signals to all roles indiscriminately or diffusely across the pipeline. However, this paper identifies that failures can often be traced to a primary responsible role, which we term the . This suggests that updating all roles may contaminate non-bottleneck roles with irrelevant updates. Therefore, an intuitive solution is to exclusively update the bottleneck role for each failed trajectory. Unfortunately, accurately , as we show that roles exhibiting visible symptoms of failure are often merely downstream victims of errors originating from other roles, rather than being the true bottleneck. To tackle this challenge, we propose ), the first method that formulates per-trajectory update-target selection as role bottleneck inference. Specifically, RISE leverages Beta-binomial lower-confidence scores to estimate role responsibility, effectively discounting uncertain evidence and suppressing updates when attribution confidence is insufficient. Furthermore, RISE decouples the visible symptom locus from the support-constrained responsibility target prior to applying selective role-level training. Empirically, across HotpotQA, 2WikiMultihopQA, and MuSiQue with Qwen2.5-7B and Llama3-8B backbones, RISE consistently outperforms the strongest learning-based baseline, with up to
PaperID: 6373, Poster
Abstract: GNNs have achieved remarkable success in graph learning, yet their black-box nature obscures the combinatorial reasoning behind their predictions. A core challenge lies in understanding how GNNs translate topological patterns (graph concepts) into logical rules. Current works only uncover hard Boolean logical rules over graph concepts, which cannot quantify the contribution of each concept to model predictions. Moreover, they are post-hoc methods that generate explanations after training via surrogate models, and thus may deviate from the true combinatorial reasoning of GNNs. In this work, we develop the graph concept bottleneck that enforces the combinatorial reasoning of GNNs to fit soft logical rules over graph concepts, thereby quantifying the contribution of each concept. To further enhance the graph concept bottleneck, we treat graph concepts as "graph words" and graphs as "graph sentences", and leverage language models to learn context-aware graph concept embeddings. Extensive experiments on multiple datasets show that our method GCBMs achieve state-of-the-art performance in both interpretability and classification.
Abstract: In this work, we introduce FLAME, a family of extremely lightweight and capable Time Series Foundation Models, which support versatile forecasting tasks via generative probabilistic modeling, while ensuring both efficiency and robustness. FLAME utilizes the Legendre Memory for strong generalization capabilities. Through adapting variants of Legendre Memory, i.e., translated Legendre (LegT) and scaled Legendre (LegS), in the Encoding and Decoding phases, FLAME can effectively capture the inherent inductive bias within data and make efficient long-range inferences. To enhance the accuracy of probabilistic forecasting while keeping efficient, FLAME adopts a Normalizing Flow based forecasting head, which can model the arbitrarily intricate distributions over the forecasting horizon in a generative manner. Comprehensive experiments on four well-recognized benchmarks, including TSFM-Bench, ProbTS, TFB, and GIFT-EVAL, demonstrate that FLAME is a strong out-of-the-box tool of decision intelligence.
PaperID: 6375, Poster
Authors:
Xuening Wang, Bin Wang, Yucen Gao, Sihui Li, Xiaochun YangAbstract: As a variant of the Multi-Agent Pathfinding (MAPF) problem, Multi-Goal Multi-Agent Pathfinding (MG-MAPF) requires planning conflict-free paths for a team of agents to visit sequences of preassigned goal vertices. Existing search-based methods suffer from poor scalability due to their reliance on centralized control, while decentralized learning-based approaches struggle with long-horizon planning under partial observability. Furthermore, existing paradigms rarely address joint goal ordering across agents, where uncoordinated goal sequences cause multiple agents to simultaneously converge on the same region, increasing congestion and conflicts. To address these challenges, we propose the Intent-aware Multi-Goal Multi-Agent Pathfinding (IMMP), a hierarchical framework for decentralized MG-MAPF that explicitly models spatial congestion at the high level and neighbor intent at the execution level. The high-level planner encodes a spatial prior via spatially-balanced goal ordering and path planning, reducing conflict pressure passed to the low-level policy. The low-level module resolves real-time conflicts via a reinforcement learning policy augmented with a neighbor intent prediction module, trained end-to-end through a prediction-guided collaborative proximal policy optimization. Extensive experiments demonstrate that IMMP achieves strong scalability and generalization, Extensive experiments demonstrate that IMMP achieves strong scalability and generalization, reducing SoC by an average of 19.77% compared to the best-performing decentralized baseline, while maintaining robust performance. Our code is available at https://anonymous.4open.science/r/IMMP-CB2E/
PaperID: 6376, Poster
Abstract: Long-horizon LLM agents solve interactive tasks by repeatedly choosing the next action from the current state. At many steps, a compact decision is sufficient to specify the action, yet the base agent still emits a full realization in its native output format (e.g., code or action text). Once the compact decision determines the executable content of the step, subsequent generation mostly adds formatting and elaboration, yielding redundant realization. This separation is reflected inside the decoder: compact decisions become locally recoverable in middle layer representations before full realizations, and attention over its token span becomes concentrated when the decision is sufficiently supported as the current step output. Building on this separation, we propose MIRA, a training-free inference method for Model-Internal Reasoning Amortization. MIRA amortizes realization through two complementary stages: Consensus-Guided Decision retrieves successful reference steps in the middle layer representation space to propose a compact decision candidate. Attention-Guided Commitment evaluates its token span using token confidence and attention concentration. The candidate is used as the step output only when both signals support it for the current state; otherwise, the agent falls back to the base agent's full realization path. Experiments on long-horizon interactive agents show that MIRA improves task performance while reducing redundant realization, yielding a 47.5% reduction in output tokens and a 24.7% reduction in inference time across tasks and agent formats.
Abstract: Structured road understanding of lane geometry, topology, and traffic element relationships is foundational to safe autonomous driving. While vision-language models (VLMs) offer promising semantic flexibility, they lack the geometric and relational grounding required for precise road reasoning. Conversely, traditional modular systems, e.g., HD maps and topological road graphs, provide structural precision but remain semantically rigid. To bridge this gap, we introduce the Combined Road Substrate (CRS), a graph-grounded framework that makes geometric road structure and open-vocabulary semantics jointly executable in a single representation. CRS enables the automatic generation of compositionally complex and linguistically varied question-answer pairs via recursive graph queries, augmented with a ``grounding for free'' mechanism that ensures logical traceability to specific map elements, and procedurally extracted chain-of-thought supervision traces. We demonstrate that state-of-the-art VLMs - including large, closed-source models - struggle significantly with structured road reasoning, yet training a small 2- or 4-billion-parameter model with as few as 20 to 80 CRS-enriched scenes yields stable gains in compositional reasoning tasks of varying depth. Analysis of model behavior via verifiable reasoning traces reveals a systematic shift in failure modes: whereas baseline models fail at relational scene understanding, CRS-trained models reduce failures to attribute recognition, suggesting that the primary bottleneck in road understanding is not model scale, but the absence of structured supervision.
PaperID: 6378, Poster
Abstract: Diffusion models have recently emerged as a powerful paradigm for generating 3D molecules. Despite showing remarkable promise, they learn atomic distributions implicitly from data, without explicitly encoding physicochemical rules into their generative dynamics. Consequently, generated samples can be statistically plausible yet chemically invalid. To address this problem, we propose Sequential Neuro-Symbolic Constrained Diffusion (SensDiff), a framework that integrates symbolic rules into generative modeling, enabling the model to internalize scientific priors during training. Central to SensDiff is the observation that the diffusion trajectory of 3D molecule generation exhibits a coarse-to-fine evolution: valid molecules arise by first mitigating macroscopic geometric conflicts and then refining structures toward microscopic chemical consistency through iterative denoising. Accordingly, SensDiff mirrors this dynamic by sequentially enforcing constraints with adaptive weighting during training, progressively steering generation toward valid molecular geometries. Moreover, SensDiff translates domain knowledge into interpretable generative controls, grounding opaque neural denoising in scientific principles. Experiments on molecular benchmarks and real-world constrained generation tasks confirm that SensDiff consistently improves chemical validity while adhering to physicochemical rules and user-specified targets.
PaperID: 6379, Poster
Abstract: Existing underwater image enhancement (UIE) methods predominantly rely on unidirectional mapping, which lacks often results in over-enhancement or under-enhancement, without user-friendly controllability. To address this, we propose a paradigm of unsupervised trajectory learning for Omnidirectional Controllable UIE that treats restoration as a omnidirectional dynamical system, namely OC-UIE. The main purpose of our work is the realization of a consistent omnidirectional mapping across the representation space, which allows the model to master complex underwater dynamics beyond traditional unidirectional constraints. To implement this, we propose an Omnidirectional Training that optimizes over arbitrary source-target positions sampled from the data distribution. We decompose this purpose into two synergistic stages: first, Representation Trajectory Flattening is employed to organize intricate degradations into a unified adaptation axis; second, Trajectory Omnidirectional Integration is introduced to model the enhancement as a path integral over the resulting manifold. By optimizing only the relative shifts, OC-UIE isolates structural preservation from fluid degradation dynamics. This approach provides a controllable space for intensity-parameterized enhancement with state-of-the-art performance.
PaperID: 6380, Poster
Authors:
Mengyuan Fan, Bokai Huang, JiaMing Pan, Xiaokun Yuan, Peizhuang Cong, Zhewen Tan, Tong YangAbstract: Vision Transformers (ViTs) achieve strong performance on image recognition and mobile vision applications, but their high-dimensional linear projections and attention computations still impose substantial storage and inference costs. Extremely low-bit quantization is a promising solution, yet ViTs often suffer severe accuracy degradation because conventional real-valued scalar codebooks are poorly matched to the directional geometry of Transformer projections. We present RPFQ-ViT, a Rotated Phase-Frame Quantization method that quantizes paired channels in two-dimensional phase planes, enabling low-bit codes to better preserve projection directions while recovering magnitude with lightweight scaling. RPFQ-ViT serves as a drop-in QAT replacement for \textttnn.Linear and does not modify the standard real-valued attention, normalization, or activation computation graph. On ImageNet-1K, RPFQ-ViT-B/16 reaches 79.33% Top-1 / 94.48% Top-5 under W2/A4, Swin-T reaches 79.30% Top-1 / 94.79% Top-5 under W2/A8, and DeiT-S reaches 77.41% Top-1 / 93.11% Top-5 under W2/A8. Ablations, phase-geometry analysis, and direction-preservation metrics show that channel pairing, learnable rotation, phase-anchor learning, and residual phase refinement each improve quantization quality. We further deploy RPFQ-ViT image-classification models on native iOS and Android runtime stacks; with 2-bit packed weights, model size shrinks by roughly 5.4--7.1× relative to FP32 and end-to-end on-device latency drops by 1.4--1.6×. All ImageNet results trained in our codebase use a matched 300-epoch recipe and are reported as mean accuracies over three independent runs. These results show that RPFQ-ViT provides a favorable trade-off among accuracy, compression, and practical mobile deployment for extremely low-bit ViTs.
PaperID: 6381, Poster
Abstract: The method of types, pioneered by Csiszar (1998), uses combinatorial arguments to approximate the probabilities of empirical measures and has become a standard tool for deriving error rates in hypothesis testing, source coding, and related information-theoretic problems. Although the technique is widely used across machine learning, computer science, and econometrics, rigorous type-theoretic guarantees for stationary Markov decision processes and reinforcement learning remain limited. This paper addresses that gap by developing a method-of-types framework for the empirical pair measure of the state--action process of a stationary Markov decision process. We establish finite-sample type bounds for the state--action process and characterize the associated entropy cost through the conditional Kullback--Leibler divergence relative to the true pair kernel. These results yield large deviation principles for learned MDP dynamics and variational rate functions for plug-in value- and Q-function estimators in offline reinforcement learning. We further develop finite-sample uncertainty-quantification and hypothesis-testing guarantees based on KL uncertainty sets, including applications to pessimistic value and Q-function estimation under model misspecification. Finally, we apply the same framework to universal source coding and obtain explicit coding error exponents for controlled Markov trajectories. Together, these results provide a unified type-theoretic approach to rare-event analysis, uncertainty quantification, offline reinforcement learning, and information-theoretic coding for finite controlled Markov systems.
PaperID: 6382, Poster
Abstract: Human motion change is currently hindered by the lack of large-scale datasets, comprehensive benchmarks, and robust algorithms capable of handling diverse motion requirements. To address these gaps, we propose AnyMotion, a novel open-world paradigm for highly controllable and identity-consistent human motion editing across arbitrary scenes and actions. Our contributions are two-fold. First, we establish a hierarchical label taxonomy covering six major categories and 527 fine-grained classes, and curate HME-1M, a large-scale dataset containing 1.04 million editing quadruplets via systematic data balancing, cleaning, and reverse construction. Second, we introduce the Skeleton Chain-of-Thought (SCoT) framework, a two-stage pipeline consisting of motion reasoning and guided generation. In the reasoning stage, a multimodal large language model serves as a cognitive planner to derive target skeletal keypoints and MetaQueries. These outputs subsequently function as explicit structural and semantic priors, providing the necessary guidance for the diffusion transformer to perform high-fidelity motion generation. Finally, built upon our label taxonomy, we present Human Motion-Bench, the first motion-centric benchmark equipped with multi-category fine-grained evaluation metrics. Extensive experiments demonstrate that AnyMotion achieves state-of-the-art performance across diverse editing scenarios.
PaperID: 6383, Poster
Abstract: Activation steering adds a vector to a language model's internal representations at inference time. It is a lightweight alternative to fine-tuning for behavioral control. The standard construction, _contrastive activation addition_ (CAA), is fragile, as on many inputs it shifts behavior in the wrong direction. To explain this fragility, in this paper we first study the geometry of the steering vector. Specifically, we introduce steering Hessian, _the matrix of second order derivative of the model's behavioral loss with respect to the steering vector_. It captures the output behavior response sensitivity of the model under small perturbations of the steering vector along each direction. Its top eigenvectors are _sharp directions_: tiny perturbations along them can cause large, erratic, input-dependent behavioral swings. On the other hand, bottom eigenvectors are _flat directions_: behavior is robust along them. We identify that removing the sharp component of a steering vector approximately preserves its direction and magnitude (cosine > 0.974, norm reduction < 3% for typical settings) while reducing the vector's directional curvature by up to 12× (median 4× across 36 conditions). We evaluate on three model families (Qwen2.5-7B, Llama-3.1-8B, Mistral-7B), with four behavioral tasks, and 36 token-level conditions. Steering along sharp directions reverses the intended behavior in every condition tested, with anti-steer rates of 53--100%. Sharp directions account for only 1--5% of the CAA vector's squared norm. However, they drive a disproportionate share of the gap between the steered and unsteered next-token distributions during autoregressive generation. We term this one-line correction as _sharp removal_; it outperforms standard CAA in open-ended generation across all 25 conditions tested. A jailbreaking experiment further shows that sharp directions act as _noise_ rather than signal: they disrupt coherent behavioral control regardless of steering polarity.
PaperID: 6384, Poster
Abstract: Video Large Language Models face severe computational and memory bottlenecks in long-video understanding. Existing visual token compression methods predominantly rely on static, query-agnostic redundancy elimination. Consequently, they lack the ability to dynamically preserve query-relevant details, failing to emulate the human "coarse-to-fine" cognitive process. To bridge this gap, we propose Q-Focus, a plug-and-play, two-stage visual token compression framework. Q-Focus first constructs a holistic semantic representation via a global compression stage. This is followed by a question-guided focusing stage driven by two novel components: the Adaptive Contrastive Module (ACM) and the Question Context Module (QCM). Specifically, ACM achieves adaptive query-visual alignment by treating the user queries as positive samples and dynamically sampling negative samples from the video's feature distribution. Complementarily, QCM leverages visual affinity matrices to diffuse semantic relevance from highly matched visual anchors to surrounding contextual regions, effectively capturing implicit yet critical details. Extensive experiments demonstrate that Q-Focus seamlessly integrates with diverse compression baselines. Remarkably, while retaining only 10% of the original visual tokens, it achieves relative accuracy gains of 7.3% and 6.6% on LLaVA-OneVision and LLaVA-Video, respectively, and consistently boosts the performance across other state-of-the-art models such as Qwen3-VL-8B-Instruct.
Authors: Benedict Russell, Chin-wing Leung, Paolo Turrini
Abstract: In social dilemmas self-interested learning agents face the choice between the societal benefit of cooperation and the immediate reward of defection. Significant evidence exists on the benefits of assortment mechanisms such as partner selection for the emergence of cooperation, but this is largely available through agent-based simulations. In this paper, we provide an analytical solution to the problem, studying the policy-gradient dynamics in a multi-agent environment with partner selection. We show how partner selection changes the opponent distribution and hence the reward landscape, and prove this promotes cooperation under simple rules known from the literature. In particular, we find that population variance is a necessary condition for cooperation to emerge. Using a two-dimensional Wiener process, we extend the dynamics to capture the stochastic effects of partner selection and the resulting opponent distribution. We derive a sufficient condition for the population to be cooperation-promoting and prove the existence of a stationary distribution. Simulations confirm that the stochastic model accurately captures the policy-gradient dynamics and clarifies how the learning rate affects the emergence of cooperation.
PaperID: 6386, Poster
Authors: Tristan Bilot, Xueyuan Han, Thomas Pasquier
Abstract: Provenance graphs capture causal interactions among operating-system entities, underpinning modern endpoint intrusion detection. Existing graph-based detectors use generic text encoders trained on small datasets to embed entities solely from their labels (e.g., file paths and process command lines). Embeddings are thus independent of the graph, missing behavioral semantics encoded in graph neighborhoods. However, much like a natural language that carries linguistic structures that can be learned once and transferred broadly, OS entities such as system binaries, configuration files, and common services engage in recurring types of activity with similar inter-entity relationships across executions. These behavioral regularities can be learned offline from provenance graphs and approximated at inference time from entity labels alone. Based on this insight, we introduce SPIDER (System Provenance-Informed Distilled Entity Representations), a text-only entity encoder distilled from provenance graphs. SPIDER uses a graph-based teacher that models who an entity interacts with via graph attention and how it interacts with them through behavioral signatures. These representations are distilled into a transformer-based student that maps raw entity text to embeddings in a single forward pass, providing pretrained behavioral priors that downstream detectors can use for real-time inference without requiring graph access. With only 5.2M parameters, SPIDER can be deployed directly on endpoints and optionally fine-tuned for specific detection tasks. We show that a single checkpoint improves detection performance as a drop-in replacement for four recent intrusion detection systems across Linux, FreeBSD, Windows, and Android, surpassing both GNN- and LLM-based baselines.
Authors: Yihan Zang, Da Li, Dominik Engel, Shinkyu Park, Ivan Viola
Abstract: Training-free weighted aggregation is widely used to lift 2D semantic features onto 3D Gaussians for open-vocabulary scene understanding, yet its theoretical role remains insufficiently understood. Existing analyses typically justify this operation from the rendering side, treating Gaussian features as linearly composable Euclidean variables for reconstructing 2D feature maps. However, this view does not match downstream 3D usage, where each Gaussian is often queried independently in a cosine-based embedding space. We revisit feature lifting from the 3D side and formulate per-Gaussian assignment as a cosine alignment problem on the CLIP unit sphere. Under this objective, the \ell_2-normalized semantic back-projected feature emerges as the closed-form solution, providing a complementary interpretation of the standard lifting rule from the perspective of per-Gaussian semantic assignment. The same formulation further yields a norm decomposition into intra-view and inter-view consistency, suggesting that feature magnitude itself can serve as a semantic reliability signal. Calibrated by effective multi-view support, this reliability score guides a mode-voting refinement that preserves CLIP feature validity by avoiding linear averaging. Experiments on open-vocabulary 3D semantic segmentation show that NormLift is an efficient, training-free framework that achieves strong performance across evaluation protocols.
PaperID: 6388, Poster
Authors: Seungwook Kim, Jongmin Lee
Abstract: Noise search at inference time improves text-to-image diffusion quality but requires many multi-step teacher rollouts per query. Distilled few-step models are now routinely released for efficient inference, but their potential as proxies inside noise search is uncharacterised. We measure proxy--teacher rank agreement across three distillation regimes and three backbones spanning DiT, MMDiT, and UNet, and find a sharp asymmetry: eight closed-form image statistics preserve rank from the distilled proxy to the teacher (Spearman \rho up to 0.80), while learned semantic scorers decorrelate (\rho \le 0.12). Null controls localise the transfer to distillation training rather than few-step denoising or initial noise characteristics. We therefore reformulate noise search as a \emphjoint constraint and propose ProxySearch: the proxy filters candidates by statistical match in 1 step, the teacher argmaxes a user-chosen semantic reward over the survivors without a tunable weight trading semantic reward against the statistical target. \textscProxySearch matches or beats teacher-only argmax on compositional accuracy and human preference, while tightening the statistical target by ~ 23% in paired statistical \ell_1 relative to teacher-only argmax at matched compute on FLUX-1024.
Abstract: Concept-bottleneck models (CBMs) are neural classifiers that compute predictions from high-level concepts extracted from the input. CBMs ensure stakeholders can understand the concepts — and the predictions they entail — by learning these from concept-level annotations, which are however seldom available. Recent CBM architectures work around this issue by obtaining annotations from Vision-Language Models (VLMs). While greatly broadening applicability, doing so can yield lower quality concepts and therefore less interpretable models. We strike for a middle ground by introducing approach that exploits both VLMs and a small amount of dense annotations. VH-CBM employs a Gaussian Process in the VLM's embedding space, which captures useful global information about the target domain, to propagate the expert's supervision to any target data point. Our empirical evaluation on four datasets showcases how predicts substantially more accurate concepts than VLM-guided CBMs even when annotating as little as
PaperID: 6390, Poster
Abstract: Diffusion-based foundation models for crystal generation have become competent priors over the manifold of plausible materials, but turning that prior into generation of promising candidates requires targeting stability while balancing several conflicting properties at once. Fine-tuning is expensive and produces a new model for every new objective, while differentiable guidance is not applicable to many real-world rewards. We instead adapt these models at \emphinference time using Diffusion Tree Search (DTS), a Monte Carlo tree search procedure over the denoising trajectory. We make two technical contributions on top of DTS that are particularly relevant for AI-for-science applications. First, we show how to use existing property-conditional foundation models (e.g., classifier-free guidance variants of MatterGen) as guided proposals inside DTS, with importance-corrected backups and selection that keep the unconditional base model as the target distribution. Second, we extend DTS to multi-objective sampling by maintaining per-objective soft value estimates and selecting expansions through a sampled scalarization, so a single search produces samples spanning a Pareto front while sharing computation across preferences. We instantiate both adaptations on de novo crystal generation: single-objective stability search with different underlying models, conditional generation of crystals belonging to rare space groups, and a multi-objective tasks using a pretrained conditional model. Across these, DTS-based inference-time adaptation improves stability rates and Pareto-front quality without modifying the base models.
PaperID: 6391, Poster
Abstract: Class-incremental learning (CIL) with pre-trained models has increasingly adopted Mixture-of-Experts (MoE) adapters, which freeze the backbone and selectively route inputs to sparse adapter subsets, achieving competitive performance through efficient parameter reuse. However, as new tasks are learned sequentially, the router's output distribution gradually shifts away from previous states, undermining consistent adapter reuse and accelerating forgetting. Through a controlled instance-level analysis, we show that this effect is highly asymmetric, where instances with large routing drift account for most of the performance degradation, while those with small drift remain stable or can even benefit from it. This finding suggests that only excessive drift requires correction, while small drift may be better preserved than suppressed, and motivates us to propose Trust-Region Projected Routing (TRPR) that adapts the trust-region principle from constrained optimization to routing stabilization. It constrains each input's routing distribution within a KL-bounded trust region around a class-level anchor, correcting only excessive deviations while leaving small drift intact. To complement class-level anchoring with per-sample adaptability, TRPR additionally introduces an Instance-Adaptive (IA) path for fine-grained per-sample specialization, fused with the projected path via a learned gate. Extensive experiments on four CIL benchmarks show that TRPR consistently outperforms nine representative PEFT-based baselines, with accuracy gains of up to 3.06% in average and final accuracy and forgetting reduced by up to 2.48%.
Abstract: Reinforcement learning methods for LLM reasoning typically assign a uniform advantage to every token in a trajectory, diluting the learning signal at pivotal reasoning steps and injecting noise at uninformative positions. Critic-free alternatives such as anchored distillation derive per-token signals from oracle-conditioned likelihood ratios, but apply the signal independently at each position. We propose Oracle-Prompted Policy Optimization (\n), which recovers exact token-level advantages through a Bayesian value recursion. Conditioning the model on the ground-truth answer during training yields per-token likelihood ratios that measure how much the answer revises the model's prediction; the ratios accumulate via Bayesian updating into a running estimate of the success probability at every position, from which the token-level advantage follows in closed form. The advantage factors as a product of a state weight, peaking where the outcome is most uncertain and vanishing where success or failure is already determined, and the per-token oracle evidence familiar from on-policy distillation, with no tunable weighting hyperparameter. The framework supports two estimator choices: a self-oracle mode in which the policy model provides the estimate through one additional forward pass, and a teacher-oracle mode in which a stronger model serves as the estimator. Experiments on two base models across seven reasoning benchmarks spanning mathematics, science, and code demonstrate consistent improvements over GRPO, DAPO, and SDPO, with the largest gains on competition-level tasks where reasoning chains are longest.
PaperID: 6393, Poster
Authors: Siquan Huang, Yijiang Li, Xingfu Yan, Ningzhi Gao, Leyu Shi, Ying Gao
Abstract: Vision Transformers (ViTs) have been widely adopted as visual encoders in multimodal models; however, the reliance on third-party pretrained checkpoints exposes these systems to backdoor attacks. Existing test-time detection methods rely on either the predicted label or intermediate embeddings, limiting their applicability across diverse tasks and incurring substantial computational overhead. In this work, we investigate how backdoor triggers affect the multi-head attention mechanism of ViTs and identify two distinct anomalous patterns: complete attention hijacking, causing the attention to exhibit abnormally high inter-head spatial overlap, and partial attention hijacking, causing an abnormally large variance in the concentration of the attention distribution. Motivated by these findings, we propose DetectViT, a test-time backdoor detection method that requires neither training data, model outputs, nor any prior knowledge of the trigger. DetectViT quantifies the two hijacking phenomena via an inter-head consistency score and an inter-head entropy variance score, with thresholds estimated from a small out-of-distribution (OOD) calibration set drawn from any source. Extensive experiments on four representative attacks, spanning diverse architectures (DeiT, CLIP, LLaVA) and downstream tasks (image classification and captioning), demonstrate that DetectViT substantially outperforms all baselines with as few as 64--128 OOD calibration samples, attaining a true positive rate of 100% and a false positive rate of 0.17% on BadCLIP. Notably, it introduces negligible additional overhead, thanks to its reliance solely on attention weights already computed during inference. We release our code at: https://anonymous.4open.science/r/DetectViT-81AC.
PaperID: 6394, Poster
Authors: Ankita Kushwaha, KIRAN RAVISH, Preeti, Pawan Kumar
Abstract: Safe multi-agent reinforcement learning usually constrains team-level safety cost, but a safe team can still assign the same agent to repeatedly wait, yield, detour, absorb intervention, or lose access to a scarce resource. We study this failure mode as burden unfairness under hard safety. Safe-Fair multi-agent constrained policy optimization (Safe-Fair MACPO) augments MACPO with agent-wise burden logging, a temporal fairness-debt state, a fairness critic, and a two-cost trust-region update. Safety is lexicographically primary: recovery may override fairness to avoid unsafe actions, while the resulting imbalance is stored as debt and penalized later. We give a safety-first Pareto theory showing that the exact safe-fair constrained selector is strongly Pareto efficient for reward, safety, and fairness, with hard-safety and temporal burden-gap certificates. Empirically, the primary reported rows keep zero hard violations while improving useful burden balance: HalfCheetah-2x3-Fair improves selected safe return from 2145.5 to 2637.4 and Jain burden from 0.903 to 0.977; Safety-Gym MultiGoal-Fair reaches the lowest matched-budget scalar fairness cost with zero safety cost; and VMAS-SafeMultiGoal improves over recovery-only attribution controls. We separate social/resource fairness from mechanical workload imbalance and reserve mixed or recovery-explained rows for qualified appendix analysis.
Authors: Sergio Mauricio Vanegas Arias, Lasse Lensu, Fredy Ruiz
Abstract: Building black-box models for dynamical systems from data is a challenging problem in machine learning, especially when asymptotic stability guarantees are required. In this paper, we introduce a novel stability-ensuring and backpropagation-compatible projection scheme based on the Schur decomposition for the state matrix of linear discrete-time state-space layers, as well as an alternative pre-factorized formulation of the methodology. The proposed methods dynamically project the quasi-triangular factor of the state matrix's real Schur decomposition onto its nearest stable peer, ensuring stable dynamics with minimal overparameterization. Experiments on synthetic linear systems demonstrate that the method achieves accuracy and convergence rates comparable to those of state-of-the-art stable-system identification techniques, despite a marginal increase in computational complexity. Furthermore, the lower weight count facilitates convergence during training without sacrificing accuracy in stacked neural-network architectures with static nonlinearities targeting real-world datasets. These results suggest that the Schur-based projection provides a numerically robust framework for identifying complex dynamics on par with the State of the Art while satisfying strict asymptotic-stability requirements.
PaperID: 6396, Poster
Abstract: Panoramic 3D reconstruction enables the holistic recovery of complete 360° scene geometry, offering significant benefits for robotics and AR/VR applications; however, direct estimation is persistently challenged by inherent spherical distortions. While decomposing panoramas into perspective views effectively circumvents these distortions and leverages established geometric priors, this strategy often fails to preserve global spatial consistency across the decomposed views. To address this fundamental limitation, we introduce an enhanced feed-forward 3D reconstruction framework tailored for the view decomposition paradigm, driven by two key innovations. First, we propose a geometry-aware attention aggregator, featuring a streamlined attention architecture equipped with a novel Geo-Attention mechanism, which explicitly exploits the intrinsic global information to guide multi-view interactions, thereby suppressing geometrically inconsistent matches. Furthermore, to jointly enforce fine-grained local details and global structural coherence, we formulate a hybrid supervision strategy. Specifically, we complement the standard local point cloud loss in perspective space with our newly proposed spatially uniform global point cloud loss in panoramic space, explicitly enforcing holistic constraints within a unified coordinate system. Experiments demonstrate that our method consistently outperforms prior single-view panoramic approaches in reconstruction accuracy, geometric consistency, and visual fidelity.
PaperID: 6397, Poster
Abstract: LLMs are widely adopted in multi-role dialogue and persona simulation. However, their strong ability to imitate role-specific linguistic and behavioral patterns can introduce privacy risks, reinforce social biases, and lead to unsafe or uncontrollable outputs. To address these issues, we introduce the Role-playing Unlearning task, which aims to selectively forget the style and knowledge associated with target roles while preserving the language ability and behaviors of the remaining roles. Unlike prior unlearning tasks that focus solely on removing factual information, our task additionally requires forgetting persona-specific characteristics, making the control over model generation more fine-grained. We propose ROLEMOE, a Mixture-of-Experts framework that enhances the separation of role-specific patterns by mapping different roles to more independent expert subspaces, providing structural interpretability. To reinforce this separation, ROLEMOE incorporates two complementary losses: a specialization loss that drives different roles to adopt different expert distributions, and a disentanglement loss that encourages different experts to develop distinct capability subspaces. This decomposed structure enables role unlearning to focus selectively on target-role experts. We finally propose RoleBench-Unlearn for evaluation, and experiments across multiple LLM architectures show that ROLEMOE achieves significantly more effective and controllable role forgetting, while better preserving the linguistic quality, knowledge, and persona consistency of non-target roles. All code and datasets are publicly available at: https://anonymous.4open.science/r/RoleMoE-6328.
Abstract: Reinforcement-learning-based post-training for large language models vastly improves capabilities on verifiable tasks however it suffers from training instability and diversity collapse. As a solution, many new advantage functions have emerged. Each method simultaneously changes which problems receive gradient, the balance between positive and negative updates, and the overall gradient scale making their comparison difficult. We propose a unifying framework that decomposes each function by their induced positive and negative gradient mass. Analyzing policy performance and the geometry of model weights reveals: symmetric methods are better but slow and asymmetric methods are fast. Combining the advantages of both, we build the FADE: Focal Adaptive Dynamic Entropy advantage which is fast, diverse and accurate at various model scales (7B, 32B).
Abstract: Modern video generators produce visually compelling clips but still struggle with physical and motion consistency, limiting their use as reliable world simulators. Existing remedies often rely on external simulators, teacher models, or curated physics-focused data. We explore a complementary self-supervised direction: extracting motion cues from the unlabeled videos already used to train video diffusion models. We propose LaMo, which formulates a latent motion prior over frame-to-frame latent changes conditioned on the current latent and prompt. This prior is exposed through two lightweight readouts: a macro motion drift used during training as a Motion Drift Loss, and a learned micro motion field used during sampling as Motion Prior Guidance. Both components are plug-and-play with existing video diffusion backbones, requiring no architectural or I/O changes. On VideoPhy and VideoPhy2, LaMo improves CogVideoX backbones and outperforms recent physics-aware baselines that use external supervision. On VBench, it preserves overall generation quality while improving motion-related dimensions. These results suggest that unlabeled video contains useful motion supervision for improving physical fidelity in modern video diffusion models.
Abstract: We study optimal design of \varepsilon-locally differentially private mechanisms for binary hypothesis testing. Each observation is drawn from one of two known distributions P_0,P_1 on a finite alphabet of size k, privatized by a mechanism Q, and then used to infer which distribution generated the data. We measure testing utility using an f-divergence—including total variation, KL, and hockey-stick divergences—between the two induced output distributions. Previous work established structural properties of optimal mechanisms, but only yielded exponential-time algorithms. We prove a sharp structure: for every \varepsilon and every f-divergence objective, after sorting the alphabet by likelihood ratio, there exists an optimal mechanism that partitions the sorted alphabet into contiguous blocks and applies randomized response to the block label. We call this class Sort–Partition–Randomize (SPR). This characterization yields an exact dynamic program that computes an optimal mechanism in O(k^3) time, and more generally in O(\ell k^2) time with an \ell-output budget. Our results make it possible to efficiently compute and characterize the exact optimum across the full privacy range, beyond asymptotic privacy regimes.
PaperID: 6401, Poster
Authors: Muhammad Usama, Summer Y Jung
Abstract: Existing depth-compression methods for large language models (LLMs) operate on layers individually or restrict grouping to adjacent or contiguous positions, implicitly assuming functional redundancy is local. We challenge this assumption by computing full pairwise Centered Kernel Alignment (CKA) similarity matrices across thirteen LLMs spanning six architecture families and 1.6B--14B parameters, revealing that non-adjacent layers exhibit 3.6--15.9× more high-similarity pairs than adjacent ones. We propose Fusion via Community Mining (FCM), which builds the full L × L similarity graph, applies spectral community detection to recover clusters of functionally equivalent layers, retains the highest-Block-Influence layer per cluster, and applies LoRA recovery. At 2× compression, FCM lowers mean perplexity below ShortGPT on twelve of thirteen models, with statistically significant improvements on ten under a Welch's t-test with Holm--Bonferroni correction (up to 50% on Qwen3-4B, p_\textcorr<0.001); two rows are ties and Mistral-7B is the lone loss (+14.3%). On three anchor models, FCM also outperforms three recent adjacent-pair methods (FlattenGPT, SWM, DeltaLLM) by 7.8--35.3% under matched calibration and recovery. The perplexity advantage transfers downstream: on a five-benchmark zero-shot reasoning suite (ARC-Easy, ARC-Challenge, HellaSwag, WinoGrande, PIQA) over four anchor models, FCM matches or exceeds ShortGPT on 15 of 20 (model, benchmark) cells, and the 2×-compressed Qwen2.5-7B realises 46% fewer parameters, 42% lower peak VRAM, and 84% higher decode throughput on a single A100. Ablations confirm spectral grouping outperforms Fisher contiguous segmentation by 41% and uniform partitioning by 51%, and a compression sweep reveals a 2× crossover beyond which FCM's advantage grows monotonically and persists at the 14B scale. Code: \urlhttps://anonymous.4open.science/r/fcm-layer-fusion-D820.
Abstract: Reinforcement learning (RL) is a key paradigm for post-training large language models (LLMs), but the widely used Group Relative Policy Optimization (GRPO) often suffers from entropy collapse: exploration quickly disappears, policies converge prematurely, and sample diversity declines, ultimately harming training effectiveness. Existing remedies, including entropy bonuses and clip-based methods, rarely keep entropy within a stable exploration regime and often introduce oscillatory entropy or reward degradation. In this work, we identify a previously overlooked asymmetry in entropy dynamics: under high-temperature sampling, positive and negative samples have opposite effects on policy entropy. Specifically, high-temperature positive samples promote entropy growth, whereas negative samples suppress it. We provide a theoretical explanation for this phenomenon: when entropy decreases during policy updates, its derivative with respect to temperature is strictly positive under positive-sample updates, indicating that high-temperature positive samples can counteract entropy decay, thereby slowing entropy collapse and potentially reversing it. Motivated by this insight, we propose SCOPE-RL, a stable and quantitative entropy control framework through a regularization term constructed from temperature-adaptive positive samples. Extensive experiments show that SCOPE-RL consistently outperforms strong RL baselines on both Pass@1 and Pass@k. Our results provide evidence that escaping entropy collapse can improve reasoning performance, while also showing that the benefit is non-monotonic, with an optimal level of exploration for RL post-training in reasoning LLMs.
Authors:
Yuxuan Jiang, Runchao Li, Shubhashis Roy Dipta, Dawei Li, Zhao YangAbstract: While recent work in Reinforcement Learning with Verifiable Rewards (RLVR) has shown that a small subset of critical tokens disproportionately drives reasoning gains, an analogous token-level understanding of On-Policy Distillation (OPD) remains largely unexplored. In this work, we investigate high-loss tokens, a token type that—as the most direct signal of student-teacher mismatch under OPD's per-token KL objective—should progressively diminish as training converges according to existing studies; however, our empirical analysis shows otherwise. Even after OPD training reaches apparent saturation, a substantial subset of tokens continues to exhibit persistently high loss; these tokens, which we term Rock Tokens, can account for up to 50% of the tokens in generated outputs. Our investigation reveals two startling paradoxes. First, despite their high occurrence frequency providing a disproportionately large share of total gradient norms, Rock Tokens themselves remain stagnant throughout training, resisting teacher-driven corrections. Second, through causal intervention, we find that these tokens provide negligible functional contribution to the model's actual reasoning performance. These findings suggest that a vast amount of optimization bandwidth is spent on structural and discourse residuals that the student model cannot or need not internalize. By deconstructing these dynamics, we demonstrate that strategically bypassing these "stumbling blocks" can significantly streamline the alignment process, challenging the necessity of uniform token weighting and offering a more efficient paradigm for large-scale model distillation.
PaperID: 6404, Poster
Abstract: Neural networks are often trained with gradient descent (GD) at the edge of stability (EoS), where step sizes exceed classical stability thresholds and standard convergence theory no longer applies. Despite exhibiting non-monotonic loss and oscillatory dynamics, this regime frequently leads to faster convergence and better generalization. Existing theoretical perspectives offer complementary but incomplete insights: one identifies a self-stabilization mechanism that drives GD toward flat regions but does not yield convergence rates; another derives rigorous rates for specific losses but does not explain the underlying stabilization mechanism; and a third unifies these views for overparameterized least squares, but relies on the existence of a minimizer manifold, which is absent in classification settings. In this work, we extend this unified geometric perspective to exponential-tail losses, including logistic and cross-entropy losses, where no finite minimizer exists. We identify a curvature-defined reference subspace that replaces the role of the minimizer manifold, yielding a coupled dynamical system in which orthogonal oscillations are damped by parallel progress along directions of decreasing sharpness. Under regularity conditions, which we verify for logistic and multiclass cross-entropy losses in symmetric data models, we derive explicit bounds for the iterates, not just the loss, covering both the EoS and stable regimes, and revealing a three-phase damping structure in the latter. Our analysis further uncovers a transient implicit bias induced by large step sizes: although GD converges asymptotically to the max-margin direction, large steps can leave persistent orthogonal residuals and, in certain geometries, amplify them.
PaperID: 6405, Poster
Abstract: Joint-Embedding Predictive Architectures (JEPAs) are prone to representation collapse, typically mitigated through empirical heuristics. We develop an early-training stability theory that unifies these heuristics. Linearising the coupled JEPA gradient flow around the trivial fixed point reveals two competing effects: a driving force (\gamma) and a decay effect (\sigma). Under approximate spectral decoupling, a per-mode stability ratio \mu_i = \gamma_i / \sigma_i factorises into independent data-side and predictor-side terms and the count of unstable modes tracks the rank of representations that can emerge. The framework predicts a phase boundary, which we confirm empirically across more than 800 Tabular-JEPA configurations. It also unifies predictor scaling, masking ratio, and EMA as distinct mechanisms for shifting \mu. Guided by this analysis, we introduce ResidualPred, a transformer predictor whose attention is biased toward the identity at initialisation; it improves both effective rank and downstream accuracy on tabular benchmarks and on I-JEPA pretraining over CIFAR-10, CIFAR-100, and STL-10. Our framework connects empirical collapse-avoidance heuristics to an explicit dynamical picture, yielding theory-driven stabilizers. Code is available as supplementary material.
PaperID: 6406, Poster
Abstract: Multimodal Large Language Models (MLLMs) have significantly advanced instruction-driven perception through the ``Embedding-as-Mask'' paradigm. However, extending this capability to Remote Sensing (RS) reasoning segmentation remains exceptionally challenging under cluttered geographical contexts and extreme scale variations, often leading to Spatial Fidelity Degradation and Target Misidentification. We identify a critical yet underexplored bottleneck: the sequential alienation of the segmentation token \texttt[SEG] from its initial visual grounding within the autoregressive generation paradigm. Produced only at the end of long reasoning chains, \texttt[SEG] relies on visual evidence that has been progressively diluted by linguistic abstraction and autoregressive noise. This information decay is particularly catastrophic for RS targets with minuscule pixel footprints, where even marginal fidelity loss leads to localization failure. To address these bottlenecks, we propose GeoReason, a novel framework designed to fortify logical deduction and preserve spatial integrity. First, we introduce an \emphAnticipatory Prior mechanism that injects the \texttt[SEG] token directly into the initial user query, shifting it from a terminal output to a primary condition. This paradigm shift transforms the token into a persistent spatial anchor, preventing information decay during long-form linguistic inference. Second, we enhance visual reasoning via a \emphLocation-aware Chain-of-Thought (CoT) that enforces coarse-to-fine localization using bounding boxes, coupled with a \emphLatent Patch-to-Token Alignment (PTA) Loss that explicitly aligns the latent segmentation token with physical RS textures. Extensive experiments demonstrate that GeoReason consistently outperforms state-of-the-art methods on the Earthreason and RRSIS-D benchmarks.
PaperID: 6407, Poster
Abstract: Automated semantic distance tests—which prompt a model to produce a set of words, and score the average embedding distance between them—are increasingly used to measure the "creativity" of large language models (LLMs). However, the validity of semantic distance tests as predictors of creativity has not yet been established, and these tests already have limited validity as predictors of human creativity. To address this problem, we conduct the first systematic study evaluating the effectiveness of semantic distance tests in predicting the creative achievement of LLMs across three constructs: creative writing, divergent thinking, and scientific ideation. We score each test on two criteria ( ), and derive a theoretical limit for the maximum attainable specificity and validity a test can achieve. We find that: Test effectiveness varies significantly by construct, and no single test predicts all constructs well. Existing tests are far below the theoretical limits, indicating meaningful room for the design of improved tests moving forward. Our findings provide clear practical takeaways and directions for future work and suggest that novel tests are needed to reliably predict scientific ideation ability.
PaperID: 6408, Poster
Abstract: Recent advances in gigapixel-level imaging have brought High-Resolution Wide shots to the forefront of research. However, these images present significant challenges: extreme sparsity of foreground, gigapixel-level resolutions and interleaving of foreground and background. This causes traditional detectors and attention mechanisms to be hindered by background, resulting in inefficiency and inaccuracy. To tackle this problem, we propose BiHashFormer, a dual sparsified branch transformer built on hashing. RoI Selector will predict the foreground proportion for top-k selecting target-containing windows. The compression branch processes windows selected and refines attention into HashAttention by discarding half of the key–value pairs, reducing cost and improving efficiency. The compensation branch uses HashMiner to perform low-cost hash searches in the remaining regions, recovering features of missed objects. Experiments confirm that the selection relationship in HashAttention is one-way and reveal the preference of different queries for keys. In experiments on the gigapixel benchmark PANDA, BiHashFormer reduces 37.5% backbone FLOPs while improving \textAP_50 to 81.0%. Moreover, this parallel dual-branch structure allows the model to avoid inference delays caused by inter-branch dependencies, while achieving much greater stability than single-branch models, specially with low rates of RoI selection.
PaperID: 6409, Poster
Abstract: Linear structural causal models (SCMs) are used to analyze relationships among random variables. In these models, directed edges encode direct causal effects, while bidirected edges capture latent confounding. The problem of generically identifying the hidden parameters from observed correlations remains open in causal inference. The half-trek criterion by Foygel, Draisma, and Drton [2012] and its generalization to the edge-wise criterion are important criteria for generic identification in linear SCMs, since the half-trek criterion can be efficiently decided by an algorithm with running time \mathcalO(n^5). We here develop a new criterion that is equivalent to the edgewise criterion. Utilizing this criterion, we design a faster randomized algorithm for deciding the half-trek criterion with running time \mathcalO(n^4). For sparse graphs, our algorithm even works in \mathcalO(n^3). We show that the algorithm is useful in practice by providing an implementation that outperforms the implementation of HTC in the state-of-the art R package SEMID.
Abstract: Despite their capabilities, Multimodal Large Language Models (MLLMs) may produce plausible but erroneous outputs, hindering reliable deployment. Accurate uncertainty metrics could enable escalation of unreliable queries to human experts or larger models for improved performance. However, existing uncertainty metrics have practical constraints, such as being designed only for specific modalities, reliant on external tools, or computationally expensive. We introduce UMPIRE, a training-free uncertainty quantification framework for MLLMs that works efficiently across various input and output modalities without external tools, relying only on the models' own internal modality features. UMPIRE computes the incoherence-adjusted semantic volume of sampled MLLM responses for a given task instance, effectively capturing both the global semantic diversity of samples and the local incoherence of responses based on internal model confidence. We provide theoretical analysis motivating UMPIRE's design. Extensive experiments show that UMPIRE consistently outperforms baseline metrics in error detection and uncertainty calibration across image, audio, and video-text benchmarks, including adversarial and out-of-distribution settings. We also demonstrate UMPIRE’s generalization to non-text output tasks, including image and audio generation.
Abstract: We introduce the Multiscale Single-Index Model, a stylized model for deep hierarchical feature learning with scale separation. Each layer extracts a shared single-index feature at one physical scale and passes it to the next, giving a tractable setting in which to study how deep architectures learn local multiscale representations. Under non-degeneracy and delocalization assumptions on the link function and planted features respectively, we prove two complementary recovery guarantees. First, for any fixed depth K and local scale d, corresponding to an input of size d^K, the first Wiener chaos expansion of the target behaves as a perturbed spiked tensor, where the perturbation comes from the non-linearity; we then leverage this fact to show that a spectral method based on tensor unfolding strongly recovers all planted directions with O(d^\lceil K/2\rceil\log d) samples. Second, for the two hidden-layer case, we analyze joint online spherical SGD and show that it achieves weak recovery from random initialization in O(d\log^2 d) samples, followed by strong recovery to accuracy \varepsilon in O(d\log(1/\varepsilon)) additional samples. The main technical challenge is that the second layer observes non-Gaussian learned features; we overcome this through a quantitative Gaussian comparison using delocalization, combined with sharp trajectory-level control of the coupled SGD dynamics.
PaperID: 6412, Poster
Abstract: Multimodal action understanding is commonly framed as a fusion problem: combine the available sensors and predict an action label. Synchronized video-IMU question answering (QA) exposes a new and fundamental challenge: The modality that matters more is not fixed. A question about scene context may be answered from appearance, a question about motion dynamics may depend on inertial signals, and a question under occlusion, corruption, or temporal misalignment may require deciding which modality to trust. Thus, the core problem is \emphlanguage-conditioned evidence selection under unreliable cross-modal observations, as opposed to multimodal fusion. We propose a structured evidence-to-Large Language Model (LLM) framework for video-IMU action QA. Given synchronized video and Inertial Measurement Unit (IMU), our model first extracts visual, sensor, and fused evidence using modality-specific encoders and bidirectional cross-modal attention. Each candidate answers the queries by independently querying these evidence sources, producing candidate-conditioned visual, sensor, and fused evidence summaries. Then, a reliability-aware router estimates observation reliability and selects how much each evidence source should contribute before projecting compact structured evidence tokens into a frozen LLM for multiple-choice answer prediction. We evaluate on a controlled MMAct-based action-QA protocol designed to isolate perception, sensor-centric reasoning, temporal ordering, reliability reasoning, and cross-modal complementarity. Our method achieves 61.20% accuracy on the held-out test split and 57.00% on the Cross-Modal Challenge subset, outperforming text-only, unimodal, naive fused-to-LLM, and shared-state evidence baselines. Ablations show that structured source identity, candidate identity, reliability tokens, and routing are all essential. Perturbation and shortcut analyses confirm that the model relies on paired multimodal evidence rather than language priors or answer-position bias. These results suggest that robust multimodal action reasoning requires moving beyond fixed sensor fusion toward query-specific, reliability-aware evidence organization for language models.
Authors: Weizhou Wang, Jonathan Weare, Aaron R Dinner
Abstract: Pretrained diffusion models provide powerful learned priors, but in scientific sampling the target distribution often depends on physical context that is not fully represented by one generative model. We introduce Generative Gibbs for Physics-Aware Sampling (GG-PA), a training-free framework that formulates the composition of learned partial priors and explicit physical context as inference over a joint target distribution in an augmented state space. We derive a Gibbs sampler for this joint target, show that it is asymptotically exact as the diffusion time approaches zero, and prove that in settings with quadratic interactions it remains exact at finite diffusion times. We further introduce replica exchange over diffusion time to accelerate mixing. Experiments on a double-well system, a \phi^4 lattice model, and atomistic peptide systems show that GG-PA recovers context-induced distribution shifts and emergent collective behavior in interacting systems using partial priors without retraining. These results demonstrate GG-PA as a practical approach for combining pretrained generative priors with explicit physical context.
Abstract: Operator learning for partial differential equations (PDEs) aims to learn solution operators on infinite-dimensional function spaces from finite-resolution data. In this setting, it is important for the learned model to be discretization-invariant, or resolution-robust, and to reflect PDE-specific structure. It is therefore natural to ask how such structure should be encoded in the model architecture, hypothesis class, or learning procedure. In this paper, we study operator learning for solution operators of nonlinear parabolic PDEs based on Duhamel--Picard iteration. We formulate Picard iteration as an abstract state-transition model and present a theoretical framework for Picard-type operator learning. We derive implementation-agnostic generalization error bounds that separate the implementation error from the estimation error associated with the abstract state-transition model induced by Picard iteration. A key consequence is that increasing the Picard depth reduces the Picard truncation error without causing an unbounded growth of the entropy-based estimation error. We also extend the analysis to long-time prediction by rolling out the same learned local model over successive time blocks. Finally, we illustrate the theory for nonlinear heat equations on the torus using a Picard-type Fourier neural operator as a concrete implementation.
PaperID: 6415, Poster
Abstract: Concept vectors aim to enhance model interpretability by linking internal representations with human-understandable semantics, but their practical utility is often limited by noisy and inconsistent activations. In this work, we uncover the SuperActivator Mechanism: a transformer dynamic that amplifies concept activation gaps, concentrating the most reliable concept evidence into a small set of high-activation tokens. To develop a theoretical understanding of this mechanism, we prove that concept-aligned attention heads multiplicatively amplify pairwise activation gaps, with already-extreme activations growing fastest. We find that this amplification is not just theoretical, but also occurs empirically on large-scale models: while in- and out-of-concept activation distributions overlap considerably, the in-concept distribution develops a positive tail clearly separated from the noise. These high-tail tokens, which we call SuperActivators, appear consistently across concept-positive samples, making them reliable indicators of concept presence. Accordingly, SuperActivator-based detection improves F1 by up to 14% over standard concept activation aggregators and prompting baselines across image and text modalities, models, layers, and concept extraction techniques, demonstrating the generality and practicality of our insights. Further empirical analysis demonstrates that the most reliable SuperActivators are sparse, with detection typically peaking when using only 5-10% of in-concept token activations, and capture more faithful localized semantics than global concept vectors.
Abstract: Confirmation bias, the tendency to seek evidence that supports rather than challenges one's belief, hinders one's reasoning ability. We examine whether large language models (LLMs) exhibit confirmation bias by adapting the rule-discovery study from human psychology: given a sequence of three numbers (a "triple"), an agent engages in an interactive feedback loop where it (1) proposes a new triple, (2) receives feedback on whether it satisfies the hidden rule, and (3) guesses the rule. Across eleven LLMs of multiple families and scales, we find that LLMs exhibit confirmation bias, often proposing triples to confirm their hypothesis rather than trying to falsify it. This leads to slower and less frequent discovery of the hidden rule. We further explore intervention strategies (e.g., encouraging the agent to consider counter examples) developed for humans. We find prompting LLMs with such instruction consistently decreases confirmation bias in LLMs, improving rule discovery rates from 42% to 56% on average. Lastly, we mitigate confirmation bias by distilling intervention-induced behavior into LLMs, showing promising generalization to a new task, the Blicket test. Our work shows that confirmation bias is a limitation of LLMs in hypothesis exploration, and that it can be mitigated via injecting interventions designed for humans.
Abstract: LiDAR scene flow estimation is essential for autonomous driving, as it provides 3D motion for each point. Self-supervised approaches use static-dynamic classification to mitigate the imbalance between static and dynamic points, deriving targeted supervision. However, existing methods rely on sparse geometric observations for this classification, making them vulnerable to data sparsity and occlusions. The resulting noisy labels provide incorrect motion guidance and degrade scene flow learning. To address this, we introduce TrackCue, a tracking-guided framework for improving dynamic object representation in LiDAR scene flow estimation. In particular, TrackCue repurposes point tracking to obtain dense image-space trajectories anchored to LiDAR points, providing motion cues beyond sparse geometric observations. Furthermore, we present a visually consistent motion compensation strategy that compares the tracked trajectories with ego-induced rigid trajectories in the image plane, effectively isolating true object motion from ego-induced apparent motion. To transfer these isolated motion cues back to the LiDAR domain, we perform visual motion cue lifting, which associates ego-compensated image trajectories with LiDAR points for static-dynamic label refinement. As a result, TrackCue produces more accurate static-dynamic classification and provides more reliable supervision for scene flow learning. Experimental results show that TrackCue significantly improves the precision and F1 score of dynamic labels, leading to performance gains in self-supervised scene flow estimation.
Abstract: Embodied agents performing long-horizon tasks require a memory representation in which the state transitions of dynamic objects remain queryable in natural language across hours-to-days observation horizons. Existing systems either drop fine-grained motion (clip-level video-language embeddings), keep it only as raw coordinates (geometric SLAM), or organise it around immediate task context (agent working memories). None of them gives the agent a per-object timeline whose state transitions are themselves queryable in language. Our key contribution is Linguistic Trajectory Encoding (LTE), which compresses dynamic object motion histories via a hybrid representation combining natural language descriptions, sparse spatial anchors, and visual anchors. LTE adapts compression to motion complexity by anchoring periods without reliable observations to the last seen location, while representing motion with geometric waypoints and linguistic descriptions to preserve accuracy. To evaluate these capabilities across extended time horizons, we construct the Spatial Memory Benchmark (SMB) from EgoLife multi-day recordings, targeting capabilities absent in existing benchmarks: semantic trajectory retrieval and long-horizon object retrieval. On SMB, the LTE-based system achieves 45.3% success in semantic trajectory retrieval and 48.7% in long-horizon object retrieval, outperforming structured-memory and VLM baselines (best prior: 31.9% and 34.4%). LTE achieves trajectory compression by factors of 8.7× to 26.1× with sub-second query latency on 24\,h video. On Ego4D natural-language queries, the system reaches 28.75% / 55.10% R@1/R@5, +15.80 / +31.30 pts over EgoVLPv2.
Authors:
Cheng-Kuang Chang, Kai-Wei Chang, Alexander Liu, Jim GlassAbstract: Full-duplex spoken language models (FD-SLMs) enable seamless speech interaction by allowing models to listen and speak simultaneously, yet the internal mechanism by which they coordinate listening and speaking remains underexplored. We analyze the predictive behavior encoded in FD-SLM hidden representations and find that they exhibit stream-specific predictive patterns: during listening, they preferentially predict the incoming user stream, whereas during speaking, they preferentially predict the model-side output stream. Building on this observation, we show that FD-SLMs dynamically modulate their internal predictive focus between two states: a generative state aligned with model-side output generation and a perceptive state aligned with incoming user input. However, this modulation can lag behind abrupt changes in conversational context. During user interruptions, the model remains transiently biased toward the generative state before transitioning into the perceptive state, causing it to miss the beginning of the incoming input. We term this delayed internal transition state inertia. To quantify its downstream impact, we introduce the Zero-Buffer Benchmark (ZBB), a diagnostic benchmark for evaluating immediate interruption comprehension when user speech begins abruptly. We evaluate this setting using response correctness and initial-word occurrence rate (IWOR). Finally, we mitigate state inertia through activation steering with a perception vector, a training-free intervention with little additional computational overhead. Across multiple state-of-the-art FD-SLMs, activation steering substantially improves interruption handling; for example, on PersonaPlex, it improves correctness from 28% to 45% and IWOR from 40% to 72% without any fine-tuning.
PaperID: 6420, Poster
Abstract: Accurate long-horizon probabilistic predictions of dynamical systems are a prerequisite for model-based decision-making. Variational Bayesian last layers (VBLLs) are an attractive model class for this setting, offering tractable uncertainty at near-deterministic cost, yet they are trained exclusively with single-step likelihood objectives that may be insufficient for reliable multi-step rollouts. We revisit VBLLs through the lens of Gibbs variational inference, which decouples posterior inference from the choice of training loss. Building on this perspective, we introduce a family of multi-step training losses that progressively incorporate long-horizon structure, robustness via continuous ranked probability scores (CRPS), and training on full autoregressive rollouts. Our approach retains the simplicity of VBLLs while directly optimizing for multi-step predictive quality. Across illustrative synthetic environments, a diverse suite of chaotic dynamical systems, and real-world data, we show that multi-step, CRPS-based objectives substantially improve long-horizon accuracy and distributional fit, consistently outperforming single-step likelihood training.
Abstract: Linear recurrent networks (LRNNs) and linear state space models (SSMs) promise computational and memory efficiency on sequence modeling tasks, yet their diagonal state transitions limit expressivity. Dense and/or nonlinear architectures (e.g., LSTMs) on the other hand are provably more expressive, but computationally costly. Here, we explore how expressivity in LRNNs can be increased via richer state mixing across time and channels while maintaining competitive efficiency. Specifically, we introduce two structured LRNN architectures: (i) Higher-order Linear Recurrent Units (H-LRU), which generalize recurrences to arbitrary order, mixing multiple past states, and (ii) Block-Diagonal LRUs (BD-LRU), which enable dense intra-block channel mixing. To ensure stable training, we introduce a selective gate normalization scheme that allows for scalable window and block sizes. To maintain efficiency, we utilize a parallel-scan implementation that keeps the throughput competitive with diagonal LRNNs for moderate orders (H-LRU) and block sizes (BD-LRU). Consistent with prior theoretical studies on the limitations of diagonal models, we empirically demonstrate in both synthetic sequence modeling and language modeling that our architectures significantly benefit from the increased expressivity of structured state mixing. Our results show that the structure of state mixing is a critical driver of performance in LRNNs, offering a practical pathway to closing the efficiency–expressivity gap in linear sequence models.
PaperID: 6422, Poster
Abstract: Dynamic MRI reconstruction from highly undersampled k-space is a severely ill-posed inverse problem that requires strong measurement consistency, fine spatial recovery, and coherent temporal dynamics. Diffusion priors provide useful population-level regularization, but task-matched spatiotemporal diffusion models are costly to train, while generic pretrained priors may be mismatched to protocols, vendors, or anatomies, leading to hallucination or over-smoothing. Untrained test-time priors such as Time-Dependent Deep Image Prior (TD-DIP) adapt to each scan and naturally couple with data consistency, but can drift toward aliasing artifacts and overfit noise. We propose Geometry-Gated TD-DIP for this weak-diffusion regime, where diffusion is used as a local geometric reference rather than a generative driver. Each TD-DIP update is decomposed into a diffusion-consistent tangent component and a measurement-supported orthogonal innovation component. A self-calibrated gate adaptively regulates the innovation to counter prior mismatch while suppressing noise amplification. We further provide a mismatch-aware analysis showing descent of a local posterior surrogate under explicit conditions. Experiments on cine MRI show consistent gains over diffusion-only, TD-DIP-only, and additive hybrid baselines, especially at high acceleration and under out-of-distribution diffusion priors.
Abstract: Integration against a probability distribution given its unnormalized density is a central task in Bayesian inference and other fields. We introduce new methods for approximating such expectations with a small set of weighted samples---i.e., a quadrature rule---constructed via an interacting particle system that minimizes maximum mean discrepancy (MMD) to the target distribution. These methods extend the classical mean shift algorithm, as well as recent algorithms for optimal quantization of empirical distributions, to the case of continuous distributions. Crucially, our approach creates dynamics for MMD minimization that are invariant to the unknown normalizing constant; they also admit both gradient-free and gradient-informed implementations. The resulting mean shift interacting particle systems converge quickly, capture anisotropy and multi-modality, avoid mode collapse, and scale to high dimensions. We demonstrate their performance on a wide range of benchmark sampling problems, including multi-modal mixtures, Bayesian hierarchical models, PDE-constrained inverse problems, and beyond.
PaperID: 6424, Poster
Authors:
Ziyu Cheng, Songtao Guo, Mingyan Li, Chang Han, Hongyu Xu, Xu LuoAbstract: Effective inference is critical for interactive large language model (LLM) serving, where both time-to-first-token (TTFT) and end-to-end (E2E) latency shape user experience. Speculative decoding has emerged as a promising solution that reduces inference latency by employing lightweight drafting and parallel verification. However, it introduces a new scheduling issue, i.e., traditional shortest remaining processing time (SRPT) scheduling based on output length estimation fails to work due to its paradigm shift in inference structure. In this paper, we explore the preemption mechanism of speculative decoding-based LLM serving systems and show that — besides the output length, a request's remaining service time depends on draft acceptance and the number of verification rounds. Building on this insight, we develop a novel termination-aware scheduling framework, named \textscAlmostDone, which first identifies requests that are almost done within several future local decoding steps. It achieves this by introducing a lightweight predictor built from runtime draft-and-verify states, which helps approximate the SRPT-like scheduling better than existing approaches. Then, \textscAlmostDone determines whether preemption is worthwhile through a switching-aware online policy to limit excessive running-set alterations. Experiments on an open-source vLLM serving system show \textscAlmostDone substantially surpasses the default schedulers, reducing mean and median TTFT by up to 3.83× and 3.39×, and mean and median E2E request latency by up to 1.39× and 2.33×. Against the oracle length-based preemption baseline TRAIL+, it still achieves up to 2.0× lower mean TTFT and 1.3× lower mean E2E request latency.
PaperID: 6425, Poster
Abstract: Human labels are important supervisory signals for preference-based reinforcement learning, but their collection is often constrained by practical considerations that standard training objectives tend to ignore. For instance, Direct Preference Optimization (DPO) usually treats observed preference labels as clean, even though real annotation pipelines rely on heterogeneous annotators whose reliability varies across individuals and with the difficulty of each comparison. This setting introduces three challenges: (1) Noisy labels introduce systematic bias into the preference optimization objective and need to be rigorously corrected. (2) The allocation of pairwise comparisons to different annotators is a design problem that directly affects the statistical efficiency of the training pipeline. (3) Non-random routing creates selection bias in annotator evaluation, complicating downstream decisions such as compensation. We address these challenges within a unified framework for DPO that combines annotator-aware posterior correction with information-guided annotator routing. Posterior correction leverages an instance-dependent noise model to infer latent ground-truth preferences from noisy labels, mitigating bias in the DPO objective. Annotator routing is formulated as an optimal experimental design problem that allocates comparisons to annotators according to their expected statistical information under a fixed annotation budget. To support reliable annotator evaluation under non-random routing, we introduce a doubly robust estimator that corrects for selection bias. Semi-synthetic experiments on the UltraFeedback dataset show that our approach improves ground-truth preference recovery and downstream AlpacaEval 2 performance over baselines, with consistent gains across multiple open-source LLM model architectures. Annotator reliability-estimation experiments further show that the doubly robust estimator substantially reduces mean squared error relative to other baselines, demonstrating its value for annotator evaluation.
PaperID: 6426, Poster
Abstract: The overconfidence of Large Language Model (LLM) agents poses a critical challenge to their reliable deployment for complex, high-stakes tasks. Current confidence estimation methods predominantly treat agent execution as a flat, one-dimensional, time-ordered sequence. This oversimplification fundamentally fails to capture the complex logical dependencies, branching paths, error propagation, and tool interactions inherent in agentic reasoning. To address this, we introduce DAG of Grounds (DoG), a novel method for agent confidence estimation via post-hoc trajectory restructuring. By transforming linear trajectories into Directed Acyclic Graphs (DAGs), DoG explicitly maps the grounding of an agent's final answer to its intermediate reasoning steps and tool outputs. We evaluate DoG on several benchmarks, GAIA, GPQA, and HLE, which involve long-horizon reasoning and iterative tool use, and show that DoG improves calibration performance across standard metrics such as ECE, Brier score, and AUROC. DoG outperforms existing calibration methods designed for general LLM outputs, performs comparably to or better than methods specifically designed for agent trajectories, and remains robust across different LLM backbones and tool use settings. Supported by an interactive diagnostic tool, our framework provides unprecedented interpretability for diagnosing opaque failure modes. Together, these contributions establish a structured, graph-based paradigm for confidence estimation that successfully "sniffs out" overconfidence in autonomous systems.
Abstract: Learning rate is a critical component of reinforcement learning (RL). This work uses global and local clocks to distinguish two types of learning rates. The former is of the standard form \alpha_t that depends only on the time step t (i.e., a global clock). The latter is of the form \alpha_\nu(S_t, t), where \nu(s, t) counts the number of visits to state s until time t (i.e., a local clock). In discounted RL, an RL algorithm that is convergent with a local clock is always also convergent with a global clock, and vice versa. We are not aware of any counterexample. The key contribution of this work is to show that this nice correspondence breaks down in average-reward RL. Specifically, we construct a counterexample showing that although differential temporal difference learning is convergent with a local clock, it can diverge with a global clock. This counterexample closes the open problem in Wan et al. [2021], Blaser et al. [2026].
Abstract: Language models typically reason via explicit chain-of-thought (CoT), generating intermediate steps token-by-token. Latent reasoning offers an alternative by performing multi-step inference in the model's hidden states, replacing decoded tokens with continuous representations. However, existing latent-reasoning methods often underperform explicit CoT, especially at scales beyond 1B parameters. We close this gap with LOTUS (Looped Transformers with supervision on latents), a framework that: (i) processes a padded latent prefix with a looped Transformer, computing all K latent blocks in parallel; and (ii) applies parallel supervision on latent thoughts: all K latent blocks are supervised simultaneously after the loop, each aligned to its corresponding CoT step. On Llama-3.2-3B-Instruct, LOTUS is, to our knowledge, the first latent-reasoning method to bridge the gap to explicit CoT on math reasoning, while reducing thought-phase latency by 2.5×. Ablations identify three ingredients necessary for this success: parallel supervision routed directly through the base model's LM head against the CoT tokens, sufficient parallel width per latent block, and sufficient sequential loop depth.
Abstract: In training Mixture-of-Experts models, sharded data parallelism splits each expert’s parameters across GPUs. Before each layer runs, GPUs must use Allgather to rebuild the full weight matrix. This communication can take a large part of each training iteration. Prior work reduces this cost with lossy compression, but lossy methods can affect accuracy. We propose Adaptive Exponent Reuse Allgather, or AdERA, a lossless Allgather compression method for sharded MoE training. Our key observation is that, after a short warmup phase, most weight exponents stay the same across iterations. AdERA stores these exponents locally and sends only the sign and mantissa when the exponent has not changed. The receiver then rebuilds the exact original weight by combining the cached exponent with the received sign and mantissa. Because parameter matrices have different shapes across layers, AdERA compresses only when compression saves time. It also overlaps compression with Allgather communication and computation. Our real experiments and large-scale trace-driven simulator show that, compared with the lossless baseline, AdERA achieves a 3.70× speedup on 16 GPUs and is projected to reach 4.28× on 128 GPUs, while preserving bitwise-exact parameter reconstruction. Compared with the lossy baseline, AdERA achieves a 3.93× speedup on 8 GPUs and is projected to reach 4.42× on 128 GPUs. The source code is publicly available.
PaperID: 6430, Poster
Abstract: Learning-rate selection is one of the most consequential and repeatedly tuned decisions in modern deep learning, and the ability to transfer it quickly across widths, depths, and architectures can substantially reduce the cost of scaling new models. Existing transfer laws explain width through maximal-update parameterization and depth through depth-based scaling laws, but many modern architectures are better described as neural computation graphs with heterogeneous merge rules rather than by depth alone. We develop a unified merge calculus for learning-rate scaling on such graphs. Its central object is a graph-aware complexity coefficient that combines local merge semantics with global path structure and induces a -1/2 power law for the target learning-rate scale. We further show that depth alone is generally insufficient once merge semantics vary across architecture families. We validate the framework on controlled graph families and lightweight representative architectures covering uniform sum, residual-aware sum, concat, weighted fusion, and hybrid designs. Across these settings, the resulting graph coefficients organize optimal learning rates more effectively than depth alone and support practical learning-rate transfer.
Abstract: Large language models (LLMs) have demonstrated impressive reasoning capabilities, but scaling their performance often relies on massive reasoning datasets that are computationally expensive to train on. Existing data selection methods aim to curate smaller, high-quality subsets but often rely on costly external models or opaque heuristics. In this work, we shift the focus from external heuristics to the model's internal mechanisms. We find that complex reasoning tasks consistently activate a sparse, specialized subset of attention heads, forming core reasoning circuits. Building on this insight, we propose CircuitSeer, a novel data selection method that quantifies the reasoning complexity of data by measuring its influence on these crucial circuits. Extensive experiments on 4 models and 9 datasets demonstrate CircuitSeer's superiority. Notably, fine-tuning Qwen2.5-Math-7B on just 10% of data selected by our method achieves a 1.2-point gain in average Pass@1 over training on the full dataset, highlighting its efficiency and effectiveness.
PaperID: 6432, Poster
Abstract: Text watermarking helps identify AI-generated content, but its effect on factual reliability remains underexplored. In this paper, we study watermarking hallucination: factual errors induced or amplified by watermarking even when the required evidence is present in the context and the unwatermarked model can answer correctly. Using a controlled RAG setting, we compare paired unwatermarked and watermarked generations under the same condition. Across representative watermarking methods, including KGW, SWEET, DiPmark, GumbelSoft, Gumbel-Max, and SynthID-style watermarking, we find that watermarking hallucination is widespread: watermarked outputs can remain fluent while introducing factual errors. We attribute this failure mode to two mechanisms: (1) direct token-level bias, which can suppress fact-consistent tokens, and (2) prefix-induced drift, which accumulates through autoregressive decoding and weakens later attention to factual context. Motivated by this analysis, we propose Fact-Preserving Token Intervention (FPTI) and Fact-Preserving Attention Intervention (FPAI), two plug-in interventions that can be integrated into existing watermarking methods to improve factuality. Our experiments show that combining FPTI and FPAI mitigates around 90% of watermark-induced hallucinations while preserving fluency and comparable decoding efficiency. Overall, this work highlights factuality as a first-class criterion in watermark evaluation, alongside detectability and robustness, and calls for careful factuality validation before deploying watermarks in fact-critical applications.
Abstract: Large language models demonstrate remarkable ability in factual recall, yet the fundamental limits of storing and retrieving input--output associations with neural networks remain unclear. We study these limits in a minimal setting: a linear associative memory that maps p input embeddings in \mathbbR^d to their corresponding~d-dimensional targets via a single layer, requiring each mapped input to be well separated from all other targets. Unlike in supervised classification, this strict separation induces p constraints per association and produces strong correlations between constraints that make a direct characterisation of the storage capacity difficult. Here, we first introduce a decoupled model in which each input has its own independent set of competing outputs, and provide numerical and analytical evidence that this decoupled model is equivalent to the original model in terms of storage capacity, spectra of the learnt weights, and storage mechanism. Using tools from statistical physics, we show that the decoupled model can store up to p_c \log p_c / d^2 = 1 / 2 associations, and generalise the computation of p_c to linear two-layer architectures. Our analysis also gives mechanistic insight into how the optimal solution improves on a naïve Hebbian learning rule: rather than boosting input-output alignments with broad fluctuations, the optimal solution raises the correct scores just above the extreme-value threshold set by the competing outputs. These findings give a sharp statistical-physics characterisation of factual storage in linear networks and provide a baseline for understanding the memory capacity of more realistic neural architectures.
Abstract: Representation Autoencoders (RAEs) leverage frozen vision foundation models (VFMs) as tokenizer encoders, providing robust high-level representations that facilitate fast convergence and high-quality generation in latent diffusion models. However, freezing the VFM inherently constrains its spatial reconstruction capacity, limiting fine-grained generation and image editing; in contrast, incorporating reconstruction-oriented signals via fine-tuning disrupts the pretrained semantic space and degrades generative fidelity. To address this trade-off, we propose DecQ, a simple yet effective framework for RAEs. Specifically, DecQ introduces lightweight detail-condensing queries that extract fine-grained information from intermediate VFM features through condenser modules. These queries are incorporated into the decoder to support reconstruction and are jointly generated with patch tokens during generative modeling. By aggregating information from both shallow and deep layers, DecQ effectively mitigates the reconstruction--generation trade-off, improving both reconstruction quality and generative performance. Our experiments demonstrate that: (1) with only 8 additional queries and 3.9% extra computation, DecQ improves reconstruction over the frozen DINOv2-based RAE, increasing PSNR from 19.13 dB to 22.76 dB; and (2) for generative modeling, DecQ achieves 3.3× faster convergence than RAE, attaining an FID of 1.41 without guidance and 1.05 with guidance.
PaperID: 6435, Poster
Authors: Ziheng Duan, Tong Wu, Xi Li, Simon D Sun, Chaoyang Wang, Junhao Liu, Martin Min, Jing Zhang
Abstract: Therapeutic target identification---determining which genes or proteins to modulate for a desired clinical effect---is a foundational step in drug discovery. However, existing computational methods that rank candidates at genome scale face three persistent challenges: (i) static feature design, which discards mechanistic knowledge locked in biomedical literature; (ii) lack of interpretability, offering no auditable gene-level rationale for individual predictions; and (iii) the false-negative problem, where unlabeled but potentially druggable genes are treated as definitive negatives. We propose TargetSage, an LLM reasoning agent for therapeutic target identification built on positive-unlabeled (PU) learning, with three modules. First, Agentic Profiling enriches each gene with LLM-gathered biomedical evidence, distilling 2,261 attributes into a 23-dimensional reasoning vocabulary. Second, Guided Reasoning produces gene-specific rationales and explicit reasoning attributes through a GRPO-trained policy rewarded by downstream Adjusted F_1. Third, PU-Aware Scoring integrates structured biological features, explicit reasoning attributes, and implicit reasoning embeddings via gated fusion under an LLM-informed class-prior to produce genome-wide druggability scores. Evaluated on 15 benchmark tasks spanning 19,032 human protein-coding genes, TargetSage consistently outperforms ten baselines, surpassing the strongest by 37% in macro-average Adjusted F_1 and generalizing to independent held-out benchmarks. Using only 2021 training data, TargetSage outperforms all ten baselines at ranking genes that gain clinical approval by 2025. On an independent validation set of Phase 2/3 clinical candidates, it further achieves a 51% higher enrichment odds ratio at the top-25 shortlist than the runner-up baseline. TargetSage provides an evidence-grounded, continuously updatable foundation for interpretable target prioritization at genome scale.
Abstract: When approximating an intractable density via variational inference VI the variational family is typically chosen as a simple parametric family that very likely does not contain the target. This raises the question: Under which conditions can we recover characteristics of the target despite misspecification? In this work, we extend previous theoretical results on robust VI with location-scale families under target symmetries in two substantial ways: (1) We open them up to a wider range of divergences by providing sufficient conditions for exact recovery of the target mean and correlation matrix when using the forward Kullback-Leibler divergence and \alpha-divergences. (2) By doing so, we find that we can drop the restrictive assumption of a log-concave target made in previous work, allowing us to give guarantees for a wider range of targets, including multi-modal ones. In our experiments, we show how our guarantees can serve as guidelines for the choice of the variational family and \alpha-value and we illustrate on a diverse set of examples how and why optimization can fail in the absence of our sufficient conditions.
Abstract: Long-horizon robot operation requires spatio-temporal memory to record the environment state and recall it for downstream reasoning. Scene graphs and retrieval-augmented systems ground VLM descriptions to persistent 3D entities with rich semantic descriptions. However, VLM captions are noisy and viewpoint-inconsistent, and existing systems treat them as an oracle with no mechanism to detect unreliable stored descriptions. We introduce object-level semantic uncertainty for multi-view VLM memory: a score that measures object-centric cross-view semantic scatter of captions and identifies semantically unresolved objects. Then, we include our uncertainty scores in an advanced spatial-semantic memory system, that we dub UQ-DAAAM. UQ-DAAAM uses this score to actively refine uncertain objects under a fixed query budget by selecting high-quality views and fusing the resulting multi-view captions into a single object description. We also derive probabilistic guarantees showing that higher-quality candidate views (as selected by our approach) are more likely to reduce uncertainty. Our experiments show that uncertainty quantification can make embodied 4D memory systems more reliable and more effective. In particular, on the OC-NaVQA benchmark, UQ-DAAAM achieves substantially larger uncertainty reduction and better spatio-temporal question answering performance than baselines.
PaperID: 6438, Poster
Abstract: Mixture-of-Experts (MoE) architectures are central to large-scale deep learning, relying on sparse execution and expert specialisation. Despite their success, their large-scale training dynamics remain poorly understood, and training is often unstable. While scaling limits such as infinite-width theory have clarified the behavior of dense models, the scaling behavior of MoEs remains largely unexplored. To this end, we derive the infinite-width limits of MoE architectures under fixed number of experts with both soft and Top-K routing, trained with SGD and Adam, using Tensor Programs. Our results show that under the Standard Parameterisation (SP), router updates freeze after one step of training in both soft and Top-K routing MoEs. This provides, at least in part, a principled explanation for the widespread use of auxiliary losses such as load-balancing or z-losses in MoE training. In dense networks, Maximal Update heuristics allow for stable and non-vanishing feature updates. However, in MoEs the Maximal Update heuristic with soft-routing and softmax gating causes the experts to collapse to the same distribution and router gradients to vanish. While sigmoid allows non-trivial router evolution, they still fail to prevent expert collapse. Thus fixed expert soft-routing MoEs do not admit a parameterisation that simultaneously ensures stability, feature learning, and expert specialisation in the infinite-width limit. We then show that, beyond computational efficiency, Top-\(K\) also acts as a symmetry breaking mechanism allowing for both feature learning and specialisation. We empirically validate these predictions, observing close agreement between finite-width training dynamics and the predicted scaling behaviour.
PaperID: 6439, Poster
Abstract: Diffusion feature — intermediate activations extracted from diffusion model backbones — has emerged as a promising approach for unifying generation and vision understanding (discrimination). Despite the advancements of diffusion backbones into Diffusion Transformers (DiT), the knowledge of diffusion feature largely remains in U-Net architectures. Hence, in this work, we aim to extend the study of diffusion feature onto DiT backbones. We find that simply applying the previous methodology of diffusion feature to DiT models yields unsatisfying results. We further show that long skip connections (LSCs) are a key missing factor behind this discrepancy and propose a mechanistic explanation for how they affect representation quality. Specifically, LSCs can provide shortcuts to help noises bypass the section of the backbone where the better features reside. This hypothesis is validated using both mutual information measuring for correlation evidence and a controlled study for causal evidence. Of the two, the controlled study demonstrates that LSCs work by transporting noise information rather than just any information, indicating that LSCs enable DiT backbones to adopt a more optimized behavior pattern. Thus, our findings suggest that representation quality in diffusion models is strongly influenced by how information is routed within the backbone.
PaperID: 6440, Poster
Abstract: Multimodal molecular representation learning aligns heterogeneous chemical and cellular signals to predict molecular properties. The dominant paradigm is instance-level contrastive alignment, such as InfoNCE and VICReg, which implicitly treats every modality as a point cloud in a shared Euclidean space. This paradigm overlooks two structural properties of molecular-cellular data. First, each modality, from atomic fingerprints to cell morphology to transcriptomic response, carries its own intrinsic geometry; directly pulling individual embeddings together can distort the relational structure each modality has learned to express. Second, the modalities form a directed biological cascade, from molecular structure to conformation, gene perturbation, cellular phenotype, and transcriptomic response, whose branching grows exponentially with depth, a regime that Euclidean spaces cannot embed without distortion. We therefore propose GeoPMR, a pre-training framework that addresses both properties through two complementary geometric modules: Metric-Guided Gromov-Wasserstein alignment (MGW) and a Hyperbolic Hierarchical Module (HHM). MGW equips each modality with a learnable diagonal Riemannian metric and aligns distance matrices rather than individual embeddings via entropic optimal transport, so cross-modal alignment preserves relational rather than merely pointwise structure. HHM lifts all modality latents onto the Poincaré ball and applies a probabilistic parent-child and sibling contrastive loss along the cascade, exploiting the exponential volume growth of negatively curved space to embed the hierarchy with low distortion. On seven molecular property prediction benchmarks spanning toxicity, bioactivity, and ADME, GeoPMR sets a new state of the art by improving the classification AUC by 1.4% and regression MAE by 1.7%.
Abstract: Large language model (LLM) agents have demonstrated strong capabilities in long-horizon tasks by interleaving reasoning with tool use. However, as these agents scale to complex workflows such as software engineering and open-ended research, context management becomes a fundamental bottleneck: interaction histories grow unbounded, become costly to maintain, and are difficult to reuse across sessions and agents. We introduce Git-Context-Controller (GCC), a structured context management framework inspired by software version control systems. GCC elevates agent context from a transient token stream to a persistent, navigable memory workspace with explicit operations---\textttCOMMIT, \textttBRANCH, \textttMERGE, and \textttCONTEXT, that enable milestone-based checkpointing, isolated exploration of alternative reasoning paths, and hierarchical retrieval of historical context. By organizing agent memory as a versioned file system, GCC allows agents to manage long-term goals, recover and transfer reasoning across sessions, and coordinate multi-trajectory problem solving in a principled manner. Empirically, agents equipped with GCC achieve state-of-the-art performance on both SWE-Bench and BrowseComp benchmarks. On SWE-Bench Verified, GCC improves task resolution by over 13% relative to strong long-context baselines and outperforms 26 existing open and commercial systems, reaching over 80% success rate. The project will be open-sourced for the research community.
Authors:
Jiaming Li, Chenyu Zhu, Zhiyuan Ma, Nanxi Yi, Youjun Bao, Li Sun, Quanying Lv, Xiang Fang, Daizong Liu, Jianjun Li, Kun He, Bowen ZhouAbstract: Reinforcement learning (RL) has shown extraordinary potential in aligning diffusion models to downstream tasks, yet most of them still suffer from significant reward hacking, which degrades generative diversity and quality by inducing visual mode collapse and amplifying unreliable rewards. We identify the root cause as the mode-seeking nature of these methods, which maximize expected reward without effectively constraining probability distribution over acceptable trajectories, causing concentration on a few high-reward paths. In contrast, we propose Trajectory Matching Policy Optimization (TMPO), which replaces scalar reward maximization with trajectory-level reward distribution matching. Specifically, TMPO introduces a Softmax Trajectory Balance (Softmax-TB) objective to match the policy probabilities of K trajectories to a reward-induced Boltzmann distribution. We prove that this objective inherits the mode-covering property of forward KL divergence, preserving coverage over all acceptable trajectories while optimizing reward. To further reduce multi-trajectory training time on large-scale flow-matching models, TMPO incorporates Dynamic Stochastic Tree Sampling, where trajectories share denoising prefixes and branch at dynamically scheduled steps, reducing redundant computation while improving training effectiveness. Extensive results across diverse alignment tasks such as human preference, compositional generation and text rendering show that TMPO improves generative diversity over state-of-the-art methods by 9.1%, and achieves competitive performance in all downstream and efficiency metrics, attaining the optimal trade-off between reward and diversity. Code is available at https://anonymous.4open.science/r/TMPO-2E2B.
PaperID: 6443, Poster
Abstract: Perturbation experiments are central to understanding cellular mechanisms, but remain costly and sparse, motivating prediction of gene expression responses for unobserved conditions. A promising recent direction leverages large language models (LLMs) as “virtual cell” simulators—using stepwise, knowledge-grounded mechanistic reasoning to infer differential expression—pointing toward an interpretable, knowledge-driven paradigm that transcends purely data-driven approaches. However, we find that plausibility is not prediction: despite producing biologically plausible explanations, these methods fail to capture perturbation-specific effects — systematically overestimating differential expression, often underperforming a simple gene‑frequency baseline in aggregate evaluations, and collapsing to chance-level performance at the per-gene level. This reveals a reliance on intrinsic gene response tendencies rather than true perturbation reasoning. We trace this failure to how evidence is presented: existing methods evaluate perturbation–gene pairs in isolation, without exposing how related perturbations differ in their effects on the same gene. To address this limitation, we introduce vidence), which reframes prediction as a comparison task by organizing evidence into positive and negative outcomes from related perturbations. Using a biomedical knowledge graph for evidence retrieval, CORE improves calibration and substantially boosts perturbation-specific prediction in both LLM-based and non-LLM settings: for example, on drug-perturbation data, CORE-Reasoning improves Qwen3.5-9B aggregate metrics by up to 28.6%, while on generic perturbation data, CORE-Voting raises macro-per-gene AUROC from chance to 0.711. This highlights contrastive evidence organization as essential to reliable LLM-based perturbation reasoning.
Abstract: Pass@k is a widely used performance metric for verifiable large language model tasks, including mathematical reasoning, code generation, and short-answer reasoning. It defines success if any of k independently sampled solutions passes a verifier. This multi-sample inference metric has motivated inference-aware fine-tuning methods that directly optimize pass@k. However, prior work reports a recurring trade-off: pass@k improves while pass@1 degrades under such methods. This trade-off is practically important because pass@1 often remains a hard operational constraint due to latency and cost budgets, imperfect verifier coverage, and the need for a reliable single-shot fallback. We study the origin of this trade-off and provide a theoretical characterization of when pass@k policy optimization can reduce pass@1 through gradient conflict induced by prompt interference. We show that pass@k policy gradients can conflict with pass@1 gradients because pass@k optimization implicitly reweights prompts toward low-success prompts; when these prompts are what we term \emphnegatively interfering, their upweighting can rotate the pass@k update direction away from the pass@1 direction. We illustrate our theoretical findings with large language model experiments on verifiable mathematical reasoning and coding tasks.
PaperID: 6445, Poster
Abstract: Machine unlearning for large language models optimizes two opposing losses: a forget objective drives the model away from targeted knowledge, while a retain objective holds it near its original behavior. Existing methods commit to a fixed schedule of learning rate, retain weight, and step count chosen by offline grid search, which is blind to how the trajectory actually evolves. We show that LLM unlearning trajectories share a method-agnostic three-phase structure, an early phase where the two gradients are nearly orthogonal, a conflict phase where they oppose each other and retain loss begins to grow, and a terminal phase where representational divergence crosses an irreversibility threshold. Each phase calls for a different control policy. We propose CLUE, a closed-loop framework whose architectural choices each correspond to a phase or transition: a dual-tower controller separates forget and retain signal pathways, a conflict-driven gate switches between them, and an explicit collapse-risk head supervises anticipatory stopping. The controller is trained by task-distribution meta-optimization. We provide a single-step bound on conflict-induced retain loss growth, a quantitative irreversibility result, and a stopping regret bound. On TOFU, WMDP, and MUSE with Llama-2/3 and Mistral models from 7B to 70B, CLUE achieves better forget-quality versus model-utility trade-offs than fixed-strategy baselines, stops within 5% of the oracle, and remains robust under adversarial prompting, relearning, and INT4 quantization.
PaperID: 6446, Poster
Abstract: Many machine learning methods aim to approximate the lower-dimensional manifold on which the data lives. A desirable feature of such methods is that they should capture the epistemic uncertainty of this learned manifold. One model that achieves this is the Gaussian Process Latent Variable Model, in which a Gaussian Process (GP) mapping from the latent space provides an estimate of the uncertainty of the manifold. However, the effectiveness of this uncertainty estimation is limited by the mean-field variational approximation between the GP inducing points and the latent variables. In this work, we use Amortized Structured Stochastic Variational Inference to allow the variational posterior for the latent space to be conditionally dependent on the value of the inducing points. We demonstrate that this more flexible variational posterior improves several metrics related to the reconstruction of points on the data manifold.
PaperID: 6447, Poster
Abstract: Reinforcement learning (RL) has become central to post-training large language models (LLMs). However, popular RL methods like GRPO incur non-negligible overhead by computing both old-policy and current-policy likelihoods to form importance sampling ratios. In this work, we propose Likelihood-Gated Policy Optimization (LGPO), which enforces a soft trust region constraint via likelihood-based gating, eliminating the need to compute old-policy likelihoods. Empirically, we show that removing the importance sampling correction term does not harm training stability, whereas removing the trust region mechanism leads to collapse. Moreover, ratio-based clipping can fail in fully on-policy training: the importance ratio stays at 1, so the ratio-based trust region constraint never activates. Under standard training settings where GRPO is stable, LGPO achieves comparable training stability and peak performance while reducing training time by ~18% on average. In fully on-policy training, where GRPO fails, LGPO remains stable, enabling more efficient and robust LLM RL post-training across training regimes.
PaperID: 6448, Poster
Abstract: Training AI agents powered by Large Language Models (LLMs) typically requires centralized access to user data, raising privacy and scalability concerns. We explore FedAgent, a decentralized reinforcement learning paradigm that collaboratively trains LLM agents across distributed clients without sharing local data. The central reliability question is: Is FedAgent effective under uniform client distribution, and more importantly, is it robust to client heterogeneity? For the former, we provide the first empirical evidence that FedAgent matches Centralized Agent Training and outperforms Local Agent Training. For the latter, we first formalize Agent Heterogeneity at two structurally distinct levels: task-level (what clients ask the agent to do) and environment-level (the dynamics in which the agent acts), anchored on the Input-Dynamics Asymmetry of task-augmented MDPs, referring to the architectural fact that tasks enter the policy through its input channel, while environments do not. Then, we theoretically establish an Asymmetric Robustness Mechanism: FedAgent is robust to task-level heterogeneity but non-robust to environment-level heterogeneity. We further identify three sufficient conditions under which FedAgent recovers robustness despite environment-level heterogeneity, and illustrate four possible training-curve patterns. On real-world agent benchmarks WebShop and ALFWorld, we empirically verify that FedAgent remains robust under extreme task-level heterogeneities and traces a stable-degrade-collapse spectrum under environment-level heterogeneities.
PaperID: 6449, Poster
Abstract: Graph Neural Networks (GNNs) face a fundamental dichotomy between expressiveness and scalability. Recent spectral and transformer-based models approach universal approximation capabilities but rely on computationally intensive global operators that scale poorly to large graphs. Conversely, scalable approaches often compromise structural fidelity through sampling or decoupling, resulting in incomplete structural acquisition. We identify that this trade-off stems from a homogeneous treatment of graph topology, where uniform computational complexity is applied to heterogeneous structures. We propose a paradigm shift towards Hierarchical Structural Acquisition and introduce S&P (Scalable & Powerful), a framework grounded in Fusion Frame Theory that aligns the operator with the intrinsic hierarchy of the data. S&P decomposes the global operator into intra-component and inter-component, and we prove that this design preserves universal spectral filtering capacity for non-degenerate inputs while remaining numerically stable and linear in complexity. Experiments on 14 datasets show that S&P achieves state-of-the-art accuracy, bridging the gap between scalability and expressiveness. Our code is available at https://anonymous.4open.science/r/HierarchicalSP.
PaperID: 6450, Poster
Authors: Mostofa Rafid Uddin, Seonghui Min, Mahek B Vora, Qifeng Wu, Muyuan Chen, Min Xu
Abstract: Emerging microscopic technologies such as cryo-electron tomography (cryo-ET) provide direct 3D visualization of macromolecules within the cell, enabling analysis of their in situ morphology. This morphology can be regarded as an SE(3)-invariant, denoised volumetric representation of subvolumes extracted from tomograms, termed as subtomograms. Morphology identification from a set of subtomograms is formulated as an inverse problem of estimating a set of template morphologies and per-subtomogram SE(3) transformations with respect to any one of the templates. The existing expectation-maximization-based solution to this end often struggles with high structural heterogeneity and requires manual selection of a large number of hyperparameters. Addressing this issue, we present a novel deep unsupervised learning framework called DUIL. Given a set of subtomograms, DUIL first models their SE(3)-invariant morphological code using a siamese-like neural network with a multi-choice learning module. The learned morphological codes are clustered and used to generate a set of template morphologies through a generator network. The generated templates are used as references to estimate SE(3) transformations for each subtomograms through latent optimization. The subtomograms with identical template morphologies are then aligned and averaged to iteratively refine the templates. Experiments on simulated and real cryo-ET datasets demonstrate clear improvements over prior methods, including the discovery of previously unidentified macromolecular morphologies.
PaperID: 6451, Poster
Abstract: LLM agents do not present a fixed inference job to an edge and cloud system. At each step, a placement decision chooses where to generate the next message, where to execute any resulting tool call, which prefix must be replayed, which cache state is updated, and whether the episode continues. As a result, the workload to be scheduled is partly produced by the scheduling decisions themselves. We formalize this problem as an (RAOP), a step level orchestration policy that first places language generation and then, after observing any tool call, places tool execution. RAOP predicts candidate action workloads from prefix replay, prefill, KV cache residency, queue, network, and tool payload features, and is trained with GRPO to optimize measured full trajectory success and latency rather than a one step latency surrogate. To make controlled evaluation auditable, the main simulator replays language and tool continuations produced by real model and tool execution while recomputing physical timing under the sampled placement sequence. As an analytical lens, we derive a Bellman regret decomposition showing that myopic placement policies incur error from workload prediction, success priors, and a controlled endogeneity term. On GSM8K, HotpotQA, WebShop, and a calibrated three tier deployment, RAOP matches the strong cloud only success reference while reducing P99 latency by about 53% and communication by about 73%. Against the strongest learned step level baseline, RAOP improves average success by 2.5 points, indicating that full trajectory optimization is central for agent orchestration.
Abstract: Decomposition-based Programming-by-example (PBE) scales performance by splitting tasks into subtasks that a learned synthesizer solves: a decomposer predicts intermediate subgoals, and a synthesizer generates programs conditioned on them. Current approaches train the decomposer to imitate ground-truth (GT) subgoals, implicitly treating decomposition quality as intrinsic to the task. We challenge this assumption: for bounded solvers with fixed inductive biases, GT decompositions reflect the annotator’s factorization choices – not the solver’s search dynamics. A decomposer trained to match GT decompositions may therefore propose subgoals that are logically valid yet intractable for the solver. We propose Solver-Aware Decomposition (SAD), a training framework that retains supervised training on GT subgoals as a structural scaffold, while additionally optimizing the decomposer via direct feedback from a frozen synthesizer. Subgoals are rewarded based on the synthesizer’s loss on the target program – a signal of subtask difficulty that encourages decompositions the solver can act on. Our experiments reveal an accuracy paradox: higher agreement with GT decompositions does not improve synthesis success – even though the synthesizer was trained on the very same GT data the decomposer is optimized to mimic. SAD instead learns decompositions that trade GT alignment for solver tractability, yielding consistent gains in synthesis and end-to-end task accuracy across two PBE domains. Moreover, SAD solves tasks that a GT decomposition oracle fails – empirical evidence that GT decompositions are not universally optimal for bounded solvers, and that decomposition quality is solver-relative, not intrinsic.
Abstract: While Gaussian Splatting-based Feature Fields (GSFFs) have shown promise for visual localization, this paper highlights that photometrically optimized GSFFs are inherently ill-suited for 2D-3D matching. The volumetric extent of each Gaussian induces many-to-one pixel-to-point mappings that destabilize PnP-based pose estimation, while photometric optimization gives rise to superfluous Gaussians devoid of multi-view consistency. To address these issues, we propose SplitGS-Loc, a localization-specialized GSFFs construction framework that disambiguates 2D-3D correspondences by exploiting Gaussian attributes. Our key design, Mixture-of-Gaussians-based splitting, decomposes each Gaussian into smaller Gaussians, replacing ambiguous many-to-one with precise one-to-one correspondences. In parallel, we exploit composition weights from GS rasterization to select Gaussians that significantly and consistently contribute across multiple views and aggregate discriminative features through strong pixel-Gaussian associations, enforcing multi-view consistency. The resulting compact yet discriminative feature fields enable stable PnP convergence, achieving state-of-the-art performance on localization benchmarks. Extensive experiments validate that SplitGS-Loc extends the utility of photometric GSFFs to accurate and efficient localization by exploiting Gaussian attributes, without per-scene training or iterative pose refinement.
PaperID: 6454, Poster
Authors:
Sujit Roy, Johannes Schmude, Ata A Asanjan, Thorsten Kurth, Rohit Lal, Kshitiz Mandal, Vishal Gaur, Harris Abdul Majid, Nikolaos Dionelis, Berkay Aydin, Himanshu Patil, Andres Munoz-Jaramillo, Campbell Watson, Juan Moreno, Manil Maskey, Rahul RamachandranAbstract: Context length remains a fundamental bottleneck for vision foundation models operating on high-resolution imagery. Current architectures rarely exceed 1M tokens, forcing practitioners to downsample inputs at the cost of fine-grained spatial information. This limitation is particularly acute in heliophysics, where satellite instruments record full-disk solar observations at native 4096 × 4096 resolution across 13 channels, and downsampling discards the small-scale magnetic structures and localized dynamics critical to understanding solar phenomena. In this work, we present the first vision foundation model for heliophysics trained at native 4K resolution on approximately 15 years of multimodal solar data (257 TB) spanning eight extreme ultraviolet channels and five magnetic field and velocity products from the Solar Dynamics Observatory (SDO). We adapt MultiMAE with a dual-view formulation: 25% of tokens are observed directly, while the remaining 75% are replaced with fixed Gaussian random Fourier projections that act as structured, non-invertible frequency views of the masked content. Using two-way parallelism (feature + sequence) with FP8 mixed precision, we scale the context window to >3 million tokens at 8 × 8 patch size, a 3× advance over prior work. We demonstrate strong scaling efficiency up to 1024 GPUs across five hardware configurations, with an architecture capable of processing sequences exceeding 10M tokens. The learned representations achieve strong zero-shot reconstruction across masking ratios up to 90% and full missing-modality scenarios. On downstream tasks (AR segmentation and EVE irradiance prediction) with a frozen encoder, we outperform SOTA by 9.5% and 35.3% respectively.
PaperID: 6455, Poster
Abstract: Machine unlearning aims to remove targeted knowledge from a trained model while preserving its general capabilities. Doing so effectively for auto-regressive language models requires identifying which tokens in a forget sample are truly relevant to forgetting, as uniformly applying unlearning across all tokens can degrade utility. Existing approaches, however, either ignore this heterogeneity or rely on auxiliary models, hand-crafted heuristics, or external annotations to approximate token-level relevance. We instead characterize token-level forget relevance through its interaction with the retain objective: tokens are forget-relevant to the extent that minimizing the forget loss on them does not conflict with the retain objective. Based on this insight, we introduce Alternating Token-Weighted Unlearning (ATWU), a framework that jointly learns token relevance and model parameters during the unlearning process. ATWU uses a lightweight linear scorer over model hidden states to predict forget-relevance scores with negligible computational overhead and no external supervision. Experiments on TOFU and RWKU show that ATWU achieves state-of-the-art forget--retain trade-offs, outperforming sample-level methods, probability-based token-weighting heuristics, and auxiliary-model-based approaches. Moreover, the learned scorer aligns with ground-truth forget-relevant spans substantially better than existing methods, suggesting that ATWU learns semantically meaningful token-level forgetting signals. Overall, ATWU shows that token-level forget-relevance can be effectively inferred from model representations during unlearning, enabling efficient and selective forgetting.
PaperID: 6456, Poster
Abstract: Learning semantically meaningful human movement representations from 3D pose sequences is essential for human behavior modeling. Existing self-supervised methods derive supervision from artificial assumptions, e.g., invariance to augmentation within a certain range or recoverability of masked regions at a certain temporal scale, introducing unintuitive hyperparameters that make the learned semantics sensitive to their settings. We propose BAL, a self-supervised framework that instead exploits supervision intrinsic to human movement: the natural ordering structure of everyday behavior. BAL autoregressively predicts in a semantically abstract latent space, analogous to how humans anticipate others' actions rather than their precise joint configurations. Bidirectional autoregressive prediction, both forward and backward in time, fully leverages these ordering constraints while enabling representation extraction using both past and future context at inference. To suppress representation collapse, BAL decomposes each prediction target into a global component capturing sequence-level semantics and a local component capturing context-independent movement. We prove that the local targets retain nontrivial angular diversity unless the representations are completely collapsed, preserving a meaningful training signal. Experiments on large-scale datasets demonstrate that BAL's representations are highly effective for motion captioning, forecasting, and interpolation, outperforming existing self-supervised methods across all tasks.
PaperID: 6457, Poster
Abstract: We study 3D visual grounding in 3D Gaussian Splatting (3DGS), where a referring expression should identify a target object as both a rendered 2D mask and a set of explicit 3D Gaussians. Prior methods typically extract 3D targets by applying heuristic thresholding or top-ratio filtering to dense primitive-level relevance scores. However, these scores often vary substantially across scenes and prompts, making such heuristics sensitive and leading to unstable targeting. ClusterSplat addresses this by reframing 3D target selection as cluster-level selection over scene-adaptive candidates. The method first learns instance-aware Gaussian features, forms feature-aware seed voxels, and merges neighboring seed voxels when a description-length criterion decreases, producing a scene-specific set of cluster candidates. Given a referring expression, a query-conditioned cluster scorer ranks the cluster candidates with a lightweight MLP, and the selected cluster directly defines the explicit 3D target while its cluster score is rendered for the 2D mask. On ScanRefer, ClusterSplat achieves the best rendered 2D segmentation and explicit 3D target metrics among the compared 3DGS baselines. It also obtains the best average Ref-LERF scores, supporting the same trend under fine-grained referring descriptions. On ScanNet v2 class name queries, ClusterSplat achieves the strongest rendered 2D results and the highest 3D matching accuracy, while maintaining competitive 3D IoU.
PaperID: 6458, Poster
Abstract: Monocular 3D detection (Mono3D) hinges on recovering depth from a single image, for which the geometric formulation z \approx f \cdot H / h_2D is widely adopted due to its simplicity and direct grounding in projective geometry. We show that this formulation harbors a systematic, projection-induced bias: substituting the 2D bounding-box height h_2D recovers the depth of an object's visible front surface rather than its center, causing size-dependent underestimation that persists even under ground-truth inputs. We address this front-surface bias with MonoPCA, a unified framework built on a projection-consistent reformulation. At its core, PCDepth retains the measurable h_2D and closes the gap to center depth through an orientation- and size-aware correction derived from perspective geometry, eliminating the bias without introducing ill-posed regression targets. Because PCDepth couples depth to predicted object size and orientation, attribute perception becomes essential. We therefore introduce GeoAlign, a teacher--student framework that transfers geometric representations from Vision Foundation Models into a lightweight student, strengthening attribute perception at no inference cost. Across four challenging benchmarks, MonoPCA establishes new SoTA performance, including a +1.44% AP3D gain on the KITTI leaderboard and +2.80% on Omni3D Out.
PaperID: 6459, Poster
Authors: Kevin Slote, Jeremie Fish
Abstract: Symbolic library --- or Koopman dictionary --- selection is a fundamental challenge in data-driven dynamical systems. Extended Dynamic Mode Decomposition (EDMD), Sparse Identification of Nonlinear Dynamics (SINDy), and Kolmogorov--Arnold Networks for Dynamics (KANDy) all require the practitioner to commit to a function library at training time; Deep-Koopman Operators avoid this commitment but produce uninterpretable latent observables. We propose Deep-Koopman-KANDy, a structured approach to post-hoc symbolic dictionary readout that combines Deep-Koopman modeling with Kolmogorov-Arnold Networks for Dynamics (KANDy). The encoder and decoder of a Deep-Koopman Operator are replaced with two-layer Kolmogorov--Arnold Networks (KANs), and a level-set construction together with a chain-rule gradient identity exposes the compositional structure of the learned observables in a basis chosen \emphafter training. We evaluate the method on the Lorenz system, the Chirikov standard map, the Ikeda map, and the Arnold cat map. On Lorenz it recovers the target dictionary \x,y,z,xy,xz\ with perfect recall and Jaccard score 0.79\pm0.06; on the standard map it recovers a low-order Fourier basis matching the analytical structure; on Ikeda---which has no sparse polynomial representation---a misspecified polynomial readout still recovers the correct foliation coordinate g\approx x^2+y^2 together with a nontrivial outer function; and on the Arnold cat map --- used as a negative control because finite-dimensional Koopman closure is provably impossible --- the method fails to find a sparse closure, as expected.
PaperID: 6460, Poster
Abstract: A kernel k(x,y) is LSHable if there exists a locality sensitive hashing scheme \mathcal H such that k(x,y)=\Pr_h~\mathcal H[h(x)=h(y)] for all x,y. This notion plays a key role in efficient kernel methods in high dimensions. In this work, we show that the p-exponential kernel k(x,y)=\exp(-\lVert x-y \rVert_p) is LSHable in bounded regions for all 1
Authors: Szu-Wei Fu, Rong Chao, Xuesong Yang, Sung-Feng Huang, Ryandhimas E Zezario, Rauf Nasretdinov, Ante Jukić, Yu Tsao, Frank Wang
Abstract: Universal Speech Enhancement (USE) aims to restore speech quality under diverse degradation conditions while preserving signal fidelity. Despite recent progress, key challenges in training target selection, the distortion--perception tradeoff, and data curation remain unresolved. In this work, we systematically address these three overlooked problems. First, we revisit the conventional practice of using early-reflected speech as the dereverberation target and show that it can degrade perceptual quality and downstream ASR performance. We instead demonstrate that time-shifted anechoic clean speech provides a superior learning target. Second, guided by the distortion--perception tradeoff theory, we propose a simple two-stage framework that achieves minimal distortion under a given level of perceptual quality. Third, we analyze the trade-off between training data scale and quality for USE, revealing that training on large uncurated corpora imposes a performance ceiling, as models struggle to remove subtle artifacts. Our method achieves state-of-the-art performance on the URGENT 2025 non-blind test set and exhibits strong language-agnostic generalization, making it effective for improving TTS training data. Code and models will be released upon acceptance. Audio samples are available at https://anonymous.4open.science/r/USE-5232/index.md.
PaperID: 6462, Poster
Abstract: While hierarchical structures are intrinsic to organizing human knowledge, current representation learning on text-attributed graphs predominantly operates on flat semantic spaces, overlooking the rich coarse-to-fine granularity inherent in real-world data. To bridge this gap, we propose SHiFT, which synergizes the reasoning power of LLMs with the structural representation capability of GNNs within an Expectation-Maximization (EM) paradigm. In the E-step, we harness LLMs as expert taxonomists to induce an explicit semantic tree, introducing a lightweight Mapping & Diagnosis strategy that reduces inference costs by surgically updating the hierarchy only when anomalies are detected. In the M-step, we align the GNN encoder with this induced hierarchy from complementary perspectives, including clustering partition, topological skeleton, and semantic concepts. Theoretical analysis interprets SHiFT as a generalized Variational EM algorithm. Extensive experiments demonstrate that SHiFT not only achieves state-of-the-art performance but also uncovers high-quality, human-readable taxonomies.
PaperID: 6463, Poster
Authors:
Zhiming Lin, Tianxiang Xu, zizhao zhang, Yixue Liu, Yumeng Ma, Yihao Zhong, ShaoYan He, CHEN RONGRU, Siyue LiuAbstract: GUI agents are increasingly expected to complete real tasks through language-based interaction, yet most supervised training still treats a single demonstrated click or action as the unique ground truth. This creates a mismatch between how agents are trained and how they are evaluated: task success depends on reaching the intended interface state, not on reproducing the demonstrator's exact surface action. We address this gap by introducing Transition-Equivalent Supervision (TES), a framework that trains GUI agents around goal-relevant state changes rather than raw action identity. TES converts demonstrated interactions into transition descriptors, supervises agents to predict the intended transition, and learns an executor that selects any verified action capable of realizing that transition. At inference time, the agent further verifies whether the executed action produces the intended state change, reducing repeated no-op or erroneous behavior. Across offline web prediction, executable web navigation, mobile control, and desktop-use benchmarks, TES consistently improves equivalent-step success, alternative-action recall, and online task completion while reducing NoOp/Error transitions. These results suggest that GUI-agent learning should move beyond click imitation toward transition-level reasoning, offering a more robust supervision principle for agents operating in diverse and dynamic interfaces.
Abstract: As large language models (LLMs) are increasingly integrated into high-stakes decision-making, the ability to reliably quantify uncertainty has become a critical requirement for safety and trust. However, current uncertainty quantification methods primarily operate at the output level, often failing to distinguish whether uncertainty arises from the model’s lack of knowledge or from ambiguity in the user’s input. While input-centric uncertainty quantification has recently emerged as a promising direction, it remains relatively underexplored and typically relies on coarse, input-level information. Consequently, users are provided with scalar uncertainty scores that offer little actionable guidance on which parts of the input should be clarified to improve reliability. To address this limitation, we propose Shapley-based input uncertainty Quantification (ShaQ), a framework for span-level attribution of input-induced uncertainty. Our approach models ambiguous spans in the input as players in a cooperative game and quantifies their contributions using Shapley values, defined via the weighted average of marginal reductions in conditional entropy obtained by clarifying each span coalition. Unlike existing input-level approaches, our formulation explicitly captures complex interactions among spans and provides a principled decomposition in which individual attributions sum exactly to the total input-induced uncertainty. We evaluate ShaQ on the AmbigQA and AmbiEnt benchmarks, where it achieves state-of-the-art performance in ambiguity detection. We further demonstrate its practical utility on a MediTOD benchmark, showing that ShaQ can precisely localize under-specified clinical utterances, facilitating more effective human-AI collaboration in high-stakes settings. Overall, our results show that ShaQ not only improves uncertainty estimation but also provides actionable insights for targeted input clarification. Codes are available https://anonymous.4open.science/r/ShaQ-0E39/README.md.
PaperID: 6465, Poster
Authors:
Lurui Wang, yongkang cheng, Can XuAbstract: Expressive performers resonate with audiences because they transform sound into motion: rhythm invites steps, prosody evokes gestures, and energy shapes the flow of the body. However, existing audio-driven character systems often reduce audio to a task-specific condition, separately optimizing for co-speech gestures or music-driven dance. This fragmented view overlooks audio as the shared perceptual driver of performance, making it difficult to generate coherent full-body motion across speaking, dancing, and their intermediate states. We present VOXPERFORMER, a unified audio-to-motion framework that moves toward a generalized performance foundation model. Our key insight is that audio is not merely an external condition, but the hidden soul of performance style, governing how motion emerges, evolves, and transitions. To realize this, VOXPERFORMER introduces a structured motion prior to organize the full-body performance space, audio-to-prior distillation to transform sound into latent motion cues, a memory-augmented prior to resolve audio-motion ambiguity, and a decoupled diffusion generator with implicit transition matching for efficient and coherent synthesis. Experiments on BEAT2 and FineDance show that VOXPERFORMER remains competitive on speech-driven gesture generation while substantially improving hand-aware music-driven dance generation. User studies further show that participants prefer VOXPERFORMER, indicating stronger audio-motion correspondence and more natural perceptual resonance.
PaperID: 6466, Poster
Abstract: Segmenting unseen anatomical structures in echocardiographic images is challenging because dense annotations are scarce and ultrasound often presents visually similar textures across different cardiac structures. Current foundation models remain limited in this setting because they treat anatomical knowledge as independent text prompts, overlooking the spatial topology that clinicians use to localize ambiguous structures from their visible neighbors. We propose Echo-SAM, a zero-shot ultrasound segmentation framework that grounds a structured cardiac knowledge graph in visual space through relation-conditioned geometric support inference. Echo-SAM localizes seen structures as anatomical anchors, propagates relational constraints through a knowledge-grounded Graph Neural Network, and enhances text prompts with image-grounded anatomical context. We further introduce a topology-to-geometry mapping that converts graph relations into geometry-aware support proposal scores, yielding posterior-guided support estimates for unseen structures. By treating relation-specific scale, thickness, and offset as latent geometric variables, Echo-SAM adapts graph-derived support proposals to image-specific anchors without relying on fixed coordinate templates. Extensive evaluations on zero-shot echocardiography segmentation benchmarks demonstrate that Echo-SAM substantially outperforms state-of-the-art foundation models, achieving up to 50.42% absolute improvement in Dice score. Code will be publicly available.
Abstract: We propose a test-time defense mechanism against adversarial attacks: Unlike existing methods that rely on feature filtering or smoothing, which can lead to information loss, we propose to ``combat noise with noise'' by leveraging stochastic resonance to enhance robustness while minimizing information loss. Our approach introduces small translational perturbations to the input image, aligns the transformed feature embeddings, and aggregates them before mapping back to the original reference image. This can be expressed in a closed-form formula, which can be deployed on diverse existing network architectures without introducing additional network modules or fine-tuning for specific attack types. The resulting method is entirely training-free, architecture-agnostic, and attack-agnostic. Empirically, the method achieves state-of-the-art robustness on image classification and provides the first generic test-time defense for dense prediction tasks, including stereo matching, optical flow, and monocular depth estimation. Across adversarial attacks, it recovers up to 68.1% of the accuracy loss on image classification, 71.9% on stereo matching (MAE), 29.2% on optical flow (EPE), 68.3% on depth estimation (SqRel), and 83.8% on vision-language alignment (cosine similarity).
PaperID: 6468, Poster
Abstract: The public belief state (PBS), which is the posterior over game histories conditioned on public information, is a fundamental abstraction for designing game-theoretically sound search algorithms for imperfect-information games. Existing sound PBS search techniques for adversarial games fall into two categories: gadget-game-based algorithms (including Libratus, DeepStack, Pluribus, Supremus, and Student of Games) and ReBeL-based algorithms. Both these categories possess disadvantages, including a reliance on discontinuous functions, an inability to re-solve subgames, and an implicit representation of policies. In this work, we propose Gravel, a new approach for PBS search---different from these previous two categories---that achieves game-theoretic soundness via regularization. Unlike the aforementioned approaches, Gravel relies on smooth functions, can initiate search at any point in the game, and outputs explicit policies directly. We empirically demonstrate the advantages of Gravel for no-limit Texas hold'em, where we show that it outperforms ReBeL both in head-to-head competitions against Slumbot and in endgame solving. We include our codebase in the submission, making Gravel the first open-source high-performance no-limit Texas hold'em AI.
Authors: Maximilian Graf, Victor Thuot
Abstract: We study multiple change point localization under bandit feedback. An unknown piecewise-constant function on a compact interval can be queried sequentially at adaptively chosen inputs, and each query returns a noisy evaluation of the function. The goal is to identify a prescribed number of discontinuities, known as change points, within a target precision \eta and confidence level 1-\delta, while using as few samples as possible. We propose an adaptive algorithm that first detects intervals likely to contain change points and then refines their locations to precision \eta. We establish non-asymptotic upper bounds on its sample budget, together with corresponding lower bounds. Prior work shows that jump magnitudes alone determine the asymptotic sample complexity as \delta\to 0. We reveal that this picture is incomplete beyond this regime. We demonstrate, both empirically and theoretically, that for general \delta and \eta, the complexity is jointly governed by the jumps and the relative positions of the change points.
PaperID: 6470, Poster
Abstract: Optimal pre-training data mixtures are vital for Large Language Models (LLMs), yet post-RL reasoning performance varies significantly across base models. This suggests that identifying "RL-friendly" pre-training data mixtures is a critical but under-explored prerequisite. This motivates us to ask: Which specific data mixtures in pre-training enhance post-RL effectiveness, and how can we evaluate a base model's suitability for RL fine-tuning? Through large-scale experiments with proxy models, we observe that (1) domain-specific data correlation strongly predicts post-RL performance, and (2) the average perplexity of the base model on k RL-generated responses, \textavg(k\text-ppl), is a more stable and correlated predictor than traditional metrics. Building on these insights, we propose BridgeRL, a framework that treats finding optimal mixing ratios as a regression task, using \textavg(k\text-ppl) as the fitting objective. We train 1M/60M-parameter proxy models for regression and scale the findings to 1B/3B-parameter models. Results across online and offline RL algorithms demonstrate that BridgeRL-optimized mixtures consistently yield superior base models for RL fine-tuning.
PaperID: 6471, Poster
Abstract: Chain-of-thought (CoT) monitoring is increasingly relied upon to detect misbehavior in LLMs, under the implicit assumption that misbehaving actors must reason about their misbehavior, leaving detectable traces in the CoT. We show this assumption fails for a class of attacks we call planning-execution decoupled attacks, in which misbehavior is injected via a corrupted plan before the actor begins reasoning. Using the investigator agent framework of Li et al. [2025], we discover that exposing an actor to a corrupted reasoning plan upstream causes it to reproduce the flawed reasoning in its own CoT, embedding misbehavior into natural-looking reasoning with few suspicious traces. The attack scales to stronger reasoning models and harder tasks, steering them toward misbehavior while evading detection. Evaluating a suite of thinking and non-thinking monitors, we find that thinking monitors substantially outperform non-thinking ones, though even the best miss a meaningful fraction of attacks. Critically, the relationship between monitor thinking budget and detection is not monotonic: extra reasoning sometimes improves detection but can also hurt it when monitors talk themselves into accepting corrupted CoTs as benign. This refines Guan et al. [2025], who show thinking budget generally helps monitoring; we find detection also depends on whether the additional reasoning is directed toward critical evaluation rather than rationalization.
PaperID: 6472, Poster
Authors: Cuong Ngoc Nguyen, Lam Ho, Vu Dinh, Georgios Karagiannis, Cuong V. Nguyen
Abstract: Bayesian neural networks (BNNs) offer principled uncertainty quantification, yet sampling from their posteriors via Hamiltonian Monte Carlo (HMC) remains challenging. While recent work theoretically identified the efficiency degradation of ReLU-based networks, smooth unbounded activations ( GELU, Swish, Mish), which are prevalent in modern architectures, were assumed to be efficient due to their smoothness. In this work, we challenge this existing justification by deriving explicit local error formulas for the leapfrog integrator. Our analyses reveal that, in addition to differentiability, third-order curvatures of the potential energy also have a significant influence on HMC efficiency of BNNs inference. Empirically, we found that even smooth unbounded activations may degrade sampling performance as severely as piecewise linear activations. Thus, optimally tuning the step size for these networks may require a comparable empirical step size scaling as their piecewise counterparts. This result sheds light on the implications for architectural design in Bayesian deep learning.
PaperID: 6473, Poster
Abstract: For two d-dimensional point sets A,B of size n, the Chamfer distance from A to B is defined as \textCH(A,B)=\sum_a \in A \min_b \in B \mathsfdist(a, b), where \mathsfdist is some underlying distance measure. We present efficient algorithms for approximating the Chamfer distance under the \ell_p norm for 1 < p \le 2. For the general case 1 < p < 2, we utilize lopsided \ell_p to \ell_1 embeddings with weak guarantees. We show that they are sufficient for preserving the Chamfer distance. In the high-dimensional regime, we use Fast Matrix Multiplication techniques to speed up the lopsided embeddings. These lead to \mathcalO(nd\cdot \mathrmpoly(\log \log n, \log(1/\varepsilon))/\varepsilon^2) total time. For the specific Euclidean case (p=2), we leverage the Fast Johnson-Lindenstrauss Transform based on Toeplitz matrices and re-analyze the previous \ell_1-specific Chamfer algorithm in \ell_2. These achieve a runtime matching the recent state-of-the-art \ell_1 result of \mathcalO(nd(\log \log n + \log(1/\varepsilon))/\varepsilon^2), improving upon the previous \mathcalO(nd \log n/\varepsilon^2) bound for \ell_2. We also give additional results for the angular distance and upper and lower bounds in the streaming setting.
Authors: Zhongjie Shi, Wenjing Liao
Abstract: This paper investigates the learning theory of Transformer networks for regression tasks on the compact Euclidean domain [0,1]^d and d-dimensional compact Riemannian manifolds. We propose a novel constructive approximation framework for Transformers that builds local approximations of the target function and aggregates them into a global approximation via softmax partition of unity. This approach leverages the attention mechanism to achieve spatial localization through affine transformations of the input. The softmax activation plays a crucial role in aggregating local approximations to a global output. From an approximation perspective, we prove that a dense Transformer equipped with only two encoder blocks and standard single-hidden-layer point-wise feed-forward networks can achieve a uniform \varepsilon-approximation error for \alpha-H\"older continuous functions with \alpha \in (0,1] using \mathcalO(\varepsilon^-d/\alpha) total parameters. Building upon this approximation guarantee, we establish a near minimax-optimal generalization error bound of order \mathcalO\big(n^-\frac2\alpha2\alpha+d \log n\big) for the empirical risk minimizer, where n is the training data size. The Transformer architecture studied in this paper is dense, shallow and wide, and employs softmax activation and sinusoidal positional encodings, closely reflecting practical implementations.
Abstract: In recent years, agentic workflows have been widely applied to solve complex human tasks. However, existing workflow construction still faces key challenges, including human-dependent workflow construction, the lack of graph-level execution feedback, and the inability to repair errors in-loop during long-horizon construction. To address these challenges, we propose FlowSteer, a new paradigm of Agent Designing Agentic Workflows — a single agent itself end-to-end designs the workflow that a downstream executor runs. To support this paradigm, we introduce the Workflow Canvas, a novel executable graph-state environment that returns syntax-checked execution feedback for every atomic edit. Built on the canvas, we further propose Reinforced Progressive Canvas Editing, in which a lightweight policy agent issues one atomic edit per turn conditioned on real canvas feedback, and is trained end-to-end via reinforcement learning. Moreover, FlowSteer provides a plug-and-play framework that supports diverse operator libraries and interchangeable LLM backends. Experimental results on twelve datasets show that FlowSteer significantly outperforms baselines across various tasks. Our code is available at https://anonymous.4open.science/r/FlowSteer-9B2E.
Abstract: Transformers are effective at inferring the latent task from context via two inference modes: recognizing a task seen during training, and adapting to a novel one. Recent interpretability studies have identified from middle-layer representations task-specific directions, or , that steer model behavior. However, a lack of rigorous foundations hinders connecting internal representations to external model behavior: existing work fails to explain how task-vector geometry is shaped by the training distribution, and what geometry enables out-of-distribution (OOD) generalization. In this paper, we study these questions in a controlled synthetic setting by training small transformers from scratch on latent-task sequence distributions, which allows a principled mathematical characterization. We show that two inference modes can coexist within a single model. In-distribution behavior is governed by Bayesian task retrieval, implemented internally through convex combinations of learned task vectors. OOD behavior, by contrast, arises through extrapolative task learning, whose representations occupy a subspace nearly orthogonal to the task-vector subspace. Taken together, our results suggest that task-vector geometry, training distributions, and generalization behaviors are closely related.
PaperID: 6477, Poster
Authors: Ray Telikani, Jaber valizadeh, Amir H Gandomi, Bao Q Vo, Ming Ding
Abstract: Neural contextual bandits support high-stakes decision systems---from clinical trial allocation to content recommendation. This paper investigates an adversarial threat, group-targeted suppression attacks (GTSA), in which an adversary selectively degrades performance for a demographic subgroup while preserving global performance metrics. We first show that under standard aggregate monitoring, GTSA with budget B = \Omega(\sqrtT) can induce \Omega(T) group regret while maintaining only o(1) deviation in observable global statistics, establishing that any defense that ignores group structure is fundamentally vulnerable. To address this vulnerability, we propose AnchorFair, a robust mitigation framework comprising three components: (i) a trust-weighted policy anchor with provable bounded drift, (ii) representation-stable demographic discovery via online clustering in a slowly evolving embedding space, and (iii) a demographic-aware exploration mechanism that adaptively amplifies learning for underrepresented or low-trust groups. We prove that \textscAnchorFair achieves sublinear group regret \tildeO(\sqrtT) for all groups under GTSA.
PaperID: 6478, Poster
Abstract: Domain Incremental Learning (DIL) aims to continuously adapt to new domains while retaining the knowledge of previous domains. Existing DIL methods are generally built upon an idealized assumption of clean supervision. However, real-world data streams are often affected by label noise, giving rise to the more challenging Noisy Domain Incremental Learning (N-DIL) scenario. Under this setting, not only is intra-domain knowledge acquisition hindered, but inter-domain knowledge conflicts are also exacerbated, which amplifies catastrophic forgetting. To address these challenges, we propose a novel Cross-Domain Knowledge Separation and Positive Transmission (ST-Prompt) framework. Specifically, to mitigate inter-domain knowledge conflicts, a Cross-domain Prompt Knowledge Contrastive Isolator is developed to enhance domain-wise knowledge separation, thereby mitigating the cross-domain knowledge interference during inference. Furthermore, to improve intra-domain knowledge acquisition, a Synergistic Knowledge Transmission scheme is introduced, which extracts reliable knowledge from previous domains to facilitate noisy data learning in the current domain. These two components are mutually reinforcing, jointly promoting effective cross-domain knowledge learning and utilization. Extensive experiments on diverse benchmarks demonstrate that ST-Prompt achieves state-of-the-art performance. Our code will be released.
PaperID: 6479, Poster
Abstract: Subject-driven video generation aims to produce videos that faithfully incorporate user-provided reference subjects. Current evaluation relies on free generation followed by reference-similarity scoring, which rewards a fundamental shortcut: models can obtain high scores by reproducing visible cues from the references while avoiding transformations that would test whether the subject is truly preserved beyond the conditioning input. We design anchored evaluation: controlled settings in which generation is anchored to held-out targets of the same subject rather than left unconstrained. Each probe withholds target content from the model's conditioning, so success requires using the references to infer content not directly visible, rather than simply copy-pasting what was given. We instantiate two probes. (1) Viewpoint control specifies target geometry via depth-based warping, leaving occluded regions that the model must complete using subject appearance from the references. (2) Degradation--recovery corrupts held-out videos showing the subject in real-world contexts that differ from the reference images, then recovers them conditioned on those references; recovery requires cross-context generalization of identity and inference of state dynamics. Both probes require only lightweight modifications to the inference loop of flow-matching models, need no retraining, and generalize across architectures. Evaluation across seven diverse, strong models reveals that rankings shift substantially: VACE-14B ranks first conventionally but Phantom-14B, ranked last, leads under anchored evaluation. Moreover, conventional scores strongly correlate with the improvement from providing ground-truth references (Pearson r = 0.91), suggesting they largely measure copying ability. Fine-grained analysis further reveals distinct failure modes: copy-paste models drift at unseen viewpoints, and no model adaptively balances copying and generalization.
Authors: Michał Sadowski, Tadija Radusinović, Maria Wyrzykowska, Lukasz Sztukiewicz, Jan Rzymkowski, Piotr Kozakowski, Mikolaj Sacha, Ruard van Workum, Stanislaw Jastrzebski
Abstract: Retrosynthesis is one of the domains transformed by the rise of generative models, and it is one where the problem of nonsensical or erroneous outputs (hallucinations) is particularly insidious: reliable assessment of synthetic plans is time-consuming, with automatic methods lacking. In this work, we present RetroTrim, a retrosynthesis system that successfully avoids nonsensical plans on a set of challenging drug-like targets. Compared to common baselines in the field, our system is not only the sole method that succeeds in filtering out hallucinated reactions, but it also results in the highest number of high-quality paths overall. The key insight behind RetroTrim is the combination of diverse reaction scoring strategies, based on machine learning models and existing chemical databases. We show that our scoring strategies capture different classes of hallucinations by analyzing them on a dataset of labeled retrosynthetic intermediates. To measure the performance of retrosynthesis systems, we propose a novel evaluation protocol for reactions and synthetic paths based on a structured review by expert chemists. Using this protocol, we compare systems on a set of 32 novel targets, curated to reflect recent trends in drug structures. While the insights behind our methodology are broadly applicable to retrosynthesis, our focus is on targets in the drug-like domain. By releasing our benchmark targets and the details of our evaluation protocol, we hope to inspire further research into reliable retrosynthesis.
PaperID: 6481, Poster
Abstract: Autoregressive language models generate under a fixed left-to-right order, whereas diffusion language models (dLLMs) denoise masked tokens through flexible any-order trajectories that may commit multiple positions per step. This makes dLLMs a distinctive setting for test-time scaling: additional inference compute can diversify both token choices and the order in which positions are committed. Existing samplers do not explicitly allocate these two sources of diversity. Confidence-based decoding is reliable but often redundant, while global temperature sampling and random remasking inject stochasticity without distinguishing useful exploration from unreliable commitments. We propose SIGMA, a training-free sampler that formulates each denoising step as a constrained allocation problem over decoding efficiency, token-level diversity, trajectory-level diversity, and model-internal commitment uncertainty. The resulting objective decomposes joint sampling entropy into token and trajectory terms, yielding a sigmoid-form stochastic gate for position selection and an adaptive per-position temperature for token sampling. Experiments on LLaDA and Dream across math and code benchmarks show that SIGMA improves the accuracy--cost Pareto frontier and can be integrated into existing dLLM test-time scaling pipelines.
PaperID: 6482, Poster
Abstract: Cross-modal scene and instance retrieval aims to retrieve a corresponding 3D environment or object given query 2D images, serving as a foundational mechanism for spatial reasoning in robotics and AR/VR. Existing methods achieve this by compressing the entire 3D map into a single global feature vector, which discards the fine-grained local geometrical structure in indoor scenes. In this work, we show that scene- and instance-level retrieval capabilities emerge from supervision on sparse, local 2D-3D correspondences, without any explicit global labels. Specifically, we propose CrossWeave, a novel framework for cross-modal retrieval that learns sparse patch-point alignment, which alone outperforms global encoding methods. Additionally, an optimal transport-based aggregator is applied over the aligned features to attain further gain by clustering structurally redundant regions, freeing descriptor capacity for semantically distinct local features. CrossWeave achieves state-of-the-art performance on both scene- and instance-level retrieval on ScanNet and 3RScan with a single model, whereas existing methods require separate models for each task. Our results suggest that whole-scene encoding is not necessary for cross-modal retrieval, and that sparse local correspondence supervision is a more effective and flexible alternative.
PaperID: 6483, Poster
Abstract: How do protein structure prediction models fold proteins? We investigate this question through causal interventions on the folding trunks of ESMFold, OpenFold, and Boltz-1. Across all three models, we find a shared two-stage computational structure. In the first stage, early blocks initialize : features like charge propagate from sequence into pairwise representations through architecture-specific pathways. In the second stage, late blocks develop : distance and contact information accumulate in the pairwise representation. We verify these mechanisms causally by showing that steering charge and distance features induces predictable structural changes. Furthermore, these representations are functionally interchangeable: pairwise states can be linearly aligned and substituted across models. Together, these results suggest that folding trunks with different architectures, inputs, and training procedures converge on a shared representational organization for mapping sequence chemistry into spatial geometry.
PaperID: 6484, Poster
Abstract: Spiking Neural Networks (SNNs) offer high energy efficiency, but direct optimization is challenging because spike generation is governed by a non-differentiable hard threshold. Practical direct-training methods, including surrogate gradients (SG) and local zeroth-order (LocalZO) estimators, maintain hard spikes in the forward pass while using surrogate or perturbation-induced derivatives in the backward pass. This forward--backward mismatch raises a central theoretical question: how do local gradient discrepancies propagate through the spatiotemporal computation graph, and can the resulting biased updates still provide guaranties for the original discrete-spike objective? We develop a unified convergence analysis framework for SG and LocalZO direct SNN training. A smoothed spike objective is introduced only as an analytical bridge, allowing us to separate the core error into two components: the gradient-level mismatch between the direct-training mean direction and the smoothed-objective gradient, and the objective-level consistency gap between the smoothed and original discrete objectives. Both components are controlled by a threshold-tail functional that measures how often hard-trajectory membrane potentials lie near the firing threshold. Under suitable regularity assumptions, we obtain a non-asymptotic guarantee for the original discrete-spike objective. Experiments across representative SNN architectures and benchmarks verify convergence consistency, the existence of errors, and demonstrate that the threshold-tail statistic is an effective diagnostic for direct SNN training.
PaperID: 6485, Poster
Abstract: Large Language Model (LLM) agents are increasingly deployed under structured generation, where grammar constraints enforce output format at inference time. Whereas existing compression methods target unconstrained generation, we first identify a fundamental mismatch: under structured generation, grammar constraints reduce the model's output distribution to valid tokens, and the logits matter most at the small set of positions that decide the agent's action. Motivated by this finding, we propose SKIM (Selective Knowledge-Informed Masking), the first pruning and distillation framework for LLM agents under structured generation, with both stages guided by token role. SKIM partitions token positions into context, decision, and format categories that reflect the input the model reads, the constrained options it chooses among, and the format realized by the constraint mechanism. Specifically, this partition drives two complementary axes: (1) category-aware pruning saliency scores that preserves parameters most relevant for decision positions; and (2) an on-policy distillation objective that aligns student and teacher on valid tokens, paired with category-aware feature matching. We evaluate SKIM on tool-calling benchmarks and within multi-step agent loops, across multiple model scales and sparsity levels. SKIM consistently improves accuracy over baselines, with its superiority becoming most striking in high-sparsity regimes where conventional methods struggle, demonstrating the potential of category-aware compression for deployment-time grammar-constrained agents.
PaperID: 6486, Poster
Abstract: Human visual similarity judgments are context-dependent. For example, two images may be similar in shape but distinct in color. Existing perceptual similarity metrics, however, collapse these nuances into a single scalar value, offering no mechanism to condition on specific aspects. To bridge this gap, we introduce a large-scale dataset of human similarity judgments over image triplets, where each triplet is annotated across multiple, free-form semantic aspects. Benchmarking a broad range of frontier vision-language models (VLMs) reveals a considerable performance gap compared to human-human agreement. Leveraging our data, we fine-tune a VLM to produce our Text-Prompted Image Perceptual Similarity (TPIPS) metric, capturing multiple senses of visual similarity depending on the specified text prompt. We demonstrate that TPIPS aligns more closely with human perception and generalizes reliably beyond the training distribution. Finally, we show that TPIPS unlocks new capabilities in text-guided retrieval, compositional search, and the fine-grained evaluation of generative models. Our code, data, and trained models will be released.
PaperID: 6487, Poster
Authors:
Julien Canitrot-Paradis, H. M AFSAR, Jean-Yves Pierron, Chokri MraidhaAbstract: The quadratic assignment problem (QAP) is widely considered one of the most difficult NP-hard combinatorial optimization problems, and one that remains challenging for neural constructive methods. We introduce CIMP, a Constructive policy trained by Imitation from Multi-expert teachers and built around an explicit Pairwise pre-conditioning block. CIMP reaches a 11.02% mean optimality gap on the 133 QAPLIB instances with published reference values, across 5 fixed same-config seeds, in a single constructive pass — without local search or per-instance refinement — with inference times between 0.1 and 3.1 seconds per instance. The model is trained on 400 synthetic instances with n ≤ 25. The most recent published neural reference reporting full QAPLIB results is SAWT [Tan and Mu, ICML 2024], a learn-to-improve (L2I) method that iteratively refines an assignment per instance after training on 5,120 synthetic instances, and reports a 26.8% mean gap. CIMP operates in a strictly more constrained learn-to-construct (L2C) regime and from a much smaller training budget, yet attains a substantially lower mean gap. The pre-block builds a state-dependent facility–location representation before scoring, a hybrid context-biased decoder mixes learned content similarity with QAP-specific algorithmic bias terms, and a sequence-distribution distillation objective is derived from a diversity-filtered pool of heuristic teachers. Family-wise and size-conditioned analyses, supported by a secondary normalized score, indicate that the gain is broadly consistent across QAPLIB families and across instance sizes beyond the training range.
PaperID: 6488, Poster
Authors: BitChan Eom, Eunchan Kim
Abstract: Early-exit networks promise inference savings by terminating computation at intermediate classifiers, but FLOPs-based gains rarely translate into on-device latency reduction under static-graph compilation. On NVIDIA Jetson Orin Nano, the per-sample exit policy of Shallow-Deep Networks (SDN) runs approximately 2× slower than vanilla ResNet-56 and misses the 60 fps deadline despite using only 40% of the FLOPs, because dynamic branching precludes layer fusion. We propose the Selective Lightweight IC Ensemble Network (SLIENet), a train-full, deploy-partial framework compiled as a single static FP16 engine. SLIENet trains the full backbone with light self-distillation, using the final classifier as a stop-gradient teacher to preserve error diversity across internal classifiers (ICs). At deployment, a sub-second calibration search selects an IC subset for softmax averaging, while overthinking late blocks are truncated at depth k. On CIFAR-100 across ResNet-56, VGG-16, and MobileNetV1, SLIENet outperforms SDN-based variants and remains competitive with ZTW. On Jetson Orin Nano, the recommended SLIE5 configuration improves accuracy by +2.46 pp while reducing p50 latency by 14% and energy by 15% over vanilla, with zero deadline misses across 10,000 test samples. Thus, multi-exit deployment should be a pre-compiled deterministic inference plan, not a sample-wise routing policy.
PaperID: 6489, Poster
Abstract: The electronic Schrödinger equation encodes a molecule's electronic structure and quantum properties, but solving it scales exponentially with the number of electrons. Neural-network variational Monte Carlo (VMC) has reached chemical accuracy on the ground state of small molecules; extending this success to excited states is constrained by two existing paradigms. Penalty-based methods are computationally scalable (O(K) Laplacian count per step) but require tuning the overlap-penalty schedule against the energy gap they aim to compute. Natural-excited-states VMC (NES-VMC) is principled and penalty-free but couples all K states through a joint \det\Psi ansatz that forces O(K^2) Laplacians per step. We resolve this trade-off via a low-rank approximation (LoRA) reparameterization of the rank-K variational objective. The reparameterized objective admits an unbiased gradient estimator via MCMC sampling from the per-state Born density, and a sequential nesting scheme recovers the bottom-K eigenfunctions in eigenvalue order without overlap penalties or post-training diagonalization. The resulting algorithm, NestedLoRA-VMC, is penalty-free, achieves O(K) Laplacian count per step, and inherits a global optimality guarantee from the LoRA principle. On the first-row atoms (Li through Ne) at K=10, it stays within the chemical-accuracy threshold of NES-VMC on every atom and outperforms the penalty-based baseline of Szabó et al. on seven of the eight.
PaperID: 6490, Poster
Abstract: Graph Domain Adaptation (GDA) has emerged as an effective technique for transferring knowledge from a label-rich source graph to an unlabeled target graph by mitigating cross-graph discrepancies. However, in real-world scenarios, the attributes of partial nodes may be unobserved (i.e., missing) due to factors such as privacy protection. This inevitably exacerbates the distribution discrepancy of nodes between graph domains, making it more challenging to transfer effective knowledge to the target graph. To tackle this issue, we propose an adaptive attribute completion approach with representation space for incomplete graph domain adaptation, namely AC-GDA. It mutually assists the attribute completion and dynamically optimizes cross-graph trustworthy embeddings in an iterative manner, promoting positive knowledge migration while completing attributes. Specifically, we first propose a joint completion mechanism guided by topology similarity, which leverages cross-graph trustworthy attribute embeddings with imputation weights to adaptively complete missing node attributes. We then enhance node embeddings with label information to provide supervision for completed nodes, making nodes of the same category distribute consistently across graph domains. Furthermore, we introduce the complement confidence score to dynamically adjust the scope of trustworthy embeddings, so as to improve the quality of subsequent cross-graph imputation. Extensive experiments on a range of cross-graph tasks validate the effectiveness of AC-GDA.
PaperID: 6491, Poster
Authors: Yongbo Zhang, Xinzhe Li, Katsuma Inoue, Mitsumasa Nakajima, Toshikazu Hashimoto, Yasuo Kuniyoshi, Kohei Nakajima
Abstract: Spiking neural networks (SNNs) are a promising paradigm for efficient neuromorphic computing, yet their training remains challenging due to the non-differentiability nature of spike firing. Backpropagation through time (BPTT) with surrogate gradients (SG) has become a mainstream approach due to its superior performance. However, this method incurs high computational and memory demands during training, while its update rule depends on specific neuronal dynamics. Furthermore, when targeting neuromorphic hardware implementation, the temporal error-propagation process and the full membrane-potential access required by SG are difficult to realize; these can introduce additional readout overhead, measurement noise, or perturbations to the system dynamics. To address these issues, we propose the input-driven derivative-free training (IDDFT) method. Rather than relying on membrane-potential-based surrogate derivatives, IDDFT constructs derivative-free error signals by applying a relaxed nonlinearity to the inputs, thereby avoiding the temporal error-propagation process and reducing the dependence of the update rule on specific neuronal dynamics. By eliminating access to internal membrane-potential states, our method reduces computational and memory costs, enhances compatibility with neuromorphic hardware, and enables training of black-box SNN models. We validate the IDDFT method through theoretical analysis and systematic experiments, demonstrating its effectiveness as an input-driven, derivative-free mechanism for constructing update signals. When integrated with a biologically inspired training strategy, IDDFT maintains comparable performance while substantially enhancing robustness to the choice of relaxed nonlinearities. Based on the resulting bio-inspired combined framework, we conduct evaluations on CIFAR-10, CIFAR-100, CIFAR10-DVS, and Tiny-ImageNet. Experimental results demonstrate that our method achieves performance comparable to BPTT and state-of-the-art methods, while exhibiting remarkable robustness under adversarial attacks.
PaperID: 6492, Poster
Abstract: Traditional backdoor attacks inject artificial trigger patterns into training data, causing models to misclassify triggered inputs while maintaining normal behavior on clean samples. Since these artificial triggers are absent in clean datasets, standard knowledge distillation (KD) typically eliminates them, leading to the common belief that KD serves as an effective purification defense. In this paper, we propose DAR (Discover-and-Relabel), a support-persistent backdoor method that retains its behavior even after clean-data distillation. DAR identifies naturally abundant patterns in low-dimensional feature subspaces using subspace clustering and applies label-only poisoning to samples that inherently satisfy the discovered rule. At test time, the backdoor is activated by locally modifying the corresponding subspace of the target images. By instantiating spatial and frequency operators, DAR achieves attack accuracy comparable to state-of-the-art backdoor attacks in ImageNet classification and CLIP-based prompt tuning. Notably, DAR keeps attack success above 96% after clean KD and remains defense-resistant in practice.
PaperID: 6493, Poster
Abstract: Text-guided image editing with diffusion models struggles with spatial edits (e.g. manipulating objects, changing viewpoints, manipulating occlusions) that depend on the 3D structure understanding of a scene. Marigold and follow-up work demonstrate that a pretrained diffusion backbone already harbors strong geometric priors that can be activated for depth estimation, yet how to redirect these priors toward editing remains unresolved. Through per-layer gradient probing, we discover that the editing and depth objectives are compatible in early transformer layers but actively conflict in late layers, explaining why naive depth-conditioned editing either ignores geometry or degrades visual quality. We present StitchEdit, which first activates the backbone's geometric capability via a lightweight depth-denoising adapter, then selectively fine-tunes this adapter for editing using a probe-guided per-layer learning rate schedule that preserves depth knowledge where it helps and prunes it where it conflicts. On the proposed StitchEditBench, a benchmark of 150 geometry-dependent edits spanning spatial transformation, occlusion manipulation, and viewpoint change, StitchEdit consistently surpasses both instruction-based and depth-conditioned baselines on perceptual quality and geometric fidelity.
PaperID: 6494, Poster
Abstract: Training-free sparse attention offers a practical acceleration solution to Diffusion Transformers (DiTs) via reducing computations without fine-tuning. It typically involves estimating the importance of query-key regions and deriving sparse masks to compute only the important candidates, which inevitably introduces approximation errors that may degrade generation quality. To better balance the efficiency-quality trade-off, we propose iCATS, integrating improved importance estimation and sparse mask construction with an efficient hardware execution strategy. Specifically, for importance estimation, unlike previous works that perform independent clustering over query and key tokens based on feature similarity to estimate attention scores, iCATS demonstrates that clustering based on query-key dot-product interactions is more accurate and further reformulates this objective as a simple quadratic form for low-cost computation. For sparse mask construction, instead of using a fixed top-p rule, we observe that tolerance to sparse approximation errors varies across denoising timesteps and therefore introduce an SNR-guided sparsity schedule to adjust sparsity dynamically, leading to higher accuracy. Finally, for hardware execution, we devise a tail-merging strategy to reduce padding overhead caused by irregular cluster sizes, improving GPU kernel utilization. Extensive experiments show that iCATS achieves 2.03× acceleration with 31.017 dB PSNR on HunyuanVideo-T2V-13B and 1.55× acceleration with 29.301 dB PSNR on Wan2.1-T2V-14B, delivering a state-of-the-art efficiency-quality trade-off.
PaperID: 6495, Poster
Abstract: Light-field (LF) spatial super-resolution hinges on exploiting cross-view correspondences to recover high-frequency details while preserving geometric structure. However, existing methods often rely on implicit feature interaction or coarse disparity maps, which suffers noticeable performance degradation in regions with large disparities. In this paper, we propose Fine-Grained Epipolar Attention Diffusion (FEAD), which adopts epipolar attention as an explicit mechanism to characterize cross-view geometric correspondence along epipolar lines. To enable more accurate feature alignment and long-range spatial–angular interaction, we introduce a diffusion-based epipolar attention refinement strategy that progressively improves the initial attention. With the refined attention as guidance, cross-view features are explicitly warped and fused to enforce geometric consistency and aggregate complementary details across views for fine-grained reconstruction. Extensive experiments demonstrate that FEAD achieves state-of-the-art performance, with particularly strong gains in large-disparity scenarios.
PaperID: 6496, Poster
Abstract: Multi-turn conversation is the predominant form of interaction with large language models (LLMs), where user queries continuously evolve. However, existing LLM routing methods are primarily designed for single-turn interactions or fixed queries settings, overlooking the dynamic nature of multi-turn conversations and the challenge of delayed rewards, thereby limiting their ability to optimize cumulative performance. To address this challenge, we move from myopic, single-turn selection to long-horizon routing for multi-turn conversation. Accordingly, we propose ConvoRouter, which first performs MCTS to explore conversation branches induced by different LLM selections and collect trajectories with high cumulative rewards. ConvoRouter then learns a lightweight routing policy from search-derived data, augmented with retrieval-based future state approximation, enabling multi-turn routing without online search. Experiments on both open-domain and domain-specific conversation tasks across diverse candidate sets of both open-source and closed-source LLMs demonstrate that ConvoRouter significantly outperforms single LLMs and existing routing baselines in task success rate, while achieving a superior performance-cost trade-off when combined with a cost-aware reward.
PaperID: 6497, Poster
Abstract: A learned latent space supports reliable scientific decisions only if decisions made in that space remain valid in the original physical variables. We study latent Kalman-type data assimilation, where an encoder--decoder pair is trained without knowledge of the sensor used to collect observations. Can representation-level diagnostics predict which sensors give accurate physical-state filtering? In general, no. We show that any such diagnostic (a smooth functional of the averaged decoder Fisher matrix, including optimal-design criteria, posterior Cram\'er--Rao surrogates, and effective-observability scores) sees the sensor only through a \emphFisher--Gram coarsening of the physical sensor Gramian, leaving a large blind spot. The failure is unconditional, an algebraic consequence of averaging the decoder Jacobian over the latent invariant measure. The positive side is conditional: when the proportionality between physical and latent filtering error is stable across sensors, latent filtering error ranks sensors accurately, with an explicit Pearson bound. Decoder smoothness provides one route to this stability, and a linear-Gaussian Riccati surrogate (built entirely from the trained latent model and decoder) empirically tracks the same ranking when full-state truth is unavailable. Experiments on five chaotic systems support both sides. Representation geometry is useful but structurally limited, and reliable sensor ranking in latent data assimilation requires filter-aware validation rather than decoder-diagnostic scoring.
PaperID: 6498, Poster
Authors:
Yijie Zhu, Rui Shao, Jie He, Wei Li, Bo Zhao, Yelin Wang, Xiaochen Yuan, Tao Tan, Miao Zhang, Xiaojiang Peng, Zitong YUAbstract: Predictive Vision–Language–Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, existing approaches often fail to realize this potential and underperform direct action prediction models. We argue that these limitations stem from modality misalignment between observations and actions, together with joint optimization conflicts that drive learning away from an action-centric objective. To this end, we introduce It aligns predictive observation and action representations by mapping both modalities into a shared discrete latent space via a unified codebook, making predictive observation latents readily usable for action generation and mitigating modality misalignment. Building upon this, it then injects predictive observation latents into action decoding as explicit predictive priors via a , enabling adaptive predictive guidance under a single action-centric objective. Extensive experiments on both simulation and real-world robotic tasks demonstrate that ATI-VLA achieves state-of-the-art performance with faster convergence.
PaperID: 6499, Poster
Abstract: RL often plateaus early: policies sharpen quickly, and further training yields little gain despite remaining model- and data-capacity. Entropy is recently used to combat this plateau, but its promotion alone is not a sufficient intervention target in the settings we study. By analyzing RL's dynamics through the optimal policy, we instead identify a sharpness factor \kappa as a more effective alternative. We then introduce a plug-in reward-budget mechanism, which allocates a fixed total reward budget to each prompt and distributes it among correct rollouts. Theoretically, this induces a more regulated \kappa-trajectory that deviates from many RL methods, and its updates are equivalent to optimizing a concave objective, approximately \mathbbE(\log R). Empirically, our method achieves significant gains on Qwen3-8B and Llama3.2-3B-Instruct (up to +12.50 on AMC23 and +10.00 on AIME24 \& AIME25). Together, these results suggest that regulating algorithm-induced sharpening is more effective for mitigating premature convergence than directly preserving entropy.
PaperID: 6500, Poster
Abstract: Vision-language models such as CLIP enable zero-shot classification from arbitrary class names, but under test-time corruptions, their visual encoder projects distorted features onto incorrect textual anchors with high confidence—a phenomenon we term modality bias. Standard entropy minimization reinforces these overconfident errors, triggering a destructive confirmation-bias loop that collapses the predicted label distribution. We propose CDIM (Cross-Modal Information Maximization), an online test-time adaptation method that breaks this loop at three levels. At the representation level, a frozen self-supervised DINO backbone is combined with CLIP via a cross-covariance logit ensemble; by projecting text prototypes into DINO's geometric space—rather than the reverse—the classifier operates free of text bias. At the optimization level, a gated information-maximization objective (GT-IM) pairs per-sample sharpening with batch-level diversity, while an adaptive gate excludes unreliable samples from the gradient. At the prediction level, an online logit adjustment corrects the residual class-distribution skew. Ultimately, CDIM validates its broad effectiveness by delivering robust, competitive performance across demanding setups, including continual TTA, mixed-domain shifts, and diverse domain generalization benchmarks.
PaperID: 6501, Poster
Authors: Yanru Zhou, Dandan Song
Abstract: Conditional agreement is non-identifying under endogenous support selection: in variable-support structured prediction, a predictor can score higher on conditional agreement while scoring lower on recall, predicted cardinality, and support coverage. This matters because agreement-style proxies are widely used as reliability evidence in counterfactual consistency, schema-constrained extraction, self-consistency, and verifier-style filtering. We formalize this risk for thresholded set-valued prediction with predictor-dependent evaluability events, and prove a strict witness theorem: a support-shrinking predictor can achieve larger conditional agreement than a coverage-preserving predictor while strictly worsening all support-aware quantities; an indifference-fiber corollary shows that even equal conditional agreement need not identify these quantities. We then use Re-DocRED as a controlled high-risk stress test, not as the full scope of the claim. Moving from No\_CF to Uniform-0.5 raises J_\mathrmcond from 0.692 to 0.877, while Silence increases by 68%, predicted cardinality drops by 20%, and recall drops by 16 points. Threshold replay, matched-cardinality controls, gold-positive margin audits, and support-state decomposition show that the gain is realized through support shrinkage rather than benign threshold choice. Boundary tasks delimit scope, and a small LLM probe is used only as reference-tier motivation. The practical recommendation is a diagnostic witness set rather than an aggregate score: report conditional agreement together with support-aware quantities.
Abstract: Understanding how deep neural networks learn representations remains a central challenge in machine learning theory. In this work, we propose a feature-centric framework for analyzing neural network training by relating weight updates to feature evolution. We introduce a simple identity, the , which identifies the weight Gram matrix as the key object capturing feature dynamics. This enables us to interpret gradient descent as implicitly inducing a hypothetical evolution of features, whose covariance structure — termed the — characterizes how representations evolve during training. Building on this perspective, we introduce , a measure quantifying the linear alignment between features and targets. By analyzing the training and layer-wise dynamics, we show that deep networks learn to transform representations toward target-linear structure. This linearization perspective provides a unified interpretation of several empirical phenomena, including Neural Collapse and linear interpolation in generative models.
PaperID: 6503, Poster
Authors:
Xueting Yang, Zhong Li, changjun jiang, Wenbin Zhang, Meikang QiuAbstract: Money laundering detection on financial transaction graphs is critical but challenging. Laundering activities often conceal illicit fund flows within dense legitimate transaction neighborhoods and disperse suspicious behaviors over extended time spans. These patterns cause weak illicit signals to be spatially diluted and temporally forgotten, limiting the effectiveness of standard graph learning methods in AML scenarios. In this paper, we propose a spatio-temporal dual denoising framework for money laundering detection on evolving graphs, named EvoDen. For spatial denoising, instead of relying on passive neighborhood aggregation, we design a risk-potential guided walk mechanism to extract denoised and suspicion-biased transaction sequences, enabling the model to isolate anomalous fund flows from noisy local neighborhoods. For temporal denoising, we propose a continual Transformer-based sequence learning method driven by a reconstruction objective, which learns denoised sequence representations while enabling latent-space replay of historical knowledge to mitigate representation drift. We conduct experiments on three transaction datasets, and the results show that EvoDen outperforms state-of-the-art baselines in accurately identifying concealed money laundering entities.
PaperID: 6504, Poster
Abstract: Tabular data synthesis aims to generate high-quality data while preserving privacy. However, we find that existing tabular generative models exhibit a clear tradeoff in the small-data regime: improving data quality typically comes at the cost of increased memorization of training samples, thereby weakening privacy protection. This tradeoff arises because small training sets make it difficult for dataset-specific generative models to distinguish generalizable structure from sample-specific patterns. To address this, we propose DiffICL, which formulates tabular data generation as an in-context learning problem. Instead of fitting each dataset from scratch, DiffICL leverages pretrained structural priors learned from a large collection of datasets, enabling it to infer data distributions from limited context rather than memorizing individual samples. We evaluate DiffICL on 14 real-world datasets. Results show that DiffICL improves both data quality and privacy, and generates synthetic data that provides effective data augmentation. Our findings suggest that the quality–privacy tradeoff can be improved through better training paradigms.
PaperID: 6505, Poster
Abstract: Volumetric radar nowcasting asks a model to forecast where storm structures move and how reflectivity evolves through height. The recent \stcgs method makes this setting practical by tracking 3D radar volumes with spatiotemporally coherent Gaussian parcels, but its forecasting head is designed to evolve all future Gaussian attributes with a unified deterministic regressor. This is a poor match to convective evolution: geometry is dominated by coherent transport, while strong-echo radiometry is local, intermittent, and uncertain. We present GeoRad-3D, a factorized forecasting head for coherent Gaussian radar states. The key idea is simple: move geometry first, predict the stable radiometric continuation next, model only the residual uncertainty, and form the final nowcast after rendering. For geometry, a height-conditioned shared field captures volume-level advection and mild coherent deformation, while a parcel-wise correction restores local parcel drift, shape change, and cross-height motion. For radiometry, a deterministic base predictor outputs the stable continuation, and conditional flow matching samples residuals around this base to model the local uncertainty concentrated around strong echoes. At inference, multiple aligned futures are rendered and aggregated through a base-anchored quantile summary in radar-product space, where threshold skill is measured. Under the published NEXRAD protocol, GeoRad-3D improves Pool4 CSI over the strongest published baseline by about 45%, 82%, and 124% at 20, 30, and 40 dBZ. It further reduces LPIPS-Radar by 31.0%, indicating better perceptual fidelity of rendered storm structures.
Abstract: The prevailing Direct Forecasting (DF) paradigm dominates Long-term Time Series Forecasting (LTSF) by forcing models to predict the entire future horizon in a single forward pass. While efficient, this rigid coupling of output and evaluation horizons necessitates computationally prohibitive re-training for every target horizon. In this work, we uncover a counter-intuitive optimization anomaly: models trained on short horizons—when coupled with our proposed Evolutionary Forecasting (EF) paradigm—significantly outperform those trained directly on long horizons. We attribute this success to the mitigation of a fundamental optimization pathology inherent in DF, where conflicting gradients from distant futures cripple the learning of local dynamics. We establish EF as a unified generative framework, proving that DF is merely a degenerate special case of EF. Extensive experiments demonstrate that a singular EF model surpasses task-specific DF ensembles across standard benchmarks and exhibits robust asymptotic stability in extreme extrapolation. This work propels a paradigm shift in LTSF: moving from passive Static Mapping' to autonomous Evolutionary Reasoning.
Authors:
Yujun Wu, Dongxu Zhang, Xinchen Li, Jinhang Xu, Yiling Duan, Yumou Liu, Jiabao Pan, Qiyuan Zhu, Xuanhe Zhou, Jingxuan Wei, Siyuan Li, Jintao Chen, Conghui He, Cheng TanAbstract: Existing research infrastructure is fundamentally document-centric, providing citation links between papers but lacking explicit representations of methodological evolution. In particular, it does not capture the structured relationships that explain how and why research methods emerge, adapt, and build upon one another. With the rise of AI-driven research agents as a new class of consumers of scientific knowledge, this limitation becomes increasingly consequential, as such agents cannot reliably reconstruct method evolution topologies from unstructured text. We introduce Intern-Atlas, a methodological evolution graph that automatically identifies method-level entities, infers lineage relationships among methodologies, and captures the bottlenecks that drive transitions between successive innovations. Built from 1,030,314 papers spanning AI conferences, journals, and arXiv preprints, the resulting graph comprises 9,410,201 semantically typed edges, forming a queryable causal network of methodological development. To operationalize this structure, we further propose a self-guided temporal tree search algorithm for constructing evolution chains that trace the progression of methods over time. We evaluate the quality of the resulting graph against expert-curated ground-truth evolution chains and observe strong alignment. In addition, we demonstrate that Intern-Atlas enables downstream applications in idea evaluation and automated idea generation. We position methodological evolution graphs as a foundational data layer for the emerging field of automated scientific discovery.
Abstract: Reinforcement learning (RL) has become central to improving language models on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain open: how does pretraining shape RL scaling, and what does RL actually do to the inherited policy? These questions are difficult to study in the standard language-model setting: web-scale pretraining corpora make it hard to attribute behaviors to pretraining versus RL, and systematic compute sweeps across pretraining and RL are prohibitively expensive. To address this challenge, we use chess as a controlled testbed for studying reasoning across the full pretraining-to-post-training pipeline, with verifiable outcomes and a controllable pretraining corpus. We pretrain language models from 5M to 1B parameters on human games, fine-tune them on synthetic reasoning traces, and run RL on chess puzzles with verifiable rewards. Using this framework, we establish a scaling law connecting pretraining and RL: pretraining loss predicts post-RL performance at fixed RL compute, while pretraining token count predicts the rate of improvement with additional RL. Beyond scaling, we find that RL does not simply sharpen the SFT policy: its effect varies qualitatively with puzzle difficulty, primarily amplifying correct moves the SFT policy already preferred on easy puzzles, while on hard puzzles RL surfaces moves that were nearly absent under SFT for a substantial fraction of states. Together, these results provide a quantitative account of the pretraining-to-RL interface and a controlled testbed for studying reasoning across the full training pipeline.
Abstract: Continual reinforcement learning (continual RL) seeks to formalize the notions of lifelong learning and endless adaptation in RL. In particular, the aim of continual RL is to develop RL agents that can maintain a careful balance between retaining useful information and adapting to new situations. To date, continual RL has been explored almost exclusively through the lens of risk-neutral decision-making, in which the agent aims to optimize the expected long-run performance. In this work, we present the first formal theoretical treatment of continual RL through the lens of risk-aware decision-making, in which the behaviour of the agent is directed towards optimizing a measure of long-run performance beyond the mean. In particular, we show that the classical theory of risk measures, widely used as a theoretical foundation in non-continual risk-aware RL, is, in its current form, incompatible with continual learning. Then, building on this insight, we extend risk measure theory into the continual setting by introducing a new class of ergodic risk measures, and showing that it is compatible with continual learning. Finally, we provide a case study of continual risk-aware learning, along with empirical results, which show the intuitive appeal of ergodic risk measures in continual settings.
Authors: Aman Sharma, Sushrut Thorat, Paras Chopra
Abstract: LLM-based coding agents are usually evaluated in familiar software settings: mainstream languages, common libraries, and public repositories. These benchmarks remain important, but they can hide how agents behave when the language itself is unfamiliar. We evaluate six contemporary coding agents on four esoteric programming languages using a sequential setup with file editing, local execution, and hidden-test grading. Our protocol exposes capability differences between these agents that mainstream coding and agentic benchmarks such as SWE-Bench Verified and Terminal-Bench 2.0 compress into much narrower bands. We observe that the strongest agents, Claude Opus 4.6 and GPT-5.4 xhigh, often avoid writing the target language directly. On Brainfuck and Befunge-98, they write Python programs that generate target-language code and debug those generators locally. Forbidding this metaprogramming strategy causes large performance drops. Text guidance distilled from this strategy does not materially improve weaker agents. In contrast, Opus-derived Python helper code for building generators, with no solved benchmark programs or hidden-test answers, sharply improves Sonnet 4.6 and GPT-5.4 mini on the same problems, while Haiku 4.5 remains low. More interpreter calls and output tokens improve stronger agents but leave weaker agents near their original performance, indicating that these resources amplify useful strategies rather than create them. Together, these results show that strong coding agents adapt to unfamiliar languages by using tools, feedback, and workspace state to build a working model of the target language. Metaprogramming is the clearest case, but the broader gap is constructing and debugging a strategy that works under the target language's rules.
PaperID: 6511, Poster
Authors:
Zhengyang Liu, Jianguo Huang, Junxi Hu, Yucheng Mi, Haojie Cai, Wenxuan Shen, zhikun zhang, Renfeng Peng, Guangtao Zhang, Zhiqiang LiuAbstract: Reconstructing physical dynamics from sparse and irregular spatiotemporal observations is a fundamental challenge in scientific research. However, sparsity leaves large regions of the field unconstrained, making the inverse problem ill-posed: a model may fit the observed sensors while producing nonphysical off-sensor structures or unreliable predictions between support times. To address this challenge, we propose PhysFTM---a physics-informed functional Tucker method that parameterizes Tucker mode factors in reproducing kernel Hilbert space (RKHS) and enforces governing physical laws through finite-difference residuals at collocation points. These physical constraints are essential for credible reconstruction, since most query locations are never directly supervised by data. Specifically, PhysFTM first reconstructs continuous spatial fields at observed times via an alternating minimization scheme, with separate treatments for the Tucker core tensor and the RKHS-based factors. To further enable continuous temporal resolution, PhysFTM constructs a continuous trajectory in Tucker-core space, which is decoded by the learned spatial RKHS-FTM representation to recover the full spatiotemporal field at arbitrary query times. Experiments on Allen--Cahn and Navier--Stokes with varying observation ratios demonstrate that PhysFTM achieves superior reconstruction accuracy, improved physical consistency, and stronger continuous-time modeling capability.
PaperID: 6512, Poster
Authors: Zhaohui Fu, Duanyu Feng, Yangshuai Wang
Abstract: Random Feature Models (RFMs) are attractive PDE surrogates because their linear-in-features form reduces training to regularized least squares. Standard RFMs, however, use one stationary spectrum over the whole domain, creating a bandwidth dilemma on heterogeneous PDE solutions: fine scales resolve sharp local gradients but oscillate in smooth regions, while coarse scales miss interface-dominated derivatives. We propose LDD-RFM, a learnable domain-decomposition framework whose local experts remain linear RFMs and whose nonlinear variables specify compact-support partition-of-unity (PoU) gates and local spectral scales. For fixed gates and spectra, expert coefficients are eliminated by ridge regression through variable projection, yielding a differentiable reduced objective with explicit linear-solve and partition-smoothness controls. We prove an H^1 error decomposition in which PoU sharpness enters through a product-rule term. Controlled regression and PDE experiments show that LDD-RFM reduces H^1 errors by 10× over global and fixed-partition RFM baselines under matched feature budgets; exact-solver, timing, and robustness studies check that these gains are not artifacts of the linear solver or evaluation protocol.
PaperID: 6513, Poster
Authors: YIhao Bu, Dan Lian, Zhenguo Gao
Abstract: Few-step diffusion sampling sharpens the tension between ODE and SDE dynamics. ODE samplers are stable but can lose diversity, while the standard SDE endpoint can become unreliable at low numbers of function evaluations (NFE). The challenge is therefore not simply whether stochasticity should be added, but which driver properties allow stochastic correction to contribute diffusion without losing finite-step control. We show that this tradeoff is governed by two driver-level quantities, the asymptotic covariance Q, which sets the effective diffusion level on the ODE-to-SDE spectrum, and the Poisson solution norm C_\chi, which controls the finite-step remainder. Gaussian drivers have C_\chi = \infty, so this finite-step remainder is not uniformly controlled by the same bound. Balancing the diffusion level set by Q with the finite-step control set by C_\chi formalizes a favorable regime for bounded structured drivers with intermediate diffusion strength. This regime directly gives a structured correction method, instantiated as Fast-Driver Sampler (FDS) with the shuffled-Lorenz driver. Across DiT-XL/2 and EDM2, FDS establishes a stronger low-NFE frontier among evaluated training-free samplers, improving ODE baselines while avoiding the degradation of the SDE endpoint.
Authors: Tyler LaBonte, Vidya Muthukumar
Abstract: Neural networks are known to be susceptible to over-reliance on spurious correlations. However, the precise mechanism by which models exploit shortcut features is not fully understood, and algorithms to mitigate this behavior rely on as yet unjustified assumptions about the learned representations. In this work, we provide the first end-to-end theoretical characterization of spurious feature learning for two-layer ReLU neural networks trained by online minibatch SGD on the logistic loss. We consider data drawn from the high-dimensional Boolean hypercube with a quadratic signal function (namely XOR) and a linear spurious correlation. We show that SGD learns the spurious feature first, and exponentially fast. Moreover, the optimization dynamics couple the spurious and signal features, with a stronger spurious component inhibiting signal feature learning. Our analysis reveals precise phase transitions in the learning dynamics. In the first phase, alignment between the signs of the spurious feature and second-layer weight drives rapid growth of the spurious feature. When the spurious correlation is maximally strong, we show theoretically that the spurious feature dominates even at the sample complexity threshold where XOR would be learned in isolation (i.e., if the spurious feature was absent). In contrast, when the correlation strength is constant, we provide preliminary empirical evidence that the model can eventually learn the XOR signal, although the spurious feature is not forgotten.
PaperID: 6515, Poster
Abstract: Architecture attribute predictors reduce the cost of neural architecture search by estimating accuracy or latency from a small set of evaluated candidates. The key difficulty is that neural architectures are operation-labeled DAGs: edge direction defines computational flow, but existing predictors often treat direction as a local structural cue. We propose DTA-GT, a Direction- and Topology-Aware Graph Transformer based on encoder-wide directional consistency, which preserves directed computational semantics across node initialization, global spectral encoding, pairwise interaction, and feature update. DTA-GT realizes this principle with direction- and topology-aware node initialization, Magnetic Laplacian spectral encoding, direction-dual structural attention, and a Direction-Sensitive MoE-FFN. Across accuracy and latency prediction on NAS-Bench-101 and NAS-Bench-201, DTA-GT consistently outperforms representative sequence-, GNN-, Transformer-, and hybrid predictors under limited-budget settings. Ablations and Magnetic spectral controls show that the improvements come from coordinated direction-aware stages and edge-orientation-sensitive spectral information, while mechanistic diagnostics confirm that the modules exhibit the intended directional behavior. These results support encoder-wide directional consistency as an effective design principle for architecture DAG representation.
Abstract: As LLM-based agentic systems grow more complex, they increasingly rely on meta-agents: higher-order agents that act on other agents, much like managers supervise employees. Yet existing agentic runtimes expose execution only as static environmental states, limiting the kinds of live and post-hoc interventions a meta-agent requires. To unlock these capabilities, we introduce Shepherd, a functional programming model that formalizes meta-agent operations on target agents as functions, with core operations mechanized in Lean. Shepherd records every agent–environment interaction as a typed event in a principled Git-like execution trace where any past state can be cheaply forked and replayed. Even at scale, Shepherd forks the agent process and its filesystem 5x faster than Docker, with >95% prompt-cache reuse on replay. We exemplify Shepherd's versatility in three use cases: (1) runtime intervention, where a live supervisor improves pair coding pass rate from 28.8% to 54.7% on CooperBench; (2) counterfactual meta-optimization, where a meta-agent branches to explore alternative paths, beating MetaHarness and GEPA across four benchmarks by up to 11 points with up to 58% lower wall-clock; and (3) Tree-RL training, where a meta-agent forks rollouts at chosen turns, lifting TerminalBench-2 from 34.2% to 39.4% on Qwen3.5-35B-A3B. These use cases show Shepherd is a performant and efficient substrate for programming meta-agents; we open-source it to advance future research.
Abstract: Evaluating and mitigating a generative system's susceptibility to jailbreak attacks is critical to its safe deployment. Given the number of deployable systems, full per-configuration evaluation and optimization is impractical. In this paper, we formalize the behavioral geometry of a population of models that, by leveraging previously evaluated and defended models, supports both efficient susceptibility prediction and effective defense transfer across a population. We apply the framework to 79 models spanning 24 providers and to 100 system configurations of a single base model. Simple methods that use the behavioral geometry reach an AUPRC of 0.94 for susceptibility detection with \approx98% fewer probes relative to a full evaluation. Using the behavioral geometry to select which model to transfer an optimized defense from outperforms same-provider assignment (+2%, p = 0.03) at no additional probe cost, with a set of three models sufficient to cover the population. Results are robust to hyperparameter selection and judge.
PaperID: 6518, Poster
Abstract: Machine learning (ML) methods have achieved notable success in enhancing individual components of mixed-integer linear programming (MILP) solvers. However, most existing approaches focus on single-component enhancement, neglecting critical interactions among components and often yielding diminishing returns when multiple approaches are integrated. The emergence of large language models (LLMs) provides a promising paradigm for jointly optimizing multiple solver components within a unified framework. In this paper, we propose Jolly-MILP, an LLM-guided framework to jointly optimize presolving, cut selection, and branching variable selection in exact MILP solvers. Our approach introduces a synergistic Coordinator to align component-level generators and a hierarchical Controller to produce adaptive, multi-stage strategies. Extensive experiments on nine MILP datasets demonstrate that Jolly-MILP consistently outperforms both ML-based approaches and existing LLM-based methods. These results show that explicit joint optimization, rather than isolated component-wise enhancement, is crucial for acquiring further gains in exact MILP solving, offering a new pathway for more advanced solver design.
PaperID: 6519, Poster
Authors: Lin Chen, Dejan Slepcev
Abstract: We introduce a framework for learning distributions supported on manifolds based on estimating the projection onto the underlying data manifold. Specifically, we define an objective functional over a class of functions whose minimizer recovers this projection. The proposed framework automatically adapts to the intrinsic dimension and allows for data supported on unions of manifolds with varying dimensions. The resulting objective can be optimized using standard neural network architectures. We present population-level results establishing recovery of the manifold projection and prove finite-sample Wasserstein generalization bounds. Furthermore, we validate the approach empirically across a range of settings. In contrast to widely used generative models such as diffusion models, whose objectives can be minimized by memorizing training samples, our formulation penalizes such behavior. Thus, returning training data does not yield low loss. We also introduce a direct memorization metric to verify that the learned projection generates new samples rather than reproducing the training set.
PaperID: 6520, Poster
Abstract: Discretizing continuous actions into skills using methods like VQ-VAE has emerged as a powerful paradigm for robotic manipulation. However, the quantization errors in discretizing continuous actions yield a suboptimal training distribution for the prior, degrading its performance. While reinforcement learning offers a path for refinement, its direct application is challenging, suffering from unstable encoder updates and a granularity dilemma in importance sampling. To address these challenges, we introduce Cascaded Skills Optimization (CSO), a two-stage post-training framework. First, to rectify the initial policy's suboptimal distribution, CSO employs Rejection-Sampling Supervised Fine-tuning to align the model's observation-to-skill mapping with the distribution of successful online trajectories via supervised fine-tuning. Second, to resolve the granularity dilemma, CSO introduces Skills Policy Optimization, which computes an independent, clipped importance ratio for each skill, enabling more stable and efficient updates. Our post-training strategy delivers highly competitive performance on challenging benchmarks like LIBERO and RoboTwin, with its effectiveness further validated on a physical robot.
PaperID: 6521, Poster
Abstract: RL post-training with methods like PPO and GRPO produce a large rollout archive, which each contain correctness-labeled completions, correct/incorrect completions for the same prompt, teacher log-probabilities, and group pass rates. Canonical distillation methods, however, ignore these traces and train based upon new data from the converged teacher. This discards the contrastive signal that helped the teacher and the metadata already produced during teacher RL, even though the teacher RL run typically generates many more tokens than the later distillation pass. However, reusing rollouts naively is non-trivial: the archive mixes early and late teacher policies, can be arbitrarily off-policy for the student, and contains many low-quality traces, and naive SFT and DPO on these rollouts underperforms the cannonical SFT baseline. In this paper, we show that rollout reuse becomes effective when it is treated as a We use two sets of information from rollouts: per-prompt pass rate, which defines an easy-to-hard curriculum and separates all-correct prompts for SFT from mixed-success prompts for DPO; and per-rollout teacher likelihood, which acts as an in-distribution proxy for student-compatible traces. Across two GRPO teachers (4B and 8B) and students from multiple families, the curated rollout datasets D
PaperID: 6522, Poster
Abstract: The standard metric LP for correlation clustering has \Theta(n^2) pairwise marginals and \Theta(n^3) triangle inequalities, while dense pairwise supervision is often the primary bottleneck in large relational pipelines. We separate three sparsification questions that are often treated together: preserving the objective of every integral clustering, rounding from a sparse set of LP marginals, and clustering when the signed weighted input itself is only sparsely observed. First, we prove that the VC dimension of the signed-edge disagreement class induced by all clusterings of n vertices is exactly n-1, so weighted edge sampling yields additive \varepsilon-coresets of size \tilde O(n/\varepsilon^2) with optimal n-dependence. Second, we introduce Sparse-LP-Pivot, which imputes missing LP marginals from triangle witnesses, and analyze it at two levels: a universal but coarse perturbation bound for every Lipschitz LP-PIVOT rule, and a sharper sparse-rounding theorem under an explicit weighted influence-stability certificate. For pseudometric-weighted CC, this gives a 10/3 exact-marginal baseline and an unconditional coarse sparse-marginal bound; a sharp same-edge bound governed by \overline\Gamma_w follows conditionally if the rounding rule is additionally proved to satisfy edge-weight-dominated influence stability. Third, we show that uniform marginal sampling has a triangle-witness threshold at m=\Theta(n^3/2) for typical pairs, while m=O(n^3/2\sqrt\log n) is sufficient for all pairs with high probability. Finally, in the stricter sparse edge-observation model, we prove an \Omega(n^3/2) lower bound on the expected approximation ratio from o(n) uniformly sampled edges for general weighted instances. Experiments illustrate the witness transition and the proposed diagnostics; the formal guarantees are stated independently of those empirical observations.
Abstract: Vision-Language Models (VLMs) have achieved strong performance on general multimodal reasoning, yet remain challenged in integrating nonlocal visual information to support semantically underdetermined visual reasoning. We describe this challenge as Fragmented Visual Reasoning. To this end, we propose Credit Assignment for Visual Evidence (CAVE), a structured process-reward method based on GRPO for interleaved visual reasoning. Specifically, CAVE evaluates the contribution of intermediate steps at the action level via three complementary reasoning process signals: belief update, evidence acquisition, and adaptive focus control, thereby guiding the model to optimize each reasoning action and learn more reliable visual reasoning strategies. Meanwhile, we construct TRACER-Bench, which covers four nonlocal and semantically confusable reasoning dimensions and provides key intermediate evidence to supervise reasoning paths. Experiments demonstrate that CAVE substantially improves performance on tasks requiring fragmented visual evidence integration, covering both public benchmarks and our newly introduced TRACER-Bench, while retaining competitive performance on general multimodal evaluations. Further analyses reveal that CAVE effectively improves the visual reasoning capacity and exhibits stronger robustness under longer-range and deeper cross-region dependencies. Our code and the proposed benchmark are available at: https://anonymous.4open.science/r/Anonymous_code-2D8B/.
PaperID: 6524, Poster
Abstract: Auction-based Federated Learning (AFL) provides a market-based mechanism to incentivize decentralized data owners (DOs) to join FL model training initiated by data consumers (DCs). However, optimizing the DC's bidding strategy remains fundamentally challenging. Existing methods predominantly rely on Reinforcement Learning (RL), which suffers from severe instability caused by the mismatch between stepwise Markovian rewards and the trajectory-level, non-decomposable feedback inherent in AFL. To address this mismatch, we transition from the stepwise RL formulation to a distributionally robust black-box optimization framework. By treating the entire recruitment-to-training process as a single high-fidelity function evaluation, this transition eliminates the need for explicit reward decomposition. We propose DR-AFL , which accounts for market non-stationarity and data distribution shifts by optimizing policies against worst-case environments within a Wasserstein ambiguity set. Our approach further incorporates an adversarial reweighting scheme to enhance robustness, alongside a dynamic safety projection that leverages real-time cost feedback to enforce budget-aware exploration. We provide theoretical analysis showing that DR-AFL induces a smooth surrogate over the discontinuous auction landscape and guarantees a robust performance lower bound under distributional shifts. Extensive experiments on 6 widely adopted datasets show that DR-AFL consistently outperforms state-of-the-art RL-based baselines, achieving a 2.43% improvement in model accuracy while reducing the number of training cycles required to reach target accuracy by 62.3%.
PaperID: 6525, Poster
Authors: Bui T Duc, Huynh Thi Thanh Binh
Abstract: Constrained preference-conditioned MORL typically assumes full task observability, treating preferences and safety budgets as stationary task signals. In real-world deployment, however, these signals are often unreliable proxies: intent may drift, safety tolerances may shift with context, and exposed descriptors may be noisy, delayed, or incomplete. We therefore formulate constrained preference-conditioned MORL as a partially observed task-specification problem, where the agent must navigate a fundamental conflict between unreliable external signals and latent interaction dynamics. We propose BECON, a belief-conditioned framework for constrained MORL with drifting preferences and budgets, which infers latent task context from recent histories to estimate preferences and safety budgets. These estimates are fused with the exposed task signal via uncertainty-dependent gates, enabling the policy to defer to observations under high epistemic uncertainty while correcting them when history is informative. The fused context jointly conditions preference optimization and context-adaptive constraint enforcement. Experiments on diverse multi-objective control tasks demonstrate that BECON consistently improves over state-of-the-art MORL baselines, achieving stronger robustness and constraint satisfaction under task drift.
PaperID: 6526, Poster
Authors: Zy Li
Abstract: Mamba's selective state-space model (SSM) achieves long-range dependence through input-dependent gating, yet existing analyses offer limited insight into why a particular parameter configuration produces a particular effective receptive field (ERF). We define per-step conductances G_t = I - \exp(A\,\Delta_t) from the zero-order-hold discretization and construct a signed spectral measure whose atoms are the per-dimension conductances weighted by the corresponding input-output residues. The impulse response at lag \tau is a discrete Laplace transform of this measure, and the ERF is governed by the per-channel squared impulse response weighted by output amplitude. Heavy tails arise from the mixture across state dimensions, not from temporal randomness within any single dimension. We extract conductances from trained Mamba models and report three empirical findings. First, replacing the signed measure with its total-variation envelope overestimates the ERF by more than three orders of magnitude. Signed cancellations among residues are what make the prediction quantitative. Second, the framework generalizes beyond the original single-layer validation setting. Applied to held-out A-initializations, a two-layer model trained from scratch, and all 48 layers of a pretrained Mamba-370M, the spectral measure tracks the frozen-parameter ERF within 0.85-1.5× on 41 of 48 pretrained layers, with log-log slopes matching to within 0.05 on 35 layers and within 0.20 on all 48. Third, varying the initialization of A causally shifts the ERF tail slope as the spectral prediction anticipates (-0.92 for slow init vs. -2.6 for default), and this shift translates to task performance. Across two long-context tasks evaluated over three random seeds, the slow initialization is the most reliable condition: it solves MQAR in 3/3 seeds (vs. 1/3 for both default and fast), and attains the highest mean selective-copy accuracy on every distance bucket (0.68\pm0.13 overall vs. 0.34\pm0.15 for fast at distances up to 1500 tokens, with the ranking preserved in a single-seed extension out to 4000 tokens). The spectral measure thus provides a controllable design lever for long-range memory. The closed-form prediction reproduces the in-sample frozen-parameter ERF at ratio 1.00 by construction, which we use as a pipeline consistency check.
PaperID: 6527, Poster
Abstract: Over the past decade, a growing body of work has established that implicit biases in learning dynamics fundamentally shape the solutions found by neural networks, governing their learned features and generalisation performance. However, despite these strides in deep learning theory, far less is known about representation learning in generative models based on dynamical systems, from normalising flows to diffusion. Importantly, continuous-time generative models can be prone to mode collapse, yet the learning dynamics underpinning such biases remain elusive. To address this, we use tools from dynamical systems and operator theory to show that bounds on the rank of the gradient of the model weights can explain how collapse occurs in generative models. We show that, in both variational inference and flow matching tasks, optimisation induces low-rank biases that encourage parsimonious representation learning but may also cause the learned distribution to collapse. Together these findings point to a trade-off between promoting feature learning while avoiding both memorisation and collapse in dynamical systems-based generative models.
PaperID: 6528, Poster
Authors: Binglin Hao, Huajun Liu
Abstract: Unsupervised domain adaptation(UDA) for semantic segmentation transfers knowledge from synthetic to real domains, where geometric cues such as depth are commonly exploited to reduce the domain gap. However, existing depth-aware methods fail to explicitly model the domain-invariant spatial structures among semantic instances, and current self-training schemes weight pseudo-labels solely by prediction confidence, ignoring geometric consistency, which limits their reliability under domain shift. To address these issues, we propose an instance-wise geometric modeling framework that captures inter-instance spatial relations, including directional concentration and dominant direction, beyond conventional depth representations. These geometry-aware features are further integrated into pseudo-label weighting to enforce geometric-semantic consistency during self-training, leading to more stable and accurate pseudo supervision. Experiments on standard benchmarks show that our method significantly outperforms state-of-the-art UDA approaches.
PaperID: 6529, Poster
Abstract: Remote sensing imagery supports diverse visual tasks, such as object detection and semantic segmentation, but these tasks vary across scenes, imaging modalities, and target objects. Since many remote sensing scenarios lack sufficient annotations, conventional supervised learning is often difficult to apply, motivating in-context learning as a flexible alternative.We present an in-context learning framework for remote sensing visual tasks, where a new task is specified by a single annotated example and applied to unseen images without retraining at inference time. Built upon diffusion models, our framework establishes target-reference associations through cross-attention and transfers the structural relationship between the reference image and its annotation to produce task-specific predictions. Unlike natural images, remote sensing objects are often small and provide weak semantic cues in early diffusion steps, making semantic alignment between target and reference images difficult. We address this limitation with a semantic-aware cross-attention that combines semantic and fine-grained details for more accurate cross-image matching. Remote sensing objects also appear at arbitrary orientations, causing large appearance variations and unstable in-context transfer. To improve robustness, we introduce a rotation-robust learning strategy that reduces sensitivity to orientation changes. Experiments show that our method achieves strong one-shot performance, establishing a flexible paradigm for adapting remote sensing models to diverse visual tasks.
Authors: Tianyi Wang, Huawei Fan, Yuanchao Shu, Peng Cheng, Cong Wang
Abstract: LLM inference is inherently expensive, even a modest slowdown can translate into substantial operating costs and severe availability risks. Recently, a growing body of research known as latency attacks focuses on crafting inputs to trigger worst-case output lengths. However, we report a contrary finding that these algorithmic-level latency attacks are largely ineffective against modern LLM serving systems. We reveal that system-level optimization such as continuous batching provides a logical isolation to mitigate contagious latency impact on co-located users. Thus, in this paper, we shift our focus from the algorithm to the system layer, and introduce a new Fill and Squeeze attack strategy targeting the state transition of the scheduler. "Fill'' first exhausts the global KV cache to induce Head-of-Line blocking, while "Squeeze'' forces the system into repetitive preemption. By manipulating output lengths using different attack prompts, and leveraging side-channel probing of memory status, we demonstrate that the attack can succeed in a practical black-box setting with much less cost. Extensive evaluations on vLLM indicate up to 75-742× TTFT degradation relative to benign baselines and 1.5-4× average slowdown on Time Per Output Token compared to existing attacks with 30-40% lower attack cost. Code: https://anonymous.4open.science/r/FS-EE97/README.md
PaperID: 6531, Poster
Abstract: Proactive task-oriented dialogue (TOD), such as outbound sales, demands a persuasive agent that actively probes the user's concerns and steers the conversation toward acceptance within a bounded number of turns. Yet post-trained LLMs are inherently conservative, and reward-shaping RL (e.g., GRPO) struggles since it only re-weights what an already passive policy samples. We show that conditioning on the user's latent concerns unlocks proactive capability that no amount of sampling can undermine, establishing these concerns as a pivotal training-time signal. To operationalize this finding, we build the Cognitive User Simulator, which models each user as a stratified persona comprising observable external traits and hidden internal concerns. The simulator produces faithful and diverse interactions, while emitting per-turn state dynamics that track persuasion progress. We then introduce Simulator-Induced Asymmetric-View Policy Optimization, which converts the modeled concerns and the simulation state transition into complementary training objectives: (1) \emphAsymmetric On-Policy Self-Distillation that transfers concern-aware behavior from a privileged view of the same policy into its deployable, conversation-only view; and (2) \emphState-Transition Policy Refinement where the final decision provides trajectory-level advantage while synchronous state transitions refine turn-level credit direction. Across two real-world food-delivery benchmarks (\emphmerchant and \emphcourier outbound recruitment), our method matches or surpasses leading proprietary LLMs, outperforms strong RL baselines, and generalizes across various LLM-based user simulators.
Authors:
Taofeng Xue, Chong Peng, Mianqiu Huang, Linsen Guo, Tiancheng Han, Haozhe Wang, Xiaocheng Zhang, Xin Yang, Dengchang Zhao, Jinrui Ding, Xiandi Ma, Yuchen Xie, Peng Pei, Xunliang Cai, Xipeng QiuAbstract: The development of native computer-use agents (CUA) represents a significant leap in multimodal AI. However, their potential is currently bottlenecked by the constraints of static data scaling. Existing paradigms relying primarily on passive imitation of static datasets struggle to capture the intricate causal dynamics inherent in long-horizon computer tasks. In this work, we introduce EvoCUA, a native agentic model that integrates data generation and policy optimization into a self-sustaining evolutionary cycle. This approach employs a verifiable synthesis engine to autonomously generate diverse tasks with executable validators, alongside a scalable infrastructure orchestrating tens of thousands of sandbox rollouts for mass experience acquisition. To internalize this experience, our iterative evolving learning strategy reinforces successful routines while transforming failure trajectories into rich supervision through error analysis and self-correction.EvoCUA achieves a 56.7% success rate on the OSWorld benchmark, establishing a new state-of-the-art for open-weights models across multiple model scales. It significantly outperforms the previous best open-source model, OpenCUA-72B (45.0%), and surpasses leading closed-weights models such as UI-TARS-2 (53.1%), offering a highly parameter-efficient and reproducible alternative to massive closed-weights models. These results demonstrate the generalizability of the evolving paradigm across foundation models of varying scales, establishing a robust and scalable path for advancing native agent capabilities.
PaperID: 6533, Poster
Abstract: Selection bias is a difficult yet widespread problem in causal discovery, occurring whenever non-random selection processes lead to data that is not representative of the underlying populations. Due to its practical importance, related works have been proposed to address the problem under local [Versteeg et al., 2022] or interventional settings [Dai et al., 2025a], but general algorithms for observational data remain elusive. In this paper, we focus on truncated data (hard selection), and provide a general result for Additive Noise Models (ANM). We show that score-based methods are a surprisingly suitable candidate due to a special property of the score function (\nabla \log p(x)): deterministic truncation preserves both the score of the joint distribution as well as the Jacobian of the score. While established score-matching algorithms still fail in most settings, we actively leverage this observation to propose a new method for causal discovery that is robust under truncated data with ANMs. Overall, our approach offers a general, robust solution that is agnostic to external information about the selection process and still achieves comparable performance to the state-of-the-art approaches when selection is not present.
PaperID: 6534, Poster
Abstract: Conformal prediction provides finite-sample marginal coverage guarantees under exchangeability, but marginal validity alone does not ensure that prediction intervals adapt to local structure in the conditional distribution. Existing adaptive methods such as conformalized quantile regression (CQR) and conformal histogram regression (CHR) improve conditional coverage by incorporating estimated quantiles or conditional densities into the conformity score, but doing so requires direct estimation of the conditional distribution of response Y given covariates X, which can be statistically and computationally demanding. We propose RLR-CHR, a residual-based local rank conformal histogram regression method that sidesteps full conditional distribution estimation by decoupling location and scale effects via a learned noise proxy and constructing histogram-based conformity scores on the standardized residuals. Local calibration is then performed through rank comparisons rather than weighted conformal quantiles, yielding a procedure that is both localized and computationally efficient. We establish finite-sample marginal coverage for RLR-CHR under exchangeability and asymptotic conditional coverage under local stability of the residual score distribution. We further prove that RLR-CHR is asymptotically equivalent to its weighted-localization counterpart RBC-CHR under suitable regularity conditions, providing a precise theoretical justification for the reduction in per-query computational cost without sacrificing asymptotic accuracy. Experiments on synthetic and real datasets demonstrate that RLR-CHR consistently improves worst-slab coverage and produces more compact prediction intervals than existing conformal approaches.
Abstract: Large language models (LLMs) are increasingly deployed in enterprise settings where they handle sensitive documents and user context, raising acute concerns over security and controllability. Conventional access control regulates whether information is accessible to the model, yet leaves how the model uses that information at generation time largely unconstrained: once sensitive content enters the context, outputs may still drift beyond a user's authorized scope. We present Permit, a novel permission-aware representation intervention framework that closes this gap by enforcing fine-grained control directly on the model's hidden states. Through exploratory analysis, we find that permission conditions induce hidden-state shifts that are (i) separable across permissions and (ii) concentrated in a small set of dominant directions. Permit exploits this geometry in two stages: it first identifies a permission-sensitive subspace from activation differences across permission conditions, and then performs lightweight interventions within this subspace to steer generation, with two concrete instantiations (offset-based and gated). Both operate atop a frozen backbone with only a handful of permission-specific parameters, achieving precise control with minimal overhead. Experimental results demonstrate that \textscPermit performs better than the state-of-the-art method across multiple permission settings while driving information leakage to near zero, achieving over (18%) F1-score improvement with (>98%) fewer trainable parameters.
PaperID: 6536, Poster
Abstract: Spatio-temporal point processes (STPPs) represent a random collection of points where each point corresponds to the time and location of an event. STPPs are widely used to describe a wide range of phenomena, from earthquakes to wildfires to crime occurrence to disease outbreaks. Generative models have recently emerged as a new powerful paradigm for STPPs, demonstrating %due to their exceptional highly competitive generalization capabilities and promising potential for systematic uncertainty quantification. However, existing generative approaches primarily rely on unimodal numerical data, overlooking the geographic context in which events occur and failing to capture critical spatial and temporal patterns within the STPP. To address these limitations, we propose Morse-aware Cross-Modality learning within Diffusion Model, or MCM-DM. MCM-DM represents a cross-modality diffusion framework based on the discrete Morse and cobordism theories that augments STPP modeling with geographic scene understanding. MCM-DM consists of three key components: (i) a vision-language model-based geographic encoder that extracts semantic embeddings from satellite tiles at each event location; (ii) an attention-based fusion mechanism to integrate critical structure representation with spatio-temporal embeddings, and (iii) a Morse-theoretic topological aligner which aligns latent representations of critical events with spatio-temporal embedding space. We demonstrate the MCM-DM utility on 8 diverse STPP datasets from Earth sciences, epidemiology, urban mobility, and crime analytics, highlighting its cross-domain versatility. The code is available at~\urlhttps://anonymous.4open.science/r/MCM-DM-CF83.
Authors: Ammar Mahran, Artavazd Maranjyan, Peter Richtarik
Abstract: Asynchronous stochastic gradient descent (ASGD) is a standard way to exploit heterogeneous compute resources in distributed learning: instead of forcing fast workers to wait for slow ones, the server updates the model whenever a gradient arrives. Vanilla ASGD applies each arriving gradient with the same weight. When local data distributions are heterogeneous, this becomes problematic: faster workers contribute more updates, and we show theoretically that the method is biased toward a frequency-weighted average of the local objectives rather than the desired global objective. Existing remedies typically move away from the simple ASGD template by introducing gathering phases, buffering, or extra memory. We show that this is unnecessary. Keeping the standard ASGD mechanism, we recover the correct objective by rescaling worker-specific stepsizes in proportion to their computation times, so that each worker contributes the same aggregate learning rate over a cycle. In the non-convex setting, under smoothness and bounded heterogeneity assumptions, we prove that the resulting method, Rescaled ASGD, converges to stationary points of the correct global objective in the fixed-computation model. Its time complexity matches the known lower bound in the leading term, while the effects of staleness and data heterogeneity appear only in lower-order terms. Experiments confirm that the method converges to the correct objective and is competitive with state-of-the-art baselines.
PaperID: 6538, Poster
Abstract: Model-based reinforcement learning (MBRL) achieves strong sample efficiency by planning within learned latent dynamics, yet its performance degrades substantially under unseen visual distractions such as background variations, lighting changes, or camera shifts. Unlike model-free RL, where encoder perturbations affect only single-step predictions, MBRL suffers from a two-level vulnerability: visual distractions first push encoder outputs out of distribution, and these errors then compound through recursive latent rollouts over the planning horizon. We propose VIsual Generalization via latent-space cOnsistency in model-based RL (VIGOR), a framework that enables zero-shot generalization to unseen visual distractions while retaining the sample efficiency of its MBRL backbone. VIGOR integrates three interdependent components: (i) asymmetric weak-to-strong augmentation, which pairs weak-only and weak-to-strong latent views within a single batch; (ii) dynamics-level consistency, which enforces augmentation-invariant transition predictions through direct latent regression; and (iii) encoder-level stabilization, which prevents encoder drift under the cross-augmentation supervision imposed by dynamics-level consistency. Evaluations on the DeepMind Control Suite (DMC) and Robosuite show that VIGOR outperforms state-of-the-art model-free and model-based baselines, surpassing the second-best baseline by 3.4% on DMC and 43.6% on Robosuite. Ablations further show that VIGOR's robustness is augmentation-agnostic: replacing the default augmentation with alternatives from distinct perturbation families preserves strong generalization, confirming that latent-space consistency, not the augmentation choice, drives robustness. The source code is available at https://anonymous.4open.science/r/vigor-B641/.
PaperID: 6539, Poster
Authors: Chashi Mahiul Islam, Samuel Jacob Chacko, Mao Nishino, Canlin Zhang, Xiuwen Liu
Abstract: Deep neural networks demonstrate remarkable generalization despite being highly over-parameterized, challenging traditional statistical learning theories. While several theoretical paradigms such as margin-based bounds, PAC-Bayesian analyses, and neural tangent kernels provide global explanations, they often lack mechanistic insight into how generalization emerges from training dynamics. We introduce Neural Expansion, a unified geometric framework that characterizes generalization through the expansion of directionally-selective ray coverage in the input space. Using Generalization Intervals (GIs), we quantify direction-specific regions around training samples where the model's predictions remain stable. Our extensive empirical studies across diverse architectures and datasets reveal that training leads to progressive expansion of ray coverage, particularly along directions with low Jacobian singular values. This mechanism helps explain a large fraction of correctly classified test inputs, adversarial examples, and even mislabeled or out-of-distribution samples through the geometry of ray coverage rather than instance-level memorization. By connecting input-space curvature to predictive behavior, our framework provides a mechanistic foundation that complements existing global theories of deep learning. The link to our code and data will be made available in the published version.
PaperID: 6540, Poster
Authors: Ricardo Luna Gutierrez, Vineet Gundecha, Rahman Ejaz, Varchas Gopalaswamy, Riccardo Betti, Sahand Ghorbanpour, Aarne Lees, Soumyendu Sarkar
Abstract: The achievement of practical fusion energy remains one of the most pressing unsolved scientific challenges, with immense implications for carbon-free power. A critical factor in the success of Inertial Confinement Fusion (ICF) is the design of a Laser Pulse Shape (LP) that can optimally drive implosions under stringent physical constraints. Traditional LP design relies on computationally expensive simulations and manual iterative refinement. We introduce the ICF Laser Pulse Shape Design System (LPDS), a generative inverse modeling framework that maps desired outcomes and target pellet configurations directly to optimized LPs. Crucially, we design a multi-objective loss function to ensure the generated LPs adhere to fundamental physical constraints and experimental feasibility. Furthermore, we present constraint-conditioning, inpainting, and gradient-based LP editing mechanisms for maximum fine-grained control over specific pulse characteristics during generation. Moreover, we validate our framework for LP design on a real-world experimental data. Our method establishes a data-driven inverse design framework for LP in ICF, contributing to the advancement of practical and sustainable fusion energy.
Abstract: Most work on improving large language models treats accuracy as the sole objective. We argue that the harness — the Python code surrounding the model that constructs prompts, routes calls, and parses outputs — is a first-class design surface whose quality is inherently multi-objective: an accurate harness that refuses no unsafe request, or that consumes an order of magnitude more tokens, is not a good harness. We present Meta-Harness, a system that casts harness design as search over three per-domain objectives — accuracy, behavioural safety, and token cost — solved by an agentic proposer (Claude Code) with full filesystem access to prior harness source, execution traces, and scoring artifacts. Our central finding is that a single-phase joint-reward proposer (MoMHa) outperforms every alternative, including a two-phase "accuracy then tokens" ablation, scalar-only feedback, and an accuracy-only baseline. We evaluate on seventeen domains: seven synthetic capability suites, seven real-world public benchmarks (HumanEval, MBPP, Spider, FEVER, MMLU-Pro, LawBench, NuminaMath), and three U-SafeBench-derived user-specific safety domains, using a 12-model fleet spanning four families. On the synthetic track MoMHa achieves a joint mean of 0.482 versus 0.198–0.422 for ten baselines, winning 7/10 per-domain columns; on the real-world track it scores 0.461 versus 0.377 for the strongest baseline (DSPy), winning 5/7 columns — demonstrating that harness strategies transfer to unseen benchmarks without retraining on 8 of 12 target models. MoMHa attains the highest measured behavioral safety composite (U-SafeBench, 0.781) and uses 95 fewer tokens per example than the two-phase alternative. We will release all harness code, evaluation infrastructure, and cross-model logs.
PaperID: 6542, Poster
Abstract: Video world models aim to generate temporally coherent visual observations conditioned on actions, but often suffer from 3D geometric inconsistencies such as spatial drift, object deformation, and perspective violations, which limit their reliability in downstream tasks. We propose WorldPrism, a post-training framework that improves 3D consistency via reinforcement learning without modifying the base model architecture. The central challenge lies in designing a reliable reward signal. To address this, we introduce Bidirectional Geometry-Flow Verification, a bidirectional cross-space mechanism that evaluates consistency through two complementary paths: (i) lifting 2D optical flow into 3D to measure position-space agreement, and (ii) projecting reconstructed 3D geometry to match the point correspondence. By leveraging these two independent and complementary signals, this design mitigates the blind spots of unilateral 2D evaluation. We further incorporate hierarchical temporal evaluation to capture both local and long-range geometric consistency. Experiments on diverse datasets demonstrate that WorldPrism achieves state-of-the-art 3D consistency while preserving visual quality and generative diversity.
Abstract: Recent work shows that preference alignment objectives can be interpreted as divergence estimators between aligned (preferred) \& unaligned (less-preferred) distributions, yielding a principled recipe for designing alignment losses. However, this view has so far been limited to preference-based supervision. We extend it to general LLM alignment, including reinforcement learning with verifiable rewards (RLVR), where alignment feedback is given only as scalar rewards. We introduce f-Group Relative Policy Optimization (f-GRPO), a class of on-policy RL objectives, and f-Hybrid Alignment Loss (f-HAL), which combines on-policy reward optimization with off-policy preference supervision. We show that these objectives estimate f-divergences between reward-aligned \& reward-unaligned distributions induced by above- \& below-average reward responses, and prove expected reward improvement after alignment. Empirically, f-GRPO improves over GRPO on math-reasoning RLVR tasks, while hybrid f-HAL mitigates reward hacking in on-policy safety alignment when verifiable rewards are unavailable and learned reward models must be used.
PaperID: 6544, Poster
Authors: Yifan Wang, Zhixiang Hao, Yu Wang, Congchao Zhu
Abstract: Photorealistic Style Transfer (PST) aims to transfer the color and tonal style of a reference to a content image while strictly preserving its structural integrity. However, existing deep learning-based methods inherently suffer from semantic entanglement caused by pre-trained image encoders, leading to unnatural spatial distortions. Moreover, current pixel-level mapping paradigms often ignore color gamut topology, resulting in color banding, while also lacking the multimodal capability for intuitive text-driven control. To address these bottlenecks, we propose StatLUT, a novel statistical feature-driven multimodal 3D LUT generation framework. First, we bypass traditional encoders and introduce a Lab-Extractor to derive spatially-agnostic statistical features, fundamentally decoupling color distributions from structural semantics to ensure artifact-free rendering. Second, we formulate LUT generation as a Transformer-based Seq2Seq translation task, utilizing a Multi-dimensional Residual Mapper (MR-Mapper) to predict topologically smooth 3D LUTs. Finally, to break the single-modal barrier, we propose the H-Diffuser, a lightweight Diffusion Transformer that directly synthesizes statistical features from natural language prompts, enabling flexible text-driven color grading. Extensive experiments on standard benchmarks demonstrate that StatLUT significantly outperforms state-of-the-art methods in both visual quality and quantitative metrics, pioneering a highly robust and flexible paradigm for multimodal photorealistic style transfer.
PaperID: 6545, Poster
Abstract: A key bottleneck in adversarial transfer is a trajectory-level geometric disconnect: ambient gradients often drift away from the intrinsic data manifold, causing surrogate-specific overfitting. To rectify this, we propose Manifold Anchored Bilevel Transfer (MABT), a unified framework that anchors adversarial trajectories to the shared semantic subspace. MABT introduces a relaxed manifold-anchoring operator as a semantic rectifier to suppress off-manifold noise. With this constraint, we cast transfer attack generation as a distributional bilevel optimization problem that learns a geometry-aligned initialization by minimizing expected transfer risk under a surrogate uncertainty distribution. We further develop a Hessian-free solver with linear-time complexity to handle the resulting hierarchy. Experiments demonstrate that MABT consistently boosts the transferability of 10 baselines across diverse attack scenarios and defense mechanisms (e.g., 63.31% average ASR \uparrow across 28 attacker combinations).
PaperID: 6546, Poster
Abstract: Estimating long-term treatment effects is essential for scientific and industrial evaluation, yet the limited follow-up duration of randomized controlled trials (RCTs) poses a significant challenge. Even the growing use of open-label extensions fails to resolve this issue, as such designs systematically censor long-term control trajectories. While supplementing RCTs with real-world evidence is a common strategy, existing methods often rely on restrictive exchangeability assumptions that fail under latent confounding. We propose a novel proximal framework to integrate long-term real-world evidence into RCTs with open-label extensions. Our approach successfully identifies long-term effects even in the presence of time-varying latent confounding effects and inter-study outcome drift. We provide a doubly robust locally efficient estimator that provides further resilience to model misspecification. Empirical results demonstrate that our framework maintains validity in complex data-generating processes that compromise standard approaches.
PaperID: 6547, Poster
Abstract: Repository-level documentation generation is vital for human comprehension and autonomous agents, yet existing approaches fail to align high-level architectural intent with granular implementation details, resulting in documentation that lacks both global coherence and local fidelity. To address this, we propose RepoScope, a repository documentation generation framework that achieves alignment via two synergistic innovations. Specifically, RepoScope employs Physico-Logical Partitioning by jointly modeling directory structures and call graphs, effectively bridging the gap between physical layouts and logical dependencies while creating coherent documentation units. Within this partitioned space, we introduce SAIL (Seeded Architectural Intent Lift), a mechanism that projects architectural intent top-down to steer granular analysis, which is then revised and lifted bottom-up via code evidence to ensure global-local consistency. Our experimental results on CodeWikiBench demonstrate state-of-the-art performance in aligning with human-authored documentation. Furthermore, the validation on real-world industrial systems shows that RepoScope yields superior correctness and completeness in empowering autonomous agents for CLI extraction compared to existing baselines.
PaperID: 6548, Poster
Abstract: Graph anomaly detection (GAD) is critical to many real-world systems. However, supervised GAD models often degrade under distribution shifts due to the closed-world training assumption. Test-time adaptation (TTA) offers a promising solution by updating pre-trained models with unlabeled test data. Yet, applying TTA to GAD poses a unique dilemma. High-confidence samples provide reliable pseudo-labels, but they are scarce and biased toward low-shift regions. Low-confidence samples better reflect target shifts, but they are too noisy for direct cross-entropy optimization and often require label-free objectives that are misaligned with anomaly discrimination. To address this dilemma, we propose SADOS, a Shifted Augmentation-based Dynamic Objective Scheduling framework for test-time GAD. SADOS turns unreliable target samples into drift-carrying signals for reliable supervised adaptation. It distills prototypes from confidence-stratified unreliable samples and injects them into reliable samples to generate label-preserving yet shift-oriented augmentations. This expands coverage over shifted target regions. In parallel, SADOS dynamically schedules the adaptation objective. It uses contrastive learning to capture early-stage drift signals, and gradually anneals its weight so that classification-oriented anomaly detection dominates later adaptation. Extensive experiments demonstrate that SADOS robustly improves test-time GAD under distribution shifts.
PaperID: 6549, Poster
Abstract: Missing modality remains a longstanding challenge in multimodal learning. Existing methods address this issue through modality recovery or adaptive strategies. However, they overlook models' internal cross-modal dependencies formed during multimodal training, which later impair robustness. We systematically characterize a counterintuitive deployment-time failure mode: models trained on full modalities can underperform unimodal models when one modality is missing at inference time. This pattern appears across diverse architectures, such as fusion models, CLIP-style two-tower models, and vision-language models. We show that such degradation is closely associated with learned cross-modal dependencies in the principal parameter subspaces. Multimodal training induces structured rotations of these subspaces, particularly in cross-modal interaction layers. These rotations are associated with reduced task-aligned margins and larger task-aware representation harm under missing-modality inputs. We propose Geodesic Unlearning (GU), a lightweight parameter-editing method that leverages Grassmannian subspace geometry for structured subspace correction to improve missing-modality robustness. It decomposes layer weights into principal and residual components, treats the principal subspace as a point on the Grassmannian, and rotates it toward a unimodal reference along a geodesic path. Experiments across architectures and datasets show that GU improves performance under missing-modality inference while preserving full-modality accuracy, outperforming strong missing-modality robustness baselines. These findings support a geometric view of deployment-time missing-modality degradation and suggest localized subspace editing as a practical route for robustness correction.
PaperID: 6550, Poster
Abstract: Evaluating large language models (LLMs) via online human feedback is prohibitively expensive and time-consuming. Off-policy evaluation (OPE) addresses this by estimating LLM policy performance from offline data. However, traditional OPE estimators such as Inverse Propensity Score (IPS) suffer from high variance. Marginalized IPS (MIPS) addresses this variance issue by reweighting the IPS weights via condensed semantic embeddings. However, MIPS assumes embeddings contain sufficient reward-predictive information, an assumption that off-the-shelf embeddings often fail to satisfy. To address this, in this paper, we propose Semantic Bottleneck Embedding (SBE), which learns representations for marginalized reweighting via a conditional information bottleneck, compressing reward-irrelevant information while preserving reward-predictive content to achieve both bias and variance reduction. Our objective targets minimizing the derived mean squared error upper bound of the resulting MIPS estimator. Empirically, evaluation on the HelpSteer2 and UltraFeedback datasets across Qwen, Gemma, and Llama policies show SBE-MIPS reduces MSE over prior baselines.
PaperID: 6551, Poster
Abstract: What if the best way to search a video with a vision-language model (VLM) is to stop asking the model to search? Recent approaches give the VLM control over temporal navigation, tying search quality to the model's reasoning ability. We propose the opposite: BeliefSearch treats the VLM as a noisy sensor and delegates navigation to an external Bayesian controller. A single belief over video segments anchors the system: the question sets the prior, each VLM observation updates it, and the same belief decides whether to search, which segment to examine next, and when to stop. Because the belief is explicit, we can trace why each choice was made. The same signal also drives training: each turn earns credit for the uncertainty it reduces, and the final reward is gated by how much the search lowered entropy. This blocks a reward-hacking shortcut where the model otherwise learns to skip search and guess from the preview frames. On long-video benchmarks, our method outperforms all prior search-based methods by up to 8.8 points while using 5 to 13× fewer frames than recent RL search baselines.
PaperID: 6552, Poster
Abstract: Industrial anomaly detection is evolving from category-specific detectors toward unified models capable of handling multiple categories simultaneously. However, in multi-class setting, greater inter-category diversity not only compels enhanced reconstruction capacity to model normal patterns but also exacerbates the “identical shortcut” problem, wherein anomalies are likewise well reconstructed. Moreover, existing unified frameworks are typically confined to single-modal (RGB) inputs and lack multimodal capability. To address these, we propose an Efficient Cross-modal Feature Reconstruction (ECFR) model for unified multi-class anomaly detection, which harnesses the inherent difficulty of cross-modal reconstruction to alleviate the shortcut issue and amplify anomaly discrimination. The framework is built on two core modules: an Adaptive Feature Interaction and Recalibration (AFIR) module and a Hybrid Attention Convolution (HAC) module. Through feature interaction, fusion, and reconstruction, our model achieves two key outcomes: learning normal patterns of multi-class objects while forcing reconstruction failures for anomalous inputs. Extensive experiments on the MVTec 3D-AD and Eyecandies datasets validate our approach, which achieves state-of-the-art performance in multimodal multi-class anomaly detection while requiring significantly fewer model parameters (44.63 M) and lower computational complexity (15.77 GFLOPs) compared to previous multimodal anomaly detection methods.
PaperID: 6553, Poster
Authors: Nikita Pospelov, Olga Ivashkina, Plusnin Viktor, Olga Rogozhnikova, Anna Ivanova, Ksenia Toropova, Konstantin Anokhin
Abstract: The information bottleneck offers a candidate general framework for learning, defined by two axes: how much a representation preserves about the external world and how much it carries about the relevant target. However, empirical evidence for information-plane dynamics in realistic-scale artificial networks and in biological systems remains limited. We address this by tracking information-plane trajectories during learning in three systems with different substrates, using a matched analytical pipeline. First, in a gradient-trained spiking ResNet-18 on CIFAR-10 --- a biologically plausible artificial substrate --- where a two-phase fitting-then-compression trajectory emerges specifically in the deep layers; the first information-plane analysis of a spiking network at this scale. Second, on the open-source dataset of mouse visual cortex during multi-week category learning, where our pipeline recovers the published cohort-level effect now reformulated in IB-plane terms. Third, in our own single-photon miniscope calcium imaging data of mouse hippocampus during multi-day place learning in open field, where the IB axes map onto the well-established egocentric-to-allocentric transformation: the input axis becomes egocentric sensorimotor state, the output axis becomes allocentric position. Behavioral and decoder-baseline controls confirm the hippocampal trajectory reflects experience-driven learning. Across all three systems, we observe coherent motion along the relevance axis with learning, while compression emerges only where task structure demands it --- most pronounced in the SNN. The information bottleneck plane therefore offers a substrate-independent coordinate for representation learning, even as the strength of compression along it remains task-dependent rather than universal.
Authors: Yifan F Zhang, Fangjun Hu, Guangkuo Liu, Mert Okyay, Xun Gao
Abstract: Diffusion models undergo a phase transition in a critical time window during generation dynamics, with two complementary diagnoses of criticality. The symmetry breaking picture views the critical window as when trajectories bifurcate into different semantic minima of the energy landscape, whereas the nonlocality picture views the critical window as when local denoising fails. We study whether two notions of such phase transitions are concurrent in modern diffusion transformers. By evaluating the dynamics and outcomes of the generation trajectory, we observe a near-simultaneous occurrence of the non-locality and symmetry breaking critical times. Our work is the first to unify the two notions of phase transitions in practice: it provides a concrete diagnostic for when and why diffusion models rely on conditioning and global denoising, enabling principled evaluation of model efficiency and guiding the design of architectures and sampling schemes that avoid unnecessary computation.
PaperID: 6555, Poster
Authors:
Andrea Ceron, Michael Schmidt, Alvaro Marcos-Ramiro, Sebastian Schmidt, Benjamin BusamAbstract: Latent LiDAR pipelines suffer from flying pixels: convolutional VAEs blur sharp LiDAR contours, yielding edge depths that back-project to points floating between surfaces. We identify this as a major, directly correctable decoder bottleneck and introduce CRISP: a pixel-space diffusion decoder with a backbone-agnostic latent adapter, DiT-based denoiser, and support mask predictor. CRISP replaces convolutional decoders commonly inherited from image and video VAEs while keeping the encoder and latent generator fixed. Across KITTI-360, SemanticKITTI, and nuScenes, replacing only the decoder reduces FSVD/FPVD by 50.5% on average across frozen backbones; for generic video VAEs, the reductions reach 71%/74%. On the LiDAR-native LiDM backbone, FRID drops by 71%, with the largest gains at depth discontinuities. Inside a pretrained LiDM world model, the same zero-shot replacement improves FSVD by 15.5%, narrowing the sim-to-real gap. Code and checkpoints will be released upon acceptance.
Abstract: Despite generating increasingly photorealistic images, text-to-image (T2I) models still exhibit localized, subtle, and structurally complex failures. Diagnosing these failures requires instance-level feedback that answers to overall image quality. While recent dense-feedback methods move beyond scalar supervision, their heatmap-centric representations still formulate diagnosis as pixel-field regression, making it difficult to localize variable-cardinality defects and bind semantic reasons to individual failures. To address this representation bottleneck, we propose , a 30K-image dataset with box-grounded annotations across four modern T2I generators, together with a dedicated evaluation protocol, . Building on this structured representation, we further present a diagnosis-to-alignment framework in which a Vision-Language Model (VLM) serves as the SDG detector, and converts predicted defect sets into box-derived, importance-weighted spatial rewards for diffusion model alignment. Extensive experiments show that our SDG detector outperforms leading proprietary VLMs on structured defect grounding, while SDG-guided rewards consistently improve T2I alignment and support localized image refinement. These results establish SDG as a unified, instance-level interface for diagnosing, evaluating, and enhancing modern generative models.
Abstract: Deep generative models have advanced rapidly across text and vision, motivating unified multimodal systems that can understand, reason over, and generate interleaved text–image sequences. Most existing approaches combine autoregressive language modeling with diffusion-based image generators, inheriting a structural mismatch between causal text generation and iterative visual denoising. We observe that autoregressive normalizing flows are autoregressive Transformers—sharing the same causal mask, KV-cache mechanism, and left-to-right structure as LLMs— making them the most natural paradigm for true unified multimodal generation. We present STARFlow2, built on the Pretzel architecture that vertically interleaves a pretrained VLM stream with a TarFlow stream via residual skip connections, both operating under the same causal mask. Combined with a deep-shallow flow design and a unified FAE latent space, STARFlow2 enables cache-friendly interleaved generation where both text and visual outputs directly enter the KV-cache without re-encoding. Experiments demonstrate strong performance across image generation and multimodal understanding benchmarks, validating autoregressive flows as a viable foundation for unified multimodal modeling.
Abstract: Tokenizer-free language models eliminate the tokenizer step of the language modeling pipeline by operating directly on bytes; patch-based variants further aggregate contiguous byte spans into patches for efficiency. However, the average patch size chosen at the model design stage governs a tight trade-off: larger patches reduce compute and KV-cache footprint, but degrade modeling quality. We trace this trade-off to patch lag: until a patch is fully observed, byte predictions within it must rely on a stale representation from the previous patch to preserve causality; this lag widens as patches grow larger. We introduce Scratchpad Patching (SP), which inserts transient scratchpads inside each patch to aggregate the bytes seen so far and refresh patch-level context for subsequent predictions. SP triggers scratchpads using next-byte prediction entropy, selectively allocating compute to information-dense regions and enabling post-hoc adjustment of inference-time compute. Across experiments on natural language and code, SP improves model quality at the same patch size; for example, even at 16 bytes per patch, SP-augmented models match or closely approach the byte-level baseline on downstream evaluations while using a 16× smaller KV cache over patches and 3–4× less inference compute.
PaperID: 6559, Poster
Abstract: Data attribution, which traces a model’s prediction back to specific training data, is an important tool for interpreting sophisticated AI models. The widely used TRAK algorithm addresses this challenge by first approximating the underlying model with a kernel machine and then leveraging techniques developed for approximating the leave-one-out (ALO) risk. Despite its strong empirical performance, the theoretical conditions under which TRAK approximations are accurate, as well as the regimes in which they break down, remain largely unexplored. In this paper, we provide a theoretical analysis of the TRAK algorithm, characterizing its performance and quantifying the errors introduced by the approximations on which the method relies. We show that although the approximations can incur significant errors, TRAK preserves the separation between highly influential and weakly influential data points. We corroborate our theoretical results through extensive simulations and empirical studies.
PaperID: 6560, Poster
Abstract: Spiking Neural Networks (SNNs) offer a distinctive paradigm for temporal modeling through their intrinsic membrane potential dynamics. However, existing SNN-based forecasting methods focus exclusively on temporal processing, lacking mechanisms to capture spatial dependencies among variables. To bridge this gap, we propose SpikeSTAG, a neuromorphic spatiotemporal architecture that integrates graph-based spatial reasoning into spike-driven temporal computation. Central to our approach is the Dendritic Graph Module, which reinterprets spectral graph convolutions as biological dendritic integration -- coupling multi-scale spatial aggregation with the temporal dynamics of spiking neurons within a single computational unit. To adaptively fuse the resulting spatial and temporal representations, we further introduce a Dual-Stream Integration Module that employs competitive gating and coincidence-based amplification, selectively enhancing consistent spatiotemporal patterns while suppressing conflicting noise. Extensive experiments demonstrate that SpikeSTAG establishes a new state-of-the-art among SNN-based models and surpasses representative ANN baselines, while reducing theoretical energy consumption by approximately 33.3% compared to Transformer architectures.
PaperID: 6561, Poster
Authors:
Jingxuan Yu, Ju Jia, Yiqian Chen, Yuchong Chen, Cong Wu, Di Wu, Siqi Ma, Jie GuiAbstract: Large language model (LLM)-driven agents are proposed for a wide range of applications. As scenarios become increasingly composite, a graph-structured multi-agent system (MAS) offers a promising solution due to the orchestration of various skills. To architect a reasonable underlying workflow, recent advances mostly investigate two structures: sequential topology and decentralized topology. The former constrains the structure into a directed acyclic graph for multi-step executions but lacks sufficient scalability, while the latter leverages node-wise subgraphs for knowledge aggregation with synchronous parallelization but lacks long-term planning. Therefore, our key insight is that their strengths are mutually complementary. Driven by this motivation, we introduce step-aware hybrid topology for a MAS, where sequential connections support inter-step ordered planning and decentralized connections promote intra-step perspective integration. Concretely, to effectively coordinate different structural components, we clarify the corresponding topological definition and the protocols for step-wise serialization and parallelization. Subsequently, to dynamically refine desired structures, we introduce the parameterized distribution of a hybrid topology, which is initialized by LLM-driven graph planning. Eventually, to further reduce token consumption and time complexity, we supplement the topological density and length normalization for end-to-end learning. Extensive experiments demonstrate that the hybrid topology outperforms single-form topology in 90.4% of tasks, our agentic initialization improves learning in 88.8% of scenarios, proposed normalizations reduce token consumption by 20%~30% and time consumption by 7%~13% with a slight utility fluctuation. The code is available at https://anonymous.4open.science/r/HTAS-D107.
PaperID: 6562, Poster
Abstract: This paper studies the problem of semi-supervised edge classification, which aims to identify edge relations using both labeled and unlabeled data. This problem is highly challenging due to the inherent asymmetry of edge relations and strong prediction biases from imbalanced homophilous and heterophilous edges. Towards this end, this paper proposes a novel approach named Complementary Tripartite Play with Bayesian CaLibration (CLEAR) for semi-supervised edge classification. The core of our CLEAR is to incorporate asymmetric semantic branches and a meta-teacher into a tripartite-play framework, enabling reliable guidance and optimization under label scarcity. In particular, our CLEAR introduces a topological branch and a contextual branch which extracts graph semantics, generating pseudo-labels in complementary views. To improve the reliability of pseudo-labels, we filter low-quality pseudo-labels using epistemic uncertainty, and then calibrate posterior distributions using the priors induced by neighborhood information. More importantly, we utilize a meta-consoler to guide the optimization of two branches by outputting the objective coefficients, ensuring a reliable pseudo-labeling process in a tripartite play framework. Extensive experiments on benchmark datasets validate the superiority of the proposed CLEAR in comparison to various baselines. Our code is available at https://anonymous.4open.science/r/CLEAR-D1D7/.
PaperID: 6563, Poster
Abstract: We propose the Grid Efficient Transport Optimization Linear Programming (GETO-LP) model, a novel formulation of the Fixed-Support Wasserstein barycenter problem for M probability distributions supported on d-dimensional grids. While standard LP models require O(MN^2) variables, the GETO-LP model requires only O(dMN^1+\frac1d), with N being the total number of points in the grid support. Despite the reduced number of variables, our formulation is \emphexact, meaning that its solution fully characterizes both the barycenter \nu and the optimal transport plans between \nu and each of the M input distributions without additional optimization steps. We validate the GETO-LP model via extensive comparisons with classical and entropic regularization solvers, demonstrating its effectiveness on both synthetic data and real-world applications such as color palette manipulation and medical data averaging. Overall, our approach provides an effective formulation for the Wasserstein barycenter problem, offering a memory-efficient alternative that naturally integrates with other existing computational frameworks.
PaperID: 6564, Poster
Abstract: Frame-wise action-controlled image-to-video generation is a promising paradigm for interactive world simulation, where each control signal should elicit an immediate visual response. However, maintaining visual fidelity and 3D consistency over long autoregressive rollouts remains challenging. Existing 3D-aware methods often suffer from catastrophic drift due to two impediments: information loss from Latent--RGB Cycling, where generated latents are repeatedly decoded to RGB and re-encoded for future conditioning, and the training--inference gap induced by the error-free hypothesis, where clean training memory fails to match prediction-corrupted inference memory. To address these challenges, we present Robust Dreamer, a memory-augmented framework built around how to design 3D memory and how to use it robustly. First, we introduce Latent Gaussian Memory, which anchors diffusion latents inherited from the generation process to Gaussian primitives and recalls them via latent-space Gaussian splatting. This provides dense, geometry-aware, view-aligned conditioning while avoiding accumulated degradation from repeated VAE conversion. Second, we propose Deviation Learning with Dynamic Deviation Archive, which synthesizes rollout-induced latent deviations through a one-step approximation, stores them by autoregressive stage and denoising timestamp, and injects them into historical memory during training. This exposes the generator to realistic corrupted memory states and teaches internal correction before inference. Experiments on ScanNet, DL3DV, and OmniWorldGame demonstrate state-of-the-art long-horizon performance.
Abstract: This paper aims at analyzing the regularization effect that data augmentation induces on supervised regression methods in the proportional regime, where the number of covariates grows proportionally to the number of samples. We provide a tight characterization of the test error, measured in mean squared error, in terms only of the population quantities of the true data, as well as first and second order statistics of the augmentation scheme. Our results are valid under misspecified feature maps, and for any network architecture where only the last readout layer is trained, and the rest of the network is either frozen or randomly initialized. We specify our results in the case of Gaussian data, and show that our asymptotic characterization is tight in this setting.
Authors:
Morunliu Yang, Ruotao Xu, Le Li, Yue Wang, Jianxin Zhang, Siwei Feng, Peifeng Li, Yihang Lou, Juntao LiAbstract: Omnimodal large language models (OmniLLMs) have recently gained increasing attention for unified audio-video understanding. However, processing long multimodal token sequences introduces substantial computational overhead, making efficient token compression crucial. Existing methods typically rely on fixed, modality-specific guidance, which fails to account for the varying importance of modalities across different queries. To address this limitation, we propose OmniSelect, a training-free, modality-adaptive token pruning framework that dynamically selects appropriate compression strategies for multimodal inputs. Specifically, we leverage a lightweight AudioCLIP model to estimate cross-modal relevance and categorize each input into three pruning regimes: Audio-Centric, Video-Centric, and Uniform pruning. Based on these relevance scores, OmniSelect further performs fine-grained token pruning within each temporal group, adaptively allocating pruning ratios to preserve informative tokens across modalities. By explicitly modeling modality preference and enabling dynamic strategy selection, OmniSelect effectively avoids the pitfalls of one-size-fits-all compression. Extensive experiments demonstrate that our method achieves efficient multimodal token reduction while maintaining strong performance, without requiring any additional training.
PaperID: 6567, Poster
Abstract: Current skeleton representation learning paradigms face distinct limitations: Contrastive Learning (CL) often overlooks fine-grained motion details, while Masked Auto-Encoders (MAE) rely on coordinate-level reconstruction. This reconstruction inherently demands dense token sequences and heavy decoders, wasting pre-training computation on discarded components and forcing downstream inference to process dense token grids. To resolve these bottlenecks, we propose SLiM (Skeleton Less is More), a compact-token framework that unifies masked feature prediction and contrastive learning via a shared encoder. By shifting the objective from raw coordinate reconstruction to decoder-free, teacher-guided feature prediction, SLiM breaks the reliance on dense tokenization and enables effective learning with a highly compact token grid. Crucially, to prevent trivial shortcut learning arising from strong inter-joint dependencies of human, we introduce Semantic Tube Masking together with Skeleton-Aware Augmentations to enforce deep skeletal-temporal reasoning and anatomical consistency. Extensive experiments across multiple downstream protocols demonstrate that SLiM achieves state-of-the-art performance while structurally reducing inference computation by 7.89× compared to dense-token MAE baselines.
PaperID: 6568, Poster
Authors: Hugo Richard
Abstract: We study regret minimization for time-homogeneous episodic tabular reinforcement learning with S states and A actions under non-interactive \epsilon-local differential privacy (LDP). In each episode k \in [K], the decision-maker sends a policy to a user, the user executes the policy for H steps, and sends an \epsilon-locally differentially private message back to the decision-maker. When H=1, the problem reduces to an LDP contextual bandit with S contexts, A arms and horizon K. We prove that the minimax regret is \min\bigg(K, \frac \sqrt\max( S, \exp(\epsilon) + 1) S A K (\exp(\epsilon) + 1)\exp(\epsilon) - 1\bigg), up to constants and a \sqrt\log(A) factor. The upper bound is achieved by an adaptation of EXP3 and a unary encoding mechanism. When H > 1, and for small \epsilon, we show a lower bound in \Omega(HS\frac\sqrtA K\epsilon) and an upper bound in \tildeO(H^2 S^2 \frac\sqrt AK\epsilon). The upper bound is achieved by an algorithm based on UCBVI and a unary encoding mechanism. For small \epsilon, these bounds improve the previously known lower bound by a factor \sqrtS and the previously known upper bound by a factor of H \sqrtA.
PaperID: 6569, Poster
Abstract: Transformer attention exhibits attention sinks, where a small subset of source tokens absorbs disproportionate attention mass across layers. Recent work shows that post-SDPA output gating strongly reduces attention-sink behavior, while separate studies link attention sinks to gradient sinks and massive activations through backward training dynamics. Yet how the same gate affects forward sink structure, backward gradient concentration, and training-time representation signatures remains unclear. We address this gap by formulating attention sinks as sink-induced low-rank modes in token-space covariance, with directions aligned to high-mass attention columns. Under this view, post-SDPA gating acts as a mode-dependent contraction: it attenuates sink-induced spectral energy in the forward pass and locally reduces value-path gradients routed through sink columns in the backward pass. Empirically, random-matrix diagnostics show that gating reduces outlier mass, weakens sink-subspace alignment, and increases effective rank, while backward analyses show active-gate attenuation of first-token value gradients. Training-time diagnostics further show that gating suppresses the co-emergence of attention concentration, value-gradient amplification, and representation dominance.
PaperID: 6570, Poster
Authors:
Michał Brzozowski, Zuzanna Dubanowska, Enrico Cassano, Neo Christopher ChungAbstract: Narrowly finetuned language models memorize implanted content verbatim, but auditing what a deployed model has been taught, without access to its weights or training data, remains an open challenge. Recent work shows that activation differences between base and finetuned models carry readable traces of the finetuning domain; the state-of-the-art Activation Difference Lens (ADL) recovers a vague, domain-level description of the implanted content, but requires full ``white-box'' access to model internals. We introduce Contrastive Decoding Diffing (CDD), a model diffing method that operates on output-level logit distributions only, with no weight access, no layer selection, and no per-model tuning, yet recovers implanted facts. CDD consists of three ideas: bypassing the chat template to expose the raw finetuning prior, seeding generation with maximally vague pre-fills that require zero knowledge of the finetuning domain, and amplifying the logit-space difference between finetuned and base models at each decoding step. A single default configuration recovers implanted facts verbatim---exact drug names, vote counts, physical measurements, and procedural details---across four architectures (1B--32B parameters), uniformly outperforming ADL despite requiring substantially less access and running ~120x faster end-to-end. Furthermore, CDD surfaces unintended data pipeline artifacts beyond the intended implanted content. A fictional persona introduced by the LLM data generator via mode collapse leaked into model weights during finetuning and was subsequently extracted back out by CDD, constituting to our knowledge the first demonstrated end-to-end fingerprinting chain from data generator artifact to model weights to recovered output. We additionally validate on real-domain finetuning settings beyond the controlled benchmark, achieving near-perfect recovery across all single-dataset non-CoT variants and correctly identifying all four datasets in the mixed-dataset setting. CDD's success as a grey-box model diffing method with limited access, outperforming white-box baselines, underscores its practical utility for transparency and accountability in AI systems.
PaperID: 6571, Poster
Abstract: LLM-based agents increasingly rely on external tools to accomplish complex tasks, yet the security of tool-calling pipelines remains poorly understood at the chain level. Prior attacks target individual tool invocations through prompt injection or metadata manipulation, but compromising a single step in a multi-step workflow is conspicuous and rarely sufficient for complex adversarial objectives. In this work, we uncover a more insidious yet realistic attack surface, tool-chain hijacking, in which an adversary constructs a coherent sequence of tools that collectively replace the agent's intended execution trace while still completing the user's task correctly, rendering the hijack invisible to both the agent and the user. To operationalize this threat, we propose ChainForge, an execution-grounded framework that mines agent execution logs to synthesize replacement chains through rollout-based optimization, then embeds adversarial payloads into the chain's code via multi-criteria iterative refinement. To systematically evaluate chain-level threats, we further construct ChainBench, a benchmark of 97 tasks across 4 domains. Experiments on four frontier LLMs show that ChainForge achieves a trace hijack rate of up to 98.54%, maintains task utility above 80.41%, transfers across models with over 84.95% success, and evades all evaluated defenses and code-safety scanners at substantially higher rates than single-tool baselines, exposing a critical blind spot in current agent security.
PaperID: 6572, Poster
Abstract: Reinforcement learning (RL) post-training has become an important paradigm for improving multimodal large language models (MLLMs), especially with automatically verifiable rewards. For vision-centric MLLMs, self-supervised visual pretext tasks provide such rewards without human annotations. However, existing self-supervised RL methods mostly exploit within-instance structure, such as spatial or temporal ordering within a single image or video, offering limited supervision for cross-instance discrimination and fine-grained comparison. We introduce Group-and-Order Self-Supervised Reinforcement Learning (GO-SSL), a cross-instance jigsaw framework for vision-centric MLLM post-training. Given two visual instances, GO-SSL mixes their local elements and trains the model to group them by instance of origin while recovering the order within each group, coupling instance-level comparison with spatial, temporal, or geometric structure recovery. This paired formulation turns the same unlabeled data into richer verifiable supervision through diverse pairings, shuffles, and hard-pair curricula. We instantiate GO-SSL mainly on cross-image jigsaw tasks and further extend it to 3D depth and video temporal ordering. Across diverse benchmarks, GO-SSL consistently improves over the base MLLM and prior baseline methods, with clear gains in fine-grained perception and spatial understanding. These results suggest a data-efficient direction for self-supervised RL in MLLMs: constructing comparative cross-instance contexts can provide richer transferable supervision than merely scaling data volume or adding isolated pretext tasks. Code will be released later.
PaperID: 6573, Poster
Authors: Keisuke Sano, Takayoshi Yamashita, Hironobu Fujiyoshi
Abstract: In fine-grained image recognition, semantic parts (e.g., wings, head) recur across classes while discriminative cues lie in within-part features (e.g., yellow wings, striped wings). The prototypical part network (ProtoPNet) is an interpretable model that provides case-based explanations for image recognition by measuring similarities between local patches in an images and their closest prototypes representing patches in a training set. However, existing ProtoPNet variants do not explicitly model this two-level structure, causing prototypes for the same anatomical part to fragment across classes. We propose a hierarchical prototype learning framework that separates shared prototypes capturing cross-class part concepts from class-specific prototypes capturing intra-part features. To model this hierarchy geometrically, prototypes are embedded in hyperbolic space and an entailment loss constrains each class-specific prototype to lie within the cone of its corresponding shared prototype. To prevent degenerate trivial solutions such as redundant prototypes, we introduce a patch-level contrastive learning using pseudo patch IDs from a vision foundation model, used only at training time, leaving the backbone choice free at inference. Unlike prior hyperbolic prototype methods where prototype hierarchies are uncontrolled, our framework defines explicit shared part to intra-part feature correspondences. Across four ProtoPNet variants and four backbones on CUB-200-2011 and Stanford Cars, our method improves classification accuracy by up to 10 points, raises hierarchical part hierarchy IoU from 11% to 75%, and consistently improves four interpretability metrics on CUB-200-2011.
PaperID: 6574, Poster
Authors:
Yuguang Li, Yichuan Deng, Zixuan Liu, Ivaylo Boyadzhiev, Xiuyuan Wu, Xinyu Zou, Linda Shapiro, Alex ColburnAbstract: Reconstructing accurate camera poses and floor-plan layouts from sparse, unposed, wide-baseline RGB panoramas remains a challenging open problem. Direct multi-room floor-plan prediction is especially difficult in this setting: pose, geometry, semantics, and global layout structure are tightly coupled, yet a complete floor-plan is hard to learn as a single global output. Our key insight is to keep the model's predictions strictly local. Multi-view Ordered Wall Instance Segmentation (MOIS) is designed around this concept. Each panorama column predicts only the wall it observes plus the next two adjacent walls in clockwise order. The global wall partition, room layouts, and shared coordinates emerge from deterministic aggregation of overlapping local predictions. Building on MOIS, we propose Multi-view floor-plan Feed-Forward Network (MvFFN), which jointly predicts coarse camera poses, dense global geometry, and per-wall semantic outputs. A downstream Layout Chaining (LC) post processing pipeline then assembles these predictions into multi-room floor-plans with shared global coordinates, consistent wall identities, and connected room structure. The dense MOIS predictions recover cross-view wall identity and local within-room connectivity more accurately than composed modular baselines, and the assembled layouts substantially outperform prior methods on wall and junction level accuracy. We also introduce the ZInD-CrossView dataset, which augments ZInD with globally unique wall instances and cross-view segmentation labels for this task.
Abstract: Existing approaches to LLM personalization focus on constructing better personalized models or inputs, while treating inference as a single-shot process. In this work, we study Test-Time Personalization (TTP) along an unexplored axis: scaling inference-time computation by sampling N candidates from a personalized policy model and selecting the best with a personalized reward model. However, standard reward models fail to realize this potential. To diagnose why, we derive a unified scaling law that decomposes any reward model's Best-of-N curve into four measurable quantities and reveals two failure modes, \emphuser-level collapse (near-constant prediction for some users) and \emphquery-level reward hacking (negative correlation with true quality for some queries). Guided by this law, we propose a probabilistic personalized reward model whose learned variance effectively mitigates both failure modes. Experiments confirm both elements of our framework: TTP delivers consistent scaling across multiple policy models and personalized text generation tasks, and our scaling law closely matches observed scaling curves across reward-model variants.
PaperID: 6576, Poster
Authors: Reo Iizuka, shioya hiroaki, Naoto YOKOYA
Abstract: Generating neural network weights in a single forward pass promises to amortize model training for fast sampling, transfer, and model-set construction. Yet accuracy alone is insufficient: existing generators can map diverse latent inputs to weights that remain functionally close to their training checkpoints. We study this functional mode collapse in the one-step regime and introduce DriftWeight, a MeanFlow-based generator with temperature-scaled repulsive drift in a checkpoint-augmented PCA proxy space. The method uses structured negatives from the current batch and optimization history to push generated weights away from observed clusters while preserving task performance. Our key finding is that repulsion is not automatically compatible with one-step inference. In a controlled ablation, applying the drift loss at the same boundary evaluation used for sampling collapses one-step accuracy even though multi-step integration remains accurate. DriftWeight mitigates this boundary-localized failure with a consistency objective that anchors the inference point to the training-weight manifold. Across MNIST, Fashion-MNIST, and CIFAR-10, DriftWeight matches or exceeds reported 100-step DeepWeightFlow accuracy in a single forward pass. On MNIST, where we directly evaluate error-set functional overlap, it reduces MaxIoU from the DeepWeightFlow regime of approximately 0.82 to 0.66. On CIFAR-100 with a ViT-Base target (~86M parameters), one-step samples remain within 1.7 percentage points of the training-population mean.
Abstract: Multimodal large language models must adapt to evolving tasks and domains, yet continual improvement under bounded deployment footprint remains difficult because repeated parameter updates or growing replay stores can accumulate adaptation state over time. We study fixed-footprint continual adaptation: the deployed adaptation state is kept under a fixed memory budget, while the backbone model is left unchanged and task-specific updates are externalized. We propose , a retrieval-based method that stores each selected training prefix as an attention-ready memory entry, consisting of a frozen retrieval key and compact layerwise key--value (KV) payloads that can be appended to the model's self-attention cache. Under a strict memory budget, constructs a compact inducing set through bilevel selection: a lightweight calibration is fit for retrieval, while the selected memory balances current-task likelihood, anchor-based retention, and coverage in the frozen retrieval space. Across task-incremental instruction tuning, continual VQA, domain-incremental adaptation, and lifelong multimodal instruction tuning, consistently improves over PEFT, MoE, replay, and prompt-retrieval baselines under matched memory budgets. We further report backbone-matched, stage-1 CoIN, compute-matched, and scalability diagnostics, showing that the gains are not due to a stronger backbone, replay alone, or an unbounded candidate pool.
PaperID: 6578, Poster
Abstract: Consistency models achieve fast image synthesis by chaining a small number of neural evaluations along a diffusion trajectory, yet the distribution error of N-step sampling for N \geq 3 has lacked a rigorous characterisation. We establish a tight Wasserstein-p error bound for arbitrary N: the total error equals a Lipschitz-weighted sum of per-step local consistency errors, \sum_k=1^N L_x^N-k \delta_k, making the propagation mechanism explicit. From this bound we derive a closed-form optimal step count N^\star whose structure bifurcates on whether the spatial Lipschitz constant L_x exceeds one: for L_x > 1, exponential error growth yields N^\star \approx (\ln L_x)^-1 \log(\beta\gamma(L_x-1)/(c \ln L_x)); for L_x < 1, geometric-series saturation gives N^\star \approx (\beta\gamma(1-L_x)/(c \, |\ln L_x|))^1/(\gamma+1). The two-regime structure accounts for ECT's best FID of 2.73 at N=2: with L_x \approx 0.85 < 1 (Case 2), the formula gives N^\star \approx 1.7. Experiments on CIFAR-10 reproduce the predicted U-shaped FID curve with N^\star = 8 on a capacity-constrained model in which error accumulation is easier to see.
Authors: Bryan Zhu, Ziang Chen
Abstract: Empirical risk minimization (ERM) can be computationally expensive, with standard solvers scaling poorly even in the convex setting. We propose a novel lossless compression framework for convex ERM based on color refinement, extending prior work from linear programs and convex quadratic programs to general differentiable convex optimization problems. We develop concrete algorithms for a range of models, including linear and polynomial regression, binary and multiclass logistic regression, regression with elastic-net regularization, and kernel methods such as kernel ridge regression and kernel logistic regression. Numerical experiments on representative datasets demonstrate the effectiveness of the proposed approach.
PaperID: 6580, Poster
Authors: Lingfei Ma, Bin Liu, Wen Li, Wentao Sun, Chaorui Liu, Haiyan Guan, Jonathan Jun LI
Abstract: 3D Referring Expression Segmentation (3D-RES) aims to segment target objects in point clouds via language. However, extending 3D-RES from indoor scenes to large-scale outdoor environments introduces three major challenges, i.e., large-scale geospatial complexity lacking in indoor benchmarks, sparse-target perception overwhelmed by massive geometric clutter and distractors, and context-induced query dilution, where highly relation-intensive expressions cause globally-initialized queries to be absorbed by salient reference objects. To address these issues, we introduce UrbanRefer, the first outdoor 3D-GRES benchmark with 195 urban road point cloud scenes and 7,840 textual descriptions for complex geospatial reasoning. We further propose LaSA-Net, a Language-guided Semantic Segmentation Network. First, to enhance sparse-target perception, LaSA-Net incorporates pixel-level dense features from DINOv3 as 2D priors and fuses them with 3D features through depth-consistent verification. Then, to alleviate query dilution, we design a semantic-adaptive query generation module that applies text-weighted farthest point sampling to focus on text-relevant regions, and performs multi-granularity sampling to maintain both local grounding and global semantic consistency. Finally, a reliability-aware decoder adaptively suppresses noisy language cues through cross-modal reliability-gated injection, generating precise prompts for accurate mask prediction. Extensive experiments demonstrate that LaSA-Net achieves 52.7%, 55.9%, and 47.9% mIoU on the ScanRefer, Multi3DRefer, and UrbanRefer datasets, surpassing prior state-of-the-art methods by 2.3%, 4.2%, and 2.8% respectively.
PaperID: 6581, Poster
Authors: Sheng Zhao, Weikai Lin, Yuhao Zhu
Abstract: Judging image quality is not only ecologically relevant to everyday humans tasks, but also underpins many machine vision tasks such as image generation. This paper proposes a framework to understand the inherent perceptual space underlying image quality judgment in humans. We propose a multi-dimensional observer model that represents images as distributions in a latent perceptual space and that models human judgment as comparing noisy samples. Being constrained by neural representations in the primate ventral stream and fit to large-scale behavioral data, the model enables analysis of perceptual structure while matching the predictive power of existing metrics. Using this model, we find that the perceptual spaces needed to account for image quality judgment in humans are extremely low-dimensional compared to the image space even when considering its sparsity. The exact structure of the space (e.g., dimensionalities, information encoded) varies between low-level and high-level quality judgments, suggesting that, despite a shared retinal encoding in the beginning, humans selectively construct task-dependent perceptual spaces in visual decision making.
PaperID: 6582, Poster
Abstract: Large language models (LLMs) perform reasoning via lengthy token-by-token generation, incurring substantial inference cost. While recent methods compress this process by enabling LLMs to reason in a latent space, they still rely on next-vector generation for sequential logic and on flat representations. We therefore present Latent Concept-Pyramid Modeling (\emphLCP), a new paradigm that reformulates LLM reasoning as hierarchical next-level concept generation in a coarse-to-fine manner. Specifically, \emphLCP generates the pyramid level by level: a single apex concept encodes the most abstract semantics, and each subsequent level refines its predecessor into finer-grained concepts that together capture distinct CoT segments at one granularity. Such next-level concept prediction approximates the reasoning trace in a manner analogous to advancing from a skeletal outline, through broad structural forms, down to local details. Experimental evaluations on zero-shot mathematical and coding reasoning benchmarks demonstrate that Qwen2.5 and Qwen3 models fine-tuned with \emphLCP achieve significant reductions in both token cost and inference latency while maintaining competitive accuracy. \emphLCP further showcases zero-shot generalization ability across different tasks. Moreover, the pyramid's hierarchical, multi-grained representational capacity enables faithful reconstruction of the original CoT, endowing \emphLCP with interpretability.
PaperID: 6583, Poster
Abstract: Punctuation Restoration and Inverse Text Normalization (ITN) are essential post-processing tasks for ASR systems. Traditionally, these are handled by separate models in a sequential pipeline, which suffers from error propagation and suboptimal latency. We propose UniForm, a unified framework that jointly solves both tasks using a single compact language model. UniForm is trained via a three-stage progressive paradigm: large-scale supervised fine-tuning on punctuation data, joint fine-tuning with ITN span masks, and reinforcement learning alignment. To enable effective reinforcement learning in this inherently low-entropy setting, we introduce Segment-Aware GRPO (SA-GRPO), which optimizes model performance through two synergistic components: a structured reward decomposition mechanism that provides dedicated signals for distinct segment types, and a token-level adaptive credit assignment mechanism that dynamically redistributes gradients toward difficult tokens based on sampling consensus and model entropy. This integrated approach effectively eliminates the need for manual sub-task weighting and ensures fine-grained optimization for both tasks. We train UniForm-0.8B on over 238 million multilingual samples. Extensive evaluations across diverse languages, domains, and code-switching scenarios demonstrate that UniForm-0.8B outperforms prior specialized models and larger general-purpose LMs on both tasks while maintaining real-time inference latency. Furthermore, our analysis reveals a strong synergistic effect, proving that joint training mutually benefits both punctuation and ITN performance.
PaperID: 6584, Poster
Authors: Yingyi Cheng, Xueqiang Han, Chen Yu
Abstract: With long-term memory, LLM-based agents have demonstrated remarkable capabilities in handling complex long-horizon tasks. However, existing memory frameworks primarily treat long-term memory as a post-admission management problem, emphasizing how experiences should be organized, updated, or controlled. This overlooks a more fundamental stage in the memory lifecycle, i.e., how transient experiences are progressively transformed into stable and task-transferable knowledge. Biological memory consolidation provides a temporal principle for memory formation, whereby long-term memory emerges through the delayed stabilization and selective transformation of recent experience, rather than through its verbatim preservation. Motivated by this, we propose NeuroMem, a neuroplastic memory framework for lifelong agents through delayed consolidation. NeuroMem formulates the formation of task-transferable knowledge in long-term memory as a delayed-reward reinforcement learning problem. Specifically, newly formed memory candidates are first maintained in a transient buffer, where a consolidation policy learns to promote, merge, retain, or discard them only after their delayed utility rewards become observable through subsequent interactions. Experiments on EHRSQL and LoCoMo show NeuroMem achieves stronger performance with a smaller memory set, improving EHRSQL task success by up to 8.7% and LoCoMo F1/BLEU-4 by up to +5.20%/+0.97% over baselines.
PaperID: 6585, Poster
Abstract: In this paper, we study distributionally robust reinforcement learning with Kullback--Leibler (KL) divergence defined uncertainty set. The goal is to find a policy that maximizes the robust value function, defined as the worst-case value over all transition kernels in the uncertainty set. Assuming access to a generative model, we aim to understand the sample complexity of finding an \epsilon-optimal robust policy. The best-known sample complexity results in the literature show a non-trivial gap of \mathcalO\\max\p_\wedge^-1(1-\gamma)^-1,(1-\gamma)^-2\\ between the upper and lower bounds, where p_\wedge denotes minimal non-zero support of the nominal transition kernel, and \gamma is the discount factor. More importantly, existing results on the lower bound only cover a limited range of the uncertainty level. In this paper, we develop tighter and complete upper and lower bounds for robust RL with KL-defined uncertainty set. Our upper bound is the tightest among all existing studies, and it improves upon the best known bound by at least the order of \mathcalO(\min\p_\wedge^-1,(1-\gamma)^-1\). Furthermore, our lower bound holds for any uncertainty level. Our upper and lower bounds (nearly) match under various cases, providing near minimax optimality results for this problem.
PaperID: 6586, Poster
Authors: Hamza Etcibasi, Ramazan Gokberk Cinbis
Abstract: A widely held assumption in knowledge distillation is that teacher and student representations must occupy a compatible or explicitly aligned feature space. We propose RPKD (Random Prototype Knowledge Distillation), which discards this assumption by projecting both networks into a shared set of randomly initialized, frozen prototype vectors, requiring no architecture-specific adapters, no class-specific alignment, and minimal architectural assumptions. Two branches handle the transfer: a logit branch aligning global prototype similarity distributions and a feature branch matching spatially-resolved prototype activation maps. Both operate within the same fixed prototype space, which can be interpreted as a random feature embedding that approximately preserves inter-sample similarity structure, keeping RPKD agnostic to the internal dimensions and inductive biases of either network. In low-category settings, decoupling the prototype vocabulary from the task label space yields richer supervision than the label space alone, an advantage absent from logit-based methods, whose supervisory signal collapses as class count shrinks. This advantage is especially relevant in real-world scenarios where the label space is typically constrained. Extensive experiments demonstrate the effectiveness of RPKD, achieving a maximum gain of \mathbf+8.26% over OFA on CIFAR-100, \mathbf+12.46% on ImageNet-100, and \mathbf+1.34% on ImageNet-1K. Source code will be released upon acceptance. Source code will be released upon acceptance.
Authors:
Kemal Oksuz, Alexandru Buburuzan, Yuhan Yao, Puneet DokaniaAbstract: State-of-the-art vision-language-action models (VLA) for autonomous driving face critical limitations: excessive parameter counts, inefficient high-resolution image processing, and lack of temporal memory. We introduce FIVE-VLA (Fast and EffectIVE VLA) to address these through two key contributions. First, we employ an efficient vision encoder that processes high-resolution (448 × 896) images while generating only 98 tokens, over 5× fewer than existing approaches, and bypass text generation entirely for single-pass trajectory prediction. Second, we propose Recurrent Action Memory (RAM), a lightweight module that conditions action prediction on previous action tokens, providing temporal context critical for manoeuvres such as overtaking and emergency braking. With only 641M parameters, FIVE-VLA completes ~10% more routes without traffic rule infractions than the previous state-of-the-art VLA on the challenging Bench2Drive closed-loop driving benchmark. Furthermore, FIVE-VLA runs at ~30 fps on an A100 and ~4 fps on a T4 GPU (proxy to an edge device), representing an 8--30× speedup over previous methods. The code will be made public upon acceptance.
Abstract: Transformers perform inference by iteratively transforming token representations across layers. This layerwise computation has been studied empirically, and recent mean-field theories of Transformer dynamics explain how attention can drive token distributions toward clustering. However, existing mean-field analyses largely treat model parameters as prescribed, leaving open how training reshapes this clustering picture. We study this question in a noisy mean-field Transformer in which only a parameter-linear FFN is trained under L^2 regularization. We find and analyze a training-induced phase in the dynamics: after initially following attention-driven clustering, the token distribution can leave the clustered regime near the final layers. Our mathematical analysis is based on an entropy-regularized interaction energy that captures the clustering bias of attention. More broadly, our results point toward a training-aware mean-field theory of Transformer dynamics, in which training and inference dynamics are treated together.
PaperID: 6589, Poster
Abstract: Current multivariate time series forecasting methods mainly rely on static linear decoders, but these often suffer from severe representational bottlenecks. In this paper, we propose a novel architecture called DecodeTS (\underlineDynamic \underlineExpert-\underlineCoupled \underlineOptimal \underlineDEcoder for Time Series Forecasting), which replaces the conventional static prediction head with a heterogeneous expert library. DecoderTS adopts a divide-and-conquer strategy to disentangle complex temporal dynamics, such as long-term trends and abrupt changes. Crucially, we introduce an optimal transport (OT)-based dynamic routing mechanism that can adaptively assign customized combinations of experts to different variables. By imposing OT marginal constraints, DecoderTS is theoretically proven to achieve collaborative load balancing among experts and effectively eliminate representation collapse. Extensive experiments on more than 15 datasets and about 30 baselines demonstrate that DecoderTS achieves state-of-the-art forecasting accuracy, which delivers a 17.06% performance gain while attaining up to about 41× inference time speedup and up to 3.61× reduction in memory footprint.
Abstract: Oversmoothing is a well-known failure mode of Graph Neural Networks (GNNs). However, most existing diagnostics rely on global aggregation measures that fail to capture the heterogeneous dynamics of message passing. Real-world graphs exhibit pronounced community structure, and message passing operates on two timescales, where representations collapse rapidly within and slowly across communities. This creates a critical gap where intra-community representations can already be indistinguishable while inter-community separation persists, which is a failure mode we refer to as Echo Chamber Effect. To quantify this, we propose to use the Echo Chamber Index (ECI), which stratifies pairwise distances by community membership and formally establishes that global energy diminishes while inter-community separation remains less impacted. The estimated ECI for a graph reveals a surprising failure of common oversmoothing remedies, where rather than escaping the echo chamber, these architectures become permanently trapped in it. The consequences depend on label structure, as when communities align with classes, the echo chamber sharpens node classification, and when they do not, the same collapse makes it provably harder. To address this, we propose Community-Aware Split Propagation (CASP), a lightweight model-agnostic plugin that decouples intra and inter-community aggregation and learns their balance from label structure. CASP consistently improves over diverse backbone GNNs across multiple benchmarks spanning homophilic and heterophilic settings. Our source code is available at: https://anonymous.4open.science/r/CASP-3059/.
PaperID: 6591, Poster
Abstract: Contrastive vision-language models such as CLIP are trained to align images and text, not to represent taxonomies. Yet empirical studies show that CLIP embeddings exhibit a hierarchical geometry aligned with the semantic hierarchy of concepts. A theoretical characterization of why such hierarchy emerges from a non-hierarchical objective remains open. In this paper, we formalize this phenomenon as Hierarchical Angular Separation (HAS), a geometric property whereby each non-leaf concept embedding is more similar to its descendants than to off-branch concepts. Empirically, we show that HAS already holds weakly in randomly initialized CLIP image embeddings, suggesting that hierarchical structure may be induced by the raw image geometry, while in text embeddings it emerges only after contrastive training. To explain this asymmetry between text and image embeddings, we analyze an unconstrained features model in which the image embeddings are fixed and only the text embeddings are trained. We show that when the image embeddings satisfy HAS with a positive margin, the global minimizer of the contrastive objective induces HAS in the text embeddings, thereby transferring hierarchical structure across modalities. The key mechanism is the hierarchical co-occurrence structure of image-text pairs, which biases text embeddings toward root-centered image prototypes. For a linear contrastive loss, we derive a closed-form solution that makes this mechanism explicit. We extend the analysis to the InfoNCE loss and establish analogous guarantees under sufficient regularization. Synthetic experiments validate the theory and show that HAS emerges even beyond the regimes covered by our analysis.
PaperID: 6592, Poster
Abstract: We identify a topological obstruction to the standard stable-attractor template for learning local Nash equilibria, also known as first-order Nash equilibria (FONE), in constrained non-concave differentiable games. We consider continuous deterministic learning dynamics that leave each FONE point fixed and aim to attract every initial condition to the full FONE set. Our main message is that this familiar convergence template can fail for a topological reason. Targets with holes, such as loops, cannot be stable global attractors for such dynamics on contractible state spaces. We then exhibit this obstruction in a simple non-concave quadratic zero-sum game on a box. Its exact FONE set is a diagonal six-edge loop together with an isolated point, so the loop creates the required topological obstruction. As a result, no continuous pointwise-stationary learning rule covered by our framework can make this exact FONE set a stable global attractor. We show that degree two is the minimal polynomial degree at which this obstruction can arise. We further amplify the construction to arbitrary homological degree. Finally, by studying the topology in sufficiently small perturbed version of the game, we show that an open positive-volume family of quadratic games retain the same homological obstruction. Thus, even in quadratic zero-sum games, this phenomenon is not a pathology of a single construction; instead, the geometry of projected first-order stationarity can create an intrinsic barrier to stable global FONE dynamics.
PaperID: 6593, Poster
Abstract: As long-video understanding becomes increasingly important, augmented reality (AR) devices offer a natural platform for deploying egocentric video intelligence. Yet first-person videos often contain substantial temporal redundancy, creating heavy memory and compute demands for resource-limited AR hardware. We present~UTOPI, a plug-and-play pre-compression module tailored to egocentric long-video processing on AR devices. UTOPI first uses user motion cues to estimate cross-frame overlap and remove redundant content through a viewpoint-shift-robust pruning strategy. It then leverages eye-tracking signals to detect user fixation regions, separating attention-relevant foreground from background and preserving important content through a fine-tuned neural adjustment module. Evaluated across multiple benchmarks and model backbones, UTOPI reduces video tokens by up to 95% while maintaining task performance.
Abstract: Diffusion Large Language Models (dLLMs) have emerged as a compelling paradigm for language modeling, offering the unique potential for highly efficient inference via multi-token parallel decoding. However, unlocking this potential remains challenging. High-quality generation typically defaults to one-token-per-step static decoding, while existing parallel algorithms rely on heuristic confidence thresholds, introducing a fragile trade-off between decoding efficiency and generation quality. To overcome this, we introduce Free Draft-and-Verification (FreeDave), a novel training-free and model-free fast decoding algorithm. FreeDave capitalizes on the inherent properties of dLLMs by treating future token predictions as "free drafts" and rigorously self-verifying them in subsequent forward passes. Theoretically, FreeDave is guaranteed to use the fewest possible model forward passes to reproduce the exact sequence generated by static decoding, achieving algorithmically lossless acceleration. Extensive evaluations on math reasoning and code generation benchmarks demonstrate that FreeDave can boost the number of generated tokens per model forward (TPF) by up to 5.48× and tokens per second (TPS) by up to 3.94×, while robustly maintaining generation quality and avoiding the task-specific performance degradation faced by heuristic methods.
PaperID: 6595, Poster
Abstract: Movable agents, such as autonomous vehicles and robots, should be able to autonomously adapt their pre-trained object recognition models to new environments without human annotations. While existing test-time adaptation (TTA) methods address label-free model adaptation, they typically treat test data as independent snapshots, without constructively exploiting a rich, unique source of supervision: the significant variance in prediction quality as an agent observes the same object from different distances and viewpoints. To capitalize on this, we propose , a model-agnostic and hyperparameter-insensitive paradigm that leverages motion-induced prediction variance for unsupervised adaptation. Specifically, we first develop mobility-calibrated filtering of pseudo-labels, which constructs reliable pseudo-labels by fusing cross-view predictions using inverse-variance weights. The filtered pseudo-labels then guide the selection of "hard samples" for effective learning. To further suppress pseudo-label corruption, we introduce , an optional extension that uses a public dataset to assign weighting factors to data samples, yielding an estimator of the clean loss with bounded bias. Extensive experiments on autonomous driving (nuScenes, KITTI) and embodied-agent datasets show that
PaperID: 6596, Poster
Abstract: Extremely sparse-view computed tomography (xSV-CT) is critical when scanning time, motion, or radiation tolerance limits each scan to only a few projections. In many such scenarios, the same acquisition limits also constrain dataset construction, making pre-reconstructed 3D volumes from dense scans unavailable. This exposes a gap in current sparse-view approaches: Optimization-based methods, which use the sparse measurements directly but lack data-driven priors to suppress artifacts, and learning-based methods, which use strong priors but usually require densely scanned training volumes acquired outside the xSV-CT regime. We study this train-sparse / infer-sparse setting by learning a pose-aware diffusion prior from one 2D projection per training sample, then guiding diffusion sampling with measured-view fidelity, shared-volume reprojection, and projection-physics consistency for tomographic reconstruction. On two public CT datasets across 1-, 4-, and 8-view settings, the proposed method achieves the best novel-view synthesis scores and best or competitive reconstruction quality compared with baselines that use dense-volume supervision.
PaperID: 6597, Poster
Abstract: Numerous distillation methods have been developed to accelerate diffusion models. Recent research indicates that hybrid strategies—blending different paradigms—yield superior results compared to single-method approaches. However, current works predominantly rely on a simple combination of objectives, leaving the question of how to effectively synergize them largely underexplored. In this paper, we specifically investigate the synergy between Score Distillation and Adversarial Distillation. We reveal that their operating zones are distinct and complementary: Score Distillation is effective in high-noise regimes for capturing the global distribution yet lacks fine-grained supervision in low-noise regions. Conversely, Adversarial Distillation is effective in low-noise regions but is prone to instability and suffers from discriminator failure when applied to high-noise intervals. Leveraging this insight, we propose Efficient Adversarial Score Distillation (EASD), a framework that synergizes these objectives via a regime-adaptive sampling strategy, directing each objective to focus on its effective interval. We further apply this paradigm to DiT-based Flow Matching models. Empirically, our method achieves a promising one-step FID of 1.39 on ImageNet 256×256 using the SiT-XL/2+REPA model. This work provides a concrete guideline for maximizing the potential of hybrid distillation.
PaperID: 6598, Poster
Authors: Shrey Singhal, Nannan Huang, Haytham Fayek
Abstract: Aligning Large Language Models (LLMs) with human preferences via Reinforcement Learning from Human Feedback (RLHF) has driven substantial improvements across a wide range of tasks, but the cost of acquiring preference labels remains the dominant bottleneck in the alignment pipeline. While many recent alignment methods improve sample efficiency through active exploration over candidate responses, they typically query the oracle uniformly across contexts, even for contexts where the learned reward estimates are already reliable enough to provide synthetic labels. We propose Adaptive Ensemble-disagreement Routing for Oracle feedback (AERO), a fully online RLHF algorithm built on a Dyna-style model-based RL architecture, in which an ensemble-based Epistemic Reward Model (ERM) serves both as a reward posterior for exploration and as a learned preference-feedback model for synthetic labeling. AERO uses ensemble vote entropy, a Query-by-Committee disagreement measure, combined with an adaptive thresholding mechanism to decide which contexts require oracle feedback and which can instead receive ERM-based synthetic preference labels, concentrating expensive labels where they are most informative. Empirical results across different model families show that AERO achieves highly competitive alignment performance while reducing the required oracle-query budget compared to online DPO and strong active-exploration baselines.
Authors: Minh T Nguyen, Jean Barbier
Abstract: We study the information-theoretic limits of learning a one-hidden-layer teacher network with hierarchical features from noisy queries, in the context of knowledge transfer to a smaller student model. We work in the high-dimensional regime where the teacher width k scales linearly with the input dimension d -- a setting that captures large-but-finite-width networks and has only recently become analytically tractable. Using a heuristic leave-one-out decoupling argument, validated numerically throughout, we derive asymptotically sharp characterizations of the Bayes-optimal generalization error and individual feature overlaps via a system of closed fixed-point equations. These equations reveal that feature learnability is governed by a sequence of sharp phase transitions: as data grows, teacher features become recoverable sequentially, each through a discontinuous jump in overlap. This sequential acquisition underlies a precise notion of effective width k_c -- the number of learnable features at a given data budget n -- which unifies two distinct scaling regimes: a feature-learning regime in which the Bayes-optimal generalization error \varepsilon^\rm BO scales as n^1/(2\beta)-1, and a refinement regime in which it scales as n^-1, where \beta>1/2 is the exponent of the power-law feature hierarchy. Both laws collapse to the single relation \varepsilon^\rm BO=\Theta(k_c d/n). We further show empirically that a student trained with \textscAdam near the effective width k_c achieves these optimal scaling laws (up to a small algorithmic gap), and provide an information-theoretic account of the associated scaling in model size.
PaperID: 6600, Poster
Authors: Bisma Amjad, Henry Wood, William Jacobs, Andrew R Mills
Abstract: In industrial machine vision, removing specular highlights is crucial for accurate surface analysis; however, existing open research remains largely confined to natural images, limiting its efficacy on highly reflective metallic surfaces. When used for zero-shot inference, these methods generate artifacts, making them unreliable for specialized industrial scenes. Since current state-of-the-art methods rely on pixel-level supervision, adapting them to novel industrial subjects is often impractical. Such adaptation requires highlight-free ground truths across hundreds of multi-view scenes, necessitating hardware-intensive cross-polarization, which is difficult to implement at scale in uncontrolled real-world settings. To address these challenges, we propose DeGlare, a flexible training framework for highlight removal in specialized industrial settings. Our approach leverages only a single-view scene by exploiting its multi-illumination observations together with a latent permutation strategy, enabling an encoder-decoder–style architecture to implicitly learn diffuse–specular decomposition. This eliminates the dependency on the burdensome manual annotation required by state-of-the-art supervised methods. We demonstrate the practical effectiveness of our approach on a real-world industrial robotic rig data set, where DeGlare achieves promising performance under limited data and annotation constraints.
PaperID: 6601, Poster
Abstract: In continual learning, most parameter-efficient fine-tuning methods rely on task-specific low-rank adapters to incrementally adapt pre-trained models to new tasks. However, they overlook two critical challenges: one is the accumulation of numerous adapters that leads to linear growth in parameters and memory, and the other is the lack of an explicit mechanism to preserve previously learned knowledge during continual adaptation. In this work, we propose a simple yet scalable framework, termed Progressive Merging of Low-Rank Adapters (PM-LoRA), to address both issues simultaneously. Specifically, PM-LoRA trains a small task-specific spectral orthogonal adapter for each new task and progressively integrates it into a single evolving LoRA adapter that encodes all prior knowledge. This merging process enables continuous knowledge integration while directly mitigating catastrophic forgetting at the parameter level. Benefiting from this design, PM-LoRA achieves true scalability by maintaining constant parameter overhead throughout continual training. Extensive experiments on four benchmark datasets demonstrate that PM-LoRA achieves state-of-the-art performance with superior memory and computation efficiency.
PaperID: 6602, Poster
Abstract: Optimization modeling is the cornerstone of Operations Research (OR). While large language models (LLMs) show promise in autoformulation, existing datasets suffer from unquantified redundancies and domain imbalances. Consequently, models overfit to frequently occurring patterns and fail on underrepresented constraints, inducing a structural capability bias. To address this, we propose Opt-Arena, a data-centric framework that formally represents optimization datasets as tripartite graphs. To map this topology, we introduce Atomic Modeling Information (AMI)—the minimal semantic-mathematical rules of optimization. This representation enables the quantification of structural coverage, difficulty, and overlap. Guided by these metrics, our core-set selection mechanism distills high-density, low-redundancy training data. To further resolve intrinsic domain absences, we extend this framework with a kernel-driven generation pipeline, leveraging foundational problem backbones to synthesize structurally feasible instances for underrepresented areas. Empirically, this approach exhibits exceptional sample efficiency: utilizing only 5k curated instances, it surpasses the 20k-budget performance plateau of conventional distance- and uncertainty-based baselines. Relying entirely on standard SFT driven by our data, lightweight empirical probes outperform both massive proprietary models and complex reasoning architectures enhanced by reinforcement learning. These results reveal that LLM formulation performance is fundamentally constrained by structural data diversity, positioning topology-aware curation as an important driver for advancing OR modeling capabilities.
PaperID: 6603, Poster
Abstract: Long-horizon robotic manipulation is often non-Markovian: the correct action can depend on earlier state changes that are no longer visible in the current observation. Vision-language-action (VLA) models provide strong generalist robot policies, but their performance can drop when a task requires remembering which earlier subtasks changed the world. Existing memory mechanisms extend context through recent-frame windows, compressed latent stores, or similarity-based keyframe retrieval. They do not explicitly distinguish which completed subtasks are relevant from which frames within those subtasks carry the needed evidence. We introduce TTC-MEMORY, a two-tier memory framework for long-horizon VLA policies. The first tier segments execution into subtasks and constructs a dependency directed acyclic graph (DAG), so retrieval is restricted to ancestors of the current subtask. The second tier adaptively stores state-change keyframes within each retained subtask and injects them into a frozen VLA backbone through graph-aware cross-attention, with fallback to recent history when planning or clustering is uncertain. We evaluate TTC-MEMORY on LIBERO, RMBench, SimplerEnv-Bridge, the SafeLab chemistry-lab simulation benchmark, and Franka real-robot tasks. The comparisons use a shared VLA backbone, training budget, simulation seed protocol, and physical-trial protocol. TTC-MEMORY improves the long-horizon and memory-dependent slices in this matched setting. On the 10-subtask SafeLab slice, it reaches 58.3% success, compared with 38.2% for the strongest matched prior memory baseline, MemER, and 45.1% for DAG-only retrieval. Real-robot trials provide supporting transfer evidence rather than the main claim. Ablations show that the two tiers are complementary: DAG retrieval increases long-range dependency recall, while adaptive keyframe selection removes redundant history and preserves attention on task-critical state changes.
PaperID: 6604, Poster
Authors:
Shuo Gao, Chenhao Zheng, Jason RenAbstract: Hour-long videos, streaming feeds, and movie-length content are pushing vision-language models (VLMs) past their breaking point—the standard recipe of encoding each frame into visual tokens yields sequences that grow linearly with video length, causing prohibitive memory consumption and inference latency. Existing approaches, such as token reduction and keyframe selection, attempt to shorten the token sequence but leave the O(n) scaling intact, failing to resolve the fundamental bottleneck. We propose \model, \empha strikingly simple, training-free frame packing framework that compresses arbitrarily long videos into a fixed-size visual representation at constant memory cost. \model requires no fine-tuning, no architectural modification, and no auxiliary models—it plugs directly into off-the-shelf VLMs at inference time, yet consistently surpasses more elaborate baselines. Drawing inspiration from FramePack in generation literature, we propose a new content-aware compression strategy that operates under a fixed token budget for visual understanding. Specifically, \model proceeds in two steps: (1) scoring each frame by its semantic similarity to the text prompt and ranking frames in a diversity-driven manner, and (2) packing the selected frames at varying resolutions into a fixed-length context window. This design concentrates the memory budget on frames that are both instruction-relevant and visually diverse. While in theory \model can pack arbitrarily many frames into the fixed budget, we identify an empirical instantiation that generalizes robustly across different settings. Extensive experiments on LongVideoBench, MLVU, and Video-MME show that \model consistently outperforms full-resolution baselines as well as prior token reduction and keyframe selection methods, across diverse model families, parameter scales, and frame budgets. Notably, on Qwen3.5-35B-A3B with a 128-frame budget, \model yields gains of +12.9%, +10.6%, and +9.7% over the official model on the three benchmarks, respectively.
PaperID: 6605, Poster
Abstract: Large Vision-Language Models (LVLMs) incur substantial inference costs due to the processing of a vast number of visual tokens. Existing methods typically struggle to model progressive visual token reduction as a multi-step decision process with sequential dependencies and often rely on hand-engineered scoring rules that lack adaptive optimization for complex reasoning trajectories. To overcome these limitations, we propose TPRL, a reinforcement learning framework that learns adaptive pruning trajectories through language-guided sequential optimization tied directly to end-task performance. We formulate visual token pruning as a sequential decision process with explicit state transitions and employ a self-supervised autoencoder to compress visual tokens into a compact state representation for efficient policy learning. The pruning policy is initialized through learning from demonstrations and subsequently fine-tuned using Proximal Policy Optimization (PPO) to jointly optimize task accuracy and computational efficiency. Our experimental results demonstrate that TPRL removes up to 66.7% of visual tokens and achieves up to a 54.2% reduction in FLOPs during inference while maintaining a near-lossless average accuracy drop of only 0.7%.
Authors:
Ruishan Guo, Yibing Liu, Guoxin Ma, Yan Wang, Yueyang Zhang, Long Xia, Kecheng Chen, Zhiyuan Sun, Daiting ShiAbstract: Scaling up model parameters has long been a prevalent training paradigm driven by the assumption that larger models yield superior generation capabilities. However, under lossy context compression in a compressor--decoder setup, we find a Size-Fidelity Paradox: increasing compressor size can lessen the faithfulness of reconstructed contexts though reconstruction error decreases. Across 27 compressor setups spanning model families, scales, and compression rates, we coin this paradox arising from two dominant factors: 1) knowledge overwriting: larger models increasingly replace source facts with their own prior beliefs, e.g., ``the white strawberry`` \to ``the red strawberry``; and 2) semantic drift: larger models tend to paraphrase or restructure content instead of reproducing it verbatim, e.g., ``Alice hit Bob`` \to ``Bob hit Alice``. Interestingly, this paradox persists across varied settings, with mid-sized compressors often outperforming larger ones in faithful recovery. By analyzing the compressed memory via embedding geometry and reconstruction determinacy, we further reveal that compressors tend to organize memory across broader semantic subspaces, yielding more ambiguous representations prone to overwriting, drift, and weakened recovery. These findings complement existing evaluations of context compression and expose a breakdown of scaling laws when the objective shifts from plausible generation to faithful preservation.
Abstract: We study posterior contraction rates for sparse Bayesian Kolmogorov-Arnold networks (KANs) over anisotropic Besov spaces, providing a statistical foundation of KANs from a Bayesian point of view. We show that sparse Bayesian KANs equipped with spike-and-slab-type sparsity priors attain the near-minimax posterior contraction. In particular, the contraction rate depends on the intrinsic anisotropic smoothness of the underlying function. Moreover, by placing a hyperprior on a single model-size parameter, the resulting posterior adapts to unknown anisotropic smoothness and still achieves the corresponding near-minimax rate. A distinctive feature of our results, compared with those for standard sparse MLP-based models, is that the KAN depth can be kept fixed: owing to the flexibility of learnable spline edge functions, the required approximation complexity is controlled through the network width, spline-grid range and size, and parameter sparsity. Our analysis develops theoretical tools tailored to sparse spline-edge architectures, including approximation and complexity bounds for Bayesian KANs. We then extend to compositional Besov spaces and show that the contraction rates depend on layerwise smoothness and the effective dimension of the underlying compositional structure, thereby effectively avoiding the curse of dimensionality. Together, the developed tools and findings advance the theoretical understanding of Bayesian neural networks and provide rigorous statistical foundations for KANs.
PaperID: 6608, Poster
Abstract: Diffusion and flow-matching based text-to-speech (TTS) models excel in naturalness but often lack explicit emotion control, as emotional signals remain entangled with speaker identity. We discover that emotion embedding emerges as a linearly decodable direction of frozen hidden states, nearly orthogonal to the direction embedding speaker identity. This inspires a plug-and-play framework DUET for emotion control over pretrained diffusion and flow-matching based TTS models. During generation, DUET unifies dual-space control to achieve fine-grained emotion intervention in a single per-step update: hidden space steering shifts generation along the target emotion direction, while mel-space guidance refines spectral details through gradients backpropagated from a differentiable vocoder. We validate DUET on five architecturally diverse pretrained TTS backbones across three datasets, where it outperforms 10 supervised state-of-the-art emotional TTS baselines across paradigms and achieves the highest human-rated emotion appropriateness. To further showcase its qualitative behavior, we deploy DUET on an Ameca humanoid robot, where it produces richly expressive emotional speech on the humanoid, demonstrating the strong potential for plug-and-play affective interaction for embodied agents. The demonstration is available at: https://anonymous.4open.science/w/duet-demo-fresh-D8E0
PaperID: 6609, Poster
Abstract: While higher-order interaction indices offer deep insights into the synergies and redundancies within tree ensembles, their exact computation is hindered by combinatorial bottlenecks. Recent advancements efficiently compute symmetric interactions, but struggle to scale for asymmetric Probabilistic Interaction Indices (PIIs), which are crucial for advanced tasks such as incorporating feature hierarchy structures. To bridge this gap, we introduce TreePII, a novel algorithm that computes exact, any-order PIIs by integrating the partial derivatives of multilinear extension using interpolatory quadrature. TreePII effectively bypasses traditional computational hurdles, demonstrating both theoretical and empirical improvements over existing baselines and opening new avenues for scalable analysis of tree ensembles.
PaperID: 6610, Poster
Authors:
Yingxin Lai, xinyuan Wang, Yufei Liu, Jialin Guo, Kailin Lyu, Zhiming Luo, Shaozi LiAbstract: Multimodal large language models (MLLMs) are increasingly used for explainable image forgery analysis, where a shared decoder is expected to predict authenticity, localize evidence, and generate natural-language explanations. Existing methods typically rely on final decoder states, leaving open whether these late, language-oriented representations preserve localized manipulation cues. We test this assumption with a token- and layer-wise probing diagnostic that uses pixel-level masks only for analysis to separate tampered-region and authentic-region vision tokens. Across three MLLM-based forgery backbones and multiple manipulation types, we identify a mid-layer forgery evidence peak: tampered-region tokens are most linearly separable at upper-middle decoder layers and become less separable afterward, whereas authentic-region tokens continue to strengthen. This region-conditioned asymmetry suggests competition between localized manipulation cues and late semantic aggregation. Motivated by this diagnostic, we propose Gated Forensic Memory (\method), a lightweight evidence relay that reads from the diagnosed peak layer and writes re-encoded vision-token states back to late vision-token states through a zero-initialized gated residual. \method keeps the LLM frozen, adds no extra vision tokens, and requires no pixel-level training supervision. Across three backbones and seven benchmarks, \method improves authenticity prediction, explanation quality, and spatial alignment without mask supervision; the best read layer matches the diagnosed peak, turning the diagnostic into a practical layer-selection criterion.
PaperID: 6611, Poster
Authors:
Yiteng Zhang, Zixiong Wang, Zhengze Wang, Ke Chen, Xingyu Li, Bin MinAbstract: Latent dynamical models can accurately fit neural population activity, yet accurate activity fitting alone does not guarantee mechanistic validity. Using synthetic benchmarks, we show that even well-fitted low-rank RNNs can yield misleading circuit interpretations when the prescribed activation function deviates from the ground-truth one. To address this limitation, we introduce gain-modulated linear dynamical systems (gmLDS), which decompose latent dynamics into a state-dependent, unit-wise gain and a static low-rank connectivity matrix, allowing the model to adapt to diverse nonlinear responses without assuming a fixed activation function. Across multiple synthetic benchmarks, gmLDS accurately recovers the local effective connectivity of the underlying model, thereby reconstructing the linearized dynamics along neural trajectories. Applied to neural recordings from perceptual and context-dependent decision-making tasks, gmLDS yields interpretable hypotheses about dynamical structure, including attractor structure in perceptual decision-making and context-dependent selection mechanisms. Together, these results support gmLDS as an effective approach for inferring computational structure from neural recordings.
PaperID: 6612, Poster
Abstract: In-context learning is usually analyzed as if the examples in the prompt were sampled before the learner is chosen. In deployed systems they are not: examples are retrieved, filtered, re-ranked, edited, and repeatedly tested after seeing the query and the data store. We formalize this gap by treating prompt construction as adaptive data analysis through a conditional stochastic-kernel calculus. Our main quantity is the \emphcontext capacity, the mutual information or approximate max-information between the data store and the final prompt transcript. For any frozen in-context learner and any adaptive context pipeline with capacity \kappa, we prove a finite-sample context generalization bound of order \sqrt\kappa/k, where k is the number of examples that enter the prompt. We strengthen it to a sub-gamma transport inequality with Bernstein curvature, prove a matching lower bound, and derive two PAC-Bayesian retrieval rules: a first-order Gibbs rule and a self-normalized second-order rule. Finally, we instantiate the theory for Bayesian linear in-context regression, including an order-statistics analysis of query-aligned top-k Gaussian retrieval. The results separate context quality from context certification: longer or more relevant contexts improve risk only when the procedure used to choose them is stable enough for selected-context evidence to transfer.
PaperID: 6613, Poster
Abstract: Projected subgradient descent (PSD) has gained popularity for solving robust Markov decision processes (RMDPs) because it applies to a broader class of uncertainty sets than traditional dynamic programming. Existing work claims that RMDPs with a general compact uncertainty set satisfy the subgradient dominance property, under which exact PSD converges to an \varepsilon-optimal policy in a polynomial number of updates (e.g., Wang et al., 2023). We show that these claims are incorrect. Even when the uncertainty set has cardinality two, the RMDP objective is not subgradient-dominant and can admit suboptimal strict local minima. Moreover, we prove that finding an \varepsilon-optimal policy can be NP-hard even in settings where subgradients are efficiently computable: (i) finite transition uncertainty sets and (ii) sa-rectangular finite transition uncertainty sets with finite cost uncertainty sets. Finally, we identify two conditions under which RMDPs do satisfy subgradient dominance: when, for each policy, either the worst-case transition kernel or the worst-case action-value function is unique.
PaperID: 6614, Poster
Authors: Maria Victoria Carro, Denise Alejandra Mester, Facundo Nieto, Oscar Stanchi, Guido Bergman, Giovanni Franco Gabriel Marraffini, Trinidad Borrell, Guido Freire, Joaquín S Machulsky, Mario Leiva, Ana Lopez, Nicolas Spinelli, Luca Gangi, Eitan Sprejer, Federico Barrera-Lemarchand, Gerardo Simari, Maria Vanina Martinez
Abstract: How can we effectively oversee AI systems that surpass human intelligence? One proposed answer is AI Safety via Debate, an approach in which two AI systems argue opposing positions and then a judge, who can be either a human or a weaker model, decides which one is correct. By facilitating supervisors in providing high-quality training signals, debate aims to encourage truthful and safe behavior from AI systems whose capabilities exceed our own. However, human judgment is not neutral: people often rely on biases which advanced AI may learn to exploit. Drawing on insights from collective intelligence and social cognition, this paper investigates whether debate can guide human and LLM supervisors toward the truth despite their prior beliefs. We introduce and evaluate several protocol configurations, including (1) the use of open-source, proprietary models and (2) human participants as judges (N=507), (3) different methods for eliciting prior beliefs, and (4) variations in claims within the same topic to test whether belief updates generalize across related subtopics. We also explore setups inspired by the wisdom of the crowds, such as the number and diversity of judges, as well as interaction mechanisms including deliberation and voting, designed to better align judgments with the truth. Overall, we find that debate increases accuracy relative to prior beliefs for human judges (from 41.3% to 57.6%), whereas the consultancy baseline, in which a single expert argues for one answer, does not. For weaker language models, however, the results are mixed. Debate does not provide similar benefits in most settings and can even degrade decision accuracy. Regarding multi-judge protocols, human deliberation yields the highest accuracy for human supervisors (62.2%), while for LLMs it mainly benefits those starting from incorrect beliefs, without producing significant improvements in aggregate performance. We provide empirical evidence supporting Debate as a viable path toward scalable human oversight, while casting doubt on the assumption that such benefits translate to weaker AI supervisors, raising concerns about the reliability of fully automated training and control pipelines in high-stakes AI safety contexts.
PaperID: 6615, Poster
Authors:
Yiqun Wu, QI WANG, Mengxian Li, Yongjun XuAbstract: Communication enhances cooperative multi-agent reinforcement learning (MARL) under partial observability. In real networks, however, each agent's available bandwidth varies over time due to changing link quality, interference, and congestion. Most existing MARL communication policies decide whether to transmit based on message importance or novelty, or are trained under fixed bandwidth conditions. As a result, they cannot adapt sending frequency to bandwidth variations, leading to packet loss when bandwidth decreases or missed information when bandwidth recovers. To address this, we propose Dual-Decoupled Adaptive Communication (DDACOM)---a framework that decouples message sending and message generation from the action policy. The message-sending module ranks each message by its deviation from teammates' information and the sender's history, and uses a dynamic token bucket to convert the probe-estimated safe rate into a real-time token budget, enabling the communication frequency to track the safe rate online. The message-generation module is trained independently with receiver-side attention feedback and counterfactual marginal value, ensuring each transmitted message meaningfully contributes to cooperation. On SMAC and MPE, DDACOM achieves leading cooperative performance among baseline methods. Under dynamic network conditions in ns-3 simulation, DDACOM sustains over (95%) safe-rate utilization and packet loss below (0.3%) while reducing communication frequency to less than half of full broadcast, confirming its adaptability to runtime network variation.
PaperID: 6616, Poster
Abstract: Many partially observed data streams are globally high-rank yet locally low-rank: user–item interactions in recommender systems span multiple preference groups, distinct operating regimes in sensor and network monitoring induce different low-dimensional structures, and the active latent subspace in adaptive systems shifts over time. We formalize this as online matrix completion on a drifting union of subspaces: each sample lies in one of K low-dimensional subspaces, only a random subset of its coordinates is observed, and the subspaces drift across epochs. We propose MoSAIC (MoE Subspace Adaptation via Incremental Completion), a method based on a routed mixture of low-rank experts trained in two phases. A base model is first pre-trained on samples from a fixed source distribution. It is then adapted on the non-stationary stream by routing each incoming sample to a single expert and updating only that expert. Our analysis identifies sufficient conditions under which the router remains correct with high probability throughout learning, while within each epoch the routed experts contract toward the current subspaces at a sublinear rate. The technical core is a uniform routing concentration argument that converts the random time steps on which each expert is updated into a deterministic time scale, reducing the per-expert analysis to a tractable stochastic recursion. Experiments on streams with repeated non-stationary changes corroborate the theory and show clear improvements over competitive baselines.
PaperID: 6617, Poster
Abstract: Multi-modal image fusion is crucial for comprehensive scene representation, yet real-world degradations compromise its efficacy, necessitating degradation-robust fusion paradigms. Existing degradation-robust multi-modal image fusion methods are often hindered by cascaded sub-optimal solution or black-box opacity. To address these challenges, we propose a theory-inspired mutual-promoting deep unfolding framework. It reformulates degradation-robust image fusion as a joint optimization problem, utilizing the degradation model and modality generation mechanism to explicitly model both single-task priors and cross-task dependencies. By decomposing the complex multi-task optimization problem into task-related iterative subproblems, the framework establishes a bidirectional reciprocal information flow between restoration and fusion. This process is further unfolded into a multi-stage deep neural network, where each component explicitly corresponds to a specific mathematical operation. In our framework, the fused cross-modal prior actively regularizes the ill-posed restoration, while the purified modality features continuously refine the fusion output. It ensures a transparent architecture that combines the merits of both model-based and data-driven methodologies. Extensive experiments demonstrate that our paradigm achieves state-of-the-art performance while providing superior interpretability.
PaperID: 6618, Poster
Abstract: Adam is the default optimizer for training modern deep neural networks, yet its adaptive behavior remains poorly understood due to the complex interaction between its first- and second-moment exponential moving averages (EMAs). We study Adam in the tied-\beta regime, where the two EMA decay rates are equal, and show that its adaptive dynamics can be expressed through a transformed ratio with approximately scale-stable behavior. Empirically, this transformed ratio exhibits a stable, heavy-tailed distribution across tasks, model scales, and training stages, in contrast to the variability of raw moment magnitudes. This empirical stability enables both practical and conceptual consequences. First, we derive a recurrence for the transformed ratio, yielding a reparameterization of Adam that replaces the second moment with a compressible state. Leveraging its stable distribution, we show that a fixed 4-bit codebook is sufficient in our experiments to store this state without auxiliary scaling, achieving performance competitive with full-precision Adam. Second, the transformed ratio view clarifies Adam's connection to sign-based methods: Adam reduces to sign momentum modulated by the transformed ratio, and replacing it with a constant recovers Signum as a limiting case. This perspective further provides a simple rule for transferring learning rates between the two methods. Together, these results suggest that tied-\beta Adam admits a simple and approximately stable ratio structure underlying its adaptive behavior and demonstrate its utility for both analysis and efficient implementation.
PaperID: 6619, Poster
Authors:
George Andriopoulos, Bimarsha Adhikari, Soyuj J Basnet, Juan Guevara, Li Guo, Keith RossAbstract: Multi-output regression can be approached either by fitting one joint vector-valued predictor or by training separate univariate predictors for each target coordinate. This holistic-versus-separate choice is classical in the largely non-neural multi-output regression literature, but is less characterized analytically for neural networks, where joint prediction arises naturally through shared features and vector-valued output layers. Because neural networks are highly non-linear and high-dimensional, exact analysis is notoriously difficult. We study two questions using the Unconstrained Feature Model (UFM) as a tractable surrogate for exact analysis: (1) how does joint-output regression compare with independent coordinate-wise regression under regularization? (2) when can target whitening or normalization improve original-scale MSE? Under UFM, the minimal training MSE depends only on the regularization product and the spectrum of the target covariance. This yields two theoretical results: joint-output regression has no larger training MSE than independent coordinate-wise regression under matched regularization, and whitening or normalizing targets can help or hurt original-scale MSE depending on the average target variance. Experiments on neural networks trained with standard weight decay across robotic imitation-learning and autonomous-driving datasets support these theoretical results, with test-MSE trends following the same qualitative pattern. Our results show that simplified analytical surrogates can provide useful guidance for neural multi-output regression.
Abstract: Autoregressive language models (ARMs) have been shown to memorize and occasionally reproduce training data verbatim, raising concerns about privacy and copyright liability. Diffusion language models (DLMs) have recently emerged as a competitive alternative, yet their memorization behavior remains largely unexplored due to fundamental differences in generation dynamics. To address this gap, we present a systematic theoretical and empirical characterization of memorization in DLMs. We propose a generalized probabilistic extraction framework that unifies prefix-conditioned decoding and diffusion-based generation under arbitrary masking patterns and stochastic sampling trajectories. Theorem 4.3 establishes a monotonic relationship between sampling resolution and memorization: increasing resolution strictly increases the probability of exact training data extraction, implying that autoregressive decoding corresponds to a limiting case of diffusion-based generation by setting the sampling resolution maximal. Extensive experiments across model scales and sampling strategies validate our theoretical predictions. Under aligned prefix-conditioned evaluations, we further demonstrate that DLMs exhibit substantially lower memorization-based leakage of personally identifiable information (PII) compared to ARMs.
Authors:
Jiawei Ge, Xintian Zhang, Jiuxin Cao, Bo Liu, Fabian Deuser, Chang Liu, Gong Wenkang, Siyou Li, Juexi Shao, Wenqing Wu, Chen Feng, Ioannis PatrasAbstract: Cross-view Referring Multi-Object Tracking (CRMOT) aims to track multiple objects specified by natural language across multiple camera views, with globally consistent identities. Despite recent progress, existing methods rely heavily on costly frame-level spatial annotations and cross-view identity supervision. To reduce such reliance, we explore CRMOT under weak supervision by leveraging the capabilities of foundation models. Our empirical study shows that directly applying foundation models such as SAM2 and SAM3, even with task-specific modifications, fails to accurately understand referring expressions and maintain consistent identities across views. Yet, such models remain effective at producing reliable object tracklets that can serve as supervision signal. We therefore repurpose foundation models as pseudo-label generators and propose a two-stage framework for weakly supervised CRMOT, using only object category labels as coarse-grained supervision. In the first stage, we design an Affinity-guided Cross-view Re-prompting strategy to refine and associate SAM3-generated tracklets across cameras, producing reliable cross-view pseudo labels for subsequent training. In the second stage, we introduce ViewSAM, a CRMOT model built upon SAM2 that explicitly models view-aware cross-modal semantics. By formulating view-induced variations as learnable conditions, ViewSAM bridges the gap between view-variant visual observations and view-invariant textual expressions, enabling robust cross-view referring tracking with only approximately 10% additional parameters. Extensive experiments demonstrate that ViewSAM achieves SOTA performance under weak supervision and remains competitive with fully supervised methods.
Authors:
Zhenlong Yuan, Yue Wang, Jing Tang, Rui Chen, Kejin Cui, Lei Sun, Dapeng Zhang, Hongwei Yu, Chengxuan Qian, Xiangxiang Chu, Shuo Li, Yuyin ZhouAbstract: Multimodal Large Language Models have shown promising capabilities in bridging visual and textual reasoning, yet their reasoning capabilities in Open-Vocabulary Human-Object Interaction (OV-HOI) are limited by cross-modal hallucinations and limited viewpoints of images. To address this, we propose ImagineAgent, an agentic framework that integrates cognitive mapping, tool-augmented reinforcement learning (RL), and generative world modeling for robust OV-HOI understanding. Specifically, we first propose an innovative CoT dataset named hicodet-6K for supervised fine-tuning (SFT), which effectively bridges the perception-to-cognition gap by structuring perceived entities into interaction pairs for comprehensive predictions. Subsequently, we develop a multimodal tool library integrating online retrieval, image cropping, and generative modeling, enabling the agent to dynamically augment reasoning with domain-specific tools to resolve visual-semantic ambiguities and hallucinations during inference. Moreover, we incorporate a generative model to reconstruct alternative viewpoints, enabling the agent to “imagine” under limited viewpoints. Finally, we propose a composite reward mechanism to jointly optimize prediction accuracy and tool efficiency. Evaluations on both SWIG-HOI and HICO-DET datasets demonstrate that our method achieves state-of-the-art performance while requiring merely 36.7% of the training data compared to existing methods, validating our robustness, empirical effectiveness and efficiency.
Abstract: Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances. Existing models can describe objects in natural language, but they struggle to precisely represent continuous object poses and generate geometrically consistent images under target viewpoints. To mitigate this, we propose \emphObject-Uni, a unified model for object-centric spatial understanding and controllable generation. Specifically, we formulate object-centric spatial intelligence as a unified problem connecting pose perception, spatial reasoning, pose-conditioned generation, and object-centric novel view synthesis. We treat object pose as an explicit geometric variable shared by understanding and generation, rather than merely a prediction label or control signal. To make pose usable by multimodal large language models, we propose a viewpoint-based orientation abstraction that maps orientation into structured viewpoint descriptions while preserving continuous geometric supervision. We further construct an object-centric spatial benchmark (UniSpatial-80K) and train a unified model with an object-token-grounded pose anchor to associate each instance with its pose state. Experiments show that our model improves object-level pose understanding and pose-controllable generation, moving unified models from describing objects toward manipulating spatial states.
Abstract: We introduce the first generative foundation model for chest radiograph synthesis trained from scratch at the billion-parameter scale. Existing radiographic AI models often suffer from poor generalisation across patient subpopulations, institutions, and acquisition settings, resulting in limited clinical utility. Controlled, high-fidelity synthesis of chest radiographs is a promising path toward diversifying clinical datasets and evaluating the robustness of diagnostic models. Therefore, we present the largest specialist generative foundation model for chest radiographs to date, with over 1.3B parameters, trained for 1.6T tokens on a curated, heterogeneous dataset comprising 1.2M radiographs and clinical expert-guided metadata. Our model supports controllable radiograph generation and editing across multiple demographic subgroups, acquisition views, and a dozen pathologies. Moreover, we significantly advance the state of the art in radiograph synthesis fidelity, producing images that are indistinguishable from real radiographs to expert pulmonologists.
PaperID: 6625, Poster
Abstract: Command-Line Interface (CLI) agents based on large language models (LLMs) demonstrate remarkable autonomous capabilities, but they also introduce significant safety and misuse risks during multi-turn interactions with external environments. Existing safety mechanisms mainly rely on external guardrails, which have a limited ability to perform fine-grained behavioral control during execution. Meanwhile, recent mechanistic interpretability methods for LLM safety are mostly confined to single-turn or jailbreak-style QA settings, limiting their ability to capture the evolving risk dynamics of multi-turn agent execution. In this paper, we investigate the safety of multi-turn CLI agents from an internal perspective. We propose Agent MechSuits (Mechanistic Subspace Intervention and Steering), a white-box defense framework that performs runtime safety detection and representation-level mitigation for CLI agents. Unlike conventional agent guardrails, Agent MechSuits detect harmful execution states from step-level hidden representations and mitigate unsafe behavior by intervening in a 10-dimensional subspace within a single layer. To support this research, we introduce the Mechanistic Agent Safety (MAS) benchmark, comprising comprehensively annotated multi-turn execution trajectories across 194 tasks using LLaMA-3.1-8B, Qwen-2.5-7B, and Gemma-2-9B. Extensive experiments show that Agent MechSuits achieves strong safety detection performance, supports lookahead risk anticipation, and substantially reduces harmful agent actions, establishing a foundation for applying mechanistic interpretability to dynamic LLM agent safety.
PaperID: 6626, Poster
Abstract: We propose a computationally efficient vine copula-based mutual information (MI) estimator. Unlike existing non-parametric density estimators that suffer from the curse of dimensionality, a non-parametric vine copula has a convergence rate that is independent of the dimensionality of data. We leverage this property to tackle the challenging task of MI estimation and propose an interpretable MI estimator. Extensive experiments on datasets with known ground-truth MI values across dimensions, data types, and (input, output) dependence structures demonstrate a superior trade-off between MI estimation error and computational time compared to SotA neural MI estimators.
Abstract: Modern video generative models produce visually impressive results, yet frequently violate basic physical principles. We propose Proprio, a training-free framework that enables a frozen video generator to assess and improve the physical plausibility of its own outputs. Inspired by proprioception, the biological sense of one’s own movement, Proprio treats the model's flow residual under controlled latent perturbations as a self-scoring signal. Samples that are better explained by the generator's learned dynamics induce smaller and more stable residuals. We aggregate this signal across timesteps and perturbations, focus it on motion-relevant regions with a dynamic spatiotemporal mask, and use it for best-of-N search, gradient-based self-refinement, or both. Across text-to-video and image-to-video benchmarks, Proprio consistently improves physical plausibility, outperforming VLM-based scoring, and external world-model baselines in several settings. On TurboWan2.2, Proprio improves Physics-IQ from 32.2 to 37.5 (+16.5%) and VideoPhy2-hard physical commonsense from 45.6 to 55.0 (+20.6%). Human evaluation further shows that raters prefer Proprio-selected or refined videos for physical plausibility in roughly two-thirds of comparisons. These results suggest that frozen video generators contain actionable internal signals for evaluating and improving the physical plausibility of their own outputs.
PaperID: 6628, Poster
Authors:
Zehui Feng, Weichuan Wang, Xiaohan Chen, Cuntai Guan, Ting HanAbstract: Recent speech brain-computer interfaces (BCIs) increasingly rely on large models and complex multi-stage pipelines for neural decoding. However, in speech BCI, where training data are scarce, noisy, and expensive to collect, the benefits of model scaling remain unclear. This work revisits a fundamental question: are large models truly necessary for accurate neural decoding? We propose BEST (Brain-Encoding-Speech-Text), a two-stage framework for efficient neural-to-text decoding. Instead of relying on large generative models or phoneme-centric pipelines, BEST introduces hierarchical intermediate representations and a compact speech-oriented decoder that directly maps neural activity to text. This design preserves essential structure while significantly reducing computational cost, enabling robust generalization in low-resource settings and practical deployment under constrained resources. On the Brain-to-Text ’24 and ’25 benchmarks, BEST achieves word error rates (WERs) of 6.19% and 1.62%, espectively, using only 3.7% of the parameters required by comparable methods. Beyond these results, we provide three key insights: (1) decoder scaling yields non-monotonic gains; (2) cross-modal alignment is more critical than model size; and (3) overfitting degards sequence-level performance. Representation analysis further reveals a consistent neural→audio→text alignment pathway.
Abstract: When an RL agent's observations contain distractors driven by the same confounders as its true state, observational data alone cannot identify which dimensions the agent controls. In our benchmarks, even state-conditioned observational selectors can collapse when distractors mimic controllable state variables. We propose Interventional Boundary Discovery (IBD), which treats the agent's own action channel as a source of randomized interventions: randomizing actions implements an interventional contrast, and per-dimension two-sample tests with FDR correction produce a binary mask over observation dimensions. Across 12 continuous-control settings with up to 100 distractors, IBD matches oracle return in 11 of 12 settings, while observational baselines including mutual information, state-conditioned forward models, and gradient-based sensitivity often underperform simply passing the full observation to SAC.
PaperID: 6630, Poster
Abstract: Latent Diffusion Models dominate generative modeling but face two-stage training complexity and the inherent reconstruction limits of pre-trained autoencoders. Pixel-space Flow Matching offers a simplified end-to-end alternative yet often encounters training instabilities when using the standard v-prediction objective. Recent research attributes this difficulty to the off-manifold nature of noised targets, prompting a widespread shift toward x-prediction. In this paper, we carefully investigate this prevailing assumption through patch-wise Principal Component Analysis (PCA), and we reveal that the failure of v-prediction is primarily a bandwidth issue. Specifically, the velocity field in minor components is mostly determined by the raw noisy input, which differs from the dynamics in the principal component space. Deep backbones must propagate this minor information to the final layer to predict the velocity field, which consumes significant model capacity that should be dedicated to principal components. This is particularly critical in pixel space where noise dimensionality is comparable to the hidden dimension. To resolve this, we simply incorporate a term containing the raw noisy input into the final layer. This architectural bypass enables the backbone to focus on high-level semantics while the final layer efficiently reconstructs the minor velocity field. Our method successfully revives v-prediction and achieves competitive performance on ImageNet 256×256 dataset compared to x-prediction baselines.
PaperID: 6631, Poster
Abstract: Robot policy pretraining based on human videos is crucial for improving the policy’s generalization ability. One main challenge for this task is the lack of explicit action-relevant representations in such unlabeled data. Recent works tend to estimate 3D hand motion trajectories from videos using 3D foundation models (3D FMs). However, they keep solely the sparse trajectories as supervision, ignoring the rich, fine-grained information produced during the trajectory extraction process. To address this issue, we propose LSA, a hierarchical robot policy pretraining framework that includes three phases: Lift, See, and Act. Specifically, the Lift phase elevates 2D video observations into dense 3D representations by adaptively aligning features in the policy's shallow layers with the intermediate representations extracted from multiple 3D FMs. Building upon these lifted representations, the See phase equips the policy with explicit geometric and interactive awareness through dual guidance. It introduces depth-map reconstruction to help the model comprehend 3D spatial layouts and utilizes hand region cues to explicitly supervise the encoder’s attention toward task-relevant interaction hotspots. Empowered by the dense 3D representation learning and precise spatial guidance, our policy achieves robust and accurate robotic manipulation in the Act phase. Extensive experiments on diverse simulation and real-world tasks demonstrate that LSA significantly outperforms current state-of-the-art approaches. The code of this work will be released soon.
PaperID: 6632, Poster
Authors: Pawan Kumar
Abstract: Diffusion language models generate text through iterative denoising, but current interpretability tools mostly analyze static hidden states and do not explain how concept information evolves across denoising time. We introduce DynaManifold-SAE, a method for discovering sparse autoencoder feature groups whose activations form time-dependent concept manifolds in masked and diffusion language models. The method builds sparse latent codes across mask ratios, constructs candidate feature groups, and evaluates them with held-out coordinate prediction, geodesic consistency, persistence across time, and causal interventions. Across BERT and LLaDA, DynaManifold-SAE reliably identifies sparse sentiment geometry that is stable across seeds and stronger than individual-feature, PCA, random-group, and graph-structured baselines. The discovered groups transfer from controlled templates to natural SST-2 examples, indicating that they capture semantic structure rather than template artifacts. We further show that these groups explain a denoising-specific mechanism: they predict token reveal order and confidence growth during LLaDA unmasking, and targeted interventions alter reveal confidence and token recovery more than matched controls. These results suggest that sparse feature manifolds provide a practical bridge between static representation geometry and the dynamics of diffusion language generation.
PaperID: 6633, Poster
Authors:
Rohit Muralitharan, Mark Squillante, Chen Wang, Xuchuang Wang, Chai Wah WuAbstract: Recent advances in online learning have achieved an O(n/\Delta_[2]) query complexity in quantum n-arm multi-armed bandits (MABs) for best-arm identification, which represents a quadratic speedup over their classical counterparts. Here and throughout, \Delta_[2] is the gap between the mean of the best and second-best arms. However, the implementations of these algorithms remain on a small scale due to various aspects of quantum hardware restrictions. One of the most significant restrictions is \emphmemory decoherence, in which the qubits quickly lose information after a short period of time. Memory decoherence can occur at \emphintra-circuit and \emphinter-circuit levels. Intra-circuit decoherence happens often in near-term quantum computers, where the noise becomes overwhelming when the circuit is deep. On the other hand, inter-circuit decoherence captures future quantum machines, where a quantum memory is promising, but objects stored in the memory for too long can still become decoherent. In this paper, we present a comprehensive treatment of both memory decoherence problems. For inter-circuit decoherence, we derive a framework based on streaming bandits, and we obtain an algorithm that finds the best arm with high constant probability using O(n/\Delta_[2]) queries and an O(\logn) decoherence window. This is the first quantum algorithm for MABs that achieves an o(n) decoherence window, which represents an exponential improvement. For intra-circuit decoherence, we observe that only O(1)-depth circuits can be used to avoid significant decoherence. Therefore, we devise various hardware optimizations and a destructive SWAP test subroutine to speed up the empirical performance. On the IBM Q System One machine with 127 qubits, our implementation scales the quantum MABs to over 1000 arms, and our algorithms consistently achieve significant efficiency improvements over classical bandit algorithms in both physical experiments and computer simulations.
PaperID: 6634, Poster
Authors: Arnab Bhattacharjee, Wayes Tushar, Tapan K Saha
Abstract: Learning effective representations from battery operational history is essential for predictive health management and control-oriented applications. Existing approaches rely on response-dependent engineered features or short early-life windows, limiting generalisability and precluding use in prospective settings where future responses are unavailable. We propose \textttNOCE-Net, a deep nested sequence model that learns control-enabling representations directly from raw cyclic operational data. Our data pipeline, \textttre-BatteryData, standardises heterogeneous aging datasets into semantically equivalent Full Equivalent Cycles (FECs) defined by absolute charge throughput, enabling end-to-end learning across diverse operational regimes. \textttNOCE-Net combines multi-patch transformer encoders for intra-cycle dynamics with a GRU backbone for inter-cycle state evolution, operating on approximately 700 batteries across ten public datasets spanning over 300 operational profiles. In zero-shot evaluation on unseen battery chemistries and usage patterns, \textttNOCE-Net achieves 13.7% MAPE in voltage response prediction, outperforming strong long-sequence baselines by over 43% while using 40 × fewer parameters. Shape and component diagnostics of the learned hidden states reveal that the GRU's principal direction of evolution achieves Spearman |\rho| \geq 0.95 with the capacity fade trajectory on average across all datasets with monotone degradation profiles, including zero-shot test batteries, providing evidence that the representations encode degradation dynamics transferable to downstream tasks beyond battery response modelling.
PaperID: 6635, Poster
Abstract: Training a single 3D Gaussian Splatting (3DGS) scene routinely requires tens of gigabytes of GPU memory, placing standard 3DGS training beyond many low-memory GPU budgets commonly found on laptops and edge devices. Existing low-memory approaches mainly follow two directions: block-wise training and frustum-culling-based host offloading. Block-wise methods train spatial blocks independently and merge them afterward. However, because 3DGS renders each image through depth-ordered alpha compositing over all visible Gaussians, independent block training reduces optimization to block-local objectives, thereby losing full-image loss supervision. Frustum-culling-based host offloading retains full-image loss supervision by jointly rendering all visible Gaussians while uploading only the view-visible subset to the GPU. However, its GPU memory consumption still scales with the size of the visible set, which can exceed the memory budget of low-end GPUs for dense scenes or wide-coverage views. We propose CC-GS, a low-memory block-wise 3DGS training framework based on CPU-GPU block-wise context compositing, featuring three key designs. i) Block-wise context compositing: CC-GS partitions Gaussians into capacity-bounded blocks and represents non-optimized visible blocks using cached per-pixel color and residual transmittance, with depth used only for ordering, enabling full-image loss supervision while processing one capacity-bounded block at a time. ii) Two-pass block-wise training: each training view is processed through a lightweight context pass for caching compositing context and a differentiable block-optimization pass for updating individual blocks. iii) Asynchronous CPU-GPU execution: CC-GS overlaps block gathering, host-device transfers, rendering, backpropagation, gradient offloading, and CPU-side updates through an asynchronous CPU-GPU pipeline. Experiments on standard 3DGS benchmarks show that CC-GS keeps the measured peak GPU memory below 2GB for all evaluated scenes while matching the rendering quality of standard 3DGS. Compared with standard 3DGS, CC-GS, incurs at most a 1.54× dataset-level training slowdown.
PaperID: 6636, Poster
Abstract: Survival models are often trained at target clinical sites where outcomes are limited and censored, while patient-level data from external cohorts may be unavailable because of institutional data-sharing constraints. In many applications, the transferable knowledge is not a dataset, a calibrated survival curve, or estimated regression coefficients, but a probability-free prognostic artifact such as a clinician-defined scorecard, guideline-based staging rule, registry-derived risk index, or coarse clinical risk group. These weak teachers may encode human expertise or evidence from prior cohorts, yet they often have only ordinal meaning and need not share a numerical scale with the target model. We propose NNCoxKL, a transfer-learning framework that trains a deep Cox model on target time-to-event data while distilling such human- or model-derived prognostic signals through Cox risk sets. For each event risk set, NNCoxKL converts both the external signal and the neural Cox risk score into Plackett--Luce distributions and regularizes the Cox partial likelihood through a temperature-scaled KL alignment term. This transforms clinical knowledge into a censoring-aware ranking prior over who is most likely to fail next, rather than treating the external score as a covariate or requiring it to represent calibrated survival probabilities. The external signal enters only through this risk-set regularizer, allowing the model to learn nonlinear target-cohort effects while borrowing relative-risk information from external clinical knowledge. Experiments on independent clinical transfer tasks and controlled survival benchmarks show that NNCoxKL improves discrimination and calibration-sensitive prediction over internal-only deep Cox models and common score-transfer baselines.
PaperID: 6637, Poster
Authors: Abdessalam Ed-dib, Amine Aboussalah
Abstract: Message-passing neural networks (MPNNs) learn node representations by iteratively aggregating information from neighbors, but stacking too many layers causes all representations to converge, a phenomenon known as oversmoothing. We propose a geometric perspective on this phenomenon. At each depth k, the propagation operator assigns to every node v a walk distribution p_v^(k), a probability distribution over all nodes describing how v distributes its attention across the graph. These walk distributions form a point cloud on the probability simplex, and oversmoothing is the contraction of this cloud to a single point. To measure this contraction, we exploit the intrinsic geometry of the simplex: the Fisher--Rao metric, the unique Riemannian metric invariant under sufficient statistics, whose chordal distance under the square-root embedding p \mapsto \sqrtp is the Hellinger distance. The resulting diagnostic, the \emphaggregation dispersion \mathcalD_k, is the variance of the embedded point cloud on the unit sphere. We prove that: (i) \mathcalD_k = 0 if and only if the k-th step propagation matrix has rank one, providing a necessary and sufficient condition for representational collapse; (ii) \mathcalD_k is monotonically non-increasing with depth, so each layer irreversibly spends a finite diversity budget; (iii) the spectral gap of the propagation matrix controls the rate of this decay; and (iv) for contractive MPNNs, the aggregation dispersion upper-bounds the representation dispersion up to architecture-dependent constants, connecting the geometry of the simplex to the geometry of the feature space. In experiments across 14 model configurations and 11 datasets, \mathcalD_k achieves the highest average Pearson correlation with accuracy degradation among all training-free diagnostics (|r| = 0.721), outperforming the post-training effective rank (|r| = 0.688).
Abstract: Large language models (LLMs) increasingly act as autonomous agents that must decide when to answer directly vs. when to invoke external tools. Prior work studying adaptive tool use has largely treated tool necessity as a model-agnostic property, annotated by human or LLM judge, and mostly cover cases where the answer is obvious (e.g., fetching the weather vs.\ paraphrasing text). However, tool necessity in the wild is more nuanced due to the divergence of capability boundaries across models: a problem solvable by a strong model on its own may still require tools for a weaker one. In this work, we introduce a model-adaptive definition of tool-necessity, grounded in each model's empirical performance. Following this definition, we compare the necessity against observed tool-call behavior across four models on arithmetic and factual QA dataset, and find substantial mismatches of 26.5–54.0% and 30.8–41.8%, respectively. To diagnose the failure, we decompose tool use into two stages: an internal cognition stage that reflects whether a model believes a tool is necessary, and an execution stage that determines whether the model actually makes a tool-call action. By probing the LLM hidden states, we find that both signals are often linearly decodable, yet their probe directions become nearly orthogonal in the late-layer, last-token regime that drives the next-token action. By tracing the trajectory of samples in the two-stage process, we further discover that the majority of mismatch is concentrated in the cognition-to-action transition, not in cognition itself. These results reveal a knowing-doing gap in LLM tool-use: improving tool-use reliability requires not only better recognition of when tools are needed, but also better translation of that recognition into action.
PaperID: 6639, Poster
Abstract: Large language models (LLMs) have achieved strong performance across diverse reasoning tasks, yet it remains unclear whether this performance reflects an understanding of underlying logical structure or a reliance on shallow semantic cues. A key indicator of such understanding is whether models treat logically equivalent inputs consistently, producing the same predictions despite differences in expression. Through systematic evaluation on curated data, we show that even strong LLMs often fail this test: they do not consistently preserve predictions across logically equivalent inputs and generalize poorly to unseen logical forms. To address this limitation, we propose LoGIcal STructure-guided Reasoning (LoGIST), a framework that guides LLM reasoning with explicit logical structure. LoGIST maps each input instance to a decision diagram, a directed acyclic graph that captures the underlying logical structure, and encodes this graph into a logical embedding that conditions the LLM’s reasoning process. Across multiple training settings, LoGIST reduces inconsistency under logical equivalence and improves generalization to unseen logical forms, suggesting that logical invariance is a useful target for building more robust and generalizable logical reasoning in LLMs.
PaperID: 6640, Poster
Abstract: Air quality prediction at unmonitored locations is challenging because historical AQI observations are unavailable, while conventional spatiotemporal models largely rely on target histories to learn spatial dependencies. Meanwhile, widely available weather covariates provide useful station-specific evidence for target-absent inference. In this paper, we propose ExtrapAir, a lightweight weather-bridged framework for air quality inference at unmonitored locations. ExtrapAir preserves variable-specific weather patterns through per-variable weather encoding and learns adaptive weather and AQI correlations for spatial transfer. To support robust inference, we introduce Activation Attention with Prior Inductive Biases, which enables non-competitive spatial aggregation, incorporates geographic and semantic priors for station-pair guidance, and avoids full station-to-station attention computation for scalability. Extensive experiments on five real-world air-quality datasets against 32 baselines show that ExtrapAir reduces MAE by up to 15.24% overall and 16.01% on unmonitored stations, while cutting training time by 84.9% and memory usage by 76.23%, achieving a strong accuracy--efficiency trade-off.
Abstract: Latent Gaussian models (LGMs) are a popular class of Bayesian hierarchical models that include Gaussian processes, as well as certain spatial models and mixed-effect models. Efficient Bayesian inference of LGMs often requires marginalizing out the latent variables. For LGMs with a non-Gaussian likelihood, exact marginalization is not possible and a popular approach is to do approximate marginalization with an integrated Laplace approximation (ILA). Using ILA produces an approximate posterior which, in some settings, can differ significantly from the correct posterior, which impacts downstream applications. We propose an importance sampling scheme to correct the error introduced by ILA. By increasing the number of samples in importance sampling, the posterior with ILA converges to the correct posterior. This idea is realized with various techniques, including pseudo-marginalization, quasi-Monte Carlo and randomized quasi-Monte Carlo. We implement our methods in an automatic differentiation framework to support gradient-based algorithms when doing inference on the hyperparameters. For the latter, we specifically consider the use of Hamiltonian Monte Carlo. We demonstrate the benefits of reduced error in various applied models.
Authors:
Jewon Yeom, Jaewon Sok, Heejun Kim, Seonghyeon Park, Jeongjae Park, Taesup KimAbstract: Hallucination is often viewed as a direct consequence of missing knowledge: a model answers incorrectly when the correct answer is absent from its generation-time distribution, and correctly when it is present. We test this assumption by introducing a semantic notion of answer availability that aggregates token-level variants expressing the same answer concept, and asks whether the correct concept is already available at the moment the model commits to an answer. Across Qwen and Llama models from 0.8B to 72B in both Instruct and Base variants, 16-47% of Instruct hallucinations occur with substantial probability mass already on the correct concept, and the rate rises monotonically with scale. Comparing such failures against correct generations with matched semantic support, the distinguishing factor is not whether the correct concept is represented, but how its probability is distributed: correct generations concentrate mass on a single surface form, hallucinations disperse it across alternatives. The same sharpening asymmetry extends across multi-token generation and is detectable in pre-generation hidden states. Together, these results identify a single mechanism: instruction tuning sharpens answer commitment with scale, making helpfulness and confident hallucination two consequences of the same underlying disposition.
PaperID: 6643, Poster
Abstract: In this paper, we study last-iterate convergence in stochastic constrained convex--concave minimax optimization. A key difficulty is that the last iterates of vanilla stochastic extragradient (S-EG) and stochastic optimistic gradient descent--ascent (S-OGDA) can fail to converge in the presence of gradient noise, even for simple bilinear problems. To address this issue, we regularize the original convex--concave problem into a strongly convex--strongly concave one. Applying S-EG and S-OGDA to the regularized problem gives two simple single-loop first-order methods, which we call perturbed S-EG and perturbed S-OGDA. By carefully choosing the regularization parameter and balancing the resulting regularization bias, stochastic error, and optimization error, we prove that both methods achieve a last-iterate convergence rate of \mathcalO(T^-1/4+\varepsilon) for any \varepsilon>0 in terms of the primal--dual gap. This improves upon the best-known \tilde\mathcalO(T^-1/7) rate under comparable constrained settings. Moreover, for unconstrained problems, we establish almost sure convergence under a single-timescale stepsize scheme, whereas prior almost sure convergence results typically rely on two-timescale stepsizes and are limited to S-EG.
Abstract: Diffusion models have become a widely used framework for probabilistic time series forecasting, modeling the distribution of future values given an observed history. In time series forecasting, however, the future continues the observed history, creating an asymmetry the standard diffusion process leaves unaddressed, with slowly-varying content largely determined by the observed continuity while higher-frequency dynamics carry most of the residual uncertainty. Existing diffusion-based forecasters decouple this asymmetry through an external rule before generation, leaving the corruption trajectory blind to which parts of the target the history can already anchor. We propose \textscDiffDiff, a diffusion framework that embeds this predictability asymmetry into the diffusion trajectory itself, so that a single end-to-end diffusion process becomes aware of which parts of the target the history can already anchor. \textscDiffDiff makes the forward operator step-dependent so that the noisy intermediate state progressively shifts from the target itself toward its second-order differenced structure, while a conditioning pathway supplies the denoiser with both value-domain and differential history information balanced by a stage-adaptive gate at each diffusion step. The terminal distribution approaches a standard Gaussian, preserving compatibility with existing samplers. On seven benchmarks across four prediction horizons, \textscDiffDiff outperforms six diffusion baselines, and our analysis confirms that \textscDiffDiff concentrates the diffusion's generative effort on the most uncertain components of the target while relieving it from rebuilding the history-anchored content.
Abstract: Large language model (LLM) agents increasingly operate over long and recurring external contexts, like document corpora and code repositories. Across invocations, existing approaches preserve either the agent's trajectory, passive access to raw material, or task-level strategies. None of them preserves what we argue is most needed for repeated same-context workloads: reusable orientation knowledge (e.g., what the context contains, how it is organized, and which entities, constants, and schemas have historically been useful) about the recurring context itself. We introduce PEEK, a system that caches and maintains this orientation knowledge as a context map: a small, constant-sized artifact in the agent's prompt that gives it a persistent peek into the external context. The map is maintained by a programmable cache policy with three modules: a Distiller that extracts transferable knowledge from inference-time signals, a Cartographer that translates it into structured edits, and a priority-based Evictor that enforces a fixed token budget. On long-context reasoning and information aggregation, PEEK improves over strong baselines by 6.3–34.0% while using 93–145 fewer iterations and incurring 1.7–5.8× lower cost than the state-of-the-art prompt-learning framework, ACE. On context learning, PEEK improves solving rate and rubric accuracy by 6.0–14.0% and 7.8–12.1%, respectively, at 1.4× lower cost than ACE. These gains generalize across LMs and agent architectures, including OpenAI Codex, a production-grade coding agent. Together, these results show that a context map helps long-context LLM agents interact with recurring external contexts more accurately and efficiently.
PaperID: 6646, Poster
Authors:
Jiacheng Huang, Long Tian, Yuan Liu, Andy W KhongAbstract: Millimeter-wave(mmWave) radar is an attractive modality for human sensing, offering all-weather, non-contact, and privacy-preserving perception. However, its inherent sparsity severely limits downstream human-centric understanding, and there is currently no public dataset that provides paired quantitative benchmarks for dense human point cloud reconstruction under occlusion. We present MIST, the first mmWave human point cloud dataset with paired quantitative benchmarks for through-obstacle dense human reconstruction. MIST captures synchronized clear-view and through-obstacle observations of the same subjects performing identical actions, separated by a physical barrier, and provides LiDAR-based 3D ground truth. We further incorporate a three-level occlusion design that converts attenuation severity into a controlled variable, covering eight subjects, thirty action classes, and 640~k paired frames. To establish a baseline on MIST and address the limitations of point cloud completion under extreme sparsity and human motion, we propose PRISM, a skeleton-guided conditional latent diffusion framework for reconstructing dense human point clouds from sparse mmWave radar alone. Three conditioning streams—skeleton joint positions, coarse body geometry, and action class embeddings—are incorporated to guide a Dynamic Transformer within a VP-SDE framework, enabling effective denoising and the reconstruction of mmWave human point clouds with near-LiDAR quality. On MIST, PRISM achieves a COV-CD of 0.362, significantly outperforming the strongest completion baseline (0.113). Notably, PRISM maintains structurally coherent reconstruction even under severe through-obstacle attenuation at approximately 50 points per frame, whereas all baselines collapse to fixed-template outputs.
PaperID: 6647, Poster
Abstract: Tree-based speculative decoding accelerates large language model inference by verifying multiple candidate branches in parallel. However, in practical serving scenarios, similar requests repeatedly induce overlapping candidate branches. While retrieval-based methods mitigate this by reusing historical drafts, the resulting candidate trees often contain many low-quality branches as the candidate tree grows, which reduces verification efficiency and wastes target-side computation. To address this, we propose ValuSpec, a plug-and-play candidate valuation mechanism that filters low-quality branches from a supplied candidate tree before expensive target-side verification. ValuSpec introduces transfer-based supervision to align filtering with the target model's behavior. A filtering threshold is estimated from the target model's greedy trajectories on a calibration split, without any parameter training. The target model then verifies the filtered tree under the output-identical greedy rule, so accepted tokens are guaranteed to match standard autoregressive decoding exactly. Experimental results on HumanEval, GSM8K, Dolly, and SpecBench show that ValuSpec achieves up to 2.38× end-to-end speedup in retrieval-based tree settings. Furthermore, integrating ValuSpec into the EAGLE-3 dynamic tree framework yields up to 2.78× end-to-end speedup without modifying its tree construction strategy, demonstrating that ValuSpec serves as a generally applicable and plug-and-play candidate valuation module.
Abstract: With LLM watermarking already being deployed commercially, practical applications increasingly require watermarks that encode more complex payloads, such as user IDs or timestamps, into the generated text. In this work, we propose a fundamentally new approach for multibit watermarking: introducing binomial encoding to directly encode every bit of the payload at every token position. We complement our approach with a that during generation dynamically redirects encoding pressure toward underencoded bits. Our evaluation against 8 baselines on up to 64-bit payloads shows that our scheme achieves superior message accuracy and robustness, with the gap to baseline methods widening in more relevant settings (i.e., large payloads and low-distortion regimes). At the same time, we challenge prior works’ evaluation metrics, highlighting their lack of practical insights, and introduce
Authors:
Yi Chen, Wonjin Shin, Shuhong Liu, Tho Mai, Jeongmo Lee, Jun Liu, Chuanbo Hua, Kun Wang, Joo-Young KimAbstract: Large foundation models (LFMs) achieve strong performance through scaling, yet current structural pruning methods derive fixed pruning decisions during inference, overlooking sparsity patterns that emerge in autoregressive token generation. We propose POP (Partition-guided Online Pruning), an efficient online structural pruning framework that enables context-conditioned dynamic pruning with minimal computational overhead. POP partitions model channels into retained, candidate, and pruned regions, where prefilling defines a coarse pruning partition, and decoding generates a fine-grained mask within the candidate region, avoiding full-channel re-evaluation. The coarse pruning partition preserves consistently important weights, while the fine-grained masking provides context-conditioned variation during decoding. Moreover, POP is a lightweight, plug-and-play method that requires no preprocessing, including offline calibration, retraining, or learning predictors. Extensive evaluations across diverse LFMs, including large language models (LLMs), mixture-of-experts models (MoEs), and vision–language models (VLMs), demonstrate that POP achieves competitive or superior accuracy over existing pruning approaches, particularly on generation tasks, while incurring minimal computational overhead and maintaining efficient inference. The source code will be publicly available.
PaperID: 6650, Poster
Abstract: Real-world problems span diverse domains, including program synthesis, symbolic reasoning, planning, and optimization, and often require fundamentally different solution paradigms. A central challenge for LLM-based reasoning is twofold: identifying an appropriate problem-solving approach and formulating it correctly. In this paper, we propose AlgoPilot, a cross-paradigm reasoning framework that enables adaptive selection of the solving approaches and structured problem formulation prior to execution. The AlgoPilot framework contains three components: 1) a SteerLM, trained via supervised fine-tuning (SFT) and reinforcement learning (RL), that adaptively selects the solving strategy (paradigm and modes) based on input problems, 2) the corresponding expert GuideLM, trained via SFT, that generates structured formulation guidance tailored to the selected strategy, and 3) an ExecutionLM that generates solutions conditioned on both the selected strategy and its formulation. Unlike prompting-based methods that rely on fixed reasoning templates, our approach learns to adaptively choose and structure problem-solving strategies based on the input. We build AlgoPilot on Qwen3-8B and evaluate across 5 problem categories and 13 benchmarks with 66 unique domains, comparing against both the latest similarly scaled small models and large LLMs. AlgoPilot significantly improves average accuracy for the same model from 38.4% to 72.9%, outperforming the best large LLM prompting baseline GLM-5 by 11.7%. Ablation studies confirm the effectiveness of each component. Notably, we show that the learned steering and formulation modules can transfer across execution models, suggesting their potential to improve other LLMs.
Abstract: We study the fundamental and timely problem of learning long sequences in autoregressive modeling and next-token prediction under model misspecification, measured by the joint Kullback--Leibler (KL) divergence. Our goal is to characterize how the sequence horizon (H) affects both approximation and estimation errors in this joint-distribution, sequence-level regime. By establishing matching upper and lower bounds, we provide, to our knowledge, the first complete characterization of long-horizon error behavior under the natural joint KL objective, with improved rates and optimality justification relative to existing work. On the approximation side, we show that joint KL admits a horizon-free approximation factor, in sharp contrast to Hellinger-based analyses that exhibit an (\Omega(H)) dependence for computationally efficient methods; this isolates the choice of divergence as the source of approximation amplification. On the estimation side, we prove a fundamental information-theoretic lower bound of order (\Omega(H)) that holds for both decomposable policy classes and fully shared policies, matching the (\widetilde O(H)) upper bounds achieved by computationally efficient algorithms. Our analysis clarifies the landscape of recent autoregressive learning results by aligning the log-loss training objective, the sequence-level evaluation metric, and the approximation metric through a sharp joint-KL oracle theory. We further show that these joint-KL guarantees imply policy learning regret bounds at rates matching prior imitation learning literature.
PaperID: 6652, Poster
Abstract: Text-to-image prompt optimization is often treated as prompt rewriting: given a user prompt, a language model expands it with richer style, composition, and photographic details. While this paradigm can improve visual appeal, it may silently weaken the user's actual intent, especially when the prompt specifies fine-grained objects, attributes, and relations. We argue that faithful prompt optimization requires a different objective: improving perceptual quality while preserving the compositional structure that makes the prompt semantically correct. We propose \ours, a single-step reinforcement learning framework that transforms noisy scene-graph feedback into a learnable prompt-editing policy. Instead of freely rewriting the entire prompt, \ours first extracts a calibrated scene graph from the input and then performs conservative deletion, reordering, and insertion actions conditioned on both text and graph representations. This design exposes structured intent to the optimizer while restricting unnecessary prompt drift. To make object--relation feedback usable for policy learning, \ours replaces sparse thresholded grounding rewards with uncertainty-aware smooth rewards and estimates advantages with environment-aware grouping over multiple image samples from the stochastic text-to-image generator. Across prompt-optimization datasets and generation backbones, \ours improves the faithfulness--aesthetics trade-off, with particularly consistent gains in relational fidelity where generic rewriting and aesthetics-oriented optimization often fail. We further evaluate with both human preference and Gemini-VQA, showing that the improvements are not confined to the SG reward used during training. Appendix additionally compares against a high-budget iterative GPT-4o structural-feedback baseline to contextualize the efficiency of amortized policy learning. More broadly, our results suggest that structural constraints should not remain only post-hoc evaluators of generated images; when properly calibrated and stabilized, they can be internalized as training signals for reliable, intent-preserving prompt policies.
PaperID: 6653, Poster
Abstract: Direct Preference Optimization (DPO) is widely used to reduce hallucination in Multimodal Large Language Models (MLLMs). Prior studies and our observations show that during DPO training, the log probabilities of chosen and rejected responses can decrease simultaneously, revealing the squeezing effect. For MLLMs, the redistributed probability mass may move toward responses driven by language priors rather than visual evidence, weakening the role of the image and amplifying hallucination. Existing squeezing-aware methods constrain rejected updates with \mathcalV-usable information or generation confidence, while methods with anchor terms regularize the DPO implicit reward of chosen responses. Neither explicitly grounds the control signal in visual evidence. We propose Visual Guided Direct Preference Optimization (VGDPO), which relaxes negative updates on rejected responses according to multimodal squeezing risk. VGDPO estimates this risk by combining token-level visual dependency with phrase-level hallucination localization, defining a hallucination ratio for each rejected response. We further apply a related reweighting principle to a visual contrastive objective, providing additional supervision for chosen responses under the original image. Experiments on four hallucination benchmarks and three MLLM backbones show that VGDPO improves DPO log probability dynamics and reduces multimodal hallucination while maintaining informative responses.
Authors:
Tianhao Hu, Xiangcheng Liu, Yuchun Miao, Youshao Xiao, Hong Yu Zang, Yang Zheng, huang xuan, Jinrui Ding, Yufei Zhang, Yu Yang, Yi-Kai Zhang, Yueqing Sun, Chengcheng Han, Xiandi Ma, Wei Wang, Qi GU, Yerui Sun, Yuchen Xie, Xunliang CaiAbstract: Asynchronous reinforcement learning (RL) has become a critical paradigm for accelerating large-scale LLM post-training in industrial settings, yet it faces a structural \emphlong-tail dilemma: rollout efficiency is bottlenecked by a small number of the longest trajectories, which are precisely the most valuable ones for RL training. Existing approaches alleviate this dilemma at the cost of either system overhead (e.g., re-prefill in partial-rollout methods) or algorithmic compromises (e.g., discarded long trajectories in replication-based methods). We trace these tradeoffs to a common assumption: the rollout cluster serves a single policy version at any moment. We propose DORA (Dynamic ORchestration for Asynchronous Rollout), which breaks this assumption by maintaining multiple policy versions concurrently within the rollout cluster. DORA combines three mechanisms: \emphmulti-version streaming training that decouples trajectory completion from batch boundaries, a centralized \emphload-balancing orchestrator that re-partitions resources across versions, and \emphzero-re-prefill migration that transfers KV-Cache directly across same-version instances. Experiments on open-source benchmarks show that DORA achieves up to 2.12× end-to-end throughput improvement and 8.2× rollout-stage acceleration over synchronous training while preserving convergence parity. In real-world production deployment with thousands of accelerators, DORA achieves up to 6.2× rollout speedup and produces competitive open-source LLMs.
PaperID: 6655, Poster
Abstract: Verbal memory retrieval carries rich diagnostic information for Alzheimer's disease (AD) through the order, hesitation, and categorical organization of recalled words, yet computational modeling of this retrieval process remains largely unexplored. A key challenge is that timing and hierarchical structure are diagnostically entangled: the diagnostic meaning of retrieval timing depends on its hierarchical context. We propose TrajLift, a spectral framework that models verbal recall as timed trajectories on a hierarchical graph. Our central mechanism is graph heat diffusion on the semantic hierarchy, which provides a unified operator for capturing both multi-scale hierarchical structure and continuous retrieval timing. Formal analysis shows that the resulting representation is provably separable into structural and temporal components with spectral selectivity across hierarchical levels. Experiments on synthetic benchmarks and a real-world clinical corpus demonstrate consistent improvements over baselines lacking joint structure-time modeling.
PaperID: 6656, Poster
Abstract: Companies that deploy generative language models on edge devices often need to train small models from scratch due to legal, data-governance, or product-customization requirements. An alternative approach is to train an enlarged model and compress it to the target size via structured pruning, but this advantage becomes unclear once the cost of pretraining the enlarged model is included. We introduce the enlarge-and-prune framework to study this question under full-cost accounting, and show that conventional enlarge-and-prune pipelines---which train the enlarged model, prune it, and recover in separate stages---do not consistently outperform direct small-model training under the same total token budget. To overcome this limitation, we propose IDEA Prune, an integrated enlarge-and-prune pipeline that unifies enlarged model pretraining, iterative structured pruning, and recovery under a single cosine annealing learning rate schedule. In controlled experiments compressing 2.8B models to 1.3B with up to 2T tokens on the same dataset, IDEA Prune consistently outperforms conventional pipelines and modestly improves overall performance relative to the scratch baseline, improving MMLU accuracy to 46.4%, compared with 31.4--33.4% for conventional pipelines. Our analysis further reveals that intermediate checkpoints provide more favorable pruning starting points than fully converged ones, and that an enlargement factor of approximately 2.6× yields the best pruned models in the tested regime.
PaperID: 6657, Poster
Abstract: Highlight detection in content-dense untrimmed videos, such as livestreams, is particularly challenging due to the prevalence of semantically similar moments and the application-driven demand of fitting a target duration budget. Existing methods typically rank segments independently based on semantic relevance, struggling to discriminate fine-grained quality among similar moments and allocating the budget to redundant selections. We therefore reformulate highlight detection as jointly selecting high-quality and complementary segments under a duration budget. To study this formulation, we introduce QB-Highlights, a large-scale benchmark of 50,000 e-commerce livestream videos with densely occurring query-relevant events and subtle quality variations. QB-Highlights is organized around two complementary dimensions: Segment Qualification, which assesses fine-grained per-segment quality; and Budgeted Composition, which evaluates set-level complementarity under a duration budget. It provides MLLM-assisted pseudo labels for training, dense human annotations for testing, and an evaluation protocol jointly measuring quality and diversity. We further propose a Qualification-then-Composition Reasoning (QCR) framework that performs structured deliberative reasoning to produce a highlight set matching the target duration. Extensive experiments show that QCR outperforms strong MLLM, moment retrieval, and highlight detection baselines on both segment quality and compositional diversity.
Abstract: Radiologist eye-tracking data provide a rich record of how experts search, compare, and accumulate evidence during image reading, yet existing methods exploit this signal only partially, either as a static spatial prior or as an auxiliary prediction target decoupled from diagnosis. We propose GazeWorld, a medical imaging world model that treats the image as the world and the radiologist's fixation sequence as a trajectory through it. GazeWorld autoregressively predicts the latent representation of the next fixated patch from all previously visited ones, while a spatial-completion branch covers unvisited regions. At inference, GazeWorld generates a sequence of patch representations from the image alone without requiring real gaze data. Frozen GazeWorld features achieve state-of-the-art diagnostic accuracy across all nine supervised settings on CheXpert, RSNA Pneumonia, and SIIM-ACR Pneumothorax, as well as the highest zero-shot accuracy on all three benchmarks. On the GazeSearch benchmark, a generic decoder trained on the same frozen features outperforms the purpose-built LogitGaze-Med by over 16% in ScanMatch and 22% in SED, despite not being explicitly trained to predict gaze. GazeWorld demonstrates that modeling how experts read, not just what they conclude, offers a promising pretraining paradigm for medical imaging AI.
Abstract: Building general-purpose role-playing agents that faithfully portray any character from a natural-language profile remains challenging. The dominant paradigm---supervised fine-tuning---encourages behavioral mimicry without deep, human-like internal thought processes, resulting in poor out-of-distribution generalization. Therefore, we propose Psy-CoT, a psychology-grounded chain-of-thought framework that decomposes pre-response reasoning into three role-specific steps---Interaction Perception, Psychological Empathy, and Logical Construction---so that the model thinks dynamically from the profile rather than merely mimicking surface patterns. While structured reasoning provides a foundation, it alone is insufficient; reinforcement learning is essential to further align the model with character fidelity. However, we observe that under LLM-based reward models, both generic phrases that hack the reward model and genuinely role-specific phrases receive identical gradient signals---this hacking accumulates over training, misleading the model into treating both as equally optimal choices. To address this, we propose Role-Aware Policy Optimization (RAPO), which uses profile--token mutual information to weight gradients asymmetrically---amplifying role-specific tokens under positive advantage while attenuating them under negative advantage. Experiments on CoSER, CharacterBench, and CharacterEval demonstrate that Psy-CoT outperforms existing role-playing CoT methods, and RAPO consistently surpasses GRPO across multiple model scales.
PaperID: 6660, Poster
Authors: Mingyu Kim, Doguk Kim
Abstract: Time series transformers have achieved strong forecasting performance, yet fundamental questions about their internal mechanisms remain unanswered: among self-attention and feed-forward networks (FFN), which component constructs temporal features? Why does a simple linear model outperform transformers? Why does patching help? We propose a temporal probing framework consisting of 15 token-level diagnostic tasks to trace information flow across transformer components. Applied to 5 architectures across 8 benchmarks, we find that temporal feature construction is not inherently tied to any single component—it emerges from the interaction between tokenization strategy and architectural design. In PatchTST, FFN dominates feature construction across all 8 datasets, with a 3.2× mean ratio over attention. Even when attention is completely removed, FFN's feature construction capability is preserved or enhanced, indicating that attention is not essential for temporal feature construction. Conversely, iTransformer's variate-level attention dominates over FFN, and CATS's cross-attention is responsible for nearly all feature transfer. Point-wise tokenization (P=1) structurally deactivates FFN, explaining why early transformers lagged behind linear models. These findings unify contradictory results across DLinear, PatchTST, attention-free models (TSMixer, PatchMLP), and CATS within a single framework.
Abstract: Activation patching is the primary tool in mechanistic interpretability. It attributes causal responsibility for a model behavior to each of its individual components by estimating its . Re-deriving the activation patching estimand from causal mediation analysis, we find that the NIE does not solely capture the causal effect through the specific component. It also contains that measure how much the component's causal effect itself depends on the state of other components in the model. A natural response may be to try to eliminate INT by adjusting the estimator or unit of analysis, but each of these potential remedies has predictable failure modes. We demonstrate these failure modes in the GPT-2 IOI circuit; components whose causal importance is conditional on the state of other components are either invisible or artificially inflated, and INT variance explains the previously documented instability of faithfulness scores. We prove that INT scales with the distance between clean and patched component activations, is negligible when the model is locally affine, and decomposes combinatorially into pairwise and higher-order group interactions. Despite its inevitability, INT is not a nuisance to be eliminated, but rather a diagnostic for interpretability studies. Its individual and group-level magnitude and sign signal when causal conclusions are prompt-dependent, and when greedy NIE-based component ranking will miss mechanisms only discoverable through combinatorial search.
Authors: Hoang Son Tran, Pranav Gupta, Rémi Bardenet, Subhroshekhar Ghosh
Abstract: Determinant point processes (DPPs) have recently emerged as a kernelized alternative to vanilla independent sampling for generating efficient minibatches, coresets and other parsimonious representations of large-scale datasets. While basic theoretical foundations and promising empirical performance have been demonstrated, there are two main challenges for current proposals for DPP-based coresets or minibatches. The first is the need for families of DPPs with certain key variance reduction properties, usually constructed in a continuous setting, of which there are few known examples. The second is the need for an ad-hoc construction of a discrete DPP defined on a given dataset, that inherits such variance reduction. In this work, we contribute to the programme of establishing DPPs as a subsampling toolbox for ML by advancing on these two fronts. First, we propose new DPPs on the Euclidean space based on wavelets, with provably better accuracy guarantees than the best known rates. Second, we introduce a general method to convert such continuous DPPs, which are more amenable to proving analytical statements, into discrete kernels, which are pertinent for subsampling tasks such as minibatch and coreset constructions. This conversion mechanism simultaneously preserves the desired variance decay and reveals a low-rank decomposition of the discrete kernel, which makes sampling the corresponding DPP computationally inexpensive. En route, we enlarge the class of ML tasks amenable to improvements via DPP-based minibatches and coresets to include highly non-smooth objective functions, with arbitrarily low regularity, and rate guarantees that adapt explicitly to this regularity.
Authors: Shourov Joarder, Diganta Sikdar, Ahsan H Akash, Binod Bhattarai, Prashnna K Gyawali
Abstract: Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning ability of LLMs, but often depends on external supervision from human annotations or gold-standard solutions. Reinforcement learning from internal feedback (RLIF) has recently emerged as a scalable unsupervised alternative, using signals extracted from the model itself. However, existing RLIF methods typically rely on a single internal reward, which can lead to reward hacking, entropy collapse, and degraded reasoning structure. We propose a multi-reward RLIF framework that decomposes the training signal into two complementary components: an answer-level reward based on cluster voting and a completion-level reward based on token-wise self-certainty. To combine these signals robustly, we apply GDPO-based normalization to reduce reward-scale imbalance. We further introduce KL-Cov regularization, which targets low-entropy token distributions responsible for disproportionate entropy reduction, preserving exploration and preventing late-stage collapse. Across mathematical reasoning and code-generation benchmarks, our method improves stability and robustness over prior unsupervised RL approaches, while achieving performance close to supervised RLVR methods. These results show that complementary internal rewards, combined with targeted regularization, can support stable long-horizon reasoning without relying on external ground-truth supervision.
PaperID: 6664, Poster
Abstract: Prior outdoor LiDAR registration works are commonly limited by the pair-wise input paradigm, which neglects temporal correlations inherent within streaming sequences. In this paper, we propose a novel multi-frame outdoor point cloud registration network for streaming LiDAR scans enhanced with pose-related memory buffers. The key observation is that long-term temporal LiDAR sequences can provide rich global contextual information to complete sparse measurements, filter outliers, and address low-overlap problems, thereby boosting registration performance. Specifically, two hybrid memory buffers are designed, including an implicit memory feature buffer and an explicit memory pose buffer, to store and dynamically update pose-related temporal features. Moreover, a novel dynamic history weighting module is developed to adaptively fuse current and history pose-related features. Extensive experiments on three outdoor datasets, including KITTI, nuScenes, and Apollo-Southbay, demonstrate state-of-the-art performance of MemReg, surpassing all previous pair-wise methods and also multi-frame SLAM systems. Our method also generalizes surprisingly well to multiview indoor registration scenarios with rather competitive performance on 3DMatch, 3DLoMatch, and ScanNet. Code will be released upon publication.
PaperID: 6665, Poster
Abstract: Robust finetuning of foundation models seeks to improve out-of-distribution (OOD) generalization during in-distribution (ID) adaptation. We revisit Mixout, a stochastic regularizer that intermittently replaces finetuned parameters with a reference anchor, and study its effectiveness for robust finetuning of vision foundation models. We reinterpret Mixout as a single-run, weight-sharing implicit ensemble and analyze its expected OOD error through a bias–variance–covariance–locality (BVCL) lens. This perspective reveals three key factors that govern the ID–OOD trade-off: the choice of masking anchor, the mask resampling frequency, and mask sparsity. Guided by this analysis, we propose GMixout, which generalizes Mixout in three ways. First, it explicitly controls the masking period through a resampling-frequency hyperparameter. Second, it replaces the fixed pretrained anchor with an exponential moving average snapshot that adapts throughout training. Third, it uses sparse kernels to update only a small subset of parameters at each iteration, introducing no inference-time overhead and enabling large-model finetuning on consumer-grade GPUs. Experiments across five vision benchmarks show that GMixout consistently outperforms Mixout, Model Soups, and strong parameter-efficient finetuning baselines under both covariate shift and class imbalance.
PaperID: 6666, Poster
Authors:
Yiting Duan, Longyan Tan, Hao Wu, Mehdi Tavakol, Yi GuoAbstract: Learning dynamics from sparse partial observations presents a fundamental dilemma: the pointwise mean squared error (MSE) training objective may induce local contraction and empirically lead to topological collapse, whereas matching the system's invariant measure preserves global geometry but lacks temporal ordering constraints. To resolve this, we introduce the Poincaré Return Induced Measure (PRIME) framework. PRIME bridges this gap by adopting a novel perspective of recurrence-level temporal organization, motivated by the return map and roof time structure of suspension flow representations. Instead of relying on unstable pointwise matching, PRIME refines the learned dynamics by leveraging recurrence-level temporal information, specifically by aligning the empirical graph measure of the induced return map over recurrence blocks. We evaluate our framework on the partially observed Lorenz63 system, sensor-observed Kuramoto-Sivashinsky equation, and flow past a cylinder. PRIME effectively reduces topological collapse across our benchmarks, improving temporal consistency while maintaining or enhancing long-term geometric and spectral fidelity.
PaperID: 6667, Poster
Abstract: Large language models are increasingly used as scientific assistants, yet we still lack reliable ways to evaluate whether they can investigate a system rather than merely answer a prompt. A useful evaluation must let agents choose experiments, interpret finite evidence, and make causal decisions, while keeping the environment coherent across repeated observations and interventions and preserving objective ground truth. We introduce ACED-BENCH (Active Causal Evidence and Decision Benchmark), a neuro-symbolic benchmark built around this requirement. The key idea is to separate the readable scientific surface from the formal causal world: LLMs generate natural-language scenarios and task text, while hidden probabilistic graphical models define the causal graph, data-generating process, intervention semantics, validation checks, and oracle answers. This yields scientific worlds that are natural for agents to investigate but exact enough for scalable, programmatic evaluation. The main leaderboard ACED-CORE contains 180 tasks: 60 set-valued causal ancestry/descendancy questions and 120 constrained causal-decision tasks. We also report ACED-QUAL, a 150-question yes/no qualification check using the same hidden-PGM interface. We evaluate seven contemporary LLMs and two earlier-generation models across three active agent designs. ACED-QUAL checks whether agents can use the interface to draw conclusions from evidence, while ACED-CORE tests whether they can decide what evidence to collect and when to commit to an answer. Evidence-ledger analysis localizes failures within the investigation: agents test the wrong intervention, lose relevant evidence, or continue after sufficient evidence has appeared. These results suggest that reliable AI co-scientists need explicit support for experiment design, evidence retention, answer commitment, and stopping.
PaperID: 6668, Poster
Authors: Tzu-Hsin Hsieh, Ricardo Marroquim
Abstract: Inserting objects into existing 3D scenes requires more than selecting a plausible location: The inserted object must also fit local geometry while preserving semantic intent and physical plausibility. Although recent Vision-Language Models (VLMs) and generative models enable semantic reasoning and visual content creation, they offer limited 3D grounding and geometric control when an inserted object must fit into constrained local spaces. We introduce ElasticFit, a VLM-guided framework for fit-aware object insertion centered on a novel scene-grounded representation. Given a language instruction and rendered scene observations, ElasticFit infers structured fitting cues that specify where the object should be grounded, what volume it should occupy, how it should be oriented, and its adaptation mode (rigid placement, uniform scaling, or elastic deformation). These cues convert high-level VLM reasoning into explicit 3D constraints that condition object generation and guide downstream geometric fitting. ElasticFit then generates a scene-conditioned object prior, reconstructs it in 3D, and refines the mesh through mode-specific fitting while enforcing collision avoidance, contact consistency, and physical grounding. In fixed-asset baseline comparisons, ElasticFit improves spatial relation success from 50.8% to 69.7% and support success from 48.3% to 91.7% over the strongest baseline, while providing novel support for generative "make-it-fit" insertions in complex scenarios.
PaperID: 6669, Poster
Authors: Tridib K. Biswas, Ruben Coen-Cagli
Abstract: Neural dynamics are a central feature of biological neural systems. In the visual cortex, the functional role of neural dynamics is poorly understood, particularly for static stimuli. The statistics of still natural images have long been used by models that successfully predict firing in the early visual cortex. Those models assume that neurons represent inferences about latent image features, but they often ignore the dynamics of said inference. Here, we test if inferential dynamics for images explain visual-cortical dynamics. To do so, we present a biologically plausible model of cortical computation that assumes the inference of features is coupled with inference about how those features are organized into meaningful parts, also known as segmentation. Given an input image, our model iterates between inferring the features and the segments, through recurrence, thus inducing dynamics in the feature representation. We make theoretical and exact predictions for the induced neural dynamics. Our model predicts both gain and variability in evoked responses to natural images, and we show that it captures classical single-neuron experimental findings such as gain dynamics and variability decay during stimulus presentation. The model also predicts that those metrics are more heterogeneous than previously thought, reflecting the complexity of inference for natural scenes, and that their population-level organization reflects the inferred segments. Finally, our analytical formulation provides a clear interpretation of those results. Insights from our normative model could generalize to segmentation problems in other experimental modalities, and help address the cost of inference in artificial networks.
PaperID: 6670, Poster
Authors: Chung-Yiu Yau, Haoming Liu, Hoi-To Wai
Abstract: This paper initiates the study of a stochastic approximation approach with the Fully Stochastic Primal Dual Algorithm (FSPDA) framework for decentralized optimization on random and time varying topologies. Our framework relies on a novel observation that randomnesses in time varying topology can be incorporated into a stochastic equality constrained optimization formulation. We derived two new algorithms supporting sparsified communication on time varying topologies --- FSPDA-SA allows agents to execute multiple local gradient steps to accelerate convergence, and FSPDA-STORM further incorporates variance reduction to improve sample complexity. For problems with smooth (possibly non-convex) objective function, within T iterations, FSPDA-SA (resp. FSPDA-STORM) finds an \mathcalO( 1/\sqrtT )-stationary (resp. \mathcalO( 1/T^2/3 )) solution. The latter shows the first near-optimal convergence rate over time varying topology.
PaperID: 6671, Poster
Abstract: Structure learning for sparse graphical models in high dimensions requires both computational scalability and selection consistency. Classical criteria such as (Extended) Bayesian Information Criterion can provide consistency, but rely on model-specific likelihood and complexity calibration. In comparison, cross-validation avoids such analytic calibration and is therefore attractive for graph learning. However, cross-validation is computationally infeasible and theoretically inconsistent in structure learning. In this work, we resolve the tension by developing a scalable graph-recovery procedure built on approximate cross-validation. We demonstrate its effectiveness in learning large graphical models both theoretically and empirically, and establish selection consistency under mild signal strength assumptions.
Authors:
Chengqian Zhang, Yucheng Jin, Duo Zhang, Tiejun Li, Han WangAbstract: Crystal generative models mainly learn what stable crystals look like, with little explicit supervision for what makes them stable. We reveal a substantial representation gap between state-of-the-art crystal generative models and pretrained universal machine learning interatomic potentials (MLIPs) via energy probing, and show this gap can be closed by a simple training-time alignment. We propose Crystal REPresentation Alignment (CrystalREPA), a plug-and-play framework that aligns the atom-wise hidden states of generative encoders with frozen MLIP representations through an element-aware contrastive objective, transferring stability-aware atomistic priors with marginal training overhead and no additional inference cost. Across three generative frameworks, ten MLIP teachers, and two benchmark datasets, CrystalREPA consistently improves the thermodynamic stability, structural validity, and structural fidelity of generated crystals. Equally important, we find that an MLIP's transfer effectiveness is poorly predicted by its accuracy on standard leaderboards (e.g., Matbench Discovery) but strongly predicted by the distinguishability of its atom-wise representation space, yielding a practical, accuracy-independent criterion for selecting MLIP teachers for generative transfer.
PaperID: 6673, Poster
Abstract: Recent advances in generative video models have significantly improved visual realism, making detection increasingly dependent on subtle temporal inconsistencies arising from imperfect frame-to-frame coherence. However, these signals are weak and spatially localized, while most of the video content is redundant. This creates a fundamental mismatch: existing detectors rely on heavy backbones to model full videos, assuming increased capacity can capture such cues, which leads to high computational cost and limited scalability. In this work, we rethink AI-generated video detection from a representation perspective. Instead of modeling entire videos, we propose to explicitly restructure them into compact temporal slice representations that isolate informative temporal evolution. Based on this idea, we introduce a lightweight framework using frequency-aware temporal slice consistency learning, where informative spatial rows and columns are aggregated over time to form slice-domain inputs. This reparameterization suppresses redundant appearance while preserving discriminative temporal dynamics, enabling efficient and targeted detection. By constructing slices from both high- and low-frequency regions, our method exposes complementary temporal artifacts, including unstable details and abnormal motion smoothness. We instantiate this idea in a Temporal Slice Consistency Network, which integrates task-driven slice localization, hierarchical positional encoding, and a lightweight Transformer to model cross-slice and cross-frequency dependencies with minimal overhead. We further introduce AI-Artist, a new benchmark with artist-curated videos from recent high-fidelity generators, including Seedance and Kling. Experiments on GenVideo and AI-Artist show that our method achieves strong performance while requiring over 400× fewer FLOPs and 3.6× faster inference than the strong video-based baseline ReStraV, enabling scalable and practical deployment.
PaperID: 6674, Poster
Abstract: Large language models are increasingly used to simulate human participants in social and behavioral studies, yet static persona prompting typically maps a participant profile and an experimental scenario directly to a response, entangling stable dispositions with situation-specific interpretations. To address this limitation, we introduce SPIN, a cognitive-affective personality system-inspired inference pipeline for social judgment and decision alignment. Specifically, SPIN implements this structured inference process through three zero-shot LLM calls that compile a task-blind participant core, elicit condition-specific cognitive-affective states, and read out decisions from those states, thereby reusing stable personality structure while routing each trial-specific response through an explicit state representation. We evaluate SPIN on two reconstructed social-psychological study families spanning uncertainty reasoning and pluralistic ignorance, across four base LLMs. Compared with blank, demographic, narrative, and chain-of-thought prompt variants, SPIN consistently delivers the strongest overall alignment performance across base LLMs and study families. Ablations and state analyses further show that both personality compilation and structured state elicitation contribute to the gains, and that the elicited states shift interpretably across informational and normative conditions. These results suggest that structured personality-state inference can improve benchmark-level behavioral alignment beyond richer persona descriptions or generic multi-step reasoning.
Abstract: Machine unlearning aims to remove targeted data or behaviors from a trained model without retraining from scratch. Yet most evaluations assume that the examples to forget are already known. In realistic language-model deployments, a requester may ask a model to stop reproducing a song or book without knowing which spans, documents, quotations, or near-duplicates in a trillion-token corpus support that behavior. We study this missing upstream problem, , a benchmark for verbatim output suppression over songs and books, with model-specific extraction profiles, content-grounded QA, and capability-retention evaluations. CleanSlate exposes two failure modes. Natural lexical and exact-substring curators often yield forget sets that lead to weak suppression. An evaluation-aware curator suppresses requested continuations almost completely, but causes collateral regression on non-requested content and model-dependent capability loss. These results show that practical unlearning is not only an optimization problem once a forget set is given: the data chosen for forgetting determines both what can be unlearnt and what else is damaged.
Authors:
Zhiheng Liu, Weiming Ren, Xiaoke Huang, Shoufa Chen, Tianhong Li, Mengzhao Chen, Yatai Ji, Sen He, Jonas Schult, Belinda Zeng, Tao Xiang, Wenhu Chen, Ping Luo, Luke Zettlemoyer, Yuren CongAbstract: Unified multimodal models typically rely on pretrained vision encoders and use separate visual representations for understanding and generation, creating misalignment between the two tasks and preventing fully end-to-end optimization from raw pixels. We introduce Tuna-2, a native unified multimodal model that performs visual understanding and generation directly based on pixel embeddings. Tuna-2 drastically simplifies the model architecture by employing simple patch embedding layers to encode visual input, completely discarding the modular vision encoder designs such as the VAE or the representation encoder. Experiments show that Tuna-2 achieves state-of-the-art performance in multimodal benchmarks, demonstrating that unified pixel-space modelling can fully compete with latent-space approaches for high-quality image generation. Moreover, while the encoder-based variant converges faster in early pretraining, Tuna-2's encoder-free design achieves stronger multimodal understanding at scale, particularly on tasks requiring fine-grained visual perception. These results show that pretrained vision encoders are not necessary for multimodal modelling, and end-to-end pixel-space learning offers a scalable path toward stronger visual representations for both generation and perception.
PaperID: 6677, Poster
Authors:
guojun lei, Hong Li, Hongbing Yang, Lixue Gong, Chi Wang, Zheng DongAbstract: Recent advances in image editing have enabled fine-grained and highly controllable visual modifications.However, achieving similar levels of controllability in video editing remains an open challenge, particularly for precise local edits, dynamic temporal effects, and consistent lighting across frames.In this paper, we propose EditDistill, a framework that distills image-level visual transformations from edits into a compact descriptor to guide video editing.Specifically, given a source video and an edit prompt, we first generate an edited version of the initial video frame using an off-the-shelf image editing model. We then distill the visual transformation of this original-edited image pair into a compact representation termed an edit descriptor.This descriptor is injected into a video diffusion model through our lightweight fusion module, which steers the denoising process toward the desired edit. To align the video trajectory with the demonstrated image edit while preserving motion and layout, we train with edit-descriptor alignment and DINO structural consistency losses, and derive an equivalent inference-time guidance formulation under flow matching.Extensive experiments demonstrate that our method achieves strong performance on standard benchmarks and produces high-quality video editing results across diverse scenarios.
PaperID: 6678, Poster
Abstract: When is it better to rely on information from others rather than experiment to discover the answer yourself? Social learning refers to the ability of humans to learn by observing, imitating, or interacting with others. In this work, we use multi-agent reinforcement learning (MARL) to investigate the conditions under which social learning achieves higher returns and improved learning efficiency relative to individual learning. We begin with a theoretical analysis and discover that, under coordinated group exploration and low skill transmission costs, social learning achieves lower sample complexity than individual learning. Motivated by this insight, we propose Selective Social Learning (SSL), a novel social learning algorithm that reduces the cost of skill transmission, and we design experiments across three MARL environments to test these predictions. Empirical results show that SSL outperforms baselines under conditions consistent with our theory. Surprisingly, we also find that in non-stationary environments with sparse rewards, which mirror the real world, social learning can emerge in standard MARL without an explicitly designed social learning mechanism. This suggests that the apparent absence of social learning in standard RL stems from limitations in training environment design rather than in RL itself. We therefore call for future MARL training environments to inherently incorporate sparsity and non-stationarity, enabling agents to naturally develop social-learning behaviors that transfer to complex real-world settings.
Authors:
Alexandra Souly, Javier Rando, Ed Chapman, Xander Davies, Burak Hasircioglu, Ezzeldin Shereen, Carlos Mougan, Vasilios Mavroudis, Erik Jones, Chris Hicks, Nicholas Carlini, Yarin Gal, Robert KirkAbstract: Poisoning attacks can compromise the safety of large language models (LLMs) by injecting malicious documents into their training data. Existing work has studied pretraining poisoning assuming adversaries control a percentage of the training corpus. However, for large models, even small percentages translate to impractically large amounts of data. This work demonstrates for the first time that poisoning attacks instead require a near-constant number of documents regardless of dataset size. We conduct the largest pretraining poisoning experiments to date, pretraining models from 600M to 13B parameters on chinchilla-optimal datasets (6B to 260B tokens). We find that 250 poisoned documents similarly compromise models across all model and dataset sizes, despite the largest models training on more than 20 times more clean data. We also run smaller-scale experiments to ablate factors that could influence attack success, including broader ratios of poisoned to clean data and non-random distributions of poisoned samples. Finally, we demonstrate the same dynamics for poisoning during fine-tuning. Altogether, our results suggest that injecting backdoors through data poisoning may be easier for large models than previously believed as the number of poisons required does not scale up with model size—highlighting the need for more research on defences to mitigate this risk in future models.
PaperID: 6680, Poster
Abstract: Existing convergence analyses of finite-width overparameterized deep neural networks under standard parameterization often rely on parameters remaining close to their initialization and typically neglect the role of noise, despite empirical evidence that stochasticity can improve optimization and generalization. To bridge this gap, we study L_2-regularized continuous-time Langevin dynamics~(CLD) for training a feedforward neural network with smooth activations. We introduce a change-of-measure inequality, inspired by the second law of thermodynamics, that enables control of key properties of the training dynamics under the time-evolving parameter distribution. Leveraging a sharpened lower bound on the minimum eigenvalue of the Neural Tangent Kernel (NTK) at initialization, we show that injected noise helps preserve NTK stability with high probability without confining the weights to a small neighborhood of their initialization. As a result, we derive a non-asymptotic upper bound on the expected training loss. Our analysis yields explicit conditions on width and noise temperature that guarantee convergence to a finite loss floor. Finally, we experimentally show the network Jacobian can change from initialization, departing the commonly analyzed lazy regime.
PaperID: 6681, Poster
Abstract: Agents based on large language models are increasingly expected to autonomously handle negotiation and transactions. However, existing benchmarks predominantly focus on text-only, bilateral, zero-sum price haggling, failing to capture the complexity of real-world commerce. To address this gap, we introduce Agenticpay, a unified framework and benchmark for evaluating how well multimodal agents reach high-welfare, multi-dimensional agreements in realistic markets. Built around four core components ( Environments, Tasks, Agents, and Metrics), Agenticpay comprises 160 multimodal tasks spanning 4 real-world business scenarios (E-commerce, Food Delivery, Ride-hailing, and Apartment Rental) and 8 market topologies, scaling from 1-to-1 bargaining to many-to-many (N-to-N) competitive markets. Beyond price haggling, agents must read product images, infer each opponent's hidden preferences, and trade off multiple binding contract terms (e.g., price, lease duration, return policy, delivery speed). Agents communicate through multi-round natural-language dialogue, with each turn proposing or revising a full contract, and outcomes are scored by a utility-based framework that rewards agreements maximizing joint welfare. Evaluations on state-of-the-art proprietary (GPT-5.4, Claude Sonnet 4.6, Gemini 3.1 Pro Preview) and open-weight (Qwen3-VL-32B-Instruct, InternVL3-38B) multimodal models reveal substantial gaps in non-zero-sum value creation and market reasoning: even the strongest agent (Gemini 3.1 Pro Preview) reaches a GlobalScore of only 42.3/100, and performance consistently degrades as markets scale from bilateral to multi-sided, with the average GlobalScore dropping by 5.7 points and open-weight models suffering the largest declines (up to 11.5 points). These findings establish Agenticpay as a foundational testbed for multimodal agentic commerce. Code and dataset are available at: https://anonymous.4open.science/r/AgenticPay-4BBB/README.md.
Abstract: We present ROCKET, a training-free model compression method that achieves state-of-the-art performance in comparison with factorization, structured-sparsification and dynamic compression baselines. Operating under a global compression budget, ROCKET comprises two key innovations: First, it formulates layer-wise compression allocation as a multi-choice knapsack problem, selecting the optimal compression level for each layer to minimize total reconstruction error while adhering to a target model size. Second, it introduces a single-step sparse matrix factorization inspired by dictionary learning: using only a small calibration set, it sparsifies weight coefficients based on activation-weights sensitivity and then updates the dictionary in closed form via least squares bypassing iterative optimization, sparse coding, or backpropagation entirely. ROCKET consistently outperforms existing compression approaches across different model architectures at 20–50% compression rates. Notably, it retains over 90% of the original model’s performance at 30% compression without any fine-tuning. We will release the code implementing ROCKET upon paper acceptance.
Authors:
Jiayao Wang, Yang Song, Zhendong Zhao, Jiale Zhang, Xiaoying Lei, Qilin Wu, Wenliang Yuan, Junwu Zhu, Dongfang ZhaoAbstract: Federated self-supervised learning (FSSL) enables collaborative training of self-supervised representation models without sharing raw unlabeled data. While it serves as a crucial paradigm for privacy-preserving learning, its security remains vulnerable to backdoor attacks, where malicious clients manipulate local training to inject targeted backdoors. Existing FSSL attack methods, however, often suffer from low utilization of poisoned samples, limited transferability, and weak persistence. To address these limitations, we propose a new backdoor attack method for FSSL, namely Hallucinated Positive Entanglement (HPE). HPE first employs hallucination-based augmentation using synthetic positive samples to enhance the encoder’s embedding of backdoor features. It then introduces feature entanglement to enforce tight binding between triggers and backdoor samples in the representation space. Finally, selective parameter poisoning and proximity-aware updates constrain the poisoned model within the vicinity of the global model, enhancing its stability and persistence. Experimental results on several FSSL scenarios and datasets show that HPE significantly outperforms existing backdoor attack methods in performance and exhibits strong robustness under various defense mechanisms.
PaperID: 6684, Poster
Abstract: Large Language Model-based Multi-Agent Systems (LLM-based MAS) have shown remarkable success in solving complex tasks by coordinating specialized agents through multi-step workflows. To overcome the high cost and specificity of manually designed workflows, recent research has shifted toward learning them automatically. However, a key limitation remains: while agent calls introduce significant latency due to model inference time and service congestion, most automated methods optimize solely for downstream performance, ignoring execution time (makespan). This issue is amplified when workflows are reused across recurring queries, making latency accumulate over time. This limits deployment in recurrent applications such as AI customer support and AI cloud troubleshooting, where users expect fast responses. To bridge this gap, we introduce a new problem, termed latency-aware workflow generation, where the objective is to minimize workflow makespan while aiming to retain downstream performance. We propose LAWA (Latency-Aware Workflow optimizAtion), a novel framework that reduces makespan via theoretically grounded graph edits targeting critical-path bottlenecks. By unifying a graph generator with this graph-editing procedure, LAWA supports both de novo workflow generation and refinement of existing ones. Empirical results across seven benchmarks demonstrate that LAWA significantly reduces makespan while achieving competitive task performance, supporting efficient reuse of multi-agent workflows for recurrent queries.
PaperID: 6685, Poster
Abstract: Video Emotion Recognition (VER) aims to identify and understand the emotional states of characters by analyzing visual and auditory information in videos. However, conventional rigid frame sampling strategies for long videos tend to fragment continuous emotional expressions. Moreover, existing fusion methods often fail to properly model temporal dependencies and may suffer from future information leakage. To address these limitations, we propose SCG-HF, a novel framework that preserves semantic consistency through adaptive grouping and hierarchical fusion. Specifically, we introduce Semantic Consistency Grouping (SCG), which dynamically segments videos into semantically coherent units based on feature similarity, thereby maintaining the integrity of emotional patterns. Furthermore, we design a three-level Hierarchical Fusion (HF) architecture to capture emotional dynamics at different temporal granularities: (i) frame-level refinement enhances subtle local emotional cues; (ii) segment-level history-guided fusion models temporal evolution across segments while strictly preventing future information leakage via attention; and (iii) sample-level contrastive alignment synchronizes audio and visual representations in a shared latent space. Extensive experiments on the VideoEmotion-8 and Ekman-6 benchmarks demonstrate that SCG-HF achieves state-of-the-art performance.
PaperID: 6686, Poster
Abstract: Raindrop removal is challenging because raindrops exhibit strong spatial non-uniformity and scene dependency. Although previous works have demonstrated that dual-pixel (DP) sensors can facilitate raindrop removal, their effectiveness in complex real-world scenarios remains limited due to local modeling paradigms and insufficient coverage of real data. In this work, we collect a large-scale real-world DP raindrop dataset that covers day/night scenes, diverse light conditions, and varying raindrop intensities. Based on this dataset, we establish a challenging benchmark for DP raindrop removal in real-world scenarios. Building upon this dataset, we develop DRWKV, which models raindrop removal as a degradation-conditioned state evolution process. DRWKV explicitly injects raindrop degradation information derived from DP sensors into the recurrent state update process, enabling the model to adaptively balance information preservation and content restoration across spatial regions while maintaining linear computational complexity. This design facilitates effective global modeling of spatially non-uniform raindrop degradations. Extensive experiments demonstrate that the proposed DRWKV outperforms prior methods on the proposed benchmark as well as existing datasets.
Authors:
Hei Yi Mak, Shadan Golestan, Hoang Le, Mehran Taghian Jazi, Yunke Peng, Yaoyuan Wang, Yao Wang, JUNSONG WANG, Tianchi Hu, Fengchen He, Guipeng Hu, Tanzila Rahman, Anandharaju D RajuAbstract: We present, to our knowledge, the first end-to-end FP4 RL post-training, in which both the rollout and training policies, including their forward and backward passes, operate at 4-bit precision. A systematic study reveals that the dominant source of degradation in FP4 RL is not training-side quantization error but rollout activation quantization: outliers stretch the dynamic range so far that a large number of activation values underflow to zero under FP4. Counterintuitively, restoring the training policy to higher precision while keeping the rollout in FP4 makes accuracy worse than full FP4 baseline, exposing rollout–training mismatch as the principal failure mode and ruling out standard pretraining-style fixes. We address this with Rollout Residual Quantization (Rollout-ResQ): a single residual correction term constrained to a hardware-friendly sparsity pattern, added only to the FP4 rollout matmul — a lightweight correction that recovers most of the precision lost to outlier-driven underflow without inflating the rollout’s compute footprint. On Qwen2.5-3B and Qwen2.5-Math-7B, Rollout-ResQ paired with the HiFloat4 (HiF4) format — whose three-level hierarchical scaling preserves resolution under FP4’s tight 4-bit budget — closes the accuracy gap to BF16 from 4.9% to 1.1%, bringing fully quantized FP4 RL within striking distance of full precision. Applied to the open-standard MXFP4, the same recipe narrows the gap from 13.6% to 5.3%, revealing that FP4 format choice is a key factor that determines the ceiling on recoverable accuracy. Together, these results establish HiF4 as the enabling format for end-to- end FP4 RL post-training, and Rollout-ResQ as the activation-side mechanism that makes the gap to BF16 closable.
Abstract: We envision a new era of AI, termed agentic organization, where agents solve complex problems by working collaboratively and concurrently, enabling outcomes beyond individual intelligence. To realize this vision, we introduce asynchronous thinking (AsyncThink) as a new paradigm of reasoning with large language models, which organizes the internal thinking process into concurrently executable structures. Specifically, we propose a thinking protocol where an organizer dynamically assigns sub-queries to workers, merges intermediate knowledge, and produces coherent solutions. More importantly, the thinking structure in this protocol can be further optimized through reinforcement learning. Experiments demonstrate that AsyncThink achieves 28% lower inference latency compared to parallel thinking while improving accuracy on mathematical reasoning. Moreover, AsyncThink generalizes its learned asynchronous thinking capabilities, effectively tackling unseen tasks without additional training.
PaperID: 6689, Poster
Authors:
Yuxi Xue, Xiangzheng Zhou, Xiaobo Chen, Jianjun Qian, Jian YangAbstract: Multi-agent trajectory prediction is crucial for ensuring autonomous driving safety. Although diffusion models show great potential in trajectory refinement, existing diffusion-based refiners suffer from severe exposure bias. This bias stems from the distribution mismatch between the noisy training endpoint and the noise-free inference starting point, which frequently causes negative optimization during the refinement process. To resolve this issue, this paper proposes the Diffusion Schrödinger Bridge Trajectory Refiner (DSBTR). By introducing a tractable Schrödinger bridge with dual-terminal constraints, DSBTR ensures that the forward evolution endpoint strictly degenerates into a deterministic and noise-free coarse prediction prior. This theoretical alignment eliminates the initial boundary error and constrains the reverse solving process to a strict residual approximation. Based on the semi-linear structure of DSBTR, we derive a tailored high-order DPM-Solver to accelerate the sampling process and meet autonomous driving application requirements. Extensive experiments on the ETH/UCY and SDD benchmarks demonstrate that DSBTR serves as an effective plug-and-play module. It consistently improves the prediction accuracy of various baseline architectures and successfully avoids negative optimization.
PaperID: 6690, Poster
Abstract: Multi-agent debate can improve reasoning, yet often fails to beat simple majority voting. Prior martingale-null theory explains such failures as debate without truth-directed signal. We develop a more complete account of when practical LLM debate succeeds: it needs both correct minority proposals and a readout that can recover them. We formalize this view through latent headroom, the gap between majority vote and proposal-oracle performance, which measures unused correct expertise in an expert society. For readout, we developLatent Verification Debate (LVD), a theory that models candidate proposals as receiving latent verified evidence before influencing final generation. For supply, we study neural-thicket expert societies, where nearby model perturbations provide heterogeneous specialists, and introduce a coverage-based construction that increases complementary proposal supply. Across matched-budget reasoning and multi-discipline benchmarks, our construction increases recoverable headroom, and controlled readout experiments show that gains arise from recovering surfaced correct minorities rather than merely adding interaction. Together, these results identify diverse proposal supply and verification-aware utilization as the two mechanisms that determine when debate helps.
Abstract: Deep search agents operate over trajectories spanning dozens of steps, yet standard reinforcement learning provides only a single outcome reward per trajectory — supervision that is far too sparse for effective credit assignment. On-policy self-distillation (OPSD) addresses this by using the model's own logits as dense token-level teachers, but extending it to search agents introduces a fundamental tension: the teacher, having access to privileged information such as the correct answer, produces a distribution that differs systematically from the student's exploration-based reasoning, and naive distillation causes the student to inherit this information asymmetry rather than learn better search strategies. We resolve this tension through two contributions. First, we construct Evidence Anchors — concise, step-level evidence snippets extracted from the web — as privileged information that captures key reasoning steps without revealing the entire answer path. Second, we propose Step-Level Self-Distilled Policy Optimization (SSPO), which converts teacher–student disagreement into step-level advantage weights within GRPO, applied exclusively to incorrect trajectories. This design decouples what to update from how much to update: the outcome reward determines the direction of policy change, while the teacher modulates its magnitude at each step. Correct trajectories are left untouched, preserving their diversity. On Qwen3-8B, SSPO consistently outperforms GRPO across BrowseComp, GAIA, and FRAMES, surpassing or matching GRPO trained with twice as many gradient steps while adding only about 5% overhead per step from a single additional forward pass.
PaperID: 6692, Poster
Authors: Shiyan Liu
Abstract: Submission timing in human evaluation systems is universally treated as an irrelevant detail. We identify a systematic temporal leniency bias: across 129,023 reviews spanning three consecutive years of a large-scale ML venue, evaluators submitting closer to the deadline assign higher within-paper scores and produce less thorough assessments. The pattern holds in all 15 robustness specifications across all three years, and early evaluators achieve measurably lower Brier scores (\Delta \approx 0.005--0.006, stable across years). We propose two complementary corrections. Optimal Timing-Calibrated Aggregation (OTCA) derives a linear weight function w(t) = a^\star t + 1 by minimising paper-level decision loss, significantly outperforming heuristic alternatives (McNemar \chi^2 = 46.3, p < 0.001). Adversarial Temporal Debiasing (ATD) learns timing-invariant representations via Gradient Reversal, reducing timing discriminability by 37.4% while preserving predictive utility. Combined in a joint scheme, the two methods improve borderline decision accuracy by up to 11.25% over uniform averaging (McNemar p < 0.001), consistently across all three years. Both methods require only submission timestamps, already logged by every major evaluation platform, making deployment cost-free.
PaperID: 6693, Poster
Abstract: Long-term time-series forecasting (LTSF) underpins applications such as energy-load prediction and traffic-flow forecasting, yet model performance often deteriorates as the horizon extends. In particular, errors accumulate with lead time and systematically amplify toward the end of the predicted sequence, a phenomenon known as tail deterioration. We posit that this instability is driven by weakly constrained fusion of cross-temporal and cross-scale context: aggressive mixing can propagate noise and early-stage bias into distant horizons, ultimately degrading tail accuracy. To address this issue, we propose TailFix, a controllable cross-scale mixing framework for long-horizon forecasting. TailFix selectively gates interactions among multi-scale representations, derives scale-level correction terms from contextual summaries, and injects them via gated residual connections to suppress error diffusion and improve long-horizon stability. We further incorporate adaptive channel enhancement and a lightweight tail-correction module to explicitly refine the error-prone tail region. Finally, we introduce the Tail Amplification Factor (TAF) to quantify tail error amplification relative to the early portion of the sequence. Across diverse benchmarks, TailFix consistently mitigates tail error amplification and improves performance on multiple long-horizon forecasting tasks.
PaperID: 6694, Poster
Abstract: Multimodal Recommender Systems (MRSs) have achieved significant success in personalized recommendations by integrating diverse modalities such as images, text, and audio. While in complex real-world industrial scenarios where MRS has usually been deployed, the presence of missing modalities is very common. Most existing MRS methods recover missing data by disentangling available modality features into general and specific representations, thereby generating missing modalities based on personalized aligned feature. However, these methods rely on implicit constraints within pre-extracted multimodal feature spaces, which usually introduce noise that prevents the effective decoupling of general and specific semantics, thereby diminishing recommendation accuracy. Moreover, unguided linear transformations of available modality features for generation lead to semantic hallucinations, which inevitably degrade the ranking quality of the recommendation results. To address these issues, we propose the Frequency-Decoupled Knowledge Distillation Framework for Multimodal Recommendation (FreDRec), which conducts effective and robust missing feature generation. Specifically, motivated by the physical characteristics of multimodal signals, FreDRec utilizes Fourier and Discrete Wavelet Transforms to explicitly decouple raw multimodal signals into low-frequency core semantics and high-frequency details, which is mathematically proven to effectively minimize noise introduction at the source. Next, to ensure robust feature generation, we pre-train a teacher model based on decoupled multi-view graph propagation to guide a student model via multi-level knowledge distillation, which establishes an end-to-end highly efficient lightweight architecture and achieves precise missing modality generation with expert collaborative bounds. Finally, extensive experiments on multiple public benchmark datasets demonstrate that FreDRec achieves superior performance under various modality missing ratios.
Abstract: Diffusion and flow-matching models scale because pretraining simply regresses against a closed-form target built by analytically noising clean data samples. In many settings, only a reward function is available, scoring how desirable a generation is, so training must proceed online from a pretrained model. Existing methods either rely on costly SDE rollouts, sometimes with reward gradients, or adopt heuristic reward-dependent variants of the pretraining objective. Under a stochastic optimal control formulation of KL-regularized reward maximization, the optimal generative process tilts only the clean-endpoint distribution and leaves the conditional noising law unchanged. Combining this path-space characterization with the adjoint-matching optimality condition and a REINFORCE estimator to avoid reward gradients, we derive Reinforce Adjoint Matching (RAM). At each step, we draw a clean endpoint from the current model with any off-the-shelf sampler, evaluate its reward, noise it to several training states, and regress against a closed-form, reward-weighted target. Like the pretraining objective, RAM is simple and scales. On Stable Diffusion 3.5M, RAM achieves the highest reward on composability, text rendering, and human preference. It is much more training efficient, reaching Flow-GRPO's peak GenEval accuracy in~50× fewer training steps.
PaperID: 6696, Poster
Abstract: Bundle construction (BC) is a critical technique in recommender systems research, which aims to select a coherent subset of items from large-scale item catalogs to either construct a complete bundle from scratch or complete a partial bundle with missing items. However, the extremely high-dimensional combinatorial spaces and heterogeneous semantic granularities involved pose significant challenges to BC benchmarking, and existing generative frameworks suffer from failing to precisely align the generated unknown slots with highly compatible known items. In this paper, we exploit conditional probability transitions in discrete state spaces and introduce a novel Non-co-progressive Markov bridge, NonMBB, a generalization of Markov bridge for constructing bundles of products. Specifically, given bundle's unique properties, NonMBB first improve Markov bridge for slot modeling through: i) categorizing item slots into high-compatibility slots and weakly-associated slots with respect to the known items, assigning them distinct timestep schedules to avoid the pitfall where high-compatibility slots must reference noisy, uninformative context at identical noise levels; ii) applying introduce schedule reweighting to constrain the maximum timestep disparity across slots, thereby preventing noise residuals caused by excessive asynchrony. Extensive experiments on Spotify and POG demonstrate that NonMBB establishes a substantially improved optimal bipartite matching over existing diffusion methods, yielding higher embedding-space alignment independently of exact-match recall.
Authors:
Yuyue Wang, Xihua Wang, Xin Cheng, Yijing Chen, Ruihua SongAbstract: Audio generation has made significant progress, yet synthesizing unified audio where speech and sounds are naturally composited remains a challenge. Current methods either rely on disjoint pipelines, which fail to capture fine-grained interactions, or require structured inputs and external text rewriting, which limits the flexibility of free-form text prompts. In this paper, we introduce a new task: Free-Form-Text-Prompt-to-Unified-Audio generation, which aims to directly synthesize unified audio containing speech, sound, and their composites from unconstrained natural language. To address this task, we propose PlanAudio, a unified, autoregressive LLM-based framework. First, it simplifies the model architecture by leveraging intrinsic LLM reasoning capability instead of traditional text encoders. Second, it introduces a semantic latent chain-of-thought mechanism, an implicit planning mechanism that bridges high-level semantic understanding and low-level acoustic synthesis. Furthermore, we create PlanAudio-Bench, a specialized benchmark for evaluating composite audio scenarios. We perform evaluations in the scenarios of speech, sound, and their composites. The results demonstrate that PlanAudio generally outperforms the existing pipeline and unified baselines, while staying competitive with models designed for a single scenario. Our analysis further reveals the superiority of semantic latent CoT over other CoT mechanisms and highlights the importance of continuous multi-scenario training curricula.
Abstract: Foundation models (FMs) have shown great promise in medical imaging, but most FMs are trained on unimodal data within isolated domains, such as brain MRI alone. Human aging and disease arise through coordinated biological processes across organs, therefore motivating multimodal FMs that learn whole-body representations. A key challenge, however, is that real-world multimodal biomedical data are often missing not at random, which can reduce power, limit generalizability, and introduce bias. We propose Pan-FM, a pan-organ foundation model pre-trained on imaging from seven organs (Brain, Heart, Adipose, Liver, Kidney, Spleen, and Pancreas) under realistic missing-organ scenarios. Pan-FM uses a unified backbone that handles organ missingness during both training and inference, and is pre-trained with masking-based self-distillation. We find that naive multimodal pre-training leads to dominant-organ shortcut learning bias, with the model over-relying on dominant organs such as adipose and heart. To address this, we introduce Saliency-Guided Masking (SGM), which uses the model attention distribution to adaptively mask dominant organs during pre-training, thus encouraging more balanced cross-organ, whole-body learning. Notably, SGM introduces negligible computational overhead and can be seamlessly integrated into existing self-supervised learning frameworks to improve multi-organ representation learning. On the UK Biobank, Pan-FM achieves stronger prediction across 13 disease categories and 14 single disease entities than single-organ and multi-organ baselines, with improved robustness under missing-organ settings. Pan-FM serves as a scalable solution to realistic modality-missingness in multimodal learning in system neuroscience and as a step toward more generalizable whole-body FMs.
PaperID: 6699, Poster
Abstract: Node-level graph anomaly detection (GAD) identifies nodes whose attributes and interactions deviate from dominant graph regularities. Existing GAD models usually encode normality and anomaly scoring indirectly through architectures, message passing, reconstruction or contrastive objectives, and tuned score families. This entangles graph trust (how strongly graph structure should define normality), graph-spectral weighting, and anomaly-score choice, yielding scores that are costly, opaque, and unstable across graph regimes. We propose EB-GAD (Empirical-Bayes GAD), a training-free framework that models normality as graph-aware generalized Ornstein--Uhlenbeck (GOU) relaxation toward a graph-filtered template. Empirical Bayes fits the graph precision and template from the residual-field likelihood; the GOU then turns scoring into closed-form transport cost from a feature-neutral node to its observed endpoint along graph-spectral relaxation. Sweeping relaxation horizon and endpoint tolerance yields a \emphtransport bank: equilibrium Mahalanobis scoring is one limit, while finite-horizon transport-energy and scale-normalized ratio scores reveal anomalies that static equilibrium scoring can mask. A label-free selector chooses the anomaly-score family from feature homophily, edge density, and feature dimension, then ranks candidates by fitted-null KS distance and rank-stability across neighboring transport configurations. On 11 benchmarks including financial fraud networks with up to 3.7M nodes, EB-GAD achieves the best or tied-best AUROC on 9 of 11 datasets; on the largest graph, eigendecomposition and scoring finish in about six minutes.
PaperID: 6700, Poster
Abstract: Memory-augmented reasoning agents typically store a single realized reasoning trajectory as the unit of experience. This abstraction is brittle under stochastic decoding: the same task can induce multiple plausible reasoning paths with different decompositions, assumptions, and failure modes. We introduce Branch-Structured Distributional Memory (BranchDIME), a framework for constructing memory from branch-structured trajectory sets rather than isolated traces. BranchDIME samples short reasoning prefixes, clusters them into prefix-induced branches, expands representative prefixes into full trajectories, and distills the resulting set into reusable memories containing positive strategies, anti-patterns, and contrastive rules. Across GPQA-DIAMOND and MATH500, BranchDIME improves over single-trajectory memory and alternative trajectory-selection baselines at comparable inference latency, yielding relative gains of 8.7% and 2.2% over single-trajectory memory, respectively. Coverage analysis shows that BranchDIME spans all discovered branches, compared with 20.0% coverage for consensus selection and 64.9% for random selection, while also achieving the highest downstream accuracy. BranchDIME further transfers across tasks and models, improving over single-trajectory memory by up to 15.4% on AIME25 and OlymMATH. These results support a shift from trajectory as experience to trajectory distribution as experience: memory should capture the local reasoning landscape induced by stochastic generation, not merely one sampled path.
PaperID: 6701, Poster
Authors: Yoel Marcu, Daphna Weinshall
Abstract: A central challenge in deep Active Learning is that optimal selection strategies depend on the labeling budget, shifting from coverage in low-budget regimes to uncertainty in high-budget regimes. Existing methods attempt to bridge this gap via heuristic switching or interpolation, but lack a principled mechanism to determine when such transitions should occur. We address this with pseudo-CDNV (pCDNV), a label-free proxy for neural collapse that tracks representation maturity. We show that pCDNV exhibits a characteristic peak that reliably signals when learned representations become suitable for uncertainty-based selection. Building on this insight, we propose Geometry-Oriented Adaptive Targeting (GOAT-AL), a unified query strategy that operates across all budget regimes. Rather than switching objectives, GOAT-AL maintains a single coverage-based objective while adapting the underlying feature space from self-supervised to task-aligned representations. Experiments on CIFAR-10, CIFAR-100, and TinyImageNet demonstrate that GOAT-AL consistently matches or outperforms state-of-the-art methods across low-, mid-, and high-budget regimes.
PaperID: 6702, Poster
Authors: Tianyi Li, Zhiqiang Shen
Abstract: Linear mode connectivity (LMC) provides a promising foundation for understanding and merging independently trained neural networks, but existing methods typically optimize the interpolation path from only one model endpoint, limiting their scalability and effectiveness for large pretrained transformers. We propose a novel and scalable framework for enabling LMC-based model merging to \em billion-parameter pretrained transformers. Our method applies properly parameterized functionality-preserving weight transformations to align functionally equivalent solutions, and introduces a dual learning procedure in which both models jointly learn their corresponding transformations toward a shared linear interpolation path. This bidirectional optimization substantially reduces interpolation barriers and enables more reliable merging across large-scale architectures. Empirically, we show that our approach achieves near-zero loss barriers on WikiText for language models with medium-sized parameters, representing, to our knowledge, the first demonstration of near-barrier-free linear connectivity at this scale. In the vision domain, ViT-L maintains above 69% ImageNet top-1 accuracy throughout the interpolation path, while modern billion-parameter LLMs exhibit only small loss barriers. These results suggest that properly resolving parameter symmetries enables large pretrained Transformers to be connected and merged through simple linear paths with substantially improved interpolation performance.
PaperID: 6703, Poster
Abstract: Modeling biomolecular conformation and dynamics is critical for elucidating biological functions, motivating generative surrogates for molecular dynamics (MD) simulations to address two complementary tasks: time-independent conformational sampling and time-dependent trajectory generation. Fundamentally, both tasks require matching the distribution of the generated collections to the MD reference, rather than reproducing individual configurations. Yet existing deep-learning approaches predominantly optimize a per-sample regression loss (score- or flow-matching), whose gradient is computed in isolation per sample and therefore carries no direct signal about the cross-sample distributional properties that define both tasks, leading to biased ensembles and long-horizon temporal drift. We address this with Collective Supervision, a training paradigm that aligns the empirical measures of the generated and reference collections via a maximum mean discrepancy on SE(3)-invariant physical observables; a single loss handles both tasks within a unified formulation. To mitigate the exposure bias inherent in autoregressive trajectory rollout, we further introduce Collective Rolling Forcing, which couples Collective Supervision with autoregressive self-rollout during training. Our framework, CoDyna, generalizes across proteins, protein--ligand and protein--protein complexes; on four all-atom MD benchmarks (ATLAS, MISATO, DynaRepo, and DynaBench), a unified model surpasses task-specialized baselines on most thermodynamic and kinetic fidelity metrics and remains structurally stable over 2k-frame rollouts.
PaperID: 6704, Poster
Abstract: Audio-visual large language models (AV-LLMs) have made strong progress in multimodal understanding, but they typically treat audio as monaural semantic content and therefore struggle to reason about where sounds originate and how they relate to the visual scene. Recent spatial audio-visual studies address parts of this problem, but often focus on specific abilities, such as spatial correspondence or direction and distance reasoning. In this paper, we propose HLR-AVSceneQA, a comprehensive benchmark for spatial audio-visual scene understanding. HLR-AVSceneQA evaluates whether models can jointly hear, localize, and reason by recognizing what is heard, grounding where it comes from in egocentric and allocentric views, and inferring how sound sources relate to visible, hidden, or nearby objects. We further introduce HLR-LLM, which extends a strong audio-visual foundation model with a binaural spatial audio branch. HLR-LLM is trained with a three-stage curriculum, applying chain-of-thought supervision in the final stage to help the model learn structured multi-step spatial reasoning. Experiments show that existing AV-LLMs remain limited in spatial grounding and relational reasoning, while HLR-LLM substantially improves spatial audio-visual scene understanding without sacrificing semantic audio-visual perception.
PaperID: 6705, Poster
Authors: Yunfei Chen, Yuchen Zhang, Hongyu Lin, Peng Liu, Zhan Yang
Abstract: The proliferation of multimodal data has underscored the imperative for efficiency in cross-modal retrieval. By mapping data into a unified Hamming space, Cross-modal hashing (CMH) significantly enhances both storage and computational efficiency, thereby establishing itself as a preferred favored paradigm for retrieval tasks. However most methods assume complete data, which rarely holds in practice. In incomplete settings, missing modalities (e.g., text or image) break semantic correspondence and make unsupervised hashing unreliable. Therefore, we propose Unsupervised Generative Completion and Graph-attention Refinement for Incomplete Cross-modal Hashing (UGGRH) method. Specifically, UGGRH leverages pretrained generative models for modality completion and employs a discriminator-guided similarity refinement module to fuse intra-modal relations with paired image-text matching signals, thereby constructing a robust cross-modal similarity matrix that yields a noise-tolerant similarity target for training. After that, it learns discriminative binary codes via hash-aware graph attention and cross-modal fusion, to refine modality embeddings and produce unified hash codes. Experiments on MIRFlickr25K, NUS-WIDE, and MS COCO datasets demonstrate that UGGRH consistently achieves superior performance under high missing rates and modality imbalance, showing significant improvements in retrieval accuracy and robustness across challenging incomplete scenarios.
Abstract: We study language generation in the limit under a global preference ordering on strings, as introduced by Kleinberg and Wei. We aim for breadth, but impose an additional requirement of timeliness: higher-ranked strings should be generated earlier. A string is then only credited if it is generated before a deadline, where its deadline is defined by a function that maps a string’s rank in the target language to the time by which it must be produced. This is in keeping with a central consideration in machine learning, where inductive bias favors "simpler" or "more plausible" outputs, all else being equal. We show that timely generation is impossible in a strong sense for eventually consistent generators—the protagonists of most prior related work. Under what is perhaps the mildest natural relaxation of consistency, a hallucination rate that vanishes over time, we show that we can circumvent our impossibility result. In particular, we can achieve optimal density with respect to any superlinear deadline function. We also show this is tight by ruling out timely generation with linear deadlines and vanishing hallucination rate.
PaperID: 6707, Poster
Authors: Madeline Kitch, Nihar Shah
Abstract: In many machine learning applications, human and LLM evaluators use assessments of relevant criteria to create an overall evaluation for an item or individual. In applications like admissions, committees assess candidates on attributes such as test scores, GPA, and research experience to evaluate their overall fit for the program. In medical care, clinicians use patient reports of symptoms to consider preliminary diagnoses and assess risks. Each case involves mapping measurable criteria to an overall evaluation—a process that reflects the evaluator's underlying preferences. We focus on the fundamental issue of learning these preferences. Many applications of this problem make specific modeling assumptions on evaluator preferences that may be substantially violated in the real world. We make the minimal assumption that the preference function is coordinate-wise non-decreasing, which is reasonable in a large number of evaluation settings. We theoretically characterize the severity of model mismatch for many common assumptions and show that it can lead to catastrophic effects for learning evaluator preferences and other important downstream tasks. We then present an algorithm for learning evaluators' preferences that is robust to model mismatch. We prove theoretically that our algorithm can learn any monotonic preference function without sacrificing performance when the linearity assumption holds. Evaluations of our algorithm with synthetic simulations and real-world data confirm its ability to learn preferences robustly and illustrate key aspects of LLM and human preferences.
Abstract: Reinforcement learning has emerged as a powerful paradigm for improving large language model (LLM) reasoning, where rollouts are sampled from the policy and reward signals computed on those rollouts are used to update the policy. However, in data-scarce scenarios, obtaining ground-truth labels to verify rollouts at scale often requires expensive human annotation or labor-intensive expert verification. For instance, evaluating mathematical proofs demands expert review, and open-ended question answering lacks definitive ground truth. When ground-truth labels are scarce, the effectiveness of reinforcement learning fine-tuning is constrained. Inspired by the success of semi-supervised learning in propagating labels from labeled to unlabeled samples, we propose MemReward, a graph-based experience memory framework that integrates reward propagation directly into online policy optimization. MemReward stores rollouts (thinking processes and final answers) from an initial LLM policy as nodes in a heterogeneous graph connected by similarity and structural edges, over which a GNN propagates rewards from labeled to unlabeled rollouts. To train such a framework, we first warm up the GNN on labeled rollouts to predict rewards via heterogeneous aggregation over query, thinking, and answer nodes. During online RL fine-tuning, unlabeled rollouts are attached to the graph by query similarity, and the GNN predicts their rewards, yielding a hybrid reward acquisition strategy that combines ground-truth and GNN-predicted rewards. Experiments on Qwen2.5-1.5B and 3B in mathematics, question answering, and code generation demonstrate that MemReward, with ground-truth rewards on only 20% of rollouts, achieves 96.6% of Oracle performance on 1.5B and 97.3% on 3B, and closely approaches Oracle on out-of-domain tasks.
PaperID: 6709, Poster
Abstract: As LLMs increasingly serve in advisory and deliberative roles, users rely on them for non-verifiable reasoning in domains lacking objective ground truths. However, traditional evaluations of LLM reasoning focus almost exclusively on fact-based domains, such as mathematics and science, leaving uncertainty over whether and to what degree models can handle ambiguous, subjective, or value-laden problems over time. To address this concern, we propose moral reasoning as a paradigmatic subdomain of non-verifiable reasoning. We define moral robustness as a model’s capacity to exhibit sound moral reasoning across time and contexts, and we introduce a scalable, adversarial, multi-turn evaluation framework to empirically measure this capability. We simulate 48,000 user-agent moral deliberations across four frontier LLMs, varying premise relevance, premise order, conversation duration, and the user's stated moral view. We find that models successfully ignore morally-irrelevant distractors, but shift their reasoning by up to 6.5%, on average, towards the user's stated preferred moral view, and varying their reasoning depending on factors such as order (altering moral judgments by order in 13-22% of the cases) and duration (altering moral judgments between single-turn and multi-turn in 10-24% of the cases). Our analysis indicates that models tailor not just their final verdicts but their underlying justifications to align with a user’s moral viewpoint — a failure mode we characterize as moral deliberative sycophancy.
Authors:
Xinglin Wang, Zishen Liu, Shaoxiong Feng, Peiwen Yuan, Yiwei Li, Jiayi Shi, Yueqi Zhang, Chuyi Tan, Ji Zhang, Boyuan Pan, Yao Hu, Prof. KanAbstract: Agentic systems increasingly solve complex user requests by executing orchestrated workflows, where subtasks are assigned to specialized models or tools and coordinated according to their dependencies. While recent work improves agent efficiency by optimizing the performance--cost--latency frontier, real deployments often impose concrete requirements: a workflow must be completed within a specified budget and before a specified deadline. This shifts the goal from average efficiency optimization to maximizing the probability that the entire workflow completes successfully under explicit budget and deadline constraints. We study constraint-driven online resource allocation for agentic workflows. Given a dependency-structured workflow and estimates of success rates and generation lengths for each subtask--model pair, the executor allocates models and parallel samples across simultaneously executable subtasks while managing the remaining budget and time. We formulate this setting as a finite-horizon stochastic online allocation problem and propose Monte Carlo Portfolio Planning (MCPP), a lightweight closed-loop planner that directly estimates constrained completion probability through simulated workflow executions and replans after observed outcomes. Experiments on CodeFlow and ProofFlow demonstrate that MCPP consistently improves constrained completion probability over strong baselines across a wide range of budget--deadline constraints.
PaperID: 6711, Poster
Authors:
Sinuo Fan, Zirui Fu, Marco Donato, Yingjie LaoAbstract: Post-training quantization is a standard tool for the efficient deployment of large language models (LLMs). In safety-aligned models, it introduces a critical and underexplored risk: preserving utility under compression does not ensure that safety alignment is retained. Small perturbations from quantization can disproportionately degrade refusal to harmful prompts, leading to increased compliance with harmful prompts or over-refusal of benign prompts, even when aggregate performance appears stable. This gap reveals a fundamental limitation of existing quantization approaches, which primarily optimize for utility while overlooking failures due to safety alignment degradation. This paper presents ACQueReLlo, a reinforcement learning framework for mixed-precision post-training quantization. It formulates block-wise bit allocation as a constrained Markov decision process and learns a quantization policy under explicit constraints on utility, harmful prompt refusal, benign prompt over-refusal, and overall bit budget. To ensure tractable training, the policy relies on proxy evaluations over sampled subsets of utility, harmful, and benign prompts. Experiments on various LLMs show that the learned policy achieves a better overall balance between compression, utility, and safety alignment than uniform quantization and other baselines.
PaperID: 6712, Poster
Abstract: How little of a pretrained large language model has to change for it to acquire new capabilities? We find that zeroing fewer than 0.05% of its weights, with no other modification, matches full-parameter fine-tuning (FPT) at scale across the three dominant LLM fine-tuning paradigms: on-policy reinforcement learning with verifiable rewards (GRPO), supervised fine-tuning (SFT), and on-policy distillation (SDFT). We call the method Bit-Mask Tuning (BMT): a learnable binary keep/zero mask over the gate projection of transformer blocks. BMT reaches FPT accuracy at our largest backbones, with the GRPO gap closing monotonically across model sizes from 0.5B to 8B. Two practical advantages follow: the trained mask is several 1000x smaller than the full-parameter checkpoint, and BMT forgets substantially less than FPT, retaining prior-task accuracy throughout training where FPT measurably drifts. To understand why masking alone suffices, we analyze the update geometry on GRPO and find that BMT's weight delta lands closer to full-parameter updates than other adapters on three measures: effective rank, spectrum drift, and principal-weight overlap. In sum, deletion alone produces a tiny yet high-performing adapter, reduced forgetting, and an update geometry that tracks FPT's more closely than any baseline we evaluate.
PaperID: 6713, Poster
Abstract: Estimating element frequencies in data streams is a fundamental problem when processing large amounts of data. Recent approaches suggest learning-augmented algorithms that can access a heavy-hitter oracle that knows the most frequent elements. In this paper, we challenge the idea of these complex oracles: First, we show that heavy-hitter predictions do not necessarily yield significant efficiency gains in frequency estimation. We prove that a classical algorithm using slightly more memory, \Theta(B \log B) instead of \Theta(B), matches the performance of learning-augmented methods with perfect predictions. Furthermore, we show that given (perfect) predictions, a trivial mechanism achieves the same performance as state-of-the-art learning-augmented algorithms. On the algorithmic side, we introduce SpaceR, a single-pass algorithm that matches the accuracy of existing two-pass prediction-based methods without requiring any predictions. SpaceR uses randomized sampling to identify and track frequent elements in one pass through the data, achieving near-perfect frequency recovery with minimal memory on both synthetic and real-world datasets.
PaperID: 6714, Poster
Abstract: Offline goal-conditioned reinforcement learning (GCRL) aims to learn a single policy that adapts its behavior to arbitrary goals from a fixed dataset, making injection of the goal information into the policy a central design question. While recent works increasingly leverage expressive conditional generative modeling architectures, which have demonstrated remarkable success in computer vision, architectural choices for goal conditioning in offline GCRL remain under-explored. We identify that goals in GCRL differ from typical conditioning signals in two ways: they are often noisy or redundant, and they are consumed within a time-sequential episode where the relevant context of a fixed goal shifts as the agent's state evolves. Motivated by these observations, we propose State-Anchored Goal Conditioning (SAGC), a modulation-based goal conditioning scheme that derives its scale and shift parameters by jointly processing state and goal representations in a learned latent space. We further introduce AdaptFlow, an end-to-end trainable offline GCRL architecture that integrates conditional flow matching with SAGC. AdaptFlow outperforms matches all baselines on 23 out of 27 OGBench tasks.
PaperID: 6715, Poster
Abstract: In recent years there has been a flurry of activity around using pretrained diffusion models as informed data priors for solving inverse problems, and more generally around steering these models towards certain reward models. Training-free methods like gradient guidance have offered simple, flexible approaches for these tasks, but when the reward is not informative enough, e.g., in inverse problems with highly compressive measurements, these techniques can veer off the data manifold, failing to produce realistic data samples. To address this challenge, we devise a simple algorithm, ReGuidance, that leverages prior methods' solutions as strong initializations and substantially enhancing their realism. Given a candidate solution x produced by a given method, we propose inverting the solution by running the unconditional probability flow ODE in reverse starting from x, and then using the resulting latent as an initialization for a deterministic steering process. Empirically, we evaluate our algorithm on difficult image restoration tasks including large box inpainting, heavily downscaled superresolution, and high noise deblurring with both linear and nonlinear blurring operations. We find that, using a wide range of baseline methods as initializations, applying our method results in much stronger samples with better realism and measurement consistency. We complement these results with rigorous proofs in commonly studied theoretical settings showing our technique boosts the reward and brings x closer to the data manifold.
PaperID: 6716, Poster
Abstract: Genomic Foundation Models (GFMs) have become a promising paradigm for decoding DNA sequences, yet their privacy risks are critically underexplored.Trained on human cohorts, GFMs can unintentionally memorize sensitive genomic signatures, that enable membership inference attacks (MIAs) and expose individuals to long-lasting privacy harms.However, existing defenses often fall short in genomics: globally applied mechanisms can substantially degrade motif-sensitive utility, and model-agnostic post-processing provides limited insight into which genomic regions drive leakage. Through a systematic privacy audit, we find that membership leakage in GFMs is highly non-uniform, concentrating on a small subset of high-risk samples and sparse token spans corresponding to biologically meaningful patterns. Motivated by this finding, we propose GenoGuard, a model- and sample-adaptive defense that shifts from blanket protection to targeted, region-aware mitigation. GenoGuard (i) localizes privacy-leaking subsequences via gradient-based token attribution and performs token-level risk reduction through selective gradient routing, (ii) applies sample-level risk reduction with risk-adaptive label smoothing guided by a reference-calibrated loss gap score, and (iii) provides region-level attribution maps for privacy auditing and biological interpretation. Experiments across representative GFMs (Mistral-DNA, Nucleotide Transformer) demonstrate that GenoGuard consistently improves privacy against multiple MIA families while preserving strong fine-tuning performance and interpretability.
Abstract: Large language models (LLMs) increasingly improve their reasoning at test time via additional computation, yet most procedures treat each problem in isolation. When problems arrive sequentially, accumulating reusable experience across them can further improve performance. Existing memory-based methods either store whole-solution templates that generalize poorly to novel problems or use heuristic step-level selection that is not optimized for final-answer correctness. Learning selection policies requires large-scale training data and fixed action spaces. We propose MILES (Modular Instruction Memory with LEarnable Selection for self-improving LLM reasoning), a framework that dynamically expands step-wise memory and applies correctness-optimized memory composition under realistic test-time constraints. MILES maintains modular memory units consisting of asymmetric pairs of sub-goal embeddings and sub-instructions, each associated with a learnable selection head. This memory structure enables a coarse-to-fine retrieval mechanism: The coarse level enables memory expansion and collects supervision for training selection heads from confident samples, while the fine stage applies learned selection heads to rerank coarse-level candidates and guide reasoning for uncertain samples. MILES consistently matches or outperforms prior methods while achieving superior accuracy–efficiency tradeoffs. Extensive experiments demonstrate its effectiveness, robustness, and transferability.
PaperID: 6718, Poster
Abstract: a prediction is made is often as important—if not more so—than achieving high accuracy. However, existing methods for sequential event prediction typically prioritize accuracy at the expense of interpretability. Neuro-symbolic approaches offer a promising direction to improve interpretability by learning symbolic rules, but extending them to next-event prediction is challenging due to the presence of temporal dependencies. We introduce RULES, a neuro-symbolic model that learns sequential rules directly from data, and combine it with a residual network to form RUNES (Rule Network for Event Sequences). RULES accurately recovers ground-truth rules from synthetic data and discovers compact, interpretable rule sets on real data. RUNES combines these rules additively with a residual network, achieving competitive accuracy while providing transparent access to active rules and how residuals adjust predictions.
PaperID: 6719, Poster
Authors:
junlong wu, Pengcheng Wei, Jia Sun, Huaiqing Wang, Weixuan Zeng, Zijun Li, Honglie Wang, Yongrui Heng, Boheng Zhang, Dewen Fan, Fan Yang, Tingting Gao, Houde Liu, Qianqian GanAbstract: Automatic e-commerce poster generation requires the joint optimization of product fidelity, visual composition, text rendering, and commercial appeal. Existing methods typically address this task through staged pipelines involving product segmentation, layout prediction, glyph rasterization, and conditional image synthesis. While effective in constrained settings, such decompositions rely on rigid preprocessing assumptions and weaken the mutual adaptation among product appearance, typography, and scene composition, especially for hand-held, worn, or context-dependent products. We reformulate this problem as \emphholistic product-aware poster editing: given a raw product image and structured metadata, the goal is to generate a complete poster directly, without segmentation masks, glyph control maps, or predefined layout boxes. To this end, we propose PosterDuet, a closed-loop framework that integrates a vision-language model (VLM), an image editing model, and reward-based optimization. The VLM first generates a holistic design prompt that specifies background style, layout arrangement, promotional copy, and typographic intent from the raw image and metadata. Conditioned on both the prompt and the original image, the image editor synthesizes the final poster in a unified editing process. To optimize both commercial effectiveness and visual quality, we introduce a mixed-reward learning framework that combines a CTR-oriented reward with a generative holistic quality reward, and use GRPO to optimize the prompt-generating VLM. We further improve textual accuracy via OCR-reward-based reinforcement tuning of the image editor, and exploit natural-language critiques from the generative reward model for iterative prompt refinement. Extensive experiments show that PosterDuet generates more coherent, text-faithful, and commercially effective posters than prior pipelined approaches.
Abstract: Text-to-image diffusion models have demonstrated remarkable capabilities in generating high-quality images, yet their tendency to reproduce undesirable concepts poses growing concerns for safe and controllable deployment, particularly regarding NSFW content, specific objects and styles. While existing concept erasure approaches primarily focus on DDPM-based diffusion models and rely on costly fine-tuning, the recent emergence of flow matching models introduces a fundamentally different generative paradigm for which prior methods are not directly applicable. In this paper, we propose , a training-free concept erasure method specifically designed for flow matching models. Our key insight is that semantic concepts are implicitly encoded in the directional structure of the velocity field governing the generative flow. Leveraging this observation, we construct a differential vector field that characterizes the directional discrepancy between a target concept and an anchor concept. During inference, DVE selectively removes concept-specific components by projecting the velocity field onto the differential direction, enabling precise concept suppression without affecting unrelated semantics. Extensive experiments on FLUX demonstrate that DVE consistently outperforms existing baselines on a wide range of concept erasure tasks, including NSFW suppression, artistic style removal, and object erasure, while preserving image quality and diversity.
Abstract: Linear attention models allow a fixed state size and a fixed amount of compute per token. However, due to their limited state size, linear attention models fall behind in long-context recall compared to softmax-attention-based transformer architectures. Increasing the state size of linear attention improves recall performance but at the cost of higher FLOPs. In this work, we introduce Sparse Delta Memory (SDM), an architecture that scales the hidden state of gated linear RNNs to orders of magnitude higher capacity using a sparse addressing scheme. SDM extends the Gated DeltaNet architecture by replacing the dense key-value outer product with sparse reads and writes to a large explicit memory. We show that, under an isoFLOP constraint and with an identical number of parameters, a higher state memory capacity significantly improves performance on in-context learning and long-context retrieval tasks. Moreover, by learning the initial state of the SDM memory and therefore using it as a parametric memory, we show that the model further improves on a wide range of common-knowledge and reasoning tasks.
Authors: Xianli Zhu, Jia Yin
Abstract: We introduce the Z-Domain Neural Operator (ZNO), a causal neural operator whose layers are stable low-rank multiple-input multiple-output (MIMO) rational filters parameterized directly in the z-plane. This operator is designed to address a critical limitation of existing operator learning methods, as most of these methods are primarily tailored for continuous-time problems, while a large class of system-identification problems is intrinsically discrete-time. The z-domain form expresses stability as a unit-disk pole constraint and makes learned discrete-time poles directly readable. The model combines low-rank channel mixing, smooth stable pole reparameterization, causal recurrence, and an optional short finite impulse response (FIR) branch in a single z-domain rational recurrent layer. Across controlled discrete system-identification experiments, ZNO's advantage is most evident when the target dynamics are stable rational systems with lightly damped poles near the unit circle. Under matched parameter budgets, ZNO is not uniformly dominant; however, with validation-selected configurations, the same architecture can achieve the lowest mean error across the controlled tasks. A five-bin difficulty sweep over near-unit-circle / long-memory dynamics further shows that ZNO has the lowest mean error across all memory regimes, from short (\approx 10 steps) to long (\approx 100-200 steps). On five public nonlinear system-identification benchmarks, ZNO is competitive with neural operator and state-space baselines, achieving the lowest mean error on benchmarks whose dynamics align with stable rational discrete-time filters, while classical or state-space baselines remain preferable on some systems. These results position ZNO as a strong model for stable rational discrete-time dynamics, especially in near-unit-circle and long-memory regimes, but not as a universal replacement for specialized system-identification methods.
PaperID: 6723, Poster
Abstract: Real-world backdoor attacks often require poisoned datasets to be stored and transmitted before they are used to compromise deep learning systems. In the era of big data, however, the inevitable use of lossy compression poses a fundamental challenge to invisible backdoor attacks. We observe that triggers embedded in RGB images can become ineffective once the images are lossily compressed into binary bitstreams, such as JPEG files, for storage and transmission. Consequently, poisoned data may lose their malicious functionality after compression, causing backdoor injection to fail. Prior compression-based attacks typically exploit compression artifacts as certain triggers to distinguish poisoned RGB samples from uncompressed benign ones, rather than addressing whether malicious information can survive a shared lossy storage-and-transmission pipeline. In this paper, we highlight the necessity of explicitly accounting for lossy compression in backdoor attacks. This requires attackers to ensure that transmitted binary bitstreams preserve malicious trigger information, such that effective triggers can be induced after decompression. Building on the region-of-interest (ROI) coding mechanism in image compression, we propose two poisoning strategies tailored to inevitable lossy compression. First, we introduce Universal Attack Reactivation, a general method that uses sample-specific ROI masks to reactivate trigger information in bitstreams for learned image compression (LIC). Second, we present Compression-Adapted Attack, a new attack strategy that employs customized ROI masks to encode trigger information into bitstreams and applies to both traditional codecs and LIC. Extensive experiments demonstrate the effectiveness of both strategies.
Abstract: Behavior cloning provides strong imitation learning guarantees when training and test environments share the same dynamics. However, in many deployment settings the test environment's transitions differ from training, and classical offline IL offers no recourse: the learner must commit to an action at every state, even when its demonstrations are uninformative and could lead to arbitrary degradation of performance. This motivates the study of _selective_ imitation, where the learner may choose to _stop_ when it cannot act reliably. We introduce a model for selective imitation under arbitrary dynamics shift: given labeled expert demonstrations from a training environment and unlabeled state trajectories from the same expert in a test environment, the learner outputs a _selective policy_ that is _complete_ (rarely stops in training) and _sound_ (incurs low regret before stopping in test). Our algorithm, \operatornameSeqRejectron, constructs a stopping rule using a small set of _validator policies_ whose size is independent of the horizon or policy class. For deterministic policies, this yields horizon-free \tildeO(\log|\Pi|/\epsilon^2) sample complexity, assuming sparse costs. For stochastic policies, we obtain analogous horizon-free guarantees using a cumulative Hellinger stopping time. We extend the framework to misspecified experts and different expert policies across train and test and obtain results that gracefully degrade with the amount of misspecification.
PaperID: 6725, Poster
Authors: Diyako Ghaderyan, Stefan Werner
Abstract: Decentralized training of deep learning models is widely used to enable data privacy and on-device learning over networks. In realistic scenarios, heterogeneity across clients' local data poses an optimization challenge and can severely degrade test accuracy, especially on sparse communication graphs. Building on recent advances in centralized primal averaging, we propose Decentralized Primal Averaging (DPA), a decentralized learning algorithm that combines primal averaging with a local quasi-global momentum estimate. DPA maintains three coupled sequences for optimization, gradient evaluation, and output averaging, while using consecutive mixed iterates to construct a network-aware momentum direction without transmitting an auxiliary tracking variable. We prove a nonconvex stationarity guarantee for DPA's averaged output model and establish a linear-speedup result up to network- and heterogeneity-dependent residual terms. We also introduce DPA-1G, a one-gossip-per-iteration variant that matches the communication budget of DSGD-style baselines. Experiments following the GUT decentralized image-classification benchmark show that DPA outperforms DSGD, gradient tracking, QG-DSGDm, momentum tracking, and QG-GUTm across datasets, architectures, topologies, and heterogeneity levels, with the largest gains under high heterogeneity. DPA-1G matches or exceeds the one-gossip baselines at the same per-iteration communication budget.
PaperID: 6726, Poster
Authors: Charles Cao
Abstract: Recent direct-sum hardness results show that computing multi-head attention by evaluating each head independently is essentially optimal under standard complexity assumptions, but leave open whether the key-value sharing in modern MQA and GQA architectures can circumvent this barrier. We prove that it cannot: even one-layer MQA with a single shared key-value head requires one full quadratic routing computation per query head in the worst case. To quantify how far practice departs from this worst case, we develop a calibration-stacked spectral framework that measures the effective number of independent routing-message directions in a layer, carefully separating signed, stochastic, and softmax-realizable compression; we also prove tight error-compounding bounds and NP-hardness of the natural route-merging compiler problem. Empirically, across fourteen open-weight GQA checkpoints from four model families (0.6B-70B parameters), eight shared routes capture over 90% of the routing-message energy in every model, and larger models use a progressively smaller fraction of their theoretical routing capacity, revealing a substantial and widening gap between worst-case hardness and the redundancy present in deployed transformers.
PaperID: 6727, Poster
Abstract: AI agents are vulnerable to prompt injection attacks, where malicious content hijacks agent behavior. Among proposed defenses, architectural isolation provides the strongest guarantees by strictly separating trusted task planning from untrusted environment observations. However, applying this design to Computer Use Agents (CUAs), which automate tasks by viewing screens and executing actions, presents a fundamental challenge. Current agents require continuous observation of UI state to determine each action, which conflicts with the isolation required for security. We resolve this tension by demonstrating that UI workflows, while dynamic, are structurally predictable. Single-shot planning, where a trusted planner emits upfront a complete branching plan covering all anticipated runtime states, provides control flow integrity guarantees against arbitrary instruction injections. We introduce NOVA (Navigating via Observation, Verification, and Action) to make this viable in the combinatorially large UI state space, where the plan can invoke a perception model to resolve runtime values such as UI coordinates. We evaluate our design on OSWorld, and retain up to 57% of the performance of frontier models while improving performance for smaller open-source models by up to 19%, demonstrating that rigorous security and utility can coexist in CUAs. Although upfront planning prevents instruction injections, we show that additional measures are needed to defend against Branch Steering attacks, where adversaries deceive the perception model into routing execution down attacker-preferred branches of the plan, such as redirecting the agent to a malicious website.
PaperID: 6728, Poster
Abstract: Reconstructing the connectome of the brain is a central goal of modern neuroscience, but the required electron microscopy (EM) imaging is extremely time-consuming. EM image super-resolution (SR) accelerates this acquisition for downstream neuron segmentation, yet existing methods do not adequately couple SR with segmentation for EM: SR lacks segmentation guidance, and segmentation cannot adapt to super-resolved features. In this work, we propose a joint SR and segmentation framework specifically designed for EM images. In the super-resolution part, we introduce segmentation as an auxiliary task via a shared encoder with task-specific decoders to provide a semantic prior, improving reconstruction fidelity. In the segmentation part, we introduce a lightweight partial diffusion model that initializes denoising from a partially noised SR embedding, producing a structural prior that closely approximates real HR features and enhances segmentation accuracy on super-resolved EM images. Extensive experiments across multiple datasets and degradation settings demonstrate the superior performance of our framework.
PaperID: 6729, Poster
Abstract: Federated Multi-view Clustering (FMVC) enables privacy-preserving cross-view fusion from distributed multi-view data, where each client holds one view of samples and participates in the server's fusion without exposing raw data. In practice, existing FMVC methods face two key challenges: (1) the collected samples often show unaligned correspondence across views, especially under distributed collection; and (2) the gap between higher data privacy and lower computation cost remains hard to reconcile. To address these issues, we propose an Anchor Graph Alignment Framework for Federated Unaligned Multi-view Clustering ( ), which employs a coarse-to-fine alignment strategy on lightweight anchor graphs to achieve misalignment-robust cross-view fusion while bridging the gap between data privacy and computation cost. Specifically, in the coarse alignment stage, we introduce an anchor-based graph generation module to generate an anchor graph on each client and upload them to the server for category-wise alignment by a designed topology-aware alignment module, which distills topology components as category-exclusive signatures to guide and achieve ideal category-wise alignment. In the next fine alignment and fusion stage, we design an attention-based alignment module to encourage each sample to pay attention to highly similar samples for sample-wise alignment, and the fine-aligned results are contrastively fused into a cross-view consistent anchor graph for clustering. Extensive experiments on multiple datasets demonstrate that Fed-AGA achieves state-of-the-art performance among related methods, while theoretical analysis offers theoretical motivation for two key components.
Abstract: Scalarization is widely used in multi-objective optimization owing to its simplicity and scalability. In many applications, the goal is to generate solutions that represent diverse user preferences, ideally with uniform coverage of the Pareto front (PF). However, uniformly sampling scalarization weights usually induces non-uniform coverage of the PF. We explain this mismatch through a geometric analysis of the scalarization path. As the scalarization weight varies, the corresponding solutions trace the PF with a generally non-uniform traversal speed. This speed induces an arc-length cumulative distribution function (CDF); inverting this map yields a principled rule for selecting weights that produce uniform PF coverage. Building on this insight, we propose ront). For structured problems, including bi-objective bandits, we derive closed-form expressions for this CDF map and the resulting PF-aware weight sampling rule. For general problems, SURF alternates between CDF reconstruction and weight sampling. Theoretically, we show that under provable conditions, SURF converges linearly to an unavoidable finite-sampling floor. Empirically, experiments on bandits and MO-Gymnasium benchmarks demonstrate that SURF efficiently achieves more uniform PF coverage than baselines.
PaperID: 6731, Poster
Authors: Jianing Zhou, Ziheng Zeng, Hongyu Gong, Suma Bhat
Abstract: Recent advancements in Automatic Speech Recognition (ASR) have significantly improved performance for high-resource languages, yet ASR systems still struggle with low-resource languages and dialects due to the scarcity of annotated training data. To address this challenge, we propose a novel framework that enhances the utilization of limited data through a two-pronged approach: (1) a novel data augmentation technique based on a multi-view pseudo-parallel strategy, and (2) a robust multi-level contrastive learning framework capable of jointly leveraging semantic and dialect-specific attributes to improve model effectiveness under noisy conditions. Our approach effectively generalizes across dialectal variations when dialect data are non-parallel, allowing acquisition of shared linguistic structures and dialectal distinctions. Furthermore, our method could be simply transferred to scenarios where dialect data are parallel. We validate our method by adapting the Whisper-large model for Swiss German, Chinese and Arabic dialect ASR tasks, demonstrating substantial performance gains. Specifically, our framework achieves up to a 48.59% reduction in WER and 23.16% reduction in CER under non-parallel conditions, outperforming state-of-the-art baselines. These results highlight the robustness and adaptability of our approach for low-resource dialectal ASR. The data and codes will be public available upon acceptance.
Abstract: Generalist neural routing solvers have shown great potential in solving diverse vehicle routing problems (VRPs) with a unified model. However, existing solvers are typically limited to symmetric settings or degrade in performance when switching to asymmetric settings due to input inconsistencies or inherent structural differences, substantially limiting their practicality in real-world scenarios that encompass both scenarios. To address this limitation, we define the spatial position of each node based on the relative distances to a specific set of pivots and further propose a Spatial Pivot-Aligned Coordinate-free Embedding (SPACE) framework that unifies node representation and solution generation across symmetric and asymmetric VRPs. Specifically, we construct a bidirectional Fréchet representation using a novel furthest pivot sampling strategy to enable invariant node representations across distinct problem settings. Furthermore, we introduce a weight-decomposed adaptive decoding mechanism that decouples geometric perception from problem representations, mitigating the overfitting of constraint decisions to a specific geometry setting. Extensive experiments on 110 VRP variants, comprising 55 symmetric problems and their asymmetric counterparts, demonstrate that SPACE achieves promising zero-shot generalization in both symmetric and asymmetric VRPs.
PaperID: 6733, Poster
Abstract: Hyperbolic representations have shown remarkable success on hierarchical and tree-like data, and Transformers have become a central architecture for representation learning. However, a principled construction of hyperbolic Transformers remains challenging because the self-attention layer decomposes into three coupled primitives, transformation, similarity, and aggregation, that should all respect the underlying geometry. Existing hyperbolic Transformer designs are typically built in the Poincar\'e or Lorentz models, where bounded domains, time-like constraints, tangent-space detours, or projection steps make it difficult to keep the full block intrinsic and scalable. To address this gap, we propose \PVFormer, to the best of our knowledge, the first Transformer architecture that operates natively in proper velocity (\PV) space. The proper velocity model provides an unconstrained representation of hyperbolic geometry, and \PVFormer exploits this structure to redesign the Transformer block in \PV space. On the transformation side, we use \PV homomorphism transformations to update query and value representations. On the similarity side, we generalize Euclidean inner-product attention with a \PV Busemann score inspired by Busemann-based hyperbolic learning and a \PV point-to-hyperplane key transformation. On the aggregation side, we replace Euclidean averaging with \PV midpoint aggregation and further derive a scalable \PV linear attention variant for large graphs. Extensive experiments support the effectiveness of \PVFormer across graph, text, and vision benchmarks. The code will be open-sourced once accepted.
Abstract: Existing attacks and defenses for diffusion-based large language models (dLLMs) target specific vulnerabilities but lack a shared framework explaining why attacks succeed. We propose one by interpreting safety alignment as shaping the denoising energy landscape: a well-aligned model routes harmful queries toward safe outputs through an energy barrier that separates the two regions. Current jailbreak attacks reduce to two strategies for circumventing this barrier: obscuring the query's safety disposition at initialisation, or intervening mid-trajectory to force the denoising path across the energy barrier. From this perspective and the result that masked diffusion models minimise kinetic energy during denoising, we derive three complementary, training-free detection signals: a step-0 ratio that reads the initial safety disposition from the logit distribution before generation begins, and two trajectory-velocity signals that track kinetic energy in complementary subspaces of the logit space. An attack must either reveal its intent at initialisation or expend kinetic energy to cross the barrier in at least one monitored subspace, so the three signals cover each other's blind spots in the energy budget by construction. Evaluation across three model families (LLaDA-8B, LLaDA-1.5, Dream-7B) confirms this complementarity. In stress tests of known attacks, every configuration that evades detection also fails to produce harmful content, suggesting that the detection and barrier-crossing thresholds are hard to separate.
Abstract: Sparse autoencoders (SAEs) operationalise the linear representation hypothesis: they reconstruct model activations as sparse linear combinations of interpretable dictionary atoms, on the implicit assumption that activation space is well approximated by a globally linear structure. Their reconstruction error varies sharply across layers in ways that existing scaling laws, fitted at single layers, do not explain. We argue that this variation is the empirical trace of a geometric mismatch: where the activation manifold is curved and its intrinsic dimension varies across layers, no sparse linear dictionary can match it uniformly, and the SAE's width-sparsity scaling becomes a layer-dependent function of manifold structure rather than a single universal law. We conduct the first cross-layer SAE scaling study, fitting and regressing on 844 residual-stream Gemma Scope SAE checkpoints across 68 layers of Gemma 2 2B and 9B. Stage 1 fits a per-layer scaling-law surface; Stage 2 regresses the fitted parameters and the derived per-layer width exponents on four layerwise geometric summaries. We find that manifold geometry predicts the per-layer width exponent in both models, and that the same regression coefficients learnt on one model predict the other model's per-layer exponents under cross-model transfer, indicating a transferable geometric law. At the showcase layers where richer width grids permit identification of the asymptotic floor, we find that the fitted floor tracks the layerwise geometric ordering: higher curvature and intrinsic dimension correspond to higher floor, consistent with the irreducible second-order residual that any sparse linear approximation of a curved manifold must leave behind. SAEs thus encounter not a finite-resource ceiling but a geometry-dependent wall, set by the manifold they are trying to reconstruct.
Abstract: In recent years, Rectified flow (RF) has gained considerable popularity largely due to its generation efficiency and state-of-the-art performance. In this paper, we investigate how well RF automatically adapts to the intrinsic low dimensionality of the support of the target distribution to accelerate sampling. We show that, using a carefully designed choice of the time-discretization scheme and with sufficiently accurate drift estimates, the RF sampler enjoys an iteration complexity of order O(k/\varepsilon) (up to log factors), where \varepsilon is the precision in total variation distance and k is the intrinsic dimension of the target distribution. In addition, we show that the denoising diffusion probabilistic model (DDPM) procedure is equivalent to a stochastic version of RF by establishing a novel connection between these processes and stochastic localization. Building on this connection, we further design a stochastic RF sampler that also adapts to the low-dimensionality of the target distribution under mild requirements on the accuracy of the drift estimates, and also with a specific time schedule. We illustrate the efficacy of newly designed time-discretization schedules with simulations on the synthetic data and text-to-image (T2I) data experiments.
PaperID: 6737, Poster
Authors: Han Liu
Abstract: Verifiable tool-interactive agents produce rich trajectories that contain far more information than final success or failure. However, standard RL post-training with verifiable rewards often collapses these trajectories into scalar outcomes and trains over a largely fixed data stream, leaving open the question of which examples are most useful for improving the current policy. We introduce TReDS, a trajectory-grounded framework for capability-conditioned training distribution shaping. TReDS estimates multidimensional task requirements from offline probing trajectories and controlled tool-view evidence, maintains the policy capability state in the same semantic space, and uses reliability-gated online correction to form a policy-conditioned requirement representation. A scheduler then converts requirement--capability mismatch, recent utility, and stabilizing pressure into a time-varying training distribution for GRPO, without modifying the GRPO objective. We evaluate TReDS with a progressive stress-test suite spanning \tau^2-Bench official evaluation, \tau-Bench native-bridge evaluation, BFCL V3, and API Bank, using only the \tau^2-Bench training split for policy optimization. TReDS improves over the GRPO baseline under the same training data and backend, and shows stronger behavior retention under same-family and cross-protocol evaluations. These results suggest that RL post-training for tool-interactive agents should optimize not only how trajectories are rewarded, but also where trajectory-revealed requirements allocate training probability mass.
PaperID: 6738, Poster
Authors: Rui Li, Shuang Cao
Abstract: Benchmark recovery is often treated as a deployment certificate for low-bit LLMs, but thresholded controller stacks ship a different object: a frozen score-to-action contract. This paper asks a narrower acceptance question: after quantization, does that interface survive independent threshold retuning against frozen FP16 labels, the strongest cheap repair an operator would plausibly try first? We study one fixed four-action controller, one frozen train/cal/report protocol, and four instruction-tuned checkpoints: Llama-3.1-8B/70B-Instruct and Qwen-2.5-7B/72B-Instruct. In the only score-carrying regime, W4A16, GPTQ + -search stays near FP16 on MMLU yet still leaves 11.9/10.9 percentage-point false tool-trigger rates and 16.1/15.4 aggregate FCR on the two large models. Because BFCL directly labels the tool frontier, that failure already establishes the paper's main acceptance claim even if FAQ and the preservation-only retrieve/stop frontiers are ignored. Aggregate FCR is reported only as interface-wide preservation context because retrieve/stop remain preservation-only. FAQ is included only as a bounded reducibility probe under the same frozen audit target, not as a generic new quantizer or the central novelty claim. FAQ lowers tool FTR and aggregate FCR without obvious MMLU or throughput collapse, but that is secondary non-collapse evidence. The contribution is therefore a single-controller deployment-acceptance counterexample under a frozen permission set: on the retained large-model W4A16 rows, benchmark-close MMLU recovery does not certify preserved tool-trigger behavior. The accompanying contribution is the freeze-and- retune audit protocol that makes that failure measurable.
PaperID: 6739, Poster
Authors: Ziyao Yi, Luca Savant Aira, Diego Valsesia, Tiziano Bianchi, Enrico Magli
Abstract: Burst image restoration aims to reconstruct a high-quality image from a sequence of low-quality frames. While state-of-the-art supervised methods achieve excellent performance, they rely on extensive paired datasets that may be difficult to acquire in practice and often fail to generalize across different camera sensors or real-world noise levels. Zero-shot diffusion-based methods offer a compelling, training-free alternative. However, they typically struggle on burst imaging due to their assumption of known forward operators. In this paper, we introduce ZSBR: Zero-Shot Burst Restoration, a zero-shot diffusion framework for burst restoration that performs MAP-style inference by combining a pretrained diffusion prior with a physically grounded burst formation model. ZSBR accounts for signal-dependent RAW noise by leveraging a continuous Poisson-Gaussian model. Moreover, unlike common diffusion-based methods that assume a fixed operator, ZSBR refines image alignment parameters online by optimizing per-frame affine matrices within the forward model during sampling. To stabilize the resulting optimization, we integrate an adaptive descent scheme into the sampling loop. Moreover, the method is flexible and allows to achieve a desired distortion-perception tradeoff. Experimental results show that our method provides a flexible and robust solution to burst restoration without requiring any paired training example.
PaperID: 6740, Poster
Authors: Maheed H Ahmed, Mahsa Ghasemi
Abstract: The increasing prevalence of preference-based learning, particularly for machine learning models interacting with human users, gives rise to scenarios where the model's decision should cater to the preferences of a number of users. Due to the inherent misalignments in such preferences, it is natural to seek a fair decision-making paradigm. Motivated by that, we pose the problem of fair multi-user dueling bandit, where each user's preferences over pairs of actions are encoded by a preference matrix unknown to the agent. We design a Nash social welfare objective based on the Borda scores of the individual user preferences. The notion of the Borda score measures the average likelihood of an action being preferred over the other actions, and importantly, it does not require the existence of a completely dominant action. Considering online learning in this general setting, we construct hard instances and establish a minimax lower bound on the achievable regret. We also design an explore-then-commit algorithm and derive an upper bound on its worst-case regret. Furthermore, we formulate a fair multi-user generalized linear dueling bandit to enable modeling large action spaces, which typically necessitate a structured representation. In this setting, too, we establish a lower bound on regret and an upper bound on regret for a proposed explore-then-commit algorithm.
PaperID: 6741, Poster
Authors: Melody Li, Taylor Webb
Abstract: Recent work has identified a set of emergent symbolic mechanisms that support abstract reasoning in large language models, but it remains unclear what factors drive the emergence of these mechanisms. To address this question, we trained neural networks from scratch on abstract sequence tasks, varying both architectural and data distributional factors, and investigated the effects of these factors on the emergence of symbolic mechanisms. Using a combination of representational, attentional, and casual mediation analyses, we first confirmed that transformer language models trained from scratch on our task developed symbolic mechanisms, partially capturing the set of mechanisms learned by large-scale pretrained models. We then investigated the effect of data diversity, operationalized as vocabulary size, on the emergence of these mechanisms, finding that mechanistic and behavioral signatures of emergent symbol processing scaled with data diversity, including downstream measures of systematic (i.e., out-of-distribution) generalization. Surprisingly, we found that the emergence of symbolic mechanisms was not driven by architectural inductive biases, as the same mechanistic and behavioral signatures emerged in modified transformer architectures and even multilayer perceptrons. These results suggest that data diversity, rather than architectural inductive biases, is the primary driver of emergent symbolic computation in neural networks.
PaperID: 6742, Poster
Abstract: High-contrast multiscale PDE problems are common in real-world applications, yet current neural PDE solvers struggle to achieve sufficient accuracy on such tasks. The Generalized Multiscale Finite Element Method (GMsFEM) addresses this by compressing fine-scale heterogeneity into localized basis functions which are used to obtain more accurate solutions than neural PDE solvers. However, constructing these basis functions requires solving many local eigenvalue problems—the major computational bottleneck. We address this issue by proposing Transolver-GMsFEM, a new hybrid framework that utilizes the efficiency of neural PDE solver for predicting multiscale basis functions while preserving the solution quality of GMsFEM. Experiments on 2D/3D steady-state and time-dependent high-contrast multiscale PDEs on irregular grids show proposed method achieves over 100× speedup for basis construction. Crucially, our experiments demonstrate that Transolver-GMsFEM significantly outperforms state-of-the-art neural PDE solvers in both accuracy and robustness, especially in out-of-distribution tasks where pure neural PDE solvers dramatically fail.
PaperID: 6743, Poster
Abstract: Video Reasoning Segmentation (VRS) aims to infer and segment target objects specified by reasoning queries in videos. Existing methods typically first infer the target and generate a few segmentation tokens for it, which are then fed into a mask decoder for mask prediction. However, such segmentation tokens tends to be both semantically monotonous and spatially ambiguous, making them largely insufficient to represent the target to guide mask prediction. In this paper, we propose an explicit visual object grounding framework for VRS. Specifically, we first propose a spatial-aware prompt generation and refinement scheme, in which we query a multimodal large language model to generate explicit spatial prompts (\ie, bounding box and point set) for the target and then refine them through multi-step dialogue. Furthermore, we introduce a query diversification and alignment module to generate auxiliary queries that describe the same target from different perspectives. We then enforce cross-query consistency between their predicted spatial prompts to alleviate the training bias caused by query limitation. Finally, we selects several reliable frames using a prompt-based keyframe selection strategy to support mask decoding and propagation. Extensive experiments on eight datasets demonstrate that our method significantly outperforms state of the art methods.
Abstract: Recent frontier large language models predominantly rely on Mixture-of-Experts (MoE) architectures. Despite empirical progress, there is still no principled understanding of how hyperparameters should scale with network width N, expert width N_e, number of experts M, sparsity K, and depth L to ensure both stability and optimal performance at scale. We take a principled step toward resolving this gap by analyzing three different scaling regimes: (I) co-scaling N\asymp N_e, (II) co-scaling N\asymp M\asymp K, and (III) full proportional scaling of N, N_e, M, and K. For each regime, we develop a novel Dynamical Mean Field Theory (DMFT) description of the limiting training dynamics of MoEs that provides a formal foundation for our analysis. Within this framework, we derive the unique parameterization for SGD and Adam satisfying all maximal-update (\mu) desiderata. We then show that the resulting \muP prescription induces neither monotonic improvement with scale nor robust learning-rate transfer. We trace these pathologies to scale-dependent observables in the aggregation dynamics, which motivates a refined set of desiderata that we term maximal scale stability. Guided by this principle, we derive a Maximally Scale-Stable Parameterization (MSSP) for both SGD and Adam in all three scaling regimes, and characterize the corresponding limiting dynamics - qualitatively distinct from the \muP limit - through a separate DMFT analysis. Experiments verify that MSSP robustly recovers learning rate transfer and monotonic improvement with scale across regimes. Combined with existing depth-scaling theory, these results provide a complete scaling prescription for MoE architectures as a function of width, depth, expert width, and number of experts.
PaperID: 6745, Poster
Abstract: Latent Diffusion Models (LDMs) typically adopt a two-stage pipeline: an auto-encoder (AE) is first pre-trained to define a latent space, then a diffusion model is trained to perform denoising within it. Such a two-stage design introduces a representation mismatch, as the latent space is optimized for reconstruction rather than adapting the denoising dynamics. We reveal that the , an end-to-end one-stage LDM training framework that eliminates the need for a separately trained tokenizer. Our key observation is that the LDM backbone actually performs a transformation at each denoising step, which can be interpreted as an internal decoding--encoding process. Leveraging this structure, we split the DiT backbone into two reciprocal components, (i.e., DiT Decoding), and impose image-space supervision on the intermediate features across all timesteps. Our model encourages the internal representation to align with the image domain throughout denoising, thereby establishing an explicit mapping, corresponding to an auto-encoding process. As a result, LDM-is-AE jointly learns latent representations and denoising dynamics in an end-to-end manner, yielding a diffusion-native latent space tailored to the generation process. Experiments demonstrate that LDM-is-AE exhibits highly competitive generation performance, achieving an FID of 1.84 on class-conditional image generation. Code and models will be released.
PaperID: 6746, Poster
Authors: Zimei Chen, Yi Ren, Wen-Hao Zhang
Abstract: Recurrent cortical circuits are hypothesized to perform Bayesian posterior sampling through their stochastic dynamics. Yet identifying which sampling algorithm a circuit intrinsically implements, as opposed to mapping manually designing algorithms into circuits, requires rigorous theoretical analysis of the biological circuit dynamics, and remains a fundamental open challenge because of the analytical intractability of nonlinear recurrent dynamics. Here, we tackle this challenge through comprehensive theoretical analysis of an analytically tractable nonlinear recurrent circuit dynamics used in neuroscience research--- the continuous attractor network (CAN). Surprisingly, we find that the circuits, without any non-Gaussian components in its connectivity or internal variability, intrinsically emerges heavy-tailed subspaces dynamics described as scale-mixture Langevin posterior sampling. The mechanism emerges from the circuit's own nonlinear dynamics: divisive normalization renders the overall population activation Gamma-distributed which multiplicatively modulates the time constant of the Langevin sampling in the stimulus feature subspace. Eventually, this scale-mixture structure produces a heavy-tailed (Student-t) sampling step size distribution, significantly accelerating the sampling speed. Moreover, the scale-mixture Langevin sampling approaches L\'evy-like regime in the weak input limit. Our results analytically identify an intrinsic circuit algorithm embedded in subspaces of nonlinear recurrent dynamics, and establish a minimal circuit mechanism for endogenous heavy-tail statistics without any explicit non-Gaussian components.
PaperID: 6747, Poster
Abstract: Reinforcement learning has become a central tool for large language model (LLM) post-training, where policy gradient methods are routinely deployed off-policy, even though vanilla policy gradient assumes on-policy sampling. We study when such off-policy deployment is justified, and what roles its core ingredients, including importance sampling, KL regularization, and baselines, play in correcting the resulting distribution shift. We show that for a general class of off-policy policy gradient objectives, only two corrections preserve the optimal policy as a stationary point: trajectory-level importance weighting, which suffers from the well-known curse of horizon, and KL regularization paired with a grouped-mean baseline. Focusing on the KL route, we identify it with the group-relative policy gradient (GRPG) update, establish an equivalence to a trajectory-level Bellman residual minimization objective, and obtain a finite-sample regret bound under all-policy coverage. This also gives a theoretical account of the empirically successful GRPO objective. We further uncover a hidden role of baselines in the off-policy setting: a constant shift in the baseline implements pessimism, the standard remedy in offline RL for partial data coverage. We instantiate this via the asymmetric REINFORCE objective, and connect it to the game-theoretic offline RL framework, showing a regret bound with single-policy coverage. Together, these results offer a unified theoretical view of off-policy policy gradient in LLM post-training and connect it to classical ideas in offline RL.
PaperID: 6748, Poster
Abstract: Generating precise complex atomistic architectures from natural language specifications is a frontier in AI-driven materials discovery, critical for realizing functional systems like catalytic active sites and hetero-interfaces. While Large Language Models (LLMs) offer a powerful interface for such tasks, mapping abstract semantics to rigorous structural constraints remains unsolved as LLMs inherently face two deficits: first, the lack of quantitative physicochemical priors leads to topologically imprecise outputs; and second, open-loop generation without physical verification fails to resolve intermediate spatial conflicts, causing error propagation to invalidate coupled multi-step assemblies. To bridge this gap, we introduce MaterialsPilot, which formulates generation as iterative code optimization. It incorporates two key mechanisms: (1) Hierarchical Retrieval to bridge domain knowledge gaps by providing precise domain schemas; and (2) a Physics-Aware Feedback Loop to master complex logic by autonomously debugging execution and physical violations. Across four complex material modeling tasks and seven diverse LLM backbones, MaterialsPilot increases the rate of fully successful generations from 22.0% to 66.4%, while also producing structures that better satisfy user-specified constraints and physical plausibility checks. These results establish a model-agnostic framework for physically reliable language-driven atomistic design.
PaperID: 6749, Poster
Abstract: Finetuning Language Models often requires enforcing constraints on individual inputs without compromising performance. However, current alignment methods typically impose constraints only on average, which can induce undesirable disparities across inputs or users. We propose a finetuning framework that addresses this limitation by enforcing alignment requirements as per-sample constraints. To handle the optimization challenges inherent in this approach, we utilize an augmented Lagrangian formulation in the dual domain. Since pointwise constraints can be overly restrictive in low-probability regions or in the presence of outliers, we introduce a learned, sample-dependent relaxation that minimally relaxes constraints to optimize objective performance. We demonstrate the versatility of our framework across three small language model tasks: safety in instruction following, preference satisfaction in function calling, and length-aware re-ranking. Across these settings, our approach reduces tail constraint violations while largely preserving or improving downstream performance
PaperID: 6750, Poster
Abstract: Accurately identifying metabolites (small molecules) from mass spectrometry data remains a core challenge in metabolomics, with broad applications in drug discovery, environmental analysis, and clinical research. We address the Molecule Retrieval task, which consists in recovering the chemical structure of a metabolite from its MS/MS spectrum given a set of candidate molecules. While the recent release of benchmark datasets such as MassSpecGym and Spectraverse has considerably accelerated the development of novel machine learning approaches, the complexity of data preprocessing pipelines and the lack of unified implementations make methods and results difficult to reproduce and compare. We make three contributions. First, we propose a unified framework encompassing recent approaches based on representation alignment and contrastive learning. Second, we introduce MSAlign, inspired by multimodal alignment in vision-language models, which learns a shared representation space by aligning two frozen foundation models (DreaMS for mass spectra and ChemBERTa for molecules) through lightweight MLP projections trained with a candidate-based contrastive objective. MSAlign is simple to implement, fast to train and consistently outperforms existing approaches across all benchmarks. Third, we investigate a long-standing evaluation problem: data splitting strategies in molecule retrieval implicitly trade off data leakage against domain shift. We formalize this tension by introducing a quantitative measure of distribution shift, and use it to evaluate splitting strategies in existing benchmarks. All datasets, splits, candidate sets, and a unified implementation of MSAlign and baselines are publicly released to support reproducible research.
Authors: Thomas S. Robinson, Ranjit Lall
Abstract: The standard constraint-based paradigm for causal discovery with incomplete data---impute first, test second---is frequently miscalibrated: any consistent conditional independence (CI) test rejects a true null with probability approaching 1 when imputation error induces spurious conditional dependence. We introduce , a nonparametric CI test that restores calibration by integrating multiple imputation directly into the inferential procedure via a paired permutation design. PAIR-CI compares cross-validated models that include and exclude the candidate variable while receiving the same imputed conditioning set, forcing imputation error to cancel in their loss difference rather than contaminate the test statistic. A provably consistent variance estimator jointly accounts for uncertainty arising from cross-validation and multiple imputation---to our knowledge, the first formal unification of these two inferential frameworks. In simulations, existing imputation-based CI tests exhibit false positive rates of 28--45% when data are missing not at random (MNAR), whereas PAIR-CI averages below the nominal 5% level across data-generating processes and missingness mechanisms. These gains are largest in nonlinear settings and grow with causal graph size: when integrated into the PC algorithm, PAIR-CI reduces structural Hamming distance by 8% on 10-variable nonlinear graphs, 15% on 30-variable equivalents, and up to 44% on the 56-variable HAILFINDER network, with stable performance in all settings.
PaperID: 6752, Poster
Abstract: Matrix multiplications dominate the inference cost of modern transformer-based vision models, yet existing efficiency techniques such as post-training quantization and mixed-precision inference are largely limited to the small set of fixed-width formats—INT4, INT8, BF16, and FP16—supported by conventional accelerators. We revisit stochastic computing (SC) as a way to lift this constraint: viewed as a dense adaptive quantizer, SC controls precision by bit-stream length L rather than a fixed datapath, while each multiplication reduces to a single AND/XNOR gate. We build a GPU library that emulates SC matrix multiplication at scale, exposes stream lengths as first-class kernel arguments, and evaluates SC end-to-end on image classification, object detection and instance segmentation, class-conditional image generation, and visual world-model planning. On top of this substrate, we develop a dynamic per-row mixed-precision policy that assigns stream length per token or group at matched average budget, requires no retraining, and uses the same SC hardware across schedules. Across tasks, SC remains competitive with fixed-format INT quantization at matched bit budgets, while per-row mixed precision helps maintain accuracy at lower average stream lengths. These results provide software-level feasibility evidence that SC can serve as a dense-precision substrate for fine-grained mixed-precision inference on modern vision transformers.
PaperID: 6753, Poster
Authors: Zhuorui Zhang, Jianqi Yan, Shanshan Feng, Fan LI
Abstract: Inference-time reasoning improves when models spend more computation through repeated sampling, search, or verification. Existing scaling strategies usually allocate this computation across prompts, candidate answers, or whole trajectories. We study a different structure: how the repair value of additional inference is distributed within a single generated reasoning trace. Across mathematical, scientific, coding, and knowledge-intensive reasoning tasks, recoverable utility is sharply concentrated: in oracle replay analysis over coarse trace windows, the top-three windows capture 83--90% of replay-recoverable utility. This concentration is not fully explained by position, length, or local uncertainty alone. Building on this finding, we introduce LEVER, a sparse local-intervention method for test-time reasoning that separates counterfactual discovery from deployable compute allocation. It uses local replay to estimate window-level recoverable utility, then distills this offline signal into a lightweight selector. At test time, the selector uses only online-safe trace features to decide which traces to activate and which windows receive additional rollouts. At a average 1.31× realized completion-token cost, LEVER improves average performance by about 4.2 points over base decoding. Across three base models and four task families, LEVER obtains the best result on 11 of 12 model--task pairs at the primary operating points, while using less compute than full-trajectory sampling baselines. These results suggest that effective test-time reasoning depends not only on how much computation is spent, but on whether it reaches the local commitments where repair value is highest.
PaperID: 6754, Poster
Abstract: Personalizing large language models is essential for user-centric applications, yet remains challenging under sparse user histories, privacy constraints, and limited on-device resources. Existing approaches either rely on prompting or retrieval without adapting model parameters, train and store separate adapters for each user, or use coarse group-level adapters that cannot fully capture individual variation. We propose PriSM (Prior-guided Shared-basis Mixture Personalization), a lightweight framework that models personalization as posterior-inspired update over mixtures of shared LoRA bases. PriSM first learns reusable basis adapters that capture group-level adaptation patterns, then uses a cluster-conditional Dirichlet prior and a user-conditioned hypernetwork to infer posterior mixture weights from sparse user profiles. The resulting posterior mean synthesizes a personalized LoRA update in a single forward pass, allowing the model to rely on group-level priors when user evidence is limited and to adapt toward user-specific behavior when sufficient evidence is available. This design enables scalable, privacy-preserving, and on-the-fly personalization without per-user training or per-user adapter storage. Code is available at \urlhttps://anonymous.4open.science/r/PriSM-D705.
PaperID: 6755, Poster
Abstract: Long-video question answering benchmarks like Video-MME-v2 typically evaluate systems under a uniform frame budget for every question, despite stark variation in the evidence each query actually demands. We analyze this mismatch on Video-MME-v2 and find that different question subsets prefer different frozen frame-budget policies, while uniformly increasing the frame count does not reliably improve all metrics. These findings motivate fixed-mean frame-budget routing, a controlled inference setting in which each question may receive a different frame budget while the average budget matches a uniform reference. We introduce FrameRouter, a training-free test-time framework that combines an evidence-demand router, a budget-constrained allocation rule, and a frozen bank of frame-sampling policies. By controlling how much visual evidence each question receives while leaving the underlying samplers unchanged, FrameRouter is orthogonal to query-aware frame selection. Across Video-MME-v2 and several long-video understanding benchmarks, FrameRouter improves long-video question answering by using the same average visual budget more selectively across questions.
Abstract: We consider the parameter estimation problem in logistic regression with Gaussian design: the estimation of a fixed unknown parameter \theta^ \in \mathbbR^d (\lVert\theta^\rVert_2\geq 1) from n i.i.d. samples \(x_i,y_i)\_i=1^n, where x_i~ N(0,I_d) and y_i\mid x_i ~ \mathrmBernoulli (1/ (1+exp(-x_i^T \theta^~ ))). Our main aim is to characterize the finite-sample estimation performance and convergence behavior of gradient descent (GD) on the maximum likelihood objective (i.e., the logistic loss). Under small O(1) stepsize and 0 initialization, it is shown that GD linearly converges to a small neighborhood of the true parameter and achieves \ell_2 error of order O(\sqrt\lVert \theta \rVert_2^5 d/n). This substantially goes beyond existing theoretical results that lack non-asymptotic estimation error rate and exhibit much slower parameter convergence. We also establish a faster local linear convergence to the same statistical error under a large stepsize. The main technical component is to show that the gradient of the logistic loss satisfies a certain approximate invertibility condition (AIC). To that end, we bound the concentration term via covering and peeling arguments, and show that the bias term is a contraction by a delicate eigenvalue analysis of the population Hessian matrices. Finally, we build upon the recent work Matsumoto and Mazumdar (2025) and devise a novel efficient estimator that attains a sharper rate in high dimensions. This indicates that the existing non-asymptotic guarantees exhibit sub-optimal dependence on \lVert\theta\rVert_2, and that in many regimes \Theta(\sqrt\lVert\theta\rVert_2d/n) is the tight estimation error rate. Numerical examples are provided to corroborate our theoretical results for GD.
PaperID: 6757, Poster
Abstract: Running dense, open-weight large language models on commercially accessible workstations or single-GPU cloud setups is increasingly desirable due to cost and privacy constraints, particularly on tasks like coding and reasoning. Starting at around 24B parameters, however, they exceed the available GPU VRAM on these setups, forcing reliance on offloading methods that limit throughput, while leaving substantial CPU compute underutilized. Heterogeneous CPU–GPU speculative decoding is the dominant lossless-acceleration framework in this regime, but existing methods are limited by what we refer to as the heterogeneity gap: their CPU and GPU do not co-execute, their draft is a separate set of weights rather than a subnetwork of the verifier, and their VRAM footprint cannot scale to fill the available budget. We observe that channel-saliency methods induce an ordering on the transformer’s FFN channels so that the top-k prefix closely approximates the full output, with graceful degradation as k decreases. Building on this observation, we introduce Speculative Tensor Parallelism (STeP), a training-free, self-speculative method that extracts a subnetwork of the verifier as a GPU-resident draft, places the remainder on the CPU, and verifies via concurrent CPU and GPU computation, provably preserving the verifier’s sampling distribution. Across dense models from 24B to 123B, GPU memory budgets from 32 to 192 GB, and five benchmarks, STeP outperforms SpecExec and SubSpec by up to 1.4 × and 1.8 × respectively, surpasses either tensor parallelism or speculative decoding alone, maximally utilizes the available VRAM, and closes the heterogeneity gap.
Abstract: Among parallel decoding paradigms, diffusion large language models (dLLMs) have emerged as a promising candidate that balances generation quality and throughput. However, their integration with Mixture-of-Experts (MoE) architectures is constrained by an expert explosion: as the number of tokens generated in parallel increases, the number of distinct experts activated grows nearly linearly. This results in substantial memory traffic that pushes inference into a memory-bound regime, negating the efficiency gains of both MoE and parallel decoding. To address this challenge, we propose Dynamic Expert Sharing (DES), a novel technique that shifts MoE optimization from token-centric pruning and conventional expert skipping methods to sequence-level coreset selection. To maximize expert reuse, DES identifies a compact, high-utility set of experts to satisfy the requirements of an entire parallel decoding block. Extensive experiments on MoE dLLMs demonstrate that DES reduces unique expert activations by over 55% and latency by up to 38%, while retaining 99% of vanilla accuracy, effectively decoupling memory overhead from the degree of parallelism.
Abstract: Large reasoning models (LRMs) improve language model capabilities by generating explicit thinking traces before final answers. In factuality-oriented question answering (QA), such thinking often improves overall performance by helping the model recover relevant knowledge and refine its answers. However, we find that this benefit is not uniform at the instance level: explicit thinking can also overturn correct non-thinking answers and lead to factual drift. We refer to this failure mode as \emphthinking-induced hallucination. To explain this phenomenon, we formulate explicit thinking in factuality QA as a thinking residual over the model's direct-answer tendency, which can either recover missing knowledge or introduce unsupported associations. Based on this formulation, we propose MARGO, \underlineMixed-Mode \underlineAdvantage \underlineRegularization for \underlineGrounded \underlineOptimization, a reinforcement learning framework that uses non-thinking rollouts as same-model references in advantage estimation. By constructing mixed-mode rollout groups with both thinking and non-thinking trajectories, MARGO evaluates whether explicit thinking adds factual value beyond direct answering, thereby suppressing hallucination-prone thinking while preserving beneficial thinking behaviors. Experiments across multiple factuality-oriented QA benchmarks demonstrate that MARGO improves factual reliability over strong baselines, while evaluations on mathematical benchmarks show that it preserves general reasoning ability.
PaperID: 6760, Poster
Abstract: Algorithmic ranking systems increasingly dictate high-stakes outcomes, yet most fairness interventions implicitly assume that protected attributes are categorical. When sensitive attributes are inherently continuous (e.g., age, income, or health scores), standard practices rely on arbitrary discretization, which discards crucial fine-grained information and destroys the natural ordinality of the data. In this paper, we study the problem of fair ranking with continuous protected attributes without relying on thresholding. We formalize fairness via the Kendall correlation between the ranking and the continuous attribute, and measure utility loss via the Kemeny distance from an initial score-based ranking. To solve this, we propose Fair Bubble Sort (FBS), a highly efficient adjacent-swap algorithm that minimizes utility loss subject to a strict continuous fairness constraint. We provide strong theoretical guarantees, proving that FBS is exactly optimal for the unweighted Kemeny distance. Extensive experiments on synthetic and real-world datasets show that FBS is scalable and consistently outperforms discretization-based baselines, achieving a superior fairness-accuracy Pareto front.
Authors:
Mengkang Hu, Bowei Xia, Yuran Wu, Ailing Yu, Yude Zou, Qiguang Chen, Shijian Wang, Jiarui Jin, Kexin Li, Wenxiang Jiao, Yuan Lu, Ping LuoAbstract: Symbolic world models (e.g., PDDL domains or executable simulators) are central to model-based planning, but training LLMs to generate such world models is limited by the lack of large-scale verifiable supervision. Current approaches rely primarily on static validation methods that fail to catch behavior-level errors arising from interactive execution. In this paper, we propose Agent2World, a tool-augmented multi-agent framework that achieves strong inference-time world-model generation and also serves as a data engine for supervised fine-tuning, by grounding generation in multi-agent feedback. Agent2World follows a three-stage pipeline: (i) A Deep Researcher agent performs knowledge synthesis by web searching to address specification gaps; (ii) A Model Developer agent implements executable world models; And (iii) a specialized Testing Team conducts adaptive unit testing and simulation-based validation. Agent2World demonstrates superior inference-time performance across three benchmarks spanning both Planning Domain Definition Language (PDDL) and executable code representations, achieving consistent state-of-the-art results. Beyond inference, Testing Team serves as an interactive environment for the Model Developer, providing behavior-aware adaptive feedback that yields multi-turn training trajectories. The model fine-tuned on these trajectories substantially improves world-model generation, yielding an average relative gain of 30.95% over the same model before training.
PaperID: 6762, Poster
Abstract: Masked diffusion language models generate text by iteratively unmasking positions, with practical samplers often selecting reveal positions from the model's own logits. The reveal order is therefore an endogenous latent variable of the rollout process with an enormous discrete latent space. Policy optimization objectives relying on pre-defined schedules or sampled trajectories introduce either bias or high variance. In this paper, we propose \textttFEROM, \emphFrontier Endogenous Reveal-Order Marginal Policy Optimization, which targets the rollout-induced marginal response policy of masked diffusion LMs. \textttFEROM derives a Rao--Blackwellized policy-gradient identity over latent reveal orders and expresses the resulting estimator as posterior edge occupancy on a reveal-state DAG. To make marginalization practical, we introduce Frontier Reveal Marginalization, a budgeted estimator that combines scorelaw local reveal estimation with frontier expansion of high-mass partial states. Integrated into a GRPO-style objective, \textttFEROM replaces single-path log-scores with a locally marginalized edge-based surrogate. Experiments on math and coding tasks show comparable or improved results over existing methods under matched compute budgets. Offline proxy study further shows potential in gains with increased budgets.
Abstract: The prevalence of missing values in data science poses a substantial risk to any further analyses. Despite a wealth of research, principled nonparametric methods to deal with general non-monotone missingness are still scarce. Instead, ad-hoc imputation methods are often used, for which it remains unclear whether the correct distribution can be recovered. In this paper, we propose FLOWGEM, a principled iterative method for generating a complete dataset from a dataset with values Missing at Random (MAR). Motivated by convergence results of the ignoring maximum likelihood estimator, our approach minimizes the expected Kullback-Leibler (KL) divergence between the observed data distribution and the distribution of the generated sample over different missingness patterns. To minimize the KL divergence, we employ a discretized particle evolution of the corresponding Wasserstein Gradient Flow, where the velocity field is approximated using a local linear estimator of the density ratio. This construction yields a data generation scheme that iteratively transports an initial particle ensemble toward the target distribution. Simulation studies and real-data benchmarks demonstrate that FLOWGEM achieves state-of-the-art performance across a range of settings, including the challenging case of non-monotone MAR mechanisms. Together, these results position FLOWGEM as a principled and practical alternative to existing imputation methods, and a decisive step towards closing the gap between theoretical rigor and empirical performance.
PaperID: 6764, Poster
Abstract: Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves the final prediction unchanged. We investigate where this failure originates by defining the temporal divergence vector \tau_l, the layer-wise representational difference induced by reversing temporal order. Tracking its magnitude across layers reveals a consistent temporal divergence profile where the divergence peaks at intermediate layers and progressively diminishes toward the output. We confirm this peak is specific to temporal reasoning and functionally critical for predictions, establishing that VideoLLMs acquire temporal information at intermediate layers but fail to maintain it to the output. This progressive fading motivates our method, Temporal Activation Injection (TAI), which extracts \tau_l at the peak of the profile for each input and reinjects it into subsequent layers following the measured decay. TAI requires no training and consistently improves temporal reasoning across three VideoLLMs and four benchmarks with negligible impact on non-temporal tasks.
PaperID: 6765, Poster
Abstract: Visual reinforcement learning (RL) is promising for real-world problems such as autonomous driving, robotic locomotion, and manipulation, but remains highly vulnerable to visual distribution shifts between training and deployment. Existing methods improve robustness through Q-consistency regularization, masking, and auxiliary objectives, yet most of them follow a coupled representation learning paradigm in which control-oriented and robustness-oriented representation learning are jointly optimized within the same training framework. We argue that this coupled structure can introduce harmful interference between control learning and robustness learning, degrading original-environment control performance and limiting generalization under visual shifts. To address this issue, we propose Separate Then Align Representations (STAR), which reformulates visual RL generalization as a problem of decoupled representation learning. Specifically, we first learn a control-relevant reference representation through online RL under weak augmentation, and then perform an offline post-representation alignment stage that maps strongly augmented inputs back to this reference representation. Experiments on RL-ViGen benchmarks spanning DMControl and Robosuite show that STAR achieves strong generalization under visual distribution shifts while better preserving performance in the original environment than prior coupled approaches. Our source codes are available in the supplementary material.
Authors:
Tianshu Zhu, Wenyu Zhang, Xiaoying Zuo, Tianlun, Haotian Zhao, yucheng-zeng, Jingnan gu, Daxiang Dong, Jianmin WuAbstract: Agentic reinforcement learning (RL) for software engineering spends much of its compute on stateful trajectories whose grouped binary rewards are highly skewed and weakly contrastive. We frame this as pass-rate control and show that the binary reward-side signal is strongest near a 50% rollout pass rate under four criteria: reward entropy, group-filtering survival, leave-one-out (RLOO) advantage energy under Group Relative Policy Optimization (GRPO), and success-failure pair count. We propose Prefix Sampling (PS), which replays self-generated trajectory prefixes to steer skewed groups toward this regime: successful prefixes give mostly failing groups a head start, while failing prefixes handicap mostly passing groups. Replayed states are reconstructed through the existing rollout path, and replayed tokens are masked from the loss so optimization applies only to current-policy continuations. On SWE-bench Verified, PS reaches the baseline high-score regime within evaluation variability while delivering 2.01×/1.55× end-to-end wall-clock speedups on Qwen3-14B/32B; the 14B peak improves from 0.274 to 0.295. AIME 2025 experiments on 4B/8B show the same pass-rate-control pattern, and 4B ablations attribute gains to replay, bidirectional coverage, and adaptive control.
PaperID: 6767, Poster
Abstract: Multimodal large language models (MLLMs) can process text presented as images, yet they often perform worse than when the same content is provided as textual tokens. We systematically diagnose this ``modality gap'' by evaluating seven MLLMs across seven benchmarks in five input modes, spanning both synthetically rendered text and realistic document images from arXiv PDFs to Wikipedia pages. We find that the gap is highly sensitive to rendering choices such as font and resolution, and that natural document images often match or exceed text-mode performance, suggesting the gap partly reflects evaluation artifacts rather than fundamental limitations. Through a grounded-theory error analysis of over 4,000 examples, we identify the primary cause: image input alone suppresses reasoning effort, with models producing 5--19× shorter outputs that skip step-by-step computation or reasoning. The reluctance to reason, not a failure of perception or knowledge retrieval, drives the performance gap, particularly on tasks requiring multi-step reasoning. We show that a simple lightweight on-policy self-distillation method by fine-tuning models on their own text-mode reasoning traces paired with image inputs closes this gap, raising image-mode accuracy to match or exceed text-mode performance with over 50% improvement, and the gains transfer to unseen benchmarks without catastrophic forgetting. Overall, our results and analyses provide a systematic understanding of the modality gap and suggest a practical path toward improving visual text understanding in multimodal language models.
PaperID: 6768, Poster
Abstract: LiDAR-based 3D Object Detection (L3OD) is a fundamental 3D perception task, yet its performance is often compromised by domain shifts arising from diverse sensor configurations and environmental conditions. Existing Test-Time Adaptation (TTA) methods primarily leverage self-supervision in the spatial domain, yet they frequently confound semantic content with domain style due to the sparse, non-Euclidean nature of point clouds. To overcome this limitation, we propose FIND (Frequency INvariance Disentanglement), a novel TTA framework that precisely disentangles domain-invariant content from domain-specific interference within the frequency domain. Our methodology utilizes B-Spline fitting to construct locally adaptive filters integrated into a dual-stream architecture: an Invariant Stream extracts robust features to guide pseudo-labeling, while a Specific Stream enforces consistency against domain perturbations. By combining invariance-driven alignment with stability-guided regularization, our approach dynamically extracts robust domain-invariant features while suppressing domain-specific interference. Extensive experiments on cross-dataset adaptation and robust corruption benchmarks demonstrate that FIND significantly outperforms state-of-the-art methods.
PaperID: 6769, Poster
Abstract: Long-horizon LLM agents trained with outcome-based RL exhibit a systematic failure mode: reasoning diverges from environmental observations at some step, and the error compounds across subsequent turns. We find that 74% of failed trajectories contain reasoning–observation inconsistency, and once it appears, 62% of subsequent turns remain inconsistent. Process rewards reduce but do not eliminate this: under free-form reasoning, agents learn to avoid penalties without genuinely tracking task progress or anticipating action outcomes. We propose ARC (Anchored Reasoning with Commitment), which requires the agent to make two explicit commitments at each step: a global progress statement that reconciles the agent's understanding of task state with accumulated observations, and a local outcome prediction that must be borne out by the next observation. Together, these commitments anchor the agent's reasoning to the actual environment state across turns, and provide well-defined targets for process supervision. To prevent interference between sparse outcome and dense process signals, we decouple their advantages via independent group normalization. On ALFWorld, WebShop, and WebArena, ARC improves over the strongest baseline by 6.4, 3.5, and 9.6 points respectively.
PaperID: 6770, Poster
Authors:
Yusen Wu, Yefan Wang, Jia Yee Tan, Guangyuan Dong, Shuang Chen, Jing Yang, Rongfeng Guo, Por L Yee, Yongtai LiuAbstract: KV-cache compression for long-horizon LLM agents is often guided by accumulated attention, yet attention frequency can retain repetitive error traces while discarding decisive instructions. We propose TaskKV, a streaming step-level retention policy with three contributions: (1) a signed directional utility label from the first-order change in action likelihood, distinguishing helpful steps from confusing distractors unlike unsigned saliency; (2) a structured residual scorer that treats attention mass and intent alignment as fixed priors and learns only a denoising correction, converging with 500 trajectories; and (3) an entropy-based regime classifier that falls back to contiguous retention when the scorer cannot confidently rank steps. With a 20% KV budget on AgentBench and AlfWorld, TaskKV preserves ~87--89% of full-cache success rate, improving over SnapKV by 5.0--7.9 SR points and over a strong heuristic by 2.9--7.3 points; on dense tasks, gains come from detecting when utility pruning should be disabled. Code is available at https://anonymous.4open.science/r/taskkv-F0D5.
PaperID: 6771, Poster
Abstract: Recovering 360^\circ HDR environment maps from unconstrained portraits is a fundamental yet ill-posed challenge in inverse rendering. Existing approaches typically suffer from a critical trade-off: traditional facial reflection-based methods yield low-fidelity, blurry results that lack practical utility, while recent generative models often resort to unconstrained panoramic hallucination (i.e., synthesizing environments with no physical grounding), which leads to severe overfitting and inaccurate lighting. In this paper, we present FaceProbe, an inpainting-driven framework that treats the human face as a reliable, physically-grounded light probe. Instead of synthesizing the environment from scratch, we reformulate the task as a masked completion and outpainting problem. By preserving the peripheral context of the input image and utilizing the portrait as a structural condition, our architecture leverages the generative priors of Diffusion Transformers (DiT) to extend sparse observations into high-fidelity, sharp 360^\circ HDR panoramas. To further eliminate color shifts and highlight inaccuracies, we introduce OLAT-GRPO, a post-training alignment mechanism based on Group Relative Policy Optimization. By subjecting the generated maps to image-based relighting on One-Light-At-a-Time (OLAT) datasets, we evaluate their physical validity through relighting consistency. This allows us to explicitly prune inconsistent denoising paths and align the model with multi-objective rewards, ensuring energy conservation and spectral accuracy. Extensive evaluations demonstrate that FaceProbe achieves a paradigm-shifting improvement. Most notably, in the critical portrait relighting task, our method reduces estimation errors by 54.4% in Angular Error and 48.5% in si-RMSE over state-of-the-art methods, delivering a robust and highly usable solution for photorealistic rendering and world model applications.
PaperID: 6772, Poster
Abstract: While dataset distillation has witnessed significant advances in image classification, its extension to dense prediction tasks such as semantic segmentation is plagued by two core bottlenecks: bi-level optimization-based approaches suffer from prohibitive computational costs stemming from pixel-wise gradient unrolling, whereas proxy-based generative methods merely infuse visual textures into fixed semantic masks derived from real data, thus failing to condense rich semantic knowledge into a compact set of informative, novel scene generations. To address both bottlenecks simultaneously, we propose a novel Sample-Set Reinforcement Learning (S^2-RL) framework for generative semantic segmentation dataset distillation (SSDD). Specifically, S^2-RL formulates SSDD as a diffusion model based text-to-image generation paradigm, enabling flexible generation of images with informative, novel semantic distributions via a concise text prompt. Furthermore, we fine-tune a diffusion policy via GRPO, leveraging a sample-set dual reward paradigm: for individual samples, it enforces semantic alignment with input text prompts and intra-class feature diversity; for sample sets, it maximizes inter-sample semantic distribution diversity by solving a contextual multi-armed bandit problem. These design choices enable S^2-RL to distill large-scale semantic segmentation datasets into a compact set of generated samples that encapsulate diverse, informative semantic knowledge without the need of pixel-wise gradient unrolling. Extensive experiments on the ADE20K and COCO datasets demonstrate that S^2-RL achieves substantial and consistent improvements over state-of-the-art baselines, establishing a new benchmark for SSDD.
Abstract: We present a novel theoretical analysis of Federated SARSA (FedSARSA) with linear function approximation and local training. We establish convergence guarantees for FedSARSA in the presence of heterogeneity, both in local transitions and rewards, providing the first sample and communication complexity bounds in this setting. At the core of our analysis is a new, exact multi-step error expansion for single-agent SARSA, which is of independent interest. Our analysis precisely quantifies the impact of heterogeneity, demonstrating the convergence of FedSARSA with multiple local updates. Crucially, we show that FedSARSA achieves linear speed-up with respect to the number of agents, up to higher-order terms due to Markovian sampling. Numerical experiments support our theoretical findings.
PaperID: 6774, Poster
Abstract: We study issues which may arise when using transformer models to formally reason about problems stated informally using natural language. In particular, we focus on transfer learning between tasks that are semantically equivalent but phrased differently. We assume that examples with a particular target label are only available for a subset of tasks, and aim to transfer zero-shot the ability to predict this label to other tasks. Using controlled experiments on propositional logic problems, we show that transformers are capable of such transfer, but that it can be impeded by the presence of task-specific tokens, due to the model picking up on correlations between these tokens and the label early in the training. As a simple solution, we show that introducing the target label later in the training dramatically improves transfer.
PaperID: 6775, Poster
Abstract: Reasoning-oriented post-training enables large language models (LLMs) to solve complex tasks via multi-step inference. Supervised fine-tuning (SFT) on high-quality reasoning traces is a particularly efficient approach, but its effectiveness depends critically on careful data curation. Existing pipelines rely on a costly generate-then-filter paradigm: large pools of long reasoning traces are produced by multiple teacher models and subsequently filtered for difficulty and diversity using additional LLMs, often requiring weeks of computation. In this work, we propose a fundamentally different approach that bypasses this pipeline by directly predicting the quality of reasoning data. We show that a brief LoRA-based adaptation, combined with evaluating loss on only 0.1–1% of each reasoning trace, suffices to estimate both difficulty and diversity. Our method is grounded in a geometric view of post-training, where pretrained LLMs lie in a low-rank, anisotropic loss basin. By probing this basin via low-rank perturbations, we construct compact loss signatures whose clustering captures diversity, while loss at curvature inflection points provides a robust proxy for difficulty. We also provide a simple criterion for teacher selection. Empirically, our approach reduces data curation time from weeks to hours while maintaining or improving performance. Remarkably, on mathematical reasoning, 5K generated examples match the highly filtered 89K-example in OpenThoughts. We further introduce a high-quality physics reasoning dataset.
PaperID: 6776, Poster
Abstract: Large language models (LLMs) have shown promise in medical question answering and clinical reasoning, yet their improvement remains constrained by static parametric knowledge and costly expert supervision. Self-evolving agents offer a promising alternative by enabling models to improve through iterative task generation and problem-solving. However, most existing self-evolving methods are designed for easily verifiable domains such as mathematics and coding, where solutions can be checked by exact answers or executable programs. Medical reasoning is fundamentally different: it is open-ended, knowledge-intensive, and often only partially verifiable. We present MedZERO, a self-evolving framework for open-ended medical reasoning. MedZERO couples an Examiner that generates frontier medical question-option pairs with a Reasoner that solves them through evidence-grounded multi-turn reasoning with external knowledge tools. To support reliable, continual improvement, MedZERO introduces controlled knowledge accumulation, which maintains temporary exploratory knowledge and curated persistent knowledge in reasoning. We evaluate MedZERO on five public medical reasoning benchmarks using 4B- and 8B-scale base models under open-ended evaluation. Across all settings, MedZERO consistently outperforms the underlying base models and prior self-evolving baselines, achieving up to 13.7 average accuracy-point gains over the next-best self-evolving baseline.
PaperID: 6777, Poster
Abstract: Efficient heterogeneous Multi-Agent Reinforcement Learning (MARL) in continuously changing environments is key to advancing MARL from simulation to the real world. However, it remains unclear whether current MARL methods can adapt under such dynamic scenarios. In this work, we show that existing methods face significant performance degradation in a novel continual heterogeneous MARL environment we construct, named slippery Multi-Agent Mujoco. We further demonstrate the connection between this phenomenon and eck (CHAIN), equipped with Heterogeneous Policy Bottleneck (HPB). In HPB, we extend the traditional information bottleneck into three parts, fitting, compression and heterogenization. Our HPB encourages agents to learn an efficient latent representation to adapt to changed environments, as well as disentangles specializations of agents in state perception, thereby encouraging heterogenization. We evaluate CHAIN on standard MA-Mujoco, our slippery MA-Mujoco and a multi-task continual MARL benchmark, MEAL, with various challenging tasks. The experimental results indicate that our CHAIN outperforms state-of-the-art heterogeneous MARL methods across key continual learning metrics.
PaperID: 6778, Poster
Abstract: Muon has recently demonstrated strong empirical performance in large language model training, but the theoretical role of momentum in Muon remains unclear. Existing analyses of Muon either remove momentum to study spectral updates in isolation, or retain momentum without explaining why it improves empirical performance. Our work bridges this gap by showing momentum in Muon acts as a spectral filter. Under a structured signal-plus-perturbation gradient model, we prove that momentum suppresses perturbations while preserving the dominant signal, thereby enlarging the spectral gap between them. This enlarged gap stabilizes the singular subspaces of the matrix passed to Muon's orthogonalization step, making the resulting update more reliable. We further show that applying momentum before orthogonalization achieves provably stronger alignment with the signal component of the gradient than either reversing this order or simply removing momentum. Experiments across diverse tasks, including LLM training, support our theoretical analysis. More broadly, our theory offers a starting point for understanding the benefits of momentum in other matrix-based optimizers.
PaperID: 6779, Poster
Authors: Dashiell Bhattacharyya
Abstract: Permutation-invariant machine learning architectures, such as DeepSets, Set Transformers, Graph Neural Networks, and k-GNNs, are each highly used and studied individually. We unify them into a single complexity-theoretic framework: noisy symmetric circuits. Each architecture falls into our hierarchy at a level k, parameterized by the order of interactions it captures. Within this framework, we tightly characterize the cost of symmetry. At level 1, where DeepSets operate, every symmetric function on boolean input computable by a neural network has an equivalent DeepSet, at the cost of an additive blowup nearly linear in the input size, and we exhibit an explicit function that requires this blowup. At level 2, where Set Transformers and GNNs operate, the same near-linear blowup suffices for WL-invariant functions, while non-WL-invariant functions are known to be uncomputable at level 2. These results show that, in the regimes we cover, WL is the only barrier to using a symmetric architecture; everywhere else, symmetry is low-cost. Experiments illustrate the near-linear threshold predicted by the theory.
PaperID: 6780, Poster
Authors: Zhangyi Wang, Bingnan Yu, Zongze Li
Abstract: Large language model (LLM) agents often keep the needed value in their prompt and still fail to pass it to a later tool call. We study this failure at the level of field routing: an earlier observation may expose an \textttorder\_id, policy flag, or nested item, and a later call must bind that field to the right argument. We formalize tool use as a boundary-indexed field-routing graph and introduce Information Utilization Rate (IUR), which measures whether each required dependency edge is realized in the later tool call. A 400-dependency expert audit, task-success correlations, and discriminant checks validate IUR as a field-use signal. On a 357-task public evaluation from ToolSandbox \textttSTATE\_DEPENDENCY (192 named scenario variants) and the original-repository \tau-bench retail/airline files (115+50 tasks), aggregate dependency-level IUR across GPT-4o, Claude-3.5-Sonnet, Llama-3.1-70B, and Qwen-2.5-72B falls from 451/548 (82.3%) at boundary distance \delta=1 to 144/283 (50.9%) at \delta=5, a 31.4-point drop, with mean context length only 3,094 tokens. Token-matched analyses and boundary transformations show that raw token length and lost-in-the-middle do not explain the curve by themselves. The same field-level view also leads to a compact intervention. Boundary-Aware Field Memory (BAFM) predicts live fields without gold dependencies at inference time and injects field reminders rather than whole observations. In paired GPT-4o runs, BAFM improves task success from 144/357 (40.3%) to 192/357 (53.8%), a 13.4-point gain over the base agent and 2.8 points over self-predicted retrieval-augmented generation. The results recast long-chain tool use as state propagation over fields: the central question is which prior values remain live for the next action.
Abstract: Non-Euclidean optimisation methods with matrix-valued updates, such as Muon and Scion, have recently shown strong empirical performance for training Transformer models, yet their theoretical advantages over Euclidean methods remain poorly understood. We address this gap in the heavy-tailed non-convex regime, where stochastic gradients have bounded p-th central moments, p \in (1,2]. We show that certain non-Euclidean methods achieve optimal sample complexity under stronger stationarity measures, while Euclidean methods incur additional dimension-dependent costs. As a consequence, for m × n matrices, Muon finds an \varepsilon-stationary point in nuclear norm within \mathcalO\left(\min\lbrace m, n\rbrace \frac\Delta_1 L\varepsilon^2 \left(\frac \sigma \varepsilon \right)^\frac p p-1\right) iterations, absorbing heavy-tailed noise without extra dimension dependence, unlike Euclidean methods. We further prove this dimension dependence is optimal for all first-order methods under nuclear-norm stationarity. Experiments on large language models support our theory. Surprisingly, our results suggest that other Schatten geometries beyond the spectral geometry of Muon can perform competitively in certain settings.
PaperID: 6782, Poster
Abstract: Multi-agent LLM systems often split work between a main agent with broad context and an executor subagent with limited local context. Communication can help recover missing information, but it can also add parsing burden, distract from the local decision, or induce protocol failures. We ask when the main agent should send a short message rather than raw context or no message, and when upgrading the main agent pays off. We formalize this as receiver-relative bounded coordination, where a message's value is the executor's next-step gain minus its protocol tax. This view yields four findings. First, compressed messages can outperform raw context when tax savings exceed losses from omitted information or decoder mismatch. Second, Bayes-sufficient compression, which preserves all information needed for the optimal decision, can still be worse than raw context when a bounded executor cannot decode or operationalize its surface form. Third, post-message failures can be localized into externalization, absorption, and action closure, with residual error concentrating in closure even after the right content reaches the executor. Fourth, a stronger upstream agent helps only when the executor's local view is weak enough for the added gain to exceed the added tax. Across six benchmarks, these regimes recur: the same multi-agent protocol raises ContextBench joint accuracy from 0.633 to 0.775 but lowers ToolSandbox from 0.889 to 0.653, helping in context-heavy regimes and hurting in locally sufficient ones. The same quantities drive an inference-time selector over communication actions, improving the accuracy-cost frontier on the two benchmarks where the full action set is evaluated. Code and results at https://anonymous.4open.science/r/Communication/
Abstract: Discrete flow matching generates text by iteratively transforming noise tokens into coherent language, but may require hundreds of forward passes. Distillation uses the multi-step trajectory to train a student to reproduce the process in a few steps. When the student underperforms, the usual explanation is insufficient capacity. We argue the opposite: the trajectory is the bottleneck, not the student. Each training trajectory is built through a chain of blind stochastic jumps with no evaluation of sequence quality; a single bad decision at an early midpoint propagates through subsequent steps, yet the student must imitate the result. Trajectory-Shaped Discrete Flow Matching (TS-DFM) replaces these blind jumps with guided navigation: a lightweight energy compass evaluates candidate continuations at each midpoint, selecting the most coherent. All shaping is training-only; inference cost is unchanged. On 170M-parameter language modeling, the shaped student at 8 steps achieves 32% lower perplexity than the 1 024-step teacher while being 128×faster, with gains consistent across source distributions and three evaluators of increasing scale. TS-DFM achieves the best perplexity of any discrete-generation baseline we compare against, including methods trained on 6×more data or using 5×larger models.
PaperID: 6784, Poster
Abstract: Cross-skeleton motion generation trains generative models to carry action structure and motion intention from one body to another. Yet an action-consistent target motion admits two observationally compatible explanations: source-preserving transfer and target-action recovery. We show that this ambiguity is structural rather than incidental: under standard generative objectives, the source-conditioned retargeting map is non-identifiable in sparse heterogeneous motion domains. Unpaired marginal matching yields \emphgauge non-identifiability: a relative gauge between skeleton-specific latent spaces allows different source-conditioned maps to induce the same training evidence. Sparse paired supervision admits the complementary failure mode, \emphconditional-mean degeneration: ambiguous action-level pairings drive squared-error objectives toward a target-action prototype that is independent of the source clip's motion intention. To make the missing evidence observable, we introduce \emphSource-Instance Fidelity (SIF), a diagnostic that tests whether the source-side geometry remains visible in the generated target motions after the target skeleton and action are fixed. Under this diagnostic, standard action-level success across animal motion domains recurrently coincides with behavior at the \emphsource-blind floor, while positive cases localize source-instance information to auxiliary motion-space constraints that narrow the observational equivalence class. The implication is that retargeting requires objectives and evaluations capable of identifying the source-conditioned map that a retargeting claim asserts.
Abstract: Dataset distillation seeks to synthesize a highly compact surrogate dataset that achieves performance comparable to the original dataset on downstream tasks. For the scenario where pre-trained self-supervised models serve as priors, traditional Linear Gradient Matching optimizes synthetic images by encouraging them to mimic the gradient updates induced by real images on the linear probe or classifier. However, this batch-level formulation requires loading thousands of real images and applying multiple differentiable augmentations to synthetic images at each distillation step, leading to substantial computational and memory overheads. In this paper, we revisit the linear gradient and theoretically derive that it is essentially a local relative distribution directed from target class centers toward non-target class centers, which we term ``flow”. This property causes instability and suboptimality, often necessitating expensive multiple augmentations to compensate. To address this, we introduce Statistical Flow Matching, a stable and efficient supervised learning framework that optimizes synthetic images by aligning global statistical flows in the original data. Our approach loads raw statistics only once and performs a single augmentation pass on the synthetic data, achieving performance comparable to or better than the state-of-the-art method with 10× less GPU memory usage and 4× faster distillation time. Moreover, increasing the number of augmentations for our method yields further performance gains while incurring lower additional cost.
PaperID: 6786, Poster
Abstract: 3D Weakly Supervised Semantic Segmentation (3D-WSSS) on point clouds aims to learn dense 3D semantics from extremely sparse point-level annotations. Existing methods follow a passive paradigm based on pseudo-label propagation or consistency regularization, leading to two key issues: 1) Lack of targeted intervention, where erroneous predictions are not explicitly identified and corrected, resulting in severe error accumulation; 2) Limited contextual awareness, where supervision is restricted to point- or view-level self-augmentation consistency, failing to capture the intricate contextual correlations embedded within the scene. To address these limitations, we propose a Policy-driven Active Residual Intervention (PARI) framework for 3D-WSSS, where an Active Intervention Agent (AIA) explicitly localizes high-risk points under the guidance of a Confidence-aware Reward Shaping (CRS) strategy, and a Bias-aware Refinement Module (BRM) subsequently executes targeted refinement by exploiting rich spatial and semantic scene contexts. Specifically, AIA actively localizes high-risk points from uncertainty maps and prediction priors, while BRM refines them through residual correction based on hybrid spatial-feature context. CRS introduces a high-confidence suppression reward that discourages unnecessary intervention on reliable predictions, and a low-confidence exploration reward that encourages intervention on uncertain predictions. Experiments on S3DIS and ScanNetV2 demonstrate state-of-the-art performance under various weak supervision settings with strong cross-backbone generalization, even surpassing the same-backbone fully supervised counterpart with 1% annotations.
PaperID: 6787, Poster
Abstract: Memory systems enable LLM agents to consolidate and retrieve relevant evidence from the factual knowledge accumulated through growing interaction histories for downstream reasoning. Existing approaches have explored diverse strategies for organizing and compressing these histories. However, balancing compression with retrieval effectiveness remains challenging: retaining too much content can cause relevant evidence to be obscured by redundant entries, while discarding too aggressively may remove content that later proves relevant. This amounts to a tradeoff between compressing redundancy and preserving enough structure to retrieve target evidence, as formalized by the information bottleneck. To this end, we propose , which organizes memory as a compression hierarchy where each level compresses redundancy further while retaining the structure needed for retrieval at that level. In this hierarchy, evidence is progressively compressed from detailed records through extracted keywords to topic groups. This enables retrieval to locate target evidence by searching across levels of the hierarchy. Comprehensive experiments demonstrate that MemCoRe outperforms existing state-of-the-art baselines.
Authors:
Zhengrong Yue, Taihang Hu, Mengting Chen, Haiyu Zhang, Zihao Pan, Tao Liu, Zikang Wang, Jinsong Lan, Xiaoyong Zhu, Bo Zheng, Yali WangAbstract: Tokenizers are a crucial component of latent diffusion models, as they define the latent space in which diffusion models operate. However, existing tokenizers are primarily designed to improve reconstruction fidelity or inherit pretrained representations, leaving unclear what kind of latent space is truly friendly for generative modeling. In this paper, we study this question from the perspective of latent manifold organization. By constructing controlled tokenizer variants, we identify three key properties of a diffusion-friendly latent manifold: coherent spatial structure, local manifold continuity, and global manifold semantics. We find that these properties are more consistent with downstream generation quality than reconstruction fidelity. Motivated by this finding, we propose the Prior-Aligned AutoEncoder (PAE), which explicitly shapes the latent manifold instead of leaving diffusion-friendly manifold to emerge indirectly from reconstruction or inheritance. Specifically, PAE leverages refined VFM-derived priors and perturbation-based regularization to turn spatial structure, local continuity, and global semantic organization into explicit training objectives. On ImageNet 256×256, PAE improves both training efficiency and generation quality over existing tokenizers, reaching comparable performance up to 13× faster than RAE under the same LightningDiT setup and achieving a new state-of-the-art gFID of 1.03. These results highlight the importance of organizing the latent manifold for latent diffusion models.
PaperID: 6789, Poster
Authors: Amir Asiaee, Samhita Pal
Abstract: Randomized controlled trials (RCTs) identify trial-anchored treatment effects but are often too small for reliable heterogeneity estimation; observational studies (OS) are larger but confounded and only partially overlap the trial in measured covariates. We propose Bayesian Calibrated ALignment under covariate Mismatch (B-CALM), a Bayesian borrowing framework for RCT-anchored conditional average treatment effect (CATE) estimation. B-CALM maps source-specific covariates into a shared latent state, jointly models trial and observational outcome surfaces, and introduces baseline-bias and comparative-bias functions that absorb how the OS departs from the trial estimand. The comparative-bias prior becomes an explicit sensitivity knob: we prove a function-valued bias-limited information bound showing that observational contrast information about the trial treatment-effect surface is capped by the prior precision of this bias function, with a scalar corollary in which the effective sample size (ESS) saturates as OS sample size grows. A PAC-Bayes-style risk decomposition separates RCT empirical risk, latent alignment, and residual calibration of the debiased OS surface. Across synthetic, semi-synthetic, and pediatric-obesity external-control studies, B-CALM delivers calibrated credible intervals and low negative transfer while pooled and forest baselines can become overconfident under comparative bias.
Abstract: Cell Painting is a high-content morphological profiling assay widely used for phenotype-based biological inference, with mechanism of action (MOA) prediction as a central application. Existing approaches largely formulate Cell Painting-based inference as representation matching, assigning predictions from nearby reference perturbations in morphological feature space. However, retrieved neighbors are often noisy and partially misleading evidence due to batch effects, non-specific cytotoxicity, phenotypic convergence, and source-dependent variability. We reformulate Cell Painting-based MOA prediction as a calibrated evidence reasoning problem, where retrieved neighbors are treated as uncertain observations that must be evaluated, compared, and sometimes rejected before supporting a mechanistic conclusion. We propose PhenoAIR, a reliability-aware multi-agent framework that maintains a candidate-centric evidence memory and performs controller-guided refinement over phenotype- and mechanism-side evidence. PhenoAIR uses offline reference-set calibration to weight evidence by source reliability, phenotype stability, and mechanism-level confusion. We evaluate PhenoAIR on a benchmark constructed from JUMP Cell Painting profiles and annotations, covering controlled, realistic, and discovery-oriented open-world MOA prediction settings. PhenoAIR outperforms representation-matching and LLM-based baselines across all settings.
Authors: Sehmimul Hoque, Roger Melko, Pooya Ronagh
Abstract: We introduce a novel technique for scalable sampling of spin-system states with continuous symmetries using diffusion models. By applying our approach to the XY model, a fundamental continuous-spin model in condensed matter physics, we show that our technique addresses the shortfalls of the Markov chain Monte Carlo (MCMC) in generalization to varying system sizes. More specifically, we show that training a temperature-conditioned diffusion model on smaller-size XY model lattices enables the generation of accurate samples in larger lattice sizes. By tracking physically important observables of the model, such as spin correlations, our experiments demonstrate that diffusion sampling followed by few MCMC steps reduces the thermalization time by an order of magnitude relative to the standard MCMC with random initialization. Our study provides valuable insight as to how generative models can be used to study continuous-state condensed matter systems at scale.
PaperID: 6792, Poster
Authors: Vezha Boboeva, Alberto Pezzotta, George Dimitriadis, Athena Akrami
Abstract: A hallmark of intelligent behavior is the ability to extract abstract relational structure from temporal sequences -- recognizing, for instance, that aab, ccd, and eef all follow the same underlying pattern, regardless of the specific elements involved. This capacity, observed across species and sensory modalities, is thought to underlie the formation of cognitive schemas: compressed internal models that support rapid generalization to novel experiences. Yet the neural circuit mechanisms by which such abstract, identity-independent representations emerge from sequential experience remain largely unknown. Here, we investigate this question using Recurrent Neural Networks (RNNs) as mechanistic models of neural circuits, trained to classify sequences based on their latent algebraic patterns (e.g., aab, aad \to AAB; aba, aca \to ABA) without supervision on intermediate transitions. We demonstrate that RNNs learn low-dimensional representations that mirror the hierarchical generative structure of the sequences, and that this abstraction is mechanistically supported by the emergence of low-rank recurrent connectivity. The leading singular component of the recurrent weights integrates relational transition information -- whether consecutive tokens are the same or different -- across time, driving the formation of a structured, tree-like geometry in the population state space. Through singular vector ablation, we establish a causal role for this component: removing it selectively erases memory for earlier transitions while leaving local, single-step sensitivity intact. Finally, while RNNs trained on next-token prediction do not acquire these abstract representations, transferring the low-rank scaffold learned from classification significantly accelerates learning and improves generalization -- an effect specific to the abstract structure of the scaffold rather than generic statistical pretraining. These findings offer a computational account of how task demands shape recurrent connectivity to support temporal abstraction, with implications for schema formation in biological brains.
PaperID: 6793, Poster
Abstract: Spiking Neural Networks (SNNs) promise energy-efficient inference, but adapting large pre-trained SNNs to new tasks is expensive because BPTT memory and runtime scale with the number of simulation steps T. We observe that weights in pre-trained SNNs are strongly concentrated near zero, and under thresholded dynamics many such synapses are functionally silent. Based on this, we propose Spike-SFT, a two-stage adaptation framework. Stage 1 (Selective Parameter Enhancement, SPE) fine-tunes only a small-magnitude subset via masked updates with low-rank regularization, progressive re-selection, and sparse-gradient storage, reducing memory while maintaining competitive accuracy. Stage 2 (PickIt) fuses multiple SPE-adapted models by alignment and interference-aware delta merging with spike-statistics calibration, yielding a single merged model with zero inference-time overhead. Across several SNN backbones and benchmarks, Spike-SFT offers a favorable accuracy--cost trade-off, improving fine-tuning efficiency (time/memory) while retaining strong accuracy, and enabling zero-overhead weight-space fusion via PickIt.
PaperID: 6794, Poster
Abstract: Harmful fine-tuning attack becomes a concerning safety risk for mainstream fine-tuning-as-a-service providers, as attackers can submit harmful data to the API to compromise the safety alignment of the large language models. In this paper, we first explore an intuitive gradient mixing solution, and derive a key property ensuring the success of defense -- \emphtaking a fine-tuning update that has higher cosine similarity with the safety gradient can mitigate harmful fine-tuning. Motivated by this key property, we design an alignment-stage defense, dubbed Tcell. The core contribution of Tcell is a regularizer that align the harmful gradient and the safety gradient during safety alignment, which ensures the harmful fine-tuning update exhibits high cosine similarity with the safety gradient, achieving \emphgradient alignment. The benefit of gradient alignment is supported by i) empirical and theoretical interpretation and ii) comparison to alternative design without gradient alignment. Code is available at https://anonymous.4open.science/r/Tcell-1C15/.
PaperID: 6795, Poster
Abstract: Deep neural networks are highly vulnerable to bit corruption in stored weights, whether from naturally occurring memory faults or deliberate fault-injection attacks. Existing defenses remain fundamentally limited: fault-tolerance methods degrade under cumulative corruption and remain vulnerable to strong white-box attacks, while ECC-based recovery methods are often tied to specific precisions or architectures and expose concentrated vulnerable surfaces. We propose a simple training-based recovery framework that endows model weights with recoverable structure through block-wise zero-sum constraints enforced by differentiable projection. The resulting method incurs no parameter-space overhead and generalizes naturally across architectures and numerical precisions. We further develop adaptive adversarial bit-flip attacks tailored to recovery-based defenses, covering both our method and prior ECC-based approaches. Across diverse architectures, precisions, datasets, and threat models, our method consistently delivers strong robustness and recoverability while preserving clean performance.
Abstract: Causal representation learning for time series has developed strong identifiability results in discrete-time latent causal models, but identifiability in continuous-time latent stochastic differential equation (SDE) models remains largely open. We address this gap using environment-induced shifts in diffusion covariance. We study additive-noise latent SDEs observed through an unknown nonlinear diffeomorphism, with shared drift but environment-specific diffusion covariance. We show that two diagonal diffusion regimes with pairwise distinct coordinate-wise variance ratios identify the latent coordinates up to permutation and scaling, without any sparsity assumption on the drift. We first prove this result for linear Ornstein--Uhlenbeck systems and then extend it to general additive-noise latent SDEs. Under mild smoothness, the instantaneous drift-Jacobian causal graph is identifiable up to the same permutation. We propose a two-stage estimator for latent disentanglement and optional graph recovery; experiments on synthetic systems confirm the predicted identifiability boundary, and an application to Hardanger Bridge monitoring data illustrates the approach on real sensor trajectories.
PaperID: 6797, Poster
Abstract: In standard acoustic anomaly detection (ASD), models are usually trained separately for each machine or domain using explicit metadata. However, in realistic deployments, machine identifiers are incomplete, unreliable, or simply unavailable. Thus, metadata-free universal ASD asks a single model to monitor mixed machine populations without machine IDs. In this paper, we argue that the main failure mode is not a weak detector family, but representation entanglement. The heterogeneous normal modes overlap in frozen pre-trained feature spaces, making anomaly scoring and approximate retrieval unreliable. We study this problem directly and propose HARMONY (Hierarchical Anchor Retrieval on Manifold for Oblivious-source acoustic aNomaly detection), an anchor-guided geometric repartitioning framework that reorganizes mixed-source features into compact Voronoi regions and reuses this structure for hierarchical retrieval. On dataset DCASE 2020 and MIMII, HARMONY improves mean AUC from 75.40% to 90.00% (+14.60% absolute improvement) over the strongest unified baseline and from 83.90% to 90.00% over a strong foundation-feature baseline. On a single CPU, its hierarchical retrieval achieves a 7.8× speedup over exhaustive search with only a 0.20% AUC drop. These results suggest that incorporating explicit geometric partitioning can improve both detection accuracy and retrieval efficiency for ASD. Code and data are available at the anonymous URL: https://anonymous.4open.science/r/HARMONY-26CC.
PaperID: 6798, Poster
Abstract: Continual learning under adversarial perturbations remains underexplored, despite its importance in security-critical deployments where attack patterns evolve over time. Existing continual adversarial defense methods mainly rely on replay or prediction-space regularization, but often fail to preserve robust representations across sequentially arriving attack types. In this paper, we study continual adversarial defense, where a model receives a sequence of tasks induced by different white-box attacks and must acquire robustness to new attacks without forgetting previously learned defenses. We address this challenging learning scenario by proposing the Learning Robust Representation Framework (LRRF), a representation-centric framework that improves robustness through hierarchical alignment at three complementary levels. First, we introduce the Dynamic Representation Matching (DRM) mechanism that aligns clean and adversarial feature distributions within each task to reduce the clean-robustness trade-off. Second, a Cross-Task Invariant Representation (CTIR) mechanism is proposed to regularize the current representation network against an accumulated network from previous tasks, encouraging attack-invariant features to persist over time. Third, a Knowledge Consolidation Optimization (KCO) mechanism is proposed to match class-consistent clean and adversarial features using replayed samples, which further stabilizes category structure across tasks. The empirical results show that the proposed approach achieves state-of-the-art performance.
PaperID: 6799, Poster
Abstract: Long-tailed recognition remains challenging not only because rare classes provide limited supervision, but also because their evidence is difficult to accumulate during standard stochastic training. We characterize this failure as tail signal erosion: in conflict-dominant regimes, frequent-class updates can dilute sparse rare-class signals before they are sufficiently reinforced. To address this issue, we propose TailCon, a memory consolidation framework that separates rare-evidence acquisition from parametric integration. During online acquisition, TailCon stores uncertain or rare representations in detached episodic memory, while a parametric prediction pathway learns global decision structure. To avoid relying only on test-time retrieval, TailCon introduces FOCUS, a periodic consolidation stage that distills stored memory evidence into the parametric prediction head through a low-capacity alignment surrogate. We provide a local optimization analysis that scopes when this separation is beneficial in sparse-recurrence, conflict-dominant regimes. Experiments on CIFAR-100-LT, CIFAR-10-LT, ImageNet-LT, and iNaturalist 2018 show consistent improvements over representative rebalancing, decoupled, expert-based, and prototype-oriented baselines under the evaluated protocols. Memory-disabled inference further shows that part of the stored evidence is transferred into the parametric prediction head rather than remaining only in the retrieval branch. An anonymized implementation is included in the Supplementary Material.
PaperID: 6800, Poster
Abstract: Existing hallucination detectors for retrieval-augmented generation (RAG), whether based on model outputs (e.g., likelihood) or internal representations, typically apply a uniform detection strategy across responses. However, we find that different detection signals are informative in different response regimes, making a uniform strategy insufficient for both short- and long-form generation. In short-form tasks such as factoid QA, hallucination typically appears as selecting an incorrect final answer among multiple context-relevant candidates. In this regime, model-output-based signals that discriminate among plausible answers are particularly important. We therefore propose a bidirectional likelihood measure that evaluates logical consistency based on the retrieved evidence and a reasoning step generated after the response. In contrast, long-form descriptive responses are more prone to gradual drift: as generation proceeds, the model increasingly relies on its own prior generations and parametric knowledge, while unsupported continuations may remain locally plausible, making output-based cues less informative. For this regime, we measure context--knowledge conflict by tracking the directional alignment of context-grounding versus internal-knowledge contributions in the hidden states. Building on these regime-specific signals, we introduce ARGUS, an adaptive hallucination detector that emphasizes the more effective signal according to the response regime. Experiments on multiple RAG benchmarks show that our method achieves strong performance across both short- and long-form settings.
PaperID: 6801, Poster
Authors: Hannah Diehl, Justin Steil, Ashia Wilson
Abstract: Semivalues are widely used to assign credit and guide data curation decisions, yet their outputs depend on a utility function that is not uniquely determined by the task. We study partition-based decisions, which directly model data curation tasks such as selecting high-quality subsets or flagging noisy examples, under two structurally unavoidable sources of utility underspecification: monotone transformations of performance scores and unconstrained small-sample behavior. We formalize outcome robustness in this setting as invariance of the induced partition across admissible utility respecifications. We establish that partition robustness is necessary for guaranteed success of semivalue-based data selection. To assess this condition in practice, we provide algorithms that certify partition robustness or produce a concrete witness utility demonstrating instability, requiring no utility evaluations beyond those already computed for the semivalues themselves, with formal correctness guarantees under both exact and approximate computation. We show that utilities admitting a coalitionally dominant top-k subset, where high-quality points contribute more than low-quality points across every possible coalition, guarantee robust partition decisions, subsuming the modular utility case previously identified as sufficient. Together, these results provide a practical framework for determining when semivalue-based selection is well-defined under utility ambiguity.
Authors: Ji Gao, Caleb Ju, Guanghui Lan, Zhaohui Tong
Abstract: Policy Dual Averaging (PDA) offers a principled Policy Mirror Descent (PMD) framework that more naturally admits value function approximation than standard PMD, enabling the use of approximate advantage (or Q-) functions while retaining strong convergence guarantees. However, applying PDA in continuous state and action spaces remains computationally challenging, since action selection involves solving an optimization sub-problem at each decision step. In this paper, we propose actor-accelerated PDA, which uses a learned policy network to approximate the solution of the optimization sub-problems, significantly reducing runtime. We provide a theoretical analysis that quantifies how actor approximation error impacts the convergence of PDA under suitable assumptions. We then evaluate its performance on several benchmarks in robotics, control, and operations research problems. Actor-accelerated PDA achieves superior performance compared to popular on-policy baselines such as Proximal Policy Optimization (PPO). Overall, our results take a significant step toward bridging the gap between the theoretical advantages of PDA and its practical deployment in continuous-action problems with function approximation.
PaperID: 6803, Poster
Abstract: This work investigates the challenge of robust graph learning within the framework of graph foundation models. Prior studies primarily rely on data-centric structural purification or adversarial augmentation during training to achieve adversarial robustness. However, in practical zero-shot inference scenarios, such methods exhibit significant vulnerability to unseen adversarial structures owing to the static nature of their defense mechanisms and the prohibitive computational cost associated with retraining. In this paper, we propose a novel framework termed Test-time gRaph recAlibration for enhancing robust zero-shot inferenCE (TRACE). The core mechanism of TRACE involves the selective learning of a structural anti-attack directly on noisy graphs during the inference phase, which neutralizes adversarial perturbations while maintaining performance on clean graphs. Specifically, TRACE introduces a non-parametric diagnostic metric based on smoothed spectral entropy to quantify the structural-semantic misalignment induced by adversarial attacks, thereby serving as an adaptive trigger for recalibration. Furthermore, a uniformity-guided optimization objective is formalized to leverage semantic anchors from pretrained text encoders, guiding the distorted graph back to the clean manifold. To ensure computational efficiency for sparse graph encoders, a first-order gradient-based edge flipping strategy is employed to reconstruct the optimal graph structure directly within the discrete domain. Extensive experiments conducted across various benchmark datasets and graph attacks demonstrate the superiority of TRACE over existing state-of-the-art baselines.
PaperID: 6804, Poster
Abstract: Streaming video understanding (SVU) requires models to answer user queries over continuously evolving video streams. Unlike offline video QA, SVU must handle queries whose evidence may lie in the past, appear in the current scene, or emerge only after future observations. Existing streaming video systems often rely on compact visual memories or fixed evidence-access schemes, which can lose fine-grained details and fail to adapt the answering process to the temporal intent of each query. In this paper, we propose TimeTraveler, a streaming VQA framework that performs temporal strategy planning with a Time Dictionary. Inspired by the need to search across different temporal regions of a stream, TimeTraveler stores each observed moment as a timestamp-indexed structured caption and uses this dictionary as an explicit source of queryable evidence. Given a query, TimeTraveler determines whether to recall past evidence, read the present context, or wait for future observations, and then applies a strategy-specific evidence acquisition process. Comparative experiments on streaming video QA benchmarks show that TimeTraveler improves streaming question answering by preserving fine-grained temporal evidence and selecting the appropriate temporal strategy for each query.
PaperID: 6805, Poster
Authors:
Seungjae Han, Joshua Y You, Soi Kim, Minho Eom, Jinho Park, Seunghee Park, Byeongwook Lee, Young-Gyu YoonAbstract: Individual neurons exhibit diverse molecular, morphological, and electrophysiological properties that shape their distinct roles in neural circuits. While these multimodal features are largely shared across individuals, they remain an underutilized source of information in analyzing neural population dynamics. We posit that these shared biological profiles provide a common reference for relating neural dynamics across subjects and recording sessions. Our framework leverages this structure by routing neurons into latent groups based on their biological profiles and analyzing population dynamics through group-level representations. Because these assignments are anchored in biological properties that are consistent across subjects, the learned organization supports transfer to unseen animals without retraining or neuron-wise correspondence. We evaluate our framework on whole-brain calcium imaging in larval zebrafish and large-scale electrophysiology in the mouse visual cortex. Across both species and recording modalities, the framework delivers state-of-the-art performance in both neural decoding and manifold analysis: it achieves superior decoding of behavior and stimulus labels while better preserving the temporal consistency and local geometric structure of latent trajectories on unseen data. Our experiments demonstrate that not only does providing biological features improve baseline performance, but our proposed framework achieves even stronger results by using these properties to explicitly structure the population rather than treating them as auxiliary inputs. Together, these results suggest that organizing neural populations by their multimodal biological features offers a promising framework for generalizing population dynamics across subjects and recording sessions.
PaperID: 6806, Poster
Authors: Sahani Pathiraja, Panos Parpas
Abstract: We introduce a new algorithm to sample from a probability density \pi \propto e^-V with V(x): \mathbbR^d \rightarrow \mathbbR a strongly multi-modal potential. Such distributions are challenging to sample from using standard gradient based methods, as exploration typically relies on inefficient diffusions. Inspired by connections between Schrodinger operators, Langevin diffusions and Morse theory, we develop a particle based method that exploits fast relaxation to transition pathways of the so-called Witten PDE on 1-forms (vector fields). Once pathways joining modes are sufficiently well explored, the method relaxes to inexpensive Langevin dynamics for sampling within approximately convex basins. Numerical experiments demonstrate superior performance over cheap gradient based sampling methods and competitive performance at lower cost compared to a state of the art method, parallel tempering.
PaperID: 6807, Poster
Abstract: We study the decoupled multi-armed bandit problem, where the player selects one arm to exploit (incurring its loss) and one arm to explore (observing its loss) at each round. We propose two algorithms: D-Exp3-Ada, based on Exp3 with an AdaGrad learning rate, and D-TINF-SPM, based on FTRL with Tsallis entropy and the stability-penalty matching learning rate. Here, K denotes the number of arms, T the time horizon, and \Delta_i the suboptimality gap of arm i. D-Exp3-Ada achieves \mathcalO(\sqrtKT\log K) regret in adversarial regimes and \mathcalO(\sum_i \neq i^\star \frac\log K\Delta_i) in stochastic regimes, combining simplicity of algorithm design with strong guarantees. D-TINF-SPM achieves the minimax optimal \mathcalO(\sqrtKT) in adversarial regimes and \mathcalO(\min\\sqrtK\sum_i \neq i^\star \frac1\Delta_i^2, \sum_i \neq i^\star \frac\log K\Delta_i\) in stochastic regimes. We also provide a refined analysis of the prior best algorithm. Finally, we derive instance-dependent lower bounds under natural monotonicity and permutation-invariance assumptions on regret upper bounds, proving that our algorithms are optimal up to \mathcalO((\log K)^2) within this class.
PaperID: 6808, Poster
Abstract: Object detection in high-resolution wide (HRW) shots, where a single image can span gigapixel resolutions and kilometre-scale fields of view, suffers from extreme foreground sparsity, dramatically varying foreground density from sparse to crowded scenes, and object clusters whose spatial extent ranges from dozens to thousands of pixels within the same dataset. State-of-the-art sparse vision transformers for gigapixel detection select a fixed top-k fraction of windows by hand-crafted variance scores, which over-promote textured but object-free background such as foliage and patterned facades, cannot adapt the keep ratio to the wide range of crowd density, and confine each kept window to a 7x7 receptive field that is far smaller than the spatial extent of many real-world object clusters. We propose SpanFormer, a sparse vision transformer that broadens the computational span at three complementary granularities: (i) Prototype Routing replaces variance with learnable foreground/background prototypes regularised by a balance constraint that prevents prototype collapse, widening the semantic span of the scoring criterion; (ii) Dynamic Top-k predicts a per-image keep ratio with a lightweight selector, making the compute span elastic so that crowded scenes receive proportionally more compute while near-empty scenes are skipped; (iii) Cluster Memory Attention performs union-find clustering over the kept windows by joint cosine similarity and Chebyshev spatial radius and augments each window's K and V with cross-window memory drawn from its cluster, extending the attention span beyond the local window with all projection parameters shared with the local branch and therefore no added projection cost. On the PANDA gigapixel benchmark, SpanFormer improves AP50 from 78.0% to 80.3% over the SparseFormer baseline while reducing backbone FLOPs by over 75%.
PaperID: 6809, Poster
Authors: Yiming Zhang, Reza Hassanpour, Georgi Gaydadjiev
Abstract: Neural network training is typically monitored through loss and accuracy curves, but these quantities do not directly reveal when the backbone begins to encode training and test data differently. A train-test negative log-likelihood (NLL) gap shows probabilistic separation between training and test examples, but it does not reveal whether this separation is present in the backbone geometry. We introduce the Volume Ratio Test (VRT), a permutation-based two-sample test for l2-normalized backbone embeddings that compares local angular neighborhood structure on the unit hypersphere. VRT converts k-nearest-neighbor angles into spherical-cap log-volume spacings and requires no kernel bandwidth selection or learned discriminator. Across 14 model configurations spanning embedding dimension D = 32 to D = 2,048, VRT matches or precedes the earliest main baseline on sustained backbone divergence. In several CIFAR-10 ResNet settings, VRT is the only method among MMD, MMDAgg, Energy Distance, and C2ST to detect sustained backbone divergence. When these baselines lag, VRT leads them by 25-60 epochs. Additional nearest-neighbor methods show that Schilling and Henze either reject 16-60 epochs later or remain null where VRT rejects. On ResNet-50/CIFAR-100, VRT first rejects at epoch 61, whereas MMD, MMDAgg, and C2ST first reject at epoch 121. On this model, the train-test NLL gap is already 0.563 at epoch 62 and never falls below this value. More generally, sustained VRT rejection is accompanied by persistent NLL separation, but large NLL gaps can occur without VRT rejection. VRT therefore distinguishes train--test separation in NLL from backbone-geometric divergence.
PaperID: 6810, Poster
Authors:
Runze Fan, Shiyi Jiang, Junsheng Wang, Shuangjun Xie, Xiaodan Li, Yong LiAbstract: Multi-agent LLM systems built on frontier backbones often fail to improve reliably as the swarm grows: under strong backbone models, performance frequently plateaus or declines beyond a small number of agents. In a budget-matched controlled intervention on SWE-bench Verified (N=8, Claude Opus 4.5 backbone, 64-call budget per instance), varying only the communication policy moves resolve rate from 80.4% under random broadcast coordination (below the 80.9% single-shot baseline) to 85.5% under a learned orchestrator, a 5.1-point spread under matched compute. We introduce OrchestraRL, an orchestration layer trained over frozen LLMs. A centralized Macro-Orchestrator reads the swarm's context entropy H(t), computed from the spread of agent output embeddings, and outputs communication mode (via learned entropy thresholds) and topology; per-agent Micro-Coordinators decide whom each agent addresses, when to speak, and at what granularity. Both are trained with REINFORCE and a learned value baseline. Across four benchmarks (SWE-bench Verified, GPQA Diamond, LiveCodeBench v6, SimpleQA), OrchestraRL is comparable to the strongest baseline at N=2 and achieves the best resolve rate at N \in \4, 8\ on every benchmark, with the gap widening as N grows. The learned policy recovers interpretable task-dependent communication patterns, including clustered exploration, star-based synthesis, and pipeline reasoning, without these structures being hand-coded.
PaperID: 6811, Poster
Authors: David Li, Angela Li
Abstract: We introduce Phase-Interference Decision Trees (PIDT), a differentiable routing model for tabular prediction in which each node propagates a bounded complex state rather than only a scalar routing probability. The main model uses two-state norm-preserving branch maps: a gate controls how much amplitude is sent to the left and right children, while learned unitary maps rotate the latent state before subsequent splits. This formulation contains ordinary soft decision trees as the phase-free scalar case, and two-state routing creates explicit score-level cross terms among latent-state trajectories. We give a self-contained PIDT definition with exact mass conservation, prove scalar phase cancellation, establish a scoped containment relation with soft trees, derive an algebraic interference decomposition, give a companion construction for shared two-state parity routing, and derive a fixed-architecture empirical Rademacher upper bound for bounded two-state PIDT scores. We evaluate PIDT on a ten-dataset OpenML suite of low-class-count tabular tasks under matched differentiable-tree controls. Across ten seeds per dataset, the main summaries track accuracy as context and report small descriptive NLL and Brier deltas relative to the scalar PIDT-1S, phase-frozen PFB, and SDT controls. The empirical study is framed as a mechanism-level study for compact differentiable trees, not as a broad tabular-performance claim.
Authors: Hyeongyeol Lim, Hongjun Yoon, Eunjin Jang, Daeky Jeong, Won J Cho, HWAMIN LEE
Abstract: Current virtual staining approaches offer the potential for time- and cost-efficient biomarker quantification in cancer diagnostics and prognostics. However, patch-wise inference for gigapixel whole slide images (WSIs) fails to maintain spatial continuity, yielding artifacts that cause catastrophic mismatches with ground-truth images. Although pathology Vision Foundation Models (VFMs) offer rich representations, their self-attention causes varying global contexts to produce inconsistent embeddings for the same physical region. We formalize and validate this ``context contamination'' as a sheaf-theoretic problem where these embeddings form a presheaf that violates the gluing axiom. To address this, we propose SheafStain, a new approach that reinterprets VFM features as sheaf-like sections for spatially and biologically coherent virtual staining. Specifically, SheafStain integrates class and patch tokens into a Schr\"odinger Bridge framework as sheaf-like sections. While the class token anchors biological consistency, patch tokens form a per-position spatial map. A backbone co-pretrained on Hematoxylin \& Eosin (H\&E) and Immunohistochemistry (IHC) yields non-degenerate cross-stain stalks, so a single VFM feature space supervises both input conditioning and output stain alignment. Departing from prior work that evaluates on isolated 256 × 256 patches and either random-crops or resizes the 1024 × 1024 ground truth, we translate at 256 × 256 and evaluate on the stitched 1024 × 1024 outputs across HER2, ER, PR, and Ki-67. SheafStain demonstrates promising results against six prior methods while mitigating patch-boundary stitching artifacts. Code will soon be released upon acceptance.
Abstract: Implicit in-weight multi-hop reasoning---composing multiple pieces of parametric knowledge within a single forward pass---is a fundamental yet challenging task of language models. Vanilla transformers consistently fail at this task, and even , whose recurrent structure naturally fits its iterative nature, generalize imperfectly, particularly out-of-distribution (OOD). We make two contributions toward closing this gap. First, on a symbolic two-hop reasoning task, we mechanistically diagnose why looped transformers fall short: after the first loop, the intermediate answer is already decodable from the residual stream with certain probability, yet the continuous hidden vector carrying it is noisy and geometrically misaligned with the clean discrete embedding of the intermediate answer that the next loop would ideally consume. A training-free intervention that re-aligns this representation nearly closes the OOD gap, indicating that the representational mismatch is a bottleneck. Second, building on this insight, we propose ntinuous hidden-state channel. DiscoLoop achieves near-perfect two-hop accuracy with substantially less training across symbolic and synthetic-language multi-hop reasoning tasks. When applied to real pretraining, DiscoLoop attains lower training loss and stronger performance on pretraining benchmarks than looped-transformer baselines, suggesting that the mixed-channel design transfers to practical language modeling.
Abstract: We consider time discretization for score-based diffusion models to generate samples from a learned reverse-time dynamic on a finite grid. Uniform and hand-crafted grids can be suboptimal given a budget on the number of time steps. We introduce Adaptive Reparameterized Time (ART), which controls the clock speed of a reparameterized time variable to redistribute computation along the sampling trajectory while preserving the terminal time, with the objective of minimizing the aggregate Euler discretization error. We derive a randomized companion ART-RL that recasts ART as a continuous-time reinforcement learning problem with Gaussian policies, and prove a two-directional bridge between the two: the deterministic ART optimum lifts to an optimal Gaussian policy, and conversely any optimal Gaussian policy must recover the ART control through its mean. This bridge turns continuous-time actor--critic learning into a principled, rather than heuristic, route to the deterministic timestep optimum. Within the official EDM pipeline, ART-RL improves FID on CIFAR--10 across a wide range of budgets; after one-time offline training, the distilled deterministic schedule transfers without retraining to AFHQv2, FFHQ, and ImageNet at no extra inference cost.
PaperID: 6815, Poster
Abstract: Unsupervised domain adaptation has become an important paradigm for mitigating distribution shift in time series classification. However, existing time series domain adaptation methods typically align source and target data at a fixed time scale and simply fuse time-frequency features before uniform alignment. Furthermore, they do not consider the impact of noisy samples on distribution alignment. To address these limitations, we propose CSPAN (Time-Frequency Decoupled Cross-Scale Partial Optimal Transport Alignment Network) for time series domain adaptation. CSPAN first constructs segment-scale time-frequency candidates, and then uses Segment-Scale Aware Top-k Fusion to select the most suitable segment-scale feature representation. It then performs cross-domain alignment after decoupling time and frequency through Temporal Partial Optimal Transport and Frequency Partial Optimal Transport, where temporal domain alignment is guided by the transferability estimated from a probing transport plan, and frequency domain alignment is guided by the discriminability estimated by an auxiliary frequency classifier. By combining segment-scale representation selection, time-frequency decoupling, and partial distribution alignment, CSPAN enables more flexible cross-domain matching. Extensive experiments on six datasets demonstrate CSPAN’s consistent superiority, achieving an average accuracy improvement of 3.16% in cross-domain scenarios. Code is available at the anonymous link: \urlhttps://anonymous.4open.science/r/CSPAN-35FB/.
PaperID: 6816, Poster
Abstract: Autoregressive generation in Large Language Models (LLMs) is constrained by the memory and computational demands of attention mechanisms. Sparse attention methods mitigate this cost by selecting only high-probability entries of the attention matrix. We observe that in many such methods, this renders the probability-value multiplication negligible, shifting the bottleneck to the query-key step. Key heads can therefore be reduced to accelerate inference, while value heads preserve capacity at no additional cost. We introduce Sparse Asymmetric Group-Query Attention (SAGA), which decouples key and value head counts to exploit this principle, and pair it with approximate top-N (Atop-N) attention, a simple GPU-agnostic sparse attention method designed to isolate sparsity's effect on decoding-optimized architectures. We formalize the benefits of this asymmetry theoretically and validate it empirically, achieving end-to-end decoding speedups exceeding 2× at long contexts across multiple model scales and benchmarks. Models trained from scratch with SAGA nearly match the quality of comparable GQA variants while delivering substantial efficiency gains. To facilitate adoption, we introduce an efficient fine-tuning method that converts pretrained models to the SAGA architecture at a small to negligible cost to quality, enabling practitioners to benefit from our approach without costly retraining.
PaperID: 6817, Poster
Abstract: Transformer-based large language models (LLMs) encode many high-level concepts as linear directions in the latent activation space. Once identified, these directions support both measurement, the quantification of a concept's presence, and intervention, the steering of the model's behavior. In practice, however, a direction learned from one context often fails when applied to prompts from a new context. This poor transfer can have two distinct sources: spurious correlation between the concept label and dataset-specific features, and genuine heterogeneity in how the concept is encoded across contexts. Methods from out-of-distribution (OOD) literature can address the first, but worsen the second by discarding context-specific structure that may carry concept signal. We introduce Context-Aided Representation Extraction (CARE), which bridges OOD generalization and mechanistic interpretability. Given activations labeled with a concept and an environment, CARE jointly learns a shared direction, optimized to be invariant across environments, and orthogonal environment-specific residuals that capture how the concept varies by context. We evaluate CARE on subject--verb agreement, refusal of harmful prompts, and toxicity. CARE produces directions that measure concepts more reliably under distribution shift and in unseen environments and datasets than existing methods while remaining effective for intervention. Its decomposition also supports a swap-cost diagnostic that identifies when context-specific structure carries concept-relevant signal.
PaperID: 6818, Poster
Abstract: Semi-dense local feature matching commonly aggregates contextual information before coarse-to-fine correspondence estimation. However, intra-image and inter-image aggregation serve different roles: the former should propagate spatial evidence within each image while retaining structured two-dimensional dependencies, whereas the latter should gather information from plausible corresponding regions while limiting noise from unrelated locations. We propose WovenAnchor Matcher (WAM), a semi-dense matching framework based on role-specialized context aggregation. WAM introduces two complementary operators. WovenMamba performs intra-image aggregation through horizontal-then-vertical state-space propagation, allowing vertical updates to operate on horizontally contextualized features and thereby encouraging a structured two-dimensional receptive-field bias. AnchorCrossAttention performs inter-image aggregation by using cross-attention to gather context from plausible corresponding regions across the image pair. To make this retrieval robust, it operates on local representative anchors rather than dense point-wise tokens, reducing sensitivity to noisy affinities that can otherwise lead to incorrect correspondence propagation. On MegaDepth, WAM achieves 66.1 pose AUC@5^\circ with a runtime of 35.4 ms on a single H100 GPU, improving over JamMa, which achieves 64.1 pose AUC@5^\circ at 43.8 ms under the same setup. Ablations show that removing or replacing either specialized component reduces accuracy, providing evidence that intra-image propagation and inter-image retrieval benefit from role-specific aggregation designs.
PaperID: 6819, Poster
Authors: Lewis Mitchell
Abstract: Iterative fine-tuning on synthetic data causes \emphmodel collapse: output diversity narrows as rare patterns are progressively lost. Existing mitigations either require model log-probabilities, an external oracle, or continued access to real human data. Here we develop a new approach grounded in mathematical information theory: the non-parametric Kontoyiannis entropy rate estimator H_K, computed entirely from raw text via match-length statistics, with no model of any kind. We show that this is in fact a superior training-data filter on text-diversity metrics in a fully-synthetic, single-lineage fine-tuning setting. In a six-generation QLoRA collapse experiment on Llama-3.1-8B, logprob-based filtering (the established gold standard) provides no significant text-diversity benefit on any metric (p > 0.23), whereas H_K-filtering yields +42% unique trigrams, +30% vocabulary, and -19% repetition (all p < 0.001). We validate H_K as a cross-domain entropy proxy (\beta = 0.924, R^2 = 0.746) and collapse detector (\rho = +0.454, p < 0.0001) across 4 domains, 3 temperatures, 2 generator--scorer model pairs, and 1,680 generated documents. Our results demonstrate that information theoretic approaches to collapse mitigation are efficient, and suggest new approaches for maintaining multi-agent diversity.
PaperID: 6820, Poster
Abstract: We present DeFlow, a decoupled offline RL framework for extracting high-value actions from a learned multi-step flow prior. Directly optimizing iterative generative policies typically requires backpropagation through ODE solvers, while shortcut policies trade off expressivity, inference cost, and policy-improvement objectives. DeFlow keeps the flow model as a behavior prior and trains a lightweight action-conditioned residual module for value improvement under an adaptive trust-region penalty. This design bypasses solver differentiation and encourages proximity to the learned prior without treating the trust region as exact data-support preservation. Empirically, DeFlow is competitive with recent generative offline RL baselines, with the clearest gains on selected multimodal manipulation and offline-to-online settings.
PaperID: 6821, Poster
Authors: Wang Xi, Yue Wang
Abstract: Data-efficient distillation of large language models depends not only on the student architecture, but also on which teacher examples are selected for training. Existing selection strategies often treat examples as independent or density-weighted samples, which can over-represent redundant high-frequency patterns while under-covering geometrically distinct regions of the teacher representation space. We study a volumetric alternative for distillation data selection. Our method, VoluCore, selects compact coresets by maximizing the regularized log-determinant of the Gram matrix formed by normalized teacher features. For any fixed regularization parameter, this objective is monotone submodular, giving a standard greedy (1-1/e)-approximation guarantee. To make the criterion practical at LLM scale, we implement greedy selection with efficient Cholesky/Gram-Schmidt rank-one updates after feature extraction. Across math, code, and instruction-following distillation settings, VoluCore reaches high-recovery performance with fewer selected examples than random, uncertainty-based, and distance-based selection baselines, while remaining competitive with gradient-based selection at substantially lower end-to-end cost. Additional spectral, layer, and task-coverage analyses indicate that VoluCore improves representation-space coverage without sacrificing common high-density capabilities. These results support log-determinant coverage of teacher representations as a simple, scalable, and theoretically grounded criterion for constructing compact distillation sets.
PaperID: 6822, Poster
Authors: Chang Shouhao, Xuan Liu, Hongye Zhu, Xinning Chen, Shigeng Zhang
Abstract: Designing effective reward functions is a critical challenge in Reinforcement Learning (RL), which traditionally requires costly manual trial-and-error. While Large Language Model (LLM)-based methods have shown promise in reward design, undifferentiated sampling and limited use of historical feedback often lead to insufficient exploration of the reward space and optimization instability. To address these challenges, we propose R^2E, a Role-driven Reward Evolutionary framework for automated reward function design that structures the process through explicit role specialization. R^2E employs multiple LLM roles with complementary objectives: an Explorer that promotes novelty to expand the search space, a Guardian that performs conservative refinements to improve training stability, and an Evolver that recombines high-performing rewards via crossover and mutation to integrate effective structures. A dynamic scheduling strategy coordinates these roles, progressively shifting the search from broad exploration to focused exploitation. To mitigate optimization instability and provide a medium for cross-role collaboration, R^2E incorporates a Global Elite Pool that retains the best-performing rewards to guide subsequent generations, thereby enabling coordinated interaction among roles and ensuring a reliable and consistent refinement process. Extensive experiments on multiple robotic tasks demonstrate the effectiveness of our method.
PaperID: 6823, Poster
Abstract: Pretraining scaling laws treat training data as a one-dimensional quantity: a token count. Agentic reinforcement learning (RL) inherits this scalar abstraction, but modern synthesis pipelines now generate tasks, environments, and trajectories independently and at very different costs, turning data into a structured design space. This raises a question pretraining never had to answer: given a synthesis budget, which axis should be scaled? We address this question end-to-end with three tightly coupled contributions. First, we decompose agentic data into four axes (task, environment, trajectory, reward) and empirically fit per-axis scaling laws on Qwen3 models (4B/8B/14B) across math, code, and web/UI domains; we observe a robust ordering of per-axis scaling exponents that holds across model sizes, RL algorithms (GRPO and PPO), and domains. Second, we use the fitted laws as a common ruler to price synthesis methods, capturing each method's efficiency as a per-axis discount factor and identifying critical synthetic-to-real ratios at which collapse begins; these ratios differ by an order of magnitude across axes. Third, we cast budget allocation as a constrained optimization under our laws and discounts, deriving a Chinchilla-style compute-optimal recipe and validating it at held-out scale. The empirical results demonstrate that following the recipe yields up to 1.7× compute efficiency over the dominant ``scale-the-trajectories'' practice.
Authors:
Shuo Lyu, Jinhao You, Zhuohang Lyu, Tanxuan Li, Zibo Zhao, Jiaxiang Hu, Kai Tang, Yichen GuoAbstract: Visual Geometry Grounded Transformer (VGGT) recovers dense 3D scene structure from multi-view images in one forward pass, but quadratic cross-frame attention limits its scalability. Existing training-free accelerators reduce computation uniformly along one axis, missing layer heterogeneity. Our spectral, probing, and causal analyses reveal three regimes: shallow layers lack cross-view structure, middle layers drive cross-view alignment, and deep layers are redundant for dense geometry yet their cross-frame attention remains essential for pose. RegimeVGGT applies layer-wise U-shaped compression along two axes: Saliency-Guided Banded Merging protects geometry- and edge-salient tokens, while Selectively Protected K/V Downsampling preserves cross-frame spatial coverage and the pose-critical path through a phase-shifted spatial grid, a reference-frame anchor, and uncompressed camera/register tokens. Training-free, RegimeVGGT achieves a 6.7× speedup over VGGT at matched reconstruction quality. Code: \urlhttps://anonymous.4open.science/status/RegimeVGGT-9477
PaperID: 6825, Poster
Abstract: Noisy gradient estimates are ubiquitous in machine learning, arising from stochastic sampling, distributed computations, or privacy- preserving mechanisms like Differential Privacy (DP). This paper introduces a general, optimizer-agnostic framework for denoising these estimates. Our approach operates as a modular wrapper that intercepts noisy gradient observations and provides denoised estimates to the optimizer, requiring no internal modifications to algorithms like SGD or Adam. We start by studying the problem of optimal denoising gradients in quadratic and cubic optimization problems. We develop maximum likelihood estimation of gradient and Hessian of the objective. We also develop other computationally-efficient order optimal algorithms for denoising gradients in such a setting. Finally, we utilize the developments for the cubic and quadratic setting to develop a more general denoising mechanisms for general smooth nonconvex optimization problems. Theoretically, by leveraging higher-order smoothness, we establish an improved convergence rate of O(T^-12/19) for smooth non-convex optimization. While our recursive algorithm requires two gradient queries per iteration, we show that the improved convergence rate yields a lower total oracle complexity than the standard O(T^-1/2) rate of SGD. This implies the gain in convergence speed asymptotically outweighs the additional computational cost. We apply our algorithm to various denoising problems, particularly nonconvex DP-training. Our experiments, including training private classifiers on CIFAR-10, demonstrate significant improvements over baselines.
Authors: Dylan Forde
Abstract: We study long-context balanced entropic optimal transport (OT) attention on TPU hardware through a stopped-base, fixed-depth tail-refinement surrogate. After a stopped T-step Sinkhorn solve, we unroll a short refinement tail and differentiate that surrogate exactly. For the production R=2 case, the backward pass contains four staircase plan factors. We prove an exact one-reference-tile schedule: the R=2 score cotangent is a single reference plan tile times an explicit modifier field built from vector cotangents and dual differences. This yields block-wise cost O((T+R)LW), O(Ld) input storage, and O(L) additional HBM usage for fixed head dimension d and band width W. We also formalize the current \textttdustbin\_block path as the same balanced surrogate on an augmented support, so the schedule lifts to the gap-aware transport path used in our TPU runs. We provide a local surrogate-bias bound, an a posteriori bias certificate, and a projective contraction certificate for strictly positive active blocks. On synthetic masked problems, the optimized kernel matches exact autodiff of the same centered surrogate to within 10^-5--10^-10. On TPU v6e-8, a four-configuration Pfam screen completes end-to-end, and a promoted balanced R=2 run sustains roughly 8.5 examples per second through a three-hour budget, reaching step 1437. Held-out Pfam test shards improve reconstruction from 3.17 to 0.99 and sparse CE from 5.86 to 5.69 relative to step 0. These results support exact fixed-depth backward theory, a theorem-matching gap-aware bridge, and trainability evidence for the production path.
Abstract: Log-likelihood is a standard metric for evaluating generative models. Unfortunately, in contrast to autoregressive models (ARMs), discrete diffusion models generally do not admit exact computation of this quantity. Existing evaluations, therefore, rely on the evidence lower bound (ELBO), leaving unclear how much higher the true value may be. We address this by introducing the ), a variational upper bound on log-likelihood that admits an unbiased Monte Carlo estimator. Our TUBE extends across latent-variable models, including masked diffusion models (MDMs), any-order ARMs (AO-ARMs), and block variants of both. Applied to block MDMs and block AO-ARMs, TUBE reveals our key empirical finding that these models lie strictly below the exact ARM baseline, showing that ARMs still dominate in likelihood.
PaperID: 6828, Poster
Abstract: Reliable progress in offline policy learning depends on a small set of methodological practices, including careful reporting, well-tuned baselines, and evaluation across diverse conditions. Prior work has noted that results can be sensitive to reporting choices, hyperparameter tuning, and dataset properties independently, but these sources of variability have not been systematically investigated at the scale needed to understand how they shape conclusions. To address this gap, we present a large-scale empirical study of offline reinforcement and imitation learning, training over 145,000 policies across 114 datasets. At this scale, no algorithm dominates: aggregate performance across top methods is often close, but the leaders differ substantially across environments. We find that proper hyperparameter tuning frequently reshuffles perceived algorithm rankings; and simple baselines, including behavior cloning, are often stronger than is commonly assumed after extensive tuning. We also study hyperparameter transfer and sensitivity across environments, identifying a simple strategy that generalizes well. From these analyses we distill practical recommendations, and release JumpStart: a resource suite of trained models, per-model hyperparameter and reward data, strong baselines across all environments, and a website to make retrieval and analysis trivial. Together, these resources aim to make offline policy learning research more reliable and to open new directions for work beyond the scope of this study.
Abstract: We analyze the query complexity of finding a local minimum in t rounds on general graphs. More precisely, given a graph G = (V,E) and oracle access to an unknown function f : V \to \mathbbR, the goal is to find a local minimum---a vertex v such that f(v) \leq f(u) for all (u,v) \in E---using at most t rounds of interaction with the oracle. The query complexity is well understood on grids, but much less is known beyond. This abstract problem captures many optimization tasks, such as finding a local minimum of a loss function during neural network training. For each graph with n \geq 2 vertices and constant integer t \geq 2, we prove a deterministic upper bound of O(n^1/t (s\Delta)^1-1/t), where s is the separation number and \Delta is the maximum degree of the graph. We complement this result with a randomized lower bound of \Omega(n^1/t) that holds for any connected graph. We also find that parallel steepest descent with a warm start provides improved bounds for graphs with high separation number and bounded degree. To obtain our results, we utilized an advanced version of Gemini at various stages of our research. We discuss our experience in the Methodology section.
Abstract: Foundation models for zero-shot time series forecasting face challenges in efficient long-horizon prediction and reproducibility, with existing synthetic-only approaches underperforming on challenging benchmarks. This paper presents TempoPFN, a univariate time series foundation model based on linear Recurrent Neural Networks (RNNs) pre-trained exclusively on synthetic data. The model uses a GatedDeltaProduct architecture with state-weaving for fully parallelizable training across sequence lengths, eliminating the need for windowing or summarization techniques while maintaining robust temporal state-tracking. Our comprehensive synthetic data pipeline unifies diverse generators, including stochastic differential equations, Gaussian processes, and audio synthesis, with novel augmentations. In zero-shot evaluations on the Gift-Eval, fev-bench and Chronos-ZS benchmarks, TempoPFN achieves top-tier competitive performance, surpassing the majority of models trained on real-world data, while being more efficient than existing baselines by leveraging fully parallelizable training and inference. We open-source our complete data generation pipeline and training code, providing a reproducible foundation for future research.
Abstract: Out-of-distribution (OOD) detection is a critical component for ensuring the reliability of deep neural networks in safety-critical applications. In this work, we present a key empirical observation: for in-distribution (ID) samples, class-wise Mahalanobis distances exhibit a pronounced sharp minimum structure, where the distance to the nearest class is small while distances to all other classes remain large, resulting in high variance across classes. In contrast, OOD samples tend to exhibit a less pronounced sharp minimum structure, producing comparatively lower variance across classes. We further provide a theoretical analysis grounding this observation in Neural Collapse geometry: under relaxed Neural Collapse assumptions on within-class compactness and inter-class separation, ID samples are shown to structurally exhibit high class-wise distance variance, offering a theoretical basis for its use as an OOD score. Motivated by this observation and its theoretical backing, we propose MahaVar, a simple and effective post-hoc OOD detector that augments the Mahalanobis distance with a class-wise distance variance term. Following the OpenOOD v1.5 benchmark protocol, MahaVar achieves state-of-the-art performance on CIFAR-100 and ImageNet, with consistent improvements in both AUROC and FPR@95 over existing Mahalanobis-based methods across all benchmarks.
Abstract: Pretraining large language models (LLMs) with next-token prediction has led to remarkable advances, yet the context-dependent nature of token embeddings in such models results in high intra-class variance and inter-class similarity, thus hindering the efficiency of representation learning. While similarity-based regularization has demonstrated benefit in supervised fine-tuning and classification tasks, its application and efficacy in large-scale LLM pretraining remains underexplored. In this work, we propose the SimReg, an embedding similarity regularization loss that explicitly encourages token representations with the same ground-truth label within each sequence to be more similar, while enforcing separation from different-label tokens via a contrastive loss. Our analysis reveals that this mechanism introduces gains by enlarging multi-classification margins, thereby enabling more efficient classification. Extensive experiments across dense and Mixture-of-Experts (MoE) architectures demonstrate that SimReg consistently accelerates training convergence by over 30% and improves average zero-shot downstream performance by over 1% across standard benchmarks. Further ablation studies and analyses offer practical insights into hyperparameter tuning and loss effectiveness.
PaperID: 6833, Poster
Abstract: We study the query complexity of approximating the top eigenvector of an asymmetric matrix in the matrix-vector product model. Our main result gives a gap-dependent lower bound for this problem in the adaptive query setting. Prior lower bounds for symmetric matrices already imply weaker hardness guarantees for the more general asymmetric problem, whereas our result yields a sharper dependence on the eigen-gap in the asymmetric case. In the inverse polynomial accuracy regime, this lower bound matches the benchmark upper bound of the power method up to lower-order factors. Our proof is based on an asymmetric spiked random matrix that serves as a hard instance. The key random-matrix ingredient is an analysis of the spike model, including both its asymptotic eigen-gap and the alignment between its top eigenvector and the planted spike. Building on an information-theoretic framework for query lower bounds, we then obtain the stated hardness result.
PaperID: 6834, Poster
Authors:
Shaofei Liu, Weichen Bi, Yun MaAbstract: Spatial intelligence is a fundamental capability of embodied Artificial Intelligence systems, and its reliable assessment and optimization require well-designed benchmark. Existing benchmarks for spatial intelligence largely rely on video datasets or procedurally generated data, which limits data scalability and ecological validity. To address these limitations, we propose to automatically extract spatial information directly from executable environments. Specifically, we design WebSpatial, a benchmarking framework built upon web runtime environments, which enables automatic acquisition of spatial data by parsing underlying code and interaction processes. It performs runtime scene graph parsing on web 3D applications to automatically extract visual attributes, world coordinates, and user interaction event sequences of objects, while constructing a dedicated spatial computation tool library. The modular generation pipeline integrates LLM-based task parsing and planning and tool-assisted execution to automatically generate question-answer pairs. We conduct experiments in method effectiveness validation and benchmark evaluation. For validation, we demonstrate that WebSpatial can reliably construct diverse and scalable spatial tasks grounded in both static and dynamic scenes, covering spatial perception including color, shape and quantity, and spatial reasoning including relation assessment, direction transformation and mental rotation. We develop a small-scale benchmark named WebSpatialBench based on WebSpatial, and use it to evaluate 10 multimodal large language models, revealing their strengths and limitations in spatial problem-solving.
PaperID: 6835, Poster
Abstract: , where minority classes become indistinguishable, is a significant challenge in imbalanced learning. This challenge is addressed by methods such as Mixup with class-balanced sampling. Although minority collapse has been mathematically analyzed using the layer-peeled model alongside Neural Collapse, no prior work has analyzed minority collapse under Mixup, particularly from the perspective of mixed labels. We investigate this overlooked factor and raise the question: Our analysis reveals that (i) mixed labels should be balanced, and (ii) in this setting, interpreting mixed labels as singletons is beneficial. Motivated by this analysis, we propose a , which balance mixed labels and treat them as singleton labels. Through theoretical analysis and experimental results, we highlight the importance of balancing mixed labels in imbalanced learning.
PaperID: 6836, Poster
Authors: Yufei Jin, Jiajun Shen, Yi He, Xingquan Zhu
Abstract: Learning from Label Proportion (LLP) is a weakly supervised learning paradigm in which only aggregated label proportions over collections of instances (i.e., bags) are provided, rather than individual labels. This allows classification while preserving privacy or reducing annotation costs. Existing LLP methods, however, have been largely restricted to i.i.d. tabular or image data. To the best of our knowledge, no solution currently addresses graphs, where instances are inherently interdependent through network structure. In this paper, we generalize LLP to the graph domain and study the problem of node classification with label proportions, where only distributional supervision is available for node bags, and the goal is to infer labels for all nodes in the graph. We argue that the lack of node-level supervision is the main challenge for LLP on graphs, and that existing methods based on i.i.d. assumptions fail to exploit topological correlations. To overcome this, we propose GLLP~(Graph Learning from Label Proportions), a framework that leverages Optimal Transport (OT) with a homophily-aware cost to generate soft pseudo-labels for individual nodes. These pseudo-labels provide stronger supervision signals for training Graph Neural Networks. We further establish theoretical guarantees showing the alignment of our cost function with the node classification objective. Extensive experiments on six homophilic graph benchmarks and the Covid real-world dataset demonstrate that GLLP consistently outperforms existing LLP baselines and variants.
PaperID: 6837, Poster
Abstract: Nonlinear independent component analysis (nICA) is generally unidentifiable due in part to measure-preserving automorphisms (MPAs), which induce indistinguishable latent representations. We show that for non-Gaussian sources with mild regularity assumptions, such MPAs are not Lipschitz, motivating Lipschitz continuity as a structural assumption. Under this condition, we prove global identifiability of nICA within the class of Orthogonal Coordinate Transformations (OCTs) with bounded sources, recovering the true components up to permutation and scaling. Our analysis reduces the learning problem to a linear program over the Birkhoff polytope, and extends to functions with symmetric Jacobians, revealing connections to optimal transport. We introduce OCTNet and SYMNet, normalizing flow architectures that enforce these constraints in practice. Experiments demonstrate strong recovery performance and robustness to model misspecification, outperforming existing approaches.
Abstract: We study Bayesian optimization (BO) through the lens of information geometry. Pulling back the Fisher information metric through the surrogate posterior map yields a local sensitivity tensor on the input space, which leads to an upper bound on the gradient of reparameterizable acquisition functions. This view explains vanishing-gradient behavior in high-dimensional BO and provides a common interpretation of heuristics such as RAASP and dimension-scaled lengthscales. Building on this analysis, we propose FITR, a trust-region-based BO method that replaces lengthscale-based scaling by local pullback-Fisher weights. FITR is not restricted to GP kernels with explicit lengthscales. On GP benchmarks with an SE kernel, experiments show competitive performance using FITR. The proposed method also easily generalizes to non-isotropic surrogates, although the gains are more task-dependent in that setting.
PaperID: 6839, Poster
Abstract: We examine generative modelling approaches based on the construction of Schrödinger bridges between Gaussian noise and a target distribution. It is known that the solution of the dynamic Schrödinger problem is a diffusion process with a drift associated with Doob's h-transform of a Schrödinger potential. Although its accurate restoration from finite samples is crucial for reliable, high-quality data generation, the existing literature lacks theoretical guarantees regarding this question. In our work, we establish theoretical upper bounds on the complexity of Schrödinger potential approximation and estimation via neural networks. These bounds are determined by the effective dimension of the target distribution. To our knowledge, this is the first result demonstrating that generative modelling methods based on Schrödinger bridges and stochastic optimal control can escape the curse of dimensionality.
Abstract: A central error measure in Gaussian DDPMs is the path-space KL divergence between the exact reverse chain and the learned Gaussian reverse process. This quantity is especially relevant for procedures such as classifier guidance, which perturb the entire reverse trajectory rather than only the terminal sample. Prior analyses show that standard isotropic reverse covariances suffer an unavoidable \Omega(1/T) path-KL error as the number of denoising steps T grows. We show that matching the full posterior covariance breaks this barrier, yielding an order-wise improvement that reduces the path KL to O(1/T^2). To make full covariance matching practical, we introduce the Lanczos Gaussian sampler, a training-free, matrix-free method for sampling from the optimal reverse covariance using only covariance-vector products, which are available through Jacobian-vector products of the posterior mean. The sampler avoids dense covariance storage and auxiliary covariance models. We prove that its approximation error decays exponentially in the number of Lanczos steps, where each Lanczos step requires a single Jacobian-vector product. Empirically, using only just three such steps improves sample quality over strong diagonal-covariance baselines, including OCM-DDPM, across standard image benchmarks. This identifies full covariance matching as both theoretically valuable and practically accessible for fast DDPM sampling.
PaperID: 6841, Poster
Abstract: Personalization histories vary across domains because surface actions and objectives differ: users watch or skip movies, click or ignore news, or request recommendation, summarization, question answering, or response generation. Rather than build a separate history encoder for each action space, we rewrite native interactions into a shared action-on-item event schema: a specific action on a specific item becomes positive, negative, or command-like evidence over a target-locus embedding whose geometry induces item-kind abstraction. Task requests become generalized commands over the active item-side abstraction. The modeling problem is how this evidence enters a user-specific state, persists, and is read out by the active command. We formulate the Multi-Timescale State Hypothesis (MTSH), where signed and command-like evidence flows through long-term stable interests, recency-sensitive short-term interests, and bursty episodic traces. We instantiate MTSH with \textttPerTIDE, an action-conditioned multi-timescale encoder that learns a user-specific preference-history state and uses command-conditioned readout to produce a task-ready state. We evaluate \textttPerTIDE on MovieLens, MIND, and PENS for prediction, and PENS/OpenAI-Reddit for generation. \textttPerTIDE improves over the strongest baselines without explicit trace factorization by +2.36/+3.47, +10.05/+9.01, and +0.90/+1.42 MRR/nDCG on MovieLens, MIND, and PENS; and over the strongest two-shot large language model by +0.10/+0.18 and +0.04/+0.19 PerSEval-JSD/PerSEval-METEOR on PENS and OpenAI-Reddit. Removing action conditioning and command-conditioned readout degrades MIND by 6.52 MRR / 9.13 HR@10 and 4.14 MRR / 4.42 HR@10. Trace diagnostics show L/S/E specialization, while the full L+S+E model is strongest under mixed stable, recency-sensitive, and episodic command evidence. These results support action-conditioned multi-timescale preference flow as a reusable user-history encoding principle.
Authors:
Zhuobin Yang, Liao Yu, Yeyao Bao, Lv L Fu, Daqing Guo, Jian Zhang, Xiaohong Li, Yunliang ZangAbstract: Spiking Neural Networks (SNNs) are promising for energy-efficient, real-time edge computing, yet their performance is often constrained by the limited adaptability of conventional leaky integrate-and-fire (LIF) neurons. Existing LIF models struggle with restricted information capacity and susceptibility to noise, leading to degraded accuracy and compromised robustness. Inspired by the dynamic self-regulation of biological potassium channels, we propose the Potassium-regulated LIF (KvLIF) neuron model. KvLIF introduces an auxiliary conductance state that integrates membrane potential and spiking history to adaptively modulate neuronal excitability and reset dynamics. This design extends the dynamic response range of neurons to varying input intensities and effectively suppresses noise-induced spikes. We extensively evaluate KvLIF on both static image and neuromorphic datasets, demonstrating consistent improvements in classification accuracy and superior robustness compared to existing LIF models. Our work bridges biological plausibility with computational efficiency, offering a neuron model that enhances SNN performance while maintaining suitability for low-power neuromorphic deployment.
Abstract: While reinforcement learning (RL) successfully enhances reasoning in large language models, its role in fostering compositional generalization (the ability to synthesize novel skills from known components) is often conflated with mere length generalization. To this end, we study what RL post-training teaches models about skill composition and how the composition structure affects the skill trasnfer. We focus on the \countdown task (given n numbers and a target, form an expression that evaluates to the target) and analyze model solutions using expression trees, where each subtree corresponds to a reusable subtask and thus can be viewed as a ``skill.'' Tracking tree shapes and their success rates over training, we find: (i) out-of-distribution (OOD) generalization to larger n and to unseen tree shapes, indicating compositional reuse of subtasks; (ii) the order of skill acquisition depends upon the composition structure---models master shallow balanced trees (workload is balanced between subtasks) before deep unbalanced ones, with persistent fragility on right-heavy structures for a fixed composition depth. Our diagnostic reveals what is learned, in what order, and where generalization fails, clarifying how RL-only post-training induces OOD generalization beyond what standard metrics such as pass@k reveal. To study the implication of our findings for more real-world tasks, we present a preliminary study on formal theorem proving.
PaperID: 6844, Poster
Abstract: Online platforms often struggle with inflated, uninformative user ratings. This issue can be mitigated by better design of the data collection mechanism, for example through changing the phrasing of user questions. Previous work optimizes a single rating scale to best distinguish the inherent qualities of items under binary feedback. In this paper, we demonstrate the power of utilizing different rating scales for different users. First, we show that a designer can always split a single rating scale into two to improve item rankings. Next, we formally characterize the optimal design of multiple scales as a series of 0-1 step functions whose threshold values are evenly spaced. Finally, we consider an adaptive model where subsequent rating scales can depend on the responses from previous users and show that adaptivity leads to even greater distinguishing power. All in all, our theoretical results show that using heterogeneous scales significantly improves item rankings on platforms.
Authors: Pedram Bakhtiarifard, Sophia Natasha Wilson, Mahmoud H. A. Afifi, Jonathan Wenshøj, Raghavendra Selvan
Abstract: Algorithmic complexity measures the intrinsic structure (or randomness) of strings. Kolmogorov-Chaitin-Solomonoff (KCS) complexity is the length of the shortest program that can output a string and halt over all possible programs; due to this reason it is uncomputable. Estimating the algorithmic probability of strings using simulations of finite Turing machines is one way of obtaining tight estimations of KCS complexity of strings. For two dimensional objects, this is currently doable only for binary strings using the block decomposition method (BDM). In this work, we present the quantized block decomposition (QuBD) method to extend the estimation of algorithmic complexity to any k-ary objects such as the weights of deep neural networks. We show theoretically that the proposed QuBD method yields better KCS complexity estimations than BDM that relies on binarization. We further study how algorithmic complexity interacts with learning in deep neural networks by tracking the evolution of weights during training. Using a variety of experiments we show that algorithmic complexity decreases as models learn, it correlates with generalization performance and can be used as a diagnostic measure to perform model compression.
Abstract: Audio Language Models (ALMs) have recently shown strong capabilities in unified reasoning over speech, sound, and natural language; yet we find that they can inherit sycophancy, the tendency to agree with user assertions even when they contradict objective evidence. This failure mode is especially concerning for audio-conditioned reasoning, where a model must preserve evidence from acoustic events, speaker characteristics, and speech rate while responding to potentially misleading user feedback. However, unlike text and vision-language sycophancy, ALM sycophancy has not been systematically studied. We therefore introduce SYAUDIO, the first benchmark dedicated to evaluating sycophancy in ALMs, consisting of 4,319 audio questions spanning Audio Perception, Audio Reasoning, Audio Math, and Audio Ethics. Built upon established audio benchmarks and augmented with TTS-generated arithmetic and moral reasoning tasks, SYAUDIO enables systematic evaluation across multiple domains and sycophancy types with carefully verified data quality, including a human-speaker validation of the TTS pipeline. Using this benchmark, we identify substantial and audio-specific sycophancy patterns under realistic conditions involving noise and speech rate, and further show that supervised fine-tuning reduces misleading susceptibility while decode-time steering reveals controllable hidden-state directions for both misleading susceptibility and correction receptiveness.
Abstract: Despite remarkable advances, today's AI systems remain narrow in scope, falling short of the flexible, adaptive, and multisensory intelligence that characterizes human capabilities. This gap has fueled longstanding debates about whether AI might one day achieve human-like generality or even consciousness, and whether theories of consciousness can inspire new architectures for AI. This paper presents an early blueprint for implementing a general AI system, CTM-AI, combining the Conscious Turing Machine (CTM), a formal machine model of consciousness, with today's foundation models. CTM-AI contains an enormous number of powerful processors ranging from specialized experts (e.g., vision-language models and APIs) to unspecialized general-purpose learners poised to develop their own expertise. Crucially, for whatever problem must be dealt with, information from many processors is selected, integrated, and exchanged appropriately to solve the task.CTM-AI achieves state-of-the-art accuracy on MUStARD (72.28) and UR-FUNNY (72.13), outperforming multimodal and multi-agent frameworks. On tool-using and agentic tasks, CTM-AI achieves 10+ points of improvement on StableToolBench and WebArena-Lite. Overall, CTM-AI offers a principled, testable blueprint for general AI inspired by a model of consciousness.
PaperID: 6848, Poster
Abstract: Large language models successfully solve classic theory-of-mind tasks adapted from cognitive science, but the mechanisms underlying this ability remain unclear. Do they deploy circuitry specialized for reasoning about other minds, memorize common false-belief vignette structures and outputs, or rely on more domain-general computations? We investigate this question in Qwen2.5-14B-Instruct using causal mediation and representational similarity analyses across matched prompt sets that systematically vary an agent’s beliefs, the state of the physical world, and the correct answer. We identify a population of mid-layer attention heads that tracks divergence between an initial representation and the current state of the world. These heads are causally relevant regardless of whether the initial representation is an agent’s belief or a photograph, despite the photograph condition involving no agent or perspective-taking. A distinct population of later-layer heads retrieves the answer token. The two populations are functionally dissociable yet combine compositionally. Together, these findings argue against both mentalizing-specific and memorization accounts, suggesting instead that LLMs solve classic false-belief problems by reusing a domain-general mechanism for detecting divergence between representations and reality.
PaperID: 6849, Poster
Authors: Kojiro Hirokane, Ryohei Ueno, Takashi Kitsukawa
Abstract: Backpropagation (BP) requires feedback (FB) weights to equal the transpose of feedforward (FF) weights at every layer and at every step. This issue is called the weight transport problem, and has been considered to be biologically implausible. Neural circuits are not built from anatomically paired FF/FB axons but from asymmetric projections that form recurrent loops shared across many computations. In multi-nuclei circuits (e.g., thalamo-cortical, cortico-cerebellar, basal-ganglia loops), the return path traverses entirely distinct nuclear structures, making any mechanism for maintaining synchronized weight transposes anatomically absent. Predictive coding (PC), originally proposed as a biologically plausible substitute for BP, silently inherits this constraint despite its biological motivation. Therefore, we propose two PC variants that eliminate it entirely. PC-DH (Dual Hebbian plasticity model) updates FF and FB weights by independent local Hebbian rules with no inter-pathway coordination. PC-RFB (Random FeedBack) fixes the FB weight with a random matrix, testing the robustness of the learning mechanism. Both match baseline PC accuracy (~85% on MNIST). We identify a phenomenon we termed Loop Alignment: the spontaneous convergence of FF and FB weights toward a transpose relationship, driven by the structure of the inference loop. We then prove that it holds exactly at every training step in both models. To test whether Loop Alignment holds in the multi-nuclei setting, we conduct a circuit-sharing experiment: two networks learn different tasks (Net 1: MNIST; Net 2: Fashion-MNIST) over a shared pathway, where each network's FB is routed through the other's FF weights. Even when the entire shared pathway is plastic and driven simultaneously by both tasks, each network solves its task independently and FF/FB alignment is stronger than in the isolated case. Loop Alignment explains how the brain can learn effectively without weight transport, and support multi-task learning over shared asymmetric circuits.
PaperID: 6850, Poster
Authors: Karan Singh
Abstract: The framework of Online Convex Optimization with Memory captures sequential decision-making settings where the learner's instantaneous loss is affected by their past choices, in addition to the most recent decision. We first give a new first-order regret bound in this framework for potentially non-smooth loss functions, with regret scaling as the square root of the loss of the best decision in hindsight, which can be much better than known regret bounds that explicitly scale with the horizon. In fact, our result holds in a generalized non-differentiable setting, which also captures other decision making problems of interest. We then give a faster gradient-based algorithm that yields an improved first-order regret bound for smooth loss functions. As an application, we consider the recently introduced problem of online nonstochastic control, which involves controlling a linear dynamical system subject to adversarial disturbances to minimize convex cost functions, where our results can be adapted to produce novel first-order regret bounds. This means whenever there is a benchmark controller with small loss in hindsight, our regret bounds are significantly smaller than those present in the literature.
Abstract: Random forests are widely used in fields involving sensitive tabular data, but existing approaches to enforcing differential privacy (DP) typically degrade performance to the point of impracticality. In this paper, we introduce Lumberjack, a differentially private random forest algorithm that achieves substantially higher utility by constructing large random decision trees and then applying aggressive, privacy-preserving pruning to retain only sufficiently populated nodes. A key component of our approach is a novel (\varepsilon,\delta)-DP heavy hitter detection algorithm for hierarchical data, whose error is O_\varepsilon,\delta(\sqrt\log h) for trees of height h and may be of independent interest. This favorable scaling enables the use of significantly deeper trees than in prior work, leading to improved expressiveness under privacy constraints. Our empirical evaluation on benchmark datasets shows that Lumberjack consistently outperforms prior DP random forest methods, establishing a new state of the art. In particular, our approach yields substantial improvements in the privacy-utility trade-off for practical privacy budgets. Our findings suggest that more carefully designed DP random forests can close much of the utility gap, highlighting a promising and underexplored direction for future research.
PaperID: 6852, Poster
Authors: Yehong Jiang, Yen-Kuang Chen, Xinmin Tian
Abstract: Block-wise post-training quantization (PTQ) compresses LLMs by minimizing per-block reconstruction error on a small calibration set. We show that the \emphcomposition of this set (200 samples) is a targeted lever at extreme bit widths, acting only on iterative methods. In a controlled study spanning six data-dependent PTQ methods plus a data-free baseline, six models (1.7B--8B) across three families plus 70B verification, three bit widths, and a 10-task evaluation with four-seed replication, we find a sharp near-cliff method-class divide. In each method's pre-cliff regime (where the model retains predictive function), iterative methods (SignRound~V2, OmniQuant, AQLM) gain 0.74--1.93pp on the 10-task average from switching Pile to a fragility-guided dataset (FragCal-DC), while single-pass analytical methods (GPTQ, AWQ) show near-zero response (|\Delta|\le0.30pp). Both classes are evaluated on functional Pile baselines (0.42--0.61, well above chance); above the cliff, every method becomes data-composition-insensitive. A reconstruction-loss paradox (the method that most reduces its proxy loss gains nothing downstream, while the method that barely moves the proxy gains the most) sharpens the mechanism; fixed-reference held-out evaluation confirms the proxy--downstream disconnect. A wrong-answer contamination ablation retains 87--91% of the improvement and a data-dependent single-pass control (RTN-opt) shows zero response, isolating the iterative optimization loop. A weight-space probe shows SignRound-V2 preserves higher per-group scale heterogeneity than GPTQ on 9 of 9 (model, width) cells; a scale-flatten knockout shows this heterogeneity is necessary for iterative fidelity. We release FragCal-DC, the discriminative-task instance used here, as a drop-in JSON for the evaluated setting.
Abstract: We study stochastic convex optimization (SCO) with heavy-tailed gradients under pure \varepsilon-differential privacy (DP). Instead of assuming a bound on the worst-case Lipschitz parameter of the loss, we assume only a bounded k-th moment. This assumption allows for unbounded, heavy-tailed stochastic gradient distributions, and can yield sharper excess risk bounds. Prior work characterized the minimax optimal rate for \rho-zero-concentrated DP SCO up to logarithmic factors in this setting, but the pure \varepsilon-DP case has remained open. We characterize the minimax optimal excess-risk rate for pure \varepsilon-DP heavy-tailed SCO up to logarithmic factors. Our algorithm achieves this rate in polynomial time with high probability. Moreover, it runs in deterministic polynomial time when the worst-case Lipschitz parameter is polynomially bounded. For important structured problem classes --- including hinge/ReLU-type and absolute-value losses on Euclidean balls, ellipsoids, and polytopes --- we achieve deterministic polynomial time even when the worst-case Lipschitz parameter is infinite. Our approach is based on a novel framework for privately optimizing Lipschitz extensions of the empirical loss. We complement our upper bound with a nearly matching high-probability lower bound.
Authors: Mohammad Tabish, Benedict Leimkuhler, Stefan Klus
Abstract: We introduce RaNNDy, a randomized neural network framework for learning transfer operators and their spectral decompositions. By fixing the weights of the hidden layers and training only the output layer, RaNNDy achieves significant reductions in computational costs---often by two orders of magnitude---with similar accuracy compared to state-of-the-art deep learning methods. The method is less sensitive to hyperparameter tuning and provides a closed-form solution for the output layer weights that directly represent eigenfunctions of the learned operator. Moreover, it is possible to estimate uncertainties associated with the computed spectral properties via ensemble learning. We demonstrate the efficacy of the proposed approach using diverse benchmark problems, including stochastic dynamical systems and protein folding processes, highlighting the strengths but also weaknesses of RaNNDy.
Authors:
Sagi Ahrac, Noya Hochwald, Mor GevaAbstract: Sparse Mixture-of-Experts (SMoE) models enable scaling language models efficiently, but training them remains challenging, as routing can collapse onto few experts and auxiliary load-balancing losses can reduce specialization. Motivated by these hurdles, we study how routing decisions in SMoEs are formed mechanistically. First, we reveal a geometric coupling between routers and their corresponding experts. For a given token, the router weights for the selected expert and the expert weights processing it receive gradients along the same input direction, differing only in scalar coefficients. Thus, matched router--expert directions accumulate the same routed token history. This theoretical coupling also appears empirically in routing dynamics. In a 1B SMoE trained from scratch, higher router scores predict stronger expert neuron activations, showing that routing decisions are mirrored inside the selected expert. Next, we analyze the effects of auxiliary load balancing on the router--expert geometric coupling, showing that such losses break this structure by spreading input-directed gradients across router weights, making distinct router directions nearly three times more similar to each other. Last, we demonstrate the centrality of geometric coupling for effective routing with a parameter-free online K-Means router, in which each expert maintains a running average of the hidden states routed to it and tokens are assigned based on cosine similarity. Compared with auxiliary-loss and loss-free balancing, this router achieves the lowest load imbalance with only a modest perplexity increase, indicating that geometric coupling captures a substantial part of what the router learns. Overall, our results explain how routers form assignment geometry that supports an effective division of labor.
Abstract: Parallel thinking improves LLM reasoning through multi-path sampling and aggregation. In standard evaluations, due to a lack of sample-specific priors, all samples share a global budget chosen to maximize dataset accuracy. However, many samples reach their best accuracy with much smaller budgets, causing low budget utilization. This . In this paper, we first provide a formal analysis of the overscaling curse and quantify its prevalence and severity in real-world systems. To break it, we propose , which probes model latent representations to predict sample-specific optimal budgets. LanBo significantly improves budget utilization while maintaining dataset accuracy. We further integrate LanBo into the full decoding pipeline, inspiring , a paradigm that allocates budgets before decoding to preserve decoding-time parallelization. LanBo substantially improves hardware-aware efficiency in latency and memory, demonstrating both its practical value and the promise of PreAda for efficient parallel decoding.
PaperID: 6857, Poster
Abstract: Multi-modal cloud removal exploits complementary optical and synthetic aperture radar (SAR) images to restore cloud-degraded remote sensing images. Most existing methods rely on black-box, data-driven architectures with limited interpretability of cross-modal interactions. Although some attempts introduce physical priors into networks, their oversimplified formulations fail to faithfully characterize the underlying physical imaging process and the networks remain loosely coupled with optimization variables. To address these issues, we propose a physics-informed deep unfolding network for cloud removal via optical-SAR image fusion, termed as PIU-CR. We first formulate cloud removal as a physics-informed optimization problem. Specifically, we fundamentally ground the optimization in the physical imaging mechanism of cloud degradation by embedding the atmospheric scattering model into data fidelity term. Then, the regularization terms establish an explicit cross-modal interaction mechanism, overcoming the single-modal information scarcity and the obscure, uninterpretable interactions in black-box models. They are dedicated to SAR-guided structure and spectral regularization, jointly introducing cross-modal structural feature alignment and coordinated spectral consistency constraints. The resulting optimization problem is then unfolded into a multi-stage neural network with dedicated structure and spectral branches. Each stage corresponds to explict optimization iteration step, enabling interpretable and effective reconstruction process. Experiments demonstrate that the proposed method outperforms state-of-the-art methods.
PaperID: 6858, Poster
Authors:
Ibrahim Omer Syed, Ginny Wong, Xiangyu ZhaoAbstract: State Space Models (SSMs) offer an efficient alternative to Transformers for sequence modeling, yet conditioning pre-trained SSMs for iterative generation typically operates outside the recurrent operator, through input injection or activation modulation. While such mechanisms expose the model to conditioning information, they leave the underlying temporal dynamics fixed. We introduce MaRK (Markov-adapted Recurrent Kernels), a dynamic operator-conditioning framework that maps context vectors directly into bounded modulations of a frozen SSM's recurrence (A), read-in (B), read-out (C), skip (D), and discretization (\Delta) parameters. Viewed through the lens of Linear Parameter-Varying systems, MaRK induces a context-indexed family of Markov parameter sequences, allowing each diffusion timestep to reshape the model's input-output memory kernel. We instantiate MaRK on a frozen 111M-parameter Hydra SSM backbone and study three adapter geometries: Hypernet, Chebyshev polynomial, and Discrete Cosine Transform kernels. Since these adapters modify the Markov parameter sequence through low-rank auxiliary maps on the frozen backbone, parameter-efficient fine-tuning arises as a structural consequence of the adaptation mechanism itself, requiring only 7.5--13M trainable auxiliary parameters to transition from a bidirectional objective to an iterative diffusion regime. The bounded recurrence parameterization further yields an analytic Affine Quadratic Stability certificate for the modulated recurrence. Through synthetic LPV recovery experiments and Markov-operator diagnostics, we show that MaRK recovers coordinate-invariant temporal operators under matched assumptions and produces distinct, stable timestep-conditioned memory profiles. Empirically, the Chebyshev variant yields the strongest performance, achieving an average validation loss of 2.63, followed by the DCT (2.66) and Hypernet (3.60) geometries. Together, these results provide initial evidence that dynamic operator modulation is a principled operator-level conditioning mechanism for adapting SSMs beyond input-stream injection and adaptive normalization.
PaperID: 6859, Poster
Authors: Haoran Liu, Guanyi Wang, Yu Yang
Abstract: Rounded capacity cuts (RCCs) are among the most effective cuts for capacitated vehicle-routing relaxations, but exact separation is computationally prohibitive at scale. Existing heuristic and neural separators can be fast, but often fail to produce sufficiently many effective cuts within practical separation budgets. We introduce SAG-Sep, a separator that scores node membership in RCC-inducing subsets across vehicle-count levels from a fractional LP solution in a single encoder forward pass, avoiding the iterative predict-and-coarsen inference used by NeuralSEP-style separators. SAG-Sep encodes the LP solution as a Sparse Augmented Graph (SAG), adding flow-informed multi-hop edges that expose long-range routing structure without densifying attention. Beyond LP-aware sparse encoding, we observe that high-value exact RCCs are typically nested, organizing into onion-like subset families, and SAG-Sep's node scores track subset nesting depth. This motivates Onion-MBP (Margin-Band Probing), a training-free subset-level search that explores nodes near the score margin while generating nested proposals, converting independent node scores into structurally coordinated RCC candidates. On the NeuralSEP benchmark with N=1000 customer instances, SAG-Sep with Onion-MBP reduces the average root gap by 17.7% relative to the strongest NeuralSEP baseline, while the single-pass SAG-Sep separator achieves a roughly 30-fold per-iteration neural separation speedup. Our findings suggest that LP-aware sparse encoding and the nested structure of high-value cuts are structural patterns worth leveraging in neural separation more broadly.
Abstract: A recurring challenge in preference fine-tuning (PFT) is handling intransitive (i.e., cyclic) preferences. Intransitive preferences often stem from either (i) inconsistent rankings along a single objective or (ii) scalarizing multiple objectives into a single metric. Regardless of their source, the downstream implication of intransitive preferences is the same: there is no well-defined optimal policy, breaking a core assumption of the standard PFT pipeline. In response, we propose a novel, game-theoretic solution concept, the Maximum Entropy Blackwell Winner (MaxEntBW), that is well-defined under multi-objective intransitive preferences. To enable computing MaxEntBWs at scale, we derive \textttPROSPER: a provably efficient PFT algorithm. Unlike prior self-play techniques, \textttPROSPER directly handles multiple objectives without requiring scalarization. We then apply \textttPROSPER to the problem of fine-tuning large language models (LLMs) from multi-objective LLM-as-a-Judge feedback (e.g., rubric-based judges), a setting where both sources of intransitivity arise. We find that \textttPROSPER outperforms all baselines considered across both instruction following and general chat benchmarks.
PaperID: 6861, Poster
Authors: Hankyul Kang, Jongbin Ryu
Abstract: SVD-based pruning and quantization have recently emerged as a promising strategy for the ultra-efficient compression of large language models. In these methods, compression is performed in two stages: components are first truncated, and the remaining ones are subsequently quantized. Although this decoupled pipeline benefits from both pruning and quantization, it requires separate optimization for each stage and fails to fully exploit their balance, which can lead to suboptimal performance under aggressive compression. To address this limitation, we propose a new LLM compression method that co-optimizes pruning and quantization in a unified framework. Our key idea is a differentiable method for learning component-wise bit-widths, allowing less important components to be assigned 0-bit precision and pruned away. Notably, our method performs favorably against two-stage baselines, even when subjected to extreme quantization settings (1.61 bits) designed for ultra-efficiency. The code will be publicly released upon acceptance.
Abstract: Multi-agent LLM decision systems for portfolio management still lack a principled way to assign credit across specialist agents, remain vulnerable to cold-start dominance under regime shifts, and offer limited transparency into how final allocations are formed. We propose Market Regime Council (MRC), a cooperative multi-agent decision system that computes exact Shapley credits across all single, pairwise, and Grand-coalition outputs for online agent weighting. Instantiated with N=3 specialist agents, at each trading period, MRC recomputes coalition-based Shapley weights from exponentially weighted performance histories, uses a Bayesian adaptive mixture to stabilize early periods, applies regime-dependent multipliers to adjust agent authority, and records each rebalance through a five-layer causal trace. Over 1,037 trading days across 13 crypto assets and five seeds, MRC achieves a Sharpe ratio of 1.51 and a cumulative return of 440.1%, ranking first on CR, SR, and IR among active baselines and attaining the lowest MDD among active methods. Ablation results show that the gains come from Shapley-weighted integration across coalition outputs rather than from any single stage in isolation. Code and demo data are included in the supplementary material.
PaperID: 6863, Poster
Abstract: Clinical decision-making is intrinsically sequential: a clinician must propose a diagnosis and treatment plan from the pre-treatment state and anticipate the post-treatment outcome in order to revise the plan. Existing medical world models address this loop only partially, through image-to-image dynamics with separate diagnostic and scoring modules, text-only trajectory models without multimodal input, or reflection agents without a world model. We propose MedWM, a unified multimodal LLM in which policy and dynamics are two modes of a single autoregressive backbone with fully shared parameters; the dynamics mode emits a text-only \emphstructured post-treatment state, free-text clinical description plus standardized key-value outcomes (pCR, RECIST, labs, survival), that admits objective evaluation without perceptual metrics. We post-train this backbone with joint policy-dynamics supervision and dynamics-grounded reinforcement learning, then wrap it in an inference-time agent harness that combines breadth-depth (N × K) search using the model's own dynamics as an internal verifier with a self-evolving case memory. On MIMIC-IV, BreastDCEDL-ISPY2 and HCC-TACE-SEG, MedWM consistently improves over same-backbone controls, matches or surpasses a specialized three-module image-based world model, and outperforms reflection-only, multi-agent, and larger closed-source LLM baselines; a blinded clinician study corroborates the quantitative gains.
PaperID: 6864, Poster
Abstract: Vision models are sensitive to object-background context, yet it remains unclear how recognition changes as the semantic distance between objects and backgrounds increases. We introduce ImageNet-OOC1k, a benchmark enabling continuous control of object-background relationships via a consensus semantic dissimilarity axis. This formulation reveals a highly consistent monotonic decline in recognition accuracy across 35 models spanning convolutional neural networks, Vision Transformers, and vision-language models. Internal analyses of Vision Transformers reveal aggregation failure, where object evidence is present at the patch level but not consolidated into the final prediction. Leveraging this characterization, a simple spatial re-ranking strategy recovers a subset of these errors and improves accuracy by up to 5.4 percentage points. These results establish context sensitivity as a continuous and predictable property of modern vision systems, with Vision Transformer analyses showing that failures can arise from limitations in aggregating competing signals rather than from missing object representations.
Abstract: Large Language Models (LLMs) are typically evaluated for safety under single-shot or low-budget adversarial prompting, which underestimates real-world risk. In practice, attackers can exploit large-scale parallel sampling to repeatedly probe a model until a harmful response is produced. While recent work shows that attack success increases with repeated sampling, principled methods for predicting large-scale adversarial risk remain limited. We propose a scaling-aware Best-of-N estimation of risk, SABER, for modeling jailbreak vulnerability under Best-of-N sampling. We model sample-level success probabilities using a Beta distribution, the conjugate prior of the Bernoulli distribution, and derive an analytic scaling law that enables reliable extrapolation of large-N attack success rates from small-budget measurements. Using only n=200 samples, our anchored estimator predicts ASR@1000 with a mean absolute error of 1.44, compared to 6.58 for the baseline, which is a 78.1% reduction in estimation error. Our results reveal heterogeneous risk scaling profiles and show that models appearing robust under standard evaluation can experience rapid nonlinear risk amplification under parallel adversarial pressure. This work provides a low-cost, scalable methodology for realistic LLM safety assessment. We provide our code in the Supplementary Material.
PaperID: 6866, Poster
Abstract: Reinforcement learning (RL) has become central to post-training large language models, but its rollout generation stage is often the dominant system bottleneck. The inefficiency comes from two sources. First, autoregressive decoding makes latency grow with response length. Second, response lengths vary widely within a batch: short sequences finish early, but the batch remains blocked by the longest sequence, leaving GPU capacity underutilized. This effect is amplified in multi-sample rollouts, where long responses can also concentrate on a few decoding workers and delay synchronization. We propose LADDERS (Length-Aware Data Distribution and Existing-Response Speculation), a lightweight framework for accelerating RL rollout generation. LADDERS first uses hidden-state-based length prediction to group prompts with similar expected response lengths, and then applies an S-shaped allocation rule to balance multi-sample requests across workers. It further reuses responses already generated by the policy as draft continuations through prompt-specific suffix trees, enabling speculative decoding without an auxiliary draft model. Experiments on Qwen3-1.7B and Qwen3-8B show that LADDERS reduces rollout generation time by up to 56% while preserving final task performance, and can be integrated into existing RL systems with minimal engineering effort.
Abstract: Finding frequently occurring subgraph patterns or network motifs in neural architectures is crucial for optimizing efficiency, accelerating design, and uncovering structural insights. However, as the subgraph size increases, enumeration-based methods are perfectly accurate but computationally prohibitive, while sampling-based methods are computationally tractable but suffer from a severe decline in discovery capability. To address these challenges, this paper proposes GraDE, a diffusion-guided search framework that ensures both computational feasibility and discovery capability. The key innovation is the \underline\textnormalGra\textnormalph \underline\textnormalD\textnormaliffusion \underline\textnormalE\textnormalstimator (GraDE), which is the first to introduce graph diffusion models to identify frequent subgraphs by scoring their typicality within the learned distribution. Comprehensive experiments demonstrate that the estimator achieves superior ranking accuracy, with up to 114% improvement compared to sampling-based baselines. Benefiting from this, the proposed framework successfully discovers large-scale frequent patterns, achieving up to 30× higher median frequency than sampling-based methods.
Abstract: Highly over-parameterized models can simultaneously memorize noisy labels and generalize well, yet how these behaviors coexist remains Highly over-parameterized models can simultaneously memorize noisy labels and generalize well, yet how these behaviors coexist remains poorly understood. In this work, we investigate the underlying mechanisms of this coexistence using modular arithmetic tasks under heavy label noise. Through extensive experiments on two-layer neural networks, we find that larger models tend to generalize better under appropriate optimization and model configurations, while noisy labels are memorized faster than clean data. Over-parameterized models internally form a generalization structure, but its expression in the output is suppressed by the need to fit noisy labels. Remarkably, even with 80% label noise, near-perfect test accuracy can be achieved by extracting this internal structure using frequency-based methods. We further propose a task-agnostic method to partition networks into generalization and memorization components. Although this subnetwork improves generalization, it is limited compared with frequency-based extraction, indicating that the generalization structure is distributed across neurons and motivating the development of new tools to retrieve generalizable knowledge from over-parameterized networks.
PaperID: 6869, Poster
Abstract: Semantic segmentation typically requires dense pixel-level annotations, creating a prohibitive bottleneck for budget-constrained applications. While foundation models and vision transformers (ViTs) have redefined visual representations, current active learning (AL) strategies for these backbones largely operate at the patch or region level, where annotation requirements remain high. Extending ViT-based AL to the pixel-level, low-budget regime is uniquely challenging: with as little as one pixel query per image, a method must simultaneously identify the most informative locations and train a reliable decoder head from an extremely sparse signal. We introduce TEAL (Token-space Efficient Active Learning), the first pixel-level AL framework built around ViT representations. TEAL uses a frozen DINOv3 backbone and performs diversity-based selection directly on its native token lattice, avoiding the artifacts introduced by interpolating token embeddings to dense pixel resolutions. Our framework applies a two-level MaxHerding strategy over multi-layer token descriptors to select representative candidates, followed by margin-based uncertainty refinement on a finer decoder grid. Across CamVid, Cityscapes, ADE20K, and Pascal Context, TEAL consistently outperforms previous baselines under extreme label scarcity. After 10 rounds of 1-pixel-per-image queries, TEAL improves over the strongest baseline by up to +14.65 mIoU on CamVid and +23.34 mIoU on Pascal Context, showing that ViT token spaces provide an effective geometry for extremely-low-budget active segmentation.
PaperID: 6870, Poster
Abstract: Large vision-language models (LVLMs) can process interleaved text and multiple images in a single context, but how they internally identify the image relevant to a question remains poorly understood. We study this grounding step as in-context image retrieval and introduce ICIR-MCQ, a controlled probe for analyzing image retrieval at attention-head granularity. Using each head's attention mass over image-token spans, we first identify sparse pointer heads that reliably point to the target image. These heads transfer beyond the controlled probe: a single pointer head outperforms external vision-language retrievers on a naturalistic multi-image retrieval benchmark, showing that LVLMs contain a strong internal image retrieval signal. However, attention alone only shows where a head reads from, not what information its output passes on to later layers. We therefore introduce a complementary output-side analysis. For each head, we train a lightweight image retriever on its value-weighted output and use it to identify writer heads, whose outputs contain enough information to identify the target image. Surprisingly, the pointer and writer criteria select substantially different head sets, with only 35 to 44 heads overlapping among the top-100 heads under each criterion. Targeted set-partition ablations show that the heads most important for standard multi-image inference are concentrated in this pointer-writer intersection, while heads selected by only one criterion contribute little beyond random-head controls. These results show that attention-based pointing alone is systematically incomplete for identifying image retrieval heads in LVLMs. Causally relevant retrieval heads must also write target-image information to the residual stream, not merely point to it.
PaperID: 6871, Poster
Authors:
Sihan Ren, Gaozheng Li, Yuanshang Quan, Yiming Qin, Fuyi Yang, Chang Liu, Lan Xu, Minye WuAbstract: Simultaneous Sign Language Translation (SLT) is critical for real-time communication, yet existing methods remain largely confined to sentence-level, offline settings that assume pre-segmented inputs. These assumptions hinder deployment in realistic scenarios involving continuous, unsegmented video streams. We present ReSCUE, a unified framework for simultaneous SLT on unsegmented long-form sign language videos that aligns training and inference with realistic streaming conditions. ReSCUE combines inference-aware training to handle partial inputs, non-signing pauses, and multi-sentence contexts, stabilized re-translation to enable low-latency yet revisable predictions with reduced output flicker, and a sentence commitment mechanism for online segmentation and memory management. Experiments on standard sentence-level benchmarks show that ReSCUE achieves lower latency and the best translation quality under low-latency settings. On long-form unsegmented datasets, ReSCUE approaches the translation quality of oracle offline systems that use ground-truth sentence boundaries, while operating at substantially lower latency, demonstrating its practicality for real-world streaming scenarios.
PaperID: 6872, Poster
Authors: Shaunak Bhandarkar, Jonathan Pillow
Abstract: A major goal in computational neuroscience is to understand how animals learn to perform a new behavior from a sequence of actions and rewards. Recent work from Liebana et al. (2025) has argued that animal learning trajectories are inconsistent with learning in shallow (1-layer) networks and are better explained by gradient descent (GD) in a deep neural network (DNN). However, previous work on inferring animal learning rules from behavioral data has focused on either 1-layer networks or fits to group behavior, highlighting the need for methods to fit DNN learning rules to individual animal learning trajectories. To address this gap, we develop a method for inferring single-animal DNN-based learning rules, parametrized by a set of initial weights, a learning rate, and a softmax parameter governing stochasticity in choice behavior. We first prove that GD in a shallow network can be arbitrarily well approximated by GD in a deep network, establishing a theoretical limit on network identifiability. Using this result to refine our inference method, we analyze two behavioral datasets from mice learning a sensory decision-making task. Specifically, we compare three model architectures: 1-layer linear (1L) and 2-layer linear (2L) models considered in Liebana et al., as well as a 2-layer nonlinear network (2NL) with a sigmoidal nonlinearity after the first layer. The 2NL model achieved the best fit for 84 out of 88 mice in both datasets (by AIC). Moreover, simulated learning trajectories from the fitted 2NL model closely matched the timecourse of key behavioral metrics such as bias, stimulus sensitivity, and accuracy. These results show that both depth and nonlinearity are critical for capturing individual learning trajectories. Strikingly, we found that the inferred learning rate and initial layer-1 weights of the 2NL model systematically predicted individual differences in accuracy and task bias, linking network initialization and learning rate to individual differences in learning performance across mice.
PaperID: 6873, Poster
Authors:
Fan zhang, Ziqi Zhang, Hossein Khalili, Neiro Cabrera, Jonathan Xue, Nader SehatbakhshAbstract: On-device machine learning exposes deep neural networks as white-box artifacts, making them vulnerable to model-stealing attacks. Trusted Execution Environments (TEEs) mitigate this risk by isolating model execution, but running entire models inside TEEs incurs prohibitive overhead. To balance security and efficiency, prior work proposes TEE-Shielded DNN Partitioning (TSDP), which executes privacy-insensitive components on GPUs while confining sensitive layers to TEEs. We demonstrate that existing TSDP schemes remain vulnerable because exposed GPU weights provide an effective warm start. This allows adversaries to exploit inevitable information leakage and reconstruct high-fidelity surrogates using only a fraction of the training data. To address this vulnerability, we propose FILOsofer, a principled defense that strategically obfuscates exposed weights using Fisher Information. By deliberately rendering leaked weights misleading, FILOsofer steers an adversary’s initialization away from the true optimum. To recover user-side accuracy, we introduce a novel cross-layer LoRA mechanism that efficiently restores performance while storing only lightweight LoRA parameters inside the TEE. Extensive evaluation in real-world settings shows that FILOsofer achieves black-box–equivalent security in the worst case, while reducing computational overhead by more than 50× compared to prior approaches.
PaperID: 6874, Poster
Abstract: Large-scale learning increasingly relies on updating only a coordinate, a block, or a low-dimensional subspace per iteration to reduce memory and communication costs, yet the statistical consequence of such reduced update coverage, beyond the well-understood optimization slowdown, remains open. Available stability analyses are restricted to coordinate-level projectors whose orthogonality drives the argument, deterministic-output assumptions incompatible with randomized updates, or uniform step sizes that ignore curvature-adaptive blockwise schedules; none explains how general subspace coverage controls population risk. We analyze this question through randomized subspace descent (RSD), a framework that recovers GD, RCD, BCD, and isotropic SSD as special cases: under joint isotropy the expected projection energy reduces to a single coverage parameter~\alpha, enabling unified stability analysis without coordinate-level structure. The convex excess risk scales as O(\alpha^-1/4N^-1/2), improving to O((N\sqrt\alpha)^-1) under low noise, substantially gentler than the O(1/\alpha) optimization slowdown, while in the nonconvex Polyak--\Lojasiewicz setting \alpha lengthens the early-stopping horizon without entering the excess-risk bound as an explicit penalty. The analysis extends to the canonical blockwise step size~1/L_i through a curvature-matched Lyapunov geometry that exposes a balance condition for exact block updates, and to nonconvex objectives by decoupling gradient and curvature moments via H\"older's inequality so that stability is controlled without an almost-sure curvature lower bound. Experiments on convex and nonconvex benchmarks confirm the predicted trade-off.
PaperID: 6875, Poster
Abstract: This paper introduces Walk3D, a novel 3D Gaussian Splatting (3DGS) framework explicitly designed for 3D referring segmentation, particularly excelling at interpreting complex, long-sentence descriptions. Our primary motivation stems from observing that existing 3DGS-based open-vocabulary methods mainly focus on simple, category-level object segmentation. These methods struggle to precisely segment targets described by highly detailed, multi-sentence queries due to weak compositional language understanding and a lack of structural scene awareness. To tackle these dense and complex linguistic constraints, we propose a spatial routing network that decodes these queries into sequential chain-of-thought instructions upon a scene graph. Specifically, we construct a static supernode graph based on the geometric and semantic properties of the 3DGS space, where a spatial Graph Neural Network (GNN) dynamically pathfinds using the decomposed linguistic chain to accurately segment the referred target. Extensive experiments on complex 3D referring segmentation benchmarks demonstrate the superiority of our proposed method in interpreting lengthy, multi-conditional queries. Our code and trained models will be publicly released.
PaperID: 6876, Poster
Abstract: Personalized dialogue requires more than recalling explicit user histories: systems also need to infer hidden user states that evolve through interaction and shape appropriate response strategies. Existing memory- and profile-based methods primarily reuse observable user information, offering limited support for modeling user-state dynamics or selecting actions based on how they shape future user states. We propose PUMA (Prospective User-state Modeling for Action selection), a framework grounded in the Free Energy Principle (FEP) that formulates personalization as decision-making under partial observability, centered on an explicit user state model that captures latent user states and their action-conditioned dynamics. At each turn, PUMA maintains a belief over the user's hidden state, refines the user state model for observation generation and action-conditioned state transition, and selects dialogue actions by minimizing expected free energy—balancing epistemic and pragmatic objectives under a unified criterion. This formulation shifts personalization from passive memory retrieval to model-based decision-making over user evolution. We instantiate PUMA on healthcare-oriented counseling and motivational interviewing benchmarks with latent state annotations for rigorous evaluation. Experiments show that PUMA improves long-horizon dialogue outcomes while maintaining strong response quality, and a cross-dataset study demonstrates more reliable user-state estimation and next-state prediction. Our code is available at: https://anonymous.4open.science/r/PUMA-4DA7/.
Authors: Congye Wang
Abstract: Trimming suspicious calibration points is a common response to contamination in conformal prediction. Its effect on clean-target coverage, however, is governed by the retained law induced by trimming, not by the contamination level alone. We analyse fixed-threshold trimming as conditioning rather than purification. It replaces the contaminated calibration law with a retained law, reducing clean-target coverage to a one-dimensional score-CDF transfer problem with an exact finite-sample identity. A componentwise bound on the transfer gap gives a population-level diagnostic. This separates a clean-side covariance cost from a retained-contamination cost, governed by the dirty-to-clean retention ratio. Trimming helps when the anomaly score separates retention probabilities while remaining score-neutral on the clean population. Otherwise, it cannot substantially reduce contamination through the retained mixture coefficient. We also give finite-sample certificate templates that provide numerical guarantees under independent audit.
PaperID: 6878, Poster
Abstract: Set-level visual difference discovery is increasingly important for dataset auditing and for understanding model behavior under distribution shift, yet manually inspecting thousands of images is impractical. This motivates the emerging task of set difference captioning, i.e., describing concepts that are more often true for one image set than another. Different from existing approaches which heavily rely on large language models (LLMs), incurring substantial computational cost in terms of time and tokens, we introduce CCDiff, a lightweight, statistically grounded, and training-free framework for set difference captioning that significantly reduces LLM dependence during core difference discovery, achieving over 2× speedup and reducing token usage to zero. CCDiff operates in three stages: constructing a shared pool of candidate concepts from domain vocabularies or lightweight captioning models, filtering candidates to obtain set-specific concept pools, and performing inverse canonical correlation analysis (CCA) across sets to identify low-correlation directions that correspond to set differences. Candidate concepts are then ranked using a CCDiff score that balances inverse correlation with within-set representativeness. We evaluate CCDiff on VisDiffBench and further validate it on additional domains including MetaShift and MIMIC-CXR. Across benchmarks and real-world settings, CCDiff matches or exceeds prior LLM-based approaches while being substantially more efficient and scalable.
Abstract: Sparse Autoencoders (SAEs) are increasingly used to interpret foundation models, but their role as an actionable intervention space remains less understood, especially in vision. We study whether sparse visual features can be used not only for post-hoc analysis, but also to steer frozen vision-language models. We introduce Visual Sparse Steering (\ours), a label-free method that trains a top-k SAE on unlabeled activations from a frozen CLIP image encoder and, at test time, constructs an interpretable steering vector by amplifying the input's active sparse features and decoding the induced change. We show that this procedure admits a closed-form decomposition as \emphcentroid-deviation steering: each input is moved along its deviation from the SAE-learned centroid. The residual term is controlled exactly by the SAE's per-sample reconstruction error, measured by FVU, yielding an FVU-based residual bound and motivating a reliability gate that falls back to zero-shot CLIP when SAE reconstruction is unreliable. With target-domain SAEs trained on unlabeled CLIP image-encoder activations, \ours improves zero-shot accuracy across nine image-classification datasets, achieving gains up to +4.12% with less than 0.1% additional inference compute. Finally, a controlled upper-bound study, \oursplus, shows that selective amplification of sparse features can yield gains up to +21.44%, exposing a \emphreconstruction-vs-task saliency gap: features salient for reconstruction need not align with features useful for downstream prediction.
Abstract: The goal of this paper is to strengthen the reasoning of Omnimodal Large Language Models (Omni-LLMs) at inference time, without additional training. These models jointly process video, audio, and text, and given the large number of tokens they consume, how attention is routed across them is central to their behaviour. We focus specifically on attention sinks, tokens that absorb a disproportionate share of attention mass regardless of their semantic content, to understand how this routing unfolds. To this end, we conduct a systematic analysis of sink behaviour in Omni-LLMs. Our analysis yields two key findings: (i) high sink attention does not solely indicate head redundancy, suggesting that sink value representations play additional functional roles; (ii) the sink value vector acts as a shared bias added to every token's output, serving as a global signal that organises the representation as a whole. Building on this, we propose OutRo, which correspondingly aligns non-sink token representations with the sink in feature space, and relaxes the causal mask for sink tokens at an early layer to sharpen this bias before the rest of decoding proceeds. This design enhances the reasoning process without requiring additional forward passes or access to attention maps. Based on extensive experiments, OutRo consistently improves performance on seven video QA benchmarks and demonstrates strong generalisation, while incurring only a 1.1\x decoding overhead.
PaperID: 6881, Poster
Abstract: Open-vocabulary semantic segmentation (OVSS) models enable dense prediction for classes specified at test time, but continual test-time adaptation (CTTA) over long, non-stationary streams can corrupt their language-aligned semantic interface, which we diagnose through synonym-prompt text alignment. We propose Cross-Foundation Complementary Learning Systems (XF-CLS), a source-free OVSS-CTTA framework that makes lightweight online adaptation stable over long horizons. Inspired by Complementary Learning Systems, which couple rapid plastic learning with slow consolidation, XF-CLS decouples plasticity from stabilization: a single fast online adaptation path updates only a small set of visual normalization parameters, a horizon-aware, source-anchored slow parameter memory consolidates stable changes and recovers from drift, and a frozen cross-foundation structural observer supplies a non-drifting structural prior for reliability-aware recovery and dense output refinement. Rather than training an additional teacher or using cross-foundation pseudo-label supervision, the frozen path preserves the open-vocabulary interface while correcting degraded visual structure. On long-horizon Cityscapes-to-ACDC and OnDA streams, XF-CLS achieves 32.92 and 30.75 mIoU, improving the NACLIP no-adapt baseline by +9.00 and +4.13 mIoU, respectively, and outperforming continual OVSS-TTA baselines. Decomposition studies and text-alignment diagnostics show that source-anchored recovery prevents long-horizon collapse, while frozen cross-foundation refinement adds consistent gains without updating extra parameters. XF-CLS provides a stable, parameter-efficient route to source-free continual open-vocabulary dense prediction, with inference overhead dominated by one additional frozen-backbone forward pass.
PaperID: 6882, Poster
Abstract: Estimating the number of clusters, k, is a fundamental prerequisite for most clustering algorithms. Traditionally, this estimation has relied on computationally expensive grid searches over internal cluster validity indices (CVIs) or adaptive clustering algorithms, both constrained by specific statistical assumptions. We introduce TabK, a zero-shot, permutation-invariant transformer architecture that frames the estimation of k as an amortized Bayesian inference problem for tabular data. TabK leverages a novel Quantile Feature Encoder (QFE) to extract position-free, content-derived representations from empirical marginal distributions. By learning a generalized structural bias over a diverse synthetic prior, TabK estimates the number of clusters in a single forward pass. Evaluated across 50 diverse real-world benchmarks, TabK outperforms its closest competitor, TabClustPFN, reducing the Mean Absolute Error (MAE) of the number of clusters estimation by 1.2 (from 2.24 to 1.04). Furthermore, TabK empirically achieves the best position on the Pareto frontier among 12 evaluated methods, providing predictions of k in an average time of less than one second, which is 2× faster than TabClustPFN and 30× faster than traditional CVIs such as Silhouette and Calinski-Harabasz. Code and data are available at: https://anonymous.4open.science/r/TabK.
PaperID: 6883, Poster
Abstract: Knowledge editing has emerged as a promising approach for efficiently updating embedded knowledge in large language models (LLMs). It first computes an ideal target hidden state that steers the LLM toward the new output, and then edits the model parameters so that the original hidden state is mapped to this target. While most existing methods focus on the second stage of parameter editing in preserving other knowledge, we find that the upstream construction of the target hidden state also plays an important role in allocating the model's limited capacity space. We first establish a theoretical connection between interference in the parameter space and interference between target hidden states, showing that the two are positively correlated. Building on this insight, we propose a simple yet effective \ell_1 regularization method that prunes away unnecessary ``tentacles'' of each knowledge vector, retaining only the essential components. With fewer extraneous tentacles, each edit affects less other knowledge, enabling better coexistence. Extensive experiments across multiple LLMs and benchmark datasets verify the effectiveness of our method.
Authors: Etienne Le Naour, Tahar Nabil, Adrien Petralia
Abstract: Foundation models mark a profound paradigm shift in time series modeling, with task-specific models being superseded by general-purpose zero-shot models. Yet, current approaches primarily focus on forecasting, while real-world time series are often irregularly and partially observed, requiring models that can jointly forecast, impute missing values, and handle degraded sampling conditions. To address these challenges, we introduce TS-ICL, a novel probabilistic In-Context Learning encoder-regressor Transformer that unifies forecasting and imputation. TS-ICL formulates time series tasks as timestamp-aligned regression and naturally incorporates covariates by training on synthetic dependency structures generated from a novel causal data prior. Empirically, TS-ICL achieves a new state-of-the-art in imputation, while remaining competitive with leading forecasting foundation models across both univariate and covariate-aware benchmarks. It shows particularly strong performance in forecasting with partially observed look-back windows.
Authors: Shengzhe Chen, Mehrdad Moradi, Kamran Paynabar, Hao Yan
Abstract: We propose Flow Mismatching, an unsupervised anomaly detection method that deliberately avoids reconstruction-based paradigms. Instead, we treat flow matching as geometric dynamics and leverage a key insight: anomalies occur at places where the learned normal flow disagrees with the geometric path toward a test image. Given a flow matching model trained only on normal images, we probe its learned velocity field along affine paths from Gaussian noise to a target image. Along each path, we compare the model-predicted velocity, which follows normal generative dynamics, with the geometric velocity toward the target, which includes any anomalous content. Anomalies induce strong local disagreement between these velocities. Aggregating the mismatch over different time steps and multiple paths yields pixel-wise heatmaps and image-level scores without test-time optimization, feature memories, or additional calibration. Our analysis shows that the population mismatch decomposes into an irreducible denoising term and a Fisher-divergence term between the test-path and normal-path score functions, which identifies the score-gap component that drives anomaly separation and explains the effectiveness of robust path aggregation. Extensive experiments on MVTec-AD and VisA demonstrate superior performance compared with SOTA reconstruction-based and recent flow matching-based approaches.
PaperID: 6886, Poster
Abstract: Neural processes (NPs) are flexible uncertainty-aware meta-learners, but often fail when meta-training tasks are scarce, losing generalization and collapsing into task memorization. We propose Context-Subset Self-Distillation Neural Processes (CSSDNP), a training framework that trains a student on random context subsets while distilling the full-context predictive distribution of an EMA teacher from the same task. This context-asymmetric objective combines reduced-context predictive learning with forward-KL distillation, enforcing conditional consistency across different context views without architectural or test-time changes. Theoretically, we show that context-subset training learns the reduced-context prediction problem while augmenting each task, and that calibrated full-context distillation preserves the asymptotic Bayes target while improving finite-sample estimation via teacher-guided shrinkage. Empirically, CSSDNP improves robustness and generalization across NP variants and tasks, especially in low-meta-data regimes.
PaperID: 6887, Poster
Authors:
Adam Shai, Thomas J Elliott, Paul RiechersAbstract: Mechanistic interpretability often studies the local features and circuits that implement model computations. What principles govern the arrangement of these features and circuits into geometric structures in activation space? To make this tractable, we study how the computational class of the training-data generator constrains the geometry of predictive states. We show that while the data distribution determines which features are required for prediction, a predictor realizes those features as beliefs about its current latent state, and the generator class determines the geometry of those beliefs. Using this theoretical insight, we design synthetic datasets whose minimal predictive representations fall into different model classes, and test which geometry neural networks learn. In particular, we train transformers, LSTMs, GRUs, and vanilla RNNs on datasets whose predictive geometries are known analytically: a classical HMM process, a quantum-realizable process with no finite-state HMM realization, and a generalized-probabilistic process with no finite-dimensional quantum realization. Across architectures, a single affine map from activations decodes the corresponding predictive representation in each case: HMM beliefs in a latent simplex, Bloch-vector quantum states, or a finite-dimensional generalized predictive vector. These representations emerge during training and fit the compact non-classical geometry far better than finite-order classical Markov baselines. These results suggest that understanding predictive representations requires asking not only which features a network represents, but what geometry organizes those features.
Abstract: Diffusion Transformers rely on static patchify tokenization, assigning the same token budget to smooth backgrounds, detailed object regions, noisy early timesteps, and late-stage refinements. We introduce the Dynamic Chunking Diffusion Transformer (DC-DiT), which replaces fixed patchification with a learned encoder-router-decoder scaffold that adaptively compresses the 2D input into a shorter token sequence through a chunking mechanism learned end-to-end with diffusion training. DC-DiT allocates fewer tokens to predictable regions and noisy timesteps, and more tokens to detailed regions and later refinement stages, yielding meaningful spatial segmentations and timestep-adaptive compression schedules without supervision. Furthermore, the router provides an importance ordering over retained tokens, enabling elastic inference: a single checkpoint can be evaluated at flexible compute budgets with a smooth quality-compute tradeoff. Additionally, DC-DiT can be upcycled from pretrained DiT checkpoints and is also compatible with orthogonal dynamic computation approaches. On class-conditional ImageNet generation, DC-DiT reduces inference FLOPs by up to 36.8% and improves FID by up to 37.8% over DiT baselines, yielding a stronger quality--compute Pareto frontier across model scales, resolutions, and guidance settings. More broadly, these results suggest that adaptive tokenization is a general mechanism for making visual generation both more efficient and more flexible at inference time.
Abstract: Large reasoning models (LRMs) reach competition-level math and coding accuracy via long autoregressive decoding, making per-token decoding cost a primary deployment concern. Weight quantization is the standard tool for acceleration, but representative recipes --- including state-of-the-art end-to-end (E2E) QAT --- lose accuracy on long-decoding reasoning benchmarks despite preserving perplexity and short-decode accuracy. Through a systematic gradient-direction analysis, we identify two factors driving this gap: (i) KV-cache fidelity preservation under the QAT loss, which E2E supervision attenuates via the softmax Fisher metric; and (ii) Hessian-subspace alignment between calibration data and the deployment distribution. We propose LookAhead Quantization (LAQuant), a layer-wise weight-only QAT method that addresses both factors without online-transform overhead by combining reasoning-domain calibration with a one-layer lookahead loss whose implicit cross-layer co-adaptation preserves the next-layer residual stream. For Qwen3-4B under W3G128 quantization, LAQuant improves AIME25 Pass@1 over ParoQuant by 15.11pp (1.93pp over ParoQuant++ at matched calibration) while achieving a 3.42× decoding speedup over FP16 on RTX A6000, compared with ParoQuant's 3.01×.
Authors: Bryce Grant, Xijia Zhao, Peng Wang
Abstract: A fine-tuned Vision-Language-Action (VLA) policy will pick up the alphabet soup and place it in the basket on demand, then drop the soup off the table when an evaluator shifts the basket five centimeters left. We use activation injection to ask what the policy is actually doing: inject task A's activations into task B's scene at the action expert, and π₀.₅ executes task A's motor trajectory in 99.6% of episodes (n=1,968); X-VLA does so in 99.8%. The injected program is bound to absolute workspace coordinates rather than the visible scene, which mechanistically explains the perturbation brittleness reported in concurrent benchmark work. The same intervention framework applied across six VLAs (π₀.₅, OpenVLA-OFT, X-VLA, SmolVLA, GR00T N1.5, and ACT as a language-free control) on LIBERO, MetaWorld, SimplerEnv, and ALOHA over 420,000+ rollouts surfaces three further findings: language is encoded by every architecture yet behaviorally ignored when vision identifies the goal; SAE pooling preference splits along architecture lines (π₀.₅ per-token, X-VLA mean-pool, SmolVLA indifferent); and pathway specialization replicates wherever expert and VLM pathways are separable. SmolVLA's interleaved fusion attenuates the headline to 52.1% LIBERO override, scoping universality. We will release 424 trained SAEs and an interactive feature-exploration platform on acceptance; an anonymized snapshot is included in the supplementary material.
Abstract: Test-Time Adaptation (TTA) for black-box models accessible only via APIs remains a largely unexplored challenge. Existing approaches such as post-hoc output refinement offer limited adaptive capacity, while Zeroth-Order Optimization (ZOO) enables input-space adaptation but faces high query costs and optimization challenges in the unsupervised TTA setting. We introduce BETA (Black-box Efficient Test-time Adaptation), a framework that addresses these limitations by employing a lightweight, local white-box steering model to create a tractable gradient pathway. Through a prediction harmonization technique combined with consistency regularization and prompt learning-oriented filtering, BETA enables stable adaptation with no additional API calls and negligible latency beyond standard inference. BETA achieves a +7.1% accuracy gain on ViT-B/16 (ImageNet-C) and +3.4% on CLIP (ImageNet-S/R), surpassing strong white-box and gray-box methods including TENT and TPT. On a commercial API, BETA achieves comparable performance to ZOO at 250x lower cost while maintaining real-time inference speed, establishing it as a practical and efficient solution for real-world black-box TTA.
Abstract: Leveraging Graph Neural Networks (GNNs) as graph encoders and aligning the resulting representations with Large Language Models (LLMs) through alignment instruction tuning has become a mainstream paradigm for constructing Graph Language Models (GLMs), combining the generalization ability of LLMs with the structural modeling capacity of GNNs. However, existing GLMs that adopt GNNs as graph encoders largely overlook the problem of aligning GNN-encoded representations across domains and tasks with the LLM token space to obtain unified graph tokens, thereby limiting their ability to generalize across diverse graph data. To bridge this gap, we aim to incorporate a multi-domain, multi-task GNN encoder into GLMs and align its representations with LLMs to enable multi-domain, multi-task graph alignment instruction tuning. This alignment problem remains underexplored and poses two key challenges: 1) learning GNN-encoded representations that are simultaneously generalizable across domains and tasks and well aligned with textual semantics is difficult, due to substantial variations in graph structures, feature distributions, and supervision signals, together with the lack of textual-semantic alignment guidance in task-specific GNN training; 2) diverse graph data and task-specific instructions can exhibit different degrees of compatibility with the LLM token space during instruction tuning, leading to varying alignment difficulty and rendering a fixed alignment strategy suboptimal. To tackle these challenges, we propose UniGraphLM, a Unified Graph Language Model that incorporates a multi-domain, multi-task GNN encoder to learn generalizable graph representations aligned with textual semantics, and then adaptively aligns these representations with the LLM. Specifically, we first develop a graph-text pair pretraining strategy with a tailored GNN encoder, trained on large-scale graph-text data spanning multiple domains and tasks to obtain generalizable representations naturally aligned with textual semantics. We further design a curriculum alignment tuning strategy that adaptively adjusts the alignment process by accounting for varying alignment difficulty across diverse graph data. Extensive experiments demonstrate that UniGraphLM consistently outperforms state-of-the-art baselines across graph datasets from different domains and tasks.
PaperID: 6893, Poster
Abstract: Language brain-computer interfaces (BCIs), including speech and handwriting BCIs, hold significant promise for restoring communication in individuals with paralysis. However, the performance of current language BCIs remains limited due to the non-stationarity of neural activity. Existing neural decoding approaches typically treat neural non-stationarity as stochastic noise. For language BCIs, however, the noise is more complex due to the influence of semantic context in language sequences. For example, the articulation of the same phoneme can vary subtly depending on its surrounding content. Such context-induced variations introduce additional noise into neural representations and can confuse neural decoders, yet this issue has been largely overlooked in previous studies. Here, we propose explicitly modeling these context-induced noises to enhance neural decoding performance in language BCIs. Our key insight is that, unlike stochastic noise, context-induced noise carries useful information that can help the decoder disambiguate linguistically similar units. Therefore, by separating and modeling context-induced noise, we achieve improved decoding performance for language BCIs. Experiments on multiple datasets demonstrate that our approach achieves state-of-the-art performance.
PaperID: 6894, Poster
Abstract: Reasoning across multiple long-form videos requires models to find sparse evidence within each video and link related entities, events, and narratives across streams. Existing long-video benchmarks mainly evaluate single-video understanding, while multi-video benchmarks typically use short clips. We introduce CLoVR-Bench, a comprehensive benchmark for Cross-Long-Video Reasoning with 2,000 expert-annotated QA pairs over 400 long-form videos. Its three-level taxonomy covers Comparative Analysis, Tracking and Retrieval, and Integrated Reasoning, spanning 14 tasks and 36 subtasks that diagnose both intra-video localization and inter-video alignment. Our evaluation of 13 representative MLLMs shows substantial gaps across this hierarchy, with performance degrading from localized comparison to long-range retrieval and integrated cross-video reasoning. We further propose HOLMES, a training-free framework that formulates cross-long-video reasoning as evidence-slot filling over a typed binding graph. HOLMES plans option-discriminating evidence slots, localizes evidence with risk-conditioned policies, records coverage certificates for failed searches, verifies cross-video bindings visually, audits constraints, and returns answers through hard graph-readout gates. Experiments show that HOLMES improves both open-source and closed-source backbones, outperforms adapted long-video baselines, and yields the largest gains on tasks requiring long-range evidence search and cross-video binding. Together, CLoVR-Bench and HOLMES provide a diagnostic testbed and an interpretable baseline for future cross-long-video reasoning.
PaperID: 6895, Poster
Abstract: Sign Language Video Generation (SLVG) aims to generate realistic and motion-accurate sign language videos from spoken text. Most existing studies heavily rely on intermediate representations (\eg poses or 3D meshes), and adopt a two-stage generation paradigm, where the text is first converted into poses or other visual modalities, which are then used as conditions to synthesize sign language video frames via diffusion models (\eg pose-to-image models). However, the end-to-end SLVG task remains largely unexplored. In this paper, we demonstrate that end-to-end SLVG is feasible without relying on intermediate representations(\eg poses), and we argue that a key challenge lies in how to generate effective and temporally consistent conditions. With well-designed condition modules, it becomes feasible to build an end-to-end SLVG model without relying on any intermediate modalities. To this end, we propose LLMCond, a first end-to-end text-to-video SLVG framework that requires no auxiliary visual cues during either training or inference. LLMCond consists of four key components: (1) a video VQVAE that compresses sign language videos into a sequence of discrete latent tokens; (2) an LLM-based conditioner that generates temporally and semantically aligned conditions corresponding to the sign sequences; (3) a Gaussian Condition Refiner designed to enforce temporal smoothness across conditions, enabling the generation of coherent and natural sign language motions; and (4) a discrete diffusion model that synthesizes motion-accurate sign language videos conditioned on the refined condition sequences. Extensive experiments on public SL datasets demonstrate that LLMCond achieves highly competitive performance, producing temporally coherent and accurate sign language motions without relying on auxiliary visual modalities such as pose, depth, or optical flow.
Authors: Annie Ulichney, Amanda L Coston
Abstract: Understanding how a prediction model will perform in a new environment before deployment is essential to preventing harm when algorithms inform decision-making. Two common sources of model performance degradation are (i) covariate shift, where the target covariate distribution differs from the source, and (ii) selective labels, where the observability of outcomes depends on historical decisions. We study pre-deployment model evaluation under the joint presence of covariate shift and labeling of outcomes selectively based on observed features. In particular, we present a double machine learning procedure for estimating the target risk of an arbitrary black-box prediction model under a general loss function. We show identification of this estimand under standard assumptions and derive a bias-corrected estimator based on the influence function of the target risk. Finally, we evaluate our estimator through experiments using the eICU electronic health records database, showing that it tracks the true target risk more accurately than methods that address either selective labels or covariate shift alone, as well as baselines that combine standard plug-in approaches.
PaperID: 6897, Poster
Authors:
Shizhuo Zhang, Nuowen Kan, Chenglin Li, Rui-Xiao Zhang, Yifan Zhao, Yuanpeng He, Huixin Zhang, Fan He, Hang Xu, Hao Zhang, Wenrui Dai, Junni Zou, Hongkai XiongAbstract: In retrieval-augmented generation (RAG) scenarios, the prefill delay of the large language model (LLM) long inputs is the dominant component of time-to-first-token (TTFT). Reusing KV caches across different LLM inputs reduces this cost, but the resulting cross-context mismatch on reused segments has a significant degradation on answer quality via propagating through prefill attention into decode-time logits. Existing selective-recomputation methods mitigate this issue by recomputing tokens chosen by empirical importance heuristics, which overlook the propagation of cache mismatches into answer-side errors, thereby failing to effectively strike the quality-TTFT tradeoff. To address this issue, we propose a perturbation-theoretic KV cache reuse framework, KVFocus, which selects reused KV caches through a perturbation-theoretic token-risk score during the prefill process. Specifically, we first develop a first-order analysis that traces reuse error in three stages: source mismatch on the reused segment, suffix contamination during prefill, and decode-time logit stability. This analysis reveals that the leading-order answer-side error decomposes into a product of source-side V-drift and downstream suffix-to-segment attention concentration. Guided by this multiplicative structure, we define a token-risk score that operationalizes the bound during the inference. A plug-and-play selector for KV cache reuse in existing RAG-based LLM serving is then designed by recomputing the top-r% tokens directly suppresses the predicted leading-order answer-side error, covering both sides of the pathway that prior one-sided heuristics miss. Empirical results across three 7-8B instruction-tuned LLMs (Mistral, Qwen2.5, Qwen3) and four multi-hop QA benchmarks demonstrate that the proposed KVFocus achieves a superior trade-off between the ansewer quality and the TTFT in comparison to state-of-the-art baselines, with up to 3× TTFT speedup over full prefill at matched answer quality.
PaperID: 6898, Poster
Abstract: When labeling is expensive or datasets exceed computing capacity, selecting an informative subset of data is a practical necessity. Identifying such subdata is a classical NP-hard problem due to its inherent discreteness. While many subdata selection methods have been proposed for parametric models, two fundamental challenges remain: first, how can we accurately assess the statistical efficiency of selected subdata relative to the theoretical optimum? Second, given N data points and a subdata size n, which n points should be selected to achieve high statistical efficiency? We address both challenges within a unified framework. We develop a new information-based subdata selection methodology grounded in optimal approximate design theory, yielding subdata that approaches the theoretical optimum. Our algorithm is general, accommodates arbitrary N and n, supports multiple optimality criteria, and is accompanied by a convergence proof. Crucially, our framework provides, for the first time, tight lower and upper bounds on the statistical efficiency of subdata selected by any method, enabling accurate and rigorous benchmarking across the literature. This benchmarking capability reveals the true performance landscape of existing methods and shows that many of them have been operating substantially below their theoretical potential. The subdata produced by our methodology is highly efficient and outperforms all existing methods.
PaperID: 6899, Poster
Abstract: The Segment Anything Model 3 (SAM 3) has achieved significant progress in Promptable Concept Segmentation (PCS) by processing short noun phrases to generate segmentation masks with unique instance identifiers. While SAM3 excels at single-concept segmentation, it falls short in complex, real-world environments where multiple concepts must be segmented simultaneously. In particular, SAM3 typically fails to simultaneously process concepts with semantic overlap, leading to semantic misclassifications and redundant mask predictions. Moreover, as the number of concepts increases, the independent inference paradigm introduces prohibitive inference latency. To address these limitations, we propose the Dense Competitive Representations SAM (DCR-SAM) framework. By decoupling cross-modal interactions, the proposed method alleviates the inference latency. To resolve semantic misclassifications, DCR-SAM dynamically aligns textual concepts with visual representations and incorporates a dense competition head to enforce explicit inter-class competition. Furthermore, a learnable background token is applied to absorb non-target objects dynamically. These mechanisms effectively suppress overlapping predictions and generate high-fidelity masks. Extensive experimental evaluations across multiple benchmarks demonstrate that the proposed framework: I) achieves a 2.3% ~ 10.7% performance improvement over state-of-the-art methods, II) delivers a 33× inference speedup compared to the native SAM 3, and III) exhibits strong robustness across diverse scenarios. All code will be released.
PaperID: 6900, Poster
Abstract: Reference-conditioned image-to-image (I2I) editing models provide a promising zero-shot approach to personalized image generation by jointly conditioning on a reference image and a text prompt. We show that recent I2I editing models, including FLUX.1 Kontext and FLUX.2, achieve strong subject identity preservation and prompt alignment, but exhibit seed-wise diversity collapse: under the same reference image and prompt, different random seeds often produce highly similar outputs. We further find that this collapse is associated with over-alignment of latent-token trajectories during denoising. Based on this observation, we introduce Selective Hidden Perturbation (SHiP), a training-free inference-time intervention that perturbs only the attention-output representations of latent tokens during early denoising while keeping reference-image and text tokens unchanged. Across diverse subjects and prompts, SHiP improves seed-wise diversity on both FLUX.1 Kontext and FLUX.2 while largely preserving identity and prompt fidelity. On FLUX.1 Kontext, SHiP increases Vendi Score from 1.546 to 2.066 and LPIPS from 0.475 to 0.621, with only minor decreases in GPT-ID and GPT-Text. Code is available at https://anonymous.4open.science/r/i2i-personalization-diversity-0260/.
Abstract: Kullback-Leibler (KL) regularization is ubiquitous in reinforcement learning (RL) algorithms in the form of reverse or forward KL. Recent studies have demonstrated \epsilon^-1-type fast rates for RL under reverse KL regularization, in contrast to the standard \epsilon^-2-type sample complexity. However, for forward-KL-regularized objectives, existing statistical analyses are either not applicable or resulting in \tildeO(\epsilon^-2) slow rates. We take the first step towards addressing this problem via a streamlined analysis of forward-KL-regularized offline contextual bandits. We give the first \tildeO(\epsilon^-1) upper bounds in tabular and general function approximation settings, both under notions of single-policy concentrability. In particular, our convex-analytical pipeline unifies these settings by exploiting the pessimism principle in a novel way and completely bypasses the proof routines in previous works based on the mean-value theorem, which might be of independent interest. Moreover, we provide rate-optimal lower bounds, manifesting the tightness of our upper bounds in terms of statistical rates. Our lower bounds also demonstrate that the forward-KL-regularized sample complexity recovers the unregularized slow rate in the low-regularization regime, similarly to the reverse-KL regularization.
PaperID: 6902, Poster
Abstract: Autonomous driving fundamentally demands dense understanding of visual, semantic, and geometric information, yet existing Vision-Language-Action (VLA) models are typically trained with sparse supervision such as language instructions or trajectory signals. This direct mapping from dense visual observations to sparse representations forces the model to compress or even discard critical scene information essential for safe driving. We argue that this limitation does not stem from insufficient model capacity, but from the lack of explicit mechanisms that encourage the formation of dense world representations in the latent space. In this spirit, we propose a novel paradigm for driving intelligence—Think Densely, Act Sparsely—and instantiate it with an efficient and interpretable framework, LECDrive (Latent Expert Cognitive Chains for Vision-Language-Action autonomous Driving). The core idea is to mimic the hierarchical perception process of humans by embedding a progressive chain of latent expert cognition within the VLA model. Concretely, we inject a compact set of task-specific latent tokens as carriers, enabling the model to internalize dense knowledge from foundational vision models, including visual representations (DINOv3), semantic structures (SAM3), and spatial geometry (DepthAnything3), while maintaining efficient inference. This interdependent cognitive chain follows a curriculum learning principle, where supervision is progressively structured from representation to semantics and geometry. Extensive experiments demonstrate that LECDrive consistently improves the performance of base VLMs across multiple autonomous driving benchmarks, covering key tasks such as object perception, state prediction, and trajectory planning. These findings suggest that constructing dense world models through structured latent expert chains is a promising direction for VLA-based autonomous driving.
Abstract: Large language models (LLMs) can generate student-like responses fluently enough to serve as simulated students, i.e., virtual learners used to train and evaluate AI tutors and human educators. Yet such simulators are typically evaluated by output similarity to real students, not by whether they behave like students with coherent misconceptions during interaction. We introduce a controlled framework for evaluating misconception faithfulness, whether a simulator maintains a misconception-driven belief state and updates selectively when feedback addresses the underlying misconception. Central to our framework is a misconception-contrastive feedback protocol that compare targeted feedback against two controls: misaligned feedback, targeting a different but plausible misconception, and generic feedback, which only signals that the answer is incorrect. We propose Selective Flip Score (SFS), which quantifies how much more often a simulator flips its answer under targeted feedback than under the contrastive controls. Across seven LLMs (4B–120B), multiple datasets, and prompting strategies, simulators exhibit near-zero SFS, correcting their answers at similarly high rates regardless of feedback relevance to the true misconception. Further analysis reveals a sycophantic failure mode: models behave less like students with stable misconceptions and more like problem solvers who treat any corrective signal as a cue to abandon the simulated misconception and recompute from internal knowledge. To improve faithfulness, we develop a post-training pipeline spanning supervised fine-tuning (SFT), preference optimization, and reinforcement learning (RL) with an SFS-aligned reward. SFT yields notable gains up to +0.56; SFS-aligned RL provides more consistent further improvements than preference optimization. Our results establish misconception faithfulness as a challenging but trainable property of student simulators, motivating a shift from static output matching toward interaction- and belief-aware student modeling.
Authors: Konstantin Riedl, Justin Sirignano, Konstantinos Spiliopoulos
Abstract: A convergence analysis is developed for the regularized Newton method for training neural networks (NNs) in the overparameterized limit. We prove that as the number of hidden units tends to infinity, the NN training dynamics converge in probability to the solution of a deterministic limit equation involving a "Newton neural tangent kernel" (NNTK). Explicit rates characterizing this convergence are provided and, in the infinite-width limit, we prove that the NN converges exponentially fast to the target data (i.e., a global minimizer with zero loss). Crucially, we show that this convergence is uniform across the frequency spectrum, addressing the spectral bias inherent in gradient descent. The eigenvalues of the neural tangent kernel for gradient descent accumulate at zero, leading to slow convergence for target data with high-frequency components. In contrast, the NNTK has uniformly lower bounded eigenvalues if the regularization parameter is selected appropriately, allowing Newton's method to converge more quickly for data with high-frequency components. Mathematical challenges that need to be addressed in our analysis include the implicit parameter update of the Newton method with a potentially indefinite Hessian matrix and the fact that the dimension of this linear system of equations tends to infinity as the NN width grows. This substantially complicates deriving the training dynamics in the overparameterized limit as well as proving the convergence of the finite-width dynamics thereto. The analysis identifies a scaling formula for selecting the regularization parameter, which we show can vanish at a suitable rate as the number of hidden units becomes larger. In addition, we prove that, for sufficiently large numbers of hidden units, the regularized Hessian remains positive definite during training and the Newton updates for individual NN parameters converge to zero, demonstrating that the model behaves as a linearization around the initialization.
Abstract: We study offline reinforcement learning under Q^\star-approximation and partial coverage, a setting that motivates practical algorithms such as Conservative Q-Learning (CQL) [KZTL20] but has received limited theoretical attention. Our work is inspired by the following open question: \emphAre Q^\star-realizability and Bellman completeness sufficient for sample-efficient offline RL under partial coverage? We answer this question in the negative through an information-theoretic lower bound. To identify additional structure that enables sample-efficient offline RL under partial coverage, we introduce a general decision-estimation framework, inspired by model-free decision-estimation coefficients (DEC) for online RL [FGO^+23, LWZ25b]. Our framework decomposes the complexity of offline RL into two parts: the \emphdecision complexity and the \emphvalue estimation error. This decomposition allows us to study the two sub-problems in a modular way. Our result not only unifies existing results in the Q^\star-approximation and partial coverage regime [CJ22, UKLS23], but further improves and generalizes them. On the decision complexity side, our improvement includes: the first \epsilon^-2 sample complexity bound for soft Q-learning under partial coverage that improves [UKLS23]'s \epsilon^-4 bound, the removal of the need for additional online interaction in the gap-dependent setting of [CJ22], and new learnable settings beyond the above two cases. On the value estimation side, we provide the first characterization of offline learnability for general low-Bellman-rank MDPs [JKA^+17, DKL^+21, JLM21], a canonical online RL setting that has remained unexplored in offline RL outside special cases. As a side contribution, our techniques give the first analysis of CQL in the function approximation setting.
Authors:
Mingwei Zheng, David OBrien, Siwei Cui, Pardis Pashakhanloo, RAJDEEP MUKHERJEE, Myeongsoo Kim, Sachit KuharAbstract: LLM coding agents operate by constructing trajectories that accumulate reasoning, tool calls, and results to enable multi-step decision-making. However, the conventional append-only trajectory architecture found in practice tightly couples file-read actions with their observations, capturing snapshots that become permanently fixed in the chronological history. As files change through agent edits or concurrent human modifications, these snapshots become stale, causing reasoning errors and causing agents to redundantly re-read files, with each re-read appending yet another copy to the trajectory. To mitigate this, we propose CORVUS, a novel trajectory architecture that decouples file-read actions from their observations by maintaining a synchronized registry of relevant files and injecting only their current contents at each reasoning cycle. This structural change produces significantly lighter-weight trajectories that remain synchronized with the actual codebase state by construction, eliminating redundant file copies and stale snapshots that bloat conventional trajectories. We evaluated CORVUS on SWE-POLYBENCH_VERIFIED and SWE-BENCH PRO across four LLMs, achieving 9–50% reduction in average input tokens per task, 15–32% shorter final prompts, and up to 37% fewer reasoning cycles while maintaining comparable pass rates.
Abstract: Constrained decoding is essential for serving LLMs, ensuring that generated outputs follow specific structures such as JSON schema-formatted function calls. Existing systems are designed for autoregressive models and assume left-to-right generation, masking out invalid next tokens at each step. Diffusion language models, however, break this assumption: they sample multiple positions simultaneously from a fully-factorized mean-field distribution at each denoising step. In this paper, we present an exact and tractable algorithm for sampling from the constrained mean-field posterior under any constraint expressible as a finite automaton. Viewing finite automata as graphical models, we obtain tractable representations of the constrained distribution that enable efficient inference. The approach guarantees constraint satisfaction by construction, supports both greedy and sampling-based decoding, and is compatible with parallel and block-wise decoding under arbitrary remasking schedules. Applying depth-reduction techniques from arithmetic circuit theory, we further reduce sampling depth from linear to logarithmic in the sequence length. Empirical evaluations on Dream-7B and LLaDA-8B show substantial accuracy gains across various tasks including function calling (xLAM, BFCL), planning (Sudoku, Countdown), text-to-SQL (Spider), and math reasoning (GSM-Symbolic), with little inference overhead relative to unconstrained decoding. For example, on BFCL-Live, our approach improves Dream-7B's greedy decoding accuracy from 63.9% to 71.5%, and stochastic sampling accuracy from 22.3% to 69.0%, where the unconstrained baseline collapses, with under 5% wall-clock overhead.
Abstract: Many applications require statistically valid inference across many related tasks, while providing only a handful of high-quality labels per task. In AI evaluation, these tasks may correspond to model behaviors across prompts, subgroups, or hypotheses; in social science surveys, they may correspond to related questions, populations, or measurement conditions. Prediction-powered inference uses inexpensive proxy measurements to improve inference from limited labels, but standard methods operate task-by-task and therefore struggle in the small-label regime. We introduce a multi-task prediction-powered inference framework that borrows strength across tasks without pooling away validity. Our methods learn surrogate outcomes using labeled data from other tasks while retaining within-task rectification for valid per-task confidence intervals. We prove that efficiency gains beyond power-tuned PPI require nonlinear structure in the proxy–ground-truth relationship: affine cross-task recalibrations are oracle-equivalent to using the original proxy. We complement our theoretical findings with experiments on semi-synthetic datasets and a case study auditing language models on election-related information during the 2024 U.S. presidential election. Using a large human-annotation study, we show that cross-task surrogate learning can substantially reduce confidence interval widths when labels are scarce.
Abstract: Diffusion large language models (dLLMs) such as LLaDA and Dream now rival autoregressive LLMs in quality while retaining native parallel decoding. To deploy them efficiently, block-wise decoding partitions generation into blocks of size B and commits a fraction of each block before moving on, but B conflates two roles: \emphhow far the model can look ahead and \emphhow many tokens get committed per step. Recent accelerators relax this trade-off with indirect heuristics, yet the underlying difficulty is already exposed at every forward pass by the model's per-step confidence. Specifically, the in-window confidence profile exhibits a context-dependent \emphcliff (a high plateau, a sigmoidal transition, and a residual floor) that directly encodes how far ahead is safe to look. We propose PACE-dLLM, a decoder that handles the two roles independently. At each step, PACE-dLLM fits the cliff in closed form on the current window's confidences and reads off the next prediction horizon at the cliff's saturation point; a separate confidence threshold decides which tokens are committed. Both rules read the same signal, add a single hyperparameter, and admit a strict-dominance guarantee over fixed-block decoding. Extensive experiments across four reasoning and code benchmarks on two open-source dLLM backbones demonstrate that PACE-dLLM achieves the best average accuracy at a roughly \mathbf4.5× wall-clock speedup over the unaccelerated semi-AR baseline, pushing the Pareto frontier outward.
PaperID: 6910, Poster
Abstract: Reliable uncertainty estimation is essential for deploying large language models in high-stakes reasoning tasks, yet most existing approaches rely on expensive sampling strategies or shallow token-level signals that ignore the internal logical structure of model outputs. We introduce Argument Graph Uncertainty, a post-hoc framework for multiple choice question tasks, that estimates model confidence directly from the logical structure of a single chain-of-thought reasoning trace. Our method uses a lightweight instruction-tuned model to segment the reasoning chain into propositional units, then applies a simple NLI model to infer relational edges and construct two complementary argument graphs over these segments. The resulting edge scores are aggregated into a Dirichlet posterior over answer options, whose concentration yields calibrated uncertainty estimates reflecting both local step-to-step coherence and global logical consistency across the reasoning trace. Evaluated on GPQA, a challenging graduate-level multiple-choice benchmark, our method consistently surpasses token-based and consistency-based uncertainty baselines across multiple reasoning models. We further demonstrate through extensive combination analyses that argument graph uncertainty captures an orthogonal signal to existing approaches, making it a complementary component in uncertainty ensembles. Crucially, our framework requires no additional calls to the evaluated model and relies entirely on small, efficient auxiliary models, producing fully interpretable, per-node uncertainty attributions at low computational cost.
PaperID: 6911, Poster
Abstract: Image super-resolution in the dark is fundamentally challenged by extreme spatial heterogeneity: severely underexposed and noisy regions demand aggressive generative enhancement to reconstruct missing details, while relatively well-exposed areas require conservative updates to prevent hallucinated textures. Standard diffusion and flow-matching models, however, rely on globally uniform integration trajectories, making them sub-optimal for such spatially varying degradations. In this paper, we propose DeLang-SR (Degradation-Language steered Super-Resolution), a one-step flow-matching framework that explicitly decouples the restoration direction and magnitude. Rather than treating low-light corruption as arbitrary latent features, DeLang-SR translates measurable physical statistics into structured degradation language, harnessing the semantic prior of pretrained text-to-image models. The degradation-language intent, complemented by image-specific visual tokens, steers the velocity field (direction) toward restoration-relevant manifolds. Simultaneously, a continuous spatial update-strength field acts as a locally varying Euler step-size field (magnitude), applying stronger integration steps in heavily degraded shadows while enforcing conservative updates in reliable regions. Experiments on paired and real dark images, together with ablations and diagnostic analyses, show that DeLang-SR provides an effective and interpretable one-step route to super-resolution in the dark.
Abstract: High-resolution 3D medical image generation remains challenging because fully volumetric models are computationally expensive, while efficient 2D slice generators often fail to preserve anatomical consistency across the third dimension. We propose LiFT, a framework for Lifted Inter-slice Feature Trajectories that factorizes 3D volume synthesis into per-slice image generation and inter-slice trajectory learning. Rather than modeling the volumetric distribution end-to-end, LiFT treats a volume as an ordered trajectory in feature space, capturing how anatomical structures appear, transform, and disappear across depth. A tri-planar drifting loss aligns the trajectory of generated slices with the trajectories of real volumes, enabling distributional learning over inter-slice progressions in unconditional generation; in paired translation, a bidirectional z-context mixer trained against the registered target supplies through-plane coherence while preserving per-slice fidelity. We evaluate LiFT on BraTS 2023 (unconditional and missing-modality MR) and SynthRAD2023 (MR-to-CT). Across these settings, LiFT preserves per-slice quality, approaches the reported cWDM missing-MR reconstruction quality, and improves through-plane coherence on MR-to-CT relative to a no-mapper ablation, demonstrating that lightweight inter-slice trajectory learning is a viable route to high-resolution 3D medical synthesis.
PaperID: 6913, Poster
Abstract: Detecting frequency trend patterns, such as items with sustained growth or decline, is an important and challenging task in high-speed data streams. State-of-the-art solutions require manual tuning of many coupled parameters, while their bloom filters suffer from high false positive rates. In this paper, we propose NeuTrend, a framework that provides accurate and automatic trend detection with high-performance in-network inference. Our key idea is that a lightweight BRNN can learn to classify incoming items as likely trending or non-trending. Thus, it eliminates unqualified items more accurately than a bloom filter and removes the need for manual parameter tuning. The BRNN uses only XNOR and popcount operations. Thus, it aligns with the strict resource constraints of data-plane switches. NeuTrend further introduces an adaptive detection module. The module uses the BRNN confidence scores to automatically tune the growth and decay thresholds. Extensive experiments on real-world datasets show that the learned filter reduces false positive rates by 23%-44% and increases effective detector occupancy from 47% to 70%. Hence, NeuTrend improves the overall F1 score by 35%-68% over existing solutions, and achieves Tbps-level line-rate processing.
PaperID: 6914, Poster
Authors: Walid Bendada, Guillaume Salha-Galvan
Abstract: Sampling from a softmax distribution is a fundamental operation in machine learning, but its linear complexity in the number of items makes exact sampling impractical at scale. Two-level softmax (2LS) sampling is a popular alternative enabling sublinear-time sampling. Assuming items are partitioned into clusters, 2LS first samples a cluster and then an item within it. In this paper, we show that, despite its advantages, 2LS introduces systematic and undesirable sampling biases, which arise from misweighting clusters by ignoring both cluster size imbalance and intra-cluster similarity dispersion. We propose two sampling methods, Size-Corrected 2LS (S-2LS) and Size- and Dispersion-Corrected 2LS (SD-2LS), which correct these biases and provide provably better softmax approximations with negligible to non-existent computational overhead. In-depth experiments on five large-scale datasets validate the improved sampling properties of our methods. We recommend their consistent use in place of standard 2LS in future work.
PaperID: 6915, Poster
Abstract: Machine learning models often suffer performance degradation under subpopulation shift, particularly when spurious correlations cause models to rely on shortcut features that fail to generalize across subgroups. A recent line of work mitigates this issue by using loss-based signals to identify informative samples, but these signals can become severely distorted under label noise: mislabeled samples may also incur large losses and contaminate subsequent reweighting or retraining. Despite its practical importance, this intersection remains largely underexplored. We propose POTER, a reweighting framework based on optimal transport that derives sample importance from the transport geometry between the training distribution and a reference distribution constructed from limited validation group annotations. By measuring alignment at the individual-sample level rather than relying on loss, POTER downweights mislabeled or strongly bias-aligned samples while assigning higher importance to samples better aligned with the reference distribution. In addition, POTER requires only a single ERM training stage, moving beyond the retraining paradigm common in recent work. Across standard benchmarks and noisy-label settings, POTER achieves state-of-the-art worst-group accuracy, including cases where label corruption is concentrated within minority subgroups.
Abstract: Transformers dominate video recognition. They split videos into tokens, and processing them has expensive superlinear computational cost. Yet videos are filled with redundancy, so we can question the need for this expense. We introduce LookWhen, a selector–extractor framework that factorizes video recognition into learning when, where, and what to compute. Our shallow selector gets a scaled-down video and quickly scores all tokens across space-time, while our deep extractor gets the top-K selected tokens to approximate full-video representations without actually processing all the tokens. A key challenge is defining effective supervision for selection and extraction. For selection pre-training, we introduce a score on representations that ranks tokens by uniqueness using a simple nearest-neighbor distance. For extraction pre-training, we distill both a video teacher and an image teacher, for which we normalize its frame-wise representations to learn what changes within videos. Through these strategies, our selector-extractor learns general and efficient representations for feature extraction or fine-tuning to a task. Through experiments on Kinetics-400, SSv2, Epic-Kitchens, Diving48, Jester, and Charades, we show that LookWhen achieves a better accuracy-computation trade-off than efficient models and upgraded baselines of similar size. LookWhen Pareto-dominates in accuracy-FLOPs on 9 of 12 cases (6 tasks X 2 settings) and roughly matches on 3. In accuracy-throughput, measuring time in practice, LookWhen is more efficient still at 6.7× faster than InternVideo2-B at equal accuracy.
PaperID: 6917, Poster
Abstract: Neural image codecs face a long-standing trade-off: distortion-optimized designs preserve pixel-wise fidelity but tend to produce blurred reconstructions at low bitrates, while generative designs close the perceptual gap with a learned prior, but with pixel fidelity saturating as bitrate increases. The latter typically requires training multiple models for various bitrate targets. This paper addresses these limitations by reframing the perceptual reconstruction problem. Starting from the quantization bound that holds for any modern multi-rate codec, we derive a transport bound that confines the perceptual reconstruction to a closed, rate-determined cell around the backbone's output. The cell is small relative to the full image manifold, casting unconstrained generative modeling as a bounded-support problem. We instantiate this as PerQ: an image codec that attaches a rate-conditioned flow-matching compensator supported on this cell to a frozen multi-rate distortion-optimized backbone. The experimental results across Kodak and CLIC2020-test datasets show that PerQ offers similar performance to generative neural image codecs (e.g., MS-ILLM) in perceptual metrics (LPIPS, FID) , while still maintaining competitive coding performance compared to distortion-optimized codecs (e.g., ELIC). Moreover, PerQ can produce compression results across a wide bitrate range and fast decoding from only a single checkpoint.
Abstract: Distributional causal inference requires estimating not only average treatment effects but also interventional outcome distributions, including quantiles, tail risks, and policy-dependent uncertainty. As a method for distributional causal inference, generative adversarial network (GAN)-based counterfactual methods are flexible tools for this task. However, these methods have several limitations. First, the objectives of certain techniques do not coincide with the statistical risk of the identifiable causal target, and therefore provide limited theoretical guarantees regarding estimable counterfactual distributions or optimality. Second, they tend to rely on unstable density-based methods, such as density ratio estimation. In this paper, we propose GANICE (GAN for Interventional Conditional Estimation) with several advantages: it (i) clarifies the conditional interventional distribution for each treatment--covariate state as the causal estimation target; (ii) estimates the conditional distribution such that its averaged Wasserstein risk is minimized; (iii) establishes minimax optimality. GANICE achieves these advantages through the introduction of the extended Wasserstein distance, the incorporation of a cellwise critic in its dual, and an optimality proof based on Besov space theory. Our experiments demonstrate that GANICE consistently outperforms existing methods.
Abstract: Understanding the internal “thinking” process of Large Language Models (LLMs) and the cause of hallucinations remains a key challenge. To this end, we introduce latent debate, a structured surrogate framework for interpreting model outputs on True/False prediction tasks through the lens of internal latent arguments and interactions amongst them. Unlike human debates, latent debate captures the hidden supporting and attacking signals that arise within a model during a single inference. We first present a model- and task-agnostic conceptual framework, and then instantiate it symbolically to approximate the thinking process of LLMs towards binary decisions. Empirical studies demonstrate that our latent debate is a faithful structured surrogate model that has highly consistent predictions with the original LLM, while providing a form of interpretability. We also demonstrate that our latent debate provides a strong baseline for hallucination detection. Specifically, we identify strong correlations between debate patterns and hallucinations, such as a high degree of disagreement in the middle layers of the latent debate surrogate is linked to a higher risk of hallucinations. Our findings suggest that latent debate shows potential to analyze internal signals in LLMs for binary decision settings.
Abstract: Recent flow matching (FM) methods improve the few-shot adaptation of vision-language models, by modeling cross-modal alignment as a continuous multi-step flow. In this paper, we argue that existing FM methods are inherently constrained by incompatible geometric priors on pre-trained cross-modal features, resulting in suboptimal adaptation performance. We first analyze these methods from a polar decomposition perspective (\ie, radial and angular sub-manifolds). Under this new geometric view, we identify three overlooked limitations in them: 1) Angular dynamics distortion: The radial-angular coupling induces non-uniform speed on the angular sub-manifold, leading to regression training difficulty and extra truncation errors. 2) Radial dynamics neglect: Feature normalization discards modality confidence, failing to distinguish out-of-distribution and in-distribution data, and abandoning crucial radial dynamics. 3) Context-agnostic unconditional flow: Dataset-specific information loss during pre-trained cross-modal feature extraction remains unrecovered. To resolve these issues, we propose warped product flow matching (WP-FM), a unified Riemannian framework that reformulates alignment on a warped product manifold. Within this framework, we derive direct product flow matching (DP-FM) by introducing a constant-warping metric, which yields a decoupled cylindrical manifold (\ie, direct product manifold). DP-FM enables independent radial evolution and constant-speed angular geodesic transport, effectively eliminating angular dynamics distortion while preserving radial consistency. Meanwhile, we incorporate classifier-free guidance by conditioning the flow on the pre-trained VLMs' hidden states to inject missing dataset-specific information. Extensive results across 11 benchmarks have demonstrated that DP-FM achieves a new state-of-the-art for multi-step few-shot adaptation.
PaperID: 6921, Poster
Abstract: Change Detection & Captioning (CDC) aims to jointly localize changed regions and describe their semantic content from bi-temporal remote sensing images. While recent approaches have demonstrated the effectiveness of unified modeling, they primarily rely on shared representations or jointly optimized architectures, leaving task-specific communication under heterogeneous supervision less explored. However, detection and captioning differ fundamentally in their supervision granularity, leading to heterogeneous representations with uneven reliability across tasks and spatial locations. In this paper, we reformulate CDC as a selective information routing problem between heterogeneous task branches, where the key challenge is to determine when and where information should be transferred. To this end, we propose CoDeRNet, a unified framework that preserves task-specific representations while enabling bidirectional exchange of complementary information. CoDeRNet introduces a Confidence-Decomposed Routing (CoDeR) mechanism, which decomposes the routing condition into receiver-side need for support and sender-side reliability cues to guide selective transfer. Experiments on LEVIR-MCI and WHU-CDC show that CoDeRNet achieves strong overall CDC performance, particularly improving caption generation while preserving competitive change localization. Extensive analyses further show that the benefits arise not merely from joint learning or generic cross-task communication, but from explicitly modeling selective cross-task information routing. These results highlight the importance of conditional information transfer for unified CDC under heterogeneous supervision.
PaperID: 6922, Poster
Abstract: Dense correspondence matching has historically been bounded by simplifying spatio-temporal priors, such as smooth motion and rigid geometry. While effective for classical tasks, these foundations collapse at the frontier of Vision-Language-guided Image Editing and Generation (VL-IEG), where transformations yield correspondences that are perceptually obvious yet physically discontinuous. To transcend spatio-temporal priors, we draw inspiration from human cognition: learning a generalized model of identity-preserving visual consistency rather than relying on physical constraints. To achieve this, we introduce FreeMatching, a unified framework that extracts identity-preserving knowledge from foundation models and diverse datasets. The model's capability is forged through a two-stage training paradigm: supervised pre-training on diverse annotated datasets, followed by weakly-supervised refinement on data without dense annotations. Experimentally, FreeMatching not only achieves performance competitive with SOTA methods on classical benchmarks but also establishes the first strong baseline for the challenging VL-IEG task. Furthermore, we demonstrate its utility as a quantitative metric for evaluating identity-preserving consistency in generative editing, aligning closely with human judgment.
Abstract: Block diffusion LLMs are an emerging paradigm for parallel language generation, but their KV caching makes memory access the dominant bottleneck in long-context inference. Sparse attention, which attends only to a small KV subset per query, can reduce this latency with minimal accuracy loss. In block diffusion, however, the B tokens of each block must share a single KV subset, and we show this per-block constraint degrades existing sparse KV estimators by up to 25% in recall. We address this challenge by exploiting a property that emerges from the block-diffusion training objective: it aligns the block-average query across denoising steps, so the All-[MASK] block at the first step already reveals the per-block KV subset for the entire trajectory. We exploit this in MAGE ([MASK]-Guided Sparse Attention), a training-free method that runs one exact attention pass at the first step and reuses its top-k index sets for all remaining steps within the block. Across three block-diffusion families on LongBench, MAGE matches Exact Attention at k=512 with near-lossless accuracy, achieves up to 6.82× end-to-end speedup at 128K context, and runs up to 3.35× and 2.28× faster than Quest and SparseD, designed for AR LLMs and fully bidirectional diffusion LLMs, respectively.
PaperID: 6924, Poster
Authors:
Zhitao FENG, Bo Ho, Huan Luo, xiaolong zou, Yuanyuan MiAbstract: Temporal sequences in working memory are hierarchically organized to support flexible behavior. Recent neural recordings suggest that this organization relies on a 1D-to-2D neural geometrical folding in working memory, whereby ordinal positions are re-encoded within an orthogonal scaffold, with two dimensions indexing global (chunk-level rank) and local ranks (within-chunk position), respectively. This findings implicate an important structural principle for memory, yet the computational mechanisms underlying this geometry remain unclear. Here, we study this question using a recurrent neural network with fast Hebbian plasticity (Hebb-RNN). When meta-trained on rank recalling tasks, the Hebb-RNN generalizes to novel item–rank bindings and hierarchical sequences, outperforming control models. Mechanistic analysis reveals that the network develops an orthogonal ordinal geometry consistent with empirical findings. This geometry transforms temporal inputs into abstract ordinal ranks that provide a memory scaffold, while Hebbian plasticity supports rapid conjunctive binding of item content to this structure. Furthermore, the pre-learned geometric scaffold significantly accelerates learning in novel tasks, supporting efficient transfer. By incorporating a predictive module, our model extends to real-world encodings and acquires hierarchical geometry without explicit chunking signals. Together, our results provide a biologically plausible framework for hierarchical sequence representation, offering new insights into structured memory organization for both neuroscience and machine learning.
PaperID: 6925, Poster
Abstract: Bayesian estimation has been used to explore optimal encoder-decoders for continuous quantities, but no such framework exists for discrete associative memory. We introduce precision-aware Hopfield retrieval, a framework that incorporates the encoder's Fisher information directly into a decoder of memory retrieval. Four results follow. First, optimal encoder-decoder information is shaped only by the prior, across both log and L^p norm loss functions. Second, the optimal information follows a general coupling law across log and L^p loss functions. Third, optimal precision-aware Hopfield retrieval uses a stimulus-dependent inverse temperature. It outperforms the encoder-blind Bayesian decoder on both retrieval loss and required coding resource. Last, the framework yields a closed-form expression for memory recall bias. It reproduces central-tendency and cognitive-load effects observed in cognitive psychology. These effects emerge only when the encoder and decoder are modeled jointly.
PaperID: 6926, Poster
Abstract: We study the problem of average reward Multi-Agent Reinforcement Learning (MARL) where agents interact within a network. Each agent in the network conducts local updates with information within its k-hop neighborhood and collaboratively maximizes the overall reward of the entire network. We first provide impossibility results in achieving convergence to the global optimal joint policies in our setting. These impossibility results highlight the challenges of decentralized learning with local information in multi-agent systems, namely the issues of multiagency, where each agent independently optimizes its policy, and partial observability, where agents have limited visibility into the global state. Given these challenges which indicate that decentralized policy optimization can generally be certified only up to stationarity, we provide finite-sample convergence guarantees for a decentralized actor-critic algorithm with linear function approximation that converges to an approximate stationary point with small gradient norm. To overcome such impossibility results, we further show that adding fixed-state entropy regularization reshapes the objective landscape which leads to even stronger global convergence guarantees for our proposed actor-critic algorithm.
PaperID: 6927, Poster
Authors: RUIWEN WANG, Thibaut Tachon, Chong Li, Raja Appuswamy
Abstract: Production training of large models alternates between model edits and hybrid-parallel strategy planning. Fixed-configuration planners solve each revision from scratch, while unvalidated reuse across edits risks returning a different plan than full re-search would. We introduce a dependency-keyed per-layer primitive interface for re-planning: closed-form layer primitives expose communication, memory, and time to the solver, and the same templates expose the input dependencies needed for reuse-equivalent re-planning. CONCORD attaches dependency sets and value signatures to primitive evaluations, enabling dependency-keyed re-planning to reuse only signature-valid primitive-evaluation work while preserving the same logical candidate set and feasibility predicate as cache-disabled full search. On production MoE strategy-quality runs—including a 4096-device MoE-438B deployment and a 32k-token long-context point—CONCORD's single-call plan achieves 1.14–1.24× training step-time speedup over an expert-tuned Megatron-style baseline (1.70× on the long-context point). Across a 9-fixture sweep at 256–4096 cards, CONCORD's warm re-planning achieves median 2.74× planner wall-clock speedup over cache-disabled full search across 54 in-dispatch cells. Full-search equivalence holds throughout: candidate and feasibility counts match in every cell, the top-1 plan matches in 54/54 cells, and stale-hit mismatches (cached value disagreeing with cold recomputation) are 0/162 over paired trials.
PaperID: 6928, Poster
Authors:
Jiayi Li, Sunny Chung, Don C Codipilly, George Marek, Varan Perananthan, Dennis L. Shung, Bradly StadieAbstract: In clinical decision-making, success depends on taking the right action to obtain the evidence that matters. One targeted clinical workup can identify a diagnosis that a battery of nonspecific tests missed. Large language models have shown considerable promise in clinical diagnosis, but most current approaches assume that all relevant patient information is available upfront, which does not reflect how diagnostic evidence is gathered in practice. Even when models gather evidence iteratively, challenges arise in determining which action to take, as they receive no feedback on which findings are diagnostically decisive. To solve this problem, we introduce TraceDx, a reinforcement learning framework for sequential clinical diagnosis. TraceDx is built on the idea that unstructured clinical records can be decomposed into atomic clinical findings, each scored by its diagnostic value. This signal rewards agents for discovering decisive evidence in addition to reaching the correct diagnosis. Consistent with this design, a case-level analysis shows that recovery of critical evidence is the only statistically significant predictor of diagnostic success. Direct case comparisons further show that TraceDx uncovers more diagnostically important evidence, and uses it to resolve the key diagnostic uncertainty. TraceDx-trained open-weight models outperform multiple frontier baselines on both MIMIC-CDM and a rare gastrointestinal disease dataset.
PaperID: 6929, Poster
Abstract: Knowledge tracing (KT) models are predominantly evaluated using aggregate metrics such as area under the curve (AUC) and accuracy. However, these global scores obscure where the remaining errors originate and fail to indicate whether a benchmark is approaching saturation. While estimating a global theoretical performance limit is challenging in realistic KT settings, it is possible to quantify local predictability. To address this, we propose an information-theoretic evaluation framework for KT benchmark diagnosis. We utilize Context Tree Weighting (CTW) as an operational causal anchor to estimate the Local Irreducible Uncertainty (LIU) of each student interaction. By projecting predictions onto this shared uncertainty coordinate, we evaluate model performance gains across distinct entropy bands rather than only at the global level.Comprehensive evaluations on NIPS Task 3/4 and Algebra 2005 reveal that model improvements are highly non-uniform. While the largest gains achieved by modern KT models consistently occur in high-entropy regions—suggesting that datasets are not yet exhausted—our framework also identifies instances where models capture inherently random noise. By surfacing these local modeling failures alongside genuine gains, this approach provides a rigorous diagnostic tool to pinpoint both the potential and the limitations of current KT benchmark and model diagnosis.
PaperID: 6930, Poster
Authors: Meng Wang, Jinshuo Liu, Weiran Pian, Feiyang Wen, Xinya Liu, Juan Deng, Jeff Pan
Abstract: Automatically migrating serial C programs to parallel CUDA is essential for high-performance computing, yet remains challenging due to the partial observability of data dependencies: critical memory access patterns cannot be reliably inferred from source code alone. Existing static analysis methods often fail on irregular and pointer-intensive loops, while LLM-based approaches cannot infer memory-access-dependent optimization strategies from source text alone. To address these limitations, we propose STParOpt, which formulates parallelism and optimization inference as an uncertainty-aware learning problem under partial observability, where ambiguity in static analysis is explicitly quantified and resolved using execution signals. STParOpt augments static program graphs with stride-level execution features via cross-attention, and employs an entropy-guided fusion mechanism that up-weights dynamic evidence when the static branch's predictive entropy over parallelizability is high. Experiments on standard HPC benchmarks demonstrate that STParOpt significantly outperforms both compiler-based and LLM-based baselines on parallelism detection and CUDA code generation. Notably, STParOpt achieves 96.4% parallelism detection accuracy and 86.1% end-to-end generation correctness, surpassing the strongest baseline by +5.0 pp and +11.4 pp respectively.
PaperID: 6931, Poster
Abstract: Gradient-based machine unlearning methods typically ``unlearn'' by optimizing forget and retain losses, but similar losses do not necessarily imply similar output-space behavior. Prior work has shown that high forget loss alone does not certify that unlearning has produced retraining-like output behavior. We therefore ask: when a gradient-based unlearner differs from the model that would have been learned on only the retain data, what structure does the resulting output-space gap have? We derive closed-form expressions for this gap in linear regression, with an extension to nonlinear models via a generalized Gauss-Newton and local linearity approximations. Our analysis puts various gradient-based methods under a unified framework parametrized by geometry, reweighting, and projection. Two leverage matrices emerge that capture the forget-retain entanglement under the retrained model geometry: forget self-leverage and retain cross-leverage. We show both theoretically and empirically that these two leverage matrices drive the impact of the forget residual on the forget and retain outputs, and existing methods account for this geometry only partially. Our framework yields precomputable leverage diagnostics for auditing forget sets to identify what type of gradient-based unlearning is likely to be successful.
PaperID: 6932, Poster
Abstract: Test-time scaling boosts the reasoning performance of large language models (LLMs) on challenging tasks but incurs substantial computational overhead. The existing training-free approaches adopt token-level confidence as a proxy and terminate the trajectories once the confidence falls below a threshold. However, these methods view all low-confidence tokens uniformly, overlooking the fact that some of them actually reflect benign generations. We introduce Calibrated Confidence (CalConf), a training-free proxy for test-time scaling in LLM reasoning. The design rests on a simple empirical observation: flawed trajectories tend to be long. CalConf therefore learns a length-calibrated token-weight map and uses it to selectively terminate degrading trajectories during generation. We theoretically prove that CalConf attains the desired early-stop rate in practice without per model tuning. Across five benchmarks (on mathematical reasoning, scientific reasoning, and code generation tasks) and three open-source LLMs, CalConf consistently matches or exceeds offline majority-vote accuracy while reducing generated tokens by 20–44%. Token-level analyses trace these gains directly to CalConf’s ability to isolate harmful patterns that uniform-confidence baselines cannot distinguish.
PaperID: 6933, Poster
Abstract: Combining two pretrained models, whether by aligning a vision encoder with a language model or by merging two fine-tuned checkpoints, is a costly bottleneck: compatibility is typically known only after joint training. We propose COMPASS (COMPatibility ASsessment Score), a theoretically grounded compatibility diagnostic that estimates whether a model pair will be compatible prior to training or merging. Building on the observation that representational convergence is most robust at the level of local neighborhoods, COMPASS decomposes model compatibility into two cheap-to-compute axes: a pair-specific ease, defined as the k-NN overlap between the two encoders, and a paradigm-level ceiling, capturing the maximum overlap their training paradigms allow. Together, these two axes provide a ranking of candidate model pairs and prescribe how to act on each pair: (i) prioritize pairs that already agree and whose paradigms are compatible, (ii) invest additional compute on pairs with compatible paradigms but low initial agreement, (iii) switch the pre-training paradigm when it cannot support sufficient agreement, or (iv) deprioritize pairs that score poorly on both axes. COMPASS is applicable to both cross-modal alignment and task-arithmetic merging, where the encoder pair is replaced by two fine-tuned checkpoints. Across 72 cross-modal alignment experiments, and 392 model merges, our compatibility estimates strongly correlate with downstream success, with Pearson correlations of 0.90 for image-to-text R@1 on Flickr30k, 0.91 for image-to-text R@1 on COCO, and 0.70 for mean retention across 8 image-classification merging tasks.
PaperID: 6934, Poster
Abstract: Video Moment Retrieval (VMR) aims to localize a temporal segment in a video based on a natural language query. While Multi-modal Large Language Models (MLLMs) are increasingly adopted for VMR due to their strong video understanding capability, they still face several limitations: (1) imprecise perception of event boundaries, (2) rigid dependency on precise object identification, and (3) inability to handle temporarily interrupted events. To address these, we propose three visual prompts: Event Segmentation (ES) prompt, Object Detection (OD) prompt, and Keyframe Marking (KM) prompt. Each prompt improves performance when applied individually, yet direct fusion degrades results because mixing prompts overcomplicates video content and inaccurate labels introduce noise. Although certain visual prompts may mislead the model, at least one of the three is likely to yield a positive effect. Even in the worst case where all three perform poorly, the prediction can fall back to the baseline. Motivated by this, we introduce a selection mechanism based on Yes/No Probability-Difference score to choose the most reliable prediction among the results produced by the three visual prompts and the baseline. We evaluate our method in a training-free setting across three datasets and four models, and obtain state-of-the-art performance.
PaperID: 6935, Poster
Authors: Chen Wang
Abstract: A multi-tool plan is correct only when its call set matches the gold API set exactly. Standard tool-use pipelines optimize first-hop relevance, but downstream verifiers require the necessary API set, creating a relevance--necessity gap. This study formalizes the gap: under correlated API indicators, symmetric decomposable selectors based only on per-API marginals can be dominated by a Bayes-optimal set selector, while whole-set exact-match supervision is a proper scoring rule for the optimal selector restricted to a candidate menu. We propose an offline refinement framework for frozen upstream traces. Under a leakage-free ToolBench G2/G3 protocol, TC-MASS improves verifier-consistent API-set recovery without additional LLM generation after the upstream trace exists. Ablations and cross-distribution probes support the central claim: necessity-aware set-level supervision, not relevance ranking alone, is the right objective for fixed-pool API-set recovery.
PaperID: 6936, Poster
Abstract: Vision-language models (VLMs) have achieved remarkable progress in 2D visual understanding tasks, yet the passive view-centric perception paradigm limits their capacity for 3D spatial reasoning. Such reasoning often requires actively acquiring spatial evidence beyond the currently visible views, such as camera poses and cross-view spatial relationships, to infer 3D relationships that are not explicitly represented in 2D images. However, existing methods either lack active exploration or merely provide additional views during inference, without explicitly acquiring such spatial evidence beyond 2D images. In this paper, we propose View-to-World (V2W), an active exploration framework that enables VLMs to actively explore 3D worlds and lifts VLMs from passive view-centric perception to active world-centric spatial reasoning. Given a spatial reasoning task with multi-view images, V2W first reconstructs an explicit 3D world with geometry, camera poses, and language-grounded semantics, and exposes it to VLMs through visual and linguistic interfaces. Through multi-turn interaction with these interfaces, VLMs iteratively acquire spatial evidence from the constructed 3D world, enabling active world-centric spatial reasoning. V2W can be integrated with existing VLMs in a training-free manner and can be further enhanced by optimizing the exploration policy via reinforcement learning. Experiments on spatial mental modeling benchmarks demonstrate that V2W substantially improves VLM baselines in the training-free setting and achieves state-of-the-art performance with the learned exploration policy, highlighting the importance of active exploration for spatial intelligence. The code will be available upon acceptance.
PaperID: 6937, Poster
Abstract: Despite recent advances in camera-based online HD map construction, these systems remain susceptible to sensor failures due to their reliance on onboard cameras. In contrast, satellite images remain available independently of real-time sensing conditions as they are cached offline. Existing methods utilizing satellite images, however, primarily leverage them under nominal conditions, without explicitly addressing as redundancy when sensor degrades. To address this gap, we propose ReSMap, a redundancy-aware camera-satellite fusion framework for robust online HD map construction. Our framework explicitly leverages satellite images as a complementary source to improve robustness against camera failures. A Predict-then-Fuse (PTF) module fuses camera and satellite BEV features by using each branch's own prediction confidence as a per-pixel reliability weight, with no learned router. An Object-centric cooperative Multi-modal Query (OCMQ) decoder then propagates this isolation to the query level: each modality cross-attends only to its own BEV, and per-modality updates are merged through a learned per-token gate. Our model achieves state-of-the-art performance on nuScenes across all splits and range settings. More importantly, it also sets the new state-of-the-art under camera-failure scenarios, demonstrating that satellite images serve not only as a strong prior but as practical redundancy for robust online HD mapping.
PaperID: 6938, Poster
Abstract: Embodied navigation and reasoning require an agent to actively acquire, organize, and verify task-relevant evidence in previously unseen environments. Recent methods typically rely on semantic maps, scene graphs, vision-language priors, or structured memory, but often underuse persistent fine-grained 3D evidence and entangle query-related ambiguity with environment-side reliability. This coupling can make evidence acquisition inefficient and final decisions vulnerable to weak or transient observations. We propose Decoupled Complementary Fields on 3D Gaussian Maps, an embodied navigation framework that maintains persistent volumetric evidence, separates query-related ambiguity from environment-side readiness, and grounds decisions through progressive verification. First, we construct a query-conditioned object-centric Gaussian navigation state built on online 3D Gaussian Splatting, which preserves fine-grained 3D evidence across viewpoints while providing an executable spatial interface for region reasoning and navigation. We further formulate Decoupled Complementary Fields, consisting of a Query-Ambiguity Field, which captures where task-relevant evidence remains unresolved, and an Environment-Readiness Field, which estimates where the environment is structurally reliable for motion and inspection. Finally, we couple these fields in a hierarchical navigation-and-verification policy, which selects informative frontiers, instantiates executable local goals, and verifies task-relevant candidates using persistent volumetric Gaussian evidence before stopping or answer handoff. Experiments on A-EQA and GOAT-Bench show consistent gains in active evidence acquisition, multi-modal lifelong navigation, and reliable final decision making.
PaperID: 6939, Poster
Authors:
Satwik Sunnam, Raghav Magazine, Vatsalya Singh, Lavanya Kotha, Xingjian Li, Min XuAbstract: Large language models consistently fail at elementary counting tasks despite strong performance on complex reasoning benchmarks. Recent work has characterized counting failures behaviorally or identified circuits through which counting succeeds, but a mechanistic account of where and why counting breaks down remains absent. In this work, we provide such an account by tracing the count signal from embedding to output and identifying the geometric bottleneck that prevents correct readout for single digit counting problems. Using a dual-metric framework combining Signal-to-Noise Ratio (SNR) with linear probing, we show that counting information is computed correctly and persists in a linearly decodable subspace through the final layer. The failure in counting is not representational but geometric i.e the digit-token columns of the unembedding matrix W_U are orthogonal to the count subspace in the residual stream, rendering correctly computed counts unreadable at the output. We term this as the Unembedding Bottleneck and to causally validate it, we apply a Probe-Informed LoRA correction targeting W_U, which restores single-digit counting accuracy by up to 80-95% across six models while updating only ~0.001% of parameters and preserving general capabilities. Our findings reveal a broader class of failure in LLMs where information is internally present but geometrically inaccessible to the output projection.
Abstract: Spatially consistent long-horizon video generation aims to maintain temporal and spatial consistency along predefined camera trajectories. Existing methods mostly entangle memory modeling with video generation, leading to inconsistent content during scene revisits and diminished generative capacity when exploring novel regions, even trained on extensive annotated data. To address these limitations, we propose a decoupled framework that separates memory conditioning from generation. Our approach significantly reduces training costs while simultaneously enhancing spatial consistency and preserving the generative capacity for novel scene exploration. Specifically, we employ a lightweight, independent memory branch to learn precise spatial consistency from historical observation. We first introduce a hybrid memory representation to capture complementary temporal and spatial cues from generated frames, then leverage a per-frame cross-attention mechanism to ensure each frame is conditioned exclusively on the most spatially relevant historical information, which is injected into the generative model to ensure spatial consistency. When generating new scenes, a camera-aware gating mechanism is proposed to mediate the interaction between memory and generation modules, enabling memory conditioning only when meaningful historical references exist. Compared with the existing method, our method is highly data-efficient, yet the experiments demonstrate that our approach achieves state-of-the-art performance in terms of both visual quality and spatial consistency.
Abstract: Reward machines (RMs) inform reinforcement learning agents about the reward structure of the environment, enabling support for non-Markovian tasks and improving sample efficiency. However, learning with RMs is ill-suited for long-horizon problems in which subtasks can be completed in any order. In such cases, the amount of information to learn increases exponentially with the number of unordered subtasks. We address this limitation by introducing three generalisations of RMs: (1) Numeric RMs allow users to express complex tasks in a compact form. (2) In Agenda RMs, states are associated with an agenda that tracks the remaining subtasks to complete. (3) Coupled RMs have coupled states associated with each subtask in the agenda. Furthermore, we introduce a new task-decomposition Q-learning-based algorithm that leverages coupled RMs and preserves global optimality guarantees: QCoRM. Our experiments across four domains---featuring both discrete and continuous action and state spaces---show that QCoRM scales better than baseline algorithms for long-horizon problems with unordered subtasks.
Abstract: The successful training of neural networks hinges on the use of first order optimization methods, yet the theoretical characterization of these methods remains incomplete. This is especially true in settings with mild overparameterization. In this work, we study the gradient flow dynamics of two-layer ReLU networks from small initialization with orthogonal training data. We prove the limiting flow converges to a saddle-to-saddle jump process as the initialization scale tends to zero, revealing an incremental learning phenomenon in which a new neuron activates at each saddle. This analysis recovers the known result of Dana et al. (2025) that the network interpolates the training data with high probability as soon as m \gtrsim \log(n), where m is the network width and n is the number of training samples. This incremental process characterization also allows us to derive a novel implicit bias result: the learned interpolator has an \ell_2-norm scaling as \sqrtn, which is within a constant factor of the minimal-norm interpolator. More broadly, our work provides the first rigorous proof of an incremental learning process for ReLU networks, whilst suggesting mildly overparameterized networks can converge to interpolating solutions whose complexity is of the same order as that of the optimal interpolator.
PaperID: 6943, Poster
Abstract: Local intrinsic dimension (LID) quantifies the degrees of freedom of the data manifold around a point and is widely used to detect memorization, diagnose distribution shift, and characterize learned representations. Existing generative-model-based LID estimators are model-specific: they rely on a score, a log-density Hessian, or an exact log-likelihood, and therefore do not apply to ODE-based generators (flow matching, rectified flows, stochastic interpolants, and CNFs) that expose only a velocity field. We propose JaSpec, a model-agnostic LID estimator that counts the singular values of the velocity-field Jacobian J_t = \partial f_\theta / \partial x that fall below a threshold \tau. JaSpec applies uniformly to score-based diffusion (via the probability-flow ODE) and to flow-based generators, and recovers Hessian-based estimators on probability-flow diffusion as a verification special case. We prove a direct dynamical identifiability theorem under a Jacobian-normal-dominance condition stated entirely in terms of J_t and verified by an empirical spectral-gap diagnostic. On VP-SDE diffusion and OT-CFM flow-matching benchmarks with known LID, JaSpec attains MAE = 0 on most settings up to D=3072, where no prior model-based estimator reaches zero error, and improves over the strongest model-based baseline by an order of magnitude on the hardest 3072-dimensional nonlinear mixture. We provide both an exact \mathcalO(D^3) path and a linear-cost stochastic Lanczos quadrature path. JaSpec is, to our knowledge, the first LID estimator applicable to flow matching, rectified flows, and generic differentiable ODE generators.
Abstract: Safety is a key problem in reinforcement learning, especially during exploration. While safety filters are promising to address this issue, they are ill-suited for high-dimensional systems with unknown dynamics. We propose Dyna-style Safety Augmented Reinforcement Learning (Dyna-SAuR), a novel algorithm that learns a control policy as well as a scalable safety filter requiring minimal domain knowledge by leveraging a learned uncertainty-aware dynamics model. A novel filter learning formulation enforces to avoid failures and areas of high model uncertainty. As joint deployment of the control policy and filter allow to safely collect data and improve the dynamics model, certain areas grow and filter conservatism decreases over time. Dyna-SAuR reduces failures compared to state-of-the-art methods by two orders of magnitude on goal-reaching CartPole and MuJoCo Walker.
PaperID: 6945, Poster
Authors:
Kaiwen Sheng, Yiqi Jiang, Yuyang Song, Gaurav Tyagi, Haotian Ye, Fangzhao Zhang, Ruiming Du, Pengxiao Wang, Stefano Ermon, Karl Deisseroth, Anish MitraAbstract: Neural responses to the same stimulus can vary dramatically with fluctuations in the brain's internal state and the organism's behavioral state. Yet studying how brain state shapes neural responses is difficult: limited experimental recordings sparsely sample this joint space, leaving many biologically valid state-stimulus pairings unobserved. Existing methods either capture state-dependent modulation with restricted model assumptions or synthesize neural activity without explicitly controlling internal brain state. We introduce , a conditional diffusion framework that models the spatiotemporal distribution of future neural activity and behavior given external stimuli and recent brain states. Motivated by two gaps in finite recordings, missing state-stimulus combinations and limited repeats of observed conditions, BrainStateDiff introduces two state-aware sampling strategies: , which recombines observed states and stimuli to synthesize plausible neural responses from unobserved pairings, and , which resamples from observed pairings to capture local trial-to-trial variability. We evaluate BrainStateDiff on our newly collected wide-field calcium imaging data in mouse V1, electrophysiology data in monkey V4, and two-photon calcium imaging data in mouse striatum. Our results show that BrainStateDiff generates biologically meaningful samples that preserve the state-dependent structure of real recordings and expand neural response space coverage beyond the original recordings. BrainStateDiff establishes state-conditioned generative modeling as a practical framework for studying how internal brain state shapes neural responses beyond the state-stimulus combinations directly observed in the experiments.
Abstract: Many modern Language Model (LM) pipelines return an averaged model, such as an exponential moving average of the training iterates, rather than the final iterate itself. This raises a fundamental question: given that we will return an iterate average, how should we change training to improve the performance of this average? We study this question by formulating optimizer design for the iterate-average estimator as an optimal-control problem. In a continuous-time stochastic quadratic model, we solve for the control strategy that minimizes the error of the returned average subject to a penalty on the size of the intervention. A practical approximation to this controller yields PACE, a lightweight wrapper around AdamW that pulls the live weights toward their exponential moving average with a clipped, per-coordinate control strength. We prove that a stylized version of PACE converges at the standard stochastic convex optimization rate, up to a factor depending on the averaging rule, while in the quadratic setting it can strictly improve the limiting squared error of the iterate-average estimator and can do so by an arbitrarily large factor on some instances. Empirically, our initial results suggest that PACE improves over AdamW and EMA-evaluated AdamW in supervised fine-tuning of 1-2B parameter LMs and in GPT-2 pretraining on FineWeb, while remaining competitive with learning-rate decay and Schedule-Free baselines.
PaperID: 6947, Poster
Abstract: Many existing defenses mainly strengthen refusal behavior to improve the safety of LLMs. Recent interpretability work shows that, in safety-aligned LLMs, harmfulness belief and refusal behavior are represented separately in the hidden states at the last token of the user instruction (t_\mathrminst) and the response-start token (t_\mathrmpost), respectively. Based on it, we find that harmfulness belief is not reliably preserved when it is carried from the t_\mathrminst to the t_\mathrmpost. As a result, even when the LLM internally recognizes a request as harmful, it may still produce a compliant response. To address this, we propose Bridging Harmfulness and Refusal (BHR), a training framework that bridges harmfulness belief and refusal behavior. Specifically, BHR trains the adapter so that harmfulness information is carried to the hidden state at the t_\mathrmpost, thereby establishing a second pathway from harmfulness belief to refusal behavior. To keep this signal reliable during fine-tuning, we introduce two additional losses. A belief loss helps the LLM maintain an accurate harmfulness judgment at the t_\mathrminst. A belief-gated refusal loss uses the LLM's harmfulness belief at t_\mathrminst to regulate refusal learning, making attacks that target refusal features less effective while reducing the risk of over-refusal on benign inputs. Experiments across multiple LLMs and safety benchmarks show that BHR substantially improves robustness against both white-box and black-box jailbreak attacks, reduces the risk of over-refusal on benign prompts, and preserves general capabilities.
Abstract: Learning high-quality latent actions from large-scale unlabeled videos, coupled with limited real-world interaction data for training an action decoder, has emerged as a promising paradigm for scalable latent policy learning. However, existing approaches typically rely on behavior cloning, which tends to collapse inherently multimodal action distributions into unimodal ones, thereby degrading the pretrained latent action structure. While flow matching provides a potential alternative, directly applying it leads to a misalignment between latent actions and physical actions during action decoder training, due to the stochastic nature of the learned policy. To address these, we propose Latent Action Flow Policy (LAFP), which leverages flow matching for latent policy learning and introduces an inference-time interpolation mechanism to mitigate stochasticity-induced misalignment. Experimental results demonstrate that LAFP consistently outperforms prior methods on downstream imitation learning tasks, achieving up to 10–15% improvement in success rate while incurring less than 1× additional inference overhead.
PaperID: 6949, Poster
Abstract: Dense large language models (LLMs) exhibit implicit sparsity at inference time: for a given token, only a small subset of FFN neurons contributes dominantly. This observation motivates , which partitions FFN neurons into experts and employs a lightweight router to sparsely activate experts for efficiency. However, existing conversion methods often incur non-trivial quality degradation under a fixed activation budget, largely due to suboptimal expert construction and unstable router learning in fine-grained settings. Motivated by this, we present , which improves Dense-to-MoE conversion with a more functionally coherent expert grouping method and a more stable and effective pairwise-ranked router training strategy, leading to markedly better quality under the same activation budget. Moreover, we further introduce an adaptive layer-wise sparsity allocation strategy to better utilize a global activation budget, and apply STE-based joint training to enhance model--router compatibility under hard routing. Experiments across various benchmarks show that ExpertNavigator consistently outperforms prior Dense-to-MoE approaches under different sparsity budgets, retaining up to of the dense model's average score while activating only 60% of FFN parameters, substantially narrowing the gap to dense models and delivering practical efficiency gains.
Abstract: The success of vision transformers—especially for generative modeling—is limited by the quadratic cost and weak spatial inductive bias of self-attention. We propose PDE-SSM, a spatial state-space block that replaces attention with a learnable convection–diffusion–reaction partial differential equation. This operator encodes a strong spatial prior by modeling information flow via physically grounded dynamics rather than all-to-all token interactions. Solving the PDE in the Fourier domain yields global coupling with near-linear complexity, delivering a principled and scalable alternative to attention. We integrate PDE-SSM into a flow-matching generative model to obtain the PDE-based Diffusion Transformer PDE-SSM-DiT. Empirically, PDE-SSM-DiT matches or exceeds the performance of state-of-the-art Diffusion Transformers while substantially reducing compute. Our results show that, analogous to 1D settings where SSMs supplant attention, multi-dimensional PDE operators provide an efficient, inductive-bias-rich foundation for next-generation vision models.
Abstract: Autoregressive language models (ARMs) suffer from the reversal curse: after learning "A is B," they often fail on the reverse query "B is A." Masked diffusion language models (MDMs) exhibit this failure in a much weaker form, but the underlying reason has remained unclear. A common explanation attributes this mitigation to their any-order masked training objective. However, observing "[\textnormalM] is B" during training teaches recovery of A from B in one positional configuration, and does not by itself explain why the learned evidence should transfer to the reverse prompt "B is [\textnormalM]." We provide a theoretical analysis showing that this transfer arises from a parameter-level coupling between forward and reverse positional conditionals: shared Transformer parameters store token-pair evidence, while relative positional encodings route attention through queries and keys without changing the value-side evidence being retrieved. In a one-layer MDM, we prove that forward masked training strengthens evidence that is reusable in reverse queries, induces correlated forward--reverse attention routes, and yields a positively aligned shared-storage gradient component that decreases the reverse loss to first order. Controlled one-layer experiments and large-scale LLaDA/Dream experiments verify these signatures and show that they translate into improved reverse prediction.
PaperID: 6952, Poster
Authors: Yibo Zhou, Yirong Xiang
Abstract: Relational out-of-distribution errors are easy to describe and hard to evaluate cleanly. An image may be natural, a caption may be fluent, and all mentioned objects may be common, yet the caption may describe the wrong relation in the image. Many evaluations intended to test this behavior can be solved without checking the pairing at all, because negatives differ in caption style, template, length, object frequency, or image statistics. We propose counterfactual pairing cycles, a blocked evaluation primitive that reuses the same two images and the same two captions on both sides of the comparison, changing only which image is paired with which caption. This makes image-only, text-only, and additive component shortcuts cancel exactly in finite samples. The same contrast also gives a training loss for relational outlier exposure and a diagnostic for bilinear vision-language scores. Across synthetic controls, GQA, and COCO, cycles separate coarse image-caption matching from stricter role-binding failures. In a shortcut-contaminated COCO outlier-exposure stress test, a model selected by a perfect shortcut validation AUROC transfers poorly to exact relational cycles, while cycle training improves CycleAcc from (0.576) to (0.911). The main lesson is simple: to test whether a model understands a pairing, the benchmark should keep the parts fixed, change the pairing, and ask only about coupling.
PaperID: 6953, Poster
Abstract: Dynamic Sparse Training (DST) is an effective paradigm for sparse to sparse training of neural network connectivity under a fixed parameter budget. Recent advances in network-science for AI introduced epitopological learning methods, such as Cannistraci-Hebb Training (CHT), demonstrating that network automata applied to the mere network topology enables gradient-free predictions of the sparse connectivity evolution that can trigger significant increase in task performance. However, these methods were currently developed for standard MLP-like fully-connected architectures and need significant rethinking to be extended to convolutional neural networks (CNNs) due to element-wise weight sharing and spatially repeated interactions. In this work, we bridge this gap by introducing epitopological learning DST for CNNs through channel-wise network modeling to address the prohibitive complexity of element-wise modeling, which treats each kernel-channel as a single network node. Results show orders-of-magnitude improvements in time complexity of link-regrowth when applied at channel-level with respect to element-level. We also enhance the channel-wise formulation with lightweight contextual modulation to improve expressivity. Extensive experiments on CIFAR-10, CIFAR-100, and Tiny-ImageNet with ResNet and VGG architectures demonstrate that our method achieves performance comparable to dense training using only 30% of the parameters, which significantly reduces computational cost. Furthermore, our approach improves robustness under noisy inputs. These results highlight the effectiveness of topology-aware sparse training and the benefit of decoupling graph construction from sparse update rules in convolutional networks.
PaperID: 6954, Poster
Abstract: While offline reinforcement learning (RL) enables policy optimization from static datasets without costly online interaction, it remains bottlenecked by the risk of executing out-of-distribution (OOD) actions. Recent generative-model-based approaches mitigate this by performing RL within a constrained latent space, but they often sacrifice policy expressiveness or suffer from vanishing gradients. In this work, we identify that entropy regularization is essential to prevent mode collapse in latent-space RL, though applying it without introducing numerical instability or compromising expressiveness remains a significant open challenge. We introduce LASER, a novel offline RL algorithm that applies latent-space adjoint matching to achieve entropy-regularized RL with expressive flow policies. By design, our approach satisfies dataset support constraints and prevents mode collapse. Through comprehensive experiments on 40 challenging OGBench tasks with varying dataset qualities, we show that LASER achieves state-of-the-art performance. Notably, LASER uses a single, constant set of hyperparameters to outperform baselines even with their hand-tuned hyperparameters, highlighting its robust applicability.
PaperID: 6955, Poster
Authors: Zeping Zhang, Zhuoya Zhao, Samy Metari, Robert Laganiere
Abstract: Current 4D occupancy world models use heavy cross-attention or deep recurrence to disentangle ego-motion from scene dynamics. This requires significant computational capacity. However, this effort is largely redundant since ego-motion is inherently known from odometry. We present FreeOcc, a lightweight continuous-time world model that introduces explicit SE(2) geometric decoupling into a conditional flow-matching framework. By mathematically factoring out self-motion, we create a purely dynamic feature space. This space supports attention-free temporal fusion and allows physics-constrained ODEs to operate strictly on residual agent kinematics. To maximize deployment efficiency, FreeOcc introduces an uncertainty-aware blending mechanism. This routes well-observed voxels to a single-pass decoder, reserving the ODE solver exclusively for ambiguous, high-uncertainty regions. Requiring only 28.3M parameters—roughly 40% of comparable models—our optimal ensemble mode (FreeOcc-E) runs at 27 FPS. On nuScenes, FreeOcc achieves state-of-the-art forecasting performance with a 0.47 m Near Field Chamfer Distance (NFCD) and 32.1% mIoU. Furthermore, the model transfers zero-shot to SemanticKITTI. It also proves highly robust, retaining 92.8% of its ideal mIoU under significant sensor packet loss and timestamp jitter.
PaperID: 6956, Poster
Abstract: Emerging super-nodes tightly couple multiple servers with symmetric high-bandwidth fabrics, substantially weakening the legacy bandwidth cliff between nodes. This makes cross-boundary TP/EP a viable part of the strategy space, shifting the optimal strategy basin away from the legacy practice of keeping TP/EP within a node. Yet adapting the parallel strategy to new platforms still requires expensive profiling and brittle manual tuning. We propose Symbolic Barrier-aware Planner (SBP), a profiling-free planner that selects DP/MP/PP factorization and activation checkpoint strength by computing a barrier-aware symbolic score from model shapes, collective semantics, and hardware specifications. On legacy clusters, SBP recovers measured-best configurations; on super-nodes, it captures the regime shift and achieves 1.40× step-time speedup and a 10.62 percentage-point MFU improvement over the best node-local Megatron-style baseline on 128-die training. SBP explores the strategy space in seconds on a single CPU, without accelerator profiling for strategy selection.
PaperID: 6957, Poster
Abstract: Recent theoretical work has shown that nonlinear solvable models exhibit scaling laws in the feature-learning regime. However, these results largely rely on the assumption of isotropic inputs. Understanding how these laws extend to anisotropic data remains a central open problem. In this work, we address this question by analyzing the learning dynamics of two-layer neural networks with quadratic (second Hermite polynomial) activations under anisotropic Gaussian inputs. We provide a sharp characterization of stochastic gradient descent (SGD) across both continuous-time dynamics and finite-sample (online) discretizations, explicitly quantifying how the covariance spectrum influences the scaling exponent and sample complexity. Furthermore, we provide evidence that normalization techniques strictly improve sample efficiency in anisotropic settings. Experiments on two-layer networks with general activations support our theoretical predictions, suggesting that these insights extend well beyond the quadratic model.
PaperID: 6958, Poster
Abstract: Model merging enables the integration of multiple expert models with different capabilities into a unified model. In practical deployments, new expert models are often continuously updated and arriving. Meanwhile, due to the training and communication overhead of large language models (LLMs) and multimodal large language models (MLLMs), Low-Rank Adaptation (LoRA) is widely used to adapt large models to specific domains due to its efficiency. This scenario brings a new challenge: continual LoRA merging of deployed MLLMs to load new task capabilities. Naive merging faces catastrophic forgetting and parameter conflicts. We first observe the distinct features of the LoRA direction and layers in MLLM LoRA tuning. Based on this, we propose our Spectral Rank Calibration Merging (SRC-Merging), a data-free and train-free continual LoRA merging method tailored for MLLMs. SRC-Merging regards the deployed LoRA as a compressed spectral memory and performs calibrated rank-level merging for each incoming adapter. It aims to preserve accumulated multimodal knowledge while avoiding over-suppression of new vision-language task updates during continual merging. Extensive experiments on multiple multimodal tasks have demonstrated that our method significantly outperforms state-of-the-art merging methods.
PaperID: 6959, Poster
Authors:
Zhihong Sun, Qi Fan, Pan Liu, Yang Yi, Chen YeAbstract: Structured generation remains a critical bottleneck for Multimodal Large Language Models (MLLMs). Current post-training paradigms struggle to resolve this: Supervised Fine-Tuning (SFT) hits a rigid ceiling, while scalar-reward Reinforcement Learning (e.g., DPO, GRPO, SimPO) suffers from a policy-level signal gap. Compressing long routing trajectories into sparse signals induces a short-completion bias and traps RL near the SFT baseline, whereas vanilla self-distillation actively regresses. To systematically quantify this joint visual-structural-relational reasoning, we introduce OracleGraph, a 10,000-page 2.5D historical document benchmark featuring a 12-dimensional diagnostic VQA suite and a structured-hallucination taxonomy. To overcome these optimization pathologies, we propose PRISM, a Policy-Reward Integrated Self-distillation framework. PRISM leverages multimodal information asymmetry via a Hindsight Teacher guided by a Diagnostic Graph Report (DGR). By synergizing Softmax Reward-Weighted Aggregation (preserving Graph-F1 rankings) with a Length-Normalized Auxiliary Preference Loss, PRISM explicitly neutralizes catastrophic truncation, yielding a +3.81pp Graph Reward gain and a +2.3pp schema-parsing improvement over SDPO (p < 0.001). While PRISM's nominal +1.02pp mean gain over SFT is not statistically significant (p = 0.12), it drastically stabilizes optimization, achieving 2.4× lower cross-seed variance. Crucially, on a cross-book out-of-distribution (OOD) probe, PRISM's stability advantage amplifies to a 14× variance reduction over SFT, confirming robust structural stabilization. Code and datasets are available at https://anonymous.4open.science/r/PRISM2026.
PaperID: 6960, Poster
Abstract: Extending 3D Gaussian Splatting (3DGS) to 4D physical simulation remains challenging. Based on the Material Point Method (MPM), existing methods either rely on manual parameter tuning or distill dynamics from video diffusion models, limiting the generalization and optimization efficiency. Recent attempts using LLMs/VLMs suffer from a text/image-to-3D perceptual gap, yielding unstable physics behavior. In addition, they often ignore the surface structure of 3DGS, leading to implausible motion. We propose FastPhysGS, a fast and robust framework for physics-based dynamic 3DGS simulation: (1) Instance-aware Particle Filling (IPF) with Monte Carlo Importance Sampling (MCIS) to efficiently populate interior particles while preserving geometric fidelity; (2) Bidirectional Graph Decoupling Optimization (BGDO), an adaptive strategy that rapidly optimizes material parameters predicted from a VLM. Experiments show FastPhysGS achieves high-fidelity physical simulation in 1 minute using only 7 GB runtime memory, outperforming prior works with broad potential applications.
PaperID: 6961, Poster
Authors: Francesco Damiani, Ruben Moreno Bote
Abstract: Building an optimal controller requires solving two coupled problems: (1) computing beliefs about the state by filtering observations, and (2) designing a control signal based on those beliefs to minimize a cost function. Ideally, these two problems can be solved independently: first, a filter is computed to estimate the state; then, a controller is built on top of that estimate. When exact inference is tractable, this approach is optimal and applies to the prominent Linear-Quadratic-Gaussian stochastic control problem. However, perfect inference is generally intractable, and filtering and control are closely intertwined problems. The question, then, is whether the commonly used first-filter-then-control approach remains optimal, and whether optimal control strategies should generally rely on internal representations that mirror the dynamics of the external world. We show that this is not the case: even in a simple setting with linear dynamics, multiplicative noise, and internal noise, the optimal linear controller is characterized by internal forward dynamics that do not match the forward dynamics of the external state. Instead, the optimal controller relies on internal representations that mix estimation and control, and this mismatch becomes more pronounced as internal noise increases. Our work challenges standard approaches that prioritize external world modeling over control.
Abstract: Unified Multimodal Large Language Models (MLLMs) offer a promising paradigm for unifying visual understanding and generation, yet they still struggle to follow complex spatial instructions and logical constraints in controllable image generation. To address this gap, we present ATLAS, a unified framework that equips MLLMs with a human-like ``Think, Plan, and Paint'' paradigm. We adopt layout as the shared representation that connects the three stages, enabling the model to reason about spatial requirements, plan explicit object arrangements, and render the final image. We further improve plan-to-image fidelity with reinforcement-learning-based layout alignment. We instantiate ATLAS at 7B and 80B scales, achieving state-of-the-art performance among MLLMs on image generation benchmarks and an average 65.31% improvement over existing layout-aware MLLMs. On spatially related tasks, ATLAS obtains a 23.06% gain over the base models. Through the same layout interface, ATLAS also supports instruction-guided editing and multimodal grounding. We further introduce ATLAS-Reasoning, a benchmark for evaluating generation under complex spatial instructions.
PaperID: 6963, Poster
Abstract: LLM agents often execute long-horizon tasks where sequential tool calls and reasoning steps compound into a final outcome. When an agent errs mid-execution, the critical question is not whether an error occurred, but whether recovery under the current policy is still possible. Existing failure analysis answers this retrospectively, making it too late for runtime intervention. We introduce Task Failure Forecasting: predicting from a partial execution trace whether continued execution under the current agent policy is likely to fail. This is hard: human experts achieve only 54% accuracy at distinguishing recoverable from failure-inducing errors, and frontier models such as Claude-4.5 Sonnet reach 59.87% weighted F1. The core obstacle is that failed traces mix recoverable mistakes with fatal ones, obscuring which steps drive downstream failure. We address this with adaptive fault injection: injects targeted perturbations into successful traces, uses bandit-based prioritization of high failure yield error types and rollout-based verification providing policy-conditioned step-level supervision. A lightweight Qwen-3-8B model trained on this synthetic data achieves 78.01% weighted F1 on human-annotated benchmarks, outperforming Claude-4.5 Sonnet by 14 weighted-F1 points at 100× lower forecasting cost. When deployed as a runtime monitor for selective inference-time intervention, the forecaster improves agent success rates across HotpotQA, MuSiQue, GAIA, AssistantBench, SWE-Bench, MBPP and EnterpriseBench, yielding absolute gains in the range of 2–16 % points across held-out and unseen tasks.
PaperID: 6964, Poster
Abstract: Scenario generation is essential for training, testing, and safety validation of autonomous vehicles (AVs), especially when real-world data coverage is limited and long-tail safety-critical events are rare. Existing methods face a three-way trade-off among open-loop prediction accuracy, closed-loop simulation robustness, and inference efficiency. Autoregressive methods are typically efficient and strong in open-loop fitting, but they are prone to error accumulation and policy-induced state-distribution shift in long-horizon closed-loop rollouts. Diffusion-based methods provide strong multimodal behavior modeling and competitive generation quality, but iterative denoising substantially increases inference latency and limits simulation throughput. To address this challenge, we propose UniDBO, a Unified Dual-Branch One-step denoising framework for autonomous driving scenario generation. UniDBO uses a shared scene encoder and two jointly optimized complementary branches with distinct roles: a continuous-scale branch (CS) for open-loop trajectory fitting and a discrete high-noise branch (DHN) for robust closed-loop simulation. Through joint training, UniDBO coordinates open-loop fitting and closed-loop robustness while preserving one-step inference efficiency. We evaluate UniDBO on multiple benchmarks, including in-domain open-loop and closed-loop evaluation, zero-shot cross-domain generalization, and system-level closed-loop AV testing in unseen scenarios. Results show that UniDBO improves the open-loop/closed-loop balance while maintaining high inference efficiency, and achieves competitive zero-shot generalization on the evaluated benchmarks.
Abstract: Video world models should maintain evolving states when evidence is unobserved, yet current generators often freeze hidden states upon interruption. This is not simply a capacity problem: pretrained video diffusion transformers already possess KV-cache mechanisms capable of non-local retrieval, but they are rarely trained to use them as dynamic memory. We introduce ReMind, a framework eliciting dynamic memory behavior via memory-oriented data, event-aware training, and cache adaptation. Organized around a taxonomy of 100+ dynamic events, we build a camera-annotated training mixture combining VLM-filtered real videos, generated hard dynamics, synthetic camera loops, and memory-interruption augmentations. Each clip is converted into a frame graph with protected anchors, degraded intervals, and explicit temporal gaps. A node-structured curriculum—including node-drop, noisy memory, frontier continuation, and reference-cache training—forces the model to retrieve relevant past states across interruptions rather than relying solely on local continuity. PM-RoPE, an elegant camera-phase RoPE extension, unlocks spatiotemporal retrieval at a single-attention cost while preserving pretrained pathways. ReMind achieves the best overall scores on STEVO-Bench and recovery tasks. Furthermore, general image-to-video evaluations confirm this curriculum avoids catastrophic forgetting. We will open-source our code, data, and models.
PaperID: 6966, Poster
Abstract: Specifying segmentation targets via text descriptions provides flexible semantic conditioning for medical image segmentation. However, existing methods still struggle to robustly translate high-level textual semantics into accurate pixel-level spatial prompts. In this paper, we propose RADAR (Residual Aggregation and Dense Alignment Representations), a parameter-efficient framework for text-guided medical image segmentation. RADAR explicitly decouples and synergizes the medical semantic anchoring capabilities of BiomedCLIP with the prompt-driven mask decoding abilities of MedSAM2, achieving low-cost task adaptation via LoRA. Specifically, RADAR introduces Guided Residual Depth Aggregation, which adaptively aggregates historical layer features under semantic guidance to mitigate textual semantic injection lag and deep feature dilution. Furthermore, we propose Structured Dense Representation Alignment, which leverages a frozen BiomedCLIP to construct crop-level semantic anchors, namely target, boundary, and context, to enhance the local discriminative power of dense features and the reliability of automated prompts. Additionally, Multi-modal Prompt Generation is designed to transform cross-modal semantic responses into composite prompts consisting of text, points, and masks. Experimental results on three public medical image segmentation datasets demonstrate that RADAR achieves competitive performance while introducing merely 0.74M trainable parameters, validating the effectiveness of the proposed approach in parameter-efficient adaptation and the enhancement of cross-layer and local representations. The source code is provided in the supplementary material.
Abstract: We study nonconvex stochastic optimization under the Blum-Gladyshev (\mathsfBG-0) noise model, where the stochastic gradient variance grows quadratically with the distance from the initialization. We consider this problem under both standard smoothness and the symmetric generalized-smoothness framework, which captures objectives whose local curvature can scale with the gradient norm. We prove that normalized stochastic gradient descent with momentum, using only one stochastic gradient per iteration, converges under \mathsfBG-0 noise with oracle complexity \mathcalO(\varepsilon^-6). This rate holds both for standard smoothness and for \alpha-symmetric generalized smoothness, showing that generalized smoothness is rate-neutral for normalized momentum in this setting. We then study a variance-reduced normalized STORM method. Under mean-square smoothness and sharp initialization, the method achieves the minimax optimal \mathcalO(\varepsilon^-4) complexity, matching the lower bound. Under expected \alpha-symmetric generalized smoothness, the STORM recursion couples gradient-dependent smoothness with distance-dependent noise, leading to complexity \mathcalO(\varepsilon^-(4+\alpha)) for \alpha \in (0,1) and \mathcalO(\varepsilon^-5) for \alpha=1. When the distance-growth parameter in the noise model vanishes, our guarantees recover the standard bounded-variance rates: \mathcalO(\varepsilon^-4) for momentum, \mathcalO(\varepsilon^-3) for variance reduction, and \mathcalO(\varepsilon^-2) in the deterministic case. To our knowledge, these are the first convergence guarantees for normalized methods in non-convex stochastic optimization under \mathsfBG-0 noise without bounded domains, increasing batch sizes, or explicit anchoring, covering both standard and generalized smoothness regimes.
PaperID: 6968, Poster
Abstract: Personalized image generation aims to synthesize text-driven images conditioned on reference images, while mainly casting the generation as image customization for foreground and style transfer for background. Previous arts of diffusion models suffers from the text misalignment with background for image customization and foreground for style transfer during the denoising process. Such facts, as we observed, rooted from the entanglement among hybrid frequency bands during the denoising process. In particular, the low-frequency band primarily captures color information, mid-frequency band encodes global structure, while high-frequency band corresponds to local texture. Besides, the early (high-noise) denoising stage mainly focuses on the low-and-mid frequency bands, while the late (low-noise) denoising stage progressively exhibits the high-frequency band, resulting in the dominance of low-and-mid over high frequency bands, especially when entangled within a single image reference, tends to reinforce each other at the early stage, hence the pitfall of overwhelming suppression of text prompt guidance. To address such salient limitation, in this paper, we study personalized generation based on dual references - customization and color and style reference - and propose a paradigm to disentangle these , to simultaneously tackle two crucial personalized image generation tasks: customization style transfer and color style transfer, by disentangling different frequency bands via mask strategy within frequency domain. For customization style transfer, we replace the mid-frequency band of the background in the style reference with that from the foreground of the customized reference. For color style transfer, we substitute the low-frequency band of the background in the style reference with that from both the foreground and background of the color reference. Both the substituted frequency bands are used as the key and value to reconstruct the query foreground and background of the denoised personalized image. We further compute the ratio of information entropy of the substituted frequency bands to adaptively modulate the denoising timesteps between the early and late stages. Extensive experiments validate the superiority of over the state-of-the-art diffusion models for personalized image generation. Our code can be accessed from the supplementary material package.
PaperID: 6969, Poster
Authors:
JUNFENG YU, Junkai Bao, Hao Zeng, XIAOWEN MO, Xin Liu, Xuejian HuangAbstract: Diffusion-based generative attacks have emerged as a promising paradigm for targeted adversarial example generation, leveraging generative priors to improve transferability and visual fidelity. However, existing methods typically entangle image reconstruction and target manipulation within the same reverse denoising trajectory, where target guidance can disrupt early structural recovery. This coupling creates a persistent trade-off among attack effectiveness, transformation robustness, and perceptual fidelity. In this paper, we propose SDTI (Stage Decoupled Target Injection), a stage-decoupled reverse diffusion framework that separates source-structure preservation from adversarial target injection along the denoising timeline. SDTI first maps the input image to an intermediate latent through null-text DDIM inversion, then recovers and anchors the source structure with a frozen null-text denoising prefix, and finally activates target-conditioned LoRA guidance only in the late denoising suffix to inject adversarial target semantics. To confine adversarial optimization to this suffix phase, we further introduce suffix-wise supervision with inter-step gradient truncation. Extensive experiments on ImageNet-NeurIPS, robust victim models, common input transformations, and MS-COCO cross-domain transfer show that SDTI achieves a favorable attack--fidelity--robustness trade-off, improving transformation robustness and perceptual quality while maintaining competitive targeted transferability. The code is available at: https://anonymous.4open.science/r/AY_code-1431.
PaperID: 6970, Poster
Abstract: Robust policy-gradient methods provide a principled framework for learning under model uncertainty, but their sample efficiency is often limited by the cost of robust policy evaluation. Unlike standard Bellman updates, robust Bellman operators are nonlinear, so sample-based critic estimates can introduce systematic error that is difficult to control. Existing finite-sample guarantees therefore often require solving the robust critic to high accuracy before each actor update. However, practical robust actor-critic methods typically use small or decaying actor stepsizes rather than aggressive increasing updates, making repeated high-accuracy critic solves particularly costly. Conservative actor updates are especially natural in robust constrained MDPs, where the optimization objective may switch between reward improvement and constraint correction during learning. We show that this high-accuracy critic requirement is overly conservative for KL-robust policy optimization. Rather than controlling the full first-order critic error, such as \mathbb E \lVert\widehat V_t - V^\pi_t\rVert, the actor analysis only requires control of the conditional critic bias, \lVert\mathbb E[\widehat V_t \mid \mathcal F_t] - V^\pi_t\rVert. We prove that, in the relevant local regime, this bias scales quadratically with the critic estimation error. Consequently, an O(\varepsilon) actor error only requires the critic mean-square error to be O(\varepsilon), rather than requiring the root-mean-square error to be O(\varepsilon). Combining this bias-based analysis with robust natural policy-gradient updates, we establish an \widetilde O(\varepsilon^-3) sample-complexity guarantee for model-free discounted robust MDPs with small or decaying actor stepsizes. We further show that the same analysis extends to primal surrogate methods for robust constrained MDPs under KL uncertainty, improving the state of the art for this setting.
Abstract: Diffusion large language models (dLLMs) generate text by iteratively denoising masked token sequences. Although dLLMs can predict all masked positions in parallel within each step, the large number of denoising iterations still makes inference expensive. This cost can be reduced spatially by unmasking multiple tokens per step, or temporally by collapsing multiple denoising steps into one verification call. We propose Parallel Speculative Decoding (PSD), a training-free framework that jointly improves inference along both axes. Using the confidence scores from a single forward pass, PSD selects positions to unmask via a configurable, adaptive unmasking policy and constructs multi-depth speculative drafts without extra model calls. A final batched verification pass then applies hierarchical acceptance, keeping the deepest draft that remains consistent with the updated predictions. Experiments on three dLLMs across reasoning and code generation tasks show that PSD achieves favorable trade-offs between inference efficiency and generation quality, reaching up to 5.5× tokens per forward pass with accuracy comparable to greedy decoding.
Authors:
Bohan Li, Shuojue Yang, Baorui Peng, Xianda Guo, Erli Zhang, Youqi Tao, Junfeng Duan, Daguang Xu, DOU QI, Xin Jin, Wenjun Zeng, Hao Zhao, Yueming JinAbstract: Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dimensional control vectors must precisely govern complex image-space evolution. In this work, we propose a kinematic-to-visual lifting paradigm that converts articulated kinematics into a unified set of five image-aligned control modalities. Building on this representation, we introduce a hierarchically routed visual control framework that selectively activates the most relevant control modalities and motion scales. Instead of uniformly applying all control signals, our model performs hierarchical routing to dynamically allocate conditioning capacity. We further design kinematic-prior-guided routing loss functions to ensure physically meaningful, temporally stable, and efficient expert utilization. To improve efficiency, we propose a budgeted training and inference scheme that leverages routing-induced sparsity. By selectively discarding low-significance control pathways during training and execution, our approach enables adaptive computation that is complementary to standard distillation. We additionally construct a new benchmark with curated articulated annotations, obtained through human-in-the-loop semantic labeling and differentiable pose tracking, providing realistic supervision for action-conditioned surgical video generation. Extensive experiments demonstrate that our method consistently improves action faithfulness, visual fidelity, and cross-domain generalization over diverse baselines. Moreover, our efficient variant achieves substantial reductions in latency while maintaining strong control accuracy.
PaperID: 6973, Poster
Abstract: Prior work on grokking—delayed generalization after memorization—has sought to shorten this delay through data augmentation or optimizer design. We take an architectural approach: rather than analyzing trained networks post-hoc, we modify architectural topology before training to test whether specific degrees of freedom prolong the memorization phase. We identify two independent factors in standard Transformers—unbounded residual magnitude and data-dependent attention routing—and introduce targeted interventions for each. A fully bounded spherical topology, enforcing L2 normalization throughout the residual stream, eliminates the memorization phase entirely on modular addition and multiplication over Z_p: training and test accuracy rise together from initialization. Independently, replacing learned attention with uniform aggregation achieves the same effect, consistent with theoretical results showing that commutative operations require only a bag-of-tokens representation. However, the same spherical constraint fails completely on non-commutative S5 permutation composition, where successful solutions rely on discrete coset structures rather than continuous Fourier features. This contrast demonstrates that bypassing the generalization delay is possible—but strictly depends on alignment between architectural priors and task symmetry. Since even the simplest non-commutative task resists the same constraint that eliminates delay on cyclic arithmetic, these results suggest that no single geometric prior can universally accelerate learning across the heterogeneous structures present in general-purpose domains.
PaperID: 6974, Poster
Abstract: AI-generated image detection has become increasingly important as generative models produce highly realistic visual content. Existing detectors mainly rely on visual artifacts or visual features extracted from pretrained multimodal models, while largely overlooking the structural differences between real and generated images in the joint image--text space. In this paper, we reveal a systematic Cross-modal Alignment Shift (CAS): generated images tend to exhibit stronger and more concentrated alignment with text than semantically matched real images. We verify this phenomenon through retrieval preference, image--text residual structure, spectral concentration, and matching entropy. Motivated by this finding, we propose a CAS-based detection framework(CASNet) that learns an amplified multimodal model with a text-guided alignment-shift objective. The amplification-induced feature increment is extracted as an explicit cross-modal cue and coupled with amplified visual representations, enabling the detector to exploit both alignment-level structural evidence and visual discriminative cues. Extensive experiments across diffusion models, GANs, autoregressive models, and commercial generators demonstrate consistent improvements over state-of-the-art detectors. In particular, on recent commercial generators, our method improves ACC and AP by 5.48% and 6.13%, respectively.
PaperID: 6975, Poster
Authors:
Yan Song, Yunfei Guo, Guyeli Yang, Li He, Chenglong Li, Huanzhen Wang, Bailin HE, Qiao Sun, Zhong Jiangwei, Shizeng Zhang, Nuo Chen, Yan Wang, Wenqiang ZhangAbstract: Active view selection for 3D Gaussian Splatting (3DGS) requires information-theoretic grounding, computational efficiency, and strong reconstruction quality, yet existing methods satisfy these only partially. Information-theoretic approaches quantify gain over the full parameter space via costly Fisher computation under a Laplace approximation; faster alternatives sacrifice the grounding they started from. We propose SpheriBED, a Bayesian experimental design framework that defines the information objective over per-Gaussian spherical appearance—the quantity each observation fundamentally measures—rather than the full parameter vector, yielding exact closed-form posteriors without Laplace approximation. Scoring a candidate view requires only a single forward rasterization pass with no backward pass or per-candidate optimization. Extensive experiments across benchmarks demonstrate that SpheriBED consistently improves active reconstruction performance and provides more reliable uncertainty quantification compared to state-of-the-art methods.
Authors:
Yuqi Li, Matthew EngelhardAbstract: In high-stakes risk prediction, interval-valued predictions provide a natural way to represent predictive uncertainty. However, standard evaluation tools such as the receiver operating characteristic (ROC) curve and the area under the curve (AUC) are designed for point-valued predictions and do not capture how uncertainty affects discrimination performance. To address this gap, we develop an uncertainty-aware ROC framework and corresponding interval-based AUC (iAUC) metrics for evaluating the performance of interval-valued risk predictions. The framework constructs two ROC-style curves with associated areas \mathrmAUC_L and \mathrmAUC_U, and shows that these quantities induce a three-region decomposition of positive-negative rankings into confidently correct, ambiguous, and confidently incorrect orderings. Under valid class-conditional coverage, \mathrmAUC_L and \mathrmAUC_U, together with the pairwise miscoverage rate, provide theoretical bounds on the optimal AUC, linking interval coverage to achievable discrimination. The framework is compatible with a range of methods for quantifying epistemic uncertainty in risk prediction, including Bayesian, bootstrap, and ensemble-based procedures, and provides a practical way to compare predictive models and interval-generation methods. Experiments on clinical benchmark data and MIMIC-IV mortality prediction show that models with similar point AUC can have different uncertainty-aware ranking profiles, and that the proposed decomposition can support risk-aware model selection and threshold-based abstention.
PaperID: 6977, Poster
Authors:
Hui Zhang, Yachao Yuan, Jiayun Wang, Yuanzhuo Li, Hongtao Wang, Yali YuanAbstract: Fine-tuning-as-a-service enables users to adapt aligned large language models (LLMs) to specialized tasks, but malicious fine-tuning can erode refusal behavior while preserving task performance on legitimate inputs. We revisit recent layer-wise safety diagnostics and find that safety sensitivity is signed: scaling different layers can strengthen refusal, weaken it, or have little effect. Motivated by this observation, we propose SLDR, a post-fine-tuning defense based on Selective Layers Recovery and Dynamic Routing. SLDR trains a LoRA recovery adapter only on the layers with the maximum and minimum sensitivity scores in the signed spectrum, and uses representation-based dynamic routing inference to activate the adapter only for malicious queries. Across four model architectures, five downstream tasks, and four harmful benchmarks, SLDR substantially reduces harmful outputs while preserving downstream utility. On Llama3.1/SST2, SLDR reduces the average harmful score from 11.54 to 0.08 while maintaining downstream accuracy, and the harmful score remains near zero under poisoning ratios up to 0.9.
PaperID: 6978, Poster
Authors: Chi Cui, Hei Victor Cheng, Jesús Cid-Sueiro
Abstract: Learning from weak labels (including noisy, partial or complementary labels) requires exploiting the statistical dependence between ground-truth classes and observed labels. Many existing methods are built around a particular transition model, often assumed to be known, accurately estimated, or instance-independent, which leads to performance degradation under model misspecification. We introduce Convex Losses for Weak Labels (CLWL), a family of losses for weak supervision based on transforming supervised one-vs-all losses. By prioritizing classification and ranking consistency over probability calibration, CLWL relaxes rigid structural assumptions on the label transition mechanism. We provide conditions for convexity and ranking consistency, characterize the transition matrices that preserve consistency for a given loss, and show that this consistency-preserving space attains the theoretical dimension bound. We propose a constructive method for generating parametric, lower-bounded, convex losses adapted to arbitrary weak-label models. Finally, we connect CLWL to existing methods, including backward correction and some losses for partial or complementary labels. Empirical results demonstrate that CLWL outperforms state-of-the-art methods when the assumed transition model is misspecified.
Abstract: Generative modeling continuous-time, continuous-space stochastic processes (e.g., videos, weather forecasts) conditioned on partial observations (e.g., first and last frames) is a fundamental challenge. Existing approaches, (e.g., diffusion models), suffer from key limitations: (1) noise-to-data evolution fails to capture structural similarity between states close in physical time and has unstable integration in low-step regimes; (2) random noise injected is insensitive to the physical process's time elapsed, resulting in incorrect dynamics; (3) they overlook conditioning on arbitrary subsets of states (e.g., irregularly sampled timesteps, future observations). We propose ABC: Any-Subset Autoregressive Models via Non-Markovian Diffusion Bridges in Continuous Time and Space. Crucially, we model the process with one continual SDE whose time variable and intermediate states track the real time and process states. This has provable advantages: (1) the starting point for generating future states is the already-close previous state, rather than uninformative noise; (2) random noise injection scales with physical time elapsed, encouraging physically plausible dynamics with similar time-adjacent states. We derive SDE dynamics via changes-of-measure on path space, yielding another advantage: (3) path-dependent conditioning on arbitrary subsets of the state history and/or future. To learn these dynamics, we derive a path- and time-dependent extension of denoising score matching. Our experiments show ABC's superiority to competing methods on multiple domains, including video generation and weather forecasting.
PaperID: 6980, Poster
Abstract: Tabular prior-fitted networks (PFNs) work well across tabular classification problems but can fail silently at deployment when distributions shift. We propose (Calibrated Anchor-Loss Disagreement), a lightweight unsupervised monitor that triggers an alarm to warn when a PFN might fail. By finetuning the decoder to maximize disagreement with the test-set prediction while maintaining accuracy on an ID anchor set, identifies harmful OOD batches through the resulting anchor loss, requiring no labels on test batches. We prove that the post-finetune anchor loss has a strictly lower expectation under harmful shift than under ID eval, yielding a one-sided test with a formal power guarantee. Across 14 datasets covering tabular and structured distribution shift, on both TabPFN and TabICL, matches or outperforms baselines on harmful shifts and is markedly more resilient to benign-shift false alarms.
PaperID: 6981, Poster
Authors: Daniel Platnick, Marjan Alirezaie, Hossein Rahnama
Abstract: Modern AI agents can plan, reflect, reason, and act over long-horizon digital tasks. Even so, they cannot plan to steer a social interaction in real time. Doing so requires anticipating how latent factors like trust and resistance will evolve under the agent's actions, fast enough to plan within a conversational turn. Generative world models approach this by narrating possible futures, but autoregressive text generation is both too slow for real-time planning and fundamentally lossy. Representation-predictive methods can be superior, but are underexplored for interaction. We build LID-Bench, the first controlled testbed with oracle latent states for interaction dynamics, enabling systematic comparison of generative and representation-predictive world models against known dynamics. A bottleneck decomposition on three generative models spanning 117M--350M parameters reveals that all learn interaction dynamics internally, with text-rendering signal losses of 89--100% observed by horizon k=5. Our representation-predictive model (Social-JEPA, 125M params, 500K trainable) bypasses text entirely, forecasting latent dynamics 2.8--3.9× more accurately while performing rollouts 1,059--2,314× faster---making real-time planning within conversational turn-taking latencies feasible.
PaperID: 6982, Poster
Authors: Ashima Khanna-Reiter, Dominik G Grimm
Abstract: Protein sequence optimization under tight oracle budgets requires methods that explore vast combinatorial spaces while making each evaluation informative. Existing reinforcement learning and off-policy generative approaches often degrade under surrogate noise, and position-agnostic mutation proposals risk disrupting functionally critical residues. We introduce SILO, a trajectory-level self-improvement imitation framework for oracle-budgeted protein design. SILO uses a hierarchical edit policy that decomposes each mutation into a position choice followed by a residue choice. In each active-learning round, the policy samples candidate trajectories via incremental stochastic beam search without replacement (SBS), and a UCB-based proxy ensemble, combined with an alanine-scan fitness score (AFS), selects candidates with functionally relevant edits for in silico oracle evaluation. The policy is then updated by next-action cross-entropy imitation on the round’s best oracle-labeled trajectories, avoiding value-function estimation. Across eight reproduced protein fitness landscapes and five strong baselines from prior work, SILO achieves the highest maximum and top-100 mean fitness on 8 of 8 landscapes within our evaluations, often exhibiting faster early-stage improvement. In low-data and noisy-proxy stress tests on two landscapes per setting, SILO remains competitive or best when several baselines degrade. Ablations show that SBS with AFS account for much of the gains, with iterative imitation providing additional improvement.
PaperID: 6983, Poster
Abstract: Transformers for temporal sequences require mapping sampled signals into sequences of vectors, called tokens. This is done by partitioning the sequence into local temporal windows (patches). We show that the tokenization step induces an approximation--estimation tradeoff. Smaller windows yield higher-variance statistical estimates because each token is supported by fewer samples, while larger windows stabilize estimation but increase the amount of information that must be summarized. Motivated by this tradeoff, we propose a lightweight tokenizer that constructs multiple estimates for each token and sums them, without changing the Transformer backbone. Across 12 datasets spanning four modalities, our method improves performance and reduces sensitivity to temporal scale across tasks.
Abstract: Generating data from discrete distributions is important for a number of application domains including text, tabular data, and genomic data. Several groups have recently used random k-satisfiability (k-SAT) as a synthetic benchmark for new generative techniques. In this paper, we show that fundamental insights from the theory of random constraint satisfaction problems have observable implications (sometime contradicting intuition) on the behavior of generative techniques on such benchmarks. More precisely, we study the problem of generating a uniformly random solution of a given (random) k-SAT or k-XORSAT formula. Among other findings, we observe that: (i) Continuous diffusions outperform masked discrete diffusions; (ii) Learned diffusions can match the theoretical `ideal' accuracy; (iii) Smart ordering of the variables can significantly improve accuracy, although not following popular heuristics.
PaperID: 6985, Poster
Authors: Lyuxin Xue, Hui Cai, Xiaoyun Feng, Xuanwei Hu, Xin Zhang
Abstract: Continual pre-training (CPT) enables injecting domain-specific knowledge into large language models during the learning rate annealing phase. However, this compute-efficient paradigm forces practitioners to navigate a challenging optimization space: simultaneously choosing the injection fraction \lambda, the data mixture ratio r, and the re-warmup peak learning rate \eta. Because existing scaling laws have yet to comprehensively integrate and couple these interacting variables, they cannot reliably forecast the resulting trade-offs and final convergence outcomes across unseen injection configurations. To bridge this gap, we introduce the Annealing-phase Domain Injection Scaling Law (ADIS-Law), a parametric framework that explicitly models late-stage distribution shifts as physical perturbations applied to the base pre-training trajectory. ADIS-Law formalizes two opposing macroscopic behaviors: general-domain performance follows a linear forgetting-recovery process, whereas target-domain adaptation adheres to a power-law adaptation trend. Across data-budget and model-scale extrapolation, ADIS-Law predicts target-domain trajectories with an average Spearman correlation of 0.947 and forecasts tail losses with relative error below 5%. We further show that ADIS-Law can predict optimal injection strategies at extrapolated scales: it ranks target-scale candidate policies with a mean Spearman correlation of 0.936. By replacing blind grid search over (\lambda, r, \eta) with a single fitted law, ADIS-Law offers a data-driven approach for strategy selection in annealing-phase domain injection.
Authors:
Haoran Sun, Wenjie Li, Yujie Zhang, Zekai Lin, Fanrui Zhang, Kaitao Chen, Xingqi He, Yichen Li, Mianxin Liu, Lei Liu, Yankai JiangAbstract: Medical agent systems are increasingly expected to support interactive clinical decision making rather than only static question answering. In such settings, effective agents must reuse prior experience across evolving cases, yet existing memory mechanisms often retain raw historical traces that are redundant, noisy, and difficult to govern. More importantly, they rarely distinguish which memories are truly useful for future reasoning. This limits their ability to accumulate compact and reliable experience for long-horizon clinical reasoning. To close this gap, we propose SkeMex, a post-deployment self-evolution framework that improves medical agents through a skill-based memory without updating model weights. SkeMex distills informative interaction trajectories into structured skills that encode reusable procedural knowledge, and organizes them into a multi-branch repository spanning general, task-specific, and action-level experience. To determine which memories should be reused and retained, SkeMex estimates context-dependent utility from environment feedback and uses it to guide value-aware retrieval and repository governance. A closed-loop ``Read--Write--Assess--Govern" lifecycle further supports continual evolution by writing new skills, updating utilities, promoting useful memories, and removing harmful entries. Experiments across diverse clinical tasks show that SkeMex consistently outperforms representative memory-based agents in both offline and online settings. It also generalizes across model backbones and supports transferable skill memory. All data and code will be released publicly.
PaperID: 6987, Poster
Abstract: A common approach for contextual optimization first trains a model to predict the unknown state from the context, and then takes the action that is optimal for the predicted state. We show that this follow-the-prediction policy can be brittle: for several standard problems, even a minimum-variance predictor with arbitrarily small RMSE can incur constant excess loss. We propose a simple policy that robustifies any black-box predictor by optimizing against the worst case in an \varepsilon-neighborhood of the prediction. We identify a problem-specific quantity, the \emphrobust optimization gap, that captures the price of this hedging, and prove a general bound on the loss. We instantiate the framework for contextual posted pricing, ski rental, and house flipping, deriving closed-form robust policies and O(\eta^2/3) loss bounds in all three settings, where \eta is the RMSE of the predictor. For posted pricing, this improves upon a result of Medina and Vassilvitskii (2017), and with a substantially simpler argument.
Authors: Tim Sinen
Abstract: We study the complexity of smoothed agnostic learning of halfspaces on \\pm 1\^n under uniform marginals in the model of \citetKM25 where each input coordinate is independently flipped with probability \sigma \in (0, 1/2). We show that L^1 polynomial regression achieves runtime and sample complexity \tildeO(n^O(\log(1/\varepsilon)/\sigma)), and prove a nearly matching Statistical Query complexity lower bound of n^\Omega(\log(1+\sigma/\varepsilon^2)/\sigma). This complements the recent work of \citetDK26, which established analogous bounds in the continuous setting under Gaussian marginals.
PaperID: 6989, Poster
Abstract: Despite advances in retrieval-augmented generation (RAG), suppressing hallucinations during complex reasoning remains a persistent challenge. We identify that existing RAG pipelines suffer from a textual bottleneck: by passing only raw text to the generator, they discard discriminative signals of the retriever. Furthermore, recent attempts to unify retrieval and generation within a single LLM induce severe objective conflict and exacerbate the rank collapse of decoder-only architectures, degrading performance on both tasks. To address these limitations, we introduce ShiftRAG, a decoupled RAG framework that operates on a single LLM backbone but isolates discriminative retrieval learning from autoregressive generation. To bridge these decoupled modes, we project retrieval alignments into continuous soft tokens, establishing a high-bandwidth interface that infuses the generator with retrieval-side signals. Concurrently, our soft orthogonality objective mitigates rank collapse, while layer-wise relevance dynamics preserve the fine-grained latent space required for highly discriminative evidence separation. Extensive evaluations demonstrate that ShiftRAG significantly improves retrieval and question answering performance over state-of-the-art baselines. ShiftRAG preserves foundational capabilities of the backbone while maintaining a favorable efficiency-performance trade-off on general text encoding and generative tasks. Our code is available.
Abstract: As state-of-the-art neural networks are deployed on reasoning and algorithmic tasks, exactness guarantees become increasingly important. However, high average-case accuracy can still mask inconsistent behaviors. This motivates exact certification, which asks for the smallest set of labeled examples needed to certify that a learned hypothesis equals the target. We show that while some hypotheses are easy to certify, even minimal overparametrization can make certification exponentially hard across several hypothesis classes. For threshold circuits of depth \ge 2, adding a single extra gate can force certificate sizes exponential in the input dimension. We show an analogous hardness result for log-precision Transformers with only constant architectural overhead. We also characterize approximate certification, showing that allowing only polynomially many mistakes still requires exponentially large certificates, whereas constant relative-error guarantees can hide exponentially many failures. Empirically, we study certification for circuits and Transformers trained to recognize binary addition and find that imperfect models can evade detection unless the certificate is exponentially large.
PaperID: 6991, Poster
Authors: Ferhat Arslan, Weihong Guo, Shuo Li
Abstract: Kolmogorov–Arnold Networks (KANs) expand each input coordinate through a fixed basis family, yet the optimization role of this basis lifting remains poorly understood. We identify basis lifting as a structural source of ill-conditioning in KANs and provide a mechanism-level analysis linking lifted feature geometry to convergence, accuracy, and seed sensitivity. Our analysis studies the evolution of lifted empirical Gram matrices under stochastic gradient descent (SGD) and shows that fixed basis lifting induces and preserves near-rank deficiency through eigenvalue and determinant drift. It further concentrates feature energy into a smaller number of directions, amplifying dominant eigenvalues and worsening condition numbers. These results show that basis selection is not merely an expressivity choice, but also an optimization and conditioning choice. Motivated by this mechanism, we introduce T-KAN, a learnable basis-coordinate transformation that aligns basis directions before the learned linear map. It reshapes lifted Gram geometries while preserving the expressivity and introducing negligible parameter overhead. In controlled small-to-moderate KAN regimes, T-KAN often improves conditioning and optimization with B-spline, Chebyshev, and RBF KANs.
PaperID: 6992, Poster
Abstract: Reinforcement learning (RL) is essential for adaptive decision-making, e.g., robust Meta-RL and domain randomization (DR), and efficient post-training of large language models (LLMs), where tasks are commonly structured as Markov decision processes or prompts. A central bottleneck in all these settings is the cost of agent-environment interactions required for task selection and policy optimization. Model predictive task sampling (MPTS) has emerged as a promising approach to improve sample efficiency by querying informative tasks via a risk predictive model (RPM) trained on optimization history. However, existing RPMs suffer from two failure modes: (i) sparse optimization histories, which limit RPMs' generalization across the task space, and (ii) a one-step lag in risk approximation, which causes biased difficulty estimation as the policy evolves. This work formalizes these failure modes, derives lagged-risk decompositions that expose the resulting temporal bias, and recasts risk prediction as an in-context sequence modeling problem. The developed RPM integrates a one-step look-ahead mechanism with a temporal-weighted sliding window to mitigate data non-stationarity and scarcity, without requiring additional rollouts. Empirically, our approach yields more reliable difficulty estimates and consistent performance gains across robust Meta-RL, DR, and prompt curriculum RL post-training of LLMs.
PaperID: 6993, Poster
Abstract: This paper studies category-level RGB-D object pose estimation, which recovers an object's 3D rotation and translation without instance-specific CAD models or reference views. This task is highly ambiguous due to intra-class variations, symmetries, and partial observations. Current solutions struggle to balance reliability and efficiency: deterministic methods are prone to error propagation from intermediate structures, while diffusion models suffer from high inference latency and complex post-hoc candidate extraction. To address this, we present F2G-Pose, a real-time direct, single-pass framework that predicts category-level pose, metric size, and camera-frame completed shape from a single RGB-D observation in one feed-forward pass. F2G-Pose lifts dense visual foundation features to partial point clouds, encodes local geometry with DGCNN tokenization, and integrates object-level structure with a Mixture-of-Experts Transformer. Joint size prediction and camera-frame shape completion encourage metric consistency through dense geometric supervision. Trained only on synthetic SOPE data, F2G-Pose achieves state-of-the-art accuracy on SOPE and ROPE while running at 28 FPS, and further demonstrates strong camera-frame geometry on HouseCat6D and real-object transfer on HANDAL. Analyses of latency, robustness, stability, shape completion, and GenPose++ candidate behavior further demonstrate the efficiency and reliability of F2G-Pose, with real-world manipulation experiments demonstrating deployment potential under segmented RGB-D inputs.
Authors: Yifan Lyu, Liang Zhang
Abstract: Understanding human personality is crucial for web applications such as personalized recommendation and mental health assessment. Existing studies on personality detection predominantly adopt a ``posts \rightarrow user vector \rightarrow labels'' modeling paradigm, which encodes social media posts into user representations for predicting personality labels (e.g., MBTI labels). While recent advances in large language models (LLMs) have improved text encoding capacities, these approaches remain constrained by limited supervision signals due to label scarcity, and under-specified semantic mappings between user language and abstract psychological constructs. We address these challenges by proposing ROME, a novel framework that explicitly injects psychological knowledge into personality detection. Inspired by standardized self-assessment tests, ROME leverages LLMs’ role-play capability to simulate user responses to validated psychometric questionnaires. These generated question-level answers transform free-form user posts into interpretable, questionnaire-grounded evidence linking linguistic cues to personality labels, thereby providing rich intermediate supervision to mitigate label scarcity while offering a semantic reasoning chain that guides and simplifies the text-to-personality mapping learning. A question-conditioned Mixture-of-Experts module then jointly routes over post and question representations, learning to answer questionnaire items under explicit supervision. The predicted answers are summarized into an interpretable answer vector and fused with the user representation for final prediction within a multi-task learning framework, where question answering serves as a powerful auxiliary task for personality detection. Experiments on two real-world datasets show that ROME outperforms state-of-the-art baselines, achieving relative improvements of 13.60% and 19.72% on Kaggle and Pandora, respectively. Our code is available at https://anonymous.4open.science/r/ROME-6641.
Abstract: While novel view synthesis (NVS) for dynamic scenes has seen significant progress, reconstructing temporally consistent geometric surfaces remains a challenge. Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) offer powerful dynamic scene rendering capabilities; however, relying solely on photometric optimization often leads to geometric ambiguities. This results in discontinuous surfaces, severe artifacts, and broken surfaces over time. To address these limitations, we present DySurface, a novel framework that bridges the effectiveness of explicit Gaussians with the geometric fidelity of implicit Signed Distance Functions (SDFs) in dynamic scenes. Our approach tackles the structural discrepancy between the forward deformation of 3DGS (canonical \rightarrow dynamic) and the backward deformation required for volumetric SDF rendering (dynamic \rightarrow canonical). Specifically, we propose the VoxGS-DSDF branch that leverages deformed Gaussians to construct a dynamic sparse voxel grid, providing explicit geometric guidance to the implicit SDF field. This explicit anchoring effectively regularizes the volumetric rendering process, significantly improving surface reconstruction quality, containing watertight boundaries and detailed representations. Quantitative and qualitative experiments demonstrate that DySurface significantly outperforms state-of-the-art baselines in geometric accuracy metrics while maintaining competitive rendering performance. Codes will be publicly available.
Abstract: To enable reliable long-term interaction, LLM agents require a memory system that can faithfully store, efficiently retrieve, and deeply reason over accumulated dialogue history. Most existing methods adopt an extracted fact based paradigm: handcrafted static prompts compress raw dialogues into atomic facts, which are then stored, matched, and injected into downstream reasoning. Nevertheless, such fact-centric designs inevitably discard fine-grained details in original dialogues and fail to support deep reasoning over scattered isolated facts. Moreover, static prompts cannot maintain consistent extraction granularity across diverse dialogue styles. To address these limitations, we propose TriMem, which maintains three coexisting representation granularities, including raw dialogue segments anchored by source identifiers for storage fidelity, extracted atomic facts for efficient memory retrieval, synthesized profiles that aggregate dispersed facts into holistic semantic understanding for deep reasoning. We further adopt TextGrad-based prompt optimization, which iteratively refines extraction and profiling prompts via response quality feedback, achieving lifelong evolution without any parameter updating. Extensive experiments on LoCoMo and PerLTQA across multiple LLM backbones demonstrate that TriMem consistently outperforms strong memory baselines.
PaperID: 6997, Poster
Authors: Simone Betteti, Matteo Martin, Morten G Pedersen, Sandro Zampieri, Giacomo Baggio
Abstract: Associative memory models classically describe retrieval as convergence to stored patterns, but the theory of sequential retrieval remains comparatively limited. We develop a two-timescale input-driven Hopfield network in which fast associative retrieval is modulated by a slow feedback variable evolving over a prescribed memory-transition graph. This yields autonomous, graph-constrained transitions while preserving analytical tractability. Using slow-fast and geometric singular perturbation theory, we characterize memory fixation via memory manifolds and identify the geometric regions that govern stability loss and escape from retrieved memories. We further derive reduced slow dynamics and explicit criteria for transition onset, self-sustained retrieval, and collapse, expressed through escape times, gain thresholds, and fixed points. Together, these results provide a tractable framework for analyzing feedback-driven sequential retrieval in Hopfield networks.
PaperID: 6998, Poster
Authors: Yongkang Xiao, Qiyi Deng, Jing Chen, Min Shi, JU MA, Ruiying Du
Abstract: Tamarin-based formal verification of security protocols requires expert-authored multiset rewriting rules and temporal-logic lemmas. Translating a natural-language protocol description into a complete formal model with an LLM is hard because syntactic, reachability, and property-level errors compound across the multi-step modeling process, and a single agent that both authors and self-evaluates cannot reliably localize them. We present MASTA, a multi-agent system that addresses these difficulties through three innovations. First, an operational--evaluative role decomposition assigns rule modeling to an Architect, property modeling to an Auditor, and pure dispatch to a dedicated Arbiter that holds no authoring duties. Second, the Arbiter implements a Tamarin-anchored scheduler that dispatches the next agent based on the latest verifier verdict and routes problem-tagged fixes through a modification list, with a six-class problem taxonomy and an LLM deduplication check preventing routing deadlock. Third, we construct a benchmark of 10 protocols paired with 139 mutation variants and five fine-grained metrics that jointly assess model correctness (M_1--M_3) and completeness (M_4--M_5). On this benchmark MASTA on Qwen-3.5-Plus improves over OpenCode on the same backbone by up to +0.27 on individual metrics and reaches the highest overall average, surpassing Codex (GPT-5.5) and Claude Code (Opus~4.7) despite using a weaker LLM. Ablations confirm that removing the Arbiter collapses property and verification scores to near zero.
Abstract: Reinforcement Learning (RL) has empowered Multimodal Large Language Models (MLLMs) to achieve superior human preference alignment in Image Quality Assessment (IQA). However, existing RL-based IQA models typically rely on coarse-grained global views, failing to capture subtle local degradations in high-resolution scenarios. While emerging "Thinking with Images" paradigms enable multi-scale visual perception via zoom-in mechanisms, their direct adaptation to IQA induces spurious "cropping-implies-degradation'' biases and misinterprets natural depth-of-field as artifacts. To address these challenges, we propose Q-Probe, the first agentic IQA framework designed to scale IQA to high resolution via context-aware probing. First, we construct Vista-Bench, a pioneering benchmark tailored for fine-grained local degradation analysis in high-resolution IQA settings. Furthermore, we propose a three-stage training paradigm that progressively aligns the model with human preferences, while simultaneously eliminating causal bias through a novel context-aware cropping strategy. Extensive experiments demonstrate that Q-Probe achieves state-of-the-art performance in high-resolution settings while maintaining superior efficacy across resolution scales.
PaperID: 7000, Poster
Abstract: We revisit regularized regret minimization under full-information and bandit feedback, where a learner optimizes an objective of the form \langle r, \pi \rangle - \eta^-1 \psi(\pi) for adversarially chosen rewards r, policies \pi \in \Delta_K, and a known strongly-convex regularizer \psi(\cdot). We establish the minimax regularized regret for a broad class of regularizers, specifically those that are relatively strongly-convex and smooth with exponent p \geq 1, Legendre, and coordinate-wise separable. For large \eta, the problem approaches unregularized online learning, recovering the standard \Theta(\sqrtT) scaling. For small \eta, the strong convexity of the regularizer dominates, and we identify the regimes that attain \Theta(\eta \log T) minimax regret under both feedback models. Notably, while prior logarithmic regret lower bounds are largely restricted to specific curved loss functions (e.g., squared loss), our lower bounds hold for any such regularizer against purely linear adversarial rewards. A key technical contribution is our proof strategy, which leverages a Bregman geometric perspective of regularized regret over the quotient space \mathbbR^K / \mathrmspan(\\\mathbf1\\) to resolve the translation invariance of the inverse mirror map. By combining this dual geometric formulation with the van Trees inequality, we reduce the adversarial online learning to Bayesian estimation, offering a unified framework for establishing logarithmic regret lower bounds.
PaperID: 7001, Poster
Authors:
Anshu Raina, Peyman Razaghi, Yuankai Chen, Cheng Yao, Yao Fu, Vidushi Goyal, Devang Patel, Wen Xie, Zhenyu Gu, Emad BarsoumAbstract: Configuring distributed LLM training across hundreds of GPUs requires jointly optimizing tensor, pipeline, expert, context, and data parallelism alongside micro-batch size, recomputation, sharding, and overlap — a combinatorial space where the boundary between memory-legal and OOM can hinge on a single byte of per-element activation accounting. Today, teams navigate it through expensive multi-node trial and error, burning GPU-hours on crashes and suboptimal recipes. We present Primus Projection, which closes this loop with two coupled components. First, a hybrid projection engine that measures what it can and simulates what it cannot: GPU microbenchmarks on as few as a single device capture compute kernels and any in-scope communication; a sub-node downscale–measure–upscale workflow analytically restores out-of-scope collectives, FSDP overlap, and the full pipeline schedule at the multi-node target. A pure simulation mode runs entirely on CPU, enabling pre-silicon planning without any GPU. Second, a tuning agent treats the projection engine as a fast scoring oracle — seconds per candidate instead of tens of minutes — and uses it to drive an intelligent search over the configuration space. A deterministic seed planner sweeps single axes; an LLM investigation loop proposes cross-axis combinations the planner cannot reach, with an analytic memory pre-filter that rejects infeasible candidates before any GPU time is spent. Validated on Llama 3.1 (dense) and Mixtral 8×22B (MoE) across two GPU generations, all projections stay within ~10% of measured multi-node throughput. In a Mixtral 8×22B case study, the agent discovers a configuration delivering +27% throughput over the published 4-node BF16 reference at the same scale — in under 30 minutes of single-node exploration, with no hand-written configs and no full-cluster profiling pass.
PaperID: 7002, Poster
Abstract: Surprise and episodic context are believed to govern how animals update their memories, with impact on how prior knowledge is retained. Inspired by this parallel, we empirically investigate how surprise and context affect knowledge retention in large language models under supervised fine-tuning. We construct 15 update datasets (230k samples, 13 newly created) spanning surprise levels across facts, ethics, and code, varying how each update is paired with an explicit temporal frame. Across GPT-2-XL, Mistral-7B, Llama-3-8B, and GPT-4.1 variants, evaluated on a held-out cross-domain sentinel set (1.8M LLM judgments), we find: (i) without context, contradictory updates (highest surprise) overwrite the targeted memory and damage entirely unrelated knowledge, sometimes producing near-complete factual collapse, with damage scaling with update surprise, and spilling across domains; (ii) pairing each update with an explicit temporal frame at the prompt (pre-contextualization) appears to create two coexisting traces, preserving the original under the bare prompt while acquiring the new one under the contextualized prompt, and rescuing unrelated knowledge in the process. Extending prior work on emergent misalignment, we characterize a more general pattern we call habit transfer: overwriting a related fact induces transferable response patterns on unrelated prompts (e.g., ``code bleeding'', i.e., answering factual questions with code, rising from 4% to 73% along the surprise axis). Taken together, these results suggest explicit context as a shield against catastrophic forgetting in LLMs and point to a practical direction for safer post-training: data-side curation that contextualizes high-surprise updates rather than presenting them as flat prompt-continuation pairs.
Abstract: Scaling robot policy learning is bottlenecked by the cost of collecting demonstrations: datasets at modern scale require thousands of skilled-operator hours on dedicated robot hardware. The language description paired with these demonstrations does not face the same bottleneck---a single short task label is cheap, but it also leaves implicit spatial relations, object interactions, embodiment state, and subgoal structure that the pixels already contain. We treat \emphlanguage density as a cheap lever for amplifying signal in a fixed demonstration corpus; whereas prior dense-language work in robotics commits to a single caption style, we instead ask which kind of dense language helps each task and learn to deliver it at deployment. We realize this as DeMiAn (Dense Multi-aspect Annotation) in two stages. First, an automatic VLM pipeline re-labels each segment of an existing demonstration along four complementary aspects---\emphphysical motion, \emphscene composition, \empharm pose, and segment-level \emphreasoning---each surfacing a distinct kind of structure that a one-line task label omits. Second, a small learned \emphinstructor, trained via supervised fine-tuning, maps the natural language task description and an initial scene snapshot to a task-appropriate annotation and runs asynchronously alongside the action policy, hiding generation latency behind the rollout. Applied to over 1M robot manipulation and 50K EgoVerse human-egocentric videos, with no new demonstrations collected, DeMiAn delivers four findings on a VLA action policy and a video-based world-action model: i) the learned instructor lifts RoboCasa success by 5 points over the no-annotation baseline, within 3 points of a per-task oracle; (ii) that oracle is non-trivial---no fixed aspect dominates, and peak performance requires selecting the right annotation aspect per task; (iii) the trained system extends usefully to composite tasks under subgoal-driven prompt switching, and to OOD scenes and objects; and (iv) dense annotation improves the compute-performance frontier in both mid-training and post-training, making re-annotation a practical scaling lever for robot policy learning.
PaperID: 7004, Poster
Authors: Haoran Xu, Wei Yuan, Xiao Han
Abstract: Flow matching is increasingly used beyond open-ended generation, including prediction tasks whose outputs are effectively deterministic, yet it remains unclear why flow matching helps when diversity is not the objective. We study this question in histology-conditioned spatial transcriptomics (ST) prediction, a challenging sparse high-dimensional multimodal task where direct regression with pathology foundation features is already competitive. We disentangle several possible explanations and find that the gain does not come from simply adding a flow objective, injecting noisy target-side states, or increasing architectural complexity. We instead identify intermediate flow states as paired, time-indexed target-side signals: they contain condition-aligned information about the target, but are mixed with source noise and residual variation. Their benefit emerges only when this information is selectively routed into histology representations. Guided by this mechanism, we instantiate conditional flow matching with a simple gated self-attention to learn interactions between histology tokens and noisy ST states. Across high-resolution ST datasets, the model improves spatial correlation over regression and flow-based baselines, with stable convergence and favorable scaling under larger and sparser gene panels. These results establish a design principle for deterministic multimodal prediction: flow matching helps when intermediate target states are treated as structured cross-modal supervision and selectively integrated into conditional representations. Code is available at https://anonymous.4open.science/r/Understanding-When-Flow-Matching-Helps-Deterministic-Multimodal-Prediction-B426/.
Authors:
Beicheng Xu, Weitong Qian, Lingching Tung, Yupeng Lu, Bin CUIAbstract: To lower the expertise barrier in machine learning, the AutoML community has focused on the CASH problem, which jointly automates algorithm selection and hyperparameter tuning. While traditional methods like Bayesian Optimization (BO) struggle with cold-start issues, Large Language Models (LLMs) can mitigate these through semantic priors. However, existing LLM-based optimizers generalize poorly to high-dimensional, structured CASH spaces. In this paper, we propose LB-MCTS, a trajectory-structured optimization framework that uses a Monte Carlo Tree Search tree as a shared state for algorithm selection, hyperparameter refinement, and BO–LLM proposer synergy. Within this shared state, BO provides algorithm-specific surrogate modeling for quantitative search, while the LLM exploits path-aware selective memory to generate semantic proposals and reflections. As the surrogate model improves, a reliability-aware proposer policy adaptively shifts from LLM-driven to BO-driven proposals within a unified search trajectory. Experiments on 104 AMLB datasets demonstrate that LB-MCTS consistently outperforms BO-based, LLM-based, and hybrid baselines.
Abstract: rocess Reward Models (PRMs) have been proven to be highly effective in guiding test-time scaling (TTS) methods, which significantly boost the capabilities of LLM-based multi-agent systems. However, existing PRMs are text-based: they re-encode the entire trajectory text from scratch. In long multi-agent rollouts, the scoring cost, growing quadratically with respect to sequence length L, creates a severe computational bottleneck, severely limiting PRMs' application in long-context scenarios. To resolve this, we introduce KV-PRM, a highly efficient process reward model that eliminates the heavy text re-encoding by directly reading the KV cache produced naturally during the LLM's generation phase. By processing a single "verify token" against the pre-existing KV cache, KV-PRM reduces the scoring cost from O(L^2) to O(L). We formally prove that the KV cache contains strictly greater information capacity than text, and is more efficient for downstream reward modeling. Empirically, across the MATH, GSM8K, and AIME benchmarks, KV-PRM matches or strictly outperforms text-PRMs under various TTS methods such as Beam Search, MCTS, and Weighted Voting, with up to a 5,000× reduction in scoring FLOPs, a 37× reduction in latency, and a 34× reduction in per-sequence memory footprint compared to text-based PRMs.
PaperID: 7007, Poster
Abstract: Large language models (LLMs) have emerged as powerful tools for automated heuristic design (AHD), enabling iterative generation and refinement of heuristics. However, the dominant paradigm embeds LLMs as narrow, fixed components, such as crossover or mutation, within heavily hand‑engineered evolutionary frameworks. We argue this misapprehends LLMs. It treats them as specialized tools rather than general reasoners, constrains them to low‑level operations, and underutilizes their autonomy. Moreover, the extensive human priors in these frameworks violate the bitter lesson principle that general methods scaling with computation surpass hand‑crafted solutions. This raises a key question: To address this, we propose metrics for AHD framework handcraftedness (AHI) and intelligence conversion efficiency (ICE). Evaluating ten LLMs across three challenging optimization problems, we obtain a notable finding that frameworks with fewer human priors consistently yield higher ICE. Based on this finding, we propose SimpleEvol, which removes nearly all structural constraints and allows the LLM to operate autonomously. SimpleEvol consistently achieves the highest ICE, often by a large margin. Our results challenge the trend toward complex AHD pipelines and point to a simpler and more model‑centric alternative, suggesting that reducing human priors is a more effective strategy to scale up with model intelligence.
Abstract: Long-horizon language agents must operate under limited runtime memory, yet existing memory mechanisms often organize experience around descriptive criteria—relevance, salience, or summary quality. For an agent, however, memory is valuable not because it faithfully describes the past, but because it preserves the distinctions between histories that must remain separated under a fixed budget to support good decisions. We cast this as a decision-centric rate-distortion problem, measuring memory quality by the loss in achievable decision quality induced by compression. This yields an exact forgetting boundary for what can be safely forgotten, and a memory-distortion frontier characterizing the optimal tradeoff between memory budget and decision quality. Motivated by this decision-centric view of memory, we propose , an online memory learner that refines its partition only when data certify that a shared state would induce decision conflict, and prove near-minimax regret guarantees. On both controlled synthetic diagnostics and long-horizon conversational benchmarks, DeMem yields consistent gains under the same runtime budget, supporting the principle that memory should preserve the distinctions that matter for decisions, not descriptions.
PaperID: 7009, Poster
Authors: Junkai Liu, Le Zhang
Abstract: Self-supervised learning on 3D medical images commonly relies on random sub-volume sampling during pretraining. However, random cropping implicitly treats all spatial regions as equally worth observing, despite anatomical priors and spatial redundancy that make regions inherently unequal in their value for representation learning. To this end, we first show through controlled crop-quality manipulations that better sub-volumes lead to better downstream transfer, revealing the observation policy as an overlooked bottleneck in 3D medical SSL. Motivated by this finding, we revisit sub-volume sampling as a where-to-look problem. We propose VolumeProbe, a plug-and-play framework that replaces random cropping with an offline coarse-to-fine probe bank guided by anatomy-awareness, informativeness, and diversity. We instantiate VolumeProbe as a lightweight Lite variant using geometric and image-statistical cues, and a diagnostic Feat variant using frozen visual features, to examine whether lightweight volume-intrinsic priors suffice or richer semantic cues reliably provide additional value. Extensive experiments across multiple 3D medical SSL backbones and downstream tasks demonstrate that VolumeProbe-Lite consistently improves transfer performance over random sampling. Mechanistic analyses further attribute these gains to more anatomically plausible and informative observations, reduced spatial redundancy, more efficient learning dynamics, and more structured representations.
PaperID: 7010, Poster
Abstract: The computational complexity of anonymous games has been a central theme in algorithmic game theory, with foundational contributions by Daskalakis and Papadimitriou [2015], Goldberg and Turchetta [2017], and Cheng et al. [2017] establishing that the landscape of equilibrium computation admits a PTAS. In this paper, we investigate the class of n-player anonymous games with linear utilities and two actions. While finding an equilibrium in anonymous games is PPAD-hard for arbitrary utility functions—even with a constant number of actions [Chen et al., 2015]—we show that this linear structure allows for significant algorithmic improvements. Exploiting the fact that payoffs depend solely on the first moment of the players’ strategy distribution, we provide the first algorithm for computing exact Nash equilibria that runs in time polynomial in the number of players n. We achieve this by establishing a novel combinatorial characterization of the Nash equilibria. Finally, we establish a novel connection between Nash equilibria in these games and non-convex non-concave min-max optimization. Despite the technical challenges posed by this structural property, we design a sequential best-response dynamic that provably converges to an \epsilon-Nash equilibrium in \mathcalO(\frac1\epsilon) steps.
PaperID: 7011, Poster
Abstract: Generative protein models are increasingly used for functional sequence design, yet the process by which biological information is organized during generation remains poorly understood. Diffusion protein language models provide a tractable setting for this question, as sequence generation proceeds through an explicit denoising trajectory from highly corrupted states to complete proteins. This trajectory enables a temporal view of interpretability: beyond identifying which biological signals are represented, it allows us to examine when these signals emerge and when specific residues become committed. Here, we analyze DPLM representations using sparse autoencoders trained across layers and noise levels. The resulting features capture biologically meaningful signals at both residue and protein scales, align with functional annotations, and preserve downstream biological information under reconstruction. Following these features along the denoising trajectory reveals a consistent functional ordering: catalytic-enriched features pre-activate at still-masked catalytic positions before residue identity is resolved, and catalytic residues are subsequently recovered earlier by the iterative denoiser. This prioritization remains significant after controlling for prediction difficulty, sequence context, amino-acid identity, structural environment, and evolutionary conservation, and is not reproduced by a random-feature null. Across ProteinGym deep-mutational-scanning assays, residues recovered earlier during denoising are also more mutation-sensitive. These results reveal a temporal organization of biological information in diffusion-based protein generation, in which functionally important residues are not only represented, but preferentially committed during sequence formation.
PaperID: 7012, Poster
Authors: Ruichen Ma, Yicai Chen, Pujun Zhou, Yue Zuo, Chongyang Li, Shuang Liu, Shaogang Hu, Guanchao Qiao
Abstract: Spiking neural networks (SNNs) offer a compelling pathway toward energy-efficient artificial intelligence, fundamentally anchored in presynaptic input-driven sparsity. However, a substantial yet chronically overlooked source of computational overhead persists: the continuous subthreshold dynamics of postsynaptic neurons. Quantitative analysis reveals that up to 44% of neurons in an ImageNet-trained ResNet-18 model reside in a deeply inhibited state, imposing a persistent computational burden despite possessing a negligible probability of firing. To eradicate this inefficiency, the Dormant Spiking Neuron (DSN) model is introduced. By incorporating a state-driven dynamic gate, the DSN allows a neuron to enter a dormant state when its membrane potential falls below a predefined negative threshold, dynamically pruning redundant computations at the instance level. Crucially, this dormancy mechanism is directly integrated into the neuronal dynamics and trained end-to-end via backpropagation through time using surrogate gradients, enabling the network to actively learn a robust dormancy policy that adaptively preserves critical information flow. Extensive experiments across static and neuromorphic datasets demonstrate that the DSN framework consistently reduces computational energy by 40% to over 70%. Notably, on the large-scale ImageNet dataset, the DSN outperforms state-of-the-art static pruning techniques by achieving equivalent energy savings while simultaneously improving upon the baseline accuracy by 1.34%. By establishing a trainable, state-driven sparsity paradigm orthogonal to conventional input-driven methods, the DSN unlocks synergistic efficiency gains, presenting a highly potent approach for advancing neuromorphic computing. Source code will be made publicly available.
PaperID: 7013, Poster
Authors: Nam Nguyen, Tuan Dam
Abstract: Top-two algorithms are simple and effective for fixed-confidence best-arm identification, but their sharp non-asymptotic behavior is still not well understood. We study this problem for Bernoulli bandits through \beta-EB-TCI, the empirical-best top-two rule of Jourdan et al. (2022), whose challenger is chosen using a Bernoulli transportation cost with a logarithmic count penalty. We prove that, after the empirical leader has become the true best arm and its sampling fraction stays close to \beta, the stopping time is T_\beta^\star(\mu)\log(1/\delta) up to lower-order concentration terms. We also show that, in this regime, every challenger is sampled linearly often. Thus, for the original algorithm without forced exploration, the main remaining difficulty is to control when the empirical leader becomes permanently correct. These results imply a non-asymptotic high-probability bound for all Bernoulli instances with a unique best arm. If the algorithm satisfies a finite-mean sufficient-exploration condition, the bound further yields the sharp expected sample complexity. In particular, this gives the sharp expectation result for the unguarded Bernoulli rule when all arm means are pairwise distinct, using the sufficient-exploration result of Jourdan et al. (2022). Finally, if we add a mild forced-exploration rule that contributes only O(\sqrtKt) pulls up to time t, we obtain a self-contained expected sample-complexity theorem for any number of arms under the unique-best-arm assumption. We also identify a limitation of proof strategies that try to handle equal suboptimal means through a single index-comparison argument.
PaperID: 7014, Poster
Abstract: Matrix-vector products (MVPs) are a key primitive in numerical linear algebra and machine learning. However, the naive quadratic running time for exact computation is prohibitive for large matrices, motivating approximate methods. In this paper, we study approximate MVPs under a natural ``lightness'' assumption, which bounds the total \ell_1-mass of a n × n input matrix A by \gamma n. For general matrices, we give an algorithm for approximating Ax up to additive \epsilon \|x\|_2 error in O(\gamma n^1.5/\epsilon) query time and polynomial preprocessing. We extend this algorithm to kernel matrices by establishing a black-box reduction to Kernel Density Estimation data structures and analyzing a noisy entry sampling scheme. For kernel matrices, our algorithm does not require any preprocessing and improves the best-known running times for Gaussian kernel matrices of [Indyk, Kapralov, Sheth, Wagner; ICLR `25] by polynomial factors in n and 1/\epsilon, and also provides the first algorithms leveraging the lightness assumption for other kernel functions as well.
PaperID: 7015, Poster
Abstract: Spatial interpolation from sparse irregular observations is a recurring task across the environmental and health sciences. Neural processes amortise this inference into a single forward pass at test time. Joint-predictive variants capture correlations across query locations but typically rely on self-attention over context and query points, incurring O(K^2) compute in the combined point count K and capping practical use at a few thousand points. State-space models offer a linear-time alternative, but existing state-space NPs compress the context into a single hidden state before reading queries, reducing the information available to the decoder. We propose MambaFlowNP, a joint-predictive NP with linear compute trained with a conditional flow-matching objective. A multi-directional Mamba-2 scan reads context and query tokens in a single shared sequence, so each query's velocity is read from its own scan position rather than from a compressed context summary. MambaFlowNP is competitive with the strongest attention-based joint-predictive baselines on synthetic Gaussian-process benchmarks, and across regional and continental-scale NOAA temperature interpolation tasks it attains the lowest RMSE on every tier and the lowest CRPS at the California and continental US tiers, while running 170--760× faster than the strongest joint-predictive baseline. A single forward pass scales to 10^6 points on one GPU in under a second. This scalability unlocks amortised joint-predictive inference for continental-scale, irregularly sampled sensor networks, spanning climate monitoring to disease vector surveillance.
PaperID: 7016, Poster
Abstract: Barrier certificates provide formal safety guarantees for neural network controlled systems by ensuring that system trajectories avoid unsafe regions over time. However, after deployment, previously unseen obstacles or constraint changes may introduce new unsafe regions, invalidating the original certificate and potentially leading to unsafe behavior. We address this setting with adaptive piecewise barrier certificates, which preserve the original certificate and policy in unchanged regions while introducing localized updates around newly emerged unsafe regions. The central question is how to perform such updates efficiently during system execution without retraining from scratch. Our key idea is to cast deployment-time adaptation as a localized certificate repair problem and reduce it to an optimization task that leverages the learned structure of previous barrier certificates, enabling fast updates without global retraining. We further design a staged repair framework that progressively transitions from localized repair to joint policy–certificate updates and, when necessary, localized retraining. Empirical results on four benchmarks with five new obstacles show that our method succeeds in all cases and completes within tens of seconds on average, whereas retraining from scratch fails in most cases. On instances where retraining succeeds, our method is 17.3× faster on average. It incurs only 16.2% performance degradation relative to the original policy.
PaperID: 7017, Poster
Abstract: Standard Message Passing Neural Networks (MPNNs) are at most 1-WL expressive, which has spurred the development of higher-order models. In particular, many higher-order designs rely on extra global structural features (e.g., positional encodings or pairwise distances), whereas the k-WL hierarchy offers a principled and quantifiable path to increased expressiveness. However, directly simulating k-WL is computationally prohibitive, and existing k-WL variants often sacrifice expressiveness to gain efficiency, failing to capture critical structural distinctions. To overcome these shortcomings, we propose a higher-order graph learning framework based on walk-induced lifted states with structure-aware state encoding. We further design a controllably sparsified and localized k-FWL-style aggregation scheme, enabling graph-level higher-order representations via simple message passing. On the recent and challenging BREC expressiveness benchmark, our model achieves state-of-the-art total distinguishing accuracy among compared higher-order GNN baselines, outperforming strong 3-WL baselines, and remains competitive on real-world graph classification tasks.
PaperID: 7018, Poster
Abstract: Normalized stochastic gradient descent (Normalized SGD) is a classical scale-invariant method that updates only along the gradient direction. While Normalized SGD has recently been analyzed in non-convex stochastic settings, its convex theory remains much less developed. We address this gap by studying normalized gradient methods for smooth convex optimization. First, we obtain the theory of deterministic approaches, including first convergence results for momentum variants. Next, we establish first high-probability convergence for mini-batch Normalized SGD with and without momentum, under heavy-tailed stochastic noise with a bounded \alpha-central moment. Addressing the key challenge -- bias introduced by normalizing a noisy mini-batch gradient -- we demonstrate that incorporation of a new bias lemma and modified inductive proof allows to control it. As a result, we obtain first high-probability convergence of Normalized SGD in function value, together with explicit oracle complexity bounds.
PaperID: 7019, Poster
Authors:
Jiayu Ying, Qijian Tian, Ruijie Xu, Xinnan Zhu, Daoguo Dong, Jiachen Xu, Xin TanAbstract: Advancing spatial intelligence in Multimodal Large Language Models (MLLMs) is bottlenecked by the scarcity of complex, scalable 3D question-answer (QA) data. While manual annotation is labor-intensive, directly utilizing LLMs to synthesize these QA pairs often fails due to their inherent deficiencies in spatial and geometric computation. We introduce Exemplar2VQA, a scalable exemplar-driven visual question answering generation framework that rapidly synthesizes large-scale spatial QA pairs in simulated environments via multi-agent coding. By equipping collaborative agents with a meticulously designed library of geometric utilities, Exemplar2VQA bypasses LLMs' spatial reasoning flaws through deterministic code execution. Crucially, the framework exhibits remarkable versatility: taking diverse static object-centric spatial query templates as exemplars, it seamlessly and autonomously scales them into massive, high-fidelity synthetic datasets. Fine-tuning Qwen2.5-VL (3B/7B) exclusively on Exemplar2VQA-generated synthetic indoor data yields significant performance improvements across various diverse benchmarks. Furthermore, its effectiveness is not limited to in-domain indoor datasets but also robustly extends to outdoor and mixed-scene benchmarks. These results establish Exemplar2VQA as a scalable and powerful paradigm for bridging the sim-to-real gap in Embodied AI. The code will be available upon acceptance.
Abstract: Recovering a low-CP-rank tensor from noisy linear measurements is a central challenge in high-dimensional data analysis, with applications spanning tensor PCA, scalar-on-tensor regression, tensor-on-tensor regression, and beyond. We exploit the intrinsic geometry of rank-one tensors by casting the recovery task as an optimization problem over the Segre manifold, the smooth Riemannian manifold of rank-one tensors. This geometric viewpoint yields two powerful algorithms: Riemannian Gradient Descent (RGD) and Riemannian Gauss-Newton (RGN), each of which preserves feasibility at every iteration. Under mild noise assumptions, we prove that RGD converges at a local linear rate, while RGN exhibits an initial local quadratic convergence phase that transitions to a linear rate as the iterates approach the statistical noise floor. Extensive synthetic experiments validate these convergence guarantees and demonstrate the practical effectiveness of our methods.
PaperID: 7021, Poster
Authors: Honghao Jia, Zhenxing Li, Jialiang Guo, Jiacheng Gan
Abstract: The KV cache is a major memory bottleneck in LLM inference, especially under large batch sizes and long contexts. While quantization is a promising solution, existing static group-wise quantization methods are constrained by a trade-off between metadata overhead and accuracy: increasing the group size reduces per-group metadata (e.g., zero-points, scaling factors, and outlier indices), but exacerbates quantization errors under the sinusoidal-like temporal drift of post-RoPE outlier channels; conversely, using smaller groups improves accuracy but incurs substantial metadata overhead. In this paper, we propose ADKV, a low-overhead adaptive delta quantization framework for KV cache in LLM inference, which leverages the temporal structure of the KV cache. ADKV tracks per-channel zero-points and scaling factors via online EMA, adapting to the sinusoidal-like temporal drift of post-RoPE outlier channels, thereby eliminating the rigidity of static group-wise quantization. This adaptive mechanism further enables delta coding to embed metadata directly into the quantized sequence, decoupling metadata overhead from context length. To minimize the MSE of attention outputs, we design an SGD-based calibration algorithm that models ADKV as an RNN. Experiments on generation and long-context benchmarks show that, while matching the accuracy of static group-wise quantization, ADKV reduces the metadata memory footprint by ~ 12.8×. A custom fused ADKV-GEMV CUDA kernel optimized for memory-bound autoregressive generation achieves ~ 1.96× throughput over the FP16 baseline.
Abstract: Data selection for language model pretraining faces a fundamental tension between quality and diversity. While quality filtering is empirically effective, it often induces diversity collapse: by favoring texts similar to high-quality reference corpora (e.g., educational or QA-style data), it systematically excludes valuable data from underrepresented domains. In contrast, diversified selection preserves domain balance and encourages robust downstream performance, yet existing methods either focus on coverage-oriented objectives that indirectly enhance diversity, or directly optimize for diversity via costly covariance matrix recomputation that limits scalability. To address these issues, we introduce Leverage Score Sampling (Lev), which iteratively selects samples that maximally expand the determinantal volume of the embedded data via leverage scores, a computationally efficient criterion that eliminates matrix recomputation and enables scalable selection. Empirically, Lev delivers up to 65× speedup and improves dataset diversity, measured by the Vendi score, by 9.2% over the strong diversification baseline DiSF. On CommonCrawl (CC) web data selection, Lev improves accuracy across seven downstream tasks by up to 1.31% over existing baselines. For domains where robust quality criteria are inherently difficult to define (e.g., code), Lev serves as an effective unsupervised curation alternative: on StarCoderData, the selected subset reduces bits-per-byte by 3.08% over DiSF. Notably, we uncover a cross-domain collapse of quality filtering: CC data filtered by DCLM-fastText fail to retain sufficient code-related content, yielding inferior code performance relative to Lev-selected data. These findings advocate for integrating diversity-aware practices into quality filtering for more effective data curation in language model pretraining.
PaperID: 7023, Poster
Authors:
Yizhak Ben-Shabat, Jiahao Zhang, Joseph Liu, Seonghyeon Moon, Young-Yoon Lee, Haomiao Jiang, Oren Jacob, Mubbasir KapadiaAbstract: While text-driven body motion generation has matured, pure-hand synthesis remains bottlenecked by data. Existing datasets often treat hands as rigid end-effectors, lack fine-grained captions, or are too limited in scale to support robust generalization. We introduce RoMo Hands, an in-the-wild, text-paired dataset of ~542K sequences (884 hours, 95.5M frames at 30\,fps) which is an order of magnitude larger than prior work. To ensure high-quality supervision, a stringent dexterity filter discards low-articulation sequences, focusing the distribution on complex bimanual movement. Each clip is paired with five captions at increasing granularities from global semantic tags to atomic finger states, organized under a hierarchical taxonomy. To bridge the gap between body and hand architectures, we propose a novel 393-D bimanual representation that adapts redundant body-motion paradigms to the high-DoF kinematics of hands. Systematic benchmarking of state-of-the-art diffusion, masked-token, and autoregressive generators reveals that body-centric design choices do not transfer cleanly to hands, exposing a sharp divergence between semantic alignment and kinematic plausibility. Our dataset, representation, and benchmark establish a rigorous foundation for the next generation of dexterous motion synthesis.
PaperID: 7024, Poster
Authors:
Wang Yifei, Wei Tian, WENHUA BAO, Jia Zhou, Hanshi Wang, Fudong Ge, Jianhua Wu, Shuai Yuan, Zhipeng Zhang, Xingliang Liu, Xianming Zeng, Jianyun Xu, Bingzhao GaoAbstract: Driving-scene reconstruction via feed-forward Gaussian splatting from sparse multi-view images remains challenging due to two factors. First, cross-camera photometric inconsistency—caused by variations in exposure, ISP, and illumination—can be incorrectly absorbed into Gaussian geometry. Second, unreliable local geometric supervision arises from normal-based constraints near depth discontinuities. To address these issues, we present , a multi-view consistent feed-forward Gaussian splatting framework for driving scenes. MCSplat constructs a canonical appearance space from depth-guided multi-camera correspondences to decouple geometry learning from camera-specific appearance. A multi-scale Bilateral Grid Transformer restores raw camera appearance from canonical renderings, enabling photometric supervision without contaminating Gaussian geometry. To improve local geometric reliability, we estimate Fisher-information based uncertainty from photometric sensitivity and use it to adaptively weight surface-oriented self-supervision. Experiments on the Waymo Dataset show that MCSplat achieves state-of-the-art performance on both the full validation set and an appearance-challenging subset with strong cross-camera photometric inconsistency, demonstrating improved reconstruction quality and more consistent novel-view geometry. Code will be released.
Abstract: Many complex networks exhibit hierarchical, tree-like structures, making hyperbolic space a natural candidate wherein to learn representations of them. Based on this observation, Hyperbolic Graph Neural Networks (HGNNs) have been widely adopted as a principled choice for representation learning on tree-like graphs. In this work, we question this paradigm by proposing the additional condition of geometry–task alignment, i.e., whether the metric structure of the target follows that of the input graph. We theoretically and empirically demonstrate the capability of HGNNs to recover low-distortion representations on regression problems, and show that their geometric inductive bias becomes helpful when the problem requires preserving metric structure. By jointly analyzing predictive performance and embedding distortion, we further show that HGNNs gain an advantage on link prediction, a naturally geometry-aligned task, whereas this advantage largely disappears on standard node classification benchmarks, which are typically not geometry-aligned. Overall, our findings shift the focus from only asking , showing that HGNNs consistently outperform Euclidean models under such alignment, while their advantage vanishes otherwise.
Abstract: Language-Assisted Image Clustering (LAIC) augments the input images with additional textual information with the help of vision-language models (VLMs) to improve clustering performance. Despite recent progress, existing LAIC methods mainly construct image-text pairs for text-modality utilization, which overlook two issues: (i) textual features constructed for images are highly similar, leading to weak inter-class discriminability; (ii) the training is restricted to pre-built image-text alignments, limiting the potential for better utilization of the text modality. To address these issues, we propose a new LAIC framework with two complementary components. First, we exploit cross-modal relations to generate more discriminative self-supervision signals for clustering, which is compatible with the pre-training mechanisms of VLMs such as CLIP. Second, we learn category-wise continuous semantic centers via prompt learning to produce the final clustering assignments. Extensive experiments on eight benchmark datasets demonstrate that our method achieves an average improvement of
PaperID: 7027, Poster
Abstract: Accurate numerical solutions of partial differential equations (PDEs) are crucial in numerous science and engineering applications. In this work, we introduce a novel neural PDE solver named AFDONet, which incorporates neural operator learning and adaptive Fourier decomposition (AFD) theory for the first time into a specifically designed variational autoencoder (VAE) structure, to solve a general class of nonlinear PDEs on smooth manifolds. AFDONet is the first neural PDE solver whose architectural and component design is fully guided by an established mathematical framework (in this case, AFD theory), turning neural operator design from an art to a science. Thus, AFDONet also exhibits exceptional mathematical explainability and groundness, and enjoys several desired properties. Furthermore, AFDONet achieves outstanding solution accuracy and competitive computational efficiency in several benchmark problems. In particular, thanks to its deep connections with AFD theory, AFDONet shows superior performance in solving PDEs on i) arbitrary (Riemannian) manifolds, and ii) datasets with sharp gradients. Overall, this work presents a new paradigm for designing explainable neural operator frameworks.
PaperID: 7028, Poster
Abstract: Learning high-quality multimodal entity representations is essential for advancing reasoning tasks such as multimodal knowledge graph completion (MMKGC). However, rigid alignment in existing fusion strategies can bias representations toward a dominant anchor modality, weakening fine-grained complementary cues and propagating noisy signals into relational reasoning. To address these limitations, we propose MANoR, a Multimodal Autonomous Negotiation framework for calibrated Representation learning. MANoR treats modalities as autonomous semantic agents that negotiate before fusion. Specifically, MANoR preserves modality-specific semantics through centroid-guided subspace contraction, which reduces unstable centroid-relative dispersion while retaining entity-level residual semantics. On this stabilized basis, negotiated attention enables selective cross-modal interaction without collapsing modality boundaries. As modalities can still differ in reliability after interaction, uncertainty calibrated fusion estimates dimension-level aleatoric uncertainty and modality-level reliability to compute relation-aware fusion weights. In this way, MANoR couples modality autonomy, selective interaction, and calibrated integration within a unified representation learning framework. Experiments on MKG-W, MKG-Y, TIVA, and KVC16K show that MANoR consistently outperforms strong MMKGC baselines, achieving a relative Hits@1 gain of up to 17.72% on KVC16K.
PaperID: 7029, Poster
Abstract: Pruning and quantization are two dominant techniques that effectively address the computational and storage burdens of large language model (LLM) inference on edge devices. Recently, one-shot pruning has gained particular attention for its ability to identify weight supports via optimization without retraining. However, integrating such methods seamlessly with quantization for further model lightweighting is non-trivial. The discreteness of quantization shatters their required variable continuity, inevitably collapsing this integration into a decoupled pruning-quantization pipeline with inherently suboptimal outcomes. To tackle this problem, we propose Quantization-aware One-shot Pruning (QOP), which directly optimizes the quantized weights under sparsity and discreteness constraints, thereby explicitly capturing the impact of quantization on pruning within its objective. Specifically, QOP generalizes the alternating direction method of multipliers to sparse-constrained discrete optimization, enabling the identification of the high-quality support and the update of quantized weights. Theoretically, via the well-conditioned approximation obtained by a slight perturbation, we establish convergence to a stationary feasible point and provide the convergence rate. Extensive experiments across various LLMs, sparsity levels, and quantization settings demonstrate that QOP consistently outperforms existing baselines in terms of both average accuracy and perplexity under weight quantization.
PaperID: 7030, Poster
Authors: Enes Arda, Atilla Eryilmaz
Abstract: Policy mirror descent (PMD) has a mature finite-time theory in discounted Markov decision processes (MDPs), but less is known in the average-reward setting, a more natural objective for many control applications. We give a finite-time, finite-sample analysis of PMD in ergodic average-reward MDPs built around a single master recursion that governs convergence under any gradient proxy, without external regularization. Its specializations yield linear rates (with a superlinear regime in the well-conditioned case) for exact, inexact-tabular, and linear function approximation (LFA) updates. We complement these convergence results with end-to-end sample complexities of order t_\mathrmmix^3/\varepsilon^2 in both tabular (|\mathcal S||\mathcal A|-dependent) and LFA (d-dependent) settings. The LFA rate sharpens the prior best t_\mathrmmix^5 mixing dependence to t_\mathrmmix^3, and matching information-theoretic lower bounds establish that the t_\mathrmmix^3/\varepsilon^2 critic core is unimprovable in both settings.
PaperID: 7031, Poster
Authors: Rahul K Sharma, Feng Liu, Xiaochen Xian
Abstract: Multimodal electrophysiological recordings provide noisy, partially observed measurements of latent neural dynamics critical for tasks like brain state estimation and neurological diagnosis. In joint EEG-intracranial EEG (iEEG) recordings, scalp EEG offers broad but spatially mixed coverage, whereas iEEG provides high-fidelity but spatially sparse measurements. This creates a partial-observation problem: directed source-level interactions must be inferred from heterogeneous sensors with incomplete coverage. A second challenge is nonstationarity: latent neural dynamics and directed interactions among brain regions can change abruptly across distinct regimes (e.g., pre-seizure, ictal, and post-seizure states). The third challenge is sparsity of significant source-level interactions compared to the high-dimensionality of the brain networks. Standard models that assume single-mode, fully observed, or stationary may therefore fail to capture regime-specific activity or spurious couplings. We propose the Sparse Multimodal Switching State-Space Model (S^4M^3), which fuses EEG and iEEG as complementary observations of a shared source-space latent process. Dynamics switch between regimes governed by separate sparse directed transition matrices, jointly addressing multimodal fusion, nonstationarity, and sparse network estimation. The estimation of the model uses the expectation-maximization (EM) algorithm, where the E-step applies the Kim filter-RTS smoother to estimate latent states and regime probabilities, and the M-step imposes elastic-net penalization on off-diagonal elements, promoting sparse inter-regional couplings while preserving self-dynamics. The learned transition matrices facilitates the assessment of regime-dependent edge changes. To address the estimation bias of elastic-net sparsity penalty, we separate edge selection from edge assessment: selected edges are refit by regime-weighted least squares, followed by Wald tests with Benjamini-Hochberg correction. On synthetic benchmarks, S^4M^3 improves Edge F1 by approximately 25% over the strongest stationary baseline and 64% over two-stage source-connectivity pipelines. On a real EEG-iEEG epilepsy case study, S^4M^3 identifies sparse hippocampal and parahippocampal outgoing pathways concordant with the clinician-confirmed seizure-onset zone, without using clinical labels.
PaperID: 7032, Poster
Abstract: Sharpness-Aware Minimization (SAM) has been reported to improve generalization while leaking more membership privacy than SGD. In this paper, we aim to theoretically explain this phenomenon under long-tailed distributions. In long-tailed settings, typical examples are explained by dominant majority features, whereas atypical tail examples depend on weaker minority features. We theoretically show that SAM moves farther than gradient descent along a minority-preserving trajectory under an overparameterized model. This bias improves the embedding-space separation between typical and atypical subclasses, reducing their representational entanglement. However, because the minority cluster is learned from a few atypical training examples, stronger minority alignment can also amplify finite-sample information from those examples. Consequently, atypical members become more aligned with the empirical atypical prototype than atypical nonmembers, creating a membership signal that can be exploited by the attacker. Finally, we validate our findings empirically. Our analysis explains the generalization--privacy tension of SAM through a single mechanism: in long-tailed learning, the tail wags the model.
PaperID: 7033, Poster
Abstract: Point cloud upsampling is essential for 3D applications but remains challenging due to sparse, under-constrained observations. Multi-view images can resolve geometric ambiguity, but existing multi-modal methods do not fully exploit visual cues. We propose ControlFlow3D, a framework that distills knowledge from pretrained foundation models into a lightweight generation network during training, enabling strong performance without multi-view images available at test time. Specifically, our method performs conditional flow matching in a structured latent space defined by a frozen 3D encoder. Multi-view image tokens and 3D latent tokens are jointly clustered into shared prototypes via unsupervised soft KMeans; the resulting prototypes guide the flow trajectory through velocity control and are aligned by a consistency loss that progressively transfers image knowledge into the base network. A Mamba-based geometry decoder then maps latent features to dense coordinates with linear complexity. We also construct ShapeNetPU, a 35K-object multi-modal benchmark that is 31times larger than prior datasets. Extensive experiments across multiple benchmarks demonstrate the effectiveness of our method and its strong generalization capability.
Abstract: Class-Incremental Learning (CIL) requires models to continuously acquire new classes without forgetting previously learned ones. A dominant paradigm involves freezing a pre-trained model and training lightweight, task-specific adapters. However, maintaining task-specific parameters hinders knowledge transfer and incurs high retrieval costs, while naive parameter fusion often leads to destructive interference and catastrophic forgetting. To address these challenges, we propose Dynamical Adapter Fusion (DAF) to construct a single robust global adapter. Grounded in the PAC-Bayes theorem, we derive a fusion mechanism that explicitly integrates three components: the optimized task-specific adapter parameters, the previous global adapter parameters, and the initialization parameters. We utilize the Taylor expansion of the loss function to derive the optimal fusion coefficients, dynamically achieving the best balance between stability and plasticity. Furthermore, we propose a Robust Initialization strategy to effectively capture global knowledge patterns. Experiments on multiple CIL benchmarks demonstrate that DAF achieves state-of-the-art (SOTA) performance.
Abstract: Training-free guidance enables pre-trained diffusion and flow models to optimize application-specific objectives using feedback from external black-box reward functions. However, existing methods are feedback-inefficient because reward feedback is used only transiently to inform a localized gradient approximation or a discrete search decision, and is subsequently discarded. To address this limitation, we propose Flow-Direct, a framework that guides the generation process via a persistent . Theoretically, this guidance field is analytically derived from the log-density ratio between the base and reward-weighted target distributions; it transports the pre-trained distribution to the target distribution. In practice, the field is implemented as a non-parametric estimator constructed from all accumulated reward-evaluated samples. As more samples are collected during optimization, this empirical guidance field becomes increasingly accurate. This persistent formulation yields two major advantages. First, Flow-Direct is highly : because every evaluated sample is used to refine the global guidance field, no reward information is wasted. Second, the framework is naturally : once optimization is complete, the collected dataset defines a reusable guidance field for generating novel target samples without additional reward evaluations, and distinct guidance fields can be combined to generate samples that simultaneously satisfy multiple objectives.
PaperID: 7036, Poster
Abstract: When language models reason in chain-of-thought or collaborate in pipelines, they serialize structured information into natural language. How much compositional structure survives this bottleneck? We propose a round-trip protocol that answers this question empirically for tree-structured symbolic expressions. A generator converts a procedurally generated arithmetic expression into a word problem, a separate extractor recovers the expression from the word problem alone, and symbolic equivalence provides an exact oracle. Evaluating all pairwise combinations of fifteen models yields a communication matrix whose marginals separate generation quality from extraction quality. Three main findings emerge. First, the channel is lossy and asymmetric: swapping which model generates and which extracts shifts accuracy by up to 53.2 points, and the best pair reaches 82.0% by combining different models on each end rather than the same model on both. Second, at least 84% of round-trip failures originate at generation, and difficulty is driven by tree structure (operator count, depth, right-branching) rather than model family. Third, the channel is trainable: ~3600 fine-tuning examples that share the evaluation's operators and tree shapes lift every open-weight model above untrained GPT-5, an upper bound under matched semantics. A disjoint-domain regime with new operators and vocabulary also raises every open-weight model, confirming the gain is not an artifact of matched semantics, though a gap to the frontier remains. Together these results identify compositional serialization as a primary limiting factor when models communicate hierarchical structure through natural language.
PaperID: 7037, Poster
Authors: Noah Mitchell, Isaac Gabriel, Alexander Wyatt
Abstract: Self-attention without positional information is permutation equivariant, therefore positional encodings are required whenever a transformer must distinguish input order. However, injecting absolute position can also weaken length generalization. We formalize this tension for additive positional encodings by introducing positional spread, the total variance of the positional embedding matrix. For additive PE in softmax attention with mean pooling, we prove an upper bound demonstrating that the equivariance breaking component of risk is controlled by \|\widetilde E\|_F^2, along with a matching lower bound for self-separable additive PEs. We finally prove a spread-expressivity tradeoff: any additive PE that distinguishes absolute positions with minimum separation \delta must have spread \Omega(\delta^2 n). The theory predicts two regimes: when the target is near permutation invariant, lower spread improves length extrapolation; when the task requires absolute position, additional spread reduction destroys expressivity. We test this prediction using controlled synthetic tasks, spread-regularized additive PE, and long context retrieval and reasoning tasks. Across low positional sensitivity tasks, spread strongly predicts OOD accuracy; on absolute position tasks, performance demonstrates the predicted non-monotone tradeoff. For score-based methods such as RoPE and ALiBi, we include a displacement based comparison rather than matching bounds, clarifying both the reach and limits of the spread theory.
PaperID: 7038, Poster
Abstract: Deadline-Constrained Cost-Aware Dynamic Workflow Scheduling (D-CADWS) aims to minimize virtual machine (VM) rental cost while maintaining high workflow deadline satisfaction in dynamic cloud environments. The problem is challenging because the deadline impact of assigning a workflow task to a VM is not directly observable at decision time. It depends on long-horizon effects, including downstream task dependencies, VM queueing, and interactions among dynamically arriving workflows. Existing methods usually fold deadline violations into a single reward or penalty term, which obscures the distinction between action safety and action cost. In this paper, we cast D-CADWS as a representation learning problem. The core idea is to learn a deadline-aware representation that can predict, before execution, whether a candidate action is risky with respect to the current task deadline, and then use it to guide cost-efficient scheduling. Based on this idea, we propose a representation-centered deep reinforcement learning (RCDRL) method. RCDRL constructs task-level deadline supervision, trains a predictive deadline model to learn deadline-aware representations, and learns a VM-cost critic on top of this representation. The final policy follows a safe-first rule that rejects risky actions before minimizing VM cost. We further propose a two-phase training strategy to keep the learned representation aligned with the evolving policy-induced data distribution. Experiments on dynamic workflow scheduling benchmarks show that RCDRL achieves substantially lower VM cost than other state-of-the-art heuristic and DRL baselines while maintaining strong workflow deadline success.
PaperID: 7039, Poster
Authors:
Boxun Xu, Jiaji Lu, Yuxuan Yin, Zihu Wang, Ziyue (Alvin) Liu, Yu Wang, Zirui Liu, Peng LiAbstract: Visual autoregressive generation spans two dominant paradigms: next-scale prediction for images and video, and next-frame prediction for video. Both decode block-by-block rather than token-by-token, yielding KV caches whose statistics diverge from those of language models. At scale, KV memory exceeds the weights and limits large-scale deployment caused by a large amount of visual tokens. Our analysis of these caches reveals three findings. First, a small set of massive-value channels carries the functional load: ablating these outlier channels collapses generation, whereas ablating a matched set of inlier channels leaves quality intact. Second, the outliers organize along two axes, feature and block, with magnitudes persistent within each cell and shifting across cells. Third, this structure arises from two independent mechanisms: projection bias and RoPE-induced scattering. Guided by these findings, we propose VAR-Q, a tuning-free KV-cache quantization framework that groups along the same two axes. The generation loop provides block boundaries for free, removing the need for outlier detection or calibration. Paired with bit-packing kernels for runtime memory reduction, VAR-Q delivers low-bit quantization across nine models spanning both paradigms, cutting KV memory by up to 78.5% with negligible quality degradation and enabling 6× larger inference batches.
PaperID: 7040, Poster
Abstract: In-context reinforcement learning (ICRL) is often described as inference-time RL in which an LLM improves by accumulating trajectory-reward pairs in its context, with the reward acting as a learning signal. This framing poses a central question: is ICRL genuine inference-time RL, or better understood as a form of in-context learning (ICL)? We investigate this in the context of direct ICRL, where the model directly uses trajectory-reward pairs. Through controlled experiments on three benchmarks across six models, we find that the reward is read, but its effect is small; randomizing or removing the reward leaves the improvement curve almost unchanged. Trajectories drive improvement, but not through their semantic content: shuffled or corrupted trajectories work as well as real ones. These patterns closely mirror those known in ICL, suggesting that direct ICRL is better understood as a special case of ICL than as inference-time RL. This reframing has implications for ICRL memory design: ICL factors such as input distribution and demonstrations may matter more than RL elements such as reward shaping and exploration.
Abstract: Attention sinks and massive activations are recurring and closely related phenomena in Transformer models. Existing explanations have largely focused on the forward pass, yet in pre-norm Transformers, large residual-stream norms play only an indirect forward role because sublayers operate on normalized inputs. We study this relationship from the perspective of backpropagation. Empirically and theoretically, we show that under causal masking, attention sinks can induce pronounced gradient concentration, which we term gradient sinks. Since the RMSNorm Jacobian attenuates gradients roughly in inverse proportion to input norm, massive activations can be understood as adaptive regulators of this localized gradient pressure during training. This interpretation predicts that attenuating sink-induced gradients should weaken massive activations. We test this prediction with V-scale, a modification that adjusts backpropagated gradients on the value path. In V-scale models, attention sinks are preserved, whereas massive activations are suppressed. These results identify gradient sinks as a backward-pass counterpart of attention sinks, and massive activations as an adaptive RMSNorm-mediated response that attenuates the resulting localized training pressure. Our code is available at https://anonymous.4open.science/r/GradientSinkCode-B309.
Abstract: Multimodal Large Language Models (MLLMs) still struggle with fine-grained visual understanding, where answers often depend on small but decisive evidence in the full image. We observe a regional-to-global perception gap: the same MLLM answers fine-grained questions more accurately when conditioned on evidence-centered crops than on the corresponding full images, suggesting that many failures stem from difficulty to focus on relevant evidence rather than insufficient local recognition ability. Motivated by this observation, we propose Vision-OPD (Vision On-Policy Distillation), a regional-to-global self-distillation framework that transfers the model's own privileged regional perception to its full-image policy. Vision-OPD instantiates two conditional policies from the same MLLM: a crop-conditioned teacher and a full-image-conditioned student. The student generates on-policy rollouts, and Vision-OPD minimizes token-level divergence between the teacher and student next-token distributions along these rollouts. This enables the model to internalize the benefit of visual zooming without external teacher models, ground-truth labels, reward verifiers, or inference-time tool use. Experiments on multiple fine-grained visual understanding benchmarks show that Vision-OPD models achieve competitive or superior performance against much larger open-source, closed-source, and "Thinking-with-Images" agentic models.
Authors: Wenjie Guan, Jelena Bradic
Abstract: Template tasks have emerged as a clean testbed for asking whether transformers reason with abstract symbols rather than concrete token names. We study the fixed-label classification version of this problem, where train and test examples share latent templates but may use disjoint vocabularies. Unlike next-token prediction, the model need not emit unseen symbols; it must learn a decision rule invariant to symbol renaming. We analyze regularized kernel logistic classification in the transformer-kernel regime. Our main result decomposes the learned predictor into an ideal template-level classifier and a finite-sample perturbation caused by accidental token overlaps in the training data. We encode these overlaps by a colored collision graph and prove high-probability margin-transfer guarantees for fresh-symbol classification. This perspective extends template-based analyses to logistic classification and refines scalar diversity conditions: vocabulary size controls the average rate of collisions, but collision geometry controls whether the ideal classification margin is preserved. More broadly, the same perturbation framework applies to abstraction-augmented inputs, yielding a general margin-versus-collision criterion for identifying when prompting strategies improve fresh-symbol generalization. Synthetic template experiments illustrate the predicted roles of regularization, sample size, and transformer-kernel structure.
PaperID: 7044, Poster
Authors: TIANJUN SHI, Haotian Xiong, Ziyu Gong, Qi Lu, Li Li
Abstract: Visual token pruning is an effective way to accelerate vision-language models and is especially useful for vision-language-action (VLA) inference, where many visual tokens must be processed before predicting robot actions. Existing pruning methods usually estimate which tokens can be pruned based on attention scores or feature diversity, retaining tokens that are either highly attended or visually different from others. However, most of them use fixed pruning schedules, such as pruning once at a preset layer or pruning at uniformly spaced layers. Such schedules can be risky for VLA models, because the model may not know which visual regions matter for the action in early layers. Tokens that look unimportant at first may become useful after the model combines visual observations with the language instruction. In this work, we propose SAPrune, a training-free visual token pruning framework for efficient VLA inference. Instead of pruning at fixed layers, SAPrune uses a small calibration set to observe how action-to-visual attention changes across layers, and chooses pruning layers only after the attention pattern becomes more reliable. At each selected layer, SAPrune applies a dual-path pruning rule: one path protects strongly attended visual tokens from pruning, while the other prevents useful surrounding context from being discarded. Experiments on LIBERO, SIMPLER, and real-world robotic tasks show that SAPrune prunes 87.5% of visual tokens and achieves up to 1.718× inference speedup while maintaining competitive task success rates.
Abstract: Scaling Federated Learning (FL) to billion-parameter models forces a challenging trade-off between privacy, scalability, and model utility. Existing solutions often tackle these challenges in isolation, sacrificing accuracy, relying on costly cryptographic tools, or introducing communication and optimization inefficiencies that affect convergence. We introduce ERIS, an FL framework centered on (FSA), a novel mechanism that partitions each client update into non-overlapping shards whose aggregation is distributed across multiple client-side aggregators. FSA removes the central aggregation bottleneck, limits the information visible to any single observer, and preserves the centralized FL update after reassembly. ERIS can further readily integrate Distributed Shifted Compression (DSC) to reduce transmitted payloads and exposed coordinates. We prove that ERIS preserves convergence under standard assumptions and bounds mutual information leakage by the observable fraction of each update, decreasing with the number of client-side aggregators, and with the compression level when DSC is enabled. Experiments across image and text tasks, including large language models, show that ERIS achieves FedAvg-level utility while substantially reducing communication bottlenecks and improving robustness to membership inference and reconstruction attacks, without relying on heavy cryptography or utility-degrading perturbations.
Abstract: In the rapidly evolving field of Electronic Design Automation (EDA), the deployment of Large Language Models (LLMs) for Register-Transfer Level (RTL) design has emerged as a promising direction. However, silicon-grade correctness remains bottlenecked by (i) limited test coverage and reliability of simulation-centric evaluation, (ii) regressions and repair hallucinations introduced by iterative debugging, and (iii) semantic drift as intent is reinterpreted across agent handoffs. In this work, we propose Veri-Sure, a multi-agent framework that establishes a design contract to align agents' intent and uses a patching mechanism guided by static dependency slicing to perform precise, localized repairs. By integrating a multi-branch verification pipeline that combines trace-driven temporal analysis with formal verification consisting of assertion-based checking and Boolean equivalence proofs, Veri-Sure helps improve functional correctness by enriching the simulation-driven debugging loop with localized diagnostic hints. We also introduce VerilogEval-v2-EXT, extending the original benchmark with 53 more industrial-grade design tasks and stratified difficulty levels, and show that Veri-Sure achieves state-of-the-art verified-correct RTL code generation performance, surpassing standalone LLMs and prior agentic systems. Code and dataset are available at https://anonymous.4open.science/r/Veri-Sure-D861.
Authors:
Shen Xu, Xiangwen Zhuge, Zhe Xu, Yingkun Hu, Zheng Yang, Yunhao LiuAbstract: LLMs often struggle with memory-constrained deployment on consumer-grade hardware due to their massive parameter sizes. While existing solutions such as model compression and offloading improve deployment feasibility, they often suffer from substantial accuracy degradation or severe throughput bottlenecks. Recent error compensation methods recover accuracy through auxiliary LoRA-style branches, and we observe that these branches are inherently amenable to offloading: they require substantial parameter storage but access only a small subset of compensation parameters during each inference step. Motivated by this opportunity, we propose HCInfer, a heterogeneous inference system that offloads residual compensation to the CPU while executing the compressed backbone on the GPU, and further introduces an asynchronous compensation pipeline and sensitivity-aware dynamic rank allocation to hide compensation overhead and maximize accuracy recovery. Experimental results show that HCInfer achieves a maximum accuracy improvement of 5.2% on downstream tasks compared to compression model and sustaining a maximum speedup of 10.4x compared to full-precision model.
PaperID: 7048, Poster
Abstract: Safety-aligned large language models often lose their safety behaviors after fine-tuning, even when safety data are included. While this fragility is well documented, its underlying cause remains unclear. We propose that formulaic refusal strings function as in the model's output distribution, i.e., low-support behaviors that are easy to fit but structurally unstable under subsequent fine-tuning, and that this is an important source of safety fragility. Controlled experiments verify this: fixed refusals are acquired and forgotten much like arbitrary constant strings, in contrast to prompt-grounded natural responses. Motivated by this insight, we improve safety retention by constructing natural, context-aware safety supervision that keeps safety responses within the model's existing output distribution and grounded in semantics. We instantiate this principle in both supervised and preference-based alignment settings, using perplexity under the original model as a practical proxy for distributional naturalness. Across multiple models and fine-tuning regimes, our approach improves both immediate safety and its retention under continued fine-tuning without sacrificing helpfulness, suggesting that naturalness is a useful principle for durable safety alignment.
PaperID: 7049, Poster
Abstract: Diffusion and flow matching have emerged as expressive generative planners for autonomous driving planning, owing to their ability to model high-fidelity and multi-modal trajectory distributions. Nevertheless, existing generative planners are predominantly optimized through imitation learning, which induces a fundamental mismatch between supervised trajectory fitting and closed-loop planning metrics, while their iterative sampling procedures often impose substantial computational overhead. In this paper, we propose FlashPlanner, a goal-conditioned flow-matching planner with online RL finetuning for closed-loop AD planning. FlashPlanner introduces a continuous future goal point as a compact navigation interface for the generative planning policy. This goal is produced by a continuous goal predictor, which is first pretrained and subsequently optimized with RL using closed-loop feedback from multiple candidate goals at each decision step. To support efficient online reinforcement fine-tuning and real-time deployment, FlashPlanner adopts a data-prediction flow-matching objective and removes redundant architectural components in existing diffusion-based planners. Experiments on the closed-loop nuPlan and interPlan benchmarks demonstrate that FlashPlanner achieves state-of-the-art planning performance while delivering 6× faster inference (166 FPS) than the previous SOTA baseline (28 FPS). We will open-source our project.
PaperID: 7050, Poster
Abstract: Reliable deployment of graph neural networks requires calibration, out-of-distribution (OOD) detection, and robustness to distribution shift, yet existing methods typically address these requirements with separate models and objectives. A key obstacle is the lack of a single learned representation whose structure can serve all three tasks through different readouts. Inspired by spectral approaches in graph signal processing and polynomial chaos theory, we model uncertainty-bearing node embeddings as random graph signals: graph Fourier filters capture structural variation, while a scalar orthogonal-polynomial chaos coordinate captures latent stochastic variation. This yields a \emphdoubly-spectral stochastic (DSS) expansion in which the mean coefficient carries evidence for classification and energy-based OOD scoring, higher-order coefficients encode structured logit variation, and quadrature averaging turns that variation into a calibration-sensitive predictive distribution. Calibration thus uses the quadrature-averaged predictive distribution, OOD detection uses the mean-logit energy score, and distribution-shift robustness uses the corresponding DSS branch as a regularized spectral residual when paired with a deterministic encoder. Under mild conditions this representation universally approximates Gaussian-latent random graph signals with exponentially decaying truncation error. The resulting model, \emphDSS-GNN (Doubly-Spectral Stochastic GNN), can be used standalone or as a residual branch alongside a deterministic encoder (\emphDSS-Hybrid). In standalone form, DSS-GNN achieves the lowest Brier score among the compared uncertainty-aware baselines on all 14 node classification benchmarks without post-hoc correction; in hybrid form, DSS-Hybrid achieves the best AUROC on most node-OOD settings, competitive cross-graph OOD performance, and the strongest shifted accuracy among compared baselines on all 7 GOOD concept-shift benchmarks under standard ERM training.
PaperID: 7051, Poster
Abstract: Deep networks memorize parts of their training data, exposing them to privacy attacks such as membership inference. The most vulnerable samples are often hard-to-learn tail examples: they are fitted late, carry high information for generalization, and are difficult to protect without sacrificing performance. We show that this privacy–utility tension can be mitigated by learning such samples earlier. To this end, we propose Weighted Loss InfoBatch (WLIB), which dynamically prunes easy samples while re-weighting hard ones according to sample-wise pruning ratios. By adjusting the influence of informative samples throughout training, WLIB jointly reduces memorization, improves robustness to membership inference attacks, and speeds up optimization. Extensive experiments demonstrate that WLIB achieves a favorable combination of privacy, utility, and training efficiency.
Abstract: Fine-tuning foundation models on new tasks inevitably suffer from catastrophic forgetting. While existing works attempt to mitigate this on the basis of parameter-efficient fine-tuning methods, they adopted an overly restrictive Subspace Orthogonality condition. In this paper, we introduce a purely , which is the necessary and sufficient condition for preserving historical performance to the first order. By projecting parameter updates into the JAcobian NUll Space (JANUS), our method significantly recovers compromised historical knowledge without interfering with the underlying fine-tuning process. To overcome the local validity of the Jacobian approximation, we further propose a Multi-step Adaptive Rectification mechanism that utilizes the JANUS shift to dynamically verify the valid trust region and adjust step sizes. Coupled with our proposed ghost projection, ghost orientation comparison, and sequence-level singular value decomposition compression techniques, JANUS also achieves great temporal and spatial efficiency. Experiments demonstrate that JANUS seamlessly integrates with various fine-tuning methods, fundamentally breaking the stability-plasticity dilemma by recovering historical knowledge while preserving downstream task adaptation.
PaperID: 7053, Poster
Abstract: Multi-agent orchestration over large language models is increasingly used to improve reasoning accuracy, especially on problems that benefit from critique, verification, specialization, or consensus, but these gains often come with inference costs that are not explicitly controlled. Many systems spend similar compute on straightforward queries and genuinely ambiguous or expert-level ones. Existing methods typically follow fixed collaboration protocols, compile task-level workflows offline, route queries among models or collaboration modes, or use learned controllers whose state does not explicitly identify which reasoning defects remain unresolved. As a result, they do not directly decide during inference whether the current multi-agent trajectory still justifies another agent call. We introduce Defect-Aware Residual Coverage (DARC), a training-free inference-time controller for adaptive multi-agent reasoning that represents the current trajectory using four residual defect dimensions: answer uncertainty, claim-level contradiction, verification failure, and aspect under-coverage. DARC selects the next agent by maximizing cost-normalized submodular marginal coverage over the remaining residuals. Each candidate agent receives a role-induced capability profile from frozen embeddings of its natural-language role description, enabling DARC to operate without task-specific supervision, learned routing parameters, or fine-tuning. Across five reasoning and multimodal suites, MMLU-Pro, GPQA-Diamond, LiveBench, MMMU-Pro, and HLE, DARC achieves the strongest accuracy under the main shared heterogeneous six-agent pool while using substantially fewer tokens and agent calls than recent workflow, adaptive orchestration, and cost-aware routing baselines, with additional large-pool experiments showing consistent scaling behavior. It also outperforms the single-call frontier reasoning models tested, showing that residual-aware selective invocation can improve both answer quality and inference efficiency.
PaperID: 7054, Poster
Abstract: Large language models may encounter factual knowledge during pre-training yet fail to reliably use that knowledge after fine-tuning. We study this gap in a stylized one-layer self-attention + MLP transformer trained by next-token prediction and subsequently fine-tuned on question-answering data. We first prove that, under suitable regularity conditions, the model reaches near-optimal pre-training loss while learning structured attention patterns. We further show that fine-tuning turns the Q&A prompt format into a trigger for pre-trained relation features, enabling the model to extract facts not revisited during fine-tuning. Our analysis reveals a relation-covering characterization for knowledge extraction: fine-tuning need not revisit every stored fact, but it must cover enough latent relation-template directions through which facts were encoded during pre-training. We prove that extraction improves with pre-training multiplicity and fine-tuning coverage, but becomes harder as the relation-template universe grows. Conversely, insufficient coverage yields a failure regime in which facts can be stored but not extracted, providing a stylized mechanism for hallucination. Our analysis covers both full and low-rank fine-tuning, and experiments on synthetic data and PopQA-based GPT-2/Llama models support the predicted trends.
Authors: Daniel Dehtyriov, Jonathan F MacArt, Justin Sirignano
Abstract: Turbulence is ubiquitous in engineering, yet direct simulation is prohibitively expensive. The Reynolds-averaged Navier-Stokes (RANS) equations provide computational savings exceeding ten orders of magnitude but introduce unclosed terms requiring modelling (the closure problem). Existing machine-learning (ML) closures suffer distribution shift when trained offline on high-fidelity data, and ML models that bypass the governing equations often have limited capability to generalise. We develop a physics-derived deep learning closure for RANS, the Deep Algebraic Reynolds Stress Model (DARSM), which can be trained on one or two flow cases and generalise across an order of magnitude in Reynolds number, to unseen geometries, and across flow regimes. A neural network maps flow invariants to empirical parameters in an implicit algebraic Reynolds stress equation, derived from the Reynolds stress transport equations under the weak-equilibrium assumption, imposing significant physics-based structure on the ML closure. End-to-end optimisation through the governing PDEs and the coupled implicit closure eliminates distribution shift, but both unrolled and implicit automatic differentiation fail on the stiff coupled solver. We derive adjoint equations that exploit the solver's implicit-explicit structure to enable computationally efficient optimisation. On the canonical square-duct and periodic-hill benchmarks, DARSM reduces test velocity error as compared to the baseline RANS by 2-4× across Reynolds number, geometries, and flow regimes, with peak case-level reductions of 12×. The model trained on attached, anisotropy-dominated flows (square duct) accurately generalises without retraining to separated flows (periodic hills), a regime change in the underlying turbulence physics. DARSM also outperforms five established ML methods: offline training, tensor-basis neural networks, field-inversion machine learning, DeepONets, and physics-informed neural networks.
PaperID: 7056, Poster
Abstract: Denoising-based Vision-Language-Action (VLA) policies are often parameterized by a single denoiser shared across denoising time. This paper shows that expert usefulness in VLA action generation can vary systematically across denoising phases, and that this variation can be exposed through controlled step-wise expert aggregation. We mix two full action experts at the velocity-field level with a fixed smooth schedule, enabling schedule-only interventions that keep expert capacity and total mixing mass matched while changing when each expert is emphasized. The empirical evidence is organized as a matched chain: semantically distinct expert pairs test denoising-phase assignment through forward versus reversed schedules, while an anchor-free isomorphic-expert control tests schedule sensitivity without manually designed expert semantics. Experiments on both LIBERO and CALVIN show the same directionality phenomenon: aligned schedules outperform their time-reversed counterparts under matched strength, indicating that when each expert is emphasized matters for action denoising. We complement these findings with a squared-error risk analysis showing that denoising-time variation in expert advantage implies a time-dependent optimal mixture, together with a schedule-weighted objective note explaining why copied experts can still functionally differentiate under non-constant schedules.
Abstract: While parameter-efficient fine-tuning methods like low-rank adaptation (LoRA) are standard for large language models, principled estimation of epistemic uncertainty remains challenging. Recent results in the LoRA regime suggest that discrete multi-mode approaches such as deep ensembles offer little benefit over single-mode methods. This contradicts broader observations in deep learning, where ensembling independent optima typically improves generalization, and linking these modes through continuous low-loss valleys further enhances Bayesian model averaging (BMA). Whether such structure exists in the LoRA space and whether it yields functional diversity missed by local or discrete methods has not been studied. We introduce , a segmented Bézier curve parameterization in the LoRA space, with two variants: a free configuration that jointly optimizes all control points, and an anchored configuration that connects independently fine-tuned LoRA optima. We prove pathwise continuity and Lipschitz regularity of the loss along the curve and empirically show, across reasoning and classification benchmarks with Qwen2.5 7B, that linear interpolation encounters loss barriers, while our anchored multi-segment curves connect independent optima through continuous low-loss valleys. Combined with flat-minima perturbations and a Jensen-Shannon divergence regularizer, yields measurably higher mutual information of the predictive distribution without sacrificing performance, and links continuous parameter-space traversal to functional diversity.
Abstract: Score-based diffusion models are typically trained by minimizing the L^2 score matching error, and standard theoretical analyses rely on this quantity to bound the sampling discrepancy between the learned and target distributions. We show the L^2 score error is not the right intrinsic measure of marginal distributional quality: a learned diffusion model can incur arbitrarily large L^2 score error while perfectly matching the target distribution. By decomposing score errors into a gradient and a solenoidal component (a Helmholtz-Hodge decomposition), we identify the geometric reason behind this: only the gradient component enters the marginal Fokker-Planck dynamics, while the solenoidal component is structurally invisible. We make this precise in three results. First, building on the corrected geometry, we prove an impossibility result: no monotone function of the L^2 score error can uniformly lower bound any divergence between the learned and target distributions. Second, we derive an upper bound on the Kullback-Leibler divergence that depends only on the observable gradient component of the error, tightening the standard Girsanov bound and identifying its looseness as the cost of operating on path-space rather than marginal-space dynamics. Third, we give a tractable estimator of the gradient component via a dual Sobolev identity, which is shown to empirically correlate substantially better with sample quality than the full L^2 error.
PaperID: 7059, Poster
Abstract: Graph Contrastive Learning (GCL) has shown strong promise for unsupervised graph representation learning, yet its effectiveness remains limited on heterophilic graphs, where connected nodes often belong to different classes. Existing methods rely on complex augmentation schemes, intricate encoders, or negative sampling, which raises the question of whether such complexity is truly necessary in this challenging setting. In this work, we revisit the foundations of supervised and unsupervised learning on graphs and uncover a simple yet effective principle for GCL: mitigating node feature noise by aggregating it with structural features derived from the graph topology. This phenomenological study suggests that the original node features and the graph structure naturally provide two complementary views for contrastive learning. Building on this insight, we propose an embarrassingly simple GCL model that uses a GCN encoder to capture structural features and an MLP encoder to isolate node feature noise. Our design requires neither data augmentation nor negative sampling, yet achieves state-of-the-art results on heterophilic benchmarks with minimal computational and memory overhead, while offering advantages in homophilic graphs in terms of complexity, scalability, and robustness. Moreover, a variant of GCN-MLP based on the same principle also achieves SOTA performance on homophilic datasets. We provide theoretical justification for our approach and validate its effectiveness through extensive experiments.
PaperID: 7060, Poster
Abstract: Spectral gradient descent (SpecGD) orthogonalizes matrix parameter updates and has inspired practical optimizers such as Muon. They often perform well in large language model training, but their dynamics remain poorly understood, especially in factorized parameterizations where the product matrix does not receive orthogonalized updates. We study such dynamics through matrix factorization (MF), where the orthogonalization is applied separately to the factor updates. We analyze spectral gradient flow (SpecGF)—a continuous-time analog of SpecGD—in the low-rank MF setting and prove "equal-rate" dynamics: all singular values grow at equal rates up to small deviations. Consequently, smaller singular values attain their target values earlier than larger ones, contrasting with the largest-first stepwise learning observed in standard gradient flow. Moreover, we prove that SpecGF in our setting converges to global minima from almost all initializations, provided the factor norms remain bounded; with \ell_2 regularization, we obtain global convergence. Empirically, we observe that LoRA fine-tuning with orthogonalization-based optimizers including Muon exhibit near-uniform growth in the product of LoRA adapters, consistent with the mechanism predicted by our MF analysis.
Abstract: Scaling Maximum Entropy Reinforcement Learning (RL) to high-dimensional humanoid control remains a fundamental challenge, as the ''curse of dimensionality'' induces severe exploration inefficiency and training instability. Consequently, highly optimized deterministic policy gradients currently dominate high-throughput regimes. We address this limitation with FastDSAC, a framework that effectively unlocks the potential of maximum entropy stochastic policies for complex continuous control. We introduce Dimension-wise Entropy Modulation (DEM) to dynamically redistribute the exploration budget, alongside a continuous distributional critic tailored to ensure accurate value estimation by mitigating both high-dimensional overestimation and discrete quantization artifacts. Extensive evaluations on HumanoidBench and a diverse set of continuous control tasks demonstrate that FastDSAC establishes state-of-the-art performance for high-dimensional stochastic policies on the evaluated benchmarks. Our method is competitive with and often outperforms strong deterministic baselines, with gains of 180% and 350% on the challenging
Authors:
Shuochen Chang, Qingyang Liu, Shaobo Wang, Bingjie Gao, Qianli Ma, Haonan Zhao, Yibo Miao, Yulin Sun, Zelin Peng, Jiangtong Li, Li NiuAbstract: Large language models achieve high reasoning performance via explicit chain-of-thought and reinforcement learning, but require long output sequences and extended inference time. Latent reasoning reduces this cost by shifting computation into a latent space; however, continuous latent methods are hard to train, suffering from unstable and uninterpretable reasoning trajectories. We argue these issues stem from a misalignment between continuous-space reasoning and discrete symbolic supervision, as continuous states lack explicit anchors for step-by-step alignment. To resolve this, we propose Discrete Latent Reasoning (DLR), the first method that converts continuous latent states into explicit discrete tokens. Inspired by render-based compression, we render textual chains of thought into images, extract visual features, and construct a discrete latent vocabulary via clustering-based fine-tuning. Expanding the vocabulary and output head enables standard autoregressive modeling over both natural language and latent tokens, supporting pretraining alignment, SFT, and RL. Experiments on five reasoning benchmarks and two model series (Qwen3-VL and LLaMA-3) confirm that DLR outperforms prior latent reasoning baselines with up to 20× compression. Furthermore, the learned latent trajectories retain an interpretable semantic structure. Overall, discrete latent tokens provide a controllable and interpretable basis for efficient latent reasoning.
PaperID: 7063, Poster
Abstract: Vision-Language Models (VLMs) show great potential for Industrial Anomaly Detection (IAD) but often overlook subtle high-frequency defects due to their semantic bias. Additionally, one-class adaptation risks rank collapse and catastrophic forgetting. To address these issues, we propose a unified framework termed SF-DST. First, we design Spatial-Frequency Dual-Stream Transformer, which explicitly captures lost details via a parallel discrete cosine transform branch. Crucially, it incorporates Asymmetric Cross-Modal Modulation (ACMM), a mechanism that leverages spatial semantics as a context gate to selectively retrieve and inject spectral features, effectively compensating for the texture blindness of VLMs. Second, Anomaly-Aware LoRA (A-LoRA) enforces Orthogonal Regularization on the adaptation subspace. This geometric constraint ensures subspace diversity to prevent mode collapse, while optimizing a hypersphere-based metric for precise anomaly discrimination. Extensive experiments on MVTec-AD, VisA and MMAD benchmarks demonstrate that our method significantly outperforms state-of-the-art approaches, successfully bridging the gap between large-model generalization and industrial-grade precision.
PaperID: 7064, Poster
Authors: Jessica E Liang, Jianbo Shi
Abstract: Understanding how information propagates within vision-language models (VLMs) and large language models (LLMs) is central to interpretability and efficient deployment. Attention weights indicate where probability mass is allocated, but do not directly measure directional influence, redundancy, or functional importance. We introduce transfer entropy (TE) as an information-theoretic measure of directed information flow in Transformers, and develop tractable TE estimators for high-dimensional hidden states at the granularity of layers, tokens, and attention heads. We first analyze multimodal models, beginning with LLaVA-1.5-7B and then CLIP ViT-B/16. In LLaVA-1.5-7B, vision\rightarrowtext TE increases with depth, suggesting that multimodal fusion is localized in the upper decoder layers. In CLIP, TE reveals late-layer redundancy in the vision tower and broader elevated TE in the text tower. We then analyze unimodal models, including RoBERTa, T5, Llama-3.2-3B, and Qwen-2.5-7B. Across these LLMs, TE reveals consistent depth-dependent structure: RoBERTa exhibits a mid-layer peak, decoder-only LLMs concentrate task-dependent computation in broad mid-depth bands, and T5 shows distinct encoder and decoder dynamics. Beyond layerwise analysis, TE also exposes token- and head-level communication roles, distinguishing sinks, broadcasters, and benign components. These signals support TE-guided layer, token, and attention-head pruning across VLMs and LLMs, consistently outperforming attention- or saliency-based heuristics under matched budgets. Overall, TE provides a unified directed-information lens for diagnosing, interpreting, and compressing Transformer-based models.
PaperID: 7065, Poster
Authors:
Kenan Li, Rongzhi Li, Linghao Zhang, Qirui Jin, Liao Zhu, XiaoSong Huang, Geng Zhang, Yikai Zhang, Shilin He, Chengxing Xie, Xin Zhang, Zijian Jin, Bowen Li, Chaoyun Zhang, Yu Kang, Yufan Huang, Elsie Nallipogu, Saravan Rajmohan, Qingwei Lin, Dongmei ZhangAbstract: Language model (LM) agents have driven substantial progress in automated software engineering (SWE), yet building and testing software repositories at scale remains a largely manual and labor-intensive bottleneck. In this work, we introduce RepoLaunch, a novel agentic framework that automatically resolves dependencies, compiles source code, and extracts test results across diverse programming languages and operating systems. RepoLaunch achieves a 78% build success rate, outperforming the Python/Linux-only prior system by 18%. To demonstrate its application, we further present a fully automated pipeline for SWE dataset creation driven by RepoLaunch, which only requires human input at the task-design stage. RepoLaunch is open-sourced, and its automated task-generation pipeline has already been adopted by several recent works on agentic benchmarking and training.
Authors:
Phu-Hoa Pham, Chi Nguyen Tran, Phu-Quy Nguyen-Lam, Dao S Minh, Trung-Kiet Huynh, Long Tran-ThanhAbstract: Streaming decision trees are natural candidates for open-world continual learning, as they perform local updates, enjoy bounded memory, and static decision boundaries. Despite these, they still fail in online class-incremental learning due to two coupled miscalibrations: (i) their split criterion grows unreliable as the class count K expands, and (ii) the absence of knowledge transfer at split time. Both failures share a common root: the range of Information Gain intrinsically scales with \log_2 K. Consequently, any Hoeffding-style confidence radius derived from it must inevitably grow with the class count, making a K-independent split criterion structurally impossible, taking away the potential benefits of applying streaming decision trees to continual learning. To fix this issue, we present \mist (McDiarmid Incremental Streaming Tree), which resolves both failures through three integrated components: (i) a tight, K-independent McDiarmid confidence radius for Gini splitting that acts as a structural regulariser; (ii) a Bayesian inheritance protocol that projects parent statistics to child nodes via truncated-Gaussian moments, with variance reduction guarantees strongest precisely when splitting is most conservative; and (iii) per-leaf KLL quantile sketches that support both continuous threshold evaluation and geometry-adaptive leaf prediction from a single data structure. On standard and stress-test tabular streams, \mist is competitive with global parametric methods on near-Gaussian benchmarks and uniquely robust on non-Gaussian geometry where SOTA benchmarks collapse.
Authors: Pierre Cesar, Sofya Dymchenko, Abhishek Purandare, Bruno Raffin
Abstract: Data-driven PDE surrogates are trained with data produced by numerical PDE solvers. However, when the surrogate's goal is to generalize across a wide range of PDE configurations (e.g., initial conditions and physical coefficients), generating a representative training set is non-trivial. Uniform sampling of configuration parameters often under-represents trajectories exhibiting challenging dynamics, leading to high prediction errors and large error variance in the trained surrogate. Online training, where data generation and surrogate training are coupled, offers a natural advantage by allowing solver parameters to be steered on-the-fly. To efficiently exploit this capability, we introduce Online Generative Active Sampling (OGAS), an active learning method that reactively learns the relationship between configuration parameters and surrogate performance to control the sampling distribution. OGAS trains a fast diffusion model in parallel to the surrogate to act as a conditional sampler, mapping a surrogate-derived difficulty signal (e.g., loss or uncertainty) to configuration parameters. By actively drawing target signals from a prior biased toward high difficulty, OGAS continuously steers data generation toward challenging regimes without delaying the training workflow. We evaluate OGAS across 2D PDEs with distinct challenging dynamics (Kuramoto-Sivashinsky, Navier-Stokes, Gray-Scott) and up to 308 parameters, using multiple surrogate architectures. Across all settings, OGAS consistently improves tail statistics, yielding substantial reductions in errors above the 99th percentile and overall error dispersion compared to uniform sampling. While prioritizing challenging trajectories introduces a trade-off with average error, OGAS effectively ensures worst-case reliability of trained surrogates with negligible wall-time overhead.
PaperID: 7068, Poster
Abstract: Low-rank factorization is a standard way to scale optimization in machine learning by replacing large matrix variables with compact factors. For positive semidefinite (PSD) variables, the symmetric Burer--Monteiro factorization (sBMF) writes Z=XX^\top with a single low-rank factor X. A recent asymmetric alternative (aBMF) writes Z=XY^\top and adds a quadratic penalty (\gamma/2)\|X-Y\|_F^2 to encourage symmetry. This split is attractive because it yields a biconvex objective with alternating convex subproblems, but its practical value depends strongly on how the penalty parameter \gamma is chosen. We study a unified regularized aBMF framework and derive an explicit lower bound on \gamma that guarantees exactness: under mild assumptions, any \gamma above this threshold makes aBMF and sBMF share the same critical points. This gives a principled way to use the asymmetric formulation without altering the critical-point structure of the symmetric problem. In particular, it answers the open question of whether an exact penalty exists for asymmetric relaxation.
Abstract: Generative models, from diffusion models to large language models, achieve remarkable performance but at a cost in training data orders of magnitude larger than what biological learners require. An alternative paradigm has emerged in which networks are trained to predict their \emphown latent representations of related views or masked regions, as in data2vec and JEPA -- an idea related to predictive-coding accounts of the cortex. Despite strong empirical results, the theoretical understanding of these methods remains limited. Central questions include: by how much does latent prediction actually improve data efficiency? Is there a benefit to stacking such methods into multi-scale hierarchies? We answer both using as data a tractable probabilistic context-free grammar that captures the compositional structure of natural language and images. We prove that latent prediction recovers the full latent tree of depth L from a number of samples scaling as m^3, where m is the number of production rules per symbol. This is much fewer than the m^L samples required by supervised learning and the m^L+1 required by token-level SSL. We confirm this bound with (i) a hierarchical clustering algorithm, (ii) an end-to-end neural network whose predictor-clusterer modules predict their own latents at each level via gradient descent, and (iii) the first sample-complexity analysis of data2vec, which we show implicitly performs hierarchical latent prediction. This suggests that explicit stacking such as H-JEPA is largely redundant.
PaperID: 7070, Poster
Abstract: Reliable map-free LiDAR relocalization requires not only accurate and efficient pose estimation but also trustworthy confidence estimation for detecting localization failures. However, estimating reliability to ensure trustworthy LiDAR relocalization remains largely unexplored. Also, existing LiDAR pose regression methods suffer from accuracy and efficiency trade-off. To address this limitation, we propose SPAR, a model-agnostic confidence estimator that reformulates localization reliability as scan-pose alignment. Given a query scan and a predicted pose, SPAR measures their compatibility in a learned embedding space and explicitly optimizes confidence scores to correlate scan with pose. In addition, we introduce VeLoc, an efficient and accurate pose regressor based on flow-matching for efficient pose refinement. Experiments on the Oxford and NCLT datasets demonstrate that the proposed framework achieves a favorable balance between efficiency and accuracy for LiDAR relocalization while enabling reliable localization failure detection across different localization backbones.
PaperID: 7071, Poster
Authors:
Dongwon Kim, Taehyoung Kim, Sungdong Lee, Joong-Ho (Johann) WonAbstract: We study scalable second-order algorithms for high-dimensional square-root Lasso problems with multiple responses. The resulting estimators preserve the scale-free tuning property of their single-response counterpart, but accurate large-scale computation is challenging because both the loss and penalty terms are nonsmooth. In this paper, we propose a computational framework that universally applies to multitask and multivariate square-root Lasso models (which includes the single-response version) based on the Douglas-Rachford splitting. We solve the nonsmooth equation induced by the splitting by a regularized semismooth Newton method with projection and fixed-point safeguards. Despite the nonsmoothness of the problem, the structure of the proximal maps and generalized Jacobian lead to a reduced Newton system in a much smaller dimension than that in the original space, promoting scalability. Our method enjoys global and locally quadratic convergence under mild conditions. Experiments on synthetic regression problems and large-scale multi-omic data show that the proposed method reaches the same fitted models and predictive performance as latest solvers tailored to individual models, while requiring fewer iterations and less wall-clock time in large-scale, high-dimensional settings.
PaperID: 7072, Poster
Abstract: Electromagnetic signal modeling is increasingly moving toward self-supervised pretraining on heterogeneous I/Q observations. However, an observed waveform is jointly shaped by signal generation, propagation environments, and acquisition processes, which mix reusable structural regularities with observation-specific variations. Existing methods usually pursue unified pretraining with reconstruction objectives on mixed observations. This heterogeneity creates two key mismatches for reconstruction pretraining: (1) appearance mismatch, induced by propagation-dependent waveform variations; and (2) temporal mismatch, induced by acquisition-dependent sampling coordinates. These mismatches make observation-specific appearances and sampling coordinates easier to reconstruct than stable signal structures, inducing observational shortcuts. We identify this failure mode as observation-biased pretraining. To address this, we propose Structural Pretraining with Time Alignment for heterogeneous EM signals. It mitigates temporal mismatch with RoPE-fs, which maps token positions to sampling-rate-normalized physical-time positions, and reduces reliance on appearance shortcuts by replacing waveform reconstruction with masked prediction of discrete structural units. These units are organized into atomic structural units and emergent units that capture variable-length compositional patterns beyond fixed patch boundaries. Across diverse EM benchmarks, our method improves both full-supervision and few-shot adaptation over existing EM pretraining baselines, showing that the learned representations better preserve structural consistency when observations drift in waveform appearance and temporal scale.
Abstract: The optimal transport (OT) map provides a geometric transformation for aligning probability distributions and has become a useful tool in machine learning. However, existing estimators of the OT map still exhibit a gap between sharp statistical guarantees and practical parametric estimation based on stable training objectives. Theoretical estimators achieve minimax optimal convergence rates, but they are typically nonparametric and can incur demanding implementation design or inference costs. Practical estimators are parametric and scalable, but their statistical guarantees remain underexplored, and their min-max, adversarial-like training objectives can be sensitive to optimization algorithms. We propose BROT (Barycentric Regression for OT), a simple two-step method that first computes the unregularized OT plan and then fits a deep neural network (DNN) to the induced barycentric targets by least-squares regression. Under standard regularity conditions, we prove that the DNN estimator of BROT attains the minimax convergence rate, when the ground-truth OT map is Lipschitz. Numerical studies on synthetic datasets and an image dataset show that BROT provides accurate map estimates, strong target distribution matching, and competitive transport costs, compared to existing estimation methods. Experiments on two downstream tasks, single-cell perturbation prediction and unsupervised domain adaptation, further suggest that the accurate estimation of BROT can translate into stronger task performance.
PaperID: 7074, Poster
Authors: Mohsen Malmir, Mohamed A Radwan, houssam nassif, Murat Bayir
Abstract: Offline evaluation of ad ranking policies in multi stage delivery systems is fundamentally challenging because action propensities are often unavailable, limiting inverse propensity scoring (IPS) estimators. While reward model based estimators such as Direct Method (DM) do not rely on propensity estimates, they can suffer from bias when candidate policies induce impression distributions that differ from the logging mixture. In this work, we formalize DM based counterfactual replay under mixture logging and propose a mixture targeted shift correction based on density ratio weighting through domain classifier. We further derive a per policy value estimation error and characterize an asymptotic error ceiling governed by the cold start mass. Experiments on both ads delivery simulator and one week of production traffic from a large scale industrial CTR ranking system demonstrate that the proposed mixture weighted reward model consistently outperforms its unweighted counterpart under distribution shift.
PaperID: 7075, Poster
Abstract: Large language model training with sequence-parallel data-parallel (SP-DP) is bottlenecked by the interplay between intra-step activation redistribution and inter-step gradient synchronization. Existing infrequent-communication methods like FedAvg keep states local or reset them lack convergence guarantees and are unstable under long-context sequence-parallel regimes where each device holds only a partial activation shard; Local Adam synchronizes all states jointly and is convergent but triples the collective payload, and its uniform synchronization period cannot distinguish the semantically mandatory intra-step All-to-All from the temporally deferrable inter-step AllReduce. We propose FAUST (Federated Asynchronous Update with Staggered Timescales), a compiler-orchestrated training system that composes compile-time sequence-parallel graph rewriting with tiered runtime synchronization, assigning independent averaging periods to parameters (K_x), first moments (K_u = 3K_x), and second moments (K_v = 6K_x) according to their impulse-response half-lives, while preserving convergence under the orthogonal collective constraint that intra-step All-to-All must retire before inter-step AllReduce begins. Our analysis shows that while first-moment drift dominates the convergence rate in-distribution, high-fidelity convergence guarantees require at least periodic synchronization of the second moment. Experiments on language models up to 1.3B parameters show that FAUST incurs 303× less wall-clock communication overhead than standard DDP and 1.78× less than the previous state-of-the-art DES-LOC, while achieving perplexity within 4.6% of the fully-synchronized baseline in 27% less wall-clock time. On bandwidth-constrained links, FAUST delivers 48--303× speedup over DDP.
PaperID: 7076, Poster
Abstract: Predicting spatial gene expression from histology images could scale spatial transcriptomics (ST) to image-only cohorts, but conventional histology-based ST prediction is trained and evaluated mainly by per-gene spatial-profile reconstruction. This objective is misaligned with a key downstream use of ST: differentially expressed gene (DEG) discovery, where genes are ranked for a biological or morphology-defined contrast by evidence of between-group expression differences. We formulate image-based differential expression ranking (IDER), which asks whether predicted expression profiles preserve the contrast-specific ranked gene list obtained from measured profiles. IDER compares gene rankings induced by differential-expression statistics, rather than raw expression magnitudes or per-gene spatial correlations. We further introduce a differentiable IDER objective that aligns these statistics across genes and can be trained with morphology-derived proxy contrasts without predefined biological group labels. Experiments on public ST datasets show improved DEG-ranking agreement and pathway-enrichment overlap over conventional reconstruction objectives, including morphology-derived and pathologist-annotated tissue-region evaluations.
Abstract: Real-world robot task planning must operate under both stochastic action execution and partial observability, yet constructing Partially Observable Markov Decision Process (POMDP) models for real robotics domains remains difficult and labor-intensive. We introduce PO-PDDL, a symbolic formulation of POMDPs that preserves the relational structure and LLM-friendly syntax of the Planning Domain Definition Language (PDDL), while explicitly modeling partial observability, stochasticity, and beliefs. Building on this formulation, we propose a demonstration-driven pipeline for learning PO-PDDL models. The proposed method reconstructs latent symbolic state trajectories from real-robot execution videos, identifies partial observability via inconsistencies between inferred states and visual observations, and learns stochastic transition and observation models accordingly. The resulting PO-PDDL domains are reusable across tasks and enable online belief-space planning under both perception and execution uncertainty. Experiments on real-world long-horizon manipulation tasks show that our method consistently outperforms existing PDDL and POMDP model-learning approaches, achieving robust task planning under uncertainty with significantly lower planning cost.
PaperID: 7078, Poster
Abstract: Style transfer aims to render target content in the style of a reference image, but existing methods often suffer from content leakage, where objects, layouts, or semantics from the style reference appear in the generated output. Although prior data-driven and training-free methods can reduce leakage, they often face a leakage--degradation dilemma: stronger content suppression may weaken style fidelity, while richer style preservation may reintroduce unwanted reference content. We identify this dilemma across the full style-transfer pipeline, including feature separation, feature-space grounding, and diffusion generation. To address these issues, we propose CLeaR, a training-free framework for content-leakage-resistant style transfer. CLeaR first uses Orthogonal Subspace Projection to define content-reduced style targets in each vision foundation model (VFM) feature space. It then performs Ensemble Inversion, which optimizes a shared pixel-space style anchor satisfying style constraints across multiple VFMs. Finally, Energy-Guided Calibration maintains style alignment during diffusion sampling by steering the denoising trajectory toward the ensemble-defined style manifold. We further provide a theoretical analysis showing that the style-anchor estimation error decreases with the number of VFMs. Experiments on StyleBench demonstrate that CLeaR improves style alignment, reduces content leakage, and achieves better LLM-as-Judge evaluation compared with existing methods.
PaperID: 7079, Poster
Abstract: Metamaterials are artificially engineered structures whose mechanical and physical behaviors are strongly shaped by geometry rather than composition. Voxel representation provides a unified format for metamaterial geometry generation, as it can express diverse classes such as truss, shell, and porous structures within a single cubic discretization. However, voxel-based generation faces a plausibility–novelty trade-off: staying close to known geometries helps preserve geometric regularities, while moving away from them is necessary for novelty but may produce degenerate geometries. To address this challenge, we propose ReGDiff, a generative framework that couples voxel representation with latent space regulation and guided diffusion. ReGDiff introduces a repel-and-sink (RAS) mechanism to smooth the latent distribution of plausible geometries, and short-range repulsion (SRR) guidance to discourage generation overly close to known samples while maintaining geometric plausibility. We further contribute a voxel-based benchmark covering truss- and shell-type metamaterial geometries, together with an evaluation module for geometric plausibility, novelty, and diversity. Experiments show that ReGDiff outperforms voxel-based generative baselines, achieving +8.9% in geometric plausibility, +46.4% in novelty, and +128.6% in diversity on average across two datasets. These results suggest that ReGDiff is a strong geometry candidate generator for downstream evaluation. Our code is provided at https://anonymous.4open.science/r/ReGDiff-6DC6.
Abstract: We propose Drifting Field Policy (DFP), a non-ODE one-step generative policy built on the drifting model paradigm. Under online RL finetuning, we frame the policy update as a reverse-KL Wasserstein-2 gradient flow toward a soft target policy, realized by DFP as a steepest-descent step on probability space. By construction, this gradient step is decomposed into an ascent toward higher action-value regions and a score matching with the anchor policy as a trust region. We further derive a simple, tractable surrogate of the otherwise intractable update loss, akin to behavior cloning on top-K critic-selected actions. We find empirically that this mechanism uniquely benefits the drifting backbone owing to its non-ODE parameterization. With one-step inference, DFP achieves state-of-the-art performance on several manipulation tasks across Robomimic and OGBench, outperforming ODE-based policies.
PaperID: 7081, Poster
Authors:
Shinya Gongyo, Ryosuke Ogasawara, Tatsuya Moe, Masafumi Mori, Yusuke Sekikawa, Mitsuru AmbaiAbstract: Mixed-precision quantization improves the accuracy-efficiency trade-off by assigning different bit-widths across a network. Since bit-width selection is a combinatorial problem, existing methods often optimize additive surrogate losses instead of directly evaluating task losses in all configurations. We show that these surrogates can suffer from surrogate-task mismatch, especially when many layers are changed simultaneously from a reference configuration. We further observe that, although allowing more layers to change improves the surrogate objective, the evaluated task loss can be minimized at an intermediate limit on the number of changed layers. Based on this observation, we propose Iterative Integer Linear Programming (I2LP), which iteratively refines a reference bit configuration. At each refinement step, I2LP solves multiple ILPs with different limits on how many layers may change from the current reference, and accepts the best candidate only when it reduces the task loss. I2LP applies to both post-training quantization and quantization-aware training, and consistently outperforms uniform-precision baselines and existing mixed-precision methods. Code will be released.
PaperID: 7082, Poster
Abstract: The reasoning trajectory of the Large Language Model (LLM) is often regarded as the verbalized description of the internal thinking. However, the unfaithfulness of the reasoning process introduces the risk of shortcut reasoning, where the model fails to reason step by step but instead relies on discovered shortcuts to reach the final answer, while post-rationalizing this decision through a seemingly coherent verbalized reasoning process. This shortcut reasoning is difficult to detect, as existing monitors and verifiers mainly inspect textual reasoning traces or final outcomes, failing to capture how the model’s answer belief forms during generation. To figure out the intrinsic pattern in shortcut reasoning, we propose ConfLens, a framework that tracks how a model's confidence in its final answer evolves throughout the reasoning process. Across three shortcut reasoning settings, we find that shortcut samples often exhibit premature confidence, characterized by high confidence in the final answer at early reasoning stages. However, reliably detecting this pattern remains challenging, as existing confidence estimation methods are limited in generalizability, reliability, and efficiency. To address this, we introduce the Distributional Answer Commitment Score (DACS), a distributional confidence estimation method that instantiates ConfLens for effective shortcut reasoning detection. DACS estimates the entropy of the model's probability distribution over answer commitment during reasoning, thereby capturing how concentrated the model's answer belief is at each reasoning step without access to the ground-truth or task-specific verifier. We further convert the detection results of ConfLens into interpretable signals to mitigate reward models’ preference for shortcut reasoning responses. Experiments on math and code reasoning tasks show that ConfLens instantiated with DACS improves shortcut reasoning detection by over 4.3% in F1 compared with strong baselines, while reducing the gap between faithfulness and correctness in reward model preferences.
PaperID: 7083, Poster
Abstract: Anomaly detection is a fundamental task for time-series analytics, with important implications for the downstream performance of many applications. Despite the large number of anomaly detection methods proposed in the literature, recent benchmark studies have shown that no single detector performs best across highly heterogeneous time series. Therefore, a practical and scalable solution is to develop a model-selection method that, for a given time series, selects the anomaly detector most likely to perform well. Nevertheless, the model selection approach proposed in the literature suffers from a significant drop in accuracy when applied in Out-of-Distribution (OOD) settings. In this paper, we tackle the aforementioned limitation and propose \textscRAMSAD, a train-free, retrieval-based framework for model selection in time-series anomaly detection. Given a new time series, our method queries a knowledge base of previously observed instances, retrieves the top-k most similar series, and transfers detector recommendations from their performance profiles. The framework can operate with standard similarity measures as well as embedding-based representations. Overall, we demonstrate that similarity-based retrieval constitutes a strong and efficient foundation for model selection in time-series anomaly detection.
PaperID: 7084, Poster
Abstract: Many high-throughput scientific experiments produce distributional outputs, requiring conditional generative models to interpret the data, and to predict outcomes at new inputs. Acquiring experimental data at new conditions (e.g., for a new antigen or a new cell type) is often expensive, motivating the need for automated experimental design methods. However, existing active learning frameworks are limited as they assume scalar or vector-valued outputs. We first show that the Wasserstein test risk of a conditional generative task decomposes into a mean term and a shape term, demonstrating why mean-based active learning is blind to distributional shape. To capture both, we propose the transport Neural Tangent Kernel (tNTK), a tractable kernel that measures how sensitive a generative model's predicted distribution is with respect to its parameters. The tNTK recovers the standard NTK in the deterministic limit and upper-bounds the posterior Fréchet variance of the prediction under small parameter perturbations. It admits a closed form for affine Gaussian transport and tractable Monte Carlo estimators otherwise. Treating the tNTK as a Gaussian process covariance function, we greedily select inputs that minimize the total GP posterior variance across the pool, yielding a distribution-aware active learning strategy. We evaluate this strategy across four classes of conditional generative models, including conditional flow matching and conditional Monge gap, on synthetic distributions, intent-labeled sentences, and single-cell perturbation responses. In these experiments, the tNTK outperforms the baseline NTK computed from the conditional mean, demonstrating the value of distribution-aware active learning for generative-modeling tasks.
PaperID: 7085, Poster
Authors: Anirban Sarkar, Alejandra Duran, Peter Koo
Abstract: Oracle-guided biological sequence design must improve predicted function without moving outside the sequence distribution where the oracle is trustworthy. We introduce Generative Population Annealing GPA, a test-time sampler for sequence design with a frozen generator and frozen oracle. GPA instantiates annealed Sequential Monte Carlo at population scale: particles are initialized from a pretrained sequence prior, reweighted by an oracle reward tilt, selectively upsampled as effective sample size falls, mutated with the pretrained generator, and returned as the design pool. The base sampler targets a reward-tilted prior; practical variants shape the proposal or modify the objective to trade activity, specificity, fidelity, and diversity. Across enhancer and promoter benchmarks, GPA scales to thousands of sequences and is competitive with inference-time samplers, gradient-based editors, tree-search diffusion, and RL-fine-tuned generators. The same inference loop is evaluated with masked discrete-diffusion and autoregressive DNA generators. GPA reaches high predicted activity while preserving strong motif fidelity and model-based likelihood, although specialized baselines retain higher diversity or stronger k-mer fidelity in some settings. Cross-oracle and sequence-level audits suggest fewer obvious reward-hacking pathologies, including shorter homopolymer runs than CTRL-DNA in the HepG2 audit (5.1 bp mean maximum run versus 20.9 bp), but all validation remains computational.
PaperID: 7086, Poster
Abstract: Recent fMRI foundation models differ substantially in the spatial scale at which they represent brain activity. ROI- and connectivity-based models are efficient but coarse, whereas voxel-level models preserve fine-grained spatial structure but require specialized architectures and costly fMRI-specific pretraining. We ask whether part of this performance gap reflects the importance of preserving cortical geometry. Motivated by evidence that macroscale brain activity is strongly constrained by brain geometry, we introduce FlatClip, a training-free surface-level baseline that renders cortical activity as geometry-aware flatmap sequences and reuses a frozen SigLIP2 image encoder with only a lightweight downstream probe. Because FlatClip is not pretrained on fMRI, it provides an out-of-domain reference for evaluating how much information can be recovered from cortical geometry without relying on fMRI-specific pretraining data. Across resting-state and visual-fMRI benchmarks, FlatClip outperforms ROI/FC-style controls and several general-purpose fMRI foundation baselines, while voxel-level models remain a strong upper comparison. Geometry-control experiments show that disrupting cortical topology reduces performance, whereas mapping ROI-level signals back into atlas-defined 4D volumes partially recovers performance but remains below full voxel inputs. Together, these results position cortical geometry as a key factor in fMRI representation learning and establish surface-level flatmap sequences as a practical middle-ground baseline between ROI and voxel models.
PaperID: 7087, Poster
Abstract: Standard training of diffusion models typically relies on a forward diffusion process anchored by ground-truth target images. However, in image fusion tasks, the scarcity of real fused images makes it difficult to formulate a forward process, thereby precluding the standard training. Fortunately, our visual analysis reveals that image fusion can be effectively modeled without standard training or reliance on the real images. Nevertheless, the misalignment of input image pairs remains a significant bottleneck for fusion quality. Recognizing that one-off registration is ill-suited to the progressive generation of diffusion models, we propose a unified registration-fusion framework driven by a time-aware joint optimization mechanism. Specifically, we design a time-aware registration network to progressively optimize registration to guide fusion, while the fusion results provide feedback constraints to the registration network at each time step. This mechanism facilitates joint optimization of registration and fusion, significantly improving the quality of fused results. Experimental results show that the proposed method performs excellently on multiple datasets, validating its effectiveness and superiority.
Abstract: Video anomaly detection (VAD) aims to automatically identify events that deviate from normal patterns in untrimmed surveillance videos. Existing methods universally depend on large-scale annotations or task-specific training procedures, severely limiting their rapid deployment to novel scenes. We observe that intermediate-layer features of pre-trained multimodal large language models (MLLMs) already encode rich anomaly semantics, yet existing approaches rely on the language output pathway and fail to exploit the geometric discriminability latent in these representations. Based on this finding, we propose SphereVAD, a fully training-free, zero-shot VAD framework that recasts anomaly discrimination as von Mises–Fisher (vMF) likelihood-ratio geodesic inference on the unit hypersphere, unleashing latent discriminability through principled geometric reasoning rather than learning new representations. Specifically, SphereVAD first applies Fréchet mean centering to unfold feature distributions and eliminate domain biases, then employs Holistic Scene Attention (HSA) to reinforce feature consistency using cross-video priors, and finally performs vMF-guided Spherical Geodesic Pulling (SGP) to align ambiguous segments with directional prototypes on the spherical manifold. This training-free pipeline requires only minimal synthetic images for calibration. SphereVAD establishes new state-of-the-art results among training-free approaches on three major benchmarks and remains competitive with fully supervised baselines. Code will be available upon acceptance.
PaperID: 7089, Poster
Abstract: LLM-based multi-agent systems (LLM-MAS) coordinate well on reliable channels but fail in poorly understood ways when communication is lossy. We measure these failures against a content-blind reference: synchronous push gossip on idempotent payloads, which admits classical bounds in diameter, conductance, and percolation. Our \emphpaired-Oracle protocol runs an LLM agent and a deterministic gossip Oracle on identical scheduler, topology, and channel realisation, so any outcome difference is attributable to per-node decision logic. On Distributed Common-Free-Slot (\textscDCFS), a set-union task isomorphic to all-to-all rumour spreading, three signatures of a \emphsemantic residual emerge as channel reliability is swept: a phase-boundary shift \Delta p^ \approx 0.35 identified directly from paired trials; a 16--46% rate of consensus on wrong answers despite complete knowledge, a regime classical gossip cannot enter; and a 3.6×--5.1× round-count overhead at perfect channels. A mediation analysis shows that coverage and failure-mode labels are jointly a sufficient statistic for accuracy --- no within-stratum residual remains --- so the LLM--Oracle gap is fully captured by a p-dependent stochastic mapping over four failure modes. This decomposition supports a four-tier mitigation hierarchy whose tier-wise gains match the externalised stratum's marginal share, and a five-parameter scaling law over an eight-model panel with leave-one-model-out RMSE 0.08.
PaperID: 7090, Poster
Abstract: Training-free safeguards for text-to-image diffusion models often rely on a reusable safety signal, such as an unsafe direction or global toxic subspace, applied broadly across prompts. We provide a controlled geometric analysis of this global-unsafety assumption and reveal a consistent coverage-selectivity trade-off: compact unsafe subspaces fail to cover heterogeneous unsafe semantics, whereas broader aggregation increasingly distorts safety-adjacent benign prompts. Motivated by this finding, we propose CALM (Counterfactual Adaptive Local Modulation), a training-free safeguard that replaces uniform global removal with prompt-local counterfactual correction. Using matched unsafe-benign anchors, CALM routes each prompt to active unsafe categories, minimally edits only violating token representations toward the safe side, and suppresses positively aligned unsafe residual components. Across broad evaluation, CALM improves unsafe-content suppression while preserving benign utility, demonstrating that local counterfactual correction provides a more selective alternative to global unsafe-signal removal.
PaperID: 7091, Poster
Authors: Rohan Ghosh, Mehul Motani
Abstract: The ability of neural networks to generalize is fundamentally shaped by their computational units. While the perceptron provides a first-order, linear inductive bias and the radial basis function (RBF) unit offers an isotropic curved bias, we identify a critical opportunity for a structured extension: the Generative Matching Unit (GMU). Each individual GMU captures complex dependencies by treating the forward pass as an inference problem; its internal generative model optimizes instance-specific latent parameters to compute a reconstruction error, effectively measuring a point-to-manifold distance as opposed to the point-to-point distance in RBFs. We focus on linear GMUs, which yield closed-form analytical expressions for fast computation and extend naturally to convolutional architectures. Our theoretical analysis demonstrates that linear GMUs are highly flexible universal approximators capable of recovering structured class posteriors. We prove that while a single GMU can exactly emulate an RBF kernel and two can emulate the decision boundary of a perceptron, emulating a single k-order GMU’s decision boundary requires a hidden layer of standard units that scales polynomially or exponentially with input dimensionality. Furthermore, we prove that the GMU’s point-to-manifold distance remains discriminative in high dimensions where standard Euclidean distances fail, offering unique robustness to the curse of dimensionality. Finally, we show that GMUs offer superior flexibility and efficiency in learning smooth manifold decision boundaries compared to MLPs and RBFs. Motivated by the theoretical results, we place GMUs in the first network layer like RBFs, where their role is to enhance linear separability for subsequent layers. This avoids the typical gradient collapse problem with stacking RBF-like units while ensuring the generalization benefits remain. We find that across extensive experiments involving 27 tabular datasets, five vision datasets evaluated across 34 test-time corruption settings, and 30 synthetic scenarios, GMU networks demonstrate statistically significant improvements in generalization and robustness.
Abstract: Whether in agentic workflows, social studies, or chat settings, large language models (LLMs) are increasingly being asked to replace humans in choosing which goals to pursue, rather than completing predefined tasks. However, the assumption that LLMs accurately reflect human preferences for goal setting remains largely untested. We assess the validity of LLMs as proxies for human goal selection in a controlled, self-directed learning task borrowed from cognitive science. Across five models (GPT-5, Gemini 2.5 Pro, Claude Sonnet 4.5, Qwen3 32B, and Centaur), we find substantial divergence from human behavior. While people gradually explore and learn to achieve goals with diversity across individuals, most models exploit a single identified solution or show surprisingly low performance, with distinct patterns across models and little variability across instances of the same model. Chain-of-thought reasoning and persona steering provide limited improvements, and our conclusions hold across experimental settings. While they await confirmation in applied settings, these findings highlight the uniqueness of human goal selection and caution against its replacement with current models.
PaperID: 7093, Poster
Abstract: Deep search capabilities have become an indispensable competency for frontier large language model (LLM) agents, yet training high-performance deep search agents remains constrained by three key challenges: as synthesized reasoning graphs grow in complexity, verifying answer correctness and uniqueness becomes increasingly difficult to scale; heavy dependence on online search APIs makes data synthesis costly and hard to scale; and the lack of fully open synthesis pipelines severely hinders reproducibility.To bridge this gap, we introduce OpenSearcher, the first deep search agent training pipeline with zero online API dependency throughout the entire data synthesis and model training process, built upon two core technical innovations: (1) Distinguishing Set-grounded QA synthesis with programmatic verification: we compute minimal Distinguishing Sets on Wikidata to uniquely identify target entities, recursively expand them into reasoning DAGs, and transform the structured graphs into natural language questions via progressive generation and fuzzification. Answer uniqueness is deterministically verified through SPARQL conjunctive queries, complemented by lightweight LLM filtering for linguistic quality. (2) Fully offline training pipeline, which builds an offline corpus of tens of millions of documents from publicly downloadable Wikidata, Wikipedia, and Web Crawl Dumps, and completes trajectory synthesis, SFT cold start, and reinforcement learning entirely offline. Experimental results demonstrate that OpenSearcher achieves strong performance across deep search benchmarks, establishing a new state of the art among fully offline methods, for example, achieving 32.3% on BrowseComp and 72.8% on GAIA. We fully open-source the complete data synthesis pipeline to democratize offline deep search agent research and foster a more transparent, collaborative ecosystem.
Abstract: The key--value (KV) cache is a major bottleneck in long-context inference, where memory and computation grow with sequence length. Existing KV eviction methods reduce this cost but typically degrade performance relative to full-cache inference. Our key insight is that full-cache attention is not always optimal: in long contexts, irrelevant tokens can dilute attention away from useful evidence, so selective, learnable eviction can improve generation rather than merely approximate the full cache. We introduce a global retention-based KV eviction method that learns each token's future utility under a unified memory budget. Lightweight retention gates assign utility scores to cached KV entries, and a shared final scoring projection calibrates these scores across all layers and heads. This enables a single global eviction policy in which tokens from different layers, heads, and modalities compete directly for cache capacity. We further provide theoretical analysis showing that preferentially retaining useful tokens reduces attention dilution, and we justify geometric retention as a query-agnostic proxy for future utility. Across long-context language and vision--language benchmarks, our method substantially reduces KV memory while matching or surpassing full-cache inference. These results suggest that learned, globally calibrated KV eviction is not only a compression technique, but also a mechanism for improving long-context reasoning.
Authors: Andre Opris
Abstract: This paper investigates the role of dynamic population sizes in evolutionary multi-objective optimization. Although such approaches are widely used in practice, their benefits remain poorly understood, and rigorous runtime analyses explaining when and why they help are still scarce. To address this, we introduce the bi-objective problem class CLIMB and analyze the runtime of GSEMO and the widely used NSGA-II on this problem. Our results show that allowing a dynamic population size for NSGA-II can lead to a moderate improvement, yielding a speedup of order \Omega(\sqrtn/\log n). In particular, we prove that GSEMO and NSGA-II-DYN, a version of NSGA-II with dynamic population sizes we propose in this paper, can find the Pareto front of CLIMB in expected O(n \log n) fitness evaluations, whereas NSGA-II with a fixed population size requires \Omega(n^1.5) fitness evaluations in expectation. To the best of our knowledge, this is the first rigorous runtime analysis in multi-objective optimization demonstrating a super-constant speedup of GSEMO over NSGA-II. Our analysis builds on concepts from single-objective optimization, like the evolution of population diversity over time, and employs the well-known family-three method to prove the lower bound.
Abstract: Generative models, including diffusion models, are increasingly used as foundation models and adapted through sequential fine-tuning, making continual learning an essential problem setting. However, continual learning in such generative models remains poorly understood: after a task change, what aspects of the learned distribution are most easily lost, and what replay samples should be prioritized? We address these questions through the modern Hopfield energy. Recent links between modern Hopfield networks (MHNs) and diffusion models allow analyses in MHNs to be transferred to diffusion models. We introduce intrinsic forgetting as an increase in Hopfield energy after the task change. In tractable settings in an MHN, we prove that high-energy, outlier-like samples undergo a larger energy increase than cluster-like samples, implying that samples located in sharp, isolated basins are more forgettable. We further analyze memory replay and show that replay is particularly effective for high-energy samples, enabling an energy-based selection of replay samples. We validate these predictions in experiments on MHNs and two diffusion models under continual-learning settings: Stable Diffusion and a pixel-space DDPM. In these diffusion models, Hopfield energy tracks reconstruction-based forgetting, and replay experiments reveal energy-dependent mitigation of forgetting that is consistent with the MHN analysis.
PaperID: 7097, Poster
Abstract: Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive LLMs by enabling non-autoregressive text generation. However, their practical deployment remains limited by inefficient inference, largely due to the absence of effective Key-Value (KV) caching and scalable parallel decoding mechanisms. Existing acceleration methods typically study KV caching and parallel decoding in isolation, overlooking the I/O bottlenecks that arise when cache reuse and parallel token verification are jointly applied. In this work, we introduce \bf Flash-dLLM, a training-free inference acceleration framework for fast and memory-efficient dLLMs. Flash-dLLM first identifies GPU memory I/O as a dominant bottleneck in KV-cache-enabled dLLM inference and addresses it with an I/O-aware fused KV-cache kernel that reduces redundant memory movement. Building on this optimized cache mechanism, Flash-dLLM further proposes an efficient KV-cache-driven draft-and-verify decoding strategy, where the dLLM itself serves as both drafter and verifier without requiring an auxiliary model. This unified design enables faster decoding while preserving generation quality and improving scalability to longer sequences and larger models. Extensive experiments on mathematical reasoning and code generation benchmarks demonstrate that Flash-dLLM consistently outperforms all existing dLLM acceleration methods in both inference speed and memory efficiency.
Abstract: Scientific discovery is constrained not only by what is true, but by what is cognitively available to the researchers currently exploring a field. Many directions are coherent in light of the literature yet unlikely to be proposed because no existing community occupies the right combination of concepts, methods, and intuitions. Modern language models inherit this bias, recombining high-density regions of the literature when prompted for novel ideas. We introduce a framework that targets the complementary region, which we call the alien space of science, where directions are plausible under the structure of existing knowledge but unlikely under the distribution of existing researchers. Our method first decomposes papers into granular conceptual units and clusters them into a shared vocabulary of \emphidea atoms. It then learns two complementary models over this vocabulary. A \emphcoherence model scores whether a combination of atoms forms a viable research direction, and an \emphavailability model scores whether any existing author community is positioned to produce a given combination. Sampling alien directions then reduces to ranking atom combinations that maximize coherence while minimizing availability. On a corpus of 16,068 peer-reviewed LLM papers from NeurIPS, ICLR, ICML, and major NLP venues, the resulting sampler explores a 3.5\text--7× broader effective atom vocabulary than frontier LLM ideation baselines without sacrificing coherence, and produces ideas that match or exceed those baselines under blind LLM, human, and downstream experimental evaluation. By separating scientific plausibility from community availability, our framework points toward AI ideation that complements rather than merely accelerates human science, expanding exploration into coherent directions the current community is unlikely to pursue.
PaperID: 7099, Poster
Authors: Han Bao
Abstract: Gradient descent has been of particular interest in modern machine learning beyond sole focus on optimization because some specific structures emerge through the optimization dynamics. Such behaviors result from optimization even though the learning objective does not explicitly encode the target structure, collectively called implicit bias, often preventing overparametrized models from fitting to spurious patterns. A typical instance is the max-margin implicit bias of a linear classifier, widely established for exponentially tailed loss functions. Even after having a given dataset separated, the parameter vector continues to evolve towards the max-margin direction asymptotically along the gradient descent dynamics. This phenomenon corroborates a frequent empirical observation of ``train longer, generalize better.'' However, the max-margin convergence is an asymptotic phenomenon, and what is worse, this asymptotic convergence rate is significantly slower than convergence in optimization. Even so, the parameter vector along gradient descent dynamics commonly correlates with the max-margin direction positively (though not exactly) within considerably fewer iterations than the asymptotic rate. By shedding another light on this classical yet profound problem, this work aims to understand the mechanism of this early-stage alignment phenomenon. Our theoretical results demonstrate that the parameter vector weakly aligns with the max-margin direction within O(\exp(\exp(-\delta))) iterations, where \delta>0 is the permissible alignment error, which is shown to be tight. By tracking the radial and tangential flows, our proof operates on the alignment dynamics directly with dataset geometry and gets rid of the asymptotic expansion, which is a key insight to enabling faster weak alignment.
PaperID: 7100, Poster
Authors:
Yumin Zheng, Qingtian Zhu, Ziyan Zhu, Kailu Song, Yinqiang Zheng, Jun DingAbstract: Spatial transcriptomics (ST) provides spatially resolved gene expression measurements, but most 3D ST datasets are acquired as independently processed 2D sections. This sliced acquisition creates a difficult reconstruction problem: measurements are dense within each section but sparse along the axial direction, and inter-section technical variation can be confounded with true biological change. Existing coordinate-based implicit neural representations (INRs) provide a continuous modeling framework, but standard isotropic coordinate encodings do not explicitly reflect this axial-to-lateral imbalance. We present HoloGene, an anisotropic conditional INR framework for estimating continuous 3D gene-expression fields from pre-registered ST sections. HoloGene first compresses high-dimensional expression profiles into a graph-autoencoder latent manifold. It then uses pseudo-label-derived biological signature conditioning to provide coarse intra-section biological context, while depth-dependent FiLM modulation models axial variation. To reduce slice-specific covariance shifts, we introduce a signature-stratified correlation-alignment regularizer over latent features. Using held-out measured sections as cross-sectional references, HoloGene improves reconstruction accuracy and biological preservation metrics compared with general-purpose INRs and adapted spatial-transcriptomics baselines. Ablation analyses show that decoupled depth conditioning improves axial generalization, while correlation alignment provides the strongest reduction in slice-specific representation leakage. These results support anisotropic conditional modeling as a practical strategy for estimating continuous 3D expression fields from discontinuous ST sections.
Abstract: Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. While such objectives support task inference from context, they tie the learned policy to the quality of offline actions and can fail when trajectories are weak or suboptimal. We propose Q-Target Pretrained Transformers (QTPT), which preserves context-conditioned inference but replaces behavior cloning with a Bellman-style Q-target objective. QTPT learns to estimate action values from contextual rewards and transitions, rather than simply imitating the behavior policy. We analyze QTPT in stochastic linear bandits and finite-horizon MDPs, deriving suboptimality bounds that separate offline-data effects from Transformer approximation error. Empirically, QTPT is most beneficial under weak offline data across controlled RL benchmarks, and controlled ablations show that context conditioning is essential while Bellman/TD targets provide additional gains. We further include D4RL experiments as higher-dimensional stress tests under matched Transformer baselines.
PaperID: 7102, Poster
Authors: Byeongju Kim, Hoonki Lee, Byungjun Kim, Inyul Ra, Yongjin Kim, Yongseong Lee, Sangyeob Kim
Abstract: Despite growing interest in on-device LLM deployment, large-scale Mixture-of-Experts (MoE) models remain impractical on resource-constrained devices, as sparse expert activation still requires all expert weights to be memory-resident. To address this, we identify three key observations for on-device MoE inference: (a) cache hit rates are fundamentally bounded even under Belady’s optimal policy; (b) these bounds are insufficient for low-latency on-device inference; and (c) prior cache-aware routing based on binary cache presence can induce routing instability. Based on these insights, we propose CECAR (Cache & Expert Co-Aware Routing), a MoE inference system that jointly considers expert routing and cache management. Unlike prior cache-aware routing methods, CECAR uses an ML-based cache policy to predict the expected reuse distance of each expert and prioritizes routing and execution toward experts with smaller predicted reuse distances. On the Qwen3-30B-A3B model, CECAR achieves 12.5 tokens/s on a consumer-grade GPU with only 16 GB of VRAM, where non-resident experts are fetched from SSD, delivering a 4.5× speedup over conventional MoE inference and exceeding the decode speed of baselines that keep the entire model fully resident in DRAM.
Authors:
Rishab Balasubramanian, Pin-Jie Lin, Rituraj Sharma, Mohit Bansal, Tu VuAbstract: We investigate whether post-trained capabilities can be transferred across model scales without retraining, and propose the Master Key Hypothesis, which states that model capabilities correspond to directions in a low-dimensional latent subspace that induce specific behaviors and are transferable across models through linear alignment. Based on the hypothesis, we introduce UNLOCK, a training-free and label-free framework that extracts a capability direction by contrasting activations between capability-present and capability-absent Source variants, aligns it with a Target model through a low-rank linear transformation, and applies it at inference time to elicit the behavior. Experiments on reasoning behaviors, including Chain-of-Thought (CoT) and mathematical reasoning, demonstrate substantial improvements across model scales without training. For example, transferring CoT reasoning from Qwen1.5-14B to Qwen1.5-7B yields an accuracy gain of 12.1% on MATH, and transferring a mathematical reasoning direction from Qwen3-4B-Base to Qwen3-14B-Base improves AGIEval Math accuracy from 61.1% to 71.3%, surpassing the 67.8% achieved by the 14B post-trained model. Our analysis shows that the success of transfer depends on the capabilities learned during pre-training, and that our intervention amplifies latent capabilities by sharpening the output distribution toward successful reasoning trajectories.
PaperID: 7104, Poster
Authors: Wesley H Holliday, Adam Lesnikowski, Roberto-Rafael Maura-Rivero
Abstract: Post-training methods for large language models (LLMs) implicitly aggregate diverse human feedback in ways that can be studied explicitly using social choice theory. In this paper, we adopt the perspective that for each post-training method, there is a corresponding reference-dependent probabilistic voting rule. Taking Direct Preference Optimization (DPO) and Nash Learning from Human Feedback (NLHF) as our case studies, we study axiomatic properties of the associated voting rules. First, we show that NLHF violates one of the central axioms of social choice, namely monotonicity, if and only if its KL regularization parameter is below a bound, whereas DPO satisfies monotonicity regardless of its regularization parameter. Second, we show that while DPO is often associated with Borda-style preference aggregation, when DPO is regarded as a probabilistic voting rule, it violates two variable population axioms that the standard probabilistic version of Borda satisfies. Third, we show that unlike DPO, NLHF violates a commutativity axiom from Bayesian epistemology, which has practical implications for staged post-training. Finally, we conduct an empirical study of the frequency and magnitude of axiom violations in the CoVal dataset using real base models.
PaperID: 7105, Poster
Authors: Zenggui Chen, Jianwei Ma
Abstract: Semi-supervised medical image segmentation leverages abundant unlabeled data to reduce the reliance on costly dense annotations, yet its dominant pseudo-labeling paradigm suffers from inevitable prediction noise, especially under scarce supervision. In this work, we observe that parametric network predictions and non-parametric prototype-based predictions exhibit complementary behaviors under pseudo-label noise: the former offers high expressiveness in capturing local structures but is sensitive to erroneous supervision, while the latter provides stable semantic guidance by aggregating class-level statistics but may lack spatial precision. Motivated by this observation, we propose a simple yet effective complementary learning framework that jointly exploits the parametric prediction space and the non-parametric prototype space for semi-supervised segmentation. Specifically, we introduce a confidence-aware selection mechanism that identifies reliable voxels from either prediction space, a prototype entropy minimization objective that sharpens the non-parametric semantic anchors, and a dual-space semantic alignment objective that encourages mutual enhancement between the two spaces. The unified objective enables reliability-aware mutual enhancement, where each space guides the other when confident and noisy alignment is suppressed when both are uncertain. Extensive experiments on multiple medical image segmentation benchmarks demonstrate that our method consistently outperforms existing approaches under various label-scarce settings.
Abstract: We introduce Sticky Jump Diffusions (SJDs), continuous-time Markov processes on \mathbb R^d with a discrete anchor set identified with token embeddings. In forward time, anchors release their mass at a hazard rate and the released mass diffuses in the continuous ambient space; time reversal couples a score-driven SDE with a sticky jump kernel whose rate and destination are fixed by flux balance with the forward law. We estimate the score and the per-anchor reverse hazards from a single denoising classifier via Denoising Hazard Matching, the hazard analogue of denoising score matching, with simulation-free cross-entropy training. SJD generalizes the classical sticky-boundary diffusions of Feller and Itô to vocabulary-sized anchor sets, and recovers masked diffusion, continuous diffusion, and hybrid diffusion as limits. Beyond these limits, the framework exposes two design axes that previous work hold fixed: the un-sticking kernel, which encodes both per-anchor geometry and cross-position structure of the corruption, and the per-anchor forward hazard, which encodes the commitment schedule. We evaluate SJD on CIFAR-10, ImageNet 64×64, Text8, and Sudoku, where it is competitive with discrete, continuous, and hybrid baselines.
Authors: Qiyuan Qiao, Ge Yuan, Can Wang, Dong Xu
Abstract: High-quality teleoperation datasets are costly to collect, particularly for hard tasks. We observe that many tasks exhibit directional asymmetry: completing the forward hard task is difficult, whereas reversing it by relaxing or disrupting the environment is comparatively easy. This suggests that reversed easy-task trajectories can serve as a scalable supervision signal for the hard task, reducing the cost of manual demonstration collection. However, reversed data can be noisy, and directly training on it may yield suboptimal policies. To enable largely automated acquisition and effective use of reversed data, we propose a teleoperation-cost effective framework for hard policy learning via temporal reversal of easy tasks, consisting of three key components: a closed-loop data collection pipeline that alternates between hard-task and easy-task policies to autonomously reset the environment and generate diverse trajectories; a hierarchical data refinement pipeline that temporally inverts easy-task rollouts and filters low-quality motion using kinematic priors and a critic-guided advantage filter; and an iterative policy learning method that trains the hard-task policy using both initial reversed easy-task demonstrations and the filtered reversed data in a continuous online learning loop. By combining automated collection, hierarchical refinement, and iterative learning, our method enables scalable, reliable training of complex, high-precision manipulation tasks. Across two simulated benchmarks and real-robot experiments, we demonstrate that our method improves hard-task success rates with higher data efficiency and more stable training compared to reversal-based and reinforcement-learning baselines, without requiring extensive hard-task teleoperation.
PaperID: 7108, Poster
Abstract: Post-training pruning of large language models (LLMs) zeroes a fraction of the weights and reconstructs the survivors without retraining, addressing the weight-memory bottleneck of LLM serving. At a fixed mask the per-row reconstruction admits a closed-form Optimal Brain Surgeon (OBS) optimum, but the cost of the exact joint solve is cubic in the layer dimension; one-pass methods such as SparseGPT and Wanda approximate it at quadratic cost without attaining the per-block joint optimum. We introduce Block-OBS-GS, which solves the per-block joint OBS reconstruction exactly within SparseGPT's complexity class via a per-row Cholesky factorisation of the in-block damped Hessian (one third the cost of an explicit inverse, reused across the algorithm) and follows it with a single cross-block Gauss--Seidel sweep. We prove three guarantees: per-layer FLOP count in SparseGPT's complexity class, with the leading mn^2 constant dropping from 1/2 (SparseGPT) to 3\rho/2 (Block-OBS-GS) where \rho=1-s, hence strictly lower for sparsity s>2/3; per-row reconstruction error at most SparseGPT's, strict whenever the activation Hessian has non-zero off-diagonal coupling; and a closed-form distance bound to the unregularised joint-OBS optimum that contracts geometrically per sweep. Across six Llama-3 and Qwen-3 backbones from 3B to 70B parameters, Block-OBS-GS attains the lowest WikiText-2 perplexity at every tested sparsity on the three Llama-3 backbones, with a 47% perplexity reduction over SparseGPT on Llama-3.1-8B at s=0.875; the highest seven-task commonsense Avg7 on six of twelve populated cells; 5-shot Massive Multitask Language Understanding (MMLU) within 0.026 of SparseGPT across populated cells; and the lowest WikiText-2 perplexity on four of five backbones under 2:4 structured sparsity.
Abstract: Agent skills are emerging as an important attack surface in LLM-based systems. Through an empirical study of existing skill scanners, we find that current defenses primarily rely on textual descriptions, manifests, and source code as the main signals for security analysis, which can leave visually conveyed malicious intent insufficiently examined. This creates a practical blind spot: harmful operational instructions hidden in images may bypass scanning while still being recoverable by multimodal agents during deployment. To systematically investigate this threat, we propose \textscSkillCamo, a document-mediated multimodal instruction attack that conceals malicious instructions within images bundled with a skill while rewriting the surrounding documentation to naturally reference those images as part of the normal workflow. Thus, the attack does not rely on the image alone, but on the joint interpretation of textual guidance and visual payload at execution time. To defend against such attacks, we further propose \textscExecScan, an execution-grounded multimodal scanning module that performs intent extraction, behavior reconstruction, abuse assessment, and deliberative execution simulation over skill artifacts. \textscExecScan jointly analyzes documentation, code, referenced resources, and visual content to recover hidden instructions, reconstruct executable behavior chains, and identify downstream risks such as exfiltration, destruction, persistence, deception, and privilege escalation. Extensive experiments show that image-hidden malicious instructions challenge existing skill scanners, while \textscExecScan can improve the skill scanning performance.
PaperID: 7110, Poster
Abstract: Vision-language models (VLMs), such as CLIP, have shown remarkable transferability in the biomedical domain, and prompt tuning enables efficient adaptation with limited supervision. In practice, biomedical images exhibit pronounced acquisition-induced style shifts, especially across imaging sites and protocols, while cross-modality transfer further introduces compound shifts from different physical acquisition processes. However, existing prompt-tuning methods ignore acquisition-induced style shifts, making the prompts sensitive to source-specific style and limiting the transferability to unseen medical domains. To address this, we formulate acquisition-induced style shifts as mixture shifts, where each domain is viewed as a mixture of latent style components. Under this formulation, the target risk depends on both mixture weights and latent component risks, motivating prompt adaptation that accounts for style-specific variation rather than relying on a single source-specific prompt. In this paper, we propose SaMoE, a Style-aware Mixture-of-Experts Multimodal Prompt Learning framework for adapting VLMs to biomedical domains under acquisition-induced style shifts. To account for latent component risks in prompt adaptation, SaMoE represents multimodal prompts with latent style-conditioned prompt transformations and composes them through image-conditioned routing, enabling each image to induce style-adapted visual prompts. By extracting style cues from intermediate visual representations, the sparse expert composition adapts style-related prompt weights without domain labels and is injected across multiple VLM layers to align with hierarchical visual-language representations. Extensive experiments on 15 medical datasets across 12 modalities and 10 organs demonstrate significant improvements in both accuracy and generalizability over state-of-the-art methods.
PaperID: 7111, Poster
Authors: Kazuma Ikeda, Kotaro Oishi, Ryosei Hara, Ryo Yoshida, Mariko Isogawa, Kentaro Yoshioka
Abstract: LiDAR is a critical sensor for autonomous driving and robotics, yet it remains vulnerable to adverse conditions such as fog, rain, transparent objects, and low-reflectance surfaces. Full-waveform LiDAR (FWL) addresses these limitations by capturing the complete return waveform as a histogram, preserving rich reflection characteristics including multi-path echoes and weak signals. While recent Transformer-based approaches have demonstrated promise on FWL data, two fundamental bottlenecks remain: high computational cost from redundant tokens in low-intensity background regions, and limited generalization due to scarce annotated training data. We address both bottlenecks by exploiting the physical structure inherent to FWL data. First, we propose FWL-ToPM, a token pruning and merging mechanism that leverages intensity distributions to selectively reduce tokens while retaining critical waveform content. Second, we introduce FWLAug, a physically consistent data augmentation framework that preserves time-of-flight waveform structure and frustum geometry to improve generalization across diverse environments. Evaluated on the Ghost-FWL benchmark, FWL-ToPM achieves over 8× speedup at real-time throughput alongside a nearly 9-point F1 gain, establishing the first real-time FWL Transformer that simultaneously improves accuracy. FWLAug further improves generalization to both standard and out-of-distribution environments. Together, these methods push FWL-based perception toward practical deployment.
Authors: Ha Dang, Sebastian Schmidt, Jürgen W Hesser
Abstract: Neural operators have achieved strong performance in learning solution operators of partial differential equations (PDEs), but their inherently continuous representations struggle to capture discontinuities and sharp transitions. Existing approaches typically approximate such features within continuous function spaces, often requiring increased model capacity and high-resolution data. In this work, we propose Cut-DeepONet, a two-stage training framework that explicitly models discontinuities while reducing learning complexity. Our approach reformulates the problem via a lifting strategy, partitioning the domain into smooth subregions while representing discontinuities as boundaries in a higher-dimensional space. This separation aligns the operator learning task with the inductive bias of neural networks and avoids directly approximating discontinuities. An additional network predicts input-dependent discontinuity locations for unseen inputs, which are then used to guide the neural operator in generating smooth components within each region. Experiments on benchmark PDEs show that Cut-DeepONet outperforms state-of-the-art methods, even when trained on low-resolution datasets. The method excels on problems with discontinuities and sharp transitions, while using fewer trainable parameters. Our results highlight the benefits of changing the representation of operator learning rather than increasing model complexity.
PaperID: 7113, Poster
Abstract: Transformer-based language models are widespread in today's society. As such, understanding the mechanisms by which they solve structured tasks and predicting how they may behave in novel scenarios is of great importance for safe deployment. We study the learning dynamics of attention heads in a controlled setting by training a decoder-only Transformer (GPT-J) on two structurally equivalent multi-hop reasoning tasks: a number task requiring positional reasoning and a letter task requiring symbolic reasoning. Using a recently introduced metric that classifies attention-head behavior as positional or symbolic for a given prompt, we show that successful learning is associated with the emergence of pure heads, i.e., heads that express themselves as either positional or symbolic. Despite the tasks' structural equivalence, they impose different mechanistic demands: the number task requires both positional and symbolic heads, whereas the letter task requires only symbolic heads. We then identify the computational roles of these heads, characterize the basic functions they implement, and give theoretical constructions showing how single-layer RoPE-based attention can realize these functions through geometrically interpretable query, key, and value operations. This analysis yields a quantitative separation between positional and symbolic mechanisms in their robustness to longer sequences, formalized through a novel notion of discrepancy. We empirically validate the resulting predictions in both controlled and real-world models, showing that symbolic mechanisms extrapolate more reliably to longer sequences while positional mechanisms face sharper limitations.
PaperID: 7114, Poster
Authors: Min H Kim, Hyun Ki, Seokbong Yoo
Abstract: Medical visual question answering (VQA) and federated learning (FL) have gained prominence as critical methods for clinical artificial intelligence, with VQA supporting image-based diagnostic reasoning and FL allowing institutions to train models jointly without sharing sensitive patient data. However, existing approaches face substantial challenges in cross-modal vertical FL settings, where each participating institution possesses unimodal, modality-specific textual data without access to paired multimodal samples. To address this limitation, we introduce SynTeX-FL, a federated framework designed to enable cross-modal text transfer for medical VQA without sharing raw data. Moreover, SynTeX-FL applies a reconstruction-based text synthesis model that extracts clinically relevant semantics from medical images to generate cross-modal textual representations. Furthermore, the proposed method incorporates modality-specialized low-rank adaptation (LoRA) modules to enhance modality-aware representations for each imaging domain. During aggregation, the central server employs discriminator-based quality scores and source-aware weighting strategies to integrate heterogeneous client contributions. Extensive experiments on multimodal medical VQA benchmarks demonstrate that SynTeX-FL significantly outperforms current FL baselines, providing a robust and efficient solution for cross-modal reasoning in decentralized clinical environments.
PaperID: 7115, Poster
Authors: Marcus Triplett, Kenneth Kay
Abstract: Humans have a remarkable ability to judge causal relationships from a limited number of unreliable observations. Past work on causal cognition has largely focused on normative accounts of human behavior, leaving unknown how biologically plausible neural systems could learn causal relationships from observations and update their representations of causal structure with additional evidence. Here, we leverage task-optimized recurrent neural networks to discover candidate implementation-level neural mechanisms of causal judgment. We propose a novel cognitive task in which a subject observes stochastic samples from an unknown causal structure (e.g. among variables A, B, and C with unknown causal relationships), and must judge whether a specific causal relationship is present given a query (e.g. "does A cause C?"). We found that, after training, recurrent neural networks perform the task with high accuracy, adopt strategies that incorporate the behavior of non-queried variables to form their judgments, and, despite being trained only on pairwise queries ("does A cause B?", "does C cause A?", etc), form implicit beliefs about the complete graphical structure underlying the observations. Lastly, we use dynamical systems analysis to identify a set of low-level neural mechanisms that implement causal judgment and representation of causal graphical structure. Together, these findings lay the groundwork for a "bottom-up" approach to causal cognition, providing a potential basis for subsequent experimental study in the brain.
PaperID: 7116, Poster
Authors:
Van K Zhi, Songnan Lin, Zhiping Lin, Bihan WenAbstract: Many existing multivariate time series anomaly detection methods formulate anomaly detection as a reconstruction problem, where detections are derived from reconstruction errors in the time domain. While effective for large deviations, reconstruction-based scoring struggles with subtle anomalies where only a small subset of features deviates slightly. In such cases, most features are well reconstructed, causing the overall reconstruction error to remain low and indistinguishable from normal data. To address these challenges, we propose Flow+Diff, a novel framework that detects anomalies via discrepancies in a latent space, where two complementary representations, induced by distinct modeling paradigms, diverge under anomalous inputs. One representation is obtained via a normalizing flow, which provides an invertible mapping that preserves temporal dependencies and captures inter-feature dependencies. The other is produced by a diffusion model that learns the distribution of normal samples in the latent space via a stochastic denoising process. When given anomalous inputs, the normalizing flow maps the inputs to latent representations that lie in low-probability regions of the learned latent distribution, while the diffusion model generates representations aligned with normal data distributions, resulting in a pronounced discrepancy between the two. We compute the anomaly score based on this latent-space discrepancy, thereby shifting anomaly detection from the time domain to the latent space. Finally, the invertibility of the normalizing flow enables reconstruction in the time domain, facilitating interpretation of anomalous features. Extensive experiments across six benchmark datasets show that Flow+Diff achieves state-of-the-art performance on four datasets against competitive baselines.
Abstract: How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead of memorizing it, remains a fundamental mystery. To address this, we introduce analytically tractable Bayesian information restricted diffusion (BIRD) models, in which each pixel observes restricted information about noisy data. A BIRD model time-reverses diffusion by inferring which past training sample produced its current restricted observation using the Bayesian posterior. This model class generalizes existing analytical diffusion models that use spatially local information restriction. We show that spatially local BIRD models closely approximate trained diffusion models early in training, across different architectures such as UNets and DiTs. Under minimal assumptions on the data distribution, we identify an information-theoretic phase boundary between memorization and generalization in the joint space of amount of training data, time in the reverse generative process, and amount of information restriction: a BIRD model memorizes when the mutual information between its restricted noisy observations and the training data exceeds the log number of training points, and it generalizes otherwise. Experiments across a range of datasets confirm our theoretically predicted location for the transition. We find that generation proceeds near the edge of memorization: both spatially local BIRD models and early-training diffusion models track the memorization-generalization phase boundary by increasingly restricting information over time. Overall, our results reveal a fundamental role for information restriction in generative AI to circumvent the curse of dimensionality.
PaperID: 7118, Poster
Authors:
Ao Chen, Yijing Lu, Pengpeng ChenAbstract: Multivariate time series often arrive with non-uniform timestamps and asynchronous per-channel observations, violating regular-grid assumptions of standard sequence models, while scarce labels and limited training resources make learning tasks more challenging. Model-space learning offers a lightweight alternative by fitting a dynamical model to each sequence and performing downstream analysis in the resulting model space. However, under irregular sampling, channels with very few observations weakly constrain the per-channel regression, leaving the fitted readout an unreliable summary. In this paper, we propose Reservoir State Statistics (RSS), a readout-complementary representation that augments the readout with three inversion-free trajectory summaries: centroid, spread, and terminal state. RSS is computed on the trajectory of the Global Decay Reservoir Network (GDRN), a continuous-time reservoir with per-channel time decay, cross-channel global pooling, and event-driven masked updates for asynchronous multivariate observations. Across multiple datasets, GDRN-RSS achieves competitive few-shot accuracy with CPU runtimes of seconds.
PaperID: 7119, Poster
Abstract: Standard embodied evaluations do not independently score whether an agent correctly commits to task completion at episode closure—a capacity we call terminal commitment. Behaviorally distinct failures—never completing the task, completing it but failing to stop, and reporting success without sufficient evidence—collapse into the same benchmark failure. We introduce VIGIL, an evaluation framework that makes terminal commitment independently measurable. Under VIGIL’s default protocol, agents observe only egocentric RGB, receive no action-success signals, and must end each episode with a semantic report checked deterministically against hidden world state. This yields two separate scores: world-state completion (W) and benchmark success (B), where B additionally requires a correct terminal report. This decoupling makes four outcome categories distinguishable: missed execution, post-attainment drift, unsupported commitment, and verified success. Across 20 models on 1,000 frozen episodes, systems with comparable W differ by up to 19.7 pp in B: one model converts achieved states into correct reports, while another with near-identical execution drifts past the goal without closing. An action-feedback intervention further tests the separation: execution-oriented signals improve W broadly, yet commitment failures persist in models that do not already ground terminal reports in the achieved state. VIGIL provides a protocol that makes terminal commitment independently visible and scorable.
Abstract: Instruction-based image editing (IIE) models have recently demonstrated strong capability in modifying specific image regions according to natural language instructions, which implicitly requires identifying where an edit should be applied. This indicates that such models inherently perform language-conditioned visual semantic grounding. In this work, we investigate whether this implicit grounding can be leveraged for zero-shot referring image segmentation (RIS), a task that requires pixel-level localization of objects described by natural language expressions. Through systematic analysis, we reveal that strong foreground-background separability emerges in the internal representations of these models at the earliest denoising timestep, well before any visible image transformation occurs. Building on this insight, we propose a training-free framework that repurposes pretrained image editing models for RIS by exploiting their intermediate representations. Our approach decomposes localization into two complementary components: attention-based spatial priors that estimate where to focus, and feature-based semantic discrimination that determines what to segment. By leveraging feature-space separability, the framework produces accurate segmentation masks using only a single denoising step, without requiring full image synthesis. Extensive experiments on RefCOCO, RefCOCO+, and RefCOCOg demonstrate that our method achieves superior performance over existing zero-shot baselines.
PaperID: 7121, Poster
Authors:
Ruiyi Fang, Gezheng Xu, QIUHAO Zeng, Zihao Jing, Jiale Cai, Zhihao Li, Hao Zheng, Xuanting Xie, Shuo Wang, Bingheng Li, RUIZHI PU, Zhao Kang, Charles Ling, Boyu WangAbstract: Graph domain adaptation (GDA) aims to transfer knowledge from a labeled source graph to an unlabeled target graph. Existing GDA methods typically assume that the source graph provides sufficiently complete information for reliable cross-domain alignment. However, this assumption is often unrealistic in practice, where only a small sampled source graph may be available due to privacy, annotation, storage, or data-collection constraints. In this paper, we study Cross-Scale Graph Domain Adaptation (CSGDA), where adaptation is performed from a limited-scale source graph to a larger target graph. Our empirical results show that CSGDA exhibits larger local and global discrepancies than conventional GDA, as the source graph provides only limited local structures and limited coverage of the global structure. To address this challenge, we propose a local-global structural alignment framework that, at the local level, augments neighborhood structures to compensate for limited local structural information. At the global level, it extracts global structural patterns and aligns source and target distributions. An adaptive GCN-MLP module then integrates the compensated local structures and extracted global patterns for domain alignment and target prediction. Experiments across multiple graph adaptation benchmarks and sampling ratios demonstrate that our method consistently improves adaptation in the presence of severe source-graph incompleteness.
PaperID: 7122, Poster
Authors: Xi Chen, Anindya De, Rocco A Servedio
Abstract: Goldwasser, Shafer, Vafa and Vaikuntanathan (STOC 2025) recently introduced a formal framework for defending against backdoors that may be planted in large-scale ML models by malicious model developers. They gave several algorithmic results in their framework for efficiently ``mitigating'' the effects of such backdoors by leveraging ideas that were developed in theoretical computer science in the 1980s, namely \emphrandom self-reducibility and \emphself-correction. However, the approaches of Goldwasser, Shafer, Vafa and Vaikuntanathan only provably achieve secure mitigation under restrictive assumptions about the ground-truth population data distribution that the ML model is trained on. In this work we apply tools that have been developed quite recently in the theoretical computer science research area known as \emphtolerant property testing to achieve secure backdoor mitigation for a much broader class of population distributions than could be handled by prior work. Our approach naturally provides a way for an ML model user to select a hypothesis class from a very wide range of possibilities for the mitigated ML model, and naturally enables the ML model user to control a tradeoff of the mitigated model's accuracy against its security, efficiency, and interpretability.
PaperID: 7123, Poster
Authors:
JUNGMIN LEE, Gwangeun Byeon, Yulhwa Kim, Seokin HongAbstract: Pruning has emerged as a promising direction for accelerating large language model (LLM) inference. However, many existing methods rely on offline calibration data, making them sensitive to distribution shifts between calibration data and inference inputs. In this paper, we introduce Token Filtering, a lightweight online pruning method that selectively skips attention computation for redundant tokens during inference, without requiring any calibration data or fine-tuning. Token Filtering identifies redundant tokens based on joint key–value (KV) similarity and bypasses their attention computation. This approach reduces both compute and KV cache size while preserving essential contextual information. To preserve accuracy while meeting the target pruning ratio, we restrict pruning to later layers, which are typically less sensitive to pruning, with a layer-wise threshold that adaptively tracks the target pruning ratio. Extensive experiments on LLaMA3.1-8B, LLaMA3.1-8B-Instruct, and Qwen3-8B demonstrate that Token Filtering consistently achieves better accuracy–efficiency trade-offs than prior methods. In long-context generation, Token Filtering achieves up to 1.5X higher accuracy than the best-performing pruning baseline while reducing latency by up to 44% compared to the dense baseline.
Abstract: Agent skills extend LLM agents with reusable instructions, tool interfaces, and executable code, and users increasingly install third-party skills from marketplaces, repositories, and community channels. Because a skill exposes both executable behavior and context-setting documentation, its deployment risk cannot be measured by single-shot audits or prompt-level red teams alone: a realistic attacker can use audit and runtime feedback to repeatedly rewrite the skill. We frame this risk as adaptive leakage—whether a budgeted attacker can iteratively revise a skill until it passes audit and produces verified runtime harm—and present Proteus, a grey-box self-evolving red-team framework for measuring it. Proteus searches a formalized five-axis skill-attack space. Each candidate is evaluated through a unified audit-sandbox-oracle pipeline that returns structured audit findings and runtime evidence to guide cross-round mutation. Beyond initial evasion, Proteus performs path expansion, which finds alternative implementations of successful attacks, and surface expansion, which transfers learned implementation patterns to new attack objectives beyond the original seed catalogue. Across eight phase-1 mutator-target-defender configurations, Proteus achieves 40-90% ASR@5 and exhibits positive learning-curve slopes on both evaluated auditors. In the full 8-cell expansion matrix, Proteus generates 438 jointly bypassing and lethal variants; SkillVetter is bypassed at \geq 93% in every expansion cell, and AI-Infra-Guard, the strongest public auditor we evaluate, still admits jointly successful variants at up to 41.3%. These results show that current skill vetting substantially underestimates residual risk when evaluated against adaptive, feedback-driven attackers. Code: \urlhttps://anonymous.4open.science/r/proteus/.
PaperID: 7125, Poster
Abstract: Vision--language models (VLMs) such as CLIP retain zero-shot recognition ability, but their open-world use often requires task-agnostic continual learning: tasks arrive sequentially, past data are unavailable, task identity is unknown, and the backbone cannot be retrained. We argue that this setting fails VLMs through three coupled pressures beyond stability--plasticity tradeoff: (1) modal-consistency drift, (2) representation-space interference, and (3) per-task capacity shortfall. We propose COMET (COllaborative Mixture-of-Expert Transfer) to resolve these three issues respectively as follows. A frozen teacher provides only a feature reference on a shared image--text pool; Tri-Affinity Distillation preserves image--image, text--text, and image--text geometry; and Student-Aware Cross-Expert Distillation aligns new experts with compatible prior experts so routing errors degrade gracefully. We design COMET to keep capacity and inference separate: a contextual bandit keeps or merges full per-task experts using student-side signals only, while a reconstruction router selects sparse experts at test time without teacher semantics. On the X-TAIL benchmark, we show that each proposed components indeed improves the continual learning performance. COMET shows that continual vision--language adaptation can preserve multimodal geometry, coordinate accumulated experts, and allocate capacity jointly when feature reference, expert shaping, routing, and capacity control are separated by construction.
Abstract: Continuous latent-space reasoning offers a compact alternative to textual chain-of-thought for multimodal models, enabling high-dimensional visual evidence to be integrated without explicit reasoning tokens. However, we identify a previously overlooked optimization pathology in existing latent visual reasoning methods: although visual latents become semantically enriched during training, their contribution to final answer prediction is systematically suppressed. Within the shared parameter space, the autoregressive objective favors shortcut reliance on direct visual input, driving latent tokens toward transition-like states rather than informative reasoning content. We term this phenomenon Silenced Visual Latents. To address it, we disentangle the two conflicting objectives by directly optimizing the latent reasoning at inference time, keeping backbone parameters frozen. In Stage I, visual latents are warmed up via query-guided contrastive latent--visual alignment, improving semantic quality while preventing latent collapse. In Stage II, the latent reasoning is further optimized via a confidence-progression reward, which incentivizes predicted token distributions along the latent span to become progressively more concentrated, routing predictions through the latent reasoning rather than bypassing it. Experiments across eight benchmarks and four model backbones show that inference-time latent optimization, without any parameter updates, effectively unleashes the suppressed reasoning capacity of visual latents.
PaperID: 7127, Poster
Abstract: Multimodal Large Language Models (MLLMs) are typically designed under the assumption that all modalities available during training will also be accessible at inference. However, many real-world settings violate this assumption, requiring models to operate under a privileged modality setting, where auxiliary modalities are available only during training. While these modalities contain valuable information, existing MLLMs largely fail to leverage them effectively, as they treat modalities as interchangeable inputs rather than sources of complementary supervision. We propose Mixture of Probes (MoP), a novel framework that disentangles modality-specific and modality-general signals within the MLLM, allowing the model to preserve modality-dependent structure while learning transferable representations across modalities. At its core, MoP achieves this through a structured probing mechanism that extracts and organizes information from intermediate representations of a shared modality encoder, rather than relying only on final-layer alignment as done in existing MLLMs. To support this disentanglement, we further introduce MoP Cross-modal Probe Training (MoP-X), a training strategy for MoP centered around a probe disentanglement loss that prevents probe collapse and encourages cross-modal learning. We evaluate MoP across two domains spanning eight tasks and four modalities under a comprehensive evaluation protocol tailored to the privileged modality setting, where each modality is independently treated as the sole input at inference time. MoP consistently outperforms strong MLLM baselines, achieving up to 65% relative improvement, demonstrating that auxiliary modalities, even when unavailable at inference, can provide substantial gains when effectively leveraged during training.
PaperID: 7128, Poster
Abstract: Transformers excel at sequence modeling, yet Softmax attention incurs quadratic complexity and unbounded KV cache growth. While linear attention offers a promising alternative, existing approaches lack systematic functional comparison with Softmax attention, rigorous error analysis, and a theoretically grounded improvement roadmap. We address this gap by framing linearization as KV cache approximation and establishing a principled pathway from Softmax attention to linear models. Our analysis identifies five critical components—redundancy elimination, token-level quantization with positional separation, positional compression, inter-layer similarity, and multi-state decomposition—each accompanied by theoretical justification and error bounds, with explicit connections to existing mechanisms. Building upon this framework, we introduce GPA, a linearized attention model that inherits pretrained weights and achieves state-of-the-art results. GPA outperforms strong baselines including MVA and GSA across multiple benchmarks, while requiring less fine-tuning resources. Our work provides both theoretical clarity and practical guidance for advancing linear attention, charting a principled course toward efficient, scalable alternatives to Softmax attention.
Authors:
Aditya Parikh, Stella Christina Frank, Sneha Das, Aasa FeragenAbstract: : group-conditional label errors that cause systematic performance disparities across demographic subgroups. Label bias in image segmentation remains underexplored, as even detecting it typically requires clean, unbiased annotations, which are not readily available. We present a data-centricadaptation of Confident Learning to segmentation, allowing detection of label bias directly in the training data without a clean, unbiased ground truth. By comparing the provided training labels to the model's confident predictions, we isolate directional errors that quantify the presence and nature of bias, where standard overlap metrics like Dice fail. We further show that label bias influences subgroup separability in the encoder's feature space, an artifact we leverage for bias mitigation rather than suppressing it. We evaluate three datasets, spanning from synthetic to real-life bias, showing how our framework reliably detects and mitigates bias without access to clean labels, achieving equitable performance across experimental conditions.
PaperID: 7130, Poster
Authors:
Jianjie Fang, Yongyan Xu, Ziyou Wang, Yuchao Huang, Zhaolu Wang, Rongze Tang, Mingyuan Jia, Baining Zhao, Weichen Zhang, Xin Zhang, Haisheng Su, Yu Shang, Chen Gao, Wei Wu, Xinlei Chen, Yong LiAbstract: With the rapid progress of interactive video generation, video generation world models have gradually emerged as one of the mainstream paradigms in world model research and are increasingly regarded as a promising path toward efficient intelligent agents. However, existing video generation world models are typically developed under different forms of control supervision and mainly focus either on interactive world modeling or embodied world modeling, leaving the compatibility of heterogeneous control signals largely unexplored. In this work, we introduce , the first training framework that enables unified learning of world models under heterogeneous supervisory controls by incorporating a Mixture-of-Experts (MoE) design into Diffusion Transformers (DiT). We further propose , a continually extensible heterogeneous training strategy for world models, which supports diverse control signals, including robotic arms, hand joints, and camera poses, within a single world model and allows the model to be progressively expanded to new control settings. By enabling joint learning from more diverse data sources, this training strategy alleviates the scaling bottleneck of current world models. Our experiments show that world models trained with MoE over heterogeneous supervision consistently outperform those trained with any single control modality alone, demonstrating clear mutual gains across different control types. UniMoE-World achieves state-of-the-art performance on the WorldArena benchmark and shows particularly strong advantages in both locomotion and hand-motion capabilities over existing methods.
PaperID: 7131, Poster
Abstract: Spiking Neural Networks (SNNs) transmit information through energy-efficient discrete spike events, unlike the continuous activations used in conventional Artificial Neural Networks (ANNs). Existing SNN robustness methods often inherit ANN-oriented objectives that suppress continuous perturbations or stabilize membrane-potential dynamics. In this paper, we argue that SNN robustness should be analyzed through the stability of spike patterns under perturbations. We formalize perturbation-induced spike flips and introduce the flip rate as a training-agnostic metric for quantifying spike-pattern instability. We further show that spike flips are closely associated with membrane potentials near the firing threshold, and introduce margin occupancy to quantify this threshold-margin vulnerability. Based on these findings, we propose Margin-Resculpted Learning (MRL), which regularizes excessive threshold-margin occupancy with an activity-adaptive bound. MRL reduces vulnerable spike flips while preserving the threshold region needed for surrogate-gradient learning. Experiments across static and event-based datasets, architectures, surrogate functions, and attack settings demonstrate consistent robustness improvements.
Abstract: Real-world image super-resolution aims to recover high-quality images from complex and unknown real-world degradations. However, existing generative Real-ISR methods largely inherit the dense latent representations and quadratic-cost global modeling paradigm developed for high-resolution image synthesis, causing computation, memory usage, and inference latency to scale unfavorably with resolution and thus limiting practical deployment. We argue that the key bottleneck lies not in insufficient restoration priors, but in excessive token redundancy and costly token interactions during high-resolution restoration. Motivated by this observation, we revisit Real-ISR from the perspectives of compact latent representation and linear-complexity modeling, and propose SANA-SR, an efficient one-step restoration framework. Specifically, SANA-SR employs a deep compression autoencoder with a 32× compression ratio to drastically reduce latent tokens while preserving restoration-relevant structures and textures. On top of this compact latent space, we introduce a linear-attention DiT with LoRA fine-tuning, enabling efficient high-resolution restoration with linear-complexity token mixing. Extensive experiments on all benchmark datasets demonstrate that SANA-SR achieves highly competitive and often superior quantitative performance against existing methods, while restoring clearer and more realistic textures. Moreover, after pruning, the deployed model runs in 0.019 s with 407.95G MACs and 344M parameters, highlighting its strong potential for practical mobile deployment.
PaperID: 7133, Poster
Abstract: Post-training alignment can be fragile: in autoregressive models (ARMs), subsequent fine-tuning on counter-aligned or distribution-shifted data can erode previously aligned behavior. While this rebound phenomenon has been studied mainly in ARMs, how such fine-tuning triggers a similar rebound in diffusion language models (DLMs) remains poorly understood. We conduct extensive fine-tuning experiments across diverse datasets, perturbation sizes, model scales, alignment algorithms, and inference hyperparameters, and find that DLMs consistently exhibit a slower rebound pattern than ARMs do. Furthermore, while ARMs fine-tuned on more positive data suffer a steeper degradation under reverse updates, DLMs exhibit ordering stability: models fine-tuned on larger positive alignment sets retain higher performance as negative data increases. We propose a compression perspective to account for this behavior. Unlike ARMs, whose compression naturally follows a fixed left-to-right order, DLMs generate text through iterative masked denoising. We formalize this distinction through a template-wise compression theory for DLMs. The resulting elasticity theory explains why counter-aligned updates lead to much slower alignment erosion in DLMs. Our experiment findings indicate that rebound is shaped not only by post-training data and objectives, but also by the underlying generation mechanism.
Authors:
Chun-Peng Chang, Shaoxiang Wang, Alain Pagani, Dariu Gavrila, Holger CaesarAbstract: Modern image encoders achieve high generalization by decoupling semantic meaning from resolution, an ability yet to be fully realized in the 3D domain. We investigate the failure of 3D point cloud encoders to achieve similar generalization and find that existing models are highly sensitive to sampling resolution and scale changes, leading to significant performance degradation. This sensitivity is a major bottleneck for real-world deployment in robotics, as it suggests models overfit to specific quantization densities and object scales rather than learning invariant semantic features. To mitigate this dependency, we propose Invaria, a point cloud encoder that achieves scale and density invariance through next-resolution prediction and receptive field calibration. While our objective is not the explicit generation of high-resolution point clouds, we find that this training objective encourages the model to learn robust, structural invariants. The resulting encoder achieves significant performance gains during resolution shifts while maintaining high efficiency through a compact model size and reduced token requirements. Specifically, on ScanNet, Invaria achieves a 56.0% higher mIoU at 3× lower resolution and a 20% improvement when the objects scale is reduced by a factor of 3. These gains are achieved with a 45% smaller model size and an average reduction of 40% in input tokens. The code will be released once the paper is published.
PaperID: 7135, Poster
Abstract: Zero-shot 3D point cloud classifiers built on large-scale pretrained models (OpenShape, Uni3D, ReCon++, DuoduoCLIP) excel on clean synthetic benchmarks such as ModelNet40 but degrade sharply on real-world scans (ScanObjectNN, ScanNet) and corrupted point clouds (ModelNet-C)–a persistent real-world generalization gap. We propose SVG-3D, the first 3D extension of Support Vector Generation (SVG), to close this gap without retraining a more capable encoder or accessing target-domain test data. SVG-3D mines a discriminative kernel-machine support set from a frozen encoder paired with DiffSplat, a text-conditioned 3D Gaussian splatting generator. It rests on two design choices that preserve strict zero-shot rigour: (i) a dynamic pair-selection rule that picks the top-K most confusable class pairs from a pseudo-confusion matrix computed entirely on generated samples, never on the test set; and (ii) a Metropolis–Hastings sampler operating on the CLIP text-embedding hypersphere with a Slerp-plus-Gaussian proposal whose density admits a closed-form ratio, yielding the same acceptance rule as SVG. On ScanObjectNN (OBJ\_ONLY/OBJ\_BG/PB\_T50\_RS) and ScanNet, SVG-3D surpasses prior zero-shot 3D classifiers built on large-scale pretrained models without consulting any test data; on ModelNet-C, a complementary corruption-aware setting in which the corruption family is assumed known further improves mean robustness over the same baselines.
PaperID: 7136, Poster
Abstract: We study online identification of linear-in-parameter dynamical systems whose true parameters drift over time. Measuring non-stationarity through the path variation V_T = \sum_t=1^T-1\|\theta_t+1^\star - \theta_t^\star\|, we make two contributions. First, we give a clean blockwise dynamic-regret analysis for continuous projected online gradient descent, yielding E[R_T] = O(T/\sqrt\tilde W + \tilde W V_T) and hence the oracle-tuned rate O(T^2/3V_T^1/3) . Second, we propose a residual-based adaptive step size and prove an oracle-proxy comparison theorem: when the residual proxy tracks the local prediction-error signal with lower-order overhead, the adaptive algorithm inherits the oracle scaling. Experiments on synthetic systems and real turbofan-engine degradation data (NASA C-MAPSS) support the matched-regime slope prediction on synthetic data, diagnose the failure regime when proxy overhead dominates, and show a clear unsupervised health-indicator trend in the residual proxy on FD001. A separate fast-rate guarantee under persistent excitation and a tube MPC integration are deferred to the appendix.
Abstract: Adaptive density control in 3D Gaussian Splatting (3DGS) repeatedly grows the Gaussian population through fixed-cardinality random splitting to discover useful scene structure. However, in vanilla 3DGS, its binary split operator requires many densification rounds to expose fine details, making it a bottleneck for efficient training schedules with fewer iterations. We introduce AdpSplit, an error-driven adaptive split operator that determines the number of split children and initializes the child parameters from L1-pixel-error region statistics, enabling fewer densification iterations, thus reduced training time, while preserving the rendering quality of full-schedule training. Across the MipNeRF360, Deep-Blending, and Tanks\&Temples datasets, AdpSplit reduces the training time of multiple accelerated 3DGS pipelines by 9.2%-22.3% as a simple drop-in replacement for the standard split operator. With FastGS, AdpSplit matches the full-schedule PSNR on MipNeRF360 while reducing training time by 16.4%, corresponding to a 12.6× acceleration over vanilla 3DGS.
PaperID: 7138, Poster
Abstract: Robot task planning in real-world environments requires mapping abstract natural language instructions to executable action sequences under long horizons and complex constraints. While large language models (LLMs) provide strong commonsense reasoning, they often fail to generate reliable and feasible plans. In contrast, symbolic planners ensure feasibility and optimality but require well-specified goals and cannot directly interpret high-level human intent. We formulate robot task planning as learning to generate verifiable subgoals in the Planning Domain Definition Language (PDDL), bridging language understanding and symbolic planning. To address the challenges of sparse and noisy supervision, we propose Trace-Guided Policy Optimization (TGPO), a reinforcement learning framework that improves structured subgoal generation through (i) verifier-grounded rewards, (ii) external correction of intermediate reasoning traces, and (iii) constrained policy updates that incorporate corrected traces into training. We evaluate TGPO on large-scale household planning tasks with long horizons, abstract instructions, and complex constraints. TGPO significantly outperforms prompting-based and reinforcement learning baselines, with the largest gains on abstract tasks. Furthermore, TGPO integrates naturally with symbolic planners and language-conditioned executors, enabling robust long-horizon planning and combinatorial generalization in realistic environments.
PaperID: 7139, Poster
Abstract: Multimodal Large Language Models (MLLMs) face significant scalability challenges in long-video understanding due to quadratic attention complexity and context window constraints. Existing retrieval-based approaches typically using supervised learning on static keyframe annotations. However, these label priors often fail to align with the actual information required by the downstream MLLM for complex reasoning, creating a fundamental misalignment between the training target and inference needs. To bridge this gap, we propose a probabilistic framework that reformulates frame selection as a Bayesian inference process. We first initialize a policy using supervised learning to capture general saliency. Then, we treat the frozen MLLM as a task environment and employ its negative log-likelihood (NLL) as dense reward evidence to update the policy via Bayesian Group Relative Policy Optimization (GRPO). This process explicitly transits the selection principle from label priors to task evidence. Extensive experiments on Video-MME, MLVU, and LongVideoBench demonstrate that our approach matches the accuracy of heavy MLLM scorers while maintaining the efficiency of lightweight models.
PaperID: 7140, Poster
Abstract: This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene structure, thereby enabling transferable multi-view geometric representations. Yet, existing encoder-based NVS methods yield poor representations. This is not because of lack of supervisory signal, but rather due to inconspicuous architectural choices: spatially expressive decoders that dilute representational capabilities of the scene encoder, and low-level pixel-space targets that hinder feature learning. We present SNAP, a self-supervised encoder-decoder transformer that addresses both through a pose-conditioned local decoder and a latent-space reconstruction objective. SNAP is general purpose, and we show that it is on par or better than special-purpose geometry-supervised methods. SNAP also outperforms self-supervised representations across five geometric tasks: visual localization, pose estimation, point correspondence, depth estimation, and robot manipulation. Remarkably, SNAP's patch tokens exhibit emergent viewpoint invariance that rivals heavily supervised models despite lower compute and data budgets. Under severe camera shifts where standard 2D representations collapse, SNAP maintains strict spatial consistency, revealing that restricting decoder expressivity actively prevents the suppression of transferable geometric structure.
Authors: Yuandong Cao, Chi C So, Jun-Min Wang, He Wang
Abstract: Physics-informed neural networks (PINNs) are powerful surrogates for differential equations but are notoriously difficult to train due to spectral bias, stiffness, and poor accuracy on high-frequency or multiscale solutions. Adversarial training based on generative adversarial networks (GANs) has recently gained surprisingly strong empirical results in improving training, but the underlying mechanisms remain elusive. To this end, we propose a new analysis framework for adversarially trained PINNs, based on the key observation of how the discriminator in GANs can influence the training dynamics of PINNs. The framework first provides a much needed theoretical grounding to why and when adversarial training is effective in PINNs, then presents a unified analysis of GANs variants in such training, and finally leads to a new, practical, efficient training algorithm for PINNs. Empirical results demonstrate that our method can significantly reduce the pathology of PINNs training, thereby providing better models with superior performances, often several magnitudes more accurate than alternative methods.
PaperID: 7142, Poster
Authors: Songyue Guo, Zhao CHEN, Caleb C Cao, Lei Chen
Abstract: Evidence-grounded medical question answering requires not only accurate answers, but also precise, verifiable, and clinically trustworthy evidence. Existing medical RAG systems typically retrieve top-(k) chunks according to item-wise relevance. Although this design improves recall, it does not explicitly optimize whether the selected evidence set is compact, sufficient, and decision-critical. In this paper, we study medical evidence-grounded QA from a different perspective: instead of retrieving evidence without optimization, the evidence-grounded medical QA system should identify a compact sufficient subset of high-density evidence snippets. We propose , a gain-aware framework for compact supporting medical evidence selection. MEGA first expands clinical queries into complementary subqueries and performs hybrid retrieval to construct a recall-oriented candidate evidence pool. It then introduces hidden-state , which uses a frozen LLM to estimate whether each candidate snippet contributes new answer-relevant information beyond surface relevance. Finally, MEGA formulates evidence selection as a budgeted utility maximization problem and proposes an to identify a compact evidence subset under an explicit token budget. We further construct two evidence-grounded medical QA benchmarks, CRC-EvidenceQA and Med-EvidenceQA, to evaluate both answer quality and evidence grounding. Across three LLM backbones and five medical RAG baselines, MEGA consistently achieves the best results, improving over the second-best baseline by (8.8%) on average in answer quality and 38.8% on average in evidence F1. Ablation studies and blinded clinical expert evaluation further validate its robustness and clinical relevance.
PaperID: 7143, Poster
Abstract: Multi-channel images encode heterogeneous channel semantics aligned over a shared spatial structure, which requires models to jointly capture globally shared structure and channel-specific variations. Recent multi-channel Vision Transformers attempt to address this challenge by augmenting the `[CLS]` token with memory tokens to organize multi-channel representations under varying channel compositions. However, we find that these context tokens collapse onto a few dominant channels under high inter-channel redundancy, producing redundant representations and a biased global summary that fails to reliably support composition-aware channel interpretation. We propose MuCa-ViT (Multi-Channel Composition-Aware Vision Transformer), a unified framework that addresses these limitations through two coupled mechanisms. Factorial Token Learning (FTL) enforces token-level decorrelation among context tokens, encouraging the `[CLS]` token to capture globally shared semantics while memory tokens preserve complementary channel-specific cues. Channel Composition-Aware Modulation (CAM) uses the FTL-refined `[CLS]` as a semantic anchor to generate sample-wise modulation signals, which enable composition-aware interpretation of channels within a unified backbone rather than relying on static channel embeddings. Experiments on JUMP-CP, CHAMMI, and So2Sat show that MuCa-ViT outperforms prior multi-channel ViTs across all three benchmarks, and the magnitude of improvement aligns with the inter-channel redundancy of each dataset (\rho \in [0.19, 0.77]), validating the diagnostic analysis that motivates our design.
PaperID: 7144, Poster
Abstract: Reconstructing speech from non-invasive brain signals offers a promising pathway for restoring communication in individuals who are cognitively intact but unable to speak. Existing EEG-to-speech approaches formulate this task as \emphacoustic reconstruction, optimizing waveform fidelity while ignoring whether the generated speech preserves high-level semantic content. In this work, we revisit this formulation and argue that EEG signals carry not only acoustic but also semantic information. We identify two key limitations of prior methods: (1) the neglect of spatial relationships between EEG electrodes, and (2) the failure to exploit the semantic structure of the N400 paradigm, where congruent and incongruent trials reflect distinct semantic processing. We propose SENSE (Semantic-EEG Neural Speech SynthEsis), which combines a graph-based EEG encoder over electrode geometry with EEG Semantic Conditioning (ESC), aligning EEG to a pretrained semantic space using only congruent trials. On the N400 dataset, SENSE consistently outperforms prior methods on both acoustic and semantic metrics, and model-internal channel attribution suggests distributed reliance on auditory, sensorimotor, and centro-parietal regions, consistent with known speech-perception neuroscience. In the unseen-subject setting, SENSE trained on only two subjects already surpasses the strongest baseline trained on all eighteen subjects in word error rate, and matches it on acoustic metrics with as few as eight subjects.
PaperID: 7145, Poster
Authors: Takahiro Kajisa
Abstract: Quantum state tomography of pure non-Gaussian states is fundamentally limited by scarce and noisy data due to measurement-induced collapse and experimental drift. We introduce a novel reconstruction framework that parameterizes the wavefunction directly via an overparameterized neural network, trained using rank-1 factored gradient descent. We rigorously prove that this formulation's loss landscape is devoid of spurious local minima and strict saddles, guaranteeing all critical points are global minima. By utilizing the max-sliced Wasserstein-1 distance, our loss perfectly maps to the structure of homodyne measurements, structurally preventing gradient vanishing in low-shot regimes. On noisy two-mode cat states, our method achieves continuous fidelity improvements, successfully bypassing the noise-floor plateaus that limit standard iterative maximum-likelihood estimation.
Abstract: Randomly initialized neural networks induce a prior over functions, but the predictor used in practice is produced only after training. We ask how much of this initial bias survives the training pipeline. To make the question measurable, we introduce initialization memory: the dependence of the validation-selected predictor on the scale of the random initialization. We perform controlled CIFAR-10 experiments on ResNets where initialization memory already sharply separates training regimes. Low-learning-rate SGD can interpolate while still remembering its initialization: on ResNet-9 with batch size b=128, test accuracy varies by 26.5 percentage points across initialization scales despite \ge99.5% training accuracy. This is not undertraining: extending the same low-learning-rate regime to 5,000 epochs leaves the spread essentially unchanged. In contrast, Adam-family methods largely erase the dependence. SGD can also be made to forget when larger learning rates are paired with explicit L_2 norm control. We interpret these findings in terms of the time scale of forgetting: gradient-flow-like dynamics can preserve initialization memory, whereas stochastic finite-step effects, explicit norm decay, and adaptive preconditioning erase it on scales governed by the size of explicit or implicit regularization. The practical inductive bias of a trained network is therefore not the architectural prior alone, but the architectural prior after being filtered by the forgetting dynamics of the training pipeline; and the same regularizers that improve generalization are precisely those that erase memory of initialization.
PaperID: 7147, Poster
Abstract: While federated reinforcement learning has traditionally guaranteed high efficiency in homogeneous environments, collaboration is often hindered when agents possess distinct environments and goals. In this paper, we flip this paradigm, framing reward heterogeneity not as a hurdle to consensus, but as a structural asset that catalyzes exploration. We first propose Personalized Federated Upper Confidence Bound Value Iteration (PF-UCBVI), which achieves optimal linear speedup by decoupling shared dynamics from personalized objectives. To eliminate the risks of explicit exploration, we then introduce Personalized Federated Exploration-Free Value Iteration (PF-EFVI), a purely greedy algorithm that leverages reward diversity to ensure state-action coverage. We prove that PF-EFVI attains logarithmic regret without explicit exploration bonuses under sufficient reward diversity. Our results show that agent disagreement is a vital resource that shifts the exploration burden from temporal complexity to spatial diversity, enabling safe and efficient collective learning.
Abstract: In-context learning (ICL) allows large language models (LLMs) to adapt to new tasks through demonstrations, yet it suffers from escalating inference costs as context length increases. While task vectors offer a promising alternative by compressing demonstrations into compact hidden-state representations, their quality has been evaluated only through downstream task accuracy. This indirect criterion provides limited insight into how to design more effective task vector extraction methods. In this paper, we posit that inference using task vectors should align their predictive distribution with that of ICL. To quantify this, we introduce d_\textNTP, a metric that measures the discrepancy in next-token probabilities between task vector-based and ICL-based inference. Our empirical analysis reveals that d_\textNTP serves as a performance proxy, exhibiting a strong negative correlation with downstream accuracy. Motivated by this, we develop Linear Task Vector (LTV), a method designed to minimize d_\textNTP via a closed-form linear mapping that estimates demonstration effects through regression. Across eight classification benchmarks and five LLMs, LTV consistently outperforms existing task vector baselines, improving average accuracy by 9.2% while reducing inference latency. We further show that LTV outperforms the baselines on regression tasks. Moreover, we investigate the transferability of LTV across different model scales; an aspect that has remained nascent in task vector research. Specifically, we empirically show that task vectors from a larger model can enhance a smaller model's performance by 6.4%, suggesting a new utility for extracted task representations.
PaperID: 7149, Poster
Abstract: Multimodal large reasoning models can reason over images and text, but they remain vulnerable to multimodal jailbreaks. Recent safety-alignment methods reduce attack success rate (ASR) by fine-tuning on reasoning traces from stronger teacher models. We show that this has a hidden cost of over-refusal. The aligned model learns to reject benign inputs that contain sensitive words or visual cues. This is especially harmful for boundary-safe inputs, where the prompt looks risky on the surface but is safe in context. We identify external teacher supervision as a key factor in this behavior. It introduces a distributional mismatch and shifts supervision away from the base model's own reasoning distribution. To address this, we propose SAGE (Safety-Aware Guided Elicitation), a self-guided data generation framework for multimodal safety alignment. SAGE guides generation in two stages: a reasoning-level guide first elicits a safety-aware trace, and a decision-level guide then elicits the final response conditioned on that trace. This factorized guidance lets SAGE shape both the reasoning process and the final safety decision without relying on external teacher traces. On LLaVA-CoT, SAGE reduces FigStep ASR from 86.2% to 2.1% and achieves an average over-refusal of 43.1, compared with 66.9--71.1 for prior reasoning-based safety-alignment methods, while preserving general utility. We will release the SAGE training dataset.
PaperID: 7150, Poster
Authors: Wenpei Shao, Ross Jacobucci
Abstract: Process reward models (PRMs) dramatically outperform outcome reward models (ORMs) for math reasoning and agent tasks, yet provide no benefit for chat and QA. No theory explains why, and practitioners choose between them by expensive trial and error. We introduce state-lift, a diagnostic that predicts whether a PRM is worth training for a new domain — in minutes, from ~500 labeled steps, before committing to expensive reward model training. State-lift measures how much step quality depends on trajectory state versus action text alone, and an accompanying effective gap distinguishes two regimes: state-noise-dominant (PRM essential) and action-gap-dominant (ORM sufficient). Across ten domains and eleven published PRM-ORM comparisons, state-lift correctly predicts PRM advantage, and a minutes-scale linear estimate matches the outcome of full Llama-3.1-8B reward model training — including CaSiNo negotiation (+0.055 lift at SL = 0.22). The accompanying two-branch decision protocol turns state-lift into an actionable PRM-vs-ORM workflow that runs in minutes on a single CPU.
Authors:
Letian Bai, Xuanming Cao, Juan Du, Chengyu TaoAbstract: Zero-shot 3D anomaly detection aims to identify anomalies without access to training data from target categories. However, existing methods mainly rely on projecting 3D observations into multi-view representations that primarily capture geometric cues rather than realistic visual semantics and process them with vision encoders pretrained on RGB data, leading to a significant domain gap between the encoder and the projected representations. To address this issue, we propose Align3D-AD, a unified two-stage framework that leverages the RGB modality from auxiliary categories as cross-modal guidance for zero-shot 3D anomaly detection. First, we introduce a cross-modal feature alignment paradigm that maps rendering features into the RGB semantic space. Unlike prior works that implicitly rely on pretrained encoders, our method enables direct semantic transfer from RGB observations. A semantic consistency reweighting strategy is further introduced to refine feature alignment by reweighting local regions according to holistic semantic consistency. Second, we propose a modality-aware prompt learning framework with dual-prompt contrastive alignment. By assigning independent prompts to RGB-aligned and rendering features, our method captures complementary semantics across modalities, while the contrastive alignment further enhances prompt representations to improve discriminability. Extensive experiments on MVTec3D-AD, Eyecandies, and Real3D-AD demonstrate that Align3D-AD consistently outperforms existing zero-shot methods under both one-vs-rest and cross-dataset settings, highlighting its generalization capability and robustness. Code and the dataset will be made available once our paper is accepted.
PaperID: 7152, Poster
Authors: Bryce Grant, Peng Wang
Abstract: A paraphrase pair is a small piece of evidence that two activations should agree. We treat that as a sheaf condition and ask what its cohomology looks like at scale. A cellular sheaf over a paraphrase graph defines a coboundary δ⁰; its kernel is the space of activations that paste consistently across the graph, and the rest of the sheaf Laplacian's spectrum reads off how badly they fail to. On a matching graph this collapses to projected within-pair covariance, so the framework only earns its keep once depth supplies cycles: stacking the same activations across the transformer's own layers turns the construction into a 2-D grid with b₁ = N(L−1) algebraic 1-cycles, one per elementary square asking whether layer-l → l+1 commutes with paraphrase equivalence. The model-derived Hodge harmonic mass at L=2 varies by four orders of magnitude across architectures (from 10⁻⁴ on Mistral-7B and 10⁻³ on Llama-3-8B to 0.08 on Llama-2-7B), and the ordering by harmonic mass matches the ordering by steering fragility we measure independently. The framework supplies a mechanistic prediction that variance-matched and probe baselines do not. Three validations follow. (i) Across nine architectures from 124M to 13B, the spectral complement of L F carries 5.6 to 26.5× the causal influence of variance-matched controls (p < 10⁻¹⁵). (ii) On held-out CounterFact across eight architectures, sheaf H⁰ at 20 dimensions beats LEACE's preserved subspace at full hidden dimension: 85.6% vs. 42.7% on GPT-2, 87.1% vs. 77.7% on Mistral-7B, mean gap +17.9 pp. (iii) On Llama-2-7B, contrastive sheaf steering preserves 7.4× more facts than random under style transfer (31.0% vs. 4.2%, n=1000, McNemar p < 10⁻⁵⁰). The same operators read off representational collapse: tr(L F) − β log det Cov reproduces VICReg's invariance + variance + covariance recipe in coordinates and holds encoder rank at 57–60/64 where naive consistency collapses to 14–42/64 across seven architectures.
PaperID: 7153, Poster
Authors:
Chongyu Qu, Ritchie Zhao, Yufan He, Dong Yang, Zhengyi Lu, Junchao Zhu, Tianyuan Yao, Juming Xiong, Junlin Guo, Yanfan Zhu, Yuechen Yang, Daguang Xu, Bennett Landman, Yucheng Tang, Yuankai HuoAbstract: Vision-language models (VLMs) applied to 3D medical scans face a structural memory wall. A single CT or MRI volume produces tens of thousands of vision tokens, and the resulting key-value (KV) cache scales linearly with both batch and context. Existing KV compression methods operate on the cache's feature or token statistics, or apply 1D and 2D transforms, leaving the three-axis spatial structure in the cache unexploited. We propose TACT-KV (Tri-Axis Cosine Transform), a calibration-free KV compression method that exploits strong spatial correlations in visual KV states along all three anatomical axes of 3D scans. This correlation concentrates the KV cache energy under a frequency-domain transform. TACT-KV applies a separable 3D discrete cosine transform to the KV cache, allocates a per-layer bit budget across frequency bins via reverse water-filling (more bits to high-energy components), and quantizes each bin with a fixed scalar codebook. We derive a close-form relation: the distortion reduction from adaptive bit allocation is dominated by the spectral concentration of the transformed cache. Across six medical VLMs and two 3D CT benchmarks, TACT-KV maintains near-lossless prediction fidelity at 1-bit. The same advantage reproduces on four general-purpose VLMs evaluated on video tasks, indicating that the benefit of exploiting three-axis structure generalizes beyond medical imaging. At equal high-bandwidth memory (HBM), TACT-KV scales the decoding batch size on a single GPU by up to 8×.
PaperID: 7154, Poster
Abstract: Modern large language models (LLMs) often exhibit overconfident errors, assigning high confidence to incorrect predictions and thereby undermining reliability in real-world use. Existing calibration methods typically rely on post-hoc adjustment or retraining, operating at the output level without addressing the underlying causes of miscalibration. We propose CoCHI (Class-oriented Causal Head Intervention), a framework that improves calibration by identifying and intervening on internal components responsible for distinct confidence behaviours. By analysing predictions through correctness–confidence patterns, CoCHI enables targeted intervention on model internals, reducing overconfident errors while preserving correct predictions. Our approach provides a mechanistic perspective on calibration, linking reliability to identifiable structures within the model rather than treating it as an output-level issue. Experiments across multiple LLMs show consistent improvements in calibration metrics, including Expected Calibration Error, Negative Log-Likelihood, and Brier Score, while maintaining accuracy on five multiple-choice question answering benchmarks. These results suggest that effective calibration requires not only adjusting outputs, but also addressing the underlying causal mechanisms within the model. Implementation details and code are provided in the supplementary material.
Abstract: Current end-to-end autonomous driving models are fundamentally constrained by the behavioral cloning ceiling of imitation learning. While reinforcement learning offers a path to smarter autonomy, it demands two missing pieces of infrastructure: (1) a cognitive foundation that understands traffic semantics and driving intent, and (2) a foresighted physical environment that can anticipate the consequences of candidate actions. To this end, we propose , we distill VLM knowledge into the BEV encoder and then discard the VLM entirely, retaining cognitive ability at zero inference cost while releasing the cognitive channel as a pluggable interface for optional human language commands. To , we build an auto-regressive BEV world model that explicitly predicts future semantic maps conditioned on candidate actions, serving as an interpretable physical sandbox from which safety metrics are directly derived. Built upon this dual infrastructure, we optimize the driving policy via GRPO with a novel dual-reward mechanism: a reward from a language-aligned scorer ensures intent compliance. Extensive experiments demonstrate that CoPhy not only achieves state-of-the-art results on NAVSIM v1 and v2 benchmarks, but also enables safer driving via cognitively informed scene compliance and flexible intent control through user-defined language instructions.
PaperID: 7156, Poster
Abstract: Existing multi-task active learning (MTL-AL) methods predominantly rely on the linear scalarization of heterogeneous objectives, such as uncertainty, diversity, and gradient conflict.We argue that this entangled formulation is fundamentally limited, as linear combination fails to capture Pareto-optimal trade-offs under conflicting objectives. To address this limitation, we propose a paradigm shift from combination to deconstruction and introduce a principled framework that rethinks MTL-AL along three orthogonal dimensions. First, we identify the Conflict Paradox: contrary to conventional wisdom, samples with high gradient conflict are not noise to be avoided but carry maximal information about unresolved inter-task trade-offs. We therefore decouple sample selection, which embraces conflict, from optimization, which resolves it via gradient surgery. Second, we introduce Orthogonal Decomposition, separating the acquisition objective into two distinct axes: Informativeness and Exploration. This two-stage process, consisting of subspace projection followed by manifold coverage, prevents ``diversity dilution'' by excluding uninformative outliers before enforcing diversity. Third, we propose Cross-Task Mutual Information (CTMI), a distribution-free proxy for inter-task dependency derived from a variational lower bound under the Maximum Entropy principle. Extensive experiments on COCO, Cityscapes, and NYUv2 — including comparisons against recent multi-task active learning baselines — demonstrate consistent improvements over state-of-the-art MTL-AL methods, highlighting that structural rethinking of acquisition yields larger gains than incremental heuristic design. Our code is publicly available at: \hrefhttps://anonymous.4open.science/r/multitask
Abstract: Training strong large language models (LLMs) requires high-quality supervision, which is often scarce. Recent work shows that paired preference data from weak–weaker model pairs (e.g., Qwen3 4B over 1.7B), despite the limited quality of individual responses, can provide an effective supervision signal through relative quality deltas, which we term a "weak" signal. This motivates a key research question: can multiple "weak" signals be constructively aggregated for improving strong models (e.g., Qwen3 8B)? To this end, we propose Preference Delta Aggregation (PDA), the first framework that derives a preference delta from each weak-weaker model pair, instantiates it as a LoRA adapter learned through preference optimization, and aggregates the resulting deltas via LoRA merging. To further mitigate directional interference during LoRA merging, we introduce Geometric Alignment Merging (GAM), a geometry-aware merging method that aligns adapter subspaces before aggregation, enabling more robust composition of diverse deltas. Evaluations on knowledge reasoning and agentic search benchmarks show that aggregating multiple "weak" signals pushes performance beyond any single signal, with further gains as additional signals are incorporated. Correspondingly, PDA with GAM improves the strong model by 6.8 and 7.3 points on average for knowledge reasoning and agentic search, respectively. It outperforms all single-delta and multi-delta baselines, exceeding the best single-delta baseline by 2.1 and 4.3 points. Further analysis attributes these gains to the effective composition of complementary capabilities encoded across distinct preference deltas.
Abstract: This paper studies the role of sinks and diagonal patterns as attention switch and anti-oversmoothing mechanisms. We analyze geometric conditions under which sinks can be represented, showing a necessary alignment between the embedding of the sink and all other embeddings. Next, we refine the current understanding of the role of sinks in oversmoothing prevention: we specify the conditions under which dense attention provably smooths more than sparse attention, and empirically verify that such conditions are often satisfied in practice. We further prove an equivalence between sinks and hard attention switch, in which the output of the attention is identically 0. Finally, we relax the hard attention switch by allowing token self-communication: we provide a quantitative comparison of the costs of representing sinks vs.\ diagonal patterns, showing why sinks are favored in pretrained transformers. The introduction and analysis of diagonal patterns and the generalization of the attention switch close the gap between what oversmoothing prevention requires and what sinks provide, while also establishing when and why attention layers act like MLPs if token communication is not necessary.
PaperID: 7159, Poster
Authors:
Yi Wen, Derong Xu, Pengyue Jia, Yichao Wang, Yingyi Zhang, Maolin Wang, Junyi Li, Wenlin Zhang, Xiaopeng Li, Yong Liu, Xiangyu ZhaoAbstract: The memory capabilities of Large Language Models (LLMs) have garnered increasing attention recently. Despite great success achieved, existing retrieval-based memory approaches typically overlook the differences between memories and employ a unified strategy to process all memories, leading to suboptimal performance. Thus, an intuitive question arises: can we categorize memory into different types and select appropriate strategies? However, given the topic-rich, scenario-complex, and boundary-blurred nature of memory scenarios, achieving precise classification of memories is not easy. To address this challenge, we propose a memory multi-class dataset in this paper, termed TriMEM, which provides precise annotations for memory types across diverse scenarios. Building upon this foundation, we propose a novel memory framework, named MemoType, which can adaptively recognize each memory and query type with the learned router model. With the memory and query routing, MemoType can retrieve the memory with corresponding query types and design tailored retrieval strategies, thereby enhancing the retrieval performance. Moreover, we theoretically prove that any single retrieval strategy is subject to a fundamental upper bound on its expected retrieval precision in multi-class corpora, leading to systematic precision degradation. Extensive experiments on three datasets demonstrate that MemoType consistently outperforms existing methods, achieving up to 16.18% improvement in Recall@1.
PaperID: 7160, Poster
Authors:
Yirong Xiong, Mingrui Cai, Gene Wen, Yuxing HanAbstract: Post-Training Quantization (PTQ) efficiently compresses Large Language Models (LLMs) without expensive retraining. However, most existing PTQ methods rely on linear quantization, which fails to fully capture the underlying characteristics of parameter distributions. Conversely, non-linear approaches, such as codebook-based methods, introduce significant complexity that hinders practical engineering deployment. To address this, we propose MixNLQ, an easy-to-use hybrid non-linear PTQ method. By combining linear and exponential terms, MixNLQ flexibly fits the "high-peak, heavy-tail" distributions typical of LLM weights and activations. It supports both calibration-free and calibration-based settings, and uses arithmetic-only dequantization to avoid memory-intensive lookup tables—ensuring highly efficient quantization and inference. Additionally, MixNLQ serves as a plug-and-play module capable of enhancing various existing PTQ frameworks. Evaluations at 3-bit and 4-bit on LLaMA, Mistral, and DeepSeek demonstrate that MixNLQ achieves competitive or superior perplexity and zero-shot accuracy compared to existing PTQ baselines, alongside promising activation quantization performance and its versatility as a plug-and-play enhancement module. Ultimately, MixNLQ offers a highly practical solution that effectively balances accuracy, algorithmic simplicity, and real-world deployment efficiency.
PaperID: 7161, Poster
Abstract: Multi-agent debate (MAD) was proposed as a promising approach for ensembling the wisdom of multiple large language models (LLMs) to improve reasoning and provide effective supervision to superhuman LLMs. However, increasing empirical evidence suggests that MAD may not outperform or even significantly underperform single-agent approaches (SA), raising doubts about the benefits of MAD. In this work, we investigate this issue by analyzing the incentive structures of popular MAD paradigms: (i) competitive MAD (CopMAD) where agents compete by holding opposing positions; (ii) consensus-seeking MAD (CosMAD) where agents are driven to seek consensus. We show that both paradigms suffer from debate hacking: CopMAD reduces to a cheap-talk game, where agents produce misleading messages to win the game, while CosMAD filters out informative disagreements for premature consensus. Consequently, agents in both CopMAD and CosMAD fail to jointly resolve the ambiguity and seek the truth. To this end, we introduce ColMAD, a collaborative protocol that reframes MAD as a non-zero-sum game to encourage agents to provide informative while truthful messages. Through extensive benchmarking on challenging tasks such as error detection, we show that ColMAD significantly outperforms previous MAD protocols up to 10 percentage points. Under the same budgets, ColMAD effectively brings non-trivial improvements over SA methods, implying that the protocol design is critical to realizing the potential of MAD.
PaperID: 7162, Poster
Abstract: Large Vision Language Models (LVLMs) have demonstrated outstanding capabilities in multimodal content understanding, yet their multimodal nature also leads to security vulnerabilities. Existing defense methods often suffer from high computational cost and inference latency. In contrast, steering-based methods avoid these issues but still depend on heuristic model representations and struggle to handle coordinated multimodal attacks, where malicious content from images or text prompts can hijack the model’s limited attention resources. We identify this phenomenon as attention distraction, where an LVLM’s attention shifts from safety reasoning toward maliciously aligned tokens. To this end, we propose DualSteer, a lightweight and training-free defense framework that restores safe reasoning via dual-space steering. DualSteer first adjusts the attention distribution to counter the influence of jailbreak distractors, and then steers hidden layer representations to address residual jailbreak effects that the model cannot resist on its own. Extensive experiments across six jailbreak benchmarks and three LVLMs of different architectures demonstrate that DualSteer consistently reduces attack success rates by over 30% relative to state-of-the-art defenses, while ensuring the model’s utility and avoiding false rejections. DualSteer provides a novel defensive perspective while offering an efficient and robust solution to defend against complex multimodal jailbreak attacks.
PaperID: 7163, Poster
Authors: Fazong Wu, Ming Yang, Xin Wang, Zhenyong Zhang, Xiaoming Wu
Abstract: Large language models (LLMs) are now deployed in a wide range of real-world applications, making it important to evaluate how reliably they resist jailbreak attacks. However, existing jailbreak evaluations mainly rely on manually collected prompts or discrete text-space optimization, which limits their coverage and makes them difficult to extend to new threat settings. We present JailBound, a jailbreak evaluation framework that combines automated benchmark construction with intent-preserving attack optimization in embedding space. JailBound organizes evaluation instances with a hierarchical threat taxonomy spanning risk categories, application domains, and attack types, and uses this structure to generate meta-attack prompts with a fine-tuned meta-attack generator. It further formulates jailbreak evaluation as an embedding-space attack optimization problem and uses a first-order loss (FOL)-guided dual-branch search to jointly identify high-value vulnerable regions and safety boundary states. Under a unified evaluation protocol, we study 46 LLMs from 13 model families. The results show that JailBound covers a broader range of risk settings than existing prompt-based benchmarks, and that its optimized attacks transfer nontrivially across model families while supporting finer-grained analysis of vulnerability patterns and safety boundary behavior. \textcolorred\textWarning: this paper includes examples that may be offensive or harmful.
PaperID: 7164, Poster
Authors:
Longyi Liu, Zhitao Wang, Jianchao Yu, Mingrui Cai, Yuxing Han, Gene WenAbstract: Inference-time activation steering provides a lightweight mechanism for aligning language model behavior without retraining, yet existing additive interventions often distort activation geometry, alter representation norms, and induce effective-rank collapse, thereby degrading open-ended generation. Recent norm-preserving rotation-based methods mitigate this pathology, but their reliance on one- or two-dimensional concept axes is insufficient for behaviors whose representations are intrinsically multi-faceted and low-rank. We propose Grassmannian Geodesic Steering (GGS), a rank-preserving framework that lifts activation control from vector directions to k-dimensional subspaces on the Grassmann manifold. We develop a theoretical account showing that additive steering provably collapses effective rank under strong intervention, that low-dimensional rotations cannot capture distributed behavioral concepts, and that orthogonal transformations along Grassmannian geodesics exactly preserve the singular-value spectrum. Guided by these results, GGS estimates concept subspaces from contrastive difference covariance, steers hidden states via a closed-form Grassmannian geodesic realized as an orthogonal operator, and adaptively modulates intervention strength using a matrix Bingham gate over subspace-valued activations. The resulting method introduces less than 1.5% wall-clock overhead per decoding step. Across LLaMA-3.1-8B and Qwen-2.5-7B on six reasoning benchmarks and an AdvBench safety evaluation, GGS consistently improves alignment accuracy while preserving generation quality and refusal behavior, reducing effective-rank degradation by two orders of magnitude relative to additive baselines, and surpasses the current state of the art among training-free inference-time steering methods.
PaperID: 7165, Poster
Abstract: Red teaming is essential for securing Large Language Models, yet current automated methods remain limited by templates and single-turn attacks. To simulate the complex, interactive nature of real-world adversarial attacks, we introduce a novel red teaming paradigm designed to maximize expected cumulative harm through strategic interaction. By formalizing red teaming as a Markov Decision Process in a hierarchical reinforcement learning framework, we navigate the challenges of sparse rewards and long-horizon planning. Our generative agent learns diverse, multi-turn attacks using a token-level harm reward, consistently uncovering vulnerabilities that bypass baselines. This approach achieves a new state of the art and reframes LLM red teaming as a principled, trajectory-based process.
PaperID: 7166, Poster
Abstract: The rapid proliferation of personalization techniques in Stable Diffusion has driven the widespread deployment of adversarial perturbations as a protective measure against unauthorized data mimicry. Existing diffusion-based purification methods primarily rely on the noising and denoising paradigm and have proven somewhat effective in neutralizing these protections. However, they suffer from a severe trade-off between purification efficacy and perceptual fidelity, especially under large perturbation budgets. To mitigate this problem, in this paper, we propose Dynamic Sparse Semantic Anchoring (DSSA), a novel purification framework. By starting the reverse denoising process from pure Gaussian noise, DSSA removes the initial adversarial perturbations and circumvents the dilemma of selecting an optimal noising level. To prevent the loss of original semantics caused by pure noise generation, we propose a dynamic sparse anchoring strategy that directly integrates sparse pixels from the adversarial sample to guide the reconstruction. Additionally, we design a perception-guided early-stopping mechanism to prevent the perturbation resurgence caused by this direct pixel integration. Extensive experiments across diverse datasets and protection schemes demonstrate that our method consistently outperforms existing baselines, achieving state-of-the-art performance in efficacy and fidelity, successfully neutralizing even advanced adaptive protections.
PaperID: 7167, Poster
Authors:
Xueying Du, Kai Yu, Chong Wang, Yi Zou, Wentai Deng, Zuoyu Ou, Xin Peng, Yiling LouAbstract: Static bug analyzers play a crucial role in ensuring software quality. However, existing analyzers for bug detection in large-scale codebases often suffer from high false positive rates. This limitation largely stems from their inadequate capabilities in performing precise path feasibility validation for complex code contexts. While recent work has explored Large Language Models (LLMs) to eliminate false positives in static bug detection, the limited reasoning capabilities of LLMs over long and complex contexts hinder their effectiveness when applied to large-scale software projects. To address this challenge, we propose PathAgent, an agent-driven framework for fine-grained path feasibility analysis. PathAgent decomposes complex inter-procedural analysis into localized symbolic range reasoning sub-tasks, and employs an on-demand contextual exploration strategy to adaptively retrieve only the necessary code context. These designs enable precise path feasibility reasoning and effectively reduce false positives reported by static bug analyzers. Evaluation on large-scale real-world projects shows that PathAgent eliminates 79% of false positives, outperforming all baselines by 23% to 44%, while maintaining strong bug detection capability with a recall of 0.94.
Abstract: Large-scale datasets and fast simulators have enabled improvements in driving policies that appear safe and robust, yet strong performance in nominal scenarios can still mask flawed reasoning and unsafe heuristics. Moreover, existing closed-loop simulations often reveal only the resulting behavior, making it difficult to determine whether driving policies truly predict the motion of surrounding vehicles or how the ego vehicle generates future plans, or merely rely on brittle heuristics that happen to succeed in nominal scenarios. To better understand the limits and weaknesses of driving policies, we focus on probing for forms of prediction, i.e., where surrounding vehicles will move next, and planning, i.e., understanding how to generate safe trajectories. We focus on these two capabilities because they reflect behaviors expected of effective driving policies, and use their presence or absence to assess policy quality across data-driven behavior cloning and simulation-driven reinforcement learning policies. To evaluate the presence of these capabilities, we investigate them as a function of scale, asking whether the closed-loop gains from larger datasets and longer simulation training reflect stronger prediction and planning or merely better behavioral heuristics. We use linear probing and targeted perturbations in both imitation learning and reinforcement learning models to track when these internal signals emerge, plateau, or fail. Despite good closed-loop performance, prediction signals are often mistimed or even degraded during simulated near-collision events. Finally, we demonstrate that these internal predictive capabilities go beyond correlation and that correcting mistaken predictions causally steers the planner toward safer, more appropriate trajectories.
PaperID: 7169, Poster
Abstract: Recent diffusion-based video generation models have enabled high-quality video customization through both tuning-based pipelines, which fine-tune a video diffusion model, and reference-based pipelines such as image-to-video generation. However, these capabilities raise serious concerns about facial privacy and security. Existing anti-diffusion protections focus on the image domain or on reference-based I2V pipelines, leaving the tuning-based video customization unexplored. Protecting videos in this setting raises three challenges: image-level perturbations are erased by temporal operations, a perturbation tied to one clip fails to generalize across videos and lengths, and inconsistent temporal noise is easily removed by temporal attack. To address these challenges, we propose Temporally Consistent Universal Adversarial Perturbations (TC-UAP), the first protection method against both reference- and tuning-based video customization. TC-UAP learns a multi-frame universal adversarial perturbation over a set of videos of the same identity, so that a single perturbation can transfer to unseen videos and arbitrary clip lengths of that identity. Besides, we further enforce consistency through intrinsic temporal modeling and an extrinsic surrogate temporal-attack loss, ensuring robustness against temporal attacks. Extensive quantitative and qualitative experiments show that TC-UAP degrades identity preservation more than existing baselines under both tuning-based and reference-based video customization, and remains robust under three unseen temporal attacks.
Abstract: Static neural point reconstructions capture a subject at high fidelity from posed images. Given such a reconstruction and a monocular fixed-viewpoint driving video of the subject, whether captured or produced by image-to-video (I2V) generation, we recover a rigged, re-posable 3D asset. Existing methods deform Gaussian splats through direct linear blend skinning (LBS) or mesh proxies, both of which are prone to joint-boundary artifacts under articulation, even with per-primitive corrections. We trace the artifact to the representation: each splat carries an individual shape calibrated in the canonical pose to tile with its neighbours. Under rigid LBS, each splat moves with its bone but cannot bend, so the canonical tiling breaks at joint boundaries into gaps and spikes. Proximity attention point rendering (PAPR) instead carries no per-primitive shape; each pixel is recomposed at render time from the deformed primitives' positions, so the surface re-forms naturally with the articulation. We present RigPAPR, which auto-rigs a static PAPR cloud and drives it under direct LBS from a single fixed-viewpoint video, without mesh proxy, pose-dependent correction, or category template. On synthetic subjects, RigPAPR matches the strongest baseline at the supervised view and exceeds mesh-based and Gaussian-splatting baselines at novel views by 3+ dB PSNR, with cleaner joint-boundary renderings of both synthetic and real subjects.
Abstract: Multimodal large language models must continually adapt to evolving tasks and domains, yet standard continual learning metrics mainly measure whether old answers remain correct, leaving the stability of multimodal grounding largely unexamined. We study this overlooked failure mode and ask whether a continually adapted MLLM can preserve not only what it answers, but also how it uses visual, textual, OCR, chart, and document evidence. We identify hidden evidence-use forgetting, where answer accuracy is retained while the model silently shifts toward different or less grounded evidence channels, and propose RCL, a replay-free reliance-constrained continual learning framework. RCL freezes the previous checkpoint as a behavioral reference, estimates teacher and student evidence-reliance profiles through counterfactual channel interventions, and jointly optimizes task learning, prediction preservation, and reliance preservation without adding inference-time cost. Across CoIN, COAST, MCITlib, and an evidence-sensitive multimodal stream, RCL consistently improves final performance and reduces forgetting over replay-free, PEFT, routing, and memory-assisted baselines, while substantially lowering modality reliance drift, dominant evidence flips, and hidden forgetting rates. These results suggest that robust continual multimodal learning requires preserving the evidence path behind correct answers, not merely the answers themselves.
Abstract: Large language models (LLMs) are increasingly accessed as remotely hosted services by edge and enterprise clients that cannot run frontier models locally. Since models vary widely in capability and price, routing queries to models that balance quality and inference cost is essential. Existing router approaches assume access to centralized query-model evaluation data. However, these data are often fragmented across clients, such as end users and organizations, and are privacy-sensitive, which makes centralizing data infeasible. Additionally, per-client router training is ineffective since local evaluation data is limited and covers only a restricted query distribution and a biased subset of model evaluations. We introduce the first federated framework for LLM routing, enabling clients to learn a shared routing policy from local offline query-model evaluation data. Our framework supports both parametric multilayer perceptron router and nonparametric K-means router under heterogeneous client query distributions and non-uniform model coverage. Across two benchmarks, federated collaboration improves the accuracy-cost frontier over client-local routers, both via increased effective model coverage and better query generalization. Our theoretical results also validate that federated training reduces routing suboptimality.
PaperID: 7173, Poster
Abstract: Multimodal large language models (MLLMs) represent images as long sequences of visual tokens, making inference costly and often redundant. The central challenge is to prune these tokens aggressively without removing evidence required for reliable reasoning. Existing pruning methods typically rank tokens by estimated importance or redundancy, implicitly assuming that low-scored tokens are safe to discard. This assumption is fragile when small, occluded, or fine-grained visual cues are uncertain yet decisive, and it becomes especially brittle under semantics-preserving prompt perturbations. We present RobustPruner, a decoder-integrated framework for prompt-robust and uncertainty-aware visual token pruning. RobustPruner predicts query-conditioned relevance, uses a determinantal point process (DPP) to construct a diverse candidate subset, and then applies uncertainty-guided refinement within that subset. By decoupling relevance and uncertainty, the method preserves ambiguous but potentially decisive evidence while remaining computationally efficient. Inserted into intermediate decoder layers, RobustPruner performs one-shot pruning and reduces KV-cache and memory costs. Across three open-source MLLMs and diverse vision-language benchmarks, RobustPruner consistently achieves a stronger accuracy-efficiency-robustness trade-off than prior pruning baselines, with especially clear gains on fine-grained and text-rich tasks. On Qwen-2.5-VL-7B, RobustPruner retains 95.22% of the original performance while reducing GPU memory, prefill FLOPs, and KV cache size to 91.43%, 50.97%, and 20.57% of the baseline.
Authors: Davide Romano, Kanak Raj, Jerrod Parker, Daniele Giofrè
Abstract: Test-time scaling (TTS) improves language model outputs by spending additional inference compute — generating multiple candidates, searching over partial sequences, or iteratively refining drafts. These techniques yield large gains on mathematics and code, but have been developed and stress-tested almost exclusively on tasks where verification is straightforward. We conduct the first compute-normalised comparison of five TTS families across five open-ended generation benchmarks spanning medicine, law, finance, general chat, and creative writing — grounded in a unified framework that decomposes the effectiveness of each method's token budget into exploration and exploitation. The answer depends on which side of that decomposition you examine. Scaling exploration works: the best candidate in the pool improves steadily with compute across all settings. What breaks is exploitation — the step that converts a rich candidate pool into a final output. With state-of-the-art generators, reward models correlate at only \hat\rho_v \approx 0.12 with true quality, rendering selection near-random regardless of budget. Tree search amplifies this failure through diversity collapse. Refinement helps on one of five benchmarks; its apparent gains elsewhere are confounded. Only synthesis across candidates (Fusion) consistently improves over single-sample baselines, yet still recovers only ~40% of available quality.The candidate pool is not the bottleneck — choosing from it is.
Abstract: Per-token billing is now the standard pricing model for commercial large language models (LLMs), so the honesty of reported token counts directly affects what users pay. We show that this kind of billing is hard to audit by design: providers hide the model, the tokenizer, and the execution to protect their IP, mitigate jailbreaks, and preserve user privacy, which means an auditor can only inspect proofs the provider supplies. The audit therefore reduces to a consistency check on the provider's own reports. We call this a trust paradox: every audit must trust some artifact, but current frameworks trust exactly the ones a provider has the strongest reason to manipulate. We study three recent token auditing frameworks and show that a provider with ordinary commercial capabilities can systematically inflate billed token counts. In the most permissive setting, hidden reasoning usage can be inflated by 1,469% on average without detection. At current frontier reasoning prices, that turns a 100 honest bill into roughly a 1,569 bill on the same query. Even when the user can see the full reasoning string, tokenization ambiguity alone still allows 50.85% over-reporting below the detection threshold. These results suggest the problem is not in any specific auditor but in any audit whose evidence comes from the audited party. Restoring honest billing will require verification that ties reported token counts to evidence the provider does not control, such as trusted execution attestation, cryptographic proofs of inference, or third-party re-execution.
Abstract: Obtaining stable diffusion-based samplers in high- and infinite-dimensional settings is challenging because errors can accumulate across high-frequency coordinates and make the dynamics unstable under refinement of the finite-dimensional approximation of the underlying function-space problem. Discretization is a typical source of such errors, and preconditioning with a suitable spectral decay is one way to control their accumulation. In this paper, we study this problem for preconditioned annealed Langevin dynamics (ALD) applied to Gaussian mixtures. We first show that Euler-Maruyama (EM) discretization, by treating the stiff linear part of the annealed score with a forward Euler step, imposes a stability constraint coupling the preconditioner with the annealed covariance scale. Together with the conditions ensuring dimension-uniform control of the annealed dynamics, this constraint forces the initial smoothed law to remain uniformly close to the target across dimensions. We then consider an exponential-integrator scheme that integrates the stiff linear part of the annealed score exactly. Under explicit spectral summability conditions coupling the smoothing covariance, the component covariance spectra, and the preconditioner, we prove a dimension-uniform Kullback-Leibler (KL) bound for this scheme. This bound can be made arbitrarily small, uniformly in dimension, by allowing enough time for annealing and then refining the time mesh accordingly. Importantly, these conditions allow regimes in which the KL divergence between the target and the initial smoothed law diverges with dimension, showing that the restrictions imposed by EM are scheme-dependent rather than intrinsic to ALD.
PaperID: 7177, Poster
Abstract: Structure-based drug design with diffusion models seeks to generate 3D molecules that bind tightly to target protein pockets, but full-step sampling is too slow for high-throughput use. Extending the recent development of training-free fast sampling methods for diffusion models (e.g., DDIM and DNDM) to high-quality 3D molecule generation is non-trivial, given that both discrete variables (atom types) and continuous variables (atom coordinates) are involved. In this paper, we propose a unified fast sampling framework that allows a shared time schedule across discrete and continuous variables for strategically accelerating the 3D ligand generation. Guided by an error-equalization principle, we combine the continuous and discrete error rates of hybrid sampling into a single step-importance density and derive a training-free coupling-aware schedule. We further propose a confidence-guided sampling strategy that integrates constraints of model prediction confidence, motion stability, and steric clash to control the acceleration based on geometric readiness. Extensive experiments demonstrate that our method achieves 18-20× speedup compared to full-step diffusion models while preserving state-of-the-art generation quality across chemical properties, binding affinity, and chemical validity.
PaperID: 7178, Poster
Authors:
Jeonghwan Cho, Minsu Kim, Junyoung Hong, Seon Joo KimAbstract: Existing 3D Gaussian Splatting (3DGS) segmentation methods produce accurate rendered masks yet often fail to select a complete set of object Gaussians: removing the predicted Gaussians leaves the rendered depth in the target region nearly unchanged. We trace this gap to the contribution-weighted supervision path of rendering-based feature learning. Because Gaussian geometry and opacity are pre-optimized for RGB reconstruction, rendered-feature losses deliver strong gradients only to high-contribution Gaussians, leaving other Gaussians in the same object region weakly supervised and ultimately unselected. We propose VODA, a plug-and-play supervision objective that addresses this bias. After RGB optimization, Gaussians of the same object portion concentrate at similar depths along the camera ray. A depth-aware rasterizer exploits this property, grouping per-pixel contributing Gaussians into depth-coherent clusters and selecting the cluster that best explains each object pixel. The rendered feature of the selected cluster is then applied as a target directly to every Gaussian within it, bypassing the contribution-weighted gradient path. A complementary pixel-level loss further sharpens boundaries between adjacent objects. To evaluate volumetric completeness, we introduce Removal-3D and the Removal Depth Error (RDE), which measure whether removing the predicted Gaussians induces the expected depth change in 3D. Across four rendering-based baselines, VODA consistently improves volumetric segmentation on Removal-3D while maintaining or improving standard performance on LERF and ScanNet.
Abstract: We study the problem of multiclass PAC learning with bandit feedback in the realizable setting. In this framework, there is an unknown data distribution over an instance space \mathcalX and a label space \mathcalY, as in classical multiclass PAC learning, but the learner does not observe the labels of the i.i.d. training examples. Instead, in each round, it receives an unlabeled instance, predicts its label, and receives bandit feedback indicating only whether the prediction is correct. Despite this restriction, the goal remains the same as in classical PAC learning. We provide a general characterization of the optimal sample complexity of this problem, sharp for every concept class up to logarithmic factors. Our characterization is based on a new combinatorial dimension, termed the bandit \mathrmDS dimension, defined via generalized combinatorial structures we call pseudo-boxes. These extend the pseudo-cubes underlying the \mathrmDS dimension by allowing a different number of neighbors in each coordinate. In contrast to the \mathrmDS dimension, which governs the full-information setting by counting the number of coordinates in the pseudo-cube, the bandit \mathrmDS dimension aggregates the number of neighbors across coordinates, leading to a characterization in which the sample complexity scales with the total number of neighbors. We also propose a general learning algorithm achieving the upper bound, based on an algorithmic principle called ListCascade, which connects bandit learning to list learning and may be of independent interest.
Abstract: Latent Chain-of-Thought aims to enable step-by-step computation without emitting long rationales, yet its internal mechanisms remain unclear. We study CODI, a continuous-thought teacher–student distillation model, on strictly sequential polynomial-iteration tasks with known intermediate states. Using logit-lens decoding, linear probes, attention analysis, and activation patching, we localize intermediate-state representations and trace how they are routed to the final readout. In short-horizon, low-hop tasks, CODI forms faithful bridge states across latent-thought positions, while the final input follows a separate near-direct route; predictions arise through late fusion at the answer readout. As task depth and difficulty increase, however, CODI does not reliably sustain a full latent rollout: it either compresses computation into a partial late-intermediate pathway or, in harder regimes, loses the latent reasoning signature altogether. To explain this transition, we show theoretically that the task’s algebraic structure controls its effective memory: compressible regimes support late-bottleneck reasoning, while incompressible regimes preserve full-history dependence and destabilize latent rollouts. Overall, our results characterize when CODI-style latent-CoT yields faithful iterative computation versus compressed or shortcut strategies.
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR), particularly GRPO, has become the standard for eliciting LLM reasoning. However, its efficiency in exploration and difficulty adaptation remains an open challenge. In this work, we identify an implicit advantage symmetry inherent in Group Relative Advantage Estimation (GRAE) as a structural property that provides a new perspective for understanding these bottlenecks. This symmetry induces two critical limitations: (i) at the group level, strict symmetry in weights between correct and incorrect trajectories leaves unsampled action logits unchanged, thereby hindering the exploration of novel correct solution. (ii) at the sample level, the algorithm implicitly prioritizes medium-difficulty samples, remaining agnostic to the non-stationary demands of difficulty focus. Through controlled experiments, we reveal that this symmetric property is sub-optimal, yielding two pivotal insights: (i) asymmetrically down-weighting the advantages of correct trajectories encourages essential exploration but risks instability; (ii) learning efficiency can be boosted by a curriculum-like transition—prioritizing simpler samples initially before gradually shifting to complex ones. Motivated by these findings, we propose Asymmetric GRAE (A-GRAE), which dynamically modulates exploration incentives and sample-difficulty focus. Experiments across seven benchmarks demonstrate that A-GRAE consistently improves GRPO and its variants across both LLMs and MLLMs.
PaperID: 7182, Poster
Abstract: Modern machine learning depends heavily on massive datasets, but obtaining high-quality annotations at scale is often expensive. As a result, learning from noisily-labeled data has become common, making accurate estimation of the label-noise transition matrix crucial. However, existing transition matrix estimators rely on the fragile estimation of class-posteriors and do not provide finite-sample performance guarantees. In this work, we propose a novel methodology to estimate the transition matrix based on one-sided selective classification. This approach bypasses class-posterior estimation, provides finite-sample performance guarantees, and leverages flexible learning methods for binary classification. Moreover, we introduce effective algorithms to implement the proposed methodology and provide their refined finite-sample performance bounds.
PaperID: 7183, Poster
Authors: Quoc D Ngo, Hung Nguyen
Abstract: Sparse autoencoders decompose activations into sparse, interpretable features and have become a central tool for mechanistic interpretability. Transcoders extend this approach from representation reconstruction to computation prediction by mapping MLP inputs to MLP outputs, making them attractive for circuit-level analysis. We ask whether a prominent sparse-autoencoder reliability fix, Matryoshka nested-prefix training, transfers to this two-hook setting. In sparse autoencoders, Matryoshka training improves absorption behavior by encouraging broad features to appear in early dictionary prefixes. In transcoders, we find that this protection does not transfer automatically. Across Gemma-2-2B MLP transcoders trained under a matched recipe, Matryoshka models form genuine prefix hierarchies, with reconstruction improving as later groups are added, but they do not inherit the absorption advantage observed for sparse autoencoders. BatchTopK achieves lower full absorption, lower FVU, and higher first-letter F1@1 in the canonical 16k comparison. Bridge controls recover the expected Matryoshka advantage when the cross-space map is removed, showing that the failure is specific to predicting MLP outputs from MLP inputs. We identify the mechanism as hierarchy misalignment. Prefix losses reward covariance with the reconstruction target, so early transcoder groups prioritize target-predictive output geometry rather than broad source-space semantic parents. These results show that hierarchical sparse-dictionary objectives do not preserve interpretability properties by default. For dictionaries that map between activation spaces, the key question is not whether a hierarchy forms, but what geometry that hierarchy is trained to organize.
PaperID: 7184, Poster
Abstract: We study the sequential decision-making problem where at each time the learner selects an assortment of bounded size and receives a ranking of the selected items. Ranking feedback arises naturally when human evaluators or automated LLM judges compare multiple options simultaneously. It is often easier and more reliable than assigning absolute scores, which are typically poorly calibrated across evaluators, while yielding substantially more information than a single pairwise comparison. We focus on rankings that are constructed according to the random utility model with linear utilities, which includes the popular Plackett-Luce model as a special case. We show that improved regret minimisation is possible under full rank breaking with a refined analysis based on 1-factorisations and Baranyai's theorem that fully exploits independence in the pairwise comparisons. This gives a tighter confidence interval around the true utility parameter which is crucial in our analysis for regret that improves with maximum assortment size. Furthermore, it matches the regret lower bound for our model in case the learner is permitted to play multi-sets of the maximum bounded size. Our work gives a principled approach to learning from ranking feedback that shows consistent improvement with the maximum assortment size.
Abstract: We study how temporal correlations in the data can make certain sparse learning problems efficiently learnable by gradient-based methods. Our focus is on Boolean k-juntas, a canonical sparse learning problem known to pose barriers for gradient-based methods under independent uniform samples. We show that this picture changes when the samples are generated by a lazy random walk on the hypercube. In this setting, the temporal dependencies can be exploited by a two-layer ReLU network trained using stylized-SGD with a temporal-difference loss, which compares target and predicted increments across consecutive samples. For every fixed k, the resulting sample complexity is essentially linear in the ambient dimension d. By contrast, we show that for large-batch gradient methods using standard convex pointwise losses, temporal correlations do not provide the same advantage.
PaperID: 7186, Poster
Abstract: The standard LLM training pipeline applies reinforcement learning (RL) only after pre-training and supervised fine-tuning (SFT). We question this status quo by training a LLM from scratch and applying RL, SFT, and SFT followed by RL directly to intermediate pre-training checkpoints. We find that RL is effective very early, and often matches the full SFT\toRL pipeline early as well. Through experiments on harder problems, we find that targeted pre-training data composition is a strong lever for RL effectiveness, even more so than model scale. Beyond reasoning accuracy, applying RL directly to base checkpoints expands the model's distribution; the sharpening effect reported in recent work arises only when RL follows SFT. The general capabilities of the model remain essentially unchanged by RL, while they degrade following SFT. Finally, we merge RL and SFT objectives by parallel averaging, which outperforms across all other training methods discussed, across metrics, while preserving general capabilities. Together, these results suggest that LLM training might benefit from an expanded use of RL.
Authors: Marte Eggen, Eirik Reiestad, Kristian Gjøsteen, Inga Strümke
Abstract: Recent cryptographic results establish that neural networks can be backdoored such that no efficient algorithm can distinguish them from a clean model. These guarantees, however, have been confined to stylised architectures of limited practical relevance, leaving open whether comparable undetectability extends to modern, end-to-end trained networks. We construct such an attack mechanism for state-of-the-art architectures, closely aligned to the cryptographic notion of undetectability, by identifying backdoor channels as learned latent directions, and show that the question of undetectability reduces to a hypothesis test between two unknown distributions over model parameters, which we conjecture to be intractable in practice. The consequence of this reframing is significant: if exploitable channels within a network's latent space are statistically indistinguishable from naturally learned directions, an attacker need not introduce foreign structure but can instead exploit the geometry the network already possesses. Demonstrating the approach on ResNet and Vision Transformer architectures trained on standard image classification datasets, the attack achieves both consistently high success rates with negligible clean accuracy degradation, and resists a comprehensive suite of post-training defences, none of which neutralise the backdoor without rendering the model unusable. Our results establish that cryptographic backdoors need not be artefacts requiring exotic architectures or artificial constructions, but identifiable as latent properties inherent to the geometry of learned representations.
PaperID: 7188, Poster
Abstract: While long chain-of-thought (CoT) distillation significantly improves the reasoning abilities of language models, its underlying mechanism remains poorly understood. Current pipelines generally rely on rejection sampling to gather reasoning traces across diverse prompts, assuming that factual correctness and broad problem coverage are indispensable. Surprisingly, we find that training exclusively on flawed reasoning trajectories or a restricted prompt set yields nearly comparable performance. Driven by this counterintuitive finding, we propose the Behavioral Hypothesis: the essence of Long-CoT distillation is behavioral alignment with structured reasoning patterns, rather than factual knowledge memorization. We then validate this by demonstrating substantial reasoning gains even when models are fine-tuned exclusively on previously mastered problems—effectively isolating behavioral alignment from novel knowledge injection. Having established that reasoning patterns are the true bottleneck, we investigate how to learn them across varying levels of complexity optimally. We reveal a critical interplay: while the recent probability-based loss excels at learning simple patterns, the standard cross-entropy loss is essential for scaling to complex CoTs. We believe these insights demystify CoT distillation and provide a principled foundation for the future development of reasoning models.
Abstract: Large language models are remarkably capable, yet how computation propagates through their layers remains poorly understood. A growing line of work treats depth as discrete time and the residual stream as a dynamical system, where each layer's nonlinear update has a local linear description. However, previous analyses have relied on scalar summaries or approximate linearizations, leaving the full spectral geometry of trained LLMs unknown. We perform full Jacobian eigendecomposition across three production--scale LLMs and show that training installs a monotonic spectral gradient through depth---from non-normal, rotation-dominated early layers to near--symmetric late layers---together with a cumulative low-rank bottleneck that funnels perturbations into a small fraction of the residual stream's effective dimensions. Our experiments reveal that this gradient and the dimensional collapse are learned rather than architectural, and is largely dissolved when structured non-normality is removed. We further show that the topological positioning of graph communities predicts whether the Jacobian amplifies or suppresses them, with the sign of the coupling determined by the local operator type, a relationship absent at initialization. These results map a learned spectral geometry in LLMs that links perturbation propagation and compression to the network's functional topology.
PaperID: 7190, Poster
Abstract: Neuroscientists have recently turned to intracranial brain recording methods, like electrocorticography (ECoG), for human experiments because of the fine spatial and temporal resolution that they afford. Models trained on this data, however, are fundamentally restricted by the patient populations that can receive the implants necessary for recording. We propose using non-invasive fMRI to bridge the gap in training data. Using spoken language representations fine-tuned on fMRI, we build encoding models of ECoG. These representations showed improved prediction performance in ECoG, even though the temporal resolution of fMRI is two orders of magnitude worse. Prediction improved in frequency bands well beyond what is directly measured in fMRI. Next, to test the procedure's generalization ability, we fine-tuned models on fMRI responses that were temporally downsampled by a factor of 2. Despite the loss in resolution, these models were able to predict fMRI and ECoG responses at levels comparable to the original fMRI-tuned models. Finally, we showed that ECoG performance steadily scales with the amount of fMRI-tuning data. Our results show that "slow" data like fMRI can be a valuable resource for building better models of "fast" brain data like ECoG. In the future, integrating across multiple recording methods may further improve performance in other applications, like decoding.
PaperID: 7191, Poster
Abstract: In image classification scenarios where both prediction and explanation efficiency are required, self-explaining models that perform both tasks in a single inference are effective. However, for users who already have prediction-only models, training a new self-explaining model from scratch imposes significant costs in terms of both labeling and computation. This study proposes a method to transfer the visual explanation capability of self-explaining Vision Transformer (ViT) models learned in a source domain to prediction-only VLM-based ViT models in a target domain via task arithmetic, without target-side explanation supervision. The proposed method endows explanation capability by adding an \emphexplainability vector induced by explanation supervision in the source domain based on task arithmetic framework. Experiments on ten diverse target datasets show that a single explainability vector learned on ImageNet-1k augmented with patch-level explanation supervision transfers consistently, improving explanation quality while largely preserving classification accuracy. Beyond the transfer itself, we further investigate whether transfer success can be anticipated \empha priori from model-internal features alone (without target-side explanation supervision), and find that a small set of such features carries non-trivial signal about the transfer outcome.
PaperID: 7192, Poster
Abstract: Multi-component embodied robots must coordinate the base, arm, camera, gripper, and local environment under shared spatial and temporal constraints. We present Q-CoMove, a coordinated motion framework that models this platform as a coupled dynamical system. Q-CoMove uses a differentiable circuit-style prior implemented with the TorchQuantum simulator on classical hardware: component states are encoded into a five-qubit register, cross-component dependence is parameterized by a 79-gate simulated circuit, and 25 measured statistics are projected into a coordination latent that conditions a physics-residual dynamics model and a joint action generator. We position this module as a compact structured coordination bottleneck rather than a claim of quantum advantage or physical necessity. Across parameter-matched Transformer and GNN baselines, residual-dynamics and joint MPC baselines, ablations, and robustness stress tests in simulation, Q-CoMove improves coordination metrics with bounded inference cost. All empirical results are limited to simulated environments, and the current study does not isolate the circuit prior against a structure-matched classical bottleneck.
PaperID: 7193, Poster
Authors: Luca Savant Aira, Diego Valsesia, Enrico Magli
Abstract: Neural Radiance Fields (NeRFs) are popular models using MLPs for continuous volumetric representations, which render novel views by means of numerical quadrature of the volumetric integral. Despite the widespread usage of NeRFs, we report subtle mistakes in their quadrature rule which have gone unnoticed for many years. In particular, we highlight mistakes present in both the original NeRF paper and its implementation as well as in the mainstream NeRFStudio implementation. We provide a theoretical analysis of rendering error bounds for NeRF and we use it to show suboptimal convergence rates in the presence of the aforementioned mistakes. We introduce a minimal fix that recovers the expected convergence order with no complexity overhead. Finally, while fixing the quadrature scheme is mostly a matter of theoretical correctness, we also provide experimental evidence that modest improvements in image quality are achieved on standard benchmarks.
PaperID: 7194, Poster
Abstract: Computational models of cognition have traditionally relied on abstract symbolic inputs that approximate sensory experiences with simplified vector representations. This simplification reflects historical computational constraints and a long-held assumption that the details of sensory realism are incidental to the study of cognitive processes. Recent advances have enabled the development of sensory-realistic models capable of performing diverse cognitive tasks, yet whether sensory realism meaningfully shapes cognition has not been formally established. We directly address this gap through a large-scale empirical comparison of abstract and sensory-realistic neural network models trained on a battery of memory-dependent decision-making tasks. Our results show that sensory-realistic models mirror human behavioral response patterns significantly more accurately than their abstract counterparts, though a substantial gap to human internal behavioral consistency remains even for the best tested models. By analyzing the internal dynamics of these systems, we demonstrate that the recurrent dynamics in abstract models is not inherently divergent from that in natural models. In fact, frozen abstract-trained weights can efficiently solve cognitive tasks from sensory-realistic inputs, however, at the cost of diverging the dynamics from their original configuration. Together, our findings argue that sensory realism is a primary determinant of human-like behavior in artificial systems, suggesting that findings obtained from abstract task models in neuroscience should be more carefully interpreted in light of their potential behavioral limitations.
Abstract: Existing Time Series Foundation Models (TSFMs) for Multivariate Time Series Anomaly Detection often overlook discrete . This inappropriate modeling approach prevents the model from fully leveraging state information and even leads to significant performance degradation when state variables are integrated. To address this limitation, this paper proposes a novel module that effectively detects anomalies inherent to the state variables themselves. Extensive experiments on real-world datasets demonstrate that STAR significantly improves the performance of existing TSFMs.
PaperID: 7196, Poster
Abstract: Accurate long-term air quality forecasting is challenging due to long-range transport, meteorological variability, and nonlinear chemical reactions underlying secondary pollutant formation. Purely data-driven models often suffer from long-horizon drift and lack physical consistency, while existing physics-informed approaches commonly follow open-loop designs that fuse neural and physical components only at the output level. We propose AeroChem, a closed-loop physics-informed state space framework for chemically reactive air quality forecasting. AeroChem adopts a two-branch co-correction design: a UDE-based process branch provides a structured but imperfect physical prior, while a Mamba-based virtual measurement branch captures long-range temporal context and produces probabilistic virtual observations. Neural Kalman Fusion couples the two branches by correcting the latent physical state at each forecasting step and feeding the posterior state back into the simulator, reducing long-horizon drift. We further introduce HanoiAir, a real-world dataset integrating pollutant concentrations, meteorology, and emission inventories. Experiments on HanoiAir and a public benchmark show that AeroChem improves long-horizon stability and sudden-change forecasting.
Authors:
Chengjun Yu, Shu XU, Jiaqi Wu, Qianben Chen, Tianrui Qin, Zhu, Qiexiang Wang, Jiayu Zhang, Xinpeng Liu, Xin Gui, Jingyi Cao, Yi Yao, WANG PIAOHONG, Dingfeng Shi, He Zhu, Tiannan Wang, Yuqing Wang, Maojia Song, Tianyu Zheng, Jian Yang, Jiaheng Liu, Minghao Liu, Eleanor Jiang, Wangchunshu ZhouAbstract: Recent deep research agents primarily improve performance by scaling reasoning depth, but this leads to high inference cost and latency in search-intensive scenarios. Moreover, generalization across heterogeneous research settings remains challenging. In this work, we propose Search\ More, Think\ Less (SMTL), a framework for long-horizon agentic search that targets both efficiency and generalization. SMTL replaces sequential reasoning with parallel evidence acquisition, enabling efficient context management under constrained context budgets. To support generalization across task types, we further introduce a unified data synthesis pipeline that constructs search tasks spanning both deterministic question answering and open-ended research scenarios with task appropriate evaluation metrics. We train an end-to-end agent using supervised fine-tuning and reinforcement learning, achieving strong and often state of the art performance across benchmarks including BrowseComp (48.6%), GAIA (75.7%), Xbench (82.0%), and DeepResearch Bench (45.9%). Compared to Mirothinker-v1.0, SMTL with maximum 100 interaction steps reduces the average number of reasoning steps on BrowseComp by 70.7%, while improving accuracy.
Abstract: We study a class of bilevel optimization problems in which both the upper- and lower-level problems are minimax problems. Existing studies on bilevel optimization primarily focus on settings where the lower-level problem is an unconstrained or constrained minimization problem, and are therefore not directly applicable to the setting of a minimax lower-level problem considered in this work. To address this gap, we develop penalty-based first-order methods for bilevel minimax optimization. In the deterministic setting, we prove that the proposed method finds an \epsilon-KKT point with \tildeO(\epsilon^-4) oracle complexity. We further show that bilevel optimization problems with constrained lower-level minimization can be reformulated, via Lagrangian duality under Slater's condition, as special cases of our framework and hence can also be solved by our method. This yields an \tildeO(\epsilon^-4) complexity bound for finding an \epsilon-KKT point, improving upon the existing \tildeO(\epsilon^-7) result. Finally, we extend our approach to the stochastic setting and establish that the proposed stochastic method finds a nearly \epsilon-KKT point with \tildeO(\epsilon^-9) oracle complexity. To the best of our knowledge, these are the first deterministic and stochastic first-order complexity results for bilevel minimax optimization.
Abstract: Camera sensor RAW data offers intrinsic advantages for object detection, including richer bit depth, preserved physical information, and freedom from image signal processor (ISP) distortions. However, varying exposure conditions, spectral sensitivities, and bit depths across devices introduce substantially larger domain gaps than sRGB, making sensor-agnostic generalization a fundamental challenge. In this study, we present a physics-guided global-local tone mapping framework for sensor-agnostic RAW object detection. By factoring sensor-induced variations into a global tonal correction and a spatially adaptive local color adjustment, both driven by RAW distribution priors, our framework enables a single network to train jointly across heterogeneous sensors. To further support cross-sensor generalization, we construct a physics-based RAW simulation pipeline that synthesizes realistic sensor outputs spanning diverse spectral sensitivities, illuminants, and sensor non-idealities. Extensive experiments across multiple RAW benchmarks, covering bit depths from 12 to 24, demonstrate state-of-the-art (SOTA) performance in single-dataset, mixed-dataset, and challenging robustness settings.
PaperID: 7200, Poster
Authors:
Yi Gu, Huacan Wang, Shuo Zhang, Yuqing Hou, Xuelei, weipeng.ming, Chen Liu, Fangzhou Yu, Kuan Li, Ronghao Chen, Sen Hu, Mou X Feng, Yi XuAbstract: Large language mode agents are moving beyond text-only interaction toward physical-world control, with smart homes as a representative domain. Real domestic interaction requires understanding ambiguous intents, operating in dynamic environments, and performing multi-turn reasoning. However, existing methods struggle to generate high-quality training data for smart home agents. We propose HomeFlow, a verifiable data flywheel for this domain. HomeFlow uses HomeEnv as a unified simulation environment and HomeMaker to procedurally generate diverse home settings. Subsequently, Blueprint compiles open-ended user intents into executable state-based success conditions, while MCTS-Flow synthesizes diverse, verifiable multi-turn trajectories through environment-guided tree search. We then optimize the agents via supervised fine-tuning and step-wise RLVE, which facilitates iterative improvement through authentic physical feedback. We further construct SmartHome-Bench to evaluate the agent across various smart home tasks. On this benchmark, HomeFlow-RL-4B and HomeFlow-RL-8B achieve task success rates of 84.60% and 87.03%. It is worth noting that HomeFlow-RL-8B even surpasses the leading GPT-5.5 by 1.23 percentage points.
Abstract: Vision encoders for retrieval are typically trained with class-label supervision: each training pair reduces to a scalar that uniformly pushes the embedding apart or pulls it together, as if every visual attribute either differed or matched. A multimodal large language model (MLLM), shown the same pair, can articulate those attributes and use them to predict whether the images share a class. We propose SAGA, a framework that turns this language-grounded, attribute-aware perception into a training signal for the encoder itself. Specifically, we use Group Relative Policy Optimization (GRPO) to reward the MLLM for correct predictions on the vision encoder's tokens. Since correct predictions require those tokens to expose the specific attributes that differ or match between the pair, the gradient pushes the encoder to encode them, replacing the uniform pair-level scalar with attribute-resolved supervision. An auxiliary attention-distillation loss anchors the encoder's embedding to tokens the MLLM attended to, and a standard metric-learning loss shapes the embedding geometry for nearest-neighbour retrieval. The MLLM is frozen throughout and discarded at inference, matching the deployment cost of a metric-learning baseline. SAGA improves Recall@1 by 3 to 6 points over state-of-the-art baselines on CUB-200-2011, Cars-196, FGVC-Aircraft, and iNaturalist Aves on zero-shot image retrieval.
PaperID: 7202, Poster
Abstract: Class-wise embeddings provide an effective strategy for multi-label learning (MLL) by capturing the distinct discriminative properties of each class. However, existing methods based on class-wise embeddings typically overlook the internal distribution of negative samples. We uncover a pivotal phenomenon: negative instances in class-wise spaces naturally organize into multiple semantically coherent sub-clusters, a structure that persists across diverse modalities even under extreme label sparsity. Leveraging this insight, we propose a Structure-Aware Pseudo-label mining framework for weakly supervised MLL (SaPu). The proposed framework mines positive pseudo-labels by aggregating cluster-level consistency across different class-wise embeddings and identifies negative pseudo-labels via neighborhood exclusion. These high-quality signals further guide class-wise contrastive learning, establishing a self-reinforcing loop that iteratively refines embeddings. Extensive experiments on image, text and audio benchmarks validate that SaPu outperforms state-of-the-art methods by an average of
Authors:
Yun Qu, Qi Wang, Yixiu Mao, Heming Zou, Yuhang Jiang, Yingyue Li, Wutong Xu, Lizhou Cai, Weijie Liu, Clive Bai, Kai Yang, Yangkun Chen, Saiyong Yang, Xiangyang JiAbstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard approach for large language models (LLMs) post-training to incentivize reasoning capacity. Among existing recipes, group-based policy gradient is prevalent, which samples a group of responses per prompt and updates the policy via group-relative advantage signals. This work reveals that these optimization strategies share a common geometric structure: each implicitly defines a target distribution on the response simplex and projects toward it via first-order approximation. Building on this insight, we propose Listwise Policy Optimization (LPO) to explicitly conduct the target-projection, which demystifies the implicit target by restricting the proximal RL objective to the response simplex, and then projects the policy via exact divergence minimization. This framework provides (i) monotonic improvement on the listwise objective with bounded, zero-sum, and self-correcting projection gradients, and (ii) flexibility in divergence selection with distinct structural properties through the decoupled projection step. On diverse reasoning tasks and LLM backbones, LPO consistently improves training performance over typical policy gradient baselines under matched targets, while intrinsically preserving optimization stability and response diversity. The code is available at https://anonymous.4open.science/r/LPO.
Abstract: Decentralized learning (DL) is an emerging machine learning paradigm where nodes collaboratively train models without a central server. However, the collaborative nature of DL makes it vulnerable to backdoor attacks, where a model is taught to behave normally on standard inputs while executing hidden, malicious actions when encountering data with specific triggers. Backdoor attacks in DL remain understudied and existing defenses often overlook DL constraints. We introduce Argus, a novel backdoor detection framework native to DL that requires neither a central coordinator nor prior knowledge of the trigger. In Argus, honest nodes locally analyze received model updates to identify potential backdoor triggers. Nodes then collectively share their triggers with their neighbors and use a structural similarity metric to separate true backdoors from false alarms induced by data heterogeneity. A key insight is that false positive triggers exhibit inconsistencies across participants while true positive ones show consistent patterns. Model updates that fail this collaborative test are rejected, and persistently malicious senders are eventually evicted. We provide the first theoretical convergence guarantees for a DL-specific backdoor detection mechanism, showing that filtering out suspicious model updates with high probability preserves a convergence rate comparable to standard DL. We implement and evaluate Argus on three standard datasets and against three state-of-the-art baselines. Across settings, Argus reduces attack success rates by up to 90 points compared to no defense, while preserving model utility within 5 percentage points of an omniscient oracle. Furthermore, the effectiveness of Argus compared to baselines improves as data heterogeneity increases.
PaperID: 7205, Poster
Abstract: Vision-Language Models for Earth Observation are bottlenecked at the visual encoder. Existing pipelines either compress multi-spectral signals into RGB pixels or collapse them into a single CLIP-style global vector, losing spectral information or spatial structure respectively. We show this is a consequence of coupling contrastive supervision to the dense feature stream, which degrades spectral fidelity. SPEXT resolves this through architectural decoupling. A wavelength-parameterized encoder produces multi-spectral patch tokens for dense tasks and VLM grounding, while a dedicated alignment token absorbs the contrastive gradient for retrieval. SPEXT simultaneously leads the PANGAEA segmentation benchmark (+2.51 mIoU), tops zero-shot retrieval among multi-spectral CLIP baselines (+4.6 mAP@100), and enables a frozen VLM, paired with a small projection adapter and LoRA, to answer free-form Earth-Observation questions directly from its tokens, substantially outperforming an RGB-pixel baseline on a downstream EO-RAG benchmark.
PaperID: 7206, Poster
Abstract: While Large Language Models (LLMs) are increasingly deployed as automated evaluators in evaluation and training pipelines, their judgements are often affected by systematic biases that conflict with human preferences. While prior work has identified several known biases and proposed methods for their detection and mitigation, they lack strong grounding in human evaluation preferences, which is essential to ensuring that the identified biases correspond to actual human judgment behavior. Moreover, they rely heavily on pre-discovered bias lists, overlook bias strength, and depend on costly interventions. In this work, we propose HUB-J, an integrated framework grounded in human evaluation preferences that detects, quantifies, and mitigates biases in LLM-as-a-Judge systems. Our approach leverages human–LLM judgement disagreement cases to automatically discover interpretable bias factors, and utilizes agreement cases to quantify bias strength through controlled input modifications and resulting shifts in model decisions. Finally, building on these quantified biases, we introduce a lightweight, training-free regression-based mitigation strategy that corrects bias-influenced judgments by removing the estimated bias effects. Empirical results show that HUB-J uncovers both known and novel bias factors, reveals meaningful differences in model susceptibility, and consistently reduces bias-driven decision flips while generalizing across models.
PaperID: 7207, Poster
Abstract: Deep learning has driven substantial progress in modeling fMRI responses, yet most existing approaches are optimized for a single objective---such as neural encoding, stimulus decoding, or post hoc interpretation---making it difficult to obtain a coherent account of neural representations. We introduce a unified framework that connects vision foundation models to brain activity while supporting neural analysis from multiple complementary perspectives. Our approach replaces black-box features with a sparse, interpretable representation composed of visual prototypes, and links these prototypes to voxel responses through a linear readout. This design makes the model transparent at both levels: prototype activations expose the visual concepts present in the stimulus, while the learned linear weights characterize each voxel's selectivity. Using this interface, we identify which visual concepts drive neural responses and where they appear in the image, reveal cortical organization through voxel tuning in prototype space, and perform controlled stimulus synthesis to probe cortical selectivity through targeted interventions. Together, these results establish prototype-based representations as a unified and interpretable foundation for studying visual representations in the human brain.
Abstract: Training loss and accuracy are the standard signals used to monitor generalization during deep neural network training. Two well-documented phenomena complicate this picture: in grokking, train loss falls rapidly while test performance improves abruptly only after a long delay; in epoch-wise double descent, train loss decreases monotonically while test loss or error rises and falls. Existing accounts are often task-specific, and a task-agnostic analysis framework for diagnosing and explaining these phenomena across realistic tasks and architectures is missing. We address this challenge by analyzing two competing processes that underlie learning dynamics: representation learning in the encoder and readout calibration in the final classifier. Using tools from representational geometry, neural tangent kernels, and linear probing, we show that both processes are active throughout training, with the fluctuations of their relative speed giving rise to seemingly anomalous generalization dynamics. Applying the representation-readout decomposition to grokking across a wide range of tasks and architectures, we find that the readout is train-biased before grokking onset, and representation learning is gradual but not absent—contrary to the lazy-to-rich account. The framework further provides diagnostic signatures distinguishing spurious from genuine generalization: in a previously reported MNIST grokking example and an epoch-wise double descent example, apparent delayed or non-monotone generalization is shown to arise from representation degradation and readout misalignment induced by non-standard training recipes. Together, these results establish the representation-readout decomposition as a top-down framework for understanding learning dynamics and revealing underlying algorithms for interpretability research.
Abstract: Training data attribution (TDA) should enable generative model interpretability and foster a variety of related downstream tasks. Nonetheless, current TDA approaches lack reliability and robustness, preventing their adoption in real-world setups. In this paper, we take a decisive step towards more reliable and robust TDA for diffusion models. We propose to perform TDA with mirrored unlearning and noise-consistent skew (MUCS). The idea is to fine-tune a second model with bounded mirrored gradient ascent, and to measure the normalized skew of this model with respect to the original one using consistent noise samples. We show that, while being conceptually simple and generic, MUCS systematically outperforms existing methods on three different datasets by a large margin. We additionally study the effect that core design choices have on final performance, and analyze novel aspects regarding the overlap of influential instances across generated items and the potential of ensembling TDA approaches. We believe that our findings may have broader implications for more general unlearning setups, as well as for tasks requiring the comparison of diffusion losses.
Abstract: With the rising need for spatially grounded tasks such as Vision-Language Navigation/Action, allocentric perception capabilities in Vision-Language Models (VLMs) are receiving growing focus. However, VLMs remain brittle on allocentric spatial queries that require explicit perspective shifts, where the answer depends on reasoning in a target-centric frame rather than the observed camera view. Thus, we introduce Allocentric Perceiver, a training-free strategy that recovers metric 3D states from one or more images with off-the-shelf geometric experts, and then instantiates a query-conditioned allocentric reference frame aligned with the instruction’s semantic intent. By deterministically transforming reconstructed geometry into the target frame and prompting the backbone VLM with structured, geometry-grounded representations, Allocentric Perceiver offloads mental rotation from implicit reasoning to explicit computation. We evaluate Allocentric Perceiver across multiple backbone families on spatial reasoning benchmarks, observing consistent and substantial gains (~10%) on allocentric tasks while maintaining strong egocentric performance, and surpassing both spatial‑perception-finetuned models and state‑of‑the‑art open‑source and proprietary models.
PaperID: 7211, Poster
Abstract: EEG foundation models (FMs) offer strong population-level priors but degrade on individual subjects due to anatomy-driven distribution shifts. We propose Rest-Tuning, a task-agnostic calibration framework that personalizes a population FM using just 1–3 minutes of unlabeled resting-state EEG. Rest-Tuning employs masked teacher-student alignment to produce a subject-calibrated backbone, reused as initialization across multiple tasks from the same subject. Across two multi-subject datasets and four EEG FMs, Rest-Tuning outperforms task-only adaptation, with up to 13.3% absolute gain in downstream task accuracy and over 40% average reduction in labeled task data. Calibration saturates quickly and requires under 3 minutes of compute per subject. Representation analyses show reduced subject-to-population distance and smoother LoRA adapter interpolation paths, indicating improved compatibility between rest- and task-adapted solutions. These results establish resting-state EEG as a practical, fast, and reusable calibration interface, enabling population EEG FMs to function as reliable, subject-specific BCIs and serve as reusable initialization for downstream task adaptation.
PaperID: 7212, Poster
Abstract: Reinforcement learning from formal specifications often separates task definition, reward shaping, curriculum design, and policy diagnostics into different mathematical objects. This fragmentation introduces additional conversion steps and discards structure already present in the specification, such as sub-task boundaries, ordering constraints, and quantitative margins. We introduce TBTCL, a framework that uses Temporal Behaviour Trees (TBTs) as a unified interface for specification-guided reinforcement learning. A single TBT specifies the task, induces dense robustness-based rewards, defines a curriculum, and supports better understanding of learnt policies - grounded in the formal semantics of the specification language. By adapting TBT syntax and semantics for specification guided reinforcement learning, we provide a formalism that is more expressive than popular specification languages such LTL and inherently hierarchical. We present algorithms to automatically extract a curriculum DAG and derive dense, robustness-based rewards directly from TBT semantics, requiring no additional user inputs like region predicates or demonstrations. Furthermore, we utilize TBT trace segmentation to provide build understandable competence profiles and automatically adapt the reward toward bottlenecks. We evaluate TBTCL on six specifications across discrete and continuous domains. The results show that TBTCL matches dense STL shaping on simpler tasks and improves over baselines on reach-avoid specifications, while also exposing interpretable sub-task failures through the same formal object used for training. We show that a Temporal Behaviour Tree act as a specification, and a reward, and a curriculum and a diagnostic.
Abstract: We propose using Vision-Language Models (VLMs) for macro placement in chip floorplanning, a complex optimization task that has recently shown promising advancements through machine learning methods. Because human designers rely heavily on spatial reasoning to arrange components on the chip canvas, we hypothesize that VLMs with strong visual reasoning abilities can effectively complement existing placement algorithms. We introduce VeoPlace (Visual Evolutionary Optimization Placement), a novel framework that uses a VLM—without any fine-tuning—to guide the actions of a base placer by constraining them to subregions of the chip canvas. The VLM proposals are iteratively optimized through an evolutionary search strategy with respect to resulting placement quality. Using global half-perimeter wirelength (gHPWL) after standard-cell placement and legalization as the evaluation metric, VeoPlace boosts ChiPFormer performance on 9 of 10 benchmarks with peak reductions exceeding 32%. We further demonstrate that VeoPlace generalizes to analytical placers, improving DREAMPlace on all 8 evaluated Superblue benchmarks with gains up to 4.3%. Our approach opens new possibilities for electronic design automation tools that leverage foundation models to solve complex physical design problems.
PaperID: 7214, Poster
Authors: Md Zarif Hossain, Awal Ahmed Fime, Ahmed Imteaj
Abstract: Robustness and accurate alignment with visual content are fundamental to the reliable deployment of large vision-language models (LVLMs). In this work, we uncover a previously overlooked trade-off: improving adversarial robustness can inadvertently increase hallucination in LVLMs. Through extensive evaluations on both open-ended generation and discriminative benchmarks, we reveal that adversarially trained LVLMs consistently produce outputs that are less aligned with the visual input, despite their improved robustness. To understand this trade-off, we analyze the visual token embeddings of robust encoders and find that adversarial training increases inter-token similarity by up to 38%. We refer to this phenomenon as spatial token homogenization. We further demonstrate that spatial token homogenization degrades visual information and amplifies decoder's reliance on language priors. Motivated by our findings, we propose Homogenization-Aware Latent Steering (HALS), a training-free inference-time method that steers decoder hidden states away from homogenization-induced hallucination without modifying the robust vision encoder. Extensive experiments across five hallucination benchmarks show that HALS consistently outperforms existing hallucination mitigation methods on robust LVLMs, reducing CHAIR_S by 31.3 points on average, improving POPE accuracy by 2.6 points and achieving a 30.6% reduction in hallucination on AMBER. Moreover, robustness evaluations confirm that HALS preserves adversarial performance, demonstrating that robust LVLMs can be made visually grounded without sacrificing their performance.
PaperID: 7215, Poster
Authors: Sebastian Persson, Giacomo Fabrini, Branwen Snelling, Fabian Fröhlich
Abstract: Neural ordinary differential equations (NODEs) and universal differential equations (UDEs) provide flexible and popular frameworks for learning interpretable dynamical systems from noisy time-series data. However, training these models remains challenging, and versatile methods that robustly handle sparse and noisy data as well as partially observed models are lacking. To address this, we introduce curriculum multiple shooting (CMS), a general-purpose training strategy for fitting ordinary differential equation (ODE) models to time-series data by integrating curriculum learning with multiple shooting. Across twelve benchmarks spanning simulated and real data, and covering NODEs, UDEs, and mechanistic ODEs, CMS accelerates and stabilises training convergence, outperforms state-of-the-art training strategies, and ranks among the best methods in generalisation. Finally, we discuss possible explanations for the strong performance of CMS in light of contemporary theories of what makes training on time-series challenging.
Abstract: Randomized smoothing provides a powerful framework for certifying robustness against adversarial perturbations by injecting randomized noise either into the training process (against poisoning attacks) or model inputs (against evasion attacks). Yet, extending these guarantees to hold against backdoor attacks, where adversaries jointly perturb both training and test data, remains challenging. In particular, certifying general mechanisms requires tight, compositional guarantees over heterogeneous randomized components, which are not jointly supported by existing approaches. We address this gap by proposing a general framework that numerically composes robustness guarantees for arbitrary mechanisms while remaining tightly connected to black-box randomized smoothing guarantees. To achieve this, we establish a connection between randomized smoothing and the dual perspective of differential privacy, enabling us to leverage advances in the analysis of differentially private mechanisms together with tight numerical composition. This yields a modular framework in which robustness guarantees are derived by reasoning about the privacy of individual components and composing them end-to-end. We instantiate our framework for DP-SGD and Deep Partition Aggregation with inference-time smoothing, deriving joint robustness guarantees against both training-time and inference-time attacks. Empirically, we demonstrate the effectiveness of our framework on MNIST and CIFAR-10. Overall, we provide a principled and general framework for certifying robustness under complex joint threat models and mechanisms, laying the groundwork for future research on unified certification methods towards guarantees under more complex real-world adversaries.
PaperID: 7217, Poster
Authors:
Ricardo Parada, Akbarkhuja Anvarkhujaev, Carlos Martin, William Chang, Long Tran-Thanh, Nam P TranAbstract: Computing Shapley values exactly requires exponentially many evaluations of a cooperative value function, making approximation essential in high-dimensional applications. Recent leverage-score-based methods yield provably accurate estimators with sample complexity scaling linearly in the number of players, but this can still be too costly when only a small subset of players is truly influential. We study Shapley value estimation under a sparse influence model, where the Shapley vector is supported on an unknown set of size s \ll n. Using the regression characterization of Shapley values, we propose a sparse leverage sampling scheme and analyze it via sparse matrix concentration. Our main result shows that the resulting estimator achieves relative-error approximation with \widetildeO(s^2) value function evaluations, up to logarithmic and accuracy factors, thereby replacing the ambient-dimension dependence of prior guarantees with dependence on the intrinsic support size. Empirically, our method matches or improves over dense Leverage SHAP on sparse synthetic and real-data benchmarks, while exhibiting better numerical stability at small sample sizes.
PaperID: 7218, Poster
Authors:
Min-Seong Kim, Jongbok Won, Jeesup Park, Yonghee Choi, Dong-Jun HanAbstract: Fine-tuning large language models (LLMs) on downstream tasks often weakens their previously aligned safety behavior, creating an inherent trade-off between task adaptation and safety preservation. Existing approaches attempt to mitigate this issue by identifying safety-relevant neurons, layers, or update directions in the original parameter space. However, safety-related information is often distributed across many directions in this space, causing downstream updates to overwrite parameters that encode safety behavior. We propose WSR-Tune, a Weight Space Reparameterization-based fine-tuning framework that preserves safety during downstream adaptation. Our approach constructs a safety-conditioned basis and reparameterizes the weight matrices such that safety-relevant information becomes concentrated in a small subset of directions. We then identify safety-critical directions and freeze them in the reparameterized space, while updating the remaining complementary directions for downstream adaptation. By explicitly structuring the parameter space to disentangle safety and task information, our method reduces interference between objectives and enables simultaneous preservation of safety and improvement in downstream performance. Experimental results show that our method achieves a more favorable trade-off between downstream performance and safety retention, demonstrating its effectiveness for reliable LLM fine-tuning.
Abstract: Gradient descent on overparameterized neural networks typically operates at the Edge of Stability (EoS), where the largest Hessian eigenvalue hovers around a step-size-dependent threshold. We study how sparse connectivity changes generalization below this threshold in two layer ReLU networks. Prior results have shown that for fully-connected networks (FCNs), generalization guarantees in this regime degrade and become vacuous on high-dimensional spherical inputs. Our analysis reveals that sparse connectivity fundamentally alters this picture. Under sparse connectivity, the network processes a collection of low-dimensional patches rather than the full input vector, so the effective constraint imposed by the stability condition is governed by the geometry of the training patch collection. We prove that when the receptive fields are small relative to the ambient dimension, the effective constraint yields non-vacuous generalization bounds in precisely the spherical regime where FCNs provably fail. The same framework also reveals a contrasting failure mode: if the patch collection lacks geometric structure, the constraint becomes unable to prevent overfitting. We corroborate this theory by analyzing the patch geometry of natural images, showing that standard convolutional designs produce patch multiset with low-dimensional structure that facilitates generalization. This provides a principled explanation for the generalization advantage of convolutional networks. Thus, our analysis yields a unified framework that identifies how architecture, data geometry, and gradient descent jointly govern generalization performance.
PaperID: 7220, Poster
Authors: Jie Hou, Qiang Zeng
Abstract: Knowledge editing aims to rewrite sensitive information in language models, yet prior work has shown that edited models remain vulnerable to extraction attacks. In this work, we identify a previously overlooked vulnerability, which we term rank leakage: although editing replaces the top-1 output with a sanitized answer, the original sensitive token often remains highly ranked in the model’s next-token distribution. Exploiting this phenomenon, we propose an effective black-box Output Attack that recovers sensitive information using only a single query. Despite its simplicity, the attack achieves strong extraction performance and remains effective even against models equipped with state-of-the-art defenses. To mitigate this vulnerability, we introduce Rank-guided Concealment Editing (RCE), a plug-in defense that enforces explicit control over the rank of sensitive tokens during editing. Although designed for the Output Attack, RCE generalizes to existing black-box and white-box extraction attacks. Extensive experiments across multiple models, datasets, and editing algorithms show that RCE consistently improves resistance to attacks while preserving edit fidelity and model utility. Our results highlight the importance of controlling the rank of sensitive tokens, rather than only the top-1 prediction, for privacy-preserving knowledge editing.
PaperID: 7221, Poster
Authors:
Benjamin Huh, Hak H Kim, Yuting Tian, Jason Peng, Christopher Kang, Soroush VosoughiAbstract: Rotary positional embeddings (RoPE) are now the dominant positional mechanism in modern large language models. By construction, RoPE encodes a shift-invariant (SI) channel in attention: head logits can depend purely on token offset rather than absolute position. We show that trained models do not just inherit this SI channel; they actively learn to exploit it. This exploitation is load-bearing: SI ablations produce systematic degradation that tracks per-head SI amplitude under converging specificity controls, and SI structure appears even in models trained with absolute positional encodings, indicating that this is learned computation rather than a purely architectural artifact. This SI exploitation is also functionally meaningful: SI ablation preferentially disrupts offset-structured retrieval, most robustly in Llama and Mistral, with directional extension to naturalistic long-context settings. At the same time, our results surface a deeper question: is the SI channel currently underutilized? Many computations are position-invariant by nature --- recognizing that a pattern holds regardless of where it appears in context --- yet our evidence suggests current models allocate SI channels mostly to narrow positional and surface routing rather than these richer invariances. Strikingly, SI exploitation amplitude varies more than sixfold, a spread that holds across eleven models spanning seven architecture families, indicating that degree of SI use is a trainable property rather than an architectural given. This points to a practical direction: pretraining pressure toward semantic invariance tasks may push models to exploit SI channels for richer computation than what we observe today.
PaperID: 7222, Poster
Abstract: Predicting where multiple agents will move next is fundamental to autonomous driving, robotics, and any system that must share space with people safely. Modern stochastic predictors read off inter-agent interactions once from the observed past and keep that view fixed throughout the generative rollout, so the reciprocal geometric constraints that emerge as the joint future sharpens are not re-examined. We address this gap with SRA (Spatial Reasoning Adapter), a modular graph denoiser that plugs into existing diffusion predictors without altering their generative cores. At every reverse step, SRA reads the host's current future estimate, builds a sparse Euclidean interaction graph on the predicted trajectories, and returns a residual correction to the host's decoding path. Our adapter consists of three core components: a sparse top-N neighborhood retrieved with an asymmetric query-key score, a future-conditioned relational feature recomputed on the evolving prediction, and uncertainty-weighted message passing where confident agents act as anchors for uncertain ones. We instantiate SRA on three structurally different hosts - LED, MID, and MoFlow - across four multi-agent benchmarks: SDD, NBA, Soccer, and Football. SRA improves every host on every dataset by 6--7% on average, indicating that spatial reasoning over the model's own evolving future is broadly useful across denoising architectures. The code will be released upon acceptance.
Abstract: Generating humorous memes is a challenging multimodal task that moves beyond direct image-to-caption supervision. It requires a nuanced reasoning over visual content, contextual cues, and subjective humor. To bridge this gap between visual perception and humorous punchline creation, we propose HUMOR, a novel framework that guides VLMs through hierarchical reasoning and aligns them with group-wise human preferences. First, HUMOR employs a hierarchical, multi-path Chain-of-Thought (CoT): the model begins by identifying a template-level intent, then explores diverse reasoning paths under different contexts, and finally anchors onto a high-quality, context-specific path. This CoT supervision, which traces back from ground-truth captions, enhances reasoning diversity. We further analyze that this multi-path exploration with anchoring maintains a high expected humor quality, under the practical condition that high-quality paths retain significant probability mass. Second, to capture subjective humor, we train a pairwise reward model that operates within groups of memes sharing the same template. Following established theory, this approach ensures a consistent and robust proxy for human preference, even with subjective and noisy labels. The reward model then enables a group-wise RL optimization, providing a theoretical guarantee for monotonic improvement within the trust region. Extensive experiments show that HUMOR empowers various VLMs with superior reasoning diversity, more reliable preference alignment, and higher overall meme quality. Beyond memes, our work presents a general paradigm for open-ended, human-aligned multimodal generation, where success is guided by comparative judgment within coherent output groups.
PaperID: 7224, Poster
Authors:
Junhee Lee, JeonDongHyeon, Taeoh Kim, beomyoung kim, MyeongAh ChoAbstract: Localizing AI-edited regions is essential for interpretable forensic analysis, but remains challenging due to subtle and spatially distributed artifacts that are misaligned with semantic or object boundaries. Existing approaches rely on pixel-level supervision from controlled editing pipelines, which is difficult to scale and can introduce misleading signals: artifacts frequently extend beyond annotated regions, while out-of-mask pixels are treated as authentic. This limits models' ability to capture transferable evidence and generalize across generators and datasets. To address these issues, we propose ReGFLoW, a Reconstruction-Guided Fake Localization framework under Weak supervision, which is the first weakly supervised approach for diffusion-edited fake region localization. ReGFLoW requires only real/fake labels at the image-level and uses diffusion reconstruction errors as dense spatial guidance to inject them into both feature and score spaces. Furthermore, by artifact-centric multiple instance learning, ReGFLoW utilizes localized diffusion evidence without relying on semantic-affinity or boundary-based pseudo-mask priors. Extensive experiments demonstrate that ReGFLoW achieves stronger out-of-domain generalization than fully supervised learning baselines.
Authors: S M RUHUL KABIR HOWLADER, Xiao Chen, Yifei Xie, Lu Liu
Abstract: Federated learning (FL) encounters substantial challenges due to heterogeneity, leading to gradient noise, client drift, and partial client participation errors, the last of which is the most pervasive but remains insufficiently addressed in current literature. In this paper, we propose FedAdaVR, a novel FL algorithm aimed at solving heterogeneity issues caused by sporadic client participation by incorporating an adaptive optimiser with a variance reduction technique. This method takes advantage of the most recent stored updates from clients, even when they are absent from the current training round, thereby emulating their presence. Furthermore, we propose FedAdaVR-Quant, which stores client updates in quantised form, significantly reducing the memory requirements (by 50%, 75%, and 87.5%) of FedAdaVR while maintaining highly competitive model performance. We analyse the convergence behaviour of FedAdaVR under general nonconvex conditions and prove that our proposed algorithm can eliminate partial client participation error. Extensive experiments conducted on multiple datasets, under both independent and identically distributed (IID) and non-IID settings, demonstrate that FedAdaVR consistently outperforms state-of-the-art baseline methods.
Abstract: We study the problem of learning with multiple correct answers, where each instance admits a set of valid labels. We primarily focus on the online setup, where in each round the learner must output a valid label for the queried example. This setting is motivated by language generation, in which a prompt may admit many acceptable completions, but not every completion is acceptable. We study this problem under three feedback models. For each model, we characterize the optimal mistake bound in the realizable setting using an appropriate combinatorial dimension. We then show that the rate of regret can be constant, linear, or sublinear across the three models in the agnostic setting. Our results also imply sample complexity bounds for the batch setup that depend on the respective combinatorial dimensions.
PaperID: 7227, Poster
Abstract: When does training language models (LMs) on explanations yield faithful introspection, rather than superficial imitation? Surprisingly, we find that LMs trained to explain the predictions of similar models frequently produce explanations more faithful to their own current behaviors than to those of their training targets. This "introspective'' coupling between the model's explanations and behaviors occurs only when the training target explanations remain sufficiently similar to model behaviors over the course of training. This alignment must be preserved throughout the training process: introspection only emerges when explanation supervision is sufficiently behaviorally compatible with the model as it changes. Finally, we show that introspection generalizes to variants of the training problem: when introspection training is run concurrently with training that shifts a model's behaviors, explanations track those behavioral shifts without requiring updated supervision. This holds across a diverse range of tasks, including sycophancy and refusal, and is robust to label noise. These results suggest that introspection training is a viable component of post-training: explanation labels need not be continually refreshed, and faithfulness extends to regions of input space not explicitly supervised.
PaperID: 7228, Poster
Abstract: While Large Vision-Language Models (LVLMs) excel at bounding-box grounding, they struggle with precise spatial tasks such as polygon grounding. We attribute this bottleneck to two properties of the vertex-level autoregressive (AR) paradigm used by current LVLMs: (1) errors in early-emitted vertices propagate uncorrected through the rest of the sequence, and (2) the model commits to local vertex placement before observing the full contour, leading to suboptimal allocation of a fixed vertex budget. We propose Diffusion Fine-Tuning (DFT), which removes the vertex-level sequential dependency by placing all 2N coordinate tokens under a single discrete-diffusion denoiser, while introducing a much shorter digit-level factorisation: each coordinate is decomposed into hundreds, tens, and units digits, and the reverse process is trained to predict them coarse-to-fine. We train this with a Hierarchical Curriculum Learning strategy that progressively refines loss supervision from macro-contour to per-pixel detail. Under matched fine-tuning protocols, DFT matches strong AR LVLMs on 2D bounding-box grounding and improves over them on 16-point polygon grounding; the same network and training recipe extend to 9-DoF monocular 3D bounding-box grounding under a 9-parameter coordinate parameterisation. A block-wise top-k decoder closes most of the latency gap to AR and recovers most of the quality lost when scaling beyond 16 vertices. We do not claim to eliminate sequential dependence in general: vertex-level O(N) AR decoding is replaced by K joint denoising steps (K=12 in this work), with the three-stage digit hierarchy entering as a training-time loss factorisation rather than an inference-time chain.
PaperID: 7229, Poster
Abstract: Vision-large language models (VLLMs) editing aims to efficiently update model responses while retaining generalization and unrelated knowledge. Recent online editors store previous edits as reusable experts, forming a mixture-of-experts editing memory for scalable VLLM correction. By retrieving relevant experts at inference time, this paradigm avoids repeatedly overwriting model parameters and enables continual reuse of accumulated editing experience. However, existing sample-level residual experts usually encode each edit instance as a whole, entangling target-relevant visual evidence with irrelevant context. This coarse expert construction weakens expert reusability, leads to inaccurate routing, and increases unintended side effects during continual editing. To address this limitation, we reformulate online VLLM editing as evidence guided expert construction and routing. A fine mask identifies edit-relevant visual evidence, from which support-suppress experts are constructed to strengthen target consistent cues and weaken edit conflicting cues. Each expert is stored with routing signals and reliability scores, and later retrieved through joint sparse routing over visual evidence, textual context, and expert quality. Experiments demonstrate consistent improvements in editing reliability, generalization, and interpretability over competitive online editing baselines.
PaperID: 7230, Poster
Abstract: This paper studies the problem of uncertainty quantification in large language models (LLMs), which is critical for reliable LLM deployment under the risk of hallucination. Existing approaches typically treat the response sets as unordered collections and rely on local pairwise similarity or clustering, which overlook relational signals among responses as well as their corresponding high-order structural semantics. Towards this end, we propose a novel approach named Complementary Signed Graph Propagation (CHAIN) for uncertainty quantification in LLMs. The core idea of CHAIN is to capture complementary signals among responses using a signed graph and then extract high-order structural information using random propagation. In particular, CHAIN builds a semantic graph where each response is considered as a node and edges represent complementary views, namely enrichment relations and contradiction relations. To acquire high-order structural signals, we update hidden states through graph-based information propagation along edges and channel fusion guided by randomized matrices. After propagation, these node states are summarized into graph-level scores for reliable uncertainty quantification. Extensive experiments on five datasets validate the effectiveness of the proposed CHAIN in comparison with competing baselines.
Abstract: We introduce Hyper Input Convex Neural Networks (HyCNNs), a novel neural network architecture designed for learning convex functions. HyCNNs combine the principles of Maxout networks with input convex neural networks (ICNNs) to create a neural network that is always convex in the input, theoretically capable of leveraging depth, and performs reliable when trained at scale compared to ICNNs. Concretely, we prove that HyCNNs require exponentially fewer parameters than ICNNs to approximate quadratic functions up to a given precision. Throughout a series of synthetic experiments, we demonstrate that HyCNNs outperform existing ICNNs and MLPs in terms of predictive performance for convex regression and interpolation tasks. We further apply HyCNNs to learn high-dimensional optimal transport maps for synthetic examples and for single-cell RNA sequencing data, where they oftentimes outperform ICNN-based neural optimal transport methods and other baselines across a wide range of settings.
Abstract: While Large Language Models (LLMs) are increasingly deployed for table-related tasks, the internal mechanisms enabling them to process linearized two-dimensional structured tables remain opaque. In this work, we investigate the process of table understanding by dissecting the atomic task of cell location. Through activation patching and complementary interpretability techniques, we delineate the table understanding mechanism into a sequential three-stage pipeline: Semantic Binding, Coordinate Localization, and Information Propagation. We demonstrate that models locate the target cell via an ordinal mechanism that counts discrete delimiters to resolve coordinates, offering mechanistic evidence across diverse table formats. Furthermore, column indices are encoded within a linear subspace that allows for precise steering of model focus through vector arithmetic. Finally, we extend our analysis to the real-world HiTab dataset and show that the same three-stage mechanism persists in hierarchical real-world tables and more complex tasks.
PaperID: 7233, Poster
Abstract: Large language models often produce fluent reasoning chains that arrive at incorrect answers, and these failures are hard to detect from output probabilities alone. We show that a strong signal of correctness is reflected in the geometry of internal computation: hidden-state trajectories across transformer layers differ between correct and incorrect reasoning, with correct solutions following smooth, efficient paths and failures exhibiting elevated curvature, abrupt directional changes, and increased geodesic deviation. We capture this signal through VANE, five parameter-free geometric features (velocity, acceleration/jerk, curvature, geodesic deviation, and token coherence) computed from a single forward pass over layer-wise hidden states. Across six models (1.5B to 72B parameters) and five benchmarks spanning mathematics, code, and verbal reasoning, VANE achieves 70.7 to 96.6 AUROC and exceeds token log-probability on all 30 model–benchmark pairs (mean +25 percentage points, up to +50 percentage points). The geometric signal is near-independent of output confidence (r = -0.26), and a base-model control shows it emerges with instruction tuning. As a downstream application, filtering geometrically unstable outputs raises accuracy to 96.1% at 50% coverage on a 4-bit quantized 72B model. These results show trajectory geometry is a reliable single-pass indicator of reasoning correctness, even when the model is confidently wrong.
PaperID: 7234, Poster
Abstract: Human values are not single labels, but distributions: they vary across individuals, contexts, and their relationships with other values. Yet most alignment approaches still compress group-level values into scalar scores or single representative targets, erasing within-group diversity and missing how value expression shifts across situations. We introduce Distributional Value Profiling, a framework for representing group-level value expression as distributions over expression modes, intensity, and inter-value relationships. Using large-scale text corpora spanning cultural, political, and religious groups, we construct profiles that capture not only which values are expressed, but also how they are expressed, how strongly they are emphasized, and how they relate to one another. We then model context-dependent shifts in these profiles using optimal transport, revealing structured redistributions of value expression across situations. Finally, we show that these profiles are actionable: a lightweight profiling model generalizes to unseen value-context combinations, and an activation-based steering method guides model outputs toward target distributions at inference time, consistently outperforming baselines across diverse benchmarks. Together, these results establish distributional value profiling as a concrete step toward pluralistic alignment, treating within-group diversity not as noise to be averaged away, but as structure to be measured, predicted, and controlled.
Abstract: Multimodal Large Language Models have achieved remarkable progress in short-form audio-video understanding, yet long-form audio-video comprehension remains challenged by limited context windows and severe information redundancy. To address these bottlenecks, we propose AVOC, a framework for long-form audio-video understanding in Omni-modal Large Language Models. AVOC introduces a learnable token compression module between the modality encoders and the LLM backbone. We reframe multimodal token compression as a top-K retrieval problem: given a fixed context budget, the module must retrieve a compact subset of tokens that best supports answering the user query. We draw inspiration from three classical Information Retrieval criteria for selecting informative units from a large candidate pool: relevance, importance, and diversity. AVOC instantiates each criterion as a tailored mechanism for audio-video understanding, and integrates them into a unified retrieval-style compression pipeline. Experiments show that AVOC achieves state-of-the-art performance on long-form audio-video benchmarks, surpassing the second-best model by 4.9 and 5.5 points in average accuracy on OmniVideoBench and LVOmniBench, respectively. Moreover, AVOC maintains robust performance on Audio-Video Needle-in-a-Haystack task at durations up to one hour.
Abstract: Denoising models such as Diffusion or Flow Matching have recently advanced generative modeling for discrete structures, yet most approaches operate directly in the discrete state space, leading to abrupt state changes. We introduce simplex denoising, a simple yet effective generative framework that operates on the probability simplex. The key idea is a non-Markovian noising scheme in which, for a given clean data point, noisy representations at different times are conditionally independent. While preserving the theoretical guarantees of denoising-based generative models, our method removes unnecessary constraints, thereby improving performance and simplifying the formulation. Empirically, \emphunrestrained simplex denoising surpasses strong discrete diffusion and flow-matching baselines across synthetic and real-world graph benchmarks.
PaperID: 7237, Poster
Abstract: Robust visual classification often depends on localizing the main foreground objects in an image while ignoring spatially-separable distractors. Surprisingly, we find that the attention maps of smaller self-supervised ViTs localize foreground objects better than those of larger ones. However, we still need large ViTs, because they extract richer representations from each patch. To get the best of both worlds, good localization _and_ rich representations, we propose A^2, a simple method that leverages this inverse scaling finding by decoupling _where to look_ (a small attention model) from _what to extract_ (a large embedding model): we crop around the attention peaks of a small model and embed the crops with a larger model. A^2 uses entirely pretrained features and does not require per-dataset attention or backbone training. Across 5 benchmarks, A^2 is competitive with backbone-matched loss-level methods like DFR, and outperforms end-to-end attention training under stronger distribution shifts.
PaperID: 7238, Poster
Abstract: Machine unlearning is becoming increasingly indispensable for meeting ever-stringent data compliance requirements, such as the ``right to be forgotten''. While Gradient Ascent (GA) based approximate methods have drawn attention for their conceptual simplicity, their practical deployment is severely constrained by three fundamental issues: first, the lack of feasibility analysis leads to divergent update directions; second, gradient conflict inevitably degrades the performance maintenance on retained data; and third, heavy reliance on the retain set restricts their applicability in real-world settings. To address these challenges, we conduct an in-depth investigation from a geometric perspective into the underlying root causes of instability in conventional GA methods. Our analysis reveals that the feasibility of unlearning intrinsically depends on specific geometric constraint relationships within the parameter space. Building upon this theoretical insight, we formally establish the methods governing unlearning feasibility. The first is the direction feasibility, which constrains the feasible update trajectory and keeps the model within reasonable parameter regions during unlearning. The second is the step-size feasibility, which further restricts the update magnitude along the feasible direction to prevent model collapse caused by large steps.Notably, we find that the direction feasibility does not rely on any information from the retain set. In light of this, we further explore viable approaches for effectively estimating the step-size feasibility under strict source-free scenarios.
Abstract: In adversarial multi-armed bandits, two performance measures are commonly used: static regret, which compares the learner to the best fixed arm, and dynamic regret, which compares it to the best sequence of arms. While optimal algorithms are known for each measure individually, there is no known algorithm achieving optimal bounds for both simultaneously. \citemarinov2021pareto first showed that such simultaneous optimality is impossible against an \emphadaptive adversary. Our work takes a first step to demonstrate its possibility against an \emphoblivious adversary when losses are \emphdeterministic. First, we extend the impossibility result of \citemarinov2021pareto to the case of deterministic losses. Then, we present an algorithm achieving optimal static and dynamic regret simultaneously against an oblivious adversary. Together, they reveal a fundamental separation between adaptive and oblivious adversaries when multiple regret benchmarks are considered simultaneously. It also provides new insight into the long open problem of simultaneously achieving optimal regret against switching benchmarks of different numbers of switches. Our algorithm uses negative static regret to compensate for the exploration overhead incurred when controlling dynamic regret, and leverages Blackwell approachability to jointly control both regrets. This yields a new model selection procedure for bandits that may be of independent interest.
Abstract: Defensive training methods such as positive preventative steering (PPS) and inoculation prompting (IP) offer surprising results through seemingly similar processes: both add trait-inducing objects to large language models (LLMs) during training, and both defend the LLM against acquiring the trait. The surprising success of these methods comes with the question: how do they work? Are PPS and IP doing the same thing? We provide behavioral and mechanistic comparisons of these two methods using "evilness" as a case-study trait. Our central finding is that PPS and IP achieve their defensive benefits through distinct mechanisms. Behaviorally, we show that neither PPS nor IP operates through a purely associative mechanism; and PPS can both defend against trait acquisition and actively reduce pre-existing expression, whereas IP is ineffective in models that were previously finetuned to express the trait. This behavioral divergence is reflected mechanistically: PPS shifts the activation gradient towards an attenuating direction along the PPS vector axis. When the PPS vector is aligned with a trait-expressing axis, it can reverse the gradient pressure, reducing rather than increasing activation along that axis. In contrast, IP continues to resist a precise mechanistic account. Direct cosine similarity analyses reveal that IP has a characteristically different gradient signature than PPS, and qualitative analyses reveal IP's gradient to be more diffuse. Furthermore, IP reduces the next-token prediction loss on trait-expressing data where PPS need not, consistent with the notion that IP "explains away" the trait-expression in the training data. Taken together, our analyses reveal distinct mechanisms by which each method operates and highlight open questions about IP's mechanistic picture.
PaperID: 7241, Poster
Authors:
Sizhuo Zhou, Xiaosong Jia, Fanrui Zhang, Qifeng Li, Junjie Li, Zirui Wang, Yukang Feng, Yu Hong, Shaofeng Zhang, Wenlong Liao, Tao He, Juyong Zhang, Junchi YanAbstract: Generating LiDAR point clouds from multi-view camera images is a valuable task with applications in controllable data synthesis, cross-modal simulation and various perception tasks. However, most existing methods are not specifically designed for this task, and the quality of their generated LiDAR remains limited. Moreover, these methods typically focus only on semantic cues in images or generate LiDAR from a monocular camera image. In this work, we introduce LiFi, a dedicated framework for generating visually and geometrically aligned LiDAR scenes from multi-view camera images. LiFi guides the latent diffusion process through two complementary branches: geometry and semantics. The geometry branch constructs scene features from images through a reverse sampling algorithm and an uncertainty-aware geometric encoder combined with DepthAnything3, whereas the semantic branch provides high-level semantic representations via a cross-view interaction module. We further introduce a dual-branch balanced classifier-free guidance (CFG) strategy, which enhances the model’s conditional generation capability while preserving the independent learning and collaborative controllability of the two branches. Extensive experiments demonstrate that LiFi outperforms state-of-the-art methods in generating high-fidelity LiDAR scenes. We will make this project publicly available.
PaperID: 7242, Poster
Abstract: Computing the unregularized Wasserstein barycenter for measure-valued data is a challenging optimization task. Recent algorithms have been tailored to either discrete measures as point clouds or continuous measures discretized on regular grids. In this work, we propose a primal mirror descent algorithm for computing the exact Wasserstein barycenter in the Fisher-Rao geometry. Our algorithm is a unified approach that is flexible enough to simultaneously cover discrete and absolutely continuous input measures, with convergence guarantees in both settings. In particular, when all input measures are discrete, our algorithm, initialized from any probability density, solves a sequence of semi-discrete optimal transport subproblems and produces absolutely continuous iterates that converge to the discrete barycenter. We use synthetic and real data examples to demonstrate the promising result in terms of accuracy and computational cost.
PaperID: 7243, Poster
Abstract: To surface "objective" fact-check notes, X's Community Notes uses Biased Matrix Factorization (BMF) on crowd-sourced ratings to separate note quality from reviewer ideology. We show that for any correctly specified model that decomposes ratings into quality and ideological alignment, this separation requires fixing an ideological anchor: a reference point for what counts as neutral. We give a characterization for the standard implementation: with L_2 regularization, BMF anchors quality at a point proportional to the mean reviewer ideology, recovering objective quality if and only if the mean reviewer is neutral. This parallels the Condorcet Jury Theorem: collective judgment recovers truth when the pool is unbiased on average. When that assumption fails, BMF's recovered ranking shifts toward maximizing aggregate reviewer satisfaction, which coincides with social welfare only when the pool is representative. This is not specific to BMF; the rating matrix admits a continuous family of gauge-equivalent decompositions, each implying a different quality ranking, so no mechanism based on ratings alone can recover quality without implicitly choosing such an anchor. What appears to be routine hyperparameter tuning is in fact a normative choice about whose preferences define quality.
Authors: Kartik Bali, Mahish Kumar Guru, Christian J Cyron, Roland Aydin
Abstract: Adaptive volumetric finite element meshing is a critical step in computer-aided engineering and analysis that dictates the computational budget of a given problem. It traditionally requires iterative PDE solvers or heavily supervised, data-driven surrogates trained on large-scale simulation data. While Multimodal Large Language Models (MLLMs) excel in 2D visual tasks, their zero-shot capability to semantically ground regions based on geometric understanding and physics remains an open question. Overall, this study explores a significant question: can the high-level semantic understanding of off-the-shelf MLLMs serve as a viable, zero-shot geometric proxy for finite element mesh refinement? To investigate this, we introduce (Geometric Reasoning Enhanced Multimodal LLMs for Finite Element Meshing), a framework that utilizes MLLMs to visually localize stress-critical regions based on physics-guided textual prompts. To bridge the gap between 2D MLLM pre-training and 3D geometries, we introduce , a view-selection module that maximizes the observability of key geometric features. We conduct an in-depth empirical evaluation across diverse CAD geometries, loading cases, and SOTA MLLMs, comparing them against a tuned geometric heuristic under a strict, matched refinement budget. Our findings reveal that MLLMs demonstrate robust zero-shot capacity to accurately follow complex spatial-physical instructions, isolating stress-relevant features with higher precision than blind heuristics. By mapping both the successes and current limitations of MLLMs in physical grounding, this study defines the frontier of foundation models as semantic assistants in automated simulation workflows.
PaperID: 7245, Poster
Authors: Xin Zhao, Pengzhan Zhou, Ziyi Li, Riheng Jia, Zhida Qin, Yu Liu
Abstract: The complex interactive dependencies between human behavior and 3D scenes make accurately predicting human intent a major challenge for human motion prediction (HMP). While recent studies use human gaze to infer intent, gaze signals suffer from ambiguity, drift, and a reliance on eye-tracking hardware. Research in cognitive science indicates that egocentric views inherently encode potential interactions, while head motion serves as a crucial precursor in motion planning. Motivated by this, we propose EgoHMP, a novel framework that uses egocentric views and head motions as robust carriers of interaction intent. By translating this intent into an envisioning of future interactions, EgoHMP achieves precise HMP in 3D scenes. EgoHMP first employs a vision-motion alignment encoder to align the visual and motion feature spaces. Based on these aligned features, an egocentric intent predictor with a modality-aware modulation mechanism adaptively suppresses invalid interaction cues to predict an accurate interaction probability map. Furthermore, to enhance the stability and realism of the generated motions, we propose a contact-aware motion decoder. Specifically, inspired by the human cognitive process of planning trajectories prior to action execution, we first predict a global trajectory to serve as a self-prompt, effectively mitigating the optimization conflict in diffusion models. Subsequently, we introduce a multi-stage denoiser, which relieves motion artifacts by incorporating human contact modeling through a multi-stage optimization strategy. Finally, a GCN-based motion decoder is employed to synthesize physically plausible and semantically consistent motions. Extensive experiments demonstrate that EgoHMP achieves state-of-the-art performance.
Authors: Ming-Yang Ho, Alberto Bartesaghi
Abstract: Open-set 3D macromolecule detection in cryogenic electron tomography eliminates the need for target-specific model retraining. However, strict VRAM constraints prohibit processing an entire 3D tomogram, forcing current methods to rely on slow sliding-window inference over extracted subvolumes. To overcome this, we propose FullTilt, an end-to-end framework that redefines 3D detection by operating directly on aligned 2D tilt-series. Because a tilt-series contains significantly fewer images than slices in a reconstructed tomogram, FullTilt eliminates redundant volumetric computation, accelerating inference by orders of magnitude. To process the entire tilt-series simultaneously, we introduce a tilt-series encoder to efficiently fuse cross-view information. We further propose a multiclass visual prompt encoder for flexible prompting, a tilt-aware query initializer to effectively anchor 3D queries, and an auxiliary geometric primitives module to enhance the model's understanding of multi-view geometry while improving robustness to adverse imaging artifacts. Extensive evaluations on three real-world datasets demonstrate that FullTilt achieves state-of-the-art zero-shot performance while drastically reducing runtime and VRAM requirements, paving the way for rapid, large-scale visual proteomics analysis. All code and data will be publicly available upon publication.
PaperID: 7247, Poster
Authors: Yuxuan Jiang, Haichuan Dong, Manning Wang
Abstract: Pathological whole slide images (WSIs) from different hospitals often exhibit severe domain shifts due to variations in scanning devices, staining protocols, and tissue preparation procedures. This leads to significant performance degradation when a trained model is applied to data from different sources. Although unsupervised domain adaptation (UDA) has achieved strong performance in transferring a model trained on a source domain to an unlabeled target domain in natural images, it is much less studied in the computational pathology field. This is due to the unique properties of WSIs, including ultra-high resolution and strong inter-slide and intra-slide heterogeneity. As a result, existing UDA methods are often unstable and may even cause negative transfer in WSI classification. To address this issue, we propose Dual-level Adversarial learning with domain-aware REgularization (DARE) for WSI Classification. Specifically, we design an adaptive pseudo-labeling framework with a teacher-student architecture to provide stable supervision for the target domain and perform adversarial alignment on both patch-level and slide-level embeddings. This enables joint optimization of patch-level and slide-level semantic embeddings for dual-level alignment across domains. We also introduce domain-aware attention masking to regularize attention and improve cross-domain generalization. We evaluate our method on four public WSI datasets through extensive experiments. The results show that our method consistently outperforms state-of-the-art approaches across various domain shifts.
Authors:
Juan A Duque, Sergio García-Heredia, Vinicius Hernandes, Eliska Greplova, Thomas Spriggs, Aaron Courville, Anna DawidAbstract: Neural quantum states (NQS) provides a flexible and scalable framework for approximating quantum many-body wavefunctions. Among NQS parameterizations, autoregressive models are especially attractive because they enable exact, independent sampling from the Born distribution, avoiding the autocorrelation and mixing issues of Markov-chain methods. Yet their optimization remains comparatively underexplored: Adam is a scalable method but ignores function space geometry, while stochastic reconfiguration is principled but costly and numerically fragile in large models. To address this gap, we show that variational energy minimization can be viewed as an advantage policy-gradient problem over the Born distribution, motivating trust-region optimization for NQS training. We introduce Proximal Wavefunction Optimization (PWO), a trust-region algorithm that clips probability-ratio changes in the amplitude channel and wrapped phase increments in the phase channel. PWO avoids explicit matrix inversion, reuses samples across inner updates, and preserves the scalability of first-order optimization. Across Ising, Heisenberg, and frustrated J_1--J_2 spin chains, PWO improves stability and wall-clock convergence over Adam, minSR, and SPRING. Finally, we fine-tune a 1.5B-parameter RWKV-7 model as a neural quantum state, demonstrating NQS optimization at a scale over three orders of magnitude beyond prior work.
PaperID: 7249, Poster
Authors: Vishal V Saley, Bhavesh Gurnani, Dinesh Raghu, Mausam
Abstract: Patient history taking is a strategic, multi-turn process that refines diagnostic hypotheses through iterative questioning, mirroring clinical reasoning where each inquiry updates beliefs over candidate diagnoses. Deploying AI in this setting is constrained by privacy, cost, and latency, making Small Language Models (SLMs) a more desirable backbone compared to larger LLMs. However, SLMs capture only the fast, intuitive "System 1" side of this process and lack explicit mechanisms for the deliberative "System 2'' reasoning needed to maintain and update diagnostic uncertainty over time. Bayesian Networks (BayesNets) offer a natural complement, providing an interpretable framework for auditable probabilistic belief tracking. Yet, clinical BayesNets are difficult to construct due to reliance on expert-curated structure, dependencies, and parameters. We introduce ProbMedTOD, that automatically synthesizes a clinical BayesNet from a small number of clinical notes in three stages: ontology building, structure prediction, and parameter estimation. ProbMedTOD then integrates the resulting BayesNet with SLMs enabling probabilistic reasoning in multi-turn dialogues. A BayesNet Agent maintains real-time posterior distributions over diagnoses, while a Policy Agent picks the diagnostically informative questions. Experiments on DDXPlus and MIMIC-IV Notes show that ProbMedTOD outperforms LLM-only and retrieval-based baselines, with especially strong gains for smaller models.
PaperID: 7250, Poster
Authors: Zhaowei Liu, zengyang zhang, HaitaoYang, Yao Shan, Xiangfu Zhao, Yongchao Song
Abstract: Fraudulent behaviors in Ethereum, such as Ponzi schemes and phishing attacks, pose persistent threats to blockchain security. Classifying high-risk accounts is crucial for mitigating these threats. Given the graph-structured nature of Ethereum transactions, Graph Neural Networks (GNNs) are widely adopted for this task. However, conventional GNNs rely on local aggregation and treat all edges equally, which amplifies structural noise and limits the ability to capture weak but semantically meaningful connections. This limitation arises from the inability to distinguish informative links from spurious ones during message passing, especially in noisy and long-range transaction graphs. To address this, we propose RAMA, a Resistance-Aware Multi-hop Aggregation framework that reweights edges using resistance distance to reduce the influence of noisy links while preserving important topological information. It also employs attention-based multi-hop aggregation to effectively integrate both local and global transaction patterns. Experiments on real-world Ethereum datasets show that RAMA consistently outperforms strong baselines in Macro-F1 and remains robust under edge perturbations, demonstrating its effectiveness for practical account classification in blockchain environments.
PaperID: 7251, Poster
Abstract: Multi-step LLM systems increasingly combine generation, retrieval, tool use, code execution, and verification into pipelines whose components consume and transform an evolving semantic state; end-to-end reliability is therefore governed by pipeline dynamics rather than by component-level accuracy alone. In such pipelines, an early ambiguity, hallucination, or reasoning error may be amplified, masked, partially corrected, or reintroduced by later components. Existing evaluations and failure taxonomies identify where LLM pipelines fail, but they do not provide a predictive theory of how semantic corruption propagates across dependent steps or how limited interventions should be allocated. To fill this gap, we introduce a latent-load framework for reliability analysis and intervention design in black-box LLM pipelines. Specifically, each intermediate state carries an unobserved semantic error load that evolves through inherited-error amplification, context dependence, and fresh error injection. The resulting recurrence decomposes final error, identifies stable, critical, and unstable regimes, and quantifies system-level risk under parameter uncertainty. We prove that the checkpoint-placement objective is monotone submodular, yielding near-optimal budgeted interventions, and provide a proxy-based partial-identification procedure for black-box estimation. Experiments show that proxy-scale propagation estimates predict relative downstream risk and select checkpoints that reduce accumulated proxy load and directionally lower failure rates under fixed budgets.
PaperID: 7252, Poster
Authors: Tom Rossa, Angus Phillips, Thomas Rainforth
Abstract: Bayesian experimental design (BED) has traditionally been based on maximising expected uncertainty reductions from prior to posterior. A major shortfall of this approach is that it leads to doubly intractable objectives that are difficult to optimise, while customising them to particular downstream tasks of interest can also be difficult. Following first principles decision theory, we demonstrate that BED can alternatively be formulated in terms of an expected future loss (EFL) on downstream actions, providing a simple and naturally task-driven framework. Critically, we then show that all such EFLs can be rearranged into using stochastic gradient descent methods. This formulation further sidesteps the need for any explicit posterior or marginal likelihood estimation and is naturally implicit, requiring only the ability to sample from the joint model over model parameters and data, and evaluate the downstream loss function. It thus allows design policies to be learned more effectively, efficiently, and simply than existing methods, while providing easy customisation to different downstream tasks and losses.
Abstract: The process of discovery requires active exploration---the act of collecting new and informative data. However, efficient autonomous exploration remains a major unsolved problem. The dominant paradigm addresses this challenge by using Reinforcement Learning (RL) to train agents with intrinsic motivation, maximizing a composite objective of extrinsic and intrinsic rewards. We suggest that this approach incurs unnecessary overhead: while policy optimization is necessary for precise task execution, employing such machinery solely to expand state coverage may be inefficient. In this paper, we propose a new approach that explicitly decouples exploration from policy optimization and bypasses RL entirely during the exploration phase. Our method uses a tree-search strategy inspired by the Go-With-The-Winner algorithm, paired with a measure of uncertainty to systematically drive exploration. By removing the overhead of policy optimization, our approach explores an order of magnitude more efficiently than standard intrinsic motivation baselines on hard exploration benchmarks. Further, we demonstrate that the trajectories discovered during exploration can be distilled into deployable policies using existing supervised backward learning algorithms, achieving state-of-the-art performance by a wide margin on Montezuma's Revenge, Pitfall!, and Venture without relying on domain-specific knowledge. Finally, we demonstrate the generality of our framework in high-dimensional continuous action spaces by solving the MuJoCo Adroit dexterous manipulation and AntMaze tasks in a \emphsparse-reward setting, directly from image observations and without expert demonstrations or offline datasets. To the best of our knowledge, this has not been achieved before.
PaperID: 7254, Poster
Authors: Fei Zuo, Yizhou Huang, Kezhi Wang
Abstract: Low-Rank Adaptation (LoRA) has emerged as the dominant paradigm for parameter-efficient fine-tuning of large language models. However, existing methods uniformly apply adaptation across all transformer layers, overlooking the heterogeneous contributions of different components. In this work, we conduct a systematic mechanistic interpretability analysis to understand how LoRA affects model internals. We observe that LoRA's adaptation effect is highly non-uniform, with significant component-wise and layer-wise sparsity. Most critically, we find that important layers are highly task-dependent, with limited overlap across tasks. Motivated by these insights, we propose Adaptive-LoRA, a calibration-driven method that automatically identifies task-specific layers via path patching, enabling efficient layer selection without exhaustive search. Extensive experiments on Qwen2.5-1.5B across 11 diverse reasoning tasks demonstrate that Adaptive-LoRA outperforms full-layer LoRA while using only 30% of layers, achieving 3.5× parameter efficiency improvement. This work bridges mechanistic interpretability and practical efficiency, providing both fundamental insights and a principled approach for adaptive fine-tuning.
PaperID: 7255, Poster
Abstract: Long-horizon planning often requires composing familiar local transitions while remaining grounded in a specified goal. We study this setting from single-stage data: on a controlled no-skip graph benchmark, training provides only adjacent-stage demonstrations while evaluation requires composing them on unseen long-horizon start–goal pairs; an event-chain diagnostic decomposes success into legality, stage-bridging, and goal-conditioned termination. Across a minimal Transformer, Qwen2.5-3B, and an external validation on MuSiQue multi-hop QA, supervised fine-tuning (SFT) learns local transitions but degrades with horizon, with failures localized to stage-bridging. Reinforcement learning (RL) narrows this composition gap, especially at longer horizons, but unanchored RL drifts—producing valid stage progress while ignoring the instructed target. Stable gains require two complementary forms of anchoring: instruction-style prompting strengthens prompt-side goal-conditioning, especially at shorter horizons, whereas KL regularization to the SFT reference policy prevents RL-induced drift across horizons; reward shaping mainly provides denser credit. Finally, we identify a Decomposition Paradox: models trained only on long compositions can reach intermediate targets yet fail the corresponding prefix subtasks because stopping at instructed intermediates is never supervised.
PaperID: 7256, Poster
Abstract: Disentangling writing style from semantic content is a fundamental challenge in literary text modeling. Style-content entanglement causes models to rely on semantic content for authorship attribution (AA) tasks and memorize author-specific content for style imitation (SI) tasks. We propose EDITORS (Editing LoRA Subspaces), a diagnose-then-deflate framework that adapts pretrained large language models and prioritizes stylistic features, making it effective for both AA and SI tasks. EDITORS first trains a diagnostic LoRA adapter using style-neutral statements derived from the training corpus to identify a content-orientedsubspace. It then trains the final style adapter under activation-space deflation that projects input activations away from the identified content-oriented directions. We demonstrate that EDITORS is empirically effective: On three AA benchmarks spanning literature, social media, and news, EDITORS achieves state-of-the-art accuracy, outperforming the strongest existing baseline by up to 29%. On SI, EDITORS generates stylistically faithful and diverse passages, outperforming baselines in achieving style similarity while adhering to content specifications.
PaperID: 7257, Poster
Authors:
Haotian Gan, Yudong Li, Xintian Li, Jiarui Han, Weidong TangAbstract: Medical vision-language models often exhibit hidden stratification, where high overall performance masks failures on subtle visual findings and overlapping clinical conditions. We trace this failure mode to modality imbalance during fine-tuning, where strong linguistic priors can suppress visually grounded learning and induce plausible but weakly grounded rationales. Existing data selection methods usually score image-text pairs as holistic units, without separating visually necessary samples from shortcut-solvable ones. To address this problem, we propose ViCoR, a counterfactual-residual-based data selection method for Med-VLM fine-tuning. ViCoR estimates visual necessity through image-dependent optimization residuals, calibrates the resulting scores with medical supervision compatibility, and constructs a diversity-aware fine-tuning subset. Across two backbones, four external generalization benchmarks, and one multi-disease robustness benchmark, ViCoR-selected 20% subsets achieve the strongest average performance compared with full-data fine-tuning and selection baselines. Gradient-level analysis and perturbation audits further show that ViCoR enriches image-dependent training signals and strengthens image-level and evidence-region dependence. The code will be released to support future research.
Authors:
Carmel Baharav, Niclas Boehmer, Bailey Flanigan, Maximilian T WittmannAbstract: AI alignment and participatory design motivate a new democratic design problem: how to collectively choose a _decision rule to use repeatedly_. We study this problem for _linear ranking rules_, which repeatedly rank items x_j within batches X=(x_1,\dots,x_m)\in(\mathbbR^d)^m, where each item's ranking is dictated by its score \langle \theta^\ast,x_j\rangle according to a fixed scoring vector \theta^\ast. Given voters' preferred scoring vectors \theta^(1),\dots,\theta^(n) and their population fractions \alpha^(1),\dots,\alpha^(n), we ask how to choose a collective vector \theta^\ast satisfying _individual proportionality (IP)_: every voter type i should agree with the resulting rankings to an \alpha^(i)-proportional degree, either on average over time (_long-run IP_) or even within each batch (_per-batch IP_). The default rule, the arithmetic mean of the \theta^(i), has been shown to be severely majoritarian; more generally, it is not clear that _any_ fixed linear rule can balance many voters' disparate opinions. Our main result is that, surprisingly, there _is_ a simple rule that does satisfy long-run IP: the _angular mean_, the spherical analog of the arithmetic mean. We then show that exact per-batch IP is impossible for fixed linear rules, but that the gap between per-batch and long-run IP shrinks quickly with batch size. Experiments on three real-world preference datasets show that all rules perform similarly when voters' preferences are homogeneous, while the angular mean substantially improves proportionality in high-disagreement regimes.
PaperID: 7259, Poster
Authors: Chang Liu
Abstract: In cooperative multi-agent reinforcement learning under partial observability, agents typically differ in both their local observations and in how much value those observations let them recover. Most existing methods treat agents symmetrically when shaping credit and exploration, ignoring this asymmetric information disadvantage. We introduce the per-agent value gap, a state- and agent-specific quantity that measures the value loss attributable to local information limits, and show that it serves as a single, measurable signal that unifies credit assignment and exploration: down-weighting policy updates where an agent is at a strong disadvantage, and rewarding behaviors that reduce the disadvantage along the trajectory. We instantiate this idea as \method, which estimates the gap online with a per-agent local value head and uses it to drive both advantage shaping and an information-seeking intrinsic reward. We further show that the resulting shaping reduces the variance of the policy-gradient estimate at high-disadvantage states under a mild correlation assumption, and that the intrinsic bonus vanishes in expectation under any stationary optimal policy, so the augmented objective remains consistent with the original cooperative objective. On sixteen tasks across five cooperative benchmarks, \method ranks first or statistically tied for first on fifteen of sixteen tasks (strictly first on twelve), with the largest gains on settings combining sparse rewards and asymmetric observability (e.g., a 24% mean relative gain over the strongest baseline averaged across the three SMAC v2 tasks).
PaperID: 7260, Poster
Abstract: Missing value imputation is a fundamental task in machine learning, with most existing methods assuming that all missing entries correspond to unobserved regular values. In many real-world datasets, however, missingness may arise from two distinct sources: some entries are meaningfully missing (intrinsically absent and semantically valid), while others are missing due to the observation process and should be imputed. We formalize this distinction as a selective imputation problem, where the goal is to jointly infer which missing entries should be preserved and which should be recovered. To address this challenge, we propose Diff-Joint, a diffusion-based framework that jointly models tabular data together with a latent missingness mask. The method alternates between conditional sampling and uncertainty-aware aggregation to iteratively refine both imputed values and missingness labels. Empirical results on synthetic and real-world datasets demonstrate that Diff-Joint effectively identifies meaningfully missing entries while achieving competitive imputation accuracy and improved downstream task performance.
Abstract: Reward modeling is not only a prediction problem: in KL-regularized policy optimization, the learned reward is exponentiated to define the deployed policy, so downstream value depends on errors in reward-tilted regions. We study this feedback in a Gaussian single-index model with r^ (x)=\sigma^ (\langle\theta^ ,x\rangle) and x~\mathcal N(0,I_d). We analyze a two-stage neural reward model that first learns the hidden direction \theta^ from reward-weighted samples and then fits the readout layer by weighted ridge regression. Exponential reward weighting changes the Hermite signal available to the first layer; for any feature-learning temperature \beta_1 above a dimension-free O(1) threshold, a constant fraction of neurons recover the hidden direction, with weak-recovery complexity governed by the generative exponent. After feature recovery, we derive tilted-policy value-gap bounds for an idealized label-weighted fit with weights e^y/\beta_2 and a more practical surrogate-weighted fit with weights e^r_a_0(x)/\beta_2. Keeping the \beta_2-dependence explicit yields an admissible set of deployment temperatures, balancing the gain from lowering \beta_2 against the learning cost amplified by exponential weighting; in the surrogate-weighted case, proxy-dependent factors shrink this admissible set.
PaperID: 7262, Poster
Authors: Jingyi Pan, Dan Xu, Qiong Luo
Abstract: 3D scene inpainting aims to recover missing or occluded regions in edited 3D scenes, while ensuring geometric and textural consistency. Existing approaches, however, typically require accurately calibrated camera poses, which restricts their applicability in casual, in-the-wild scenarios and introduces additional preprocessing overhead. To overcome this limitation, we present FreeInpaint, a novel feed-forward framework that generates complete and 3D-consistent scenes directly from unposed multi-view images with masked regions. At its core, FreeInpaint extends a 3D foundation model to propagate masked regions from a reference view to other unposed views, bridging 3D reconstruction and scene inpainting while preserving the model's native ability to recover camera poses and scene geometry. Our method addresses two key challenges in adapting feed-forward 3D foundation models to masked inputs. First, masked regions can corrupt cross-view correspondence reasoning, degrading pose estimation and geometry recovery. To address this, we introduce a Learnable Mask Attention mechanism that preserves the spatial anchoring of reliable observations while allowing masked regions to progressively absorb useful context in deeper layers. Second, under severe occlusions, a single forward pass often lacks sufficient appearance evidence for high-fidelity completion. Therefore, we propose a Support Token Refinement strategy, which injects diffusion-generated support evidence as confidence-weighted auxiliary tokens to refine under-observed regions while preserving the original spatial anchor. Extensive experiments across diverse datasets demonstrate that FreeInpaint achieves superior inpainting quality, eliminating the reliance on pre-computed camera poses while keeping a fast inference speed.
PaperID: 7263, Poster
Abstract: Existing unified 4D reconstruction and point tracking approaches typically rely on heuristic interpolations or just predict at integer timestamps, lacking kinematic coherence and failing to model dynamics at any arbitrary timestamp. In this paper, we propose Uni4R, a framework that unifies these tasks by learning continuous velocity fields through the synergy of Optimal Transport (OT) and Ordinary Differential Equation (ODE). Importantly, this continuous velocity field acts as a kinematic prior that mutually benefits both 4D reconstruction and point tracking. Specifically, we propose the Flow Matching Guided Decoder (FMGD). A global velocity branch first extracts anchor features that capture the global dynamic state of the sequence. Then, FMGD leverages Flow Matching (FM) theory to formulate a probability path defined by OT on the anchor feature manifold, instantiating it as FM-guided velocity features for velocity prediction. This establishes a robust kinematic inductive bias. Meanwhile, a point reconstruction branch provides geometric features. The local velocity prediction module then joint above features and time embeddings, to decode velocities at arbitrary timestamps. To overcome the absence of high-quality ground-truth velocities in fractional frames, we propose an integral-consistency training strategy. This strategy uses an ODE solver to integrate velocities to recover target pointmaps, enabling the model to be supervised end-to-end directly from integer timestamps. Experimental results demonstrate that Uni4R achieves SOTA performance in both 4D reconstruction and point tracking, and achieves SOTA in our new kinematics-aware benchmark at continuous time.
PaperID: 7264, Poster
Abstract: LLM unlearning is often implemented by assigning surrogate targets, such as refusal, uniformity, or likelihood suppression, to forget-related prompts. We argue that this surrogate assignment view is misaligned with the counterfactual goal of unlearning: recovering the retain-only predictive posterior. This mismatch is especially problematic in mixed-evidence regimes, where a response may be associated with forgotten evidence while still being partially supported by retained evidence. We propose EASE: Evidence Attribution and Subtraction Estimator, a logit-space posterior-correction method that estimates the forget-induced component of the deployed model's predictive support and subtracts it at inference time. EASE uses lightweight deletion and compensation assistants to remove forget-neighbour evidence while restoring nearby retain-supported evidence. We formalize the suppression-scope dilemma and provide guarantees showing when posterior correction recovers the retain-only posterior. Experiments on TOFU and MUSE show that EASE improves forgetting--retention trade-offs over strong baselines, achieving the best aggregate score in 5/6 main TOFU settings, reducing mixed-query semantic leakage by 22.2%, and matching or exceeding full-retain baselines with 9 times fewer retain examples. Our code is available at: https://anonymous.4open.science/r/EASE-9675.
PaperID: 7265, Poster
Abstract: Current foundation models of electroencephalography (EEG) sought to characterize the fluctuations in neural signals, yet this objective stands fundamentally at odds with the intrinsic physiological mechanisms governing brain activity. As with their LLM predecessors, these models pack millions of neurons, even those contributing negligibly small weights, yielding a computational footprint incompatible with resource-constrained portable brain-computer interfaces (BCIs). To overcome these challenges, we propose , an energy-preserving and biologically-realistic EEG foundation model operating in the regime of oscillatory neural synchronization. Drawing on principles from cognitive neuroscience, our foundation model is designed to recover the masked EEG signals through the transient spikes of spatially coupled neurons, from which self-organized cortical synchronization (aka. ) emerges to reciprocally modulate ongoing neural firings. Under the hood, we introduce a macro-micro learning mechanism that bridges two scales of neural computation: lower layers capture neural oscillations via a spiking neural network (SNN), while upper layers model oscillatory synchronization across large neuronal populations, with cross-scale interactions learned end-to-end. Taken together, our sheds new light on learning intrinsic representation of functional neural fluctuations. We pre-train the model to reconstruct segmented EEG signals directly from the synchronized spiking representations using a purely self-supervised objective on a diverse corpus of public datasets. To demonstrate its universal representational power, we conduct massive-scale evaluations across ten heterogeneous downstream datasets, including clinical pathology, sleep staging, and cognitive assessment. not only facilitates deployment on portable edge devices through its lightweight architecture and energy-preserving inference, but also provides profound neuroscience insights. By further capturing pre- and post-synaptic firing, unlocks a new pathway for characterizing effective connectivity and the oscillatory synchronization that underpins cognition.
Abstract: Post-training quantization (PTQ) compresses large language models to low bit-widths using a small calibration set, making calibration selection a consequential but under-specified part of the quantization pipeline. We study calibration selection for weight-only LLM PTQ backends such as AWQ and GPTQ. In this setting, calibration data does not define an inference-time activation clipping threshold; instead, it determines activation-conditioned quantities used for weight reconstruction, saliency estimation, and scale search. We identify a failure mode in which a calibration set misses rare, high-magnitude input channels of quantized linear modules, causing the backend to underrepresent reconstruction-sensitive directions. Motivated by this observation, we formulate calibration selection as weighted maximum coverage over module-level activation outlier channels. Each candidate sequence covers the module/channel pairs it activates above a reference threshold, and each covered pair is weighted by a dimensionally consistent surrogate for its potential contribution to squared output reconstruction error. The resulting objective is monotone submodular, so greedy selection admits the classical (1-1/e) approximation guarantee for the coverage objective. We instantiate this formulation in COVERCAL, a backend-agnostic calibration selector operating on cached activation masks. Across LLaMA-2, LLaMA-3, and Mistral models, under AWQ and GPTQ INT4 backends and five downstream evaluations, COVERCAL improves over the reported fixed-pool calibration baselines, with the largest gains at small calibration budgets. Under AWQ at K=128, COVERCAL improves MMLU by 1.2--1.5 points over random calibration; at K=64, it matches or exceeds random calibration at K=256 on LLaMA-3-8B. The contribution is not a new quantizer, but a module-correct coverage formulation, an efficient greedy selector, and empirical evidence that weighted outlier coverage is a useful calibration signal for existing weight-only PTQ backends.
PaperID: 7267, Poster
Abstract: Solubility depends on intermolecular interactions between solutes and solvents and is central to drug discovery, chemical and materials process engineering. Its prediction requires out-of-distribution (OOD) generalization to unseen solutes, solvents, and solvent mixtures, where supervised methods suffer from data scarcity, and existing self-supervised pretraining methods do not explicitly learn chemically meaningful relations. Learning chemically meaningful relational information without labels can alleviate data scarcity and improve OOD generalizations, but remains challenging. We introduce Prior Relational Information for Self-supervised Modeling (PRISM) and apply it to solubility OOD generalizations. PRISM is a label-free pretraining framework that converts expert-defined chemical priors into relational teachers. For solubility prediction, it trains four encoders to preserve the population-level relational information from priors capturing hydrophobic, electrostatic, polarizability, and substructural effects, using Kullback-Leibler (KL) divergence to match pairwise distance-based distributions during pretraining. For scaffold, solute, and solvent OOD generalizations, PRISM matches or surpasses supervised and self-supervised baselines. For mixed-solvent OOD generalizations, freezing the four PRISM encoders and training a light fusion-and-composition head outperforms the strongest supervised baselines and greatly reduces prediction variations. Ablations and control studies confirm that gains stem from chemically meaningful relational structures baked into the encoders by PRISM. We expect this strategy of turning relational priors into self-supervision signals to extend beyond solubility prediction to domains where chemically or biologically meaningful relational priors are available.
PaperID: 7268, Poster
Abstract: Recovering dynamic 3D scene geometry, appearance, and the physical motion states that govern scene evolution from multi-view videos is important yet challenging. A key difficulty is that appearance can be deceptive. When trained primarily with rendering supervision, existing methods may fit the observed appearance well by exploiting appearance shortcuts, rather than recovering the true evolving geometry and motion. This issue is especially severe in scenes with fast or locally complex motion, where inaccurate trajectories can still produce plausible images over the observed frames, but lead to unreliable reconstruction and discontinuous future prediction. In this paper, we propose PhyMo, a framework for learning evolving dynamics from multi-view videos. We introduce a Motion-Aware Trajectory Representation that drives the model to recover accurate motion from deformation, rather than relying on appearance shortcuts to explain observed motion. On top of this, we propose a Continuous Motion Evolution module with bidirectional kinematic continuity constraint to promote temporally consistent evolution of displacement, velocity, and acceleration. Together, these designs yield geometrically faithful interpolation and continuous motion extrapolation. Experiments on four dynamic scene benchmarks demonstrate state-of-the-art performance.
PaperID: 7269, Poster
Abstract: Vision-language-action (VLA) policies increasingly generate natural-language rationales before executing embodied actions. These rationales are attractive as monitors for safety-critical behavior because users or automated checkers can inspect whether the model appears to understand the scene and intend a safe action before execution. We show that this signal can be misleading. In many reasoning VLAs, the rationale generator and action head share continuous hidden states, while the generated text reveals only part of that representation. This creates a Say the Same, Act Differently failure mode, where rationale text remains stable while the action output changes substantially. We first study this phenomenon by introducing a new statistical method to separate text-sensitive and action-sensitive directions in hidden-state space of a reasoning VLA. Across the driving VLA Alpamayo and the manipulation VLA InstructVLA, we show that the dominant action direction lies largely outside the active text span, with mean projection ratios of only 1.9% and 5.3%. We define the residual as a text-neutral action-subspace (TNAS) direction, which tests whether action-relevant directions remain after removing measured text-sensitive directions. We then use TNAS to guide white-box pixel-space projected-gradient-descent (PGD) attacks on Alpamayo. TNAS-guided PGD successfully uses a localized two-forward-camera patch to induce a 30.5 m mean ADE shift in trajectory planning while preserving the exact rationale in 86.7% of cases. Moreover, across multiple text-preserving PGD objectives, the optimization direction consistently aligns with TNAS. These results reveal a fundamental shared-cache monitorability gap in reasoning VLAs: stable rationale text should not be treated as sufficient evidence that the action output remains stable.
PaperID: 7270, Poster
Authors: Martin Carrasco, Caio Deberaldini Netto, Ehimare Okoyomon, Aneeqa Mehrab, Vahan Martirosyan, Caterina Graziani
Abstract: Understanding the interplay between generalization, expressivity, and the geometry of the input space is a central challenge in graph learning. The expressivity of Graph Neural Networks (GNNs) is typically characterized through their correspondence with graph invariants, such as those from the Weisfeiler-Leman (WL) hierarchy. While more expressive GNNs can distinguish a richer set of graphs, they are also associated with weaker generalization guarantees. Previous works have addressed this trade-off using the VC dimension, a purely combinatorial measure, independent of the training data. In this work, we adopt a data-dependent measure of generalization, the empirical Rademacher complexity, and derive tight generalization bounds that jointly consider the expressive power of GNNs and the geometry of the underlying input space. Specifically, any graph invariant that upper-bounds a GNN's expressive power partitions the input space into equivalence classes, and we show that the empirical Rademacher complexity is controlled by the distribution of training samples across these classes. Moving beyond discrete partitions, we incorporate the geometry of the input space and derive covering-number bounds under Lipschitz continuity, showing that the complexity cost can be mitigated when the hypothesis class remains smooth over the data geometry. In addition, we prove that the empirical Rademacher complexity is Lipschitz continuous with respect to the Wasserstein distance between empirical measures supported on different datasets. This yields robustness and generalization guarantees under sampling variability. Importantly, our framework is not restricted to message-passing GNNs or WL, but extends to arbitrary GNN architectures and their associated invariants, providing a step toward a unified theory of GNN generalization.
PaperID: 7271, Poster
Authors:
Guo Yue, Yang Liu, Donghui Zhang, Tang Qingkang, Rong Fu, Li AoyuAbstract: Modern LLM agents repeatedly fall into a familiar trap: they execute irreversible actions whose consequences they did not anticipate, such as confirming a non-refundable booking or sending an unrecoverable message. World models are a natural remedy, yet two dominant lines of work both fail this use case. Generative pixel and HTML world models pay a heavy compute cost to produce vivid but often unfaithful continuations of the full observation, while reward and value models only summarise how good an action is and tell the agent nothing about what it will change or whether the change can be undone. We argue that what an agent really needs is a model of what changes, not of what comes next. We introduce Causal Delta World Models (CDWM), which represent each transition as a typed structured delta over an entity, attribute and relation graph, and explicitly factor out reversibility, irreversible cost and prediction confidence. A learnable causal sparsity mask binds each action class to the delta dimensions it can plausibly affect, providing a strong inductive bias for unseen actions and environments. We post-train CDWM with a verifiable composite reward that aligns delta accuracy, action following, calibration error and closed-loop task success. The resulting world model is consumed by a risk-aware closed-loop planner that prunes catastrophic candidates, performs short-horizon model-predictive control, and abstains when confidence is low. On four interactive benchmarks (WebArena, Mind2Web, ALFWorld, ScienceWorld) and two visual control suites (Procgen and Atari 100k), CDWM trained on four NVIDIA A100 (80 GB) GPUs improves task success by 6.0 to 11.7 absolute points over strong agentic and world model baselines, reduces irreversible failures by 47 to 71 percent, lowers token cost per successful trajectory by 27 to 45 percent, and shows substantially better generalisation to unseen websites and rule perturbations. We conclude that the right unit for an agent's world model is the structured delta and not the full observa
PaperID: 7272, Poster
Abstract: Many language-model tasks reduce to selecting among structured alternatives, yet we still lack a simple picture of how a decoder-only transformer organizes an entire selection prompt across depth. We introduce token covariance maps, a prompt-level geometric view that tracks how token positions co-vary in hidden-state space at each layer. This turns a structured prompt into a depth-indexed sequence of maps whose spectral and block structure can be aligned with instructions, context, and candidate answers. Across 14 decoder-only models from 4 families and 12 structured-selection benchmarks, token covariance maps reveal a robust compress-then-compare organization. Early layers represent prompts diffusely, middle layers compress them into a shared low-rank bottleneck, and late layers re-expand into context-option and option-option geometry. This organization is stable under larger fixed-template samples, deterministic prompt rewrites, pairwise reranking reformulations, and controlled answer-set manipulations, while also emerging during training from an initially flat depth profile. Its timing aligns with behavior: useful answer discrimination appears well after the bottleneck, and low-rank intervention around the bottleneck preferentially disrupts late comparison geometry. Together, these results identify a compact layerwise account of structured selection in decoder-only language models: prompts are first compressed into a shared mid-layer representation and then re-expanded into late answer-comparison geometry. Token covariance maps provide a reusable diagnostic for comparing prompts, checkpoints, and model families, and reveal a reproducible geometric organization underlying compare-and-select behavior.
Abstract: Last-iterate convergence of learning dynamics in games has attracted significant recent attention. In two-player zero-sum games with bandit feedback, where only the loss of the selected action pair is observed, Fiegel et al. [2025] show a separation between average-iterate and last-iterate convergence in duality gap: while the optimal t^-\frac12 rate after t rounds is achievable for the former via standard no-regret algorithms, the latter cannot converge faster than t^-\frac13 in expectation or t^-\frac14 with high probability. However, in many practical settings (such as preference learning), the players observe not only their loss but also the opponent’s action. This raises a natural question: can such additional information enable faster last-iterate convergence? We answer this question affirmatively, showing that t^-\frac12 last-iterate convergence is achievable with high probability in this setting, via an efficient algorithm that updates its strategy infrequently by solving an estimated log-barrier-regularized game. We identify fundamental obstacles preventing standard analysis for multi-armed bandits (the single-player case) from generalizing to games, and develop a novel analysis to overcome them. Experiments confirm that our algorithm indeed converges faster than naive baselines and prior methods that do not exploit opponent-action feedback. Finally, we note that our results also improve those for dueling bandits, a special case with skew-symmetric game matrices.
PaperID: 7274, Poster
Abstract: Long-context inference is increasingly important for privacy-sensitive local LLM deployment. In this setting, the prefill stage dominates Time-to-First-Token (TTFT), yet existing acceleration methods struggle to balance speed and quality. Sparse-attention methods accelerate only the attention module, leaving Feed-Forward Network (FFN) costs unchanged, while auxiliary-model-based compression often compromises quality under tight memory budgets. Recent layer-wise token pruning methods reduce both attention and FFN computation, but they typically rely on heuristic pruning schedules or token-recovery mechanisms that introduce decoding overhead. We propose FastPrefill, a recovery-free layer-wise token pruning system for efficient local long-prompt inference. FastPrefill employs an offline data-driven optimizer to determine layer-wise pruning schedules that maximize fidelity under a latency constraint, and executes these schedules via an algorithm–system co-design supporting head-specific sparse attention at runtime, where each head dynamically attends to its own selected key/value tokens. Evaluations on LongBench and RULER show that FastPrefill maintains comparable inference quality while achieving up to 2.13x TTFT speedup and 3.03x end-to-end speedup over state-of-the-art baselines.
PaperID: 7275, Poster
Abstract: We study how much a linear program (LP) can be compressed when solved repeatedly, given prior knowledge about its objective function. Existing data-driven projection methods learn low-dimensional surrogate LPs with approximate objective-value guarantees, but cannot provably identify the optimal projection for a prescribed compression budget. We instead ask a sharper question: how far can an LP be compressed into a lower-dimensional equivalent while exactly preserving optimality, enabling faster repeated solves with no loss in solution quality? We provide an exact geometric characterization of such compressed LPs, together with a tractable sample-based learning algorithm that comes with fast-rate guarantees: the compressed LP recovers the optimal solution of an unseen instance with probability at least 1-\widetilde O(d^\star/n), where d^\star is the dimension of the decision-relevant subspace, and n the number of available historical LP samples. This 1/n dependence is sharper than the \widetilde O(1/\sqrt n) uniform-convergence rates of approximate projection methods. Our framework further exposes a tunable tradeoff between the dimension of the compressed LP and the probability of recovering the optimal solution, allowing the user to trade compression for accuracy.
PaperID: 7276, Poster
Abstract: We study the existence of uniformly optimal stationary policies, which are optimal for any starting distribution in Markov Decision Processes (MDPs) under epistemic (parameter) uncertainty within two complementary frameworks. For robust MDPs, we establish that (s- and sa-) rectangular models are equivalent to the existence of uniformly optimal stationary policies and provide several interesting counterexamples, including a simple r-rectangular instance with only history-dependent optimal policies. For ambiguity-averse MDPs where the uncertain transition kernel is a random variable, and the return is evaluated via a law-invariant risk measure, we show that uniformly optimal stationary policies can only exist for the essential supremum or essential infimum (under mild monotone continuity assumptions), thus excluding other risk measures such as CVaR, VaR, and the Entropic Risk Measure. We conclude our work by relating dynamic programming and the existence of uniformly optimal policies. One of the main takeaways of our impossibility results is the need for new models to go beyond uniformly optimal stationary policies.
PaperID: 7277, Poster
Abstract: Decoder-only Large View Synthesis Models (LVSMs) with KV-cache have recently achieved state-of-the-art quality by regressing novel views from neural networks without reconstructing 3D geometry, while Test-Time Training (TTT) layers can be introduced to save computation, but sacrifice quality. We propose ), a strategy that learns to bake global scene information from cross-view attention into a lightweight MLP, enabling fast and high-fidelity novel-view decoding. The first key is to repurpose the fast-weight MLP of TTT to learn the cross-view attention mapping explicitly from its input-output pairs, rather than replacing attention entirely, achieving the same rendering efficiency as TTT while delivering substantially higher quality. The second key addresses a modality misalignment between context prefilling and novel-view rendering: we disentangle context tokens into camera-only context rendering tokens to serve as TTB input, and context ground-truth tokens with image content to serve as the TTB mapping target, ensuring the baked MLPs are relieved of the modality generalization challenge between context and novel-view processing. Our method achieves up to
PaperID: 7278, Poster
Abstract: Large language models and their vision counterparts have achieved remarkable capabilities, yet their wider deployment is increasingly constrained by the computational and memory demands of Transformer architectures. Much recent work improves Transformer efficiency, but the query-key pathway is still typically parameterized by fully independent per-head projections, which place a substantial parameter burden on attention. Motivated by the view that token selection may admit a shared low-dimensional geometry, we propose Basis-Mediated Bilinear Attention (BMB), a training-time reparameterization of the query-key pathway in which all heads within a layer share a latent basis while retaining head-specific interactions. A practical explicit variant, BMB-UV, materializes per-head query and key tensors through a shared basis and lightweight head-specific factors, reducing the query-key parameter count from 2d^2 to dr+2Hrs. We evaluate our method family on ViT-Base image classification on ImageNet-1K and on GPT-2 small pretraining on the C4 RealNewsLike subset, and compare it against recent query-key parameter-sharing and low-rank baselines. Relative to standard multi-head attention, our method family reduces query-key parameters by up to 95.83 % while remaining competitive with the baseline on both benchmarks.
Abstract: Preference optimization is a standard alignment method for generative models, yet extending it to continuous-time dynamics remains non-trivial. In flow matching, reward-driven updates modify the transport trajectories but lack inherent constraints to pretrained data manifold, often pushing generated (terminal) samples off the pretrained support. We formalize this failure mode as , a phenomenon where preference-aligned trajectories diverge from the pretrained support. Theoretically, we prove that while optimal flow matching exactly preserves the terminal manifold, existing methods like FlowDPO fail to do so whenever reward-driven updates contain components normal to the manifold surface. As a remedy, we propose , a temperature-controlled objective that anchors pairwise preference optimization on preferred samples. Across temperature regimes, this objective connects rejection sampling fine-tuning and FlowDPO while upper-bounding a reconstruction-based surrogate for manifold drift. To counteract diminished signals at low temperatures, we further introduce a weighted variant, , which we evaluate across all experiments for its superior trade-off between preference alignment and manifold preservation. Empirically, demonstrates a better trade-off on synthetic manifolds than FlowDPO variants. On real-world image benchmarks, it improves OCR-oriented alignment over the pretrained model and FlowDPO variants while remaining competitive with strong baselines in held-out model-based and human evaluations without visible sample degradation.
PaperID: 7280, Poster
Abstract: Initialization determines the optimization ``fate'' of a model, going far beyond just setting the scale of its weights. It establishes a fundamental initial geometry over the data, permanently dictating which examples are considered close, which directions are easy to change, and which representations gradient descent will naturally refine first. Because this starting point dictates the model's future trajectory, we ask: can this initial geometry be explicitly chosen by referencing a completely different architecture? To achieve this, we introduce representational similarity as an optimization prior. This is an initialization-only procedure that aligns a target network to the representational geometry of a randomly initialized guide network before any downstream training occurs. Crucially, the guide network transfers no learned knowledge: it is frozen, never sees labels. We demonstrate that this cross-architecture transfer of fate works in two key domains: instilling the optimization benefits of depth into shallow targets by aligning them to deeper, randomly initialized guides, and transferring representational granularity by using fine-resolution guides to shape the initialization geometry of networks built with coarser representations. Ultimately, this establishes that a model's architectural destiny can be decoupled from its physical structure and explicitly programmed at initialization.
PaperID: 7281, Poster
Authors:
Zehao Du, Zheng Huang, Rui Li, Hongyu Liu, Jiude Wei, Jiachun Bao, Cewu Lu, Jianhua SunAbstract: Robot manipulation under partial observability requires policies to act from incomplete visual evidence while still reasoning about object structure, spatial layout, and physical relations. Existing imitation learning methods mainly encode currently visible observations, which makes them brittle when task-critical parts, contact regions, or support relations are occluded. We introduce Scene Map, a structured scene representation that models object shape, part structure, spatial layout, and physical constraints as a graph of manipulation-relevant primitives and relations. Scene Map turns occluded but task-relevant structure into explicit prediction targets and provides structural priors for inferring unseen components from partial observations. Building on Scene Map, we propose MapPolicy, an imitation learning framework that fuses structured map features with visual and robot state features for action prediction. MapPolicy uses structure-modulated graph attention to propagate global structural and spatial information through primitive features, and a physical constraint loss to regularize learned features with geometric and kinematic consistency. Across 49 manipulation tasks in three simulation benchmarks and five real world tasks, MapPolicy consistently improves over strong imitation learning baselines, with especially clear gains under severe occlusion. Our code will be released publicly to facilitate reproducibility and further research.
Abstract: Tool-augmented language models are bounded by the APIs humans bothered to write; existing tool-creation systems patch this by prompting a frozen LLM at inference time, leaving the model that writes a tool decoupled from the one that uses it, with no signal that the schemas it produces are schemas it can invoke. We propose SMITH (Schema-grounded Multi-task Iterative Tool Honing), a reinforcement learning framework that jointly trains tool creation and tool use inside a single policy. Each rollout is either a build task (write a tool from a few examples) or a use task (invoke a pooled tool on a held-out question). Three separate reward axes catch schema, code, and outcome failures independently, so each failure mode contributes its own gradient. A 4B Qwen3 trained with SMITH on 13 procedural reasoning tasks with exact verifiers reaches 79.8 macro-average accuracy on held-out tasks, the best across all evaluated methods and ahead of an untrained 30B-A3B tool-writer. It also reaches 40.4 on TabMWP-Hard and 42.6 on out-of-domain GQA (+7.6 over the best same-backbone inference-time baseline), without any visual or tabular training data. When invoked by a frozen 350M student, tools written by our 4B match those produced by a writer an order of magnitude larger. The same recipe also lifts Qwen3-8B and Granite-3.3-8B without modification.
Authors: Jonghyun Shin, Sejun Park
Abstract: We analyze generalization error, uniform stability, and uniform argument stability of gradient descent (GD) and stochastic gradient descent (SGD) over discrete parameter spaces, where each update involves deterministic or stochastic rounding. We show that deterministic rounding degrades the generalization error of GD on convex, Lipschitz, and smooth loss functions, increasing the rate from O(T/n) to O(T/\sqrtn), and establish matching lower bounds. We further prove that uniform stability of GD becomes \Omega(T), showing that stability-based generalization bounds are vacuous in this setting. In contrast, for the same losses, stochastic gradient descent with deterministic rounding admits nontrivial uniform stability guarantees, which differ qualitatively from the real-valued case and exhibit distinct dependencies on the number of iterations and the dimension: we prove tight bounds O(T/n) for one dimension and O(T^2/n) for higher dimensions. We also show that stochastic rounding can introduce generalization error that increases with the dimension; such a phenomenon is absent in standard real-valued optimization and in the deterministic rounding case. Finally, we provide upper bounds on uniform argument stability for stochastic rounding schemes and show that these bounds are tight when the loss can be represented as a sum of coordinate-wise functions.
PaperID: 7284, Poster
Abstract: Protein design generates amino acid sequences first and then translates them into DNA post-hoc. Control over the nucleotide sequence via codon choice at design time is thus limited. We introduce NuCaliby, a structure-conditioned model that jointly designs a protein's amino acid, mRNA, and DNA sequences at nucleotide resolution. It predicts a Potts model over a nucleotide graph, trained using amino acid supervision by marginalizing over synonymous codons. Its energy-based formulation supports composable inference-time guidance from both differentiable and non-differentiable objectives. NuCaliby preserves residue-level designability on par with Caliby while learning codon-aware nucleotide couplings directly from structure. Under organism-specific tRNA adaptation guidance, NuCaliby improves adaptation to host tRNA pools beyond synonymous sequence space while preserving protein designability. NuCaliby also embeds functional RNA motifs, directly into coding sequences, a problem that is combinatorially infeasible with extensive sampling alone. In wet lab experiments NuCaliby produces expressible and soluble proteins on par with state of the art methods, although it enables more complex nucleotide-level control and optimization. Together, these results extend protein design beyond the residue level, enabling joint optimization across the central dogma of molecular biology.
PaperID: 7285, Poster
Abstract: Federated Graph Learning (FGL) on Text-Attributed Graphs often encounters the missing neighbor issue. This problem arises from structural fragmentation across dispersed clients, which restricts the effective message passing of Graph Neural Networks. Current completion methods frequently struggle in cross-domain, non-IID scenarios because they lack access to external knowledge. To address these limitations, we propose FedProG, an asymmetric framework that utilizes a frozen server-side LLM to reconstruct missing structural contexts. Furthermore, deploying Large Language Models locally to augment these graphs incurs prohibitive computational costs for edge devices. We introduce a split architecture where a lightweight client-side generator maps graph features into the LLM’s embedding space. This allows the server to produce virtual neighbors via uploaded soft prompts. To ensure system reliability and privacy, we implement a dual-guard mechanism: a server-side uncertainty filter rejects low-confidence generations through multi-path sampling, while a client-side semantic consistency projector maintains utility under differential privacy. Experiments show that FedProG improves robustness and generalization over baselines while drastically reducing client computational overhead.
PaperID: 7286, Poster
Abstract: Score matching and flow matching exhibit powerful expressive capabilities in continuous control tasks. However, since the log probability densities of their generated actions cannot be directly accessed, they introduce substantial challenges to the implementation and optimization of Maximum-Entropy Reinforcement Learning (RL). To address this critical limitation, the RL community has developed a wide range of diffusion-based RL algorithms, leading to a widespread perception that the integration of diffusion models and RL has achieved remarkable success. Yet, is this optimistic conclusion truly well-founded? In this paper, we propose a systematic taxonomy for existing mainstream diffusion RL methods that consists of three distinct categories, and we conduct in-depth theoretical analyses of their inherent limitations and core defects.Our analysis reveals that existing methods either collapse the intended stochastic generative policy into action maximization, require costly long-horizon sampling in high-dimensional action spaces, or rely heavily on proposal distributions whose coverage fundamentally limits policy improvement. Building on the comprehensive analysis above, we propose a novel method Copuled Flow . The core insight of our method is to leverage linear ordinary differential equations to maintain fully tractable log-likelihood calculation, while adopting coupled generation mechanisms to effectively capture complex distributions. In the experiments, we evaluate the Copuled Flow on 3 HumanoidBench tasks, 5 MuJoCo tasks, and 7 tasks from the DeepMind Control (DMC) Suite. Empirically, Copuled Flow achieves superior performance on high-dimensional benchmark tasks compared to competitive strong baselines. Specifically, it significantly outperforms existing diffusion-based RL methods by 100% on the HumanoidBench benchmark.
PaperID: 7287, Poster
Abstract: How much of zero-shot benchmark accuracy is captured by word-level alignment between a pre-training corpus and the benchmark target? We study this question through word-level unigram cross-entropy (UCE), a tokenizer-agnostic, model-free measure of lexical alignment between a reference corpus and a target benchmark. Likelihood training allocates probability mass according to the pre-training corpus, and likelihood-based zero-shot evaluation partly reads out the lexical alignment between this mass and the benchmark target. Cleanly tracing how the lexical imprint of pre-training data scales with corpus--benchmark alignment requires controlled from-scratch pre-training that varies corpus identity while holding the training recipe fixed. Across this controlled sweep, spanning 11 zero-shot benchmarks, 4 pre-training corpora, and 3 model scales, we find that corpus--benchmark lexical alignment, UCE, consistently tracks zero-shot accuracy on every benchmark, with corpora more aligned with the benchmark (lower UCE) achieving higher accuracy. The same signal extends from measurement to control at finer scopes. At the document level, strengthening the imprint by selecting training documents to maximize lexical alignment with the benchmark yields a data-selection rule that improves accuracy over same-size random sampling and an importance-resampling baseline across sources and scales. At the prompt level, matching the imprint per prompt by routing each prompt to the drafter whose domain corpus is most lexically aligned with it gives a training-free routing rule for speculative decoding and improves throughput over a generalist-drafter baseline. Taken together, pre-training corpora leave a measurable lexical imprint on zero-shot behavior, and UCE provides a model-free probe of lexical alignment between reference and target data.
Abstract: Probabilistic Circuits (PCs) are deep generative models that support exact and efficient probabilistic inference. Yet in autoregressive language modeling, PCs still lag behind Transformer-based large language models (LLMs), suggesting an important expressivity gap. In this work, we compare PCs and LLMs under a unified autoregressive formulation. First, an output bottleneck: PCs parameterize predictions as convex combinations in probability space, which struggles to represent the sharp distributions typical of language; adopting a logit-space parameterization substantially narrows this gap. Second, a context-encoding bottleneck: we prove that structured-decomposable PCs can match Transformer separation rank on vtree-aligned partitions, but show, both theoretically and empirically, that this capacity is limited to partitions aligned with the fixed routing structure, leading to severe degradation when the data exhibits heterogeneous dependency topologies. We further prove that decomposable PCs are strictly more expressive than structured-decomposable ones, though effectively optimizing them remains an open challenge.
PaperID: 7289, Poster
Abstract: Sign Language Translation (SLT) research has typically focused on pre-segmented clips, a constraint that disconnects models from the continuous reality of real-world communication. To bridge the gap from clips to streams, we formally define Streaming SLT, a realistic task demanding the simultaneous discovery and translation of linguistic events from continuous, untrimmed video inputs. To address this challenge, we propose StreamSLST, the first end-to-end framework for streaming Sign Language Segmentation and Translation via Deformable Transformer architecture. Unlike cascaded approaches that suffer from error propagation, StreamSLST employs a parallel decoding strategy to jointly optimize temporal segmentation and translation in a single pass. To enable scalable training on long-form videos, we incorporate a Sentence-aware Sliding Window Sampler that crops continuous streams into trainable units. We further design a Tri-modal Visual-Language Pretraining stage that aligns pose dynamics with textual semantics before joint optimization. To support comprehensive evaluation, we conduct experiments on two real-world streaming datasets, BOBSL and How2Sign, together with two newly constructed synthetic streaming benchmarks, Streaming-CSL-Daily and Streaming-Phoenix-2014T. The results demonstrate that StreamSLST consistently outperforms cascaded segmentation-then-translation pipelines. Interestingly, our analysis reveals that learned sentence boundaries can outperform ground-truth boundaries for translation, suggesting that joint optimization drives localization toward semantically informative regions rather than exact temporal endpoints. Our work establishes the first robust benchmark for continuous sign language understanding and paves the way for accessible real-world communication interfaces.
PaperID: 7290, Poster
Abstract: Multimodal Recommender Systems (MMRS) leverage rich item features to mitigate data sparsity, yet their success largely hinges on the foundational homophily assumption. This paradigm makes existing models adept at recommending similar substitutes but poorly equipped to handle heterophilic complementarity. Consequently, complementary items often become geometrically isolated, leading to fragmented “Similarity Islands” and semantic drift when standard propagation is applied. To overcome these limitations, we propose MAGE (Manifold-Augmented Graph Embedding), a novel framework that decouples complementarity modeling into offline geometric warping and online topological propagation. The offline phase enhances item representations by extracting complement-aware residual directions via a Soft Conceptor mechanism and injecting them under a local gap-preserving constraint, enabling the model to capture functional relationships while reducing the risk of semantic drift. Synergistically, the online phase employs an asymmetric dual-stream GNN that treats visual features as stable anchors and augmented textual features as flexible probes, fused through a self-calibrated cross-modal mechanism to reconcile shared semantics and modality-specific residual cues. Extensive experiments on four Amazon datasets show that MAGE yields consistent improvements over strong multimodal recommendation baselines. Together with the supporting analyses in the appendix, these results suggest that geometry-aware augmentation can alleviate the semantic mismatch that arises in complementary recommendation settings. Our code is available at
PaperID: 7291, Poster
Abstract: Concept unlearning in text-to-image diffusion models aims to suppress a target concept, e.g., \texttthorse, while preserving the model's ability to generate semantically related but distinct content, like \textttdonkey. Yet existing methods either leak under indirect prompts or visibly degrade remaining concepts. \em We show that these failure modes arise naturally from overlapping concept representations. Our main contribution is demonstrating that the observed trade-offs stem from the geometry of concept representations, rather than from weaknesses in any particular algorithm. We support this claim through both theoretical analysis and empirical observations. By formalizing concepts as activation-space regions, we show that overlap between a target and other concepts lower-bounds the unavoidable degradation on those other concepts when the target is erased. Our analysis suggests that strong erasure and preservation become fundamentally coupled when concepts occupy overlapping activation regions, with the trade-off scaling linearly in the degree of overlap. We verify this trade-off on seven unlearning methods spanning fine-tuning, adversarially-robust fine-tuning, and inference-time interventions. Averaged across target concepts, STEREO almost completely suppresses the target concept under indirect prompts but cuts the model's ability to generate semantically related concepts by more than 75%. On the other hand, sparse inference-time methods (SAeUron, SEOT) better preserve utility but leave substantial target leakage. The trade-off is steep for concepts whose internal representations heavily overlap with their evaluated semantic neighborhoods, (e.g., \texttthorse with \textttcat, \textttdog, and \textttbear) and milder for \textttcastle, which is comparatively less entangled within our concept pool and probe layer in activation space. Across targets, damage to related concepts scales monotonically with our activation-overlap measure. These results suggest that perfect unlearning is the wrong target for entangled concepts. The field should evaluate on the Pareto frontier our theorem establishes; current benchmarks, which decouple unlearning accuracy from utility preservation, hide this trade-off and need to be revised.
Abstract: Vision-language models (VLMs) can locate an image region referred to by a text prompt and route the corresponding visual evidence to the output, yet the internal mechanism behind this behavior is not understood. Inspired by retrieval heads in large language models, we ask whether VLMs contain an analogous mechanism for visual retrieval. We answer affirmatively by introducing Visual Retrieval Heads(VRHs), a small subset of attention heads (about 1.7–2.6%) that are causally responsible for grounding text descriptions to image regions. To find them, we recast existing head-scoring methods under a unified design space over query tokens, key aggregation, and cross-sample aggregation. We then show that scoring attention from output prediction tokens with a sum over the ground-truth referent region most reliably identifies causal heads. Across four VLMs and five referring-expression benchmarks, masking only the top 20 VRHs reduces grounding accuracy by up to 80 percentage points, while masking the same number of random heads has little effect. Beyond replicating the causal-sparse-universal triad established for text retrieval heads, VRHs exhibit several properties not previously reported: they generalize across visual reference tasks, remaining causal on attribute, spatial, counting, and visual-math benchmarks despite being discovered through bounding-box prediction; they are functionally specific, preserving output format while corrupting localization; and they are architecturally shared, transferring causally across VLMs that share an LLM backbone but differ in vision encoder, projector, and instruction tuning
PaperID: 7293, Poster
Abstract: Equipping large language models (LLMs) with external memory banks is a critical stepping stone towards autonomous agents. Current approaches have concentrated substantial efforts on the design of memory construction and retrieval mechanisms. However, they typically relegate memory utilization to the simplistic concatenation of explicit texts for LLM inputs, thereby remaining constrained by the fundamental bottlenecks of context overload and semantic dilution. In response to this dilemma, we present UniMem, a unified memory adaptation framework that empowers LLM agents to holistically harness memory information, extending beyond textual signals into latent and parametric dimensions. Operating along these pathways, UniMem not only elevates contextual density through latent compression, but also facilitates semantic grounding via parametric modulation. Extensive experiments show that UniMem (I) achieves superior performance across diverse benchmarks, including long-term dialogue, memory personalization, and open-domain question answering; and (II) can seamlessly complement prevailing memory construction and retrieval mechanisms for synergistic enhancements.
PaperID: 7294, Poster
Abstract: Reconstruction-based provenance intrusion detection systems (PIDS) detect malicious activity by classifying nodes in system-call provenance graphs from per-edge reconstruction errors produced by benign-trained encoder--decoders. Because these detectors run on the endpoints they monitor, adversarial robustness is a critical deployment concern, yet it remains rarely evaluated systematically. We introduce ProvRL, a reinforcement-learning framework for automated red-teaming of these detectors, and use it to characterise the robustness of four leading PIDS spanning the four encoder families used in the field. ProvRL casts red-teaming as a semantically constrained graph-editing MDP, with a factored autoregressive policy masked by transition relations mined from benign data and a GRU belief state, to model the multi-step consequences of message passing on GNN-based detectors. ProvRL causes targeted false negatives in white-, grey-, and black-box settings on three of four encoder families (GNN, linear, VAE) using 5-99x fewer queries than exhaustive search; direct-reconstruction encoders resist insertion attacks structurally, suggesting an architectural direction for robust detection. No threshold-adaptive aggregation policy maintains both a deployable false-positive rate and meaningful evasion resistance, leaving adversarial retraining as the only viable defence on the vulnerable families. ProvRL trains on CPU within hours and produces attack chains that replay as real Linux syscalls. We release ProvRL as an open-source red-teaming tool for PIDS.
Abstract: Neural PDE simulators often receive only a single observed field at deployment. In this setting, a field-to-future predictor can collapse distinct latent problem states into the same deterministic interface, losing the ambiguity needed for reliable rollout and downstream decisions. We propose : first infer a posterior over the minimal task-sufficient problem state, then condition prediction on that posterior. The resulting theory connects the object, the learning target, and the failure mode: Bayes downstream values factor through this posterior, refinement labels make it learnable by proper scoring rules, and deterministic collapse incurs an ambiguity barrier whenever the true posterior is non-Dirac. Synthetic exact-ambiguity experiments show that point-versus-posterior gaps track the predicted barrier. On metadata-hidden PDEBench tasks, posterior recovery reduces pooled rollout nRMSE from 0.175 to 0.132, closing 59.4% of the direct-to-oracle gap. These results suggest that single-observation neural PDE simulation should be posterior-first rather than monolithic field-to-future prediction.
PaperID: 7296, Poster
Abstract: Accurate traversal (travel) time distribution estimation on temporal graphs benefits from satisfying six structural properties: (non-Markov, snapshot). No existing temporal graph learning method jointly satisfies these properties. We propose PATH-HAZ, a discrete-time hazard model that provides guarantees of all six requirements. It composes per-step PMFs via dynamic programming for spatially consistent predictions across shared paths, while cross-attention over historical graph snapshots encodes confounder-induced dependencies between non-adjacent steps on the path. Evaluated on simulated logistics and real-world high-speed rail data, PATH-HAZ surpasses both point-estimate (19.3% RMSE improvement) and probabilistic baselines (56% CRPS reduction), while also allowing for detailed structural explainability of delays.
PaperID: 7297, Poster
Authors: Riccardo Saporiti, Fabio Nobile
Abstract: One of the primary challenges in Bayesian inference on the parameters of a diffusion model from discrete observations is the unavailability of an analytical expression for the transition density function between consecutive observation times, which is needed to derive the likelihood function. Extending previous studies that solve Fokker-Planck (FP) type partial differential equations with Normalizing Flows, we propose a new Normalizing Flow architecture to learn the transition density function of the diffusion process between two observation times. We do so by solving in a Neural Galerkin framework the associated FP equation with a Dirac mass as initial condition, over a specified training distribution of the initial datum and the coefficients of the diffusion. We specifically focus on processes whose diffusion matrix vanishes in certain inaccessible boundary regions, such as Stochastic Volatility models that satisfy a Feller condition. The product of the obtained transition densities evaluated along the observed trajectory approximates the likelihood function, thereby enabling cheap posterior sampling via Markov chain Monte Carlo (MCMC). After the offline training phase, inference becomes significantly more efficient, as it avoids the need to solve the FP equation in real time for each parameter proposed by the MCMC sampler or to rely on other likelihood-free methods for Bayesian inference that involve repeated simulation of diffusion bridges.
PaperID: 7298, Poster
Abstract: Targeting RNA with small molecules is severely bottlenecked by epistemic uncertainty in high-throughput screening data, formally manifesting as the "Unknown-as-Negative" (UaN) bias. We demonstrate that naive Empirical Risk Minimization (ERM) under UaN contamination structurally misaligns optimization trajectories, introducing adversarial decoy gradients that force a topological rupture of the chemical manifold and shatter Structure-Activity Relationship (SAR) continuity. To transcend this epistemic barrier, we propose Topology-preserving Robust Adaptive Contrastive Embedding (TRACE), shifting the paradigm from heuristic data-fitting to theoretically guaranteed manifold reconstruction. TRACE neutralizes topological collapse through three mathematically bounded mechanisms: (i) confidence-weighted importance sampling dynamically rectifies the surrogate risk under posterior drift, establishing a uniform convergence bound to selectively rescue unannotated bioactive scaffolds; (ii) thermodynamic temperature regulation strictly bounds the backpropagated Jacobian factor, enforcing local Lipschitz continuity to geometrically shield SAR against stochastic gradient explosions; and (iii) margin-aware contrastive regularization provides explicit exponential tail guarantees on misordering probabilities, ensuring discriminative sub-Gaussian alignment between genuine binders and decoys. Grounded entirely in these theoretical pillars, empirical validations strictly reflect our formal bounds. Beyond establishing state-of-the-art generalization, yielding an 18.2% relative AUPRC improvement on SMRTnet, TRACE sustains 1.44x higher topological resilience under extreme 50% noise corruption, and fundamentally resolves UaN-induced structural overconfidence by slashing Expected Calibration Error (ECE) by 83.9% and reducing Brier score by 3.4%.
PaperID: 7299, Poster
Abstract: We introduce BLARM, a feed-forward method for video-driven 3D mesh animation. Given a monocular video and a static object mesh, BLARM predicts a temporally coherent animated mesh whose motion follows the video. Rather than relying on explicit rigs or directly regressing high-dimensional vertex motion, we represent animation using a compact set of learned, time-varying rigid motion components and time-invariant vertex-to-component skinning weights. This yields a low-dimensional deformation space without requiring skeletons, cages, skinning weights, or rig annotations. Our architecture conditions geometry-derived deformation latents on video features through factorized spatial-temporal attention, then decodes rigid transformations blended by predicted skinning weights. Trained with trajectory reconstruction, entropy regularization, and motion-aware contrastive learning, BLARM produces accurate and temporally stable animations while recovering compact, interpretable motion structure from monocular video.
Authors: Yuhan Ye
Abstract: We study the parameterized complexity of testing approximate first-order stationarity at a prescribed point for continuous piecewise-affine (PA) functions, a basic task in nonsmooth optimization. PA functions form a canonical model for nonsmooth stationarity testing and capture the local polyhedral geometry that appears in ReLU-type training losses. Recent work of Tian and So (2025) shows that testing approximate stationarity notions for PA functions is computationally intractable in the worst case, and identifies fixed-dimensional tractability as an open direction. We address this direction from the viewpoint of parameterized complexity, with the ambient dimension d as the parameter. In this paper, we give XP algorithms in fixed dimension for the tractable sides, and prove W[1]-hardness for the complementary sides. Moreover, lower bounds under the Exponential Time Hypothesis rule out algorithms running in time \rho(d)s^o(d) for any computable function \rho, where s denotes the total binary encoding length of the stationarity-testing instance. As a further consequence, our results yield the corresponding parameterized complexity picture for testing local minimality of continuous PA functions. We further extend our hardness results to a family of shallow ReLU CNN training losses, with stationarity tested in the trainable weight space. Thus, the same parameterized-complexity picture also appears for simple CNN training losses.
PaperID: 7301, Poster
Abstract: Inference-time search offers a principled way to improve the reliability of large language models (LLMs) by exploring and scoring intermediate trajectories. However, widely used search paradigms exhibit complementary failure modes: sequence search scales linearly but commits early and cannot recover from mistaken prefixes, whereas tree search can backtrack but often becomes impractical for deep search due to the exponential growth of the search space with depth. These limitations become increasingly pronounced in long-chain-of-thought responses. We propose Width-Limited Tree Search (WLTS), which reconciles these strengths by preserving the tree-search structure that enables backtracking while using a dynamic expandability criterion to enforce a strict per-layer width constraint. This design keeps generated search-tree growth linear in depth and sustains verifier supervision across the full reasoning process. WLTS further supports two optional deployment components: prefix deletion for diversity and agreement-based early stopping to reduce computation on tasks with canonicalizable outputs. Across mathematical, coding, and logical reasoning benchmarks, and under both classifier and generative verifier settings, WLTS achieves higher accuracy and better inference-budget efficiency than well-established sequence- and tree-search baselines.
PaperID: 7302, Poster
Authors: Minsu Kim, Ye-Sung Kim, Hyeseong Jeon, Wooseok Hyung, Joshua Lee, Chang-Hwan Im
Abstract: Electroencephalography (EEG) provides a non-invasive measure of ongoing neural activity, but building general-purpose EEG models remains challenging due to the heterogeneity of subjects, devices, and electrode montages. Existing EEG foundation models predominantly rely on input-space reconstruction, which can bias the encoder toward memorizing noise and artifacts. We introduce SPERA (Spherical Prior EEG Representation Architecture), an EEG foundation model that predicts in latent space following the joint-embedding predictive architecture (JEPA). SPERA combines three components tailored to EEG: (i) a hybrid attention backbone that interleaves factorized temporal and spatial attention with periodic full-attention layers, (ii) a Legendre-polynomial spatial prior incorporated into attention to encode varying scalp electrode geometries, and (iii) a relational spectral regularizer that aligns latent similarity structure with spectral views, inducing frequency-aware latent representations. Pretrained on approximately 90,213 hours of EEG from 31,771 subjects across 106 datasets, SPERA achieves the highest average performance across nine downstream tasks spanning clinical, cognitive, and BCI applications. SPERA further exhibits strong parameter efficiency and robustness across varying recording conditions, suggesting its potential as a general-purpose backbone for diverse EEG analyses.
PaperID: 7303, Poster
Abstract: Long-horizon tasks require identifying which past decisions actually mattered for distant outcomes. In deep achievement chains, a single early decision can determine success thousands of steps later, far beyond the reach of short-horizon imagination. We introduce WHAM, a world-model agent for causal credit assignment in long-horizon tasks. WHAM ascends Pearl's causal hierarchy within a learned world model to identify decision-critical bottleneck states, prioritizes replay and imagination around them, and propagates value through sparse bottleneck chains using replay-derived returns. A hierarchical bottleneck critic bridges credit across thousands of timesteps with provably controlled error. On Crafter, WHAM achieves 67.1% (+8.2 over a state-of-the-art baseline), with gains that grow systematically with achievement-chain depth: diamond collection rises 17× to 8.5%, and the best seed reaches 12.5%, exceeding the human expert rate. Long-horizon credit flows where actions shape distant outcomes.
PaperID: 7304, Poster
Abstract: Predicting high-order tensor properties for crystalline materials is crucial for various scientific and engineering applications. Crystal symmetry is one of the primary factors influencing high-order tensor properties, such as elasticity and piezoelectricity, making strict adherence to symmetry constraints essential. However, exactly guaranteeing symmetry compliance remains challenging. Recent approaches rely on enforcing symmetry but often fail to strictly preserve symmetry. In this work, we propose a novel method that guarantees exact symmetry compliance by predicting symmetry-constrained irreducible components of high-order tensors. Specifically, we first develop a computational procedure to identify the basis tensors corresponding to symmetry-constrained irreducible components under various symmetry conditions. This symmetry-constrained basis guarantees that the assembled full tensor strictly adheres to the required symmetry constraints. To predict the numerical values for these irreducible components, we then propose a spherical-harmonic convolutional neural network designed to effectively capture essential high-order tensor information. Extensive experiments validate that our method achieves exact symmetry compliance without compromising prediction accuracy, thereby outperforming state-of-the-art approaches.
PaperID: 7305, Poster
Abstract: Mixture-of-Experts (MoE) models have become prevailing, yet their distributed deployment suffers from severe dynamic load imbalance due to the conflict between static expert placement and inherent routing dynamism. While host DRAM offers capacity to avoid static device storage for load balancing, we identify existing offloading works fundamentally fail in distributed inference due to two overlooked challenges: (1) coupled performance impacts of runtime Host-to-Device transfers due to load dynamism and shared PCIe topology; (2) system-level memory overhead incurred by DRAM usage at scale. In this paper, we propose vExpert that shifts DRAM offloading paradigm from partial residency to full redundancy by virtualized expert storage. We pioneer the first systematic performance model that jointly captures H2D latency, PCIe contention, and load imbalance, building a lightweight allocator to dynamically map physical experts to virtual slots. A novel disaggregated manager is designed to reduce memory overhead in multi-process environments. Evaluated on DeepSeek-V3.1 within an 8-node H100 cluster, vExpert improves load balance ratio by up to 55% and achieves 1.2× end-to-end speedup. Code is available at https://anonymous.4open.science/r/vExpert-B743/.
PaperID: 7306, Poster
Abstract: Detecting abnormal changes in server operational data is an important task in modern computing systems, as unexpected shifts in traffic, workload, or performance metrics may indicate service degradation, abnormal access patterns, system failures, or potential security risks. However, server monitoring data are often naturally observed as functional time series, with complex temporal patterns, high dimensionality, and heterogeneous structures, making traditional change-point detection methods difficult to apply reliably in practice. To address this problem, we propose a general change-point detection framework based on a Functional Embedding Neural Network (FENet), designed to provide robust detection across different data types and structural conditions without relying on strong prior assumptions. The proposed method is motivated by the connection between the classical functional CUSUM (FCUSUM) statistic and neural network representations, which allows FENet to learn effective detection rules directly from data while retaining useful statistical intuition. We provide theoretical analysis to formalize this connection and study key properties of the proposed approach. Extensive simulation studies show that FENet achieves competitive performance compared with existing methods and remains robust under limited training data and distributional mismatch between training and test settings. We further apply FENet to real-world server operational metrics, demonstrating its practical value in detecting and localizing abnormal changes in complex server monitoring systems.
PaperID: 7307, Poster
Authors:
Shurui Zheng, Fanhong Li, Zhixing Huang, Zixi Li, Lei JiAbstract: Prior work has shown that grokked Transformers implement discrete Fourier transforms or group character representations for modular arithmetic. We unify these findings under the Wedderburn--Artin decomposition: using interchange intervention accuracy (IIA), a causal method stronger than ablation, we show that grokked models decompose computation along the Wedderburn components of the target algebra. Across 11 commutative algebras over \F_p (p = 2,3,5,7)---including 6 with nontrivial Jacobson radical---all grokked models achieve raw IIA \geq 0.97. When models predict incorrectly, errors remain compartmentalized by component (3.8× chance), confirming that independence is structural, not a byproduct of accuracy. Training dynamics reveal that Wedderburn alignment emerges synchronously with grokking. For non-commutative groups, models exhibit \emphselective Wedderburn alignment; non-commutativity raises a capacity threshold that larger models overcome.
PaperID: 7308, Poster
Abstract: In this paper, we move beyond the conventional setting of motion in-betweening that relies on structured pose inputs, and instead propose a new paradigm that is conditioned on egocentric (first-person) visual observations, thereby reducing the requirements on both users and upstream systems; we term this task Egocentric Human Motion Bridging. This is motivated by realistic consumer scenarios in VR/AR and robotics, where users typically have access only to raw egocentric imagery (e.g., from wearable devices) without reliable human pose annotations or motion capture signals. Unlike conventional motion interpolation in human pose-conditioned settings or motion generation conditioned on a complete temporal language or image context, the egocentric perspective introduces large viewpoint variations, partial body visibility, and severe temporal ill-posedness due to cross-modal endpoint constraints. To this end, we propose Ego-HMB, a unified diffusion-based model that internalizes both endpoint human pose estimation capability and temporally consistent 3D motion infilling ability within a single model through two training stages. Specifically, we first introduce Egocentric-to-Canonical Pretraining (pretraining stage), where we design a hierarchical curriculum with progressively increasing temporal spans. Then, we present a novel Motion Bridge Diffusion Training with endpoint and physics awareness as a fine-tuning stage, where our bridge diffusion diffuse from neighboring poses to enhance temporal continuity while preserving the global constraints between the two endpoints of a 3D motion sequence. Extensive experiments and comparisons with one-stage baselines (a generative model conditioned directly on two egocentric images) and two-stage baselines (estimation models followed by interpolation models) demonstrate that our approach consistently outperforms these baselines.
Abstract: We study signal propagation in linear recurrent models at finite width. While existing signal propagation theory relies predominantly on the infinite-width limit, it remains unclear for how long that approximation remains accurate when recurrent depth t grows jointly with width n. This question is especially relevant for modern recurrent sequence models, whose natural operating regime involves long input sequences, i.e., large t. We derive exact finite-width formulas for the hidden state signal energies in linear recurrences under complex Gaussian initialization. Using these formulas, we identify the joint depth-width scaling regimes that govern signal propagation: (i) a subcritical regime t=o(\sqrt n), in which the infinite-width approximation remains valid; (ii) a critical regime t~ c\sqrt n, in which non-negligible deviations from infinite-width predictions appear and a nontrivial joint scaling limit emerges; and (iii) a supercritical regime t\gg \sqrt n, in which finite-width effects dominate. Thus, our results pinpoint the precise recurrent depth scale at which infinite-width theory breaks down in long-range linear recurrences. In turn, this shows when standard initialization schemes, such as Glorot, become unstable. More broadly, our results demonstrate that finite-width effects accumulate more rapidly with depth in recurrent models than in feedforward ones, leading to qualitatively different signal propagation behavior.
PaperID: 7310, Poster
Abstract: Parkinson's disease and other gait-related neurological disorders manifest in subtle but biomechanically meaningful changes in how patients move, making computational gait phenotyping a valuable tool for clinical assessment. Yet existing latent representations of motion are typically learned from data alone and miss the biomechanical structure that gives clinical motion its diagnostic meaning, leaving them vulnerable when training cohorts are small or acquisition sites differ. We propose REINS (Riemannian Embedding for INertia-induced motion Spaces), an Euler-Lagrange-induced motion representation that uses the generalized inertia matrix of articulated systems to define a biomechanically grounded geometry on configuration space. Because kinetic energy is itself a quadratic form in velocity weighted by the inertia matrix, this metric is a mechanically determined Riemannian structure, and we learn a latent embedding of the resulting structured space. By defining the representation space itself through rigid-body dynamics rather than enforcing physics as a post-hoc penalty, REINS aligns the geometry of motion representation with the same physical laws clinicians rely on to interpret gait. We apply the framework on Parkinson's disease gait analysis, where deviation from healthy motion is measured as a geodesic distance under the inertia-induced metric, and report results on healthy-control vs.\ PD discrimination and UPDRS-gait severity correlation.
PaperID: 7311, Poster
Abstract: Universal adversarial perturbations (UAPs) aim to learn a single input-agnostic perturbation that consistently induces misclassification across diverse inputs, exposing systemic vulnerabilities of deep models. To enhance practical applicability, data-free UAP methods replace real data with samples drawn from handcrafted synthetic priors and optimize UAPs via activation maximization. However, existing approaches rely on a single synthetic prior throughout training, which biases optimization toward a narrow feature subspace and necessitates costly per-prior tuning. Moreover, the commonly used stochastic gradient ascent (SGA) updates are inherently unstable and tend to converge to sharp regions of the loss landscape, weakening black-box transferability. To overcome these limitations, we propose HPSR-UAP, a data-free framework that couples hybrid-prior reweighting with gradient-guided sharpness regularization. Specifically, we develop a dynamic reweighting mechanism over multiple synthetic priors to mitigate prior bias and avoid overfitting to any single distribution. Building upon this, we introduce a gradient-guided sharpness regularization that smooths the prior-weighted activation loss landscape within a local perturbation neighborhood, promoting updates towards flatter and more transferable directions. Extensive experiments on ImageNet demonstrate that HPSR-UAP outperforms state-of-the-art methods in both white-box and black-box data-free attack settings. We further introduce a multi-granularity block-wise similarity metric that quantifies spatial repetitiveness within UAPs, revealing a strong correlation between structural regularity and attack performance.
PaperID: 7312, Poster
Abstract: , a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning. Instead of asking the MLLM to judge a generated image or answer decomposed verification questions, SpectraReward measures how well the original prompt can be recovered from the generated image through a single image-conditioned, teacher-forced forward pass. We use the average image-conditioned prompt log-likelihood as the reward, directly reusing the MLLM's pretrained image-text alignment ability without preference labels, reward-model fine-tuning. We further introduce a special case for unified multimodal models where the policy's own understanding branch serves as the reward model for its generation branch, forming a closed-loop self-improving framework without external reward models or external knowledge. Extensive experiments validate SpectraReward through a broad image-generation RL study covering two diffusion models, three RL algorithms, nine reward MLLM backbones from four MLLM families spanning 4B to 235B parameters, and five out-of-distribution text-to-image benchmarks. Results show that both SpectraReward and Self-SpectraReward significantly and consistently improve generation performance and outperform prior MLLM-derived reward training methods. Further analysis reveals that larger reward MLLMs are not always better, while Self-SpectraReward can match or surpass much larger external reward models, suggesting that reward-policy alignment is a key factor for effective image-generation RL.
Abstract: Large Reasoning Language Models (LRLMs) leverage long Chain-of-Thought (CoT) reasoning to solve complex tasks, yet often engage in overthinking. Specifically, they continue to generate reasoning steps that provide meaningless contribution to or even degrade the final answer, even after sufficient reasoning has been produced. Early-exit strategies are proposed to dynamically terminate meaningless or even harmful reasoning, aiming to improve both efficiency and accuracy. However, existing methods either suffer from over-truncation that degrades accuracy below vanilla CoT, or yield only spurious efficiency gains, which reduce token counts but increase wall-clock inference time. We observe that when an LRLM's reasoning path deviates into meaningless or even harmful wandering, this deviation is accompanied by an anomalous surge of high-entropy transition tokens. Building on this insight, we propose RPDI-EE, a training-free early-exit method that monitors the Reasoning Path Deviation Index (RPDI), defined as the ratio of local to global average entropy, which serves as a proxy for reasoning path deviation. By tracking this index, RPDI-EE detects and terminates likely meaningless or even harmful reasoning while preserving productive inference. Experiments across multiple benchmarks and four LRLMs of varying types and scales show that RPDI-EE achieves the largest average accuracy improvement over vanilla CoT among all tested early-exit methods, while mitigating the spurious efficiency gains of existing alternatives.
PaperID: 7314, Poster
Abstract: While large language models have significantly advanced Text-to-SQL generation, it remains largely an open-loop paradigm due to the absence of a reliable validator capable of systematically diagnosing generated queries. Training a dedicated diagnostic validator to provide targeted, reflective feedback is crucial for closing this loop. However, employing standard Reinforcement Learning paradigms, particularly Group Relative Policy Optimization (GRPO), reveals fundamental structural misalignments. Through exploratory analysis, we identify a critical limitation of standard GRPO in this context: a severe credit assignment failure, where uniform sequence-level rewards inadvertently reinforce hallucinated reasoning alongside correct sub-answers in complex structured outputs. Concurrently, we discover profound structural dependencies among SQL errors, revealing inherent co-occurrence semantics dictating that certain errors frequently appear together. To address the limitation and leverage this insight, we propose DiagSQL, a novel framework that utilizes an improved GRPO algorithm tailored for diagnostic Text-to-SQL validation. DiagSQL introduces Fine-grained Token-level Reward Allocation (FTRA), which parses structured responses to precisely distribute rewards at the token level, effectively isolating and penalizing erroneous reasoning. Furthermore, it incorporates Co-occurrence-aware Reward Shaping (CORS), which leverages our discovered pre-computed error co-occurrence matrix to dynamically adjust optimization objectives, encouraging logical error combinations while penalizing structural violations. Extensive experiments demonstrate that DiagSQL significantly enhances the robustness and accuracy of SQL error diagnosis, providing actionable, high-quality feedback that effectively transforms Text-to-SQL into a closed-loop paradigm. Our code is publicly available at https://anonymous.4open.science/r/DiagSQL.
Abstract: Kullback-Leibler (KL) regularization is widely used in offline decision-making and offers several benefits, motivating recent work on the sample complexity of offline learning with respect to \emphKL-regularized performance metrics. Nevertheless, the exact sample complexity of KL-regularized offline learning remains largely from fully characterized. In this paper, we study this question in the setting of multi-armed bandits (MABs). We provide a sharp analysis of KL-PCB~(Zhao et al., 2026), showing that it achieves a sample complexity of \tildeO(\eta SAC^\pi^\/\epsilon) under large regularization \eta = \tildeO(\epsilon^-1), and a sample complexity of \tilde\Omega(SAC^\pi^\/\epsilon^2) under small regularization \eta = \tilde\Omega(\epsilon^-1), where \eta is the regularization parameter, S is the number of contexts, A is the number of arms, C^\pi^\ is the policy coverage coefficient at the optimal policy \pi^, \epsilon is the desired sub-optimality, and \tildeO and \tilde\Omega hide all poly-logarithmic factors. We further provide a pair of sharper sample complexity lower bounds, which matches the upper bounds over the entire range of regularization strengths. Overall, our results provide a nearly complete characterization of offline multi-armed bandits with KL regularization.
PaperID: 7316, Poster
Abstract: While the maximal coding rate reduction (MCR^2) objective has become a useful principle for learning compact and discriminative representations, a theoretical understanding of its optimization geometry remains limited, especially on why first-order methods often exhibit fast convergence near its local maximizers. In this work, we show that the MCR^2 objective satisfies an error-bound property in a neighborhood of its local maximizers. This property guarantees that the distance to the local maximizer set can be bounded above by the gradient norm at any point in the neighborhood, directly implying regularity conditions such as the Polyak-Lojasiewicz inequality and quadratic growth. Leveraging this property, we prove that first-order methods converge linearly to a local maximizer of the MCR^2 objective under mild conditions. Numerical experiments validate the predicted local convergence behavior and provide empirical evidence that optimizing the MCR^2 objective improves representation geometry while maintaining competitive downstream accuracy.
PaperID: 7317, Poster
Abstract: RNA function arises from a hierarchical folding process in which primary sequences form secondary structures that govern biological activity. While foundation models have proven effective for learning transferable RNA representations, most existing approaches rely on primary sequence alone or incorporate structural information only superficially during pre-training. Three problems limit current approaches: experimentally determined secondary structures are scarce, computational pseudo-labels are noisy, and no existing model integrates structure at multiple levels across masking, architecture, and supervision. Here, we present N\"uwa.RNA, a pre-trained foundation model built on a \emphstructure-driven multi-level pre-training framework that integrates secondary structure at all three stages of pre-training simultaneously. N\"uwa.RNA introduces dual-granularity masking that combines single-token masking with structure motif masking, treating entire structural elements as indivisible units; a noisy structural attention bias that injects curriculum-corrupted contact maps into every Transformer layer to model pseudo-label uncertainty; and a structure denoising objective that recovers clean base-pairing states from corrupted input, completing a closed denoising loop. Pre-trained on a large-scale non-coding RNA corpus, N\"uwa.RNA achieves state-of-the-art performance on 10 of 13 BEACON benchmark tasks, with up to +7.83 F1 improvement on structure prediction, and attains the highest zero-shot secondary structure prediction accuracy among single-sequence RNA foundation models.
Authors: Yuhuan You, Lai Wei, Tianshu Qu
Abstract: Large audio-language models have made rapid progress in recognizing what is present in an audio clip, yet spatial audio-language understanding still lacks a clear task interface. A model must not only identify sound events, but also decide where they occur, which semantic and spatial attributes belong to the same auditory object, how multiple objects are arranged, and whether a scene-level answer is physically plausible. We formalize this missing capability as audio scene analysis (ASA), a three-level problem spanning atomic perception, relational integration, and cognitive reasoning. We propose (TWNM), a framework that instantiates this definition by equipping audio-language models with explicit spatial evidence. TWNM uses physically grounded First-Order Ambisonics (FOA) simulation to obtain controllable supervision, learns slot-regularized spatial representations from multichannel audio, and fuses these representations with semantic audio features before reasoning with a language model. We further train the model with a progressive curriculum, ending with preference optimization over metadata-derived correct answers and auxiliary format/evidence rewards. To operationalize the ASA definition, we build a controlled benchmark from scene metadata, covering localization, attribute binding, spatial comparison, scene abduction, and counterfactual reasoning. On this ASA benchmark, TWNM achieves 70.8% overall accuracy, 66.4% on spatial-family tasks, and 79.76% on mixed L3 scene-level question answering (QA) under exact multiple-choice question answering (MCQA) scoring. We also audit monaural and binaural reference systems as diagnostic references with explicit audit labels, because they differ in spatial input, training interface, and output format. The supported claim is that a clearly defined ASA task hierarchy, FOA-conditioned spatial representations, and metadata-grounded training together enable controlled, auditable spatial audio-language reasoning, with STARSS23 providing a limited real-recording diagnostic.
Authors:
Trevor Chen, Ariel Dai, Jason Yang, Riccardo De Santi, Daniel Khalil, Wenda Chu, Nate Gruver, Pranav Murugan, Alexander Goldberg, Maruan Al-Shedivat, Yisong YueAbstract: We study how to design online optimization loops for molecular optimization via adaptation of pre-trained generative models. At test time, we aim to utilize limited oracle feedback to steer generation toward top-molecules achieving high rewards. This creates a design problem coupled across several dimensions, including which candidates receive oracle evaluations, how observed rewards are utilized for model update, and how to overcome pre-trained model biases–to effectively explore over complex design spaces. Despite recent algorithmic progress on each individual component, it remains unclear how they interact in practice, within real-world online adaptation loops. For instance, top-K optimization objectives may make adaptation overly greedy and thereby reduce exploration, while model-debiasing techniques may be unnecessary when exploration is already induced by exploratory acquisition functions, e.g., Thompson sampling. To address this, we conduct a controlled study on discrete-diffusion molecular optimization, finding that well-designed components remain beneficial when combined, indicating that they tackle complementary issues. Together, these components yield an online fine-tuning recipe that outperforms offline fine-tuning and search-augmented baselines across several small-molecule binding affinity and protein fitness optimization tasks, under equal oracle-call budgets and GPU-hour accounting.
Abstract: Multimodal Large Language Models (MLLMs) are prone to hallucination as their generation preferences are insufficiently calibrated to visual evidence, causing them to fall back on linguistic priors, rather than faithful grounding. In this work, we start from an empirical observation: when query-relevant visual evidence is explicitly strengthened using the model’s own attention, generation becomes more accurate, suggesting that many failures do not arise solely from missing perception, but from an insufficient tendency to trust the evidence the model has already attended to. Motivated by this finding, we propose Oriented Pickup Preference Optimization (\textttOPPO), an evidence-aware alignment objective that learns preferences over the strength of visual evidence, rather than only response quality. Concretely, \textttOPPO contrasts the same faithful response under stronger, anchored, weaker-evidence views, turning naive visual preference into ordered visual-evidence alignment. We further combine this objective with fine-grained span-level and token-level regularization to stabilize the training. Besides, we provide a theoretical analysis showing that ordered evidence margins induce a positive lower bound on local visual sensitivity. Extensive evaluations across hallucination and general-purpose benchmarks demonstrate that \textttOPPO consistently outperforms baseline methods.
Abstract: LLM agents, which often comprise parallel inference tasks, are commonly adopted to solve real-world problems. When serving such task-parallel LLM agents in shared GPU servers, the scheduler is expected to attain fast agent completion with guaranteed worst-case performance. For that objective, our insight is to selectively pampering agents based on their completion order under idealized fair-sharing. We design Justitia, a fair and also efficient scheduler for task-parallel LLM agents. Noticing that memory is prevalently a bottleneck in LLM serving, Justitia quantifies the true agent cost in a memory-centric manner. It also adopts a light-weight yet accurate method to predict agent costs. Finally, Justitia adopts a virtual-time based fair queuing algorithm to reduce the overall performance with guaranteed worst-case delay. We have implemented Justitia atop vLLM, and experimental results involving diverse agents show that it can substantially enhance the scheduling efficiency with fairness preserved.
Abstract: Multimodal large language models (MLLMs) have demonstrated remarkable capabilities across vision–language tasks, yet their large-scale deployment raises pressing concerns about memorized private data, outdated knowledge, and harmful content. Existing unlearning approaches for MLLMs typically adapt training-based strategies such as gradient ascent or preference optimization, but these methods are computationally expensive, irreversible, and often distort retained knowledge. In this work, we propose MLLMEraser, an input-aware, training-free framework for test-time unlearning. Our approach leverages activation steering to enable dynamic knowledge erasure without parameter updates. Specifically, we construct a multimodal erasure direction by contrasting adversarially perturbed, knowledge-recall image–text pairs with knowledge-erasure counterparts, capturing both textual and visual discrepancies. To prevent unnecessary interference, we further design an input-aware steering mechanism that adaptively determines when and how the erasure direction should be applied, preserving utility on retained knowledge while enforcing forgetting on designated content. Experiments on LLaVA-1.5 and Qwen-2.5-VL demonstrate that MLLMEraser consistently outperforms state-of-the-art MLLM unlearning baselines, achieving stronger forgetting performance with lower computational cost and minimal utility degradation.
Authors: Othmane Mazhar, Huyen PHAM
Abstract: We study nonparametric estimation of Schrödinger bridge (SB) drifts from i.i.d data observed on a single time interval. Starting from the conditional-ratio form of the Schrödinger bridge time-series (SBTS) drift formula, we analyze a direct Nadaraya--Watson plug-in estimator built from kernelized numerator and denominator terms. Unlike recent SB analyses based on entropic-OT potentials, Sinkhorn iterations, or iterative bridge solvers, our approach works directly at the drift level and isolates \emphstatistical error from optimization, approximation, and discretization error. Under Hölder regularity, a marginal-density floor, and bounded support, we prove a uniform non-asymptotic bound for admissible bandwidth pairs, a pointwise CLT under genuine undersmoothing, and an adaptive bandwidth selector satisfying an oracle inequality. We also prove a pivot-local minimax lower bound which, through an explicit uniform pivot, yields a global minimax lower bound under transparent compatibility conditions; hence the adaptive selector is minimax-rate optimal up to logarithmic factors. Synthetic experiments provide theorem-targeted diagnostics for finite-sample scaling, Gaussian approximation, and adaptive behavior.
PaperID: 7324, Poster
Authors: Ziwei Zhao, Yuhang Liu
Abstract: Long-video understanding with multimodal LLMs is fundamentally constrained by the mismatch between video length and the model's limited visual context budget, making frame selection essential for practical reasoning. Existing query-aware methods typically rely on heuristic pipelines: they assign one-dimensional query-conditioned scores to frames and then apply sampling, ranking, or extraction rules to form the final subset. While effective in practice, such methods are generally not derived from an explicit optimization objective and struggle to capture the heterogeneous evidence required for complex reasoning. In this paper, we revisit long-video frame selection from a principled optimization perspective. We formulate it as a structured evidence-allocation problem and propose OTFS, an optimal transport framework that maps multiple evidence sources onto video frames. To reflect the asymmetry of this problem---important evidence should be preserved, whereas redundant frames need not absorb mass---we instantiate this view with a semi-unbalanced entropic optimal transport objective and an efficient solver, OTFS-Sinkhorn. Combined with an exact dynamic program for coverage-aware subset extraction, OTFS yields a practical two-stage training-free pipeline. We further show that, in the single-source case, our formulation induces a Gibbs-type distribution over frames, making Q-Frame's temperature-scaled distribution a special case of our framework while also providing a broader perspective for understanding related prior methods. Experiments on three long-video understanding benchmarks show that OTFS consistently outperforms strong training-free baselines. Code will be released.
Authors: Lukas Billera, Hedwig N Nordlinder, Jack C Ryder, Anton Oresten, Aron Stålmarck, Theodor M Björk, Ben Murrell
Abstract: Diffusion and flow matching approaches to generative modeling have shown promise in domains where the number of elements in a state is fixed in advance (e.g. images), but require solutions when, for example, the length of a response from a large language model, the number of atoms in a molecule, or the number of amino acids in a protein chain is not known . Here we propose Branching Flows, a generative modeling framework that, like diffusion and flow matching approaches, transports a simple distribution to the data distribution. But in Branching Flows, the elements in the state evolve over a forest of binary trees, branching and dying stochastically with rates that are learned by the model. This allows the model to control, during generation, the number of elements in the sequence. We show that Branching Flows can compose with any flow matching base process on discrete sets, continuous Euclidean spaces, Riemannian manifolds, and "multimodal" product spaces that mix these components, and we demonstrate distribution matching on small molecules and antibody sequences, and that this scales to complicated domains such as protein structures.
PaperID: 7326, Poster
Abstract: Large language models (LLMs) have demonstrated strong capabilities on complex reasoning tasks. However, recent studies~\citepgekhman2026thinking,song2026large,jin2025disentangling,cheng2024understanding suggest that when handling knowledge-intensive reasoning tasks, LLMs often fail to effectively utilize the knowledge acquired during pre-training, which limits their reasoning performance. To investigate how internal knowledge (also termed memory) is used by an LLM when dealing with a reasoning query, we propose the Memory-to-Reasoning Alignment (MRA) metric to measure the memory engagement level of an LLM on the reasoning query. We empirically verify that failures in reasoning may correspond to insufficient memory engagement. Inspired by this, we propose a training-free Memory-Guided Activation Injection (MGAI) method to increase the memory engagement level of an LLM, in order to improve the reasoning capability of an LLM. Specifically, given a query, MGAI generates an intervention vector and applies it to the reasoning representation to inject the corresponding memory information into the reasoning representation of the LLM. Furthermore, we propose an adaptive intervention strategy to control which queries require intervention and the magnitude of the intervention. Experiments show that our method consistently improves performance across multiple reasoning benchmarks while largely preserving the memory and general capabilities of LLMs.
PaperID: 7327, Poster
Authors: Yiyu Gui, Mingzhi Chen, Guibo Luo, Yuchao Yang
Abstract: Brain signals exhibit substantial multi-center heterogeneity, strong inter-subject variability, and highly diverse task settings, which makes it difficult to build models that generalize beyond a single task. While recent brain signal foundation models leverage large-scale pretraining to learn generalizable representations, they often fall short in practice because they still require target-data-specific supervised adaptation, lack multi-task zero-shot capability, and rarely provide verifiable evidence to support their predictions. In this work, we present , a multi-task zero-shot Brain signal foundation model with neurometric-anchored Chain-of-Thought (CoT) reasoning. BrainCoT unifies heterogeneous tasks into an instruction-driven framework so that, after pretraining, a single model can be directly applied to multiple tasks in zero-shot settings without additional adaptation. It further introduces neurometric-anchored chain-of-thought grounded in neurometric evidence for verifiable reasoning, and adopts Decision-Consistent Rationale Training to keep process learning aligned with correct predictions. Across extensive evaluations, BrainCoT achieves an average ACC of 67.41% and AUC of 74.33% in zero-shot classification without downstream training data, outperforming the strongest linear-probing baseline trained with 1% labeled downstream data by 3.85% in ACC and 5.97% in AUC. Code will be released upon acceptance.
Abstract: Motivated by the challenge of stabilizing unknown linear dynamical systems (LDS) from observations, we study the fundamental prerequisite of online prediction. Our goal is to achieve sublinear regret with a memory footprint that adapts to the intrinsic complexity of the dynamics rather than the full hidden-state dimension. We focus on the practically central regime of systems with low instability complexity—eigenvalues outside the real stable interval that do not decay rapidly, together with non-semisimple modes—potentially embedded in an otherwise stable real spectrum of much higher dimension; we write k for this count. This regime is the primary setting in which stabilization is plausible: we show that many systems with high instability complexity cannot be stabilized without exponentially large controls. Thus, prediction is meaningful for stabilization precisely when the instability complexity is small. Within this regime, we introduce a unified online algorithm that handles every LDS (including systems with complex or exploding modes) with a learnable parameter count of \widetildeO(k^2), completely independent of the number of stable real modes. Finally, we show that any bounded-coefficient filter-based predictor requires at least k filter directions.
Authors:
Zhuofeng Li, Haoxiang Zhang, Cong Wei, Pan Lu, Ping Nie, Yuyang Bai, Shangbin Feng, Hangxiao Zhu, Ming Zhong, Yuyu Zhang, Jianwen Xie, Yejin Choi, James Zou, Jiawei Han, Wenhu Chen, Jimmy Lin, Dongfu Jiang, Yu ZhangAbstract: Modern retrieval-augmented systems typically access a corpus through a fixed similarity interface (either sparse or dense) that returns only the top-k most similar documents for a query. While efficient, this design becomes a bottleneck for agentic search: exact lexical matching, conjunctions of weak clues, local context verification, and multi-step hypothesis refinement are difficult to realize with a conventional off-the-shelf retriever, and evidence filtered out early cannot be recovered downstream. This problem becomes particularly acute when agents must iteratively discover intermediate entities, combine partial evidence, and revise their search strategy. To address this problem, we study direct corpus interaction (DCI), a simple alternative in which an agent searches the raw corpus directly using general-purpose terminal tools (e.g., grep, file reads, shell commands, and lightweight scripts), without relying on embedding models, vector indexes, or retrieval APIs. This approach requires no offline indexing and adapts naturally to evolving local corpora. Across benchmarks for information retrieval (i.e., BRIGHT and BEIR) and end-to-end agentic search (i.e., BrowseComp-Plus), this simple setup substantially outperforms strong sparse, dense, and reranking baselines. It also attains strong accuracy on multi-hop question answering without relying on any conventional semantic retriever. Our results suggest that as language agents grow more capable, retrieval quality depends not only on reasoning ability, but also on how richly the corpus interface supports interaction. DCI therefore points to a broader interface design space for agentic search. Our code is available at https://anonymous.4open.science/r/dci-agent/.
Authors: Furkan Mert Algan, Eckehard Steinbach
Abstract: 3D shape completion from partial scans remains challenging for unseen categories and noisy real-world observations, where geometry alone is often insufficient for inferring missing structure. We present DinoComplete, a deterministic and efficient shape completion framework that augments geometric reconstruction with voxel-aligned semantic priors distilled from DINO features. First, we construct multi-view DINO feature volumes aligned with ShapeNet data and train a student network to predict dense semantic features directly from incomplete shapes. These predicted features capture global structure and part-aware semantic context while remaining aligned with the underlying geometry. We then integrate these distilled features into a completion network, where geometric and semantic voxel representations are fused through voxel state-space modeling. To enable efficient long-range reasoning without sacrificing resolution, we introduce a multi-scale voxel Mamba module that refines the fused features by combining full-grid and chunk-wise sequence modeling. Experiments on unseen ShapeNet categories and ScanNet objects show that DinoComplete achieves stronger completion quality than prior deterministic and generative based completion methods while using fewer parameters, requiring lower memory, and achieving faster inference. Our results demonstrate that distilling semantic priors from visual foundation models improves generalization and robustness in 3D shape completion.
PaperID: 7331, Poster
Authors: William W Yang, Peter E Latham, Andrew Saxe
Abstract: In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space. In deep learning, this idea is known as the linear representation hypothesis and underpins many interpretability and control methods based on linear probes, from concept detection to activation steering. Yet while prior work has studied whether such directions should exist after training, the dynamics of how they emerge during training remain poorly understood. Here, we develop a framework to study the alignment of concept directions during training -- a process we call "abstraction". In a minimal linear network setting, we obtain exact solutions for the full trajectory of abstraction. These solutions reveal key analytic principles governing abstraction: (i) data and target geometry jointly determine terminal abstraction, (ii) abstraction improves with network depth, and (iii) initialization scale controls the maximum abstraction reached during training. Extending our theory to nonlinear networks, we analyze how the choice of nonlinearity affects abstraction dynamics: erf networks approximate the linear theory, while abstraction in ReLU networks depends less on target geometry and more on input geometry. Across both, we prove a striking attenuation law: both nonlinearities weaken abstraction in activations relative to preactivations. We find evidence for this law in open models (DINOv3, Gemma 4) and apply our theory to improve linear probe generalization in LLMs. Together, our results provide a dynamical theory of abstraction with implications for interpretability and control.
PaperID: 7332, Poster
Authors: Markus Pettersen, Nicolai Haug, Joakim Bergli, Thomas M Surowiec, Mikkel Lepperød
Abstract: A foundational challenge in neuroscience and AI is understanding how physical space is mapped into neural representations. While artificial neural networks can generate brain-like spatial representations, such as those observed in place and grid cells, their "black-box" nature makes it difficult to determine if these representations arise as general solutions or as artifacts of a chosen architecture, objective function, or training protocol. Critically, these models offer no guarantee that learned solutions for core navigational tasks, like path integration (updating position from self-motion), will generalize beyond their training data. To address these challenges, we introduce a first-principles framework based on the exponential map. Bypassing gradient-based optimization entirely, we use generator matrices to map physical locations to neural population vectors. Within the proposed formalism, essential navigational capabilities emerge analytically as intrinsic geometric properties. Exact, trajectory-invariant path integration naturally dictates commuting generators. Furthermore, skew-symmetric generators inherently yield equinorm representations with translationally invariant similarities. Translational similarity invariance is critical for stable egocentric navigation in open fields and empirical models typically enforce them via explicit regularization or architectural constraints. Extending our geometric perspective, we demonstrate that preserving the metric of flat space restricts the spatial wavevectors (derived from the generator eigenvalues) to form sets of roots of unity. The proposed framework supports diverse, biologically relevant spatial tuning profiles, including periodic grid fields, localized place fields, and context-dependent remapping. Beyond explaining mechanisms that underpin recent deep learning models of spatially-tuned cells, our work provides a formal connection between continuous attractor dynamics and oscillatory interference models. Finally, we show how the framework natively supports goal-oriented, multi-map navigation by identifying superpositions of representations alongside remapping as hyperdimensional computing operations. By grounding spatial representations in strict algebraic laws, this work offers a transparent alternative to empirical representation learning, revealing the exact conditions required for a coherent neural map of space.
PaperID: 7333, Poster
Abstract: We present a novel GPU-accelerated global solver for constrained nonlinear programming (NLP) inspired by the efficient linear bound propagation framework in neural network (NN) verification. Existing work on NN verification, a problem that can be cast as feasibility detection for a nonlinear objective, demonstrated orders-of-magnitude speedups with the linear bound propagation-based framework when compared to traditional solvers. Extending this paradigm to global nonlinear programming requires addressing many challenges: a global constrained NLP solver must effectively exploit all, possibly unstructured, nonlinear constraints to improve the objective, whereas NN verification typically handles simple input box constraints over a well-structured NN and does not directly optimize an objective function. We address these challenges by fully exploiting the linear bounds produced by bound propagation: these linear bounds can be viewed as relaxations of nonlinear constraints and used to provide dual objectives, check infeasibility, and guide the search for a primal solution. In addition, we proposed tighter linear relaxations (essential for linear bound propagation) for bilinear functions commonly used in NLP problems, and a projected augmented Lagrangian method tightly coupled to our dual solving procedure to produce primal solutions. Our method solves 399 of 505 bounded and continuous NLP problems in the GAMS and MINLPLib benchmarks under a 180-second limit, outperforming many strong global solvers (including commercial ones) such as SCIP, BARON, and LindoGlobal. Compared to SCIP (the strongest solver on these problems), we solve 104 problems exclusively, including large-scale, highly nonlinear ones that are often out of reach for traditional approaches. Our method demonstrates a distinct, highly complementary regime for GPU-accelerated global optimization of nonlinear programming problems.
PaperID: 7334, Poster
Authors:
Valay Bundele, Susmit Agrawal, Mehran Hosseinzadeh, The Nam Nguyen, Hendrik PA LenschAbstract: Recent segmentation-based trackers built on foundation models such as Segment Anything Model (SAM) achieve strong performance through memory-based mask propagation. However, despite their strong generalization ability, these methods remain brittle under occlusion, reappearance, and distractor scenarios, where errors accumulate and lead to identity drift. Existing approaches attempt to improve robustness through better memory design and temporal modeling, yet they implicitly assume that propagated predictions remain reliable over time, lacking a mechanism to identify and correct failures. To address this limitation, we propose a failure-aware tracking framework that augments segmentation-based propagation with explicit reasoning and selective recovery. Our key idea is to decouple tracking into two complementary processes: verifying prediction reliability and recovering the target when failures occur. Specifically, we introduce a dual-agent architecture consisting of a vision-language verification agent and a correction agent. The verification agent reasons over spatio-temporal context to assess whether current predictions remain consistent with the target object, enabling the detection of identity switches, missed recoveries, and distractor-induced failures. When unreliable predictions are identified, the correction agent is selectively activated to recover the target using temporal cues. To further improve robustness under ambiguous scenarios, we introduce targeted training perturbations that simulate identity switches and and distractor-induced drift during training. By explicitly modeling failure detection and recovery within tracking loop, our framework transforms tracking from a purely propagation-based process into a self-correcting system. Experimental results on LVOS-v2 and SecVOS show that the proposed framework improves SAM3 by 1.2 J&F and 2.7 J&F respectively, demonstrating the value of explicit reasoning and selective recovery in segmentation-based tracking.
Abstract: Real-world time series forecasting faces the fundamental challenge of non-stationary statistical properties, including shifts in mean and variance over time. While reversible instance normalization (RevIN) has shown promise by stationarizing inputs and denormalizing outputs, it relies on the strong assumption that historical and future distributions remain identical. We observe that in many practical applications, distribution shifts follow cyclical patterns that correlate with periodic positions (e.g., seasonal and holiday volatility). To this end, we propose PAMod, a lightweight yet powerful framework that models cyclical distribution shifts via Phase-Amplitude Modulation in the normalized feature space. PAMod learns periodic embeddings to modulate representations: phase modulation captures mean shifts, while amplitude modulation adapts to variance changes. Crucially, we prove mathematically that modulating in normalized space is equivalent to applying dynamic denormalization, offering an elegant unification of distribution adaptation and representation learning. Extensive experiments on twelve real-world benchmarks demonstrate that PAMod achieves state-of-the-art performance with fewer computational resources. Furthermore, our modulation mechanism, as a novel plug-and-play technique, can improve existing time-series forecasting methods with simple integration.
Authors: Duc-Cuong Dang, Roman Kalkreuth, Andre Opris
Abstract: Genetic Programming (GP) is a search paradigm inspired from natural evolution for the automated discovery of expressions, functions, and computer programs. Cartesian GP (CGP) is a flavor of GP that uses a graph-based model to encode candidate programs, as contrasted to the conventional Tree-based GP (TGP). Since its inception two decades ago, CGP has predominantly been analyzed empirically, thus little is known regarding its theoretical performance guarantees. This paper analyzes CGP for the MAX problem, which has been only rigorously studied for TGP with a proven expected runtime superlinear in the optimal program length n. We prove that the expected time for the (1+1) CGP employing common mutation operators to evolve an equivalent optimal of the same output is polylogarithmic in n, specifically at most O(\log^9n). This is remarkable, as it shows an exponential speedup by switching to CGP. Our results shed light on the benefit and the compactness of the graph-based representation for programs, and on how CGP can navigate its search and fitness spaces efficiently. This is a first proven performance guarantee for CGP on an established benchmark problem. Experiments complement our theoretical findings.
PaperID: 7337, Poster
Authors: Michal Prusek, Adam Novozámský, Filip Sroubek
Abstract: Low-budget molecular optimization is often framed as Bayesian optimization in a learned latent space, but standard benchmarks typically allow hundreds or thousands of evaluations, whereas a realistic wet-lab round may allow only a few dozen. We revisit this regime using a SELFIES molecular variational autoencoder and GuacaMol objectives at N=30 oracle calls. The main finding is not a new surrogate model, but a geometric mismatch in candidate generation. The training latents of the VAE concentrate near a narrow spherical shell, while many high-dimensional optimizers propose candidates from boxes or box-shaped trust regions. Such candidates can leave the decoder's training distribution even when their coordinates look individually plausible. We study a simple geometric constraint, the on-sphere scaffold, that restricts candidate generation to the empirical shell occupied by the VAE training latents. Uniform random search on this shell already outperforms most of the classical box-supported Bayesian optimization baselines. We then introduce two minimal optimizers to probe what remains useful once the support is fixed. GeoWalk performs geodesic local search on the empirical sphere and outperforms every published Bayesian optimization baseline tested. GeoCMA keeps the scaffold but adds covariance learning in a random subspace, becoming preferable when the search representation contains learnable directional or semantic structure, as in LassoBench, NAS-Bench-101 graph-VAE optimization, and GSM8K text-latent prompt optimization. A controlled ablation of three recent Gaussian-process recipes, changing only whether candidates are proposed on the empirical sphere or on the published box-supported domain, shows that the scaffold rather than the surrogate is the load-bearing component. The practical lesson is simple: first match the empirical search support; add a surrogate or covariance learner only when the remaining search space contains directional structure that can be learned from the available budget.
Authors:
Harsh Raj, Niranjan Orkat, Suvrorup Mukherjee, Aritra Guha, Cheryl Flynn, Subhabrata MajumdarAbstract: This paper establishes a rigorous measurement science for AI agent reliability, providing a foundational framework for quantifying consistency under semantically preserving perturbations. By leveraging U-statistics for output-level reliability and kernel-based metrics for trajectory-level stability, we offer a principled approach to evaluating agents across diverse operating conditions. Our proposal highlights the important distinction between the core capability and execution robustness of an agent, showing that minor task-level variations can induce complete strategy breakdowns despite the agent possessing the requisite knowledge for the task. We validate our framework through extensive experiments on three agentic benchmarks, demonstrating that trajectory-level consistency metrics provide far greater diagnostic sensitivity than traditional pass@1 rates. By providing the mathematical tools to isolate where and why agents deviate, we enable the identification and rectification of architectural concerns that hinder the deployment of agents in high-stakes, real-world environments.
Authors:
Siqi Luo, Jianghan Shen, Yi Xin, Huayu Zheng, Haoxing Chen, Yan Tai, Yue Li, Junjun He, Yihao Liu, Guangtao Zhai, Yuewen Cao, Xiaohong LiuAbstract: Diffusion Multi-Modal Large Language Models (dMLLMs) are powerful for image generation, but optimizing them through reinforcement learning (RL) remains a major challenge. One primary difficulty is that a single image can be generated through many different unmasking sequences, which makes calculating importance ratios is often intractable. Additionally, existing methods tend to ignore the hierarchical generation process of dMLLMs, where early tokens define the global layout and later tokens focus on local details. By assigning uniform rewards to all tokens, these current methods fail to reflect the actual contribution of each token to the final image. To address these issues, we propose Hierarchical Token GRPO (HT-GRPO), which integrates this hierarchy directly into the policy optimization process. Our approach features a Sketch-Then-Paint training scheme that organizes updates into three distinct stages: global, structure, and refinement. We also use a prompt-conditioned estimator to calculate importance ratios starting from a fully masked state. Furthermore, we introduce a Hierarchical Credit Assignment mechanism that prioritizes key structural tokens to ensure accurate reward propagation. Experiments using two popular dMLLM backbones, MMaDA and Lumina-DiMOO, demonstrate that HT-GRPO achieves substantial gains on the GenEval and DPG benchmarks. Evaluations across six additional metrics confirm significant improvements in image quality, aesthetics, and human preference.
PaperID: 7340, Poster
Abstract: Network-level traffic signal control (TSC) involves coordination under partial observability and strict real-time constraints. In real-world urban networks with dozens or hundreds of intersections, direct coordination incurs an intractably large joint action space, making centralized cooperation difficult to deploy in practice. We propose CorridorLight, which tackles this challenge by casting cooperation as task negotiation. A high-level Corridor Agent represents the network as a direction-level graph and periodically generates a small set of junction-disjoint corridor tasks via a seed-and-grow procedure. These tasks provide lane-level preference guidance, while Intersection Agents maintain fast decentralized control and selectively support the active corridor task. By learning when to cooperate versus act selfishly, CorridorLight enables an adaptive trade-off between local delay minimization and corridor-level relief, leading to improved global performance. To this end, we derive a closed-form, Lipschitz-continuous gating function by maximizing a counterfactual net gain between cooperative and selfish behavior, yielding smooth, interpretable cooperation without brittle heuristics. Experiments on synthetic grids and real-world networks show consistent system-level improvements and outperform cooperative TSC baselines across diverse topologies and demand patterns.
PaperID: 7341, Poster
Authors: Luan T Trinh, Atsuki Osanai, Kenji Doi
Abstract: Subject-driven diffusion transformers preserve coarse object identity but routinely lose identity-defining details such as logos, printed text, textures, and small structural cues. We show that this failure traces to a structural mismatch: identity sensitivity is sparse and non-uniform across layers, denoising stages, and spatial regions, yet reference influence is typically applied as a global signal. We introduce VitalScore, a training-free method that discovers this identity-vital structure through offline probing and uses it to build a reference-attention redistribution tensor \Lambda(l,t,s) coordinated with a stage-aware prompt schedule (P(t). The spatial axis is grounded in a two-phase VLM blueprint calibrated against human identity annotations, ensuring the controller reflects the cues people actually use to judge subject identity. We evaluate across Diptych Prompting, FLUX Kontext, and FLUX 2.0 on DreamBench++, where VitalScore improves DINO-v2 by up to 15%, and on a new product identity benchmark targeting fine-grained commercial cues that standard benchmarks overlook, where gains reach 16.5% on the hardest prompt axis---all without parameter updates.
Abstract: Progress in LLMs is increasingly measured through standardized benchmarks, where state-of-the-art improvements are often separated by fractions of a percentage point. At the same time, the computational cost of evaluating modern LLMs has driven widespread adoption of specialized inference backends, software systems that execute trained models efficiently at inference time. While critical for scalability, system-level optimizations, such as custom CUDA kernels and reduced-precision arithmetic, can alter token probabilities and introduce non-determinism, possibly cascading into divergent generation. In this work, we first survey the inference landscape, identifying 200 distinct engines, and analyze 35,000 ML publications, finding that the specific inference stack is rarely reported despite this widespread diversity. We then present a systematic empirical study of how inference backends affect LLM benchmark results. Holding model weights, decoding parameters, and hardware constant, we evaluate five widely used inference engines, including vLLM, SGLang, and llama.cpp, across multiple open-weight models and established benchmarks. We show that the choice of backend alone can shift benchmark scores by up to 16.6 percentage points and induce high rates of output disagreement. By isolating backend optimizations and tracing the execution pipeline, we find this divergence is driven by system-level optimizations like prefix caching and CUDA graphs, custom kernels, and engine-specific defaults in logit processing. Our findings identify the inference backend as a previously unreported but consequential hyperparameter in the evaluation of LLM and advocate standardized reporting of inference stacks to improve the reproducibility and interpretability of benchmark comparisons.
PaperID: 7343, Poster
Abstract: Online data selection is typically driven by per-sample scores such as loss, uncertainty, or gradient norm, which measure how strongly the model reacts to a candidate but not what parameter update the candidate would induce. What an online learner should select is therefore not the most striking example, but the one whose update is most worth taking. To make this concrete, we score a candidate by the movement induced by its update, using a Fisher-whitened information gain that evaluates the update in the loss geometry rather than by raw gradient size. Movement, however, is not yet progress, so we pass this score through a \emphDescent Alignment Gate, a soft factor that retains it only when the induced update agrees with the descent direction. This yields a sample-level score favoring updates that are both informative and descent-aligned. To make the criterion compatible with training, we extend it from samples to batches, where selected updates should not only score well individually but also complement one another. We capture this with the \emphGated Fisher Volume, a log-determinant batch objective that reduces to the per-sample score on a singleton, grows with the volume spanned by the chosen updates in the Fisher geometry, and admits a greedy algorithm with a constant-factor approximation guarantee. Across supervised fine-tuning benchmarks, our selector consistently outperforms online selection baselines under matched training budgets. More broadly, the framework reframes data selection itself: rather than a spotlight on conspicuous examples, it acts as a compass toward updates worth taking.
PaperID: 7344, Poster
Abstract: Learning across multiple inherently conflicting objectives requires selecting desirable trade-offs from the Pareto set. In this work, we study preference-guided multi-objective optimization, where a predefined preference function is minimized over the weak Pareto set of multiple nonconvex objectives. Leveraging a merit function that vanishes on the weak Pareto set, we propose a penalty-based min-max-min formulation, termed the M^3 problem, which unifies preference optimization and the enforcement of weak Pareto optimality in a single objective. We further develop ParetoM^3, a single-loop first-order algorithm for computing stationary solutions of the resulting M^3 problem. Under a local Kurdyka-\Lojasiewicz assumption, we prove that ParetoM^3 finds an \epsilon-stationary point of the M^3 problem that is also approximately weakly Pareto optimal for the original multi-objective problem when the penalty parameter is sufficiently large, with explicit non-asymptotic convergence rates. We also show that this stationarity notion recovers existing optimality criteria, including approximate preference stationarity and approximate Karush-Kuhn-Tucker conditions, under the same assumptions used in prior works. Experiments on a synthetic benchmark, image classification, regression, and math reasoning demonstrate that ParetoM^3 reliably attains preferred weak Pareto optimal solutions while achieving strong practical performance.
PaperID: 7345, Poster
Abstract: Feed-forward 3D reconstruction models have demonstrated strong generalization under large-scale pre-training, yet they struggle in long-tail scenarios such as low-light reconstruction, joint human-scene reconstruction, and test-time adaptation. We observe relational attention structure collapse, manifested as the failure of cross-view correspondences, as a primary cause of these limitations. While adaptation offers a practical way to bridge this gap, existing parameter-efficient fine-tuning (PEFT) approaches mainly perform channel-wise feature adjustment and lack mechanisms to explicitly adapt relational structures, leaving correspondence failures unresolved. In this work, we formulate adaptation of 3D reconstruction models as a problem of correspondence recovery under domain shift and propose 3R-Adapter, a PEFT paradigm for long-tail 3D reconstruction scenarios. Our method consists of three components: Deformable Retrieval for task-adaptive relation candidate retrieval, Relational Rewiring for reconstructing cross-view relational structures, and Consistency Refinement for injecting the rewired relations into stable predictions. Extensive experiments on low-light reconstruction, joint human-scene reconstruction, and test-time adaptation demonstrate that our approach significantly improves robustness over existing PEFT methods while maintaining efficient lightweight adaptation.
PaperID: 7346, Poster
Abstract: Deploying deep learning models across heterogeneous hardware requires each model to jointly satisfy a bundle of coupled constraints on accuracy, latency, memory, and energy. Maintaining a specialist for every deployment scenario is operationally prohibitive, so practitioners need a small portfolio of K models that collectively covers many deployment profiles. Existing hardware-aware NAS and Pareto-based approaches target marginal objectives and do not directly address this downstream decision. We formalize it as Deployment Portfolio Selection over a fixed pre-measured archive, where a profile is covered only when a single retained model simultaneously satisfies its entire constraint bundle. The induced objective is weighted maximum coverage, which is NP-hard but monotone submodular, so greedy selection inherits a (1-1/e) guarantee, and sample-then-optimize admits a finite-sample bound. Across an 18-family sweep on HW-NAS-Bench, direct bundled-coverage optimization is most valuable in the small-portfolio, low-coverability regime: gains concentrate on the hardest profiles, are most robust at K = 3, and persist under three non-uniform demand models. Complementary analyses of remeasurement noise, runtime, and compact feasibility storage clarify when bundled coverage is preferable to scalarized or Pareto-based alternatives.
Authors: Navjot Singh, Kipton Barros, Sherry Li
Abstract: Objectives involving bilinear forms (u^\top f(A(\theta))v) for Hermitian (A) arise widely in scientific computing and probabilistic machine learning. For large matrices, Lanczos efficiently approximates these quantities, but differentiating them with respect to (\theta) is challenging. Existing approaches either backpropagate through the Lanczos recurrence, requiring reorthogonalization for stability, or apply Arnoldi to an augmented block matrix of twice the original size. Both introduce extra computation and orthogonalization costs that can limit performance on modern hardware. We propose a forward-only gradient approximation that reuses the Lanczos pass and adds very minimal overhead in most cases. We prove that its error is proportional to the Lanczos residual norm, the same quantity controlling the forward approximation. Whereas a traditional adjoint-based calculation would be unstable without reorthogonalization, the new method appears unconditionally stable in our tests. It is also faster than existing state-of-the-art approaches.
Authors:
Ziqing Qian, Haohang Chen, Shengqi Dang, Yuhan Xiong, Canyu Shen, Jiaying Lei, Nan CaoAbstract: Understanding user preferences from noisy and temporally evolving social media behaviors is fundamentally challenging due to interest drift, where user preferences shift across time and exhibit both multi-scale temporal patterns and diverse co-existing interests. To address this, we propose DUMoE, a unified framework for drift-aware multimodal user representation learning. Our model consists of (i) a temporal dynamics-aware backbone that captures and integrates static profiles, short-term behavioral signals, and long-term dependencies into a coherent representation, and (ii) a sparse mixture-of-experts (MoE) interest adapter that disentangles multiple latent interests via expert specialization and adaptive routing. Each expert models a distinct interest subspace, while a gating network dynamically selects and aggregates a sparse subset of relevant experts for each user. To enable stable and effective optimization, we further introduce a three-stage training strategy that decouples backbone learning, expert specialization, and gating optimization. Extensive experiments on real-world social media datasets show that DUMoE consistently outperforms state-of-the-art methods on both user interest prediction and interaction prediction tasks.
PaperID: 7349, Poster
Abstract: Large language models (LLMs) show strong performance in multi-document question answering, but their practical deployment is limited by unreliable and unfaithful attribution to supporting evidence. Existing prompting and training-based methods often suffer from hallucinated citations and lack interpretability in how evidence is selected. In this work, we investigate whether attribution signals are inherently encoded within transformer attention mechanisms. We introduce a sensitivity-based diagnostic that identifies a small subset of attention heads, termed , which are highly responsive to perturbations in supporting documents. Through causal interventions and semantic analysis, we show that these heads play a significant role in evidence identification and exhibit alignment with document-level entailment signals. Building on this finding, we propose a training-free Attention-based Attribution framework that extracts evidence signals directly from Evidence Head attention using global and local strategies. Extensive experiments show that our method consistently outperforms strong baselines while remaining lightweight and interpretable. Overall, our results suggest that structured attribution signals are implicitly encoded in LLM attention, and can be effectively leveraged for faithful multi-document reasoning.
Authors:
Jingyuan Li, Xiaoyi Jiang, Fukang Wen, WEI LIU, Renqian Luo, Yi Zhu, Zuoqiang Shi, Pipi HuAbstract: Discrete diffusion models based on continuous-time Markov chains (CTMCs) have shown strong performance on language and discrete data generation, yet existing approaches typically parameterize the reverse rate matrix monolithically---through proxies such as concrete scores (SEDD) or clean-data predictions (MDLM, GIDD)---rather than aligning the parameterization with the intrinsic CTMC decomposition into jump timing and jump direction. We propose Neural CTMC, which exploits the underlying Poisson structure of CTMC dynamics by separately parameterizing the reverse process through an \emphexit rate (when to jump) and a \emphjump distribution (where to jump) via two dedicated network heads. We show that the evidence lower bound (ELBO) reduces to a path-space KL divergence between the true and learned reverse processes that factorizes into a Poisson KL for timing and a categorical KL for direction, and admits a tractable, gradient-equivalent and consistent loss. Experimentally, scored by Gemma2-9B, our pure-uniform Neural CTMC achieves 16.36 generative perplexity on TinyStories (vs.\ GIDD 37.60 and MDLM 42.66). On OpenWebText, it attains the best perplexity at the same training-token budget across 16--128 sampling steps among the methods we compare (e.g., at 128 steps: Neural CTMC 183.6 vs.\ MDLM 210.5 and GIDD 249.8).
PaperID: 7351, Poster
Abstract: Graph-level anomaly detection (GLAD) is increasingly deployed in high-stakes settings, where its performance is evaluated on static, clean test graphs at deployment. This creates a critical gap in practice, as anomalous graph inputs may be structurally modified after deployment to evade a fixed detector. We formalize this challenge as test-time graph cloaking: an unsupervised score-minimization attack within a bounded graph-edit budget. We introduce score-directed graph cloaking (SGC), an architecture-agnostic white-box attack that operates solely on the scalar anomaly score, making it applicable across diverse GLAD architectures. To isolate adaptive evasion from generic perturbation sensitivity, we pair SGC with budget-matched random edits and identify a distinct directedness gap. While random structural noise largely preserves anomaly separation, score-directed edits can drive anomalous graphs into the normal-score region. To mitigate this, we propose adversarial score training (AST), a normal-only defense that smooths the score landscape via adversarial local score regularization on normal training graphs. Our experiments across diverse benchmarks and GLAD detectors demonstrate that AST empirically improves cloaking robustness against adaptive evasion without relying on anomaly labels.
Abstract: Multimodal large language models are increasingly expected to perform \emphthinking with images, yet existing visual latent reasoning methods still rely on explicit textual chain-of-thought interleaved with visual latent tokens. This interleaved design limits efficiency and keeps reasoning fragmented across separate text and vision channels. We propose , a unified visual latent reasoning framework that treats textual reasoning and auxiliary visual evidence as a shared visual workspace. Instead of preserving text CoT as an independent inference-time path, UniVLR renders reasoning traces together with auxiliary images and learns to compress this unified representation into compact visual latent tokens. At inference time, the model reasons only through visual latents and directly decodes the final answer, avoiding both external tool calls and verbose text reasoning. Experiments on real-world perception and visual reasoning tasks show that UniVLR outperforms prior visual latent reasoning methods while using substantially fewer generated reasoning tokens, suggesting a more unified and efficient paradigm for visual thinking in MLLMs. Our code is available at: https://anonymous.4open.science/r/UniVLR-1B0C/.
PaperID: 7353, Poster
Authors:
Mingcan Yuan, He Li, Zhiyi Ju, Mang Ye, Qingxiong TanAbstract: Drug-target interaction (DTI) prediction is pivotal for accelerating drug discovery, yet existing methods struggle with generalization under cross-domain distribution shifts. Conventional approaches typically rely on static protein representations and assume distributional consistency between source and target domains, often resulting in poor calibration when encountering novel protein families or scaffolds. To address this issue, we present Bi-view Fold-aware prediction (BiFold), a novel unified interaction-adaptive framework that enhances generalization via dual-view representation learning and calibration-aware adaptation. BiFold synergistically models sequence and structural evidence, introducing a routing module to dynamically fuse the two views at the feature-channel level and exploit complementary information. Furthermore, we design a calibration-aware training principle that aligns domain statistics and encourages flatter source solutions. Crucially, target supervision is introduced only after a forward-only warm-up phase to prevent reinforcement of miscalibrated early predictions, effectively mitigating confirmation bias in pseudo-label learning. Extensive experiments demonstrate that BiFold consistently outperforms the state-of-the-art methods across diverse datasets in both in-domain and cross-domain settings. Additionally, interpretability analysis reveals that the routing module partitions the feature space into view-dominant and view-neutral subspaces, providing concrete evidence of adaptive multi-modal complementarity.
PaperID: 7354, Poster
Abstract: Diffusion Transformers enable powerful in-context image editing from a reference image and a text prompt, but also make it easier to maliciously edit public images while preserving the private identity of the reference. Existing DiT-oriented protections, such as DeContext, mainly suppresses context-to-target attention to weaken reference utilization. However, this strategy can be insufficient for complex edits where prompt semantics already dominate generation. Instead, we study proactive protection for DiT-based editors from a conditional-flow perspective, and use a local geometric interpretation of in-context editing, where the denoising trajectory is empirically characterized as being steered by prompt-consistent and reference-consistent directions. Thus, effective protection should not simply corrupt the output, but should drive it away from the original context manifold. To this end, we propose Conditional Flow Hijacking, which adds imperceptible perturbations to the reference image. Instead of suppressing context attention, our method preserves the context pathway, but redirects the context-induced steering signals in the intermediate transformer layers, causing the final trajectory to drift away from the reference consistent solution region. Experiments on FLUX.1-Kontext with different datasets and prompts including attribute and complex scene editing show that our method significantly suppresses reference identity leakage while maintaining image quality and prompt alignment, outperforming prior protection baselines, especially under complex scene editing prompts. We further verify the robustness of the proposed protection method as well as its transferability across other models, highlighting the need for proactive defenses for in-context generative models.
PaperID: 7355, Poster
Authors: Idan Lev-Yehudi, Vadim Indelman
Abstract: Planning under uncertainty in continuous domains poses significant challenges, yet remains essential for autonomous systems. Tree-based search methods such as Monte Carlo Tree Search (MCTS) remain popular, but their branching structure can require sampling budgets that grow exponentially with lookahead depth in the worst case. From a tree perspective, continuous state or action spaces become especially challenging, since the planner must decide where to search in an infinite branching hierarchy. We propose Graph Sparse Sampling (GSS), an online planning algorithm that shares sampled futures across many candidate decisions, rather than sampling separate successors for each candidate action. This branch-free graph exposes large GPU-friendly batches, while using heuristics to focus computation. We prove finite-sample performance guarantees for GSS covering full-rank or low-rank generative simulators via smoothed backups, and discrete or sampled continuous action spaces. Under suitable overlap, regularity, and action-coverage conditions, these bounds have polynomial dependence on the planning horizon, formalizing when shared futures can avoid the exponential horizon dependence of tree-shaped sparse sampling. We demonstrate continuous-control simulations where GSS substantially outperforms tree-based planners on long horizons or achieves near-optimal performance, supporting no-branching graph planning as a useful complementary design principle for online control.
PaperID: 7356, Poster
Authors: Joery A de Vries, Neil Lawrence, Zhenwen Dai
Abstract: Critic-free reinforcement fine-tuning (RFT) for agentic large language models is dominated by the GRPO and DPO families, yet these methods often update too aggressively on the high-variance samples from agentic tasks. We propose (FTW), a critic-free policy-learning algorithm that adapts the cross-entropy method to RFT, obtaining a more conservative policy-learning method. Unlike prior exponentially concentrating approaches, FTW induces polynomial concentration in the order statistic of returns. Through a control-as-inference lens, we show that FTW induces a mild risk-seeking bias that scales linearly with return uncertainty: less aggressive than DPO’s quadratic risk bias, yet more exploratory than GRPO’s risk-neutral profile. In deterministic-reward agentic settings, where uncertainty is primarily epistemic, this linear risk bonus encourages knowledge-seeking without entrenching on noisy samples. Consistent with this perspective, FTW matches GRPO on the perfect-information Sokoban domain while outperforming it on the imperfect-information Search-R1 domain.
PaperID: 7357, Poster
Authors: Ali Raza
Abstract: Diffusion models require a noise schedule, yet for a given objective it is not generally clear before training whether schedule choice should affect trained-model quality. This leaves a diagnostic question open: given a training objective and a dataset, does the objective admit a pre-training spectral criterion for schedule choice? We show that the answer is determined by the structure of the per-step loss. Product-form training objectives admit a Schedule Mismatch Index, a scalar measuring how unevenly a schedule distributes the training signal across timesteps, computable from data-covariance eigenvalues alone, that strictly ranks schedules by proxy-estimator variance; point-evaluation training objectives do not admit the same construction (Theorem 1). This matches the empirical contrast between VLB training, where cosine beats linear by 1.4% on CIFAR-10, and \epsilon-prediction, where the top three schedules fall within 0.5%. The same framework yields a training-aligned timestep sampler q^\star_\mathrmtrain that improves CIFAR-10 by 2.61% under VLB, with concordant gains on Fisher–Rao-weighted evaluator and FID. At 100k iterations across CIFAR-10 and ImageNet-32 with U-Net and DiT-S backbones, the improvement persists in all four dataset-architecture settings (1.99%–2.14%), while the closed-form proxy remains indistinguishable from uniform. These results show that schedule and timestep-sampling choices must be analyzed relative to the trained objective, not the data spectrum alone.
PaperID: 7358, Poster
Authors:
Xiaoxuan Gong, Zhilin Zheng, Tiancheng Lin, Qinji Yu, Jianpeng Zhang, Shaoteng Zhang, Kai Cao, Xiaoli Yin, Yu Shi, Jie Ma, Ling Zhang, Yingda XiaAbstract: The success of deep learning in medical image analysis is often hindered by the long-tail distribution of disease cases, where rare or early-stage tumors are critically under-represented. While controllable synthesis offers a potential solution, existing methods suffer from a specification bottleneck, as low-dimensional signals like text or binary masks fail to capture complex textures and morphological variations. In this paper, we propose an exemplar-driven synthesis framework that utilizes real tumor samples as high-dimensional templates to enable precise, case-specific generation. To the best of our knowledge, this is the first work to introduce reinforcement learning (RL) into the field of 3D medical image generation, called . We formulate the unpaired tumor synthesis task as a constrained trajectory optimization problem and leverage Flow-GRPO to find an optimal generation path. By designing a multi-objective reward system encompassing texture, shape, and semantic consistency, our model effectively balances pathological fidelity with anatomical coherence without requiring paired supervision. Rigorous evaluations across four major tumor types and the external AbdomenAtlas 2.0 dataset demonstrate that our method produces high-fidelity synthetic data that significantly enhances the performance of downstream diagnostic models.
PaperID: 7359, Poster
Abstract: Membership inference attacks (MIAs) are commonly formulated as single-sample hypothesis tests with false positive rate control. However, realistic adversaries often test many candidate samples and aggregate the declared members, making membership inference a multiple testing problem. In this regime, per-sample calibration can produce unreliable sets of inferred members. We develop a false discovery rate (FDR)-controlled framework for membership inference based on the classical local FDR formulation. We instantiate the framework for high dimensional ridge regression. Under a conditional sufficiency condition, we establish FDR guarantee of the proposed method, and derive an asymptotic characterization of its detection power in high-dimensional regimes under an isotropic Gaussian design. Our analysis shows that, under FDR control, stronger ridge regularization reduces membership inference risk. Experiments on synthetic and a real-world dataset validate the theoretical findings and demonstrate that the proposed methods provide more reliable membership discoveries than existing attacks.
PaperID: 7360, Poster
Abstract: Large language models have achieved remarkable capabilities across diverse domains, yet their safety alignment remains vulnerable to jailbreak attacks. In this work, we identify a previously underexplored failure mode—safety generalization lag—where alignment trained predominantly on natural language fails to transfer to the code domain. We show that this lag induces a code-completion blind spot, allowing malicious intent embedded within syntactically valid code to evade safety mechanisms. To exploit this vulnerability, we propose CodeMimicry, a fully automated black-box jailbreak framework that generates structured, object-oriented code prompts to induce harmful outputs via code completion. Experiments on 8 state-of-the-art commercial LLMs demonstrate that CodeMimicry achieves a 96.25% attack success rate with 1.51 queries on average, significantly outperforming both template-based and optimization-based baselines. Beyond empirical performance, we provide a mechanistic analysis of code-based jailbreaks through latent space representations, including projection onto refusal-related directions and activation steering. This analysis offers an explanation of how CodeMimicry bypasses safety mechanisms in code-related domains. Our findings reveal a weakness in current safety alignment and highlight the need for robust alignments in structured domains such as code.
PaperID: 7361, Poster
Authors:
Siva Kumar Sastry Hari, Vignesh Balaji, Sana Damani, Qijing Huang, Christos KozyrakisAbstract: LLM agents can optimize GPU kernels, but each candidate requires generation, compilation, correctness testing, and profiling. Our objective is to reduce the number of expensive attempts needed to find a correct, fast kernel. We study two complementary ways to improve this attempt efficiency by changing the representation that the agent writes in and adding a first-principles performance signal that estimates remaining headroom. We introduce \muCUTLASS, a compact in-context learnable DSL that exposes high-impact CUTLASS optimization choices while hiding template plumbing, and pair it with Speed-of-Light (SOL) guidance for search steering, budget allocation, and integrity checking. On 59 KernelBench problems on H100, switching GPT-5-mini from low-level code generation to \muCUTLASS changes a 0.40× geomean regression versus PyTorch into a 1.27× speedup. Adding SOL-guided steering, it reaches 1.56×. Across model tiers, this combination lets weaker models match or exceed stronger raw-code baselines, and SOL-guided scheduling saves 19--43% of tokens while retaining at least 95% geomean speedup. We also show that integrity filtering is essential because, without SOL-assisted checks, reported speedups can be inflated by up to 1.9× by benchmark-gaming or PyTorch-only solutions.
PaperID: 7362, Poster
Abstract: Flow Matching enables high-quality visual generation via continuous-time dynamics, but inference remains costly due to multiple sequential function evaluations. Existing acceleration methods reduce the number of function evaluations but often introduce additional training overhead, degrade quality, or fail to account for input-dependent variability. We propose an inference-time method, CoFlow, that adaptively selects the step counts for each generation based on the prompt features. Our context-aware CoFlow is trained online with an unsupervised reward that balances efficiency and fidelity. Our method is plug-and-play, requiring no retraining of the generative model. It generalizes to image and video generation, achieving over 2.5× speedup while preserving perceptual and semantic quality. We also provide theoretical insights into the link between adaptive step allocation and discretization error. The anonymized source code is available at \urlhttps://anonymous.4open.science/r/Contextual_Flow_Matching-9970.
PaperID: 7363, Poster
Authors: Piyush Sao, Keita Teranishi, Sudip K Seal, Pedro Valero-Lara, Narasinga R Miniskar
Abstract: Gaussian process hyperparameter optimization via marginal likelihood is notoriously brittle. We identify a precise, removable cause of a common source of line-search inefficiency in unprofiled exact GP training: mismatch between the current scalar log-amplitude and its profiled optimum. The GP negative log-likelihood decomposes exactly as F(a, x) = f(x) + \fracn2(e^-r + r - 1), where r = a - a^(x) is the mismatch between the current log-amplitude a = \log \sigma_f^2 and its analytical optimum a^(x) for the shape parameters x = (\log \ell, \log \tau). The residual \Psi_n(r) = \fracn2(e^-r + r - 1) is a convex, nonnegative penalty that (i) introduces pole singularities at analytic continuations of the kernel matrix past its singular boundary, and (ii) classifies 80% of rejected Armijo trials overall, and up to 100% in high-SNR regimes, as amplitude-caused in our L-BFGS experiments. Profiling out \sigma_f^2 eliminates this term exactly, reducing poles to logarithmic branch points and converting stiff line profiles into nearly flat ones. We introduce an Armijo failure diagnostic that exactly decomposes each rejected unprofiled trial into profiled-shape and amplitude-mismatch contributions, determining whether a particular rejection was caused by scalar amplitude mismatch or by the profiled shape objective. Experiments on synthetic regression tasks show that, in the tested L-BFGS/Armijo settings, amplitude profiling reduces the number of objective evaluations needed to reach a target marginal likelihood by 2-5x, with no change to the per-evaluation cost or the global optimum.
Authors:
Chris Dong, Sonja Kraiczy, Rohit Prashanth Vasishta, Markus Brill, Wesley H Holliday, Niclas BoehmerAbstract: Society constantly has to determine the boundaries of what it deems acceptable, from legislative decisions to the guardrails governing autonomous systems. We initiate the axiomatic study of : given individuals' attitudes toward options, which options should receive societal consent? We organize our analysis around three principles: sufficient support, minority protection, and dominance by decisively better options. Each captures a distinct reason for withholding societal consent from an option. For each principle, we develop a corresponding solution concept that transparently implements the principle and is canonical in a mathematically precise sense. For example, for minority protection, the resulting concept is a consent-adapted version of Moulin's Proportional Veto Core. Balancing multiple principles simultaneously is more challenging. To address this, we develop a game-theoretic characterization of our veto core that naturally gives rise to a family of related concepts. From this family, we identify the as particularly desirable. By making minorities' blocking power depend on the approval support of the options being challenged, it smoothly interpolates between proportional minority protection and majority support. Experiments on five datasets spanning decision-making (such as political elections, ethical AI evaluations, and moral decision-making) show that solution concepts violating a principle in theory also violate it empirically.
PaperID: 7365, Poster
Abstract: Best arm identification, also known as pure exploration, is a fundamental problem in the multi-armed bandit (MAB) model. Recent quantum algorithms achieve O(n/\Delta_[2]) query complexity for identifying the best arm with high constant probability, but all of them require \Omega(\log(1/\Delta_[2])) or \Omega(\log n) adaptive rounds. Motivated by the practical cost of adaptivity on near-term quantum hardware, where each round incurs substantial overhead from state re-preparation and classical--quantum communication, we study quantum best arm identification with no or very limited adaptivity. On the lower bound side, we prove that any non-adaptive quantum algorithm must make \Omega(n\log n) queries when \Delta_[2]=O(1), establishing a separation between adaptive and non-adaptive algorithms. To complement this lower bound, we present an algorithm that identifies the best arm with high constant probability using O(n/\Delta_[2]) queries and only \log^(n)+1 adaptive rounds. En route to the lower bound, we show that the instances and techniques used to prove adaptivity lower bounds for classical bandits, such as those of Agarwal et al. [COLT'17], do not extend to the quantum setting. As such, we employ the polynomial method to prove our adaptivity lower bounds. To the best of our knowledge, this represents the first application of the polynomial method in this context, which may be of independent interest.
PaperID: 7366, Poster
Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress, yet visual hallucination remains a critical barrier to reliable deployment. This paper aims to answer a fundamental question: when MLLMs hallucinate, where are they actually looking? By tracing the inter-layer consistency of visual attention, we reveal that hallucinated reasoning trajectories bifurcate into two pathological states: Attention Locking, where the gaze becomes overly focal and rigidly anchored to limited visual evidence, and Attention Collapse, where attention becomes excessively dispersed and fails to ground reasoning in meaningful cues. In contrast, faithful reasoning maintains a dynamic equilibrium between established visual evidence and broader peripheral exploration. Building on this insight, we propose ReGaze, a training-free framework that redirects the model's gaze toward this healthy equilibrium to mitigate hallucinations. Specifically, ReGaze monitors layer-wise gaze consistency and triggers timely interventions whenever an unhealthy gaze state emerges. For the locked state, ReGaze redistributes attention from over-dominant critical tokens to peripheral regions, encouraging the model to explore richer visual details. For the collapsed state, it gathers scattered attention from non-critical tokens and injects it into critical anchors, amplifying key visual signals. Experimental results demonstrate that ReGaze enables MLLMs to process visual information more effectively and significantly mitigates hallucinations.
Abstract: Recently, Generative Recommenders (GRs), characterized by a unified end-to-end framework, have exhibited astonishing potential in transforming the recommendation paradigm. Despite their effectiveness, we recognize that GRs are still susceptible to the long-standing issue of popularity bias that has pervaded the recommendation community. Although a few studies have attempted to extend traditional debiasing methods to GRs, their effectiveness is marginal, and the fundamental reason why GRs suffer from popularity bias remains under-explored. To bridge this gap, this study focuses on two core aspects in GRs: the optimization of generative framework and the item tokenization based on semantic index. Based on theoretical analyses, we identify that the severe popularity bias emerges from the confluence of a token-level optimization flaw and the undifferentiated property of item tokenization. Accordingly, this study develops a novel generative recommender system, called Ghost, by designing the asymmetric unlikelihood optimization and the skeleton-founded tokenization. Extensive empirical evaluations across three datasets, alongside multiple SOTA baselines, reveal that Ghost substantially alleviates popularity bias and promotes fairer recommendations, while incurring slight degradation to the overall recommendation utility.
Abstract: Optimizing functionals over the space of probability measures is now ubiquitous in machine learning. A widely used approach is to perform the optimization directly over the Wasserstein space, but many objective functionals of practical interest are non-convex along Wasserstein geodesics, making the analysis of standard first-order methods challenging. In this work, we study a class of objectives over the Wasserstein space that admit a difference-of-convex (DC) decomposition and we lift the classical convex-concave procedure (CCCP) to this setting. Under smoothness and strong convexity assumptions on the convex components of the decomposition, we prove almost stationarity along the iterates of the resulting algorithm. Our main focus is on the Maximum Mean Discrepancy (MMD) and the Energy Distance (ED) functionals, for which we develop explicit Wasserstein DC decompositions, and establish local convergence of the scheme under mild assumptions. Empirically, we show that well-chosen DC decompositions yield faster and more stable convergence than Wasserstein gradient descent on these MMD objectives.
Abstract: Designing continuous trajectories whose time-averaged occupancy provably matches a prescribed spatial density (the ergodic coverage problem) is central to UAV-assisted data collection and sensing, robotic exploration, and mobile monitoring. For flying agents in particular, this challenge is acute: trajectories must balance coverage fidelity against tight energy budgets, no-fly zones, and acceleration limits. Existing methods either re-optimize each trajectory online (with cost growing in the horizon and re-running for every target, agent, and realization) or rely on bespoke analytical constructions that must be re-derived for each new constraint. We propose a pushforward framework that decouples ergodicity from density matching: an analytic latent trajectory provides exact uniform ergodicity on a simple annular domain, and a single map, learned offline by optimal-transport conditional flow matching, transports this latent occupancy onto the prescribed target density. The composed trajectory is then asymptotically ergodic with respect to the learned pushforward distribution, with deviation from the target controlled by the flow-matching training loss. Once trained for a given target density and constraint set, the map serves an unbounded number of trajectories and a multi-agent fleet without per-agent retraining, and many differentiable operational constraints (no-fly zones, acceleration ceilings, or fairness penalties) enter as additive soft penalties in the training loss without re-deriving the design. We prove three results (an acceleration-energy bound, an O(1/\sqrtK) ergodic convergence rate in the number of trajectory cycles K, and an approximation-error bound) that combine into an end-to-end coverage bound estimable from CFM training diagnostics (certified given an architectural Lipschitz bound on v_\theta). Experiments on synthetic targets empirically support each theoretical envelope; on a real UAV-coverage dataset, the proposed method gives the strongest observed coverage--energy and coverage--constraint tradeoffs among the evaluated time-warping, optimization, and concurrent flow-matching baselines.
Authors: Hamid Kazemi, Atoosa Malemir Chegini, Maria Safi
Abstract: Safety alignment in language models operates through two mechanistically distinct systems: refusal neurons that gate whether harmful knowledge is expressed, and concept neurons that encode the harmful knowledge itself. By targeting a single neuron in each system, we demonstrate both directions of failure --- bypassing safety on explicit harmful requests via suppression, and inducing harmful content from innocent prompts via amplification --- across seven models spanning two families and 1.7B to 70B parameters, without any training or prompt engineering. Our findings suggest that safety alignment is not robustly distributed across model weights but is mediated by individual neurons that are each causally sufficient to gate refusal behavior --- suppressing any refusal neuron bypasses safety alignment across diverse harmful requests.
PaperID: 7371, Poster
Abstract: Multimodal reasoning systems increasingly assume that longer chain-of-thought uniformly improves downstream grounding. We show that this assumption fails in referring audio-visual segmentation: simple queries can be harmed by excessive reasoning, creating an “overthinking trap” that dilutes visual focus, while genuinely ambiguous queries still require multi-step reasoning. To systematically study this, we construct counterfactual reasoning-budget labels under a rigorous, leakage-free evaluation protocol. Our results recast adaptive multimodal reasoning as a pre-decisional representation problem: before the first generated token, the generation-onset state contains linearly decodable information predictive of whether additional reasoning will improve segmentation. A linear probe predicts binary reasoning need with 70.1% accuracy, substantially outperforming text-only features and uncertainty-based proxy signals. As a router, the onset-state controller improves over downstream confidence routing and closely matches Short-then-Decide while using far fewer tokens. Mechanistic decomposition suggests that this signal is better captured by distributed high-dimensional representations than by isolated manual heuristics. Reusing this onset state for lightweight, single forward-pass routing reduces autoregressive token consumption by 60% while retaining ~96% of always-long performance. It further transfers to out-of-distribution (OOD) benchmarks without target-split tuning. Comprehensive error analysis shows that bypassing unnecessary reasoning can reduce phrase and localization drift. Ultimately, our results indicate that adaptive multimodal reasoning can be reliably and efficiently routed from pre-decisional internal representations.
PaperID: 7372, Poster
Abstract: Novel View Synthesis(NVS) for thermal infrared scenarios has become increasingly important in practical applications. Recent advances in 3D Gaussian Splatting(3DGS) have demonstrated its strong performance for NVS, and a handful of preliminary efforts have explored its extension to thermal NVS. However, existing 3DGS-based methods for thermal NVS often suffer from severe local blur in rendered images. Theoretical and empirical analyses reveal that the single-channel nature of thermal data allows intensity discrepancy to be compensated through opacity adjustment, which can erroneously amplify the opacity of floating Gaussians, especially when the background behind them is insufficiently represented. To address this issue, we propose a method termed Floating Gaussian Suppression (FGS) for thermal NVS. Specifically, we leverage an over-saturation of luminance to introduce an implicit gradient penalty, which guides the optimizer to suppress the opacity of floaters to match the ground truth. By effectively reducing the opacity of these floaters, this mechanism alleviates the occlusion of distant details, thereby mitigating local blur. Moreover, to preserve rendering fidelity, the luminance parameters of Gaussians are optimized concurrently. Experimental results demonstrate that our method effectively mitigates local blur and produces sharper and more faithful thermal reconstruction compared with the baseline methods.
Abstract: Continual model merging addresses a more realistic setting in which task-specific models arrive sequentially. Naive extension of standard model merging approaches to the continual setting leads to task balance collapse and substantial performance degradation. To address this issue, we propose Global Singular Subspace Separation and Restoration (GS3R), a data-free and optimization-compatible framework for continual model merging. GS3R preserves task balance by globally separating singular components and introducing vector restoration matrices for perturbation-free recovery. To further enhance efficiency, we introduce Pre-merge Restoration and Flushing (PRF), which guarantees peak memory usage comparable to that of other methods. Experiments on vision and language tasks demonstrate that GS3R consistently outperforms prior continual model merging methods.
PaperID: 7374, Poster
Authors: CHIU-CHANG CHENG, Ya-Ning Chang, Chao-Hung Wang
Abstract: Spiking neural networks (SNNs) with learnable membrane time constants can improve temporal processing by adapting neuronal integration timescales, but they also turn the decay factor \beta into an optimized dynamical parameter that is fragile under hardware-induced parameter mismatch. We develop a perturbation-based account of how time-constant variation affects trained SNNs. In controlled software simulations, this fragility is strongly task-dependent: Spiking Heidelberg Digits (SHD) and Spiking Speech Commands (SSC) settings lose 7\text--20 percentage points of accuracy under 20% coefficient-of-variation (CV) perturbations in \beta, whereas DVS-Gesture and CIFAR-10 lose less than 3 pp. We show that, under a surrogate-linearized first-order analysis, the failure mode is governed by the membrane sensitivity \Psi. This quantity admits a closed-form online recursion requiring only one additional state per layer and no extra temporal storage. The analysis identifies time-averaged sensitivity energy as the controllable term in deployment fragility, leading to MARS (Membrane-Aware Robustness through Sensitivity Regularization), a training objective derived from perturbation analysis rather than heuristic noise augmentation. We prove a conservative worst-case output-perturbation bound and introduce a typical-case diagnostic explaining why sensitivity remains predictive when worst-case constants are vacuous. At 20% CV perturbation, MARS reduces the accuracy drop in the sensitive settings to at most 0.5 pp while preserving nominal accuracy, with minimal effect on DVS-Gesture and CIFAR-10.
PaperID: 7375, Poster
Authors:
Yuchen Fan, Chang Wei, Pao-Hsiung Chiu, Chin Chun Ooi, Heyang Wang, Jian Cheng WongAbstract: Scientific machine learning has opened new avenues for solving parameterized partial differential equations (PDEs), enabling models to learn a family of PDEs and generalize to unseen instances. In this context, data-driven operator learning methods typically require large training datasets, while physics-informed neural networks (PINNs) suffer from difficult optimization and limited generalization, especially for nonlinear PDEs. We propose Newton-PINet, a physics-informed network enhanced by Newton linearization, offering an effective meta-learning framework for nonlinear PDEs. Newton-PINet (i) employs a physics-informed multilayer network with skip connections, where the output-layer weights are solved by least squares; (ii) adopts a two-stage learning strategy that first leverages gradient-based training to learn robust representations from the available training tasks, and then performs gradient-free fine-tuning on the output layer for fast task-specific generalization; and (iii) incorporates a Newton linearization method to speed up the least-squares iteration for nonlinear PDE problems. On a challenging nonlinear reaction-diffusion benchmark, Newton-PINet achieves up to three orders of magnitude lower relative error than recent neural solvers, while using 16× fewer training tasks and over an order of magnitude less training time (under 5 minutes versus several hours). This work advances the meta-learning of PINNs toward data-efficient, fast, and generalizable physics solvers. The datasets and code are provided in the supplementary material.
PaperID: 7376, Poster
Abstract: In non-stationary multi-agent settings, concurrent policy updates continually reshape inter-agent coordination patterns. In cooperative tasks, an agent’s current action value is closely tied to teammates’ current and future behaviors, making it important to infer how these behaviors evolve and how they influence local action-value estimation. However, existing methods often lack a value-level representation that connects evolving teammate-dependent interactions to local action-value learning. Firstly, we theoretically show that the individual utilities under Centralized Training with Decentralized Execution(CTDE) paradigm struggle to faithfully characterize coordination-dependent action values, making it necessary to introduce a computable proxy for the missing coordination information. Then we propose Hierarchical Inference Agent-Centric World Model (HIA), a framework that incorporates role information into a Transformer-based world model to predict future teammate interactions from each agent’s perspective. The predicted teammate-action rollouts are transformed into auxiliary interaction utilities and integrated into individual value estimates, enabling each agent to capture how teammates’ future actions modulate its current action value. Experiments on StarCraft Multi-Agent Challenge (SMAC) and SMACv2 show that HIA achieves superior cooperative performance over strong baselines.
Abstract: This paper studies the convergence of the Optimistic Multiplicative Weights Update algorithm (OMWU) in two-player zero-sum games. Recent works have identified instances on which the last-iterate of OMWU can converge arbitrarily slowly, but understanding when and why this slow convergence occurs has remained open. In this work, we develop a new analysis framework that gives sharp, quantitative explanations for this behavior. Our analysis is based on viewing the algorithm's dual iterates as an optimistic skew-gradient descent with respect to an energy function. We prove over the dual iterates that energy is dissipative, and by establishing tight bounds on the magnitude of dissipation, our analysis quantifies the geometric bottlenecks that arise when the corresponding primal iterates are close to the simplex boundary. This further translates into a new linear last-iterate convergence rate in KL divergence on games with a unique and interior Nash equilibrium. Compared to prior work, this new rate contains a much sharper dependence on game-specific constants, and we prove this dependence is optimal. Moreover, these geometric insights further translate into new separations on uniform convergence rates for OMWU. On the one hand, we prove constant lower bounds on the uniform best-iterate convergence rate in KL divergence and Total Variation distance from Nash. On the other hand, we establish for the 2× 2 setting a new \widetilde O(T^-1/2) best-iterate rate in duality gap, improving substantially over prior work. Together, this shows in general that uniform convergence rate guarantees do not transfer across different measures of distance to Nash.
PaperID: 7378, Poster
Abstract: Edge-cloud collaboration between large language models (LLMs) and small language models (SLMs) offers a promising way to leverage the problem-solving capability of LLMs while keeping private data on device. Existing methods typically have the LLM generate guidance from an initial user query, which the on-device SLM combines with local private data to produce the final response. However, the LLM's general guidance often does not align with the user's personalized needs, limiting the utilization of the LLM's capabilities. It has been observed that the on-device SLM exhibits metacognitive capabilities, which allow it to recognize its limitations and adjust its reasoning process. Motivated by this, we transform the on-device SLM into a metacognitive inquirer agent that can assess its internal state and actively decide to query the LLM for personalized guidance. Specifically, we propose Active Inquiry via Metacognitive Actions (AIMA). AIMA learns an on-device policy that jointly optimizes the selection of high-level metacognitive actions and the action-conditioned query refinement. We further introduce a privacy-constrained rewriting mechanism that detects and eliminates sensitive information leakage in the refined query. Extensive experiments and theoretical analysis demonstrate that our method outperforms existing methods.
PaperID: 7379, Poster
Abstract: Visual state space models (SSMs) have emerged as efficient alternatives to attention-based backbones, yet their performance remains sensitive to how visual evidence is sampled and propagated. Deformable visual Mamba methods address fixed scanning by predicting adaptive sampling locations, but offset prediction is still driven primarily by local responses, producing locally plausible samples that are weakly coupled to the SSM's recurrent state evolution, particularly in off-center or multi-object scenes. We present , a scan-path-aware deformable visual Mamba framework that reformulates offset prediction as recurrent scan-path prediction. Two components drive this design. First, an organizes visual tokens along a center-to-periphery path, providing a stable propagation prior that does not assume object centering. Second, a lightweight Mamba module embedded within the offset prediction network forms the , which predicts offsets from hidden states accumulated along the ring path. RS-OPN tightly couples deformable sampling with recurrent state evolution and produces coherent, object-aware sampling trajectories. Foveal-Mamba consistently outperforms strong baselines across multiple benchmarks. On ImageNet-1K, Foveal-Mamba-Tiny achieves 84.3% Top-1 accuracy, surpassing DAMamba-Tiny by 0.5% with fewer parameters and FLOPs. On ADE20K semantic segmentation, Tiny/Small variants reach 50.8/51.8 mIoU. On COCO, it achieves 48.9/50.2 box AP and 43.7/44.9 mask AP for object detection and instance segmentation, respectively. Qualitative results further demonstrate earlier object focus and more stable activations in multi-object scenes.
Authors: Humera Sabir, Fatima Farooq, Ashraf Aboulnaga
Abstract: Knowledge Graphs (KGs) are a rich source of structured data, and graph neural networks (GNNs) are the dominant tool for learning over them. However, existing message-passing GNNs struggle to scale to large KGs because they rely on the iterative message passing process to learn graph structure. This process is inefficient, especially under mini-batch training, where a node sees only a partial view of its neighborhood. In this paper, we address this problem and present gHAWK, a novel and scalable GNN training framework for large KGs. The key idea is to precompute structural features for each node that capture its local and global structure before GNN training even begins. Specifically, gHAWK introduces a preprocessing step that computes: (a) Bloom filters to compactly encode local neighborhood structure, and (b) TransE embeddings to represent each node's global position in the graph. These features are then fused with any domain-specific features (e.g., text embeddings), producing a node feature vector that can be used with any GNN backbone. Equipped with these structural priors, the GNN no longer needs to rediscover graph structure, significantly improving memory usage, convergence, and model accuracy. Extensive experiments on large datasets from the Open Graph Benchmark demonstrate that gHAWK improves accuracy on both node property prediction and link prediction tasks. Remarkably, gHAWK enables simple GNNs to outperform more complex relation-aware ones, suggesting that careful preprocessing can substitute for architectural complexity in KG learning.
PaperID: 7381, Poster
Abstract: Large language models frequently ignore in-context evidence that contradicts their parametric priors, producing answers that reflect default beliefs rather than the supplied passage, a phenomenon termed . This failure is commonly attributed to insufficient encoding of the contextual signal. We challenge this explanation. Through systematic, layer-by-layer analysis in a controlled multi-hop question-answering setting, we show that models faithfully encode the context-supported answer in their intermediate representations, yet suppress it in the final layers before generation. We term this : the representational support for the evidence answer rises through the middle layers, then collapses in the final ~40% of the network. To explain this phenomenon, we identify a , recovered by a learned linear probe, encodes the contextually supported answer; a nearly orthogonal , aligned with the unembedding geometry, determines the model's prediction. These two directions are causally severed: interventions along the knowledge direction produce no change in output, whereas targeted suppression of prior-restoring computation along the output direction reliably causes the model to follow the contextual evidence. Our findings hold across nine model-dataset combinations and reveal that knowledge conflict failures arise not because models fail to represent contextual knowledge, but because their generation pathway declines to consult it.
Authors: Thomas Tulinski, Simona Cocco, Remi Monasson, Jorge FERNANDEZ-DE-COSSIO-DIAZ
Abstract: Energy-based models (EBMs) are flexible generative architectures inspired by statistical physics, but their learning and generative properties remain poorly understood. Here, we analyze a solvable EBM in the high-dimensional limit: the spherical Boltzmann machine (SBM). Combining tools from random matrix theory and dynamical mean-field theory, we: solve exact equations describing the training dynamics of the SBM; compute the Bayesian evidence, which acts as a partition function in parameter space and encodes global properties of the trained model; and uncover cascades of phase transitions that occur both during training and as a function of hyperparameters, related to successive alignment and condensation of the top modes of the coupling matrix to the data. We connect these transitions to sampling-time generative phenomena in a teacher-student scenario, including: sampling temperature tuning, double descent as a function of regularization strength, tempered posterior effects, and out-of-equilibrium effects during training that induce biases in the trained model. We provide numerical evidence demonstrating that all these phenomena appear in standard generative architectures, beyond the SBM.
PaperID: 7383, Poster
Authors: Boyang Li, Matthew Kim, Sylvia Herbert
Abstract: Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. A popular line of research in safe RL relaxes safety to a soft expected-cost constraint and solves the resulting Constrained Markov Decision Process (CMDP) via primal-dual Lagrangian updates that only enforce safety on average. To address this limitation, hard, state-wise constraints are introduced and often imposed through Hamilton-Jacobi (HJ) reachability. Yet such constraints require solving different objectives in the feasible and infeasible regions of the state space: reward maximization in the former, recovery toward the feasible regions in the latter. The resulting target action distributions are inherently multimodal, and this structure poses a fundamental challenge for the Gaussian or deterministic actors used in existing HJ-based safe RL, which often collapse onto suboptimal modes. Diffusion policies provide the expressiveness needed to represent such distributions, and recent work on Q-score matching offers a route to training them for online RL by score regression---but has been applied only to reward maximization. Building on this framework, we propose Safe Score Matching (SSM), an off-policy actor-critic method that adapts Q-score matching to hard-constrained safe RL by gating a two-branch score target with HJ reachability: inside the feasible set, the denoising process degenerates to Q-score matching on actions classified as viable by the HJ critic, encouraging reward maximization; outside, a recovery branch biases denoising toward regions with lower worst-case violation, guided by the negative action gradient of the HJ safety value. Across high-dimensional quadrotor and fixed-wing trajectory-tracking and stabilize-and-avoid benchmarks, SSM achieves competitive performance without sacrificing safety, in contrast to primal-dual and reachability-based baselines, which tend to trade off performance for safety and are thus overly conservative.
Authors: Yaacov Pariente, Vadim Indelman
Abstract: Online POMDP planners optimize the expected cumulative cost, which can mask dangerous states when the belief places significant mass on high-cost states. Existing risk-averse methods apply static or dynamic Conditional Value at Risk (CVaR) to the value function, capturing trajectory-level risk, but share two gaps: (i) by retaining the immediate cost as an expectation of a state-dependent cost over the belief, the risk \emphwithin the belief is left unaddressed; and (ii) by modifying the value function, they require new tailored algorithms rather than reusing existing expectation-based planners. We instead apply CVaR to the immediate cost over the belief at each step, directly targeting per-step uncertainty about the current state. The standard expected cumulative return is retained as the objective, so the resulting problem has a standard MDP structure: any expectation-based POMDP planner can be made risk-sensitive by changing only the cost computation. We inherit finite-time guarantees for policy evaluation and sparse sampling---with estimation error independent of the risk level---and, as our central theoretical result, prove a finite-time bound on the gap between the particle belief MDP surrogate and the original POMDP, which together yield an end-to-end guarantee from the true POMDP value to the algorithmic estimate. In the risk-neutral limit, the formulation recovers standard expectation-based planning.
Abstract: Cell-level dense prediction is central to computational pathology, but remains challenging due to fine-grained histological structures, strong domain shifts, and costly dense annotations. Existing ViT-based pathology foundation models rely on patch tokenization, which can disrupt spatial continuity and weaken local morphological details needed for cell-level prediction. To address this, we propose Masked-Diffusion Convolutional Foundation Models, termed ConvNeXt Masked-Diffusion (CMD), a self-supervised convolutional generative pretraining framework for dense pathology representation learning. CMD uses a fully convolutional ConvNeXt-UNet backbone, performs masked-diffusion pretraining in pixel space, and incorporates frozen pathology foundation model features through adaptive normalization. Experimental results demonstrate that CMD consistently outperforms existing ViT-based pathology foundation models and even surpasses state-of-the-art end-to-end segmentation methods while fine-tuning only a small number of task-specific parameters across multiple pathology dense prediction tasks. The advantage is particularly pronounced under limited annotation settings, where CMD exhibits stronger robustness and generalization ability. Our findings suggest that purely convolutional architectures can also serve as competitive pathology foundation models for cell-level dense prediction, achieving leading performance within the current ViT-dominated paradigm and providing a scalable, high-performance solution that better preserves histological structural priors for fine-grained pathology understanding. Project page: https://anonymous.4open.science/r/ConvNeXt-Masked-Diffusion-5BC6.
PaperID: 7386, Poster
Abstract: Deep narrow neural networks underlie modern architectures from ResNets to Neural ODEs, yet their fundamental computational properties at constant width remain poorly understood. We prove that, for every deterministic Turing machine M, there is a single ReLU network N_M \colon [0,1]^4 \to [0,1]^4 of width 4 and depth \mathrmpoly(|M|), polynomial-time constructible, whose iterates simulate M on every infinite binary tape encoded in a \emphconvex initial region of [0,1]^4, with iteration count proportional to Turing-machine time. This establishes a tight correspondence between depth in width-4 networks and time on a Turing machine. The correspondence has three immediate consequences. First, verification of width-4 ReLU networks spans the full complexity hierarchy along the single axis of depth: it is NP-complete for unary-encoded time bound, NEXP-complete for binary-encoded time bound, and RE-complete (undecidable) for unbounded time. Second, the hardness extends with two extra neurons of width to every common smooth activation: at width 6, the robust verification problem (a promise problem with \delta/2\delta gap) is NP-complete for \tanh, sigmoid, GELU, Swish, ELU, and any polynomial-time-computable continuous non-polynomial activation admitting a polynomial-time computable Lipschitz constant. A corollary is a discretized Neural ODE analogue at dimension 5. The simulation and statements are carried out on convex inputs, covering a regime previously inaccessible because of a continuity obstruction inherent to \ReLU networks.
PaperID: 7387, Poster
Abstract: CLIP serves as a foundational vision-language model and the de facto vision encoder for downstream VLMs such as LLaVA. Post-training offers a resource-efficient route to refine CLIP, but recent work argues that the standard contrastive loss is unsuitable for post-training due to catastrophic forgetting under small batches, motivating designs that abandon the contrastive objective in favor of distillation. We revisit this premise and find that the reported forgetting is not caused by insufficient negatives, but by an inappropriate magnitude of the contrastive temperature \tau: with \tau set sufficiently small, contrastive loss becomes the strongest single-loss post-training objective. Building on this finding, we propose ComCLIP, a complementary post-training framework that freezes CLIP's text encoder---preserving compatibility with downstream VLMs---and trains the vision encoder under three complementary losses operating in the shared contrastive space: a properly-tempered contrastive loss for image-text alignment, an MSE anchoring loss against the original CLIP for preserving the pretrained feature manifold, and a relational distillation loss from DINOv2 for fine-grained visual knowledge injection. With only one epoch on small data, ComCLIP consistently outperforms prior post-training methods on ViT-B/16 and ViT-L/14 across zero-shot classification, retrieval, linear probing, and MMVP, achieving up to \mathbf+7.41 points on MMVP. Furthermore, ComCLIP serves as a drop-in vision encoder for LLaVA-7B, improving average performance across 8 standard VLM benchmarks without any re-alignment of the LLM or projector.
PaperID: 7388, Poster
Abstract: Long-form video understanding (LVU) requires models to reason over hour-long narratives while preserving fine-grained spatiotemporal details across extended temporal horizons. Although recent multimodal large language models (MLLMs) have achieved strong performance on short-video tasks, their reasoning capabilities degrade substantially as video duration and information density increase. We present VideoSleuth, a narrative-centric agentic framework for LVU built on video- and audio-native MLLMs. VideoSleuth organizes its memory around the four narrative elements-time, place, person, and event-and equips the agent with tools explicitly designed to construct and query these elements within a ReAct-style reasoning loop. Unlike prior LVU agents that rely on dense offline preprocessing or image-only VLMs, VideoSleuth performs perception on demand using a video-audio native perceptual backbone, jointly leveraging visual and auditory cues to support dialogue-based identity grounding, long-range event localization, and fine-grained evidence retrieval. On four standard LVU benchmarks, VideoSleuth consistently outperforms prior on-demand grounding agents and achieves competitive accuracy with the strongest dense-preprocessing baseline, while consuming at least 5.2× fewer end-to-end tokens than DVD-style pipelines spend on offline preprocessing alone.
Authors:
Pengfei He, Zhenwei Dai, Xianfeng Tang, Yue XING, Hui Liu, Jingying Zeng, Qiankun Peng, Shrivats Agrawal, Samarth Varshney, Suhang Wang, Jiliang Tang, Qi HeAbstract: Large Language Model-based Multi-Agent Systems (LLM-MAS) have demonstrated strong capabilities in solving complex tasks but remain vulnerable when agents receive unreliable messages. This vulnerability stems from a fundamental gap: LLM agents treat all incoming messages equally without evaluating their trustworthiness. While some existing studies approach trustworthiness, they focus on a single type of harmfulness rather than analyze it in a holistic approach from multiple trustworthiness perspectives. We address this gap by proposing a comprehensive definition of trustworthiness inspired by human communication theory \citepgrice1975logic. Our definition identifies six orthogonal trust dimensions that provide interpretable measures of trustworthiness. Building on this definition, we introduce the Attention Trust Score (A‑Trust), a lightweight, attention‑based method for evaluating the trustworthiness of messages. We then develop a principled trust management system (TMS) for LLM‑MAS that supports both message‑level and agent‑level trust assessments. Experiments across diverse multi‑agent settings and tasks demonstrate that our TMS significantly improves robustness against malicious inputs.
PaperID: 7390, Poster
Authors: Daeyoung Choi, Gyuejeong Lee
Abstract: Label distribution leakage in federated learning (FL) poses a critical privacy threat, exposing population-level patterns beyond individual data breaches. Building on known correlations between class sample frequencies and classifier weight magnitudes, we demonstrate how adversaries can exploit these relationships to infer label distributions from shared model parameters in FL. To characterize this threat, we propose WLIA (Weight Norm-based Label Distribution Inference Attack), which infers client label distributions by analyzing variations in classifier weight norms across communication rounds. WLIA operates on the classifier layer without requiring auxiliary data, synthetic samples, or meta-classifiers, enabling broader applicability across heterogeneous model architectures. To counter this threat, we introduce RNR (Randomized Weight Norm Regularization), a defense mechanism that disrupts frequency-weight correlations by randomly regularizing weight norms during local training. RNR adds only a penalty term to the cross-entropy loss, making it efficient and easy to integrate. Comprehensive experiments across six datasets demonstrate that WLIA achieves superior inference accuracy compared to existing attacks, while RNR effectively mitigates leakage with minimal impact on utility.
PaperID: 7391, Poster
Authors: Zhongtian He, Shuyang chu, Jingang Shi
Abstract: Remote photoplethysmography (rPPG) enables non-contact physiological measurement from facial videos, but performance often degrades under unseen environments due to severe domain shifts. Continual learning offers a practical paradigm for updating rPPG models over sequential domains. However, it is challenged by two coupled issues: First, adapting to new domains may overwrite previously acquired physiological representations, resulting in the degradation of physiological consistency across domains. Second, drastic environmental variations induce heterogeneous facial video distributions, making the adaptation process unstable and prone to suboptimal convergence. To address these challenges, we propose ApexPhys, a continual rPPG framework that simultaneously anchors physiological invariance and expands domain plasticity. Specifically, we introduce a Backbone Consistency Representation Preservation (BCRP) mechanism, which leverages the Fisher Information Matrix to identify and preserve parameters critical to physiological consistency. To enhance adaptation flexibility under diverse environmental shifts, we further propose a Hierarchical Expansion Strategy (HES) that autonomously perceives layer-wise distribution discrepancies and dynamically expands hierarchical adaptation branches. Additionally, we develop a Prior-Guided Pathfinding (PGP) strategy, which leverages parameter inheritance to guide the adaptation along a stable trajectory. Experiments on five public datasets demonstrate that ApexPhys achieves strong anti-forgetting performance and robust adaptation under continual learning.
PaperID: 7392, Poster
Authors: James Ghawaly
Abstract: Most deep-learning approaches to file fragment classification have been evaluated on datasets with fewer than 100 file types, a small fraction of the format diversity seen in real-world forensic recovery. We introduce , a 1.9 billion-fragment dataset covering 619 file types collected from permissively licensed public sources, and , a compact 9.7M-parameter architecture that processes byte fragments through three parallel encoders with distinct inductive biases: convolutional networks for local byte motifs, statistical features for position-agnostic distributional signals, and a hierarchical transformer for long-range sequential context. An attentive fusion module combines these representations via cross-attention with learned, input-dependent weights. We address Tesserae's severe class imbalance through decoupled representation learning followed by class-balanced classifier retraining. On the 619-class evaluation, MoDiCo achieves 89.12% balanced top-1 accuracy at 8KB and outperforms the strongest baseline by up to 17 percentage points on balanced accuracy at 4KB, with most of the improvement concentrated on rare classes. A single MoDiCo model trained jointly on 512B and 4KB fragments generalizes without retraining to 8KB inputs, where it exceeds its 4KB performance. We release Tesserae, model weights, and training and evaluation code under permissive open-source licenses.
Abstract: Random sampling is a fundamental tool in modern machine learning and numerical linear algebra for reducing the computational cost of large-scale matrix problems. Existing analyses, however, rely primarily on subspace embedding guarantees, which do random oblique projections induced by sampling, which arises ubiquitously in subsampled least squares and fast low-rank approximation methods. Because (pseudo)inversion is even when the underlying sketch is unbiased, thereby introducing hidden bias into downstream least squares and low-rank approximation solutions. In this work, we develop a unified non-asymptotic theory for random oblique projections in high dimensions. We show that standard random sampling schemes generally induce a systematic statistical overlooked by classical subspace embedding-style analyses, and we propose a principled debiasing framework to correct it. We illustrate the power of the theory through two canonical applications. For subsampled least squares, we obtain sharp bias--variance characterizations, reveal previously unrecognized statistical in widely used sampling schemes, and identify when debiasing yields provable improvements. For fast CUR decomposition, we develop a debiased approach with improved approximation accuracy. Numerical experiments further validate our theoretical findings.
PaperID: 7394, Poster
Authors:
Benjamin Doerr, Moaad Elmoutassim, Martin S. Krejca, Guy M KibuluAbstract: Besides domination-based algorithms such as the NSGA-II, the decomposition-based MOEA/D is one of the most popular and successful optimization heuristic for multi-objective problems. It decomposes the problem into several, single-objective subproblems. By solving these in a co-evolutionary manner and archiving non-dominated solutions, it approximates or computes the Pareto front. Interestingly, so far no advantages from the co-evolutionary communication of subproblem solutions were proven. In this first mathematical runtime analysis addressing this research question, we investigate how the mutation-based MOEA/D solves the classic OMM benchmark when solutions are exchanged among subproblems in different distances r. We show, for the first time, that the MOEA/D profits from an exchange of solutions, and that already very small communication radii r are sufficient to maximally profit from this feature of the algorithm. More precisely, for all values r \ge 1 we prove a tight runtime guarantee of \Theta(n^2), where n is the problem size. This is asymptotically faster than the known bound of O(n^2 \log n) guarantee for r=0, which we complement with a matching bound of \Omega(n^2 \log n).
PaperID: 7395, Poster
Authors: Amal Alphonse, Pavel Dvurechenskii, Clemens Sirotenko
Abstract: Second-order methods are provably faster than first-order methods, and their efficient implementations for large-scale optimization problems have attracted significant attention. Yet, optimization problems in ML often have nonsmooth derivatives, which makes the existing convergence rate theory of second-order methods inapplicable. In this paper, we propose a new semismooth Newton method (SSN) that enjoys both global convergence rates and asymptotic superlinear convergence without requiring second-order differentiability. Crucially, our method does not require (generalized) Hessians to be evaluated at each iteration but only periodically, and it reuses stale Hessians otherwise (i.e., it performs lazy Hessian updates), saving compute cost and often leading to significant speedups in time, whilst still maintaining strong global and local convergence rate guarantees. We develop our theory in an infinite-dimensional setting and illustrate it with numerical experiments on matrix factorization and neural networks with Lipschitz constraints.
Authors: Paarth Gulati, Ilya Nemenman
Abstract: Self-supervised methods that learn representations and predict dynamics fully in the latent space, such as JEPA, have been shown to confuse slowly varying noise with the dynamical signals they aim to capture. Specifically, when noise features remain approximately constant within each trajectory, contrastive predictive objectives preferentially encode these features instead of the true latent variables governing the system. The learned representation then becomes dominated by trajectory-specific noise, so downstream performance degrades with noise strength and does not improve even as the number and duration of training trajectories increase. We argue that this failure is a property of the objective itself, shared by a long line of contrastive predictive objectives that sample negatives across trajectories. To illustrate this generality, we study the failure mode and its remedy in two settings: a standard SimCLR-style JEPA on a synthetic moving-dot dataset, and DySIB, a recently introduced method designed for extracting physically interpretable representations of dynamics, on movies of a rigid-body pendulum. When negatives are instead sampled within a single trajectory, the slow noise can no longer distinguish frames within that trajectory, removing the predictive shortcut. Training one encoder simultaneously on many such trajectories then forces it to encode the variables relevant for the dynamics, with longer trajectories yielding better representations even for strong slow noise. Our results point toward principles for designing contrastive predictive objectives in dynamical representation learning, especially for physical systems with noisy experimental observations.
Abstract: Large language models (LLMs) encode rich semantic knowledge that can be useful for supervised learning, but their outputs are unreliable as statistical priors: they may be noisy, misspecified, or hallucinated. Existing LLM-informed learning methods either trust such signals directly, leaving predictions vulnerable to unreliable LLM guidance, or restrict semantic integration to a single model class. We introduce , a validated framework for learning when to trust LLM-derived semantic priors in supervised statistical learning. Statsformer maps LLM-derived feature scores into a family of learner-specific prior-injection mechanisms across a heterogeneous library of linear and nonlinear predictors. It then uses out-of-fold validation to adaptively calibrate the influence of each prior-informed learner, allowing useful semantic information to improve prediction while attenuating weak, misspecified, or adversarial priors. This yields a guardrailed statistical learning system with an oracle-style guarantee: up to statistical error, the final predictor performs no worse than the best convex combination of its in-library candidates, including prior-free learners. Across diverse prediction tasks, informative LLM priors improve performance, while unreliable priors are automatically downweighted. These results position as a reliability-oriented approach to LLM-informed statistical learning: rather than trusting LLM knowledge directly, it validates semantic priors against data before allowing them to influence the final predictor.
PaperID: 7398, Poster
Authors: Yi Feng, Jiaqi Wang, Wenxuan Zhang
Abstract: Humans often say what sounds acceptable rather than what they truly think when facing social pressure, a well-documented phenomenon called Social Desirability Bias (SDB). Trained on human-generated text, LLMs are steeped in the same social fabric that produces these pressures, yet whether and how SDB shapes their responses remains unknown. We present PRESS (Psychologically-grounded Representation Extraction for Social-desirability Steering), the first framework to identify, localise, and steer SDB inside open-weight LLMs. We decompose SDB into four observable and relatively orthogonal mechanisms: \emphAudience, \emphAccountability, \emphSelf-Monitoring, and \emphSocial Norms, and operationalise each as paired High/Low pressure conditions across 9 cultures in cultural values where social pressure bites hardest. Our experiments show that SDB forms a decodable and low-dimensional direction inside LLMs, organised by specific psychological mechanism rather than cultural origin. Steering along it improves all 58 downstream cultural content moderation tasks across three model families (Qwen +5.7%,Gemma +7.9%, Llama +16.3% mean F1 across cultures), and a direction extracted from one culture can transfer to others without re-extraction. Ablation confirms the direction is causally necessary, and evaluation on capability benchmarks further shows PRESS shifts cultural stance without disrupting factual knowledge.We hope this work makes social pressure in LLMs more transparent and opens new directions toward more honest and socially aware language models.
PaperID: 7399, Poster
Abstract: Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages---prior-training, in-training, post-training, and inference-time---and share one of two underlying strategies: either (by repairing model weights or gating inputs after a fully backdoored model has formed). We propose a third strategy, : allow backdoor formation during training but route it into a designated, quarantined component that can be disabled at deployment. To this end, we propose hutdown (QES), a computationally efficient containment strategy built in a regularization-steered MoE-like setting. Specifically, given a poisoned dataset, QES augments a Transformer-based language model with routed expert-specific LoRA branches and lightweight routers, and uses auxiliary routing objectives to attract trigger-conditioned behavior into a designated expert while preserving benign capability elsewhere. At deployment, mitigation reduces to a single constant-time operation: zeroing the quarantined expert's routing weight, without trigger screening or further updating model weights. Empirically, our methods reduce the attack success rate (ASR) from 100% to 0-10% on most settings across two tasks, three attacks, and four model families, while downstream utility is often preserved or only modestly affected. These results establish
PaperID: 7400, Poster
Authors: PRAKUL S HIREMATH, PeerAhammad M Bagawan, Sahil Bhekane
Abstract: For K=2 regimes, scalar thresholding of the Bayesian posterior is Bayes-optimal. For K \geq 3, the posterior evolves on a (K-1)-dimensional simplex \Delta^K-1, yet virtually all deployed detection systems reduce it to a scalar score without theoretical justification. We resolve this question completely. We introduce Extended Decision Sufficiency (EDS)—three linear constraints on emission densities and transition dynamics (Rank-One Emissions, Markov Factorisation, Normal Factorisation), verifiable in O(K^2) operations from model parameters alone—and prove the following exact characterisation: Main result. EDS holds if and only if the Bayes-optimal stopping value function is constant on every level set of a scalar statistic \phi : \Delta^K-1 \to \mathbbR. Sufficiency follows from algebraic closure of the Bellman operator \mathcalB under EDS. Necessity is proved by a coupling-based fixed-point contradiction: for each failure mode we exhibit an explicit belief pair with equal \phi-values but a strictly positive total-variation gap on predictive densities, which propagates to a value-function gap via the Lipschitz bound on \mathcalB and the \rho-contraction of \mathcalT, without assuming any structure on V beyond continuity. Three consequences are sharp and quantitative: (i) When EDS fails, every scalar rule—regardless of architecture, capacity, or training data—incurs a strictly positive, computable, algorithm-independent lead-time loss. (ii) When EDS holds, optimal expected lead time is given in closed form by a hitting-time functional of a scalar Markov chain. (iii) Under approximate EDS with deviation \eta, performance degrades at rate O(\eta/(1-\rho)), making \eta a practical model-diagnostic with direct performance implications. Across 72 synthetic configurations and the CICIDS2017 benchmark, predicted and observed lead times agree within 0.4 steps (MAE). A controlled experiment with matched parameter counts confirms that the lead-time advantage is structural, not a regularisation artefact, thereby validating the theory as the sole causal mechanism.
Abstract: Geometric foundation models show promise in 3D reconstruction, yet their progress is severely constrained by the scarcity of diverse, large-scale 3D annotations. While Internet videos offer virtually unlimited raw data, utilizing them as a scaling source for geometric learning is challenging due to the absence of ground-truth geometry and the presence of observational noise. To address this, we propose SAGE, a framework for Scalable Adaptation of GEometric foundation models from raw video streams. SAGE leverages a hierarchical mining pipeline to transform videos into training trajectories and hybrid supervision: (1) Informative training trajectory selection; (2) Sparse Geometric Anchoring via SfM point clouds for global structural guidance; and (3) Dense Differentiable Consistency via 3D Gaussian rendering for multi-view constraints. To prevent catastrophic forgetting, we introduce a regularization strategy using anchor data. Experiments on applying SAGE to two architecturally distinct backbones (including both MV-DUSt3R and VGGT) show that Chamfer Distance is consistently reduced by 20-42% on unseen benchmarks (7Scenes, TUM-RGBD, Matterport3D). SAGE establishes Internet video as a viable and scalable adaptation source for geometric foundation models.
Abstract: Dropout is common in clinical studies, with up to half of patients leaving early due to side effects or other reasons. When dropout is informative (i.e., dependent on survival time), it introduces censoring bias, because of which treatment effect estimates are also biased. In this paper, we propose an assumption-lean framework to assess the robustness of conditional average treatment effect (CATE) estimates in survival analysis when facing censoring bias. Unlike existing works that rely on strong assumptions, such as non-informative censoring, to obtain point estimation, we use partial identification to derive informative bounds on the CATE. Thereby, our framework helps to identify patient subgroups where treatment is effective despite informative censoring. We further propose a novel model-agnostic meta-learner, called , to estimate the bounds that can be used in combination with arbitrary machine-learning models, and that has favorable theoretical properties such as double-robustness and quasi-oracle efficiency. We finally demonstrate the effectiveness of our meta-learner across various experiments using both simulated and real-world data.
PaperID: 7403, Poster
Abstract: Estimating semantic correspondences between different object instances of similar categories in real-world images is a fundamental challenge in computer vision. While the recent progress in foundation models has significantly advanced solutions for such difficult image correspondence problems, respective methods are often insufficiently regularised. For example, common nearest-neighbour matching with foundation model features often results in noisy correspondences, which is particularly prominent when solving for dense correspondences. In this work, we tackle image matching via functional maps, which have been popularised for 3D shape matching due to their powerful spectral formalism utilising an efficient low-dimensional linear representation. However, functional maps require that domains (in our cases the rectangular image grids) are equipped with an informative and non-trivial geometry. To define a meaningful manifold structure for images, we use Galerkin's method from finite elements to derive a discrete Laplace-Beltrami operator (LBO) based on a manifold embedding of deep image features. Further, we leverage a neural adjoint map that allows to represent non-linear mappings. We demonstrate that our method sets the new state of the art among zero-shot image correspondence methods on multiple dense correspondence benchmarks of real-world images.
Authors: Won J Cho, Daeky Jeong, Hyeongyeol Lim, Hongjun Yoon
Abstract: Synthetic histopathology image generation addresses critical challenges in computational pathology, including patient privacy and the growing need for large-scale training data for foundation models. Latent diffusion models have dominated the image generation domain, with recent works emphasizing that the choice of latent space is critical to the quality of generated images. Existing state-of-the-art generative models in histopathology use pretrained Vision Foundation Models (VFMs) as conditioning signals, and we observe that this leads to ``conditioning collapse'', where the conditioning signal dominates the latent space and lowers the quality and diversity of generated samples. Therefore, we instead use pretrained histopathology VFMs as the latent space itself, leveraging their patch-token features that encode rich semantic information. We empirically show that these features are \ell_2-normalized and lie on the unit hypersphere \mathcalS^d-1 with strong angular dominance and intrinsic curvature, making them naturally suited for a Riemannian formulation. We therefore present STREAM, the first framework to apply Riemannian flow matching in the pathology domain. STREAM consists of two stages: 1) a bridge-type stochastic perturbation that establishes per-token rectifiability on \mathcalS^d-1 for training a Diffusion Transformer (DiT) in latent space, and 2) a novel anisotropic decoder that allocates robustness to data-sparse directions while preserving fidelity along data-dense ones. Together, STREAM achieves state-of-the-art reconstruction and generation performance on breast and colorectal cancer datasets.
PaperID: 7405, Poster
Abstract: In autonomous driving, multi-agent collaboration enhances individual perception capabilities through information sharing. However, in real-world applications, differences in models and training data among heterogeneous agents inevitably lead to domain gaps between shared features, bringing challenges to collaboration. To address this, existing methods typically rely on a one-to-one adaptation paradigm, necessitating specific training or fine-tuning for every new agent. This results in a large accumulation of adapters on the vehicle, which is unsustainable for resource-constrained edge devices. To overcome these limitations, we propose UniMTO, a Unified adaptation framework that enables Many-To-One adaptation across heterogeneous agents using a single, fixed adapter. Specifically, it treats domain adaptation as a pattern transformation task, disentangling domain-invariant content and domain-specific pattern from features, and achieving source-to-target pattern transformation through recombination based on a Mixture-of-Experts (MoE) mechanism. Extensive experiments on V2V4Real and V2X-Real demonstrate that UniMTO achieves superior generalization, and exhibits strong zero-shot adaptation capabilities, outperforming all state-of-the-art methods, enabling many-to-one domain adaptation with a single, fixed adapter.
PaperID: 7406, Poster
Abstract: World models are central to model-based reinforcement learning and planning, but their latent dimensions, architectural spectra, and downstream behavior remain poorly understood. In this work, we develop a spectral theory of linear world models in three settings. First, for shared multi-horizon prediction in linear dynamical systems, we identify the finite-horizon block-Krylov operator \mathcal K_H as the exact latent bottleneck and give a perturbation condition for recovering its empirical rank elbow. Second, for stable diagonal state-space models, we prove a Rademacher complexity bound governed by the spectral memory energy \sum_j L_j(T)^2, rather than the parameter count alone. Third, for jointly trained diagonal latent dynamics, we show that disjoint factor spectra imply factor-aligned latent coordinates, with a Riesz-projector stability extension for approximate recovery and a local gradient-flow companion. The supporting results connect these spectral quantities to compositional OOD prediction, rollout envelopes, and prediction-to-control mismatch. Further, generated-system experiments verify the predicted spectral behavior and illustrate why pixel prediction and control returns can rank world models differently.
PaperID: 7407, Poster
Abstract: Generating accurate radiology reports from 3D medical volumes remains prone to hallucination. Existing grounding strategies often depend on external visual foundation models or on fixed medical phrase queues. The former passes segmenter-derived evidence to generation, where segmentation errors are hard to suppress. The latter has limited coverage of long-tail and compositional abnormalities. We therefore propose HeatRAD, a glance-before-tell framework that first routes suspicious slice evidence into an organ-ordered visual prefix, then generates reports from this localized evidence without explicit grounding supervision. HeatRAD uses a slice-centric signed heat-conduction encoder that performs efficient spectral mixing while separating smooth anatomy from lesion-like residues. \method then scores suspicious slices, binds their evidence to coarse organ contexts, and assembles an organ-ordered visual prefix, guiding the decoder toward localized abnormalities instead of redundant volumetric context. Experiments on three 3D radiology benchmarks show consistent gains in clinical fidelity and hallucination-related metrics, supporting anomaly-first evidence selection as an annotation-efficient paradigm for 3D report generation.
Abstract: Real-world clinical diagnosis is a complex process in which the doctor is required to obtain information from both interaction with the patient and conducting medical exams. Additionally, the doctor needs to adapt to different patient personas, as well as noisy and incomplete information that can happen at any time during the process. However, existing benchmarks for medical LLMs and methods for automatic diagnosis largely simplify this process by reducing it to single-turn question answering, noise-free conversations, or sequential exam making, etc., ignoring the interactive and uncertain nature of clinical diagnosis. In this paper, we aim to address this gap by formalizing clinical diagnosis as a Partially Observable Markov Decision Process (POMDP) with three action types: questioning the patient, ordering medical exams as tool calls, and issuing a diagnosis. We also introduce a systematic noise model comprising seven patient noise types and three exam noise types. Using our proposed environment, we train an effective diagnosis agent, MedExAgent, through a two-stage pipeline that first performs supervised finetuning on synthetic conversations structured after the Calgary-Cambridge model for clinical interviews, and then applies DAPO to optimize a composite reward capturing diagnostic accuracy, tool call quality, and exam cost including financial cost and patient discomfort. Through extensive experiments and ablation studies, we demonstrate that MedExAgent achieves diagnostic performance comparable to larger models while maintaining cost-efficient examination strategies.
PaperID: 7409, Poster
Abstract: Continual tuning is essential for adapting multimodal large language models to evolving tasks and domains, yet existing methods often assume that every image-present example provides valid supervision for the visual pathway. We address the resulting gap: how to learn from all incoming multimodal instructions while preventing visually unnecessary supervision from drifting the visual interface. We propose Visual-Necessity-Gated Continual Tuning (VNG-CT), a path-aware framework that estimates sample-level dependence on visual evidence by comparing target likelihoods under the original image and a counterfactual null image. The resulting gate routes gradients to different parameter paths: all examples update the language path, visually necessary examples update the visual path, and low-necessity examples are absorbed by a lightweight calibration path that promotes invariance to irrelevant visual evidence. Across CoIN, MLLM-CL, and UCIT-style continual instruction streams, VNG-CT improves final and average performance while reducing visual forgetting. These results suggest that visual necessity offers a practical principle for stabilizing continual multimodal adaptation without discarding useful language-side supervision.
Abstract: Embodied agents are expected to assist humans by actively exploring unknown environments and reasoning about spatial contexts. When deployed in real life, agents often face sequential tasks where each new task follows the completion of the previous one and may include infeasible objectives, such as searching for non-existent objects. However, most existing research focuses on isolated goals, overlooking the core challenge of sequential tasks: the ability to reuse spatial knowledge accumulated from previous explorations to guide subsequent reasoning and exploration. In this work, we investigate this underexplored yet practically significant embodied AI challenge. Specifically, we propose 3DSPMR, a 3D SPatial Memory Reasoning framework that utilizes Field-of-View (FoV) coverage as an explicit geometric prior. By integrating FoV-based constraints, 3DSPMR significantly enhances an agent’s memory, reasoning, and exploration capabilities across sequential tasks. To facilitate research in this area, we further introduce SEER-Bench, a novel Sequential Embodied Exploration and Reasoning Benchmark that spans two foundational tasks: Embodied Question Answering (EQA) and Embodied Multi-modal Navigation (EMN). SEER-Bench uniquely incorporates both feasible and infeasible tasks to provide a rigorous and comprehensive evaluation of agent performance. Extensive experiments verify that 3DSPMR achieves substantial performance gains on both sequential EQA and EMN tasks.
Abstract: We study the problem of learning to bid when the bidder’s value is dynamic, i.e., when the current value depends on past outcomes. Specifically, we consider a bidder participating in repeated second-price auctions whose value depends on the time elapsed since their last successful bid, with auctions arriving in continuous time and only aggregated feedback revealed at the end of the horizon. Such a bidder must (1) balance the immediate benefit of winning the current auction against its impact on future values and (2) learn unknown environmental parameters. We derive regret bounds for a class of learning methods that combine plug-in estimators with a differential-equation characterization of the optimal policy, and show that a specific confidence bound algorithm learns the optimal policy with a near optimal regret of \tilde\mathcalO(\log N) for piecewise linear primitives, and \tilde\mathcalO(N^1/3) for general, smooth primitives, achieving these regrets without explicit randomization. These theoretical results are supported by numerical experiments.
Abstract: Constructing mechanistic models of neural circuits is a fundamental goal of neuroscience, yet verifying such models is limited by the lack of ground truth. To rigorously test model discovery, we establish an in silico testbed using neuromechanical simulations of a larval zebrafish as a transparent ground truth. We find that LLM-based tree search autonomously discovers predictive models that significantly outperform established forecasting baselines. However, in-distribution accuracy is a poor proxy for faithful system identification, as models exploit statistical shortcuts. Structural priors prove essential for enabling robust out-of-distribution generalization and recovery of interpretable mechanistic models. Our insights provide guidance for modeling real-world neural recordings and offer a broader template for AI-driven scientific discovery.
PaperID: 7413, Poster
Abstract: Therapeutic antibody screening requires prioritizing antibody candidates across specificity, affinity, and developability. Current pipelines typically silo these objectives and underuse structural information that governs antibody-antigen recognition. To address these inefficacies, we introduce FoldAbS, a unified supervised screening framework that reconceptualizes protein folding models as frozen foundation encoders, exploiting their internal representations to capture rich sequence, geometric, and interaction dynamics well beyond basic structure prediction. By coupling the folding model with lightweight, task-specific heads, FoldAbS seamlessly evaluates all three therapeutic criteria within a single architecture. Instantiating this framework with the open-source Protenix model (FoldAbS-Px) yields consistent improvements over baseline approaches, achieving significant performance gains of up to 13.1%, 9.2%, and 6.2% in specificity, affinity, and developability, respectively. Collectively, these results reposition protein folding models from structure predictors and confidence scorers into reusable foundation encoders for comprehensive and supervised antibody screening.
PaperID: 7414, Poster
Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in Information Extraction (IE) tasks; however, their performance in zero-shot Named Entity Recognition (NER) remains suboptimal. We argue that this limitation stems not from intrinsic defects within the models, but rather from the failure of existing instruction paradigms to effectively elicit their latent capabilities. Drawing inspiration from cognitive science, we propose an instruction optimization framework termed Instruction Semantic Elaboration (ISE), which fully elicits models' latent NER capabilities by mimicking the semantic elaboration process—specifically by contrasting related concepts, decomposing constituent elements, and illustrating with concrete examples. We evaluate our framework under a zero-shot setting across four domain-specific datasets. Results demonstrate that ISE effectively and stably unlocks the intrinsic NER capabilities of LLMs, achieving a 10.99% improvement in F1 score over initial instructions and outperforming strong baselines by 6.89%. Furthermore, the framework generalizes well across model architectures and parameter scales, while remaining robust to the quality of initial instructions. This study offers a novel perspective on activating NER capabilities in LLMs, facilitating their efficient and stable deployment in real-world scenarios.
PaperID: 7415, Poster
Abstract: Studies of human reasoning have shown that people are typically stronger at evaluating reasoning than producing it from scratch. In contrast, large reasoning models (LRMs) are trained to excel at producing long chains of reasoning to solve complex problems. How then do LRMs perform at evaluating reasons? We investigate this with the Valid-Answer-Invalid-Reasoning (VAIR) dataset: math problems and solutions with trivial reasoning flaws but valid answers, designed to isolate reasoning evaluation from the confound of reasoning production. Unlike humans, who we find are only 6% worse at grading than solving such problems, we find a substantial production-evaluation gap in LRMs: frontier models score as low as 48% when evaluating VAIR solutions, despite near-perfect solution production. Why this enigma? Through chain-of-thought (CoT) analysis, we find evidence of an answer confirmation bias: LRMs often produce then check for the correct answer instead of carefully verifying each step, fabricating rationalizations even when noticing anomalous reasoning. Linear probes corroborate this, showing that while LRM activations encode some representation of valid reasoning, they fail to robustly represent VAIR solutions as invalid. Causal patching of the final answer's representations causes LRM verdicts and activations to flip, demonstrating that answer validity is responsible for models' confirmation biases. These findings indicate an outstanding limitation in dominant approaches to reasoning training, which incentivize LRMs to produce and confirm reasoning towards correct answers, but not to robustly evaluate the underlying reasons.
PaperID: 7416, Poster
Abstract: We study fair multi-armed bandits under the Nash Social Welfare (NSW) objective, which measures performance via the geometric mean of accumulated rewards as a fairness-respecting performance metric. Existing formulations define Nash regret as NR_T = \mu^\star - (\prod_t=1^T \mathbbE\mu_I_t)^1/T, where \mu_I_t is the mean reward of the recommended arm I_t and T is the learning horizon. We observe that, since ensemble Nash regret operates on per-round marginal expectations before applying the geometric mean, it does not capture the joint distribution of rewards across rounds — leaving an aspect of the NSW fairness motivation unaddressed at the trajectory level. To address this, we propose \emphtrajectory-wise Nash regret \widetildeNR_T= \mu^\star - \mathbbE[(\prod_t=1^T \mu_I_t)^1/T], where the geometric mean is computed over complete sample paths before taking expectations. This formulation captures the fairness properties of NSW more faithfully and requires controlling the joint distribution of rewards across rounds, rather than merely per-round marginals. By Jensen's inequality applied to the concave geometric mean, \widetildeNR_T \geq NR_T, making it a strictly stronger metric. We additionally introduce \emphhigh probability Nash regret \widehatNR_T = \mu^\star - (\prod_t \mu_I_t)^1/T, obtaining the first high probability regret bounds in the fair bandits literature. We propose a two-phase algorithm, Round Robin Nash Confidence Bound (\textttRR-NCB), combining structured round robin exploration with a Nash confidence bound index policy. We establish that \widetildeNR_T \leq \widetilde\mathcalO(\sqrtk\log T/T) and \widehatNR_T \leq \widetilde\mathcalO(\sqrtk\log(kT/\delta)/T) with probability 1-\delta, matching the optimal \widetilde\mathcalO(\sqrtk/T) scaling despite operating under strictly stronger metrics. Optimality follows from a lower bound chain via AM-GM and standard k-armed bandit minimax arguments. We validate our theoretical findings through simulations.
PaperID: 7417, Poster
Abstract: Unlike real-world datasets, utilizing synthetic data from generative models as train dataset necessitates a rigorous consideration of the training data's informativeness. In this paper, we propose a novel framework designed to steer diffusion models based on Bayesian epistemic uncertainty to train the target model. To ensure computational tractability, we estimate the Bayesian Active Learning by Disagreement (BALD) score by applying the Laplace approximation specifically to the last-layer parameters of the downstream model. Furthermore, we incorporate Feynman-Kac Steering (FKS) for diffusion to enhance the structural diversity of the synthetic samples and mitigate the emergence of undesirable artifacts. Extensive experiments on Imagenette and ImageNet-100 for the classification task and MS-COCO for the object detection task demonstrate that our method significantly outperforms existing baselines in downstream accuracy and OOD generalization. Furthermore, our analysis of pairwise distances between inter-class and intra-class confirms that our method effectively expands the downstream model's knowledge manifold by generating diverse and informative training signals.
PaperID: 7418, Poster
Authors: Chao Yan
Abstract: We present the first nearly optimal differentially private PAC learner for any concept class with VC dimension 1 and Littlestone dimension d. Our algorithm achieves the sample complexity of \tildeO_\varepsilon,\delta,\alpha,\beta(\log^d), nearly matching the lower bound of \Omega(\log^d) proved by Alon et al. [STOC19]. Prior to our work, the best known upper bound is \tildeO(VC\cdot d^5) for general VC classes, as shown by Ghazi et al. [STOC21]. The main idea is to combine the tree structure of VC-dimension-1 classes with the private median primitive for interior-point selection. We first privately locate a good depth in the tree, then privately identify a good node on that layer. For proper learning, we show how to descend from the improper output to a proper leaf while controlling the additional false positives.
PaperID: 7419, Poster
Abstract: In-context learning (ICL) allows large language models (LLMs) to adapt to new tasks from a few examples without parameter updates. Previous work and our empirical studies suggest two modes in ICL: Task Recognition, which applies pre-trained knowledge to familiar tasks, and Task Learning, which generalizes to novel mappings during inference. However, a unified theoretical account of these behaviors remains lacking. We propose a new Bayesian prediction framework that models pre-training data as a Pitman–Yor mixture over latent tasks, which captures the growing and heavy-tailed structure of natural language. By formulating an information-theoretic optimization problem, we derive the structure of the optimal ICL predictor under explicit capacity constraints. Our analysis reveals a phase transition in optimal capacity allocation: a water-filling strategy that prioritizes high-frequency tasks while assigning zero task-specific information to rare ones. This gives rise to two regimes: for frequent tasks, the model performs Task Recognition, yielding exponentially fast error decay with prompt length. For rare or unseen tasks, the model performs Task Learning by executing an implicit inference algorithm, resulting in a slower power-law convergence rate. Our results provide a principled explanation for the dual nature of ICL and establish a direct connection between pre-training data distributions, model capacity, and in-context generalization behavior.
Abstract: The "Locate-then-Update" paradigm has become a predominant approach in the post-training of large language models (LLMs), identifying critical components via mechanistic interpretability for targeted parameter updates. However, this paradigm rests on a fundamental yet unverified assumption: can mechanisms derived from current static parameters reliably guide future dynamic parameter updates? To investigate this, we systematically track the structural evolution of Transformer circuits throughout the supervised fine-tuning (SFT) process, revealing the underlying dynamics of task mechanisms. We introduce three novel metrics—Circuit Distance, Circuit Stability, and Circuit Conflict—to analyze circuit evolution across three dimensions: neural migration, semantic stability, and cross-task interference. Our empirical results reveal that circuits inherently exhibit "Free Evolution" during parameter updates. Consequently, static mechanisms extracted from current states inevitably suffer from temporal latency, making them fundamentally inadequate for guiding future states. Moreover, by deconstructing the "illusion of effectiveness" in existing methods, this work underscores the necessity of "foresight" in mechanistic localization and proposes a predictive framework for future research.
PaperID: 7421, Poster
Abstract: Diffusion Transformers (DiTs) have emerged as the state-of-the-art architecture in high-fidelity generative modeling, but their massive computational demands and multi-step inference process limit their deployment in practical scenarios. Post-training quantization (PTQ) offers a promising solution for accelerating inference, but DiTs suffer from severe activation outlier issues, which can lead to catastrophic accuracy loss in integer formats. Although some recent outlier mitigation methods employ orthogonal transformations or diagonal scaling to smooth the activation distribution, they either solve the transformation matrix solely for activations or fail to fully balance the inter-channel variance between activations and weights. Additionally, the construction of the calibration set by taking the weighted average of the activation values for each time step introduces spurious temporal correlations. To address these issues, we propose a joint diagonalization of the activation-weight covariance matrix based on the solution of a generalized eigenvalue problem (GEVP). We also transform the feature fusion of the calibration set's temporal dimension into probabilistic sampling and assembly in the spatial dimension using Uniform Token Splicing (UTS). Specifically, compared to suboptimal baseline methods, our approach reduces the FID by 16.06 and 6.49 on both the W3A4-quantized DiT-XL/2 and PixArt-\Sigma models, respectively. This demonstrates that our method successfully suppresses outliers in activations and weights, maintaining DiT's performance under low-bit quantization conditions. Code is available at the anonymous repository \urlhttps://anonymous.4open.science/r/JDUS-DiT-74E1.
PaperID: 7422, Poster
Abstract: Stochastic Gradient Descent (SGD) is a workhorse algorithm for continuous optimization. It is a folklore belief among deep learning practitioners that learning rate decay in SGD substantially improves performance. However, it has remained unclear whether learning rate decay offers an asymptotic advantage for \emphglobal nonconvex optimization. We answer this question in the affirmative by demonstrating a natural class of nonconvex functions for which we prove that SGD with learning rate decay requires \emphexponentially fewer gradient queries for global optimization than SGD with any fixed learning rate. To the best of our knowledge, this is the first exponential advantage that has been demonstrated for learning rate decay. The class of functions we consider includes many popular benchmark nonconvex functions. Our results provide a robust explanatory theory for the empirically observed benefits of learning rate decay. Our technical results are built upon a new discretization analysis of SGD with decaying step size, that allows us to establish a clean correspondence with continuous-time annealing of Langevin diffusions. We then show that this annealed diffusion performs the non-logconcave sampling task associated to the optimization problem in polynomial time, via a novel application of weak Poincar\'e inequalities.
Abstract: Mean-field control (MFC) offers a scalable solution to the curse of dimensionality in multi-agent systems but traditionally hinges on the restrictive assumption of exchangeability via dense, all-to-all interactions. In this work, we bridge the gap to real-world network structures by proposing a rigorous framework for MFC on large sparse graphs. We redefine the system state as a probability measure over decorated rooted neighborhoods, effectively capturing local heterogeneity. Our central contribution is a theoretical foundation for scalable reinforcement learning in this setting. We prove horizon-dependent locality: for finite-horizon problems, an agent's optimal policy at time t depends strictly on its (T-t)-hop neighborhood. This result renders the infinite-dimensional control problem tractable and underpins a novel Dynamic Programming Principle (DPP) on the lifted space of neighborhood distributions. Furthermore, we formally and experimentally justify the use of Graph Neural Networks (GNNs) for actor-critic algorithms in this context. Our framework naturally recovers classical MFC as a degenerate case while enabling efficient, theoretically grounded control on complex sparse topologies.
Abstract: Large language models are increasingly deployed as long-lived agents that must adapt across users, tasks, domains, modalities, and feedback regimes without access to model weights. Existing black-box adaptation methods typically optimize a single prompt, maintain an undifferentiated memory, or rely on repeated rollout-heavy search. However, these designs struggle when streams of input are nonstationary, feedback is sparse, and failures from one task family can contaminate behavior on another. We introduce RIZZ ( ones), a continual adaptation framework for compound language-model systems that learns entirely through verifier-gated memory, routing, and prompt compilation. RIZZ organizes input streams into dynamically spawned memory branches. At inference time, either while online or offline, a context-aware router selects or creates a branch that retrieves branch-local, global, graph-structured, and working-memory context, which is compiled into a bounded prompt together with retrieved task evidence. After the model acts, task verifiers score the output, and only verified interactions can update memory, promote reusable rules, demote harmful rules, or create anti-patterns. This yields a black-box agent that improves through persistent natural-language feedback while explicitly controlling interference. RIZZ targets the regime where adaptation must occur online under context budgets. Finally, we demonstrate the effectiveness of our framework against state-of-the-art baselines on competitive benchmarks.
PaperID: 7425, Poster
Authors: Georgia Dimaki, Martin Villanueva Dimakis
Abstract: Recent work has shown that large language models form linear representations of emotion concepts that causally influence alignment-relevant behavior, including blackmail, reward hacking, and sycophancy. These findings raise a natural question: can such representations be used not only to characterize model behavior in controlled evaluations, but also as practical tools for monitoring and guiding models during deployment? We explore this possibility, investigating how emotion-vector activations can serve as interpretable signals of a model's internal state over the course of a conversation or agentic task. We discuss how these signals might support debugging, early detection of drift toward misaligned behavior, and lightweight steering interventions to nudge the model back toward desired behavior. Beyond the Assistant, the same representations apply to other characters in the model's context — including the user — suggesting a broader family of applications in which internal representations of emotional state inform how systems respond. We outline the opportunities and limitations of this approach and highlight open questions around validation, robustness, and responsible use.
Abstract: Clinical events captured in Electronic Health Records (EHR) are irregularly sampled and may consist of a mixture of discrete events and numerical measurements, such as laboratory values or treatment dosages. The sequential nature of EHR, analogous to natural language, has motivated the use of next-token prediction to train prior EHR Foundation Models (FMs) over events. However, this training fails to capture the full structure of EHR. When a given event occurs must be captured, but the event value (abnormal lab) also modulates the likelihood of other clinical events. Most existing EHR FMs do not jointly model this likelihood and are unable to capture the full observation process, impacting downstream capabilities. We propose ORA, a marked time-to-event pretraining objective that jointly models event timing and associated measurements. Across multiple datasets, downstream tasks, and model architectures, this objective consistently yields more generalizable representations than next-token prediction and pretraining losses that ignore continuous measurements. Importantly, the proposed objective yields improvements beyond traditional classification evaluation, including better regression and time-to-event prediction. Beyond introducing a new family of FMs, our ablations suggest a broader takeaway: pretraining objectives that account for EHR structure are critical for expanding downstream capabilities and generalizability.
PaperID: 7427, Poster
Authors:
Yongji Fu, Yi Zhou, Gaojie Jin, Guanqun CaoAbstract: Scaling neural solvers to large Vehicle Routing Problems (VRPs) runs into two recurring difficulties. To stay within compute budgets, scalable solvers tend to rely on partitioning, local candidate restriction, or staged decisions, which limit the global structure that ultimately drives solution quality. To improve the final solution, many neural pipelines fall back on a classical search, but use it as a one-shot post-processor. The model predicts, search repairs, and no sustained feedback is formed between them. We propose LSSL (Learning to Search, Searching to Learn), a closed-loop learning-search framework for large-scale VRP that addresses both difficulties in one framework. LSSL predicts a search-friendly guidance heatmap on a sparse candidate graph. A classical heuristic backend then uses this heatmap to refine the incumbent solution, and the structural state returned by the backend is fed back to recondition the model. As a result, learning and search interact within a closed loop, rather than operating as two detached stages. Built on a sparse diffusion transformer over an O(NK) candidate graph, the loop scales to ultra-large instances on a single consumer-grade GPU without hard decomposition. Experiments on large-scale TSP and CVRP benchmarks demonstrate that LSSL achieves superior scalability, efficiency, and solution quality. Ablation studies further reveal that the closed learning–search loop accounts for a substantial portion of the observed performance gains.
PaperID: 7428, Poster
Authors: Hoan M Tran, Xin Wang, Xuechen Liu
Abstract: Speech foundation models pretrained via self-supervised learning provide powerful representations for deepfake detection, yet we find that standard fine-tuning exploits only the deepest transformer layers — leaving the majority of the network's capacity discriminatively inert. Layer-wise probing of XLS-R-128 reveals that the top two-thirds of the transformer stack contribute negligibly to anti-spoofing performance after standard fine-tuning, a phenomenon we term shallow-layer dormancy. We connect this observation to the representational redundancy of pretrained transformers: adjacent layers encode highly similar information, and fine-tuning without explicit intermediate supervision reinforces rather than breaks this redundancy. To address this, we propose Progressive Layer-wise Supervision (PLS), a training framework that attaches a shared linear classifier to every transformer layer and schedules layer-wise loss contributions via an exponentially annealed power weighting. Early training concentrates supervision on deep layers to preserve the pretrained representational hierarchy; later training progressively redistributes supervision to shallower layers, converting dormant earlier representations into discriminative ones. Competitive detection (\leq9.25% average EER) is achievable on \emphXLS-R-128 at layer 7, which reduces the parameter count from 315M to 101M and exemplifies an early exit mechanism. Experiments across twelve benchmarks spanning in-domain, and out-of-domain deepfake scenarios show PLS achieves a new state-of-the-art with 239M parameters, outperforming prior methods with larger and more complex architectures.
Authors: Aratrika Mustafi, Soumya Mukherjee, Bharath Sriperumbudur
Abstract: We develop a gradient flow on the space of probability measures defined on matrix-valued parameters induced by regularized Muon, an analytically smoothed version of the idealized Muon optimizer. The key observation is that the regularized orthogonalization map is the gradient of a smooth Fenchel-dual smoothing of the nuclear norm. This identifies the (regularized) Muon update as a mirror/prox step in the update variable, with momentum acting as the dual coordinate. We use this structure to lift Muon from a single matrix parameter to finite-particle probability objectives of the form J(\rho)=R\left(\int F d \rho\right), a setting motivated by mean-field descriptions of neural-network training, and derive the inertial continuous-time limit. Using this structure, we derive the finite-particle continuous-time limit under the inertial scaling of step size and momentum, and then pass to a phase-space mean-field equation over probability laws on parameter-momentum pairs. The resulting flow can be shown to be a damped Hamiltonian probability dynamics whose kinetic energy is induced by the regularized Muon mirror potential. We prove an exact Hamiltonian dissipation identity, showing that the Hamiltonian energy decreases monotonically- leading to the descent of the target objective functional. Under appropriate gradient dominance assumptions, we obtain continuous and discrete time exponential convergence rates. We also study the well-posedness of the mean field limit equation and establish propagation of chaos guarantees for the interacting particle system. Finally, we extend the formulation to Hilbert-valued feature maps on product matrix spaces, yielding a blockwise Muon probability flow applicable to smooth transformer mixture-of-experts models.
Abstract: Generalized planning aims to learn policies that generalize across large collections of instances within a classical planning domain. Recent approaches using Graph Neural Networks (GNNs) have shown the ability to learn nearly perfect policies for several domains. This work improves on the recently published idea of Iterated Width (IW) policies. Therein, the policy broadens its successor scope through an IW-lookahead search that offers to "jump" over multiple transitions, simplifying the problem structure. Yet, each transition is evaluated individually leading to unscalable compute and expressivity limitations. Furthermore, while an IW(1) search is attractive due to scaling linearly with the number of atoms in a problem, it still becomes inefficient once thousands of objects are considered like in the International Planning Competition (IPC) 2023 benchmark. In this work, we address both limitations. Firstly, we introduce a vastly more efficient holistic encoding of the entire search tree. It jointly represents IW(1)-reachable states only by their relational differences to the current state, which enables Relational GNNs (R-GNNs) to score all transitions in a single forward pass. Secondly, we define Abstracted IW(1) to improve scaling through relational abstraction during novelty checks. Rather than testing fully instantiated atoms, it abstracts each atom by replacing all but one of its arguments with their types. The original atom is then novel if any of its abstracted forms is. This structural compression shifts the scaling of the novelty search to be linear in the number of objects, rather than atoms, still capturing meaningful subgoal structure. Our contributions are evaluated on the hyperscaling IPC 2023 benchmark and across a broad range of domains, including those that require features beyond the C_2 logic fragment. The results show that our policies achieve a new state-of-the-art performance, significantly surpassing prior work, including classical planners like LAMA.
Abstract: While large language models (LLMs) can solve advanced reasoning problems in seconds, we show that even frontier models fail to perform a much simpler operation: exactly copying an input string that lies well within their context windows. We attribute this failure to positional encodings in Transformer architectures, whose inductive bias favors copying through a shortcut based on matching local contexts rather than carefully locating the corresponding input positions. To address this issue, we introduce 2D-RoPE, which organizes text into a 2D grid rather than a 1D sequence and assigns each token a row ID and a column ID. Under this view, copying becomes simply retrieving input tokens from the same column, which makes the task easy to learn. In synthetic copy experiments, shallow Transformers with 2D-RoPE achieve perfect copying at input lengths hundreds of times longer than those seen during training, whereas standard positional encodings fall far behind. We further pretrain 2D-RoPE language models on DCLM at scales up to 1.4B parameters and show that 2D-RoPE substantially improves performance on copy and NIAH tasks. Overall, our results suggest that viewing text in 2D can benefit language modeling, and we hope this encourages future work to further explore the potential of 2D positional encodings.
Abstract: Humans systematically misrepresent probability in a stereotyped inverse-S pattern. It has been documented for decades, but its origin remains unexplained. We propose a Bayesian encoding–decoding account in which probabilities are represented by noisy internal signals and decoded by Bayes-risk minimization. For bounded probability stimuli, we show that distortion decomposes into boundary regression, likelihood repulsion, and prior attraction, yielding a key prediction: the classic inverse-S-shaped weighting pattern implies a U-shaped allocation of encoding precision with greater sensitivity near 0 and 1. Across judgment of relative frequency, lottery pricing, and risky choice, this U-shape is recovered from data without imposing any functional form on the encoding, and our framework outperforms deterministic weighting functions, bounded log-odds models, uniform-encoding Bayesian accounts, and matched efficient-coding models on held-out data. In a new dot probability estimation experiment with bimodal stimulus statistics, the recovered prior tracks the new distribution while the recovered encoding remains U-shaped. Together, these results identify the inverse-S-shaped probability weighting function as the joint product of a stable U-shaped encoding and a flexible prior, integrated by optimal Bayesian decoding.
PaperID: 7433, Poster
Authors:
Yongjae Kim, Sungbin Park, Hoon Ji, Sukmin Yun, Yeonjoon LeeAbstract: Transfer-based black-box adversarial attacks provide a practical way to evaluate the robustness of Vision-Language Pre-training (VLP) models, which remain vulnerable to adversarial perturbations that mislead cross-modal matching and downstream predictions. For strong transfer against VLP models, attacks should depend on prediction-relevant evidence localization and model-agnostic perturbation design. Existing methods fall short in both aspects, typically perturbing images broadly or relying on raw attention maps, which imprecisely localize decisive patches and derive perturbation patterns from the surrogate model's own representations, hindering transfer across VLP architectures. In this paper, we propose Graph-Guided Transferable (GGT) attack, a transferable multimodal adversarial attack that decouples what to perturb from how to propagate: the former by cross-modal sensitivity, the latter guided by intrinsic image structure. This separation enables better localization by capturing VLP-specific cross-modal evidence, while enforcing perturbation propagation through model-agnostic image structure, thereby improving transfer across unseen architectures. GGT first identifies crux patches decisive for cross-modal prediction by combining gradient-weighted attention and attention entropy, which capture how indispensable and how broadly influential each patch is. It then propagates perturbations from these crux patches through a model-agnostic graph constructed from spatial adjacency and vision-only patch similarity, with hop-wise attenuation. Extensive experiments demonstrate that GGT consistently outperforms strong baselines in black-box transfer success rate across diverse VLP architectures, improving over the strongest baseline by up to 23.66% in text retrieval and 16.06% in image retrieval.
Abstract: Optimizing non-convex functions is a fundamental challenge across machine learning and combinatorial optimization. We introduce and study \gamma-weakly \theta-up-concavity, a novel first-order condition that characterizes a broad class of such functions. This condition provides a powerful unifying framework, strictly generalizing both DR-submodular and One-Sided Smooth (OSS) functions while capturing broader forms of scale-dependent curvature, including accumulating-then-diminishing returns and flat-start behavior. Our central theoretical contribution demonstrates that \gamma-weakly \theta-up-concave functions are upper-linearizable: for any feasible point, we can construct a linear surrogate whose gains provably approximate the original non-linear objective. A key technical contribution is a nonuniform upper-linearization argument yielding approximation coefficients that depend explicitly on the curvature parameters and the geometry of the feasible region. This linearizability yields immediate and unified approximation guarantees for a wide range of problems. Specifically, we obtain unified approximation guarantees for offline optimization as well as static and dynamic regret bounds in online settings via standard reductions to linear optimization. Moreover, our framework recovers the optimal approximation coefficient for DR-submodular maximization and improves existing approximation coefficients for OSS optimization, particularly over matroid constraints.
PaperID: 7435, Poster
Abstract: Flow matching has shown strong potential in generative tasks, while its optimization via Reinforcement Learning (RL) remains underexplored, especially in continuous-discrete mixed settings. Moreover, flow matching is formulated as an Ordinary Differential Equation (ODE), whose deterministic trajectories do not naturally support RL optimization. Introducing stochasticity by converting the entire trajectory into a Stochastic Differential Equation (SDE) is a common practice. Nevertheless, not all timesteps contribute to the final outcomes, and redundant exploration introduces additional noise. Furthermore, in multi-objective settings, some objectives may dominate the optimization while others receive insufficient updates. In this study, we propose an RL-guided continuous-discrete flow matching framework with temporal selective exploration for 3D de novo molecular design. Specifically, our method jointly optimizes continuous and discrete flow matching for atomic coordinates and molecular identities, respectively, while restricting exploration to timesteps that primarily contribute to target properties, reducing ineffective exploration. We also introduce a reward-based advantage calibration mechanism to mitigate objective dominance in multi-objective optimization. The framework is applied to design inhibitors for both the well-studied target protein EGFR and the challenging Pin1 protein. We identify some novel candidate inhibitors, supported by in silico validation.
PaperID: 7436, Poster
Authors:
Enqiang Zhu, Ke Wang, Yu Zhang, Shengzhi Wang, Xianhang Luo, Chanjuan LiuAbstract: Constraint satisfaction problems (CSPs) serve as critical benchmarks for neural algorithmic reasoning, requiring accurate multi-step constraint propagation to determine variable values that satisfy all specified constraints. Existing neural solvers often rely on extensive training datasets, struggle with distribution shifts, or allow fixed clues to drift in continuous spaces. We introduce CSP-RSB, a data-efficient recurrent Transformer that directly incorporates constraint topology into the attention mechanism through a learnable relational bias. Additionally, we propose a hard axiom clamping strategy that fixes clues after each recurrence constraint projection and uses the QK-Norm to stabilize a two-layer block. In 9 × 9 Sudoku, CSP-RSB achieves 99.94% board accuracy on the RRN-test, 99.13% on unfiltered Kaggle, and 99.78% on filtered Kaggle, surpassing the previous best, Trial & Error 2025, which reported accuracies of 99.40%, 98.90%, and 99.50%, respectively. Furthermore, it successfully solves Erd\Hos--R\'enyi graph coloring with 100% accuracy in 9 epochs and maintains 99.5-100% accuracy on 10 × 10 binary puzzles. By incorporating explicit topological structure, CSP-RSB reduces the need for depth and data requirements for effective algorithmic reasoning.
Authors:
Geo Ahn, Jiwook Han, Youngrae Kim, Joonseok Lee, Jinwoo ChoiAbstract: Fine-tuning MLLMs for Video Temporal Grounding (VTG) often improves in-domain performance but degrades sharply under domain shift. In this work, we find that this failure is primarily driven not just by unseen query concepts, but by visual domain shift, which prevents the model from coupling its learned temporal localization knowledge with its inherent entity-attention capability. To address this, we introduce EVIDENT, a parameter-efficient adaptation framework that anchors temporal grounding in the inherent entity-attention of pre-trained MLLMs by routing VTG adaptation through explicit visual entity evidence. EVIDENT consists of three components: (i) an Entity Bottleneck Adapter that transforms dense visual tokens into compact entity-level slots, (ii) an Entity-Binding Distillation loss that instills objectness priors into the semantically unstructured MLLM visual space, guiding each slot to bind to a coherent entity, and (iii) an Entity-to-eVidence gating mechanism that leverages the captured entities as evidence, steering the model to localize moments containing query-relevant entities. Together, these components enable VTG fine-tuning to rely on entity-grounded evidence rather than brittle dataset shortcuts. Experiments on cross-domain VTG benchmarks show that EVIDENT consistently improves out-of-domain robustness while preserving competitive in-domain performance with modest parameter overhead. These results suggest that entity-level grounding is an effective inductive bias for generalizable temporal localization.
PaperID: 7438, Poster
Authors: Rohit Sinha, Saroj Kumar
Abstract: Graph neural networks exhibit a fragile relationship with depth, as adding layers leads to rapid homogenization of node representations (over-smoothing). Recent manifold-constrained approaches using doubly stochastic matrices address this issue, but the fundamental principle underlying their success remains unclear. This study investigates whether doubly stochasticity is essential or exemplifies a broader geometric principle for stable residual propagation in deep GNNs.We formalize three axioms to define residual mixing matrices as contractive monoids. We instantiate this framework through three manifolds (Orthogonal Group, Diagonal Contraction Monoid, Spectral Norm Ball) and evaluate them across nine node-classification benchmarks using four GNN backbones. Results demonstrate that contractive-monoid constraints maintain stable accuracy and low representational collapse at extreme depths, while standard GNNs degrade sharply beyond moderate depth. This work establishes that stable residual propagation arises from contractive, compositional structure anchored at identity, revealing a significantly larger design space for deep GNN architectures while maintaining theoretical guarantees.
PaperID: 7439, Poster
Abstract: The linear quadratic regulator (LQR) is a canonical optimal control problem where recent work has established global convergence guarantees for model-free policy gradient methods. However, designing model-free algorithms that enforce hard constraints throughout learning remains an open challenge. Existing approaches are either partially model-based, enforce soft constraints through Lagrangian relaxations, or project pretrained policies onto the constraint set, sacrificing hard constraint satisfaction or convergence guarantees of the underlying algorithms. In this work, we develop principled model-free policy gradient methods for LQR problems with hard constraints. Our key idea is to parameterize the policy class using a truncated stochastic policy so that hard constraints are satisfied by construction, then perform policy optimization entirely within this class. We prove that the resulting LQR objective over the proposed truncated policy class satisfies a weak gradient-dominance property in both the undiscounted and discounted settings. We leverage this property to develop model-free policy-gradient algorithms that maintain zero violation of the encoded hard constraints at each iteration, and converge to an approximately globally optimal policy within the proposed safe policy class.
Abstract: We measure how much one recurrence is worth to a looped (depth-recurrent) transformer, in equivalent unique parameters. From an iso-depth pretraining sweep across recurrence counts r \in \1, 2, 4, 8\ spanning ~50× in training compute, we fit a joint scaling law L = E + A\,(N_\textonce + r^\varphi N_\textrec)^-\alpha + B\,D^-\beta and measure a recurrence-equivalence exponent \varphi = 0.46. Intuitively, \varphi tells us whether looping a block r times is equivalent in validation loss to r unique blocks of a non-looped model (full equivalence, \varphi=1) or to a single block run repeatedly with no capacity gain (\varphi=0). Our \varphi = 0.46 sits in between, so replacing unique blocks with shared recurrences increases validation loss at matched training compute. For example, at r=4 a 410M looped model performs on par with a 580M non-looped model, but incurs the training cost of a 1B non-looped one. We demonstrate the utility of \varphi as a diagnostic tool on two case studies: commonly used truncated backpropagation lowers \varphi to 0.38, indicating that the loop mechanism is poorly trained under truncation, even though validation loss decreases. Conversely, hyperconnections raise \varphi to 0.65, a genuine capacity gain. Our method separates true loop improvements from training-side gains, a distinction raw validation loss cannot make.
PaperID: 7441, Poster
Authors: MingCai Chen, Heng-yang Lu, Yuntao Du, Baoming Zhang, Hao Zhou
Abstract: Pre-trained vision-language models (VLMs) enable zero-shot classification by matching images with text prompts, yet their raw confidence scores are unreliable for out-of-distribution (OOD) detection. We identify a confidence bottleneck: zero-shot VLMs produce posteriors whose maximum confidence yields much weaker ID/OOD separation than supervised classifiers trained on the target label space. To theoretically explain this gap, we relate VLM confidence to a well-separated supervised reference through a Markov-kernel transformation and analyze single-threshold detection using the Youden/KS index. Our analysis shows why raw VLM confidence is unlikely to recover the clean separation achieved by supervised classifiers. Motivated by the observation that class-wise posterior distributions contain useful structure beyond maximum confidence, we propose Self-Reinforced Optimal Transport (SROT), a label-free test-time framework that reshapes VLM posteriors with optimal transport and adaptively trains a lightweight OOD detector from rectified pseudo-labels. The detector further feeds open-set evidence back into posterior alignment, forming a self-reinforcing loop between distribution-level rectification and feature-level OOD learning. Extensive experiments on standard zero-shot OOD detection benchmarks show that SROT achieves state-of-the-art performance, with ablations validating posterior alignment, adaptive detector training, and self-reinforcement.
PaperID: 7442, Poster
Authors:
Wanrong Yang, Lingfang Li, Dominik Wojtczak, Yihang Zhou, Danli Shi, Yalin ZhengAbstract: A language model agent acts on a context assembled by retrieval, summarisation, prompt templates, and memory. How many decision-relevant bits must this context carry before the agent can act with low regret? We give an answer that does not depend on the model, decoder, or scaffold. Modelling the agent as a Markov chain H→C→A and treating Bayes regret (the value gap between the chosen and the optimal action) as the distortion measure, we define a regret rate-distortion (RRD) function Rℓ(ε) whose converse is a one-line consequence of the data-processing inequality (DPI). A context that carries fewer than Rℓ(ε) bits cannot drive expected regret below ε, regardless of decoder, scaffold, scale, or chain-of-thought. For finite M-ary decisions under zero-one loss the bound reduces to the Fano floor in closed form, requiring ≈3.14 bits at (M=16, ε=0.1) and ≈6.73 bits at (M=256, ε=0.1). We instantiate the theory on 441 controlled single-step tool-selection cells, a balanced 16-way classification proxy spanning three open-source large language models (Qwen3-32B, Gemma3-27B, GLM-4-32B) and three public benchmarks (API-BANK, τ-bench, ToolBench). Because the gold tool index X=f(H) is a deterministic function of the history, the action-side rate IMM(X; A) lower-bounds the context capacity by data processing, so an action-side test of the converse is strictly tighter than a context-side one. Every cell respects the predicted floor (median margin +1.07 bits). Across 629 matched-rate pairs, 97.5% share overlapping regret confidence intervals (CIs), and a chain-of-thought supplement slides cells along the frontier rather than across it. The same data exposes a tokens-versus-bits gap reaching ∼1.7×10⁶:1, with the model extracting fewer than 0.05 bits about the gold action on a 32,768-token ToolBench prompt. In the single-step regime, the framework supplies a falsifiable, model-agnostic stopping criterion at IMM=Rℓ(ε) and reframes context engineering as an information-allocation problem.
PaperID: 7443, Poster
Authors: Seong Hah Cho, Junyi Li, Anna Leshinskaya
Abstract: Value alignment of Large Language Models (LLMs) requires us to empirically measure these models' actual, acquired representation of value. Among the characteristics of value representation in humans is that they distinguish among value of different kinds. We investigate whether LLMs likewise distinguish three different kinds of good: moral, grammatical, and economic. By probing model behavior, embeddings, and residual stream activations, we report pervasive cases of value entanglement: a conflation between these distinct representations of value. Specifically, both grammatical and economic valuation was found to be overly influenced by moral value, relative to human norms. This conflation was repaired by selective ablation of the activation vectors associated with morality.
PaperID: 7444, Poster
Abstract: Solving algorithmic tasks that involve many sequential steps typically requires scaling up transformer depth or chain-of-thought length, incurring unfavorable parameter and inference-compute costs, respectively, as tasks deepen. Looped transformers promise to recover that depth at a fraction of the cost by reusing a single block across iterations, but it remains unclear how the loop actually performs its computation, or whether it merely simulates a deeper untied stack. We test this on \emphdirected permutation traversal (\dePo), a synthetic task in the physics-of-LMs lineage in which solving an H-hop query requires following H pointers along a fixed permutation of N=16 entities, with H up to 15. We demonstrate that, above a minimum effective depth, a 3.2\,M-parameter looped transformer trained with hop length sampled uniformly per batch traverses all 15 hops across all seeds, while a 25.8\,M-parameter fixed-depth baseline at the same effective depth fails to exceed three hops; weight tying, not effective depth, predicts capability. In the same models, linear probes on the residual stream uncover a clean \emphstaircase: at iteration k the stream linearly decodes hop k, with slope exactly one hop per iteration, and each hop stays decodable for roughly two further iterations before being overwritten. The model's own logit lens, applied through the trained final layers, traces the same one-hop-per-iteration ordering at lower amplitude. Taken together, our results recast looped transformers as iterative program learners: weight tying forces the optimizer to discover a single per-iteration update rule, and that rule corresponds to one algorithmic step of the underlying task. Code is available at \urlhttps://anonymous.4open.science/r/insidetheloop-3ED0.
PaperID: 7445, Poster
Authors: Nguyen Vu Nguyen
Abstract: We introduce DIAL (Directional Intervention, Asymptotically Limited), a parameter-efficient adapter that adds a bounded, monotone, scalar-controlled residual \delta(s,c) = \sigma(g) \odot \tanh(c/T) \odot \tanh(Dz) with z = D^\top \phi(s) to a base language model's hidden state. The architectural form yields a closed-form upper bound \lVert \delta \rVert_\infty \le \sigma(g_\max) on the residual L_\infty norm, computable from model weights alone. Across 5 base models (3.8B to 70B parameters, 4 architecture families; one checkpoint per family, with 5-seed csweep variance at Llama-3.1-8B), the residual saturates at the predicted bound with mean ratio 1.000 \pm 0.001 (worst-case family deviation 0.2 percentage points). On Llama-3.1-8B under GPT-4o judge (n=100 per bench, zero content-filter blocks), DIAL exceeds a FiLM baseline trained with the same multi-c objective on every benchmark and judge. Semantic vs strict judge: HarmBench DIAL +0.52 / +0.55 vs FiLM +0.44 / +0.39; AdvBench DIAL +0.63 / +0.66 vs FiLM +0.46 / +0.50; XSTest DIAL +0.33 / +0.25 vs FiLM +0.26 / +0.22; both judges agree on the ranking. Outside training range, DIAL's residual saturates at \sigma(g_\max) while FiLM's grows unboundedly: at c = 10^4, FiLM's residual L_\infty norm is 534 times larger than DIAL's. The bound holds across all control values, not only trained ones; concrete deployment scenarios (configuration error, multi-tenant policy ceilings, adversarial composition) are where this distinction matters in practice. The bounded-monotone form produces analogous saturation behavior on two continuous-control RL environments (HC-Vel n=35, Ant-Goal n=30) under matched oracle-c; the prior fails when c must be inferred online (meta-RL). Ablations confirm both the bound and the saturating temperature are necessary; a frozen-gate variant retains the bound at lower controllability magnitude. We release adapter train reports and canonical bench generations, algebraic derivations (with a SymPy reproducibility script), the full GPT-4o judge cache (semantic and strict), and a human-labeled validation set (n=50, Cohen's kappa = 0.895). Full adapter checkpoints are deferred to camera-ready release.
PaperID: 7446, Poster
Abstract: Quantifying mouse behavior from pose is a foundational tool in neuroscience, ethology, drug discovery, and animal welfare. Self-supervised foundation models are a natural fit, but a foundation model for mouse behavior must handle social behavior - many behaviors that matter (aggression, courtship, mounting, allogrooming) are inherently multi-animal. Existing self-supervised pose models fall short: they encode each animal independently and discard interaction signal, hard-code the animal count, and pretrain one dataset at a time. We introduce MICE (Multi-animal Interaction Context Encoder), a hierarchical foundation model for mouse behavior that pretrains jointly on six mouse-pose corpora spanning one to four mice and 7 to 27 keypoints. MICE combines a hierarchical individual encoder capturing per-mouse kinematics at multiple temporal scales with a Perceiver-style social encoder whose learned latent codebook decouples parameter count from animal count, so a single checkpoint serves any group size without retraining. Across four evaluation protocols on all six datasets, MICE matches or exceeds prior keypoint-based baselines, and layer-wise probing shows the social stage incrementally improves discriminability beyond the individual encoder. Code and weights will be released.
PaperID: 7447, Poster
Abstract: Pretrained EEG backbones improve transfer performance, but downstream diagnostic heads remain difficult to audit: predictions are typically computed from unconstrained hidden representations, while explanations are often generated only post hoc. We introduce tPoE-EIB, an evidence-information bottleneck head for adapting EEG backbones under an evidence-only prediction constraint. tPoE-EIB selects temporal and channel-specific evidence, maps the resulting summaries to Gaussian experts over a shared latent variable, and fuses them via a tempered product-of-experts posterior. The classifier conditions exclusively on this latent variable, yielding an explicit decision pathway whose information flow is constrained by the expected posterior KL divergence. This formulation induces a tractable supervised objective with an information-rate penalty, while the closed-form tempered posterior mitigates overconfident aggregation of correlated evidence sources. We evaluate tPoE-EIB on pretrained EEG foundation model backbones across six diagnostic tasks: event-type classification, abnormality detection, seizure detection, cognitive decline staging, depression screening, and cerebrovascular disease classification. The evaluation spans public benchmarks and in-house clinical cohorts, covering both binary screening and fine-grained staging, as well as sparse and dense electrode montages. tPoE-EIB maintains competitive balanced accuracy while outperforming representative post-hoc explanation methods on selection-faithfulness audits, including insertion–deletion and gate-causality tests. Its structured posterior further supports integration-faithfulness evaluations, including expert-drop, posterior-reliance, and expert-disagreement tests. Overall, these results suggest that evidence-only, rate-limited fusion is a practical approach to building auditable diagnostic models on top of frozen EEG foundation backbones.
Authors: Aishik Nagar, Arun-Kumar Kaliya-Perumal, Yu-Hsuan Han, Andrew Sheng-Han Huang, Kristen Kee, Yushi Cao, Yiming Chen, Hongchao Jiang
Abstract: Inpatient clinical reasoning is a sequential decision under partial observability: the clinician sees the admission so far and must choose the next action whose downstream consequences are not yet visible. Existing clinical-LLM evaluations and reinforcement-learning (RL) rewards collapse this structure into closed-form retrieval, full clinical journey leakage, or unanchored LLM-as-judge scoring. Here, we introduce CLR-voyance, an end-to-end framework that reformulates inpatient reasoning as a Partially Observable Markov Decision Process (POMDP) and supervises it with rewards that are simultaneously outcome-grounded and clinician-validated. We instantiate the formulation as CLR-POMDP, which partitions successful patient journeys into a policy-visible past and an oracle-only future. Using the past information, an oracle LLM generates a case-specific query-answer pair, and the first adaptive rubric for clinical reasoning which is verifiable in the future of the patient journey. These rubrics are used for both RL post-training and evaluation of models for inpatient clinical reasoning. CLR-voyance post-trains Qwen3-8B and MedGemma-4B with GRPO followed by model merging, yielding state-of-the-art inpatient clinical reasoning while retaining generalist capabilities. CLR-voyance-8B achieves 84.91% on CLR-POMDP, ahead of frontier medical reasoning models like GPT-5 (77.83%) and MedGemma-27B (66.66%) and has comparable or better performance on existing medical benchmarks. To ensure that our setting is clinically meaningful, we conduct a large-scale clinician alignment study, where physicians curate per-case rubrics, grade candidate responses against them, and provide blinded pairwise preferences of model reasoning. This study provides insights on clinical LLM-as-a-judge and clinical preference-model selection, which can inform the community at large. CLR-voyance has been deployed for 6+ months at a partner public hospital, drafting thousands of reasoning-heavy inpatient notes and 1.03T+ tokens processed to date.
PaperID: 7449, Poster
Abstract: Open-world task planning requires not only generating feasible plans but also acquiring the knowledge needed to execute them. Existing task planners for egocentric agents either attempt to exhaustively gather information from the environment or generate plans under prior assumptions of the world and replan as new observations arrive. Both strategies are impractical in complex environments and can potentially expose agents to undesirable or unsafe situations. To enable effective planning in open-world settings, we propose Plan2Sense, a framework that grounds planning in the agent’s epistemic state to determine what information is required for plan feasibility and how to acquire it. We formalize decision-making as a three-valued first-order Knowledge-State Markov Decision Process (KS-MDP) and develop a knowledge state goal regression algorithm KSR to construct sensing-aware conditional plans in the open-world setting. Our formulation supports sensing actions for both predicate disambiguation and knowledge acquisition via object binding without assuming a fully known initial state or a closed set of predefined objects. Across three open-world egocentric benchmarks requiring active knowledge acquisition, \textscPlan2Sense achieves higher success rates and improved efficiency over existing baselines while consistently avoiding undesirable and potentially unsafe states.
PaperID: 7450, Poster
Abstract: Confronted by the existing prototype-based methods that often fail to capture complex intra-class variations and inter-class semantic relationships, weakly supervised semantic segmentation (WSSS) faces persistent challenges due to the inherent incompleteness and boundary inaccuracies of Class Activation Maps (CAMs). To tackle these critical limitations, we propose a relationally optimized prototype Memory bank (RO-PMB) framework in this paper, where the integration of Graph Neural Networks (GNNs) with a dynamically optimized global prototype memory bank is pioneered to explicitly model and learn the intricate semantic relationships among all prototypes. Specifically, our Prototype Relationship Diagram Construction (PRDC) module leverages GNNs for contextual message passing over a learned prototype graph, enriching global prototypes with crucial associative knowledge. As a result, a subsequent context-aware refinement process is sustained to inject structured semantic information into local image-specific prototypes, thereby generating CAMs with unparalleled completeness and boundary precisions. Extensive experiments validate that, compared with the existing representative state of the arts in multi-stage WSSS baselines, our proposed RO-PMB achieves competitive and superior performances on a range of challenging benchmarks, such as PASCAL VOC and MS COCO etc. Code will be released.
Abstract: conditioned on desired semantic properties is an emerging challenge in modern generative modeling. Metropolis-Hastings (MH) provides a principled route to conditional sampling, but requires access to exact pointwise target-density evaluations, which are not available in generative settings. Meanwhile, pairwise comparisons by humans or model ``judges'' are highly accessible and have proved valuable across diverse applications. We introduce Pref-MH, a general exact MH sampler for judge-induced conditional distributions using only stochastic binary pairwise comparisons. Our key observation is that the MH unnormalized density ratio matches the preference odds of BT choice models. The central challenge is that while MH requires precise ratio computation, BT judges provide only sampled binary feedback. To this end, we develop a valid accept/reject rule whose resulting Markov chain provably converges to the target distribution. We further show that, for a fixed proposal kernel and budget, Pref-MH is optimal in the Peskun-Tierney sense among this class of exact reversible acceptance rules. Experiments on text generation with LLM judges and image generation with VLM judges demonstrate that Pref-MH provides a practical and flexible approach to conditional sampling from comparative feedback.
PaperID: 7452, Poster
Abstract: We propose a label-free approach to adapt powerful but generic vision foundation models to specialized scientific domains. Standard supervised fine-tuning is often ill-suited to these settings: labels are scarce, and task-specific training can collapse the model's generality and hurt robustness. We instead leverage metadata to adapt representations to new domains in a self-supervised manner. Our method, META-DINO, combines a standard self-supervised objective with flexible metadata guidance that handles both highly granular discrete metadata and continuous metadata. It encourages the representation to preserve informative factors while suppressing spurious ones. Across subcellular fluorescence microscopy, Earth observation, wildlife monitoring, and medical imaging, META-DINO consistently outperforms standard unsupervised domain adaptation and fully supervised adaptation. It also exceeds highly-specialized domain-specific state of the art, while using no task labels for backbone adaptation and only lightweight probes for supervision.
PaperID: 7453, Poster
Authors:
Zhifeng Yang, Wei Luo, Yongjian Chen, Shiqi OuAbstract: Large language models (LLMs) generally lack deep causal reasoning capabilities in temporal reasoning. Inspired by document layout reordering, this paper proposes a paradigm shift from discrete narratives to causal topologies. It constructs two architectures: a training-free neuro-symbolic framework, temporal causal reasoning agent (TCR-Agent), which achieves interpretable causal deduction and counterfactual truncation through external causal graph instantiation and a symbolic intervention engine; and an end-to-end model, temporal causal reasoning former (TCR-Former), which internalizes these mechanisms into the Transformer attention space via a causal-topological attention mask and temporal span biases. Experiments on the TempoCausal benchmark demonstrate that TCR-Agent comprehensively surpasses existing training-free baselines. At the same time, TCR-Former achieves an accuracy 16.5% higher than that of state-of-the-art (SOTA) methods, outperforming large-scale models such as GPT-5.4 and DeepSeek-V3.2 with only approximately 8.6B parameters. Further analysis reveals a structure-reasoning disconnection phenomenon in existing methods: external symbolic constraints and chain-of-thought (CoT) reasoning can improve reasoning steps, yet fail to enhance deep causal deductive capability. By internalizing temporal causal reasoning capabilities into the model, TCR-Former effectively bridges this gap, providing an efficient and reliable end-to-end foundation for complex reasoning scenarios.
Abstract: Generative modeling and self-supervised representation learning (SSL) optimize structurally different objectives: generative training rewards distributional fidelity, while SSL rewards semantic coherence. Yet recent work repeatedly finds that SSL features improve generative training, though the mechanism of this synergy remains unclear. Here, we study the benefits of SSL in generative modeling in the framework of one-step generation where the role of representation is explicit: frozen SSL features are used to match generated samples to real data. We use the Sinkhorn divergence in that feature space, providing a tractable surrogate for the Wasserstein distance, the population-level discrepancy approximated by Fr\'echet-style evaluation metrics (such as FID). We find that this objective becomes highly effective when computed in a semantically structured SSL feature space (a 39× reduction in ImageNet FID). We trace this behavior primarily to matching estimation: semantic SSL features that suppress nuisance reconstruction details induce a more compact geometry, making distribution matching more tractable. As a consequence, the best training SSL features need not match the features used by the evaluation metric. In particular, we show that using Inception as the feature extractor can improve FID while degrading matching stability and sample quality, revealing a form of metric hacking. Using extensive experiments on ImageNet, we identify which SSL feature families lead to best generation performance and show that matching stability is a quantitative criterion for selecting them.
Abstract: Aligning Large Language Models (LLMs) with human preferences typically relies on external supervision, which faces critical limitations: human annotations are scarce and subjective, reward models are vulnerable to reward hacking, and self-evaluation methods suffer from prompt sensitivity and biases. In this work, we propose , an intrinsic, annotation-free quality signal derived from model representations. Stable rank measures the effective dimensionality of hidden states by computing the ratio of total variance to dominant-direction variance, capturing quality through how information distributes across representation dimensions. Empirically, stable rank achieves 84.04% accuracy on RewardBench and improves task accuracy by an average of 14.1% relative to greedy decoding via Best-of-N sampling. Leveraging this insight, we introduce , which uses stable rank as a reward signal for reinforcement learning. Without external supervision, SR-GRPO improves Qwen2.5-1.5B-Instruct by 10% on STEM and 19% on mathematical reasoning, outperforming both learned reward models and self-evaluation baselines. Our findings demonstrate that quality signals can be extracted from internal model geometry, offering a path toward scalable alignment without external supervision.
PaperID: 7456, Poster
Authors: Sungho Chun, Ju Yong Chang
Abstract: Estimating 6-degrees-of-freedom (6DoF) head pose from a single RGB image remains challenging under strong perspective distortion. We observe that the widely used design choice of rectangular face cropping introduces a geometric inconsistency. This leads not only to image--label mismatch under 2D rotation augmentation but also exacerbates perspective ambiguity, making it difficult to distinguish projection distortion from the actual appearance and pose of faces located near the image periphery. To address this issue, we propose HOGWARTS, a geometry-driven framework that replaces rectangular cropping with homography warping induced by a newly defined virtual camera space. By explicitly defining a virtual camera that shares the optical center with the original camera and warping the input image into this space, our method compensates for off-axis distortion in a geometrically consistent manner. Furthermore, to accurately estimate 3D translation in the virtual space, we introduce a pinhole-camera-based translation formulation that explicitly compensates for projection scale variation caused by the virtual camera transformation. Consequently, HOGWARTS enables more geometrically consistent and robust 6DoF pose estimation across varying image locations. To the best of our knowledge, this is the first study to estimate the 6DoF pose of a target within a virtual camera space explicitly designed to mitigate perspective distortion. Experimental results on the ARKitFace and BIWI datasets demonstrate that HOGWARTS consistently outperforms prior methods under severe perspective distortion as well as in cross-domain settings. Code will be made publicly available.
PaperID: 7457, Poster
Abstract: Partial Label Learning (PLL) trains models from instances with candidate label sets containing the ground-truth, while Semi-Supervised Partial Label Learning (SSPLL) further leverages the available unlabeled data to prompt the performance. The previous SSPLL methods implicitly assume the unlabeled data are class-balanced distributed. However, in realistic scenarios, the class distribution of unlabeled data is unknown, resulting in performance collapse on minority classes for the existing SSPLL methods. We identify two coupled biases behind this failure: (i) pseudo-label selection bias, where high-confidence filtering tends to over-select samples from dominant classes, causing the training data to become increasingly skewed toward these classes; and (ii) disambiguation bias, where dominant classes are more likely to win within candidate sets, even when they are incorrect. Theoretically, we show that these two biases reinforce each other over training, forming a feedback loop that degrades generalization performance. To weaken this loop, we propose AdapMatch, which uses DAS to reduce the skew in training data caused by pseudo-label selection bias via prior-aware pseudo-label admission and CAD to lower the disambiguation error by suppressing candidate-set domination. Across five benchmarks, AdapMatch yields gains up to 41.31% at most , showing strong robustness to varying scale, imbalance, and ambiguity. The code is available at \urlhttps://anonymous.4open.science/r/AdapMatch-8BC5.
PaperID: 7458, Poster
Abstract: Large language models (LLMs) are driving autonomous agents toward real-world environments, with the Web emerging as one of the most representative interaction scenarios. However, training Web agents with naive exploration remains challenging: successful trajectories are rare, and the sparse terminal rewards make it difficult for agents to obtain useful feedback from failed trajectories. Inspired by skill acquisition theory, we propose Skill-Composed Progressive Exploration (SCOPE), a short-to-long learning framework for long-horizon Web agents that decomposes complex task exploration into a gradual expansion process from local skills to complete task-solving capability. Specifically, SCOPE iteratively explores around local skills: Skill Discovery extracts reusable local skills and failure experiences from historical rollouts; Skill Orchestration composes discovered skills into high-level skill chains; andSkill Expansion leverages failure experiences aligned with the current skill chain to guide the agent toward exploring new subskills. Through this process, SCOPE progressively constructs task-completing skill chains across rollouts and internalizes them into the policy parameters through agentic on-policy distillation. Extensive experiments on WebArena, WebVoyager, and WebShop show that SCOPE consistently improves exploration dynamics and achieves 3.9%--18.6% relative gains over the strongest baselines in each setting. Our code is publicly available at https://anonymous.4open.science/r/SCOPE-14D3.
PaperID: 7459, Poster
Abstract: Time Series Anomaly Detection (TSAD) is critical for ensuring reliability in real-world systems such as industrial monitoring, healthcare, and online services. However, learning-based TSAD methods trained on a single normal-only TargetSet, namely the target dataset, often suffer from incomplete and biased knowledge of normality, causing domain-level false alarms when unseen but valid normal patterns appear at test time. A natural extension is to leverage DomainSet, composed of datasets from one domain, to enrich domain-level normality. Yet, our analysis shows that simply incorporating DomainSet could make domain-admissible normal variations and TargetSet-specific anomalies highly entangled in the learned representations, weakening anomaly separability. Motivated by human experts who establish normal patterns, calibrate admissible variations, and identify true anomalies, we propose dStructAD, a domain-level structured representation framework for time series anomaly detection, to address this challenge. Specifically, dStructAD first builds structured domain-level normality based on the Kolmogorov-Arnold network, and then calibrates admissible variations via two phases. Phase 1 learns shared domain-level semi-structural normality knowledge from DomainSet, while Phase 2 performs structure-preserving TargetSet calibration to absorb admissible variations without collapsing anomaly separability. Across six benchmarks, dStructAD delivers consistent improvements over SOTAs, with 3% average gain in overall performance. Code could be available at
PaperID: 7460, Poster
Abstract: Object-centric representation learning (OCL) has been widely proposed as a principled solution to the binding problem in deep learning, promising improved compositionality, reasoning, and robustness. However, its effectiveness in modern foundation model settings remains unclear. In this work, we systematically evaluate whether explicit object-centric (OC) representations meaningfully improve binding in pretrained foundation models. Departing from prior studies, we assess OC representations in a realistic regime by pairing them with pretrained foundation LLMs and analyze their performance on diverse, open-ended VQA and grounding benchmarks. Our findings challenge the prevailing narrative: slot-based OC representations consistently underperform standard dense patch-based features across multiple benchmarks, including compositional reasoning and hallucination-sensitive tasks. To understand this behavior, we analyze the slot representations through the lens of core binding components, namely scene segregation and object representation,and uncover three key limitations: (i) inferior scene segregation, where slot assignments fail to cleanly disentangle objects compared to the implicit grouping already present in pretrained visual backbones; (ii) inherent information loss during slot encoding, which degrades downstream performance of VLMs despite adapting them to the resulting feature space of OC; and (iii) weak attribute encoding, where slots struggle to preserve fine-grained properties such as category, color, and spatial position. Further analysis reveals that these issues are not solely attributable to slot-based learning: while dense patch-based representations exhibit stronger binding, they too are imperfect. This suggests that explicit object-centric modeling, as currently instantiated, introduces bottlenecks that discard useful information without sufficiently improving structural reasoning. Together, our results expose a fundamental disconnect between the theoretical appeal of OCL and its practical utility in the era of foundation models. They point toward a need for rethinking binding, not as explicit object-factorization alone, but as a balance between structure and information preservation, potentially through new mechanisms that retain the richness of dense representations while enabling more reliable compositional reasoning.
PaperID: 7461, Poster
Abstract: Direct Preference Optimization (DPO) has been widely adopted for aligning language models with human preferences. Recent empirical work has revealed that DPO induces remarkably sparse parameter updates, yet no theoretical explanation for this phenomenon has been offered. We analyze DPO's optimization dynamics under a neural tangent kernel linearization and show that the logistic structure of the DPO loss gives rise to an implicit maximum-margin bias. Building on this observation, we introduce the concept of dual sparsity: the DPO update direction is a sparse linear combination (data-level sparsity, driven by support vectors) of sparse gradient feature vectors (parameter-level sparsity, arising from gradient cancellation between preferred and dispreferred responses). We formalize this structure in our main theorem and present a framework relating it to weight disentanglement---a condition favorable for model merging. Experiments on GPT-2 Small and Medium with the Anthropic HH-RLHF and UltraFeedback datasets confirm both sources of sparsity and show that DPO reduces cross-task interference by up to 35 × compared to supervised fine-tuning.
Authors: Pritika Vig, Renchin Wu, William Lotter
Abstract: Vision foundation models trained on discretely sampled images achieve strong performance on classification benchmarks, yet whether their representations encode the continuous processes underlying their training data remains unclear. This question is especially pertinent in computational pathology, where we posit that models whose latent representations implicitly capture continuous disease progression may better reflect underlying biology, support more robust generalization, and enable quantitative analyses of features associated with disease transitions. Using diffusion pseudotime, a method developed to infer developmental trajectories from single-cell transcriptomics, we probe whether foundation models organize disease states along coherent progression directions in representation space. Across four cancer progressions and six models, we find that all pathology-specific models recover trajectory orderings significantly exceeding null baselines, with vision-only models achieving the highest fidelities (\tau > 0.78 on CRC-Serrated). Model rankings by trajectory fidelity on reference diseases strongly predict few-shot classification performance on held-out diseases (\rho = 0.92), and exploratory analysis shows cell-type composition varies smoothly along inferred trajectories in patterns consistent with known stromal remodeling. Together, these results demonstrate that vision foundation models can implicitly learn to represent continuous processes from independent static observations, and that trajectory fidelity provides a complementary measure of representation quality beyond downstream performance. While demonstrated in pathology, this framework could be applied to other domains where continuous processes are observed through static snapshots.
PaperID: 7463, Poster
Authors: Collin Cranston, Zhichao Wang, Todd Kemp
Abstract: In the infinite-width limit, training a neural network (NN) via gradient descent reduces to kernel ridge regression (KRR) on a deterministic kernel, the Neural Tangent Kernel (NTK). While the isotropic-input case of the NTK is well understood, the spectra of real world inputs always exhibit power-law decay, and the alignment between this spectrum and the target signal has been shown to govern the generalization behavior of the model. Motivated by this, we study finite-width NTK regression under a dual power-law dataset, where data \boldsymbolx \in \mathbbR^p is drawn with covariance \boldsymbol\Sigma_jj = j^-\alpha and labels are generated from power-law truth vector \boldsymbol\beta_j = j^-r. We introduce a novel technique combining random matrix theory (RMT) and stochastic matrix calculus to derive deterministic equivalents for random matrix functionals that determine the prediction risk, including a new signal-weighted resolvent, which yields an explicit formula for the bias-variance decomposition of NTK regression in the high-dimensional proportional regime. As an application we prove a scaling law for finite-width NTK regression, explicitly characterizing the decay rate of the optimally-tuned excess risk jointly in the data and width resources.
PaperID: 7464, Poster
Authors:
Binze Wang, Jinyu Tian, Xingrun Wang, Jianqing LIAbstract: Unlearnable Examples (UEs) aim to protect data from unauthorized model training by injecting imperceptible perturbations that degrade model generalization. Existing methods typically rely on well-trained surrogate models to generate such perturbations, implicitly assuming that stronger surrogates yield better degradation. In this work, we challenge this widely adopted assumption and show that it is fundamentally suboptimal. We provide theoretical insights by establishing an upper bound on the target model's test loss, showing that perturbation effectiveness is closely related to the surrogate’s sensitivity to input perturbations. Specifically, a less robust surrogate model leads to stronger poisoning performance. Motivated by this analysis, we propose a simple yet effective stability metric to select optimal surrogate models for single-level methods. We further extend this principle to bi-level optimization frameworks. Beyond revealing the bi-level method's implicit reliance on non-robust surrogates, we substantially boost its poisoning performance by the stability-based surrogate selection as well. Experiments on five datasets and four representative UE methods show that our selected surrogates reduce the target model's test accuracy by an average of 18% compared with well-trained surrogates. Notably, this is the first work to reduce CIFAR-10 test accuracy to 10% under a perturbation budget of 1/255, whereas prior UE methods typically rely on 8/255.
PaperID: 7465, Poster
Abstract: Transformers trained via Reinforcement Learning (RL) with outcome-based supervision can spontaneously develop the ability to generate intermediate reasoning steps (Chain-of-Thought). Yet the mechanism by which sparse rewards drive policy gradient to discover such systematic reasoning remains poorly understood. We address this by analyzing the policy gradient dynamics of single-layer Transformers on a synthetic graph traversal task that cannot be solved without Chain-of-Thought but admits a simple iterative solution. We prove that despite training solely on final-answer correctness, policy gradient drives the Transformer to converge to a structured, interpretable algorithm that iteratively traverses the graph vertex-by-vertex. We characterize the distributional properties required for this emergence, identifying the critical role of ``simple examples'': instances requiring fewer reasoning steps. When the training distribution places sufficient mass on these simpler examples, the Transformer learns a generalizable traversal strategy that extrapolates to longer chains; when this mass vanishes, policy gradient learning becomes infeasible. We corroborate our theoretical results through experiments on synthetic data and with real-world language models on mathematical reasoning tasks, validating that our theoretical findings carry over to practical settings.
PaperID: 7466, Poster
Authors:
Rui Hou, Yao Liu, Ruilin Jiang, Mengyao Lu, Jingbo Wang, Lanyi Zhang, Qiao LiuAbstract: Non-stationary time series forecasting (NTSF) requires modeling evolving trends and seasonal patterns. Although linear decomposition operators (e.g., FFT) help isolate these patterns, they often introduce operator-dependent biases and idiosyncratic noise. To address this issue, we propose a novel Multi-view Decomposed Kalman Mixture of Experts (MDK-MoE) framework, which improves forecasting by strategically integrating multiple complementary, yet biased, linear decompositions. Specifically, a Multi-view Decomposition (MVD) extracts frequency components from three complementary views. To mitigate view-specific noise, a Maximizing Mutual Information (MMI) enhances cross-view consistency, maximizes task-relevant information across diverse views, and suppresses view-dependent task-irrelevant information. To address the conflicts inherent in multi-source observations, the efficient Kalman MoE adaptively fuses distinct multi-view features. Instead of adopting traditional O(d^3) covariance inversion, our learnable Kalman gain design enhances numerical stability and enables reliable state estimation under non-stationary scenarios. Extensive experiments on 13 real-world datasets demonstrate that MDK-MoE consistently achieves state-of-the-art performance.
Abstract: Recent advances in diffusion transformers (DiTs) have enabled promising single-turn image editing capabilities. However, multi-turn editing often leads to progressive semantic drift and quality degradation. In this work, we study this problem from a latent-space frequency perspective by decomposing the editing process into two functional components: VAE and DiT. Through systematic analysis in the VAE latent space, we uncover that the DiT introduces dominant low-frequency drift that accumulates as semantic misalignment across editing rounds, while the VAE contributes comparatively stable reconstruction bias. Based on this insight, we propose (Low Frequency Alignment), a training-free, plug-and-play method that performs alignment in VAE latent space. VAE-LFA decomposes latent discrepancies across editing rounds via low-pass filtering, and aligns low-frequency statistics to an exponential moving average of previous rounds, effectively suppressing accumulated semantic drift while preserving high-frequency details. Our method requires , making it applicable to both white-box and black-box DiT editors. For white-box models, VAE-LFA is seamlessly integrated into the editing pipeline by eliminating redundant VAE round trips; for black-box models, it operates via an off-the-shelf VAE to perform inter-round latent alignment. Extensive experiments demonstrate that VAE-LFA improves semantic consistency and visual fidelity across diverse multi-turn editing scenarios, including both controlled and in-the-wild images.
PaperID: 7468, Poster
Authors:
Qiuyu Chen, Liang Xu, Yunnan Wang, Mingqi Yuan, Qi Wang, Jiahao Li, Zhicheng Wang, Hao Zheng, Mulin Chen, Baao Xie, Xin Jin, Wenjun ZengAbstract: Disentangled representation learning (DRL) aims to identify and decompose the interpretable underlying factors of observations. However, most DRL methods rely on independence-oriented regularization to pursue statistical factorization, leaving learned factors without explicit semantic grounding. Such independence-oriented regularization induces an intrinsic trade-off between disentanglement and reconstruction. To address these issues, we establish a spectral learnability threshold for independence-oriented DRL, theoretically revealing a mismatch between the factors favored by the learning objective and human-perceivable semantics. We further introduce SAGE, a semantic-guided DRL framework that leverages MLLMs to automatically extract semantic signals from unlabeled data and uses a triangular latent-geometry constraint to anchor each semantic factor to a designated latent dimension, enabling factor-wise control over human-perceivable semantic variations in generated images. Extensive experiments support our theoretical analysis and demonstrate that SAGE achieves state-of-the-art performance across multiple benchmarks. In addition, our framework exhibits high flexibility and can be integrated with most advanced generative models.
PaperID: 7469, Poster
Abstract: Existing audio deepfake detection (ADD) methods frequently struggle with limited generalization, as fixed bottleneck architectures fail to adapt to the heterogeneous information densities inherent in diverse spoofing artifacts. Moreover, standard augmentation often incurs a pathological robustness-accuracy trade-off. In this paper, we propose a unified ADD framework to address these challenges. Our approach integrates: (1) an Adaptive Audio Feature Integration (AAFI) module that dynamically selects optimal latent dimensions via a compression pool and a stochastic exploration mechanism, ensuring flexible representations across multi-type dynamic attacks; and (2) a Mel-Band Adversarial Perturbation (Mel-BAP) strategy that applies targeted regularization in the mel-spectrogram domain to encourage the learning of invariant decision boundaries. Extensive evaluations on benchmarks including ADD 2022, FakeOrReal, and In-the-Wild demonstrate that our Whisper-small based model achieves a state-of-the-art average Equal Error Rate (EER) of 1.51%. Compared to existing SOTA models such as the 3-billion-parameter Resemble-Detect-3B-Omni, our framework reduces parameter count by up to 91.6% while improving average EER by 60.7%. Notably, the model exhibits exceptional zero-shot transferability, achieving 3.91--5.65% EER on unseen languages and maintaining high resilience against real-world acoustic distortions. This work provides a scalable and efficient “plug-in” solution for enhancing trust and security in the audio Web ecosystem.
Abstract: Text-guided image editing with visual autoregressive (VAR) generators requires controlling both what the model samples and where the sampled change is written back into the image. Existing VAR editors mainly operate on token streams, features, or flat next-token logits, leaving two native structures of bitwise-residual VAR models underused: the per-bit Bernoulli prediction head and the additive multi-scale residual code field. We propose BitResEdit, a training-free editor for bitwise-residual VAR generators such as \textscInfinity. BitEdit performs source-negative guidance by tilting the post-CFG per-bit log-odds along a source--target contrast computed on a shared edited prefix, then projects each update into a closed-form Bernoulli-KL trust region around the clean CFG sampler. ResEdit converts the resulting sampled bits into per-scale continuous-code residuals, gates them with a localization mask, and re-injects them through the generator's native sum-of-scales. The two components couple decision-time bit guidance with combination-time code composition, preserving masked-out latent features exactly while applying localized, scale-aware edits inside the target region. On PIE-Bench with Infinity-2B, BitResEdit achieves the strongest text alignment among same-backbone VAR editors and competitive background preservation, setting a new state of the art among training-free VAR editing methods. Ablations show that BitEdit and ResEdit play complementary roles in target alignment and background preservation. Detailed ablations show that both code-space residual editing and bitwise phrase-contrast guidance are necessary for robust autoregressive image editing.
PaperID: 7471, Poster
Abstract: We investigate the automated generation of executable robot reinforcement learning (RL) policies from natural-language task descriptions. Rather than using large language models (LLMs) as execution-time decision makers, we use them as designers that construct RL pipelines before deployment, including Markov decision process (MDP) formulation and simulator implementation. We identify a phase-wise structure in this problem: MDP formulation is dependency-structured and sequential, whereas simulator implementation becomes modular once the formulation is fixed. Based on this structure, we propose ARPG, a dependency-aware agentic framework that performs sequential MDP modeling, localized verification and correction, and simulator-structured implementation while externalizing intermediate artifacts. ARPG targets a key failure mode of automated policy generation: executable code may still yield non-learnable policies when state, action, reward, and termination artifacts are not semantically aligned. Experiments in Isaac Sim show that ARPG improves modeling correctness, environment executability, and policy-generation success compared with monolithic LLM baselines and off-the-shelf agentic systems. Additional demonstrations on custom robotics, standard RL, and engineering-domain MDPs further suggest that the same modeling pipeline can be reused beyond the main benchmark suite.
Abstract: Long-context inference in large language models is bottlenecked by the quadratic cost of full attention. Existing efficient alternatives often rely either on native sparse training or on heuristic token eviction, creating an undesirable trade-off among efficiency, training cost, and accuracy. In this work, we show that full-attention LLMs are already intrinsically sparse and can be transformed into highly sparse models with only minimal adaptation. Our approach is built on three observations: (1) only a small subset of attention heads truly requires full long-context processing; (2) long-range retrieval is governed primarily by a low-dimensional subspace, allowing relevant tokens to be retrieved efficiently with a 16-dimensional indexer; and (3) the useful token budget is strongly query-dependent, making dynamic top-p selection more suitable than fixed top-k sparsification. Based on these insights, we propose RTPurbo, which retains the full KV cache only for retrieval heads and introduces a lightweight token indexer for sparse attention. By exploiting the model's intrinsic sparsity, RTPurbo achieves sparsification with only a few hundred training steps. Experiments on long-context benchmarks and reasoning tasks show that RTPurbo preserves near-lossless accuracy while delivering substantial efficiency gains, including up to a 9.36× prefill speedup at 1M context and about a 2.01× decode speedup. These results suggest that strong sparse inference can be obtained from standard full-attention training without expensive native sparse pretraining.
PaperID: 7473, Poster
Authors:
Boxiang Tao, Lei Guo, binwang, Zexin Wang, HongboDouAbstract: Model-based reinforcement learning (MBRL) promises high sample efficiency but is often crippled by compounding model errors: small prediction inaccuracies accumulate rapidly over multi-step rollouts, leading to biased policy optimization. We propose FTA, a trajectory-level generative framework using conditional flow matching (CFM) instead of step-wise dynamics. To avoid curved, energy-inefficient paths, we introduce an optimal transport regularizer grounded in the Benamou–Brenier formula. This regularizer minimizes kinetic energy while enforcing exact state displacement, yielding straight, low-energy trajectories with provable linear error propagation. Since a static generator fails in online RL due to (i) shifting data distributions and (ii) mismatch between training (historical) and desired (high-return) conditionals, we add periodic retraining from the replay buffer and a value-guided ODE sampler that biases generation toward high returns via the current Q-function. Theoretically, the regularizer drives the flow toward the Wasserstein-2 optimal transport map, and value-guided sampling reduces asymptotic Q-estimate bias. On MuJoCo benchmarks, FTA consistently outperforms strong MBRL and model-free baselines in sample efficiency and final performance, and generalizes across tasks. By mitigating compounding errors at the trajectory level and adapting to online shifts, FTA offers a robust, data-efficient alternative to traditional dynamics models.
PaperID: 7474, Poster
Abstract: To advance the development of embodied intelligence, Open-Environment Active 3D Mapping has attracted increasing attention, aiming to perform a long-horizon and shortest trajectory exploration for reconstructing unseen scenarios. Since only limited information about unseen environments is available, methods built on the closed-set assumption, i.e., assuming that the test environments are similar to those seen during training, cannot generalize satisfactorily. In existing active mapping methods, long-horizon exploration is often guided by predicting a coarse long-range goal and then converting it into an executable path. However, this stage is usually formulated as single-point prediction. Under partial observability, the same local observation may correspond to multiple plausible exploration directions, making such deterministic prediction prone to brittle decisions and degraded performance in unseen scenarios. Our experiments further verify that this is a key factor underlying their weak generalization. To address this issue, we reformulate long-horizon target prediction as conditional multimodal anchor generation using Conditional Flow Matching. Instead of predicting a single goal, our method learns a conditional distribution over coarse exploration anchors from the current mapping state. These anchors are first converted into executable candidate paths through obstacle-aware planning. We then apply exploration-mode clustering to compress geometrically similar trajectories and reduce candidate redundancy. Finally, a hierarchical selection module selects the most promising mode and reranks paths within it to produce the final executable trajectory. Experiments show that our method improves generalization and reconstruction efficiency in open environments.
Authors: Haris Moazam Sheikh
Abstract: Multi-objective Bayesian optimization (MOBO) is commonly approached through specialized acquisition functions or scalarization schemes designed to explicitly account for trade-offs among non-preferential objectives. In this work, we show that such complexity might be unnecessary. We propose a framework that extends standard single-objective acquisition functions directly to the multi-objective setting through a hypervolume-based transformation. We further extend hedge strategies for acquisition functions, which are typically used only in single-objective optimization, to the multi-objective regime. Our approach requires minimal modification to existing Bayesian optimization pipelines and avoids the need for bespoke multi-objective formulations. We demonstrate how a broad class of commonly used single-objective acquisition functions and hedge strategies can be adapted in a principled manner to handle multiple objectives, while preserving their intuitive interpretation and computational efficiency. Empirically, we evaluate the proposed methods across a range of synthetic and real-world multi-objective benchmarks. Despite their simplicity, our extensions consistently match or outperform more complex state-of-the-art MOBO methods in terms of optimization performance and sample efficiency. These results suggest that effective multi-objective Bayesian optimization can be achieved by reusing and carefully extending well-established single-objective acquisition strategies, offering a simpler and more flexible alternative to existing approaches.
PaperID: 7476, Poster
Abstract: Reinforcement learning has rapidly become the dominant paradigm for training software engineering (SWE) agents on long-horizon, multi-turn tasks such as resolving real-world GitHub issues. Yet existing pipelines suffer from a fundamental credit-assignment bottleneck: rewards arrive only at the end of trajectories that may span dozens of file edits, shell commands, and test invocations, leaving every intermediate decision indistinguishable from luck. We observe that containerised SWE tasks provide a particularly useful form of environment forkability: a commit hash plus a container image specifies a real executable software state that can be restored cheaply and exactly, with auditable provenance and without a simulator-to-reality gap. We exploit this property by introducing Counterfactual Rollout Replay (CRR), a training-time procedure that re-executes the environment from selected decision points with alternative actions and assigns each selected step an advantage equal to the sampled on-policy versus counterfactual return differential. CRR requires no learned process reward model, no human rubric annotation, and no oracle hindsight signal; it is free of auxiliary process labels and reward-model training, while still incurring extra environment-replay compute. Applied to a 14B base model on SWE-Gym and evaluated on SWE-bench Verified and SWE-rebench, CRR-trained agents reach higher resolution rates with substantially fewer trajectories than outcome-only GRPO and complement orthogonal advances such as PRM-based scoring and trajectory search.
Abstract: A plethora of protein language models have been released in recent years. Yet comparatively little work has addressed how to best sample from them to optimize desired biological properties. We fill this gap by proposing a flexible, effective sampling method for masked language models (MLMs), and by systematically evaluating models and methods both in silico and in vitro on actual antibody therapeutics campaigns. Firstly, we propose sampling with stochastic beam search, exploiting the fact that MLMs are surprisingly efficient at evaluating the pseudo-perplexity of the entire 1-edit neighborhood of a sequence. Reframing generation in terms of entire-sequence evaluation enables flexible guidance with multiple optimization objectives. Secondly, we report results from our extensive in vitro head-to-head evaluation for the antibody engineering setting. This reveals that choice of sampling method is at least as impactful as the model used, motivating future research into this under-explored area.
Abstract: While multimodal large reasoning models (MLRMs) have exhibited impressive capabilities, they remain prone to hallucinations, and effective solutions remain underexplored. In this paper, we investigate the underlying causes of hallucinations in MLRMs and accordingly propose a reasoning-centric training framework for mitigation. Specifically, we find that introducing reasoning mechanisms exacerbates models' reliance on language priors and overlooks visual inputs, leading to CoTs with reduced visual cues but redundant text tokens. Guided by the Information Bottleneck (IB) principle, we propose selectively filtering redundant thinking tokens to obtain a more compact, signal-efficient CoT that preserves task-relevant information while suppressing noise. We also observe that the quality of the reasoning trace largely determines whether hallucination emerges in final answers. To leverage this, we introduce a reasoning-enhanced preference optimization scheme that constructs training pairs using high-quality AI feedback. We further propose generating high-quality negative samples for contrastive optimization via a multimodal hallucination-inducing mechanism that exposes model hallucination behaviors through carefully crafted visual and textual inducers. By modeling CoT chains as intermediate representations between inputs and final answers, we establish an IB-based theoretical justification for our method. Extensive experiments reveal consistent hallucination reduction across diverse MLRMs and benchmarks.
Abstract: Graph Neural Networks (GNNs) often fail to capture the link-specific structural patterns essential for accurate link prediction, since their node-centric message passing might overlook the subgraph structures connecting two nodes. Prior attempts to inject such structural context either suffer from high computational cost or rely on oversimplified heuristics (e.g., common neighbor counts) that cannot capture multi-hop dependencies. We propose SP4LP (Shortest Path for Link Prediction), a new framework that integrates GNN-based node encodings with sequence modeling over shortest paths. Specifically, SP4LP first computes node representations with a GNN, then extracts the shortest path between each candidate node pair and processes the sequence of node embeddings with a sequence model. This design allows SP4LP to efficiently capture expressive multi-hop relational patterns. Theoretically, we show that SP4LP is strictly more expressive than both standard message-passing GNNs and several leading structural feature methods, positioning it as a general and principled framework for link prediction in graphs. Empirically, SP4LP sets a new state of the art on many standard link prediction benchmarks.
PaperID: 7480, Poster
Abstract: Efficient long-video understanding with vision-language models (VLMs) has largely been treated as informative token selection at fixed native resolution: which frames or which visual tokens to retain under a token budget. We argue this framing leaves two axes unused --- per-frame resolution itself can be traded for denser temporal coverage, and front-end decoding latency scales with the candidate pool, not the token budget. Through an empirical study across multiple VLMs and long-video benchmarks, we distill three lessons: (i) dense low-resolution sampling outperforms sparse native-resolution sampling at matched token budgets; (ii) some tasks are resolution-sensitive and benefit from high-resolution frames; and (iii) front-end decoding dominates wall time on hour-scale clips. These lessons motivate LoHi, a training-free, single-pass framework that pairs a dense low-resolution video stream with a sparse set of high-resolution image streams, processed through the VLM's native video and image pathways. Two plug-and-play selectors choose Hi-I frames at near-zero or low overhead: LoHi-Anchor uses codec-level I-frame metadata, and LoHi-SemDiv uses a query-relevance and visual-diversity DPP over CLIP features. On three long-video benchmarks, LoHi improves over the vanilla native-resolution baseline by +10.6% on average at matched token budget and over the strongest prior efficiency methods by +5.2%, while reducing front-end decoding latency by up to 7x on hour-scale clips.
PaperID: 7481, Poster
Authors:
Guanyu Cui, Yuhe Guo, Zhewei Wei, Hsin-Hao SuAbstract: Message-passing graph neural networks (MPGNNs) are commonly compared with the Weisfeiler-Lehman (WL) color-refinement procedure, but this comparison does not quantify the resource parameters a network needs to realize color refinement with bounded-size messages and finite numerical precision. We study the cost of simulating a single color-refinement step on unattributed graphs. We distinguish input-independent, or oblivious, simulation from instance-dependent simulation. In the former, the parameters, or their distributions in randomized models, are fixed before the input instance is known. Our results show that the local form of WL color refinement hides a global relabeling problem. In the oblivious setting, deterministic and zero-error randomized MPGNNs cannot solve this problem in the worst case using only shallow networks with small messages. We complement this lower bound with a nearly matching construction in a stronger rooted, port-aware model. By contrast, when the color set is large, bounded-error randomness can greatly reduce the cost, and a one-layer MPGNN with messages of logarithmic size and a logarithmic number of random bits suffices. We show that this logarithmic number of random bits is essentially necessary for shallow, small-message simulations. When the color set is small, we still obtain a rooted, port-aware simulation, but this construction requires more layers or larger messages. We also prove that this extra cost is partly unavoidable, as small color sets force a nontrivial trade-off between the number of layers and the message size. Finally, instance-dependent simulation can be much shallower, but the required instance-specific parameters are not necessarily easy to find. Together, these results reveal quantitative structure hidden behind the statement that MPGNNs match WL color refinement.
PaperID: 7482, Poster
Authors: Zhikang Xu, Jiarui Xing, Jian Wang
Abstract: We present a robust deep clustering framework for unsupervised visual discovery that bridges the gap between statistical robustness and mathematical tractability. Current generative approaches typically rely on either noise-prone Gaussian priors or computationally intensive diffusion models; however, neither paradigm simultaneously offers explicit latent structures and exact analytic optimization. While diffusion models generate strong priors, their reliance on iterative sampling and intractable likelihoods limits their scalability and interpretability in clustering tasks. Our approach addresses these limitations by constructing a latent space governed by a Student's t distribution via a Bayesian scale-mixture formulation. This yields the first fully closed-form augmented evidence lower bound (ELBO) for this model class, effectively eliminating the biased variational approximations and high computational costs inherent in prior heavy-tailed or diffusion-based frameworks. By treating precision as a latent variable, our objective automatically modulates the influence of each data point; it downweights outliers to learn a data-dependent measure of trust. This mechanism leads to superior robustness and more stable manifold learning compared to state-of-the-art approaches. We demonstrate the resilience of our framework on standard vision benchmarks and complex neuroimaging tasks. In real-world brain magnetic resonance imaging (MRI)-driven neurodegenerative disease analysis, our model successfully recovers clean and anatomically coherent clusters where clinical noise typically collapses rare subtypes. These results underscore the ability of our framework to preserve clinically crucial morphological subtypes, allowing precise discovery within clinical interventions.
Authors:
Hongjie Jiang, Di LuoAbstract: Accurately solving time-dependent partial differential equations (PDEs) with neural networks remains challenging due to long-time error accumulation and the difficulty of enforcing general boundary conditions. We introduce TENG-BC, a high-precision neural PDE solver based on the Time-Evolving Natural Gradient, designed to perform under generic boundary constraints. At each time step, TENG-BC performs a boundary-aware optimization that jointly enforces interior dynamics and boundary conditions, accommodating Dirichlet, Neumann, Robin, and mixed types within a unified framework. This formulation admits a natural-gradient interpretation, enabling stable time evolution without delicate penalty tuning. Across benchmarks over diffusion, transport, and nonlinear PDEs with various boundary conditions, TENG-BC achieves solver-level accuracy under comparable sampling budgets, outperforming conventional solvers and physics-informed neural network (PINN) baselines.
Abstract: Dynamic link prediction requires determining whether the historical interactions of a source node provide reliable structural support for a specific target. Existing target-conditioned methods capture such support through structural heuristics such as repeated interactions or co-neighbor overlap, but typically instantiate them as identity-based encodings and inject them as passive token-level features before aggregation. This design is brittle when exact structural encodings are sparse, and the resulting structural signals can be diluted after being concatenated with high-dimensional time and edge features, while their dynamics remain unmodeled. In this paper, we revisit target-conditioned dynamic link prediction from a structural dynamics modeling perspective and propose TCSD, which formulates structural signals as target-conditioned states, evolves them over recent histories, and uses them to guide aggregation. Concretely, we construct hybrid structural states from exact indicators that preserve time-aware repeat and co-neighbor matches, and relaxed indicators that recover latent support beyond exact identity matching. TCSD further models the dynamics of these states to capture temporal consistency within indicators and complementarity across indicators. Instead of treating structural states as passive features, TCSD uses the learned dynamics as aggregation controllers to amplify target-relevant interactions and suppress irrelevant ones. Experiments on sixteen dynamic graphs show that TCSD outperforms ten baselines, achieving up to a 14.27% relative improvement in MRR.
PaperID: 7485, Poster
Abstract: Value iteration (VI) is the unit-step explicit-Euler discretization of a continuous-time Bellman residual flow \dot V=\operatornameBR(V), and its standard linear dependence on the effective horizon 1/(1-\gamma) can be viewed as reflecting the step-size restriction that explicit integration imposes on stiff dynamics. We propose Implicit-Euler Value Iteration (IE-VI): a planning routine equivalent to the implicit-Euler discretization of the same flow, in which each step is unconditionally stable and contracts at rate 1/(1+h(1-\gamma)) for every step size h>0. The implicit step requires solving an equation in the true Bellman operator, which we realize at finite cost by iterative defect correction: a short sequence of cheap planning rounds in an approximate model \hat\mathcalP \approx \mathcalP (a simulator or a learned dynamics model), with the true model queried once per round to refresh a residual correction term. For Policy Evaluation and Control, IE-VI converges to the value function of \mathcalP for any \hat\mathcalP, however inaccurate---in contrast to model-based planners, which generally converge to a biased fixed point. When the model error shrinks at a sublinear-in-horizon rate, IE-VI attains polylogarithmic iteration complexity in 1/(1-\gamma), an exponential improvement over VI; outside this regime its complexity stays within a logarithmic factor of VI. We also give an adaptive variant A-IE-VI, which removes the need to know the model-error scale, and a Dyna-style sample-based variant IE-Dyna for tabular RL.
PaperID: 7486, Poster
Abstract: How does data shape language model (LM) behavior throughout pretraining? We investigate this question through a case study on entity comparison, e.g., . We begin with controlled experiments in which we train small LMs (124M parameters) on mixtures of natural text from pretraining corpora and synthetic data from entity comparison tasks. We identify three distinct phases of learning: (1) an early phase where the LM selects entities by frequency, (2) a middle phase where the LM selects entities by position in a prompt (first vs. last), and (3) a late phase where the LM selects the entity that is the correct answer to the question. We show that the emergence of these three phases is controlled by statistical properties of the training data. With small amounts of task-specific synthetic data, we observe only the first two phases and the model fails to learn the task; with large amounts, the model jumps directly from the frequency-based heuristic to solving the task correctly. Moreover, if we modify the frequency of entities in data from a naturally occurring Zipfian distribution (a small number of entities are very common and the vast majority are rare) to a uniform distribution, the first phase disappears and the model learns the task more quickly. Finally, we find the same three phases of learning in the pretraining of open-sourced OLMo models. Together, our findings demonstrate that properties of pretraining data are causal drivers of heuristic learning and show that small-scale synthetic experiments can predict training dynamics at larger scales.
PaperID: 7487, Poster
Abstract: Chain-of-thought (CoT) prompting and related strategies that elicit intermediate reasoning traces have fundamentally changed how language models are used, evaluated, and developed. CoT now touches many areas of contemporary research beyond prompting, from advanced test-time inference and model tuning to interpretability, yet much of this work treats CoT traces in task-specific ways, without a shared formal account of what they are or how they compose and support inference. We argue that probabilistic programs — ordinary programs extended with stochastic choices and probabilistic conditioning — provide a natural framework for modeling CoT reasoning. To demonstrate this, we derive a small discrete probabilistic programming language for CoT in which reasoning steps are stochastic choices, branches encode dependencies, trace likelihoods score executions, and answers are return values. Even with a minimal reachability-based semantics, this language enables a first-principles analysis of trace likelihood, exposing why likelihood alone can be a misleading model of reasoning. Motivated by this limitation, we add differentiable operators for composing scores and feedback over CoT traces. The resulting framework links several uses of CoT — scoring, evidence aggregation, feedback propagation, model tuning, and step-level sensitivity analysis — within a single probabilistic program. Through case studies, we show how this approach yields new analysis techniques and test-time inference strategies, offering a promising direction for programming-language-based approaches to language-model reasoning.
PaperID: 7488, Poster
Abstract: Optimal transport (OT) provides a principled framework for mapping between probability distributions. Despite extensive progress, applying OT to large-scale data remains computationally demanding, and the resulting pointwise transport plans are often difficult to interpret. We introduce Optimal Mixture Transport (OMT), a scalable framework that shifts the transport paradigm from individual samples to mixtures of subpopulations, reformulating the transport problem as a strictly biconvex optimization with a unique global minimizer. We further establish theoretical guarantees on the stability of the OMT map, showing that bounded perturbations of the underlying distributions lead to bounded changes in the transport plan. By formulating subpopulations as exponential-family distributions, OMT decouples computational complexity from the sample size, scaling solely with the number of mixture components. We demonstrate the effectiveness and practicality of OMT on a wide range of synthetic benchmarks and real-world datasets, including image data and large-scale single-cell RNA sequencing measurements.
Abstract: State Space Models (SSMs) have emerged as a compelling alternative to Transformers, enabling sequence modeling with constant memory and linear compute. While SSMs excel in certain settings like length generalization, they continue to lag behind Transformers on tasks that require in-context learning and precise retrieval, slowing their adoption for large-scale language modeling. In this work, we demonstrate that both the success and failure of SSMs in these domains can be explained by studying the role of the gating mechanism, a prevalent component in modern recurrent networks. Specifically, we show through theory and experiments that this gating mechanism causes SSMs to first learn an in-weights ``memorization'' solution, while delaying, or even preventing, convergence to a correct in-context learning solution. Importantly, this happens even in cases where there are no fundamental limitations due to the architecture or its memory capacity. On the other hand, we find that gating is often beneficial for improving generalization to long sequence lengths. Our results illuminate the crucial role of the gating mechanism in shaping both the training dynamics and generalization of SSMs, and provide a basis for understanding and improving linear-time models.
Authors: Salini Yadav, Taveena Lotey, Pravendra Singh, Partha P Roy
Abstract: Zero-shot visual decoding from electroencephalography (EEG) aims to infer visual semantics from non-invasive neural recordings, but remains challenging due to the low signal-to-noise ratio, non-stationarity, and limited spatial resolution of EEG. Existing EEG–vision alignment methods often rely on holistic EEG embeddings, which can obscure the complementary temporal, spectral, and spatial structure underlying visual perception. We introduce a unified multiview EEG representation learning framework for aligning brain responses with visual semantic embeddings. Our method builds an EEG encoder that jointly models three complementary views: input-conditioned state-space temporal dynamics, learnable wavelet-based spectral decomposition for sample-adaptive frequency modeling, and attention-modulated graph learning for structured electrode interactions. The resulting multiview EEG embeddings are fused and aligned with pretrained visual representations in a shared semantic space using contrastive learning with EEG-specific regularization, enabling 200-way zero-shot visual classification. Experiments on THINGS-EEG benchmark show that our method achieves state-of-the-art performance, with 54.8% Top-1 and 85.6% Top-5 accuracy in the within-subject setting and 15.3% Top-1 and 45.4% Top-5 accuracy in the cross-subject setting. We further present the first systematic cross session EEG–image decoding evaluation, achieving 40.8% Top-1 and 78.0% Top-5 accuracy. These results suggest that explicitly modeling multiview neural structure improves both semantic alignment and generalization in EEG-based visual decoding.
Abstract: Rehearsal-free continual learning with parameter-efficient adapters can be cast as a sequence of task-vector write-in operations: for each new task, a low-rank adapter is learned and merged into a running model. We propose Proximity Regularized Merging (PRM), a minimal modification to sequential LoRA merging that adds a proximal penalty during task-vector training without changing the subsequent write-in rule. PRM acts as a robust task-vector regularizer: in the reported Base\rightarrow+Prox diagnostics, it improves AAA across multiple write-in rules, backbones, and class-incremental settings, while its fixed-coefficient variant remains competitive with strong coefficient-based baselines. Mechanistically, matched-prefix norm controls and proximal-strength sweeps show that proximal training shrinks the task-vector radius, lowers Fisher-weighted interference, broadens the coefficient plateau, and exposes a stability--plasticity trade-off. Together, these results suggest that the effectiveness of sequential LoRA merging depends not only on how much of a task vector is written in, but also on whether the task vector itself has been trained to be mergeable.
Abstract: Diffusion language models (DLMs) promise parallel, order-agnostic generation, but on standard benchmarks they have historically lagged behind autoregressive models in sample quality and diversity. Recent continuous flow-matching approaches over one-hot and embedding spaces have narrowed this gap, suggesting continuous state spaces are highly effective for language. In this work, we further close the autoregressive gap by modeling text as a continuous diffusion process over fixed-width bitstreams. Our approach represents semantic tokens as analog bit sequences and utilizes a matched-filter residual parameterization to isolate contextual learning from analytic independent-bit posteriors. Crucially, our empirical comparisons reveal that while deterministic continuous DLMs are competitive, they can be overly contractive, undershooting real-data entropy. We demonstrate that the remaining quality-diversity gap is bridged by a stochastic sampler that applies Langevin-type corrections gated by the entropy-rate profile, automatically concentrating stochasticity in high-information regions while remaining nearly deterministic elsewhere. On the One Billion Word Benchmark (LM1B), our 130M-parameter bitstream model reaches a generative perplexity (GenPPL) of 59.76 at matched real-data entropy (4.31) using 256 neural function evaluations (NFEs), decisively outperforming prior DLM baselines and reaching the autoregressive reference. On OpenWebText (OWT), our stochastic sampler establishes a new continuous-DLM Pareto frontier, achieving GenPPL = 27.06 at an entropy of 5.26 using 4x fewer steps than previous 1024-NFE baselines. As an additional architectural benefit, bitstream diffusion removes the O(V) vocabulary scaling bottleneck shared by standard DLMs. By predicting O(log V) bitwise logits via semantic bit-patching, our model yields a reduced memory footprint and higher throughput, demonstrating a scalable paradigm for language generation as vocabulary sizes grow.
Authors: Younes Essafouri, Laure Raynaud, Luciano DROZDA, Laurent Risser
Abstract: As the demand to integrate Artificial Intelligence into high-stakes environments continues to grow, explaining the reasoning behind neural-network predictions has shifted from a theoretical curiosity to a strict operational requirement. Our work is motivated by the explanations of autoregressive neural predictions on dynamic physical fields, as in weather forecasting. Gradient-based feature attribution methods are widely used to explain the predictions on such data, in particular due to their scalability to high-dimensional inputs. It is also interesting to remark that gradient-based techniques such as SmoothGrad are now standard on images to robustify the explanations using pointwise averages of the attribution maps obtained from several noised inputs. Our goal is to efficiently adapt this aggregation strategy to dynamic physical fields. To do so, our first contribution is to identify a fundamental failure mode when averaging perturbed attribution maps on dynamic physical fields: stochastic input perturbations do not induce stationary amplitude noise in attribution maps, but instead cause a geometric displacement of the attributions. Consequently, pointwise averaging blurs these spatially misaligned features. To tackle this issue, we introduce WassersteinGrad, which extracts a geometric consensus of perturbed attribution maps by computing their entropic Wasserstein barycenter. The results, obtained on regional weather data and a meteorologist-validated neural model, demonstrate promising explainability properties of WassersteinGrad over gradient-based baselines across both single-step and autoregressive forecasting settings.
Abstract: Abstract: Mode collapse --- the failure to capture one or more modes when targetting a multimodal distribution --- is a central challenge in modern variational inference. In this work, we provide a mathematical analysis of annealing-based strategies for mitigating mode collapse in a tractable setting: learning a Gaussian mixture, where mode collapse is known to arise. Leveraging a low-dimensional summary statistics description, we precisely characterize the interplay between the initial temperature and the annealing rate, and derive a sharp formula for the probability of mode collapse. Our analysis shows that an appropriately chosen annealing scheme can robustly prevent mode collapse. Finally, we present numerical evidence that these theoretical trade-offs qualitatively extend to neural network–based models, RealNVP normalizing flows, providing guidance for designing annealing strategies mitigating mode collapse in practical variational inference pipelines.
PaperID: 7495, Poster
Abstract: Time series forecasting is an actively researched problem with diverse applications. Recent work has shown that simple linear models can compete with complex deep learning architectures in terms of forecasting accuracy. This has led to the design of several small models, yet the reasons behind their effectiveness remain insufficiently understood. We address this gap through a theoretical framework grounded in reduced-rank regression and kernel analysis. We show both analytically and empirically that the intrinsic complexity of many TSF tasks is lower than commonly assumed and that the performance of linear models is linked to their alignment with low-rank data structures. Building on these insights, we propose a new design for low-rank neural networks that incur lower computational cost than linear models. Specifically, we introduce a simple filter bank architecture that provides a principled way to control model complexity through tunable and interpretable hyperparameters that directly correspond to the rank of the forecasting model. The proposed architecture also serves as a versatile building block for constructing deeper networks. Our evaluation shows that filter bank architectures achieve state-of-the-art results on most long-term TSF benchmarks, with lower computational cost than recently proposed solutions.
PaperID: 7496, Poster
Authors: EunChan Kim, Jeongwan Shin, Jae-Ho Choi
Abstract: Masked Autoencoders (MAE) have emerged as a powerful self-supervised paradigm in vision, yet their direct application to radar perception remains bottlenecked by a fundamental domain gap. Specifically, while conventional MAEs are optimized for RGB images using isotropic grid patching and standard mean squared error (MSE) loss, radar range-azimuth (RA) heatmaps exhibit sinc-oriented anisotropic spatial structures as well as complex physical artifacts, such as heavy-tailed multipath interference and sidelobe spikes. In this paper, we propose RadarMAE, the first radar-centric MAE tailored specifically to the physical properties of radar signals. We present an axial tokenization strategy that preserves the resolution disparities and continuous sinc-shaped lobes of RA maps. Furthermore, considering that the range and azimuth dimensions exhibit distinct physical artifacts, we decouple their reconstruction objectives by introducing a multipath distribution loss along the range axis to capture heavy-tailed multipath reflections as well as a sidelobe distribution loss along the azimuth axis to learn deterministic sidelobe spikes. Our extensive experiments demonstrate that injecting these domain-specific physical priors enables robust representation learning for radar object detection. Notably, our method achieves a state-of-the-art 64.17% bird's-eye-view (BEV) AP on the K-Radar dataset using only a single-frame RA map setting, significantly outperforming the previous baselines.
PaperID: 7497, Poster
Authors:
Evelyn Zhang, Fufu Yu, Hanjun Li, Aoqi Wu, Ke Yan, Shouhong Ding, Tianhe Ren, Xinting Hu, Xiaojuan Qi, Li Jiang, Jiaya JiaAbstract: Existing multimodal large language models (MLLMs) typically formulate object detection as a conditional sequence generation task. Under this paradigm, spatial locations are commonly represented by either textual coordinates or special quantized learnable tokens. However, textual coordinates are tokenized as unrelated symbols, spatially adjacent points (e.g., x=199 vs. y=200) yield entirely different tokens, so the cross-entropy loss is misaligned with true spatial distance. Special quantized location tokens, on the other hand, inevitably introduce quantization error. Reducing this error requires linearly more tokens, but a larger vocabulary is harder to train and converges slower. We argue that both limitations stem from the formulation of spatial representation. To address this issue, we propose a simple yet novel localization representation method, HiLoc, which predicts locations in a coarse-to-fine manner with special tokens: first a grid index for coarse localization, then a cell index to refine the position within the grid. Under this paradigm, accuracy can be scaled up by simply adding the hierarchy level. We also build HiRef-500K, a dataset consisting negative and complex referring samples to strengthen rejection and complex referring ability. Experiments show that our method achieves strong performance on both detection and grounding benchmarks, especially on small objects, under high IOU thresholds and complex referring tasks.
Authors:
Jasper Lu, Zhenhao Shen, Yuanfei Wang, Shugao Liu, Shengqiang Xu, Shawn Xie, Jingkai Xu, Feng Jiang, Jade Yang, Chen Xie, Ruihai WuAbstract: Learning robust robot policies in real-world environments requires diverse data augmentation, yet scaling real-world data collection is costly due to the need for acquiring physical assets and reconfiguring environments. Therefore, augmenting real-world scenes into simulation has become a practical augmentation for efficient learning and evaluation. We present WorldComposer, an automated generative Real2Sim2Real framework that maps real-world panoramas into high-fidelity, task-ready simulation environments and further synthesizes diverse Digital Cousins through scene and object variations. Combined with high-quality physics engines and realistic assets, WorldComposer supports interactive manipulation across rigid, articulated, and deformable objects. Additionally, we incorporate multi-room stitching to construct consistent large-scale environments for long-horizon tasks. Experiments demonstrate a strong sim-to-real correlation validating our platform's fidelity, and show that extensively scaling up data generation leads to significantly better generalization to unseen scene and object variations, demonstrating the effectiveness of Digital Cousins for generalizable robot learning and evaluation.
PaperID: 7499, Poster
Abstract: Chain-of-Thought reasoning improves the explainability of large language models, yet verifying such reasoning remains difficult when intermediate steps express epistemic uncertainty rather than deterministic entailment. Neuro-symbolic verifiers typically translate reasoning steps into rigid first-order or SMT-style constraints. This forces each proposition into a binary truth regime and creates a Rigidity--Ambiguity Mismatch: uncertainty-marked claims such as may, likely, or suggests are either over-hardened into deterministic assertions or rejected as unverifiable. We propose ProbVeri, a probabilistic soft-logic framework for formal verification under uncertainty. The key idea is to represent uncertainty-marked reasoning steps as weighted soft logical constraints over continuous truth assignments, so that partial evidential support remains expressible. Building on this soft-logic representation, we formulate verification as a contrastive energy minimization problem: a step is considered supported when its natural interpretation has lower verification energy than its negated counterpart under the same evidence, soft rules, and assumption regularization. Probabilistic Soft Logic defines the logical energy through weighted hinge-loss constraints. We introduce a support-grounded assumption penalty that discourages unsupported latent bridge assumptions between context and claim. The resulting energy gap provides a confidence-scored verification signal that separates unsupported steps from weakly but consistently supported inferences. Experiments on ProofWriter, BioASQ, and LegalBench-SARA show that ProbVeri recovers more verified-correct reasoning traces than rigid verifiers while reducing over-acceptance compared with soft-only PSL. These results suggest that uncertainty-aware verification benefits from soft logical formalization and contrastive comparison between natural and negated interpretations.
PaperID: 7500, Poster
Abstract: Modern end-to-end autonomous driving systems often suffer from fragile out-of-distribution (OOD) generalization and limited decision transparency, largely due to their reliance on raw sensory observations and susceptibility to spurious visual correlations. Although memory-augmented approaches have been explored to improve generalization, existing methods typically store raw observations, symbolic rules, or entangled latent features, making retrieval sensitive to surface-level similarity rather than decision-relevant driving structure. To address this limitation, we propose DriveMem, a plug-and-play framework that moves beyond raw observations by distilling and storing invariant driving memories for generalizable driving systems. DriveMem integrates two core components: Invariant Feature Abstraction that first distills perturbation-stable, decision-relevant representations from visual tokens, and Prototype Memory Bank that then stores and retrieves prototypical driving experiences grounded in invariant representations. Extensive open-loop and closed-loop experiments demonstrate that DriveMem improves OOD generalization and safety-related planning metrics across diverse unseen scenarios. Notably, under severe sensor disruptions, DriveMem achieves up to a 39.5% relative reduction in Collision Rate compared with a strong domain-generalization baseline. These results suggest that building prototype memory in an invariant representation space is a promising approach for improving robustness, while offering prototype-grounded traceability for model behavior.
Authors:
Jérémie Dentan, Alexi Canesse, Mahammed E Sharkawy, Sonia VanierAbstract: Claim-level Uncertainty Quantification (UQ) aims to mitigate the lack of reliability of Large Language Models (LLMs) by evaluating the factuality of each claim in their outputs. We introduce Kernel Token Contradiction (KTC), a lightweight approach to compute claim-level UQ under realistic white-box conditions. KTC represents the candidate tokens involved in LLM generation as a positive semi-definite kernel that integrates both the LLM’s conditional distribution and a token contradiction score. We then use the Von Neumann entropy to quantify the uncertainty of this kernel. To estimate token contradiction, we develop a new approach based on frequency statistics from the Wikipedia corpus. Although CPU-only, our approach achieves over an 8.2× speedup compared to state-of-the-art GPU-accelerated methods based on cross-encoders, and over a 65× speedup compared to CPU-only methods with comparable performance. Our evaluation spans two benchmarks across four European languages and 16 different models. KTC not only matches the average performance of existing methods but also outperforms them in high-precision regimes. This combination of computational efficiency and accuracy makes real-time monitoring of LLM outputs practical in production.
Abstract: Interactive streaming music generation promises the use of generative models for live performance and co-creation that is impossible with offline models. However, SOTA models exist in the discrete-AR regime, requiring industrial levels of compute for both training and inference. In this work, we investigate whether audio diffusion models, with their wide support in the open-source community but non-streaming bidirectional nature, can be repurposed efficiently into interactive models accessible on consumer hardware. By taking a critical look at the modern pipeline for block-wise outpainting diffusion, we identify critical inefficiencies during inference that result in strictly worse computational efficiency than their discrete-AR counterparts. We propose Live Music Diffusion Models (LMDMs), a simple modification of the generative diffusion process that recovers, and then outperforms, the inference complexity of the discrete Live Music Models (LMMs) through block-wise KV Caching. Unlike LMMs, LMDMs further enable stable post-training alignment through our novel ARC-Forcing paradigm, reducing error accumulation without any explicit RL or reward models. We demonstrate the application of LMDMs in a number of creative domains, including text-conditioned generation, sketch-based music synthesis, and jamming. We finally show how LMDMs can be used as a generative instrument in a real artist-AI collaboration, utilizing LMDMs as a “generative delay” to transform musicians’ improvisation live for variable timbral effects while running locally on a consumer gaming laptop.
PaperID: 7503, Poster
Authors:
Fang Li, Shaoqing Xu, Yuechen Luo, Hanbing Li, Zhi-Xin YangAbstract: Vision-Language-Action (VLA) models offer a promising paradigm for end-to-end autonomous driving by connecting visual scene understanding, linguistic reasoning, and low-level trajectory generation. However, existing driving VLAs typically use language only as static commands or supervised rationales, while treating trajectory prediction as an open-loop generation problem. As a result, they lack an explicit mechanism for understanding the consequences of their own actions, which limits robustness under closed-loop distribution shift and makes reinforcement learning inefficient in long-tail scenarios. We propose SelfCritic-VLA, a self-evaluative VLA framework that unifies trajectory generation, language-grounded action critique, and trajectory refinement. Given a driving scene, the model first generates an initial trajectory, then produces a language-grounded critique that describes the potential safety, progress, and rule-compliance consequences of executing that trajectory, and finally refines the trajectory conditioned on this critique. To train this capability, we construct a large-scale counterfactual trajectory dataset with multi-source candidate trajectories and semantic critiques, and perform hybrid supervised fine-tuning for planning, action evaluation, and refinement. Building on this initialization, we introduce a self-critic reinforcement learning procedure based on GRPO, where critique-conditioned refinements are filtered by grounded driving rewards and distilled back into the base policy. Experiments on NAVSIM and Bench2Drive show that SelfCritic-VLA consistently improves both open-loop and closed-loop driving performance over strong baselines, demonstrating the effectiveness of language-guided self-refinement for driving policy optimization.
PaperID: 7504, Poster
Abstract: Poisoning attacks against GraphRAG focus on knowledge pollution at the node and community levels, but conventional poisoning defenses mainly rely on sentence-level semantic features, which makes them less effective against such attacks. To address this issue, we analyze the poisoning strategy from a game-theoretic perspective and find that attackers prioritize fabricating query-related facts while ignoring supporting background knowledge, which in turn induces a systematic topological discrepancy between poisoned and clean subgraphs. Based on this insight, we propose the lightweight oisoning Attack on GraphRAG (TDP). TDP constructs a pair of conflicting candidate subgraphs from the retrieved evidence and trains a pairwise topology-ranking discriminator to distinguish clean evidence with dense cross-validation from poisoned evidence with sparse structural support, thereby removing poisoned subgraphs. Notably, the topological patterns captured by TDP reflect structural preferences induced by the poisoning game rather than dataset-specific distributional biases, allowing it to be transferred as a plug-and-play module after pretraining without end-to-end retraining. To the best of our knowledge, this is the first systematic work on poisoning defense for GraphRAG, and experiments show that TDP achieves state-of-the-art defense performance across multiple benchmarks and poisoning attacks, demonstrating strong generalization.
PaperID: 7505, Poster
Authors: Ruilan Wang, Francisco A Reina, Katerina Papadaki
Abstract: Double Oracle (DO) and PSRO approximate Nash equilibria in two-player zero-sum games by iteratively expanding a strategy population with best responses. Online Double Oracle (ODO) interleaves multiplicative weights updates (MWU) with discovery and obtains an O(\sqrtk \ln k / T) rate. We identify the \emphreinitialisation step as an underexplored design choice: at each expansion, MWU-based methods reset newly discovered strategies to a uniform prior. We introduce Virtual Double Oracle (VDO), which exploits a bilinear identity---a strategy's full hindsight loss history collapses to its utility against the time-averaged opponent---to initialise each new strategy with the MWU state it would have had from the start. Conditional on the discovered population, VDO recovers the fixed-population rate O(\sqrt\ln k / T) asymptotically, removing ODO's leading-order \sqrtk window-decomposition penalty. Discounted VDO adds a single parameter \beta to control finite-horizon prior strength. Experiments across matrix games, poker, and Goofspiel validate both contributions: pure VDO confirms the asymptotic rate, and discounted VDO gives the strongest finite-horizon performance.
Abstract: While language models remain frozen at their training state, the world evolves continuously. Knowledge editing has emerged as a key alternative to full retraining, but its deployment is bottlenecked by the erosion of core capabilities: mathematical and programmatic reasoning collapse while encyclopedic recall remains intact. We trace this asymmetric degradation to a distributional mismatch. Covariance-based editors preserve only the subspaces spanned by their reference corpus, but fail to capture the operative distribution shaped by post-training such as SFT and DPO. Static external corpora, including Wikipedia and even the original pretraining mixture, cannot recover this shifted manifold. We propose Memoir, which estimates the preservation covariance C directly from the model itself by sampling from its own decoding distribution. Seeding generation with a single random vocabulary token bypasses the instruction-following templates that otherwise dominate sampled outputs, exposing the broader subspaces the model has internalized. Memoir requires no external data and serves as a drop-in component for any covariance-based editor, a practical advantage given that the pre- and post-training corpora of most modern LLMs are not publicly accessible. Across OLMo-2, Llama-3.1, and Qwen-3 (7--8B), under both MEMIT and AlphaEdit and in batch and sequential regimes, Memoir consistently extends preservation in the most vulnerable domains, most strikingly on Qwen3-8B after 20,000 AlphaEdit batch edits, it retains 79.9% GSM8K accuracy compared to 10.9% with the Wikipedia baseline. These results suggest that aligning the preservation distribution with the model's operative distribution is a key factor in non-destructive editing, and that the model itself may be the most accessible source of that distribution for deployed systems.
PaperID: 7507, Poster
Authors: Rafael Silva, Maxime Sermesant
Abstract: Atrial fibrillation (AFib) is the most common sustained cardiac arrhythmia and a major risk factor for stroke, yet it often remains undetected until a first clinical event. Existing deep learning approaches typically estimate AFib risk from a single, cross-sectional ECG, discarding the longitudinal information contained in serial recordings. We propose a transformer-based survival model, SurviFormer, that leverages sequences of ECG representations to model patient-specific AFib risk over time. Our approach encodes irregular inter-visit time gaps as pairwise relative attention biases within a causally masked transformer, and outputs a discrete hazard function trained with a multi-landmark strategy. We train and evaluate our model on the CODE dataset (140k patients, 517k ECGs), with external validation on MIMIC-IV. Compared to static baselines (DeepSurv and DeepHit applied to ResNet-50 embeddings) and a dynamic competing-risk model (DynamicDeepHit), our model achieves a C-index of 0.853, a mean TD-AUC of 0.867, and an IBS of 0.030 on the internal test set, and a C-index of 0.805, a mean TD-AUC of 0.848, and an IBS of 0.106 on the external cohort, demonstrating that longitudinal ECG trajectories provide prognostic value beyond a single recording.
PaperID: 7508, Poster
Abstract: Reward design remains a central bottleneck in reinforcement learning (RL), particularly in sparse-reward, long-horizon, and partially observable settings requiring sequential dependency chaining, memory, and deceptive affordance disambiguation. Recent foundation-model approaches reduce manual reward engineering; however, most assume explicit goal-conditioning, privileged task descriptions, expert demonstrations, or access to ground-truth metrics. These assumptions are particularly problematic in partially observable environments, where task-relevant objectives, affordances, and dynamics must be discovered through interaction rather than disclosed a priori. We formalise this problem setting by introducing WATERFALL, an iterative population-based workflow for automated reward discovery operating under the strict observability constraints inherent in POMDP environments. WATERFALL does so without access to semantic task descriptors, pre-disclosed environment dynamics, goal conditioning, or ground-truth metrics. WATERFALL achieves this by (i) synthesising diverse programmatic reward candidates via persona-conditioned generation, (ii) evaluating candidate behaviour from raw visual rollouts utilising a Swiss-system evolutionary tournament workflow judged by vision-language models to identify elites, and (iii) leveraging longitudinal assessment histories to iteratively mutate, escalate, or prune candidates. We evaluate WATERFALL across fully observable and reveal-gated partially observable environments against sparse-reward RL, intrinsic-motivation baselines, as well as state-of-the-art foundation-model reward-synthesis baselines. Our results indicate that WATERFALL can synthesise, evaluate, and iteratively refine reward programs without access to privileged semantic information or ground-truth scalar supervision. Enforcing this strict information contract is especially important in the intended reveal-gated POMDP regime.
PaperID: 7509, Poster
Abstract: World Action Models have recently emerged as a promising paradigm for improving action generalization in autonomous driving by leveraging future video prediction as dense supervision for scene dynamics and temporal causality. However, it remains unclear how video modeling should be integrated with action generation to maximally transfer these benefits. Existing cascaded or dual-Diffusion Transformer architectures decouple video imagination from trajectory prediction, limiting knowledge transfer: the action model may still overfit dataset-specific driving priors, while the video model only indirectly regularizes planning. In this paper, we propose UNIVERSE, a unified video-action model built upon a single mask-modulated Diffusion Transformer. By co-training future video latents and ego-trajectory tokens within shared generative parameters, UNIVERSE allows dense video supervision to directly shape trajectory denoising, leading to stronger cross-domain action generalization. To ensure causal validity and efficient deployment, we introduce a Modality-Decoupling Visibility Mask, which shares historical context across modalities while blocking mutual attention between future video and trajectory tokens. This prevents future-modality dependency and enables trajectory-only inference by removing future-video denoising at test time, achieving a 4.3× speedup over joint video--trajectory rollout while maintaining comparable planning accuracy. The same model also supports video-only and joint video--trajectory rollouts. Experiments show that UNIVERSE achieves 91.0 PDMS on NAVSIM and strong zero-shot transfer to nuScenes and Bench2Drive without fine-tuning, while ablations confirm the importance of single-DiT unification, video co-training, and mask-based modality decoupling.
PaperID: 7510, Poster
Abstract: Event-based optical flow estimation achieves high temporal resolution and dynamic range by exploiting asynchronous events, while self-supervised learning provides an effective solution to the absence of ground-truth labels. Spiking neural networks (SNNs), as brain-inspired models, process information through asynchronous spikes, naturally aligning with optical flow events. However, the temporal irregularity across event streams hinders the compact extraction of spatiotemporal features, while the absence of explicit labels leads to spurious correlations, thereby increasing redundancy within latent representations. In this paper, we propose a novel “Compactly Compress, Time-varying Focus” training strategy. We present SICAF (\textSpikes \textInformation \textCompression \textAnd \textFocus) framework that formulates the learning process as a constrained optimization with temporal modulation. It encourages self-supervised SNNs to extract compact latent representations while preventing over-compression that would discard temporal saliency. To better enable feature selectivity across time steps, we design an estimator with learnable temporal attention, achieving context awareness with cross-time dependencies. Experimental results on several datasets demonstrate that SICAF achieves state-of-the-art performance in self-supervised SNNs, with a 8.14% improvement in AEE on the MVSEC dataset and a 13.49% increase in robustness under white-box adversarial attacks.
Abstract: Speculative Decoding (SD) has significantly accelerated Large Language Model (LLM) inference, yet existing approaches face a fundamental tradeoff between two drafting strategies: neural drafting and context-based copying. Neural drafts (e.g., EAGLE3) provide robust performance across diverse text settings, while copy-based methods achieve higher speedups in copy-intensive regimes by generating candidates faster and exploiting long repetition spans for near-perfect speculation. We analyze existing copy-based methods and find that they are prone to accidental repetitions where surface-level n-gram overlap does not reflect a structural intent to copy, leading to false-positive triggers that ultimately degrade throughput. We introduce SwitchSD, an adaptive framework that treats copying as a latent control signal of the LLM. By training lightweight probes on the target model's internal representations, SwitchSD identifies genuine copy-intent with high precision (AUC > 0.99). This allows the system to dynamically switch between neural drafting (e.g., EAGLE) and context-based copying. Our results across Llama and Qwen families demonstrate throughput gains of up to 15% over state-of-the-art baselines like EAGLE3, effectively turning "copying" from a noisy heuristic into a principled, model-aware decoding regime.
PaperID: 7512, Poster
Authors: Nathaniel Kang
Abstract: Predictive models on structured tabular data often face imbalanced targets: training emphasizes common values while rare outcomes stay underrepresented. Distributionally robust optimization hedges against shift, but standard ambiguity sets use a single radius over the entire distribution, so adversarial mass concentrates on dense regions and rare target ranges remain weakly protected; worst-case statements then track bulk behavior rather than tails. We introduce target-conditional distributionally robust optimization (TC-DRO), which defines ambiguity per target region with radii that grow where data are scarce. We establish coverage, per-region excess risk that tightens with local sample size when standard DRO yields vacuous tail guarantees, and a dual as weighted empirical risk with theory-derived density-adaptive weights. Experiments on seven tabular benchmarks show that Wasserstein DRO collapses to empirical risk minimization under uniform weights, while global \chi^2-DRO can remain close to ERM in practice; TC-DRO improves tail MSE and SERA relative to both standard training and recent imbalanced-regression methods.
Authors: Corey Adams, Rishikesh Ranade, Sheel Nidhan, Ram Cherukuri, Sanjay Choudhry, Mohammad Amin Nabian, Deepak Akhare
Abstract: We present GeoTransolver, a multiscale geometry-aware physics attention transformer for Computer Aided Engineering (CAE). GeoTransolver extends the Transolver backbone with GALE (Geometry-Aware Latent Embeddings) attention, which pairs physics-aware self-attention on learned state slices with cross-attention to a shared geometry and global context computed via multi-scale ball queries (inspired by Domino) and reused in every block. Implemented and released in NVIDIA PhysicsNeMo, GeoTransolver persistently projects geometry and global parameters, into physical state spaces to anchor computations to domain structure and operating regimes. We benchmark on DrivAerML, SHIFT-SUV, and SHIFT-Wing against Domino, Transolver (PhysicsNeMo implementation), and literature-reported AB-UPT, evaluating drag/lift R^2 and relative L_1 errors on field variables. As an additional nonlinear structural mechanics application, we also report Transolver and GeoTransolver results on bumper-beam and full-vehicle Body-in-White (BIW) crash-dynamics benchmarks, evaluating relative L_2 trajectory error and probe-level kinematic MSE. GeoTransolver delivers improved accuracy, robustness to geometry and regime shifts, and favorable data efficiency; we include DrivAerML ablations and qualitative contour and design-trend results, advancing operator learning for high-fidelity surrogates on complex, irregular, non-linear domains.
PaperID: 7514, Poster
Authors: Ramakrishnan Krishnamurthy
Abstract: More data is commonly assumed to improve learning performance. We show that this intuition can fail in sequential decision-making. Motivated by multi-agent collaborative learning, we study stochastic bandits with exogenous side information, and ask if free and unbiased reward samples are guaranteed to reduce regret. We introduce a formal notion of non-degradation under side information and demonstrate that the widely used UCB algorithm violates this property: there exist benign side-information designs under which UCB incurs strictly higher regret than when run without side information. We contrast this behavior with Thompson Sampling and Explore-Then-Commit, which benefit from the same information. Constructively, we further propose a meta-algorithm that enforces non-degradation for a broad class of bandit algorithms. Finally, experiments in recommender systems scenarios show that data sharing among different groups of users under UCB can yield diminishing and even negative returns for individual groups.
PaperID: 7515, Poster
Abstract: This work proposes the Adaptive Federated Hamiltonian Monte Carlo (AFHMC) for Bayesian federated learning (FL) tasks. The prior work on HMC applications for FL achieves O(\sqrtd/\epsilon) communication cost under log-convex distributions, already demonstrating its advantage over the Langevin dynamic-based counterpart. Based on a refined analysis, this paper further reveals the benefit of HMC for the FL regime by incorporating a control mechanism for local node shifts. We prove that AFHMC at most requires O((d/\epsilon^2)^1/3) communication cost up to a logarithmic term, which is significantly better than the existing results. Moreover, the proposed method adapts to client heterogeneity. Under low-heterogeneity settings (i.e., the target sampling precision \epsilon^2 is higher than the heterogeneity level), the communication rate of AFHMC can be as good as O(\log(d/\epsilon^2)), resembling the communication costs for FL optimization tools. To the best of our knowledge, this provides the first logarithmic communication complexity result for federated sampling. Our code is available at https://anonymous.4open.science/r/Adaptive-Federated-HMC.
PaperID: 7516, Poster
Abstract: Multimodal Large Language Models (MLLMs) have achieved impressive performance across a wide range of reasoning tasks, yet it remains unclear whether their internal representations and decision criteria align with the latent structure of human cognition. In this work, we propose a novel framework to align MLLMs with the multidimensional mental representations underlying human similarity judgements. We build on a previously established behavior-derived semantic embedding space, in which objects are represented along 66 interpretable dimensions derived from 4.70 million human Odd-One-Out (O1O) judgements. For each triplet, we infer the most salient decision dimension and convert it into a dimension-grounded linguistic rationale. This enables MLLMs to learn both the final choice and a justification consistent with the inferred behavioral criterion. Experimental results demonstrate that our approach improves model consistency with human similarity judgements while retaining competitive performance on standard multimodal benchmarks. Furthermore, using searchlight Representational Similarity Analysis (RSA) on an independent fMRI Natural Scenes Dataset (NSD), we observe increased representational alignment between model and human. Ultimately, our findings point to a data-driven route for incorporating human cognitive structure into MLLMs alignment.
Authors: Ningkang Peng, Xuanming Chen, Yanhui Gu
Abstract: Long-tailed out-of-distribution (LT-OOD) detection is often addressed with specialized training, including auxiliary out-of-distribution (OOD) data, abstention heads, contrastive objectives, energy losses, or gradient-conflict control. We show that these training mechanisms can obscure a simpler issue: frozen long-tailed representations may already contain useful OOD evidence, but raw Mahalanobis distance is distorted by frequency-coupled feature radius and poorly supported tail covariance. We propose \emphHyperspherical Pooled Mahalanobis (\HPM), a post-hoc detector that normalizes features onto the unit sphere and replaces class-specific covariance with a pooled, ridge-regularized metric while keeping class means as semantic anchors. In CIFAR-LT experiments and an ImageNet-100-LT near-OOD boundary analysis, \HPM improves raw Mahalanobis scoring; for Prior-Calibrated ERM (PC-ERM), it raises AUROC from 46.49 to 85.67 on CIFAR-10-LT and from 50.40 to 78.35 on CIFAR-100-LT. This simple PC-ERM+\HPM pipeline also achieves the best Log Efficiency Score (LES; 3.08) on CIFAR-100-LT, retaining roughly 95% of the best CIFAR-100-LT AUROC observed among the compared post-hoc scores at substantially lower training-time cost. These results argue for evaluating representation quality, detector geometry, and training complexity as separate factors in LT-OOD detection.
PaperID: 7518, Poster
Abstract: Sparse object detection remains difficult when weak target evidence occupies only a few pixels and shares local statistics with surrounding clutter. This failure is not merely due to limited model capacity, but to a representation bottleneck: spatial domain features mix target edges, clutter textures, and scene context before they can be reliably separated. We propose PRISM, a wavelet domain detector that treats this problem as frequency separated representation learning followed by priority ordered feature interaction. PRISM decomposes each scene into directional high frequency evidence and global low frequency context, models the two streams with dedicated encoders, and fuses them through a learned content adaptive causal scan that lets likely target regions guide feature interaction before clutter dominated regions participate. Multi scale features are then reconstructed by inverse wavelet synthesis with adaptive gates and passed to a standard detection head. On low altitude UAV maritime benchmarks, including SeaDronesSee and AFO, PRISM achieves state-of-the-art performance with the largest gains on the smallest targets. Without architectural modification, it also transfers to aerial urban imagery on VisDrone, suggesting that priority ordered fusion of directional frequency evidence is a generalizable design principle for sparse visual signal detection.
PaperID: 7519, Poster
Authors:
Íñigo Zubeldia, Boris Bolliet, Francisco Villaescusa, Pablo Villanueva-Domingo, Miles CranmerAbstract: LLM-based peer review is now used at production scale, but rigorous evaluation remains limited. We introduce a benchmark for evaluating AI peer reviewers’ detection of injected errors and quantifying false-positive rates across error type, severity, prompting strategy, and reviewer architecture. We also present skepthical, a multi-agent AI reviewer that verifies citations by retrieving and reading cited papers, checks mathematics with computer algebra, audits numerical claims with code, and uses a cross-model ensemble for general scientific issues. On 15 astrophysics preprints containing 177 injected errors, we benchmark 12 reviewer systems: skepthical; six single-LLM reviewers from three providers, each run with free-form and structured prompts; two Claude Code reviewers using the same Claude model and tools, one with a free-form prompt and one with a structured prompt; and three external reviewers. Detection performance varies sharply by model, prompt, tool use, and error type. Under the same free-form prompt, three frontier LLMs differ by 48.6 pp in total detection rate; a structured prompt halves this spread by improving the weaker models. Within Claude Code, changing only the prompt from free-form to structured raises detection by 32.2 pp. skepthical achieves the highest overall detection rate, 70.6±6.2%, with the best performance on general issues, 82.2%, and citation errors, 46.7%. On numerical errors, it matches both GPT-5.5 configurations at 80.0%, behind structured Claude Code at 91.1%; on mathematics, its 73.8% rate overlaps with both GPT-5.5 configurations within confidence intervals. False-positive rates also vary substantially: free-form Opus 4.7 has the lowest overall rate at 1.0%, while skepthical is the only reviewer with no false-positive failure mode jointly across mathematics, numerics, and citations, yielding a 2.8% overall rate. skepthical is deployed publicly at skepthical-ai.org, and we release the evaluation set as a benchmark.
PaperID: 7520, Poster
Abstract: On-device deployment of large language models (LLMs) is increasingly adopted due to latency and privacy requirements, but exposes proprietary models to reverse engineering and intellectual property theft. Trusted Execution Environments (TEEs) provide hardware-enforced isolation for on-device inference, yet most TEE-based parameter obfuscation relies on unilateral matrix transformations that obfuscate only output side. We identify a vulnerability in this design, termed the Solution Determinacy Property (SDP): the obfuscation secret is uniquely determined, exposing the structure of the true key and enabling attackers to break the protection. Under our attacks, representative methods GroupCover and ArrowCloak are reduced to negligible protection, with attack accuracy reaching over 2.4× the full-shielding black-box baseline and recovering near-complete model functionality. To address this vulnerability, we propose BiShield-TEE, a bilateral framework that induces Ambiguous Solutions through joint obfuscation of input and output dimensions, hiding the true key among many equally valid candidates. Experiments across diverse models and datasets show that BiShield-TEE confines attack accuracy to the full-shielding black-box baseline, while achieving higher inference efficiency than prior methods on real hardware. Our code is available at https://anonymous.4open.science/r/BiShield-TEE-28DF.
PaperID: 7521, Poster
Abstract: Existing approaches to optimization modeling using large language models (LLMs) treat tasks through from-scratch construction, generating variables, objectives, and constraints without leveraging reusable structural knowledge. However, in real-world scenarios, expert modelers typically begin by identifying the canonical type or core structure of a problem, followed by iterative refinements based on specific task requirements. To incorporate this expert incremental logic, we propose a novel two-stage framework. In the first stage, we conduct Logic-Anchored Data Synthesis starting from 17 canonical prototypes, preserving critical bottleneck constraints and generating synthetic problem–formulation pairs. These pairs, together with the OR-Instruct dataset, are used for supervised fine-tuning (SFT) to initialize a policy. Separately, for each generated problem, we produce multiple modeling trajectories that follow a prototype-grounded order with variables preceding expressions. These trajectories are ranked according to criteria of accuracy and efficiency, and subsequently used to train a Logical Reward Model (LRM). In the second stage, we introduce Logic-Test-Time Group Relative Policy Optimization (Logic-TGRPO), which leverages the LRM during test-time reinforcement learning to reward accurate prototype identification and disciplined structural adherence while penalizing illogical patterns. Evaluated across seven benchmarks, our 8B parameter model achieves an average accuracy of 79.9%, outperforming comparably scaled baselines and competing effectively with heavyweight multi-agent systems with far fewer parameters. Strong performance on out-of-distribution tasks confirms that the model exhibits genuine incremental reasoning rather than mere memorization. Ablation studies further validate that both stages in our framework are essential.
PaperID: 7522, Poster
Authors: Christophe Thomassin, Marçal Comajoan Cara, Margarita Geleta, David Bonet, Benet O Sabat, Daniel Mas Montserrat, Alexander Ioannidis
Abstract: Precision medicine relies on integrating diverse patient data, including genetic variants, clinical phenotypes, and medical imaging, to inform personalized prevention, diagnosis, and treatment. However, learning general patient-level representations from these heterogeneous modalities remains challenging due to their high dimensionality and pervasive missingness. We propose PM1, a multimodal foundation model that learns meaningful patient-level representations across genetic, clinical, and imaging data from 337,129 individuals in the UK Biobank, with a focus on retinal imaging and ophthalmic traits. To address modality dimensionality, PM1 adopts an intermediate fusion architecture with modality-specific encoders and a shared Transformer backbone, where a novel pathway-based genotype encoder enables the incorporation of over 600,000 genetic variants by reducing dimensionality while preserving trait-relevant variation. In addition, a contrastive loss invariant to modality missingness allows PM1 to learn robust patient-level representations despite only 6% of UK Biobank participants having complete modality coverage. By partitioning modalities into interpretable tokens, PM1 enables interpretability methods to investigate known and potentially novel associations across genetic, clinical, and imaging features, and their interactions. Evaluated across downstream tasks, including phenotype prediction, conditioned image generation, and attribution analyses, PM1 demonstrates the value of multimodal patient representations, improving predictive performance over unimodal baselines and competing methods while enabling interpretable analyses in settings with heterogeneous and incomplete patient data.
Abstract: Language models (LMs) raise an alluring alternative to vector-based retrieval: \emphgenerating a relevant answer, entirely based on an in-context corpus. In spite of this appeal, prior studies consider proprietary systems or the smaller-scale reranking task, leaving corpus-scale in-context retrieval largely unexplored. In this work, we present the first systematic study of in-context retrieval on two scales practical retrievers demand: \emphmillion-token corpora and \emphlength-generalization far beyond training-time sizes. We first introduce BlockSearch, a 0.6B LCLM retriever whose architectural and training modifications improve over prior LM baselines and length-generalize up to 10× beyond its training regime. However, its retrieval still collapses under more extreme extrapolation. We trace this failure to an \emphattention dilution effect: as the corpus grows, irrelevant documents dominate the softmax denominator, minimizing the normalized mass on the gold document even as its pre-softmax score stays high. Motivated by this analysis, we introduce \emphlength-aware adjustments to the attention softmax and \emphdocument-level sparse attention, improving retrieval at million-token scale, better performance than the concurrent 7×-larger MSA system, and comparable performance to dense retrieval. Together, our results position in-context retrieval as a viable alternative to classical retrieval, emphasizing attention control under extreme context growth as a new challenge.
PaperID: 7524, Poster
Authors: Wang Xi, Shijia Xu
Abstract: Autoregressive models are often used as sequential decision makers by asking them to rewrite the full intermediate state before producing the next action. We study a failure mode of this interface: exact high-entropy state copying can interfere with goal-directed computation. We formulate this as an operational copy--reason interference hypothesis rather than as a directly measurable capacity law. Matched controls separate copying from length: on controlled state-tracking tasks, Full-State Generation degrades much more sharply than an Iso-Length Control and similarly to a Complex-Copy Control. Activation patching on Llama-3.1 8B and 70B recovers target-action evidence when clean residual activations are inserted into late layers of full-state runs, suggesting that goal features remain recoverable but are poorly routed under the copy-heavy interface. Delta-style alternatives reduce this burden. At the 2,000-token Formal Math anchor, Scratchpad Residual improves accuracy from 41.2--68.5% under Full-State Generation to 97.1--99.4%; on SWE-bench Lite, it improves pass rates from 12.4--26.8% to 18.5--34.8%. The gain has a boundary: probe and task accuracy both decay with long dependency distance. These results argue for separating state maintenance from action generation in long-horizon autoregressive systems.
Abstract: 3D object grounding localizes referred objects in a 3D scene from natural language. Unified instance-centric 3D-LLMs aim to solve grounding together with dialog, QA, and captioning, yet many rely on a single pointer-style grounding decision that compresses a relational instruction into one selection. This is brittle for fine-grained queries where multiple same-class candidates must be ruled out by context objects and spatial relations. We propose Structured Spatial Reasoning 3D-LLM (SSR3D-LLM), a structured grounding interface for unified 3D-LLMs. Given fixed Mask3D object proposals, the LLM writes a sequence of latent spatial reasoning steps and memory tokens from the query, and a geometry-aware scorer reads these latent steps in order to refine candidate rankings step by step with step-length masking. The latent steps are learned from standard benchmark target supervision with auxiliary referential-cue supervision during training, while inference uses only the input query and Mask3D proposals. Across ReferIt3D, ScanRefer, and Multi3DRef, SSR3D-LLM achieves the strongest results among unified 3D-LLM baselines, with substantial gains over the single-pointer QPG baseline on fine-grained grounding and consistent improvements over prior unified 3D-LLMs, while preserving the default language-task route.
PaperID: 7526, Poster
Abstract: In-context learning enables large language models to adapt to tasks directly from input sequences, without parameter updates. We investigate the mechanism underlying in-context learning of autoregressive processes in transformers under heterogeneous second-order moments of prompt distributions. Our study considers a two-layer architecture consisting of a linear attention head and a nonlinear MLP, with the outer layer trained via gradient descent. We prove that there exists a parametrization under which the model converges to the unique function that approximates the optimal second-order moment-based AR(1) estimator. We analyze bounds on approximation and convergence rates and confirm our findings experimentally. These results extend the first-moment-centric understanding of estimators realized by in-context learning to second-order moment-based estimators, offering further insight into ICL's robustness to new tasks.
PaperID: 7527, Poster
Abstract: We study dynamic pricing where a seller repeatedly interacts with a strategic, non-myopic buyer who has a fixed private valuation and discounts future utility. Prior work focused exclusively on posted-price mechanisms, where the seller gives a take-it-or-leave-it offer. For our first result, we show that menu mechanisms consisting of allocation-payment achieve O(T_\gamma \log T_\gamma) regret, where T_\gamma is the buyer's effective discounted time horizon. We also establish a \Omega(T_\gamma) lower bound, demonstrating the bound is tight up to \log factors. Considering the geometric discounting buyer with a constant discount factor, our bound is O(1), while prior bounds using posted-price mechanisms incur an unavoidable \Omega(\log\log T) factor in regret. Our second contribution is more conceptual in nature. The problem of dynamic pricing sits at the intersection of two paradigms: learning with strategic agents in computer science / machine learning and revelation-principle-based mechanism design in economics, yet their relationship has remained unclear. We establish a fundamental equivalence: indirect learning-based mechanisms and direct revelation mechanisms achieve identical optimal regret. The adaptive, data-driven algorithms of online learning and explicit type elicitation are two languages towards solving the same problem.
PaperID: 7528, Poster
Authors: zhenglong chen
Abstract: We develop a Koopman-operator analysis of self-attention that resolves a long-standing puzzle: why does linear attention diverge on long sequences while softmax attention remains stable, even though both share the same next-step prediction objective? We show that attention stability is intrinsically a two-layer problem. The architectural layer concerns the spectrum of the n × n token-mixing matrix and is data-independent; the dynamical layer concerns the spectrum of the d × d dual operator induced by training and should be data-determined. We prove three concrete results within this framework: (i) under a shared encoder, the dual operator C^ = W_Q W_K^\top G W_V of linear attention coincides with the EDMD Koopman estimator K_\textEDMD = G^-1A at the global optimum, and its spectral radius is unconstrained; (ii) causal softmax attention enjoys an architectural guarantee \rho(A^\textsoft) \le 1 via row-stochasticity, while its value projection W_V leaves the dynamical layer entirely unconstrained; (iii) LayerNorm makes the empirical Gram matrix G_\textLN rank-deficient with kernel \textspan(\mathbf1), producing a d-parameter family of loss-equivalent but spectrally distinct EDMD solutions, while RMSNorm preserves invertibility. The framework yields an immediate design principle—constrain the architectural layer, free the dynamical layer—which we instantiate as two stabilized linear-attention variants (SN-LA, RN-LA). Numerical experiments reveal a hierarchy of architectural stability: vanilla linear attention has neither bound and diverges immediately under autoregressive rollout; our stabilized variants enforce only the single-pass spectral bound \rho(\tildeA) \le 1 and delay divergence; softmax attains the strictly stronger pointwise contraction \Vert A v\Vert_\infty \le \Vert v\Vert_\infty via row-stochasticity and remains bounded. On real time-series benchmarks (ETT, Weather), our methods close \ge 80% of the vanilla-to-softmax gap and surpass softmax on 3/5 datasets at long horizons, while preserving the dynamical-layer freedom that constraints on the dual operator forfeit.
PaperID: 7529, Poster
Abstract: Kernel methods, and Gaussian Processes (GPs) in particular, require a conditionally negative-definite (CND) distance measure to guarantee positive semi-definiteness (PSD) of the kernel matrix — a condition that fails for many natural input spaces, including smooth manifolds and spaces of probability distributions. We propose the Sparse Landmark Embedding (SLE) kernel, which eliminates this requirement entirely. Each input is embedded into a sparse feature vector via compactly supported bump functions centered at all training points |\mathcalD|; applying any standard PSD kernel in this embedding space yields a kernel that is provably PSD for arbitrary distance measures. The compact support automatically controls embedding sparsity, keeping kernel matrices well-conditioned and computationally tractable despite the high ambient dimension. We provide theoretical guarantees on PSD, sparsity, stability, and universal approximation, and demonstrate, using geodesic and Wasserstein distances, that the SLE kernel matches or substantially exceeds domain-specific baselines in both predictive accuracy and uncertainty quantification.
PaperID: 7530, Poster
Authors: Yujian Ye, Siqi Qian, Yizhi Wu, Tianxiang Cui, Goran Strbac
Abstract: Deep decarbonization is transforming power grids into safety-critical, low-carbon cyber-physical systems, where AI controllers must jointly manage economic efficiency, carbon reduction, and operational security. This operational energy trilemma is challenging for reinforcement learning: cost and carbon objectives are naturally optimized in expectation, whereas security violations are rare, heavy-tailed events whose consequences can be catastrophic and therefore require explicit tail-risk control. Existing safe reinforcement learning methods either constrain safety only in expectation or rely on light-tailed approximations, which can underestimate rare but severe grid violations. We propose a heterogeneous risk-constrained reinforcement learning framework for low-carbon grid operation. The key idea is to assign different risk semantics to different objectives: safety-critical security constraints are enforced through Conditional Value-at-Risk (CVaR), while economic and carbon-related objectives are optimized in expectation to preserve operational flexibility. To better capture extreme scenarios, a mixture distribution model is introduced to characterize heavy-tailed constraint violations. We further develop a residual policy learning scheme built on a distributionally robust chance-constrained reference module: the robust module provides a feasible and economically efficient baseline policy, and the learned residual refines it toward lower-carbon operation while respecting tail-risk bounds. Experiments on power-system operation tasks show that the proposed framework reduces economic cost by over 25% compared with expectation-based CMDP baselines and improves safety by two orders of magnitude over Gaussian-tail methods. These results suggest that risk-aware reinforcement learning can support reliable decarbonization by reconciling efficiency, sustainability, and security within a unified decision-making framework.
PaperID: 7531, Poster
Authors:
Tommaso Castellani, Jiaqi Li, Muriah D Wheelock, Rezwana R Razzaque, Helmut Laufs, Hong Chen, Claire DonnatAbstract: Graph Neural Networks (GNNs) are an important tool for fMRI-based prediction, where empirical functional connectivity matrices serve both as node features and as the basis for graph construction. However, recent studies (and our own experiments) show that these methods often fail to outperform simpler graph-agnostic baselines. We propose a statistical explanation for this failure by formulating connectome prediction as an errors-in-variables (EIV) problem: the true connectome is latent, while the observed covariance or correlation matrix is a noisy finite-time-series estimate. In a linear asymptotic setting, we decompose the effect of graph convolution into three mechanisms: smoothing-induced bias and variance reduction, graph estimation error, and a double-dipping effect arising when the same noisy matrix determines both topology and features. Under a latent community model adapted to this neuroscience problem, we identify a narrow favorable regime in which message passing can improve prediction: the regression target must be community-smooth, and the time-series length must lie in a window between the graph-estimation and raw-noise limits. Simulations validate this decomposition by isolating oracle, independently estimated, and same-sample graphs, while real-data experiments illustrate that standard connectome GNN pipelines often fall outside this regime. Our results clarify when graph convolutions can help brain connectome prediction, and why they often do not.
Authors: Richard Dewey, Janos Botyanszki, Ciamac C Moallemi, Andrew Zheng
Abstract: AI researchers have long focused on poker-like games as a testbed for environments characterized by multi-player dynamics, imperfect information, and reasoning under uncertainty. While recent breakthroughs have matched elite human play at no-limit Texas hold'em, the multi-player dynamics are subdued: most hands converge quickly with only two players engaged through multiple rounds of bidding. In this paper, we present Solly, the first AI agent to achieve elite human play in reduced-format Liar’s Poker, a game characterized by extensive multi-player engagement. We trained Solly using self-play with Regularized Nash Dynamics (R-NaD), a model-free, actor-critic, deep reinforcement learning algorithm, and we successfully extend R-NaD to the multi-player scenario for the first time. Solly played at an elite human level as measured by win rate (won over 50% of hands) and equity (money won) in heads-up and multi-player Liar’s Poker. Solly also outperformed large language models (LLMs), including those with reasoning abilities, on the same metrics. Solly developed novel bidding strategies, randomized play effectively, and was not easily exploitable by world-class human players.
PaperID: 7533, Poster
Abstract: Generative object detection, introduced by DiffusionDet, is defined by two coupled properties: (P1) data-independent random-box initialization with a learned noise-to-data transport, and (P2) test-time scaling (TTS), the ability to trade inference compute for accuracy at a fixed checkpoint by adjusting the number of proposals or refinement steps. This paper asks: \emphHow can we preserve both properties with a minimal conceptual design? We dissect DiffusionDet component by component and find that, in our controlled interventions, the diffusion schedule, the initial noise distribution, and DDIM sampling stochasticity have only a small effect on inference results. In contrast, detection performance is more sensitive to the denoising head, iterative refinement, and proposal coverage. Based on these diagnostic results, we propose LinearDet, a compact generative detection prototype: \emphlinear noising at training, linear denoising at inference---no diffusion-specific schedule, no diffusion ODE/SDE solver, only a symmetric pair of linear mixtures. On COCO, LinearDet matches DiffusionDet under standard settings, is more robust under sparse proposal coverage, outperforms it in zero-shot CrowdHuman transfer, and produces smoother refinement trajectories that admit clean heatmap visualization. More broadly, we hope this diagnosis and compact prototype offer the object detection community a cleaner baseline for studying generative detection.
PaperID: 7534, Poster
Abstract: Recent works have proposed incorporating heavy-tailed (HT) noise into diffusion- and flow-based generative models, with the goals of better recovering the tails of target distributions and improving generative diversity. This motivation is intuitive: if the data are heavy-tailed, HT noise may appear better matched than light-tailed (LT) Gaussian noise. However, replacing Gaussian noise by HT noise also changes the underlying estimation problem. In this paper, we revisit this paradigm through a combined theoretical and empirical study, establishing sampling-error bounds for two representative diffusion models driven by HT and LT noise. We show that HT noise makes the statistical estimation problem harder, leading to less favorable sampling-error bounds. We support these findings with experiments on synthetic and real-world datasets, empirically recovering the predicted error trade-off. Our results call into question a growing design trend in generative modeling and challenge the use of HT noise to improve rare-region exploration.
PaperID: 7535, Poster
Authors:
Haoxiang Hu, Cangjun Gao, Zhengming Zhang, Yaxian Shan, Zeyuan Huang, Ran Zuo, Qiang He, Qingkun Li, Xiaoming Deng, Cuixia Ma, Yu-Kun Lai, Yong-jin Liu, Hongan WangAbstract: Sketch-Text image retrieval aims to find natural images using both free-hand sketches and text descriptions. Although these two modalities provide complementary information, the task is challenging due to their large representational differences and the difficulty of matching fine details in complex scenes. Most existing methods rely on simple feature fusion or global alignment, which often fails to preserve important discriminative cues and leads to inaccurate retrieval results. To tackle these issues, we propose ST-Bridge, a high-precision sketch-text image retrieval framework that enables reliable coarse-to-fine matching through an effective semantic bridging mechanism. Specifically, we introduce LLM as a semantic bridge to generate sketch representations that are closer to the textual semantic space. We further enhance these representations and the original sketch features via gated attention and feature normalization, substantially reducing the sketch--text modality gap. Building upon this, we design a coarse-to-fine alignment strategy that supports accurate sketch--text image retrieval. Extensive experiments on scene-level sketch--text image retrieval benchmarks demonstrate that ST-Bridge significantly outperforms existing methods and achieves state-of-the-art performance, validating the effectiveness of large-model-based semantic bridging for cross-modal retrieval tasks. Our code and models will be released upon acceptance.
PaperID: 7536, Poster
Authors: Dahun Choi, Inseong Hwang, Hyun Kim
Abstract: Vision–language models (VLMs) exhibit strong multimodal reasoning, yet their growing parameter counts incur prohibitive memory footprints and inference latency, hindering real-time deployment. Post-training quantization alleviates these costs, but aggressive low-bit quantization often degrades accuracy by disrupting visual–linguistic alignment. Meanwhile, lossless compression preserves accuracy but suffers from metadata overhead and sequential decoding, limiting GPU parallelism. We propose Comp²VLM, a hybrid framework that jointly designs quantization and lossless compression for efficient VLM serving. We introduce Duplication-Enhanced Quantization (DEQ), which adaptively selects group-wise scales to reduce the entropy of quantized representations while preserving reconstruction fidelity, thereby improving compressibility. To realize runtime gains, we present Quantized Data-aware Lossless Compression (QPress), a hardware-friendly scheme based on a universal Huffman table that enables block-wise metadata sharing and GPU-parallel decoding. Across diverse VLMs and benchmarks, Comp²VLM maintains competitive accuracy while reducing the effective precision to W3.5A3.2~4.9. On LLaVA-onevision-7B, it achieves +4.7% TextVQA and +7.7% OCRBench accuracy over W4A8-based SOTA at an average bit-width of W3.5A4.6. DEQ yields up to 33% entropy reduction, and QPress improves compression efficiency by up to 13.3% with only ~4% end-to-end latency overhead. To our knowledge, Comp²VLM is the first unified framework to co-design quantization and lossless compression for VLMs, demonstrating their complementary benefits beyond standalone methods.
PaperID: 7537, Poster
Authors: Luyao Liu, Wenlong Dong, Chengjie Zhang, Hong Zhang
Abstract: Generative modeling has become a mainstream paradigm for learning based robot manipulation research. Among these, diffusion-based policies achieve strong performance but typically require multiple iterative denoising steps to generate final action, leading to inefficient inference. Although flow-based methods eliminate the need for iterative denoising process, existing approaches often face challenges in maintaining stable optimization and high-quality action generation. In this paper, we propose SwiftFlow, an efficient one-step policy learning framework built upon improved mean flow for robotic manipulation. Specifically, SwiftFlow formulates policy learning as a conditional generation problem by jointly leveraging 3D point cloud observations and proprioceptive robot states. It stably learns mean velocity field through a principled regression objective and optimization process, enabling more reliable action generation while preserving efficient inference. To further regularize the learned implicit generation dynamics, we construct a mean velocity smoothness regularization term, which constrains rapid local variations in the learned mean velocity field. Since this regularization is imposed only during training, it hardly introduces additional inference overhead. Extensive experiments on Adroit and MetaWorld benchmarks demonstrate that our proposed method achieves stronger task performance together with highly efficient one-step inference, validating its effectiveness for robotic manipulation tasks.
Abstract: Recently, Group Relative Policy Optimization (GRPO) and its variants have been developed for policy optimization and demonstrated notable performance gains. However, these methods usually incur substantial computational overhead due to per-question multi-rollout sampling and repeated per-token probability evaluation across rollouts. Furthermore, low-information or highly homogeneous trajectories can degrade downstream learning signal efficiency, hindering model optimization and limiting final performance. To address these issues, we propose FastRL, a novel plug-and-play reinforcement learning framework that simultaneously improves training efficiency and the effectiveness of policy learning. Specifically, 1) We introduce an advantage-aware pruning strategy to selectively preserve high-advantage trajectories while maximizing inter-trajectory gradient diversity. 2) Then, we design an adaptive rollout sampling mechanism to dynamically adjust the sampling scale across different training stages based on historical pruning distributions, balancing exploration adequacy and computational efficiency. Experiments demonstrate that FastRL can be seamlessly integrated into GRPO, DAPO, and GSPO variants, achieving an average 2.07x training speedup on Geometry3K and GeoQA8K-R1V, along with an approximately 1.64% improvement in average accuracy on visual reasoning benchmarks. Source codes will be available at https://anonymous.4open.science/r/FastRL0.
PaperID: 7539, Poster
Abstract: Video anomaly detection (VAD) is critical for identifying safety risks in surveillance and autonomous systems. Recent large-model-based VAD methods generate textual rationales and directly output anomaly scores. However, mapping rich semantic descriptions and analysis directly to absolute scalar scores can be poorly grounded and inconsistent, leading to scores that are not well aligned with anomaly severity. With the observation that for humans, it is usually easier to compare things than to rate things directly, we explore whether it also holds for large models in VAD by proposing CA-Judge (Comparative Anomaly Judge), a training-free framework that replaces direct scoring with comparative order inference. CA-Judge treats model-perceived anomaly severity as a latent coordinate, constructs this coordinate from self-comparisons among generated normal/anomalous reference events, and models how anomaly likelihood varies along this coordinate using the generated normal/anomalous reference identities. At test time, CA-Judge avoids exhaustive comparison by maintaining a posterior belief over query severity and selecting a few query--reference comparisons by their expected information about the anomaly status. The resulting comparison trace is aggregated under a Bradley--Terry likelihood and read out through the learned severity--anomaly association to produce the anomaly probability. Without fine-tuning or training data, CA-Judge substantially improves over direct-scoring baselines and surpasses prior training-free state-of-the-art methods on UCF-Crime and XD-Violence.
PaperID: 7540, Poster
Abstract: Fairness-accuracy trade-offs are a central concern in the deployment of fairness-aware machine learning methods. When sensitive attributes are unavailable at inference time–the so called unawareness setting, principled methods for obtaining accurate predictions under relaxed fairness constraints are largely missing. In this work, we address this gap by formulating regression under a demographic parity penalty as an optimal transport problem. Our framework unifies both the \emphaware and \emphunaware settings and characterizes optimal prediction functions via optimal transport maps, under both squared Wasserstein-2 and Total Variation penalties. These results reveal that the choice of penalty reflects fundamentally different fairness philosophies: the Wasserstein penalty induces a smooth, population-wide compromise, while Total Variation enforces exact parity for a subset of individuals. Building on these theoretical characterizations, we propose an algorithm that is simple to implement, computationally efficient, and consistently matches or outperforms state-of-the-art baselines on real-world benchmarks.
PaperID: 7541, Poster
Abstract: Modern classifiers remain sensitive to small input perturbations, while strong model-agnostic robustness certificates such as randomized smoothing are often too expensive for deployment, requiring an extensive number (e.g. 2000) of model forward passes per input. We propose a framework that delegates certification to a cheap proxy function while statistically controlling its deviation from the original sound certificate. Given an exchangeable unlabeled calibration set, we tune the proxy via robust conformal risk control to guarantee a bounded marginal false-approval rate: the probability that the proxy accepts an input whose prediction is flipped due to perturbation. The resulting proxy certificate can be used either as a standalone marginal certificate or as a fallback mechanism that invokes the expensive certificate only when the proxy rejects. We provide both single-sample and sample-heavy calibration recipes. Empirically, our method substantially reduces test-time certification cost (to a single model forward) while retaining much of the certified performance of randomized smoothing, enabling cheap yet effective robustness guarantees for larger architectures such as vision transformers and language models.
PaperID: 7542, Poster
Authors:
Kaibo Liao, Yazhou Ren, Song Wu, Kecheng Chen, Xiaorong Pu, Lixin DuanAbstract: Semi-supervised medical image segmentation has gained increasing attention for its outstanding performance with limited annotations. Most methods follow the Mean Teacher framework, where Exponential Moving Average (EMA) is primarily used to stabilize predictions, while its potential roles beyond stabilization are largely overlooked. In this work, we interpret EMA as a first-order infinite impulse response (IIR) low-pass filter with inherent frequency selectivity, which enables the preservation of semantic information at different levels. Motivated by this insight, we propose a dual-EMA teachers framework, including a long-term teacher focusing on global structures and a short-term teacher emphasizing local details. Furthermore, to help the student better exploit the guidance from both teachers, a First-In-First-Out replay queue is designed to reuse historical high-confidence pseudo-labeled pairs, leveraging pseudo labels provided at different training stages and reducing the sensitivity to sampling bias. Meanwhile, we impose feature-level global semantic and local relation consistency constraints to alleviate the impact of noisy pseudo labels, facilitating robust transfer of complementary representations from dual teachers to the student model. Experiments on three public medical image datasets demonstrate that our method outperforms current SoTA methods, verifying its effectiveness.
PaperID: 7543, Poster
Abstract: Structured pruning offers hardware-agnostic acceleration for Large Language Models (LLMs) but typically degrades performance by strictly adhering to the original static sequential execution graph. In this paper, we challenge this rigidity with LoopPrune, an evolutionary framework that transforms model compression into an architectural re-discovery process. By treating pre-trained layers as a library of reusable primitives, LoopPrune optimizes a dynamic execution path where modules can be skipped, executed once, or iterated via controlled loops. This approach effectively decouples effective model depth from parameter count, allowing compressed models to maintain high reasoning capacity. To navigate the vast combinatorial search space, we introduce a hybrid evolutionary strategy enhanced by a Multi-Tier Population Initialization mechanism, which stratifies the search start-point using cross-model elite transfer, structural priors, and Wanda score importance estimation. Extensive experiments on diverse LLM families demonstrate that LoopPrune sets a new state-of-the-art across all sparsity levels. Notably, LoopPrune identifies efficient iterative patterns that allow compressed models to match or even surpass the zero-shot reasoning accuracy of their dense counterparts on specific benchmarks, while maintaining near-lossless perplexity at high retention rates.
PaperID: 7544, Poster
Abstract: Feature absorption is a structural failure mode of sparse autoencoders, where a single dictionary atom captures multiple distinct concepts rather than one, reducing interpretability. While prior work has shown that absorption configurations are spurious local minima of the sparse dictionary learning loss, little is known about the geometric mechanisms that determine whether a coherence penalty can escape them. We study the curvature geometry of pairwise coherence penalties at absorption configurations on the unit sphere. We show that (1) every coherence penalty in the power family \mathcalRp = \suml 0, yet the efficiency of this force depends critically on the penalty's curvature; (2) a phase transition at p = 2 governs this efficiency: the Riemannian Hessian in the splitting direction equals p(p-2) \cdot k^1-p/2 in closed form, so concave penalties (p < 2) receive curvature assistance and convert absorption into a strict saddle point, while convex penalties (p > 2) face curvature resistance and deepen the local minimum; (3) the boundary at p = 2 corresponds precisely to a Parseval invariance that renders the quadratic penalty geometrically neutral; and (4) under feature correlation \mu^, convex penalties reverse above \mu^_crit = (p-2)/(p+k-2), giving a closed-form threshold for when resistance can be overcome. Experiments on Pythia-160M and Gemma-2-2B confirm the predicted \lambda-efficiency ordering for competitive-activation SAEs, with absorption reduction verified as feature-level improvement rather than feature suppression. More broadly, our results point toward a principled framework for selecting coherence penalties based on the curvature geometry of the loss landscape rather than empirical tuning.
PaperID: 7545, Poster
Abstract: End-to-end speech LLMs often answer the same question less accurately from speech than from text, even when the spoken content is semantically matched to the text prompt. We study this modality gap as a layer-wise inference problem rather than a static embedding mismatch. Across four open-weight speech LLMs on SpeechMMLU and VoiceBench BBH, cross-layer Centered Kernel Alignment (CKA) with speech–text token alignment shows that speech can become geometrically text-like in middle layers while still failing to form stable late-layer answer margins. The central finding is that mid-layer alignment does not imply functional equivalence. ASR-to-LLM controls show that recognition explains part of the gap in some settings but not all of it. Matched-instance margin analysis localizes the failure to weak late-layer answer separation. Attention diagnostics further show that speech evidence remains more diffuse at the decision token. Speech-faithful temporal compaction improves BBH accuracy from 59.9% to 62.1% with tempo 1.2× and to 61.9% with silence trimming, while recovering 34.3% and 38.2% of text-correct/speech-wrong failures. These conditional repairs reveal a 2.35–2.62× asymmetry between temporal compaction and quantity-only KV merging on the same gap subset, motivating post-projection temporal abstraction rather than only input-level feature matching.
Authors: Hans Gundlach, Alex Fogelson, Jayson Lynch, Ana Trisovic, Jonathan S Rosenfeld, Anmol Rattan Singh Sandhu, Neil Thompson
Abstract: Algorithms have been estimated to increase AI pretraining FLOP efficiency by a factor of 22,000 between 2012 and 2023 (Ho et al., 2024) . Running ablation experiments on key innovations from this time period, we are able to account for less than 10× of these gains at small compute scales. Surveying the broader literature, we estimate that additional innovations not included in our ablations also account for less than 10×, yielding less than 100× efficiency gains total in low-compute regimes. This leads us to conduct scaling experiments, which reveal that much of this efficiency gap can be explained by algorithms with scale-dependent efficiency improvements. In particular, we conduct scaling experiments between LSTMs and Transformers, finding exponent differences in their compute-optimal scaling law while finding little scaling difference for many other innovations. These experiments demonstrate that -- contrary to standard assumptions -- an algorithm's efficiency gains are tied to compute scale. Using experimental extrapolation and literature estimates, we account for 3,360× efficiency gains over the same time period, with scale-dependent innovations accounting for close to 90% of log-efficiency gains by 2023. Our results indicate that algorithmic progress for small models has been far slower than previously assumed, and that measures of algorithmic efficiency are strongly reference-dependent.
Abstract: Unified multimodal models (UMMs) with interleaved reasoning, which generate both textual and visual steps into intermediate reasoning traces, have demonstrated great potential for visual mathematical reasoning tasks. However, we identify a key insight in this paradigm: generating intermediate visual reasoning steps is not always beneficial and can even be harmful, as self-generated visual steps may introduce erroneous visual evidence that misleads subsequent reasoning. Moreover, frequently triggering visual steps during reasoning incurs substantial computational and memory overhead, degrading inference efficiency. To address these accuracy and efficiency challenges, we observe that the model's internal signals can indicate whether a visual step will benefit reasoning before the entire visual generation is completed. Specifically, this work identifies two internal signals: i) Generation Intent, which reflects whether the model has a concrete textual plan for what to draw, and ii) Visual Fidelity, which measures whether the visual generation remains grounded in the original input image. Leveraging these internal signals, we propose AdaViG, a training-free adaptive visual gating method for unified multimodal reasoning. AdaViG dynamically evaluates each triggered visual step at an early visual generation stage and aborts it when both signals are weak, thereby preventing misleading visual evidence from entering the reasoning trace while avoiding unnecessary computation. Comprehensive experiments demonstrate that AdaViG improves accuracy by up to 5.7% while reducing visual generation FLOPs by 25.0%-91.0% and wall-clock latency by 15.4%-45.6%.
Abstract: Persistence diagrams are common representations in topological data analysis, yet they lack both the vector space structure required by machine learning algorithms and a principled framework for statistical comparison. We introduce STRAND (Survival Topological Representation ANalysis of Diagrams), which treats (collections of) PDs as survival data: each topological feature with persistence value p = d - b is a fully observed time-to-event, and the persistence survival function S(t) = \mathbbP(p > t) is the central object for comparing diagrams. From this single representation we derive (i) a non-parametric two-sample test with calibrated Type~I error and high power from a small number of diagrams; (ii) interpretable effect sizes; and (iii) a 1-Wasserstein-stable feature vector for downstream machine learning. We validate calibration and power on synthetic manifolds with controlled topology, demonstrate competitive vectorisation across 14 graph and 3D point cloud benchmarks, and apply the method to study functional brain connectivity in fMRI/neuroscience data. To our knowledge, STRAND is the first method to provide hypothesis testing and vectorisation for persistence diagrams from a single coherent and interpretable representation.
Abstract: The Maximum Mean Discrepancy (MMD) is a cornerstone statistic for nonparametric two-sample testing, but its test power is dictated entirely by the chosen kernel. Because any fixed kernel inherently fails to distinguish certain distributions, the kernel must be dynamically optimised. However, data-driven optimisation violates the foundational i.i.d. assumption, forcing a strict trade-off in existing frameworks. Ratio criteria ignore this dependence, inducing overfitting and variance collapse on rich kernel classes. Conversely, aggregation methods bypass the dependence using finite grids, but this strategy cannot scale to continuous search spaces like deep kernels. To break this dichotomy, we establish data-driven kernel selection as a model selection problem. We propose Complexity-Penalised MMD (CP-MMD), a criterion derived from a uniform concentration inequality (extending Maurer's framework to the two-sample setting) that bounds the post-optimisation distribution by penalising the empirical MMD with the complexity of the kernel search space. Because this penalty mathematically absorbs the cost of optimisation, CP-MMD enables direct, grid-free maximisation over continuous parametric classes, including scalar bandwidths, polynomial coefficients, and deep network parameters. By formally accounting for optimisation complexity, we theoretically guarantee that CP-MMD maximises true test power while ensuring unconditional Type-I validity. Consequently, CP-MMD enables grid-free kernel selection across linear, polynomial, and deep regimes, matching or exceeding state-of-the-art test power.
Abstract: Causal representation learning (CRL) seeks to recover latent variables with identifiability guarantees, typically up to permutation and component-wise reparameterization under appropriate assumptions. However, identifiability does not imply interpretability: latent semantics are typically assigned post hoc by alignment with known ground-truth factors. This limitation is particularly acute in scientific time series, where underlying mechanisms are unknown and discovering interpretable structure is a primary goal. In contrast, scientific observations (such as residue-pair distances, climate indices, or process sensors) are inherently semantic, as they correspond to named physical quantities. This raises a key question: can the interpretability of observations be transferred to the identifiable latent space? We propose MOSAIC (Module discovery via Sparse Additive Identifiable Causal learning), a sparse temporal VAE that integrates temporal CRL identifiability with support recovery over observed variables. MOSAIC identifies latent variables via regime-conditioned temporal variation, and recovers for each latent a sparse set of associated observations through an additive decoder, yielding module-level interpretability. We show that ANOVA main-effect supports are identifiable under general smooth mixing functions, and provide finite-sample recovery guarantees for a tractable sparse-additive variant. Empirically, MOSAIC recovers domain-consistent variable groups across RNA molecular dynamics, solar wind, ENSO climate, the Tennessee Eastman process, and a synthetic tokamak benchmark, enabling interpretable discovery of latent mechanisms in scientific time series.
Abstract: Reinforcement learning (RL) has become a prevalent paradigm for training tool calling agents, which typically requires online interactive environments. Existing approaches either rely on training data with ground truth annotations or require advanced proprietary language models (LMs) to synthesize environments that keep fixed once created. In this work, we propose TRUSTEE, a cost-friendly method for training tool calling agents with dynamic environments fully simulated by free open-source LMs that can be as small as 8B, including task generation, user simulation, tool simulation and trajectory evaluation, paired with an adaptive curriculum learning mechanism that controls task difficulty during training. Our empirical results show that TRUSTEE outperforms baselines which require extra external resources in most cases. These confirm that, with a sufficiently sophisticated design, even simulated environments with a local 8B LM as the backbone could set a strong baseline for tool learning. We hope our proposed paradigm could democratize tool learning and inspire future research on environment scaling with limited resources.
Abstract: Recent work has identified incremental learning in shallow networks trained on single-index and multi-index models. However, existing analyses often rely on simplifying settings, such as small initialization, correlation loss, or layer-wise training. These choices reduce neuron interactions and leave some feature learning dynamics under standard initialization unexplored. We study gradient flow dynamics for polynomial-width two-layer networks learning orthogonal multi-index targets under standard initialization using polynomially many samples. We first prove that incremental learning still occurs: the loss decreases sequentially according to the Hermite expansion of the target, with lower-order components learned before higher-order components recover the individual target directions. In this standard initialization regime, training also shows a competitive reallocation of parameter mass: after the total mass fits the target mean and stabilizes, mass shifts into the target subspace and then concentrates on aligned neurons. Technically, we introduce a symmetry-based finite-width approximation via symmetrized networks, rather than comparing directly with an infinite-width limit. This yields better control of approximation errors and may be of independent interest.
Authors: Johnson Zhou, Daniel Tanneberg, Forough Habibollahi, Alon Loeffler, Kiaran Lawson, Valentina Baccetti, Kwaku D Abu-Bonsrah, Candice Desouza, Finn Doensen, Bradley Watmuff, Daria Kornienko, Azin Azadi, Justin L Bourke, Bernhard Sendhoff, Brett J. Kagan
Abstract: Biological neural networks (BNNs) have been established as a powerful and adaptive substrate that offer the potential for incredibly energy and data efficient information processing with distinct learning mechanisms. Yet a core challenge to utilizing BNN for neurocomputation is determining the optimal encoding and decoding mechanisms between the traditional silicon computing interface and the living biology. Here, we propose an Embodied Neurocomputation framework as a systems-level approach to this multi-variable optimization encoding/decoding problem. We operationalize this approach through the first large-scale parameter optimization of encoding configurations for a BNN agent performing closed-loop navigation along an odor-style gradient in a simulated grid-world. Despite the relative simplicity of the task, the biological interactions gave rise to a massive multi-combinatorial search space for optimal parameters. By considering how the components of the system are interconnected and parameterized, we evaluated approximately 1,300 parameter combinations, over 4,000 hours of real-time agent-environment interactions, to identify 12 configurations that consistently demonstrated learning across multiple episodes. These configurations achieved significantly higher task performances than optimized silicon-based DQN agents under the same interaction budget. These findings represent an initial step toward robust and scalable goal-oriented learning using BNNs. Our framework establishes a foundation for applying task-driven neurocomputing and supports the development of field-wide benchmarks. In the long term, this work supports the development of hybrid bio-silicon architectures capable of efficient, adaptive and real-time computation, including the potential for robotic control applications.
Abstract: While post-training backdoor detection and trigger inversion schemes have been developed for AIs used e.g. for images, there is a paucity of such methods for LLMs. First, the LLM input space is discrete, with up to 150,000^k k-tuples to consider with k the token-length of a putative trigger. Second, one must blacklist tokens typical of the putative target response (class) of an attack, as such tokens may give false detection signals. However, a comprehensive blacklist is not available, in general, for a given domain. We develop a highly effective detection and inversion framework for LLMs treated as classifiers. Central to our approach is class subspace orthogonalization (CSO), a novel plug-and-play paradigm for backdoor detection that serves two fundamental roles when applied to LLMs: (i) it enhances both sensitivity and specificity of a baseline detector; (ii) it provides a form of implicit blacklisting, as it penalizes against inclusion, in a candidate trigger, of tokens that induce signal perturbations "in the direction of" the putative target class of an attack. One version of our detector performs continuous optimization in token embedding space, while a companion trigger-inversion and detection method performs greedy accretion in discrete token space. Our methods give both strong detection performance and accurate inversion of ground-truth triggers on several LLM classification domains, and for several different LLM architectures.
PaperID: 7555, Poster
Abstract: Neural circuits are infamously more robust than their deep network counterparts. Biology often explains this through degeneracy, where structurally distinct configurations support the same function. Deep learning studies a parallel phenomenon through mode connectivity, which links same-objective networks by low-loss curves, sheets, or volumes. Yet real biological systems undergo continual structural change, a more flexible form of degeneracy in which specialized circuits gradually adapt toward new competencies through localized changes while preserving existing function. The analogous question in deep learning is harder than ordinary mode connectivity, asking whether many independently initialized models trained on distinct tasks can be connected by a high-dimensional region of competent multi-task solutions. We introduce Nimbus, which learns a high-dimensional Bézier manifold connecting N specialists, going beyond the one-dimensional paths and two-dimensional sheets of prior mode connectivity work. A naive parameterization would scale combinatorially with N, but Nimbus requires only a small number of learned correction tensors shared per interaction order, showing that the connecting manifold has surprisingly low complexity yet supports multi-task competencies. The resulting manifold from Nimbus forms a navigable "Marauder's Map" connecting independently trained specialists through a shared geometry of competent solutions. Across vision (CIFAR-100) and language modeling (five disjoint text domains), Nimbus produces manifolds that are both functionally and locally weight-flat. Ensembles sampled from them degrade more gracefully than baselines under input and weight perturbations, with the largest gains in negative log-likelihood and expected calibration error. Beyond its empirical behavior, Nimbus supplies a deep learning interpretation of flexible degeneracy, recasting the connectivity between specialized circuits from a low-dimensional path between distant solutions into a high-dimensional volume of nearby ones.
PaperID: 7556, Poster
Abstract: Low-rank decomposition has emerged as a promising paradigm for compressing Large Language Models (LLMs). However, current methods universally rely on greedy, layer-wise reconstruction objective. This approach induces cascading error propagation, where localized truncation errors amplify across deep architectures. To mitigate this limitation, we formulate the low-rank decomposition problem from a global perspective using the Alternating Direction Method of Multipliers (ADMM). By decoupling the non-convex rank constraints from the holistic language modeling objective, our method preserves model quality. Systemically, the standard sequential execution of factorized matrices often fails to yield practical wall-clock speedups due to kernel launch overheads and activation memory traffic. We bridge the gap by developing a fused GPU operator. Extensive experiments across diverse architectures demonstrate that our algorithm-agnostic method significantly benefits existing methods. For example, at a 30% compression ratio on Qwen-3-8B, our method achieves a perplexity of 10.05—nearly indistinguishable from the uncompressed model's 9.85—while further improving zero-shot reasoning accuracy. Furthermore, by integrating our custom fused operator into the widely adopted inference ecosystems, we translate theoretical FLOPs reduction into practical inference speedup.
Abstract: Visual reasoning through reinforcement learning with verifiable rewards (RLVR) has achieved remarkable progress. However, when dealing with multi-source inputs, existing approaches tend to treat them as a mere accumulation of information, lacking explicit mechanisms to distinguish whether integrating additional sources yields information gain or introduces interference. Therefore, they struggle to effectively model dynamic interaction when integrating multiple sources, particularly when they differ significantly in physical properties and semantics, \eg, infrared and depth, leading to inferior performance to mono-source reasoning when a certain source holds the dominant signal. To address this issue, we propose MARS, a novel mono-anchored multi-source reasoning framework that models each visual modality as an independent information source. Specifically, by treating mono-source rewards as dynamic anchors, our method explicitly incorporates the information gain introduced by multi-source fusion into advantage normalization and adaptively emphasizes mutual promotion between sources while suppressing potential noise or conflicts during RLVR. From theoretical analysis, our method effectively quantifies information gain introduced by multi-source integration in gradient estimation, enabling consistent modality regulation. Empirical results also show impressive 3.2% and 4.9% performance gains on GRPO and DAPO across diverse datasets, confirming the effectiveness of our method.
Abstract: Speech-driven talking character animation seeks to generate life-like portrait videos that convey natural conversation behavior, aligning facial motion with spoken audio. Although recent advances in video generation have substantially improved realism in video-based animation, achieving both accurate lip articulation and expressive behavior remains challenging. Existing approaches typically trade off precise phoneme-to-lip synchronization against dynamic facial expressions and head motion, yielding animations that are either accurate yet rigid, or expressive but poorly synchronized. We address this challenge by proposing ReFree-S2V, a flow-matching speech-to-portrait animation framework that builds upon a pretrained video generation model to achieve fine-grained speech articulation and high-level expressive cues in speech-driven portrait animation. This model introduces a multi-level speech representation capturing phonetic and prosodic information at both local and global granularities. These representations are selectively injected into transformer blocks via learnable level selectors, enabling both accurate lip synchronization and natural expressive motion. To achieve natural head movements, we further introduce a novel reward-free reinforcement learning scheme into flow-matching training to discourage perceptually implausible motion without relying on handcrafted synchronization metrics or reward models, or the high cost of human preference annotation. Extensive experiments demonstrate that ReFree-S2V achieves state-of-the-art performance, significantly outperforming existing methods in both quantitative lip-sync accuracy and qualitative human evaluations of naturalness and expressivity.
Abstract: Current large reasoning models (LRMs) have shown strong ability on challenging tasks after reinforcement learning (RL) based post-training. However, previous work mainly focuses on English reasoning in expectation of the strongest performance, despite the demonstrated potential advantage of multilingual thinking, as well as the requirement for native thinking traces by global users. In this paper, we propose ExpLang, a novel LLM post-training pipeline that enables on-policy thinking language selection to improve exploration and exploitation during RL with the use of multiple languages. The results show that our method steadily outperforms English-only training with the same training budget, while showing high thinking language compliance for both seen and unseen languages. Analysis shows that, by enabling on-policy thinking language selection as an action during RL, ExpLang effectively extends the RL exploration space with diversified language preference and improves the RL exploitation outcome. The method is orthogonal to most RL algorithms and opens up a new perspective on using multilinguality to improve LRMs.
Abstract: Recently, rubrics have been used to guide LLM judges in capturing nuanced, multi-dimensional human preferences, and have further been extended as reward signals for reinforcement fine-tuning (RFT). However, open-ended generation typically lacks a unique canonical target, and its evaluation rubrics are latent rather than directly observed. This makes rubric generation under-determined and difficult to control: rubrics often lack coverage, conflate dimensions, misalign preference direction, and contain highly correlated criteria, degrading judge accuracy and producing suboptimal rewards during RFT. We propose RRD, a practical framework for rubric refinement built on a recursive decompose–filter cycle. RRD decomposes coarse rubrics into fine-grained, discriminative criteria, expanding coverage while sharpening separation between responses. A complementary filtering mechanism removes misaligned and redundant rubrics, and a correlation-aware weighting scheme to prevent over-representing highly correlated criteria, yielding rubric sets that are informative, comprehensive, and non-redundant. Empirically, RRD delivers large, consistent gains across all evaluation and training: it improves preference-judgment accuracy on JudgeBench and PPE for both GPT-4o and Llama3.1-405B judges, achieving top performance in all settings with up to +17.7 points on JudgeBench. When used as the reward source for RFT on WildChat, it yields substantially stronger and more stable learning signals, boosting reward by up to 160% (Qwen3-4B) and 60% (Llama3.1-8B) versus ∼10--20% for prior rubric baselines, with gains that transfer to HealthBench-Hard and BiGGen Bench. Overall, RRD establish recursive rubric refinement as a scalable and interpretable foundation for LLM judging and reward modeling in open-ended domains.
PaperID: 7561, Poster
Abstract: Post-discharge risk prediction from electronic health records (EHRs) is difficult because many dependencies that link discharge-time observations to downstream complications, such as comorbidity cascades and drug-disease interactions, are absent from the record. External medical knowledge graphs (KGs) can supply these missing dependencies, but tracing them demands three properties: KG exploration must remain cost-bounded, retrieved evidence must be differentiated by source quality, and the resulting rationale must be citable for retrospective review. Large language models (LLMs) can plan and verify over structured evidence, making them natural candidates for KG reasoning, but existing LLM-based methods do not satisfy these three properties jointly. In this paper, we propose BAR, a Budget-Aware LLM Reasoning framework over medical KGs with three components. First, BAR refines the raw KG into disease-specific evidence graphs whose edges carry support scores and provenance records, turning the KG into a quality-annotated reasoning space rather than a static feature source. Second, an LLM then reasons over this graph through a plan-navigate-verify loop that decomposes the question into steps, retrieves evidence under a patient-specific budget, and revises when verification fails. Third, a reasoning policy is trained with a reward that compares predictions with and without acquired evidence, combined with acquisition cost and citation-integrity terms. Across 8 diseases and 3 prediction horizons on MIMIC-III and MIMIC-IV, BAR improves AUPRC by 3.4 points over the strongest baseline, raises citation precision from 59.8% to 77.9%, and consumes only 62-65% of the budget cap. The code is available at https://anonymous.4open.science/r/BAR/.
PaperID: 7562, Poster
Abstract: Linking brain activity to computational representations of visual perception is a central goal at the intersection of neuroscience and machine learning. However, cross-subject decoding from fMRI remains challenging because neural responses vary substantially across individuals, while existing methods often rely on subject-specific fine-tuning or parameter-heavy decoders. This limits generalization to unseen subjects and hinders efficient deployment. We propose BrainTC--Brain Tensor Cores, a lightweight framework for cross-subject brain visual decoding based on shared tensor decomposition. BrainTC separates shared functional structure from subject-specific variation through a shared tensor-core alignment module that maps ROI-wise responses into a common functional space with shared tensor bases and residual subject factors. This supports both zero-shot decoding for unseen subjects and few-shot adaptation by updating only residual parameters. We then develop a Brain Tensor-Transformer with tensorized attention to compactly model inter-ROI dependencies, together with a hierarchical neural-to-visual mapping module for multi-level visual prediction. Experiments show that BrainTC achieves competitive cross-subject decoding performance with substantially improved parameter efficiency, particularly in zero-shot generalization to unseen subjects.
PaperID: 7563, Poster
Abstract: Effective team collaboration hinges on consensus formation through information exchange, a principle equally critical in multi-agent reinforcement learning (MARL). However, existing communication-based and consensus-learning methods often struggle to form a coherent global understanding from agents' limited local views. We propose STAGE, a local-to-global progressive consensus framework based on hierarchical communication. By first grouping agents according to their perceptual focuses, STAGE enables intra-group communication to form local consensuses, then selects group leaders to exchange these consensuses across groups, progressively expanding them into a unified global consensus with reduced communication redundancy. To further improve consensus quality, we introduce KL-divergence constraints for consensus alignment and a variational autoencoder (VAE) objective for preserving task-relevant global information. Extensive experiments on challenging MARL benchmarks show that STAGE consistently outperforms state-of-the-art baselines, especially on more difficult tasks and larger-scale multi-agent systems.
Authors:
Zisu Huang, Jingwen Xu, Yifan Yang, Ziyang Gong, Qihao Yang, Muzhao Tian, Xiaohua Wang, Changze Lv, Xuemei Gao, Qi Dai, Bei Liu, Kai Qiu, Xue Yang, Dongdong Chen, Xiaoqing Zheng, Chong LuoAbstract: Language agents increasingly improve by reusing \emphskills---structured procedural artifacts distilled from past experience. In particular, \emphdomain-level and \emphmodel-generated skills are especially promising. They offer fast adaptation within a domain by encoding domain-specific recurring procedures, and they scale beyond labor-intensive hand-crafting. However, while extraction methods continue to proliferate, understanding remains limited, with no comprehensive study spanning the full skill lifecycle---experience generation, skill extraction, and skill consumption---to ask whether such skills actually work, when they work, and what makes them succeed or fail. To close this gap, we build a utility-grounded evaluation framework that provides systematic experimental results across extractors and target agents, covering five diverse agentic task domains. We find that model-generated skills are beneficial on average but exhibit non-trivial negative transfer, and that neither extractors nor targets behave uniformly. A model can be a strong extractor yet a weak consumer, or vice versa, with skill utility independent of model scale or baseline task strength. To explain these patterns, we then dissect each lifecycle stage in depth, analyzing how experience composition shapes skill quality, what properties characterize useful skills, and how the same skill transfers across different consumers. Finally, we translate these findings into a concrete \emphmeta-skill that guides skill extraction toward the features tied to actual utility, which consistently improves skill quality across domains and substantially reduces negative transfer.
PaperID: 7565, Poster
Authors: Xinyu Guo, Jiajia Xie, Xin He, Christophe Ye, Batuhan Nursal, Hongyu Xue, Cassie Mitchell
Abstract: Modeling neurodegenerative pathology from observational cohorts is challenging because subjects are typically observed as unpaired endpoint snapshots rather than aligned longitudinal trajectories. Existing approaches to pathology transport are either endpoint-conditioned or population-level, yet neither extreme alone yields the best terminal-distribution fit. We introduce the Doob-Bridge Posterior Sampler (DBPS), a controlled-SDE posterior sampler with balanced endpoint-population control that blends individual endpoint guidance with population-level guidance. This ρ-weighted controller, obtained as the Cole–Hopf log-gradient of a geometric interpolation between endpoint and population desirability functions, is the exact optimal feedback policy of a linearly-solvable stochastic optimal control problem with closed-form terminal and running costs. Endpoint-conditioned bridge sampling and population Doob sampling are the ρ = 0 and ρ = 1 boundary cases; ρ ∈ (0, 1) realizes the genuine balanced regime. We evaluate DBPS on tau PET pathology transport in Alzheimer’s disease across ADNI and NACC cohorts (84 brain regions). Compared to representative baselines from all major unpaired-transport families, DBPS achieves state-of-the-art performance across multi-cohort settings (union training: SWD 0.760 vs. 0.821; held-out ADNI: 0.822 vs. 0.859), and remains competitive in single-cohort training (ADNI: 0.868 vs. 0.876). Beyond distribution matching, the learned controls recover canonical Braak-stage pathology organization without anatomical supervision, suggesting that balanced SOC objectives capture biologically meaningful structure in latent neurodegenerative progression.
PaperID: 7566, Poster
Abstract: In this paper, we investigate the long-run behavior of discounted regularized learning with stochastic gradient feedback in general, non-concave games. Specifically, we focus on a family of implicitly regularized exponential / multiplicative weight update schemes, and we seek to determine which actions\textemdash or, more generally, which recurrent patterns of play\textemdash are more likely to arise in the long run. We approach this question through the lens of large deviations theory and randomly perturbed dynamical systems, and we obtain a precise characterization of the distribution of the process: in the long run, it follows a Boltzmann-Gibbs law with temperature equal to the method's step-size, and energy levels determined by the game and the statistics of the noise. Concretely, we show that the distribution of play concentrates exponentially around the dynamics' (ICT) sets - i.e., irreducible invariant sets containing no smaller attractors - and the probability of visiting such a set depends exponentially on its energy. As a result, unstable ICT sets are exponentially less likely to be visited than stable ones, and the sequence of play is exponentially concentrated around the problem's ``ground state'', where energy is minimized. In this manner, learning acts as a selection mechanism: with exponentially high probability, the ground state is the only outcome observed in the long run, even in the presence of multiple equilibria and other attractors.
PaperID: 7567, Poster
Authors: Cagri Gungor, Qingshuang Chen, Hongda Mao, Chi Zhang, Yelin Kim
Abstract: Grounding natural language queries to target objects in images requires both high-level reasoning about intentions and precise spatial localization. Multimodal large language models (MLLMs) excel at reasoning about which object is desired given an implicit intention query but turning this understanding into precise bounding box predictions remains challenging. Conversely, specialized detectors such as GroundingDINO offer robust localization capabilities but lack the inherent capacity to understand implicit intentions. Existing hybrid methods couple these components only at the semantic level, such as by passing object names or a single global token to the detector, thereby discarding the rich, dense spatial signals encoded in MLLM attention maps. We propose Reasoning-based Spatial Prior (RSP), a reasoning-based detection framework that converts the MLLM’s internal attention into a learnable spatial prior and injects it directly into the detector. Concretely, we learn input-adaptive gating over attention heads, supervise the fused spatial prior with a SoftIoU localization loss, and integrate it into GroundingDINO via a prior-guided deformable fusion module that biases sampling toward regions highlighted by the reasoning process. Our method consistently outperforms strong MLLM-only, detector-only, and hybrid baselines on the EgoIntention and RIO benchmarks, achieving state-of-the-art performance with substantial gains across both frequent and uncommon categories.
PaperID: 7568, Poster
Authors: Ferhat Arslan, Weihong Guo, Shuo Li
Abstract: Multilayer perceptron (MLP)-based model compression often overlooks two distinct sources of error: the hypothesis gap and the capacity gap. The hypothesis gap arises from approximation error induced by selecting a student hypothesis class that is mismatched to the teacher function, while the capacity gap captures the additional error introduced by enforcing a limited parameter budget within the chosen class. We show that Kolmogorov–Arnold Networks (KANs) can reduce the hypothesis gap for smooth low-dimensional target functions. Furthermore, width and depth constraints in KANs allow within-family compression to be formulated as a convex projection problem that directly targets the capacity gap. Building on this formulation, we derive projection-based algorithms for width and depth reduction, as well as a function transfer method that maps a trained MLP teacher into a compact KAN student. These are unified in the Gap-aware Projections via KANs (ProjKAN) framework, which integrates the error decomposition, theoretical guarantees, and compression algorithms into a single methodology. Empirically, in regimes where pruning and standard knowledge distillation degrade sharply, ProjKAN achieves improved accuracy–size tradeoffs, reduced train–test gaps, and greater robustness in data-scarce settings.
PaperID: 7569, Poster
Abstract: In the computer vision world, vision-language and self-supervised encoders represent two unique paradigms characterized by global language alignment and high-quality vision-only patch representations, respectively. Multimodal large language models (MLLMs) often use the former because of language alignment, despite the recent rise in vision-centric tasks resembling traditional dense vision tasks that would likely benefit from the high-quality patch representations of the latter. This raises the question whether self-supervised models could aid MLLMs in vision-centric tasks. To investigate this, we define the concept of semantic (in)consistency of patch tokens produced by language-aligned encoders and identify two sets of tokens, semantically consistent (SCTs) and semantically inconsistent (SITs). Through systematic probing via the removal of SCTs or SITs from the input of modern MLLMs, we show higher reliance on SCTs than SITs, as removal of the former induces larger performance drops. This demonstrates that MLLMs benefit from semantic consistency of vision tokens. With this in mind, we devise a simple yet effective extension of the MLLM training objective that combines traditional language modeling with semantic guidance via minimization of semantic inconsistency of vision tokens. This strategy, named Semantically-Guided Training (SGT), combines the signals of both loss terms to push the vision encoder towards providing semantically consistent tokens, while preserving language alignment. Through extensive evaluation, we show that SGT drastically outperforms all baselines and MLLMs with specialized vision encoders on all vision-centric benchmarks while keeping strong scores on text-centric benchmarks, thus setting state of the art when averaged across all tasks. Our results highlight the importance of combining language alignment and semantic consistency through the lens of MLLMs rather than independently from them and pave the way towards the design of better MLLM vision encoders.
PaperID: 7570, Poster
Abstract: On-policy distillation transfers reasoning capabilities by training a student model on its own generated trajectories using token-level feedback from a teacher. However, we identify a critical bottleneck, Supervision Fidelity Decay (SFD): as student-generated prefixes lengthen, the teacher’s next-token distribution becomes less confident and less discriminative. Consequently, the teacher-dependent corrective signal in reverse-KL distillation weakens, causing student drift to compound across long reasoning chains. To mitigate SFD, we introduce Lookahead Group Reward (LGR). Building on the insight that next-step teacher confidence reflects the discriminative strength of future reverse-KL supervision, LGR evaluates the student’s top-K candidate tokens by the teacher confidence they induce at the subsequent step and assigns a group-normalized reward. To maintain computational efficiency, we further design an entropy-triggered tree-attention mechanism. Across six math and code benchmarks, LGR improves mean@8 by 2.57 points over OPD for a 7B student, with gains increasing in longer-generation and reaching 4.92 points on AIME-26 at 39k tokens.
PaperID: 7571, Poster
Abstract: LLM chain-of-thought reasoning improves with additional compute, but standard 2 inference scaling [Wei et al., 2022, Wang et al., 2022, Snell et al., 2024, Brown 3 et al., 2024] treats all steps equally. We ask: which step in a reasoning chain benefits 4 most from extra computation? We make three layered contributions, each scoped 5 explicitly. (I) Measurement. We define expansion utility Uk as the marginal 6 gain in correctness probability from resampling from step k onward, measured via 7 a step-oracle sweep. Across nine model families (8B–123B), seven benchmarks 8 (math, science, code), and 6,500 chains, every one of the 41 healthy (model, bench) 9 cells shows positive oracle gap at 10% step budget (95% Wilson CI [0.914, 1.000]; 10 mean +22.0 pp, range +1 to +48 pp). (II) Structure. Per-step utility is sparse 11 (97.6% of cells: Gini ≥ 0.80); Gini and mean oracle gap are tightly anti-correlated (Pearson r = −0.913, p = 9.5 × 10−17 12 ). A resample-stability simulation and 13 a segmentation-rule sensitivity check (Appendix J) bound the headline against 14 noise and rule-choice confounders. (III) Routing proof-of-concept. A cross15 model router trained with a within-example listwise loss [Cao et al., 2007] and a 16 learned depth prior, combined at inference with a bench-conditional depth mask 17 and a confidence gate, recovers 5-seed-averaged gated GR@10% = +0.149 (95% 18 bootstrap CI [+0.085, +0.217]) of the cross-model oracle gap at matched step19 budget, with no architectural change to the underlying model. We frame this as 20 proof the framework is actionable cross-model, not as deployment-ready scaling; 21 matched-token comparison to self-consistency / best-of-N requires a separate 22 generation campaign and is left to follow-up (§7). Methodological observation. 23 Within-example pointwise top-k BCE under cross-model pooling produces a high24 AUC discriminator (AUC@10 = 0.70) that selects the wrong steps (gated GR 25 = −0.15); the listwise loss is shift-invariant to per-(model, bench) utility scale and 26 removes this inversion.
PaperID: 7572, Poster
Authors:
Zhihao Zhan, Jiaying Zhou, Likui Zhang, Qinhan Lyu, Hao Liu, Weizheng Li, jusheng zhang, Ziliang Chen, Tianshui Chen, Ruifeng zhai, Keze Wang, Liang Lin, Guangrun WangAbstract: Vision–Language–Action (VLA) models offer a unified framework for robotic manipulation by integrating visual perception, language understanding, and control generation. However, existing VLA systems still struggle to generalize across diverse tasks, scenes, and camera viewpoints, and often produce coarse or unstable actions. We argue that these limitations are closely tied to the structural properties of actions in VLA settings, including the inherent multi-peaked nature of action distributions, the token-based symbolic reasoning of pretrained VLM/VLA backbones, and the effective finite resolution imposed by real-world robotic control. Motivated by these properties, we introduce E0, a tweedie discrete diffusion framework that formulates action generation as iterative denoising over quantized action tokens. By operating in a discrete action space with a principled diffusion process, E0 naturally aligns with token-based reasoning, supports fine-grained yet executable action control, and avoids the distributional mismatch of masking-based discrete diffusion. Experiments on LIBERO, VLABench, ManiSkill, and a real-world Franka arm demonstrate that E0 achieves state-of-the-art performance across diverse environments, outperforming strong baselines by 8.3% on average.
PaperID: 7573, Poster
Abstract: LLMs are widely used, yet they remain prone to factual errors that erode trust and limit adoption in high-risk settings. One approach to mitigate this risk is to equip models with uncertainty estimation mechanisms that abstain when confidence is low. However, this binary "all-or-nothing" approach is excessively restrictive in long-form settings, often discarding valuable information. We introduce , a framework that enables LLMs to trade specificity for reliability by selectively reducing the detail of uncertain content. We formalize SA through selective risk and coverage, and propose , a claim-level instantiation that decomposes responses into atomic claims and replaces uncertain atoms with higher-confidence, less specific abstractions. We develop a novel end-to-end pipeline for evaluating long-form generation that instantiates risk as factual correctness and measures coverage as retained information. Our comprehensive empirical evaluation across five LLMs and two long-form factual benchmarks shows that atom-wise SA consistently improves the risk-coverage trade-off, outperforming seven baselines, including claim removal, self-revision, sampling-based aggregation, and chain-of-verification. It improves AURC by up to 35.61% over claim removal, demonstrating that reducing specificity boosts reliability while preserving substantial information. We further provide a risk-guided procedure for selecting a confidence threshold that satisfies a target risk.
PaperID: 7574, Poster
Authors:
Zouquan Chen, Yifei Yang, Lidong Zheng, Zhiming Fang, Yuchen ZhengAbstract: Federated Spiking Neural Networks (FedSNNs) provide an energy-efficient and privacy-preserving learning paradigm for edge intelligence. However, their robustness under Non-IID data is limited by a coupled failure mode: label shift distorts membrane potential distributions, drives neurons away from gradient-sensitive regions, and further amplifies surrogate gradient errors, which destabilizes federated optimization. To address this problem, this paper proposes Spatio-Temporal Spiking Adaptive Thresholds (ST-SAT), a framework that mitigates the coupling between label shift and surrogate gradient errors through coordinated neuron-level adaptation and client--server knowledge alignment. At the neuron level, ST-SAT introduces the Spatio-Temporal Parametric Leaky Integrate-and-Fire (ST-PLIF) neuron, which adaptively calibrates firing thresholds, membrane potentials, and decay dynamics to keep local neuronal responses within effective gradient regions. At the collaborative interaction level, ST-SAT adopts a Dual-Scale Knowledge Distillation strategy that combines adaptive logit-level distillation with feature-level alignment of neuronal statistics, thereby reducing class distribution bias and improving knowledge consistency across heterogeneous clients. Extensive experiments on three benchmark datasets show that ST-SAT consistently improves both global model performance and client-local performance, with particularly strong gains under severe label shift. These results validate the effectiveness of the proposed framework for robust federated SNN learning. The anonymous code link is https://anonymous.4open.science/r/ST-SAT-4771.
PaperID: 7575, Poster
Abstract: Large Language Model-based Automated Algorithm Design (LLM-AAD) has shown promising results across diverse domains by coupling LLMs with iterative search. Some recent efforts have begun to fine-tune LLMs for algorithm design through reinforcement learning. However, these attempts still rely on outcome-level supervision, leaving intermediate ideas and plans without direct credit, even though they determine the conceptual direction of algorithm design. In this paper, we propose Algorithm Tree Policy Optimization (ATPO), which utilizes tree-structured rollouts to decouple the generation of conceptual ideas from their specific code implementations. By applying group-relative advantage estimation, ATPO assigns credit to both strategic ideas and algorithmic implementations, thereby optimizing the policy across multiple levels of abstraction. Integration with FunSearch, EoH, and OpenEvolve demonstrates improvements over their original counterparts. Notably, the learned policy transfers effectively to unseen problem instances and alternative search methods, indicating that it can improve the generalization of algorithm design policies.
PaperID: 7576, Poster
Authors:
Zehao Liu, Huashuo Lei, Xi Lin, Haoang Li, Yuliang Chen, Huarui Zhang, Yifan WangAbstract: As AI models evolve from text and vision to physical agents, embodied AI faces a fundamentally different attack surface. Prior attacks on embodied systems have largely focused on semantic or instruction-level manipulations, such as prompts, adversarial images, and action commands; by contrast, we study environment-only perturbations that require no access to the model or instructions. We reveal a new attack surface: the physical environment itself. An adversary who rearranges objects, without modifying instructions or accessing models, can cause hazardous consequences. In this work, we propose EnvTrap, a diagnostic pipeline that constructs paired safe, trap, and null-trap (benign but misleading) embodied scenarios. We demonstrate that environment-only perturbations raise hazardous-action rates to an average of 86.5% across multiple vision-language-action (VLA) models, with similar vulnerability patterns confirmed in world models. On a consequence-prediction task, model accuracy remains near chance, while human evaluators succeed easily. We further propose a consequence-aware defense that reduces trap trigger rates by an average of 76.3% across VLA models in simulation and by 61.7% on physical robots. This vulnerability arises because current embodied models can recognize scene state but often fail to predict action consequences under altered layouts. Our findings establish environment integrity as a prerequisite for safe embodied AI deployment. Our code and data are available at https://anonymous.4open.science/r/Envtrap-1BA4/
PaperID: 7577, Poster
Authors: Jiarong Fan, Juhyun Park, Thi Phuong Thuy Vo, Nicolas Brunel
Abstract: Conformal prediction (CP) provides a distribution-free framework for uncertainty quantification with finite-sample marginal coverage guarantees. Yet, missing values introduce distributional heterogeneity, where marginal coverage can hide systematic undercoverage for certain missingness patterns. At the same time, existing methods for conditional guarantees often fail when missing values are present, as they rely on geometric continuity of the covariate space, which is broken by missingness. Although imputation restores continuity, it obscures the missingness patterns and introduces task-dependent overhead. We develop a new CP framework that is \emphadaptive to missingness patterns, without imputing covariates. We rely on data augmentation by feature embedding that accounts for missingness. Our framework represents each input with missing values by a tree embedding together with its raw missingness mask, and then learns a regularized quantile score function. Varying the choices for the penalty function and the augmentation enables us to characterize three practically relevant validity levels: exact coordinate-wise validity, exact validity on observed masks, and approximate missingness-conditional guarantees for general mechanisms (MCAR, MAR, MNAR). Experiments on synthetic and real datasets support these guarantees and show that the proposed method reduces miscoverage risk while maintaining informative intervals, especially in non-MCAR settings. These results highlight the practical value of uncertainty quantification that adapts to missingness.
Abstract: Data attribution and valuation are critical for understanding data-model synergy for Large Language Models (LLMs), yet existing gradient-based methods suffer from scalability challenges on LLMs. Inspired by human cognition, where decision making relies on a focused readout of relevant memories rather than replaying all pathways, we introduce RISE (Readout Influence Sketching Estimator). Instead of computing and indexing gradients across the entire LLM, RISE focuses on influence hotspots at the output layer, where influence signals concentrate, and the gradient admits a decomposed outer-product form. This enables a dual-channel representation combining a lexical residual channel (RH) and a semantic projected-error channel (GH). Applying CountSketch projections to these channels achieves strong compression while maintaining accurate attribution. Across the OLMo (1B–32B) and Pythia (14M–6.9B) families, RISE reduces index storage by up to 112× compared to RapidIn and scales to 32B parameters LLM, where gradient-based baselines such as RapidIn and ZO-Inf become memory-infeasible. We evaluate RISE on two paradigms: (1) retrospective attribution, retrieving influential training examples for specific predictions, and (2) prospective valuation, scoring candidate data utility zero-shot. We validate RISE on three tasks: Howdy backdoor data detection, Finance-Medical domain separation, and Brain Rot high-quality data selection. In a closed-loop Brain Rot study, continued pretraining on RISE-selected data yields consistent downstream improvements. Overall, RISE provides a practical and scalable primitive for influence analysis and training-data selection in modern large language models.
PaperID: 7579, Poster
Abstract: Hybrid models and neural differential equations (NDEs) are becoming increasingly important for modelling physical systems; however, they often encounter stability and accuracy issues during long-term integration. Training on unrolled trajectories is known to limit these divergences but quickly becomes too computationally expensive due to the need to compute gradients over an iterative process, making this method impractical for large-scale physical systems (e.g. weather and climate forecasting). Two popular paradigms for improving the stability of learned dynamics are to train on interpolated or noisy samples or to regularize the Jacobian of the learned representation. In this work, we demonstrate that these paradigms can be understood under a common lens and that they both correspond to an approximation of the gradients obtained in an unrolled learning strategy. Building on this understanding, we derive a principled and complete approach to regularize training and demonstrate its impact on the stability of learned dynamical systems.
PaperID: 7580, Poster
Abstract: Forensic verification has converged on a binary "real vs. fake" label that conflates fully synthetic, tampered, and AI-retouched images despite their very different consequences. We argue that GenAI manipulations decompose along two complementary channels: a camera channel of sensor fingerprints (destroyed by neural processing), and a semantic channel of scene-level coherence (disrupted by content edits). Each manipulation type leaves a distinctive two-channel signature. We instantiate this taxonomy in 2CAP (2-Channel Authenticity Protocol), pairing a contrastively-trained camera encoder with a frozen semantic encoder to serve two applications on a shared backbone: (i) four-class authenticity classification via cross-attention fusion with learnable reliability weights; and (ii) evidence generation, where an authentic image serves as a reference for a query and the two encoders produce per-channel alignment scores together with a patch-level saliency map. These signals are supplied as privileged context to a frozen Vision Language Model (VLM) through an Observe-Generate-Refine loop, yielding explanations with localization without forensic fine-tuning. On a multi-source benchmark, 2CAP attains the best classification performance, including strong retouching detection where existing methods fail, and improves explanation quality.
Abstract: Intracranial electrocorticography (ECoG) offers high–signal-to-noise access to cortical activity for brain–computer interfaces, yet limited per-patient data has led most prior work to rely on small, subject-specific decoders that neglect information shared across patients. We investigate whether large pretrained scalp-EEG foundation models (EEG FMs) can be adapted to ECoG, enabling cross-patient learning and competitive decoding performance while calibrating to a held-out patient in 10--30 minutes on a single GPU. We introduce CORTEG, a cross-modality transfer framework that combines a pretrained EEG FM backbone, an electrode-aware KNNSoftFourier spatial adapter, a dual-stream tokenizer for low-frequency and high-gamma activity, and a leave-one-subject-out fine-tuning strategy. We evaluate CORTEG on two challenging regression tasks: public finger trajectory regression (n=9) and private audio envelope regression (n=16). CORTEG matches or exceeds the strongest task-specific baselines on both tasks: it reaches the highest mean correlation among compared methods on the public finger benchmark (gain not statistically significant on n=9 subjects), with larger and statistically significant gains on the audio task and in low-data per-patient calibration. Feature analyses align with neurophysiology, and latent manifolds capture low-dimensional finger-movement structure. CORTEG provides systematic evidence that scalp-EEG pretraining can be repurposed for ECoG decoding, enabling data-efficient intracranial BCIs that can adapt to new patients.
PaperID: 7582, Poster
Abstract: Video Reasoning Segmentation (VRS) demands both complex semantic reasoning and robust spatiotemporal mask propagation. Existing paradigms typically compress reasoning into an isolated, opaque token (e.g., \texttt[SEG]) via expensive fine-tuning, or rely on training-free prompt-then-track pipelines that inevitably suffer from semantic degradation and identity switches when encountering hard distractors. To fundamentally address these vulnerabilities, we propose Counterfactual Logic and Explicit Anchoring Reasoning (CLEAR) framework, a training-free VRS framework inspired by human dual-system cognition. Specifically, we introduce a Causal-Counterfactual Inference mechanism that transforms error-prone coordinate regression into a tractable discrete instance selection task, establishing highly reliable initial anchors via rigorous bidirectional reasoning. During mask propagation, we design an adaptive mode-switching mechanism bridged by a Heterogeneous Event Monitor. The fast-thinking mode utilizes a purely visual engine for efficient temporal propagation, while the Monitor continuously evaluates its reliability across multiple dimensions. Upon detecting propagation anomalies, the slow-thinking mode intervenes and explicitly replaces \texttt[SEG] tokens with a structured Contrastive Concept Dictionary for closed-loop identity recovery..Extensive experiments demonstrate that CLEAR significantly outperforms existing training-free methods and achieves performance comparable to state-of-the-art training-based approaches on comprehensive referring and reasoning VOS benchmarks. Code will be available after acceptance.
PaperID: 7583, Poster
Abstract: Despite rapid progress in indoor vision-and-language navigation (VLN), its aerial counterpart remains significantly more challenging and underexplored. In this setting, unmanned aerial vehicles (UAVs) must interpret free-form instructions and traverse kilometer-scale outdoor environments. Although recent UAV-VLN methods have shown promising results, they often fail to fully exploit contextual signals across modality, time, and prior experience. This limitation appears in three aspects: the predominant reliance on RGB-only inputs, the use of discrete frame inputs that fail to capture temporal dynamics, and the absence of episodic memory, resulting in unreliable termination decisions. To this end, we propose STEMFly, a unified framework grounded in Sensor signals, Temporal diversity, and Episodic Memory. STEMFly comprises three components. First, Sensor-Augmented Prompt Injection (SAPI) incorporates multi-modal sensor signals as semantically grounded language prompts that can be interpreted by the language model. Second, Temporal-Diversity Frame Selection (TDFS) constructs a temporally informative observation sequence via entropy-guided frame sampling. Third, Memory-Augmented Success Verification (MASV) improves termination reliability by validating navigation outcomes through episodic memory retrieval. In addition to methodological contributions, we identify and rectify systematic annotation inconsistencies in the OpenUAV benchmark, releasing a revised evaluation protocol. On the original OpenUAV benchmark, STEMFly outperforms the prior state-of-the-art TravelUAV, achieving absolute improvements of 8.26% in Success Rate (SR) and 6.70% in Success weighted by Path Length (SPL), while reducing Navigation Error (NE) by 9.55 meters. Similar consistent gains are also observed on our rectified benchmark. These results demonstrate that holistic contextual grounding, which integrates multi-modal sensing, temporal modeling, and episodic memory, is crucial for improving UAV-VLN performance.
Abstract: Multi-Agent Debate (MAD) improves the reasoning performance of Large Language Models (LLMs) through multi-round interaction. However, LLMs in MAD are highly susceptible to blind conformity. Existing individual evaluation methods, typically based on confidence or perplexity, fail to reflect the correctness of reasoning and may even exacerbate blind conformity. To address this, we shift the perspective from individual evaluation to group interaction. We define mutual referencing among LLMs as Debate Relationships and recognize that regulating these relationships is the key to mitigating blind conformity. In this paper, we propose a novel framework for Dynamically rEgulating debAte Relationships (DEAR) from the group perspective. At first, DEAR quantifies consensus and divergence as group evidence to capture the debate state. Then, DEAR operates through three stages: 1) What: perceiving group consultation tendency and uncertainty; 2) Who: introducing a Selection RL-Agent to dynamically select reference peers; and 3) How: adopting a Behavior RL-Agent to adaptively adjust generation behaviors.
PaperID: 7585, Poster
Abstract: Link prediction is a standard pretraining objective for graph neural networks: hide edges, train an encoder--decoder to predict them, and reuse the node embeddings for downstream node tasks. Despite its empirical success, there is little quantitative theory explaining why an edge-level objective should produce embeddings that are linearly useful for node classification.We study this question in the finite-sample stochastic block model (SBM). For any bounded GNN encoder paired with a symmetric bilinear decoder, we prove a conditional reduction: small population held-out LP excess risk implies small community-classification error for a transductive linear classifier on the learned embeddings.The bound decomposes into an optimization residual, an architectural approximation residual, and explicit statistical terms of order \widetildeO(\sqrtd/n) in bounded dense regimes. We further give a constructive few-label classifier: nearest-centroid classification under the PSD quadratic form G = W Z^\top Z W, whose label error is coupled to LP quality.The framework is architecture-agnostic through the approximation residual, and we instantiate this residual as O(\log n / n) for a spectral-positional-encoding GIN with a bilinear decoder. Synthetic SBM experiments validate the theorem chain from LP risk to score error to classification error, and show when the worst-case constants become conservative.
PaperID: 7586, Poster
Authors:
Niloufar Mehrabi, Sayedpedram Haeri Boroujeni, Abolfazl RaziAbstract: The emergence of specialized, domain-tuned Large Language Models (LLMs) has demonstrated that smaller models can achieve expert-level performance in specific tasks, while struggling in out-of-domain settings. Current ensemble methods to combine their complementary expertise primarily rely on iterative re-prompting or cross-model refinement. These approaches suffer from high computational costs and latency because they require repeated LLM inference calls. Furthermore, naive aggregation often leads to anchor corruption, in which noise propagated from weaker models degrades the performance of the most accurate expert. To address these challenges, we propose a framework that integrates model predictions at the semantic layer using a bipartite factor graph. In this architecture, individual LLMs are represented as variable nodes, while a set of check nodes assess their consistency based on diverse epistemic criteria. We develop a message-passing protocol inspired by error-recovery systems to resolve disagreements iteratively. Furthermore, we introduce an asymmetric damping mechanism that protects high-reliability anchor nodes from being overridden by the ensemble majority. Unlike existing methods, our approach operates entirely on output distributions and requires no additional LLM calls during the refinement phase. Evaluating on four benchmarks, including MMLU, MMLU-Pro, GPQA, and MedMCQA, our method demonstrates a 97% reduction in token usage and up to a six-fold decrease in API calls, reducing inference time from several minutes to mere milliseconds while consistently outperforming leading multi-agent baselines. These results suggest that graph-based belief propagation is a robust, high-speed, and scalable alternative to the current multi-agent LLM systems. The full pipeline and code will be made public.
PaperID: 7587, Poster
Authors: Zihua Wang, Xiaopei Jiao, Yunfeng Cai
Abstract: We propose a physics-informed method for solving high-dimensional optimal control problems that requires only control labels, not value function labels. A neural network parameterizes the value function, and the control policy is derived from its gradient. The network is trained by minimizing a weighted sum of the standard Hamilton–Jacobi–Bellman (HJB) residual and the mismatch between the network’s control output and control labels, which are precomputed by a state-dependent Riccati equation (SDRE) solver. Theoretically, with the help of the viscous HJB equation, we establish error bounds for the value function and its gradient, providing a rigorous performance guarantee for the resulting policy. Numerical experiments show that the learned model accurately approximates the optimal value function, closely matches the cost-to-go of the SDRE reference, and yields a neural controller that is significantly faster while exhibiting excellent stability.
Abstract: Autoregressive (AR) video diffusion is a powerful paradigm for streaming and interactive video generation. However, its reliance on softmax self-attention leads to quadratic compute complexity in sequence length and increasing memory usage due to key-value caching, which fundamentally limits its scalability to long video horizons. Existing remedies (e.g., sparse attention and KV-cache compression) reduce per-step cost but still rely on a linearly growing cache or irreversibly discard past context, and thus fail to simultaneously address linear memory growth and streaming context management. To address this scalability bottleneck, we propose ARL^2 (Attend Locally, Remember Linearly), a hybrid attention module that replaces quadratic cross-frame attention with a fixed-size recurrent state. Specifically, we decompose self-attention into two branches: an intra-frame softmax branch for spatial detail and local dependencies, and an inter-frame gated recurrent linear branch that maintains a fixed-size state for streaming context. Our key insight is that softmax attention captures fine-grained local interactions, while a recurrent state provides a controllable mechanism for long-range memory. This design achieves linear-time scaling with constant memory while improving temporal consistency over the full-softmax model. To prevent noisy intermediate states from corrupting long-range memory, we update the recurrent state only after the denoised pass. To avoid within-frame information asymmetry, all tokens share the same pre-update state rather than sequential updates. To the best of our knowledge, this is the first work to convert a pretrained AR video diffusion model into a hybrid linear attention architecture, through an efficient two-stage training scheme adapted to the AR video setting. With 75% of layers replaced by hybrid linear attention, the model achieves up to 2.26× wall-clock speedup and 54% memory reduction, while maintaining comparable quality with improving temporal consistency.
Abstract: Model distillation can transfer not only intended capabilities but also hidden traits from the teacher. A teacher biased by a system prompt can generate semantically clean training data, such as numeric sequences, that still causes a downstream student to inherit the hidden preference, a phenomenon known as subliminal learning. Prior work establishes this phenomenon, but the transfer mechanism remains unclear, making targeted mitigation difficult. We propose and validate trait-direction drift as the mechanism underlying subliminal learning: (1) on the teacher side, biased teacher samples retain a measurable preference gap toward the target trait even when their semantic content is unrelated to that trait; (2) on the student side, under a low-rank logit-linear approximation, samples with larger preference gaps induce stronger trait-aligned updates during supervised fine-tuning, and these updates accumulate over training. Guided by this mechanism, we propose probe-space corridor, a regularizer that constrains drift along a calibrated trait readout during distillation. The method substantially reduces hidden-trait transfer while preserving task performance: for example, it lowers malicious-response transfer from 29.55% to 6.47% with low main-task accuracy cost, and consistently suppresses animal-preference transfer across the main Qwen-7B preference settings. These results provide both a validated mechanistic explanation of subliminal learning and a targeted recipe for more controllable distillation.
PaperID: 7590, Poster
Abstract: Token-level dynamic sparse attention exemplified by DeepSeek Sparse Attention (DSA) selects the globally most relevant key-value tokens via an exact Top-K operator, achieving superior model quality over block-level alternatives. However, this exact selection creates a severe distributed inference bottleneck: enforcing an exact global Top-K across GPUs inevitably incurs either redundant full-context retrieval or costly multi-stage cross-device synchronization, which largely negates the computational advantages of DSA at long context lengths. We first show that the exact Top-K bound is unnecessary during inference: once the truly critical tokens are recalled, admitting additional context preserves or even improves accuracy. Leveraging this insight, we propose Interleaved DeepSeek Sparse Attention (IDSA), which distributes tokens across GPUs in an interleaved layout so that each device performs only a relaxed local Top-m selection. Under this layout, the union of independent per-GPU Top-m selections near-completely covers the globally most relevant Top-K tokens. This allows each device to proceed with its local selection with minimal cross-GPU overhead while avoiding both expensive full-context Top-K computation and multi-stage cross-GPU merging, enabling a not only distributed but also synchronization-efficient inference pipeline. Without any retraining, IDSA delivers dramatic throughput gains for context lengths exceeding 100K tokens on both DeepSeek-V3.2 and GLM-5, while preserving equivalent or better reasoning performance on the AIME and Needle-In-A-Haystack benchmarks.
PaperID: 7591, Poster
Authors: Hugo Schmutz, Hachem Kadri, Thierry Artières
Abstract: We address the problem of evaluating machine learning models under a limited labelling budget in streaming environments, where data arrive sequentially. This setting is particularly critical for embedded and continuously deployed systems, such as autonomous driving or computer-aided diagnosis, where models must reliably assess their own performance online. We focus on a constrained online setting, where labelling decisions must be made immediately, and data cannot be stored. We introduce an \emphonline active testing (OAT) framework based on an adaptive importance sampling estimator that provides unbiased risk estimation for both regression and classification tasks. While importance sampling has been extensively studied in pool-based settings, we establish novel theoretical ground for the fully online setting, including deviation bounds and consistency guarantees. We derive an optimal sampling strategy in order to minimise the estimator's variance, which is driven by the expected second moment of the loss. This strategy is expressed for common loss functions, including mean squared error, mean absolute error, 0–1 loss, and log-loss, and analysed from a statistical perspective. We empirically demonstrate the effectiveness of OAT on synthetic benchmarks, deep learning tasks, and a real-world application involving the evaluation of machine learning models for molecular simulations, where data are generated sequentially, and labels require expensive quantum chemistry computations, showing improved efficiency and reliable performance estimation under limited labelling budgets. In particular, on the real-world task, OAT requires only ~30% of labels to match the accuracy of a naïve uniform baseline, compared to ~23% for an oracle method that has access to the labels.
Abstract: The inference efficiency of diffusion large language models (dLLMs) is constrained by two challenges: bidirectional attention precludes efficient KV-cache reuse, while increasing decoding parallelism with static confidence thresholds can compromise generation quality. We observe that both challenges arise from a shared phenomenon: as tokens are decoded, their contextual integration through bidirectional attention causes token representations to drift (evolve) across decoding steps. This insight motivates Polestar, a training-free inference framework that uses token representation drift as a unified signal to jointly address both challenges. Polestar comprises two components: Polestar-Cache, which identifies stale KV-cache positions via drift and performs sparse KV-cache refreshes to enable efficient reuse, and Polestar-Commit, which detects sharp drift events to reliably identify commit-ready tokens. Across mathematics and coding benchmarks on several dLLM families, Polestar sets a new state of the art on the accuracy-throughput Pareto frontier, achieving up to 10.73% accuracy improvement, up to 3.7x higher throughput, and high decoding parallelism of 3.67 tokens per forward pass over existing baselines.
PaperID: 7593, Poster
Abstract: Graph domain adaptation (GDA) has emerged as an important problem in graph machine learning when the distribution of the source graph used for training differs from that of the target graph used for testing. While much of the prior work on GDA has focused on aligning node representations across source and target domains, recent studies show that such approaches can be suboptimal in the presence of graph structure shift, where the underlying connection patterns between nodes change across domains. In this work, we develop a unified pairwise distribution matching framework for mitigating conditional structure shift (CSS), a specific form of graph structure shift in which conditional edge distributions vary across domains. The framework recovers existing GDA methods as instances of moment matching and motivates PLSA, a new likelihood-based method that uses calibrated probabilistic predictors on connected target node pairs. Theoretically, we establish finite-sample guarantees for both the likelihood-based and moment-matching estimators under the contextual stochastic block model. Our analysis uses Bernstein-type concentration bounds for edge-weighted U-statistics, leading to error bounds that reflect both the effective number of observed edges and the conditioning of the corresponding likelihood or moment-matching problem. We complement our theoretical results with empirical studies that demonstrate the effectiveness of the proposed framework.
Abstract: LLM/VLM-based digital agents have advanced rapidly thanks to scalable sandboxes for coding, web navigation, and computer use, which provide rich interactive training grounds. In contrast, embodied agents still lack abundant, diverse, and automatically generated 3D environments for interactive learning. Existing embodied simulators rely on manually crafted scenes or procedural templates, while recent LLM-based 3D generation systems mainly produce static scenes rather than deployable environments with verifiable tasks and standard learning interfaces. We introduce SimWorld Studio, an open-source platform built on Unreal Engine 5 for generating evolving embodied learning environments. At its core is SimCoder, a tool/skill-augmented coding agent that writes and executes engine-level code to construct physically grounded 3D worlds from language/image instructions. SimCoder self-evolves by using verifier feedback (e.g., compilation errors, physics checks, VLM critiques) to revise environments and autonomously add reusable tools and skills to its library. Generated worlds are exported as Gym-style environments for embodied agent learning. SimWorld Studio further enables co-evolution between environment generation and embodied learning: agent performance feedback guides SimCoder to generate adaptive curricula near the learner’s capability frontier, so that environments become increasingly challenging as the embodied agent improves. Three case studies on embodied navigation show that self-evolution improves generation reliability, generated environments substantially improve embodied agent performance that generalizes to unseen benchmarks, and co-evolution yields an 18-point success-rate gain over fixed-environment learning and a 40-point gain over an untrained agent.
Abstract: We propose a deterministic adjoint matching framework that formulates human preference alignment for flow-based generative models as an optimal control problem over velocity fields. One can directly regress the control toward a value-gradient-induced target under the current policy, leading to a simple and stable training objective. Building on this perspective, we introduce a truncated adjoint scheme that focuses computation on the terminal portion of the trajectory, where reward-relevant signals concentrate, which yields substantial computational savings while preserving alignment quality. We further generalize the framework beyond standard KL-based regularization, allowing more flexible trade-offs between alignment strength and distributional preservation. Experiments on SiT-XL/2 and FLUX.2-Klein-4B demonstrate consistent gains across multiple alignment metrics, along with substantially improved diversity and mode preservation.
Authors: JingChuan Guan, Tomoyuki Kubota, Yasuo Kuniyoshi, Kohei Nakajima
Abstract: State space models have shown strong potential to outperform Transformers, but their learning mechanisms remain poorly understood, particularly the conditions for successful parameter updates that preserve teacher information. Prior work attributes the difficulty of such updates to the vanishing gradient during backpropagation through time, in which error signals decay exponentially through repeated gradient multiplication. In this work, by expressing error signals explicitly as functions of inputs, we show that the conventional vanishing gradient framework is insufficient to explain the loss of information. Using this representation of error, we reveal that the memory of the input encoded in teacher information does not decay exponentially, and we analytically demonstrate that the loss of supervisory information cannot be alleviated through training. This result theoretically confirms the importance of the initial weights of SSMs and suggests that the recurrent layer does not require training. By performing experiments on both linear SSMs and a representative nonlinear SSM, including language modeling tasks, we confirm that appropriate initialization enables us to fix the recurrence, leading to more stable training of the entire network and higher performance than the conventional training scheme. Our results partially elucidate the learning mechanisms of SSMs and are expected to contribute to the development of improved RNN-based models.
PaperID: 7597, Poster
Abstract: Data Shapley has emerged as a premier framework for data valuation, yet its practical utility is severely hampered by exponential computational complexity and the need for costly model retraining. Unlike existing methods that compute values independently for each dataset, we propose GenSHAP, a novel graph-based paradigm that reformulates data valuation as a transferable learning task. GenSHAP learns an amortised deep model capable of predicting Data Shapley values for any unseen classification dataset in a single forward pass. Our framework utilises a relation-based dataset representation method and a hybrid GIN-Transformer architecture to capture local and global inter-sample dependencies. We further introduce GenSHAP-S, which augments these structural representations with low-budget marginal contribution probes to achieve high-fidelity valuation. Extensive experiments demonstrate that GenSHAP closely approximates high-budget permutation-Shapley reference estimates and excels in downstream tasks such as data removal and addition while reducing valuation time from hours to seconds. The source code is available in the anonymous repository~\footnote\urlhttps://anonymous.4open.science/r/GenSHAP/.
PaperID: 7598, Poster
Abstract: Clinical Decision Making (CDM) requires integrating heterogeneous information across sequential diagnostic stages while interacting with patients whose behaviors are often non-stationary and unreliable. However, existing medical agents are typically trained against cooperative patient simulators and tackle each clinical task in isolation, making them fragile to dynamic patient behaviors and unable to capture inter-task dependencies essential for globally coherent decision-making. To address these issues, we propose CoE-Agent, a co-evolving patient–doctor agentic framework that learns globally coherent CDM strategies via interactive policy graph optimization. First, Dynamic Patient–Doctor Co-Evolutionary Learning pairs a Dynamic Patient Agent, which maintains an evolving cognition state to simulate non-stationary behaviors, with an Interactive Policy Graph-Based Doctor Agent that generates DAG-structured policy graphs unifying interactions, tool invocations, and reasoning steps. The doctor agent is then optimized via Interactive Policy Graph Optimization (IPGO), which jointly enforces structural validity and task-aligned correctness through reinforcement learning. Second, Cross-Task Policy Graph Consolidation and Unified Model Learning prunes low-value nodes, aggregates high-value nodes across tasks while filtering redundancy, and distills a unified policy planner with task-specific decision models for globally coherent decision-making. Experiments on MedChain and ClinicalBench show that CoE-Agent surpasses state-of-the-art baselines by 11.13% and 23.63% in average score, while achieving up to 19.5× faster inference. Source code is to be released.
PaperID: 7599, Poster
Abstract: Low-rank adaptation (LoRA) is the de facto method for parameter-efficient fine-tuning of large neural networks, achieving strong task adaptation with minimal trainable parameters and low optimization cost. Beyond this efficiency, practitioners have observed two striking properties: robustness to catastrophic forgetting and the ability to merge independently trained adapters into a single model that performs competitively across multiple tasks. These properties are central to continual and multi-task learning, yet remain poorly understood. In this work, we provide a theoretical explanation through a unified margin-based perspective. We analyze LoRA in multiclass linear classification under a near-orthogonal task regime with l_2-regularization, and characterize the optimal LoRA adapter across regularization regimes. Our analysis shows that, in an intermediate regime of regularization parameters, the optimal adapter aligns with the max-margin solution on the fine-tuning data. Building on this characterization, we derive two key consequences. First, we obtain closed-form expressions for the margins on pre-training and fine-tuning data, revealing a precise margin trade-off: the regularization parameter controls the balance between retention and adaptation. Second, we analyze adapter merging, proving that merged models achieve positive margin on each task and deriving optimal mixing coefficients that balance margins across tasks and maximize the margin over their union. These results lead to a simple, training-free merging rule, which we term margin-based LoRA merging (MaLoRA). Experiments on modern architectures and real datasets validate our theoretical predictions, showing that MaLoRA matches or outperforms several adapter-merging baselines across a range of vision and language classification tasks.
PaperID: 7600, Poster
Authors:
Hyeongwoo Nam, Woongje Cho, Juwon Kim, Jongeun ChoiAbstract: Large language model (LLM)-based robot task planning is promising for open-ended instruction following, but degrades on long-horizon tasks in large environments. When spatial information is conveyed to the LLM through text, the model can fail to capture spatial context, and token cost grows with environment size. Generating action sequences directly with an LLM also makes it difficult to satisfy the current world state and action preconditions. We address this with an ontology-grounded scene representation that aligns objects, spaces, relations, and states in a shared symbolic vocabulary for spatial reasoning and task planning, and with OntoPlan, an agentic framework that interprets instructions, selectively retrieves task-relevant information, formalizes goals and constraints, and produces executable plans. Across 150 benchmark tasks spanning five indoor environments and three scene scales, OntoPlan achieves 0.83 average task success, compared with 0.30 for the strongest baseline, while using 17.2k total tokens per task on average, about 5.8× fewer than the most efficient baseline. These advantages persist as scene scale increases, whereas prior methods degrade more sharply in both success and token cost. OntoPlan also responds appropriately to ambiguous or infeasible instructions by asking follow-up questions or reporting insufficient information rather than committing to invalid plans. Code available at https://anonymous.4open.science/r/OntoPlan.
PaperID: 7601, Poster
Authors: yang sheng, Jie Fu
Abstract: Circuit extraction identifies model components that preserve a target behavior under ablation, yet it remains unclear which parts of the reported circuit—components, edges, or coarser summaries—remain stable across reasonable extraction choices. We study this question in a Lean-style, rule-generated tactic-prediction benchmark spanning atomic and compositional proof tasks. The proof-state rules and task structure are fixed by construction, and dense and weight-sparse variants of one transformer architecture are compared. In this controlled setting, we compare three ways to describe a circuit: a compact prediction-preserving circuit, a broader graph that keeps read, write, and routing components around that circuit, and a pruning graph whose size is set by a target post-ablation loss budget. We vary model sparsity, whether the query and key sides of attention are represented together or separately, the post-ablation loss budget for the pruning graph, and which supervised checkpoint initializes reinforcement learning (RL). When the same extraction method is used, dense and weight-sparse checkpoints overlap far more in the set of attention heads identified than in exact component-to-component edge lists. This head-level overlap is well above a random top-k baseline; the exact-edge overlap is not consistently so. Tracking query and key sides separately, and varying the loss budget across the tested range, preserves the ordering of the RL initialization conditions when each graph is summarized by how many components it selects. Among these RL conditions, the largest gains on compositional tasks are accompanied by the largest fraction of compositional-task circuit nodes that lie outside the matched atomic-task circuits. These findings show that, even in this controlled setting, behavior preservation under ablation does not determine a unique circuit-level conclusion: the conclusion depends on which graph is reported, how it is extracted, and at which level of description the comparison is made.
PaperID: 7602, Poster
Abstract: Fine-tuning is widely used to align large language models (LLMs) with desired behavioral norms such as safety, yet it is typically evaluated using aggregate metrics that assume uniform improvements across the model. However, fine-tuning operates within a constrained parameter space, and its entangled effects on localized knowledge remain poorly understood. To address this, we investigate fine-tuning at the level of individual knowledge concepts using a controlled framework with representative concepts and paired safe/unsafe samples under varying data compositions. By jointly analyzing behavioral safety via model outputs and internal knowledge accessibility via perplexity, we obtain a unified view of how knowledge is modified and expressed. Our results reveal that fine-tuning does not uniformly improve safety. Instead, it (1) exhibits strong concept-level heterogeneity, including persistent low-safety concepts and robustly safe concepts with clear domain specificity; (2) improves safe–unsafe discrimination on average, but coexists with familiarity on both safe and unsafe content or even loses previously learned safe knowledge; and (3) decouples internal knowledge from external behavior, as changes in familiarity do not reliably translate into safer outputs, especially due to early-stage disruption of existing safety structures. Our findings position consistency as a distinct objective beyond conventional aggregate safety and demonstrate that reducing unsafe knowledge exposure under uniform concept-level coverage actively homogenizes safety distributions across model knowledge, offering a principled path toward alignment that is not merely effective on average, but uniformly reliable, predictable, and robust across knowledge concepts.
Abstract: Latent diffusion models offer an attractive alternative to discrete diffusion for non-autoregressive text generation by operating on continuous text representations and denoising entire sequences in parallel. The major challenge in latent diffusion modeling is constructing a suitable latent space. In this work, we present the Latent Diffusion Language Model (LDLM), in which the latent encoder, diffusion model, and decoder are trained jointly. LDLM builds its latent space by reshaping the representations of a pre-trained language model with a trainable encoder, yielding latents that are easy to both denoise and decode into tokens. We show that naive joint training produces a low-quality diffusion model, and propose a simple training recipe consisting of an MSE decoder loss, diffusion-to-encoder warmup, adaptive timestep sampling, and decoder-input noise. Ablations show that each component substantially impacts generation performance. On OpenWebText and LM1B, LDLM achieves better generation performance than existing discrete and continuous diffusion language models while being 2\text -13× faster, indicating that jointly learning the latent space is a key step toward making latent diffusion competitive for text generation.
Authors: Leonel Aguilar, Jan Nagler, Christoph Hoelscher, Nino Antulov-Fantulin
Abstract: Successful deep neural networks discover salient features of data. We show when and why they fail to learn out-of-distribution (OOD)-relevant representations from an in-distribution (ID) training window. This requires decoupling feature learning from data-generating-process (DGP) identifiability. From a single training window, OOD extrapolation is non-identifiable: infinitely many DGPs are \varepsilon-observationally equivalent on the training data but diverge arbitrarily outside it, and no in-distribution criterion alone reliably breaks the tie. A structural commitment, the feature map, label map, and model class (\varphi, \psi, \mathcalM), dictates the assumed DGP and governs OOD generalization while leaving ID performance essentially unchanged. When architecture, pretraining, augmentation, input formats, or domain knowledge implicitly inject the missing commitment, the model succeeds. When it cannot infer OOD-relevant structure from ID evidence, it fails. Changing only the representation can make the same architecture, at the same in-distribution loss, differ by ~520× out of distribution. When the commitment is correct \em and identifiable, OOD error vanishes. For example, Fourier coordinates turn periodic extrapolation into interpolation on \mathbbS^1. The same mechanism predicts outcomes in three natural-science settings (mass-action chemistry; Kepler's-third-law exoplanet prediction, n=2,362; and cross-species coding-DNA detection) and in a 264-run positional-encoding study across Transformer, Mamba, and S4D. Finally, a controlled study shows: correct features are necessary but not sufficient. The model class must express the target, and the transformed training data must cover the relevant representation space. Thus, feature engineering is not a departure from what makes deep learning successful. It is the explicit structural commitment that makes extrapolation identifiable.
PaperID: 7605, Poster
Abstract: The Model-as-a-Service (MaaS) paradigm enables resource-constrained edge users to access cloud-based large language model (LLM) services but raises significant privacy concerns. Existing approaches perturb user queries under local differential privacy (LDP) before sending them to LLMs. However, naively applying these methods to multi-query scenarios leads to linear growth in privacy budget consumption. To address this, we propose LDPCache, the first cache-enhanced LDP framework for multi-query LLM processing. By leveraging semantic correlations across consecutive user queries, LDPCache maintains a cache of historical responses and intelligently decides whether to reuse a cached result or query the cloud-based LLM, thereby reducing privacy budget consumption. The core of LDPCache is an LDP-compliant scheduler that evaluates cache reusability based on both semantic similarity and the expected accuracy gains from LLM access. This scheduler enables hit decisions without actual LLM access while ensuring that the scheduling process itself satisfies LDP guarantees. To accommodate the limited memory of edge devices, we further design a multi-factor cache eviction strategy that balances query diversity, response accuracy, and recency to improve the hit ratio. Extensive experiments on real-world datasets demonstrate that LDPCache significantly improves response accuracy over baseline methods.
PaperID: 7606, Poster
Abstract: Visual Foundation Models (VFMs) have demonstrated strong capabilities in visual-geometric tasks such as image matching and 3D reconstruction. Recent work has revealed the potential of VFMs in label-free indoor RGB-D registration, where the visual modality is intrinsic to the task. A natural question follows: can this success extend to more challenging outdoor scenarios characterized by sparse, irregular LiDAR data with only partial visual coverage? We answer this with PROMOTER, a novel teacher–student framework for unsupervised LiDAR registration that bridges visual priors and sparse LiDAR geometry. To leverage pre-trained geometric priors for 3D registration, we must disentangle them from confounding signals introduced by modality and task gaps under partial visual coverage. To this end, we introduce a lightweight Adaptive Prior Integrator that extracts and integrates these priors into an enhanced teacher, together with a Perceptual-Geometric Labeler that produces high-quality pseudo-labels. For student training, we propose Matchability-Anchored Training to improve robustness under noisy supervision. Experiments on KITTI and nuScenes demonstrate state-of-the-art performance and improved scalability across registration models and VFMs, while incurring no additional parameters at inference. Code will be released.
PaperID: 7607, Poster
Authors: Isaiah Freeman, Joed Ngangmeni
Abstract: Representational Engineering (RepE) aims to expose and steer high-level concept directions in language model activations, and recent work shows that hidden states can be transferred across heterogeneous architectures via trained neural adapters. Whether specific concepts survive such transfer remains an open question. We study one such concept, truth, by mapping source-model activations through autoencoder adapters into target-model representation space, varying injection strength \alpha \in [0, 1], and reading the result with the target's native truth probe across 21 cross-family trajectories from four model families, with comparison against a closed-form rigid alignment baseline on a representative subset. We find that ordinal truth ranking is preserved through partial injection (median AUROC near the native ceiling through \alpha \in [0, 0.7]) but collapses sharply at full injection (\alpha=1), with parametric separability dropping ~14× while rank order partially survives. The apparent accuracy overshoot under partial injection decomposes cleanly into two effects: a large recalibration component dominated by native probe miscalibration (R^2 = 0.75), and a small but universal genuine separability gain present in all 21 trajectories. These findings characterize cross-family truth structure as partially compatible, ordinally stable, but non-isomorphic: preserved enough under partial injection that learned and rigid alignment methods both surface meaningful truth signal, but with full replacement representing a hard failure mode that practitioners should treat as a deployment boundary.
PaperID: 7608, Poster
Abstract: We present GAMMA, a scalable feed-forward model that reconstructs 4D Gaussian Splatting from monocular videos, enabling dynamic novel view synthesis in real-time. Most existing approaches for dynamic scene modeling require multi-view videos as input or costly per-scene optimization. In contrast, GAMMA learns a unified 4D scene representation from monocular videos and generatively predicts 4D Gaussian primitives within seconds. The key insight of GAMMA lies in its unified Gaussian reconstruction model, which jointly estimates temporally consistent geometry, appearance, and motion from the input video. We further explore its scalability through large-scale training on a comprehensive dataset encompassing both synthetic and real-world videos. The reconstructed GAMMA representation supports interactive scene exploration via real-time rendering across timesteps and local viewpoints. Extensive experiments demonstrate that GAMMA outperforms existing methods in both reconstruction fidelity and efficiency.
PaperID: 7609, Poster
Abstract: Regular time-series generation from irregular time series is typically reduced to one-shot imputation: complete each partial sequence once, freeze the completions, and train a generator on the resulting surrogate dataset. We show that this open-loop strategy is structurally flawed. Point imputers collapse ambiguous missing regions to conditional averages, while stochastic imputers become stale as the generator evolves; the binding constraint is therefore iteration, not imputer design. We introduce a co-evolving Monte Carlo EM framework that closes the imputer--generator loop. Each E-step samples missing values from the posterior of the current generator, and each M-step retrains the generator on these completions, preserving unconditional generator training. Since our diffusion prior operates in a lifted 2D representation while observations live in time-series space, we further introduce Posterior Sampling for Lifted Representations (PSLR), which enforces operator-consistent conditioning, uncertainty preservation, and manifold consistency during the E-step. Across nine datasets and four corruption regimes, our method achieves state-of-the-art performance, improving over the strongest baselines by 59% in discriminative score, 15% in predictive score, 73% in Context-FID, and 71% in feature-correlation error. Ablations show that our approach reduces average discriminative score by 73% over single-step baselines, closing most of the remaining gap to the clean-data oracle.
Abstract: We study classification and regression tasks under label differential privacy (label-DP), a setting where only the labels are considered sensitive information. We quantify the "price of locality" by revealing fundamental statistical gaps between local and central privacy models. For classification, we establish a strict separation in accuracy. Specifically, we prove that the minimax excess risk in the local model is lower bounded by \Omega(\frac1\epsilon\sqrt\textVC(\mathcalH)/n) for classification on size-n dataset, a fundamental limitation that holds even under relaxed approximate (\epsilon, \delta)-local DP. Overcoming this barrier, we propose the first efficient central-model algorithm that operates under the stricter pure \epsilon-DP, yet achieves a significantly improved upper bound of \tildeO(\sqrt\textVC(\mathcalH)/(\epsilon n)). For linear regression, we show that local randomizers or naive central-DP approach algorithms are statistically inefficient. Instead, we turn to the Matrix Mechanism and use an encoder-decoder framework built on ridge regression, which projects the data into a spherical latent space before injecting noise. This structural design effectively isolates and exponentially suppresses the severe noise inflation typically caused by ill-conditioned feature covariance matrices in full-DP (Cai et al., The Annals of Statistics, 49(5)), achieving an optimal privacy penalty of \tildeO(d^2/(n^2\epsilon^2)) where d is the feature dimension.
Abstract: Reinforcement learning (RL) has emerged as a paradigm for fine-tuning large-scale generative models, such as diffusion and flow models, to align with complex human preferences and user-specified tasks. A fundamental limitation remains the curse of diversity collapse, where the objective formulation and optimization landscape inherently collapse the policy to a Dirac delta distribution. To address this challenge, we propose DRIFT (DiveRsity-Incentivized Reinforcement Fine-Tuning for Versatile Image Generation), an innovative framework that systematically incentivizes output diversity throughout the on-policy fine-tuning process, reconciling strong task alignment with high generation diversity to enhance versatility essential for applications that demand diverse candidate generations. We approach the problem across three representative perspectives: i) sampling a reward-concentrated subset that filters out reward outliers to prevent premature collapse; ii) prompting with stochastic variations to expand the conditioning space, and iii) optimization of the intra-group diversity with a potential-based reward shaping mechanism. Experimental results show that DRIFT exhibits clear Pareto dominance in task alignment and generation diversity, achieving 7.19%~93.40% higher diversity at matched alignment and 13.23%~60.13% higher alignment at matched diversity.
PaperID: 7612, Poster
Abstract: Automated Heuristic Design (AHD) with Large Language Models (LLMs) has shown promise in solving optimization tasks. However, existing methods mainly rely on iterative search frameworks and do not systematically structure or reuse task-specific design knowledge elicited from LLM priors and accumulated during online search. Moreover, multi-objective AHD introduces additional challenges, as conflicting objectives such as solution quality and computational efficiency must be balanced during heuristic search. We propose Multi-objective Evolution of Heuristics with Insight-Driven Search (MEoH-IDS), which uses dynamically maintained design insights to guide LLM-based heuristic generation. MEoH-IDS maintains an insight pool by extracting insights from successful heuristics, evaluating their search utility, and filtering less effective ones over time. It selects promising insight combinations using a UCB-inspired policy, with rewards based on non-dominated status in the current Pareto archive. We validate MEoH-IDS on three AHD tasks, including the Traveling Salesman Problem (TSP), Capacitated Vehicle Routing Problem (CVRP), and Vehicle Routing Problem with Time Windows (VRPTW), optimizing both solution quality and computational efficiency. Experimental results show that MEoH-IDS accelerates convergence and discovers heuristics with improved quality--efficiency trade-offs compared with representative LLM-based AHD baselines.
PaperID: 7613, Poster
Authors: Feiran Li, Jiahao Lyu, Ke Jiang, Yu Zhou, Jinqiao Shi
Abstract: Scene text encodes high-density semantic information, raising privacy risks for automated vision systems. Existing privacy-preserving adversarial methods mainly target Scene Text Recognition (STR) by inducing character-level misrecognition, yet text regions remain detectable, enabling manual inspection or downstream recovery. We argue that stronger privacy requires suppressing Scene Text Detection (STD) during localization, hiding text as background to bypass subsequent OCR pipelines. To this end, we propose AR-Attacker (Attention-guided Reinforcement learning text Attacker), a black-box attack framework that generates sparse, visually imperceptible perturbations to effectively suppress STD outputs. We formulate the attack as a Markov Decision Process (MDP) and design an attention-guided action space, which leverages a frozen surrogate model to focus perturbations on salient text regions. Furthermore, AR-Attacker employs a discrete-continuous hybrid policy optimized with Proximal Policy Optimization (PPO), which jointly selects perturbation locations and magnitudes. Empirical results show that AR-Attacker surpasses state-of-the-art baselines by 30% in text removal rate while simultaneously reducing query counts by 40%. Remarkably, the framework maintains high visual fidelity (PSNR > 45 dB), advancing the "visible to humans, invisible to machines" privacy paradigm through a robust RL-based approach.
Abstract: Large Language Models (LLMs) have achieved remarkable progress on complex reasoning tasks, yet the mechanisms by which they acquire reasoning capabilities and the efficiency of their reasoning remain poorly understood. In this paper, we investigate the path-finding problem over directed graphs—a symbolic abstraction of multi-step reasoning. We first characterize the architectural requirements for this task, proving that one-layer transformers require \Omega(N^2) size on dense graphs to implement the key DFS child-selection primitive. While a two-layer transformer implements full DFS with O(N\log(N)) size. Thus, a second layer is necessary and sufficient for near-linear DFS reasoning. We then extend this expressivity analysis to deeper models, proving the existence of an L-layer transformer that performs DFS with a look-ahead horizon of 2^L-3, establishing that increasing depth yields an exponential gain in search efficiency. Finally, going beyond the expressivity, given appropriately curated training data, we show that two-layer transformers can provably learn to execute DFS via gradient flow and generalize to unseen graphs. Together, our results provide a mechanistic account of how transformers implement graph search, how depth improves reasoning efficiency, and how such reasoning algorithms can emerge through training.
Abstract: Vision-language models such as CLIP often struggle to faithfully understand long, detail-rich captions, relying on dominant scene cues while overlooking fine-grained visual evidence. We propose a hierarchical vision-language learning principle for understanding scenes as part-to-whole compositions: before forming a whole-scene representation, a model should uncover what semantic parts appear where in the image. To this end, we propose CAFT (Cross-domain Alignment of Forests and Trees), a vision-language model that jointly learns local text-region alignment at intermediate representations and global image-text alignment at the final representation. Exploiting the organization of long captions, where local descriptions often correspond to scene parts, CAFT employs a fine-to-coarse image encoder and a part-whole text encoder to discover localized part semantics and progressively compose them into a global image-text representation. Trained on 30M image-text pairs, CAFT achieves state-of-the-art performance on six long-text retrieval benchmarks and exhibits strong scaling behavior. Experiments show that CAFT learns fine-grained representations that localize textual semantics in image regions without explicit region-level supervision.
PaperID: 7616, Poster
Authors:
Yang Gao, Yingjing Xiao, Junbin Ren, Chenxu Zhang, Wenbo Zhang, Zhanpeng JinAbstract: We present a Neuro-inspired Structured Model (NiSM) for fine-grained finger-level contact force estimation during dexterous grasping from sparse wearable sensing. NiSM uses a thumb-mounted inertial measurement unit (IMU) and a single-channel wrist electromyography (EMG) sensor, while deriving grasp context from an IMU-based hand-shape latent. This work studies the trade-off between sensing density and structural inductive bias: when sensor coverage is limited, inferred grasp context, kinematic coupling, and muscle-effort modulation can provide useful constraints for force prediction. Inspired by human sensorimotor control, NiSM organizes computation into complementary pathways: a context-conditioned feedforward pathway maps thumb kinematics to finger-specific estimates through sparse learned gating, while an EMG-conditioned modulation pathway provides effort-dependent scaling. By explicitly encoding these structural priors, NiSM reduces reliance on dense sensor arrays and large generic black-box models. Evaluations on multi-user, multi-object datasets show that NiSM achieves performance competitive with a dense-sensing reference where comparable dense inputs are available, while using substantially fewer parameters, and is more robust than generic models under limited-data settings. These results indicate that structured model design can recover part of the information typically supplied by denser sensing, offering a practical direction for wearable finger force estimation under real-world deployment constraints.
Authors:
Tianqi Shen, Jinji Yang, Runze SHI, Jianhao Ma, Jiaye Teng, Ziye MaAbstract: Recently, Muon has gained substantial attention as an appealing alternative to Adam, with many works highlighting its advantages through spectral normalization and improved conditioning. Yet this positive theoretical narrative contrasts with its empirical performance in large language model (LLM) training, where Muon’s gains over Adam/AdamW are often mixed, schedule-sensitive, and not uniformly superior. To address this gap, we develop a trajectory-level theory characterizing both the strengths and limitations of Muon. We introduce a mixed-spiked matrix sensing model whose sensing operator decomposes into signal, spike, and bulk components, capturing a mixture of anisotropic structure and long-tail information reminiscent of LLM training. On top of it, we adopted a river-valley perspective in which we view the landscape as composed of a river direction flowing to the desired solution and hill directions encoding nuisance or task-irrelevant information. In the momentum-free setting, we show that Muon moves faster along the information-bearing river direction during early optimization, but can converge much more slowly near the river bottom than gradient descent. We then extend the river-valley perspective to general nonconvex objectives with momentum by studying points on the spectral river. There, while Muon converges faster early on, its orthogonalized update removes residual scale information, making it prone to overshooting and oscillation near the target solution. Together, these results suggest that our characterizations extend beyond spiked matrix sensing and motivate switching to GD-like refinement optimizers in the final phase, rather than relying only on a fixed learning-rate schedule for Muon. We also provide preliminary evidence supporting this two-stage approach in language model training experiments.
PaperID: 7618, Poster
Authors:
mengqi Han, Bo Yang, Qi Liu, Mengshuo Jia, Chen Gong, Liu Yuxiang, Sicheng Liu, Mingxuan CaiAbstract: Large Language Models (LLMs) present a novel paradigm for power system dispatch (PSD), enabling operators to derive optimal dispatch strategies from multi-objective requirements and grid dynamics. However, isolated paradigms focusing on modeling or solving often yield decision bias or physical infeasibility. Meanwhile, end-to-end paradigms couple modeling-solving with reflection-feedback loops, obscuring whether failures stem from inaccurate modeling, invalid solutions, or insufficient feedback, thereby causing ineffective iterations. To address these challenges, we propose Agile Dispatch Agent (ADA), an agent based on two-time-scale evolution that decouples the optimization process into two asynchronous loops. On the fast time scale, the Planner actively acquires information to clarify complex requirements, while the Solver selects algorithms to adapt to the problem mathematical characteristics. On the slow time scale, the Summarizer module updates the modeling feasible region based on evaluations from the Judger. Under a finite-horizon Planner and standard two-time-scale stochastic approximation assumptions, our analysis provides local tracking guarantees for the fast decision process and the slow knowledge evolution. Experiments on L2RPN benchmarks with ambiguous requirements and renewable dynamics show that ADA outperforms state-of-the-art baselines.
PaperID: 7619, Poster
Abstract: Multimodal discrete diffusion language models (dLLMs) generate responses through iterative mask prediction and offer a promising alternative to autoregressive large vision-language models. However, they still suffer from hallucination, producing outputs that are plausible under language priors but insufficiently grounded in the image. We argue that hallucination in multimodal dLLMs is not only a token prediction problem, but also a premature commitment problem: during iterative denoising, some masked positions are unmasked too early because their confidence is driven more by language priors than by visual evidence, and these early commitments can bias later generation. To address this issue, we propose Visual Gain Ordering (VIGOR), a training-free and plug-and-play inference strategy that reorders masked positions according to counterfactual visual evidence. At each denoising step, VIGOR compares the confidence of each candidate token under the original image-conditioned pass and an additional visual-ablation pass, and prioritizes positions whose predictions are more strongly supported by the image. Experiments on LLaDA-V and Lumina-DiMOO show that VIGOR consistently reduces hallucination and improves multimodal reasoning performance across benchmarks. Anonymous code is available at https://anonymous.4open.science/r/519_VIGOR-4051.
Abstract: The recent surge of interest in Time Series Foundation Models has rapidly advanced the field. Yet, heterogeneous training setups across studies make it difficult to determine whether performance improvements came from architectural innovations or effective data engineering. In this work, we investigate whether a standard patch Transformer architecture can achieve state-of-the-art zero-shot forecasting performance when paired with a straightforward training protocol. We conduct a comprehensive ablation study covering model scaling, data composition, and training techniques to isolate the essential ingredients for high performance. Our findings identify the key drivers of forecasting accuracy, confirming that the generic architecture itself can deliver high performance with good scalability. By strictly controlling these experimental variables, we provide comprehensive empirical results on model scaling across multiple dimensions. Furthermore, we establish that the effectiveness of this training protocol generalizes robustly across benchmarks and alternative architectures. We release our open-source model and detailed findings to establish a transparent, reproducible baseline for future research.
PaperID: 7621, Poster
Abstract: Black-box distillation is a practical route for transferring capabilities from API-accessible large language models that expose only text outputs into smaller student models. Recent on-policy adversarial methods such as GAD improve over SeqKD by forming an adversarial loop between a critic and a student, where the critic provides rewards for GRPO-based student policy optimization over the student's sampled responses. However, GRPO computes advantages from the within-group relative rewards of student samples for the same prompt, whereas the critic is trained primarily to distinguish teacher responses from student responses. This objective mismatch can produce reward groups with collapsed scale or fragile margins, leading to brittle grouped optimization signals. We propose Groupwise Reward Geometry Conditioning (GRGC), a two-stage framework that improves advantage construction by shaping student-side reward groups during both critic training and policy optimization. To improve critic-side conditioning, Gaussian groupwise Optimal Transport calibration regularizes the critic during training to produce reward groups with non-collapsed spread and smooth rank-wise gaps by matching sorted prompt-wise rewards to group-centered Gaussian quantiles. Building on this conditioned reward geometry, policy-side group power modulation reshapes the prompt-wise reward groups before they are converted into advantages, preserving the critic-induced ordering while increasing optimization-relevant margin separability. Extensive experiments across diverse teachers, student model families and scales, and training datasets demonstrate the effectiveness of GRGC on both in-distribution and out-of-distribution evaluations, while introducing negligible overhead over GAD. The code is available at https://anonymous.4open.science/r/GRGC.
PaperID: 7622, Poster
Abstract: We consider linear equations with min and max operators~(LEMMs) that contain many subproblems ranging from optimization, learning, to games. Recently, Chatterjee et al. [2025] gave a systematic study of the complexity of different subclasses. Three key subclasses--(I) halting branching process, (II) absolutely halting LEMMs, and (III) halting LEMMs--are proved to be in UP \cap coUP while generalizing stochastic games (SSGs). In this work, we study the classic algorithms of Policy Iteration (PI) and Value Iteration (VI) for these general subclasses. First, we simplify the problem hierarchy by showing the equivalence between halting branching process and SSGs. Then, we show that while PI diverges for absolutely halting LEMMs due to the loss of monotonicity, VI remains convergent, a result we establish via diagonal rescaling. Finally, we show that neither PI nor VI converges for general halting LEMMs and, to this end, propose variants of simple policy iteration that ensure convergence across all subclasses.
PaperID: 7623, Poster
Abstract: Passive matching networks are critical components in radio-frequency integrated circuits (RFICs), but their design is bottlenecked by expensive finite-difference-based electromagnetic (EM) simulation. We propose an end-to-end neural framework for free-form RFIC passives that predicts both EM response and downstream design performance under varying frequency and circuit contexts, combining a frequency-conditioned neural simulator with a circuit-conditioned neural solver. To improve design performance prediction, the simulator outputs a neural EM response consisting of both the explicit EM response and a complementary response, together with a dependence regularizer that encourages the two responses to capture orthogonal information. To evaluate the proposed framework, we construct a large-scale dataset of free-form passive layouts, synthesized from both expert-designed contours and random paths. Our 21M-parameter model achieves high accuracy across a broad frequency range, circuit contexts and passive matching settings, with R^2>0.98 for almost all targets, while providing orders-of-magnitude faster evaluation than traditional EM simulation.
Abstract: Reinforcement learning (RL) effectively optimizes Large Language Model (LLM)-based recommenders by contrasting positive and negative items. Empirically, training with beam-search negatives consistently outperforms random negatives, yet the mechanism is not well understood. We address this gap by analyzing the induced optimization objective and show that: (i) Under binary reward feedback, optimizing LLM recommenders with Group Relative Policy Optimization (GRPO) is theoretically equivalent to maximizing the Area Under the ROC Curve (AUC), which is often misaligned with Top-K recommendation; and (ii) Replacing random negatives with beam-search negatives reshapes the objective toward partial AUC, improving alignment with Top-K metrics. Motivated by this perspective, we introduce Windowed Partial AUC (WPAUC), which constrains the false positive rate (FPR) to a window [\alpha,\alpha+d] to more directly align with Top-K metrics. We further propose an efficient Threshold-Adjusted Windowed reweighting (TAWin) RL method for its optimization, enabling explicit control over the targeted Top-K performance. Experiments on four real-world datasets validate the theory and deliver consistent state-of-the-art performance.
PaperID: 7625, Poster
Authors:
Milan Bhan, Jean-Noël Vittaut, Nicolas CHESNEAU, Sarath Chandar, Marie-Jeanne LesotAbstract: Large Language Models (LLMs) can generate plausible free text self-explanations to justify their answers. However, these natural language explanations may not accurately reflect the model's actual reasoning process, indicating a lack of faithfulness. Existing faithfulness evaluation methods rely primarily on behavioral tests or computational block analysis without examining the semantic content of internal neural representations. This paper proposes NeuroFaith, a flexible framework that measures the faithfulness of LLM free text self-explanation by identifying key concepts within explanations and mechanistically testing whether these concepts actually influence the model's predictions. We show the versatility of NeuroFaith across 2-hop reasoning and classification tasks. Additionally, we develop a linear faithfulness probe based on NeuroFaith to detect unfaithful self-explanations from representation space and improve faithfulness through steering. NeuroFaith provides a principled approach to evaluating and enhancing the faithfulness of LLM free text self-explanations, addressing critical needs for trustworthy AI systems.
PaperID: 7626, Poster
Abstract: We study the minimax regret of full-information online decision making in extensive-form games. For prediction with expert advice, the optimal regret is \Theta(\sqrtT\log K), where K is the number of experts, or equivalently, the number of normal-form actions. For treeplex strategy spaces, existing efficient algorithms, including KOMWU and online mirror descent with the weight-one dilated entropy regularizer, achieve regret \mathcalO(\sqrtT\log|\mathcalV|), where \mathcalV is the set of reduced normal-form strategies. We prove that this logarithmic dependence is minimax optimal in the worst-case full-information reward model. The main difficulty is that pure plans in a treeplex are not independent experts: their rewards are correlated through the recursive structure of decisions and observations. To capture this correlation, we introduce tropical Gaussian formulae, recursive Gaussian expressions built from maxima and normalized sums which preserve the variance. The index count of such a formula matches the number of reduced normal-form continuation plans in the corresponding subtree. Our main analytic result shows that every log-balanced tropical Gaussian formula with index count N has Gaussian width \Omega(\sqrt\log N). While direct approach struggles to recover this lower bound, our proof uses a spherical entropy profile to measure scale-dependent separation among the Gaussian directions, and then applies Tur\'an's theorem and Sudakov minoration to obtain the width lower bound. Applying this result to treeplexes gives an oblivious hard distribution with valid transition outcome and Rademacher terminal rewards, establishing a regret lower bound of \Omega(\sqrtT\log|\mathcalV|). This matches known upper bounds up to universal constants and shows that the DilEnt/KOMWU dependence on the reduced normal-form complexity is unavoidable.
PaperID: 7627, Poster
Abstract: Humanoid whole-body loco-manipulation requires scalable and physically reliable demonstrations, yet existing human-object interaction datasets lack direct contact supervision and egocentric observability, often resulting in physically inconsistent interactions. We present DeXHOI, a multimodal egocentric HOI benchmark for whole-body dexterous interaction, capturing synchronized body motion, object trajectories, tactile sensing, and egocentric RGB-D across diverse spatial layouts and long-horizon activities. To improve physical consistency, we introduce a tactile-guided refinement pipeline that leverages measured pressure to correct hand-object interactions while preserving global motion fidelity. We further establish an Egocentric HOI Recovery Benchmark for contact-aware reconstruction from first-person observations, along with EgoTacti, a tactile prediction model. Together, EgoTacti provides a scalable and physically grounded foundation for learning humanoid whole-body dexterous interaction, pointing toward low-cost egocentric data collection for embodied intelligence.
Abstract: Activation steering is a practical post-training model alignment technique to enhance the utility of Large Language Models (LLMs). Prior to deploying a model as a service, developers can steer a pre-trained model toward specific behavioral objectives, such as better truthfulness, or reasoning ability, without the need for retraining. Conceptually, these methods implement behavior control through hidden-state interventions, without changing the underlying model parameters. However, this capability unintentionally introduces critical and under-explored safety risks. We identify a phenomenon termed , where steering vectors derived from benign datasets—such as reducing harmless refusals, improving structured-output following, truthfulness, and reasoning performance—inadvertently erode safety guardrails. Experiments reveal that these interventions act as a force multiplier, creating new vulnerabilities to jailbreaks and increasing attack success rates to over 80% on standard benchmarks by bypassing the initial safety alignment. Ultimately, our results expose a critical blind spot in deployment: benign activation steering can erode the ``safety margin,'' rendering models more vulnerable to black-box attacks and indicating that inference-time utility improvements must be rigorously audited for unintended safety externalities.
PaperID: 7629, Poster
Authors: Mingqi Yang, Yanming Shen
Abstract: Large-graph spectral learning promises global, trainable propagation beyond local neighborhoods, but explicit spectral layers remain difficult to use in stochastic large-graph training: their bases are commonly obtained by a full-graph spectralization step and their graph-wide contexts are naturally full-batch. Matrix-free polynomial filters avoid eigensolvers and train stochastically, but tie the response to finite propagation bases; explicit spectral-subspace models offer more flexible global responses, but usually pay for full-graph basis construction and context evaluation. We introduce Randomized Adaptive Spectral Estimation (RASE), a matrix-free stochastic training framework for explicit low-rank spectral-context layers. RASE constructs a refreshable randomized block-Krylov/Ritz basis from probe vectors and sparse matrix--vector products, reuses probe-derived summaries as control variates for mini-batch estimates of the spectral context, and keeps the response module replaceable across polynomial, Chebyshev, and transformer-style instantiations. We analyze Krylov/Ritz polynomial-action exactness for the probe-generated subspace and a control-variate variance identity that isolates when the mini-batch context estimator reduces variance. Across 14 node-classification benchmarks, RASE matches its eigenbasis-trained reference within seed noise where the eigensolver is feasible (~67× faster basis on PubMed), and trains the same backbone on the largest graphs where eigensolver preprocessing exceeds a 30-minute budget. The contribution is therefore not a new spectral filter family, but a training route that makes explicit low-rank spectral-context layers matrix-free and mini-batch trainable.
PaperID: 7630, Poster
Abstract: Structure-based drug design increasingly employs LLM agents to iteratively refine ligands against a target pocket, yet a viable ligand must satisfy two often-conflicting objectives---binding affinity and druggability---which single optimization steps rarely improve together. To quantify this difficulty, we introduce two diagnostic metrics: the first measures how often a single edit improves both objectives, and the second measures how often a gain on one objective comes with a loss on the other. Applying these diagnostics to current LLM-agent pipelines exposes a consistent failure mode: the agent performs molecular editing without knowing how the pocket-ligand complex responds to local modifications, thus rarely achieving joint improvement. Inspired by medicinal chemists, who probe the pocket-ligand complex with controlled analog edits before choosing an optimization direction, we propose PROBE, an optimization framework built around edit--response probing. PROBE first decomposes the ligand into editable sites and builds a pocket-specific site map that flags where joint gains are plausible, where the two objectives are likely in tension, and where liability substructures should be changed; it then performs controlled probe edits whose responses are distilled into an EditManual. Guided by the site map and EditManual, PROBE runs an iterative multi-agent loop in which an affinity agent, a druggability agent, and a co-optimization agent jointly produce edits. On the CrossDocked2020 benchmark, PROBE achieves state-of-the-art performance and substantially mitigates the failure modes exposed by our diagnostics metrics.
Authors:
Sining Ang, Yuguang Yang, Chenxu Dang, Canyu Chen, Cheng Chi, Liu Haiyan, Xuanyao Mao, jason bao, Xuliang, Bingchuan Sun, Yan WangAbstract: Vision-Language-Action (VLA) driving augments end-to-end (E2E) planning with language-enabled visual backbones, but how VLM representations differ from vision-only encoders after policy learning, and whether such differences matter for planning, remains unclear. Under a unified VLM-hidden + diffusion-policy paradigm, we compare multiple VLM families/scales (InternVL3 and Qwen3VL) with standard vision-only encoders (ResNet, ViT, and EVA-CLIP) while keeping the downstream planner fixed. We study representation, behavior, and system design. CKA/CCA and Shared--Unique SAE show that policy learning enlarges a common decision subspace, but both branches retain non-transferable residual factors. Latent-intervention policies and scenario-level analysis further show that these residuals are behaviorally meaningful: vision-only encoders are stronger in simple geometry-dominant scenes, whereas VLMs are more effective in semantically complex and interaction-heavy long-tail cases. The two branches also exhibit distinct progress--braking and path-choice tendencies, and an oracle best-of-two VLM+ViT selector reaches 93.58 PDMS on NAVSIM. We convert this complementarity into two lightweight systems: HybridDriveVLA, which selects from a compact cross-model candidate set using a learned trajectory scorer and improves PDMS from 90.80 to 92.10, and DualDriveVLA, a fast--slow variant that invokes the VLM in only 15% of scenarios, achieving 91.00 PDMS with about 1.9× lower latency than the VLM baseline. Code will be released.
PaperID: 7632, Poster
Abstract: We study the implicit bias of Riemannian gradient flow for hyperbolic multiclass classification with fixed class prototypes in hyperbolic space \mathbb H^n. Our framework accommodates general permutation invariant relative margin (PERM) losses, a class that includes cross entropy and other standard multiclass losses. Our analysis is based on a decomposition: at large radius, the distance to each prototype splits into a radial term and a direction dependent term described by the Busemann function. This yields two main results. First, we prove a radial dichotomy: the sign of a drift coefficient \mu determines whether the radius is pushed toward the ideal boundary or back toward the interior; if the positive drift persists, then r(t)=\tfrac12\log t+O(1), while persistent negative drift returns the trajectory to the large radius threshold in finite time. Second, we show that the boundary direction converges to a critical point of the Busemann risk on \partial\mathbb H^n. These results provide a rigorous asymptotic perspective on two phenomena we refer to as boundary saturation and near-boundary clustering in hyperbolic representation learning.
PaperID: 7633, Poster
Authors: Sanjit Dandapanthula, Nicholas Boffi
Abstract: Reward guidance algorithms steer a learned generative process toward the reward-tilted measure at inference time. While empirically powerful, these methods are prone to reward hacking: the guided model over-optimizes the reward at the cost of fidelity to the learned distribution. Prior work has attributed this to the complexity of neural reward functions or implicit biases in diffusion training, but its fundamental origins remain poorly understood. We show that reward hacking arises from an approximation made in most practical implementations of reward-guided diffusion---finite-particle plug-in estimation of the Doob h-function---even in the simplest non-trivial settings of Gaussian and Gaussian mixture targets with quadratic rewards. In closed form, we isolate two distinct failure modes of the plug-in estimator: it leads to reward hacking within each mode and it cannot select high-reward modes. We propose a closed-form reward damping schedule that corrects the within-mode bias with no additional compute, and clarify the role of best-of-n sampling in compensating for the mode selection failure. Experiments on Gaussian mixture targets, a 2D checkerboard, and FLUX.1 text-to-image generation confirm that our theoretical insights carry over to practical settings.
Authors: Octave Oliviers, Glenn Vinnicombe
Abstract: The asymptotic behaviour of Monte Carlo Exploring Starts (MCES) is a long-standing open question in reinforcement learning, even in the tabular setting. We investigated the convergence properties of tabular MCES by constructing examples in which the algorithm converges to suboptimal solutions. This paper presents new counterexamples for both initial-visit and first-visit MCES and gives a convergence-restoring modification for the initial-visit case. We show that stable suboptimal solutions may exist for initial-visit MCES with sample-average updates even when greedy actions are updated more often than non-greedy actions on average. However, by scaling learning rates inversely to update frequencies on a state-by-state basis, convergence to optimality is guaranteed. Unlike previous uniformisation methods, this modification is applicable to large-scale problems that require approximating the estimated value function. We then extend the example to show that sample-average first-visit MCES may also converge to suboptimal solutions. This largely settles a fundamental open problem and shows that exploring starts alone do not guarantee convergence to optimality. More broadly, these results highlight that convergence depends critically on the relative size and frequency of updates applied to different actions, making the choice of learning rates and the balance between exploration and exploitation central to the analysis of MCES and the implementation of scalable Monte Carlo control methods.
PaperID: 7635, Poster
Authors:
Yuanze Hu, Zhichao Yang, Junwei Jing, Xin Yang, Wenxuan Liu, Tinghai Zhang, Jia Fu, Hongzhi Zhang, Xiangyu Wu, Ruiming Tang, Tingting Gao, Han LiAbstract: Detailed image captioning is usually framed as an amount problem: a model should mention more visual facts while avoiding hallucination. We argue that this framing hides a basic structure of long-form captioning: a caption is an ordered sequence, and different failures appear at different positions. We introduce a position-aware rubric representation that decomposes dense human references into atomic visual claims, verifies generated captions against those claims, and induces a quartile-wise composition of supported, hallucinated, and generic content. This lens reveals two qualitatively different failures in long captions: models tend to produce unsupported specificity in the middle, while the final portion often escapes into vague, non-evidential language rather than simply accumulating more hallucination. Dense human references do not exhibit the same position-dependent pattern under the identical pipeline, suggesting that the effect is a model behavior rather than an artifact of long descriptive text. Motivated by this diagnosis, we propose einforcement), a GRPO training objective with separate signals for supported fine-grained middle evidence, evidence-bearing tail continuation, claim coverage, and content-aware length control. The result is a closed loop: the same ordered claim representation identifies where long captions fail and supplies the position-specific reward needed to repair those failures. Experiments on detailed-caption benchmarks show that PAGER improves position-wise grounding and pairwise caption preference while preserving competitive global caption quality.
PaperID: 7636, Poster
Abstract: Owing to enhanced capabilities to capture a broad range of higher-order interactions, simplicial neural networks (SNNs) have emerged as a new powerful methodology for graph learning. However, prevailing SNNs are often limited in efficient exploration of heterogeneous graphs, thereby restricting the SNN utility on downstream tasks. To address this challenge, we introduce the concept of L\'evy flights over simplices, offering a new efficient alternative to learning higher-order graph (sub)structures under heterogeneous scenarios. Specifically, we develop a new Fractional Hodge Laplacian Simplicial Neural Network (FHL-SNN), a novel approach leveraging fractional powers of the Hodge-Laplacian and allowing for a more efficient exploration of the underlying higher-order graph organization in conjunction with the link prediction. We establish a theoretical connection between a fractional diffusion process and graph exploration over simplices. In particular, we prove that the L\'evy flight over simplices yields a lower mixing time than that of its non-fractional counterpart. Our extensive experiments on both directed and undirected link prediction tasks illuminate the power of non-local search of graph space via fractional dynamics, yielding relative gains of up to 9%. Finally, we show that the idea of L\'evy flights over simplices is versatile and that the integration of fractional dynamics to existing SNNs has the potential to further boost the model performance.
Abstract: Modern RL post-training methods such as GRPO and DAPO train on N response sequences of R tokens sampled from a shared prompt of P tokens, but standard FlashAttention replicates all P prompt tokens N times across both forward and backward passes --- duplicating compute and memory on identical hidden states. In large-rollout, long-context RL training (N\geq16, P\geq8\textK), this redundancy dominates the policy update cost. We observe that in decoder-only models, causal masking makes prompt representations invariant across sequences at every layer, so all per-token operations (norms, projections, MLP) and attention can process the prompt once --- a property not yet exploited at the kernel level for training. We propose DualKV, the first FlashAttention kernel variant that eliminates shared-prompt replication during RL training, via (1) fused CUDA forward and backward kernels that iterate over two disjoint KV regions --- shared context and per-sequence response --- in a single kernel launch, and (2)~a data-pipeline redesign in veRL that repacks N(P+R) tokens into P+NR tokens per micro-batch, extending the token reduction from attention to the entire model by a factor \rho = N(P+R)/(P+NR). DualKV is mathematically equivalent to standard attention and introduces no approximation. On Qwen3-8B GRPO training with 8×H100 GPUs (N=32, 8K-context), DualKV achieves 1.63--2.10× policy-update speedup, enables 2× larger micro-batches, and raises MFU from 36% to 76%. Similar gains hold for DAPO (2.47× speedup, 77% MFU). At 30B MoE scale on 16×H100, DualKV achieves 3.82× policy-update and 3.38× end-to-end step speedup over FlashAttention (which requires 4-way Ulysses sequence parallelism to avoid OOM).
PaperID: 7638, Poster
Abstract: Tool calling, the ability to invoke external tools on demand, is central to agentic LLMs. Within an agent, the LLM decides whether and when to use these tools, greatly extending the agent’s capability boundaries. However, the mechanism that determines whether a model calls a tool or responds directly remains poorly understood. Agentic prompts are often long and heavily scaffolded, combining role instructions, tool schemas, format templates, and user requests across thousands of tokens. This creates a noisy and highly entangled context that makes it difficult to identify a single controllable variable for mechanistic analysis.To obtain such a variable, we propose a method to convert complex agentic prompts into minimal contrastive pairs, in which a single request verb determines the tool-call decision. Replacing an execution verb, such as write, with an analysis verb, such as discuss, reliably flips whether the model chooses to call a tool. This suggests that the choice is mediated by a compact internal state. 1,500 paired prompts across Python, Java, and C++, where 1,200 are for mechanistic analysis, and 300 are held out for evaluation, are constructed.We trace the tool-calling decision to a single vector, \mu_\Delta, and show that it is both causally necessary and sufficient.Transcoder analysis reveals how this vector forms: the scaffold makes tool calling the default, while analysis verbs suppress this default by activating features that signal no tool is needed. Execution verbs activate no analogous tool-use features, so \mu_\Delta reflects the scaffold-induced default rather than a positive execution-verb signal. Downstream attention and MLP layers then read out this vector into the first output token. The mechanism replicates across diverse model families. Our code is available at \urlhttps://anonymous.4open.science/r/MI4ToolCalling.
PaperID: 7639, Poster
Abstract: Large language models (LLMs) provide a powerful reasoning backbone for speech understanding, but integrating continuous acoustic signals into a frozen LLM remains challenging. Existing speech-to-LLM interfaces typically operate at two extremes: either enforcing near-discrete token alignment, which benefits transcription but loses paralinguistic information, or learning unconstrained continuous representations, which can drift away from the LLM’s input space and degrade autoregressive decoding. In this work, we propose Convex Gate (C-Gate), a speech–LLM bridge that constrains all speech representations to lie within the LLM’s input embedding manifold with an architectural convex-hull constraint. Concretely, each frame is represented as a convex combination of token embeddings, ensuring compatibility with the pretrained LLM while preserving continuous expressivity. Across automatic speech recognition (ASR) and emotion recognition, C-Gate achieves strong joint performance, improving LibriSpeech WER by up to 48.7% relative while matching or exceeding single-task emotion accuracy. Beyond performance, our analysis reveals a key insight: information is not carried by discrete token identities, but by time-resolved trajectories in the embedding space. Causal interventions confirm that both the trajectory structure and alignment to the pretrained embedding manifold are critical for performance. These results suggest that geometry, rather than token discreteness, is the fundamental design factor in speech–LLM interfaces, and provide a controlled regime for studying multimodal integration in frozen LLMs. We release the checkpoint, per-sample outputs, mechanism dumps, and intervention suite for replication.
PaperID: 7640, Poster
Abstract: Hallucinations in large language models (LLMs) that are plausible-sounding but factually incorrect or unsupported pose a major challenge for deploying these models in high-stakes applications such as medical diagnosis, legal reasoning, and knowledge-based question answering. Existing detection methods primarily rely on either generating multiple responses to check consistency or supervised training on labeled hallucinations. Multiple-response approaches tend to be computationally expensive and may be less practical for real-time scenarios, while supervised methods are limited to known hallucination types and cannot generalize to unseen cases. To address these limitations, we propose an efficient unsupervised hallucination detection method. Our approach measures the embedding discrepancy between the LLM’s response considered independently and the response contextualized with the question. A large discrepancy indicates that the question significantly influences the response, suggesting it is grounded, whereas a small discrepancy signals a higher likelihood of hallucination. This method is efficient, interpretable, and does not require labeled hallucinations. Extensive experiments across multiple tasks demonstrate that our approach consistently achieves superior or comparable detection performance while reducing the detection cost compared to the multiple-response methods.
Abstract: Enforcing alignment between the internal representations of diffusion or flow-based generative models and those of pretrained self-supervised encoders has recently been shown to provide a powerful inductive bias, improving both convergence and sample quality. In this work, we extend this idea to inverse problems, where pretrained generative models are employed as priors. We propose applying representation alignment (Repa) between diffusion or flow-based models and a DINOv2 visual encoder, to guide the reconstruction process at inference time. Although ground-truth signals are unavailable in inverse problems, we empirically show that aligning model representations of approximate target features can substantially enhance reconstruction quality and perceptual realism. We provide theoretical results showing (a) that Repa regularization can be viewed as a variational approach for minimizing a divergence measure in the DINOv2 space, and (b) how under certain regularity assumptions Repa updates steer the latent diffusion states toward those of the clean image. We integrate Repa into multiple state-of-the-art inverse problem solvers, and provide extensive experiments on super-resolution, box inpainting, Gaussian deblurring, and motion deblurring confirming that our method consistently improves reconstruction quality, while also reducing the number of discretization steps required to reach the same performance level as the underlying solver.
Abstract: Local SGD, also known as Federated Averaging, is a widely used distributed optimization algorithm. Although it frequently outperforms alternatives such as Mini-batch SGD in practice, its theoretical advantage under realistic data heterogeneity remains only partially understood. Recent work demonstrates that bounded second-order heterogeneity accounts for the benefits of Local SGD for strongly convex objectives and conjectures similar advantages for general convex objectives. In this paper, we establish such gains for general convex objectives by providing an improved convergence guarantee for Local SGD under bounded second-order heterogeneity. We also improve the best-known lower bounds for Local SGD in this setting, showing that our upper bounds are almost tight. Using our techniques, we also improve the convergence guarantee of SGD-with-replacement under bounded second-order heterogeneity and obtain an almost-matching lower bound.
Abstract: Diffusion models achieve strong performance in generative modeling, but their success often relies heavily on classifier-free guidance (CFG), an inference-time heuristic that modifies the sampling trajectory. In theory, diffusion models trained with standard denoising score matching (DSM) should recover the target data distribution, raising two fundamental questions: (i) why is inference-time guidance necessary in practice, and (ii) can its underlying effect be internalized into a principled training objective? In this work, we argue that a key limitation of standard DSM is insufficient inter-class separation. To address this issue, we propose MCLR, an alignment objective that explicitly maximizes inter-class likelihood-ratios during training. Fine-tuning diffusion models with MCLR induces CFG-like improvements under standard sampling, substantially improving guidance-free conditional generation and narrowing the gap to inference-time CFG. Beyond these empirical benefits, we show theoretically that the CFG-guided score is exactly the optimal solution to a sample-adaptive weighted MCLR objective. This result connects CFG to alignment-based objectives, providing a mechanistic interpretation of CFG as an implicit inference-time contrastive alignment procedure.
Abstract: We study adversarial multi-armed bandits with and without delayed feedback under a safety-aware goal: achieving minimax-optimal worst-case regret while keeping nearly constant regret relative to a designated "safe" baseline policy. Existing approaches can balance this trade-off with immediate feedback for smooth comparators, but arbitrary delays can mistime transitions between conservatism and exploration, endangering the safety guarantee. To bridge this gap, we propose PRUDENT-BANKER, a novel algorithm that combines a delay-adapted variant of Online Mirror Descent with a modified phased-aggression mechanism. Its key technical contribution is a delay-calibrated restart threshold that rigorously accounts for the worst-case distortion induced by unobserved feedback and reliably detects comparator suboptimality. We also establish new lower bounds for safety-constrained adversarial delayed bandits, showing that the regret guarantees of PRUDENT-BANKER are unimprovable, up to logarithmic factors, under the baseline-safety requirement. To the best of our knowledge, PRUDENT-BANKER is the first algorithm to achieve the optimal safety–robustness trade-off: pseudo-regret \tildeO(\sqrtT+\sqrtD) together with \tildeO(1) regret against the safe comparator, both with and without delays. Experiments across diverse delay distributions show that, unlike standard delay-robust baselines, PRUDENT-BANKER effectively balances safety and learning.
PaperID: 7645, Poster
Abstract: Characterizing the properties of neuronal cell types is central to understanding brain function. Yet, biophysically realistic neuron models remain difficult to scale to large populations while preserving the richness of cellular dynamics and within-cell-type variability. We present a cell-type-specific neural-operator framework for scalable neuronal population modeling. Rather than training separate surrogate models for individual neurons, the proposed framework learns a shared operator across multiple modeled neurons from the same cell type, conditioned on biologically interpretable electrophysiological features. This enables efficient generation of large ensembles of somatic voltage responses while preserving variability across cells and model configurations. We evaluate the approach using membrane dynamics metrics, action potential waveform metrics, spike timing accuracy, input-output firing relationships, and electrophysiological feature distributions. Our framework reproduces subthreshold membrane dynamics, spike waveforms, and firing-rate responses across major cortical inhibitory neuron types, while accurately preserving spike timing despite the sensitivity of threshold-crossing dynamics. The predicted electrophysiological feature distributions show strong agreement with detailed simulator outputs and human experimental recordings, including high explained variability in properties such as inter-spike intervals, firing frequency, action-potential shape, and post-spike recovery. By enabling fast ensemble simulation of cell-type-specific neuronal responses, the proposed framework provides a scalable tool for studying cellular diversity, probing variability near functional thresholds, and generating synthetic but biophysically grounded neuronal populations.
Abstract: Large language models confidently produce outdated answers, and no existing method can detect them. We show this is not an engineering failure but a structural one: temporal drift, whether a stored fact has changed since training, is encoded as a direction in the residual stream geometrically orthogonal to both correctness and uncertainty. Any method operating on correctness or uncertainty signals is therefore blind to drift by construction. We verify this across six instruction-tuned models. A linear probe trained directly on drift labels achieves AUROC 0.83--0.95; methods based on token entropy, semantic entropy, CCS, and SAPLMA all remain near chance (0.49--0.57). Five tests confirm the geometric orthogonality: weight cosines (|\cos| \leq 0.14), score correlations (|r| \leq 0.20), bidirectional null-space projection (|\Delta| \leq 0.008), iterative null-space projection with k=10, and difference-of-means dissociation. Mechanistically, the MLP retrieval circuit produces identical dynamics for stale recall and confabulation (r > 0.81, six models), explaining why output confidence cannot separate them. A cross-cutoff experiment holds inputs constant and varies only the model: the probe fires on the model whose training predates the fact's transition and stays silent otherwise (P(A>B) = 0.975--0.998, twelve model pairs), confirming it reads model-internal knowledge state rather than input properties. Our code and datasets will be publicly released.
Abstract: Graph-structured optimization with linear constraints is fundamental to critical infrastructure but faces scalability limits due to massive strict hard constraints and high dimensionality. While recent projection-based methods such as Trainable Sampling Kaczmarz-Motzkin Net (T-SKM-Net) guarantee feasibility, they face high computational costs in dynamic environments by processing the entire constraint set and requiring expensive matrix factorizations. To bridge this gap, we propose the Accelerated Trainable-SKM (AT-SKM) Net framework. To concentrate computation on the active constraints and eliminate redundant calculations, we introduce a hybrid sampling strategy guided by a topology-aware heterogeneous GNN model. To efficiently handle topological shifts in graph-based constraints, we employ a Cholesky Update mechanism that theoretically reduces the equality projection complexity from \mathcalO(N^3) to \mathcalO(N^2) under low-rank perturbations. Experiments on random geometric graphs, N-1 Security-Constrained DC-OPF, and minimum-cost gas transport problem demonstrate that AT-SKM reduces iteration counts by up to 85% and achieves 2.95×-7.29× SKM layer speedups, while maintaining zero constraint violations. Our code is available at [https://anonymous.4open.science/status/Anonymous_Submission-CFC7](https://anonymous.4open.science/status/Anonymous_Submission-CFC7)
Authors:
Juan D. Guerra, Thomas Garbay, Numa Dancause, Guillaume Lajoie, Marco BonizzatoAbstract: Hierarchical Gaussian Process (H-GP) models divide problems into different subtasks, allowing for different components to address each part, making them well-suited for problems with inherent compositional structure. However, existing H-GP frameworks typically employ one-way information sharing — either top-down or bottom-up — which limits sample efficiency and slows convergence. We propose Bidirectional Information Flow (BIF), which establishes continuous two-way communication. BIF retains the modular structure of hierarchical models — the parent conditions its own posterior on child summaries, treating them as structured priors — while introducing top-down feedback to softly decompose environment observations from the parent into sub-responses. This mutual exchange improves sample efficiency, enables robust training, and allows modular reuse of learned subtask models. We prove analytically the regret of a GP with a learned kernel scales linearly with the mismatch to the true kernel, tightening in the hierarchical case to the sum of child-level errors. Ablation shows removing the downward pathway collapses child R^2 by up to 58%. Across synthetic, neurostimulation, and HPO benchmarks, BIF achieves up to 4× higher parent R^2 and ~100% AUC improvement over vanilla GPBO, and outscores all hierarchical state-of-the-art methods on child R^2 given the correct acquisition function, while supporting modular child transfer to novel composite tasks.
PaperID: 7649, Poster
Abstract: Production retrieval systems require three properties jointly: domain adaptation, preservation of general-retrieval quality, and compatibility with existing vector indices. Existing adaptation methods achieve at most two. Standard parameter-efficient fine-tuning (PEFT) adapts to a domain but breaks the index and degrades general retrieval; backward-compatible training (BCT) preserves the index approximately via distillation while sacrificing general retrieval under domain shift; frozen-backbone side networks (L2C) preserve general retrieval but still modify the deployed embedding via additive combining, so existing indices are not preserved. We introduce CoSE (Complementary Subspace Expansion), a side-branch LoRA whose frozen pathway is structurally untouched, so any index built from the base encoder remains valid --- the adapter is \emphpluggable, added on top of an already-deployed index without modifying existing entries. A joint training objective converts the frozen+LoRA concatenation into Pareto-efficient retrieval. Across four distinct domains (Korean broadcast, medical radiology, art paintings, remote sensing) and 3 seeds, CoSE at 6.4M trainable parameters is the only configuration that simultaneously stays within \pm 0.048 of the frozen base on Flickr, COCO, and ImageNet at the deployment-default \alpha=0.5, is domain-competitive with standard LoRA (ties on RSICD, within ~ 2\sigma on HAN and SemArt, -0.05 on ROCO), and is backward-compatible by construction. Deployment-time knobs --- \alpha blending, Matryoshka truncation, partial branching --- expose further cost/quality trade-offs without retraining.
Abstract: In Large Language Model (LLM) fine-tuning, parameter and data selection are common strategies for reducing fine-tuning cost, yet they are typically driven by separate scoring mechanisms. When a parameter mask and data subset jointly determine restricted fine-tuning, this separation incurs redundant overhead and makes coordinated selection difficult. We cast parameter and data selection as two bilevel selection problems under a common validation objective and derive a shared local response-surrogate scoring rule. Under first- and second-order validation-improvement approximations, parameter importance and data utility emerge as column-wise and row-wise aggregations of a single gradient interaction matrix, yielding a closed-form row--column correspondence for co-extracting both signals. Building on this structure, we propose DualSFT (Dual-Selection Fine-Tuning), a one-shot dual-scoring algorithm that produces a parameter mask and data subset from shared gradient statistics. On 3B-9B LLMs, single-axis DualSFT variants strengthen target-task performance and stability--plasticity trade-offs within their comparison groups, while full DualSFT yields a more favorable joint-constrained trade-off than sequential hybrid baselines under matched budgets.
PaperID: 7651, Poster
Authors: Junkai Liu, Haofan Wu, Nay Aung, Joao A Lima, Steffen E Petersen, Le Zhang
Abstract: Dense cardiac dynamics are essential for assessing cardiac function, yet in practice clinicians often observe only extremely sparse measurements. However, recovering full dynamics from such sparsity is a highly ill-posed problem due to the lack of temporal cues. While auxiliary time-series signals like ECG provide physiological guidance, effectively exploiting them is challenging due to the cross-modal representation gap and phase misalignment. In this work, we present CardioFAD, a Focus--Align--Diffuse paradigm that reformulates sparse cardiac dynamic synthesis as a recoverability--alignability anchor discovery problem. Our core insight is that resolving extreme temporal ambiguity requires a compact set of anchor states that are both sufficient to recover the visual trajectory and reliable for ECG--visual alignment. We therefore learn motion-salient keyframes under two complementary principles: trajectory sufficiency, which preserves trajectory-critical intra-modal motion content, and anchor alignment, which promotes consistent inter-modal correspondence at these anchors. We implement the Focus and Align modules optimized by the two aspects.Diffuse then performs two-stage generation by first synthesizing keyframes and subsequently interpolating the remaining frames. Extensive experiments on cardiac MRI and echocardiography videos demonstrate superior visual quality, temporal coherence, and clinical fidelity under the extreme one-shot regime.
PaperID: 7652, Poster
Authors:
Mauricio Byrd Victorica, Ezzeldin Shereen, György DánAbstract: Visual adversarial examples are a well-known vulnerability of deep learning (DL) systems. The emergence of vision-language models (VLMs) further expands the attack surface through multimodal interactions. Despite extensive research on adversarial defenses over the past decade, existing works on VLM robustness often overlook adaptive white-box attacks, where the adversary has full access to both the model and the defense mechanism. We show that recent defenses can be effectively bypassed under such adaptive settings and propose \method, an adaptive test-time detection method that leverages gradient information induced by a self-targeted attack maximizing the likelihood of the VLM’s generated output. Our approach consistently outperforms existing defenses across multiple victim VLMs, attack formulations, and benchmark datasets. Our code is open-source and available online.
Abstract: We study the max-margin solutions reached by mirror flow in deep neural networks with homogeneous activation functions. Extend classical results on gradient flow, we derive a novel balance equation for mirror flow from convex duality, enabling a characterization of the horizon function governing the induced margin. We further establish max-margin characterizations together with convergence rates and norm growth estimates. Finally, we support our theory through experiments on synthetic datasets and standard vision tasks. Concretely, we show that: (1) distinct non-homogeneous mirror maps can induce the same max-margin solution; (2) convergence can be extremely slow, including exponentially slow regimes; and (3) although all considered mirror maps exhibit feature learning, they can produce markedly different representations, ranging from sparse to dense neuron activations. Together, these results provide a unified perspective on sparse and dense feature learning in homogeneous neural networks, highlighting how mirror maps shape both optimization dynamics and the geometry of the learned classifiers.
Authors: Armand Rousselot, Joran Wendebourg, Ullrich Köthe
Abstract: The performance of machine learning models depends on discovering features in the data that are relevant to solving the given task while ignoring irrelevant variation. To check and visualize whether this is the case, we propose ‘invariance auditing’, a diagnostic tool to analyze feature extractors. Invariance auditing generates diverse inputs which all yield the same features as a given query input, and are therefore treated as identical by the model. Using external knowledge about the task at hand, this enables assessing if the learned invariances do or do not conform to the task's required semantics. Unlike existing work where a dedicated generative model is trained for each feature extractor, our algorithm is training-free and exploits a pretrained diffusion or flow-matching model to sample invariant inputs. Our fiber loss -- which penalizes feature mismatch -- guides the denoising process so the output shares the same representation as the query input. This replaces days of training with a single guided generation procedure at the same quality. Experiments on popular datasets and model types demonstrate that our auditing method reveals invariances spanning very desirable and concerning behavior. For instance, it detects cases where Qwen-2B places patients with situs inversus (heart on the right side) in the same fiber as typical anatomy.
Abstract: Reusing a held-out benchmark adaptively should, in principle, invite overfitting. Yet benchmark-driven machine learning (ML) has produced surprisingly little overfitting in practice. An attractive hypothesis is that successful ML strategies are highly compressible. We study this in the setting of LLM-driven research agents, where the hypothesis becomes directly testable via two complementary information bottlenecks. In \emphoutput compression, an exploration agent adaptively searches for high-performance models using a validation set, and we test whether a fresh ``reproducer agent'' can reproduce its performance given only an extremely short prompt and the training data. In \emphinput compression, the explorer receives only one-bit feedback indicating whether each submitted model improves on the running best. Across 8 datasets spanning tabular classification, vision, language modeling, diffusion modeling, and reward modeling, we find that these bottlenecks have little effect on performance: short prompts and compressible feedback are sufficient to reproduce and find high-performance models. The hypothesis is falsifiable: when we deliberately induce validation-set overfitting, the results fail to reproduce with short prompts. Taken together, our results support a description-length explanation for the lack of overfitting in benchmark-driven ML: successful strategies occupy a low-complexity region of strategy space.
PaperID: 7656, Poster
Abstract: While large reasoning models (LRMs) can solve complex tasks via generating extended Chain-of-Thought traces (Long-CoT), longer traces often increase confidence without increasing correctness. In this work, we formalize such failure via a \emphself-conditioned evidence model: generated tokens provide information about the model's current working hypothesis rather than the ground-truth answer. This yields a simple diagnostic principle: entropy, fluency, and answer stability can certify within-trajectory self-consistency while remaining causally disconnected from correctness. Building on these findings, we propose Faithful Pandora Inference (FPI), a minimal training-free controller that keeps the base LRM frozen and fits only a small calibration map on held-out data. At each checkpoint, FPI estimates calibrated correctness r_t and the marginal evidence gain m_t of continuing the current trajectory. It stops only when r_t reaches a user-specified reliability target; if r_t remains insufficient and m_t falls below cost, it reallocates the remaining budget to a diversified restart. Across six reasoning benchmarks and multiple DeepSeek-R1 distilled models, FPI improves the accuracy-efficiency frontier, reduces expected calibration error, and improves selective accuracy on high-confidence outputs. Ablations on confidently wrong trajectories show that calibrated stopping and marginal-gain restarts address different failure modes and that the gains are not explained by additional samples alone.
PaperID: 7657, Poster
Authors: Karim E Kanbi, Yannis Cattan, Aziz Fouché, Charlotte Claye, Pierre Marschall, Julien Duquesne
Abstract: Recent work has reported two failure modes of transcriptomic foundation models: validation loss plateaus beyond 100M parameters, and consistent underperformance against simple linear baselines on clinically relevant tasks. We show that both findings invert under a specific pretraining recipe. We introduce EVA-RNA, a transformer pretrained on 545k human and mouse samples spanning bulk RNA-seq, microarray, and pseudobulked single-cell data, scoped to immunology and inflammation, one of the largest therapeutic areas in clinical development. EVA-RNA exhibits clean power-law scaling from 7M to 300M parameters, with no plateau emerging within our scale range. On a benchmark co-designed with immunologists and drug development experts, EVA-RNA outperforms existing foundation models on every task category, spanning drug discovery, preclinical-to-clinical translation and patient stratification. Mechanistically, EVA-RNA learns species-invariant representations in which orthologous genes progressively align across layers without supervision. We interpret these results as evidence that scope and data composition are sufficient, with conventional architecture and knowledge-informed gene embeddings, to achieve clinical utility and scaling in I&I. We also release EVA-RNA 60M model weights to support continued investigation.
Abstract: Compositional generalization, the ability to generate novel combinations of known concepts, is a key ingredient for visual generative models. Yet, not all mechanisms that enable or inhibit it are fully understood. In this work, we conduct a systematic study of which design choices critically determine compositional generalization in image and video generation. By isolating independent design axes, we identify two key factors that determine compositional success: (i) whether the training objective operates on a discrete or continuous distribution, and (ii) the completeness of conditioning information about constituent factors during training. We also show that relaxing the discrete loss with an auxiliary continuous latent objective can partially recover compositional performance in discrete models like MaskGIT. Our findings, corroborated by diverse compositional tasks and preliminary evidence in world models and LLMs, motivate a shift toward continuous objectives for compositional generalization.
PaperID: 7659, Poster
Abstract: Modern video diffusion models generate increasingly realistic and temporally coherent videos, motivating their use as candidate world simulators. Yet it remains unclear whether these models internally encode physical structure, or merely reproduce motion patterns seen during training. We study this question by probing video diffusion models along latent trajectories corresponding to real videos with known physical plausibility. To obtain such trajectories, we approximately invert the deterministic sampling process by integrating the learned velocity field backward from a clean video latent to noise, giving access to the model’s intermediate states and attention maps. Using these recovered trajectories, we show that physical plausibility is linearly decodable from diffusion transformer states across IntPhys and InfLevel, reaching around 81% average accuracy and outperforming dedicated representation-learning baselines such as V-JEPA and VideoMAE. Surprisingly, this signal is absent from the VAE latent input and emerges inside the denoising transformer itself, despite the model not being trained with a self-supervised predictive objective. These findings suggest that physically meaningful representations can arise as a byproduct of generative denoising.
PaperID: 7660, Poster
Abstract: Action Quality Assessment (AQA) for diving aims to quantify the execution quality of diving sequences and has long been a prominent research topic. However, existing diving AQA methods typically regress solely the final diving score, leading to poor interpretability. To address this challenge, we propose Diving-R1, the first attempt to leverage a Multimodal Large Language Model (MLLM) for interpretable diving AQA, enabling the integration of comprehensive information, thereby supporting progressive reasoning and traceable assessment. We construct a novel dataset, DivingThink, as the foundation, which contains progressively structured reasoning chains that follow judging logic and explicitly describe the execution quality of each diving sub-action, providing clear evidence for score determination. Furthermore, we establish a new benchmark, DivingInterp, which evaluates the performance of MLLMs on interpretable diving AQA from multiple complementary perspectives, including score estimation, textual similarity, and reasoning quality. Additionally, we carefully design a three-stage training paradigm, combined with a dedicated composite reward function, which gradually relaxes the annotation requirements for the training corpus at each stage to incrementally enrich the diversity of the training data, achieving continuous improvement. Extensive experiments demonstrate the superiority of Diving-R1 over existing MLLMs in interpretable diving AQA. The dataset and code will be available on GitHub.
Authors: Yigit Turkmen, Baturalp Buyukates, Melih Bastopcu
Abstract: Large language models (LLMs) are often ensembled together to improve overall reliability and robustness, but in practice models are strongly correlated. This raises a fundamental question: which models should be selected when forming an LLM ensemble? We formulate budgeted ensemble selection as maximizing the mutual information between the true label and predictions of the selected models. Furthermore, to explain why performance can saturate even with many models, we model the correlated errors of the models using a Gaussian-copula and show an information-theoretic error floor for the performance of the ensemble. Motivated by these, we propose a simple greedy mutual-information selection algorithm that estimates the required information terms directly from data and iteratively builds an ensemble under a query budget. We test our approach on four datasets spanning binary and multi-class classification and sentiment analysis: MEDMCQA, MMLU, AG News, and IMDB. Across all datasets, we observe that our method consistently outperforms strong baselines under the same query budget.
Abstract: We study a data-dependent notion of diffusion-model generalization: when a model does not memorize the training set, where do its generated samples go relative to the geometry induced by the data? To answer this, we introduce a time-dependent family of log-density ridge manifolds constructed from the smoothed empirical distribution, and use it to characterize reverse-time inference. Our main result shows that generated samples evolve by a mechanism: they first enter a neighborhood of the ridge, then their distance to the ridge is controlled by the normal component of training error, and finally their motion along the ridge is controlled by the tangential component. We further connect this geometric picture to training dynamics through directional decompositions of the learned error, and make this link explicit for random feature models, where architectural bias and optimization error can be separated quantitatively. Experiments on synthetic multimodal data and MNIST latent diffusion support the predicted geometric behavior in both low and high dimensions.
PaperID: 7663, Poster
Abstract: AI-Generated Image (AIGI) Detection is becoming critical as generative models proliferate, yet existing detectors often generalize poorly by relying on narrow, dataset-specific artifacts. Many prior works improve generalization through additions, such as introducing extra artifact cues, pretraining models, training constraints, or test-time operations, which may increase inference cost and model complexity. In this paper, we take the opposite route and study generalization through subtractions and propose AID--, a plug-and-play framework that improves transformer-based AIGI detectors from both feature and parameter perspectives. At the feature level, we introduce high-order dependency ranking to identify and prune redundant or potentially misleading tokens for progressive token reduction; and at the parameter level, we design relative-advantage head supervision and a class-asymmetric exit rule to train depth-specific early-exit heads. AID-- can be seamlessly integrated into transformer-based AIGI detectors to improve both efficiency and out-of-distribution generalizability, which is experimentally validated on six public datasets. For example, AID-- helps the detector Effort achieve up to a 2.2× speedup while improving generalization accuracy by 3.6% on GenImages.
PaperID: 7664, Poster
Abstract: Large Language Models (LLMs) are increasingly deployed as automated judges to scale supervision for data curation, reinforcement learning (RL), and agentic systems. While existing works have extensively explored bias in pretrained LLMs, the origins of such inherent biases remain largely untraced. We trace these biased tendencies to cognitively biased patterns (e.g., Authority, Bandwagon) latent in training corpora, and expose that such patterns are naturally prevalent in large-scale pretraining corpora yet remain entirely undetectable by existing data cleaning pipelines. To demonstrate that this underexplored threat can be deliberately exploited, we introduce BiasTrojan, a framework that concentrates and injects these naturally occurring bias patterns into training samples via context-aware bias cues and contrastive preference pairs, augmented with counterfeited reasoning chains for efficient injection. Experiments across several LLMs from 7B to 70B on human-preference and fact-related datasets show that mere hundreds of deliberately biased samples suffice to compromise LLMs into biased evaluators, overriding their factual knowledge. The injected biases generalize robustly out-of-domain and persist despite massive continual post-training. These findings reveal that this underexplored latent threat poses far greater risks than commonly recognized: biased LLM judges whose evaluations propagate irreversibly downstream, underscoring the critical need for bias-aware auditing and strict scrutiny of LLM training data.
PaperID: 7665, Poster
Authors:
Ziji Sheng, Guiyao Tie, Weidong Wang, Jiawen Shi, Junhao Dong, Xiaoye Qu, Shanshan Ye, Daizong Liu, Dengpan Ye, Pan Zhou, Jing ZhangAbstract: Large vision-language models (LVLMs) achieve strong multimodal reasoning performance but remain prone to object hallucination, where generated objects or existence judgments are not faithfully grounded in the image. Existing mitigation methods either require additional training or rely on external visual experts, which limits plug-and-play usability and increases inference overhead. In this paper, we propose SIGMA, a training-free framework for hallucination mitigation based on Self-vision Information-theoretic Gating and dual-Mask Assembly. SIGMA operates directly on the LVLM's own vision encoder, constructing complementary support and complement masks from image-driven concept relevance, question-aware entity cues, and CLS--patch shortcut suppression. These masks form global, support-masked, and complement-masked visual streams for query-conditioned evidence separation. During decoding, SIGMA activates correction only under high uncertainty, using full-vocabulary and candidate-answer entropy to perform support-first refinement and complement-on-demand suppression. The resulting streams are combined through logit-level assembly with lightweight positive calibration and a plausibility constraint. Experiments across multiple LVLM backbones and hallucination-oriented benchmarks show that SIGMA effectively reduces object hallucination while maintaining competitive perception performance and a selective low-overhead inference path.
Abstract: Pretraining for electroencephalogram (EEG) foundation models has predominantly relied on self-supervised masked reconstruction, a paradigm largely adapted from and inspired by the success of vision and language foundation models. However, unlike images and text, EEG datasets are notoriously expensive to collect and characterized by low signal-to-noise ratio. These challenges introduce difficulties in scaling the EEG foundation models and capturing the underlying neural semantics through reconstruction. In this work, we ask the question: can we stand on the shoulders of well-established foundation models from well-represented modalities to bootstrap the pretraining of EEG foundation models? We first demonstrate that mainstream foundation models, such as those from vision and time series, transfer surprisingly well to EEG domain. Motivated by this observation, we propose the multi-teacher distillation pretraining (MTDP) framework for pretraining EEG foundation models via two-stage multi-teacher distillation. In the first stage, we introduce a learnable gating network to fuse representations from diverse teachers (e.g., DINOv3 and Chronos) via a masked latent denoising objective. In the second stage, we distill the fused representation into an EEG foundation model. Extensive evaluations across 2 backbone architectures, 9 downstream tasks and 12 datasets demonstrate that multi-teacher distillation pretraining consistently outperforms self-supervised pretraining for EEG foundation model and remains competitive even when the pretraining data is reduced to 25%.
Authors: Mohsen Dehghankar, Abolfazl Asudeh
Abstract: Sparse attention improves LLM inference efficiency by selecting a subset of key–value (KV) entries, but at the cost of potential accuracy degradation. In particular, omitting critical KV entries can induce substantial errors in model outputs. Existing methods typically operate under fixed or adaptive token budgets and provide empirical robustness or partial theoretical guarantees, yet they do not ensure zero false negatives across decoding steps, particularly since the set of relevant tokens is both query- and step-dependent. Our empirical observations confirm that missing even one critical key can lead to sharp error spikes, especially in long-output reasoning tasks where the set of important tokens varies throughout decoding. This observation motivates the need for indexing methods that dynamically adapt to these variations across decoding steps while guaranteeing a full recall of the relevant keys above a certain threshold. We address this challenge by reformulating sparse attention as the computational geometry problem of halfspace range searching. However, existing range searching data structures are not suitable for modern LLM inference due to their computational and implementation overheads. To overcome this, we introduce Louver, a novel index structure tailored for efficient KV cache retrieval. Louver (i) guarantees zero false negatives with respect to a specified threshold in both theory and practice, (ii) is lightweight to integrate into existing LLM inference pipelines, and (iii) incorporates hardware-aware optimizations for both CPU and GPU executions. Our experiments demonstrate that Louver outperforms prior sparse attention methods in both accuracy and runtime, and is faster than highly optimized dense implementations such as FlashAttention. These results highlight that recall guarantees are a critical and overlooked dimension of sparse attention, and open a new direction for building theoretically grounded, efficient KV cache indices.
PaperID: 7668, Poster
Authors: Chanho Kim, Kyeonghwan Gwak, Muhammad Salman Ali, Muhammad Shaheryar, Incheol Park, Jeongwan On, Seungryul Baek
Abstract: Recent hand avatar reconstruction methods achieve high-quality results through multi-view observations or per-subject optimization, limiting scalability and practical applicability. We present \emphFF3DGS-Hand, a feed-forward framework that reconstructs animatable hand avatars from a single monocular image using 3D Gaussian Splatting. Our method establishes a stable canonical geometry for thin and articulated hand structures, then reconstructs appearance by leveraging source image evidence instead of relying solely on global latent features. We further introduce a rendering-aware, source-conditioned appearance framework that selectively updates Gaussians according to their rendering contribution, recovering input-specific details while suppressing artifacts in unobserved regions. The resulting representation is animatable via standard hand pose parameters and supports efficient rendering. Experiments demonstrate strong visual quality and quantitative performance in the single-image setting.
Abstract: We study Nash equilibrium learning in partially observable Markov games (POMGs), a multi-agent reinforcement learning framework in which agents cannot fully observe the underlying state. Prior work in this setting relies on centralization or information sharing, and suffers from sample and computational complexity that scales exponentially in the number of players. We focus on a subclass of POMGs with independent state transitions, where agents remain coupled through their rewards, and assume that the underlying fully observed Markov game is a Markov potential game. For this class, we present an independent learning algorithm in which players, observing only their own actions and observations and without communication, jointly converge to an approximate Nash equilibrium. Due to partial observability, optimal policies may in general depend on the full action-observation history. Under a filter stability assumption, we show that policies based on finite history windows provide sufficient approximation guarantees. This enables us to approximate the POMG by a surrogate Markov game that is near-potential, leading to quasi-polynomial sample and computational complexity for independent Nash equilibrium learning in the underlying POMG.
PaperID: 7670, Poster
Abstract: Satellite constellations now image the Earth daily, capturing the lifecycle of events from deforestation to urban construction. Anticipating these events before completion enables timely intervention, yet existing systems cannot jointly identify where a change occurs, what it will become, and when it will finish across an open vocabulary. Building such a model requires diverse event data, but events are like needles in a haystack, making manual annotation at global scale infeasible. We address this problem by representing change as the vector difference between vision-language model embeddings at distinct time points. This arithmetic approach allows us to search for semantic transformations directly in the latent space, enabling (1) an automated data engine for global event discovery without manual annotation and (2) the largest global satellite event dataset to our knowledge, comprising 21,000+ locations across 64+ event types with ~2M images from PlanetScope and Sentinel-2 spanning six continents, to train (3) a novel global-scale event anticipation model. Our model detects change with a 93.5% max F1 score, outperforms baselines 3.6× on predicted event retrieval, and forecasts completion dates with a 3-day median error, even nine months before completion.
PaperID: 7671, Poster
Abstract: Reinforcement learning with verifiable rewards (RLVR) has greatly advanced large reasoning models (LRMs), but it requires timely training on a huge fully-annotated dataset. To this end, data-efficient RLVR methods have been widely studied from two perspectives: (i) data selection methods identify a small subset of "golden" samples that yield near-full-data performance, but they rely on a pre-existing pool of labeled data. (ii) unsupervised RLVR methods train the model using its own internal supervision signals on large-scale unlabeled data, yet they exhibit suboptimal performance. Accordingly, we investigate the "pick in the dark" setup for RLVR, which aims to select, without prior supervision, unlabeled samples that are most beneficial for training and worthy of annotation. Through systematic analysis, we demonstrate that smart picks hinge on a well-calibrated uncertainty estimator to enable strategic partitioning of data for adaptive training regimes. Building on this insight, we propose PivotTrace, a three-way data triage framework that leverages attention dynamics to trace metacognitive pivots during reasoning. By precisely quantifying uncertainty through pivot density, PivotTrace achieves automated data routing to synergistically maximize both annotation and training efficiency. Empirically, PivotTrace surpasses the fully supervised LRM with only 29.3% annotated samples and 2.75× faster convergence.
PaperID: 7672, Poster
Abstract: Vision-Language Models (VLMs) such as CLIP have demonstrated strong performance across tasks yet remain highly vulnerable to adversarial attacks. While adversarial fine-tuning has been demonstrated to be effective against such malicious inputs, its multi-step adversary generation scheme during fine-tuning further induces prohibitive computational costs for VLM backbones. Previous works in unimodal models have introduced single-step strategies to reduce this burden. However, we find naive single-step strategies fail on VLMs due to more severe catastrophic overfitting. In addition, we identify a novel failure mode termed , where VLMs overfit to the specific training attack budget while paradoxically becoming fragile to weaker perturbations. We trace these failures to growing angular misalignment and magnitude surge of the local gradients, which compromise the local linearity essential for accurate single-step approximation. Guided by this analysis, we introduce a joint adversarial optimization scheme that actively rectifies the local geometry by suppressing both angular misalignment and magnitude surge along the perturbation path. Our approach consistently outperforms existing single-step methods across various datasets and VLM architectures, achieving state-of-the-art robustness while preserving efficiency. To our knowledge, this is the first work to systematically examine single-step adversarial fine-tuning on VLMs and establish the geometric foundations of their robustness.
PaperID: 7673, Poster
Abstract: Object hallucination, where large vision-language models (LVLMs) generate objects not supported by the visual input, is a persistent challenge caused by visual uncertainty during decoding. Existing methods reduce hallucinations using contrastive signals, but they rely on heuristics and lack principled control of false positives at the image level. To address this, we propose False Discovery Rate-COntRol of Object HALlucination ( ), a training-free framework that models visual uncertainty using an uncertainty-aware visual data splitting strategy and leverages mirror statistics to quantify visual contrast during decoding. By computing mirror statistics from paired, symmetrically perturbed visual inputs, estimates spurious object predictions and sets a data-driven threshold to control the expected fraction of false discoveries per image, suppressing hallucinations while retaining high power for truly grounded objects. The framework is flexible, supports multiple LVLMs, and mitigates hallucinations without retraining or supervision. Extensive experiments on multiple benchmarks with several evaluation metrics demonstrate that consistently outperforms state-of-the-art methods, providing more reliable and robust hallucination control.
PaperID: 7674, Poster
Authors: Alessandro Licciardi
Abstract: Clustered Federated Learning (FL) partitions a client population into groups of similar local distributions and trains one specialized model per cluster, mitigating client drift that degrades single-model methods under non-IID data. Prior methods discover cluster structure inside the training loop through gradient similarity, loss evaluation, or EM-style updates, thus increasing communication overhead, exposing gradients to inversion attacks, and providing no mechanism to assign clients absent from training. We propose Ripple, a clustered FL framework in which cluster assignment is computed entirely offline from a spectral characterization of each client's local data: a variance-weighted principal-component prototype embedded via the Wavelet Scattering Transform and decoded by a Gaussian Mixture VAE trained server-side on synthetic client populations before federation begins. Per-round communication cost matches FedAvg exactly, and a client absent from training obtains a personalized model from a single forward pass, without gradient computation, model evaluation, or extra communication round. We prove that the gap between Ripple's surrogate clustered objective and the oracle is bounded by a computable quantity decaying with client sample size and independent of federation duration; per-cluster convergence matches the minimax-optimal rate for non-convex smooth objectives. Across five benchmarks spanning controlled and realistic heterogeneity, Ripple consistently outperforms all baselines, with margins growing on the most realistic partitions.
PaperID: 7675, Poster
Authors:
Huamin Chen, Xunzhuo Liu, Yuhan Liu, Junchen Jiang, Steve Liu, Chenxu Niu, Bowei HeAbstract: Energy efficiency in LLM inference is often treated as a property of the model, the hardware, or the serving software. We show that, for decode-dominated serving, a simpler systems variable dominates: the serving context window. Because KV-cache memory is finite, increasing the context window reduces the number of concurrent sequences a GPU can hold. Under continuous batching, this concurrency limit directly sets decode throughput, while GPU power remains nearly flat over most of the operating range. The result is a simple 1/\mathcalW law: tokens per watt scale inversely with the serving context window \mathcalW. We derive this law from KV-cache capacity and a logistic GPU power model, and empirically verify it for Llama-3.1-70B inference over the 2K--128K context range. On identical H100 hardware and software, tok/W varies by nearly 40× across this range. We then show that the same mechanism determines GPU fleet-level energy efficiency. Context-length routing keeps short requests on high-concurrency short-window pools, while newer hardware shifts the power and memory curve upward. These two operator-controllable levers are approximately multiplicative: on the short-context-dominant Azure trace, two-pool context-length routing improves tok/W by 2.5× over a homogeneous NVIDIA H100 fleet, a calibrated NVIDIA B200 deployment improves tok/W by 2.0× at fixed topology, and combining them reaches 5.76×. Finally, we analyze active-parameter weight streaming in MoE models as a third, architecture-level lever, giving an upper-bound gain of up to 5.1×. Together, these results suggest a deployment order for energy-efficient LLM serving: route by context length first, upgrade hardware second, and exploit sparse architectures where dispatch costs are controlled. We release code, data, and simulator in anonymous [repository](https://anonymous.4open.science/r/gpu-fleet-sim-1BEB) to facilitate reproduction and further research.
PaperID: 7676, Poster
Authors: Denis Larsen, Kazi S Ripon, Anis Yazidi, Gustavo B Moreno e Mello
Abstract: In artificial systems, neuroevolution with synaptic plasticity promises agents that adapt during their lifetime, yet standard fitness evaluations do not distinguish inherited static competence from within-lifetime adaptation. Evolutionary algorithms can produce \emphstatic solvers: genomes whose inherited topology performs well even when synaptic plasticity is disabled and weights remain at their lifetime initialization, confounding adaptation with evolved structure. To expose this confound, we introduce a two-phase evaluation protocol with an explicit static-solver extinction mechanism that measures each genome both with plasticity disabled (frozen-weight) and enabled (plastic-weight) on the same task. The first phase quantifies the static competence of the inherited network structure, while the second adds online learning through spike-timing-dependent plasticity (STDP). The \emphextinction mechanism removes static solvers from the gene pool, i.e., any genome whose frozen-weight fitness exceeds a threshold, thereby exerting evolutionary pressure for learning and against static solving. We combine this protocol with a complete 2×2 factorial design that independently toggles extinction pressure and within-lifetime learning. The method is instantiated on a mutable cart-pole benchmark using NEAT-evolved spiking networks with inherited topology and per-connection STDP parameters, while synaptic weights are reinitialized each lifetime and STDP updates run continuously at inference time during evaluation. Across 30 replicates per condition, the protocol reveals a clear static-solver fingerprint: plasticity-free neuroevolution reaches high training fitness but suffers a roughly 33-point train--test gap on held-out pole lengths. Learning-enabled agents reduce this gap to about 9 points and achieve the highest held-out AUC. Extinction does not improve mean test performance when learning is available; instead, it reduces replicate variance and strengthens attribution by preventing high-fitness frozen-weight genomes from dominating selection. These results position extinction not as a generic performance booster, but as a diagnostic intervention for disentangling inherited structure from genuine within-lifetime adaptation.
Abstract: A key problem in the modern study of Deep Learning is predicting and understanding emergent capabilities in models during training. Inspired by methods for studying reactions in quantum chemistry, we present the ``2-datapoint reduced density matrix" (2RDM). We show that this object provides a basis for computationally efficient, unified observables of phase transitions during training. First is the spectral heat capacity, which we prove provides early warning signals for learning events. Second is the participation ratio, which reveals the dimensionality of the underlying reorganization associated with a learning event. Remarkably, the top eigenvectors of the 2RDM are directly interpretable, making it straightforward to study the nature of the transitions. We validate across four distinct settings: deep linear networks, induction head formation, grokking, and emergent misalignment. We then discuss directions for future work using the 2RDM.
Authors: Fernando Martin-Maroto, Nabil Abderrahaman-Elena, David Mendez, Gonzalo de Polavieja
Abstract: We propose a machine-learning framework based on the principle that learning can be achieved through algebraic decomposition rather than numerical optimization. A learning task is encoded as axioms of an algebra, and learning proceeds by computing a subdirect decomposition of the algebra into subdirectly irreducible components. We show that suitable subsets of these components form models that generalize to unseen data because they approximate the underlying rule in the training data. The method has no dataset-specific hyperparameters to tune and shows no evidence of classical overfitting in our experiments, so we run it without a validation dataset. We compare against multilayer perceptrons as generic parametric baselines without task-specific inductive biases. Across 24 classification datasets, the algebraic method is applied only once to each training set without hyperparameter tuning, whereas the multi-layer perceptrons are selected from many configurations using validation data or cross-validation. Nevertheless, the resulting test accuracies are statistically indistinguishable. Moreover, adding a statistical readout on top of the algebraic representation, using untuned logistic regression, improves over the best multilayer perceptrons selected by validation or cross-validation. More broadly, because task axioms may encode examples, formal knowledge, or both, this approach provides a common algebraic framework and the same algorithm for data-driven problems such as classification, formally specified problems such as computing Hamiltonian cycles from specifications, and mixed settings combining data with formal constraints.
PaperID: 7679, Poster
Abstract: The Hessian of the energy with respect to the nuclear positions is indispensable in atomistic modelling. However, constructing this matrix requires O(N) Hessian vector products, traditionally limiting high-accuracy Hessians to small systems. Machine learning interatomic potentials (MLIPs) have accelerated atomistic modelling by providing highly accurate energies and forces at O(N) cost, yet the resulting O(N^2) cost of Hessians remains a practical bottleneck for large systems. Based on the insight that we can derive the sparsity pattern for an MLIP's Hessians in closed form, we show in this paper how to use techniques from sparse automatic differentiation to reduce the cost of a local MLIP's Hessians to a system-size-independent number of Hessian-vector products, yielding overall O(N) total cost without any approximations. We benchmark our approach on a variety of systems ranging from alkane chains to water clusters to A\beta40 conformers. Depending on the MLIP configuration, we achieve the linear scaling regime already on relatively small systems, resulting in large runtime reductions between 2×-15× for these systems. This opens up the possibility of scaling high-accuracy MLIP Hessians to very large systems, such as proteins that were previously inaccessible.
Abstract: The performance of machine learning models often relies on large labeled datasets; however, data collected from diverse sources can contain label noise. Recent work has shown that, in noisy settings, there may exist a subset of the training data on which models can achieve performance comparable to training on a noise-free dataset. A widely used method for identifying such subsets is cutstats, which employs k-nearest neighbors (k-NN) to detect low-noise samples. However, its performance on high-dimensional data remains largely unexplored. In this work, we formally establish that the performance of a classifier trained on a subset of a noisy dataset selected via cutstats is influenced by the accuracy of k-NN. We further demonstrate that, in noisy environments, exploiting data invariance and knowledge of underlying symmetries can significantly enhance the performance of k-NN, bringing it closer to the Bayes optimal classifier even in high-dimensional regimes. Finally, we show that for real-world scenarios, where information about the underlying invariance is only partially known, learnt invariant representations can still facilitate the identification of near-optimal subsets.
PaperID: 7681, Poster
Abstract: Recently, reinforcement learning from verifiable rewards (RLVR) is practically very important, where the leading approach is group relative policy optimization (GRPO) that typically faces a trade-off between rollout efficiency and downstream performance. To improve efficiency while preserving performance, existing selective rollout methods mainly exploit only source-side training dynamics to optimize the discrete distribution over source prompts. However, in practice, the source training dataset is rarely identically distributed with the target test dataset. When such a mismatch exists, due to legal or privacy constraints, it is usually not acceptable to share target prompts themselves with the RLVR trainer. Nevertheless, it is still possible to share some target-side feedback such as accuracy on a small held-out data. In this paper, we propose learning-to-reweight rollout (L2RR) that explicitly incorporates target-side performance feedback into GRPO. Specifically, L2RR learns a discrete sampling distribution over source prompts by optimizing a bi-level objective that maximizes accuracy on a held-out target dataset, while simultaneously penalizing rollouts that provide little incremental learning information. Empirical studies on six math reasoning benchmarks and three model scales show that L2RR improves target-domain performance and increases rollout efficiency.
PaperID: 7682, Poster
Abstract: Recent text-to-image diffusion models (T2I DMs) depend on vast, uncurated datasets that often include copyrighted artworks and personal images, risking the generation of unwanted content. While concept erasure methods have emerged to suppress such outputs, they predominantly rely on text prompts, limiting their effectiveness for visually nuanced or ineffable styles that are difficult to articulate verbally. To bridge this gap, we propose Quiet Prompt (QuP), a reference-based concept erasure method that identifies complex concepts using fewer than ten reference images. QuP operates in two stages: (1) Concept Embedding Generation (CEG), which captures the intricate visual morphology of an unwanted concept into a single latent embedding; and (2) Concept leakage-aware Negative Guidance (CNG), which utilizes this embedding as a negative semantic condition to steer the diffusion process. In addition, to overcome the limitations of conventional fixed-scale negative guidance (negative prompting)—which often suffers from structural distortion and unintended content loss—we introduce concept leakage-aware negative guidance scale n_g(t). This mechanism adaptively modulates the guidance strength at each denoising step based on a semantic measurement of concept leakage between the intermediate generated image and the target concept in a joint embedding space. Extensive evaluations on idiosyncratic art styles and object erasure tasks demonstrate that QuP achieves superior concept suppression and content preservation, outperforming state-of-the-art text-based baselines in both image fidelity and erasure precision. The source code will be publicly provided to ensure reproducibility of our work.
PaperID: 7683, Poster
Abstract: The Lipschitz constant of a neural network is a basic stability quantity appearing in optimization guarantees, generalization bounds, and robustness certificates. Deep networks naturally admit two such notions, \emphinput Lipschitzness, measuring sensitivity to data perturbations, and \emphparameter Lipschitzness, measuring sensitivity to weight perturbations. We prove that, for deep networks with normalization layers, with batch normalization as the canonical example, \emphboth Lipschitz constants can grow exponentially with depth and hence with parameter dimension, even under strong per-layer norm control such as \\|W_\ell\\|_2\le 1, and already for linear activations. Thus, a mechanism widely used to stabilize training can create severe worst-case instabilities that are invisible from layerwise norm bounds alone. Parameter-Lipschitz explosion turns Lipschitz-dependent nonsmooth optimization guarantees into exponential-in-dimension bounds for deep normalized networks, while input-Lipschitz explosion makes worst-case Lipschitz-based generalization and robustness certificates vacuous. The latter also yields a concrete rank-separation mechanism for adversarial vulnerability. The theory predicts that perturbations introducing new singular directions should be amplified much more strongly than equal-energy perturbations that remain within the input's original singular subspace. Experiments on MNIST, Fashion-MNIST, and CIFAR-10 support this prediction, showing that rank-creating perturbations cause substantially sharper drops in accuracy, confidence, and margins than same-subspace perturbations.
Abstract: Establishing almost sure convergence rates for stochastic approximation and reinforcement learning under Markovian noise is a fundamental theoretical challenge. We make progress towards this challenge for a class of stochastic approximation algorithms whose expected updates are contractive, a setting that arises in many reinforcement learning algorithms such as Q-learning and linear temporal difference learning. Specifically, for a power-law learning rate \mathcalO(n^-\eta) with \eta \in (1/2, 1), we obtain an almost sure convergence rate arbitrarily close to o(n^1 - 2\eta). For a harmonic learning rate \mathcalO(n^-1), we obtain an almost sure convergence rate arbitrarily close to o(n^-1), which we argue is a strong result because it is close to the optimal rate \mathcalO(n^-1\log\log n) given by the law of the iterated logarithm (for a special case of i.i.d. noise). Key to our analysis is a novel Lyapunov drift construction that applies a Poisson-equation based correction for Markovian noise to the well-established Moreau-envelope smoothing for the contractive mapping.
Abstract: Recent work has shown that well-optimized individual decision trees can match complex black box models in some settings, primarily in noisy domains. For the remaining settings, however, complex ensembled compositions of trees often achieve higher accuracy at the cost of interpretability, leaving practitioners with difficult modeling decisions along an accuracy-interpretability tradeoff. Ideally, we would like to classify as much of the data as possible with one or a small number of trees, achieving interpretability for most samples while maintaining state-of-the-art accuracy. We introduce Multistage Defer Trees: a sequence of sparse decision trees that each make predictions for most samples, while deferring a small proportion to the next tree in the sequence or, ultimately, to a black box. We demonstrate that we can train this model class to match the performance of complex tree-based ensembles while routing most samples through only one or a small number of sparse decision trees. We discuss a range of techniques for training these models while maintaining simplicity. Our method expands the accuracy--interpretability frontier in settings where single-tree methods remain insufficient, demonstrating that even when complex models are necessary, they need not be fully opaque.
Authors: Xinchen Han, Qiuyang Fang, Hossam Afifi, Michel Marot
Abstract: Offline Reinforcement Learning (RL) relies on policy constraints to mitigate extrapolation error, where both the constraint form and constraint strength critically shape performance. However, most existing methods commit to a single constraint family: weighted behavior cloning, density regularization, or support constraints, without a unified principle that explains their connections or trade-offs. In this work, we propose Continuous Constraint Interpolation (CCI), a unified optimization framework in which these three constraint families arise as special cases along a common constraint spectrum. The CCI framework introduces a single interpolation parameter that enables smooth transitions and principled combinations across constraint types. Building on CCI, we develop Automatic Constraint Policy Optimization (ACPO), a practical primal--dual algorithm that adapts the interpolation parameter via a Lagrangian dual update. Moreover, we establish a maximum-entropy performance difference lemma and derive performance lower bounds for both the closed-form optimal policy and its parametric projection. Experiments on D4RL and NeoRL2 demonstrate robust gains across diverse domains, achieving state-of-the-art performance overall.
Abstract: Current neural architecture search (NAS) methods are often limited by their predefined, restrictive search spaces. While recent large language model (LLM)-assisted NAS methods enable open-ended search spaces, they often suffer from inefficient exploration due to biased or low-quality design ideas. To address these issues, we propose to semi-automatically structure model design knowledge to guide the search process. Our approach first defines a high-level structural template of architectural attributes. An LLM then populates this template by analyzing papers, creating a rich and diverse search space that embodies this structured design knowledge. To efficiently explore this vast space, we introduce FairNAD, using a multi-type mutation that enables broad exploration through mutation with fair idea sampling, Pareto-aware mutation, LLM-driven iterative mutation, and a fine-grained feedback loop. We demonstrate the effectiveness of FairNAD in discovering high-performing architectures that yield 0.84, 2.17, and 2.35 points improvement on CIFAR-10, CIFAR-100, and ImageNet16-120, respectively, compared to current state-of-the-art methods.
PaperID: 7688, Poster
Abstract: We provide theoretical guarantees for convergence of discrete-time policy mirror descent with inexact advantage functions updated using temporal difference (TD) learning for entropy regularised MDPs in Polish state and action spaces. We rigorously derive sufficient conditions under which the single-loop actor-critic scheme is stable and convergent. To weaken these conditions, we introduce a variant that performs multiple TD steps per policy update and derive an explicit lower bound on the number of TD steps required to ensure stability. Finally, we establish sub-linear convergence when the number of TD steps grows logarithmically with the number of policy updates, and linear convergence when it grows linearly under a concentrability assumption.
PaperID: 7689, Poster
Authors: Mohsen K Rad, Sotiris Moschoyiannis, Roman Bauer
Abstract: Spiking neural networks (SNNs) offer a compelling path to energy-efficient, biologically plausible computation. Yet, a persistent gap remains: fully spiking models are difficult to train and incompatible with standard deep learning frameworks, while ANN-to-SNN conversion and hybrid approaches sacrifice either biological plausibility or deployment efficiency. We address this gap by introducing SpikeCore, a recurrent spiking unit designed to be natively compatible with ANN training frameworks, differentiable, gradient-friendly, and composable within standard pipelines, while retaining event-driven, spike-based computation. SpikeCore achieves this through learnable membrane decay and spike-history-dependent threshold adaptation, enabling rich temporal dynamics without surrogate gradient instability. Building on SpikeCore, we present a full spatio-temporal pipeline comprising: a two-stage learnable temporal delay hierarchy motivated by differential retinal ganglion latency and axonal conduction delays; a Multi-Branch Squeeze-Excitation (MBSE) backbone that extracts spatial features at multiple scales with channel-wise recalibration, using 11.8× fewer parameters than equivalent residual architectures; and a Parameter-Free Temporal Gain Control (PFTGC) module that modulates each channel's temporal response without added parameters. Evaluated across five diverse benchmarks, the proposed architecture converges 5--6× faster and requires \approx3× fewer timesteps than competitive baselines while matching or surpassing state-of-the-art models on accuracy.
Abstract: Training LLMs to think and reason for longer has become a key ingredient in building state-of-the-art models that can solve complex problems previously out of reach. Recent efforts pursue this in different ways, such as RL fine-tuning to elicit long CoT or scaling latent reasoning through architectural recurrence. This makes reasoning length an important scaling knob. In this work, we identify a novel phenomenon (both theoretically and experimentally): under outcome-only supervision, out-of-distribution (OOD) performance can continue improving as training-time reasoning length (e.g., the token budget in RL, or the loop count in looped Transformers) increases, even after in-distribution (ID) performance has saturated. This suggests that robustness may require a larger budget than ID validation alone would indicate. We provide theoretical explanations via two mechanisms: (i) self-iteration can induce a stronger inductive bias in the hypothesis class, reshaping ID-optimal solutions in ways that improve OOD generalization; and (ii) when shortcut solutions that work for ID samples but not for OOD samples persist in the hypothesis class, regularization can reduce the learned solution's reliance on these shortcuts as the number of self-iterations increases. We complement the theory with empirical evidence from two realizations of scaling training-time reasoning length: increasing the number of loops in looped Transformers on a synthetic task, and increasing token budgets during RL fine-tuning of LLMs on mathematical reasoning, coding, and cross-representation tasks.
Abstract: A central challenge in reinforcement learning (RL) is to learn models that generalize beyond the tasks on which they are trained, a goal traditionally pursued through multi-task and meta RL. Recently, transformer architectures have emerged as a promising approach, enabling adaptation to new tasks via in-context learning without explicit parameter updates. From a functional perspective, a transformer can be viewed as a that maps a context to a task-specific function. It is thus fundamental to understand and design this operator to support stronger generalization in RL. In this work, we address this resulting question of generalization from a kernel-based perspective by establishing a connection between non-linear transformers and kernel-based temporal difference learning. By interpreting the transformer as performing regression in a (RKHS), we show that value functions from different domains can be represented using a shared set of weights, provided they lie within the same RKHS. Experiments on multiple MetaWorld domains support this interpretation, demonstrating convergence of the temporal-difference objective.
Abstract: Score- and flow-matching models often rely on preference-based reinforcement learning for two purposes: aligning with subjective preferences and, surprisingly, recovering properties---such as visual realism and coherent object structure---that matching-based training is intended to learn from the data itself. We argue that this reflects a structural mismatch. Matching losses measure \ell_2 regression error on the velocity or score field under training-time marginals, a proxy poorly aligned with the visual and semantic properties that determine sample quality at inference. Given a reward aligned with these properties, RL sidesteps the mismatch by evaluating the model on its own samples and following the reward landscape directly. The challenge is to obtain such a reward without relying on human references, which are expensive and conflate data realism with annotator inclinations. We propose Discriminator-Guided RL (DRL). DRL trains a discriminator to separate data from base-model samples in a pretrained representation space and uses its logit as the reward in KL-regularized RL. The pretrained space restricts the discriminator to perceptually meaningful directions, and the logit estimates the log-likelihood ratio between data and model, which is the optimal reward for targeting the data distribution. Across SiT, JiT, REPA, and RAE, DRL reduces guidance-free FID (e.g., 9.38 \to 2.62 on SiT) and semantic-space FD (e.g., 88.2 \to 19.3 on DINOv3 for SiT), with consistent gains across all backbones, and improves human-preference rewards without training on them. It also yields a better Pareto frontier between preference reward and image fidelity under subsequent preference-based post-training, increasing alignment while reducing low-level artifacts such as oversaturation and excessive brightness.
PaperID: 7693, Poster
Abstract: It is now a common practice to use PyTorch’s official Fully Sharded Data Parallel (FSDP) framework to train large language models (LLMs) on large-scale GPU clusters, typically together with pipeline parallelism (PP) to achieve better performance. However, existing popular systems like TorchTitan with FSDP and PP are inefficient due to the large bubbles and peak memory footprint caused by suboptimal task scheduling. In this paper, we propose PipeFSDP, an efficient training framework for FSDP combined with PP that alleviates pipeline bubbles and resource contention, thus improve training efficiency. Specifically, we design four fine-grained optimizations: dynamic context-aware prefetching, communication contention mitigation, redundant reshard elimination, and communication order optimization. In addition, we develop a heuristic model that selects the FSDP and PP degrees based on the given hardware and model configurations, thereby enhancing the practicality of PipeFSDP. Empirical evaluations on both dense and sparse LLMs (Llama3, Qwen3, and Qwen3-MoE) across 32-GPU and 64-GPU clusters demonstrate that PipeFSDP significantly outperforms the Torchtitan system by up to 1.28× in Model FLOPs Utilization (MFU) and 1.87× in memory consumption.
Abstract: Knowledge graph learning provides a powerful framework for representing and inferring structured knowledge, with broad practical applications. However, the scarcity of relation-specific labeled triples per entity hinders the training of expressive models, and the ad hoc design of scoring functions limits generalizability and lacks theoretical grounding. We address both issues with a theoretically grounded, end-to-end training framework that extends and subsumes existing methods. Our framework is a two-stage procedure: unsupervised pretraining over heterogeneous corpora followed by supervised learning with diverse relationship types. We establish a nonasymptotic risk bound that disentangles pretraining approximation error from labeled-sample complexity, formally quantifying the benefit of large-scale unlabeled data for downstream knowledge prediction. Synthetic experiments validate each theoretical component, and real-world experiments confirm the effectiveness of our approach on large-scale knowledge graph benchmarks.
PaperID: 7695, Poster
Abstract: Layer-wise depth dependence is known for speech SSL, but it remains unclear how readout depth changes across audio pretraining paradigms and how that choice can support pruning rather than only post hoc probing. We present the Audio Interpretability Atlas, a readout-centered diagnostic study over a wide variety of frozen encoders and complementary sound, music, and speech benchmarks. The atlas connects layer wise transfer, CKA, geometry priors, classical descriptors, sparse-feature concentration, transcoder routing, robustness, steering, and targeted feature ablation on the same encoder-layer-task cells. Its operational goal is to identify the earliest layer or prefix that preserves task evidence, so later blocks can be treated as pruning candidates and, when labels are available, learned layer weighting can be restricted to the selected prefix. A consistent pretraining-associated pattern emerges: audio-text encoders expose compact upper-stage category features, ASR-supervised encoders route speech information densely into mid-to-late layers, and masked or denoising SSL exposes reusable acoustic structure earlier. Final-layer extraction loses at least 10 score points in about half of evaluated encoder-task settings, with the largest single-layer gaps reaching 24--38 points; speech-affect readouts, by contrast, often justify later depth. In low-resource ASR, \textscQuickLayer's label-free geometry mode, based on isotropy and participation ratio, improves 11 of 12 language-encoder settings, averaging 21.3% relative CER reduction, while few-shot probes recover 11--13 points over final-layer extraction. Sparse features, transcoders, perturbations, and interventions then turn selected layers into audit records, identifying when a pruned readout is accurate, stable, localized, and editable.
PaperID: 7696, Poster
Authors: Yangyi Fang, Haolin Shi
Abstract: Training multi-turn LLM agents with reinforcement learning requires distributing a single episode reward across many interaction steps. Recent group-based methods tackle this credit assignment challenge by finding steps where the agent visits the "same" state across different rollouts and comparing their outcomes. However, determining when two states are "the same" remains an open problem: existing methods compare raw text observations via exact string matching or bag-of-words similarity, both of which ignore whether state differences actually matter for future returns. We show that this is a state abstraction problem with a precise bias-variance tradeoff: matching on value-irrelevant features creates unnecessarily small groups (high variance), while ignoring value-relevant features merges states that should be distinguished (high bias). Guided by this analysis, we propose Structured State Abstraction (SSA), a step-level credit assignment method that decomposes observations into typed fields using environment-provided metadata and learns per-field importance weights from return variance. Rather than hard grouping, SSA uses a kernel-smoothed advantage estimator where all states contribute to each step's baseline proportionally to their structured similarity, yielding lower MSE than threshold-based grouping. Experiments on ALFWorld and WebShop show consistent improvements over GiGPO and HGPO across both benchmarks and model scales; notably, kernel smoothing eliminates the singleton problem that causes existing methods to lose step-level credit for 30-36% of steps. Our implementation is provided in the supplementary materials and will be open-sourced upon acceptance.
Authors:
Chih-Hsuan Yang, Jingyan Jiang, Vikram Vasudevan, Cheng-Hau Yang, Huihuo Zheng, Benoit Cote, Le Chen, Eliu Huerta, Venkatram Vishwanath, Ian Foster, Rajeev ThakurAbstract: Many recent math- and science-oriented agent systems adopt hierarchical designs with specialized reviewer roles, motivated by the idea that routing critique through a dedicated review stage should help turn wrong candidates into correct ones on hard problems. We study that expectation on 4,181 verifier-grounded Omni-MATH problems, a hard ten-tier benchmark with enough headroom to separate protocols, using matched gpt-oss-120b actors as the primary actor family. On tiers 1-2, multi-agent collaboration adds at most about 2 percentage points over the matched single-agent baseline. From tier 4 onward, the gains open sharply, reaching about 10 to 20 percentage points on tiers 6-9. In that regime, broadcast-style peer discussion attains higher final accuracy than a planner-executor-reviewer pipeline (PER). We use that divergence as a starting observation and ask when reviewer quality translates into effective solver updates. On this benchmark, the PER-broadcast accuracy gap is not accounted for by reviewer precision alone. Here reviewer precision asks how often a reviewer warning truly points to a real error. PER's reviewer has higher precision (0.861 vs. 0.644), yet correct critique is much less likely to change the next candidate the protocol carries forward, and reviewer-guided repair is correspondingly lower. These results show that reviewer detection quality and critique uptake are empirically separable. Within matched PER interventions, forcing explicit acknowledgment lowers final accuracy instead of improving follow-through, while placing reviewer guidance directly in the solver's working context partially improves follow-through without closing the gap. Taken together, these interventions point in the same direction: critique appears more likely to be acted upon when it is presented more directly in the solver's working context, although this evidence is directional rather than causal. Under reviewer-centric evaluation, a system can look strong at spotting errors yet still fail to solve more problems if the protocol does not act on those critiques.
PaperID: 7698, Poster
Abstract: Does language modeling on image-text data truly improve vision, or does it merely adapt visual features for language alignment? Despite rapid progress in multimodal large language models (MLLMs), this question remains poorly understood. We present the first systematic analysis of visual representations in MLLMs by probing popular model families across a broad suite of dense visual tasks and comparing each MLLM to its exact pre-MLLM vision encoder. We find that language modeling largely improves visual representations for semantic tasks but can degrade performance on fine-grained geometric tasks, such as monocular depth. We further show that scaling the language model consistently improves the quality of visual representations. Given the improved MLLM representations, we next examine how to best extract and aggregate these features. Since no single layer provides a universally strong representation and concatenating features across layers is impractical, we propose a lightweight adapter that efficiently combines features across MLLM layers, turning frozen MLLMs into competitive visual backbones. Finally, we study whether MLLM representations can improve downstream applications, including text-to-image generation and inverse dynamics modeling, and show that they can outperform pre-MLLM visual features as general-purpose representations. Together, our results provide a representation-centric perspective for understanding and leveraging MLLMs as visual foundation models.
PaperID: 7699, Poster
Abstract: Federated Continual Learning (FCL) suffers from spatial-temporal catastrophic forgetting caused by sequential task learning on individual clients and the aggregation of heterogeneous client knowledge. However, most existing methods either overemphasize preserving old knowledge, which restricts model plasticity for new tasks, or rely on training complex generative models for replay, incurring substantial computational overhead. In the human immune system, immune memory is preserved implicitly and can be reactivated through vaccination even after long periods of dormancy. Motivated by this, we propose , a novel framework that explicitly permits forgetting to preserve high plasticity for new tasks, while efficiently recovering past knowledge through re-learning a compact vaccine set. Specifically, the vaccine set for each task is synthesized by compress task gradient information under the guidance of the fine-tuned local model, capturing representative data characteristics and task-specific discriminative patterns. Two novel components are also introduced: , which encodes task-specific knowledge in a compact spectral form, priming the model with essential information and enabling efficient memory boosting via vaccine set replay after forgetting; and , which aggregates vaccine sets from all clients to train a globally generalized model without accessing any raw private data. Extensive experiments on three benchmarks demonstrate that FedVaccine achieves competitive performance, enables rapid knowledge recovery after forgetting, and incurs no additional overhead from training generative models.
PaperID: 7700, Poster
Abstract: Spiking Neural Networks (SNNs) offer a promising energy-efficient alternative to Artificial Neural Networks (ANNs) through event-driven, multiplication-free computation. However, when applied to downstream tasks such as Neural Radiance Fields (NeRF), SNNs suffer from significant information loss, resulting in a noticeable performance gap. In this paper, we present Spik-NeRF v2 to advance SNN-based neural rendering. We propose the \pmI-LIF spiking neuron, which extends I-LIF to support signed integer values during training, thereby mitigating information loss. During inference, it is converted to ternary spikes \-1, 0, 1\, preserving event-driven properties with addition-only operations. Furthermore, we introduce a re-parameterization technique that transforms a trained I-LIF-based Spik-NeRF with t timesteps into an equivalent \pmI-LIF-based model with t/2 timesteps. This enables faster inference while preserving rendering quality, overcoming the suboptimal performance of directly trained \pmI-LIF models with reduced timesteps. Extensive experiments on synthetic and realistic datasets demonstrate that Spik-NeRF v2 surpasses existing SNN-based NeRF methods and achieves rendering quality comparable to ANN-based approaches.
PaperID: 7701, Poster
Authors: Zeyu Wang, Jiayu Wang, Haiyu Song, Haoran Duan
Abstract: Multimodal image fusion (MMIF) aims to integrate complementary information from different modalities into a high-quality fused image and support downstream tasks. Recently, feature decomposition has become an important paradigm by separating source images into common and modality-specific unique features. However, existing methods lack clear supervision because ground-truth (GT) decomposition feature maps are unavailable. They usually combine multiple image-level metrics as losses, which are inherently incomplete and may conflict since each pixel couples attributes such as texture, edge, and contour. To address this, we propose a 1D signal-level self-supervised feature decomposition paradigm. Our core insight is to reformulate feature decomposition from unclear 2D image-level supervision into an integral-driven 1D signal-level optimization problem. This objective-level reformulation uses the 1D signal form to compute the integral constraint. The decomposer is optimized by the integral area between common and original signals, enabling more stable optimization with a clear optimization objective. Our model follows a two-stage SSL framework. Stage I designs dual pretext tasks for integral-driven decomposition at the signal level and structure-preserving reconstruction at the image level. Stage II fuses unique features and combines them with common features to reconstruct the fused image. Experiments on representative MMIF tasks show state-of-the-art(SOTA) performance. Our code will publicly available.
PaperID: 7702, Poster
Abstract: Structured video generation (e.g., autonomous driving) faces a fundamental domain gap between simple human control and dense physical reality. Text-to-Video models accept accessible inputs but struggle with precise spatial guidance, while dedicated autonomous driving world models require expensive, dense 3D annotations that are difficult for users to provide. To bridge this gap, we introduce Motion Forcing, a structured video generation framework that achieves precise multi-agent control using extremely sparse 2D inputs. Our key insight is to explicitly bridge human intent and visual synthesis via a hierarchical ``Point-Shape-Appearance'' paradigm. This approach decomposes generation into verifiable stages: modeling user intent as sparse geometric points, expanding them into dynamic depth maps to explicitly resolve 3D geometry and decouple ego-motion from object dynamics, and finally rendering high-fidelity textures. Furthermore, to elevate the model from passive instruction-following to active physical reasoning, we employ a Masked Point Recovery strategy. By forcing the reconstruction of dynamic depth from occluded trajectories, the model internalizes latent physical laws, enabling causal inference. Extensive experiments demonstrate that Motion Forcing significantly outperforms state-of-the-art baselines on large-scale autonomous driving benchmarks. Additional validation in rigid-body physics and robotic manipulation confirms the structural integrity and generalizability of our framework. To facilitate future research, we have released our model weights and code.
PaperID: 7703, Poster
Abstract: Vision-language models (VLMs) have achieved remarkable progress. However, processing high-resolution images and long videos generates a massive number of vision tokens, creating a severe inference bottleneck due to the quadratic complexity of attention mechanisms. Previous efforts to mitigate this issue have primarily focused on accelerating the prefill stage by reducing the number of vision tokens, often yielding limited practical speedups while neglecting the increasingly time-consuming decoding stage when the generated response is long. In addition, existing approaches heavily rely on attention maps or similarity maps that are not compatible with efficient attention implementations. To comprehensively address these limitations, we first propose FlashPruner, a lightweight module that selects the most helpful vision tokens without relying on attention maps, significantly reducing prefill time. For the decoding stage, we adopt self-speculative decoding and theoretically prove that on-policy distillation guarantees a lower bound for the speculative acceptance rate. Guided by this formulation, we introduce LoFT, an on-policy distillation method that effectively trains a highly reliable and efficient draft model including FlashPruner and early layers of the language model. Extensive experiments demonstrate that after training with LoFT, FlashPruner achieves substantial inference speedups while preserving task performance. Notably, on MVBench, FlashPruner improves the time-to-first-token by 1.5× and decoding speed by 2×, maintaining 96.8% relative accuracy with only 10% vision tokens.
PaperID: 7704, Poster
Abstract: Denoising score matching (DSM) is the cornerstone of scalable diffusion model training. However, DSM becomes inapplicable on curved manifolds: the curved geometry makes the forward diffusion process nonlinear and the transition kernel is no longer tractable. Existing approaches therefore revert to implicit score matching (ISM) with expensive Brownian simulation, approximate the intractable heat kernel, or bypass the transition kernel altogether via flow matching with the logarithm map, each incurring significant computational cost or approximation error on general manifolds. We propose \emphsplitting diffusion, a new framework that restores DSM-like training on general Riemannian manifolds by replacing the intractable Brownian forward process with a splitting scheme that alternates two steps: (1) an Ornstein-Uhlenbeck velocity update in the tangent space, whose transition kernel has a closed-form solution; (2) a deterministic geodesic transport step. This splitting yields a fully tractable transition kernel locally, enabling DSM-style training while requiring only the exponential map as a geometric oracle. The effectiveness of splitting diffusion is demonstrated on complicated manifolds and high-dimensional data.
PaperID: 7705, Poster
Abstract: Films introduce characters incrementally as they progress. We study the problem of identifying principal characters and assigning their names directly from the movie itself, without using external cast lists, actor photographs, or pre-built character banks. We introduce the character naming task: given a name mentioned in dialogue and the associated video context, the model must determine which visible character bears that name. This task is challenging because name resolution often depends on cinematic cues such as shot–reverse-shot structure, self-introduction, direct address, third-party reference, gaze, and temporal continuity. We provide a dataset and annotations for evaluating character naming and first character appearances, develop a video-and-dialogue model for the task, and compare it with proprietary multimodal large language models. We further introduce a progressive framework for building a character bank as the film unfolds. Together, these contributions support online, human-like movie understanding in which character identities are inferred from the film itself.
Authors: Leyla Roksan Caglar, Pedro A.M Mediano, Baihan Lin
Abstract: Humans and modern vision models can reach similar classification accuracy while making systematically different kinds of mistakes - differing not in how often they err, but in who gets mistaken for whom and in which direction. We show that these directional confusions reveal distinct inductive biases that are invisible to accuracy alone. Using matched human and deep vision model responses on a natural-image categorization task under 12 perturbation types, we quantify asymmetry in confusion matrices and link it to the shape of the information–error trade-off. We characterize this trade-off geometry using three signatures extracted from the rate–distortion (RD) frontier: slope (beta), curvature (kappa) and efficiency (AUC). We find that humans exhibit broad but weak asymmetries, whereas deep vision models show sparser, stronger directional collapses. Robustness training reduces global asymmetry but fails to recover the human-like breadth–strength profile of graded similarity. Mechanistic simulations further show that different asymmetry organizations shift the RD frontier in opposite directions, even when matched for performance. Together, these results position directional confusions and RD geometry as compact, interpretable signatures of inductive bias under distribution shift.
PaperID: 7707, Poster
Authors: Arun Ponnambalam, Krishnan V Pottore
Abstract: The ventral visual stream is widely modeled as a serial feedforward hierarchy in which V1, V4, and IT population codes develop sequentially during object recognition. We ask whether a second, concurrent coding mode exists—one organized not by anatomical order but by joint population structure across areas. Using Partial Information Decomposition applied to simultaneous multielectrode spiking recordings across all three areas at millisecond resolution—the first simultaneous three-area spiking PID analysis of the primate ventral stream—in two macaque monkeys viewing 25,000+ natural images, we decompose population coding into serial (unique per area) and synergistic (joint across areas) components at 5 ms resolution across five CNN target representations spanning low-level spatial features to high-level object identity. Three findings replicate across both animals and all five representations. First, synergistic inter-area coupling emerges before IT carries any unique object-related information—a dissociation of 15–65 ms that replicates in direction without exception across both animals—such that the joint population integrates before the apex encodes; moreover, V1–IT synergy persists for over 120 ms after V1’s unique information reaches zero. Second, although V1↔IT and V1↔V4 coupling emerge simultaneously and rise in parallel, V1↔IT exhibits stronger peak synergy at mid-to-high-level targets in both animals, suggesting a dominant role for non-serial joint coding. Third, when V1 and V4 are treated as an integrated feedforward block, their synergistic coupling with IT emerges last across all tested conditions—the feedforward foundation is the final component to join the synergistic mode, not the first. Together, these results show that serial and synergistic population codes co-occur in the same recordings, overlap in time, but follow different organizational principles, Providing a new level of nuance in our understanding of the primate ventral stream and introducing concrete constraints for biologically grounded models of vision.
PaperID: 7708, Poster
Abstract: Conventional real-valued neural networks struggle to explicitly capture phase-frequency coupling, while long-horizon time series forecasting requires models to characterize non-stationary amplitude variations, phase shifts, and frequency-dependent temporal evolution. This mismatch limits the stability and expressiveness of neural models on complex forecasting tasks. To address this challenge, we propose SDHilb, a structure-preserving forecasting framework that lifts real-valued observations into an adaptive complex state space and models temporal evolution as a process jointly governed by amplitude and phase information. Through the Hilbert transform, phase-related structures are explicitly encoded, leading to a coupled evolution of real and quadrature components that motivates Schrödinger-Hamiltonian structured dynamics in the lifted complex space. Based on this formulation, SDHilb integrates Schrödinger-guided linear dynamics with an adaptive multi-scale Hilbert transform and introduces three key components: (1) a structure-preserving linear dynamic module with symplectic discretization, (2) an adaptive multi-scale Hilbert module for constructing data-dependent quadrature components and refined time-frequency decomposition, and (3) a theoretically grounded parameter reuse mechanism derived from the closed-form recurrence structure. Extensive experiments demonstrate that SDHilb achieves competitive or superior performance on both short- and long-term forecasting tasks while reducing horizon-specific parameterization through recurrence-guided parameter reuse.
PaperID: 7709, Poster
Abstract: Vision-Language-Action (VLA) models often struggle with precise manipulation due to their lack of explicit spatial awareness. To bridge this gap, we propose GIVLA, a framework that internalizes geometric priors into the VLA backbone through a deep-coupled architecture. GIVLA implements a two-stage gradient-informed training paradigm to resolve task-level interference and ensure precise execution, a design grounded in the perception-action duality identified through our analysis of training dynamics. Extensive experiments on LIBERO and physical robot platforms demonstrate that GIVLA achieves superior accuracy and robustness with high parameter efficiency.
Abstract: In-context reinforcement learning (ICRL) leverages the in-context learning capabilities of transformer models (TMs) to efficiently generalize to unseen sequential decision-making tasks without parameter updates. However, existing ICRL methods rely on explicit reward signals during pretraining, which limits their applicability when rewards are ambiguous, hard to specify, or costly to obtain. To overcome this limitation, we propose a new learning paradigm, \emphIn-Context Preference-based Reinforcement Learning (ICPRL), in which both pretraining and deployment rely solely on preference feedback, eliminating the need for reward supervision. We study two variants that differ in the granularity of feedback: \emphImmediate Preference-based RL (I-PRL) with per-step preferences that generalize dueling bandits, and \emphTrajectory Preference-based RL (T-PRL) with trajectory-level comparisons. We first show that supervised pretraining, a standard approach in ICRL, remains effective under preference-only context datasets, demonstrating the feasibility of in-context solving unseen tasks using only preference signals. To further improve data efficiency, we introduce alternative preference-native frameworks for I-PRL and T-PRL that directly optimize TM policies from preference data without requiring reward signals nor optimal action labels. Experiments on dueling bandits, navigation, and continuous control tasks demonstrate that ICPRL enables strong in-context generalization to unseen tasks, achieving performance comparable to ICRL methods trained with full reward supervision.
Abstract: The key-value (KV) cache dominates memory bandwidth and footprint in long-context autoregressive inference. Recent rotation-preconditioned codecs (TurboQuant, PolarQuant) show that a structured random rotation followed by a per-coordinate scalar quantizer matched to an analytically tractable marginal is a near-optimal recipe for KV compression. OCTOPUS advances this paradigm through joint quantization of rotated coordinate triplets. Each triplet's direction is mapped to a square via an octahedral parameterization, and the two resulting coordinates and the triplet norm are Lloyd-Max quantized against implementation-matched marginals. Optimizing the per-triplet squared error gives a strictly non-uniform bit allocation depending only on the total dimensionality of the keys. We find the finite-dimensional quality optimum with sweeps to be constant on every real decoder we test. The codec is data-oblivious, online, and deterministic given a seed. Across text, video, and audio, OCTOPUS matches or beats every prior rotation codec at every reported bit width and metric, with a lead that grows as bits drop for extreme compression. Furthermore, a fused Triton implementation reconstructs keys on the fly without materializing the uncompressed key, so the codec adds no decode-time bandwidth or latency over the existing dequantization.
Authors: Kyungwon Jeong, Won-Gi Paeng, Honggyo Suh
Abstract: Weight matrices in deep networks exhibit geometric continuity---principal singular vectors of adjacent layers point in similar directions. While this property has been widely observed, its origin remains unexplained. Through experiments on toy MLPs and small transformers, we identify two mechanisms: residual connections create cross-layer gradient coherence that aligns weight updates across layers, and symmetry-breaking nonlinearities constrain all layers to a shared coordinate frame, preventing the rotation drift that would otherwise destabilize weight structure. Crucially, a nonlinear but rotation-preserving activation fails to retain continuity, isolating symmetry breaking---not nonlinearity itself---as the active ingredient. Activation and normalization play distinct roles: activation concentrates continuity in the leading singular direction, while normalization distributes it across multiple directions. In transformers, continuity is \emphprojection-specific: Q, K, Gate, and Up (which read from the residual stream) develop input-space (\mathbfv_1) continuity; O and Down (which write to it) develop output-space (\mathbfu_1) continuity; V alone, lacking an adjacent nonlinearity, develops only low continuity.
PaperID: 7713, Poster
Abstract: Workload forecasting is a fundamental task to realizing electricity-computing synergy in modern data centers, where operational planning must account for the risk induced by fluctuating computing demand. Yet production workloads exhibit substantial volatility, making deterministic point forecasts insufficient for uncertainty-aware decision making. Existing methods mainly model workload uncertainty from temporal correlations in utilization traces, overlooking the job-induced cross-machine dependencies shaped by the scheduling layer. We empirically find that workload volatility exhibits structured and time-varying cross-machine patterns whose strength aligns with job co-occurrence. This observation motivates GD-Diff, a job-aware graph-conditioned diffusion framework for probabilistic workload forecasting. GD-Diff constructs a job-induced weighted dynamic graph from scheduling logs, transforming job co-occurrence into time-varying inter-machine dependencies. Built on this graph, GD-Diff uses a dual-branch noise prediction network within the diffusion process, where the temporal branch captures regular temporal evolution and the graph branch models job-induced dynamics. Experiments on Alibaba and Google cluster traces show that GD-Diff achieves state-of-the-art performance in both point and probabilistic forecasting, and ablation studies further validate the job-induced dynamic graph and dual-branch architecture. The code is available at \urlhttps://anonymous.4open.science/r/GD-Diff-F5E8/.
PaperID: 7714, Poster
Authors: Rohan Maskrey
Abstract: Reservoirs, equilibrium models, neural ODEs, and recurrent test-time compute all perform inference by running a state trajectory. Standard practice collapses that trajectory to a terminal readout, implicitly assuming that the final state is the right summary of the computation. We show that this assumption can fail: in supercritical nonlinear dynamics, terminal representations can degrade while earlier trajectory prefixes still support accurate prediction. We formalise the usable prefix as an accuracy-defined grace period and compare fixed early-window readout, norm-ratio stopping, and GRACE (Growth-Regulated Adaptive State Extraction). In complex modReLU reservoirs, non-terminal readouts preserve near-stable MNIST accuracy across an unstable spectral sweep while terminal readout collapses. The best decoder is nonlinearity-dependent: GRACE is strongest in fixed unbounded dynamics and under injected dynamical noise, while simple early-window readout is strongest in trained LeakyReLU and LSTM checks. The same endpoint-vs.-trajectory gap appears on CIFAR-10, including a frozen-feature variant where it survives translation to an 85%-accuracy substrate, and cross-system comparisons show that the regime is severe in unbounded nonlinear dynamics, milder in bounded tanh systems, and absent in linear systems. Terminal failure, therefore, is not necessarily computational failure; it can be a failure of when we choose to read the system out.
PaperID: 7715, Poster
Abstract: We present AnyEdit, a unified framework for speech and singing voice editing in real-world acoustic environments. Recent unified generative models already achieve strong speech and singing generation in clean conditions, but real-world vocal editing remains challenging because the model must revise content while simultaneously preserving speaker identity, prosody, and the surrounding acoustic scene. A straightforward solution is to directly model noisy edited audio end to end, yet this entangles semantic editing with environmental acoustics and can weaken the core vocal generation capability. We therefore adopt a two-stage design that preserves clean-domain editing ability while deferring acoustic rendering to a separate stage. First, an instruction-guided autoregressive editor predicts clean prosody and content-style tokens. To make this stage robust to degraded inputs, we introduce a teacher-student noise-robust prosody tokenizer that distills clean prosodic tokens from noisy recordings. Second, a flow-matching acoustic model performs in-context acoustic rendering conditioned on the source acoustic context, enabling the edited segment to inherit surrounding noise, reverberation, and background sound while maintaining speaker timbre consistently. To further stabilize the unedited regions, we introduce a unmodified-area-aware supervision mechanism. After pre-training, we also apply direct preference optimization (DPO) using real singing editing pairs mined via dynamic time warping to achieve better singing editing quality. To enable realistic evaluation, we introduce AnyEditBench, the first benchmark covering both speech and singing voice editing under diverse acoustic conditions. Experiments on Chinese and English speech and singing datasets showed that AnyEdit consistently outperforms strong baselines in content accuracy, subjective quality, and background consistency under noisy and reverberant conditions. Demo audios can be found at https://any-edit.github.io.
Authors:
Jishnu S Nair, Patrice Bechard, Rishabh Maheshwary, SRAVAN RAMACHANDRAN, Surajit Dasgupta, Aakash Bhagat, Shruthan Radhakrishna, Pulkit Pattnaik, Johan Obando Ceron, Shiva Krishna Reddy Malay, Sagar Davasam, Seganrasan Subramanian, Vipul Mittal, Sridhar K Nemala, Chris Pal, Srinivas Sunkara, Sai Rajeswar MudumbaAbstract: World models enable agents to anticipate the effects of their actions by internalizing environment dynamics. In enterprise systems, however, these dynamics are often defined by tenant-specific business logic that varies across deployments and evolves over time, making models trained on historical transitions brittle under deployment shift. We ask a question the world-models literature has not addressed: when the rules can be read at inference time, does an agent still need to learn them? We argue, and demonstrate empirically, that in settings where transition dynamics are configurable and readable, runtime discovery complements offline training by grounding predictions in the active system instance. We propose enterprise discovery agents, which recover relevant transition dynamics at runtime by reading the system's configuration rather than relying solely on internalized representations. We introduce CascadeBench, a reasoning-focused benchmark for enterprise cascade prediction that adopts the evaluation methodology of World of Workflows on diverse synthetic environments, and use it together with deployment-shift evaluation to show that offline-trained world models can perform well in-distribution but degrade as dynamics change, whereas discovery-based agents are more robust under shift by grounding their predictions in the current instance. Our findings suggest that, in configurable enterprise environments, agents should not rely solely on fixed internalized dynamics, but should incorporate mechanisms for discovering relevant transition logic at runtime.
PaperID: 7717, Poster
Abstract: Reconstructing audio from brain activity remains challenging. A key limitation lies in current audio diffusion models, which are primarily designed for text-based conditioning and lack mechanisms to incorporate non-linguistic continuous signals such as neural responses. While recent work in the visual domain has explored direct conditioning on brain activity, this approach has not yet been extended to audio diffusion models. We introduce a framework that directly guides a latent audio diffusion model with fMRI signals, avoiding intermediate feature prediction for the conditioning stage. Central to our approach is a transformer-compatible adapter that enables the integration of non-linguistic representations into the cross-attention layers of diffusion transformers (DiTs). Although motivated by neural decoding, this adapter provides a general mechanism for conditioning audio diffusion models on arbitrary continuous inputs, not inherently limited to neural signals. We demonstrate consistent improvements over baselines, including two-stage fMRI-to-latent conditioning, across both acoustic and semantic evaluation metrics. In addition, we introduce a fidelity metric grounded in a computational model of auditory neural processing, which allows us to quantify the contribution of individual auditory cortical regions to reconstructed audio representations.
Abstract: Large language models (LLMs) often fail to reason under temporal cutoffs: when prompted to answer from the standpoint of an earlier time, they exploit knowledge that became available only later. We study this failure through the lens of ex-ante reasoning, where a model must rely exclusively on information knowable before a cutoff. Through a systematic analysis of prompt-level interventions, we find that temporal leakage is highly sensitive to cutoff formulation and instruction placement: explicit cutoff statements outperform implicit historical framings, and prefix constraints reduce leakage more effectively than suffix constraints. These findings indicate that prompting can steer models into a temporal frame, but does not endow them with the ability to verify whether a response is temporally admissible. We further argue that supervised fine-tuning is insufficient, since ex-ante correctness is not an intrinsic property of an answer, but a relation between the answer and the cutoff. To address this gap, we propose TCFT, a \underlineTemporal \underlineCritique \underlineFine-\underlineTuning framework that trains models to acquire cutoff-aware temporal verification. Given a query, a cutoff, and a candidate response, TCFT teaches the model to identify post-cutoff leakage, explain temporal boundary violations, and judge temporal admissibility. Experiments with Qwen2.5-7B-Instruct and Qwen2.5-14B-Instruct show that TCFT consistently outperforms prompting and SFT baselines, reducing average leakage by 41.89 and 37.79 percentage points, respectively.
Abstract: Self-attention has greatly contributed to the success of the widely used Transformer architecture by enabling learning from data with long-range dependencies. In an effort to improve performance, a gated attention model that leverages a gating mechanism within the multi-head self-attention has recently been proposed as a promising alternative. Gated attention has been empirically demonstrated to increase the expressiveness of low-rank mapping in standard attention and even to eliminate the attention sink phenomenon. Despite its efficacy, a clear theoretical understanding of gated attention's benefits remains lacking in the literature. To close this gap, we rigorously show that each entry in a gated attention matrix or a multi-head self-attention matrix can be written as a hierarchical mixture of experts. By recasting learning as an expert estimation problem, we demonstrate that gated attention is more sample-efficient than multi-head self-attention. In particular, while the former needs only a polynomial number of data points to estimate an expert, the latter requires exponentially many data points to achieve the same estimation error. Furthermore, our analysis also provides a theoretical justification for why gated attention yields higher performance when a gate is placed at the output of the scaled dot product attention or the value map rather than at other positions in the multi-head self-attention architecture.
PaperID: 7720, Poster
Abstract: Mixture-of-Experts (MoE) large language models (LLMs) suffer significant inference and storage overheads, which makes extremely low-bit quantization (sub 2-bit) highly desirable. However, existing methods for dense LLMs typically quantize each expert independently and require it to approximate the original mapping over the entire input space. We argue that this formulation is overly conservative and structurally mismatched for MoE: 1) experts in the same MoE layer are not fully independent but operate on a shared hidden representation space. Under low-bit budgets, independently binarizing each expert will distort the activation-sensitive input geometry, which leads to severe performance degradation, and 2) due to routing-induced specialization, each expert only serves a concentrated subset of tokens, whose activations typically occupy a much narrower subspace than that of dense layers. More importantly, when the weights approach 1-bit, the dominant deployment bottleneck is no longer the weights themselves, but the metadata (e.g., scaling factors, bitmap, grouping index) which can raise the real effective inference precision to around 3-bit. This issue, however, has largely been overlooked in prior works. Building on these perspectives, we introduce a novel extremely low-bit quantization framework for MoE-LLMs, called TriS-MoE, which considers Structure, Subspace and System. Specifically, we first extract a shared high-precision input-side backbone to preserve activation-sensitive input geometry, while binarizing the remaining components to save costs. Secondly, we propose to redirect expert quantization errors into the null space of routed activations to minimize output perturbation on the subspace actually used by each expert. Finally, we co-design a bitmap compression method based on Golomb-Rice coding which approaches the Shannon entropy lower bound and a specialized streaming loading mechanism to reduce metadata overhead during inference. Extensive experiments on MoE-LLMs demonstrate our TriS-MoE outperforms the strongest baseline by 8.31% average accuracies, and also achieves an average 6.32× reduction in inference memory on real systems, enabling Qwen3.5-35B-A3B and Mixtral-8×7B to run on a single consumer-GPU.
PaperID: 7721, Poster
Abstract: Generating 3D assets with explicit parts requires two properties at once: geometric detail inside each part and structural coherence across the whole object. Recent image-conditioned 3D generation models adopt a \emphstructured 3D representation that anchors tokens to explicit spatial positions, capturing both properties, but their output remains a single fused mesh. We ask how a part-decomposed generation model can inherit both properties of these holistic priors, and answer with STRUCT-PARTS, a framework whose core technical contribution is a dual-frame coordinate interleaving mechanism: each shape latent is interpreted under both a part-local and an object-global reference frame, and processed through a global-local interleaved transformer that simultaneously reuses the prior's geometric detail at the part scale and its structural coherence at the whole-object scale. As its structural input, STRUCT-PARTS uses a segmented mesh scaffold, a coarse mesh paired with a face-level segmentation mask kept entirely within the prior's native 3D representation. With only lightweight finetuning, both intra-part detail and inter-part coherence emerge from the same pretrained weights. On PartObjaverse-Tiny, STRUCT-PARTS matches or surpasses prior part-decomposed models at a fraction of their training cost, while naturally supporting part-mesh refinement and part-level editing.
PaperID: 7722, Poster
Authors:
Xiaokang Wang, Zihe Liu, Xinghan Qin, Zihao Yin, Jidong Yuan, Qinxuan Zhang, Xiushuo HuAbstract: Deep learning has emerged as a powerful paradigm for constructing nonlinear asset pricing factor models. However, temporal distribution shifts such as bull-bear transitions and industry rotations undermine the generalization of models trained under the i.i.d. assumption. As a result, existing models tend to exploit environment-specific correlations that are highly predictive in-sample but unstable across market regimes. To address this challenge, we propose CASH, a causality-inspired asset pricing factor framework designed to learn invariant and minimally sufficient representations under temporal distribution shifts. CASH formulates an information-theoretic adversarial learning objective that systematically discards transient market noise by optimizing against an adversary designed to exploit spurious correlations, thereby distilling invariant factor-return relationships. The resulting invariant representation is then integrated into a conditional factor pricing model via a dynamic factor exposure network. Experiments on real-world stock datasets across multiple markets demonstrate that CASH consistently outperforms state-of-the-art baselines in out-of-sample return prediction and exhibits superior robustness under pronounced temporal regime shifts.
Abstract: Reconstructing dynamic scenes from Vehicle-to-Infrastructure Cooperative Autonomous Driving (VICAD) data is fundamentally complicated by temporal asynchrony: vehicle and infrastructure cameras operate on independent clocks, capturing the same dynamic agent such as cars and pedestrians at different physical times. Existing Gaussian Scene Graph methods implicitly assume synchronized observations and assign a single pose per agent per frame, which is an assumption that breaks in cooperative settings, where the resulting gradient conflicts cause severe ghosting on dynamic agents. We identify this as a representation-level failure, not an optimization artifact: we prove that any single-timeline formulation incurs an irreducible photometric loss scaling quadratically with agent velocity and cross-source time offset. To resolve this, we propose Dust (DecoUpled Spatio-Temporal) Gaussian Scene Graph for 4D Cooperative Driving Reconstruction. DUST Gaussian Scene Graph shares a canonical Gaussian set per agent for appearance consistency, while maintaining decouple pose trajectories aligned to each source's true capture timestamps. We prove that this decoupling enables the pose-gradient kernel block-diagonal, eliminating cross-source interference entirely. To make Dust practical, we further introduce a static anchor-based pose correction pipeline that corrects spatio misalignment between vehicle and infrastructure annotations, and a pose-regularized joint optimization scheme that prevents trajectory jitter and drift during early training. On 26 sequences from V2X-Seq, DUST achieves state-of-the-art performance, improving dynamic-area PSNR by 3.2 dB over the strongest baseline and reducing Fréchet Video Distance by 37.7%, with keeping robustness under larger temporal asynchrony. Code is available at https://anonymous.4open.science/r/DUST-6A55.
PaperID: 7724, Poster
Authors: Arya Farahi, Ritwik Vashistha
Abstract: Real-world machine learning systems often operate on imperfect measurements: sensors are noisy, annotations are produced by non-experts, and scientific observations are often indirect. In these settings, classic distribution comparison tools such as maximum mean discrepancy (MMD) can be systematically biased, matching artifacts of the measurement process rather than the underlying mechanisms of interest. We study a practical regime in which noisy observations are abundant, but only a small audited subset contains ground-truth measurements. We propose measurement-error-corrected MMD (MEC-MMD), an audit-based estimator that debiases kernel evaluations without requiring a noise model. MEC-MMD is exactly unbiased for the oracle MMD under simple random auditing, admits an explicit variance decomposition, and is consistent for the population MMD, achieving the standard O_p(N^-1/2) rate when the audit size scales with the dataset. We demonstrate the utility of distributional matching with MEC-MMD loss for both parameter estimation and unsupervised applications, showing that MEC-MMD consistently outperforms MMD and alternative methods. These results provide a principled foundation for robust distribution comparison, simulation-based inference, and learning in noisy real-world machine learning pipelines.
PaperID: 7725, Poster
Authors:
Jiale Dai, Hongcan Deng, Liuxian Ma, Xiaoke Niu, Guojie SongAbstract: Activation steering can change an LLM's behavior without updating its weights, but a direction intended to change safety or value stance often also shifts topic, phrasing, and task content. We study a narrower and directly testable question: can a frozen residual-stream state be equipped with an editable interface that separates semantic content from value framing well enough to reduce this collateral damage? We propose a lightweight dual-code interface with a one-way semantic\rightarrowvalue path. Semantic information may ground value recognition, while stop-gradient gating, topic de-confounding, and swap consistency discourage value supervision from rewriting the semantic factor. At inference time, we edit only the value code through a residual delta update. Across two instruction-tuned backbones, the learned interface improves the steering--damage trade-off over dense and sparse steering baselines: it preserves semantics better, reduces topic leakage and benign false refusals, and remains stable under noise, distribution shift, and recomposer controls.
Authors: Ruoqi Wang, Haitao Wang, Shaojie Guo, Qiong Luo
Abstract: Out-of-domain (OOD) robustness is challenging to achieve in real-world computer vision, especially in unsupervised domain adaptation scenarios, where shifts in image background, style, and acquisition instruments often degrade model performance. Generic augmentations show inconsistent gains under such shifts, whereas dataset-specific augmentations require expert knowledge and prior analysis. Moreover, prior studies show that neural networks adapt poorly to domain shifts because they exhibit a learning bias to domain-specific frequency components. Perturbing frequency values can mitigate such bias but overlooks pixel-level details, leading to suboptimal performance. To address these limitations, we propose D-GAP, a Dataset-agnostic and Gradient-guided augmentation method for the Amplitude spectrum (in frequency space) and the Pixel values. Unlike conventional handcrafted augmentations, D-GAP computes sensitivity maps in the frequency space from task gradients, which reflect how strongly the deep models respond to different frequency components, and uses the maps to adaptively interpolate amplitudes between source and target samples. We further propose a dual-space augmentation that jointly controls spectral bias and spatial fidelity by introducing a complementary pixel-space blending branch. This way, D-GAP turns augmentation from fixed, random, or manually designed perturbation into a model-response-adaptive intervention. Extensive experimental results show that the proposed method consistently outperforms both generic and dataset-specific domain adaptation methods, improving average OOD performance by +5.3% on four real-world datasets and +1.9% on three benchmark datasets.
PaperID: 7727, Poster
Abstract: Large Language Models (LLMs) have made remarkable progress on solving competition-level mathematics, yet their ability to verifying natural-language mathematical proofs remains relatively underexplored. Current proof verification primarily relies on expert inspection that is costly to scale and Olympiad-level problems represent a significant challenge in this area, as they require the meticulous evaluation of every claim in the reasoning chain. In this paper, we present a systematic study of LLM-based verification in this setting. Specifically, we introduce a human-curated dataset of 48 USAMO 2025 candidate solutions with expert grading, conduct a broad study of existing LLM verifiers in Olympiad mathematics across frontier models. We also propose an iterative self-critique pipeline TROJAN to generates high-fidelity adversarial proofs at scale. Our evaluation spans 20 models on 7 metrics across two prompt styles with three inference methods and three prompt templates, illustrating that the state-of-the-art models still have more room for improvement. To address this gap, we propose MABGrader, a Multi-Armed Bandit framework that reframes grading as arm selection over the discrete score space. MABGrader outperforms the SOTA ProofGrader by 26.4% in Quadratic Weighted Kappa and reduces grading error by 14%.
PaperID: 7728, Poster
Abstract: Providing external information beyond the query is a common way to improve large language model (LLM) inference, as seen in in-context learning (ICL), retrieval-augmented generation (RAG), and memory-based methods (Mem). Existing work increasingly suggests that useful external information should compensate for the gap between LLMs and user queries, implicitly assuming that stronger gap-correction can lead to better performance. In this paper, we challenge this assumption and show that more gap-corrective external information can instead harm performance. We analyze this phenomenon in Transformers through the lens of reasoning error, defined as the difference between the predicted answer vector and the ground-truth answer vector. Our analysis reveals that external information can be not only under-corrective but also over-corrective, where excessively strong correction increases reasoning error. We further show that the induced correction vector is jointly determined by attention and the complementarity between the external information and the query, and derive conditions under which information is properly corrective in both direction and magnitude. Experiments on four mainstream LLMs and seven reasoning benchmarks across ICL, RAG, and Mem validate our theory and support a theory-guided information selection method.
Abstract: Scaling large recommendation systems requires advancing three major frontiers: processing longer user histories, expanding candidate sets, and increasing model capacity. While promising, transformers' computational cost scales quadratically with the user sequence length and linearly with the number of candidates. This trade-off makes it prohibitively expensive to expand candidate sets or increase sequence length at inference, despite the significant performance improvements. We introduce LIME, a novel architecture that resolves this trade-off. Through two key innovations, LIME fundamentally reduces computational complexity. First, low-rank ``link embeddings" enable pre-computation of attention weights by decoupling user and candidate interactions, making the inference cost nearly independent of candidate set size. Second, a linear attention mechanism, LIME-XOR, reduces the complexity with respect to user sequence length from quadratic (O(N^2)) to linear (O(N)). Experiments on public and industrial datasets show LIME achieves near-parity with state-of-the-art transformers but with a 10× inference speedup on large candidate sets or long sequence lengths. When tested on a major recommendation platform, LIME improved user engagement while maintaining minimal inference costs with respect to candidate set size and user history length, establishing a new paradigm for efficient and expressive recommendation systems.
PaperID: 7730, Poster
Authors: Aditi Aravind, Konstantinos Ladakis, Mario A Savaglio, Stelios Smirnakis, Maria Papadopouli
Abstract: We investigate how internal representations emerge across hierarchical processing systems by introducing a neuroscience-inspired framework for analyzing deep spiking neural networks (SNN) through the lens of functional connectivity. Drawing on concepts from systems neuroscience and information theory, we form the first-order functionally connected (1FC) group of a neuron based on its statistically significant pairwise correlations with neurons from the previous layer of a trained SNN architecture. We then track its response properties during inference under various conditions. Our analysis shows that several principles of functional connectivity previously observed in biological cortex are preserved in spiking ResNet architectures. These 1FC ensembles display interesting properties: their aggregate cofiring reliably predicts downstream neuronal responses through a robust, ReLU-like input–output relationship, whose gain scales systematically with ensemble size. Reliable encoding of the presented class emerges only during high 1FC cofiring events, which themselves occur infrequently, indicating that informative representations are concentrated in rare but highly coordinated activity patterns. Under uniform random noise or adversarial perturbations, these response profiles are disrupted, particularly in early and intermediate layers. This enables a targeted high-resolution interrogation at specific nodes and pathways when large 1FC cofirings occur. We showed that the functional connectivity structure is shaped by learning and this structure breaks under weight permutation. This establishes 1FC ensembles as a functionally meaningful substrate for input encoding and information transfer, enabling targeted fine-grained diagnostics on the information flow.
PaperID: 7731, Poster
Authors: Farid Talibli, Awais Khan, Christopher Zimmer, Sungyong Park, Jihoon Yang, Youngjae Kim
Abstract: Large-scale AI training runs span days within fixed HPC allocations, making checkpointing essential for bounded recovery after failure. Asynchronous checkpointing hides checkpoint latency from the training loop, but does not guarantee that the most recent state has reached a durable target: data may still be in flight to the shared parallel filesystem (PFS) when a failure occurs, forcing rollback to an older checkpoint. We formalize this durability gap through a distributional model that combines hardware crossover analysis with statistical models of PFS completion time under scale and contention. Rank-local completion times follow a lognormal body distribution, while wall-clock completion is governed by an extreme-value process determined by the slowest writer. The resulting analysis yields a contention-dependent threshold,~C^\star_p(n), beyond which the PFS is no longer the fastest path to durable persistence. Guided by this analysis, we design CacheX, a transparent staging layer that composes with PyTorch Distributed Checkpoint~(DCP). CacheX stages bulk data to local DRAM and replicates it to a neighboring node's NVMe, providing allocation-local durability while preserving eventual persistence to the PFS. On Frontier, across IOR runs up to 256~nodes and end-to-end FSDP+DCP training up to 64~nodes, CacheX reduces write-time variability by up to~43×, achieves a replica-durable effective throughput of~~\,2.5 to 2.8\,GiB/s per node decoupled from PFS contention, and reduces worst-case rollback exposure. Our results show that staging is not a universal accelerator; rather, it is a conditional policy for selecting the first durable checkpoint boundary under measured scale, load, and tail behavior.
Abstract: Cross-view localization classically asks: where does this ground image lie on the satellite tile? Existing methods are typically limited to 3-DoF estimates---an (x,y) position and a yaw angle---because nadir satellite imagery provides no direct cues for roll, pitch, or altitude, forcing a reliance on planar-motion and zero-tilt assumptions. These assumptions break on real terrain with slopes, ramps, and tilted camera mounts. To overcome this, we introduce a single UAV image as an intermediate viewpoint: it reveals the 3D structure invisible from nadir, supplies the cues for roll, pitch, and altitude that the satellite alone cannot provide, and needs only spatial overlap with the ground camera---no known relative pose is required. Building on this insight, we propose Cross3R, a flexible feed-forward model that ingests a satellite tile together with a UAV image, a ground image, or both, and, in a single forward pass, recovers a cross-view 3D point cloud, the 6-DoF poses of every input camera, and the on-tile (x,y) position and yaw of each perspective camera. For training and evaluation, we also construct CrossGeo, a 278K-image tri-view dataset spanning 85 scenes across every continent except Antarctica. On CrossGeo, Cross3R consistently outperforms feed-forward 3D baselines in point-cloud reconstruction, 6-DoF camera-pose estimation, and cross-view localization. On KITTI, it outperforms dedicated cross-view methods trained on KITTI on most metrics, despite having no KITTI training itself.
PaperID: 7733, Poster
Abstract: Generative modeling for science is data-scarce: each sample is expensive but arrives with several coupled measurements of the same system. Standard conditional models often use single-direction generation, discarding supervision from the rest. We argue that this asymmetry should not be a constraint during training: Under random target/conditioning masking, every measurement can act as a training target, hence one expensive sample yields multiple supervised denoising tasks. We present MOSAIC, a multimodal flow matching transformer over molecular 3D structure and UV, IR, and Raman spectra. On QM9S, MOSAIC achieves state-of-the-art performance with 88.8% Acc@1 for spectra-to-structure task, outperforming the strongest baseline by over 20% with 20× fewer integration steps. MOSAIC enables any-to-any generation from any combination of input to any combination of output modalities common in chemical sciences. We provide ablations to demonstrate the effectiveness of MOSAIC and show consistent performance scaling with increasing model and data size.
Abstract: Production deployments often swap between different-sized models in a family for cost-quality cascading, mid-conversation switching, and routing, and each swap forces the receiver to repay the prefill from scratch. We propose cross-model KV cache transfer, where the receiver reuses the source’s KV cache, skipping prefill. We find that cross-model KV has substantial linear structure across matched-KV pairs, where source and target share KV head count and per-head dimension. On Qwen3 14B→32B, one source layer explains 56% of variance in the target’s keys and 32% in values, rising to 79% and 65% with multiple source layers. Building on this, we design a closed-form ridge mapper that operates per head and proceeds in three steps. First, for each target layer we select the top-kmost predictive source layers and concatenate their KV as input. Second, we strip RoPE from the keys before mapping, so the fit is position-free and reusable across context lengths. Third, we fit ridge regression on a small calibration set of 500 FineWeb-Edu sequences of 1,024 tokens each. Surprisingly, across six pairs in three families, this linear mapper retains 71–98% of the receiver’s standalone-prefill accuracy on four pairs, while two degrade sharply. A nonlinear MLP recovers up to +22 pp HellaSwag retention on the failures. The mapper runs 3–31×faster than re-prefill and remains stable across multi-turn handoff, making cross-model KV cache transfer practical.
PaperID: 7735, Poster
Abstract: Semantic ID-based generative recommendation represents each item as a sequence of discrete tokens, enabling structured modeling of item semantics. A critical challenge is constructing semantic IDs that are both semantically expressive and computationally efficient. While recent approaches favor complex learned quantization, simple hashing-based methods such as SimHash are widely regarded as fundamentally inferior. In this work, we challenge this consensus by showing that the apparent performance gap does not stem from inherent limitations of hashing, but rather from a structural mismatch with autoregressive decoding, coupled with the inevitable information loss during rigid discretization. Based on this insight, we propose FLASH, a two-stage framework that revitalizes training-free SimHash tokenization through parallel decoding and explicit semantic alignment. Despite its simplicity, FLASH achieves state-of-the-art performance across multiple datasets without requiring any tokenizer training, while exhibiting stronger generalization in cold-start scenarios. Notably, we demonstrate that semantic alignment acts as a universally effective mechanism across diverse paradigms. Our findings suggest that, with compatible decoding and semantic grounding, simple and efficient tokenizers can outperform complex learned counterparts. Our code and data are available: https://anonymous.4open.science/r/Flash-014A/.
Abstract: Deep Research Agents (DRAs) answer complex questions by searching the web, checking evidence, and synthesizing conclusions across heterogeneous sources. We introduce a category-theoretic framework for evaluating such agents. The framework treats deep research as a structured mapping from user intent to evidence-grounded conclusions, making retrieval traces, cross-source alignment, and final synthesis explicit. Guided by this view, we build a mechanism-aware benchmark of 296 bilingual questions covering four structural skills: following multi-hop evidence chains, verifying claims across sources, re-ordering fragmented information, and rejecting unsupported assumptions. We evaluate 16 systems with human verification and find that these tasks remain difficult: the best system reaches 19.9% average accuracy. The results reveal complementary strengths across systems, but also persistent weaknesses in long-horizon retrieval and intersection-heavy verification. We further instantiate two theory-guided interventions, tracked search and category tools, in API-based agents. These variants improve over their corresponding baselines, suggesting that the framework is useful not only for diagnosis but also for modest, targeted system design.
PaperID: 7737, Poster
Authors: Eric Gilerson
Abstract: Large language models spend enormous parameter capacity approximating arithmetic, a deterministic function with an exact solution, and still fail reliably on multi-digit operations. We propose Arithmetic Residual Blocks (ARBs): frozen, differentiable modules inserted into the transformer forward pass that compute exact integer arithmetic using Residue Number System (RNS) encoding on unit circles. The model learns only the interface, a gated injection pathway and a low-rank adapter on the language model head, while the base model and all arithmetic computation remain frozen. The architecture embodies a design principle we call Don't Learn What You Can Compute: any deterministic function expressible as tensor operations can be embedded as a frozen residual block, and the model will learn to route through it because the gradient reward for exact answers dominates internal approximation. We validate this principle by inserting ARBs into a frozen SmolLM2-360M model, training only 1.7M parameters (0.47% of the base model). On exact-match evaluation across addition, subtraction, multiplication, and division with operands up to 3 digits, the augmented model achieves 99.9%+ accuracy across all operations, compared to 4.6% mean accuracy for the unmodified base model. A logit analysis of the frozen base model confirms that the injection pathway overrides the base model's prior uniformly across operations, and that the residual errors (0.047%) concentrate at later digit positions in autoregressive generation, not at any particular operation. These findings motivate the hypothesis that pre-training with ARBs, where no misaligned frozen representations exist, would yield ceiling performance across all operations.
Authors: Zirui Ren, Ziming Liu
Abstract: Hierarchical reasoning model (HRM) achieves extraordinary performance on various reasoning tasks, significantly outperforming large language model-based reasoners. To understand the strengths and potential failure modes of HRM, we conduct a mechanistic study on its reasoning patterns and find three surprising facts: (a) , e.g., HRM can fail on a puzzle with only one unknown cell. We attribute this failure to the violation of the fixed point property, a fundamental assumption of HRM. (b) , i.e., the answer is not improved uniformly, but instead there is a critical reasoning step that suddenly makes the answer correct; (c) . HRM "guesses" the first fixed point, which could be incorrect, and gets trapped there for a while or forever. All facts imply that . Leveraging this "guessing" picture, we propose three strategies to scale HRM's guesses: data augmentation (scaling the quality of guesses), input perturbation (scaling the number of guesses by leveraging inference randomness), and model bootstrapping (scaling the number of guesses by leveraging training randomness). On the practical side, by combining all methods, we develop
PaperID: 7739, Poster
Authors: Shenglei Li, Tomoji Kishi
Abstract: Millimeter-wave radar is promising for privacy-preserving human sensing, yet current radar-to-language pipelines typically map sparse point clouds into either global clip features or discrete token codes that are only weakly matched to the structure expected by large language models (LLMs). We present RILA, a radar-native structured interface layer for large-language-model adaptation, which converts sparse mmWave radar point clouds into LLM-readable event-aware tokens. RILA combines kinematic phase-space tokenization, flow-aware dual-order serialization, dual-scan selective state-space encoding, and an event-aware token abstraction that exposes both global clip context and temporally localized motion units to the language model. The same interface supports three language tasks through a shared decoder: clip summary generation, ordered event description, and temporal question answering. To stabilize the interface before instruction tuning, we introduce a structured interface-language alignment objective over clip and event tokens. We evaluate RILA using physics-aware synthetic pretraining and real-world mmWave adaptation under controlled radar-native baselines and matched decoder settings. This formulation reframes radar-to-LLM modeling as an interface design problem and studies how radar-native structure, rather than only generic cross-modal projection, determines whether sparse mmWave point clouds can be effectively understood by LLMs.
Abstract: Setting the learning rate for a deep learning model is a critical part of successful training, yet choosing this hyperparameter is often done empirically with trial and error. In this work, we explore a solvable model of optimal learning rate schedules for a powerlaw random feature model trained with stochastic gradient descent (SGD). We consider the optimal schedule \eta_T^\star(t) where t is the current iterate and T is the total training horizon. This schedule is computed both as a numerical optimization problem and also analytically (when possible) using optimal control theory. Our analysis reveals two regimes which we term the easy phase and hard phase. In the easy phase the optimal schedule is a polynomial decay \eta_T^\star(t) ~eq T^-\xi (1-t/T)^\delta where \xi and \delta depend on the properties of the features and task. In the hard phase, the optimal schedule resembles warmup-stable-decay with constant (in T) initial learning rate and annealing performed over a vanishing (in T) fraction of training steps. We investigate joint optimization of learning rate and batch size, identifying a degenerate optimality condition and showing that batch ramps can improve the scaling in wall-clock time in the easy phase. Beyond SGD, we derive optimal schedules for the momentum parameter \beta(t) and show that momentum achieves a better loss-scaling exponent in the hard phase. We compare our optimal schedule to various benchmarks in our task including (1) optimal constant learning rates \eta_T(t) ~ T^-\xi (2) optimal power laws \eta_T(t) ~ T^-\xi t^-\chi, finding that our schedule achieves better rates than either of these. Our theory suggests that learning rate transfer across training horizon depends on the structure of the model and task. For ResNet image classification on CIFAR-5M, the learning curves exhibit hard-phase behavior where optimal base learning rates are constant under sufficient annealing. GPT-2 style transformers trained in language modeling exhibit easy-phase behavior where optimal learning rates shift even under annealing.
Abstract: Recent work has shown that flow matching models can be trained without explicit time conditioning, challenging the standard view that the interpolation time is needed to disambiguate velocity targets. But why should a time-blind model work at all? Decomposing the time-blind flow matching loss, we identify two sources of irreducible error: a coupling variance, which arises from ambiguous velocity targets induced by how noise and data points are paired, and the time-blindness gap, which is the additional error caused by ignoring time. This gap shows that time-blind training is strictly harder than conventional training, reinforcing the puzzle that time-blind models work so well in practice. We resolve this tension by showing that the geometry of high-dimensional data makes time identifiable directly from noisy observations. When data concentrates near a k-dimensional subspace, time can be recovered from the statistical structure of noisy interpolants in directions orthogonal to the data; under a spiked-covariance model, this yields a closed-form estimator that recovers t from a single observation z at rate O(1/\sqrtd-k) for ambient dimension d. As a consequence, we prove that the time-blindness gap is asymptotically negligible relative to the coupling variance. We empirically demonstrate our identifiability result on real-world data and show that changing the coupling has a much larger effect on loss and sample quality than removing time conditioning across CIFAR-10, CelebA-HQ, and FFHQ. These results explain why time-blind flow matching works and show that the main practical lever is the choice of coupling, not explicit time conditioning.
PaperID: 7742, Poster
Abstract: Learning Gaussian mixture models (GMMs) using the Expectation-Maximization (EM) algorithm and its gradient-based variants is a fundamental problem in machine learning. It is known that random initialized (gradient) EM fails to learn multi-component GMMs in the exact-parameterized setting, where the number of components matches that of the ground-truth GMM. Recently, global convergence of gradient EM has been established in the over-parameterized setting, where more components are used, provided that the ground-truth components are well separated. In particular, the minimum separation between ground-truth components is required to scale as \Omega(\sqrtd), where d is the dimension. In this paper, we show that this dimensional dependence is unavoidable in high-dimensional settings. Specifically, for any \epsilon > 0, we prove that when the dimension is sufficiently large, a separation of order \Omega(d^0.5-\epsilon) is insufficient to guarantee global convergence of population gradient EM in sub-exponential time under random initialization, even in the over-parameterized regime. Our result establishes an almost optimal lower bound on the ground-truth separation required for learning Gaussian mixtures via gradient EM in high dimensions.
Abstract: We introduce a meshfree exterior calculus (MEEC) for learning structure-preserving descriptions of physics on point clouds, and use it to build MEEC-Net, a data-efficient surrogate that transfers across resolutions, geometries, and physical parameters. MEEC equips an \epsilon-ball graph with virtual node and edge measures via a single sparse Schur complement solve; the resulting complex satisfies discrete conservation exactly, is end-to-end differentiable in the point positions, and exposes a direct geometry-to-physics link without the mesh-generation step required by conventional structure-preserving discretizations. MEEC-Net learns unknown physics as a shared edge-wise flux law in an SO(d)-invariant local frame, so the same kernel produces compatible fluxes on any point cloud whose features lie in the training range. We prove a solution-error bound that splits into discretization and kernel-approximation terms which is independent of problem geometry, explaining the observed transfer from very few examples. We show that single-solution training transfers to unseen geometries, boundary conditions, and physical parameters. On five canonical PDE benchmarks MEEC-Net achieves 1–2 orders of magnitude lower out-of-distribution error than baseline neural-operator approaches. On the SimJEB structural-bracket benchmark it achieves competitive error while using substantially fewer training geometries.
PaperID: 7744, Poster
Abstract: Generating 3D indoor objects under fine-grained structural control is crucial and has broad applications in robotics, gaming, and simulation. Such fine-grained structural control encompasses an object’s spatial layout, geometric icons, and relative size, which together constitute the object’s core structural properties. However, existing controllable 3D generation methods do not take structural control as a primary objective, and they commonly rely on suboptimal conditions in representing structure or inadequate algorithms for maintaining structural consistency. Therefore, they fail to achieve fine-grained structural control during generation. To address this, we present structure-grounded 3D object generation, a novel perspective on 3D object generation, which decomposes 3D shapes into structural configurations and geometric details, directly conditions on the object’s structural configurations and synthesizes realistic geometric details to produce high-quality 3D objects. Building on this perspective, we introduce StructBridge, which first employs a diffusion bridge in the 3D latent space to transform comprehensive structural conditions into structurally accurate coarse meshes, followed by a structure-preserving normal refiner that enriches geometric details while preserving structure. We conduct extensive experiments and comprehensive comparisons with various conditional generation methods, demonstrating that StructBridge achieves state-of-the-art performance in both generation quality and structural control capabilities.
PaperID: 7745, Poster
Authors: Yuan Sheng Fang, Dmitri E Nikonov, Ikenna Odinaka, Alan Kalitsov, Roza Kotlyar, Arnab Kabiraj, Benjamin W Chen, Shundan Xiao, Brian Demsky, Chryston B Meng, Chang W Kang, Teck L Tan, Jin Hongmei, Debo Olaosebikan, Sasikanth Manipatruni, Amrita Mathuriya
Abstract: Discovery of novel materials for physical memory is essential for superior memory systems enabling AI hardware scaling. Among them, ferroelectrics hold promise for highly scalable DRAM. Yet their discovery remains driven largely by human intuition and exhaustive database screening. We present Ferrogen, a generative pipeline for the targeted discovery of novel ferroelectric materials for memory by fine-tuning the diffusion-based crystal structure generator Mattergen on ferroelectric-relevant properties. To enable property-conditioned generation, we construct a ML-labeled dataset using a suite of fast machine learning estimators, including a novel two-stage ensemble polarization predictor, dramatically reducing reliance on high-throughput Density Field Theory (DFT) during training and screening. The base Mattergen model is fine-tuned via lightweight adapter layers conditioned on polarization, switching energy, and metal-probability embeddings, with classifier-free guidance enabling targeted generation of candidates with high polarization, low but finite switching barriers, and insulating character. Generated candidates are screened by the same ML estimators and validated through rigorous DFT calculations, including structural relaxation, band gap verification, and Berry phase polarization computation. Against the ML screens, Ferrogen achieves a roughly 80 × improvement in search efficiency over exhaustive database search. The screened candidates are then finally validated using DFT and made publicly available as a database of theoretically confirmed novel ferroelectrics. The pipeline is readily extensible to other functional material classes, establishing generative models as a powerful paradigm for accelerating electronic materials discovery for high performance memory.
PaperID: 7746, Poster
Authors: Peng Jiang, Xiaoxuan Jia
Abstract: Forward prediction of neural population activity is a prerequisite for closed-loop brain-computer interfaces and a stringent test of learned neural dynamics. Unlike neural decoding or masked reconstruction, this task requires a model to forecast future spiking activity from neural history alone and to remain stable when its own predictions become part of the conditioning context. We introduce NeuroHorizon, an encoder-decoder architecture that combines event-level tokenization of spike trains with an autoregressive causal decoder for population-level firing-rate prediction. The decoder uses a hierarchical tail+segment memory to retain recent predictions at high resolution while compressing longer prediction history, and scheduled sampling reduces the mismatch between teacher-forced training and autoregressive rollout. Evaluated on a multi-horizon motor-cortex benchmark with additional motor- and visual-cortex datasets for scaling and cross-population transfer, NeuroHorizon consistently outperforms strong baselines and remains stable under multi-step rollout where prior approaches collapse. These results show that long-horizon neural forecasting becomes feasible when models are explicitly trained and evaluated under autoregressive rollout, providing a foundation for closed-loop BCI and multi-step neural-state prediction.
PaperID: 7747, Poster
Authors: Hiroyuki Kasai
Abstract: Most theories of in-context learning (ICL) treat the prompt as fixed. This fixed-prompt view misses the key difficulty in memory-updated retrieval systems such as Memory Mosaics. An inference-time write can help later queries, but it can also shift softmax weights and create interference. The question is when the benefit exceeds the later cost. We introduce TRIM (Theory of Retrieval with Incremental Memory), a theorem-level theory of append-only softmax retrieval under a write-independence abstraction. The five-theorem spine links retrieval-weight concentration, local write-induced defect, recurrence-level accumulation, ideal squared-loss benefit, and a benefit-cost comparison with critical-horizon regimes. TRIM places helpful memory and harmful interference within a single frame, so the harmful length scale is derived rather than tuned. We record a first-order dynamic correction for state-dependent writes separately, and it is not part of this spine. We give a theorem-facing empirical protocol. Exact synthetic checks test the spine identities and inequalities. Learned base-MM analyses measure calibrated analogues of concentration, local defect, horizon worsening, and realized benefit. These analyses test whether TRIM observables organize learned systems along the same routes.
PaperID: 7748, Poster
Abstract: Medical image segmentation remains fragmented along two axes: segmentation paradigms and data dimensionality. Existing methods are typically developed separately for semantic, in-context, and interactive segmentation, and are further specialized to either native 2D images or 3D volumetric data. In clinical practice, however, segmentation workflows take many forms: a case may be initialized by semantic prediction, reference-guided segmentation, or user interaction. Regardless of how it begins, fine-grained refinement is naturally performed on 2D views; for volumetric scans, such 2D edits must propagate coherently to the rest of the volume. We present pagation to extend 2D segmentation to 3D volumes. Our key insight is that volumetric propagation and in-context segmentation share the same reference-conditioned prediction mechanism, differing only in whether the reference image–mask pairs come from other cases or from previously segmented neighboring slices. Building on this view, UniPro supports semantic, in-context, interactive, and propagation-based segmentation within a single slice-based framework, using class priors, reference exemplars, user clicks, and neighboring-slice predictions as mode-specific conditioning inputs. To improve propagation reliability, UniPro further incorporates bidirectional and 3D supervision to regularize slice-wise propagation beyond per-slice losses. Extensive experiments across diverse modalities and anatomies show that UniPro achieves strong performance across all segmentation settings, enabling annotation-efficient 3D segmentation from sparse 2D initialization and reducing slice-by-slice correction effort.
PaperID: 7749, Poster
Authors: Arvid Ernst Gollwitzer, David de Gruijl
Abstract: Microbial ecosystems govern human health, agricultural productivity, and biogeochemical cycling. Two coupled modalities determine their function: (metabolomic content). Existing biological foundation models cannot represent this joint state: they operate on a single modality, on assembled single-organism genomes, and on the within-DNA axis of generalization, leaving the multi-omic, mixed-organism, mixed-quality regime that governs clinical and surveillance use unaddressed. We present , a 7-billion-parameter multi-omic foundation model whose primary axis of generalization spans both modalities (metagenome + metabolome) and organizational scales (read → community → host phenotype). We combine a Mamba–Transformer hybrid backbone (24 of 32 layers Mamba) with a hypergraph neural network over KEGG metabolic-reaction hyperedges and bidirectional cross-modal attention modules that align genomic and metabolomic representations in a shared 40,192-token embedding; an Aitchison-space compositional consistency loss respects the simplex geometry of microbial-abundance data at the output. We pretrain on 8 T base pairs of metagenomics, 250K LC-MS/MS metabolite profiles, and 2M KEGG/GO functional annotations, preparing the metagenomic corpus with a published quality-aware tokenization framework. Against the matched 7B genomic baseline METAGENE-1, we report 94.5 MCC on Pathogen Detection (vs. 93.0) and 0.98 F1 on CAMI species-level metagenomic profiling; on a December 2025 SRA outbreak benchmark post-dating the pretraining cutoff, we report 91.0 MCC vs. 81.2. We unlock four clinical tasks single-modality genomic models cannot reach: IBD AUC 0.947, T2D AUC 0.883, antibiotic-resistance AUC 0.910, and metabolic-pathway prediction wF1 0.91, beating the strongest multi-omic baseline (MOGONet) by +3.0–+3.7 AUC. Token-matched scaling from 100M to 70B yields monotone gains with no saturation on the multi-omic clinical AUCs. We further validate predictions in a prospective wet-lab pilot (n = 87) at predicted-vs-observed Pearson r = 0.72. Sparse-autoencoder probing recovers 14 high-coherence biological features tracking KEGG pathways, NCBI taxa, and CARD AMR genes (r = 0.68–0.74); zeroing the butyrate-producer feature drops IBD AUC from 0.947 to 0.928. We conclude that multi-omic, multi-scale modeling is tractable at 7B parameters, and we release the model card, training-data manifest, evaluation pipelines, and a 100-trajectory MetaOmics-10T causal-pilot dataset.
Authors: Anmol Gulati, Hariom Gupta, Elias Lumer, Sahil Sen, Vamse K Subbiah
Abstract: Long-horizon AI agents execute complex workflows spanning hundreds of sequential actions, yet a single wrong assumption early on can cascade into irreversible errors. When instructions are incomplete, the agent must decide not only whether to ask for clarification but when, and no prior work measures how clarification value changes over the course of execution. We introduce a forced-injection framework that provides ground-truth clarifications at controlled points in the agent's trajectory across four information dimensions (goal, input, constraint, context), three agent benchmarks, and four frontier models (three per benchmark; one on a single benchmark only; 84 task variants; 6,000+ runs). Counter to the common intuition that ``earlier is always better,'' we find that the value of clarification depends sharply on what information is missing: goal clarification loses nearly all value after 10% of execution (pass@3 drops from 0.78 to baseline), while input clarification retains value through roughly 50%. Deferring any clarification type past mid-trajectory degrades performance below never asking at all. Cross-model Kendall τ correlations (0.78-0.87 among models sharing identical task coverage; 0.34-0.67 across the full 4-model panel) confirm these timing profiles are substantially task-intrinsic. A complementary study of 300 unscripted sessions reveals that no current frontier model asks within the empirically optimal window, with strategies ranging from over-asking (52% of sessions) to never asking at all. These empirical demand curves provide the quantitative foundation that existing theoretical frameworks require but have lacked, and establish concrete design targets for timing-aware clarification policies. Code and data will be publicly released.
Authors:
Suyoung Bae, Jaehoon Lee, Changkyu Choi, YunSeok Choi, Jee-Hyong LeeAbstract: Automated code documentation is essential for modern software development, providing the contextual grounding that both human developers and coding agents rely on to navigate large codebases. Existing repository-level approaches process components independently, causing redundant retrieval and conflicting descriptions across documents while producing outputs that lack hierarchical structure. Therefore, we propose MemDocAgent, a long-horizon agentic framework that generates documentation within a single, integrated context spanning the entire repository. It combines two components: (i) Dependency-Aware Traversal Guiding that predetermines a traversal order respecting dependency and granularity hierarchies; (ii) Memory-Guided Agentic Interaction, in which the agent interacts with RepoMemory, a shared memory accumulating prior work traces through read, write, and verify operations. Through an in-depth multi-criteria evaluation, MemDocAgent achieves the best performance over both open and closed-source baselines and demonstrates practical applicability in real software development workflows.
PaperID: 7752, Poster
Authors: Blaz Bertalanic, Carolina Fortuna
Abstract: Inference-time multi-agent LLM scaling lacks a shared unit: counting nominal agents conflates cost with independent evidence. We derive a two-parameter scaling law R(N) = N_\texteff/N = 1/(1+c(N-1)N^-\beta) where the regime exponent \beta classifies any configuration into one of three asymptotic regimes --- hard-ceiling at 1/c (\beta = 0), sublinear at N^\beta/c (0 < \beta < 1), or linear (\beta \ge 1), and a mean-field theorem predicts that peer count k and rounds\tau during agent debate enter the dynamics only through their product k\tau. The law applies at two levels: answer diversity and correctness redundancy. Across 44 (model × task × condition) cells spanning peer debate, self-correction, random-noise placebo, self-consistency, three open-weight families (Qwen, Llama, Ministral) at scales from 7B to 32B with a frontier API check (Gemini), thinking models, heterogeneous teams, and sparse communication, the functional form fits every condition at R^2 > 0.99; only (c, \beta) shifts. On free-form math, dense peer influence collapses the answer-level regime from sublinear into hard-ceiling; correctness-level fits remain hard-ceiling throughout. Three findings have practical implications. \emph(i)~Thirty dense debating agents produce no more answer diversity than one on MMLU-Hard. \emph(ii)~A noise placebo tracks self-correction on free-form math and at 4× scale, so within homogeneous teams the gain commonly attributed to ``debate'' comes from re-evaluation, not peer content. \emph(iii)~A single N \le 5 pilot predicts the N=30 structural ceiling, and within the configurations tested only architectural diversity (heterogeneous teams) lowers c and escapes the hard-ceiling regime, communication-mode interventions do not.
Abstract: Generative models trained on finite data face a fundamental tension: their score-matching or next-token objective converges to the empirical training distribution rather than the population distribution we seek to learn. Using rule-valid synthetic tasks, we trace this tension across two training timescales: \tau_rule, the step at which generations first become rule-valid, and \tau_mem, the step at which models begin reproducing training samples. Focusing on parity and extending to other binary rules and combinatorial puzzles, we characterize how these two clocks \tau_rule, \tau_mem depend on key aspects of the learning setup. Specifically, we show that \tau_rule increases with rule complexity and decreases with model capacity, while \tau_mem is approximately invariant to the rule and scales nearly linearly with dataset size N. We define the \emphinnovation window as the interval [\tau_rule, \tau_mem]. This window widens with increasing N and narrows with rule complexity, and may vanish entirely when \tau_rule \geq \tau_mem. The same two-clock structure arises in both diffusion (DiT) and autoregressive (GPT) models, with architecture-dependent offsets. Dissecting the learned score of DiT models reveals a corresponding evolution of the optimization landscapes, where attractors emerge at both timescales: rule-valid samples' basins expand substantially around \tau_rule, while training samples' basins begin to dominate around \tau_mem. Together, these results yield a unified and predictive account of when and how generative models exhibit genuine innovation.
PaperID: 7754, Poster
Abstract: We study the problem of efficient online reinforcement learning in the infinite horizon setting when there is an offline dataset to start with. We assume that the offline dataset is generated by an expert but with unknown level of competence, i.e., it is not perfect and not necessarily using the optimal policy. We show that if the learning agent models the behavioral policy (parameterized by a competence parameter) used by the expert, it can do substantially better in terms of minimizing cumulative regret, than if it doesn't do that. We establish an upper bound on regret of the exact informed PSRL (iPSRL) algorithm that scales as \tilde\mathcalO(\sqrtT). This requires a novel prior-dependent regret analysis of Bayesian online learning algorithms for the infinite horizon setting. We then propose the informed RLSVI (iRLSVI) algorithm to efficiently approximate the iPSRL algorithm. Empirical results show that small offline expert datasets are sufficient for iRLSVI to outperform online-only algorithms on hard exploration domains.
Abstract: Inference-time reward guidance is a simple way to improve generative models when completed outputs can be checked or scored by an external verifier. Best-of-N sampling is a particularly strong baseline: it is parallelizable and often highly improves with more samples. However, when the reference model assigns low probability to reward-satisfying samples, selection among independent full rollouts becomes sample-inefficient. We introduce MDM-VGB, a reward-guided discrete diffusion sampler that augments masked generation with value-guided re-masking. MDM-VGB extends the autoregressive VGB backtracking chain from fixed left-to-right prefix trees to any-order masked-state graphs, allowing the sampler to reveal and revise tokens at any position. The resulting Markov chain favors local reveal and re-mask moves that lead to higher-value partial masked states, enabling both root-start reward tilting and leaf-start editing of low-reward samples. We further introduce a flow-cancelled momentum lift of MDM-VGB that preserves the projected target law while reducing oscillatory reveal/re-mask dynamics under finite budgets. We prove that the sampler has the correct reward-tilted leaf law, places non-negligible stationary mass on complete outputs, and remains robust to partial-state verifier error. Across structured reward satisfaction and repair tasks, MDM-VGB improves the quality--cost frontier over reference sampling, best-of-N, and forward-only value-guided rollout.
PaperID: 7756, Poster
Authors:
Lei Liu, Yunji Liang, Xiaowen Zhang, Jingqi Liu, Qi Li, Bin Guo, Zhiwen YuAbstract: Large language model (LLM) agents are highly vulnerable to memory-poisoning attacks. Existing defenses primarily rely on external modules for either static memory isolation or continuous online auditing, resulting in low memory efficiency and high computational overhead. To address these limitations, inspired by neuroscientific mechanisms that weak reactivation can destabilize memories and induce decay in the absence of reinforcement, we propose an agent-intrinsic training-free memory repair framework, DREAM (Dynamic Reactivation for Engram Attenuation in Memory), that enables selective functional forgetting of poisoned memories. Specifically, DREAM implements a three-stage pipeline: perturbation-induced implicit reactivation to reveal the structural fragility of malicious memories, adaptive anomaly diagnosis based on multi-dimensional activation patterns to detect topological anomalies, and a dynamic memory repair module that selectively suppresses harmful memories while reinforcing benign ones. Extensive experiments across diverse attackers and real-world tasks demonstrate that DREAM reduces attack success rates by over 95% against backdoor poisoning while preserving strong benign utility. DREAM also achieves competitive robustness against injection attacks. In terms of runtime and token consumption, DREAM achieves up to a 2.13× speedup and a 37.7% reduction in token consumption compared with A-MemGuard. Furthermore, DREAM also achieves a task success rate of 92.96% on poisoned multi-agent systems.
PaperID: 7757, Poster
Abstract: Multimodal Large Language Models (MLLMs) require continuous factual updates to stay current, yet lifelong knowledge editing remains a daunting challenge. The primary obstacle is the severe Locality Erosion and Catastrophic Forgetting, where sequential edits inevitably interfere with pre-trained capabilities and prior updates. In this paper, we propose ROSE, a Risk-aware Orthogonal Subspace navigation framework for lifelong knowledge Editing. ROSE addresses forgetting by restricting parameter updates to the orthogonal complement of previously edited subspaces. Crucially, it introduces a risk-guided mechanism that utilizes a dynamically computed risk mask to safeguard specific factual parameter subspaces essential for model stability. During inference, ROSE employs a multi-granularity knowledge integration strategy, featuring prototype-anchored semantic gating and perturbation-driven subspace fusion to precisely activate edited knowledge for relevant queries while strictly maintaining the integrity of the frozen backbone for unrelated inputs. Extensive evaluations on two major multimodal benchmarks demonstrate that ROSE consistently outperforms state-of-the-art methods across five key metrics, exhibiting exceptional robustness over 1,000 sequential updates. Our code is available in the supplementary material.
Abstract: Data-parallel (DP) load balancing has emerged as a first-order bottleneck in large-scale LLM serving. When a model is sharded across devices via tensor parallelism (TP) or expert parallelism (EP) and replicated across many DP workers, every decode step ends in a synchronization barrier whose latency is set by the most heavily loaded worker; even modest persistent imbalance across DP workers compounds, step after step, into a substantial fraction of wasted compute. The problem is hard for reasons specific to LLM decoding: assignments are sticky (KV caches cannot be migrated), per-request loads grow over time, arrivals are non-stationary, and the router must decide within a sub-100ms decode budget over hundreds of waiting requests and tens of workers. We present BalanceRoute, a family of practical online routing algorithms that target this bottleneck. The first, BR-0, requires no prediction infrastructure and uses a piecewise-linear F-score that captures the sharp asymmetry between admissions that fill safe margin and those that overflow into the envelope; a two-stage decomposition keeps per-step cost compatible with millisecond-scale scheduling. The second, BR-H, generalizes BR-0 with a short, constant lookahead H and a lightweight termination-classifier interface, extending the F-score to a horizon-discounted form. We deploy BalanceRoute on a 144-NPU cluster and evaluate against vLLM baselines on both a proprietary production trace and the public Azure-2024 trace. Across both workloads, BalanceRoute substantially reduces average DP imbalance and improves end-to-end serving throughput. An anonymized open-source release is available at https://anonymous.4open.science/r/BR-Family-on-vllm-acend-0F95/.
PaperID: 7759, Poster
Authors: Chanyoung Chung, Jeemin Kim, Joyce Whang
Abstract: Mapping software issue descriptions to relevant code segments, a task known as code localization, is a primary bottleneck in automated software engineering. While recent attempts have automated it using code graphs that capture syntax-level dependencies among code elements, including functions and modules, their code-centric approaches often overlook the design rationales and prior fix patterns present in a codebase's development history, such as pull requests and commits. We propose Co-Locator, a collaborative multi-agent code localization framework that integrates code graphs reflecting functional dependencies and knowledge graphs constructed from triplets capturing design rationales and past code modifications scattered across the development history. During localization, a Code Graph (CG) agent explores functional hierarchies and consults a Knowledge Graph (KG) agent to gather insights into developer intents and prior bug fixes. Combined with a Retriever agent that assists CG and KG agents, Co-Locator leverages a holistic view of the code's functional and historical context, considering what the code implements and how it was changed, to identify locations to fix. Experiments show that Co-Locator significantly outperforms 12 state-of-the-art baseline methods.
Abstract: Understanding the curvature evolution of the loss landscape is fundamental to analyzing the training dynamics of neural networks. The most commonly studied measure, Hessian sharpness (\lambda_\max^H) \textemdash the largest eigenvalue of the loss Hessian \textemdash determines local training stability and interacts with the learning rate throughout training. Despite its significance in analyzing training dynamics, direct measurement of Hessian sharpness remains prohibitive for Large Language Models (LLMs) due to high computational cost. We analyze \emphcritical sharpness (\lambda_c), a computationally efficient measure requiring fewer than 10 forward passes given the update direction \Delta \bm\theta. Critically, this measure captures well-documented Hessian sharpness phenomena, including progressive sharpening and Edge of Stability. Using this measure, we provide the first demonstration of these sharpness phenomena at scale, up to 7B parameters, spanning both pre-training and mid-training of OLMo-2 models. We further introduce \emphrelative critical sharpness (\lambda_c^1\to 2), which quantifies the curvature of one loss landscape while optimizing another, to analyze the transition from pre-training to fine-tuning and guide data mixing strategies. Critical sharpness provides practitioners with a practical tool for diagnosing curvature dynamics and informing data composition choices at scale. More broadly, our work shows that scalable curvature measures can provide actionable insights for large-scale training.
Abstract: Vision-Language Model (VLM) driving agents promise explainable end-to-end autonomy by first producing natural-language reasoning and then predicting trajectory planning. However, whether planning is driven by this reasoning remains a critical but unverified assumption. To investigate this, we build DriveMind, a large-scale driving Visual Question Answering corpus with plan-aligned Chain-of-Thought (CoT), automatically generated from nuPlan. Our data generation process converts sensors and annotations into structured inputs and, crucially, separates priors from to-be-reasoned signals, enabling clean information ablations. Using DriveMind, we train representative VLM agents with Supervised Fine-Tuning and Group Relative Policy Optimization and evaluate them with nuPlan’s metrics. Our results, unfortunately, indicate a consistent in reasoning-planning: removing ego/navigation priors causes large drops in planning scores, whereas removing CoT produces only minor changes. Attention analysis further shows that planning primarily focuses on priors rather than the CoT. Based on this evidence, we propose the Reasoning-Planning Decoupling Hypothesis, positing that the training-yielded reasoning is an ancillary byproduct rather than a causal mediator. To enable efficient diagnosis, we introduce a novel, training-free probe that measures an agent's reliance on priors by evaluating its planning robustness against minor input perturbations. In summary, we provide the community with a new dataset and a diagnostic tool to evaluate the
PaperID: 7762, Poster
Authors:
JiahaoXu, Peiyuan Wang, Hanzhuo Zhang, Zihao Yu, Tianyu Fu, Hao Chen, Xuanhao Xiang, Zixuan Li, Jianbo Yu, Chenchen Fu, Wanyuan WangAbstract: In robotic manipulation, the tight coupling between grasping and motion planning often obscures the true source of failure, leading to inefficient trial-and-error. To enable efficient long-horizon manipulation, we propose GTP-FA (Grasp-Then-Plan with Failure Attribution), a task-oriented two-stage ‘grasp-then-plan’ framework that generates grasp candidates and performs downstream motion planning conditioned on the selected grasp. Given a failed manipulation trajectory, we learn a failure attribution model that generalizes to unseen grasps and produces a stable distribution over failure modes for diagnosis-guided optimization. Based on these attribution results, we then optimize both modules in a diagnosis-driven manner: on the grasping side, we inject task-level priors and risk penalties into grasp candidate scoring and optimization to suppress unstable or task-incompatible grasps; on the planning side, we target high-risk initial states through data collection and fine-tuning to address genuine planning bottlenecks. We evaluate the proposed framework in both simulation and real-robot experiments, and show that GTP-FA improves the corresponding base learners across RL, IL, diffusion-policy, and VLA-based settings, achieving substantially higher overall task success rates. Project page: here.
PaperID: 7763, Poster
Abstract: Associative Recall (AR) is the cognitive ability to learn and retrieve links between items in memory. In NLP, AR is used as a benchmark for evaluating the in-context memory capacity of architectures such as Mamba, and has been found to strongly correlate with language modeling performance. This paper explores AR from the perspective of mechanistic interpretability, aiming to reverse-engineer the exact internal algorithm used by Mamba to perform recall. Our key insight is that Mamba performs recall by implicitly learning linear hash functions, and we identify the low-level circuit that enables this behavior. Building on these findings and inspired by theoretical tools in similarity-preserving hashing, such as the Johnson–Lindenstrauss (JL) lemma, we develop a theoretical framework for analyzing AR, which we term recall scaling laws. For example, given a model’s state, context length, and embedding dimensions, our theory predicts a lower bound on the largest vocabulary size that can be perfectly recalled in the AR task. Empirical results show that this bound is tight and predictive, offering insights into how AR capacity scales with vocabulary size, state size, embedding size, and model architecture.
PaperID: 7764, Poster
Abstract: Large language model (LLM) unlearning has emerged as an essential post-training mechanism for erasing specific knowledge or undesirable behaviors. However, forgetting target data often causes an unintended degradation in overall model utility. Although various advanced methods have explored different learning objectives to mitigate the trade-off, it remains unclear how the highly entangled internal representations in LLMs contribute to unlearning. In this work, we introduce the notion of latent knowledge fragility to explore the vulnerability of retained knowledge to unlearning. We develop a unified analytical approach via component-wise parameter patching that isolates and quantifies fragility in terms of different transformer blocks. We observe that the LLM encodes different levels of abstraction, from surface syntax in shallow layers to complex semantics in deeper layers, which align with different degrees of representation disruption and utility degradation. Based on the insights, we propose a lightweight framework called Component-wise Replacement Unlearning (CRU) that restores fragile layers (also extendable to other components) from the original model based on post-hoc validation, which allows us to obtain a hybrid model without additional training. Extensive experiments on various aspects verify that CRU generally improves the trade-off between removal and retention with non-uniform unlearning influence.
PaperID: 7765, Poster
Abstract: As Large Language Models (LLMs) deploy into mission-critical domains (e.g., finance, medicine, and law), output reproducibility has become a strict system requirement. While practitioners use greedy decoding to eliminate algorithmic stochasticity, empirical deployments with 16-bit precisions still exhibit catastrophic output divergence across heterogeneous GPUs. Through SASS-level profiling, we reveal that this inconsistency is fundamentally driven by truncation errors introduced during downcasting at kernel boundaries. However, achieving reproducibility via a global FP32 pipeline incurs prohibitive system penalties: bypassing 16-bit hardware accelerators hurts compute efficiency, while upcasting the KV cache doubles memory overhead. To bridge this gap, we propose Hybrid Error ALleviation (HEAL), a targeted intervention that approximates FP32 precision while resolving hardware constraints through two targeted mechanisms. First, recognizing that floating-point formats underutilize their bit-width for Q, K, V tensors, HEAL applies INT16 quantization that preserves numerical stability without expanding the KV cache footprint. Second, HEAL synthesizes high-precision matrix multiplications via an algebraic error compensation strategy, executing entirely on high-throughput 16-bit Tensor Cores. To evaluate our approach practically, we introduce MCR-Bench, a benchmark targeting reproducibility in mission-critical tasks. HEAL achieves the same level of reproducibility on downstream tasks as the FP32 baseline while reducing the performance overhead by up to 6.7×.
Abstract: Real-world datasets are inherently heterogeneous, yet how per-class structural differences and sampling imbalance shape the training dynamics of diffusion models—and potentially exacerbate disparities—remains poorly understood. While models typically transition from an initial phase of generalization to memorizing the training set, existing theory assumes homogeneous data, leaving open how class imbalance and heterogeneity reshape these dynamics. In this work, we develop a high-dimensional analytical framework to study class-dependent learning in score-based diffusion models. Analyzing a random-features model trained on Gaussian mixtures, we derive the feature-covariance spectrum to characterize per-class generalization and memorization times. We reveal the explicit hierarchy governing these dynamics: class variance is the primary determinant of learning order—consistently favoring higher-variance classes—while centroid geometry plays a secondary role. Sampling imbalance acts as a modulator that can reverse this ordering and, under strong imbalance, forces minority classes to acquire distinct, delayed speciation times during backward diffusion. Together, these results suggest that diffusion models can memorize some classes while others remain insufficiently learned. We validate our theoretical predictions empirically using U-Net models trained on Fashion MNIST.
Abstract: Missing data imputation, where a model is trained on observed data to estimate unobserved values, is a fundamental problem in machine learning. In this paper, we rigorously formulate imputation model learning as a mean-squared error risk minimisation problem. We show that when the probability of missingness depends on the data, many state-of-the-art methods fail to account for the resulting distribution shift between the observed data used for training and the full data distribution used for evaluation. Consequently, these approaches do not minimise mean-squared error on the full data distribution. Instead, we propose a novel imputation algorithm designed to learn an imputation model from the observed data while explicitly accounting for this distribution shift. Simulation studies show consistent improvements over otherwise identical uncorrected baselines, with average reductions of 3% in RMSE and 7% in Wasserstein distance.
PaperID: 7768, Poster
Abstract: Detecting training data memorization in diffusion models is important for copyright protection and privacy auditing. We view memorization as an abnormally local concentration process: during reverse-time generation, a memorized instance acts as a point-like attractor whose probability basin is sharper and more position-sensitive than that of a generalized concept. Starting from the Fokker–Planck equation, we derive an exact evolution law for the score field \partial_t \mathbfs_t and show that its leading geometric contribution in score-dominated local regimes is \mathbfH_t \mathbfs_t, linking temporal concentration dynamics to spatial curvature. We then show that standard sampling along a shared guided trajectory already exposes this curvature information through a complementary transport-curvature response, without requiring explicit Hessian computation. Motivated by this analysis, we propose the Dynamical Singularity Metric (DSM), an on-trajectory detector that measures the pathwise score-evolution discrepancy between conditional and unconditional branches. DSM requires no additional backpropagation or network evaluations beyond sampling itself. Experiments on Stable Diffusion show that DSM matches or exceeds curvature-based baselines while being substantially cheaper and effective at the earliest reverse step.
PaperID: 7769, Poster
Authors: Stephane Zsoldos, Therice Morris, Varundev Sukhil, Benjamin Wetherfield
Abstract: The analysis of electrostatic induction effects on transmission lines in real-world power distribution systems requires inferring induced surface charge density \sigma on overhead-line corridors at deployment scale, where dense \mathcalO(N^3) boundary-element solvers for the underlying Poisson boundary integral equation (BIE) are infeasible. We present a recipe for training cheap PDE surrogates from canonical-case closed forms (flat ground, grounded sphere) plus sparse dense-solver fine-tuning where the canonical case breaks like in our case study. The recipe also avoids a structural degeneracy between training operator A_\theta and solution \sigma in self-supervised operator learning that tends to lead to internally consistent but non-physical solutions. The recipe combines a closed-form local-tangent-plane baseline plus a small neural correction; a frozen analytic kernel; an SE(3)- and scale-invariant dimensionless feature stack derived from a 2-term identity \sigma = -(\nabla\phi\cdot\hatn)/(2\pi) + H\phi/(4\pi), exact on flat ground and the grounded sphere and supervised training against this identity. An 8-parameter linear head recovers the analytic curvature coefficient +1/(4\pi) from data to within 2%; holds to 1% relative RMS in \sigma across 12 orders of magnitude in source-charge scale, where DeepONet, Set Transformer, and FNO baselines miss the analytic answer by orders of magnitude out-of-distribution; and matches direct-collocation boundary-element ground truth on a synthetic curved patch to 3.8% relative RMS. A regime-of-validity result delineates when the canonical 2-term identity applies; outside it (terrain curvature radius \kappa^-1 much greater than source clearance height |h|, the corridor regime), we extend the recipe with sparse dense-BEM-target supervision. On several industrial transmission-line corridors with native triangulated meshes, the extended recipe attains 9.2% median relative RMS to dense BEM on held-out corridors, 20× tighter than the analytic 2-term plug-in. Deployment: a 1.7kB MLP, seconds per corridor on a single CPU core.
Abstract: We develop a reduction-based framework for online learning with delayed feedback that recovers and improves upon existing results for both first-order and bandit convex optimization. Our approach introduces a continuous-time model under which regret decomposes into a delay-independent learning term and a delay-induced drift term, yielding a delay-adaptive reduction that converts any algorithm for online linear optimization into one that handles round-dependent delays. For bandit convex optimization, we significantly improve existing regret bounds, with delay-dependent terms matching optimal first-order rates. For first-order feedback, we recover optimal regret bounds via a simpler, unified analysis. Quantitatively, for bandit convex optimization we obtain O(\sqrtd_\texttot + T^\frac34\sqrtk) regret, improving the delay-dependent term from O(\min\\sqrtT d_\textmax,(Td_\texttot)^\frac13\) in previous work to O(\sqrtd_\texttot). Here, k, T, d_\textmax, and d_\texttot denote the dimension, time horizon, maximum delay, and total delay, respectively. Under strong convexity, we achieve O(\min\\sigma_\textmax \ln T, \sqrtd_\texttot\ + (T^2\ln T)^\frac13 k^\frac23), improving the delay-dependent term from O(d_\textnormalmax \ln T) in previous work to O(\min\\sigma_\textmax \ln T, \sqrtd_\texttot\), where \sigma_\textmax denotes the maximum number of outstanding observations and may be considerably smaller than d_\textmax.
Abstract: Abstract concepts - justice, theory, availability - have no single perceivable referent; in the human brain, their meaning emerges from a web of experiences, affect, and social context. Do large language models (LLMs) ground abstract concepts in a similar way? We study this by replicating property-generation experiments from cognitive science on 21 frontier and open-source LLMs. Across models and experiments, we find a consistent pattern: when compared to humans, models rely too heavily on word associations, and underproduce properties tied to emotion and internal states. This yields a large and consistent grounding gap: no model exceeds a Pearson correlation r=0.37 with human responses, compared to a human-to-human ceiling above r=0.9. To better interpret this gap, we also replicate a rating experiment on grounding categories and find that here LLMs align more closely with human judgment, and alignment improves as models get larger. We then use sparse autoencoders (SAEs) to determine whether this information is also reflected in the models' internal features, and identify features connected to grounding dimensions such as "sensorimotor" and "social". These findings suggest that current LLMs can recover grounding dimensions when explicitly queried, but do not recruit them in a human-like way when words are generated freely.
PaperID: 7772, Poster
Abstract: Modern autonomous driving systems widely adopt multi-view Bird’s-Eye-View (BEV) based 3D perception models due to their superior performance. Despite their success, the robustness of these models against adversarial camouflage attacks remains largely unexplored. Existing camouflage attacks primarily target single-view 2D detectors, limiting their effectiveness in multi-view 3D perception settings. To address this gap, we propose BEVCA, a novel adversarial camouflage generation framework tailored for multi-view BEV-based 3D perception. Our framework integrates a BEV-feature-based adversarial loss with a multi-view neural rendering module, enabling effective and transferable camouflage attacks across different models and tasks. Extensive experiments in both digital and physical settings demonstrate that BEVCA outperforms state-of-the-art baselines and exhibits strong robustness under real-world conditions.
Abstract: Transformers have become a central architecture for \emphin-context learning, particularly through their strong empirical performance in large language models. This success suggests that transformers can extract task-relevant structure directly from prompts, even when the underlying data have complex geometric structure. However, existing theoretical analyses of transformer-based in-context learning are largely confined to simplified settings, such as Euclidean domains or single-manifold models. In this paper, we study in-context nonparametric regression under growing geometric complexity, modeled by a sample-size-dependent mixture of manifolds with heterogeneous local geometry. For this model class, we establish a minimax lower bound that captures the aggregate difficulty of its local components. We then construct an oracle tangent local-polynomial estimator and prove a matching upper bound by exploiting local geometric structure. The main technical step is to connect this estimator to a transformer architecture. We construct a two-stage linear-attention transformer consisting of a geometric preconditioner and all-chart reduced local-polynomial solvers, and show that its approximation error is negligible relative to the minimax regression rate. We also prove an in-context generalization bound for \varepsilon-near empirical risk minimizers within this transformer class. Together, the lower bound, oracle construction, transformer approximation, and generalization analysis identify the conditions under which the resulting in-context predictor adapts to local geometry and attains the aggregate minimax rate.
PaperID: 7774, Poster
Abstract: Pretrained sequential recommenders and language models are routinely extended with new entries by jointly fine-tuning old and new embeddings. We show this has a hidden failure mode: old-entry quality degrades while new entries still improve, forcing premature early-stopping. We propose population-specific low-rank subspaces: base and new entries are parameterized with separate shared projection matrices, decoupling their optimization while providing implicit regularization proportional to data availability. Three instantiations (Freeze-SV, Freeze1-SV, Dual-SV) prevent the failure while maintaining or improving new-entry quality. On sequential recommendation (two architectures, two large-scale datasets) our methods Pareto-dominate joint fine-tuning and continual-learning baselines (EWC, ADER) on the base-vs-new quality tradeoff. On LLM vocabulary expansion (three model sizes across two domains), at least one frozen variant wins on overall perplexity in 8 of 9 model-scale cells on both domains, and our low-rank parameterization matches or beats full-rank quality at 4× fewer new-entry parameters (r/d = 25%).
Abstract: In environmental monitoring, data collection is often costly, sparse, and shaped by urgent public-health needs. This is particularly true for cancer-causing PFAS (Per- and polyfluoroalkyl substances) contamination, where discussions with domain experts and environmental organizations highlight the need to strategically identify high-risk, under-observed regions under tight sampling budgets. More broadly, similar challenges arise in disaster response and public health settings, where dynamic environments make it essential to efficiently uncover hidden targets from limited ground truth. Yet sparse and biased geospatial labels limit the applicability of existing learning-based methods, such as reinforcement learning. To address this, we propose a unified geospatial discovery framework that integrates active learning, online meta-learning, and concept-guided reasoning. Our approach introduces two key innovations built on a shared notion of , where uncertainty is modulated by learned relevance from readily available concepts such as land cover and source proximity; and a that promotes semantic diversity during online-meta updates, improving generalization in dynamic environments. We evaluate our framework on PFAS contamination discovery as a real-world inspired environmental monitoring task, demonstrating robust target discovery under limited data and changing conditions.
Abstract: Recent progress in computer vision has produced a wide range of powerful specialized models for detection, segmentation, counting, and other visual tasks. However, these models are usually optimized for isolated task formulations, making it difficult to directly support general-purpose visual intelligence, especially when a task requires complex language understanding and dense small-object perception. In this paper, we propose VisHarness, a trainable visual agent that decouples high-level perception, reasoning, and decision-making from low-level task execution. Instead of training a model to solve a specific visual task, VisHarness learns to harness a set of carefully designed heterogeneous visual experts. This paradigm preserves the general intelligence of the agent while fully leveraging the precision advantages of specialized visual models in concrete visual tasks. With only lightweight training, VisHarness learns a generalizable visual expert-harnessing policy and can solve common fundamental vision tasks under various complex conditions through multi-turn interactions with visual expert models. To enable efficient on-policy reinforcement learning training in a live environment, we introduce dynamic visual memory archiving, which mitigates the rapidly accumulating visual-token overhead caused by multi-turn interactions with visual expert models. Experiments on four representative benchmarks covering reasoning segmentation, generalized referring segmentation, dense small-object detection, and referring counting demonstrate that VisHarness substantially outperforms existing general-purpose models and achieves competitive or superior performance compared with task-specific models.
Authors:
钟 关, Yongjian Guo, Haoran Sun, Wen Huang, shuai di, Xiong J Wu, Likang Wu, Hongke ZhaoAbstract: Asynchronous reinforcement learning improves rollout throughput for large language model agents by decoupling sample generation from policy optimization, but it also introduces a critical failure mode for PPO-style off-policy correction. In heterogeneous training systems, the total importance ratio should ideally be decomposed into two semantically distinct factors: a \emphtraining--inference discrepancy term that aligns inference-side and training-side distributions at the same behavior-policy version, and a \emphpolicy-staleness term that constrains the update from the historical policy to the current policy. We show that practical asynchronous pipelines with delayed updates and partial rollouts often lose the required historical training-side logits, or old logits. This missing-old-logit problem entangles discrepancy repair with staleness correction, breaks the intended semantics of decoupled correction, and makes clipping and masking thresholds interact undesirably. To address this issue, we study both exact and approximate correction routes. We propose three exact old-logit acquisition strategies: snapshot-based version tracking, a dedicated old-logit model, and synchronization via partial rollout interruption, and compare their system trade-offs. From the perspective of approximate correction, we focus on preserving the benefits of decoupled correction through a more appropriate approximate policy when exact old logits cannot be recovered at low cost, without incurring extra system overhead. Following this analysis, we adopt a revised PPO-EWMA method, which achieves significant gains in both training speed and optimization performance. Code at \urlhttps://anonymous.4open.science/r/ROLL-8138/.
PaperID: 7778, Poster
Abstract: A growing failure mode in agent evaluation and training is that models can achieve high evaluation scores by exploiting shortcuts instead of solving the intended task, producing \emphdeceptive performance. This makes evaluation scores unreliable as measures of true task-solving ability. We propose CapCode, a framework for constructing coding datasets with randomized tests whose best achievable non-cheating performance is deliberately capped below one. This capped-performance design gives evaluation scores a clearer interpretation: scores substantially above the cap are implausible and therefore provide evidence of cheating. To prevent cheating, we propose CapReward, a reward design based on the CapCode principle to discourage optimization beyond the cap. Experiments across multiple datasets show that CapCode detects cheating while preserving performance ranking, and CapReward reduces cheating behavior, yielding models that better follow the intended task specification.
PaperID: 7779, Poster
Abstract: End-to-end autonomous driving systems are known to suffer from shortcut learning and causal confusion, limiting robust generalization under distribution shift. To address this, recent Vision-Language Model (VLM) approaches introduce intermediate linguistic reasoning steps including chain-of-thought, chain-of-causation, and counterfactual (CF) reasoning, and report consistent improvements in closed-loop driving. Among these, CF reasoning has been emphasized as enabling self-reflective, high-level causal understanding, yet whether these gains reflect genuine causal reasoning or language-induced distributional bias remains unclear. In this work, we examine what CF reasoning actually induces in VLM-based autonomous driving using accident scenarios with explicit causal structure and a discrete meta-action space. We observe that CF reasoning consistently shifts the action distribution toward more conservative behaviors, accompanied by increased perplexity and policy entropy, and that this effect persists under reinforcement fine-tuning, suggesting a structural rather than transient artifact. However, visual attention analysis and causal intervention experiments show that CF reasoning does not meaningfully alter visual grounding, and VLMs respond similarly to causal and non-causal perturbations, unlike humans who selectively react to causal factors. Overall, these results suggest that CF reasoning does not induce genuine causal reasoning grounded in visual structure. Instead, it acts as a linguistic mechanism that reshapes the action distribution, teaching models what to say but not how to reason, highlighting a fundamental gap between language-based reasoning and causal understanding in VLMs.
Abstract: Behavior Foundation Models (BFMs) enable scalable imitation learning (IL) by pretraining task-agnostic representations that can be rapidly adapted to new tasks. However, existing BFMs assume fixed environment dynamics, limiting their robustness under real-world shifts such as changes in friction, actuation, or sensor noise. We address this by formulating BFM task-inference as a robust minimax optimization problem, enabling adaptation to worst-case dynamics perturbations without modifying pretraining. To the best of our knowledge, this is the first BFM-based framework that achieves robustness to dynamics shifts while relying solely on offline data from a single nominal environment. Our approach significantly outperforms standard BFM and robust offline IL baselines under dynamics shifts. These results demonstrate that robust policy can be achieved entirely at task-inference time, improving the practicality of BFMs in dynamic settings.
PaperID: 7781, Poster
Authors: Alexandre Garcia-Duran, Alejandro Rodriguez-Garcia, Manuel Molano-Mazon, Alexandre HYAFIL, Srikanth Ramaswamy, Janne Lappalainen
Abstract: Biological visual systems compute reliably under tight metabolic and wiring constraints despite noisy inputs and stochastic circuit components. Artificial vision systems tend to invert both properties: They consume vastly more energy and can be sensitive to adversarial perturbations. Whether artificial networks with structural properties of real brain wiring are more robust at lower computational cost is unclear. We address this with FlyNet (Lappalainen et al. 2024), a task-optimized, continuous-time neural network model of the fly visual system constrained by its synaptic connectome. We benchmark it against CNN, RAFT (Teed and Deng, 2020), and a sparse continuous-time recurrent neural network without connectome constraints (CTRNN) on a looming-based collision-detection task within a differentiable 3D rendering pipeline. We find that FlyNet is more robust than CNN controls to flicker, unstructured noise, and motion jitter. Compared with RAFT controls, FlyNet is more robust to flicker and motion jitter but less robust to all gradient-based attacks. Yet, FlyNet uses up to five orders of magnitude fewer FLOPs than RAFT, achieving higher robustness on common video corruptions at far lower cost. FlyNet is less robust than CTRNN across all tested perturbations, suggesting that FlyNet's connectome constraints contribute to cost, not robustness. To probe FlyNet's failure modes, we use counterfactual restoration: textures equally far from training as adversarial textures that recover correct classification also recover position and velocity representations, ruling out unfamiliar textures as the cause of failure. Classical reverse-phi illusions degrade the same downstream position and velocity readouts, suggesting that both perturbations may exploit similar computational vulnerabilities. Together, these results show that FlyNet is robust for its computational cost on an ecologically important task, and indicate that different architectures specialized on the same task have specific robustness profiles across perturbation types. Our differentiable pipeline enables computational tests of this hypothesis and generates targeted adversarial stimuli for in vivo experiments.
PaperID: 7782, Poster
Authors:
huanghaohe zou, Peng Han, Emad Nazerian, Mafu Zhang, Zhicheng Guo, Alex HuangAbstract: Most LLM code-synthesis benchmarks rely on unit tests as the reward oracle, but PCB schematic design has none: correctness is defined by structured physical constraints over real IC packages and pin-level assignments, per-task golden references are unavailable, and SPICE simulation does not validate schematic-level correctness. We introduce PCBSchemaGen, a training-free inference-time framework that turns a frozen LLM into a verifiable, repairable PCB schematic generator. The framework induces a domain schema from IC datasheets to ground LLM decoding, pairs it with a deterministic 5-layer continuous-reward verifier with pin-level error localization, and refines candidates through a Thompson Sampling arm-acquiring bandit. We evaluate on 2 PCB benchmarks covering 227 real-IC tasks across 22 unified circuit domains, including a public-schematic-derived suite that serves as a fully held-out generalization test (verifier, KG library, and prompts frozen before any evaluation). Under our framework, an open-weight 31B model (Gemma-4-31B) passes 81.3% of PCBBench tasks on average, and the same framework transfers across both benchmarks with zero verifier code changes; a Circuitron-style inference-time prompting baseline on the same Gemma-4-31B backbone collapses on hard system-level designs. This suggests inference-time refinement under a deterministic structural verifier is a general recipe for reference-free LLM code synthesis in domains without unit-test oracles. All code, datasets, and benchmarks: \urlhttps://anonymous.4open.science/r/e7165e48e92b.
Abstract: We study adversarial online learning with hidden-convex losses, i.e., nonconvex losses that become convex after a nonlinear reparameterization. Ghai, Lu, and Hazan (NeurIPS 2022) proved that, under geometric and smoothness assumptions, online gradient descent (OGD) on such nonconvex losses approximately simulates online mirror descent (OMD) on the underlying convex losses with a suitable regularizer, yielding \mathcalO(T^2/3) regret. They left open whether the optimal \Theta(\sqrtT) regret from online convex optimization can be recovered in this hidden-convex setting. We answer this question affirmatively. More specifically, via a sharper discrete-time algorithmic equivalence argument, we prove that OGD achieves \mathcalO(\sqrtT) regret under the same assumptions, matching the optimal worst-case rate for adversarial online convex optimization. We also address another open question of Ghai, Lu, and Hazan by clarifying the geometry required for this algorithmic equivalence. We replace the diagonal-Jacobian sufficient condition with a necessary-and-sufficient Hessian compatibility condition, thereby expanding the class of admissible reparameterizations. We complement our tight regret bound with a lower bound showing that the Hessian compatibility assumption is essential for OGD; when it fails, we construct a smooth reparameterization and an adversarial sequence of hidden-convex losses for which OGD suffers \Omega(T) regret. Finally, we extend our analysis to one-point bandit feedback and prove a \mathcalO(T^3/4) expected regret bound for bandit OGD with spherical smoothing, matching its classical rate on convex losses.
Abstract: Recent large language models achieve strong performance on complex reasoning tasks, where reinforcement learning with Group Relative Policy Optimization (GRPO) has emerged as a leading paradigm for optimizing models on self-generated trajectories. However, the on-policy nature of GRPO bounds the model to the reasoning skills it can already produce, restricting to learn more advanced capabilities. Prior works inject privileged reasoning traces from a stronger teacher policy to guide training, yet these traces are inherently out of distribution with respect to the student policy. We observe that this off-policy mismatch causes gradient clipping on semantically critical reasoning tokens, ultimately rewarding correct answers while leaving the reasoning that justifies them unlearned. Hence, we propose Echo-GRPO, a framework that lets the model reason in the words it speaks. Rather than imitating low-probability privileged traces from the teacher model, Echo-GRPO rewrites them into the student policy's own idiolect, that is, its own characteristic vocabulary and expression patterns, while preserving their semantics via Dual-Reference Decoding. We instantiate this framework as VideoEcho-R1 for video reasoning distillation, achieving consistent improvements across three multimodal LLM backbones and five benchmarks. Finally, we show that our idiolectal paraphrasing is a plug-in module that consistently improves both reinforcement learning and supervised fine-tuning frameworks for reasoning distillation, demonstrating that policy-aligned supervision extends beyond GRPO.
Authors:
Yanting Miao, Yutao Sun, Dexin Wang, Pascal Poupart, Mengyu Zhou, Lei Lv, Qi Zhao, Li Wang, Hao Li, xiaoxi jiang, Guanjun JiangAbstract: Visual latent reasoning lets a multimodal large language model (MLLM) create intermediate visual evidence as continuous tokens, avoiding external tools or image generators. However, existing methods usually follow an output-as-input latent paradigm and yield unstable gains. We identify a feature-space mismatch behind this instability: dominant visual-latent models build on pre-norm MLLMs and reuse decoder hidden states as predicted latent inputs, even though these states occupy a substantially different norm regime from the input embeddings the model was trained to consume~\citepxie2025mhc,li2026siamesenorm,team2026attention. This mismatch makes direct latent feedback unreliable. Motivated by this diagnosis, we propose GAP, a Granular Alignment Paradigm for visual latent modeling. GAP aligns visual latent reasoning at three levels: feature-level alignment maps decoder outputs into input-compatible visual latents through a lightweight PCA-aligned latent head; context-level alignment grounds latent targets with inspectable auxiliary visual supervision; and capacity-guided alignment assigns latent supervision selectively to examples where the base MLLM struggles. On Qwen2.5-VL 7B, the resulting model achieves the strongest average perception and reasoning performance among our supervised variants. Intervention probing further indicates that generated latents provide task-relevant visual signal rather than merely extra token slots.
PaperID: 7786, Poster
Authors:
Ruitao Wang, Jinyi Liu, Hongyao Tang, Rong Cheng, Yi Ma, Shaojin Ma, Yiwen Zhu, Hebin Liang, Gangyi Zhao, Pengyi Li, YAN ZHENG, Jianye HaoAbstract: Difficulty perception is essential for adaptive reasoning in large language models (LLMs). Previous studies rely on training auxiliary models or using extra reasoning rollouts to estimate difficulty, which incurs high computational costs. In this paper, we identify an intrinsic property of LLMs: their internal representations, even before explicit reasoning, encode an informative and useful signal that correlates with problem difficulty. Inspired by this property, we propose Referential Latent Difficulty Perception (RLDP), a training-free and rollout-free method that estimates difficulty directly from hidden activations in a single forward pass, and requires only minimal reference problems. Additionally, we introduce RLDP-AdaSwitch, a lightweight controller that dynamically allocates reasoning effort based on the difficulty signals provided by RLDP, enabling efficient trade-offs between accuracy and compute. Our experimental results across multiple LLMs and diverse datasets, including math reasoning, code generation, and QA, demonstrate that RLDP provides stable and effective difficulty discrimination. This further powers RLDP-AdaSwitch to achieve 1.34×–2.00× efficiency compared to rollout-based methods, while matching the performance of training-based methods.
PaperID: 7787, Poster
Abstract: Recent studies have explored attention dynamics in Large Vision-Language Models (LVLMs). However, most existing approaches aggregate attention across heads, potentially obscuring head-specific visual information and weakening visual grounding, leading to hallucinations. In this work, we revisit hallucination from the perspective of how visual information is distributed and utilized across attention heads. We observe that a small subset of visual tokens accounts for most of the attention within each attention head, with these tokens—defined as head-wise visual cues—being complementary across heads, suggesting that effective grounding requires preserving head-specific support. We further find that the proportion of visual cues within visual attention declines during generation, leading to key visual information loss and hallucination. In addition, hallucinated tokens show weaker utilization of long-range textual context compared to correctly generated tokens. Building on these findings, we propose ), a training-free method to mitigate hallucinations. Specifically, we remove redundant visual tokens via head-adaptive de-redundancy, suppress local textual bias through distance-aware attenuation, and reinforce visual cues by redistributing attention. Experiments across multiple LVLMs and benchmarks demonstrate that VerDA consistently outperforms existing methods in mitigating hallucinations, with negligible inference overhead.
PaperID: 7788, Poster
Abstract: Spatial resolved inference (SRI) is a critical technique in biomedical research, which aims to measure RNA sequence abundances by associating them with whole-tissue images at a fine-grained molecular scale. However, the extremely low throughput and distributional discrepancies involved pose significant challenges to cellular morphologies captured in hematoxylin and eosin (H&E) connection benchmarking, and existing SRI frameworks suffer from coverage gaps. In this paper, we propose MollFlow, a novel generative Mollified Flow matching model for spatial resolution inference, which is designed to learn the probabilistic dependency between the distributions of whole-slide images and expression profiles. We innovatively reframe the anchors of flow matching as a sequential informative prior derived from whole-slide images rather than isotropic Gaussian noise. During inference, MollFlow progressively refines the prior trajectory, ultimately converging to biologically plausible inferences. For the challenging coupling, which requires enabling intermediate states to explore pathologically aligned subspaces, we carry out a boundary mollified strategy to stabilize the velocity field, rather than being constrained by rigid trajectories. Endowed by this, our unique modulation generation perspective allows the model to retain coupling alignment for manifold incommensurability and prevent coverage gaps. Empirically, we find that MollFlow is capable of accurately inferring gene expression while exhibiting excellent performance in a variety of application scenarios.
Authors:
Shuaizhi Cheng, Xiang Shi, Zhiwei Zhang, Mingwei LiAbstract: Hypernetwork-based methods such as Doc-to-LoRA internalize a document into an LLM's weights in a single forward pass, but they fail systematically when the document contradicts pretraining knowledge: accuracy drops to 46.4% on the deepest conflicts. We show that this failure is primarily a magnitude problem rather than a representational one. The generated adapter reaches the relevant layers, but its adapter margin is not calibrated to the strength of the contradicted prior; as the pretrained margin grows, deep conflicts lose the override competition. This account predicts that failure should track prior strength. Sorting 194 conflicts by the base model's log-probability on the contradicted fact, baseline accuracy falls from 68% on weak-prior questions to 16% on strong-prior questions, a 52 percentage-point gap. We propose two training-free corrections. Selective Layer Boosting (SLB) scales the adapter at its highest-activity layers, and Conflict-Aware Internalization (CA) applies stronger boosting only when the base model appears confident. Together they raise deep-conflict accuracy from 46.4% to 71.0% on Gemma-2B and from 53.6% to 72.5% on Mistral-7B, while preserving novel-knowledge recall at 97.1%. They outperform vanilla retrieval-augmented generation on Gemma medium conflicts by 18 percentage points, though explicit conflict-aware prompts remain stronger when the user already knows a conflict exists. We release KID-Bench, a 489-question benchmark that separates novel recall, cross-knowledge combination, and prior-graded conflicts.
PaperID: 7790, Poster
Abstract: Training-signal imbalance in medical image segmentation is commonly diagnosed as a data-distribution problem, such as class frequency imbalance, foreground sparsity, or domain shift, and addressed through reweighting, sampling, or domain adaptation. We argue that this diagnosis is incomplete: a more fundamental source is structural misalignment between supervision targets and segmentation objectives. Weidentify two operative mechanisms underlying this phenomenon: gradient-mass dilution, where dominant classes consume gradient capacity needed for foreground discrimination, and inter-class supervision noise, where weak subtype boundaries generate unreliable optimization signals. On a private renal pathology dataset, four standard segmentation architectures trained under conventional four-class supervision underperform a foreground-focused supervision regime (a one-line modification of the loss) by 7.87–16.05 fg-mIoU points. The same effect replicates on three public datasets, suggesting foreground-focused supervision as a low-cost, general-purpose recipe for any pixel-imbalanced segmentation task. Building on this observation, we propose Progressive Signal Calibration (PSC), a framework that calibrates training signals at three abstraction levels. At the annotation level, PSC adaptively selects the supervision granularity that maximizes signal qual ity. At the architectural level, PSC introduces a boundary-aware feature gating mechanism for structure-sensitive feature aggregation. At the regularization level, PSC employs uncertainty- and ratio-aware constraints to stabilize optimization and prevent late-stage class-proportion drift. Importantly, PSC satisfies a con ditional gain monotonicity property: when structural misalignment is weak or absent, its regularization terms naturally attenuate, avoiding systematic degradation. Across 18 public histopathology benchmarks, PSC achieves statistically significant improvement on six datasets with no statistically significant regression on any benchmark. On the private renal pathology dataset, PSC reaches 96.0 fg-mIoU, ex ceeding the strongest matched-backbone baseline by +2.56 points under identical FG3 supervision—a substantial gain in a regime where the FG3 baseline is already above 93 fg-mIoU
PaperID: 7791, Poster
Authors: William Grolleau, Astrid Sabourin, Guillaume Lapouge, Catherine Achard
Abstract: In this paper, we address the Cross-View Object Re-Identification (ReID) task, aiming to match identities across drastically different viewpoints. We focus specifically on Single-View Aerial-Ground ReID (SV AG-ReID), where models train on one viewpoint and must generalize to another at inference. To bridge the viewpoint gap of Cross-View ReID, related work employs generative models to synthesize novel views from source images as additional training data. We leverage recent advances in large-scale single-view 3D mesh reconstruction to lift source images into textured meshes and to render them from new viewpoints. However, this pipeline produces unreliable data in two distinct ways, and addressing both is vital for synthetic data to serve as an effective training signal. First, the reconstruction may fail and produce geometrically incorrect meshes, resulting in severe corruption in rendered novel views. Second, even when the geometry is correct, the rendered novel views contain hallucinated textures in regions not visible in the source image, which may degrade training if used naively. We address both shortcomings by exploiting a signal that prior generative methods do not naturally expose. Indeed, the reconstructed mesh carries a per-pixel partition between regions visible from the source view and regions that the reconstruction model had to hallucinate, available as a deterministic byproduct of rendering. We introduce this per-pixel visibility map and exploit it in a two-stage design. Cross-View Consistency Filtering (CVCF) discards geometrically incorrect meshes by ensuring multi-view consistency between source and novel views over visible patches. Visibility-Aware Attention Supervision (VAAS) leverages the visibility map to guide the model's attention toward reliable regions of the rendered view during training. Combined, these two components enable our method, 3D-VITAL, to better leverage the synthetic signal from single-view 3D mesh reconstruction, benefiting from wider viewpoint augmentation. 3D-VITAL sets a new state-of-the-art on three SV AG-ReID benchmarks (AG-ReID, AG-ReID.v2, MOO). Code and data will be released at \urlurl
PaperID: 7792, Poster
Authors: Pawan Kumar
Abstract: Multi-view 3D vision-language-action (VLA) policies are brittle to small sensor perturbations: depth jitter, partial occlusion, calibration drift and per-view dropout can move the predicted action chunk by far more than the perturbation itself. We trace the brittleness to the fusion stage: concatenating per-view tokens lets a small geometric shift flip arbitrary token indices, turning a continuous-mass-transport problem into a discrete indexing problem. We propose a single substitution: replace concat-then-attention fusion with a differentiable Wasserstein barycentric tokenizer that fuses K view measures into M barycentric tokens by transport. We prove a chain of Lipschitz-style stability bounds that hold conditional on a declared OT perturbation set, calibrated Lipschitz constants, and bounded Sinkhorn-solver residuals. Empirically, the barycenter is the single dominant mechanism behind the gain: (i) the proposed model has the lowest mean clean \ell_2 on synthetic, LIBERO, RLBench and CALVIN ABC→D (3-of-4 with non-overlapping 95% CIs); (ii) a controlled ablation on synthetic regresses by 46× on the corruption-stability metric when the barycenter is removed and by <1.5× when any other component is; (iii) the same mechanism transfers to LIBERO with a +37% clean \ell_2 regression and +62% action-OT regression on real data; (iv) in synthetic closed-loop the concat-fusion baseline collapses from 85% clean to 0% under per-view dropout while the barycenter keeps a 92% worst-case across all five conditions (3 seeds, narrowest CI of all four methods).
PaperID: 7793, Poster
Abstract: Multiple instance learning (MIL) learns from bag-level labels, but a bag’s label is often supported by only a few hidden instances. In heterogeneous bags, these useful instances are rare, diverse, and surrounded by many weak or noisy observations. We propose Sparse Support Grounding, a MIL framework that learns instance influence directly. The key idea is simple: an instance should not matter just because it can be matched somewhere; it should matter when the match provides reliable label evidence. Our method forms a compact, bag-specific evidence support, compares instances with label-consistent supports, and uses unbalanced transport to estimate how much source mass each instance should retain. This retained mass becomes the strength of its grounding signal, so reliable evidence shapes training more than nuisance content, while the bag classifier still uses the full bag. Thus, transport is turned from a matching rule into a way of learning which instances should drive supervision. Experiments on synthetic MIL, whole-slide images, and clinical endomicroscopy show consistent gains, with the largest improvements when sparse diagnostic evidence is embedded in heterogeneous nuisance content.
Abstract: Quantifying variability in a target population relative to a reference population is central to many scientific and clinical problems (e.g., diseased vs. healthy). Yet, without paired data and in the presence of heterogeneous target variation, existing methods struggle to separate multiple modes of target-specific variation. We propose CASL-VAE, a deep contrastive latent variable model that learns structured latent generative factors from unpaired data. CASL-VAE factorizes variation into continuous common latent factors shared across populations and hierarchical salient latent factors that model target-specific heterogeneity as discrete subtypes and continuous within-subtype variation. Using variational inference, we show how approximate joint likelihood optimization over reference and target domains can be performed using unpaired data, providing a principled basis for paired-sample generation and cross-domain analysis. We validate CASL-VAE on semi-synthetic neuroimaging data, demonstrating improved subtype recovery and paired-sample generation compared to baseline clustering and generative models. We also validate its ability to reveal biologically plausible heterogeneity in Alzheimer's disease. Code will be released upon acceptance.
PaperID: 7795, Poster
Authors: Drew Tufto, Daniel Muthukrishna
Abstract: Many scientific instruments produce structured noise that contaminates the signal of interest and cannot be observed in isolation, ruling out standard supervised denoising. We propose BEAM (Background Estimation via Additive Modeling), a generative framework for modeling complex noise in these settings. As ground truth noise observations are unavailable, BEAM learns a generative model of the clean signal and uses it to score candidate noise estimates by how plausible the underlying signal looks. Score matching on signal data turns this into a tractable likelihood that we sample from with Langevin dynamics, which refines a coarse classical estimate of the noise into a more accurate one. We further show that when metadata governing the noise is known in advance, a separate conditional flow model predicts the noise from that metadata alone, enabling pre-observation forecasting of the noise. We instantiate BEAM on stray light in NASA's Transiting Exoplanet Survey Satellite (TESS), where contamination from the Earth and Moon obscures stellar signals. BEAM produces cleaner stray light estimates than the filtering tool used in current TESS pipelines and forecasts stray-light contamination across the observing window for the interstellar comet 3I/ATLAS using only spacecraft--Earth--Moon geometry, supporting prospective scheduling of time-critical observations and a path to recovering faint signals that current pipelines discard.
PaperID: 7796, Poster
Abstract: Partial differential equation (PDE) surrogate modeling often relies on either dense observations or access to governing equations/physics constraints. In practice, observations can be sparse and partial, and the underlying physics may be partially specified or unavailable. We introduce SBNO, an equation-free paired‑data diffusion bridge matching framework that learns conditional operators for both forward and inverse PDE problems under partial observations. SBNO learns a pair of time-dependent forward/backward drifts via diffusion Schrödinger bridge matching, using a unified endpoint coupling that preserves empirical input-solution pairing, enabling both directions to be handled within a single framework. At inference time, SBNO performs deterministic reconstruction by integrating the probability-flow ODE of the learned bridge, eliminating the need for inference-time hyperparameter tuning. On five PDE benchmarks with 5% observations, SBNO achieves superior or comparable accuracy to baselines while demonstrating lower test-instance error dispersion. Furthermore, SBNO is at least 9.7× faster and reduces GPU memory by at least 4.3× compared to diffusion-based baselines.
PaperID: 7797, Poster
Abstract: Collaboratively fine-tuning (FT) large language models (LLMs) over heterogeneous mobile devices fosters immense potential applications of personalized intelligence. However, such a vision faces critical system challenges. Existing federated LLM fine-tuning approaches remain limited by prohibitive on-device resource overhead, straggler-prone synchronous aggregation, and the assumption of a unified pretrained backbone across heterogeneous devices. In this paper, we propose EC-MobiLLM, a novel design for efficient collaborative LLM fine-tuning across heterogeneous mobile devices and heterogeneous backbones. EC-MobiLLM adopts a server-assisted side-tuning paradigm to minimize on-device overhead and pioneers non-blocking asynchronous collaboration to accelerate training. Moreover, EC-MobiLLM introduces adaptive feature alignment to support collaboration across heterogeneous backbones, aligning with practical needs in real-world deployments. Extensive experimental results demonstrate that EC-MobiLLM can maintain robust fine-tuning performance while achieving extremely low on-device memory, with at least 95.2% reduction in computation overhead, 93.2% reduction in communication costs and 5.1× faster convergence compared to existing methods, validating its efficacy for practical LLM adaptation over heterogeneous mobile devices.
Abstract: Current safety alignment of foundation models largely follows a \emphone-size-fits-all paradigm, applying the same refusal policy across users and contexts. As a result, models may refuse requests that are unsafe for general users but legitimate for authorized professionals, limiting helpfulness in specialized professional settings. Existing approaches either require costly realignment or rely on inference-time steering that suffers from imprecise control and added latency. To this end, we propose \textscPalette, a modular, controllable, and efficient framework that selectively relaxes refusal behavior on authorized target domains while preserving standard safety elsewhere. Our method identifies a refusal direction via multi-objective search and internalizes it into the model through lightweight adaptation. \textscPalette further supports modular composition: it learns domain-specific safety controls independently and composes them through parameter merging, enabling on-demand multi-domain authorization without retraining. Experiments across four safety benchmarks, multiple model variants, and both LLMs and VLMs show that \textscPalette delivers precise safety control without sacrificing general utility, offering a practical path toward foundation models that adapt to diverse professional needs.
Abstract: Modern training pipelines for language model agents begin with a supervised fine-tuning stage in which a small student imitates a costly teacher. Recent work mitigates the covariate shift of pure imitation learning by collecting teacher feedback at states the student itself reaches, with a prevailing trend toward elaborate filters on the teacher's responses. We frame this design choice as a budget-allocation problem and compare three constructions of supervised training data: short unfiltered teacher continuations at learner-induced states; full teacher trajectories filtered for success (the rejection-sampling step in recent on-policy expert-correction work); or those further restricted to tasks the student cannot solve on its own. Across three agentic benchmarks (HotpotQA, ALFWorld, and Terminal-Bench-Dev), short unfiltered teacher continuations beat pure behavioral cloning at matched supervision budgets, and on HotpotQA also match or exceed the filtered alternatives. On Terminal-Bench-Dev, this construction at one-tenth of the corpus budget and with no reinforcement-learning stage matches the OpenThoughts-Agent baseline that uses the full corpus together with reinforcement learning; the optimal continuation length is task-dependent. The same supervised-learning checkpoints yield faster early-stage gains under subsequent reinforcement learning, with mixed evidence on whether this advantage persists later in training.
PaperID: 7800, Poster
Authors: Yunxiao Li, Difeng Gao, Yubin Zheng, JIJIAO ZENG
Abstract: Strain engineering for industrial fermentation faces a structural design-test scale gap: combinatorial pathway perturbations of 2–3 enzymes already span 10^4–10^6 candidate strains, while bench-scale fermentation throughput remains on the order of 10^2 runs per laboratory per year. One way to narrow this gap is in silico evaluation by perturbing a learned gene-expression \to phenotype mapping (conditioned on process state); the precondition is that such a mapping can be learned from bench-scale-sized datasets. As a proof-of-concept, we ask whether a genome-scale metabolic model (GEM)-graph mechanistic prior supports learning such a mapping from ~44 bioreactor runs of an engineered Yarrowia lipolytica astaxanthin-producing strain. We propose a spatio-temporal GNN encoder structured as a Mass Flow Graph over the GEM, paired with a multimodal input head over gene-expression snapshots, sensor streams, time-series assays, and run-level metadata. The encoder is pretrained by masked-flux prediction on Flux Balance Analysis (FBA) samples, distilling the GEM's stoichiometric and mass-balance constraints into it before fermentation data are seen. Under 5-fold cross-validation across 5 seeds, the surrogate attains the lowest mean composite test RMSE among the baselines we compare against. Structural ablations show that the GEM-graph prior is necessary—an order-of-magnitude RMSE divergence when removed—while pretraining together with the gene-expression modality make a significant joint contribution, supporting the proof-of-concept claim.
PaperID: 7801, Poster
Authors:
Yi Liu, Xiangyue Zhang, Jia Ma, Jianfang Li, Yisheng HE, Jianqiang RenAbstract: Co-speech motion generation remains challenging over long horizons due to the need for temporal coherence and consistent style. Existing methods typically generate long motions segment by segment, either conditioning on a short seed of recent frames, which gradually drifts, or retaining ever-longer explicit histories, which are costly to maintain and hard to learn from limited data. We trace these limitations to a key asymmetry between past and present: past motion carries slowly varying style that should be compactly summarized, whereas the current segment exhibits fast-changing local dynamics that require fine-grained modeling. This asymmetry motivates a simple Past-As-State, Present-as-Attention principle for long-horizon co-speech motion generation. Following this principle, we propose PASPA, which summarizes past history with an inter-block persistent state and models the current segment with intra-block bidirectional self-attention. We instantiate PASPA within a blockwise flow matching framework, with multi-block supervision for effective training over variable histories, prefill-decode state caching for efficient rollout, and hybrid classifier-free guidance for enhanced history and speech conditioning. Extensive public-benchmark experiments show that PASPA achieves state-of-the-art FGD performance and preserves motion style across long-sequence generation.
PaperID: 7802, Poster
Authors:
Mingyu Jin, Yutong Yin, Jingcheng Niu, Qingcheng Zeng, Wujiang Xu, Xinyuan Song, Wei Cheng, Mengnan Du, Zhaoran Wang, Tianlong Chen, Dimitris MetaxasAbstract: In this work, we investigate how the internal representations of Large Language Models (LLMs) change with inputs of increasing difficulty. Although sparsity changes may not be obvious for individual samples, dataset-level results reveal a consistent statistical trend: as task difficulty increases (e.g., harder questions, longer contexts, or more answer choices), the last hidden states of LLMs become systematically sparser. In short, \emphthe harder the task, the sparser the representation. This sparsity--difficulty relation is observable across diverse models and domains, suggesting that harder inputs drive more concentrated activation patterns in the last hidden state. Through a series of controlled analyses and learning-dynamics experiments, we show that representational density is a learned property of data familiarity: models develop rich, distributed representations for mastered patterns, while unfamiliar inputs default to sparser activations. Our finding is also actionable. We illustrate this with Sparsity-Guided Curriculum In-Context Learning (SG-ICL), which uses sparsity to select few-shot demonstrations matched to the query's difficulty, outperforming standard CoT and Auto-CoT baselines on MATH-500. Our study provides new insights into how LLM representations reflect task difficulty and how this signal can be leveraged for inference.
PaperID: 7803, Poster
Abstract: Time series forecasting (TSF) models suffer from a critical rigidity: the inability to handle Multiple Input and Output Lengths (MIOL) within a single architecture. Current workarounds, such as padding or autoregression, invariably incur computational redundancy or error accumulation. We propose OrangeTree, a linear and tree-based architecture that decouples model parameters from sequence lengths. It employs a Segment Tree Encoder and an Inverse Tree Decoder to translate multi-length input history into multi-length output predictions within a unified framework. Crucially, a Range Weighter and Feature Fuser bridge these components, dynamically selecting and fusing relevant ranges into optimal contexts for forecasting. Experiments on six benchmarks confirm OrangeTree achieves SOTA performance, reducing MSE by 3.4% compared to the previous SOTA. In MIOL settings, it matches the accuracy of length-fixed models without the accuracy degradation and latency increase typical of heuristic methods. Code is available at https://anonymous.4open.science/r/OrangeTreeTSF/.
PaperID: 7804, Poster
Abstract: Open-vocabulary indoor scene synthesis aims to generate plausible layouts from arbitrary user instructions while ensuring physical feasibility and semantic consistency. Existing methods directly infer spatial relations between objects from instructions and solve for layouts based on predefined rules, achieving progress in physical feasibility. However, they often struggle to precisely align with complex instructions due to insufficient comprehension of spatial relations and limited capability to translate them into layouts. In this paper, we propose a Hierarchical Reasoning and Reflection agent framework (HierRR), which leverages a relational hierarchy to enable multi-granularity spatial relation comprehension and attributable layout reasoning, improving semantic consistency while maintaining physical feasibility. Specifically, HierRR introduces the spatial transformation chain‑of‑thought reasoning module that explicitly models the mapping from spatial relations to executable geometric placements, mitigating semantic drift during reasoning. Furthermore, the semantic- and vision-guided reflection module is employed to iteratively refine spatial relations and layout reasoning from local to global along the hierarchy, effectively bridging the gap between language instructions and geometric layouts. Extensive comparative experiments and ablation studies demonstrate that HierRR generates more plausible and semantically consistent indoor scene layouts than existing methods, while generalizing effectively to complex scenes with numerous objects and diverse spatial constraints.
PaperID: 7805, Poster
Abstract: Despite being crucial for effective LLM alignment, data selection remains understudied. Prior work on reward model (RM) training and policy optimization (e.g., DPO, GRPO) identifies example difficulty, the reward gap between chosen and rejected responses, as a key factor, but findings conflict: some favor easier examples with larger gaps, others harder ones. To isolate difficulty from confounders, we assume access to a reference RM and systematically study data selection across RM, DPO, and GRPO training. When difficulty is measured via the reference RM, easier pairs consistently outperform harder ones, especially for smaller base models: using only the top 20% easiest examples often matches or exceeds full-dataset performance while cutting post-training costs 5×. However, this advantage hinges on reward estimation quality. As the difficulty signal is corrupted by noise or estimated with a weak proxy RM, the easy-example advantage shrinks and can reverse. A signal-to-noise analysis explains why: larger-gap examples yield more reliable gradient directions, an advantage that weakens with noisier reward estimates. These results suggest that conflicting findings in prior work partly stem from differences in reward reliability and signal mismatch.
PaperID: 7806, Poster
Authors:
Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva, Anastasia Voznyuk, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev, Serguei BarannikovAbstract: Large language models (LLMs) routinely fail to output the correct option in multiple-choice question answering (MCQA) while encoding the answer internally. We expose this latent knowledge via the Query–Key (QK) score, defined for an attention head as the inner product between the last-token query and the key at the end-of-line token following option i, evaluated before rotary positional embedding is applied. Its argmax identifies a universal class of select-and-copy heads in middle layers that perform option selection through semantic query–key alignment, mechanistically distinct from induction and copy-suppression heads (Olsson et al., 2022): they are invariant to label symbols, and solve a synthetic task with zero surface overlap—properties no positional-copy account explains and that critically require stripping RoPE. Across 24 models from 1.5\mathrmB to 72\mathrmB parameters (LLaMA-2/3/3.1/3.3, Qwen-2.5, Gemma, Phi-3.5, DeepSeek-R1-Distill), a single head's QK-score exceeds the model's own zero-shot accuracy by up to +27.4 pp on HellaSwag and +49.8 pp on HaluDialogue; causal zero-ablation collapses MCQA accuracy to near-random. To remove any dependence on labeled validation data, we introduce an unsupervised HeadScore that ranks heads from unlabeled inputs and recovers the supervised top-k heads on every tested model. Against four positional-debiasing baselines (e.g., PriDe, Wiegreffe, Wang), QK-score is complementary by construction: debiasing re-weights output logits, whereas QK-score reads the model's selection from a middle-layer head before decoding. We release a one-line drop-in \mathttHeadScore script and per-model head indices, making every result one-command reproducible across all 24 models and four benchmarks.
PaperID: 7807, Poster
Authors: Hung Phan, Thuy T Nguyen, Minh N Dinh, Nhat-Quang Tran
Abstract: Time series foundation models (TSFMs) commonly adapt to new data by attaching a single trainable head to a frozen backbone, a one-size-fits-all setup that underfits heterogeneous regimes. Replacing the head with a mixture of experts is the standard upgrade, but on instance-normalized backbones (the dominant TSFM design class) it fails: routing entropy collapses to zero and one expert absorbs every input, a failure we call normalization-induced routing collapse. Standard MoE rescue mechanisms do not repair it, because the cause is in the router's input, not its optimization. Pre-encoder normalization strips the statistics a router would need to tell regimes apart. A mutual-information decomposition makes this precise and yields a signal-ratio that, computed before training, predicts dataset vulnerability (Spearman \rho = -0.88). Eight causal controls, including a vision-modality replication, isolate instance normalization as the cause. The prescription is a minimal causal intervention: Raw-Routed Mixture of Adapters (RR-MoA), which routes on the raw, pre-normalization input. Under a strictly frozen backbone, RR-MoA wins 54/54 comparisons against the strongest fixed adapter and significantly outperforms LoRA, TRACE, AdaMix, and full fine-tuning. The effect generalizes across six backbones and an imputation task. Frozen RR-MoA also beats full fine-tuning by 12–79% (the Frozen Paradox); two architecturally distinct variants confirm the principle generalizes beyond this specific router. Code is provided in the supplementary material.
PaperID: 7808, Poster
Abstract: While the Log-Sum-Exp (LSE) function is a foundational numerical primitive for stable log-domain computation in modern machine learning, its fast and efficient computation on generic platforms (\ie, CPUs and GPUs) has been underexplored, potentially creating execution bottlenecks for large workloads. On GPU, the de facto implementation torch.logsumexp dispatches a chain of nine sub-kernels per call, incurring redundant High Bandwidth Memory (HBM) traffic and launch overhead. On CPU, exp and log operations are hardware-hungry, limiting the computing efficiency. To address these bottlenecks, we propose two drop-in kernels to significantly accelerate LSE computation on both GPUs and CPUs: FuseLSE, a single-pass fused kernel that evaluates LSE's exp and log functions on the GPU's Special Function Unit (SFU); and IntLSE, which further replaces the LSE computation with approximated integer arithmetic for CPU computation. Our experiments, validated by benchmarks across workload sizes and hardware platforms, show that both kernels are ~10-11× faster than PyTorch at small workloads and ~6× faster on larger ones on GPUs, where the speed-up is attributed to reduced launch overhead and HBM traffic. On CPU, where no SFU is available, LSE computation dominates the cost, allowing IntLSE to gain ~3-6× execution speed over FuseLSE.
PaperID: 7809, Poster
Authors: Morteza Mahdiani, Catherine Landry, Jasper J van den Bosch, Frédéric Gosselin, Ian Charest
Abstract: The visual features that support vivid mental images remain poorly understood. Here, we collected more than 229,000 vividness judgments for the 73,000 images in the Natural Scenes Dataset in a large online experiment (n = 1,991 participants). On each trial, participants viewed two images, were cued to mentally recreate one of them for 4 seconds and then rated the vividness of their mental image on a continuous scale from 0 to 100, followed by a memory task to encourage compliance. Using these data, we trained lightweight MLP readout heads on frozen vision-model representations to predict imagery vividness directly from image content. Several models predicted human vividness judgments above a low-level features baseline, with the best individual model reaching r = 0.376 [0.355, 0.396]. A single principal component of model predictions captured most of the cross-model variance and matched or exceeded individual models in predicting human judgments (r = 0.378). Even after removing human-aligned variance, residual predictions remained highly correlated across architectures, indicating a strong shared structure in model predictions. Attempts to extract additional human-aligned signal through residual probes, fresh layer-cached probes, and fine-tuning failed; fine-tuning instead further amplified the shared component. This shared axis also correlated more strongly with image memorability than with vividness itself, while observer-aware models generalized poorly to held-out raters, suggesting that models rely on a generic image-affordance signal rather than a vividness-specific representation. Together, these results show that subjective imagery vividness is partially predictable from visual content alone, while revealing a systematic gap between predictive performance and construct-specific representation in current vision models.
PaperID: 7810, Poster
Abstract: Real-world image dehazing (RID) aims to remove haze-induced degradation from real scenes. This task remains challenging due to non-uniform haze distribution, spatially varying color shifts, and the scarcity of paired real hazy-clean data. In PRISM, we propose Proximal Scattering Atmosphere Reconstruction (PSAR), a physically structured framework that jointly reconstructs the clear scene and scattering variables under the atmospheric scattering model, making the restoration process more interpretable in complex real-world conditions. To bridge the synthetic-to-real gap, we design an online non-uniform haze synthesis pipeline and a Selective Self-Distillation Adaptation (SSDA) scheme for unpaired real-world scenarios, which enables the model to selectively learn from high-quality perceptual targets while leveraging its intrinsic scattering understanding to audit residual haze and guide self-refinement. Experiments on real-world benchmarks demonstrate that PRISM achieves competitive performance on RID tasks.
PaperID: 7811, Poster
Authors:
Jiachen Shen, Xingke Yang, Hui Zhong, Aohan Li, TOMOAKI OHTSUKI, Xin Fu, Miao Pan, Zhu HanAbstract: Personalized fine-tuning of Large Language Models (LLMs) enhances user experience but incurs substantial computational costs. Recent studies demonstrate that initializing training from style-similar fine-tuned models significantly reduces adaptation overhead, turning these models into reusable, tradable assets. However, emerging model marketplaces suffer from inherent information asymmetry: providers may misrepresent model capabilities, while model buyers cannot verify utility prior to purchase, leading to adverse selection and potential market collapse. We introduce FairTune Market to mitigate these risks. Our framework combines a posted-price mechanism with a truncated proper scoring rule, conditioning payments on performance under text-generation verification. This design incentivizes truthful reporting as an equilibrium strategy while bounding downside risk from stochastic evaluation. Compared with strong market baselines, FairTune helps buyers find better-matched starting checkpoints, reducing total downstream fine-tuning cost by 72.4% while recovering 96.3% of the welfare of an Oracle market with perfect style information. It also improves market health: provider misreporting drops by 96.5% (from 0.424 to 0.015), and all providers remain profitable and willing to participate in the default setting. These gains are robust across market scales, evaluation noise levels, and niche-provider scenarios. A semi-real reuse experiment on sentence-transformer embeddings from 10 text domains further shows that FairTune achieves near-Oracle matching in a realistic style-embedding space.
Authors:
Polina Gordienko, Georg Schollmeyer, Frauke Kreuter, Christoph JansenAbstract: Multi-task benchmarks have become a central pillar of machine learning research, yet their growing influence has incentivised benchmark gaming -- strategic actions taken to improve the leaderboard rank of a specific model. Treating datasets as voters and models as candidates, we consider benchmark-specific training -- the inclusion of benchmark data in training -- as a form of election manipulation. For any ordinal benchmark, the problem of choosing datasets to train on so that a target model becomes top-ranked corresponds to shift bribery, a class of manipulation problems from computational social choice. Leveraging this identification, we show that the benchmark-specific training problem is NP-hard under Borda count and mean win rate. Complementing this worst-case perspective, we introduce the instance-level robustness, the minimum number of datasets a model developer must include in training to top a given leaderboard, and derive expressions for it under arithmetic mean, median, mean win rate and pairwise majority. We evaluate these expressions on MMLU under HELM and on BIG-Bench Hard (BBH) under the Open LLM Leaderboard. Across both suites, mean win rate is hardest to manipulate: this gap is clear on BBH (24 tasks, 4507 models), where its median robustness is 22 tasks (92%), compared with 13 (54%) under arithmetic mean and 12 (50%) under median and pairwise majority.
PaperID: 7813, Poster
Abstract: Cooperative multi-agent reinforcement learning under task variability requires both identifying which task an episode belongs to and adapting the learning algorithm to that identity. Recent class-aware methods address the first half of this problem by clustering trajectories into latent task classes and feeding the class label to the agent policy. They leave the second half largely untouched: the Bellman update and the credit assignment performed by the value mixer remain identical across all task classes. We show empirically that, even when task identity is recovered with high accuracy, the Bellman dynamics across classes diverge sharply, with simple classes saturating early and harder classes sustaining large temporal-difference residuals and large cross-agent value disagreement. Building on this observation, we propose Know Your Task, Learn It Right (KYT), a task-aware extension of value factorization that injects per-class learning signal into the Bellman update itself. KYT maintains a lightweight tracker that summarizes the per-class TD-error and cross-agent value spread without any extra network, and uses the tracker in two places: an optimistic target that adds a bounded per-class exploration bonus, and a class-conditioned mixer that modulates credit assignment through a softplus-gated affine layer preserving the IGM property. Across StarCraft II micromanagement benchmarks including SMACv2 and the unit-combination suite SurComb, KYT outperforms strong class-aware and exploration-aware baselines, with the largest gains concentrated on the hardest task classes that current methods underfit.
PaperID: 7814, Poster
Abstract: Standard single-cell differential expression (DE) analysis identifies genes that change across conditions, but it usually overlooks two key sources of heterogeneity that single-cell data are uniquely positioned to reveal: which cells within a donor are truly disease-affected, and how disease effects depend on cell state or subject-level context. Assigning every cell the diagnosis of its donor can contaminate DE signal when disease penetrance is partial, while modeling each gene with a single disease effect cannot capture heterogeneous responses across cell populations. We present PRISM (Phenotype-Resolved Inference in Single-cell Mixed models), a negative-binomial mixed-effects framework that augments standard DE analysis with three outputs: a context-DE vector \theta_g that tests whether disease effects vary along a biological axis z_ij (such as cell state, sex, or age); a cell-level disease posterior q_ij that provides unsupervised disease annotation; and a subject-level disease burden \rho_i. Naive likelihood-only inference of q_ij is not identified onto the disease axis when nuisance variation (batch, cell-cycle, dominant subtypes) competes with disease; we resolve this with a closed-form 1-D Wasserstein-2 projection of the disease-arm posterior onto a bimodal reference marginal that enforces the \healthy, affected\ structure implied by the disease label.
Abstract: Reinforcement learning from verifiable rewards (RLVR) has shown strong promise for LLM reasoning, but typical outcome-based RLVR methods remain inefficient on hard problems. Correct final-answer rollouts are rare, and standard sample-level credit assignment fails to leverage partial reasoning progress embedded in unsuccessful attempts. To address this problem, we introduce SCRL (Subproblem Curriculum Reinforcement Learning), a curriculum reinforcement learning framework built on verifiable subproblems derived from reasoning chains. Given a reference solution, SCRL derives a series of verifiable subproblems and constructs a subproblem curriculum, with the final subproblem fixed as the original problem. This converts partial progress on hard problems into verifiable learning signals. Algorithmically, we propose subproblem-level normalization, a training technique based on RLVR that normalizes rewards independently at each subproblem position within the rollout group. By assigning the resulting advantages to the corresponding answer spans, we enable finer-grained credit assignment without external rubrics or reward models. Our theoretical analysis shows that this subproblem curriculum makes hard problems more learnable by lifting them out of gradient dead zones, with larger relative gains as the original problem becomes harder. Across seven mathematical reasoning benchmarks, SCRL outperforms strong curriculum-learning baselines, yielding +4.1 and +1.9 average-point gains compared to GRPO on Qwen3-4B-Base and Qwen3-14B-Base respectively. On three hard benchmarks (AIME24, AIME25, and IMO-Bench), SCRL further yields point gains of +3.7 in pass@1 and +4.6 in pass@64 on Qwen3-4B-Base, suggesting improved exploration on hard reasoning problems. Further ablations show that the proposed credit assignment is effective and that the gains do not require highly curated subproblems or strong external generators, highlighting SCRL as a practical curriculum framework that enables fine-grained credit assignment for LLM reasoning.
Authors:
Yusuf Kesmen, Fay Elhassan, Jiayi Ma, Julien Stalhandske, Yena Chang, Alexandra V Kulinkina, David Sasu, Akhil Arora, Lars H Klein, Mary-Anne HartleyAbstract: Large language models (LLMs) are increasingly used for conversational clinical decision support, yet they conflate next token prediction with probabilistic decision making. We argue that this conflation reflects an architectural limitation: such systems lack explicit posterior tracking, controllable abstention thresholds, and auditable reasoning chains. We introduce MoBayes, a Modular Bayesian dialogue framework that separates reasoning from language. The LLM acts only as a language interface, parsing patient conversation into structured observations, while a Bayesian module performs probabilistic inference over these observations to update posteriors, select follow-up questions via expected-information-gain and determine when to stop or defer through calibrated decision thresholds. This design enables explicit posterior tracking, controllable selective decision-making, and replaceable population-specific statistical backends without retraining the language model. Across empirical and LLM-generated knowledge bases, MoBayes outperforms standalone frontier LLM doctors, including matched model-family comparisons where inexpensive sensor models paired with MoBayes exceed larger autonomous models at lower cost. The advantage persists under adversarial patient communication styles and across varying diagnostic scenarios. These results suggest that reliable conversational clinical decision support systems should separate probabilistic reasoning from language generation rather than scaling model size alone. Code is available at https://anonymous.4open.science/r/MoBayes/.
PaperID: 7817, Poster
Authors: Hanpei Liu, Samit S Miftah, Dipali Jain, Dan Fiumara, Pradip Bose, Sumit K Mandal, Kanad Basu
Abstract: Chiplet-based GPUs are a promising alternative to monolithic designs for scaling compute and memory resources in machine learning (ML) workloads. However, distributing computation and memory across multiple chiplets introduces a hierarchical Non-Uniform Memory Access (NUMA) environment, where performance depends on the alignment between kernel execution and data placement. In deep neural network (DNN) inference, kernels execute repeatedly over shared weights and activations, and poor alignment leads to sustained inter-chiplet traffic that degrades performance. Existing approaches optimize scheduling or data placement in isolation, failing to coordinate these decisions across dependent kernel executions. We present CHoRD (Chiplet-Hierarchy-oriented Residency and Dispatch), a compiler-assisted, profile-guided framework for joint scheduling and data placement in chiplet-based GPUs. CHoRD reconstructs program structure from compiled kernels and host code, and supports two modes: a static mode that derives scheduling and placement decisions without profiling, and a profiling mode that performs a single lightweight run to collect execution statistics. Using this information, CHoRD determines chiplet-level kernel mapping, enables safe inter-kernel overlap, improves intra-chiplet parallelism, and assigns data to chiplets based on access patterns. These coordinated decisions reduce inter-chiplet communication and improve compute utilization. We evaluate CHoRD on 4-chiplet and 8-chiplet MGPUSim substrates across nine CNN and transformer-based inference workloads. On the 4-chiplet system, CHoRD-static and CHoRD-profiled achieve 2.57× and 2.94× geometric-mean end-to-end speedups over RR, and 1.46× and 1.66× over the placement-only CLAP baseline, respectively, with benefits increasing as chiplet count scales to 8.
PaperID: 7818, Poster
Abstract: A central yet underexplored question in Multimodal Large Language Models (MLLMs) is where to connect — how visual information from a hierarchical vision encoder should be wired into the layer-wise semantics of a large language model. Mainstream MLLMs inject a fixed single-layer visual representation into the LLM and treat all LLM layers as a single consumer, ignoring that different layers demand different visual granularities. Recent multi-layer fusion methods either compress hierarchical features on the encoder side or rely on predefined sparse or hierarchical layer-to-layer connections, and therefore stop short of modeling adaptive cross-layer interactions between the ViT and LLM hierarchies. We address this gap with DGP (Dynamic Gated Pathways), a routing mechanism that establishes all-to-all connections between every ViT layer and every LLM layer, with each LLM layer dynamically gating its preference over visual layers conditioned on its own semantic state. Built on top of LLaVA-1.5-7B and 13B, DGP delivers consistent gains over the LLaVA-1.5-7B baseline (e.g., +4.2 on SQA^I, +3.3 on MM-Vet, +3.1 on VizWiz and LLaVA^W) and surpasses prior multi-layer visual fusion methods on most benchmarks. Layer-wise routing analyses further show that different LLM layers do select different visual granularities on demand, supporting both the effectiveness and interpretability of the proposed dynamic gated pathways. We will open-source code/weight/demo soon.
Authors:
Shubhra Mishra, Yuka Machino, Gabriel Poesia, Albert Q. Jiang, Joy Hsu, Adrian Weller, Challenger Mishra, David Broman, Josh Tenenbaum, Mateja Jamnik, Cedegao (Ced) Zhang, Katie CollinsAbstract: The evolution of mathematics is shaped importantly by interestingness: researchers choose which problems to pursue, and students choose which problems to engage with, based on expectations of interest and challenge. As AI systems, particularly large language models (LLMs) that operate flexibly over natural language and formal mathematics, are increasingly used in mathematics research and education, it becomes crucial to characterize how closely their judgments align with people from different mathematical backgrounds. We study whether LLMs align with human interestingness judgments by comparing LLM ratings with those of two populations, crowdsourced participants with college math experience and International Math Olympiad competitors. Although many LLMs broadly agree with human notions of interestingness, they largely fail to match the distribution of human judgments. They also weakly align with why humans find problems interesting, with low correlation to human-selected rationales. Finally, we evaluate LLMs’ ability to generate interesting problems and find that, after filtering for validity, LLMs are able to generate engaging problems. We conclude with takeaways, including the need for multi-LLM human-AI collaborative systems, that highlight both the promise and current limits of LLMs as partners in mathematical reasoning.
PaperID: 7820, Poster
Abstract: Aiming to modify textual content while preserving the original style, Scene Text Editing (STE) has significant practical value in many applications. STE typically requires paired images before and after editing for training. However, recent approaches rely heavily on synthetic paired data and complex disentanglement strategies. They struggle to faithfully capture the diversify and complexity of natural scenes, often producing artifacts or hallucinations. In this work, we challenge the common practice and investigate whether STE can be driven by turning each real sample into its own paired supervision. Concretely, we explore multiple ways to disrupt the original image (e.g., Crop, Shuffle, Mask) to construct self-paired training data, and build a strong baseline (TextRIRO) upon a conditional diffusion architecture. Our experiments show that combining character-column shuffling with image cropping (Shuffle&Crop) effectively destroys text semantics while preserving key style cues, enabling the model to learn editing behaviors directly from real text images. With this concise yet real-oriented design, TextRIRO achieves promising performance in both accuracy and visual realism across real-world benchmarks, substantially mitigating the synthetic-to-real domain gap. Furthermore, we leverage TextRIRO to build TextRIRO-3M, a large-scale realistic dataset for long-text recognition by concatenating edited and original images. Experimental results show that training recognition models with TextRIRO-3M significantly improves recognition accuracy, demonstrating that TextRIRO provides highly useful training resources for downstream tasks. Overall, TextRIRO lifts STE to realistic, reliable, and widely applicable new levels. Code is provided in the supplement, and data will be released upon acceptance.
PaperID: 7821, Poster
Authors: Mengtian Hong, Jason Lee, Qian Yu
Abstract: Zeroth-order stochastic optimization is a fundamental formulation that arises in real-world design problems where gradients are inaccessible. A central challenge in this pursuit is to design gradient estimators and optimization algorithms under noisy, function-only feedback that exploit local Hessian geometry to achieve optimal sample efficiency. We introduce the Spectrally Grouped Estimator (SGE), a novel gradient estimator that samples over a non-convex union of sphere sections, and utilize it to build an algorithm that achieves order-wise improved Hessian-dependent simple regrets over second-order smooth, strongly convex functions compared to conventional convex-sets sampling baseline methods. We complement these results with the first tight analyses of the baseline schemes, revealing a shared performance bottleneck and thus emphasizing the necessity of non-convex sampling for optimality. We further establish matching converse bounds that tightly characterize the optimal sample complexities within general subclasses of functions sharing the same Hessian at the global minimum, thereby proving the universal optimality of SGE with respect to the Hessian-dependent rates. This fully resolves an open conjecture from the preliminary version of this work. This fully resolves an open conjecture in the preliminary version of this work.
PaperID: 7822, Poster
Authors: Zixin Zhuang, Benjamin Mashford, Dan Andrews
Abstract: Flow cytometry is the primary assay for single-cell immune profiling, yet cross-study modelling remains under-explored. The challenge is two-fold: Studies measure different marker panels, inducing feature-space inconsistency, while batch effects shift signals even on shared markers, making transfer harder than in vision or language where input spaces are stable. We conduct a large-scale empirical study, pretraining an encoder on 89,324 samples from 112 public studies spanning 525 distinct marker panels. We evaluate it frozen --- isolating representation quality from task-specific adaptation, on five clinical tasks under a taxonomy of panel-overlap conditions, including three panel-zero-shot tasks whose marker panels were unseen during pretraining, and one negative control with no expected discriminative signal. We further characterise transfer across scaling axes and probe robustness to marker removal at test time. Despite heterogeneity that disrupts direct transfer, the frozen encoder matches or exceeds in-distribution specialist encoders on 4 of 5 tasks, including all three novel-panel tasks, and behaves as expected on the negative control. Scaling analyses reveal saturation in both encoder capacity and random corpus subsampling, and representations degrade gracefully under marker reduction. These results suggest a pretrained encoder can serve as a general-purpose feature extractor for clinical cytometry data, without task-specific retraining.
PaperID: 7823, Poster
Authors: Wen Dong
Abstract: Modern neural-network training is governed by local interactions: residual streams carry signal, normalization changes scale, attention and state-space blocks route information, augmentation changes which directions generalize, and optimizer state determines the actual checkpoint update. Parameter-space curvature and full-weight posterior views reveal global geometry, but they do not show which layer, block, or interface created it. We introduce a computation-graph view in which parameters \theta and intermediate states u define a local Markov--Gibbs system. A potential energy encodes module relations, data loss, and regularization; its Gibbs law gives a probability reference, and its checkpoint-local Fisher/Gauss--Newton precision matrix gives the curvature geometry. This precision matrix is sparse because variable groups interact only when they share a local module or loss factor. Eliminating intermediate states recovers the reduced parameter curvature. From this common object we prove two results: separator transfers give two-sided criticality certificates for upper stability and active plasticity, and a resolved KL-to-Gibbs bound yields bounded-loss PAC-Bayes certificates on directions explored by training. Checkpoint measurements on CIFAR-10 and GPT-20M/FineWeb-Edu show that these quantities diagnose learning-rate stability, plasticity, and generalization geometry without dense Hessian, Fisher, or covariance matrices.
PaperID: 7824, Poster
Abstract: Online imitation learning (IL), particularly on-policy distillation, has emerged as a strong LLM post-training approach, often outperforming offline supervised fine-tuning (SFT). Yet a principled understanding of when and why online helps remains unclear. In this work, we challenge the view that error accumulation is the main source of online IL's advantage, and instead show that the benefits of online interaction depend critically on whether the setting is realizable, i.e., whether the student policy class can represent the expert policy. Under realizability, we empirically find that offline IL already matches expert performance. In contrast, in non-realizable (misspecified) settings, we prove that offline IL encounters an information-theoretic bottleneck even when horizon H=1, and propose a structural characterization of misspecification relative to the reward, under which online IL provably achieves high performance despite large expert-student discrepancy.
PaperID: 7825, Poster
Abstract: Large language model agents increasingly act through tools, browsers, code interpreters, and external APIs, turning safety from a single-output problem into a trace-level problem. We identify \emphcompositional emergent risk (CER), a failure mode where individually safe actions interact across time to produce unsafe outcomes. We show that bounded-window safety filters can miss CER whenever the risky dependency lies outside their visible context. To study this failure systematically, we introduce \textscCER-Bench, a controlled long-horizon benchmark with 440 tasks and 19,525 actions across 5--500 steps, spanning five risk domains and two compositional mechanisms. At the largest tier, each trace induces a 79M+ candidate risk-composition search space. Across 15 frontier and open-source LLM agents, every model exhibits a non-zero compositionality gap (CG), ranging from 44.4% to 100.0% with a mean of 82.2%. Moreover, 86.1% of compliant executions contain caution language yet still proceed, showing that verbal risk awareness does not reliably prevent unsafe composition. We introduce \textscRiskShield, a conformal-calibrated runtime shield that learns trace-risk boundaries, evaluates planned continuations before commitment, and substitutes risky suffixes while preserving safe prefixes. On 7 held-out \textscCER-Bench models, \textscRiskShield reduces mean \textscCER-5 CG from 69.9% to 0.0%, with no shielded pass observed in 763 evaluations and a 95% upper confidence bound of 1.6%, while preserving 78.4% step-level task completion. It outperforms cumulative-threshold, sliding-window, full-trace, and published safety baselines, with 63.3% of tasks requiring no extra API call and additional robustness across longer traces, risk domains, and external agent-safety benchmarks. Code and benchmark are available at \urlhttps://anonymous.4open.science/r/RiskShield-2284.
Abstract: Multimodal Large Language Models (MLLMs) can listen and see, but how do audio and visual signals actually travel through the network to shape an answer? Despite their growing role in research and real-world applications, the internal pathways through which audio and visual tokens influence the final prediction remain poorly understood. In this study, we examine audio-visual information flow inside Audio-Visual Large Language Models (AVLLMs), tracing how AVLLMs route, utilize, and integrate audio and visual information across two input configurations, audio-visual video and multiple interleaved audio-visual items. We find that for audio-visual video, AVLLMs follow the sequential information flow pathway established for VLMs and VideoLLMs, with audio and visual contribution flowing along this pathway in proportion to the task's reliance on each modality. In settings with multiple interleaved audio-visual items, this routing shifts to different parallel streams. Furthermore, we demonstrate that audio-visual and other token types can be discarded once their information is transferred to LLM, with minimal impact on the model's prediction or even slight improvement, generalizing across multiple tasks and datasets, enabling more efficient inference. These findings hold across multiple models and scales, Qwen2.5-Omni and Video-SALMONN2 Plus at 3B and 7B scales, leading to hypotheses on why these flow structures emerge. Together, these results deliver the first coherent picture of how AVLLMs orchestrate sound and sight inside the network and lay the groundwork for the next wave of interpretability, design, and efficiency advances in audio-visual and broader MLLMs.
PaperID: 7827, Poster
Authors:
Qi Wang, Yunfeng Min, Liwei Huang, Zeyu Zhang, Di Gai, Jun LiuAbstract: Despite the remarkable progress of large Vision-Language Models (VLMs) in Radiology Report Generation (RRG), existing black-box autoregressive generation paradigms lack intermediate clinical logical reasoning. This deficiency renders such models prone to generating linguistically fluent reports with erroneous diagnostic findings, which manifest as plausible hallucinations and thereby undermine clinical reliability. To this end, we propose a novel RRG paradigm termed "Look-Before-You-Leap", aimed at enhancing the clinical logical reasoning capabilities of the model to improve diagnostic accuracy. Specifically, we first construct a self-evolving Radiological Skill (RadSkill) repository to guide the teacher model in distilling high-confidence explicit clinical reasoning, and further internalize this reasoning capability into the implicit diagnostic intuition of the RRG baseline. Building upon this foundation, a novel preference optimization method termed Psychometric Preference Optimization (PSY-PO) is proposed. Grounded in Item Response Theory (IRT) from psychometrics, this method precisely samples clinically misleading dispreferred reports for preference optimization, thereby effectively mitigating plausible hallucinations. Extensive experimental results demonstrate that our proposed method significantly enhances both the generation quality and trustworthiness of radiology reports.
PaperID: 7828, Poster
Authors: Qi Teng, Xueer Wang
Abstract: Sliding-window human activity recognition (HAR) splits continuous sensor streams into fixed-length, overlapping windows and trains a classifier to assign one activity label to each window. Window length and stride are chosen from sampling rate, latency or computation constraints, and assumptions about typical motion duration, but they need not align with motion cycles or activity changes. Some windows therefore contain only part of a repeated motion, fragments shared by related activities, or mixed evidence near a transition. Yet each window still receives a single label, creating a granularity mismatch between activity-level supervision and local motion evidence. We propose PSDNet (Phase-Script Deliberation Network), which separates local phase interpretation from final activity classification. PSDNet encodes the current window into latent phase primitives, combines them with recent history, and produces an initial classification with an uncertainty score. When uncertainty remains high, a deliberation module compares the current phase evidence with learned class-specific phase scripts, i.e., compact templates of short-term phase patterns, and updates classification scores over a small candidate set. An auxiliary boundary objective encourages sensitivity to possible activity changes without serving as a hard routing rule. This design concentrates extra computation where direct classification is least reliable while keeping easy cases efficient. Experiments on eight public HAR benchmarks show that PSDNet consistently improves accuracy and weighted F1 over strong CNN- and sequence-based baselines. Additional analyses on ambiguous windows and similar-activity confusion pairs support phase-aware selective deliberation for windows whose local evidence is insufficient for reliable classification.
Abstract: Self-play reinforcement learning trains language models on their own generated tasks, co-evolving a proposer and solver without human labels. Recent systems report strong reasoning gains, but collapse and instability are widely observed and poorly understood. The dominant response treats this as a reward-design problem. We argue instead that self-play stability is governed by two distinct levers: a data-level gate that decides which proposer-generated tasks enter the training pool, and the reward signal that updates the policy on tasks already admitted. Through controlled experiments on a Python output-prediction task and a deterministic-DSL twin task that strips pretraining priors, output ambiguity, and executor noise, we find the two levers are asymmetric. A strict gate is sufficient for stability under every reward variant we test, including a self-consistency reward with no access to ground truth; while no reward variant is sufficient once the gate is removed. This asymmetry exposes a counter-intuitive coupling we call the Grounded Proposer Paradox: a proposer with ground-truth access accelerates collapse faster than an ungrounded one when paired with a self-consistency solver, by concentrating training on clean tasks that form the fastest path to a spurious self-consistent attractor. Replacing the binary gate with a continuous strictness parameter \varepsilon further reveals a two-stage phase transition: training-side metrics decouple at low \varepsilon, while validation accuracy holds until \varepsilon is much higher. Data-level gating, not reward calibration, is the binding constraint on self-play stability.
Authors:
Chunze Yang, Qidong Liu, Wenjie Zhao, Yue Tang, Jiusong Ge, Di Zhang, Jiashuai Liu, Lei Wu, Junbo Lu, Ni Zhang, Xian Wu, Zeyu Gao, Chen LiAbstract: Whole-slide image visual question answering (WSI-VQA) frames pathology as an extreme-context search problem: to answer a free-form clinical query, a system must first navigate a gigapixel slide under a strict inspection budget to locate sparse, high-resolution evidence. Existing approaches largely fall into two paradigms: i) supervised pathology multimodal large language models (MLLMs) and agents can absorb localization and reasoning into learned modules, but they often couple navigation to task-specific supervision and retraining, limiting their practicality; ii) training-free pathology agents avoid this cost by keeping core models frozen, but often follow a question-first design, i.e. constructing the initial candidate set mainly from query-conditioned relevance. This can miss decisive morphology that is not named in the question, and force heavier inference-time scaffolding. To address this challenge, we introduce PathNavigate, a training-free pathology agent built around a scan-search-readout routine. Before question matching, PathNavigate scans the current slide at low magnification with a shared online memory module over frozen pathology features, producing a slide-specific surprise field that marks an abnormal-region pool. It then applies question-conditioned PLIP relevance only within this pool to select high-magnification search targets. Finally, it extracts local high-magnification evidence and answers with a frozen perceptor-adjudicator stack, using the same online memory as slide-level context. Experiments on WSI-VQA and SlideBench-BCNB show that the proposed scan-search-readout design improves answer accuracy and yields more interpretable evidence-selection trajectories with higher efficiency. To ease reproducibility, we have released the code online.
Authors: Mushan Li, Kihyun Han, Yanyuan Ma
Abstract: Importance weights are essential in domain adaptation under label shift, yet their utility is often undermined by the finite sample uncertainty associated with their estimation. Existing methods typically analyze this uncertainty through Gaussian elimination on interval-valued linear systems, which leads to overly conservative confidence regions and inefficient downstream applications. We propose a paradigm shift from inversion-based inference to a direct matrix constraint framework. We use this framework to define a joint confidence region and extract marginal intervals via linear programming, deriving provably tighter bounds for importance weights while maintaining exact finite-sample validity. Furthermore, we analyze the confidence region's geometry and provide the theoretical results for its diameter bounds. Evaluated across text, image, multimodal benchmarks, including AGNews, MNIST, CIFAR-10, N24News, and a real-world autonomous driving dataset, nuImages, our approach consistently yields shorter confidence intervals and smaller prediction sets than inversion-based methods.
PaperID: 7832, Poster
Abstract: Most modern deep neural networks are trained with Adam and related coordinate-wise adaptive optimizers, which are efficient but do not respect the natural matrix symmetries, geometry, and spectral structure of neural network parameters. We develop a symmetry-based principle for matrix-gradient optimization, showing that orthogonally equivariant optimizers, especially spectral optimizers such as stochastic spectral descent, Muon, Scion, and polar gradient methods, are better aligned with matrix layers than coordinate-wise updates. We show that coordinate-wise adaptivity breaks orthogonal equivariance, discards gradient spectral structure, and can be viewed as applying spectral updates to a pathological diagonal lifting of the parameter space. Motivated by permutation and orthogonal symmetries, we derive symmetry-compatible optimizers for different layer types, including linear layers, embedding and LM head matrices, and MoE routers. This yields one-sided spectral optimizers, row-norm optimizers, and hybrid row-norm/one-sided-spectral variants. Pre-training experiments on dense and sparse MoE language models, including Qwen3-0.6B-style, Gemma 3 1B-style, OLMoE-1B-7B-style, and downsized gpt-oss models, show that replacing AdamW on large vocabulary-indexed matrices and router matrices with symmetry-compatible optimizers consistently improves final validation loss and sometimes training stability.
PaperID: 7833, Poster
Authors: Hiroyuki Kasai
Abstract: Many prediction problems involve a hidden predictive state: a latent task, environment, rule, or data-generating mechanism that changes the query-optimal act. For a fixed query, recovering the full hidden predictive state is often stronger than necessary. The relevant target is the Bayes-act distinction among hidden predictive states that preserves oracle-gain adaptation. This question is central for in-context learning (ICL): a prompt may reveal part of an unseen task or rule, but an ICL system needs only to recover the prediction-relevant quotient state label for the query. We formalize this target through the canonical predictive quotient (CPQ), a problem-relative Bayes-act quotient of hidden predictive states. The theory characterizes the value of adaptation, the quotient that preserves full oracle gain, the sharp zero/positive-gain boundary, and a Bayes barrier showing why context-only prediction cannot recover a positive oracle gain at the risk level. We give two exact examples. A finite-state example makes the quotient structure, sharp positivity boundary, and a strict-loss mechanism explicit. A linear-Gaussian prior-family witness gives closed forms for a collapsed Gaussian baseline and a state-aware oracle, a restricted zero-gain boundary, and an exact prediction-level disagreement identity for the restricted gain. Together, the results give a decision-theoretic target for ICL: not full latent complexity, but the prediction-relevant quotient state label. We support the theory with exact checks, realized bridge experiments over prediction-relevant quotient state labels, and external validity tests on released pretrained in-context predictors.
Authors: Pablo Mercader-Perez, Carolina Cuesta Lazaro, Daniel Muthukrishna, Jeroen Audenaert, V Villar, David W Hogg, Marc Huertas-Company, Bill Freeman
Abstract: Data collected from the physical world is always a combination of multiple sources: an underlying signal from the physical process of interest and a signal from measurement-dependent artifacts from the sensor or instrument. This secondary signal acts as a confounding factor, limiting our ability to extract information about the underlying physics. Moreover, it poses significant challenges for combining data in heterogeneous or multi-instrument frameworks. To disentangle these factors of variation, we propose a dual-encoder architecture with a counterfactual generation objective that leverages overlapping observations. The resulting representations explicitly separate intrinsic signals from sensor-specific distortions and noise, and can be used for counterfactual view generation, parameter inference, and instrument-independent similarity search---all unconfounded by measurement artifacts. We demonstrate the effectiveness of our approach in a multi-instrument setting on astrophysical galaxy images from the DESI Legacy Imaging Survey (Legacy) and the Hyper Suprime-Cam (HSC) Survey. This framework provides a general recipe for scientific self-supervised pretraining: construct training pairs from overlapping observations of the same physical system, treat sensor- or modality-specific effects as augmentations, and learn invariant representations through counterfactual generation.
PaperID: 7835, Poster
Abstract: Predicting spatial gene expression from histopathology images enables transcriptomic analysis on archival H&E samples where matched spatial transcriptomics data is unavailable. Gene expression is organized through cyclic regulatory circuits: transcription factor complexes are themselves products of regulated genes, aggregating upstream gene activity and modulating co-expressed downstream programs (e.g., the ISGF3 complex coordinating hundreds of interferon-stimulated genes through shared cis-regulatory elements). Existing approaches to histology-to-expression prediction, both deterministic and generative, employ generic architectures that ignore this structure. We introduce regulator attention, a mechanism that encodes this cycle by pooling gene representations into a small set of content-dependent tokens via a learned selector and routing them back through cross-attention to modulate gene-level features, in contrast to learnable token arrays whose identity is independent of the input. Within a flow matching framework conditioned on pathology foundation model embeddings, regulator attention consistently outperforms deterministic and generative baselines on correlation-based metrics across four tissue types spanning human cancer, normal tissue, and cross-species samples, with substantial gains on highly variable genes. Beyond predictive accuracy, 87.5% of regulator tokens align with known biological pathways without supervision. Replacing them with content-independent learned queries eliminates this specialization and degrades predictive performance, isolating content-dependent token generation as the causal mechanism.
PaperID: 7836, Poster
Authors:
Zhenze Yang, Xiaowen Cai, Junhao Dong, Guangke Chen, Guiyao Tie, Daizong Liu, Hongyang He, Zhiyuan Ma, Xiang Fang, Dengpan Ye, Jing ZhangAbstract: Large Vision-Language Models (LVLMs) have achieved remarkable multimodal reasoning capabilities but remain critically vulnerable to adversarial perturbations. In practical black-box scenarios, severe architectural discrepancies, misaligned feature spaces, and divergent multimodal fusion strategies often undermine adversarial transferability, preventing perturbations crafted on surrogate models from generalizing to unseen target models. Existing ensemble-based LVLM attacks attempt to mitigate this by integrating multiple surrogates, yet they often fail to bridge the structural generalization gap across heterogeneous LVLM families. In this work, we propose a novel framework that enhances cross-model adversarial transferability via model diversity enrichment and stochastic parameter sampling. We show that transfer success is fundamentally governed by cross-model gradient variance. To mitigate this, we introduce (1) a diversity-induced continuous surrogate manifold that captures cross-model decision variability, and (2) stochastic parameter sampling during optimization to approximate the gradient distribution of unseen LVLM models. We provide a theoretical analysis bounding the expected adversarial transfer gain by gradient variance and prove that stochastic sampling reduces adversarial overfitting to surrogate decision boundaries. Extensive experiments across multiple state-of-the-art LVLMs demonstrate substantial improvements in black-box transferability under strict evaluation protocols.
PaperID: 7837, Poster
Abstract: dLLMs usually denoise by predicting masked tokens independently, but it is unclear when and whether explicit token coupling through data structure can improve their theoretical guarantees and translate into practical gains. In this work, we prove the first learning-error analysis of dLLMs with structured reverse processes based on graphical models, including tree CRFs, bounded-treewidth clique-tree extensions, and Hidden Markov Model probabilistic circuits (HMM-PCs), covering a broad spectrum of text-data structures. Our main result first decouples the dLLM learning error into estimation, optimization, and structural approximation, and then analyzes the role of each term. For the approximation error, a path-sensitive analysis bounds the contribution of an edge missing from the tree by the geometric decay of correlations along its tree-distance detour, yielding a structure-selection criterion that can disagree with Chow--Liu when correlations are weak. We also identify a tradeoff between approximation bias and estimation error, which yields a slowly-increasing preferred width w^\star=\Theta(\log n) as a data-to-structure-capacity scaling regime. From the empirical perspective, we first conduct controlled simulation experiments on synthetic and small-text settings (WikiText-103) to validate the predicted finite-sample signatures of our theory, including exponential path decay on chain MRFs, a bound-optimal tree that disagrees with Chow--Liu, and the balance between different error terms. Then, in the real-world setting, across models ranging from 13M to 1.1B parameters, we show that the same HMM-PC teacher is helpful in from-scratch training, mildly negative in SFT distillation at pretrained scale, and substantially helpful in inference-time guidance, supporting the conclusion that both the structural prior and the deployment interface are critical, with the prior providing token-coupling structure and the interface determining whether that structure transfers positively or negatively at a given scale.
Authors:
Payal Chandak, Victoria Alkin, David Wu, Maya Dagan, Taposh D Roy, Maria C Menezes, Ayush Noori, Nirali Somia, John S Brownstein, Ran Balicer, Rebecca W Brendel, Noa Dagan, Isaac S Kohane, Gabriel A BratAbstract: Medicine is inherently pluralistic. Principles such as autonomy, beneficence, nonmaleficence, and justice routinely conflict, and such ethical dilemmas often sharply divide reasonable physicians. Good clinical practice navigates these tensions in concert with each patient's values rather than imposing a single ethical stance. The ethical values that large language models bring to medical advice, however, have not been systematically examined. We present a framework for auditing value pluralism in medical AI, comprising a benchmark of clinician-verified dilemmas and an attribution method that recovers value priorities directly from decisions. The ecosystem of frontier models spans physician-level value heterogeneity, and models discuss competing values in their reasoning (Overton pluralism) before committing to a decision. However, individual model decisions are near-deterministic across repeated sampling and semantic variations, failing to reproduce the distributional pluralism of the physician panel. Across benchmark cases, these consistent decisions reflect committed, systematic value preferences. While most model priorities fall within the natural range of inter-physician variation, some significantly underweight patient autonomy. A single LLM deployed without regard for its value priorities could amplify those priorities at scale to every patient it serves. Without explicit efforts to balance ethical perspectives with one or multiple models, these tools risk replacing clinical pluralism with a deployment monoculture.
PaperID: 7839, Poster
Authors: Mien B Wang, Liu Mengxing, Kai Hwang, Michael Halassa
Abstract: Human cognition operates under capacity limitations, requiring computational adaptations for solving complex tasks. Although such adaptations have been described algorithmically, their neural substrates and mechanistic implementations remain poorly understood. Here, we study a hierarchical decision-making task in which subjects must simultaneously infer latent rule changes and resolve sensory uncertainty. By comparing human behavior with Bayesian observer models, we identify three systematic deviations: humans make decisions hierarchically, persist in the previous context after a context switch, and preferentially explore a new context within the same sensory modality. We argue that these deviations are jointly explained by a cognitive-constraint perspective in which hierarchical decisions are less cognitively demanding and switching context or modality incurs additional cost. Incorporating these constraints substantially improves behavioral fits and reveals correlated subject-level context and modality stickiness parameters, suggesting a shared underlying cost mechanism. We then asked how neural representations support hierarchical decisions. Because hierarchical decisions require selecting a latent context before choosing an action, context representations provide a natural substrate for reducing the effective decision space. Neuroimaging data revealed that context information is represented in mediodorsal thalamus, consistent with a role for thalamus in task-state compression. Motivated by the context signal in thalamus, we trained a thalamocortical recurrent neural network in which a thalamic module and faster corticothalamic plasticity provided an inductive bias for task compression, while switch-dependent working-memory noise instantiated switching costs. These constrained networks better matched human behaviorial bias than standard task-optimized networks and reproduced key neural signatures observed in the data, including context encoding in mediodorsal thalamus and cue/rule representations in prefrontal cortex. Together, these results suggest that structured biases in hierarchical decision-making arise from cognitive constraints and identify thalamocortical dynamics as a candidate mechanism for efficient task-state inference.
PaperID: 7840, Poster
Abstract: While Reinforcement Learning (RL) enhances Large Language Model reasoning, on-policy algorithms like GRPO are sample-inefficient as they discard past rollouts. Existing experience replay methods address this by reusing accurate samples for direct policy updates, but this often incurs high computational costs and causes mode collapse via overfitting. We argue that historical data should prioritize sustaining local policy plasticity rather than simply reinforcing accuracy. We define local policy plasticity as a persistent ability to remain backtrackable and switchable with respect to recent successful trajectories: throughout training, the policy maintains non-negligible support on recently reachable alternatives, so it can smoothly reallocate probability mass when needed without being locked into a single dominant path. To this end, we propose Dynamic Jensen-Shannon Replay (DyJR), a simple yet effective regularization framework using a dynamic reference distribution from recent trajectories. DyJR introduces two innovations: (1) A Time-Sensitive Dynamic Buffer that uses FIFO and adaptive sizing to retain only temporally proximal samples, synchronizing with model evolution; and (2) Jensen-Shannon Divergence Regularization, which replaces direct replay updates with a distributional constraint to prevent premature over-commitment to a single Rank-1 trajectory. Experiments on mathematical reasoning and Text-to-SQL benchmarks demonstrate that DyJR significantly outperforms GRPO as well as baselines such as RLEP and Ex-GRPO, while maintaining training efficiency comparable to the original GRPO. Furthermore, from the perspective of Rank-k token probability evolution, we show that DyJR improves policy plasticity by reducing over-reliance on Rank-1 tokens and preserving recently reachable alternatives, elucidating how specific sub-modules of DyJR influence the training dynamics.
PaperID: 7841, Poster
Abstract: Offline multi-task reinforcement learning (Offline MTRL) heavily depends on the quality and distribution of pre-collected data. However, existing methods mainly focus on algorithmic optimization, with less emphasis on data-level improvements to enhance learning ability and generalization performance. This paper, from a data perspective, reveals three key bottlenecks that limit Offline MTRL performance:(i) ineffective utilization of prompts length under diverse task complexities, and (ii) semantic irrelevance of randomly sampled prompt segments, (iii) misleading supervision induced by fragmented and discontinuous trajectories. To address these challenges, we propose DaCe-DT, a robust offline MTRL framework designed to be insensitive to heterogeneous task complexities and data quality, featuring length-gated prompt masking (LGPM), retrieval-augmented prompt construction (RAPC), and value-adaptive return calibration (VARC). Together, these mechanisms enable DaCe-DT to deliver data-centric prompt adaptation and trajectory refinement, resulting in robust multi-task generalization and stable policy learning amid heterogeneous offline data and tasks. Experimental results on Meta-World show that DaCe-DT consistently outperforms state-of-the-art methods, achieving an average improvement of 11.73% on optimal datasets and an improvement of 13.34% on suboptimal datasets, demonstrating its effectiveness in learning stably from imperfect data and improving overall multi-task performance.
Authors: Soham Bakshi, Sunrit Chakraborty
Abstract: The problem of model collapse has presented new challenges in iterative training of generative models, where such training with synthetic data leads to an overall degradation of performance. This paper looks at the problem from a statistical viewpoint, illustrating that one can actually hope for improvement when models are trained on data contaminated with synthetic samples, as long as there is some amount of fresh information from the true target distribution. In particular, we consider iterative training on samples sourced from a mixture of the true target and synthetic distributions. We analyze the entire iterative evolution in a next-token prediction language model, capturing how the interplay between the mixture weights and the sample size controls the overall long-term performance. With non-trivial mixture weight of the true distribution, even if it decays over time, simply training the model in a contamination-agnostic manner with appropriate sample sizes can avoid collapse and even recover the true target distribution under certain conditions. Simulation studies support our findings and also show that such behavior is more general for other classes of models.
PaperID: 7843, Poster
Abstract: Class-incremental learning systems trained with \ell_2-regularized cross-entropy and within-task softmax exhibit a structural bias on cross-task discriminators that no amount of data or regularization tuning can remove. We characterize this bias in closed form: at the population optimum, the cross-task discriminator under sequential within-task training differs from the joint-optimal by exactly \Sigma^-1(\bar\mu^(j) - \bar\mu^(i)), an offset arising from the per-task gauge fix imposed by within-task softmax, identical for all C^2 class pairs in a task block---this invariance, not the closed form alone, is what makes the diagnosis falsifiable and the correction one-line. Under a separation condition (\kappa > 1, cross-class distances exceeding inter-task distances) we prove a margin-based bound on task confusion at rate \exp(-(\kappa - 1)^2 \gamma_ij/8). By explicit construction we show this bound cannot be uniform: when separation fails, cross-task misclassification remains \Omega(1) regardless of inter-task margin, and on configurations with extreme within-task class-norm asymmetry it converges to 1. We propose \emphpost-hoc de-centering, an evaluation-time correction that adds the per-task offset \Sigma^-1\bar\mu^(t) to each weight and a corresponding scalar to each threshold; it recovers the Bayes-optimal cross-task discriminator at the population exactly and yields a bound on task confusion that does not require the separation condition. The same operation appears in shared-covariance LDA-style heads (SLDA, FeCAM, NCM); our contribution is the gauge-theoretic derivation identifying it as the unique per-task affine correction to the structural bias of within-task cross-entropy training, the falsifiable empirical diagnostic across five published CIL methods, and the finite-sample analysis. Experiments on CIFAR-100, ImageNet-100, ImageNet-1K, ImageNet-R, and DomainNet across five frozen backbones show post-hoc de-centering recovers 82--94% of the joint--sequential gap (mean 90.5%).
PaperID: 7844, Poster
Abstract: Few-shot CLIP adaptation is often treated as a parameterization problem, where full fine-tuning is assumed to overfit, so parameter-efficient fine-tuning (PEFT) methods are preferred. In this paper, we show that this conclusion conflates parameterization with optimization. Across 11 datasets, 3 CLIP backbones, and 5 shot levels, dense full fine-tuning with plain SGD outperforms representative PEFT methods, whereas Adam/AdamW collapses. We trace this collapse to unreliable support-set preconditioning. The few-shot second moment systematically underestimates the corresponding full-data statistic, already at initialization. Adam’s inverse-square-root normalization converts this error into an oversized adaptive update. Matching Adam’s first-step update budget to SGD resolves most of the collapse, whereas numerator-side alignment does not explain the failure. Shampoo and SOAP also fail, showing that the problem is not merely Adam’s diagonal approximation but a broader mismatch between support-set and full-data preconditioning geometry. Theory formalizes how few-shot estimation error is amplified into unsafe updates. SGD succeeds because it avoids unreliable preconditioners, stays close to pretrained CLIP, and follows broad, connected, low-curvature corridors. These results identify optimizer-induced preconditioning, not dense full fine-tuning itself, as the main source of failure in few-shot CLIP adaptation. Our code is included in the supplementary materials and will be made public.
Abstract: Tokenizer transplant in cross-vocabulary model composition reconstructs donor-only embedding rows as weighted combinations over shared lexical anchors and reuses those coefficients on the base. We identify a structural geometric property of this reconstruction: the same coefficient vector reaches different sets in the donor and base anchor spans, an asymmetric realizability gap. Across 65 donor-base pairs under OMP, with cross-operator validation on CLP, WECHSEL, and FOCUS, we construct breaker tokens: single coefficient vectors that remain statistically inert in the donor anchor span while producing a high-salience reconstruction in the base. The same Gemma-2-2B donor checkpoint admits this construction against 13 different downstream bases drawn from five model families. The planted direction passes weight-merging with a clean reference unchanged. In a deployer case study, standard LoRA fine-tuning suppresses the breaker primarily on prompts whose distribution matches the training corpus and is not a sufficient mitigation against this attack family in our setting. The tested spectral filters miss the asymmetry. We discuss potential misuse in the open-weight composition supply chain.
PaperID: 7846, Poster
Authors: Tianyi Xu, Daniel Pimentel-Alarcón, Zuzana Buřivalová, Claudia Solis-Lemus
Abstract: Passive acoustic monitoring can measure biodiversity at larger scales, but time--frequency annotation of animal vocalizations is expensive, site-specific, and difficult to sustain at scale. We present a label-efficient sound detection framework that combines masked audio pretraining with a lightweight detector on mel spectrograms, then further improves robustness through iterative self-training on unlabeled audio. We first pretrain a ViT-based encoder on unlabeled recordings via masked reconstruction and transfer the encoder to a detection backbone. To better separate animal sounds from confounding background, we add a box-level contrastive loss that pulls matched event regions together while pushing noisy negatives apart. We then apply a two-stage pseudo-labeling curriculum to exploit large unlabeled pools without additional annotation. We evaluate the performance on two ecologically distinct domains: tropical rainforest soundscapes (Indonesia) and bird vocalizations in Mediterranean habitats (Spain). On both domains, masked audio pretraining and contrastive learning consistently improve time--frequency detection under temporal and cross-site distribution shift, and self-training yields further gains in out-of-distribution performance. On the rainforest domain, MAST with self-training achieves +0.22 mAP and +0.24 F1 over the strongest baseline under cross-site shift. On the bird domain, self-training achieves +0.12 mAP and +0.10 F1 over the strongest baseline under cross-site shift. Overall, our results show that MAST can effectively extend self-supervised audio representations from clip-level tasks to robust box-level localization across diverse bioacoustic settings, providing a practical path for biodiversity monitoring with limited labels.
PaperID: 7847, Poster
Abstract: We study N-armed stochastic dueling bandits under the Condorcet-winner assumption, where three widely adopted objectives are considered: best-arm identification (BAI), weak regret, and strong regret. We propose Tree-Guided Identify-Then-Exploit (TG-ITE), the first unified framework to tackle all these objectives to our knowledge. Without requiring stronger assumptions, we propose a shared tree-guided identification approach to find a high-confidence incumbent within O(N) comparisons. We further propose varied exploitation strategies to utilize this warm-start stage to optimize the specific objectives at hand. This methodology enables our approach to (1) achieve O(N) sample complexity in BAI without commonly adopted stronger assumptions; (2) build the first winner-stays-style algorithm to achieve O(N) weak regret; (3) enjoy the same O(N \log T) guarantee as specialized strong-regret approaches; (4) realize the joint optimization of BAI and weak regret with O(N) guarantees for both, eliminating the sub-optimal gap of O(\log N) in the existing approach. Our results provide evidence that the trade-off between BAI and regret minimization is relatively benign in dueling bandits.
PaperID: 7848, Poster
Abstract: Spiking reservoir computing, and reservoir computing more generally, is a powerful and efficient framework for neuromorphic and biological applications, in which a fixed random reservoir drives a trained readout. Yet its performance depends critically on the reservoir initialization. Enriching the reservoir with adaptive, unsupervised plasticity rules offers a natural solution to this limitation. However, it is unclear which rules work, or why, and how to find the best-suited rules for specific tasks without exhaustive and expensive search. Here, we learn how to play the Atari game Pong in a plastic spiking reservoir network. We systematically characterize a large family of local plasticity rules that were meta-learned in prior work. We then show that their performance is predictable and structured: high-scoring rules are characterized by strong differentiation between neurons encoding the ball trajectory and the background, with stable weight dynamics, and consistent readout alignment across time. Rule performance can be estimated from rule parameters alone. Finally, conditioning simulation-based inference (SBI) on high scores allows us to sample directly from promising regions of the rule space, to discover rules that exceed the performance of those found in the prior distribution and reveal stable, high-performing configurations. Together, these results offer competitive performance against classical reservoir computing while providing a transparent, interpretable account of what makes a plasticity rule useful for neuromorphic hardware and biological computing.
PaperID: 7849, Poster
Authors: Zhi Yang, Chengpu Yu, Gonzalo Ferrer
Abstract: Zero-shot object goal navigation (ObjectNav) requires an embodied agent to locate a target object in an unseen environment without any task-specific training. Recent methods leverage large vision-language models (VLMs) to inject semantic priors, but still suffer from two structural issues: (i) semantic cues are treated uniformly across space, ignoring how the spatial extent and hierarchy of contextual entities modulate target likelihood; and (ii) semantic exploitation and geometric exploration are decoupled, producing brittle behavior whenever semantic cues are weak, conflicting, or absent. We address these issues by reformulating navigation as a unified value estimation problem over frontier candidates. Specifically, we propose InfoNav, which jointly models (1) a hierarchical semantic value map that instantiates a spatially-weighted total-probability decomposition of target likelihood through a multi-level spatial influence field; (2) a VLM-based estimator that approximates the intractable long-horizon information gain by grounding it in visual-semantic connectivity; and (3) a memory-augmented verification mechanism that replaces fixed confidence thresholds with a temporally adaptive schedule with deferred re-examination of uncertain detections. InfoNav establishes the highest zero-shot success rate on all three major ObjectNav benchmarks: 62.0% on HM3Dv1 (+0.6 over BeliefMapNav), 81.5% on HM3Dv2 (the only method above 80%), and 41.8% on MP3D (+4.5 over BeliefMapNav), while requiring more than 96% fewer multimodal-model calls than representative LLM-driven baselines. Extensive ablations confirm that the semantic and information-theoretic value fields are complementary, and that the unified formulation is what enables the agent to remain decisive under strong cues and exploratory under weak ones.
PaperID: 7850, Poster
Abstract: Cluttered indoor environments, where large and heavy objects are scattered across diverse surfaces, require humanoid robots to sequentially navigate, grasp, transport, and accurately place each item at its target location within a single uninterrupted episode. This long-horizon, whole-body loco-manipulation task remains a significant challenge for current methods. Previous approaches often suffer from two main issues: easy-reward bias, where training overemphasizes early transport stages at the expense of later ones, and catastrophic forgetting, where focusing on later stages leads to a decline in earlier-stage performance. In this work, we introduce Humanoid Horizon, a unified policy framework designed to overcome these limitations through three interrelated mechanisms. The Parallel Training Strategy expands N scenes into S × N concurrent stage streams governed by a shared policy, ensuring all transport stages receive continuous gradient updates and removing the bottleneck of sequential optimization. The Dynamic Starting Mechanism updates each environment's initial state with terminal states from upstream rollouts, gradually broadening transition coverage and enhancing robustness at stage boundaries. Reward Gating sets the reward to zero in later-stage streams if any previously placed object is displaced beyond a set threshold, thus maintaining object placement throughout the episode without extra reward terms. Collectively, these strategies achieve per-stage success rates exceeding 80% on the LHM-Humanoid benchmark (350 training scenes, 66 held-out scenes), with performance remaining stable even as the number of sequentially transported objects increases beyond two—unlike the sharp drop seen in all baselines. Additionally, we demonstrate that the RL teacher policy can be distilled into a Vision-Language-Action student via DAgger. When provided only with egocentric RGB or depth observations and natural language instructions, the student successfully replicates the teacher's long-horizon, multi-object behaviors, highlighting the potential for perception-driven deployment on real humanoid robots.
PaperID: 7851, Poster
Abstract: In multimodal crack segmentation for industrial facilities, the key challenge is preventing missing modalities from degrading pixel-level performance while keeping computation low. Existing methods struggle to harmonize missing modality effects and to efficiently, adaptively perceive cross-modal topology. We propose Prism, resilient to any modality miss, leveraging implicit structural alignment to harmonize missing data for high-quality segmentation at low computational cost. Prism comprises Modal Reconstructor (MoRe), Prompt-as-Knowledge Distillation (PaKD), a Sparse Gated Pulse Propagation Mixer (PulseMixer), and an Uncertainty-guided Connectivity Propagation Fusion module (UCPF). MoRe utilizes the manifolds of available modalities to probabilistically harmonize missing data, empowered by PaKD to master where to look via prompt distillation. PulseMixer efficiently breaks the static parameterization bottleneck and perceives anisotropic textures via Pulse Propagation Insight (PPI) and sparse gating. UCPF fuses cross-modal cues using uncertainty estimation to produce clear segmentation maps while suppressing background noise. Experiments on three multimodal crack datasets demonstrate strong performance under diverse modality-miss scenarios. On the depth dataset with 90% missing depth modal, Prism achieves 0.8217 in F1 and 0.8478 in mIoU with only 2.62M parameters.
Abstract: Imitation Learning (IL) is a natural framework for learning in sequential decision-making systems and has emerged as the dominant paradigm through which we understand language model training. A central puzzle is that, while in theory offline IL can be horizon-free and optimal, in practice online methods such as on-policy distillation (OPD) often outperform offline methods such as supervised fine-tuning (SFT). We propose a noisy expert model to explain this gap, in which the learner only has access to a noisy version of the expert's policy, but wishes to compete against the reward achieved by a clean expert, motivated by the fact that in many applications, e.g. training language models to perform long chains of thought, the expert is often imperfect. In this setting, we show a sharp separation between offline and online IL. Offline learning from noisy trajectories is fundamentally hard: to compete with the clean expert, the sample complexity must grow exponentially, in contradistinction to the clean expert setting where no explicit horizon dependence exists. In contrast, we prove that online interaction with the noisy expert via a novel variant of OPD enables horizon-free guarantees in some settings and polynomial dependence on horizon in general. Our analysis leads to an alternative loss function form that is commonly considered empirically for LM training. We further provide algorithms and lower bounds, and extend our results to the more realistic setting of unknown corruption when the clean expert is deterministic, thereby providing a theoretical foundation for why on-policy distillation can outperform standard supervised fine-tuning when training language models from imperfect teachers. We complement our theoretical results with experiments on synthetic and natural-language tasks, showing that the OPD variant suggested by our theory outperforms both offline BC and existing OPD objectives under noisy expert feedback.
PaperID: 7853, Poster
Abstract: Retrieval-augmented generation (RAG) is increasingly used to ground Large language Models (LLMs) with up-to-date or private information, and has become a key component for knowledge-intensive LLM systems (e.g., LLM agents). However, RAG with LLMs often hallucinate due to unreliable retrieved information or outdated internal knowledge, making hallucination detection crucial for reliable LLM workflows. Recently, reasoning-intensive LLMs that produce explicit reasoning trajectories before answering have shown promise in RAG by examining and synthesizing retrieved evidence. Yet, they often suffer from overthinking: generating redundant or unproductive reasoning with limited performance gains. While prior work mainly studies overthinking in closed-book reasoning tasks that mainly rely on internal knowledge, overthinking in RAG-based question answering (QA) that requires open-book reasoning over both internal knowledge and external context remains underexplored. Thereby, we conduct empirical studies to investigate the relationship between overthinking and hallucination in RAG-based QA and find: (1) Overthinking in reasoning LLMs produces lengthy reasoning trajectories that are correlated with hallucinations; (2) This overthinking is often associated with knowledge conflict, where LLMs struggle to reconcile retrieved context with internal knowledge. Based on our findings, we propose a knowledge conflict-aware hallucination detection framework. We train a classifier to identify the three knowledge conflict reasoning patterns, design a knowledge conflict score to measure knowledge conflict severity, and apply tailored detection strategies with different prompts for different conflict levels. Unlike existing methods, our approach requires neither LLM parameter access nor retrieval reliability, and only needs the LLM output in a single run. Experiments on three QA datasets and three LLMs confirm our framework's effectiveness.
PaperID: 7854, Poster
Abstract: Classifier-Free Guidance (CFG) is essential for conditional diffusion, yet constant guidance scales often fail to balance alignment with over-exposure and extreme sample issues. Existing heuristic schedules, like guidance truncation, lack rigorous principles and sacrifice classification confidence, while theoretical studies remain limited to constant guidance at final denoising stages, leaving the full sampling trajectory largely unanalyzed. In this work, we provide a theoretical foundation for dynamic guidance by analyzing Gaussian Mixture Models (GMM) within a general SDE framework. We characterize the dynamics of classification confidence through three distinct stages: convergence, mixed, and divergence, and rigorously prove the impact of varying guidance scale \omega at reverse time k. We analyze the Convergence Boundary \Omega(k) as the maximum guidance scale that ensures the reverse trajectory remains within the high-density regions of the data distribution. We prove that during denoising, the permissible guidance scale \omega must decay to prevent the trajectory from escaping into the divergent phase, especially at the end of the denoising. Although different noise schedules have different decay rates, we found that VP(Variance Preserving) has more stringent decay rate requirement compared to VE(Variance Exploding), which is why VP is most prone to overexposure. Specifically, \Omega(k) simplifies to an exponential decay O(e^T-k) in VPSDE and an algebraic decay O(T-k) in VESDE. This provides a unified explanation for why constant guidance inevitably triggers artifacts and demonstrates that guidance truncation is essentially a coarse approximations of this stability limit. We can achieve dynamic optimization based on the decay rate, which better restricts the scale within the convergence region, thereby effectively eliminating overexposure while maintaining excellent semantic alignment and improving generation quality.
PaperID: 7855, Poster
Abstract: Agentic memory research has accumulated extraction, supersession, reflection, and graph synthesis on top of retrieval-augmented generation. Yet on the field's three established benchmarks, almost none of those primitives earn their cost: We show this with Bridge, a deliberately minimal RAG memory with chained observation notes, hybrid BM25 + cosine retrieval, and chronological rendering, which reaches parity within bootstrap CI with the prior published SOTA on LongMemEval-S and BEAM, and beats the best memory baseline on EverMemBench by +5.7-6.3pp at matched reader. It runs at ~5-7K reader-prompt tokens per query and a single LLM call for each ingestion or query. We argue this parity is itself the result: standard benchmarks score what memory retrieves, not what memory enables when the agent acts. To test the latter, we introduce five adversarial probes (Rule Application \pm, Pattern Induction style / procedural, Dynamic Multi-turn) judged at generation time by a cross-vendor verification chain; each can only be solved by applying a remembered rule or pattern the probe never mentions. Across 5 probes × 20 scenarios, no architecture wins universally: atomic-fact extractors win rule application (LangMem leads Bridge by +45pp on RA^+, with non-overlapping bootstrap CIs), while Bridge's block-summary memory wins pattern induction (leads the strongest baseline by +33pp on PI-P). The right memory architecture is mechanism-dependent, and the field's evaluation paradigm fails on both ends: standard benchmarks fail to distinguish a minimal memory from a complex one, and application-grade probes that do distinguish find no universal winner.
Authors: Nick Bettencourt, Xiaowei Ding, Kay Giesecke
Abstract: As high-quality public web corpora become increasingly exhausted, clean long-context documents have become a scarce and expensive source of training data for large language models (LLMs). Existing long-context corpora are often proprietary and costly to acquire, synthetically generated, or concentrated in narrow domains such as programming. We introduce the Stanford EDGAR Filings Dataset (SFD), an open reconstruction of SEC filings into layout-faithful MultiMarkdown for financial language modeling and evaluation. SFD makes audited financial statements, risk disclosures, ownership reports, accounting notes, and market-moving event filings usable as long-context pretraining data and as a basis for financial reasoning, forecasting, compliance, and document understanding. The resulting corpus is token-efficient, model-ready, and has less than 0.1% overlap with Common Crawl-derived corpora. We release SFD-v1, a 152B-token snapshot of a larger 18.4M-filing corpus estimated at 500B tokens. We further introduce two SFD-derived benchmarks: EDGAR-Forecast, which evaluates filing-grounded numerical forecasting after model knowledge cutoffs, and EDGAR-OCR, which evaluates transcription of complex financial tables.
PaperID: 7857, Poster
Abstract: Dynamic Scene Graph Generation (DSGG) aims to capture objects and their evolving relations in videos. Despite recent progress, the practicality and quality of generated scene graphs remain limited compared to the rapid advances in Multimodal Large Language Models (MLLMs). In this work, we revisit DSGG from two fundamental perspectives: task setup and model design. From the task setup perspective, we identify two key limitations of the current recall-oriented evaluation protocol: (i) a severe precision–recall trade-off, and (ii) uninformative and redundant relation generation. To better assess practical usefulness, we introduce four additional metrics that provide a more comprehensive evaluation of the quality of generated dynamic scene graphs. From the model design perspective, we explore directly using MLLMs for scene graph generation and establish a strong MLLM-based DSGG baseline through three design changes. First, we replace the conventional bottom-up pipeline with a top-down reason-then-locate strategy. Second, we reformulate frame-wise dynamic graphs as Temporal Relation Set (TRS) prediction, improving both efficiency and performance. Third, we introduce Importance-Aware Finetuning (IAF) to encourage more relevant and diverse relation generation. Extensive experiments on Action Genome, VidVRD, and PVSG show that our approach consistently achieves state-of-the-art performance across multiple metrics.
PaperID: 7858, Poster
Authors:
Guangjian Zeng, Siyuan Tao, Chengzhi Zhao, Xiaoqing Li, Na Tang, Wenting Huang, Shenying FangAbstract: Multiple instance learning (MIL) has been increasingly used to analyze histopathology whole slide images (WSIs) and has achieved accurate diagnostic classification and a certain degree of automatic segmentation capability. However, segmentation performance without annotation remains suboptimal due to the immense size of WSIs, despite its essential role in assisting pathologists’ diagnosis. Moreover, existing approaches treat classification and segmentation as independent tasks, with segmentation primarily serving as an interpretability section. To address this limitation, we propose a novel weakly supervised learning algorithm that unifies classification and segmentation within a framework. Specifically, we explore the interconnections between patches and introduce class activation maps in WSI analysis through the proposed Sparse-CAM backbone to activate the segmentation capability, and further enhance feature representation using the proposed subset-matching module to improve both classification and segmentation performance. The proposed method demonstrated superior performance in diagnosis and achieved state-of-the-art performance in segmentation compared to other MIL methods across multiple datasets. In addition, it offered high-precision segmentation that is sensitive to micro-tumor regions without pixel-level annotations. The code will be publicly available on GitHub upon acceptance.
PaperID: 7859, Poster
Authors: Gabriel Moryoussef, Alon Kipnis
Abstract: We consider classification in high-dimensional feature spaces where only a vanishing fraction of features are useful and the number of classes may be large. While this rare-useful-feature regime is well understood in the binary setting, its multiclass version raises new questions: a feature can be useful for class discrimination without separating any specific pair of classes. We focus on a rare feature diversity model, in which most features are useless, but a small fraction promote discrimination among classes in a way that can be detected by a per-feature diversity test, such as one-way ANOVA or the Kruskal--Wallis test. For this setting, we propose Diversity Pursuit Higher Criticism (DP-HC), a feature-screening procedure that applies a higher-criticism threshold to the collection of p-values obtained from per-feature diversity tests. We justify DP-HC under a multiclass rare/weak Gaussian feature model with class-mean contrasts of fixed magnitude in random directions, extending the binary model of Donoho and Jin (2008). When the number of classes is fixed, our model recovers the binary phase transition. When the number of classes grows logarithmically with the number of features, however, we derive a new and strictly harder phase transition, driven by a second-order variance term that the fixed-class analysis does not capture. We show that whenever classification beyond chance is possible, DP-HC with ANOVA, followed by a simple classifier based on similarity to class means, achieves asymptotically perfect accuracy. Experiments on real multiclass classification benchmarks demonstrate substantial improvements, and numerical simulations confirm that the derived phase-transition curve accurately predicts the finite-sample behavior of DP-HC in the growing-class regime.
Abstract: Offline RL can learn effective policies from fixed datasets, but deployment objectives often shift after training, while the trained actor may be frozen due to data, cost, or governance constraints. We study deployment-time adaptation of frozen offline actors using Product-of-Experts (PoE) composition with a goal-conditioned prior. Our main practical finding is graceful degradation rather than universal improvement: with degraded or random priors, precision-weighted PoE remains anchored to the frozen actor, whereas additive and prior-only adaptation often collapse, and a KL-budget selector typically reaches a near-oracle operating point. We also clarify that, for diagonal-Gaussian actors and priors, PoE with coefficient \alpha induces the same deterministic policy as KL-regularized adaptation with \beta=\alpha/(1-\alpha), with posterior covariance differences cancelling under deterministic action selection. Across D4RL MuJoCo and AntMaze diagnostics, results reveal an actor-competence ceiling: composition helps or preserves performance when the frozen actor is sufficiently competent, but cannot recover when the actor itself fails. Swept CQL/IQL-guided and SF/GPI baselines are competitive in some settings, but do not eliminate this regime boundary. Thus, PoE/KL-Reg is best viewed as an actor-anchored safety layer for deployment-time objective shift, not as a general return-maximization method, with effectiveness governed by actor competence rather than knob choice.
Abstract: Speech large language models (SLMs) are typically built from text large language model (TLM) checkpoints, yet they still suffer from a substantial modality gap. Prior work has mainly attempted to reduce this gap from the output side by making speech generation more text-like, but the gap remains. We argue that the key remaining bottleneck lies on the input side. We propose TextPro-SLM, an SLM that makes spoken input more closely resemble that of a prosody-aware text LLM. TextPro-SLM combines WhisperPro, a unified speech encoder that produces synchronized text tokens and prosody embeddings, with an LLM backbone trained to preserve the semantic capabilities of the original TLM while learning paralinguistic understanding. Experiments show that TextPro-SLM achieves the lowest modality gap among leading SLMs at both 3B and 7B scales, while also delivering strong overall performance on paralinguistic understanding tasks. These gains are achieved with only roughly 1,000 hours of audio training data, suggesting that reducing the modality gap from the input side is both effective and data-efficient.
PaperID: 7862, Poster
Abstract: Reinforcement learning with verifiable rewards (RLVR) often adopts GRPO-style group-relative updates, sampling multiple rollouts per prompt to construct normalized learning signals. However, merely increasing the number of rollouts does not reliably strengthen learning: under GRPO-style group normalization, per-rollout policy-gradient features can concentrate into a low-rank, signed geometry, causing substantial cancellation during aggregation and weakening the effective update. We address this failure mode with SALT, a Subspace-Adaptive geometry pLug-in componenT that uses sample-wise gradient geometry to reweight the coefficients of group-relative updates. SALT estimates a dominant shared subspace from the mini-batch Gram geometry, decomposes group-relative coefficients into shared and residual channels, and adaptively amplifies the residual channel when signed cancellation is severe. Across diverse reasoning-oriented RLVR benchmarks and model scales, SALT improves effective update geometry and performance without modifying the reward model or the rollout sampling procedure.
PaperID: 7863, Poster
Authors: Mihaly Petreczky, Mohamad Al Ahdab, John Leth
Abstract: In this paper we derive a Probably Approximately Correct (PAC)-Bayesian error bound for partially observed linear time-invariant (LTI) stochastic dynamical systems in state-space form with inputs. Such bounds are widespread in machine learning, and they are useful for characterizing the predictive power of models learned from finitely many data points. The bound derived in this paper relates the expectation of prediction errors with the prediction error generated by the model on the data used for learning. In addition, we show that it can also be used to derive bounds for the parameter estimation error. In turn, this allows us to provide finite-sample error bounds for the prediction error and parameter estimation error for a wide class of system identification algorithms. Furthermore, as LTI systems are a sub-class of recurrent neural networks (RNNs), these error bounds could be a first step towards PAC-Bayesian bounds for RNNs.
Authors:
Lit S Tan, Junzhe Chen, Xiaolong Fu, Lichen Ma, Junshi Huang, Jianzhong Shi, Yan Li, Lijie WenAbstract: Test-time scaling (TTS) improves text-to-image (T2I) generation in unified multimodal models (UMMs), but existing TTS methods typically rely on frozen-parameter search or sampling, yielding only ephemeral, instance-level improvements. We propose Meta-TTRL, a metacognitive test-time reinforcement learning (TTRL) framework that converts test-time generation experience into learning signals, enabling UMMs to self-improve during T2I generation without external reward models. Meta-TTRL formulates test-time learning as a closed-loop interaction between an object-level generator and a meta-level introspector. The introspector decomposes prompts into structured verification rubrics and produces confidence-enhanced intrinsic monitoring signals, which are used to optimize the generator policy through reinforcement learning (RL). Extensive experiments demonstrate that Meta-TTRL generalizes well across three representative UMMs, including Janus-Pro-7B, BAGEL, and Qwen-Image, achieving significant gains on compositional reasoning tasks and multiple T2I benchmarks with limited data. Further analyses with external introspectors, RL leakage, and alternative monitoring signals reveal a key principle for effective TTRL: metacognitive synergy, where monitoring signals must align with the model's own optimization regime.
Abstract: Compositional priors describe the generic properties of layered functions in deep Bayesian models, where deep neural networks with random weights are a canonical example. In the wide-network limit, the prior is a Gaussian process with a depth-dependent kernel, and its behaviour as depth grows has been extensively studied through this kernel. Here, we study another case, where each layer itself is a vector valued Gaussian process, and our aim is similarly to understand the limiting behaviour of the prior as depth grows. Previous GP work has established that for the RBF kernel and a certain range of bandwidths r, the prior degenerates in the limit, converging to the set of constant functions --- which is not useful as a probabilistic model. In this paper we establish several new results. First, we identify a sharp bandwidth threshold r_c(d) = \Theta(\sqrtd) above which the limit is degenerate, strengthening the earlier bounds. Second, and more importantly, we show that for r below the threshold r_c(d) the prior converges to a limit distribution \pi_\barZ. We also prove that these distributions are non-degenerate and non-Gaussian, with non-vanishing dependence between coordinates. In contrast to the previously known degenerate regime, deep Gaussian process priors can therefore admit non-trivial limits. Empirically, we verify the threshold across a range of dimensions d, and demonstrate a complex multimodal behaviour of the limit distributions \pi_\barZ --- a regime that becomes increasingly narrow with d and would be hard to identify without knowing the threshold.
Authors: Ryo Yoshida, Takami Sato, Wenlun Zhang, Yuki Hayakawa, Shota Nagai, Kentaro Yoshioka
Abstract: LiDAR sensors are critical for autonomous driving perception, yet remain vulnerable to spoofing attacks. Jamming attacks inject high-frequency laser pulses that completely blind LiDAR sensors by overwhelming authentic returns with malicious signals. We discover that while point clouds become randomized, the underlying full-waveform data retains distinguishable signatures between attack and legitimate signals. In this work, we propose PULSAR-Net, capable of reconstructing authentic point clouds under jamming attacks by leveraging previously underutilized intermediate full-waveform representations and simultaneous laser sensing in modern LiDAR systems. PULSAR-Net adopts a novel U-Net architecture with axial spatial attention mechanisms specifically designed to identify attack-induced signals from authentic object returns in the full-waveform representation. To address the lack of full-waveform representations in existing LiDAR datasets under jamming attacks, we introduce a physics-aware dataset generation pipeline that synthesizes realistic full-waveform representations under jamming attacks. Despite being trained exclusively on synthetic data, PULSAR-Net achieves reconstruction rates of 92% and 73% for vehicles obscured by jamming attacks in real-world static and driving scenarios, respectively.
PaperID: 7867, Poster
Abstract: Text-to-image diffusion models fail on semantically rare attribute--object compositions, but it is unclear whether this brittleness arises during denoising or is already present in conditioning representations. We study this question at three levels of granularity. First, using a controlled stress protocol across SD1.5, SDXL, and SD3, we show that a Compositional Stress Index derived from text-side non-factorization residuals separates anchor from stress prompts with perfect discrimination (AUC=1.0) cross-model, establishing a robust pre-denoising risk axis. Second, embedding-level low-rank repair reduces representation residuals without any positive semantic improvement, while a downstream cross-attention intervention succeeds on the same subset---a site-sensitive boundary between diagnosis and control. Third, we decompose this boundary at block-group resolution under three intervention modalities. Diagnostic accessibility peaks at the deep encoder (D2) across all conditions, but token-level intervention of either sign produces no measurable improvement at any block. Only broad cross-attention ablation moves semantics at D2, while decoder blocks respond to selective boost and collapse under ablation. Together, these results establish that compositional defects are diagnosable before denoising, but where to look is not how to fix: the diagnostic site and the control site require different intervention modalities, and even at the correct site, the wrong modality is futile.
Abstract: Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a dominant paradigm. Although recent LLM-based ASR models have shown promising performance on public benchmarks, it remains challenging to balance recognition quality with latency and overhead, while hallucinations further limit real-world deployment. In this study, we revisit LLM-based ASR from an entropy allocation perspective and introduce three encoder-side diagnostics to characterize how training paradigms shape uncertainty reduction across the speech encoder-LLM interface. To remedy entropy-allocation inefficiencies in prevailing approaches, we propose a capability-boundary-aware multi-stage training strategy that targets parameter efficiency and robustness to hallucinations. Specifically, we redesign the pretraining strategy to alleviate the speech-text modality gap, and further introduce an iterative asynchronous SFT stage between alignment and joint SFT to preserve functional decoupling and constrain encoder representation drift. Experiments on various benchmarks show that our method achieves competitive performance against state-of-the-art models using only 2.3B parameters, while also effectively mitigating hallucinations through our decoupling-oriented design.
PaperID: 7869, Poster
Authors: Nischal Subedi, Cencheng Shen, Peng Zhao
Abstract: Low-Rank Adaptation (LoRA) constrains weight updates to a low-rank subspace, making the choice of that subspace central to adaptation. We show that the value of LoRA initialization direction is regime-dependent. It helps when the representation dimension is large relative to the labeled set size and the estimated activation subspace is aligned with task structure, and is otherwise inconsequential. We formalize this view with a spiked-covariance Davis--Kahan argument and a first-step LoRA analysis showing that the initial learning signal flows through projected target-token features. From this analysis we derive an alignment coefficient \rho, computable from a single forward pass, that closely tracks the empirical gain from direction-aware initialization across tasks. Motivated by this view, we propose PivotLoRA, an initialization that estimates LoRA adapters from target-token PCA and applies a single output-preserving orthogonal pivot after early training drift. Across 30 BERT-base few-shot settings and 13 compared methods, PivotLoRA achieves the best average rank and improves over vanilla LoRA on nearly every setting, with the largest gains concentrated on high-alignment tasks. Decoder-scale classification on LLaMA-2-7B and Llama-3.1-8B reproduces the low-shot pattern, while full-data and generation controls confirm the predicted regime specificity.
Abstract: GNNs have become a standard tool for learning on relational data, yet they remain highly vulnerable to backdoor attacks. Prior defenses often depend on inspecting specific subgraph patterns or node features, and thus can be circumvented by adaptive attackers. We propose PRAETORIAN, a new defense that targets intrinsic requirements of effective GNN backdoors rather than surface-level cues. Our key observation is that flipping a victim node's prediction requires substantial influence on the victim: attackers must either inject many trigger nodes or rely on a small set of highly influential ones. Building on this observation, PRAETORIAN (i) analyzes internal correlations within potential trigger subgraphs to detect abnormally large injected structures, and (ii) quantifies external node influence to identify triggers with disproportionate impact. Across our evaluations, PRAETORIAN reduces the average attack success rate (ASR) to 0.55% with only a 0.68% drop in clean accuracy (CA), whereas state-of-the-art defenses still yield an average ASR of \geq20% and a CA drop of \geq3% under the same conditions. Moreover, PRAETORIAN remains effective against a range of adaptive attacks, forcing adversaries to either inject many trigger nodes to achieve high ASR (\geq 80%), which incurs a >10% CA drop, or preserve CA at the cost of limiting ASR to 18.1%. Overall, PRAETORIAN constrains attackers to an unfavorable trade-off between efficacy and detectability.
PaperID: 7871, Poster
Authors: Fereshteh Abedini, Soheil Damangir
Abstract: Cell-free DNA (cfDNA) liquid biopsy has been clinically routine for cancer detection and monitoring for years but has yet to produce applications in chronic diseases of parenchymal organs such as fibrotic liver, kidney, and lung disease, despite diseased organs shedding orders of magnitude more cfDNA into the bloodstream than tumors do. This gap is structural: methods developed for cancer have been adapted for chronic disease without a theoretical framework establishing when this should work or where it should break down. We formalize cfDNA classification as recovering a disease-induced shift on a low-dimensional submanifold of the simplex over molecular configurations, observed through sparse genomic windows. The formalization reveals that cancer and chronic disease occupy structurally distinct regimes parameterized by one dimensionless quantity: the inverse participation ratio c(v) of the disease perturbation across windows. Cancer is small-c(v): signal on a few genomic hotspots. Chronic disease is large-c(v): signal spread across many co-regulated regulatory programs. We prove an information-theoretic lower bound: any classifier observing k windows recovers at most a \sqrtk/c(v) fraction of the available information in the local-asymptotic-normality regime, regardless of per-cfDNA model complexity. The bound applies architecture-agnostically to classifiers that focus deep computation on a small number of pre-selected genomic regions, i.e. the architectural shape that has succeeded in cancer detection. Such classifiers degrade sharply rather than gradually as c(v) grows, with the transition provable from the geometry of the problem. Chronic-disease classification requires an architecture matched to its regime: compute allocated across a very large set of windows rather than into deeper per-cfDNA models, with k \gtrsim c(v). We propose a concrete architecture and verify on synthetic experiments that the regime transition predicted by the theorem is observed, and the architecture outperforms cancer-style baselines on chronic-like signal. We then validate on real patient data, where the architecture exceeds both cancer-tuned methods and standard-of-care clinical tests.
PaperID: 7872, Poster
Abstract: RAW image denoising and demosaicing play a critical role in the early stages of the image signal processing (ISP) pipeline. Previous methods handle them sequentially, which can introduce error propagation. Unified restoration, such as joint denoising and demosaicing (JDD), has become a widely adopted paradigm and achieves impressive performance. Nonetheless, under constraints of model capacity and low-light conditions, these methods tend to apply overly aggressive denoising, causing notable smearing effects and the erosion of fine textures. Researchers attempt to moderate the denoising strength to better preserve details, but this often results in noticeable artifacts. These observations suggest a key insight: Not All Noise is Harmful—retaining an appropriate amount of noise can be preferable to over-smoothing and is often aligned with user aesthetics. Moreover, most existing methods are mainly trained on paired clean–degraded data, yielding fixed, hard-to-control denoising behavior that fails to meet evolving user demands. In this paper, we propose JDD-PRO, a Perception-awaRe and contrOllable framework for RAW image JDD. Built on a controllable JDD training scheme and a HVS-inspired perceptual adapter, our method allows users to tune the denoising strength at inference to strike their desired balance between noise removal and detail preservation. Comprehensive experiments on real-world datasets validate the effectiveness and robustness of our proposed method.
PaperID: 7873, Poster
Abstract: Despite rapid progress in Multimodal Large Language Models (MLLMs), visual spatial reasoning remains unreliable when correct answers depend on how a scene would appear under unseen or alternative viewpoints. Recent work addresses this by augmenting reasoning with world models for visual imagination, but questions such as when imagination is actually necessary, how much of it is beneficial, and when it becomes harmful, remain poorly understood. In practice, indiscriminate imagination can increase computation and even degrade performance by introducing misleading evidence. In this work, we present an in-depth analysis of test-time visual imagination as a controllable resource for spatial reasoning. We first study when static visual evidence is sufficient, when imagination improves reasoning, and how excessive or unnecessary imagination affects accuracy and efficiency. To support this analysis, we then introduce AVIC, an adaptive test-time framework with world models that explicitly reasons about the sufficiency of current visual evidence before selectively invoking and scaling visual imagination. Finally, to further learn this gating and planning behavior without any annotation of when and how much to imagine, we introduce AVIC-R, which trains the policy end-to-end via GRPO from QA-correctness rewards and penalties by imagination cost. Across spatial reasoning benchmarks (SAT, MMSI) and an embodied navigation benchmark (R2R), our results reveal clear scenarios where imagination is critical, marginal, or detrimental, and show that selective control can match or outperform fixed imagination strategies with substantially fewer world-model calls and language tokens. Our AVIC-R surpasses strong proprietary baselines including GPT-4o and GPT-4.1 while invoking the world model less often. Overall, our findings highlight the importance of analyzing and controlling test-time imagination for efficient and reliable spatial reasoning.
Authors:
Patrick T Kon, Archana Pradeep, Ang Chen, Alexander P Ellis, Warren Hunt, Zijian Wang, John Yang, Samuel ThompsonAbstract: Small language models (SLMs) offer compelling advantages in cost, latency, and adaptability, but have so far lagged behind larger models on long-horizon software engineering tasks such as SWE-bench, where they suffer from pervasive action looping and low resolution rates. We introduce SWE-Protégé, a post-training framework that reframes software repair as an expert–protégé collaboration problem. In SWE-Protégé, an SLM remains the sole decision-maker while learning to selectively seek guidance from a strong expert model, recognize stalled states, and follow through on expert feedback. Our approach combines supervised fine-tuning on expert-augmented trajectories with agentic reinforcement learning that explicitly discourages degenerative looping and unproductive expert collaboration. We lightly post-train Qwen2.5-Coder-7B-Instruct to achieve 42.4% Pass@1 on SWE-bench Verified, a +25.4% improvement over the prior SLM state of the art, while using expert assistance sparsely (≈ 4 calls per task and 11% of total tokens).
PaperID: 7875, Poster
Abstract: Large-scale foundation models (FMs) in remote sensing (RS) (denoted as RS FMs) are developed following paradigms established in computer vision (CV), yet the validity of transferring CV scaling laws to RS has not been systematically examined. We hypothesize that RS FMs enter an overparameterized regime at substantially smaller scales than their CV counterparts, with task-relevant information encoded redundantly across model dimensions. To test this hypothesis, we apply post-hoc slimmability, uniform width reduction of pretrained encoder transformer blocks, as a tool to measure representational redundancy across eight state-of-the-art RS FMs on classification, segmentation, and change detection tasks. RS FMs retain 69% to 109% relative accuracy on RS datasets under aggressive width reduction, while masked autoencoder (MAE) and DINOv2 pretrained on natural images (denoted as CV MAE and CV DINOv2) degrade sharply on ImageNet subsets of matched class count over the same range of computational requirements. A CV MAE evaluated directly on the same RS datasets narrows but does not close the gap, indicating that both dataset characteristics and domain-specific pretraining contribute to the differences between the models. Mechanistic analyses such as feature correlation, explained variance, and effective dimensionality indicate that task-relevant variance concentrates in few principal components and is redundantly encoded across model dimensions. We further show that learned slimmable training improves over post-hoc slimmability for contrastive objectives, while reconstruction-based objectives do not benefit from current slimmable training protocols. Our findings establish post-hoc slimming as a practical deployment strategy for resource-constrained RS applications and as a diagnostic tool for representational redundancy in RS FMs. Upon acceptance, we will publish all code.
PaperID: 7876, Poster
Abstract: It is widely believed that stochastic gradient descent (SGD) performs significantly worse than adaptive optimizers such as Adam in pre-training Large Language Models (LLMs). Yet the underlying reason for this gap remains unclear. In this work, we attribute a large part of the discrepancy to SGD's inability to sustain learning rates comparable to Adam's much larger effective learning rates. Through empirical and theoretical analysis of LLM pre-training dynamics, we identify that training is characterized by small gradient norms and large weight-to-gradient ratios, an effect that becomes more pronounced with larger batch sizes typical in pre-training, necessitating such large effective learning rates. However, we find that output-layer gradient magnitudes become highly uneven across token classes, and that large gradient spikes frequently occur during training. Together, these effects severely restrict the admissible learning rate of SGD. Guided by this understanding, we show that simple clipping mechanisms that stabilize SGD at large learning rates enable it to recover most of Adam's performance. In our large-scale experiments, the validation loss gap between large-learning-rate SGD and Adam shrinks from more than 50% to only about 3.5% when pre-training a 1B-parameter LLaMA model with a 1M-token batch size.
Abstract: We revisit the regret loss framework introduced in [Park, 2024], which uses decision-theoretic regret as a direct loss function for training models to make better decisions, through the lens of probability-simplex policies. Our first result shows that a single-layer self-attention model trained with regret loss admits a stationary point whose forward-pass exactly matches \emphsmoothed fictitious play with the appropriate stepsize that ensures no-regret behavior—i.e., for any given policy input, the model outputs the same update that smoothed fictitious play would produce. In parallel, we also newly introduce a swap-regret loss function, which extends the regret-loss framework beyond external regret and enables models to directly optimize for swap-deviation robustness. We further show that this swap-regret loss admits a stationary point whose forward pass implements the corresponding swap-regret update induced by classical Blum–Mansour no-swap-regret algorithm, with each head implementing an external-regret update via smoothed fictitious play. Together, these results show that regret-trained attention can realize differentiable mechanisms whose deployment induces equilibrium behavior in games: external-regret dynamics lead to coarse correlated equilibrium, while swap-regret dynamics lead to correlated equilibrium. Thus, regret-based objectives steer minimal attention architectures toward online-learning dynamics with game-theoretic guarantees, without supervised traces of those algorithms.
Authors: Daniel Boharon, Yehuda Dar
Abstract: We study transfer learning for a linear regression task using several least-squares pretrained models that can be overparameterized. We formulate the target learning task as optimization that minimizes squared errors on the target dataset with penalty on the distance of the learned model from the pretrained models. We analytically formulate the test error of the learned target model and provide the corresponding empirical evaluations. Our results elucidate when using more pretrained models can improve transfer learning. Specifically, if the pretrained models are overparameterized, using sufficiently many of them is important for beneficial transfer learning. However, the learning may be compromised by overparameterization bias of pretrained models, i.e., the minimum L2-norm solution's restriction to a small subspace spanned by the training examples in the high-dimensional parameter space. We propose a simple debiasing via multiplicative correction factor that can reduce the overparameterization bias and leverage more pretrained models to learn a target predictor.
PaperID: 7879, Poster
Authors: Shuchang Ye, Jinqiang Yu, Zhujun Xiao, Yajing Kong, Yist Y Lin, Yang Ma, Jiaxi Liu, Xiaolei XU, Zheng Yu
Abstract: Industry-scale video and live-streaming moderation imposes requirements that are difficult to satisfy with generic pretrained public models or external APIs, including adaptation to platform-specific data distributions, policy-specific objectives, and product-level safety constraints. As a result, platforms must undertake internal model development, naturally turning to shared public research for guidance. However, existing multimodal foundation-model studies primarily report architectures, training recipes, data scaling strategies, and benchmark results, but provide less systematic guidance on how failures should be localized and translated into targeted model-development interventions. Interventions are essential because deployment failures are rarely self-explanatory. Similar failures can originate from different causes. Without targeted interventions, improvement reduces to heuristic trial-and-error, where benchmark improvements are weakly attributable, and failures are difficult to trace to their underlying causes. To address this gap, we present a diagnostic methodology for industry-scale Audio-Visual-Language Models~(AVLM) development. The methodology maps model failures into a taxonomy of observable failure signatures and links each class of failure to an intervention space. We instantiate this methodology across the development and alignment lifecycle of an AVLM foundation model for a large-scale video and live-streaming platform. The resulting system supports over 100 regions and is designed for noisy, ambiguous, and highly diverse content drawn from global platform traffic.
Authors: Olivia Jullian Parra, Sara Zoccheddu, David C Cerezo, Tom Forzy, Franziska S Ulrich, William Sutcliffe, Jakob M Burgstaller, Oliver Senn, Patrick Owen, Nicola Serra
Abstract: Target trial emulation (TTE) provides a framework for answering causal questions using observational data when randomized controlled trials (RCTs) are infeasible. However, standard methods for treatment effect estimation have been developed in isolation, failing to jointly address the compounding challenges inherent to analyses of observational data such as electronic health records (EHRs). In particular, these challenges include time-varying confounding and missing-not-at-random (MNAR) missingness reaching 50%-80% for critical biomarkers. To address this gap, we propose a two-stage pipeline that jointly handles MNAR missingness and causal structure across time. CausalFlow-T, a Directed Acyclic Graph (DAG)-constrained normalizing flow with Long Short-Term Memory (LSTM)-encoded patient history, performs exact invertible counterfactual inference, eliminating the approximation errors and confounding biases where existing variational and adversarial methods fail silently. Ablations on four synthetic and one semi-synthetic dataset with known counterfactuals confirm that its two core design choices address strictly non-overlapping failure modes (DAG constraints for confounding separation, exact inference for structural propagation) with neither compensating for the absence of the other. To handle the incomplete data CausalFlow-T receives as input, we propose an LLM-driven evolutionary imputer and evaluate it with three LLM backends, including two open-source models. Across 30%-80% MNAR missingness, the imputer achieves the best pooled rank across biomarker and causal metrics, leading on point-wise accuracy and temporal extrapolation while maintaining average treatment effect (ATE) recovery where statistical methods progressively degrade. Applied to a cohort of adults with type 2 diabetes in Swiss primary care initiating a GLP-1 receptor agonist or SGLT-2 inhibitor, the pipeline recovers a per-protocol weight-loss difference of -0.98 kg [95% CI -1.01, -0.96] favoring GLP-1 receptor agonists, consistent with RCT evidence and estimated directly from realistically incomplete real-world EHR data.
PaperID: 7881, Poster
Abstract: The electrocardiogram (ECG) is a time series that records the electrical activity driving cardiac motion through electromechanical coupling (EMC), where the depolarization captured by ECG triggers each heartbeat. The temporal delay and low spatial resolution of such multichannel dynamic data often cause temporal and spatial misalignments along the physical motion, limiting the application of ECG data to event/segment-based downstream clinic practice. Although many efforts have been made to recover spatiotemporal trajectories conditioned on ECG waveforms, the intrinsic principle of EMC that results in various temporal delays remains underexplored in current generative models. To address this limitation, we propose Time2Track, a mechanics foundation model that generates individualized 4D heart motion from same-day ECG recordings by encoding electrophysiology as delay-aware token embeddings. Specifically, a learnable wavelet decomposition extracts the electromechanical delay, which is then formulated as a physics-based conditioning signal for motion diffusion. Time2Track generates 4D motion with individual-specific kinematics, avoiding the collapse to a population average, i.e., a common failure scenario when diffusing motions whose inter-subject differences are subtle. On n=40,577 subjects from the UK Biobank, our Time2Track accurately synthesizes 4D cardiac trajectories from ECG and preserves inter-subject diversity, a regime in which generic motion diffusion baselines collapse to a population average.
PaperID: 7882, Poster
Authors: Siddharth Chaini, federica bianco, Ashish Mahabal
Abstract: Astrophysical observations from Earth are subject to weather, environmental, and scientific constraints that lead to characteristically sparse and unevenly sampled multi-wavelength time series (light curves), severely limiting the performance on downstream tasks. Interpolating light curves help in mitigating this problem, where Gaussian Processes (GPs) are the current standard. However, GPs struggle with cross-wavelength correlations, require a priori kernel specification, and cannot handle ensembles of diverse transient and variable classes. Furthermore, GPs and other existing models (e.g., Bayesian neural networks), scale poorly and must be fit to each light curve individually, transferring no information between light curves. We introduce Attentive Neural Processes (ANPs) for light curve reconstruction, combining the probabilistic framework of GPs with the scalability of deep learning, and we demonstrate their use as a fast, data-driven, universal light curve interpolator. Unlike prior processes tailored to single object classes, our approach is morphologically agnostic. By meta-learning on a diverse population of simulated transients, ANPs shift the computation cost to the training phase, enabling amortized inference - particularly important at the eve of the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST), which will produce massive datasets for transient astrophysics. We evaluate our models on realistic LSST cadences for its 6-filters across 14 transient classes. The ANP outperforms all our benchmark models, including a suite of GP and neural-network based models, on all regression quality, feature recovery, and rigorous probabilistic calibration metrics. Our ANP performs probabilistic interpolation on all filters in ~ 10^-6 seconds -- over four orders of magnitude faster than the next-best neural benchmark and over five orders of magnitude faster than GPs -- making them suitable for on-line use on the LSST alert stream estimated to carry ~10^7 events/night. The ANP avoids neural networks' overconfidence and GPs' underconfidence, delivering sharp, highly calibrated uncertainties. This work establishes the neural process family as a scalable, probabilistic foundation for real-time transient science in the LSST era.
PaperID: 7883, Poster
Abstract: With the rapid advancement of generative models, AI-generated images (AIGIs) have become increasingly realistic, posing new challenges for reliable detection. In this work, we revisit the root causes of poor generalization in AIGI detectors from the perspective of shortcut learning. We find that detectors often rely on spurious semantic cues or a single prominent artifact, causing the feature space to collapse toward low diversity and severely limiting generalization. To characterize this issue, we propose the Generative Artifact Manifold Clustering (GAMC) hypothesis, which shows that although generative artifacts are multi-dimensional, different generators occupy distinct distributions along these dimensions, making detectors that rely on local artifact cues inherently fragile. To address this, we shift AIGI detection from “single-cue reliance” to “multi-dimensional discrimination” and introduce the Semantic–Artifact Self-Adversarial Feature Learning (SA-AFL) framework. SA-AFL progressively disentangles semantics from artifacts and employs a self-adversarial mechanism to balance learning across artifact dimensions, enabling the model to retain richer and more shared generative features. As a result, SA-AFL reduces shortcut reliance and learns more transferable evidence for cross-generator detection. Comprehensive experiments on images from 23 generative models show that SA-AFL improves ACC by 6.84% and mAP by 8.25% over state-of-the-art methods.
Abstract: Human cognition emerges from coordinated spiking dynamics in distributed neural circuits, where information is encoded via both firing rates and precise spike timing determined by brain rhythms. Inspired by this notion, we propose a brain-inspired learning primitive in which cognition-level neural synchrony emerges through iterative interactions between micro-scale dynamics of spiking neurons and a macro-scale mechanism of oscillatory synchronization. Specifically, we model each parcel (e.g., a cortical region or an image pixel) in the target system as a spiking neuron embedded in a predefined connectivity scaffold. Low-level information is encoded in a spatiotemporal domain, where neurons are selectively grouped and fire spontaneously over time through self-organized dynamics. In the bottom-up route, oscillatory synchronization is formed from past spiking activity accumulated over a finite memory window. Since brain dynamics operate in a regime of partial and transient synchronization rather than global phase locking, we model oscillatory coordination using a time-delayed synchronization formulation, which enables a top-down modulation of heterogeneous neural spiking for a large-scale distributed system. Together, we devise a spiking-by-synchronization neural network (S²-Net) that uses rhythmic timing as a control mechanism for efficient information processing. Promising results have been achieved across a broad range of tasks, including neural activity decoding, energy-efficient signal processing, temporal binding and semantic reasoning.
PaperID: 7885, Poster
Abstract: Harmful-content detectors are increasingly deployed in safety-critical settings, but they remain vulnerable to subtle rhetorical reframings that preserve the underlying claim while altering its presentation. Existing adversarial robustness methods are typically token-level or prompt-driven, hindering mechanism-level attribution and targeted repair. We propose Adaptive Rhetorical Adversarial Optimization (ARAO), which operationalizes rhetorical theory as an axis-aligned strategy space with directional and interpretable controls for adversarial rewriting. A planner-rewriter pipeline converts strategy vectors into non-contradictory rewrites, enabling mechanism-level failure attribution and axis-targeted repair. Experiments on misinformation and extremist-rhetoric benchmarks show that ARAO improves robustness by 2.35 ROC-AUC points on average over the strongest baselines on attacked sets, while achieving the best performance on original content.
PaperID: 7886, Poster
Authors: Ali Mehrabian, Mahdi Mazloum
Abstract: A two-branch routed regressor realizes exactly the function class of ordinary quadratic regression. Because routing does not enlarge the observable function class here, the routed model differs from quadratic regression only through the prior induced by routed coordinates. We compute this induced prior in closed form and show that fiber integration along the routing map creates a logarithmically singular density on the linear submodel. The induced posterior over quadratic coefficients then coincides exactly with the quadratic-regression posterior under the pushed-forward prior, isolating the role of prior geometry in this proxy model. For compactly supported priors on linear truth (q_2=0), the support-relative global real log canonical threshold (RLCT) is (3/2,2) when the support reaches the collapsed stratum and (3/2,1) when it avoids it; under the full-support Gaussian-coordinate prior, the global RLCT is (3/2,2). Against smooth coefficient-space baselines with the same known noise variance on linear truth, this singular-versus-smooth prior comparison yields a rigorous \log\log n Bayes-factor advantage. The same multiplicity-2 product singularity transfers locally to sigmoid gates on linear truth and extends locally to d-dimensional inputs.
PaperID: 7887, Poster
Abstract: Class-incremental learning (class-IL) faces a structural obstacle: discriminative sequential training optimises only the diagonal blocks of a pairwise loss matrix, leaving the inter-task blocks---task confusion (TC)---unminimised. A recent Infeasibility Theorem makes this precise, proving that no discriminative class-IL learner can attain the joint-training optimum even when catastrophic forgetting (CF) is perfectly addressed. Yet the state of the art consists almost entirely of discriminative methods built on pretrained foundation models that approach joint-training performance. We explain this regime by showing that pretraining \emphexponentially attenuates the obstacle. For any pretrained backbone \phi, the TC gap of any CF-optimal prototype-based discriminative learner---a class that captures the test-time classifier of SLDA, RanPAC, FeCAM, and the prototype branches of EASE and InfLoRA---is bounded by \exp(-\gamma_\min(\phi)/8), where \gamma_\min(\phi) is the worst-case Mahalanobis margin between task centroids in feature space: a single, computable scalar that determines how much TC remains. Combined with a finite-sample convergence analysis under heavy-tailed gradient noise and a representation drift bound governing backbone fine-tuning, this yields a diagnostic decomposition of the class-IL excess risk into E_\mathrmCF + E_\mathrmTC, with disjoint direct dependencies on optimisation budget and backbone quality. Across 12 (backbone, dataset) configurations spanning ViT-IN21K, DINOv2, and CLIP plus four random-backbone anchor points, the pooled empirical slope of \log(\mathrmTC\ gap) versus \widehat\gamma_\min on softmax cross-entropy heads is -0.119 (bootstrap 95% CI [-0.134, -0.105]), consistent with the predicted -1/8 and excluding the conservative fall-back -1/16. The slope fit extends the rate beyond the formal LDA/QDA scope. E_\mathrmCF scales as K^-(\alpha-1)/\alpha as predicted, with the rate steepening by more than 3× under gradient clipping; representation drift scales linearly in the backbone learning rate (r = +0.998). A theory-derived training procedure---whose schedule is determined by \widehat\gamma_\min, the tail index \widehat\alpha, and warm-up estimates of smoothness and curvature---is competitive with the strongest pretrained-CIL baselines across the four standard benchmarks.
Authors:
Jing Liang, Hongyao Tang, Yi Ma, Yancheng He, Weixun Wang, Xiaoyang Li, Wenbo Su, Jinyi Liu, YAN ZHENG, Jianye Hao, Bo ZhengAbstract: Reinforcement learning (RL) has gained growing attention in large language model (LLM) post-training, yet RL training remains fragile and can suffer from instability or collapse. One vital cause is training-inference mismatch: % different from traditional RL, the engineering design of LLM adopts separate inference and training engines for generation efficiency and training precision, which in practice exhibits inconsistent probabilities for the same trajectories on training and inference sides, even with synchronized model parameters. This naturally induces a special type of off-policyness ever existing and poisoning the training. Prior works have made various efforts in addressing the off-policyness to stabilize the training policies under the mismatch. In this paper, we point out the object misalignment neglected by existing works that an effective update to the policy in the training engine not necessarily ensures the improvement of the inference policy, i.e., the one used in deployment. To this end, we propose a new policy optimization objective for LLM RL, named Monotonic Inference Policy Improvement (MIPI). Following this principle, we introduce Monotonic Inference Policy Update (MIPU), a two-step LLM RL framework that constructs sampler-referenced candidate updates and selectively accepts synchronized candidates using an inference-side gap proxy. Experiments under high mismatch show that MIPU improves average reasoning performance and training stability.
Authors: Franziska Heeg, Jonas Sauer, Petra Mutzel, Ingo Scholtes
Abstract: An important characteristic of temporal graphs is how the directed arrow of time influences their causal topology, i.e., which nodes can possibly influence each other causally via time-respecting paths. The resulting patterns are often neglected by temporal graph neural networks (TGNNs). To formally analyze the expressive power of TGNNs, we lack a generalization of graph isomorphism to temporal graphs that captures their causal topology. Addressing this gap, we introduce consistent event graph isomorphism, which utilizes a time-unfolded representation of time-respecting paths in temporal graphs. We compare this definition with existing notions of temporal graph isomorphisms. We highlight the advantages of our approach and develop a temporal generalization of the Weisfeiler-Leman algorithm to heuristically distinguish non-isomorphic temporal graphs. Building on this theoretical foundation, we derive a novel message passing scheme for temporal graph neural networks that operates on the event graph representation of temporal graphs and experimentally evaluate it in a temporal graph classification setting.
PaperID: 7890, Poster
Authors: Yibo Zhou, Yirong Xiang
Abstract: Source labels tell us where a sample came from, but not how that source oriented its coordinates. In a sparse Gaussian multi-source model with source-specific sign gauges, that distinction changes the learning problem. Without calibration, pooling across sources erases the mean signal, and the source-labelled experiment reduces to sparse PCA. With task-independent relative-gauge metadata, synchronization aligns sources and the same observations become a sparse mean problem. We formalize the value of metadata through aligned mass, which quantifies how much first-order signal a metadata-induced alignment restores. Source labels create only limited aligned mass, whereas valid calibration can create aligned mass proportional to the full sample size. This yields a statistical separation in micro-source regimes and, under standard sparse-PCA hardness assumptions, a computational separation as well. Experiments confirm the predicted phase transition, the failure of invalid metadata, reuse of one calibration graph across many tasks, and controlled real-feature settings with imposed source gauges.
PaperID: 7891, Poster
Abstract: Classification-trained sequential glimpse policies have attained competitive performance and efficiency on large-scale recognition benchmarks, and are motivated by analogy to the saccadic architecture of biological active vision. Yet whether these policies reproduce the behavioral priors that organize human free viewing, or merely approximate their spatial distribution as a byproduct of classification reward, has not been systematically examined. We present a diagnostic behavioral audit of AdaptiveNN across three free-viewing datasets, evaluating five dimensions associated with known properties of biological gaze: spatial priors, face prioritization, local return, saccade amplitude structure, and semantic guidance. We show that AdaptiveNN exhibits: (i) spatial behavior that falls below a simple center-bias baseline; (ii) scanpath geometry unlike human scanpaths in amplitude and sequential structure; (iii) revisitation dynamics inconsistent with human free viewing and not reducible to step-size geometry; and (iv) selectively attenuated semantic guidance, with face-specific prioritization center-explained, social sampling attenuated, and object-level rates preserved. In AdaptiveNN, this dissociation suggests that classification reward alone cannot reproduce several behavioral priors organizing human free viewing, even when object-level sampling is preserved. These findings identify which behavioral priors require constraints beyond classification utility, offering targets for designing cognitively aligned glimpse policies.
Abstract: Multi-agent reinforcement learning (MARL) holds great potential but faces robustness challenges due to environmental uncertainty. To address this, distributionally robust Markov games (RMGs) optimize worst-case performance when the environment deviates from the nominal model within a uncertainty set. Beyond robustness, an equally urgent goal for MARL is data efficiency—sampling from vast state and action spaces that grow exponentially with the number of agents potentially leads to the curse of multiagency. However, current provably data-efficient algorithms for RMGs are limited to tabular settings with finite state and action spaces, which are only computationally manageable for small-scale problems, leaving RMGs with large-scale (or infinite) state spaces largely unexplored. The only existing work beyond tabular settings focuses on linear function approximation (LFA) for a restrictive class of RMGs using vanish minimal value assumption and still suffers from sample complexity with the curse of multiagency. In this work, we focuses on general RMGs with LFA. For uncertainty sets defined by total variation distance, we develop provably data-efficient algorithms that break the curse of multiagency in both the generative model setting and a newly proposed online interactive setting. To our knowledge, our results are the first to break the curse of multiagency of sample complexity for RMGs with large (possibly infinite) state spaces, regardless of the uncertainty set construction.
PaperID: 7893, Poster
Abstract: Multi-omics drug-drug interaction (DDI) prediction forecasts DDI outcomes from heterogeneous multi-omics data and complex structured rules, where multi-omics interpretability of hierarchical biological factor interactions is extremely crucial, which remains unexplored in literature. In this paper, we are the first to study the problem of Multi-Omics Interpretable DDI Prediction, which is highly non-trivial with three challenges: i) how to identify interaction-relevant biological entities from large-scale knowledge sources, which lay the foundation for interpretability; ii) how to capture multi-omics cross-modal connections that elucidate underlying mechanisms under diverse interaction contexts; and iii) how to explicitly enable mechanism-aware reasoning over structured biological rules. To tackle the challenges, we propose a Multi-omics Interpretable prediction framework with joint optimization of graph structure, Neural architecture, and symbolic rules to hanDle DDI prediction (MIND-DDI), which follows a novel interpretable learning paradigm that tightly couples biomedical graph sampling, hierarchical neural representation learning, and symbolic rule modeling under a unified optimization scheme. Specifically, we propose a curriculum-guided biological subgraph sampler for progressively constructing informative subgraphs, then introduce a hierarchical multi-omics neural architecture search module to capture multi-omics alignment as well as cross-modal interactions, and finally present a mechanism-aware pathway-based symbolic reasoner for multi-hop reasoning over a unified neural-symbolic space. Extensive experiments show the superiority of MIND-DDI in both predictive accuracy and interpretability.
PaperID: 7894, Poster
Authors: Beier Meng
Abstract: Large-scale video–text contrastive learning has become the de facto standard for video understanding, yet mainstream methods almost universally adopt the linear patch embedding of vision Transformers (ViT) or evenly sliced tubelets as the front end. We revisit this front end from the biological principles of the primary visual cortex (area V1) and introduce V1DynamicVisionSystem (V1-DVS), a fully end-to-end trainable, bio-plausible video embedding framework. Its main contributions are: (i) a three-pathway parallel architecture that maps onto layer 4C orientation selectivity, layer 4B motion direction selectivity, and the cytochrome-oxidase Blob color-opponent system; (ii) an explicit shape–color decoupling, with isotropic Blob processing running in parallel with anisotropic Gabor orientation processing; (iii) long-range spatiotemporal convolution that captures inter-frame optical flow through 3D filters and the Adelson–Bergen motion-energy model rather than naive frame stacking; and (iv) bidirectional cross-pathway attention, allowing motion and color features to mutually modulate, as in V1→V4 inter-areal feedback. Pretrained on CC3M, WebVid-10M, ActivityNet and Ego4D and evaluated under a fully self-supervised objective(SimCLR + temporal-order verification + masked-token prediction), V1-DVS outperforms four strong baselines—linear patch embedding, mean temporal pooling, tubelet embedding, and spatiotemporal latent encoding—on six metrics spanning retrieval and embedding-space geometry, validating the practical value of biologically grounded structure for video embedding.
Abstract: Knowledge editing in Large Language Models (LLMs) aims to update the model’s prediction for a specific query to a desired target while preserving its behavior on all other inputs. This process typically involves two stages: identifying the layer to edit and performing the parameter update. Intuitively, different queries may localize knowledge at different depths of the model, resulting in different sample-wise editing performance for a fixed editing layer. In this work, we hypothesize the existence of fixed golden layers that can achieve near-optimal editing performance similar to sample-wise optimal layers. To validate this hypothesis, we provide empirical evidence by comparing golden layers against ground-truth sample-wise optimal layers. Furthermore, we show that golden layers can be reliably identified using a proxy dataset and generalize effectively to unseen test set queries across datasets. Finally, we propose a novel method, namely Layer Gradient Analysis (LGA), that estimates golden layers efficiently via gradient-attribution, avoiding extensive trial-and-error across multiple editing runs. Extensive experiments on several benchmark datasets demonstrate the effectiveness and robustness of our LGA approach across different LLM types and various knowledge editing methods.
PaperID: 7896, Poster
Abstract: Monte Carlo Tree Search (MCTS) performance degrades significantly in Reentrant Blocking Hybrid Flow Shop (RBHFS) scheduling, which arises in modern automated scientific laboratories. The interaction of blocking and reentrant constraints causes rollout simulations to consistently end in a “dense deadlock.” As a result, infeasible branches become indistinguishable, which undermines the exploitation capability of standard MCTS and effectively reduces it to uniform random exploration. However, we find that deadlocks are not equally uninformative. We observe that the average rollout deadlock depth—the number of steps executed before failure—across different action branches shows a strong positive correlation with true feasibility probability, making it a reliable proxy for the latent likelihood that a branch contains feasible solutions. Guided by this insight, we propose Deadlock-Depth Guided MCTS (D^2-MCTS). By integrating deadlock-depth statistics into the evaluation mechanism, D^2-MCTS is able to distinguish among otherwise indistinguishable infeasible branches. Empirical results demonstrate that D^2-MCTS significantly outperforms state-of-the-art baselines, demonstrating superior robustness and efficiency.
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) enhances large language model (LLM) reasoning, yet growing evidence indicates an \emphexploration ceiling: they tend to reweight existing solution traces and produce low-diversity rollouts. This restricts the discovery of novel solution perspective and weakens out-of-distribution generalization. We address this limitation by directly perturbing the rollout policy's parameters while updating the clean policy model as usual, an approach that achieves better temporal consistency and yields more complex behavior patterns. To stabilize learning from these perturbed rollouts, which are inherently off-policy, we incorporate truncated importance sampling. Furthermore, we introduce a lightweight, dynamic noise scheduler based on semantic diversity and model self-certainty to optimally adjust the noise scale. Instantiated on GRPO, our method (PSN-GRPO) demonstrates its efficacy through four key contributions: (1) it improves high-budget pass@k, increases semantic and operational diversity, increases CoT consistency, and remains complementary to exploration-oriented RLVR techniques; (2) it validates efficacy across multiple models, including Qwen2.5, Qwen3, and Llama-3.1; (3) it shows robust performance on out-of-distribution science and coding benchmarks; and (4) it uncovers novel solution perspectives for problems the original model could not solve, as shown by qualitative analysis.
Abstract: Stopgrads are widely used in training machine learning models, but stopgrads can alter the gradient, stationary points and convergence guarantees of the original objective, which can make stopgrad training theoretically ungrounded. We introduce a stopgrad regression principle, which identifies a general template for stopgrad objectives with a closed-form characterization of stationary points, and unifies stopgrad objectives for flow maps, reinforcement learning, and diffusion samplers. We provide theoretical grounding for optimizing stopgrad flow map objectives by characterizing stationary points and proving convergence results. Remarkably, we show that the learned flow map has a closed-form expression composing the initial flow map and the true flow map, under functional semi-gradient flow. We use our stopgrad regression principle to propose novel modified stopgrad placements for flow map objectives which speed up training by ~2.5×.
PaperID: 7899, Poster
Abstract: Image Forgery Localization (IFL) aims to identify and localize the tampered regions within edited images. Many studies employ a dual-branch backbone to extract tampering features from dual modalities, followed by feature fusion at the final stage. In this process, the extraction and fusion of dual-modality features is relatively independent, which fails to fully leverage the complementarity between different modalities and thus diminishes sensitivity to tampering artifacts. Inspired by the way humans continuously integrate multi-faceted knowledge to understand the world, we propose MAMQ-Net, which contains a novel Multi-stage Alternating Feature Extraction and Interaction architecture. At each stage, we deeply explore the intrinsic relationships and mappings between different modality features. Feature extraction and interaction are performed alternately, constructing complementary dual-modality tampering feature representations and enhancing sensitivity to tampering artifacts. Additionally, we introduce a lightweight, Query-driven Multi-level Feature Decoding. This mechanism progressively aggregates key information from multi-level dual-modality tampering features through multiple sets of learnable tamper-aware queries, effectively filtering out irrelevant features. Finally, multi-level queries are used to refine discriminative features, enabling precise localization of tampered regions. Extensive experiments demonstrate that our framework outperforms current state-of-the-art models in localization accuracy and robustness across multiple public datasets, achieving a favorable balance between performance and efficiency.
PaperID: 7900, Poster
Abstract: Language understanding in the brain is context-dependent, varying across experimental stimuli and individuals, which makes it difficult to build computational models that generalize across both. This calls for a foundation model of language-evoked brain activity that can capture shared structure while adapting efficiently to new participants and inputs. We introduce RABBiT (Rapidly Adaptive BOLD foundation model via BraIn-Tuning), a compact audio-to-fMRI encoder designed for accurate zero- and few-shot prediction. A comprehensive evaluation on 324 participants across multiple unseen fMRI datasets shows that RABBiT enables accurate zero-shot prediction of fMRI responses to natural speech across auditory and language-selective regions, surpassing the SOTA foundation model for fMRI and predictions based on group averages. With as little as 10 minutes of participant-specific data, RABBiT further improves performance via parameter-efficient tuning, substantially outperforming per-participant linear models. RABBiT's performance is driven by two key innovations: (1) learned region-specific attention, and (2) a decomposition of brain responses into shared and subject-specific components, combined with a brain-tuned speech backbone. In addition to supporting strong predictive accuracy, the structured, region-specific representations that RABBiT learns enable interpretability. By eliminating the need for extensive per-participant data and model fitting, RABBiT enables scalable population-level analyses of language in the human brain.